{"thread":{"id":"60273","subject":"[PATCH 00/30] Initial support for multiple hash functions","startedAt":"2023-09-27T19:50:09Z","lastAt":"2024-02-17T01:59:31Z","messageCount":104,"participants":["Eric W. Biederman","Eric Sunshine","Junio C Hamano","Eric Biederman","Linus Arver","Patrick Steinhardt","Jean-Noël Avila","Kristoffer Haugsbakk"],"isPatch":true,"patchVersion":1,"patchTotal":30},"messages":[{"id":"482380","messageId":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":null,"subject":"[PATCH 00/30] Initial support for multiple hash functions","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:49:57Z","receivedAt":"2023-09-27T19:50:09Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"\nI have been going over and over this patchset trying to figure\nout if it is ready to be merged.  I don't know of any deficiencies\nso it is at a point it could benefit from a set of eyes that\nare not mine.\n\nI had planned to wait a little bit longer but there are some on-going\nconversations that could benefit from people seeing what it means for a\nrepository to support two hash functions at the same time.\n\n\nA key part of the hash function transition plan is a way that a single\ngit repository can inter-operate with git repositories whose storage\nhash function is SHA-1 and git repositories whose storage hash function\nis SHA-256.\n\nThis interoperability can defined in terms of two repositories one whose\nstorage hash function is SHA-1 and another whose storage hash function\nis SHA-256.  Those two repositories receive exactly the same objects,\nbut they store them in different but equivalent ways.\n\nFor a repository that has one storage hash function to inter-operate\nwith a repository that has a different storage hash function requires\nthe first repository to be able produce it's objects as if they were\nstored in the second hash function.\n\nThis series of changes focuses on implementing the pieces that allow\na repository that uses one storage hash function to produce the objects\nthat would have been stored with a second storage hash function.\n\nThe final patch in this series is the addition of a test that creates\ntwo repositories one that uses SHA-1 as it's storage hash function\nand the other that uses SHA-256 as it's storage hash function.\nIdentical operations are performed on the two repositories, and their\ncompatibility objects are compared to verify they are the same.\nAKA the SHA-1 repository on the fly generates the objects store in\nthe SHA-256 repository, and the SHA-256 repository on the fly generates\nthe objects that are stored in the SHA-1 repository.\n\nThere are two fundamental technologies for enabling this.\n- The ability to convert a stored object into the object the\n  other repository would have stored.\n- The ability to remember a mapping between SHA-1 and SHA-256 oids\n  of equivalent objects.\n\nWith such technologies it is very easy to implement user facing changes.\nTo avoid locking git into poor decisions by accident I have done my best\nto minimize the user facing changes, while still building the internal\ninfrastructure that is needed for interoperability.\n\nAll of this work is inspired by earlier work on interoperability by\n\"brian m. carlson\" and some of the key pieces of code are still his.\n\nTo get to the point where I can test if a SHA-1 and a SHA-256 repository\ncan on the fly generate each other, I have made some small user-facing\nchanges.\n\ngit rev-parse now supports --output-object-format as a way to query\nthe internal mapping tables between oids and report the equivalent\noid of the other format.\n\ngit cat-file when given a oid that does not match the repositories\nstorage format will now attempt to find the oids equivalent object that\nis stored in the repository and if found dynamically generate the object\nthat would have been stored in a repository with a different storage\nhash function and display the object.\n\nAn additional file loose-object-index will be stored in \".git/objects/\".\n\nAn additional option \"extensions.compatObjectFormat\" is implemented,\nthat generates and stores mappings between the oids of objects stored in\nthe repository and oids of the equivalent objects that would be stored\nin a repository show storage format was extensions.compatObjectFormat.\n\nEric W. Biederman (23):\n      object-file-convert: Stubs for converting from one object format to another\n      oid-array: Teach oid-array to handle multiple kinds of oids\n      object-names: Support input of oids in any supported hash\n      repository: add a compatibility hash algorithm\n      loose: Compatibilty short name support\n      object-file: Update the loose object map when writing loose objects\n      object-file: Add a compat_oid_in parameter to write_object_file_flags\n      commit: Convert mergetag before computing the signature of a commit\n      commit: Export add_header_signature to support handling signatures on tags\n      tag: sign both hashes\n      object: Factor out parse_mode out of fast-import and tree-walk into in object.h\n      object-file-convert: Don't leak when converting tag objects\n      object-file-convert: Convert commits that embed signed tags\n      object-file: Update object_info_extended to reencode objects\n      rev-parse: Add an --output-object-format parameter\n      builtin/cat-file:  Let the oid determine the output algorithm\n      tree-walk: init_tree_desc take an oid to get the hash algorithm\n      object-file: Handle compat objects in check_object_signature\n      builtin/ls-tree: Let the oid determine the output algorithm\n      test-lib: Compute the compatibility hash so tests may use it\n      t1006: Rename sha1 to oid\n      t1006: Test oid compatibility with cat-file\n      t1016-compatObjectFormat: Add tests to verify the conversion between objects\n\nbrian m. carlson (7):\n      loose: add a mapping between SHA-1 and SHA-256 for loose objects\n      commit: write commits for both hashes\n      cache: add a function to read an OID of a specific algorithm\n      object-file-convert: add a function to convert trees between algorithms\n      object-file-convert: convert tag objects when writing\n      object-file-convert: convert commit objects when writing\n      repository: Implement extensions.compatObjectFormat\n\n Documentation/config/extensions.txt |  12 ++\n Documentation/git-rev-parse.txt     |  12 ++\n Makefile                            |   3 +\n archive.c                           |   3 +-\n builtin/am.c                        |   6 +-\n builtin/cat-file.c                  |  12 +-\n builtin/checkout.c                  |   8 +-\n builtin/clone.c                     |   2 +-\n builtin/commit.c                    |   2 +-\n builtin/fast-import.c               |  18 +-\n builtin/grep.c                      |   8 +-\n builtin/ls-tree.c                   |   5 +-\n builtin/merge.c                     |   3 +-\n builtin/pack-objects.c              |   6 +-\n builtin/read-tree.c                 |   2 +-\n builtin/rev-parse.c                 |  25 ++-\n builtin/stash.c                     |   5 +-\n builtin/tag.c                       |  45 ++++-\n cache-tree.c                        |   4 +-\n commit.c                            | 219 ++++++++++++++++-----\n commit.h                            |   1 +\n delta-islands.c                     |   2 +-\n diff-lib.c                          |   2 +-\n fsck.c                              |   6 +-\n hash-ll.h                           |   1 +\n hash.h                              |   9 +-\n http-push.c                         |   2 +-\n list-objects.c                      |   2 +-\n loose.c                             | 258 ++++++++++++++++++++++++\n loose.h                             |  22 +++\n match-trees.c                       |   4 +-\n merge-ort.c                         |  11 +-\n merge-recursive.c                   |   2 +-\n merge.c                             |   3 +-\n object-file-convert.c               | 277 ++++++++++++++++++++++++++\n object-file-convert.h               |  24 +++\n object-file.c                       | 212 ++++++++++++++++++--\n object-name.c                       |  49 +++--\n object-name.h                       |   3 +-\n object-store-ll.h                   |   7 +-\n object.c                            |   2 +\n object.h                            |  18 ++\n oid-array.c                         |  12 +-\n pack-bitmap-write.c                 |   2 +-\n packfile.c                          |   3 +-\n reflog.c                            |   2 +-\n repository.c                        |  14 ++\n repository.h                        |   4 +\n revision.c                          |   4 +-\n setup.c                             |  22 +++\n setup.h                             |   1 +\n t/helper/test-delete-gpgsig.c       |  62 ++++++\n t/helper/test-tool.c                |   1 +\n t/helper/test-tool.h                |   1 +\n t/t1006-cat-file.sh                 | 379 +++++++++++++++++++++---------------\n t/t1016-compatObjectFormat.sh       | 280 ++++++++++++++++++++++++++\n t/t1016/gpg                         |   2 +\n t/test-lib-functions.sh             |  17 +-\n tree-walk.c                         |  58 +++---\n tree-walk.h                         |   7 +-\n tree.c                              |   2 +-\n walker.c                            |   2 +-\n 62 files changed, 1843 insertions(+), 349 deletions(-)\n\nEric\n"},{"id":"482381","messageId":"20230927195537.1682-1-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 01/30] object-file-convert: Stubs for converting from one object format to another","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:08Z","receivedAt":"2023-09-27T19:55:57Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nTwo basic functions are provided:\n- convert_object_file Takes an object file it's type and hash algorithm\n  and converts it into the equivalent object file that would\n  have been generated with hash algorithm \"to\".\n\n  For blob objects there is no converstion to be done and it is an\n  error to use this function on them.\n\n  For commit, tree, and tag objects embedded oids are replaced by the\n  oids of the objects they refer to with those objects and their\n  object ids reencoded in with the hash algorithm \"to\".  Signatures\n  are rearranged so that they remain valid after the object has\n  been reencoded.\n\n- repo_oid_to_algop which takes an oid that refers to an object file\n  and returns the oid of the equavalent object file generated\n  with the target hash algorithm.\n\nThe pair of files object-file-convert.c and object-file-convert.h are\nintroduced to hold as much of this logic as possible to keep this\nconversion logic cleanly separated from everything else and in the\nhopes that someday the code will be clean enough git can support\ncompiling out support for sha1 and the various conversion functions.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Makefile              |  1 +\n object-file-convert.c | 57 +++++++++++++++++++++++++++++++++++++++++++\n object-file-convert.h | 24 ++++++++++++++++++\n 3 files changed, 82 insertions(+)\n create mode 100644 object-file-convert.c\n create mode 100644 object-file-convert.h\n\ndiff --git a/Makefile b/Makefile\nindex 577630936535..f7e824f25cda 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1073,6 +1073,7 @@ LIB_OBJS += notes-cache.o\n LIB_OBJS += notes-merge.o\n LIB_OBJS += notes-utils.o\n LIB_OBJS += notes.o\n+LIB_OBJS += object-file-convert.o\n LIB_OBJS += object-file.o\n LIB_OBJS += object-name.o\n LIB_OBJS += object.o\ndiff --git a/object-file-convert.c b/object-file-convert.c\nnew file mode 100644\nindex 000000000000..ba3e18f6af44\n--- /dev/null\n+++ b/object-file-convert.c\n@@ -0,0 +1,57 @@\n+#include \"git-compat-util.h\"\n+#include \"gettext.h\"\n+#include \"strbuf.h\"\n+#include \"repository.h\"\n+#include \"hash-ll.h\"\n+#include \"object.h\"\n+#include \"object-file-convert.h\"\n+\n+int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n+\t\t      const struct git_hash_algo *to, struct object_id *dest)\n+{\n+\t/*\n+\t * If the source alogirthm is not set, then we're using the\n+\t * default hash algorithm for that object.\n+\t */\n+\tconst struct git_hash_algo *from =\n+\t\tsrc->algo ? &hash_algos[src->algo] : repo->hash_algo;\n+\n+\tif (from == to) {\n+\t\tif (src != dest)\n+\t\t\toidcpy(dest, src);\n+\t\treturn 0;\n+\t}\n+\treturn -1;\n+}\n+\n+int convert_object_file(struct strbuf *outbuf,\n+\t\t\tconst struct git_hash_algo *from,\n+\t\t\tconst struct git_hash_algo *to,\n+\t\t\tconst void *buf, size_t len,\n+\t\t\tenum object_type type,\n+\t\t\tint gentle)\n+{\n+\tint ret;\n+\n+\t/* Don't call this function when no conversion is necessary */\n+\tif ((from == to) || (type == OBJ_BLOB))\n+\t\tdie(\"Refusing noop object file conversion\");\n+\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\tcase OBJ_TAG:\n+\tdefault:\n+\t\t/* Not implemented yet, so fail. */\n+\t\tret = -1;\n+\t\tbreak;\n+\t}\n+\tif (!ret)\n+\t\treturn 0;\n+\tif (gentle) {\n+\t\tstrbuf_release(outbuf);\n+\t\treturn ret;\n+\t}\n+\tdie(_(\"Failed to convert object from %s to %s\"),\n+\t\tfrom->name, to->name);\n+}\ndiff --git a/object-file-convert.h b/object-file-convert.h\nnew file mode 100644\nindex 000000000000..a4f802aa8eea\n--- /dev/null\n+++ b/object-file-convert.h\n@@ -0,0 +1,24 @@\n+#ifndef OBJECT_CONVERT_H\n+#define OBJECT_CONVERT_H\n+\n+struct repository;\n+struct object_id;\n+struct git_hash_algo;\n+struct strbuf;\n+#include \"object.h\"\n+\n+int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n+\t\t      const struct git_hash_algo *to, struct object_id *dest);\n+\n+/*\n+ * Convert an object file from one hash algorithm to another algorithm.\n+ * Return -1 on failure, 0 on success.\n+ */\n+int convert_object_file(struct strbuf *outbuf,\n+\t\t\tconst struct git_hash_algo *from,\n+\t\t\tconst struct git_hash_algo *to,\n+\t\t\tconst void *buf, size_t len,\n+\t\t\tenum object_type type,\n+\t\t\tint gentle);\n+\n+#endif /* OBJECT_CONVERT_H */\n-- \n2.41.0\n\n"},{"id":"482382","messageId":"20230927195537.1682-2-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 02/30] oid-array: Teach oid-array to handle multiple kinds of oids","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:09Z","receivedAt":"2023-09-27T19:55:58Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nWhile looking at how to handle input of both SHA-1 and SHA-256 oids in\nget_oid_with_context, I realized that the oid_array in\nrepo_for_each_abbrev might have more than one kind of oid stored in it\nsimulataneously.\n\nUpdate to oid_array_append to ensure that oids added to an oid array\nalways have an algorithm set.\n\nUpdate void_hashcmp to first verify two oids use the same hash algorithm\nbefore comparing them to each other.\n\nWith that oid-array should be safe to use with differnt kinds of\noids simultaneously.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n oid-array.c | 12 ++++++++++--\n 1 file changed, 10 insertions(+), 2 deletions(-)\n\ndiff --git a/oid-array.c b/oid-array.c\nindex 8e4717746c31..1f36651754ed 100644\n--- a/oid-array.c\n+++ b/oid-array.c\n@@ -6,12 +6,20 @@ void oid_array_append(struct oid_array *array, const struct object_id *oid)\n {\n \tALLOC_GROW(array->oid, array->nr + 1, array->alloc);\n \toidcpy(&array->oid[array->nr++], oid);\n+\tif (!oid->algo)\n+\t\toid_set_algo(&array->oid[array->nr - 1], the_hash_algo);\n \tarray->sorted = 0;\n }\n \n-static int void_hashcmp(const void *a, const void *b)\n+static int void_hashcmp(const void *va, const void *vb)\n {\n-\treturn oidcmp(a, b);\n+\tconst struct object_id *a = va, *b = vb;\n+\tint ret;\n+\tif (a->algo == b->algo)\n+\t\tret = oidcmp(a, b);\n+\telse\n+\t\tret = a->algo > b->algo ? 1 : -1;\n+\treturn ret;\n }\n \n void oid_array_sort(struct oid_array *array)\n-- \n2.41.0\n\n"},{"id":"482383","messageId":"20230927195537.1682-4-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 04/30] repository: add a compatibility hash algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:11Z","receivedAt":"2023-09-27T19:55:59Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nWe currently have support for using a full stage 4 SHA-256\nimplementation.  However, we'd like to support interoperability with\nSHA-1 repositories as well.  The transition plan anticipates a\ncompatibility hash algorithm configuration option that we can use to\nimplement support for this.  Let's add an element to the repository\nstructure that indicates the compatibility hash algorithm so we can use\nit when we need to consider interoperability between algorithms.\n\nAdd a helper function repo_set_compat_hash_algo that takes a\ncompatibility hash algorithm and sets \"repo->compat_hash_algo\".  If\nGIT_HASH_UNKNOWN is passed as the compatibilty hash algorithm\n\"repo->compat_hash_algo\" is set to NULL.\n\nFor now, the code always results in \"repo->compat_hash_algo\" always\nbeing set to NULL, but that will change once a configuration option\nis added.\n\nInspired-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n repository.c | 8 ++++++++\n repository.h | 4 ++++\n setup.c      | 3 +++\n 3 files changed, 15 insertions(+)\n\ndiff --git a/repository.c b/repository.c\nindex a7679ceeaa45..80252b79e93e 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -104,6 +104,13 @@ void repo_set_hash_algo(struct repository *repo, int hash_algo)\n \trepo->hash_algo = &hash_algos[hash_algo];\n }\n \n+void repo_set_compat_hash_algo(struct repository *repo, int algo)\n+{\n+\tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n+\t\tBUG(\"hash_algo and compat_hash_algo match\");\n+\trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n+}\n+\n /*\n  * Attempt to resolve and set the provided 'gitdir' for repository 'repo'.\n  * Return 0 upon success and a non-zero value upon failure.\n@@ -184,6 +191,7 @@ int repo_init(struct repository *repo,\n \t\tgoto error;\n \n \trepo_set_hash_algo(repo, format.hash_algo);\n+\trepo_set_compat_hash_algo(repo, GIT_HASH_UNKNOWN);\n \trepo->repository_format_worktree_config = format.worktree_config;\n \n \t/* take ownership of format.partial_clone */\ndiff --git a/repository.h b/repository.h\nindex 5f18486f6465..bf3fc601cc53 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -160,6 +160,9 @@ struct repository {\n \t/* Repository's current hash algorithm, as serialized on disk. */\n \tconst struct git_hash_algo *hash_algo;\n \n+\t/* Repository's compatibility hash algorithm. */\n+\tconst struct git_hash_algo *compat_hash_algo;\n+\n \t/* A unique-id for tracing purposes. */\n \tint trace2_repo_id;\n \n@@ -199,6 +202,7 @@ void repo_set_gitdir(struct repository *repo, const char *root,\n \t\t     const struct set_gitdir_args *extra_args);\n void repo_set_worktree(struct repository *repo, const char *path);\n void repo_set_hash_algo(struct repository *repo, int algo);\n+void repo_set_compat_hash_algo(struct repository *repo, int compat_algo);\n void initialize_the_repository(void);\n RESULT_MUST_BE_USED\n int repo_init(struct repository *r, const char *gitdir, const char *worktree);\ndiff --git a/setup.c b/setup.c\nindex ef9f79b8885e..deb5a33fe9e1 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1572,6 +1572,8 @@ const char *setup_git_directory_gently(int *nongit_ok)\n \t\t}\n \t\tif (startup_info->have_repository) {\n \t\t\trepo_set_hash_algo(the_repository, repo_fmt.hash_algo);\n+\t\t\trepo_set_compat_hash_algo(the_repository,\n+\t\t\t\t\t\t  GIT_HASH_UNKNOWN);\n \t\t\tthe_repository->repository_format_worktree_config =\n \t\t\t\trepo_fmt.worktree_config;\n \t\t\t/* take ownership of repo_fmt.partial_clone */\n@@ -1665,6 +1667,7 @@ void check_repository_format(struct repository_format *fmt)\n \tcheck_repository_format_gently(get_git_dir(), fmt, NULL);\n \tstartup_info->have_repository = 1;\n \trepo_set_hash_algo(the_repository, fmt->hash_algo);\n+\trepo_set_compat_hash_algo(the_repository, GIT_HASH_UNKNOWN);\n \tthe_repository->repository_format_worktree_config =\n \t\tfmt->worktree_config;\n \tthe_repository->repository_format_partial_clone =\n-- \n2.41.0\n\n"},{"id":"482384","messageId":"20230927195537.1682-3-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 03/30] object-names: Support input of oids in any supported hash","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:10Z","receivedAt":"2023-09-27T19:56:00Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nSupport short oids encoded in any algorithm, while ensuring enough of\nthe oid is specified to disambiguate between all of the oids in the\nrepository encoded in any algorithm.\n\nBy default have the code continue to only accept oids specified in the\nstorage hash algorithm of the repository, but when something is\nambiguous display all of the possible oids from any oid encoding.\n\nA new flag is added GET_OID_HASH_ANY that when supplied causes the\ncode to accept oids specified in any hash algorithm, and to return the\noids that were resolved.\n\nThis implements the functionality that allows both SHA-1 and SHA-256\nobject names, from the \"Object names on the command line\" section of\nthe hash function transition document.\n\nCare is taken in get_short_oid so that when the result is ambiguous\nthe output remains the same of GIT_OID_HASH_ANY was not supplied.\nIf GET_OID_HASH_ANY was supplied objects of any hash algorithm\nthat match the prefix are displayed.\n\nThis required updating repo_for_each_abbrev to give it a parameter\nso that it knows to look at all hash algorithms.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/rev-parse.c |  2 +-\n hash-ll.h           |  1 +\n object-name.c       | 49 +++++++++++++++++++++++++++++++++++----------\n object-name.h       |  3 ++-\n 4 files changed, 42 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex fde8861ca4e0..43e96765400c 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -882,7 +882,7 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tif (skip_prefix(arg, \"--disambiguate=\", &arg)) {\n-\t\t\t\trepo_for_each_abbrev(the_repository, arg,\n+\t\t\t\trepo_for_each_abbrev(the_repository, arg, the_hash_algo,\n \t\t\t\t\t\t     show_abbrev, NULL);\n \t\t\t\tcontinue;\n \t\t\t}\ndiff --git a/hash-ll.h b/hash-ll.h\nindex 10d84cc20888..2cfde63ae1cf 100644\n--- a/hash-ll.h\n+++ b/hash-ll.h\n@@ -145,6 +145,7 @@ struct object_id {\n #define GET_OID_RECORD_PATH     0200\n #define GET_OID_ONLY_TO_DIE    04000\n #define GET_OID_REQUIRE_PATH  010000\n+#define GET_OID_HASH_ANY      020000\n \n #define GET_OID_DISAMBIGUATORS \\\n \t(GET_OID_COMMIT | GET_OID_COMMITTISH | \\\ndiff --git a/object-name.c b/object-name.c\nindex 0bfa29dbbfe9..976b7106821b 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -25,6 +25,7 @@\n #include \"midx.h\"\n #include \"commit-reach.h\"\n #include \"date.h\"\n+#include \"object-file-convert.h\"\n \n static int get_oid_oneline(struct repository *r, const char *, struct object_id *, struct commit_list *);\n \n@@ -49,6 +50,7 @@ struct disambiguate_state {\n \n static void update_candidates(struct disambiguate_state *ds, const struct object_id *current)\n {\n+\t/* The hash algorithm of the current has already been filtered */\n \tif (ds->always_call_fn) {\n \t\tds->ambiguous = ds->fn(ds->repo, current, ds->cb_data) ? 1 : 0;\n \t\treturn;\n@@ -134,6 +136,8 @@ static void unique_in_midx(struct multi_pack_index *m,\n {\n \tuint32_t num, i, first = 0;\n \tconst struct object_id *current = NULL;\n+\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n+\t\tds->repo->hash_algo->hexsz : ds->len;\n \tnum = m->num_objects;\n \n \tif (!num)\n@@ -149,7 +153,7 @@ static void unique_in_midx(struct multi_pack_index *m,\n \tfor (i = first; i < num && !ds->ambiguous; i++) {\n \t\tstruct object_id oid;\n \t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n-\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, current->hash))\n+\t\tif (!match_hash(len, ds->bin_pfx.hash, current->hash))\n \t\t\tbreak;\n \t\tupdate_candidates(ds, current);\n \t}\n@@ -159,6 +163,8 @@ static void unique_in_pack(struct packed_git *p,\n \t\t\t   struct disambiguate_state *ds)\n {\n \tuint32_t num, i, first = 0;\n+\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n+\t\tds->repo->hash_algo->hexsz : ds->len;\n \n \tif (p->multi_pack_index)\n \t\treturn;\n@@ -177,7 +183,7 @@ static void unique_in_pack(struct packed_git *p,\n \tfor (i = first; i < num && !ds->ambiguous; i++) {\n \t\tstruct object_id oid;\n \t\tnth_packed_object_id(&oid, p, i);\n-\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, oid.hash))\n+\t\tif (!match_hash(len, ds->bin_pfx.hash, oid.hash))\n \t\t\tbreak;\n \t\tupdate_candidates(ds, &oid);\n \t}\n@@ -188,6 +194,10 @@ static void find_short_packed_object(struct disambiguate_state *ds)\n \tstruct multi_pack_index *m;\n \tstruct packed_git *p;\n \n+\t/* Skip, unless oids from the storage hash algorithm are wanted */\n+\tif (ds->bin_pfx.algo && (&hash_algos[ds->bin_pfx.algo] != ds->repo->hash_algo))\n+\t\treturn;\n+\n \tfor (m = get_multi_pack_index(ds->repo); m && !ds->ambiguous;\n \t     m = m->next)\n \t\tunique_in_midx(m, ds);\n@@ -326,11 +336,12 @@ int set_disambiguate_hint_config(const char *var, const char *value)\n \n static int init_object_disambiguation(struct repository *r,\n \t\t\t\t      const char *name, int len,\n+\t\t\t\t      const struct git_hash_algo *algo,\n \t\t\t\t      struct disambiguate_state *ds)\n {\n \tint i;\n \n-\tif (len < MINIMUM_ABBREV || len > the_hash_algo->hexsz)\n+\tif (len < MINIMUM_ABBREV || len > GIT_MAX_HEXSZ)\n \t\treturn -1;\n \n \tmemset(ds, 0, sizeof(*ds));\n@@ -357,6 +368,7 @@ static int init_object_disambiguation(struct repository *r,\n \tds->len = len;\n \tds->hex_pfx[len] = '\\0';\n \tds->repo = r;\n+\tds->bin_pfx.algo = algo ? hash_algo_by_ptr(algo) : GIT_HASH_UNKNOWN;\n \tprepare_alt_odb(r);\n \treturn 0;\n }\n@@ -491,9 +503,10 @@ static int repo_collect_ambiguous(struct repository *r UNUSED,\n \treturn collect_ambiguous(oid, data);\n }\n \n-static int sort_ambiguous(const void *a, const void *b, void *ctx)\n+static int sort_ambiguous(const void *va, const void *vb, void *ctx)\n {\n \tstruct repository *sort_ambiguous_repo = ctx;\n+\tconst struct object_id *a = va, *b = vb;\n \tint a_type = oid_object_info(sort_ambiguous_repo, a, NULL);\n \tint b_type = oid_object_info(sort_ambiguous_repo, b, NULL);\n \tint a_type_sort;\n@@ -503,8 +516,13 @@ static int sort_ambiguous(const void *a, const void *b, void *ctx)\n \t * Sorts by hash within the same object type, just as\n \t * oid_array_for_each_unique() would do.\n \t */\n-\tif (a_type == b_type)\n-\t\treturn oidcmp(a, b);\n+\tif (a_type == b_type) {\n+\t\t/* Is the hash algorithm the same? */\n+\t\tif (a->algo == b->algo)\n+\t\t\treturn oidcmp(a, b);\n+\t\telse\n+\t\t\treturn a->algo > b->algo ? 1 : -1;\n+\t}\n \n \t/*\n \t * Between object types show tags, then commits, and finally\n@@ -533,8 +551,12 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \tint status;\n \tstruct disambiguate_state ds;\n \tint quietly = !!(flags & GET_OID_QUIETLY);\n+\tconst struct git_hash_algo *algo = r->hash_algo;\n+\n+\tif (flags & GET_OID_HASH_ANY)\n+\t\talgo = NULL;\n \n-\tif (init_object_disambiguation(r, name, len, &ds) < 0)\n+\tif (init_object_disambiguation(r, name, len, algo, &ds) < 0)\n \t\treturn -1;\n \n \tif (HAS_MULTI_BITS(flags & GET_OID_DISAMBIGUATORS))\n@@ -553,6 +575,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \telse\n \t\tds.fn = default_disambiguate_hint;\n \n+\n \tfind_short_object_filename(&ds);\n \tfind_short_packed_object(&ds);\n \tstatus = finish_object_disambiguation(&ds, oid);\n@@ -588,7 +611,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t\tif (!ds.ambiguous)\n \t\t\tds.fn = NULL;\n \n-\t\trepo_for_each_abbrev(r, ds.hex_pfx, collect_ambiguous, &collect);\n+\t\trepo_for_each_abbrev(r, ds.hex_pfx, algo, collect_ambiguous, &collect);\n \t\tsort_ambiguous_oid_array(r, &collect);\n \n \t\tif (oid_array_for_each(&collect, show_ambiguous_object, &out))\n@@ -610,15 +633,17 @@ static enum get_oid_result get_short_oid(struct repository *r,\n }\n \n int repo_for_each_abbrev(struct repository *r, const char *prefix,\n+\t\t\t const struct git_hash_algo *algo,\n \t\t\t each_abbrev_fn fn, void *cb_data)\n {\n \tstruct oid_array collect = OID_ARRAY_INIT;\n \tstruct disambiguate_state ds;\n \tint ret;\n \n-\tif (init_object_disambiguation(r, prefix, strlen(prefix), &ds) < 0)\n+\tif (init_object_disambiguation(r, prefix, strlen(prefix), algo, &ds) < 0)\n \t\treturn -1;\n \n+\tds.bin_pfx.algo = GIT_HASH_UNKNOWN;\n \tds.always_call_fn = 1;\n \tds.fn = repo_collect_ambiguous;\n \tds.cb_data = &collect;\n@@ -787,10 +812,12 @@ void strbuf_add_unique_abbrev(struct strbuf *sb, const struct object_id *oid,\n int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \t\t\t      const struct object_id *oid, int len)\n {\n+\tconst struct git_hash_algo *algo =\n+\t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n \tstruct disambiguate_state ds;\n \tstruct min_abbrev_data mad;\n \tstruct object_id oid_ret;\n-\tconst unsigned hexsz = r->hash_algo->hexsz;\n+\tconst unsigned hexsz = algo->hexsz;\n \n \tif (len < 0) {\n \t\tunsigned long count = repo_approximate_object_count(r);\n@@ -826,7 +853,7 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \n \tfind_abbrev_len_packed(&mad);\n \n-\tif (init_object_disambiguation(r, hex, mad.cur_len, &ds) < 0)\n+\tif (init_object_disambiguation(r, hex, mad.cur_len, algo, &ds) < 0)\n \t\treturn -1;\n \n \tds.fn = repo_extend_abbrev_len;\ndiff --git a/object-name.h b/object-name.h\nindex 9ae522307148..064ddc97d1fe 100644\n--- a/object-name.h\n+++ b/object-name.h\n@@ -67,7 +67,8 @@ enum get_oid_result get_oid_with_context(struct repository *repo, const char *st\n \n \n typedef int each_abbrev_fn(const struct object_id *oid, void *);\n-int repo_for_each_abbrev(struct repository *r, const char *prefix, each_abbrev_fn, void *);\n+int repo_for_each_abbrev(struct repository *r, const char *prefix,\n+\t\t\t const struct git_hash_algo *algo, each_abbrev_fn, void *);\n \n int set_disambiguate_hint_config(const char *var, const char *value);\n \n-- \n2.41.0\n\n"},{"id":"482385","messageId":"20230927195537.1682-5-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 05/30] loose: add a mapping between SHA-1 and SHA-256 for loose objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:12Z","receivedAt":"2023-09-27T19:56:02Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nAs part of the transition plan, we'd like to add a file in the .git\ndirectory that maps loose objects between SHA-1 and SHA-256.  Let's\nimplement the specification in the transition plan and store this data\non a per-repository basis in struct repository.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n Makefile              |   1 +\n loose.c               | 245 ++++++++++++++++++++++++++++++++++++++++++\n loose.h               |  22 ++++\n object-file-convert.c |  14 ++-\n object-store-ll.h     |   3 +\n object.c              |   2 +\n repository.c          |   6 ++\n 7 files changed, 292 insertions(+), 1 deletion(-)\n create mode 100644 loose.c\n create mode 100644 loose.h\n\ndiff --git a/Makefile b/Makefile\nindex f7e824f25cda..3c18664def9a 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1053,6 +1053,7 @@ LIB_OBJS += list-objects-filter.o\n LIB_OBJS += list-objects.o\n LIB_OBJS += lockfile.o\n LIB_OBJS += log-tree.o\n+LIB_OBJS += loose.o\n LIB_OBJS += ls-refs.o\n LIB_OBJS += mailinfo.o\n LIB_OBJS += mailmap.o\ndiff --git a/loose.c b/loose.c\nnew file mode 100644\nindex 000000000000..28d11593b2ea\n--- /dev/null\n+++ b/loose.c\n@@ -0,0 +1,245 @@\n+#include \"git-compat-util.h\"\n+#include \"hash.h\"\n+#include \"path.h\"\n+#include \"object-store.h\"\n+#include \"hex.h\"\n+#include \"wrapper.h\"\n+#include \"gettext.h\"\n+#include \"loose.h\"\n+#include \"lockfile.h\"\n+\n+static const char *loose_object_header = \"# loose-object-idx\\n\";\n+\n+static inline int should_use_loose_object_map(struct repository *repo)\n+{\n+\treturn repo->compat_hash_algo && repo->gitdir;\n+}\n+\n+void loose_object_map_init(struct loose_object_map **map)\n+{\n+\tstruct loose_object_map *m;\n+\tm = xmalloc(sizeof(**map));\n+\tm->to_compat = kh_init_oid_map();\n+\tm->to_storage = kh_init_oid_map();\n+\t*map = m;\n+}\n+\n+static int insert_oid_pair(kh_oid_map_t *map, const struct object_id *key, const struct object_id *value)\n+{\n+\tkhiter_t pos;\n+\tint ret;\n+\tstruct object_id *stored;\n+\n+\tpos = kh_put_oid_map(map, *key, &ret);\n+\n+\t/* This item already exists in the map. */\n+\tif (ret == 0)\n+\t\treturn 0;\n+\n+\tstored = xmalloc(sizeof(*stored));\n+\toidcpy(stored, value);\n+\tkh_value(map, pos) = stored;\n+\treturn 1;\n+}\n+\n+static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n+{\n+\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\tFILE *fp;\n+\n+\tif (!dir->loose_map)\n+\t\tloose_object_map_init(&dir->loose_map);\n+\n+\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n+\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n+\n+\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n+\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n+\n+\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n+\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n+\n+\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n+\tfp = fopen(path.buf, \"rb\");\n+\tif (!fp)\n+\t\treturn 0;\n+\n+\terrno = 0;\n+\tif (strbuf_getwholeline(&buf, fp, '\\n') || strcmp(buf.buf, loose_object_header))\n+\t\tgoto err;\n+\twhile (!strbuf_getline_lf(&buf, fp)) {\n+\t\tconst char *p;\n+\t\tstruct object_id oid, compat_oid;\n+\t\tif (parse_oid_hex_algop(buf.buf, &oid, &p, repo->hash_algo) ||\n+\t\t    *p++ != ' ' ||\n+\t\t    parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n+\t\t    p != buf.buf + buf.len)\n+\t\t\tgoto err;\n+\t\tinsert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n+\t\tinsert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n+\t}\n+\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn errno ? -1 : 0;\n+err:\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn -1;\n+}\n+\n+int repo_read_loose_object_map(struct repository *repo)\n+{\n+\tstruct object_directory *dir;\n+\n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n+\tprepare_alt_odb(repo);\n+\n+\tfor (dir = repo->objects->odb; dir; dir = dir->next) {\n+\t\tif (load_one_loose_object_map(repo, dir) < 0) {\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\treturn 0;\n+}\n+\n+int repo_write_loose_object_map(struct repository *repo)\n+{\n+\tkh_oid_map_t *map = repo->objects->odb->loose_map->to_compat;\n+\tstruct lock_file lock;\n+\tint fd;\n+\tkhiter_t iter;\n+\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n+\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n+\tfd = hold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n+\titer = kh_begin(map);\n+\tif (write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n+\t\tgoto errout;\n+\n+\tfor (; iter != kh_end(map); iter++) {\n+\t\tif (kh_exist(map, iter)) {\n+\t\t\tif (oideq(&kh_key(map, iter), the_hash_algo->empty_tree) ||\n+\t\t\t    oideq(&kh_key(map, iter), the_hash_algo->empty_blob))\n+\t\t\t\tcontinue;\n+\t\t\tstrbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(&kh_key(map, iter)), oid_to_hex(kh_value(map, iter)));\n+\t\t\tif (write_in_full(fd, buf.buf, buf.len) < 0)\n+\t\t\t\tgoto errout;\n+\t\t\tstrbuf_reset(&buf);\n+\t\t}\n+\t}\n+\tstrbuf_release(&buf);\n+\tif (commit_lock_file(&lock) < 0) {\n+\t\terror_errno(_(\"could not write loose object index %s\"), path.buf);\n+\t\tstrbuf_release(&path);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_release(&path);\n+\treturn 0;\n+errout:\n+\trollback_lock_file(&lock);\n+\tstrbuf_release(&buf);\n+\terror_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n+\tstrbuf_release(&path);\n+\treturn -1;\n+}\n+\n+static int write_one_object(struct repository *repo, const struct object_id *oid,\n+\t\t\t    const struct object_id *compat_oid)\n+{\n+\tstruct lock_file lock;\n+\tint fd;\n+\tstruct stat st;\n+\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\n+\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n+\thold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n+\n+\tfd = open(path.buf, O_WRONLY | O_CREAT | O_APPEND, 0666);\n+\tif (fd < 0)\n+\t\tgoto errout;\n+\tif (fstat(fd, &st) < 0)\n+\t\tgoto errout;\n+\tif (!st.st_size && write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n+\t\tgoto errout;\n+\n+\tstrbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(oid), oid_to_hex(compat_oid));\n+\tif (write_in_full(fd, buf.buf, buf.len) < 0)\n+\t\tgoto errout;\n+\tif (close(fd))\n+\t\tgoto errout;\n+\tadjust_shared_perm(path.buf);\n+\trollback_lock_file(&lock);\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn 0;\n+errout:\n+\terror_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n+\tclose(fd);\n+\trollback_lock_file(&lock);\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn -1;\n+}\n+\n+int repo_add_loose_object_map(struct repository *repo, const struct object_id *oid,\n+\t\t\t      const struct object_id *compat_oid)\n+{\n+\tint inserted = 0;\n+\n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n+\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_compat, oid, compat_oid);\n+\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_storage, compat_oid, oid);\n+\tif (inserted)\n+\t\treturn write_one_object(repo, oid, compat_oid);\n+\treturn 0;\n+}\n+\n+int repo_loose_object_map_oid(struct repository *repo,\n+\t\t\t      const struct object_id *src,\n+\t\t\t      const struct git_hash_algo *to,\n+\t\t\t      struct object_id *dest)\n+\n+{\n+\tstruct object_directory *dir;\n+\tkh_oid_map_t *map;\n+\tkhiter_t pos;\n+\n+\tfor (dir = repo->objects->odb; dir; dir = dir->next) {\n+\t\tstruct loose_object_map *loose_map = dir->loose_map;\n+\t\tif (!loose_map)\n+\t\t\tcontinue;\n+\t\tmap = (to == repo->compat_hash_algo) ?\n+\t\t\tloose_map->to_compat :\n+\t\t\tloose_map->to_storage;\n+\t\tpos = kh_get_oid_map(map, *src);\n+\t\tif (pos < kh_end(map)) {\n+\t\t\toidcpy(dest, kh_value(map, pos));\n+\t\t\treturn 0;\n+\t\t}\n+\t}\n+\treturn -1;\n+}\n+\n+void loose_object_map_clear(struct loose_object_map **map)\n+{\n+\tstruct loose_object_map *m = *map;\n+\tstruct object_id *oid;\n+\n+\tif (!m)\n+\t\treturn;\n+\n+\tkh_foreach_value(m->to_compat, oid, free(oid));\n+\tkh_foreach_value(m->to_storage, oid, free(oid));\n+\tkh_destroy_oid_map(m->to_compat);\n+\tkh_destroy_oid_map(m->to_storage);\n+\tfree(m);\n+\t*map = NULL;\n+}\ndiff --git a/loose.h b/loose.h\nnew file mode 100644\nindex 000000000000..2c2957072c5f\n--- /dev/null\n+++ b/loose.h\n@@ -0,0 +1,22 @@\n+#ifndef LOOSE_H\n+#define LOOSE_H\n+\n+#include \"khash.h\"\n+\n+struct loose_object_map {\n+\tkh_oid_map_t *to_compat;\n+\tkh_oid_map_t *to_storage;\n+};\n+\n+void loose_object_map_init(struct loose_object_map **map);\n+void loose_object_map_clear(struct loose_object_map **map);\n+int repo_loose_object_map_oid(struct repository *repo,\n+\t\t\t      const struct object_id *src,\n+\t\t\t      const struct git_hash_algo *dest_algo,\n+\t\t\t      struct object_id *dest);\n+int repo_add_loose_object_map(struct repository *repo, const struct object_id *oid,\n+\t\t\t      const struct object_id *compat_oid);\n+int repo_read_loose_object_map(struct repository *repo);\n+int repo_write_loose_object_map(struct repository *repo);\n+\n+#endif\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex ba3e18f6af44..4d62ed192bf0 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -4,6 +4,7 @@\n #include \"repository.h\"\n #include \"hash-ll.h\"\n #include \"object.h\"\n+#include \"loose.h\"\n #include \"object-file-convert.h\"\n \n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n@@ -21,7 +22,18 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \t\t\toidcpy(dest, src);\n \t\treturn 0;\n \t}\n-\treturn -1;\n+\tif (repo_loose_object_map_oid(repo, src, to, dest)) {\n+\t\t/*\n+\t\t * We may have loaded the object map at repo initialization but\n+\t\t * another process (perhaps upstream of a pipe from us) may have\n+\t\t * written a new object into the map.  If the object is missing,\n+\t\t * let's reload the map to see if the object has appeared.\n+\t\t */\n+\t\trepo_read_loose_object_map(repo);\n+\t\tif (repo_loose_object_map_oid(repo, src, to, dest))\n+\t\t\treturn -1;\n+\t}\n+\treturn 0;\n }\n \n int convert_object_file(struct strbuf *outbuf,\ndiff --git a/object-store-ll.h b/object-store-ll.h\nindex 26a3895c821c..bc76d6bec80d 100644\n--- a/object-store-ll.h\n+++ b/object-store-ll.h\n@@ -26,6 +26,9 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/* Map between object IDs for loose objects. */\n+\tstruct loose_object_map *loose_map;\n+\n \t/*\n \t * This is a temporary object store created by the tmp_objdir\n \t * facility. Disable ref updates since the objects in the store\ndiff --git a/object.c b/object.c\nindex 2c61e4c86217..186a0a47c0fb 100644\n--- a/object.c\n+++ b/object.c\n@@ -13,6 +13,7 @@\n #include \"alloc.h\"\n #include \"packfile.h\"\n #include \"commit-graph.h\"\n+#include \"loose.h\"\n \n unsigned int get_max_object_index(void)\n {\n@@ -540,6 +541,7 @@ void free_object_directory(struct object_directory *odb)\n {\n \tfree(odb->path);\n \todb_clear_loose_cache(odb);\n+\tloose_object_map_clear(&odb->loose_map);\n \tfree(odb);\n }\n \ndiff --git a/repository.c b/repository.c\nindex 80252b79e93e..6214f61cf4e7 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -14,6 +14,7 @@\n #include \"read-cache-ll.h\"\n #include \"remote.h\"\n #include \"setup.h\"\n+#include \"loose.h\"\n #include \"submodule-config.h\"\n #include \"sparse-index.h\"\n #include \"trace2.h\"\n@@ -109,6 +110,8 @@ void repo_set_compat_hash_algo(struct repository *repo, int algo)\n \tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n \t\tBUG(\"hash_algo and compat_hash_algo match\");\n \trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n+\tif (repo->compat_hash_algo)\n+\t\trepo_read_loose_object_map(repo);\n }\n \n /*\n@@ -201,6 +204,9 @@ int repo_init(struct repository *repo,\n \tif (worktree)\n \t\trepo_set_worktree(repo, worktree);\n \n+\tif (repo->compat_hash_algo)\n+\t\trepo_read_loose_object_map(repo);\n+\n \tclear_repository_format(&format);\n \treturn 0;\n \n-- \n2.41.0\n\n"},{"id":"482386","messageId":"20230927195537.1682-6-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 06/30] loose: Compatibilty short name support","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:13Z","receivedAt":"2023-09-27T19:56:05Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpdate loose_objects_cache when udpating the loose objects map.  This\noidtree is used to discover which oids are possibilities when\nresolving short names, and it can support a mixture of sha1\nand sha256 oids.\n\nWith this any oid recorded objects/loose-objects-idx is usable\nfor resolving an oid to an object.\n\nTo make this maintainable a helper insert_loose_map is factored\nout of load_one_loose_object_map and repo_add_loose_object_map,\nand then modified to also update the loose_objects_cache.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n loose.c | 37 +++++++++++++++++++++++++------------\n 1 file changed, 25 insertions(+), 12 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 28d11593b2ea..bb50d43cd1d9 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -7,6 +7,7 @@\n #include \"gettext.h\"\n #include \"loose.h\"\n #include \"lockfile.h\"\n+#include \"oidtree.h\"\n \n static const char *loose_object_header = \"# loose-object-idx\\n\";\n \n@@ -42,6 +43,21 @@ static int insert_oid_pair(kh_oid_map_t *map, const struct object_id *key, const\n \treturn 1;\n }\n \n+static int insert_loose_map(struct object_directory *odb,\n+\t\t\t    const struct object_id *oid,\n+\t\t\t    const struct object_id *compat_oid)\n+{\n+\tstruct loose_object_map *map = odb->loose_map;\n+\tint inserted = 0;\n+\n+\tinserted |= insert_oid_pair(map->to_compat, oid, compat_oid);\n+\tinserted |= insert_oid_pair(map->to_storage, compat_oid, oid);\n+\tif (inserted)\n+\t\toidtree_insert(odb->loose_objects_cache, compat_oid);\n+\n+\treturn inserted;\n+}\n+\n static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n {\n \tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n@@ -49,15 +65,14 @@ static int load_one_loose_object_map(struct repository *repo, struct object_dire\n \n \tif (!dir->loose_map)\n \t\tloose_object_map_init(&dir->loose_map);\n+\tif (!dir->loose_objects_cache) {\n+\t\tALLOC_ARRAY(dir->loose_objects_cache, 1);\n+\t\toidtree_init(dir->loose_objects_cache);\n+\t}\n \n-\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n-\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n-\n-\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n-\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n-\n-\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n-\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n+\tinsert_loose_map(dir, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n+\tinsert_loose_map(dir, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n+\tinsert_loose_map(dir, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n \n \tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n \tfp = fopen(path.buf, \"rb\");\n@@ -75,8 +90,7 @@ static int load_one_loose_object_map(struct repository *repo, struct object_dire\n \t\t    parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n \t\t    p != buf.buf + buf.len)\n \t\t\tgoto err;\n-\t\tinsert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n-\t\tinsert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n+\t\tinsert_loose_map(dir, &oid, &compat_oid);\n \t}\n \n \tstrbuf_release(&buf);\n@@ -195,8 +209,7 @@ int repo_add_loose_object_map(struct repository *repo, const struct object_id *o\n \tif (!should_use_loose_object_map(repo))\n \t\treturn 0;\n \n-\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_compat, oid, compat_oid);\n-\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_storage, compat_oid, oid);\n+\tinserted = insert_loose_map(repo->objects->odb, oid, compat_oid);\n \tif (inserted)\n \t\treturn write_one_object(repo, oid, compat_oid);\n \treturn 0;\n-- \n2.41.0\n\n"},{"id":"482387","messageId":"20230927195537.1682-8-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 08/30] object-file: Add a compat_oid_in parameter to write_object_file_flags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:15Z","receivedAt":"2023-09-27T19:56:07Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nTo create the proper signatures for commit objects both versions of\nthe commit object need to be generated and signed.  After that it is\na waste to throw away the work of generating the compatibility hash\nso update write_object_file_flags to take a compatibility hash input\nparameter that it can use to skip the work of generating the\ncompatability hash.\n\nUpdate the places that don't generate the compatability hash to\npass NULL so it is easy to tell write_object_file_flags should\nnot attempt to use their compatability hash.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n cache-tree.c      | 2 +-\n object-file.c     | 6 ++++--\n object-store-ll.h | 4 ++--\n 3 files changed, 7 insertions(+), 5 deletions(-)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 641427ed410a..ddc7d3d86959 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -448,7 +448,7 @@ static int update_one(struct cache_tree *it,\n \t\thash_object_file(the_hash_algo, buffer.buf, buffer.len,\n \t\t\t\t OBJ_TREE, &it->oid);\n \t} else if (write_object_file_flags(buffer.buf, buffer.len, OBJ_TREE,\n-\t\t\t\t\t   &it->oid, flags & WRITE_TREE_SILENT\n+\t\t\t\t\t   &it->oid, NULL, flags & WRITE_TREE_SILENT\n \t\t\t\t\t   ? HASH_SILENT : 0)) {\n \t\tstrbuf_release(&buffer);\n \t\treturn -1;\ndiff --git a/object-file.c b/object-file.c\nindex 4ad31f25c555..d66d11890696 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2235,7 +2235,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n-\t\t\t    unsigned flags)\n+\t\t\t    struct object_id *compat_oid_in, unsigned flags)\n {\n \tstruct repository *repo = the_repository;\n \tconst struct git_hash_algo *algo = repo->hash_algo;\n@@ -2246,7 +2246,9 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \n \t/* Generate compat_oid */\n \tif (compat) {\n-\t\tif (type == OBJ_BLOB)\n+\t\tif (compat_oid_in)\n+\t\t\toidcpy(&compat_oid, compat_oid_in);\n+\t\telse if (type == OBJ_BLOB)\n \t\t\thash_object_file(compat, buf, len, type, &compat_oid);\n \t\telse {\n \t\t\tstruct strbuf converted = STRBUF_INIT;\ndiff --git a/object-store-ll.h b/object-store-ll.h\nindex bc76d6bec80d..c5f2bb2fc2fe 100644\n--- a/object-store-ll.h\n+++ b/object-store-ll.h\n@@ -255,11 +255,11 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n-\t\t\t    unsigned flags);\n+\t\t\t    struct object_id *comapt_oid_in, unsigned flags);\n static inline int write_object_file(const void *buf, unsigned long len,\n \t\t\t\t    enum object_type type, struct object_id *oid)\n {\n-\treturn write_object_file_flags(buf, len, type, oid, 0);\n+\treturn write_object_file_flags(buf, len, type, oid, NULL, 0);\n }\n \n int write_object_file_literally(const void *buf, unsigned long len,\n-- \n2.41.0\n\n"},{"id":"482388","messageId":"20230927195537.1682-7-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 07/30] object-file: Update the loose object map when writing loose objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:14Z","receivedAt":"2023-09-27T19:56:09Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nTo implement SHA1 compatibility on SHA256 repositories the loose\nobject map needs to be updated whenver a loose object is written.\nUpdating the loose object map this way allows git to support\nthe old hash algorithm in constant time.\n\nThe functions write_loose_object, and stream_loose_object are\nthe only two functions that write to the loose object store.\n\nUpdate stream_loose_object to compute the compatibiilty hash, update\nthe loose object, and then call repo_add_loose_object_map to update\nthe loose object map.\n\nUpdate write_object_file_flags to convert the object into\nit's compatibility encoding, hash the compatibility encoding,\nwrite the object, and then update the loose object map.\n\nUpdate force_object_loose to lookup the hash of the compatibility\nencoding, write the loose object, and then update the loose object\nmap.\n\nUpdate write_object_file_literally to convert the object into it's\ncompatibility hash encoding, hash the compatibility enconding, write\nthe object, and then update the loose object map, when the type string\nis a known type.  For objects with an unknown type this results in a\npartially broken repository, as the objects are not mapped.\n\nThe point of write_object_file_literally is to generate a partially\nbroken repository for testing.  For testing skipping writing the loose\nobject map is much more useful than refusing to write the broken\nobject at all.\n\nExcept that the loose objects are updated before the loose object map\nI have not done any analysis to see how robust this scheme is in the\nevent of failure.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n object-file.c | 113 ++++++++++++++++++++++++++++++++++++++++++--------\n 1 file changed, 95 insertions(+), 18 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 7c7afe579364..4ad31f25c555 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -43,6 +43,8 @@\n #include \"setup.h\"\n #include \"submodule.h\"\n #include \"fsck.h\"\n+#include \"loose.h\"\n+#include \"object-file-convert.h\"\n \n /* The maximum size for an object header. */\n #define MAX_HEADER_LEN 32\n@@ -1952,9 +1954,12 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \t\t\t\t     const char *filename, unsigned flags,\n \t\t\t\t     git_zstream *stream,\n \t\t\t\t     unsigned char *buf, size_t buflen,\n-\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     git_hash_ctx *c, git_hash_ctx *compat_c,\n \t\t\t\t     char *hdr, int hdrlen)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tint fd;\n \n \tfd = create_tmpfile(tmp_file, filename);\n@@ -1974,14 +1979,18 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \tgit_deflate_init(stream, zlib_compression_level);\n \tstream->next_out = buf;\n \tstream->avail_out = buflen;\n-\tthe_hash_algo->init_fn(c);\n+\talgo->init_fn(c);\n+\tif (compat && compat_c)\n+\t\tcompat->init_fn(compat_c);\n \n \t/*  Start to feed header to zlib stream */\n \tstream->next_in = (unsigned char *)hdr;\n \tstream->avail_in = hdrlen;\n \twhile (git_deflate(stream, 0) == Z_OK)\n \t\t; /* nothing */\n-\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\talgo->update_fn(c, hdr, hdrlen);\n+\tif (compat && compat_c)\n+\t\tcompat->update_fn(compat_c, hdr, hdrlen);\n \n \treturn fd;\n }\n@@ -1990,16 +1999,21 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n  * Common steps for the inner git_deflate() loop for writing loose\n  * objects. Returns what git_deflate() returns.\n  */\n-static int write_loose_object_common(git_hash_ctx *c,\n+static int write_loose_object_common(git_hash_ctx *c, git_hash_ctx *compat_c,\n \t\t\t\t     git_zstream *stream, const int flush,\n \t\t\t\t     unsigned char *in0, const int fd,\n \t\t\t\t     unsigned char *compressed,\n \t\t\t\t     const size_t compressed_len)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tint ret;\n \n \tret = git_deflate(stream, flush ? Z_FINISH : 0);\n-\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n+\talgo->update_fn(c, in0, stream->next_in - in0);\n+\tif (compat && compat_c)\n+\t\tcompat->update_fn(compat_c, in0, stream->next_in - in0);\n \tif (write_in_full(fd, compressed, stream->next_out - compressed) < 0)\n \t\tdie_errno(_(\"unable to write loose object file\"));\n \tstream->next_out = compressed;\n@@ -2014,15 +2028,21 @@ static int write_loose_object_common(git_hash_ctx *c,\n  * - End the compression of zlib stream.\n  * - Get the calculated oid to \"oid\".\n  */\n-static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n-\t\t\t\t   struct object_id *oid)\n+static int end_loose_object_common(git_hash_ctx *c, git_hash_ctx *compat_c,\n+\t\t\t\t   git_zstream *stream, struct object_id *oid,\n+\t\t\t\t   struct object_id *compat_oid)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tint ret;\n \n \tret = git_deflate_end_gently(stream);\n \tif (ret != Z_OK)\n \t\treturn ret;\n-\tthe_hash_algo->final_oid_fn(oid, c);\n+\talgo->final_oid_fn(oid, c);\n+\tif (compat && compat_c)\n+\t\tcompat->final_oid_fn(compat_oid, compat_c);\n \n \treturn Z_OK;\n }\n@@ -2046,7 +2066,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n \t\t\t\t       &stream, compressed, sizeof(compressed),\n-\t\t\t\t       &c, hdr, hdrlen);\n+\t\t\t\t       &c, NULL, hdr, hdrlen);\n \tif (fd < 0)\n \t\treturn -1;\n \n@@ -2056,14 +2076,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n \n-\t\tret = write_loose_object_common(&c, &stream, 1, in0, fd,\n+\t\tret = write_loose_object_common(&c, NULL, &stream, 1, in0, fd,\n \t\t\t\t\t\tcompressed, sizeof(compressed));\n \t} while (ret == Z_OK);\n \n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n-\tret = end_loose_object_common(&c, &stream, &parano_oid);\n+\tret = end_loose_object_common(&c, NULL, &stream, &parano_oid, NULL);\n \tif (ret != Z_OK)\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n@@ -2108,10 +2128,12 @@ static int freshen_packed_object(const struct object_id *oid)\n int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t\t\tstruct object_id *oid)\n {\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n+\tstruct object_id compat_oid;\n \tint fd, ret, err = 0, flush = 0;\n \tunsigned char compressed[4096];\n \tgit_zstream stream;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, compat_c;\n \tstruct strbuf tmp_file = STRBUF_INIT;\n \tstruct strbuf filename = STRBUF_INIT;\n \tint dirlen;\n@@ -2135,7 +2157,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t */\n \tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n \t\t\t\t       &stream, compressed, sizeof(compressed),\n-\t\t\t\t       &c, hdr, hdrlen);\n+\t\t\t\t       &c, &compat_c, hdr, hdrlen);\n \tif (fd < 0) {\n \t\terr = -1;\n \t\tgoto cleanup;\n@@ -2153,7 +2175,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t\t\tif (in_stream->is_finished)\n \t\t\t\tflush = 1;\n \t\t}\n-\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n+\t\tret = write_loose_object_common(&c, &compat_c, &stream, flush, in0, fd,\n \t\t\t\t\t\tcompressed, sizeof(compressed));\n \t\t/*\n \t\t * Unlike write_loose_object(), we do not have the entire\n@@ -2176,7 +2198,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t */\n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n-\tret = end_loose_object_common(&c, &stream, oid);\n+\tret = end_loose_object_common(&c, &compat_c, &stream, oid, &compat_oid);\n \tif (ret != Z_OK)\n \t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n \tclose_loose_object(fd, tmp_file.buf);\n@@ -2203,6 +2225,8 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t}\n \n \terr = finalize_object_file(tmp_file.buf, filename.buf);\n+\tif (!err && compat)\n+\t\terr = repo_add_loose_object_map(the_repository, oid, &compat_oid);\n cleanup:\n \tstrbuf_release(&tmp_file);\n \tstrbuf_release(&filename);\n@@ -2213,17 +2237,38 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n \t\t\t    unsigned flags)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct object_id compat_oid;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen = sizeof(hdr);\n \n+\t/* Generate compat_oid */\n+\tif (compat) {\n+\t\tif (type == OBJ_BLOB)\n+\t\t\thash_object_file(compat, buf, len, type, &compat_oid);\n+\t\telse {\n+\t\t\tstruct strbuf converted = STRBUF_INIT;\n+\t\t\tconvert_object_file(&converted, algo, compat,\n+\t\t\t\t\t    buf, len, type, 0);\n+\t\t\thash_object_file(compat, converted.buf, converted.len,\n+\t\t\t\t\t type, &compat_oid);\n+\t\t\tstrbuf_release(&converted);\n+\t\t}\n+\t}\n+\n \t/* Normally if we have it in the pack then we do not bother writing\n \t * it out into .git/objects/??/?{38} file.\n \t */\n-\twrite_object_file_prepare(the_hash_algo, buf, len, type, oid, hdr,\n-\t\t\t\t  &hdrlen);\n+\twrite_object_file_prepare(algo, buf, len, type, oid, hdr, &hdrlen);\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\treturn 0;\n-\treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n+\tif (write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags))\n+\t\treturn -1;\n+\tif (compat)\n+\t\treturn repo_add_loose_object_map(repo, oid, &compat_oid);\n+\treturn 0;\n }\n \n int write_object_file_literally(const void *buf, unsigned long len,\n@@ -2231,7 +2276,27 @@ int write_object_file_literally(const void *buf, unsigned long len,\n \t\t\t\tunsigned flags)\n {\n \tchar *header;\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct object_id compat_oid;\n \tint hdrlen, status = 0;\n+\tint compat_type = -1;\n+\n+\tif (compat) {\n+\t\tcompat_type = type_from_string_gently(type, -1, 1);\n+\t\tif (compat_type == OBJ_BLOB)\n+\t\t\thash_object_file(compat, buf, len, compat_type,\n+\t\t\t\t\t &compat_oid);\n+\t\telse if (compat_type != -1) {\n+\t\t\tstruct strbuf converted = STRBUF_INIT;\n+\t\t\tconvert_object_file(&converted, algo, compat,\n+\t\t\t\t\t    buf, len, compat_type, 0);\n+\t\t\thash_object_file(compat, converted.buf, converted.len,\n+\t\t\t\t\t compat_type, &compat_oid);\n+\t\t\tstrbuf_release(&converted);\n+\t\t}\n+\t}\n \n \t/* type string, SP, %lu of the length plus NUL must fit this */\n \thdrlen = strlen(type) + MAX_HEADER_LEN;\n@@ -2244,6 +2309,8 @@ int write_object_file_literally(const void *buf, unsigned long len,\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\tgoto cleanup;\n \tstatus = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n+\tif (compat_type != -1)\n+\t\treturn repo_add_loose_object_map(repo, oid, &compat_oid);\n \n cleanup:\n \tfree(header);\n@@ -2252,9 +2319,12 @@ int write_object_file_literally(const void *buf, unsigned long len,\n \n int force_object_loose(const struct object_id *oid, time_t mtime)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tvoid *buf;\n \tunsigned long len;\n \tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct object_id compat_oid;\n \tenum object_type type;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n@@ -2267,8 +2337,15 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \toi.contentp = &buf;\n \tif (oid_object_info_extended(the_repository, oid, &oi, 0))\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n+\tif (compat) {\n+\t\tif (repo_oid_to_algop(repo, oid, compat, &compat_oid))\n+\t\t\treturn error(_(\"cannot map object %s to %s\"),\n+\t\t\t\t     oid_to_hex(oid), compat->name);\n+\t}\n \thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n+\tif (!ret && compat)\n+\t\tret = repo_add_loose_object_map(the_repository, oid, &compat_oid);\n \tfree(buf);\n \n \treturn ret;\n-- \n2.41.0\n\n"},{"id":"482389","messageId":"20230927195537.1682-9-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 09/30] commit: write commits for both hashes","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:16Z","receivedAt":"2023-09-27T19:56:11Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nWhen we write a commit, we include data that is specific to the hash\nalgorithm, such as parents and the root tree.  In order to write both a\nSHA-1 commit and a SHA-256 version, we need to convert between them.\n\nHowever, a straightforward conversion isn't necessarily what we want.\nWhen we sign a commit, we sign its data, so if we create a commit for\nSHA-256 and then write a SHA-1 version, we'll still have only signed the\nSHA-256 data.  While this is valid, it would be better to sign both\nforms of data so people using SHA-1 can verify the signatures as well.\n\nConsequently, we don't want to use the standard mapping that occurs when\nwe write an object.  Instead, let's move most of the writing of the\ncommit into a separate function which is agnostic of the hash algorithm\nand which simply writes into a buffer and specify both versions of the\nobject ourselves.\n\nWe can then call this function twice: once with the SHA-256 contents,\nand if SHA-1 is enabled, once with the SHA-1 contents.  If we're signing\nthe commit, we then sign both versions and append both signatures to\nboth buffers.  To produce a consistent hash, we always append the\nsignatures in the order in which Git implemented them: first SHA-1, then\nSHA-256.\n\nIn order to make this signing code work, we split the commit signing\ncode into two functions, one which signs the buffer, and one which\nappends the signature.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n commit.c | 179 +++++++++++++++++++++++++++++++++++++++++--------------\n 1 file changed, 134 insertions(+), 45 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex b3223478bc2a..46696ede8981 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -28,6 +28,7 @@\n #include \"shallow.h\"\n #include \"tree.h\"\n #include \"hook.h\"\n+#include \"object-file-convert.h\"\n \n static struct commit_extra_header *read_commit_extra_header_lines(const char *buf, size_t len, const char **);\n \n@@ -1100,12 +1101,11 @@ static const char *gpg_sig_headers[] = {\n \t\"gpgsig-sha256\",\n };\n \n-int sign_with_header(struct strbuf *buf, const char *keyid)\n+static int add_commit_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo)\n {\n-\tstruct strbuf sig = STRBUF_INIT;\n \tint inspos, copypos;\n \tconst char *eoh;\n-\tconst char *gpg_sig_header = gpg_sig_headers[hash_algo_by_ptr(the_hash_algo)];\n+\tconst char *gpg_sig_header = gpg_sig_headers[hash_algo_by_ptr(algo)];\n \tint gpg_sig_header_len = strlen(gpg_sig_header);\n \n \t/* find the end of the header */\n@@ -1115,15 +1115,8 @@ int sign_with_header(struct strbuf *buf, const char *keyid)\n \telse\n \t\tinspos = eoh - buf->buf + 1;\n \n-\tif (!keyid || !*keyid)\n-\t\tkeyid = get_signing_key();\n-\tif (sign_buffer(buf, &sig, keyid)) {\n-\t\tstrbuf_release(&sig);\n-\t\treturn -1;\n-\t}\n-\n-\tfor (copypos = 0; sig.buf[copypos]; ) {\n-\t\tconst char *bol = sig.buf + copypos;\n+\tfor (copypos = 0; sig->buf[copypos]; ) {\n+\t\tconst char *bol = sig->buf + copypos;\n \t\tconst char *eol = strchrnul(bol, '\\n');\n \t\tint len = (eol - bol) + !!*eol;\n \n@@ -1136,11 +1129,17 @@ int sign_with_header(struct strbuf *buf, const char *keyid)\n \t\tinspos += len;\n \t\tcopypos += len;\n \t}\n-\tstrbuf_release(&sig);\n \treturn 0;\n }\n \n-\n+static int sign_commit_to_strbuf(struct strbuf *sig, struct strbuf *buf, const char *keyid)\n+{\n+\tif (!keyid || !*keyid)\n+\t\tkeyid = get_signing_key();\n+\tif (sign_buffer(buf, sig, keyid))\n+\t\treturn -1;\n+\treturn 0;\n+}\n \n int parse_signed_commit(const struct commit *commit,\n \t\t\tstruct strbuf *payload, struct strbuf *signature,\n@@ -1599,70 +1598,160 @@ N_(\"Warning: commit message did not conform to UTF-8.\\n\"\n    \"You may want to amend it after fixing the message, or set the config\\n\"\n    \"variable i18n.commitEncoding to the encoding your project uses.\\n\");\n \n-int commit_tree_extended(const char *msg, size_t msg_len,\n-\t\t\t const struct object_id *tree,\n-\t\t\t struct commit_list *parents, struct object_id *ret,\n-\t\t\t const char *author, const char *committer,\n-\t\t\t const char *sign_commit,\n-\t\t\t struct commit_extra_header *extra)\n+static void write_commit_tree(struct strbuf *buffer, const char *msg, size_t msg_len,\n+\t\t\t      const struct object_id *tree,\n+\t\t\t      const struct object_id *parents, size_t parents_len,\n+\t\t\t      const char *author, const char *committer,\n+\t\t\t      struct commit_extra_header *extra)\n {\n-\tint result;\n \tint encoding_is_utf8;\n-\tstruct strbuf buffer;\n-\n-\tassert_oid_type(tree, OBJ_TREE);\n-\n-\tif (memchr(msg, '\\0', msg_len))\n-\t\treturn error(\"a NUL byte in commit log message not allowed.\");\n+\tsize_t i;\n \n \t/* Not having i18n.commitencoding is the same as having utf-8 */\n \tencoding_is_utf8 = is_encoding_utf8(git_commit_encoding);\n \n-\tstrbuf_init(&buffer, 8192); /* should avoid reallocs for the headers */\n-\tstrbuf_addf(&buffer, \"tree %s\\n\", oid_to_hex(tree));\n+\tstrbuf_init(buffer, 8192); /* should avoid reallocs for the headers */\n+\tstrbuf_addf(buffer, \"tree %s\\n\", oid_to_hex(tree));\n \n \t/*\n \t * NOTE! This ordering means that the same exact tree merged with a\n \t * different order of parents will be a _different_ changeset even\n \t * if everything else stays the same.\n \t */\n-\twhile (parents) {\n-\t\tstruct commit *parent = pop_commit(&parents);\n-\t\tstrbuf_addf(&buffer, \"parent %s\\n\",\n-\t\t\t    oid_to_hex(&parent->object.oid));\n-\t}\n+\tfor (i = 0; i < parents_len; i++)\n+\t\tstrbuf_addf(buffer, \"parent %s\\n\", oid_to_hex(&parents[i]));\n \n \t/* Person/date information */\n \tif (!author)\n \t\tauthor = git_author_info(IDENT_STRICT);\n-\tstrbuf_addf(&buffer, \"author %s\\n\", author);\n+\tstrbuf_addf(buffer, \"author %s\\n\", author);\n \tif (!committer)\n \t\tcommitter = git_committer_info(IDENT_STRICT);\n-\tstrbuf_addf(&buffer, \"committer %s\\n\", committer);\n+\tstrbuf_addf(buffer, \"committer %s\\n\", committer);\n \tif (!encoding_is_utf8)\n-\t\tstrbuf_addf(&buffer, \"encoding %s\\n\", git_commit_encoding);\n+\t\tstrbuf_addf(buffer, \"encoding %s\\n\", git_commit_encoding);\n \n \twhile (extra) {\n-\t\tadd_extra_header(&buffer, extra);\n+\t\tadd_extra_header(buffer, extra);\n \t\textra = extra->next;\n \t}\n-\tstrbuf_addch(&buffer, '\\n');\n+\tstrbuf_addch(buffer, '\\n');\n \n \t/* And add the comment */\n-\tstrbuf_add(&buffer, msg, msg_len);\n+\tstrbuf_add(buffer, msg, msg_len);\n+}\n \n-\t/* And check the encoding */\n-\tif (encoding_is_utf8 && !verify_utf8(&buffer))\n-\t\tfprintf(stderr, _(commit_utf8_warn));\n+int commit_tree_extended(const char *msg, size_t msg_len,\n+\t\t\t const struct object_id *tree,\n+\t\t\t struct commit_list *parents, struct object_id *ret,\n+\t\t\t const char *author, const char *committer,\n+\t\t\t const char *sign_commit,\n+\t\t\t struct commit_extra_header *extra)\n+{\n+\tstruct repository *r = the_repository;\n+\tint result = 0;\n+\tint encoding_is_utf8;\n+\tstruct strbuf buffer, compat_buffer;\n+\tstruct strbuf sig = STRBUF_INIT, compat_sig = STRBUF_INIT;\n+\tstruct object_id *parent_buf = NULL, *compat_oid = NULL;\n+\tstruct object_id compat_oid_buf;\n+\tsize_t i, nparents;\n+\n+\t/* Not having i18n.commitencoding is the same as having utf-8 */\n+\tencoding_is_utf8 = is_encoding_utf8(git_commit_encoding);\n+\n+\tassert_oid_type(tree, OBJ_TREE);\n+\n+\tif (memchr(msg, '\\0', msg_len))\n+\t\treturn error(\"a NUL byte in commit log message not allowed.\");\n+\n+\tnparents = commit_list_count(parents);\n+\tparent_buf = xcalloc(nparents, sizeof(*parent_buf));\n+\tfor (i = 0; i < nparents; i++) {\n+\t\tstruct commit *parent = pop_commit(&parents);\n+\t\toidcpy(&parent_buf[i], &parent->object.oid);\n+\t}\n+\n+\t/* should avoid reallocs for the headers */\n+\tstrbuf_init(&buffer, 8192);\n+\tstrbuf_init(&compat_buffer, 8192);\n \n-\tif (sign_commit && sign_with_header(&buffer, sign_commit)) {\n+\twrite_commit_tree(&buffer, msg, msg_len, tree, parent_buf, nparents, author, committer, extra);\n+\tif (sign_commit && sign_commit_to_strbuf(&sig, &buffer, sign_commit)) {\n \t\tresult = -1;\n \t\tgoto out;\n \t}\n+\tif (r->compat_hash_algo) {\n+\t\tstruct object_id mapped_tree;\n+\t\tstruct object_id *mapped_parents = xcalloc(nparents, sizeof(*mapped_parents));\n+\t\tif (repo_oid_to_algop(r, tree, r->compat_hash_algo, &mapped_tree)) {\n+\t\t\tresult = -1;\n+\t\t\tfree(mapped_parents);\n+\t\t\tgoto out;\n+\t\t}\n+\t\tfor (i = 0; i < nparents; i++)\n+\t\t\tif (repo_oid_to_algop(r, &parent_buf[i], r->compat_hash_algo, &mapped_parents[i])) {\n+\t\t\t\tresult = -1;\n+\t\t\t\tfree(mapped_parents);\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\t\twrite_commit_tree(&compat_buffer, msg, msg_len, &mapped_tree,\n+\t\t\t\t  mapped_parents, nparents, author, committer, extra);\n \n-\tresult = write_object_file(buffer.buf, buffer.len, OBJ_COMMIT, ret);\n+\t\tif (sign_commit && sign_commit_to_strbuf(&compat_sig, &compat_buffer, sign_commit)) {\n+\t\t\tresult = -1;\n+\t\t\tgoto out;\n+\t\t}\n+\t}\n+\n+\tif (sign_commit) {\n+\t\tstruct sig_pairs {\n+\t\t\tstruct strbuf *sig;\n+\t\t\tconst struct git_hash_algo *algo;\n+\t\t} bufs [2] = {\n+\t\t\t{ &compat_sig, r->compat_hash_algo },\n+\t\t\t{ &sig, r->hash_algo },\n+\t\t};\n+\t\tint i;\n+\n+\t\t/*\n+\t\t * We write algorithms in the order they were implemented in\n+\t\t * Git to produce a stable hash when multiple algorithms are\n+\t\t * used.\n+\t\t */\n+\t\tif (r->compat_hash_algo && hash_algo_by_ptr(bufs[0].algo) > hash_algo_by_ptr(bufs[1].algo))\n+\t\t\tSWAP(bufs[0], bufs[1]);\n+\n+\t\t/*\n+\t\t * We traverse each algorithm in order, and apply the signature\n+\t\t * to each buffer.\n+\t\t */\n+\t\tfor (i = 0; i < ARRAY_SIZE(bufs); i++) {\n+\t\t\tif (!bufs[i].algo)\n+\t\t\t\tcontinue;\n+\t\t\tadd_commit_signature(&buffer, bufs[i].sig, bufs[i].algo);\n+\t\t\tif (r->compat_hash_algo)\n+\t\t\t\tadd_commit_signature(&compat_buffer, bufs[i].sig, bufs[i].algo);\n+\t\t}\n+\t}\n+\n+\t/* And check the encoding. */\n+\tif (encoding_is_utf8 && (!verify_utf8(&buffer) || !verify_utf8(&compat_buffer)))\n+\t\tfprintf(stderr, _(commit_utf8_warn));\n+\n+\tif (r->compat_hash_algo) {\n+\t\thash_object_file(r->compat_hash_algo, compat_buffer.buf, compat_buffer.len,\n+\t\t\tOBJ_COMMIT, &compat_oid_buf);\n+\t\tcompat_oid = &compat_oid_buf;\n+\t}\n+\n+\tresult = write_object_file_flags(buffer.buf, buffer.len, OBJ_COMMIT,\n+\t\t\t\t\t ret, compat_oid, 0);\n out:\n \tstrbuf_release(&buffer);\n+\tstrbuf_release(&compat_buffer);\n+\tstrbuf_release(&sig);\n+\tstrbuf_release(&compat_sig);\n \treturn result;\n }\n \n-- \n2.41.0\n\n"},{"id":"482390","messageId":"20230927195537.1682-10-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 10/30] commit: Convert mergetag before computing the signature of a commit","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:17Z","receivedAt":"2023-09-27T19:56:13Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nIt so happens that commit mergetag lines embed a tag object.  So to\ncompute the compatible signature of a commit object that has mergetag\nlines the compatible embedded tag must be computed first.\n\nImplement this by duplicating and converting the commit extra headers\ninto the compatible version of the commit extra headers, that need\nto be passed to commit_tree_extended.\n\nTo handle merge tags only the compatible extra headers need to be\ncomputed.\n\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n commit.c | 42 +++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 41 insertions(+), 1 deletion(-)\n\ndiff --git a/commit.c b/commit.c\nindex 46696ede8981..a6dac9a1957b 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1355,6 +1355,39 @@ void append_merge_tag_headers(struct commit_list *parents,\n \t}\n }\n \n+static int convert_commit_extra_headers(struct commit_extra_header *orig,\n+\t\t\t\t\tstruct commit_extra_header **result)\n+{\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n+\tconst struct git_hash_algo *algo = the_repository->hash_algo;\n+\tstruct commit_extra_header *extra = NULL, **tail = &extra;\n+\tstruct strbuf out = STRBUF_INIT;\n+\twhile (orig) {\n+\t\tstruct commit_extra_header *new;\n+\t\tCALLOC_ARRAY(new, 1);\n+\t\tif (!strcmp(orig->key, \"mergetag\")) {\n+\t\t\tif (convert_object_file(&out, algo, compat,\n+\t\t\t\t\t\torig->value, orig->len,\n+\t\t\t\t\t\tOBJ_TAG, 1)) {\n+\t\t\t\tfree(new);\n+\t\t\t\tfree_commit_extra_headers(extra);\n+\t\t\t\treturn -1;\n+\t\t\t}\n+\t\t\tnew->key = xstrdup(\"mergetag\");\n+\t\t\tnew->value = strbuf_detach(&out, &new->len);\n+\t\t} else {\n+\t\t\tnew->key = xstrdup(orig->key);\n+\t\t\tnew->len = orig->len;\n+\t\t\tnew->value = xmemdupz(orig->value, orig->len);\n+\t\t}\n+\t\t*tail = new;\n+\t\ttail = &new->next;\n+\t\torig = orig->next;\n+\t}\n+\t*result = extra;\n+\treturn 0;\n+}\n+\n static void add_extra_header(struct strbuf *buffer,\n \t\t\t     struct commit_extra_header *extra)\n {\n@@ -1682,6 +1715,7 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n \t\tgoto out;\n \t}\n \tif (r->compat_hash_algo) {\n+\t\tstruct commit_extra_header *compat_extra = NULL;\n \t\tstruct object_id mapped_tree;\n \t\tstruct object_id *mapped_parents = xcalloc(nparents, sizeof(*mapped_parents));\n \t\tif (repo_oid_to_algop(r, tree, r->compat_hash_algo, &mapped_tree)) {\n@@ -1695,8 +1729,14 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n \t\t\t\tfree(mapped_parents);\n \t\t\t\tgoto out;\n \t\t\t}\n+\t\tif (convert_commit_extra_headers(extra, &compat_extra)) {\n+\t\t\tresult = -1;\n+\t\t\tfree(mapped_parents);\n+\t\t\tgoto out;\n+\t\t}\n \t\twrite_commit_tree(&compat_buffer, msg, msg_len, &mapped_tree,\n-\t\t\t\t  mapped_parents, nparents, author, committer, extra);\n+\t\t\t\t  mapped_parents, nparents, author, committer, compat_extra);\n+\t\tfree_commit_extra_headers(compat_extra);\n \n \t\tif (sign_commit && sign_commit_to_strbuf(&compat_sig, &compat_buffer, sign_commit)) {\n \t\t\tresult = -1;\n-- \n2.41.0\n\n"},{"id":"482391","messageId":"20230927195537.1682-11-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 11/30] commit: Export add_header_signature to support handling signatures on tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:18Z","receivedAt":"2023-09-27T19:56:14Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nRename add_commit_signature as add_header_signature, and expose it so\nthat it can be used for converting tags from one object format to\nanother.\n\nInspired-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n commit.c | 6 +++---\n commit.h | 1 +\n 2 files changed, 4 insertions(+), 3 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex a6dac9a1957b..e75270171bc3 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1101,7 +1101,7 @@ static const char *gpg_sig_headers[] = {\n \t\"gpgsig-sha256\",\n };\n \n-static int add_commit_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo)\n+int add_header_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo)\n {\n \tint inspos, copypos;\n \tconst char *eoh;\n@@ -1769,9 +1769,9 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n \t\tfor (i = 0; i < ARRAY_SIZE(bufs); i++) {\n \t\t\tif (!bufs[i].algo)\n \t\t\t\tcontinue;\n-\t\t\tadd_commit_signature(&buffer, bufs[i].sig, bufs[i].algo);\n+\t\t\tadd_header_signature(&buffer, bufs[i].sig, bufs[i].algo);\n \t\t\tif (r->compat_hash_algo)\n-\t\t\t\tadd_commit_signature(&compat_buffer, bufs[i].sig, bufs[i].algo);\n+\t\t\t\tadd_header_signature(&compat_buffer, bufs[i].sig, bufs[i].algo);\n \t\t}\n \t}\n \ndiff --git a/commit.h b/commit.h\nindex 28928833c544..03edcec0129f 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -370,5 +370,6 @@ int parse_buffer_signed_by_header(const char *buffer,\n \t\t\t\t  struct strbuf *payload,\n \t\t\t\t  struct strbuf *signature,\n \t\t\t\t  const struct git_hash_algo *algop);\n+int add_header_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo);\n \n #endif /* COMMIT_H */\n-- \n2.41.0\n\n"},{"id":"482392","messageId":"20230927195537.1682-12-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 12/30] tag: sign both hashes","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:19Z","receivedAt":"2023-09-27T19:56:20Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nWhen we write a tag the object oid is specific to the hash algorithm.\n\nThis matters when a tag is signed.  The hash transition plan calls for\nsignatures on both the sha1 form and the sha256 form of the object,\nand for both of those signatures to live in the tag object.\n\nTo generate tag object with multiple signatures, first compute the\nunsigned form of the tag, and then if the tag is being signed compute\nthe unsigned form of the tag with the compatibilityr hash.  Then\ncompute compute the signatures of both buffers.\n\nOnce the signatures are computed add them to both buffers.  This\nallows computing the compatibility hash in do_sign, saving\nwrite_object_file the expense of recomputing the compatibility tag\njust to compute it's hash.\n\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n builtin/tag.c | 45 +++++++++++++++++++++++++++++++++++++++++----\n 1 file changed, 41 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/tag.c b/builtin/tag.c\nindex 3918eacbb57b..8c4bc28952c2 100644\n--- a/builtin/tag.c\n+++ b/builtin/tag.c\n@@ -28,6 +28,7 @@\n #include \"ref-filter.h\"\n #include \"date.h\"\n #include \"write-or-die.h\"\n+#include \"object-file-convert.h\"\n \n static const char * const git_tag_usage[] = {\n \tN_(\"git tag [-a | -s | -u <key-id>] [-f] [-m <msg> | -F <file>] [-e]\\n\"\n@@ -174,9 +175,43 @@ static int verify_tag(const char *name, const char *ref UNUSED,\n \treturn 0;\n }\n \n-static int do_sign(struct strbuf *buffer)\n+static int do_sign(struct strbuf *buffer, struct object_id **compat_oid,\n+\t\t   struct object_id *compat_oid_buf)\n {\n-\treturn sign_buffer(buffer, buffer, get_signing_key());\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n+\tstruct strbuf sig = STRBUF_INIT, compat_sig = STRBUF_INIT;\n+\tstruct strbuf compat_buf = STRBUF_INIT;\n+\tconst char *keyid = get_signing_key();\n+\tint ret = -1;\n+\n+\tif (sign_buffer(buffer, &sig, keyid))\n+\t\treturn -1;\n+\n+\tif (compat) {\n+\t\tconst struct git_hash_algo *algo = the_repository->hash_algo;\n+\n+\t\tif (convert_object_file(&compat_buf, algo, compat,\n+\t\t\t\t\tbuffer->buf, buffer->len, OBJ_TAG, 1))\n+\t\t\tgoto out;\n+\t\tif (sign_buffer(&compat_buf, &compat_sig, keyid))\n+\t\t\tgoto out;\n+\t\tadd_header_signature(&compat_buf, &sig, algo);\n+\t\tstrbuf_addbuf(&compat_buf, &compat_sig);\n+\t\thash_object_file(compat, compat_buf.buf, compat_buf.len,\n+\t\t\t\t OBJ_TAG, compat_oid_buf);\n+\t\t*compat_oid = compat_oid_buf;\n+\t}\n+\n+\tif (compat_sig.len)\n+\t\tadd_header_signature(buffer, &compat_sig, compat);\n+\n+\tstrbuf_addbuf(buffer, &sig);\n+\tret = 0;\n+out:\n+\tstrbuf_release(&sig);\n+\tstrbuf_release(&compat_sig);\n+\tstrbuf_release(&compat_buf);\n+\treturn ret;\n }\n \n static const char tag_template[] =\n@@ -249,9 +284,11 @@ static void write_tag_body(int fd, const struct object_id *oid)\n \n static int build_tag_object(struct strbuf *buf, int sign, struct object_id *result)\n {\n-\tif (sign && do_sign(buf) < 0)\n+\tstruct object_id *compat_oid = NULL, compat_oid_buf;\n+\tif (sign && do_sign(buf, &compat_oid, &compat_oid_buf) < 0)\n \t\treturn error(_(\"unable to sign the tag\"));\n-\tif (write_object_file(buf->buf, buf->len, OBJ_TAG, result) < 0)\n+\tif (write_object_file_flags(buf->buf, buf->len, OBJ_TAG, result,\n+\t\t\t\t    compat_oid, 0) < 0)\n \t\treturn error(_(\"unable to write tag file\"));\n \treturn 0;\n }\n-- \n2.41.0\n\n"},{"id":"482393","messageId":"20230927195537.1682-14-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 14/30] object: Factor out parse_mode out of fast-import and tree-walk into in object.h","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:21Z","receivedAt":"2023-09-27T19:56:22Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nbuiltin/fast-import.c and tree-walk.c have almost identical version of\nget_mode.  The two functions started out the same but have diverged\nslightly.  The version in fast-import changed mode to a uint16_t to\nsave memory.  The version in tree-walk started erroring if no mode was\npresent.\n\nAs far as I can tell both of these changes are valid for both of the\ncallers, so add the both changes and place the common parsing helper\nin object.h\n\nRename the helper from get_mode to parse_mode so it does not\nconflict with another helper named get_mode in diff-no-index.c\n\nThis will be used shortly in a new helper decode_tree_entry_raw\nwhich is used to compute cmpatibility objects as part of\nthe sha256 transition.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/fast-import.c | 18 ++----------------\n object.h              | 18 ++++++++++++++++++\n tree-walk.c           | 22 +++-------------------\n 3 files changed, 23 insertions(+), 35 deletions(-)\n\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 4dbb10aff3da..2c645fcfbe3f 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -1235,20 +1235,6 @@ static void *gfi_unpack_entry(\n \treturn unpack_entry(the_repository, p, oe->idx.offset, &type, sizep);\n }\n \n-static const char *get_mode(const char *str, uint16_t *modep)\n-{\n-\tunsigned char c;\n-\tuint16_t mode = 0;\n-\n-\twhile ((c = *str++) != ' ') {\n-\t\tif (c < '0' || c > '7')\n-\t\t\treturn NULL;\n-\t\tmode = (mode << 3) + (c - '0');\n-\t}\n-\t*modep = mode;\n-\treturn str;\n-}\n-\n static void load_tree(struct tree_entry *root)\n {\n \tstruct object_id *oid = &root->versions[1].oid;\n@@ -1286,7 +1272,7 @@ static void load_tree(struct tree_entry *root)\n \t\tt->entries[t->entry_count++] = e;\n \n \t\te->tree = NULL;\n-\t\tc = get_mode(c, &e->versions[1].mode);\n+\t\tc = parse_mode(c, &e->versions[1].mode);\n \t\tif (!c)\n \t\t\tdie(\"Corrupt mode in %s\", oid_to_hex(oid));\n \t\te->versions[0].mode = e->versions[1].mode;\n@@ -2275,7 +2261,7 @@ static void file_change_m(const char *p, struct branch *b)\n \tstruct object_id oid;\n \tuint16_t mode, inline_data = 0;\n \n-\tp = get_mode(p, &mode);\n+\tp = parse_mode(p, &mode);\n \tif (!p)\n \t\tdie(\"Corrupt mode: %s\", command_buf.buf);\n \tswitch (mode) {\ndiff --git a/object.h b/object.h\nindex 114d45954d08..70c8d4ae63dc 100644\n--- a/object.h\n+++ b/object.h\n@@ -190,6 +190,24 @@ void *create_object(struct repository *r, const struct object_id *oid, void *obj\n \n void *object_as_type(struct object *obj, enum object_type type, int quiet);\n \n+\n+static inline const char *parse_mode(const char *str, uint16_t *modep)\n+{\n+\tunsigned char c;\n+\tunsigned int mode = 0;\n+\n+\tif (*str == ' ')\n+\t\treturn NULL;\n+\n+\twhile ((c = *str++) != ' ') {\n+\t\tif (c < '0' || c > '7')\n+\t\t\treturn NULL;\n+\t\tmode = (mode << 3) + (c - '0');\n+\t}\n+\t*modep = mode;\n+\treturn str;\n+}\n+\n /*\n  * Returns the object, having parsed it to find out what it is.\n  *\ndiff --git a/tree-walk.c b/tree-walk.c\nindex 29ead71be173..3af50a01c2c7 100644\n--- a/tree-walk.c\n+++ b/tree-walk.c\n@@ -10,27 +10,11 @@\n #include \"pathspec.h\"\n #include \"json-writer.h\"\n \n-static const char *get_mode(const char *str, unsigned int *modep)\n-{\n-\tunsigned char c;\n-\tunsigned int mode = 0;\n-\n-\tif (*str == ' ')\n-\t\treturn NULL;\n-\n-\twhile ((c = *str++) != ' ') {\n-\t\tif (c < '0' || c > '7')\n-\t\t\treturn NULL;\n-\t\tmode = (mode << 3) + (c - '0');\n-\t}\n-\t*modep = mode;\n-\treturn str;\n-}\n-\n static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned long size, struct strbuf *err)\n {\n \tconst char *path;\n-\tunsigned int mode, len;\n+\tunsigned int len;\n+\tuint16_t mode;\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \n \tif (size < hashsz + 3 || buf[size - (hashsz + 1)]) {\n@@ -38,7 +22,7 @@ static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned l\n \t\treturn -1;\n \t}\n \n-\tpath = get_mode(buf, &mode);\n+\tpath = parse_mode(buf, &mode);\n \tif (!path) {\n \t\tstrbuf_addstr(err, _(\"malformed mode in tree entry\"));\n \t\treturn -1;\n-- \n2.41.0\n\n"},{"id":"482394","messageId":"20230927195537.1682-13-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 13/30] cache: add a function to read an OID of a specific algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:20Z","receivedAt":"2023-09-27T19:56:26Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nCurrently, we always read a object ID of the current algorithm with\noidread.  However, once we start converting objects, we'll need to\nconsider what happens when we want to read an object ID of a specific\nalgorithm, such as the compatibility algorithm.  To make this easier,\nlet's define oidread_algop, which specifies which algorithm we should\nuse for our object ID, and define oidread in terms of it.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n hash.h | 9 +++++++--\n 1 file changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/hash.h b/hash.h\nindex 615ae0691d07..e064807c1733 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -73,10 +73,15 @@ static inline void oidclr(struct object_id *oid)\n \toid->algo = hash_algo_by_ptr(the_hash_algo);\n }\n \n+static inline void oidread_algop(struct object_id *oid, const unsigned char *hash, const struct git_hash_algo *algop)\n+{\n+\tmemcpy(oid->hash, hash, algop->rawsz);\n+\toid->algo = hash_algo_by_ptr(algop);\n+}\n+\n static inline void oidread(struct object_id *oid, const unsigned char *hash)\n {\n-\tmemcpy(oid->hash, hash, the_hash_algo->rawsz);\n-\toid->algo = hash_algo_by_ptr(the_hash_algo);\n+\toidread_algop(oid, hash, the_hash_algo);\n }\n \n static inline int is_empty_blob_sha1(const unsigned char *sha1)\n-- \n2.41.0\n\n"},{"id":"482395","messageId":"20230927195537.1682-15-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 15/30] object-file-convert: add a function to convert trees between algorithms","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:22Z","receivedAt":"2023-09-27T19:56:28Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nIn the future, we're going to want to provide SHA-256 repositories that\nhave compatibility support for SHA-1 as well.  In order to do so, we'll\nneed to be able to convert tree objects from SHA-256 to SHA-1 by writing\na tree with each SHA-256 object ID mapped to a SHA-1 object ID.\n\nWe implement a function, convert_tree_object, that takes an existing\ntree buffer and writes it to a new strbuf, converting between\nalgorithms.  Let's make this function generic, because while we only\nneed it to convert from the main algorithm to the compatibility\nalgorithm now, we may need to do the other way around in the future,\nsuch as for transport.\n\nWe avoid reusing the code in decode_tree_entry because that code\nnormalizes data, and we don't want that here.  We want to produce a\ncomplete round trip of data, so if, for example, the old entry had a\nwrongly zero-padded mode, we'd want to preserve that when converting to\nensure a stable hash value.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 51 ++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 50 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 4d62ed192bf0..a9e7a208a707 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -1,8 +1,10 @@\n #include \"git-compat-util.h\"\n #include \"gettext.h\"\n #include \"strbuf.h\"\n+#include \"hex.h\"\n #include \"repository.h\"\n #include \"hash-ll.h\"\n+#include \"hash.h\"\n #include \"object.h\"\n #include \"loose.h\"\n #include \"object-file-convert.h\"\n@@ -36,6 +38,51 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \treturn 0;\n }\n \n+static int decode_tree_entry_raw(struct object_id *oid, const char **path,\n+\t\t\t\t size_t *len, const struct git_hash_algo *algo,\n+\t\t\t\t const char *buf, unsigned long size)\n+{\n+\tuint16_t mode;\n+\tconst unsigned hashsz = algo->rawsz;\n+\n+\tif (size < hashsz + 3 || buf[size - (hashsz + 1)]) {\n+\t\treturn -1;\n+\t}\n+\n+\t*path = parse_mode(buf, &mode);\n+\tif (!*path || !**path)\n+\t\treturn -1;\n+\t*len = strlen(*path) + 1;\n+\n+\toidread_algop(oid, (const unsigned char *)*path + *len, algo);\n+\treturn 0;\n+}\n+\n+static int convert_tree_object(struct strbuf *out,\n+\t\t\t       const struct git_hash_algo *from,\n+\t\t\t       const struct git_hash_algo *to,\n+\t\t\t       const char *buffer, size_t size)\n+{\n+\tconst char *p = buffer, *end = buffer + size;\n+\n+\twhile (p < end) {\n+\t\tstruct object_id entry_oid, mapped_oid;\n+\t\tconst char *path = NULL;\n+\t\tsize_t pathlen;\n+\n+\t\tif (decode_tree_entry_raw(&entry_oid, &path, &pathlen, from, p,\n+\t\t\t\t\t  end - p))\n+\t\t\treturn error(_(\"failed to decode tree entry\"));\n+\t\tif (repo_oid_to_algop(the_repository, &entry_oid, to, &mapped_oid))\n+\t\t\treturn error(_(\"failed to map tree entry for %s\"), oid_to_hex(&entry_oid));\n+\t\tstrbuf_add(out, p, path - p);\n+\t\tstrbuf_add(out, path, pathlen);\n+\t\tstrbuf_add(out, mapped_oid.hash, to->rawsz);\n+\t\tp = path + pathlen + from->rawsz;\n+\t}\n+\treturn 0;\n+}\n+\n int convert_object_file(struct strbuf *outbuf,\n \t\t\tconst struct git_hash_algo *from,\n \t\t\tconst struct git_hash_algo *to,\n@@ -50,8 +97,10 @@ int convert_object_file(struct strbuf *outbuf,\n \t\tdie(\"Refusing noop object file conversion\");\n \n \tswitch (type) {\n-\tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n+\t\tret = convert_tree_object(outbuf, from, to, buf, len);\n+\t\tbreak;\n+\tcase OBJ_COMMIT:\n \tcase OBJ_TAG:\n \tdefault:\n \t\t/* Not implemented yet, so fail. */\n-- \n2.41.0\n\n"},{"id":"482396","messageId":"20230927195537.1682-16-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 16/30] object-file-convert: convert tag objects when writing","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:23Z","receivedAt":"2023-09-27T19:56:35Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nWhen writing a tag object in a repository with both SHA-1 and SHA-256,\nwe'll need to convert our commit objects so that we can write the hash\nvalues for both into the repository.  To do so, let's add a function to\nconvert tag objects.\n\nNote that signatures for tag objects in the current algorithm trail the\nmessage, and those for the alternate algorithm are in headers.\nTherefore, we parse the tag object for both a trailing signature and a\nheader and then, when writing the other format, swap the two around.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 52 ++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 51 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex a9e7a208a707..777ae5b58036 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -7,6 +7,8 @@\n #include \"hash.h\"\n #include \"object.h\"\n #include \"loose.h\"\n+#include \"commit.h\"\n+#include \"gpg-interface.h\"\n #include \"object-file-convert.h\"\n \n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n@@ -83,6 +85,52 @@ static int convert_tree_object(struct strbuf *out,\n \treturn 0;\n }\n \n+static int convert_tag_object(struct strbuf *out,\n+\t\t\t      const struct git_hash_algo *from,\n+\t\t\t      const struct git_hash_algo *to,\n+\t\t\t      const char *buffer, size_t size)\n+{\n+\tstruct strbuf payload = STRBUF_INIT, temp = STRBUF_INIT, oursig = STRBUF_INIT, othersig = STRBUF_INIT;\n+\tsize_t payload_size;\n+\tstruct object_id oid, mapped_oid;\n+\tconst char *p;\n+\n+\t/* Add some slop for longer signature header in the new algorithm. */\n+\tstrbuf_grow(out, size + 7);\n+\n+\t/* Is there a signature for our algorithm? */\n+\tpayload_size = parse_signed_buffer(buffer, size);\n+\tstrbuf_add(&payload, buffer, payload_size);\n+\tif (payload_size != size) {\n+\t\t/* Yes, there is. */\n+\t\tstrbuf_add(&oursig, buffer + payload_size, size - payload_size);\n+\t}\n+\t/* Now, is there a signature for the other algorithm? */\n+\tif (parse_buffer_signed_by_header(payload.buf, payload.len, &temp, &othersig, to)) {\n+\t\t/* Yes, there is. */\n+\t\tstrbuf_swap(&payload, &temp);\n+\t\tstrbuf_release(&temp);\n+\t}\n+\n+\t/*\n+\t * Our payload is now in payload and we may have up to two signatrures\n+\t * in oursig and othersig.\n+\t */\n+\tif (strncmp(payload.buf, \"object \", 7) || payload.buf[from->hexsz + 7] != '\\n')\n+\t\treturn error(\"bogus tag object\");\n+\tif (parse_oid_hex_algop(payload.buf + 7, &oid, &p, from) < 0)\n+\t\treturn error(\"bad tag object ID\");\n+\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\treturn error(\"unable to map tree %s in tag object\",\n+\t\t\t     oid_to_hex(&oid));\n+\tstrbuf_addf(out, \"object %s\", oid_to_hex(&mapped_oid));\n+\tstrbuf_add(out, p, payload.len - (p - payload.buf));\n+\tstrbuf_addbuf(out, &othersig);\n+\tif (oursig.len)\n+\t\tadd_header_signature(out, &oursig, from);\n+\treturn 0;\n+}\n+\n int convert_object_file(struct strbuf *outbuf,\n \t\t\tconst struct git_hash_algo *from,\n \t\t\tconst struct git_hash_algo *to,\n@@ -100,8 +148,10 @@ int convert_object_file(struct strbuf *outbuf,\n \tcase OBJ_TREE:\n \t\tret = convert_tree_object(outbuf, from, to, buf, len);\n \t\tbreak;\n-\tcase OBJ_COMMIT:\n \tcase OBJ_TAG:\n+\t\tret = convert_tag_object(outbuf, from, to, buf, len);\n+\t\tbreak;\n+\tcase OBJ_COMMIT:\n \tdefault:\n \t\t/* Not implemented yet, so fail. */\n \t\tret = -1;\n-- \n2.41.0\n\n"},{"id":"482397","messageId":"20230927195537.1682-17-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 17/30] object-file-convert: Don't leak when converting tag objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:24Z","receivedAt":"2023-09-27T19:56:37Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpon close examination I discovered that while brian's code to convert\ntag objects was functionally correct, it leaked memory.\n\nRearrange the code so that all error checking happens before any\nmemory is allocated.\n\nAdd code to release the temporary strbufs the code uses.\n\nThe code pretty much assumes the tag object ends with a newline,\nso add an explict test to verify that is the case.\n\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 45 ++++++++++++++++++++++++-------------------\n 1 file changed, 25 insertions(+), 20 deletions(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 777ae5b58036..822be9d0fdb8 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -90,44 +90,49 @@ static int convert_tag_object(struct strbuf *out,\n \t\t\t      const struct git_hash_algo *to,\n \t\t\t      const char *buffer, size_t size)\n {\n-\tstruct strbuf payload = STRBUF_INIT, temp = STRBUF_INIT, oursig = STRBUF_INIT, othersig = STRBUF_INIT;\n+\tstruct strbuf payload = STRBUF_INIT, oursig = STRBUF_INIT, othersig = STRBUF_INIT;\n+\tconst int entry_len = from->hexsz + 7;\n \tsize_t payload_size;\n \tstruct object_id oid, mapped_oid;\n \tconst char *p;\n \n-\t/* Add some slop for longer signature header in the new algorithm. */\n-\tstrbuf_grow(out, size + 7);\n+\t/* Consume the object line */\n+\tif ((entry_len >= size) ||\n+\t    memcmp(buffer, \"object \", 7) || buffer[entry_len] != '\\n')\n+\t\treturn error(\"bogus tag object\");\n+\tif (parse_oid_hex_algop(buffer + 7, &oid, &p, from) < 0)\n+\t\treturn error(\"bad tag object ID\");\n+\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\treturn error(\"unable to map tree %s in tag object\",\n+\t\t\t     oid_to_hex(&oid));\n+\tsize -= ((p + 1) - buffer);\n+\tbuffer = p + 1;\n \n \t/* Is there a signature for our algorithm? */\n \tpayload_size = parse_signed_buffer(buffer, size);\n-\tstrbuf_add(&payload, buffer, payload_size);\n \tif (payload_size != size) {\n \t\t/* Yes, there is. */\n \t\tstrbuf_add(&oursig, buffer + payload_size, size - payload_size);\n \t}\n-\t/* Now, is there a signature for the other algorithm? */\n-\tif (parse_buffer_signed_by_header(payload.buf, payload.len, &temp, &othersig, to)) {\n-\t\t/* Yes, there is. */\n-\t\tstrbuf_swap(&payload, &temp);\n-\t\tstrbuf_release(&temp);\n-\t}\n \n+\t/* Now, is there a signature for the other algorithm? */\n+\tparse_buffer_signed_by_header(buffer, payload_size, &payload, &othersig, to);\n \t/*\n \t * Our payload is now in payload and we may have up to two signatrures\n \t * in oursig and othersig.\n \t */\n-\tif (strncmp(payload.buf, \"object \", 7) || payload.buf[from->hexsz + 7] != '\\n')\n-\t\treturn error(\"bogus tag object\");\n-\tif (parse_oid_hex_algop(payload.buf + 7, &oid, &p, from) < 0)\n-\t\treturn error(\"bad tag object ID\");\n-\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n-\t\treturn error(\"unable to map tree %s in tag object\",\n-\t\t\t     oid_to_hex(&oid));\n-\tstrbuf_addf(out, \"object %s\", oid_to_hex(&mapped_oid));\n-\tstrbuf_add(out, p, payload.len - (p - payload.buf));\n-\tstrbuf_addbuf(out, &othersig);\n+\n+\t/* Add some slop for longer signature header in the new algorithm. */\n+\tstrbuf_grow(out, (7 + to->hexsz + 1) + size + 7);\n+\tstrbuf_addf(out, \"object %s\\n\", oid_to_hex(&mapped_oid));\n+\tstrbuf_addbuf(out, &payload);\n \tif (oursig.len)\n \t\tadd_header_signature(out, &oursig, from);\n+\tstrbuf_addbuf(out, &othersig);\n+\n+\tstrbuf_release(&payload);\n+\tstrbuf_release(&othersig);\n+\tstrbuf_release(&oursig);\n \treturn 0;\n }\n \n-- \n2.41.0\n\n"},{"id":"482398","messageId":"20230927195537.1682-18-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 18/30] object-file-convert: convert commit objects when writing","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:25Z","receivedAt":"2023-09-27T19:56:38Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nWhen writing a commit object in a repository with both SHA-1 and\nSHA-256, we'll need to convert our commit objects so that we can write\nthe hash values for both into the repository.  To do so, let's add a\nfunction to convert commit objects.\n\nRead the commit object and map the tree value and any of the parent\nvalues, and copy the rest of the commit through unmodified.  Note that\nwe don't need to modify the signature headers, because they are the\nsame under both algorithms.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 46 ++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 45 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 822be9d0fdb8..f53e14e5a170 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -136,6 +136,48 @@ static int convert_tag_object(struct strbuf *out,\n \treturn 0;\n }\n \n+static int convert_commit_object(struct strbuf *out,\n+\t\t\t\t const struct git_hash_algo *from,\n+\t\t\t\t const struct git_hash_algo *to,\n+\t\t\t\t const char *buffer, size_t size)\n+{\n+\tconst char *tail = buffer;\n+\tconst char *bufptr = buffer;\n+\tconst int tree_entry_len = from->hexsz + 5;\n+\tconst int parent_entry_len = from->hexsz + 7;\n+\tstruct object_id oid, mapped_oid;\n+\tconst char *p;\n+\n+\ttail += size;\n+\tif (tail <= bufptr + tree_entry_len + 1 || memcmp(bufptr, \"tree \", 5) ||\n+\t\t\tbufptr[tree_entry_len] != '\\n')\n+\t\treturn error(\"bogus commit object\");\n+\tif (parse_oid_hex_algop(bufptr + 5, &oid, &p, from) < 0)\n+\t\treturn error(\"bad tree pointer\");\n+\n+\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\treturn error(\"unable to map tree %s in commit object\",\n+\t\t\t     oid_to_hex(&oid));\n+\tstrbuf_addf(out, \"tree %s\\n\", oid_to_hex(&mapped_oid));\n+\tbufptr = p + 1;\n+\n+\twhile (bufptr + parent_entry_len < tail && !memcmp(bufptr, \"parent \", 7)) {\n+\t\tif (tail <= bufptr + parent_entry_len + 1 ||\n+\t\t    parse_oid_hex_algop(bufptr + 7, &oid, &p, from) ||\n+\t\t    *p != '\\n')\n+\t\t\treturn error(\"bad parents in commit\");\n+\n+\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\t\treturn error(\"unable to map parent %s in commit object\",\n+\t\t\t\t     oid_to_hex(&oid));\n+\n+\t\tstrbuf_addf(out, \"parent %s\\n\", oid_to_hex(&mapped_oid));\n+\t\tbufptr = p + 1;\n+\t}\n+\tstrbuf_add(out, bufptr, tail - bufptr);\n+\treturn 0;\n+}\n+\n int convert_object_file(struct strbuf *outbuf,\n \t\t\tconst struct git_hash_algo *from,\n \t\t\tconst struct git_hash_algo *to,\n@@ -150,13 +192,15 @@ int convert_object_file(struct strbuf *outbuf,\n \t\tdie(\"Refusing noop object file conversion\");\n \n \tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\t\tret = convert_commit_object(outbuf, from, to, buf, len);\n+\t\tbreak;\n \tcase OBJ_TREE:\n \t\tret = convert_tree_object(outbuf, from, to, buf, len);\n \t\tbreak;\n \tcase OBJ_TAG:\n \t\tret = convert_tag_object(outbuf, from, to, buf, len);\n \t\tbreak;\n-\tcase OBJ_COMMIT:\n \tdefault:\n \t\t/* Not implemented yet, so fail. */\n \t\tret = -1;\n-- \n2.41.0\n\n"},{"id":"482399","messageId":"20230927195537.1682-19-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 19/30] object-file-convert: Convert commits that embed signed tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:26Z","receivedAt":"2023-09-27T19:56:40Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nAs mentioned in the hash function transition plan commit mergetag\nlines need to be handled.  The commit mergetag lines embed an entire\ntag object in a commit object.\n\nKeep the implementation sane if not fast by unembedding the tag\nobject, converting the tag object, and embedding the new tag object,\nin the new commit object.\n\nIn the long run I don't expect any other approach is maintainable, as\ntag objects may be extended in ways that require additional\ntranslation.\n\nTo keep the implementation of convert_commit_object maintainable I\nhave modified convert_commit_object to process the lines in any order,\nand to fail on unknown lines.  We can't know ahead of time if a new\nline might embed something that needs translation or not so it is\nbetter to fail and require the code to be updated instead of silently\nmistranslating objects.\n\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 104 +++++++++++++++++++++++++++++++++---------\n 1 file changed, 82 insertions(+), 22 deletions(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex f53e14e5a170..8ede9889a7ab 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -146,35 +146,95 @@ static int convert_commit_object(struct strbuf *out,\n \tconst int tree_entry_len = from->hexsz + 5;\n \tconst int parent_entry_len = from->hexsz + 7;\n \tstruct object_id oid, mapped_oid;\n-\tconst char *p;\n+\tconst char *p, *eol;\n \n \ttail += size;\n-\tif (tail <= bufptr + tree_entry_len + 1 || memcmp(bufptr, \"tree \", 5) ||\n-\t\t\tbufptr[tree_entry_len] != '\\n')\n-\t\treturn error(\"bogus commit object\");\n-\tif (parse_oid_hex_algop(bufptr + 5, &oid, &p, from) < 0)\n-\t\treturn error(\"bad tree pointer\");\n \n-\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n-\t\treturn error(\"unable to map tree %s in commit object\",\n-\t\t\t     oid_to_hex(&oid));\n-\tstrbuf_addf(out, \"tree %s\\n\", oid_to_hex(&mapped_oid));\n-\tbufptr = p + 1;\n+\twhile ((bufptr < tail) && (*bufptr != '\\n')) {\n+\t\teol = memchr(bufptr, '\\n', tail - bufptr);\n+\t\tif (!eol)\n+\t\t\treturn error(_(\"bad %s in commit\"), \"line\");\n+\n+\t\tif (((bufptr + 5) < eol) && !memcmp(bufptr, \"tree \", 5))\n+\t\t{\n+\t\t\tif (((bufptr + tree_entry_len) != eol) ||\n+\t\t\t    parse_oid_hex_algop(bufptr + 5, &oid, &p, from) ||\n+\t\t\t    (p != eol))\n+\t\t\t\treturn error(_(\"bad %s in commit\"), \"tree\");\n+\n+\t\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\t\t\treturn error(_(\"unable to map %s %s in commit object\"),\n+\t\t\t\t\t     \"tree\", oid_to_hex(&oid));\n+\t\t\tstrbuf_addf(out, \"tree %s\\n\", oid_to_hex(&mapped_oid));\n+\t\t}\n+\t\telse if (((bufptr + 7) < eol) && !memcmp(bufptr, \"parent \", 7))\n+\t\t{\n+\t\t\tif (((bufptr + parent_entry_len) != eol) ||\n+\t\t\t    parse_oid_hex_algop(bufptr + 7, &oid, &p, from) ||\n+\t\t\t    (p != eol))\n+\t\t\t\treturn error(_(\"bad %s in commit\"), \"parent\");\n \n-\twhile (bufptr + parent_entry_len < tail && !memcmp(bufptr, \"parent \", 7)) {\n-\t\tif (tail <= bufptr + parent_entry_len + 1 ||\n-\t\t    parse_oid_hex_algop(bufptr + 7, &oid, &p, from) ||\n-\t\t    *p != '\\n')\n-\t\t\treturn error(\"bad parents in commit\");\n+\t\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\t\t\treturn error(_(\"unable to map %s %s in commit object\"),\n+\t\t\t\t\t     \"parent\", oid_to_hex(&oid));\n \n-\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n-\t\t\treturn error(\"unable to map parent %s in commit object\",\n-\t\t\t\t     oid_to_hex(&oid));\n+\t\t\tstrbuf_addf(out, \"parent %s\\n\", oid_to_hex(&mapped_oid));\n+\t\t}\n+\t\telse if (((bufptr + 9) < eol) && !memcmp(bufptr, \"mergetag \", 9))\n+\t\t{\n+\t\t\tstruct strbuf tag = STRBUF_INIT, new_tag = STRBUF_INIT;\n \n-\t\tstrbuf_addf(out, \"parent %s\\n\", oid_to_hex(&mapped_oid));\n-\t\tbufptr = p + 1;\n+\t\t\t/* Recover the tag object from the mergetag */\n+\t\t\tstrbuf_add(&tag, bufptr + 9, (eol - (bufptr + 9)) + 1);\n+\n+\t\t\tbufptr = eol + 1;\n+\t\t\twhile ((bufptr < tail) && (*bufptr == ' ')) {\n+\t\t\t\teol = memchr(bufptr, '\\n', tail - bufptr);\n+\t\t\t\tif (!eol) {\n+\t\t\t\t\tstrbuf_release(&tag);\n+\t\t\t\t\treturn error(_(\"bad %s in commit\"), \"mergetag continuation\");\n+\t\t\t\t}\n+\t\t\t\tstrbuf_add(&tag, bufptr + 1, (eol - (bufptr + 1)) + 1);\n+\t\t\t\tbufptr = eol + 1;\n+\t\t\t}\n+\n+\t\t\t/* Compute the new tag object */\n+\t\t\tif (convert_tag_object(&new_tag, from, to, tag.buf, tag.len)) {\n+\t\t\t\tstrbuf_release(&tag);\n+\t\t\t\tstrbuf_release(&new_tag);\n+\t\t\t\treturn -1;\n+\t\t\t}\n+\n+\t\t\t/* Write the new mergetag */\n+\t\t\tstrbuf_addstr(out, \"mergetag\");\n+\t\t\tstrbuf_add_lines(out, \" \", new_tag.buf, new_tag.len);\n+\t\t\tstrbuf_release(&tag);\n+\t\t\tstrbuf_release(&new_tag);\n+\t\t}\n+\t\telse if (((bufptr + 7) < tail) && !memcmp(bufptr, \"author \", 7))\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\telse if (((bufptr + 10) < tail) && !memcmp(bufptr, \"committer \", 10))\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\telse if (((bufptr + 9) < tail) && !memcmp(bufptr, \"encoding \", 9))\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\telse if (((bufptr + 6) < tail) && !memcmp(bufptr, \"gpgsig\", 6))\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\telse {\n+\t\t\t/* Unknown line fail it might embed an oid */\n+\t\t\treturn -1;\n+\t\t}\n+\t\t/* Consume any trailing continuation lines */\n+\t\tbufptr = eol + 1;\n+\t\twhile ((bufptr < tail) && (*bufptr == ' ')) {\n+\t\t\teol = memchr(bufptr, '\\n', tail - bufptr);\n+\t\t\tif (!eol)\n+\t\t\t\treturn error(_(\"bad %s in commit\"), \"continuation\");\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\t\tbufptr = eol + 1;\n+\t\t}\n \t}\n-\tstrbuf_add(out, bufptr, tail - bufptr);\n+\tif (bufptr < tail)\n+\t\tstrbuf_add(out, bufptr, tail - bufptr);\n \treturn 0;\n }\n \n-- \n2.41.0\n\n"},{"id":"482400","messageId":"20230927195537.1682-20-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 20/30] object-file: Update object_info_extended to reencode objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:27Z","receivedAt":"2023-09-27T19:56:41Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\noid_object_info_extended is updated to detect an oid encoding that\ndoes not match the current repository, use repo_oid_to_algop to find\nthe correspoding oid in the current repository and to return the data\nfor the oid.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n object-file.c | 91 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 91 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex d66d11890696..1601d624c9fd 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1662,10 +1662,101 @@ static int do_oid_object_info_extended(struct repository *r,\n \treturn 0;\n }\n \n+static int oid_object_info_convert(struct repository *r,\n+\t\t\t\t   const struct object_id *input_oid,\n+\t\t\t\t   struct object_info *input_oi, unsigned flags)\n+{\n+\tconst struct git_hash_algo *input_algo = &hash_algos[input_oid->algo];\n+\tint do_die = flags & OBJECT_INFO_DIE_IF_CORRUPT;\n+\tstruct strbuf type_name = STRBUF_INIT;\n+\tstruct object_id oid, delta_base_oid;\n+\tstruct object_info new_oi, *oi;\n+\tunsigned long size;\n+\tvoid *content;\n+\tint ret;\n+\n+\tif (repo_oid_to_algop(r, input_oid, the_hash_algo, &oid)) {\n+\t\tif (do_die)\n+\t\t\tdie(_(\"missing mapping of %s to %s\"),\n+\t\t\t    oid_to_hex(input_oid), the_hash_algo->name);\n+\t\treturn -1;\n+\t}\n+\n+\t/* Is new_oi needed? */\n+\toi = input_oi;\n+\tif (input_oi && (input_oi->delta_base_oid || input_oi->sizep ||\n+\t\t\t input_oi->contentp)) {\n+\t\tnew_oi = *input_oi;\n+\t\t/* Does delta_base_oid need to be converted? */\n+\t\tif (input_oi->delta_base_oid)\n+\t\t\tnew_oi.delta_base_oid = &delta_base_oid;\n+\t\t/* Will the attributes differ when converted? */\n+\t\tif (input_oi->sizep || input_oi->contentp) {\n+\t\t\tnew_oi.contentp = &content;\n+\t\t\tnew_oi.sizep = &size;\n+\t\t\tnew_oi.type_name = &type_name;\n+\t\t}\n+\t\toi = &new_oi;\n+\t}\n+\n+\tret = oid_object_info_extended(r, &oid, oi, flags);\n+\tif (ret)\n+\t\treturn -1;\n+\tif (oi == input_oi)\n+\t\treturn ret;\n+\n+\tif (new_oi.contentp) {\n+\t\tstruct strbuf outbuf = STRBUF_INIT;\n+\t\tenum object_type type;\n+\n+\t\ttype = type_from_string_gently(type_name.buf, type_name.len,\n+\t\t\t\t\t       !do_die);\n+\t\tif (type == -1)\n+\t\t\treturn -1;\n+\t\tif (type != OBJ_BLOB) {\n+\t\t\tret = convert_object_file(&outbuf,\n+\t\t\t\t\t\t  the_hash_algo, input_algo,\n+\t\t\t\t\t\t  content, size, type, !do_die);\n+\t\t\tif (ret == -1)\n+\t\t\t\treturn -1;\n+\t\t\tfree(content);\n+\t\t\tsize = outbuf.len;\n+\t\t\tcontent = strbuf_detach(&outbuf, NULL);\n+\t\t}\n+\t\tif (input_oi->sizep)\n+\t\t\t*input_oi->sizep = size;\n+\t\tif (input_oi->contentp)\n+\t\t\t*input_oi->contentp = content;\n+\t\telse\n+\t\t\tfree(content);\n+\t\tif (input_oi->type_name)\n+\t\t\t*input_oi->type_name = type_name;\n+\t\telse\n+\t\t\tstrbuf_release(&type_name);\n+\t}\n+\tif (new_oi.delta_base_oid == &delta_base_oid) {\n+\t\tif (repo_oid_to_algop(r, &delta_base_oid, input_algo,\n+\t\t\t\t input_oi->delta_base_oid)) {\n+\t\t\tif (do_die)\n+\t\t\t\tdie(_(\"missing mapping of %s to %s\"),\n+\t\t\t\t    oid_to_hex(&delta_base_oid),\n+\t\t\t\t    input_algo->name);\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\tinput_oi->whence = new_oi.whence;\n+\tinput_oi->u = new_oi.u;\n+\treturn ret;\n+}\n+\n int oid_object_info_extended(struct repository *r, const struct object_id *oid,\n \t\t\t     struct object_info *oi, unsigned flags)\n {\n \tint ret;\n+\n+\tif (oid->algo && (hash_algo_by_ptr(r->hash_algo) != oid->algo))\n+\t\treturn oid_object_info_convert(r, oid, oi, flags);\n+\n \tobj_read_lock();\n \tret = do_oid_object_info_extended(r, oid, oi, flags);\n \tobj_read_unlock();\n-- \n2.41.0\n\n"},{"id":"482401","messageId":"20230927195537.1682-22-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 22/30] rev-parse: Add an --output-object-format parameter","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:29Z","receivedAt":"2023-09-27T19:56:43Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nThe new --output-object-format parameter returns the oid in the\nspecified format.\n\nThis is a generally useful plumbing facility.  It is useful for writing\ntest cases and for directly querying the translation maps.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Documentation/git-rev-parse.txt | 12 ++++++++++++\n builtin/rev-parse.c             | 23 +++++++++++++++++++++++\n 2 files changed, 35 insertions(+)\n\ndiff --git a/Documentation/git-rev-parse.txt b/Documentation/git-rev-parse.txt\nindex f26a7591e373..f0f9021f2a5a 100644\n--- a/Documentation/git-rev-parse.txt\n+++ b/Documentation/git-rev-parse.txt\n@@ -159,6 +159,18 @@ for another option.\n \tunfortunately named tag \"master\"), and show them as full\n \trefnames (e.g. \"refs/heads/master\").\n \n+--output-object-format=(sha1|sha256|storage)::\n+\n+\tAllow oids to be input from any object format that the current\n+\trepository supports.\n+\n+\tSpecifying \"sha1\" translates if necessary and returns a sha1 oid.\n+\n+\tSpecifying \"sha256\" translates if necessary and returns a sha256 oid.\n+\n+\tSpecifying \"storage\" translates if necessary and returns an oid in\n+\tencoded in the storage hash algorithm.\n+\n Options for Objects\n ~~~~~~~~~~~~~~~~~~~\n \ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex 43e96765400c..0ef3e658cc5b 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -25,6 +25,7 @@\n #include \"submodule.h\"\n #include \"commit-reach.h\"\n #include \"shallow.h\"\n+#include \"object-file-convert.h\"\n \n #define DO_REVS\t\t1\n #define DO_NOREV\t2\n@@ -675,6 +676,8 @@ static void print_path(const char *path, const char *prefix, enum format_type fo\n int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n {\n \tint i, as_is = 0, verify = 0, quiet = 0, revs_count = 0, type = 0;\n+\tconst struct git_hash_algo *output_algo = NULL;\n+\tconst struct git_hash_algo *compat = NULL;\n \tint did_repo_setup = 0;\n \tint has_dashdash = 0;\n \tint output_prefix = 0;\n@@ -746,6 +749,7 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \n \t\t\tprepare_repo_settings(the_repository);\n \t\t\tthe_repository->settings.command_requires_full_index = 0;\n+\t\t\tcompat = the_repository->compat_hash_algo;\n \t\t}\n \n \t\tif (!strcmp(arg, \"--\")) {\n@@ -833,6 +837,22 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t\t\tflags |= GET_OID_QUIETLY;\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (opt_with_value(arg, \"--output-object-format\", &arg)) {\n+\t\t\t\tif (!arg)\n+\t\t\t\t\tdie(_(\"no object format specified\"));\n+\t\t\t\tif (!strcmp(arg, the_hash_algo->name) ||\n+\t\t\t\t    !strcmp(arg, \"storage\")) {\n+\t\t\t\t\tflags |= GET_OID_HASH_ANY;\n+\t\t\t\t\toutput_algo = the_hash_algo;\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t\telse if (compat && !strcmp(arg, compat->name)) {\n+\t\t\t\t\tflags |= GET_OID_HASH_ANY;\n+\t\t\t\t\toutput_algo = compat;\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t\telse die(_(\"unsupported object format: %s\"), arg);\n+\t\t\t}\n \t\t\tif (opt_with_value(arg, \"--short\", &arg)) {\n \t\t\t\tfilter &= ~(DO_FLAGS|DO_NOREV);\n \t\t\t\tverify = 1;\n@@ -1083,6 +1103,9 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t}\n \t\tif (!get_oid_with_context(the_repository, name,\n \t\t\t\t\t  flags, &oid, &unused)) {\n+\t\t\tif (output_algo)\n+\t\t\t\trepo_oid_to_algop(the_repository, &oid,\n+\t\t\t\t\t\t  output_algo, &oid);\n \t\t\tif (verify)\n \t\t\t\trevs_count++;\n \t\t\telse\n-- \n2.41.0\n\n"},{"id":"482402","messageId":"20230927195537.1682-21-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 21/30] repository: Implement extensions.compatObjectFormat","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:28Z","receivedAt":"2023-09-27T19:56:49Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nAdd a configuration option to enable updating and reading from\ncompatibility hash maps when git accesses the reposotiry.\n\nCall the helper function repo_set_compat_hash_algo with the value\nthat compatObjectFormat is set to.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Documentation/config/extensions.txt | 12 ++++++++++++\n repository.c                        |  2 +-\n setup.c                             | 23 +++++++++++++++++++++--\n setup.h                             |  1 +\n 4 files changed, 35 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/config/extensions.txt b/Documentation/config/extensions.txt\nindex bccaec7a9636..9f72e6d9f4f1 100644\n--- a/Documentation/config/extensions.txt\n+++ b/Documentation/config/extensions.txt\n@@ -7,6 +7,18 @@ Note that this setting should only be set by linkgit:git-init[1] or\n linkgit:git-clone[1].  Trying to change it after initialization will not\n work and will produce hard-to-diagnose issues.\n \n+extensions.compatObjectFormat::\n+\n+\tSpecify a compatitbility hash algorithm to use.  The acceptable values\n+\tare `sha1` and `sha256`.  The value specified must be different from the\n+\tvalue of extensions.objectFormat.  This allows client level\n+\tinteroperability between git repositories whose objectFormat matches\n+\tthis compatObjectFormat.  In particular when fully implemented the\n+\tpushes and pulls from a repository in whose objectFormat matches\n+\tcompatObjectFormat.  As well as being able to use oids encoded in\n+\tcompatObjectFormat in addition to oids encoded with objectFormat to\n+\tlocally specify objects.\n+\n extensions.worktreeConfig::\n \tIf enabled, then worktrees will load config settings from the\n \t`$GIT_DIR/config.worktree` file in addition to the\ndiff --git a/repository.c b/repository.c\nindex 6214f61cf4e7..9d91536b613b 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -194,7 +194,7 @@ int repo_init(struct repository *repo,\n \t\tgoto error;\n \n \trepo_set_hash_algo(repo, format.hash_algo);\n-\trepo_set_compat_hash_algo(repo, GIT_HASH_UNKNOWN);\n+\trepo_set_compat_hash_algo(repo, format.compat_hash_algo);\n \trepo->repository_format_worktree_config = format.worktree_config;\n \n \t/* take ownership of format.partial_clone */\ndiff --git a/setup.c b/setup.c\nindex deb5a33fe9e1..87b40472dbc5 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -598,6 +598,25 @@ static enum extension_result handle_extension(const char *var,\n \t\t}\n \t\tdata->hash_algo = format;\n \t\treturn EXTENSION_OK;\n+\t} else if (!strcmp(ext, \"compatobjectformat\")) {\n+\t\tstruct string_list_item *item;\n+\t\tint format;\n+\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tformat = hash_algo_by_name(value);\n+\t\tif (format == GIT_HASH_UNKNOWN)\n+\t\t\treturn error(_(\"invalid value for '%s': '%s'\"),\n+\t\t\t\t     \"extensions.compatobjectformat\", value);\n+\t\t/* For now only support compatObjectFormat being specified once. */\n+\t\tfor_each_string_list_item(item, &data->v1_only_extensions) {\n+\t\t\tif (!strcmp(item->string, \"compatobjectformat\"))\n+\t\t\t\treturn error(_(\"'%s' already specified as '%s'\"),\n+\t\t\t\t\t\"extensions.compatobjectformat\",\n+\t\t\t\t\thash_algos[data->compat_hash_algo].name);\n+\t\t}\n+\t\tdata->compat_hash_algo = format;\n+\t\treturn EXTENSION_OK;\n \t}\n \treturn EXTENSION_UNKNOWN;\n }\n@@ -1573,7 +1592,7 @@ const char *setup_git_directory_gently(int *nongit_ok)\n \t\tif (startup_info->have_repository) {\n \t\t\trepo_set_hash_algo(the_repository, repo_fmt.hash_algo);\n \t\t\trepo_set_compat_hash_algo(the_repository,\n-\t\t\t\t\t\t  GIT_HASH_UNKNOWN);\n+\t\t\t\t\t\t  repo_fmt.compat_hash_algo);\n \t\t\tthe_repository->repository_format_worktree_config =\n \t\t\t\trepo_fmt.worktree_config;\n \t\t\t/* take ownership of repo_fmt.partial_clone */\n@@ -1667,7 +1686,7 @@ void check_repository_format(struct repository_format *fmt)\n \tcheck_repository_format_gently(get_git_dir(), fmt, NULL);\n \tstartup_info->have_repository = 1;\n \trepo_set_hash_algo(the_repository, fmt->hash_algo);\n-\trepo_set_compat_hash_algo(the_repository, GIT_HASH_UNKNOWN);\n+\trepo_set_compat_hash_algo(the_repository, fmt->compat_hash_algo);\n \tthe_repository->repository_format_worktree_config =\n \t\tfmt->worktree_config;\n \tthe_repository->repository_format_partial_clone =\ndiff --git a/setup.h b/setup.h\nindex 58fd2605dd26..5d678ceb8caa 100644\n--- a/setup.h\n+++ b/setup.h\n@@ -86,6 +86,7 @@ struct repository_format {\n \tint worktree_config;\n \tint is_bare;\n \tint hash_algo;\n+\tint compat_hash_algo;\n \tint sparse_index;\n \tchar *work_tree;\n \tstruct string_list unknown_extensions;\n-- \n2.41.0\n\n"},{"id":"482403","messageId":"20230927195537.1682-23-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 23/30] builtin/cat-file: Let the oid determine the output algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:30Z","receivedAt":"2023-09-27T19:56:50Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUse GET_OID_HASH_ANY when calling get_oid_with_context.  This\nimplements the semi-obvious behaviour that specifying a sha1 oid shows\nthe output for a sha1 encoded object, and specifying a sha256 oid\nshows the output for a sha256 encoded object.\n\nThis is useful for testing the the conversion of an object to an\nequivalent object encoded with a different hash function.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/cat-file.c | 12 +++++++++---\n 1 file changed, 9 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 694c8538df2f..e615d1f8e0da 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -107,7 +107,10 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct strbuf sb = STRBUF_INIT;\n \tunsigned flags = OBJECT_INFO_LOOKUP_REPLACE;\n-\tunsigned get_oid_flags = GET_OID_RECORD_PATH | GET_OID_ONLY_TO_DIE;\n+\tunsigned get_oid_flags =\n+\t\tGET_OID_RECORD_PATH |\n+\t\tGET_OID_ONLY_TO_DIE |\n+\t\tGET_OID_HASH_ANY;\n \tconst char *path = force_path;\n \tconst int opt_cw = (opt == 'c' || opt == 'w');\n \tif (!path && opt_cw)\n@@ -223,7 +226,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \t\t\t\t\t\t\t\t     &size);\n \t\t\t\tconst char *target;\n \t\t\t\tif (!skip_prefix(buffer, \"object \", &target) ||\n-\t\t\t\t    get_oid_hex(target, &blob_oid))\n+\t\t\t\t    get_oid_hex_algop(target, &blob_oid,\n+\t\t\t\t\t\t      &hash_algos[oid.algo]))\n \t\t\t\t\tdie(\"%s not a valid tag\", oid_to_hex(&oid));\n \t\t\t\tfree(buffer);\n \t\t\t} else\n@@ -512,7 +516,9 @@ static void batch_one_object(const char *obj_name,\n \t\t\t     struct expand_data *data)\n {\n \tstruct object_context ctx;\n-\tint flags = opt->follow_symlinks ? GET_OID_FOLLOW_SYMLINKS : 0;\n+\tint flags =\n+\t\tGET_OID_HASH_ANY |\n+\t\t(opt->follow_symlinks ? GET_OID_FOLLOW_SYMLINKS : 0);\n \tenum get_oid_result result;\n \n \tresult = get_oid_with_context(the_repository, obj_name,\n-- \n2.41.0\n\n"},{"id":"482404","messageId":"20230927195537.1682-25-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 25/30] object-file: Handle compat objects in check_object_signature","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:32Z","receivedAt":"2023-09-27T19:56:58Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpdate check_object_signature to find the hash algorithm the exising\nsignature uses, and to use the same hash algorithm when recomputing it\nto check the signature is valid.\n\nThis will be useful when teaching git ls-tree to display objects\nencoded with the compat hash algorithm.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n object-file.c | 4 +++-\n 1 file changed, 3 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 1601d624c9fd..df49d2239f24 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1094,9 +1094,11 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n \t\t\t   void *buf, unsigned long size,\n \t\t\t   enum object_type type)\n {\n+\tconst struct git_hash_algo *algo =\n+\t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n \tstruct object_id real_oid;\n \n-\thash_object_file(r->hash_algo, buf, size, type, &real_oid);\n+\thash_object_file(algo, buf, size, type, &real_oid);\n \n \treturn !oideq(oid, &real_oid) ? -1 : 0;\n }\n-- \n2.41.0\n\n"},{"id":"482405","messageId":"20230927195537.1682-26-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 26/30] builtin/ls-tree: Let the oid determine the output algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:33Z","receivedAt":"2023-09-27T19:56:59Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpdate cmd_ls_tree to call get_oid_with_context and pass\nGET_OID_HASH_ANY instead of calling the simpler repo_get_oid.\n\nThis implments in ls-tree the behavior that asking to display a sha1\nhash displays the corrresponding sha1 encoded object and asking to\ndisplay a sha256 hash displayes the corresponding sha256 encoded\nobject.\n\nThis is useful for testing the conversion of an object to an\nequivlanet object encoded with a different hash function.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/ls-tree.c | 5 ++++-\n 1 file changed, 4 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/ls-tree.c b/builtin/ls-tree.c\nindex f558db5f3b80..71281ab705b6 100644\n--- a/builtin/ls-tree.c\n+++ b/builtin/ls-tree.c\n@@ -376,6 +376,7 @@ int cmd_ls_tree(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \tstruct ls_tree_cmdmode_to_fmt *m2f = ls_tree_cmdmode_format;\n+\tstruct object_context obj_context;\n \tint ret;\n \n \tgit_config(git_default_config, NULL);\n@@ -407,7 +408,9 @@ int cmd_ls_tree(int argc, const char **argv, const char *prefix)\n \t\t\tls_tree_usage, ls_tree_options);\n \tif (argc < 1)\n \t\tusage_with_options(ls_tree_usage, ls_tree_options);\n-\tif (repo_get_oid(the_repository, argv[0], &oid))\n+\tif (get_oid_with_context(the_repository, argv[0],\n+\t\t\t\t GET_OID_HASH_ANY, &oid,\n+\t\t\t\t &obj_context))\n \t\tdie(\"Not a valid object name %s\", argv[0]);\n \n \t/*\n-- \n2.41.0\n\n"},{"id":"482407","messageId":"20230927195537.1682-24-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 24/30] tree-walk: init_tree_desc take an oid to get the hash algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:31Z","receivedAt":"2023-09-27T19:57:01Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nTo make it possible for git ls-tree to display the tree encoded\nin the hash algorithm of the oid specified to git ls-tree, update\ninit_tree_desc to take as a parameter the oid of the tree object.\n\nUpdate all callers of init_tree_desc and init_tree_desc_gently\nto pass the oid of the tree object.\n\nUse the oid of the tree object to discover the hash algorithm\nof the oid and store that hash algorithm in struct tree_desc.\n\nUse the hash algorithm in decode_tree_entry and\nupdate_tree_entry_internal to handle reading a tree object encoded in\na hash algorithm that differs from the repositories hash algorithm.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n archive.c              |  3 ++-\n builtin/am.c           |  6 +++---\n builtin/checkout.c     |  8 +++++---\n builtin/clone.c        |  2 +-\n builtin/commit.c       |  2 +-\n builtin/grep.c         |  8 ++++----\n builtin/merge.c        |  3 ++-\n builtin/pack-objects.c |  6 ++++--\n builtin/read-tree.c    |  2 +-\n builtin/stash.c        |  5 +++--\n cache-tree.c           |  2 +-\n delta-islands.c        |  2 +-\n diff-lib.c             |  2 +-\n fsck.c                 |  6 ++++--\n http-push.c            |  2 +-\n list-objects.c         |  2 +-\n match-trees.c          |  4 ++--\n merge-ort.c            | 11 ++++++-----\n merge-recursive.c      |  2 +-\n merge.c                |  3 ++-\n pack-bitmap-write.c    |  2 +-\n packfile.c             |  3 ++-\n reflog.c               |  2 +-\n revision.c             |  4 ++--\n tree-walk.c            | 36 +++++++++++++++++++++---------------\n tree-walk.h            |  7 +++++--\n tree.c                 |  2 +-\n walker.c               |  2 +-\n 28 files changed, 80 insertions(+), 59 deletions(-)\n\ndiff --git a/archive.c b/archive.c\nindex ca11db185b15..b10269aee7be 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -339,7 +339,8 @@ int write_archive_entries(struct archiver_args *args,\n \t\topts.src_index = args->repo->index;\n \t\topts.dst_index = args->repo->index;\n \t\topts.fn = oneway_merge;\n-\t\tinit_tree_desc(&t, args->tree->buffer, args->tree->size);\n+\t\tinit_tree_desc(&t, &args->tree->object.oid,\n+\t\t\t       args->tree->buffer, args->tree->size);\n \t\tif (unpack_trees(1, &t, &opts))\n \t\t\treturn -1;\n \t\tgit_attr_set_direction(GIT_ATTR_INDEX);\ndiff --git a/builtin/am.c b/builtin/am.c\nindex 8bde034fae68..4dfd714b910e 100644\n--- a/builtin/am.c\n+++ b/builtin/am.c\n@@ -1991,8 +1991,8 @@ static int fast_forward_to(struct tree *head, struct tree *remote, int reset)\n \topts.reset = reset ? UNPACK_RESET_PROTECT_UNTRACKED : 0;\n \topts.preserve_ignored = 0; /* FIXME: !overwrite_ignore */\n \topts.fn = twoway_merge;\n-\tinit_tree_desc(&t[0], head->buffer, head->size);\n-\tinit_tree_desc(&t[1], remote->buffer, remote->size);\n+\tinit_tree_desc(&t[0], &head->object.oid, head->buffer, head->size);\n+\tinit_tree_desc(&t[1], &remote->object.oid, remote->buffer, remote->size);\n \n \tif (unpack_trees(2, t, &opts)) {\n \t\trollback_lock_file(&lock_file);\n@@ -2026,7 +2026,7 @@ static int merge_tree(struct tree *tree)\n \topts.dst_index = &the_index;\n \topts.merge = 1;\n \topts.fn = oneway_merge;\n-\tinit_tree_desc(&t[0], tree->buffer, tree->size);\n+\tinit_tree_desc(&t[0], &tree->object.oid, tree->buffer, tree->size);\n \n \tif (unpack_trees(1, t, &opts)) {\n \t\trollback_lock_file(&lock_file);\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex f53612f46870..03eff73fd031 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -701,7 +701,7 @@ static int reset_tree(struct tree *tree, const struct checkout_opts *o,\n \t\t\t       info->commit ? &info->commit->object.oid : null_oid(),\n \t\t\t       NULL);\n \tparse_tree(tree);\n-\tinit_tree_desc(&tree_desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&tree_desc, &tree->object.oid, tree->buffer, tree->size);\n \tswitch (unpack_trees(1, &tree_desc, &opts)) {\n \tcase -2:\n \t\t*writeout_error = 1;\n@@ -815,10 +815,12 @@ static int merge_working_tree(const struct checkout_opts *opts,\n \t\t\tdie(_(\"unable to parse commit %s\"),\n \t\t\t\toid_to_hex(old_commit_oid));\n \n-\t\tinit_tree_desc(&trees[0], tree->buffer, tree->size);\n+\t\tinit_tree_desc(&trees[0], &tree->object.oid,\n+\t\t\t       tree->buffer, tree->size);\n \t\tparse_tree(new_tree);\n \t\ttree = new_tree;\n-\t\tinit_tree_desc(&trees[1], tree->buffer, tree->size);\n+\t\tinit_tree_desc(&trees[1], &tree->object.oid,\n+\t\t\t       tree->buffer, tree->size);\n \n \t\tret = unpack_trees(2, trees, &topts);\n \t\tclear_unpack_trees_porcelain(&topts);\ndiff --git a/builtin/clone.c b/builtin/clone.c\nindex c6357af94989..79ceefb93995 100644\n--- a/builtin/clone.c\n+++ b/builtin/clone.c\n@@ -737,7 +737,7 @@ static int checkout(int submodule_progress, int filter_submodules)\n \tif (!tree)\n \t\tdie(_(\"unable to parse commit %s\"), oid_to_hex(&oid));\n \tparse_tree(tree);\n-\tinit_tree_desc(&t, tree->buffer, tree->size);\n+\tinit_tree_desc(&t, &tree->object.oid, tree->buffer, tree->size);\n \tif (unpack_trees(1, &t, &opts) < 0)\n \t\tdie(_(\"unable to checkout working tree\"));\n \ndiff --git a/builtin/commit.c b/builtin/commit.c\nindex 7da5f924484d..537319932b65 100644\n--- a/builtin/commit.c\n+++ b/builtin/commit.c\n@@ -340,7 +340,7 @@ static void create_base_index(const struct commit *current_head)\n \tif (!tree)\n \t\tdie(_(\"failed to unpack HEAD tree object\"));\n \tparse_tree(tree);\n-\tinit_tree_desc(&t, tree->buffer, tree->size);\n+\tinit_tree_desc(&t, &tree->object.oid, tree->buffer, tree->size);\n \tif (unpack_trees(1, &t, &opts))\n \t\texit(128); /* We've already reported the error, finish dying */\n }\ndiff --git a/builtin/grep.c b/builtin/grep.c\nindex 50e712a18479..0c2b8a376f8e 100644\n--- a/builtin/grep.c\n+++ b/builtin/grep.c\n@@ -530,7 +530,7 @@ static int grep_submodule(struct grep_opt *opt,\n \t\tstrbuf_addstr(&base, filename);\n \t\tstrbuf_addch(&base, '/');\n \n-\t\tinit_tree_desc(&tree, data, size);\n+\t\tinit_tree_desc(&tree, oid, data, size);\n \t\thit = grep_tree(&subopt, pathspec, &tree, &base, base.len,\n \t\t\t\tobject_type == OBJ_COMMIT);\n \t\tstrbuf_release(&base);\n@@ -574,7 +574,7 @@ static int grep_cache(struct grep_opt *opt,\n \n \t\t\tdata = repo_read_object_file(the_repository, &ce->oid,\n \t\t\t\t\t\t     &type, &size);\n-\t\t\tinit_tree_desc(&tree, data, size);\n+\t\t\tinit_tree_desc(&tree, &ce->oid, data, size);\n \n \t\t\thit |= grep_tree(opt, pathspec, &tree, &name, 0, 0);\n \t\t\tstrbuf_setlen(&name, name_base_len);\n@@ -670,7 +670,7 @@ static int grep_tree(struct grep_opt *opt, const struct pathspec *pathspec,\n \t\t\t\t    oid_to_hex(&entry.oid));\n \n \t\t\tstrbuf_addch(base, '/');\n-\t\t\tinit_tree_desc(&sub, data, size);\n+\t\t\tinit_tree_desc(&sub, &entry.oid, data, size);\n \t\t\thit |= grep_tree(opt, pathspec, &sub, base, tn_len,\n \t\t\t\t\t check_attr);\n \t\t\tfree(data);\n@@ -714,7 +714,7 @@ static int grep_object(struct grep_opt *opt, const struct pathspec *pathspec,\n \t\t\tstrbuf_add(&base, name, len);\n \t\t\tstrbuf_addch(&base, ':');\n \t\t}\n-\t\tinit_tree_desc(&tree, data, size);\n+\t\tinit_tree_desc(&tree, &obj->oid, data, size);\n \t\thit = grep_tree(opt, pathspec, &tree, &base, base.len,\n \t\t\t\tobj->type == OBJ_COMMIT);\n \t\tstrbuf_release(&base);\ndiff --git a/builtin/merge.c b/builtin/merge.c\nindex de68910177fb..718165d45917 100644\n--- a/builtin/merge.c\n+++ b/builtin/merge.c\n@@ -704,7 +704,8 @@ static int read_tree_trivial(struct object_id *common, struct object_id *head,\n \tcache_tree_free(&the_index.cache_tree);\n \tfor (i = 0; i < nr_trees; i++) {\n \t\tparse_tree(trees[i]);\n-\t\tinit_tree_desc(t+i, trees[i]->buffer, trees[i]->size);\n+\t\tinit_tree_desc(t+i, &trees[i]->object.oid,\n+\t\t\t       trees[i]->buffer, trees[i]->size);\n \t}\n \tif (unpack_trees(nr_trees, t, &opts))\n \t\treturn -1;\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex d2a162d52804..d34902002656 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1756,7 +1756,8 @@ static void add_pbase_object(struct tree_desc *tree,\n \t\t\ttree = pbase_tree_get(&entry.oid);\n \t\t\tif (!tree)\n \t\t\t\treturn;\n-\t\t\tinit_tree_desc(&sub, tree->tree_data, tree->tree_size);\n+\t\t\tinit_tree_desc(&sub, &tree->oid,\n+\t\t\t\t       tree->tree_data, tree->tree_size);\n \n \t\t\tadd_pbase_object(&sub, down, downlen, fullname);\n \t\t\tpbase_tree_put(tree);\n@@ -1816,7 +1817,8 @@ static void add_preferred_base_object(const char *name)\n \t\t}\n \t\telse {\n \t\t\tstruct tree_desc tree;\n-\t\t\tinit_tree_desc(&tree, it->pcache.tree_data, it->pcache.tree_size);\n+\t\t\tinit_tree_desc(&tree, &it->pcache.oid,\n+\t\t\t\t       it->pcache.tree_data, it->pcache.tree_size);\n \t\t\tadd_pbase_object(&tree, name, cmplen, name);\n \t\t}\n \t}\ndiff --git a/builtin/read-tree.c b/builtin/read-tree.c\nindex 1fec702a04fa..24d6d156d3a2 100644\n--- a/builtin/read-tree.c\n+++ b/builtin/read-tree.c\n@@ -264,7 +264,7 @@ int cmd_read_tree(int argc, const char **argv, const char *cmd_prefix)\n \tfor (i = 0; i < nr_trees; i++) {\n \t\tstruct tree *tree = trees[i];\n \t\tparse_tree(tree);\n-\t\tinit_tree_desc(t+i, tree->buffer, tree->size);\n+\t\tinit_tree_desc(t+i, &tree->object.oid, tree->buffer, tree->size);\n \t}\n \tif (unpack_trees(nr_trees, t, &opts))\n \t\treturn 128;\ndiff --git a/builtin/stash.c b/builtin/stash.c\nindex fe64cde9ce30..9ee52af4d28e 100644\n--- a/builtin/stash.c\n+++ b/builtin/stash.c\n@@ -285,7 +285,7 @@ static int reset_tree(struct object_id *i_tree, int update, int reset)\n \tif (parse_tree(tree))\n \t\treturn -1;\n \n-\tinit_tree_desc(t, tree->buffer, tree->size);\n+\tinit_tree_desc(t, &tree->object.oid, tree->buffer, tree->size);\n \n \topts.head_idx = 1;\n \topts.src_index = &the_index;\n@@ -871,7 +871,8 @@ static void diff_include_untracked(const struct stash_info *info, struct diff_op\n \t\ttree[i] = parse_tree_indirect(oid[i]);\n \t\tif (parse_tree(tree[i]) < 0)\n \t\t\tdie(_(\"failed to parse tree\"));\n-\t\tinit_tree_desc(&tree_desc[i], tree[i]->buffer, tree[i]->size);\n+\t\tinit_tree_desc(&tree_desc[i], &tree[i]->object.oid,\n+\t\t\t       tree[i]->buffer, tree[i]->size);\n \t}\n \n \tunpack_tree_opt.head_idx = -1;\ndiff --git a/cache-tree.c b/cache-tree.c\nindex ddc7d3d86959..334973a01cee 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -770,7 +770,7 @@ static void prime_cache_tree_rec(struct repository *r,\n \n \toidcpy(&it->oid, &tree->object.oid);\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \tcnt = 0;\n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (!S_ISDIR(entry.mode))\ndiff --git a/delta-islands.c b/delta-islands.c\nindex 5de5759f3f13..1ff3506b10f2 100644\n--- a/delta-islands.c\n+++ b/delta-islands.c\n@@ -289,7 +289,7 @@ void resolve_tree_islands(struct repository *r,\n \t\tif (!tree || parse_tree(tree) < 0)\n \t\t\tdie(_(\"bad tree object %s\"), oid_to_hex(&ent->idx.oid));\n \n-\t\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\t\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \t\twhile (tree_entry(&desc, &entry)) {\n \t\t\tstruct object *obj;\n \ndiff --git a/diff-lib.c b/diff-lib.c\nindex 6b0c6a7180cc..add323f5628d 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -558,7 +558,7 @@ static int diff_cache(struct rev_info *revs,\n \topts.pathspec = &revs->diffopt.pathspec;\n \topts.pathspec->recursive = 1;\n \n-\tinit_tree_desc(&t, tree->buffer, tree->size);\n+\tinit_tree_desc(&t, &tree->object.oid, tree->buffer, tree->size);\n \treturn unpack_trees(1, &t, &opts);\n }\n \ndiff --git a/fsck.c b/fsck.c\nindex 2b1e348005b7..6b492a48da82 100644\n--- a/fsck.c\n+++ b/fsck.c\n@@ -313,7 +313,8 @@ static int fsck_walk_tree(struct tree *tree, void *data, struct fsck_options *op\n \t\treturn -1;\n \n \tname = fsck_get_object_name(options, &tree->object.oid);\n-\tif (init_tree_desc_gently(&desc, tree->buffer, tree->size, 0))\n+\tif (init_tree_desc_gently(&desc, &tree->object.oid,\n+\t\t\t\t  tree->buffer, tree->size, 0))\n \t\treturn -1;\n \twhile (tree_entry_gently(&desc, &entry)) {\n \t\tstruct object *obj;\n@@ -583,7 +584,8 @@ static int fsck_tree(const struct object_id *tree_oid,\n \tconst char *o_name;\n \tstruct name_stack df_dup_candidates = { NULL };\n \n-\tif (init_tree_desc_gently(&desc, buffer, size, TREE_DESC_RAW_MODES)) {\n+\tif (init_tree_desc_gently(&desc, tree_oid, buffer, size,\n+\t\t\t\t  TREE_DESC_RAW_MODES)) {\n \t\tretval += report(options, tree_oid, OBJ_TREE,\n \t\t\t\t FSCK_MSG_BAD_TREE,\n \t\t\t\t \"cannot be parsed as a tree\");\ndiff --git a/http-push.c b/http-push.c\nindex a704f490fdb2..81c35b5e96f7 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -1308,7 +1308,7 @@ static struct object_list **process_tree(struct tree *tree,\n \tobj->flags |= SEEN;\n \tp = add_one_object(obj, p);\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry))\n \t\tswitch (object_type(entry.mode)) {\ndiff --git a/list-objects.c b/list-objects.c\nindex e60a6cd5b46e..312335c8a7f2 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -97,7 +97,7 @@ static void process_tree_contents(struct traversal_context *ctx,\n \tenum interesting match = ctx->revs->diffopt.pathspec.nr == 0 ?\n \t\tall_entries_interesting : entry_not_interesting;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (match != all_entries_interesting) {\ndiff --git a/match-trees.c b/match-trees.c\nindex 0885ac681cd5..3412b6a1401d 100644\n--- a/match-trees.c\n+++ b/match-trees.c\n@@ -63,7 +63,7 @@ static void *fill_tree_desc_strict(struct tree_desc *desc,\n \t\tdie(\"unable to read tree (%s)\", oid_to_hex(hash));\n \tif (type != OBJ_TREE)\n \t\tdie(\"%s is not a tree\", oid_to_hex(hash));\n-\tinit_tree_desc(desc, buffer, size);\n+\tinit_tree_desc(desc, hash, buffer, size);\n \treturn buffer;\n }\n \n@@ -194,7 +194,7 @@ static int splice_tree(const struct object_id *oid1, const char *prefix,\n \tbuf = repo_read_object_file(the_repository, oid1, &type, &sz);\n \tif (!buf)\n \t\tdie(\"cannot read tree %s\", oid_to_hex(oid1));\n-\tinit_tree_desc(&desc, buf, sz);\n+\tinit_tree_desc(&desc, oid1, buf, sz);\n \n \trewrite_here = NULL;\n \twhile (desc.size) {\ndiff --git a/merge-ort.c b/merge-ort.c\nindex 8631c997002d..3a5729c91e48 100644\n--- a/merge-ort.c\n+++ b/merge-ort.c\n@@ -1679,9 +1679,10 @@ static int collect_merge_info(struct merge_options *opt,\n \tparse_tree(merge_base);\n \tparse_tree(side1);\n \tparse_tree(side2);\n-\tinit_tree_desc(t + 0, merge_base->buffer, merge_base->size);\n-\tinit_tree_desc(t + 1, side1->buffer, side1->size);\n-\tinit_tree_desc(t + 2, side2->buffer, side2->size);\n+\tinit_tree_desc(t + 0, &merge_base->object.oid,\n+\t\t       merge_base->buffer, merge_base->size);\n+\tinit_tree_desc(t + 1, &side1->object.oid, side1->buffer, side1->size);\n+\tinit_tree_desc(t + 2, &side2->object.oid, side2->buffer, side2->size);\n \n \ttrace2_region_enter(\"merge\", \"traverse_trees\", opt->repo);\n \tret = traverse_trees(NULL, 3, t, &info);\n@@ -4400,9 +4401,9 @@ static int checkout(struct merge_options *opt,\n \tunpack_opts.fn = twoway_merge;\n \tunpack_opts.preserve_ignored = 0; /* FIXME: !opts->overwrite_ignore */\n \tparse_tree(prev);\n-\tinit_tree_desc(&trees[0], prev->buffer, prev->size);\n+\tinit_tree_desc(&trees[0], &prev->object.oid, prev->buffer, prev->size);\n \tparse_tree(next);\n-\tinit_tree_desc(&trees[1], next->buffer, next->size);\n+\tinit_tree_desc(&trees[1], &next->object.oid, next->buffer, next->size);\n \n \tret = unpack_trees(2, trees, &unpack_opts);\n \tclear_unpack_trees_porcelain(&unpack_opts);\ndiff --git a/merge-recursive.c b/merge-recursive.c\nindex 6a4081bb0f52..93df9eecdd95 100644\n--- a/merge-recursive.c\n+++ b/merge-recursive.c\n@@ -411,7 +411,7 @@ static inline int merge_detect_rename(struct merge_options *opt)\n static void init_tree_desc_from_tree(struct tree_desc *desc, struct tree *tree)\n {\n \tparse_tree(tree);\n-\tinit_tree_desc(desc, tree->buffer, tree->size);\n+\tinit_tree_desc(desc, &tree->object.oid, tree->buffer, tree->size);\n }\n \n static int unpack_trees_start(struct merge_options *opt,\ndiff --git a/merge.c b/merge.c\nindex b60925459c29..86179c34102d 100644\n--- a/merge.c\n+++ b/merge.c\n@@ -81,7 +81,8 @@ int checkout_fast_forward(struct repository *r,\n \t}\n \tfor (i = 0; i < nr_trees; i++) {\n \t\tparse_tree(trees[i]);\n-\t\tinit_tree_desc(t+i, trees[i]->buffer, trees[i]->size);\n+\t\tinit_tree_desc(t+i, &trees[i]->object.oid,\n+\t\t\t       trees[i]->buffer, trees[i]->size);\n \t}\n \n \tmemset(&opts, 0, sizeof(opts));\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex f6757c3cbf20..9211e08f0127 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -366,7 +366,7 @@ static int fill_bitmap_tree(struct bitmap *bitmap,\n \tif (parse_tree(tree) < 0)\n \t\tdie(\"unable to load tree object %s\",\n \t\t    oid_to_hex(&tree->object.oid));\n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry)) {\n \t\tswitch (object_type(entry.mode)) {\ndiff --git a/packfile.c b/packfile.c\nindex 9cc0a2e37a83..1fae0fcdd9e7 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2250,7 +2250,8 @@ static int add_promisor_object(const struct object_id *oid,\n \t\tstruct tree *tree = (struct tree *)obj;\n \t\tstruct tree_desc desc;\n \t\tstruct name_entry entry;\n-\t\tif (init_tree_desc_gently(&desc, tree->buffer, tree->size, 0))\n+\t\tif (init_tree_desc_gently(&desc, &tree->object.oid,\n+\t\t\t\t\t  tree->buffer, tree->size, 0))\n \t\t\t/*\n \t\t\t * Error messages are given when packs are\n \t\t\t * verified, so do not print any here.\ndiff --git a/reflog.c b/reflog.c\nindex 9ad50e7d93e4..c6992a19268f 100644\n--- a/reflog.c\n+++ b/reflog.c\n@@ -40,7 +40,7 @@ static int tree_is_complete(const struct object_id *oid)\n \t\ttree->buffer = data;\n \t\ttree->size = size;\n \t}\n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \tcomplete = 1;\n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (!repo_has_object_file(the_repository, &entry.oid) ||\ndiff --git a/revision.c b/revision.c\nindex 2f4c53ea207b..a60dfc23a2a5 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -82,7 +82,7 @@ static void mark_tree_contents_uninteresting(struct repository *r,\n \tif (parse_tree_gently(tree, 1) < 0)\n \t\treturn;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \twhile (tree_entry(&desc, &entry)) {\n \t\tswitch (object_type(entry.mode)) {\n \t\tcase OBJ_TREE:\n@@ -189,7 +189,7 @@ static void add_children_by_path(struct repository *r,\n \tif (parse_tree_gently(tree, 1) < 0)\n \t\treturn;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \twhile (tree_entry(&desc, &entry)) {\n \t\tswitch (object_type(entry.mode)) {\n \t\tcase OBJ_TREE:\ndiff --git a/tree-walk.c b/tree-walk.c\nindex 3af50a01c2c7..0b44ec7c75ff 100644\n--- a/tree-walk.c\n+++ b/tree-walk.c\n@@ -15,7 +15,7 @@ static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned l\n \tconst char *path;\n \tunsigned int len;\n \tuint16_t mode;\n-\tconst unsigned hashsz = the_hash_algo->rawsz;\n+\tconst unsigned hashsz = desc->algo->rawsz;\n \n \tif (size < hashsz + 3 || buf[size - (hashsz + 1)]) {\n \t\tstrbuf_addstr(err, _(\"too-short tree object\"));\n@@ -37,15 +37,19 @@ static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned l\n \tdesc->entry.path = path;\n \tdesc->entry.mode = (desc->flags & TREE_DESC_RAW_MODES) ? mode : canon_mode(mode);\n \tdesc->entry.pathlen = len - 1;\n-\toidread(&desc->entry.oid, (const unsigned char *)path + len);\n+\toidread_algop(&desc->entry.oid, (const unsigned char *)path + len,\n+\t\t      desc->algo);\n \n \treturn 0;\n }\n \n-static int init_tree_desc_internal(struct tree_desc *desc, const void *buffer,\n-\t\t\t\t   unsigned long size, struct strbuf *err,\n+static int init_tree_desc_internal(struct tree_desc *desc,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   const void *buffer, unsigned long size,\n+\t\t\t\t   struct strbuf *err,\n \t\t\t\t   enum tree_desc_flags flags)\n {\n+\tdesc->algo = (oid && oid->algo) ? &hash_algos[oid->algo] : the_hash_algo;\n \tdesc->buffer = buffer;\n \tdesc->size = size;\n \tdesc->flags = flags;\n@@ -54,19 +58,21 @@ static int init_tree_desc_internal(struct tree_desc *desc, const void *buffer,\n \treturn 0;\n }\n \n-void init_tree_desc(struct tree_desc *desc, const void *buffer, unsigned long size)\n+void init_tree_desc(struct tree_desc *desc, const struct object_id *tree_oid,\n+\t\t    const void *buffer, unsigned long size)\n {\n \tstruct strbuf err = STRBUF_INIT;\n-\tif (init_tree_desc_internal(desc, buffer, size, &err, 0))\n+\tif (init_tree_desc_internal(desc, tree_oid, buffer, size, &err, 0))\n \t\tdie(\"%s\", err.buf);\n \tstrbuf_release(&err);\n }\n \n-int init_tree_desc_gently(struct tree_desc *desc, const void *buffer, unsigned long size,\n+int init_tree_desc_gently(struct tree_desc *desc, const struct object_id *oid,\n+\t\t\t  const void *buffer, unsigned long size,\n \t\t\t  enum tree_desc_flags flags)\n {\n \tstruct strbuf err = STRBUF_INIT;\n-\tint result = init_tree_desc_internal(desc, buffer, size, &err, flags);\n+\tint result = init_tree_desc_internal(desc, oid, buffer, size, &err, flags);\n \tif (result)\n \t\terror(\"%s\", err.buf);\n \tstrbuf_release(&err);\n@@ -85,7 +91,7 @@ void *fill_tree_descriptor(struct repository *r,\n \t\tif (!buf)\n \t\t\tdie(\"unable to read tree %s\", oid_to_hex(oid));\n \t}\n-\tinit_tree_desc(desc, buf, size);\n+\tinit_tree_desc(desc, oid, buf, size);\n \treturn buf;\n }\n \n@@ -102,7 +108,7 @@ static void entry_extract(struct tree_desc *t, struct name_entry *a)\n static int update_tree_entry_internal(struct tree_desc *desc, struct strbuf *err)\n {\n \tconst void *buf = desc->buffer;\n-\tconst unsigned char *end = (const unsigned char *)desc->entry.path + desc->entry.pathlen + 1 + the_hash_algo->rawsz;\n+\tconst unsigned char *end = (const unsigned char *)desc->entry.path + desc->entry.pathlen + 1 + desc->algo->rawsz;\n \tunsigned long size = desc->size;\n \tunsigned long len = end - (const unsigned char *)buf;\n \n@@ -611,7 +617,7 @@ int get_tree_entry(struct repository *r,\n \t\tretval = -1;\n \t} else {\n \t\tstruct tree_desc t;\n-\t\tinit_tree_desc(&t, tree, size);\n+\t\tinit_tree_desc(&t, tree_oid, tree, size);\n \t\tretval = find_tree_entry(r, &t, name, oid, mode);\n \t}\n \tfree(tree);\n@@ -654,7 +660,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \tstruct tree_desc t;\n \tint follows_remaining = GET_TREE_ENTRY_FOLLOW_SYMLINKS_MAX_LINKS;\n \n-\tinit_tree_desc(&t, NULL, 0UL);\n+\tinit_tree_desc(&t, NULL, NULL, 0UL);\n \tstrbuf_addstr(&namebuf, name);\n \toidcpy(&current_tree_oid, tree_oid);\n \n@@ -690,7 +696,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \t\t\t\tgoto done;\n \n \t\t\t/* descend */\n-\t\t\tinit_tree_desc(&t, tree, size);\n+\t\t\tinit_tree_desc(&t, &current_tree_oid, tree, size);\n \t\t}\n \n \t\t/* Handle symlinks to e.g. a//b by removing leading slashes */\n@@ -724,7 +730,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \t\t\tfree(parent->tree);\n \t\t\tparents_nr--;\n \t\t\tparent = &parents[parents_nr - 1];\n-\t\t\tinit_tree_desc(&t, parent->tree, parent->size);\n+\t\t\tinit_tree_desc(&t, &parent->oid, parent->tree, parent->size);\n \t\t\tstrbuf_remove(&namebuf, 0, remainder ? 3 : 2);\n \t\t\tcontinue;\n \t\t}\n@@ -804,7 +810,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \t\t\tcontents_start = contents;\n \n \t\t\tparent = &parents[parents_nr - 1];\n-\t\t\tinit_tree_desc(&t, parent->tree, parent->size);\n+\t\t\tinit_tree_desc(&t, &parent->oid, parent->tree, parent->size);\n \t\t\tstrbuf_splice(&namebuf, 0, len,\n \t\t\t\t      contents_start, link_len);\n \t\t\tif (remainder)\ndiff --git a/tree-walk.h b/tree-walk.h\nindex 74cdceb3fed2..cf54d01019e9 100644\n--- a/tree-walk.h\n+++ b/tree-walk.h\n@@ -26,6 +26,7 @@ struct name_entry {\n  * A semi-opaque data structure used to maintain the current state of the walk.\n  */\n struct tree_desc {\n+\tconst struct git_hash_algo *algo;\n \t/*\n \t * pointer into the memory representation of the tree. It always\n \t * points at the current entry being visited.\n@@ -85,9 +86,11 @@ int update_tree_entry_gently(struct tree_desc *);\n  * size parameters are assumed to be the same as the buffer and size\n  * members of `struct tree`.\n  */\n-void init_tree_desc(struct tree_desc *desc, const void *buf, unsigned long size);\n+void init_tree_desc(struct tree_desc *desc, const struct object_id *tree_oid,\n+\t\t    const void *buf, unsigned long size);\n \n-int init_tree_desc_gently(struct tree_desc *desc, const void *buf, unsigned long size,\n+int init_tree_desc_gently(struct tree_desc *desc, const struct object_id *oid,\n+\t\t\t  const void *buf, unsigned long size,\n \t\t\t  enum tree_desc_flags flags);\n \n /*\ndiff --git a/tree.c b/tree.c\nindex c745462f968e..44bcf728f10a 100644\n--- a/tree.c\n+++ b/tree.c\n@@ -27,7 +27,7 @@ int read_tree_at(struct repository *r,\n \tif (parse_tree(tree))\n \t\treturn -1;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (retval != all_entries_interesting) {\ndiff --git a/walker.c b/walker.c\nindex 65002a7220ad..c0fd632d921c 100644\n--- a/walker.c\n+++ b/walker.c\n@@ -45,7 +45,7 @@ static int process_tree(struct walker *walker, struct tree *tree)\n \tif (parse_tree(tree))\n \t\treturn -1;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \twhile (tree_entry(&desc, &entry)) {\n \t\tstruct object *obj = NULL;\n \n-- \n2.41.0\n\n"},{"id":"482406","messageId":"20230927195537.1682-27-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 27/30] test-lib: Compute the compatibility hash so tests may use it","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:34Z","receivedAt":"2023-09-27T19:57:02Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nInspired-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n t/test-lib-functions.sh | 17 ++++++++++++++++-\n 1 file changed, 16 insertions(+), 1 deletion(-)\n\ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 2f8868caa171..92b462e2e711 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -1599,7 +1599,16 @@ test_set_hash () {\n \n # Detect the hash algorithm in use.\n test_detect_hash () {\n-\ttest_hash_algo=\"${GIT_TEST_DEFAULT_HASH:-sha1}\"\n+\tcase \"$GIT_TEST_DEFAULT_HASH\" in\n+\t\"sha256\")\n+\t    test_hash_algo=sha256\n+\t    test_compat_hash_algo=sha1\n+\t    ;;\n+\t*)\n+\t    test_hash_algo=sha1\n+\t    test_compat_hash_algo=sha256\n+\t    ;;\n+\tesac\n }\n \n # Load common hash metadata and common placeholder object IDs for use with\n@@ -1651,6 +1660,12 @@ test_oid () {\n \tlocal algo=\"${test_hash_algo}\" &&\n \n \tcase \"$1\" in\n+\t--hash=storage)\n+\t\talgo=\"$test_hash_algo\" &&\n+\t\tshift;;\n+\t--hash=compat)\n+\t\talgo=\"$test_compat_hash_algo\" &&\n+\t\tshift;;\n \t--hash=*)\n \t\talgo=\"${1#--hash=}\" &&\n \t\tshift;;\n-- \n2.41.0\n\n"},{"id":"482408","messageId":"20230927195537.1682-29-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 29/30] t1006: Test oid compatibility with cat-file","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:36Z","receivedAt":"2023-09-27T19:57:04Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpdate the existing tests that are oid based to test that cat-file\nworks correctly with the normal oid and the compat_oid.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n t/t1006-cat-file.sh | 251 +++++++++++++++++++++++++++-----------------\n 1 file changed, 154 insertions(+), 97 deletions(-)\n\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex 9b018b538950..23d3d37283bb 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -236,27 +236,38 @@ hello_size=$(strlen \"$hello_content\")\n hello_oid=$(echo_without_newline \"$hello_content\" | git hash-object --stdin)\n \n test_expect_success \"setup\" '\n+\tgit config core.repositoryformatversion 1 &&\n+\tgit config extensions.objectformat $test_hash_algo &&\n+\tgit config extensions.compatobjectformat $test_compat_hash_algo &&\n \techo_without_newline \"$hello_content\" > hello &&\n \tgit update-index --add hello\n '\n \n-run_tests 'blob' $hello_oid $hello_size \"$hello_content\" \"$hello_content\"\n+run_blob_tests () {\n+    oid=$1\n \n-test_expect_success '--batch-command --buffer with flush for blob info' '\n-\techo \"$hello_oid blob $hello_size\" >expect &&\n-\ttest_write_lines \"info $hello_oid\" \"flush\" |\n+    run_tests 'blob' $oid $hello_size \"$hello_content\" \"$hello_content\"\n+\n+    test_expect_success '--batch-command --buffer with flush for blob info' '\n+\techo \"$oid blob $hello_size\" >expect &&\n+\ttest_write_lines \"info $oid\" \"flush\" |\n \tGIT_TEST_CAT_FILE_NO_FLUSH_ON_EXIT=1 \\\n \tgit cat-file --batch-command --buffer >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success '--batch-command --buffer without flush for blob info' '\n+    test_expect_success '--batch-command --buffer without flush for blob info' '\n \ttouch output &&\n-\ttest_write_lines \"info $hello_oid\" |\n+\ttest_write_lines \"info $oid\" |\n \tGIT_TEST_CAT_FILE_NO_FLUSH_ON_EXIT=1 \\\n \tgit cat-file --batch-command --buffer >>output &&\n \ttest_must_be_empty output\n-'\n+    '\n+}\n+\n+hello_compat_oid=$(git rev-parse --output-object-format=$test_compat_hash_algo $hello_oid)\n+run_blob_tests $hello_oid\n+run_blob_tests $hello_compat_oid\n \n test_expect_success '--batch-check without %(rest) considers whole line' '\n \techo \"$hello_oid blob $hello_size\" >expect &&\n@@ -267,35 +278,58 @@ test_expect_success '--batch-check without %(rest) considers whole line' '\n '\n \n tree_oid=$(git write-tree)\n+tree_compat_oid=$(git rev-parse --output-object-format=$test_compat_hash_algo $tree_oid)\n tree_size=$(($(test_oid rawsz) + 13))\n+tree_compat_size=$(($(test_oid --hash=compat rawsz) + 13))\n tree_pretty_content=\"100644 blob $hello_oid\thello${LF}\"\n+tree_compat_pretty_content=\"100644 blob $hello_compat_oid\thello${LF}\"\n \n run_tests 'tree' $tree_oid $tree_size \"\" \"$tree_pretty_content\"\n+run_tests 'tree' $tree_compat_oid $tree_compat_size \"\" \"$tree_compat_pretty_content\"\n \n commit_message=\"Initial commit\"\n commit_oid=$(echo_without_newline \"$commit_message\" | git commit-tree $tree_oid)\n+commit_compat_oid=$(git rev-parse --output-object-format=$test_compat_hash_algo $commit_oid)\n commit_size=$(($(test_oid hexsz) + 137))\n+commit_compat_size=$(($(test_oid --hash=compat hexsz) + 137))\n commit_content=\"tree $tree_oid\n author $GIT_AUTHOR_NAME <$GIT_AUTHOR_EMAIL> $GIT_AUTHOR_DATE\n committer $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL> $GIT_COMMITTER_DATE\n \n $commit_message\"\n \n+commit_compat_content=\"tree $tree_compat_oid\n+author $GIT_AUTHOR_NAME <$GIT_AUTHOR_EMAIL> $GIT_AUTHOR_DATE\n+committer $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL> $GIT_COMMITTER_DATE\n+\n+$commit_message\"\n+\n run_tests 'commit' $commit_oid $commit_size \"$commit_content\" \"$commit_content\"\n+run_tests 'commit' $commit_compat_oid $commit_compat_size \"$commit_compat_content\" \"$commit_compat_content\"\n \n-tag_header_without_timestamp=\"object $hello_oid\n-type blob\n+tag_header_without_oid=\"type blob\n tag hellotag\n tagger $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL>\"\n+tag_header_without_timestamp=\"object $hello_oid\n+$tag_header_without_oid\"\n+tag_compat_header_without_timestamp=\"object $hello_compat_oid\n+$tag_header_without_oid\"\n tag_description=\"This is a tag\"\n tag_content=\"$tag_header_without_timestamp 0 +0000\n \n+$tag_description\"\n+tag_compat_content=\"$tag_compat_header_without_timestamp 0 +0000\n+\n $tag_description\"\n \n tag_oid=$(echo_without_newline \"$tag_content\" | git hash-object -t tag --stdin -w)\n tag_size=$(strlen \"$tag_content\")\n \n+tag_compat_oid=$(git rev-parse --output-object-format=$test_compat_hash_algo $tag_oid)\n+tag_compat_size=$(strlen \"$tag_compat_content\")\n+\n run_tests 'tag' $tag_oid $tag_size \"$tag_content\" \"$tag_content\"\n+run_tests 'tag' $tag_compat_oid $tag_compat_size \"$tag_compat_content\" \"$tag_compat_content\"\n \n test_expect_success \"Reach a blob from a tag pointing to it\" '\n \techo_without_newline \"$hello_content\" >expect &&\n@@ -303,37 +337,43 @@ test_expect_success \"Reach a blob from a tag pointing to it\" '\n \ttest_cmp expect actual\n '\n \n-for batch in batch batch-check batch-command\n+for oid in $hello_oid $hello_compat_oid\n do\n-    for opt in t s e p\n+    for batch in batch batch-check batch-command\n     do\n+\tfor opt in t s e p\n+\tdo\n \ttest_expect_success \"Passing -$opt with --$batch fails\" '\n-\t    test_must_fail git cat-file --$batch -$opt $hello_oid\n+\t    test_must_fail git cat-file --$batch -$opt $oid\n \t'\n \n \ttest_expect_success \"Passing --$batch with -$opt fails\" '\n-\t    test_must_fail git cat-file -$opt --$batch $hello_oid\n+\t    test_must_fail git cat-file -$opt --$batch $oid\n \t'\n-    done\n+\tdone\n \n-    test_expect_success \"Passing <type> with --$batch fails\" '\n-\ttest_must_fail git cat-file --$batch blob $hello_oid\n-    '\n+\ttest_expect_success \"Passing <type> with --$batch fails\" '\n+\ttest_must_fail git cat-file --$batch blob $oid\n+\t'\n \n-    test_expect_success \"Passing --$batch with <type> fails\" '\n-\ttest_must_fail git cat-file blob --$batch $hello_oid\n-    '\n+\ttest_expect_success \"Passing --$batch with <type> fails\" '\n+\ttest_must_fail git cat-file blob --$batch $oid\n+\t'\n \n-    test_expect_success \"Passing oid with --$batch fails\" '\n-\ttest_must_fail git cat-file --$batch $hello_oid\n-    '\n+\ttest_expect_success \"Passing oid with --$batch fails\" '\n+\ttest_must_fail git cat-file --$batch $oid\n+\t'\n+    done\n done\n \n-for opt in t s e p\n+for oid in $hello_oid $hello_compat_oid\n do\n-    test_expect_success \"Passing -$opt with --follow-symlinks fails\" '\n-\t    test_must_fail git cat-file --follow-symlinks -$opt $hello_oid\n+    for opt in t s e p\n+    do\n+\ttest_expect_success \"Passing -$opt with --follow-symlinks fails\" '\n+\t    test_must_fail git cat-file --follow-symlinks -$opt $oid\n \t'\n+    done\n done\n \n test_expect_success \"--batch-check for a non-existent named object\" '\n@@ -386,112 +426,102 @@ test_expect_success 'empty --batch-check notices missing object' '\n \ttest_cmp expect actual\n '\n \n-batch_input=\"$hello_oid\n-$commit_oid\n-$tag_oid\n+batch_tests () {\n+    boid=$1\n+    loid=$2\n+    lsize=$3\n+    coid=$4\n+    csize=$5\n+    ccontent=$6\n+    toid=$7\n+    tsize=$8\n+    tcontent=$9\n+\n+    batch_input=\"$boid\n+$coid\n+$toid\n deadbeef\n \n \"\n \n-printf \"%s\\0\" \\\n-\t\"$hello_oid blob $hello_size\" \\\n+    printf \"%s\\0\" \\\n+\t\"$boid blob $hello_size\" \\\n \t\"$hello_content\" \\\n-\t\"$commit_oid commit $commit_size\" \\\n-\t\"$commit_content\" \\\n-\t\"$tag_oid tag $tag_size\" \\\n-\t\"$tag_content\" \\\n+\t\"$coid commit $csize\" \\\n+\t\"$ccontent\" \\\n+\t\"$toid tag $tsize\" \\\n+\t\"$tcontent\" \\\n \t\"deadbeef missing\" \\\n \t\" missing\" >batch_output\n \n-test_expect_success '--batch with multiple oids gives correct format' '\n+    test_expect_success '--batch with multiple oids gives correct format' '\n \ttr \"\\0\" \"\\n\" <batch_output >expect &&\n \techo_without_newline \"$batch_input\" >in &&\n \tgit cat-file --batch <in >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success '--batch, -z with multiple oids gives correct format' '\n+    test_expect_success '--batch, -z with multiple oids gives correct format' '\n \techo_without_newline_nul \"$batch_input\" >in &&\n \ttr \"\\0\" \"\\n\" <batch_output >expect &&\n \tgit cat-file --batch -z <in >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success '--batch, -Z with multiple oids gives correct format' '\n+    test_expect_success '--batch, -Z with multiple oids gives correct format' '\n \techo_without_newline_nul \"$batch_input\" >in &&\n \tgit cat-file --batch -Z <in >actual &&\n \ttest_cmp batch_output actual\n-'\n+    '\n \n-batch_check_input=\"$hello_oid\n-$tree_oid\n-$commit_oid\n-$tag_oid\n+batch_check_input=\"$boid\n+$loid\n+$coid\n+$toid\n deadbeef\n \n \"\n \n-printf \"%s\\0\" \\\n-\t\"$hello_oid blob $hello_size\" \\\n-\t\"$tree_oid tree $tree_size\" \\\n-\t\"$commit_oid commit $commit_size\" \\\n-\t\"$tag_oid tag $tag_size\" \\\n+    printf \"%s\\0\" \\\n+\t\"$boid blob $hello_size\" \\\n+\t\"$loid tree $lsize\" \\\n+\t\"$coid commit $csize\" \\\n+\t\"$toid tag $tsize\" \\\n \t\"deadbeef missing\" \\\n \t\" missing\" >batch_check_output\n \n-test_expect_success \"--batch-check with multiple oids gives correct format\" '\n+    test_expect_success \"--batch-check with multiple oids gives correct format\" '\n \ttr \"\\0\" \"\\n\" <batch_check_output >expect &&\n \techo_without_newline \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check <in >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success \"--batch-check, -z with multiple oids gives correct format\" '\n+    test_expect_success \"--batch-check, -z with multiple oids gives correct format\" '\n \ttr \"\\0\" \"\\n\" <batch_check_output >expect &&\n \techo_without_newline_nul \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check -z <in >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success \"--batch-check, -Z with multiple oids gives correct format\" '\n+    test_expect_success \"--batch-check, -Z with multiple oids gives correct format\" '\n \techo_without_newline_nul \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check -Z <in >actual &&\n \ttest_cmp batch_check_output actual\n-'\n-\n-test_expect_success FUNNYNAMES 'setup with newline in input' '\n-\ttouch -- \"newline${LF}embedded\" &&\n-\tgit add -- \"newline${LF}embedded\" &&\n-\tgit commit -m \"file with newline embedded\" &&\n-\ttest_tick &&\n-\n-\tprintf \"HEAD:newline${LF}embedded\" >in\n-'\n-\n-test_expect_success FUNNYNAMES '--batch-check, -z with newline in input' '\n-\tgit cat-file --batch-check -z <in >actual &&\n-\techo \"$(git rev-parse \"HEAD:newline${LF}embedded\") blob 0\" >expect &&\n-\ttest_cmp expect actual\n-'\n-\n-test_expect_success FUNNYNAMES '--batch-check, -Z with newline in input' '\n-\tgit cat-file --batch-check -Z <in >actual &&\n-\tprintf \"%s\\0\" \"$(git rev-parse \"HEAD:newline${LF}embedded\") blob 0\" >expect &&\n-\ttest_cmp expect actual\n-'\n+    '\n \n-batch_command_multiple_info=\"info $hello_oid\n-info $tree_oid\n-info $commit_oid\n-info $tag_oid\n+batch_command_multiple_info=\"info $boid\n+info $loid\n+info $coid\n+info $toid\n info deadbeef\"\n \n-test_expect_success '--batch-command with multiple info calls gives correct format' '\n+    test_expect_success '--batch-command with multiple info calls gives correct format' '\n \tcat >expect <<-EOF &&\n-\t$hello_oid blob $hello_size\n-\t$tree_oid tree $tree_size\n-\t$commit_oid commit $commit_size\n-\t$tag_oid tag $tag_size\n+\t$boid blob $hello_size\n+\t$loid tree $lsize\n+\t$coid commit $csize\n+\t$toid tag $tsize\n \tdeadbeef missing\n \tEOF\n \n@@ -510,22 +540,22 @@ test_expect_success '--batch-command with multiple info calls gives correct form\n \tgit cat-file --batch-command --buffer -Z <in >actual &&\n \n \ttest_cmp expect_nul actual\n-'\n+    '\n \n-batch_command_multiple_contents=\"contents $hello_oid\n-contents $commit_oid\n-contents $tag_oid\n+batch_command_multiple_contents=\"contents $boid\n+contents $coid\n+contents $toid\n contents deadbeef\n flush\"\n \n-test_expect_success '--batch-command with multiple command calls gives correct format' '\n+    test_expect_success '--batch-command with multiple command calls gives correct format' '\n \tprintf \"%s\\0\" \\\n-\t\t\"$hello_oid blob $hello_size\" \\\n+\t\t\"$boid blob $hello_size\" \\\n \t\t\"$hello_content\" \\\n-\t\t\"$commit_oid commit $commit_size\" \\\n-\t\t\"$commit_content\" \\\n-\t\t\"$tag_oid tag $tag_size\" \\\n-\t\t\"$tag_content\" \\\n+\t\t\"$coid commit $csize\" \\\n+\t\t\"$ccontent\" \\\n+\t\t\"$toid tag $tsize\" \\\n+\t\t\"$tcontent\" \\\n \t\t\"deadbeef missing\" >expect_nul &&\n \ttr \"\\0\" \"\\n\" <expect_nul >expect &&\n \n@@ -543,6 +573,33 @@ test_expect_success '--batch-command with multiple command calls gives correct f\n \tgit cat-file --batch-command --buffer -Z <in >actual &&\n \n \ttest_cmp expect_nul actual\n+    '\n+\n+}\n+\n+batch_tests $hello_oid $tree_oid $tree_size $commit_oid $commit_size \"$commit_content\" $tag_oid $tag_size \"$tag_content\"\n+batch_tests $hello_compat_oid $tree_compat_oid $tree_compat_size $commit_compat_oid $commit_compat_size \"$commit_compat_content\" $tag_compat_oid $tag_compat_size \"$tag_compat_content\"\n+\n+\n+test_expect_success FUNNYNAMES 'setup with newline in input' '\n+\ttouch -- \"newline${LF}embedded\" &&\n+\tgit add -- \"newline${LF}embedded\" &&\n+\tgit commit -m \"file with newline embedded\" &&\n+\ttest_tick &&\n+\n+\tprintf \"HEAD:newline${LF}embedded\" >in\n+'\n+\n+test_expect_success FUNNYNAMES '--batch-check, -z with newline in input' '\n+\tgit cat-file --batch-check -z <in >actual &&\n+\techo \"$(git rev-parse \"HEAD:newline${LF}embedded\") blob 0\" >expect &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success FUNNYNAMES '--batch-check, -Z with newline in input' '\n+\tgit cat-file --batch-check -Z <in >actual &&\n+\tprintf \"%s\\0\" \"$(git rev-parse \"HEAD:newline${LF}embedded\") blob 0\" >expect &&\n+\ttest_cmp expect actual\n '\n \n test_expect_success 'setup blobs which are likely to delta' '\n-- \n2.41.0\n\n"},{"id":"482409","messageId":"20230927195537.1682-28-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 28/30] t1006: Rename sha1 to oid","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:35Z","receivedAt":"2023-09-27T19:57:06Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nBefore I extend this test, changing the naming of the relevant\nhash from sha1 to oid.  Calling the hash sha1 is incorrect today\nas it can be either sha1 or sha256 depending on the value of\nGIT_DEFAULT_HASH_FUNCTION when the test is called.\n\nI plan to test sha1 and sha256 simultaneously in the same repository.\nHaving a name like sha1 will be even more confusing.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n t/t1006-cat-file.sh | 220 ++++++++++++++++++++++----------------------\n 1 file changed, 110 insertions(+), 110 deletions(-)\n\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex d73a0be1b9d1..9b018b538950 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -112,65 +112,65 @@ strlen () {\n \n run_tests () {\n     type=$1\n-    sha1=$2\n+    oid=$2\n     size=$3\n     content=$4\n     pretty_content=$5\n \n-    batch_output=\"$sha1 $type $size\n+    batch_output=\"$oid $type $size\n $content\"\n \n     test_expect_success \"$type exists\" '\n-\tgit cat-file -e $sha1\n+\tgit cat-file -e $oid\n     '\n \n     test_expect_success \"Type of $type is correct\" '\n \techo $type >expect &&\n-\tgit cat-file -t $sha1 >actual &&\n+\tgit cat-file -t $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"Size of $type is correct\" '\n \techo $size >expect &&\n-\tgit cat-file -s $sha1 >actual &&\n+\tgit cat-file -s $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"Type of $type is correct using --allow-unknown-type\" '\n \techo $type >expect &&\n-\tgit cat-file -t --allow-unknown-type $sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"Size of $type is correct using --allow-unknown-type\" '\n \techo $size >expect &&\n-\tgit cat-file -s --allow-unknown-type $sha1 >actual &&\n+\tgit cat-file -s --allow-unknown-type $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test -z \"$content\" ||\n     test_expect_success \"Content of $type is correct\" '\n \techo_without_newline \"$content\" >expect &&\n-\tgit cat-file $type $sha1 >actual &&\n+\tgit cat-file $type $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"Pretty content of $type is correct\" '\n \techo_without_newline \"$pretty_content\" >expect &&\n-\tgit cat-file -p $sha1 >actual &&\n+\tgit cat-file -p $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test -z \"$content\" ||\n     test_expect_success \"--batch output of $type is correct\" '\n \techo \"$batch_output\" >expect &&\n-\techo $sha1 | git cat-file --batch >actual &&\n+\techo $oid | git cat-file --batch >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"--batch-check output of $type is correct\" '\n-\techo \"$sha1 $type $size\" >expect &&\n-\techo_without_newline $sha1 | git cat-file --batch-check >actual &&\n+\techo \"$oid $type $size\" >expect &&\n+\techo_without_newline $oid | git cat-file --batch-check >actual &&\n \ttest_cmp expect actual\n     '\n \n@@ -179,33 +179,33 @@ $content\"\n \ttest -z \"$content\" ||\n \t\ttest_expect_success \"--batch-command $opt output of $type content is correct\" '\n \t\techo \"$batch_output\" >expect &&\n-\t\ttest_write_lines \"contents $sha1\" | git cat-file --batch-command $opt >actual &&\n+\t\ttest_write_lines \"contents $oid\" | git cat-file --batch-command $opt >actual &&\n \t\ttest_cmp expect actual\n \t'\n \n \ttest_expect_success \"--batch-command $opt output of $type info is correct\" '\n-\t\techo \"$sha1 $type $size\" >expect &&\n-\t\ttest_write_lines \"info $sha1\" |\n+\t\techo \"$oid $type $size\" >expect &&\n+\t\ttest_write_lines \"info $oid\" |\n \t\tgit cat-file --batch-command $opt >actual &&\n \t\ttest_cmp expect actual\n \t'\n     done\n \n     test_expect_success \"custom --batch-check format\" '\n-\techo \"$type $sha1\" >expect &&\n-\techo $sha1 | git cat-file --batch-check=\"%(objecttype) %(objectname)\" >actual &&\n+\techo \"$type $oid\" >expect &&\n+\techo $oid | git cat-file --batch-check=\"%(objecttype) %(objectname)\" >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"custom --batch-command format\" '\n-\techo \"$type $sha1\" >expect &&\n-\techo \"info $sha1\" | git cat-file --batch-command=\"%(objecttype) %(objectname)\" >actual &&\n+\techo \"$type $oid\" >expect &&\n+\techo \"info $oid\" | git cat-file --batch-command=\"%(objecttype) %(objectname)\" >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success '--batch-check with %(rest)' '\n \techo \"$type this is some extra content\" >expect &&\n-\techo \"$sha1    this is some extra content\" |\n+\techo \"$oid    this is some extra content\" |\n \t\tgit cat-file --batch-check=\"%(objecttype) %(rest)\" >actual &&\n \ttest_cmp expect actual\n     '\n@@ -216,7 +216,7 @@ $content\"\n \t\techo \"$size\" &&\n \t\techo \"$content\"\n \t} >expect &&\n-\techo $sha1 | git cat-file --batch=\"%(objectsize)\" >actual &&\n+\techo $oid | git cat-file --batch=\"%(objectsize)\" >actual &&\n \ttest_cmp expect actual\n     '\n \n@@ -226,25 +226,25 @@ $content\"\n \t\techo \"$type\" &&\n \t\techo \"$content\"\n \t} >expect &&\n-\techo $sha1 | git cat-file --batch=\"%(objecttype)\" >actual &&\n+\techo $oid | git cat-file --batch=\"%(objecttype)\" >actual &&\n \ttest_cmp expect actual\n     '\n }\n \n hello_content=\"Hello World\"\n hello_size=$(strlen \"$hello_content\")\n-hello_sha1=$(echo_without_newline \"$hello_content\" | git hash-object --stdin)\n+hello_oid=$(echo_without_newline \"$hello_content\" | git hash-object --stdin)\n \n test_expect_success \"setup\" '\n \techo_without_newline \"$hello_content\" > hello &&\n \tgit update-index --add hello\n '\n \n-run_tests 'blob' $hello_sha1 $hello_size \"$hello_content\" \"$hello_content\"\n+run_tests 'blob' $hello_oid $hello_size \"$hello_content\" \"$hello_content\"\n \n test_expect_success '--batch-command --buffer with flush for blob info' '\n-\techo \"$hello_sha1 blob $hello_size\" >expect &&\n-\ttest_write_lines \"info $hello_sha1\" \"flush\" |\n+\techo \"$hello_oid blob $hello_size\" >expect &&\n+\ttest_write_lines \"info $hello_oid\" \"flush\" |\n \tGIT_TEST_CAT_FILE_NO_FLUSH_ON_EXIT=1 \\\n \tgit cat-file --batch-command --buffer >actual &&\n \ttest_cmp expect actual\n@@ -252,38 +252,38 @@ test_expect_success '--batch-command --buffer with flush for blob info' '\n \n test_expect_success '--batch-command --buffer without flush for blob info' '\n \ttouch output &&\n-\ttest_write_lines \"info $hello_sha1\" |\n+\ttest_write_lines \"info $hello_oid\" |\n \tGIT_TEST_CAT_FILE_NO_FLUSH_ON_EXIT=1 \\\n \tgit cat-file --batch-command --buffer >>output &&\n \ttest_must_be_empty output\n '\n \n test_expect_success '--batch-check without %(rest) considers whole line' '\n-\techo \"$hello_sha1 blob $hello_size\" >expect &&\n-\tgit update-index --add --cacheinfo 100644 $hello_sha1 \"white space\" &&\n+\techo \"$hello_oid blob $hello_size\" >expect &&\n+\tgit update-index --add --cacheinfo 100644 $hello_oid \"white space\" &&\n \ttest_when_finished \"git update-index --remove \\\"white space\\\"\" &&\n \techo \":white space\" | git cat-file --batch-check >actual &&\n \ttest_cmp expect actual\n '\n \n-tree_sha1=$(git write-tree)\n+tree_oid=$(git write-tree)\n tree_size=$(($(test_oid rawsz) + 13))\n-tree_pretty_content=\"100644 blob $hello_sha1\thello${LF}\"\n+tree_pretty_content=\"100644 blob $hello_oid\thello${LF}\"\n \n-run_tests 'tree' $tree_sha1 $tree_size \"\" \"$tree_pretty_content\"\n+run_tests 'tree' $tree_oid $tree_size \"\" \"$tree_pretty_content\"\n \n commit_message=\"Initial commit\"\n-commit_sha1=$(echo_without_newline \"$commit_message\" | git commit-tree $tree_sha1)\n+commit_oid=$(echo_without_newline \"$commit_message\" | git commit-tree $tree_oid)\n commit_size=$(($(test_oid hexsz) + 137))\n-commit_content=\"tree $tree_sha1\n+commit_content=\"tree $tree_oid\n author $GIT_AUTHOR_NAME <$GIT_AUTHOR_EMAIL> $GIT_AUTHOR_DATE\n committer $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL> $GIT_COMMITTER_DATE\n \n $commit_message\"\n \n-run_tests 'commit' $commit_sha1 $commit_size \"$commit_content\" \"$commit_content\"\n+run_tests 'commit' $commit_oid $commit_size \"$commit_content\" \"$commit_content\"\n \n-tag_header_without_timestamp=\"object $hello_sha1\n+tag_header_without_timestamp=\"object $hello_oid\n type blob\n tag hellotag\n tagger $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL>\"\n@@ -292,14 +292,14 @@ tag_content=\"$tag_header_without_timestamp 0 +0000\n \n $tag_description\"\n \n-tag_sha1=$(echo_without_newline \"$tag_content\" | git hash-object -t tag --stdin -w)\n+tag_oid=$(echo_without_newline \"$tag_content\" | git hash-object -t tag --stdin -w)\n tag_size=$(strlen \"$tag_content\")\n \n-run_tests 'tag' $tag_sha1 $tag_size \"$tag_content\" \"$tag_content\"\n+run_tests 'tag' $tag_oid $tag_size \"$tag_content\" \"$tag_content\"\n \n test_expect_success \"Reach a blob from a tag pointing to it\" '\n \techo_without_newline \"$hello_content\" >expect &&\n-\tgit cat-file blob $tag_sha1 >actual &&\n+\tgit cat-file blob $tag_oid >actual &&\n \ttest_cmp expect actual\n '\n \n@@ -308,31 +308,31 @@ do\n     for opt in t s e p\n     do\n \ttest_expect_success \"Passing -$opt with --$batch fails\" '\n-\t    test_must_fail git cat-file --$batch -$opt $hello_sha1\n+\t    test_must_fail git cat-file --$batch -$opt $hello_oid\n \t'\n \n \ttest_expect_success \"Passing --$batch with -$opt fails\" '\n-\t    test_must_fail git cat-file -$opt --$batch $hello_sha1\n+\t    test_must_fail git cat-file -$opt --$batch $hello_oid\n \t'\n     done\n \n     test_expect_success \"Passing <type> with --$batch fails\" '\n-\ttest_must_fail git cat-file --$batch blob $hello_sha1\n+\ttest_must_fail git cat-file --$batch blob $hello_oid\n     '\n \n     test_expect_success \"Passing --$batch with <type> fails\" '\n-\ttest_must_fail git cat-file blob --$batch $hello_sha1\n+\ttest_must_fail git cat-file blob --$batch $hello_oid\n     '\n \n-    test_expect_success \"Passing sha1 with --$batch fails\" '\n-\ttest_must_fail git cat-file --$batch $hello_sha1\n+    test_expect_success \"Passing oid with --$batch fails\" '\n+\ttest_must_fail git cat-file --$batch $hello_oid\n     '\n done\n \n for opt in t s e p\n do\n     test_expect_success \"Passing -$opt with --follow-symlinks fails\" '\n-\t    test_must_fail git cat-file --follow-symlinks -$opt $hello_sha1\n+\t    test_must_fail git cat-file --follow-symlinks -$opt $hello_oid\n \t'\n done\n \n@@ -360,12 +360,12 @@ test_expect_success \"--batch-check for a non-existent hash\" '\n \n test_expect_success \"--batch for an existent and a non-existent hash\" '\n \tcat >expect <<-EOF &&\n-\t$tag_sha1 tag $tag_size\n+\t$tag_oid tag $tag_size\n \t$tag_content\n \t0000000000000000000000000000000000000000 missing\n \tEOF\n \n-\tprintf \"$tag_sha1\\n0000000000000000000000000000000000000000\" >in &&\n+\tprintf \"$tag_oid\\n0000000000000000000000000000000000000000\" >in &&\n \tgit cat-file --batch <in >actual &&\n \ttest_cmp expect actual\n '\n@@ -386,74 +386,74 @@ test_expect_success 'empty --batch-check notices missing object' '\n \ttest_cmp expect actual\n '\n \n-batch_input=\"$hello_sha1\n-$commit_sha1\n-$tag_sha1\n+batch_input=\"$hello_oid\n+$commit_oid\n+$tag_oid\n deadbeef\n \n \"\n \n printf \"%s\\0\" \\\n-\t\"$hello_sha1 blob $hello_size\" \\\n+\t\"$hello_oid blob $hello_size\" \\\n \t\"$hello_content\" \\\n-\t\"$commit_sha1 commit $commit_size\" \\\n+\t\"$commit_oid commit $commit_size\" \\\n \t\"$commit_content\" \\\n-\t\"$tag_sha1 tag $tag_size\" \\\n+\t\"$tag_oid tag $tag_size\" \\\n \t\"$tag_content\" \\\n \t\"deadbeef missing\" \\\n \t\" missing\" >batch_output\n \n-test_expect_success '--batch with multiple sha1s gives correct format' '\n+test_expect_success '--batch with multiple oids gives correct format' '\n \ttr \"\\0\" \"\\n\" <batch_output >expect &&\n \techo_without_newline \"$batch_input\" >in &&\n \tgit cat-file --batch <in >actual &&\n \ttest_cmp expect actual\n '\n \n-test_expect_success '--batch, -z with multiple sha1s gives correct format' '\n+test_expect_success '--batch, -z with multiple oids gives correct format' '\n \techo_without_newline_nul \"$batch_input\" >in &&\n \ttr \"\\0\" \"\\n\" <batch_output >expect &&\n \tgit cat-file --batch -z <in >actual &&\n \ttest_cmp expect actual\n '\n \n-test_expect_success '--batch, -Z with multiple sha1s gives correct format' '\n+test_expect_success '--batch, -Z with multiple oids gives correct format' '\n \techo_without_newline_nul \"$batch_input\" >in &&\n \tgit cat-file --batch -Z <in >actual &&\n \ttest_cmp batch_output actual\n '\n \n-batch_check_input=\"$hello_sha1\n-$tree_sha1\n-$commit_sha1\n-$tag_sha1\n+batch_check_input=\"$hello_oid\n+$tree_oid\n+$commit_oid\n+$tag_oid\n deadbeef\n \n \"\n \n printf \"%s\\0\" \\\n-\t\"$hello_sha1 blob $hello_size\" \\\n-\t\"$tree_sha1 tree $tree_size\" \\\n-\t\"$commit_sha1 commit $commit_size\" \\\n-\t\"$tag_sha1 tag $tag_size\" \\\n+\t\"$hello_oid blob $hello_size\" \\\n+\t\"$tree_oid tree $tree_size\" \\\n+\t\"$commit_oid commit $commit_size\" \\\n+\t\"$tag_oid tag $tag_size\" \\\n \t\"deadbeef missing\" \\\n \t\" missing\" >batch_check_output\n \n-test_expect_success \"--batch-check with multiple sha1s gives correct format\" '\n+test_expect_success \"--batch-check with multiple oids gives correct format\" '\n \ttr \"\\0\" \"\\n\" <batch_check_output >expect &&\n \techo_without_newline \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check <in >actual &&\n \ttest_cmp expect actual\n '\n \n-test_expect_success \"--batch-check, -z with multiple sha1s gives correct format\" '\n+test_expect_success \"--batch-check, -z with multiple oids gives correct format\" '\n \ttr \"\\0\" \"\\n\" <batch_check_output >expect &&\n \techo_without_newline_nul \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check -z <in >actual &&\n \ttest_cmp expect actual\n '\n \n-test_expect_success \"--batch-check, -Z with multiple sha1s gives correct format\" '\n+test_expect_success \"--batch-check, -Z with multiple oids gives correct format\" '\n \techo_without_newline_nul \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check -Z <in >actual &&\n \ttest_cmp batch_check_output actual\n@@ -480,18 +480,18 @@ test_expect_success FUNNYNAMES '--batch-check, -Z with newline in input' '\n \ttest_cmp expect actual\n '\n \n-batch_command_multiple_info=\"info $hello_sha1\n-info $tree_sha1\n-info $commit_sha1\n-info $tag_sha1\n+batch_command_multiple_info=\"info $hello_oid\n+info $tree_oid\n+info $commit_oid\n+info $tag_oid\n info deadbeef\"\n \n test_expect_success '--batch-command with multiple info calls gives correct format' '\n \tcat >expect <<-EOF &&\n-\t$hello_sha1 blob $hello_size\n-\t$tree_sha1 tree $tree_size\n-\t$commit_sha1 commit $commit_size\n-\t$tag_sha1 tag $tag_size\n+\t$hello_oid blob $hello_size\n+\t$tree_oid tree $tree_size\n+\t$commit_oid commit $commit_size\n+\t$tag_oid tag $tag_size\n \tdeadbeef missing\n \tEOF\n \n@@ -512,19 +512,19 @@ test_expect_success '--batch-command with multiple info calls gives correct form\n \ttest_cmp expect_nul actual\n '\n \n-batch_command_multiple_contents=\"contents $hello_sha1\n-contents $commit_sha1\n-contents $tag_sha1\n+batch_command_multiple_contents=\"contents $hello_oid\n+contents $commit_oid\n+contents $tag_oid\n contents deadbeef\n flush\"\n \n test_expect_success '--batch-command with multiple command calls gives correct format' '\n \tprintf \"%s\\0\" \\\n-\t\t\"$hello_sha1 blob $hello_size\" \\\n+\t\t\"$hello_oid blob $hello_size\" \\\n \t\t\"$hello_content\" \\\n-\t\t\"$commit_sha1 commit $commit_size\" \\\n+\t\t\"$commit_oid commit $commit_size\" \\\n \t\t\"$commit_content\" \\\n-\t\t\"$tag_sha1 tag $tag_size\" \\\n+\t\t\"$tag_oid tag $tag_size\" \\\n \t\t\"$tag_content\" \\\n \t\t\"deadbeef missing\" >expect_nul &&\n \ttr \"\\0\" \"\\n\" <expect_nul >expect &&\n@@ -569,7 +569,7 @@ test_expect_success 'confirm that neither loose blob is a delta' '\n # we will check only that one of the two objects is a delta\n # against the other, but not the order. We can do so by just\n # asking for the base of both, and checking whether either\n-# sha1 appears in the output.\n+# oid appears in the output.\n test_expect_success '%(deltabase) reports packed delta bases' '\n \tgit repack -ad &&\n \tgit cat-file --batch-check=\"%(deltabase)\" <blobs >actual &&\n@@ -583,12 +583,12 @@ test_expect_success 'setup bogus data' '\n \tbogus_short_type=\"bogus\" &&\n \tbogus_short_content=\"bogus\" &&\n \tbogus_short_size=$(strlen \"$bogus_short_content\") &&\n-\tbogus_short_sha1=$(echo_without_newline \"$bogus_short_content\" | git hash-object -t $bogus_short_type --literally -w --stdin) &&\n+\tbogus_short_oid=$(echo_without_newline \"$bogus_short_content\" | git hash-object -t $bogus_short_type --literally -w --stdin) &&\n \n \tbogus_long_type=\"abcdefghijklmnopqrstuvwxyz1234679\" &&\n \tbogus_long_content=\"bogus\" &&\n \tbogus_long_size=$(strlen \"$bogus_long_content\") &&\n-\tbogus_long_sha1=$(echo_without_newline \"$bogus_long_content\" | git hash-object -t $bogus_long_type --literally -w --stdin)\n+\tbogus_long_oid=$(echo_without_newline \"$bogus_long_content\" | git hash-object -t $bogus_long_type --literally -w --stdin)\n '\n \n for arg1 in '' --allow-unknown-type\n@@ -608,9 +608,9 @@ do\n \n \t\t\tif test \"$arg1\" = \"--allow-unknown-type\"\n \t\t\tthen\n-\t\t\t\tgit cat-file $arg1 $arg2 $bogus_short_sha1\n+\t\t\t\tgit cat-file $arg1 $arg2 $bogus_short_oid\n \t\t\telse\n-\t\t\t\ttest_must_fail git cat-file $arg1 $arg2 $bogus_short_sha1 >out 2>actual &&\n+\t\t\t\ttest_must_fail git cat-file $arg1 $arg2 $bogus_short_oid >out 2>actual &&\n \t\t\t\ttest_must_be_empty out &&\n \t\t\t\ttest_cmp expect actual\n \t\t\tfi\n@@ -620,21 +620,21 @@ do\n \t\t\tif test \"$arg2\" = \"-p\"\n \t\t\tthen\n \t\t\t\tcat >expect <<-EOF\n-\t\t\t\terror: header for $bogus_long_sha1 too long, exceeds 32 bytes\n-\t\t\t\tfatal: Not a valid object name $bogus_long_sha1\n+\t\t\t\terror: header for $bogus_long_oid too long, exceeds 32 bytes\n+\t\t\t\tfatal: Not a valid object name $bogus_long_oid\n \t\t\t\tEOF\n \t\t\telse\n \t\t\t\tcat >expect <<-EOF\n-\t\t\t\terror: header for $bogus_long_sha1 too long, exceeds 32 bytes\n+\t\t\t\terror: header for $bogus_long_oid too long, exceeds 32 bytes\n \t\t\t\tfatal: git cat-file: could not get object info\n \t\t\t\tEOF\n \t\t\tfi &&\n \n \t\t\tif test \"$arg1\" = \"--allow-unknown-type\"\n \t\t\tthen\n-\t\t\t\tgit cat-file $arg1 $arg2 $bogus_short_sha1\n+\t\t\t\tgit cat-file $arg1 $arg2 $bogus_short_oid\n \t\t\telse\n-\t\t\t\ttest_must_fail git cat-file $arg1 $arg2 $bogus_long_sha1 >out 2>actual &&\n+\t\t\t\ttest_must_fail git cat-file $arg1 $arg2 $bogus_long_oid >out 2>actual &&\n \t\t\t\ttest_must_be_empty out &&\n \t\t\t\ttest_cmp expect actual\n \t\t\tfi\n@@ -668,28 +668,28 @@ do\n done\n \n test_expect_success '-e is OK with a broken object without --allow-unknown-type' '\n-\tgit cat-file -e $bogus_short_sha1\n+\tgit cat-file -e $bogus_short_oid\n '\n \n test_expect_success '-e can not be combined with --allow-unknown-type' '\n-\ttest_expect_code 128 git cat-file -e --allow-unknown-type $bogus_short_sha1\n+\ttest_expect_code 128 git cat-file -e --allow-unknown-type $bogus_short_oid\n '\n \n test_expect_success '-p cannot print a broken object even with --allow-unknown-type' '\n-\ttest_must_fail git cat-file -p $bogus_short_sha1 &&\n-\ttest_expect_code 128 git cat-file -p --allow-unknown-type $bogus_short_sha1\n+\ttest_must_fail git cat-file -p $bogus_short_oid &&\n+\ttest_expect_code 128 git cat-file -p --allow-unknown-type $bogus_short_oid\n '\n \n test_expect_success '<type> <hash> does not work with objects of broken types' '\n \tcat >err.expect <<-\\EOF &&\n \tfatal: invalid object type \"bogus\"\n \tEOF\n-\ttest_must_fail git cat-file $bogus_short_type $bogus_short_sha1 2>err.actual &&\n+\ttest_must_fail git cat-file $bogus_short_type $bogus_short_oid 2>err.actual &&\n \ttest_cmp err.expect err.actual\n '\n \n test_expect_success 'broken types combined with --batch and --batch-check' '\n-\techo $bogus_short_sha1 >bogus-oid &&\n+\techo $bogus_short_oid >bogus-oid &&\n \n \tcat >err.expect <<-\\EOF &&\n \tfatal: invalid object type\n@@ -711,52 +711,52 @@ test_expect_success 'the --allow-unknown-type option does not consider replaceme\n \tcat >expect <<-EOF &&\n \t$bogus_short_type\n \tEOF\n-\tgit cat-file -t --allow-unknown-type $bogus_short_sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $bogus_short_oid >actual &&\n \ttest_cmp expect actual &&\n \n \t# Create it manually, as \"git replace\" will die on bogus\n \t# types.\n \thead=$(git rev-parse --verify HEAD) &&\n-\ttest_when_finished \"test-tool ref-store main delete-refs 0 msg refs/replace/$bogus_short_sha1\" &&\n-\ttest-tool ref-store main update-ref msg \"refs/replace/$bogus_short_sha1\" $head $ZERO_OID REF_SKIP_OID_VERIFICATION &&\n+\ttest_when_finished \"test-tool ref-store main delete-refs 0 msg refs/replace/$bogus_short_oid\" &&\n+\ttest-tool ref-store main update-ref msg \"refs/replace/$bogus_short_oid\" $head $ZERO_OID REF_SKIP_OID_VERIFICATION &&\n \n \tcat >expect <<-EOF &&\n \tcommit\n \tEOF\n-\tgit cat-file -t --allow-unknown-type $bogus_short_sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $bogus_short_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success \"Type of broken object is correct\" '\n \techo $bogus_short_type >expect &&\n-\tgit cat-file -t --allow-unknown-type $bogus_short_sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $bogus_short_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success \"Size of broken object is correct\" '\n \techo $bogus_short_size >expect &&\n-\tgit cat-file -s --allow-unknown-type $bogus_short_sha1 >actual &&\n+\tgit cat-file -s --allow-unknown-type $bogus_short_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success 'clean up broken object' '\n-\trm .git/objects/$(test_oid_to_path $bogus_short_sha1)\n+\trm .git/objects/$(test_oid_to_path $bogus_short_oid)\n '\n \n test_expect_success \"Type of broken object is correct when type is large\" '\n \techo $bogus_long_type >expect &&\n-\tgit cat-file -t --allow-unknown-type $bogus_long_sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $bogus_long_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success \"Size of large broken object is correct when type is large\" '\n \techo $bogus_long_size >expect &&\n-\tgit cat-file -s --allow-unknown-type $bogus_long_sha1 >actual &&\n+\tgit cat-file -s --allow-unknown-type $bogus_long_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success 'clean up broken object' '\n-\trm .git/objects/$(test_oid_to_path $bogus_long_sha1)\n+\trm .git/objects/$(test_oid_to_path $bogus_long_oid)\n '\n \n test_expect_success 'cat-file -t and -s on corrupt loose object' '\n@@ -853,7 +853,7 @@ test_expect_success 'prep for symlink tests' '\n \ttest_ln_s_add loop2 loop1 &&\n \tgit add morx dir/subdir/ind2 dir/ind1 &&\n \tgit commit -am \"test\" &&\n-\techo $hello_sha1 blob $hello_size >found\n+\techo $hello_oid blob $hello_size >found\n '\n \n test_expect_success 'git cat-file --batch-check --follow-symlinks works for non-links' '\n@@ -941,7 +941,7 @@ test_expect_success 'git cat-file --batch-check --follow-symlinks works for dir/\n \techo HEAD:dirlink/morx >>expect &&\n \techo HEAD:dirlink/morx | git cat-file --batch-check --follow-symlinks >actual &&\n \ttest_cmp expect actual &&\n-\techo $hello_sha1 blob $hello_size >expect &&\n+\techo $hello_oid blob $hello_size >expect &&\n \techo HEAD:dirlink/ind1 | git cat-file --batch-check --follow-symlinks >actual &&\n \ttest_cmp expect actual\n '\n-- \n2.41.0\n\n"},{"id":"482410","messageId":"20230927195537.1682-30-ebiederm@gmail.com","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH 30/30] t1016-compatObjectFormat: Add tests to verify the conversion between objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-27T19:55:37Z","receivedAt":"2023-09-27T19:57:07Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nFor now my strategy is simple.  Create two identical repositories one\nin each format.  Use fixed timestamps. Verify the dynamically computed\ncompatibility objects from one repository match the objects stored in\nthe other repository.\n\nA general limitation of this strategy is that the git when generating\nsigned tags and commits with compatObjectFormat enabled will generate\na signature for both formats.  To overcome this limitation I have\nadded \"test-tool delete-gpgsig\" that when fed an signed commit or tag\nwith two signatures deletes one of the signatures.\n\nWith that in place I can have \"git commit\" and  \"git tag\" generate\nsigned objects, have my tool delete one, and feed the new object\ninto \"git hash-object\" to create the kinds of commits and tags\ngit without compatObjectFormat enabled will generate.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Makefile                      |   1 +\n t/helper/test-delete-gpgsig.c |  62 ++++++++\n t/helper/test-tool.c          |   1 +\n t/helper/test-tool.h          |   1 +\n t/t1016-compatObjectFormat.sh | 280 ++++++++++++++++++++++++++++++++++\n t/t1016/gpg                   |   2 +\n 6 files changed, 347 insertions(+)\n create mode 100644 t/helper/test-delete-gpgsig.c\n create mode 100755 t/t1016-compatObjectFormat.sh\n create mode 100755 t/t1016/gpg\n\ndiff --git a/Makefile b/Makefile\nindex 3c18664def9a..3e4444fb9ab2 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -790,6 +790,7 @@ TEST_BUILTINS_OBJS += test-crontab.o\n TEST_BUILTINS_OBJS += test-csprng.o\n TEST_BUILTINS_OBJS += test-ctype.o\n TEST_BUILTINS_OBJS += test-date.o\n+TEST_BUILTINS_OBJS += test-delete-gpgsig.o\n TEST_BUILTINS_OBJS += test-delta.o\n TEST_BUILTINS_OBJS += test-dir-iterator.o\n TEST_BUILTINS_OBJS += test-drop-caches.o\ndiff --git a/t/helper/test-delete-gpgsig.c b/t/helper/test-delete-gpgsig.c\nnew file mode 100644\nindex 000000000000..e36831af03f6\n--- /dev/null\n+++ b/t/helper/test-delete-gpgsig.c\n@@ -0,0 +1,62 @@\n+#include \"test-tool.h\"\n+#include \"gpg-interface.h\"\n+#include \"strbuf.h\"\n+\n+\n+int cmd__delete_gpgsig(int argc, const char **argv)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tconst char *pattern = \"gpgsig\";\n+\tconst char *bufptr, *tail, *eol;\n+\tint deleting = 0;\n+\tsize_t plen;\n+\n+\tif (argc >= 2) {\n+\t\tpattern = argv[1];\n+\t\targv++;\n+\t\targc--;\n+\t}\n+\n+\tplen = strlen(pattern);\n+\tstrbuf_read(&buf, 0, 0);\n+\n+\tif (!strcmp(pattern, \"trailer\")) {\n+\t\tsize_t payload_size = parse_signed_buffer(buf.buf, buf.len);\n+\t\tfwrite(buf.buf, 1, payload_size, stdout);\n+\t\tfflush(stdout);\n+\t\treturn 0;\n+\t}\n+\n+\tbufptr = buf.buf;\n+\ttail = bufptr + buf.len;\n+\n+\twhile (bufptr < tail) {\n+\t\t/* Find the end of the line */\n+\t\teol = memchr(bufptr, '\\n', tail - bufptr);\n+\t\tif (!eol)\n+\t\t\teol = tail;\n+\n+\t\t/* Drop continuation lines */\n+\t\tif (deleting && (bufptr < eol) && (bufptr[0] == ' ')) {\n+\t\t\tbufptr = eol + 1;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tdeleting = 0;\n+\n+\t\t/* Does the line match the prefix? */\n+\t\tif (((bufptr + plen) < eol) &&\n+\t\t    !memcmp(bufptr, pattern, plen) &&\n+\t\t    (bufptr[plen] == ' ')) {\n+\t\t\tdeleting = 1;\n+\t\t\tbufptr = eol + 1;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\t/* Print all other lines */\n+\t\tfwrite(bufptr, 1, (eol - bufptr) + 1, stdout);\n+\t\tbufptr = eol + 1;\n+\t}\n+\tfflush(stdout);\n+\n+\treturn 0;\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex abe8a785eb65..8b6c84f202d6 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -21,6 +21,7 @@ static struct test_cmd cmds[] = {\n \t{ \"csprng\", cmd__csprng },\n \t{ \"ctype\", cmd__ctype },\n \t{ \"date\", cmd__date },\n+\t{ \"delete-gpgsig\", cmd__delete_gpgsig },\n \t{ \"delta\", cmd__delta },\n \t{ \"dir-iterator\", cmd__dir_iterator },\n \t{ \"drop-caches\", cmd__drop_caches },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex ea2672436c9a..76baaece35b9 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -15,6 +15,7 @@ int cmd__csprng(int argc, const char **argv);\n int cmd__ctype(int argc, const char **argv);\n int cmd__date(int argc, const char **argv);\n int cmd__delta(int argc, const char **argv);\n+int cmd__delete_gpgsig(int argc, const char **argv);\n int cmd__dir_iterator(int argc, const char **argv);\n int cmd__drop_caches(int argc, const char **argv);\n int cmd__dump_cache_tree(int argc, const char **argv);\ndiff --git a/t/t1016-compatObjectFormat.sh b/t/t1016-compatObjectFormat.sh\nnew file mode 100755\nindex 000000000000..bb558a1d562a\n--- /dev/null\n+++ b/t/t1016-compatObjectFormat.sh\n@@ -0,0 +1,280 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2023 Eric Biederman\n+#\n+\n+test_description='Test how well compatObjectFormat works'\n+\n+. ./test-lib.sh\n+. \"$TEST_DIRECTORY\"/lib-gpg.sh\n+\n+# All of the follow variables must be defined in the environment:\n+# GIT_AUTHOR_NAME\n+# GIT_AUTHOR_EMAIL\n+# GIT_AUTHOR_DATE\n+# GIT_COMMITTER_NAME\n+# GIT_COMMITTER_EMAIL\n+# GIT_COMMITTER_DATE\n+#\n+# The test relies on these variables being set so that the two\n+# different commits in two different repositories encoded with two\n+# different hash functions result in the same content in the commits.\n+# This means that when the commit is translated between hash functions\n+# the commit is identical to the commit in the other repository.\n+\n+compat_hash () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"sha256\"\n+\t;;\n+    \"sha256\")\n+\techo \"sha1\"\n+\t;;\n+    esac\n+}\n+\n+hello_oid () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"$hello_sha1_oid\"\n+\t;;\n+    \"sha256\")\n+\techo \"$hello_sha256_oid\"\n+\t;;\n+    esac\n+}\n+\n+tree_oid () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"$tree_sha1_oid\"\n+\t;;\n+    \"sha256\")\n+\techo \"$tree_sha256_oid\"\n+\t;;\n+    esac\n+}\n+\n+commit_oid () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"$commit_sha1_oid\"\n+\t;;\n+    \"sha256\")\n+\techo \"$commit_sha256_oid\"\n+\t;;\n+    esac\n+}\n+\n+commit2_oid () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"$commit2_sha1_oid\"\n+\t;;\n+    \"sha256\")\n+\techo \"$commit2_sha256_oid\"\n+\t;;\n+    esac\n+}\n+\n+del_sigcommit () {\n+    local delete=$1\n+\n+    if test \"$delete\" = \"sha256\" ; then\n+\tlocal pattern=\"gpgsig-sha256\"\n+    else\n+\tlocal pattern=\"gpgsig\"\n+    fi\n+    test-tool delete-gpgsig \"$pattern\"\n+}\n+\n+\n+del_sigtag () {\n+    local storage=$1\n+    local delete=$2\n+\n+    if test \"$storage\" = \"$delete\" ; then\n+\tlocal pattern=\"trailer\"\n+    elif test \"$storage\" = \"sha256\" ; then\n+\tlocal pattern=\"gpgsig\"\n+    else\n+\tlocal pattern=\"gpgsig-sha256\"\n+    fi\n+    test-tool delete-gpgsig \"$pattern\"\n+}\n+\n+base=$(pwd)\n+for hash in sha1 sha256\n+do\n+\tcd \"$base\"\n+\tmkdir -p repo-$hash\n+\tcd repo-$hash\n+\n+\ttest_expect_success \"setup $hash repository\" '\n+\t\tgit init --object-format=$hash &&\n+\t\tgit config core.repositoryformatversion 1 &&\n+\t\tgit config extensions.objectformat $hash &&\n+\t\tgit config extensions.compatobjectformat $(compat_hash $hash) &&\n+\t\tgit config gpg.program $TEST_DIRECTORY/t1016/gpg &&\n+\t\techo \"Hellow World!\" > hello &&\n+\t\teval hello_${hash}_oid=$(git hash-object hello) &&\n+\t\tgit update-index --add hello &&\n+\t\tgit commit -m \"Initial commit\" &&\n+\t\teval commit_${hash}_oid=$(git rev-parse HEAD) &&\n+\t\teval tree_${hash}_oid=$(git rev-parse HEAD^{tree})\n+\t'\n+\ttest_expect_success \"create a $hash  tagged blob\" '\n+\t\tgit tag --no-sign -m \"This is a tag\" hellotag $(hello_oid $hash) &&\n+\t\teval hellotag_${hash}_oid=$(git rev-parse hellotag)\n+\t'\n+\ttest_expect_success \"create a $hash tagged tree\" '\n+\t\tgit tag --no-sign -m \"This is a tag\" treetag $(tree_oid $hash) &&\n+\t\teval treetag_${hash}_oid=$(git rev-parse treetag)\n+\t'\n+\ttest_expect_success \"create a $hash tagged commit\" '\n+\t\tgit tag --no-sign -m \"This is a tag\" committag $(commit_oid $hash) &&\n+\t\teval committag_${hash}_oid=$(git rev-parse committag)\n+\t'\n+\ttest_expect_success GPG2 \"create a $hash signed commit\" '\n+\t\tgit commit --gpg-sign --allow-empty -m \"This is a signed commit\" &&\n+\t\teval signedcommit_${hash}_oid=$(git rev-parse HEAD)\n+\t'\n+\ttest_expect_success GPG2 \"create a $hash signed tag\" '\n+\t\tgit tag -s -m \"This is a signed tag\" signedtag HEAD &&\n+\t\teval signedtag_${hash}_oid=$(git rev-parse signedtag)\n+\t'\n+\ttest_expect_success \"create a $hash branch\" '\n+\t\tgit checkout -b branch $(commit_oid $hash) &&\n+\t\techo \"More more more give me more!\" > more &&\n+\t\teval more_${hash}_oid=$(git hash-object more) &&\n+\t\techo \"Another and another and another\" > another &&\n+\t\teval another_${hash}_oid=$(git hash-object another) &&\n+\t\tgit update-index --add more another &&\n+\t\tgit commit -m \"Add more files!\" &&\n+\t\teval commit2_${hash}_oid=$(git rev-parse HEAD) &&\n+\t\teval tree2_${hash}_oid=$(git rev-parse HEAD^{tree})\n+\t'\n+\ttest_expect_success GPG2 \"create another $hash signed tag\" '\n+\t\tgit tag -s -m \"This is another signed tag\" signedtag2 $(commit2_oid $hash) &&\n+\t\teval signedtag2_${hash}_oid=$(git rev-parse signedtag2)\n+\t'\n+\ttest_expect_success GPG2 \"merge the $hash branches together\" '\n+\t\tgit merge -S -m \"merge some signed tags together\" signedtag signedtag2 &&\n+\t\teval signedcommit2_${hash}_oid=$(git rev-parse HEAD)\n+\t'\n+\ttest_expect_success GPG2 \"create additional $hash signed commits\" '\n+\t\tgit commit --gpg-sign --allow-empty -m \"This is an additional signed commit\" &&\n+\t\tgit cat-file commit HEAD | del_sigcommit sha256 > \"../${hash}_signedcommit3\" &&\n+\t\tgit cat-file commit HEAD | del_sigcommit sha1 > \"../${hash}_signedcommit4\" &&\n+\t\teval signedcommit3_${hash}_oid=$(git hash-object -t commit -w ../${hash}_signedcommit3) &&\n+\t\teval signedcommit4_${hash}_oid=$(git hash-object -t commit -w ../${hash}_signedcommit4)\n+\t'\n+\ttest_expect_success GPG2 \"create additional $hash signed tags\" '\n+\t\tgit tag -s -m \"This is an additional signed tag\" signedtag34 HEAD &&\n+\t\tgit cat-file tag signedtag34 | del_sigtag \"${hash}\" sha256 > ../${hash}_signedtag3 &&\n+\t\tgit cat-file tag signedtag34 | del_sigtag \"${hash}\" sha1 > ../${hash}_signedtag4 &&\n+\t\teval signedtag3_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag3) &&\n+\t\teval signedtag4_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag4)\n+\t'\n+done\n+cd \"$base\"\n+\n+compare_oids () {\n+    test \"$#\" = 5 && { local PREREQ=$1; shift; } || PREREQ=\n+    local type=\"$1\"\n+    local name=\"$2\"\n+    local sha1_oid=\"$3\"\n+    local sha256_oid=\"$4\"\n+\n+    echo ${sha1_oid} > ${name}_sha1_expected\n+    echo ${sha256_oid} > ${name}_sha256_expected\n+    echo ${type} > ${name}_type_expected\n+\n+    git --git-dir=repo-sha1/.git rev-parse --output-object-format=sha256 ${sha1_oid} > ${name}_sha1_sha256_found\n+    git --git-dir=repo-sha256/.git rev-parse --output-object-format=sha1 ${sha256_oid} > ${name}_sha256_sha1_found\n+    local sha1_sha256_oid=$(cat ${name}_sha1_sha256_found)\n+    local sha256_sha1_oid=$(cat ${name}_sha256_sha1_found)\n+\n+    test_expect_success $PREREQ \"Verify ${type} ${name}'s sha1 oid\" '\n+\tgit --git-dir=repo-sha256/.git rev-parse --output-object-format=sha1 ${sha256_oid} > ${name}_sha1 &&\n+\ttest_cmp ${name}_sha1 ${name}_sha1_expected\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${type} ${name}'s sha256 oid\" '\n+\tgit --git-dir=repo-sha1/.git rev-parse --output-object-format=sha256 ${sha1_oid} > ${name}_sha256 &&\n+\ttest_cmp ${name}_sha256 ${name}_sha256_expected\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha1 type\" '\n+\tgit --git-dir=repo-sha1/.git cat-file -t ${sha1_oid} > ${name}_type1 &&\n+\tgit --git-dir=repo-sha256/.git cat-file -t ${sha256_sha1_oid} > ${name}_type2 &&\n+\ttest_cmp ${name}_type1 ${name}_type2 &&\n+\ttest_cmp ${name}_type1 ${name}_type_expected\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha256 type\" '\n+\tgit --git-dir=repo-sha256/.git cat-file -t ${sha256_oid} > ${name}_type3 &&\n+\tgit --git-dir=repo-sha1/.git cat-file -t ${sha1_sha256_oid} > ${name}_type4 &&\n+\ttest_cmp ${name}_type3 ${name}_type4 &&\n+\ttest_cmp ${name}_type3 ${name}_type_expected\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha1 size\" '\n+\tgit --git-dir=repo-sha1/.git cat-file -s ${sha1_oid} > ${name}_size1 &&\n+\tgit --git-dir=repo-sha256/.git cat-file -s ${sha256_sha1_oid} > ${name}_size2 &&\n+\ttest_cmp ${name}_size1 ${name}_size2\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha256 size\" '\n+\tgit --git-dir=repo-sha256/.git cat-file -s ${sha256_oid} > ${name}_size3 &&\n+\tgit --git-dir=repo-sha1/.git cat-file -s ${sha1_sha256_oid} > ${name}_size4 &&\n+\ttest_cmp ${name}_size3 ${name}_size4\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha1 pretty content\" '\n+\tgit --git-dir=repo-sha1/.git cat-file -p ${sha1_oid} > ${name}_content1 &&\n+\tgit --git-dir=repo-sha256/.git cat-file -p ${sha256_sha1_oid} > ${name}_content2 &&\n+\ttest_cmp ${name}_content1 ${name}_content2\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha256 pretty content\" '\n+\tgit --git-dir=repo-sha256/.git cat-file -p ${sha256_oid} > ${name}_content3 &&\n+\tgit --git-dir=repo-sha1/.git cat-file -p ${sha1_sha256_oid} > ${name}_content4 &&\n+\ttest_cmp ${name}_content3 ${name}_content4\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha1 content\" '\n+\tgit --git-dir=repo-sha1/.git cat-file ${type} ${sha1_oid} > ${name}_content5 &&\n+\tgit --git-dir=repo-sha256/.git cat-file ${type} ${sha256_sha1_oid} > ${name}_content6 &&\n+\ttest_cmp ${name}_content5 ${name}_content6\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha256 content\" '\n+\tgit --git-dir=repo-sha256/.git cat-file ${type} ${sha256_oid} > ${name}_content7 &&\n+\tgit --git-dir=repo-sha1/.git cat-file ${type} ${sha1_sha256_oid} > ${name}_content8 &&\n+\ttest_cmp ${name}_content7 ${name}_content8\n+'\n+\n+}\n+\n+compare_oids 'blob' hello \"$hello_sha1_oid\" \"$hello_sha256_oid\"\n+compare_oids 'tree' tree \"$tree_sha1_oid\" \"$tree_sha256_oid\"\n+compare_oids 'commit' commit \"$commit_sha1_oid\" \"$commit_sha256_oid\"\n+compare_oids GPG2 'commit' signedcommit \"$signedcommit_sha1_oid\" \"$signedcommit_sha256_oid\"\n+compare_oids 'tag' hellotag \"$hellotag_sha1_oid\" \"$hellotag_sha256_oid\"\n+compare_oids 'tag' treetag \"$treetag_sha1_oid\" \"$treetag_sha256_oid\"\n+compare_oids 'tag' committag \"$committag_sha1_oid\" \"$committag_sha256_oid\"\n+compare_oids GPG2 'tag' signedtag \"$signedtag_sha1_oid\" \"$signedtag_sha256_oid\"\n+\n+compare_oids 'blob' more \"$more_sha1_oid\" \"$more_sha256_oid\"\n+compare_oids 'blob' another \"$another_sha1_oid\" \"$another_sha256_oid\"\n+compare_oids 'tree' tree2 \"$tree2_sha1_oid\" \"$tree2_sha256_oid\"\n+compare_oids 'commit' commit2 \"$commit2_sha1_oid\" \"$commit2_sha256_oid\"\n+compare_oids GPG2 'tag' signedtag2 \"$signedtag2_sha1_oid\" \"$signedtag2_sha256_oid\"\n+compare_oids GPG2 'commit' signedcommit2 \"$signedcommit2_sha1_oid\" \"$signedcommit2_sha256_oid\"\n+compare_oids GPG2 'commit' signedcommit3 \"$signedcommit3_sha1_oid\" \"$signedcommit3_sha256_oid\"\n+compare_oids GPG2 'commit' signedcommit4 \"$signedcommit4_sha1_oid\" \"$signedcommit4_sha256_oid\"\n+compare_oids GPG2 'tag' signedtag3 \"$signedtag3_sha1_oid\" \"$signedtag3_sha256_oid\"\n+compare_oids GPG2 'tag' signedtag4 \"$signedtag4_sha1_oid\" \"$signedtag4_sha256_oid\"\n+\n+test_done\ndiff --git a/t/t1016/gpg b/t/t1016/gpg\nnew file mode 100755\nindex 000000000000..2601cb18a5b3\n--- /dev/null\n+++ b/t/t1016/gpg\n@@ -0,0 +1,2 @@\n+#!/bin/sh\n+exec gpg --faked-system-time \"20230918T154812\" \"$@\"\n-- \n2.41.0\n\n"},{"id":"482414","messageId":"CAPig+cRshiUNXfU=ZY4nZXgBgTJ_wF0WVDxWpqkEKPAT9pjX_w@mail.gmail.com","threadId":"60273","inReplyTo":"20230927195537.1682-1-ebiederm@gmail.com","subject":"Re: [PATCH 01/30] object-file-convert: Stubs for converting from one object format to another","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2023-09-27T20:42:03Z","receivedAt":"2023-09-27T20:42:18Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Wed, Sep 27, 2023 at 3:55 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n> Two basic functions are provided:\n> - convert_object_file Takes an object file it's type and hash algorithm\n>   and converts it into the equivalent object file that would\n>   have been generated with hash algorithm \"to\".\n>\n>   For blob objects there is no converstion to be done and it is an\n>   error to use this function on them.\n\ns/converstion/conversion/\n\n>   For commit, tree, and tag objects embedded oids are replaced by the\n>   oids of the objects they refer to with those objects and their\n>   object ids reencoded in with the hash algorithm \"to\".  Signatures\n>   are rearranged so that they remain valid after the object has\n>   been reencoded.\n>\n> - repo_oid_to_algop which takes an oid that refers to an object file\n>   and returns the oid of the equavalent object file generated\n>   with the target hash algorithm.\n\ns/equavalent/equivalent/\n\n> The pair of files object-file-convert.c and object-file-convert.h are\n> introduced to hold as much of this logic as possible to keep this\n> conversion logic cleanly separated from everything else and in the\n> hopes that someday the code will be clean enough git can support\n> compiling out support for sha1 and the various conversion functions.\n>\n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nJust some minor comments below, many of which are subjective\nstyle-related observations, thus not necessarily actionable, but also\none or two legitimate questions.\n\n> diff --git a/object-file-convert.c b/object-file-convert.c\n> @@ -0,0 +1,57 @@\n> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n> +                     const struct git_hash_algo *to, struct object_id *dest)\n> +{\n> +       /*\n> +        * If the source alogirthm is not set, then we're using the\n> +        * default hash algorithm for that object.\n> +        */\n\ns/alogirthm/algorithm/\n\n> +       const struct git_hash_algo *from =\n> +               src->algo ? &hash_algos[src->algo] : repo->hash_algo;\n> +\n> +       if (from == to) {\n> +               if (src != dest)\n> +                       oidcpy(dest, src);\n> +               return 0;\n> +       }\n> +       return -1;\n> +}\n\nOn this project, we usually get the simple cases out of the way first,\nwhich often reduces the indentation level, making the code easier to\ndigest at a glance. So, it would be typical to write this as:\n\n    if (from != to)\n        return -1\n    if (src != dest)\n        oidcpy(dest, src);\n    return 0;\n\nor even:\n\n    if (from != to)\n        return -1\n    if (src == dest)\n        return 0;\n    oidcpy(dest, src);\n    return 0;\n\nThis way, for instance, the reader doesn't get to the end of the\nfunction and then have to scan backward to understand the condition of\nthe `return -1`.\n\n> +int convert_object_file(struct strbuf *outbuf,\n> +                       const struct git_hash_algo *from,\n> +                       const struct git_hash_algo *to,\n> +                       const void *buf, size_t len,\n> +                       enum object_type type,\n> +                       int gentle)\n> +{\n> +       int ret;\n> +\n> +       /* Don't call this function when no conversion is necessary */\n> +       if ((from == to) || (type == OBJ_BLOB))\n> +               die(\"Refusing noop object file conversion\");\n\nSeveral comments...\n\nStyle: we usually reduce the noise level by dropping the extra parentheses:\n\n    if (from == to || type == OBJ_BLOB)\n\nDoes this condition represent a programming error or a runtime error\ntriggerable by some input? If a programming error, then use BUG()\nrather than die().\n\nIf a triggerable runtime error, then...\n\n* start user-facing messages with lowercase rather than capitalized word\n\n* make the user-facing message localizable so readers of other\nlanguages can digest it\n\n    die(_(\"refusing do-nothing object conversion\"));\n\nOn the other hand, don't make BUG() messages localizable.\n\n> +       switch (type) {\n> +       case OBJ_COMMIT:\n> +       case OBJ_TREE:\n> +       case OBJ_TAG:\n> +       default:\n> +               /* Not implemented yet, so fail. */\n> +               ret = -1;\n> +               break;\n> +       }\n> +       if (!ret)\n> +               return 0;\n> +       if (gentle) {\n> +               strbuf_release(outbuf);\n> +               return ret;\n> +       }\n\nThis function appears to be a mere skeleton at the moment, so it's\ndifficult to judge at this point whether you are using `outbuf` as a\nbag of bytes or as a legitimate string container. If the latter, then\nthe API may be reasonable, but if you're using it as a bag-of-bytes,\nthen it feels like you're leaking an implementation detail into the\nAPI.\n\n> +       die(_(\"Failed to convert object from %s to %s\"),\n> +               from->name, to->name);\n\ns/Failed/failed/\n\nFor people trying to diagnose this problem, would it be helpful to\npresent more information about the failed conversion, such as object\ntype and perhaps even its OID?\n\n> diff --git a/object-file-convert.h b/object-file-convert.h\n> @@ -0,0 +1,24 @@\n> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n> +                     const struct git_hash_algo *to, struct object_id *dest);\n\nI suppose the function name is pretty much self-explanatory to those\nfamiliar with the underlying concepts, but it might still be helpful\nto add a comment explaining what the function does.\n\n> +/*\n> + * Convert an object file from one hash algorithm to another algorithm.\n> + * Return -1 on failure, 0 on success.\n> + */\n> +int convert_object_file(struct strbuf *outbuf,\n> +                       const struct git_hash_algo *from,\n> +                       const struct git_hash_algo *to,\n> +                       const void *buf, size_t len,\n> +                       enum object_type type,\n> +                       int gentle);\n"},{"id":"482415","messageId":"xmqqjzsbl2v9.fsf@gitster.g","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"Re: [PATCH 00/30] Initial support for multiple hash functions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-27T21:31:38Z","receivedAt":"2023-09-27T21:31:54Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> I have been going over and over this patchset trying to figure\n> out if it is ready to be merged.  I don't know of any deficiencies\n> so it is at a point it could benefit from a set of eyes that\n> are not mine.\n>\n> I had planned to wait a little bit longer but there are some on-going\n> conversations that could benefit from people seeing what it means for a\n> repository to support two hash functions at the same time.\n\nThanks.  On top of what commit are these patches expected to be applied?\n"},{"id":"482416","messageId":"xmqqfs2zl2iy.fsf@gitster.g","threadId":"60273","inReplyTo":"20230927195537.1682-21-ebiederm@gmail.com","subject":"Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-27T21:39:01Z","receivedAt":"2023-09-27T21:39:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> diff --git a/setup.c b/setup.c\n> index deb5a33fe9e1..87b40472dbc5 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -598,6 +598,25 @@ static enum extension_result handle_extension(const char *var,\n>  \t\t}\n\nThis line in the pre-context needed fuzzing, but otherwise the\nseries applied cleanly on top of v2.42.0.\n\n> Subject: Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat\n\n\"Implement\" -> \"implement\" (many other patches share the same\nproblem, none of which I fixed while queueing).\n\n"},{"id":"482417","messageId":"CAPig+cR5mGZ7-4t1YBW-=j3FyWGvBRBN7eogQb1BYiw8QM1UKA@mail.gmail.com","threadId":"60273","inReplyTo":"20230927195537.1682-2-ebiederm@gmail.com","subject":"Re: [PATCH 02/30] oid-array: Teach oid-array to handle multiple kinds of oids","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2023-09-27T23:20:27Z","receivedAt":"2023-09-27T23:20:47Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Wed, Sep 27, 2023 at 3:55 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n> While looking at how to handle input of both SHA-1 and SHA-256 oids in\n> get_oid_with_context, I realized that the oid_array in\n> repo_for_each_abbrev might have more than one kind of oid stored in it\n> simulataneously.\n\ns/simulataneously/simultaneously/\n\n> Update to oid_array_append to ensure that oids added to an oid array\n> always have an algorithm set.\n>\n> Update void_hashcmp to first verify two oids use the same hash algorithm\n> before comparing them to each other.\n>\n> With that oid-array should be safe to use with differnt kinds of\n> oids simultaneously.\n\ns/differnt/different/\n\n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n"},{"id":"482418","messageId":"CAPig+cRHBxPZYQ5XYA5Un7LeS21NgqxZGg=Q8D+aQckrw9Ymtg@mail.gmail.com","threadId":"60273","inReplyTo":"20230927195537.1682-3-ebiederm@gmail.com","subject":"Re: [PATCH 03/30] object-names: Support input of oids in any supported hash","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2023-09-27T23:29:53Z","receivedAt":"2023-09-27T23:30:09Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Wed, Sep 27, 2023 at 3:56 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n> Support short oids encoded in any algorithm, while ensuring enough of\n> the oid is specified to disambiguate between all of the oids in the\n> repository encoded in any algorithm.\n>\n> By default have the code continue to only accept oids specified in the\n> storage hash algorithm of the repository, but when something is\n> ambiguous display all of the possible oids from any oid encoding.\n>\n> A new flag is added GET_OID_HASH_ANY that when supplied causes the\n> code to accept oids specified in any hash algorithm, and to return the\n> oids that were resolved.\n>\n> This implements the functionality that allows both SHA-1 and SHA-256\n> object names, from the \"Object names on the command line\" section of\n> the hash function transition document.\n>\n> Care is taken in get_short_oid so that when the result is ambiguous\n> the output remains the same of GIT_OID_HASH_ANY was not supplied.\n\ns/of/as if/\n\n> If GET_OID_HASH_ANY was supplied objects of any hash algorithm\n> that match the prefix are displayed.\n>\n> This required updating repo_for_each_abbrev to give it a parameter\n> so that it knows to look at all hash algorithms.\n>\n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n> diff --git a/object-name.c b/object-name.c\n> @@ -49,6 +50,7 @@ struct disambiguate_state {\n>  static void update_candidates(struct disambiguate_state *ds, const struct object_id *current)\n>  {\n> +       /* The hash algorithm of the current has already been filtered */\n\nIs there a word missing after \"current\"?\n\n> @@ -503,8 +516,13 @@ static int sort_ambiguous(const void *a, const void *b, void *ctx)\n> -       if (a_type == b_type)\n> -               return oidcmp(a, b);\n> +       if (a_type == b_type) {\n> +               /* Is the hash algorithm the same? */\n> +               if (a->algo == b->algo)\n> +                       return oidcmp(a, b);\n> +               else\n> +                       return a->algo > b->algo ? 1 : -1;\n> +       }\n\nNit: unnecessary comment (\"Is the hash algorithm...\") is merely\nrepeating what the code itself already says clearly enough\n\n> @@ -553,6 +575,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>         else\n>                 ds.fn = default_disambiguate_hint;\n>\n> +\n>         find_short_object_filename(&ds);\n\nNit: unnecessary new blank line\n"},{"id":"482427","messageId":"CAPig+cS02ushqgw+u39Tmnoy3rgp8BzqT4T9D=-01m5fsLxC6Q@mail.gmail.com","threadId":"60273","inReplyTo":"20230927195537.1682-5-ebiederm@gmail.com","subject":"Re: [PATCH 05/30] loose: add a mapping between SHA-1 and SHA-256 for loose objects","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2023-09-28T07:14:31Z","receivedAt":"2023-09-28T07:15:29Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Wed, Sep 27, 2023 at 3:56 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n> As part of the transition plan, we'd like to add a file in the .git\n> directory that maps loose objects between SHA-1 and SHA-256.  Let's\n> implement the specification in the transition plan and store this data\n> on a per-repository basis in struct repository.\n>\n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> Signed-off-by: Eric W. Biederman <ebiederm@xmission.com>\n> ---\n> diff --git a/loose.c b/loose.c\n> @@ -0,0 +1,245 @@\n> +static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n> +{\n> +       struct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n> +       FILE *fp;\n> +\n> +       if (!dir->loose_map)\n> +               loose_object_map_init(&dir->loose_map);\n> +\n> +       insert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n> +       insert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n> +\n> +       insert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n> +       insert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n> +\n> +       insert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n> +       insert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n> +\n> +       strbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n> +       fp = fopen(path.buf, \"rb\");\n> +       if (!fp)\n> +               return 0;\n\nThis early return leaks `path`. At minimum, call\n`strbuf_release(&path)` before returning.\n\n> +       errno = 0;\n> +       if (strbuf_getwholeline(&buf, fp, '\\n') || strcmp(buf.buf, loose_object_header))\n> +               goto err;\n> +       while (!strbuf_getline_lf(&buf, fp)) {\n> +               const char *p;\n> +               struct object_id oid, compat_oid;\n> +               if (parse_oid_hex_algop(buf.buf, &oid, &p, repo->hash_algo) ||\n> +                   *p++ != ' ' ||\n> +                   parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n> +                   p != buf.buf + buf.len)\n> +                       goto err;\n> +               insert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n> +               insert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n> +       }\n> +\n> +       strbuf_release(&buf);\n> +       strbuf_release(&path);\n> +       return errno ? -1 : 0;\n> +err:\n> +       strbuf_release(&buf);\n> +       strbuf_release(&path);\n> +       return -1;\n> +}\n> +\n> +int repo_write_loose_object_map(struct repository *repo)\n> +{\n> +       kh_oid_map_t *map = repo->objects->odb->loose_map->to_compat;\n> +       struct lock_file lock;\n> +       int fd;\n> +       khiter_t iter;\n> +       struct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n> +\n> +       if (!should_use_loose_object_map(repo))\n> +               return 0;\n> +\n> +       strbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n> +       fd = hold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n> +       iter = kh_begin(map);\n> +       if (write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n> +               goto errout;\n> +\n> +       for (; iter != kh_end(map); iter++) {\n> +               if (kh_exist(map, iter)) {\n> +                       if (oideq(&kh_key(map, iter), the_hash_algo->empty_tree) ||\n> +                           oideq(&kh_key(map, iter), the_hash_algo->empty_blob))\n> +                               continue;\n> +                       strbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(&kh_key(map, iter)), oid_to_hex(kh_value(map, iter)));\n> +                       if (write_in_full(fd, buf.buf, buf.len) < 0)\n> +                               goto errout;\n> +                       strbuf_reset(&buf);\n> +               }\n> +       }\n\nNit: If you call strbuf_reset() immediately before strbuf_addf(), then\nyou save the reader of the code the effort of having to scan backward\nthrough the function to verify that `buf` is empty the first time\nthrough the loop.\n\n> +       strbuf_release(&buf);\n> +       if (commit_lock_file(&lock) < 0) {\n> +               error_errno(_(\"could not write loose object index %s\"), path.buf);\n> +               strbuf_release(&path);\n> +               return -1;\n> +       }\n> +       strbuf_release(&path);\n> +       return 0;\n> +errout:\n> +       rollback_lock_file(&lock);\n> +       strbuf_release(&buf);\n> +       error_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n> +       strbuf_release(&path);\n> +       return -1;\n> +}\n> +\n> +int repo_loose_object_map_oid(struct repository *repo,\n> +                             const struct object_id *src,\n> +                             const struct git_hash_algo *to,\n> +                             struct object_id *dest)\n> +\n> +{\n\nStyle: unnecessary blank line before opening `{`\n\n> +       struct object_directory *dir;\n> +       kh_oid_map_t *map;\n> +       khiter_t pos;\n> +\n> +       for (dir = repo->objects->odb; dir; dir = dir->next) {\n> +               struct loose_object_map *loose_map = dir->loose_map;\n> +               if (!loose_map)\n> +                       continue;\n> +               map = (to == repo->compat_hash_algo) ?\n> +                       loose_map->to_compat :\n> +                       loose_map->to_storage;\n> +               pos = kh_get_oid_map(map, *src);\n> +               if (pos < kh_end(map)) {\n> +                       oidcpy(dest, kh_value(map, pos));\n> +                       return 0;\n> +               }\n> +       }\n> +       return -1;\n> +}\n"},{"id":"482438","messageId":"xmqqbkdmjbkp.fsf@gitster.g","threadId":"60273","inReplyTo":"xmqqfs2zl2iy.fsf@gitster.g","subject":"Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-28T20:18:46Z","receivedAt":"2023-09-28T20:19:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n>\n>> diff --git a/setup.c b/setup.c\n>> index deb5a33fe9e1..87b40472dbc5 100644\n>> --- a/setup.c\n>> +++ b/setup.c\n>> @@ -598,6 +598,25 @@ static enum extension_result handle_extension(const char *var,\n>>  \t\t}\n>\n> This line in the pre-context needed fuzzing, but otherwise the\n> series applied cleanly on top of v2.42.0.\n>\n>> Subject: Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat\n>\n> \"Implement\" -> \"implement\" (many other patches share the same\n> problem, none of which I fixed while queueing).\n\n\nThe topic when merged near the tip of 'seen' seems to break a few CI\njobs here and there.  The log from the broken run can be seen at\n\n    https://github.com/git/git/actions/runs/6331978214\n\nYou may have to log-in there before you can view the details.\n\nThanks.\n\n"},{"id":"482446","messageId":"409703AB-9349-4AF4-ADE8-750EFA768074@gmail.com","threadId":"60273","inReplyTo":"xmqqbkdmjbkp.fsf@gitster.g","subject":"Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat","fromName":"Eric Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-29T00:50:56Z","receivedAt":"2023-09-29T00:51:12Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"\n\nOn September 28, 2023 3:18:46 PM CDT, Junio C Hamano <gitster@pobox.com> wrote:\n>Junio C Hamano <gitster@pobox.com> writes:\n>\n>> \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n>>\n>>> diff --git a/setup.c b/setup.c\n>>> index deb5a33fe9e1..87b40472dbc5 100644\n>>> --- a/setup.c\n>>> +++ b/setup.c\n>>> @@ -598,6 +598,25 @@ static enum extension_result handle_extension(const char *var,\n>>>  \t\t}\n>>\n>> This line in the pre-context needed fuzzing, but otherwise the\n>> series applied cleanly on top of v2.42.0.\n>>\n>>> Subject: Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat\n>>\n>> \"Implement\" -> \"implement\" (many other patches share the same\n>> problem, none of which I fixed while queueing).\n>\n>\n>The topic when merged near the tip of 'seen' seems to break a few CI\n>jobs here and there.  The log from the broken run can be seen at\n>\n>    https://github.com/git/git/actions/runs/6331978214\n>\n>You may have to log-in there before you can view the details.\n\nThanks.\n\nIt might take me a couple of days before I can dig into this, but I will dig in and see if I can understand and fix the build failures.\n\nWith any luck it will be something simple like forgetting that {} != {0}.\n\nEric \n"},{"id":"482451","messageId":"87bkdkhq4s.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"xmqqbkdmjbkp.fsf@gitster.g","subject":"Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-09-29T16:59:31Z","receivedAt":"2023-09-29T16:59:43Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n>>\n>>> diff --git a/setup.c b/setup.c\n>>> index deb5a33fe9e1..87b40472dbc5 100644\n>>> --- a/setup.c\n>>> +++ b/setup.c\n>>> @@ -598,6 +598,25 @@ static enum extension_result handle_extension(const char *var,\n>>>  \t\t}\n>>\n>> This line in the pre-context needed fuzzing, but otherwise the\n>> series applied cleanly on top of v2.42.0.\n>>\n>>> Subject: Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat\n>>\n>> \"Implement\" -> \"implement\" (many other patches share the same\n>> problem, none of which I fixed while queueing).\n>\n>\n> The topic when merged near the tip of 'seen' seems to break a few CI\n> jobs here and there.  The log from the broken run can be seen at\n>\n>     https://github.com/git/git/actions/runs/6331978214\n>\n> You may have to log-in there before you can view the details.\n\nDid you have any manual merge conflicts you had to resolve?\nIf so it is possible to see the merge result you had?\n\nThere is a static failure in commit.c of oidcpy because it thinks the\narray is zero size.  That is weird, but once I get a test environment\nsetup I expect I can figure out what it is talking about.\n\nThere in linux-leaks it lists a bunch of test failures, and unless I see\nwhat code is actually failing I am not certain I can figure it out.\n\nThanks,\nEric\n\n"},{"id":"482452","messageId":"xmqqedigizms.fsf@gitster.g","threadId":"60273","inReplyTo":"87bkdkhq4s.fsf@gmail.froward.int.ebiederm.org","subject":"Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-29T18:48:59Z","receivedAt":"2023-09-29T18:49:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> Did you have any manual merge conflicts you had to resolve?\n> If so it is possible to see the merge result you had?\n\nThe only merge-fix I had to apply to make everything compile was\nthis:\n\ndiff --git a/bloom.c b/bloom.c\nindex ff131893cd..59eb0a0481 100644\n--- a/bloom.c\n+++ b/bloom.c\n@@ -278,7 +278,7 @@ static int has_entries_with_high_bit(struct repository *r, struct tree *t)\n \t\tstruct tree_desc desc;\n \t\tstruct name_entry entry;\n \n-\t\tinit_tree_desc(&desc, t->buffer, t->size);\n+\t\tinit_tree_desc(&desc, &t->object.oid, t->buffer, t->size);\n \t\twhile (tree_entry(&desc, &entry)) {\n \t\t\tsize_t i;\n \t\t\tfor (i = 0; i < entry.pathlen; i++) {\n\nas one topic changed the function signature while the other topic\nadded a new callsite.\n\nEverything else was pretty-much auto resolved, I think.\n\nOutput from \"git show --cc seen\" matches my recollection.  The above\ndoes appear as an evil merge.\n\n"},{"id":"482479","messageId":"87r0mdetmv.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"xmqqedigizms.fsf@gitster.g","subject":"Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T00:48:56Z","receivedAt":"2023-10-02T00:49:03Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n>\n>> Did you have any manual merge conflicts you had to resolve?\n>> If so it is possible to see the merge result you had?\n>\n> The only merge-fix I had to apply to make everything compile was\n> this:\n>\n> diff --git a/bloom.c b/bloom.c\n> index ff131893cd..59eb0a0481 100644\n> --- a/bloom.c\n> +++ b/bloom.c\n> @@ -278,7 +278,7 @@ static int has_entries_with_high_bit(struct repository *r, struct tree *t)\n>  \t\tstruct tree_desc desc;\n>  \t\tstruct name_entry entry;\n>  \n> -\t\tinit_tree_desc(&desc, t->buffer, t->size);\n> +\t\tinit_tree_desc(&desc, &t->object.oid, t->buffer, t->size);\n>  \t\twhile (tree_entry(&desc, &entry)) {\n>  \t\t\tsize_t i;\n>  \t\t\tfor (i = 0; i < entry.pathlen; i++) {\n>\n> as one topic changed the function signature while the other topic\n> added a new callsite.\n>\n> Everything else was pretty-much auto resolved, I think.\n>\n> Output from \"git show --cc seen\" matches my recollection.  The above\n> does appear as an evil merge.\n\nThanks, and I found all of this on your seen branch.\n\nAfter tracking all of these down it appears all of the errors\ncame from my branch, I will be resending the patches as soon\nas I finish going through the review comments.\n\n\nLooking at the build errors pretty much all of the all of the\nautomatic test failures came from commit_tree_extended.\n\n\nThere was a strbuf that did\nstrbuf_init(&buf, 8192);\nstrbuf_init(&buf, 8192);\ntwice.\n\nPlus there was another buffer that was allocated and not freed,\nin commit_tree_extended.\n\n\nThe leaks were a bit tricky to track down as building with SANITIZE=leak\ncauses tests to fail somewhat randomly for me with \"gcc (Debian\n12.2.0-14) 12.2.0\".\n\n\nThere was one smatch static-analysis error that suggested using\nCALLOC_ARRAY instead of xcalloc.\n\n\n\nThe \"win\" build and \"linux-gcc-default (ubuntu-lastest)\" build failed\nbecause of an over eager gcc warning -Werror=array-bounds.\nClaiming:\n\n In file included from /usr/include/string.h:535,\n                  from git-compat-util.h:228,\n                  from commit.c:1:\n In function ‘memcpy’,\n     inlined from ‘oidcpy’ at hash-ll.h:272:2,\n     inlined from ‘commit_tree_extended’ at commit.c:1705:3:\n ##[error]/usr/include/x86_64-linux-gnu/bits/string_fortified.h:29:10: ‘__builtin_memcpy’ offset [0, 31] is out of the bounds [0, 0] [-Werror=array-bounds]\n    29 |   return __builtin___memcpy_chk (__dest, __src, __len,\n       |          ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n    30 |                                  __glibc_objsize0 (__dest));\n       |                                  ~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nPreviously in commit_tree_extended the structure was:\n\n\twhile (parents) {\n\t\tstruct commit *parent = pop_commit(&parents);\n\t\tstrbuf_addf(&buffer, \"parent %s\\n\",\n\t\t\t    oid_to_hex(&parent->object.oid));\n\t}\n\nbrian had changed it to:\n\n\tnparents = commit_list_count(parents);\n\tparent_buf = xcalloc(nparents, sizeof(*parent_buf));\n\tfor (i = 0; i < nparents; i++) {\n\t\tstruct commit *parent = pop_commit(&parents);\n\t\toidcpy(&parent_buf[i], &parent->object.oid);\n\t}\n\nWhich is perfectly sound code.\n\nI changed the structure of the loop to:\n\n\tnparents = commit_list_count(parents);\n\tparent_buf = xcalloc(nparents, sizeof(*parent_buf));\n\ti = 0;\n\twhile (parents) {\n\t\tstruct commit *parent = pop_commit(&parents);\n\t\toidcpy(&parent_buf[i++], &parent->object.oid);\n\t}\n\nAnd the \"array-bounds\" warning had no problems with the code.\nSo it looks like the error was actually that array-bounds thought\nthere was a potential NULL pointer dereference at which point\nit would not have array bounds, and then it complained about\nthe array bounds, instead of the NULL pointer dereference.\n\nI am going to fix the patch.  If array-bounds causes further\nproblems you may want to think about disabling it.\n\nEric\n\n"},{"id":"482480","messageId":"87o7hhbyxi.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"CAPig+cRshiUNXfU=ZY4nZXgBgTJ_wF0WVDxWpqkEKPAT9pjX_w@mail.gmail.com","subject":"Re: [PATCH 01/30] object-file-convert: Stubs for converting from one object format to another","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T01:22:49Z","receivedAt":"2023-10-02T01:23:43Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"Eric Sunshine <sunshine@sunshineco.com> writes:\n\n> On Wed, Sep 27, 2023 at 3:55 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n>> Two basic functions are provided:\n>> - convert_object_file Takes an object file it's type and hash algorithm\n>>   and converts it into the equivalent object file that would\n>>   have been generated with hash algorithm \"to\".\n>>\n>>   For blob objects there is no converstion to be done and it is an\n>>   error to use this function on them.\n>\n> s/converstion/conversion/\n>\n>>   For commit, tree, and tag objects embedded oids are replaced by the\n>>   oids of the objects they refer to with those objects and their\n>>   object ids reencoded in with the hash algorithm \"to\".  Signatures\n>>   are rearranged so that they remain valid after the object has\n>>   been reencoded.\n>>\n>> - repo_oid_to_algop which takes an oid that refers to an object file\n>>   and returns the oid of the equavalent object file generated\n>>   with the target hash algorithm.\n>\n> s/equavalent/equivalent/\n>\n>> The pair of files object-file-convert.c and object-file-convert.h are\n>> introduced to hold as much of this logic as possible to keep this\n>> conversion logic cleanly separated from everything else and in the\n>> hopes that someday the code will be clean enough git can support\n>> compiling out support for sha1 and the various conversion functions.\n>>\n>> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>\n> Just some minor comments below, many of which are subjective\n> style-related observations, thus not necessarily actionable, but also\n> one or two legitimate questions.\n>\n>> diff --git a/object-file-convert.c b/object-file-convert.c\n>> @@ -0,0 +1,57 @@\n>> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n>> +                     const struct git_hash_algo *to, struct object_id *dest)\n>> +{\n>> +       /*\n>> +        * If the source alogirthm is not set, then we're using the\n>> +        * default hash algorithm for that object.\n>> +        */\n>\n> s/alogirthm/algorithm/\n>\n>> +       const struct git_hash_algo *from =\n>> +               src->algo ? &hash_algos[src->algo] : repo->hash_algo;\n>> +\n>> +       if (from == to) {\n>> +               if (src != dest)\n>> +                       oidcpy(dest, src);\n>> +               return 0;\n>> +       }\n>> +       return -1;\n>> +}\n>\n> On this project, we usually get the simple cases out of the way first,\n> which often reduces the indentation level, making the code easier to\n> digest at a glance. So, it would be typical to write this as:\n>\n>     if (from != to)\n>         return -1\n>     if (src != dest)\n>         oidcpy(dest, src);\n>     return 0;\n>\n> or even:\n>\n>     if (from != to)\n>         return -1\n>     if (src == dest)\n>         return 0;\n>     oidcpy(dest, src);\n>     return 0;\n>\n> This way, for instance, the reader doesn't get to the end of the\n> function and then have to scan backward to understand the condition of\n> the `return -1`.\n\nThe \"return -1\" is there only because it is a stub, and it is there\nwhere the rest of the code needs to go.\n\nAs for simple cases the \"if (from == to)\" case is a simple case I am\ngetting out of the way.  It unfortunately is cluttered by the fact\nthat \"oidcpy(&oid, &oid)\" is not valid so it has to guard the copy.\n\nif (from == to) {\n\tif (src == dest)\n        \treturn 0;\n        oidcpy(dest, src);\n        return 0;\n}\n\nCould be used but is wordier.  And duplicates the return code\nfor the same case so I am not enthusiastic about it.\n\n\n>> +int convert_object_file(struct strbuf *outbuf,\n>> +                       const struct git_hash_algo *from,\n>> +                       const struct git_hash_algo *to,\n>> +                       const void *buf, size_t len,\n>> +                       enum object_type type,\n>> +                       int gentle)\n>> +{\n>> +       int ret;\n>> +\n>> +       /* Don't call this function when no conversion is necessary */\n>> +       if ((from == to) || (type == OBJ_BLOB))\n>> +               die(\"Refusing noop object file conversion\");\n>\n> Several comments...\n>\n> Style: we usually reduce the noise level by dropping the extra parentheses:\n>\n>     if (from == to || type == OBJ_BLOB)\n\nI honestly can not be confident of C code that does that.\n\nThe precedence of the operators in C has been wrong for longer than I\nhave been programming, and I can never remember exactly how the\nprecedence is wrong.  So for the last 30 years I have been adding enough\nparenthesis that I don't have to remember.\n\n> Does this condition represent a programming error or a runtime error\n> triggerable by some input? If a programming error, then use BUG()\n> rather than die().\n\nAgreed BUG would be better there.\n\n> If a triggerable runtime error, then...\n>\n> * start user-facing messages with lowercase rather than capitalized word\n>\n> * make the user-facing message localizable so readers of other\n> languages can digest it\n>\n>     die(_(\"refusing do-nothing object conversion\"));\n>\n> On the other hand, don't make BUG() messages localizable.\n>\n>> +       switch (type) {\n>> +       case OBJ_COMMIT:\n>> +       case OBJ_TREE:\n>> +       case OBJ_TAG:\n>> +       default:\n>> +               /* Not implemented yet, so fail. */\n>> +               ret = -1;\n>> +               break;\n>> +       }\n>> +       if (!ret)\n>> +               return 0;\n>> +       if (gentle) {\n>> +               strbuf_release(outbuf);\n>> +               return ret;\n>> +       }\n>\n> This function appears to be a mere skeleton at the moment, so it's\n> difficult to judge at this point whether you are using `outbuf` as a\n> bag of bytes or as a legitimate string container. If the latter, then\n> the API may be reasonable, but if you're using it as a bag-of-bytes,\n> then it feels like you're leaking an implementation detail into the\n> API.\n\nIt is a string that represents the entire object.\n\nIt might be arguable if tree objects are text given that trees represent\noids in binary, but for tag and commit objects they are definitely one\nbig text string.\n\nThis is an implementation detail that makes the code simpler, and less\nerror prone.\n\nI have not encountered anything where a string buffer would not be\na reasonable fit.\n\n>> +       die(_(\"Failed to convert object from %s to %s\"),\n>> +               from->name, to->name);\n>\n> s/Failed/failed/\n\nI don't understand wanting to start a sentence with a lower case letter.\nCan you explain?\n\n> For people trying to diagnose this problem, would it be helpful to\n> present more information about the failed conversion, such as object\n> type and perhaps even its OID?\n\nI expect some of that context will come from the conversion functions\nthemselves.  Those messages are almost certain to give the type\ninformation one way or another because they are type specific.\n\nThat said it requires some version of corrupt repository to reach this\nerror.  Either missing mapping tables, or a corrupt object.\n\nSo I don't know how much it matters to get this perfect the first\ntime.  We can improve the error message we gain experience.\n\nI would agree that including the OID makes sense.  Unfortunately the\ncode does not have the OID at this point.\n\n>> diff --git a/object-file-convert.h b/object-file-convert.h\n>> @@ -0,0 +1,24 @@\n>> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n>> +                     const struct git_hash_algo *to, struct object_id *dest);\n>\n> I suppose the function name is pretty much self-explanatory to those\n> familiar with the underlying concepts, but it might still be helpful\n> to add a comment explaining what the function does.\n\nI could use words that repeat what is in the function signature.\nBut I don't think I could add anything.\n\nI would have to say something like:\n\nLook up the oid that an equivalent object would have in a repository\nwhose object format is \"to\".\n\nIs that helpful?\n\nEric\n"},{"id":"482481","messageId":"8734ytbyiz.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"xmqqfs2zl2iy.fsf@gitster.g","subject":"Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T01:31:32Z","receivedAt":"2023-10-02T01:33:56Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n>\n>> Subject: Re: [PATCH 21/30] repository: Implement extensions.compatObjectFormat\n>\n> \"Implement\" -> \"implement\" (many other patches share the same\n> problem, none of which I fixed while queueing).\n\nWhy shouldn't sentences begin with a capital letter?\n\nI agree it isn't a very extensive sentence just two words but it is a\ncomplete thought, and thus is a sentence.\n\nThere is great value in uniformity in a project so I will make the\nchange, but it seems very weird to me.\n\nEric\n"},{"id":"482482","messageId":"87il7paiwa.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"CAPig+cRHBxPZYQ5XYA5Un7LeS21NgqxZGg=Q8D+aQckrw9Ymtg@mail.gmail.com","subject":"Re: [PATCH 03/30] object-names: Support input of oids in any supported hash","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T01:54:29Z","receivedAt":"2023-10-02T01:57:23Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"Eric Sunshine <sunshine@sunshineco.com> writes:\n\n> On Wed, Sep 27, 2023 at 3:56 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n>> Support short oids encoded in any algorithm, while ensuring enough of\n>> the oid is specified to disambiguate between all of the oids in the\n>> repository encoded in any algorithm.\n>>\n>> By default have the code continue to only accept oids specified in the\n>> storage hash algorithm of the repository, but when something is\n>> ambiguous display all of the possible oids from any oid encoding.\n>>\n>> A new flag is added GET_OID_HASH_ANY that when supplied causes the\n>> code to accept oids specified in any hash algorithm, and to return the\n>> oids that were resolved.\n>>\n>> This implements the functionality that allows both SHA-1 and SHA-256\n>> object names, from the \"Object names on the command line\" section of\n>> the hash function transition document.\n>>\n>> Care is taken in get_short_oid so that when the result is ambiguous\n>> the output remains the same of GIT_OID_HASH_ANY was not supplied.\n>\n> s/of/as if/\n\nActually s/of/if/\n\nThank you for catching that.  When reviewing this to understand what I\nwas trying to say I found a bug.  The function repo_for_each_abbrev was\npassing algo to init_object_disambiguation to properly initialize\nds.bin_pfx.algo, and then was resetting ds.bin_pfx.algo to GIT_HASH_ANY.\nOops!\n\n>> If GET_OID_HASH_ANY was supplied objects of any hash algorithm\n>> that match the prefix are displayed.\n>>\n>> This required updating repo_for_each_abbrev to give it a parameter\n>> so that it knows to look at all hash algorithms.\n>>\n>> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>> ---\n>> diff --git a/object-name.c b/object-name.c\n>> @@ -49,6 +50,7 @@ struct disambiguate_state {\n>>  static void update_candidates(struct disambiguate_state *ds, const struct object_id *current)\n>>  {\n>> +       /* The hash algorithm of the current has already been filtered */\n>\n> Is there a word missing after \"current\"?\n\nNo.  I forgot to remove \"the\" before current.\n\n>> @@ -503,8 +516,13 @@ static int sort_ambiguous(const void *a, const void *b, void *ctx)\n>> -       if (a_type == b_type)\n>> -               return oidcmp(a, b);\n>> +       if (a_type == b_type) {\n>> +               /* Is the hash algorithm the same? */\n>> +               if (a->algo == b->algo)\n>> +                       return oidcmp(a, b);\n>> +               else\n>> +                       return a->algo > b->algo ? 1 : -1;\n>> +       }\n>\n> Nit: unnecessary comment (\"Is the hash algorithm...\") is merely\n> repeating what the code itself already says clearly enough\n\nFair enough.\n\n>\n>> @@ -553,6 +575,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>>         else\n>>                 ds.fn = default_disambiguate_hint;\n>>\n>> +\n>>         find_short_object_filename(&ds);\n>\n> Nit: unnecessary new blank line\n\nNaughty thing how did that sneak in there ;)\n\n\nEric\n"},{"id":"482483","messageId":"87v8bp93j5.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"CAPig+cS02ushqgw+u39Tmnoy3rgp8BzqT4T9D=-01m5fsLxC6Q@mail.gmail.com","subject":"Re: [PATCH 05/30] loose: add a mapping between SHA-1 and SHA-256 for loose objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:11:42Z","receivedAt":"2023-10-02T02:17:12Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"Eric Sunshine <sunshine@sunshineco.com> writes:\n\n> On Wed, Sep 27, 2023 at 3:56 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n>> As part of the transition plan, we'd like to add a file in the .git\n>> directory that maps loose objects between SHA-1 and SHA-256.  Let's\n>> implement the specification in the transition plan and store this data\n>> on a per-repository basis in struct repository.\n>>\n>> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n>> Signed-off-by: Eric W. Biederman <ebiederm@xmission.com>\n>> ---\n>> diff --git a/loose.c b/loose.c\n>> @@ -0,0 +1,245 @@\n>> +static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n>> +{\n>> +       struct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n>> +       FILE *fp;\n>> +\n>> +       if (!dir->loose_map)\n>> +               loose_object_map_init(&dir->loose_map);\n>> +\n>> +       insert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n>> +       insert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n>> +\n>> +       insert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n>> +       insert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n>> +\n>> +       insert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n>> +       insert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n>> +\n>> +       strbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n>> +       fp = fopen(path.buf, \"rb\");\n>> +       if (!fp)\n>> +               return 0;\n>\n> This early return leaks `path`. At minimum, call\n> `strbuf_release(&path)` before returning.\n\nyep.  Fixed.\n\n>> +       errno = 0;\n>> +       if (strbuf_getwholeline(&buf, fp, '\\n') || strcmp(buf.buf, loose_object_header))\n>> +               goto err;\n>> +       while (!strbuf_getline_lf(&buf, fp)) {\n>> +               const char *p;\n>> +               struct object_id oid, compat_oid;\n>> +               if (parse_oid_hex_algop(buf.buf, &oid, &p, repo->hash_algo) ||\n>> +                   *p++ != ' ' ||\n>> +                   parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n>> +                   p != buf.buf + buf.len)\n>> +                       goto err;\n>> +               insert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n>> +               insert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n>> +       }\n>> +\n>> +       strbuf_release(&buf);\n>> +       strbuf_release(&path);\n>> +       return errno ? -1 : 0;\n>> +err:\n>> +       strbuf_release(&buf);\n>> +       strbuf_release(&path);\n>> +       return -1;\n>> +}\n>> +\n>> +int repo_write_loose_object_map(struct repository *repo)\n>> +{\n>> +       kh_oid_map_t *map = repo->objects->odb->loose_map->to_compat;\n>> +       struct lock_file lock;\n>> +       int fd;\n>> +       khiter_t iter;\n>> +       struct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n>> +\n>> +       if (!should_use_loose_object_map(repo))\n>> +               return 0;\n>> +\n>> +       strbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n>> +       fd = hold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n>> +       iter = kh_begin(map);\n>> +       if (write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n>> +               goto errout;\n>> +\n>> +       for (; iter != kh_end(map); iter++) {\n>> +               if (kh_exist(map, iter)) {\n>> +                       if (oideq(&kh_key(map, iter), the_hash_algo->empty_tree) ||\n>> +                           oideq(&kh_key(map, iter), the_hash_algo->empty_blob))\n>> +                               continue;\n>> +                       strbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(&kh_key(map, iter)), oid_to_hex(kh_value(map, iter)));\n>> +                       if (write_in_full(fd, buf.buf, buf.len) < 0)\n>> +                               goto errout;\n>> +                       strbuf_reset(&buf);\n>> +               }\n>> +       }\n>\n> Nit: If you call strbuf_reset() immediately before strbuf_addf(), then\n> you save the reader of the code the effort of having to scan backward\n> through the function to verify that `buf` is empty the first time\n> through the loop.\n\nI am actually perplexed why the code works this way.\n\nMy gut says we should create the entire buffer in memory and then\nwrite it to disk all in a single system call, and not perform\nthis line buffering.\n\nDoing that would remove the need for the strbuf_reset entirely,\nand would just require the buffer to be primed with the\nloose_object_header.\n\nBut what is there works and I will leave it for now.\n\nIt isn't a bug, and it can be improved with a follow up patch.\n\n>> +       strbuf_release(&buf);\n>> +       if (commit_lock_file(&lock) < 0) {\n>> +               error_errno(_(\"could not write loose object index %s\"), path.buf);\n>> +               strbuf_release(&path);\n>> +               return -1;\n>> +       }\n>> +       strbuf_release(&path);\n>> +       return 0;\n>> +errout:\n>> +       rollback_lock_file(&lock);\n>> +       strbuf_release(&buf);\n>> +       error_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n>> +       strbuf_release(&path);\n>> +       return -1;\n>> +}\n>> +\n>> +int repo_loose_object_map_oid(struct repository *repo,\n>> +                             const struct object_id *src,\n>> +                             const struct git_hash_algo *to,\n>> +                             struct object_id *dest)\n>> +\n>> +{\n>\n> Style: unnecessary blank line before opening `{`\n\nYep. Fixed.\n>\n>> +       struct object_directory *dir;\n>> +       kh_oid_map_t *map;\n>> +       khiter_t pos;\n>> +\n>> +       for (dir = repo->objects->odb; dir; dir = dir->next) {\n>> +               struct loose_object_map *loose_map = dir->loose_map;\n>> +               if (!loose_map)\n>> +                       continue;\n>> +               map = (to == repo->compat_hash_algo) ?\n>> +                       loose_map->to_compat :\n>> +                       loose_map->to_storage;\n>> +               pos = kh_get_oid_map(map, *src);\n>> +               if (pos < kh_end(map)) {\n>> +                       oidcpy(dest, kh_value(map, pos));\n>> +                       return 0;\n>> +               }\n>> +       }\n>> +       return -1;\n>> +}\n\nEric\n"},{"id":"482484","messageId":"CAPig+cRxJ1422BMKWZ7bXASMwist0tU+Nb1ZrhHPdt6iknv2kQ@mail.gmail.com","threadId":"60273","inReplyTo":"87o7hhbyxi.fsf@gmail.froward.int.ebiederm.org","subject":"Re: [PATCH 01/30] object-file-convert: Stubs for converting from one object format to another","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2023-10-02T02:27:09Z","receivedAt":"2023-10-02T02:27:25Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Sun, Oct 1, 2023 at 9:22 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n> Eric Sunshine <sunshine@sunshineco.com> writes:\n> > On Wed, Sep 27, 2023 at 3:55 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n> >> +       die(_(\"Failed to convert object from %s to %s\"),\n> >> +               from->name, to->name);\n> >\n> > s/Failed/failed/\n>\n> I don't understand wanting to start a sentence with a lower case letter.\n> Can you explain?\n\nConsistency with most of the rest of the messages emitted by Git.\nCodingGuidelines call for lowercase and omission of the full-stop\n(period).\n\nThe choice, of course, is subjective, but these are the guidelines\nupon which the project has settled for better or worse. (You can\ncertainly find older code, which predates the guideline, using\ncapitalized messages, but it's good to adhere to the guideline for new\ncode if possible.)\n\n> >> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n> >> +                     const struct git_hash_algo *to, struct object_id *dest);\n> >\n> > I suppose the function name is pretty much self-explanatory to those\n> > familiar with the underlying concepts, but it might still be helpful\n> > to add a comment explaining what the function does.\n>\n> I could use words that repeat what is in the function signature.\n> But I don't think I could add anything.\n>\n> I would have to say something like:\n>\n> Look up the oid that an equivalent object would have in a repository\n> whose object format is \"to\".\n>\n> Is that helpful?\n\nThat would help clarify the function a bit for me. Explaining the\nfunction's return value could also be useful.\n"},{"id":"482485","messageId":"CAPig+cS2TVZJQmfQNQud5dRcROFN3kTeu-OBTmCW=7aNjoJGxw@mail.gmail.com","threadId":"60273","inReplyTo":"87v8bp93j5.fsf@gmail.froward.int.ebiederm.org","subject":"Re: [PATCH 05/30] loose: add a mapping between SHA-1 and SHA-256 for loose objects","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2023-10-02T02:36:31Z","receivedAt":"2023-10-02T02:36:46Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Sun, Oct 1, 2023 at 10:11 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n> Eric Sunshine <sunshine@sunshineco.com> writes:\n> > On Wed, Sep 27, 2023 at 3:56 PM Eric W. Biederman <ebiederm@gmail.com> wrote:\n> >> +       for (; iter != kh_end(map); iter++) {\n> >> +               if (kh_exist(map, iter)) {\n> >> +                       if (oideq(&kh_key(map, iter), the_hash_algo->empty_tree) ||\n> >> +                           oideq(&kh_key(map, iter), the_hash_algo->empty_blob))\n> >> +                               continue;\n> >> +                       strbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(&kh_key(map, iter)), oid_to_hex(kh_value(map, iter)));\n> >> +                       if (write_in_full(fd, buf.buf, buf.len) < 0)\n> >> +                               goto errout;\n> >> +                       strbuf_reset(&buf);\n> >> +               }\n> >> +       }\n> >\n> > Nit: If you call strbuf_reset() immediately before strbuf_addf(), then\n> > you save the reader of the code the effort of having to scan backward\n> > through the function to verify that `buf` is empty the first time\n> > through the loop.\n>\n> I am actually perplexed why the code works this way.\n>\n> My gut says we should create the entire buffer in memory and then\n> write it to disk all in a single system call, and not perform\n> this line buffering.\n\nI think I had a similar question/thought while reading the patch but...\n\n> Doing that would remove the need for the strbuf_reset entirely,\n> and would just require the buffer to be primed with the\n> loose_object_header.\n>\n> But what is there works and I will leave it for now.\n\n... came to this same conclusion.\n\n> It isn't a bug, and it can be improved with a follow up patch.\n\nExactly.\n\nI haven't thought through the consequences, but perhaps brian was\nworried about the memory footprint of buffering the whole thing before\nwriting to disk.\n"},{"id":"482486","messageId":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"87jzsbjt0a.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 00/30] initial support for multiple hash functions","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:39:09Z","receivedAt":"2023-10-02T02:39:15Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"\nThis addresses all of the known test failures from v1 of this set of\nchanges.  In particular I have reworked commit_tree_extended which\nwas flagged by smatch, -Werror=array-bounds, and the leak detector.\n\nOne functional bug was fixed in repo_for_each_abbrev where it was\nmistakenly displaying too many ambiguous oids.\n\nI am posting this so that people review and testing of this patchset\nwon't be distracted by the known and fixed issues.\n\n\n\nA key part of the hash function transition plan is a way that a single\ngit repository can inter-operate with git repositories whose storage\nhash function is SHA-1 and git repositories whose storage hash function\nis SHA-256.\n\nThis interoperability can defined in terms of two repositories one whose\nstorage hash function is SHA-1 and another whose storage hash function\nis SHA-256.  Those two repositories receive exactly the same objects,\nbut they store them in different but equivalent ways.\n\nFor a repository that has one storage hash function to inter-operate\nwith a repository that has a different storage hash function requires\nthe first repository to be able produce it's objects as if they were\nstored in the second hash function.\n\nThis series of changes focuses on implementing the pieces that allow\na repository that uses one storage hash function to produce the objects\nthat would have been stored with a second storage hash function.\n\nThe final patch in this series is the addition of a test that creates\ntwo repositories one that uses SHA-1 as it's storage hash function\nand the other that uses SHA-256 as it's storage hash function.\nIdentical operations are performed on the two repositories, and their\ncompatibility objects are compared to verify they are the same.\nAKA the SHA-1 repository on the fly generates the objects store in\nthe SHA-256 repository, and the SHA-256 repository on the fly generates\nthe objects that are stored in the SHA-1 repository.\n\nThere are two fundamental technologies for enabling this.\n- The ability to convert a stored object into the object the\n  other repository would have stored.\n- The ability to remember a mapping between SHA-1 and SHA-256 oids\n  of equivalent objects.\n\nWith such technologies it is very easy to implement user facing changes.\nTo avoid locking git into poor decisions by accident I have done my best\nto minimize the user facing changes, while still building the internal\ninfrastructure that is needed for interoperability.\n\nAll of this work is inspired by earlier work on interoperability by\n\"brian m. carlson\" and some of the key pieces of code are still his.\n\nTo get to the point where I can test if a SHA-1 and a SHA-256 repository\ncan on the fly generate each other, I have made some small user-facing\nchanges.\n\ngit rev-parse now supports --output-object-format as a way to query\nthe internal mapping tables between oids and report the equivalent\noid of the other format.\n\ngit cat-file when given a oid that does not match the repositories\nstorage format will now attempt to find the oids equivalent object that\nis stored in the repository and if found dynamically generate the object\nthat would have been stored in a repository with a different storage\nhash function and display the object.\n\nAn additional file loose-object-index will be stored in \".git/objects/\".\n\nAn additional option \"extensions.compatObjectFormat\" is implemented,\nthat generates and stores mappings between the oids of objects stored in\nthe repository and oids of the equivalent objects that would be stored\nin a repository show storage format was extensions.compatObjectFormat.\n\nEric W. Biederman (23):\n      object-file-convert: stubs for converting from one object format to another\n      oid-array: teach oid-array to handle multiple kinds of oids\n      object-names: support input of oids in any supported hash\n      repository: add a compatibility hash algorithm\n      loose: compatibilty short name support\n      object-file: update the loose object map when writing loose objects\n      object-file: add a compat_oid_in parameter to write_object_file_flags\n      commit: convert mergetag before computing the signature of a commit\n      commit: export add_header_signature to support handling signatures on tags\n      tag: sign both hashes\n      object: factor out parse_mode out of fast-import and tree-walk into in object.h\n      object-file-convert: don't leak when converting tag objects\n      object-file-convert: convert commits that embed signed tags\n      object-file: update object_info_extended to reencode objects\n      rev-parse: add an --output-object-format parameter\n      builtin/cat-file: let the oid determine the output algorithm\n      tree-walk: init_tree_desc take an oid to get the hash algorithm\n      object-file: handle compat objects in check_object_signature\n      builtin/ls-tree: let the oid determine the output algorithm\n      test-lib: compute the compatibility hash so tests may use it\n      t1006: rename sha1 to oid\n      t1006: test oid compatibility with cat-file\n      t1016-compatObjectFormat: add tests to verify the conversion between objects\n\nbrian m. carlson (7):\n      loose: add a mapping between SHA-1 and SHA-256 for loose objects\n      commit: write commits for both hashes\n      cache: add a function to read an OID of a specific algorithm\n      object-file-convert: add a function to convert trees between algorithms\n      object-file-convert: convert tag objects when writing\n      object-file-convert: convert commit objects when writing\n      repository: implement extensions.compatObjectFormat\n\n Documentation/config/extensions.txt |  12 ++\n Documentation/git-rev-parse.txt     |  12 ++\n Makefile                            |   3 +\n archive.c                           |   3 +-\n builtin/am.c                        |   6 +-\n builtin/cat-file.c                  |  12 +-\n builtin/checkout.c                  |   8 +-\n builtin/clone.c                     |   2 +-\n builtin/commit.c                    |   2 +-\n builtin/fast-import.c               |  18 +-\n builtin/grep.c                      |   8 +-\n builtin/ls-tree.c                   |   5 +-\n builtin/merge.c                     |   3 +-\n builtin/pack-objects.c              |   6 +-\n builtin/read-tree.c                 |   2 +-\n builtin/rev-parse.c                 |  25 ++-\n builtin/stash.c                     |   5 +-\n builtin/tag.c                       |  45 ++++-\n cache-tree.c                        |   4 +-\n commit.c                            | 221 ++++++++++++++++-----\n commit.h                            |   1 +\n delta-islands.c                     |   2 +-\n diff-lib.c                          |   2 +-\n fsck.c                              |   6 +-\n hash-ll.h                           |   1 +\n hash.h                              |   9 +-\n http-push.c                         |   2 +-\n list-objects.c                      |   2 +-\n loose.c                             | 259 ++++++++++++++++++++++++\n loose.h                             |  22 +++\n match-trees.c                       |   4 +-\n merge-ort.c                         |  11 +-\n merge-recursive.c                   |   2 +-\n merge.c                             |   3 +-\n object-file-convert.c               | 277 ++++++++++++++++++++++++++\n object-file-convert.h               |  24 +++\n object-file.c                       | 212 ++++++++++++++++++--\n object-name.c                       |  46 +++--\n object-name.h                       |   3 +-\n object-store-ll.h                   |   7 +-\n object.c                            |   2 +\n object.h                            |  18 ++\n oid-array.c                         |  12 +-\n pack-bitmap-write.c                 |   2 +-\n packfile.c                          |   3 +-\n reflog.c                            |   2 +-\n repository.c                        |  14 ++\n repository.h                        |   4 +\n revision.c                          |   4 +-\n setup.c                             |  22 +++\n setup.h                             |   1 +\n t/helper/test-delete-gpgsig.c       |  62 ++++++\n t/helper/test-tool.c                |   1 +\n t/helper/test-tool.h                |   1 +\n t/t1006-cat-file.sh                 | 379 +++++++++++++++++++++---------------\n t/t1016-compatObjectFormat.sh       | 281 ++++++++++++++++++++++++++\n t/t1016/gpg                         |   2 +\n t/test-lib-functions.sh             |  17 +-\n tree-walk.c                         |  58 +++---\n tree-walk.h                         |   7 +-\n tree.c                              |   2 +-\n walker.c                            |   2 +-\n 62 files changed, 1844 insertions(+), 349 deletions(-)\n\nEric\n"},{"id":"482487","messageId":"20231002024034.2611-1-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 01/30] object-file-convert: stubs for converting from one object format to another","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:05Z","receivedAt":"2023-10-02T02:40:43Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nTwo basic functions are provided:\n- convert_object_file Takes an object file it's type and hash algorithm\n  and converts it into the equivalent object file that would\n  have been generated with hash algorithm \"to\".\n\n  For blob objects there is no conversation to be done and it is an\n  error to use this function on them.\n\n  For commit, tree, and tag objects embedded oids are replaced by the\n  oids of the objects they refer to with those objects and their\n  object ids reencoded in with the hash algorithm \"to\".  Signatures\n  are rearranged so that they remain valid after the object has\n  been reencoded.\n\n- repo_oid_to_algop which takes an oid that refers to an object file\n  and returns the oid of the equivalent object file generated\n  with the target hash algorithm.\n\nThe pair of files object-file-convert.c and object-file-convert.h are\nintroduced to hold as much of this logic as possible to keep this\nconversion logic cleanly separated from everything else and in the\nhopes that someday the code will be clean enough git can support\ncompiling out support for sha1 and the various conversion functions.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Makefile              |  1 +\n object-file-convert.c | 57 +++++++++++++++++++++++++++++++++++++++++++\n object-file-convert.h | 24 ++++++++++++++++++\n 3 files changed, 82 insertions(+)\n create mode 100644 object-file-convert.c\n create mode 100644 object-file-convert.h\n\ndiff --git a/Makefile b/Makefile\nindex 577630936535..f7e824f25cda 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1073,6 +1073,7 @@ LIB_OBJS += notes-cache.o\n LIB_OBJS += notes-merge.o\n LIB_OBJS += notes-utils.o\n LIB_OBJS += notes.o\n+LIB_OBJS += object-file-convert.o\n LIB_OBJS += object-file.o\n LIB_OBJS += object-name.o\n LIB_OBJS += object.o\ndiff --git a/object-file-convert.c b/object-file-convert.c\nnew file mode 100644\nindex 000000000000..4777aba83636\n--- /dev/null\n+++ b/object-file-convert.c\n@@ -0,0 +1,57 @@\n+#include \"git-compat-util.h\"\n+#include \"gettext.h\"\n+#include \"strbuf.h\"\n+#include \"repository.h\"\n+#include \"hash-ll.h\"\n+#include \"object.h\"\n+#include \"object-file-convert.h\"\n+\n+int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n+\t\t      const struct git_hash_algo *to, struct object_id *dest)\n+{\n+\t/*\n+\t * If the source algorithm is not set, then we're using the\n+\t * default hash algorithm for that object.\n+\t */\n+\tconst struct git_hash_algo *from =\n+\t\tsrc->algo ? &hash_algos[src->algo] : repo->hash_algo;\n+\n+\tif (from == to) {\n+\t\tif (src != dest)\n+\t\t\toidcpy(dest, src);\n+\t\treturn 0;\n+\t}\n+\treturn -1;\n+}\n+\n+int convert_object_file(struct strbuf *outbuf,\n+\t\t\tconst struct git_hash_algo *from,\n+\t\t\tconst struct git_hash_algo *to,\n+\t\t\tconst void *buf, size_t len,\n+\t\t\tenum object_type type,\n+\t\t\tint gentle)\n+{\n+\tint ret;\n+\n+\t/* Don't call this function when no conversion is necessary */\n+\tif ((from == to) || (type == OBJ_BLOB))\n+\t\tBUG(\"Refusing noop object file conversion\");\n+\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\tcase OBJ_TAG:\n+\tdefault:\n+\t\t/* Not implemented yet, so fail. */\n+\t\tret = -1;\n+\t\tbreak;\n+\t}\n+\tif (!ret)\n+\t\treturn 0;\n+\tif (gentle) {\n+\t\tstrbuf_release(outbuf);\n+\t\treturn ret;\n+\t}\n+\tdie(_(\"Failed to convert object from %s to %s\"),\n+\t\tfrom->name, to->name);\n+}\ndiff --git a/object-file-convert.h b/object-file-convert.h\nnew file mode 100644\nindex 000000000000..a4f802aa8eea\n--- /dev/null\n+++ b/object-file-convert.h\n@@ -0,0 +1,24 @@\n+#ifndef OBJECT_CONVERT_H\n+#define OBJECT_CONVERT_H\n+\n+struct repository;\n+struct object_id;\n+struct git_hash_algo;\n+struct strbuf;\n+#include \"object.h\"\n+\n+int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n+\t\t      const struct git_hash_algo *to, struct object_id *dest);\n+\n+/*\n+ * Convert an object file from one hash algorithm to another algorithm.\n+ * Return -1 on failure, 0 on success.\n+ */\n+int convert_object_file(struct strbuf *outbuf,\n+\t\t\tconst struct git_hash_algo *from,\n+\t\t\tconst struct git_hash_algo *to,\n+\t\t\tconst void *buf, size_t len,\n+\t\t\tenum object_type type,\n+\t\t\tint gentle);\n+\n+#endif /* OBJECT_CONVERT_H */\n-- \n2.41.0\n\n"},{"id":"482488","messageId":"20231002024034.2611-2-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:06Z","receivedAt":"2023-10-02T02:40:45Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nWhile looking at how to handle input of both SHA-1 and SHA-256 oids in\nget_oid_with_context, I realized that the oid_array in\nrepo_for_each_abbrev might have more than one kind of oid stored in it\nsimultaneously.\n\nUpdate to oid_array_append to ensure that oids added to an oid array\nalways have an algorithm set.\n\nUpdate void_hashcmp to first verify two oids use the same hash algorithm\nbefore comparing them to each other.\n\nWith that oid-array should be safe to use with different kinds of\noids simultaneously.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n oid-array.c | 12 ++++++++++--\n 1 file changed, 10 insertions(+), 2 deletions(-)\n\ndiff --git a/oid-array.c b/oid-array.c\nindex 8e4717746c31..1f36651754ed 100644\n--- a/oid-array.c\n+++ b/oid-array.c\n@@ -6,12 +6,20 @@ void oid_array_append(struct oid_array *array, const struct object_id *oid)\n {\n \tALLOC_GROW(array->oid, array->nr + 1, array->alloc);\n \toidcpy(&array->oid[array->nr++], oid);\n+\tif (!oid->algo)\n+\t\toid_set_algo(&array->oid[array->nr - 1], the_hash_algo);\n \tarray->sorted = 0;\n }\n \n-static int void_hashcmp(const void *a, const void *b)\n+static int void_hashcmp(const void *va, const void *vb)\n {\n-\treturn oidcmp(a, b);\n+\tconst struct object_id *a = va, *b = vb;\n+\tint ret;\n+\tif (a->algo == b->algo)\n+\t\tret = oidcmp(a, b);\n+\telse\n+\t\tret = a->algo > b->algo ? 1 : -1;\n+\treturn ret;\n }\n \n void oid_array_sort(struct oid_array *array)\n-- \n2.41.0\n\n"},{"id":"482489","messageId":"20231002024034.2611-3-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 03/30] object-names: support input of oids in any supported hash","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:07Z","receivedAt":"2023-10-02T02:40:46Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nSupport short oids encoded in any algorithm, while ensuring enough of\nthe oid is specified to disambiguate between all of the oids in the\nrepository encoded in any algorithm.\n\nBy default have the code continue to only accept oids specified in the\nstorage hash algorithm of the repository, but when something is\nambiguous display all of the possible oids from any accepted oid\nencoding.\n\nA new flag is added GET_OID_HASH_ANY that when supplied causes the\ncode to accept oids specified in any hash algorithm, and to return the\noids that were resolved.\n\nThis implements the functionality that allows both SHA-1 and SHA-256\nobject names, from the \"Object names on the command line\" section of\nthe hash function transition document.\n\nCare is taken in get_short_oid so that when the result is ambiguous\nthe output remains the same if GIT_OID_HASH_ANY was not supplied.  If\nGET_OID_HASH_ANY was supplied objects of any hash algorithm that match\nthe prefix are displayed.\n\nThis required updating repo_for_each_abbrev to give it a parameter so\nthat it knows to look at all hash algorithms.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/rev-parse.c |  2 +-\n hash-ll.h           |  1 +\n object-name.c       | 46 ++++++++++++++++++++++++++++++++++-----------\n object-name.h       |  3 ++-\n 4 files changed, 39 insertions(+), 13 deletions(-)\n\ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex fde8861ca4e0..43e96765400c 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -882,7 +882,7 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t\t\tcontinue;\n \t\t\t}\n \t\t\tif (skip_prefix(arg, \"--disambiguate=\", &arg)) {\n-\t\t\t\trepo_for_each_abbrev(the_repository, arg,\n+\t\t\t\trepo_for_each_abbrev(the_repository, arg, the_hash_algo,\n \t\t\t\t\t\t     show_abbrev, NULL);\n \t\t\t\tcontinue;\n \t\t\t}\ndiff --git a/hash-ll.h b/hash-ll.h\nindex 10d84cc20888..2cfde63ae1cf 100644\n--- a/hash-ll.h\n+++ b/hash-ll.h\n@@ -145,6 +145,7 @@ struct object_id {\n #define GET_OID_RECORD_PATH     0200\n #define GET_OID_ONLY_TO_DIE    04000\n #define GET_OID_REQUIRE_PATH  010000\n+#define GET_OID_HASH_ANY      020000\n \n #define GET_OID_DISAMBIGUATORS \\\n \t(GET_OID_COMMIT | GET_OID_COMMITTISH | \\\ndiff --git a/object-name.c b/object-name.c\nindex 0bfa29dbbfe9..7dd6e5e47566 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -25,6 +25,7 @@\n #include \"midx.h\"\n #include \"commit-reach.h\"\n #include \"date.h\"\n+#include \"object-file-convert.h\"\n \n static int get_oid_oneline(struct repository *r, const char *, struct object_id *, struct commit_list *);\n \n@@ -49,6 +50,7 @@ struct disambiguate_state {\n \n static void update_candidates(struct disambiguate_state *ds, const struct object_id *current)\n {\n+\t/* The hash algorithm of current has already been filtered */\n \tif (ds->always_call_fn) {\n \t\tds->ambiguous = ds->fn(ds->repo, current, ds->cb_data) ? 1 : 0;\n \t\treturn;\n@@ -134,6 +136,8 @@ static void unique_in_midx(struct multi_pack_index *m,\n {\n \tuint32_t num, i, first = 0;\n \tconst struct object_id *current = NULL;\n+\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n+\t\tds->repo->hash_algo->hexsz : ds->len;\n \tnum = m->num_objects;\n \n \tif (!num)\n@@ -149,7 +153,7 @@ static void unique_in_midx(struct multi_pack_index *m,\n \tfor (i = first; i < num && !ds->ambiguous; i++) {\n \t\tstruct object_id oid;\n \t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n-\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, current->hash))\n+\t\tif (!match_hash(len, ds->bin_pfx.hash, current->hash))\n \t\t\tbreak;\n \t\tupdate_candidates(ds, current);\n \t}\n@@ -159,6 +163,8 @@ static void unique_in_pack(struct packed_git *p,\n \t\t\t   struct disambiguate_state *ds)\n {\n \tuint32_t num, i, first = 0;\n+\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n+\t\tds->repo->hash_algo->hexsz : ds->len;\n \n \tif (p->multi_pack_index)\n \t\treturn;\n@@ -177,7 +183,7 @@ static void unique_in_pack(struct packed_git *p,\n \tfor (i = first; i < num && !ds->ambiguous; i++) {\n \t\tstruct object_id oid;\n \t\tnth_packed_object_id(&oid, p, i);\n-\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, oid.hash))\n+\t\tif (!match_hash(len, ds->bin_pfx.hash, oid.hash))\n \t\t\tbreak;\n \t\tupdate_candidates(ds, &oid);\n \t}\n@@ -188,6 +194,10 @@ static void find_short_packed_object(struct disambiguate_state *ds)\n \tstruct multi_pack_index *m;\n \tstruct packed_git *p;\n \n+\t/* Skip, unless oids from the storage hash algorithm are wanted */\n+\tif (ds->bin_pfx.algo && (&hash_algos[ds->bin_pfx.algo] != ds->repo->hash_algo))\n+\t\treturn;\n+\n \tfor (m = get_multi_pack_index(ds->repo); m && !ds->ambiguous;\n \t     m = m->next)\n \t\tunique_in_midx(m, ds);\n@@ -326,11 +336,12 @@ int set_disambiguate_hint_config(const char *var, const char *value)\n \n static int init_object_disambiguation(struct repository *r,\n \t\t\t\t      const char *name, int len,\n+\t\t\t\t      const struct git_hash_algo *algo,\n \t\t\t\t      struct disambiguate_state *ds)\n {\n \tint i;\n \n-\tif (len < MINIMUM_ABBREV || len > the_hash_algo->hexsz)\n+\tif (len < MINIMUM_ABBREV || len > GIT_MAX_HEXSZ)\n \t\treturn -1;\n \n \tmemset(ds, 0, sizeof(*ds));\n@@ -357,6 +368,7 @@ static int init_object_disambiguation(struct repository *r,\n \tds->len = len;\n \tds->hex_pfx[len] = '\\0';\n \tds->repo = r;\n+\tds->bin_pfx.algo = algo ? hash_algo_by_ptr(algo) : GIT_HASH_UNKNOWN;\n \tprepare_alt_odb(r);\n \treturn 0;\n }\n@@ -491,9 +503,10 @@ static int repo_collect_ambiguous(struct repository *r UNUSED,\n \treturn collect_ambiguous(oid, data);\n }\n \n-static int sort_ambiguous(const void *a, const void *b, void *ctx)\n+static int sort_ambiguous(const void *va, const void *vb, void *ctx)\n {\n \tstruct repository *sort_ambiguous_repo = ctx;\n+\tconst struct object_id *a = va, *b = vb;\n \tint a_type = oid_object_info(sort_ambiguous_repo, a, NULL);\n \tint b_type = oid_object_info(sort_ambiguous_repo, b, NULL);\n \tint a_type_sort;\n@@ -503,8 +516,12 @@ static int sort_ambiguous(const void *a, const void *b, void *ctx)\n \t * Sorts by hash within the same object type, just as\n \t * oid_array_for_each_unique() would do.\n \t */\n-\tif (a_type == b_type)\n-\t\treturn oidcmp(a, b);\n+\tif (a_type == b_type) {\n+\t\tif (a->algo == b->algo)\n+\t\t\treturn oidcmp(a, b);\n+\t\telse\n+\t\t\treturn a->algo > b->algo ? 1 : -1;\n+\t}\n \n \t/*\n \t * Between object types show tags, then commits, and finally\n@@ -533,8 +550,12 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \tint status;\n \tstruct disambiguate_state ds;\n \tint quietly = !!(flags & GET_OID_QUIETLY);\n+\tconst struct git_hash_algo *algo = r->hash_algo;\n+\n+\tif (flags & GET_OID_HASH_ANY)\n+\t\talgo = NULL;\n \n-\tif (init_object_disambiguation(r, name, len, &ds) < 0)\n+\tif (init_object_disambiguation(r, name, len, algo, &ds) < 0)\n \t\treturn -1;\n \n \tif (HAS_MULTI_BITS(flags & GET_OID_DISAMBIGUATORS))\n@@ -588,7 +609,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t\tif (!ds.ambiguous)\n \t\t\tds.fn = NULL;\n \n-\t\trepo_for_each_abbrev(r, ds.hex_pfx, collect_ambiguous, &collect);\n+\t\trepo_for_each_abbrev(r, ds.hex_pfx, algo, collect_ambiguous, &collect);\n \t\tsort_ambiguous_oid_array(r, &collect);\n \n \t\tif (oid_array_for_each(&collect, show_ambiguous_object, &out))\n@@ -610,13 +631,14 @@ static enum get_oid_result get_short_oid(struct repository *r,\n }\n \n int repo_for_each_abbrev(struct repository *r, const char *prefix,\n+\t\t\t const struct git_hash_algo *algo,\n \t\t\t each_abbrev_fn fn, void *cb_data)\n {\n \tstruct oid_array collect = OID_ARRAY_INIT;\n \tstruct disambiguate_state ds;\n \tint ret;\n \n-\tif (init_object_disambiguation(r, prefix, strlen(prefix), &ds) < 0)\n+\tif (init_object_disambiguation(r, prefix, strlen(prefix), algo, &ds) < 0)\n \t\treturn -1;\n \n \tds.always_call_fn = 1;\n@@ -787,10 +809,12 @@ void strbuf_add_unique_abbrev(struct strbuf *sb, const struct object_id *oid,\n int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \t\t\t      const struct object_id *oid, int len)\n {\n+\tconst struct git_hash_algo *algo =\n+\t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n \tstruct disambiguate_state ds;\n \tstruct min_abbrev_data mad;\n \tstruct object_id oid_ret;\n-\tconst unsigned hexsz = r->hash_algo->hexsz;\n+\tconst unsigned hexsz = algo->hexsz;\n \n \tif (len < 0) {\n \t\tunsigned long count = repo_approximate_object_count(r);\n@@ -826,7 +850,7 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \n \tfind_abbrev_len_packed(&mad);\n \n-\tif (init_object_disambiguation(r, hex, mad.cur_len, &ds) < 0)\n+\tif (init_object_disambiguation(r, hex, mad.cur_len, algo, &ds) < 0)\n \t\treturn -1;\n \n \tds.fn = repo_extend_abbrev_len;\ndiff --git a/object-name.h b/object-name.h\nindex 9ae522307148..064ddc97d1fe 100644\n--- a/object-name.h\n+++ b/object-name.h\n@@ -67,7 +67,8 @@ enum get_oid_result get_oid_with_context(struct repository *repo, const char *st\n \n \n typedef int each_abbrev_fn(const struct object_id *oid, void *);\n-int repo_for_each_abbrev(struct repository *r, const char *prefix, each_abbrev_fn, void *);\n+int repo_for_each_abbrev(struct repository *r, const char *prefix,\n+\t\t\t const struct git_hash_algo *algo, each_abbrev_fn, void *);\n \n int set_disambiguate_hint_config(const char *var, const char *value);\n \n-- \n2.41.0\n\n"},{"id":"482490","messageId":"20231002024034.2611-4-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 04/30] repository: add a compatibility hash algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:08Z","receivedAt":"2023-10-02T02:40:51Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nWe currently have support for using a full stage 4 SHA-256\nimplementation.  However, we'd like to support interoperability with\nSHA-1 repositories as well.  The transition plan anticipates a\ncompatibility hash algorithm configuration option that we can use to\nimplement support for this.  Let's add an element to the repository\nstructure that indicates the compatibility hash algorithm so we can use\nit when we need to consider interoperability between algorithms.\n\nAdd a helper function repo_set_compat_hash_algo that takes a\ncompatibility hash algorithm and sets \"repo->compat_hash_algo\".  If\nGIT_HASH_UNKNOWN is passed as the compatibility hash algorithm\n\"repo->compat_hash_algo\" is set to NULL.\n\nFor now, the code results in \"repo->compat_hash_algo\" always being set\nto NULL, but that will change once a configuration option is added.\n\nInspired-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n repository.c | 8 ++++++++\n repository.h | 4 ++++\n setup.c      | 3 +++\n 3 files changed, 15 insertions(+)\n\ndiff --git a/repository.c b/repository.c\nindex a7679ceeaa45..80252b79e93e 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -104,6 +104,13 @@ void repo_set_hash_algo(struct repository *repo, int hash_algo)\n \trepo->hash_algo = &hash_algos[hash_algo];\n }\n \n+void repo_set_compat_hash_algo(struct repository *repo, int algo)\n+{\n+\tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n+\t\tBUG(\"hash_algo and compat_hash_algo match\");\n+\trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n+}\n+\n /*\n  * Attempt to resolve and set the provided 'gitdir' for repository 'repo'.\n  * Return 0 upon success and a non-zero value upon failure.\n@@ -184,6 +191,7 @@ int repo_init(struct repository *repo,\n \t\tgoto error;\n \n \trepo_set_hash_algo(repo, format.hash_algo);\n+\trepo_set_compat_hash_algo(repo, GIT_HASH_UNKNOWN);\n \trepo->repository_format_worktree_config = format.worktree_config;\n \n \t/* take ownership of format.partial_clone */\ndiff --git a/repository.h b/repository.h\nindex 5f18486f6465..bf3fc601cc53 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -160,6 +160,9 @@ struct repository {\n \t/* Repository's current hash algorithm, as serialized on disk. */\n \tconst struct git_hash_algo *hash_algo;\n \n+\t/* Repository's compatibility hash algorithm. */\n+\tconst struct git_hash_algo *compat_hash_algo;\n+\n \t/* A unique-id for tracing purposes. */\n \tint trace2_repo_id;\n \n@@ -199,6 +202,7 @@ void repo_set_gitdir(struct repository *repo, const char *root,\n \t\t     const struct set_gitdir_args *extra_args);\n void repo_set_worktree(struct repository *repo, const char *path);\n void repo_set_hash_algo(struct repository *repo, int algo);\n+void repo_set_compat_hash_algo(struct repository *repo, int compat_algo);\n void initialize_the_repository(void);\n RESULT_MUST_BE_USED\n int repo_init(struct repository *r, const char *gitdir, const char *worktree);\ndiff --git a/setup.c b/setup.c\nindex 18927a847b86..aa8bf5da5226 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1564,6 +1564,8 @@ const char *setup_git_directory_gently(int *nongit_ok)\n \t\t}\n \t\tif (startup_info->have_repository) {\n \t\t\trepo_set_hash_algo(the_repository, repo_fmt.hash_algo);\n+\t\t\trepo_set_compat_hash_algo(the_repository,\n+\t\t\t\t\t\t  GIT_HASH_UNKNOWN);\n \t\t\tthe_repository->repository_format_worktree_config =\n \t\t\t\trepo_fmt.worktree_config;\n \t\t\t/* take ownership of repo_fmt.partial_clone */\n@@ -1657,6 +1659,7 @@ void check_repository_format(struct repository_format *fmt)\n \tcheck_repository_format_gently(get_git_dir(), fmt, NULL);\n \tstartup_info->have_repository = 1;\n \trepo_set_hash_algo(the_repository, fmt->hash_algo);\n+\trepo_set_compat_hash_algo(the_repository, GIT_HASH_UNKNOWN);\n \tthe_repository->repository_format_worktree_config =\n \t\tfmt->worktree_config;\n \tthe_repository->repository_format_partial_clone =\n-- \n2.41.0\n\n"},{"id":"482491","messageId":"20231002024034.2611-6-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 06/30] loose: compatibilty short name support","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:10Z","receivedAt":"2023-10-02T02:40:52Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpdate loose_objects_cache when udpating the loose objects map.  This\noidtree is used to discover which oids are possibilities when\nresolving short names, and it can support a mixture of sha1\nand sha256 oids.\n\nWith this any oid recorded objects/loose-objects-idx is usable\nfor resolving an oid to an object.\n\nTo make this maintainable a helper insert_loose_map is factored\nout of load_one_loose_object_map and repo_add_loose_object_map,\nand then modified to also update the loose_objects_cache.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n loose.c | 37 +++++++++++++++++++++++++------------\n 1 file changed, 25 insertions(+), 12 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 6ba73cc84dca..f6faa6216a08 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -7,6 +7,7 @@\n #include \"gettext.h\"\n #include \"loose.h\"\n #include \"lockfile.h\"\n+#include \"oidtree.h\"\n \n static const char *loose_object_header = \"# loose-object-idx\\n\";\n \n@@ -42,6 +43,21 @@ static int insert_oid_pair(kh_oid_map_t *map, const struct object_id *key, const\n \treturn 1;\n }\n \n+static int insert_loose_map(struct object_directory *odb,\n+\t\t\t    const struct object_id *oid,\n+\t\t\t    const struct object_id *compat_oid)\n+{\n+\tstruct loose_object_map *map = odb->loose_map;\n+\tint inserted = 0;\n+\n+\tinserted |= insert_oid_pair(map->to_compat, oid, compat_oid);\n+\tinserted |= insert_oid_pair(map->to_storage, compat_oid, oid);\n+\tif (inserted)\n+\t\toidtree_insert(odb->loose_objects_cache, compat_oid);\n+\n+\treturn inserted;\n+}\n+\n static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n {\n \tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n@@ -49,15 +65,14 @@ static int load_one_loose_object_map(struct repository *repo, struct object_dire\n \n \tif (!dir->loose_map)\n \t\tloose_object_map_init(&dir->loose_map);\n+\tif (!dir->loose_objects_cache) {\n+\t\tALLOC_ARRAY(dir->loose_objects_cache, 1);\n+\t\toidtree_init(dir->loose_objects_cache);\n+\t}\n \n-\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n-\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n-\n-\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n-\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n-\n-\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n-\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n+\tinsert_loose_map(dir, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n+\tinsert_loose_map(dir, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n+\tinsert_loose_map(dir, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n \n \tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n \tfp = fopen(path.buf, \"rb\");\n@@ -77,8 +92,7 @@ static int load_one_loose_object_map(struct repository *repo, struct object_dire\n \t\t    parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n \t\t    p != buf.buf + buf.len)\n \t\t\tgoto err;\n-\t\tinsert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n-\t\tinsert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n+\t\tinsert_loose_map(dir, &oid, &compat_oid);\n \t}\n \n \tstrbuf_release(&buf);\n@@ -197,8 +211,7 @@ int repo_add_loose_object_map(struct repository *repo, const struct object_id *o\n \tif (!should_use_loose_object_map(repo))\n \t\treturn 0;\n \n-\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_compat, oid, compat_oid);\n-\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_storage, compat_oid, oid);\n+\tinserted = insert_loose_map(repo->objects->odb, oid, compat_oid);\n \tif (inserted)\n \t\treturn write_one_object(repo, oid, compat_oid);\n \treturn 0;\n-- \n2.41.0\n\n"},{"id":"482492","messageId":"20231002024034.2611-5-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 05/30] loose: add a mapping between SHA-1 and SHA-256 for loose objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:09Z","receivedAt":"2023-10-02T02:40:53Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nAs part of the transition plan, we'd like to add a file in the .git\ndirectory that maps loose objects between SHA-1 and SHA-256.  Let's\nimplement the specification in the transition plan and store this data\non a per-repository basis in struct repository.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n Makefile              |   1 +\n loose.c               | 246 ++++++++++++++++++++++++++++++++++++++++++\n loose.h               |  22 ++++\n object-file-convert.c |  14 ++-\n object-store-ll.h     |   3 +\n object.c              |   2 +\n repository.c          |   6 ++\n 7 files changed, 293 insertions(+), 1 deletion(-)\n create mode 100644 loose.c\n create mode 100644 loose.h\n\ndiff --git a/Makefile b/Makefile\nindex f7e824f25cda..3c18664def9a 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1053,6 +1053,7 @@ LIB_OBJS += list-objects-filter.o\n LIB_OBJS += list-objects.o\n LIB_OBJS += lockfile.o\n LIB_OBJS += log-tree.o\n+LIB_OBJS += loose.o\n LIB_OBJS += ls-refs.o\n LIB_OBJS += mailinfo.o\n LIB_OBJS += mailmap.o\ndiff --git a/loose.c b/loose.c\nnew file mode 100644\nindex 000000000000..6ba73cc84dca\n--- /dev/null\n+++ b/loose.c\n@@ -0,0 +1,246 @@\n+#include \"git-compat-util.h\"\n+#include \"hash.h\"\n+#include \"path.h\"\n+#include \"object-store.h\"\n+#include \"hex.h\"\n+#include \"wrapper.h\"\n+#include \"gettext.h\"\n+#include \"loose.h\"\n+#include \"lockfile.h\"\n+\n+static const char *loose_object_header = \"# loose-object-idx\\n\";\n+\n+static inline int should_use_loose_object_map(struct repository *repo)\n+{\n+\treturn repo->compat_hash_algo && repo->gitdir;\n+}\n+\n+void loose_object_map_init(struct loose_object_map **map)\n+{\n+\tstruct loose_object_map *m;\n+\tm = xmalloc(sizeof(**map));\n+\tm->to_compat = kh_init_oid_map();\n+\tm->to_storage = kh_init_oid_map();\n+\t*map = m;\n+}\n+\n+static int insert_oid_pair(kh_oid_map_t *map, const struct object_id *key, const struct object_id *value)\n+{\n+\tkhiter_t pos;\n+\tint ret;\n+\tstruct object_id *stored;\n+\n+\tpos = kh_put_oid_map(map, *key, &ret);\n+\n+\t/* This item already exists in the map. */\n+\tif (ret == 0)\n+\t\treturn 0;\n+\n+\tstored = xmalloc(sizeof(*stored));\n+\toidcpy(stored, value);\n+\tkh_value(map, pos) = stored;\n+\treturn 1;\n+}\n+\n+static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n+{\n+\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\tFILE *fp;\n+\n+\tif (!dir->loose_map)\n+\t\tloose_object_map_init(&dir->loose_map);\n+\n+\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n+\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n+\n+\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n+\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n+\n+\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n+\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n+\n+\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n+\tfp = fopen(path.buf, \"rb\");\n+\tif (!fp) {\n+\t\tstrbuf_release(&path);\n+\t\treturn 0;\n+\t}\n+\n+\terrno = 0;\n+\tif (strbuf_getwholeline(&buf, fp, '\\n') || strcmp(buf.buf, loose_object_header))\n+\t\tgoto err;\n+\twhile (!strbuf_getline_lf(&buf, fp)) {\n+\t\tconst char *p;\n+\t\tstruct object_id oid, compat_oid;\n+\t\tif (parse_oid_hex_algop(buf.buf, &oid, &p, repo->hash_algo) ||\n+\t\t    *p++ != ' ' ||\n+\t\t    parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n+\t\t    p != buf.buf + buf.len)\n+\t\t\tgoto err;\n+\t\tinsert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n+\t\tinsert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n+\t}\n+\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn errno ? -1 : 0;\n+err:\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn -1;\n+}\n+\n+int repo_read_loose_object_map(struct repository *repo)\n+{\n+\tstruct object_directory *dir;\n+\n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n+\tprepare_alt_odb(repo);\n+\n+\tfor (dir = repo->objects->odb; dir; dir = dir->next) {\n+\t\tif (load_one_loose_object_map(repo, dir) < 0) {\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\treturn 0;\n+}\n+\n+int repo_write_loose_object_map(struct repository *repo)\n+{\n+\tkh_oid_map_t *map = repo->objects->odb->loose_map->to_compat;\n+\tstruct lock_file lock;\n+\tint fd;\n+\tkhiter_t iter;\n+\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n+\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n+\tfd = hold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n+\titer = kh_begin(map);\n+\tif (write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n+\t\tgoto errout;\n+\n+\tfor (; iter != kh_end(map); iter++) {\n+\t\tif (kh_exist(map, iter)) {\n+\t\t\tif (oideq(&kh_key(map, iter), the_hash_algo->empty_tree) ||\n+\t\t\t    oideq(&kh_key(map, iter), the_hash_algo->empty_blob))\n+\t\t\t\tcontinue;\n+\t\t\tstrbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(&kh_key(map, iter)), oid_to_hex(kh_value(map, iter)));\n+\t\t\tif (write_in_full(fd, buf.buf, buf.len) < 0)\n+\t\t\t\tgoto errout;\n+\t\t\tstrbuf_reset(&buf);\n+\t\t}\n+\t}\n+\tstrbuf_release(&buf);\n+\tif (commit_lock_file(&lock) < 0) {\n+\t\terror_errno(_(\"could not write loose object index %s\"), path.buf);\n+\t\tstrbuf_release(&path);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_release(&path);\n+\treturn 0;\n+errout:\n+\trollback_lock_file(&lock);\n+\tstrbuf_release(&buf);\n+\terror_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n+\tstrbuf_release(&path);\n+\treturn -1;\n+}\n+\n+static int write_one_object(struct repository *repo, const struct object_id *oid,\n+\t\t\t    const struct object_id *compat_oid)\n+{\n+\tstruct lock_file lock;\n+\tint fd;\n+\tstruct stat st;\n+\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\n+\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n+\thold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n+\n+\tfd = open(path.buf, O_WRONLY | O_CREAT | O_APPEND, 0666);\n+\tif (fd < 0)\n+\t\tgoto errout;\n+\tif (fstat(fd, &st) < 0)\n+\t\tgoto errout;\n+\tif (!st.st_size && write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n+\t\tgoto errout;\n+\n+\tstrbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(oid), oid_to_hex(compat_oid));\n+\tif (write_in_full(fd, buf.buf, buf.len) < 0)\n+\t\tgoto errout;\n+\tif (close(fd))\n+\t\tgoto errout;\n+\tadjust_shared_perm(path.buf);\n+\trollback_lock_file(&lock);\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn 0;\n+errout:\n+\terror_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n+\tclose(fd);\n+\trollback_lock_file(&lock);\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn -1;\n+}\n+\n+int repo_add_loose_object_map(struct repository *repo, const struct object_id *oid,\n+\t\t\t      const struct object_id *compat_oid)\n+{\n+\tint inserted = 0;\n+\n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n+\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_compat, oid, compat_oid);\n+\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_storage, compat_oid, oid);\n+\tif (inserted)\n+\t\treturn write_one_object(repo, oid, compat_oid);\n+\treturn 0;\n+}\n+\n+int repo_loose_object_map_oid(struct repository *repo,\n+\t\t\t      const struct object_id *src,\n+\t\t\t      const struct git_hash_algo *to,\n+\t\t\t      struct object_id *dest)\n+{\n+\tstruct object_directory *dir;\n+\tkh_oid_map_t *map;\n+\tkhiter_t pos;\n+\n+\tfor (dir = repo->objects->odb; dir; dir = dir->next) {\n+\t\tstruct loose_object_map *loose_map = dir->loose_map;\n+\t\tif (!loose_map)\n+\t\t\tcontinue;\n+\t\tmap = (to == repo->compat_hash_algo) ?\n+\t\t\tloose_map->to_compat :\n+\t\t\tloose_map->to_storage;\n+\t\tpos = kh_get_oid_map(map, *src);\n+\t\tif (pos < kh_end(map)) {\n+\t\t\toidcpy(dest, kh_value(map, pos));\n+\t\t\treturn 0;\n+\t\t}\n+\t}\n+\treturn -1;\n+}\n+\n+void loose_object_map_clear(struct loose_object_map **map)\n+{\n+\tstruct loose_object_map *m = *map;\n+\tstruct object_id *oid;\n+\n+\tif (!m)\n+\t\treturn;\n+\n+\tkh_foreach_value(m->to_compat, oid, free(oid));\n+\tkh_foreach_value(m->to_storage, oid, free(oid));\n+\tkh_destroy_oid_map(m->to_compat);\n+\tkh_destroy_oid_map(m->to_storage);\n+\tfree(m);\n+\t*map = NULL;\n+}\ndiff --git a/loose.h b/loose.h\nnew file mode 100644\nindex 000000000000..2c2957072c5f\n--- /dev/null\n+++ b/loose.h\n@@ -0,0 +1,22 @@\n+#ifndef LOOSE_H\n+#define LOOSE_H\n+\n+#include \"khash.h\"\n+\n+struct loose_object_map {\n+\tkh_oid_map_t *to_compat;\n+\tkh_oid_map_t *to_storage;\n+};\n+\n+void loose_object_map_init(struct loose_object_map **map);\n+void loose_object_map_clear(struct loose_object_map **map);\n+int repo_loose_object_map_oid(struct repository *repo,\n+\t\t\t      const struct object_id *src,\n+\t\t\t      const struct git_hash_algo *dest_algo,\n+\t\t\t      struct object_id *dest);\n+int repo_add_loose_object_map(struct repository *repo, const struct object_id *oid,\n+\t\t\t      const struct object_id *compat_oid);\n+int repo_read_loose_object_map(struct repository *repo);\n+int repo_write_loose_object_map(struct repository *repo);\n+\n+#endif\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 4777aba83636..1ec945eaa17f 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -4,6 +4,7 @@\n #include \"repository.h\"\n #include \"hash-ll.h\"\n #include \"object.h\"\n+#include \"loose.h\"\n #include \"object-file-convert.h\"\n \n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n@@ -21,7 +22,18 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \t\t\toidcpy(dest, src);\n \t\treturn 0;\n \t}\n-\treturn -1;\n+\tif (repo_loose_object_map_oid(repo, src, to, dest)) {\n+\t\t/*\n+\t\t * We may have loaded the object map at repo initialization but\n+\t\t * another process (perhaps upstream of a pipe from us) may have\n+\t\t * written a new object into the map.  If the object is missing,\n+\t\t * let's reload the map to see if the object has appeared.\n+\t\t */\n+\t\trepo_read_loose_object_map(repo);\n+\t\tif (repo_loose_object_map_oid(repo, src, to, dest))\n+\t\t\treturn -1;\n+\t}\n+\treturn 0;\n }\n \n int convert_object_file(struct strbuf *outbuf,\ndiff --git a/object-store-ll.h b/object-store-ll.h\nindex 26a3895c821c..bc76d6bec80d 100644\n--- a/object-store-ll.h\n+++ b/object-store-ll.h\n@@ -26,6 +26,9 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/* Map between object IDs for loose objects. */\n+\tstruct loose_object_map *loose_map;\n+\n \t/*\n \t * This is a temporary object store created by the tmp_objdir\n \t * facility. Disable ref updates since the objects in the store\ndiff --git a/object.c b/object.c\nindex 2c61e4c86217..186a0a47c0fb 100644\n--- a/object.c\n+++ b/object.c\n@@ -13,6 +13,7 @@\n #include \"alloc.h\"\n #include \"packfile.h\"\n #include \"commit-graph.h\"\n+#include \"loose.h\"\n \n unsigned int get_max_object_index(void)\n {\n@@ -540,6 +541,7 @@ void free_object_directory(struct object_directory *odb)\n {\n \tfree(odb->path);\n \todb_clear_loose_cache(odb);\n+\tloose_object_map_clear(&odb->loose_map);\n \tfree(odb);\n }\n \ndiff --git a/repository.c b/repository.c\nindex 80252b79e93e..6214f61cf4e7 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -14,6 +14,7 @@\n #include \"read-cache-ll.h\"\n #include \"remote.h\"\n #include \"setup.h\"\n+#include \"loose.h\"\n #include \"submodule-config.h\"\n #include \"sparse-index.h\"\n #include \"trace2.h\"\n@@ -109,6 +110,8 @@ void repo_set_compat_hash_algo(struct repository *repo, int algo)\n \tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n \t\tBUG(\"hash_algo and compat_hash_algo match\");\n \trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n+\tif (repo->compat_hash_algo)\n+\t\trepo_read_loose_object_map(repo);\n }\n \n /*\n@@ -201,6 +204,9 @@ int repo_init(struct repository *repo,\n \tif (worktree)\n \t\trepo_set_worktree(repo, worktree);\n \n+\tif (repo->compat_hash_algo)\n+\t\trepo_read_loose_object_map(repo);\n+\n \tclear_repository_format(&format);\n \treturn 0;\n \n-- \n2.41.0\n\n"},{"id":"482493","messageId":"20231002024034.2611-7-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 07/30] object-file: update the loose object map when writing loose objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:11Z","receivedAt":"2023-10-02T02:40:55Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nTo implement SHA1 compatibility on SHA256 repositories the loose\nobject map needs to be updated whenver a loose object is written.\nUpdating the loose object map this way allows git to support\nthe old hash algorithm in constant time.\n\nThe functions write_loose_object, and stream_loose_object are\nthe only two functions that write to the loose object store.\n\nUpdate stream_loose_object to compute the compatibiilty hash, update\nthe loose object, and then call repo_add_loose_object_map to update\nthe loose object map.\n\nUpdate write_object_file_flags to convert the object into\nit's compatibility encoding, hash the compatibility encoding,\nwrite the object, and then update the loose object map.\n\nUpdate force_object_loose to lookup the hash of the compatibility\nencoding, write the loose object, and then update the loose object\nmap.\n\nUpdate write_object_file_literally to convert the object into it's\ncompatibility hash encoding, hash the compatibility enconding, write\nthe object, and then update the loose object map, when the type string\nis a known type.  For objects with an unknown type this results in a\npartially broken repository, as the objects are not mapped.\n\nThe point of write_object_file_literally is to generate a partially\nbroken repository for testing.  For testing skipping writing the loose\nobject map is much more useful than refusing to write the broken\nobject at all.\n\nExcept that the loose objects are updated before the loose object map\nI have not done any analysis to see how robust this scheme is in the\nevent of failure.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n object-file.c | 113 ++++++++++++++++++++++++++++++++++++++++++--------\n 1 file changed, 95 insertions(+), 18 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 7dc0c4bfbba8..4e55f475b3b4 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -43,6 +43,8 @@\n #include \"setup.h\"\n #include \"submodule.h\"\n #include \"fsck.h\"\n+#include \"loose.h\"\n+#include \"object-file-convert.h\"\n \n /* The maximum size for an object header. */\n #define MAX_HEADER_LEN 32\n@@ -1952,9 +1954,12 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \t\t\t\t     const char *filename, unsigned flags,\n \t\t\t\t     git_zstream *stream,\n \t\t\t\t     unsigned char *buf, size_t buflen,\n-\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     git_hash_ctx *c, git_hash_ctx *compat_c,\n \t\t\t\t     char *hdr, int hdrlen)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tint fd;\n \n \tfd = create_tmpfile(tmp_file, filename);\n@@ -1974,14 +1979,18 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \tgit_deflate_init(stream, zlib_compression_level);\n \tstream->next_out = buf;\n \tstream->avail_out = buflen;\n-\tthe_hash_algo->init_fn(c);\n+\talgo->init_fn(c);\n+\tif (compat && compat_c)\n+\t\tcompat->init_fn(compat_c);\n \n \t/*  Start to feed header to zlib stream */\n \tstream->next_in = (unsigned char *)hdr;\n \tstream->avail_in = hdrlen;\n \twhile (git_deflate(stream, 0) == Z_OK)\n \t\t; /* nothing */\n-\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\talgo->update_fn(c, hdr, hdrlen);\n+\tif (compat && compat_c)\n+\t\tcompat->update_fn(compat_c, hdr, hdrlen);\n \n \treturn fd;\n }\n@@ -1990,16 +1999,21 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n  * Common steps for the inner git_deflate() loop for writing loose\n  * objects. Returns what git_deflate() returns.\n  */\n-static int write_loose_object_common(git_hash_ctx *c,\n+static int write_loose_object_common(git_hash_ctx *c, git_hash_ctx *compat_c,\n \t\t\t\t     git_zstream *stream, const int flush,\n \t\t\t\t     unsigned char *in0, const int fd,\n \t\t\t\t     unsigned char *compressed,\n \t\t\t\t     const size_t compressed_len)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tint ret;\n \n \tret = git_deflate(stream, flush ? Z_FINISH : 0);\n-\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n+\talgo->update_fn(c, in0, stream->next_in - in0);\n+\tif (compat && compat_c)\n+\t\tcompat->update_fn(compat_c, in0, stream->next_in - in0);\n \tif (write_in_full(fd, compressed, stream->next_out - compressed) < 0)\n \t\tdie_errno(_(\"unable to write loose object file\"));\n \tstream->next_out = compressed;\n@@ -2014,15 +2028,21 @@ static int write_loose_object_common(git_hash_ctx *c,\n  * - End the compression of zlib stream.\n  * - Get the calculated oid to \"oid\".\n  */\n-static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n-\t\t\t\t   struct object_id *oid)\n+static int end_loose_object_common(git_hash_ctx *c, git_hash_ctx *compat_c,\n+\t\t\t\t   git_zstream *stream, struct object_id *oid,\n+\t\t\t\t   struct object_id *compat_oid)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tint ret;\n \n \tret = git_deflate_end_gently(stream);\n \tif (ret != Z_OK)\n \t\treturn ret;\n-\tthe_hash_algo->final_oid_fn(oid, c);\n+\talgo->final_oid_fn(oid, c);\n+\tif (compat && compat_c)\n+\t\tcompat->final_oid_fn(compat_oid, compat_c);\n \n \treturn Z_OK;\n }\n@@ -2046,7 +2066,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n \t\t\t\t       &stream, compressed, sizeof(compressed),\n-\t\t\t\t       &c, hdr, hdrlen);\n+\t\t\t\t       &c, NULL, hdr, hdrlen);\n \tif (fd < 0)\n \t\treturn -1;\n \n@@ -2056,14 +2076,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n \n-\t\tret = write_loose_object_common(&c, &stream, 1, in0, fd,\n+\t\tret = write_loose_object_common(&c, NULL, &stream, 1, in0, fd,\n \t\t\t\t\t\tcompressed, sizeof(compressed));\n \t} while (ret == Z_OK);\n \n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n-\tret = end_loose_object_common(&c, &stream, &parano_oid);\n+\tret = end_loose_object_common(&c, NULL, &stream, &parano_oid, NULL);\n \tif (ret != Z_OK)\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n@@ -2108,10 +2128,12 @@ static int freshen_packed_object(const struct object_id *oid)\n int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t\t\tstruct object_id *oid)\n {\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n+\tstruct object_id compat_oid;\n \tint fd, ret, err = 0, flush = 0;\n \tunsigned char compressed[4096];\n \tgit_zstream stream;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, compat_c;\n \tstruct strbuf tmp_file = STRBUF_INIT;\n \tstruct strbuf filename = STRBUF_INIT;\n \tint dirlen;\n@@ -2135,7 +2157,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t */\n \tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n \t\t\t\t       &stream, compressed, sizeof(compressed),\n-\t\t\t\t       &c, hdr, hdrlen);\n+\t\t\t\t       &c, &compat_c, hdr, hdrlen);\n \tif (fd < 0) {\n \t\terr = -1;\n \t\tgoto cleanup;\n@@ -2153,7 +2175,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t\t\tif (in_stream->is_finished)\n \t\t\t\tflush = 1;\n \t\t}\n-\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n+\t\tret = write_loose_object_common(&c, &compat_c, &stream, flush, in0, fd,\n \t\t\t\t\t\tcompressed, sizeof(compressed));\n \t\t/*\n \t\t * Unlike write_loose_object(), we do not have the entire\n@@ -2176,7 +2198,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t */\n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n-\tret = end_loose_object_common(&c, &stream, oid);\n+\tret = end_loose_object_common(&c, &compat_c, &stream, oid, &compat_oid);\n \tif (ret != Z_OK)\n \t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n \tclose_loose_object(fd, tmp_file.buf);\n@@ -2203,6 +2225,8 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t}\n \n \terr = finalize_object_file(tmp_file.buf, filename.buf);\n+\tif (!err && compat)\n+\t\terr = repo_add_loose_object_map(the_repository, oid, &compat_oid);\n cleanup:\n \tstrbuf_release(&tmp_file);\n \tstrbuf_release(&filename);\n@@ -2213,17 +2237,38 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n \t\t\t    unsigned flags)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct object_id compat_oid;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen = sizeof(hdr);\n \n+\t/* Generate compat_oid */\n+\tif (compat) {\n+\t\tif (type == OBJ_BLOB)\n+\t\t\thash_object_file(compat, buf, len, type, &compat_oid);\n+\t\telse {\n+\t\t\tstruct strbuf converted = STRBUF_INIT;\n+\t\t\tconvert_object_file(&converted, algo, compat,\n+\t\t\t\t\t    buf, len, type, 0);\n+\t\t\thash_object_file(compat, converted.buf, converted.len,\n+\t\t\t\t\t type, &compat_oid);\n+\t\t\tstrbuf_release(&converted);\n+\t\t}\n+\t}\n+\n \t/* Normally if we have it in the pack then we do not bother writing\n \t * it out into .git/objects/??/?{38} file.\n \t */\n-\twrite_object_file_prepare(the_hash_algo, buf, len, type, oid, hdr,\n-\t\t\t\t  &hdrlen);\n+\twrite_object_file_prepare(algo, buf, len, type, oid, hdr, &hdrlen);\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\treturn 0;\n-\treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n+\tif (write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags))\n+\t\treturn -1;\n+\tif (compat)\n+\t\treturn repo_add_loose_object_map(repo, oid, &compat_oid);\n+\treturn 0;\n }\n \n int write_object_file_literally(const void *buf, unsigned long len,\n@@ -2231,7 +2276,27 @@ int write_object_file_literally(const void *buf, unsigned long len,\n \t\t\t\tunsigned flags)\n {\n \tchar *header;\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct object_id compat_oid;\n \tint hdrlen, status = 0;\n+\tint compat_type = -1;\n+\n+\tif (compat) {\n+\t\tcompat_type = type_from_string_gently(type, -1, 1);\n+\t\tif (compat_type == OBJ_BLOB)\n+\t\t\thash_object_file(compat, buf, len, compat_type,\n+\t\t\t\t\t &compat_oid);\n+\t\telse if (compat_type != -1) {\n+\t\t\tstruct strbuf converted = STRBUF_INIT;\n+\t\t\tconvert_object_file(&converted, algo, compat,\n+\t\t\t\t\t    buf, len, compat_type, 0);\n+\t\t\thash_object_file(compat, converted.buf, converted.len,\n+\t\t\t\t\t compat_type, &compat_oid);\n+\t\t\tstrbuf_release(&converted);\n+\t\t}\n+\t}\n \n \t/* type string, SP, %lu of the length plus NUL must fit this */\n \thdrlen = strlen(type) + MAX_HEADER_LEN;\n@@ -2244,6 +2309,8 @@ int write_object_file_literally(const void *buf, unsigned long len,\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\tgoto cleanup;\n \tstatus = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n+\tif (compat_type != -1)\n+\t\treturn repo_add_loose_object_map(repo, oid, &compat_oid);\n \n cleanup:\n \tfree(header);\n@@ -2252,9 +2319,12 @@ int write_object_file_literally(const void *buf, unsigned long len,\n \n int force_object_loose(const struct object_id *oid, time_t mtime)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tvoid *buf;\n \tunsigned long len;\n \tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct object_id compat_oid;\n \tenum object_type type;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n@@ -2267,8 +2337,15 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \toi.contentp = &buf;\n \tif (oid_object_info_extended(the_repository, oid, &oi, 0))\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n+\tif (compat) {\n+\t\tif (repo_oid_to_algop(repo, oid, compat, &compat_oid))\n+\t\t\treturn error(_(\"cannot map object %s to %s\"),\n+\t\t\t\t     oid_to_hex(oid), compat->name);\n+\t}\n \thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n+\tif (!ret && compat)\n+\t\tret = repo_add_loose_object_map(the_repository, oid, &compat_oid);\n \tfree(buf);\n \n \treturn ret;\n-- \n2.41.0\n\n"},{"id":"482494","messageId":"20231002024034.2611-8-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 08/30] object-file: add a compat_oid_in parameter to write_object_file_flags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:12Z","receivedAt":"2023-10-02T02:40:57Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nTo create the proper signatures for commit objects both versions of\nthe commit object need to be generated and signed.  After that it is\na waste to throw away the work of generating the compatibility hash\nso update write_object_file_flags to take a compatibility hash input\nparameter that it can use to skip the work of generating the\ncompatability hash.\n\nUpdate the places that don't generate the compatability hash to\npass NULL so it is easy to tell write_object_file_flags should\nnot attempt to use their compatability hash.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n cache-tree.c      | 2 +-\n object-file.c     | 6 ++++--\n object-store-ll.h | 4 ++--\n 3 files changed, 7 insertions(+), 5 deletions(-)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 641427ed410a..ddc7d3d86959 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -448,7 +448,7 @@ static int update_one(struct cache_tree *it,\n \t\thash_object_file(the_hash_algo, buffer.buf, buffer.len,\n \t\t\t\t OBJ_TREE, &it->oid);\n \t} else if (write_object_file_flags(buffer.buf, buffer.len, OBJ_TREE,\n-\t\t\t\t\t   &it->oid, flags & WRITE_TREE_SILENT\n+\t\t\t\t\t   &it->oid, NULL, flags & WRITE_TREE_SILENT\n \t\t\t\t\t   ? HASH_SILENT : 0)) {\n \t\tstrbuf_release(&buffer);\n \t\treturn -1;\ndiff --git a/object-file.c b/object-file.c\nindex 4e55f475b3b4..820810a5f4b3 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2235,7 +2235,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n-\t\t\t    unsigned flags)\n+\t\t\t    struct object_id *compat_oid_in, unsigned flags)\n {\n \tstruct repository *repo = the_repository;\n \tconst struct git_hash_algo *algo = repo->hash_algo;\n@@ -2246,7 +2246,9 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \n \t/* Generate compat_oid */\n \tif (compat) {\n-\t\tif (type == OBJ_BLOB)\n+\t\tif (compat_oid_in)\n+\t\t\toidcpy(&compat_oid, compat_oid_in);\n+\t\telse if (type == OBJ_BLOB)\n \t\t\thash_object_file(compat, buf, len, type, &compat_oid);\n \t\telse {\n \t\t\tstruct strbuf converted = STRBUF_INIT;\ndiff --git a/object-store-ll.h b/object-store-ll.h\nindex bc76d6bec80d..c5f2bb2fc2fe 100644\n--- a/object-store-ll.h\n+++ b/object-store-ll.h\n@@ -255,11 +255,11 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n-\t\t\t    unsigned flags);\n+\t\t\t    struct object_id *comapt_oid_in, unsigned flags);\n static inline int write_object_file(const void *buf, unsigned long len,\n \t\t\t\t    enum object_type type, struct object_id *oid)\n {\n-\treturn write_object_file_flags(buf, len, type, oid, 0);\n+\treturn write_object_file_flags(buf, len, type, oid, NULL, 0);\n }\n \n int write_object_file_literally(const void *buf, unsigned long len,\n-- \n2.41.0\n\n"},{"id":"482495","messageId":"20231002024034.2611-10-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 10/30] commit: convert mergetag before computing the signature of a commit","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:14Z","receivedAt":"2023-10-02T02:40:59Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nIt so happens that commit mergetag lines embed a tag object.  So to\ncompute the compatible signature of a commit object that has mergetag\nlines the compatible embedded tag must be computed first.\n\nImplement this by duplicating and converting the commit extra headers\ninto the compatible version of the commit extra headers, that need\nto be passed to commit_tree_extended.\n\nTo handle merge tags only the compatible extra headers need to be\ncomputed.\n\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n commit.c | 42 +++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 41 insertions(+), 1 deletion(-)\n\ndiff --git a/commit.c b/commit.c\nindex 6765f3a82b9d..913e015966b4 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1355,6 +1355,39 @@ void append_merge_tag_headers(struct commit_list *parents,\n \t}\n }\n \n+static int convert_commit_extra_headers(struct commit_extra_header *orig,\n+\t\t\t\t\tstruct commit_extra_header **result)\n+{\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n+\tconst struct git_hash_algo *algo = the_repository->hash_algo;\n+\tstruct commit_extra_header *extra = NULL, **tail = &extra;\n+\tstruct strbuf out = STRBUF_INIT;\n+\twhile (orig) {\n+\t\tstruct commit_extra_header *new;\n+\t\tCALLOC_ARRAY(new, 1);\n+\t\tif (!strcmp(orig->key, \"mergetag\")) {\n+\t\t\tif (convert_object_file(&out, algo, compat,\n+\t\t\t\t\t\torig->value, orig->len,\n+\t\t\t\t\t\tOBJ_TAG, 1)) {\n+\t\t\t\tfree(new);\n+\t\t\t\tfree_commit_extra_headers(extra);\n+\t\t\t\treturn -1;\n+\t\t\t}\n+\t\t\tnew->key = xstrdup(\"mergetag\");\n+\t\t\tnew->value = strbuf_detach(&out, &new->len);\n+\t\t} else {\n+\t\t\tnew->key = xstrdup(orig->key);\n+\t\t\tnew->len = orig->len;\n+\t\t\tnew->value = xmemdupz(orig->value, orig->len);\n+\t\t}\n+\t\t*tail = new;\n+\t\ttail = &new->next;\n+\t\torig = orig->next;\n+\t}\n+\t*result = extra;\n+\treturn 0;\n+}\n+\n static void add_extra_header(struct strbuf *buffer,\n \t\t\t     struct commit_extra_header *extra)\n {\n@@ -1679,6 +1712,7 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n \t\tgoto out;\n \t}\n \tif (r->compat_hash_algo) {\n+\t\tstruct commit_extra_header *compat_extra = NULL;\n \t\tstruct object_id mapped_tree;\n \t\tstruct object_id *mapped_parents;\n \n@@ -1695,8 +1729,14 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n \t\t\t\tfree(mapped_parents);\n \t\t\t\tgoto out;\n \t\t\t}\n+\t\tif (convert_commit_extra_headers(extra, &compat_extra)) {\n+\t\t\tresult = -1;\n+\t\t\tfree(mapped_parents);\n+\t\t\tgoto out;\n+\t\t}\n \t\twrite_commit_tree(&compat_buffer, msg, msg_len, &mapped_tree,\n-\t\t\t\t  mapped_parents, nparents, author, committer, extra);\n+\t\t\t\t  mapped_parents, nparents, author, committer, compat_extra);\n+\t\tfree_commit_extra_headers(compat_extra);\n \t\tfree(mapped_parents);\n \n \t\tif (sign_commit && sign_commit_to_strbuf(&compat_sig, &compat_buffer, sign_commit)) {\n-- \n2.41.0\n\n"},{"id":"482496","messageId":"20231002024034.2611-9-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 09/30] commit: write commits for both hashes","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:13Z","receivedAt":"2023-10-02T02:41:01Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nWhen we write a commit, we include data that is specific to the hash\nalgorithm, such as parents and the root tree.  In order to write both a\nSHA-1 commit and a SHA-256 version, we need to convert between them.\n\nHowever, a straightforward conversion isn't necessarily what we want.\nWhen we sign a commit, we sign its data, so if we create a commit for\nSHA-256 and then write a SHA-1 version, we'll still have only signed the\nSHA-256 data.  While this is valid, it would be better to sign both\nforms of data so people using SHA-1 can verify the signatures as well.\n\nConsequently, we don't want to use the standard mapping that occurs when\nwe write an object.  Instead, let's move most of the writing of the\ncommit into a separate function which is agnostic of the hash algorithm\nand which simply writes into a buffer and specify both versions of the\nobject ourselves.\n\nWe can then call this function twice: once with the SHA-256 contents,\nand if SHA-1 is enabled, once with the SHA-1 contents.  If we're signing\nthe commit, we then sign both versions and append both signatures to\nboth buffers.  To produce a consistent hash, we always append the\nsignatures in the order in which Git implemented them: first SHA-1, then\nSHA-256.\n\nIn order to make this signing code work, we split the commit signing\ncode into two functions, one which signs the buffer, and one which\nappends the signature.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n commit.c | 181 +++++++++++++++++++++++++++++++++++++++++--------------\n 1 file changed, 136 insertions(+), 45 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex b3223478bc2a..6765f3a82b9d 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -28,6 +28,7 @@\n #include \"shallow.h\"\n #include \"tree.h\"\n #include \"hook.h\"\n+#include \"object-file-convert.h\"\n \n static struct commit_extra_header *read_commit_extra_header_lines(const char *buf, size_t len, const char **);\n \n@@ -1100,12 +1101,11 @@ static const char *gpg_sig_headers[] = {\n \t\"gpgsig-sha256\",\n };\n \n-int sign_with_header(struct strbuf *buf, const char *keyid)\n+static int add_commit_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo)\n {\n-\tstruct strbuf sig = STRBUF_INIT;\n \tint inspos, copypos;\n \tconst char *eoh;\n-\tconst char *gpg_sig_header = gpg_sig_headers[hash_algo_by_ptr(the_hash_algo)];\n+\tconst char *gpg_sig_header = gpg_sig_headers[hash_algo_by_ptr(algo)];\n \tint gpg_sig_header_len = strlen(gpg_sig_header);\n \n \t/* find the end of the header */\n@@ -1115,15 +1115,8 @@ int sign_with_header(struct strbuf *buf, const char *keyid)\n \telse\n \t\tinspos = eoh - buf->buf + 1;\n \n-\tif (!keyid || !*keyid)\n-\t\tkeyid = get_signing_key();\n-\tif (sign_buffer(buf, &sig, keyid)) {\n-\t\tstrbuf_release(&sig);\n-\t\treturn -1;\n-\t}\n-\n-\tfor (copypos = 0; sig.buf[copypos]; ) {\n-\t\tconst char *bol = sig.buf + copypos;\n+\tfor (copypos = 0; sig->buf[copypos]; ) {\n+\t\tconst char *bol = sig->buf + copypos;\n \t\tconst char *eol = strchrnul(bol, '\\n');\n \t\tint len = (eol - bol) + !!*eol;\n \n@@ -1136,11 +1129,17 @@ int sign_with_header(struct strbuf *buf, const char *keyid)\n \t\tinspos += len;\n \t\tcopypos += len;\n \t}\n-\tstrbuf_release(&sig);\n \treturn 0;\n }\n \n-\n+static int sign_commit_to_strbuf(struct strbuf *sig, struct strbuf *buf, const char *keyid)\n+{\n+\tif (!keyid || !*keyid)\n+\t\tkeyid = get_signing_key();\n+\tif (sign_buffer(buf, sig, keyid))\n+\t\treturn -1;\n+\treturn 0;\n+}\n \n int parse_signed_commit(const struct commit *commit,\n \t\t\tstruct strbuf *payload, struct strbuf *signature,\n@@ -1599,70 +1598,162 @@ N_(\"Warning: commit message did not conform to UTF-8.\\n\"\n    \"You may want to amend it after fixing the message, or set the config\\n\"\n    \"variable i18n.commitEncoding to the encoding your project uses.\\n\");\n \n-int commit_tree_extended(const char *msg, size_t msg_len,\n-\t\t\t const struct object_id *tree,\n-\t\t\t struct commit_list *parents, struct object_id *ret,\n-\t\t\t const char *author, const char *committer,\n-\t\t\t const char *sign_commit,\n-\t\t\t struct commit_extra_header *extra)\n+static void write_commit_tree(struct strbuf *buffer, const char *msg, size_t msg_len,\n+\t\t\t      const struct object_id *tree,\n+\t\t\t      const struct object_id *parents, size_t parents_len,\n+\t\t\t      const char *author, const char *committer,\n+\t\t\t      struct commit_extra_header *extra)\n {\n-\tint result;\n \tint encoding_is_utf8;\n-\tstruct strbuf buffer;\n-\n-\tassert_oid_type(tree, OBJ_TREE);\n-\n-\tif (memchr(msg, '\\0', msg_len))\n-\t\treturn error(\"a NUL byte in commit log message not allowed.\");\n+\tsize_t i;\n \n \t/* Not having i18n.commitencoding is the same as having utf-8 */\n \tencoding_is_utf8 = is_encoding_utf8(git_commit_encoding);\n \n-\tstrbuf_init(&buffer, 8192); /* should avoid reallocs for the headers */\n-\tstrbuf_addf(&buffer, \"tree %s\\n\", oid_to_hex(tree));\n+\tstrbuf_grow(buffer, 8192); /* should avoid reallocs for the headers */\n+\tstrbuf_addf(buffer, \"tree %s\\n\", oid_to_hex(tree));\n \n \t/*\n \t * NOTE! This ordering means that the same exact tree merged with a\n \t * different order of parents will be a _different_ changeset even\n \t * if everything else stays the same.\n \t */\n-\twhile (parents) {\n-\t\tstruct commit *parent = pop_commit(&parents);\n-\t\tstrbuf_addf(&buffer, \"parent %s\\n\",\n-\t\t\t    oid_to_hex(&parent->object.oid));\n-\t}\n+\tfor (i = 0; i < parents_len; i++)\n+\t\tstrbuf_addf(buffer, \"parent %s\\n\", oid_to_hex(&parents[i]));\n \n \t/* Person/date information */\n \tif (!author)\n \t\tauthor = git_author_info(IDENT_STRICT);\n-\tstrbuf_addf(&buffer, \"author %s\\n\", author);\n+\tstrbuf_addf(buffer, \"author %s\\n\", author);\n \tif (!committer)\n \t\tcommitter = git_committer_info(IDENT_STRICT);\n-\tstrbuf_addf(&buffer, \"committer %s\\n\", committer);\n+\tstrbuf_addf(buffer, \"committer %s\\n\", committer);\n \tif (!encoding_is_utf8)\n-\t\tstrbuf_addf(&buffer, \"encoding %s\\n\", git_commit_encoding);\n+\t\tstrbuf_addf(buffer, \"encoding %s\\n\", git_commit_encoding);\n \n \twhile (extra) {\n-\t\tadd_extra_header(&buffer, extra);\n+\t\tadd_extra_header(buffer, extra);\n \t\textra = extra->next;\n \t}\n-\tstrbuf_addch(&buffer, '\\n');\n+\tstrbuf_addch(buffer, '\\n');\n \n \t/* And add the comment */\n-\tstrbuf_add(&buffer, msg, msg_len);\n+\tstrbuf_add(buffer, msg, msg_len);\n+}\n \n-\t/* And check the encoding */\n-\tif (encoding_is_utf8 && !verify_utf8(&buffer))\n-\t\tfprintf(stderr, _(commit_utf8_warn));\n+int commit_tree_extended(const char *msg, size_t msg_len,\n+\t\t\t const struct object_id *tree,\n+\t\t\t struct commit_list *parents, struct object_id *ret,\n+\t\t\t const char *author, const char *committer,\n+\t\t\t const char *sign_commit,\n+\t\t\t struct commit_extra_header *extra)\n+{\n+\tstruct repository *r = the_repository;\n+\tint result = 0;\n+\tint encoding_is_utf8;\n+\tstruct strbuf buffer = STRBUF_INIT, compat_buffer = STRBUF_INIT;\n+\tstruct strbuf sig = STRBUF_INIT, compat_sig = STRBUF_INIT;\n+\tstruct object_id *parent_buf = NULL, *compat_oid = NULL;\n+\tstruct object_id compat_oid_buf;\n+\tsize_t i, nparents;\n+\n+\t/* Not having i18n.commitencoding is the same as having utf-8 */\n+\tencoding_is_utf8 = is_encoding_utf8(git_commit_encoding);\n+\n+\tassert_oid_type(tree, OBJ_TREE);\n+\n+\tif (memchr(msg, '\\0', msg_len))\n+\t\treturn error(\"a NUL byte in commit log message not allowed.\");\n \n-\tif (sign_commit && sign_with_header(&buffer, sign_commit)) {\n+\tnparents = commit_list_count(parents);\n+\tCALLOC_ARRAY(parent_buf, nparents);\n+\ti = 0;\n+\twhile (parents) {\n+\t\tstruct commit *parent = pop_commit(&parents);\n+\t\toidcpy(&parent_buf[i++], &parent->object.oid);\n+\t}\n+\n+\twrite_commit_tree(&buffer, msg, msg_len, tree, parent_buf, nparents, author, committer, extra);\n+\tif (sign_commit && sign_commit_to_strbuf(&sig, &buffer, sign_commit)) {\n \t\tresult = -1;\n \t\tgoto out;\n \t}\n+\tif (r->compat_hash_algo) {\n+\t\tstruct object_id mapped_tree;\n+\t\tstruct object_id *mapped_parents;\n+\n+\t\tCALLOC_ARRAY(mapped_parents, nparents);\n+\n+\t\tif (repo_oid_to_algop(r, tree, r->compat_hash_algo, &mapped_tree)) {\n+\t\t\tresult = -1;\n+\t\t\tfree(mapped_parents);\n+\t\t\tgoto out;\n+\t\t}\n+\t\tfor (i = 0; i < nparents; i++)\n+\t\t\tif (repo_oid_to_algop(r, &parent_buf[i], r->compat_hash_algo, &mapped_parents[i])) {\n+\t\t\t\tresult = -1;\n+\t\t\t\tfree(mapped_parents);\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\t\twrite_commit_tree(&compat_buffer, msg, msg_len, &mapped_tree,\n+\t\t\t\t  mapped_parents, nparents, author, committer, extra);\n+\t\tfree(mapped_parents);\n+\n+\t\tif (sign_commit && sign_commit_to_strbuf(&compat_sig, &compat_buffer, sign_commit)) {\n+\t\t\tresult = -1;\n+\t\t\tgoto out;\n+\t\t}\n+\t}\n+\n+\tif (sign_commit) {\n+\t\tstruct sig_pairs {\n+\t\t\tstruct strbuf *sig;\n+\t\t\tconst struct git_hash_algo *algo;\n+\t\t} bufs [2] = {\n+\t\t\t{ &compat_sig, r->compat_hash_algo },\n+\t\t\t{ &sig, r->hash_algo },\n+\t\t};\n+\t\tint i;\n+\n+\t\t/*\n+\t\t * We write algorithms in the order they were implemented in\n+\t\t * Git to produce a stable hash when multiple algorithms are\n+\t\t * used.\n+\t\t */\n+\t\tif (r->compat_hash_algo && hash_algo_by_ptr(bufs[0].algo) > hash_algo_by_ptr(bufs[1].algo))\n+\t\t\tSWAP(bufs[0], bufs[1]);\n+\n+\t\t/*\n+\t\t * We traverse each algorithm in order, and apply the signature\n+\t\t * to each buffer.\n+\t\t */\n+\t\tfor (i = 0; i < ARRAY_SIZE(bufs); i++) {\n+\t\t\tif (!bufs[i].algo)\n+\t\t\t\tcontinue;\n+\t\t\tadd_commit_signature(&buffer, bufs[i].sig, bufs[i].algo);\n+\t\t\tif (r->compat_hash_algo)\n+\t\t\t\tadd_commit_signature(&compat_buffer, bufs[i].sig, bufs[i].algo);\n+\t\t}\n+\t}\n \n-\tresult = write_object_file(buffer.buf, buffer.len, OBJ_COMMIT, ret);\n+\t/* And check the encoding. */\n+\tif (encoding_is_utf8 && (!verify_utf8(&buffer) || !verify_utf8(&compat_buffer)))\n+\t\tfprintf(stderr, _(commit_utf8_warn));\n+\n+\tif (r->compat_hash_algo) {\n+\t\thash_object_file(r->compat_hash_algo, compat_buffer.buf, compat_buffer.len,\n+\t\t\tOBJ_COMMIT, &compat_oid_buf);\n+\t\tcompat_oid = &compat_oid_buf;\n+\t}\n+\n+\tresult = write_object_file_flags(buffer.buf, buffer.len, OBJ_COMMIT,\n+\t\t\t\t\t ret, compat_oid, 0);\n out:\n+\tfree(parent_buf);\n \tstrbuf_release(&buffer);\n+\tstrbuf_release(&compat_buffer);\n+\tstrbuf_release(&sig);\n+\tstrbuf_release(&compat_sig);\n \treturn result;\n }\n \n-- \n2.41.0\n\n"},{"id":"482497","messageId":"20231002024034.2611-11-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 11/30] commit: export add_header_signature to support handling signatures on tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:15Z","receivedAt":"2023-10-02T02:41:03Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nRename add_commit_signature as add_header_signature, and expose it so\nthat it can be used for converting tags from one object format to\nanother.\n\nInspired-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n commit.c | 6 +++---\n commit.h | 1 +\n 2 files changed, 4 insertions(+), 3 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex 913e015966b4..2b61a4d0aa11 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1101,7 +1101,7 @@ static const char *gpg_sig_headers[] = {\n \t\"gpgsig-sha256\",\n };\n \n-static int add_commit_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo)\n+int add_header_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo)\n {\n \tint inspos, copypos;\n \tconst char *eoh;\n@@ -1770,9 +1770,9 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n \t\tfor (i = 0; i < ARRAY_SIZE(bufs); i++) {\n \t\t\tif (!bufs[i].algo)\n \t\t\t\tcontinue;\n-\t\t\tadd_commit_signature(&buffer, bufs[i].sig, bufs[i].algo);\n+\t\t\tadd_header_signature(&buffer, bufs[i].sig, bufs[i].algo);\n \t\t\tif (r->compat_hash_algo)\n-\t\t\t\tadd_commit_signature(&compat_buffer, bufs[i].sig, bufs[i].algo);\n+\t\t\t\tadd_header_signature(&compat_buffer, bufs[i].sig, bufs[i].algo);\n \t\t}\n \t}\n \ndiff --git a/commit.h b/commit.h\nindex 28928833c544..03edcec0129f 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -370,5 +370,6 @@ int parse_buffer_signed_by_header(const char *buffer,\n \t\t\t\t  struct strbuf *payload,\n \t\t\t\t  struct strbuf *signature,\n \t\t\t\t  const struct git_hash_algo *algop);\n+int add_header_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo);\n \n #endif /* COMMIT_H */\n-- \n2.41.0\n\n"},{"id":"482498","messageId":"20231002024034.2611-12-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 12/30] tag: sign both hashes","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:16Z","receivedAt":"2023-10-02T02:41:04Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nWhen we write a tag the object oid is specific to the hash algorithm.\n\nThis matters when a tag is signed.  The hash transition plan calls for\nsignatures on both the sha1 form and the sha256 form of the object,\nand for both of those signatures to live in the tag object.\n\nTo generate tag object with multiple signatures, first compute the\nunsigned form of the tag, and then if the tag is being signed compute\nthe unsigned form of the tag with the compatibilityr hash.  Then\ncompute compute the signatures of both buffers.\n\nOnce the signatures are computed add them to both buffers.  This\nallows computing the compatibility hash in do_sign, saving\nwrite_object_file the expense of recomputing the compatibility tag\njust to compute it's hash.\n\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n builtin/tag.c | 45 +++++++++++++++++++++++++++++++++++++++++----\n 1 file changed, 41 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/tag.c b/builtin/tag.c\nindex 3918eacbb57b..8c4bc28952c2 100644\n--- a/builtin/tag.c\n+++ b/builtin/tag.c\n@@ -28,6 +28,7 @@\n #include \"ref-filter.h\"\n #include \"date.h\"\n #include \"write-or-die.h\"\n+#include \"object-file-convert.h\"\n \n static const char * const git_tag_usage[] = {\n \tN_(\"git tag [-a | -s | -u <key-id>] [-f] [-m <msg> | -F <file>] [-e]\\n\"\n@@ -174,9 +175,43 @@ static int verify_tag(const char *name, const char *ref UNUSED,\n \treturn 0;\n }\n \n-static int do_sign(struct strbuf *buffer)\n+static int do_sign(struct strbuf *buffer, struct object_id **compat_oid,\n+\t\t   struct object_id *compat_oid_buf)\n {\n-\treturn sign_buffer(buffer, buffer, get_signing_key());\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n+\tstruct strbuf sig = STRBUF_INIT, compat_sig = STRBUF_INIT;\n+\tstruct strbuf compat_buf = STRBUF_INIT;\n+\tconst char *keyid = get_signing_key();\n+\tint ret = -1;\n+\n+\tif (sign_buffer(buffer, &sig, keyid))\n+\t\treturn -1;\n+\n+\tif (compat) {\n+\t\tconst struct git_hash_algo *algo = the_repository->hash_algo;\n+\n+\t\tif (convert_object_file(&compat_buf, algo, compat,\n+\t\t\t\t\tbuffer->buf, buffer->len, OBJ_TAG, 1))\n+\t\t\tgoto out;\n+\t\tif (sign_buffer(&compat_buf, &compat_sig, keyid))\n+\t\t\tgoto out;\n+\t\tadd_header_signature(&compat_buf, &sig, algo);\n+\t\tstrbuf_addbuf(&compat_buf, &compat_sig);\n+\t\thash_object_file(compat, compat_buf.buf, compat_buf.len,\n+\t\t\t\t OBJ_TAG, compat_oid_buf);\n+\t\t*compat_oid = compat_oid_buf;\n+\t}\n+\n+\tif (compat_sig.len)\n+\t\tadd_header_signature(buffer, &compat_sig, compat);\n+\n+\tstrbuf_addbuf(buffer, &sig);\n+\tret = 0;\n+out:\n+\tstrbuf_release(&sig);\n+\tstrbuf_release(&compat_sig);\n+\tstrbuf_release(&compat_buf);\n+\treturn ret;\n }\n \n static const char tag_template[] =\n@@ -249,9 +284,11 @@ static void write_tag_body(int fd, const struct object_id *oid)\n \n static int build_tag_object(struct strbuf *buf, int sign, struct object_id *result)\n {\n-\tif (sign && do_sign(buf) < 0)\n+\tstruct object_id *compat_oid = NULL, compat_oid_buf;\n+\tif (sign && do_sign(buf, &compat_oid, &compat_oid_buf) < 0)\n \t\treturn error(_(\"unable to sign the tag\"));\n-\tif (write_object_file(buf->buf, buf->len, OBJ_TAG, result) < 0)\n+\tif (write_object_file_flags(buf->buf, buf->len, OBJ_TAG, result,\n+\t\t\t\t    compat_oid, 0) < 0)\n \t\treturn error(_(\"unable to write tag file\"));\n \treturn 0;\n }\n-- \n2.41.0\n\n"},{"id":"482499","messageId":"20231002024034.2611-13-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 13/30] cache: add a function to read an OID of a specific algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:17Z","receivedAt":"2023-10-02T02:41:14Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nCurrently, we always read a object ID of the current algorithm with\noidread.  However, once we start converting objects, we'll need to\nconsider what happens when we want to read an object ID of a specific\nalgorithm, such as the compatibility algorithm.  To make this easier,\nlet's define oidread_algop, which specifies which algorithm we should\nuse for our object ID, and define oidread in terms of it.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n hash.h | 9 +++++++--\n 1 file changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/hash.h b/hash.h\nindex 615ae0691d07..e064807c1733 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -73,10 +73,15 @@ static inline void oidclr(struct object_id *oid)\n \toid->algo = hash_algo_by_ptr(the_hash_algo);\n }\n \n+static inline void oidread_algop(struct object_id *oid, const unsigned char *hash, const struct git_hash_algo *algop)\n+{\n+\tmemcpy(oid->hash, hash, algop->rawsz);\n+\toid->algo = hash_algo_by_ptr(algop);\n+}\n+\n static inline void oidread(struct object_id *oid, const unsigned char *hash)\n {\n-\tmemcpy(oid->hash, hash, the_hash_algo->rawsz);\n-\toid->algo = hash_algo_by_ptr(the_hash_algo);\n+\toidread_algop(oid, hash, the_hash_algo);\n }\n \n static inline int is_empty_blob_sha1(const unsigned char *sha1)\n-- \n2.41.0\n\n"},{"id":"482500","messageId":"20231002024034.2611-14-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 14/30] object: factor out parse_mode out of fast-import and tree-walk into in object.h","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:18Z","receivedAt":"2023-10-02T02:41:16Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nbuiltin/fast-import.c and tree-walk.c have almost identical version of\nget_mode.  The two functions started out the same but have diverged\nslightly.  The version in fast-import changed mode to a uint16_t to\nsave memory.  The version in tree-walk started erroring if no mode was\npresent.\n\nAs far as I can tell both of these changes are valid for both of the\ncallers, so add the both changes and place the common parsing helper\nin object.h\n\nRename the helper from get_mode to parse_mode so it does not\nconflict with another helper named get_mode in diff-no-index.c\n\nThis will be used shortly in a new helper decode_tree_entry_raw\nwhich is used to compute cmpatibility objects as part of\nthe sha256 transition.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/fast-import.c | 18 ++----------------\n object.h              | 18 ++++++++++++++++++\n tree-walk.c           | 22 +++-------------------\n 3 files changed, 23 insertions(+), 35 deletions(-)\n\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 4dbb10aff3da..2c645fcfbe3f 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -1235,20 +1235,6 @@ static void *gfi_unpack_entry(\n \treturn unpack_entry(the_repository, p, oe->idx.offset, &type, sizep);\n }\n \n-static const char *get_mode(const char *str, uint16_t *modep)\n-{\n-\tunsigned char c;\n-\tuint16_t mode = 0;\n-\n-\twhile ((c = *str++) != ' ') {\n-\t\tif (c < '0' || c > '7')\n-\t\t\treturn NULL;\n-\t\tmode = (mode << 3) + (c - '0');\n-\t}\n-\t*modep = mode;\n-\treturn str;\n-}\n-\n static void load_tree(struct tree_entry *root)\n {\n \tstruct object_id *oid = &root->versions[1].oid;\n@@ -1286,7 +1272,7 @@ static void load_tree(struct tree_entry *root)\n \t\tt->entries[t->entry_count++] = e;\n \n \t\te->tree = NULL;\n-\t\tc = get_mode(c, &e->versions[1].mode);\n+\t\tc = parse_mode(c, &e->versions[1].mode);\n \t\tif (!c)\n \t\t\tdie(\"Corrupt mode in %s\", oid_to_hex(oid));\n \t\te->versions[0].mode = e->versions[1].mode;\n@@ -2275,7 +2261,7 @@ static void file_change_m(const char *p, struct branch *b)\n \tstruct object_id oid;\n \tuint16_t mode, inline_data = 0;\n \n-\tp = get_mode(p, &mode);\n+\tp = parse_mode(p, &mode);\n \tif (!p)\n \t\tdie(\"Corrupt mode: %s\", command_buf.buf);\n \tswitch (mode) {\ndiff --git a/object.h b/object.h\nindex 114d45954d08..70c8d4ae63dc 100644\n--- a/object.h\n+++ b/object.h\n@@ -190,6 +190,24 @@ void *create_object(struct repository *r, const struct object_id *oid, void *obj\n \n void *object_as_type(struct object *obj, enum object_type type, int quiet);\n \n+\n+static inline const char *parse_mode(const char *str, uint16_t *modep)\n+{\n+\tunsigned char c;\n+\tunsigned int mode = 0;\n+\n+\tif (*str == ' ')\n+\t\treturn NULL;\n+\n+\twhile ((c = *str++) != ' ') {\n+\t\tif (c < '0' || c > '7')\n+\t\t\treturn NULL;\n+\t\tmode = (mode << 3) + (c - '0');\n+\t}\n+\t*modep = mode;\n+\treturn str;\n+}\n+\n /*\n  * Returns the object, having parsed it to find out what it is.\n  *\ndiff --git a/tree-walk.c b/tree-walk.c\nindex 29ead71be173..3af50a01c2c7 100644\n--- a/tree-walk.c\n+++ b/tree-walk.c\n@@ -10,27 +10,11 @@\n #include \"pathspec.h\"\n #include \"json-writer.h\"\n \n-static const char *get_mode(const char *str, unsigned int *modep)\n-{\n-\tunsigned char c;\n-\tunsigned int mode = 0;\n-\n-\tif (*str == ' ')\n-\t\treturn NULL;\n-\n-\twhile ((c = *str++) != ' ') {\n-\t\tif (c < '0' || c > '7')\n-\t\t\treturn NULL;\n-\t\tmode = (mode << 3) + (c - '0');\n-\t}\n-\t*modep = mode;\n-\treturn str;\n-}\n-\n static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned long size, struct strbuf *err)\n {\n \tconst char *path;\n-\tunsigned int mode, len;\n+\tunsigned int len;\n+\tuint16_t mode;\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \n \tif (size < hashsz + 3 || buf[size - (hashsz + 1)]) {\n@@ -38,7 +22,7 @@ static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned l\n \t\treturn -1;\n \t}\n \n-\tpath = get_mode(buf, &mode);\n+\tpath = parse_mode(buf, &mode);\n \tif (!path) {\n \t\tstrbuf_addstr(err, _(\"malformed mode in tree entry\"));\n \t\treturn -1;\n-- \n2.41.0\n\n"},{"id":"482501","messageId":"20231002024034.2611-15-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 15/30] object-file-convert: add a function to convert trees between algorithms","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:19Z","receivedAt":"2023-10-02T02:41:18Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nIn the future, we're going to want to provide SHA-256 repositories that\nhave compatibility support for SHA-1 as well.  In order to do so, we'll\nneed to be able to convert tree objects from SHA-256 to SHA-1 by writing\na tree with each SHA-256 object ID mapped to a SHA-1 object ID.\n\nWe implement a function, convert_tree_object, that takes an existing\ntree buffer and writes it to a new strbuf, converting between\nalgorithms.  Let's make this function generic, because while we only\nneed it to convert from the main algorithm to the compatibility\nalgorithm now, we may need to do the other way around in the future,\nsuch as for transport.\n\nWe avoid reusing the code in decode_tree_entry because that code\nnormalizes data, and we don't want that here.  We want to produce a\ncomplete round trip of data, so if, for example, the old entry had a\nwrongly zero-padded mode, we'd want to preserve that when converting to\nensure a stable hash value.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 51 ++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 50 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 1ec945eaa17f..70b80fb61e54 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -1,8 +1,10 @@\n #include \"git-compat-util.h\"\n #include \"gettext.h\"\n #include \"strbuf.h\"\n+#include \"hex.h\"\n #include \"repository.h\"\n #include \"hash-ll.h\"\n+#include \"hash.h\"\n #include \"object.h\"\n #include \"loose.h\"\n #include \"object-file-convert.h\"\n@@ -36,6 +38,51 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \treturn 0;\n }\n \n+static int decode_tree_entry_raw(struct object_id *oid, const char **path,\n+\t\t\t\t size_t *len, const struct git_hash_algo *algo,\n+\t\t\t\t const char *buf, unsigned long size)\n+{\n+\tuint16_t mode;\n+\tconst unsigned hashsz = algo->rawsz;\n+\n+\tif (size < hashsz + 3 || buf[size - (hashsz + 1)]) {\n+\t\treturn -1;\n+\t}\n+\n+\t*path = parse_mode(buf, &mode);\n+\tif (!*path || !**path)\n+\t\treturn -1;\n+\t*len = strlen(*path) + 1;\n+\n+\toidread_algop(oid, (const unsigned char *)*path + *len, algo);\n+\treturn 0;\n+}\n+\n+static int convert_tree_object(struct strbuf *out,\n+\t\t\t       const struct git_hash_algo *from,\n+\t\t\t       const struct git_hash_algo *to,\n+\t\t\t       const char *buffer, size_t size)\n+{\n+\tconst char *p = buffer, *end = buffer + size;\n+\n+\twhile (p < end) {\n+\t\tstruct object_id entry_oid, mapped_oid;\n+\t\tconst char *path = NULL;\n+\t\tsize_t pathlen;\n+\n+\t\tif (decode_tree_entry_raw(&entry_oid, &path, &pathlen, from, p,\n+\t\t\t\t\t  end - p))\n+\t\t\treturn error(_(\"failed to decode tree entry\"));\n+\t\tif (repo_oid_to_algop(the_repository, &entry_oid, to, &mapped_oid))\n+\t\t\treturn error(_(\"failed to map tree entry for %s\"), oid_to_hex(&entry_oid));\n+\t\tstrbuf_add(out, p, path - p);\n+\t\tstrbuf_add(out, path, pathlen);\n+\t\tstrbuf_add(out, mapped_oid.hash, to->rawsz);\n+\t\tp = path + pathlen + from->rawsz;\n+\t}\n+\treturn 0;\n+}\n+\n int convert_object_file(struct strbuf *outbuf,\n \t\t\tconst struct git_hash_algo *from,\n \t\t\tconst struct git_hash_algo *to,\n@@ -50,8 +97,10 @@ int convert_object_file(struct strbuf *outbuf,\n \t\tBUG(\"Refusing noop object file conversion\");\n \n \tswitch (type) {\n-\tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n+\t\tret = convert_tree_object(outbuf, from, to, buf, len);\n+\t\tbreak;\n+\tcase OBJ_COMMIT:\n \tcase OBJ_TAG:\n \tdefault:\n \t\t/* Not implemented yet, so fail. */\n-- \n2.41.0\n\n"},{"id":"482502","messageId":"20231002024034.2611-16-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 16/30] object-file-convert: convert tag objects when writing","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:20Z","receivedAt":"2023-10-02T02:41:19Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nWhen writing a tag object in a repository with both SHA-1 and SHA-256,\nwe'll need to convert our commit objects so that we can write the hash\nvalues for both into the repository.  To do so, let's add a function to\nconvert tag objects.\n\nNote that signatures for tag objects in the current algorithm trail the\nmessage, and those for the alternate algorithm are in headers.\nTherefore, we parse the tag object for both a trailing signature and a\nheader and then, when writing the other format, swap the two around.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 52 ++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 51 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 70b80fb61e54..089b68442de8 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -7,6 +7,8 @@\n #include \"hash.h\"\n #include \"object.h\"\n #include \"loose.h\"\n+#include \"commit.h\"\n+#include \"gpg-interface.h\"\n #include \"object-file-convert.h\"\n \n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n@@ -83,6 +85,52 @@ static int convert_tree_object(struct strbuf *out,\n \treturn 0;\n }\n \n+static int convert_tag_object(struct strbuf *out,\n+\t\t\t      const struct git_hash_algo *from,\n+\t\t\t      const struct git_hash_algo *to,\n+\t\t\t      const char *buffer, size_t size)\n+{\n+\tstruct strbuf payload = STRBUF_INIT, temp = STRBUF_INIT, oursig = STRBUF_INIT, othersig = STRBUF_INIT;\n+\tsize_t payload_size;\n+\tstruct object_id oid, mapped_oid;\n+\tconst char *p;\n+\n+\t/* Add some slop for longer signature header in the new algorithm. */\n+\tstrbuf_grow(out, size + 7);\n+\n+\t/* Is there a signature for our algorithm? */\n+\tpayload_size = parse_signed_buffer(buffer, size);\n+\tstrbuf_add(&payload, buffer, payload_size);\n+\tif (payload_size != size) {\n+\t\t/* Yes, there is. */\n+\t\tstrbuf_add(&oursig, buffer + payload_size, size - payload_size);\n+\t}\n+\t/* Now, is there a signature for the other algorithm? */\n+\tif (parse_buffer_signed_by_header(payload.buf, payload.len, &temp, &othersig, to)) {\n+\t\t/* Yes, there is. */\n+\t\tstrbuf_swap(&payload, &temp);\n+\t\tstrbuf_release(&temp);\n+\t}\n+\n+\t/*\n+\t * Our payload is now in payload and we may have up to two signatrures\n+\t * in oursig and othersig.\n+\t */\n+\tif (strncmp(payload.buf, \"object \", 7) || payload.buf[from->hexsz + 7] != '\\n')\n+\t\treturn error(\"bogus tag object\");\n+\tif (parse_oid_hex_algop(payload.buf + 7, &oid, &p, from) < 0)\n+\t\treturn error(\"bad tag object ID\");\n+\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\treturn error(\"unable to map tree %s in tag object\",\n+\t\t\t     oid_to_hex(&oid));\n+\tstrbuf_addf(out, \"object %s\", oid_to_hex(&mapped_oid));\n+\tstrbuf_add(out, p, payload.len - (p - payload.buf));\n+\tstrbuf_addbuf(out, &othersig);\n+\tif (oursig.len)\n+\t\tadd_header_signature(out, &oursig, from);\n+\treturn 0;\n+}\n+\n int convert_object_file(struct strbuf *outbuf,\n \t\t\tconst struct git_hash_algo *from,\n \t\t\tconst struct git_hash_algo *to,\n@@ -100,8 +148,10 @@ int convert_object_file(struct strbuf *outbuf,\n \tcase OBJ_TREE:\n \t\tret = convert_tree_object(outbuf, from, to, buf, len);\n \t\tbreak;\n-\tcase OBJ_COMMIT:\n \tcase OBJ_TAG:\n+\t\tret = convert_tag_object(outbuf, from, to, buf, len);\n+\t\tbreak;\n+\tcase OBJ_COMMIT:\n \tdefault:\n \t\t/* Not implemented yet, so fail. */\n \t\tret = -1;\n-- \n2.41.0\n\n"},{"id":"482503","messageId":"20231002024034.2611-17-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 17/30] object-file-convert: don't leak when converting tag objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:21Z","receivedAt":"2023-10-02T02:41:20Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpon close examination I discovered that while brian's code to convert\ntag objects was functionally correct, it leaked memory.\n\nRearrange the code so that all error checking happens before any\nmemory is allocated.\n\nAdd code to release the temporary strbufs the code uses.\n\nThe code pretty much assumes the tag object ends with a newline,\nso add an explict test to verify that is the case.\n\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 45 ++++++++++++++++++++++++-------------------\n 1 file changed, 25 insertions(+), 20 deletions(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 089b68442de8..79e8e211ff95 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -90,44 +90,49 @@ static int convert_tag_object(struct strbuf *out,\n \t\t\t      const struct git_hash_algo *to,\n \t\t\t      const char *buffer, size_t size)\n {\n-\tstruct strbuf payload = STRBUF_INIT, temp = STRBUF_INIT, oursig = STRBUF_INIT, othersig = STRBUF_INIT;\n+\tstruct strbuf payload = STRBUF_INIT, oursig = STRBUF_INIT, othersig = STRBUF_INIT;\n+\tconst int entry_len = from->hexsz + 7;\n \tsize_t payload_size;\n \tstruct object_id oid, mapped_oid;\n \tconst char *p;\n \n-\t/* Add some slop for longer signature header in the new algorithm. */\n-\tstrbuf_grow(out, size + 7);\n+\t/* Consume the object line */\n+\tif ((entry_len >= size) ||\n+\t    memcmp(buffer, \"object \", 7) || buffer[entry_len] != '\\n')\n+\t\treturn error(\"bogus tag object\");\n+\tif (parse_oid_hex_algop(buffer + 7, &oid, &p, from) < 0)\n+\t\treturn error(\"bad tag object ID\");\n+\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\treturn error(\"unable to map tree %s in tag object\",\n+\t\t\t     oid_to_hex(&oid));\n+\tsize -= ((p + 1) - buffer);\n+\tbuffer = p + 1;\n \n \t/* Is there a signature for our algorithm? */\n \tpayload_size = parse_signed_buffer(buffer, size);\n-\tstrbuf_add(&payload, buffer, payload_size);\n \tif (payload_size != size) {\n \t\t/* Yes, there is. */\n \t\tstrbuf_add(&oursig, buffer + payload_size, size - payload_size);\n \t}\n-\t/* Now, is there a signature for the other algorithm? */\n-\tif (parse_buffer_signed_by_header(payload.buf, payload.len, &temp, &othersig, to)) {\n-\t\t/* Yes, there is. */\n-\t\tstrbuf_swap(&payload, &temp);\n-\t\tstrbuf_release(&temp);\n-\t}\n \n+\t/* Now, is there a signature for the other algorithm? */\n+\tparse_buffer_signed_by_header(buffer, payload_size, &payload, &othersig, to);\n \t/*\n \t * Our payload is now in payload and we may have up to two signatrures\n \t * in oursig and othersig.\n \t */\n-\tif (strncmp(payload.buf, \"object \", 7) || payload.buf[from->hexsz + 7] != '\\n')\n-\t\treturn error(\"bogus tag object\");\n-\tif (parse_oid_hex_algop(payload.buf + 7, &oid, &p, from) < 0)\n-\t\treturn error(\"bad tag object ID\");\n-\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n-\t\treturn error(\"unable to map tree %s in tag object\",\n-\t\t\t     oid_to_hex(&oid));\n-\tstrbuf_addf(out, \"object %s\", oid_to_hex(&mapped_oid));\n-\tstrbuf_add(out, p, payload.len - (p - payload.buf));\n-\tstrbuf_addbuf(out, &othersig);\n+\n+\t/* Add some slop for longer signature header in the new algorithm. */\n+\tstrbuf_grow(out, (7 + to->hexsz + 1) + size + 7);\n+\tstrbuf_addf(out, \"object %s\\n\", oid_to_hex(&mapped_oid));\n+\tstrbuf_addbuf(out, &payload);\n \tif (oursig.len)\n \t\tadd_header_signature(out, &oursig, from);\n+\tstrbuf_addbuf(out, &othersig);\n+\n+\tstrbuf_release(&payload);\n+\tstrbuf_release(&othersig);\n+\tstrbuf_release(&oursig);\n \treturn 0;\n }\n \n-- \n2.41.0\n\n"},{"id":"482504","messageId":"20231002024034.2611-19-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 19/30] object-file-convert: convert commits that embed signed tags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:23Z","receivedAt":"2023-10-02T02:41:30Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nAs mentioned in the hash function transition plan commit mergetag\nlines need to be handled.  The commit mergetag lines embed an entire\ntag object in a commit object.\n\nKeep the implementation sane if not fast by unembedding the tag\nobject, converting the tag object, and embedding the new tag object,\nin the new commit object.\n\nIn the long run I don't expect any other approach is maintainable, as\ntag objects may be extended in ways that require additional\ntranslation.\n\nTo keep the implementation of convert_commit_object maintainable I\nhave modified convert_commit_object to process the lines in any order,\nand to fail on unknown lines.  We can't know ahead of time if a new\nline might embed something that needs translation or not so it is\nbetter to fail and require the code to be updated instead of silently\nmistranslating objects.\n\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 104 +++++++++++++++++++++++++++++++++---------\n 1 file changed, 82 insertions(+), 22 deletions(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 0da081104ed4..4f6189095be8 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -146,35 +146,95 @@ static int convert_commit_object(struct strbuf *out,\n \tconst int tree_entry_len = from->hexsz + 5;\n \tconst int parent_entry_len = from->hexsz + 7;\n \tstruct object_id oid, mapped_oid;\n-\tconst char *p;\n+\tconst char *p, *eol;\n \n \ttail += size;\n-\tif (tail <= bufptr + tree_entry_len + 1 || memcmp(bufptr, \"tree \", 5) ||\n-\t\t\tbufptr[tree_entry_len] != '\\n')\n-\t\treturn error(\"bogus commit object\");\n-\tif (parse_oid_hex_algop(bufptr + 5, &oid, &p, from) < 0)\n-\t\treturn error(\"bad tree pointer\");\n \n-\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n-\t\treturn error(\"unable to map tree %s in commit object\",\n-\t\t\t     oid_to_hex(&oid));\n-\tstrbuf_addf(out, \"tree %s\\n\", oid_to_hex(&mapped_oid));\n-\tbufptr = p + 1;\n+\twhile ((bufptr < tail) && (*bufptr != '\\n')) {\n+\t\teol = memchr(bufptr, '\\n', tail - bufptr);\n+\t\tif (!eol)\n+\t\t\treturn error(_(\"bad %s in commit\"), \"line\");\n+\n+\t\tif (((bufptr + 5) < eol) && !memcmp(bufptr, \"tree \", 5))\n+\t\t{\n+\t\t\tif (((bufptr + tree_entry_len) != eol) ||\n+\t\t\t    parse_oid_hex_algop(bufptr + 5, &oid, &p, from) ||\n+\t\t\t    (p != eol))\n+\t\t\t\treturn error(_(\"bad %s in commit\"), \"tree\");\n+\n+\t\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\t\t\treturn error(_(\"unable to map %s %s in commit object\"),\n+\t\t\t\t\t     \"tree\", oid_to_hex(&oid));\n+\t\t\tstrbuf_addf(out, \"tree %s\\n\", oid_to_hex(&mapped_oid));\n+\t\t}\n+\t\telse if (((bufptr + 7) < eol) && !memcmp(bufptr, \"parent \", 7))\n+\t\t{\n+\t\t\tif (((bufptr + parent_entry_len) != eol) ||\n+\t\t\t    parse_oid_hex_algop(bufptr + 7, &oid, &p, from) ||\n+\t\t\t    (p != eol))\n+\t\t\t\treturn error(_(\"bad %s in commit\"), \"parent\");\n \n-\twhile (bufptr + parent_entry_len < tail && !memcmp(bufptr, \"parent \", 7)) {\n-\t\tif (tail <= bufptr + parent_entry_len + 1 ||\n-\t\t    parse_oid_hex_algop(bufptr + 7, &oid, &p, from) ||\n-\t\t    *p != '\\n')\n-\t\t\treturn error(\"bad parents in commit\");\n+\t\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\t\t\treturn error(_(\"unable to map %s %s in commit object\"),\n+\t\t\t\t\t     \"parent\", oid_to_hex(&oid));\n \n-\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n-\t\t\treturn error(\"unable to map parent %s in commit object\",\n-\t\t\t\t     oid_to_hex(&oid));\n+\t\t\tstrbuf_addf(out, \"parent %s\\n\", oid_to_hex(&mapped_oid));\n+\t\t}\n+\t\telse if (((bufptr + 9) < eol) && !memcmp(bufptr, \"mergetag \", 9))\n+\t\t{\n+\t\t\tstruct strbuf tag = STRBUF_INIT, new_tag = STRBUF_INIT;\n \n-\t\tstrbuf_addf(out, \"parent %s\\n\", oid_to_hex(&mapped_oid));\n-\t\tbufptr = p + 1;\n+\t\t\t/* Recover the tag object from the mergetag */\n+\t\t\tstrbuf_add(&tag, bufptr + 9, (eol - (bufptr + 9)) + 1);\n+\n+\t\t\tbufptr = eol + 1;\n+\t\t\twhile ((bufptr < tail) && (*bufptr == ' ')) {\n+\t\t\t\teol = memchr(bufptr, '\\n', tail - bufptr);\n+\t\t\t\tif (!eol) {\n+\t\t\t\t\tstrbuf_release(&tag);\n+\t\t\t\t\treturn error(_(\"bad %s in commit\"), \"mergetag continuation\");\n+\t\t\t\t}\n+\t\t\t\tstrbuf_add(&tag, bufptr + 1, (eol - (bufptr + 1)) + 1);\n+\t\t\t\tbufptr = eol + 1;\n+\t\t\t}\n+\n+\t\t\t/* Compute the new tag object */\n+\t\t\tif (convert_tag_object(&new_tag, from, to, tag.buf, tag.len)) {\n+\t\t\t\tstrbuf_release(&tag);\n+\t\t\t\tstrbuf_release(&new_tag);\n+\t\t\t\treturn -1;\n+\t\t\t}\n+\n+\t\t\t/* Write the new mergetag */\n+\t\t\tstrbuf_addstr(out, \"mergetag\");\n+\t\t\tstrbuf_add_lines(out, \" \", new_tag.buf, new_tag.len);\n+\t\t\tstrbuf_release(&tag);\n+\t\t\tstrbuf_release(&new_tag);\n+\t\t}\n+\t\telse if (((bufptr + 7) < tail) && !memcmp(bufptr, \"author \", 7))\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\telse if (((bufptr + 10) < tail) && !memcmp(bufptr, \"committer \", 10))\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\telse if (((bufptr + 9) < tail) && !memcmp(bufptr, \"encoding \", 9))\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\telse if (((bufptr + 6) < tail) && !memcmp(bufptr, \"gpgsig\", 6))\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\telse {\n+\t\t\t/* Unknown line fail it might embed an oid */\n+\t\t\treturn -1;\n+\t\t}\n+\t\t/* Consume any trailing continuation lines */\n+\t\tbufptr = eol + 1;\n+\t\twhile ((bufptr < tail) && (*bufptr == ' ')) {\n+\t\t\teol = memchr(bufptr, '\\n', tail - bufptr);\n+\t\t\tif (!eol)\n+\t\t\t\treturn error(_(\"bad %s in commit\"), \"continuation\");\n+\t\t\tstrbuf_add(out, bufptr, (eol - bufptr) + 1);\n+\t\t\tbufptr = eol + 1;\n+\t\t}\n \t}\n-\tstrbuf_add(out, bufptr, tail - bufptr);\n+\tif (bufptr < tail)\n+\t\tstrbuf_add(out, bufptr, tail - bufptr);\n \treturn 0;\n }\n \n-- \n2.41.0\n\n"},{"id":"482505","messageId":"20231002024034.2611-18-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 18/30] object-file-convert: convert commit objects when writing","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:22Z","receivedAt":"2023-10-02T02:41:32Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nWhen writing a commit object in a repository with both SHA-1 and\nSHA-256, we'll need to convert our commit objects so that we can write\nthe hash values for both into the repository.  To do so, let's add a\nfunction to convert commit objects.\n\nRead the commit object and map the tree value and any of the parent\nvalues, and copy the rest of the commit through unmodified.  Note that\nwe don't need to modify the signature headers, because they are the\nsame under both algorithms.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 46 ++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 45 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 79e8e211ff95..0da081104ed4 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -136,6 +136,48 @@ static int convert_tag_object(struct strbuf *out,\n \treturn 0;\n }\n \n+static int convert_commit_object(struct strbuf *out,\n+\t\t\t\t const struct git_hash_algo *from,\n+\t\t\t\t const struct git_hash_algo *to,\n+\t\t\t\t const char *buffer, size_t size)\n+{\n+\tconst char *tail = buffer;\n+\tconst char *bufptr = buffer;\n+\tconst int tree_entry_len = from->hexsz + 5;\n+\tconst int parent_entry_len = from->hexsz + 7;\n+\tstruct object_id oid, mapped_oid;\n+\tconst char *p;\n+\n+\ttail += size;\n+\tif (tail <= bufptr + tree_entry_len + 1 || memcmp(bufptr, \"tree \", 5) ||\n+\t\t\tbufptr[tree_entry_len] != '\\n')\n+\t\treturn error(\"bogus commit object\");\n+\tif (parse_oid_hex_algop(bufptr + 5, &oid, &p, from) < 0)\n+\t\treturn error(\"bad tree pointer\");\n+\n+\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\treturn error(\"unable to map tree %s in commit object\",\n+\t\t\t     oid_to_hex(&oid));\n+\tstrbuf_addf(out, \"tree %s\\n\", oid_to_hex(&mapped_oid));\n+\tbufptr = p + 1;\n+\n+\twhile (bufptr + parent_entry_len < tail && !memcmp(bufptr, \"parent \", 7)) {\n+\t\tif (tail <= bufptr + parent_entry_len + 1 ||\n+\t\t    parse_oid_hex_algop(bufptr + 7, &oid, &p, from) ||\n+\t\t    *p != '\\n')\n+\t\t\treturn error(\"bad parents in commit\");\n+\n+\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\t\treturn error(\"unable to map parent %s in commit object\",\n+\t\t\t\t     oid_to_hex(&oid));\n+\n+\t\tstrbuf_addf(out, \"parent %s\\n\", oid_to_hex(&mapped_oid));\n+\t\tbufptr = p + 1;\n+\t}\n+\tstrbuf_add(out, bufptr, tail - bufptr);\n+\treturn 0;\n+}\n+\n int convert_object_file(struct strbuf *outbuf,\n \t\t\tconst struct git_hash_algo *from,\n \t\t\tconst struct git_hash_algo *to,\n@@ -150,13 +192,15 @@ int convert_object_file(struct strbuf *outbuf,\n \t\tBUG(\"Refusing noop object file conversion\");\n \n \tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\t\tret = convert_commit_object(outbuf, from, to, buf, len);\n+\t\tbreak;\n \tcase OBJ_TREE:\n \t\tret = convert_tree_object(outbuf, from, to, buf, len);\n \t\tbreak;\n \tcase OBJ_TAG:\n \t\tret = convert_tag_object(outbuf, from, to, buf, len);\n \t\tbreak;\n-\tcase OBJ_COMMIT:\n \tdefault:\n \t\t/* Not implemented yet, so fail. */\n \t\tret = -1;\n-- \n2.41.0\n\n"},{"id":"482506","messageId":"20231002024034.2611-20-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 20/30] object-file: update object_info_extended to reencode objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:24Z","receivedAt":"2023-10-02T02:41:33Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\noid_object_info_extended is updated to detect an oid encoding that\ndoes not match the current repository, use repo_oid_to_algop to find\nthe correspoding oid in the current repository and to return the data\nfor the oid.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n object-file.c | 91 +++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 91 insertions(+)\n\ndiff --git a/object-file.c b/object-file.c\nindex 820810a5f4b3..b2d43d009898 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1662,10 +1662,101 @@ static int do_oid_object_info_extended(struct repository *r,\n \treturn 0;\n }\n \n+static int oid_object_info_convert(struct repository *r,\n+\t\t\t\t   const struct object_id *input_oid,\n+\t\t\t\t   struct object_info *input_oi, unsigned flags)\n+{\n+\tconst struct git_hash_algo *input_algo = &hash_algos[input_oid->algo];\n+\tint do_die = flags & OBJECT_INFO_DIE_IF_CORRUPT;\n+\tstruct strbuf type_name = STRBUF_INIT;\n+\tstruct object_id oid, delta_base_oid;\n+\tstruct object_info new_oi, *oi;\n+\tunsigned long size;\n+\tvoid *content;\n+\tint ret;\n+\n+\tif (repo_oid_to_algop(r, input_oid, the_hash_algo, &oid)) {\n+\t\tif (do_die)\n+\t\t\tdie(_(\"missing mapping of %s to %s\"),\n+\t\t\t    oid_to_hex(input_oid), the_hash_algo->name);\n+\t\treturn -1;\n+\t}\n+\n+\t/* Is new_oi needed? */\n+\toi = input_oi;\n+\tif (input_oi && (input_oi->delta_base_oid || input_oi->sizep ||\n+\t\t\t input_oi->contentp)) {\n+\t\tnew_oi = *input_oi;\n+\t\t/* Does delta_base_oid need to be converted? */\n+\t\tif (input_oi->delta_base_oid)\n+\t\t\tnew_oi.delta_base_oid = &delta_base_oid;\n+\t\t/* Will the attributes differ when converted? */\n+\t\tif (input_oi->sizep || input_oi->contentp) {\n+\t\t\tnew_oi.contentp = &content;\n+\t\t\tnew_oi.sizep = &size;\n+\t\t\tnew_oi.type_name = &type_name;\n+\t\t}\n+\t\toi = &new_oi;\n+\t}\n+\n+\tret = oid_object_info_extended(r, &oid, oi, flags);\n+\tif (ret)\n+\t\treturn -1;\n+\tif (oi == input_oi)\n+\t\treturn ret;\n+\n+\tif (new_oi.contentp) {\n+\t\tstruct strbuf outbuf = STRBUF_INIT;\n+\t\tenum object_type type;\n+\n+\t\ttype = type_from_string_gently(type_name.buf, type_name.len,\n+\t\t\t\t\t       !do_die);\n+\t\tif (type == -1)\n+\t\t\treturn -1;\n+\t\tif (type != OBJ_BLOB) {\n+\t\t\tret = convert_object_file(&outbuf,\n+\t\t\t\t\t\t  the_hash_algo, input_algo,\n+\t\t\t\t\t\t  content, size, type, !do_die);\n+\t\t\tif (ret == -1)\n+\t\t\t\treturn -1;\n+\t\t\tfree(content);\n+\t\t\tsize = outbuf.len;\n+\t\t\tcontent = strbuf_detach(&outbuf, NULL);\n+\t\t}\n+\t\tif (input_oi->sizep)\n+\t\t\t*input_oi->sizep = size;\n+\t\tif (input_oi->contentp)\n+\t\t\t*input_oi->contentp = content;\n+\t\telse\n+\t\t\tfree(content);\n+\t\tif (input_oi->type_name)\n+\t\t\t*input_oi->type_name = type_name;\n+\t\telse\n+\t\t\tstrbuf_release(&type_name);\n+\t}\n+\tif (new_oi.delta_base_oid == &delta_base_oid) {\n+\t\tif (repo_oid_to_algop(r, &delta_base_oid, input_algo,\n+\t\t\t\t input_oi->delta_base_oid)) {\n+\t\t\tif (do_die)\n+\t\t\t\tdie(_(\"missing mapping of %s to %s\"),\n+\t\t\t\t    oid_to_hex(&delta_base_oid),\n+\t\t\t\t    input_algo->name);\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\tinput_oi->whence = new_oi.whence;\n+\tinput_oi->u = new_oi.u;\n+\treturn ret;\n+}\n+\n int oid_object_info_extended(struct repository *r, const struct object_id *oid,\n \t\t\t     struct object_info *oi, unsigned flags)\n {\n \tint ret;\n+\n+\tif (oid->algo && (hash_algo_by_ptr(r->hash_algo) != oid->algo))\n+\t\treturn oid_object_info_convert(r, oid, oi, flags);\n+\n \tobj_read_lock();\n \tret = do_oid_object_info_extended(r, oid, oi, flags);\n \tobj_read_unlock();\n-- \n2.41.0\n\n"},{"id":"482507","messageId":"20231002024034.2611-21-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 21/30] repository: implement extensions.compatObjectFormat","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:25Z","receivedAt":"2023-10-02T02:41:35Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nAdd a configuration option to enable updating and reading from\ncompatibility hash maps when git accesses the reposotiry.\n\nCall the helper function repo_set_compat_hash_algo with the value\nthat compatObjectFormat is set to.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Documentation/config/extensions.txt | 12 ++++++++++++\n repository.c                        |  2 +-\n setup.c                             | 23 +++++++++++++++++++++--\n setup.h                             |  1 +\n 4 files changed, 35 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/config/extensions.txt b/Documentation/config/extensions.txt\nindex bccaec7a9636..9f72e6d9f4f1 100644\n--- a/Documentation/config/extensions.txt\n+++ b/Documentation/config/extensions.txt\n@@ -7,6 +7,18 @@ Note that this setting should only be set by linkgit:git-init[1] or\n linkgit:git-clone[1].  Trying to change it after initialization will not\n work and will produce hard-to-diagnose issues.\n \n+extensions.compatObjectFormat::\n+\n+\tSpecify a compatitbility hash algorithm to use.  The acceptable values\n+\tare `sha1` and `sha256`.  The value specified must be different from the\n+\tvalue of extensions.objectFormat.  This allows client level\n+\tinteroperability between git repositories whose objectFormat matches\n+\tthis compatObjectFormat.  In particular when fully implemented the\n+\tpushes and pulls from a repository in whose objectFormat matches\n+\tcompatObjectFormat.  As well as being able to use oids encoded in\n+\tcompatObjectFormat in addition to oids encoded with objectFormat to\n+\tlocally specify objects.\n+\n extensions.worktreeConfig::\n \tIf enabled, then worktrees will load config settings from the\n \t`$GIT_DIR/config.worktree` file in addition to the\ndiff --git a/repository.c b/repository.c\nindex 6214f61cf4e7..9d91536b613b 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -194,7 +194,7 @@ int repo_init(struct repository *repo,\n \t\tgoto error;\n \n \trepo_set_hash_algo(repo, format.hash_algo);\n-\trepo_set_compat_hash_algo(repo, GIT_HASH_UNKNOWN);\n+\trepo_set_compat_hash_algo(repo, format.compat_hash_algo);\n \trepo->repository_format_worktree_config = format.worktree_config;\n \n \t/* take ownership of format.partial_clone */\ndiff --git a/setup.c b/setup.c\nindex aa8bf5da5226..85259a259be3 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -590,6 +590,25 @@ static enum extension_result handle_extension(const char *var,\n \t\t\t\t     \"extensions.objectformat\", value);\n \t\tdata->hash_algo = format;\n \t\treturn EXTENSION_OK;\n+\t} else if (!strcmp(ext, \"compatobjectformat\")) {\n+\t\tstruct string_list_item *item;\n+\t\tint format;\n+\n+\t\tif (!value)\n+\t\t\treturn config_error_nonbool(var);\n+\t\tformat = hash_algo_by_name(value);\n+\t\tif (format == GIT_HASH_UNKNOWN)\n+\t\t\treturn error(_(\"invalid value for '%s': '%s'\"),\n+\t\t\t\t     \"extensions.compatobjectformat\", value);\n+\t\t/* For now only support compatObjectFormat being specified once. */\n+\t\tfor_each_string_list_item(item, &data->v1_only_extensions) {\n+\t\t\tif (!strcmp(item->string, \"compatobjectformat\"))\n+\t\t\t\treturn error(_(\"'%s' already specified as '%s'\"),\n+\t\t\t\t\t\"extensions.compatobjectformat\",\n+\t\t\t\t\thash_algos[data->compat_hash_algo].name);\n+\t\t}\n+\t\tdata->compat_hash_algo = format;\n+\t\treturn EXTENSION_OK;\n \t}\n \treturn EXTENSION_UNKNOWN;\n }\n@@ -1565,7 +1584,7 @@ const char *setup_git_directory_gently(int *nongit_ok)\n \t\tif (startup_info->have_repository) {\n \t\t\trepo_set_hash_algo(the_repository, repo_fmt.hash_algo);\n \t\t\trepo_set_compat_hash_algo(the_repository,\n-\t\t\t\t\t\t  GIT_HASH_UNKNOWN);\n+\t\t\t\t\t\t  repo_fmt.compat_hash_algo);\n \t\t\tthe_repository->repository_format_worktree_config =\n \t\t\t\trepo_fmt.worktree_config;\n \t\t\t/* take ownership of repo_fmt.partial_clone */\n@@ -1659,7 +1678,7 @@ void check_repository_format(struct repository_format *fmt)\n \tcheck_repository_format_gently(get_git_dir(), fmt, NULL);\n \tstartup_info->have_repository = 1;\n \trepo_set_hash_algo(the_repository, fmt->hash_algo);\n-\trepo_set_compat_hash_algo(the_repository, GIT_HASH_UNKNOWN);\n+\trepo_set_compat_hash_algo(the_repository, fmt->compat_hash_algo);\n \tthe_repository->repository_format_worktree_config =\n \t\tfmt->worktree_config;\n \tthe_repository->repository_format_partial_clone =\ndiff --git a/setup.h b/setup.h\nindex 58fd2605dd26..5d678ceb8caa 100644\n--- a/setup.h\n+++ b/setup.h\n@@ -86,6 +86,7 @@ struct repository_format {\n \tint worktree_config;\n \tint is_bare;\n \tint hash_algo;\n+\tint compat_hash_algo;\n \tint sparse_index;\n \tchar *work_tree;\n \tstruct string_list unknown_extensions;\n-- \n2.41.0\n\n"},{"id":"482508","messageId":"20231002024034.2611-22-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 22/30] rev-parse: add an --output-object-format parameter","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:26Z","receivedAt":"2023-10-02T02:41:36Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nThe new --output-object-format parameter returns the oid in the\nspecified format.\n\nThis is a generally useful plumbing facility.  It is useful for writing\ntest cases and for directly querying the translation maps.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Documentation/git-rev-parse.txt | 12 ++++++++++++\n builtin/rev-parse.c             | 23 +++++++++++++++++++++++\n 2 files changed, 35 insertions(+)\n\ndiff --git a/Documentation/git-rev-parse.txt b/Documentation/git-rev-parse.txt\nindex f26a7591e373..f0f9021f2a5a 100644\n--- a/Documentation/git-rev-parse.txt\n+++ b/Documentation/git-rev-parse.txt\n@@ -159,6 +159,18 @@ for another option.\n \tunfortunately named tag \"master\"), and show them as full\n \trefnames (e.g. \"refs/heads/master\").\n \n+--output-object-format=(sha1|sha256|storage)::\n+\n+\tAllow oids to be input from any object format that the current\n+\trepository supports.\n+\n+\tSpecifying \"sha1\" translates if necessary and returns a sha1 oid.\n+\n+\tSpecifying \"sha256\" translates if necessary and returns a sha256 oid.\n+\n+\tSpecifying \"storage\" translates if necessary and returns an oid in\n+\tencoded in the storage hash algorithm.\n+\n Options for Objects\n ~~~~~~~~~~~~~~~~~~~\n \ndiff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\nindex 43e96765400c..0ef3e658cc5b 100644\n--- a/builtin/rev-parse.c\n+++ b/builtin/rev-parse.c\n@@ -25,6 +25,7 @@\n #include \"submodule.h\"\n #include \"commit-reach.h\"\n #include \"shallow.h\"\n+#include \"object-file-convert.h\"\n \n #define DO_REVS\t\t1\n #define DO_NOREV\t2\n@@ -675,6 +676,8 @@ static void print_path(const char *path, const char *prefix, enum format_type fo\n int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n {\n \tint i, as_is = 0, verify = 0, quiet = 0, revs_count = 0, type = 0;\n+\tconst struct git_hash_algo *output_algo = NULL;\n+\tconst struct git_hash_algo *compat = NULL;\n \tint did_repo_setup = 0;\n \tint has_dashdash = 0;\n \tint output_prefix = 0;\n@@ -746,6 +749,7 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \n \t\t\tprepare_repo_settings(the_repository);\n \t\t\tthe_repository->settings.command_requires_full_index = 0;\n+\t\t\tcompat = the_repository->compat_hash_algo;\n \t\t}\n \n \t\tif (!strcmp(arg, \"--\")) {\n@@ -833,6 +837,22 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t\t\tflags |= GET_OID_QUIETLY;\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (opt_with_value(arg, \"--output-object-format\", &arg)) {\n+\t\t\t\tif (!arg)\n+\t\t\t\t\tdie(_(\"no object format specified\"));\n+\t\t\t\tif (!strcmp(arg, the_hash_algo->name) ||\n+\t\t\t\t    !strcmp(arg, \"storage\")) {\n+\t\t\t\t\tflags |= GET_OID_HASH_ANY;\n+\t\t\t\t\toutput_algo = the_hash_algo;\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t\telse if (compat && !strcmp(arg, compat->name)) {\n+\t\t\t\t\tflags |= GET_OID_HASH_ANY;\n+\t\t\t\t\toutput_algo = compat;\n+\t\t\t\t\tcontinue;\n+\t\t\t\t}\n+\t\t\t\telse die(_(\"unsupported object format: %s\"), arg);\n+\t\t\t}\n \t\t\tif (opt_with_value(arg, \"--short\", &arg)) {\n \t\t\t\tfilter &= ~(DO_FLAGS|DO_NOREV);\n \t\t\t\tverify = 1;\n@@ -1083,6 +1103,9 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n \t\t}\n \t\tif (!get_oid_with_context(the_repository, name,\n \t\t\t\t\t  flags, &oid, &unused)) {\n+\t\t\tif (output_algo)\n+\t\t\t\trepo_oid_to_algop(the_repository, &oid,\n+\t\t\t\t\t\t  output_algo, &oid);\n \t\t\tif (verify)\n \t\t\t\trevs_count++;\n \t\t\telse\n-- \n2.41.0\n\n"},{"id":"482509","messageId":"20231002024034.2611-23-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 23/30] builtin/cat-file: let the oid determine the output algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:27Z","receivedAt":"2023-10-02T02:41:38Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUse GET_OID_HASH_ANY when calling get_oid_with_context.  This\nimplements the semi-obvious behaviour that specifying a sha1 oid shows\nthe output for a sha1 encoded object, and specifying a sha256 oid\nshows the output for a sha256 encoded object.\n\nThis is useful for testing the the conversion of an object to an\nequivalent object encoded with a different hash function.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/cat-file.c | 12 +++++++++---\n 1 file changed, 9 insertions(+), 3 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 694c8538df2f..e615d1f8e0da 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -107,7 +107,10 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct strbuf sb = STRBUF_INIT;\n \tunsigned flags = OBJECT_INFO_LOOKUP_REPLACE;\n-\tunsigned get_oid_flags = GET_OID_RECORD_PATH | GET_OID_ONLY_TO_DIE;\n+\tunsigned get_oid_flags =\n+\t\tGET_OID_RECORD_PATH |\n+\t\tGET_OID_ONLY_TO_DIE |\n+\t\tGET_OID_HASH_ANY;\n \tconst char *path = force_path;\n \tconst int opt_cw = (opt == 'c' || opt == 'w');\n \tif (!path && opt_cw)\n@@ -223,7 +226,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \t\t\t\t\t\t\t\t     &size);\n \t\t\t\tconst char *target;\n \t\t\t\tif (!skip_prefix(buffer, \"object \", &target) ||\n-\t\t\t\t    get_oid_hex(target, &blob_oid))\n+\t\t\t\t    get_oid_hex_algop(target, &blob_oid,\n+\t\t\t\t\t\t      &hash_algos[oid.algo]))\n \t\t\t\t\tdie(\"%s not a valid tag\", oid_to_hex(&oid));\n \t\t\t\tfree(buffer);\n \t\t\t} else\n@@ -512,7 +516,9 @@ static void batch_one_object(const char *obj_name,\n \t\t\t     struct expand_data *data)\n {\n \tstruct object_context ctx;\n-\tint flags = opt->follow_symlinks ? GET_OID_FOLLOW_SYMLINKS : 0;\n+\tint flags =\n+\t\tGET_OID_HASH_ANY |\n+\t\t(opt->follow_symlinks ? GET_OID_FOLLOW_SYMLINKS : 0);\n \tenum get_oid_result result;\n \n \tresult = get_oid_with_context(the_repository, obj_name,\n-- \n2.41.0\n\n"},{"id":"482510","messageId":"20231002024034.2611-25-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 25/30] object-file: handle compat objects in check_object_signature","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:29Z","receivedAt":"2023-10-02T02:41:49Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpdate check_object_signature to find the hash algorithm the exising\nsignature uses, and to use the same hash algorithm when recomputing it\nto check the signature is valid.\n\nThis will be useful when teaching git ls-tree to display objects\nencoded with the compat hash algorithm.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n object-file.c | 4 +++-\n 1 file changed, 3 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex b2d43d009898..5fa4b14baee0 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1094,9 +1094,11 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n \t\t\t   void *buf, unsigned long size,\n \t\t\t   enum object_type type)\n {\n+\tconst struct git_hash_algo *algo =\n+\t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n \tstruct object_id real_oid;\n \n-\thash_object_file(r->hash_algo, buf, size, type, &real_oid);\n+\thash_object_file(algo, buf, size, type, &real_oid);\n \n \treturn !oideq(oid, &real_oid) ? -1 : 0;\n }\n-- \n2.41.0\n\n"},{"id":"482511","messageId":"20231002024034.2611-26-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 26/30] builtin/ls-tree: let the oid determine the output algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:30Z","receivedAt":"2023-10-02T02:41:52Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpdate cmd_ls_tree to call get_oid_with_context and pass\nGET_OID_HASH_ANY instead of calling the simpler repo_get_oid.\n\nThis implments in ls-tree the behavior that asking to display a sha1\nhash displays the corrresponding sha1 encoded object and asking to\ndisplay a sha256 hash displayes the corresponding sha256 encoded\nobject.\n\nThis is useful for testing the conversion of an object to an\nequivlanet object encoded with a different hash function.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/ls-tree.c | 5 ++++-\n 1 file changed, 4 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/ls-tree.c b/builtin/ls-tree.c\nindex f558db5f3b80..71281ab705b6 100644\n--- a/builtin/ls-tree.c\n+++ b/builtin/ls-tree.c\n@@ -376,6 +376,7 @@ int cmd_ls_tree(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \tstruct ls_tree_cmdmode_to_fmt *m2f = ls_tree_cmdmode_format;\n+\tstruct object_context obj_context;\n \tint ret;\n \n \tgit_config(git_default_config, NULL);\n@@ -407,7 +408,9 @@ int cmd_ls_tree(int argc, const char **argv, const char *prefix)\n \t\t\tls_tree_usage, ls_tree_options);\n \tif (argc < 1)\n \t\tusage_with_options(ls_tree_usage, ls_tree_options);\n-\tif (repo_get_oid(the_repository, argv[0], &oid))\n+\tif (get_oid_with_context(the_repository, argv[0],\n+\t\t\t\t GET_OID_HASH_ANY, &oid,\n+\t\t\t\t &obj_context))\n \t\tdie(\"Not a valid object name %s\", argv[0]);\n \n \t/*\n-- \n2.41.0\n\n"},{"id":"482512","messageId":"20231002024034.2611-24-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 24/30] tree-walk: init_tree_desc take an oid to get the hash algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:28Z","receivedAt":"2023-10-02T02:41:53Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nTo make it possible for git ls-tree to display the tree encoded\nin the hash algorithm of the oid specified to git ls-tree, update\ninit_tree_desc to take as a parameter the oid of the tree object.\n\nUpdate all callers of init_tree_desc and init_tree_desc_gently\nto pass the oid of the tree object.\n\nUse the oid of the tree object to discover the hash algorithm\nof the oid and store that hash algorithm in struct tree_desc.\n\nUse the hash algorithm in decode_tree_entry and\nupdate_tree_entry_internal to handle reading a tree object encoded in\na hash algorithm that differs from the repositories hash algorithm.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n archive.c              |  3 ++-\n builtin/am.c           |  6 +++---\n builtin/checkout.c     |  8 +++++---\n builtin/clone.c        |  2 +-\n builtin/commit.c       |  2 +-\n builtin/grep.c         |  8 ++++----\n builtin/merge.c        |  3 ++-\n builtin/pack-objects.c |  6 ++++--\n builtin/read-tree.c    |  2 +-\n builtin/stash.c        |  5 +++--\n cache-tree.c           |  2 +-\n delta-islands.c        |  2 +-\n diff-lib.c             |  2 +-\n fsck.c                 |  6 ++++--\n http-push.c            |  2 +-\n list-objects.c         |  2 +-\n match-trees.c          |  4 ++--\n merge-ort.c            | 11 ++++++-----\n merge-recursive.c      |  2 +-\n merge.c                |  3 ++-\n pack-bitmap-write.c    |  2 +-\n packfile.c             |  3 ++-\n reflog.c               |  2 +-\n revision.c             |  4 ++--\n tree-walk.c            | 36 +++++++++++++++++++++---------------\n tree-walk.h            |  7 +++++--\n tree.c                 |  2 +-\n walker.c               |  2 +-\n 28 files changed, 80 insertions(+), 59 deletions(-)\n\ndiff --git a/archive.c b/archive.c\nindex ca11db185b15..b10269aee7be 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -339,7 +339,8 @@ int write_archive_entries(struct archiver_args *args,\n \t\topts.src_index = args->repo->index;\n \t\topts.dst_index = args->repo->index;\n \t\topts.fn = oneway_merge;\n-\t\tinit_tree_desc(&t, args->tree->buffer, args->tree->size);\n+\t\tinit_tree_desc(&t, &args->tree->object.oid,\n+\t\t\t       args->tree->buffer, args->tree->size);\n \t\tif (unpack_trees(1, &t, &opts))\n \t\t\treturn -1;\n \t\tgit_attr_set_direction(GIT_ATTR_INDEX);\ndiff --git a/builtin/am.c b/builtin/am.c\nindex 8bde034fae68..4dfd714b910e 100644\n--- a/builtin/am.c\n+++ b/builtin/am.c\n@@ -1991,8 +1991,8 @@ static int fast_forward_to(struct tree *head, struct tree *remote, int reset)\n \topts.reset = reset ? UNPACK_RESET_PROTECT_UNTRACKED : 0;\n \topts.preserve_ignored = 0; /* FIXME: !overwrite_ignore */\n \topts.fn = twoway_merge;\n-\tinit_tree_desc(&t[0], head->buffer, head->size);\n-\tinit_tree_desc(&t[1], remote->buffer, remote->size);\n+\tinit_tree_desc(&t[0], &head->object.oid, head->buffer, head->size);\n+\tinit_tree_desc(&t[1], &remote->object.oid, remote->buffer, remote->size);\n \n \tif (unpack_trees(2, t, &opts)) {\n \t\trollback_lock_file(&lock_file);\n@@ -2026,7 +2026,7 @@ static int merge_tree(struct tree *tree)\n \topts.dst_index = &the_index;\n \topts.merge = 1;\n \topts.fn = oneway_merge;\n-\tinit_tree_desc(&t[0], tree->buffer, tree->size);\n+\tinit_tree_desc(&t[0], &tree->object.oid, tree->buffer, tree->size);\n \n \tif (unpack_trees(1, t, &opts)) {\n \t\trollback_lock_file(&lock_file);\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex f53612f46870..03eff73fd031 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -701,7 +701,7 @@ static int reset_tree(struct tree *tree, const struct checkout_opts *o,\n \t\t\t       info->commit ? &info->commit->object.oid : null_oid(),\n \t\t\t       NULL);\n \tparse_tree(tree);\n-\tinit_tree_desc(&tree_desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&tree_desc, &tree->object.oid, tree->buffer, tree->size);\n \tswitch (unpack_trees(1, &tree_desc, &opts)) {\n \tcase -2:\n \t\t*writeout_error = 1;\n@@ -815,10 +815,12 @@ static int merge_working_tree(const struct checkout_opts *opts,\n \t\t\tdie(_(\"unable to parse commit %s\"),\n \t\t\t\toid_to_hex(old_commit_oid));\n \n-\t\tinit_tree_desc(&trees[0], tree->buffer, tree->size);\n+\t\tinit_tree_desc(&trees[0], &tree->object.oid,\n+\t\t\t       tree->buffer, tree->size);\n \t\tparse_tree(new_tree);\n \t\ttree = new_tree;\n-\t\tinit_tree_desc(&trees[1], tree->buffer, tree->size);\n+\t\tinit_tree_desc(&trees[1], &tree->object.oid,\n+\t\t\t       tree->buffer, tree->size);\n \n \t\tret = unpack_trees(2, trees, &topts);\n \t\tclear_unpack_trees_porcelain(&topts);\ndiff --git a/builtin/clone.c b/builtin/clone.c\nindex c6357af94989..79ceefb93995 100644\n--- a/builtin/clone.c\n+++ b/builtin/clone.c\n@@ -737,7 +737,7 @@ static int checkout(int submodule_progress, int filter_submodules)\n \tif (!tree)\n \t\tdie(_(\"unable to parse commit %s\"), oid_to_hex(&oid));\n \tparse_tree(tree);\n-\tinit_tree_desc(&t, tree->buffer, tree->size);\n+\tinit_tree_desc(&t, &tree->object.oid, tree->buffer, tree->size);\n \tif (unpack_trees(1, &t, &opts) < 0)\n \t\tdie(_(\"unable to checkout working tree\"));\n \ndiff --git a/builtin/commit.c b/builtin/commit.c\nindex 7da5f924484d..537319932b65 100644\n--- a/builtin/commit.c\n+++ b/builtin/commit.c\n@@ -340,7 +340,7 @@ static void create_base_index(const struct commit *current_head)\n \tif (!tree)\n \t\tdie(_(\"failed to unpack HEAD tree object\"));\n \tparse_tree(tree);\n-\tinit_tree_desc(&t, tree->buffer, tree->size);\n+\tinit_tree_desc(&t, &tree->object.oid, tree->buffer, tree->size);\n \tif (unpack_trees(1, &t, &opts))\n \t\texit(128); /* We've already reported the error, finish dying */\n }\ndiff --git a/builtin/grep.c b/builtin/grep.c\nindex 50e712a18479..0c2b8a376f8e 100644\n--- a/builtin/grep.c\n+++ b/builtin/grep.c\n@@ -530,7 +530,7 @@ static int grep_submodule(struct grep_opt *opt,\n \t\tstrbuf_addstr(&base, filename);\n \t\tstrbuf_addch(&base, '/');\n \n-\t\tinit_tree_desc(&tree, data, size);\n+\t\tinit_tree_desc(&tree, oid, data, size);\n \t\thit = grep_tree(&subopt, pathspec, &tree, &base, base.len,\n \t\t\t\tobject_type == OBJ_COMMIT);\n \t\tstrbuf_release(&base);\n@@ -574,7 +574,7 @@ static int grep_cache(struct grep_opt *opt,\n \n \t\t\tdata = repo_read_object_file(the_repository, &ce->oid,\n \t\t\t\t\t\t     &type, &size);\n-\t\t\tinit_tree_desc(&tree, data, size);\n+\t\t\tinit_tree_desc(&tree, &ce->oid, data, size);\n \n \t\t\thit |= grep_tree(opt, pathspec, &tree, &name, 0, 0);\n \t\t\tstrbuf_setlen(&name, name_base_len);\n@@ -670,7 +670,7 @@ static int grep_tree(struct grep_opt *opt, const struct pathspec *pathspec,\n \t\t\t\t    oid_to_hex(&entry.oid));\n \n \t\t\tstrbuf_addch(base, '/');\n-\t\t\tinit_tree_desc(&sub, data, size);\n+\t\t\tinit_tree_desc(&sub, &entry.oid, data, size);\n \t\t\thit |= grep_tree(opt, pathspec, &sub, base, tn_len,\n \t\t\t\t\t check_attr);\n \t\t\tfree(data);\n@@ -714,7 +714,7 @@ static int grep_object(struct grep_opt *opt, const struct pathspec *pathspec,\n \t\t\tstrbuf_add(&base, name, len);\n \t\t\tstrbuf_addch(&base, ':');\n \t\t}\n-\t\tinit_tree_desc(&tree, data, size);\n+\t\tinit_tree_desc(&tree, &obj->oid, data, size);\n \t\thit = grep_tree(opt, pathspec, &tree, &base, base.len,\n \t\t\t\tobj->type == OBJ_COMMIT);\n \t\tstrbuf_release(&base);\ndiff --git a/builtin/merge.c b/builtin/merge.c\nindex de68910177fb..718165d45917 100644\n--- a/builtin/merge.c\n+++ b/builtin/merge.c\n@@ -704,7 +704,8 @@ static int read_tree_trivial(struct object_id *common, struct object_id *head,\n \tcache_tree_free(&the_index.cache_tree);\n \tfor (i = 0; i < nr_trees; i++) {\n \t\tparse_tree(trees[i]);\n-\t\tinit_tree_desc(t+i, trees[i]->buffer, trees[i]->size);\n+\t\tinit_tree_desc(t+i, &trees[i]->object.oid,\n+\t\t\t       trees[i]->buffer, trees[i]->size);\n \t}\n \tif (unpack_trees(nr_trees, t, &opts))\n \t\treturn -1;\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex d2a162d52804..d34902002656 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1756,7 +1756,8 @@ static void add_pbase_object(struct tree_desc *tree,\n \t\t\ttree = pbase_tree_get(&entry.oid);\n \t\t\tif (!tree)\n \t\t\t\treturn;\n-\t\t\tinit_tree_desc(&sub, tree->tree_data, tree->tree_size);\n+\t\t\tinit_tree_desc(&sub, &tree->oid,\n+\t\t\t\t       tree->tree_data, tree->tree_size);\n \n \t\t\tadd_pbase_object(&sub, down, downlen, fullname);\n \t\t\tpbase_tree_put(tree);\n@@ -1816,7 +1817,8 @@ static void add_preferred_base_object(const char *name)\n \t\t}\n \t\telse {\n \t\t\tstruct tree_desc tree;\n-\t\t\tinit_tree_desc(&tree, it->pcache.tree_data, it->pcache.tree_size);\n+\t\t\tinit_tree_desc(&tree, &it->pcache.oid,\n+\t\t\t\t       it->pcache.tree_data, it->pcache.tree_size);\n \t\t\tadd_pbase_object(&tree, name, cmplen, name);\n \t\t}\n \t}\ndiff --git a/builtin/read-tree.c b/builtin/read-tree.c\nindex 1fec702a04fa..24d6d156d3a2 100644\n--- a/builtin/read-tree.c\n+++ b/builtin/read-tree.c\n@@ -264,7 +264,7 @@ int cmd_read_tree(int argc, const char **argv, const char *cmd_prefix)\n \tfor (i = 0; i < nr_trees; i++) {\n \t\tstruct tree *tree = trees[i];\n \t\tparse_tree(tree);\n-\t\tinit_tree_desc(t+i, tree->buffer, tree->size);\n+\t\tinit_tree_desc(t+i, &tree->object.oid, tree->buffer, tree->size);\n \t}\n \tif (unpack_trees(nr_trees, t, &opts))\n \t\treturn 128;\ndiff --git a/builtin/stash.c b/builtin/stash.c\nindex fe64cde9ce30..9ee52af4d28e 100644\n--- a/builtin/stash.c\n+++ b/builtin/stash.c\n@@ -285,7 +285,7 @@ static int reset_tree(struct object_id *i_tree, int update, int reset)\n \tif (parse_tree(tree))\n \t\treturn -1;\n \n-\tinit_tree_desc(t, tree->buffer, tree->size);\n+\tinit_tree_desc(t, &tree->object.oid, tree->buffer, tree->size);\n \n \topts.head_idx = 1;\n \topts.src_index = &the_index;\n@@ -871,7 +871,8 @@ static void diff_include_untracked(const struct stash_info *info, struct diff_op\n \t\ttree[i] = parse_tree_indirect(oid[i]);\n \t\tif (parse_tree(tree[i]) < 0)\n \t\t\tdie(_(\"failed to parse tree\"));\n-\t\tinit_tree_desc(&tree_desc[i], tree[i]->buffer, tree[i]->size);\n+\t\tinit_tree_desc(&tree_desc[i], &tree[i]->object.oid,\n+\t\t\t       tree[i]->buffer, tree[i]->size);\n \t}\n \n \tunpack_tree_opt.head_idx = -1;\ndiff --git a/cache-tree.c b/cache-tree.c\nindex ddc7d3d86959..334973a01cee 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -770,7 +770,7 @@ static void prime_cache_tree_rec(struct repository *r,\n \n \toidcpy(&it->oid, &tree->object.oid);\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \tcnt = 0;\n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (!S_ISDIR(entry.mode))\ndiff --git a/delta-islands.c b/delta-islands.c\nindex 5de5759f3f13..1ff3506b10f2 100644\n--- a/delta-islands.c\n+++ b/delta-islands.c\n@@ -289,7 +289,7 @@ void resolve_tree_islands(struct repository *r,\n \t\tif (!tree || parse_tree(tree) < 0)\n \t\t\tdie(_(\"bad tree object %s\"), oid_to_hex(&ent->idx.oid));\n \n-\t\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\t\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \t\twhile (tree_entry(&desc, &entry)) {\n \t\t\tstruct object *obj;\n \ndiff --git a/diff-lib.c b/diff-lib.c\nindex 6b0c6a7180cc..add323f5628d 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -558,7 +558,7 @@ static int diff_cache(struct rev_info *revs,\n \topts.pathspec = &revs->diffopt.pathspec;\n \topts.pathspec->recursive = 1;\n \n-\tinit_tree_desc(&t, tree->buffer, tree->size);\n+\tinit_tree_desc(&t, &tree->object.oid, tree->buffer, tree->size);\n \treturn unpack_trees(1, &t, &opts);\n }\n \ndiff --git a/fsck.c b/fsck.c\nindex 2b1e348005b7..6b492a48da82 100644\n--- a/fsck.c\n+++ b/fsck.c\n@@ -313,7 +313,8 @@ static int fsck_walk_tree(struct tree *tree, void *data, struct fsck_options *op\n \t\treturn -1;\n \n \tname = fsck_get_object_name(options, &tree->object.oid);\n-\tif (init_tree_desc_gently(&desc, tree->buffer, tree->size, 0))\n+\tif (init_tree_desc_gently(&desc, &tree->object.oid,\n+\t\t\t\t  tree->buffer, tree->size, 0))\n \t\treturn -1;\n \twhile (tree_entry_gently(&desc, &entry)) {\n \t\tstruct object *obj;\n@@ -583,7 +584,8 @@ static int fsck_tree(const struct object_id *tree_oid,\n \tconst char *o_name;\n \tstruct name_stack df_dup_candidates = { NULL };\n \n-\tif (init_tree_desc_gently(&desc, buffer, size, TREE_DESC_RAW_MODES)) {\n+\tif (init_tree_desc_gently(&desc, tree_oid, buffer, size,\n+\t\t\t\t  TREE_DESC_RAW_MODES)) {\n \t\tretval += report(options, tree_oid, OBJ_TREE,\n \t\t\t\t FSCK_MSG_BAD_TREE,\n \t\t\t\t \"cannot be parsed as a tree\");\ndiff --git a/http-push.c b/http-push.c\nindex a704f490fdb2..81c35b5e96f7 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -1308,7 +1308,7 @@ static struct object_list **process_tree(struct tree *tree,\n \tobj->flags |= SEEN;\n \tp = add_one_object(obj, p);\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry))\n \t\tswitch (object_type(entry.mode)) {\ndiff --git a/list-objects.c b/list-objects.c\nindex e60a6cd5b46e..312335c8a7f2 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -97,7 +97,7 @@ static void process_tree_contents(struct traversal_context *ctx,\n \tenum interesting match = ctx->revs->diffopt.pathspec.nr == 0 ?\n \t\tall_entries_interesting : entry_not_interesting;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (match != all_entries_interesting) {\ndiff --git a/match-trees.c b/match-trees.c\nindex 0885ac681cd5..3412b6a1401d 100644\n--- a/match-trees.c\n+++ b/match-trees.c\n@@ -63,7 +63,7 @@ static void *fill_tree_desc_strict(struct tree_desc *desc,\n \t\tdie(\"unable to read tree (%s)\", oid_to_hex(hash));\n \tif (type != OBJ_TREE)\n \t\tdie(\"%s is not a tree\", oid_to_hex(hash));\n-\tinit_tree_desc(desc, buffer, size);\n+\tinit_tree_desc(desc, hash, buffer, size);\n \treturn buffer;\n }\n \n@@ -194,7 +194,7 @@ static int splice_tree(const struct object_id *oid1, const char *prefix,\n \tbuf = repo_read_object_file(the_repository, oid1, &type, &sz);\n \tif (!buf)\n \t\tdie(\"cannot read tree %s\", oid_to_hex(oid1));\n-\tinit_tree_desc(&desc, buf, sz);\n+\tinit_tree_desc(&desc, oid1, buf, sz);\n \n \trewrite_here = NULL;\n \twhile (desc.size) {\ndiff --git a/merge-ort.c b/merge-ort.c\nindex 8631c997002d..3a5729c91e48 100644\n--- a/merge-ort.c\n+++ b/merge-ort.c\n@@ -1679,9 +1679,10 @@ static int collect_merge_info(struct merge_options *opt,\n \tparse_tree(merge_base);\n \tparse_tree(side1);\n \tparse_tree(side2);\n-\tinit_tree_desc(t + 0, merge_base->buffer, merge_base->size);\n-\tinit_tree_desc(t + 1, side1->buffer, side1->size);\n-\tinit_tree_desc(t + 2, side2->buffer, side2->size);\n+\tinit_tree_desc(t + 0, &merge_base->object.oid,\n+\t\t       merge_base->buffer, merge_base->size);\n+\tinit_tree_desc(t + 1, &side1->object.oid, side1->buffer, side1->size);\n+\tinit_tree_desc(t + 2, &side2->object.oid, side2->buffer, side2->size);\n \n \ttrace2_region_enter(\"merge\", \"traverse_trees\", opt->repo);\n \tret = traverse_trees(NULL, 3, t, &info);\n@@ -4400,9 +4401,9 @@ static int checkout(struct merge_options *opt,\n \tunpack_opts.fn = twoway_merge;\n \tunpack_opts.preserve_ignored = 0; /* FIXME: !opts->overwrite_ignore */\n \tparse_tree(prev);\n-\tinit_tree_desc(&trees[0], prev->buffer, prev->size);\n+\tinit_tree_desc(&trees[0], &prev->object.oid, prev->buffer, prev->size);\n \tparse_tree(next);\n-\tinit_tree_desc(&trees[1], next->buffer, next->size);\n+\tinit_tree_desc(&trees[1], &next->object.oid, next->buffer, next->size);\n \n \tret = unpack_trees(2, trees, &unpack_opts);\n \tclear_unpack_trees_porcelain(&unpack_opts);\ndiff --git a/merge-recursive.c b/merge-recursive.c\nindex 6a4081bb0f52..93df9eecdd95 100644\n--- a/merge-recursive.c\n+++ b/merge-recursive.c\n@@ -411,7 +411,7 @@ static inline int merge_detect_rename(struct merge_options *opt)\n static void init_tree_desc_from_tree(struct tree_desc *desc, struct tree *tree)\n {\n \tparse_tree(tree);\n-\tinit_tree_desc(desc, tree->buffer, tree->size);\n+\tinit_tree_desc(desc, &tree->object.oid, tree->buffer, tree->size);\n }\n \n static int unpack_trees_start(struct merge_options *opt,\ndiff --git a/merge.c b/merge.c\nindex b60925459c29..86179c34102d 100644\n--- a/merge.c\n+++ b/merge.c\n@@ -81,7 +81,8 @@ int checkout_fast_forward(struct repository *r,\n \t}\n \tfor (i = 0; i < nr_trees; i++) {\n \t\tparse_tree(trees[i]);\n-\t\tinit_tree_desc(t+i, trees[i]->buffer, trees[i]->size);\n+\t\tinit_tree_desc(t+i, &trees[i]->object.oid,\n+\t\t\t       trees[i]->buffer, trees[i]->size);\n \t}\n \n \tmemset(&opts, 0, sizeof(opts));\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex f6757c3cbf20..9211e08f0127 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -366,7 +366,7 @@ static int fill_bitmap_tree(struct bitmap *bitmap,\n \tif (parse_tree(tree) < 0)\n \t\tdie(\"unable to load tree object %s\",\n \t\t    oid_to_hex(&tree->object.oid));\n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry)) {\n \t\tswitch (object_type(entry.mode)) {\ndiff --git a/packfile.c b/packfile.c\nindex 9cc0a2e37a83..1fae0fcdd9e7 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2250,7 +2250,8 @@ static int add_promisor_object(const struct object_id *oid,\n \t\tstruct tree *tree = (struct tree *)obj;\n \t\tstruct tree_desc desc;\n \t\tstruct name_entry entry;\n-\t\tif (init_tree_desc_gently(&desc, tree->buffer, tree->size, 0))\n+\t\tif (init_tree_desc_gently(&desc, &tree->object.oid,\n+\t\t\t\t\t  tree->buffer, tree->size, 0))\n \t\t\t/*\n \t\t\t * Error messages are given when packs are\n \t\t\t * verified, so do not print any here.\ndiff --git a/reflog.c b/reflog.c\nindex 9ad50e7d93e4..c6992a19268f 100644\n--- a/reflog.c\n+++ b/reflog.c\n@@ -40,7 +40,7 @@ static int tree_is_complete(const struct object_id *oid)\n \t\ttree->buffer = data;\n \t\ttree->size = size;\n \t}\n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \tcomplete = 1;\n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (!repo_has_object_file(the_repository, &entry.oid) ||\ndiff --git a/revision.c b/revision.c\nindex 2f4c53ea207b..a60dfc23a2a5 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -82,7 +82,7 @@ static void mark_tree_contents_uninteresting(struct repository *r,\n \tif (parse_tree_gently(tree, 1) < 0)\n \t\treturn;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \twhile (tree_entry(&desc, &entry)) {\n \t\tswitch (object_type(entry.mode)) {\n \t\tcase OBJ_TREE:\n@@ -189,7 +189,7 @@ static void add_children_by_path(struct repository *r,\n \tif (parse_tree_gently(tree, 1) < 0)\n \t\treturn;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \twhile (tree_entry(&desc, &entry)) {\n \t\tswitch (object_type(entry.mode)) {\n \t\tcase OBJ_TREE:\ndiff --git a/tree-walk.c b/tree-walk.c\nindex 3af50a01c2c7..0b44ec7c75ff 100644\n--- a/tree-walk.c\n+++ b/tree-walk.c\n@@ -15,7 +15,7 @@ static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned l\n \tconst char *path;\n \tunsigned int len;\n \tuint16_t mode;\n-\tconst unsigned hashsz = the_hash_algo->rawsz;\n+\tconst unsigned hashsz = desc->algo->rawsz;\n \n \tif (size < hashsz + 3 || buf[size - (hashsz + 1)]) {\n \t\tstrbuf_addstr(err, _(\"too-short tree object\"));\n@@ -37,15 +37,19 @@ static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned l\n \tdesc->entry.path = path;\n \tdesc->entry.mode = (desc->flags & TREE_DESC_RAW_MODES) ? mode : canon_mode(mode);\n \tdesc->entry.pathlen = len - 1;\n-\toidread(&desc->entry.oid, (const unsigned char *)path + len);\n+\toidread_algop(&desc->entry.oid, (const unsigned char *)path + len,\n+\t\t      desc->algo);\n \n \treturn 0;\n }\n \n-static int init_tree_desc_internal(struct tree_desc *desc, const void *buffer,\n-\t\t\t\t   unsigned long size, struct strbuf *err,\n+static int init_tree_desc_internal(struct tree_desc *desc,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   const void *buffer, unsigned long size,\n+\t\t\t\t   struct strbuf *err,\n \t\t\t\t   enum tree_desc_flags flags)\n {\n+\tdesc->algo = (oid && oid->algo) ? &hash_algos[oid->algo] : the_hash_algo;\n \tdesc->buffer = buffer;\n \tdesc->size = size;\n \tdesc->flags = flags;\n@@ -54,19 +58,21 @@ static int init_tree_desc_internal(struct tree_desc *desc, const void *buffer,\n \treturn 0;\n }\n \n-void init_tree_desc(struct tree_desc *desc, const void *buffer, unsigned long size)\n+void init_tree_desc(struct tree_desc *desc, const struct object_id *tree_oid,\n+\t\t    const void *buffer, unsigned long size)\n {\n \tstruct strbuf err = STRBUF_INIT;\n-\tif (init_tree_desc_internal(desc, buffer, size, &err, 0))\n+\tif (init_tree_desc_internal(desc, tree_oid, buffer, size, &err, 0))\n \t\tdie(\"%s\", err.buf);\n \tstrbuf_release(&err);\n }\n \n-int init_tree_desc_gently(struct tree_desc *desc, const void *buffer, unsigned long size,\n+int init_tree_desc_gently(struct tree_desc *desc, const struct object_id *oid,\n+\t\t\t  const void *buffer, unsigned long size,\n \t\t\t  enum tree_desc_flags flags)\n {\n \tstruct strbuf err = STRBUF_INIT;\n-\tint result = init_tree_desc_internal(desc, buffer, size, &err, flags);\n+\tint result = init_tree_desc_internal(desc, oid, buffer, size, &err, flags);\n \tif (result)\n \t\terror(\"%s\", err.buf);\n \tstrbuf_release(&err);\n@@ -85,7 +91,7 @@ void *fill_tree_descriptor(struct repository *r,\n \t\tif (!buf)\n \t\t\tdie(\"unable to read tree %s\", oid_to_hex(oid));\n \t}\n-\tinit_tree_desc(desc, buf, size);\n+\tinit_tree_desc(desc, oid, buf, size);\n \treturn buf;\n }\n \n@@ -102,7 +108,7 @@ static void entry_extract(struct tree_desc *t, struct name_entry *a)\n static int update_tree_entry_internal(struct tree_desc *desc, struct strbuf *err)\n {\n \tconst void *buf = desc->buffer;\n-\tconst unsigned char *end = (const unsigned char *)desc->entry.path + desc->entry.pathlen + 1 + the_hash_algo->rawsz;\n+\tconst unsigned char *end = (const unsigned char *)desc->entry.path + desc->entry.pathlen + 1 + desc->algo->rawsz;\n \tunsigned long size = desc->size;\n \tunsigned long len = end - (const unsigned char *)buf;\n \n@@ -611,7 +617,7 @@ int get_tree_entry(struct repository *r,\n \t\tretval = -1;\n \t} else {\n \t\tstruct tree_desc t;\n-\t\tinit_tree_desc(&t, tree, size);\n+\t\tinit_tree_desc(&t, tree_oid, tree, size);\n \t\tretval = find_tree_entry(r, &t, name, oid, mode);\n \t}\n \tfree(tree);\n@@ -654,7 +660,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \tstruct tree_desc t;\n \tint follows_remaining = GET_TREE_ENTRY_FOLLOW_SYMLINKS_MAX_LINKS;\n \n-\tinit_tree_desc(&t, NULL, 0UL);\n+\tinit_tree_desc(&t, NULL, NULL, 0UL);\n \tstrbuf_addstr(&namebuf, name);\n \toidcpy(&current_tree_oid, tree_oid);\n \n@@ -690,7 +696,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \t\t\t\tgoto done;\n \n \t\t\t/* descend */\n-\t\t\tinit_tree_desc(&t, tree, size);\n+\t\t\tinit_tree_desc(&t, &current_tree_oid, tree, size);\n \t\t}\n \n \t\t/* Handle symlinks to e.g. a//b by removing leading slashes */\n@@ -724,7 +730,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \t\t\tfree(parent->tree);\n \t\t\tparents_nr--;\n \t\t\tparent = &parents[parents_nr - 1];\n-\t\t\tinit_tree_desc(&t, parent->tree, parent->size);\n+\t\t\tinit_tree_desc(&t, &parent->oid, parent->tree, parent->size);\n \t\t\tstrbuf_remove(&namebuf, 0, remainder ? 3 : 2);\n \t\t\tcontinue;\n \t\t}\n@@ -804,7 +810,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \t\t\tcontents_start = contents;\n \n \t\t\tparent = &parents[parents_nr - 1];\n-\t\t\tinit_tree_desc(&t, parent->tree, parent->size);\n+\t\t\tinit_tree_desc(&t, &parent->oid, parent->tree, parent->size);\n \t\t\tstrbuf_splice(&namebuf, 0, len,\n \t\t\t\t      contents_start, link_len);\n \t\t\tif (remainder)\ndiff --git a/tree-walk.h b/tree-walk.h\nindex 74cdceb3fed2..cf54d01019e9 100644\n--- a/tree-walk.h\n+++ b/tree-walk.h\n@@ -26,6 +26,7 @@ struct name_entry {\n  * A semi-opaque data structure used to maintain the current state of the walk.\n  */\n struct tree_desc {\n+\tconst struct git_hash_algo *algo;\n \t/*\n \t * pointer into the memory representation of the tree. It always\n \t * points at the current entry being visited.\n@@ -85,9 +86,11 @@ int update_tree_entry_gently(struct tree_desc *);\n  * size parameters are assumed to be the same as the buffer and size\n  * members of `struct tree`.\n  */\n-void init_tree_desc(struct tree_desc *desc, const void *buf, unsigned long size);\n+void init_tree_desc(struct tree_desc *desc, const struct object_id *tree_oid,\n+\t\t    const void *buf, unsigned long size);\n \n-int init_tree_desc_gently(struct tree_desc *desc, const void *buf, unsigned long size,\n+int init_tree_desc_gently(struct tree_desc *desc, const struct object_id *oid,\n+\t\t\t  const void *buf, unsigned long size,\n \t\t\t  enum tree_desc_flags flags);\n \n /*\ndiff --git a/tree.c b/tree.c\nindex c745462f968e..44bcf728f10a 100644\n--- a/tree.c\n+++ b/tree.c\n@@ -27,7 +27,7 @@ int read_tree_at(struct repository *r,\n \tif (parse_tree(tree))\n \t\treturn -1;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (retval != all_entries_interesting) {\ndiff --git a/walker.c b/walker.c\nindex 65002a7220ad..c0fd632d921c 100644\n--- a/walker.c\n+++ b/walker.c\n@@ -45,7 +45,7 @@ static int process_tree(struct walker *walker, struct tree *tree)\n \tif (parse_tree(tree))\n \t\treturn -1;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \twhile (tree_entry(&desc, &entry)) {\n \t\tstruct object *obj = NULL;\n \n-- \n2.41.0\n\n"},{"id":"482513","messageId":"20231002024034.2611-27-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 27/30] test-lib: compute the compatibility hash so tests may use it","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:31Z","receivedAt":"2023-10-02T02:41:54Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nInspired-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n t/test-lib-functions.sh | 17 ++++++++++++++++-\n 1 file changed, 16 insertions(+), 1 deletion(-)\n\ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 2f8868caa171..92b462e2e711 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -1599,7 +1599,16 @@ test_set_hash () {\n \n # Detect the hash algorithm in use.\n test_detect_hash () {\n-\ttest_hash_algo=\"${GIT_TEST_DEFAULT_HASH:-sha1}\"\n+\tcase \"$GIT_TEST_DEFAULT_HASH\" in\n+\t\"sha256\")\n+\t    test_hash_algo=sha256\n+\t    test_compat_hash_algo=sha1\n+\t    ;;\n+\t*)\n+\t    test_hash_algo=sha1\n+\t    test_compat_hash_algo=sha256\n+\t    ;;\n+\tesac\n }\n \n # Load common hash metadata and common placeholder object IDs for use with\n@@ -1651,6 +1660,12 @@ test_oid () {\n \tlocal algo=\"${test_hash_algo}\" &&\n \n \tcase \"$1\" in\n+\t--hash=storage)\n+\t\talgo=\"$test_hash_algo\" &&\n+\t\tshift;;\n+\t--hash=compat)\n+\t\talgo=\"$test_compat_hash_algo\" &&\n+\t\tshift;;\n \t--hash=*)\n \t\talgo=\"${1#--hash=}\" &&\n \t\tshift;;\n-- \n2.41.0\n\n"},{"id":"482514","messageId":"20231002024034.2611-29-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 29/30] t1006: test oid compatibility with cat-file","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:33Z","receivedAt":"2023-10-02T02:41:57Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nUpdate the existing tests that are oid based to test that cat-file\nworks correctly with the normal oid and the compat_oid.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n t/t1006-cat-file.sh | 251 +++++++++++++++++++++++++++-----------------\n 1 file changed, 154 insertions(+), 97 deletions(-)\n\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex 9b018b538950..23d3d37283bb 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -236,27 +236,38 @@ hello_size=$(strlen \"$hello_content\")\n hello_oid=$(echo_without_newline \"$hello_content\" | git hash-object --stdin)\n \n test_expect_success \"setup\" '\n+\tgit config core.repositoryformatversion 1 &&\n+\tgit config extensions.objectformat $test_hash_algo &&\n+\tgit config extensions.compatobjectformat $test_compat_hash_algo &&\n \techo_without_newline \"$hello_content\" > hello &&\n \tgit update-index --add hello\n '\n \n-run_tests 'blob' $hello_oid $hello_size \"$hello_content\" \"$hello_content\"\n+run_blob_tests () {\n+    oid=$1\n \n-test_expect_success '--batch-command --buffer with flush for blob info' '\n-\techo \"$hello_oid blob $hello_size\" >expect &&\n-\ttest_write_lines \"info $hello_oid\" \"flush\" |\n+    run_tests 'blob' $oid $hello_size \"$hello_content\" \"$hello_content\"\n+\n+    test_expect_success '--batch-command --buffer with flush for blob info' '\n+\techo \"$oid blob $hello_size\" >expect &&\n+\ttest_write_lines \"info $oid\" \"flush\" |\n \tGIT_TEST_CAT_FILE_NO_FLUSH_ON_EXIT=1 \\\n \tgit cat-file --batch-command --buffer >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success '--batch-command --buffer without flush for blob info' '\n+    test_expect_success '--batch-command --buffer without flush for blob info' '\n \ttouch output &&\n-\ttest_write_lines \"info $hello_oid\" |\n+\ttest_write_lines \"info $oid\" |\n \tGIT_TEST_CAT_FILE_NO_FLUSH_ON_EXIT=1 \\\n \tgit cat-file --batch-command --buffer >>output &&\n \ttest_must_be_empty output\n-'\n+    '\n+}\n+\n+hello_compat_oid=$(git rev-parse --output-object-format=$test_compat_hash_algo $hello_oid)\n+run_blob_tests $hello_oid\n+run_blob_tests $hello_compat_oid\n \n test_expect_success '--batch-check without %(rest) considers whole line' '\n \techo \"$hello_oid blob $hello_size\" >expect &&\n@@ -267,35 +278,58 @@ test_expect_success '--batch-check without %(rest) considers whole line' '\n '\n \n tree_oid=$(git write-tree)\n+tree_compat_oid=$(git rev-parse --output-object-format=$test_compat_hash_algo $tree_oid)\n tree_size=$(($(test_oid rawsz) + 13))\n+tree_compat_size=$(($(test_oid --hash=compat rawsz) + 13))\n tree_pretty_content=\"100644 blob $hello_oid\thello${LF}\"\n+tree_compat_pretty_content=\"100644 blob $hello_compat_oid\thello${LF}\"\n \n run_tests 'tree' $tree_oid $tree_size \"\" \"$tree_pretty_content\"\n+run_tests 'tree' $tree_compat_oid $tree_compat_size \"\" \"$tree_compat_pretty_content\"\n \n commit_message=\"Initial commit\"\n commit_oid=$(echo_without_newline \"$commit_message\" | git commit-tree $tree_oid)\n+commit_compat_oid=$(git rev-parse --output-object-format=$test_compat_hash_algo $commit_oid)\n commit_size=$(($(test_oid hexsz) + 137))\n+commit_compat_size=$(($(test_oid --hash=compat hexsz) + 137))\n commit_content=\"tree $tree_oid\n author $GIT_AUTHOR_NAME <$GIT_AUTHOR_EMAIL> $GIT_AUTHOR_DATE\n committer $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL> $GIT_COMMITTER_DATE\n \n $commit_message\"\n \n+commit_compat_content=\"tree $tree_compat_oid\n+author $GIT_AUTHOR_NAME <$GIT_AUTHOR_EMAIL> $GIT_AUTHOR_DATE\n+committer $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL> $GIT_COMMITTER_DATE\n+\n+$commit_message\"\n+\n run_tests 'commit' $commit_oid $commit_size \"$commit_content\" \"$commit_content\"\n+run_tests 'commit' $commit_compat_oid $commit_compat_size \"$commit_compat_content\" \"$commit_compat_content\"\n \n-tag_header_without_timestamp=\"object $hello_oid\n-type blob\n+tag_header_without_oid=\"type blob\n tag hellotag\n tagger $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL>\"\n+tag_header_without_timestamp=\"object $hello_oid\n+$tag_header_without_oid\"\n+tag_compat_header_without_timestamp=\"object $hello_compat_oid\n+$tag_header_without_oid\"\n tag_description=\"This is a tag\"\n tag_content=\"$tag_header_without_timestamp 0 +0000\n \n+$tag_description\"\n+tag_compat_content=\"$tag_compat_header_without_timestamp 0 +0000\n+\n $tag_description\"\n \n tag_oid=$(echo_without_newline \"$tag_content\" | git hash-object -t tag --stdin -w)\n tag_size=$(strlen \"$tag_content\")\n \n+tag_compat_oid=$(git rev-parse --output-object-format=$test_compat_hash_algo $tag_oid)\n+tag_compat_size=$(strlen \"$tag_compat_content\")\n+\n run_tests 'tag' $tag_oid $tag_size \"$tag_content\" \"$tag_content\"\n+run_tests 'tag' $tag_compat_oid $tag_compat_size \"$tag_compat_content\" \"$tag_compat_content\"\n \n test_expect_success \"Reach a blob from a tag pointing to it\" '\n \techo_without_newline \"$hello_content\" >expect &&\n@@ -303,37 +337,43 @@ test_expect_success \"Reach a blob from a tag pointing to it\" '\n \ttest_cmp expect actual\n '\n \n-for batch in batch batch-check batch-command\n+for oid in $hello_oid $hello_compat_oid\n do\n-    for opt in t s e p\n+    for batch in batch batch-check batch-command\n     do\n+\tfor opt in t s e p\n+\tdo\n \ttest_expect_success \"Passing -$opt with --$batch fails\" '\n-\t    test_must_fail git cat-file --$batch -$opt $hello_oid\n+\t    test_must_fail git cat-file --$batch -$opt $oid\n \t'\n \n \ttest_expect_success \"Passing --$batch with -$opt fails\" '\n-\t    test_must_fail git cat-file -$opt --$batch $hello_oid\n+\t    test_must_fail git cat-file -$opt --$batch $oid\n \t'\n-    done\n+\tdone\n \n-    test_expect_success \"Passing <type> with --$batch fails\" '\n-\ttest_must_fail git cat-file --$batch blob $hello_oid\n-    '\n+\ttest_expect_success \"Passing <type> with --$batch fails\" '\n+\ttest_must_fail git cat-file --$batch blob $oid\n+\t'\n \n-    test_expect_success \"Passing --$batch with <type> fails\" '\n-\ttest_must_fail git cat-file blob --$batch $hello_oid\n-    '\n+\ttest_expect_success \"Passing --$batch with <type> fails\" '\n+\ttest_must_fail git cat-file blob --$batch $oid\n+\t'\n \n-    test_expect_success \"Passing oid with --$batch fails\" '\n-\ttest_must_fail git cat-file --$batch $hello_oid\n-    '\n+\ttest_expect_success \"Passing oid with --$batch fails\" '\n+\ttest_must_fail git cat-file --$batch $oid\n+\t'\n+    done\n done\n \n-for opt in t s e p\n+for oid in $hello_oid $hello_compat_oid\n do\n-    test_expect_success \"Passing -$opt with --follow-symlinks fails\" '\n-\t    test_must_fail git cat-file --follow-symlinks -$opt $hello_oid\n+    for opt in t s e p\n+    do\n+\ttest_expect_success \"Passing -$opt with --follow-symlinks fails\" '\n+\t    test_must_fail git cat-file --follow-symlinks -$opt $oid\n \t'\n+    done\n done\n \n test_expect_success \"--batch-check for a non-existent named object\" '\n@@ -386,112 +426,102 @@ test_expect_success 'empty --batch-check notices missing object' '\n \ttest_cmp expect actual\n '\n \n-batch_input=\"$hello_oid\n-$commit_oid\n-$tag_oid\n+batch_tests () {\n+    boid=$1\n+    loid=$2\n+    lsize=$3\n+    coid=$4\n+    csize=$5\n+    ccontent=$6\n+    toid=$7\n+    tsize=$8\n+    tcontent=$9\n+\n+    batch_input=\"$boid\n+$coid\n+$toid\n deadbeef\n \n \"\n \n-printf \"%s\\0\" \\\n-\t\"$hello_oid blob $hello_size\" \\\n+    printf \"%s\\0\" \\\n+\t\"$boid blob $hello_size\" \\\n \t\"$hello_content\" \\\n-\t\"$commit_oid commit $commit_size\" \\\n-\t\"$commit_content\" \\\n-\t\"$tag_oid tag $tag_size\" \\\n-\t\"$tag_content\" \\\n+\t\"$coid commit $csize\" \\\n+\t\"$ccontent\" \\\n+\t\"$toid tag $tsize\" \\\n+\t\"$tcontent\" \\\n \t\"deadbeef missing\" \\\n \t\" missing\" >batch_output\n \n-test_expect_success '--batch with multiple oids gives correct format' '\n+    test_expect_success '--batch with multiple oids gives correct format' '\n \ttr \"\\0\" \"\\n\" <batch_output >expect &&\n \techo_without_newline \"$batch_input\" >in &&\n \tgit cat-file --batch <in >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success '--batch, -z with multiple oids gives correct format' '\n+    test_expect_success '--batch, -z with multiple oids gives correct format' '\n \techo_without_newline_nul \"$batch_input\" >in &&\n \ttr \"\\0\" \"\\n\" <batch_output >expect &&\n \tgit cat-file --batch -z <in >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success '--batch, -Z with multiple oids gives correct format' '\n+    test_expect_success '--batch, -Z with multiple oids gives correct format' '\n \techo_without_newline_nul \"$batch_input\" >in &&\n \tgit cat-file --batch -Z <in >actual &&\n \ttest_cmp batch_output actual\n-'\n+    '\n \n-batch_check_input=\"$hello_oid\n-$tree_oid\n-$commit_oid\n-$tag_oid\n+batch_check_input=\"$boid\n+$loid\n+$coid\n+$toid\n deadbeef\n \n \"\n \n-printf \"%s\\0\" \\\n-\t\"$hello_oid blob $hello_size\" \\\n-\t\"$tree_oid tree $tree_size\" \\\n-\t\"$commit_oid commit $commit_size\" \\\n-\t\"$tag_oid tag $tag_size\" \\\n+    printf \"%s\\0\" \\\n+\t\"$boid blob $hello_size\" \\\n+\t\"$loid tree $lsize\" \\\n+\t\"$coid commit $csize\" \\\n+\t\"$toid tag $tsize\" \\\n \t\"deadbeef missing\" \\\n \t\" missing\" >batch_check_output\n \n-test_expect_success \"--batch-check with multiple oids gives correct format\" '\n+    test_expect_success \"--batch-check with multiple oids gives correct format\" '\n \ttr \"\\0\" \"\\n\" <batch_check_output >expect &&\n \techo_without_newline \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check <in >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success \"--batch-check, -z with multiple oids gives correct format\" '\n+    test_expect_success \"--batch-check, -z with multiple oids gives correct format\" '\n \ttr \"\\0\" \"\\n\" <batch_check_output >expect &&\n \techo_without_newline_nul \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check -z <in >actual &&\n \ttest_cmp expect actual\n-'\n+    '\n \n-test_expect_success \"--batch-check, -Z with multiple oids gives correct format\" '\n+    test_expect_success \"--batch-check, -Z with multiple oids gives correct format\" '\n \techo_without_newline_nul \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check -Z <in >actual &&\n \ttest_cmp batch_check_output actual\n-'\n-\n-test_expect_success FUNNYNAMES 'setup with newline in input' '\n-\ttouch -- \"newline${LF}embedded\" &&\n-\tgit add -- \"newline${LF}embedded\" &&\n-\tgit commit -m \"file with newline embedded\" &&\n-\ttest_tick &&\n-\n-\tprintf \"HEAD:newline${LF}embedded\" >in\n-'\n-\n-test_expect_success FUNNYNAMES '--batch-check, -z with newline in input' '\n-\tgit cat-file --batch-check -z <in >actual &&\n-\techo \"$(git rev-parse \"HEAD:newline${LF}embedded\") blob 0\" >expect &&\n-\ttest_cmp expect actual\n-'\n-\n-test_expect_success FUNNYNAMES '--batch-check, -Z with newline in input' '\n-\tgit cat-file --batch-check -Z <in >actual &&\n-\tprintf \"%s\\0\" \"$(git rev-parse \"HEAD:newline${LF}embedded\") blob 0\" >expect &&\n-\ttest_cmp expect actual\n-'\n+    '\n \n-batch_command_multiple_info=\"info $hello_oid\n-info $tree_oid\n-info $commit_oid\n-info $tag_oid\n+batch_command_multiple_info=\"info $boid\n+info $loid\n+info $coid\n+info $toid\n info deadbeef\"\n \n-test_expect_success '--batch-command with multiple info calls gives correct format' '\n+    test_expect_success '--batch-command with multiple info calls gives correct format' '\n \tcat >expect <<-EOF &&\n-\t$hello_oid blob $hello_size\n-\t$tree_oid tree $tree_size\n-\t$commit_oid commit $commit_size\n-\t$tag_oid tag $tag_size\n+\t$boid blob $hello_size\n+\t$loid tree $lsize\n+\t$coid commit $csize\n+\t$toid tag $tsize\n \tdeadbeef missing\n \tEOF\n \n@@ -510,22 +540,22 @@ test_expect_success '--batch-command with multiple info calls gives correct form\n \tgit cat-file --batch-command --buffer -Z <in >actual &&\n \n \ttest_cmp expect_nul actual\n-'\n+    '\n \n-batch_command_multiple_contents=\"contents $hello_oid\n-contents $commit_oid\n-contents $tag_oid\n+batch_command_multiple_contents=\"contents $boid\n+contents $coid\n+contents $toid\n contents deadbeef\n flush\"\n \n-test_expect_success '--batch-command with multiple command calls gives correct format' '\n+    test_expect_success '--batch-command with multiple command calls gives correct format' '\n \tprintf \"%s\\0\" \\\n-\t\t\"$hello_oid blob $hello_size\" \\\n+\t\t\"$boid blob $hello_size\" \\\n \t\t\"$hello_content\" \\\n-\t\t\"$commit_oid commit $commit_size\" \\\n-\t\t\"$commit_content\" \\\n-\t\t\"$tag_oid tag $tag_size\" \\\n-\t\t\"$tag_content\" \\\n+\t\t\"$coid commit $csize\" \\\n+\t\t\"$ccontent\" \\\n+\t\t\"$toid tag $tsize\" \\\n+\t\t\"$tcontent\" \\\n \t\t\"deadbeef missing\" >expect_nul &&\n \ttr \"\\0\" \"\\n\" <expect_nul >expect &&\n \n@@ -543,6 +573,33 @@ test_expect_success '--batch-command with multiple command calls gives correct f\n \tgit cat-file --batch-command --buffer -Z <in >actual &&\n \n \ttest_cmp expect_nul actual\n+    '\n+\n+}\n+\n+batch_tests $hello_oid $tree_oid $tree_size $commit_oid $commit_size \"$commit_content\" $tag_oid $tag_size \"$tag_content\"\n+batch_tests $hello_compat_oid $tree_compat_oid $tree_compat_size $commit_compat_oid $commit_compat_size \"$commit_compat_content\" $tag_compat_oid $tag_compat_size \"$tag_compat_content\"\n+\n+\n+test_expect_success FUNNYNAMES 'setup with newline in input' '\n+\ttouch -- \"newline${LF}embedded\" &&\n+\tgit add -- \"newline${LF}embedded\" &&\n+\tgit commit -m \"file with newline embedded\" &&\n+\ttest_tick &&\n+\n+\tprintf \"HEAD:newline${LF}embedded\" >in\n+'\n+\n+test_expect_success FUNNYNAMES '--batch-check, -z with newline in input' '\n+\tgit cat-file --batch-check -z <in >actual &&\n+\techo \"$(git rev-parse \"HEAD:newline${LF}embedded\") blob 0\" >expect &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success FUNNYNAMES '--batch-check, -Z with newline in input' '\n+\tgit cat-file --batch-check -Z <in >actual &&\n+\tprintf \"%s\\0\" \"$(git rev-parse \"HEAD:newline${LF}embedded\") blob 0\" >expect &&\n+\ttest_cmp expect actual\n '\n \n test_expect_success 'setup blobs which are likely to delta' '\n-- \n2.41.0\n\n"},{"id":"482515","messageId":"20231002024034.2611-28-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 28/30] t1006: rename sha1 to oid","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:32Z","receivedAt":"2023-10-02T02:41:59Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nBefore I extend this test, changing the naming of the relevant\nhash from sha1 to oid.  Calling the hash sha1 is incorrect today\nas it can be either sha1 or sha256 depending on the value of\nGIT_DEFAULT_HASH_FUNCTION when the test is called.\n\nI plan to test sha1 and sha256 simultaneously in the same repository.\nHaving a name like sha1 will be even more confusing.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n t/t1006-cat-file.sh | 220 ++++++++++++++++++++++----------------------\n 1 file changed, 110 insertions(+), 110 deletions(-)\n\ndiff --git a/t/t1006-cat-file.sh b/t/t1006-cat-file.sh\nindex d73a0be1b9d1..9b018b538950 100755\n--- a/t/t1006-cat-file.sh\n+++ b/t/t1006-cat-file.sh\n@@ -112,65 +112,65 @@ strlen () {\n \n run_tests () {\n     type=$1\n-    sha1=$2\n+    oid=$2\n     size=$3\n     content=$4\n     pretty_content=$5\n \n-    batch_output=\"$sha1 $type $size\n+    batch_output=\"$oid $type $size\n $content\"\n \n     test_expect_success \"$type exists\" '\n-\tgit cat-file -e $sha1\n+\tgit cat-file -e $oid\n     '\n \n     test_expect_success \"Type of $type is correct\" '\n \techo $type >expect &&\n-\tgit cat-file -t $sha1 >actual &&\n+\tgit cat-file -t $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"Size of $type is correct\" '\n \techo $size >expect &&\n-\tgit cat-file -s $sha1 >actual &&\n+\tgit cat-file -s $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"Type of $type is correct using --allow-unknown-type\" '\n \techo $type >expect &&\n-\tgit cat-file -t --allow-unknown-type $sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"Size of $type is correct using --allow-unknown-type\" '\n \techo $size >expect &&\n-\tgit cat-file -s --allow-unknown-type $sha1 >actual &&\n+\tgit cat-file -s --allow-unknown-type $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test -z \"$content\" ||\n     test_expect_success \"Content of $type is correct\" '\n \techo_without_newline \"$content\" >expect &&\n-\tgit cat-file $type $sha1 >actual &&\n+\tgit cat-file $type $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"Pretty content of $type is correct\" '\n \techo_without_newline \"$pretty_content\" >expect &&\n-\tgit cat-file -p $sha1 >actual &&\n+\tgit cat-file -p $oid >actual &&\n \ttest_cmp expect actual\n     '\n \n     test -z \"$content\" ||\n     test_expect_success \"--batch output of $type is correct\" '\n \techo \"$batch_output\" >expect &&\n-\techo $sha1 | git cat-file --batch >actual &&\n+\techo $oid | git cat-file --batch >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"--batch-check output of $type is correct\" '\n-\techo \"$sha1 $type $size\" >expect &&\n-\techo_without_newline $sha1 | git cat-file --batch-check >actual &&\n+\techo \"$oid $type $size\" >expect &&\n+\techo_without_newline $oid | git cat-file --batch-check >actual &&\n \ttest_cmp expect actual\n     '\n \n@@ -179,33 +179,33 @@ $content\"\n \ttest -z \"$content\" ||\n \t\ttest_expect_success \"--batch-command $opt output of $type content is correct\" '\n \t\techo \"$batch_output\" >expect &&\n-\t\ttest_write_lines \"contents $sha1\" | git cat-file --batch-command $opt >actual &&\n+\t\ttest_write_lines \"contents $oid\" | git cat-file --batch-command $opt >actual &&\n \t\ttest_cmp expect actual\n \t'\n \n \ttest_expect_success \"--batch-command $opt output of $type info is correct\" '\n-\t\techo \"$sha1 $type $size\" >expect &&\n-\t\ttest_write_lines \"info $sha1\" |\n+\t\techo \"$oid $type $size\" >expect &&\n+\t\ttest_write_lines \"info $oid\" |\n \t\tgit cat-file --batch-command $opt >actual &&\n \t\ttest_cmp expect actual\n \t'\n     done\n \n     test_expect_success \"custom --batch-check format\" '\n-\techo \"$type $sha1\" >expect &&\n-\techo $sha1 | git cat-file --batch-check=\"%(objecttype) %(objectname)\" >actual &&\n+\techo \"$type $oid\" >expect &&\n+\techo $oid | git cat-file --batch-check=\"%(objecttype) %(objectname)\" >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success \"custom --batch-command format\" '\n-\techo \"$type $sha1\" >expect &&\n-\techo \"info $sha1\" | git cat-file --batch-command=\"%(objecttype) %(objectname)\" >actual &&\n+\techo \"$type $oid\" >expect &&\n+\techo \"info $oid\" | git cat-file --batch-command=\"%(objecttype) %(objectname)\" >actual &&\n \ttest_cmp expect actual\n     '\n \n     test_expect_success '--batch-check with %(rest)' '\n \techo \"$type this is some extra content\" >expect &&\n-\techo \"$sha1    this is some extra content\" |\n+\techo \"$oid    this is some extra content\" |\n \t\tgit cat-file --batch-check=\"%(objecttype) %(rest)\" >actual &&\n \ttest_cmp expect actual\n     '\n@@ -216,7 +216,7 @@ $content\"\n \t\techo \"$size\" &&\n \t\techo \"$content\"\n \t} >expect &&\n-\techo $sha1 | git cat-file --batch=\"%(objectsize)\" >actual &&\n+\techo $oid | git cat-file --batch=\"%(objectsize)\" >actual &&\n \ttest_cmp expect actual\n     '\n \n@@ -226,25 +226,25 @@ $content\"\n \t\techo \"$type\" &&\n \t\techo \"$content\"\n \t} >expect &&\n-\techo $sha1 | git cat-file --batch=\"%(objecttype)\" >actual &&\n+\techo $oid | git cat-file --batch=\"%(objecttype)\" >actual &&\n \ttest_cmp expect actual\n     '\n }\n \n hello_content=\"Hello World\"\n hello_size=$(strlen \"$hello_content\")\n-hello_sha1=$(echo_without_newline \"$hello_content\" | git hash-object --stdin)\n+hello_oid=$(echo_without_newline \"$hello_content\" | git hash-object --stdin)\n \n test_expect_success \"setup\" '\n \techo_without_newline \"$hello_content\" > hello &&\n \tgit update-index --add hello\n '\n \n-run_tests 'blob' $hello_sha1 $hello_size \"$hello_content\" \"$hello_content\"\n+run_tests 'blob' $hello_oid $hello_size \"$hello_content\" \"$hello_content\"\n \n test_expect_success '--batch-command --buffer with flush for blob info' '\n-\techo \"$hello_sha1 blob $hello_size\" >expect &&\n-\ttest_write_lines \"info $hello_sha1\" \"flush\" |\n+\techo \"$hello_oid blob $hello_size\" >expect &&\n+\ttest_write_lines \"info $hello_oid\" \"flush\" |\n \tGIT_TEST_CAT_FILE_NO_FLUSH_ON_EXIT=1 \\\n \tgit cat-file --batch-command --buffer >actual &&\n \ttest_cmp expect actual\n@@ -252,38 +252,38 @@ test_expect_success '--batch-command --buffer with flush for blob info' '\n \n test_expect_success '--batch-command --buffer without flush for blob info' '\n \ttouch output &&\n-\ttest_write_lines \"info $hello_sha1\" |\n+\ttest_write_lines \"info $hello_oid\" |\n \tGIT_TEST_CAT_FILE_NO_FLUSH_ON_EXIT=1 \\\n \tgit cat-file --batch-command --buffer >>output &&\n \ttest_must_be_empty output\n '\n \n test_expect_success '--batch-check without %(rest) considers whole line' '\n-\techo \"$hello_sha1 blob $hello_size\" >expect &&\n-\tgit update-index --add --cacheinfo 100644 $hello_sha1 \"white space\" &&\n+\techo \"$hello_oid blob $hello_size\" >expect &&\n+\tgit update-index --add --cacheinfo 100644 $hello_oid \"white space\" &&\n \ttest_when_finished \"git update-index --remove \\\"white space\\\"\" &&\n \techo \":white space\" | git cat-file --batch-check >actual &&\n \ttest_cmp expect actual\n '\n \n-tree_sha1=$(git write-tree)\n+tree_oid=$(git write-tree)\n tree_size=$(($(test_oid rawsz) + 13))\n-tree_pretty_content=\"100644 blob $hello_sha1\thello${LF}\"\n+tree_pretty_content=\"100644 blob $hello_oid\thello${LF}\"\n \n-run_tests 'tree' $tree_sha1 $tree_size \"\" \"$tree_pretty_content\"\n+run_tests 'tree' $tree_oid $tree_size \"\" \"$tree_pretty_content\"\n \n commit_message=\"Initial commit\"\n-commit_sha1=$(echo_without_newline \"$commit_message\" | git commit-tree $tree_sha1)\n+commit_oid=$(echo_without_newline \"$commit_message\" | git commit-tree $tree_oid)\n commit_size=$(($(test_oid hexsz) + 137))\n-commit_content=\"tree $tree_sha1\n+commit_content=\"tree $tree_oid\n author $GIT_AUTHOR_NAME <$GIT_AUTHOR_EMAIL> $GIT_AUTHOR_DATE\n committer $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL> $GIT_COMMITTER_DATE\n \n $commit_message\"\n \n-run_tests 'commit' $commit_sha1 $commit_size \"$commit_content\" \"$commit_content\"\n+run_tests 'commit' $commit_oid $commit_size \"$commit_content\" \"$commit_content\"\n \n-tag_header_without_timestamp=\"object $hello_sha1\n+tag_header_without_timestamp=\"object $hello_oid\n type blob\n tag hellotag\n tagger $GIT_COMMITTER_NAME <$GIT_COMMITTER_EMAIL>\"\n@@ -292,14 +292,14 @@ tag_content=\"$tag_header_without_timestamp 0 +0000\n \n $tag_description\"\n \n-tag_sha1=$(echo_without_newline \"$tag_content\" | git hash-object -t tag --stdin -w)\n+tag_oid=$(echo_without_newline \"$tag_content\" | git hash-object -t tag --stdin -w)\n tag_size=$(strlen \"$tag_content\")\n \n-run_tests 'tag' $tag_sha1 $tag_size \"$tag_content\" \"$tag_content\"\n+run_tests 'tag' $tag_oid $tag_size \"$tag_content\" \"$tag_content\"\n \n test_expect_success \"Reach a blob from a tag pointing to it\" '\n \techo_without_newline \"$hello_content\" >expect &&\n-\tgit cat-file blob $tag_sha1 >actual &&\n+\tgit cat-file blob $tag_oid >actual &&\n \ttest_cmp expect actual\n '\n \n@@ -308,31 +308,31 @@ do\n     for opt in t s e p\n     do\n \ttest_expect_success \"Passing -$opt with --$batch fails\" '\n-\t    test_must_fail git cat-file --$batch -$opt $hello_sha1\n+\t    test_must_fail git cat-file --$batch -$opt $hello_oid\n \t'\n \n \ttest_expect_success \"Passing --$batch with -$opt fails\" '\n-\t    test_must_fail git cat-file -$opt --$batch $hello_sha1\n+\t    test_must_fail git cat-file -$opt --$batch $hello_oid\n \t'\n     done\n \n     test_expect_success \"Passing <type> with --$batch fails\" '\n-\ttest_must_fail git cat-file --$batch blob $hello_sha1\n+\ttest_must_fail git cat-file --$batch blob $hello_oid\n     '\n \n     test_expect_success \"Passing --$batch with <type> fails\" '\n-\ttest_must_fail git cat-file blob --$batch $hello_sha1\n+\ttest_must_fail git cat-file blob --$batch $hello_oid\n     '\n \n-    test_expect_success \"Passing sha1 with --$batch fails\" '\n-\ttest_must_fail git cat-file --$batch $hello_sha1\n+    test_expect_success \"Passing oid with --$batch fails\" '\n+\ttest_must_fail git cat-file --$batch $hello_oid\n     '\n done\n \n for opt in t s e p\n do\n     test_expect_success \"Passing -$opt with --follow-symlinks fails\" '\n-\t    test_must_fail git cat-file --follow-symlinks -$opt $hello_sha1\n+\t    test_must_fail git cat-file --follow-symlinks -$opt $hello_oid\n \t'\n done\n \n@@ -360,12 +360,12 @@ test_expect_success \"--batch-check for a non-existent hash\" '\n \n test_expect_success \"--batch for an existent and a non-existent hash\" '\n \tcat >expect <<-EOF &&\n-\t$tag_sha1 tag $tag_size\n+\t$tag_oid tag $tag_size\n \t$tag_content\n \t0000000000000000000000000000000000000000 missing\n \tEOF\n \n-\tprintf \"$tag_sha1\\n0000000000000000000000000000000000000000\" >in &&\n+\tprintf \"$tag_oid\\n0000000000000000000000000000000000000000\" >in &&\n \tgit cat-file --batch <in >actual &&\n \ttest_cmp expect actual\n '\n@@ -386,74 +386,74 @@ test_expect_success 'empty --batch-check notices missing object' '\n \ttest_cmp expect actual\n '\n \n-batch_input=\"$hello_sha1\n-$commit_sha1\n-$tag_sha1\n+batch_input=\"$hello_oid\n+$commit_oid\n+$tag_oid\n deadbeef\n \n \"\n \n printf \"%s\\0\" \\\n-\t\"$hello_sha1 blob $hello_size\" \\\n+\t\"$hello_oid blob $hello_size\" \\\n \t\"$hello_content\" \\\n-\t\"$commit_sha1 commit $commit_size\" \\\n+\t\"$commit_oid commit $commit_size\" \\\n \t\"$commit_content\" \\\n-\t\"$tag_sha1 tag $tag_size\" \\\n+\t\"$tag_oid tag $tag_size\" \\\n \t\"$tag_content\" \\\n \t\"deadbeef missing\" \\\n \t\" missing\" >batch_output\n \n-test_expect_success '--batch with multiple sha1s gives correct format' '\n+test_expect_success '--batch with multiple oids gives correct format' '\n \ttr \"\\0\" \"\\n\" <batch_output >expect &&\n \techo_without_newline \"$batch_input\" >in &&\n \tgit cat-file --batch <in >actual &&\n \ttest_cmp expect actual\n '\n \n-test_expect_success '--batch, -z with multiple sha1s gives correct format' '\n+test_expect_success '--batch, -z with multiple oids gives correct format' '\n \techo_without_newline_nul \"$batch_input\" >in &&\n \ttr \"\\0\" \"\\n\" <batch_output >expect &&\n \tgit cat-file --batch -z <in >actual &&\n \ttest_cmp expect actual\n '\n \n-test_expect_success '--batch, -Z with multiple sha1s gives correct format' '\n+test_expect_success '--batch, -Z with multiple oids gives correct format' '\n \techo_without_newline_nul \"$batch_input\" >in &&\n \tgit cat-file --batch -Z <in >actual &&\n \ttest_cmp batch_output actual\n '\n \n-batch_check_input=\"$hello_sha1\n-$tree_sha1\n-$commit_sha1\n-$tag_sha1\n+batch_check_input=\"$hello_oid\n+$tree_oid\n+$commit_oid\n+$tag_oid\n deadbeef\n \n \"\n \n printf \"%s\\0\" \\\n-\t\"$hello_sha1 blob $hello_size\" \\\n-\t\"$tree_sha1 tree $tree_size\" \\\n-\t\"$commit_sha1 commit $commit_size\" \\\n-\t\"$tag_sha1 tag $tag_size\" \\\n+\t\"$hello_oid blob $hello_size\" \\\n+\t\"$tree_oid tree $tree_size\" \\\n+\t\"$commit_oid commit $commit_size\" \\\n+\t\"$tag_oid tag $tag_size\" \\\n \t\"deadbeef missing\" \\\n \t\" missing\" >batch_check_output\n \n-test_expect_success \"--batch-check with multiple sha1s gives correct format\" '\n+test_expect_success \"--batch-check with multiple oids gives correct format\" '\n \ttr \"\\0\" \"\\n\" <batch_check_output >expect &&\n \techo_without_newline \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check <in >actual &&\n \ttest_cmp expect actual\n '\n \n-test_expect_success \"--batch-check, -z with multiple sha1s gives correct format\" '\n+test_expect_success \"--batch-check, -z with multiple oids gives correct format\" '\n \ttr \"\\0\" \"\\n\" <batch_check_output >expect &&\n \techo_without_newline_nul \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check -z <in >actual &&\n \ttest_cmp expect actual\n '\n \n-test_expect_success \"--batch-check, -Z with multiple sha1s gives correct format\" '\n+test_expect_success \"--batch-check, -Z with multiple oids gives correct format\" '\n \techo_without_newline_nul \"$batch_check_input\" >in &&\n \tgit cat-file --batch-check -Z <in >actual &&\n \ttest_cmp batch_check_output actual\n@@ -480,18 +480,18 @@ test_expect_success FUNNYNAMES '--batch-check, -Z with newline in input' '\n \ttest_cmp expect actual\n '\n \n-batch_command_multiple_info=\"info $hello_sha1\n-info $tree_sha1\n-info $commit_sha1\n-info $tag_sha1\n+batch_command_multiple_info=\"info $hello_oid\n+info $tree_oid\n+info $commit_oid\n+info $tag_oid\n info deadbeef\"\n \n test_expect_success '--batch-command with multiple info calls gives correct format' '\n \tcat >expect <<-EOF &&\n-\t$hello_sha1 blob $hello_size\n-\t$tree_sha1 tree $tree_size\n-\t$commit_sha1 commit $commit_size\n-\t$tag_sha1 tag $tag_size\n+\t$hello_oid blob $hello_size\n+\t$tree_oid tree $tree_size\n+\t$commit_oid commit $commit_size\n+\t$tag_oid tag $tag_size\n \tdeadbeef missing\n \tEOF\n \n@@ -512,19 +512,19 @@ test_expect_success '--batch-command with multiple info calls gives correct form\n \ttest_cmp expect_nul actual\n '\n \n-batch_command_multiple_contents=\"contents $hello_sha1\n-contents $commit_sha1\n-contents $tag_sha1\n+batch_command_multiple_contents=\"contents $hello_oid\n+contents $commit_oid\n+contents $tag_oid\n contents deadbeef\n flush\"\n \n test_expect_success '--batch-command with multiple command calls gives correct format' '\n \tprintf \"%s\\0\" \\\n-\t\t\"$hello_sha1 blob $hello_size\" \\\n+\t\t\"$hello_oid blob $hello_size\" \\\n \t\t\"$hello_content\" \\\n-\t\t\"$commit_sha1 commit $commit_size\" \\\n+\t\t\"$commit_oid commit $commit_size\" \\\n \t\t\"$commit_content\" \\\n-\t\t\"$tag_sha1 tag $tag_size\" \\\n+\t\t\"$tag_oid tag $tag_size\" \\\n \t\t\"$tag_content\" \\\n \t\t\"deadbeef missing\" >expect_nul &&\n \ttr \"\\0\" \"\\n\" <expect_nul >expect &&\n@@ -569,7 +569,7 @@ test_expect_success 'confirm that neither loose blob is a delta' '\n # we will check only that one of the two objects is a delta\n # against the other, but not the order. We can do so by just\n # asking for the base of both, and checking whether either\n-# sha1 appears in the output.\n+# oid appears in the output.\n test_expect_success '%(deltabase) reports packed delta bases' '\n \tgit repack -ad &&\n \tgit cat-file --batch-check=\"%(deltabase)\" <blobs >actual &&\n@@ -583,12 +583,12 @@ test_expect_success 'setup bogus data' '\n \tbogus_short_type=\"bogus\" &&\n \tbogus_short_content=\"bogus\" &&\n \tbogus_short_size=$(strlen \"$bogus_short_content\") &&\n-\tbogus_short_sha1=$(echo_without_newline \"$bogus_short_content\" | git hash-object -t $bogus_short_type --literally -w --stdin) &&\n+\tbogus_short_oid=$(echo_without_newline \"$bogus_short_content\" | git hash-object -t $bogus_short_type --literally -w --stdin) &&\n \n \tbogus_long_type=\"abcdefghijklmnopqrstuvwxyz1234679\" &&\n \tbogus_long_content=\"bogus\" &&\n \tbogus_long_size=$(strlen \"$bogus_long_content\") &&\n-\tbogus_long_sha1=$(echo_without_newline \"$bogus_long_content\" | git hash-object -t $bogus_long_type --literally -w --stdin)\n+\tbogus_long_oid=$(echo_without_newline \"$bogus_long_content\" | git hash-object -t $bogus_long_type --literally -w --stdin)\n '\n \n for arg1 in '' --allow-unknown-type\n@@ -608,9 +608,9 @@ do\n \n \t\t\tif test \"$arg1\" = \"--allow-unknown-type\"\n \t\t\tthen\n-\t\t\t\tgit cat-file $arg1 $arg2 $bogus_short_sha1\n+\t\t\t\tgit cat-file $arg1 $arg2 $bogus_short_oid\n \t\t\telse\n-\t\t\t\ttest_must_fail git cat-file $arg1 $arg2 $bogus_short_sha1 >out 2>actual &&\n+\t\t\t\ttest_must_fail git cat-file $arg1 $arg2 $bogus_short_oid >out 2>actual &&\n \t\t\t\ttest_must_be_empty out &&\n \t\t\t\ttest_cmp expect actual\n \t\t\tfi\n@@ -620,21 +620,21 @@ do\n \t\t\tif test \"$arg2\" = \"-p\"\n \t\t\tthen\n \t\t\t\tcat >expect <<-EOF\n-\t\t\t\terror: header for $bogus_long_sha1 too long, exceeds 32 bytes\n-\t\t\t\tfatal: Not a valid object name $bogus_long_sha1\n+\t\t\t\terror: header for $bogus_long_oid too long, exceeds 32 bytes\n+\t\t\t\tfatal: Not a valid object name $bogus_long_oid\n \t\t\t\tEOF\n \t\t\telse\n \t\t\t\tcat >expect <<-EOF\n-\t\t\t\terror: header for $bogus_long_sha1 too long, exceeds 32 bytes\n+\t\t\t\terror: header for $bogus_long_oid too long, exceeds 32 bytes\n \t\t\t\tfatal: git cat-file: could not get object info\n \t\t\t\tEOF\n \t\t\tfi &&\n \n \t\t\tif test \"$arg1\" = \"--allow-unknown-type\"\n \t\t\tthen\n-\t\t\t\tgit cat-file $arg1 $arg2 $bogus_short_sha1\n+\t\t\t\tgit cat-file $arg1 $arg2 $bogus_short_oid\n \t\t\telse\n-\t\t\t\ttest_must_fail git cat-file $arg1 $arg2 $bogus_long_sha1 >out 2>actual &&\n+\t\t\t\ttest_must_fail git cat-file $arg1 $arg2 $bogus_long_oid >out 2>actual &&\n \t\t\t\ttest_must_be_empty out &&\n \t\t\t\ttest_cmp expect actual\n \t\t\tfi\n@@ -668,28 +668,28 @@ do\n done\n \n test_expect_success '-e is OK with a broken object without --allow-unknown-type' '\n-\tgit cat-file -e $bogus_short_sha1\n+\tgit cat-file -e $bogus_short_oid\n '\n \n test_expect_success '-e can not be combined with --allow-unknown-type' '\n-\ttest_expect_code 128 git cat-file -e --allow-unknown-type $bogus_short_sha1\n+\ttest_expect_code 128 git cat-file -e --allow-unknown-type $bogus_short_oid\n '\n \n test_expect_success '-p cannot print a broken object even with --allow-unknown-type' '\n-\ttest_must_fail git cat-file -p $bogus_short_sha1 &&\n-\ttest_expect_code 128 git cat-file -p --allow-unknown-type $bogus_short_sha1\n+\ttest_must_fail git cat-file -p $bogus_short_oid &&\n+\ttest_expect_code 128 git cat-file -p --allow-unknown-type $bogus_short_oid\n '\n \n test_expect_success '<type> <hash> does not work with objects of broken types' '\n \tcat >err.expect <<-\\EOF &&\n \tfatal: invalid object type \"bogus\"\n \tEOF\n-\ttest_must_fail git cat-file $bogus_short_type $bogus_short_sha1 2>err.actual &&\n+\ttest_must_fail git cat-file $bogus_short_type $bogus_short_oid 2>err.actual &&\n \ttest_cmp err.expect err.actual\n '\n \n test_expect_success 'broken types combined with --batch and --batch-check' '\n-\techo $bogus_short_sha1 >bogus-oid &&\n+\techo $bogus_short_oid >bogus-oid &&\n \n \tcat >err.expect <<-\\EOF &&\n \tfatal: invalid object type\n@@ -711,52 +711,52 @@ test_expect_success 'the --allow-unknown-type option does not consider replaceme\n \tcat >expect <<-EOF &&\n \t$bogus_short_type\n \tEOF\n-\tgit cat-file -t --allow-unknown-type $bogus_short_sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $bogus_short_oid >actual &&\n \ttest_cmp expect actual &&\n \n \t# Create it manually, as \"git replace\" will die on bogus\n \t# types.\n \thead=$(git rev-parse --verify HEAD) &&\n-\ttest_when_finished \"test-tool ref-store main delete-refs 0 msg refs/replace/$bogus_short_sha1\" &&\n-\ttest-tool ref-store main update-ref msg \"refs/replace/$bogus_short_sha1\" $head $ZERO_OID REF_SKIP_OID_VERIFICATION &&\n+\ttest_when_finished \"test-tool ref-store main delete-refs 0 msg refs/replace/$bogus_short_oid\" &&\n+\ttest-tool ref-store main update-ref msg \"refs/replace/$bogus_short_oid\" $head $ZERO_OID REF_SKIP_OID_VERIFICATION &&\n \n \tcat >expect <<-EOF &&\n \tcommit\n \tEOF\n-\tgit cat-file -t --allow-unknown-type $bogus_short_sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $bogus_short_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success \"Type of broken object is correct\" '\n \techo $bogus_short_type >expect &&\n-\tgit cat-file -t --allow-unknown-type $bogus_short_sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $bogus_short_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success \"Size of broken object is correct\" '\n \techo $bogus_short_size >expect &&\n-\tgit cat-file -s --allow-unknown-type $bogus_short_sha1 >actual &&\n+\tgit cat-file -s --allow-unknown-type $bogus_short_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success 'clean up broken object' '\n-\trm .git/objects/$(test_oid_to_path $bogus_short_sha1)\n+\trm .git/objects/$(test_oid_to_path $bogus_short_oid)\n '\n \n test_expect_success \"Type of broken object is correct when type is large\" '\n \techo $bogus_long_type >expect &&\n-\tgit cat-file -t --allow-unknown-type $bogus_long_sha1 >actual &&\n+\tgit cat-file -t --allow-unknown-type $bogus_long_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success \"Size of large broken object is correct when type is large\" '\n \techo $bogus_long_size >expect &&\n-\tgit cat-file -s --allow-unknown-type $bogus_long_sha1 >actual &&\n+\tgit cat-file -s --allow-unknown-type $bogus_long_oid >actual &&\n \ttest_cmp expect actual\n '\n \n test_expect_success 'clean up broken object' '\n-\trm .git/objects/$(test_oid_to_path $bogus_long_sha1)\n+\trm .git/objects/$(test_oid_to_path $bogus_long_oid)\n '\n \n test_expect_success 'cat-file -t and -s on corrupt loose object' '\n@@ -853,7 +853,7 @@ test_expect_success 'prep for symlink tests' '\n \ttest_ln_s_add loop2 loop1 &&\n \tgit add morx dir/subdir/ind2 dir/ind1 &&\n \tgit commit -am \"test\" &&\n-\techo $hello_sha1 blob $hello_size >found\n+\techo $hello_oid blob $hello_size >found\n '\n \n test_expect_success 'git cat-file --batch-check --follow-symlinks works for non-links' '\n@@ -941,7 +941,7 @@ test_expect_success 'git cat-file --batch-check --follow-symlinks works for dir/\n \techo HEAD:dirlink/morx >>expect &&\n \techo HEAD:dirlink/morx | git cat-file --batch-check --follow-symlinks >actual &&\n \ttest_cmp expect actual &&\n-\techo $hello_sha1 blob $hello_size >expect &&\n+\techo $hello_oid blob $hello_size >expect &&\n \techo HEAD:dirlink/ind1 | git cat-file --batch-check --follow-symlinks >actual &&\n \ttest_cmp expect actual\n '\n-- \n2.41.0\n\n"},{"id":"482516","messageId":"20231002024034.2611-30-ebiederm@gmail.com","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"[PATCH v2 30/30] t1016-compatObjectFormat: add tests to verify the conversion between objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2023-10-02T02:40:34Z","receivedAt":"2023-10-02T02:42:07Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nFor now my strategy is simple.  Create two identical repositories one\nin each format.  Use fixed timestamps. Verify the dynamically computed\ncompatibility objects from one repository match the objects stored in\nthe other repository.\n\nA general limitation of this strategy is that the git when generating\nsigned tags and commits with compatObjectFormat enabled will generate\na signature for both formats.  To overcome this limitation I have\nadded \"test-tool delete-gpgsig\" that when fed an signed commit or tag\nwith two signatures deletes one of the signatures.\n\nWith that in place I can have \"git commit\" and  \"git tag\" generate\nsigned objects, have my tool delete one, and feed the new object\ninto \"git hash-object\" to create the kinds of commits and tags\ngit without compatObjectFormat enabled will generate.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Makefile                      |   1 +\n t/helper/test-delete-gpgsig.c |  62 ++++++++\n t/helper/test-tool.c          |   1 +\n t/helper/test-tool.h          |   1 +\n t/t1016-compatObjectFormat.sh | 281 ++++++++++++++++++++++++++++++++++\n t/t1016/gpg                   |   2 +\n 6 files changed, 348 insertions(+)\n create mode 100644 t/helper/test-delete-gpgsig.c\n create mode 100755 t/t1016-compatObjectFormat.sh\n create mode 100755 t/t1016/gpg\n\ndiff --git a/Makefile b/Makefile\nindex 3c18664def9a..3e4444fb9ab2 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -790,6 +790,7 @@ TEST_BUILTINS_OBJS += test-crontab.o\n TEST_BUILTINS_OBJS += test-csprng.o\n TEST_BUILTINS_OBJS += test-ctype.o\n TEST_BUILTINS_OBJS += test-date.o\n+TEST_BUILTINS_OBJS += test-delete-gpgsig.o\n TEST_BUILTINS_OBJS += test-delta.o\n TEST_BUILTINS_OBJS += test-dir-iterator.o\n TEST_BUILTINS_OBJS += test-drop-caches.o\ndiff --git a/t/helper/test-delete-gpgsig.c b/t/helper/test-delete-gpgsig.c\nnew file mode 100644\nindex 000000000000..e36831af03f6\n--- /dev/null\n+++ b/t/helper/test-delete-gpgsig.c\n@@ -0,0 +1,62 @@\n+#include \"test-tool.h\"\n+#include \"gpg-interface.h\"\n+#include \"strbuf.h\"\n+\n+\n+int cmd__delete_gpgsig(int argc, const char **argv)\n+{\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tconst char *pattern = \"gpgsig\";\n+\tconst char *bufptr, *tail, *eol;\n+\tint deleting = 0;\n+\tsize_t plen;\n+\n+\tif (argc >= 2) {\n+\t\tpattern = argv[1];\n+\t\targv++;\n+\t\targc--;\n+\t}\n+\n+\tplen = strlen(pattern);\n+\tstrbuf_read(&buf, 0, 0);\n+\n+\tif (!strcmp(pattern, \"trailer\")) {\n+\t\tsize_t payload_size = parse_signed_buffer(buf.buf, buf.len);\n+\t\tfwrite(buf.buf, 1, payload_size, stdout);\n+\t\tfflush(stdout);\n+\t\treturn 0;\n+\t}\n+\n+\tbufptr = buf.buf;\n+\ttail = bufptr + buf.len;\n+\n+\twhile (bufptr < tail) {\n+\t\t/* Find the end of the line */\n+\t\teol = memchr(bufptr, '\\n', tail - bufptr);\n+\t\tif (!eol)\n+\t\t\teol = tail;\n+\n+\t\t/* Drop continuation lines */\n+\t\tif (deleting && (bufptr < eol) && (bufptr[0] == ' ')) {\n+\t\t\tbufptr = eol + 1;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tdeleting = 0;\n+\n+\t\t/* Does the line match the prefix? */\n+\t\tif (((bufptr + plen) < eol) &&\n+\t\t    !memcmp(bufptr, pattern, plen) &&\n+\t\t    (bufptr[plen] == ' ')) {\n+\t\t\tdeleting = 1;\n+\t\t\tbufptr = eol + 1;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\t/* Print all other lines */\n+\t\tfwrite(bufptr, 1, (eol - bufptr) + 1, stdout);\n+\t\tbufptr = eol + 1;\n+\t}\n+\tfflush(stdout);\n+\n+\treturn 0;\n+}\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex abe8a785eb65..8b6c84f202d6 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -21,6 +21,7 @@ static struct test_cmd cmds[] = {\n \t{ \"csprng\", cmd__csprng },\n \t{ \"ctype\", cmd__ctype },\n \t{ \"date\", cmd__date },\n+\t{ \"delete-gpgsig\", cmd__delete_gpgsig },\n \t{ \"delta\", cmd__delta },\n \t{ \"dir-iterator\", cmd__dir_iterator },\n \t{ \"drop-caches\", cmd__drop_caches },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex ea2672436c9a..76baaece35b9 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -15,6 +15,7 @@ int cmd__csprng(int argc, const char **argv);\n int cmd__ctype(int argc, const char **argv);\n int cmd__date(int argc, const char **argv);\n int cmd__delta(int argc, const char **argv);\n+int cmd__delete_gpgsig(int argc, const char **argv);\n int cmd__dir_iterator(int argc, const char **argv);\n int cmd__drop_caches(int argc, const char **argv);\n int cmd__dump_cache_tree(int argc, const char **argv);\ndiff --git a/t/t1016-compatObjectFormat.sh b/t/t1016-compatObjectFormat.sh\nnew file mode 100755\nindex 000000000000..8132cd37b8c8\n--- /dev/null\n+++ b/t/t1016-compatObjectFormat.sh\n@@ -0,0 +1,281 @@\n+#!/bin/sh\n+#\n+# Copyright (c) 2023 Eric Biederman\n+#\n+\n+test_description='Test how well compatObjectFormat works'\n+\n+TEST_PASSES_SANITIZE_LEAK=true\n+. ./test-lib.sh\n+. \"$TEST_DIRECTORY\"/lib-gpg.sh\n+\n+# All of the follow variables must be defined in the environment:\n+# GIT_AUTHOR_NAME\n+# GIT_AUTHOR_EMAIL\n+# GIT_AUTHOR_DATE\n+# GIT_COMMITTER_NAME\n+# GIT_COMMITTER_EMAIL\n+# GIT_COMMITTER_DATE\n+#\n+# The test relies on these variables being set so that the two\n+# different commits in two different repositories encoded with two\n+# different hash functions result in the same content in the commits.\n+# This means that when the commit is translated between hash functions\n+# the commit is identical to the commit in the other repository.\n+\n+compat_hash () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"sha256\"\n+\t;;\n+    \"sha256\")\n+\techo \"sha1\"\n+\t;;\n+    esac\n+}\n+\n+hello_oid () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"$hello_sha1_oid\"\n+\t;;\n+    \"sha256\")\n+\techo \"$hello_sha256_oid\"\n+\t;;\n+    esac\n+}\n+\n+tree_oid () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"$tree_sha1_oid\"\n+\t;;\n+    \"sha256\")\n+\techo \"$tree_sha256_oid\"\n+\t;;\n+    esac\n+}\n+\n+commit_oid () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"$commit_sha1_oid\"\n+\t;;\n+    \"sha256\")\n+\techo \"$commit_sha256_oid\"\n+\t;;\n+    esac\n+}\n+\n+commit2_oid () {\n+    case \"$1\" in\n+    \"sha1\")\n+\techo \"$commit2_sha1_oid\"\n+\t;;\n+    \"sha256\")\n+\techo \"$commit2_sha256_oid\"\n+\t;;\n+    esac\n+}\n+\n+del_sigcommit () {\n+    local delete=$1\n+\n+    if test \"$delete\" = \"sha256\" ; then\n+\tlocal pattern=\"gpgsig-sha256\"\n+    else\n+\tlocal pattern=\"gpgsig\"\n+    fi\n+    test-tool delete-gpgsig \"$pattern\"\n+}\n+\n+\n+del_sigtag () {\n+    local storage=$1\n+    local delete=$2\n+\n+    if test \"$storage\" = \"$delete\" ; then\n+\tlocal pattern=\"trailer\"\n+    elif test \"$storage\" = \"sha256\" ; then\n+\tlocal pattern=\"gpgsig\"\n+    else\n+\tlocal pattern=\"gpgsig-sha256\"\n+    fi\n+    test-tool delete-gpgsig \"$pattern\"\n+}\n+\n+base=$(pwd)\n+for hash in sha1 sha256\n+do\n+\tcd \"$base\"\n+\tmkdir -p repo-$hash\n+\tcd repo-$hash\n+\n+\ttest_expect_success \"setup $hash repository\" '\n+\t\tgit init --object-format=$hash &&\n+\t\tgit config core.repositoryformatversion 1 &&\n+\t\tgit config extensions.objectformat $hash &&\n+\t\tgit config extensions.compatobjectformat $(compat_hash $hash) &&\n+\t\tgit config gpg.program $TEST_DIRECTORY/t1016/gpg &&\n+\t\techo \"Hellow World!\" > hello &&\n+\t\teval hello_${hash}_oid=$(git hash-object hello) &&\n+\t\tgit update-index --add hello &&\n+\t\tgit commit -m \"Initial commit\" &&\n+\t\teval commit_${hash}_oid=$(git rev-parse HEAD) &&\n+\t\teval tree_${hash}_oid=$(git rev-parse HEAD^{tree})\n+\t'\n+\ttest_expect_success \"create a $hash  tagged blob\" '\n+\t\tgit tag --no-sign -m \"This is a tag\" hellotag $(hello_oid $hash) &&\n+\t\teval hellotag_${hash}_oid=$(git rev-parse hellotag)\n+\t'\n+\ttest_expect_success \"create a $hash tagged tree\" '\n+\t\tgit tag --no-sign -m \"This is a tag\" treetag $(tree_oid $hash) &&\n+\t\teval treetag_${hash}_oid=$(git rev-parse treetag)\n+\t'\n+\ttest_expect_success \"create a $hash tagged commit\" '\n+\t\tgit tag --no-sign -m \"This is a tag\" committag $(commit_oid $hash) &&\n+\t\teval committag_${hash}_oid=$(git rev-parse committag)\n+\t'\n+\ttest_expect_success GPG2 \"create a $hash signed commit\" '\n+\t\tgit commit --gpg-sign --allow-empty -m \"This is a signed commit\" &&\n+\t\teval signedcommit_${hash}_oid=$(git rev-parse HEAD)\n+\t'\n+\ttest_expect_success GPG2 \"create a $hash signed tag\" '\n+\t\tgit tag -s -m \"This is a signed tag\" signedtag HEAD &&\n+\t\teval signedtag_${hash}_oid=$(git rev-parse signedtag)\n+\t'\n+\ttest_expect_success \"create a $hash branch\" '\n+\t\tgit checkout -b branch $(commit_oid $hash) &&\n+\t\techo \"More more more give me more!\" > more &&\n+\t\teval more_${hash}_oid=$(git hash-object more) &&\n+\t\techo \"Another and another and another\" > another &&\n+\t\teval another_${hash}_oid=$(git hash-object another) &&\n+\t\tgit update-index --add more another &&\n+\t\tgit commit -m \"Add more files!\" &&\n+\t\teval commit2_${hash}_oid=$(git rev-parse HEAD) &&\n+\t\teval tree2_${hash}_oid=$(git rev-parse HEAD^{tree})\n+\t'\n+\ttest_expect_success GPG2 \"create another $hash signed tag\" '\n+\t\tgit tag -s -m \"This is another signed tag\" signedtag2 $(commit2_oid $hash) &&\n+\t\teval signedtag2_${hash}_oid=$(git rev-parse signedtag2)\n+\t'\n+\ttest_expect_success GPG2 \"merge the $hash branches together\" '\n+\t\tgit merge -S -m \"merge some signed tags together\" signedtag signedtag2 &&\n+\t\teval signedcommit2_${hash}_oid=$(git rev-parse HEAD)\n+\t'\n+\ttest_expect_success GPG2 \"create additional $hash signed commits\" '\n+\t\tgit commit --gpg-sign --allow-empty -m \"This is an additional signed commit\" &&\n+\t\tgit cat-file commit HEAD | del_sigcommit sha256 > \"../${hash}_signedcommit3\" &&\n+\t\tgit cat-file commit HEAD | del_sigcommit sha1 > \"../${hash}_signedcommit4\" &&\n+\t\teval signedcommit3_${hash}_oid=$(git hash-object -t commit -w ../${hash}_signedcommit3) &&\n+\t\teval signedcommit4_${hash}_oid=$(git hash-object -t commit -w ../${hash}_signedcommit4)\n+\t'\n+\ttest_expect_success GPG2 \"create additional $hash signed tags\" '\n+\t\tgit tag -s -m \"This is an additional signed tag\" signedtag34 HEAD &&\n+\t\tgit cat-file tag signedtag34 | del_sigtag \"${hash}\" sha256 > ../${hash}_signedtag3 &&\n+\t\tgit cat-file tag signedtag34 | del_sigtag \"${hash}\" sha1 > ../${hash}_signedtag4 &&\n+\t\teval signedtag3_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag3) &&\n+\t\teval signedtag4_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag4)\n+\t'\n+done\n+cd \"$base\"\n+\n+compare_oids () {\n+    test \"$#\" = 5 && { local PREREQ=$1; shift; } || PREREQ=\n+    local type=\"$1\"\n+    local name=\"$2\"\n+    local sha1_oid=\"$3\"\n+    local sha256_oid=\"$4\"\n+\n+    echo ${sha1_oid} > ${name}_sha1_expected\n+    echo ${sha256_oid} > ${name}_sha256_expected\n+    echo ${type} > ${name}_type_expected\n+\n+    git --git-dir=repo-sha1/.git rev-parse --output-object-format=sha256 ${sha1_oid} > ${name}_sha1_sha256_found\n+    git --git-dir=repo-sha256/.git rev-parse --output-object-format=sha1 ${sha256_oid} > ${name}_sha256_sha1_found\n+    local sha1_sha256_oid=$(cat ${name}_sha1_sha256_found)\n+    local sha256_sha1_oid=$(cat ${name}_sha256_sha1_found)\n+\n+    test_expect_success $PREREQ \"Verify ${type} ${name}'s sha1 oid\" '\n+\tgit --git-dir=repo-sha256/.git rev-parse --output-object-format=sha1 ${sha256_oid} > ${name}_sha1 &&\n+\ttest_cmp ${name}_sha1 ${name}_sha1_expected\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${type} ${name}'s sha256 oid\" '\n+\tgit --git-dir=repo-sha1/.git rev-parse --output-object-format=sha256 ${sha1_oid} > ${name}_sha256 &&\n+\ttest_cmp ${name}_sha256 ${name}_sha256_expected\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha1 type\" '\n+\tgit --git-dir=repo-sha1/.git cat-file -t ${sha1_oid} > ${name}_type1 &&\n+\tgit --git-dir=repo-sha256/.git cat-file -t ${sha256_sha1_oid} > ${name}_type2 &&\n+\ttest_cmp ${name}_type1 ${name}_type2 &&\n+\ttest_cmp ${name}_type1 ${name}_type_expected\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha256 type\" '\n+\tgit --git-dir=repo-sha256/.git cat-file -t ${sha256_oid} > ${name}_type3 &&\n+\tgit --git-dir=repo-sha1/.git cat-file -t ${sha1_sha256_oid} > ${name}_type4 &&\n+\ttest_cmp ${name}_type3 ${name}_type4 &&\n+\ttest_cmp ${name}_type3 ${name}_type_expected\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha1 size\" '\n+\tgit --git-dir=repo-sha1/.git cat-file -s ${sha1_oid} > ${name}_size1 &&\n+\tgit --git-dir=repo-sha256/.git cat-file -s ${sha256_sha1_oid} > ${name}_size2 &&\n+\ttest_cmp ${name}_size1 ${name}_size2\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha256 size\" '\n+\tgit --git-dir=repo-sha256/.git cat-file -s ${sha256_oid} > ${name}_size3 &&\n+\tgit --git-dir=repo-sha1/.git cat-file -s ${sha1_sha256_oid} > ${name}_size4 &&\n+\ttest_cmp ${name}_size3 ${name}_size4\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha1 pretty content\" '\n+\tgit --git-dir=repo-sha1/.git cat-file -p ${sha1_oid} > ${name}_content1 &&\n+\tgit --git-dir=repo-sha256/.git cat-file -p ${sha256_sha1_oid} > ${name}_content2 &&\n+\ttest_cmp ${name}_content1 ${name}_content2\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha256 pretty content\" '\n+\tgit --git-dir=repo-sha256/.git cat-file -p ${sha256_oid} > ${name}_content3 &&\n+\tgit --git-dir=repo-sha1/.git cat-file -p ${sha1_sha256_oid} > ${name}_content4 &&\n+\ttest_cmp ${name}_content3 ${name}_content4\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha1 content\" '\n+\tgit --git-dir=repo-sha1/.git cat-file ${type} ${sha1_oid} > ${name}_content5 &&\n+\tgit --git-dir=repo-sha256/.git cat-file ${type} ${sha256_sha1_oid} > ${name}_content6 &&\n+\ttest_cmp ${name}_content5 ${name}_content6\n+'\n+\n+    test_expect_success $PREREQ \"Verify ${name}'s sha256 content\" '\n+\tgit --git-dir=repo-sha256/.git cat-file ${type} ${sha256_oid} > ${name}_content7 &&\n+\tgit --git-dir=repo-sha1/.git cat-file ${type} ${sha1_sha256_oid} > ${name}_content8 &&\n+\ttest_cmp ${name}_content7 ${name}_content8\n+'\n+\n+}\n+\n+compare_oids 'blob' hello \"$hello_sha1_oid\" \"$hello_sha256_oid\"\n+compare_oids 'tree' tree \"$tree_sha1_oid\" \"$tree_sha256_oid\"\n+compare_oids 'commit' commit \"$commit_sha1_oid\" \"$commit_sha256_oid\"\n+compare_oids GPG2 'commit' signedcommit \"$signedcommit_sha1_oid\" \"$signedcommit_sha256_oid\"\n+compare_oids 'tag' hellotag \"$hellotag_sha1_oid\" \"$hellotag_sha256_oid\"\n+compare_oids 'tag' treetag \"$treetag_sha1_oid\" \"$treetag_sha256_oid\"\n+compare_oids 'tag' committag \"$committag_sha1_oid\" \"$committag_sha256_oid\"\n+compare_oids GPG2 'tag' signedtag \"$signedtag_sha1_oid\" \"$signedtag_sha256_oid\"\n+\n+compare_oids 'blob' more \"$more_sha1_oid\" \"$more_sha256_oid\"\n+compare_oids 'blob' another \"$another_sha1_oid\" \"$another_sha256_oid\"\n+compare_oids 'tree' tree2 \"$tree2_sha1_oid\" \"$tree2_sha256_oid\"\n+compare_oids 'commit' commit2 \"$commit2_sha1_oid\" \"$commit2_sha256_oid\"\n+compare_oids GPG2 'tag' signedtag2 \"$signedtag2_sha1_oid\" \"$signedtag2_sha256_oid\"\n+compare_oids GPG2 'commit' signedcommit2 \"$signedcommit2_sha1_oid\" \"$signedcommit2_sha256_oid\"\n+compare_oids GPG2 'commit' signedcommit3 \"$signedcommit3_sha1_oid\" \"$signedcommit3_sha256_oid\"\n+compare_oids GPG2 'commit' signedcommit4 \"$signedcommit4_sha1_oid\" \"$signedcommit4_sha256_oid\"\n+compare_oids GPG2 'tag' signedtag3 \"$signedtag3_sha1_oid\" \"$signedtag3_sha256_oid\"\n+compare_oids GPG2 'tag' signedtag4 \"$signedtag4_sha1_oid\" \"$signedtag4_sha256_oid\"\n+\n+test_done\ndiff --git a/t/t1016/gpg b/t/t1016/gpg\nnew file mode 100755\nindex 000000000000..2601cb18a5b3\n--- /dev/null\n+++ b/t/t1016/gpg\n@@ -0,0 +1,2 @@\n+#!/bin/sh\n+exec gpg --faked-system-time \"20230918T154812\" \"$@\"\n-- \n2.41.0\n\n"},{"id":"488192","messageId":"xmqqv86z5359.fsf@gitster.g","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"Re: [PATCH v2 00/30] initial support for multiple hash functions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2024-02-07T22:18:58Z","receivedAt":"2024-02-07T22:19:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> This addresses all of the known test failures from v1 of this set of\n> changes.  In particular I have reworked commit_tree_extended which\n> was flagged by smatch, -Werror=array-bounds, and the leak detector.\n>\n> One functional bug was fixed in repo_for_each_abbrev where it was\n> mistakenly displaying too many ambiguous oids.\n>\n> I am posting this so that people review and testing of this patchset\n> won't be distracted by the known and fixed issues.\n\nWe haven't seen any reviews on this second round, and have had it\noutside 'next' for too long.  I am tempted to say that we merge it\nto 'next' and see if anybody screams at this point.\n\nThanks.\n"},{"id":"488198","messageId":"owlymssbn6qa.fsf@fine.c.googlers.com","threadId":"60273","inReplyTo":"xmqqv86z5359.fsf@gitster.g","subject":"Re: [PATCH v2 00/30] initial support for multiple hash functions","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2024-02-08T00:24:13Z","receivedAt":"2024-02-08T00:24:16Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n>\n>> This addresses all of the known test failures from v1 of this set of\n>> changes.  In particular I have reworked commit_tree_extended which\n>> was flagged by smatch, -Werror=array-bounds, and the leak detector.\n>>\n>> One functional bug was fixed in repo_for_each_abbrev where it was\n>> mistakenly displaying too many ambiguous oids.\n>>\n>> I am posting this so that people review and testing of this patchset\n>> won't be distracted by the known and fixed issues.\n>\n> We haven't seen any reviews on this second round, and have had it\n> outside 'next' for too long.  I am tempted to say that we merge it\n> to 'next' and see if anybody screams at this point.\n\nFWIW out of all the \"Needs review\" topics this one seemed like the most\ndeserving of another pair of eyes, and I was planning to review some of\nthe patches here this week + the weekend. If my review takes too long\n(taking longer than this weekend) I can give another update next week\nsaying \"too hard for me, please don't wait for me\" to unblock you from\nmerging to next.\n\nThanks.\n"},{"id":"488210","messageId":"ZcRwlUdKbM6bEVmr@tanuki","threadId":"60273","inReplyTo":"owlymssbn6qa.fsf@fine.c.googlers.com","subject":"Re: [PATCH v2 00/30] initial support for multiple hash functions","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-02-08T06:11:33Z","receivedAt":"2024-02-08T06:11:40Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Feb 07, 2024 at 04:24:13PM -0800, Linus Arver wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n> > \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n> >\n> >> This addresses all of the known test failures from v1 of this set of\n> >> changes.  In particular I have reworked commit_tree_extended which\n> >> was flagged by smatch, -Werror=array-bounds, and the leak detector.\n> >>\n> >> One functional bug was fixed in repo_for_each_abbrev where it was\n> >> mistakenly displaying too many ambiguous oids.\n> >>\n> >> I am posting this so that people review and testing of this patchset\n> >> won't be distracted by the known and fixed issues.\n> >\n> > We haven't seen any reviews on this second round, and have had it\n> > outside 'next' for too long.  I am tempted to say that we merge it\n> > to 'next' and see if anybody screams at this point.\n> \n> FWIW out of all the \"Needs review\" topics this one seemed like the most\n> deserving of another pair of eyes, and I was planning to review some of\n> the patches here this week + the weekend. If my review takes too long\n> (taking longer than this weekend) I can give another update next week\n> saying \"too hard for me, please don't wait for me\" to unblock you from\n> merging to next.\n\nI completely lost track of this patch series. So same for me: I don't\nwant to hold it up, but would be happy to give it a pair of eyes next\nweek.\n\nPatrick\n"},{"id":"488218","messageId":"owlyfry3cqkp.fsf@fine.c.googlers.com","threadId":"60273","inReplyTo":"20231002024034.2611-1-ebiederm@gmail.com","subject":"Re: [PATCH v2 01/30] object-file-convert: stubs for converting from one object format to another","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2024-02-08T08:23:18Z","receivedAt":"2024-02-08T08:23:20Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>\n> Two basic functions are provided:\n> - convert_object_file Takes an object file it's type and hash algorithm\n>   and converts it into the equivalent object file that would\n>   have been generated with hash algorithm \"to\".\n\nShould probably be\n\n    - convert_object_file takes an object file, its type, and hash algorithm\n      and converts it into the equivalent object file using hash\n      algorithm \"to\".\n\nIt would be nice if you gave the function name some clue of what sort of\nconversion is being done though. Something like\n\"convert_object_file_hash\" or \"switch_object_file_hash\". I like \"switch\"\nbecause the word \"hash\" itself means to convert some input bytes into\nanother set of (aka \"hashed\") bytes, and the indirect shadowing of\n\"convert\" in this way can be avoided here by using a different word.The\nwhole point of this function is to switch the hashing scheme from one to\nanother, while still keeping everything else the same, so \"switch\" seems\nmore appropriate.\n\nUnless, of course, we already pervasively use \"convert\" this way\nelsewhere in the codebase (I have not checked).\n\n>   For blob objects there is no conversation to be done and it is an\n\ns/conversation/conversion\n\n>   error to use this function on them.\n\nCould you explain why no conversion is needed for blob objects, and also\nwhy it should be an error (and not just a NOP)?\n\nIn the code we can also call BUG() if the from/to algos are the same.\nIt's probably worth mentioning in here as well?\n\nAlso for such detailed explanations, I think it's much better\nto place them as comments directly above the function (and only mention\nthe important bits about these helper functions, other than the fact\nthat they will come in handy in later patches, in the log message).\n\n>   For commit, tree, and tag objects embedded oids are replaced by the\n>   oids of the objects they refer to with those objects and their\n>   object ids reencoded in with the hash algorithm \"to\".\n\nThat's a little wordy. I assume embedded oids just mean oids that these\nobjects refer to (e.g., commit objects have oids, but for example the\ntree referred by a commit would be an example of an embedded oid). If\nso, then how about just\n\n    For commit, tree, and tag objects both their oids and embedded\n    (dependent) oids are converted using hash algorithm \"to\".\n\nI dropped \"reencoded\" in favor of \"converted\" btecause that's the verb\nyou use in your function name \"convert_object_file()\".\n\n>   Signatures\n>   are rearranged so that they remain valid after the object has\n>   been reencoded.\n\nMaybe s/reencoded/converted here as well? But also, it sounds odd to me\nthat signatures are simply \"rearranged\" and not \"regenerated\" or\n\"recreated\" becaues \"rearranged\" means keeping most of the old stuff\naround but just repositioning them, which doesn't sound like it's doing\njustice to the meaning of a hash algo transition.\n\n> - repo_oid_to_algop which takes an oid that refers to an object file\n>   and returns the oid of the equivalent object file generated\n>   with the target hash algorithm.\n\nThe name is odd to me because \"repo_oid\" doesn't make sense (a repo\nis not an object so it doesn't have an oid), but also the \"to_algop\"\nname (AFAICS \"algop\" just means \"pointer to algorithm\" in the codebase,\nfor examle in <hash-ll.h>).\n\nHere's a possible rewording:\n\n    - switch_oid_hash takes an object file's oid and returns a new one\n      using hash algorithm \"to\".\n\n> The pair of files object-file-convert.c and object-file-convert.h are\n> introduced to hold as much of this logic as possible to keep this\n> conversion logic cleanly separated from everything else and in the\n> hopes that someday the code will be clean enough git can support\n\nDid you mean \"clean enough so that Git ...\"?\n\n> compiling out support for sha1 and the various conversion functions.\n\nFYI you can cut down this ambitious sentence into two for readability,\nlike this:\n\n    The new files object-file-convert.{c,h} hold as much of this logic\n    as possible to keep this conversion logic cleanly separated from\n    everything else. This separation will help us easily phase out SHA1\n    support (perhaps as a compile-time flag) in the future.\n\n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n>  Makefile              |  1 +\n>  object-file-convert.c | 57 +++++++++++++++++++++++++++++++++++++++++++\n>  object-file-convert.h | 24 ++++++++++++++++++\n>  3 files changed, 82 insertions(+)\n>  create mode 100644 object-file-convert.c\n>  create mode 100644 object-file-convert.h\n>\n> diff --git a/Makefile b/Makefile\n> index 577630936535..f7e824f25cda 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -1073,6 +1073,7 @@ LIB_OBJS += notes-cache.o\n>  LIB_OBJS += notes-merge.o\n>  LIB_OBJS += notes-utils.o\n>  LIB_OBJS += notes.o\n> +LIB_OBJS += object-file-convert.o\n>  LIB_OBJS += object-file.o\n>  LIB_OBJS += object-name.o\n>  LIB_OBJS += object.o\n> diff --git a/object-file-convert.c b/object-file-convert.c\n> new file mode 100644\n> index 000000000000..4777aba83636\n> --- /dev/null\n> +++ b/object-file-convert.c\n> @@ -0,0 +1,57 @@\n> +#include \"git-compat-util.h\"\n> +#include \"gettext.h\"\n> +#include \"strbuf.h\"\n> +#include \"repository.h\"\n> +#include \"hash-ll.h\"\n> +#include \"object.h\"\n> +#include \"object-file-convert.h\"\n> +\n> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n> +\t\t      const struct git_hash_algo *to, struct object_id *dest)\n> +{\n> +\t/*\n> +\t * If the source algorithm is not set, then we're using the\n> +\t * default hash algorithm for that object.\n> +\t */\n> +\tconst struct git_hash_algo *from =\n> +\t\tsrc->algo ? &hash_algos[src->algo] : repo->hash_algo;\n\nHmm, if we're only using \"repo\" to grab its hash_algo member as a\nfallback, should this function be broken down into 2, one to get the\nfallback algo and another that has the meat of this function without the\n\"repo\" parameter?\n\nI haven't read the rest of the series yet, let me keep reading.\n\n> +\n> +\tif (from == to) {\n> +\t\tif (src != dest)\n> +\t\t\toidcpy(dest, src);\n> +\t\treturn 0;\n> +\t}\n> +\treturn -1;\n> +}\n\nIt's curious to me that you treat same-hash-algo as a NOP here, but as a\nBUG() in the other helper. Why the difference (perhaps add a comment)?\n\n> +int convert_object_file(struct strbuf *outbuf,\n> +\t\t\tconst struct git_hash_algo *from,\n> +\t\t\tconst struct git_hash_algo *to,\n> +\t\t\tconst void *buf, size_t len,\n> +\t\t\tenum object_type type,\n> +\t\t\tint gentle)\n\nIs there a particular reason you chose to put the out parameter (outbuf)\nat the beginning, rather than at the end? In the other helper function\nthe \"dest\" out param comes at the end. Style conflict...?\n\nAlso, you don't use buf or len, so you could have kept them out from\nthis patch (so that you only add them later as needed).\n\n> +{\n> +\tint ret;\n> +\n> +\t/* Don't call this function when no conversion is necessary */\n\nPlease avoid double-negation. How about simply\n\n    /* Refuse nonsensical conversion */\n\nor simply drop the comment (as the BUG() description already serves the\nsame purpose?)\n\n> +\tif ((from == to) || (type == OBJ_BLOB))\n> +\t\tBUG(\"Refusing noop object file conversion\");\n\nI think it would be better if you separated these out and gave them\ndifferent BUG() messages. If we barf with a BUG() it would be so much\nmore helpful if the message we get is as specific as possible, rather\nthan leaving us guessing whether (as in this case) we passed in\nidentical hash algos or whether we tried to handle a blob object. Since\nyou already have the switch statement below, the OBJ_BLOB case could go\nthere easily enough.\n\nNit: I think BUG() messages are not supposed to be capitalized.\n\n> +\tswitch (type) {\n> +\tcase OBJ_COMMIT:\n> +\tcase OBJ_TREE:\n> +\tcase OBJ_TAG:\n> +\tdefault:\n> +\t\t/* Not implemented yet, so fail. */\n> +\t\tret = -1;\n> +\t\tbreak;\n> +\t}\n> +\tif (!ret)\n> +\t\treturn 0;\n> +\tif (gentle) {\n> +\t\tstrbuf_release(outbuf);\n\nWhat does \"gentle\" mean here? But also if you are just freeing the\noutbuf, then why not just name this param \"free_outbuf\"? But also it\nseems a bit odd that the caller (who presumably owns the outbuf) is\nasking a conversion function to possibly free it before returning.\n\n> +\t\treturn ret;\n> +\t}\n> +\tdie(_(\"Failed to convert object from %s to %s\"),\n> +\t\tfrom->name, to->name);\n> +}\n> diff --git a/object-file-convert.h b/object-file-convert.h\n> new file mode 100644\n> index 000000000000..a4f802aa8eea\n> --- /dev/null\n> +++ b/object-file-convert.h\n> @@ -0,0 +1,24 @@\n> +#ifndef OBJECT_CONVERT_H\n> +#define OBJECT_CONVERT_H\n> +\n> +struct repository;\n> +struct object_id;\n> +struct git_hash_algo;\n> +struct strbuf;\n> +#include \"object.h\"\n\nIt looks a bit odd that the forward declarations of these structs come\nbefore the #include line. I don't think that's the pattern in our\ncodebase (maybe I'm wrong?).\n\nIn hindsight, the log message could have added that switching back\nand forth between different hash algos is a fundamental (but currently\nmissing) operation, and that this is the reason why these conversion\nfunctions (currently unfinished) are necessary as the first step for the\npatch series. I realize that you've stated as much in your cover letter,\nbut it is always nice to have the intent embedded in the log message(s)\nwhere applicable (such as the case in this patch that introduces a brand\nnew header file) to save future developers the hassle of looking up the\nrelevant cover letter.\n\nThanks.\n\n> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n> +\t\t      const struct git_hash_algo *to, struct object_id *dest);\n> +\n> +/*\n> + * Convert an object file from one hash algorithm to another algorithm.\n> + * Return -1 on failure, 0 on success.\n> + */\n> +int convert_object_file(struct strbuf *outbuf,\n> +\t\t\tconst struct git_hash_algo *from,\n> +\t\t\tconst struct git_hash_algo *to,\n> +\t\t\tconst void *buf, size_t len,\n> +\t\t\tenum object_type type,\n> +\t\t\tint gentle);\n> +\n> +#endif /* OBJECT_CONVERT_H */\n> -- \n> 2.41.0\n"},{"id":"488235","messageId":"02a018e7-1c6f-4fa0-8a19-0c1fdc202eed@gmail.com","threadId":"60273","inReplyTo":"20231002024034.2611-22-ebiederm@gmail.com","subject":"Re: [PATCH v2 22/30] rev-parse: add an --output-object-format parameter","fromName":"Jean-Noël Avila","fromEmail":"avila.jn@gmail.com","sentAt":"2024-02-08T16:25:40Z","receivedAt":"2024-02-08T16:25:43Z","isPatch":true,"sender":{"key":"jn.avila@free.fr","avatar":"https://avatars.githubusercontent.com/u/156172?v=4"},"body":"Le 02/10/2023 à 04:40, Eric W. Biederman a écrit :\n> +\t\t\t\telse die(_(\"unsupported object format: %s\"), arg);\n\n\"unsupported object format '%s'\" already exist in translated strings.\n\nOne less similar string to translate.\n\n\nThank you.\n\n> +\t\t\t}\n>   \t\t\tif (opt_with_value(arg, \"--short\", &arg)) {\n>   \t\t\t\tfilter &= ~(DO_FLAGS|DO_NOREV);\n\n\n\n"},{"id":"488506","messageId":"owlya5o4dbj1.fsf@fine.c.googlers.com","threadId":"60273","inReplyTo":"20231002024034.2611-2-ebiederm@gmail.com","subject":"Re: [PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2024-02-13T08:16:34Z","receivedAt":"2024-02-13T08:16:37Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>\n> While looking at how to handle input of both SHA-1 and SHA-256 oids in\n> get_oid_with_context, I realized that the oid_array in\n> repo_for_each_abbrev might have more than one kind of oid stored in it\n> simultaneously.\n>\n> Update to oid_array_append to ensure that oids added to an oid array\n\ns/Update to/Update\n\n> always have an algorithm set.\n>\n> Update void_hashcmp to first verify two oids use the same hash algorithm\n> before comparing them to each other.\n>\n> With that oid-array should be safe to use with different kinds of\n\ns/oid-array/oid_array\n\n> oids simultaneously.\n>\n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n>  oid-array.c | 12 ++++++++++--\n>  1 file changed, 10 insertions(+), 2 deletions(-)\n>\n> diff --git a/oid-array.c b/oid-array.c\n> index 8e4717746c31..1f36651754ed 100644\n> --- a/oid-array.c\n> +++ b/oid-array.c\n> @@ -6,12 +6,20 @@ void oid_array_append(struct oid_array *array, const struct object_id *oid)\n>  {\n>  \tALLOC_GROW(array->oid, array->nr + 1, array->alloc);\n>  \toidcpy(&array->oid[array->nr++], oid);\n> +\tif (!oid->algo)\n> +\t\toid_set_algo(&array->oid[array->nr - 1], the_hash_algo);\n\nHow come we can't set oid->algo _before_ we call oidcpy()? It seems odd\nthat we do the copy first and then modify what we just copied after the\nfact, instead of making sure that the thing we want to copy is correct\nbefore doing the copy.\n\nBut also, if we are going to make the oid object \"correct\" before\ninvoking oidcpy(), we might as well do it when the oid is first\ncreated/used (in the caller(s) of this function). I don't demand that\nyou find/demonstrate where all these places are in this series (maybe\nthat's a hairy problem to tackle?), but it seems cleaner in principle to\nfix the creation of oid objects instead of having to make oid users\nclean up their act like this after using them.\n\n>  \tarray->sorted = 0;\n>  }\n>  \n> -static int void_hashcmp(const void *a, const void *b)\n> +static int void_hashcmp(const void *va, const void *vb)\n>  {\n> -\treturn oidcmp(a, b);\n> +\tconst struct object_id *a = va, *b = vb;\n> +\tint ret;\n> +\tif (a->algo == b->algo)\n> +\t\tret = oidcmp(a, b);\n\nThis makes sense (per the commit message description) ...\n\n> +\telse\n> +\t\tret = a->algo > b->algo ? 1 : -1;\n\n... but this seems to go against it? I thought you wanted to only ever\ncompare hashes if they were of the same algo? It would be good to add a\ncomment explaining why this is OK (we are no longer doing a byte-by-byte\ncomparison of these oids any more here like we do for oidcmp() above\nwhich boils down to calling memcmp()).\n\n> +\treturn ret;\n\nAlso, in terms of style I think the \"early return for errors\" style\nwould be simpler to read. I.e.\n\n    if (a->algo > b->algo)\n        return 1;\n\n    if (a->algo < b->algo)\n        return -1;\n\n    return oidcmd(a, b);\n\n>  }\n>  \n>  void oid_array_sort(struct oid_array *array)\n> -- \n> 2.41.0\n"},{"id":"488507","messageId":"023394e2-5f64-4a59-af96-b77dafb20051@app.fastmail.com","threadId":"60273","inReplyTo":"20231002024034.2611-2-ebiederm@gmail.com","subject":"Re: [PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids","fromName":"Kristoffer Haugsbakk","fromEmail":"code@khaugsbakk.name","sentAt":"2024-02-13T08:31:22Z","receivedAt":"2024-02-13T08:33:29Z","isPatch":true,"sender":{"key":"code@khaugsbakk.name","avatar":"https://avatars.githubusercontent.com/u/2229597?v=4"},"body":"On Mon, Oct 2, 2023, at 04:40, Eric W. Biederman wrote:\n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n\nMost of your patches have this sign-off line with your name quoted.\n\n-- \nKristoffer Haugsbakk\n\n\n"},{"id":"488522","messageId":"owly7cj8d7y6.fsf@fine.c.googlers.com","threadId":"60273","inReplyTo":"20231002024034.2611-3-ebiederm@gmail.com","subject":"Re: [PATCH v2 03/30] object-names: support input of oids in any supported hash","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2024-02-13T09:33:53Z","receivedAt":"2024-02-13T09:33:56Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>\n> Support short oids encoded in any algorithm, while ensuring enough of\n> the oid is specified to disambiguate between all of the oids in the\n> repository encoded in any algorithm.\n>\n> By default have the code continue to only accept oids specified in the\n> storage hash algorithm of the repository, but when something is\n> ambiguous display all of the possible oids from any accepted oid\n> encoding.\n>\n> A new flag is added GET_OID_HASH_ANY that when supplied causes the\n> code to accept oids specified in any hash algorithm, and to return the\n> oids that were resolved.\n>\n> This implements the functionality that allows both SHA-1 and SHA-256\n> object names, from the \"Object names on the command line\" section of\n> the hash function transition document.\n>\n> Care is taken in get_short_oid so that when the result is ambiguous\n> the output remains the same if GIT_OID_HASH_ANY was not supplied.  If\n> GET_OID_HASH_ANY was supplied objects of any hash algorithm that match\n> the prefix are displayed.\n>\n> This required updating repo_for_each_abbrev to give it a parameter so\n> that it knows to look at all hash algorithms.\n>\n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n>  builtin/rev-parse.c |  2 +-\n>  hash-ll.h           |  1 +\n>  object-name.c       | 46 ++++++++++++++++++++++++++++++++++-----------\n>  object-name.h       |  3 ++-\n>  4 files changed, 39 insertions(+), 13 deletions(-)\n>\n> diff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\n> index fde8861ca4e0..43e96765400c 100644\n> --- a/builtin/rev-parse.c\n> +++ b/builtin/rev-parse.c\n> @@ -882,7 +882,7 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n>  \t\t\t\tcontinue;\n>  \t\t\t}\n>  \t\t\tif (skip_prefix(arg, \"--disambiguate=\", &arg)) {\n> -\t\t\t\trepo_for_each_abbrev(the_repository, arg,\n> +\t\t\t\trepo_for_each_abbrev(the_repository, arg, the_hash_algo,\n>  \t\t\t\t\t\t     show_abbrev, NULL);\n>  \t\t\t\tcontinue;\n>  \t\t\t}\n> diff --git a/hash-ll.h b/hash-ll.h\n> index 10d84cc20888..2cfde63ae1cf 100644\n> --- a/hash-ll.h\n> +++ b/hash-ll.h\n> @@ -145,6 +145,7 @@ struct object_id {\n>  #define GET_OID_RECORD_PATH     0200\n>  #define GET_OID_ONLY_TO_DIE    04000\n>  #define GET_OID_REQUIRE_PATH  010000\n> +#define GET_OID_HASH_ANY      020000\n\nSo far the \"GET_OID_*\" flags sound like the \"GET_OID\" is the prefix and\nthe remaining text describes the behavior. So in this sense, the\nbehavior here would be \"HASH_ANY\" and I think \"ANY_HASH\" is better. But\nalso, going by the description of this flag in the commit message, I\nthink \"ACCEPT_ANY_HASH\" is still better.\n>\n>  #define GET_OID_DISAMBIGUATORS \\\n>  \t(GET_OID_COMMIT | GET_OID_COMMITTISH | \\\n> diff --git a/object-name.c b/object-name.c\n> index 0bfa29dbbfe9..7dd6e5e47566 100644\n> --- a/object-name.c\n> +++ b/object-name.c\n> @@ -25,6 +25,7 @@\n>  #include \"midx.h\"\n>  #include \"commit-reach.h\"\n>  #include \"date.h\"\n> +#include \"object-file-convert.h\"\n>\n>  static int get_oid_oneline(struct repository *r, const char *, struct object_id *, struct commit_list *);\n>\n> @@ -49,6 +50,7 @@ struct disambiguate_state {\n>\n>  static void update_candidates(struct disambiguate_state *ds, const struct object_id *current)\n>  {\n> +\t/* The hash algorithm of current has already been filtered */\n\nIs this new comment describing existing behavior before this patch or\nafter?\n\n>  \tif (ds->always_call_fn) {\n>  \t\tds->ambiguous = ds->fn(ds->repo, current, ds->cb_data) ? 1 : 0;\n>  \t\treturn;\n> @@ -134,6 +136,8 @@ static void unique_in_midx(struct multi_pack_index *m,\n>  {\n>  \tuint32_t num, i, first = 0;\n>  \tconst struct object_id *current = NULL;\n> +\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n> +\t\tds->repo->hash_algo->hexsz : ds->len;\n\nNit: So this is just trying to use the shorter length between ds->len\nand hexsz. Would it be good to encode this information into the variable\nname, such as \"len_short\" or similar? And if so, flipping the \">\" to\n\"<\" would be more natural, like\n\n    int shorter_len = ds->len < ds->repo->hash_algo->hexsz ?\n                      ds->len : ds->repo->hash_algo->hexsz;\n\nbecause it reads \"if ds->len is shorter, use it; otherwise use hexsz\".\n\n>  \tnum = m->num_objects;\n>\n>  \tif (!num)\n> @@ -149,7 +153,7 @@ static void unique_in_midx(struct multi_pack_index *m,\n>  \tfor (i = first; i < num && !ds->ambiguous; i++) {\n>  \t\tstruct object_id oid;\n>  \t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n> -\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, current->hash))\n> +\t\tif (!match_hash(len, ds->bin_pfx.hash, current->hash))\n>  \t\t\tbreak;\n>  \t\tupdate_candidates(ds, current);\n>  \t}\n> @@ -159,6 +163,8 @@ static void unique_in_pack(struct packed_git *p,\n>  \t\t\t   struct disambiguate_state *ds)\n>  {\n>  \tuint32_t num, i, first = 0;\n> +\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n> +\t\tds->repo->hash_algo->hexsz : ds->len;\n\nDitto.\n\n>\n>  \tif (p->multi_pack_index)\n>  \t\treturn;\n> @@ -177,7 +183,7 @@ static void unique_in_pack(struct packed_git *p,\n>  \tfor (i = first; i < num && !ds->ambiguous; i++) {\n>  \t\tstruct object_id oid;\n>  \t\tnth_packed_object_id(&oid, p, i);\n> -\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, oid.hash))\n> +\t\tif (!match_hash(len, ds->bin_pfx.hash, oid.hash))\n>  \t\t\tbreak;\n>  \t\tupdate_candidates(ds, &oid);\n>  \t}\n> @@ -188,6 +194,10 @@ static void find_short_packed_object(struct disambiguate_state *ds)\n>  \tstruct multi_pack_index *m;\n>  \tstruct packed_git *p;\n>\n> +\t/* Skip, unless oids from the storage hash algorithm are wanted */\n\nPerhaps a simpler phrasing would be\n\n    /* Only accept oids specified in the storage hash algorithm of the repository. */\n\nwhich is closer to the wording of the commit message?\n\n> +\tif (ds->bin_pfx.algo && (&hash_algos[ds->bin_pfx.algo] != ds->repo->hash_algo))\n> +\t\treturn;\n> +\n>  \tfor (m = get_multi_pack_index(ds->repo); m && !ds->ambiguous;\n>  \t     m = m->next)\n>  \t\tunique_in_midx(m, ds);\n> @@ -326,11 +336,12 @@ int set_disambiguate_hint_config(const char *var, const char *value)\n>\n>  static int init_object_disambiguation(struct repository *r,\n>  \t\t\t\t      const char *name, int len,\n> +\t\t\t\t      const struct git_hash_algo *algo,\n>  \t\t\t\t      struct disambiguate_state *ds)\n>  {\n>  \tint i;\n>\n> -\tif (len < MINIMUM_ABBREV || len > the_hash_algo->hexsz)\n> +\tif (len < MINIMUM_ABBREV || len > GIT_MAX_HEXSZ)\n>  \t\treturn -1;\n>\n>  \tmemset(ds, 0, sizeof(*ds));\n> @@ -357,6 +368,7 @@ static int init_object_disambiguation(struct repository *r,\n>  \tds->len = len;\n>  \tds->hex_pfx[len] = '\\0';\n>  \tds->repo = r;\n> +\tds->bin_pfx.algo = algo ? hash_algo_by_ptr(algo) : GIT_HASH_UNKNOWN;\n>  \tprepare_alt_odb(r);\n>  \treturn 0;\n>  }\n> @@ -491,9 +503,10 @@ static int repo_collect_ambiguous(struct repository *r UNUSED,\n>  \treturn collect_ambiguous(oid, data);\n>  }\n>\n> -static int sort_ambiguous(const void *a, const void *b, void *ctx)\n> +static int sort_ambiguous(const void *va, const void *vb, void *ctx)\n>  {\n>  \tstruct repository *sort_ambiguous_repo = ctx;\n> +\tconst struct object_id *a = va, *b = vb;\n>  \tint a_type = oid_object_info(sort_ambiguous_repo, a, NULL);\n>  \tint b_type = oid_object_info(sort_ambiguous_repo, b, NULL);\n>  \tint a_type_sort;\n> @@ -503,8 +516,12 @@ static int sort_ambiguous(const void *a, const void *b, void *ctx)\n>  \t * Sorts by hash within the same object type, just as\n>  \t * oid_array_for_each_unique() would do.\n>  \t */\n> -\tif (a_type == b_type)\n> -\t\treturn oidcmp(a, b);\n> +\tif (a_type == b_type) {\n> +\t\tif (a->algo == b->algo)\n> +\t\t\treturn oidcmp(a, b);\n> +\t\telse\n> +\t\t\treturn a->algo > b->algo ? 1 : -1;\n> +\t}\n\nThis is duplicated from the previous patch (void_hashcmp). Do we want to\navoid a performance penalty from calling out to a common function?\n\n>  \t/*\n>  \t * Between object types show tags, then commits, and finally\n> @@ -533,8 +550,12 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>  \tint status;\n>  \tstruct disambiguate_state ds;\n>  \tint quietly = !!(flags & GET_OID_QUIETLY);\n> +\tconst struct git_hash_algo *algo = r->hash_algo;\n\nI see some existing uses of the name \"algop\" (presumably to mean\npointer-to-algorithm) and wonder if we should do the same here, for\nconsistency.\n\n> +\n> +\tif (flags & GET_OID_HASH_ANY)\n> +\t\talgo = NULL;\n\nIf we look at the change to init_object_disambiguation() above, it looks\nlike the \"algo\" here is only used to find a GIT_HASH_* constant. So it\nseems a bit roundabout to use a NULL here, then inside\ninit_object_disambiguation() do\n\n    ds->bin_pfx.algo = algo ? hash_algo_by_ptr(algo) : GIT_HASH_UNKNOWN;\n\nto convert the NULL to GIT_HASH_UNKNOWN when we could probably just pass\nin the GIT_HASH_* constant directly from the caller.\n\n> -\tif (init_object_disambiguation(r, name, len, &ds) < 0)\n> +\tif (init_object_disambiguation(r, name, len, algo, &ds) < 0)\n>  \t\treturn -1;\n>\n>  \tif (HAS_MULTI_BITS(flags & GET_OID_DISAMBIGUATORS))\n> @@ -588,7 +609,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>  \t\tif (!ds.ambiguous)\n>  \t\t\tds.fn = NULL;\n>\n> -\t\trepo_for_each_abbrev(r, ds.hex_pfx, collect_ambiguous, &collect);\n> +\t\trepo_for_each_abbrev(r, ds.hex_pfx, algo, collect_ambiguous, &collect);\n>  \t\tsort_ambiguous_oid_array(r, &collect);\n>\n>  \t\tif (oid_array_for_each(&collect, show_ambiguous_object, &out))\n> @@ -610,13 +631,14 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>  }\n>\n>  int repo_for_each_abbrev(struct repository *r, const char *prefix,\n> +\t\t\t const struct git_hash_algo *algo,\n>  \t\t\t each_abbrev_fn fn, void *cb_data)\n>  {\n>  \tstruct oid_array collect = OID_ARRAY_INIT;\n>  \tstruct disambiguate_state ds;\n>  \tint ret;\n>\n> -\tif (init_object_disambiguation(r, prefix, strlen(prefix), &ds) < 0)\n> +\tif (init_object_disambiguation(r, prefix, strlen(prefix), algo, &ds) < 0)\n>  \t\treturn -1;\n>\n>  \tds.always_call_fn = 1;\n> @@ -787,10 +809,12 @@ void strbuf_add_unique_abbrev(struct strbuf *sb, const struct object_id *oid,\n>  int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n>  \t\t\t      const struct object_id *oid, int len)\n>  {\n> +\tconst struct git_hash_algo *algo =\n> +\t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n\nIt would be nice to have symmetry with the style in get_short_oid() and\ninstead do\n\n   const struct git_hash_algo *algo = r->hash_algo;\n\n   if (oid->algo)\n       algo = &hash_algos[oid->algo];\n\n>  \tstruct disambiguate_state ds;\n>  \tstruct min_abbrev_data mad;\n>  \tstruct object_id oid_ret;\n> -\tconst unsigned hexsz = r->hash_algo->hexsz;\n> +\tconst unsigned hexsz = algo->hexsz;\n>\n>  \tif (len < 0) {\n>  \t\tunsigned long count = repo_approximate_object_count(r);\n> @@ -826,7 +850,7 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n>\n>  \tfind_abbrev_len_packed(&mad);\n>\n> -\tif (init_object_disambiguation(r, hex, mad.cur_len, &ds) < 0)\n> +\tif (init_object_disambiguation(r, hex, mad.cur_len, algo, &ds) < 0)\n>  \t\treturn -1;\n>\n>  \tds.fn = repo_extend_abbrev_len;\n> diff --git a/object-name.h b/object-name.h\n> index 9ae522307148..064ddc97d1fe 100644\n> --- a/object-name.h\n> +++ b/object-name.h\n> @@ -67,7 +67,8 @@ enum get_oid_result get_oid_with_context(struct repository *repo, const char *st\n>\n>\n>  typedef int each_abbrev_fn(const struct object_id *oid, void *);\n> -int repo_for_each_abbrev(struct repository *r, const char *prefix, each_abbrev_fn, void *);\n> +int repo_for_each_abbrev(struct repository *r, const char *prefix,\n> +\t\t\t const struct git_hash_algo *algo, each_abbrev_fn, void *);\n>\n>  int set_disambiguate_hint_config(const char *var, const char *value);\n>\n> --\n> 2.41.0\n"},{"id":"488524","messageId":"owly4jecd6n6.fsf@fine.c.googlers.com","threadId":"60273","inReplyTo":"20231002024034.2611-4-ebiederm@gmail.com","subject":"Re: [PATCH v2 04/30] repository: add a compatibility hash algorithm","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2024-02-13T10:02:05Z","receivedAt":"2024-02-13T10:02:08Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>\n> We currently have support for using a full stage 4 SHA-256\n> implementation.  However, we'd like to support interoperability with\n> SHA-1 repositories as well.  The transition plan anticipates a\n> compatibility hash algorithm configuration option that we can use to\n> implement support for this.\n\nPerhaps add\n\n    See section \"Object names on the command line\" in\n    git/Documentation/technical/hash-function-transition.txt .\n\n? That section does not use the language \"compatibility hash algorithm\"\nthough, and I think \"hash compatibility option\" is easier to say.\n\nHmm, or are you talking about \"compatObjectFormat\" discussed in that doc?\n\n> Let's add an element to the repository\n> structure that indicates the compatibility hash algorithm so we can use\n> it when we need to consider interoperability between algorithms.\n\nHow about just\n\n    Add a hash compatibility option to the repository structure to\n    consider interoperability between hash algorithms.\n\n?\n\nAside: already we are seeing multiple keywords \"compatibility\",\n\"transition\", \"interoperability\" to all mean roughly similar things. I\nhope we can settle on just one (ideally) in the codebase by the end of\nthis series.\n\n> Add a helper function repo_set_compat_hash_algo that takes a\n> compatibility hash algorithm and sets \"repo->compat_hash_algo\".  If\n> GIT_HASH_UNKNOWN is passed as the compatibility hash algorithm\n> \"repo->compat_hash_algo\" is set to NULL.\n>\n> For now, the code results in \"repo->compat_hash_algo\" always being set\n> to NULL, but that will change once a configuration option is added.\n\nIt's not clear to me whether you are talking about a config option to\ndescribe the different stages of transition around algorithms, or a hash\nalgorithm itself (SHA1, SHA256, UNKNOWN).\n\n> Inspired-by: brian m. carlson <sandals@crustytoothpaste.net>\n> Signed-off-by: Eric W. Biederman <ebiederm@xmission.com>\n> ---\n>  repository.c | 8 ++++++++\n>  repository.h | 4 ++++\n>  setup.c      | 3 +++\n>  3 files changed, 15 insertions(+)\n>\n> diff --git a/repository.c b/repository.c\n> index a7679ceeaa45..80252b79e93e 100644\n> --- a/repository.c\n> +++ b/repository.c\n> @@ -104,6 +104,13 @@ void repo_set_hash_algo(struct repository *repo, int hash_algo)\n>  \trepo->hash_algo = &hash_algos[hash_algo];\n>  }\n>  \n> +void repo_set_compat_hash_algo(struct repository *repo, int algo)\n> +{\n> +\tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n> +\t\tBUG(\"hash_algo and compat_hash_algo match\");\n> +\trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n> +}\n\nAh, OK. So we are talking about an algorithm itself. Looking at this\ncode it seems like a compat_hash_algo is something like \"the hash\nalgorithm I want my repository to start using but which has not\nalready\". Such a description would have been useful in the commit\nmessage.\n\nNit: I think \n\n    BUG(\"compat_hash_algo may not be the same as hash_algo\");\n\nis more natural because the error message should explain the badness of\nthe behavior rather than merely reflect the triggering condition. And\nthe \"star of the show\" here is the new compat_hash_algo member, so it\nmakes sense to emphasize that more as the only subject of the sentence\ninstead of grouping it together with hash_algo (given them equal\nimportance).\n\n> +\n>  /*\n>   * Attempt to resolve and set the provided 'gitdir' for repository 'repo'.\n>   * Return 0 upon success and a non-zero value upon failure.\n> @@ -184,6 +191,7 @@ int repo_init(struct repository *repo,\n>  \t\tgoto error;\n>  \n>  \trepo_set_hash_algo(repo, format.hash_algo);\n> +\trepo_set_compat_hash_algo(repo, GIT_HASH_UNKNOWN);\n>  \trepo->repository_format_worktree_config = format.worktree_config;\n>  \n>  \t/* take ownership of format.partial_clone */\n> diff --git a/repository.h b/repository.h\n> index 5f18486f6465..bf3fc601cc53 100644\n> --- a/repository.h\n> +++ b/repository.h\n> @@ -160,6 +160,9 @@ struct repository {\n>  \t/* Repository's current hash algorithm, as serialized on disk. */\n>  \tconst struct git_hash_algo *hash_algo;\n>  \n> +\t/* Repository's compatibility hash algorithm. */\n\nPerhaps add \"May not be the same as hash_algo.\" ?\n\n> +\tconst struct git_hash_algo *compat_hash_algo;\n> +\n>  \t/* A unique-id for tracing purposes. */\n>  \tint trace2_repo_id;\n>  \n> @@ -199,6 +202,7 @@ void repo_set_gitdir(struct repository *repo, const char *root,\n>  \t\t     const struct set_gitdir_args *extra_args);\n>  void repo_set_worktree(struct repository *repo, const char *path);\n>  void repo_set_hash_algo(struct repository *repo, int algo);\n> +void repo_set_compat_hash_algo(struct repository *repo, int compat_algo);\n>  void initialize_the_repository(void);\n>  RESULT_MUST_BE_USED\n>  int repo_init(struct repository *r, const char *gitdir, const char *worktree);\n> diff --git a/setup.c b/setup.c\n> index 18927a847b86..aa8bf5da5226 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -1564,6 +1564,8 @@ const char *setup_git_directory_gently(int *nongit_ok)\n>  \t\t}\n>  \t\tif (startup_info->have_repository) {\n>  \t\t\trepo_set_hash_algo(the_repository, repo_fmt.hash_algo);\n> +\t\t\trepo_set_compat_hash_algo(the_repository,\n> +\t\t\t\t\t\t  GIT_HASH_UNKNOWN);\n>  \t\t\tthe_repository->repository_format_worktree_config =\n>  \t\t\t\trepo_fmt.worktree_config;\n>  \t\t\t/* take ownership of repo_fmt.partial_clone */\n> @@ -1657,6 +1659,7 @@ void check_repository_format(struct repository_format *fmt)\n>  \tcheck_repository_format_gently(get_git_dir(), fmt, NULL);\n>  \tstartup_info->have_repository = 1;\n>  \trepo_set_hash_algo(the_repository, fmt->hash_algo);\n> +\trepo_set_compat_hash_algo(the_repository, GIT_HASH_UNKNOWN);\n>  \tthe_repository->repository_format_worktree_config =\n>  \t\tfmt->worktree_config;\n>  \tthe_repository->repository_format_partial_clone =\n> -- \n> 2.41.0\n"},{"id":"488612","messageId":"owlybk8ja4w2.fsf@fine.c.googlers.com","threadId":"60273","inReplyTo":"20231002024034.2611-5-ebiederm@gmail.com","subject":"Re: [PATCH v2 05/30] loose: add a mapping between SHA-1 and SHA-256 for loose objects","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2024-02-14T07:20:29Z","receivedAt":"2024-02-14T07:20:31Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n>\n> As part of the transition plan, we'd like to add a file in the .git\n> directory that maps loose objects between SHA-1 and SHA-256.  Let's\n> implement the specification in the transition plan and store this data\n> on a per-repository basis in struct repository.\n\nCould you explain a bit what the specification is, exactly? That would\nsave reviewers the trouble of comparing the large chunk of new code here\nwith the transition plan (which I assume is still\nDocumentation/technical/hash-function-transition.txt.\n\nAlso, are there any slight deviations from the specification for reasons\nthat may not be obvious to reviewers?\n\nI would prefer if this patch is split up into smaller preparatory\npatches, starting with the core essentials and then building it up\npiece-by-piece to make it easier to review.\n\n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> Signed-off-by: Eric W. Biederman <ebiederm@xmission.com>\n> ---\n>  Makefile              |   1 +\n>  loose.c               | 246 ++++++++++++++++++++++++++++++++++++++++++\n>  loose.h               |  22 ++++\n>  object-file-convert.c |  14 ++-\n>  object-store-ll.h     |   3 +\n>  object.c              |   2 +\n>  repository.c          |   6 ++\n>  7 files changed, 293 insertions(+), 1 deletion(-)\n>  create mode 100644 loose.c\n>  create mode 100644 loose.h\n>\n> diff --git a/Makefile b/Makefile\n> index f7e824f25cda..3c18664def9a 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -1053,6 +1053,7 @@ LIB_OBJS += list-objects-filter.o\n>  LIB_OBJS += list-objects.o\n>  LIB_OBJS += lockfile.o\n>  LIB_OBJS += log-tree.o\n> +LIB_OBJS += loose.o\n\nThe name \"loose\" appears to be a bit too generic for something with such\na specialized role (_mapping_ of loose objects). Would\n\"loose-object-map\" be a better name?\n\n>  LIB_OBJS += ls-refs.o\n>  LIB_OBJS += mailinfo.o\n>  LIB_OBJS += mailmap.o\n> diff --git a/loose.c b/loose.c\n> new file mode 100644\n> index 000000000000..6ba73cc84dca\n> --- /dev/null\n> +++ b/loose.c\n> @@ -0,0 +1,246 @@\n> +#include \"git-compat-util.h\"\n> +#include \"hash.h\"\n> +#include \"path.h\"\n> +#include \"object-store.h\"\n> +#include \"hex.h\"\n> +#include \"wrapper.h\"\n> +#include \"gettext.h\"\n> +#include \"loose.h\"\n> +#include \"lockfile.h\"\n> +\n> +static const char *loose_object_header = \"# loose-object-idx\\n\";\n\nIDK what the \"loose-object-idx\" is vs the \"loose-object-map\", but I\nguess I need to read more of the code.\n\nBut also, I am at my limits here and am unable to review this patch as\nis (too big for me to chew at once, sorry).\n\nI'll pause my review of this series here to give Eric B some time to\nrespond. Thanks.\n"},{"id":"488613","messageId":"owly7cj7a464.fsf@fine.c.googlers.com","threadId":"60273","inReplyTo":"owlymssbn6qa.fsf@fine.c.googlers.com","subject":"Re: [PATCH v2 00/30] initial support for multiple hash functions","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2024-02-14T07:36:03Z","receivedAt":"2024-02-14T07:36:05Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"Linus Arver <linusa@google.com> writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n>>\n>>> This addresses all of the known test failures from v1 of this set of\n>>> changes.  In particular I have reworked commit_tree_extended which\n>>> was flagged by smatch, -Werror=array-bounds, and the leak detector.\n>>>\n>>> One functional bug was fixed in repo_for_each_abbrev where it was\n>>> mistakenly displaying too many ambiguous oids.\n>>>\n>>> I am posting this so that people review and testing of this patchset\n>>> won't be distracted by the known and fixed issues.\n>>\n>> We haven't seen any reviews on this second round, and have had it\n>> outside 'next' for too long.  I am tempted to say that we merge it\n>> to 'next' and see if anybody screams at this point.\n>\n> FWIW out of all the \"Needs review\" topics this one seemed like the most\n> deserving of another pair of eyes, and I was planning to review some of\n> the patches here this week + the weekend. If my review takes too long\n> (taking longer than this weekend) I can give another update next week\n> saying \"too hard for me, please don't wait for me\" to unblock you from\n> merging to next.\n>\n> Thanks.\n\nUnfortunately I don't think I can finish reviewing the rest of the\nseries (after all this time I've only been able to review just 4 out of\n30 patches). I'm also stuck on trying to understand patch 5, as there is\na lot going on there.\n\nFWIW a lot (perhaps all?) of my comments so far were around readability\nand not material to the actual design or approach of anything AFAICS.\nSo, it's time for me to say \"don't bother waiting for me\" as I said\n(predicted?) earlier.\n\nDon't bother waiting for me. Thanks.\n"},{"id":"488684","messageId":"8734tupa00.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"owlybk8ja4w2.fsf@fine.c.googlers.com","subject":"Re: [PATCH v2 05/30] loose: add a mapping between SHA-1 and SHA-256 for loose objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2024-02-15T05:33:19Z","receivedAt":"2024-02-15T05:33:23Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"Linus Arver <linusa@google.com> writes:\n\n> I'll pause my review of this series here to give Eric B some time to\n> respond. Thanks.\n\nI will respond shortly.  The re-awakening of the review process came\njust as I am in the middle of something else that is taking a lot of\ncycles.\n\nEric\n\n"},{"id":"488687","messageId":"8734tumekr.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"owlya5o4dbj1.fsf@fine.c.googlers.com","subject":"Re: [PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2024-02-15T06:22:44Z","receivedAt":"2024-02-15T06:22:47Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"Linus Arver <linusa@google.com> writes:\n\n> \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n>\n>> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>>\n>> While looking at how to handle input of both SHA-1 and SHA-256 oids in\n>> get_oid_with_context, I realized that the oid_array in\n>> repo_for_each_abbrev might have more than one kind of oid stored in it\n>> simultaneously.\n>>\n>> Update to oid_array_append to ensure that oids added to an oid array\n>\n> s/Update to/Update\n>\n>> always have an algorithm set.\n>>\n>> Update void_hashcmp to first verify two oids use the same hash algorithm\n>> before comparing them to each other.\n>>\n>> With that oid-array should be safe to use with different kinds of\n>\n> s/oid-array/oid_array\n>\n>> oids simultaneously.\n>>\n>> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>> ---\n>>  oid-array.c | 12 ++++++++++--\n>>  1 file changed, 10 insertions(+), 2 deletions(-)\n>>\n>> diff --git a/oid-array.c b/oid-array.c\n>> index 8e4717746c31..1f36651754ed 100644\n>> --- a/oid-array.c\n>> +++ b/oid-array.c\n>> @@ -6,12 +6,20 @@ void oid_array_append(struct oid_array *array, const struct object_id *oid)\n>>  {\n>>  \tALLOC_GROW(array->oid, array->nr + 1, array->alloc);\n>>  \toidcpy(&array->oid[array->nr++], oid);\n>> +\tif (!oid->algo)\n>> +\t\toid_set_algo(&array->oid[array->nr - 1], the_hash_algo);\n>\n> How come we can't set oid->algo _before_ we call oidcpy()? It seems odd\n> that we do the copy first and then modify what we just copied after the\n> fact, instead of making sure that the thing we want to copy is correct\n> before doing the copy.\n>\n> But also, if we are going to make the oid object \"correct\" before\n> invoking oidcpy(), we might as well do it when the oid is first\n> created/used (in the caller(s) of this function). I don't demand that\n> you find/demonstrate where all these places are in this series (maybe\n> that's a hairy problem to tackle?), but it seems cleaner in principle to\n> fix the creation of oid objects instead of having to make oid users\n> clean up their act like this after using them.\n\nThere is a hairy problem here.\n\nI believe for reasons of simplicity when the algo field was added to\nstruct object_id it was allowed to be zero for users that don't\nparticularly care about the hash algorithm, and are happy to use the git\ndefault hash algorithm.\n\nMe experience working on this set of change set showed that there\nare oids without their algo set in all kinds of places in the tree.\n\nI could not think of any sure way to go through the entire tree\nand find those users, so I just made certain that oid array handled\nthat case.\n\nI need algo to be set properly in the oids in the oid array so I\ncould extend oid_array to hold multiple kinds of oids at the same\ntime.  To allow multiple kinds of oids at the same time void_hashcmp\nneeds a simple and reliable way to tell what the algorithm is of\nany given oid.\n\n>\n>>  \tarray->sorted = 0;\n>>  }\n>>  \n>> -static int void_hashcmp(const void *a, const void *b)\n>> +static int void_hashcmp(const void *va, const void *vb)\n>>  {\n>> -\treturn oidcmp(a, b);\n>> +\tconst struct object_id *a = va, *b = vb;\n>> +\tint ret;\n>> +\tif (a->algo == b->algo)\n>> +\t\tret = oidcmp(a, b);\n>\n> This makes sense (per the commit message description) ...\n>\n>> +\telse\n>> +\t\tret = a->algo > b->algo ? 1 : -1;\n>\n> ... but this seems to go against it? I thought you wanted to only ever\n> compare hashes if they were of the same algo? It would be good to add a\n> comment explaining why this is OK (we are no longer doing a byte-by-byte\n> comparison of these oids any more here like we do for oidcmp() above\n> which boils down to calling memcmp()).\n\nSo the goal of this change is for oid_array to be able to hold hashes\nfrom multiple algorithms at the same time.\n\nA key part of oid_array is oid_array_sort that allows functions such\nas oid_array_lookup and oid_array_for_each_unique.\n\nTo that end there needs to be a total ordering of oids.\n\nThe function oidcmp is only defined when two oids are of the same\nalgorithm, it does not even test to detect the case of comparing\nmismatched algorithms.\n\nTherefore to get a total ordering of oids.  I must use oidcmp\nwhen the algorithm is the same (the common case) or simply order\nthe oids by algorithm when the algorithms are different.\n\n\n\nAll of this is relevant to get_oid_with_context as get_oid_with_context\nand it's helper functions contain the logic that determines what\nwe do when a hex string that is ambiguous is specified.\n\nIn the ambiguous case all of the possible candidates are placed in\nan oid_array, sorted and then displayed.\n\n\nWith a repository that can knows both the sha1 and the sha256 oid\nof it's objects it is possible for a short oid to match both\nsome sha1 oids and some sha256 oids.\n\n>> +\treturn ret;\n>\n> Also, in terms of style I think the \"early return for errors\" style\n> would be simpler to read. I.e.\n>\n>     if (a->algo > b->algo)\n>         return 1;\n>\n>     if (a->algo < b->algo)\n>         return -1;\n>\n>     return oidcmd(a, b);\n>\n\nI can see doing:\n\tif (a->algo == b->algo)\n        \treturn oidcmp(a,b);\n\n\tif (a->algo > b->algo)\n        \treturn 1;\n        else\n        \treturn -1;\n\nOr even:\n\tif (a->algo == b->algo)\n        \treturn oidcmp(a,b);\n\n\treturn a->algo - b->algo;\n\nAlthough I suspect using subtraction is a bit too clever.\n\nComparing for less than, and greater than, and then assuming\nthe values are equal hides what is important before calling\noidcmp which is that the algo values are equal.\n\n\n>>  }\n>>  \n>>  void oid_array_sort(struct oid_array *array)\n>> -- \n>> 2.41.0\n\nEric\n"},{"id":"488688","messageId":"87v86qkzwy.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"023394e2-5f64-4a59-af96-b77dafb20051@app.fastmail.com","subject":"Re: [PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2024-02-15T06:24:45Z","receivedAt":"2024-02-15T06:24:47Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"\"Kristoffer Haugsbakk\" <code@khaugsbakk.name> writes:\n\n> On Mon, Oct 2, 2023, at 04:40, Eric W. Biederman wrote:\n>> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>\n> Most of your patches have this sign-off line with your name quoted.\n\nAt least for email syntax from which Signed-off-by syntax descends\nhaving a period after my middle initial requires the name to be in\nquotes.\n\nEric\n\n\n\n"},{"id":"488724","messageId":"Zc3zyi42slXWGJTC@tanuki","threadId":"60273","inReplyTo":"20231002024034.2611-1-ebiederm@gmail.com","subject":"Re: [PATCH v2 01/30] object-file-convert: stubs for converting from one object format to another","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-02-15T11:21:46Z","receivedAt":"2024-02-15T11:21:54Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sun, Oct 01, 2023 at 09:40:05PM -0500, Eric W. Biederman wrote:\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> \n> Two basic functions are provided:\n> - convert_object_file Takes an object file it's type and hash algorithm\n>   and converts it into the equivalent object file that would\n>   have been generated with hash algorithm \"to\".\n> \n>   For blob objects there is no conversation to be done and it is an\n>   error to use this function on them.\n> \n>   For commit, tree, and tag objects embedded oids are replaced by the\n>   oids of the objects they refer to with those objects and their\n>   object ids reencoded in with the hash algorithm \"to\".  Signatures\n>   are rearranged so that they remain valid after the object has\n>   been reencoded.\n> \n> - repo_oid_to_algop which takes an oid that refers to an object file\n>   and returns the oid of the equivalent object file generated\n>   with the target hash algorithm.\n> \n> The pair of files object-file-convert.c and object-file-convert.h are\n> introduced to hold as much of this logic as possible to keep this\n> conversion logic cleanly separated from everything else and in the\n> hopes that someday the code will be clean enough git can support\n> compiling out support for sha1 and the various conversion functions.\n> \n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n>  Makefile              |  1 +\n>  object-file-convert.c | 57 +++++++++++++++++++++++++++++++++++++++++++\n>  object-file-convert.h | 24 ++++++++++++++++++\n>  3 files changed, 82 insertions(+)\n>  create mode 100644 object-file-convert.c\n>  create mode 100644 object-file-convert.h\n> \n> diff --git a/Makefile b/Makefile\n> index 577630936535..f7e824f25cda 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -1073,6 +1073,7 @@ LIB_OBJS += notes-cache.o\n>  LIB_OBJS += notes-merge.o\n>  LIB_OBJS += notes-utils.o\n>  LIB_OBJS += notes.o\n> +LIB_OBJS += object-file-convert.o\n>  LIB_OBJS += object-file.o\n>  LIB_OBJS += object-name.o\n>  LIB_OBJS += object.o\n> diff --git a/object-file-convert.c b/object-file-convert.c\n> new file mode 100644\n> index 000000000000..4777aba83636\n> --- /dev/null\n> +++ b/object-file-convert.c\n> @@ -0,0 +1,57 @@\n> +#include \"git-compat-util.h\"\n> +#include \"gettext.h\"\n> +#include \"strbuf.h\"\n> +#include \"repository.h\"\n> +#include \"hash-ll.h\"\n> +#include \"object.h\"\n> +#include \"object-file-convert.h\"\n> +\n> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n> +\t\t      const struct git_hash_algo *to, struct object_id *dest)\n> +{\n> +\t/*\n> +\t * If the source algorithm is not set, then we're using the\n> +\t * default hash algorithm for that object.\n> +\t */\n> +\tconst struct git_hash_algo *from =\n> +\t\tsrc->algo ? &hash_algos[src->algo] : repo->hash_algo;\n> +\n> +\tif (from == to) {\n> +\t\tif (src != dest)\n> +\t\t\toidcpy(dest, src);\n> +\t\treturn 0;\n> +\t}\n> +\treturn -1;\n> +}\n\nIn it's current form, `repo_oid_to_algop()` basically never does\nanything except for copying over the object ID because we do not handle\nthe case where object hashes are different. I assume this is intended,\nas we basically only provide stubs in this commit. But still, it would\nhelp to document this in-code as well with a comment.\n\n> +int convert_object_file(struct strbuf *outbuf,\n> +\t\t\tconst struct git_hash_algo *from,\n> +\t\t\tconst struct git_hash_algo *to,\n> +\t\t\tconst void *buf, size_t len,\n> +\t\t\tenum object_type type,\n> +\t\t\tint gentle)\n> +{\n> +\tint ret;\n> +\n> +\t/* Don't call this function when no conversion is necessary */\n> +\tif ((from == to) || (type == OBJ_BLOB))\n> +\t\tBUG(\"Refusing noop object file conversion\");\n\nThe extra braces around comparisons are unneeded and to the best of my\nknowledge not customary in our code base. Also, error messages should\nstart with a lower-case letter.\n\n> +\tswitch (type) {\n> +\tcase OBJ_COMMIT:\n> +\tcase OBJ_TREE:\n> +\tcase OBJ_TAG:\n> +\tdefault:\n> +\t\t/* Not implemented yet, so fail. */\n> +\t\tret = -1;\n> +\t\tbreak;\n> +\t}\n\nIt's a bit weird that we handle all object types except for blobs\nseparately, and then still have a `default` statement. I would've\nthought that we should handle the object types specifically and set `ret\n= -1` for all of them, and then the `default` case would instead call\n`BUG()` due to an unknown object type.\n\n> +\tif (!ret)\n> +\t\treturn 0;\n> +\tif (gentle) {\n> +\t\tstrbuf_release(outbuf);\n> +\t\treturn ret;\n> +\t}\n\nDo you really intend to call `strbuf_release()` on the caller provided\nbuffer, or should this rather be `strbuf_reset()`? Memory management of\nsuch an in/out parameter should typically be handled by the caller, not\nthe callee.\n\n> +\tdie(_(\"Failed to convert object from %s to %s\"),\n> +\t\tfrom->name, to->name);\n> +}\n\nThe error message should start with a lower-case letter.\n\n> diff --git a/object-file-convert.h b/object-file-convert.h\n> new file mode 100644\n> index 000000000000..a4f802aa8eea\n> --- /dev/null\n> +++ b/object-file-convert.h\n> @@ -0,0 +1,24 @@\n> +#ifndef OBJECT_CONVERT_H\n> +#define OBJECT_CONVERT_H\n> +\n> +struct repository;\n> +struct object_id;\n> +struct git_hash_algo;\n> +struct strbuf;\n> +#include \"object.h\"\n> +\n> +int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n> +\t\t      const struct git_hash_algo *to, struct object_id *dest);\n> +\n> +/*\n> + * Convert an object file from one hash algorithm to another algorithm.\n> + * Return -1 on failure, 0 on success.\n> + */\n> +int convert_object_file(struct strbuf *outbuf,\n> +\t\t\tconst struct git_hash_algo *from,\n> +\t\t\tconst struct git_hash_algo *to,\n> +\t\t\tconst void *buf, size_t len,\n> +\t\t\tenum object_type type,\n> +\t\t\tint gentle);\n\nIt would be nice to document what `gentle` does.\n\nPatrick\n\n> +#endif /* OBJECT_CONVERT_H */\n> -- \n> 2.41.0\n> \n"},{"id":"488725","messageId":"Zc3zz0hJFShTZp3M@tanuki","threadId":"60273","inReplyTo":"20231002024034.2611-2-ebiederm@gmail.com","subject":"Re: [PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-02-15T11:21:51Z","receivedAt":"2024-02-15T11:21:55Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sun, Oct 01, 2023 at 09:40:06PM -0500, Eric W. Biederman wrote:\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> \n> While looking at how to handle input of both SHA-1 and SHA-256 oids in\n> get_oid_with_context, I realized that the oid_array in\n> repo_for_each_abbrev might have more than one kind of oid stored in it\n> simultaneously.\n> \n> Update to oid_array_append to ensure that oids added to an oid array\n> always have an algorithm set.\n> \n> Update void_hashcmp to first verify two oids use the same hash algorithm\n> before comparing them to each other.\n> \n> With that oid-array should be safe to use with different kinds of\n> oids simultaneously.\n> \n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n>  oid-array.c | 12 ++++++++++--\n>  1 file changed, 10 insertions(+), 2 deletions(-)\n> \n> diff --git a/oid-array.c b/oid-array.c\n> index 8e4717746c31..1f36651754ed 100644\n> --- a/oid-array.c\n> +++ b/oid-array.c\n> @@ -6,12 +6,20 @@ void oid_array_append(struct oid_array *array, const struct object_id *oid)\n>  {\n>  \tALLOC_GROW(array->oid, array->nr + 1, array->alloc);\n>  \toidcpy(&array->oid[array->nr++], oid);\n> +\tif (!oid->algo)\n> +\t\toid_set_algo(&array->oid[array->nr - 1], the_hash_algo);\n\nI feel like it's a design wart that `oid_array_append()` now started to\ndepend on repository discovery, adding an external dependency to it that\nmay cause very confusing behaviour. Are there for example ever cases\nwhere we populate such an OID array before we have discovered the repo?\nCan it happen that we use OID arrays in the context of a submodule that\nhas a different object ID than the main repository?\n\n>  \tarray->sorted = 0;\n>  }\n>  \n> -static int void_hashcmp(const void *a, const void *b)\n> +static int void_hashcmp(const void *va, const void *vb)\n>  {\n> -\treturn oidcmp(a, b);\n> +\tconst struct object_id *a = va, *b = vb;\n> +\tint ret;\n> +\tif (a->algo == b->algo)\n> +\t\tret = oidcmp(a, b);\n> +\telse\n> +\t\tret = a->algo > b->algo ? 1 : -1;\n\nOkay, so we basically end up sorting first by the algorithm, and then we\nsort all object IDs of a specific algorithm relative to each other.\n\nPatrick\n\n> +\treturn ret;\n>  }\n>  \n>  void oid_array_sort(struct oid_array *array)\n> -- \n> 2.41.0\n> \n"},{"id":"488726","messageId":"Zc3z09woCStanYlP@tanuki","threadId":"60273","inReplyTo":"20231002024034.2611-3-ebiederm@gmail.com","subject":"Re: [PATCH v2 03/30] object-names: support input of oids in any supported hash","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-02-15T11:21:55Z","receivedAt":"2024-02-15T11:21:59Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sun, Oct 01, 2023 at 09:40:07PM -0500, Eric W. Biederman wrote:\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> \n> Support short oids encoded in any algorithm, while ensuring enough of\n> the oid is specified to disambiguate between all of the oids in the\n> repository encoded in any algorithm.\n> \n> By default have the code continue to only accept oids specified in the\n> storage hash algorithm of the repository, but when something is\n> ambiguous display all of the possible oids from any accepted oid\n> encoding.\n> \n> A new flag is added GET_OID_HASH_ANY that when supplied causes the\n> code to accept oids specified in any hash algorithm, and to return the\n> oids that were resolved.\n> \n> This implements the functionality that allows both SHA-1 and SHA-256\n> object names, from the \"Object names on the command line\" section of\n> the hash function transition document.\n> \n> Care is taken in get_short_oid so that when the result is ambiguous\n> the output remains the same if GIT_OID_HASH_ANY was not supplied.  If\n> GET_OID_HASH_ANY was supplied objects of any hash algorithm that match\n> the prefix are displayed.\n> \n> This required updating repo_for_each_abbrev to give it a parameter so\n> that it knows to look at all hash algorithms.\n> \n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n>  builtin/rev-parse.c |  2 +-\n>  hash-ll.h           |  1 +\n>  object-name.c       | 46 ++++++++++++++++++++++++++++++++++-----------\n>  object-name.h       |  3 ++-\n>  4 files changed, 39 insertions(+), 13 deletions(-)\n> \n> diff --git a/builtin/rev-parse.c b/builtin/rev-parse.c\n> index fde8861ca4e0..43e96765400c 100644\n> --- a/builtin/rev-parse.c\n> +++ b/builtin/rev-parse.c\n> @@ -882,7 +882,7 @@ int cmd_rev_parse(int argc, const char **argv, const char *prefix)\n>  \t\t\t\tcontinue;\n>  \t\t\t}\n>  \t\t\tif (skip_prefix(arg, \"--disambiguate=\", &arg)) {\n> -\t\t\t\trepo_for_each_abbrev(the_repository, arg,\n> +\t\t\t\trepo_for_each_abbrev(the_repository, arg, the_hash_algo,\n>  \t\t\t\t\t\t     show_abbrev, NULL);\n>  \t\t\t\tcontinue;\n>  \t\t\t}\n> diff --git a/hash-ll.h b/hash-ll.h\n> index 10d84cc20888..2cfde63ae1cf 100644\n> --- a/hash-ll.h\n> +++ b/hash-ll.h\n> @@ -145,6 +145,7 @@ struct object_id {\n>  #define GET_OID_RECORD_PATH     0200\n>  #define GET_OID_ONLY_TO_DIE    04000\n>  #define GET_OID_REQUIRE_PATH  010000\n> +#define GET_OID_HASH_ANY      020000\n>  \n>  #define GET_OID_DISAMBIGUATORS \\\n>  \t(GET_OID_COMMIT | GET_OID_COMMITTISH | \\\n> diff --git a/object-name.c b/object-name.c\n> index 0bfa29dbbfe9..7dd6e5e47566 100644\n> --- a/object-name.c\n> +++ b/object-name.c\n> @@ -25,6 +25,7 @@\n>  #include \"midx.h\"\n>  #include \"commit-reach.h\"\n>  #include \"date.h\"\n> +#include \"object-file-convert.h\"\n>  \n>  static int get_oid_oneline(struct repository *r, const char *, struct object_id *, struct commit_list *);\n>  \n> @@ -49,6 +50,7 @@ struct disambiguate_state {\n>  \n>  static void update_candidates(struct disambiguate_state *ds, const struct object_id *current)\n>  {\n> +\t/* The hash algorithm of current has already been filtered */\n>  \tif (ds->always_call_fn) {\n>  \t\tds->ambiguous = ds->fn(ds->repo, current, ds->cb_data) ? 1 : 0;\n>  \t\treturn;\n> @@ -134,6 +136,8 @@ static void unique_in_midx(struct multi_pack_index *m,\n>  {\n>  \tuint32_t num, i, first = 0;\n>  \tconst struct object_id *current = NULL;\n> +\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n> +\t\tds->repo->hash_algo->hexsz : ds->len;\n\n`hexsz` is not an `int`, but a `size_t`. `match_hash()` of course uses\na third type `unsigned` instead, adding to the confusion.\n\n>  \tnum = m->num_objects;\n>  \n>  \tif (!num)\n> @@ -149,7 +153,7 @@ static void unique_in_midx(struct multi_pack_index *m,\n>  \tfor (i = first; i < num && !ds->ambiguous; i++) {\n>  \t\tstruct object_id oid;\n>  \t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n> -\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, current->hash))\n> +\t\tif (!match_hash(len, ds->bin_pfx.hash, current->hash))\n>  \t\t\tbreak;\n>  \t\tupdate_candidates(ds, current);\n>  \t}\n> @@ -159,6 +163,8 @@ static void unique_in_pack(struct packed_git *p,\n>  \t\t\t   struct disambiguate_state *ds)\n>  {\n>  \tuint32_t num, i, first = 0;\n> +\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n> +\t\tds->repo->hash_algo->hexsz : ds->len;\n>  \n>  \tif (p->multi_pack_index)\n>  \t\treturn;\n> @@ -177,7 +183,7 @@ static void unique_in_pack(struct packed_git *p,\n>  \tfor (i = first; i < num && !ds->ambiguous; i++) {\n>  \t\tstruct object_id oid;\n>  \t\tnth_packed_object_id(&oid, p, i);\n> -\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, oid.hash))\n> +\t\tif (!match_hash(len, ds->bin_pfx.hash, oid.hash))\n>  \t\t\tbreak;\n>  \t\tupdate_candidates(ds, &oid);\n>  \t}\n> @@ -188,6 +194,10 @@ static void find_short_packed_object(struct disambiguate_state *ds)\n>  \tstruct multi_pack_index *m;\n>  \tstruct packed_git *p;\n>  \n> +\t/* Skip, unless oids from the storage hash algorithm are wanted */\n> +\tif (ds->bin_pfx.algo && (&hash_algos[ds->bin_pfx.algo] != ds->repo->hash_algo))\n> +\t\treturn;\n> +\n>  \tfor (m = get_multi_pack_index(ds->repo); m && !ds->ambiguous;\n>  \t     m = m->next)\n>  \t\tunique_in_midx(m, ds);\n> @@ -326,11 +336,12 @@ int set_disambiguate_hint_config(const char *var, const char *value)\n>  \n>  static int init_object_disambiguation(struct repository *r,\n>  \t\t\t\t      const char *name, int len,\n> +\t\t\t\t      const struct git_hash_algo *algo,\n>  \t\t\t\t      struct disambiguate_state *ds)\n>  {\n>  \tint i;\n>  \n> -\tif (len < MINIMUM_ABBREV || len > the_hash_algo->hexsz)\n> +\tif (len < MINIMUM_ABBREV || len > GIT_MAX_HEXSZ)\n>  \t\treturn -1;\n\nIsn't this loosening things up a bit too much? I'd have expected that we\nwould compare with `algo->hexsz`, unless `GET_OID_HASH_ANY` is set and\nthus `algo == NULL`.\n\nPatrick\n\n>  \tmemset(ds, 0, sizeof(*ds));\n> @@ -357,6 +368,7 @@ static int init_object_disambiguation(struct repository *r,\n>  \tds->len = len;\n>  \tds->hex_pfx[len] = '\\0';\n>  \tds->repo = r;\n> +\tds->bin_pfx.algo = algo ? hash_algo_by_ptr(algo) : GIT_HASH_UNKNOWN;\n>  \tprepare_alt_odb(r);\n>  \treturn 0;\n>  }\n> @@ -491,9 +503,10 @@ static int repo_collect_ambiguous(struct repository *r UNUSED,\n>  \treturn collect_ambiguous(oid, data);\n>  }\n>  \n> -static int sort_ambiguous(const void *a, const void *b, void *ctx)\n> +static int sort_ambiguous(const void *va, const void *vb, void *ctx)\n>  {\n>  \tstruct repository *sort_ambiguous_repo = ctx;\n> +\tconst struct object_id *a = va, *b = vb;\n>  \tint a_type = oid_object_info(sort_ambiguous_repo, a, NULL);\n>  \tint b_type = oid_object_info(sort_ambiguous_repo, b, NULL);\n>  \tint a_type_sort;\n> @@ -503,8 +516,12 @@ static int sort_ambiguous(const void *a, const void *b, void *ctx)\n>  \t * Sorts by hash within the same object type, just as\n>  \t * oid_array_for_each_unique() would do.\n>  \t */\n> -\tif (a_type == b_type)\n> -\t\treturn oidcmp(a, b);\n> +\tif (a_type == b_type) {\n> +\t\tif (a->algo == b->algo)\n> +\t\t\treturn oidcmp(a, b);\n> +\t\telse\n> +\t\t\treturn a->algo > b->algo ? 1 : -1;\n> +\t}\n>  \n>  \t/*\n>  \t * Between object types show tags, then commits, and finally\n> @@ -533,8 +550,12 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>  \tint status;\n>  \tstruct disambiguate_state ds;\n>  \tint quietly = !!(flags & GET_OID_QUIETLY);\n> +\tconst struct git_hash_algo *algo = r->hash_algo;\n> +\n> +\tif (flags & GET_OID_HASH_ANY)\n> +\t\talgo = NULL;\n>  \n> -\tif (init_object_disambiguation(r, name, len, &ds) < 0)\n> +\tif (init_object_disambiguation(r, name, len, algo, &ds) < 0)\n>  \t\treturn -1;\n>  \n>  \tif (HAS_MULTI_BITS(flags & GET_OID_DISAMBIGUATORS))\n> @@ -588,7 +609,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>  \t\tif (!ds.ambiguous)\n>  \t\t\tds.fn = NULL;\n>  \n> -\t\trepo_for_each_abbrev(r, ds.hex_pfx, collect_ambiguous, &collect);\n> +\t\trepo_for_each_abbrev(r, ds.hex_pfx, algo, collect_ambiguous, &collect);\n>  \t\tsort_ambiguous_oid_array(r, &collect);\n>  \n>  \t\tif (oid_array_for_each(&collect, show_ambiguous_object, &out))\n> @@ -610,13 +631,14 @@ static enum get_oid_result get_short_oid(struct repository *r,\n>  }\n>  \n>  int repo_for_each_abbrev(struct repository *r, const char *prefix,\n> +\t\t\t const struct git_hash_algo *algo,\n>  \t\t\t each_abbrev_fn fn, void *cb_data)\n>  {\n>  \tstruct oid_array collect = OID_ARRAY_INIT;\n>  \tstruct disambiguate_state ds;\n>  \tint ret;\n>  \n> -\tif (init_object_disambiguation(r, prefix, strlen(prefix), &ds) < 0)\n> +\tif (init_object_disambiguation(r, prefix, strlen(prefix), algo, &ds) < 0)\n>  \t\treturn -1;\n>  \n>  \tds.always_call_fn = 1;\n> @@ -787,10 +809,12 @@ void strbuf_add_unique_abbrev(struct strbuf *sb, const struct object_id *oid,\n>  int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n>  \t\t\t      const struct object_id *oid, int len)\n>  {\n> +\tconst struct git_hash_algo *algo =\n> +\t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n>  \tstruct disambiguate_state ds;\n>  \tstruct min_abbrev_data mad;\n>  \tstruct object_id oid_ret;\n> -\tconst unsigned hexsz = r->hash_algo->hexsz;\n> +\tconst unsigned hexsz = algo->hexsz;\n>  \n>  \tif (len < 0) {\n>  \t\tunsigned long count = repo_approximate_object_count(r);\n> @@ -826,7 +850,7 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n>  \n>  \tfind_abbrev_len_packed(&mad);\n>  \n> -\tif (init_object_disambiguation(r, hex, mad.cur_len, &ds) < 0)\n> +\tif (init_object_disambiguation(r, hex, mad.cur_len, algo, &ds) < 0)\n>  \t\treturn -1;\n>  \n>  \tds.fn = repo_extend_abbrev_len;\n> diff --git a/object-name.h b/object-name.h\n> index 9ae522307148..064ddc97d1fe 100644\n> --- a/object-name.h\n> +++ b/object-name.h\n> @@ -67,7 +67,8 @@ enum get_oid_result get_oid_with_context(struct repository *repo, const char *st\n>  \n>  \n>  typedef int each_abbrev_fn(const struct object_id *oid, void *);\n> -int repo_for_each_abbrev(struct repository *r, const char *prefix, each_abbrev_fn, void *);\n> +int repo_for_each_abbrev(struct repository *r, const char *prefix,\n> +\t\t\t const struct git_hash_algo *algo, each_abbrev_fn, void *);\n>  \n>  int set_disambiguate_hint_config(const char *var, const char *value);\n>  \n> -- \n> 2.41.0\n> \n"},{"id":"488727","messageId":"Zc3z2F7r2oMSlOW-@tanuki","threadId":"60273","inReplyTo":"20231002024034.2611-4-ebiederm@gmail.com","subject":"Re: [PATCH v2 04/30] repository: add a compatibility hash algorithm","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-02-15T11:22:00Z","receivedAt":"2024-02-15T11:22:04Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sun, Oct 01, 2023 at 09:40:08PM -0500, Eric W. Biederman wrote:\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> \n> We currently have support for using a full stage 4 SHA-256\n> implementation.\n\nWhat is a \"full stage 4 SHA-256 implementation\"? I was assuming that you\nreferred to \"Documentation/technical/hash-function-transition.txt\", but\nit does not mention stages either.\n\n> However, we'd like to support interoperability with\n> SHA-1 repositories as well.  The transition plan anticipates a\n> compatibility hash algorithm configuration option that we can use to\n> implement support for this.  Let's add an element to the repository\n> structure that indicates the compatibility hash algorithm so we can use\n> it when we need to consider interoperability between algorithms.\n> \n> Add a helper function repo_set_compat_hash_algo that takes a\n> compatibility hash algorithm and sets \"repo->compat_hash_algo\".  If\n> GIT_HASH_UNKNOWN is passed as the compatibility hash algorithm\n> \"repo->compat_hash_algo\" is set to NULL.\n> \n> For now, the code results in \"repo->compat_hash_algo\" always being set\n> to NULL, but that will change once a configuration option is added.\n> \n> Inspired-by: brian m. carlson <sandals@crustytoothpaste.net>\n> Signed-off-by: Eric W. Biederman <ebiederm@xmission.com>\n> ---\n>  repository.c | 8 ++++++++\n>  repository.h | 4 ++++\n>  setup.c      | 3 +++\n>  3 files changed, 15 insertions(+)\n> \n> diff --git a/repository.c b/repository.c\n> index a7679ceeaa45..80252b79e93e 100644\n> --- a/repository.c\n> +++ b/repository.c\n> @@ -104,6 +104,13 @@ void repo_set_hash_algo(struct repository *repo, int hash_algo)\n>  \trepo->hash_algo = &hash_algos[hash_algo];\n>  }\n>  \n> +void repo_set_compat_hash_algo(struct repository *repo, int algo)\n> +{\n> +\tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n> +\t\tBUG(\"hash_algo and compat_hash_algo match\");\n> +\trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n> +}\n> +\n>  /*\n>   * Attempt to resolve and set the provided 'gitdir' for repository 'repo'.\n>   * Return 0 upon success and a non-zero value upon failure.\n> @@ -184,6 +191,7 @@ int repo_init(struct repository *repo,\n>  \t\tgoto error;\n>  \n>  \trepo_set_hash_algo(repo, format.hash_algo);\n> +\trepo_set_compat_hash_algo(repo, GIT_HASH_UNKNOWN);\n>  \trepo->repository_format_worktree_config = format.worktree_config;\n>  \n>  \t/* take ownership of format.partial_clone */\n> diff --git a/repository.h b/repository.h\n> index 5f18486f6465..bf3fc601cc53 100644\n> --- a/repository.h\n> +++ b/repository.h\n> @@ -160,6 +160,9 @@ struct repository {\n>  \t/* Repository's current hash algorithm, as serialized on disk. */\n>  \tconst struct git_hash_algo *hash_algo;\n>  \n> +\t/* Repository's compatibility hash algorithm. */\n> +\tconst struct git_hash_algo *compat_hash_algo;\n> +\n>  \t/* A unique-id for tracing purposes. */\n>  \tint trace2_repo_id;\n>  \n> @@ -199,6 +202,7 @@ void repo_set_gitdir(struct repository *repo, const char *root,\n>  \t\t     const struct set_gitdir_args *extra_args);\n>  void repo_set_worktree(struct repository *repo, const char *path);\n>  void repo_set_hash_algo(struct repository *repo, int algo);\n> +void repo_set_compat_hash_algo(struct repository *repo, int compat_algo);\n>  void initialize_the_repository(void);\n>  RESULT_MUST_BE_USED\n>  int repo_init(struct repository *r, const char *gitdir, const char *worktree);\n> diff --git a/setup.c b/setup.c\n> index 18927a847b86..aa8bf5da5226 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -1564,6 +1564,8 @@ const char *setup_git_directory_gently(int *nongit_ok)\n>  \t\t}\n>  \t\tif (startup_info->have_repository) {\n>  \t\t\trepo_set_hash_algo(the_repository, repo_fmt.hash_algo);\n> +\t\t\trepo_set_compat_hash_algo(the_repository,\n> +\t\t\t\t\t\t  GIT_HASH_UNKNOWN);\n>  \t\t\tthe_repository->repository_format_worktree_config =\n>  \t\t\t\trepo_fmt.worktree_config;\n>  \t\t\t/* take ownership of repo_fmt.partial_clone */\n> @@ -1657,6 +1659,7 @@ void check_repository_format(struct repository_format *fmt)\n>  \tcheck_repository_format_gently(get_git_dir(), fmt, NULL);\n>  \tstartup_info->have_repository = 1;\n>  \trepo_set_hash_algo(the_repository, fmt->hash_algo);\n> +\trepo_set_compat_hash_algo(the_repository, GIT_HASH_UNKNOWN);\n>  \tthe_repository->repository_format_worktree_config =\n>  \t\tfmt->worktree_config;\n>  \tthe_repository->repository_format_partial_clone =\n\nThere's also `init_db()`, where we call `repo_set_hash_algo()`. Would we\nhave to call `repo_set_compat_hash_algo()` there, too? There are some\nother locations when handling remotes or clones, but I don't think those\nare relevant right now.\n\nPatrick\n"},{"id":"488728","messageId":"Zc3z3YhIzCcUxXlI@tanuki","threadId":"60273","inReplyTo":"20231002024034.2611-5-ebiederm@gmail.com","subject":"Re: [PATCH v2 05/30] loose: add a mapping between SHA-1 and SHA-256 for loose objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-02-15T11:22:05Z","receivedAt":"2024-02-15T11:22:08Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sun, Oct 01, 2023 at 09:40:09PM -0500, Eric W. Biederman wrote:\n> From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n> \n> As part of the transition plan, we'd like to add a file in the .git\n> directory that maps loose objects between SHA-1 and SHA-256.  Let's\n> implement the specification in the transition plan and store this data\n> on a per-repository basis in struct repository.\n> \n> Signed-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n> Signed-off-by: Eric W. Biederman <ebiederm@xmission.com>\n> ---\n>  Makefile              |   1 +\n>  loose.c               | 246 ++++++++++++++++++++++++++++++++++++++++++\n>  loose.h               |  22 ++++\n>  object-file-convert.c |  14 ++-\n>  object-store-ll.h     |   3 +\n>  object.c              |   2 +\n>  repository.c          |   6 ++\n>  7 files changed, 293 insertions(+), 1 deletion(-)\n>  create mode 100644 loose.c\n>  create mode 100644 loose.h\n> \n> diff --git a/Makefile b/Makefile\n> index f7e824f25cda..3c18664def9a 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -1053,6 +1053,7 @@ LIB_OBJS += list-objects-filter.o\n>  LIB_OBJS += list-objects.o\n>  LIB_OBJS += lockfile.o\n>  LIB_OBJS += log-tree.o\n> +LIB_OBJS += loose.o\n>  LIB_OBJS += ls-refs.o\n>  LIB_OBJS += mailinfo.o\n>  LIB_OBJS += mailmap.o\n> diff --git a/loose.c b/loose.c\n> new file mode 100644\n> index 000000000000..6ba73cc84dca\n> --- /dev/null\n> +++ b/loose.c\n\nWhen reading \"loose\" I immediately think about loose objects, only. I\nwould not consider this about mapping object IDs, which I expect would\nalso happen for packed objects?\n\nIt very much seems like you explicitly only care about loose objects in\nthe code here, which is weird to me. If that is in fact intentional\nbecause we learn to store the compat object hash in pack files over the\ncourse of this patch seires then it would make sense to explain this a\nbit more in depth.\n\n> @@ -0,0 +1,246 @@\n> +#include \"git-compat-util.h\"\n> +#include \"hash.h\"\n> +#include \"path.h\"\n> +#include \"object-store.h\"\n> +#include \"hex.h\"\n> +#include \"wrapper.h\"\n> +#include \"gettext.h\"\n> +#include \"loose.h\"\n> +#include \"lockfile.h\"\n> +\n> +static const char *loose_object_header = \"# loose-object-idx\\n\";\n> +\n> +static inline int should_use_loose_object_map(struct repository *repo)\n> +{\n> +\treturn repo->compat_hash_algo && repo->gitdir;\n> +}\n> +\n> +void loose_object_map_init(struct loose_object_map **map)\n> +{\n> +\tstruct loose_object_map *m;\n> +\tm = xmalloc(sizeof(**map));\n> +\tm->to_compat = kh_init_oid_map();\n> +\tm->to_storage = kh_init_oid_map();\n> +\t*map = m;\n> +}\n> +\n> +static int insert_oid_pair(kh_oid_map_t *map, const struct object_id *key, const struct object_id *value)\n> +{\n> +\tkhiter_t pos;\n> +\tint ret;\n> +\tstruct object_id *stored;\n> +\n> +\tpos = kh_put_oid_map(map, *key, &ret);\n> +\n> +\t/* This item already exists in the map. */\n> +\tif (ret == 0)\n> +\t\treturn 0;\n\nShould we safeguard this and compare whether the key's value matches the\npassed-in value? One of the more general themes that I'm worried about\nis what happens when we hit hash collisions (e.g. two objects mapping to\nthe same SHA1, but different SHA256 hashes), and safeguarding us against\nthis possibility feels sensible to me.\n\n> +\tstored = xmalloc(sizeof(*stored));\n> +\toidcpy(stored, value);\n> +\tkh_value(map, pos) = stored;\n> +\treturn 1;\n> +}\n> +\n> +static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n> +{\n> +\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n> +\tFILE *fp;\n> +\n> +\tif (!dir->loose_map)\n> +\t\tloose_object_map_init(&dir->loose_map);\n> +\n> +\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n> +\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n> +\n> +\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n> +\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n> +\n> +\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n> +\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n> +\n> +\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n> +\tfp = fopen(path.buf, \"rb\");\n> +\tif (!fp) {\n> +\t\tstrbuf_release(&path);\n> +\t\treturn 0;\n\nI think we should discern ENOENT from other errors. Failing gracefully\nwhen the file doesn't exist may be sensible, but not when we failed due\nto something like an I/O error.\n\n> +\t}\n> +\n> +\terrno = 0;\n> +\tif (strbuf_getwholeline(&buf, fp, '\\n') || strcmp(buf.buf, loose_object_header))\n> +\t\tgoto err;\n> +\twhile (!strbuf_getline_lf(&buf, fp)) {\n> +\t\tconst char *p;\n> +\t\tstruct object_id oid, compat_oid;\n> +\t\tif (parse_oid_hex_algop(buf.buf, &oid, &p, repo->hash_algo) ||\n> +\t\t    *p++ != ' ' ||\n> +\t\t    parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n> +\t\t    p != buf.buf + buf.len)\n> +\t\t\tgoto err;\n> +\t\tinsert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n> +\t\tinsert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n> +\t}\n\nIs the actual format specified anywhere? I have to wonder about the\nscalability of such a format that uses a simple line-based format for\nevery object ID. Two main concerns:\n\n  1. If the format is unsorted and we simply append to it whenever the\n     repo gains new objects then we are forced to always load the\n     complete map into memory. This would be quite inefficient in larger\n     repositories that have millions of objects. Every line contains two\n     object hashes as well as two whitespace characters, which amounts\n     to `(2 + 40 + 64) * $numobjects` many bytes.\n\n     For linux.git with more than 10 million objects, the map would thus\n     be around 1GB in size. Loading that into memory and converting it\n     into maps feels prohibitively expensive to me.\n\n  2. If the format was sorted then we could perform binary searches\n     inside the format to look up object IDs because we know that each\n     line has a fixed length. On the other hand, adding new objects\n     would require us to rewrite the whole file every time.\n\nI think loading the complete object map into memory is simply too\nexpensive in any larger \"real-world\" repository. But rewriting a sorted\nfile format every time we add new objects feels sufficiently expensive,\ntoo. Neither of these properties sounds like it would be feasible to use\nfor larger Git hosting platforms. So I think we should put some more\nthought into this.\n\nSome proposals:\n\n  - We shouldn't store hex characters but raw object IDs, thus reducing\n    the size of the file by almost half.\n\n  - We should store the file sorted so that we can avoid loading it into\n    memory and do binary searches.\n\n  - We might grow this into a \"stack\" of object maps so that it becomes\n    easier to add new objects to the map without having to rewrite it\n    every time. With geometric repacking this should be somewhat\n    manageable.\n\nWe don't have to do all of this right from the beginning, I just want to\nstart the discussion around this.\n\n> +\tstrbuf_release(&buf);\n> +\tstrbuf_release(&path);\n> +\treturn errno ? -1 : 0;\n\nIt feels quite fragile to me to check for `errno` in this way. Should we\ninstead check `ferror(fp)`?\n\n> +err:\n> +\tstrbuf_release(&buf);\n> +\tstrbuf_release(&path);\n> +\treturn -1;\n> +}\n\nWe could deduplicate the error paths by storing the return value into an\n`int ret`.\n\n> +int repo_read_loose_object_map(struct repository *repo)\n> +{\n> +\tstruct object_directory *dir;\n> +\n> +\tif (!should_use_loose_object_map(repo))\n> +\t\treturn 0;\n> +\n> +\tprepare_alt_odb(repo);\n> +\n> +\tfor (dir = repo->objects->odb; dir; dir = dir->next) {\n> +\t\tif (load_one_loose_object_map(repo, dir) < 0) {\n> +\t\t\treturn -1;\n> +\t\t}\n> +\t}\n\nThe braces here are not needed.\n\n> +\treturn 0;\n> +}\n> +\n> +int repo_write_loose_object_map(struct repository *repo)\n> +{\n> +\tkh_oid_map_t *map = repo->objects->odb->loose_map->to_compat;\n> +\tstruct lock_file lock;\n> +\tint fd;\n> +\tkhiter_t iter;\n> +\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n> +\n> +\tif (!should_use_loose_object_map(repo))\n> +\t\treturn 0;\n> +\n> +\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n> +\tfd = hold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n> +\titer = kh_begin(map);\n> +\tif (write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n> +\t\tgoto errout;\n> +\n> +\tfor (; iter != kh_end(map); iter++) {\n> +\t\tif (kh_exist(map, iter)) {\n> +\t\t\tif (oideq(&kh_key(map, iter), the_hash_algo->empty_tree) ||\n> +\t\t\t    oideq(&kh_key(map, iter), the_hash_algo->empty_blob))\n> +\t\t\t\tcontinue;\n> +\t\t\tstrbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(&kh_key(map, iter)), oid_to_hex(kh_value(map, iter)));\n> +\t\t\tif (write_in_full(fd, buf.buf, buf.len) < 0)\n> +\t\t\t\tgoto errout;\n> +\t\t\tstrbuf_reset(&buf);\n> +\t\t}\n> +\t}\n> +\tstrbuf_release(&buf);\n> +\tif (commit_lock_file(&lock) < 0) {\n> +\t\terror_errno(_(\"could not write loose object index %s\"), path.buf);\n> +\t\tstrbuf_release(&path);\n> +\t\treturn -1;\n> +\t}\n> +\tstrbuf_release(&path);\n> +\treturn 0;\n> +errout:\n> +\trollback_lock_file(&lock);\n> +\tstrbuf_release(&buf);\n> +\terror_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n> +\tstrbuf_release(&path);\n> +\treturn -1;\n\nSame here, we should be able to combine cleanup of both the successful\nand error paths. It's safe to call `rollback_lock_file()` even if the\nfile has already been committed.\n\n> +}\n> +\n> +static int write_one_object(struct repository *repo, const struct object_id *oid,\n> +\t\t\t    const struct object_id *compat_oid)\n> +{\n> +\tstruct lock_file lock;\n> +\tint fd;\n> +\tstruct stat st;\n> +\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n> +\n> +\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n> +\thold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n> +\n> +\tfd = open(path.buf, O_WRONLY | O_CREAT | O_APPEND, 0666);\n> +\tif (fd < 0)\n> +\t\tgoto errout;\n> +\tif (fstat(fd, &st) < 0)\n> +\t\tgoto errout;\n> +\tif (!st.st_size && write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n> +\t\tgoto errout;\n> +\n> +\tstrbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(oid), oid_to_hex(compat_oid));\n> +\tif (write_in_full(fd, buf.buf, buf.len) < 0)\n> +\t\tgoto errout;\n> +\tif (close(fd))\n> +\t\tgoto errout;\n\nIt's not safe to update the file in-place like this. A concurrent reader\nmay end up seeing partial lines and error out. Also, if we were to crash\nwe might easily end up with a corrupted mapping file.\n\n> +\tadjust_shared_perm(path.buf);\n> +\trollback_lock_file(&lock);\n> +\tstrbuf_release(&buf);\n> +\tstrbuf_release(&path);\n> +\treturn 0;\n> +errout:\n> +\terror_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n> +\tclose(fd);\n> +\trollback_lock_file(&lock);\n> +\tstrbuf_release(&buf);\n> +\tstrbuf_release(&path);\n> +\treturn -1;\n\nSame.\n\n> +}\n> +\n> +int repo_add_loose_object_map(struct repository *repo, const struct object_id *oid,\n> +\t\t\t      const struct object_id *compat_oid)\n> +{\n> +\tint inserted = 0;\n> +\n> +\tif (!should_use_loose_object_map(repo))\n> +\t\treturn 0;\n> +\n> +\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_compat, oid, compat_oid);\n> +\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_storage, compat_oid, oid);\n> +\tif (inserted)\n> +\t\treturn write_one_object(repo, oid, compat_oid);\n> +\treturn 0;\n> +}\n> +\n> +int repo_loose_object_map_oid(struct repository *repo,\n> +\t\t\t      const struct object_id *src,\n> +\t\t\t      const struct git_hash_algo *to,\n> +\t\t\t      struct object_id *dest)\n> +{\n> +\tstruct object_directory *dir;\n> +\tkh_oid_map_t *map;\n> +\tkhiter_t pos;\n> +\n> +\tfor (dir = repo->objects->odb; dir; dir = dir->next) {\n> +\t\tstruct loose_object_map *loose_map = dir->loose_map;\n> +\t\tif (!loose_map)\n> +\t\t\tcontinue;\n> +\t\tmap = (to == repo->compat_hash_algo) ?\n> +\t\t\tloose_map->to_compat :\n> +\t\t\tloose_map->to_storage;\n> +\t\tpos = kh_get_oid_map(map, *src);\n> +\t\tif (pos < kh_end(map)) {\n> +\t\t\toidcpy(dest, kh_value(map, pos));\n> +\t\t\treturn 0;\n> +\t\t}\n> +\t}\n> +\treturn -1;\n> +}\n> +\n> +void loose_object_map_clear(struct loose_object_map **map)\n\nNit: I'd rather call it `loose_object_map_release()`. `clear` typically\nindicates that we clear contents, but do not end up freeing the\ncontaining structure.\n\n> +{\n> +\tstruct loose_object_map *m = *map;\n> +\tstruct object_id *oid;\n> +\n> +\tif (!m)\n> +\t\treturn;\n> +\n> +\tkh_foreach_value(m->to_compat, oid, free(oid));\n> +\tkh_foreach_value(m->to_storage, oid, free(oid));\n> +\tkh_destroy_oid_map(m->to_compat);\n> +\tkh_destroy_oid_map(m->to_storage);\n> +\tfree(m);\n> +\t*map = NULL;\n> +}\n> diff --git a/loose.h b/loose.h\n> new file mode 100644\n> index 000000000000..2c2957072c5f\n> --- /dev/null\n> +++ b/loose.h\n> @@ -0,0 +1,22 @@\n> +#ifndef LOOSE_H\n> +#define LOOSE_H\n> +\n> +#include \"khash.h\"\n> +\n> +struct loose_object_map {\n> +\tkh_oid_map_t *to_compat;\n> +\tkh_oid_map_t *to_storage;\n> +};\n\nAny specific reason why you don't use `struct oidmap` here?\n\nPatrick\n\n> +void loose_object_map_init(struct loose_object_map **map);\n> +void loose_object_map_clear(struct loose_object_map **map);\n> +int repo_loose_object_map_oid(struct repository *repo,\n> +\t\t\t      const struct object_id *src,\n> +\t\t\t      const struct git_hash_algo *dest_algo,\n> +\t\t\t      struct object_id *dest);\n> +int repo_add_loose_object_map(struct repository *repo, const struct object_id *oid,\n> +\t\t\t      const struct object_id *compat_oid);\n> +int repo_read_loose_object_map(struct repository *repo);\n> +int repo_write_loose_object_map(struct repository *repo);\n> +\n> +#endif\n> diff --git a/object-file-convert.c b/object-file-convert.c\n> index 4777aba83636..1ec945eaa17f 100644\n> --- a/object-file-convert.c\n> +++ b/object-file-convert.c\n> @@ -4,6 +4,7 @@\n>  #include \"repository.h\"\n>  #include \"hash-ll.h\"\n>  #include \"object.h\"\n> +#include \"loose.h\"\n>  #include \"object-file-convert.h\"\n>  \n>  int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n> @@ -21,7 +22,18 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n>  \t\t\toidcpy(dest, src);\n>  \t\treturn 0;\n>  \t}\n> -\treturn -1;\n> +\tif (repo_loose_object_map_oid(repo, src, to, dest)) {\n> +\t\t/*\n> +\t\t * We may have loaded the object map at repo initialization but\n> +\t\t * another process (perhaps upstream of a pipe from us) may have\n> +\t\t * written a new object into the map.  If the object is missing,\n> +\t\t * let's reload the map to see if the object has appeared.\n> +\t\t */\n> +\t\trepo_read_loose_object_map(repo);\n> +\t\tif (repo_loose_object_map_oid(repo, src, to, dest))\n> +\t\t\treturn -1;\n> +\t}\n> +\treturn 0;\n>  }\n>  \n>  int convert_object_file(struct strbuf *outbuf,\n> diff --git a/object-store-ll.h b/object-store-ll.h\n> index 26a3895c821c..bc76d6bec80d 100644\n> --- a/object-store-ll.h\n> +++ b/object-store-ll.h\n> @@ -26,6 +26,9 @@ struct object_directory {\n>  \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n>  \tstruct oidtree *loose_objects_cache;\n>  \n> +\t/* Map between object IDs for loose objects. */\n> +\tstruct loose_object_map *loose_map;\n> +\n>  \t/*\n>  \t * This is a temporary object store created by the tmp_objdir\n>  \t * facility. Disable ref updates since the objects in the store\n> diff --git a/object.c b/object.c\n> index 2c61e4c86217..186a0a47c0fb 100644\n> --- a/object.c\n> +++ b/object.c\n> @@ -13,6 +13,7 @@\n>  #include \"alloc.h\"\n>  #include \"packfile.h\"\n>  #include \"commit-graph.h\"\n> +#include \"loose.h\"\n>  \n>  unsigned int get_max_object_index(void)\n>  {\n> @@ -540,6 +541,7 @@ void free_object_directory(struct object_directory *odb)\n>  {\n>  \tfree(odb->path);\n>  \todb_clear_loose_cache(odb);\n> +\tloose_object_map_clear(&odb->loose_map);\n>  \tfree(odb);\n>  }\n>  \n> diff --git a/repository.c b/repository.c\n> index 80252b79e93e..6214f61cf4e7 100644\n> --- a/repository.c\n> +++ b/repository.c\n> @@ -14,6 +14,7 @@\n>  #include \"read-cache-ll.h\"\n>  #include \"remote.h\"\n>  #include \"setup.h\"\n> +#include \"loose.h\"\n>  #include \"submodule-config.h\"\n>  #include \"sparse-index.h\"\n>  #include \"trace2.h\"\n> @@ -109,6 +110,8 @@ void repo_set_compat_hash_algo(struct repository *repo, int algo)\n>  \tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n>  \t\tBUG(\"hash_algo and compat_hash_algo match\");\n>  \trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n> +\tif (repo->compat_hash_algo)\n> +\t\trepo_read_loose_object_map(repo);\n>  }\n>  \n>  /*\n> @@ -201,6 +204,9 @@ int repo_init(struct repository *repo,\n>  \tif (worktree)\n>  \t\trepo_set_worktree(repo, worktree);\n>  \n> +\tif (repo->compat_hash_algo)\n> +\t\trepo_read_loose_object_map(repo);\n> +\n>  \tclear_repository_format(&format);\n>  \treturn 0;\n>  \n> -- \n> 2.41.0\n> \n"},{"id":"488729","messageId":"Zc3z4YWYybC0Xi2u@tanuki","threadId":"60273","inReplyTo":"20231002024034.2611-6-ebiederm@gmail.com","subject":"Re: [PATCH v2 06/30] loose: compatibilty short name support","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-02-15T11:22:09Z","receivedAt":"2024-02-15T11:22:12Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sun, Oct 01, 2023 at 09:40:10PM -0500, Eric W. Biederman wrote:\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> \n> Update loose_objects_cache when udpating the loose objects map.  This\n> oidtree is used to discover which oids are possibilities when\n> resolving short names, and it can support a mixture of sha1\n> and sha256 oids.\n> \n> With this any oid recorded objects/loose-objects-idx is usable\n> for resolving an oid to an object.\n> \n> To make this maintainable a helper insert_loose_map is factored\n> out of load_one_loose_object_map and repo_add_loose_object_map,\n> and then modified to also update the loose_objects_cache.\n> \n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n>  loose.c | 37 +++++++++++++++++++++++++------------\n>  1 file changed, 25 insertions(+), 12 deletions(-)\n> \n> diff --git a/loose.c b/loose.c\n> index 6ba73cc84dca..f6faa6216a08 100644\n> --- a/loose.c\n> +++ b/loose.c\n> @@ -7,6 +7,7 @@\n>  #include \"gettext.h\"\n>  #include \"loose.h\"\n>  #include \"lockfile.h\"\n> +#include \"oidtree.h\"\n>  \n>  static const char *loose_object_header = \"# loose-object-idx\\n\";\n>  \n> @@ -42,6 +43,21 @@ static int insert_oid_pair(kh_oid_map_t *map, const struct object_id *key, const\n>  \treturn 1;\n>  }\n>  \n> +static int insert_loose_map(struct object_directory *odb,\n> +\t\t\t    const struct object_id *oid,\n> +\t\t\t    const struct object_id *compat_oid)\n\nI think it would've been nice to fold this into the preceding patch\nalready. At least I wanted to propose adding such a function to avoid\nthe duplication down below.\n\nPatrick\n\n> +{\n> +\tstruct loose_object_map *map = odb->loose_map;\n> +\tint inserted = 0;\n> +\n> +\tinserted |= insert_oid_pair(map->to_compat, oid, compat_oid);\n> +\tinserted |= insert_oid_pair(map->to_storage, compat_oid, oid);\n> +\tif (inserted)\n> +\t\toidtree_insert(odb->loose_objects_cache, compat_oid);\n> +\n> +\treturn inserted;\n> +}\n> +\n>  static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n>  {\n>  \tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n> @@ -49,15 +65,14 @@ static int load_one_loose_object_map(struct repository *repo, struct object_dire\n>  \n>  \tif (!dir->loose_map)\n>  \t\tloose_object_map_init(&dir->loose_map);\n> +\tif (!dir->loose_objects_cache) {\n> +\t\tALLOC_ARRAY(dir->loose_objects_cache, 1);\n> +\t\toidtree_init(dir->loose_objects_cache);\n> +\t}\n>  \n> -\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n> -\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n> -\n> -\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n> -\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n> -\n> -\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n> -\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n> +\tinsert_loose_map(dir, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n> +\tinsert_loose_map(dir, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n> +\tinsert_loose_map(dir, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n>  \n>  \tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n>  \tfp = fopen(path.buf, \"rb\");\n> @@ -77,8 +92,7 @@ static int load_one_loose_object_map(struct repository *repo, struct object_dire\n>  \t\t    parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n>  \t\t    p != buf.buf + buf.len)\n>  \t\t\tgoto err;\n> -\t\tinsert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n> -\t\tinsert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n> +\t\tinsert_loose_map(dir, &oid, &compat_oid);\n>  \t}\n>  \n>  \tstrbuf_release(&buf);\n> @@ -197,8 +211,7 @@ int repo_add_loose_object_map(struct repository *repo, const struct object_id *o\n>  \tif (!should_use_loose_object_map(repo))\n>  \t\treturn 0;\n>  \n> -\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_compat, oid, compat_oid);\n> -\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_storage, compat_oid, oid);\n> +\tinserted = insert_loose_map(repo->objects->odb, oid, compat_oid);\n>  \tif (inserted)\n>  \t\treturn write_one_object(repo, oid, compat_oid);\n>  \treturn 0;\n> -- \n> 2.41.0\n> \n"},{"id":"488730","messageId":"Zc3z53gXllPafrFr@tanuki","threadId":"60273","inReplyTo":"20231002024034.2611-7-ebiederm@gmail.com","subject":"Re: [PATCH v2 07/30] object-file: update the loose object map when writing loose objects","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-02-15T11:22:15Z","receivedAt":"2024-02-15T11:22:19Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sun, Oct 01, 2023 at 09:40:11PM -0500, Eric W. Biederman wrote:\n> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> \n> To implement SHA1 compatibility on SHA256 repositories the loose\n> object map needs to be updated whenver a loose object is written.\n\nNot only when loose objects are written, but also when packfiles are\nwritten e.g. when accepting a push via git-receive-pack(1). Basically,\nwhenever an object gets written into the main object database.\n\nThis also brings up another interesting angle: how will this work in the\ncontext of alternate object directories? We have no control over new\nobjects being written into those, and thus the object mapping that we\nhave in our satellite repository that uses the alternate would be out of\ndate.\n\nI think this is another indicator that stacking might be the right way\nto go. Like that, the stack of object maps would be the main stack plus\nall stack of alternates concatenated. Finding a mapping would then have\nto go through all of these maps to find the desired object.\n\n> Updating the loose object map this way allows git to support\n> the old hash algorithm in constant time.\n\nAs mentioned before, appending objects is constant-time, but the reading\nside is unfortunately not. It's probably more something like `O(nlogn)`\nbecause we have to load all objects and add each of the objects into the\nmap, which I expect to be `O(logn)`. So the reading time isn't even\nlinear.\n\nPatrick\n\n> The functions write_loose_object, and stream_loose_object are\n> the only two functions that write to the loose object store.\n> \n> Update stream_loose_object to compute the compatibiilty hash, update\n> the loose object, and then call repo_add_loose_object_map to update\n> the loose object map.\n> \n> Update write_object_file_flags to convert the object into\n> it's compatibility encoding, hash the compatibility encoding,\n> write the object, and then update the loose object map.\n> \n> Update force_object_loose to lookup the hash of the compatibility\n> encoding, write the loose object, and then update the loose object\n> map.\n> \n> Update write_object_file_literally to convert the object into it's\n> compatibility hash encoding, hash the compatibility enconding, write\n> the object, and then update the loose object map, when the type string\n> is a known type.  For objects with an unknown type this results in a\n> partially broken repository, as the objects are not mapped.\n> \n> The point of write_object_file_literally is to generate a partially\n> broken repository for testing.  For testing skipping writing the loose\n> object map is much more useful than refusing to write the broken\n> object at all.\n> \n> Except that the loose objects are updated before the loose object map\n> I have not done any analysis to see how robust this scheme is in the\n> event of failure.\n> \n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n>  object-file.c | 113 ++++++++++++++++++++++++++++++++++++++++++--------\n>  1 file changed, 95 insertions(+), 18 deletions(-)\n> \n> diff --git a/object-file.c b/object-file.c\n> index 7dc0c4bfbba8..4e55f475b3b4 100644\n> --- a/object-file.c\n> +++ b/object-file.c\n> @@ -43,6 +43,8 @@\n>  #include \"setup.h\"\n>  #include \"submodule.h\"\n>  #include \"fsck.h\"\n> +#include \"loose.h\"\n> +#include \"object-file-convert.h\"\n>  \n>  /* The maximum size for an object header. */\n>  #define MAX_HEADER_LEN 32\n> @@ -1952,9 +1954,12 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n>  \t\t\t\t     const char *filename, unsigned flags,\n>  \t\t\t\t     git_zstream *stream,\n>  \t\t\t\t     unsigned char *buf, size_t buflen,\n> -\t\t\t\t     git_hash_ctx *c,\n> +\t\t\t\t     git_hash_ctx *c, git_hash_ctx *compat_c,\n>  \t\t\t\t     char *hdr, int hdrlen)\n>  {\n> +\tstruct repository *repo = the_repository;\n> +\tconst struct git_hash_algo *algo = repo->hash_algo;\n> +\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n>  \tint fd;\n>  \n>  \tfd = create_tmpfile(tmp_file, filename);\n> @@ -1974,14 +1979,18 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n>  \tgit_deflate_init(stream, zlib_compression_level);\n>  \tstream->next_out = buf;\n>  \tstream->avail_out = buflen;\n> -\tthe_hash_algo->init_fn(c);\n> +\talgo->init_fn(c);\n> +\tif (compat && compat_c)\n> +\t\tcompat->init_fn(compat_c);\n>  \n>  \t/*  Start to feed header to zlib stream */\n>  \tstream->next_in = (unsigned char *)hdr;\n>  \tstream->avail_in = hdrlen;\n>  \twhile (git_deflate(stream, 0) == Z_OK)\n>  \t\t; /* nothing */\n> -\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n> +\talgo->update_fn(c, hdr, hdrlen);\n> +\tif (compat && compat_c)\n> +\t\tcompat->update_fn(compat_c, hdr, hdrlen);\n>  \n>  \treturn fd;\n>  }\n> @@ -1990,16 +1999,21 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n>   * Common steps for the inner git_deflate() loop for writing loose\n>   * objects. Returns what git_deflate() returns.\n>   */\n> -static int write_loose_object_common(git_hash_ctx *c,\n> +static int write_loose_object_common(git_hash_ctx *c, git_hash_ctx *compat_c,\n>  \t\t\t\t     git_zstream *stream, const int flush,\n>  \t\t\t\t     unsigned char *in0, const int fd,\n>  \t\t\t\t     unsigned char *compressed,\n>  \t\t\t\t     const size_t compressed_len)\n>  {\n> +\tstruct repository *repo = the_repository;\n> +\tconst struct git_hash_algo *algo = repo->hash_algo;\n> +\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n>  \tint ret;\n>  \n>  \tret = git_deflate(stream, flush ? Z_FINISH : 0);\n> -\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n> +\talgo->update_fn(c, in0, stream->next_in - in0);\n> +\tif (compat && compat_c)\n> +\t\tcompat->update_fn(compat_c, in0, stream->next_in - in0);\n>  \tif (write_in_full(fd, compressed, stream->next_out - compressed) < 0)\n>  \t\tdie_errno(_(\"unable to write loose object file\"));\n>  \tstream->next_out = compressed;\n> @@ -2014,15 +2028,21 @@ static int write_loose_object_common(git_hash_ctx *c,\n>   * - End the compression of zlib stream.\n>   * - Get the calculated oid to \"oid\".\n>   */\n> -static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n> -\t\t\t\t   struct object_id *oid)\n> +static int end_loose_object_common(git_hash_ctx *c, git_hash_ctx *compat_c,\n> +\t\t\t\t   git_zstream *stream, struct object_id *oid,\n> +\t\t\t\t   struct object_id *compat_oid)\n>  {\n> +\tstruct repository *repo = the_repository;\n> +\tconst struct git_hash_algo *algo = repo->hash_algo;\n> +\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n>  \tint ret;\n>  \n>  \tret = git_deflate_end_gently(stream);\n>  \tif (ret != Z_OK)\n>  \t\treturn ret;\n> -\tthe_hash_algo->final_oid_fn(oid, c);\n> +\talgo->final_oid_fn(oid, c);\n> +\tif (compat && compat_c)\n> +\t\tcompat->final_oid_fn(compat_oid, compat_c);\n>  \n>  \treturn Z_OK;\n>  }\n> @@ -2046,7 +2066,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \n>  \tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n>  \t\t\t\t       &stream, compressed, sizeof(compressed),\n> -\t\t\t\t       &c, hdr, hdrlen);\n> +\t\t\t\t       &c, NULL, hdr, hdrlen);\n>  \tif (fd < 0)\n>  \t\treturn -1;\n>  \n> @@ -2056,14 +2076,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n>  \tdo {\n>  \t\tunsigned char *in0 = stream.next_in;\n>  \n> -\t\tret = write_loose_object_common(&c, &stream, 1, in0, fd,\n> +\t\tret = write_loose_object_common(&c, NULL, &stream, 1, in0, fd,\n>  \t\t\t\t\t\tcompressed, sizeof(compressed));\n>  \t} while (ret == Z_OK);\n>  \n>  \tif (ret != Z_STREAM_END)\n>  \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n>  \t\t    ret);\n> -\tret = end_loose_object_common(&c, &stream, &parano_oid);\n> +\tret = end_loose_object_common(&c, NULL, &stream, &parano_oid, NULL);\n>  \tif (ret != Z_OK)\n>  \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n>  \t\t    ret);\n> @@ -2108,10 +2128,12 @@ static int freshen_packed_object(const struct object_id *oid)\n>  int stream_loose_object(struct input_stream *in_stream, size_t len,\n>  \t\t\tstruct object_id *oid)\n>  {\n> +\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n> +\tstruct object_id compat_oid;\n>  \tint fd, ret, err = 0, flush = 0;\n>  \tunsigned char compressed[4096];\n>  \tgit_zstream stream;\n> -\tgit_hash_ctx c;\n> +\tgit_hash_ctx c, compat_c;\n>  \tstruct strbuf tmp_file = STRBUF_INIT;\n>  \tstruct strbuf filename = STRBUF_INIT;\n>  \tint dirlen;\n> @@ -2135,7 +2157,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n>  \t */\n>  \tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n>  \t\t\t\t       &stream, compressed, sizeof(compressed),\n> -\t\t\t\t       &c, hdr, hdrlen);\n> +\t\t\t\t       &c, &compat_c, hdr, hdrlen);\n>  \tif (fd < 0) {\n>  \t\terr = -1;\n>  \t\tgoto cleanup;\n> @@ -2153,7 +2175,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n>  \t\t\tif (in_stream->is_finished)\n>  \t\t\t\tflush = 1;\n>  \t\t}\n> -\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n> +\t\tret = write_loose_object_common(&c, &compat_c, &stream, flush, in0, fd,\n>  \t\t\t\t\t\tcompressed, sizeof(compressed));\n>  \t\t/*\n>  \t\t * Unlike write_loose_object(), we do not have the entire\n> @@ -2176,7 +2198,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n>  \t */\n>  \tif (ret != Z_STREAM_END)\n>  \t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n> -\tret = end_loose_object_common(&c, &stream, oid);\n> +\tret = end_loose_object_common(&c, &compat_c, &stream, oid, &compat_oid);\n>  \tif (ret != Z_OK)\n>  \t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n>  \tclose_loose_object(fd, tmp_file.buf);\n> @@ -2203,6 +2225,8 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n>  \t}\n>  \n>  \terr = finalize_object_file(tmp_file.buf, filename.buf);\n> +\tif (!err && compat)\n> +\t\terr = repo_add_loose_object_map(the_repository, oid, &compat_oid);\n>  cleanup:\n>  \tstrbuf_release(&tmp_file);\n>  \tstrbuf_release(&filename);\n> @@ -2213,17 +2237,38 @@ int write_object_file_flags(const void *buf, unsigned long len,\n>  \t\t\t    enum object_type type, struct object_id *oid,\n>  \t\t\t    unsigned flags)\n>  {\n> +\tstruct repository *repo = the_repository;\n> +\tconst struct git_hash_algo *algo = repo->hash_algo;\n> +\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n> +\tstruct object_id compat_oid;\n>  \tchar hdr[MAX_HEADER_LEN];\n>  \tint hdrlen = sizeof(hdr);\n>  \n> +\t/* Generate compat_oid */\n> +\tif (compat) {\n> +\t\tif (type == OBJ_BLOB)\n> +\t\t\thash_object_file(compat, buf, len, type, &compat_oid);\n> +\t\telse {\n> +\t\t\tstruct strbuf converted = STRBUF_INIT;\n> +\t\t\tconvert_object_file(&converted, algo, compat,\n> +\t\t\t\t\t    buf, len, type, 0);\n> +\t\t\thash_object_file(compat, converted.buf, converted.len,\n> +\t\t\t\t\t type, &compat_oid);\n> +\t\t\tstrbuf_release(&converted);\n> +\t\t}\n> +\t}\n> +\n>  \t/* Normally if we have it in the pack then we do not bother writing\n>  \t * it out into .git/objects/??/?{38} file.\n>  \t */\n> -\twrite_object_file_prepare(the_hash_algo, buf, len, type, oid, hdr,\n> -\t\t\t\t  &hdrlen);\n> +\twrite_object_file_prepare(algo, buf, len, type, oid, hdr, &hdrlen);\n>  \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n>  \t\treturn 0;\n> -\treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n> +\tif (write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags))\n> +\t\treturn -1;\n> +\tif (compat)\n> +\t\treturn repo_add_loose_object_map(repo, oid, &compat_oid);\n> +\treturn 0;\n>  }\n>  \n>  int write_object_file_literally(const void *buf, unsigned long len,\n> @@ -2231,7 +2276,27 @@ int write_object_file_literally(const void *buf, unsigned long len,\n>  \t\t\t\tunsigned flags)\n>  {\n>  \tchar *header;\n> +\tstruct repository *repo = the_repository;\n> +\tconst struct git_hash_algo *algo = repo->hash_algo;\n> +\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n> +\tstruct object_id compat_oid;\n>  \tint hdrlen, status = 0;\n> +\tint compat_type = -1;\n> +\n> +\tif (compat) {\n> +\t\tcompat_type = type_from_string_gently(type, -1, 1);\n> +\t\tif (compat_type == OBJ_BLOB)\n> +\t\t\thash_object_file(compat, buf, len, compat_type,\n> +\t\t\t\t\t &compat_oid);\n> +\t\telse if (compat_type != -1) {\n> +\t\t\tstruct strbuf converted = STRBUF_INIT;\n> +\t\t\tconvert_object_file(&converted, algo, compat,\n> +\t\t\t\t\t    buf, len, compat_type, 0);\n> +\t\t\thash_object_file(compat, converted.buf, converted.len,\n> +\t\t\t\t\t compat_type, &compat_oid);\n> +\t\t\tstrbuf_release(&converted);\n> +\t\t}\n> +\t}\n>  \n>  \t/* type string, SP, %lu of the length plus NUL must fit this */\n>  \thdrlen = strlen(type) + MAX_HEADER_LEN;\n> @@ -2244,6 +2309,8 @@ int write_object_file_literally(const void *buf, unsigned long len,\n>  \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n>  \t\tgoto cleanup;\n>  \tstatus = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n> +\tif (compat_type != -1)\n> +\t\treturn repo_add_loose_object_map(repo, oid, &compat_oid);\n>  \n>  cleanup:\n>  \tfree(header);\n> @@ -2252,9 +2319,12 @@ int write_object_file_literally(const void *buf, unsigned long len,\n>  \n>  int force_object_loose(const struct object_id *oid, time_t mtime)\n>  {\n> +\tstruct repository *repo = the_repository;\n> +\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n>  \tvoid *buf;\n>  \tunsigned long len;\n>  \tstruct object_info oi = OBJECT_INFO_INIT;\n> +\tstruct object_id compat_oid;\n>  \tenum object_type type;\n>  \tchar hdr[MAX_HEADER_LEN];\n>  \tint hdrlen;\n> @@ -2267,8 +2337,15 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n>  \toi.contentp = &buf;\n>  \tif (oid_object_info_extended(the_repository, oid, &oi, 0))\n>  \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n> +\tif (compat) {\n> +\t\tif (repo_oid_to_algop(repo, oid, compat, &compat_oid))\n> +\t\t\treturn error(_(\"cannot map object %s to %s\"),\n> +\t\t\t\t     oid_to_hex(oid), compat->name);\n> +\t}\n>  \thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n>  \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n> +\tif (!ret && compat)\n> +\t\tret = repo_add_loose_object_map(the_repository, oid, &compat_oid);\n>  \tfree(buf);\n>  \n>  \treturn ret;\n> -- \n> 2.41.0\n> \n"},{"id":"488732","messageId":"Zc31H77MyE1WLf_L@tanuki","threadId":"60273","inReplyTo":"878r8l929e.fsf@gmail.froward.int.ebiederm.org","subject":"Re: [PATCH v2 00/30] initial support for multiple hash functions","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2024-02-15T11:27:27Z","receivedAt":"2024-02-15T11:27:32Z","isPatch":true,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Sun, Oct 01, 2023 at 09:39:09PM -0500, Eric W. Biederman wrote:\n> \n> This addresses all of the known test failures from v1 of this set of\n> changes.  In particular I have reworked commit_tree_extended which\n> was flagged by smatch, -Werror=array-bounds, and the leak detector.\n> \n> One functional bug was fixed in repo_for_each_abbrev where it was\n> mistakenly displaying too many ambiguous oids.\n> \n> I am posting this so that people review and testing of this patchset\n> won't be distracted by the known and fixed issues.\n\nThanks! I've reviewed this patch series up to patch 7.\n\nI think the most important question mark I currently have is scalability\nof the proposed object mapping format. The complexity to load the object\nmappings is currently O(nlogn) and requires reading a file that is\neasily hundreds of megabytes or even gigabytes in size.\n\nI have a feeling that this needs to be addressed before such an object\nmapping would be feasible for production use, or otherwise it would\nincur too high a cost to be useful. I'm afraid that this will make the\nwhole series more complex to implement -- I'm sorry about that.\n\nI've added comments and ideas regarding this issue on patch 6 and 7. It\ncould totally be that I'm missing the obvious though, or that my ideas\nsuck. Please don't hesitate to point that out if that is the case.\n\nPatrick\n"},{"id":"488773","messageId":"owly1q9d9sau.fsf@fine.c.googlers.com","threadId":"60273","inReplyTo":"8734tumekr.fsf@gmail.froward.int.ebiederm.org","subject":"Re: [PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2024-02-16T00:16:57Z","receivedAt":"2024-02-16T00:16:59Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> Linus Arver <linusa@google.com> writes:\n>\n>> \"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n>>\n>>> From: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>>>\n>>> While looking at how to handle input of both SHA-1 and SHA-256 oids in\n>>> get_oid_with_context, I realized that the oid_array in\n>>> repo_for_each_abbrev might have more than one kind of oid stored in it\n>>> simultaneously.\n>>>\n>>> Update to oid_array_append to ensure that oids added to an oid array\n>>\n>> s/Update to/Update\n>>\n>>> always have an algorithm set.\n>>>\n>>> Update void_hashcmp to first verify two oids use the same hash algorithm\n>>> before comparing them to each other.\n>>>\n>>> With that oid-array should be safe to use with different kinds of\n>>\n>> s/oid-array/oid_array\n>>\n>>> oids simultaneously.\n>>>\n>>> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>>> ---\n>>>  oid-array.c | 12 ++++++++++--\n>>>  1 file changed, 10 insertions(+), 2 deletions(-)\n>>>\n>>> diff --git a/oid-array.c b/oid-array.c\n>>> index 8e4717746c31..1f36651754ed 100644\n>>> --- a/oid-array.c\n>>> +++ b/oid-array.c\n>>> @@ -6,12 +6,20 @@ void oid_array_append(struct oid_array *array, const struct object_id *oid)\n>>>  {\n>>>  \tALLOC_GROW(array->oid, array->nr + 1, array->alloc);\n>>>  \toidcpy(&array->oid[array->nr++], oid);\n>>> +\tif (!oid->algo)\n>>> +\t\toid_set_algo(&array->oid[array->nr - 1], the_hash_algo);\n>>\n>> How come we can't set oid->algo _before_ we call oidcpy()? It seems odd\n>> that we do the copy first and then modify what we just copied after the\n>> fact, instead of making sure that the thing we want to copy is correct\n>> before doing the copy.\n>>\n>> But also, if we are going to make the oid object \"correct\" before\n>> invoking oidcpy(), we might as well do it when the oid is first\n>> created/used (in the caller(s) of this function). I don't demand that\n>> you find/demonstrate where all these places are in this series (maybe\n>> that's a hairy problem to tackle?), but it seems cleaner in principle to\n>> fix the creation of oid objects instead of having to make oid users\n>> clean up their act like this after using them.\n>\n> There is a hairy problem here.\n>\n> I believe for reasons of simplicity when the algo field was added to\n> struct object_id it was allowed to be zero for users that don't\n> particularly care about the hash algorithm, and are happy to use the git\n> default hash algorithm.\n>\n> Me experience working on this set of change set showed that there\n> are oids without their algo set in all kinds of places in the tree.\n\nAh, I see. Thanks for the clarification.\n\n> I could not think of any sure way to go through the entire tree\n> and find those users, so I just made certain that oid array handled\n> that case.\n>\n> I need algo to be set properly in the oids in the oid array so I\n> could extend oid_array to hold multiple kinds of oids at the same\n> time.  To allow multiple kinds of oids at the same time void_hashcmp\n> needs a simple and reliable way to tell what the algorithm is of\n> any given oid.\n\nMakes sense.\n\n>>\n>>>  \tarray->sorted = 0;\n>>>  }\n>>>  \n>>> -static int void_hashcmp(const void *a, const void *b)\n>>> +static int void_hashcmp(const void *va, const void *vb)\n>>>  {\n>>> -\treturn oidcmp(a, b);\n>>> +\tconst struct object_id *a = va, *b = vb;\n>>> +\tint ret;\n>>> +\tif (a->algo == b->algo)\n>>> +\t\tret = oidcmp(a, b);\n>>\n>> This makes sense (per the commit message description) ...\n>>\n>>> +\telse\n>>> +\t\tret = a->algo > b->algo ? 1 : -1;\n>>\n>> ... but this seems to go against it? I thought you wanted to only ever\n>> compare hashes if they were of the same algo? It would be good to add a\n>> comment explaining why this is OK (we are no longer doing a byte-by-byte\n>> comparison of these oids any more here like we do for oidcmp() above\n>> which boils down to calling memcmp()).\n>\n> So the goal of this change is for oid_array to be able to hold hashes\n> from multiple algorithms at the same time.\n>\n> A key part of oid_array is oid_array_sort that allows functions such\n> as oid_array_lookup and oid_array_for_each_unique.\n>\n> To that end there needs to be a total ordering of oids.\n>\n> The function oidcmp is only defined when two oids are of the same\n> algorithm, it does not even test to detect the case of comparing\n> mismatched algorithms.\n>\n> Therefore to get a total ordering of oids.  I must use oidcmp\n> when the algorithm is the same (the common case) or simply order\n> the oids by algorithm when the algorithms are different.\n>\n>\n>\n> All of this is relevant to get_oid_with_context as get_oid_with_context\n> and it's helper functions contain the logic that determines what\n> we do when a hex string that is ambiguous is specified.\n>\n> In the ambiguous case all of the possible candidates are placed in\n> an oid_array, sorted and then displayed.\n>\n>\n> With a repository that can knows both the sha1 and the sha256 oid\n> of it's objects it is possible for a short oid to match both\n> some sha1 oids and some sha256 oids.\n\nThanks for the additional clarification. I think a lot of this could\nhave been added as comments or perhaps in the commit message. The \"short\nid can match both sha1 or sha256\" is a very real scenario we need to\nconsider in the sha1+sha256 world, indeed.\n\n>>> +\treturn ret;\n>>\n>> Also, in terms of style I think the \"early return for errors\" style\n>> would be simpler to read. I.e.\n>>\n>>     if (a->algo > b->algo)\n>>         return 1;\n>>\n>>     if (a->algo < b->algo)\n>>         return -1;\n>>\n>>     return oidcmd(a, b);\n>>\n>\n> I can see doing:\n> \tif (a->algo == b->algo)\n>         \treturn oidcmp(a,b);\n>\n> \tif (a->algo > b->algo)\n>         \treturn 1;\n>         else\n>         \treturn -1;\n>\n> Or even:\n> \tif (a->algo == b->algo)\n>         \treturn oidcmp(a,b);\n>\n> \treturn a->algo - b->algo;\n>\n> Although I suspect using subtraction is a bit too clever.\n\nAgreed.\n\n> Comparing for less than, and greater than, and then assuming\n> the values are equal hides what is important before calling\n> oidcmp which is that the algo values are equal.\n\nI would still prefer the \"early return for errors\" style even in this\ncase. This is because I much prefer to have the question \"how can things\ngo wrong?\" answered first, and dealt with, such that as I read\ntop-to-bottom I am left with less things I have to consider to\nunderstand the \"happy path\". WRT emphasizing the \"algos equal each\nother\" concern, a simple comment like\n\n     /* Only compare equal algorithms. */\n     return oidcmp(a, b);\n\nseems sufficient.\n\nBut, of course it is possible (perhaps even likely) that my preferred\nstyle is in the minority. Up to you. Thanks.\n"},{"id":"488781","messageId":"87le7lkoai.fsf@gmail.froward.int.ebiederm.org","threadId":"60273","inReplyTo":"owly1q9d9sau.fsf@fine.c.googlers.com","subject":"Re: [PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids","fromName":"Eric W. Biederman","fromEmail":"ebiederm@gmail.com","sentAt":"2024-02-16T04:48:05Z","receivedAt":"2024-02-16T04:48:08Z","isPatch":true,"sender":{"key":"ebiederm@gmail.com","avatar":null},"body":"Linus Arver <linusa@google.com> writes:\n\n>\n> I would still prefer the \"early return for errors\" style even in this\n> case. This is because I much prefer to have the question \"how can things\n> go wrong?\" answered first, and dealt with, such that as I read\n> top-to-bottom I am left with less things I have to consider to\n> understand the \"happy path\". WRT emphasizing the \"algos equal each\n> other\" concern, a simple comment like\n\nvoid_hashcmp is a function that reports how two elements are ordered.\nThere is no error handling.\n\nThere are in fact two cases that need to be handled with oid_array being\nable to contain more then one kind of hash at a time.\n\nThe two entries are ordered by oidcmp.\nThe two entries are ordered by hash algorithm.\n\nThe order that is maintained is first everything is ordered by hash\nalgorithm, then for entries of identical hash algorithm they are ordered\nby oidcmp.\n\nSo I don't think the concept of early return for errors can ever apply\nto any version of void_hashcmp.\n\n\nEric\n"},{"id":"488836","messageId":"owly8r3j97ge.fsf@fine.c.googlers.com","threadId":"60273","inReplyTo":"87le7lkoai.fsf@gmail.froward.int.ebiederm.org","subject":"Re: [PATCH v2 02/30] oid-array: teach oid-array to handle multiple kinds of oids","fromName":"Linus Arver","fromEmail":"linusa@google.com","sentAt":"2024-02-17T01:59:29Z","receivedAt":"2024-02-17T01:59:31Z","isPatch":true,"sender":{"key":"linus@ucla.edu","avatar":null},"body":"\"Eric W. Biederman\" <ebiederm@gmail.com> writes:\n\n> Linus Arver <linusa@google.com> writes:\n>\n>>\n>> I would still prefer the \"early return for errors\" style even in this\n>> case. This is because I much prefer to have the question \"how can things\n>> go wrong?\" answered first, and dealt with, such that as I read\n>> top-to-bottom I am left with less things I have to consider to\n>> understand the \"happy path\". WRT emphasizing the \"algos equal each\n>> other\" concern, a simple comment like\n>\n> void_hashcmp is a function that reports how two elements are ordered.\n> There is no error handling.\n>\n> There are in fact two cases that need to be handled with oid_array being\n> able to contain more then one kind of hash at a time.\n>\n> The two entries are ordered by oidcmp.\n> The two entries are ordered by hash algorithm.\n>\n> The order that is maintained is first everything is ordered by hash\n> algorithm, then for entries of identical hash algorithm they are ordered\n> by oidcmp.\n>\n> So I don't think the concept of early return for errors can ever apply\n> to any version of void_hashcmp.\n\nAh, yes. It appears that I simply read the code wrong. Thank you for the\nclarification, and I agree with your analysis and preferred style.\n\nSorry for the noise.\n"}]}