{"thread":{"id":"60212","subject":"[RFC][PATCH 0/32] SHA256 and SHA1 interoperability","startedAt":"2023-09-08T23:06:29Z","lastAt":"2023-10-05T18:15:23Z","messageCount":59,"participants":["Eric W. Biederman","brian m. carlson","Junio C Hamano","Oswald Buddenhagen","Taylor Blau"],"isPatch":true,"patchVersion":1,"patchTotal":32},"messages":[{"id":"481582","messageId":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":null,"subject":"[RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:05:52Z","receivedAt":"2023-09-08T23:06:29Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\nI would like to see the SHA256 transition happen so I started playing\nwith the k2204-transition-interop branch of brian m. carlson's tree.\n\nBefore I go farther I need to some other folks to look at this and see\nif this is a general direction that the git project can stand.\n\nThis patchset is not complete it does not implement converting a\nreceived pack of the compatibility hash into the hash function of the\nrepository, nor have I written any automated tests.  Both need to happen\nbefore this is finalized.\n\nThat said I think I have working implementations of all of the\ninteresting cases.  In particular I have \"git index-pack\" computing the\ncompatibility hash of every object in a pack file, and I can tell you\nthe sha256 of every sha1 in the git://git.kernel.org/pub/scm/git/git.git\n\nTo get there I have tweaked the transition plan a little.\n\nSo far I have just aimed for code that works, so there is doubtless\nroom for improvement.  My hope is that I have implemented enough\nthat people can play with this, and that people can see all of the\nweird little details that need to be taken care of to make this work.\n\nWhat do everyone else think?  Does this direction look plausible?\n\nEric W. Biederman (24):\n      doc hash-file-transition: A map file for mapping between sha1 and sha256\n      doc hash-function-transition: Replace compatObjectFormat with compatMap\n      object-file-convert:  Stubs for converting from one object format to another\n      object-name: Initial support for ^{sha1} and ^{sha256}\n      repository: add a compatibility hash algorithm\n      loose: Compatibilty short name support\n      object-file: Update the loose object map when writing loose objects\n      bulk-checkin: Only accept blobs\n      pack: Communicate the compat_oid through struct pack_idx_entry\n      object-file: Add a compat_oid_in parameter to write_object_file_flags\n      object: Factor out parse_mode out of fast-import and tree-walk into in object.h\n      builtin/cat-file:  Let the oid determine the output algorithm\n      tree-walk: init_tree_desc take an oid to get the hash algorithm\n      object-file: Handle compat objects in check_object_signature\n      builtin/ls-tree: Let the oid determine the output algorithm\n      builtin/pack-objects:  Communicate the compatibility hash through struct pack_idx_entry\n      pack-compat-map:  Add support for .compat files of a packfile\n      object-file-convert: Implement convert_object_file_{begin,step,end}\n      builtin/fast-import: compute compatibility hashs for imported objects\n      builtin/index-pack:  Add a simple oid index\n      builtin/index-pack:  Compute the compatibility hash\n      builtin/index-pack: Make the stack in compute_compat_oid explicit\n      unpack-objects: Update to compute and write the compatibility hashes\n      object-file-convert: Implement repo_submodule_oid_to_algop\n\nbrian m. carlson (8):\n      repository: Implement core.compatMap\n      loose: add a mapping between SHA-1 and SHA-256 for loose objects\n      bulk-checkin: hash object with compatibility algorithm\n      commit: write commits for both hashes\n      cache: add a function to read an OID of a specific algorithm\n      object-file-convert: add a function to convert trees between algorithms\n      object-file-convert: convert commit objects when writing\n      object-file-convert: convert tag commits when writing\n\n\n Documentation/config/core.txt                      |   6 +\n .../technical/hash-function-transition.txt         |  56 ++-\n Makefile                                           |   4 +\n archive.c                                          |   3 +-\n builtin.h                                          |   1 +\n builtin/am.c                                       |   6 +-\n builtin/cat-file.c                                 |   8 +-\n builtin/checkout.c                                 |   8 +-\n builtin/clone.c                                    |   2 +-\n builtin/commit.c                                   |   2 +-\n builtin/fast-import.c                              | 110 +++--\n builtin/grep.c                                     |   8 +-\n builtin/index-pack.c                               | 441 ++++++++++++++++++++-\n builtin/ls-tree.c                                  |   5 +-\n builtin/merge.c                                    |   3 +-\n builtin/pack-objects.c                             |  13 +-\n builtin/read-tree.c                                |   2 +-\n builtin/show-compat-map.c                          | 139 +++++++\n builtin/stash.c                                    |   5 +-\n builtin/unpack-objects.c                           |  14 +-\n bulk-checkin.c                                     |  55 ++-\n bulk-checkin.h                                     |   6 +-\n cache-tree.c                                       |   4 +-\n commit.c                                           | 176 +++++---\n commit.h                                           |   1 +\n delta-islands.c                                    |   2 +-\n diff-lib.c                                         |   2 +-\n fsck.c                                             |   6 +-\n git.c                                              |   1 +\n hash-ll.h                                          |   3 +\n hash.h                                             |   9 +-\n http-push.c                                        |   2 +-\n list-objects.c                                     |   2 +-\n loose.c                                            | 256 ++++++++++++\n loose.h                                            |  20 +\n match-trees.c                                      |   4 +-\n merge-ort.c                                        |  11 +-\n merge-recursive.c                                  |   2 +-\n merge.c                                            |   3 +-\n object-file-convert.c                              | 366 +++++++++++++++++\n object-file-convert.h                              |  50 +++\n object-file.c                                      | 197 +++++++--\n object-name.c                                      |  77 +++-\n object-store-ll.h                                  |  13 +-\n object.c                                           |   2 +\n object.h                                           |  18 +\n pack-bitmap-write.c                                |   2 +-\n pack-compat-map.c                                  | 334 ++++++++++++++++\n pack-compat-map.h                                  |  27 ++\n pack-write.c                                       | 158 ++++++++\n pack.h                                             |   1 +\n packfile.c                                         |  15 +-\n reflog.c                                           |   2 +-\n repository.c                                       |  17 +\n repository.h                                       |   4 +\n revision.c                                         |   4 +-\n setup.c                                            |   5 +\n setup.h                                            |   1 +\n tree-walk.c                                        |  58 ++-\n tree-walk.h                                        |   7 +-\n tree.c                                             |   2 +-\n walker.c                                           |   2 +-\n 62 files changed, 2525 insertions(+), 238 deletions(-)\n\nEric\n"},{"id":"481583","messageId":"20230908231049.2035003-2-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 02/32] doc hash-function-transition: Replace compatObjectFormat with compatMap","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:19Z","receivedAt":"2023-09-08T23:11:34Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Ir makes a lot of sense for the hash algorithm that determines how all\nof the objects in the repostiory be an extension so that versions of\ngit that don't know about it won't even try.\n\nFor implementing the compatiblity maps that really is not the case.\nAn version of git that does not recognizes the won't care and continue\nto use the repository as is.  The mapping functionality simply won't be\npresent.\n\nSimilarly if all of the objects are not mapped this could cause\nsome practical difficulties but it will not cause anything to perform\nthe wrong actions to the repository.  Some commands just won't work.\nIn the worst case all that needs to happen is for the compatibilty\nmaps to be rebuilt.\n\nSo let's use an option that forces unnecessary breakage of existing\ntools.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n .../technical/hash-function-transition.txt       | 16 ++++++++--------\n 1 file changed, 8 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\nindex 4b937480848a..10572c5794f9 100644\n--- a/Documentation/technical/hash-function-transition.txt\n+++ b/Documentation/technical/hash-function-transition.txt\n@@ -148,14 +148,14 @@ Detailed Design\n Repository format extension\n ~~~~~~~~~~~~~~~~~~~~~~~~~~~\n A SHA-256 repository uses repository format version `1` (see\n-Documentation/technical/repository-version.txt) with extensions\n-`objectFormat` and `compatObjectFormat`:\n+Documentation/technical/repository-version.txt) with the extension\n+`objectFormat`, and an optional core.compatMap configuration.\n \n \t[core]\n \t\trepositoryFormatVersion = 1\n+\t\tcompatMap = on\n \t[extensions]\n \t\tobjectFormat = sha256\n-\t\tcompatObjectFormat = sha1\n \n The combination of setting `core.repositoryFormatVersion=1` and\n populating `extensions.*` ensures that all versions of Git later than\n@@ -682,7 +682,7 @@ Some initial steps can be implemented independently of one another:\n - adding support for the PSRC field and safer object pruning\n \n The first user-visible change is the introduction of the objectFormat\n-extension (without compatObjectFormat). This requires:\n+extension. This requires:\n \n - teaching fsck about this mode of operation\n - using the hash function API (vtable) when computing object names\n@@ -690,7 +690,7 @@ extension (without compatObjectFormat). This requires:\n - rejecting attempts to fetch from or push to an incompatible\n   repository\n \n-Next comes introduction of compatObjectFormat:\n+Next comes introduction of compatMap:\n \n - implementing the loose-object-idx\n - translating object names between object formats\n@@ -724,9 +724,9 @@ Over time projects would encourage their users to adopt the \"early\n transition\" and then \"late transition\" modes to take advantage of the\n new, more futureproof SHA-256 object names.\n \n-When objectFormat and compatObjectFormat are both set, commands\n-generating signatures would generate both SHA-1 and SHA-256 signatures\n-by default to support both new and old users.\n+When objectFormat and compatMap are both set, commands generating\n+signatures would generate both SHA-1 and SHA-256 signatures by default\n+to support both new and old users.\n \n In projects using SHA-256 heavily, users could be encouraged to adopt\n the \"post-transition\" mode to avoid accidentally making implicit use\n-- \n2.41.0\n\n"},{"id":"481584","messageId":"20230908231049.2035003-4-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 04/32] object-name: Initial support for ^{sha1} and ^{sha256}","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:21Z","receivedAt":"2023-09-08T23:11:44Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"In Documentation/technical/hash-function-transition.txt it suggests\nsupporting references like abac87a^{sha1} and f787cac^{sha256}.\n\nThis changes goes a step farther and supports a short oid in any\nalgorithm, and to just ensures enough of the oid is present to\ndisambiguate between all possible oids in any algorithm.\n\nSupport for suffixes of ^{sha1} and ^{sha256} is implemented as it is\neasy, and can be handy for testing.  To support this mode of operation\ntwo flags are added: GET_OID_SHA1, and GET_OID_SHA256.\n\nBy default when an oid is specified in an algorithm that does not\nmatch the algorithm of the repository, the oid is translated to the\noid that matches the hash algorithm of the repository.  This ensures\noids that don't match the repository hash algorithm can be used\neverywhere oids can currently be used.\n\nA new flag is added GET_OID_UNTRANSLATED that suppresses the\ntranslation of an oid into the repositories hash algorithm.\nThis is useful for testing and raw tools like git cat-file.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n hash-ll.h     |  3 +++\n object-name.c | 59 +++++++++++++++++++++++++++++++++++++++++++++------\n 2 files changed, 55 insertions(+), 7 deletions(-)\n\ndiff --git a/hash-ll.h b/hash-ll.h\nindex 10d84cc20888..2a4f72d70c3f 100644\n--- a/hash-ll.h\n+++ b/hash-ll.h\n@@ -143,8 +143,11 @@ struct object_id {\n #define GET_OID_BLOB             040\n #define GET_OID_FOLLOW_SYMLINKS 0100\n #define GET_OID_RECORD_PATH     0200\n+#define GET_OID_SHA1           01000\n+#define GET_OID_SHA256         02000\n #define GET_OID_ONLY_TO_DIE    04000\n #define GET_OID_REQUIRE_PATH  010000\n+#define GET_OID_UNTRANSLATED  020000\n \n #define GET_OID_DISAMBIGUATORS \\\n \t(GET_OID_COMMIT | GET_OID_COMMITTISH | \\\ndiff --git a/object-name.c b/object-name.c\nindex 0bfa29dbbfe9..ebe87f5c4fdd 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -25,6 +25,7 @@\n #include \"midx.h\"\n #include \"commit-reach.h\"\n #include \"date.h\"\n+#include \"object-file-convert.h\"\n \n static int get_oid_oneline(struct repository *r, const char *, struct object_id *, struct commit_list *);\n \n@@ -32,6 +33,7 @@ typedef int (*disambiguate_hint_fn)(struct repository *, const struct object_id\n \n struct disambiguate_state {\n \tint len; /* length of prefix in hex chars */\n+\tint algo;\n \tchar hex_pfx[GIT_MAX_HEXSZ + 1];\n \tstruct object_id bin_pfx;\n \n@@ -49,6 +51,10 @@ struct disambiguate_state {\n \n static void update_candidates(struct disambiguate_state *ds, const struct object_id *current)\n {\n+\t/* Is the oid encoded in the desired algo? */\n+\tif (ds->algo && (current->algo != ds->algo))\n+\t\treturn;\n+\n \tif (ds->always_call_fn) {\n \t\tds->ambiguous = ds->fn(ds->repo, current, ds->cb_data) ? 1 : 0;\n \t\treturn;\n@@ -134,6 +140,8 @@ static void unique_in_midx(struct multi_pack_index *m,\n {\n \tuint32_t num, i, first = 0;\n \tconst struct object_id *current = NULL;\n+\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n+\t\tds->repo->hash_algo->hexsz : ds->len;\n \tnum = m->num_objects;\n \n \tif (!num)\n@@ -149,7 +157,7 @@ static void unique_in_midx(struct multi_pack_index *m,\n \tfor (i = first; i < num && !ds->ambiguous; i++) {\n \t\tstruct object_id oid;\n \t\tcurrent = nth_midxed_object_oid(&oid, m, i);\n-\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, current->hash))\n+\t\tif (!match_hash(len, ds->bin_pfx.hash, current->hash))\n \t\t\tbreak;\n \t\tupdate_candidates(ds, current);\n \t}\n@@ -159,6 +167,8 @@ static void unique_in_pack(struct packed_git *p,\n \t\t\t   struct disambiguate_state *ds)\n {\n \tuint32_t num, i, first = 0;\n+\tint len = ds->len > ds->repo->hash_algo->hexsz ?\n+\t\tds->repo->hash_algo->hexsz : ds->len;\n \n \tif (p->multi_pack_index)\n \t\treturn;\n@@ -177,7 +187,7 @@ static void unique_in_pack(struct packed_git *p,\n \tfor (i = first; i < num && !ds->ambiguous; i++) {\n \t\tstruct object_id oid;\n \t\tnth_packed_object_id(&oid, p, i);\n-\t\tif (!match_hash(ds->len, ds->bin_pfx.hash, oid.hash))\n+\t\tif (!match_hash(len, ds->bin_pfx.hash, oid.hash))\n \t\t\tbreak;\n \t\tupdate_candidates(ds, &oid);\n \t}\n@@ -188,6 +198,10 @@ static void find_short_packed_object(struct disambiguate_state *ds)\n \tstruct multi_pack_index *m;\n \tstruct packed_git *p;\n \n+\t/* Skip, unless oids from the repository algorithm are wanted */\n+\tif (ds->algo && (&hash_algos[ds->algo] != ds->repo->hash_algo))\n+\t\treturn;\n+\n \tfor (m = get_multi_pack_index(ds->repo); m && !ds->ambiguous;\n \t     m = m->next)\n \t\tunique_in_midx(m, ds);\n@@ -330,7 +344,7 @@ static int init_object_disambiguation(struct repository *r,\n {\n \tint i;\n \n-\tif (len < MINIMUM_ABBREV || len > the_hash_algo->hexsz)\n+\tif (len < MINIMUM_ABBREV || len > GIT_MAX_HEXSZ)\n \t\treturn -1;\n \n \tmemset(ds, 0, sizeof(*ds));\n@@ -357,6 +371,7 @@ static int init_object_disambiguation(struct repository *r,\n \tds->len = len;\n \tds->hex_pfx[len] = '\\0';\n \tds->repo = r;\n+\tds->algo = GIT_HASH_UNKNOWN;\n \tprepare_alt_odb(r);\n \treturn 0;\n }\n@@ -491,9 +506,10 @@ static int repo_collect_ambiguous(struct repository *r UNUSED,\n \treturn collect_ambiguous(oid, data);\n }\n \n-static int sort_ambiguous(const void *a, const void *b, void *ctx)\n+static int sort_ambiguous(const void *va, const void *vb, void *ctx)\n {\n \tstruct repository *sort_ambiguous_repo = ctx;\n+\tconst struct object_id *a = va, *b = vb;\n \tint a_type = oid_object_info(sort_ambiguous_repo, a, NULL);\n \tint b_type = oid_object_info(sort_ambiguous_repo, b, NULL);\n \tint a_type_sort;\n@@ -503,8 +519,13 @@ static int sort_ambiguous(const void *a, const void *b, void *ctx)\n \t * Sorts by hash within the same object type, just as\n \t * oid_array_for_each_unique() would do.\n \t */\n-\tif (a_type == b_type)\n-\t\treturn oidcmp(a, b);\n+\tif (a_type == b_type) {\n+\t\t/* Is the hash algorithm the same? */\n+\t\tif (a->algo == b->algo)\n+\t\t\treturn oidcmp(a, b);\n+\t\telse\n+\t\t\treturn a->algo < b->algo ? -1 : 1;\n+\t}\n \n \t/*\n \t * Between object types show tags, then commits, and finally\n@@ -553,6 +574,11 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \telse\n \t\tds.fn = default_disambiguate_hint;\n \n+\tif (flags & GET_OID_SHA1)\n+\t\tds.algo = GIT_HASH_SHA1;\n+\telse if (flags & GET_OID_SHA256)\n+\t\tds.algo = GIT_HASH_SHA256;\n+\n \tfind_short_object_filename(&ds);\n \tfind_short_packed_object(&ds);\n \tstatus = finish_object_disambiguation(&ds, oid);\n@@ -606,6 +632,15 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t\tstrbuf_release(&out.sb);\n \t}\n \n+\t/* Ensure oid->algo is set */\n+\tif (oid->algo == GIT_HASH_UNKNOWN)\n+\t\toid->algo = hash_algo_by_ptr(r->hash_algo);\n+\n+\t/* Return oids using the repository's hash algorithm */\n+\tif ((&hash_algos[oid->algo] != r->hash_algo) &&\n+\t    !(flags & GET_OID_UNTRANSLATED))\n+\t\trepo_oid_to_algop(r, oid, r->hash_algo, oid);\n+\n \treturn status;\n }\n \n@@ -787,10 +822,12 @@ void strbuf_add_unique_abbrev(struct strbuf *sb, const struct object_id *oid,\n int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \t\t\t      const struct object_id *oid, int len)\n {\n+\tconst struct git_hash_algo *algo =\n+\t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n \tstruct disambiguate_state ds;\n \tstruct min_abbrev_data mad;\n \tstruct object_id oid_ret;\n-\tconst unsigned hexsz = r->hash_algo->hexsz;\n+\tconst unsigned hexsz = algo->hexsz;\n \n \tif (len < 0) {\n \t\tunsigned long count = repo_approximate_object_count(r);\n@@ -1158,6 +1195,14 @@ static int peel_onion(struct repository *r, const char *name, int len,\n \t\treturn -1;\n \n \tsp++; /* beginning of type name, or closing brace for empty */\n+\n+\tif (starts_with(sp, \"sha1}\"))\n+\t\treturn get_short_oid(r, name, len - 7, oid,\n+\t\t\t\t     lookup_flags | GET_OID_SHA1);\n+\telse if (starts_with(sp, \"sha256\"))\n+\t\treturn get_short_oid(r, name, len - 9, oid,\n+\t\t\t\t     lookup_flags | GET_OID_SHA256);\n+\n \tif (starts_with(sp, \"commit}\"))\n \t\texpected_type = OBJ_COMMIT;\n \telse if (starts_with(sp, \"tag}\"))\n-- \n2.41.0\n\n"},{"id":"481585","messageId":"20230908231049.2035003-6-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 06/32] repository: Implement core.compatMap","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:23Z","receivedAt":"2023-09-08T23:11:46Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nAdd a configuration option to enable updating and reading from\ncompatibility hash maps when git accesses the reposotiry.\n\nAdd a helper function repo_enable_compat_map that when passed false\ndisables the compatiblily hash algorithm and when passed true computes\nthe compatibilty hash algorithm and sets \"repo->compat_hash_algo\".\n\nFor now the option is limited to being specified in \".git/config\".\nPerhaps in the future we can allow specifying it in \".gitconfig\" as\nwell.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Documentation/config/core.txt |  6 ++++++\n repository.c                  | 11 +++++++++++\n repository.h                  |  1 +\n setup.c                       |  5 +++++\n setup.h                       |  1 +\n 5 files changed, 24 insertions(+)\n\ndiff --git a/Documentation/config/core.txt b/Documentation/config/core.txt\nindex dfbdaf00b8bc..a9eb2006cc32 100644\n--- a/Documentation/config/core.txt\n+++ b/Documentation/config/core.txt\n@@ -736,3 +736,9 @@ core.abbrev::\n \tIf set to \"no\", no abbreviation is made and the object names\n \tare shown in their full length.\n \tThe minimum length is 4.\n+\n+core.compatMap::\n+\tEnables the use of a compat map to recored the hash in the\n+\tother object format.  This allows repositories in different\n+\tobjects formats to interoperate.  It allows looking up old oids\n+\tin a repository that has been converted from sha1 to sha256.\ndiff --git a/repository.c b/repository.c\nindex a7679ceeaa45..de620d82bfc6 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -104,6 +104,16 @@ void repo_set_hash_algo(struct repository *repo, int hash_algo)\n \trepo->hash_algo = &hash_algos[hash_algo];\n }\n \n+void repo_enable_compat_map(struct repository *repo, int enable_compat)\n+{\n+\tconst struct git_hash_algo *other_algo =\n+\t\t&hash_algos[(hash_algo_by_ptr(repo->hash_algo) == GIT_HASH_SHA1) ?\n+\t\t\tGIT_HASH_SHA256 :\n+\t\t\tGIT_HASH_SHA1];\n+\n+\trepo->compat_hash_algo = enable_compat ? other_algo : NULL;\n+}\n+\n /*\n  * Attempt to resolve and set the provided 'gitdir' for repository 'repo'.\n  * Return 0 upon success and a non-zero value upon failure.\n@@ -184,6 +194,7 @@ int repo_init(struct repository *repo,\n \t\tgoto error;\n \n \trepo_set_hash_algo(repo, format.hash_algo);\n+\trepo_enable_compat_map(repo, format.use_compat_map);\n \trepo->repository_format_worktree_config = format.worktree_config;\n \n \t/* take ownership of format.partial_clone */\ndiff --git a/repository.h b/repository.h\nindex 6c4130f0c36e..03cadf6d9a98 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -202,6 +202,7 @@ void repo_set_gitdir(struct repository *repo, const char *root,\n \t\t     const struct set_gitdir_args *extra_args);\n void repo_set_worktree(struct repository *repo, const char *path);\n void repo_set_hash_algo(struct repository *repo, int algo);\n+void repo_enable_compat_map(struct repository *repo, int enable_compat);\n void initialize_the_repository(void);\n RESULT_MUST_BE_USED\n int repo_init(struct repository *r, const char *gitdir, const char *worktree);\ndiff --git a/setup.c b/setup.c\nindex 18927a847b86..b4d32bd820f1 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -623,6 +623,8 @@ static int check_repo_format(const char *var, const char *value,\n \t\t\treturn 0;\n \t\t}\n \t}\n+\telse if (strcmp(var, \"core.compatmap\") == 0)\n+\t\tdata->use_compat_map = git_config_bool(var, value);\n \n \treturn read_worktree_config(var, value, ctx, vdata);\n }\n@@ -1564,8 +1566,10 @@ const char *setup_git_directory_gently(int *nongit_ok)\n \t\t}\n \t\tif (startup_info->have_repository) {\n \t\t\trepo_set_hash_algo(the_repository, repo_fmt.hash_algo);\n+\t\t\trepo_enable_compat_map(the_repository, repo_fmt.use_compat_map);\n \t\t\tthe_repository->repository_format_worktree_config =\n \t\t\t\trepo_fmt.worktree_config;\n+\n \t\t\t/* take ownership of repo_fmt.partial_clone */\n \t\t\tthe_repository->repository_format_partial_clone =\n \t\t\t\trepo_fmt.partial_clone;\n@@ -1657,6 +1661,7 @@ void check_repository_format(struct repository_format *fmt)\n \tcheck_repository_format_gently(get_git_dir(), fmt, NULL);\n \tstartup_info->have_repository = 1;\n \trepo_set_hash_algo(the_repository, fmt->hash_algo);\n+\trepo_enable_compat_map(the_repository, fmt->use_compat_map);\n \tthe_repository->repository_format_worktree_config =\n \t\tfmt->worktree_config;\n \tthe_repository->repository_format_partial_clone =\ndiff --git a/setup.h b/setup.h\nindex 58fd2605dd26..afa05b2b64f3 100644\n--- a/setup.h\n+++ b/setup.h\n@@ -86,6 +86,7 @@ struct repository_format {\n \tint worktree_config;\n \tint is_bare;\n \tint hash_algo;\n+\tint use_compat_map;\n \tint sparse_index;\n \tchar *work_tree;\n \tstruct string_list unknown_extensions;\n-- \n2.41.0\n\n"},{"id":"481586","messageId":"20230908231049.2035003-7-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 07/32] loose: add a mapping between SHA-1 and SHA-256 for loose objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:24Z","receivedAt":"2023-09-08T23:11:52Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nAs part of the transition plan, we'd like to add a file in the .git\ndirectory that maps loose objects between SHA-1 and SHA-256.  Let's\nimplement the specification in the transition plan and store this data\non a per-repository basis in struct repository.\n\n****\n- split repo_object_map between repo_loose_object_map_oid and\n  repo_oid_to_algop.\n- Verified the loose_map is set in repo_loose_object_map_oid\n\n-- EWB\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n Makefile              |   1 +\n loose.c               | 243 ++++++++++++++++++++++++++++++++++++++++++\n loose.h               |  20 ++++\n object-file-convert.c |  14 ++-\n object-store-ll.h     |   3 +\n object.c              |   2 +\n repository.c          |   6 ++\n 7 files changed, 288 insertions(+), 1 deletion(-)\n create mode 100644 loose.c\n create mode 100644 loose.h\n\ndiff --git a/Makefile b/Makefile\nindex f7e824f25cda..3c18664def9a 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1053,6 +1053,7 @@ LIB_OBJS += list-objects-filter.o\n LIB_OBJS += list-objects.o\n LIB_OBJS += lockfile.o\n LIB_OBJS += log-tree.o\n+LIB_OBJS += loose.o\n LIB_OBJS += ls-refs.o\n LIB_OBJS += mailinfo.o\n LIB_OBJS += mailmap.o\ndiff --git a/loose.c b/loose.c\nnew file mode 100644\nindex 000000000000..8ddb7112a541\n--- /dev/null\n+++ b/loose.c\n@@ -0,0 +1,243 @@\n+#include \"git-compat-util.h\"\n+#include \"hash.h\"\n+#include \"path.h\"\n+#include \"object-store.h\"\n+#include \"hex.h\"\n+#include \"wrapper.h\"\n+#include \"gettext.h\"\n+#include \"loose.h\"\n+#include \"lockfile.h\"\n+\n+static const char *loose_object_header = \"# loose-object-idx\\n\";\n+\n+static inline int should_use_loose_object_map(struct repository *repo)\n+{\n+\treturn repo->compat_hash_algo && repo->gitdir;\n+}\n+\n+void loose_object_map_init(struct loose_object_map **map)\n+{\n+\tstruct loose_object_map *m;\n+\tm = xmalloc(sizeof(**map));\n+\tm->to_compat = kh_init_oid_map();\n+\tm->to_storage = kh_init_oid_map();\n+\t*map = m;\n+}\n+\n+static int insert_oid_pair(kh_oid_map_t *map, const struct object_id *key, const struct object_id *value)\n+{\n+\tkhiter_t pos;\n+\tint ret;\n+\tstruct object_id *stored;\n+\n+\tpos = kh_put_oid_map(map, *key, &ret);\n+\n+\t/* This item already exists in the map. */\n+\tif (ret == 0)\n+\t\treturn 0;\n+\n+\tstored = xmalloc(sizeof(*stored));\n+\toidcpy(stored, value);\n+\tkh_value(map, pos) = stored;\n+\treturn 1;\n+}\n+\n+static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n+{\n+\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\tFILE *fp;\n+\n+\tif (!dir->loose_map)\n+\t\tloose_object_map_init(&dir->loose_map);\n+\n+\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n+\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n+\n+\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n+\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n+\n+\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n+\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n+\n+\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n+\tfp = fopen(path.buf, \"rb\");\n+\tif (!fp)\n+\t\treturn 0;\n+\n+\terrno = 0;\n+\tif (strbuf_getwholeline(&buf, fp, '\\n') || strcmp(buf.buf, loose_object_header))\n+\t\tgoto err;\n+\twhile (!strbuf_getline_lf(&buf, fp)) {\n+\t\tconst char *p;\n+\t\tstruct object_id oid, compat_oid;\n+\t\tif (parse_oid_hex_algop(buf.buf, &oid, &p, repo->hash_algo) ||\n+\t\t    *p++ != ' ' ||\n+\t\t    parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n+\t\t    p != buf.buf + buf.len)\n+\t\t\tgoto err;\n+\t\tinsert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n+\t\tinsert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n+\t}\n+\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn errno ? -1 : 0;\n+err:\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn -1;\n+}\n+\n+int repo_read_loose_object_map(struct repository *repo)\n+{\n+\tstruct object_directory *dir;\n+\n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n+\tprepare_alt_odb(repo);\n+\n+\tfor (dir = repo->objects->odb; dir; dir = dir->next) {\n+\t\tif (load_one_loose_object_map(repo, dir) < 0) {\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\treturn 0;\n+}\n+\n+int repo_write_loose_object_map(struct repository *repo)\n+{\n+\tkh_oid_map_t *map = repo->objects->odb->loose_map->to_compat;\n+\tstruct lock_file lock;\n+\tint fd;\n+\tkhiter_t iter;\n+\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n+\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n+\tfd = hold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n+\titer = kh_begin(map);\n+\tif (write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n+\t\tgoto errout;\n+\n+\tfor (; iter != kh_end(map); iter++) {\n+\t\tif (kh_exist(map, iter)) {\n+\t\t\tif (oideq(&kh_key(map, iter), the_hash_algo->empty_tree) ||\n+\t\t\t    oideq(&kh_key(map, iter), the_hash_algo->empty_blob))\n+\t\t\t\tcontinue;\n+\t\t\tstrbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(&kh_key(map, iter)), oid_to_hex(kh_value(map, iter)));\n+\t\t\tif (write_in_full(fd, buf.buf, buf.len) < 0)\n+\t\t\t\tgoto errout;\n+\t\t\tstrbuf_reset(&buf);\n+\t\t}\n+\t}\n+\tstrbuf_release(&buf);\n+\tif (commit_lock_file(&lock) < 0) {\n+\t\terror_errno(_(\"could not write loose object index %s\"), path.buf);\n+\t\tstrbuf_release(&path);\n+\t\treturn -1;\n+\t}\n+\tstrbuf_release(&path);\n+\treturn 0;\n+errout:\n+\trollback_lock_file(&lock);\n+\tstrbuf_release(&buf);\n+\terror_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n+\tstrbuf_release(&path);\n+\treturn -1;\n+}\n+\n+static int write_one_object(struct repository *repo, const struct object_id *oid,\n+\t\t\t    const struct object_id *compat_oid)\n+{\n+\tstruct lock_file lock;\n+\tint fd;\n+\tstruct stat st;\n+\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\n+\tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n+\thold_lock_file_for_update_timeout(&lock, path.buf, LOCK_DIE_ON_ERROR, -1);\n+\n+\tfd = open(path.buf, O_WRONLY | O_CREAT | O_APPEND, 0666);\n+\tif (fd < 0)\n+\t\tgoto errout;\n+\tif (fstat(fd, &st) < 0)\n+\t\tgoto errout;\n+\tif (!st.st_size && write_in_full(fd, loose_object_header, strlen(loose_object_header)) < 0)\n+\t\tgoto errout;\n+\n+\tstrbuf_addf(&buf, \"%s %s\\n\", oid_to_hex(oid), oid_to_hex(compat_oid));\n+\tif (write_in_full(fd, buf.buf, buf.len) < 0)\n+\t\tgoto errout;\n+\tif (close(fd))\n+\t\tgoto errout;\n+\tadjust_shared_perm(path.buf);\n+\trollback_lock_file(&lock);\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn 0;\n+errout:\n+\terror_errno(_(\"failed to write loose object index %s\\n\"), path.buf);\n+\tclose(fd);\n+\trollback_lock_file(&lock);\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&path);\n+\treturn -1;\n+}\n+\n+int repo_add_loose_object_map(struct repository *repo, const struct object_id *oid,\n+\t\t\t      const struct object_id *compat_oid)\n+{\n+\tint inserted = 0;\n+\n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n+\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_compat, oid, compat_oid);\n+\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_storage, compat_oid, oid);\n+\tif (inserted)\n+\t\treturn write_one_object(repo, oid, compat_oid);\n+\treturn 0;\n+}\n+\n+int repo_loose_object_map_oid(struct repository *repo, struct object_id *dest,\n+\t\t\t      const struct git_hash_algo *to,\n+\t\t\t      const struct object_id *src)\n+{\n+\tstruct object_directory *dir;\n+\tkh_oid_map_t *map;\n+\tkhiter_t pos;\n+\n+\tfor (dir = repo->objects->odb; dir; dir = dir->next) {\n+\t\tstruct loose_object_map *loose_map = dir->loose_map;\n+\t\tif (!loose_map)\n+\t\t\tcontinue;\n+\t\tmap = (to == repo->compat_hash_algo) ?\n+\t\t\tloose_map->to_compat :\n+\t\t\tloose_map->to_storage;\n+\t\tpos = kh_get_oid_map(map, *src);\n+\t\tif (pos < kh_end(map)) {\n+\t\t\toidcpy(dest, kh_value(map, pos));\n+\t\t\treturn 0;\n+\t\t}\n+\t}\n+\treturn -1;\n+}\n+\n+void loose_object_map_clear(struct loose_object_map **map)\n+{\n+\tstruct loose_object_map *m = *map;\n+\tstruct object_id *oid;\n+\n+\tif (!m)\n+\t\treturn;\n+\n+\tkh_foreach_value(m->to_compat, oid, free(oid));\n+\tkh_foreach_value(m->to_storage, oid, free(oid));\n+\tkh_destroy_oid_map(m->to_compat);\n+\tkh_destroy_oid_map(m->to_storage);\n+\tfree(m);\n+\t*map = NULL;\n+}\ndiff --git a/loose.h b/loose.h\nnew file mode 100644\nindex 000000000000..061c6937aead\n--- /dev/null\n+++ b/loose.h\n@@ -0,0 +1,20 @@\n+#ifndef LOOSE_H\n+#define LOOSE_H\n+\n+#include \"khash.h\"\n+\n+struct loose_object_map {\n+\tkh_oid_map_t *to_compat;\n+\tkh_oid_map_t *to_storage;\n+};\n+\n+void loose_object_map_init(struct loose_object_map **map);\n+void loose_object_map_clear(struct loose_object_map **map);\n+int repo_loose_object_map_oid(struct repository *repo, struct object_id *dest,\n+\tconst struct git_hash_algo *dest_algo, const struct object_id *src);\n+int repo_add_loose_object_map(struct repository *repo, const struct object_id *oid,\n+\t\t\t      const struct object_id *compat_oid);\n+int repo_read_loose_object_map(struct repository *repo);\n+int repo_write_loose_object_map(struct repository *repo);\n+\n+#endif\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 9f4d5b354f5f..e7c62434016d 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -4,6 +4,7 @@\n #include \"repository.h\"\n #include \"hash-ll.h\"\n #include \"object.h\"\n+#include \"loose.h\"\n #include \"object-file-convert.h\"\n \n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n@@ -21,7 +22,18 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \t\t\toidcpy(dest, src);\n \t\treturn 0;\n \t}\n-\treturn -1;\n+\tif (repo_loose_object_map_oid(repo, dest, to, src)) {\n+\t\t/*\n+\t\t * We may have loaded the object map at repo initialization but\n+\t\t * another process (perhaps upstream of a pipe from us) may have\n+\t\t * written a new object into the map.  If the object is missing,\n+\t\t * let's reload the map to see if the object has appeared.\n+\t\t */\n+\t\trepo_read_loose_object_map(repo);\n+\t\tif (repo_loose_object_map_oid(repo, dest, to, src))\n+\t\t\treturn -1;\n+\t}\n+\treturn 0;\n }\n \n int convert_object_file(struct strbuf *outbuf,\ndiff --git a/object-store-ll.h b/object-store-ll.h\nindex 26a3895c821c..bc76d6bec80d 100644\n--- a/object-store-ll.h\n+++ b/object-store-ll.h\n@@ -26,6 +26,9 @@ struct object_directory {\n \tuint32_t loose_objects_subdir_seen[8]; /* 256 bits */\n \tstruct oidtree *loose_objects_cache;\n \n+\t/* Map between object IDs for loose objects. */\n+\tstruct loose_object_map *loose_map;\n+\n \t/*\n \t * This is a temporary object store created by the tmp_objdir\n \t * facility. Disable ref updates since the objects in the store\ndiff --git a/object.c b/object.c\nindex 2c61e4c86217..186a0a47c0fb 100644\n--- a/object.c\n+++ b/object.c\n@@ -13,6 +13,7 @@\n #include \"alloc.h\"\n #include \"packfile.h\"\n #include \"commit-graph.h\"\n+#include \"loose.h\"\n \n unsigned int get_max_object_index(void)\n {\n@@ -540,6 +541,7 @@ void free_object_directory(struct object_directory *odb)\n {\n \tfree(odb->path);\n \todb_clear_loose_cache(odb);\n+\tloose_object_map_clear(&odb->loose_map);\n \tfree(odb);\n }\n \ndiff --git a/repository.c b/repository.c\nindex de620d82bfc6..4ab44d3b0344 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -14,6 +14,7 @@\n #include \"read-cache-ll.h\"\n #include \"remote.h\"\n #include \"setup.h\"\n+#include \"loose.h\"\n #include \"submodule-config.h\"\n #include \"sparse-index.h\"\n #include \"trace2.h\"\n@@ -112,6 +113,8 @@ void repo_enable_compat_map(struct repository *repo, int enable_compat)\n \t\t\tGIT_HASH_SHA1];\n \n \trepo->compat_hash_algo = enable_compat ? other_algo : NULL;\n+\tif (enable_compat)\n+\t\trepo_read_loose_object_map(repo);\n }\n \n /*\n@@ -204,6 +207,9 @@ int repo_init(struct repository *repo,\n \tif (worktree)\n \t\trepo_set_worktree(repo, worktree);\n \n+\tif (repo->compat_hash_algo)\n+\t\trepo_read_loose_object_map(repo);\n+\n \tclear_repository_format(&format);\n \treturn 0;\n \n-- \n2.41.0\n\n"},{"id":"481587","messageId":"20230908231049.2035003-19-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 19/32] object-file-convert: convert tag commits when writing","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:36Z","receivedAt":"2023-09-08T23:12:18Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nWhen writing a tag object in a repository with both SHA-1 and SHA-256,\nwe'll need to convert our commit objects so that we can write the hash\nvalues for both into the repository.  To do so, let's add a function to\nconvert tag objects.\n\nNote that signatures for tag objects in the current algorithm trail the\nmessage, and those for the alternate algorithm are in headers.\nTherefore, we parse the tag object for both a trailing signature and a\nheader and then, when writing the other format, swap the two around.\n\nWe expose the add_commit_signature function, which we rename now that it\nis useful for tags as well, and use it to add the header.\n\n****\n- Moved convert_tag_object into object-file-convert.c and\n  made it static\n- Adjusted how convert_object_file calls convert_tag_object\n--EWB\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n commit.c              |  6 +++---\n commit.h              |  1 +\n object-file-convert.c | 50 +++++++++++++++++++++++++++++++++++++++++++\n 3 files changed, 54 insertions(+), 3 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex 522ebb4b3002..54f19ed0328c 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1101,7 +1101,7 @@ static const char *gpg_sig_headers[] = {\n \t\"gpgsig-sha256\",\n };\n \n-static int add_commit_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo)\n+int add_header_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo)\n {\n \tint inspos, copypos;\n \tconst char *eoh;\n@@ -1732,9 +1732,9 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n \t\tfor (i = 0; i < ARRAY_SIZE(bufs); i++) {\n \t\t\tif (!bufs[i].algo)\n \t\t\t\tcontinue;\n-\t\t\tadd_commit_signature(&buffer, bufs[i].sig, bufs[i].algo);\n+\t\t\tadd_header_signature(&buffer, bufs[i].sig, bufs[i].algo);\n \t\t\tif (r->compat_hash_algo)\n-\t\t\t\tadd_commit_signature(&compat_buffer, bufs[i].sig, bufs[i].algo);\n+\t\t\t\tadd_header_signature(&compat_buffer, bufs[i].sig, bufs[i].algo);\n \t\t}\n \t}\n \ndiff --git a/commit.h b/commit.h\nindex 28928833c544..03edcec0129f 100644\n--- a/commit.h\n+++ b/commit.h\n@@ -370,5 +370,6 @@ int parse_buffer_signed_by_header(const char *buffer,\n \t\t\t\t  struct strbuf *payload,\n \t\t\t\t  struct strbuf *signature,\n \t\t\t\t  const struct git_hash_algo *algop);\n+int add_header_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo);\n \n #endif /* COMMIT_H */\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 9c715a9864d5..d381d3d2ea65 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -7,6 +7,8 @@\n #include \"hash.h\"\n #include \"object.h\"\n #include \"loose.h\"\n+#include \"commit.h\"\n+#include \"gpg-interface.h\"\n #include \"object-file-convert.h\"\n \n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n@@ -125,6 +127,52 @@ static int convert_commit_object(struct strbuf *out,\n \treturn 0;\n }\n \n+static int convert_tag_object(struct strbuf *out,\n+\t\t\t      const struct git_hash_algo *from,\n+\t\t\t      const struct git_hash_algo *to,\n+\t\t\t      const char *buffer, size_t size)\n+{\n+\tstruct strbuf payload = STRBUF_INIT, temp = STRBUF_INIT, oursig = STRBUF_INIT, othersig = STRBUF_INIT;\n+\tsize_t payload_size;\n+\tstruct object_id oid, mapped_oid;\n+\tconst char *p;\n+\n+\t/* Add some slop for longer signature header in the new algorithm. */\n+\tstrbuf_grow(out, size + 7);\n+\n+\t/* Is there a signature for our algorithm? */\n+\tpayload_size = parse_signed_buffer(buffer, size);\n+\tstrbuf_add(&payload, buffer, payload_size);\n+\tif (payload_size != size) {\n+\t\t/* Yes, there is. */\n+\t\tstrbuf_add(&oursig, buffer + payload_size, size - payload_size);\n+\t}\n+\t/* Now, is there a signature for the other algorithm? */\n+\tif (parse_buffer_signed_by_header(payload.buf, payload.len, &temp, &othersig, to)) {\n+\t\t/* Yes, there is. */\n+\t\tstrbuf_swap(&payload, &temp);\n+\t\tstrbuf_release(&temp);\n+\t}\n+\n+\t/*\n+\t * Our payload is now in payload and we may have up to two signatrures\n+\t * in oursig and othersig.\n+\t */\n+\tif (strncmp(payload.buf, \"object \", 7) || payload.buf[from->hexsz + 7] != '\\n')\n+\t\treturn error(\"bogus tag object\");\n+\tif (parse_oid_hex_algop(payload.buf + 7, &oid, &p, from) < 0)\n+\t\treturn error(\"bad tag object ID\");\n+\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\treturn error(\"unable to map tree %s in tag object\",\n+\t\t\t     oid_to_hex(&oid));\n+\tstrbuf_addf(out, \"object %s\\n\", oid_to_hex(&mapped_oid));\n+\tstrbuf_add(out, p, payload.len - (p - payload.buf));\n+\tstrbuf_addbuf(out, &othersig);\n+\tif (oursig.len)\n+\t\tadd_header_signature(out, &oursig, from);\n+\treturn 0;\n+}\n+\n int convert_object_file(struct strbuf *outbuf,\n \t\t\tconst struct git_hash_algo *from,\n \t\t\tconst struct git_hash_algo *to,\n@@ -146,6 +194,8 @@ int convert_object_file(struct strbuf *outbuf,\n \t\tret = convert_commit_object(outbuf, from, to, buf, len);\n \t\tbreak;\n \tcase OBJ_TAG:\n+\t\tret = convert_tag_object(outbuf, from, to, buf, len);\n+\t\tbreak;\n \tdefault:\n \t\t/* Not implemented yet, so fail. */\n \t\tret = -1;\n-- \n2.41.0\n\n"},{"id":"481588","messageId":"20230908231049.2035003-20-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 20/32] builtin/cat-file: Let the oid determine the output algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:37Z","receivedAt":"2023-09-08T23:12:19Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Use GET_OID_UNTRANSLATED when calling get_oid_with_context.  This\nimplements the semi-obvious behaviour that specifying a sha1 oid shows\nthe output for a sha1 encoded object, and specifying a sha256 oid\nshows the output for a sha256 encoded object.\n\nThis is useful for testing the the conversion of an object to an\nequivalent object encoded with a different hash function.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/cat-file.c | 8 ++++++--\n 1 file changed, 6 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 694c8538df2f..7c9600292376 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -107,7 +107,10 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \tstruct object_info oi = OBJECT_INFO_INIT;\n \tstruct strbuf sb = STRBUF_INIT;\n \tunsigned flags = OBJECT_INFO_LOOKUP_REPLACE;\n-\tunsigned get_oid_flags = GET_OID_RECORD_PATH | GET_OID_ONLY_TO_DIE;\n+\tunsigned get_oid_flags =\n+\t\tGET_OID_RECORD_PATH |\n+\t\tGET_OID_ONLY_TO_DIE |\n+\t\tGET_OID_UNTRANSLATED;\n \tconst char *path = force_path;\n \tconst int opt_cw = (opt == 'c' || opt == 'w');\n \tif (!path && opt_cw)\n@@ -223,7 +226,8 @@ static int cat_one_file(int opt, const char *exp_type, const char *obj_name,\n \t\t\t\t\t\t\t\t     &size);\n \t\t\t\tconst char *target;\n \t\t\t\tif (!skip_prefix(buffer, \"object \", &target) ||\n-\t\t\t\t    get_oid_hex(target, &blob_oid))\n+\t\t\t\t    get_oid_hex_algop(target, &blob_oid,\n+\t\t\t\t\t\t      &hash_algos[oid.algo]))\n \t\t\t\t\tdie(\"%s not a valid tag\", oid_to_hex(&oid));\n \t\t\t\tfree(buffer);\n \t\t\t} else\n-- \n2.41.0\n\n"},{"id":"481589","messageId":"20230908231049.2035003-22-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 22/32] object-file: Handle compat objects in check_object_signature","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:39Z","receivedAt":"2023-09-08T23:12:20Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Update check_object_signature to find the hash algorithm the exising\nsignature uses, and to use the same hash algorithm when recomputing it\nto check the signature is valid.\n\nThis will be useful when teaching git ls-tree to display objects\nencoded with the compat hash algorithm.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n object-file.c | 4 +++-\n 1 file changed, 3 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex fd420dd303df..d6140ebccaf1 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -1094,9 +1094,11 @@ int check_object_signature(struct repository *r, const struct object_id *oid,\n \t\t\t   void *buf, unsigned long size,\n \t\t\t   enum object_type type)\n {\n+\tconst struct git_hash_algo *algo =\n+\t\toid->algo ? &hash_algos[oid->algo] : r->hash_algo;\n \tstruct object_id real_oid;\n \n-\thash_object_file(r->hash_algo, buf, size, type, &real_oid);\n+\thash_object_file(algo, buf, size, type, &real_oid);\n \n \treturn !oideq(oid, &real_oid) ? -1 : 0;\n }\n-- \n2.41.0\n\n"},{"id":"481590","messageId":"20230908231049.2035003-26-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 26/32] object-file-convert: Implement convert_object_file_{begin,step,end}","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:43Z","receivedAt":"2023-09-08T23:12:29Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"When converting trees, commits, and tags the objects they reference\nneed to be converted before the objects themselves can be converted.\n\nSplit convert_objet_file_convert into a couple of pieces that are\neffectively an iterator over the oids that need to be converted.  This\nallows the objects to be processed depth first when being converted\nand it allows changing the logic to map oids.  In cases like \"git\nindex-pack\" none of the oids will be mapped in any of the existing\nmapping tables so an in-memory table needs to be converted and\nconsulted, and this allows that.\n\nNot having to update the existing object id mapping mechanisms\nis particularly nice as it makes it easy to avoid having to\nintroduce new locks to syncrhonize the update of internal\nmapping mechanisms.\n\nThis was inspired by a similar change by \"brian m. carlson\"\nwhere he modified convert_object_file to return the\nunmmaped oids.\n\nInspired-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n object-file-convert.c | 226 ++++++++++++++++++++++++++++++------------\n object-file-convert.h |  21 ++++\n 2 files changed, 186 insertions(+), 61 deletions(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 7978aa63dfa9..3fd080ebc112 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -67,55 +67,74 @@ static int decode_tree_entry_raw(struct object_id *oid, const char **path,\n \treturn 0;\n }\n \n-static int convert_tree_object(struct strbuf *out,\n-\t\t\t       const struct git_hash_algo *from,\n-\t\t\t       const struct git_hash_algo *to,\n-\t\t\t       const char *buffer, size_t size)\n+static int convert_tree_object_step(struct object_file_convert_state *state)\n {\n-\tconst char *p = buffer, *end = buffer + size;\n+\tconst char *buf = state->buf, *p, *end = buf + state->buf_len;\n+\tconst struct git_hash_algo *from = state->from;\n+\tconst struct git_hash_algo *to = state->to;\n+\tstruct strbuf *out = state->outbuf;\n+\n+\t/* The current position */\n+\tp = buf + state->buf_pos;\n \n \twhile (p < end) {\n-\t\tstruct object_id entry_oid, mapped_oid;\n+\t\tstruct object_id entry_oid;\n \t\tconst char *path = NULL;\n \t\tsize_t pathlen;\n \n \t\tif (decode_tree_entry_raw(&entry_oid, &path, &pathlen, from, p,\n \t\t\t\t\t  end - p))\n \t\t\treturn error(_(\"failed to decode tree entry\"));\n-\t\tif (repo_oid_to_algop(the_repository, &entry_oid, to, &mapped_oid))\n-\t\t\treturn error(_(\"failed to map tree entry for %s\"), oid_to_hex(&entry_oid));\n+\n+\t\tif (!state->mapped_oid.algo) {\n+\t\t\toidcpy(&state->oid, &entry_oid);\n+\t\t\treturn 1;\n+\t\t}\n+\t\telse if (!oideq(&entry_oid, &state->oid))\n+\t\t\treturn error(_(\"bad object_file_convert_state oid\"));\n+\n \t\tstrbuf_add(out, p, path - p);\n \t\tstrbuf_add(out, path, pathlen);\n-\t\tstrbuf_add(out, mapped_oid.hash, to->rawsz);\n+\t\tstrbuf_add(out, state->mapped_oid.hash, to->rawsz);\n+\t\tstate->mapped_oid.algo = 0;\n \t\tp = path + pathlen + from->rawsz;\n+\t\tstate->buf_pos = p - buf;\n \t}\n \treturn 0;\n }\n \n-static int convert_commit_object(struct strbuf *out,\n-\t\t\t\t const struct git_hash_algo *from,\n-\t\t\t\t const struct git_hash_algo *to,\n-\t\t\t\t const char *buffer, size_t size)\n+static int convert_commit_object_step(struct object_file_convert_state *state)\n {\n-\tconst char *tail = buffer;\n-\tconst char *bufptr = buffer;\n+\tconst struct git_hash_algo *from = state->from;\n+\tstruct strbuf *out = state->outbuf;\n+\tconst char *buf = state->buf;\n+\tconst char *tail = buf + state->buf_len;\n+\tconst char *bufptr = buf + state->buf_pos;\n \tconst int tree_entry_len = from->hexsz + 5;\n \tconst int parent_entry_len = from->hexsz + 7;\n-\tstruct object_id oid, mapped_oid;\n+\tstruct object_id oid;\n \tconst char *p;\n \n-\ttail += size;\n-\tif (tail <= bufptr + tree_entry_len + 1 || memcmp(bufptr, \"tree \", 5) ||\n-\t\t\tbufptr[tree_entry_len] != '\\n')\n-\t\treturn error(\"bogus commit object\");\n-\tif (parse_oid_hex_algop(bufptr + 5, &oid, &p, from) < 0)\n-\t\treturn error(\"bad tree pointer\");\n+\tif (state->buf_pos == 0) {\n+\t\tif (tail <= bufptr + tree_entry_len + 1 || memcmp(bufptr, \"tree \", 5) ||\n+\t\t    bufptr[tree_entry_len] != '\\n')\n+\t\t\treturn error(\"bogus commit object\");\n+\n+\t\tif (parse_oid_hex_algop(bufptr + 5, &oid, &p, from) < 0)\n+\t\t\treturn error(\"bad tree pointer\");\n \n-\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n-\t\treturn error(\"unable to map tree %s in commit object\",\n-\t\t\t     oid_to_hex(&oid));\n-\tstrbuf_addf(out, \"tree %s\\n\", oid_to_hex(&mapped_oid));\n-\tbufptr = p + 1;\n+\t\tif (!state->mapped_oid.algo) {\n+\t\t\toidcpy(&state->oid, &oid);\n+\t\t\treturn 1;\n+\t\t}\n+\t\telse if (!oideq(&oid, &state->oid))\n+\t\t\treturn error(_(\"bad object_file_convert_state oid\"));\n+\n+\t\tstrbuf_addf(out, \"tree %s\\n\", oid_to_hex(&state->mapped_oid));\n+\t\tstate->mapped_oid.algo = 0;\n+\t\tbufptr = p + 1;\n+\t\tstate->buf_pos = bufptr - buf;\n+\t}\n \n \twhile (bufptr + parent_entry_len < tail && !memcmp(bufptr, \"parent \", 7)) {\n \t\tif (tail <= bufptr + parent_entry_len + 1 ||\n@@ -123,26 +142,44 @@ static int convert_commit_object(struct strbuf *out,\n \t\t    *p != '\\n')\n \t\t\treturn error(\"bad parents in commit\");\n \n-\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n-\t\t\treturn error(\"unable to map parent %s in commit object\",\n-\t\t\t\t     oid_to_hex(&oid));\n+\t\tif (!state->mapped_oid.algo) {\n+\t\t\toidcpy(&state->oid, &oid);\n+\t\t\treturn 1;\n+\t\t}\n+\t\telse if (!oideq(&oid, &state->oid))\n+\t\t\treturn error(_(\"bad object_file_convert_state oid\"));\n \n-\t\tstrbuf_addf(out, \"parent %s\\n\", oid_to_hex(&mapped_oid));\n+\t\tstrbuf_addf(out, \"parent %s\\n\", oid_to_hex(&state->mapped_oid));\n+\t\tstate->mapped_oid.algo = 0;\n \t\tbufptr = p + 1;\n+\t\tstate->buf_pos = bufptr - buf;\n \t}\n \tstrbuf_add(out, bufptr, tail - bufptr);\n \treturn 0;\n }\n \n-static int convert_tag_object(struct strbuf *out,\n-\t\t\t      const struct git_hash_algo *from,\n-\t\t\t      const struct git_hash_algo *to,\n-\t\t\t      const char *buffer, size_t size)\n+static int convert_tag_object_step(struct object_file_convert_state *state)\n {\n \tstruct strbuf payload = STRBUF_INIT, temp = STRBUF_INIT, oursig = STRBUF_INIT, othersig = STRBUF_INIT;\n-\tsize_t payload_size;\n-\tstruct object_id oid, mapped_oid;\n+\tconst struct git_hash_algo *from = state->from;\n+\tconst struct git_hash_algo *to = state->to;\n+\tstruct strbuf *out = state->outbuf;\n+\tconst char *buffer = state->buf;\n+\tsize_t payload_size, size = state->buf_len;;\n+\tstruct object_id oid;\n \tconst char *p;\n+\tint ret = 0;\n+\n+\tif (!state->mapped_oid.algo) {\n+\t\tif (strncmp(buffer, \"object \", 7) ||\n+\t\t    buffer[from->hexsz + 7] != '\\n')\n+\t\t\treturn error(\"bogus tag object\");\n+\t\tif (parse_oid_hex_algop(buffer + 7, &oid, &p, from) < 0)\n+\t\t\treturn error(\"bad tag object ID\");\n+\n+\t\toidcpy(&state->oid, &oid);\n+\t\treturn 1;\n+\t}\n \n \t/* Add some slop for longer signature header in the new algorithm. */\n \tstrbuf_grow(out, size + 7);\n@@ -165,52 +202,119 @@ static int convert_tag_object(struct strbuf *out,\n \t * Our payload is now in payload and we may have up to two signatrures\n \t * in oursig and othersig.\n \t */\n-\tif (strncmp(payload.buf, \"object \", 7) || payload.buf[from->hexsz + 7] != '\\n')\n-\t\treturn error(\"bogus tag object\");\n-\tif (parse_oid_hex_algop(payload.buf + 7, &oid, &p, from) < 0)\n-\t\treturn error(\"bad tag object ID\");\n-\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n-\t\treturn error(\"unable to map tree %s in tag object\",\n-\t\t\t     oid_to_hex(&oid));\n-\tstrbuf_addf(out, \"object %s\\n\", oid_to_hex(&mapped_oid));\n+\tif (strncmp(payload.buf, \"object \", 7) || payload.buf[from->hexsz + 7] != '\\n') {\n+\t\tret = error(\"bogus tag object\");\n+\t\tgoto out;\n+\t}\n+\tif (parse_oid_hex_algop(payload.buf + 7, &oid, &p, from) < 0) {\n+\t\tret = error(\"bad tag object ID\");\n+\t\tgoto out;\n+\t}\n+\tif (!oideq(&oid, &state->oid)) {\n+\t\tret = error(_(\"bad object_file_convert_state oid\"));\n+\t\tgoto out;\n+\t}\n+\n+\tstrbuf_addf(out, \"object %s\\n\", oid_to_hex(&state->mapped_oid));\n \tstrbuf_add(out, p, payload.len - (p - payload.buf));\n \tstrbuf_addbuf(out, &othersig);\n \tif (oursig.len)\n \t\tadd_header_signature(out, &oursig, from);\n-\treturn 0;\n+out:\n+\tstrbuf_release(&oursig);\n+\tstrbuf_release(&othersig);\n+\tstrbuf_release(&payload);\n+\treturn ret;\n }\n \n-int convert_object_file(struct strbuf *outbuf,\n-\t\t\tconst struct git_hash_algo *from,\n-\t\t\tconst struct git_hash_algo *to,\n-\t\t\tconst void *buf, size_t len,\n-\t\t\tenum object_type type,\n-\t\t\tint gentle)\n+void convert_object_file_begin(struct object_file_convert_state *state,\n+\t\t\t      struct strbuf *outbuf,\n+\t\t\t      const struct git_hash_algo *from,\n+\t\t\t      const struct git_hash_algo *to,\n+\t\t\t      const void *buf, size_t len,\n+\t\t\t      enum object_type type)\n {\n-\tint ret;\n+\tmemset(state, 0, sizeof(*state));\n+\tstate->outbuf = outbuf;\n+\tstate->from = from;\n+\tstate->to = to;\n+\tstate->buf = buf;\n+\tstate->buf_len = len;\n+\tstate->buf_pos = 0;\n+\tstate->type = type;\n+\n \n \t/* Don't call this function when no conversion is necessary */\n \tif ((from == to) || (type == OBJ_BLOB))\n-\t\tdie(\"Refusing noop object file conversion\");\n+\t\tBUG(\"Attempting noop object file conversion\");\n \n \tswitch (type) {\n \tcase OBJ_TREE:\n-\t\tret = convert_tree_object(outbuf, from, to, buf, len);\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TAG:\n+\t\tbreak;\n+\tdefault:\n+\t\t/* Not implemented yet, so fail. */\n+\t\tBUG(\"Unknown object file type found in conversion\");\n+\t}\n+}\n+\n+int convert_object_file_step(struct object_file_convert_state *state)\n+{\n+\tint ret;\n+\n+\tswitch(state->type) {\n+\tcase OBJ_TREE:\n+\t\tret = convert_tree_object_step(state);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n-\t\tret = convert_commit_object(outbuf, from, to, buf, len);\n+\t\tret = convert_commit_object_step(state);\n \t\tbreak;\n \tcase OBJ_TAG:\n-\t\tret = convert_tag_object(outbuf, from, to, buf, len);\n+\t\tret = convert_tag_object_step(state);\n \t\tbreak;\n \tdefault:\n-\t\t/* Not implemented yet, so fail. */\n \t\tret = -1;\n \t\tbreak;\n \t}\n-\tif (!ret)\n-\t\treturn 0;\n-\tif (gentle)\n+\treturn ret;\n+}\n+\n+void convert_object_file_end(struct object_file_convert_state *state, int ret)\n+{\n+\tif (ret != 0) {\n+\t\tstrbuf_release(state->outbuf);\n+\t}\n+\tmemset(state, 0, sizeof(*state));\n+}\n+\n+int convert_object_file(struct strbuf *outbuf,\n+\t\t\tconst struct git_hash_algo *from,\n+\t\t\tconst struct git_hash_algo *to,\n+\t\t\tconst void *buf, size_t len,\n+\t\t\tenum object_type type,\n+\t\t\tint gentle)\n+{\n+\tstruct object_file_convert_state state;\n+\tint ret;\n+\n+\tconvert_object_file_begin(&state, outbuf, from, to, buf, len, type);\n+\n+\tfor (;;) {\n+\t\tret = convert_object_file_step(&state);\n+\t\tif (ret != 1)\n+\t\t\tbreak;\n+\t\tret = repo_oid_to_algop(the_repository, &state.oid, state.to,\n+\t\t\t\t\t&state.mapped_oid);\n+\t\tif (ret) {\n+\t\t\terror(_(\"failed to map %s entry for %s\"),\n+\t\t\t      type_name(type), oid_to_hex(&state.oid));\n+\t\t\tbreak;\n+\t\t}\n+\t}\n+\n+\tconvert_object_file_end(&state, ret);\n+\tif (!ret || gentle)\n \t\treturn ret;\n \tdie(_(\"Failed to convert object from %s to %s\"),\n \t\tfrom->name, to->name);\ndiff --git a/object-file-convert.h b/object-file-convert.h\nindex a4f802aa8eea..da032d7a91ef 100644\n--- a/object-file-convert.h\n+++ b/object-file-convert.h\n@@ -10,6 +10,27 @@ struct strbuf;\n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \t\t      const struct git_hash_algo *to, struct object_id *dest);\n \n+struct object_file_convert_state {\n+\tstruct strbuf *outbuf;\n+\tconst struct git_hash_algo *from;\n+\tconst struct git_hash_algo *to;\n+\tconst void *buf;\n+\tsize_t buf_len;\n+\tsize_t buf_pos;\n+\tenum object_type type;\n+\tstruct object_id oid;\n+\tstruct object_id mapped_oid;\n+};\n+\n+void convert_object_file_begin(struct object_file_convert_state *state,\n+\t\t\t       struct strbuf *outbuf,\n+\t\t\t       const struct git_hash_algo *from,\n+\t\t\t       const struct git_hash_algo *to,\n+\t\t\t       const void *buf, size_t len,\n+\t\t\t       enum object_type type);\n+int convert_object_file_step(struct object_file_convert_state *state);\n+void convert_object_file_end(struct object_file_convert_state *state, int ret);\n+\n /*\n  * Convert an object file from one hash algorithm to another algorithm.\n  * Return -1 on failure, 0 on success.\n-- \n2.41.0\n\n"},{"id":"481591","messageId":"20230908231049.2035003-27-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 27/32] builtin/fast-import: compute compatibility hashs for imported objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:44Z","receivedAt":"2023-09-08T23:12:35Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"When the code is in dual hash mode for every object fast-import\ncreates compute the standard oid and it's compatibility mapping.  The\ncompatibility mapping is stored in struct pack_idx_entry so that it\ncan be used when an index is created.\n\nFor fast-import the code needs to be careful because when a new object\nonly refers to other newly created objects the compatibility mapping\nfor those new objects is not stored anywhere permanently.  So have the\ncode first look the the compatibility oid in the newly created\nobjects, and then look for the compatibilty oid in the standard\nmapping tables.\n\nAs fast-import requires objects to be specified before the\nobjects that reference them nothing special needs to happen\nto deal with out of order objects.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/fast-import.c | 89 +++++++++++++++++++++++++++++++++++++------\n 1 file changed, 77 insertions(+), 12 deletions(-)\n\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 2c645fcfbe3f..f1c250dd3c8f 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -26,6 +26,8 @@\n #include \"commit-reach.h\"\n #include \"khash.h\"\n #include \"date.h\"\n+#include \"object-file-convert.h\"\n+#include \"pack-compat-map.h\"\n \n #define PACK_ID_BITS 16\n #define MAX_PACK_ID ((1<<PACK_ID_BITS)-1)\n@@ -775,9 +777,14 @@ static void start_packfile(void)\n \tall_packs[pack_id] = p;\n }\n \n-static const char *create_index(void)\n+struct pack_index_names {\n+\tconst char *index_name;\n+\tconst char *compat_name;\n+};\n+\n+static struct pack_index_names create_index(void)\n {\n-\tconst char *tmpfile;\n+\tstruct pack_index_names tmp = {};\n \tstruct pack_idx_entry **idx, **c, **last;\n \tstruct object_entry *e;\n \tstruct object_entry_pool *o;\n@@ -793,13 +800,15 @@ static const char *create_index(void)\n \tif (c != last)\n \t\tdie(\"internal consistency error creating the index\");\n \n-\ttmpfile = write_idx_file(NULL, idx, object_count, &pack_idx_opts,\n-\t\t\t\t pack_data->hash);\n+\ttmp.index_name = write_idx_file(NULL, idx, object_count, &pack_idx_opts,\n+\t\t\t\t\tpack_data->hash);\n+\ttmp.compat_name = write_compat_map_file(NULL, idx, object_count,\n+\t\t\t\t\t\tpack_data->hash);\n \tfree(idx);\n-\treturn tmpfile;\n+\treturn tmp;\n }\n \n-static char *keep_pack(const char *curr_index_name)\n+static char *keep_pack(struct pack_index_names curr)\n {\n \tstatic const char *keep_msg = \"fast-import\";\n \tstruct strbuf name = STRBUF_INIT;\n@@ -818,9 +827,17 @@ static char *keep_pack(const char *curr_index_name)\n \t\tdie(\"cannot store pack file\");\n \n \todb_pack_name(&name, pack_data->hash, \"idx\");\n-\tif (finalize_object_file(curr_index_name, name.buf))\n+\tif (finalize_object_file(curr.index_name, name.buf))\n \t\tdie(\"cannot store index file\");\n-\tfree((void *)curr_index_name);\n+\n+\tif (curr.compat_name) {\n+\t\todb_pack_name(&name, pack_data->hash, \"compat\");\n+\t\tif (finalize_object_file(curr.compat_name, name.buf))\n+\t\t\tdie(\"cannot store compatibility map file\");\n+\t}\n+\n+\tfree((void *)curr.index_name);\n+\tfree((void *)curr.compat_name);\n \treturn strbuf_detach(&name, NULL);\n }\n \n@@ -943,6 +960,8 @@ static int store_object(\n \tstruct object_id *oidout,\n \tuintmax_t mark)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tvoid *out, *delta;\n \tstruct object_entry *e;\n \tunsigned char hdr[96];\n@@ -966,8 +985,7 @@ static int store_object(\n \tif (e->idx.offset) {\n \t\tduplicate_count_by_type[type]++;\n \t\treturn 1;\n-\t} else if (find_sha1_pack(oid.hash,\n-\t\t\t\t  get_all_packs(the_repository))) {\n+\t} else if (find_sha1_pack(oid.hash, get_all_packs(repo))) {\n \t\te->type = type;\n \t\te->pack_id = MAX_PACK_ID;\n \t\te->idx.offset = 1; /* just not zero! */\n@@ -1026,6 +1044,42 @@ static int store_object(\n \te->type = type;\n \te->pack_id = pack_id;\n \te->idx.offset = pack_size;\n+\tif (compat && (type == OBJ_BLOB)) {\n+\t\tcompat->init_fn(&c);\n+\t\tcompat->update_fn(&c, hdr, hdrlen);\n+\t\tcompat->update_fn(&c, dat->buf, dat->len);\n+\t\tcompat->final_oid_fn(&e->idx.compat_oid, &c);\n+\t} else if (compat) {\n+\t\tstruct object_file_convert_state state;\n+\t\tstruct strbuf out = STRBUF_INIT;\n+\t\tint ret;\n+\n+\t\tconvert_object_file_begin(&state, &out, the_hash_algo, compat,\n+\t\t\t\t\t  dat->buf, dat->len, type);\n+\t\tfor (;;) {\n+\t\t\tstruct object_entry *pobj;\n+\n+\t\t\tconvert_object_file_step(&state);\n+\t\t\tif (ret != 1)\n+\t\t\t\tbreak;\n+\n+\t\t\tret = -1;\n+\t\t\tpobj = find_object(&state.oid);\n+\t\t\tif (pobj && pobj->idx.compat_oid.algo)\n+\t\t\t\toidcpy(&state.mapped_oid, &pobj->idx.compat_oid);\n+\t\t\telse if (pobj)\n+\t\t\t\tbreak;\n+\t\t\telse if (repo_oid_to_algop(repo, &state.oid, compat,\n+\t\t\t\t\t\t   &state.mapped_oid))\n+\t\t\t\tbreak;\n+\t\t}\n+\t\tconvert_object_file_end(&state, ret);\n+\t\tif (ret)\n+\t\t\tdie(_(\"No mapping for %s to %s\\n\"),\n+\t\t\t    oid_to_hex(&state.oid), compat->name);\n+\t\thash_object_file(compat, out.buf, out.len, type, &e->idx.compat_oid);\n+\t\tstrbuf_release(&out);\n+\t}\n \tobject_count++;\n \tobject_count_by_type[type]++;\n \n@@ -1084,14 +1138,15 @@ static void truncate_pack(struct hashfile_checkpoint *checkpoint)\n \n static void stream_blob(uintmax_t len, struct object_id *oidout, uintmax_t mark)\n {\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n \tsize_t in_sz = 64 * 1024, out_sz = 64 * 1024;\n \tunsigned char *in_buf = xmalloc(in_sz);\n \tunsigned char *out_buf = xmalloc(out_sz);\n \tstruct object_entry *e;\n-\tstruct object_id oid;\n+\tstruct object_id oid, compat_oid;\n \tunsigned long hdrlen;\n \toff_t offset;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, compat_c;\n \tgit_zstream s;\n \tstruct hashfile_checkpoint checkpoint;\n \tint status = Z_OK;\n@@ -1109,6 +1164,10 @@ static void stream_blob(uintmax_t len, struct object_id *oidout, uintmax_t mark)\n \n \tthe_hash_algo->init_fn(&c);\n \tthe_hash_algo->update_fn(&c, out_buf, hdrlen);\n+\tif (compat) {\n+\t\tcompat->init_fn(&compat_c);\n+\t\tcompat->update_fn(&compat_c, out_buf, hdrlen);\n+\t}\n \n \tcrc32_begin(pack_file);\n \n@@ -1127,6 +1186,8 @@ static void stream_blob(uintmax_t len, struct object_id *oidout, uintmax_t mark)\n \t\t\t\tdie(\"EOF in data (%\" PRIuMAX \" bytes remaining)\", len);\n \n \t\t\tthe_hash_algo->update_fn(&c, in_buf, n);\n+\t\t\tif (compat)\n+\t\t\t\tcompat->update_fn(&compat_c, in_buf, n);\n \t\t\ts.next_in = in_buf;\n \t\t\ts.avail_in = n;\n \t\t\tlen -= n;\n@@ -1153,6 +1214,8 @@ static void stream_blob(uintmax_t len, struct object_id *oidout, uintmax_t mark)\n \t}\n \tgit_deflate_end(&s);\n \tthe_hash_algo->final_oid_fn(&oid, &c);\n+\tif (compat)\n+\t\tcompat->final_oid_fn(&compat_oid, &compat_c);\n \n \tif (oidout)\n \t\toidcpy(oidout, &oid);\n@@ -1180,6 +1243,8 @@ static void stream_blob(uintmax_t len, struct object_id *oidout, uintmax_t mark)\n \t\te->pack_id = pack_id;\n \t\te->idx.offset = offset;\n \t\te->idx.crc32 = crc32_end(pack_file);\n+\t\tif (compat)\n+\t\t\toidcpy(&e->idx.compat_oid, &compat_oid);\n \t\tobject_count++;\n \t\tobject_count_by_type[OBJ_BLOB]++;\n \t}\n-- \n2.41.0\n\n"},{"id":"481592","messageId":"20230908231049.2035003-29-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 29/32] builtin/index-pack: Compute the compatibility hash","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:46Z","receivedAt":"2023-09-08T23:12:37Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"When a pack is received encoded with the same algorithm as our\nrepository it is necessary to compute it's hash values and create it's\nindexes.  That is the job of \"git index-pack\".  To compute the primary\nhash values of the objects, the objects must be loaded in memory.\nWith the objects loaded into memory this is the perfect time to also\ncompute the compatiblity hash values of the objects as loading the\nobjects into memory is the primary cost of that operation.\n\nThis is limited by the the fact that to compute the compatiblity hash\nfor tree objects, commit objects, and tag objects the objects need to\nencoded into their compatbiilty form which requires replacing\nreferences to objects encoded with the primary hash to references to\nthe same objects encoded with the compatibility hash.\n\nWhich means that before the compatibility hash for a tree object,\ncommit object or tag object can be computed the compatibility hash\nfor all objects to which they refer must be computed first.\n\nIn general this requires an extra pass so that the dependencies between\nobjects can be resolved.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/index-pack.c | 335 ++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 328 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 75c2113e455c..f5da671ed82d 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -18,12 +18,15 @@\n #include \"thread-utils.h\"\n #include \"packfile.h\"\n #include \"pack-revindex.h\"\n+#include \"pack-compat-map.h\"\n #include \"object-file.h\"\n #include \"object-store-ll.h\"\n #include \"oid-array.h\"\n #include \"replace-object.h\"\n #include \"promisor-remote.h\"\n #include \"setup.h\"\n+#include \"strbuf.h\"\n+#include \"object-file-convert.h\"\n \n static const char index_pack_usage[] =\n \"git index-pack [-v] [-o <index-file>] [--keep | --keep=<msg>] [--[no-]rev-index] [--verify] [--strict] (<pack-file> | --stdin [--fix-thin] [<pack-file>])\";\n@@ -124,6 +127,8 @@ static int nr_ofs_deltas;\n static int nr_ref_deltas;\n static int ref_deltas_alloc;\n static int nr_resolved_deltas;\n+static int nr_pending_mappings;\n+static int nr_resolved_mappings;\n static int nr_threads;\n \n static int32_t *oid_index;\n@@ -505,28 +510,76 @@ static void prune_base_data(struct base_data *retain)\n \t}\n }\n \n+static int compat_hash_object_file(const void *buf, size_t len, enum object_type type,\n+\t\t\t\t   struct object_id *oid)\n+{\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct object_file_convert_state state;\n+\tstruct strbuf out = STRBUF_INIT;\n+\tint ret;\n+\n+\tconvert_object_file_begin(&state, &out, algo, compat,\n+\t\t\t\t  buf, len, type);\n+\tfor (;;) {\n+\t\tstruct object_entry *pobj;\n+\t\tret = convert_object_file_step(&state);\n+\t\tif (ret != 1)\n+\t\t\tbreak;\n+\n+\t\tpobj = find_in_oid_index(&state.oid);\n+\n+\t\tret = -1;\n+\t\tif (pobj && pobj->idx.compat_oid.algo)\n+\t\t\toidcpy(&state.mapped_oid, &pobj->idx.compat_oid);\n+\t\telse if (pobj)\n+\t\t\tbreak;\n+\t\telse if (repo_oid_to_algop(repo, &state.oid, compat,\n+\t\t\t\t\t   &state.mapped_oid))\n+\t\t\tbreak;\n+\t}\n+\tconvert_object_file_end(&state, ret);\n+\tif (ret == 0) {\n+\t\thash_object_file(compat, out.buf, out.len, type, oid);\n+\t\tstrbuf_release(&out);\n+\t}\n+\treturn ret;\n+}\n+\n static int is_delta_type(enum object_type type)\n {\n \treturn (type == OBJ_REF_DELTA || type == OBJ_OFS_DELTA);\n }\n \n static void *unpack_entry_data(off_t offset, unsigned long size,\n-\t\t\t       enum object_type type, struct object_id *oid)\n+\t\t\t       enum object_type type, struct object_id *oid,\n+\t\t\t       struct object_id *compat_oid)\n {\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n \tstatic char fixed_buf[8192];\n \tint status;\n \tgit_zstream stream;\n \tvoid *buf;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, compat_c;\n \tchar hdr[32];\n \tint hdrlen;\n \n+\tif (!compat)\n+\t\tcompat_oid = NULL;\n+\n \tif (!is_delta_type(type)) {\n \t\thdrlen = format_object_header(hdr, sizeof(hdr), type, size);\n \t\tthe_hash_algo->init_fn(&c);\n \t\tthe_hash_algo->update_fn(&c, hdr, hdrlen);\n-\t} else\n+\t\tif (compat_oid && (type == OBJ_BLOB)) {\n+\t\t\tcompat->init_fn(&compat_c);\n+\t\t\tcompat->update_fn(&compat_c, hdr, hdrlen);\n+\t\t}\n+\t} else {\n \t\toid = NULL;\n+\t\tcompat_oid = NULL;\n+\t}\n \tif (type == OBJ_BLOB && size > big_file_threshold)\n \t\tbuf = fixed_buf;\n \telse\n@@ -545,6 +598,8 @@ static void *unpack_entry_data(off_t offset, unsigned long size,\n \t\tuse(input_len - stream.avail_in);\n \t\tif (oid)\n \t\t\tthe_hash_algo->update_fn(&c, last_out, stream.next_out - last_out);\n+\t\tif (compat_oid && (type == OBJ_BLOB))\n+\t\t\tcompat->update_fn(&compat_c, last_out, stream.next_out - last_out);\n \t\tif (buf == fixed_buf) {\n \t\t\tstream.next_out = buf;\n \t\t\tstream.avail_out = sizeof(fixed_buf);\n@@ -555,13 +610,20 @@ static void *unpack_entry_data(off_t offset, unsigned long size,\n \tgit_inflate_end(&stream);\n \tif (oid)\n \t\tthe_hash_algo->final_oid_fn(oid, &c);\n+\tif (compat_oid && (type == OBJ_BLOB))\n+\t\tcompat->final_oid_fn(compat_oid, &compat_c);\n+\telse if (compat_oid &&\n+\t\t compat_hash_object_file(buf, size, type, compat_oid)) {\n+\t\tnr_pending_mappings++;\n+\t}\n \treturn buf == fixed_buf ? NULL : buf;\n }\n \n static void *unpack_raw_entry(struct object_entry *obj,\n \t\t\t      off_t *ofs_offset,\n \t\t\t      struct object_id *ref_oid,\n-\t\t\t      struct object_id *oid)\n+\t\t\t      struct object_id *oid,\n+\t\t\t      struct object_id *compat_oid)\n {\n \tunsigned char *p;\n \tunsigned long size, c;\n@@ -620,7 +682,8 @@ static void *unpack_raw_entry(struct object_entry *obj,\n \t}\n \tobj->hdr_size = consumed_bytes - obj->idx.offset;\n \n-\tdata = unpack_entry_data(obj->idx.offset, obj->size, obj->type, oid);\n+\tdata = unpack_entry_data(obj->idx.offset, obj->size, obj->type, oid,\n+\t\t\t\t compat_oid);\n \tobj->idx.crc32 = input_crc32;\n \treturn data;\n }\n@@ -1023,9 +1086,11 @@ static struct base_data *make_base(struct object_entry *obj,\n static struct base_data *resolve_delta(struct object_entry *delta_obj,\n \t\t\t\t       struct base_data *base)\n {\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n \tvoid *delta_data, *result_data;\n \tstruct base_data *result;\n \tunsigned long result_size;\n+\tint pending_map = 0;\n \n \tif (show_stat) {\n \t\tint i = delta_obj - objects;\n@@ -1046,6 +1111,16 @@ static struct base_data *resolve_delta(struct object_entry *delta_obj,\n \t\tbad_object(delta_obj->idx.offset, _(\"failed to apply delta\"));\n \thash_object_file(the_hash_algo, result_data, result_size,\n \t\t\t delta_obj->real_type, &delta_obj->idx.oid);\n+\tif (compat && (delta_obj->real_type == OBJ_BLOB))\n+\t\thash_object_file(compat, result_data, result_size,\n+\t\t\t\t delta_obj->real_type, &delta_obj->idx.compat_oid);\n+\telse if (compat &&\n+\t\t compat_hash_object_file(result_data, result_size,\n+\t\t\t\t\t delta_obj->real_type,\n+\t\t\t\t\t &delta_obj->idx.compat_oid)) {\n+\t\tpending_map = 1;\n+\t}\n+\n \tplace_in_oid_index(delta_obj);\n \tsha1_object(result_data, NULL, result_size, delta_obj->real_type,\n \t\t    &delta_obj->idx.oid);\n@@ -1056,6 +1131,8 @@ static struct base_data *resolve_delta(struct object_entry *delta_obj,\n \n \tcounter_lock();\n \tnr_resolved_deltas++;\n+\tif (pending_map)\n+\t\tnr_pending_mappings++;\n \tcounter_unlock();\n \n \treturn result;\n@@ -1236,7 +1313,8 @@ static void parse_pack_objects(unsigned char *hash)\n \t\tstruct object_entry *obj = &objects[i];\n \t\tvoid *data = unpack_raw_entry(obj, &ofs_delta->offset,\n \t\t\t\t\t      &ref_delta_oid,\n-\t\t\t\t\t      &obj->idx.oid);\n+\t\t\t\t\t      &obj->idx.oid,\n+\t\t\t\t\t      &obj->idx.compat_oid);\n \t\tobj->real_type = obj->type;\n \t\tif (obj->type == OBJ_OFS_DELTA) {\n \t\t\tnr_ofs_deltas++;\n@@ -1578,6 +1656,7 @@ static void rename_tmp_packfile(const char **final_name,\n static void final(const char *final_pack_name, const char *curr_pack_name,\n \t\t  const char *final_index_name, const char *curr_index_name,\n \t\t  const char *final_rev_index_name, const char *curr_rev_index_name,\n+\t\t  const char *final_compat_index_name, const char *curr_compat_index_name,\n \t\t  const char *keep_msg, const char *promisor_msg,\n \t\t  unsigned char *hash)\n {\n@@ -1585,6 +1664,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n \tstruct strbuf pack_name = STRBUF_INIT;\n \tstruct strbuf index_name = STRBUF_INIT;\n \tstruct strbuf rev_index_name = STRBUF_INIT;\n+\tstruct strbuf compat_index_name = STRBUF_INIT;\n \tint err;\n \n \tif (!from_stdin) {\n@@ -1608,6 +1688,9 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n \tif (curr_rev_index_name)\n \t\trename_tmp_packfile(&final_rev_index_name, curr_rev_index_name,\n \t\t\t\t    &rev_index_name, hash, \"rev\", 1);\n+\tif (curr_compat_index_name)\n+\t\trename_tmp_packfile(&final_compat_index_name, curr_compat_index_name,\n+\t\t\t\t    &compat_index_name, hash, \"compat\", 1);\n \trename_tmp_packfile(&final_index_name, curr_index_name, &index_name,\n \t\t\t    hash, \"idx\", 1);\n \n@@ -1640,6 +1723,7 @@ static void final(const char *final_pack_name, const char *curr_pack_name,\n \t\t}\n \t}\n \n+\tstrbuf_release(&compat_index_name);\n \tstrbuf_release(&rev_index_name);\n \tstrbuf_release(&index_name);\n \tstrbuf_release(&pack_name);\n@@ -1789,16 +1873,236 @@ static void show_pack_info(int stat_only)\n \tfree(chain_histogram);\n }\n \n+static int compare_ofs_delta_entry_obj_no(const void *a, const void *b)\n+{\n+\tconst struct ofs_delta_entry *delta_a = a;\n+\tconst struct ofs_delta_entry *delta_b = b;\n+\n+\treturn delta_a->obj_no < delta_b->obj_no ? -1 :\n+\t       delta_a->obj_no > delta_b->obj_no ?  1 :\n+\t       0;\n+}\n+\n+static int compare_ref_delta_entry_obj_no(const void *a, const void *b)\n+{\n+\tconst struct ref_delta_entry *delta_a = a;\n+\tconst struct ref_delta_entry *delta_b = b;\n+\n+\treturn delta_a->obj_no < delta_b->obj_no ? -1 :\n+\t       delta_a->obj_no > delta_b->obj_no ?  1 :\n+\t       0;\n+}\n+\n+static struct ofs_delta_entry *find_ofs_delta_obj_no(int obj_no)\n+{\n+\tint first = 0, last = nr_ofs_deltas;\n+\n+\twhile (first < last) {\n+\t\tint next = first + (last - first) / 2;\n+\t\tstruct ofs_delta_entry *entry = &ofs_deltas[next];\n+\n+\t\tif (obj_no == entry->obj_no)\n+\t\t\treturn entry;\n+\t\tif (obj_no < entry->obj_no) {\n+\t\t\tlast = next;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tfirst = next + 1;\n+\t}\n+\treturn NULL;\n+}\n+\n+static struct ref_delta_entry *find_ref_delta_obj_no(int obj_no)\n+{\n+\tint first = 0, last = nr_ref_deltas;\n+\n+\twhile (first < last) {\n+\t\tint next = first + (last - first) / 2;\n+\t\tstruct ref_delta_entry *entry = &ref_deltas[next];\n+\n+\t\tif (obj_no == entry->obj_no)\n+\t\t\treturn entry;\n+\t\tif (obj_no < entry->obj_no) {\n+\t\t\tlast = next;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tfirst = next + 1;\n+\t}\n+\treturn NULL;\n+}\n+\n+static struct object_entry *find_obj_offset(off_t offset)\n+{\n+\tint first = 0, last = nr_objects;\n+\n+\twhile (first < last) {\n+\t\tint next = first + (last - first) / 2;\n+\t\tstruct object_entry *entry = &objects[next];\n+\n+\t\tif (offset == entry->idx.offset)\n+\t\t\treturn entry;\n+\t\tif (offset < entry->idx.offset) {\n+\t\t\tlast = next;\n+\t\t\tcontinue;\n+\t\t}\n+\t\tfirst = next + 1;\n+\t}\n+\treturn NULL;\n+}\n+\n+static void *get_object_data(struct object_entry *obj, size_t *result_size)\n+{\n+\t/* Allow random reading objects */\n+\tvoid *data;\n+\n+\tif (!is_delta_type(obj->type)) {\n+\t\tdata = get_data_from_pack(obj);\n+\t\t*result_size = obj->size;\n+\t\treturn data;\n+\t}\n+\tif (obj->type == OBJ_OFS_DELTA) {\n+\t\tstruct ofs_delta_entry *delta;\n+\t\tstruct object_entry *bobj;\n+\t\tsize_t base_size;\n+\t\tvoid *base, *raw;\n+\n+\t\tdelta = find_ofs_delta_obj_no(obj - objects);\n+\t\tif (!delta)\n+\t\t\tBUG(\"Delta object without ofs_delta entry\");\n+\n+\t\tbobj = find_obj_offset(delta->offset);\n+\t\tif (!bobj)\n+\t\t\tBUG(\"Delta object without object entry\");\n+\n+\t\tbase = get_object_data(bobj, &base_size);\n+\t\traw = get_data_from_pack(obj);\n+\t\tdata = patch_delta(\n+\t\t\tbase, base_size,\n+\t\t\traw, obj->size,\n+\t\t\tresult_size);\n+\t\tif (!data)\n+\t\t\tBUG(\"patch_delta failed\");\n+\t\tfree(raw);\n+\t\tfree(base);\n+\t\treturn data;\n+\t}\n+\tif (obj->type == OBJ_REF_DELTA) {\n+\t\tstruct ref_delta_entry *delta;\n+\t\tenum object_type base_type;\n+\t\tsize_t base_size;\n+\t\tvoid *base, *raw;\n+\n+\t\tdelta = find_ref_delta_obj_no(obj - objects);\n+\t\tif (!delta)\n+\t\t\tBUG(\"Delta object without ref_delta entry\");\n+\n+\t\tbase = repo_read_object_file(the_repository, &delta->oid,\n+\t\t\t\t\t     &base_type, &base_size);\n+\t\tif (!base)\n+\t\t\tBUG(\"ref_delta oid %s not present in repository\",\n+\t\t\t    oid_to_hex(&delta->oid));\n+\t\traw = get_data_from_pack(obj);\n+\t\tdata = patch_delta(\n+\t\t\tbase, base_size,\n+\t\t\traw, obj->size,\n+\t\t\tresult_size);\n+\t\tif (!data)\n+\t\t\tBUG(\"patch_delta failed\");\n+\t\tfree(raw);\n+\t\tfree(base);\n+\t\treturn data;\n+\t}\n+\treturn NULL; /* The code never reaches here */\n+}\n+\n+static void compute_compat_oid(struct object_entry *obj)\n+{\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct object_file_convert_state state;\n+\tstruct strbuf out = STRBUF_INIT;\n+\tsize_t data_size;\n+\tvoid *data;\n+\tint ret;\n+\n+\tif (obj->idx.compat_oid.algo)\n+\t\treturn;\n+\n+\tif (obj->real_type == OBJ_BLOB)\n+\t\tdie(\"Blob object not converted\");\n+\n+\tdata = get_object_data(obj, &data_size);\n+\n+\tconvert_object_file_begin(&state, &out, algo, compat,\n+\t\t\t\t  data, data_size, obj->real_type);\n+\n+\tfor (;;) {\n+\t\tstruct object_entry *pobj;\n+\t\tret = convert_object_file_step(&state);\n+\t\tif (ret != 1)\n+\t\t\tbreak;\n+\t\t/* Does it name an object in the pack? */\n+\t\tpobj = find_in_oid_index(&state.oid);\n+\t\tif (pobj) {\n+\t\t\tcompute_compat_oid(pobj);\n+\t\t\toidcpy(&state.mapped_oid, &pobj->idx.compat_oid);\n+\t\t} else if (repo_oid_to_algop(repo, &state.oid, compat,\n+\t\t\t\t\t     &state.mapped_oid))\n+\t\t\tdie(_(\"No mapping for oid %s to %s\\n\"),\n+\t\t\t    oid_to_hex(&state.oid), compat->name);\n+\t}\n+\tconvert_object_file_end(&state, ret);\n+\tif (ret != 0)\n+\t\tdie(_(\"Bad object %s\\n\"), oid_to_hex(&obj->idx.oid));\n+\thash_object_file(compat, out.buf, out.len, obj->real_type,\n+\t\t\t &obj->idx.compat_oid);\n+\tstrbuf_release(&out);\n+\n+\tfree(data);\n+\n+\tnr_resolved_mappings++;\n+\tdisplay_progress(progress, nr_resolved_mappings);\n+}\n+\n+static void compute_compat_oids(void)\n+{\n+\tunsigned i;\n+\n+\tif (verbose)\n+\t\tprogress = start_progress(_(\"Mapping objects\"),\n+\t\t\tnr_pending_mappings);\n+\n+\t/* Sort deltas by obj_no for fast searching */\n+\tQSORT(ofs_deltas, nr_ofs_deltas, compare_ofs_delta_entry_obj_no);\n+\tQSORT(ref_deltas, nr_ref_deltas, compare_ref_delta_entry_obj_no);\n+\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\tstruct object_entry *obj = &objects[i];\n+\t\tif (obj->idx.compat_oid.algo)\n+\t\t\tcontinue;\n+\t\tif (is_delta_type(obj->real_type))\n+\t\t\tcontinue;\n+\t\tcompute_compat_oid(obj);\n+\t}\n+\n+\tstop_progress(&progress);\n+}\n+\n int cmd_index_pack(int argc, const char **argv, const char *prefix)\n {\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n \tint i, fix_thin_pack = 0, verify = 0, stat_only = 0, rev_index;\n \tconst char *curr_index;\n \tconst char *curr_rev_index = NULL;\n+\tconst char *curr_compat_index = NULL;\n \tconst char *index_name = NULL, *pack_name = NULL, *rev_index_name = NULL;\n+\tconst char *compat_index_name = NULL;\n \tconst char *keep_msg = NULL;\n \tconst char *promisor_msg = NULL;\n \tstruct strbuf index_name_buf = STRBUF_INIT;\n \tstruct strbuf rev_index_name_buf = STRBUF_INIT;\n+\tstruct strbuf compat_index_name_buf = STRBUF_INIT;\n \tstruct pack_idx_entry **idx_objects;\n \tstruct pack_idx_option opts;\n \tunsigned char pack_hash[GIT_MAX_RAWSZ];\n@@ -1946,6 +2250,12 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \t\t\t\t\t\t\t \"idx\", \"rev\",\n \t\t\t\t\t\t\t &rev_index_name_buf);\n \t}\n+\tif (compat) {\n+\t\tif (index_name)\n+\t\t\tcompat_index_name = derive_filename(index_name,\n+\t\t\t\t\t\t\t    \"idx\", \"compat\",\n+\t\t\t\t\t\t\t    &compat_index_name_buf);\n+\t}\n \n \tif (verify) {\n \t\tif (!index_name)\n@@ -1989,6 +2299,8 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \t\twrite_in_full(2, \"\\0\", 1);\n \tresolve_deltas();\n \tconclude_pack(fix_thin_pack, curr_pack, pack_hash);\n+\tif (compat)\n+\t\tcompute_compat_oids();\n \tfree(ofs_deltas);\n \tfree(ref_deltas);\n \tif (strict)\n@@ -1999,18 +2311,24 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \n \tALLOC_ARRAY(idx_objects, nr_objects);\n \tfor (i = 0; i < nr_objects; i++)\n-\t\tidx_objects[i] = &objects[i].idx;\n+\t\tidx_objects[i] = (struct pack_idx_entry *)&objects[i].idx;\n \tcurr_index = write_idx_file(index_name, idx_objects, nr_objects, &opts, pack_hash);\n \tif (rev_index)\n \t\tcurr_rev_index = write_rev_file(rev_index_name, idx_objects,\n \t\t\t\t\t\tnr_objects, pack_hash,\n \t\t\t\t\t\topts.flags);\n+\n+\tif (compat)\n+\t\tcurr_compat_index = write_compat_map_file(\n+\t\t\t(opts.flags & WRITE_IDX_VERIFY) ? compat_index_name : NULL,\n+\t\t\tidx_objects, nr_objects, pack_hash);\n \tfree(idx_objects);\n \n \tif (!verify)\n \t\tfinal(pack_name, curr_pack,\n \t\t      index_name, curr_index,\n \t\t      rev_index_name, curr_rev_index,\n+\t\t      compat_index_name, curr_compat_index,\n \t\t      keep_msg, promisor_msg,\n \t\t      pack_hash);\n \telse\n@@ -2023,12 +2341,15 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \tfree(objects);\n \tstrbuf_release(&index_name_buf);\n \tstrbuf_release(&rev_index_name_buf);\n+\tstrbuf_release(&compat_index_name_buf);\n \tif (!pack_name)\n \t\tfree((void *) curr_pack);\n \tif (!index_name)\n \t\tfree((void *) curr_index);\n \tif (!rev_index_name)\n \t\tfree((void *) curr_rev_index);\n+\tif (!compat_index_name)\n+\t\tfree((void *) curr_compat_index);\n \n \t/*\n \t * Let the caller know this pack is not self contained\n-- \n2.41.0\n\n"},{"id":"481593","messageId":"20230908231049.2035003-31-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 31/32] unpack-objects: Update to compute and write the compatibility hashes","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:48Z","receivedAt":"2023-09-08T23:12:48Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"To properly generate the compatibility hash objects that are referred\nto must be written before the objects that refer to them.  When\n--strict is set the unpack-objects already writes objects in that\norder.\n\nWhen a compatibilty hash is desired force use of the same code\npath that --strict uses.   If --strict is not wanted don't\nactually fsck the object buffers, just use fsck_walk to\nwalk to the parents of the objects recursively.\n\nUnlike in index-pack nothing special needs to be done when an object\nis written.  The guarantee that referred to objects are written to the\nloose object store before their refers ensures that the object\nmappings are in the loose object map.  The object mapings being in the\nloose object map guarantees that the call to convert_object_file can\nfind all of the mappings of the referred to objects.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/unpack-objects.c | 14 +++++++-------\n 1 file changed, 7 insertions(+), 7 deletions(-)\n\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex 32505255a009..834551142cd8 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -241,7 +241,8 @@ static int check_object(struct object *obj, enum object_type type,\n \tobj_buf = lookup_object_buffer(obj);\n \tif (!obj_buf)\n \t\tdie(\"Whoops! Cannot find object '%s'\", oid_to_hex(&obj->oid));\n-\tif (fsck_object(obj, obj_buf->buffer, obj_buf->size, &fsck_options))\n+\tif (strict &&\n+\t    fsck_object(obj, obj_buf->buffer, obj_buf->size, &fsck_options))\n \t\tdie(\"fsck error in packed object\");\n \tfsck_options.walk = check_object;\n \tif (fsck_walk(obj, NULL, &fsck_options))\n@@ -270,7 +271,7 @@ static void added_object(unsigned nr, enum object_type type,\n static void write_object(unsigned nr, enum object_type type,\n \t\t\t void *buf, unsigned long size)\n {\n-\tif (!strict) {\n+\tif (!strict && !the_repository->compat_hash_algo) {\n \t\tif (write_object_file(buf, size, type,\n \t\t\t\t      &obj_list[nr].oid) < 0)\n \t\t\tdie(\"failed to write object\");\n@@ -409,7 +410,7 @@ static void stream_blob(unsigned long size, unsigned nr)\n \t\tdie(_(\"inflate returned (%d)\"), data.status);\n \tgit_inflate_end(&zstream);\n \n-\tif (strict) {\n+\tif (strict || the_repository->compat_hash_algo) {\n \t\tstruct blob *blob = lookup_blob(the_repository, &info->oid);\n \n \t\tif (!blob)\n@@ -670,11 +671,10 @@ int cmd_unpack_objects(int argc, const char **argv, const char *prefix UNUSED)\n \tunpack_all();\n \tthe_hash_algo->update_fn(&ctx, buffer, offset);\n \tthe_hash_algo->final_oid_fn(&oid, &ctx);\n-\tif (strict) {\n+\tif (strict || the_repository->compat_hash_algo)\n \t\twrite_rest();\n-\t\tif (fsck_finish(&fsck_options))\n-\t\t\tdie(_(\"fsck error in pack objects\"));\n-\t}\n+\tif (strict && fsck_finish(&fsck_options))\n+\t\tdie(_(\"fsck error in pack objects\"));\n \tif (!hasheq(fill(the_hash_algo->rawsz), oid.hash))\n \t\tdie(\"final sha1 did not match\");\n \tuse(the_hash_algo->rawsz);\n-- \n2.41.0\n\n"},{"id":"481594","messageId":"20230908231049.2035003-16-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 16/32] object: Factor out parse_mode out of fast-import and tree-walk into in object.h","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:33Z","receivedAt":"2023-09-08T23:16:06Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"builtin/fast-import.c and tree-walk.c have almost identical version of\nget_mode.  The two functions started out the same but have diverged\nslightly.  The version in fast-import changed mode to a uint16_t to\nsave memory.  The version in tree-walk started erroring if no mode was\npresent.\n\nAs far as I can tell both of these changes are valid for both of the\ncallers, so add the both changes and place the common parsing helper\nin object.h\n\nRename the helper from get_mode to parse_mode so it does not\nconflict with another helper named get_mode in diff-no-index.c\n\nThis will be used shortly in a new helper decode_tree_entry_raw\nwhich is used to compute cmpatibility objects as part of\nthe sha256 transition.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/fast-import.c | 18 ++----------------\n object.h              | 18 ++++++++++++++++++\n tree-walk.c           | 22 +++-------------------\n 3 files changed, 23 insertions(+), 35 deletions(-)\n\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 4dbb10aff3da..2c645fcfbe3f 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -1235,20 +1235,6 @@ static void *gfi_unpack_entry(\n \treturn unpack_entry(the_repository, p, oe->idx.offset, &type, sizep);\n }\n \n-static const char *get_mode(const char *str, uint16_t *modep)\n-{\n-\tunsigned char c;\n-\tuint16_t mode = 0;\n-\n-\twhile ((c = *str++) != ' ') {\n-\t\tif (c < '0' || c > '7')\n-\t\t\treturn NULL;\n-\t\tmode = (mode << 3) + (c - '0');\n-\t}\n-\t*modep = mode;\n-\treturn str;\n-}\n-\n static void load_tree(struct tree_entry *root)\n {\n \tstruct object_id *oid = &root->versions[1].oid;\n@@ -1286,7 +1272,7 @@ static void load_tree(struct tree_entry *root)\n \t\tt->entries[t->entry_count++] = e;\n \n \t\te->tree = NULL;\n-\t\tc = get_mode(c, &e->versions[1].mode);\n+\t\tc = parse_mode(c, &e->versions[1].mode);\n \t\tif (!c)\n \t\t\tdie(\"Corrupt mode in %s\", oid_to_hex(oid));\n \t\te->versions[0].mode = e->versions[1].mode;\n@@ -2275,7 +2261,7 @@ static void file_change_m(const char *p, struct branch *b)\n \tstruct object_id oid;\n \tuint16_t mode, inline_data = 0;\n \n-\tp = get_mode(p, &mode);\n+\tp = parse_mode(p, &mode);\n \tif (!p)\n \t\tdie(\"Corrupt mode: %s\", command_buf.buf);\n \tswitch (mode) {\ndiff --git a/object.h b/object.h\nindex 114d45954d08..70c8d4ae63dc 100644\n--- a/object.h\n+++ b/object.h\n@@ -190,6 +190,24 @@ void *create_object(struct repository *r, const struct object_id *oid, void *obj\n \n void *object_as_type(struct object *obj, enum object_type type, int quiet);\n \n+\n+static inline const char *parse_mode(const char *str, uint16_t *modep)\n+{\n+\tunsigned char c;\n+\tunsigned int mode = 0;\n+\n+\tif (*str == ' ')\n+\t\treturn NULL;\n+\n+\twhile ((c = *str++) != ' ') {\n+\t\tif (c < '0' || c > '7')\n+\t\t\treturn NULL;\n+\t\tmode = (mode << 3) + (c - '0');\n+\t}\n+\t*modep = mode;\n+\treturn str;\n+}\n+\n /*\n  * Returns the object, having parsed it to find out what it is.\n  *\ndiff --git a/tree-walk.c b/tree-walk.c\nindex 29ead71be173..3af50a01c2c7 100644\n--- a/tree-walk.c\n+++ b/tree-walk.c\n@@ -10,27 +10,11 @@\n #include \"pathspec.h\"\n #include \"json-writer.h\"\n \n-static const char *get_mode(const char *str, unsigned int *modep)\n-{\n-\tunsigned char c;\n-\tunsigned int mode = 0;\n-\n-\tif (*str == ' ')\n-\t\treturn NULL;\n-\n-\twhile ((c = *str++) != ' ') {\n-\t\tif (c < '0' || c > '7')\n-\t\t\treturn NULL;\n-\t\tmode = (mode << 3) + (c - '0');\n-\t}\n-\t*modep = mode;\n-\treturn str;\n-}\n-\n static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned long size, struct strbuf *err)\n {\n \tconst char *path;\n-\tunsigned int mode, len;\n+\tunsigned int len;\n+\tuint16_t mode;\n \tconst unsigned hashsz = the_hash_algo->rawsz;\n \n \tif (size < hashsz + 3 || buf[size - (hashsz + 1)]) {\n@@ -38,7 +22,7 @@ static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned l\n \t\treturn -1;\n \t}\n \n-\tpath = get_mode(buf, &mode);\n+\tpath = parse_mode(buf, &mode);\n \tif (!path) {\n \t\tstrbuf_addstr(err, _(\"malformed mode in tree entry\"));\n \t\treturn -1;\n-- \n2.41.0\n\n"},{"id":"481595","messageId":"20230908231049.2035003-10-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 10/32] bulk-checkin: Only accept blobs","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:27Z","receivedAt":"2023-09-08T23:21:39Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"As the code is written today bulk_checkin only accepts blobs.  When\ndealing with multiple hash algorithms it is necessary to distinguish\nbetween blobs and object types that have embedded oids.  For object\nthat embed oids a completely new object needs to be generated to\ncompute the compatibility hash on.  For blobs however all that is\nneeded is to compute the compatibility hash on the same blob as the\ndefault hash.\n\nAs the code will soon need the compatiblity hash from\na bulk checkin remove support for a bulk checking of\nanything except blobs.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n bulk-checkin.c | 35 +++++++++++++++++------------------\n bulk-checkin.h |  6 +++---\n object-file.c  | 12 ++++++------\n 3 files changed, 26 insertions(+), 27 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 73bff3a23d27..223562b4e748 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -155,10 +155,10 @@ static int already_written(struct bulk_checkin_packfile *state, struct object_id\n  * status before calling us just in case we ask it to call us again\n  * with a new pack.\n  */\n-static int stream_to_pack(struct bulk_checkin_packfile *state,\n-\t\t\t  git_hash_ctx *ctx, off_t *already_hashed_to,\n-\t\t\t  int fd, size_t size, enum object_type type,\n-\t\t\t  const char *path, unsigned flags)\n+static int stream_blob_to_pack(struct bulk_checkin_packfile *state,\n+\t\t\t       git_hash_ctx *ctx, off_t *already_hashed_to,\n+\t\t\t       int fd, size_t size, const char *path,\n+\t\t\t       unsigned flags)\n {\n \tgit_zstream s;\n \tunsigned char ibuf[16384];\n@@ -170,7 +170,7 @@ static int stream_to_pack(struct bulk_checkin_packfile *state,\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n-\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), type, size);\n+\thdrlen = encode_in_pack_object_header(obuf, sizeof(obuf), OBJ_BLOB, size);\n \ts.next_out = obuf + hdrlen;\n \ts.avail_out = sizeof(obuf) - hdrlen;\n \n@@ -247,11 +247,10 @@ static void prepare_to_stream(struct bulk_checkin_packfile *state,\n \t\tdie_errno(\"unable to write pack header\");\n }\n \n-static int deflate_to_pack(struct bulk_checkin_packfile *state,\n-\t\t\t   struct object_id *result_oid,\n-\t\t\t   int fd, size_t size,\n-\t\t\t   enum object_type type, const char *path,\n-\t\t\t   unsigned flags)\n+static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n+\t\t\t\tstruct object_id *result_oid,\n+\t\t\t\tint fd, size_t size,\n+\t\t\t\tconst char *path, unsigned flags)\n {\n \toff_t seekback, already_hashed_to;\n \tgit_hash_ctx ctx;\n@@ -265,7 +264,7 @@ static int deflate_to_pack(struct bulk_checkin_packfile *state,\n \t\treturn error(\"cannot find the current offset\");\n \n \theader_len = format_object_header((char *)obuf, sizeof(obuf),\n-\t\t\t\t\t  type, size);\n+\t\t\t\t\t  OBJ_BLOB, size);\n \tthe_hash_algo->init_fn(&ctx);\n \tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n \n@@ -282,8 +281,8 @@ static int deflate_to_pack(struct bulk_checkin_packfile *state,\n \t\t\tidx->offset = state->offset;\n \t\t\tcrc32_begin(state->f);\n \t\t}\n-\t\tif (!stream_to_pack(state, &ctx, &already_hashed_to,\n-\t\t\t\t    fd, size, type, path, flags))\n+\t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n+\t\t\t\t\t fd, size, path, flags))\n \t\t\tbreak;\n \t\t/*\n \t\t * Writing this object to the current pack will make\n@@ -350,12 +349,12 @@ void fsync_loose_object_bulk_checkin(int fd, const char *filename)\n \t}\n }\n \n-int index_bulk_checkin(struct object_id *oid,\n-\t\t       int fd, size_t size, enum object_type type,\n-\t\t       const char *path, unsigned flags)\n+int index_blob_bulk_checkin(struct object_id *oid,\n+\t\t\t    int fd, size_t size,\n+\t\t\t    const char *path, unsigned flags)\n {\n-\tint status = deflate_to_pack(&bulk_checkin_packfile, oid, fd, size, type,\n-\t\t\t\t     path, flags);\n+\tint status = deflate_blob_to_pack(&bulk_checkin_packfile, oid, fd, size,\n+\t\t\t\t\t  path, flags);\n \tif (!odb_transaction_nesting)\n \t\tflush_bulk_checkin_packfile(&bulk_checkin_packfile);\n \treturn status;\ndiff --git a/bulk-checkin.h b/bulk-checkin.h\nindex 48fe9a6e9171..aa7286a7b3e1 100644\n--- a/bulk-checkin.h\n+++ b/bulk-checkin.h\n@@ -9,9 +9,9 @@\n void prepare_loose_object_bulk_checkin(void);\n void fsync_loose_object_bulk_checkin(int fd, const char *filename);\n \n-int index_bulk_checkin(struct object_id *oid,\n-\t\t       int fd, size_t size, enum object_type type,\n-\t\t       const char *path, unsigned flags);\n+int index_blob_bulk_checkin(struct object_id *oid,\n+\t\t\t    int fd, size_t size,\n+\t\t\t    const char *path, unsigned flags);\n \n /*\n  * Tell the object database to optimize for adding\ndiff --git a/object-file.c b/object-file.c\nindex 6a14b8875343..6cc4ae1fd957 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2587,11 +2587,11 @@ static int index_core(struct index_state *istate,\n  * binary blobs, they generally do not want to get any conversion, and\n  * callers should avoid this code path when filters are requested.\n  */\n-static int index_stream(struct object_id *oid, int fd, size_t size,\n-\t\t\tenum object_type type, const char *path,\n-\t\t\tunsigned flags)\n+static int index_blob_stream(struct object_id *oid, int fd, size_t size,\n+\t\t\t     const char *path,\n+\t\t\t     unsigned flags)\n {\n-\treturn index_bulk_checkin(oid, fd, size, type, path, flags);\n+\treturn index_blob_bulk_checkin(oid, fd, size, path, flags);\n }\n \n int index_fd(struct index_state *istate, struct object_id *oid,\n@@ -2613,8 +2613,8 @@ int index_fd(struct index_state *istate, struct object_id *oid,\n \t\tret = index_core(istate, oid, fd, xsize_t(st->st_size),\n \t\t\t\t type, path, flags);\n \telse\n-\t\tret = index_stream(oid, fd, xsize_t(st->st_size), type, path,\n-\t\t\t\t   flags);\n+\t\tret = index_blob_stream(oid, fd, xsize_t(st->st_size), path,\n+\t\t\t\t\tflags);\n \tclose(fd);\n \treturn ret;\n }\n-- \n2.41.0\n\n"},{"id":"481596","messageId":"20230908231049.2035003-23-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 23/32] builtin/ls-tree: Let the oid determine the output algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:40Z","receivedAt":"2023-09-08T23:22:44Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Update cmd_ls_tree to call get_oid_with_context and pass\nGET_OID_UNTRANSLATED instead of calling the simpler repo_get_oid.\n\nThis implments in ls-tree the behavior that asking to display a sha1\nhash displays the corrresponding sha1 encoded object and asking to\ndisplay a sha256 hash displayes the corresponding sha256 encoded\nobject.\n\nThis is useful for testing the conversion of an object to an\nequivlanet object encoded with a different hash function.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/ls-tree.c | 5 ++++-\n 1 file changed, 4 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/ls-tree.c b/builtin/ls-tree.c\nindex f558db5f3b80..346e3fd812eb 100644\n--- a/builtin/ls-tree.c\n+++ b/builtin/ls-tree.c\n@@ -376,6 +376,7 @@ int cmd_ls_tree(int argc, const char **argv, const char *prefix)\n \t\tOPT_END()\n \t};\n \tstruct ls_tree_cmdmode_to_fmt *m2f = ls_tree_cmdmode_format;\n+\tstruct object_context obj_context;\n \tint ret;\n \n \tgit_config(git_default_config, NULL);\n@@ -407,7 +408,9 @@ int cmd_ls_tree(int argc, const char **argv, const char *prefix)\n \t\t\tls_tree_usage, ls_tree_options);\n \tif (argc < 1)\n \t\tusage_with_options(ls_tree_usage, ls_tree_options);\n-\tif (repo_get_oid(the_repository, argv[0], &oid))\n+\tif (get_oid_with_context(the_repository, argv[0],\n+\t\t\t\t GET_OID_UNTRANSLATED, &oid,\n+\t\t\t\t &obj_context))\n \t\tdie(\"Not a valid object name %s\", argv[0]);\n \n \t/*\n-- \n2.41.0\n\n"},{"id":"481597","messageId":"20230908231049.2035003-12-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 12/32] bulk-checkin: hash object with compatibility algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:29Z","receivedAt":"2023-09-08T23:30:30Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nAny time we write an object into the repository when we're in dual hash\nmode, we need to compute both algorithms.  We already do this when we\nwrite a loose object into the repository, but we also need to do so in\nthe other case we write an object, which is the bulk check-in code.\n\n****\n\nWrite the compatibility hash into idx->compat_oid so it is available\nfor code that generates indexes that include the compatibilty\nmappings.\n\n--EWB\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n bulk-checkin.c | 24 ++++++++++++++++++++----\n 1 file changed, 20 insertions(+), 4 deletions(-)\n\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 223562b4e748..3206412a19e0 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -156,7 +156,8 @@ static int already_written(struct bulk_checkin_packfile *state, struct object_id\n  * with a new pack.\n  */\n static int stream_blob_to_pack(struct bulk_checkin_packfile *state,\n-\t\t\t       git_hash_ctx *ctx, off_t *already_hashed_to,\n+\t\t\t       git_hash_ctx *ctx, git_hash_ctx *compat_ctx,\n+\t\t\t       off_t *already_hashed_to,\n \t\t\t       int fd, size_t size, const char *path,\n \t\t\t       unsigned flags)\n {\n@@ -167,6 +168,7 @@ static int stream_blob_to_pack(struct bulk_checkin_packfile *state,\n \tint status = Z_OK;\n \tint write_object = (flags & HASH_WRITE_OBJECT);\n \toff_t offset = 0;\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n \n \tgit_deflate_init(&s, pack_compression_level);\n \n@@ -188,8 +190,11 @@ static int stream_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tsize_t hsize = offset - *already_hashed_to;\n \t\t\t\tif (rsize < hsize)\n \t\t\t\t\thsize = rsize;\n-\t\t\t\tif (hsize)\n+\t\t\t\tif (hsize) {\n \t\t\t\t\tthe_hash_algo->update_fn(ctx, ibuf, hsize);\n+\t\t\t\t\tif (compat)\n+\t\t\t\t\t\tcompat->update_fn(compat_ctx, ibuf, hsize);\n+\t\t\t\t}\n \t\t\t\t*already_hashed_to = offset;\n \t\t\t}\n \t\t\ts.next_in = ibuf;\n@@ -253,11 +258,13 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\tconst char *path, unsigned flags)\n {\n \toff_t seekback, already_hashed_to;\n-\tgit_hash_ctx ctx;\n+\tgit_hash_ctx ctx, compat_ctx;\n \tunsigned char obuf[16384];\n \tunsigned header_len;\n \tstruct hashfile_checkpoint checkpoint = {0};\n \tstruct pack_idx_entry *idx = NULL;\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n+\tstruct object_id compat_oid = {};\n \n \tseekback = lseek(fd, 0, SEEK_CUR);\n \tif (seekback == (off_t) -1)\n@@ -267,6 +274,10 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\t\t\t  OBJ_BLOB, size);\n \tthe_hash_algo->init_fn(&ctx);\n \tthe_hash_algo->update_fn(&ctx, obuf, header_len);\n+\tif (compat) {\n+\t\tcompat->init_fn(&compat_ctx);\n+\t\tcompat->update_fn(&compat_ctx, obuf, header_len);\n+\t}\n \n \t/* Note: idx is non-NULL when we are writing */\n \tif ((flags & HASH_WRITE_OBJECT) != 0)\n@@ -281,7 +292,8 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\tidx->offset = state->offset;\n \t\t\tcrc32_begin(state->f);\n \t\t}\n-\t\tif (!stream_blob_to_pack(state, &ctx, &already_hashed_to,\n+\t\tif (!stream_blob_to_pack(state, &ctx, &compat_ctx,\n+\t\t\t\t\t &already_hashed_to,\n \t\t\t\t\t fd, size, path, flags))\n \t\t\tbreak;\n \t\t/*\n@@ -298,6 +310,8 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\t\treturn error(\"cannot seek back\");\n \t}\n \tthe_hash_algo->final_oid_fn(result_oid, &ctx);\n+\tif (compat)\n+\t\tcompat->final_oid_fn(&compat_oid, &compat_ctx);\n \tif (!idx)\n \t\treturn 0;\n \n@@ -308,6 +322,8 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \t\tfree(idx);\n \t} else {\n \t\toidcpy(&idx->oid, result_oid);\n+\t\tif (compat)\n+\t\t\toidcpy(&idx->compat_oid, &compat_oid);\n \t\tALLOC_GROW(state->written,\n \t\t\t   state->nr_written + 1,\n \t\t\t   state->alloc_written);\n-- \n2.41.0\n\n"},{"id":"481598","messageId":"20230908231049.2035003-14-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 14/32] commit: write commits for both hashes","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:31Z","receivedAt":"2023-09-08T23:30:33Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nWhen we write a commit, we include data that is specific to the hash\nalgorithm, such as parents and the root tree.  In order to write both a\nSHA-1 commit and a SHA-256 version, we need to convert between them.\n\nHowever, a straightforward conversion isn't necessarily what we want.\nWhen we sign a commit, we sign its data, so if we create a commit for\nSHA-256 and then write a SHA-1 version, we'll still have only signed the\nSHA-256 data.  While this is valid, it would be better to sign both\nforms of data so people using SHA-1 can verify the signatures as well.\n\nConsequently, we don't want to use the standard mapping that occurs when\nwe write an object.  Instead, let's move most of the writing of the\ncommit into a separate function which is agnostic of the hash algorithm\nand which simply writes into a buffer and specify both versions of the\nobject ourselves.\n\nWe can then call this function twice: once with the SHA-256 contents,\nand if SHA-1 is enabled, once with the SHA-1 contents.  If we're signing\nthe commit, we then sign both versions and append both signatures to\nboth buffers.  To produce a consistent hash, we always append the\nsignatures in the order in which Git implemented them: first SHA-1, then\nSHA-256.\n\nIn order to make this signing code work, we split the commit signing\ncode into two functions, one which signs the buffer, and one which\nappends the signature.\n\n*****\n\nUpdated to use write_object_file_flags and repo_oid_to_algop\n\n-- EWB\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n commit.c | 176 +++++++++++++++++++++++++++++++++++++++++--------------\n 1 file changed, 131 insertions(+), 45 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex b3223478bc2a..522ebb4b3002 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -28,6 +28,7 @@\n #include \"shallow.h\"\n #include \"tree.h\"\n #include \"hook.h\"\n+#include \"object-file-convert.h\"\n \n static struct commit_extra_header *read_commit_extra_header_lines(const char *buf, size_t len, const char **);\n \n@@ -1100,12 +1101,11 @@ static const char *gpg_sig_headers[] = {\n \t\"gpgsig-sha256\",\n };\n \n-int sign_with_header(struct strbuf *buf, const char *keyid)\n+static int add_commit_signature(struct strbuf *buf, struct strbuf *sig, const struct git_hash_algo *algo)\n {\n-\tstruct strbuf sig = STRBUF_INIT;\n \tint inspos, copypos;\n \tconst char *eoh;\n-\tconst char *gpg_sig_header = gpg_sig_headers[hash_algo_by_ptr(the_hash_algo)];\n+\tconst char *gpg_sig_header = gpg_sig_headers[hash_algo_by_ptr(algo)];\n \tint gpg_sig_header_len = strlen(gpg_sig_header);\n \n \t/* find the end of the header */\n@@ -1115,15 +1115,8 @@ int sign_with_header(struct strbuf *buf, const char *keyid)\n \telse\n \t\tinspos = eoh - buf->buf + 1;\n \n-\tif (!keyid || !*keyid)\n-\t\tkeyid = get_signing_key();\n-\tif (sign_buffer(buf, &sig, keyid)) {\n-\t\tstrbuf_release(&sig);\n-\t\treturn -1;\n-\t}\n-\n-\tfor (copypos = 0; sig.buf[copypos]; ) {\n-\t\tconst char *bol = sig.buf + copypos;\n+\tfor (copypos = 0; sig->buf[copypos]; ) {\n+\t\tconst char *bol = sig->buf + copypos;\n \t\tconst char *eol = strchrnul(bol, '\\n');\n \t\tint len = (eol - bol) + !!*eol;\n \n@@ -1136,11 +1129,17 @@ int sign_with_header(struct strbuf *buf, const char *keyid)\n \t\tinspos += len;\n \t\tcopypos += len;\n \t}\n-\tstrbuf_release(&sig);\n \treturn 0;\n }\n \n-\n+static int sign_commit_to_strbuf(struct strbuf *sig, struct strbuf *buf, const char *keyid)\n+{\n+\tif (!keyid || !*keyid)\n+\t\tkeyid = get_signing_key();\n+\tif (sign_buffer(buf, sig, keyid))\n+\t\treturn -1;\n+\treturn 0;\n+}\n \n int parse_signed_commit(const struct commit *commit,\n \t\t\tstruct strbuf *payload, struct strbuf *signature,\n@@ -1599,70 +1598,157 @@ N_(\"Warning: commit message did not conform to UTF-8.\\n\"\n    \"You may want to amend it after fixing the message, or set the config\\n\"\n    \"variable i18n.commitEncoding to the encoding your project uses.\\n\");\n \n-int commit_tree_extended(const char *msg, size_t msg_len,\n-\t\t\t const struct object_id *tree,\n-\t\t\t struct commit_list *parents, struct object_id *ret,\n-\t\t\t const char *author, const char *committer,\n-\t\t\t const char *sign_commit,\n-\t\t\t struct commit_extra_header *extra)\n+static void write_commit_tree(struct strbuf *buffer, const char *msg, size_t msg_len,\n+\t\t\t      const struct object_id *tree,\n+\t\t\t      const struct object_id *parents, size_t parents_len,\n+\t\t\t      const char *author, const char *committer,\n+\t\t\t      struct commit_extra_header *extra)\n {\n-\tint result;\n \tint encoding_is_utf8;\n-\tstruct strbuf buffer;\n-\n-\tassert_oid_type(tree, OBJ_TREE);\n-\n-\tif (memchr(msg, '\\0', msg_len))\n-\t\treturn error(\"a NUL byte in commit log message not allowed.\");\n+\tsize_t i;\n \n \t/* Not having i18n.commitencoding is the same as having utf-8 */\n \tencoding_is_utf8 = is_encoding_utf8(git_commit_encoding);\n \n-\tstrbuf_init(&buffer, 8192); /* should avoid reallocs for the headers */\n-\tstrbuf_addf(&buffer, \"tree %s\\n\", oid_to_hex(tree));\n+\tstrbuf_init(buffer, 8192); /* should avoid reallocs for the headers */\n+\tstrbuf_addf(buffer, \"tree %s\\n\", oid_to_hex(tree));\n \n \t/*\n \t * NOTE! This ordering means that the same exact tree merged with a\n \t * different order of parents will be a _different_ changeset even\n \t * if everything else stays the same.\n \t */\n-\twhile (parents) {\n-\t\tstruct commit *parent = pop_commit(&parents);\n-\t\tstrbuf_addf(&buffer, \"parent %s\\n\",\n-\t\t\t    oid_to_hex(&parent->object.oid));\n-\t}\n+\tfor (i = 0; i < parents_len; i++)\n+\t\tstrbuf_addf(buffer, \"parent %s\\n\", oid_to_hex(&parents[i]));\n \n \t/* Person/date information */\n \tif (!author)\n \t\tauthor = git_author_info(IDENT_STRICT);\n-\tstrbuf_addf(&buffer, \"author %s\\n\", author);\n+\tstrbuf_addf(buffer, \"author %s\\n\", author);\n \tif (!committer)\n \t\tcommitter = git_committer_info(IDENT_STRICT);\n-\tstrbuf_addf(&buffer, \"committer %s\\n\", committer);\n+\tstrbuf_addf(buffer, \"committer %s\\n\", committer);\n \tif (!encoding_is_utf8)\n-\t\tstrbuf_addf(&buffer, \"encoding %s\\n\", git_commit_encoding);\n+\t\tstrbuf_addf(buffer, \"encoding %s\\n\", git_commit_encoding);\n \n \twhile (extra) {\n-\t\tadd_extra_header(&buffer, extra);\n+\t\tadd_extra_header(buffer, extra);\n \t\textra = extra->next;\n \t}\n-\tstrbuf_addch(&buffer, '\\n');\n+\tstrbuf_addch(buffer, '\\n');\n \n \t/* And add the comment */\n-\tstrbuf_add(&buffer, msg, msg_len);\n+\tstrbuf_add(buffer, msg, msg_len);\n+}\n \n-\t/* And check the encoding */\n-\tif (encoding_is_utf8 && !verify_utf8(&buffer))\n-\t\tfprintf(stderr, _(commit_utf8_warn));\n+int commit_tree_extended(const char *msg, size_t msg_len,\n+\t\t\t const struct object_id *tree,\n+\t\t\t struct commit_list *parents, struct object_id *ret,\n+\t\t\t const char *author, const char *committer,\n+\t\t\t const char *sign_commit,\n+\t\t\t struct commit_extra_header *extra)\n+{\n+\tstruct repository *r = the_repository;\n+\tint result = 0;\n+\tint encoding_is_utf8;\n+\tstruct strbuf buffer, compat_buffer;\n+\tstruct strbuf sig = STRBUF_INIT, compat_sig = STRBUF_INIT;\n+\tstruct object_id *parent_buf = NULL;\n+\tstruct object_id compat_oid = {};\n+\tsize_t i, nparents;\n+\n+\t/* Not having i18n.commitencoding is the same as having utf-8 */\n+\tencoding_is_utf8 = is_encoding_utf8(git_commit_encoding);\n+\n+\tassert_oid_type(tree, OBJ_TREE);\n+\n+\tif (memchr(msg, '\\0', msg_len))\n+\t\treturn error(\"a NUL byte in commit log message not allowed.\");\n+\n+\tnparents = commit_list_count(parents);\n+\tparent_buf = xcalloc(nparents, sizeof(*parent_buf));\n+\tfor (i = 0; i < nparents; i++) {\n+\t\tstruct commit *parent = pop_commit(&parents);\n+\t\toidcpy(&parent_buf[i], &parent->object.oid);\n+\t}\n \n-\tif (sign_commit && sign_with_header(&buffer, sign_commit)) {\n+\t/* should avoid reallocs for the headers */\n+\tstrbuf_init(&buffer, 8192);\n+\tstrbuf_init(&compat_buffer, 8192);\n+\n+\twrite_commit_tree(&buffer, msg, msg_len, tree, parent_buf, nparents, author, committer, extra);\n+\tif (sign_commit && sign_commit_to_strbuf(&sig, &buffer, sign_commit)) {\n \t\tresult = -1;\n \t\tgoto out;\n \t}\n+\tif (r->compat_hash_algo) {\n+\t\tstruct object_id mapped_tree;\n+\t\tstruct object_id *mapped_parents = xcalloc(nparents, sizeof(*mapped_parents));\n+\t\tif (repo_oid_to_algop(r, tree, r->compat_hash_algo, &mapped_tree)) {\n+\t\t\tresult = -1;\n+\t\t\tfree(mapped_parents);\n+\t\t\tgoto out;\n+\t\t}\n+\t\tfor (i = 0; i < nparents; i++)\n+\t\t\tif (repo_oid_to_algop(r, &parent_buf[i], r->compat_hash_algo, &mapped_parents[i])) {\n+\t\t\t\tresult = -1;\n+\t\t\t\tfree(mapped_parents);\n+\t\t\t\tgoto out;\n+\t\t\t}\n+\t\twrite_commit_tree(&compat_buffer, msg, msg_len, &mapped_tree,\n+\t\t\t\t  mapped_parents, nparents, author, committer, extra);\n+\n+\t\thash_object_file(r->compat_hash_algo, compat_buffer.buf, compat_buffer.len,\n+\t\t\t\t OBJ_COMMIT, &compat_oid);\n \n-\tresult = write_object_file(buffer.buf, buffer.len, OBJ_COMMIT, ret);\n+\t\tif (sign_commit && sign_commit_to_strbuf(&compat_sig, &compat_buffer, sign_commit)) {\n+\t\t\tresult = -1;\n+\t\t\tgoto out;\n+\t\t}\n+\t}\n+\n+\tif (sign_commit) {\n+\t\tstruct sig_pairs {\n+\t\t\tstruct strbuf *sig;\n+\t\t\tconst struct git_hash_algo *algo;\n+\t\t} bufs [2] = {\n+\t\t\t{ &compat_sig, r->compat_hash_algo },\n+\t\t\t{ &sig, r->hash_algo },\n+\t\t};\n+\t\tint i;\n+\n+\t\t/*\n+\t\t * We write algorithms in the order they were implemented in\n+\t\t * Git to produce a stable hash when multiple algorithms are\n+\t\t * used.\n+\t\t */\n+\t\tif (r->compat_hash_algo && hash_algo_by_ptr(bufs[0].algo) > hash_algo_by_ptr(bufs[1].algo))\n+\t\t\tSWAP(bufs[0], bufs[1]);\n+\n+\t\t/*\n+\t\t * We traverse each algorithm in order, and apply the signature\n+\t\t * to each buffer.\n+\t\t */\n+\t\tfor (i = 0; i < ARRAY_SIZE(bufs); i++) {\n+\t\t\tif (!bufs[i].algo)\n+\t\t\t\tcontinue;\n+\t\t\tadd_commit_signature(&buffer, bufs[i].sig, bufs[i].algo);\n+\t\t\tif (r->compat_hash_algo)\n+\t\t\t\tadd_commit_signature(&compat_buffer, bufs[i].sig, bufs[i].algo);\n+\t\t}\n+\t}\n+\n+\t/* And check the encoding. */\n+\tif (encoding_is_utf8 && (!verify_utf8(&buffer) || !verify_utf8(&compat_buffer)))\n+\t\tfprintf(stderr, _(commit_utf8_warn));\n+\n+\tresult = write_object_file_flags(buffer.buf, buffer.len, OBJ_COMMIT,\n+\t\t\t\t\t ret, &compat_oid, 0);\n out:\n \tstrbuf_release(&buffer);\n+\tstrbuf_release(&compat_buffer);\n+\tstrbuf_release(&sig);\n+\tstrbuf_release(&compat_sig);\n \treturn result;\n }\n \n-- \n2.41.0\n\n"},{"id":"481599","messageId":"20230908231049.2035003-3-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 03/32] object-file-convert: Stubs for converting from one object format to another","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:20Z","receivedAt":"2023-09-08T23:30:38Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Two basic functions are provided:\n- convert_object_file Takes an object file it's type and hash algorithm\n  and converts it into the equivalent object file that would\n  have been generated with hash algorithm \"to\".\n\n  For blob objects there is no converstion to be done and it is an\n  error to use this function on them.\n\n  For commit, tree and tag objects that embedded oids are replaced by\n  the oids of the objects they refer to if those objects had been\n  generated with the hash \"to\".\n\n- repo_oid_to_algop which takes an oid that refers to an object file\n  and returns the oid of the equavalent object file generated\n  with the target hash algorithm.\n\nTwo core functions are modified:\n- oid_object_info_extended is updated to detect an oid encoding\n  that does not match the current repository, use repo_oid_to_algop\n  to find the correspoding oid in the current repository and to return\n  the data for the oid.\n\nThe pair of files object-file-convert.c and object-file-convert.h\nis introduced to hold as much of this logic as possible to keep\nthis conversion logic cleanly separated from everything else\nand in the hopes that someday the code will be clean enough\ngit can support compiling out support for sha1 and the\nvarious conversion functions.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Makefile              |  1 +\n object-file-convert.c | 55 ++++++++++++++++++++++++++++\n object-file-convert.h | 24 +++++++++++++\n object-file.c         | 83 +++++++++++++++++++++++++++++++++++++++++++\n 4 files changed, 163 insertions(+)\n create mode 100644 object-file-convert.c\n create mode 100644 object-file-convert.h\n\ndiff --git a/Makefile b/Makefile\nindex 577630936535..f7e824f25cda 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1073,6 +1073,7 @@ LIB_OBJS += notes-cache.o\n LIB_OBJS += notes-merge.o\n LIB_OBJS += notes-utils.o\n LIB_OBJS += notes.o\n+LIB_OBJS += object-file-convert.o\n LIB_OBJS += object-file.o\n LIB_OBJS += object-name.o\n LIB_OBJS += object.o\ndiff --git a/object-file-convert.c b/object-file-convert.c\nnew file mode 100644\nindex 000000000000..9f4d5b354f5f\n--- /dev/null\n+++ b/object-file-convert.c\n@@ -0,0 +1,55 @@\n+#include \"git-compat-util.h\"\n+#include \"gettext.h\"\n+#include \"strbuf.h\"\n+#include \"repository.h\"\n+#include \"hash-ll.h\"\n+#include \"object.h\"\n+#include \"object-file-convert.h\"\n+\n+int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n+\t\t      const struct git_hash_algo *to, struct object_id *dest)\n+{\n+\t/*\n+\t * If the source alogirthm is not set, then we're using the\n+\t * default hash algorithm for that object.\n+\t */\n+\tconst struct git_hash_algo *from =\n+\t\tsrc->algo ? &hash_algos[src->algo] : repo->hash_algo;\n+\n+\tif (from == to) {\n+\t\tif (src != dest)\n+\t\t\toidcpy(dest, src);\n+\t\treturn 0;\n+\t}\n+\treturn -1;\n+}\n+\n+int convert_object_file(struct strbuf *outbuf,\n+\t\t\tconst struct git_hash_algo *from,\n+\t\t\tconst struct git_hash_algo *to,\n+\t\t\tconst void *buf, size_t len,\n+\t\t\tenum object_type type,\n+\t\t\tint gentle)\n+{\n+\tint ret;\n+\n+\t/* Don't call this function when no conversion is necessary */\n+\tif ((from == to) || (type == OBJ_BLOB))\n+\t\tdie(\"Refusing noop object file conversion\");\n+\n+\tswitch (type) {\n+\tcase OBJ_COMMIT:\n+\tcase OBJ_TREE:\n+\tcase OBJ_TAG:\n+\tdefault:\n+\t\t/* Not implemented yet, so fail. */\n+\t\tret = -1;\n+\t\tbreak;\n+\t}\n+\tif (!ret)\n+\t\treturn 0;\n+\tif (gentle)\n+\t\treturn ret;\n+\tdie(_(\"Failed to convert object from %s to %s\"),\n+\t\tfrom->name, to->name);\n+}\ndiff --git a/object-file-convert.h b/object-file-convert.h\nnew file mode 100644\nindex 000000000000..a4f802aa8eea\n--- /dev/null\n+++ b/object-file-convert.h\n@@ -0,0 +1,24 @@\n+#ifndef OBJECT_CONVERT_H\n+#define OBJECT_CONVERT_H\n+\n+struct repository;\n+struct object_id;\n+struct git_hash_algo;\n+struct strbuf;\n+#include \"object.h\"\n+\n+int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n+\t\t      const struct git_hash_algo *to, struct object_id *dest);\n+\n+/*\n+ * Convert an object file from one hash algorithm to another algorithm.\n+ * Return -1 on failure, 0 on success.\n+ */\n+int convert_object_file(struct strbuf *outbuf,\n+\t\t\tconst struct git_hash_algo *from,\n+\t\t\tconst struct git_hash_algo *to,\n+\t\t\tconst void *buf, size_t len,\n+\t\t\tenum object_type type,\n+\t\t\tint gentle);\n+\n+#endif /* OBJECT_CONVERT_H */\ndiff --git a/object-file.c b/object-file.c\nindex 7dc0c4bfbba8..7f24f19b8a68 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -36,6 +36,7 @@\n #include \"quote.h\"\n #include \"packfile.h\"\n #include \"object-file.h\"\n+#include \"object-file-convert.h\"\n #include \"object-store.h\"\n #include \"oidtree.h\"\n #include \"path.h\"\n@@ -1660,10 +1661,92 @@ static int do_oid_object_info_extended(struct repository *r,\n \treturn 0;\n }\n \n+static int oid_object_info_convert(struct repository *r,\n+\t\t\t\t   const struct object_id *input_oid,\n+\t\t\t\t   struct object_info *input_oi, unsigned flags)\n+{\n+\tconst struct git_hash_algo *input_algo = &hash_algos[input_oid->algo];\n+\tint do_die = flags & OBJECT_INFO_DIE_IF_CORRUPT;\n+\tstruct strbuf type_name = STRBUF_INIT;\n+\tstruct object_info oi = *input_oi;\n+\tstruct object_id oid, delta_base_oid;\n+\tunsigned long size;\n+\tvoid *content;\n+\tint ret;\n+\n+\tif (repo_oid_to_algop(r, input_oid, the_hash_algo, &oid)) {\n+\t\tif (do_die)\n+\t\t\tdie(_(\"missing mapping of %s to %s\"),\n+\t\t\t    oid_to_hex(input_oid), the_hash_algo->name);\n+\t\treturn -1;\n+\t}\n+\n+\t/* Do we need to convert the delta base oid? */\n+\tif (oi.delta_base_oid)\n+\t\toi.delta_base_oid = &delta_base_oid;\n+\n+\t/* Do we need attributes that differ when converted? */\n+\tif (oi.sizep || oi.contentp) {\n+\t\toi.contentp = &content;\n+\t\toi.sizep = &size;\n+\t\toi.type_name = &type_name;\n+\t}\n+\n+\tret = oid_object_info_extended(r, &oid, &oi, flags);\n+\tif (ret)\n+\t\treturn -1;\n+\n+\tif (oi.contentp == &content) {\n+\t\tstruct strbuf outbuf = STRBUF_INIT;\n+\t\tenum object_type type;\n+\n+\t\ttype = type_from_string_gently(type_name.buf, type_name.len,\n+\t\t\t\t\t       !do_die);\n+\t\tif (type == -1)\n+\t\t\treturn -1;\n+\t\tif (type != OBJ_BLOB) {\n+\t\t\tret = convert_object_file(&outbuf,\n+\t\t\t\t\t\t  the_hash_algo, input_algo,\n+\t\t\t\t\t\t  content, size, type, !do_die);\n+\t\t\tif (ret == -1)\n+\t\t\t\treturn -1;\n+\t\t\tsize = outbuf.len;\n+\t\t\tcontent = strbuf_detach(&outbuf, NULL);\n+\t\t}\n+\t\tif (input_oi->sizep)\n+\t\t\t*input_oi->sizep = size;\n+\t\tif (input_oi->contentp)\n+\t\t\t*input_oi->contentp = content;\n+\t\telse\n+\t\t\tfree(content);\n+\t\tif (input_oi->type_name)\n+\t\t\t*input_oi->type_name = type_name;\n+\t\telse\n+\t\t\tstrbuf_release(&type_name);\n+\t}\n+\tif (oi.delta_base_oid == &delta_base_oid) {\n+\t\tif (repo_oid_to_algop(r, &delta_base_oid, input_algo,\n+\t\t\t\t input_oi->delta_base_oid)) {\n+\t\t\tif (do_die)\n+\t\t\t\tdie(_(\"missing mapping of %s to %s\"),\n+\t\t\t\t    oid_to_hex(&delta_base_oid),\n+\t\t\t\t    input_algo->name);\n+\t\t\treturn -1;\n+\t\t}\n+\t}\n+\tinput_oi->whence = oi.whence;\n+\tinput_oi->u = oi.u;\n+\treturn ret;\n+}\n+\n int oid_object_info_extended(struct repository *r, const struct object_id *oid,\n \t\t\t     struct object_info *oi, unsigned flags)\n {\n \tint ret;\n+\n+\tif (oid->algo && (hash_algo_by_ptr(r->hash_algo) != oid->algo))\n+\t\treturn oid_object_info_convert(r, oid, oi, flags);\n+\n \tobj_read_lock();\n \tret = do_oid_object_info_extended(r, oid, oi, flags);\n \tobj_read_unlock();\n-- \n2.41.0\n\n"},{"id":"481600","messageId":"20230908231049.2035003-8-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 08/32] loose: Compatibilty short name support","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:25Z","receivedAt":"2023-09-08T23:30:42Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Update loose_objects_cache when udpating the loose objects map.  This\noidtree is used to discover which oids are possibilities when\nresolving short names, and it can support a mixture of sha1\nand sha256 oids.\n\nWith this any oid recorded objects/loose-objects-idx is usable\nfor resolving an oid to an object.\n\nTo make this maintainable a helper insert_loose_map is factored\nout of load_one_loose_object_map and repo_add_loose_object_map,\nand then modified to also update the loose_objects_cache.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n loose.c | 37 +++++++++++++++++++++++++------------\n 1 file changed, 25 insertions(+), 12 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 8ddb7112a541..81ed0e6b0c1e 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -7,6 +7,7 @@\n #include \"gettext.h\"\n #include \"loose.h\"\n #include \"lockfile.h\"\n+#include \"oidtree.h\"\n \n static const char *loose_object_header = \"# loose-object-idx\\n\";\n \n@@ -42,6 +43,21 @@ static int insert_oid_pair(kh_oid_map_t *map, const struct object_id *key, const\n \treturn 1;\n }\n \n+static int insert_loose_map(struct object_directory *odb,\n+\t\t\t    const struct object_id *oid,\n+\t\t\t    const struct object_id *compat_oid)\n+{\n+\tstruct loose_object_map *map = odb->loose_map;\n+\tint inserted = 0;\n+\n+\tinserted |= insert_oid_pair(map->to_compat, oid, compat_oid);\n+\tinserted |= insert_oid_pair(map->to_storage, compat_oid, oid);\n+\tif (inserted)\n+\t\toidtree_insert(odb->loose_objects_cache, compat_oid);\n+\n+\treturn inserted;\n+}\n+\n static int load_one_loose_object_map(struct repository *repo, struct object_directory *dir)\n {\n \tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n@@ -49,15 +65,14 @@ static int load_one_loose_object_map(struct repository *repo, struct object_dire\n \n \tif (!dir->loose_map)\n \t\tloose_object_map_init(&dir->loose_map);\n+\tif (!dir->loose_objects_cache) {\n+\t\tALLOC_ARRAY(dir->loose_objects_cache, 1);\n+\t\toidtree_init(dir->loose_objects_cache);\n+\t}\n \n-\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n-\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_tree, repo->hash_algo->empty_tree);\n-\n-\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n-\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->empty_blob, repo->hash_algo->empty_blob);\n-\n-\tinsert_oid_pair(dir->loose_map->to_compat, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n-\tinsert_oid_pair(dir->loose_map->to_storage, repo->compat_hash_algo->null_oid, repo->hash_algo->null_oid);\n+\tinsert_loose_map(dir, repo->hash_algo->empty_tree, repo->compat_hash_algo->empty_tree);\n+\tinsert_loose_map(dir, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n+\tinsert_loose_map(dir, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n \n \tstrbuf_git_common_path(&path, repo, \"objects/loose-object-idx\");\n \tfp = fopen(path.buf, \"rb\");\n@@ -75,8 +90,7 @@ static int load_one_loose_object_map(struct repository *repo, struct object_dire\n \t\t    parse_oid_hex_algop(p, &compat_oid, &p, repo->compat_hash_algo) ||\n \t\t    p != buf.buf + buf.len)\n \t\t\tgoto err;\n-\t\tinsert_oid_pair(dir->loose_map->to_compat, &oid, &compat_oid);\n-\t\tinsert_oid_pair(dir->loose_map->to_storage, &compat_oid, &oid);\n+\t\tinsert_loose_map(dir, &oid, &compat_oid);\n \t}\n \n \tstrbuf_release(&buf);\n@@ -195,8 +209,7 @@ int repo_add_loose_object_map(struct repository *repo, const struct object_id *o\n \tif (!should_use_loose_object_map(repo))\n \t\treturn 0;\n \n-\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_compat, oid, compat_oid);\n-\tinserted |= insert_oid_pair(repo->objects->odb->loose_map->to_storage, compat_oid, oid);\n+\tinserted = insert_loose_map(repo->objects->odb, oid, compat_oid);\n \tif (inserted)\n \t\treturn write_one_object(repo, oid, compat_oid);\n \treturn 0;\n-- \n2.41.0\n\n"},{"id":"481601","messageId":"20230908231049.2035003-1-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 01/32] doc hash-file-transition: A map file for mapping between sha1 and sha256","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:18Z","receivedAt":"2023-09-08T23:31:00Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"The v3 pack index file as documented has a lot of complexity making it\ndifficult to implement correctly.  I worked with bryan's preliminary\nimplementation and it took several passes to get the bugs out.\n\nThe complexity also requires multiple table look-ups to find all of\nthe information that is needed to translate from one kind of oid to\nanother.  Which can't be good for cache locality.\n\nEven worse coming up with a new index file version requires making\nchanges that have the potentialy to break anything that uses the index\nof a pack file.\n\nInstead of continuing to deal with the chance of braking things\nbesides the oid mapping functionality, the additional complexity in\nthe file format, and worry if the performance would be reasonable I\nstripped down the problem to it's fundamental complexity and came up\nwith a file format that is exactly about mapping one kind of oid to\nanother, and only supports two kinds of oids.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n .../technical/hash-function-transition.txt    | 40 +++++++++++++++++++\n 1 file changed, 40 insertions(+)\n\ndiff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\nindex ed574810891c..4b937480848a 100644\n--- a/Documentation/technical/hash-function-transition.txt\n+++ b/Documentation/technical/hash-function-transition.txt\n@@ -209,6 +209,46 @@ format described in linkgit:gitformat-pack[5], just like\n today. The content that is compressed and stored uses SHA-256 content\n instead of SHA-1 content.\n \n+Per Pack Mapping Table\n+~~~~~~~~~~~~~~~~~~~~~~\n+A pack compat map file (.compat) files have the following format:\n+\n+HEADER:\n+\t4-byte signature:\n+\t    The signature is: {'C', 'M', 'A', 'P'}\n+\t1-byte version number:\n+\t    Git only writes or recognizes version 1.\n+\t1-byte First Object Id Version\n+\t    We infer the length of object IDs (OIDs) from this value:\n+\t\t1 => SHA-1\n+\t\t2 => SHA-256\n+\t1-byte Second Object Id Version\n+\t    We infer the length of object IDs (OIDs) from this value:\n+\t\t1 => SHA-1\n+\t\t2 => SHA-256\n+\t1-byte reserved (must be zero)\n+\t4-byte number of objects names contained in this mapping\n+\t1-byte length in bytes of shorted object names for the first object id.\n+\t       This is the shortest possible length needed to make the\n+\t       first object names unambigious.\n+\t1-byte reserved (must be zero)\n+\t1-byte length in bytes of shorted object names for the second object id.\n+\t       This is the shortest possible length needed to make the\n+\t       second object names unambigious.\n+\t1-byte reserved (must be zero)\n+\n+OBJECT NAME TABLES:\n+\t[Object name raw length + 4]*Number of object names\n+\t   This table is sorted by object name\n+\t   Each entry in the table is formated as:\n+\t\t[20 or 32 byte] Object name\n+\t\t4-byte index into the other object name table\n+\n+TRAILER:\n+\tchecksum of the corresponding packfile, and\n+\n+\tchecksum of all of the above.\n+\n Pack index\n ~~~~~~~~~~\n Pack index (.idx) files use a new v3 format that supports multiple\n-- \n2.41.0\n\n"},{"id":"481602","messageId":"20230908231049.2035003-15-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 15/32] cache: add a function to read an OID of a specific algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:32Z","receivedAt":"2023-09-08T23:31:01Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nCurrently, we always read a object ID of the current algorithm with\noidread.  However, once we start converting objects, we'll need to\nconsider what happens when we want to read an object ID of a specific\nalgorithm, such as the compatibility algorithm.  To make this easier,\nlet's define oidread_algop, which specifies which algorithm we should\nuse for our object ID, and define oidread in terms of it.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n hash.h | 9 +++++++--\n 1 file changed, 7 insertions(+), 2 deletions(-)\n\ndiff --git a/hash.h b/hash.h\nindex 615ae0691d07..e064807c1733 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -73,10 +73,15 @@ static inline void oidclr(struct object_id *oid)\n \toid->algo = hash_algo_by_ptr(the_hash_algo);\n }\n \n+static inline void oidread_algop(struct object_id *oid, const unsigned char *hash, const struct git_hash_algo *algop)\n+{\n+\tmemcpy(oid->hash, hash, algop->rawsz);\n+\toid->algo = hash_algo_by_ptr(algop);\n+}\n+\n static inline void oidread(struct object_id *oid, const unsigned char *hash)\n {\n-\tmemcpy(oid->hash, hash, the_hash_algo->rawsz);\n-\toid->algo = hash_algo_by_ptr(the_hash_algo);\n+\toidread_algop(oid, hash, the_hash_algo);\n }\n \n static inline int is_empty_blob_sha1(const unsigned char *sha1)\n-- \n2.41.0\n\n"},{"id":"481603","messageId":"20230908231049.2035003-32-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 32/32] object-file-convert: Implement repo_submodule_oid_to_algop","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:49Z","receivedAt":"2023-09-08T23:31:07Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"From time to time git tree objects contain gitlinks.  These gitlinks\ncontain the oid of an object in another git repository.  To\nsuccesfully translate these oids it is necessary to look at the\nmapping tables in the submodules where the mapping tables live.\n\nLimiting myself to submodule interfaces I can see in the code\nrepo_submodule_oid_to_algop is the best I can figure out how to do,\nfor a gitlink agnostic implementation.\n\nThe big downsides are that the code as implemented is not thread\nsafe, it depends upon a worktree, and it always walks through\nall of the submodules.\n\nThere are interfaces in the code to lookup the submodule for an\nindividual gitlink.  As such iterating all of the submodules could be\navoided if care was taken to compute the path to the gitlink and to\nrecognizes the code is translating a gitlink.\n\nThe dependency on a worktree, and the thread safety issues\nI do not see a solution to short of reworking how git\ndeals with submodules.\n\nFor now repo_oid_to_algop does not call repo_submodule_oid_to_algop to\nallow avoiding the thread safety issues.\n\nUpdate callers of repo_oid_to_algop that can benefit from a submodule\ntranslation to also call repo_sumodule_oid_to_algop.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/fast-import.c |  5 ++++-\n builtin/index-pack.c  |  4 +++-\n object-file-convert.c | 45 +++++++++++++++++++++++++++++++++++++++++++\n object-file-convert.h |  5 +++++\n 4 files changed, 57 insertions(+), 2 deletions(-)\n\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex f1c250dd3c8f..66c471bc730e 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -1070,7 +1070,10 @@ static int store_object(\n \t\t\telse if (pobj)\n \t\t\t\tbreak;\n \t\t\telse if (repo_oid_to_algop(repo, &state.oid, compat,\n-\t\t\t\t\t\t   &state.mapped_oid))\n+\t\t\t\t\t\t   &state.mapped_oid) &&\n+\t\t\t\t repo_submodule_oid_to_algop(repo, &state.oid,\n+\t\t\t\t\t\t\t     compat,\n+\t\t\t\t\t\t\t     &state.mapped_oid))\n \t\t\t\tbreak;\n \t\t}\n \t\tconvert_object_file_end(&state, ret);\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 6827d14b91ce..4100fd56a845 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -2093,7 +2093,9 @@ static void compute_compat_oid(struct object_entry *obj)\n \t\telse if (pobj)\n \t\t\tcco = cco_push(cco, pobj);\n \t\telse if (repo_oid_to_algop(repo, &cco->state.oid, compat,\n-\t\t\t\t\t   &cco->state.mapped_oid))\n+\t\t\t\t\t   &cco->state.mapped_oid) &&\n+\t\t\t repo_submodule_oid_to_algop(repo, &cco->state.oid, compat,\n+\t\t\t\t\t\t     &cco->state.mapped_oid))\n \t\t\tdie(_(\"When converting %s no mapping for oid %s to %s\\n\"),\n \t\t\t    oid_to_hex(&cco->obj->idx.oid),\n \t\t\t    oid_to_hex(&cco->state.oid),\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 3fd080ebc112..2306e17dd57e 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -11,6 +11,45 @@\n #include \"gpg-interface.h\"\n #include \"pack-compat-map.h\"\n #include \"object-file-convert.h\"\n+#include \"read-cache.h\"\n+#include \"submodule-config.h\"\n+\n+int repo_submodule_oid_to_algop(struct repository *repo,\n+\t\t\t\tconst struct object_id *src,\n+\t\t\t\tconst struct git_hash_algo *to,\n+\t\t\t\tstruct object_id *dest)\n+{\n+\tint i;\n+\n+\tif (repo_read_index(repo) < 0)\n+\t\tdie(_(\"index file corrupt\"));\n+\n+\tfor (i = 0; i < repo->index->cache_nr; i++) {\n+\t\tconst struct cache_entry *ce = repo->index->cache[i];\n+\t\tstruct repository subrepo = {};\n+\t\tint ret;\n+\n+\t\tif (!S_ISGITLINK(ce->ce_mode))\n+\t\t\tcontinue;\n+\n+\t\twhile (i + 1 < repo->index->cache_nr &&\n+\t\t       !strcmp(ce->name, repo->index->cache[i + 1]->name))\n+\t\t\t/*\n+\t\t\t * Skip entries with the same name in different stages\n+\t\t\t * to make sure an entry is returned only once.\n+\t\t\t */\n+\t\t\ti++;\n+\n+\t\tif (repo_submodule_init(&subrepo, repo, ce->name, null_oid()))\n+\t\t\tcontinue;\n+\n+\t\tret = repo_oid_to_algop(&subrepo, src, to, dest);\n+\t\trepo_clear(&subrepo);\n+\t\tif (ret == 0)\n+\t\t\treturn 0;\n+\t}\n+\treturn -1;\n+}\n \n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \t\t      const struct git_hash_algo *to, struct object_id *dest)\n@@ -34,6 +73,7 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \t\t */\n \t\tif (!repo_packed_oid_to_algop(repo, src, to, dest))\n \t\t\treturn 0;\n+\n \t\t/*\n \t\t * We may have loaded the object map at repo initialization but\n \t\t * another process (perhaps upstream of a pipe from us) may have\n@@ -306,6 +346,11 @@ int convert_object_file(struct strbuf *outbuf,\n \t\t\tbreak;\n \t\tret = repo_oid_to_algop(the_repository, &state.oid, state.to,\n \t\t\t\t\t&state.mapped_oid);\n+\t\tif (ret)\n+\t\t\tret = repo_submodule_oid_to_algop(the_repository,\n+\t\t\t\t\t\t\t  &state.oid,\n+\t\t\t\t\t\t\t  state.to,\n+\t\t\t\t\t\t\t  &state.mapped_oid);\n \t\tif (ret) {\n \t\t\terror(_(\"failed to map %s entry for %s\"),\n \t\t\t      type_name(type), oid_to_hex(&state.oid));\ndiff --git a/object-file-convert.h b/object-file-convert.h\nindex da032d7a91ef..7a19feda5f0c 100644\n--- a/object-file-convert.h\n+++ b/object-file-convert.h\n@@ -10,6 +10,11 @@ struct strbuf;\n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \t\t      const struct git_hash_algo *to, struct object_id *dest);\n \n+int repo_submodule_oid_to_algop(struct repository *repo,\n+\t\t\t\tconst struct object_id *src,\n+\t\t\t\tconst struct git_hash_algo *to,\n+\t\t\t\tstruct object_id *dest);\n+\n struct object_file_convert_state {\n \tstruct strbuf *outbuf;\n \tconst struct git_hash_algo *from;\n-- \n2.41.0\n\n"},{"id":"481604","messageId":"20230908231049.2035003-30-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 30/32] builtin/index-pack: Make the stack in compute_compat_oid explicit","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:47Z","receivedAt":"2023-09-08T23:31:15Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Testing index-pack generating the compatibilty hashes on a large\nrepository (in this case the linux kernel) resulted in a stack\noverflow.  I confirmed this by using ulimit -s to force the stack to a\nmuch larger size, and rerunning the test and the code succeeded.\nStill it is not a good look to overflow the stack in the default\nconfiguration.\n\nIdeally the objects would be ordered such that no object has any\nreferences to any object that comes after it.  With such an ordering\nconvert_object_file followed by hash_object_file to could just be on\nevery object in order to compute the compatibility hashes for every\nobject.\n\nUnfortunately the work to compute such an order is roughly equaivalent\nto the depth first processing compute_compat_oid is doing.  The\nobjects have to be loaded to get which other objects they reference.\nKnowning which objects reference which others is necessary to compute\nsuch an order.\n\nLong story short I can see how to move the depth first traversal into\na topological sort, but that just moves the problem that caused the\ndeep recursion into another function, and makes everything more\nexpensive by requiring reading the objects yet another time.\n\nAvoid stack overflow by using an explicitly stack made of heap\nallocated objects instead of using the C call stack.\n\nTo get a feel for how much this explicit stack consumes I instrumented\nup the code.  Testing against a linux kernel 2.16GiB packfile.  This\npackfile had 9,033,248 objects, and 7,470,317 deltas.  There were\n6,543,758 mappings that cound not be computed opportunistically when\nthe data was first read.  In the function compute_compat_oid I\nmeasured a maximum cco stack depth of 66,415.  I measured a maximum\nmemory consumption of 103,783,520 bytes, or about 1563 bytes per level\nof the stack. In short call it 100MiB extra to compute the mappings in\na 2GiB packfile.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/index-pack.c | 106 +++++++++++++++++++++++++++++--------------\n 1 file changed, 71 insertions(+), 35 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex f5da671ed82d..6827d14b91ce 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -2015,54 +2015,90 @@ static void *get_object_data(struct object_entry *obj, size_t *result_size)\n \treturn NULL; /* The code never reaches here */\n }\n \n-static void compute_compat_oid(struct object_entry *obj)\n-{\n-\tstruct repository *repo = the_repository;\n-\tconst struct git_hash_algo *algo = repo->hash_algo;\n-\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+struct cco {\n+\tstruct cco *prev;\n+\tstruct object_entry *obj;\n \tstruct object_file_convert_state state;\n-\tstruct strbuf out = STRBUF_INIT;\n+\tstruct strbuf out;\n \tsize_t data_size;\n \tvoid *data;\n-\tint ret;\n+};\n \n-\tif (obj->idx.compat_oid.algo)\n-\t\treturn;\n+static struct cco *cco_push(struct cco *prev, struct object_entry *obj)\n+{\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct cco *cco;\n \n \tif (obj->real_type == OBJ_BLOB)\n-\t\tdie(\"Blob object not converted\");\n+\t\tBUG(\"Blob object not converted\");\n \n-\tdata = get_object_data(obj, &data_size);\n+\tcco = xmallocz(sizeof(*cco));\n+\tcco->prev = prev;\n+\tcco->obj = obj;\n+\tstrbuf_init(&cco->out, 0);\n \n-\tconvert_object_file_begin(&state, &out, algo, compat,\n-\t\t\t\t  data, data_size, obj->real_type);\n+\tcco->data = get_object_data(obj, &cco->data_size);\n \n-\tfor (;;) {\n-\t\tstruct object_entry *pobj;\n-\t\tret = convert_object_file_step(&state);\n-\t\tif (ret != 1)\n-\t\t\tbreak;\n-\t\t/* Does it name an object in the pack? */\n-\t\tpobj = find_in_oid_index(&state.oid);\n-\t\tif (pobj) {\n-\t\t\tcompute_compat_oid(pobj);\n-\t\t\toidcpy(&state.mapped_oid, &pobj->idx.compat_oid);\n-\t\t} else if (repo_oid_to_algop(repo, &state.oid, compat,\n-\t\t\t\t\t     &state.mapped_oid))\n-\t\t\tdie(_(\"No mapping for oid %s to %s\\n\"),\n-\t\t\t    oid_to_hex(&state.oid), compat->name);\n-\t}\n-\tconvert_object_file_end(&state, ret);\n-\tif (ret != 0)\n-\t\tdie(_(\"Bad object %s\\n\"), oid_to_hex(&obj->idx.oid));\n-\thash_object_file(compat, out.buf, out.len, obj->real_type,\n-\t\t\t &obj->idx.compat_oid);\n-\tstrbuf_release(&out);\n+\tconvert_object_file_begin(&cco->state, &cco->out, algo, compat,\n+\t\t\t\t  cco->data, cco->data_size, obj->real_type);\n+\treturn cco;\n+}\n \n-\tfree(data);\n+static struct cco *cco_pop(struct cco *cco, int ret)\n+{\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct cco *prev = cco->prev;\n+\n+\tconvert_object_file_end(&cco->state, ret);\n+\tif (ret != 0)\n+\t\tdie(_(\"Bad object %s\\n\"), oid_to_hex(&cco->obj->idx.oid));\n+\thash_object_file(compat, cco->out.buf, cco->out.len,\n+\t\t\t cco->obj->real_type, &cco->obj->idx.compat_oid);\n+\tstrbuf_release(&cco->out);\n+\tif (prev)\n+\t\toidcpy(&prev->state.mapped_oid, &cco->obj->idx.compat_oid);\n \n \tnr_resolved_mappings++;\n \tdisplay_progress(progress, nr_resolved_mappings);\n+\n+\tfree(cco->data);\n+\tfree(cco);\n+\n+\treturn prev;\n+}\n+\n+static void compute_compat_oid(struct object_entry *obj)\n+{\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct cco *cco;\n+\n+\tcco = cco_push(NULL, obj);\n+\tfor (;cco;) {\n+\t\tstruct object_entry *pobj;\n+\n+\t\tint ret = convert_object_file_step(&cco->state);\n+\t\tif (ret != 1) {\n+\t\t\tcco = cco_pop(cco, ret);\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\t/* Does it name an object in the pack? */\n+\t\tpobj = find_in_oid_index(&cco->state.oid);\n+\t\tif (pobj && pobj->idx.compat_oid.algo)\n+\t\t\toidcpy(&cco->state.mapped_oid, &pobj->idx.compat_oid);\n+\t\telse if (pobj)\n+\t\t\tcco = cco_push(cco, pobj);\n+\t\telse if (repo_oid_to_algop(repo, &cco->state.oid, compat,\n+\t\t\t\t\t   &cco->state.mapped_oid))\n+\t\t\tdie(_(\"When converting %s no mapping for oid %s to %s\\n\"),\n+\t\t\t    oid_to_hex(&cco->obj->idx.oid),\n+\t\t\t    oid_to_hex(&cco->state.oid),\n+\t\t\t    compat->name);\n+\t}\n }\n \n static void compute_compat_oids(void)\n-- \n2.41.0\n\n"},{"id":"481605","messageId":"20230908231049.2035003-28-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 28/32] builtin/index-pack: Add a simple oid index","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:45Z","receivedAt":"2023-09-08T23:31:19Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"To support computing the compatibility hash a way to lookup objects by\ntheir oid is needed.  This adds a simple hash table to enable looking\nup objects by their oid.  The implementation is inspired by the hash\ntable for looking up object_entries by their oid in struct packing_data,\nand implemented in pack-objects.c\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/index-pack.c | 68 +++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 67 insertions(+), 1 deletion(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex 006ffdc9c550..75c2113e455c 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -126,6 +126,9 @@ static int ref_deltas_alloc;\n static int nr_resolved_deltas;\n static int nr_threads;\n \n+static int32_t *oid_index;\n+static uint32_t oid_index_size;\n+\n static int from_stdin;\n static int strict;\n static int do_fsck_object;\n@@ -183,6 +186,62 @@ static inline void unlock_mutex(pthread_mutex_t *mutex)\n \t\tpthread_mutex_unlock(mutex);\n }\n \n+static uint32_t locate_oid_index(const struct object_id *oid, int *found)\n+{\n+\tuint32_t i, mask = (oid_index_size - 1);\n+\n+\ti = oidhash(oid) & mask;\n+\n+\twhile (oid_index[i] > 0) {\n+\t\tuint32_t pos = oid_index[i] - 1;\n+\n+\t\tif (oideq(oid, &objects[pos].idx.oid)) {\n+\t\t\t*found = 1;\n+\t\t\treturn i;\n+\t\t}\n+\n+\t\ti = (i + 1) & mask;\n+\t}\n+\n+\t*found = 0;\n+\treturn i;\n+}\n+\n+static void place_in_oid_index(struct object_entry *obj)\n+{\n+\tint found;\n+\tuint32_t pos = locate_oid_index(&obj->idx.oid, &found);\n+\n+\t/* Ignore duplicates */\n+\tif (found)\n+\t\treturn;\n+\n+\toid_index[pos] = (obj - objects) + 1;\n+}\n+\n+static struct object_entry *find_in_oid_index(struct object_id *oid)\n+{\n+\tuint32_t i;\n+\tint found;\n+\n+\ti = locate_oid_index(oid, &found);\n+\tif (!found)\n+\t\treturn NULL;\n+\n+\treturn &objects[oid_index[i] - 1];\n+}\n+\n+static inline uint32_t closest_pow2(uint32_t v)\n+{\n+\tv = v - 1;\n+\tv |= v >> 1;\n+\tv |= v >> 2;\n+\tv |= v >> 4;\n+\tv |= v >> 8;\n+\tv |= v >> 16;\n+\treturn v + 1;\n+}\n+\n /*\n  * Mutex and conditional variable can't be statically-initialized on Windows.\n  */\n@@ -987,6 +1046,7 @@ static struct base_data *resolve_delta(struct object_entry *delta_obj,\n \t\tbad_object(delta_obj->idx.offset, _(\"failed to apply delta\"));\n \thash_object_file(the_hash_algo, result_data, result_size,\n \t\t\t delta_obj->real_type, &delta_obj->idx.oid);\n+\tplace_in_oid_index(delta_obj);\n \tsha1_object(result_data, NULL, result_size, delta_obj->real_type,\n \t\t    &delta_obj->idx.oid);\n \n@@ -1188,12 +1248,16 @@ static void parse_pack_objects(unsigned char *hash)\n \t\t\tref_deltas[nr_ref_deltas].obj_no = i;\n \t\t\tnr_ref_deltas++;\n \t\t} else if (!data) {\n+\t\t\tplace_in_oid_index(obj);\n+\n \t\t\t/* large blobs, check later */\n \t\t\tobj->real_type = OBJ_BAD;\n \t\t\tnr_delays++;\n-\t\t} else\n+\t\t} else {\n+\t\t\tplace_in_oid_index(obj);\n \t\t\tsha1_object(data, NULL, obj->size, obj->type,\n \t\t\t\t    &obj->idx.oid);\n+\t\t}\n \t\tfree(data);\n \t\tdisplay_progress(progress, i+1);\n \t}\n@@ -1918,6 +1982,8 @@ int cmd_index_pack(int argc, const char **argv, const char *prefix)\n \tif (show_stat)\n \t\tCALLOC_ARRAY(obj_stat, st_add(nr_objects, 1));\n \tCALLOC_ARRAY(ofs_deltas, nr_objects);\n+\toid_index_size = closest_pow2(nr_objects * 3);\n+\tCALLOC_ARRAY(oid_index, oid_index_size);\n \tparse_pack_objects(pack_hash);\n \tif (report_end_of_input)\n \t\twrite_in_full(2, \"\\0\", 1);\n-- \n2.41.0\n\n"},{"id":"481606","messageId":"20230908231049.2035003-25-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 25/32] pack-compat-map: Add support for .compat files of a packfile","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:42Z","receivedAt":"2023-09-08T23:31:28Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"These .compat files hold a bidirectional mapping between the names of\nstored objects between sha1 and sha256.\n\nCare has been taken so that index-pack --verify can be supported to\nvalidate an existing compat map file is not currupted.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n Makefile                  |   2 +\n builtin.h                 |   1 +\n builtin/show-compat-map.c | 139 ++++++++++++++++\n git.c                     |   1 +\n object-file-convert.c     |   7 +\n object-name.c             |  18 ++\n object-store-ll.h         |   6 +\n pack-compat-map.c         | 334 ++++++++++++++++++++++++++++++++++++++\n pack-compat-map.h         |  27 +++\n pack-write.c              | 158 ++++++++++++++++++\n packfile.c                |  12 ++\n 11 files changed, 705 insertions(+)\n create mode 100644 builtin/show-compat-map.c\n create mode 100644 pack-compat-map.c\n create mode 100644 pack-compat-map.h\n\ndiff --git a/Makefile b/Makefile\nindex 3c18664def9a..b3f3dbe7bfeb 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1088,6 +1088,7 @@ LIB_OBJS += pack-check.o\n LIB_OBJS += pack-mtimes.o\n LIB_OBJS += pack-objects.o\n LIB_OBJS += pack-revindex.o\n+LIB_OBJS += pack-compat-map.o\n LIB_OBJS += pack-write.o\n LIB_OBJS += packfile.o\n LIB_OBJS += pager.o\n@@ -1299,6 +1300,7 @@ BUILTIN_OBJS += builtin/send-pack.o\n BUILTIN_OBJS += builtin/shortlog.o\n BUILTIN_OBJS += builtin/show-branch.o\n BUILTIN_OBJS += builtin/show-index.o\n+BUILTIN_OBJS += builtin/show-compat-map.o\n BUILTIN_OBJS += builtin/show-ref.o\n BUILTIN_OBJS += builtin/sparse-checkout.o\n BUILTIN_OBJS += builtin/stash.o\ndiff --git a/builtin.h b/builtin.h\nindex d560baa6618a..25882d281dd2 100644\n--- a/builtin.h\n+++ b/builtin.h\n@@ -223,6 +223,7 @@ int cmd_shortlog(int argc, const char **argv, const char *prefix);\n int cmd_show(int argc, const char **argv, const char *prefix);\n int cmd_show_branch(int argc, const char **argv, const char *prefix);\n int cmd_show_index(int argc, const char **argv, const char *prefix);\n+int cmd_show_compat_map(int argc, const char **argv, const char *prefix);\n int cmd_sparse_checkout(int argc, const char **argv, const char *prefix);\n int cmd_status(int argc, const char **argv, const char *prefix);\n int cmd_stash(int argc, const char **argv, const char *prefix);\ndiff --git a/builtin/show-compat-map.c b/builtin/show-compat-map.c\nnew file mode 100644\nindex 000000000000..8cc10bdaab61\n--- /dev/null\n+++ b/builtin/show-compat-map.c\n@@ -0,0 +1,139 @@\n+#include \"builtin.h\"\n+#include \"gettext.h\"\n+#include \"hash.h\"\n+#include \"hex.h\"\n+#include \"pack.h\"\n+#include \"parse-options.h\"\n+#include \"repository.h\"\n+\n+static const char *const show_compat_map_usage[] = {\n+\t\"git show-compat-map [--verbose] \",\n+\tNULL\n+};\n+\n+struct pack_compat_map_header {\n+\tuint8_t sig[4];\n+\tuint8_t version;\n+\tuint8_t first_oid_version;\n+\tuint8_t second_oid_version;\n+\tuint8_t mbz1;\n+\tuint32_t nr_objects;\n+\tuint8_t first_abbrev_len;\n+\tuint8_t mbz2;\n+\tuint8_t second_abbrev_len;\n+\tuint8_t mbz3;\n+};\n+\n+struct map_entry {\n+\tstruct object_id oid;\n+\tuint32_t index;\n+};\n+\n+static const struct git_hash_algo *from_oid_version(unsigned oid_version)\n+{\n+\tif (oid_version == 1) {\n+\t\treturn &hash_algos[GIT_HASH_SHA1];\n+\t} else if (oid_version == 2) {\n+\t\treturn &hash_algos[GIT_HASH_SHA256];\n+\t}\n+\tdie(\"unknown oid version %u\\n\", oid_version);\n+}\n+\n+static void read_half_map(struct map_entry *map, unsigned nr,\n+\t\t     const struct git_hash_algo *algo)\n+{\n+\tunsigned i;\n+\tfor (i = 0; i < nr; i++) {\n+\t\tuint32_t index;\n+\t\tif (fread(map[i].oid.hash, algo->rawsz, 1, stdin) != 1)\n+\t\t\tdie(\"unable to read hash of %s entry %u/%u\",\n+\t\t\t    algo->name, i, nr);\n+\t\tif (fread(&index, 4, 1, stdin) != 1)\n+\t\t\tdie(\"unable to read index of %s entry %u/%u\",\n+\t\t\t    algo->name, i, nr);\n+\t\tmap[i].oid.algo = hash_algo_by_ptr(algo);\n+\t\tmap[i].index = ntohl(index);\n+\t}\n+}\n+\n+static void print_half_map(const struct map_entry *map,\n+\t\t\t   unsigned nr)\n+{\n+\tunsigned i;\n+\tfor (i = 0; i < nr; i++) {\n+\t\tprintf(\"%s %\"PRIu32\"\\n\",\n+\t\t       oid_to_hex(&map[i].oid),\n+\t\t       map[i].index);\n+\t}\n+}\n+\n+static void print_map(const struct map_entry *map,\n+\t\t      const struct map_entry *compat_map,\n+\t\t      unsigned nr)\n+{\n+\tunsigned i;\n+\tfor (i = 0; i < nr; i++) {\n+\t\tprintf(\"%s \",\n+\t\t       oid_to_hex(&map[i].oid));\n+\t\tprintf(\"%s\\n\",\n+\t\t       oid_to_hex(&compat_map[map[i].index].oid));\n+\t}\n+}\n+\n+int cmd_show_compat_map(int argc, const char **argv, const char *prefix)\n+{\n+\tconst struct git_hash_algo *algo = NULL, *compat = NULL;\n+\tunsigned nr;\n+\tstruct pack_compat_map_header hdr;\n+\tstruct map_entry *map, *compat_map;\n+\tint verbose = 0;\n+\tconst struct option show_comapt_map_options[] = {\n+\t\tOPT_BOOL(0, \"verbose\", &verbose,\n+\t\t\t N_(\"print implementation details of the map file\")),\n+\t\tOPT_END()\n+\t};\n+\n+\targc = parse_options(argc, argv, prefix, show_comapt_map_options,\n+\t\t\t     show_compat_map_usage, 0);\n+\n+\tif (fread(&hdr, sizeof(hdr), 1, stdin) != 1)\n+\t\tdie(\"unable to read header\");\n+\tif ((hdr.sig[0] != 'C') ||\n+\t    (hdr.sig[1] != 'M') ||\n+\t    (hdr.sig[2] != 'A') ||\n+\t    (hdr.sig[3] != 'P'))\n+\t\tdie(\"Missing map signature\");\n+\tif (hdr.version != 1)\n+\t\tdie(\"Unknown map version\");\n+\tif ((hdr.mbz1 != 0) ||\n+\t    (hdr.mbz2 != 0) ||\n+\t    (hdr.mbz3 != 0))\n+\t\tdie(\"Must be zero fields non-zero\");\n+\n+\tnr = ntohl(hdr.nr_objects);\n+\n+\talgo = from_oid_version(hdr.first_oid_version);\n+\tcompat = from_oid_version(hdr.second_oid_version);\n+\n+\n+\tif (verbose) {\n+\t\tprintf(\"Map v%u for %u objects from %s to %s abbrevs (%u:%u)\\n\",\n+\t\t       hdr.version,\n+\t\t       nr,\n+\t\t       algo->name, compat->name,\n+\t\t       hdr.first_abbrev_len,\n+\t\t       hdr.second_abbrev_len);\n+\t}\n+\tALLOC_ARRAY(map, nr);\n+\tALLOC_ARRAY(compat_map, nr);\n+\tread_half_map(map, nr, algo);\n+\tread_half_map(compat_map, nr, compat);\n+\tif (verbose) {\n+\t\tprint_half_map(map, nr);\n+\t\tprint_half_map(compat_map, nr);\n+\t}\n+\tprint_map(map, compat_map, nr);\n+\tfree(compat_map);\n+\tfree(map);\n+\treturn 0;\n+}\ndiff --git a/git.c b/git.c\nindex c67e44dd82d2..bfaeece5ae0e 100644\n--- a/git.c\n+++ b/git.c\n@@ -606,6 +606,7 @@ static struct cmd_struct commands[] = {\n \t{ \"show\", cmd_show, RUN_SETUP },\n \t{ \"show-branch\", cmd_show_branch, RUN_SETUP },\n \t{ \"show-index\", cmd_show_index, RUN_SETUP_GENTLY },\n+\t{ \"show-compat-map\", cmd_show_compat_map, RUN_SETUP_GENTLY },\n \t{ \"show-ref\", cmd_show_ref, RUN_SETUP },\n \t{ \"sparse-checkout\", cmd_sparse_checkout, RUN_SETUP },\n \t{ \"stage\", cmd_add, RUN_SETUP | NEED_WORK_TREE },\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex d381d3d2ea65..7978aa63dfa9 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -9,6 +9,7 @@\n #include \"loose.h\"\n #include \"commit.h\"\n #include \"gpg-interface.h\"\n+#include \"pack-compat-map.h\"\n #include \"object-file-convert.h\"\n \n int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n@@ -27,6 +28,12 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \t\treturn 0;\n \t}\n \tif (repo_loose_object_map_oid(repo, dest, to, src)) {\n+\t\t/*\n+\t\t * It's not in the loose object map, so let's see if it's in a\n+\t\t * pack.\n+\t\t */\n+\t\tif (!repo_packed_oid_to_algop(repo, src, to, dest))\n+\t\t\treturn 0;\n \t\t/*\n \t\t * We may have loaded the object map at repo initialization but\n \t\t * another process (perhaps upstream of a pipe from us) may have\ndiff --git a/object-name.c b/object-name.c\nindex ebe87f5c4fdd..d33c82bc96ba 100644\n--- a/object-name.c\n+++ b/object-name.c\n@@ -26,6 +26,7 @@\n #include \"commit-reach.h\"\n #include \"date.h\"\n #include \"object-file-convert.h\"\n+#include \"pack-compat-map.h\"\n \n static int get_oid_oneline(struct repository *r, const char *, struct object_id *, struct commit_list *);\n \n@@ -210,6 +211,19 @@ static void find_short_packed_object(struct disambiguate_state *ds)\n \t\tunique_in_pack(p, ds);\n }\n \n+static void find_short_packed_compat_object(struct disambiguate_state *ds)\n+{\n+\tstruct packed_git *p;\n+\n+\t/* Skip, unless compatibility oids are wanted */\n+\tif (!ds->algo && (&hash_algos[ds->algo] != ds->repo->compat_hash_algo))\n+\t\treturn;\n+\n+\tfor (p = get_packed_git(ds->repo); p && !ds->ambiguous; p = p->next)\n+\t\tpack_compat_map_each(ds->repo, p, ds->bin_pfx.hash, ds->len,\n+\t\t\t\t     match_prefix, ds);\n+}\n+\n static int finish_object_disambiguation(struct disambiguate_state *ds,\n \t\t\t\t\tstruct object_id *oid)\n {\n@@ -581,6 +595,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \n \tfind_short_object_filename(&ds);\n \tfind_short_packed_object(&ds);\n+\tfind_short_packed_compat_object(&ds);\n \tstatus = finish_object_disambiguation(&ds, oid);\n \n \t/*\n@@ -592,6 +607,7 @@ static enum get_oid_result get_short_oid(struct repository *r,\n \t\treprepare_packed_git(r);\n \t\tfind_short_object_filename(&ds);\n \t\tfind_short_packed_object(&ds);\n+\t\tfind_short_packed_compat_object(&ds);\n \t\tstatus = finish_object_disambiguation(&ds, oid);\n \t}\n \n@@ -659,6 +675,7 @@ int repo_for_each_abbrev(struct repository *r, const char *prefix,\n \tds.cb_data = &collect;\n \tfind_short_object_filename(&ds);\n \tfind_short_packed_object(&ds);\n+\tfind_short_packed_compat_object(&ds);\n \n \tret = oid_array_for_each_unique(&collect, fn, cb_data);\n \toid_array_clear(&collect);\n@@ -871,6 +888,7 @@ int repo_find_unique_abbrev_r(struct repository *r, char *hex,\n \tds.cb_data = (void *)&mad;\n \n \tfind_short_object_filename(&ds);\n+\tfind_short_packed_compat_object(&ds);\n \t(void)finish_object_disambiguation(&ds, &oid_ret);\n \n \thex[mad.cur_len] = 0;\ndiff --git a/object-store-ll.h b/object-store-ll.h\nindex c5f2bb2fc2fe..c37c19ada0c3 100644\n--- a/object-store-ll.h\n+++ b/object-store-ll.h\n@@ -135,6 +135,12 @@ struct packed_git {\n \t */\n \tconst uint32_t *mtimes_map;\n \tsize_t mtimes_size;\n+\n+\tconst void *compat_mapping;\n+\tsize_t compat_mapping_size;\n+\tconst uint8_t *hash_map;\n+\tconst uint8_t *compat_hash_map;\n+\n \t/* something like \".git/objects/pack/xxxxx.pack\" */\n \tchar pack_name[FLEX_ARRAY]; /* more */\n };\ndiff --git a/pack-compat-map.c b/pack-compat-map.c\nnew file mode 100644\nindex 000000000000..3a992095ebe3\n--- /dev/null\n+++ b/pack-compat-map.c\n@@ -0,0 +1,334 @@\n+#include \"git-compat-util.h\"\n+#include \"gettext.h\"\n+#include \"hex.h\"\n+#include \"hash-ll.h\"\n+#include \"hash.h\"\n+#include \"object-store.h\"\n+#include \"object-file.h\"\n+#include \"packfile.h\"\n+#include \"pack-compat-map.h\"\n+#include \"packfile.h\"\n+\n+struct pack_compat_map_header {\n+\tuint8_t sig[4];\n+\tuint8_t version;\n+\tuint8_t first_oid_version;\n+\tuint8_t second_oid_version;\n+\tuint8_t mbz1;\n+\tuint32_t nr_objects;\n+\tuint8_t first_abbrev_len;\n+\tuint8_t mbz2;\n+\tuint8_t second_abbrev_len;\n+\tuint8_t mbz3;\n+};\n+\n+static char *pack_compat_map_filename(struct packed_git *p)\n+{\n+\tsize_t len;\n+\tif (!strip_suffix(p->pack_name, \".pack\", &len))\n+\t\tBUG(\"pack_name does not end in .pack\");\n+\treturn xstrfmt(\"%.*s.compat\", (int)len, p->pack_name);\n+}\n+\n+static int oid_version_match(const char *filename,\n+\t\t\t     unsigned oid_version,\n+\t\t\t     const struct git_hash_algo *algo)\n+{\n+\tconst struct git_hash_algo *found = NULL;\n+\tint ret = 0;\n+\n+\tif (oid_version == 1) {\n+\t\tfound = &hash_algos[GIT_HASH_SHA1];\n+\t} else if (oid_version == 2) {\n+\t\tfound = &hash_algos[GIT_HASH_SHA256];\n+\t}\n+\tif (found == NULL) {\n+\t\tret = error(_(\"compat map file %s hash version %u unknown\"),\n+\t\t\t    filename, oid_version);\n+\t}\n+\telse if (found != algo) {\n+\t\tret = error(_(\"compat map file %s found hash %s expected hash %s\"),\n+\t\t\t    filename, found->name, algo->name);\n+\t}\n+\treturn ret;\n+}\n+\n+\n+static int load_pack_compat_map_file(char *compat_map_file,\n+\t\t\t\t     struct repository *repo,\n+\t\t\t\t     struct packed_git *p)\n+{\n+\tconst struct pack_compat_map_header *hdr;\n+\tunsigned compat_map_objects = 0;\n+\tconst uint8_t *data = NULL;\n+\tconst uint8_t *packs_hash = NULL;\n+\tint fd, ret = 0;\n+\tstruct stat st;\n+\tsize_t size, map1sz, map2sz, expected_size;\n+\n+\tfd = git_open(compat_map_file);\n+\n+\tif (fd < 0) {\n+\t\tret = -1;\n+\t\tgoto cleanup;\n+\t}\n+\tif (fstat(fd, &st)) {\n+\t\tret = error_errno(_(\"failed to read %s\"), compat_map_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tsize = xsize_t(st.st_size);\n+\n+\tif (size < sizeof(struct pack_compat_map_header)) {\n+\t\tret = error(_(\"compat map file %s is too small\"), compat_map_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tdata = xmmap(NULL, size, PROT_READ, MAP_PRIVATE, fd, 0);\n+\n+\thdr = (const struct pack_compat_map_header *)data;\n+\tif ((hdr->sig[0] != 'C') ||\n+\t    (hdr->sig[1] != 'M') ||\n+\t    (hdr->sig[2] != 'A') ||\n+\t    (hdr->sig[3] != 'P')) {\n+\t\tret = error(_(\"compat map file %s has unknown signature\"),\n+\t\t\t    compat_map_file);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tif (hdr->version != 1) {\n+\t\tret = error(_(\"compat map file %s has unsupported version %\"PRIu8),\n+\t\t\t    compat_map_file, hdr->version);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tret = oid_version_match(compat_map_file, hdr->first_oid_version, repo->hash_algo);\n+\tif (ret)\n+\t\tgoto cleanup;\n+\tret = oid_version_match(compat_map_file, hdr->second_oid_version, repo->compat_hash_algo);\n+\tif (ret)\n+\t\tgoto cleanup;\n+\tcompat_map_objects = ntohl(hdr->nr_objects);\n+\tif (compat_map_objects != p->num_objects) {\n+\t\tret = error(_(\"compat map file %s number of objects found %u wanted %u\"),\n+\t\t\t    compat_map_file, compat_map_objects, p->num_objects);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tmap1sz = st_mult(repo->hash_algo->rawsz + 4, compat_map_objects);\n+\tmap2sz = st_mult(repo->compat_hash_algo->rawsz + 4, compat_map_objects);\n+\n+\texpected_size = sizeof(struct pack_compat_map_header);\n+\texpected_size = st_add(expected_size, map1sz);\n+\texpected_size = st_add(expected_size, map2sz);\n+\texpected_size = st_add(expected_size, 2 * repo->hash_algo->rawsz);\n+\n+\tif (size != expected_size) {\n+\t\tret = error(_(\"compat map file %s is corrupt size %zu expected %zu objects %u sz1 %zu sz2 %zu\"),\n+\t\t\t    compat_map_file, size, expected_size, compat_map_objects,\n+\t\t\t    map1sz, map2sz\n+\t\t\t);\n+\t\tgoto cleanup;\n+\t}\n+\n+\tpacks_hash = data + sizeof(struct pack_compat_map_header) + map1sz + map2sz;\n+\tif (hashcmp(packs_hash, p->hash)) {\n+\t\tret = error(_(\"compat map file %s does not match pack %s\\n\"),\n+\t\t\t      compat_map_file, hash_to_hex(p->hash));\n+\t}\n+\n+\n+\tp->compat_mapping = data;\n+\tp->compat_mapping_size = size;\n+\n+\tp->hash_map = data + sizeof(struct pack_compat_map_header);\n+\tp->compat_hash_map = p->hash_map + map1sz;\n+\n+cleanup:\n+\tif (ret) {\n+\t\tif (data) {\n+\t\t\tmunmap((void *)data, size);\n+\t\t}\n+\t}\n+\tif (fd >= 0)\n+\t\tclose(fd);\n+\treturn ret;\n+}\n+\n+int load_pack_compat_map(struct repository *repo, struct packed_git *p)\n+{\n+\tchar *compat_map_name = NULL;\n+\tint ret = 0;\n+\n+\tif (p->compat_mapping)\n+\t\treturn ret;\t/* already loaded */\n+\n+\tif (!repo->compat_hash_algo)\n+\t\treturn 1;\t\t/* Nothing to do */\n+\n+\tret = open_pack_index(p);\n+\tif (ret < 0)\n+\t\tgoto cleanup;\n+\n+\tcompat_map_name = pack_compat_map_filename(p);\n+\tret = load_pack_compat_map_file(compat_map_name, repo, p);\n+cleanup:\n+\tfree(compat_map_name);\n+\treturn ret;\n+}\n+\n+static int keycmp(const unsigned char *a, const unsigned char *b,\n+\t\t  size_t key_hex_size)\n+{\n+\tsize_t key_byte_size = key_hex_size / 2;\n+\tunsigned a_last, b_last, mask = (key_hex_size & 1) ? 0xf0 : 0;\n+\tint cmp = memcmp(a, b, key_byte_size);\n+\tif (cmp)\n+\t\treturn cmp;\n+\n+\ta_last = a[key_byte_size] & mask;\n+\tb_last = b[key_byte_size] & mask;\n+\n+\tif (a_last == b_last)\n+\t\tcmp = 0;\n+\telse if (a_last < b_last)\n+\t\tcmp = -1;\n+\telse\n+\t\tcmp = 1;\n+\n+\treturn cmp;\n+}\n+\n+static const uint8_t *bsearch_map(const unsigned char *hash,\n+\t\t\t\t  const uint8_t *table, unsigned nr,\n+\t\t\t\t  size_t entry_size, size_t key_hex_size)\n+{\n+\tuint32_t hi, lo;\n+\n+\thi = nr - 1;\n+\tlo = 0;\n+\twhile (lo < hi) {\n+\t\tunsigned mi = lo + ((hi - lo) / 2);\n+\t\tconst unsigned char *entry = table + (mi * entry_size);\n+\t\tint cmp = keycmp(entry, hash, key_hex_size);\n+\t\tif (!cmp)\n+\t\t\treturn entry;\n+\t\tif (cmp > 0)\n+\t\t\thi = mi;\n+\t\telse\n+\t\t\tlo = mi + 1;\n+\t}\n+\tif (lo == hi) {\n+\t\tconst unsigned char *entry = table + (lo * entry_size);\n+\t\tint cmp = keycmp(entry, hash, key_hex_size);\n+\t\tif (!cmp)\n+\t\t\treturn entry;\n+\t}\n+\treturn NULL;\n+}\n+\n+static void map_each(const struct git_hash_algo *compat,\n+\t\t     const unsigned char *prefix, size_t prefix_hexsz,\n+\t\t     const uint8_t *table, unsigned nr, size_t entry_bytes,\n+\t\t     compat_map_iter_t iter, void *data)\n+{\n+\tconst uint8_t *found, *last = table + (entry_bytes * nr);\n+\n+\tfound = bsearch_map(prefix, table, nr, entry_bytes, prefix_hexsz);\n+\tif (!found)\n+\t\treturn;\n+\n+\t/* Visit each matching key */\n+\tdo {\n+\t\tstruct object_id oid;\n+\n+\t\tif (keycmp(found, prefix, prefix_hexsz) != 0)\n+\t\t\tbreak;\n+\n+\t\toidread_algop(&oid, found, compat);\n+\t\tif (iter(&oid, data) == CB_BREAK)\n+\t\t\tbreak;\n+\n+\t\tfound = found + entry_bytes;\n+\t} while (found < last);\n+}\n+\n+void pack_compat_map_each(struct repository *repo, struct packed_git *p,\n+\t\t\t const unsigned char *prefix, size_t prefix_hexsz,\n+\t\t\t compat_map_iter_t iter, void *data)\n+{\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\n+\tif (!p->num_objects ||\n+\t    (!p->compat_mapping && load_pack_compat_map(repo, p)))\n+\t\treturn;\n+\n+\tif (prefix_hexsz > compat->hexsz)\n+\t\tprefix_hexsz = compat->hexsz;\n+\n+\tmap_each(compat, prefix, prefix_hexsz,\n+\t\t p->compat_hash_map, p->num_objects, compat->rawsz + 4,\n+\t\t iter, data);\n+}\n+\n+static int compat_map_to_algop(const struct object_id *src,\n+\t\t\t       const struct git_hash_algo *to,\n+\t\t\t       const struct git_hash_algo *from,\n+\t\t\t       const uint8_t *to_table,\n+\t\t\t       const uint8_t *from_table,\n+\t\t\t       unsigned nr,\n+\t\t\t       struct object_id *dest)\n+{\n+\tconst uint8_t *found;\n+\tuint32_t index;\n+\n+\tif (src->algo != hash_algo_by_ptr(from))\n+\t\treturn -1;\n+\n+\tfound = bsearch_map(src->hash,\n+\t\t\t    from_table, nr,\n+\t\t\t    from->rawsz + 4,\n+\t\t\t    from->hexsz);\n+\tif (!found)\n+\t\treturn -1;\n+\n+\tindex = ntohl(*(uint32_t *)(found + from->rawsz));\n+\toidread_algop(dest, to_table + index * (to->rawsz + 4), to);\n+\treturn 0;\n+}\n+\n+static int pack_to_algop(struct repository *repo, struct packed_git *p,\n+\t\t\t const struct object_id *src,\n+\t\t\t const struct git_hash_algo *to, struct object_id *dest)\n+{\n+\tif (!p->compat_mapping && load_pack_compat_map(repo, p))\n+\t\treturn -1;\n+\n+\tif (to == repo->hash_algo) {\n+\t\treturn compat_map_to_algop(src, to, repo->compat_hash_algo,\n+\t\t\t\t\t   p->hash_map,\n+\t\t\t\t\t   p->compat_hash_map,\n+\t\t\t\t\t   p->num_objects, dest);\n+\t}\n+\telse if (to == repo->compat_hash_algo) {\n+\t\treturn compat_map_to_algop(src, to, repo->hash_algo,\n+\t\t\t\t\t   p->compat_hash_map,\n+\t\t\t\t\t   p->hash_map,\n+\t\t\t\t\t   p->num_objects, dest);\n+\t}\n+\telse\n+\t\treturn -1;\n+}\n+\n+int repo_packed_oid_to_algop(struct repository *repo,\n+\t\t\t     const struct object_id *src,\n+\t\t\t     const struct git_hash_algo *to,\n+\t\t\t     struct object_id *dest)\n+{\n+\tstruct packed_git *p;\n+\tfor (p = get_packed_git(repo); p; p = p->next) {\n+\t\tif (!pack_to_algop(repo, p, src, to, dest))\n+\t\t\treturn 0;\n+\t}\n+\treturn -1;\n+}\ndiff --git a/pack-compat-map.h b/pack-compat-map.h\nnew file mode 100644\nindex 000000000000..2a4561ffdff6\n--- /dev/null\n+++ b/pack-compat-map.h\n@@ -0,0 +1,27 @@\n+#ifndef PACK_COMPAT_MAP_H\n+#define PACK_COMPAT_MAP_H\n+\n+#include \"cbtree.h\"\n+struct repository;\n+struct packed_git;\n+struct object_id;\n+struct git_hash_algo;\n+struct pack_idx_entry;\n+\n+int load_pack_compat_map(struct repository *repo, struct packed_git *p);\n+\n+typedef enum cb_next (*compat_map_iter_t)(const struct object_id *, void *data);\n+void pack_compat_map_each(struct repository *repo, struct packed_git *p,\n+\t\t\t const unsigned char *prefix, size_t prefix_hexsz,\n+\t\t\t compat_map_iter_t, void *data);\n+\n+int repo_packed_oid_to_algop(struct repository *repo,\n+\t\t\t     const struct object_id *src,\n+\t\t\t     const struct git_hash_algo *to,\n+\t\t\t     struct object_id *dest);\n+\n+const char *write_compat_map_file(const char *compat_map_name,\n+\t\t\t\t  struct pack_idx_entry **objects,\n+\t\t\t\t  int nr_objects, const unsigned char *hash);\n+\n+#endif /* PACK_COMPAT_MAP_H */\ndiff --git a/pack-write.c b/pack-write.c\nindex b19ddf15b284..f22eea964f77 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -12,6 +12,7 @@\n #include \"pack-revindex.h\"\n #include \"path.h\"\n #include \"strbuf.h\"\n+#include \"object-file-convert.h\"\n \n void reset_pack_idx_option(struct pack_idx_option *opts)\n {\n@@ -345,6 +346,157 @@ static char *write_mtimes_file(struct packing_data *to_pack,\n \treturn mtimes_name;\n }\n \n+struct map_entry {\n+\tconst struct pack_idx_entry *idx;\n+\tuint32_t oid_index;\n+\tuint32_t compat_oid_index;\n+};\n+\n+static int map_oid_cmp(const void *_a, const void *_b)\n+{\n+\tstruct map_entry *a = *(struct map_entry **)_a;\n+\tstruct map_entry *b = *(struct map_entry **)_b;\n+\treturn oidcmp(&a->idx->oid, &b->idx->oid);\n+}\n+\n+static int map_compat_oid_cmp(const void *_a, const void *_b)\n+{\n+\tstruct map_entry *a = *(struct map_entry **)_a;\n+\tstruct map_entry *b = *(struct map_entry **)_b;\n+\treturn oidcmp(&a->idx->compat_oid, &b->idx->compat_oid);\n+}\n+\n+struct pack_compat_map_header {\n+\tuint8_t sig[4];\n+\tuint8_t version;\n+\tuint8_t first_oid_version;\n+\tuint8_t second_oid_version;\n+\tuint8_t mbz1;\n+\tuint32_t nr_objects;\n+\tuint8_t first_abbrev_len;\n+\tuint8_t mbz2;\n+\tuint8_t second_abbrev_len;\n+\tuint8_t mbz3;\n+};\n+\n+static inline unsigned last_matching_offset(const struct object_id *a,\n+\t\t\t\t\t    const struct object_id *b,\n+\t\t\t\t\t    const struct git_hash_algo *algop)\n+{\n+\tunsigned i;\n+\tfor (i = 0; i < algop->rawsz; i++)\n+\t\tif (a->hash[i] != b->hash[i])\n+\t\t\treturn i;\n+\t/* We should never hit this case. */\n+\treturn i;\n+}\n+\n+/*\n+ * The *hash contains the pack content hash.\n+ * The objects array is passed in sorted.\n+ */\n+const char *write_compat_map_file(const char *compat_map_name,\n+\t\t\t\t  struct pack_idx_entry **objects,\n+\t\t\t\t  int nr_objects, const unsigned char *hash)\n+{\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tunsigned short_name_len, compat_short_name_len;\n+\tstruct hashfile *f;\n+\tstruct map_entry *map_entries, **map;\n+\tstruct pack_compat_map_header hdr;\n+\tunsigned i;\n+\tint fd;\n+\n+\tif (!compat || !nr_objects)\n+\t\treturn NULL;\n+\n+\tALLOC_ARRAY(map_entries, nr_objects);\n+\tALLOC_ARRAY(map, nr_objects);\n+\tshort_name_len = 1;\n+\tfor (i = 0; i < nr_objects; ++i) {\n+\t\tunsigned offset;\n+\n+\t\tmap[i] = &map_entries[i];\n+\t\tmap_entries[i].idx = objects[i];\n+\t\tif (!objects[i]->compat_oid.algo)\n+\t\t\tBUG(\"No mapping from %s to %s\\n\",\n+\t\t\t    oid_to_hex(&objects[i]->oid),\n+\t\t\t    compat->name);\n+\n+\t\tmap_entries[i].oid_index = i;\n+\t\tmap_entries[i].compat_oid_index = 0;\n+\t\tif (i == 0)\n+\t\t\tcontinue;\n+\n+\t\toffset = last_matching_offset(&map_entries[i].idx->oid,\n+\t\t\t\t\t      &map_entries[i - 1].idx->oid,\n+\t\t\t\t\t      algo);\n+\t\tif (offset > short_name_len)\n+\t\t\tshort_name_len = offset;\n+\t}\n+\tQSORT(map, nr_objects, map_compat_oid_cmp);\n+\tcompat_short_name_len = 1;\n+\tfor (i = 0; i < nr_objects; ++i) {\n+\t\tunsigned offset;\n+\n+\t\tmap[i]->compat_oid_index = i;\n+\n+\t\tif (i == 0)\n+\t\t\tcontinue;\n+\n+\t\toffset = last_matching_offset(&map[i]->idx->compat_oid,\n+\t\t\t\t\t      &map[i - 1]->idx->compat_oid,\n+\t\t\t\t\t      compat);\n+\t\tif (offset > compat_short_name_len)\n+\t\t\tcompat_short_name_len = offset;\n+\t}\n+\n+\tif (compat_map_name) {\n+\t\t/* Verify an existing compat map file */\n+\t\tf = hashfd_check(compat_map_name);\n+\t} else {\n+\t\tstruct strbuf tmp_file = STRBUF_INIT;\n+\t\tfd = odb_mkstemp(&tmp_file, \"pack/tmp_compat_map_XXXXXX\");\n+\t\tcompat_map_name = strbuf_detach(&tmp_file, NULL);\n+\t\tf = hashfd(fd, compat_map_name);\n+\t}\n+\n+\thdr.sig[0] = 'C';\n+\thdr.sig[1] = 'M';\n+\thdr.sig[2] = 'A';\n+\thdr.sig[3] = 'P';\n+\thdr.version = 1;\n+\thdr.first_oid_version = oid_version(algo);\n+\thdr.second_oid_version = oid_version(compat);\n+\thdr.mbz1 = 0;\n+\thdr.nr_objects = htonl(nr_objects);\n+\thdr.first_abbrev_len = short_name_len;\n+\thdr.mbz2 = 0;\n+\thdr.second_abbrev_len = compat_short_name_len;\n+\thdr.mbz3 = 0;\n+\thashwrite(f, &hdr, sizeof(hdr));\n+\n+\tQSORT(map, nr_objects, map_oid_cmp);\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\thashwrite(f, map[i]->idx->oid.hash, algo->rawsz);\n+\t\thashwrite_be32(f, map[i]->compat_oid_index);\n+\t}\n+\tQSORT(map, nr_objects, map_compat_oid_cmp);\n+\tfor (i = 0; i < nr_objects; i++) {\n+\t\thashwrite(f, map[i]->idx->compat_oid.hash, compat->rawsz);\n+\t\thashwrite_be32(f, map[i]->oid_index);\n+\t}\n+\n+\thashwrite(f, hash, algo->rawsz);\n+\tfinalize_hashfile(f, NULL, FSYNC_COMPONENT_PACK_METADATA,\n+\t\t\t  CSUM_HASH_IN_STREAM | CSUM_CLOSE | CSUM_FSYNC);\n+\tfree(map);\n+\tfree(map_entries);\n+\treturn compat_map_name;\n+}\n+\n off_t write_pack_header(struct hashfile *f, uint32_t nr_entries)\n {\n \tstruct pack_header hdr;\n@@ -548,6 +700,7 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n {\n \tconst char *rev_tmp_name = NULL;\n \tchar *mtimes_tmp_name = NULL;\n+\tconst char *compat_map_tmp_name = NULL;\n \n \tif (adjust_shared_perm(pack_tmp_name))\n \t\tdie_errno(\"unable to make temporary pack file readable\");\n@@ -566,11 +719,16 @@ void stage_tmp_packfiles(struct strbuf *name_buffer,\n \t\t\t\t\t\t    hash);\n \t}\n \n+\tcompat_map_tmp_name = write_compat_map_file(NULL, written_list,\n+\t\t\t\t\t\t    nr_written, hash);\n+\n \trename_tmp_packfile(name_buffer, pack_tmp_name, \"pack\");\n \tif (rev_tmp_name)\n \t\trename_tmp_packfile(name_buffer, rev_tmp_name, \"rev\");\n \tif (mtimes_tmp_name)\n \t\trename_tmp_packfile(name_buffer, mtimes_tmp_name, \"mtimes\");\n+\tif (compat_map_tmp_name)\n+\t\trename_tmp_packfile(name_buffer, compat_map_tmp_name, \"compat\");\n \n \tfree((char *)rev_tmp_name);\n \tfree(mtimes_tmp_name);\ndiff --git a/packfile.c b/packfile.c\nindex 1fae0fcdd9e7..c1a6bd9bc6b3 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -349,6 +349,17 @@ static void close_pack_mtimes(struct packed_git *p)\n \tp->mtimes_map = NULL;\n }\n \n+static void close_pack_compat_map(struct packed_git *p)\n+{\n+\tif (!p->compat_mapping)\n+\t\treturn;\n+\n+\tmunmap((void *)p->compat_mapping, p->compat_mapping_size);\n+\tp->compat_mapping = NULL;\n+\tp->hash_map = NULL;\n+\tp->compat_hash_map = NULL;\n+}\n+\n void close_pack(struct packed_git *p)\n {\n \tclose_pack_windows(p);\n@@ -356,6 +367,7 @@ void close_pack(struct packed_git *p)\n \tclose_pack_index(p);\n \tclose_pack_revindex(p);\n \tclose_pack_mtimes(p);\n+\tclose_pack_compat_map(p);\n \toidset_clear(&p->bad_objects);\n }\n \n-- \n2.41.0\n\n"},{"id":"481607","messageId":"20230908231049.2035003-21-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 21/32] tree-walk: init_tree_desc take an oid to get the hash algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:38Z","receivedAt":"2023-09-08T23:31:37Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"To make it possible for git ls-tree to display the tree encoded\nin the hash algorithm of the oid specified to git ls-tree, update\ninit_tree_desc to take as a parameter the oid of the tree object.\n\nUpdate all callers of init_tree_desc and init_tree_desc_gently\nto pass the oid of the tree object.\n\nUse the oid of the tree object to discover the hash algorithm\nof the oid and store that hash algorithm in struct tree_desc.\n\nUse the hash algorithm in decode_tree_entry and\nupdate_tree_entry_internal to handle reading a tree object encoded in\na hash algorithm that differs from the repositories hash algorithm.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n archive.c              |  3 ++-\n builtin/am.c           |  6 +++---\n builtin/checkout.c     |  8 +++++---\n builtin/clone.c        |  2 +-\n builtin/commit.c       |  2 +-\n builtin/grep.c         |  8 ++++----\n builtin/merge.c        |  3 ++-\n builtin/pack-objects.c |  6 ++++--\n builtin/read-tree.c    |  2 +-\n builtin/stash.c        |  5 +++--\n cache-tree.c           |  2 +-\n delta-islands.c        |  2 +-\n diff-lib.c             |  2 +-\n fsck.c                 |  6 ++++--\n http-push.c            |  2 +-\n list-objects.c         |  2 +-\n match-trees.c          |  4 ++--\n merge-ort.c            | 11 ++++++-----\n merge-recursive.c      |  2 +-\n merge.c                |  3 ++-\n pack-bitmap-write.c    |  2 +-\n packfile.c             |  3 ++-\n reflog.c               |  2 +-\n revision.c             |  4 ++--\n tree-walk.c            | 36 +++++++++++++++++++++---------------\n tree-walk.h            |  7 +++++--\n tree.c                 |  2 +-\n walker.c               |  2 +-\n 28 files changed, 80 insertions(+), 59 deletions(-)\n\ndiff --git a/archive.c b/archive.c\nindex ca11db185b15..b10269aee7be 100644\n--- a/archive.c\n+++ b/archive.c\n@@ -339,7 +339,8 @@ int write_archive_entries(struct archiver_args *args,\n \t\topts.src_index = args->repo->index;\n \t\topts.dst_index = args->repo->index;\n \t\topts.fn = oneway_merge;\n-\t\tinit_tree_desc(&t, args->tree->buffer, args->tree->size);\n+\t\tinit_tree_desc(&t, &args->tree->object.oid,\n+\t\t\t       args->tree->buffer, args->tree->size);\n \t\tif (unpack_trees(1, &t, &opts))\n \t\t\treturn -1;\n \t\tgit_attr_set_direction(GIT_ATTR_INDEX);\ndiff --git a/builtin/am.c b/builtin/am.c\nindex 8bde034fae68..4dfd714b910e 100644\n--- a/builtin/am.c\n+++ b/builtin/am.c\n@@ -1991,8 +1991,8 @@ static int fast_forward_to(struct tree *head, struct tree *remote, int reset)\n \topts.reset = reset ? UNPACK_RESET_PROTECT_UNTRACKED : 0;\n \topts.preserve_ignored = 0; /* FIXME: !overwrite_ignore */\n \topts.fn = twoway_merge;\n-\tinit_tree_desc(&t[0], head->buffer, head->size);\n-\tinit_tree_desc(&t[1], remote->buffer, remote->size);\n+\tinit_tree_desc(&t[0], &head->object.oid, head->buffer, head->size);\n+\tinit_tree_desc(&t[1], &remote->object.oid, remote->buffer, remote->size);\n \n \tif (unpack_trees(2, t, &opts)) {\n \t\trollback_lock_file(&lock_file);\n@@ -2026,7 +2026,7 @@ static int merge_tree(struct tree *tree)\n \topts.dst_index = &the_index;\n \topts.merge = 1;\n \topts.fn = oneway_merge;\n-\tinit_tree_desc(&t[0], tree->buffer, tree->size);\n+\tinit_tree_desc(&t[0], &tree->object.oid, tree->buffer, tree->size);\n \n \tif (unpack_trees(1, t, &opts)) {\n \t\trollback_lock_file(&lock_file);\ndiff --git a/builtin/checkout.c b/builtin/checkout.c\nindex f53612f46870..03eff73fd031 100644\n--- a/builtin/checkout.c\n+++ b/builtin/checkout.c\n@@ -701,7 +701,7 @@ static int reset_tree(struct tree *tree, const struct checkout_opts *o,\n \t\t\t       info->commit ? &info->commit->object.oid : null_oid(),\n \t\t\t       NULL);\n \tparse_tree(tree);\n-\tinit_tree_desc(&tree_desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&tree_desc, &tree->object.oid, tree->buffer, tree->size);\n \tswitch (unpack_trees(1, &tree_desc, &opts)) {\n \tcase -2:\n \t\t*writeout_error = 1;\n@@ -815,10 +815,12 @@ static int merge_working_tree(const struct checkout_opts *opts,\n \t\t\tdie(_(\"unable to parse commit %s\"),\n \t\t\t\toid_to_hex(old_commit_oid));\n \n-\t\tinit_tree_desc(&trees[0], tree->buffer, tree->size);\n+\t\tinit_tree_desc(&trees[0], &tree->object.oid,\n+\t\t\t       tree->buffer, tree->size);\n \t\tparse_tree(new_tree);\n \t\ttree = new_tree;\n-\t\tinit_tree_desc(&trees[1], tree->buffer, tree->size);\n+\t\tinit_tree_desc(&trees[1], &tree->object.oid,\n+\t\t\t       tree->buffer, tree->size);\n \n \t\tret = unpack_trees(2, trees, &topts);\n \t\tclear_unpack_trees_porcelain(&topts);\ndiff --git a/builtin/clone.c b/builtin/clone.c\nindex c6357af94989..79ceefb93995 100644\n--- a/builtin/clone.c\n+++ b/builtin/clone.c\n@@ -737,7 +737,7 @@ static int checkout(int submodule_progress, int filter_submodules)\n \tif (!tree)\n \t\tdie(_(\"unable to parse commit %s\"), oid_to_hex(&oid));\n \tparse_tree(tree);\n-\tinit_tree_desc(&t, tree->buffer, tree->size);\n+\tinit_tree_desc(&t, &tree->object.oid, tree->buffer, tree->size);\n \tif (unpack_trees(1, &t, &opts) < 0)\n \t\tdie(_(\"unable to checkout working tree\"));\n \ndiff --git a/builtin/commit.c b/builtin/commit.c\nindex 7da5f924484d..537319932b65 100644\n--- a/builtin/commit.c\n+++ b/builtin/commit.c\n@@ -340,7 +340,7 @@ static void create_base_index(const struct commit *current_head)\n \tif (!tree)\n \t\tdie(_(\"failed to unpack HEAD tree object\"));\n \tparse_tree(tree);\n-\tinit_tree_desc(&t, tree->buffer, tree->size);\n+\tinit_tree_desc(&t, &tree->object.oid, tree->buffer, tree->size);\n \tif (unpack_trees(1, &t, &opts))\n \t\texit(128); /* We've already reported the error, finish dying */\n }\ndiff --git a/builtin/grep.c b/builtin/grep.c\nindex 50e712a18479..0c2b8a376f8e 100644\n--- a/builtin/grep.c\n+++ b/builtin/grep.c\n@@ -530,7 +530,7 @@ static int grep_submodule(struct grep_opt *opt,\n \t\tstrbuf_addstr(&base, filename);\n \t\tstrbuf_addch(&base, '/');\n \n-\t\tinit_tree_desc(&tree, data, size);\n+\t\tinit_tree_desc(&tree, oid, data, size);\n \t\thit = grep_tree(&subopt, pathspec, &tree, &base, base.len,\n \t\t\t\tobject_type == OBJ_COMMIT);\n \t\tstrbuf_release(&base);\n@@ -574,7 +574,7 @@ static int grep_cache(struct grep_opt *opt,\n \n \t\t\tdata = repo_read_object_file(the_repository, &ce->oid,\n \t\t\t\t\t\t     &type, &size);\n-\t\t\tinit_tree_desc(&tree, data, size);\n+\t\t\tinit_tree_desc(&tree, &ce->oid, data, size);\n \n \t\t\thit |= grep_tree(opt, pathspec, &tree, &name, 0, 0);\n \t\t\tstrbuf_setlen(&name, name_base_len);\n@@ -670,7 +670,7 @@ static int grep_tree(struct grep_opt *opt, const struct pathspec *pathspec,\n \t\t\t\t    oid_to_hex(&entry.oid));\n \n \t\t\tstrbuf_addch(base, '/');\n-\t\t\tinit_tree_desc(&sub, data, size);\n+\t\t\tinit_tree_desc(&sub, &entry.oid, data, size);\n \t\t\thit |= grep_tree(opt, pathspec, &sub, base, tn_len,\n \t\t\t\t\t check_attr);\n \t\t\tfree(data);\n@@ -714,7 +714,7 @@ static int grep_object(struct grep_opt *opt, const struct pathspec *pathspec,\n \t\t\tstrbuf_add(&base, name, len);\n \t\t\tstrbuf_addch(&base, ':');\n \t\t}\n-\t\tinit_tree_desc(&tree, data, size);\n+\t\tinit_tree_desc(&tree, &obj->oid, data, size);\n \t\thit = grep_tree(opt, pathspec, &tree, &base, base.len,\n \t\t\t\tobj->type == OBJ_COMMIT);\n \t\tstrbuf_release(&base);\ndiff --git a/builtin/merge.c b/builtin/merge.c\nindex de68910177fb..718165d45917 100644\n--- a/builtin/merge.c\n+++ b/builtin/merge.c\n@@ -704,7 +704,8 @@ static int read_tree_trivial(struct object_id *common, struct object_id *head,\n \tcache_tree_free(&the_index.cache_tree);\n \tfor (i = 0; i < nr_trees; i++) {\n \t\tparse_tree(trees[i]);\n-\t\tinit_tree_desc(t+i, trees[i]->buffer, trees[i]->size);\n+\t\tinit_tree_desc(t+i, &trees[i]->object.oid,\n+\t\t\t       trees[i]->buffer, trees[i]->size);\n \t}\n \tif (unpack_trees(nr_trees, t, &opts))\n \t\treturn -1;\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex d2a162d52804..d34902002656 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -1756,7 +1756,8 @@ static void add_pbase_object(struct tree_desc *tree,\n \t\t\ttree = pbase_tree_get(&entry.oid);\n \t\t\tif (!tree)\n \t\t\t\treturn;\n-\t\t\tinit_tree_desc(&sub, tree->tree_data, tree->tree_size);\n+\t\t\tinit_tree_desc(&sub, &tree->oid,\n+\t\t\t\t       tree->tree_data, tree->tree_size);\n \n \t\t\tadd_pbase_object(&sub, down, downlen, fullname);\n \t\t\tpbase_tree_put(tree);\n@@ -1816,7 +1817,8 @@ static void add_preferred_base_object(const char *name)\n \t\t}\n \t\telse {\n \t\t\tstruct tree_desc tree;\n-\t\t\tinit_tree_desc(&tree, it->pcache.tree_data, it->pcache.tree_size);\n+\t\t\tinit_tree_desc(&tree, &it->pcache.oid,\n+\t\t\t\t       it->pcache.tree_data, it->pcache.tree_size);\n \t\t\tadd_pbase_object(&tree, name, cmplen, name);\n \t\t}\n \t}\ndiff --git a/builtin/read-tree.c b/builtin/read-tree.c\nindex 1fec702a04fa..24d6d156d3a2 100644\n--- a/builtin/read-tree.c\n+++ b/builtin/read-tree.c\n@@ -264,7 +264,7 @@ int cmd_read_tree(int argc, const char **argv, const char *cmd_prefix)\n \tfor (i = 0; i < nr_trees; i++) {\n \t\tstruct tree *tree = trees[i];\n \t\tparse_tree(tree);\n-\t\tinit_tree_desc(t+i, tree->buffer, tree->size);\n+\t\tinit_tree_desc(t+i, &tree->object.oid, tree->buffer, tree->size);\n \t}\n \tif (unpack_trees(nr_trees, t, &opts))\n \t\treturn 128;\ndiff --git a/builtin/stash.c b/builtin/stash.c\nindex fe64cde9ce30..9ee52af4d28e 100644\n--- a/builtin/stash.c\n+++ b/builtin/stash.c\n@@ -285,7 +285,7 @@ static int reset_tree(struct object_id *i_tree, int update, int reset)\n \tif (parse_tree(tree))\n \t\treturn -1;\n \n-\tinit_tree_desc(t, tree->buffer, tree->size);\n+\tinit_tree_desc(t, &tree->object.oid, tree->buffer, tree->size);\n \n \topts.head_idx = 1;\n \topts.src_index = &the_index;\n@@ -871,7 +871,8 @@ static void diff_include_untracked(const struct stash_info *info, struct diff_op\n \t\ttree[i] = parse_tree_indirect(oid[i]);\n \t\tif (parse_tree(tree[i]) < 0)\n \t\t\tdie(_(\"failed to parse tree\"));\n-\t\tinit_tree_desc(&tree_desc[i], tree[i]->buffer, tree[i]->size);\n+\t\tinit_tree_desc(&tree_desc[i], &tree[i]->object.oid,\n+\t\t\t       tree[i]->buffer, tree[i]->size);\n \t}\n \n \tunpack_tree_opt.head_idx = -1;\ndiff --git a/cache-tree.c b/cache-tree.c\nindex ddc7d3d86959..334973a01cee 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -770,7 +770,7 @@ static void prime_cache_tree_rec(struct repository *r,\n \n \toidcpy(&it->oid, &tree->object.oid);\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \tcnt = 0;\n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (!S_ISDIR(entry.mode))\ndiff --git a/delta-islands.c b/delta-islands.c\nindex 5de5759f3f13..1ff3506b10f2 100644\n--- a/delta-islands.c\n+++ b/delta-islands.c\n@@ -289,7 +289,7 @@ void resolve_tree_islands(struct repository *r,\n \t\tif (!tree || parse_tree(tree) < 0)\n \t\t\tdie(_(\"bad tree object %s\"), oid_to_hex(&ent->idx.oid));\n \n-\t\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\t\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \t\twhile (tree_entry(&desc, &entry)) {\n \t\t\tstruct object *obj;\n \ndiff --git a/diff-lib.c b/diff-lib.c\nindex 6b0c6a7180cc..add323f5628d 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -558,7 +558,7 @@ static int diff_cache(struct rev_info *revs,\n \topts.pathspec = &revs->diffopt.pathspec;\n \topts.pathspec->recursive = 1;\n \n-\tinit_tree_desc(&t, tree->buffer, tree->size);\n+\tinit_tree_desc(&t, &tree->object.oid, tree->buffer, tree->size);\n \treturn unpack_trees(1, &t, &opts);\n }\n \ndiff --git a/fsck.c b/fsck.c\nindex 2b1e348005b7..6b492a48da82 100644\n--- a/fsck.c\n+++ b/fsck.c\n@@ -313,7 +313,8 @@ static int fsck_walk_tree(struct tree *tree, void *data, struct fsck_options *op\n \t\treturn -1;\n \n \tname = fsck_get_object_name(options, &tree->object.oid);\n-\tif (init_tree_desc_gently(&desc, tree->buffer, tree->size, 0))\n+\tif (init_tree_desc_gently(&desc, &tree->object.oid,\n+\t\t\t\t  tree->buffer, tree->size, 0))\n \t\treturn -1;\n \twhile (tree_entry_gently(&desc, &entry)) {\n \t\tstruct object *obj;\n@@ -583,7 +584,8 @@ static int fsck_tree(const struct object_id *tree_oid,\n \tconst char *o_name;\n \tstruct name_stack df_dup_candidates = { NULL };\n \n-\tif (init_tree_desc_gently(&desc, buffer, size, TREE_DESC_RAW_MODES)) {\n+\tif (init_tree_desc_gently(&desc, tree_oid, buffer, size,\n+\t\t\t\t  TREE_DESC_RAW_MODES)) {\n \t\tretval += report(options, tree_oid, OBJ_TREE,\n \t\t\t\t FSCK_MSG_BAD_TREE,\n \t\t\t\t \"cannot be parsed as a tree\");\ndiff --git a/http-push.c b/http-push.c\nindex a704f490fdb2..81c35b5e96f7 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -1308,7 +1308,7 @@ static struct object_list **process_tree(struct tree *tree,\n \tobj->flags |= SEEN;\n \tp = add_one_object(obj, p);\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry))\n \t\tswitch (object_type(entry.mode)) {\ndiff --git a/list-objects.c b/list-objects.c\nindex e60a6cd5b46e..312335c8a7f2 100644\n--- a/list-objects.c\n+++ b/list-objects.c\n@@ -97,7 +97,7 @@ static void process_tree_contents(struct traversal_context *ctx,\n \tenum interesting match = ctx->revs->diffopt.pathspec.nr == 0 ?\n \t\tall_entries_interesting : entry_not_interesting;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (match != all_entries_interesting) {\ndiff --git a/match-trees.c b/match-trees.c\nindex 0885ac681cd5..3412b6a1401d 100644\n--- a/match-trees.c\n+++ b/match-trees.c\n@@ -63,7 +63,7 @@ static void *fill_tree_desc_strict(struct tree_desc *desc,\n \t\tdie(\"unable to read tree (%s)\", oid_to_hex(hash));\n \tif (type != OBJ_TREE)\n \t\tdie(\"%s is not a tree\", oid_to_hex(hash));\n-\tinit_tree_desc(desc, buffer, size);\n+\tinit_tree_desc(desc, hash, buffer, size);\n \treturn buffer;\n }\n \n@@ -194,7 +194,7 @@ static int splice_tree(const struct object_id *oid1, const char *prefix,\n \tbuf = repo_read_object_file(the_repository, oid1, &type, &sz);\n \tif (!buf)\n \t\tdie(\"cannot read tree %s\", oid_to_hex(oid1));\n-\tinit_tree_desc(&desc, buf, sz);\n+\tinit_tree_desc(&desc, oid1, buf, sz);\n \n \trewrite_here = NULL;\n \twhile (desc.size) {\ndiff --git a/merge-ort.c b/merge-ort.c\nindex 8631c997002d..3a5729c91e48 100644\n--- a/merge-ort.c\n+++ b/merge-ort.c\n@@ -1679,9 +1679,10 @@ static int collect_merge_info(struct merge_options *opt,\n \tparse_tree(merge_base);\n \tparse_tree(side1);\n \tparse_tree(side2);\n-\tinit_tree_desc(t + 0, merge_base->buffer, merge_base->size);\n-\tinit_tree_desc(t + 1, side1->buffer, side1->size);\n-\tinit_tree_desc(t + 2, side2->buffer, side2->size);\n+\tinit_tree_desc(t + 0, &merge_base->object.oid,\n+\t\t       merge_base->buffer, merge_base->size);\n+\tinit_tree_desc(t + 1, &side1->object.oid, side1->buffer, side1->size);\n+\tinit_tree_desc(t + 2, &side2->object.oid, side2->buffer, side2->size);\n \n \ttrace2_region_enter(\"merge\", \"traverse_trees\", opt->repo);\n \tret = traverse_trees(NULL, 3, t, &info);\n@@ -4400,9 +4401,9 @@ static int checkout(struct merge_options *opt,\n \tunpack_opts.fn = twoway_merge;\n \tunpack_opts.preserve_ignored = 0; /* FIXME: !opts->overwrite_ignore */\n \tparse_tree(prev);\n-\tinit_tree_desc(&trees[0], prev->buffer, prev->size);\n+\tinit_tree_desc(&trees[0], &prev->object.oid, prev->buffer, prev->size);\n \tparse_tree(next);\n-\tinit_tree_desc(&trees[1], next->buffer, next->size);\n+\tinit_tree_desc(&trees[1], &next->object.oid, next->buffer, next->size);\n \n \tret = unpack_trees(2, trees, &unpack_opts);\n \tclear_unpack_trees_porcelain(&unpack_opts);\ndiff --git a/merge-recursive.c b/merge-recursive.c\nindex 6a4081bb0f52..93df9eecdd95 100644\n--- a/merge-recursive.c\n+++ b/merge-recursive.c\n@@ -411,7 +411,7 @@ static inline int merge_detect_rename(struct merge_options *opt)\n static void init_tree_desc_from_tree(struct tree_desc *desc, struct tree *tree)\n {\n \tparse_tree(tree);\n-\tinit_tree_desc(desc, tree->buffer, tree->size);\n+\tinit_tree_desc(desc, &tree->object.oid, tree->buffer, tree->size);\n }\n \n static int unpack_trees_start(struct merge_options *opt,\ndiff --git a/merge.c b/merge.c\nindex b60925459c29..86179c34102d 100644\n--- a/merge.c\n+++ b/merge.c\n@@ -81,7 +81,8 @@ int checkout_fast_forward(struct repository *r,\n \t}\n \tfor (i = 0; i < nr_trees; i++) {\n \t\tparse_tree(trees[i]);\n-\t\tinit_tree_desc(t+i, trees[i]->buffer, trees[i]->size);\n+\t\tinit_tree_desc(t+i, &trees[i]->object.oid,\n+\t\t\t       trees[i]->buffer, trees[i]->size);\n \t}\n \n \tmemset(&opts, 0, sizeof(opts));\ndiff --git a/pack-bitmap-write.c b/pack-bitmap-write.c\nindex f6757c3cbf20..9211e08f0127 100644\n--- a/pack-bitmap-write.c\n+++ b/pack-bitmap-write.c\n@@ -366,7 +366,7 @@ static int fill_bitmap_tree(struct bitmap *bitmap,\n \tif (parse_tree(tree) < 0)\n \t\tdie(\"unable to load tree object %s\",\n \t\t    oid_to_hex(&tree->object.oid));\n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry)) {\n \t\tswitch (object_type(entry.mode)) {\ndiff --git a/packfile.c b/packfile.c\nindex 9cc0a2e37a83..1fae0fcdd9e7 100644\n--- a/packfile.c\n+++ b/packfile.c\n@@ -2250,7 +2250,8 @@ static int add_promisor_object(const struct object_id *oid,\n \t\tstruct tree *tree = (struct tree *)obj;\n \t\tstruct tree_desc desc;\n \t\tstruct name_entry entry;\n-\t\tif (init_tree_desc_gently(&desc, tree->buffer, tree->size, 0))\n+\t\tif (init_tree_desc_gently(&desc, &tree->object.oid,\n+\t\t\t\t\t  tree->buffer, tree->size, 0))\n \t\t\t/*\n \t\t\t * Error messages are given when packs are\n \t\t\t * verified, so do not print any here.\ndiff --git a/reflog.c b/reflog.c\nindex 9ad50e7d93e4..c6992a19268f 100644\n--- a/reflog.c\n+++ b/reflog.c\n@@ -40,7 +40,7 @@ static int tree_is_complete(const struct object_id *oid)\n \t\ttree->buffer = data;\n \t\ttree->size = size;\n \t}\n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \tcomplete = 1;\n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (!repo_has_object_file(the_repository, &entry.oid) ||\ndiff --git a/revision.c b/revision.c\nindex 2f4c53ea207b..a60dfc23a2a5 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -82,7 +82,7 @@ static void mark_tree_contents_uninteresting(struct repository *r,\n \tif (parse_tree_gently(tree, 1) < 0)\n \t\treturn;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \twhile (tree_entry(&desc, &entry)) {\n \t\tswitch (object_type(entry.mode)) {\n \t\tcase OBJ_TREE:\n@@ -189,7 +189,7 @@ static void add_children_by_path(struct repository *r,\n \tif (parse_tree_gently(tree, 1) < 0)\n \t\treturn;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \twhile (tree_entry(&desc, &entry)) {\n \t\tswitch (object_type(entry.mode)) {\n \t\tcase OBJ_TREE:\ndiff --git a/tree-walk.c b/tree-walk.c\nindex 3af50a01c2c7..0b44ec7c75ff 100644\n--- a/tree-walk.c\n+++ b/tree-walk.c\n@@ -15,7 +15,7 @@ static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned l\n \tconst char *path;\n \tunsigned int len;\n \tuint16_t mode;\n-\tconst unsigned hashsz = the_hash_algo->rawsz;\n+\tconst unsigned hashsz = desc->algo->rawsz;\n \n \tif (size < hashsz + 3 || buf[size - (hashsz + 1)]) {\n \t\tstrbuf_addstr(err, _(\"too-short tree object\"));\n@@ -37,15 +37,19 @@ static int decode_tree_entry(struct tree_desc *desc, const char *buf, unsigned l\n \tdesc->entry.path = path;\n \tdesc->entry.mode = (desc->flags & TREE_DESC_RAW_MODES) ? mode : canon_mode(mode);\n \tdesc->entry.pathlen = len - 1;\n-\toidread(&desc->entry.oid, (const unsigned char *)path + len);\n+\toidread_algop(&desc->entry.oid, (const unsigned char *)path + len,\n+\t\t      desc->algo);\n \n \treturn 0;\n }\n \n-static int init_tree_desc_internal(struct tree_desc *desc, const void *buffer,\n-\t\t\t\t   unsigned long size, struct strbuf *err,\n+static int init_tree_desc_internal(struct tree_desc *desc,\n+\t\t\t\t   const struct object_id *oid,\n+\t\t\t\t   const void *buffer, unsigned long size,\n+\t\t\t\t   struct strbuf *err,\n \t\t\t\t   enum tree_desc_flags flags)\n {\n+\tdesc->algo = (oid && oid->algo) ? &hash_algos[oid->algo] : the_hash_algo;\n \tdesc->buffer = buffer;\n \tdesc->size = size;\n \tdesc->flags = flags;\n@@ -54,19 +58,21 @@ static int init_tree_desc_internal(struct tree_desc *desc, const void *buffer,\n \treturn 0;\n }\n \n-void init_tree_desc(struct tree_desc *desc, const void *buffer, unsigned long size)\n+void init_tree_desc(struct tree_desc *desc, const struct object_id *tree_oid,\n+\t\t    const void *buffer, unsigned long size)\n {\n \tstruct strbuf err = STRBUF_INIT;\n-\tif (init_tree_desc_internal(desc, buffer, size, &err, 0))\n+\tif (init_tree_desc_internal(desc, tree_oid, buffer, size, &err, 0))\n \t\tdie(\"%s\", err.buf);\n \tstrbuf_release(&err);\n }\n \n-int init_tree_desc_gently(struct tree_desc *desc, const void *buffer, unsigned long size,\n+int init_tree_desc_gently(struct tree_desc *desc, const struct object_id *oid,\n+\t\t\t  const void *buffer, unsigned long size,\n \t\t\t  enum tree_desc_flags flags)\n {\n \tstruct strbuf err = STRBUF_INIT;\n-\tint result = init_tree_desc_internal(desc, buffer, size, &err, flags);\n+\tint result = init_tree_desc_internal(desc, oid, buffer, size, &err, flags);\n \tif (result)\n \t\terror(\"%s\", err.buf);\n \tstrbuf_release(&err);\n@@ -85,7 +91,7 @@ void *fill_tree_descriptor(struct repository *r,\n \t\tif (!buf)\n \t\t\tdie(\"unable to read tree %s\", oid_to_hex(oid));\n \t}\n-\tinit_tree_desc(desc, buf, size);\n+\tinit_tree_desc(desc, oid, buf, size);\n \treturn buf;\n }\n \n@@ -102,7 +108,7 @@ static void entry_extract(struct tree_desc *t, struct name_entry *a)\n static int update_tree_entry_internal(struct tree_desc *desc, struct strbuf *err)\n {\n \tconst void *buf = desc->buffer;\n-\tconst unsigned char *end = (const unsigned char *)desc->entry.path + desc->entry.pathlen + 1 + the_hash_algo->rawsz;\n+\tconst unsigned char *end = (const unsigned char *)desc->entry.path + desc->entry.pathlen + 1 + desc->algo->rawsz;\n \tunsigned long size = desc->size;\n \tunsigned long len = end - (const unsigned char *)buf;\n \n@@ -611,7 +617,7 @@ int get_tree_entry(struct repository *r,\n \t\tretval = -1;\n \t} else {\n \t\tstruct tree_desc t;\n-\t\tinit_tree_desc(&t, tree, size);\n+\t\tinit_tree_desc(&t, tree_oid, tree, size);\n \t\tretval = find_tree_entry(r, &t, name, oid, mode);\n \t}\n \tfree(tree);\n@@ -654,7 +660,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \tstruct tree_desc t;\n \tint follows_remaining = GET_TREE_ENTRY_FOLLOW_SYMLINKS_MAX_LINKS;\n \n-\tinit_tree_desc(&t, NULL, 0UL);\n+\tinit_tree_desc(&t, NULL, NULL, 0UL);\n \tstrbuf_addstr(&namebuf, name);\n \toidcpy(&current_tree_oid, tree_oid);\n \n@@ -690,7 +696,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \t\t\t\tgoto done;\n \n \t\t\t/* descend */\n-\t\t\tinit_tree_desc(&t, tree, size);\n+\t\t\tinit_tree_desc(&t, &current_tree_oid, tree, size);\n \t\t}\n \n \t\t/* Handle symlinks to e.g. a//b by removing leading slashes */\n@@ -724,7 +730,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \t\t\tfree(parent->tree);\n \t\t\tparents_nr--;\n \t\t\tparent = &parents[parents_nr - 1];\n-\t\t\tinit_tree_desc(&t, parent->tree, parent->size);\n+\t\t\tinit_tree_desc(&t, &parent->oid, parent->tree, parent->size);\n \t\t\tstrbuf_remove(&namebuf, 0, remainder ? 3 : 2);\n \t\t\tcontinue;\n \t\t}\n@@ -804,7 +810,7 @@ enum get_oid_result get_tree_entry_follow_symlinks(struct repository *r,\n \t\t\tcontents_start = contents;\n \n \t\t\tparent = &parents[parents_nr - 1];\n-\t\t\tinit_tree_desc(&t, parent->tree, parent->size);\n+\t\t\tinit_tree_desc(&t, &parent->oid, parent->tree, parent->size);\n \t\t\tstrbuf_splice(&namebuf, 0, len,\n \t\t\t\t      contents_start, link_len);\n \t\t\tif (remainder)\ndiff --git a/tree-walk.h b/tree-walk.h\nindex 74cdceb3fed2..cf54d01019e9 100644\n--- a/tree-walk.h\n+++ b/tree-walk.h\n@@ -26,6 +26,7 @@ struct name_entry {\n  * A semi-opaque data structure used to maintain the current state of the walk.\n  */\n struct tree_desc {\n+\tconst struct git_hash_algo *algo;\n \t/*\n \t * pointer into the memory representation of the tree. It always\n \t * points at the current entry being visited.\n@@ -85,9 +86,11 @@ int update_tree_entry_gently(struct tree_desc *);\n  * size parameters are assumed to be the same as the buffer and size\n  * members of `struct tree`.\n  */\n-void init_tree_desc(struct tree_desc *desc, const void *buf, unsigned long size);\n+void init_tree_desc(struct tree_desc *desc, const struct object_id *tree_oid,\n+\t\t    const void *buf, unsigned long size);\n \n-int init_tree_desc_gently(struct tree_desc *desc, const void *buf, unsigned long size,\n+int init_tree_desc_gently(struct tree_desc *desc, const struct object_id *oid,\n+\t\t\t  const void *buf, unsigned long size,\n \t\t\t  enum tree_desc_flags flags);\n \n /*\ndiff --git a/tree.c b/tree.c\nindex c745462f968e..44bcf728f10a 100644\n--- a/tree.c\n+++ b/tree.c\n@@ -27,7 +27,7 @@ int read_tree_at(struct repository *r,\n \tif (parse_tree(tree))\n \t\treturn -1;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \n \twhile (tree_entry(&desc, &entry)) {\n \t\tif (retval != all_entries_interesting) {\ndiff --git a/walker.c b/walker.c\nindex 65002a7220ad..c0fd632d921c 100644\n--- a/walker.c\n+++ b/walker.c\n@@ -45,7 +45,7 @@ static int process_tree(struct walker *walker, struct tree *tree)\n \tif (parse_tree(tree))\n \t\treturn -1;\n \n-\tinit_tree_desc(&desc, tree->buffer, tree->size);\n+\tinit_tree_desc(&desc, &tree->object.oid, tree->buffer, tree->size);\n \twhile (tree_entry(&desc, &entry)) {\n \t\tstruct object *obj = NULL;\n \n-- \n2.41.0\n\n"},{"id":"481608","messageId":"20230908231049.2035003-24-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 24/32] builtin/pack-objects: Communicate the compatibility hash through struct pack_idx_entry","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:41Z","receivedAt":"2023-09-08T23:31:44Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"When pack-objects is run all objects in the repository should already\nhave a compatibilty hash computed so it is just necessary to read\nthe existing mappings and store the value in struct pack_idx_entry.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n builtin/pack-objects.c | 7 +++++++\n 1 file changed, 7 insertions(+)\n\ndiff --git a/builtin/pack-objects.c b/builtin/pack-objects.c\nindex d34902002656..ff04660a18fd 100644\n--- a/builtin/pack-objects.c\n+++ b/builtin/pack-objects.c\n@@ -42,6 +42,7 @@\n #include \"promisor-remote.h\"\n #include \"pack-mtimes.h\"\n #include \"parse-options.h\"\n+#include \"object-file-convert.h\"\n \n /*\n  * Objects we are going to pack are collected in the `to_pack` structure.\n@@ -1547,10 +1548,16 @@ static struct object_entry *create_object_entry(const struct object_id *oid,\n \t\t\t\t\t\tstruct packed_git *found_pack,\n \t\t\t\t\t\toff_t found_offset)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tstruct object_entry *entry;\n \n \tentry = packlist_alloc(&to_pack, oid);\n \tentry->hash = hash;\n+\tif (compat &&\n+\t    repo_oid_to_algop(repo, &entry->idx.oid, compat,\n+\t\t\t      &entry->idx.compat_oid))\n+\t\tdie(_(\"can't map object %s while writing pack\"), oid_to_hex(oid));\n \toe_set_type(entry, type);\n \tif (exclude)\n \t\tentry->preferred_base = 1;\n-- \n2.41.0\n\n"},{"id":"481609","messageId":"20230908231049.2035003-18-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 18/32] object-file-convert: convert commit objects when writing","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:35Z","receivedAt":"2023-09-08T23:31:45Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nWhen writing a commit object in a repository with both SHA-1 and\nSHA-256, we'll need to convert our commit objects so that we can write\nthe hash values for both into the repository.  To do so, let's add a\nfunction to convert commit objects.\n\nRead the commit object and map the tree value and any of the parent\nvalues, and copy the rest of the commit through unmodified.  Note that\nwe don't need to modify the signature headers, because they are the same\nunder both algorithms.\n\n****\n- made static and moved to object-file-convert.c\n- Renamed the variable compat_oid to mapped_oid for clarity\n- Replaced repo_map_object with oid_to_algop\n-- EWB\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 44 +++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 44 insertions(+)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex f266c8c6cc95..9c715a9864d5 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -83,6 +83,48 @@ static int convert_tree_object(struct strbuf *out,\n \treturn 0;\n }\n \n+static int convert_commit_object(struct strbuf *out,\n+\t\t\t\t const struct git_hash_algo *from,\n+\t\t\t\t const struct git_hash_algo *to,\n+\t\t\t\t const char *buffer, size_t size)\n+{\n+\tconst char *tail = buffer;\n+\tconst char *bufptr = buffer;\n+\tconst int tree_entry_len = from->hexsz + 5;\n+\tconst int parent_entry_len = from->hexsz + 7;\n+\tstruct object_id oid, mapped_oid;\n+\tconst char *p;\n+\n+\ttail += size;\n+\tif (tail <= bufptr + tree_entry_len + 1 || memcmp(bufptr, \"tree \", 5) ||\n+\t\t\tbufptr[tree_entry_len] != '\\n')\n+\t\treturn error(\"bogus commit object\");\n+\tif (parse_oid_hex_algop(bufptr + 5, &oid, &p, from) < 0)\n+\t\treturn error(\"bad tree pointer\");\n+\n+\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\treturn error(\"unable to map tree %s in commit object\",\n+\t\t\t     oid_to_hex(&oid));\n+\tstrbuf_addf(out, \"tree %s\\n\", oid_to_hex(&mapped_oid));\n+\tbufptr = p + 1;\n+\n+\twhile (bufptr + parent_entry_len < tail && !memcmp(bufptr, \"parent \", 7)) {\n+\t\tif (tail <= bufptr + parent_entry_len + 1 ||\n+\t\t    parse_oid_hex_algop(bufptr + 7, &oid, &p, from) ||\n+\t\t    *p != '\\n')\n+\t\t\treturn error(\"bad parents in commit\");\n+\n+\t\tif (repo_oid_to_algop(the_repository, &oid, to, &mapped_oid))\n+\t\t\treturn error(\"unable to map parent %s in commit object\",\n+\t\t\t\t     oid_to_hex(&oid));\n+\n+\t\tstrbuf_addf(out, \"parent %s\\n\", oid_to_hex(&mapped_oid));\n+\t\tbufptr = p + 1;\n+\t}\n+\tstrbuf_add(out, bufptr, tail - bufptr);\n+\treturn 0;\n+}\n+\n int convert_object_file(struct strbuf *outbuf,\n \t\t\tconst struct git_hash_algo *from,\n \t\t\tconst struct git_hash_algo *to,\n@@ -101,6 +143,8 @@ int convert_object_file(struct strbuf *outbuf,\n \t\tret = convert_tree_object(outbuf, from, to, buf, len);\n \t\tbreak;\n \tcase OBJ_COMMIT:\n+\t\tret = convert_commit_object(outbuf, from, to, buf, len);\n+\t\tbreak;\n \tcase OBJ_TAG:\n \tdefault:\n \t\t/* Not implemented yet, so fail. */\n-- \n2.41.0\n\n"},{"id":"481610","messageId":"20230908231049.2035003-17-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 17/32] object-file-convert: add a function to convert trees between algorithms","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:34Z","receivedAt":"2023-09-08T23:31:46Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"From: \"brian m. carlson\" <sandals@crustytoothpaste.net>\n\nIn the future, we're going to want to provide SHA-256 repositories that\nhave compatibility support for SHA-1 as well.  In order to do so, we'll\nneed to be able to convert tree objects from SHA-256 to SHA-1 by writing\na tree with each SHA-256 object ID mapped to a SHA-1 object ID.\n\nWe implement a function, convert_tree_object, that takes an existing\ntree buffer and writes it to a new strbuf, converting between\nalgorithms.  Let's make this function generic, because while we only\nneed it to convert from the main algorithm to the compatibility\nalgorithm now, we may need to do the other way around in the future,\nsuch as for transport.\n\nWe avoid reusing the code in decode_tree_entry because that code\nnormalizes data, and we don't want that here.  We want to produce a\ncomplete round trip of data, so if, for example, the old entry had a\nwrongly zero-padded mode, we'd want to preserve that when converting to\nensure a stable hash value.\n\n****\n- Removed the repository parameter to convert_tree_object\n- Removed setting from and to defaults in convert_tree_object\n- Replaced repo_map_object with oid_to_algop\n- Replaced get_mode with parse_mode\n- Made convert_tree_object static.\n- Called convert_tree_object from convert_object_file.\n\n-- EWB\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n object-file-convert.c | 51 ++++++++++++++++++++++++++++++++++++++++++-\n 1 file changed, 50 insertions(+), 1 deletion(-)\n\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex e7c62434016d..f266c8c6cc95 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -1,8 +1,10 @@\n #include \"git-compat-util.h\"\n #include \"gettext.h\"\n #include \"strbuf.h\"\n+#include \"hex.h\"\n #include \"repository.h\"\n #include \"hash-ll.h\"\n+#include \"hash.h\"\n #include \"object.h\"\n #include \"loose.h\"\n #include \"object-file-convert.h\"\n@@ -36,6 +38,51 @@ int repo_oid_to_algop(struct repository *repo, const struct object_id *src,\n \treturn 0;\n }\n \n+static int decode_tree_entry_raw(struct object_id *oid, const char **path,\n+\t\t\t\t size_t *len, const struct git_hash_algo *algo,\n+\t\t\t\t const char *buf, unsigned long size)\n+{\n+\tuint16_t mode;\n+\tconst unsigned hashsz = algo->rawsz;\n+\n+\tif (size < hashsz + 3 || buf[size - (hashsz + 1)]) {\n+\t\treturn -1;\n+\t}\n+\n+\t*path = parse_mode(buf, &mode);\n+\tif (!*path || !**path)\n+\t\treturn -1;\n+\t*len = strlen(*path) + 1;\n+\n+\toidread_algop(oid, (const unsigned char *)*path + *len, algo);\n+\treturn 0;\n+}\n+\n+static int convert_tree_object(struct strbuf *out,\n+\t\t\t       const struct git_hash_algo *from,\n+\t\t\t       const struct git_hash_algo *to,\n+\t\t\t       const char *buffer, size_t size)\n+{\n+\tconst char *p = buffer, *end = buffer + size;\n+\n+\twhile (p < end) {\n+\t\tstruct object_id entry_oid, mapped_oid;\n+\t\tconst char *path = NULL;\n+\t\tsize_t pathlen;\n+\n+\t\tif (decode_tree_entry_raw(&entry_oid, &path, &pathlen, from, p,\n+\t\t\t\t\t  end - p))\n+\t\t\treturn error(_(\"failed to decode tree entry\"));\n+\t\tif (repo_oid_to_algop(the_repository, &entry_oid, to, &mapped_oid))\n+\t\t\treturn error(_(\"failed to map tree entry for %s\"), oid_to_hex(&entry_oid));\n+\t\tstrbuf_add(out, p, path - p);\n+\t\tstrbuf_add(out, path, pathlen);\n+\t\tstrbuf_add(out, mapped_oid.hash, to->rawsz);\n+\t\tp = path + pathlen + from->rawsz;\n+\t}\n+\treturn 0;\n+}\n+\n int convert_object_file(struct strbuf *outbuf,\n \t\t\tconst struct git_hash_algo *from,\n \t\t\tconst struct git_hash_algo *to,\n@@ -50,8 +97,10 @@ int convert_object_file(struct strbuf *outbuf,\n \t\tdie(\"Refusing noop object file conversion\");\n \n \tswitch (type) {\n-\tcase OBJ_COMMIT:\n \tcase OBJ_TREE:\n+\t\tret = convert_tree_object(outbuf, from, to, buf, len);\n+\t\tbreak;\n+\tcase OBJ_COMMIT:\n \tcase OBJ_TAG:\n \tdefault:\n \t\t/* Not implemented yet, so fail. */\n-- \n2.41.0\n\n"},{"id":"481611","messageId":"20230908231049.2035003-9-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 09/32] object-file: Update the loose object map when writing loose objects","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:26Z","receivedAt":"2023-09-09T00:02:49Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"To implement SHA1 compatibility on SHA256 repositories the loose\nobject map needs to be updated whenver a loose object is written.\nUpdating the loose object map this way allows git to support\nthe old hash algorithm in constant time.\n\nThe functions write_loose_object, and stream_loose_object are\nthe only two functions that write to the loose object store.\n\nUpdate stream_loose_object to compute the compatibiilty hash, update\nthe loose object, and then call repo_add_loose_object_map to update\nthe loose object map.\n\nUpdate write_object_file_flags to convert the object into\nit's compatibility encoding, hash the compatibility encoding,\nwrite the object, and then update the loose object map.\n\nUpdate force_object_loose to lookup the hash of the compatibility\nencoding, write the loose object, and then update the loose object\nmap.\n\nUpdate write_object_file_litterally to refuse to write any objects\nwhen a compatibility encoding is enabled.  The problem is that\nwrite_object_file_literally is frequently used to write ill-formed\nobjects.  Especially when the type of those objects is changed there\nis by definition no possibile way to convert them, as no converstion\nhas been defined.\n\nSince a compatibilty encoding can not be found and a compatibility\nmapping can not be written the cleanest behavior is to simply\ndisallow write_object_file_literraly from writing files.\n\nExcept that the loose objects are updated before the loose object map\nI have not done any analysis to see how robust this scheme is in the\nevent of failure.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n object-file.c | 94 +++++++++++++++++++++++++++++++++++++++++----------\n 1 file changed, 76 insertions(+), 18 deletions(-)\n\ndiff --git a/object-file.c b/object-file.c\nindex 7f24f19b8a68..6a14b8875343 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -44,6 +44,7 @@\n #include \"setup.h\"\n #include \"submodule.h\"\n #include \"fsck.h\"\n+#include \"loose.h\"\n \n /* The maximum size for an object header. */\n #define MAX_HEADER_LEN 32\n@@ -2035,9 +2036,12 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \t\t\t\t     const char *filename, unsigned flags,\n \t\t\t\t     git_zstream *stream,\n \t\t\t\t     unsigned char *buf, size_t buflen,\n-\t\t\t\t     git_hash_ctx *c,\n+\t\t\t\t     git_hash_ctx *c, git_hash_ctx *compat_c,\n \t\t\t\t     char *hdr, int hdrlen)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tint fd;\n \n \tfd = create_tmpfile(tmp_file, filename);\n@@ -2057,14 +2061,18 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n \tgit_deflate_init(stream, zlib_compression_level);\n \tstream->next_out = buf;\n \tstream->avail_out = buflen;\n-\tthe_hash_algo->init_fn(c);\n+\talgo->init_fn(c);\n+\tif (compat && compat_c)\n+\t\tcompat->init_fn(compat_c);\n \n \t/*  Start to feed header to zlib stream */\n \tstream->next_in = (unsigned char *)hdr;\n \tstream->avail_in = hdrlen;\n \twhile (git_deflate(stream, 0) == Z_OK)\n \t\t; /* nothing */\n-\tthe_hash_algo->update_fn(c, hdr, hdrlen);\n+\talgo->update_fn(c, hdr, hdrlen);\n+\tif (compat && compat_c)\n+\t\tcompat->update_fn(compat_c, hdr, hdrlen);\n \n \treturn fd;\n }\n@@ -2073,16 +2081,21 @@ static int start_loose_object_common(struct strbuf *tmp_file,\n  * Common steps for the inner git_deflate() loop for writing loose\n  * objects. Returns what git_deflate() returns.\n  */\n-static int write_loose_object_common(git_hash_ctx *c,\n+static int write_loose_object_common(git_hash_ctx *c, git_hash_ctx *compat_c,\n \t\t\t\t     git_zstream *stream, const int flush,\n \t\t\t\t     unsigned char *in0, const int fd,\n \t\t\t\t     unsigned char *compressed,\n \t\t\t\t     const size_t compressed_len)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tint ret;\n \n \tret = git_deflate(stream, flush ? Z_FINISH : 0);\n-\tthe_hash_algo->update_fn(c, in0, stream->next_in - in0);\n+\talgo->update_fn(c, in0, stream->next_in - in0);\n+\tif (compat && compat_c)\n+\t\tcompat->update_fn(compat_c, in0, stream->next_in - in0);\n \tif (write_in_full(fd, compressed, stream->next_out - compressed) < 0)\n \t\tdie_errno(_(\"unable to write loose object file\"));\n \tstream->next_out = compressed;\n@@ -2097,15 +2110,21 @@ static int write_loose_object_common(git_hash_ctx *c,\n  * - End the compression of zlib stream.\n  * - Get the calculated oid to \"oid\".\n  */\n-static int end_loose_object_common(git_hash_ctx *c, git_zstream *stream,\n-\t\t\t\t   struct object_id *oid)\n+static int end_loose_object_common(git_hash_ctx *c, git_hash_ctx *compat_c,\n+\t\t\t\t   git_zstream *stream, struct object_id *oid,\n+\t\t\t\t   struct object_id *compat_oid)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tint ret;\n \n \tret = git_deflate_end_gently(stream);\n \tif (ret != Z_OK)\n \t\treturn ret;\n-\tthe_hash_algo->final_oid_fn(oid, c);\n+\talgo->final_oid_fn(oid, c);\n+\tif (compat && compat_c)\n+\t\tcompat->final_oid_fn(compat_oid, compat_c);\n \n \treturn Z_OK;\n }\n@@ -2129,7 +2148,7 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \n \tfd = start_loose_object_common(&tmp_file, filename.buf, flags,\n \t\t\t\t       &stream, compressed, sizeof(compressed),\n-\t\t\t\t       &c, hdr, hdrlen);\n+\t\t\t\t       &c, NULL, hdr, hdrlen);\n \tif (fd < 0)\n \t\treturn -1;\n \n@@ -2139,14 +2158,14 @@ static int write_loose_object(const struct object_id *oid, char *hdr,\n \tdo {\n \t\tunsigned char *in0 = stream.next_in;\n \n-\t\tret = write_loose_object_common(&c, &stream, 1, in0, fd,\n+\t\tret = write_loose_object_common(&c, NULL, &stream, 1, in0, fd,\n \t\t\t\t\t\tcompressed, sizeof(compressed));\n \t} while (ret == Z_OK);\n \n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to deflate new object %s (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n-\tret = end_loose_object_common(&c, &stream, &parano_oid);\n+\tret = end_loose_object_common(&c, NULL, &stream, &parano_oid, NULL);\n \tif (ret != Z_OK)\n \t\tdie(_(\"deflateEnd on object %s failed (%d)\"), oid_to_hex(oid),\n \t\t    ret);\n@@ -2191,10 +2210,12 @@ static int freshen_packed_object(const struct object_id *oid)\n int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t\t\tstruct object_id *oid)\n {\n+\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n+\tstruct object_id compat_oid;\n \tint fd, ret, err = 0, flush = 0;\n \tunsigned char compressed[4096];\n \tgit_zstream stream;\n-\tgit_hash_ctx c;\n+\tgit_hash_ctx c, compat_c;\n \tstruct strbuf tmp_file = STRBUF_INIT;\n \tstruct strbuf filename = STRBUF_INIT;\n \tint dirlen;\n@@ -2218,7 +2239,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t */\n \tfd = start_loose_object_common(&tmp_file, filename.buf, 0,\n \t\t\t\t       &stream, compressed, sizeof(compressed),\n-\t\t\t\t       &c, hdr, hdrlen);\n+\t\t\t\t       &c, &compat_c, hdr, hdrlen);\n \tif (fd < 0) {\n \t\terr = -1;\n \t\tgoto cleanup;\n@@ -2236,7 +2257,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t\t\tif (in_stream->is_finished)\n \t\t\t\tflush = 1;\n \t\t}\n-\t\tret = write_loose_object_common(&c, &stream, flush, in0, fd,\n+\t\tret = write_loose_object_common(&c, &compat_c, &stream, flush, in0, fd,\n \t\t\t\t\t\tcompressed, sizeof(compressed));\n \t\t/*\n \t\t * Unlike write_loose_object(), we do not have the entire\n@@ -2259,7 +2280,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t */\n \tif (ret != Z_STREAM_END)\n \t\tdie(_(\"unable to stream deflate new object (%d)\"), ret);\n-\tret = end_loose_object_common(&c, &stream, oid);\n+\tret = end_loose_object_common(&c, &compat_c, &stream, oid, &compat_oid);\n \tif (ret != Z_OK)\n \t\tdie(_(\"deflateEnd on stream object failed (%d)\"), ret);\n \tclose_loose_object(fd, tmp_file.buf);\n@@ -2286,6 +2307,8 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \t}\n \n \terr = finalize_object_file(tmp_file.buf, filename.buf);\n+\tif (!err && compat)\n+\t\terr = repo_add_loose_object_map(the_repository, oid, &compat_oid);\n cleanup:\n \tstrbuf_release(&tmp_file);\n \tstrbuf_release(&filename);\n@@ -2296,17 +2319,38 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n \t\t\t    unsigned flags)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *algo = repo->hash_algo;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n+\tstruct object_id compat_oid;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen = sizeof(hdr);\n \n+\t/* Generate compat_oid */\n+\tif (compat) {\n+\t\tif (type == OBJ_BLOB)\n+\t\t\thash_object_file(compat, buf, len, type, &compat_oid);\n+\t\telse {\n+\t\t\tstruct strbuf converted = STRBUF_INIT;\n+\t\t\tconvert_object_file(&converted, algo, compat,\n+\t\t\t\t\t    buf, len, type, 0);\n+\t\t\thash_object_file(compat, converted.buf, converted.len,\n+\t\t\t\t\t type, &compat_oid);\n+\t\t\tstrbuf_release(&converted);\n+\t\t}\n+\t}\n+\n \t/* Normally if we have it in the pack then we do not bother writing\n \t * it out into .git/objects/??/?{38} file.\n \t */\n-\twrite_object_file_prepare(the_hash_algo, buf, len, type, oid, hdr,\n-\t\t\t\t  &hdrlen);\n+\twrite_object_file_prepare(algo, buf, len, type, oid, hdr, &hdrlen);\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\treturn 0;\n-\treturn write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags);\n+\tif (write_loose_object(oid, hdr, hdrlen, buf, len, 0, flags))\n+\t\treturn -1;\n+\tif (compat)\n+\t\treturn repo_add_loose_object_map(repo, oid, &compat_oid);\n+\treturn 0;\n }\n \n int write_object_file_literally(const void *buf, unsigned long len,\n@@ -2324,6 +2368,10 @@ int write_object_file_literally(const void *buf, unsigned long len,\n \n \tif (!(flags & HASH_WRITE_OBJECT))\n \t\tgoto cleanup;\n+\telse if (the_repository->compat_hash_algo) {\n+\t\tstatus = -1;\n+\t\tgoto cleanup;\n+\t}\n \tif (freshen_packed_object(oid) || freshen_loose_object(oid))\n \t\tgoto cleanup;\n \tstatus = write_loose_object(oid, header, hdrlen, buf, len, 0, 0);\n@@ -2335,9 +2383,12 @@ int write_object_file_literally(const void *buf, unsigned long len,\n \n int force_object_loose(const struct object_id *oid, time_t mtime)\n {\n+\tstruct repository *repo = the_repository;\n+\tconst struct git_hash_algo *compat = repo->compat_hash_algo;\n \tvoid *buf;\n \tunsigned long len;\n \tstruct object_info oi = OBJECT_INFO_INIT;\n+\tstruct object_id compat_oid;\n \tenum object_type type;\n \tchar hdr[MAX_HEADER_LEN];\n \tint hdrlen;\n@@ -2350,8 +2401,15 @@ int force_object_loose(const struct object_id *oid, time_t mtime)\n \toi.contentp = &buf;\n \tif (oid_object_info_extended(the_repository, oid, &oi, 0))\n \t\treturn error(_(\"cannot read object for %s\"), oid_to_hex(oid));\n+\tif (compat) {\n+\t\tif (repo_oid_to_algop(repo, oid, compat, &compat_oid))\n+\t\t\treturn error(_(\"cannot map object %s to %s\"),\n+\t\t\t\t     oid_to_hex(oid), compat->name);\n+\t}\n \thdrlen = format_object_header(hdr, sizeof(hdr), type, len);\n \tret = write_loose_object(oid, hdr, hdrlen, buf, len, mtime, 0);\n+\tif (!ret && compat)\n+\t\tret = repo_add_loose_object_map(the_repository, oid, &compat_oid);\n \tfree(buf);\n \n \treturn ret;\n-- \n2.41.0\n\n"},{"id":"481612","messageId":"20230908231049.2035003-11-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 11/32] pack: Communicate the compat_oid through struct pack_idx_entry","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:28Z","receivedAt":"2023-09-09T00:05:19Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Add compat_oid into struct pack_idx_entry to allow communicating the\nthe compat hash value of the objects being indexed to the code that\nbuilds the indexes for a pack.\n\nHaving a mechanism that communicates the compat_oid from the code\nbuilding the pack is necessary for bulk-checkin, fast-import, and\nindex-pack.  Only pack-objects could rely on the existing\ncomaptibility mappings, but there is not point since the\nother creators of indexes can't.\n\nUnfortunately this adds a 4 byte hole into struct pack_idx_entry.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n pack.h | 1 +\n 1 file changed, 1 insertion(+)\n\ndiff --git a/pack.h b/pack.h\nindex 3ab9e3f60c0b..321d38374f70 100644\n--- a/pack.h\n+++ b/pack.h\n@@ -75,6 +75,7 @@ struct pack_idx_header {\n  */\n struct pack_idx_entry {\n \tstruct object_id oid;\n+\tstruct object_id compat_oid;\n \tuint32_t crc32;\n \toff_t offset;\n };\n-- \n2.41.0\n\n"},{"id":"481613","messageId":"20230908231049.2035003-5-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 05/32] repository: add a compatibility hash algorithm","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:22Z","receivedAt":"2023-09-09T00:19:18Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"We currently have support for using a full stage 4 SHA-256\nimplementation.  However, we'd like to support interoperability with\nSHA-1 repositories as well.  The transition plan anticipates a\ncompatibility hash algorithm configuration option that we can use to\nimplement support for this.  Let's add an element to the repository\nstructure that indicates the compatibility hash algorithm so we can use\nit when we need to consider interoperability between algorithms.\n\nFor now, we always set it to NULL, but we'll initialize it differently\nin the future.\n\nInspired-by: brian m. carlson <sandals@crustytoothpaste.net>\nSigned-off-by: Eric W. Biederman <ebiederm@xmission.com>\n---\n repository.h | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/repository.h b/repository.h\nindex 5f18486f6465..6c4130f0c36e 100644\n--- a/repository.h\n+++ b/repository.h\n@@ -160,6 +160,9 @@ struct repository {\n \t/* Repository's current hash algorithm, as serialized on disk. */\n \tconst struct git_hash_algo *hash_algo;\n \n+\t/* Repository's compatibility hash algorithm. */\n+\tconst struct git_hash_algo *compat_hash_algo;\n+\n \t/* A unique-id for tracing purposes. */\n \tint trace2_repo_id;\n \n-- \n2.41.0\n\n"},{"id":"481614","messageId":"20230908231049.2035003-13-ebiederm@xmission.com","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"[PATCH 13/32] object-file: Add a compat_oid_in parameter to write_object_file_flags","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-08T23:10:30Z","receivedAt":"2023-09-09T00:40:47Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"To create the proper signatures for commit objects both versions of\nthe commit object need to be generated and signed.  After that it is\na waste to throw away the work of generating the compatibility hash\nso update write_object_file_flags to take a compatibility hash input\nparameter that it can use to skip the work of generating the\ncompatability hash.\n\nUpdate the places that don't generate the compatability hash to\npass NULL so it is easy to tell write_object_file_flags should\nnot attempt to use their compatability hash.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n cache-tree.c      | 2 +-\n object-file.c     | 6 ++++--\n object-store-ll.h | 4 ++--\n 3 files changed, 7 insertions(+), 5 deletions(-)\n\ndiff --git a/cache-tree.c b/cache-tree.c\nindex 641427ed410a..ddc7d3d86959 100644\n--- a/cache-tree.c\n+++ b/cache-tree.c\n@@ -448,7 +448,7 @@ static int update_one(struct cache_tree *it,\n \t\thash_object_file(the_hash_algo, buffer.buf, buffer.len,\n \t\t\t\t OBJ_TREE, &it->oid);\n \t} else if (write_object_file_flags(buffer.buf, buffer.len, OBJ_TREE,\n-\t\t\t\t\t   &it->oid, flags & WRITE_TREE_SILENT\n+\t\t\t\t\t   &it->oid, NULL, flags & WRITE_TREE_SILENT\n \t\t\t\t\t   ? HASH_SILENT : 0)) {\n \t\tstrbuf_release(&buffer);\n \t\treturn -1;\ndiff --git a/object-file.c b/object-file.c\nindex 6cc4ae1fd957..fd420dd303df 100644\n--- a/object-file.c\n+++ b/object-file.c\n@@ -2317,7 +2317,7 @@ int stream_loose_object(struct input_stream *in_stream, size_t len,\n \n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n-\t\t\t    unsigned flags)\n+\t\t\t    struct object_id *compat_oid_in, unsigned flags)\n {\n \tstruct repository *repo = the_repository;\n \tconst struct git_hash_algo *algo = repo->hash_algo;\n@@ -2328,7 +2328,9 @@ int write_object_file_flags(const void *buf, unsigned long len,\n \n \t/* Generate compat_oid */\n \tif (compat) {\n-\t\tif (type == OBJ_BLOB)\n+\t\tif (compat_oid_in)\n+\t\t\toidcpy(&compat_oid, compat_oid_in);\n+\t\telse if (type == OBJ_BLOB)\n \t\t\thash_object_file(compat, buf, len, type, &compat_oid);\n \t\telse {\n \t\t\tstruct strbuf converted = STRBUF_INIT;\ndiff --git a/object-store-ll.h b/object-store-ll.h\nindex bc76d6bec80d..c5f2bb2fc2fe 100644\n--- a/object-store-ll.h\n+++ b/object-store-ll.h\n@@ -255,11 +255,11 @@ void hash_object_file(const struct git_hash_algo *algo, const void *buf,\n \n int write_object_file_flags(const void *buf, unsigned long len,\n \t\t\t    enum object_type type, struct object_id *oid,\n-\t\t\t    unsigned flags);\n+\t\t\t    struct object_id *comapt_oid_in, unsigned flags);\n static inline int write_object_file(const void *buf, unsigned long len,\n \t\t\t\t    enum object_type type, struct object_id *oid)\n {\n-\treturn write_object_file_flags(buf, len, type, oid, 0);\n+\treturn write_object_file_flags(buf, len, type, oid, NULL, 0);\n }\n \n int write_object_file_literally(const void *buf, unsigned long len,\n-- \n2.41.0\n\n"},{"id":"481626","messageId":"87msxvlczq.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"Re: [RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-09T12:58:33Z","receivedAt":"2023-09-09T12:59:11Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\nI forgot to mention the patches are against 2.42.\n\nEric\n"},{"id":"481651","messageId":"ZP3Rr8Ei0sG0lg0R@tapette.crustytoothpaste.net","threadId":"60212","inReplyTo":"20230908231049.2035003-1-ebiederm@xmission.com","subject":"Re: [PATCH 01/32] doc hash-file-transition: A map file for mapping between sha1 and sha256","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-09-10T14:24:47Z","receivedAt":"2023-09-10T14:24:52Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-09-08 at 23:10:18, Eric W. Biederman wrote:\n> The v3 pack index file as documented has a lot of complexity making it\n> difficult to implement correctly.  I worked with bryan's preliminary\n> implementation and it took several passes to get the bugs out.\n> \n> The complexity also requires multiple table look-ups to find all of\n> the information that is needed to translate from one kind of oid to\n> another.  Which can't be good for cache locality.\n> \n> Even worse coming up with a new index file version requires making\n> changes that have the potentialy to break anything that uses the index\n> of a pack file.\n> \n> Instead of continuing to deal with the chance of braking things\n> besides the oid mapping functionality, the additional complexity in\n> the file format, and worry if the performance would be reasonable I\n> stripped down the problem to it's fundamental complexity and came up\n> with a file format that is exactly about mapping one kind of oid to\n> another, and only supports two kinds of oids.\n\nI think this is a fine approach, and as I'm sure you noticed from my\nseries, it's a lot more robust than trying to implement pack v3.  I'd be\nfine with going with this approach instead of pack v3.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"481652","messageId":"ZP3UCQf+9D/J3wqT@tapette.crustytoothpaste.net","threadId":"60212","inReplyTo":"20230908231049.2035003-2-ebiederm@xmission.com","subject":"Re: [PATCH 02/32] doc hash-function-transition: Replace compatObjectFormat with compatMap","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-09-10T14:34:49Z","receivedAt":"2023-09-10T14:34:53Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-09-08 at 23:10:19, Eric W. Biederman wrote:\n> Ir makes a lot of sense for the hash algorithm that determines how all\n\nMinor nit: \"It\".\n\n> diff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\n> index 4b937480848a..10572c5794f9 100644\n> --- a/Documentation/technical/hash-function-transition.txt\n> +++ b/Documentation/technical/hash-function-transition.txt\n> @@ -148,14 +148,14 @@ Detailed Design\n>  Repository format extension\n>  ~~~~~~~~~~~~~~~~~~~~~~~~~~~\n>  A SHA-256 repository uses repository format version `1` (see\n> -Documentation/technical/repository-version.txt) with extensions\n> -`objectFormat` and `compatObjectFormat`:\n> +Documentation/technical/repository-version.txt) with the extension\n> +`objectFormat`, and an optional core.compatMap configuration.\n>  \n>  \t[core]\n>  \t\trepositoryFormatVersion = 1\n> +\t\tcompatMap = on\n>  \t[extensions]\n>  \t\tobjectFormat = sha256\n> -\t\tcompatObjectFormat = sha1\n\nWhile I'm in favour of an approach that uses the compat map, the\nsituation we've implemented here doesn't specify the extra hash\nalgorithm.  We want this approach to work just as well for moving from\nSHA-1 to SHA-256 as it might for a future transition from SHA-256 to,\nsay, SHA-3-512, if that becomes necessary.\n\nMaking a future transition easier has been a goal of my SHA-256 work\n(because who wants to write several hundred patches in such a case?), so\nmy hope is we can keep that here as well by explicitly naming the\nalgorithm we're using.\n\nI also wonder if an approach that doesn't use an extension is going to\nbe helpful.  Say, that I have a repository that is using Git 3.x, which\nsupports interop, but I also need to use Git 2.x, which does not.  While\nit's true that Git 2.x can read my SHA-256 repository, it won't write\nthe appropriate objects into the map, and thus it will be practically\nvery difficult to actually use Git 3.x to push data to a repository of a\ndifferent hash function.  We might well prefer to have Git 2.x not work\nwith the repository at all rather than have incomplete data preventing\nus from, well, interoperating.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"481653","messageId":"ZP3i9WdpDKlsWuNP@tapette.crustytoothpaste.net","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"Re: [RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-09-10T15:38:29Z","receivedAt":"2023-09-10T15:38:43Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-09-08 at 23:05:52, Eric W. Biederman wrote:\n> \n> I would like to see the SHA256 transition happen so I started playing\n> with the k2204-transition-interop branch of brian m. carlson's tree.\n> \n> Before I go farther I need to some other folks to look at this and see\n> if this is a general direction that the git project can stand.\n\nI'm really excited to see this and I think it's a great way forward.\nI've taken a brief look at each patch, and I don't see anything that\nshould be a dealbreaker.  I left a few comments, although I think your\nmailserver is blocking mine at the moment, so you may not have received\nthem (hopefully you can read them on the list in the interim).\n\nYou may also feel free to simply adjust the commit message for the\npatches of mine you've modified without needing to document that you've\nchanged them.  I expect that you will have changed them when you submit\nthem, if only to resolve conflicts.  After all, Junio does so all the\ntime.\n\n> This patchset is not complete it does not implement converting a\n> received pack of the compatibility hash into the hash function of the\n> repository, nor have I written any automated tests.  Both need to happen\n> before this is finalized.\n\nSpeaking of tests, one set of tests I had intended to write and think\nshould be written, but had not yet implemented, is tests for\nround-tripping objects.  That is, the SHA-1 value we get for a revision\nin a pure SHA-1 repository should obviously be the same as the SHA-1\nvalue we get in a SHA-256 repository in interop mode, and we should be\nable to use the `test_oid_cache` functionality to hard-code the desired\nobjects.  I think it would be also helpful to do this for fixed objects\nthat are doubly-signed (with both algorithms) as well, since that's a\ntricky edge case that we'll want to avoid breaking.  Other edge cases\nwill include things like merge commits, including octopus merges.\n\nBut overall, I think this is a great improvement, and I'm very excited\nto see someone picking up some of this work and moving it forward.\nThanks for doing so.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"481654","messageId":"87bke9hprr.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"ZP3UCQf+9D/J3wqT@tapette.crustytoothpaste.net","subject":"Re: [PATCH 02/32] doc hash-function-transition: Replace compatObjectFormat with compatMap","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-10T18:00:40Z","receivedAt":"2023-09-10T18:01:19Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> On 2023-09-08 at 23:10:19, Eric W. Biederman wrote:\n>> Ir makes a lot of sense for the hash algorithm that determines how all\n>\n> Minor nit: \"It\".\n>\n>> diff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\n>> index 4b937480848a..10572c5794f9 100644\n>> --- a/Documentation/technical/hash-function-transition.txt\n>> +++ b/Documentation/technical/hash-function-transition.txt\n>> @@ -148,14 +148,14 @@ Detailed Design\n>>  Repository format extension\n>>  ~~~~~~~~~~~~~~~~~~~~~~~~~~~\n>>  A SHA-256 repository uses repository format version `1` (see\n>> -Documentation/technical/repository-version.txt) with extensions\n>> -`objectFormat` and `compatObjectFormat`:\n>> +Documentation/technical/repository-version.txt) with the extension\n>> +`objectFormat`, and an optional core.compatMap configuration.\n>>  \n>>  \t[core]\n>>  \t\trepositoryFormatVersion = 1\n>> +\t\tcompatMap = on\n>>  \t[extensions]\n>>  \t\tobjectFormat = sha256\n>> -\t\tcompatObjectFormat = sha1\n>\n> While I'm in favour of an approach that uses the compat map, the\n> situation we've implemented here doesn't specify the extra hash\n> algorithm.  We want this approach to work just as well for moving from\n> SHA-1 to SHA-256 as it might for a future transition from SHA-256 to,\n> say, SHA-3-512, if that becomes necessary.\n>\n> Making a future transition easier has been a goal of my SHA-256 work\n> (because who wants to write several hundred patches in such a case?), so\n> my hope is we can keep that here as well by explicitly naming the\n> algorithm we're using.\n>\n> I also wonder if an approach that doesn't use an extension is going to\n> be helpful.  Say, that I have a repository that is using Git 3.x, which\n> supports interop, but I also need to use Git 2.x, which does not.  While\n> it's true that Git 2.x can read my SHA-256 repository, it won't write\n> the appropriate objects into the map, and thus it will be practically\n> very difficult to actually use Git 3.x to push data to a repository of a\n> different hash function.  We might well prefer to have Git 2.x not work\n> with the repository at all rather than have incomplete data preventing\n> us from, well, interoperating.\n\nFirst it is my hope that we can get a command such as \"git gc\" to scan\nthe repository and fill in all of the missing compatibility hashes.\n\nNot so much for day to day work, but for people able to enable\ncompatibility hashes on an existing repository.  Enabling compatibility\nhashes on a sha1 repository is going to be necessary to create a sha256\nrepository from it.  A depth first walk, or a topological sort of the\nobjects pretty much has to happen as a separate pass.  So it makes sense\njust to require all of the objects have their compatibility hash\ncomputed before attempting to generate a pack in the compatibility\nformat.\n\nI say all of that and I feel silly.\n\nThe core and optimized path is what whatever receive pack does to deal\nwith a pack in the repositories compatibility format.  Once that is\nbuilt we can create a sha256 repository from a sha1 repository just\nby cloning it, and letting receive-pack figure out the details.\n\nBefore we can generate a sha256 pack from a sha1 pack we still need\nto compute the sha256 hash of every object, but that can be very\noptimized and local to the case of receiving a non-native pack.  So a\nrepository that generates a compatibility hash for all of it's objects\nis not necessary to transition to another hash algorithm.  All we need\nis another repository in the other format.\n\n\nThat said there is value in being able to add compatibility hashes\nto an existing repository.  The upstream repository can just convert\nto the new hash function and all of the downstream repositories\ncan compute their compatibility hashes and convert when they are ready.\n\nBasically once a git with transition support exists any repository can\nconvert at any time without creating a problem for other repositories.\n\nIn my head it seems cheaper/safer to compute the compatibility hash of\nevery object in an existing repository than it does to convert a\nrepository.  Is it?\n\nI think that if the first pull from a repository in another format can\ntrigger the initial computation of the compatibility hash (like the\nfirst use of a reverse index triggers the creation of the reverse\nindex), then it will definitely be easier to just enable compatibility\nhashes in an existing repository.\n\nThe additional hash computation step every pull from upstream (even when\nwell optimized) should be an incentive for people to fully convert their\nrepositories after the upstream has converted.\n\n\nThat is when things get tricky and the transition plan has not talked\nabout.  There are references to existing oid's in email, bug trackers,\nand commit comments.  Digging through the history and dealing with those\nreferences is something that developers are going to need to do for the\nrest of the life of a project.\n\nWhich means eventually we will need to support a mode where we have some\npacks with a ``.compat'' index but we no longer compute or generate the\nold hash for new objects.\n\nIn summary.  I agree that compatMap is likely insufficient. So far I\nthink it is too cheap/easy to generate the missing mappings to make it a\nmandatory requirement that all operations always generate them.\n\nI also agree that making the configuration resilient foreseeable future\ndemands is a good idea.\n\nSo I will push this change farther out in the patch series.\n\nEric\n"},{"id":"481655","messageId":"87zg1tgaw6.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"ZP3Rr8Ei0sG0lg0R@tapette.crustytoothpaste.net","subject":"Re: [PATCH 01/32] doc hash-file-transition: A map file for mapping between sha1 and sha256","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-10T18:07:21Z","receivedAt":"2023-09-10T18:07:34Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> On 2023-09-08 at 23:10:18, Eric W. Biederman wrote:\n>> The v3 pack index file as documented has a lot of complexity making it\n>> difficult to implement correctly.  I worked with bryan's preliminary\n>> implementation and it took several passes to get the bugs out.\n>> \n>> The complexity also requires multiple table look-ups to find all of\n>> the information that is needed to translate from one kind of oid to\n>> another.  Which can't be good for cache locality.\n>> \n>> Even worse coming up with a new index file version requires making\n>> changes that have the potentialy to break anything that uses the index\n>> of a pack file.\n>> \n>> Instead of continuing to deal with the chance of braking things\n>> besides the oid mapping functionality, the additional complexity in\n>> the file format, and worry if the performance would be reasonable I\n>> stripped down the problem to it's fundamental complexity and came up\n>> with a file format that is exactly about mapping one kind of oid to\n>> another, and only supports two kinds of oids.\n>\n> I think this is a fine approach, and as I'm sure you noticed from my\n> series, it's a lot more robust than trying to implement pack v3.  I'd be\n> fine with going with this approach instead of pack v3.\n\nI think I got your pack v3 working but it was at a minimum a serious\ndistraction.\n\nI worry a little bit that this might leave some performance on the\ntable, with something like a 256 way jump table like we have in the\nindex file.\n\nStill I figure we can start simple and when we start optimizing and\nprofiling we can revisit the format if it shows up as a performance\nissue.\n\nEric\n\n"},{"id":"481656","messageId":"878r9dgaad.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"ZP3i9WdpDKlsWuNP@tapette.crustytoothpaste.net","subject":"Re: [RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-10T18:20:26Z","receivedAt":"2023-09-10T18:20:39Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> On 2023-09-08 at 23:05:52, Eric W. Biederman wrote:\n>> \n>> I would like to see the SHA256 transition happen so I started playing\n>> with the k2204-transition-interop branch of brian m. carlson's tree.\n>> \n>> Before I go farther I need to some other folks to look at this and see\n>> if this is a general direction that the git project can stand.\n>\n> I'm really excited to see this and I think it's a great way forward.\n> I've taken a brief look at each patch, and I don't see anything that\n> should be a dealbreaker.  I left a few comments, although I think your\n> mailserver is blocking mine at the moment, so you may not have received\n> them (hopefully you can read them on the list in the interim).\n\nI can.  I will see if I can figure out what is happening with direct\nreception tomorrow.\n\n> You may also feel free to simply adjust the commit message for the\n> patches of mine you've modified without needing to document that you've\n> changed them.  I expect that you will have changed them when you submit\n> them, if only to resolve conflicts.  After all, Junio does so all the\n> time.\n\nThanks.  I was doing my best at striking a balance between giving credit\nwhere is credit is due, and pointing out the bugs are probably mine.\n\n>> This patchset is not complete it does not implement converting a\n>> received pack of the compatibility hash into the hash function of the\n>> repository, nor have I written any automated tests.  Both need to happen\n>> before this is finalized.\n>\n> Speaking of tests, one set of tests I had intended to write and think\n> should be written, but had not yet implemented, is tests for\n> round-tripping objects.  That is, the SHA-1 value we get for a revision\n> in a pure SHA-1 repository should obviously be the same as the SHA-1\n> value we get in a SHA-256 repository in interop mode, and we should be\n> able to use the `test_oid_cache` functionality to hard-code the desired\n> objects.  I think it would be also helpful to do this for fixed objects\n> that are doubly-signed (with both algorithms) as well, since that's a\n> tricky edge case that we'll want to avoid breaking.  Other edge cases\n> will include things like merge commits, including octopus merges.\n\nYes.  I think we can use cat-file to do that.  Have two repositories one\nin each format.  Verify that when cat-file prints out an object given\nthe native oid cat-file prints out what was put in.  Similarly verify\nthat when cat-file prints out an object given the compatibility oid\ncat-file prints out the expected conversion.  That logic performed in\nboth repositories should work.\n\n> But overall, I think this is a great improvement, and I'm very excited\n> to see someone picking up some of this work and moving it forward.\n> Thanks for doing so.\n\nThanks.\n\nThen next goal is to get enough merged that I can test the round-trip\nconversions.  More than anything else we need to know the conversion\nfunctionality is solid.\n\nPlus I expect that while 32 patches were important to show the scope of\nthe work, but a bit much to fully review and merge all at once.\n\nEric\n"},{"id":"481666","messageId":"xmqqy1hdi6hp.fsf@gitster.g","threadId":"60212","inReplyTo":"ZP3UCQf+9D/J3wqT@tapette.crustytoothpaste.net","subject":"Re: [PATCH 02/32] doc hash-function-transition: Replace compatObjectFormat with compatMap","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T06:11:46Z","receivedAt":"2023-09-11T06:11:50Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n>> +Documentation/technical/repository-version.txt) with the extension\n>> +`objectFormat`, and an optional core.compatMap configuration.\n>>  \n>>  \t[core]\n>>  \t\trepositoryFormatVersion = 1\n>> +\t\tcompatMap = on\n>>  \t[extensions]\n>>  \t\tobjectFormat = sha256\n>> -\t\tcompatObjectFormat = sha1\n>\n> While I'm in favour of an approach that uses the compat map, the\n> situation we've implemented here doesn't specify the extra hash\n> algorithm.  We want this approach to work just as well for moving from\n> SHA-1 to SHA-256 as it might for a future transition from SHA-256 to,\n> say, SHA-3-512, if that becomes necessary.\n>\n> Making a future transition easier has been a goal of my SHA-256 work\n> (because who wants to write several hundred patches in such a case?), so\n> my hope is we can keep that here as well by explicitly naming the\n> algorithm we're using.\n>\n> I also wonder if an approach that doesn't use an extension is going to\n> be helpful.  Say, that I have a repository that is using Git 3.x, which\n> supports interop, but I also need to use Git 2.x, which does not.  While\n> it's true that Git 2.x can read my SHA-256 repository, it won't write\n> the appropriate objects into the map, and thus it will be practically\n> very difficult to actually use Git 3.x to push data to a repository of a\n> different hash function.  We might well prefer to have Git 2.x not work\n> with the repository at all rather than have incomplete data preventing\n> us from, well, interoperating.\n\nVery sensible line of thought and suggestion to move the topic\nforward.  Very much appreciated.\n\n"},{"id":"481667","messageId":"xmqqsf7li686.fsf@gitster.g","threadId":"60212","inReplyTo":"20230908231049.2035003-12-ebiederm@xmission.com","subject":"Re: [PATCH 12/32] bulk-checkin: hash object with compatibility algorithm","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T06:17:29Z","receivedAt":"2023-09-11T06:17:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@xmission.com> writes:\n\n>  \tstruct hashfile_checkpoint checkpoint = {0};\n>  \tstruct pack_idx_entry *idx = NULL;\n> +\tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n> +\tstruct object_id compat_oid = {};\n\n\nbulk-checkin.c:267:39: error: ISO C forbids empty initializer braces [-Werror=pedantic]\n\n"},{"id":"481668","messageId":"xmqqo7i9i5v8.fsf@gitster.g","threadId":"60212","inReplyTo":"20230908231049.2035003-14-ebiederm@xmission.com","subject":"Re: [PATCH 14/32] commit: write commits for both hashes","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T06:25:15Z","receivedAt":"2023-09-11T06:25:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@xmission.com> writes:\n\n> +\tstruct strbuf sig = STRBUF_INIT, compat_sig = STRBUF_INIT;\n> +\tstruct object_id *parent_buf = NULL;\n> +\tstruct object_id compat_oid = {};\n\nDitto.\n\n\tstruct object_id compat_oid = { 0 };\n\nwould be our zero-initialization convention.\n\nThanks.\n"},{"id":"481669","messageId":"xmqqjzsxi5qm.fsf@gitster.g","threadId":"60212","inReplyTo":"20230908231049.2035003-26-ebiederm@xmission.com","subject":"Re: [PATCH 26/32] object-file-convert: Implement convert_object_file_{begin,step,end}","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T06:28:01Z","receivedAt":"2023-09-11T06:28:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@xmission.com> writes:\n\n> +\tconst struct git_hash_algo *from = state->from;\n> +\tconst struct git_hash_algo *to = state->to;\n> +\tstruct strbuf *out = state->outbuf;\n> +\tconst char *buffer = state->buf;\n> +\tsize_t payload_size, size = state->buf_len;;\n\nThe excess ';' at the end is an empty statment, hence ...\n\n> +\tstruct object_id oid;\n>  \tconst char *p;\n> +\tint ret = 0;\n\n... these three violate our \"no declaration after statement\" house rule.\n"},{"id":"481670","messageId":"xmqqfs3li5ly.fsf@gitster.g","threadId":"60212","inReplyTo":"20230908231049.2035003-25-ebiederm@xmission.com","subject":"Re: [PATCH 25/32] pack-compat-map: Add support for .compat files of a packfile","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T06:30:49Z","receivedAt":"2023-09-11T06:30:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@xmission.com> writes:\n\n> diff --git a/pack-write.c b/pack-write.c\n> index b19ddf15b284..f22eea964f77 100644\n> --- a/pack-write.c\n> +++ b/pack-write.c\n> @@ -12,6 +12,7 @@\n>  #include \"pack-revindex.h\"\n>  #include \"path.h\"\n>  #include \"strbuf.h\"\n> +#include \"object-file-convert.h\"\n> ...\n> +/*\n> + * The *hash contains the pack content hash.\n> + * The objects array is passed in sorted.\n> + */\n> +const char *write_compat_map_file(const char *compat_map_name,\n> +\t\t\t\t  struct pack_idx_entry **objects,\n> +\t\t\t\t  int nr_objects, const unsigned char *hash)\n\nInclude \"pack-compat-map.h\"; otherwise the compiler would complain\nfor missing prototypes.\n"},{"id":"481671","messageId":"xmqq8r9di5ba.fsf@gitster.g","threadId":"60212","inReplyTo":"87sf7ol0z3.fsf@email.froward.int.ebiederm.org","subject":"Re: [RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-11T06:37:13Z","receivedAt":"2023-09-11T06:37:21Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@xmission.com> writes:\n\n> I would like to see the SHA256 transition happen so I started playing\n> with the k2204-transition-interop branch of brian m. carlson's tree.\n\nI needed these tweaks to build the series standalone on 'master' (or\n2.42).  There are semantic merge conflicts with some topics in flight\nwhen this is merged to 'seen', so it may take me a bit more time to\npush the integration result.\n\nThanks.\n\n builtin/fast-import.c | 2 +-\n bulk-checkin.c        | 2 +-\n commit.c              | 2 +-\n object-file-convert.c | 4 ++--\n pack-write.c          | 1 +\n 5 files changed, 6 insertions(+), 5 deletions(-)\n\ndiff --git a/builtin/fast-import.c b/builtin/fast-import.c\nindex 66c471bc73..93cc4a491c 100644\n--- a/builtin/fast-import.c\n+++ b/builtin/fast-import.c\n@@ -784,7 +784,7 @@ struct pack_index_names {\n \n static struct pack_index_names create_index(void)\n {\n-\tstruct pack_index_names tmp = {};\n+\tstruct pack_index_names tmp = { 0 };\n \tstruct pack_idx_entry **idx, **c, **last;\n \tstruct object_entry *e;\n \tstruct object_entry_pool *o;\ndiff --git a/bulk-checkin.c b/bulk-checkin.c\nindex 3206412a19..d63b3ffa01 100644\n--- a/bulk-checkin.c\n+++ b/bulk-checkin.c\n@@ -264,7 +264,7 @@ static int deflate_blob_to_pack(struct bulk_checkin_packfile *state,\n \tstruct hashfile_checkpoint checkpoint = {0};\n \tstruct pack_idx_entry *idx = NULL;\n \tconst struct git_hash_algo *compat = the_repository->compat_hash_algo;\n-\tstruct object_id compat_oid = {};\n+\tstruct object_id compat_oid = { 0 };\n \n \tseekback = lseek(fd, 0, SEEK_CUR);\n \tif (seekback == (off_t) -1)\ndiff --git a/commit.c b/commit.c\nindex 54f19ed032..2e2b805d5e 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1654,7 +1654,7 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n \tstruct strbuf buffer, compat_buffer;\n \tstruct strbuf sig = STRBUF_INIT, compat_sig = STRBUF_INIT;\n \tstruct object_id *parent_buf = NULL;\n-\tstruct object_id compat_oid = {};\n+\tstruct object_id compat_oid = { 0 };\n \tsize_t i, nparents;\n \n \t/* Not having i18n.commitencoding is the same as having utf-8 */\ndiff --git a/object-file-convert.c b/object-file-convert.c\nindex 2306e17dd5..148e61d24f 100644\n--- a/object-file-convert.c\n+++ b/object-file-convert.c\n@@ -26,7 +26,7 @@ int repo_submodule_oid_to_algop(struct repository *repo,\n \n \tfor (i = 0; i < repo->index->cache_nr; i++) {\n \t\tconst struct cache_entry *ce = repo->index->cache[i];\n-\t\tstruct repository subrepo = {};\n+\t\tstruct repository subrepo = { 0 };\n \t\tint ret;\n \n \t\tif (!S_ISGITLINK(ce->ce_mode))\n@@ -205,7 +205,7 @@ static int convert_tag_object_step(struct object_file_convert_state *state)\n \tconst struct git_hash_algo *to = state->to;\n \tstruct strbuf *out = state->outbuf;\n \tconst char *buffer = state->buf;\n-\tsize_t payload_size, size = state->buf_len;;\n+\tsize_t payload_size, size = state->buf_len;\n \tstruct object_id oid;\n \tconst char *p;\n \tint ret = 0;\ndiff --git a/pack-write.c b/pack-write.c\nindex f22eea964f..b2ec09737e 100644\n--- a/pack-write.c\n+++ b/pack-write.c\n@@ -7,6 +7,7 @@\n #include \"remote.h\"\n #include \"chunk-format.h\"\n #include \"pack-mtimes.h\"\n+#include \"pack-compat-map.h\"\n #include \"oidmap.h\"\n #include \"pack-objects.h\"\n #include \"pack-revindex.h\"\n-- \n2.42.0-158-g94e83dcf5b\n\n"},{"id":"481689","messageId":"87cyyoeli0.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"xmqq8r9di5ba.fsf@gitster.g","subject":"Re: [RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-11T16:13:27Z","receivedAt":"2023-09-11T21:39:00Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Eric W. Biederman\" <ebiederm@xmission.com> writes:\n>\n>> I would like to see the SHA256 transition happen so I started playing\n>> with the k2204-transition-interop branch of brian m. carlson's tree.\n>\n> I needed these tweaks to build the series standalone on 'master' (or\n> 2.42).  There are semantic merge conflicts with some topics in flight\n> when this is merged to 'seen', so it may take me a bit more time to\n> push the integration result.\n\nJunio, brian for the very warm reception of this, it is very\nencouraging.\n\nI am not worried about what it will take time to get the changes I\nposted into the integration.  I had only envisioned them as good enough\nto get the technical ideas across, and had never envisioned them as\nbeing accepted as is.\n\nWhat I am envisioning as my future directions are:\n\n- Post non controversial cleanups, so they can be merged.\n  (I can only see about 4 of them the most significant is:\n   bulk-checkin: Only accept blobs)\n\n- Sort out the configuration options\n\n- Post the smallest patchset I can that will allow testing the code in\n  object-file-convert.c.  Unfortunately for that I need configuration\n  options to enable the mapping.\n\n  In starting to write the tests I have already found a bug in\n  the conversion of tags (an extra newline is added), and I haven't\n  even gotten to testing the tricky bits with signatures.\n\n- Once the object file conversion is tested and is solid work on\n  the more substantial pieces.\n\nDoes that sound like a reasonable plan?\n\nEric\n"},{"id":"481700","messageId":"87sf7kd5xg.fsf_-_@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"xmqqy1hdi6hp.fsf@gitster.g","subject":"[PATCH v2 02/32] doc hash-function-transition: Replace compatObjectFormat with mapObjectFormat","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-11T16:35:07Z","receivedAt":"2023-09-11T21:39:18Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\nDeeply and fundamentally the plan is to only operate one one hash\nfunction for the core of git, to use only one hash function for what\nis stored in the repository.\n\nTo avoid requring a flag day to transition hash functions for naming\nobjects, and to support being able to access objects using legacy object\nnames a mapping functionality will be provided.\n\nWe want to provide user facing configuration that is robust enough\nthat it can accomodate multiple different scenarios on how git\nevolves and how people use their repositories.\n\nThere are two different ways it is envisioned to use mapped object\nids.  The first is to require every object in the repository to have a\nmapping, so that pushes and pulls from repositories using a different\nhash algorithm can work.  The second is to have an incomplete mapping\nof object ids so that old references to objects in emails, commit\nmessages, bug trackers and are usable in a read-only manner\nwith tools like \"git show\".\n\nThe first way fundamentally needs every object in the repository to\nhave a mapping, which requires the repository to be marked incompatible\nfor writes fron older versions of git.  Thus the mapObjectFormat option\nis placed in [extensions].\n\nThe ext2 family of filesystems has 3 ways of describing new features\ncompatible, read-only-compatible, and incompatible.  The current git\nconfigurtation has compat (any feature mentioned anywhere in the\nconfiguration outside of [extensions] section), and incompatible (any\nconfiguration inside of the [extensions] section.  It would be nice to\nhave a read-only compatible section for the mandatory mapping\nfunction.  Would it be worth adding it now so that we have it for\nfuture extensions?\n\nHaving a mapping that is just used in a read-only mode for looking up\nold objects with old object ids will be needed post-transition.  Such\na mode does not require computing the old hash function or even\nsupport automatically writing any new mappings.  So it is completely\nsafe to enable in a backwards compatible mode.  Fort that let's\nuse core.readObjectMap to make it clear the mappings only read.\n\nI have documented that both of the options readObjectMap and\nmapObjectFormat can be specified multiple times if that is needed to\nsupport the desired configuration of git.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n\nPosting this to hopefully move the conversation forward.  Unfortunately\nI need something like this so I can tests so I guess now is the time to\nresolve this detail.\n\n .../technical/hash-function-transition.txt    | 49 ++++++++++++++++---\n 1 file changed, 43 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\nindex 4b937480848a..9f5c672d9ad1 100644\n--- a/Documentation/technical/hash-function-transition.txt\n+++ b/Documentation/technical/hash-function-transition.txt\n@@ -149,13 +149,13 @@ Repository format extension\n ~~~~~~~~~~~~~~~~~~~~~~~~~~~\n A SHA-256 repository uses repository format version `1` (see\n Documentation/technical/repository-version.txt) with extensions\n-`objectFormat` and `compatObjectFormat`:\n+`objectFormat` and `mapObjectFormat`:\n \n \t[core]\n \t\trepositoryFormatVersion = 1\n \t[extensions]\n \t\tobjectFormat = sha256\n-\t\tcompatObjectFormat = sha1\n+\t\tmapObjectFormat = sha1\n \n The combination of setting `core.repositoryFormatVersion=1` and\n populating `extensions.*` ensures that all versions of Git later than\n@@ -171,6 +171,43 @@ repository, instead producing an error message.\n \t\tobjectformat\n \t\tcompatobjectformat\n \n+Configurate for a future hash function transition would be:\n+\n+\t[core]\n+\t\trepositoryFormatVersion = 1\n+\t[extensions]\n+\t\tobjectFormat = futureHash\n+\t\tmapObjectFormat = sha256\n+\t\tmapObjectFormat = sha1\n+\n+Or possibly:\n+\n+\t[core]\n+\t\trepositoryFormatVersion = 1\n+\t\treadObjectMap = sha1\n+\t[extensions]\n+\t\tobjectFormat = futureHash\n+\t\tmapObjectFormat = sha256\n+\n+Or post transition to futureHash:\n+\n+\t[core]\n+\t\trepositoryFormatVersion = 1\n+\t\treadObjectMap = sha1\n+\t\treadObjectMap = sha256\n+\t[extensions]\n+\t\tobjectFormat = futureHash\n+\n+The difference between mapObjectFormat and readObjectMap would be that\n+mapObjectFormat would ask git to read existing maps, but would not ask\n+git to write or create them.  Which is enough to support looking up\n+old oids post transition, when they are only needed to support\n+references in commit logs, bug trackers, emails and the like.\n+\n+Meanwhile with mapObjectFormat set every object in the entire\n+repository would be required to have a bi-directional mapping from the\n+the mapped object format to the repositories storage hash function.\n+\n See the \"Transition plan\" section below for more details on these\n repository extensions.\n \n@@ -682,7 +719,7 @@ Some initial steps can be implemented independently of one another:\n - adding support for the PSRC field and safer object pruning\n \n The first user-visible change is the introduction of the objectFormat\n-extension (without compatObjectFormat). This requires:\n+extension. This requires:\n \n - teaching fsck about this mode of operation\n - using the hash function API (vtable) when computing object names\n@@ -690,7 +727,7 @@ extension (without compatObjectFormat). This requires:\n - rejecting attempts to fetch from or push to an incompatible\n   repository\n \n-Next comes introduction of compatObjectFormat:\n+Next comes introduction of mapObjectFormat:\n \n - implementing the loose-object-idx\n - translating object names between object formats\n@@ -724,9 +761,9 @@ Over time projects would encourage their users to adopt the \"early\n transition\" and then \"late transition\" modes to take advantage of the\n new, more futureproof SHA-256 object names.\n \n-When objectFormat and compatObjectFormat are both set, commands\n+When objectFormat and mapObjectFormat are both set, commands\n generating signatures would generate both SHA-1 and SHA-256 signatures\n by default to support both new and old users.\n \n In projects using SHA-256 heavily, users could be encouraged to adopt\n the \"post-transition\" mode to avoid accidentally making implicit use\n-- \n2.41.0\n\n"},{"id":"481721","messageId":"ZP+tTFK3Ly4sqlsq@tapette.crustytoothpaste.net","threadId":"60212","inReplyTo":"20230908231049.2035003-1-ebiederm@xmission.com","subject":"Re: [PATCH 01/32] doc hash-file-transition: A map file for mapping between sha1 and sha256","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-09-12T00:14:04Z","receivedAt":"2023-09-12T00:45:17Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-09-08 at 23:10:18, Eric W. Biederman wrote:\n> The v3 pack index file as documented has a lot of complexity making it\n> difficult to implement correctly.  I worked with bryan's preliminary\n> implementation and it took several passes to get the bugs out.\n> \n> The complexity also requires multiple table look-ups to find all of\n> the information that is needed to translate from one kind of oid to\n> another.  Which can't be good for cache locality.\n> \n> Even worse coming up with a new index file version requires making\n> changes that have the potentialy to break anything that uses the index\n> of a pack file.\n> \n> Instead of continuing to deal with the chance of braking things\n> besides the oid mapping functionality, the additional complexity in\n> the file format, and worry if the performance would be reasonable I\n> stripped down the problem to it's fundamental complexity and came up\n> with a file format that is exactly about mapping one kind of oid to\n> another, and only supports two kinds of oids.\n> \n> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n> ---\n>  .../technical/hash-function-transition.txt    | 40 +++++++++++++++++++\n>  1 file changed, 40 insertions(+)\n> \n> diff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\n> index ed574810891c..4b937480848a 100644\n> --- a/Documentation/technical/hash-function-transition.txt\n> +++ b/Documentation/technical/hash-function-transition.txt\n> @@ -209,6 +209,46 @@ format described in linkgit:gitformat-pack[5], just like\n>  today. The content that is compressed and stored uses SHA-256 content\n>  instead of SHA-1 content.\n>  \n> +Per Pack Mapping Table\n> +~~~~~~~~~~~~~~~~~~~~~~\n> +A pack compat map file (.compat) files have the following format:\n> +\n> +HEADER:\n> +\t4-byte signature:\n> +\t    The signature is: {'C', 'M', 'A', 'P'}\n> +\t1-byte version number:\n> +\t    Git only writes or recognizes version 1.\n> +\t1-byte First Object Id Version\n> +\t    We infer the length of object IDs (OIDs) from this value:\n> +\t\t1 => SHA-1\n> +\t\t2 => SHA-256\n\nOne thing I forgot to mention here, is that we have 32-bit format IDs\nfor these in the structure, so we should use them here and below.  These\nare GIT_SHA1_FORMAT_ID and GIT_SHA256_FORMAT_ID.\n\nNot that I would encourage distributing such software, but it makes it\nmuch easier for people to experiment with additional hash algorithms (in\nterms of performance, etc.) if we make the space a little sparser.\n\n> +\t1-byte Second Object Id Version\n> +\t    We infer the length of object IDs (OIDs) from this value:\n> +\t\t1 => SHA-1\n> +\t\t2 => SHA-256\n\nIn your new patch for the next part, you consider that there might be\nmultiple compatibility hash algorithms.  I had anticipated only one at\na time in my series, but I'm not opposed to multiple if you want to\nsupport that.\n\nHowever, here you're making the assumption that there are only two.  If\nyou want to support multiple values, we need to explicitly consider that\nboth here (where we need a count of object ID version and multiple\ntables, one for each algorithm), and in the follow-up series.\n\nI had not considered more than two algorithms because it substantially\ncomplicates the code and requires us to develop n*(n-1) tables, but I'm\nnot the one volunteering to do most of the work here, so I'll defer to\nyour preference.  (I do intend to send a patch or two, though.)\n\nIt's also possible we could be somewhat provident and define the on-disk\nformats for multiple algorithms and then punt on the code until later if\nyou prefer that.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"481729","messageId":"87ledcb7ec.fsf_-_@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"87sf7kd5xg.fsf_-_@email.froward.int.ebiederm.org","subject":"[PATCH v3 02/32] doc hash-function-transition: Augment compatObjectFormat with readCompatMap","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-11T23:46:19Z","receivedAt":"2023-09-12T01:42:37Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\nDeeply and fundamentally the plan is to only operate one one hash\nfunction for the core of git, to use only one hash function for what\nis stored in the repository.\n\nTo avoid requring a flag day to transition hash functions for naming\nobjects, and to support being able to access objects using legacy object\nnames a mapping functionality will be provided.\n\nWe want to provide user facing configuration that is robust enough\nthat it can accomodate multiple different scenarios on how git\nevolves and how people use their repositories.\n\nThere are two different ways it is envisioned to use mapped object\nids.  The first is to require every object in the repository to have a\nmapping, so that pushes and pulls from repositories using a different\nhash algorithm can work.  The second is to have an incomplete mapping\nof object ids so that old references to objects in emails, commit\nmessages, bug trackers and are usable in a read-only manner\nwith tools like \"git show\".\n\nThe first way fundamentally needs every object in the repository to\nhave a mapping, which requires the repository to be marked incompatible\nfor writes fron older versions of git.  Thus the compatObjectFormat option\nis placed in [extensions].\n\nThe ext2 family of filesystems has 3 ways of describing new features\ncompatible, read-only-compatible, and incompatible.  The current git\nconfigurtation has compat (any feature mentioned anywhere in the\nconfiguration outside of [extensions] section), and incompatible (any\nconfiguration inside of the [extensions] section.  It would be nice to\nhave a read-only compatible section for the mandatory mapping\nfunction.  Would it be worth adding it now so that we have it for\nfuture extensions?\n\nHaving a mapping that is just used in a read-only mode for looking up\nold objects with old object ids will be needed post-transition.  Such\na mode does not require computing the old hash function or even\nsupport automatically writing any new mappings.  So it is completely\nsafe to enable in a backwards compatible mode.  Fort that let's\nuse core.readCompatMap to make it clear the mappings only read.\n\nI have documented that both of the options readCompatMap and\ncompatObjectFormat can be specified multiple times if that is needed to\nsupport the desired configuration of git.\n\nSigned-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n---\n\nMy v2 version was just silly.  Changing the name of the option in\nthe [extensions] section made practical sense.  It was just me being\ncontrary for no good reason.  I still think we should have an additional\noption for reading old hashes and to document that we expect multiple of\nthese.\n\nSo here is my proposal for extending the documentation along those\nlines.\n\nAdditionally just accepting the existing option name means I am not\nbottlenecked for writing tests convert_object_file which is the\nimportant part right now.\n\nMy apologies for all of the noise.\n\n .../technical/hash-function-transition.txt    | 37 +++++++++++++++++++\n 1 file changed, 37 insertions(+)\n\ndiff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\nindex 4b937480848a..26dfc3138b3b 100644\n--- a/Documentation/technical/hash-function-transition.txt\n+++ b/Documentation/technical/hash-function-transition.txt\n@@ -171,6 +171,43 @@ repository, instead producing an error message.\n \t\tobjectformat\n \t\tcompatobjectformat\n \n+Configurate for a future hash function transition would be:\n+\n+\t[core]\n+\t\trepositoryFormatVersion = 1\n+\t[extensions]\n+\t\tobjectFormat = futureHash\n+\t\tcompatObjectFormat = sha256\n+\t\tcompatObjectFormat = sha1\n+\n+Or possibly:\n+\n+\t[core]\n+\t\trepositoryFormatVersion = 1\n+\t\treadCompatMap = sha1\n+\t[extensions]\n+\t\tobjectFormat = futureHash\n+\t\tcompatObjectFormat = sha256\n+\n+Or post transition to futureHash:\n+\n+\t[core]\n+\t\trepositoryFormatVersion = 1\n+\t\treadCompatMap = sha1\n+\t\treadComaptMap = sha256\n+\t[extensions]\n+\t\tobjectFormat = futureHash\n+\n+The difference between compatObjectFormat and readCompatMap would be that\n+compatObjectFormat would ask git to read existing maps, but would not ask\n+git to write or create them.  Which is enough to support looking up\n+old oids post transition, when they are only needed to support\n+references in commit logs, bug trackers, emails and the like.\n+\n+Meanwhile with compatObjectFormat set every object in the entire\n+repository would be required to have a bi-directional mapping from the\n+the mapped object format to the repositories storage hash function.\n+\n See the \"Transition plan\" section below for more details on these\n repository extensions.\n \n-- \n2.41.0\n\n"},{"id":"481733","messageId":"ZP+PKa3N1N1PXROM@tapette.crustytoothpaste.net","threadId":"60212","inReplyTo":"87cyyoeli0.fsf@email.froward.int.ebiederm.org","subject":"Re: [RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2023-09-11T22:05:29Z","receivedAt":"2023-09-12T02:25:52Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2023-09-11 at 16:13:27, Eric W. Biederman wrote:\n> Junio, brian for the very warm reception of this, it is very\n> encouraging.\n> \n> I am not worried about what it will take time to get the changes I\n> posted into the integration.  I had only envisioned them as good enough\n> to get the technical ideas across, and had never envisioned them as\n> being accepted as is.\n> \n> What I am envisioning as my future directions are:\n> \n> - Post non controversial cleanups, so they can be merged.\n>   (I can only see about 4 of them the most significant is:\n>    bulk-checkin: Only accept blobs)\n> \n> - Sort out the configuration options\n> \n> - Post the smallest patchset I can that will allow testing the code in\n>   object-file-convert.c.  Unfortunately for that I need configuration\n>   options to enable the mapping.\n> \n>   In starting to write the tests I have already found a bug in\n>   the conversion of tags (an extra newline is added), and I haven't\n>   even gotten to testing the tricky bits with signatures.\n\nI wonder if unit tests are a possibility here now that we're starting to\nuse them.  They're not obligatory, of course, but it may be more\nconvenient for you if they turn out to be a suitable option.  If not, no\nbig deal.\n\n> - Once the object file conversion is tested and is solid work on\n>   the more substantial pieces.\n> \n> Does that sound like a reasonable plan?\n\nYeah, that seems fine.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"481746","messageId":"ZQAZ19gyvt8ab7f6@ugly","threadId":"60212","inReplyTo":"87ledcb7ec.fsf_-_@email.froward.int.ebiederm.org","subject":"Re: [PATCH v3 02/32] doc hash-function-transition: Augment compatObjectFormat with readCompatMap","fromName":"Oswald Buddenhagen","fromEmail":"oswald.buddenhagen@gmx.de","sentAt":"2023-09-12T07:57:11Z","receivedAt":"2023-09-12T07:57:26Z","isPatch":true,"sender":{"key":"oswald.buddenhagen@gmx.de","avatar":"https://avatars.githubusercontent.com/u/812380?v=4"},"body":"On Mon, Sep 11, 2023 at 06:46:19PM -0500, Eric W. Biederman wrote:\n>+The difference between compatObjectFormat and readCompatMap would be that\n>+compatObjectFormat would ask git to read existing maps, but would not ask\n>+git to write or create them.\n> \nthe argument makes sense, but the asymmetry in the naming bugs me. in \nparticular \"[read]compatMap\" seems too non-descript.\n\nregards\n"},{"id":"481758","messageId":"87msxr8uc1.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"ZQAZ19gyvt8ab7f6@ugly","subject":"Re: [PATCH v3 02/32] doc hash-function-transition: Augment compatObjectFormat with readCompatMap","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-12T12:11:26Z","receivedAt":"2023-09-12T12:12:23Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Oswald Buddenhagen <oswald.buddenhagen@gmx.de> writes:\n\n> On Mon, Sep 11, 2023 at 06:46:19PM -0500, Eric W. Biederman wrote:\n>>+The difference between compatObjectFormat and readCompatMap would be that\n>>+compatObjectFormat would ask git to read existing maps, but would not ask\n>>+git to write or create them.\n>> \n> the argument makes sense, but the asymmetry in the naming bugs me. in particular\n> \"[read]compatMap\" seems too non-descript.\n\nI am open to suggestions for better names.\n\nFrom a code point of view I am intending readCompatMap only supporting\nthe things that can be support with just the mapping functions aka\nrepo_oid_to_algop for the \"readComatMap\" case.\n\nWhile the compatObjectFormat case includes what can be done with using\nthe compatible hash algorithm and convert_object_file.\n\nThere is quite a large variation.  So there is some fundamental\nasymmetry in the implementation.   I am just not certain how to name it.\n\nEric\n\n"},{"id":"481763","messageId":"87zg1r1pk9.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"ZP+tTFK3Ly4sqlsq@tapette.crustytoothpaste.net","subject":"Re: [PATCH 01/32] doc hash-file-transition: A map file for mapping between sha1 and sha256","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-12T13:36:22Z","receivedAt":"2023-09-12T13:36:35Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> On 2023-09-08 at 23:10:18, Eric W. Biederman wrote:\n>> The v3 pack index file as documented has a lot of complexity making it\n>> difficult to implement correctly.  I worked with bryan's preliminary\n>> implementation and it took several passes to get the bugs out.\n>> \n>> The complexity also requires multiple table look-ups to find all of\n>> the information that is needed to translate from one kind of oid to\n>> another.  Which can't be good for cache locality.\n>> \n>> Even worse coming up with a new index file version requires making\n>> changes that have the potentialy to break anything that uses the index\n>> of a pack file.\n>> \n>> Instead of continuing to deal with the chance of braking things\n>> besides the oid mapping functionality, the additional complexity in\n>> the file format, and worry if the performance would be reasonable I\n>> stripped down the problem to it's fundamental complexity and came up\n>> with a file format that is exactly about mapping one kind of oid to\n>> another, and only supports two kinds of oids.\n>> \n>> Signed-off-by: \"Eric W. Biederman\" <ebiederm@xmission.com>\n>> ---\n>>  .../technical/hash-function-transition.txt    | 40 +++++++++++++++++++\n>>  1 file changed, 40 insertions(+)\n>> \n>> diff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\n>> index ed574810891c..4b937480848a 100644\n>> --- a/Documentation/technical/hash-function-transition.txt\n>> +++ b/Documentation/technical/hash-function-transition.txt\n>> @@ -209,6 +209,46 @@ format described in linkgit:gitformat-pack[5], just like\n>>  today. The content that is compressed and stored uses SHA-256 content\n>>  instead of SHA-1 content.\n>>  \n>> +Per Pack Mapping Table\n>> +~~~~~~~~~~~~~~~~~~~~~~\n>> +A pack compat map file (.compat) files have the following format:\n>> +\n>> +HEADER:\n>> +\t4-byte signature:\n>> +\t    The signature is: {'C', 'M', 'A', 'P'}\n>> +\t1-byte version number:\n>> +\t    Git only writes or recognizes version 1.\n>> +\t1-byte First Object Id Version\n>> +\t    We infer the length of object IDs (OIDs) from this value:\n>> +\t\t1 => SHA-1\n>> +\t\t2 => SHA-256\n>\n> One thing I forgot to mention here, is that we have 32-bit format IDs\n> for these in the structure, so we should use them here and below.  These\n> are GIT_SHA1_FORMAT_ID and GIT_SHA256_FORMAT_ID.\n>\n> Not that I would encourage distributing such software, but it makes it\n> much easier for people to experiment with additional hash algorithms (in\n> terms of performance, etc.) if we make the space a little sparser.\n\nUnfortunately that ship has already sailed. If you look at pack reverse\nindices, pack mtime files, multi-pack-index files, they all use an\noid_version field.  So to experiment with a new hash function a new\nnumber has to be picked.\n\nThe only use I can find of your 4 byte format_id's is in the reftable\ncode.\n\nUsing a 4 byte magic number in this case also conflicts with basic\nsimplicity.  With a one byte field I can specify it easily, and read it\nback with no special tools, and understand what it means at a glance.\n\nI admit I can only understand what a oid version field means at a glance\nbecause the variation of object id's is low, but that is fundamental.\nWe require global agreement on names.  Fundamentally git can not\nsupport many object id transitions.  Names are just too expensive.\n\nWhen I come to how the map file is specified a single byte has real\nadvantages.  A single byte never needs byte swapping.  So it won't\nbe misread.  Using a single byte for each format allows me to\nkeep the header for the file at 16 bytes.  Which guarantees good\nalignment of everything in the file without having to be clever.\n\nAll of this is for a file that is strictly local and the entire function\nof the bytes is a sanity check to make certain that something weird is\nnot going on, or to assist recover if something bad happens.\n\nSo in this case I don't see any additional agility provided by longer\nnames helping.\n\n>> +\t1-byte Second Object Id Version\n>> +\t    We infer the length of object IDs (OIDs) from this value:\n>> +\t\t1 => SHA-1\n>> +\t\t2 => SHA-256\n>\n> In your new patch for the next part, you consider that there might be\n> multiple compatibility hash algorithms.  I had anticipated only one at\n> a time in my series, but I'm not opposed to multiple if you want to\n> support that.\n>\n> However, here you're making the assumption that there are only two.  If\n> you want to support multiple values, we need to explicitly consider that\n> both here (where we need a count of object ID version and multiple\n> tables, one for each algorithm), and in the follow-up series.\n>\n> I had not considered more than two algorithms because it substantially\n> complicates the code and requires us to develop n*(n-1) tables, but I'm\n> not the one volunteering to do most of the work here, so I'll defer to\n> your preference.  (I do intend to send a patch or two, though.)\n>\n> It's also possible we could be somewhat provident and define the on-disk\n> formats for multiple algorithms and then punt on the code until later if\n> you prefer that.\n\nIn the long term I anticipate people disabling compatObjectFormat and\nswitching to readCompatMap so they still have access to their old\nobjects by their original names, but they don't have the over head\nof computing a compatibility hash.\n\nIn a world where there is a transition to futureHash I anticipate\nthe files associated with an old pack looking something like:\npack-abcdefg.compat12\npack-abcdefg.compat32\n\nFor a repository still using hash version sha256 for storage, with a\nmapping to some sha1 names, and a mapping of everything to new\nnames for compatibility with futureHash.\n\nAfter transitioning to the futureHash those files would look like:\npack-abcdefg.compat13\npack-abcdefg.compat23\n\nI deeply and fundamentally care about having some way to look up\nold names because I do that all of the time.\n\nIn my work on the linux-kernel I have found myself frequently digging\ninto old issues.  I have on my hard drive tglx's git import of the\nold bitkeeper tree.  I also have an import of all of the old kernel\nreleases into git from before the code was stored in bitkeeper.  I find\nmyself actually using all of those trees when digging into issues.\n\nSo I think the idea that we will ever be able to get rid of the mapping\nfor old converted repositories is unlikely.  We have entirely too many\nreferences out there.\n\nWhich means that for every hash format conversion a repository goes\nthrough I am going to have another collection of old names.\n\n\nI don't honestly anticipate ever needing to have multiple\ncompatObjectFormat entries specified for a single repository.  I do\nagree that if we are going to worry about forward and backward\ncompatibility we should be robust and have a configuration file syntax\nthat can handle the possibility.\n\nI do very much anticipate needing to have multiple readCompatMap\nentries, and pretty much only using them in get_short_oid in\nobject-name.c.  It will make the loop in find_short_packed_compat_object\na little longer but that is about all that will need to be implemented\nand maintained long term.\n\n\nI view this compat map format a lot like the loose objects.  It is\nsimple and good enough to get us started.  If it turns out we need\nto optimize it's simplicity means all of the interfaces in the code\nto use it have already been built, and we can just concentrate on\noptimizing.\n\nEric\n\n"},{"id":"481770","messageId":"xmqqil8fqs6o.fsf@gitster.g","threadId":"60212","inReplyTo":"87cyyoeli0.fsf@email.froward.int.ebiederm.org","subject":"Re: [RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-09-12T16:20:31Z","receivedAt":"2023-09-12T16:20:41Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Eric W. Biederman\" <ebiederm@xmission.com> writes:\n\n> I am not worried about what it will take time to get the changes I\n> posted into the integration.  I had only envisioned them as good enough\n> to get the technical ideas across, and had never envisioned them as\n> being accepted as is.\n\nAh, no worries.  By \"integration\" I did not mean \"patches considered\nperfect, they are accepted, and are now part of the Git codebase\".\n\nAll that happens when the patches become part of the 'master'\nbranch, but before that, patches that prove testable and worthy of\ngetting tested will be merged to the 'next' branch and spend about a\nweek there.  What I meant to refer to is a step _before_ that, i.e.\nbefore the patches probe to be testable.  New patches first appear\non the 'seen' branch that merges \"everything else\" to see the\ninteraction with all the topics \"in flight\" (i.e.  not yet in\n'master').  The 'seen' branch is reassembled from the latest\niteration of the patches twice of thrice per day, and some patches\nare merged to 'next' and down to 'master', these \"merging to prepare\n'master', 'next' and 'seen' branches for publishing\" was what I\nmeant by \"integration\".  In short, being queued on 'seen' does not\nmean all that much.  It gives project participants an easy access to\nview how topics look in the larger picture, potentially interacting\nwith other topics in flight, but the patches in there can be\nreplaced wholesale or even dropped if they do not turn out to be\ndesirable.\n\nI resolved textual conflicts and also compiler detectable semantic\nconflicts (e.g. some in-flight topics may have added callsites to a\nfunction your topic changes the function sigunature, or vice versa)\nto the point that the result compiles while merging this topic to\n'seen', but tests are broken the big time, it seems, even though the\ntopic by itself seems to pass the tests standalone.\n\n> What I am envisioning as my future directions are:\n> ...\n> Does that sound like a reasonable plan?\n\nNice.\n"},{"id":"481793","messageId":"87v8cfxf72.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"ZP+PKa3N1N1PXROM@tapette.crustytoothpaste.net","subject":"Re: [RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-12T21:19:13Z","receivedAt":"2023-09-12T21:19:26Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> On 2023-09-11 at 16:13:27, Eric W. Biederman wrote:\n>> Junio, brian for the very warm reception of this, it is very\n>> encouraging.\n>> \n>> I am not worried about what it will take time to get the changes I\n>> posted into the integration.  I had only envisioned them as good enough\n>> to get the technical ideas across, and had never envisioned them as\n>> being accepted as is.\n>> \n>> What I am envisioning as my future directions are:\n>> \n>> - Post non controversial cleanups, so they can be merged.\n>>   (I can only see about 4 of them the most significant is:\n>>    bulk-checkin: Only accept blobs)\n>> \n>> - Sort out the configuration options\n>> \n>> - Post the smallest patchset I can that will allow testing the code in\n>>   object-file-convert.c.  Unfortunately for that I need configuration\n>>   options to enable the mapping.\n>> \n>>   In starting to write the tests I have already found a bug in\n>>   the conversion of tags (an extra newline is added), and I haven't\n>>   even gotten to testing the tricky bits with signatures.\n>\n> I wonder if unit tests are a possibility here now that we're starting to\n> use them.  They're not obligatory, of course, but it may be more\n> convenient for you if they turn out to be a suitable option.  If not, no\n> big deal.\n\nI believe you mean using test-tool and making a very narrow focused test\non just the functionality.\n\nIf the number of patches I have to go through before I can test anything\nbecomes a problem I might go there.  Unfortunately it would take some\nrefactoring to make object-file-convert independent of the object\nmapping layer, and that is extra work and is likely to introduces bugs\nas anything.\n\nI have managed to get a set of tests working.  I am just going through\nnow and plugging the holes.\n\nMy big strategy for testing convert_object_file is to build two\nrepositories one sha1 and the other sha256 both with compatibility\nsupport enabled.  I add a series of objects to those repositories and\ncompare them to ensure the objects are identical.\n\nIt is working well and is finding bugs not just in convert_object_file\nbut in code such as commit and tag that perform interesting work with\nsigned commits.\n\nI discovered I had bungled the placement of hash_object_file for\nthe compatibility hash in commit.\n\nI found that git tag did not yet support building tags with both\nhash algorithms.\n\nRight now I am looking at commits with mergetag lines.  It is not fun.\nThe mergetag instead of pointing to the tag objects they include the\nbody of the tag objects in the commit object.  So I have convert\nthe embedded tag objects from one hash function to another.\nWhich given the presence of preceding space on every line of\nthe embedded tag object makes it doubly interesting.  Perhaps\nsomeone has already written code to extract the embedded tag.\n\nOr in short everything is moving along steadily.\n\nEric\n"},{"id":"481800","messageId":"ZQFukZ4q2Ehen8Yn@ugly","threadId":"60212","inReplyTo":"87msxr8uc1.fsf@email.froward.int.ebiederm.org","subject":"Re: [PATCH v3 02/32] doc hash-function-transition: Augment compatObjectFormat with readCompatMap","fromName":"Oswald Buddenhagen","fromEmail":"oswald.buddenhagen@gmx.de","sentAt":"2023-09-13T08:10:57Z","receivedAt":"2023-09-13T08:11:20Z","isPatch":true,"sender":{"key":"oswald.buddenhagen@gmx.de","avatar":"https://avatars.githubusercontent.com/u/812380?v=4"},"body":"On Tue, Sep 12, 2023 at 07:11:26AM -0500, Eric W. Biederman wrote:\n>Oswald Buddenhagen <oswald.buddenhagen@gmx.de> writes:\n>> On Mon, Sep 11, 2023 at 06:46:19PM -0500, Eric W. Biederman wrote:\n>>>+The difference between compatObjectFormat and readCompatMap would be that\n>>>+compatObjectFormat would ask git to read existing maps, but would not ask\n>>>+git to write or create them.\n>>> \n>> the argument makes sense, but the asymmetry in the naming bugs me. in particular\n>> \"[read]compatMap\" seems too non-descript.\n>\n>I am open to suggestions for better names.\n>\nisn't readCompatObjectFormat an obvious choice?\n(and for symmetry, the other then would be writeCompatObjectFormat, i \nguess.)\n\nregards\n"},{"id":"481850","messageId":"871qf0wmsd.fsf@email.froward.int.ebiederm.org","threadId":"60212","inReplyTo":"xmqqil8fqs6o.fsf@gitster.g","subject":"Re: [RFC][PATCH 0/32] SHA256 and SHA1 interoperability","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2023-09-14T19:57:22Z","receivedAt":"2023-09-14T19:57:34Z","isPatch":true,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> \"Eric W. Biederman\" <ebiederm@xmission.com> writes:\n>\n>> I am not worried about what it will take time to get the changes I\n>> posted into the integration.  I had only envisioned them as good enough\n>> to get the technical ideas across, and had never envisioned them as\n>> being accepted as is.\n>\n> Ah, no worries.  By \"integration\" I did not mean \"patches considered\n> perfect, they are accepted, and are now part of the Git codebase\".\n>\n> All that happens when the patches become part of the 'master'\n> branch, but before that, patches that prove testable and worthy of\n> getting tested will be merged to the 'next' branch and spend about a\n> week there.  What I meant to refer to is a step _before_ that, i.e.\n> before the patches probe to be testable.  New patches first appear\n> on the 'seen' branch that merges \"everything else\" to see the\n> interaction with all the topics \"in flight\" (i.e.  not yet in\n> 'master').  The 'seen' branch is reassembled from the latest\n> iteration of the patches twice of thrice per day, and some patches\n> are merged to 'next' and down to 'master', these \"merging to prepare\n> 'master', 'next' and 'seen' branches for publishing\" was what I\n> meant by \"integration\".  In short, being queued on 'seen' does not\n> mean all that much.  It gives project participants an easy access to\n> view how topics look in the larger picture, potentially interacting\n> with other topics in flight, but the patches in there can be\n> replaced wholesale or even dropped if they do not turn out to be\n> desirable.\n>\n> I resolved textual conflicts and also compiler detectable semantic\n> conflicts (e.g. some in-flight topics may have added callsites to a\n> function your topic changes the function sigunature, or vice versa)\n> to the point that the result compiles while merging this topic to\n> 'seen', but tests are broken the big time, it seems, even though the\n> topic by itself seems to pass the tests standalone.\n\nThat the tests are broken is very unfortunate.\n\nI took at look at What's cooking in git.git and I did not see my topic\nmentioned.  So I presume I would have to perform the test merge myself\nto have a sense of what the conflicts were.\n\nIs there a time when in flight topics is low?  I had a hunch that basing\nmy work on a brand new release would achieve that but I saw a lot of\ntopics in your \"What's cooking\" email.\n\nI am just trying to figure out a good plan to deal with conflicts,\nbecause the bugs need to be hunted down.\n\nEric\n\n\n\n"},{"id":"482705","messageId":"ZR79DqE91z+6+ZSd@nand.local","threadId":"60212","inReplyTo":"xmqqfs3li5ly.fsf@gitster.g","subject":"Re: [PATCH 25/32] pack-compat-map: Add support for .compat files of a packfile","fromName":"Taylor Blau","fromEmail":"me@ttaylorr.com","sentAt":"2023-10-05T18:14:38Z","receivedAt":"2023-10-05T18:15:23Z","isPatch":true,"sender":{"key":"me@ttaylorr.com","avatar":"https://avatars.githubusercontent.com/u/301000140?v=4"},"body":"On Sun, Sep 10, 2023 at 11:30:49PM -0700, Junio C Hamano wrote:\n> \"Eric W. Biederman\" <ebiederm@xmission.com> writes:\n>\n> > diff --git a/pack-write.c b/pack-write.c\n> > index b19ddf15b284..f22eea964f77 100644\n> > --- a/pack-write.c\n> > +++ b/pack-write.c\n> > @@ -12,6 +12,7 @@\n> >  #include \"pack-revindex.h\"\n> >  #include \"path.h\"\n> >  #include \"strbuf.h\"\n> > +#include \"object-file-convert.h\"\n> > ...\n> > +/*\n> > + * The *hash contains the pack content hash.\n> > + * The objects array is passed in sorted.\n> > + */\n> > +const char *write_compat_map_file(const char *compat_map_name,\n> > +\t\t\t\t  struct pack_idx_entry **objects,\n> > +\t\t\t\t  int nr_objects, const unsigned char *hash)\n>\n> Include \"pack-compat-map.h\"; otherwise the compiler would complain\n> for missing prototypes.\n\nLikewise this is missing an entry in the .gitignore:\n\n--- >8 ---\ndiff --git a/.gitignore b/.gitignore\nindex 5e56e471b3..7f5a93a6f6 100644\n--- a/.gitignore\n+++ b/.gitignore\n@@ -152,6 +152,7 @@\n /git-shortlog\n /git-show\n /git-show-branch\n+/git-show-compat-map\n /git-show-index\n /git-show-ref\n /git-sparse-checkout\n--- 8< ---\n\nThanks,\nTaylor\n"}]}