{"thread":{"id":"66056","subject":"[PATCH 0/5] odb: make creation of object database pluggable","startedAt":"2026-07-24T03:48:57Z","lastAt":"2026-08-07T09:10:34Z","messageCount":68,"participants":["Patrick Steinhardt","Junio C Hamano","Justin Tobler","Toon Claes"],"isPatch":true,"patchVersion":1,"patchTotal":5},"messages":[{"id":"548853","messageId":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","threadId":"66056","inReplyTo":null,"subject":"[PATCH 0/5] odb: make creation of object database pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-24T03:48:39Z","receivedAt":"2026-07-24T03:48:57Z","isPatch":true,"body":"Hi,\n\nwhen creating a new repository we create a couple of on-disk data\nstructures for the object database. This includes the \"objects/\"\ndirectory hierarchy with \"objects/info\" and \"objects/pack\", which are\nspecific to the backend.\n\nThis patch series makes the creation of the on-disk data structures\npluggable. While we continue to always create \"objects/\" regardless of\nthe backend (it's required for a repository to be recognized as such),\nthe other subdirectories are now created by the backend. This will allow\nother backends to plug in their own logic.\n\nThe series starts with a small detour into the loose-object map. This\ndetour is required so that we can defer initialization of the object\ndatabase itself to a later point in time.\n\nThe series is based on 9a0c4701dc (The 7th batch, 2026-07-22).\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (5):\n      loose: load loose object map for the correct source\n      setup: detangle loading of loose object maps\n      setup: defer object database creation\n      odb/source: introduce function to map source type to name\n      odb: make creation of on-disk structures pluggable\n\n loose.c               | 25 ++++++++++----------\n loose.h               |  1 +\n odb/source-files.c    | 19 +++++++++++++++\n odb/source-files.h    |  4 +++-\n odb/source-inmemory.h |  4 +++-\n odb/source-loose.c    |  2 ++\n odb/source-loose.h    |  4 +++-\n odb/source-packed.h   |  4 +++-\n odb/source.c          | 19 +++++++++++++++\n odb/source.h          | 29 +++++++++++++++++++++++\n repository.c          |  2 --\n setup.c               | 65 +++++++++++++++++++++++++++++++++++----------------\n setup.h               |  9 +++++++\n 13 files changed, 149 insertions(+), 38 deletions(-)\n\n\n---\nbase-commit: 9a0c4701dcd5725c4184599322b52933ff5005ca\nchange-id: 20260710-pks-odb-create-on-disk-ae8757861c69\n\n"},{"id":"548854","messageId":"20260724-pks-odb-create-on-disk-v1-1-3b3d265d979b@pks.im","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","subject":"[PATCH 1/5] loose: load loose object map for the correct source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-24T03:48:40Z","receivedAt":"2026-07-24T03:49:01Z","isPatch":true,"body":"When loading the loose object map via `load_one_loose_object_map()` we\npass in both a repository and the corresponding source. We ultimately\ndon't really respect the passed-in source though as we instead always\nload the map via the common directory. This doesn't make any sense\nthough, as the function is called in a loop through all sources, and as\nsuch the expectation is that we'll load the map that belongs to the\ngiven source.\n\nFix this bug by instead loading the map via the loose source's path.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c | 18 ++++++++++--------\n 1 file changed, 10 insertions(+), 8 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex bf01d3e42d..9dad75373b 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,9 +61,11 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct repository *repo, struct odb_source_loose *loose)\n+static int load_one_loose_object_map(struct odb_source_loose *loose)\n {\n-\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\tstruct repository *repo = loose->base.odb->repo;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar *path;\n \tFILE *fp;\n \tint ret = -1;\n \n@@ -78,10 +80,10 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n \tinsert_loose_map(loose, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n \tinsert_loose_map(loose, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n \n-\trepo_common_path_replace(repo, &path, \"objects/loose-object-idx\");\n-\tfp = fopen(path.buf, \"rb\");\n+\tpath = xstrfmt(\"%s/loose-object-idx\", loose->base.path);\n+\tfp = fopen(path, \"rb\");\n \tif (!fp) {\n-\t\tstrbuf_release(&path);\n+\t\tfree(path);\n \t\treturn 0;\n \t}\n \n@@ -102,7 +104,7 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n err:\n \tfclose(fp);\n \tstrbuf_release(&buf);\n-\tstrbuf_release(&path);\n+\tfree(path);\n \treturn ret;\n }\n \n@@ -117,10 +119,10 @@ int repo_read_loose_object_map(struct repository *repo)\n \n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(repo, files->loose) < 0) {\n+\t\tif (load_one_loose_object_map(files->loose) < 0)\n \t\t\treturn -1;\n-\t\t}\n \t}\n+\n \treturn 0;\n }\n \n\n-- \n2.55.0.407.g700c83d4f3.dirty\n\n"},{"id":"548855","messageId":"20260724-pks-odb-create-on-disk-v1-2-3b3d265d979b@pks.im","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","subject":"[PATCH 2/5] setup: detangle loading of loose object maps","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-24T03:48:41Z","receivedAt":"2026-07-24T03:49:06Z","isPatch":true,"body":"When a repository is configured to use a compatibility hash function\nthen we load the loose object map when we initialize the repository.\nThis object map provides the mappings between the canonical object hash\nand the compatibility object hash.\n\nLoading the object map happens in `repo_set_compat_hash_algo()`, which\ncalls `repo_read_loose_object_map()` in case the compatibility object\nhash is non-zero. This setup sequence has two major downsides:\n\n  - We assume that the primary object database is the \"files\" object\n    database so that we can extract its \"loose\" backend. This stops\n    working with pluggable object databases.\n\n  - We require the object database to already have been initialized when\n    configuring the object database. This means that we must intermix\n    configuration of the repository and initialization of its\n    sub-structures in a weird way.\n\nRefactor the logic so that we instead load the loose object map via the\n\"loose\" backend, which fixes both of the above issues.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c            | 11 +++++------\n loose.h            |  1 +\n odb/source-loose.c |  2 ++\n repository.c       |  2 --\n setup.c            |  5 +++--\n 5 files changed, 11 insertions(+), 10 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 9dad75373b..a3b2dcedc2 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,7 +61,7 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct odb_source_loose *loose)\n+int loose_object_map_load(struct odb_source_loose *loose)\n {\n \tstruct repository *repo = loose->base.odb->repo;\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -69,6 +69,9 @@ static int load_one_loose_object_map(struct odb_source_loose *loose)\n \tFILE *fp;\n \tint ret = -1;\n \n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n \tif (!loose->map)\n \t\tloose_object_map_init(&loose->map);\n \tif (!loose->cache) {\n@@ -112,14 +115,10 @@ int repo_read_loose_object_map(struct repository *repo)\n {\n \tstruct odb_source *source;\n \n-\tif (!should_use_loose_object_map(repo))\n-\t\treturn 0;\n-\n \todb_prepare_alternates(repo->objects);\n-\n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(files->loose) < 0)\n+\t\tif (loose_object_map_load(files->loose) < 0)\n \t\t\treturn -1;\n \t}\n \ndiff --git a/loose.h b/loose.h\nindex 6c9b3f4571..ed663ac550 100644\n--- a/loose.h\n+++ b/loose.h\n@@ -13,6 +13,7 @@ struct loose_object_map {\n \n void loose_object_map_init(struct loose_object_map **map);\n void loose_object_map_clear(struct loose_object_map **map);\n+int loose_object_map_load(struct odb_source_loose *loose);\n int repo_loose_object_map_oid(struct repository *repo,\n \t\t\t      const struct object_id *src,\n \t\t\t      const struct git_hash_algo *dest_algo,\ndiff --git a/odb/source-loose.c b/odb/source-loose.c\nindex 3f7d04a56e..812ca1c138 100644\n--- a/odb/source-loose.c\n+++ b/odb/source-loose.c\n@@ -727,5 +727,7 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n \tif (!is_absolute_path(loose->base.path))\n \t\tchdir_notify_register(NULL, odb_source_loose_reparent, loose);\n \n+\tloose_object_map_load(loose);\n+\n \treturn loose;\n }\ndiff --git a/repository.c b/repository.c\nindex 2ef0778846..6d633002b4 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -201,8 +201,6 @@ void repo_set_compat_hash_algo(struct repository *repo MAYBE_UNUSED, uint32_t al\n \tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n \t\tBUG(\"hash_algo and compat_hash_algo match\");\n \trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n-\tif (repo->compat_hash_algo)\n-\t\trepo_read_loose_object_map(repo);\n #else\n \tif (algo)\n \t\tdie(_(\"compatibility hash algorithm support requires Rust\"));\ndiff --git a/setup.c b/setup.c\nindex d31808130b..825572f5f1 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1788,8 +1788,6 @@ int apply_repository_format(struct repository *repo,\n \n \trepo->bare_cfg = format->is_bare;\n \trepo_set_hash_algo(repo, format->hash_algo);\n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n \trepo_set_compat_hash_algo(repo, format->compat_hash_algo);\n \trepo_set_ref_storage_format(repo,\n \t\t\t\t    format->ref_storage_format,\n@@ -1805,6 +1803,9 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n+\n \tfree(alternate_object_directories);\n \tfree(object_directory);\n \treturn 0;\n\n-- \n2.55.0.407.g700c83d4f3.dirty\n\n"},{"id":"548856","messageId":"20260724-pks-odb-create-on-disk-v1-3-3b3d265d979b@pks.im","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","subject":"[PATCH 3/5] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-24T03:48:42Z","receivedAt":"2026-07-24T03:49:10Z","isPatch":true,"body":"In a subsequent commit we'll make the creation of the on-disk data\nstructures of an object database pluggable. This will lead to an\nin-between state where we have already configured the repository's\nobject database, but it's not usable yet until we eventually call\n`create_object_directory()`.\n\nDefer the object database creation so that we handle both steps in the\nsame function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n setup.c | 35 +++++++++++++++++++++++++++--------\n setup.h |  9 +++++++++\n 2 files changed, 36 insertions(+), 8 deletions(-)\n\ndiff --git a/setup.c b/setup.c\nindex 825572f5f1..a7b1b9eaef 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1760,6 +1760,13 @@ enum discovery_result discover_git_directory_reason(struct strbuf *commondir,\n \treturn result;\n }\n \n+static void get_object_directories(char **object_directory,\n+\t\t\t\t   char **alternate_object_directories)\n+{\n+\t*object_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n+\t*alternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n+}\n+\n int apply_repository_format(struct repository *repo,\n \t\t\t    const struct repository_format *format,\n \t\t\t    enum apply_repository_format_flags flags,\n@@ -1779,8 +1786,9 @@ int apply_repository_format(struct repository *repo,\n \tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV) {\n \t\tconst char *shallow_file;\n \n-\t\tobject_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n-\t\talternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n+\t\tget_object_directories(&object_directory,\n+\t\t\t\t       &alternate_object_directories);\n+\n \t\tshallow_file = getenv(GIT_SHALLOW_FILE_ENVIRONMENT);\n \t\tif (shallow_file)\n \t\t\tset_alternate_shallow_file(repo, shallow_file);\n@@ -1803,8 +1811,9 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n+\tif (!(flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION))\n+\t\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\t\talternate_object_directories);\n \n \tfree(alternate_object_directories);\n \tfree(object_directory);\n@@ -2654,11 +2663,16 @@ static int create_default_files(struct repository *repo,\n \treturn reinit;\n }\n \n-static void create_object_directory(struct repository *repo)\n+static void create_object_database(struct repository *repo)\n {\n+\tchar *object_directory, *alternate_object_directories;\n \tstruct strbuf path = STRBUF_INIT;\n \tsize_t baselen;\n \n+\tget_object_directories(&object_directory, &alternate_object_directories);\n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n+\n \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n \tbaselen = path.len;\n \n@@ -2672,6 +2686,8 @@ static void create_object_directory(struct repository *repo)\n \tstrbuf_addstr(&path, \"/info\");\n \tsafe_create_dir(repo, path.buf, 1);\n \n+\tfree(alternate_object_directories);\n+\tfree(object_directory);\n \tstrbuf_release(&path);\n }\n \n@@ -2867,9 +2883,10 @@ int init_db(struct repository *repo,\n \t */\n \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n-\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n+\tif (apply_repository_format(repo, &repo_fmt,\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV |\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION, &err) < 0)\n \t\tdie(\"%s\", err.buf);\n-\tstartup_info->have_repository = 1;\n \n \t/*\n \t * Ensure `core.hidedotfiles` is processed. This must happen after we\n@@ -2885,7 +2902,9 @@ int init_db(struct repository *repo,\n \n \tif (!(flags & INIT_DB_SKIP_REFDB))\n \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n-\tcreate_object_directory(repo);\n+\tcreate_object_database(repo);\n+\n+\tstartup_info->have_repository = 1;\n \n \tif (repo_settings_get_shared_repository(repo)) {\n \t\tchar buf[10];\ndiff --git a/setup.h b/setup.h\nindex 654f10e059..e55d647b70 100644\n--- a/setup.h\n+++ b/setup.h\n@@ -241,6 +241,15 @@ enum apply_repository_format_flags {\n \t * relate to the object database.\n \t */\n \tAPPLY_REPOSITORY_FORMAT_HONOR_ENV = (1 << 0),\n+\n+\t/*\n+\t * Usually, the object database is created after the repository format\n+\t * was applied. This step is skipped if this flag is set, which leaves\n+\t * us with a partially-working repository.\n+\t *\n+\t * This is useful when initializing a new repository.\n+\t */\n+\tAPPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION = (1 << 1),\n };\n \n /*\n\n-- \n2.55.0.407.g700c83d4f3.dirty\n\n"},{"id":"548857","messageId":"20260724-pks-odb-create-on-disk-v1-4-3b3d265d979b@pks.im","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","subject":"[PATCH 4/5] odb/source: introduce function to map source type to name","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-24T03:48:43Z","receivedAt":"2026-07-24T03:49:15Z","isPatch":true,"body":"Introduce a new function that maps an object source's type to a\nhuman-readable name. Use the function to provide better human-readable\nerror messages for the downcasting functions.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.h    |  4 +++-\n odb/source-inmemory.h |  4 +++-\n odb/source-loose.h    |  4 +++-\n odb/source-packed.h   |  4 +++-\n odb/source.c          | 19 +++++++++++++++++++\n odb/source.h          |  6 ++++++\n 6 files changed, 37 insertions(+), 4 deletions(-)\n\ndiff --git a/odb/source-files.h b/odb/source-files.h\nindex d7ac3c1c81..6a803afdda 100644\n--- a/odb/source-files.h\n+++ b/odb/source-files.h\n@@ -28,7 +28,9 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n static inline struct odb_source_files *odb_source_files_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_FILES)\n-\t\tBUG(\"trying to downcast source of type '%d' to files\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_FILES));\n \treturn container_of(source, struct odb_source_files, base);\n }\n \ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex a88fc2e320..adbad23e8b 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -26,7 +26,9 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_INMEMORY)\n-\t\tBUG(\"trying to downcast source of type '%d' to in-memory\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_INMEMORY));\n \treturn container_of(source, struct odb_source_inmemory, base);\n }\n \ndiff --git a/odb/source-loose.h b/odb/source-loose.h\nindex 6070aaf3ce..3cf2e1f8f1 100644\n--- a/odb/source-loose.h\n+++ b/odb/source-loose.h\n@@ -41,7 +41,9 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n static inline struct odb_source_loose *odb_source_loose_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_LOOSE)\n-\t\tBUG(\"trying to downcast source of type '%d' to loose\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_LOOSE));\n \treturn container_of(source, struct odb_source_loose, base);\n }\n \ndiff --git a/odb/source-packed.h b/odb/source-packed.h\nindex 77309ddd09..a0f6b5096d 100644\n--- a/odb/source-packed.h\n+++ b/odb/source-packed.h\n@@ -78,7 +78,9 @@ struct odb_source_packed *odb_source_packed_new(struct object_database *odb,\n static inline struct odb_source_packed *odb_source_packed_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_PACKED)\n-\t\tBUG(\"trying to downcast source of type '%d' to packed\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_PACKED));\n \treturn container_of(source, struct odb_source_packed, base);\n }\n \ndiff --git a/odb/source.c b/odb/source.c\nindex 7993dcbd65..c300e836f6 100644\n--- a/odb/source.c\n+++ b/odb/source.c\n@@ -4,6 +4,25 @@\n #include \"odb/source.h\"\n #include \"packfile.h\"\n \n+static const char * const odb_source_names_by_type[] = {\n+\t[ODB_SOURCE_UNKNOWN] = \"unknown\",\n+\t[ODB_SOURCE_FILES] = \"files\",\n+\t[ODB_SOURCE_LOOSE] = \"loose\",\n+\t[ODB_SOURCE_PACKED] = \"packed\",\n+\t[ODB_SOURCE_INMEMORY] = \"inmemory\",\n+};\n+\n+const char *odb_source_type_to_name(enum odb_source_type type)\n+{\n+\tconst char *name;\n+\tif (type < 0 || type >= ARRAY_SIZE(odb_source_names_by_type))\n+\t\ttype = ODB_SOURCE_UNKNOWN;\n+\tname = odb_source_names_by_type[type];\n+\tif (!name)\n+\t\tBUG(\"name missing in `odb_source_names_by_type` for '%d'\", type);\n+\treturn name;\n+}\n+\n struct odb_source *odb_source_new(struct object_database *odb,\n \t\t\t\t  const char *path,\n \t\t\t\t  bool local)\ndiff --git a/odb/source.h b/odb/source.h\nindex cd63dba91f..ab16d152f4 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -25,6 +25,12 @@ enum odb_source_type {\n \tODB_SOURCE_INMEMORY,\n };\n \n+/*\n+ * Convert between the enum and its name. Returns the equivalent of \"unknown\"\n+ * for unknown types.\n+ */\n+const char *odb_source_type_to_name(enum odb_source_type type);\n+\n struct object_id;\n struct odb_read_stream;\n struct strvec;\n\n-- \n2.55.0.407.g700c83d4f3.dirty\n\n"},{"id":"548858","messageId":"20260724-pks-odb-create-on-disk-v1-5-3b3d265d979b@pks.im","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","subject":"[PATCH 5/5] odb: make creation of on-disk structures pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-24T03:48:44Z","receivedAt":"2026-07-24T03:49:19Z","isPatch":true,"body":"When creating a new \"files\" object database source we have to create a\ncouple of directories. These directories are of course specific to this\nparticular backend, and a different backend may require a setup that is\ncompletely different.\n\nMake the creation of on-disk structures pluggable to accommodate for\nthis.\n\nNote that there is one exception though: the \"objects\" directory must\nexist in a repository regardless of which backend is in use. If it\ndoesn't exist then the repository is not treated as a Git repository at\nall. Consequently, we create this directory regardless of the backend.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.c | 19 +++++++++++++++++++\n odb/source.h       | 23 +++++++++++++++++++++++\n setup.c            | 35 ++++++++++++++++++++---------------\n 3 files changed, 62 insertions(+), 15 deletions(-)\n\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex 4138758511..0db6e681fe 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -9,6 +9,7 @@\n #include \"odb/source-files.h\"\n #include \"odb/source-loose.h\"\n #include \"packfile.h\"\n+#include \"path.h\"\n #include \"strbuf.h\"\n #include \"write-or-die.h\"\n \n@@ -41,6 +42,23 @@ static void odb_source_files_close(struct odb_source *source)\n \todb_source_close(&files->packed->base);\n }\n \n+static int odb_source_files_create_on_disk(struct odb_source *source)\n+{\n+\tstruct strbuf path = STRBUF_INIT;\n+\n+\tsafe_create_dir(source->odb->repo, source->path, 1);\n+\n+\tstrbuf_addf(&path, \"%s/pack\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_reset(&path);\n+\tstrbuf_addf(&path, \"%s/info\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_release(&path);\n+\treturn 0;\n+}\n+\n static void odb_source_files_prepare(struct odb_source *source,\n \t\t\t\t     enum odb_prepare_flags flags)\n {\n@@ -271,6 +289,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \n \tfiles->base.free = odb_source_files_free;\n \tfiles->base.close = odb_source_files_close;\n+\tfiles->base.create_on_disk = odb_source_files_create_on_disk;\n \tfiles->base.prepare = odb_source_files_prepare;\n \tfiles->base.read_object_info = odb_source_files_read_object_info;\n \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\ndiff --git a/odb/source.h b/odb/source.h\nindex ab16d152f4..4abc418bdd 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -89,6 +89,18 @@ struct odb_source {\n \t */\n \tvoid (*close)(struct odb_source *source);\n \n+\t/*\n+\t * This callback is expected to create on-disk data structures that are\n+\t * required for this source to operate.\n+\t *\n+\t * The callback is expected to return 0 on success, a negative error\n+\t * code otherwise.\n+\t *\n+\t * This callback may be NULL in case the source does not need any\n+\t * on-disk setup.\n+\t */\n+\tint (*create_on_disk)(struct odb_source *source);\n+\n \t/*\n \t * This callback is expected to prepare the source so that it becomes\n \t * ready for use. It optionally clears underlying caches of the object\n@@ -316,6 +328,17 @@ static inline void odb_source_close(struct odb_source *source)\n \tsource->close(source);\n }\n \n+/*\n+ * Create on-disk data structures that are required for this source to operate\n+ * correctly. Returns 0 on success, a negative error code otherwise.\n+ */\n+static inline int odb_source_create_on_disk(struct odb_source *source)\n+{\n+\tif (!source->create_on_disk)\n+\t\treturn 0;\n+\treturn source->create_on_disk(source);\n+}\n+\n /*\n  * Prepare the object database source and clear any caches. Depending on the\n  * backend used this may have the effect that concurrently-written objects\ndiff --git a/setup.c b/setup.c\nindex a7b1b9eaef..14ef119cb7 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -2666,29 +2666,34 @@ static int create_default_files(struct repository *repo,\n static void create_object_database(struct repository *repo)\n {\n \tchar *object_directory, *alternate_object_directories;\n-\tstruct strbuf path = STRBUF_INIT;\n-\tsize_t baselen;\n \n \tget_object_directories(&object_directory, &alternate_object_directories);\n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n \n-\tstrbuf_addstr(&path, repo_get_object_directory(repo));\n-\tbaselen = path.len;\n-\n-\tsafe_create_dir(repo, path.buf, 1);\n+\t/*\n+\t * Create the \"objects\" directory in the common directory. This is done\n+\t * so that the repository can be discovered regardless of the backend\n+\t * used.\n+\t *\n+\t * Note that we only do this in case the object directory wasn't\n+\t * overwritten via an environment variable. If it _is_ being overridden\n+\t * then we skip this step, as the repository won't be discoverable\n+\t * anyway without the environment variable.\n+\t */\n+\tif (!object_directory) {\n+\t\tstruct strbuf objects_dir = STRBUF_INIT;\n+\t\trepo_common_path_append(repo, &objects_dir, \"objects\");\n+\t\tsafe_create_dir(repo, objects_dir.buf, 1);\n+\t\tstrbuf_release(&objects_dir);\n+\t}\n \n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/pack\");\n-\tsafe_create_dir(repo, path.buf, 1);\n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n \n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/info\");\n-\tsafe_create_dir(repo, path.buf, 1);\n+\tif (odb_source_create_on_disk(repo->objects->sources) < 0)\n+\t\tdie(\"failed creating object database\");\n \n \tfree(alternate_object_directories);\n \tfree(object_directory);\n-\tstrbuf_release(&path);\n }\n \n static void separate_git_dir(const char *git_dir, const char *git_link)\n\n-- \n2.55.0.407.g700c83d4f3.dirty\n\n"},{"id":"548901","messageId":"xmqq5x248fk9.fsf@gitster.g","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-1-3b3d265d979b@pks.im","subject":"Re: [PATCH 1/5] loose: load loose object map for the correct source","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-24T17:26:14Z","receivedAt":"2026-07-24T17:26:20Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> When loading the loose object map via `load_one_loose_object_map()` we\n> pass in both a repository and the corresponding source. We ultimately\n> don't really respect the passed-in source though as we instead always\n> load the map via the common directory. This doesn't make any sense\n> though, as the function is called in a loop through all sources, and as\n> such the expectation is that we'll load the map that belongs to the\n> given source.\n>\n> Fix this bug by instead loading the map via the loose source's path.\n\nMakes perfect sense.  We still need access to the 'repo' to learn\nthe hash algorithm used in the repository along with built-in object\nnames, but they are now obtained from the repository associated with\nthe loose object source, which is far more consistent.\n\n\n"},{"id":"548905","messageId":"xmqqh5lo6xi2.fsf@gitster.g","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-2-3b3d265d979b@pks.im","subject":"Re: [PATCH 2/5] setup: detangle loading of loose object maps","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-24T18:41:41Z","receivedAt":"2026-07-24T18:41:45Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> When a repository is configured to use a compatibility hash function\n> then we load the loose object map when we initialize the repository.\n> This object map provides the mappings between the canonical object hash\n> and the compatibility object hash.\n>\n> Loading the object map happens in `repo_set_compat_hash_algo()`, which\n> calls `repo_read_loose_object_map()` in case the compatibility object\n> hash is non-zero. This setup sequence has two major downsides:\n>\n>   - We assume that the primary object database is the \"files\" object\n>     database so that we can extract its \"loose\" backend. This stops\n>     working with pluggable object databases.\n\nI am not sure if I understand this sentence, especially \"we can\nextract its loose backend\" part.  Do you mean 'extract the object\nmap from the loose backend'?  Or something else?\n\n>   - We require the object database to already have been initialized when\n>     configuring the object database. This means that we must intermix\n>     configuration of the repository and initialization of its\n>     sub-structures in a weird way.\n>\n> Refactor the logic so that we instead load the loose object map via the\n> \"loose\" backend, which fixes both of the above issues.\n\nIt does make sense to have loose_object_map_load() that is very much\nspecific to the loose object odb source to odb_source_loose_new().\nThat way set_compat_hash_algo() does not have to assume that files\nbackend is used as the object store.\n\n> @@ -112,14 +115,10 @@ int repo_read_loose_object_map(struct repository *repo)\n>  {\n>  \tstruct odb_source *source;\n>  \n> -\tif (!should_use_loose_object_map(repo))\n> -\t\treturn 0;\n> -\n>  \todb_prepare_alternates(repo->objects);\n> -\n>  \tfor (source = repo->objects->sources; source; source = source->next) {\n>  \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n> -\t\tif (load_one_loose_object_map(files->loose) < 0)\n> +\t\tif (loose_object_map_load(files->loose) < 0)\n>  \t\t\treturn -1;\n\nIf this particular source in the list of sources is not backed by\nthe files backend, would downcast signal the fact (e.g., by\nreturning NULL) so that we can skip the next call instead?\n\nOr would the next step in refactoring be to define \"load object map\"\nmethod that is generic to odb_source so that this part does not have\nto do any of these and instead simply do\n\n\tfor (source = ...) {\n\t\tif (odb_source_object_map_load(source))\n                \treturn -1;\n        }\n\nor something?\n\n"},{"id":"548906","messageId":"xmqq5x246x35.fsf@gitster.g","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-3-3b3d265d979b@pks.im","subject":"Re: [PATCH 3/5] setup: defer object database creation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-24T18:50:38Z","receivedAt":"2026-07-24T18:50:40Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> In a subsequent commit we'll make the creation of the on-disk data\n> structures of an object database pluggable. This will lead to an\n> in-between state where we have already configured the repository's\n> object database, but it's not usable yet until we eventually call\n> `create_object_directory()`.\n>\n> Defer the object database creation so that we handle both steps in the\n> same function.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  setup.c | 35 +++++++++++++++++++++++++++--------\n>  setup.h |  9 +++++++++\n>  2 files changed, 36 insertions(+), 8 deletions(-)\n>\n> diff --git a/setup.c b/setup.c\n> index 825572f5f1..a7b1b9eaef 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -1760,6 +1760,13 @@ enum discovery_result discover_git_directory_reason(struct strbuf *commondir,\n>  \treturn result;\n>  }\n>  \n> +static void get_object_directories(char **object_directory,\n> +\t\t\t\t   char **alternate_object_directories)\n> +{\n> +\t*object_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n> +\t*alternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n> +}\n> +\n>  int apply_repository_format(struct repository *repo,\n>  \t\t\t    const struct repository_format *format,\n>  \t\t\t    enum apply_repository_format_flags flags,\n> @@ -1779,8 +1786,9 @@ int apply_repository_format(struct repository *repo,\n>  \tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV) {\n>  \t\tconst char *shallow_file;\n>  \n> -\t\tobject_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n> -\t\talternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n> +\t\tget_object_directories(&object_directory,\n> +\t\t\t\t       &alternate_object_directories);\n> +\n>  \t\tshallow_file = getenv(GIT_SHALLOW_FILE_ENVIRONMENT);\n>  \t\tif (shallow_file)\n>  \t\t\tset_alternate_shallow_file(repo, shallow_file);\n\nHONOR_ENV still means we read the environment variable to learn where\nthe object directory (which is admittedly a files backend specific\nconcept) and alternate object directories (ditto) are.\n\n> @@ -1803,8 +1811,9 @@ int apply_repository_format(struct repository *repo,\n>  \trepo->repository_format_precious_objects =\n>  \t\tformat->precious_objects;\n>  \n> -\trepo->objects = odb_new(repo, object_directory,\n> -\t\t\t\talternate_object_directories);\n> +\tif (!(flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION))\n> +\t\trepo->objects = odb_new(repo, object_directory,\n> +\t\t\t\t\talternate_object_directories);\n\nAnd SKIP_ODB_CREATION can tell apply_repository_format() not to\ncreate an odb there.\n\n> -static void create_object_directory(struct repository *repo)\n> +static void create_object_database(struct repository *repo)\n>  {\n> +\tchar *object_directory, *alternate_object_directories;\n>  \tstruct strbuf path = STRBUF_INIT;\n>  \tsize_t baselen;\n>  \n> +\tget_object_directories(&object_directory, &alternate_object_directories);\n> +\trepo->objects = odb_new(repo, object_directory,\n> +\t\t\t\talternate_object_directories);\n> +\n>  \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n>  \tbaselen = path.len;\n>  \n> @@ -2672,6 +2686,8 @@ static void create_object_directory(struct repository *repo)\n>  \tstrbuf_addstr(&path, \"/info\");\n>  \tsafe_create_dir(repo, path.buf, 1);\n>  \n> +\tfree(alternate_object_directories);\n> +\tfree(object_directory);\n>  \tstrbuf_release(&path);\n>  }\n>  \n\n\n> @@ -2867,9 +2883,10 @@ int init_db(struct repository *repo,\n>  \t */\n>  \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n>  \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n> -\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n> +\tif (apply_repository_format(repo, &repo_fmt,\n> +\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV |\n> +\t\t\t\t    APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION, &err) < 0)\n>  \t\tdie(\"%s\", err.buf);\n> -\tstartup_info->have_repository = 1;\n\nEarly in initialization, we no longer recreate the ODB when calling\napply_repository_format(), and we defer declaring that we have a\nrepository until we call create_object_database().\n\n> @@ -2885,7 +2902,9 @@ int init_db(struct repository *repo,\n>  \n>  \tif (!(flags & INIT_DB_SKIP_REFDB))\n>  \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n> -\tcreate_object_directory(repo);\n> +\tcreate_object_database(repo);\n> +\n> +\tstartup_info->have_repository = 1;\n\nInstead we call create_object_database() rather late, after we\nfinish creating leading directories and default files and processing\nthe configuration.  I guess this is a prelude to specifying \"no, we\nare not doing the files backend but are using this new thing\" in the\nglobal configuration?\n\n> diff --git a/setup.h b/setup.h\n> index 654f10e059..e55d647b70 100644\n> --- a/setup.h\n> +++ b/setup.h\n> @@ -241,6 +241,15 @@ enum apply_repository_format_flags {\n>  \t * relate to the object database.\n>  \t */\n>  \tAPPLY_REPOSITORY_FORMAT_HONOR_ENV = (1 << 0),\n> +\n> +\t/*\n> +\t * Usually, the object database is created after the repository format\n> +\t * was applied. This step is skipped if this flag is set, which leaves\n> +\t * us with a partially-working repository.\n> +\t *\n> +\t * This is useful when initializing a new repository.\n> +\t */\n> +\tAPPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION = (1 << 1),\n>  };\n\nOK.\n"},{"id":"549045","messageId":"xmqqfr15v6ba.fsf@gitster.g","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-4-3b3d265d979b@pks.im","subject":"Re: [PATCH 4/5] odb/source: introduce function to map source type to name","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-26T20:34:17Z","receivedAt":"2026-07-26T20:34:19Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Introduce a new function that maps an object source's type to a\n> human-readable name. Use the function to provide better human-readable\n> error messages for the downcasting functions.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb/source-files.h    |  4 +++-\n>  odb/source-inmemory.h |  4 +++-\n>  odb/source-loose.h    |  4 +++-\n>  odb/source-packed.h   |  4 +++-\n>  odb/source.c          | 19 +++++++++++++++++++\n>  odb/source.h          |  6 ++++++\n>  6 files changed, 37 insertions(+), 4 deletions(-)\n\nOK.\n\n> +static const char * const odb_source_names_by_type[] = {\n> +\t[ODB_SOURCE_UNKNOWN] = \"unknown\",\n> +\t[ODB_SOURCE_FILES] = \"files\",\n> +\t[ODB_SOURCE_LOOSE] = \"loose\",\n> +\t[ODB_SOURCE_PACKED] = \"packed\",\n> +\t[ODB_SOURCE_INMEMORY] = \"inmemory\",\n> +};\n\nThis is a trivially obvious implementation for mapping in either\ndirection.\n\n'inmemory' should probably be spelled 'in-memory', though.\n\nThanks.\n\n> +const char *odb_source_type_to_name(enum odb_source_type type)\n> +{\n> +\tconst char *name;\n> +\tif (type < 0 || type >= ARRAY_SIZE(odb_source_names_by_type))\n> +\t\ttype = ODB_SOURCE_UNKNOWN;\n> +\tname = odb_source_names_by_type[type];\n> +\tif (!name)\n> +\t\tBUG(\"name missing in `odb_source_names_by_type` for '%d'\", type);\n> +\treturn name;\n> +}\n> +\n>  struct odb_source *odb_source_new(struct object_database *odb,\n>  \t\t\t\t  const char *path,\n>  \t\t\t\t  bool local)\n> diff --git a/odb/source.h b/odb/source.h\n> index cd63dba91f..ab16d152f4 100644\n> --- a/odb/source.h\n> +++ b/odb/source.h\n> @@ -25,6 +25,12 @@ enum odb_source_type {\n>  \tODB_SOURCE_INMEMORY,\n>  };\n>  \n> +/*\n> + * Convert between the enum and its name. Returns the equivalent of \"unknown\"\n> + * for unknown types.\n> + */\n> +const char *odb_source_type_to_name(enum odb_source_type type);\n> +\n>  struct object_id;\n>  struct odb_read_stream;\n>  struct strvec;\n"},{"id":"549046","messageId":"xmqqbjbtv5y6.fsf@gitster.g","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-5-3b3d265d979b@pks.im","subject":"Re: [PATCH 5/5] odb: make creation of on-disk structures pluggable","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-26T20:42:09Z","receivedAt":"2026-07-26T20:42:12Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Note that there is one exception though: the \"objects\" directory must\n> exist in a repository regardless of which backend is in use. If it\n> doesn't exist then the repository is not treated as a Git repository at\n> all. Consequently, we create this directory regardless of the backend.\n\nVery good thing to leave a note in the log message for.\n\nPerhaps in Git 4.0 ;-)\n\n> @@ -271,6 +289,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n>  \n>  \tfiles->base.free = odb_source_files_free;\n>  \tfiles->base.close = odb_source_files_close;\n> +\tfiles->base.create_on_disk = odb_source_files_create_on_disk;\n>  \tfiles->base.prepare = odb_source_files_prepare;\n>  \tfiles->base.read_object_info = odb_source_files_read_object_info;\n>  \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\n\nIf we are going to write a brand new object backing store that does\nnot use an on-disk filesystem (or a network filesystem, for that\nmatter) but still requires some sort of \"initialization\", for\nexample, an object database in the cloud that needs provisioning\nbefore its first use, would this virtual function be the ideal place\nto do so?\n\nI wonder if we can give it a name better suited to its purpose by\nmoving away from the '_on_disk' suffix.\n\n"},{"id":"549163","messageId":"amkMipjGA_7cwpOR@denethor","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-1-3b3d265d979b@pks.im","subject":"Re: [PATCH 1/5] loose: load loose object map for the correct source","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-07-28T20:14:05Z","receivedAt":"2026-07-28T20:14:09Z","isPatch":true,"body":"On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n> When loading the loose object map via `load_one_loose_object_map()` we\n> pass in both a repository and the corresponding source. We ultimately\n> don't really respect the passed-in source though as we instead always\n> load the map via the common directory. This doesn't make any sense\n> though, as the function is called in a loop through all sources, and as\n> such the expectation is that we'll load the map that belongs to the\n> given source.\n> \n> Fix this bug by instead loading the map via the loose source's path.\n\nIIUC the primary source is always being used, does this mean that\nrepositories using a compat hash and alternates are currently broken?\n\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  loose.c | 18 ++++++++++--------\n>  1 file changed, 10 insertions(+), 8 deletions(-)\n> \n> diff --git a/loose.c b/loose.c\n> index bf01d3e42d..9dad75373b 100644\n> --- a/loose.c\n> +++ b/loose.c\n> @@ -61,9 +61,11 @@ static int insert_loose_map(struct odb_source_loose *loose,\n>  \treturn inserted;\n>  }\n>  \n> -static int load_one_loose_object_map(struct repository *repo, struct odb_source_loose *loose)\n> +static int load_one_loose_object_map(struct odb_source_loose *loose)\n>  {\n> -\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n> +\tstruct repository *repo = loose->base.odb->repo;\n\nOk, we really only need the repository to know the hash algo, but we can\nget this from the loose source.\n\n> +\tstruct strbuf buf = STRBUF_INIT;\n> +\tchar *path;\n>  \tFILE *fp;\n>  \tint ret = -1;\n>  \n> @@ -78,10 +80,10 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n>  \tinsert_loose_map(loose, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n>  \tinsert_loose_map(loose, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n>  \n> -\trepo_common_path_replace(repo, &path, \"objects/loose-object-idx\");\n> -\tfp = fopen(path.buf, \"rb\");\n> +\tpath = xstrfmt(\"%s/loose-object-idx\", loose->base.path);\n\nNow we use the correct path per source. Looks good.\n\n-Justin\n"},{"id":"549164","messageId":"amkOb3rvWFUpnT28@denethor","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-2-3b3d265d979b@pks.im","subject":"Re: [PATCH 2/5] setup: detangle loading of loose object maps","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-07-28T20:32:27Z","receivedAt":"2026-07-28T20:32:31Z","isPatch":true,"body":"On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n> When a repository is configured to use a compatibility hash function\n> then we load the loose object map when we initialize the repository.\n> This object map provides the mappings between the canonical object hash\n> and the compatibility object hash.\n> \n> Loading the object map happens in `repo_set_compat_hash_algo()`, which\n> calls `repo_read_loose_object_map()` in case the compatibility object\n> hash is non-zero. This setup sequence has two major downsides:\n> \n>   - We assume that the primary object database is the \"files\" object\n>     database so that we can extract its \"loose\" backend. This stops\n>     working with pluggable object databases.\n\nSo IIUC, does this mean that `repo_set_compat_hash_algo()` is directly\nreaching into the loose object source to load the compatibility object\nmap? I suppose it should be the responsibility of the respective ODB\nbackend to handle object compatibility.\n\n>   - We require the object database to already have been initialized when\n>     configuring the object database. This means that we must intermix\n>     configuration of the repository and initialization of its\n>     sub-structures in a weird way.\n\nIf there any reason we need to eagerly load compatibility object\nmappings?\n\n> Refactor the logic so that we instead load the loose object map via the\n> \"loose\" backend, which fixes both of the above issues.\n\nSounds reasonable.\n\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  loose.c            | 11 +++++------\n>  loose.h            |  1 +\n>  odb/source-loose.c |  2 ++\n>  repository.c       |  2 --\n>  setup.c            |  5 +++--\n>  5 files changed, 11 insertions(+), 10 deletions(-)\n> \n> diff --git a/loose.c b/loose.c\n> index 9dad75373b..a3b2dcedc2 100644\n> --- a/loose.c\n> +++ b/loose.c\n> @@ -61,7 +61,7 @@ static int insert_loose_map(struct odb_source_loose *loose,\n>  \treturn inserted;\n>  }\n>  \n> -static int load_one_loose_object_map(struct odb_source_loose *loose)\n> +int loose_object_map_load(struct odb_source_loose *loose)\n>  {\n>  \tstruct repository *repo = loose->base.odb->repo;\n>  \tstruct strbuf buf = STRBUF_INIT;\n> @@ -69,6 +69,9 @@ static int load_one_loose_object_map(struct odb_source_loose *loose)\n>  \tFILE *fp;\n>  \tint ret = -1;\n>  \n> +\tif (!should_use_loose_object_map(repo))\n> +\t\treturn 0;\n\nPreviously the above condition has asserted in\n`repo_read_loose_object_map()` which calls `loose_object_map_load()` for\neach source. Do we expect each source to potentially answer differently\nthough?\n\n> +\n>  \tif (!loose->map)\n>  \t\tloose_object_map_init(&loose->map);\n>  \tif (!loose->cache) {\n> @@ -112,14 +115,10 @@ int repo_read_loose_object_map(struct repository *repo)\n>  {\n>  \tstruct odb_source *source;\n>  \n> -\tif (!should_use_loose_object_map(repo))\n> -\t\treturn 0;\n> -\n>  \todb_prepare_alternates(repo->objects);\n> -\n>  \tfor (source = repo->objects->sources; source; source = source->next) {\n>  \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n> -\t\tif (load_one_loose_object_map(files->loose) < 0)\n> +\t\tif (loose_object_map_load(files->loose) < 0)\n>  \t\t\treturn -1;\n>  \t}\n>  \n> diff --git a/loose.h b/loose.h\n> index 6c9b3f4571..ed663ac550 100644\n> --- a/loose.h\n> +++ b/loose.h\n> @@ -13,6 +13,7 @@ struct loose_object_map {\n>  \n>  void loose_object_map_init(struct loose_object_map **map);\n>  void loose_object_map_clear(struct loose_object_map **map);\n> +int loose_object_map_load(struct odb_source_loose *loose);\n>  int repo_loose_object_map_oid(struct repository *repo,\n>  \t\t\t      const struct object_id *src,\n>  \t\t\t      const struct git_hash_algo *dest_algo,\n> diff --git a/odb/source-loose.c b/odb/source-loose.c\n> index 3f7d04a56e..812ca1c138 100644\n> --- a/odb/source-loose.c\n> +++ b/odb/source-loose.c\n> @@ -727,5 +727,7 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n>  \tif (!is_absolute_path(loose->base.path))\n>  \t\tchdir_notify_register(NULL, odb_source_loose_reparent, loose);\n>  \n> +\tloose_object_map_load(loose);\n\nNow we load the loose object map for the specific source when its\ncreated.\n\n> +\n>  \treturn loose;\n>  }\n> diff --git a/repository.c b/repository.c\n> index 2ef0778846..6d633002b4 100644\n> --- a/repository.c\n> +++ b/repository.c\n> @@ -201,8 +201,6 @@ void repo_set_compat_hash_algo(struct repository *repo MAYBE_UNUSED, uint32_t al\n>  \tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n>  \t\tBUG(\"hash_algo and compat_hash_algo match\");\n>  \trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n> -\tif (repo->compat_hash_algo)\n> -\t\trepo_read_loose_object_map(repo);\n\nThe loose object map is no longer read eagerly.\n\n>  #else\n>  \tif (algo)\n>  \t\tdie(_(\"compatibility hash algorithm support requires Rust\"));\n> diff --git a/setup.c b/setup.c\n> index d31808130b..825572f5f1 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -1788,8 +1788,6 @@ int apply_repository_format(struct repository *repo,\n>  \n>  \trepo->bare_cfg = format->is_bare;\n>  \trepo_set_hash_algo(repo, format->hash_algo);\n> -\trepo->objects = odb_new(repo, object_directory,\n> -\t\t\t\talternate_object_directories);\n>  \trepo_set_compat_hash_algo(repo, format->compat_hash_algo);\n>  \trepo_set_ref_storage_format(repo,\n>  \t\t\t\t    format->ref_storage_format,\n> @@ -1805,6 +1803,9 @@ int apply_repository_format(struct repository *repo,\n>  \trepo->repository_format_precious_objects =\n>  \t\tformat->precious_objects;\n>  \n> +\trepo->objects = odb_new(repo, object_directory,\n> +\t\t\t\talternate_object_directories);\n\nWe now defer creating the ODB until after the compat hash is configured.\nMakes sense.\n\n-Justin\n"},{"id":"549166","messageId":"amkXcmwzbBYsMgjc@denethor","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-3-3b3d265d979b@pks.im","subject":"Re: [PATCH 3/5] setup: defer object database creation","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-07-28T21:13:42Z","receivedAt":"2026-07-28T21:13:46Z","isPatch":true,"body":"On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n> In a subsequent commit we'll make the creation of the on-disk data\n> structures of an object database pluggable. This will lead to an\n> in-between state where we have already configured the repository's\n> object database, but it's not usable yet until we eventually call\n> `create_object_directory()`.\n>\n> Defer the object database creation so that we handle both steps in the\n> same function.\n\nSo IIUC, the repository gets configured via `apply_repository_format()`\nwhich invokes `odb_new()`. In this patch a\nAPPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION flag is introduced to allow\nthe creation of the ODB to be delayed until after source specific\non-disk state has been created.\n\nNaive question: would it be simpler to just require invoking `odb_new()`\nexplicitly after `apply_repository_format()` in all cases? There doesn't\nappear to be too many callsites.\n\n-Justin\n"},{"id":"549167","messageId":"amkcNhMTKqWdLXwX@denethor","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-5-3b3d265d979b@pks.im","subject":"Re: [PATCH 5/5] odb: make creation of on-disk structures pluggable","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-07-28T21:23:41Z","receivedAt":"2026-07-28T21:23:43Z","isPatch":true,"body":"On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n> When creating a new \"files\" object database source we have to create a\n> couple of directories. These directories are of course specific to this\n> particular backend, and a different backend may require a setup that is\n> completely different.\n> \n> Make the creation of on-disk structures pluggable to accommodate for\n> this.\n\nOk.\n\n> Note that there is one exception though: the \"objects\" directory must\n> exist in a repository regardless of which backend is in use. If it\n> doesn't exist then the repository is not treated as a Git repository at\n> all. Consequently, we create this directory regardless of the backend.\n\nMakes sense.\n\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb/source-files.c | 19 +++++++++++++++++++\n>  odb/source.h       | 23 +++++++++++++++++++++++\n>  setup.c            | 35 ++++++++++++++++++++---------------\n>  3 files changed, 62 insertions(+), 15 deletions(-)\n> \n> diff --git a/odb/source-files.c b/odb/source-files.c\n> index 4138758511..0db6e681fe 100644\n> --- a/odb/source-files.c\n> +++ b/odb/source-files.c\n> @@ -9,6 +9,7 @@\n>  #include \"odb/source-files.h\"\n>  #include \"odb/source-loose.h\"\n>  #include \"packfile.h\"\n> +#include \"path.h\"\n>  #include \"strbuf.h\"\n>  #include \"write-or-die.h\"\n>  \n> @@ -41,6 +42,23 @@ static void odb_source_files_close(struct odb_source *source)\n>  \todb_source_close(&files->packed->base);\n>  }\n>  \n> +static int odb_source_files_create_on_disk(struct odb_source *source)\n> +{\n> +\tstruct strbuf path = STRBUF_INIT;\n> +\n> +\tsafe_create_dir(source->odb->repo, source->path, 1);\n> +\n> +\tstrbuf_addf(&path, \"%s/pack\", source->path);\n> +\tsafe_create_dir(source->odb->repo, path.buf, 1);\n> +\n> +\tstrbuf_reset(&path);\n> +\tstrbuf_addf(&path, \"%s/info\", source->path);\n> +\tsafe_create_dir(source->odb->repo, path.buf, 1);\n> +\n> +\tstrbuf_release(&path);\n> +\treturn 0;\n> +}\n\nThis is the callback to create on-disk state specific to the \"files\"\nsource and matches the current set of created files.\n\n> +\n>  static void odb_source_files_prepare(struct odb_source *source,\n>  \t\t\t\t     enum odb_prepare_flags flags)\n>  {\n> @@ -271,6 +289,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n>  \n>  \tfiles->base.free = odb_source_files_free;\n>  \tfiles->base.close = odb_source_files_close;\n> +\tfiles->base.create_on_disk = odb_source_files_create_on_disk;\n>  \tfiles->base.prepare = odb_source_files_prepare;\n>  \tfiles->base.read_object_info = odb_source_files_read_object_info;\n>  \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\n> diff --git a/odb/source.h b/odb/source.h\n> index ab16d152f4..4abc418bdd 100644\n> --- a/odb/source.h\n> +++ b/odb/source.h\n> @@ -89,6 +89,18 @@ struct odb_source {\n>  \t */\n>  \tvoid (*close)(struct odb_source *source);\n>  \n> +\t/*\n> +\t * This callback is expected to create on-disk data structures that are\n> +\t * required for this source to operate.\n> +\t *\n> +\t * The callback is expected to return 0 on success, a negative error\n> +\t * code otherwise.\n> +\t *\n> +\t * This callback may be NULL in case the source does not need any\n> +\t * on-disk setup.\n> +\t */\n> +\tint (*create_on_disk)(struct odb_source *source);\n> +\n>  \t/*\n>  \t * This callback is expected to prepare the source so that it becomes\n>  \t * ready for use. It optionally clears underlying caches of the object\n> @@ -316,6 +328,17 @@ static inline void odb_source_close(struct odb_source *source)\n>  \tsource->close(source);\n>  }\n>  \n> +/*\n> + * Create on-disk data structures that are required for this source to operate\n> + * correctly. Returns 0 on success, a negative error code otherwise.\n> + */\n> +static inline int odb_source_create_on_disk(struct odb_source *source)\n> +{\n> +\tif (!source->create_on_disk)\n> +\t\treturn 0;\n> +\treturn source->create_on_disk(source);\n> +}\n> +\n>  /*\n>   * Prepare the object database source and clear any caches. Depending on the\n>   * backend used this may have the effect that concurrently-written objects\n> diff --git a/setup.c b/setup.c\n> index a7b1b9eaef..14ef119cb7 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -2666,29 +2666,34 @@ static int create_default_files(struct repository *repo,\n>  static void create_object_database(struct repository *repo)\n>  {\n>  \tchar *object_directory, *alternate_object_directories;\n> -\tstruct strbuf path = STRBUF_INIT;\n> -\tsize_t baselen;\n>  \n>  \tget_object_directories(&object_directory, &alternate_object_directories);\n> -\trepo->objects = odb_new(repo, object_directory,\n> -\t\t\t\talternate_object_directories);\n>  \n> -\tstrbuf_addstr(&path, repo_get_object_directory(repo));\n> -\tbaselen = path.len;\n> -\n> -\tsafe_create_dir(repo, path.buf, 1);\n> +\t/*\n> +\t * Create the \"objects\" directory in the common directory. This is done\n> +\t * so that the repository can be discovered regardless of the backend\n> +\t * used.\n> +\t *\n> +\t * Note that we only do this in case the object directory wasn't\n> +\t * overwritten via an environment variable. If it _is_ being overridden\n> +\t * then we skip this step, as the repository won't be discoverable\n> +\t * anyway without the environment variable.\n> +\t */\n> +\tif (!object_directory) {\n> +\t\tstruct strbuf objects_dir = STRBUF_INIT;\n> +\t\trepo_common_path_append(repo, &objects_dir, \"objects\");\n> +\t\tsafe_create_dir(repo, objects_dir.buf, 1);\n> +\t\tstrbuf_release(&objects_dir);\n> +\t}\n\nHere we always create the objects directory regardless of the backend.\nLooks good.\n\n> -\tstrbuf_setlen(&path, baselen);\n> -\tstrbuf_addstr(&path, \"/pack\");\n> -\tsafe_create_dir(repo, path.buf, 1);\n> +\trepo->objects = odb_new(repo, object_directory,\n> +\t\t\t\talternate_object_directories);\n>  \n> -\tstrbuf_setlen(&path, baselen);\n> -\tstrbuf_addstr(&path, \"/info\");\n> -\tsafe_create_dir(repo, path.buf, 1);\n> +\tif (odb_source_create_on_disk(repo->objects->sources) < 0)\n> +\t\tdie(\"failed creating object database\");\n\nHere we invoke the pluggable callback to create source specific on-disk\nstate. Part of me does wonder if this would be better to include this\ninside of `odb_new()` and enable it with a specific flag, but having it\nas a explicit separate step is probably fine too.\n\n-Justin\n"},{"id":"549297","messageId":"87tspgd4p5.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"amkMipjGA_7cwpOR@denethor","subject":"Re: [PATCH 1/5] loose: load loose object map for the correct source","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-07-30T12:47:50Z","receivedAt":"2026-07-30T12:47:58Z","isPatch":true,"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n>> When loading the loose object map via `load_one_loose_object_map()` we\n>> pass in both a repository and the corresponding source. We ultimately\n>> don't really respect the passed-in source though as we instead always\n>> load the map via the common directory. This doesn't make any sense\n>> though, as the function is called in a loop through all sources, and as\n>> such the expectation is that we'll load the map that belongs to the\n>> given source.\n>> \n>> Fix this bug by instead loading the map via the loose source's path.\n>\n> IIUC the primary source is always being used, does this mean that\n> repositories using a compat hash and alternates are currently broken?\n\nYeah, the commit message seems to undersell this fix.\n\nI think it wouldn't hurt to add a small test for this:\n\n    test_expect_success 'rev-parse maps oid of object borrowed from alternate' '\n    \ttest_when_finished rm -rf alt borrow &&\n    \n    \tgit init --object-format=sha256 alt &&\n    \tgit -C alt config extensions.compatObjectFormat sha1 &&\n    \ttest_commit -C alt A &&\n    \n    \tgit init --object-format=sha256 borrow &&\n    \tgit -C borrow config extensions.compatObjectFormat sha1 &&\n    \techo \"$PWD/alt/.git/objects\" >borrow/.git/objects/info/alternates &&\n    \n    \toid=$(git -C alt rev-parse HEAD) &&\n    \tgit -C alt    rev-parse --output-object-format=sha1 \"$oid\" >expect &&\n    \tgit -C borrow rev-parse --output-object-format=sha1 \"$oid\" >actual &&\n    \ttest_cmp expect actual\n    '\n\n-- \nCheers,\nToon\n"},{"id":"549308","messageId":"87pl04d03q.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"amkOb3rvWFUpnT28@denethor","subject":"Re: [PATCH 2/5] setup: detangle loading of loose object maps","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-07-30T14:27:05Z","receivedAt":"2026-07-30T14:27:14Z","isPatch":true,"body":"Justin Tobler <jltobler@gmail.com> writes:\n\n> On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n>> When a repository is configured to use a compatibility hash function\n>> then we load the loose object map when we initialize the repository.\n>> This object map provides the mappings between the canonical object hash\n>> and the compatibility object hash.\n>> \n>> Loading the object map happens in `repo_set_compat_hash_algo()`, which\n>> calls `repo_read_loose_object_map()` in case the compatibility object\n>> hash is non-zero. This setup sequence has two major downsides:\n>> \n>>   - We assume that the primary object database is the \"files\" object\n>>     database so that we can extract its \"loose\" backend. This stops\n>>     working with pluggable object databases.\n>\n> So IIUC, does this mean that `repo_set_compat_hash_algo()` is directly\n> reaching into the loose object source to load the compatibility object\n> map? I suppose it should be the responsibility of the respective ODB\n> backend to handle object compatibility.\n>\n>>   - We require the object database to already have been initialized when\n>>     configuring the object database. This means that we must intermix\n>>     configuration of the repository and initialization of its\n>>     sub-structures in a weird way.\n>\n> If there any reason we need to eagerly load compatibility object\n> mappings?\n>\n>> Refactor the logic so that we instead load the loose object map via the\n>> \"loose\" backend, which fixes both of the above issues.\n>\n> Sounds reasonable.\n>\n>> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n>> ---\n>>  loose.c            | 11 +++++------\n>>  loose.h            |  1 +\n>>  odb/source-loose.c |  2 ++\n>>  repository.c       |  2 --\n>>  setup.c            |  5 +++--\n>>  5 files changed, 11 insertions(+), 10 deletions(-)\n>> \n>> diff --git a/loose.c b/loose.c\n>> index 9dad75373b..a3b2dcedc2 100644\n>> --- a/loose.c\n>> +++ b/loose.c\n>> @@ -61,7 +61,7 @@ static int insert_loose_map(struct odb_source_loose *loose,\n>>  \treturn inserted;\n>>  }\n>>  \n>> -static int load_one_loose_object_map(struct odb_source_loose *loose)\n>> +int loose_object_map_load(struct odb_source_loose *loose)\n>>  {\n>>  \tstruct repository *repo = loose->base.odb->repo;\n>>  \tstruct strbuf buf = STRBUF_INIT;\n>> @@ -69,6 +69,9 @@ static int load_one_loose_object_map(struct odb_source_loose *loose)\n>>  \tFILE *fp;\n>>  \tint ret = -1;\n>>  \n>> +\tif (!should_use_loose_object_map(repo))\n>> +\t\treturn 0;\n>\n> Previously the above condition has asserted in\n> `repo_read_loose_object_map()` which calls `loose_object_map_load()` for\n> each source. Do we expect each source to potentially answer differently\n> though?\n\nI've been wondering about this as well. The reason for this change is to\nalso have this guard when odb_source_loose_new(), in source-loose.c (see\nfurther down in the patch), calls this function too.\n\n>> +\n>>  \tif (!loose->map)\n>>  \t\tloose_object_map_init(&loose->map);\n>>  \tif (!loose->cache) {\n>> @@ -112,14 +115,10 @@ int repo_read_loose_object_map(struct repository *repo)\n>>  {\n>>  \tstruct odb_source *source;\n>>  \n>> -\tif (!should_use_loose_object_map(repo))\n>> -\t\treturn 0;\n>> -\n>>  \todb_prepare_alternates(repo->objects);\n>> -\n>>  \tfor (source = repo->objects->sources; source; source = source->next) {\n>>  \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n>> -\t\tif (load_one_loose_object_map(files->loose) < 0)\n>> +\t\tif (loose_object_map_load(files->loose) < 0)\n>>  \t\t\treturn -1;\n>>  \t}\n>>  \n>> diff --git a/loose.h b/loose.h\n>> index 6c9b3f4571..ed663ac550 100644\n>> --- a/loose.h\n>> +++ b/loose.h\n>> @@ -13,6 +13,7 @@ struct loose_object_map {\n>>  \n>>  void loose_object_map_init(struct loose_object_map **map);\n>>  void loose_object_map_clear(struct loose_object_map **map);\n>> +int loose_object_map_load(struct odb_source_loose *loose);\n>>  int repo_loose_object_map_oid(struct repository *repo,\n>>  \t\t\t      const struct object_id *src,\n>>  \t\t\t      const struct git_hash_algo *dest_algo,\n>> diff --git a/odb/source-loose.c b/odb/source-loose.c\n>> index 3f7d04a56e..812ca1c138 100644\n>> --- a/odb/source-loose.c\n>> +++ b/odb/source-loose.c\n>> @@ -727,5 +727,7 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n>>  \tif (!is_absolute_path(loose->base.path))\n>>  \t\tchdir_notify_register(NULL, odb_source_loose_reparent, loose);\n>>  \n>> +\tloose_object_map_load(loose);\n>\n> Now we load the loose object map for the specific source when its\n> created.\n\nHere.\n\n-- \nCheers,\nToon\n"},{"id":"549520","messageId":"anGS37r_67pWr7u0@pks.im","threadId":"66056","inReplyTo":"87tspgd4p5.fsf@emacs.iotcl.com","subject":"Re: [PATCH 1/5] loose: load loose object map for the correct source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T07:21:03Z","receivedAt":"2026-08-04T07:21:17Z","isPatch":true,"body":"On Thu, Jul 30, 2026 at 02:47:50PM +0200, Toon Claes wrote:\n> Justin Tobler <jltobler@gmail.com> writes:\n> \n> > On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n> >> When loading the loose object map via `load_one_loose_object_map()` we\n> >> pass in both a repository and the corresponding source. We ultimately\n> >> don't really respect the passed-in source though as we instead always\n> >> load the map via the common directory. This doesn't make any sense\n> >> though, as the function is called in a loop through all sources, and as\n> >> such the expectation is that we'll load the map that belongs to the\n> >> given source.\n> >> \n> >> Fix this bug by instead loading the map via the loose source's path.\n> >\n> > IIUC the primary source is always being used, does this mean that\n> > repositories using a compat hash and alternates are currently broken?\n> \n> Yeah, the commit message seems to undersell this fix.\n> \n> I think it wouldn't hurt to add a small test for this:\n> \n>     test_expect_success 'rev-parse maps oid of object borrowed from alternate' '\n>     \ttest_when_finished rm -rf alt borrow &&\n>     \n>     \tgit init --object-format=sha256 alt &&\n>     \tgit -C alt config extensions.compatObjectFormat sha1 &&\n>     \ttest_commit -C alt A &&\n>     \n>     \tgit init --object-format=sha256 borrow &&\n>     \tgit -C borrow config extensions.compatObjectFormat sha1 &&\n>     \techo \"$PWD/alt/.git/objects\" >borrow/.git/objects/info/alternates &&\n>     \n>     \toid=$(git -C alt rev-parse HEAD) &&\n>     \tgit -C alt    rev-parse --output-object-format=sha1 \"$oid\" >expect &&\n>     \tgit -C borrow rev-parse --output-object-format=sha1 \"$oid\" >actual &&\n>     \ttest_cmp expect actual\n>     '\n\nGood idea indeed, will do. Thanks!\n\nPatrick\n"},{"id":"549521","messageId":"anGS69L4vh3TsDlP@pks.im","threadId":"66056","inReplyTo":"xmqqbjbtv5y6.fsf@gitster.g","subject":"Re: [PATCH 5/5] odb: make creation of on-disk structures pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T07:21:15Z","receivedAt":"2026-08-04T07:21:20Z","isPatch":true,"body":"On Sun, Jul 26, 2026 at 01:42:09PM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > @@ -271,6 +289,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n> >  \n> >  \tfiles->base.free = odb_source_files_free;\n> >  \tfiles->base.close = odb_source_files_close;\n> > +\tfiles->base.create_on_disk = odb_source_files_create_on_disk;\n> >  \tfiles->base.prepare = odb_source_files_prepare;\n> >  \tfiles->base.read_object_info = odb_source_files_read_object_info;\n> >  \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\n> \n> If we are going to write a brand new object backing store that does\n> not use an on-disk filesystem (or a network filesystem, for that\n> matter) but still requires some sort of \"initialization\", for\n> example, an object database in the cloud that needs provisioning\n> before its first use, would this virtual function be the ideal place\n> to do so?\n\nIt would, even though...\n\n> I wonder if we can give it a name better suited to its purpose by\n> moving away from the '_on_disk' suffix.\n\n... the name is admittedly a bit misleading. I couldn't really come up\nwith a better name though, and the `on_disk()` suffix is what we already\nuse in the reference subsystem, too (see `ref_store_create_on_disk()`).\nSo I'm inclined to leave the name as-is for the sake of consistency.\n\nPatrick\n"},{"id":"549522","messageId":"anGS8aD_j6rq5YeT@pks.im","threadId":"66056","inReplyTo":"xmqqfr15v6ba.fsf@gitster.g","subject":"Re: [PATCH 4/5] odb/source: introduce function to map source type to name","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T07:21:21Z","receivedAt":"2026-08-04T07:21:27Z","isPatch":true,"body":"On Sun, Jul 26, 2026 at 01:34:17PM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > Introduce a new function that maps an object source's type to a\n> > human-readable name. Use the function to provide better human-readable\n> > error messages for the downcasting functions.\n> >\n> > Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> > ---\n> >  odb/source-files.h    |  4 +++-\n> >  odb/source-inmemory.h |  4 +++-\n> >  odb/source-loose.h    |  4 +++-\n> >  odb/source-packed.h   |  4 +++-\n> >  odb/source.c          | 19 +++++++++++++++++++\n> >  odb/source.h          |  6 ++++++\n> >  6 files changed, 37 insertions(+), 4 deletions(-)\n> \n> OK.\n> \n> > +static const char * const odb_source_names_by_type[] = {\n> > +\t[ODB_SOURCE_UNKNOWN] = \"unknown\",\n> > +\t[ODB_SOURCE_FILES] = \"files\",\n> > +\t[ODB_SOURCE_LOOSE] = \"loose\",\n> > +\t[ODB_SOURCE_PACKED] = \"packed\",\n> > +\t[ODB_SOURCE_INMEMORY] = \"inmemory\",\n> > +};\n> \n> This is a trivially obvious implementation for mapping in either\n> direction.\n> \n> 'inmemory' should probably be spelled 'in-memory', though.\n\nFair, that reads better indeed. Will adapt.\n\nPatrick\n"},{"id":"549523","messageId":"anGS-IOKHMo5VUJm@pks.im","threadId":"66056","inReplyTo":"xmqqh5lo6xi2.fsf@gitster.g","subject":"Re: [PATCH 2/5] setup: detangle loading of loose object maps","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T07:21:28Z","receivedAt":"2026-08-04T07:21:33Z","isPatch":true,"body":"On Fri, Jul 24, 2026 at 11:41:41AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > When a repository is configured to use a compatibility hash function\n> > then we load the loose object map when we initialize the repository.\n> > This object map provides the mappings between the canonical object hash\n> > and the compatibility object hash.\n> >\n> > Loading the object map happens in `repo_set_compat_hash_algo()`, which\n> > calls `repo_read_loose_object_map()` in case the compatibility object\n> > hash is non-zero. This setup sequence has two major downsides:\n> >\n> >   - We assume that the primary object database is the \"files\" object\n> >     database so that we can extract its \"loose\" backend. This stops\n> >     working with pluggable object databases.\n> \n> I am not sure if I understand this sentence, especially \"we can\n> extract its loose backend\" part.  Do you mean 'extract the object\n> map from the loose backend'?  Or something else?\n\nYeah, this is a bit awkward. Rewritten like this:\n\n  - We assume that the primary object database is the \"files\" object\n    database and unconditionally downcast it. This will BUG in case a\n    different object database type was used together with a compat hash\n    algorithm.\n\n> > @@ -112,14 +115,10 @@ int repo_read_loose_object_map(struct repository *repo)\n> >  {\n> >  \tstruct odb_source *source;\n> >  \n> > -\tif (!should_use_loose_object_map(repo))\n> > -\t\treturn 0;\n> > -\n> >  \todb_prepare_alternates(repo->objects);\n> > -\n> >  \tfor (source = repo->objects->sources; source; source = source->next) {\n> >  \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n> > -\t\tif (load_one_loose_object_map(files->loose) < 0)\n> > +\t\tif (loose_object_map_load(files->loose) < 0)\n> >  \t\t\treturn -1;\n> \n> If this particular source in the list of sources is not backed by\n> the files backend, would downcast signal the fact (e.g., by\n> returning NULL) so that we can skip the next call instead?\n\nNo, the downcast will BUG in case it's not the \"files\" backend.\n\n> Or would the next step in refactoring be to define \"load object map\"\n> method that is generic to odb_source so that this part does not have\n> to do any of these and instead simply do\n> \n> \tfor (source = ...) {\n> \t\tif (odb_source_object_map_load(source))\n>                 \treturn -1;\n>         }\n> \n> or something?\n\nThis patch series is rather moving into the direction of making the\nobject map an internal implementation detail. Ideally, callers shouldn't\neven have to be aware that such an object map exists. And by making the\nloose object source load it automatically we get closer to that state.\n\nThere's only one more caller that calls `repo_read_loose_object_map()`\ndirectly, in \"object-file-convert.c\", and that caller only calls it to\nreload the map in case a concurrent process may have rewritten it. If we\nmake the backends handle this via `odb_source_prepare(FLUSH_CACHES)`\nthen we could also get rid of that caller.\n\nPatrick\n"},{"id":"549524","messageId":"anGS_njkfklt9gbd@pks.im","threadId":"66056","inReplyTo":"amkOb3rvWFUpnT28@denethor","subject":"Re: [PATCH 2/5] setup: detangle loading of loose object maps","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T07:21:34Z","receivedAt":"2026-08-04T07:21:40Z","isPatch":true,"body":"On Tue, Jul 28, 2026 at 03:32:27PM -0500, Justin Tobler wrote:\n> On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n> > When a repository is configured to use a compatibility hash function\n> > then we load the loose object map when we initialize the repository.\n> > This object map provides the mappings between the canonical object hash\n> > and the compatibility object hash.\n> > \n> > Loading the object map happens in `repo_set_compat_hash_algo()`, which\n> > calls `repo_read_loose_object_map()` in case the compatibility object\n> > hash is non-zero. This setup sequence has two major downsides:\n> > \n> >   - We assume that the primary object database is the \"files\" object\n> >     database so that we can extract its \"loose\" backend. This stops\n> >     working with pluggable object databases.\n> \n> So IIUC, does this mean that `repo_set_compat_hash_algo()` is directly\n> reaching into the loose object source to load the compatibility object\n> map? I suppose it should be the responsibility of the respective ODB\n> backend to handle object compatibility.\n> \n> >   - We require the object database to already have been initialized when\n> >     configuring the object database. This means that we must intermix\n> >     configuration of the repository and initialization of its\n> >     sub-structures in a weird way.\n> \n> If there any reason we need to eagerly load compatibility object\n> mappings?\n\nI'm not familiar enough with the compatibility mappings to really be\nable to say. Naively I'd say \"no\", but I'm rather erring on the side of\ncaution and want to leave this as-is.\n\n> > diff --git a/loose.c b/loose.c\n> > index 9dad75373b..a3b2dcedc2 100644\n> > --- a/loose.c\n> > +++ b/loose.c\n> > @@ -69,6 +69,9 @@ static int load_one_loose_object_map(struct odb_source_loose *loose)\n> >  \tFILE *fp;\n> >  \tint ret = -1;\n> >  \n> > +\tif (!should_use_loose_object_map(repo))\n> > +\t\treturn 0;\n> \n> Previously the above condition has asserted in\n> `repo_read_loose_object_map()` which calls `loose_object_map_load()` for\n> each source. Do we expect each source to potentially answer differently\n> though?\n\nNot really, no. But as it's now the source that loads the object map it\nhas to verify for itself whether it should or should not load it.\n\nPatrick\n"},{"id":"549525","messageId":"anGTBQIYpDl7HbXf@pks.im","threadId":"66056","inReplyTo":"amkXcmwzbBYsMgjc@denethor","subject":"Re: [PATCH 3/5] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T07:21:41Z","receivedAt":"2026-08-04T07:21:47Z","isPatch":true,"body":"On Tue, Jul 28, 2026 at 04:13:42PM -0500, Justin Tobler wrote:\n> On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n> > In a subsequent commit we'll make the creation of the on-disk data\n> > structures of an object database pluggable. This will lead to an\n> > in-between state where we have already configured the repository's\n> > object database, but it's not usable yet until we eventually call\n> > `create_object_directory()`.\n> >\n> > Defer the object database creation so that we handle both steps in the\n> > same function.\n> \n> So IIUC, the repository gets configured via `apply_repository_format()`\n> which invokes `odb_new()`. In this patch a\n> APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION flag is introduced to allow\n> the creation of the ODB to be delayed until after source specific\n> on-disk state has been created.\n> \n> Naive question: would it be simpler to just require invoking `odb_new()`\n> explicitly after `apply_repository_format()` in all cases? There doesn't\n> appear to be too many callsites.\n\nI don't think it would, mostly because the logic to figure out the\nobject directory and the alternate object directory requires a bunch of\nlogic.\n\nI think it'll ultimately become simpler though once we move into the\ndirection of what we've discussed in [1], where we said that we want to\nmove handling of those environment variables into the \"files\" backend,\ntoo. And then it might make sense to revisit this.\n\nPatrick\n"},{"id":"549526","messageId":"anGTC81J4q76fUr1@pks.im","threadId":"66056","inReplyTo":"xmqq5x246x35.fsf@gitster.g","subject":"Re: [PATCH 3/5] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T07:21:47Z","receivedAt":"2026-08-04T07:21:53Z","isPatch":true,"body":"On Fri, Jul 24, 2026 at 11:50:38AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > diff --git a/setup.c b/setup.c\n> > index 825572f5f1..a7b1b9eaef 100644\n> > --- a/setup.c\n> > +++ b/setup.c\n> > @@ -2885,7 +2902,9 @@ int init_db(struct repository *repo,\n> >  \n> >  \tif (!(flags & INIT_DB_SKIP_REFDB))\n> >  \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n> > -\tcreate_object_directory(repo);\n> > +\tcreate_object_database(repo);\n> > +\n> > +\tstartup_info->have_repository = 1;\n> \n> Instead we call create_object_database() rather late, after we\n> finish creating leading directories and default files and processing\n> the configuration.  I guess this is a prelude to specifying \"no, we\n> are not doing the files backend but are using this new thing\" in the\n> global configuration?\n\nYes, exactly. Many of the refactorings I'm doing in \"setup.c\" ultimately\nhave the goal to detangle the setup and configuration of repository\nextensions. It's been painful back when I introduced the \"refStorage\"\nextension, and it's still painful now with the planned \"objectStorage\"\nextension. So this time around I decided to detangle the logic before\nintroducing the extension to make the infra easier to understand going\nforward.\n\nPatrick\n"},{"id":"549535","messageId":"anGUo7NZZE0ysS6D@pks.im","threadId":"66056","inReplyTo":"anGTBQIYpDl7HbXf@pks.im","subject":"Re: [PATCH 3/5] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T07:28:35Z","receivedAt":"2026-08-04T07:28:43Z","isPatch":true,"body":"On Tue, Aug 04, 2026 at 09:21:45AM +0200, Patrick Steinhardt wrote:\n> On Tue, Jul 28, 2026 at 04:13:42PM -0500, Justin Tobler wrote:\n> > On 26/07/24 05:48AM, Patrick Steinhardt wrote:\n> > > In a subsequent commit we'll make the creation of the on-disk data\n> > > structures of an object database pluggable. This will lead to an\n> > > in-between state where we have already configured the repository's\n> > > object database, but it's not usable yet until we eventually call\n> > > `create_object_directory()`.\n> > >\n> > > Defer the object database creation so that we handle both steps in the\n> > > same function.\n> > \n> > So IIUC, the repository gets configured via `apply_repository_format()`\n> > which invokes `odb_new()`. In this patch a\n> > APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION flag is introduced to allow\n> > the creation of the ODB to be delayed until after source specific\n> > on-disk state has been created.\n> > \n> > Naive question: would it be simpler to just require invoking `odb_new()`\n> > explicitly after `apply_repository_format()` in all cases? There doesn't\n> > appear to be too many callsites.\n> \n> I don't think it would, mostly because the logic to figure out the\n> object directory and the alternate object directory requires a bunch of\n> logic.\n> \n> I think it'll ultimately become simpler though once we move into the\n> direction of what we've discussed in [1], where we said that we want to\n> move handling of those environment variables into the \"files\" backend,\n> too. And then it might make sense to revisit this.\n> \n> Patrick\n\n[1]: <amLgMqkqxR8mKIbT@pks.im>\n"},{"id":"549539","messageId":"20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","subject":"[PATCH v2 0/5] odb: make creation of object database pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T08:29:05Z","receivedAt":"2026-08-04T08:29:57Z","isPatch":true,"body":"Hi,\n\nwhen creating a new repository we create a couple of on-disk data\nstructures for the object database. This includes the \"objects/\"\ndirectory hierarchy with \"objects/info\" and \"objects/pack\", which are\nspecific to the backend.\n\nThis patch series makes the creation of the on-disk data structures\npluggable. While we continue to always create \"objects/\" regardless of\nthe backend (it's required for a repository to be recognized as such),\nthe other subdirectories are now created by the backend. This will allow\nother backends to plug in their own logic.\n\nThe series starts with a small detour into the loose-object map. This\ndetour is required so that we can defer initialization of the object\ndatabase itself to a later point in time.\n\nThe series is based on 9a0c4701dc (The 7th batch, 2026-07-22).\n\nChanges in v2:\n  - Add a testcase that demonstrates the bug fixed with alternate loose\n    object maps.\n  - Rename the \"inmemory\" bakcend to \"in-memory\".\n  - Clarify some commit messages.\n  - Link to v1: https://patch.msgid.link/20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (5):\n      loose: load loose object map for the correct source\n      setup: detangle loading of loose object maps\n      setup: defer object database creation\n      odb/source: introduce function to map source type to name\n      odb: make creation of on-disk structures pluggable\n\n loose.c                       | 25 +++++++++--------\n loose.h                       |  1 +\n odb/source-files.c            | 19 +++++++++++++\n odb/source-files.h            |  4 ++-\n odb/source-inmemory.h         |  4 ++-\n odb/source-loose.c            |  2 ++\n odb/source-loose.h            |  4 ++-\n odb/source-packed.h           |  4 ++-\n odb/source.c                  | 19 +++++++++++++\n odb/source.h                  | 29 +++++++++++++++++++\n repository.c                  |  2 --\n setup.c                       | 65 ++++++++++++++++++++++++++++++-------------\n setup.h                       |  9 ++++++\n t/t1016-compatObjectFormat.sh | 18 ++++++++++++\n 14 files changed, 167 insertions(+), 38 deletions(-)\n\nRange-diff versus v1:\n\n1:  c126882da3 ! 1:  087bbd9fa7 loose: load loose object map for the correct source\n    @@ Commit message\n         load the map via the common directory. This doesn't make any sense\n         though, as the function is called in a loop through all sources, and as\n         such the expectation is that we'll load the map that belongs to the\n    -    given source.\n    +    given source. The consequence is that we'll ignore loose object maps of\n    +    any configured alternates.\n     \n         Fix this bug by instead loading the map via the loose source's path.\n     \n    +    Helped-by: Toon Claes <toon@iotcl.com>\n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n     \n      ## loose.c ##\n    @@ loose.c: int repo_read_loose_object_map(struct repository *repo)\n      \treturn 0;\n      }\n      \n    +\n    + ## t/t1016-compatObjectFormat.sh ##\n    +@@ t/t1016-compatObjectFormat.sh: do\n    + \t\teval signedtag3_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag3) &&\n    + \t\teval signedtag4_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag4)\n    + \t'\n    ++\n    ++\ttest_expect_success 'rev-parse maps oid of object borrowed from alternate' '\n    ++\t\tfor repo in alt borrow\n    ++\t\tdo\n    ++\t\t\ttest_when_finished \"rm -rf $repo\" &&\n    ++\t\t\tgit init --object-format=$hash $repo &&\n    ++\t\t\tgit -C $repo config set core.repositoryformatversion 1 &&\n    ++\t\t\tgit -C $repo config set extensions.compatObjectFormat $(compat_hash $hash) || exit 1\n    ++\t\tdone &&\n    ++\n    ++\t\tgit -C alt commit --allow-empty --message A &&\n    ++\t\techo \"$(pwd)/alt/.git/objects\" >borrow/.git/objects/info/alternates &&\n    ++\n    ++\t\toid=$(git -C alt rev-parse HEAD) &&\n    ++\t\tgit -C alt    rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >expect &&\n    ++\t\tgit -C borrow rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >actual &&\n    ++\t\ttest_cmp expect actual\n    ++\t'\n    + done\n    + cd \"$base\"\n    + \n2:  6e06a82905 ! 2:  00a693dd72 setup: detangle loading of loose object maps\n    @@ Commit message\n         hash is non-zero. This setup sequence has two major downsides:\n     \n           - We assume that the primary object database is the \"files\" object\n    -        database so that we can extract its \"loose\" backend. This stops\n    -        working with pluggable object databases.\n    +        database and unconditionally downcast it. This will cause us to BUG\n    +        in case a different object database type was used together with a\n    +        compat hash algorithm.\n     \n           - We require the object database to already have been initialized when\n             configuring the object database. This means that we must intermix\n3:  183ed0f34c = 3:  1dc1f83d73 setup: defer object database creation\n4:  eb997d22d7 ! 4:  de1555ee1f odb/source: introduce function to map source type to name\n    @@ odb/source.c\n     +\t[ODB_SOURCE_FILES] = \"files\",\n     +\t[ODB_SOURCE_LOOSE] = \"loose\",\n     +\t[ODB_SOURCE_PACKED] = \"packed\",\n    -+\t[ODB_SOURCE_INMEMORY] = \"inmemory\",\n    ++\t[ODB_SOURCE_INMEMORY] = \"in-memory\",\n     +};\n     +\n     +const char *odb_source_type_to_name(enum odb_source_type type)\n5:  3303124a7d = 5:  cadf131e70 odb: make creation of on-disk structures pluggable\n\n---\nbase-commit: 9a0c4701dcd5725c4184599322b52933ff5005ca\nchange-id: 20260710-pks-odb-create-on-disk-ae8757861c69\n\n"},{"id":"549540","messageId":"20260804-pks-odb-create-on-disk-v2-1-ddf8b59bd207@pks.im","threadId":"66056","inReplyTo":"20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im","subject":"[PATCH v2 1/5] loose: load loose object map for the correct source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T08:29:06Z","receivedAt":"2026-08-04T08:29:58Z","isPatch":true,"body":"When loading the loose object map via `load_one_loose_object_map()` we\npass in both a repository and the corresponding source. We ultimately\ndon't really respect the passed-in source though as we instead always\nload the map via the common directory. This doesn't make any sense\nthough, as the function is called in a loop through all sources, and as\nsuch the expectation is that we'll load the map that belongs to the\ngiven source. The consequence is that we'll ignore loose object maps of\nany configured alternates.\n\nFix this bug by instead loading the map via the loose source's path.\n\nHelped-by: Toon Claes <toon@iotcl.com>\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c                       | 18 ++++++++++--------\n t/t1016-compatObjectFormat.sh | 18 ++++++++++++++++++\n 2 files changed, 28 insertions(+), 8 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex bf01d3e42d..9dad75373b 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,9 +61,11 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct repository *repo, struct odb_source_loose *loose)\n+static int load_one_loose_object_map(struct odb_source_loose *loose)\n {\n-\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\tstruct repository *repo = loose->base.odb->repo;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar *path;\n \tFILE *fp;\n \tint ret = -1;\n \n@@ -78,10 +80,10 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n \tinsert_loose_map(loose, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n \tinsert_loose_map(loose, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n \n-\trepo_common_path_replace(repo, &path, \"objects/loose-object-idx\");\n-\tfp = fopen(path.buf, \"rb\");\n+\tpath = xstrfmt(\"%s/loose-object-idx\", loose->base.path);\n+\tfp = fopen(path, \"rb\");\n \tif (!fp) {\n-\t\tstrbuf_release(&path);\n+\t\tfree(path);\n \t\treturn 0;\n \t}\n \n@@ -102,7 +104,7 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n err:\n \tfclose(fp);\n \tstrbuf_release(&buf);\n-\tstrbuf_release(&path);\n+\tfree(path);\n \treturn ret;\n }\n \n@@ -117,10 +119,10 @@ int repo_read_loose_object_map(struct repository *repo)\n \n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(repo, files->loose) < 0) {\n+\t\tif (load_one_loose_object_map(files->loose) < 0)\n \t\t\treturn -1;\n-\t\t}\n \t}\n+\n \treturn 0;\n }\n \ndiff --git a/t/t1016-compatObjectFormat.sh b/t/t1016-compatObjectFormat.sh\nindex 92d48b96a1..9cafcee509 100755\n--- a/t/t1016-compatObjectFormat.sh\n+++ b/t/t1016-compatObjectFormat.sh\n@@ -187,6 +187,24 @@ do\n \t\teval signedtag3_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag3) &&\n \t\teval signedtag4_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag4)\n \t'\n+\n+\ttest_expect_success 'rev-parse maps oid of object borrowed from alternate' '\n+\t\tfor repo in alt borrow\n+\t\tdo\n+\t\t\ttest_when_finished \"rm -rf $repo\" &&\n+\t\t\tgit init --object-format=$hash $repo &&\n+\t\t\tgit -C $repo config set core.repositoryformatversion 1 &&\n+\t\t\tgit -C $repo config set extensions.compatObjectFormat $(compat_hash $hash) || exit 1\n+\t\tdone &&\n+\n+\t\tgit -C alt commit --allow-empty --message A &&\n+\t\techo \"$(pwd)/alt/.git/objects\" >borrow/.git/objects/info/alternates &&\n+\n+\t\toid=$(git -C alt rev-parse HEAD) &&\n+\t\tgit -C alt    rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >expect &&\n+\t\tgit -C borrow rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >actual &&\n+\t\ttest_cmp expect actual\n+\t'\n done\n cd \"$base\"\n \n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549541","messageId":"20260804-pks-odb-create-on-disk-v2-2-ddf8b59bd207@pks.im","threadId":"66056","inReplyTo":"20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im","subject":"[PATCH v2 2/5] setup: detangle loading of loose object maps","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T08:29:07Z","receivedAt":"2026-08-04T08:30:03Z","isPatch":true,"body":"When a repository is configured to use a compatibility hash function\nthen we load the loose object map when we initialize the repository.\nThis object map provides the mappings between the canonical object hash\nand the compatibility object hash.\n\nLoading the object map happens in `repo_set_compat_hash_algo()`, which\ncalls `repo_read_loose_object_map()` in case the compatibility object\nhash is non-zero. This setup sequence has two major downsides:\n\n  - We assume that the primary object database is the \"files\" object\n    database and unconditionally downcast it. This will cause us to BUG\n    in case a different object database type was used together with a\n    compat hash algorithm.\n\n  - We require the object database to already have been initialized when\n    configuring the object database. This means that we must intermix\n    configuration of the repository and initialization of its\n    sub-structures in a weird way.\n\nRefactor the logic so that we instead load the loose object map via the\n\"loose\" backend, which fixes both of the above issues.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c            | 11 +++++------\n loose.h            |  1 +\n odb/source-loose.c |  2 ++\n repository.c       |  2 --\n setup.c            |  5 +++--\n 5 files changed, 11 insertions(+), 10 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 9dad75373b..a3b2dcedc2 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,7 +61,7 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct odb_source_loose *loose)\n+int loose_object_map_load(struct odb_source_loose *loose)\n {\n \tstruct repository *repo = loose->base.odb->repo;\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -69,6 +69,9 @@ static int load_one_loose_object_map(struct odb_source_loose *loose)\n \tFILE *fp;\n \tint ret = -1;\n \n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n \tif (!loose->map)\n \t\tloose_object_map_init(&loose->map);\n \tif (!loose->cache) {\n@@ -112,14 +115,10 @@ int repo_read_loose_object_map(struct repository *repo)\n {\n \tstruct odb_source *source;\n \n-\tif (!should_use_loose_object_map(repo))\n-\t\treturn 0;\n-\n \todb_prepare_alternates(repo->objects);\n-\n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(files->loose) < 0)\n+\t\tif (loose_object_map_load(files->loose) < 0)\n \t\t\treturn -1;\n \t}\n \ndiff --git a/loose.h b/loose.h\nindex 6c9b3f4571..ed663ac550 100644\n--- a/loose.h\n+++ b/loose.h\n@@ -13,6 +13,7 @@ struct loose_object_map {\n \n void loose_object_map_init(struct loose_object_map **map);\n void loose_object_map_clear(struct loose_object_map **map);\n+int loose_object_map_load(struct odb_source_loose *loose);\n int repo_loose_object_map_oid(struct repository *repo,\n \t\t\t      const struct object_id *src,\n \t\t\t      const struct git_hash_algo *dest_algo,\ndiff --git a/odb/source-loose.c b/odb/source-loose.c\nindex 3f7d04a56e..812ca1c138 100644\n--- a/odb/source-loose.c\n+++ b/odb/source-loose.c\n@@ -727,5 +727,7 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n \tif (!is_absolute_path(loose->base.path))\n \t\tchdir_notify_register(NULL, odb_source_loose_reparent, loose);\n \n+\tloose_object_map_load(loose);\n+\n \treturn loose;\n }\ndiff --git a/repository.c b/repository.c\nindex 2ef0778846..6d633002b4 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -201,8 +201,6 @@ void repo_set_compat_hash_algo(struct repository *repo MAYBE_UNUSED, uint32_t al\n \tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n \t\tBUG(\"hash_algo and compat_hash_algo match\");\n \trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n-\tif (repo->compat_hash_algo)\n-\t\trepo_read_loose_object_map(repo);\n #else\n \tif (algo)\n \t\tdie(_(\"compatibility hash algorithm support requires Rust\"));\ndiff --git a/setup.c b/setup.c\nindex d31808130b..825572f5f1 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1788,8 +1788,6 @@ int apply_repository_format(struct repository *repo,\n \n \trepo->bare_cfg = format->is_bare;\n \trepo_set_hash_algo(repo, format->hash_algo);\n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n \trepo_set_compat_hash_algo(repo, format->compat_hash_algo);\n \trepo_set_ref_storage_format(repo,\n \t\t\t\t    format->ref_storage_format,\n@@ -1805,6 +1803,9 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n+\n \tfree(alternate_object_directories);\n \tfree(object_directory);\n \treturn 0;\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549542","messageId":"20260804-pks-odb-create-on-disk-v2-3-ddf8b59bd207@pks.im","threadId":"66056","inReplyTo":"20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im","subject":"[PATCH v2 3/5] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T08:29:08Z","receivedAt":"2026-08-04T08:30:07Z","isPatch":true,"body":"In a subsequent commit we'll make the creation of the on-disk data\nstructures of an object database pluggable. This will lead to an\nin-between state where we have already configured the repository's\nobject database, but it's not usable yet until we eventually call\n`create_object_directory()`.\n\nDefer the object database creation so that we handle both steps in the\nsame function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n setup.c | 35 +++++++++++++++++++++++++++--------\n setup.h |  9 +++++++++\n 2 files changed, 36 insertions(+), 8 deletions(-)\n\ndiff --git a/setup.c b/setup.c\nindex 825572f5f1..a7b1b9eaef 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1760,6 +1760,13 @@ enum discovery_result discover_git_directory_reason(struct strbuf *commondir,\n \treturn result;\n }\n \n+static void get_object_directories(char **object_directory,\n+\t\t\t\t   char **alternate_object_directories)\n+{\n+\t*object_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n+\t*alternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n+}\n+\n int apply_repository_format(struct repository *repo,\n \t\t\t    const struct repository_format *format,\n \t\t\t    enum apply_repository_format_flags flags,\n@@ -1779,8 +1786,9 @@ int apply_repository_format(struct repository *repo,\n \tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV) {\n \t\tconst char *shallow_file;\n \n-\t\tobject_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n-\t\talternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n+\t\tget_object_directories(&object_directory,\n+\t\t\t\t       &alternate_object_directories);\n+\n \t\tshallow_file = getenv(GIT_SHALLOW_FILE_ENVIRONMENT);\n \t\tif (shallow_file)\n \t\t\tset_alternate_shallow_file(repo, shallow_file);\n@@ -1803,8 +1811,9 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n+\tif (!(flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION))\n+\t\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\t\talternate_object_directories);\n \n \tfree(alternate_object_directories);\n \tfree(object_directory);\n@@ -2654,11 +2663,16 @@ static int create_default_files(struct repository *repo,\n \treturn reinit;\n }\n \n-static void create_object_directory(struct repository *repo)\n+static void create_object_database(struct repository *repo)\n {\n+\tchar *object_directory, *alternate_object_directories;\n \tstruct strbuf path = STRBUF_INIT;\n \tsize_t baselen;\n \n+\tget_object_directories(&object_directory, &alternate_object_directories);\n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n+\n \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n \tbaselen = path.len;\n \n@@ -2672,6 +2686,8 @@ static void create_object_directory(struct repository *repo)\n \tstrbuf_addstr(&path, \"/info\");\n \tsafe_create_dir(repo, path.buf, 1);\n \n+\tfree(alternate_object_directories);\n+\tfree(object_directory);\n \tstrbuf_release(&path);\n }\n \n@@ -2867,9 +2883,10 @@ int init_db(struct repository *repo,\n \t */\n \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n-\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n+\tif (apply_repository_format(repo, &repo_fmt,\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV |\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION, &err) < 0)\n \t\tdie(\"%s\", err.buf);\n-\tstartup_info->have_repository = 1;\n \n \t/*\n \t * Ensure `core.hidedotfiles` is processed. This must happen after we\n@@ -2885,7 +2902,9 @@ int init_db(struct repository *repo,\n \n \tif (!(flags & INIT_DB_SKIP_REFDB))\n \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n-\tcreate_object_directory(repo);\n+\tcreate_object_database(repo);\n+\n+\tstartup_info->have_repository = 1;\n \n \tif (repo_settings_get_shared_repository(repo)) {\n \t\tchar buf[10];\ndiff --git a/setup.h b/setup.h\nindex 654f10e059..e55d647b70 100644\n--- a/setup.h\n+++ b/setup.h\n@@ -241,6 +241,15 @@ enum apply_repository_format_flags {\n \t * relate to the object database.\n \t */\n \tAPPLY_REPOSITORY_FORMAT_HONOR_ENV = (1 << 0),\n+\n+\t/*\n+\t * Usually, the object database is created after the repository format\n+\t * was applied. This step is skipped if this flag is set, which leaves\n+\t * us with a partially-working repository.\n+\t *\n+\t * This is useful when initializing a new repository.\n+\t */\n+\tAPPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION = (1 << 1),\n };\n \n /*\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549543","messageId":"20260804-pks-odb-create-on-disk-v2-4-ddf8b59bd207@pks.im","threadId":"66056","inReplyTo":"20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im","subject":"[PATCH v2 4/5] odb/source: introduce function to map source type to name","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T08:29:09Z","receivedAt":"2026-08-04T08:30:10Z","isPatch":true,"body":"Introduce a new function that maps an object source's type to a\nhuman-readable name. Use the function to provide better human-readable\nerror messages for the downcasting functions.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.h    |  4 +++-\n odb/source-inmemory.h |  4 +++-\n odb/source-loose.h    |  4 +++-\n odb/source-packed.h   |  4 +++-\n odb/source.c          | 19 +++++++++++++++++++\n odb/source.h          |  6 ++++++\n 6 files changed, 37 insertions(+), 4 deletions(-)\n\ndiff --git a/odb/source-files.h b/odb/source-files.h\nindex d7ac3c1c81..6a803afdda 100644\n--- a/odb/source-files.h\n+++ b/odb/source-files.h\n@@ -28,7 +28,9 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n static inline struct odb_source_files *odb_source_files_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_FILES)\n-\t\tBUG(\"trying to downcast source of type '%d' to files\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_FILES));\n \treturn container_of(source, struct odb_source_files, base);\n }\n \ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex a88fc2e320..adbad23e8b 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -26,7 +26,9 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_INMEMORY)\n-\t\tBUG(\"trying to downcast source of type '%d' to in-memory\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_INMEMORY));\n \treturn container_of(source, struct odb_source_inmemory, base);\n }\n \ndiff --git a/odb/source-loose.h b/odb/source-loose.h\nindex 6070aaf3ce..3cf2e1f8f1 100644\n--- a/odb/source-loose.h\n+++ b/odb/source-loose.h\n@@ -41,7 +41,9 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n static inline struct odb_source_loose *odb_source_loose_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_LOOSE)\n-\t\tBUG(\"trying to downcast source of type '%d' to loose\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_LOOSE));\n \treturn container_of(source, struct odb_source_loose, base);\n }\n \ndiff --git a/odb/source-packed.h b/odb/source-packed.h\nindex 77309ddd09..a0f6b5096d 100644\n--- a/odb/source-packed.h\n+++ b/odb/source-packed.h\n@@ -78,7 +78,9 @@ struct odb_source_packed *odb_source_packed_new(struct object_database *odb,\n static inline struct odb_source_packed *odb_source_packed_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_PACKED)\n-\t\tBUG(\"trying to downcast source of type '%d' to packed\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_PACKED));\n \treturn container_of(source, struct odb_source_packed, base);\n }\n \ndiff --git a/odb/source.c b/odb/source.c\nindex 7993dcbd65..30188b806d 100644\n--- a/odb/source.c\n+++ b/odb/source.c\n@@ -4,6 +4,25 @@\n #include \"odb/source.h\"\n #include \"packfile.h\"\n \n+static const char * const odb_source_names_by_type[] = {\n+\t[ODB_SOURCE_UNKNOWN] = \"unknown\",\n+\t[ODB_SOURCE_FILES] = \"files\",\n+\t[ODB_SOURCE_LOOSE] = \"loose\",\n+\t[ODB_SOURCE_PACKED] = \"packed\",\n+\t[ODB_SOURCE_INMEMORY] = \"in-memory\",\n+};\n+\n+const char *odb_source_type_to_name(enum odb_source_type type)\n+{\n+\tconst char *name;\n+\tif (type < 0 || type >= ARRAY_SIZE(odb_source_names_by_type))\n+\t\ttype = ODB_SOURCE_UNKNOWN;\n+\tname = odb_source_names_by_type[type];\n+\tif (!name)\n+\t\tBUG(\"name missing in `odb_source_names_by_type` for '%d'\", type);\n+\treturn name;\n+}\n+\n struct odb_source *odb_source_new(struct object_database *odb,\n \t\t\t\t  const char *path,\n \t\t\t\t  bool local)\ndiff --git a/odb/source.h b/odb/source.h\nindex cd63dba91f..ab16d152f4 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -25,6 +25,12 @@ enum odb_source_type {\n \tODB_SOURCE_INMEMORY,\n };\n \n+/*\n+ * Convert between the enum and its name. Returns the equivalent of \"unknown\"\n+ * for unknown types.\n+ */\n+const char *odb_source_type_to_name(enum odb_source_type type);\n+\n struct object_id;\n struct odb_read_stream;\n struct strvec;\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549544","messageId":"20260804-pks-odb-create-on-disk-v2-5-ddf8b59bd207@pks.im","threadId":"66056","inReplyTo":"20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im","subject":"[PATCH v2 5/5] odb: make creation of on-disk structures pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-04T08:29:10Z","receivedAt":"2026-08-04T08:30:14Z","isPatch":true,"body":"When creating a new \"files\" object database source we have to create a\ncouple of directories. These directories are of course specific to this\nparticular backend, and a different backend may require a setup that is\ncompletely different.\n\nMake the creation of on-disk structures pluggable to accommodate for\nthis.\n\nNote that there is one exception though: the \"objects\" directory must\nexist in a repository regardless of which backend is in use. If it\ndoesn't exist then the repository is not treated as a Git repository at\nall. Consequently, we create this directory regardless of the backend.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.c | 19 +++++++++++++++++++\n odb/source.h       | 23 +++++++++++++++++++++++\n setup.c            | 35 ++++++++++++++++++++---------------\n 3 files changed, 62 insertions(+), 15 deletions(-)\n\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex 4138758511..0db6e681fe 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -9,6 +9,7 @@\n #include \"odb/source-files.h\"\n #include \"odb/source-loose.h\"\n #include \"packfile.h\"\n+#include \"path.h\"\n #include \"strbuf.h\"\n #include \"write-or-die.h\"\n \n@@ -41,6 +42,23 @@ static void odb_source_files_close(struct odb_source *source)\n \todb_source_close(&files->packed->base);\n }\n \n+static int odb_source_files_create_on_disk(struct odb_source *source)\n+{\n+\tstruct strbuf path = STRBUF_INIT;\n+\n+\tsafe_create_dir(source->odb->repo, source->path, 1);\n+\n+\tstrbuf_addf(&path, \"%s/pack\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_reset(&path);\n+\tstrbuf_addf(&path, \"%s/info\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_release(&path);\n+\treturn 0;\n+}\n+\n static void odb_source_files_prepare(struct odb_source *source,\n \t\t\t\t     enum odb_prepare_flags flags)\n {\n@@ -271,6 +289,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \n \tfiles->base.free = odb_source_files_free;\n \tfiles->base.close = odb_source_files_close;\n+\tfiles->base.create_on_disk = odb_source_files_create_on_disk;\n \tfiles->base.prepare = odb_source_files_prepare;\n \tfiles->base.read_object_info = odb_source_files_read_object_info;\n \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\ndiff --git a/odb/source.h b/odb/source.h\nindex ab16d152f4..4abc418bdd 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -89,6 +89,18 @@ struct odb_source {\n \t */\n \tvoid (*close)(struct odb_source *source);\n \n+\t/*\n+\t * This callback is expected to create on-disk data structures that are\n+\t * required for this source to operate.\n+\t *\n+\t * The callback is expected to return 0 on success, a negative error\n+\t * code otherwise.\n+\t *\n+\t * This callback may be NULL in case the source does not need any\n+\t * on-disk setup.\n+\t */\n+\tint (*create_on_disk)(struct odb_source *source);\n+\n \t/*\n \t * This callback is expected to prepare the source so that it becomes\n \t * ready for use. It optionally clears underlying caches of the object\n@@ -316,6 +328,17 @@ static inline void odb_source_close(struct odb_source *source)\n \tsource->close(source);\n }\n \n+/*\n+ * Create on-disk data structures that are required for this source to operate\n+ * correctly. Returns 0 on success, a negative error code otherwise.\n+ */\n+static inline int odb_source_create_on_disk(struct odb_source *source)\n+{\n+\tif (!source->create_on_disk)\n+\t\treturn 0;\n+\treturn source->create_on_disk(source);\n+}\n+\n /*\n  * Prepare the object database source and clear any caches. Depending on the\n  * backend used this may have the effect that concurrently-written objects\ndiff --git a/setup.c b/setup.c\nindex a7b1b9eaef..14ef119cb7 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -2666,29 +2666,34 @@ static int create_default_files(struct repository *repo,\n static void create_object_database(struct repository *repo)\n {\n \tchar *object_directory, *alternate_object_directories;\n-\tstruct strbuf path = STRBUF_INIT;\n-\tsize_t baselen;\n \n \tget_object_directories(&object_directory, &alternate_object_directories);\n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n \n-\tstrbuf_addstr(&path, repo_get_object_directory(repo));\n-\tbaselen = path.len;\n-\n-\tsafe_create_dir(repo, path.buf, 1);\n+\t/*\n+\t * Create the \"objects\" directory in the common directory. This is done\n+\t * so that the repository can be discovered regardless of the backend\n+\t * used.\n+\t *\n+\t * Note that we only do this in case the object directory wasn't\n+\t * overwritten via an environment variable. If it _is_ being overridden\n+\t * then we skip this step, as the repository won't be discoverable\n+\t * anyway without the environment variable.\n+\t */\n+\tif (!object_directory) {\n+\t\tstruct strbuf objects_dir = STRBUF_INIT;\n+\t\trepo_common_path_append(repo, &objects_dir, \"objects\");\n+\t\tsafe_create_dir(repo, objects_dir.buf, 1);\n+\t\tstrbuf_release(&objects_dir);\n+\t}\n \n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/pack\");\n-\tsafe_create_dir(repo, path.buf, 1);\n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n \n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/info\");\n-\tsafe_create_dir(repo, path.buf, 1);\n+\tif (odb_source_create_on_disk(repo->objects->sources) < 0)\n+\t\tdie(\"failed creating object database\");\n \n \tfree(alternate_object_directories);\n \tfree(object_directory);\n-\tstrbuf_release(&path);\n }\n \n static void separate_git_dir(const char *git_dir, const char *git_link)\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549577","messageId":"anIU0ivwnjn026wa@denethor","threadId":"66056","inReplyTo":"20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im","subject":"Re: [PATCH v2 0/5] odb: make creation of object database pluggable","fromName":"Justin Tobler","fromEmail":"jltobler@gmail.com","sentAt":"2026-08-04T16:36:11Z","receivedAt":"2026-08-04T16:36:17Z","isPatch":true,"body":"On 26/08/04 10:29AM, Patrick Steinhardt wrote:\n> Changes in v2:\n>   - Add a testcase that demonstrates the bug fixed with alternate loose\n>     object maps.\n>   - Rename the \"inmemory\" bakcend to \"in-memory\".\n>   - Clarify some commit messages.\n>   - Link to v1: https://patch.msgid.link/20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im\n\nFrom the range-diff, this version of the series looks good to me.\n\n-Justin\n"},{"id":"549601","messageId":"87bjbh67sl.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"20260804-pks-odb-create-on-disk-v2-3-ddf8b59bd207@pks.im","subject":"Re: [PATCH v2 3/5] setup: defer object database creation","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-08-04T18:48:42Z","receivedAt":"2026-08-04T18:48:50Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> In a subsequent commit we'll make the creation of the on-disk data\n> structures of an object database pluggable. This will lead to an\n> in-between state where we have already configured the repository's\n> object database, but it's not usable yet until we eventually call\n> `create_object_directory()`.\n>\n> Defer the object database creation so that we handle both steps in the\n> same function.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  setup.c | 35 +++++++++++++++++++++++++++--------\n>  setup.h |  9 +++++++++\n>  2 files changed, 36 insertions(+), 8 deletions(-)\n>\n> diff --git a/setup.c b/setup.c\n> index 825572f5f1..a7b1b9eaef 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -1760,6 +1760,13 @@ enum discovery_result discover_git_directory_reason(struct strbuf *commondir,\n>  \treturn result;\n>  }\n>  \n> +static void get_object_directories(char **object_directory,\n> +\t\t\t\t   char **alternate_object_directories)\n> +{\n> +\t*object_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n> +\t*alternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n> +}\n\nWould it make sense to wrap these in a APPLY_REPOSITORY_FORMAT_HONOR_ENV\nguard?\n\nI mean, below we call this function *only* when flags has that bit set.\nBut the return values of that function are used at the bottom of\napply_repository_format(), that's a bit awkard.\n\nSo can I suggest the following patch instead? That would remove the\nweird double pointer passing around, which feels a bit unneeded.\n\n\n--- >8 ---\nSubject: [PATCH] setup: defer object database creation\n\nIn a subsequent commit we'll make the creation of the on-disk data\nstructures of an object database pluggable. This will lead to an\nin-between state where we have already configured the repository's\nobject database, but it's not usable yet until we eventually call\n`create_object_directory()`.\n\nDefer the object database creation so that we handle both steps in the\nsame function.\n\nSigned-off-by: Toon Claes <toon@iotcl.com>\n---\n setup.c | 35 +++++++++++++++++++++++++++--------\n setup.h |  9 +++++++++\n 2 files changed, 36 insertions(+), 8 deletions(-)\n\ndiff --git a/setup.c b/setup.c\nindex 825572f5f1..2e9bc92481 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1760,13 +1760,28 @@ enum discovery_result discover_git_directory_reason(struct strbuf *commondir,\n \treturn result;\n }\n \n+static void setup_objects_odb_new(struct repository *repo,\n+\t\t\t\t  bool from_env)\n+{\n+\tchar *object_directory = NULL, *alternate_object_directories = NULL;\n+\n+\tif (from_env) {\n+\t\tobject_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n+\t\talternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n+\t}\n+\n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n+\n+\tfree(alternate_object_directories);\n+\tfree(object_directory);\n+}\n+\n int apply_repository_format(struct repository *repo,\n \t\t\t    const struct repository_format *format,\n \t\t\t    enum apply_repository_format_flags flags,\n \t\t\t    struct strbuf *err)\n {\n-\tchar *object_directory = NULL, *alternate_object_directories = NULL;\n-\n \tif (verify_repository_format(format, err) < 0)\n \t\treturn -1;\n \n@@ -1779,8 +1794,6 @@ int apply_repository_format(struct repository *repo,\n \tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV) {\n \t\tconst char *shallow_file;\n \n-\t\tobject_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n-\t\talternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n \t\tshallow_file = getenv(GIT_SHALLOW_FILE_ENVIRONMENT);\n \t\tif (shallow_file)\n \t\t\tset_alternate_shallow_file(repo, shallow_file);\n@@ -1803,11 +1816,11 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n+\tif (flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION)\n+\t\treturn 0;\n+\n+\tsetup_objects_odb_new(repo, flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV);\n \n-\tfree(alternate_object_directories);\n-\tfree(object_directory);\n \treturn 0;\n }\n \n@@ -2654,11 +2667,13 @@ static int create_default_files(struct repository *repo,\n \treturn reinit;\n }\n \n-static void create_object_directory(struct repository *repo)\n+static void create_object_database(struct repository *repo)\n {\n \tstruct strbuf path = STRBUF_INIT;\n \tsize_t baselen;\n \n+\tsetup_objects_odb_new(repo, true);\n+\n \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n \tbaselen = path.len;\n \n@@ -2867,9 +2882,10 @@ int init_db(struct repository *repo,\n \t */\n \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n-\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n+\tif (apply_repository_format(repo, &repo_fmt,\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV |\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION, &err) < 0)\n \t\tdie(\"%s\", err.buf);\n-\tstartup_info->have_repository = 1;\n \n \t/*\n \t * Ensure `core.hidedotfiles` is processed. This must happen after we\n@@ -2885,7 +2901,9 @@ int init_db(struct repository *repo,\n \n \tif (!(flags & INIT_DB_SKIP_REFDB))\n \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n-\tcreate_object_directory(repo);\n+\tcreate_object_database(repo);\n+\n+\tstartup_info->have_repository = 1;\n \n \tif (repo_settings_get_shared_repository(repo)) {\n \t\tchar buf[10];\ndiff --git a/setup.h b/setup.h\nindex 654f10e059..e55d647b70 100644\n--- a/setup.h\n+++ b/setup.h\n@@ -241,6 +241,15 @@ enum apply_repository_format_flags {\n \t * relate to the object database.\n \t */\n \tAPPLY_REPOSITORY_FORMAT_HONOR_ENV = (1 << 0),\n+\n+\t/*\n+\t * Usually, the object database is created after the repository format\n+\t * was applied. This step is skipped if this flag is set, which leaves\n+\t * us with a partially-working repository.\n+\t *\n+\t * This is useful when initializing a new repository.\n+\t */\n+\tAPPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION = (1 << 1),\n };\n \n /*\n-- \n2.55.0.629.g250fe7f194\n\n"},{"id":"549647","messageId":"anLl8Cy6Bkv5XA7-@pks.im","threadId":"66056","inReplyTo":"87bjbh67sl.fsf@emacs.iotcl.com","subject":"Re: [PATCH v2 3/5] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T07:27:59Z","receivedAt":"2026-08-05T07:28:12Z","isPatch":true,"body":"On Tue, Aug 04, 2026 at 08:48:42PM +0200, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > In a subsequent commit we'll make the creation of the on-disk data\n> > structures of an object database pluggable. This will lead to an\n> > in-between state where we have already configured the repository's\n> > object database, but it's not usable yet until we eventually call\n> > `create_object_directory()`.\n> >\n> > Defer the object database creation so that we handle both steps in the\n> > same function.\n> >\n> > Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> > ---\n> >  setup.c | 35 +++++++++++++++++++++++++++--------\n> >  setup.h |  9 +++++++++\n> >  2 files changed, 36 insertions(+), 8 deletions(-)\n> >\n> > diff --git a/setup.c b/setup.c\n> > index 825572f5f1..a7b1b9eaef 100644\n> > --- a/setup.c\n> > +++ b/setup.c\n> > @@ -1760,6 +1760,13 @@ enum discovery_result discover_git_directory_reason(struct strbuf *commondir,\n> >  \treturn result;\n> >  }\n> >  \n> > +static void get_object_directories(char **object_directory,\n> > +\t\t\t\t   char **alternate_object_directories)\n> > +{\n> > +\t*object_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n> > +\t*alternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n> > +}\n> \n> Would it make sense to wrap these in a APPLY_REPOSITORY_FORMAT_HONOR_ENV\n> guard?\n> \n> I mean, below we call this function *only* when flags has that bit set.\n> But the return values of that function are used at the bottom of\n> apply_repository_format(), that's a bit awkard.\n> \n> So can I suggest the following patch instead? That would remove the\n> weird double pointer passing around, which feels a bit unneeded.\n\nYou're right, this is somewhat awkward. I have a different proposal\nthough: instead of creating a separate function, we can move handling of\nenvironment variables into `odb_new()` itself. This also paves the way\nfor moving handling of these environment variables into the backend,\nwhich is something I want to do soonish [1].\n\nPatrick\n\n[1]: https://lore.kernel.org/git/amLgMqkqxR8mKIbT@pks.im/\n"},{"id":"549669","messageId":"20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","subject":"[PATCH v3 0/6] odb: make creation of object database pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T09:28:50Z","receivedAt":"2026-08-05T09:29:00Z","isPatch":true,"body":"Hi,\n\nwhen creating a new repository we create a couple of on-disk data\nstructures for the object database. This includes the \"objects/\"\ndirectory hierarchy with \"objects/info\" and \"objects/pack\", which are\nspecific to the backend.\n\nThis patch series makes the creation of the on-disk data structures\npluggable. While we continue to always create \"objects/\" regardless of\nthe backend (it's required for a repository to be recognized as such),\nthe other subdirectories are now created by the backend. This will allow\nother backends to plug in their own logic.\n\nThe series starts with a small detour into the loose-object map. This\ndetour is required so that we can defer initialization of the object\ndatabase itself to a later point in time.\n\nThe series is based on 9a0c4701dc (The 7th batch, 2026-07-22).\n\nChanges in v3:\n  - Move handling of GIT_OBJECT_DIRECTORY and\n    GIT_ALTERNATE_OBJECT_DIRECTORIES into `odb_new()` itself. This\n    deduplicates some of the logic and also preps us for a future where\n    alternates are handled in the \"files\" backend itself.\n  - Link to v2: https://patch.msgid.link/20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im\n\nChanges in v2:\n  - Add a testcase that demonstrates the bug fixed with alternate loose\n    object maps.\n  - Rename the \"inmemory\" bakcend to \"in-memory\".\n  - Clarify some commit messages.\n  - Link to v1: https://patch.msgid.link/20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (6):\n      loose: load loose object map for the correct source\n      setup: detangle loading of loose object maps\n      setup: handle ODB-related environment variables in `odb_new()`\n      setup: defer object database creation\n      odb/source: introduce function to map source type to name\n      odb: make creation of on-disk structures pluggable\n\n loose.c                       | 25 +++++++++---------\n loose.h                       |  1 +\n odb.c                         | 20 +++++++++------\n odb.h                         | 17 ++++++++++--\n odb/source-files.c            | 19 ++++++++++++++\n odb/source-files.h            |  4 ++-\n odb/source-inmemory.h         |  4 ++-\n odb/source-loose.c            |  2 ++\n odb/source-loose.h            |  4 ++-\n odb/source-packed.h           |  4 ++-\n odb/source.c                  | 19 ++++++++++++++\n odb/source.h                  | 29 +++++++++++++++++++++\n repository.c                  |  2 --\n setup.c                       | 60 ++++++++++++++++++++++++-------------------\n setup.h                       |  9 +++++++\n t/t1016-compatObjectFormat.sh | 18 +++++++++++++\n t/unit-tests/u-odb-inmemory.c |  2 +-\n 17 files changed, 183 insertions(+), 56 deletions(-)\n\nRange-diff versus v2:\n\n1:  b0beb61a74 = 1:  d384dd0635 loose: load loose object map for the correct source\n2:  097bdcad14 = 2:  0ee1b3c032 setup: detangle loading of loose object maps\n-:  ---------- > 3:  f52992b9bd setup: handle ODB-related environment variables in `odb_new()`\n3:  06645224ef ! 4:  4524fc5ec4 setup: defer object database creation\n    @@ Commit message\n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n     \n      ## setup.c ##\n    -@@ setup.c: enum discovery_result discover_git_directory_reason(struct strbuf *commondir,\n    - \treturn result;\n    - }\n    - \n    -+static void get_object_directories(char **object_directory,\n    -+\t\t\t\t   char **alternate_object_directories)\n    -+{\n    -+\t*object_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n    -+\t*alternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n    -+}\n    -+\n    - int apply_repository_format(struct repository *repo,\n    - \t\t\t    const struct repository_format *format,\n    - \t\t\t    enum apply_repository_format_flags flags,\n     @@ setup.c: int apply_repository_format(struct repository *repo,\n    - \tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV) {\n    - \t\tconst char *shallow_file;\n    + \t\t\t    enum apply_repository_format_flags flags,\n    + \t\t\t    struct strbuf *err)\n    + {\n    +-\tenum odb_new_flags odb_new_flags = 0;\n    +-\n    + \tif (verify_repository_format(format, err) < 0)\n    + \t\treturn -1;\n      \n    --\t\tobject_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n    --\t\talternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n    -+\t\tget_object_directories(&object_directory,\n    -+\t\t\t\t       &alternate_object_directories);\n    -+\n    - \t\tshallow_file = getenv(GIT_SHALLOW_FILE_ENVIRONMENT);\n    - \t\tif (shallow_file)\n    - \t\t\tset_alternate_shallow_file(repo, shallow_file);\n     @@ setup.c: int apply_repository_format(struct repository *repo,\n      \trepo->repository_format_precious_objects =\n      \t\tformat->precious_objects;\n      \n    --\trepo->objects = odb_new(repo, object_directory,\n    --\t\t\t\talternate_object_directories);\n    -+\tif (!(flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION))\n    -+\t\trepo->objects = odb_new(repo, object_directory,\n    -+\t\t\t\t\talternate_object_directories);\n    +-\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n    +-\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n    +-\trepo->objects = odb_new(repo, odb_new_flags);\n    ++\tif (!(flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION)) {\n    ++\t\tenum odb_new_flags odb_new_flags = 0;\n    ++\t\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n    ++\t\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n    ++\t\trepo->objects = odb_new(repo, odb_new_flags);\n    ++\t}\n      \n    - \tfree(alternate_object_directories);\n    - \tfree(object_directory);\n    + \treturn 0;\n    + }\n     @@ setup.c: static int create_default_files(struct repository *repo,\n      \treturn reinit;\n      }\n    @@ setup.c: static int create_default_files(struct repository *repo,\n     -static void create_object_directory(struct repository *repo)\n     +static void create_object_database(struct repository *repo)\n      {\n    -+\tchar *object_directory, *alternate_object_directories;\n      \tstruct strbuf path = STRBUF_INIT;\n      \tsize_t baselen;\n      \n    -+\tget_object_directories(&object_directory, &alternate_object_directories);\n    -+\trepo->objects = odb_new(repo, object_directory,\n    -+\t\t\t\talternate_object_directories);\n    ++\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n     +\n      \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n      \tbaselen = path.len;\n      \n    -@@ setup.c: static void create_object_directory(struct repository *repo)\n    - \tstrbuf_addstr(&path, \"/info\");\n    - \tsafe_create_dir(repo, path.buf, 1);\n    - \n    -+\tfree(alternate_object_directories);\n    -+\tfree(object_directory);\n    - \tstrbuf_release(&path);\n    - }\n    - \n     @@ setup.c: int init_db(struct repository *repo,\n      \t */\n      \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n4:  46ad0386bb = 5:  c526fd526b odb/source: introduce function to map source type to name\n5:  3063325cf9 ! 6:  d752e48eba odb: make creation of on-disk structures pluggable\n    @@ odb/source.h: static inline void odb_source_close(struct odb_source *source)\n     \n      ## setup.c ##\n     @@ setup.c: static int create_default_files(struct repository *repo,\n    + \n      static void create_object_database(struct repository *repo)\n      {\n    - \tchar *object_directory, *alternate_object_directories;\n     -\tstruct strbuf path = STRBUF_INIT;\n     -\tsize_t baselen;\n    - \n    - \tget_object_directories(&object_directory, &alternate_object_directories);\n    --\trepo->objects = odb_new(repo, object_directory,\n    --\t\t\t\talternate_object_directories);\n    - \n    --\tstrbuf_addstr(&path, repo_get_object_directory(repo));\n    --\tbaselen = path.len;\n    --\n    --\tsafe_create_dir(repo, path.buf, 1);\n     +\t/*\n     +\t * Create the \"objects\" directory in the common directory. This is done\n     +\t * so that the repository can be discovered regardless of the backend\n    @@ setup.c: static int create_default_files(struct repository *repo,\n     +\t * then we skip this step, as the repository won't be discoverable\n     +\t * anyway without the environment variable.\n     +\t */\n    -+\tif (!object_directory) {\n    ++\tif (!getenv(DB_ENVIRONMENT)) {\n     +\t\tstruct strbuf objects_dir = STRBUF_INIT;\n     +\t\trepo_common_path_append(repo, &objects_dir, \"objects\");\n     +\t\tsafe_create_dir(repo, objects_dir.buf, 1);\n     +\t\tstrbuf_release(&objects_dir);\n     +\t}\n      \n    + \trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n    + \n    +-\tstrbuf_addstr(&path, repo_get_object_directory(repo));\n    +-\tbaselen = path.len;\n    +-\n    +-\tsafe_create_dir(repo, path.buf, 1);\n    +-\n     -\tstrbuf_setlen(&path, baselen);\n     -\tstrbuf_addstr(&path, \"/pack\");\n     -\tsafe_create_dir(repo, path.buf, 1);\n    -+\trepo->objects = odb_new(repo, object_directory,\n    -+\t\t\t\talternate_object_directories);\n    - \n    +-\n     -\tstrbuf_setlen(&path, baselen);\n     -\tstrbuf_addstr(&path, \"/info\");\n     -\tsafe_create_dir(repo, path.buf, 1);\n    +-\n    +-\tstrbuf_release(&path);\n     +\tif (odb_source_create_on_disk(repo->objects->sources) < 0)\n     +\t\tdie(\"failed creating object database\");\n    - \n    - \tfree(alternate_object_directories);\n    - \tfree(object_directory);\n    --\tstrbuf_release(&path);\n      }\n      \n      static void separate_git_dir(const char *git_dir, const char *git_link)\n\n---\nbase-commit: 9a0c4701dcd5725c4184599322b52933ff5005ca\nchange-id: 20260710-pks-odb-create-on-disk-ae8757861c69\n\n"},{"id":"549670","messageId":"20260805-pks-odb-create-on-disk-v3-1-c0ee3ac5141f@pks.im","threadId":"66056","inReplyTo":"20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im","subject":"[PATCH v3 1/6] loose: load loose object map for the correct source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T09:28:51Z","receivedAt":"2026-08-05T09:29:01Z","isPatch":true,"body":"When loading the loose object map via `load_one_loose_object_map()` we\npass in both a repository and the corresponding source. We ultimately\ndon't really respect the passed-in source though as we instead always\nload the map via the common directory. This doesn't make any sense\nthough, as the function is called in a loop through all sources, and as\nsuch the expectation is that we'll load the map that belongs to the\ngiven source. The consequence is that we'll ignore loose object maps of\nany configured alternates.\n\nFix this bug by instead loading the map via the loose source's path.\n\nHelped-by: Toon Claes <toon@iotcl.com>\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c                       | 18 ++++++++++--------\n t/t1016-compatObjectFormat.sh | 18 ++++++++++++++++++\n 2 files changed, 28 insertions(+), 8 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex bf01d3e42d..9dad75373b 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,9 +61,11 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct repository *repo, struct odb_source_loose *loose)\n+static int load_one_loose_object_map(struct odb_source_loose *loose)\n {\n-\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\tstruct repository *repo = loose->base.odb->repo;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar *path;\n \tFILE *fp;\n \tint ret = -1;\n \n@@ -78,10 +80,10 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n \tinsert_loose_map(loose, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n \tinsert_loose_map(loose, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n \n-\trepo_common_path_replace(repo, &path, \"objects/loose-object-idx\");\n-\tfp = fopen(path.buf, \"rb\");\n+\tpath = xstrfmt(\"%s/loose-object-idx\", loose->base.path);\n+\tfp = fopen(path, \"rb\");\n \tif (!fp) {\n-\t\tstrbuf_release(&path);\n+\t\tfree(path);\n \t\treturn 0;\n \t}\n \n@@ -102,7 +104,7 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n err:\n \tfclose(fp);\n \tstrbuf_release(&buf);\n-\tstrbuf_release(&path);\n+\tfree(path);\n \treturn ret;\n }\n \n@@ -117,10 +119,10 @@ int repo_read_loose_object_map(struct repository *repo)\n \n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(repo, files->loose) < 0) {\n+\t\tif (load_one_loose_object_map(files->loose) < 0)\n \t\t\treturn -1;\n-\t\t}\n \t}\n+\n \treturn 0;\n }\n \ndiff --git a/t/t1016-compatObjectFormat.sh b/t/t1016-compatObjectFormat.sh\nindex 92d48b96a1..9cafcee509 100755\n--- a/t/t1016-compatObjectFormat.sh\n+++ b/t/t1016-compatObjectFormat.sh\n@@ -187,6 +187,24 @@ do\n \t\teval signedtag3_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag3) &&\n \t\teval signedtag4_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag4)\n \t'\n+\n+\ttest_expect_success 'rev-parse maps oid of object borrowed from alternate' '\n+\t\tfor repo in alt borrow\n+\t\tdo\n+\t\t\ttest_when_finished \"rm -rf $repo\" &&\n+\t\t\tgit init --object-format=$hash $repo &&\n+\t\t\tgit -C $repo config set core.repositoryformatversion 1 &&\n+\t\t\tgit -C $repo config set extensions.compatObjectFormat $(compat_hash $hash) || exit 1\n+\t\tdone &&\n+\n+\t\tgit -C alt commit --allow-empty --message A &&\n+\t\techo \"$(pwd)/alt/.git/objects\" >borrow/.git/objects/info/alternates &&\n+\n+\t\toid=$(git -C alt rev-parse HEAD) &&\n+\t\tgit -C alt    rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >expect &&\n+\t\tgit -C borrow rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >actual &&\n+\t\ttest_cmp expect actual\n+\t'\n done\n cd \"$base\"\n \n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549671","messageId":"20260805-pks-odb-create-on-disk-v3-2-c0ee3ac5141f@pks.im","threadId":"66056","inReplyTo":"20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im","subject":"[PATCH v3 2/6] setup: detangle loading of loose object maps","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T09:28:52Z","receivedAt":"2026-08-05T09:29:04Z","isPatch":true,"body":"When a repository is configured to use a compatibility hash function\nthen we load the loose object map when we initialize the repository.\nThis object map provides the mappings between the canonical object hash\nand the compatibility object hash.\n\nLoading the object map happens in `repo_set_compat_hash_algo()`, which\ncalls `repo_read_loose_object_map()` in case the compatibility object\nhash is non-zero. This setup sequence has two major downsides:\n\n  - We assume that the primary object database is the \"files\" object\n    database and unconditionally downcast it. This will cause us to BUG\n    in case a different object database type was used together with a\n    compat hash algorithm.\n\n  - We require the object database to already have been initialized when\n    configuring the object database. This means that we must intermix\n    configuration of the repository and initialization of its\n    sub-structures in a weird way.\n\nRefactor the logic so that we instead load the loose object map via the\n\"loose\" backend, which fixes both of the above issues.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c            | 11 +++++------\n loose.h            |  1 +\n odb/source-loose.c |  2 ++\n repository.c       |  2 --\n setup.c            |  5 +++--\n 5 files changed, 11 insertions(+), 10 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 9dad75373b..a3b2dcedc2 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,7 +61,7 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct odb_source_loose *loose)\n+int loose_object_map_load(struct odb_source_loose *loose)\n {\n \tstruct repository *repo = loose->base.odb->repo;\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -69,6 +69,9 @@ static int load_one_loose_object_map(struct odb_source_loose *loose)\n \tFILE *fp;\n \tint ret = -1;\n \n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n \tif (!loose->map)\n \t\tloose_object_map_init(&loose->map);\n \tif (!loose->cache) {\n@@ -112,14 +115,10 @@ int repo_read_loose_object_map(struct repository *repo)\n {\n \tstruct odb_source *source;\n \n-\tif (!should_use_loose_object_map(repo))\n-\t\treturn 0;\n-\n \todb_prepare_alternates(repo->objects);\n-\n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(files->loose) < 0)\n+\t\tif (loose_object_map_load(files->loose) < 0)\n \t\t\treturn -1;\n \t}\n \ndiff --git a/loose.h b/loose.h\nindex 6c9b3f4571..ed663ac550 100644\n--- a/loose.h\n+++ b/loose.h\n@@ -13,6 +13,7 @@ struct loose_object_map {\n \n void loose_object_map_init(struct loose_object_map **map);\n void loose_object_map_clear(struct loose_object_map **map);\n+int loose_object_map_load(struct odb_source_loose *loose);\n int repo_loose_object_map_oid(struct repository *repo,\n \t\t\t      const struct object_id *src,\n \t\t\t      const struct git_hash_algo *dest_algo,\ndiff --git a/odb/source-loose.c b/odb/source-loose.c\nindex 3f7d04a56e..812ca1c138 100644\n--- a/odb/source-loose.c\n+++ b/odb/source-loose.c\n@@ -727,5 +727,7 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n \tif (!is_absolute_path(loose->base.path))\n \t\tchdir_notify_register(NULL, odb_source_loose_reparent, loose);\n \n+\tloose_object_map_load(loose);\n+\n \treturn loose;\n }\ndiff --git a/repository.c b/repository.c\nindex 2ef0778846..6d633002b4 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -201,8 +201,6 @@ void repo_set_compat_hash_algo(struct repository *repo MAYBE_UNUSED, uint32_t al\n \tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n \t\tBUG(\"hash_algo and compat_hash_algo match\");\n \trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n-\tif (repo->compat_hash_algo)\n-\t\trepo_read_loose_object_map(repo);\n #else\n \tif (algo)\n \t\tdie(_(\"compatibility hash algorithm support requires Rust\"));\ndiff --git a/setup.c b/setup.c\nindex d31808130b..825572f5f1 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1788,8 +1788,6 @@ int apply_repository_format(struct repository *repo,\n \n \trepo->bare_cfg = format->is_bare;\n \trepo_set_hash_algo(repo, format->hash_algo);\n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n \trepo_set_compat_hash_algo(repo, format->compat_hash_algo);\n \trepo_set_ref_storage_format(repo,\n \t\t\t\t    format->ref_storage_format,\n@@ -1805,6 +1803,9 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n+\n \tfree(alternate_object_directories);\n \tfree(object_directory);\n \treturn 0;\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549672","messageId":"20260805-pks-odb-create-on-disk-v3-3-c0ee3ac5141f@pks.im","threadId":"66056","inReplyTo":"20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im","subject":"[PATCH v3 3/6] setup: handle ODB-related environment variables in `odb_new()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T09:28:53Z","receivedAt":"2026-08-05T09:29:08Z","isPatch":true,"body":"When initializing a repository's object database we have to respect the\nGIT_OBJECT_DIRECTORY and GIT_ALTERNATE_OBJECT_DIRECTORIES environment\nvariables, which can be set by the user to override the default location\nof where we write objects to and read objects from.\n\nThis is handled in `apply_repository_format()`, which is fine. But in a\nsubsequent commit we'll have to defer constructing the object database\nto a later point in some cases, and that will require a second site\nwhere we call `odb_new()`. And of course, that second site would have to\nhandle those environment variables, as well.\n\nIt would be somewhat awkward to duplicate the logic though. But there's\na better alternative: instead of handling this logic in \"setup.c\", we\ncan easily handle environment variables in `odb_new()` itself. This\nensures that object database creation is neatly self-contained, and we\ndon't have to duplicate any of the logic.\n\nAnother benefit is that in a future patch series we plan to move\nhandling of alternates into the backends themselves [1], and that will\nrequire us to also handle those environment variables in the \"files\"\nbackend itself. So moving the logic into the ODB level already gets us\none step closer to that goal.\n\nRefactor the logic accordingly.\n\n[1]: https://lore.kernel.org/git/amLgMqkqxR8mKIbT@pks.im/\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                         | 20 ++++++++++++--------\n odb.h                         | 17 +++++++++++++++--\n setup.c                       | 11 ++++-------\n t/unit-tests/u-odb-inmemory.c |  2 +-\n 4 files changed, 32 insertions(+), 18 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex cf6e7938c0..b463afa072 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -1004,26 +1004,30 @@ int odb_write_object_stream(struct object_database *odb,\n }\n \n struct object_database *odb_new(struct repository *repo,\n-\t\t\t\tconst char *primary_source,\n-\t\t\t\tconst char *secondary_sources)\n+\t\t\t\tenum odb_new_flags flags)\n {\n-\tstruct object_database *o = xmalloc(sizeof(*o));\n-\tchar *to_free = NULL;\n+\tchar *primary_source = NULL, *secondary_sources = NULL;\n+\tstruct object_database *o;\n \n-\tmemset(o, 0, sizeof(*o));\n+\tCALLOC_ARRAY(o, 1);\n \to->repo = repo;\n \tpthread_mutex_init(&o->replace_mutex, NULL);\n \tstring_list_init_dup(&o->submodule_source_paths);\n \n+\tif (flags & ODB_NEW_HONOR_ENV) {\n+\t\tprimary_source = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n+\t\tsecondary_sources = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n+\t}\n \tif (!primary_source)\n-\t\tprimary_source = to_free = xstrfmt(\"%s/objects\", repo->commondir);\n+\t\tprimary_source = xstrfmt(\"%s/objects\", repo->commondir);\n+\n \to->sources = odb_source_new(o, primary_source, true);\n \to->sources_tail = &o->sources->next;\n \to->alternate_db = xstrdup_or_null(secondary_sources);\n \to->inmemory_objects = &odb_source_inmemory_new(o)->base;\n \n-\tfree(to_free);\n-\n+\tfree(secondary_sources);\n+\tfree(primary_source);\n \treturn o;\n }\n \ndiff --git a/odb.h b/odb.h\nindex 7995bed97b..8ec335c7f7 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -100,6 +100,20 @@ struct object_database {\n \tstruct string_list submodule_source_paths;\n };\n \n+enum odb_new_flags {\n+\t/*\n+\t * Honor environment variables when constructing the object database\n+\t * sources. This makes us respect the following environment variables:\n+\t *\n+\t *   - GIT_OBJECT_DIRECTORY to override the primary object directory.\n+\t *\n+\t *   - GIT_ALTERNATE_OBJECT_DIRECTORIES to override alternates.\n+\t *\n+\t * Environment variables may be backend-specific.\n+\t */\n+\tODB_NEW_HONOR_ENV = (1 << 0),\n+};\n+\n /*\n  * Create a new object database for the given repository.\n  *\n@@ -112,8 +126,7 @@ struct object_database {\n  * Returns the newly created object database.\n  */\n struct object_database *odb_new(struct repository *repo,\n-\t\t\t\tconst char *primary_source,\n-\t\t\t\tconst char *alternate_sources);\n+\t\t\t\tenum odb_new_flags flags);\n \n /* Free the object database and release all resources. */\n void odb_free(struct object_database *o);\ndiff --git a/setup.c b/setup.c\nindex 825572f5f1..5dfab3e79e 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1765,7 +1765,7 @@ int apply_repository_format(struct repository *repo,\n \t\t\t    enum apply_repository_format_flags flags,\n \t\t\t    struct strbuf *err)\n {\n-\tchar *object_directory = NULL, *alternate_object_directories = NULL;\n+\tenum odb_new_flags odb_new_flags = 0;\n \n \tif (verify_repository_format(format, err) < 0)\n \t\treturn -1;\n@@ -1779,8 +1779,6 @@ int apply_repository_format(struct repository *repo,\n \tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV) {\n \t\tconst char *shallow_file;\n \n-\t\tobject_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n-\t\talternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n \t\tshallow_file = getenv(GIT_SHALLOW_FILE_ENVIRONMENT);\n \t\tif (shallow_file)\n \t\t\tset_alternate_shallow_file(repo, shallow_file);\n@@ -1803,11 +1801,10 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n+\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n+\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n+\trepo->objects = odb_new(repo, odb_new_flags);\n \n-\tfree(alternate_object_directories);\n-\tfree(object_directory);\n \treturn 0;\n }\n \ndiff --git a/t/unit-tests/u-odb-inmemory.c b/t/unit-tests/u-odb-inmemory.c\nindex 6844bfc37c..db323e10fd 100644\n--- a/t/unit-tests/u-odb-inmemory.c\n+++ b/t/unit-tests/u-odb-inmemory.c\n@@ -38,7 +38,7 @@ static void cl_assert_object_info(struct odb_source_inmemory *source,\n \n void test_odb_inmemory__initialize(void)\n {\n-\todb = odb_new(&repo, \"\", \"\");\n+\todb = odb_new(&repo, 0);\n }\n \n void test_odb_inmemory__cleanup(void)\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549673","messageId":"20260805-pks-odb-create-on-disk-v3-4-c0ee3ac5141f@pks.im","threadId":"66056","inReplyTo":"20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im","subject":"[PATCH v3 4/6] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T09:28:54Z","receivedAt":"2026-08-05T09:29:10Z","isPatch":true,"body":"In a subsequent commit we'll make the creation of the on-disk data\nstructures of an object database pluggable. This will lead to an\nin-between state where we have already configured the repository's\nobject database, but it's not usable yet until we eventually call\n`create_object_directory()`.\n\nDefer the object database creation so that we handle both steps in the\nsame function.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n setup.c | 24 +++++++++++++++---------\n setup.h |  9 +++++++++\n 2 files changed, 24 insertions(+), 9 deletions(-)\n\ndiff --git a/setup.c b/setup.c\nindex 5dfab3e79e..d85171f3b6 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1765,8 +1765,6 @@ int apply_repository_format(struct repository *repo,\n \t\t\t    enum apply_repository_format_flags flags,\n \t\t\t    struct strbuf *err)\n {\n-\tenum odb_new_flags odb_new_flags = 0;\n-\n \tif (verify_repository_format(format, err) < 0)\n \t\treturn -1;\n \n@@ -1801,9 +1799,12 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n-\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n-\trepo->objects = odb_new(repo, odb_new_flags);\n+\tif (!(flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION)) {\n+\t\tenum odb_new_flags odb_new_flags = 0;\n+\t\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n+\t\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n+\t\trepo->objects = odb_new(repo, odb_new_flags);\n+\t}\n \n \treturn 0;\n }\n@@ -2651,11 +2652,13 @@ static int create_default_files(struct repository *repo,\n \treturn reinit;\n }\n \n-static void create_object_directory(struct repository *repo)\n+static void create_object_database(struct repository *repo)\n {\n \tstruct strbuf path = STRBUF_INIT;\n \tsize_t baselen;\n \n+\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n+\n \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n \tbaselen = path.len;\n \n@@ -2864,9 +2867,10 @@ int init_db(struct repository *repo,\n \t */\n \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n-\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n+\tif (apply_repository_format(repo, &repo_fmt,\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV |\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION, &err) < 0)\n \t\tdie(\"%s\", err.buf);\n-\tstartup_info->have_repository = 1;\n \n \t/*\n \t * Ensure `core.hidedotfiles` is processed. This must happen after we\n@@ -2882,7 +2886,9 @@ int init_db(struct repository *repo,\n \n \tif (!(flags & INIT_DB_SKIP_REFDB))\n \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n-\tcreate_object_directory(repo);\n+\tcreate_object_database(repo);\n+\n+\tstartup_info->have_repository = 1;\n \n \tif (repo_settings_get_shared_repository(repo)) {\n \t\tchar buf[10];\ndiff --git a/setup.h b/setup.h\nindex 654f10e059..e55d647b70 100644\n--- a/setup.h\n+++ b/setup.h\n@@ -241,6 +241,15 @@ enum apply_repository_format_flags {\n \t * relate to the object database.\n \t */\n \tAPPLY_REPOSITORY_FORMAT_HONOR_ENV = (1 << 0),\n+\n+\t/*\n+\t * Usually, the object database is created after the repository format\n+\t * was applied. This step is skipped if this flag is set, which leaves\n+\t * us with a partially-working repository.\n+\t *\n+\t * This is useful when initializing a new repository.\n+\t */\n+\tAPPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION = (1 << 1),\n };\n \n /*\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549674","messageId":"20260805-pks-odb-create-on-disk-v3-5-c0ee3ac5141f@pks.im","threadId":"66056","inReplyTo":"20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im","subject":"[PATCH v3 5/6] odb/source: introduce function to map source type to name","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T09:28:55Z","receivedAt":"2026-08-05T09:29:13Z","isPatch":true,"body":"Introduce a new function that maps an object source's type to a\nhuman-readable name. Use the function to provide better human-readable\nerror messages for the downcasting functions.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.h    |  4 +++-\n odb/source-inmemory.h |  4 +++-\n odb/source-loose.h    |  4 +++-\n odb/source-packed.h   |  4 +++-\n odb/source.c          | 19 +++++++++++++++++++\n odb/source.h          |  6 ++++++\n 6 files changed, 37 insertions(+), 4 deletions(-)\n\ndiff --git a/odb/source-files.h b/odb/source-files.h\nindex d7ac3c1c81..6a803afdda 100644\n--- a/odb/source-files.h\n+++ b/odb/source-files.h\n@@ -28,7 +28,9 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n static inline struct odb_source_files *odb_source_files_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_FILES)\n-\t\tBUG(\"trying to downcast source of type '%d' to files\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_FILES));\n \treturn container_of(source, struct odb_source_files, base);\n }\n \ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex a88fc2e320..adbad23e8b 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -26,7 +26,9 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_INMEMORY)\n-\t\tBUG(\"trying to downcast source of type '%d' to in-memory\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_INMEMORY));\n \treturn container_of(source, struct odb_source_inmemory, base);\n }\n \ndiff --git a/odb/source-loose.h b/odb/source-loose.h\nindex 6070aaf3ce..3cf2e1f8f1 100644\n--- a/odb/source-loose.h\n+++ b/odb/source-loose.h\n@@ -41,7 +41,9 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n static inline struct odb_source_loose *odb_source_loose_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_LOOSE)\n-\t\tBUG(\"trying to downcast source of type '%d' to loose\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_LOOSE));\n \treturn container_of(source, struct odb_source_loose, base);\n }\n \ndiff --git a/odb/source-packed.h b/odb/source-packed.h\nindex 77309ddd09..a0f6b5096d 100644\n--- a/odb/source-packed.h\n+++ b/odb/source-packed.h\n@@ -78,7 +78,9 @@ struct odb_source_packed *odb_source_packed_new(struct object_database *odb,\n static inline struct odb_source_packed *odb_source_packed_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_PACKED)\n-\t\tBUG(\"trying to downcast source of type '%d' to packed\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_PACKED));\n \treturn container_of(source, struct odb_source_packed, base);\n }\n \ndiff --git a/odb/source.c b/odb/source.c\nindex 7993dcbd65..30188b806d 100644\n--- a/odb/source.c\n+++ b/odb/source.c\n@@ -4,6 +4,25 @@\n #include \"odb/source.h\"\n #include \"packfile.h\"\n \n+static const char * const odb_source_names_by_type[] = {\n+\t[ODB_SOURCE_UNKNOWN] = \"unknown\",\n+\t[ODB_SOURCE_FILES] = \"files\",\n+\t[ODB_SOURCE_LOOSE] = \"loose\",\n+\t[ODB_SOURCE_PACKED] = \"packed\",\n+\t[ODB_SOURCE_INMEMORY] = \"in-memory\",\n+};\n+\n+const char *odb_source_type_to_name(enum odb_source_type type)\n+{\n+\tconst char *name;\n+\tif (type < 0 || type >= ARRAY_SIZE(odb_source_names_by_type))\n+\t\ttype = ODB_SOURCE_UNKNOWN;\n+\tname = odb_source_names_by_type[type];\n+\tif (!name)\n+\t\tBUG(\"name missing in `odb_source_names_by_type` for '%d'\", type);\n+\treturn name;\n+}\n+\n struct odb_source *odb_source_new(struct object_database *odb,\n \t\t\t\t  const char *path,\n \t\t\t\t  bool local)\ndiff --git a/odb/source.h b/odb/source.h\nindex cd63dba91f..ab16d152f4 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -25,6 +25,12 @@ enum odb_source_type {\n \tODB_SOURCE_INMEMORY,\n };\n \n+/*\n+ * Convert between the enum and its name. Returns the equivalent of \"unknown\"\n+ * for unknown types.\n+ */\n+const char *odb_source_type_to_name(enum odb_source_type type);\n+\n struct object_id;\n struct odb_read_stream;\n struct strvec;\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549675","messageId":"20260805-pks-odb-create-on-disk-v3-6-c0ee3ac5141f@pks.im","threadId":"66056","inReplyTo":"20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im","subject":"[PATCH v3 6/6] odb: make creation of on-disk structures pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T09:28:56Z","receivedAt":"2026-08-05T09:29:16Z","isPatch":true,"body":"When creating a new \"files\" object database source we have to create a\ncouple of directories. These directories are of course specific to this\nparticular backend, and a different backend may require a setup that is\ncompletely different.\n\nMake the creation of on-disk structures pluggable to accommodate for\nthis.\n\nNote that there is one exception though: the \"objects\" directory must\nexist in a repository regardless of which backend is in use. If it\ndoesn't exist then the repository is not treated as a Git repository at\nall. Consequently, we create this directory regardless of the backend.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.c | 19 +++++++++++++++++++\n odb/source.h       | 23 +++++++++++++++++++++++\n setup.c            | 34 ++++++++++++++++++----------------\n 3 files changed, 60 insertions(+), 16 deletions(-)\n\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex 4138758511..0db6e681fe 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -9,6 +9,7 @@\n #include \"odb/source-files.h\"\n #include \"odb/source-loose.h\"\n #include \"packfile.h\"\n+#include \"path.h\"\n #include \"strbuf.h\"\n #include \"write-or-die.h\"\n \n@@ -41,6 +42,23 @@ static void odb_source_files_close(struct odb_source *source)\n \todb_source_close(&files->packed->base);\n }\n \n+static int odb_source_files_create_on_disk(struct odb_source *source)\n+{\n+\tstruct strbuf path = STRBUF_INIT;\n+\n+\tsafe_create_dir(source->odb->repo, source->path, 1);\n+\n+\tstrbuf_addf(&path, \"%s/pack\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_reset(&path);\n+\tstrbuf_addf(&path, \"%s/info\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_release(&path);\n+\treturn 0;\n+}\n+\n static void odb_source_files_prepare(struct odb_source *source,\n \t\t\t\t     enum odb_prepare_flags flags)\n {\n@@ -271,6 +289,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \n \tfiles->base.free = odb_source_files_free;\n \tfiles->base.close = odb_source_files_close;\n+\tfiles->base.create_on_disk = odb_source_files_create_on_disk;\n \tfiles->base.prepare = odb_source_files_prepare;\n \tfiles->base.read_object_info = odb_source_files_read_object_info;\n \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\ndiff --git a/odb/source.h b/odb/source.h\nindex ab16d152f4..4abc418bdd 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -89,6 +89,18 @@ struct odb_source {\n \t */\n \tvoid (*close)(struct odb_source *source);\n \n+\t/*\n+\t * This callback is expected to create on-disk data structures that are\n+\t * required for this source to operate.\n+\t *\n+\t * The callback is expected to return 0 on success, a negative error\n+\t * code otherwise.\n+\t *\n+\t * This callback may be NULL in case the source does not need any\n+\t * on-disk setup.\n+\t */\n+\tint (*create_on_disk)(struct odb_source *source);\n+\n \t/*\n \t * This callback is expected to prepare the source so that it becomes\n \t * ready for use. It optionally clears underlying caches of the object\n@@ -316,6 +328,17 @@ static inline void odb_source_close(struct odb_source *source)\n \tsource->close(source);\n }\n \n+/*\n+ * Create on-disk data structures that are required for this source to operate\n+ * correctly. Returns 0 on success, a negative error code otherwise.\n+ */\n+static inline int odb_source_create_on_disk(struct odb_source *source)\n+{\n+\tif (!source->create_on_disk)\n+\t\treturn 0;\n+\treturn source->create_on_disk(source);\n+}\n+\n /*\n  * Prepare the object database source and clear any caches. Depending on the\n  * backend used this may have the effect that concurrently-written objects\ndiff --git a/setup.c b/setup.c\nindex d85171f3b6..af02cd965c 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -2654,25 +2654,27 @@ static int create_default_files(struct repository *repo,\n \n static void create_object_database(struct repository *repo)\n {\n-\tstruct strbuf path = STRBUF_INIT;\n-\tsize_t baselen;\n+\t/*\n+\t * Create the \"objects\" directory in the common directory. This is done\n+\t * so that the repository can be discovered regardless of the backend\n+\t * used.\n+\t *\n+\t * Note that we only do this in case the object directory wasn't\n+\t * overwritten via an environment variable. If it _is_ being overridden\n+\t * then we skip this step, as the repository won't be discoverable\n+\t * anyway without the environment variable.\n+\t */\n+\tif (!getenv(DB_ENVIRONMENT)) {\n+\t\tstruct strbuf objects_dir = STRBUF_INIT;\n+\t\trepo_common_path_append(repo, &objects_dir, \"objects\");\n+\t\tsafe_create_dir(repo, objects_dir.buf, 1);\n+\t\tstrbuf_release(&objects_dir);\n+\t}\n \n \trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n \n-\tstrbuf_addstr(&path, repo_get_object_directory(repo));\n-\tbaselen = path.len;\n-\n-\tsafe_create_dir(repo, path.buf, 1);\n-\n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/pack\");\n-\tsafe_create_dir(repo, path.buf, 1);\n-\n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/info\");\n-\tsafe_create_dir(repo, path.buf, 1);\n-\n-\tstrbuf_release(&path);\n+\tif (odb_source_create_on_disk(repo->objects->sources) < 0)\n+\t\tdie(\"failed creating object database\");\n }\n \n static void separate_git_dir(const char *git_dir, const char *git_link)\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549695","messageId":"878q6k66ha.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"20260805-pks-odb-create-on-disk-v3-3-c0ee3ac5141f@pks.im","subject":"Re: [PATCH v3 3/6] setup: handle ODB-related environment variables in `odb_new()`","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-08-05T13:29:21Z","receivedAt":"2026-08-05T13:29:31Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> When initializing a repository's object database we have to respect the\n> GIT_OBJECT_DIRECTORY and GIT_ALTERNATE_OBJECT_DIRECTORIES environment\n> variables, which can be set by the user to override the default location\n> of where we write objects to and read objects from.\n>\n> This is handled in `apply_repository_format()`, which is fine. But in a\n> subsequent commit we'll have to defer constructing the object database\n> to a later point in some cases, and that will require a second site\n> where we call `odb_new()`. And of course, that second site would have to\n> handle those environment variables, as well.\n>\n> It would be somewhat awkward to duplicate the logic though. But there's\n> a better alternative: instead of handling this logic in \"setup.c\", we\n> can easily handle environment variables in `odb_new()` itself. This\n> ensures that object database creation is neatly self-contained, and we\n> don't have to duplicate any of the logic.\n>\n> Another benefit is that in a future patch series we plan to move\n> handling of alternates into the backends themselves [1], and that will\n> require us to also handle those environment variables in the \"files\"\n> backend itself. So moving the logic into the ODB level already gets us\n> one step closer to that goal.\n>\n> Refactor the logic accordingly.\n\nI like this!\n\n>\n> [1]: https://lore.kernel.org/git/amLgMqkqxR8mKIbT@pks.im/\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb.c                         | 20 ++++++++++++--------\n>  odb.h                         | 17 +++++++++++++++--\n>  setup.c                       | 11 ++++-------\n>  t/unit-tests/u-odb-inmemory.c |  2 +-\n>  4 files changed, 32 insertions(+), 18 deletions(-)\n>\n> diff --git a/odb.c b/odb.c\n> index cf6e7938c0..b463afa072 100644\n> --- a/odb.c\n> +++ b/odb.c\n> @@ -1004,26 +1004,30 @@ int odb_write_object_stream(struct object_database *odb,\n>  }\n>  \n>  struct object_database *odb_new(struct repository *repo,\n> -\t\t\t\tconst char *primary_source,\n> -\t\t\t\tconst char *secondary_sources)\n> +\t\t\t\tenum odb_new_flags flags)\n>  {\n> -\tstruct object_database *o = xmalloc(sizeof(*o));\n> -\tchar *to_free = NULL;\n> +\tchar *primary_source = NULL, *secondary_sources = NULL;\n> +\tstruct object_database *o;\n>  \n> -\tmemset(o, 0, sizeof(*o));\n> +\tCALLOC_ARRAY(o, 1);\n>  \to->repo = repo;\n>  \tpthread_mutex_init(&o->replace_mutex, NULL);\n>  \tstring_list_init_dup(&o->submodule_source_paths);\n>  \n> +\tif (flags & ODB_NEW_HONOR_ENV) {\n> +\t\tprimary_source = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n> +\t\tsecondary_sources = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n> +\t}\n>  \tif (!primary_source)\n> -\t\tprimary_source = to_free = xstrfmt(\"%s/objects\", repo->commondir);\n> +\t\tprimary_source = xstrfmt(\"%s/objects\", repo->commondir);\n> +\n>  \to->sources = odb_source_new(o, primary_source, true);\n>  \to->sources_tail = &o->sources->next;\n>  \to->alternate_db = xstrdup_or_null(secondary_sources);\n\nI'd say this xstrdup_or_null() is not needed no more, and so is the\nfree() of that variable below.\n\n>  \to->inmemory_objects = &odb_source_inmemory_new(o)->base;\n>  \n> -\tfree(to_free);\n> -\n> +\tfree(secondary_sources);\n> +\tfree(primary_source);\n>  \treturn o;\n>  }\n\n-- \nCheers,\nToon\n"},{"id":"549701","messageId":"8733ws6424.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"20260805-pks-odb-create-on-disk-v3-4-c0ee3ac5141f@pks.im","subject":"Re: [PATCH v3 4/6] setup: defer object database creation","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-08-05T14:21:39Z","receivedAt":"2026-08-05T14:21:52Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> In a subsequent commit we'll make the creation of the on-disk data\n> structures of an object database pluggable. This will lead to an\n> in-between state where we have already configured the repository's\n> object database, but it's not usable yet until we eventually call\n> `create_object_directory()`.\n>\n> Defer the object database creation so that we handle both steps in the\n> same function.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  setup.c | 24 +++++++++++++++---------\n>  setup.h |  9 +++++++++\n>  2 files changed, 24 insertions(+), 9 deletions(-)\n>\n> diff --git a/setup.c b/setup.c\n> index 5dfab3e79e..d85171f3b6 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -1765,8 +1765,6 @@ int apply_repository_format(struct repository *repo,\n>  \t\t\t    enum apply_repository_format_flags flags,\n>  \t\t\t    struct strbuf *err)\n>  {\n> -\tenum odb_new_flags odb_new_flags = 0;\n> -\n>  \tif (verify_repository_format(format, err) < 0)\n>  \t\treturn -1;\n>  \n> @@ -1801,9 +1799,12 @@ int apply_repository_format(struct repository *repo,\n>  \trepo->repository_format_precious_objects =\n>  \t\tformat->precious_objects;\n>  \n> -\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n> -\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n> -\trepo->objects = odb_new(repo, odb_new_flags);\n> +\tif (!(flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION)) {\n> +\t\tenum odb_new_flags odb_new_flags = 0;\n> +\t\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n> +\t\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n> +\t\trepo->objects = odb_new(repo, odb_new_flags);\n> +\t}\n>  \n>  \treturn 0;\n>  }\n> @@ -2651,11 +2652,13 @@ static int create_default_files(struct repository *repo,\n>  \treturn reinit;\n>  }\n>  \n> -static void create_object_directory(struct repository *repo)\n> +static void create_object_database(struct repository *repo)\n>  {\n>  \tstruct strbuf path = STRBUF_INIT;\n>  \tsize_t baselen;\n>  \n> +\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n> +\n>  \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n>  \tbaselen = path.len;\n>  \n> @@ -2864,9 +2867,10 @@ int init_db(struct repository *repo,\n>  \t */\n>  \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n>  \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n> -\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n> +\tif (apply_repository_format(repo, &repo_fmt,\n> +\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV |\n> +\t\t\t\t    APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION, &err) < 0)\n>  \t\tdie(\"%s\", err.buf);\n> -\tstartup_info->have_repository = 1;\n>  \n>  \t/*\n>  \t * Ensure `core.hidedotfiles` is processed. This must happen after we\n> @@ -2882,7 +2886,9 @@ int init_db(struct repository *repo,\n>  \n>  \tif (!(flags & INIT_DB_SKIP_REFDB))\n>  \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n> -\tcreate_object_directory(repo);\n> +\tcreate_object_database(repo);\n> +\n> +\tstartup_info->have_repository = 1;\n>  \n>  \tif (repo_settings_get_shared_repository(repo)) {\n>  \t\tchar buf[10];\n> diff --git a/setup.h b/setup.h\n> index 654f10e059..e55d647b70 100644\n> --- a/setup.h\n> +++ b/setup.h\n> @@ -241,6 +241,15 @@ enum apply_repository_format_flags {\n>  \t * relate to the object database.\n>  \t */\n>  \tAPPLY_REPOSITORY_FORMAT_HONOR_ENV = (1 << 0),\n> +\n> +\t/*\n> +\t * Usually, the object database is created after the repository format\n> +\t * was applied. This step is skipped if this flag is set, which leaves\n> +\t * us with a partially-working repository.\n> +\t *\n> +\t * This is useful when initializing a new repository.\n> +\t */\n> +\tAPPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION = (1 << 1),\n>  };\n>  \n>  /*\n>\n> -- \n> 2.55.0.679.g6767b8d81c.dirty\n>\n\nWith [PATCH v3 3/6], Justin's objection[1] is stronger now:\n\n> Naive question: would it be simpler to just require invoking `odb_new()`\n> explicitly after `apply_repository_format()` in all cases? There doesn't\n> appear to be too many callsites.\n\nAs a matter of fact, I've given this a try and see these changes on top\nof this series below.\n\n[1]: <amkXcmwzbBYsMgjc@denethor>\n\n--- >8 ---\n\ndiff --git a/repository.c b/repository.c\nindex 6d633002b4..9eee74113c 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -295,6 +295,8 @@ int repo_init(struct repository *repo,\n \t\tgoto error;\n \t}\n \n+\trepo->objects = odb_new(repo, 0);\n+\n \tif (worktree)\n \t\trepo_set_worktree(repo, worktree);\n \ndiff --git a/setup.c b/setup.c\nindex af02cd965c..1106f38bb0 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1799,13 +1799,6 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\tif (!(flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION)) {\n-\t\tenum odb_new_flags odb_new_flags = 0;\n-\t\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n-\t\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n-\t\trepo->objects = odb_new(repo, odb_new_flags);\n-\t}\n-\n \treturn 0;\n }\n \n@@ -1889,6 +1882,7 @@ const char *enter_repo(struct repository *repo, const char *path, unsigned flags\n \t\tread_and_verify_repository_format(&fmt, \".\", NULL);\n \t\tif (apply_repository_format(repo, &fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n \t\tstartup_info->have_repository = 1;\n \n \t\tclear_repository_format(&fmt);\n@@ -2092,6 +2086,8 @@ const char *setup_git_directory_gently(struct repository *repo, int *nongit_ok)\n \t\t\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n \t\t\t\tdie(\"%s\", err.buf);\n \n+\t\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n+\n \t\t\tclear_repository_format(&discovery.format);\n \t\t\tstrbuf_release(&err);\n \t\t}\n@@ -2870,8 +2866,7 @@ int init_db(struct repository *repo,\n \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n \tif (apply_repository_format(repo, &repo_fmt,\n-\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV |\n-\t\t\t\t    APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION, &err) < 0)\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n \t\tdie(\"%s\", err.buf);\n \n \t/*\ndiff --git a/setup.h b/setup.h\nindex e55d647b70..654f10e059 100644\n--- a/setup.h\n+++ b/setup.h\n@@ -241,15 +241,6 @@ enum apply_repository_format_flags {\n \t * relate to the object database.\n \t */\n \tAPPLY_REPOSITORY_FORMAT_HONOR_ENV = (1 << 0),\n-\n-\t/*\n-\t * Usually, the object database is created after the repository format\n-\t * was applied. This step is skipped if this flag is set, which leaves\n-\t * us with a partially-working repository.\n-\t *\n-\t * This is useful when initializing a new repository.\n-\t */\n-\tAPPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION = (1 << 1),\n };\n \n /*\n\n\n"},{"id":"549727","messageId":"87zez04l1x.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"20260805-pks-odb-create-on-disk-v3-6-c0ee3ac5141f@pks.im","subject":"Re: [PATCH v3 6/6] odb: make creation of on-disk structures pluggable","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-08-05T15:57:30Z","receivedAt":"2026-08-05T15:57:39Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> When creating a new \"files\" object database source we have to create a\n> couple of directories. These directories are of course specific to this\n> particular backend, and a different backend may require a setup that is\n> completely different.\n>\n> Make the creation of on-disk structures pluggable to accommodate for\n> this.\n>\n> Note that there is one exception though: the \"objects\" directory must\n> exist in a repository regardless of which backend is in use. If it\n> doesn't exist then the repository is not treated as a Git repository at\n> all. Consequently, we create this directory regardless of the backend.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  odb/source-files.c | 19 +++++++++++++++++++\n>  odb/source.h       | 23 +++++++++++++++++++++++\n>  setup.c            | 34 ++++++++++++++++++----------------\n>  3 files changed, 60 insertions(+), 16 deletions(-)\n>\n> diff --git a/odb/source-files.c b/odb/source-files.c\n> index 4138758511..0db6e681fe 100644\n> --- a/odb/source-files.c\n> +++ b/odb/source-files.c\n> @@ -9,6 +9,7 @@\n>  #include \"odb/source-files.h\"\n>  #include \"odb/source-loose.h\"\n>  #include \"packfile.h\"\n> +#include \"path.h\"\n>  #include \"strbuf.h\"\n>  #include \"write-or-die.h\"\n>  \n> @@ -41,6 +42,23 @@ static void odb_source_files_close(struct odb_source *source)\n>  \todb_source_close(&files->packed->base);\n>  }\n>  \n> +static int odb_source_files_create_on_disk(struct odb_source *source)\n> +{\n> +\tstruct strbuf path = STRBUF_INIT;\n> +\n> +\tsafe_create_dir(source->odb->repo, source->path, 1);\n> +\n> +\tstrbuf_addf(&path, \"%s/pack\", source->path);\n> +\tsafe_create_dir(source->odb->repo, path.buf, 1);\n> +\n> +\tstrbuf_reset(&path);\n> +\tstrbuf_addf(&path, \"%s/info\", source->path);\n> +\tsafe_create_dir(source->odb->repo, path.buf, 1);\n> +\n> +\tstrbuf_release(&path);\n> +\treturn 0;\n> +}\n> +\n>  static void odb_source_files_prepare(struct odb_source *source,\n>  \t\t\t\t     enum odb_prepare_flags flags)\n>  {\n> @@ -271,6 +289,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n>  \n>  \tfiles->base.free = odb_source_files_free;\n>  \tfiles->base.close = odb_source_files_close;\n> +\tfiles->base.create_on_disk = odb_source_files_create_on_disk;\n>  \tfiles->base.prepare = odb_source_files_prepare;\n>  \tfiles->base.read_object_info = odb_source_files_read_object_info;\n>  \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\n> diff --git a/odb/source.h b/odb/source.h\n> index ab16d152f4..4abc418bdd 100644\n> --- a/odb/source.h\n> +++ b/odb/source.h\n> @@ -89,6 +89,18 @@ struct odb_source {\n>  \t */\n>  \tvoid (*close)(struct odb_source *source);\n>  \n> +\t/*\n> +\t * This callback is expected to create on-disk data structures that are\n> +\t * required for this source to operate.\n> +\t *\n> +\t * The callback is expected to return 0 on success, a negative error\n> +\t * code otherwise.\n> +\t *\n> +\t * This callback may be NULL in case the source does not need any\n> +\t * on-disk setup.\n> +\t */\n> +\tint (*create_on_disk)(struct odb_source *source);\n> +\n>  \t/*\n>  \t * This callback is expected to prepare the source so that it becomes\n>  \t * ready for use. It optionally clears underlying caches of the object\n> @@ -316,6 +328,17 @@ static inline void odb_source_close(struct odb_source *source)\n>  \tsource->close(source);\n>  }\n>  \n> +/*\n> + * Create on-disk data structures that are required for this source to operate\n> + * correctly. Returns 0 on success, a negative error code otherwise.\n> + */\n> +static inline int odb_source_create_on_disk(struct odb_source *source)\n> +{\n> +\tif (!source->create_on_disk)\n> +\t\treturn 0;\n> +\treturn source->create_on_disk(source);\n> +}\n> +\n>  /*\n>   * Prepare the object database source and clear any caches. Depending on the\n>   * backend used this may have the effect that concurrently-written objects\n> diff --git a/setup.c b/setup.c\n> index d85171f3b6..af02cd965c 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -2654,25 +2654,27 @@ static int create_default_files(struct repository *repo,\n>  \n>  static void create_object_database(struct repository *repo)\n>  {\n> -\tstruct strbuf path = STRBUF_INIT;\n> -\tsize_t baselen;\n> +\t/*\n> +\t * Create the \"objects\" directory in the common directory. This is done\n> +\t * so that the repository can be discovered regardless of the backend\n> +\t * used.\n> +\t *\n> +\t * Note that we only do this in case the object directory wasn't\n> +\t * overwritten via an environment variable. If it _is_ being overridden\n> +\t * then we skip this step, as the repository won't be discoverable\n> +\t * anyway without the environment variable.\n> +\t */\n> +\tif (!getenv(DB_ENVIRONMENT)) {\n\nIt's a bit sad that [PATCH 3/6] removed the use of DB_ENVIRONMENT from\nthis file, and now we're re-adding it. Although, I don't see how else we\ncan do this.\n\n> +\t\tstruct strbuf objects_dir = STRBUF_INIT;\n> +\t\trepo_common_path_append(repo, &objects_dir, \"objects\");\n> +\t\tsafe_create_dir(repo, objects_dir.buf, 1);\n> +\t\tstrbuf_release(&objects_dir);\n> +\t}\n>  \n>  \trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n>  \n> -\tstrbuf_addstr(&path, repo_get_object_directory(repo));\n> -\tbaselen = path.len;\n> -\n> -\tsafe_create_dir(repo, path.buf, 1);\n> -\n> -\tstrbuf_setlen(&path, baselen);\n> -\tstrbuf_addstr(&path, \"/pack\");\n> -\tsafe_create_dir(repo, path.buf, 1);\n> -\n> -\tstrbuf_setlen(&path, baselen);\n> -\tstrbuf_addstr(&path, \"/info\");\n> -\tsafe_create_dir(repo, path.buf, 1);\n> -\n> -\tstrbuf_release(&path);\n> +\tif (odb_source_create_on_disk(repo->objects->sources) < 0)\n> +\t\tdie(\"failed creating object database\");\n\nThis error isn't translatable.\n\n>  }\n>  \n>  static void separate_git_dir(const char *git_dir, const char *git_link)\n>\n> -- \n> 2.55.0.679.g6767b8d81c.dirty\n>\n\n-- \nCheers,\nToon\n"},{"id":"549787","messageId":"anQjhlnvvhKLOFPV@pks.im","threadId":"66056","inReplyTo":"8733ws6424.fsf@emacs.iotcl.com","subject":"Re: [PATCH v3 4/6] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T06:02:46Z","receivedAt":"2026-08-06T06:02:58Z","isPatch":true,"body":"On Wed, Aug 05, 2026 at 04:21:39PM +0200, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > In a subsequent commit we'll make the creation of the on-disk data\n> > structures of an object database pluggable. This will lead to an\n> > in-between state where we have already configured the repository's\n> > object database, but it's not usable yet until we eventually call\n> > `create_object_directory()`.\n> >\n> > Defer the object database creation so that we handle both steps in the\n> > same function.\n> \n> With [PATCH v3 3/6], Justin's objection[1] is stronger now:\n> \n> > Naive question: would it be simpler to just require invoking `odb_new()`\n> > explicitly after `apply_repository_format()` in all cases? There doesn't\n> > appear to be too many callsites.\n> \n> As a matter of fact, I've given this a try and see these changes on top\n> of this series below.\n\nThe reason I was hesitant to do this is that I want to move\n`apply_repository_format()` into `repo_init()` eventually. But I guess\nmoving the call to `odb_new()` out of it doesn't really prevent that.\nSo... fine, I'll do it.\n\nPatrick\n"},{"id":"549788","messageId":"anQj_ww0Y2guJDcM@pks.im","threadId":"66056","inReplyTo":"878q6k66ha.fsf@emacs.iotcl.com","subject":"Re: [PATCH v3 3/6] setup: handle ODB-related environment variables in `odb_new()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T06:04:47Z","receivedAt":"2026-08-06T06:04:53Z","isPatch":true,"body":"On Wed, Aug 05, 2026 at 03:29:21PM +0200, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > diff --git a/odb.c b/odb.c\n> > index cf6e7938c0..b463afa072 100644\n> > --- a/odb.c\n> > +++ b/odb.c\n> > @@ -1004,26 +1004,30 @@ int odb_write_object_stream(struct object_database *odb,\n> >  }\n> >  \n> >  struct object_database *odb_new(struct repository *repo,\n> > -\t\t\t\tconst char *primary_source,\n> > -\t\t\t\tconst char *secondary_sources)\n> > +\t\t\t\tenum odb_new_flags flags)\n> >  {\n> > -\tstruct object_database *o = xmalloc(sizeof(*o));\n> > -\tchar *to_free = NULL;\n> > +\tchar *primary_source = NULL, *secondary_sources = NULL;\n> > +\tstruct object_database *o;\n> >  \n> > -\tmemset(o, 0, sizeof(*o));\n> > +\tCALLOC_ARRAY(o, 1);\n> >  \to->repo = repo;\n> >  \tpthread_mutex_init(&o->replace_mutex, NULL);\n> >  \tstring_list_init_dup(&o->submodule_source_paths);\n> >  \n> > +\tif (flags & ODB_NEW_HONOR_ENV) {\n> > +\t\tprimary_source = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n> > +\t\tsecondary_sources = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n> > +\t}\n> >  \tif (!primary_source)\n> > -\t\tprimary_source = to_free = xstrfmt(\"%s/objects\", repo->commondir);\n> > +\t\tprimary_source = xstrfmt(\"%s/objects\", repo->commondir);\n> > +\n> >  \to->sources = odb_source_new(o, primary_source, true);\n> >  \to->sources_tail = &o->sources->next;\n> >  \to->alternate_db = xstrdup_or_null(secondary_sources);\n> \n> I'd say this xstrdup_or_null() is not needed no more, and so is the\n> free() of that variable below.\n\nTrue indeed.\n\nPatrick\n"},{"id":"549802","messageId":"20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","subject":"[PATCH v4 0/6] odb: make creation of object database pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T07:50:58Z","receivedAt":"2026-08-06T07:51:11Z","isPatch":true,"body":"Hi,\n\nwhen creating a new repository we create a couple of on-disk data\nstructures for the object database. This includes the \"objects/\"\ndirectory hierarchy with \"objects/info\" and \"objects/pack\", which are\nspecific to the backend.\n\nThis patch series makes the creation of the on-disk data structures\npluggable. While we continue to always create \"objects/\" regardless of\nthe backend (it's required for a repository to be recognized as such),\nthe other subdirectories are now created by the backend. This will allow\nother backends to plug in their own logic.\n\nThe series starts with a small detour into the loose-object map. This\ndetour is required so that we can defer initialization of the object\ndatabase itself to a later point in time.\n\nThe series is based on 9a0c4701dc (The 7th batch, 2026-07-22).\n\nChanges in v4:\n  - Drop `APPLY_REPOSITOY_FORMAT_SKIP_ODB_CREATION` in favor of explicit\n    calls to `odb_new()`.\n  - Remove a useless call to `xstrdup()`.\n  - Mark a string as translatable.\n  - Link to v3: https://patch.msgid.link/20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im\n\nChanges in v3:\n  - Move handling of GIT_OBJECT_DIRECTORY and\n    GIT_ALTERNATE_OBJECT_DIRECTORIES into `odb_new()` itself. This\n    deduplicates some of the logic and also preps us for a future where\n    alternates are handled in the \"files\" backend itself.\n  - Link to v2: https://patch.msgid.link/20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im\n\nChanges in v2:\n  - Add a testcase that demonstrates the bug fixed with alternate loose\n    object maps.\n  - Rename the \"inmemory\" bakcend to \"in-memory\".\n  - Clarify some commit messages.\n  - Link to v1: https://patch.msgid.link/20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (6):\n      loose: load loose object map for the correct source\n      setup: detangle loading of loose object maps\n      setup: handle ODB-related environment variables in `odb_new()`\n      setup: defer object database creation\n      odb/source: introduce function to map source type to name\n      odb: make creation of on-disk structures pluggable\n\n loose.c                       | 25 ++++++++++----------\n loose.h                       |  1 +\n odb.c                         | 21 +++++++++--------\n odb.h                         | 17 ++++++++++++--\n odb/source-files.c            | 19 +++++++++++++++\n odb/source-files.h            |  4 +++-\n odb/source-inmemory.h         |  4 +++-\n odb/source-loose.c            |  2 ++\n odb/source-loose.h            |  4 +++-\n odb/source-packed.h           |  4 +++-\n odb/source.c                  | 19 +++++++++++++++\n odb/source.h                  | 29 +++++++++++++++++++++++\n repository.c                  |  3 +--\n setup.c                       | 54 +++++++++++++++++++++----------------------\n t/t1016-compatObjectFormat.sh | 18 +++++++++++++++\n t/unit-tests/u-odb-inmemory.c |  2 +-\n 16 files changed, 169 insertions(+), 57 deletions(-)\n\nRange-diff versus v3:\n\n1:  e1a585a3f7 = 1:  6dd8d575c6 loose: load loose object map for the correct source\n2:  1f1200f7ba = 2:  1e7adada64 setup: detangle loading of loose object maps\n3:  af02e520a2 ! 3:  2265f38695 setup: handle ODB-related environment variables in `odb_new()`\n    @@ odb.c: int odb_write_object_stream(struct object_database *odb,\n     +\n      \to->sources = odb_source_new(o, primary_source, true);\n      \to->sources_tail = &o->sources->next;\n    - \to->alternate_db = xstrdup_or_null(secondary_sources);\n    +-\to->alternate_db = xstrdup_or_null(secondary_sources);\n    ++\to->alternate_db = secondary_sources;\n      \to->inmemory_objects = &odb_source_inmemory_new(o)->base;\n      \n     -\tfree(to_free);\n     -\n    -+\tfree(secondary_sources);\n     +\tfree(primary_source);\n      \treturn o;\n      }\n4:  2c794be101 ! 4:  5274ee6bab setup: defer object database creation\n    @@ Commit message\n         object database, but it's not usable yet until we eventually call\n         `create_object_directory()`.\n     \n    -    Defer the object database creation so that we handle both steps in the\n    -    same function.\n    +    Lift the call to `odb_new()` out of `apply_repository_format()` so that\n    +    callers have more wiggle room with when exactly they call it, and adapt\n    +    them accordingly. The only exception is `init_db()`, where we now defer\n    +    creating the object database until we call `create_object_database()`.\n    +\n    +    With this change, initializing and creating the object database on disk\n    +    is now neatly encapsulated in a single function, which will make it\n    +    easier for a subsequent commit to move creation of the on-disk data\n    +    structures into the `struct odb_source` backends.\n     \n         Signed-off-by: Patrick Steinhardt <ps@pks.im>\n     \n    + ## repository.c ##\n    +@@ repository.c: int repo_init(struct repository *repo,\n    + \t\twarning(\"%s\", err.buf);\n    + \t\tgoto error;\n    + \t}\n    ++\trepo->objects = odb_new(repo, 0);\n    + \n    + \tif (worktree)\n    + \t\trepo_set_worktree(repo, worktree);\n    +\n      ## setup.c ##\n     @@ setup.c: int apply_repository_format(struct repository *repo,\n      \t\t\t    enum apply_repository_format_flags flags,\n    @@ setup.c: int apply_repository_format(struct repository *repo,\n     -\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n     -\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n     -\trepo->objects = odb_new(repo, odb_new_flags);\n    -+\tif (!(flags & APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION)) {\n    -+\t\tenum odb_new_flags odb_new_flags = 0;\n    -+\t\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n    -+\t\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n    -+\t\trepo->objects = odb_new(repo, odb_new_flags);\n    -+\t}\n    - \n    +-\n      \treturn 0;\n      }\n    + \n    +@@ setup.c: const char *enter_repo(struct repository *repo, const char *path, unsigned flags\n    + \t\tread_and_verify_repository_format(&fmt, \".\", NULL);\n    + \t\tif (apply_repository_format(repo, &fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n    + \t\t\tdie(\"%s\", err.buf);\n    ++\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n    + \t\tstartup_info->have_repository = 1;\n    + \n    + \t\tclear_repository_format(&fmt);\n    +@@ setup.c: const char *setup_git_directory_gently(struct repository *repo, int *nongit_ok)\n    + \t\t\tif (apply_repository_format(repo, &discovery.format,\n    + \t\t\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n    + \t\t\t\tdie(\"%s\", err.buf);\n    ++\t\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n    + \n    + \t\t\tclear_repository_format(&discovery.format);\n    + \t\t\tstrbuf_release(&err);\n     @@ setup.c: static int create_default_files(struct repository *repo,\n      \treturn reinit;\n      }\n    @@ setup.c: int init_db(struct repository *repo,\n      \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n     -\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n     +\tif (apply_repository_format(repo, &repo_fmt,\n    -+\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV |\n    -+\t\t\t\t    APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION, &err) < 0)\n    ++\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n      \t\tdie(\"%s\", err.buf);\n     -\tstartup_info->have_repository = 1;\n      \n    @@ setup.c: int init_db(struct repository *repo,\n      \n      \tif (repo_settings_get_shared_repository(repo)) {\n      \t\tchar buf[10];\n    -\n    - ## setup.h ##\n    -@@ setup.h: enum apply_repository_format_flags {\n    - \t * relate to the object database.\n    - \t */\n    - \tAPPLY_REPOSITORY_FORMAT_HONOR_ENV = (1 << 0),\n    -+\n    -+\t/*\n    -+\t * Usually, the object database is created after the repository format\n    -+\t * was applied. This step is skipped if this flag is set, which leaves\n    -+\t * us with a partially-working repository.\n    -+\t *\n    -+\t * This is useful when initializing a new repository.\n    -+\t */\n    -+\tAPPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION = (1 << 1),\n    - };\n    - \n    - /*\n5:  7397c760df = 5:  b444a314a6 odb/source: introduce function to map source type to name\n6:  7049e41a73 ! 6:  acb48f1072 odb: make creation of on-disk structures pluggable\n    @@ setup.c: static int create_default_files(struct repository *repo,\n     -\n     -\tstrbuf_release(&path);\n     +\tif (odb_source_create_on_disk(repo->objects->sources) < 0)\n    -+\t\tdie(\"failed creating object database\");\n    ++\t\tdie(_(\"failed creating object database\"));\n      }\n      \n      static void separate_git_dir(const char *git_dir, const char *git_link)\n\n---\nbase-commit: 9a0c4701dcd5725c4184599322b52933ff5005ca\nchange-id: 20260710-pks-odb-create-on-disk-ae8757861c69\n\n"},{"id":"549803","messageId":"20260806-pks-odb-create-on-disk-v4-1-ba8b4fdd2e3c@pks.im","threadId":"66056","inReplyTo":"20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im","subject":"[PATCH v4 1/6] loose: load loose object map for the correct source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T07:50:59Z","receivedAt":"2026-08-06T07:51:12Z","isPatch":true,"body":"When loading the loose object map via `load_one_loose_object_map()` we\npass in both a repository and the corresponding source. We ultimately\ndon't really respect the passed-in source though as we instead always\nload the map via the common directory. This doesn't make any sense\nthough, as the function is called in a loop through all sources, and as\nsuch the expectation is that we'll load the map that belongs to the\ngiven source. The consequence is that we'll ignore loose object maps of\nany configured alternates.\n\nFix this bug by instead loading the map via the loose source's path.\n\nHelped-by: Toon Claes <toon@iotcl.com>\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c                       | 18 ++++++++++--------\n t/t1016-compatObjectFormat.sh | 18 ++++++++++++++++++\n 2 files changed, 28 insertions(+), 8 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex bf01d3e42d..9dad75373b 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,9 +61,11 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct repository *repo, struct odb_source_loose *loose)\n+static int load_one_loose_object_map(struct odb_source_loose *loose)\n {\n-\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\tstruct repository *repo = loose->base.odb->repo;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar *path;\n \tFILE *fp;\n \tint ret = -1;\n \n@@ -78,10 +80,10 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n \tinsert_loose_map(loose, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n \tinsert_loose_map(loose, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n \n-\trepo_common_path_replace(repo, &path, \"objects/loose-object-idx\");\n-\tfp = fopen(path.buf, \"rb\");\n+\tpath = xstrfmt(\"%s/loose-object-idx\", loose->base.path);\n+\tfp = fopen(path, \"rb\");\n \tif (!fp) {\n-\t\tstrbuf_release(&path);\n+\t\tfree(path);\n \t\treturn 0;\n \t}\n \n@@ -102,7 +104,7 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n err:\n \tfclose(fp);\n \tstrbuf_release(&buf);\n-\tstrbuf_release(&path);\n+\tfree(path);\n \treturn ret;\n }\n \n@@ -117,10 +119,10 @@ int repo_read_loose_object_map(struct repository *repo)\n \n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(repo, files->loose) < 0) {\n+\t\tif (load_one_loose_object_map(files->loose) < 0)\n \t\t\treturn -1;\n-\t\t}\n \t}\n+\n \treturn 0;\n }\n \ndiff --git a/t/t1016-compatObjectFormat.sh b/t/t1016-compatObjectFormat.sh\nindex 92d48b96a1..9cafcee509 100755\n--- a/t/t1016-compatObjectFormat.sh\n+++ b/t/t1016-compatObjectFormat.sh\n@@ -187,6 +187,24 @@ do\n \t\teval signedtag3_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag3) &&\n \t\teval signedtag4_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag4)\n \t'\n+\n+\ttest_expect_success 'rev-parse maps oid of object borrowed from alternate' '\n+\t\tfor repo in alt borrow\n+\t\tdo\n+\t\t\ttest_when_finished \"rm -rf $repo\" &&\n+\t\t\tgit init --object-format=$hash $repo &&\n+\t\t\tgit -C $repo config set core.repositoryformatversion 1 &&\n+\t\t\tgit -C $repo config set extensions.compatObjectFormat $(compat_hash $hash) || exit 1\n+\t\tdone &&\n+\n+\t\tgit -C alt commit --allow-empty --message A &&\n+\t\techo \"$(pwd)/alt/.git/objects\" >borrow/.git/objects/info/alternates &&\n+\n+\t\toid=$(git -C alt rev-parse HEAD) &&\n+\t\tgit -C alt    rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >expect &&\n+\t\tgit -C borrow rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >actual &&\n+\t\ttest_cmp expect actual\n+\t'\n done\n cd \"$base\"\n \n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549804","messageId":"20260806-pks-odb-create-on-disk-v4-2-ba8b4fdd2e3c@pks.im","threadId":"66056","inReplyTo":"20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im","subject":"[PATCH v4 2/6] setup: detangle loading of loose object maps","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T07:51:00Z","receivedAt":"2026-08-06T07:51:14Z","isPatch":true,"body":"When a repository is configured to use a compatibility hash function\nthen we load the loose object map when we initialize the repository.\nThis object map provides the mappings between the canonical object hash\nand the compatibility object hash.\n\nLoading the object map happens in `repo_set_compat_hash_algo()`, which\ncalls `repo_read_loose_object_map()` in case the compatibility object\nhash is non-zero. This setup sequence has two major downsides:\n\n  - We assume that the primary object database is the \"files\" object\n    database and unconditionally downcast it. This will cause us to BUG\n    in case a different object database type was used together with a\n    compat hash algorithm.\n\n  - We require the object database to already have been initialized when\n    configuring the object database. This means that we must intermix\n    configuration of the repository and initialization of its\n    sub-structures in a weird way.\n\nRefactor the logic so that we instead load the loose object map via the\n\"loose\" backend, which fixes both of the above issues.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c            | 11 +++++------\n loose.h            |  1 +\n odb/source-loose.c |  2 ++\n repository.c       |  2 --\n setup.c            |  5 +++--\n 5 files changed, 11 insertions(+), 10 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 9dad75373b..a3b2dcedc2 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,7 +61,7 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct odb_source_loose *loose)\n+int loose_object_map_load(struct odb_source_loose *loose)\n {\n \tstruct repository *repo = loose->base.odb->repo;\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -69,6 +69,9 @@ static int load_one_loose_object_map(struct odb_source_loose *loose)\n \tFILE *fp;\n \tint ret = -1;\n \n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n \tif (!loose->map)\n \t\tloose_object_map_init(&loose->map);\n \tif (!loose->cache) {\n@@ -112,14 +115,10 @@ int repo_read_loose_object_map(struct repository *repo)\n {\n \tstruct odb_source *source;\n \n-\tif (!should_use_loose_object_map(repo))\n-\t\treturn 0;\n-\n \todb_prepare_alternates(repo->objects);\n-\n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(files->loose) < 0)\n+\t\tif (loose_object_map_load(files->loose) < 0)\n \t\t\treturn -1;\n \t}\n \ndiff --git a/loose.h b/loose.h\nindex 6c9b3f4571..ed663ac550 100644\n--- a/loose.h\n+++ b/loose.h\n@@ -13,6 +13,7 @@ struct loose_object_map {\n \n void loose_object_map_init(struct loose_object_map **map);\n void loose_object_map_clear(struct loose_object_map **map);\n+int loose_object_map_load(struct odb_source_loose *loose);\n int repo_loose_object_map_oid(struct repository *repo,\n \t\t\t      const struct object_id *src,\n \t\t\t      const struct git_hash_algo *dest_algo,\ndiff --git a/odb/source-loose.c b/odb/source-loose.c\nindex 3f7d04a56e..812ca1c138 100644\n--- a/odb/source-loose.c\n+++ b/odb/source-loose.c\n@@ -727,5 +727,7 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n \tif (!is_absolute_path(loose->base.path))\n \t\tchdir_notify_register(NULL, odb_source_loose_reparent, loose);\n \n+\tloose_object_map_load(loose);\n+\n \treturn loose;\n }\ndiff --git a/repository.c b/repository.c\nindex 2ef0778846..6d633002b4 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -201,8 +201,6 @@ void repo_set_compat_hash_algo(struct repository *repo MAYBE_UNUSED, uint32_t al\n \tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n \t\tBUG(\"hash_algo and compat_hash_algo match\");\n \trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n-\tif (repo->compat_hash_algo)\n-\t\trepo_read_loose_object_map(repo);\n #else\n \tif (algo)\n \t\tdie(_(\"compatibility hash algorithm support requires Rust\"));\ndiff --git a/setup.c b/setup.c\nindex d31808130b..825572f5f1 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1788,8 +1788,6 @@ int apply_repository_format(struct repository *repo,\n \n \trepo->bare_cfg = format->is_bare;\n \trepo_set_hash_algo(repo, format->hash_algo);\n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n \trepo_set_compat_hash_algo(repo, format->compat_hash_algo);\n \trepo_set_ref_storage_format(repo,\n \t\t\t\t    format->ref_storage_format,\n@@ -1805,6 +1803,9 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n+\n \tfree(alternate_object_directories);\n \tfree(object_directory);\n \treturn 0;\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549805","messageId":"20260806-pks-odb-create-on-disk-v4-3-ba8b4fdd2e3c@pks.im","threadId":"66056","inReplyTo":"20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im","subject":"[PATCH v4 3/6] setup: handle ODB-related environment variables in `odb_new()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T07:51:01Z","receivedAt":"2026-08-06T07:51:18Z","isPatch":true,"body":"When initializing a repository's object database we have to respect the\nGIT_OBJECT_DIRECTORY and GIT_ALTERNATE_OBJECT_DIRECTORIES environment\nvariables, which can be set by the user to override the default location\nof where we write objects to and read objects from.\n\nThis is handled in `apply_repository_format()`, which is fine. But in a\nsubsequent commit we'll have to defer constructing the object database\nto a later point in some cases, and that will require a second site\nwhere we call `odb_new()`. And of course, that second site would have to\nhandle those environment variables, as well.\n\nIt would be somewhat awkward to duplicate the logic though. But there's\na better alternative: instead of handling this logic in \"setup.c\", we\ncan easily handle environment variables in `odb_new()` itself. This\nensures that object database creation is neatly self-contained, and we\ndon't have to duplicate any of the logic.\n\nAnother benefit is that in a future patch series we plan to move\nhandling of alternates into the backends themselves [1], and that will\nrequire us to also handle those environment variables in the \"files\"\nbackend itself. So moving the logic into the ODB level already gets us\none step closer to that goal.\n\nRefactor the logic accordingly.\n\n[1]: https://lore.kernel.org/git/amLgMqkqxR8mKIbT@pks.im/\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                         | 21 ++++++++++++---------\n odb.h                         | 17 +++++++++++++++--\n setup.c                       | 11 ++++-------\n t/unit-tests/u-odb-inmemory.c |  2 +-\n 4 files changed, 32 insertions(+), 19 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex cf6e7938c0..ed1d63f4bd 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -1004,26 +1004,29 @@ int odb_write_object_stream(struct object_database *odb,\n }\n \n struct object_database *odb_new(struct repository *repo,\n-\t\t\t\tconst char *primary_source,\n-\t\t\t\tconst char *secondary_sources)\n+\t\t\t\tenum odb_new_flags flags)\n {\n-\tstruct object_database *o = xmalloc(sizeof(*o));\n-\tchar *to_free = NULL;\n+\tchar *primary_source = NULL, *secondary_sources = NULL;\n+\tstruct object_database *o;\n \n-\tmemset(o, 0, sizeof(*o));\n+\tCALLOC_ARRAY(o, 1);\n \to->repo = repo;\n \tpthread_mutex_init(&o->replace_mutex, NULL);\n \tstring_list_init_dup(&o->submodule_source_paths);\n \n+\tif (flags & ODB_NEW_HONOR_ENV) {\n+\t\tprimary_source = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n+\t\tsecondary_sources = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n+\t}\n \tif (!primary_source)\n-\t\tprimary_source = to_free = xstrfmt(\"%s/objects\", repo->commondir);\n+\t\tprimary_source = xstrfmt(\"%s/objects\", repo->commondir);\n+\n \to->sources = odb_source_new(o, primary_source, true);\n \to->sources_tail = &o->sources->next;\n-\to->alternate_db = xstrdup_or_null(secondary_sources);\n+\to->alternate_db = secondary_sources;\n \to->inmemory_objects = &odb_source_inmemory_new(o)->base;\n \n-\tfree(to_free);\n-\n+\tfree(primary_source);\n \treturn o;\n }\n \ndiff --git a/odb.h b/odb.h\nindex 7995bed97b..8ec335c7f7 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -100,6 +100,20 @@ struct object_database {\n \tstruct string_list submodule_source_paths;\n };\n \n+enum odb_new_flags {\n+\t/*\n+\t * Honor environment variables when constructing the object database\n+\t * sources. This makes us respect the following environment variables:\n+\t *\n+\t *   - GIT_OBJECT_DIRECTORY to override the primary object directory.\n+\t *\n+\t *   - GIT_ALTERNATE_OBJECT_DIRECTORIES to override alternates.\n+\t *\n+\t * Environment variables may be backend-specific.\n+\t */\n+\tODB_NEW_HONOR_ENV = (1 << 0),\n+};\n+\n /*\n  * Create a new object database for the given repository.\n  *\n@@ -112,8 +126,7 @@ struct object_database {\n  * Returns the newly created object database.\n  */\n struct object_database *odb_new(struct repository *repo,\n-\t\t\t\tconst char *primary_source,\n-\t\t\t\tconst char *alternate_sources);\n+\t\t\t\tenum odb_new_flags flags);\n \n /* Free the object database and release all resources. */\n void odb_free(struct object_database *o);\ndiff --git a/setup.c b/setup.c\nindex 825572f5f1..5dfab3e79e 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1765,7 +1765,7 @@ int apply_repository_format(struct repository *repo,\n \t\t\t    enum apply_repository_format_flags flags,\n \t\t\t    struct strbuf *err)\n {\n-\tchar *object_directory = NULL, *alternate_object_directories = NULL;\n+\tenum odb_new_flags odb_new_flags = 0;\n \n \tif (verify_repository_format(format, err) < 0)\n \t\treturn -1;\n@@ -1779,8 +1779,6 @@ int apply_repository_format(struct repository *repo,\n \tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV) {\n \t\tconst char *shallow_file;\n \n-\t\tobject_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n-\t\talternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n \t\tshallow_file = getenv(GIT_SHALLOW_FILE_ENVIRONMENT);\n \t\tif (shallow_file)\n \t\t\tset_alternate_shallow_file(repo, shallow_file);\n@@ -1803,11 +1801,10 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n+\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n+\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n+\trepo->objects = odb_new(repo, odb_new_flags);\n \n-\tfree(alternate_object_directories);\n-\tfree(object_directory);\n \treturn 0;\n }\n \ndiff --git a/t/unit-tests/u-odb-inmemory.c b/t/unit-tests/u-odb-inmemory.c\nindex 6844bfc37c..db323e10fd 100644\n--- a/t/unit-tests/u-odb-inmemory.c\n+++ b/t/unit-tests/u-odb-inmemory.c\n@@ -38,7 +38,7 @@ static void cl_assert_object_info(struct odb_source_inmemory *source,\n \n void test_odb_inmemory__initialize(void)\n {\n-\todb = odb_new(&repo, \"\", \"\");\n+\todb = odb_new(&repo, 0);\n }\n \n void test_odb_inmemory__cleanup(void)\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549806","messageId":"20260806-pks-odb-create-on-disk-v4-4-ba8b4fdd2e3c@pks.im","threadId":"66056","inReplyTo":"20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im","subject":"[PATCH v4 4/6] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T07:51:02Z","receivedAt":"2026-08-06T07:51:20Z","isPatch":true,"body":"In a subsequent commit we'll make the creation of the on-disk data\nstructures of an object database pluggable. This will lead to an\nin-between state where we have already configured the repository's\nobject database, but it's not usable yet until we eventually call\n`create_object_directory()`.\n\nLift the call to `odb_new()` out of `apply_repository_format()` so that\ncallers have more wiggle room with when exactly they call it, and adapt\nthem accordingly. The only exception is `init_db()`, where we now defer\ncreating the object database until we call `create_object_database()`.\n\nWith this change, initializing and creating the object database on disk\nis now neatly encapsulated in a single function, which will make it\neasier for a subsequent commit to move creation of the on-disk data\nstructures into the `struct odb_source` backends.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n repository.c |  1 +\n setup.c      | 20 ++++++++++----------\n 2 files changed, 11 insertions(+), 10 deletions(-)\n\ndiff --git a/repository.c b/repository.c\nindex 6d633002b4..5ec264e607 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -294,6 +294,7 @@ int repo_init(struct repository *repo,\n \t\twarning(\"%s\", err.buf);\n \t\tgoto error;\n \t}\n+\trepo->objects = odb_new(repo, 0);\n \n \tif (worktree)\n \t\trepo_set_worktree(repo, worktree);\ndiff --git a/setup.c b/setup.c\nindex 5dfab3e79e..e39a1646bb 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1765,8 +1765,6 @@ int apply_repository_format(struct repository *repo,\n \t\t\t    enum apply_repository_format_flags flags,\n \t\t\t    struct strbuf *err)\n {\n-\tenum odb_new_flags odb_new_flags = 0;\n-\n \tif (verify_repository_format(format, err) < 0)\n \t\treturn -1;\n \n@@ -1801,10 +1799,6 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n-\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n-\trepo->objects = odb_new(repo, odb_new_flags);\n-\n \treturn 0;\n }\n \n@@ -1888,6 +1882,7 @@ const char *enter_repo(struct repository *repo, const char *path, unsigned flags\n \t\tread_and_verify_repository_format(&fmt, \".\", NULL);\n \t\tif (apply_repository_format(repo, &fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n \t\tstartup_info->have_repository = 1;\n \n \t\tclear_repository_format(&fmt);\n@@ -2090,6 +2085,7 @@ const char *setup_git_directory_gently(struct repository *repo, int *nongit_ok)\n \t\t\tif (apply_repository_format(repo, &discovery.format,\n \t\t\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n \t\t\t\tdie(\"%s\", err.buf);\n+\t\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n \n \t\t\tclear_repository_format(&discovery.format);\n \t\t\tstrbuf_release(&err);\n@@ -2651,11 +2647,13 @@ static int create_default_files(struct repository *repo,\n \treturn reinit;\n }\n \n-static void create_object_directory(struct repository *repo)\n+static void create_object_database(struct repository *repo)\n {\n \tstruct strbuf path = STRBUF_INIT;\n \tsize_t baselen;\n \n+\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n+\n \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n \tbaselen = path.len;\n \n@@ -2864,9 +2862,9 @@ int init_db(struct repository *repo,\n \t */\n \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n-\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n+\tif (apply_repository_format(repo, &repo_fmt,\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n \t\tdie(\"%s\", err.buf);\n-\tstartup_info->have_repository = 1;\n \n \t/*\n \t * Ensure `core.hidedotfiles` is processed. This must happen after we\n@@ -2882,7 +2880,9 @@ int init_db(struct repository *repo,\n \n \tif (!(flags & INIT_DB_SKIP_REFDB))\n \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n-\tcreate_object_directory(repo);\n+\tcreate_object_database(repo);\n+\n+\tstartup_info->have_repository = 1;\n \n \tif (repo_settings_get_shared_repository(repo)) {\n \t\tchar buf[10];\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549807","messageId":"20260806-pks-odb-create-on-disk-v4-5-ba8b4fdd2e3c@pks.im","threadId":"66056","inReplyTo":"20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im","subject":"[PATCH v4 5/6] odb/source: introduce function to map source type to name","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T07:51:03Z","receivedAt":"2026-08-06T07:51:24Z","isPatch":true,"body":"Introduce a new function that maps an object source's type to a\nhuman-readable name. Use the function to provide better human-readable\nerror messages for the downcasting functions.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.h    |  4 +++-\n odb/source-inmemory.h |  4 +++-\n odb/source-loose.h    |  4 +++-\n odb/source-packed.h   |  4 +++-\n odb/source.c          | 19 +++++++++++++++++++\n odb/source.h          |  6 ++++++\n 6 files changed, 37 insertions(+), 4 deletions(-)\n\ndiff --git a/odb/source-files.h b/odb/source-files.h\nindex d7ac3c1c81..6a803afdda 100644\n--- a/odb/source-files.h\n+++ b/odb/source-files.h\n@@ -28,7 +28,9 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n static inline struct odb_source_files *odb_source_files_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_FILES)\n-\t\tBUG(\"trying to downcast source of type '%d' to files\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_FILES));\n \treturn container_of(source, struct odb_source_files, base);\n }\n \ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex a88fc2e320..adbad23e8b 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -26,7 +26,9 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_INMEMORY)\n-\t\tBUG(\"trying to downcast source of type '%d' to in-memory\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_INMEMORY));\n \treturn container_of(source, struct odb_source_inmemory, base);\n }\n \ndiff --git a/odb/source-loose.h b/odb/source-loose.h\nindex 6070aaf3ce..3cf2e1f8f1 100644\n--- a/odb/source-loose.h\n+++ b/odb/source-loose.h\n@@ -41,7 +41,9 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n static inline struct odb_source_loose *odb_source_loose_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_LOOSE)\n-\t\tBUG(\"trying to downcast source of type '%d' to loose\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_LOOSE));\n \treturn container_of(source, struct odb_source_loose, base);\n }\n \ndiff --git a/odb/source-packed.h b/odb/source-packed.h\nindex 77309ddd09..a0f6b5096d 100644\n--- a/odb/source-packed.h\n+++ b/odb/source-packed.h\n@@ -78,7 +78,9 @@ struct odb_source_packed *odb_source_packed_new(struct object_database *odb,\n static inline struct odb_source_packed *odb_source_packed_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_PACKED)\n-\t\tBUG(\"trying to downcast source of type '%d' to packed\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_PACKED));\n \treturn container_of(source, struct odb_source_packed, base);\n }\n \ndiff --git a/odb/source.c b/odb/source.c\nindex 7993dcbd65..30188b806d 100644\n--- a/odb/source.c\n+++ b/odb/source.c\n@@ -4,6 +4,25 @@\n #include \"odb/source.h\"\n #include \"packfile.h\"\n \n+static const char * const odb_source_names_by_type[] = {\n+\t[ODB_SOURCE_UNKNOWN] = \"unknown\",\n+\t[ODB_SOURCE_FILES] = \"files\",\n+\t[ODB_SOURCE_LOOSE] = \"loose\",\n+\t[ODB_SOURCE_PACKED] = \"packed\",\n+\t[ODB_SOURCE_INMEMORY] = \"in-memory\",\n+};\n+\n+const char *odb_source_type_to_name(enum odb_source_type type)\n+{\n+\tconst char *name;\n+\tif (type < 0 || type >= ARRAY_SIZE(odb_source_names_by_type))\n+\t\ttype = ODB_SOURCE_UNKNOWN;\n+\tname = odb_source_names_by_type[type];\n+\tif (!name)\n+\t\tBUG(\"name missing in `odb_source_names_by_type` for '%d'\", type);\n+\treturn name;\n+}\n+\n struct odb_source *odb_source_new(struct object_database *odb,\n \t\t\t\t  const char *path,\n \t\t\t\t  bool local)\ndiff --git a/odb/source.h b/odb/source.h\nindex cd63dba91f..ab16d152f4 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -25,6 +25,12 @@ enum odb_source_type {\n \tODB_SOURCE_INMEMORY,\n };\n \n+/*\n+ * Convert between the enum and its name. Returns the equivalent of \"unknown\"\n+ * for unknown types.\n+ */\n+const char *odb_source_type_to_name(enum odb_source_type type);\n+\n struct object_id;\n struct odb_read_stream;\n struct strvec;\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549808","messageId":"20260806-pks-odb-create-on-disk-v4-6-ba8b4fdd2e3c@pks.im","threadId":"66056","inReplyTo":"20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im","subject":"[PATCH v4 6/6] odb: make creation of on-disk structures pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T07:51:04Z","receivedAt":"2026-08-06T07:51:27Z","isPatch":true,"body":"When creating a new \"files\" object database source we have to create a\ncouple of directories. These directories are of course specific to this\nparticular backend, and a different backend may require a setup that is\ncompletely different.\n\nMake the creation of on-disk structures pluggable to accommodate for\nthis.\n\nNote that there is one exception though: the \"objects\" directory must\nexist in a repository regardless of which backend is in use. If it\ndoesn't exist then the repository is not treated as a Git repository at\nall. Consequently, we create this directory regardless of the backend.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.c | 19 +++++++++++++++++++\n odb/source.h       | 23 +++++++++++++++++++++++\n setup.c            | 34 ++++++++++++++++++----------------\n 3 files changed, 60 insertions(+), 16 deletions(-)\n\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex 4138758511..0db6e681fe 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -9,6 +9,7 @@\n #include \"odb/source-files.h\"\n #include \"odb/source-loose.h\"\n #include \"packfile.h\"\n+#include \"path.h\"\n #include \"strbuf.h\"\n #include \"write-or-die.h\"\n \n@@ -41,6 +42,23 @@ static void odb_source_files_close(struct odb_source *source)\n \todb_source_close(&files->packed->base);\n }\n \n+static int odb_source_files_create_on_disk(struct odb_source *source)\n+{\n+\tstruct strbuf path = STRBUF_INIT;\n+\n+\tsafe_create_dir(source->odb->repo, source->path, 1);\n+\n+\tstrbuf_addf(&path, \"%s/pack\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_reset(&path);\n+\tstrbuf_addf(&path, \"%s/info\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_release(&path);\n+\treturn 0;\n+}\n+\n static void odb_source_files_prepare(struct odb_source *source,\n \t\t\t\t     enum odb_prepare_flags flags)\n {\n@@ -271,6 +289,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \n \tfiles->base.free = odb_source_files_free;\n \tfiles->base.close = odb_source_files_close;\n+\tfiles->base.create_on_disk = odb_source_files_create_on_disk;\n \tfiles->base.prepare = odb_source_files_prepare;\n \tfiles->base.read_object_info = odb_source_files_read_object_info;\n \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\ndiff --git a/odb/source.h b/odb/source.h\nindex ab16d152f4..4abc418bdd 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -89,6 +89,18 @@ struct odb_source {\n \t */\n \tvoid (*close)(struct odb_source *source);\n \n+\t/*\n+\t * This callback is expected to create on-disk data structures that are\n+\t * required for this source to operate.\n+\t *\n+\t * The callback is expected to return 0 on success, a negative error\n+\t * code otherwise.\n+\t *\n+\t * This callback may be NULL in case the source does not need any\n+\t * on-disk setup.\n+\t */\n+\tint (*create_on_disk)(struct odb_source *source);\n+\n \t/*\n \t * This callback is expected to prepare the source so that it becomes\n \t * ready for use. It optionally clears underlying caches of the object\n@@ -316,6 +328,17 @@ static inline void odb_source_close(struct odb_source *source)\n \tsource->close(source);\n }\n \n+/*\n+ * Create on-disk data structures that are required for this source to operate\n+ * correctly. Returns 0 on success, a negative error code otherwise.\n+ */\n+static inline int odb_source_create_on_disk(struct odb_source *source)\n+{\n+\tif (!source->create_on_disk)\n+\t\treturn 0;\n+\treturn source->create_on_disk(source);\n+}\n+\n /*\n  * Prepare the object database source and clear any caches. Depending on the\n  * backend used this may have the effect that concurrently-written objects\ndiff --git a/setup.c b/setup.c\nindex e39a1646bb..1f65f69534 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -2649,25 +2649,27 @@ static int create_default_files(struct repository *repo,\n \n static void create_object_database(struct repository *repo)\n {\n-\tstruct strbuf path = STRBUF_INIT;\n-\tsize_t baselen;\n+\t/*\n+\t * Create the \"objects\" directory in the common directory. This is done\n+\t * so that the repository can be discovered regardless of the backend\n+\t * used.\n+\t *\n+\t * Note that we only do this in case the object directory wasn't\n+\t * overwritten via an environment variable. If it _is_ being overridden\n+\t * then we skip this step, as the repository won't be discoverable\n+\t * anyway without the environment variable.\n+\t */\n+\tif (!getenv(DB_ENVIRONMENT)) {\n+\t\tstruct strbuf objects_dir = STRBUF_INIT;\n+\t\trepo_common_path_append(repo, &objects_dir, \"objects\");\n+\t\tsafe_create_dir(repo, objects_dir.buf, 1);\n+\t\tstrbuf_release(&objects_dir);\n+\t}\n \n \trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n \n-\tstrbuf_addstr(&path, repo_get_object_directory(repo));\n-\tbaselen = path.len;\n-\n-\tsafe_create_dir(repo, path.buf, 1);\n-\n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/pack\");\n-\tsafe_create_dir(repo, path.buf, 1);\n-\n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/info\");\n-\tsafe_create_dir(repo, path.buf, 1);\n-\n-\tstrbuf_release(&path);\n+\tif (odb_source_create_on_disk(repo->objects->sources) < 0)\n+\t\tdie(_(\"failed creating object database\"));\n }\n \n static void separate_git_dir(const char *git_dir, const char *git_link)\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549854","messageId":"87tsp749be.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"20260806-pks-odb-create-on-disk-v4-4-ba8b4fdd2e3c@pks.im","subject":"Re: [PATCH v4 4/6] setup: defer object database creation","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-08-06T14:23:17Z","receivedAt":"2026-08-06T14:23:27Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> In a subsequent commit we'll make the creation of the on-disk data\n> structures of an object database pluggable. This will lead to an\n> in-between state where we have already configured the repository's\n> object database, but it's not usable yet until we eventually call\n> `create_object_directory()`.\n>\n> Lift the call to `odb_new()` out of `apply_repository_format()` so that\n> callers have more wiggle room with when exactly they call it, and adapt\n> them accordingly. The only exception is `init_db()`, where we now defer\n> creating the object database until we call `create_object_database()`.\n>\n> With this change, initializing and creating the object database on disk\n> is now neatly encapsulated in a single function, which will make it\n> easier for a subsequent commit to move creation of the on-disk data\n> structures into the `struct odb_source` backends.\n>\n> Signed-off-by: Patrick Steinhardt <ps@pks.im>\n> ---\n>  repository.c |  1 +\n>  setup.c      | 20 ++++++++++----------\n>  2 files changed, 11 insertions(+), 10 deletions(-)\n>\n> diff --git a/repository.c b/repository.c\n> index 6d633002b4..5ec264e607 100644\n> --- a/repository.c\n> +++ b/repository.c\n> @@ -294,6 +294,7 @@ int repo_init(struct repository *repo,\n>  \t\twarning(\"%s\", err.buf);\n>  \t\tgoto error;\n>  \t}\n> +\trepo->objects = odb_new(repo, 0);\n>  \n>  \tif (worktree)\n>  \t\trepo_set_worktree(repo, worktree);\n> diff --git a/setup.c b/setup.c\n> index 5dfab3e79e..e39a1646bb 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -1765,8 +1765,6 @@ int apply_repository_format(struct repository *repo,\n>  \t\t\t    enum apply_repository_format_flags flags,\n>  \t\t\t    struct strbuf *err)\n\nI've noticed the docs in setup.h say:\n\n    /*\n     * Apply the given repository format to the repo. This initializes extensions\n     * and basic data structures required for normal operation. Returns 0 on\n     * success, a negative error code when the format is not valid as determined by\n     * `verify_repository_format()`.\n     */\n\nI'm not sure that's still applicable, now odb_new() isn't called no\nmore.\n\n>  {\n> -\tenum odb_new_flags odb_new_flags = 0;\n> -\n>  \tif (verify_repository_format(format, err) < 0)\n>  \t\treturn -1;\n>  \n> @@ -1801,10 +1799,6 @@ int apply_repository_format(struct repository *repo,\n>  \trepo->repository_format_precious_objects =\n>  \t\tformat->precious_objects;\n>  \n> -\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n> -\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n> -\trepo->objects = odb_new(repo, odb_new_flags);\n> -\n>  \treturn 0;\n>  }\n>  \n> @@ -1888,6 +1882,7 @@ const char *enter_repo(struct repository *repo, const char *path, unsigned flags\n>  \t\tread_and_verify_repository_format(&fmt, \".\", NULL);\n>  \t\tif (apply_repository_format(repo, &fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n>  \t\t\tdie(\"%s\", err.buf);\n> +\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n>  \t\tstartup_info->have_repository = 1;\n>  \n>  \t\tclear_repository_format(&fmt);\n> @@ -2090,6 +2085,7 @@ const char *setup_git_directory_gently(struct repository *repo, int *nongit_ok)\n>  \t\t\tif (apply_repository_format(repo, &discovery.format,\n>  \t\t\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n>  \t\t\t\tdie(\"%s\", err.buf);\n> +\t\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n>  \n>  \t\t\tclear_repository_format(&discovery.format);\n>  \t\t\tstrbuf_release(&err);\n> @@ -2651,11 +2647,13 @@ static int create_default_files(struct repository *repo,\n>  \treturn reinit;\n>  }\n>  \n> -static void create_object_directory(struct repository *repo)\n> +static void create_object_database(struct repository *repo)\n>  {\n>  \tstruct strbuf path = STRBUF_INIT;\n>  \tsize_t baselen;\n>  \n> +\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n> +\n>  \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n>  \tbaselen = path.len;\n>  \n> @@ -2864,9 +2862,9 @@ int init_db(struct repository *repo,\n>  \t */\n>  \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n>  \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n> -\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n> +\tif (apply_repository_format(repo, &repo_fmt,\n> +\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n\nNit: Not sure why this formatting change was needed. I would have\nassumed to have all apply_repository_format() calls formatted the same,\nbut I've noticed at line 1883 in enter_repo() it's still a single-line\ncall.\n\n>  \t\tdie(\"%s\", err.buf);\n> -\tstartup_info->have_repository = 1;\n>  \n>  \t/*\n>  \t * Ensure `core.hidedotfiles` is processed. This must happen after we\n> @@ -2882,7 +2880,9 @@ int init_db(struct repository *repo,\n>  \n>  \tif (!(flags & INIT_DB_SKIP_REFDB))\n>  \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n> -\tcreate_object_directory(repo);\n> +\tcreate_object_database(repo);\n> +\n> +\tstartup_info->have_repository = 1;\n>  \n>  \tif (repo_settings_get_shared_repository(repo)) {\n>  \t\tchar buf[10];\n>\n> -- \n> 2.55.0.679.g6767b8d81c.dirty\n>\n\n-- \nCheers,\nToon\n"},{"id":"549855","messageId":"87qzkb495s.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im","subject":"Re: [PATCH v4 0/6] odb: make creation of object database pluggable","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-08-06T14:26:39Z","receivedAt":"2026-08-06T14:26:50Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Hi,\n>\n> when creating a new repository we create a couple of on-disk data\n> structures for the object database. This includes the \"objects/\"\n> directory hierarchy with \"objects/info\" and \"objects/pack\", which are\n> specific to the backend.\n>\n> This patch series makes the creation of the on-disk data structures\n> pluggable. While we continue to always create \"objects/\" regardless of\n> the backend (it's required for a repository to be recognized as such),\n> the other subdirectories are now created by the backend. This will allow\n> other backends to plug in their own logic.\n>\n> The series starts with a small detour into the loose-object map. This\n> detour is required so that we can defer initialization of the object\n> database itself to a later point in time.\n>\n> The series is based on 9a0c4701dc (The 7th batch, 2026-07-22).\n>\n> Changes in v4:\n>   - Drop `APPLY_REPOSITOY_FORMAT_SKIP_ODB_CREATION` in favor of explicit\n>     calls to `odb_new()`.\n>   - Remove a useless call to `xstrdup()`.\n>   - Mark a string as translatable.\n>   - Link to v3: https://patch.msgid.link/20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im\n\nStructurally I'm very happy about this version. Only had some nits about\ncomments and formatting, but overall this version looks good to me.\n\nThanks!\n\n-- \nCheers,\nToon\n"},{"id":"549856","messageId":"anSgJ4pHuwJ5hylE@pks.im","threadId":"66056","inReplyTo":"87tsp749be.fsf@emacs.iotcl.com","subject":"Re: [PATCH v4 4/6] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-06T14:54:31Z","receivedAt":"2026-08-06T14:54:39Z","isPatch":true,"body":"On Thu, Aug 06, 2026 at 04:23:17PM +0200, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> > diff --git a/setup.c b/setup.c\n> > index 5dfab3e79e..e39a1646bb 100644\n> > --- a/setup.c\n> > +++ b/setup.c\n> > @@ -1765,8 +1765,6 @@ int apply_repository_format(struct repository *repo,\n> >  \t\t\t    enum apply_repository_format_flags flags,\n> >  \t\t\t    struct strbuf *err)\n> \n> I've noticed the docs in setup.h say:\n> \n>     /*\n>      * Apply the given repository format to the repo. This initializes extensions\n>      * and basic data structures required for normal operation. Returns 0 on\n>      * success, a negative error code when the format is not valid as determined by\n>      * `verify_repository_format()`.\n>      */\n> \n> I'm not sure that's still applicable, now odb_new() isn't called no\n> more.\n\nFair enough.\n\n> > @@ -2864,9 +2862,9 @@ int init_db(struct repository *repo,\n> >  \t */\n> >  \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n> >  \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n> > -\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n> > +\tif (apply_repository_format(repo, &repo_fmt,\n> > +\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n> \n> Nit: Not sure why this formatting change was needed. I would have\n> assumed to have all apply_repository_format() calls formatted the same,\n> but I've noticed at line 1883 in enter_repo() it's still a single-line\n> call.\n\nIt's an artifact from previous versions.\n\nI'll send a (hopefully last) reroll in a bit. Thanks!\n\nPatrick\n"},{"id":"549875","messageId":"xmqqh5l7jghw.fsf@gitster.g","threadId":"66056","inReplyTo":"anSgJ4pHuwJ5hylE@pks.im","subject":"Re: [PATCH v4 4/6] setup: defer object database creation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-06T17:39:07Z","receivedAt":"2026-08-06T17:39:09Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> It's an artifact from previous versions.\n>\n> I'll send a (hopefully last) reroll in a bit. Thanks!\n>\n> Patrick\n\nWith Toon's <87qzkb495s.fsf@emacs.iotcl.com> and this message, I'll\nmark the topic as \"Expecting a (hopefully small and final) reroll.\"\n\nThanks.\n"},{"id":"549916","messageId":"20260807-pks-odb-create-on-disk-v5-0-399da0b0b140@pks.im","threadId":"66056","inReplyTo":"20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im","subject":"[PATCH v5 0/6] odb: make creation of object database pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-07T03:34:24Z","receivedAt":"2026-08-07T03:34:35Z","isPatch":true,"body":"Hi,\n\nwhen creating a new repository we create a couple of on-disk data\nstructures for the object database. This includes the \"objects/\"\ndirectory hierarchy with \"objects/info\" and \"objects/pack\", which are\nspecific to the backend.\n\nThis patch series makes the creation of the on-disk data structures\npluggable. While we continue to always create \"objects/\" regardless of\nthe backend (it's required for a repository to be recognized as such),\nthe other subdirectories are now created by the backend. This will allow\nother backends to plug in their own logic.\n\nThe series starts with a small detour into the loose-object map. This\ndetour is required so that we can defer initialization of the object\ndatabase itself to a later point in time.\n\nThe series is based on 9a0c4701dc (The 7th batch, 2026-07-22).\n\nChanges in v5:\n  - Remove a leftover formatting change.\n  - Fix a stale comment.\n  - Link to v4: https://patch.msgid.link/20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im\n\nChanges in v4:\n  - Drop `APPLY_REPOSITOY_FORMAT_SKIP_ODB_CREATION` in favor of explicit\n    calls to `odb_new()`.\n  - Remove a useless call to `xstrdup()`.\n  - Mark a string as translatable.\n  - Link to v3: https://patch.msgid.link/20260805-pks-odb-create-on-disk-v3-0-c0ee3ac5141f@pks.im\n\nChanges in v3:\n  - Move handling of GIT_OBJECT_DIRECTORY and\n    GIT_ALTERNATE_OBJECT_DIRECTORIES into `odb_new()` itself. This\n    deduplicates some of the logic and also preps us for a future where\n    alternates are handled in the \"files\" backend itself.\n  - Link to v2: https://patch.msgid.link/20260804-pks-odb-create-on-disk-v2-0-ddf8b59bd207@pks.im\n\nChanges in v2:\n  - Add a testcase that demonstrates the bug fixed with alternate loose\n    object maps.\n  - Rename the \"inmemory\" bakcend to \"in-memory\".\n  - Clarify some commit messages.\n  - Link to v1: https://patch.msgid.link/20260724-pks-odb-create-on-disk-v1-0-3b3d265d979b@pks.im\n\nThanks!\n\nPatrick\n\n---\nPatrick Steinhardt (6):\n      loose: load loose object map for the correct source\n      setup: detangle loading of loose object maps\n      setup: handle ODB-related environment variables in `odb_new()`\n      setup: defer object database creation\n      odb/source: introduce function to map source type to name\n      odb: make creation of on-disk structures pluggable\n\n loose.c                       | 25 +++++++++++----------\n loose.h                       |  1 +\n odb.c                         | 21 ++++++++++--------\n odb.h                         | 17 +++++++++++++--\n odb/source-files.c            | 19 ++++++++++++++++\n odb/source-files.h            |  4 +++-\n odb/source-inmemory.h         |  4 +++-\n odb/source-loose.c            |  2 ++\n odb/source-loose.h            |  4 +++-\n odb/source-packed.h           |  4 +++-\n odb/source.c                  | 19 ++++++++++++++++\n odb/source.h                  | 29 ++++++++++++++++++++++++\n repository.c                  |  3 +--\n setup.c                       | 51 +++++++++++++++++++++----------------------\n setup.h                       |  4 ++--\n t/t1016-compatObjectFormat.sh | 18 +++++++++++++++\n t/unit-tests/u-odb-inmemory.c |  2 +-\n 17 files changed, 169 insertions(+), 58 deletions(-)\n\nRange-diff versus v4:\n\n1:  40ca0d1345 = 1:  3a0fbf9498 loose: load loose object map for the correct source\n2:  d18ddec5dd = 2:  7ba250f4d7 setup: detangle loading of loose object maps\n3:  9b6fbc510f = 3:  fbe755388b setup: handle ODB-related environment variables in `odb_new()`\n4:  f27f8d45a4 ! 4:  4d7a12e3cb setup: defer object database creation\n    @@ setup.c: static int create_default_files(struct repository *repo,\n      \tbaselen = path.len;\n      \n     @@ setup.c: int init_db(struct repository *repo,\n    - \t */\n    - \tread_and_verify_repository_format(&repo_fmt, repo_get_git_dir(repo), NULL);\n      \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n    --\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n    -+\tif (apply_repository_format(repo, &repo_fmt,\n    -+\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n    + \tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n      \t\tdie(\"%s\", err.buf);\n     -\tstartup_info->have_repository = 1;\n      \n    @@ setup.c: int init_db(struct repository *repo,\n      \n      \tif (repo_settings_get_shared_repository(repo)) {\n      \t\tchar buf[10];\n    +\n    + ## setup.h ##\n    +@@ setup.h: enum apply_repository_format_flags {\n    + \n    + /*\n    +  * Apply the given repository format to the repo. This initializes extensions\n    +- * and basic data structures required for normal operation. Returns 0 on\n    +- * success, a negative error code when the format is not valid as determined by\n    ++ * required for normal operation. Returns 0 on success, a negative error code\n    ++ * when the format is not valid as determined by\n    +  * `verify_repository_format()`.\n    +  */\n    + int apply_repository_format(struct repository *repo,\n5:  1c0afb893f = 5:  6bb4ecc76d odb/source: introduce function to map source type to name\n6:  387fe6e204 = 6:  806f399c63 odb: make creation of on-disk structures pluggable\n\n---\nbase-commit: 9a0c4701dcd5725c4184599322b52933ff5005ca\nchange-id: 20260710-pks-odb-create-on-disk-ae8757861c69\n\n"},{"id":"549917","messageId":"20260807-pks-odb-create-on-disk-v5-1-399da0b0b140@pks.im","threadId":"66056","inReplyTo":"20260807-pks-odb-create-on-disk-v5-0-399da0b0b140@pks.im","subject":"[PATCH v5 1/6] loose: load loose object map for the correct source","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-07T03:34:25Z","receivedAt":"2026-08-07T03:34:37Z","isPatch":true,"body":"When loading the loose object map via `load_one_loose_object_map()` we\npass in both a repository and the corresponding source. We ultimately\ndon't really respect the passed-in source though as we instead always\nload the map via the common directory. This doesn't make any sense\nthough, as the function is called in a loop through all sources, and as\nsuch the expectation is that we'll load the map that belongs to the\ngiven source. The consequence is that we'll ignore loose object maps of\nany configured alternates.\n\nFix this bug by instead loading the map via the loose source's path.\n\nHelped-by: Toon Claes <toon@iotcl.com>\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c                       | 18 ++++++++++--------\n t/t1016-compatObjectFormat.sh | 18 ++++++++++++++++++\n 2 files changed, 28 insertions(+), 8 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex bf01d3e42d..9dad75373b 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,9 +61,11 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct repository *repo, struct odb_source_loose *loose)\n+static int load_one_loose_object_map(struct odb_source_loose *loose)\n {\n-\tstruct strbuf buf = STRBUF_INIT, path = STRBUF_INIT;\n+\tstruct repository *repo = loose->base.odb->repo;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tchar *path;\n \tFILE *fp;\n \tint ret = -1;\n \n@@ -78,10 +80,10 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n \tinsert_loose_map(loose, repo->hash_algo->empty_blob, repo->compat_hash_algo->empty_blob);\n \tinsert_loose_map(loose, repo->hash_algo->null_oid, repo->compat_hash_algo->null_oid);\n \n-\trepo_common_path_replace(repo, &path, \"objects/loose-object-idx\");\n-\tfp = fopen(path.buf, \"rb\");\n+\tpath = xstrfmt(\"%s/loose-object-idx\", loose->base.path);\n+\tfp = fopen(path, \"rb\");\n \tif (!fp) {\n-\t\tstrbuf_release(&path);\n+\t\tfree(path);\n \t\treturn 0;\n \t}\n \n@@ -102,7 +104,7 @@ static int load_one_loose_object_map(struct repository *repo, struct odb_source_\n err:\n \tfclose(fp);\n \tstrbuf_release(&buf);\n-\tstrbuf_release(&path);\n+\tfree(path);\n \treturn ret;\n }\n \n@@ -117,10 +119,10 @@ int repo_read_loose_object_map(struct repository *repo)\n \n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(repo, files->loose) < 0) {\n+\t\tif (load_one_loose_object_map(files->loose) < 0)\n \t\t\treturn -1;\n-\t\t}\n \t}\n+\n \treturn 0;\n }\n \ndiff --git a/t/t1016-compatObjectFormat.sh b/t/t1016-compatObjectFormat.sh\nindex 92d48b96a1..9cafcee509 100755\n--- a/t/t1016-compatObjectFormat.sh\n+++ b/t/t1016-compatObjectFormat.sh\n@@ -187,6 +187,24 @@ do\n \t\teval signedtag3_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag3) &&\n \t\teval signedtag4_${hash}_oid=$(git hash-object -t tag -w ../${hash}_signedtag4)\n \t'\n+\n+\ttest_expect_success 'rev-parse maps oid of object borrowed from alternate' '\n+\t\tfor repo in alt borrow\n+\t\tdo\n+\t\t\ttest_when_finished \"rm -rf $repo\" &&\n+\t\t\tgit init --object-format=$hash $repo &&\n+\t\t\tgit -C $repo config set core.repositoryformatversion 1 &&\n+\t\t\tgit -C $repo config set extensions.compatObjectFormat $(compat_hash $hash) || exit 1\n+\t\tdone &&\n+\n+\t\tgit -C alt commit --allow-empty --message A &&\n+\t\techo \"$(pwd)/alt/.git/objects\" >borrow/.git/objects/info/alternates &&\n+\n+\t\toid=$(git -C alt rev-parse HEAD) &&\n+\t\tgit -C alt    rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >expect &&\n+\t\tgit -C borrow rev-parse --output-object-format=$(compat_hash $hash) \"$oid\" >actual &&\n+\t\ttest_cmp expect actual\n+\t'\n done\n cd \"$base\"\n \n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549918","messageId":"20260807-pks-odb-create-on-disk-v5-2-399da0b0b140@pks.im","threadId":"66056","inReplyTo":"20260807-pks-odb-create-on-disk-v5-0-399da0b0b140@pks.im","subject":"[PATCH v5 2/6] setup: detangle loading of loose object maps","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-07T03:34:26Z","receivedAt":"2026-08-07T03:34:39Z","isPatch":true,"body":"When a repository is configured to use a compatibility hash function\nthen we load the loose object map when we initialize the repository.\nThis object map provides the mappings between the canonical object hash\nand the compatibility object hash.\n\nLoading the object map happens in `repo_set_compat_hash_algo()`, which\ncalls `repo_read_loose_object_map()` in case the compatibility object\nhash is non-zero. This setup sequence has two major downsides:\n\n  - We assume that the primary object database is the \"files\" object\n    database and unconditionally downcast it. This will cause us to BUG\n    in case a different object database type was used together with a\n    compat hash algorithm.\n\n  - We require the object database to already have been initialized when\n    configuring the object database. This means that we must intermix\n    configuration of the repository and initialization of its\n    sub-structures in a weird way.\n\nRefactor the logic so that we instead load the loose object map via the\n\"loose\" backend, which fixes both of the above issues.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n loose.c            | 11 +++++------\n loose.h            |  1 +\n odb/source-loose.c |  2 ++\n repository.c       |  2 --\n setup.c            |  5 +++--\n 5 files changed, 11 insertions(+), 10 deletions(-)\n\ndiff --git a/loose.c b/loose.c\nindex 9dad75373b..a3b2dcedc2 100644\n--- a/loose.c\n+++ b/loose.c\n@@ -61,7 +61,7 @@ static int insert_loose_map(struct odb_source_loose *loose,\n \treturn inserted;\n }\n \n-static int load_one_loose_object_map(struct odb_source_loose *loose)\n+int loose_object_map_load(struct odb_source_loose *loose)\n {\n \tstruct repository *repo = loose->base.odb->repo;\n \tstruct strbuf buf = STRBUF_INIT;\n@@ -69,6 +69,9 @@ static int load_one_loose_object_map(struct odb_source_loose *loose)\n \tFILE *fp;\n \tint ret = -1;\n \n+\tif (!should_use_loose_object_map(repo))\n+\t\treturn 0;\n+\n \tif (!loose->map)\n \t\tloose_object_map_init(&loose->map);\n \tif (!loose->cache) {\n@@ -112,14 +115,10 @@ int repo_read_loose_object_map(struct repository *repo)\n {\n \tstruct odb_source *source;\n \n-\tif (!should_use_loose_object_map(repo))\n-\t\treturn 0;\n-\n \todb_prepare_alternates(repo->objects);\n-\n \tfor (source = repo->objects->sources; source; source = source->next) {\n \t\tstruct odb_source_files *files = odb_source_files_downcast(source);\n-\t\tif (load_one_loose_object_map(files->loose) < 0)\n+\t\tif (loose_object_map_load(files->loose) < 0)\n \t\t\treturn -1;\n \t}\n \ndiff --git a/loose.h b/loose.h\nindex 6c9b3f4571..ed663ac550 100644\n--- a/loose.h\n+++ b/loose.h\n@@ -13,6 +13,7 @@ struct loose_object_map {\n \n void loose_object_map_init(struct loose_object_map **map);\n void loose_object_map_clear(struct loose_object_map **map);\n+int loose_object_map_load(struct odb_source_loose *loose);\n int repo_loose_object_map_oid(struct repository *repo,\n \t\t\t      const struct object_id *src,\n \t\t\t      const struct git_hash_algo *dest_algo,\ndiff --git a/odb/source-loose.c b/odb/source-loose.c\nindex 3f7d04a56e..812ca1c138 100644\n--- a/odb/source-loose.c\n+++ b/odb/source-loose.c\n@@ -727,5 +727,7 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n \tif (!is_absolute_path(loose->base.path))\n \t\tchdir_notify_register(NULL, odb_source_loose_reparent, loose);\n \n+\tloose_object_map_load(loose);\n+\n \treturn loose;\n }\ndiff --git a/repository.c b/repository.c\nindex 2ef0778846..6d633002b4 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -201,8 +201,6 @@ void repo_set_compat_hash_algo(struct repository *repo MAYBE_UNUSED, uint32_t al\n \tif (hash_algo_by_ptr(repo->hash_algo) == algo)\n \t\tBUG(\"hash_algo and compat_hash_algo match\");\n \trepo->compat_hash_algo = algo ? &hash_algos[algo] : NULL;\n-\tif (repo->compat_hash_algo)\n-\t\trepo_read_loose_object_map(repo);\n #else\n \tif (algo)\n \t\tdie(_(\"compatibility hash algorithm support requires Rust\"));\ndiff --git a/setup.c b/setup.c\nindex d31808130b..825572f5f1 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1788,8 +1788,6 @@ int apply_repository_format(struct repository *repo,\n \n \trepo->bare_cfg = format->is_bare;\n \trepo_set_hash_algo(repo, format->hash_algo);\n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n \trepo_set_compat_hash_algo(repo, format->compat_hash_algo);\n \trepo_set_ref_storage_format(repo,\n \t\t\t\t    format->ref_storage_format,\n@@ -1805,6 +1803,9 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n+\trepo->objects = odb_new(repo, object_directory,\n+\t\t\t\talternate_object_directories);\n+\n \tfree(alternate_object_directories);\n \tfree(object_directory);\n \treturn 0;\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549919","messageId":"20260807-pks-odb-create-on-disk-v5-3-399da0b0b140@pks.im","threadId":"66056","inReplyTo":"20260807-pks-odb-create-on-disk-v5-0-399da0b0b140@pks.im","subject":"[PATCH v5 3/6] setup: handle ODB-related environment variables in `odb_new()`","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-07T03:34:27Z","receivedAt":"2026-08-07T03:34:43Z","isPatch":true,"body":"When initializing a repository's object database we have to respect the\nGIT_OBJECT_DIRECTORY and GIT_ALTERNATE_OBJECT_DIRECTORIES environment\nvariables, which can be set by the user to override the default location\nof where we write objects to and read objects from.\n\nThis is handled in `apply_repository_format()`, which is fine. But in a\nsubsequent commit we'll have to defer constructing the object database\nto a later point in some cases, and that will require a second site\nwhere we call `odb_new()`. And of course, that second site would have to\nhandle those environment variables, as well.\n\nIt would be somewhat awkward to duplicate the logic though. But there's\na better alternative: instead of handling this logic in \"setup.c\", we\ncan easily handle environment variables in `odb_new()` itself. This\nensures that object database creation is neatly self-contained, and we\ndon't have to duplicate any of the logic.\n\nAnother benefit is that in a future patch series we plan to move\nhandling of alternates into the backends themselves [1], and that will\nrequire us to also handle those environment variables in the \"files\"\nbackend itself. So moving the logic into the ODB level already gets us\none step closer to that goal.\n\nRefactor the logic accordingly.\n\n[1]: https://lore.kernel.org/git/amLgMqkqxR8mKIbT@pks.im/\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb.c                         | 21 ++++++++++++---------\n odb.h                         | 17 +++++++++++++++--\n setup.c                       | 11 ++++-------\n t/unit-tests/u-odb-inmemory.c |  2 +-\n 4 files changed, 32 insertions(+), 19 deletions(-)\n\ndiff --git a/odb.c b/odb.c\nindex cf6e7938c0..ed1d63f4bd 100644\n--- a/odb.c\n+++ b/odb.c\n@@ -1004,26 +1004,29 @@ int odb_write_object_stream(struct object_database *odb,\n }\n \n struct object_database *odb_new(struct repository *repo,\n-\t\t\t\tconst char *primary_source,\n-\t\t\t\tconst char *secondary_sources)\n+\t\t\t\tenum odb_new_flags flags)\n {\n-\tstruct object_database *o = xmalloc(sizeof(*o));\n-\tchar *to_free = NULL;\n+\tchar *primary_source = NULL, *secondary_sources = NULL;\n+\tstruct object_database *o;\n \n-\tmemset(o, 0, sizeof(*o));\n+\tCALLOC_ARRAY(o, 1);\n \to->repo = repo;\n \tpthread_mutex_init(&o->replace_mutex, NULL);\n \tstring_list_init_dup(&o->submodule_source_paths);\n \n+\tif (flags & ODB_NEW_HONOR_ENV) {\n+\t\tprimary_source = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n+\t\tsecondary_sources = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n+\t}\n \tif (!primary_source)\n-\t\tprimary_source = to_free = xstrfmt(\"%s/objects\", repo->commondir);\n+\t\tprimary_source = xstrfmt(\"%s/objects\", repo->commondir);\n+\n \to->sources = odb_source_new(o, primary_source, true);\n \to->sources_tail = &o->sources->next;\n-\to->alternate_db = xstrdup_or_null(secondary_sources);\n+\to->alternate_db = secondary_sources;\n \to->inmemory_objects = &odb_source_inmemory_new(o)->base;\n \n-\tfree(to_free);\n-\n+\tfree(primary_source);\n \treturn o;\n }\n \ndiff --git a/odb.h b/odb.h\nindex 7995bed97b..8ec335c7f7 100644\n--- a/odb.h\n+++ b/odb.h\n@@ -100,6 +100,20 @@ struct object_database {\n \tstruct string_list submodule_source_paths;\n };\n \n+enum odb_new_flags {\n+\t/*\n+\t * Honor environment variables when constructing the object database\n+\t * sources. This makes us respect the following environment variables:\n+\t *\n+\t *   - GIT_OBJECT_DIRECTORY to override the primary object directory.\n+\t *\n+\t *   - GIT_ALTERNATE_OBJECT_DIRECTORIES to override alternates.\n+\t *\n+\t * Environment variables may be backend-specific.\n+\t */\n+\tODB_NEW_HONOR_ENV = (1 << 0),\n+};\n+\n /*\n  * Create a new object database for the given repository.\n  *\n@@ -112,8 +126,7 @@ struct object_database {\n  * Returns the newly created object database.\n  */\n struct object_database *odb_new(struct repository *repo,\n-\t\t\t\tconst char *primary_source,\n-\t\t\t\tconst char *alternate_sources);\n+\t\t\t\tenum odb_new_flags flags);\n \n /* Free the object database and release all resources. */\n void odb_free(struct object_database *o);\ndiff --git a/setup.c b/setup.c\nindex 825572f5f1..5dfab3e79e 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1765,7 +1765,7 @@ int apply_repository_format(struct repository *repo,\n \t\t\t    enum apply_repository_format_flags flags,\n \t\t\t    struct strbuf *err)\n {\n-\tchar *object_directory = NULL, *alternate_object_directories = NULL;\n+\tenum odb_new_flags odb_new_flags = 0;\n \n \tif (verify_repository_format(format, err) < 0)\n \t\treturn -1;\n@@ -1779,8 +1779,6 @@ int apply_repository_format(struct repository *repo,\n \tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV) {\n \t\tconst char *shallow_file;\n \n-\t\tobject_directory = xstrdup_or_null(getenv(DB_ENVIRONMENT));\n-\t\talternate_object_directories = xstrdup_or_null(getenv(ALTERNATE_DB_ENVIRONMENT));\n \t\tshallow_file = getenv(GIT_SHALLOW_FILE_ENVIRONMENT);\n \t\tif (shallow_file)\n \t\t\tset_alternate_shallow_file(repo, shallow_file);\n@@ -1803,11 +1801,10 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\trepo->objects = odb_new(repo, object_directory,\n-\t\t\t\talternate_object_directories);\n+\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n+\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n+\trepo->objects = odb_new(repo, odb_new_flags);\n \n-\tfree(alternate_object_directories);\n-\tfree(object_directory);\n \treturn 0;\n }\n \ndiff --git a/t/unit-tests/u-odb-inmemory.c b/t/unit-tests/u-odb-inmemory.c\nindex 6844bfc37c..db323e10fd 100644\n--- a/t/unit-tests/u-odb-inmemory.c\n+++ b/t/unit-tests/u-odb-inmemory.c\n@@ -38,7 +38,7 @@ static void cl_assert_object_info(struct odb_source_inmemory *source,\n \n void test_odb_inmemory__initialize(void)\n {\n-\todb = odb_new(&repo, \"\", \"\");\n+\todb = odb_new(&repo, 0);\n }\n \n void test_odb_inmemory__cleanup(void)\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549920","messageId":"20260807-pks-odb-create-on-disk-v5-4-399da0b0b140@pks.im","threadId":"66056","inReplyTo":"20260807-pks-odb-create-on-disk-v5-0-399da0b0b140@pks.im","subject":"[PATCH v5 4/6] setup: defer object database creation","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-07T03:34:28Z","receivedAt":"2026-08-07T03:34:46Z","isPatch":true,"body":"In a subsequent commit we'll make the creation of the on-disk data\nstructures of an object database pluggable. This will lead to an\nin-between state where we have already configured the repository's\nobject database, but it's not usable yet until we eventually call\n`create_object_directory()`.\n\nLift the call to `odb_new()` out of `apply_repository_format()` so that\ncallers have more wiggle room with when exactly they call it, and adapt\nthem accordingly. The only exception is `init_db()`, where we now defer\ncreating the object database until we call `create_object_database()`.\n\nWith this change, initializing and creating the object database on disk\nis now neatly encapsulated in a single function, which will make it\neasier for a subsequent commit to move creation of the on-disk data\nstructures into the `struct odb_source` backends.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n repository.c |  1 +\n setup.c      | 17 ++++++++---------\n setup.h      |  4 ++--\n 3 files changed, 11 insertions(+), 11 deletions(-)\n\ndiff --git a/repository.c b/repository.c\nindex 6d633002b4..5ec264e607 100644\n--- a/repository.c\n+++ b/repository.c\n@@ -294,6 +294,7 @@ int repo_init(struct repository *repo,\n \t\twarning(\"%s\", err.buf);\n \t\tgoto error;\n \t}\n+\trepo->objects = odb_new(repo, 0);\n \n \tif (worktree)\n \t\trepo_set_worktree(repo, worktree);\ndiff --git a/setup.c b/setup.c\nindex 5dfab3e79e..97338cbc51 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1765,8 +1765,6 @@ int apply_repository_format(struct repository *repo,\n \t\t\t    enum apply_repository_format_flags flags,\n \t\t\t    struct strbuf *err)\n {\n-\tenum odb_new_flags odb_new_flags = 0;\n-\n \tif (verify_repository_format(format, err) < 0)\n \t\treturn -1;\n \n@@ -1801,10 +1799,6 @@ int apply_repository_format(struct repository *repo,\n \trepo->repository_format_precious_objects =\n \t\tformat->precious_objects;\n \n-\tif (flags & APPLY_REPOSITORY_FORMAT_HONOR_ENV)\n-\t\todb_new_flags |= ODB_NEW_HONOR_ENV;\n-\trepo->objects = odb_new(repo, odb_new_flags);\n-\n \treturn 0;\n }\n \n@@ -1888,6 +1882,7 @@ const char *enter_repo(struct repository *repo, const char *path, unsigned flags\n \t\tread_and_verify_repository_format(&fmt, \".\", NULL);\n \t\tif (apply_repository_format(repo, &fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n \t\t\tdie(\"%s\", err.buf);\n+\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n \t\tstartup_info->have_repository = 1;\n \n \t\tclear_repository_format(&fmt);\n@@ -2090,6 +2085,7 @@ const char *setup_git_directory_gently(struct repository *repo, int *nongit_ok)\n \t\t\tif (apply_repository_format(repo, &discovery.format,\n \t\t\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n \t\t\t\tdie(\"%s\", err.buf);\n+\t\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n \n \t\t\tclear_repository_format(&discovery.format);\n \t\t\tstrbuf_release(&err);\n@@ -2651,11 +2647,13 @@ static int create_default_files(struct repository *repo,\n \treturn reinit;\n }\n \n-static void create_object_directory(struct repository *repo)\n+static void create_object_database(struct repository *repo)\n {\n \tstruct strbuf path = STRBUF_INIT;\n \tsize_t baselen;\n \n+\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n+\n \tstrbuf_addstr(&path, repo_get_object_directory(repo));\n \tbaselen = path.len;\n \n@@ -2866,7 +2864,6 @@ int init_db(struct repository *repo,\n \trepository_format_configure(&repo_fmt, hash, ref_storage_format);\n \tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n \t\tdie(\"%s\", err.buf);\n-\tstartup_info->have_repository = 1;\n \n \t/*\n \t * Ensure `core.hidedotfiles` is processed. This must happen after we\n@@ -2882,7 +2879,9 @@ int init_db(struct repository *repo,\n \n \tif (!(flags & INIT_DB_SKIP_REFDB))\n \t\tcreate_reference_database(repo, initial_branch, flags & INIT_DB_QUIET);\n-\tcreate_object_directory(repo);\n+\tcreate_object_database(repo);\n+\n+\tstartup_info->have_repository = 1;\n \n \tif (repo_settings_get_shared_repository(repo)) {\n \t\tchar buf[10];\ndiff --git a/setup.h b/setup.h\nindex 654f10e059..763fd384e8 100644\n--- a/setup.h\n+++ b/setup.h\n@@ -245,8 +245,8 @@ enum apply_repository_format_flags {\n \n /*\n  * Apply the given repository format to the repo. This initializes extensions\n- * and basic data structures required for normal operation. Returns 0 on\n- * success, a negative error code when the format is not valid as determined by\n+ * required for normal operation. Returns 0 on success, a negative error code\n+ * when the format is not valid as determined by\n  * `verify_repository_format()`.\n  */\n int apply_repository_format(struct repository *repo,\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549921","messageId":"20260807-pks-odb-create-on-disk-v5-5-399da0b0b140@pks.im","threadId":"66056","inReplyTo":"20260807-pks-odb-create-on-disk-v5-0-399da0b0b140@pks.im","subject":"[PATCH v5 5/6] odb/source: introduce function to map source type to name","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-07T03:34:29Z","receivedAt":"2026-08-07T03:34:49Z","isPatch":true,"body":"Introduce a new function that maps an object source's type to a\nhuman-readable name. Use the function to provide better human-readable\nerror messages for the downcasting functions.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.h    |  4 +++-\n odb/source-inmemory.h |  4 +++-\n odb/source-loose.h    |  4 +++-\n odb/source-packed.h   |  4 +++-\n odb/source.c          | 19 +++++++++++++++++++\n odb/source.h          |  6 ++++++\n 6 files changed, 37 insertions(+), 4 deletions(-)\n\ndiff --git a/odb/source-files.h b/odb/source-files.h\nindex d7ac3c1c81..6a803afdda 100644\n--- a/odb/source-files.h\n+++ b/odb/source-files.h\n@@ -28,7 +28,9 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n static inline struct odb_source_files *odb_source_files_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_FILES)\n-\t\tBUG(\"trying to downcast source of type '%d' to files\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_FILES));\n \treturn container_of(source, struct odb_source_files, base);\n }\n \ndiff --git a/odb/source-inmemory.h b/odb/source-inmemory.h\nindex a88fc2e320..adbad23e8b 100644\n--- a/odb/source-inmemory.h\n+++ b/odb/source-inmemory.h\n@@ -26,7 +26,9 @@ struct odb_source_inmemory *odb_source_inmemory_new(struct object_database *odb)\n static inline struct odb_source_inmemory *odb_source_inmemory_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_INMEMORY)\n-\t\tBUG(\"trying to downcast source of type '%d' to in-memory\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_INMEMORY));\n \treturn container_of(source, struct odb_source_inmemory, base);\n }\n \ndiff --git a/odb/source-loose.h b/odb/source-loose.h\nindex 6070aaf3ce..3cf2e1f8f1 100644\n--- a/odb/source-loose.h\n+++ b/odb/source-loose.h\n@@ -41,7 +41,9 @@ struct odb_source_loose *odb_source_loose_new(struct object_database *odb,\n static inline struct odb_source_loose *odb_source_loose_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_LOOSE)\n-\t\tBUG(\"trying to downcast source of type '%d' to loose\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_LOOSE));\n \treturn container_of(source, struct odb_source_loose, base);\n }\n \ndiff --git a/odb/source-packed.h b/odb/source-packed.h\nindex 77309ddd09..a0f6b5096d 100644\n--- a/odb/source-packed.h\n+++ b/odb/source-packed.h\n@@ -78,7 +78,9 @@ struct odb_source_packed *odb_source_packed_new(struct object_database *odb,\n static inline struct odb_source_packed *odb_source_packed_downcast(struct odb_source *source)\n {\n \tif (source->type != ODB_SOURCE_PACKED)\n-\t\tBUG(\"trying to downcast source of type '%d' to packed\", source->type);\n+\t\tBUG(\"trying to downcast source of type '%s' to '%s'\",\n+\t\t    odb_source_type_to_name(source->type),\n+\t\t    odb_source_type_to_name(ODB_SOURCE_PACKED));\n \treturn container_of(source, struct odb_source_packed, base);\n }\n \ndiff --git a/odb/source.c b/odb/source.c\nindex 7993dcbd65..30188b806d 100644\n--- a/odb/source.c\n+++ b/odb/source.c\n@@ -4,6 +4,25 @@\n #include \"odb/source.h\"\n #include \"packfile.h\"\n \n+static const char * const odb_source_names_by_type[] = {\n+\t[ODB_SOURCE_UNKNOWN] = \"unknown\",\n+\t[ODB_SOURCE_FILES] = \"files\",\n+\t[ODB_SOURCE_LOOSE] = \"loose\",\n+\t[ODB_SOURCE_PACKED] = \"packed\",\n+\t[ODB_SOURCE_INMEMORY] = \"in-memory\",\n+};\n+\n+const char *odb_source_type_to_name(enum odb_source_type type)\n+{\n+\tconst char *name;\n+\tif (type < 0 || type >= ARRAY_SIZE(odb_source_names_by_type))\n+\t\ttype = ODB_SOURCE_UNKNOWN;\n+\tname = odb_source_names_by_type[type];\n+\tif (!name)\n+\t\tBUG(\"name missing in `odb_source_names_by_type` for '%d'\", type);\n+\treturn name;\n+}\n+\n struct odb_source *odb_source_new(struct object_database *odb,\n \t\t\t\t  const char *path,\n \t\t\t\t  bool local)\ndiff --git a/odb/source.h b/odb/source.h\nindex cd63dba91f..ab16d152f4 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -25,6 +25,12 @@ enum odb_source_type {\n \tODB_SOURCE_INMEMORY,\n };\n \n+/*\n+ * Convert between the enum and its name. Returns the equivalent of \"unknown\"\n+ * for unknown types.\n+ */\n+const char *odb_source_type_to_name(enum odb_source_type type);\n+\n struct object_id;\n struct odb_read_stream;\n struct strvec;\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549922","messageId":"20260807-pks-odb-create-on-disk-v5-6-399da0b0b140@pks.im","threadId":"66056","inReplyTo":"20260807-pks-odb-create-on-disk-v5-0-399da0b0b140@pks.im","subject":"[PATCH v5 6/6] odb: make creation of on-disk structures pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-07T03:34:30Z","receivedAt":"2026-08-07T03:34:52Z","isPatch":true,"body":"When creating a new \"files\" object database source we have to create a\ncouple of directories. These directories are of course specific to this\nparticular backend, and a different backend may require a setup that is\ncompletely different.\n\nMake the creation of on-disk structures pluggable to accommodate for\nthis.\n\nNote that there is one exception though: the \"objects\" directory must\nexist in a repository regardless of which backend is in use. If it\ndoesn't exist then the repository is not treated as a Git repository at\nall. Consequently, we create this directory regardless of the backend.\n\nSigned-off-by: Patrick Steinhardt <ps@pks.im>\n---\n odb/source-files.c | 19 +++++++++++++++++++\n odb/source.h       | 23 +++++++++++++++++++++++\n setup.c            | 34 ++++++++++++++++++----------------\n 3 files changed, 60 insertions(+), 16 deletions(-)\n\ndiff --git a/odb/source-files.c b/odb/source-files.c\nindex 4138758511..0db6e681fe 100644\n--- a/odb/source-files.c\n+++ b/odb/source-files.c\n@@ -9,6 +9,7 @@\n #include \"odb/source-files.h\"\n #include \"odb/source-loose.h\"\n #include \"packfile.h\"\n+#include \"path.h\"\n #include \"strbuf.h\"\n #include \"write-or-die.h\"\n \n@@ -41,6 +42,23 @@ static void odb_source_files_close(struct odb_source *source)\n \todb_source_close(&files->packed->base);\n }\n \n+static int odb_source_files_create_on_disk(struct odb_source *source)\n+{\n+\tstruct strbuf path = STRBUF_INIT;\n+\n+\tsafe_create_dir(source->odb->repo, source->path, 1);\n+\n+\tstrbuf_addf(&path, \"%s/pack\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_reset(&path);\n+\tstrbuf_addf(&path, \"%s/info\", source->path);\n+\tsafe_create_dir(source->odb->repo, path.buf, 1);\n+\n+\tstrbuf_release(&path);\n+\treturn 0;\n+}\n+\n static void odb_source_files_prepare(struct odb_source *source,\n \t\t\t\t     enum odb_prepare_flags flags)\n {\n@@ -271,6 +289,7 @@ struct odb_source_files *odb_source_files_new(struct object_database *odb,\n \n \tfiles->base.free = odb_source_files_free;\n \tfiles->base.close = odb_source_files_close;\n+\tfiles->base.create_on_disk = odb_source_files_create_on_disk;\n \tfiles->base.prepare = odb_source_files_prepare;\n \tfiles->base.read_object_info = odb_source_files_read_object_info;\n \tfiles->base.read_object_stream = odb_source_files_read_object_stream;\ndiff --git a/odb/source.h b/odb/source.h\nindex ab16d152f4..4abc418bdd 100644\n--- a/odb/source.h\n+++ b/odb/source.h\n@@ -89,6 +89,18 @@ struct odb_source {\n \t */\n \tvoid (*close)(struct odb_source *source);\n \n+\t/*\n+\t * This callback is expected to create on-disk data structures that are\n+\t * required for this source to operate.\n+\t *\n+\t * The callback is expected to return 0 on success, a negative error\n+\t * code otherwise.\n+\t *\n+\t * This callback may be NULL in case the source does not need any\n+\t * on-disk setup.\n+\t */\n+\tint (*create_on_disk)(struct odb_source *source);\n+\n \t/*\n \t * This callback is expected to prepare the source so that it becomes\n \t * ready for use. It optionally clears underlying caches of the object\n@@ -316,6 +328,17 @@ static inline void odb_source_close(struct odb_source *source)\n \tsource->close(source);\n }\n \n+/*\n+ * Create on-disk data structures that are required for this source to operate\n+ * correctly. Returns 0 on success, a negative error code otherwise.\n+ */\n+static inline int odb_source_create_on_disk(struct odb_source *source)\n+{\n+\tif (!source->create_on_disk)\n+\t\treturn 0;\n+\treturn source->create_on_disk(source);\n+}\n+\n /*\n  * Prepare the object database source and clear any caches. Depending on the\n  * backend used this may have the effect that concurrently-written objects\ndiff --git a/setup.c b/setup.c\nindex 97338cbc51..ace3c59d18 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -2649,25 +2649,27 @@ static int create_default_files(struct repository *repo,\n \n static void create_object_database(struct repository *repo)\n {\n-\tstruct strbuf path = STRBUF_INIT;\n-\tsize_t baselen;\n+\t/*\n+\t * Create the \"objects\" directory in the common directory. This is done\n+\t * so that the repository can be discovered regardless of the backend\n+\t * used.\n+\t *\n+\t * Note that we only do this in case the object directory wasn't\n+\t * overwritten via an environment variable. If it _is_ being overridden\n+\t * then we skip this step, as the repository won't be discoverable\n+\t * anyway without the environment variable.\n+\t */\n+\tif (!getenv(DB_ENVIRONMENT)) {\n+\t\tstruct strbuf objects_dir = STRBUF_INIT;\n+\t\trepo_common_path_append(repo, &objects_dir, \"objects\");\n+\t\tsafe_create_dir(repo, objects_dir.buf, 1);\n+\t\tstrbuf_release(&objects_dir);\n+\t}\n \n \trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n \n-\tstrbuf_addstr(&path, repo_get_object_directory(repo));\n-\tbaselen = path.len;\n-\n-\tsafe_create_dir(repo, path.buf, 1);\n-\n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/pack\");\n-\tsafe_create_dir(repo, path.buf, 1);\n-\n-\tstrbuf_setlen(&path, baselen);\n-\tstrbuf_addstr(&path, \"/info\");\n-\tsafe_create_dir(repo, path.buf, 1);\n-\n-\tstrbuf_release(&path);\n+\tif (odb_source_create_on_disk(repo->objects->sources) < 0)\n+\t\tdie(_(\"failed creating object database\"));\n }\n \n static void separate_git_dir(const char *git_dir, const char *git_link)\n\n-- \n2.55.0.679.g6767b8d81c.dirty\n\n"},{"id":"549924","messageId":"xmqq1pcah7vg.fsf@gitster.g","threadId":"66056","inReplyTo":"20260807-pks-odb-create-on-disk-v5-4-399da0b0b140@pks.im","subject":"Re: [PATCH v5 4/6] setup: defer object database creation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-07T04:28:19Z","receivedAt":"2026-08-07T04:28:22Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> diff --git a/setup.c b/setup.c\n> index 5dfab3e79e..97338cbc51 100644\n> --- a/setup.c\n> +++ b/setup.c\n> @@ -1888,6 +1882,7 @@ const char *enter_repo(struct repository *repo, const char *path, unsigned flags\n>  \t\tread_and_verify_repository_format(&fmt, \".\", NULL);\n>  \t\tif (apply_repository_format(repo, &fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n>  \t\t\tdie(\"%s\", err.buf);\n> +\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n>  \t\tstartup_info->have_repository = 1;\n>  \n>  \t\tclear_repository_format(&fmt);\n\nThe previous round corrected the overly long line while at it, but\nit is no longer done here.\n\nWhich is OK either way.\n\n> @@ -2090,6 +2085,7 @@ const char *setup_git_directory_gently(struct repository *repo, int *nongit_ok)\n>  \t\t\tif (apply_repository_format(repo, &discovery.format,\n>  \t\t\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n>  \t\t\t\tdie(\"%s\", err.buf);\n> +\t\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n\nLooks like the differences since the last round is truly minimum ;-)\n"},{"id":"549946","messageId":"87mruy4czd.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"xmqq1pcah7vg.fsf@gitster.g","subject":"Re: [PATCH v5 4/6] setup: defer object database creation","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-08-07T07:16:22Z","receivedAt":"2026-08-07T07:16:38Z","isPatch":true,"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Patrick Steinhardt <ps@pks.im> writes:\n>\n>> diff --git a/setup.c b/setup.c\n>> index 5dfab3e79e..97338cbc51 100644\n>> --- a/setup.c\n>> +++ b/setup.c\n>> @@ -1888,6 +1882,7 @@ const char *enter_repo(struct repository *repo, const char *path, unsigned flags\n>>  \t\tread_and_verify_repository_format(&fmt, \".\", NULL);\n>>  \t\tif (apply_repository_format(repo, &fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n>>  \t\t\tdie(\"%s\", err.buf);\n>> +\t\trepo->objects = odb_new(repo, ODB_NEW_HONOR_ENV);\n>>  \t\tstartup_info->have_repository = 1;\n>>  \n>>  \t\tclear_repository_format(&fmt);\n>\n> The previous round corrected the overly long line while at it, but\n> it is no longer done here.\n\nYeah, I've asked about this. In [PATCH v3 4/6] this change existed:\n\n-\tif (apply_repository_format(repo, &repo_fmt, APPLY_REPOSITORY_FORMAT_HONOR_ENV, &err) < 0)\n+\tif (apply_repository_format(repo, &repo_fmt,\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_HONOR_ENV |\n+\t\t\t\t    APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION, &err) < 0)\n\nBut adding APPLY_REPOSITORY_FORMAT_SKIP_ODB_CREATION was reverted in v4,\nbut that version still had the reformatting change (fixing the overly\nlong line).\n\nThere are multiple occurrences of this overly long line, but only this\none was changed in v4. So Patrick reverted changing the overly long line\nin v5, which I think is better.\n\n-- \nCheers,\nToon\n"},{"id":"549947","messageId":"87jyq24cxm.fsf@emacs.iotcl.com","threadId":"66056","inReplyTo":"20260807-pks-odb-create-on-disk-v5-0-399da0b0b140@pks.im","subject":"Re: [PATCH v5 0/6] odb: make creation of object database pluggable","fromName":"Toon Claes","fromEmail":"toon@iotcl.com","sentAt":"2026-08-07T07:17:25Z","receivedAt":"2026-08-07T07:17:33Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> Hi,\n>\n> when creating a new repository we create a couple of on-disk data\n> structures for the object database. This includes the \"objects/\"\n> directory hierarchy with \"objects/info\" and \"objects/pack\", which are\n> specific to the backend.\n>\n> This patch series makes the creation of the on-disk data structures\n> pluggable. While we continue to always create \"objects/\" regardless of\n> the backend (it's required for a repository to be recognized as such),\n> the other subdirectories are now created by the backend. This will allow\n> other backends to plug in their own logic.\n>\n> The series starts with a small detour into the loose-object map. This\n> detour is required so that we can defer initialization of the object\n> database itself to a later point in time.\n>\n> The series is based on 9a0c4701dc (The 7th batch, 2026-07-22).\n>\n> Changes in v5:\n>   - Remove a leftover formatting change.\n>   - Fix a stale comment.\n>   - Link to v4: https://patch.msgid.link/20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im\n\nI'm completely happy with this version, thanks for bearing with me.\n\n-- \nCheers,\nToon\n"},{"id":"549967","messageId":"anWhA5zZK2eg1h47@pks.im","threadId":"66056","inReplyTo":"87jyq24cxm.fsf@emacs.iotcl.com","subject":"Re: [PATCH v5 0/6] odb: make creation of object database pluggable","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-07T09:10:27Z","receivedAt":"2026-08-07T09:10:34Z","isPatch":true,"body":"On Fri, Aug 07, 2026 at 09:17:25AM +0200, Toon Claes wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > Hi,\n> >\n> > when creating a new repository we create a couple of on-disk data\n> > structures for the object database. This includes the \"objects/\"\n> > directory hierarchy with \"objects/info\" and \"objects/pack\", which are\n> > specific to the backend.\n> >\n> > This patch series makes the creation of the on-disk data structures\n> > pluggable. While we continue to always create \"objects/\" regardless of\n> > the backend (it's required for a repository to be recognized as such),\n> > the other subdirectories are now created by the backend. This will allow\n> > other backends to plug in their own logic.\n> >\n> > The series starts with a small detour into the loose-object map. This\n> > detour is required so that we can defer initialization of the object\n> > database itself to a later point in time.\n> >\n> > The series is based on 9a0c4701dc (The 7th batch, 2026-07-22).\n> >\n> > Changes in v5:\n> >   - Remove a leftover formatting change.\n> >   - Fix a stale comment.\n> >   - Link to v4: https://patch.msgid.link/20260806-pks-odb-create-on-disk-v4-0-ba8b4fdd2e3c@pks.im\n> \n> I'm completely happy with this version, thanks for bearing with me.\n\nThanks for your reviews!\n\nPatrick\n"}]}