{"thread":{"id":"66423","subject":"[PATCH RFC 0/5] Add --dry-run option to git-backfill(1)","startedAt":"2026-09-30T00:21:56Z","lastAt":"2026-09-30T20:03:45Z","messageCount":18,"participants":["Pablo Sabater","Karthik Nayak","Junio C Hamano","Derrick Stolee"],"isPatch":true,"patchVersion":1,"patchTotal":5},"messages":[{"id":"553653","messageId":"20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com","threadId":"66423","inReplyTo":null,"subject":"[PATCH RFC 0/5] Add --dry-run option to git-backfill(1)","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T00:21:45Z","receivedAt":"2026-09-30T00:21:56Z","isPatch":true,"body":"[Cc'd Derrick Stolee for his work in the backfill(1) command]\n\nThis series adds a --dry-run option to git-backfill(1) that reports how\nmany missing blobs would be fetched and, when the remote server\nsupports the object-info capability, their total size:\n\n        $ git backfill --dry-run\n        After backfill, 48 blobs would be fetched (1.20 KiB).\n\nIf the server does not advertise object-info, only the count is shown.\n\nI am not a git-backfill(1) user myself, but it seemed useful for users\nto know how much data a backfill would bring in before running it.\n\nThe number of missing blobs is the sum of the number of blobs to be\nfetched in each batch. The object-info capability lets us ask the server\nfor the size of each blob without downloading it, so summing them gives\nan estimate of the total.\n\nNote that this is an upper bound rather than the exact disk usage:\nobject-info reports the uncompressed size of each object, while the\nobjects end up stored compressed and possibly deltified in a packfile,\nso the space actually used on disk will usually be smaller.\n\nSince I do not use backfill, feedback on whether this is useful, and on\nthe output format, is very welcome.\n\nPatches 1-3 are preparatory:\n\n  [1/5] transport-internal: update fetch_object_info comment\n        Fixes an outdated comment: object-info supports type as well as\n        size.\n\n  [2/5] fetch-object-info: add enum for fetch_object_info() statuses\n  [3/5] fetch-object-info: return a status instead of dying\n        Teach fetch_object_info() to return a status instead of dying\n        when the server does not advertise object-info. The die() is\n        kept in cat-file's remote-object-info path, so its behavior is\n        unchanged.\n\nPatches 4-5 add the option in two steps:\n\n  [4/5] backfill: add --dry-run option\n        Prints only the number of blobs that would be fetched.\n\n  [5/5] backfill: report total size of missing blobs in --dry-run\n        Also prints their total size when the server supports\n        object-info, and falls back to the count alone otherwise.\n\nThanks.\n\nSigned-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n---\nPablo Sabater (5):\n      transport-internal: update fetch_object_info comment\n      fetch-object-info: add enum for fetch_object_info() statuses\n      fetch-object-info: return a status instead of dying\n      backfill: add --dry-run option\n      backfill: report total size of missing blobs in --dry-run\n\n Documentation/git-backfill.adoc | 10 ++++-\n builtin/backfill.c              | 97 +++++++++++++++++++++++++++++++++++++++--\n builtin/cat-file.c              |  4 ++\n fetch-object-info.c             | 17 ++++----\n fetch-object-info.h             | 24 +++++++---\n t/t5620-backfill.sh             | 48 ++++++++++++++++++++\n transport-helper.c              |  6 +--\n transport-internal.h            | 12 ++---\n transport.c                     | 29 ++++++------\n transport.h                     |  7 +--\n 10 files changed, 208 insertions(+), 46 deletions(-)\n\n\n---\nbase-commit: 12cb6293d6288865c1a133cf22accbaf99d13eb6\nchange-id: 20260914-backfill-dryrun-fac997003322\n\n"},{"id":"553654","messageId":"20260930-backfill-dryrun-v1-1-1128f247ee01@gmail.com","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com","subject":"[PATCH RFC 1/5] transport-internal: update fetch_object_info comment","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T00:21:46Z","receivedAt":"2026-09-30T00:21:57Z","isPatch":true,"body":"The comment describing the fetch_object_info() callback in struct\ntransport_vtable says that only the object size can be fetched. The\nobject-info capability can now also report the object type, or\nneither, to only check whether an object exists on the remote.\n\nUpdate the comment.\n\nSigned-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n---\n transport-internal.h | 4 ++--\n 1 file changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/transport-internal.h b/transport-internal.h\nindex a10b27cc81..626ceaae2b 100644\n--- a/transport-internal.h\n+++ b/transport-internal.h\n@@ -48,8 +48,8 @@ struct transport_vtable {\n \tint (*fetch_refs)(struct transport *transport, int refs_nr, struct ref **refs);\n \n \t/*\n-\t * Fetch object info (only size currently) from remote without\n-\t * downloading the objects.\n+\t * Fetch object info (size, type, or none of them to only check\n+\t * for existence) from the remote without downloading the objects.\n \t *\n \t * Uses object-info capability of v2 protocol.\n \t */\n\n-- \n2.54.0\n\n"},{"id":"553655","messageId":"20260930-backfill-dryrun-v1-2-1128f247ee01@gmail.com","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com","subject":"[PATCH RFC 2/5] fetch-object-info: add enum for fetch_object_info() statuses","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T00:21:47Z","receivedAt":"2026-09-30T00:21:58Z","isPatch":true,"body":"fetch_object_info() dies when the server does not advertise the\nobject-info capability. That is fine for git cat-file\nremote-object-info command, which cannot work without it. However\na subsequent commit needs fetch_object_info() to not die, to be\nable to fallback.\n\nAdd \"enum fetch_object_info_status\" so that fetch_object_info() can\nreport this case to its callers. It is used in a subsequent commit.\n\nSigned-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n---\n fetch-object-info.h | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/fetch-object-info.h b/fetch-object-info.h\nindex 2fba96c6f7..663a7f3ae7 100644\n--- a/fetch-object-info.h\n+++ b/fetch-object-info.h\n@@ -16,6 +16,12 @@ struct fetch_object_info_results {\n \n #define FETCH_OBJECT_INFO_RESULTS_INIT { 0 }\n \n+enum fetch_object_info_status {\n+\tFETCH_OBJECT_INFO_OK = 0,\n+\tFETCH_OBJECT_INFO_ERR = -1,\n+\tFETCH_OBJECT_INFO_NOT_ENABLED = -2,\n+};\n+\n struct oid_array;\n /*\n  * Sends git-cat-file object-info command into the request buf and reads the\n\n-- \n2.54.0\n\n"},{"id":"553656","messageId":"20260930-backfill-dryrun-v1-3-1128f247ee01@gmail.com","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com","subject":"[PATCH RFC 3/5] fetch-object-info: return a status instead of dying","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T00:21:48Z","receivedAt":"2026-09-30T00:21:59Z","isPatch":true,"body":"A subsequent commit needs fetch_object_info() not to die() when the\nobject-info capability is not enabled on the server, so that it can\nfall back.\n\nMake fetch_object_info() return FETCH_OBJECT_INFO_NOT_ENABLED instead\nof die()'ing when the server does not advertise the object-info\ncapability, and propagate the status through the transport layer so\nthat callers of transport_fetch_object_info() can act on it. It is now\nup to them whether to die() or fall back.\n\ncat-file now dies by itself on FETCH_OBJECT_INFO_NOT_ENABLED, so its\nbehavior is unchanged.\n\nSigned-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n---\n builtin/cat-file.c   |  4 ++++\n fetch-object-info.c  | 17 +++++++++--------\n fetch-object-info.h  | 18 +++++++++++-------\n transport-helper.c   |  6 +++---\n transport-internal.h |  8 ++++----\n transport.c          | 29 +++++++++++++++--------------\n transport.h          |  7 ++++---\n 7 files changed, 50 insertions(+), 39 deletions(-)\n\ndiff --git a/builtin/cat-file.c b/builtin/cat-file.c\nindex 8870a210ec..f4758f2203 100644\n--- a/builtin/cat-file.c\n+++ b/builtin/cat-file.c\n@@ -726,6 +726,10 @@ static int get_remote_info(int argc,\n \n \tretval = transport_fetch_object_info(gtransport, object_info_oids,\n \t\t\t\t\t     results);\n+\n+\tif (retval == FETCH_OBJECT_INFO_NOT_ENABLED)\n+\t\tdie(_(\"object-info capability is not enabled on the server\"));\n+\n cleanup:\n \ttransport_disconnect(gtransport);\n \treturn retval;\ndiff --git a/fetch-object-info.c b/fetch-object-info.c\nindex 0a58308f9b..7e4c922d27 100644\n--- a/fetch-object-info.c\n+++ b/fetch-object-info.c\n@@ -52,13 +52,13 @@ static int parse_object_size(const char *s, size_t *res)\n \treturn 0;\n }\n \n-void fetch_object_info(const enum protocol_version version,\n-\t\t       const struct string_list *server_options,\n-\t\t       const struct oid_array *oids,\n-\t\t       struct packet_reader *reader,\n-\t\t       struct fetch_object_info_results *results,\n-\t\t       const int stateless_rpc,\n-\t\t       const int fd_out)\n+enum fetch_object_info_status fetch_object_info(const enum protocol_version version,\n+\t\t\t\t\t\tconst struct string_list *server_options,\n+\t\t\t\t\t\tconst struct oid_array *oids,\n+\t\t\t\t\t\tstruct packet_reader *reader,\n+\t\t\t\t\t\tstruct fetch_object_info_results *results,\n+\t\t\t\t\t\tconst int stateless_rpc,\n+\t\t\t\t\t\tconst int fd_out)\n {\n \tunsigned ask_size = 0;\n \tunsigned ask_type = 0;\n@@ -72,7 +72,7 @@ void fetch_object_info(const enum protocol_version version,\n \tswitch (version) {\n \tcase protocol_v2:\n \t\tif (!server_supports_v2(\"object-info\"))\n-\t\t\tdie(_(\"object-info capability is not enabled on the server\"));\n+\t\t\treturn FETCH_OBJECT_INFO_NOT_ENABLED;\n \n \t\tif (results->wants_size &&\n \t\t    server_supports_feature(\"object-info\", \"size\", 0))\n@@ -188,6 +188,7 @@ void fetch_object_info(const enum protocol_version version,\n \t\t    (uintmax_t)oids->nr);\n \n \tcheck_stateless_delimiter(stateless_rpc, reader, \"stateless delimiter expected\");\n+\treturn FETCH_OBJECT_INFO_OK;\n }\n \n void free_fetch_object_info_results(struct fetch_object_info_results *results)\ndiff --git a/fetch-object-info.h b/fetch-object-info.h\nindex 663a7f3ae7..9d2750bc93 100644\n--- a/fetch-object-info.h\n+++ b/fetch-object-info.h\n@@ -32,14 +32,18 @@ struct oid_array;\n  * the server both advertised and answered with. An array left NULL means the\n  * attribute is not available.\n  * Release them with free_fetch_object_info_results().\n+ *\n+ * Returns FETCH_OBJECT_INFO_NOT_ENABLED if the server does not advertise the\n+ * object-info capability, FETCH_OBJECT_INFO_OK otherwise.\n+ * die()'s on any other error.\n  */\n-void fetch_object_info(enum protocol_version version,\n-\t\t       const struct string_list *server_options,\n-\t\t       const struct oid_array *oids,\n-\t\t       struct packet_reader *reader,\n-\t\t       struct fetch_object_info_results *results,\n-\t\t       int stateless_rpc,\n-\t\t       int fd_out);\n+enum fetch_object_info_status fetch_object_info(enum protocol_version version,\n+\t\t\t\t\t\tconst struct string_list *server_options,\n+\t\t\t\t\t\tconst struct oid_array *oids,\n+\t\t\t\t\t\tstruct packet_reader *reader,\n+\t\t\t\t\t\tstruct fetch_object_info_results *results,\n+\t\t\t\t\t\tint stateless_rpc,\n+\t\t\t\t\t\tint fd_out);\n \n void free_fetch_object_info_results(struct fetch_object_info_results *results);\n \ndiff --git a/transport-helper.c b/transport-helper.c\nindex d5a064d386..855b53da59 100644\n--- a/transport-helper.c\n+++ b/transport-helper.c\n@@ -786,9 +786,9 @@ static int fetch_refs(struct transport *transport,\n \treturn -1;\n }\n \n-static int fetch_object_info_helper(struct transport *transport,\n-\t\t\t\t    const struct oid_array *oids,\n-\t\t\t\t    struct fetch_object_info_results *results)\n+static enum fetch_object_info_status fetch_object_info_helper(struct transport *transport,\n+\t\t\t\t\t\t\t      const struct oid_array *oids,\n+\t\t\t\t\t\t\t      struct fetch_object_info_results *results)\n {\n \tget_helper(transport);\n \tif (process_connect(transport, 0))\ndiff --git a/transport-internal.h b/transport-internal.h\nindex 626ceaae2b..067134081c 100644\n--- a/transport-internal.h\n+++ b/transport-internal.h\n@@ -2,13 +2,13 @@\n #define TRANSPORT_INTERNAL_H\n \n #include \"connect.h\"\n+#include \"fetch-object-info.h\"\n \n struct ref;\n struct transport;\n struct strvec;\n struct transport_ls_refs_options;\n struct oid_array;\n-struct fetch_object_info_results;\n \n struct transport_vtable {\n \t/**\n@@ -53,9 +53,9 @@ struct transport_vtable {\n \t *\n \t * Uses object-info capability of v2 protocol.\n \t */\n-\tint (*fetch_object_info)(struct transport *transport,\n-\t\t\t\t const struct oid_array *oids,\n-\t\t\t\t struct fetch_object_info_results *results);\n+\tenum fetch_object_info_status (*fetch_object_info)(struct transport *transport,\n+\t\t\t\t\t\t\t   const struct oid_array *oids,\n+\t\t\t\t\t\t\t   struct fetch_object_info_results *results);\n \n \t/**\n \t * Push the objects and refs. Send the necessary objects, and\ndiff --git a/transport.c b/transport.c\nindex 25e2c14a7b..561764cb6a 100644\n--- a/transport.c\n+++ b/transport.c\n@@ -433,11 +433,11 @@ static int get_bundle_uri(struct transport *transport)\n \t\t\t\t     transport->bundles, stateless_rpc);\n }\n \n-static int fetch_object_info_via_pack(struct transport *transport,\n-\t\t\t\t      const struct oid_array *oids,\n-\t\t\t\t      struct fetch_object_info_results *results)\n+static enum fetch_object_info_status fetch_object_info_via_pack(struct transport *transport,\n+\t\t\t\t\t\t\t\tconst struct oid_array *oids,\n+\t\t\t\t\t\t\t\tstruct fetch_object_info_results *results)\n {\n-\tint ret = 0;\n+\tenum fetch_object_info_status ret = FETCH_OBJECT_INFO_OK;\n \tstruct git_transport_data *data = transport->data;\n \tstruct packet_reader reader;\n \n@@ -450,26 +450,27 @@ static int fetch_object_info_via_pack(struct transport *transport,\n \tdata->version = discover_version(&reader);\n \ttransport->hash_algo = reader.hash_algo;\n \n-\tfetch_object_info(data->version,\n-\t\t\t  transport->server_options,\n-\t\t\t  oids,\n-\t\t\t  &reader,\n-\t\t\t  results,\n-\t\t\t  transport->stateless_rpc, data->fd[1]);\n+\tret = fetch_object_info(data->version,\n+\t\t\t\ttransport->server_options,\n+\t\t\t\toids,\n+\t\t\t\t&reader,\n+\t\t\t\tresults,\n+\t\t\t\ttransport->stateless_rpc,\n+\t\t\t\tdata->fd[1]);\n \n \tclose(data->fd[0]);\n \tif (data->fd[1] >= 0)\n \t\tclose(data->fd[1]);\n \tif (finish_connect(data->conn))\n-\t\tret = -1;\n+\t\tret = FETCH_OBJECT_INFO_ERR;\n \tdata->conn = NULL;\n \n \treturn ret;\n }\n \n-int transport_fetch_object_info(struct transport *transport,\n-\t\t\t\tconst struct oid_array *oids,\n-\t\t\t\tstruct fetch_object_info_results *results)\n+enum fetch_object_info_status transport_fetch_object_info(struct transport *transport,\n+\t\t\t\t\t\t\t  const struct oid_array *oids,\n+\t\t\t\t\t\t\t  struct fetch_object_info_results *results)\n {\n \tif (!transport->vtable->fetch_object_info)\n \t\tdie(_(\"remote does not support object-info\"));\ndiff --git a/transport.h b/transport.h\nindex 39193d0077..c1671639d6 100644\n--- a/transport.h\n+++ b/transport.h\n@@ -1,6 +1,7 @@\n #ifndef TRANSPORT_H\n #define TRANSPORT_H\n \n+#include \"fetch-object-info.h\"\n #include \"run-command.h\"\n #include \"remote.h\"\n #include \"list-objects-filter-options.h\"\n@@ -314,9 +315,9 @@ int transport_fetch_refs(struct transport *transport, struct ref *refs);\n /*\n  * Fetch the object info from remote\n  */\n-int transport_fetch_object_info(struct transport *transport,\n-\t\t\t\tconst struct oid_array *oids,\n-\t\t\t\tstruct fetch_object_info_results *results);\n+enum fetch_object_info_status transport_fetch_object_info(struct transport *transport,\n+\t\t\t\t\t\t\t  const struct oid_array *oids,\n+\t\t\t\t\t\t\t  struct fetch_object_info_results *results);\n \n /*\n  * If this flag is set, unlocking will avoid to call non-async-signal-safe\n\n-- \n2.54.0\n\n"},{"id":"553657","messageId":"20260930-backfill-dryrun-v1-4-1128f247ee01@gmail.com","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com","subject":"[PATCH RFC 4/5] backfill: add --dry-run option","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T00:21:49Z","receivedAt":"2026-09-30T00:22:00Z","isPatch":true,"body":"Users have no way to know how many blobs git backfill is going to\ndownload before running it.\n\nAdd a new --dry-run option to the backfill command. The objects are\nwalked as usual, but instead of fetching each batch of missing blobs\nthey are only counted, and the total is printed at the end.\n\nA subsequent commit will also print their size when the server supports\nthe object-info capability.\n\nSigned-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n---\n Documentation/git-backfill.adoc |  6 +++++-\n builtin/backfill.c              | 37 +++++++++++++++++++++++++++++++++----\n t/t5620-backfill.sh             | 26 ++++++++++++++++++++++++++\n 3 files changed, 64 insertions(+), 5 deletions(-)\n\ndiff --git a/Documentation/git-backfill.adoc b/Documentation/git-backfill.adoc\nindex 82d6a1969d..08f19fea17 100644\n--- a/Documentation/git-backfill.adoc\n+++ b/Documentation/git-backfill.adoc\n@@ -9,7 +9,7 @@ git-backfill - Download missing objects in a partial clone\n SYNOPSIS\n --------\n [synopsis]\n-git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]\n+git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]\n \n DESCRIPTION\n -----------\n@@ -70,6 +70,10 @@ OPTIONS\n \t--onto TARGET A..B`, where A..B normally excludes A but you need\n \tthe blobs from A as well.  `--include-edges` is the default.\n \n+`--dry-run`::\n+\tDo not download any objects. Instead, print the number of\n+\tmissing blobs that would be downloaded.\n+\n `<revision-range>`::\n \tBackfill only blobs reachable from commits in the specified\n \trevision range.  When no _<revision-range>_ is specified, it\ndiff --git a/builtin/backfill.c b/builtin/backfill.c\nindex e71e0f4742..6019112966 100644\n--- a/builtin/backfill.c\n+++ b/builtin/backfill.c\n@@ -26,7 +26,7 @@\n #include \"path-walk.h\"\n \n static const char * const builtin_backfill_usage[] = {\n-\tN_(\"git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]\"),\n+\tN_(\"git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]\"),\n \tNULL\n };\n \n@@ -36,6 +36,8 @@ struct backfill_context {\n \tsize_t min_batch_size;\n \tint sparse;\n \tint include_edges;\n+\tint dry_run;\n+\tsize_t total_batch_nr;\n \tstruct rev_info revs;\n };\n \n@@ -58,6 +60,15 @@ static void download_batch(struct backfill_context *ctx)\n \todb_reprepare(ctx->repo->objects);\n }\n \n+static void dry_run_batch(struct backfill_context *ctx)\n+{\n+\tif (!ctx->current_batch.nr)\n+\t\treturn;\n+\n+\tctx->total_batch_nr += ctx->current_batch.nr;\n+\toid_array_clear(&ctx->current_batch);\n+}\n+\n static int fill_missing_blobs(const char *path UNUSED,\n \t\t\t      struct oid_array *list,\n \t\t\t      enum object_type type,\n@@ -73,8 +84,12 @@ static int fill_missing_blobs(const char *path UNUSED,\n \t\t\toid_array_append(&ctx->current_batch, &list->oid[i]);\n \t}\n \n-\tif (ctx->current_batch.nr >= ctx->min_batch_size)\n-\t\tdownload_batch(ctx);\n+\tif (ctx->current_batch.nr >= ctx->min_batch_size) {\n+\t\tif (ctx->dry_run)\n+\t\t\tdry_run_batch(ctx);\n+\t\telse\n+\t\t\tdownload_batch(ctx);\n+\t}\n \n \treturn 0;\n }\n@@ -131,10 +146,23 @@ static int do_backfill(struct backfill_context *ctx)\n \n \tret = walk_objects_by_path(&info);\n \n+\tif (ret)\n+\t\tgoto end;\n+\n \t/* Download the objects that did not fill a batch. */\n-\tif (!ret)\n+\tif (!ctx->dry_run) {\n \t\tdownload_batch(ctx);\n+\t\tgoto end;\n+\t}\n+\n+\tdry_run_batch(ctx);\n+\n+\tprintf(Q_(\"After backfill, %\" PRIuMAX \" blob would be fetched.\\n\",\n+\t\t  \"After backfill, %\" PRIuMAX \" blobs would be fetched.\\n\",\n+\t\t  (unsigned long)ctx->total_batch_nr),\n+\t       (uintmax_t)ctx->total_batch_nr);\n \n+end:\n \tpath_walk_info_clear(&info);\n \treturn ret;\n }\n@@ -157,6 +185,7 @@ int cmd_backfill(int argc, const char **argv, const char *prefix, struct reposit\n \t\t\t N_(\"Restrict the missing objects to the current sparse-checkout\")),\n \t\tOPT_BOOL(0, \"include-edges\", &ctx.include_edges,\n \t\t\t N_(\"Include blobs from boundary commits in the backfill\")),\n+\t\tOPT__DRY_RUN(&ctx.dry_run, N_(\"Preview the number of blobs to be fetched\")),\n \t\tOPT_END(),\n \t};\n \tstruct repo_config_values *cfg = repo_config_values(the_repository);\ndiff --git a/t/t5620-backfill.sh b/t/t5620-backfill.sh\nindex 7462280470..e76fa6081b 100755\n--- a/t/t5620-backfill.sh\n+++ b/t/t5620-backfill.sh\n@@ -141,6 +141,32 @@ test_expect_success 'do partial clone 2, backfill min batch size' '\n \ttest_line_count = 0 revs2\n '\n \n+test_expect_success '--dry-run reports missing blobs without fetching them' '\n+\ttest_when_finished \"rm -rf backfill-dry-run dry-trace\" &&\n+\tgit clone --no-checkout --filter=blob:none \\\n+\t\t--single-branch --branch=main \\\n+\t\t\"file://$(pwd)/srv.bare\" backfill-dry-run &&\n+\n+\tGIT_TRACE2_EVENT=\"$(pwd)/dry-trace\" git \\\n+\t\t-C backfill-dry-run backfill --dry-run >out &&\n+\n+\ttest_grep \"48 blobs would be fetched\" out &&\n+\ttest_grep ! fetch_count dry-trace &&\n+\tgit -C backfill-dry-run rev-list --quiet --objects --missing=print HEAD >missing &&\n+\ttest_line_count = 48 missing\n+'\n+\n+test_expect_success '--dry-run with no missing blobs' '\n+\ttest_when_finished rm -rf backfill-dry-run &&\n+\tgit clone --no-checkout --filter=blob:none \\\n+\t\t--single-branch --branch=main \\\n+\t\t\"file://$(pwd)/srv.bare\" backfill-dry-run &&\n+\tgit -C backfill-dry-run backfill &&\n+\n+\tgit -C backfill-dry-run backfill --dry-run >out &&\n+\ttest_grep \"0 blobs would be fetched\" out\n+'\n+\n test_expect_success 'backfill --sparse without sparse-checkout fails' '\n \tgit init not-sparse &&\n \ttest_must_fail git -C not-sparse backfill --sparse 2>err &&\n\n-- \n2.54.0\n\n"},{"id":"553658","messageId":"20260930-backfill-dryrun-v1-5-1128f247ee01@gmail.com","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com","subject":"[PATCH RFC 5/5] backfill: report total size of missing blobs in --dry-run","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T00:21:50Z","receivedAt":"2026-09-30T00:22:02Z","isPatch":true,"body":"The number of missing blobs says little about how much data a backfill\nwill transfer.  In --dry-run mode, if the server has the object-info\ncapability enabled, fetch the size of the objects to be fetched and show\nthe total size, if the server doesn't have the capability enabled, fall\nback to printing only the count.\n\nThe reported size is the sum of the uncompressed object sizes, so it\nis an upper bound: once fetched, the blobs are stored compressed and\npossibly deltified in a packfile, and usually take less space on disk.\n\nSigned-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n---\n Documentation/git-backfill.adoc |  6 +++-\n builtin/backfill.c              | 70 ++++++++++++++++++++++++++++++++++++++---\n t/t5620-backfill.sh             | 22 +++++++++++++\n 3 files changed, 92 insertions(+), 6 deletions(-)\n\ndiff --git a/Documentation/git-backfill.adoc b/Documentation/git-backfill.adoc\nindex 08f19fea17..660e1c17c0 100644\n--- a/Documentation/git-backfill.adoc\n+++ b/Documentation/git-backfill.adoc\n@@ -72,7 +72,11 @@ OPTIONS\n \n `--dry-run`::\n \tDo not download any objects. Instead, print the number of\n-\tmissing blobs that would be downloaded.\n+\tmissing blobs that would be downloaded and, if the promisor\n+\tremote supports the `object-info` capability, their total\n+\tsize. This is the sum of the uncompressed sizes of the blobs,\n+\tso the space used on disk after a real backfill is usually\n+\tsmaller.\n \n `<revision-range>`::\n \tBackfill only blobs reachable from commits in the specified\ndiff --git a/builtin/backfill.c b/builtin/backfill.c\nindex 6019112966..caf1a64e95 100644\n--- a/builtin/backfill.c\n+++ b/builtin/backfill.c\n@@ -24,6 +24,9 @@\n #include \"progress.h\"\n #include \"packfile.h\"\n #include \"path-walk.h\"\n+#include \"transport.h\"\n+#include \"remote.h\"\n+#include \"fetch-object-info.h\"\n \n static const char * const builtin_backfill_usage[] = {\n \tN_(\"git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]\"),\n@@ -37,7 +40,11 @@ struct backfill_context {\n \tint sparse;\n \tint include_edges;\n \tint dry_run;\n+\tint object_info_enabled;\n+\tsize_t total_batch_size;\n \tsize_t total_batch_nr;\n+\tstruct transport *object_info_transport;\n+\tstruct fetch_object_info_results object_info_results;\n \tstruct rev_info revs;\n };\n \n@@ -62,10 +69,47 @@ static void download_batch(struct backfill_context *ctx)\n \n static void dry_run_batch(struct backfill_context *ctx)\n {\n+\tstruct fetch_object_info_results *results = &ctx->object_info_results;\n+\tenum fetch_object_info_status status;\n+\n \tif (!ctx->current_batch.nr)\n \t\treturn;\n \n \tctx->total_batch_nr += ctx->current_batch.nr;\n+\n+\tif (!ctx->object_info_enabled)\n+\t\tgoto cleanup;\n+\n+\tif (!ctx->object_info_transport) {\n+\t\tstruct promisor_remote *promise =\n+\t\t\trepo_promisor_remote_find(ctx->repo, NULL);\n+\t\tstruct remote *remote = NULL;\n+\n+\t\tif (!promise || !(remote = remote_get(promise->name)))\n+\t\t\tdie(_(\"--dry-run requires a promisor remote\"));\n+\n+\t\tctx->object_info_transport = transport_get(remote, NULL);\n+\n+\t\tif (!ctx->object_info_transport->smart_options)\n+\t\t\tdie(_(\"failed to get object info: smart options required\"));\n+\t}\n+\n+\tresults->wants_size = 1;\n+\tstatus = transport_fetch_object_info(ctx->object_info_transport,\n+\t\t\t\t\t     &ctx->current_batch,\n+\t\t\t\t\t     results);\n+\n+\tif (status == FETCH_OBJECT_INFO_NOT_ENABLED ||\n+\t    !results->sizes) {\n+\t\tctx->object_info_enabled = 0;\n+\t\tgoto cleanup;\n+\t}\n+\n+\tfor (size_t i = 0; i < results->nr; i++)\n+\t\tctx->total_batch_size += results->sizes[i];\n+\n+cleanup:\n+\tfree_fetch_object_info_results(&ctx->object_info_results);\n \toid_array_clear(&ctx->current_batch);\n }\n \n@@ -157,12 +201,25 @@ static int do_backfill(struct backfill_context *ctx)\n \n \tdry_run_batch(ctx);\n \n-\tprintf(Q_(\"After backfill, %\" PRIuMAX \" blob would be fetched.\\n\",\n-\t\t  \"After backfill, %\" PRIuMAX \" blobs would be fetched.\\n\",\n-\t\t  (unsigned long)ctx->total_batch_nr),\n-\t       (uintmax_t)ctx->total_batch_nr);\n+\tif (ctx->object_info_enabled && ctx->total_batch_nr) {\n+\t\tstruct strbuf size = STRBUF_INIT;\n+\n+\t\tstrbuf_humanise_bytes(&size, ctx->total_batch_size);\n+\t\tprintf(Q_(\"After backfill, %\" PRIuMAX \" blob would be fetched (%s).\\n\",\n+\t\t\t  \"After backfill, %\" PRIuMAX \" blobs would be fetched (%s).\\n\",\n+\t\t\t  (unsigned long)ctx->total_batch_nr),\n+\t\t       (uintmax_t)ctx->total_batch_nr, size.buf);\n+\t\tstrbuf_release(&size);\n+\t} else {\n+\t\tprintf(Q_(\"After backfill, %\" PRIuMAX \" blob would be fetched.\\n\",\n+\t\t\t  \"After backfill, %\" PRIuMAX \" blobs would be fetched.\\n\",\n+\t\t\t  (unsigned long)ctx->total_batch_nr),\n+\t\t       (uintmax_t)ctx->total_batch_nr);\n+\t}\n \n end:\n+\tif (ctx->object_info_transport)\n+\t\ttransport_disconnect(ctx->object_info_transport);\n \tpath_walk_info_clear(&info);\n \treturn ret;\n }\n@@ -177,6 +234,8 @@ int cmd_backfill(int argc, const char **argv, const char *prefix, struct reposit\n \t\t.sparse = -1,\n \t\t.revs = REV_INFO_INIT,\n \t\t.include_edges = 1,\n+\t\t.object_info_results = FETCH_OBJECT_INFO_RESULTS_INIT,\n+\t\t.object_info_enabled = 1,\n \t};\n \tstruct option options[] = {\n \t\tOPT_UNSIGNED(0, \"min-batch-size\", &ctx.min_batch_size,\n@@ -185,7 +244,8 @@ int cmd_backfill(int argc, const char **argv, const char *prefix, struct reposit\n \t\t\t N_(\"Restrict the missing objects to the current sparse-checkout\")),\n \t\tOPT_BOOL(0, \"include-edges\", &ctx.include_edges,\n \t\t\t N_(\"Include blobs from boundary commits in the backfill\")),\n-\t\tOPT__DRY_RUN(&ctx.dry_run, N_(\"Preview the number of blobs to be fetched\")),\n+\t\tOPT__DRY_RUN(&ctx.dry_run, N_(\"Preview the number of blobs and their total \"\n+\t\t\t\t\t      \"size to be fetched\")),\n \t\tOPT_END(),\n \t};\n \tstruct repo_config_values *cfg = repo_config_values(the_repository);\ndiff --git a/t/t5620-backfill.sh b/t/t5620-backfill.sh\nindex e76fa6081b..b21feb625c 100755\n--- a/t/t5620-backfill.sh\n+++ b/t/t5620-backfill.sh\n@@ -167,6 +167,28 @@ test_expect_success '--dry-run with no missing blobs' '\n \ttest_grep \"0 blobs would be fetched\" out\n '\n \n+test_expect_success '--dry-run reports total size with object-info' '\n+\ttest_config -C srv.bare transfer.advertiseobjectinfo true &&\n+\ttest_when_finished rm -rf backfill-dry-run &&\n+\tgit clone --no-checkout --filter=blob:none \\\n+\t\t--single-branch --branch=main \\\n+\t\t\"file://$(pwd)/srv.bare\" backfill-dry-run &&\n+\n+\tgit -C backfill-dry-run backfill --dry-run >out &&\n+\ttest_grep \"48 blobs would be fetched (.*)\" out\n+'\n+\n+test_expect_success '--dry-run reports only the count without object-info' '\n+\ttest_config -C srv.bare transfer.advertiseobjectinfo false &&\n+\ttest_when_finished rm -rf backfill-dry-run &&\n+\tgit clone --no-checkout --filter=blob:none \\\n+\t\t--single-branch --branch=main \\\n+\t\t\"file://$(pwd)/srv.bare\" backfill-dry-run &&\n+\n+\tgit -C backfill-dry-run backfill --dry-run >out &&\n+\ttest_grep \"48 blobs would be fetched\\.$\" out\n+'\n+\n test_expect_success 'backfill --sparse without sparse-checkout fails' '\n \tgit init not-sparse &&\n \ttest_must_fail git -C not-sparse backfill --sparse 2>err &&\n\n-- \n2.54.0\n\n"},{"id":"553680","messageId":"CAOLa=ZQ8hvAxpU6SQ-KvN4eK5xMtJh4PoLXfDiNPmG89GQkMYQ@mail.gmail.com","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-2-1128f247ee01@gmail.com","subject":"Re: [PATCH RFC 2/5] fetch-object-info: add enum for fetch_object_info() statuses","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-09-30T10:56:40Z","receivedAt":"2026-09-30T10:56:47Z","isPatch":true,"body":"Pablo Sabater <pabloosabaterr@gmail.com> writes:\n\n> fetch_object_info() dies when the server does not advertise the\n> object-info capability. That is fine for git cat-file\n> remote-object-info command, which cannot work without it. However\n> a subsequent commit needs fetch_object_info() to not die, to be\n> able to fallback.\n>\n\nNit: The last sentence reads a little weird, perhaps:\n\n  However a subsequent commit uses fetch_object_info() optionally and\n  requires it to not die.\n\nOr something?\n\nBut I think we should just squash this into the next commit. It doesn't\nreally need to standout on its own.\n\n> Add \"enum fetch_object_info_status\" so that fetch_object_info() can\n> report this case to its callers. It is used in a subsequent commit.\n>\n> Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n> ---\n>  fetch-object-info.h | 6 ++++++\n>  1 file changed, 6 insertions(+)\n>\n> diff --git a/fetch-object-info.h b/fetch-object-info.h\n> index 2fba96c6f7..663a7f3ae7 100644\n> --- a/fetch-object-info.h\n> +++ b/fetch-object-info.h\n> @@ -16,6 +16,12 @@ struct fetch_object_info_results {\n>\n>  #define FETCH_OBJECT_INFO_RESULTS_INIT { 0 }\n>\n> +enum fetch_object_info_status {\n> +\tFETCH_OBJECT_INFO_OK = 0,\n> +\tFETCH_OBJECT_INFO_ERR = -1,\n> +\tFETCH_OBJECT_INFO_NOT_ENABLED = -2,\n> +};\n> +\n>  struct oid_array;\n>  /*\n>   * Sends git-cat-file object-info command into the request buf and reads the\n>\n> --\n> 2.54.0\n"},{"id":"553681","messageId":"CAOLa=ZR2Ka+5o8HxgZtnO98oH6vo6B_HmjVekYT8hC3HTMq-QQ@mail.gmail.com","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-4-1128f247ee01@gmail.com","subject":"Re: [PATCH RFC 4/5] backfill: add --dry-run option","fromName":"Karthik Nayak","fromEmail":"karthik.188@gmail.com","sentAt":"2026-09-30T11:06:10Z","receivedAt":"2026-09-30T11:06:24Z","isPatch":true,"body":"Pablo Sabater <pabloosabaterr@gmail.com> writes:\n\n> Users have no way to know how many blobs git backfill is going to\n> download before running it.\n>\n> Add a new --dry-run option to the backfill command. The objects are\n> walked as usual, but instead of fetching each batch of missing blobs\n> they are only counted, and the total is printed at the end.\n>\n> A subsequent commit will also print their size when the server supports\n> the object-info capability.\n>\n> Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n> ---\n>  Documentation/git-backfill.adoc |  6 +++++-\n>  builtin/backfill.c              | 37 +++++++++++++++++++++++++++++++++----\n>  t/t5620-backfill.sh             | 26 ++++++++++++++++++++++++++\n>  3 files changed, 64 insertions(+), 5 deletions(-)\n>\n> diff --git a/Documentation/git-backfill.adoc b/Documentation/git-backfill.adoc\n> index 82d6a1969d..08f19fea17 100644\n> --- a/Documentation/git-backfill.adoc\n> +++ b/Documentation/git-backfill.adoc\n> @@ -9,7 +9,7 @@ git-backfill - Download missing objects in a partial clone\n>  SYNOPSIS\n>  --------\n>  [synopsis]\n> -git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]\n> +git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]\n>\n>  DESCRIPTION\n>  -----------\n> @@ -70,6 +70,10 @@ OPTIONS\n>  \t--onto TARGET A..B`, where A..B normally excludes A but you need\n>  \tthe blobs from A as well.  `--include-edges` is the default.\n>\n> +`--dry-run`::\n> +\tDo not download any objects. Instead, print the number of\n> +\tmissing blobs that would be downloaded.\n> +\n>  `<revision-range>`::\n>  \tBackfill only blobs reachable from commits in the specified\n>  \trevision range.  When no _<revision-range>_ is specified, it\n> diff --git a/builtin/backfill.c b/builtin/backfill.c\n> index e71e0f4742..6019112966 100644\n> --- a/builtin/backfill.c\n> +++ b/builtin/backfill.c\n> @@ -26,7 +26,7 @@\n>  #include \"path-walk.h\"\n>\n>  static const char * const builtin_backfill_usage[] = {\n> -\tN_(\"git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]\"),\n> +\tN_(\"git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]\"),\n>  \tNULL\n>  };\n>\n> @@ -36,6 +36,8 @@ struct backfill_context {\n>  \tsize_t min_batch_size;\n>  \tint sparse;\n>  \tint include_edges;\n> +\tint dry_run;\n\nNit: This could be a bool, since `OPT__DRY_RUN` uses `OPT_BOOL` internally.\n\n> +\tsize_t total_batch_nr;\n\n\n>  \tstruct rev_info revs;\n>  };\n>\n> @@ -58,6 +60,15 @@ static void download_batch(struct backfill_context *ctx)\n>  \todb_reprepare(ctx->repo->objects);\n>  }\n>\n> +static void dry_run_batch(struct backfill_context *ctx)\n> +{\n\nWhile it is used during dry_run, probably makes more sense to rename it\nto `count_batch()` since that's what it does.\n\n> +\tif (!ctx->current_batch.nr)\n> +\t\treturn;\n> +\n> +\tctx->total_batch_nr += ctx->current_batch.nr;\n> +\toid_array_clear(&ctx->current_batch);\n> +}\n> +\n>  static int fill_missing_blobs(const char *path UNUSED,\n>  \t\t\t      struct oid_array *list,\n>  \t\t\t      enum object_type type,\n> @@ -73,8 +84,12 @@ static int fill_missing_blobs(const char *path UNUSED,\n>  \t\t\toid_array_append(&ctx->current_batch, &list->oid[i]);\n>  \t}\n>\n> -\tif (ctx->current_batch.nr >= ctx->min_batch_size)\n> -\t\tdownload_batch(ctx);\n> +\tif (ctx->current_batch.nr >= ctx->min_batch_size) {\n> +\t\tif (ctx->dry_run)\n> +\t\t\tdry_run_batch(ctx);\n> +\t\telse\n> +\t\t\tdownload_batch(ctx);\n> +\t}\n>\n>  \treturn 0;\n>  }\n> @@ -131,10 +146,23 @@ static int do_backfill(struct backfill_context *ctx)\n>\n>  \tret = walk_objects_by_path(&info);\n>\n> +\tif (ret)\n> +\t\tgoto end;\n> +\n>  \t/* Download the objects that did not fill a batch. */\n> -\tif (!ret)\n> +\tif (!ctx->dry_run) {\n>  \t\tdownload_batch(ctx);\n> +\t\tgoto end;\n> +\t}\n> +\n> +\tdry_run_batch(ctx);\n> +\n> +\tprintf(Q_(\"After backfill, %\" PRIuMAX \" blob would be fetched.\\n\",\n> +\t\t  \"After backfill, %\" PRIuMAX \" blobs would be fetched.\\n\",\n> +\t\t  (unsigned long)ctx->total_batch_nr),\n> +\t       (uintmax_t)ctx->total_batch_nr);\n>\n\nNit: Okay so we have a goto inside the first if(...), which skips this\nsection. I would have found it easier to read if it was\n\nif (dry_run)\n   count()\nelse\n   download()\n\n> +end:\n>  \tpath_walk_info_clear(&info);\n>  \treturn ret;\n>  }\n> @@ -157,6 +185,7 @@ int cmd_backfill(int argc, const char **argv, const char *prefix, struct reposit\n>  \t\t\t N_(\"Restrict the missing objects to the current sparse-checkout\")),\n>  \t\tOPT_BOOL(0, \"include-edges\", &ctx.include_edges,\n>  \t\t\t N_(\"Include blobs from boundary commits in the backfill\")),\n> +\t\tOPT__DRY_RUN(&ctx.dry_run, N_(\"Preview the number of blobs to be fetched\")),\n>  \t\tOPT_END(),\n>  \t};\n>  \tstruct repo_config_values *cfg = repo_config_values(the_repository);\n> diff --git a/t/t5620-backfill.sh b/t/t5620-backfill.sh\n> index 7462280470..e76fa6081b 100755\n> --- a/t/t5620-backfill.sh\n> +++ b/t/t5620-backfill.sh\n> @@ -141,6 +141,32 @@ test_expect_success 'do partial clone 2, backfill min batch size' '\n>  \ttest_line_count = 0 revs2\n>  '\n>\n> +test_expect_success '--dry-run reports missing blobs without fetching them' '\n> +\ttest_when_finished \"rm -rf backfill-dry-run dry-trace\" &&\n> +\tgit clone --no-checkout --filter=blob:none \\\n> +\t\t--single-branch --branch=main \\\n> +\t\t\"file://$(pwd)/srv.bare\" backfill-dry-run &&\n> +\n> +\tGIT_TRACE2_EVENT=\"$(pwd)/dry-trace\" git \\\n> +\t\t-C backfill-dry-run backfill --dry-run >out &&\n> +\n> +\ttest_grep \"48 blobs would be fetched\" out &&\n> +\ttest_grep ! fetch_count dry-trace &&\n> +\tgit -C backfill-dry-run rev-list --quiet --objects --missing=print HEAD >missing &&\n> +\ttest_line_count = 48 missing\n> +'\n> +\n> +test_expect_success '--dry-run with no missing blobs' '\n> +\ttest_when_finished rm -rf backfill-dry-run &&\n> +\tgit clone --no-checkout --filter=blob:none \\\n> +\t\t--single-branch --branch=main \\\n> +\t\t\"file://$(pwd)/srv.bare\" backfill-dry-run &&\n> +\tgit -C backfill-dry-run backfill &&\n> +\n> +\tgit -C backfill-dry-run backfill --dry-run >out &&\n> +\ttest_grep \"0 blobs would be fetched\" out\n> +'\n> +\n>  test_expect_success 'backfill --sparse without sparse-checkout fails' '\n>  \tgit init not-sparse &&\n>  \ttest_must_fail git -C not-sparse backfill --sparse 2>err &&\n>\n> --\n> 2.54.0\n"},{"id":"553684","messageId":"DLSMTVIXW120.3F87OEHJW1LEC@gmail.com","threadId":"66423","inReplyTo":"CAOLa=ZQ8hvAxpU6SQ-KvN4eK5xMtJh4PoLXfDiNPmG89GQkMYQ@mail.gmail.com","subject":"Re: [PATCH RFC 2/5] fetch-object-info: add enum for fetch_object_info() statuses","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T11:59:21Z","receivedAt":"2026-09-30T11:59:24Z","isPatch":true,"body":"On Wed Sep 30, 2026 at 11:56 AM WEST, Karthik Nayak wrote:\n> Pablo Sabater <pabloosabaterr@gmail.com> writes:\n>\n>> fetch_object_info() dies when the server does not advertise the\n>> object-info capability. That is fine for git cat-file\n>> remote-object-info command, which cannot work without it. However\n>> a subsequent commit needs fetch_object_info() to not die, to be\n>> able to fallback.\n>>\n>\n> Nit: The last sentence reads a little weird, perhaps:\n>\n>   However a subsequent commit uses fetch_object_info() optionally and\n>   requires it to not die.\n>\n> Or something?\n>\n> But I think we should just squash this into the next commit. It doesn't\n> really need to standout on its own.\n\nOkay, squashing them makes sense.  I only split them to avoid having the\nenum buried among the signature changes.\n\nWill fix, thanks.\n\n>\n>> Add \"enum fetch_object_info_status\" so that fetch_object_info() can\n>> report this case to its callers. It is used in a subsequent commit.\n>>\n>> Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n>> ---\n>>  fetch-object-info.h | 6 ++++++\n>>  1 file changed, 6 insertions(+)\n>>\n>> diff --git a/fetch-object-info.h b/fetch-object-info.h\n>> index 2fba96c6f7..663a7f3ae7 100644\n>> --- a/fetch-object-info.h\n>> +++ b/fetch-object-info.h\n>> @@ -16,6 +16,12 @@ struct fetch_object_info_results {\n>>\n>>  #define FETCH_OBJECT_INFO_RESULTS_INIT { 0 }\n>>\n>> +enum fetch_object_info_status {\n>> +\tFETCH_OBJECT_INFO_OK = 0,\n>> +\tFETCH_OBJECT_INFO_ERR = -1,\n>> +\tFETCH_OBJECT_INFO_NOT_ENABLED = -2,\n>> +};\n>> +\n>>  struct oid_array;\n>>  /*\n>>   * Sends git-cat-file object-info command into the request buf and reads the\n>>\n>> --\n>> 2.54.0\n\n"},{"id":"553688","messageId":"DLSN94VUFI80.15JU19PJ3RPK1@gmail.com","threadId":"66423","inReplyTo":"CAOLa=ZR2Ka+5o8HxgZtnO98oH6vo6B_HmjVekYT8hC3HTMq-QQ@mail.gmail.com","subject":"Re: [PATCH RFC 4/5] backfill: add --dry-run option","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T12:19:17Z","receivedAt":"2026-09-30T12:19:20Z","isPatch":true,"body":"On Wed Sep 30, 2026 at 12:06 PM WEST, Karthik Nayak wrote:\n> Pablo Sabater <pabloosabaterr@gmail.com> writes:\n>\n>> Users have no way to know how many blobs git backfill is going to\n>> download before running it.\n>>\n>> Add a new --dry-run option to the backfill command. The objects are\n>> walked as usual, but instead of fetching each batch of missing blobs\n>> they are only counted, and the total is printed at the end.\n>>\n>> A subsequent commit will also print their size when the server supports\n>> the object-info capability.\n>>\n>> Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>\n>> ---\n>>  Documentation/git-backfill.adoc |  6 +++++-\n>>  builtin/backfill.c              | 37 +++++++++++++++++++++++++++++++++----\n>>  t/t5620-backfill.sh             | 26 ++++++++++++++++++++++++++\n>>  3 files changed, 64 insertions(+), 5 deletions(-)\n>>\n>> diff --git a/Documentation/git-backfill.adoc b/Documentation/git-backfill.adoc\n>> index 82d6a1969d..08f19fea17 100644\n>> --- a/Documentation/git-backfill.adoc\n>> +++ b/Documentation/git-backfill.adoc\n>> @@ -9,7 +9,7 @@ git-backfill - Download missing objects in a partial clone\n>>  SYNOPSIS\n>>  --------\n>>  [synopsis]\n>> -git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]\n>> +git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]\n>>\n>>  DESCRIPTION\n>>  -----------\n>> @@ -70,6 +70,10 @@ OPTIONS\n>>  \t--onto TARGET A..B`, where A..B normally excludes A but you need\n>>  \tthe blobs from A as well.  `--include-edges` is the default.\n>>\n>> +`--dry-run`::\n>> +\tDo not download any objects. Instead, print the number of\n>> +\tmissing blobs that would be downloaded.\n>> +\n>>  `<revision-range>`::\n>>  \tBackfill only blobs reachable from commits in the specified\n>>  \trevision range.  When no _<revision-range>_ is specified, it\n>> diff --git a/builtin/backfill.c b/builtin/backfill.c\n>> index e71e0f4742..6019112966 100644\n>> --- a/builtin/backfill.c\n>> +++ b/builtin/backfill.c\n>> @@ -26,7 +26,7 @@\n>>  #include \"path-walk.h\"\n>>\n>>  static const char * const builtin_backfill_usage[] = {\n>> -\tN_(\"git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]\"),\n>> +\tN_(\"git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]\"),\n>>  \tNULL\n>>  };\n>>\n>> @@ -36,6 +36,8 @@ struct backfill_context {\n>>  \tsize_t min_batch_size;\n>>  \tint sparse;\n>>  \tint include_edges;\n>> +\tint dry_run;\n>\n> Nit: This could be a bool, since `OPT__DRY_RUN` uses `OPT_BOOL` internally.\n\nWill do, didn't know that using bool was a thing ;).  Is it also prefered\nfor 0/1 functions?\n\n>\n>> +\tsize_t total_batch_nr;\n>\n>\n>>  \tstruct rev_info revs;\n>>  };\n>>\n>> @@ -58,6 +60,15 @@ static void download_batch(struct backfill_context *ctx)\n>>  \todb_reprepare(ctx->repo->objects);\n>>  }\n>>\n>> +static void dry_run_batch(struct backfill_context *ctx)\n>> +{\n>\n> While it is used during dry_run, probably makes more sense to rename it\n> to `count_batch()` since that's what it does.\n\ncount_batch() works for me, I think that \"count\" doesn't fit too well\nfor summing the size of the objects, might opt for another name if I\nthink of a better name.\n\nProb a comment helps.\n\n>\n>> +\tif (!ctx->current_batch.nr)\n>> +\t\treturn;\n>> +\n>> +\tctx->total_batch_nr += ctx->current_batch.nr;\n>> +\toid_array_clear(&ctx->current_batch);\n>> +}\n>> +\n>>  static int fill_missing_blobs(const char *path UNUSED,\n>>  \t\t\t      struct oid_array *list,\n>>  \t\t\t      enum object_type type,\n>> @@ -73,8 +84,12 @@ static int fill_missing_blobs(const char *path UNUSED,\n>>  \t\t\toid_array_append(&ctx->current_batch, &list->oid[i]);\n>>  \t}\n>>\n>> -\tif (ctx->current_batch.nr >= ctx->min_batch_size)\n>> -\t\tdownload_batch(ctx);\n>> +\tif (ctx->current_batch.nr >= ctx->min_batch_size) {\n>> +\t\tif (ctx->dry_run)\n>> +\t\t\tdry_run_batch(ctx);\n>> +\t\telse\n>> +\t\t\tdownload_batch(ctx);\n>> +\t}\n>>\n>>  \treturn 0;\n>>  }\n>> @@ -131,10 +146,23 @@ static int do_backfill(struct backfill_context *ctx)\n>>\n>>  \tret = walk_objects_by_path(&info);\n>>\n>> +\tif (ret)\n>> +\t\tgoto end;\n>> +\n>>  \t/* Download the objects that did not fill a batch. */\n>> -\tif (!ret)\n>> +\tif (!ctx->dry_run) {\n>>  \t\tdownload_batch(ctx);\n>> +\t\tgoto end;\n>> +\t}\n>> +\n>> +\tdry_run_batch(ctx);\n>> +\n>> +\tprintf(Q_(\"After backfill, %\" PRIuMAX \" blob would be fetched.\\n\",\n>> +\t\t  \"After backfill, %\" PRIuMAX \" blobs would be fetched.\\n\",\n>> +\t\t  (unsigned long)ctx->total_batch_nr),\n>> +\t       (uintmax_t)ctx->total_batch_nr);\n>>\n>\n> Nit: Okay so we have a goto inside the first if(...), which skips this\n> section. I would have found it easier to read if it was\n>\n> if (dry_run)\n>    count()\n> else\n>    download()\n\nWill change it, thanks.\n\n>\n\n[snip]\n\n"},{"id":"553729","messageId":"xmqqwls2brca.fsf@gitster.g","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-3-1128f247ee01@gmail.com","subject":"Re: [PATCH RFC 3/5] fetch-object-info: return a status instead of dying","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-09-30T17:07:33Z","receivedAt":"2026-09-30T17:07:35Z","isPatch":true,"body":"Pablo Sabater <pabloosabaterr@gmail.com> writes:\n\n> A subsequent commit needs fetch_object_info() not to die() when the\n> object-info capability is not enabled on the server, so that it can\n> fall back.\n>\n> Make fetch_object_info() return FETCH_OBJECT_INFO_NOT_ENABLED instead\n> of die()'ing when the server does not advertise the object-info\n> capability, and propagate the status through the transport layer so\n> that callers of transport_fetch_object_info() can act on it. It is now\n> up to them whether to die() or fall back.\n\nIt may be just me but unless the client can tell between the server\nnot supporting (i.e., they are unable to enable it even if they\nwanted to) and not enabling (i.e., they are capable, but are not\nwilling to give it to you), it may make sense to report it as \"not\navailable\".  \"not enabled\" sounds as if we know that it is the\nlatter and not the former.\n\nThe code change looks very cleanly done.\n"},{"id":"553730","messageId":"xmqqqziabqw8.fsf@gitster.g","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-5-1128f247ee01@gmail.com","subject":"Re: [PATCH RFC 5/5] backfill: report total size of missing blobs in --dry-run","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-09-30T17:17:11Z","receivedAt":"2026-09-30T17:17:13Z","isPatch":true,"body":"Pablo Sabater <pabloosabaterr@gmail.com> writes:\n\n> +\n> +\tif (!ctx->object_info_enabled)\n> +\t\tgoto cleanup;\n> +\n> +\tif (!ctx->object_info_transport) {\n> +\t\tstruct promisor_remote *promise =\n> +\t\t\trepo_promisor_remote_find(ctx->repo, NULL);\n> +\t\tstruct remote *remote = NULL;\n> +\n> +\t\tif (!promise || !(remote = remote_get(promise->name)))\n> +\t\t\tdie(_(\"--dry-run requires a promisor remote\"));\n> +\n> +\t\tctx->object_info_transport = transport_get(remote, NULL);\n> +\n> +\t\tif (!ctx->object_info_transport->smart_options)\n> +\t\t\tdie(_(\"failed to get object info: smart options required\"));\n> +\t}\n> +\n> +\tresults->wants_size = 1;\n> +\tstatus = transport_fetch_object_info(ctx->object_info_transport,\n> +\t\t\t\t\t     &ctx->current_batch,\n> +\t\t\t\t\t     results);\n> +\n> +\tif (status == FETCH_OBJECT_INFO_NOT_ENABLED ||\n> +\t    !results->sizes) {\n> +\t\tctx->object_info_enabled = 0;\n\nYuck.\n\nBecause we cannot tell if they allow you to look at the information,\nthis cannot be helped, but it means anybody that looks at this\nctx->object_info_enabled member to decide what to do must be careful.\n\nIs it guaranteed that results.sizes[] have been populated as long as\nstatus is not FETCH_OBJECT_INFO_NOT_ENABLED?  Can there be other\nerrors that makes result.sizes[] unusable?  If that is the case,\nthen it would be cleaner to have a dedicated ctx->sizes_valid member\nrather than relying on ctx->object_info_enabled member to carry this\ninformation ...\n\n> +\t\tgoto cleanup;\n> +\t}\n> +\n> +\tfor (size_t i = 0; i < results->nr; i++)\n> +\t\tctx->total_batch_size += results->sizes[i];\n> +\n> +cleanup:\n> +\tfree_fetch_object_info_results(&ctx->object_info_results);\n>  \toid_array_clear(&ctx->current_batch);\n>  }\n>  \n> @@ -157,12 +201,25 @@ static int do_backfill(struct backfill_context *ctx)\n>  \n>  \tdry_run_batch(ctx);\n>  \n> -\tprintf(Q_(\"After backfill, %\" PRIuMAX \" blob would be fetched.\\n\",\n> -\t\t  \"After backfill, %\" PRIuMAX \" blobs would be fetched.\\n\",\n> -\t\t  (unsigned long)ctx->total_batch_nr),\n> -\t       (uintmax_t)ctx->total_batch_nr);\n> +\tif (ctx->object_info_enabled && ctx->total_batch_nr) {\n\n... and use it here.  Within the design presented in this series, we\nknow we have asked the other end at this point, and the above\nfunction may have turned ctx->object_info member off if the\ninformation is not there, so this may be safe.  But as I said, I am\nnot sure what happens when fetch-object-info returned other kind of\nerrors.\n\n"},{"id":"553734","messageId":"DLSUKMRZHHGO.HVX98MB8KLSF@gmail.com","threadId":"66423","inReplyTo":"xmqqwls2brca.fsf@gitster.g","subject":"Re: [PATCH RFC 3/5] fetch-object-info: return a status instead of dying","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T18:03:26Z","receivedAt":"2026-09-30T18:03:29Z","isPatch":true,"body":"On Wed Sep 30, 2026 at 6:07 PM WEST, Junio C Hamano wrote:\n> Pablo Sabater <pabloosabaterr@gmail.com> writes:\n>\n>> A subsequent commit needs fetch_object_info() not to die() when the\n>> object-info capability is not enabled on the server, so that it can\n>> fall back.\n>>\n>> Make fetch_object_info() return FETCH_OBJECT_INFO_NOT_ENABLED instead\n>> of die()'ing when the server does not advertise the object-info\n>> capability, and propagate the status through the transport layer so\n>> that callers of transport_fetch_object_info() can act on it. It is now\n>> up to them whether to die() or fall back.\n>\n> It may be just me but unless the client can tell between the server\n> not supporting (i.e., they are unable to enable it even if they\n> wanted to) and not enabling (i.e., they are capable, but are not\n> willing to give it to you), it may make sense to report it as \"not\n> available\".  \"not enabled\" sounds as if we know that it is the\n> latter and not the former.\n>\n> The code change looks very cleanly done.\n\nMakes sense, I'll rename it to FETCH_OBJECT_INFO_NOT_AVAILABLE.\nThe git cat-file remote-object-info command path die()'d with this\nmessage:\n\n\tdie(_(\"object-info capability is not enabled on the server\"));\n\nI'll update the die() message as well.\n\nThanks.\n\n"},{"id":"553736","messageId":"ed1b9048-d438-4143-a224-fa0e28d4fd42@gmail.com","threadId":"66423","inReplyTo":"20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com","subject":"Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-09-30T18:14:46Z","receivedAt":"2026-09-30T18:14:49Z","isPatch":true,"body":"On 9/29/2026 8:21 PM, Pablo Sabater wrote:\n> [Cc'd Derrick Stolee for his work in the backfill(1) command]\n> \n> This series adds a --dry-run option to git-backfill(1) that reports how\n> many missing blobs would be fetched and, when the remote server\n> supports the object-info capability, their total size:\n> \n>         $ git backfill --dry-run\n>         After backfill, 48 blobs would be fetched (1.20 KiB).\n> \n> If the server does not advertise object-info, only the count is shown.\n\nThis is a helpful capability, but I'm not sure the size check counts\nas a \"dry run\" because it involves a network call (and possibly many\ndepending on --min-batch-size).\n\nPerhaps a different argument would be better, such as --info=(count|size)\nto make it clear what level of information you want to know in advance\nand thus how much effort are you willing to put in to discover this. \n> I am not a git-backfill(1) user myself, but it seemed useful for users\n> to know how much data a backfill would bring in before running it.\n\nI'm not sure that we want to add a feature based on speculation. Git\nis a collection of \"itches\" that the contributors needed scratched.\nThe work is motivated by real needs.\n\nWhile I can see some benefit to curiosity, I'm not sure how much this\nwould prevent users from making their decision as to whether they\nshould run backfill or not.\n\n> The number of missing blobs is the sum of the number of blobs to be\n> fetched in each batch. The object-info capability lets us ask the server\n> for the size of each blob without downloading it, so summing them gives\n> an estimate of the total.\n> \n> Note that this is an upper bound rather than the exact disk usage:\n> object-info reports the uncompressed size of each object, while the\n> objects end up stored compressed and possibly deltified in a packfile,\n> so the space actually used on disk will usually be smaller.\n\nI don't think the uncompressed size is a useful metric here, as it is\nlikely astronomically larger than what will be downloaded. How will\nthis help a user make a decision?\n\nThanks,\n-Stolee\n\n"},{"id":"553738","messageId":"DLSV653N0668.1OZOWE9P5ZJV6@gmail.com","threadId":"66423","inReplyTo":"xmqqqziabqw8.fsf@gitster.g","subject":"Re: [PATCH RFC 5/5] backfill: report total size of missing blobs in --dry-run","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T18:31:31Z","receivedAt":"2026-09-30T18:31:35Z","isPatch":true,"body":"On Wed Sep 30, 2026 at 6:17 PM WEST, Junio C Hamano wrote:\n> Pablo Sabater <pabloosabaterr@gmail.com> writes:\n>\n>> +\n>> +\tif (!ctx->object_info_enabled)\n>> +\t\tgoto cleanup;\n>> +\n>> +\tif (!ctx->object_info_transport) {\n>> +\t\tstruct promisor_remote *promise =\n>> +\t\t\trepo_promisor_remote_find(ctx->repo, NULL);\n>> +\t\tstruct remote *remote = NULL;\n>> +\n>> +\t\tif (!promise || !(remote = remote_get(promise->name)))\n>> +\t\t\tdie(_(\"--dry-run requires a promisor remote\"));\n>> +\n>> +\t\tctx->object_info_transport = transport_get(remote, NULL);\n>> +\n>> +\t\tif (!ctx->object_info_transport->smart_options)\n>> +\t\t\tdie(_(\"failed to get object info: smart options required\"));\n>> +\t}\n>> +\n>> +\tresults->wants_size = 1;\n>> +\tstatus = transport_fetch_object_info(ctx->object_info_transport,\n>> +\t\t\t\t\t     &ctx->current_batch,\n>> +\t\t\t\t\t     results);\n>> +\n>> +\tif (status == FETCH_OBJECT_INFO_NOT_ENABLED ||\n>> +\t    !results->sizes) {\n>> +\t\tctx->object_info_enabled = 0;\n>\n> Yuck.\n>\n> Because we cannot tell if they allow you to look at the information,\n> this cannot be helped, but it means anybody that looks at this\n> ctx->object_info_enabled member to decide what to do must be careful.\n>\n> Is it guaranteed that results.sizes[] have been populated as long as\n> status is not FETCH_OBJECT_INFO_NOT_ENABLED?  Can there be other\n> errors that makes result.sizes[] unusable?  If that is the case,\n> then it would be cleaner to have a dedicated ctx->sizes_valid member\n> rather than relying on ctx->object_info_enabled member to carry this\n> information ...\n\nNot quite, transport_fetch_object_info() can also return\nFETCH_OBJECT_INFO_ERR when finish_connect() fails in\nfetch_object_info_via_pack(). I'll change the condition to:\n\n\tif (status != FETCH_OBJECT_INFO_OK || !results->sizes) {\n\nWith FETCH_OBJECT_INFO_OK, fetch_object_info() only allocates\nresults->sizes when the server answers with the \"size\" attribute,\nand it die()s on any incomplete or malformed response, so a\nnon-NULL sizes[] is always fully populated. A NULL sizes[] with\nFETCH_OBJECT_INFO_OK means object-info is available but the server\ndoes not support \"size\".\n\nAgreed, the member is also cleared in that last case, where\nobject-info does work but \"size\" is not advertised, so I get that it can\nbe misleading. Since its only use is deciding whether to report the \ntotal size, I'll rename it to ctx->sizes_valid instead of adding a \nseparate member.\n\n>\n>> +\t\tgoto cleanup;\n>> +\t}\n>> +\n>> +\tfor (size_t i = 0; i < results->nr; i++)\n>> +\t\tctx->total_batch_size += results->sizes[i];\n>> +\n>> +cleanup:\n>> +\tfree_fetch_object_info_results(&ctx->object_info_results);\n>>  \toid_array_clear(&ctx->current_batch);\n>>  }\n>>  \n>> @@ -157,12 +201,25 @@ static int do_backfill(struct backfill_context *ctx)\n>>  \n>>  \tdry_run_batch(ctx);\n>>  \n>> -\tprintf(Q_(\"After backfill, %\" PRIuMAX \" blob would be fetched.\\n\",\n>> -\t\t  \"After backfill, %\" PRIuMAX \" blobs would be fetched.\\n\",\n>> -\t\t  (unsigned long)ctx->total_batch_nr),\n>> -\t       (uintmax_t)ctx->total_batch_nr);\n>> +\tif (ctx->object_info_enabled && ctx->total_batch_nr) {\n>\n> ... and use it here.  Within the design presented in this series, we\n> know we have asked the other end at this point, and the above\n> function may have turned ctx->object_info member off if the\n> information is not there, so this may be safe.  But as I said, I am\n> not sure what happens when fetch-object-info returned other kind of\n> errors.\n\nThe condition above is checked on every batch, and any status other\nthan FETCH_OBJECT_INFO_OK or a NULL sizes[] clears ctx->sizes_valid\n(currently ctx->object_info_enabled), so by the time we get here it\nis only set when all the batches returned valid sizes.\n\n"},{"id":"553744","messageId":"DLSVWT1QTELK.17U19NG169AW9@gmail.com","threadId":"66423","inReplyTo":"ed1b9048-d438-4143-a224-fa0e28d4fd42@gmail.com","subject":"Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)","fromName":"Pablo Sabater","fromEmail":"pabloosabaterr@gmail.com","sentAt":"2026-09-30T19:06:21Z","receivedAt":"2026-09-30T19:06:25Z","isPatch":true,"body":"On Wed Sep 30, 2026 at 7:14 PM WEST, Derrick Stolee wrote:\n> On 9/29/2026 8:21 PM, Pablo Sabater wrote:\n>> [Cc'd Derrick Stolee for his work in the backfill(1) command]\n>> \n>> This series adds a --dry-run option to git-backfill(1) that reports how\n>> many missing blobs would be fetched and, when the remote server\n>> supports the object-info capability, their total size:\n>> \n>>         $ git backfill --dry-run\n>>         After backfill, 48 blobs would be fetched (1.20 KiB).\n>> \n>> If the server does not advertise object-info, only the count is shown.\n>\n> This is a helpful capability, but I'm not sure the size check counts\n> as a \"dry run\" because it involves a network call (and possibly many\n> depending on --min-batch-size).\n>\n> Perhaps a different argument would be better, such as --info=(count|size)\n> to make it clear what level of information you want to know in advance\n> and thus how much effort are you willing to put in to discover this. \n\nMakes sense to have it as an --info option.\n\n>> I am not a git-backfill(1) user myself, but it seemed useful for users\n>> to know how much data a backfill would bring in before running it.\n>\n> I'm not sure that we want to add a feature based on speculation. Git\n> is a collection of \"itches\" that the contributors needed scratched.\n> The work is motivated by real needs.\n>\n> While I can see some benefit to curiosity, I'm not sure how much this\n> would prevent users from making their decision as to whether they\n> should run backfill or not.\n>\n>> The number of missing blobs is the sum of the number of blobs to be\n>> fetched in each batch. The object-info capability lets us ask the server\n>> for the size of each blob without downloading it, so summing them gives\n>> an estimate of the total.\n>> \n>> Note that this is an upper bound rather than the exact disk usage:\n>> object-info reports the uncompressed size of each object, while the\n>> objects end up stored compressed and possibly deltified in a packfile,\n>> so the space actually used on disk will usually be smaller.\n>\n> I don't think the uncompressed size is a useful metric here, as it is\n> likely astronomically larger than what will be downloaded. How will\n> this help a user make a decision?\n\nYes, that's one of the itches I have with the object-info protocol: it\ncannot give you a reliable compressed size, and that's why only the\ntotal size is supported.\n\nThe object-info protocol could be extended to support the\nobjectsize:disk attribute, either by having the server know what we\nalready have and the oids that we want, or by directly having the\nserver send us the compressed size of its local copy (to avoid too much\nwork).  Even with the first option, because it goes in batches, it\nwould still be an estimate, just a closer one.\n\nGiven that I don't use backfill, I Cc'd you because I wasn't sure if it\nwas really useful, and the main motivation was the \"I'm going to check\n--dry-run before backfilling\" case, so it helps to decide.  If it's not\nthat helpful and seems to end up as a decoration option, it might be\nbetter to drop it.\n\n>\n> Thanks,\n> -Stolee\n\nThanks for taking a look,\nPablo\n\n"},{"id":"553750","messageId":"xmqq1paaa4pf.fsf@gitster.g","threadId":"66423","inReplyTo":"DLSUKMRZHHGO.HVX98MB8KLSF@gmail.com","subject":"Re: [PATCH RFC 3/5] fetch-object-info: return a status instead of dying","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-09-30T20:01:48Z","receivedAt":"2026-09-30T20:01:50Z","isPatch":true,"body":"\"Pablo Sabater\" <pabloosabaterr@gmail.com> writes:\n\n> On Wed Sep 30, 2026 at 6:07 PM WEST, Junio C Hamano wrote:\n>> Pablo Sabater <pabloosabaterr@gmail.com> writes:\n>>\n>>> A subsequent commit needs fetch_object_info() not to die() when the\n>>> object-info capability is not enabled on the server, so that it can\n>>> fall back.\n>>>\n>>> Make fetch_object_info() return FETCH_OBJECT_INFO_NOT_ENABLED instead\n>>> of die()'ing when the server does not advertise the object-info\n>>> capability, and propagate the status through the transport layer so\n>>> that callers of transport_fetch_object_info() can act on it. It is now\n>>> up to them whether to die() or fall back.\n>>\n>> It may be just me but unless the client can tell between the server\n>> not supporting (i.e., they are unable to enable it even if they\n>> wanted to) and not enabling (i.e., they are capable, but are not\n>> willing to give it to you), it may make sense to report it as \"not\n>> available\".  \"not enabled\" sounds as if we know that it is the\n>> latter and not the former.\n>>\n>> The code change looks very cleanly done.\n>\n> Makes sense, I'll rename it to FETCH_OBJECT_INFO_NOT_AVAILABLE.\n\nMake it UNAVAILABLE instead.\n"},{"id":"553751","messageId":"xmqqwls28q1t.fsf@gitster.g","threadId":"66423","inReplyTo":"ed1b9048-d438-4143-a224-fa0e28d4fd42@gmail.com","subject":"Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-09-30T20:03:42Z","receivedAt":"2026-09-30T20:03:45Z","isPatch":true,"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> This is a helpful capability, but I'm not sure the size check counts\n> as a \"dry run\" because it involves a network call (and possibly many\n> depending on --min-batch-size).\n> ...\n> I'm not sure that we want to add a feature based on speculation. Git\n> is a collection of \"itches\" that the contributors needed scratched.\n> The work is motivated by real needs.\n> ...\n> I don't think the uncompressed size is a useful metric here, as it is\n> likely astronomically larger than what will be downloaded. How will\n> this help a user make a decision?\n\nI agree with you on all counts.\n"}]}