Volume XXII, number 279Tuesday, October 6, 2026Latest message 35 minutes ago

The Git List

News and archive of git@vger.kernel.org, since April 2005

RFC patch, 5 partsAdd --dry-run option to git-backfill(1)

18 messages between Sep 30, 2026 and Sep 30, 2026, from Pablo Sabater, Karthik Nayak, Junio C Hamano, Derrick Stolee.

Plain Markdown or JSON for tools and agents. Diffs are folded; open one to read it.

Pablo SabaterSep 30, 2026, 00:21 UTC on lore
[Cc'd Derrick Stolee for his work in the backfill(1) command]

This series adds a --dry-run option to git-backfill(1) that reports how many missing blobs would be fetched and, when the remote server supports the object-info capability, their total size:

        $ git backfill --dry-run
        After backfill, 48 blobs would be fetched (1.20 KiB).
If the server does not advertise object-info, only the count is shown.

I am not a git-backfill(1) user myself, but it seemed useful for users to know how much data a backfill would bring in before running it.

The number of missing blobs is the sum of the number of blobs to be fetched in each batch. The object-info capability lets us ask the server for the size of each blob without downloading it, so summing them gives an estimate of the total.

Note that this is an upper bound rather than the exact disk usage: object-info reports the uncompressed size of each object, while the objects end up stored compressed and possibly deltified in a packfile, so the space actually used on disk will usually be smaller.

Since I do not use backfill, feedback on whether this is useful, and on the output format, is very welcome.

Patches 1-3 are preparatory:
  [1/5] transport-internal: update fetch_object_info comment
        Fixes an outdated comment: object-info supports type as well as
        size.
  [2/5] fetch-object-info: add enum for fetch_object_info() statuses
  [3/5] fetch-object-info: return a status instead of dying
        Teach fetch_object_info() to return a status instead of dying
        when the server does not advertise object-info. The die() is
        kept in cat-file's remote-object-info path, so its behavior is
        unchanged.
Patches 4-5 add the option in two steps:
  [4/5] backfill: add --dry-run option
        Prints only the number of blobs that would be fetched.
  [5/5] backfill: report total size of missing blobs in --dry-run
        Also prints their total size when the server supports
        object-info, and falls back to the count alone otherwise.
Thanks.
Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
---
Pablo Sabater (5):
      transport-internal: update fetch_object_info comment
      fetch-object-info: add enum for fetch_object_info() statuses
      fetch-object-info: return a status instead of dying
      backfill: add --dry-run option
      backfill: report total size of missing blobs in --dry-run
 Documentation/git-backfill.adoc | 10 ++++-
 builtin/backfill.c              | 97 +++++++++++++++++++++++++++++++++++++++--
 builtin/cat-file.c              |  4 ++
 fetch-object-info.c             | 17 ++++----
 fetch-object-info.h             | 24 +++++++---
 t/t5620-backfill.sh             | 48 ++++++++++++++++++++
 transport-helper.c              |  6 +--
 transport-internal.h            | 12 ++---
 transport.c                     | 29 ++++++------
 transport.h                     |  7 +--
 10 files changed, 208 insertions(+), 46 deletions(-)

--- base-commit: 12cb6293d6288865c1a133cf22accbaf99d13eb6 change-id: 20260914-backfill-dryrun-fac997003322

Pablo SabaterSep 30, 2026, 00:21 UTC in reply to Pablo Sabater on lore

[PATCH RFC 1/5] transport-internal: update fetch_object_info comment

The comment describing the fetch_object_info() callback in struct transport_vtable says that only the object size can be fetched. The object-info capability can now also report the object type, or neither, to only check whether an object exists on the remote.

Update the comment.
Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
---
 transport-internal.h | 4 ++--
 1 file changed, 2 insertions(+), 2 deletions(-)
Show changes to transport-internal.h +2 −2
diff --git a/transport-internal.h b/transport-internal.h
index a10b27cc81..626ceaae2b 100644
--- a/transport-internal.h
+++ b/transport-internal.h
@@ -48,8 +48,8 @@ struct transport_vtable {
 	int (*fetch_refs)(struct transport *transport, int refs_nr, struct ref **refs);
 
 	/*
-	 * Fetch object info (only size currently) from remote without
-	 * downloading the objects.
+	 * Fetch object info (size, type, or none of them to only check
+	 * for existence) from the remote without downloading the objects.
 	 *
 	 * Uses object-info capability of v2 protocol.
 	 */
-- 
2.54.0
Pablo SabaterSep 30, 2026, 00:21 UTC in reply to Pablo Sabater on lore

[PATCH RFC 2/5] fetch-object-info: add enum for fetch_object_info() statuses

fetch_object_info() dies when the server does not advertise the object-info capability. That is fine for git cat-file remote-object-info command, which cannot work without it. However a subsequent commit needs fetch_object_info() to not die, to be able to fallback.

Add "enum fetch_object_info_status" so that fetch_object_info() can report this case to its callers. It is used in a subsequent commit.

Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
---
 fetch-object-info.h | 6 ++++++
 1 file changed, 6 insertions(+)
Show changes to fetch-object-info.h +6 −0
diff --git a/fetch-object-info.h b/fetch-object-info.h
index 2fba96c6f7..663a7f3ae7 100644
--- a/fetch-object-info.h
+++ b/fetch-object-info.h
@@ -16,6 +16,12 @@ struct fetch_object_info_results {
 
 #define FETCH_OBJECT_INFO_RESULTS_INIT { 0 }
 
+enum fetch_object_info_status {
+	FETCH_OBJECT_INFO_OK = 0,
+	FETCH_OBJECT_INFO_ERR = -1,
+	FETCH_OBJECT_INFO_NOT_ENABLED = -2,
+};
+
 struct oid_array;
 /*
  * Sends git-cat-file object-info command into the request buf and reads the
-- 
2.54.0
Pablo SabaterSep 30, 2026, 00:21 UTC in reply to Pablo Sabater on lore

[PATCH RFC 3/5] fetch-object-info: return a status instead of dying

A subsequent commit needs fetch_object_info() not to die() when the object-info capability is not enabled on the server, so that it can fall back.

Make fetch_object_info() return FETCH_OBJECT_INFO_NOT_ENABLED instead of die()'ing when the server does not advertise the object-info capability, and propagate the status through the transport layer so that callers of transport_fetch_object_info() can act on it. It is now up to them whether to die() or fall back.

cat-file now dies by itself on FETCH_OBJECT_INFO_NOT_ENABLED, so its behavior is unchanged.

Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
---
 builtin/cat-file.c   |  4 ++++
 fetch-object-info.c  | 17 +++++++++--------
 fetch-object-info.h  | 18 +++++++++++-------
 transport-helper.c   |  6 +++---
 transport-internal.h |  8 ++++----
 transport.c          | 29 +++++++++++++++--------------
 transport.h          |  7 ++++---
 7 files changed, 50 insertions(+), 39 deletions(-)
Show changes to 7 files +50 −39

builtin/cat-file.c, fetch-object-info.c, fetch-object-info.h, transport-helper.c, transport-internal.h, transport.c, transport.h

diff --git a/builtin/cat-file.c b/builtin/cat-file.c
index 8870a210ec..f4758f2203 100644
--- a/builtin/cat-file.c
+++ b/builtin/cat-file.c
@@ -726,6 +726,10 @@ static int get_remote_info(int argc,
 
 	retval = transport_fetch_object_info(gtransport, object_info_oids,
 					     results);
+
+	if (retval == FETCH_OBJECT_INFO_NOT_ENABLED)
+		die(_("object-info capability is not enabled on the server"));
+
 cleanup:
 	transport_disconnect(gtransport);
 	return retval;
diff --git a/fetch-object-info.c b/fetch-object-info.c
index 0a58308f9b..7e4c922d27 100644
--- a/fetch-object-info.c
+++ b/fetch-object-info.c
@@ -52,13 +52,13 @@ static int parse_object_size(const char *s, size_t *res)
 	return 0;
 }
 
-void fetch_object_info(const enum protocol_version version,
-		       const struct string_list *server_options,
-		       const struct oid_array *oids,
-		       struct packet_reader *reader,
-		       struct fetch_object_info_results *results,
-		       const int stateless_rpc,
-		       const int fd_out)
+enum fetch_object_info_status fetch_object_info(const enum protocol_version version,
+						const struct string_list *server_options,
+						const struct oid_array *oids,
+						struct packet_reader *reader,
+						struct fetch_object_info_results *results,
+						const int stateless_rpc,
+						const int fd_out)
 {
 	unsigned ask_size = 0;
 	unsigned ask_type = 0;
@@ -72,7 +72,7 @@ void fetch_object_info(const enum protocol_version version,
 	switch (version) {
 	case protocol_v2:
 		if (!server_supports_v2("object-info"))
-			die(_("object-info capability is not enabled on the server"));
+			return FETCH_OBJECT_INFO_NOT_ENABLED;
 
 		if (results->wants_size &&
 		    server_supports_feature("object-info", "size", 0))
@@ -188,6 +188,7 @@ void fetch_object_info(const enum protocol_version version,
 		    (uintmax_t)oids->nr);
 
 	check_stateless_delimiter(stateless_rpc, reader, "stateless delimiter expected");
+	return FETCH_OBJECT_INFO_OK;
 }
 
 void free_fetch_object_info_results(struct fetch_object_info_results *results)
diff --git a/fetch-object-info.h b/fetch-object-info.h
index 663a7f3ae7..9d2750bc93 100644
--- a/fetch-object-info.h
+++ b/fetch-object-info.h
@@ -32,14 +32,18 @@ struct oid_array;
  * the server both advertised and answered with. An array left NULL means the
  * attribute is not available.
  * Release them with free_fetch_object_info_results().
+ *
+ * Returns FETCH_OBJECT_INFO_NOT_ENABLED if the server does not advertise the
+ * object-info capability, FETCH_OBJECT_INFO_OK otherwise.
+ * die()'s on any other error.
  */
-void fetch_object_info(enum protocol_version version,
-		       const struct string_list *server_options,
-		       const struct oid_array *oids,
-		       struct packet_reader *reader,
-		       struct fetch_object_info_results *results,
-		       int stateless_rpc,
-		       int fd_out);
+enum fetch_object_info_status fetch_object_info(enum protocol_version version,
+						const struct string_list *server_options,
+						const struct oid_array *oids,
+						struct packet_reader *reader,
+						struct fetch_object_info_results *results,
+						int stateless_rpc,
+						int fd_out);
 
 void free_fetch_object_info_results(struct fetch_object_info_results *results);
 
diff --git a/transport-helper.c b/transport-helper.c
index d5a064d386..855b53da59 100644
--- a/transport-helper.c
+++ b/transport-helper.c
@@ -786,9 +786,9 @@ static int fetch_refs(struct transport *transport,
 	return -1;
 }
 
-static int fetch_object_info_helper(struct transport *transport,
-				    const struct oid_array *oids,
-				    struct fetch_object_info_results *results)
+static enum fetch_object_info_status fetch_object_info_helper(struct transport *transport,
+							      const struct oid_array *oids,
+							      struct fetch_object_info_results *results)
 {
 	get_helper(transport);
 	if (process_connect(transport, 0))
diff --git a/transport-internal.h b/transport-internal.h
index 626ceaae2b..067134081c 100644
--- a/transport-internal.h
+++ b/transport-internal.h
@@ -2,13 +2,13 @@
 #define TRANSPORT_INTERNAL_H
 
 #include "connect.h"
+#include "fetch-object-info.h"
 
 struct ref;
 struct transport;
 struct strvec;
 struct transport_ls_refs_options;
 struct oid_array;
-struct fetch_object_info_results;
 
 struct transport_vtable {
 	/**
@@ -53,9 +53,9 @@ struct transport_vtable {
 	 *
 	 * Uses object-info capability of v2 protocol.
 	 */
-	int (*fetch_object_info)(struct transport *transport,
-				 const struct oid_array *oids,
-				 struct fetch_object_info_results *results);
+	enum fetch_object_info_status (*fetch_object_info)(struct transport *transport,
+							   const struct oid_array *oids,
+							   struct fetch_object_info_results *results);
 
 	/**
 	 * Push the objects and refs. Send the necessary objects, and
diff --git a/transport.c b/transport.c
index 25e2c14a7b..561764cb6a 100644
--- a/transport.c
+++ b/transport.c
@@ -433,11 +433,11 @@ static int get_bundle_uri(struct transport *transport)
 				     transport->bundles, stateless_rpc);
 }
 
-static int fetch_object_info_via_pack(struct transport *transport,
-				      const struct oid_array *oids,
-				      struct fetch_object_info_results *results)
+static enum fetch_object_info_status fetch_object_info_via_pack(struct transport *transport,
+								const struct oid_array *oids,
+								struct fetch_object_info_results *results)
 {
-	int ret = 0;
+	enum fetch_object_info_status ret = FETCH_OBJECT_INFO_OK;
 	struct git_transport_data *data = transport->data;
 	struct packet_reader reader;
 
@@ -450,26 +450,27 @@ static int fetch_object_info_via_pack(struct transport *transport,
 	data->version = discover_version(&reader);
 	transport->hash_algo = reader.hash_algo;
 
-	fetch_object_info(data->version,
-			  transport->server_options,
-			  oids,
-			  &reader,
-			  results,
-			  transport->stateless_rpc, data->fd[1]);
+	ret = fetch_object_info(data->version,
+				transport->server_options,
+				oids,
+				&reader,
+				results,
+				transport->stateless_rpc,
+				data->fd[1]);
 
 	close(data->fd[0]);
 	if (data->fd[1] >= 0)
 		close(data->fd[1]);
 	if (finish_connect(data->conn))
-		ret = -1;
+		ret = FETCH_OBJECT_INFO_ERR;
 	data->conn = NULL;
 
 	return ret;
 }
 
-int transport_fetch_object_info(struct transport *transport,
-				const struct oid_array *oids,
-				struct fetch_object_info_results *results)
+enum fetch_object_info_status transport_fetch_object_info(struct transport *transport,
+							  const struct oid_array *oids,
+							  struct fetch_object_info_results *results)
 {
 	if (!transport->vtable->fetch_object_info)
 		die(_("remote does not support object-info"));
diff --git a/transport.h b/transport.h
index 39193d0077..c1671639d6 100644
--- a/transport.h
+++ b/transport.h
@@ -1,6 +1,7 @@
 #ifndef TRANSPORT_H
 #define TRANSPORT_H
 
+#include "fetch-object-info.h"
 #include "run-command.h"
 #include "remote.h"
 #include "list-objects-filter-options.h"
@@ -314,9 +315,9 @@ int transport_fetch_refs(struct transport *transport, struct ref *refs);
 /*
  * Fetch the object info from remote
  */
-int transport_fetch_object_info(struct transport *transport,
-				const struct oid_array *oids,
-				struct fetch_object_info_results *results);
+enum fetch_object_info_status transport_fetch_object_info(struct transport *transport,
+							  const struct oid_array *oids,
+							  struct fetch_object_info_results *results);
 
 /*
  * If this flag is set, unlocking will avoid to call non-async-signal-safe
-- 
2.54.0
Pablo SabaterSep 30, 2026, 00:21 UTC in reply to Pablo Sabater on lore

[PATCH RFC 4/5] backfill: add --dry-run option

Users have no way to know how many blobs git backfill is going to download before running it.

Add a new --dry-run option to the backfill command. The objects are walked as usual, but instead of fetching each batch of missing blobs they are only counted, and the total is printed at the end.

A subsequent commit will also print their size when the server supports the object-info capability.

Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
---
 Documentation/git-backfill.adoc |  6 +++++-
 builtin/backfill.c              | 37 +++++++++++++++++++++++++++++++++----
 t/t5620-backfill.sh             | 26 ++++++++++++++++++++++++++
 3 files changed, 64 insertions(+), 5 deletions(-)
Show changes to 3 files +64 −5

Documentation/git-backfill.adoc, builtin/backfill.c, t/t5620-backfill.sh

diff --git a/Documentation/git-backfill.adoc b/Documentation/git-backfill.adoc
index 82d6a1969d..08f19fea17 100644
--- a/Documentation/git-backfill.adoc
+++ b/Documentation/git-backfill.adoc
@@ -9,7 +9,7 @@ git-backfill - Download missing objects in a partial clone
 SYNOPSIS
 --------
 [synopsis]
-git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]
+git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]
 
 DESCRIPTION
 -----------
@@ -70,6 +70,10 @@ OPTIONS
 	--onto TARGET A..B`, where A..B normally excludes A but you need
 	the blobs from A as well.  `--include-edges` is the default.
 
+`--dry-run`::
+	Do not download any objects. Instead, print the number of
+	missing blobs that would be downloaded.
+
 `<revision-range>`::
 	Backfill only blobs reachable from commits in the specified
 	revision range.  When no _<revision-range>_ is specified, it
diff --git a/builtin/backfill.c b/builtin/backfill.c
index e71e0f4742..6019112966 100644
--- a/builtin/backfill.c
+++ b/builtin/backfill.c
@@ -26,7 +26,7 @@
 #include "path-walk.h"
 
 static const char * const builtin_backfill_usage[] = {
-	N_("git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]"),
+	N_("git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]"),
 	NULL
 };
 
@@ -36,6 +36,8 @@ struct backfill_context {
 	size_t min_batch_size;
 	int sparse;
 	int include_edges;
+	int dry_run;
+	size_t total_batch_nr;
 	struct rev_info revs;
 };
 
@@ -58,6 +60,15 @@ static void download_batch(struct backfill_context *ctx)
 	odb_reprepare(ctx->repo->objects);
 }
 
+static void dry_run_batch(struct backfill_context *ctx)
+{
+	if (!ctx->current_batch.nr)
+		return;
+
+	ctx->total_batch_nr += ctx->current_batch.nr;
+	oid_array_clear(&ctx->current_batch);
+}
+
 static int fill_missing_blobs(const char *path UNUSED,
 			      struct oid_array *list,
 			      enum object_type type,
@@ -73,8 +84,12 @@ static int fill_missing_blobs(const char *path UNUSED,
 			oid_array_append(&ctx->current_batch, &list->oid[i]);
 	}
 
-	if (ctx->current_batch.nr >= ctx->min_batch_size)
-		download_batch(ctx);
+	if (ctx->current_batch.nr >= ctx->min_batch_size) {
+		if (ctx->dry_run)
+			dry_run_batch(ctx);
+		else
+			download_batch(ctx);
+	}
 
 	return 0;
 }
@@ -131,10 +146,23 @@ static int do_backfill(struct backfill_context *ctx)
 
 	ret = walk_objects_by_path(&info);
 
+	if (ret)
+		goto end;
+
 	/* Download the objects that did not fill a batch. */
-	if (!ret)
+	if (!ctx->dry_run) {
 		download_batch(ctx);
+		goto end;
+	}
+
+	dry_run_batch(ctx);
+
+	printf(Q_("After backfill, %" PRIuMAX " blob would be fetched.\n",
+		  "After backfill, %" PRIuMAX " blobs would be fetched.\n",
+		  (unsigned long)ctx->total_batch_nr),
+	       (uintmax_t)ctx->total_batch_nr);
 
+end:
 	path_walk_info_clear(&info);
 	return ret;
 }
@@ -157,6 +185,7 @@ int cmd_backfill(int argc, const char **argv, const char *prefix, struct reposit
 			 N_("Restrict the missing objects to the current sparse-checkout")),
 		OPT_BOOL(0, "include-edges", &ctx.include_edges,
 			 N_("Include blobs from boundary commits in the backfill")),
+		OPT__DRY_RUN(&ctx.dry_run, N_("Preview the number of blobs to be fetched")),
 		OPT_END(),
 	};
 	struct repo_config_values *cfg = repo_config_values(the_repository);
diff --git a/t/t5620-backfill.sh b/t/t5620-backfill.sh
index 7462280470..e76fa6081b 100755
--- a/t/t5620-backfill.sh
+++ b/t/t5620-backfill.sh
@@ -141,6 +141,32 @@ test_expect_success 'do partial clone 2, backfill min batch size' '
 	test_line_count = 0 revs2
 '
 
+test_expect_success '--dry-run reports missing blobs without fetching them' '
+	test_when_finished "rm -rf backfill-dry-run dry-trace" &&
+	git clone --no-checkout --filter=blob:none \
+		--single-branch --branch=main \
+		"file://$(pwd)/srv.bare" backfill-dry-run &&
+
+	GIT_TRACE2_EVENT="$(pwd)/dry-trace" git \
+		-C backfill-dry-run backfill --dry-run >out &&
+
+	test_grep "48 blobs would be fetched" out &&
+	test_grep ! fetch_count dry-trace &&
+	git -C backfill-dry-run rev-list --quiet --objects --missing=print HEAD >missing &&
+	test_line_count = 48 missing
+'
+
+test_expect_success '--dry-run with no missing blobs' '
+	test_when_finished rm -rf backfill-dry-run &&
+	git clone --no-checkout --filter=blob:none \
+		--single-branch --branch=main \
+		"file://$(pwd)/srv.bare" backfill-dry-run &&
+	git -C backfill-dry-run backfill &&
+
+	git -C backfill-dry-run backfill --dry-run >out &&
+	test_grep "0 blobs would be fetched" out
+'
+
 test_expect_success 'backfill --sparse without sparse-checkout fails' '
 	git init not-sparse &&
 	test_must_fail git -C not-sparse backfill --sparse 2>err &&
-- 
2.54.0
Pablo SabaterSep 30, 2026, 00:21 UTC in reply to Pablo Sabater on lore

[PATCH RFC 5/5] backfill: report total size of missing blobs in --dry-run

The number of missing blobs says little about how much data a backfill will transfer. In --dry-run mode, if the server has the object-info capability enabled, fetch the size of the objects to be fetched and show the total size, if the server doesn't have the capability enabled, fall back to printing only the count.

The reported size is the sum of the uncompressed object sizes, so it is an upper bound: once fetched, the blobs are stored compressed and possibly deltified in a packfile, and usually take less space on disk.

Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
---
 Documentation/git-backfill.adoc |  6 +++-
 builtin/backfill.c              | 70 ++++++++++++++++++++++++++++++++++++++---
 t/t5620-backfill.sh             | 22 +++++++++++++
 3 files changed, 92 insertions(+), 6 deletions(-)
Show changes to 3 files +92 −6

Documentation/git-backfill.adoc, builtin/backfill.c, t/t5620-backfill.sh

diff --git a/Documentation/git-backfill.adoc b/Documentation/git-backfill.adoc
index 08f19fea17..660e1c17c0 100644
--- a/Documentation/git-backfill.adoc
+++ b/Documentation/git-backfill.adoc
@@ -72,7 +72,11 @@ OPTIONS
 
 `--dry-run`::
 	Do not download any objects. Instead, print the number of
-	missing blobs that would be downloaded.
+	missing blobs that would be downloaded and, if the promisor
+	remote supports the `object-info` capability, their total
+	size. This is the sum of the uncompressed sizes of the blobs,
+	so the space used on disk after a real backfill is usually
+	smaller.
 
 `<revision-range>`::
 	Backfill only blobs reachable from commits in the specified
diff --git a/builtin/backfill.c b/builtin/backfill.c
index 6019112966..caf1a64e95 100644
--- a/builtin/backfill.c
+++ b/builtin/backfill.c
@@ -24,6 +24,9 @@
 #include "progress.h"
 #include "packfile.h"
 #include "path-walk.h"
+#include "transport.h"
+#include "remote.h"
+#include "fetch-object-info.h"
 
 static const char * const builtin_backfill_usage[] = {
 	N_("git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]"),
@@ -37,7 +40,11 @@ struct backfill_context {
 	int sparse;
 	int include_edges;
 	int dry_run;
+	int object_info_enabled;
+	size_t total_batch_size;
 	size_t total_batch_nr;
+	struct transport *object_info_transport;
+	struct fetch_object_info_results object_info_results;
 	struct rev_info revs;
 };
 
@@ -62,10 +69,47 @@ static void download_batch(struct backfill_context *ctx)
 
 static void dry_run_batch(struct backfill_context *ctx)
 {
+	struct fetch_object_info_results *results = &ctx->object_info_results;
+	enum fetch_object_info_status status;
+
 	if (!ctx->current_batch.nr)
 		return;
 
 	ctx->total_batch_nr += ctx->current_batch.nr;
+
+	if (!ctx->object_info_enabled)
+		goto cleanup;
+
+	if (!ctx->object_info_transport) {
+		struct promisor_remote *promise =
+			repo_promisor_remote_find(ctx->repo, NULL);
+		struct remote *remote = NULL;
+
+		if (!promise || !(remote = remote_get(promise->name)))
+			die(_("--dry-run requires a promisor remote"));
+
+		ctx->object_info_transport = transport_get(remote, NULL);
+
+		if (!ctx->object_info_transport->smart_options)
+			die(_("failed to get object info: smart options required"));
+	}
+
+	results->wants_size = 1;
+	status = transport_fetch_object_info(ctx->object_info_transport,
+					     &ctx->current_batch,
+					     results);
+
+	if (status == FETCH_OBJECT_INFO_NOT_ENABLED ||
+	    !results->sizes) {
+		ctx->object_info_enabled = 0;
+		goto cleanup;
+	}
+
+	for (size_t i = 0; i < results->nr; i++)
+		ctx->total_batch_size += results->sizes[i];
+
+cleanup:
+	free_fetch_object_info_results(&ctx->object_info_results);
 	oid_array_clear(&ctx->current_batch);
 }
 
@@ -157,12 +201,25 @@ static int do_backfill(struct backfill_context *ctx)
 
 	dry_run_batch(ctx);
 
-	printf(Q_("After backfill, %" PRIuMAX " blob would be fetched.\n",
-		  "After backfill, %" PRIuMAX " blobs would be fetched.\n",
-		  (unsigned long)ctx->total_batch_nr),
-	       (uintmax_t)ctx->total_batch_nr);
+	if (ctx->object_info_enabled && ctx->total_batch_nr) {
+		struct strbuf size = STRBUF_INIT;
+
+		strbuf_humanise_bytes(&size, ctx->total_batch_size);
+		printf(Q_("After backfill, %" PRIuMAX " blob would be fetched (%s).\n",
+			  "After backfill, %" PRIuMAX " blobs would be fetched (%s).\n",
+			  (unsigned long)ctx->total_batch_nr),
+		       (uintmax_t)ctx->total_batch_nr, size.buf);
+		strbuf_release(&size);
+	} else {
+		printf(Q_("After backfill, %" PRIuMAX " blob would be fetched.\n",
+			  "After backfill, %" PRIuMAX " blobs would be fetched.\n",
+			  (unsigned long)ctx->total_batch_nr),
+		       (uintmax_t)ctx->total_batch_nr);
+	}
 
 end:
+	if (ctx->object_info_transport)
+		transport_disconnect(ctx->object_info_transport);
 	path_walk_info_clear(&info);
 	return ret;
 }
@@ -177,6 +234,8 @@ int cmd_backfill(int argc, const char **argv, const char *prefix, struct reposit
 		.sparse = -1,
 		.revs = REV_INFO_INIT,
 		.include_edges = 1,
+		.object_info_results = FETCH_OBJECT_INFO_RESULTS_INIT,
+		.object_info_enabled = 1,
 	};
 	struct option options[] = {
 		OPT_UNSIGNED(0, "min-batch-size", &ctx.min_batch_size,
@@ -185,7 +244,8 @@ int cmd_backfill(int argc, const char **argv, const char *prefix, struct reposit
 			 N_("Restrict the missing objects to the current sparse-checkout")),
 		OPT_BOOL(0, "include-edges", &ctx.include_edges,
 			 N_("Include blobs from boundary commits in the backfill")),
-		OPT__DRY_RUN(&ctx.dry_run, N_("Preview the number of blobs to be fetched")),
+		OPT__DRY_RUN(&ctx.dry_run, N_("Preview the number of blobs and their total "
+					      "size to be fetched")),
 		OPT_END(),
 	};
 	struct repo_config_values *cfg = repo_config_values(the_repository);
diff --git a/t/t5620-backfill.sh b/t/t5620-backfill.sh
index e76fa6081b..b21feb625c 100755
--- a/t/t5620-backfill.sh
+++ b/t/t5620-backfill.sh
@@ -167,6 +167,28 @@ test_expect_success '--dry-run with no missing blobs' '
 	test_grep "0 blobs would be fetched" out
 '
 
+test_expect_success '--dry-run reports total size with object-info' '
+	test_config -C srv.bare transfer.advertiseobjectinfo true &&
+	test_when_finished rm -rf backfill-dry-run &&
+	git clone --no-checkout --filter=blob:none \
+		--single-branch --branch=main \
+		"file://$(pwd)/srv.bare" backfill-dry-run &&
+
+	git -C backfill-dry-run backfill --dry-run >out &&
+	test_grep "48 blobs would be fetched (.*)" out
+'
+
+test_expect_success '--dry-run reports only the count without object-info' '
+	test_config -C srv.bare transfer.advertiseobjectinfo false &&
+	test_when_finished rm -rf backfill-dry-run &&
+	git clone --no-checkout --filter=blob:none \
+		--single-branch --branch=main \
+		"file://$(pwd)/srv.bare" backfill-dry-run &&
+
+	git -C backfill-dry-run backfill --dry-run >out &&
+	test_grep "48 blobs would be fetched\.$" out
+'
+
 test_expect_success 'backfill --sparse without sparse-checkout fails' '
 	git init not-sparse &&
 	test_must_fail git -C not-sparse backfill --sparse 2>err &&
-- 
2.54.0
Karthik NayakSep 30, 2026, 10:56 UTC in reply to Pablo Sabater on lore

Re: [PATCH RFC 2/5] fetch-object-info: add enum for fetch_object_info() statuses

Pablo Sabater <pabloosabaterr@gmail.com> writes:
Show 6 quoted lines
> fetch_object_info() dies when the server does not advertise the
> object-info capability. That is fine for git cat-file
> remote-object-info command, which cannot work without it. However
> a subsequent commit needs fetch_object_info() to not die, to be
> able to fallback.
>
Nit: The last sentence reads a little weird, perhaps:
  However a subsequent commit uses fetch_object_info() optionally and
  requires it to not die.
Or something?

But I think we should just squash this into the next commit. It doesn't really need to standout on its own.

Show 28 quoted lines
> Add "enum fetch_object_info_status" so that fetch_object_info() can
> report this case to its callers. It is used in a subsequent commit.
>
> Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
> ---
>  fetch-object-info.h | 6 ++++++
>  1 file changed, 6 insertions(+)
>
> diff --git a/fetch-object-info.h b/fetch-object-info.h
> index 2fba96c6f7..663a7f3ae7 100644
> --- a/fetch-object-info.h
> +++ b/fetch-object-info.h
> @@ -16,6 +16,12 @@ struct fetch_object_info_results {
>
>  #define FETCH_OBJECT_INFO_RESULTS_INIT { 0 }
>
> +enum fetch_object_info_status {
> +	FETCH_OBJECT_INFO_OK = 0,
> +	FETCH_OBJECT_INFO_ERR = -1,
> +	FETCH_OBJECT_INFO_NOT_ENABLED = -2,
> +};
> +
>  struct oid_array;
>  /*
>   * Sends git-cat-file object-info command into the request buf and reads the
>
> --
> 2.54.0
Karthik NayakSep 30, 2026, 11:06 UTC in reply to Pablo Sabater on lore

Re: [PATCH RFC 4/5] backfill: add --dry-run option

Pablo Sabater <pabloosabaterr@gmail.com> writes:
Show 59 quoted lines
> Users have no way to know how many blobs git backfill is going to
> download before running it.
>
> Add a new --dry-run option to the backfill command. The objects are
> walked as usual, but instead of fetching each batch of missing blobs
> they are only counted, and the total is printed at the end.
>
> A subsequent commit will also print their size when the server supports
> the object-info capability.
>
> Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
> ---
>  Documentation/git-backfill.adoc |  6 +++++-
>  builtin/backfill.c              | 37 +++++++++++++++++++++++++++++++++----
>  t/t5620-backfill.sh             | 26 ++++++++++++++++++++++++++
>  3 files changed, 64 insertions(+), 5 deletions(-)
>
> diff --git a/Documentation/git-backfill.adoc b/Documentation/git-backfill.adoc
> index 82d6a1969d..08f19fea17 100644
> --- a/Documentation/git-backfill.adoc
> +++ b/Documentation/git-backfill.adoc
> @@ -9,7 +9,7 @@ git-backfill - Download missing objects in a partial clone
>  SYNOPSIS
>  --------
>  [synopsis]
> -git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]
> +git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]
>
>  DESCRIPTION
>  -----------
> @@ -70,6 +70,10 @@ OPTIONS
>  	--onto TARGET A..B`, where A..B normally excludes A but you need
>  	the blobs from A as well.  `--include-edges` is the default.
>
> +`--dry-run`::
> +	Do not download any objects. Instead, print the number of
> +	missing blobs that would be downloaded.
> +
>  `<revision-range>`::
>  	Backfill only blobs reachable from commits in the specified
>  	revision range.  When no _<revision-range>_ is specified, it
> diff --git a/builtin/backfill.c b/builtin/backfill.c
> index e71e0f4742..6019112966 100644
> --- a/builtin/backfill.c
> +++ b/builtin/backfill.c
> @@ -26,7 +26,7 @@
>  #include "path-walk.h"
>
>  static const char * const builtin_backfill_usage[] = {
> -	N_("git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]"),
> +	N_("git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]"),
>  	NULL
>  };
>
> @@ -36,6 +36,8 @@ struct backfill_context {
>  	size_t min_batch_size;
>  	int sparse;
>  	int include_edges;
> +	int dry_run;
Nit: This could be a bool, since `OPT__DRY_RUN` uses `OPT_BOOL` internally.
> +	size_t total_batch_nr;
Show 9 quoted lines
>  	struct rev_info revs;
>  };
>
> @@ -58,6 +60,15 @@ static void download_batch(struct backfill_context *ctx)
>  	odb_reprepare(ctx->repo->objects);
>  }
>
> +static void dry_run_batch(struct backfill_context *ctx)
> +{

While it is used during dry_run, probably makes more sense to rename it to `count_batch()` since that's what it does.

Show 46 quoted lines
> +	if (!ctx->current_batch.nr)
> +		return;
> +
> +	ctx->total_batch_nr += ctx->current_batch.nr;
> +	oid_array_clear(&ctx->current_batch);
> +}
> +
>  static int fill_missing_blobs(const char *path UNUSED,
>  			      struct oid_array *list,
>  			      enum object_type type,
> @@ -73,8 +84,12 @@ static int fill_missing_blobs(const char *path UNUSED,
>  			oid_array_append(&ctx->current_batch, &list->oid[i]);
>  	}
>
> -	if (ctx->current_batch.nr >= ctx->min_batch_size)
> -		download_batch(ctx);
> +	if (ctx->current_batch.nr >= ctx->min_batch_size) {
> +		if (ctx->dry_run)
> +			dry_run_batch(ctx);
> +		else
> +			download_batch(ctx);
> +	}
>
>  	return 0;
>  }
> @@ -131,10 +146,23 @@ static int do_backfill(struct backfill_context *ctx)
>
>  	ret = walk_objects_by_path(&info);
>
> +	if (ret)
> +		goto end;
> +
>  	/* Download the objects that did not fill a batch. */
> -	if (!ret)
> +	if (!ctx->dry_run) {
>  		download_batch(ctx);
> +		goto end;
> +	}
> +
> +	dry_run_batch(ctx);
> +
> +	printf(Q_("After backfill, %" PRIuMAX " blob would be fetched.\n",
> +		  "After backfill, %" PRIuMAX " blobs would be fetched.\n",
> +		  (unsigned long)ctx->total_batch_nr),
> +	       (uintmax_t)ctx->total_batch_nr);
>
Nit: Okay so we have a goto inside the first if(...), which skips this
section. I would have found it easier to read if it was
if (dry_run)
   count()
else
   download()
Show 52 quoted lines
> +end:
>  	path_walk_info_clear(&info);
>  	return ret;
>  }
> @@ -157,6 +185,7 @@ int cmd_backfill(int argc, const char **argv, const char *prefix, struct reposit
>  			 N_("Restrict the missing objects to the current sparse-checkout")),
>  		OPT_BOOL(0, "include-edges", &ctx.include_edges,
>  			 N_("Include blobs from boundary commits in the backfill")),
> +		OPT__DRY_RUN(&ctx.dry_run, N_("Preview the number of blobs to be fetched")),
>  		OPT_END(),
>  	};
>  	struct repo_config_values *cfg = repo_config_values(the_repository);
> diff --git a/t/t5620-backfill.sh b/t/t5620-backfill.sh
> index 7462280470..e76fa6081b 100755
> --- a/t/t5620-backfill.sh
> +++ b/t/t5620-backfill.sh
> @@ -141,6 +141,32 @@ test_expect_success 'do partial clone 2, backfill min batch size' '
>  	test_line_count = 0 revs2
>  '
>
> +test_expect_success '--dry-run reports missing blobs without fetching them' '
> +	test_when_finished "rm -rf backfill-dry-run dry-trace" &&
> +	git clone --no-checkout --filter=blob:none \
> +		--single-branch --branch=main \
> +		"file://$(pwd)/srv.bare" backfill-dry-run &&
> +
> +	GIT_TRACE2_EVENT="$(pwd)/dry-trace" git \
> +		-C backfill-dry-run backfill --dry-run >out &&
> +
> +	test_grep "48 blobs would be fetched" out &&
> +	test_grep ! fetch_count dry-trace &&
> +	git -C backfill-dry-run rev-list --quiet --objects --missing=print HEAD >missing &&
> +	test_line_count = 48 missing
> +'
> +
> +test_expect_success '--dry-run with no missing blobs' '
> +	test_when_finished rm -rf backfill-dry-run &&
> +	git clone --no-checkout --filter=blob:none \
> +		--single-branch --branch=main \
> +		"file://$(pwd)/srv.bare" backfill-dry-run &&
> +	git -C backfill-dry-run backfill &&
> +
> +	git -C backfill-dry-run backfill --dry-run >out &&
> +	test_grep "0 blobs would be fetched" out
> +'
> +
>  test_expect_success 'backfill --sparse without sparse-checkout fails' '
>  	git init not-sparse &&
>  	test_must_fail git -C not-sparse backfill --sparse 2>err &&
>
> --
> 2.54.0
Pablo SabaterSep 30, 2026, 11:59 UTC in reply to Karthik Nayak on lore

Re: [PATCH RFC 2/5] fetch-object-info: add enum for fetch_object_info() statuses

On Wed Sep 30, 2026 at 11:56 AM WEST, Karthik Nayak wrote:
Show 18 quoted lines
> Pablo Sabater <pabloosabaterr@gmail.com> writes:
>
>> fetch_object_info() dies when the server does not advertise the
>> object-info capability. That is fine for git cat-file
>> remote-object-info command, which cannot work without it. However
>> a subsequent commit needs fetch_object_info() to not die, to be
>> able to fallback.
>>
>
> Nit: The last sentence reads a little weird, perhaps:
>
>   However a subsequent commit uses fetch_object_info() optionally and
>   requires it to not die.
>
> Or something?
>
> But I think we should just squash this into the next commit. It doesn't
> really need to standout on its own.

Okay, squashing them makes sense. I only split them to avoid having the enum buried among the signature changes.

Will fix, thanks.
Show 29 quoted lines
>
>> Add "enum fetch_object_info_status" so that fetch_object_info() can
>> report this case to its callers. It is used in a subsequent commit.
>>
>> Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
>> ---
>>  fetch-object-info.h | 6 ++++++
>>  1 file changed, 6 insertions(+)
>>
>> diff --git a/fetch-object-info.h b/fetch-object-info.h
>> index 2fba96c6f7..663a7f3ae7 100644
>> --- a/fetch-object-info.h
>> +++ b/fetch-object-info.h
>> @@ -16,6 +16,12 @@ struct fetch_object_info_results {
>>
>>  #define FETCH_OBJECT_INFO_RESULTS_INIT { 0 }
>>
>> +enum fetch_object_info_status {
>> +	FETCH_OBJECT_INFO_OK = 0,
>> +	FETCH_OBJECT_INFO_ERR = -1,
>> +	FETCH_OBJECT_INFO_NOT_ENABLED = -2,
>> +};
>> +
>>  struct oid_array;
>>  /*
>>   * Sends git-cat-file object-info command into the request buf and reads the
>>
>> --
>> 2.54.0
Pablo SabaterSep 30, 2026, 12:19 UTC in reply to Karthik Nayak on lore

Re: [PATCH RFC 4/5] backfill: add --dry-run option

On Wed Sep 30, 2026 at 12:06 PM WEST, Karthik Nayak wrote:
Show 63 quoted lines
> Pablo Sabater <pabloosabaterr@gmail.com> writes:
>
>> Users have no way to know how many blobs git backfill is going to
>> download before running it.
>>
>> Add a new --dry-run option to the backfill command. The objects are
>> walked as usual, but instead of fetching each batch of missing blobs
>> they are only counted, and the total is printed at the end.
>>
>> A subsequent commit will also print their size when the server supports
>> the object-info capability.
>>
>> Signed-off-by: Pablo Sabater <pabloosabaterr@gmail.com>
>> ---
>>  Documentation/git-backfill.adoc |  6 +++++-
>>  builtin/backfill.c              | 37 +++++++++++++++++++++++++++++++++----
>>  t/t5620-backfill.sh             | 26 ++++++++++++++++++++++++++
>>  3 files changed, 64 insertions(+), 5 deletions(-)
>>
>> diff --git a/Documentation/git-backfill.adoc b/Documentation/git-backfill.adoc
>> index 82d6a1969d..08f19fea17 100644
>> --- a/Documentation/git-backfill.adoc
>> +++ b/Documentation/git-backfill.adoc
>> @@ -9,7 +9,7 @@ git-backfill - Download missing objects in a partial clone
>>  SYNOPSIS
>>  --------
>>  [synopsis]
>> -git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]
>> +git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]
>>
>>  DESCRIPTION
>>  -----------
>> @@ -70,6 +70,10 @@ OPTIONS
>>  	--onto TARGET A..B`, where A..B normally excludes A but you need
>>  	the blobs from A as well.  `--include-edges` is the default.
>>
>> +`--dry-run`::
>> +	Do not download any objects. Instead, print the number of
>> +	missing blobs that would be downloaded.
>> +
>>  `<revision-range>`::
>>  	Backfill only blobs reachable from commits in the specified
>>  	revision range.  When no _<revision-range>_ is specified, it
>> diff --git a/builtin/backfill.c b/builtin/backfill.c
>> index e71e0f4742..6019112966 100644
>> --- a/builtin/backfill.c
>> +++ b/builtin/backfill.c
>> @@ -26,7 +26,7 @@
>>  #include "path-walk.h"
>>
>>  static const char * const builtin_backfill_usage[] = {
>> -	N_("git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [<revision-range>]"),
>> +	N_("git backfill [--min-batch-size=<n>] [--[no-]sparse] [--[no-]include-edges] [--dry-run] [<revision-range>]"),
>>  	NULL
>>  };
>>
>> @@ -36,6 +36,8 @@ struct backfill_context {
>>  	size_t min_batch_size;
>>  	int sparse;
>>  	int include_edges;
>> +	int dry_run;
>
> Nit: This could be a bool, since `OPT__DRY_RUN` uses `OPT_BOOL` internally.

Will do, didn't know that using bool was a thing ;). Is it also prefered for 0/1 functions?

Show 16 quoted lines
>
>> +	size_t total_batch_nr;
>
>
>>  	struct rev_info revs;
>>  };
>>
>> @@ -58,6 +60,15 @@ static void download_batch(struct backfill_context *ctx)
>>  	odb_reprepare(ctx->repo->objects);
>>  }
>>
>> +static void dry_run_batch(struct backfill_context *ctx)
>> +{
>
> While it is used during dry_run, probably makes more sense to rename it
> to `count_batch()` since that's what it does.

count_batch() works for me, I think that "count" doesn't fit too well for summing the size of the objects, might opt for another name if I think of a better name.

Prob a comment helps.
Show 55 quoted lines
>
>> +	if (!ctx->current_batch.nr)
>> +		return;
>> +
>> +	ctx->total_batch_nr += ctx->current_batch.nr;
>> +	oid_array_clear(&ctx->current_batch);
>> +}
>> +
>>  static int fill_missing_blobs(const char *path UNUSED,
>>  			      struct oid_array *list,
>>  			      enum object_type type,
>> @@ -73,8 +84,12 @@ static int fill_missing_blobs(const char *path UNUSED,
>>  			oid_array_append(&ctx->current_batch, &list->oid[i]);
>>  	}
>>
>> -	if (ctx->current_batch.nr >= ctx->min_batch_size)
>> -		download_batch(ctx);
>> +	if (ctx->current_batch.nr >= ctx->min_batch_size) {
>> +		if (ctx->dry_run)
>> +			dry_run_batch(ctx);
>> +		else
>> +			download_batch(ctx);
>> +	}
>>
>>  	return 0;
>>  }
>> @@ -131,10 +146,23 @@ static int do_backfill(struct backfill_context *ctx)
>>
>>  	ret = walk_objects_by_path(&info);
>>
>> +	if (ret)
>> +		goto end;
>> +
>>  	/* Download the objects that did not fill a batch. */
>> -	if (!ret)
>> +	if (!ctx->dry_run) {
>>  		download_batch(ctx);
>> +		goto end;
>> +	}
>> +
>> +	dry_run_batch(ctx);
>> +
>> +	printf(Q_("After backfill, %" PRIuMAX " blob would be fetched.\n",
>> +		  "After backfill, %" PRIuMAX " blobs would be fetched.\n",
>> +		  (unsigned long)ctx->total_batch_nr),
>> +	       (uintmax_t)ctx->total_batch_nr);
>>
>
> Nit: Okay so we have a goto inside the first if(...), which skips this
> section. I would have found it easier to read if it was
>
> if (dry_run)
>    count()
> else
>    download()
Will change it, thanks.
>
[snip]
Junio C HamanoSep 30, 2026, 17:07 UTC in reply to Pablo Sabater on lore

Re: [PATCH RFC 3/5] fetch-object-info: return a status instead of dying

Pablo Sabater <pabloosabaterr@gmail.com> writes:
Show 9 quoted lines
> A subsequent commit needs fetch_object_info() not to die() when the
> object-info capability is not enabled on the server, so that it can
> fall back.
>
> Make fetch_object_info() return FETCH_OBJECT_INFO_NOT_ENABLED instead
> of die()'ing when the server does not advertise the object-info
> capability, and propagate the status through the transport layer so
> that callers of transport_fetch_object_info() can act on it. It is now
> up to them whether to die() or fall back.

It may be just me but unless the client can tell between the server not supporting (i.e., they are unable to enable it even if they wanted to) and not enabling (i.e., they are capable, but are not willing to give it to you), it may make sense to report it as "not available". "not enabled" sounds as if we know that it is the latter and not the former.

The code change looks very cleanly done.
Junio C HamanoSep 30, 2026, 17:17 UTC in reply to Pablo Sabater on lore

Re: [PATCH RFC 5/5] backfill: report total size of missing blobs in --dry-run

Pablo Sabater <pabloosabaterr@gmail.com> writes:
Show 26 quoted lines
> +
> +	if (!ctx->object_info_enabled)
> +		goto cleanup;
> +
> +	if (!ctx->object_info_transport) {
> +		struct promisor_remote *promise =
> +			repo_promisor_remote_find(ctx->repo, NULL);
> +		struct remote *remote = NULL;
> +
> +		if (!promise || !(remote = remote_get(promise->name)))
> +			die(_("--dry-run requires a promisor remote"));
> +
> +		ctx->object_info_transport = transport_get(remote, NULL);
> +
> +		if (!ctx->object_info_transport->smart_options)
> +			die(_("failed to get object info: smart options required"));
> +	}
> +
> +	results->wants_size = 1;
> +	status = transport_fetch_object_info(ctx->object_info_transport,
> +					     &ctx->current_batch,
> +					     results);
> +
> +	if (status == FETCH_OBJECT_INFO_NOT_ENABLED ||
> +	    !results->sizes) {
> +		ctx->object_info_enabled = 0;
Yuck.

Because we cannot tell if they allow you to look at the information, this cannot be helped, but it means anybody that looks at this ctx->object_info_enabled member to decide what to do must be careful.

Is it guaranteed that results.sizes[] have been populated as long as status is not FETCH_OBJECT_INFO_NOT_ENABLED? Can there be other errors that makes result.sizes[] unusable? If that is the case, then it would be cleaner to have a dedicated ctx->sizes_valid member rather than relying on ctx->object_info_enabled member to carry this information ...

Show 20 quoted lines
> +		goto cleanup;
> +	}
> +
> +	for (size_t i = 0; i < results->nr; i++)
> +		ctx->total_batch_size += results->sizes[i];
> +
> +cleanup:
> +	free_fetch_object_info_results(&ctx->object_info_results);
>  	oid_array_clear(&ctx->current_batch);
>  }
>  
> @@ -157,12 +201,25 @@ static int do_backfill(struct backfill_context *ctx)
>  
>  	dry_run_batch(ctx);
>  
> -	printf(Q_("After backfill, %" PRIuMAX " blob would be fetched.\n",
> -		  "After backfill, %" PRIuMAX " blobs would be fetched.\n",
> -		  (unsigned long)ctx->total_batch_nr),
> -	       (uintmax_t)ctx->total_batch_nr);
> +	if (ctx->object_info_enabled && ctx->total_batch_nr) {

... and use it here. Within the design presented in this series, we know we have asked the other end at this point, and the above function may have turned ctx->object_info member off if the information is not there, so this may be safe. But as I said, I am not sure what happens when fetch-object-info returned other kind of errors.

Pablo SabaterSep 30, 2026, 18:03 UTC in reply to Junio C Hamano on lore

Re: [PATCH RFC 3/5] fetch-object-info: return a status instead of dying

On Wed Sep 30, 2026 at 6:07 PM WEST, Junio C Hamano wrote:
Show 20 quoted lines
> Pablo Sabater <pabloosabaterr@gmail.com> writes:
>
>> A subsequent commit needs fetch_object_info() not to die() when the
>> object-info capability is not enabled on the server, so that it can
>> fall back.
>>
>> Make fetch_object_info() return FETCH_OBJECT_INFO_NOT_ENABLED instead
>> of die()'ing when the server does not advertise the object-info
>> capability, and propagate the status through the transport layer so
>> that callers of transport_fetch_object_info() can act on it. It is now
>> up to them whether to die() or fall back.
>
> It may be just me but unless the client can tell between the server
> not supporting (i.e., they are unable to enable it even if they
> wanted to) and not enabling (i.e., they are capable, but are not
> willing to give it to you), it may make sense to report it as "not
> available".  "not enabled" sounds as if we know that it is the
> latter and not the former.
>
> The code change looks very cleanly done.

Makes sense, I'll rename it to FETCH_OBJECT_INFO_NOT_AVAILABLE. The git cat-file remote-object-info command path die()'d with this message:

	die(_("object-info capability is not enabled on the server"));
I'll update the die() message as well.
Thanks.
Derrick StoleeSep 30, 2026, 18:14 UTC in reply to Pablo Sabater on lore

Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)

On 9/29/2026 8:21 PM, Pablo Sabater wrote:
Show 10 quoted lines
> [Cc'd Derrick Stolee for his work in the backfill(1) command]
> 
> This series adds a --dry-run option to git-backfill(1) that reports how
> many missing blobs would be fetched and, when the remote server
> supports the object-info capability, their total size:
> 
>         $ git backfill --dry-run
>         After backfill, 48 blobs would be fetched (1.20 KiB).
> 
> If the server does not advertise object-info, only the count is shown.

This is a helpful capability, but I'm not sure the size check counts as a "dry run" because it involves a network call (and possibly many depending on --min-batch-size).

Perhaps a different argument would be better, such as --info=(count|size) to make it clear what level of information you want to know in advance and thus how much effort are you willing to put in to discover this.

> I am not a git-backfill(1) user myself, but it seemed useful for users
> to know how much data a backfill would bring in before running it.

I'm not sure that we want to add a feature based on speculation. Git is a collection of "itches" that the contributors needed scratched. The work is motivated by real needs.

While I can see some benefit to curiosity, I'm not sure how much this would prevent users from making their decision as to whether they should run backfill or not.

Show 9 quoted lines
> The number of missing blobs is the sum of the number of blobs to be
> fetched in each batch. The object-info capability lets us ask the server
> for the size of each blob without downloading it, so summing them gives
> an estimate of the total.
> 
> Note that this is an upper bound rather than the exact disk usage:
> object-info reports the uncompressed size of each object, while the
> objects end up stored compressed and possibly deltified in a packfile,
> so the space actually used on disk will usually be smaller.

I don't think the uncompressed size is a useful metric here, as it is likely astronomically larger than what will be downloaded. How will this help a user make a decision?

Thanks, -Stolee

Pablo SabaterSep 30, 2026, 18:31 UTC in reply to Junio C Hamano on lore

Re: [PATCH RFC 5/5] backfill: report total size of missing blobs in --dry-run

On Wed Sep 30, 2026 at 6:17 PM WEST, Junio C Hamano wrote:
Show 41 quoted lines
> Pablo Sabater <pabloosabaterr@gmail.com> writes:
>
>> +
>> +	if (!ctx->object_info_enabled)
>> +		goto cleanup;
>> +
>> +	if (!ctx->object_info_transport) {
>> +		struct promisor_remote *promise =
>> +			repo_promisor_remote_find(ctx->repo, NULL);
>> +		struct remote *remote = NULL;
>> +
>> +		if (!promise || !(remote = remote_get(promise->name)))
>> +			die(_("--dry-run requires a promisor remote"));
>> +
>> +		ctx->object_info_transport = transport_get(remote, NULL);
>> +
>> +		if (!ctx->object_info_transport->smart_options)
>> +			die(_("failed to get object info: smart options required"));
>> +	}
>> +
>> +	results->wants_size = 1;
>> +	status = transport_fetch_object_info(ctx->object_info_transport,
>> +					     &ctx->current_batch,
>> +					     results);
>> +
>> +	if (status == FETCH_OBJECT_INFO_NOT_ENABLED ||
>> +	    !results->sizes) {
>> +		ctx->object_info_enabled = 0;
>
> Yuck.
>
> Because we cannot tell if they allow you to look at the information,
> this cannot be helped, but it means anybody that looks at this
> ctx->object_info_enabled member to decide what to do must be careful.
>
> Is it guaranteed that results.sizes[] have been populated as long as
> status is not FETCH_OBJECT_INFO_NOT_ENABLED?  Can there be other
> errors that makes result.sizes[] unusable?  If that is the case,
> then it would be cleaner to have a dedicated ctx->sizes_valid member
> rather than relying on ctx->object_info_enabled member to carry this
> information ...

Not quite, transport_fetch_object_info() can also return FETCH_OBJECT_INFO_ERR when finish_connect() fails in fetch_object_info_via_pack(). I'll change the condition to:

	if (status != FETCH_OBJECT_INFO_OK || !results->sizes) {

With FETCH_OBJECT_INFO_OK, fetch_object_info() only allocates results->sizes when the server answers with the "size" attribute, and it die()s on any incomplete or malformed response, so a non-NULL sizes[] is always fully populated. A NULL sizes[] with FETCH_OBJECT_INFO_OK means object-info is available but the server does not support "size".

Agreed, the member is also cleared in that last case, where object-info does work but "size" is not advertised, so I get that it can be misleading. Since its only use is deciding whether to report the total size, I'll rename it to ctx->sizes_valid instead of adding a separate member.

Show 28 quoted lines
>
>> +		goto cleanup;
>> +	}
>> +
>> +	for (size_t i = 0; i < results->nr; i++)
>> +		ctx->total_batch_size += results->sizes[i];
>> +
>> +cleanup:
>> +	free_fetch_object_info_results(&ctx->object_info_results);
>>  	oid_array_clear(&ctx->current_batch);
>>  }
>>  
>> @@ -157,12 +201,25 @@ static int do_backfill(struct backfill_context *ctx)
>>  
>>  	dry_run_batch(ctx);
>>  
>> -	printf(Q_("After backfill, %" PRIuMAX " blob would be fetched.\n",
>> -		  "After backfill, %" PRIuMAX " blobs would be fetched.\n",
>> -		  (unsigned long)ctx->total_batch_nr),
>> -	       (uintmax_t)ctx->total_batch_nr);
>> +	if (ctx->object_info_enabled && ctx->total_batch_nr) {
>
> ... and use it here.  Within the design presented in this series, we
> know we have asked the other end at this point, and the above
> function may have turned ctx->object_info member off if the
> information is not there, so this may be safe.  But as I said, I am
> not sure what happens when fetch-object-info returned other kind of
> errors.

The condition above is checked on every batch, and any status other than FETCH_OBJECT_INFO_OK or a NULL sizes[] clears ctx->sizes_valid (currently ctx->object_info_enabled), so by the time we get here it is only set when all the batches returned valid sizes.

Pablo SabaterSep 30, 2026, 19:06 UTC in reply to Derrick Stolee on lore

Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)

On Wed Sep 30, 2026 at 7:14 PM WEST, Derrick Stolee wrote:
Show 19 quoted lines
> On 9/29/2026 8:21 PM, Pablo Sabater wrote:
>> [Cc'd Derrick Stolee for his work in the backfill(1) command]
>> 
>> This series adds a --dry-run option to git-backfill(1) that reports how
>> many missing blobs would be fetched and, when the remote server
>> supports the object-info capability, their total size:
>> 
>>         $ git backfill --dry-run
>>         After backfill, 48 blobs would be fetched (1.20 KiB).
>> 
>> If the server does not advertise object-info, only the count is shown.
>
> This is a helpful capability, but I'm not sure the size check counts
> as a "dry run" because it involves a network call (and possibly many
> depending on --min-batch-size).
>
> Perhaps a different argument would be better, such as --info=(count|size)
> to make it clear what level of information you want to know in advance
> and thus how much effort are you willing to put in to discover this. 
Makes sense to have it as an --info option.
Show 24 quoted lines
>> I am not a git-backfill(1) user myself, but it seemed useful for users
>> to know how much data a backfill would bring in before running it.
>
> I'm not sure that we want to add a feature based on speculation. Git
> is a collection of "itches" that the contributors needed scratched.
> The work is motivated by real needs.
>
> While I can see some benefit to curiosity, I'm not sure how much this
> would prevent users from making their decision as to whether they
> should run backfill or not.
>
>> The number of missing blobs is the sum of the number of blobs to be
>> fetched in each batch. The object-info capability lets us ask the server
>> for the size of each blob without downloading it, so summing them gives
>> an estimate of the total.
>> 
>> Note that this is an upper bound rather than the exact disk usage:
>> object-info reports the uncompressed size of each object, while the
>> objects end up stored compressed and possibly deltified in a packfile,
>> so the space actually used on disk will usually be smaller.
>
> I don't think the uncompressed size is a useful metric here, as it is
> likely astronomically larger than what will be downloaded. How will
> this help a user make a decision?

Yes, that's one of the itches I have with the object-info protocol: it cannot give you a reliable compressed size, and that's why only the total size is supported.

The object-info protocol could be extended to support the objectsize:disk attribute, either by having the server know what we already have and the oids that we want, or by directly having the server send us the compressed size of its local copy (to avoid too much work). Even with the first option, because it goes in batches, it would still be an estimate, just a closer one.

Given that I don't use backfill, I Cc'd you because I wasn't sure if it was really useful, and the main motivation was the "I'm going to check --dry-run before backfilling" case, so it helps to decide. If it's not that helpful and seems to end up as a decoration option, it might be better to drop it.

>
> Thanks,
> -Stolee

Thanks for taking a look, Pablo

Junio C HamanoSep 30, 2026, 20:01 UTC in reply to Pablo Sabater on lore

Re: [PATCH RFC 3/5] fetch-object-info: return a status instead of dying

"Pablo Sabater" <pabloosabaterr@gmail.com> writes:
Show 23 quoted lines
> On Wed Sep 30, 2026 at 6:07 PM WEST, Junio C Hamano wrote:
>> Pablo Sabater <pabloosabaterr@gmail.com> writes:
>>
>>> A subsequent commit needs fetch_object_info() not to die() when the
>>> object-info capability is not enabled on the server, so that it can
>>> fall back.
>>>
>>> Make fetch_object_info() return FETCH_OBJECT_INFO_NOT_ENABLED instead
>>> of die()'ing when the server does not advertise the object-info
>>> capability, and propagate the status through the transport layer so
>>> that callers of transport_fetch_object_info() can act on it. It is now
>>> up to them whether to die() or fall back.
>>
>> It may be just me but unless the client can tell between the server
>> not supporting (i.e., they are unable to enable it even if they
>> wanted to) and not enabling (i.e., they are capable, but are not
>> willing to give it to you), it may make sense to report it as "not
>> available".  "not enabled" sounds as if we know that it is the
>> latter and not the former.
>>
>> The code change looks very cleanly done.
>
> Makes sense, I'll rename it to FETCH_OBJECT_INFO_NOT_AVAILABLE.
Make it UNAVAILABLE instead.
Junio C HamanoSep 30, 2026, 20:03 UTC in reply to Derrick Stolee on lore

Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)

Derrick Stolee <stolee@gmail.com> writes:
Show 11 quoted lines
> This is a helpful capability, but I'm not sure the size check counts
> as a "dry run" because it involves a network call (and possibly many
> depending on --min-batch-size).
> ...
> I'm not sure that we want to add a feature based on speculation. Git
> is a collection of "itches" that the contributors needed scratched.
> The work is motivated by real needs.
> ...
> I don't think the uncompressed size is a useful metric here, as it is
> likely astronomically larger than what will be downloaded. How will
> this help a user make a decision?
I agree with you on all counts.

Back to recent threads