Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)
- From
Derrick Stolee <stolee@gmail.com>
- Date
- Sep 30, 2026, 18:14 UTC
- Message-ID
- <ed1b9048-d438-4143-a224-fa0e28d4fd42@gmail.com>
- In-Reply-To
- <20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com>
On 9/29/2026 8:21 PM, Pablo Sabater wrote:
Show 10 quoted lines
> [Cc'd Derrick Stolee for his work in the backfill(1) command] > > This series adds a --dry-run option to git-backfill(1) that reports how > many missing blobs would be fetched and, when the remote server > supports the object-info capability, their total size: > > $ git backfill --dry-run > After backfill, 48 blobs would be fetched (1.20 KiB). > > If the server does not advertise object-info, only the count is shown.
This is a helpful capability, but I'm not sure the size check counts as a "dry run" because it involves a network call (and possibly many depending on --min-batch-size).
Perhaps a different argument would be better, such as --info=(count|size) to make it clear what level of information you want to know in advance and thus how much effort are you willing to put in to discover this.
> I am not a git-backfill(1) user myself, but it seemed useful for users > to know how much data a backfill would bring in before running it.
I'm not sure that we want to add a feature based on speculation. Git is a collection of "itches" that the contributors needed scratched. The work is motivated by real needs.
While I can see some benefit to curiosity, I'm not sure how much this would prevent users from making their decision as to whether they should run backfill or not.
Show 9 quoted lines
> The number of missing blobs is the sum of the number of blobs to be > fetched in each batch. The object-info capability lets us ask the server > for the size of each blob without downloading it, so summing them gives > an estimate of the total. > > Note that this is an upper bound rather than the exact disk usage: > object-info reports the uncompressed size of each object, while the > objects end up stored compressed and possibly deltified in a packfile, > so the space actually used on disk will usually be smaller.
I don't think the uncompressed size is a useful metric here, as it is likely astronomically larger than what will be downloaded. How will this help a user make a decision?
Thanks, -Stolee