From: Derrick Stolee Date: Wed, 30 Sep 2026 18:14:46 GMT Subject: Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1) Message-ID: In-Reply-To: <20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com> On 9/29/2026 8:21 PM, Pablo Sabater wrote: > [Cc'd Derrick Stolee for his work in the backfill(1) command] > > This series adds a --dry-run option to git-backfill(1) that reports how > many missing blobs would be fetched and, when the remote server > supports the object-info capability, their total size: > > $ git backfill --dry-run > After backfill, 48 blobs would be fetched (1.20 KiB). > > If the server does not advertise object-info, only the count is shown. This is a helpful capability, but I'm not sure the size check counts as a "dry run" because it involves a network call (and possibly many depending on --min-batch-size). Perhaps a different argument would be better, such as --info=(count|size) to make it clear what level of information you want to know in advance and thus how much effort are you willing to put in to discover this. > I am not a git-backfill(1) user myself, but it seemed useful for users > to know how much data a backfill would bring in before running it. I'm not sure that we want to add a feature based on speculation. Git is a collection of "itches" that the contributors needed scratched. The work is motivated by real needs. While I can see some benefit to curiosity, I'm not sure how much this would prevent users from making their decision as to whether they should run backfill or not. > The number of missing blobs is the sum of the number of blobs to be > fetched in each batch. The object-info capability lets us ask the server > for the size of each blob without downloading it, so summing them gives > an estimate of the total. > > Note that this is an upper bound rather than the exact disk usage: > object-info reports the uncompressed size of each object, while the > objects end up stored compressed and possibly deltified in a packfile, > so the space actually used on disk will usually be smaller. I don't think the uncompressed size is a useful metric here, as it is likely astronomically larger than what will be downloaded. How will this help a user make a decision? Thanks, -Stolee