git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH RFC 0/5] Add --dry-run option to git-backfill(1)

From
Derrick Stolee <stolee@gmail.com>
Date
Sep 30, 2026, 18:14 UTC
Message-ID
<ed1b9048-d438-4143-a224-fa0e28d4fd42@gmail.com>
In-Reply-To
<20260930-backfill-dryrun-v1-0-1128f247ee01@gmail.com>
On 9/29/2026 8:21 PM, Pablo Sabater wrote:
Show 10 quoted lines
> [Cc'd Derrick Stolee for his work in the backfill(1) command]
> 
> This series adds a --dry-run option to git-backfill(1) that reports how
> many missing blobs would be fetched and, when the remote server
> supports the object-info capability, their total size:
> 
>         $ git backfill --dry-run
>         After backfill, 48 blobs would be fetched (1.20 KiB).
> 
> If the server does not advertise object-info, only the count is shown.

This is a helpful capability, but I'm not sure the size check counts as a "dry run" because it involves a network call (and possibly many depending on --min-batch-size).

Perhaps a different argument would be better, such as --info=(count|size) to make it clear what level of information you want to know in advance and thus how much effort are you willing to put in to discover this.

> I am not a git-backfill(1) user myself, but it seemed useful for users
> to know how much data a backfill would bring in before running it.

I'm not sure that we want to add a feature based on speculation. Git is a collection of "itches" that the contributors needed scratched. The work is motivated by real needs.

While I can see some benefit to curiosity, I'm not sure how much this would prevent users from making their decision as to whether they should run backfill or not.

Show 9 quoted lines
> The number of missing blobs is the sum of the number of blobs to be
> fetched in each batch. The object-info capability lets us ask the server
> for the size of each blob without downloading it, so summing them gives
> an estimate of the total.
> 
> Note that this is an upper bound rather than the exact disk usage:
> object-info reports the uncompressed size of each object, while the
> objects end up stored compressed and possibly deltified in a packfile,
> so the space actually used on disk will usually be smaller.

I don't think the uncompressed size is a useful metric here, as it is likely astronomically larger than what will be downloaded. How will this help a user make a decision?

Thanks, -Stolee

Previous: Pablo SabaterNext: Pablo Sabater
Message 16 of 18 in “Add --dry-run option to git-backfill(1)”
  1. 0/5 Add --dry-run option to git-backfill(1)Pablo Sabater, Sep 30, 2026
  2. 1/5 transport-internal: update fetch_object_info commentPablo Sabater, Sep 30, 2026
  3. 2/5 fetch-object-info: add enum for fetch_object_info() statusesPablo Sabater, Sep 30, 2026
  4. Karthik NayakSep 30, 2026
  5. Pablo SabaterSep 30, 2026
  6. 3/5 fetch-object-info: return a status instead of dyingPablo Sabater, Sep 30, 2026
  7. Junio C HamanoSep 30, 2026
  8. Pablo SabaterSep 30, 2026
  9. Junio C HamanoSep 30, 2026
  10. 4/5 backfill: add --dry-run optionPablo Sabater, Sep 30, 2026
  11. Karthik NayakSep 30, 2026
  12. Pablo SabaterSep 30, 2026
  13. 5/5 backfill: report total size of missing blobs in --dry-runPablo Sabater, Sep 30, 2026
  14. Junio C HamanoSep 30, 2026
  15. Pablo SabaterSep 30, 2026
  16. Derrick StoleeSep 30, 2026
  17. Pablo SabaterSep 30, 2026
  18. Junio C HamanoSep 30, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.