Re: [PATCH] fetch, clone: add fetch.blobSizeLimit config
- From
Patrick Steinhardt <ps@pks.im>
- Date
- Mar 2, 2026, 11:53 UTC
- Message-ID
- <aaV6PLJCrpb2mQnq@pks.im>
- In-Reply-To
- <pull.2058.git.1772383499900.gitgitgadget@gmail.com>
On Sun, Mar 01, 2026 at 04:44:59PM +0000, Alan Braithwaite via GitGitGadget wrote:
Show 10 quoted lines
> From: Alan Braithwaite <alan@braithwaite.dev> > > External tools like git-lfs and git-fat use the filter clean/smudge > mechanism to manage large binary objects, but this requires pointer > files, a separate storage backend, and careful coordination. Git's > partial clone infrastructure provides a more native approach: large > blobs can be excluded at the protocol level during fetch and lazily > retrieved on demand. However, enabling this requires passing > `--filter=blob:limit=<size>` on every clone, which is not > discoverable and cannot be set as a global default.
I'm not sure that we should make blob size limiting the default. The problem with specifying a limit is that this is comparatively expensive to compute on the server side: we have to look up each blob so that we can determine its size. Unfortunately, such requests cannot (currently) be optimized via for example bitmaps, or any other cache that we have.
So if we want to make any filter the default, I'd propose that we should rather think about filters that are computationally less expensive, like for example `--filter=blob:none`. This can be computed efficiently via bitmaps.
The downside is of course that in this case we have to do way more backfill fetches compared to the case where we only leave out a couple of blobs. But unless we figure out a way to serve the size limit filter in a more efficient way I'm not sure about proper alternatives.
Another question to consider: is it really sensible to set this setting globally? It is very much dependent on the forge that you're connecting to, as forges may not even allow object filters at all, or only a subset of them.
Thanks!
Patrick