Re: [PATCH] fetch, clone: add fetch.blobSizeLimit config
- From
Jeff King <peff@peff.net>
- Date
- Mar 2, 2026, 18:28 UTC
- Message-ID
- <20260302182838.GI28275@coredump.intra.peff.net>
- In-Reply-To
- <aaV6PLJCrpb2mQnq@pks.im>
On Mon, Mar 02, 2026 at 12:53:32PM +0100, Patrick Steinhardt wrote:
Show 17 quoted lines
> On Sun, Mar 01, 2026 at 04:44:59PM +0000, Alan Braithwaite via GitGitGadget wrote: > > From: Alan Braithwaite <alan@braithwaite.dev> > > > > External tools like git-lfs and git-fat use the filter clean/smudge > > mechanism to manage large binary objects, but this requires pointer > > files, a separate storage backend, and careful coordination. Git's > > partial clone infrastructure provides a more native approach: large > > blobs can be excluded at the protocol level during fetch and lazily > > retrieved on demand. However, enabling this requires passing > > `--filter=blob:limit=<size>` on every clone, which is not > > discoverable and cannot be set as a global default. > > I'm not sure that we should make blob size limiting the default. The > problem with specifying a limit is that this is comparatively expensive > to compute on the server side: we have to look up each blob so that we > can determine its size. Unfortunately, such requests cannot (currently) > be optimized via for example bitmaps, or any other cache that we have.
We actually can do blob:limit filters with bitmaps. See 84243da129 (pack-bitmap: implement BLOB_LIMIT filtering, 2020-02-14). It's more expensive than blob:none, but not much. Once we have the list of blobs we can get their sizes directly from the packfile. It's stuff like path-limiting that is truly expensive, because it requires a traversal.
All that said, I'd be wary of turning on partial clones like this by default. I feel like there are still a lot of performance gotchas lurking (and possibly some correctness ones, too).
-Peff