From: Jeff King Date: Mon, 02 Mar 2026 18:28:38 GMT Subject: Re: [PATCH] fetch, clone: add fetch.blobSizeLimit config Message-ID: <20260302182838.GI28275@coredump.intra.peff.net> In-Reply-To: On Mon, Mar 02, 2026 at 12:53:32PM +0100, Patrick Steinhardt wrote: > On Sun, Mar 01, 2026 at 04:44:59PM +0000, Alan Braithwaite via GitGitGadget wrote: > > From: Alan Braithwaite > > > > External tools like git-lfs and git-fat use the filter clean/smudge > > mechanism to manage large binary objects, but this requires pointer > > files, a separate storage backend, and careful coordination. Git's > > partial clone infrastructure provides a more native approach: large > > blobs can be excluded at the protocol level during fetch and lazily > > retrieved on demand. However, enabling this requires passing > > `--filter=blob:limit=` on every clone, which is not > > discoverable and cannot be set as a global default. > > I'm not sure that we should make blob size limiting the default. The > problem with specifying a limit is that this is comparatively expensive > to compute on the server side: we have to look up each blob so that we > can determine its size. Unfortunately, such requests cannot (currently) > be optimized via for example bitmaps, or any other cache that we have. We actually can do blob:limit filters with bitmaps. See 84243da129 (pack-bitmap: implement BLOB_LIMIT filtering, 2020-02-14). It's more expensive than blob:none, but not much. Once we have the list of blobs we can get their sizes directly from the packfile. It's stuff like path-limiting that is truly expensive, because it requires a traversal. All that said, I'd be wary of turning on partial clones like this by default. I feel like there are still a lot of performance gotchas lurking (and possibly some correctness ones, too). -Peff