Re: [PATCH v5] clone: add clone.<url>.defaultObjectFilter config
- From
- Alan Braithwaite <alan@braithwaite.dev>
- Date
- Mar 15, 2026, 01:33 UTC
- Message-ID
- <9b67801b-ce07-42b6-b2c6-2e7f0e5fd5f7@app.fastmail.com>
- In-Reply-To
- <abEdTQrRtAveH1rB@pks.im>
Thanks for the review, Patrick.
> `url_normalize()` will return a `NULL` pointer in case > it cannot parse the URL. We need to be prepared for > this, otherwise we might segfault.
Good catch. The updated patch guards on the return value and skips the urlmatch lookup entirely when the URL cannot be normalized. Today `match_urls()` happens to handle this safely (it returns 0 when `url->url` is NULL), but an explicit NULL check guards against future regressions in that code path.
> Do we want to "test_when_finished rm -rf > default-filter-clone" here and for all the subsequent > tests?
Done -- added `test_when_finished` cleanup to each test.
Patch incoming. :)
Thanks, - Alan
On Wed, Mar 11, 2026, at 00:44, Patrick Steinhardt wrote:
Show 126 quoted lines
> On Sat, Mar 07, 2026 at 01:33:56AM +0000, Alan Braithwaite via
> GitGitGadget wrote:
>> diff --git a/Documentation/config/clone.adoc b/Documentation/config/clone.adoc
>> index 0a10efd174..1d6c0957a0 100644
>> --- a/Documentation/config/clone.adoc
>> +++ b/Documentation/config/clone.adoc
>> @@ -21,3 +21,37 @@ endif::[]
>> If a partial clone filter is provided (see `--filter` in
>> linkgit:git-rev-list[1]) and `--recurse-submodules` is used, also apply
>> the filter to submodules.
>> +
>> +`clone.defaultObjectFilter`::
>> +`clone.<url>.defaultObjectFilter`::
>> + When set to a filter spec string (e.g., `blob:limit=1m`,
>> + `blob:none`, `tree:0`), linkgit:git-clone[1] will automatically
>> + use `--filter=<value>` to enable partial clone behavior.
>> + Objects matching the filter are excluded from the initial
>> + transfer and lazily fetched on demand (e.g., during checkout).
>> + Subsequent fetches inherit the filter via the per-remote config
>> + that is written during the clone.
>> ++
>> +The bare `clone.defaultObjectFilter` applies to all clones. The
>> +URL-qualified form `clone.<url>.defaultObjectFilter` restricts the
>> +setting to clones whose URL matches `<url>`, following the same
>> +rules as `http.<url>.*` (see linkgit:git-config[1]). The most
>> +specific URL match wins. You can match a domain, a namespace, or a
>> +specific project:
>> ++
>> +----
>> +[clone]
>> + defaultObjectFilter = blob:limit=1m
>> +
>> +[clone "https://github.com/"]
>> + defaultObjectFilter = blob:limit=5m
>> +
>> +[clone "https://internal.corp.com/large-project/"]
>> + defaultObjectFilter = blob:none
>> +----
>> ++
>> +An explicit `--filter` option on the command line takes precedence
>> +over this config, and `--no-filter` defeats it entirely to force a
>> +full clone. Only affects the initial clone; it has no effect on
>> +later fetches into an existing repository. If the server does not
>> +support object filtering, the setting is silently ignored.
>
> This all reads good to me.
>
>> diff --git a/builtin/clone.c b/builtin/clone.c
>> index 45d8fa0eed..1207655815 100644
>> --- a/builtin/clone.c
>> +++ b/builtin/clone.c
>> @@ -757,6 +758,47 @@ static int git_clone_config(const char *k, const char *v,
>> return git_default_config(k, v, ctx, cb);
>> }
>>
>> +static int clone_filter_collect(const char *var, const char *value,
>> + const struct config_context *ctx UNUSED,
>> + void *cb)
>> +{
>> + char **filter_spec_p = cb;
>> +
>> + if (!strcmp(var, "clone.defaultobjectfilter")) {
>> + if (!value)
>> + return config_error_nonbool(var);
>> + free(*filter_spec_p);
>> + *filter_spec_p = xstrdup(value);
>> + }
>> + return 0;
>> +}
>> +
>> +/*
>> + * Look up clone.defaultObjectFilter or clone.<url>.defaultObjectFilter
>> + * using the urlmatch infrastructure. A URL-qualified entry that matches
>> + * the clone URL takes precedence over the bare form, following the same
>> + * rules as http.<url>.* configuration variables.
>> + */
>> +static char *get_default_object_filter(const char *url)
>> +{
>> + struct urlmatch_config config = URLMATCH_CONFIG_INIT;
>> + char *filter_spec = NULL;
>> + char *normalized_url;
>> +
>> + config.section = "clone";
>> + config.key = "defaultobjectfilter";
>> + config.collect_fn = clone_filter_collect;
>> + config.cb = &filter_spec;
>> +
>> + normalized_url = url_normalize(url, &config.url);
>
> `url_normalize()` will return a `NULL` pointer in case it cannot parse
> the URL. We need to be prepared for this, otherwise we might segfault.
> I guess the best route is to simply ignore the URL in that case --
> otherwise, we would always error out in case the remote has a weird URL
> configured.
>
>> diff --git a/t/t5616-partial-clone.sh b/t/t5616-partial-clone.sh
>> index 1e354e057f..1254901f3e 100755
>> --- a/t/t5616-partial-clone.sh
>> +++ b/t/t5616-partial-clone.sh
>> @@ -722,6 +722,124 @@ test_expect_success 'after fetching descendants of non-promisor commits, gc work
>> git -C partial gc --prune=now
>> '
>>
>> +# Test clone.<url>.defaultObjectFilter config
>> +
>> +test_expect_success 'setup for clone.defaultObjectFilter tests' '
>> + git init default-filter-src &&
>> + echo "small" >default-filter-src/small.txt &&
>> + dd if=/dev/zero of=default-filter-src/large.bin bs=1024 count=100 2>/dev/null &&
>> + git -C default-filter-src add . &&
>> + git -C default-filter-src commit -m "initial" &&
>> +
>> + git clone --bare "file://$(pwd)/default-filter-src" default-filter-srv.bare &&
>> + git -C default-filter-srv.bare config --local uploadpack.allowfilter 1 &&
>> + git -C default-filter-srv.bare config --local uploadpack.allowanysha1inwant 1
>> +'
>> +
>> +test_expect_success 'clone with clone.<url>.defaultObjectFilter applies filter' '
>> + SERVER_URL="file://$(pwd)/default-filter-srv.bare" &&
>> + git -c "clone.$SERVER_URL.defaultObjectFilter=blob:limit=1k" clone \
>> + "$SERVER_URL" default-filter-clone &&
>
> Do we want to "test_when_finished rm -rf default-filter-clone" here and
> for all the subsequent tests?
>
> Patrick