Volume XXII, number 279Tuesday, October 6, 2026Latest message 10 minutes ago

The Git List

News and archive of git@vger.kernel.org, since April 2005

patchodb/files: be less aggressive with geometric repacking

8 messages between Aug 11, 2026 and Aug 21, 2026, from Patrick Steinhardt, Justin Tobler, Elijah Newren.

Plain Markdown or JSON for tools and agents. Diffs are folded; open one to read it.

Patrick SteinhardtAug 11, 2026, 09:04 UTC on lore

When performing auto-maintenance with geometric repacking we have two conditions that may trigger a repack:

  - Either the geometric sequence of packfiles is invalidated.
  - Or we have too many loose objects.

The first condition shouldn't trigger all that often: it may be hit when we fetch a new packfile, but users tend to not do that all the time. The second condition is what typically triggers more regularly though, as every command that ends up writing new objects may cause us to cross the threshold of loose objects. It is thus preferable to not be too aggressive here, as otherwise we may end up repacking objects quite often.

For the geometric-repacking strategy though we have a default of 100 objects, only. As we're approximating the count of objects by only reading the "objects/17/" shared, we'd only need 2 objects in there before we perform a repack by default, which is quite aggressive. git-gc(1) on the other hand has a default of 6700, so it is quite a bit more conservative here.

Being this aggressive is also causing problems as reported by our users. When running lots of concurrent writers, those writes will constantly end up spawning maintenance jobs that end up repacking objects. As we also prune objects, a concurrently running process that tries to write an object may see that the sharding directories get removed under their feet. While we try re-creating such leading directories, we only do so a single time, and it may happen that the directory vanishes again before we had the chance to create the loose object. This is not a new problem, but it is exacerbated by us running maintenance this aggressively.

Improve the status quo by reducing the frequency at which we pack loose objects to the same frequency that git-gc(1) uses.

Reported-by: Stefan Haller <lists@haller-berlin.de>
Signed-off-by: Patrick Steinhardt <ps@pks.im>
---
Hi,
as reported by Stefan at [1]. Thanks!
Patrick
[1]: <4f6a96ac-d993-4872-b3c4-30d899f61ca9@haller-berlin.de>
---
 Documentation/config/maintenance.adoc | 2 +-
 odb/source-files.c                    | 2 +-
 2 files changed, 2 insertions(+), 2 deletions(-)
Show changes to 2 files +2 −2

Documentation/config/maintenance.adoc, odb/source-files.c

diff --git a/Documentation/config/maintenance.adoc b/Documentation/config/maintenance.adoc
index b578856dde..da8be9f812 100644
--- a/Documentation/config/maintenance.adoc
+++ b/Documentation/config/maintenance.adoc
@@ -101,7 +101,7 @@ maintenance.geometric-repack.auto::
 	there are packfiles that need to be merged together to retain the
 	geometric progression, or when there are at least this many loose
 	objects that would be written into a new packfile. The default value is
-	100.
+	6700.
 
 maintenance.geometric-repack.splitFactor::
 	This integer config option controls the factor used for the geometric
diff --git a/odb/source-files.c b/odb/source-files.c
index 5a68af7d84..555e466145 100644
--- a/odb/source-files.c
+++ b/odb/source-files.c
@@ -521,7 +521,7 @@ bool odb_source_files_optimize_required(struct odb_source *source,
 		};
 		struct existing_packs existing_packs = EXISTING_PACKS_INIT;
 		struct string_list kept_packs = STRING_LIST_INIT_DUP;
-		int auto_value = 100;
+		int auto_value = 6700;
 		bool ret;
 
 		repo_config_get_int(repo, "maintenance.geometric-repack.auto",

---
base-commit: 010afd3166ddc64c9863b1506f12cbcdda0d4ea1
change-id: 20260810-pks-geometric-maintenance-reduce-frequency-5c1c9423ceb3
Justin ToblerAug 11, 2026, 20:44 UTC in reply to Patrick Steinhardt on lore

Re: [PATCH] odb/files: be less aggressive with geometric repacking

On 26/08/11 11:04AM, Patrick Steinhardt wrote:
Show 21 quoted lines
> When performing auto-maintenance with geometric repacking we have two
> conditions that may trigger a repack:
> 
>   - Either the geometric sequence of packfiles is invalidated.
> 
>   - Or we have too many loose objects.
> 
> The first condition shouldn't trigger all that often: it may be hit when
> we fetch a new packfile, but users tend to not do that all the time. The
> second condition is what typically triggers more regularly though, as
> every command that ends up writing new objects may cause us to cross the
> threshold of loose objects. It is thus preferable to not be too
> aggressive here, as otherwise we may end up repacking objects quite
> often.
> 
> For the geometric-repacking strategy though we have a default of 100
> objects, only. As we're approximating the count of objects by only
> reading the "objects/17/" shared, we'd only need 2 objects in there
> before we perform a repack by default, which is quite aggressive.
> git-gc(1) on the other hand has a default of 6700, so it is quite a bit
> more conservative here.

Ok IIUC, the reason two loose objects can potentially trigger repacking is because the heuristic used to estimate the number of loose objects only counts objects in "objects/17/" and multiples it by 256 (the maximum number of directories that are fanned-out). That makes sense and indeed seems like it could lead to repacking processes be spawned more frequently than desired.

My first thought is whether the heuristic itself should be updated to capture a more accurate estimate for the number of objects. That would of course require looking up more objects and thus be more expensive. If the goal here is just for a very rough estimate anyways, maybe it wouldn't be worth it though.

Increasing the loose object threshold here to be more conservative seems like a reasonable approach. I'm not sure exactly why 6700 was chosen here. 6700 / 256 ~= 26.2 which means "objects/17/" would have to contain at least 27 objects before repacking is triggered. That is certainly much more conservative. I see that 6700 has also been chosen else where in the codebase as the threshold too. It might be nice to explain the reasoning a bit more in the commit message though.

Show 12 quoted lines
> Being this aggressive is also causing problems as reported by our users.
> When running lots of concurrent writers, those writes will constantly
> end up spawning maintenance jobs that end up repacking objects. As we
> also prune objects, a concurrently running process that tries to write
> an object may see that the sharding directories get removed under their
> feet. While we try re-creating such leading directories, we only do so a
> single time, and it may happen that the directory vanishes again before
> we had the chance to create the loose object. This is not a new problem,
> but it is exacerbated by us running maintenance this aggressively.
> 
> Improve the status quo by reducing the frequency at which we pack loose
> objects to the same frequency that git-gc(1) uses.
Makes sense and the patch itself looks trivially correct.
-Justin
Patrick SteinhardtAug 12, 2026, 05:44 UTC in reply to Justin Tobler on lore

Re: [PATCH] odb/files: be less aggressive with geometric repacking

On Tue, Aug 11, 2026 at 03:44:12PM -0500, Justin Tobler wrote:
Show 35 quoted lines
> On 26/08/11 11:04AM, Patrick Steinhardt wrote:
> > When performing auto-maintenance with geometric repacking we have two
> > conditions that may trigger a repack:
> > 
> >   - Either the geometric sequence of packfiles is invalidated.
> > 
> >   - Or we have too many loose objects.
> > 
> > The first condition shouldn't trigger all that often: it may be hit when
> > we fetch a new packfile, but users tend to not do that all the time. The
> > second condition is what typically triggers more regularly though, as
> > every command that ends up writing new objects may cause us to cross the
> > threshold of loose objects. It is thus preferable to not be too
> > aggressive here, as otherwise we may end up repacking objects quite
> > often.
> > 
> > For the geometric-repacking strategy though we have a default of 100
> > objects, only. As we're approximating the count of objects by only
> > reading the "objects/17/" shared, we'd only need 2 objects in there
> > before we perform a repack by default, which is quite aggressive.
> > git-gc(1) on the other hand has a default of 6700, so it is quite a bit
> > more conservative here.
> 
> Ok IIUC, the reason two loose objects can potentially trigger repacking
> is because the heuristic used to estimate the number of loose objects
> only counts objects in "objects/17/" and multiples it by 256 (the
> maximum number of directories that are fanned-out). That makes sense and
> indeed seems like it could lead to repacking processes be spawned more
> frequently than desired.
> 
> My first thought is whether the heuristic itself should be updated to
> capture a more accurate estimate for the number of objects. That would
> of course require looking up more objects and thus be more expensive. If
> the goal here is just for a very rough estimate anyways, maybe it
> wouldn't be worth it though.

That wouldn't really solve the problem though. The problem is not really that the estimation can be wrong, it's rather that even if it was always correct we're still being too aggressive with packing the loose objects. Because ultimately, a 100 objects is a comparatively small threshold, and leads to 67 times more repacking compared to git-gc(1).

Show 7 quoted lines
> Increasing the loose object threshold here to be more conservative seems
> like a reasonable approach. I'm not sure exactly why 6700 was chosen
> here. 6700 / 256 ~= 26.2 which means "objects/17/" would have to contain
> at least 27 objects before repacking is triggered. That is certainly
> much more conservative. I see that 6700 has also been chosen else where
> in the codebase as the threshold too. It might be nice to explain the
> reasoning a bit more in the commit message though.

Hmm, don't I already do that? In the paragraph you're responding to I'm saying that git-gc(1) already had that default forever, so I'm adjusting our heuristic to match that.

Thanks!
Patrick
Justin ToblerAug 18, 2026, 22:34 UTC in reply to Patrick Steinhardt on lore

Re: [PATCH] odb/files: be less aggressive with geometric repacking

On 26/08/12 07:44AM, Patrick Steinhardt wrote:
Show 42 quoted lines
> On Tue, Aug 11, 2026 at 03:44:12PM -0500, Justin Tobler wrote:
> > On 26/08/11 11:04AM, Patrick Steinhardt wrote:
> > > When performing auto-maintenance with geometric repacking we have two
> > > conditions that may trigger a repack:
> > > 
> > >   - Either the geometric sequence of packfiles is invalidated.
> > > 
> > >   - Or we have too many loose objects.
> > > 
> > > The first condition shouldn't trigger all that often: it may be hit when
> > > we fetch a new packfile, but users tend to not do that all the time. The
> > > second condition is what typically triggers more regularly though, as
> > > every command that ends up writing new objects may cause us to cross the
> > > threshold of loose objects. It is thus preferable to not be too
> > > aggressive here, as otherwise we may end up repacking objects quite
> > > often.
> > > 
> > > For the geometric-repacking strategy though we have a default of 100
> > > objects, only. As we're approximating the count of objects by only
> > > reading the "objects/17/" shared, we'd only need 2 objects in there
> > > before we perform a repack by default, which is quite aggressive.
> > > git-gc(1) on the other hand has a default of 6700, so it is quite a bit
> > > more conservative here.
> > 
> > Ok IIUC, the reason two loose objects can potentially trigger repacking
> > is because the heuristic used to estimate the number of loose objects
> > only counts objects in "objects/17/" and multiples it by 256 (the
> > maximum number of directories that are fanned-out). That makes sense and
> > indeed seems like it could lead to repacking processes be spawned more
> > frequently than desired.
> > 
> > My first thought is whether the heuristic itself should be updated to
> > capture a more accurate estimate for the number of objects. That would
> > of course require looking up more objects and thus be more expensive. If
> > the goal here is just for a very rough estimate anyways, maybe it
> > wouldn't be worth it though.
> 
> That wouldn't really solve the problem though. The problem is not really
> that the estimation can be wrong, it's rather that even if it was always
> correct we're still being too aggressive with packing the loose objects.
> Because ultimately, a 100 objects is a comparatively small threshold,
> and leads to 67 times more repacking compared to git-gc(1).
Ok, that makes sense.
Show 11 quoted lines
> > Increasing the loose object threshold here to be more conservative seems
> > like a reasonable approach. I'm not sure exactly why 6700 was chosen
> > here. 6700 / 256 ~= 26.2 which means "objects/17/" would have to contain
> > at least 27 objects before repacking is triggered. That is certainly
> > much more conservative. I see that 6700 has also been chosen else where
> > in the codebase as the threshold too. It might be nice to explain the
> > reasoning a bit more in the commit message though.
> 
> Hmm, don't I already do that? In the paragraph you're responding to I'm
> saying that git-gc(1) already had that default forever, so I'm adjusting
> our heuristic to match that.

I think I was just curious as to why 6700 was the chosen number for git-gc(1) as well, but its probably just good to be consistent here. I think this patch is fine as is.

-Justin
Patrick SteinhardtAug 19, 2026, 04:52 UTC in reply to Justin Tobler on lore

Re: [PATCH] odb/files: be less aggressive with geometric repacking

On Tue, Aug 18, 2026 at 05:34:54PM -0500, Justin Tobler wrote:
Show 17 quoted lines
> On 26/08/12 07:44AM, Patrick Steinhardt wrote:
> > On Tue, Aug 11, 2026 at 03:44:12PM -0500, Justin Tobler wrote:
> > > Increasing the loose object threshold here to be more conservative seems
> > > like a reasonable approach. I'm not sure exactly why 6700 was chosen
> > > here. 6700 / 256 ~= 26.2 which means "objects/17/" would have to contain
> > > at least 27 objects before repacking is triggered. That is certainly
> > > much more conservative. I see that 6700 has also been chosen else where
> > > in the codebase as the threshold too. It might be nice to explain the
> > > reasoning a bit more in the commit message though.
> > 
> > Hmm, don't I already do that? In the paragraph you're responding to I'm
> > saying that git-gc(1) already had that default forever, so I'm adjusting
> > our heuristic to match that.
> 
> I think I was just curious as to why 6700 was the chosen number for
> git-gc(1) as well, but its probably just good to be consistent here. I
> think this patch is fine as is.

That's a good question. It has been introduced all the way back in 2c3c439947 (Implement git gc --auto, 2007-09-05), but that commit does not mention any reasoning for the 6700 limit either.

Digging in history a bit surfaces this nugget [1]. So the limit was chosen so that git-gc(1) would not trigger for a fully unpacked Git v0.99, would trigger for v1.0, but not triggering when doing an incremental gc after going from v0.99 to v1.0. This is of course quite arbitrary, but as the mail points out, "[t]he default threshold is arbitrarily set by yours truly" (Junio).

Patrick
[1]: https://lore.kernel.org/git/7vr6lcj2zi.fsf@gitster.siamese.dyndns.org/
Patrick SteinhardtAug 21, 2026, 06:34 UTC in reply to Patrick Steinhardt on lore

Re: [PATCH] odb/files: be less aggressive with geometric repacking

Cc'ing Junio, as I haven't seen this topic in "What's cooking" yet. I assume it must've fallen through the cracks. Thanks!

Patrick
Elijah NewrenAug 21, 2026, 06:40 UTC in reply to Patrick Steinhardt on lore

Re: [PATCH] odb/files: be less aggressive with geometric repacking

Heh, looks like a typed up a response and got distracted just before the end and never came back and sent it. Sending now due to Patrick's ping about this not showing up in What's Cooking; maybe an extra review will help. :-)

On Tue, Aug 11, 2026 at 2:17 AM Patrick Steinhardt <ps@pks.im> wrote:
Show 22 quoted lines
>
> When performing auto-maintenance with geometric repacking we have two
> conditions that may trigger a repack:
>
>   - Either the geometric sequence of packfiles is invalidated.
>
>   - Or we have too many loose objects.
>
> The first condition shouldn't trigger all that often: it may be hit when
> we fetch a new packfile, but users tend to not do that all the time. The
> second condition is what typically triggers more regularly though, as
> every command that ends up writing new objects may cause us to cross the
> threshold of loose objects. It is thus preferable to not be too
> aggressive here, as otherwise we may end up repacking objects quite
> often.
>
> For the geometric-repacking strategy though we have a default of 100
> objects, only. As we're approximating the count of objects by only
> reading the "objects/17/" shared, we'd only need 2 objects in there
> before we perform a repack by default, which is quite aggressive.
> git-gc(1) on the other hand has a default of 6700, so it is quite a bit
> more conservative here.

2? Wouldn't you only need 1 (or if you could have fractional numbers of objects, only 0.390625 of them)? <looks around...> Oh, huh:

        /*
         * This is weird, but stems from legacy behaviour: the GC auto
         * threshold was always essentially interpreted as if it was rounded up
         * to the next multiple 256 of, so we retain this behaviour for now.
         */
        return loose_count > (DIV_ROUND_UP(((unsigned long) limit), 256) * 256);
So, indeed, you need 2.
Show 9 quoted lines
> Being this aggressive is also causing problems as reported by our users.
> When running lots of concurrent writers, those writes will constantly
> end up spawning maintenance jobs that end up repacking objects. As we
> also prune objects, a concurrently running process that tries to write
> an object may see that the sharding directories get removed under their
> feet. While we try re-creating such leading directories, we only do so a
> single time, and it may happen that the directory vanishes again before
> we had the chance to create the loose object. This is not a new problem,
> but it is exacerbated by us running maintenance this aggressively.

Unrelated to this patch...but should git avoid pruning the loose object sharding directories?

> Improve the status quo by reducing the frequency at which we pack loose
> objects to the same frequency that git-gc(1) uses.
Makes sense.
Show 45 quoted lines
> Reported-by: Stefan Haller <lists@haller-berlin.de>
> Signed-off-by: Patrick Steinhardt <ps@pks.im>
> ---
> Hi,
>
> as reported by Stefan at [1]. Thanks!
>
> Patrick
>
> [1]: <4f6a96ac-d993-4872-b3c4-30d899f61ca9@haller-berlin.de>
> ---
>  Documentation/config/maintenance.adoc | 2 +-
>  odb/source-files.c                    | 2 +-
>  2 files changed, 2 insertions(+), 2 deletions(-)
>
> diff --git a/Documentation/config/maintenance.adoc b/Documentation/config/maintenance.adoc
> index b578856dde..da8be9f812 100644
> --- a/Documentation/config/maintenance.adoc
> +++ b/Documentation/config/maintenance.adoc
> @@ -101,7 +101,7 @@ maintenance.geometric-repack.auto::
>         there are packfiles that need to be merged together to retain the
>         geometric progression, or when there are at least this many loose
>         objects that would be written into a new packfile. The default value is
> -       100.
> +       6700.
>
>  maintenance.geometric-repack.splitFactor::
>         This integer config option controls the factor used for the geometric
> diff --git a/odb/source-files.c b/odb/source-files.c
> index 5a68af7d84..555e466145 100644
> --- a/odb/source-files.c
> +++ b/odb/source-files.c
> @@ -521,7 +521,7 @@ bool odb_source_files_optimize_required(struct odb_source *source,
>                 };
>                 struct existing_packs existing_packs = EXISTING_PACKS_INIT;
>                 struct string_list kept_packs = STRING_LIST_INIT_DUP;
> -               int auto_value = 100;
> +               int auto_value = 6700;
>                 bool ret;
>
>                 repo_config_get_int(repo, "maintenance.geometric-repack.auto",
>
> ---
> base-commit: 010afd3166ddc64c9863b1506f12cbcdda0d4ea1
> change-id: 20260810-pks-geometric-maintenance-reduce-frequency-5c1c9423ceb3
Looks good to me.
Patrick SteinhardtAug 21, 2026, 11:40 UTC in reply to Elijah Newren on lore

Re: [PATCH] odb/files: be less aggressive with geometric repacking

On Thu, Aug 20, 2026 at 11:40:46PM -0700, Elijah Newren wrote:
> On Tue, Aug 11, 2026 at 2:17 AM Patrick Steinhardt <ps@pks.im> wrote:
[snip]
Show 12 quoted lines
> > Being this aggressive is also causing problems as reported by our users.
> > When running lots of concurrent writers, those writes will constantly
> > end up spawning maintenance jobs that end up repacking objects. As we
> > also prune objects, a concurrently running process that tries to write
> > an object may see that the sharding directories get removed under their
> > feet. While we try re-creating such leading directories, we only do so a
> > single time, and it may happen that the directory vanishes again before
> > we had the chance to create the loose object. This is not a new problem,
> > but it is exacerbated by us running maintenance this aggressively.
> 
> Unrelated to this patch...but should git avoid pruning the loose
> object sharding directories?

I was wondering about that, too. There are two contradicting arguments to make here:

  - Pruning the sharding directories allows us to quickly determine that
    an empty shard cannot have an object.
  - Not pruning the sharding directories may avoid a lot of write churn.

The question is how large the impact of these two individual arguments is.

By gut feeling, I think that the first argument is somewhat weak. Not having empty directories means that looking up a loose object by its path will be slightly faster because we have to walk one less directory in the hierarchy. But this really only matters in the case where we look for a nonexistent object, which does not happen all that often because we prefer searching packfiles first.

Furthermore, iterating through all objects in the object database will be faster, as we don't have to open each of the directories only to find them empty. But again, that's not really something that we do all that frequently.

On the other hand, we _do_ have to recreate the loose object shards quite frequently as that's how we write data into a repository. And as we've seen, pruning those shards can easily cause races.

So in the end I think it could be a useful thing to explore. The only thing I wonder is whether there's a good reason for why we prune those that I miss.

Patrick

Back to recent threads