git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC/PATCH] gc: run more pre-detach operations under lock

From
Ævar Arnfjörð Bjarmason <avarab@gmail.com>
Date
Jun 20, 2019, 21:49 UTC
Message-ID
<87ef3o7ws1.fsf@evledraar.gmail.com>
In-Reply-To
<CACsJy8AjXXOpcKrSV4z6kEM=eyFDWSyf==tJZzvDyEN591XdGw@mail.gmail.com>
On Thu, Jun 20 2019, Duy Nguyen wrote:
Show 58 quoted lines
> On Thu, Jun 20, 2019 at 5:49 AM Ævar Arnfjörð Bjarmason
> <avarab@gmail.com> wrote:
>>
>>
>> On Wed, Jun 19 2019, Jeff King wrote:
>>
>> > On Wed, Jun 19, 2019 at 08:01:55PM +0200, Ævar Arnfjörð Bjarmason wrote:
>> >
>> >> > You could sort of avoid the problem here too with
>> >> >
>> >> > parallel 'git fetch --no-auto-gc {}' ::: $(git remote)
>> >> > git gc --auto
>> >> >
>> >> > It's definitely simpler, but of course we have to manually add
>> >> > --no-auto-gc in everywhere we need, so not quite as elegant.
>> >> >
>> >> > Actually you could already do that with 'git -c gc.auto=false fetch', I guess.
>> >>
>> >> The point of the 'parallel' example is to show disconnected git
>> >> commands, think trying to run 'git' in a terminal while your editor
>> >> asynchronously runs a polling 'fetch', or a server with multiple
>> >> concurrent clients running 'gc --auto'.
>> >>
>> >> That's the question my RFC patch raises. As far as I can tell the
>> >> approach in your patch is only needed because our locking for gc is
>> >> buggy, rather than introduce the caveat that an fetch(N) operation won't
>> >> do "gc" until it's finished (we may have hundreds, thousands of remotes,
>> >> I use that for some more obscure use-cases) shouldn't we just fix the
>> >> locking?
>> >
>> > I think there may be room for both approaches. Yours fixes the repeated
>> > message in the more general case, but Duy's suggestion is the most
>> > efficient thing.
>> >
>> > I agree that the "thousands of remotes" case means we might want to gc
>> > in the interim. But we probably ought to do that deterministically
>> > rather than hoping that the pattern of lock contention makes sense.
>>
>> We do it deterministically, when gc.auto thresholds et al are exceeded
>> we kick one off without waiting for other stuff, if we can get the lock.
>>
>> I don't think this desire to just wait a bit until all the fetches are
>> complete makes sense as a special-case.
>>
>> If, as you noted in <20190619190845.GD28145@sigill.intra.peff.net>, the
>> desire is to reduce GC CPU use then you're better off just tweaking the
>> limits upwards. Then you get that with everything, like when you run
>> "commit" in a for-loop, not just this one special case of "fetch".
>>
>> We have existing potentially long-running operations like "fetch",
>> "rebase" and "git svn fetch" that run "gc --auto" for their incremental
>> steps, and that's a feature.
>
> gc --auto is added at arbitrary points to help garbage collection. I
> don't think it's ever intended to "do gc at this and that exact
> moment", just "hey this command has taken a lot of time already (i.e.
> no instant response needed) and it may have added a bit more garbage,
> let's just check real quick".

I don't mean we can't ever change the algorithm, but that we've documented:

    When common porcelain operations that create objects are run, they
    will check whether the repository has grown substantially since the
    last maintenance[...]

The "fetch" command is a common porcelain operation, when it fetches from N remotes it just runs an invocation of itself, so thus far it's both worked & been intuitive that if we needed (potentially multiple) gc's while doing that we'd just go ahead and run it then, even if something concurrent was happening.

No that's not optimal in many cases, but at least doesn't create caveats we don't have now where we have runaway object growth.

Show 7 quoted lines
>> It keeps "gc --auto" dumb enough to avoid a pathological case where
>> we'll have a ballooning objects dir because we figure we can run
>> something "at the end", when "the end" could be hours away, and we're
>> adding a new pack or hundreds of loose objects every second.
>
> Are we optimizing for a rare (large scale) case? Such setup requires
> tuning regardless to me.

At least for me it doesn't require custom tuning before this patch of yours.

I.e. now "gc --auto" is dumb enough that you can run it on everything from stuff that just does "commit" from cron, user's laptops, massive rebases that take forever, and e.g. "stats" like jobs where for <reasons> I'll add thousands of repos and "fetch --all" them (so e.g. I can run "log --author=<x> --all").

Yeah of course I'm an advanced user and I can just grumble and manually invoke fetch, actually I'll probably submit a follow-up patch to add a gc.* config to disable this thing.

But I think even if the use-case is rather obscure it's important that if at all possible we keep "gc" elastic enough to work for pretty much all combinations of object-adding porcelain commands, and I think in this case we're better off doing things differently...

Show 6 quoted lines
>> So I don't think Duy's patch is a good way to go.
>
> This reminds me of being perfect is the enemy of the good. A normal
> user has a couple remotes at most, finishing fast (enough) and in such
> case it's a good idea to wait until everything is in before running
> gc.

Even a user with two remotes will run into issues with your patch where "gc" will print things twice, or outright error due to concurrent access by another process, which as has been discussed here on-list is *very* common e.g. with editor integration.

So the extent of my complains about this is:
 1) The edge case with runaway object growth (obscure)
 2) It doesn't really fix the bug except for the narrow case of users
    who invoke things one-terminal-at-a-time and don't e.g. have an
    editor with "git" integration & a terminal (not obscure).
Maybe I should have led with #2 :)

Anyway, I'm not *just* complaining. I have patches too, but so far I'm up to 8 patches on what's probably just the first one-third of it. *Sigh*.

But the 2/3 of that if you want to dig through my crappy WIP code on GitHub is instrumenting the test suite to demonstrate that with an approach like what your patch does we still get these GC race conditions, because our locking still sucks, which brings me to...

> Of course making git-gc more robust wrt. parallel access is great, but
> it's hard work. Dealing with locks is always tricky, especially when
> new locks can come up any time.

I've poked at it a bit now, and it's really not hard, I think it's just that nobody looked at it hard enough before.

The issue is that currently we do:
    1. parent: do_we_need_gc();
    2. parent: say_way_will_gc();
    3. parent: lock();
    4. parent: do_a_bit_of_work();
    5. parent: unlock();
    6. parent: fork();
    7. child: lock();
    8. child: do_a_lot_of_work();
    9. child: unlock();
My RFC patch currently changes that to:
    1. parent: do_we_need_gc();
    2. parent: lock_or_silently_exit();
    3. parent: say_way_will_gc();
    4. parent: do_a_bit_of_work();
    5. parent: unlock();
    6. parent: fork();
    7. child: lock();
    8. child: do_a_lot_of_work();
    9. child: unlock();

I.e. we won't duplicate the message, but *do* introduce the caveat that it's in principle possible nobody gc's, but in practice when we fail to get the lock in lock_or_silently_exit() it's because we lost the race to a sister process that's going to do the actual GC, so all is well.

But we are left with the brief race when we fork. My RFC patch proposed some elaborate PID hand-over dance to deal with this. But having looked at it again I think we can easily get rid of that race too with just:

    1. parent: do_we_need_gc();
    2. parent: lock_or_silently_exit();
    3. parent: say_way_will_gc();
    4. parent: do_a_bit_of_work();
    5. parent: fork();
    6. parent: <writes child pid to gc.lock>
    7. parent: <exits without unlocking gc.lock>
    8. child: do_a_lot_of_work();
    9. child: unlock();

Which can fail in cases where the child segfaults, or manages to exit earlier than the parent etc, hence my earlier proposed elaborate pid hand-over dance.

But looking at it again we only usurp an existing gc.lock if the mtime is >12hrs, so it's OK if we have very rare cases where the PID info got corrupted, we can still back ourselves out of it, which is what I was paranoid about.

Furthermore the gc.lock contains the <hostname><pid> of the working process, but we can in a backwards-compatible way add new entries to that file, i.e. list both the child & parent pid. Older clients will just read whichever one comes first, but if we make new versions check both we can be more paranoid going forward.

> Having said that, I don't mind if my patch gets dropped. It was just a
> "hey that multiple gc output looks strange, hah the fix is quite
> simple" moment for me.
Previous: Duy NguyenNext: Jeff King
Message 55 of 61 in “fetch: only run 'gc' once when fetching multiple remotes”
  1. fetch: only run 'gc' once when fetching multiple remotesNguyễn Thái Ngọc Duy, Jun 19, 2019
  2. gc: run more pre-detach operations under lockÆvar Arnfjörð Bjarmason, Jun 19, 2019
  3. Duy NguyenJun 19, 2019
  4. Ævar Arnfjörð BjarmasonJun 19, 2019
  5. Jeff KingJun 19, 2019
  6. Ævar Arnfjörð BjarmasonJun 19, 2019
  7. 0/6 Change <non-empty?> GIT_TEST_* variables to <boolean>Ævar Arnfjörð Bjarmason, Jun 19, 2019
  8. Junio C HamanoJun 20, 2019
  9. Ævar Arnfjörð BjarmasonJun 20, 2019
  10. Junio C HamanoJun 20, 2019
  11. 0/8 Change <non-empty?> GIT_TEST_* variables to <boolean>Ævar Arnfjörð Bjarmason, Jun 20, 2019
  12. 1/8 config tests: simplify include cycle testÆvar Arnfjörð Bjarmason, Jun 21, 2019
  13. 0/8 Change <non-empty?> GIT_TEST_* variables to <boolean>Ævar Arnfjörð Bjarmason, Jun 21, 2019
  14. 2/8 env--helper: new undocumented builtin wrapping git_env_*()Ævar Arnfjörð Bjarmason, Jun 21, 2019
  15. Junio C HamanoJun 21, 2019
  16. 3/8 config.c: refactor die_bad_number() to not call gettext() earlyÆvar Arnfjörð Bjarmason, Jun 21, 2019
  17. 4/8 t6040 test: stop using global "script" variableÆvar Arnfjörð Bjarmason, Jun 21, 2019
  18. 6/8 tests README: re-flow a previously changed paragraphÆvar Arnfjörð Bjarmason, Jun 21, 2019
  19. 5/8 tests: make GIT_TEST_GETTEXT_POISON a booleanÆvar Arnfjörð Bjarmason, Jun 21, 2019
  20. Junio C HamanoJun 24, 2019
  21. 7/8 tests: replace test_tristate with "git env--helper"Ævar Arnfjörð Bjarmason, Jun 21, 2019
  22. 1/2 t/lib-git-svn.sh: check GIT_TEST_SVN_HTTPD when running SVN HTTP testsSZEDER Gábor, Sep 6, 2019
  23. 2/2 ci: restore running httpd testsSZEDER Gábor, Sep 6, 2019
  24. Junio C HamanoSep 6, 2019
  25. Jeff KingSep 6, 2019
  26. SZEDER GáborSep 7, 2019
  27. 0/2 tests: catch non-bool GIT_TEST_* valuesSZEDER Gábor, Nov 22, 2019
  28. 1/2 tests: add 'test_bool_env' to catch non-bool GIT_TEST_* valuesSZEDER Gábor, Nov 22, 2019
  29. Jeff KingNov 25, 2019
  30. 2/2 t5608-clone-2gb.sh: turn GIT_TEST_CLONE_2GB into a boolSZEDER Gábor, Nov 22, 2019
  31. Jeff KingNov 25, 2019
  32. 8/8 tests: make GIT_TEST_FAIL_PREREQS a booleanÆvar Arnfjörð Bjarmason, Jun 21, 2019
  33. 1/8 config tests: simplify include cycle testÆvar Arnfjörð Bjarmason, Jun 20, 2019
  34. 2/8 env--helper: new undocumented builtin wrapping git_env_*()Ævar Arnfjörð Bjarmason, Jun 20, 2019
  35. Junio C HamanoJun 20, 2019
  36. Junio C HamanoJun 20, 2019
  37. Ævar Arnfjörð BjarmasonJun 21, 2019
  38. Junio C HamanoJun 21, 2019
  39. 3/8 config.c: refactor die_bad_number() to not call gettext() earlyÆvar Arnfjörð Bjarmason, Jun 20, 2019
  40. 4/8 t6040 test: stop using global "script" variableÆvar Arnfjörð Bjarmason, Jun 20, 2019
  41. 5/8 tests: make GIT_TEST_GETTEXT_POISON a booleanÆvar Arnfjörð Bjarmason, Jun 20, 2019
  42. 7/8 tests: replace test_tristate with "git env--helper"Ævar Arnfjörð Bjarmason, Jun 20, 2019
  43. 6/8 tests README: re-flow a previously changed paragraphÆvar Arnfjörð Bjarmason, Jun 20, 2019
  44. 8/8 tests: make GIT_TEST_FAIL_PREREQS a booleanÆvar Arnfjörð Bjarmason, Jun 20, 2019
  45. 1/6 env--helper: new undocumented builtin wrapping git_env_*()Ævar Arnfjörð Bjarmason, Jun 19, 2019
  46. Junio C HamanoJun 20, 2019
  47. 2/6 t6040 test: stop using global "script" variableÆvar Arnfjörð Bjarmason, Jun 19, 2019
  48. Junio C HamanoJun 20, 2019
  49. 3/6 tests: make GIT_TEST_GETTEXT_POISON a booleanÆvar Arnfjörð Bjarmason, Jun 19, 2019
  50. Junio C HamanoJun 20, 2019
  51. 5/6 tests: replace test_tristate with "git env--helper"Ævar Arnfjörð Bjarmason, Jun 19, 2019
  52. 4/6 tests README: re-flow a previously changed paragraphÆvar Arnfjörð Bjarmason, Jun 19, 2019
  53. 6/6 tests: make GIT_TEST_FAIL_PREREQS a booleanÆvar Arnfjörð Bjarmason, Jun 19, 2019
  54. Duy NguyenJun 20, 2019
  55. Ævar Arnfjörð BjarmasonJun 20, 2019
  56. Jeff KingJun 20, 2019
  57. Junio C HamanoJun 20, 2019
  58. Jeff KingJun 19, 2019
  59. Jeff KingJun 19, 2019
  60. Duy NguyenJun 20, 2019
  61. Jeff KingJun 20, 2019

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.