{"thread":{"id":"46311","subject":"Flurries of 'git reflog expire'","startedAt":"2017-07-04T08:10:57Z","lastAt":"2017-07-12T21:10:12Z","messageCount":15,"participants":["Andreas Krey","Ævar Arnfjörð Bjarmason","Jeff King","Bryan Turner","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"323796","messageId":"20170704075758.GA22249@inner.h.apk.li","threadId":"46311","inReplyTo":null,"subject":"Flurries of 'git reflog expire'","fromName":"Andreas Krey","fromEmail":"a.krey@gmx.de","sentAt":"2017-07-04T07:57:58Z","receivedAt":"2017-07-04T08:10:57Z","isPatch":false,"sender":{"key":"a.krey@gmx.de","avatar":"https://avatars.githubusercontent.com/u/37810?v=4"},"body":"Hi everyone,\n\nhow is 'git reflog expire' triggered? We're occasionally seeing a lot\nof the running in parallel on a single of our repos (atlassian bitbucket),\nand this usually means that the machine is not very responsive for\ntwenty minutes, the repo being as big as it is.\n\nThe server is still on git 2.6.2 (and bitbucket 4.14.5).\n\nQuestions:\n\nWhat can be done about this? Cronjob 'git reflog expire' at midnight,\nso the heuristic don't trigger during the day? (The relnotes don't\nmention anything after 2.4.0, so I suppose a git upgrade won't help.)\n\nWhat is the actual cause? Bad heuristics in git itself, or does\nbitbucket run them too often (improbable)?\n\nAndreas\n\n-- \n\"Totally trivial. Famous last words.\"\nFrom: Linus Torvalds <torvalds@*.org>\nDate: Fri, 22 Jan 2010 07:29:21 -0800\n"},{"id":"323800","messageId":"87podgbkqi.fsf@gmail.com","threadId":"46311","inReplyTo":"20170704075758.GA22249@inner.h.apk.li","subject":"Re: Flurries of 'git reflog expire'","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2017-07-04T09:43:33Z","receivedAt":"2017-07-04T09:43:44Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Jul 04 2017, Andreas Krey jotted:\n\n> Hi everyone,\n>\n> how is 'git reflog expire' triggered? We're occasionally seeing a lot\n> of the running in parallel on a single of our repos (atlassian bitbucket),\n> and this usually means that the machine is not very responsive for\n> twenty minutes, the repo being as big as it is.\n\nAssuming Linux, what does 'ps auxf' look like when this happens? Is the\nparent a 'git gc --auto'?\n\n> The server is still on git 2.6.2 (and bitbucket 4.14.5).\n\nYou might want to upgrade, we've had a bunch of changes since then,\nmaybe some of this fixes it:\n\n    git log --reverse -p -L'/^static.*lock_repo_for/,/^}/:builtin/gc.c'\n\n> Questions:\n>\n> What can be done about this? Cronjob 'git reflog expire' at midnight,\n> so the heuristic don't trigger during the day? (The relnotes don't\n> mention anything after 2.4.0, so I suppose a git upgrade won't help.)\n>\n> What is the actual cause? Bad heuristics in git itself, or does\n> bitbucket run them too often (improbable)?\n\nYou can set gc.auto=0 in the repo to disable auto-gc, and play with\ne.g. the reflog expire values, see the git-gc manpage.\n\nBut then you need to run your own gc, which is not a bad idea anyway\nwith a dedicated git server.\n\nBut it would be good to get to the bottom of this, we shouldn't be\nrunning these concurrently.\n"},{"id":"323830","messageId":"20170705082027.ujddejajjlvto7bp@sigill.intra.peff.net","threadId":"46311","inReplyTo":"20170704075758.GA22249@inner.h.apk.li","subject":"Re: Flurries of 'git reflog expire'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-07-05T08:20:27Z","receivedAt":"2017-07-05T08:20:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jul 04, 2017 at 09:57:58AM +0200, Andreas Krey wrote:\n\n> Questions:\n> \n> What can be done about this? Cronjob 'git reflog expire' at midnight,\n> so the heuristic don't trigger during the day? (The relnotes don't\n> mention anything after 2.4.0, so I suppose a git upgrade won't help.)\n> \n> What is the actual cause? Bad heuristics in git itself, or does\n> bitbucket run them too often (improbable)?\n\nIf it's using --expire-unreachable (which a default \"git gc\" does), that\nmeans the we have to traverse the entire history to see what is\nreachable and what is not. Added on to a normal git-gc, that's usually\nnot a big deal (it has to do that traversal and much more for the\nrepack). But if bitbucket is triggering it for other operations, that\ncould be related (I don't think anything but gc should ever run it\notherwise).\n\nI seem to recall that using --stale-fix is also extremely expensive,\ntoo. What do the command line arguments for the slow commands look like?\nAnd what does the process tree look like?\n\n-Peff\n"},{"id":"323892","messageId":"20170706132744.GA1216@inner.h.apk.li","threadId":"46311","inReplyTo":"87podgbkqi.fsf@gmail.com","subject":"Re: Flurries of 'git reflog expire'","fromName":"Andreas Krey","fromEmail":"a.krey@gmx.de","sentAt":"2017-07-06T13:27:44Z","receivedAt":"2017-07-06T13:27:55Z","isPatch":false,"sender":{"key":"a.krey@gmx.de","avatar":"https://avatars.githubusercontent.com/u/37810?v=4"},"body":"On Tue, 04 Jul 2017 11:43:33 +0000, Ævar Arnfjörð Bjarmason wrote:\n...\n> You can set gc.auto=0 in the repo to disable auto-gc, and play with\n> e.g. the reflog expire values, see the git-gc manpage.\n> \n> But then you need to run your own gc, which is not a bad idea anyway\n> with a dedicated git server.\n\nActually, bitbucket should be doing this. Although I can't quite\nrule out the possibility that we reenabled GC in this repo some\ntime ago.\n\n> But it would be good to get to the bottom of this, we shouldn't be\n> running these concurrently.\n\nIndeed. Unfortunately this isn't easily reproduced in the test instance,\nso I will need to get a newer git under the production bitbucket.\n\nThere are quite some of\n\n          \\_ /usr/bin/git receive-pack /opt/apps/atlassian/bitbucket-data/shared/data/repositories/68\n          |   \\_ git gc --auto --quiet\n          |       \\_ git reflog expire --all\n\nin the process tree, apparently a new one gets started even though previous\nones are still running. The number of running expires grew slowly, in the\norder of many minutes.\n\nAndreas\n\n-- \n\"Totally trivial. Famous last words.\"\nFrom: Linus Torvalds <torvalds@*.org>\nDate: Fri, 22 Jan 2010 07:29:21 -0800\n"},{"id":"323893","messageId":"20170706133124.GB1216@inner.h.apk.li","threadId":"46311","inReplyTo":"20170705082027.ujddejajjlvto7bp@sigill.intra.peff.net","subject":"Re: Flurries of 'git reflog expire'","fromName":"Andreas Krey","fromEmail":"a.krey@gmx.de","sentAt":"2017-07-06T13:31:24Z","receivedAt":"2017-07-06T13:31:36Z","isPatch":false,"sender":{"key":"a.krey@gmx.de","avatar":"https://avatars.githubusercontent.com/u/37810?v=4"},"body":"On Wed, 05 Jul 2017 04:20:27 +0000, Jeff King wrote:\n> On Tue, Jul 04, 2017 at 09:57:58AM +0200, Andreas Krey wrote:\n...\n> I seem to recall that using --stale-fix is also extremely expensive,\n> too. What do the command line arguments for the slow commands look like?\n\nThe problem isn't that the expire is slow, it is that there are\nmany of them, waiting for disk writes.\n\n> And what does the process tree look like?\n\nLots (~ 10) of\n\n          \\_ /usr/bin/git receive-pack /opt/apps/atlassian/bitbucket-data/shared/data/repositories/68\n          |   \\_ git gc --auto --quiet\n          |       \\_ git reflog expire --all\n\nplus another dozen gc/expire pairs where the parent is already gone.\nAll with the same arguments - auto GC.\n\nI'd wager that each push sees that a GC is in order,\nand doesn't notice that there is one already running.\n\n- Andreas\n\n-- \n\"Totally trivial. Famous last words.\"\nFrom: Linus Torvalds <torvalds@*.org>\nDate: Fri, 22 Jan 2010 07:29:21 -0800\n"},{"id":"323902","messageId":"CAGyf7-FnaWM=XNb_Skb1qR4vu_jAw-5swkgWpEDQqwM0NNq3YQ@mail.gmail.com","threadId":"46311","inReplyTo":"20170706133124.GB1216@inner.h.apk.li","subject":"Re: Flurries of 'git reflog expire'","fromName":"Bryan Turner","fromEmail":"bturner@atlassian.com","sentAt":"2017-07-06T17:01:05Z","receivedAt":"2017-07-06T17:01:12Z","isPatch":false,"sender":{"key":"bturner@atlassian.com","avatar":"https://gravatar.com/avatar/16bcf3167981c1ef7c804e502642366d888a35b0d0b0a4ca01fdc442aa1acb1e?d=mp&s=160"},"body":"I'm one of the Bitbucket Server developers. My apologies; I just\nnoticed this thread or I would have jumped in sooner!\n\nOn Thu, Jul 6, 2017 at 6:31 AM, Andreas Krey <a.krey@gmx.de> wrote:\n> On Wed, 05 Jul 2017 04:20:27 +0000, Jeff King wrote:\n>> On Tue, Jul 04, 2017 at 09:57:58AM +0200, Andreas Krey wrote:\n> ...\n>> And what does the process tree look like?\n>\n> Lots (~ 10) of\n>\n>           \\_ /usr/bin/git receive-pack /opt/apps/atlassian/bitbucket-data/shared/data/repositories/68\n>           |   \\_ git gc --auto --quiet\n>           |       \\_ git reflog expire --all\n>\n> plus another dozen gc/expire pairs where the parent is already gone.\n> All with the same arguments - auto GC.\n\nDo you know what version of Bitbucket Server is in use? Based on the\nfact that it's \"git gc --auto\" triggered from a \"git receive-pack\",\nthat implies two things:\n- You're on a 4.x version of Bitbucket Server\n- The repository (68) has never been forked\n\nDepending on your Bitbucket Server version (this being the reason I\nasked), there are a couple different fixes available:\n\n- Fork the repository. You don't need to _use_ the fork, but having a\nfork existing will trigger Bitbucket Server to disable auto GC and\nfully manage that itself. That includes managing both _concurrency_\nand _frequency_ of GC. This works on all versions of Bitbucket Server.\n\n- Run \"git config gc.auto 0\" in\n/opt/apps/atlassian/bitbucket-data/shared/data/repositories/68 to\ndisable auto GC yourself. This may be preferable to forking the\nrepository, which, in addition to disabling auto GC, also disables\nobject pruning. However, you must be running at least Bitbucket Server\n4.6.0 for this approach to work. Otherwise auto GC will simply be\nreenabled the first time Bitbucket Server goes to trigger GC, when it\ndetects that the repository has no forks.\n\nAssuming you're on 4.6.0 or newer, either approach should fix the\nissue. If you're on 4.5 or older, forking is the only viable approach\nunless you upgrade Bitbucket Server first.\n\nI also want to add that Bitbucket Server 5.x includes totally\nrewritten GC handling. 5.0.x automatically disables auto GC in all\nrepositories and manages it explicitly, and 5.1.x fully removes use of\n\"git gc\" in favor of running relevant plumbing commands directly. We\nmoved away from \"git gc\" specifically to avoid the \"git reflog expire\n--all\", because there's no config setting that _fully disables_\nforking that process. By default our bare clones only have reflogs for\npull request refs, and we've explicitly configured those to never\nexpire, so all \"git reflog expire --all\" can do is use up I/O and,\nquite frequently, fail because refs are updated. Since we stopped\nrunning \"git gc\", we've not yet seen any GC failures on our internal\nBitbucket Server clusters.\n\nBitbucket Server 5.1.x also includes a new \"gc.log\" (not to be\nconfused with the one Git itself writes) which retains a record of\nevery GC-related process we run in each repository, and how long that\nprocess took to complete. That can be useful for getting clearer\ninsight into both how often GC work is being done, and how long it's\ntaking.\n\nUpgrading to 5.x can be a bit of an undertaking, since the major\nversion brings API changes, so it's totally understandable that many\norganizations haven't upgraded yet. I'm just noting that these\nimprovements are there for when such an upgrade becomes viable.\n\nHope this helps!\nBryan\n\n>\n> I'd wager that each push sees that a GC is in order,\n> and doesn't notice that there is one already running.\n>\n> - Andreas\n>\n> --\n> \"Totally trivial. Famous last words.\"\n> From: Linus Torvalds <torvalds@*.org>\n> Date: Fri, 22 Jan 2010 07:29:21 -0800\n"},{"id":"324150","messageId":"20170711044553.GG3786@inner.h.apk.li","threadId":"46311","inReplyTo":"CAGyf7-FnaWM=XNb_Skb1qR4vu_jAw-5swkgWpEDQqwM0NNq3YQ@mail.gmail.com","subject":"Re: Flurries of 'git reflog expire'","fromName":"Andreas Krey","fromEmail":"a.krey@gmx.de","sentAt":"2017-07-11T04:45:53Z","receivedAt":"2017-07-11T04:46:07Z","isPatch":false,"sender":{"key":"a.krey@gmx.de","avatar":"https://avatars.githubusercontent.com/u/37810?v=4"},"body":"On Thu, 06 Jul 2017 10:01:05 +0000, Bryan Turner wrote:\n....\n> Do you know what version of Bitbucket Server is in use?\n\nWe're on the newest 4.x.\n\n...\n> - Run \"git config gc.auto 0\" in\n\nGoing that route.\n\n...\n> I also want to add that Bitbucket Server 5.x includes totally\n> rewritten GC handling. 5.0.x automatically disables auto GC in all\n> repositories and manages it explicitly, and 5.1.x fully removes use of\n> \"git gc\" in favor of running relevant plumbing commands directly.\n\nThat's the part that irks me. This shouldn't be necessary - git itself\nshould make sure auto GC isn't run in parallel. Now I probably can't\nevaluate whether a git upgrade would fix this, but given that you\nare going the do-gc-ourselves route I suppose it wouldn't.\n\n...\n> Upgrading to 5.x can be a bit of an undertaking, since the major\n> version brings API changes,\n\nThe upgrade is on my todo list, but there are plugins that don't\nappear to be ready for 5.0, notable the jenkins one.\n\nAndreas\n\n-- \n\"Totally trivial. Famous last words.\"\nFrom: Linus Torvalds <torvalds@*.org>\nDate: Fri, 22 Jan 2010 07:29:21 -0800\n"},{"id":"324156","messageId":"20170711072536.ijpldg4uxb5pbtdw@sigill.intra.peff.net","threadId":"46311","inReplyTo":"20170711044553.GG3786@inner.h.apk.li","subject":"[BUG] detached auto-gc does not respect lock for 'reflog expire', was Re: Flurries of 'git reflog expire'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-07-11T07:25:37Z","receivedAt":"2017-07-11T07:25:44Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"[Updating the subject since I think this really is a bug].\n\nOn Tue, Jul 11, 2017 at 06:45:53AM +0200, Andreas Krey wrote:\n\n> > I also want to add that Bitbucket Server 5.x includes totally\n> > rewritten GC handling. 5.0.x automatically disables auto GC in all\n> > repositories and manages it explicitly, and 5.1.x fully removes use of\n> > \"git gc\" in favor of running relevant plumbing commands directly.\n> \n> That's the part that irks me. This shouldn't be necessary - git itself\n> should make sure auto GC isn't run in parallel. Now I probably can't\n> evaluate whether a git upgrade would fix this, but given that you\n> are going the do-gc-ourselves route I suppose it wouldn't.\n\nIt's _supposed_ to take a lock, even in older versions. See 64a99eb47\n(gc: reject if another gc is running, unless --force is given,\n2013-08-08).\n\nBut it looks like before we take that lock, we sometimes run pack-refs\nand reflog expire. This is due to 62aad1849 (gc --auto: do not lock refs\nin the background, 2014-05-25). IMHO this is buggy; it should be\nchecking the lock before calling gc_before_repack() and daemonizing.\n\nAnnoyingly, the lock code interacts badly with daemonizing because that\nlatter will fork to a new process. So the simple solution like:\n\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex 2ba50a287..79480124a 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -414,6 +414,9 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t\tif (report_last_gc_error())\n \t\t\t\treturn -1;\n \n+\t\t\tif (lock_repo_for_gc(force, &pid))\n+\t\t\t\treturn 0;\n+\n \t\t\tif (gc_before_repack())\n \t\t\t\treturn -1;\n \t\t\t/*\n\nmeans that anybody looking at the lockfile will report the wrong pid\n(and thus think the lock is invalid). I guess we'd need to update it in\nplace after daemonizing.\n\n-Peff\n"},{"id":"324157","messageId":"20170711073125.bddyb3ledzasbhtl@sigill.intra.peff.net","threadId":"46311","inReplyTo":"CAGyf7-FnaWM=XNb_Skb1qR4vu_jAw-5swkgWpEDQqwM0NNq3YQ@mail.gmail.com","subject":"Re: Flurries of 'git reflog expire'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-07-11T07:31:25Z","receivedAt":"2017-07-11T07:31:43Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jul 06, 2017 at 10:01:05AM -0700, Bryan Turner wrote:\n\n> I also want to add that Bitbucket Server 5.x includes totally\n> rewritten GC handling. 5.0.x automatically disables auto GC in all\n> repositories and manages it explicitly, and 5.1.x fully removes use of\n> \"git gc\" in favor of running relevant plumbing commands directly. We\n> moved away from \"git gc\" specifically to avoid the \"git reflog expire\n> --all\", because there's no config setting that _fully disables_\n> forking that process.\n\nFWIW, I think auto-gc in general is not a good way to handle maintenance\non a busy hosting server. Repacking can be very resource hungry (both\nCPU and memory), and it needs to be throttled. You _could_ throttle with\nan auto-gc hook, but that isn't very elegant when it comes to\nre-queueing jobs which fail or timeout.\n\nThe right model IMHO (and what GitHub uses, and what I'm guessing\nBitbucket is doing in more recent versions) is to make note of write\noperations in a data structure, then use that data to schedule\nmaintenance in a job queue. But that can never really be part of Git\nitself, as the notion of a system job queue is outside its scope.\n\n-Peff\n"},{"id":"324160","messageId":"CAGyf7-F-zG7NR_bd_sVLdVM5xT9UbJrW_=5imiihNu15E5OsYg@mail.gmail.com","threadId":"46311","inReplyTo":"20170711044553.GG3786@inner.h.apk.li","subject":"Re: Flurries of 'git reflog expire'","fromName":"Bryan Turner","fromEmail":"bturner@atlassian.com","sentAt":"2017-07-11T07:35:50Z","receivedAt":"2017-07-11T07:36:09Z","isPatch":false,"sender":{"key":"bturner@atlassian.com","avatar":"https://gravatar.com/avatar/16bcf3167981c1ef7c804e502642366d888a35b0d0b0a4ca01fdc442aa1acb1e?d=mp&s=160"},"body":"On Mon, Jul 10, 2017 at 9:45 PM, Andreas Krey <a.krey@gmx.de> wrote:\n> On Thu, 06 Jul 2017 10:01:05 +0000, Bryan Turner wrote:\n> ....\n>> I also want to add that Bitbucket Server 5.x includes totally\n>> rewritten GC handling. 5.0.x automatically disables auto GC in all\n>> repositories and manages it explicitly, and 5.1.x fully removes use of\n>> \"git gc\" in favor of running relevant plumbing commands directly.\n>\n> That's the part that irks me. This shouldn't be necessary - git itself\n> should make sure auto GC isn't run in parallel. Now I probably can't\n> evaluate whether a git upgrade would fix this, but given that you\n> are going the do-gc-ourselves route I suppose it wouldn't.\n>\n\nI believe I've seen some commits on the mailing list that suggest \"git\ngc --auto\" manages its concurrency better in newer versions than it\nused to, but even then it can only manage its concurrency within a\nsingle repository. For a hosting server with thousands, or tens of\nthousands, of active repositories, there still wouldn't be any\nprotection against \"git gc --auto\" running concurrently in dozens of\nthem at the same time.\n\nBut it's not only about concurrency. \"git gc\" (and by extension \"git\ngc --auto\") is a general purpose tool, designed to generally do what\nyou need, and to mostly stay out of your way while it does it. I'd\nhazard to say it's not really designed for managing heavily-trafficked\nrepositories on busy hosting services, though, and as a result, there\nare things it can't do.\n\nFor example, I can configure auto GC to run based on how many loose\nobjects or packs I have, but there's no heuristic to make it repack\nrefs when I have a lot of loose ones, or configure it to _only_ pack\nrefs without repacking objects or pruning reflogs. There are knobs for\nvarious things (like \"gc.*.reflogExpire\"), but those don't give\ncomplete control. Even if I set \"gc.reflogExpire=never\", \"git gc\"\nstill forks \"git reflog expire --all\" (compared to\n\"gc.packRefs=false\", which completely prevents forking \"git\npack-refs\").\n\nA trace on \"git gc\" shows this:\n$ GIT_TRACE=1 git gc\n00:10:45.058066 git.c:437               trace: built-in: git 'gc'\n00:10:45.067075 run-command.c:369       trace: run_command:\n'pack-refs' '--all' '--prune'\n00:10:45.077086 git.c:437               trace: built-in: git\n'pack-refs' '--all' '--prune'\n00:10:45.084098 run-command.c:369       trace: run_command: 'reflog'\n'expire' '--all'\n00:10:45.093102 git.c:437               trace: built-in: git 'reflog'\n'expire' '--all'\n00:10:45.097088 run-command.c:369       trace: run_command: 'repack'\n'-d' '-l' '-A' '--unpack-unreachable=2.weeks.ago'\n00:10:45.106096 git.c:437               trace: built-in: git 'repack'\n'-d' '-l' '-A' '--unpack-unreachable=2.weeks.ago'\n00:10:45.107098 run-command.c:369       trace: run_command:\n'pack-objects' '--keep-true-parents' '--honor-pack-keep' '--non-empty'\n'--all' '--reflog' '--indexed-objects'\n'--unpack-unreachable=2.weeks.ago' '--local' '--delta-base-offset'\n'objects/pack/.tmp-15212-pack'\n00:10:45.127117 git.c:437               trace: built-in: git\n'pack-objects' '--keep-true-parents' '--honor-pack-keep' '--non-empty'\n'--all' '--reflog' '--indexed-objects'\n'--unpack-unreachable=2.weeks.ago' '--local' '--delta-base-offset'\n'objects/pack/.tmp-15212-pack'\nCounting objects: 6, done.\nDelta compression using up to 16 threads.\nCompressing objects: 100% (2/2), done.\nWriting objects: 100% (6/6), done.\nTotal 6 (delta 0), reused 6 (delta 0)\n00:10:45.173161 run-command.c:369       trace: run_command: 'prune'\n'--expire' '2.weeks.ago'\n00:10:45.184171 git.c:437               trace: built-in: git 'prune'\n'--expire' '2.weeks.ago'\n00:10:45.199202 run-command.c:369       trace: run_command: 'worktree'\n'prune' '--expire' '3.months.ago'\n00:10:45.208193 git.c:437               trace: built-in: git\n'worktree' 'prune' '--expire' '3.months.ago'\n00:10:45.212198 run-command.c:369       trace: run_command: 'rerere' 'gc'\n00:10:45.221223 git.c:437               trace: built-in: git 'rerere' 'gc'\n\nThe bare repositories used by Bitbucket Server:\n* Don't have reflogs enabled generally, and for the ones that are\nenabled \"gc.*.reflogExpire\" is set to \"never\"\n* Never have worktrees, so they don't need to be pruned\n* Never use rerere, so that doesn't need to GC\n* Have pruning disabled if they've been forked, due to using\nalternates to manage disk space\n\nThat means of all the commands \"git gc\" runs, under the covers, at\nmost only \"pack-refs\", \"repack\" and sometimes \"prune\" have any value.\n\"reflog expire --all\" in particular is extremely likely to fail. Which\nbrings up another consideration.\n\n\"git gc --auto\" has no sense of context, or adjacent behavior. Even if\nit correctly guards against concurrency, it still doesn't know what\nelse is going on. Immediately after a push, Bitbucket Server has many\nother housekeeping tasks it performs, especially around pull requests.\nThat means pull request refs are disproportionately likely to be\n\"moving\" immediately after a push completes--exactly when \"git gc\n--auto\" tries to run. (Which tends to be why \"reflog expire --all\"\nfails, due ref locking issues with pull request refs.) Bitbucket\nServer, on the other hand, better understands the context GC is\nrunning in. So it can defer GC processing for a period of time after a\npush completes, to increase the likelihood that the repository is\n\"quiet\" and GC can complete without issue.\n\nAnother limitation is that you can't configure \"negative\" heuristics,\nlike \"Don't run GC more than once per day.\". If \"git gc --auto\"'s\nheuristics are exceeded, it'll run GC. Depending, for example, on how\nrapidly a repository generates unreachable objects, it's entirely\npossible to get to a point where \"git gc --auto\" wants to run after\nevery single push, sometimes for days in a row, while it waits for\nobjects to hit the prune threshold. By managing GC ourselves, we gain\nthe ability to enforce \"cooldowns\" to prevent continuous GC.\n\n\"git gc --auto\" also has a tendency to run \"attached\" to the \"git\nreceive-pack\" process, which means both that pushing users can have\ntheir local process \"delayed\" while it runs, and that they sometimes\nget to see \"scary\" errors that they can't fix (or, often, understand).\nNewer versions of Git have increased the likelihood that \"git gc\n--auto\" will run detached, but that doesn't always happen. (Up to and\nincluding 2.13.2, the \"git config\" documentation for \"gc.autoDetach\"\nis qualified with \"if the system supports it.\") Managing GC in\nBitbucket Server guarantees that it's _always_ detached from user\nprocesses.\n\nThat's a few of the reasons we've switched over. I'd imagine most\nhosting providers take a similarly \"hands on\" approach to controlling\ntheir GC. Beyond a certain scale, it seems almost unavoidable. Git\nnever has more than a repository-level view of the world; only the\nhosting provider can see the big picture.\n\nBest regards,\nBryan Turner\n\n> ...\n>> Upgrading to 5.x can be a bit of an undertaking, since the major\n>> version brings API changes,\n>\n> The upgrade is on my todo list, but there are plugins that don't\n> appear to be ready for 5.0, notable the jenkins one.\n>\n> Andreas\n>\n> --\n> \"Totally trivial. Famous last words.\"\n> From: Linus Torvalds <torvalds@*.org>\n> Date: Fri, 22 Jan 2010 07:29:21 -0800\n"},{"id":"324163","messageId":"20170711074539.xqtqqee7a7fbemyn@sigill.intra.peff.net","threadId":"46311","inReplyTo":"CAGyf7-F-zG7NR_bd_sVLdVM5xT9UbJrW_=5imiihNu15E5OsYg@mail.gmail.com","subject":"Re: Flurries of 'git reflog expire'","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-07-11T07:45:40Z","receivedAt":"2017-07-11T07:45:51Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jul 11, 2017 at 12:35:50AM -0700, Bryan Turner wrote:\n\n> That's a few of the reasons we've switched over. I'd imagine most\n> hosting providers take a similarly \"hands on\" approach to controlling\n> their GC. Beyond a certain scale, it seems almost unavoidable. Git\n> never has more than a repository-level view of the world; only the\n> hosting provider can see the big picture.\n\nThanks for writing this out. I agree with all of the reasons given (in\nmy email which I suspect crossed with yours, I just said \"throttling\",\nbut there really are a lot of other reasons).\n\n-Peff\n"},{"id":"324168","messageId":"20170711090635.swowex7yry7kqb7v@sigill.intra.peff.net","threadId":"46311","inReplyTo":"20170711072536.ijpldg4uxb5pbtdw@sigill.intra.peff.net","subject":"[PATCH] gc: run pre-detach operations under lock","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-07-11T09:06:35Z","receivedAt":"2017-07-11T09:06:42Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Jul 11, 2017 at 03:25:36AM -0400, Jeff King wrote:\n\n> Annoyingly, the lock code interacts badly with daemonizing because that\n> latter will fork to a new process. So the simple solution like:\n> \n> diff --git a/builtin/gc.c b/builtin/gc.c\n> index 2ba50a287..79480124a 100644\n> --- a/builtin/gc.c\n> +++ b/builtin/gc.c\n> @@ -414,6 +414,9 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n>  \t\t\tif (report_last_gc_error())\n>  \t\t\t\treturn -1;\n>  \n> +\t\t\tif (lock_repo_for_gc(force, &pid))\n> +\t\t\t\treturn 0;\n> +\n>  \t\t\tif (gc_before_repack())\n>  \t\t\t\treturn -1;\n>  \t\t\t/*\n> \n> means that anybody looking at the lockfile will report the wrong pid\n> (and thus think the lock is invalid). I guess we'd need to update it in\n> place after daemonizing.\n\nUpdating it in place is a bit tricky. I came up with this hack that\nmakes it work, but I'm not sure if the reasoning is too gross.\n\n-- >8 --\nSubject: [PATCH] gc: run pre-detach operations under lock\n\nWe normally try to avoid having two auto-gc operations run\nat the same time, because it wastes resources. This was done\nlong ago in 64a99eb47 (gc: reject if another gc is running,\nunless --force is given, 2013-08-08).\n\nWhen we do a detached auto-gc, we run the ref-related\ncommands _before_ detaching, to avoid confusing lock\ncontention. This was done by 62aad1849 (gc --auto: do not\nlock refs in the background, 2014-05-25).\n\nThese two features do not interact well. The pre-detach\noperations are run before we check the gc.pid lock, meaning\nthat on a busy repository we may run many of them\nconcurrently. Ideally we'd take the lock before spawning any\noperations, and hold it for the duration of the program.\n\nThis is tricky, though, with the way the pid-file interacts\nwith the daemonize() process.  Other processes will check\nthat the pid recorded in the pid-file still exists. But\ndetaching causes us to fork and continue running under a\nnew pid. So if we take the lock before detaching, the\npid-file will have a bogus pid in it. We'd have to go back\nand update it with the new pid after detaching. We'd also\nhave to play some tricks with the tempfile subsystem to\ntweak the \"owner\" field, so that the parent process does not\nclean it up on exit, but the child process does.\n\nInstead, we can do something a bit simpler: take the lock\nonly for the duration of the pre-detach work, then detach,\nthen take it again for the post-detach work. Technically,\nthis means that the post-detach lock could lose to another\nprocess doing pre-detach work. But in the long run this\nworks out.\n\nThat second process would then follow-up by doing\npost-detach work. Unless it was in turn blocked by a third\nprocess doing pre-detach work, and so on. This could in\ntheory go on indefinitely, as the pre-detach work does not\nrepack, and so need_to_gc() will continue to trigger.  But\nin each round we are racing between the pre- and post-detach\nlocks. Eventually, one of the post-detach locks will win the\nrace and complete the full gc. So in the worst case, we may\nracily repeat the pre-detach work, but we would never do so\nsimultaneously (it would happen via a sequence of serialized\nrace-wins).\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n builtin/gc.c  |  4 ++++\n t/t6500-gc.sh | 21 +++++++++++++++++++++\n 2 files changed, 25 insertions(+)\n\ndiff --git a/builtin/gc.c b/builtin/gc.c\nindex bd91f136f..5a535040f 100644\n--- a/builtin/gc.c\n+++ b/builtin/gc.c\n@@ -414,8 +414,12 @@ int cmd_gc(int argc, const char **argv, const char *prefix)\n \t\t\tif (report_last_gc_error())\n \t\t\t\treturn -1;\n \n+\t\t\tif (lock_repo_for_gc(force, &pid))\n+\t\t\t\treturn 0;\n \t\t\tif (gc_before_repack())\n \t\t\t\treturn -1;\n+\t\t\tdelete_tempfile(&pidfile);\n+\n \t\t\t/*\n \t\t\t * failure to daemonize is ok, we'll continue\n \t\t\t * in foreground\ndiff --git a/t/t6500-gc.sh b/t/t6500-gc.sh\nindex cc7acd101..41b0be575 100755\n--- a/t/t6500-gc.sh\n+++ b/t/t6500-gc.sh\n@@ -95,6 +95,27 @@ test_expect_success 'background auto gc does not run if gc.log is present and re\n \ttest_line_count = 1 packs\n '\n \n+test_expect_success 'background auto gc respects lock for all operations' '\n+\t# make sure we run a background auto-gc\n+\ttest_commit make-pack &&\n+\tgit repack &&\n+\ttest_config gc.autopacklimit 1 &&\n+\ttest_config gc.autodetach true &&\n+\n+\t# create a ref whose loose presence we can use to detect a pack-refs run\n+\tgit update-ref refs/heads/should-be-loose HEAD &&\n+\ttest_path_is_file .git/refs/heads/should-be-loose &&\n+\n+\t# now fake a concurrent gc that holds the lock; we can use our\n+\t# shell pid so that it looks valid.\n+\thostname=$(hostname || echo unknown) &&\n+\tprintf \"$$ %s\" \"$hostname\" >.git/gc.pid &&\n+\n+\t# our gc should exit zero without doing anything\n+\trun_and_wait_for_auto_gc &&\n+\ttest_path_is_file .git/refs/heads/should-be-loose\n+'\n+\n # DO NOT leave a detached auto gc process running near the end of the\n # test script: it can run long enough in the background to racily\n # interfere with the cleanup in 'test_done'.\n-- \n2.13.2.1149.g835628f6b\n\n"},{"id":"324267","messageId":"xmqqvamx1u3i.fsf@gitster.mtv.corp.google.com","threadId":"46311","inReplyTo":"20170711090635.swowex7yry7kqb7v@sigill.intra.peff.net","subject":"Re: [PATCH] gc: run pre-detach operations under lock","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-07-12T16:46:25Z","receivedAt":"2017-07-12T16:46:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Instead, we can do something a bit simpler: take the lock\n> only for the duration of the pre-detach work, then detach,\n> then take it again for the post-detach work. Technically,\n> this means that the post-detach lock could lose to another\n> process doing pre-detach work. But in the long run this\n> works out.\n\nYou might have found this part gross, but I actually don't.  It\nlooks like a reasonable practical compromise, and I tried to think\nof a scenario that this would do a wrong thing but I didn't---it is\nnot like we carry information off-disk from the pre-detach to\npost-detach work to cause the latter make decisions on it, so this\n\"split into two phrases\" looks fairly safe.\n\n"},{"id":"324268","messageId":"20170712165817.xcq4we5ynl3opm37@sigill.intra.peff.net","threadId":"46311","inReplyTo":"xmqqvamx1u3i.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH] gc: run pre-detach operations under lock","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-07-12T16:58:17Z","receivedAt":"2017-07-12T16:58:24Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Jul 12, 2017 at 09:46:25AM -0700, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > Instead, we can do something a bit simpler: take the lock\n> > only for the duration of the pre-detach work, then detach,\n> > then take it again for the post-detach work. Technically,\n> > this means that the post-detach lock could lose to another\n> > process doing pre-detach work. But in the long run this\n> > works out.\n> \n> You might have found this part gross, but I actually don't.  It\n> looks like a reasonable practical compromise, and I tried to think\n> of a scenario that this would do a wrong thing but I didn't---it is\n> not like we carry information off-disk from the pre-detach to\n> post-detach work to cause the latter make decisions on it, so this\n> \"split into two phrases\" looks fairly safe.\n\nAnytime I have to spend a few paragraphs saying \"well, it looks like\nthis might behave terribly, but it doesn't because...\" I get worried\nthat my analysis is missing a case. And that writing it in a way that\navoids that analysis might be safer, even if it's a little more work.\n\nI gave it some more thought after sending the earlier message. And I\nreally think it's not \"a little more work\". Even if we decided to keep\nthe same file and replace the PID in it with the daemonized one, I think\nthat still isn't quite right. Because we don't do so atomically unless\nwe take gc.pid.lock again. But we may actually conflict with somebody\nelse on that! Even though that somebody would just pick up the lock,\nread gc.pid and say \"well, looks like somebody else is running\" and\nrelease it again. So we'd have to either hold the lock the whole time,\nor do some kind of retry loop to race with other processes picking up\nthe lock.\n\nIt's definitely possible, but it's fighting an uphill battle against the\nway our locking and tempfile code works. So I came to the conclusion\nthat it's not worth the trouble, and what I posted is probably a good\ncompromise.\n\n-Peff\n"},{"id":"324303","messageId":"xmqq8tjt1hw1.fsf@gitster.mtv.corp.google.com","threadId":"46311","inReplyTo":"20170712165817.xcq4we5ynl3opm37@sigill.intra.peff.net","subject":"Re: [PATCH] gc: run pre-detach operations under lock","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-07-12T21:10:06Z","receivedAt":"2017-07-12T21:10:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> ... And I\n> really think it's not \"a little more work\". Even if we decided to keep\n> the same file and replace the PID in it with the daemonized one, I think\n> that still isn't quite right. Because we don't do so atomically unless\n> we take gc.pid.lock again.\n\nExactly.\n"}]}