{"thread":{"id":"65847","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","startedAt":"2026-06-20T15:33:26Z","lastAt":"2026-07-06T11:37:49Z","messageCount":42,"participants":["Michael Montalbo","Jeff King","Patrick Steinhardt","Junio C Hamano","Kristofer Karlsson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"546035","messageId":"CAC2Qwm+9sh=ks1fuux415JGdDJ38Jq6eZrSH7-qzQxYCoy+Aug@mail.gmail.com","threadId":"65847","inReplyTo":null,"subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Michael Montalbo","fromEmail":"mmontalbo@gmail.com","sentAt":"2026-06-20T15:33:13Z","receivedAt":"2026-06-20T15:33:26Z","isPatch":false,"body":"Patrick Steinhardt <ps@pks.im> writes:\n> So I strongly suspect that it most be one of the t555* tests.\n> [...]\n> Maybe this is something that's specific to GitHub's environment...\n\nI think you're right it's t5551/t5559. The runs Junio linked:\n\n  osx-clang     cancelled  360min\n  osx-gcc       cancelled  360min\n  osx-reftable  success     35min\n  osx-meson     success     61min\n\nAll four run the same t5551/t5559 under EXPENSIVE. The two that\nfinished differ in just two ways, which look like the levers:\nosx-reftable generates the 100k-ref advertisement in ~24ms vs ~1.2s\nfor loose refs on macOS (so much less time mid-response), and\nosx-meson runs tests at nproc while the prove jobs hardcode --jobs=10\non a 3-core runner (over recent master/next the prove jobs hang ~40%,\nmeson ~10%).\n\nWhen it is wedged the whole chain sits at 0% CPU. upload-pack is\nblocked in write() on the ls-refs advertisement, curl blocked in\nselect(). So it looks like an HTTP/2 flow-control stall on the\nresponse side. The same stall resets itself after ~60-85s on my Linux\nbox and on a bare-metal Mac, but not on the GitHub runner; I haven't\npinned down why yet.\n\nOn the chance those two levers are the fix, a branch off master:\n\n  https://github.com/mmontalbo/git/tree/mm/macos-ci-hang-fix\n\n  - pack the refs in t5551's enormous-ref-negotiation test (doesn't\n    change what it checks on the wire, just avoids re-reading 100k loose\n    files to advertise them, like reftable already does)\n  - use the core count for $JOBS on the GitHub macOS path, matching the\n    GitLab branch in the same ci/lib.sh and what meson does\n\nI ran the two macOS jobs under EXPENSIVE about eight times with these\nand they all finished in ~30-44min instead of hanging. Happy to send\nout a patch if it's helpful.\n"},{"id":"546092","messageId":"20260621213407.GC2297179@coredump.intra.peff.net","threadId":"65847","inReplyTo":"CAC2Qwm+9sh=ks1fuux415JGdDJ38Jq6eZrSH7-qzQxYCoy+Aug@mail.gmail.com","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-21T21:34:07Z","receivedAt":"2026-06-21T21:34:09Z","isPatch":false,"body":"On Sat, Jun 20, 2026 at 08:33:13AM -0700, Michael Montalbo wrote:\n\n> Patrick Steinhardt <ps@pks.im> writes:\n> > So I strongly suspect that it most be one of the t555* tests.\n> > [...]\n> > Maybe this is something that's specific to GitHub's environment...\n> \n> I think you're right it's t5551/t5559. The runs Junio linked:\n> \n>   osx-clang     cancelled  360min\n>   osx-gcc       cancelled  360min\n>   osx-reftable  success     35min\n>   osx-meson     success     61min\n> \n> All four run the same t5551/t5559 under EXPENSIVE. The two that\n> finished differ in just two ways, which look like the levers:\n> osx-reftable generates the 100k-ref advertisement in ~24ms vs ~1.2s\n> for loose refs on macOS (so much less time mid-response), and\n> osx-meson runs tests at nproc while the prove jobs hardcode --jobs=10\n> on a 3-core runner (over recent master/next the prove jobs hang ~40%,\n> meson ~10%).\n\nIf the problem is a racy deadlock, there is a reasonable chance that\nsome jobs may simply be lucky. Even if things like packing refs help, I\nsuspect the problem may still be lurking. Maybe I'm just a pessimist,\nthough. ;)\n\n> When it is wedged the whole chain sits at 0% CPU. upload-pack is\n> blocked in write() on the ls-refs advertisement, curl blocked in\n> select(). So it looks like an HTTP/2 flow-control stall on the\n> response side. The same stall resets itself after ~60-85s on my Linux\n> box and on a bare-metal Mac, but not on the GitHub runner; I haven't\n> pinned down why yet.\n\nWe had some HTTP/2 stalls/deadlocks in the past, and they were dependent\non libcurl and apache (actually h2_mod) versions. IIRC some of the\nnon-TLS code paths for HTTP/2 were not well tested, which led to\n8f2146dbf1 (t5559: make SSL/TLS the default, 2023-02-23). Of course\nafter that commit those cleartext code paths should not be a problem, so\nthat is probably not exactly the issue now.\n\nBut it might be worth checking the versions you're running locally\nversus what's in the GitHub runner.\n\n-Peff\n"},{"id":"546102","messageId":"aji9MOE-NTHKXYqn@pks.im","threadId":"65847","inReplyTo":"20260621213407.GC2297179@coredump.intra.peff.net","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-22T04:42:24Z","receivedAt":"2026-06-22T04:42:33Z","isPatch":false,"body":"On Sun, Jun 21, 2026 at 05:34:07PM -0400, Jeff King wrote:\n> On Sat, Jun 20, 2026 at 08:33:13AM -0700, Michael Montalbo wrote:\n> \n> > Patrick Steinhardt <ps@pks.im> writes:\n> > > So I strongly suspect that it most be one of the t555* tests.\n> > > [...]\n> > > Maybe this is something that's specific to GitHub's environment...\n> > \n> > I think you're right it's t5551/t5559. The runs Junio linked:\n> > \n> >   osx-clang     cancelled  360min\n> >   osx-gcc       cancelled  360min\n> >   osx-reftable  success     35min\n> >   osx-meson     success     61min\n> > \n> > All four run the same t5551/t5559 under EXPENSIVE. The two that\n> > finished differ in just two ways, which look like the levers:\n> > osx-reftable generates the 100k-ref advertisement in ~24ms vs ~1.2s\n> > for loose refs on macOS (so much less time mid-response), and\n> > osx-meson runs tests at nproc while the prove jobs hardcode --jobs=10\n> > on a 3-core runner (over recent master/next the prove jobs hang ~40%,\n> > meson ~10%).\n> \n> If the problem is a racy deadlock, there is a reasonable chance that\n> some jobs may simply be lucky. Even if things like packing refs help, I\n> suspect the problem may still be lurking. Maybe I'm just a pessimist,\n> though. ;)\n\nI had the same thought.\n\n> > When it is wedged the whole chain sits at 0% CPU. upload-pack is\n> > blocked in write() on the ls-refs advertisement, curl blocked in\n> > select(). So it looks like an HTTP/2 flow-control stall on the\n> > response side. The same stall resets itself after ~60-85s on my Linux\n> > box and on a bare-metal Mac, but not on the GitHub runner; I haven't\n> > pinned down why yet.\n> \n> We had some HTTP/2 stalls/deadlocks in the past, and they were dependent\n> on libcurl and apache (actually h2_mod) versions. IIRC some of the\n> non-TLS code paths for HTTP/2 were not well tested, which led to\n> 8f2146dbf1 (t5559: make SSL/TLS the default, 2023-02-23). Of course\n> after that commit those cleartext code paths should not be a problem, so\n> that is probably not exactly the issue now.\n> \n> But it might be worth checking the versions you're running locally\n> versus what's in the GitHub runner.\n\nI didn't observe any similar hangs in GitLab's CI systems, so I wonder\nwhether this is because of different versions of curl. And indeed we use\ndifferent versions:\n\n  - On GitHub we use 8.6.0.\n\n  - On GitLab we use 8.7.1.\n\nNow this of course doesn't mean that updating the curl version is the\nfix to this whole issue, as there's a ton of other factors that could\nplay a role in whether or not the test hangs. So while we could just\nupgrade parts of the stack and cross our fingers, but that feels rather\nunsatisfactory. Still, one place to start could be to update our build\nimages to macOS 15.\n\nBut the big question to me is whether the hang is because of a bug in\nGit with how we drive curl, a bug in curl itself, or a bug in Apache.\n\nPatrick\n"},{"id":"546105","messageId":"xmqqqzlz412x.fsf@gitster.g","threadId":"65847","inReplyTo":"20260621213407.GC2297179@coredump.intra.peff.net","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-22T05:05:10Z","receivedAt":"2026-06-22T05:05:12Z","isPatch":false,"body":"Jeff King <peff@peff.net> writes:\n\n> If the problem is a racy deadlock, there is a reasonable chance that\n> some jobs may simply be lucky. Even if things like packing refs help, I\n> suspect the problem may still be lurking. Maybe I'm just a pessimist,\n> though. ;)\n\nI share the pessimism X-<.\n\n> We had some HTTP/2 stalls/deadlocks in the past, and they were dependent\n> on libcurl and apache (actually h2_mod) versions. IIRC some of the\n> non-TLS code paths for HTTP/2 were not well tested, which led to\n> 8f2146dbf1 (t5559: make SSL/TLS the default, 2023-02-23). Of course\n> after that commit those cleartext code paths should not be a problem, so\n> that is probably not exactly the issue now.\n>\n> But it might be worth checking the versions you're running locally\n> versus what's in the GitHub runner.\n\nTrue.\n"},{"id":"546156","messageId":"ajkEzhdqzmAePk_P@pks.im","threadId":"65847","inReplyTo":"aji9MOE-NTHKXYqn@pks.im","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-22T09:47:58Z","receivedAt":"2026-06-22T09:48:06Z","isPatch":false,"body":"On Mon, Jun 22, 2026 at 06:42:24AM +0200, Patrick Steinhardt wrote:\n> On Sun, Jun 21, 2026 at 05:34:07PM -0400, Jeff King wrote:\n> > On Sat, Jun 20, 2026 at 08:33:13AM -0700, Michael Montalbo wrote:\n[snip]\n> > > When it is wedged the whole chain sits at 0% CPU. upload-pack is\n> > > blocked in write() on the ls-refs advertisement, curl blocked in\n> > > select(). So it looks like an HTTP/2 flow-control stall on the\n> > > response side. The same stall resets itself after ~60-85s on my Linux\n> > > box and on a bare-metal Mac, but not on the GitHub runner; I haven't\n> > > pinned down why yet.\n> > \n> > We had some HTTP/2 stalls/deadlocks in the past, and they were dependent\n> > on libcurl and apache (actually h2_mod) versions. IIRC some of the\n> > non-TLS code paths for HTTP/2 were not well tested, which led to\n> > 8f2146dbf1 (t5559: make SSL/TLS the default, 2023-02-23). Of course\n> > after that commit those cleartext code paths should not be a problem, so\n> > that is probably not exactly the issue now.\n> > \n> > But it might be worth checking the versions you're running locally\n> > versus what's in the GitHub runner.\n> \n> I didn't observe any similar hangs in GitLab's CI systems, so I wonder\n> whether this is because of different versions of curl. And indeed we use\n> different versions:\n> \n>   - On GitHub we use 8.6.0.\n> \n>   - On GitLab we use 8.7.1.\n> \n> Now this of course doesn't mean that updating the curl version is the\n> fix to this whole issue, as there's a ton of other factors that could\n> play a role in whether or not the test hangs. So while we could just\n> upgrade parts of the stack and cross our fingers, but that feels rather\n> unsatisfactory. Still, one place to start could be to update our build\n> images to macOS 15.\n> \n> But the big question to me is whether the hang is because of a bug in\n> Git with how we drive curl, a bug in curl itself, or a bug in Apache.\n\nI noticed that a osx-clang job failed today in t5551 [1]. This time it\ndidn't hang, but produced an actual error:\n\n    2026-06-22T09:25:45.1984230Z ++ git -C too-many-refs fetch -q --tags\n    2026-06-22T09:25:45.1984420Z error: RPC failed; curl 18 transfer closed with outstanding read data remaining\n    2026-06-22T09:25:45.1984520Z fatal: expected flush after ref listing\n    2026-06-22T09:25:45.1984610Z error: last command exited with $?=128\n    2026-06-22T09:25:45.1984660Z ++ rm -f tags\n    2026-06-22T09:25:45.1984710Z ++ :\n    2026-06-22T09:25:45.1984830Z not ok 35 - http can handle enormous ref negotiation\n\nThere was a second test failing similarly.\n\nPatrick\n\n[1]: https://github.com/git/git/actions/runs/27940620478/job/82672854726\n"},{"id":"546157","messageId":"ajkGkB2ckf3p43QR@pks.im","threadId":"65847","inReplyTo":"ajkEzhdqzmAePk_P@pks.im","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-22T09:55:28Z","receivedAt":"2026-06-22T09:55:35Z","isPatch":false,"body":"On Mon, Jun 22, 2026 at 11:48:01AM +0200, Patrick Steinhardt wrote:\n> On Mon, Jun 22, 2026 at 06:42:24AM +0200, Patrick Steinhardt wrote:\n> > On Sun, Jun 21, 2026 at 05:34:07PM -0400, Jeff King wrote:\n> > > On Sat, Jun 20, 2026 at 08:33:13AM -0700, Michael Montalbo wrote:\n> [snip]\n> > > > When it is wedged the whole chain sits at 0% CPU. upload-pack is\n> > > > blocked in write() on the ls-refs advertisement, curl blocked in\n> > > > select(). So it looks like an HTTP/2 flow-control stall on the\n> > > > response side. The same stall resets itself after ~60-85s on my Linux\n> > > > box and on a bare-metal Mac, but not on the GitHub runner; I haven't\n> > > > pinned down why yet.\n> > > \n> > > We had some HTTP/2 stalls/deadlocks in the past, and they were dependent\n> > > on libcurl and apache (actually h2_mod) versions. IIRC some of the\n> > > non-TLS code paths for HTTP/2 were not well tested, which led to\n> > > 8f2146dbf1 (t5559: make SSL/TLS the default, 2023-02-23). Of course\n> > > after that commit those cleartext code paths should not be a problem, so\n> > > that is probably not exactly the issue now.\n> > > \n> > > But it might be worth checking the versions you're running locally\n> > > versus what's in the GitHub runner.\n> > \n> > I didn't observe any similar hangs in GitLab's CI systems, so I wonder\n> > whether this is because of different versions of curl. And indeed we use\n> > different versions:\n> > \n> >   - On GitHub we use 8.6.0.\n> > \n> >   - On GitLab we use 8.7.1.\n> > \n> > Now this of course doesn't mean that updating the curl version is the\n> > fix to this whole issue, as there's a ton of other factors that could\n> > play a role in whether or not the test hangs. So while we could just\n> > upgrade parts of the stack and cross our fingers, but that feels rather\n> > unsatisfactory. Still, one place to start could be to update our build\n> > images to macOS 15.\n> > \n> > But the big question to me is whether the hang is because of a bug in\n> > Git with how we drive curl, a bug in curl itself, or a bug in Apache.\n> \n> I noticed that a osx-clang job failed today in t5551 [1]. This time it\n> didn't hang, but produced an actual error:\n> \n>     2026-06-22T09:25:45.1984230Z ++ git -C too-many-refs fetch -q --tags\n>     2026-06-22T09:25:45.1984420Z error: RPC failed; curl 18 transfer closed with outstanding read data remaining\n>     2026-06-22T09:25:45.1984520Z fatal: expected flush after ref listing\n>     2026-06-22T09:25:45.1984610Z error: last command exited with $?=128\n>     2026-06-22T09:25:45.1984660Z ++ rm -f tags\n>     2026-06-22T09:25:45.1984710Z ++ :\n>     2026-06-22T09:25:45.1984830Z not ok 35 - http can handle enormous ref negotiation\n> \n> There was a second test failing similarly.\n\nOh, and Linux is also failing in the same test suite [1], even though\nthe job logs are truncated, so it's hard to say whether it's the same\nfailure or not.\n\nThere certainly seems to be a deeper issue here. We could of course just\ndisable the test again, but by now I do wonder whether this would paper\nover an actual bug.\n\nPatrick\n\n[1]: https://github.com/git/git/actions/runs/27940620478/job/82672854864\n"},{"id":"546161","messageId":"ajkOoRhqaAcy6gBg@pks.im","threadId":"65847","inReplyTo":"ajkGkB2ckf3p43QR@pks.im","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-22T10:29:53Z","receivedAt":"2026-06-22T10:30:00Z","isPatch":false,"body":"On Mon, Jun 22, 2026 at 11:55:31AM +0200, Patrick Steinhardt wrote:\n> On Mon, Jun 22, 2026 at 11:48:01AM +0200, Patrick Steinhardt wrote:\n> > On Mon, Jun 22, 2026 at 06:42:24AM +0200, Patrick Steinhardt wrote:\n> > > On Sun, Jun 21, 2026 at 05:34:07PM -0400, Jeff King wrote:\n> > > > On Sat, Jun 20, 2026 at 08:33:13AM -0700, Michael Montalbo wrote:\n> > [snip]\n> > > > > When it is wedged the whole chain sits at 0% CPU. upload-pack is\n> > > > > blocked in write() on the ls-refs advertisement, curl blocked in\n> > > > > select(). So it looks like an HTTP/2 flow-control stall on the\n> > > > > response side. The same stall resets itself after ~60-85s on my Linux\n> > > > > box and on a bare-metal Mac, but not on the GitHub runner; I haven't\n> > > > > pinned down why yet.\n> > > > \n> > > > We had some HTTP/2 stalls/deadlocks in the past, and they were dependent\n> > > > on libcurl and apache (actually h2_mod) versions. IIRC some of the\n> > > > non-TLS code paths for HTTP/2 were not well tested, which led to\n> > > > 8f2146dbf1 (t5559: make SSL/TLS the default, 2023-02-23). Of course\n> > > > after that commit those cleartext code paths should not be a problem, so\n> > > > that is probably not exactly the issue now.\n> > > > \n> > > > But it might be worth checking the versions you're running locally\n> > > > versus what's in the GitHub runner.\n> > > \n> > > I didn't observe any similar hangs in GitLab's CI systems, so I wonder\n> > > whether this is because of different versions of curl. And indeed we use\n> > > different versions:\n> > > \n> > >   - On GitHub we use 8.6.0.\n> > > \n> > >   - On GitLab we use 8.7.1.\n> > > \n> > > Now this of course doesn't mean that updating the curl version is the\n> > > fix to this whole issue, as there's a ton of other factors that could\n> > > play a role in whether or not the test hangs. So while we could just\n> > > upgrade parts of the stack and cross our fingers, but that feels rather\n> > > unsatisfactory. Still, one place to start could be to update our build\n> > > images to macOS 15.\n> > > \n> > > But the big question to me is whether the hang is because of a bug in\n> > > Git with how we drive curl, a bug in curl itself, or a bug in Apache.\n> > \n> > I noticed that a osx-clang job failed today in t5551 [1]. This time it\n> > didn't hang, but produced an actual error:\n> > \n> >     2026-06-22T09:25:45.1984230Z ++ git -C too-many-refs fetch -q --tags\n> >     2026-06-22T09:25:45.1984420Z error: RPC failed; curl 18 transfer closed with outstanding read data remaining\n> >     2026-06-22T09:25:45.1984520Z fatal: expected flush after ref listing\n> >     2026-06-22T09:25:45.1984610Z error: last command exited with $?=128\n> >     2026-06-22T09:25:45.1984660Z ++ rm -f tags\n> >     2026-06-22T09:25:45.1984710Z ++ :\n> >     2026-06-22T09:25:45.1984830Z not ok 35 - http can handle enormous ref negotiation\n> > \n> > There was a second test failing similarly.\n> \n> Oh, and Linux is also failing in the same test suite [1], even though\n> the job logs are truncated, so it's hard to say whether it's the same\n> failure or not.\n> \n> There certainly seems to be a deeper issue here. We could of course just\n> disable the test again, but by now I do wonder whether this would paper\n> over an actual bug.\n> \n> Patrick\n> \n> [1]: https://github.com/git/git/actions/runs/27940620478/job/82672854864\n\nSorry for the repeated spam.\n\nI think the issue is rather simple: we're hitting timeouts in Apache. If\nyou apply the following diff:\n\ndiff --git a/t/lib-httpd/apache.conf b/t/lib-httpd/apache.conf\nindex 40a690b0bb..4054fe008f 100644\n--- a/t/lib-httpd/apache.conf\n+++ b/t/lib-httpd/apache.conf\n@@ -302,3 +302,5 @@ RewriteRule ^/half-auth-complete/ - [E=AUTHREQUIRED:yes]\n \t\tSVNPath \"${LIB_HTTPD_SVNPATH}\"\n \t</Location>\n </IfDefine>\n+\n+Timeout 1\n\nThen you'll see the same errors locally:\n\n    $ GIT_TEST_LONG=Yes meson test t5551-http-fetch-smart --test-args=-ix -i\n    Failed to clone 'sub'. Retry scheduled\n    Cloning into '/home/pks/Development/git/build/test-output/trash directory.t5551-http-fetch-smart/sub'...\n    error: RPC failed; curl 18 transfer closed with outstanding read data remaining\n    fatal: early EOF\n    fatal: fetch-pack: invalid index-pack output\n    fatal: clone of 'http://127.0.0.1:5551/smart_headers/repo.git' into submodule path '/home/pks/Development/git/build/test-output/trash directory.t5551-http-fetch-smart/sub' failed\n    Failed to clone 'sub' a second time, aborting\n    error: last command exited with $?=1\n    not ok 36 - custom http headers\n    #\t\n    #\t\ttest_must_fail git -c http.extraheader=\"x-magic-two: cadabra\" \\\n    #\t\t\tfetch \"$HTTPD_URL/smart_headers/repo.git\" &&\n    #\t\tgit -c http.extraheader=\"x-magic-one: abra\" \\\n    #\t\t    -c http.extraheader=\"x-magic-two: cadabra\" \\\n    #\t\t    fetch \"$HTTPD_URL/smart_headers/repo.git\" &&\n    #\t\tgit update-index --add --cacheinfo 160000,$(git rev-parse HEAD),sub &&\n    #\t\tgit config -f .gitmodules submodule.sub.path sub &&\n    #\t\tgit config -f .gitmodules submodule.sub.url \\\n    #\t\t\t\"$HTTPD_URL/smart_headers/repo.git\" &&\n    #\t\tgit submodule init sub &&\n    #\t\ttest_must_fail git submodule update sub &&\n    #\t\tgit -c http.extraheader=\"x-magic-one: abra\" \\\n    #\t\t    -c http.extraheader=\"x-magic-two: cadabra\" \\\n    #\t\t\tsubmodule update sub\n    #\t\n    1..36\n\nAnd Apache also logs this as a timeout:\n\n    [Mon Jun 22 10:26:52.115717 2026] [cgi:warn] [pid 3686957:tid 3686957] [client 127.0.0.1:55114] AH01220: Timeout waiting for output from CGI script /home/pks/Development/git/build/git-http-backend\n    [Mon Jun 22 10:26:52.115748 2026] [core:error] [pid 3686957:tid 3686957] (70007)The timeout specified has expired: [client 127.0.0.1:55114] AH00574: ap_content_length_filter: apr_bucket_read() failed\n    [Mon Jun 22 10:27:01.567533 2026] [cgi:warn] [pid 3686958:tid 3686958] [client 127.0.0.1:54384] AH01220: Timeout waiting for output from CGI script /home/pks/Development/git/build/git-http-backend\n    [Mon Jun 22 10:27:01.567559 2026] [core:error] [pid 3686958:tid 3686958] (70007)The timeout specified has expired: [client 127.0.0.1:54384] AH00574: ap_content_length_filter: apr_bucket_read() failed\n\nThis is because our keepalive mechanisms aren't helping:\n\n  - The TCP-level keepalives don't help with Apache.\n\n  - The application-level sideband keepalives don't apply to the\n    \"ls-refs\" endpoint.\n\nWhether that's the same issue like we see in macOS sometimes is a\ndifferent question.\n\nPatrick\n"},{"id":"546438","messageId":"CAC2QwmJA2TH6BmO0O61qRYvV2pqURUk0dTXpkJtb9e-TZNZDZQ@mail.gmail.com","threadId":"65847","inReplyTo":"ajkOoRhqaAcy6gBg@pks.im","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Michael Montalbo","fromEmail":"mmontalbo@gmail.com","sentAt":"2026-06-26T03:27:35Z","receivedAt":"2026-06-26T03:27:48Z","isPatch":false,"body":"Patrick Steinhardt <ps@pks.im> writes:\n> I think the issue is rather simple: we're hitting timeouts in Apache.\n> [...]\n> This is because our keepalive mechanisms aren't helping [...]\n> Whether that's the same issue like we see in macOS sometimes is a\n> different question.\n\nI think that is the trigger for issues we've been seeing. I spent\nsome time investigating the Apache side over the last week and maybe\nfound a mod_http2 bug, which I filed upstream with a potential fix:\n\n  bug:  https://bz.apache.org/bugzilla/show_bug.cgi?id=70131\n  fix:  https://github.com/mmontalbo/httpd/pull/2\n\nTo Patrick's earlier question of whether this is a Git, curl, or Apache\nbug: as best I can tell it's Apache. I could reproduce it with no Git\ninvolved at all (just Apache and a small CGI that goes quiet past the\nTimeout), and across several curl versions (8.6.0, which is what the\nGitHub runners use, up to 8.20.0), so I don't think bumping curl would\nhelp. It also seems to wear two faces from the same trigger: over\nHTTP/1.1 Apache closes the connection and curl bails with the\n\"transfer closed\" error (which looks like what you hit with Timeout=1,\nand the recent failures on both macOS and Linux), and over HTTP/2 it\ndoes not reliably reset the stream, so the client just waits, which is\nthe six-hour macOS hang. I share the pessimism from earlier in the\nthread, though: I think the real fix is upstream in Apache, and\nanything we do on our side mostly just bounds the symptom in the\nmeantime.\n\nGiven there could be a potential reliability issue with an upstream\ndependency like Apache, I was considering what mitigation strategies\nmight help:\n\n  - Enforce some kind of lower bound speed limit and a client-side\n    timeout so runs that wedge fail fast (and loudly) instead of\n    hanging.\n\n  - Potentially provide some affordance for retrying flaky tests\n    that might fail due to upstream dependencies. Git already has\n    some HTTP retry support (http.maxRetries and friends, added\n    recently), but as far as I can tell it only triggers on HTTP 429\n    rate limiting, so it would not catch a stall like this on its\n    own. A test-level retry is not something I like that much, since\n    it might encourage papering over flakiness that should be\n    resolved, but it was a consideration vs requiring a fresh CI run\n    to resolve the flake.\n\n  - Make slow tests faster by optimizing the test itself and/or\n    the test runner configuration (e.g., job number matching\n    cores) so wedges become less likely.\n\nFor the first one, I think Git already provides some affordances. There\nis a stall-based timeout that just ships disabled: as I understand it\nhttp.lowSpeedLimit sets a bytes/sec floor and http.lowSpeedTime how long\na transfer can sit below it before curl gives up, so it would catch a\nwedged connection without punishing one that is just slow. Enabling it\nfor the http tests might look something like:\n\n    diff --git a/t/lib-httpd.sh b/t/lib-httpd.sh\n    @@ GIT_TRACE=$GIT_TRACE; export GIT_TRACE\n    +# Abort a transfer that makes essentially no progress for a while,\n    +# so a wedged connection fails in seconds instead of hanging to the\n    +# job cap. Tiny limit, generous window, so it only trips on a true\n    +# stall; override either var, or set the limit to 0, to disable.\n    +GIT_HTTP_LOW_SPEED_LIMIT=${GIT_HTTP_LOW_SPEED_LIMIT-1}\n    +GIT_HTTP_LOW_SPEED_TIME=${GIT_HTTP_LOW_SPEED_TIME-60}\n    +export GIT_HTTP_LOW_SPEED_LIMIT GIT_HTTP_LOW_SPEED_TIME\n\nI went conservative on the values on purpose: a floor of 1 byte/sec\nshould only really fire on a true zero-progress stall, not on something\nthat is just crawling on a slow runner, and the 60s window is generous\nfor the same reason. When I tried it locally against a stall-proxy it\ndid turn an otherwise indefinite hang into a bounded abort (a tighter\nlimit/window brings that down to single-digit seconds). It probably does\nnot need to be suite-wide either; it could be scoped per-command with\ngit -c, which the http tests already lean on for this kind of thing\n(t5551 passes http.postbuffer and http.extraheader that way), if a\nnarrower blast radius feels safer.\n\nI only dug into the first option in any depth, since I wanted to\nsanity-check the direction before writing patches. Does turning on a\nstall timeout for the http tests seem reasonable? Are there other\nstrategies that we should implement?\n"},{"id":"546440","messageId":"20260626051657.GB3138423@coredump.intra.peff.net","threadId":"65847","inReplyTo":"CAC2QwmJA2TH6BmO0O61qRYvV2pqURUk0dTXpkJtb9e-TZNZDZQ@mail.gmail.com","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-26T05:16:57Z","receivedAt":"2026-06-26T05:17:00Z","isPatch":false,"body":"On Thu, Jun 25, 2026 at 08:27:35PM -0700, Michael Montalbo wrote:\n\n> I think that is the trigger for issues we've been seeing. I spent\n> some time investigating the Apache side over the last week and maybe\n> found a mod_http2 bug, which I filed upstream with a potential fix:\n> \n>   bug:  https://bz.apache.org/bugzilla/show_bug.cgi?id=70131\n>   fix:  https://github.com/mmontalbo/httpd/pull/2\n\nThanks both of you for digging into this. I'm not familiar enough with\nApache's code to pass confident judgement, but your findings certainly\nconvinced me that this is just an apache bug.\n\n> Given there could be a potential reliability issue with an upstream\n> dependency like Apache, I was considering what mitigation strategies\n> might help:\n> [...]\n\nDepending on how widespread the Apache bug is, another option might just\nbe: do nothing and wait for it to get fixed.\n\nTrying to make the wedged state fail fast and loudly is mostly just\npunting on the problem. We'd still see spurious failures. We've so far\nresisted the urge to do any automatic flaky-test retries, preferring\ninstead to just try to root out the flakes. I'm a little hesitant to\nstart now, because I think our strategy has mostly been good so far, and\nI've seen some horrible counter-examples where flakes and retries become\na routine drag on development (and I'm afraid that accommodating flakes\nmight make them more common).\n\n>   - Make slow tests faster by optimizing the test itself and/or\n>     the test runner configuration (e.g., job number matching\n>     cores) so wedges become less likely.\n\nIt sounds like the bad state is triggered when Apache hits a timeout,\nand we hit that timeout because the system is slow or busy. We could try\nto make things less slow, but would it work equally well to increase\nthat timeout?\n\n-Peff\n"},{"id":"546453","messageId":"aj5ZaZK7xylfs4Xw@pks.im","threadId":"65847","inReplyTo":"20260626051657.GB3138423@coredump.intra.peff.net","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-26T10:50:17Z","receivedAt":"2026-06-26T10:50:29Z","isPatch":false,"body":"On Fri, Jun 26, 2026 at 01:16:57AM -0400, Jeff King wrote:\n> On Thu, Jun 25, 2026 at 08:27:35PM -0700, Michael Montalbo wrote:\n> \n> > I think that is the trigger for issues we've been seeing. I spent\n> > some time investigating the Apache side over the last week and maybe\n> > found a mod_http2 bug, which I filed upstream with a potential fix:\n> > \n> >   bug:  https://bz.apache.org/bugzilla/show_bug.cgi?id=70131\n> >   fix:  https://github.com/mmontalbo/httpd/pull/2\n> \n> Thanks both of you for digging into this. I'm not familiar enough with\n> Apache's code to pass confident judgement, but your findings certainly\n> convinced me that this is just an apache bug.\n\nThe bug manifests both with HTTP/1.1 and HTTP/2 though, so this wouldn't\nfully fix the flakes we see, right?\n\n> > Given there could be a potential reliability issue with an upstream\n> > dependency like Apache, I was considering what mitigation strategies\n> > might help:\n> > [...]\n> \n> Depending on how widespread the Apache bug is, another option might just\n> be: do nothing and wait for it to get fixed.\n> \n> Trying to make the wedged state fail fast and loudly is mostly just\n> punting on the problem. We'd still see spurious failures. We've so far\n> resisted the urge to do any automatic flaky-test retries, preferring\n> instead to just try to root out the flakes. I'm a little hesitant to\n> start now, because I think our strategy has mostly been good so far, and\n> I've seen some horrible counter-examples where flakes and retries become\n> a routine drag on development (and I'm afraid that accommodating flakes\n> might make them more common).\n\nI agree. I'm not a fan of retry logic, as every flaky test may mask an\nactual bug that we haven't fully investigated yet.\n\n> >   - Make slow tests faster by optimizing the test itself and/or\n> >     the test runner configuration (e.g., job number matching\n> >     cores) so wedges become less likely.\n> \n> It sounds like the bad state is triggered when Apache hits a timeout,\n> and we hit that timeout because the system is slow or busy. We could try\n> to make things less slow, but would it work equally well to increase\n> that timeout?\n\nI was also wondering whether we can maybe work around the issue by\nincreasing the Apache timeout value. That sounds like an easy potential\nsolution to try, and from all we've discovered so far it doesn't feel\nlike this is something we can address on the Git side.\n\nPatrick\n"},{"id":"546469","messageId":"xmqq1pdte7pz.fsf@gitster.g","threadId":"65847","inReplyTo":"aj5ZaZK7xylfs4Xw@pks.im","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-26T13:45:12Z","receivedAt":"2026-06-26T13:45:14Z","isPatch":false,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> Trying to make the wedged state fail fast and loudly is mostly just\n>> punting on the problem. We'd still see spurious failures. We've so far\n>> resisted the urge to do any automatic flaky-test retries, preferring\n>> instead to just try to root out the flakes. I'm a little hesitant to\n>> start now, because I think our strategy has mostly been good so far, and\n>> I've seen some horrible counter-examples where flakes and retries become\n>> a routine drag on development (and I'm afraid that accommodating flakes\n>> might make them more common).\n>\n> I agree. I'm not a fan of retry logic, as every flaky test may mask an\n> actual bug that we haven't fully investigated yet.\n\nCan't agree more.\n\n> I was also wondering whether we can maybe work around the issue by\n> increasing the Apache timeout value. That sounds like an easy potential\n> solution to try, and from all we've discovered so far it doesn't feel\n> like this is something we can address on the Git side.\n\nThanks, all, for looking into this.\n"},{"id":"546521","messageId":"CAC2QwmLkHUymvtYbjY8aQO9_VogvaSXdbb1_DSZtcBttGfN0tg@mail.gmail.com","threadId":"65847","inReplyTo":"aj5ZaZK7xylfs4Xw@pks.im","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Michael Montalbo","fromEmail":"mmontalbo@gmail.com","sentAt":"2026-06-26T23:26:28Z","receivedAt":"2026-06-26T23:26:41Z","isPatch":false,"body":"On Fri, Jun 26, 2026 at 3:50 AM Patrick Steinhardt <ps@pks.im> wrote:\n> The bug manifests both with HTTP/1.1 and HTTP/2 though, so this wouldn't\n> fully fix the flakes we see, right?\n\nYes you are right. The linked fix would just prevent the hanging after timeout\nfor HTTP/2 tests, but still leaves HTTP/1.1 fakes.\n\n> I was also wondering whether we can maybe work around the issue by\n> increasing the Apache timeout value. That sounds like an easy potential\n> solution to try, and from all we've discovered so far it doesn't feel\n> like this is something we can address on the Git side.\n\nI think Peff and Patrick's suggestion to just increase the Apache timeout\nmakes sense. I ran some experiments using a really long timeout with an\nartificially slowed down CI runner and all the jobs made progress\n(if slowly) without stalling, and eventually completed successfully:\n\nhttps://github.com/mmontalbo/git/actions/runs/28267019651\n\nI haven't spent a lot of time trying to figure out what the right timeout\nvalue should be. An hour definitely seems like overkill, with something\non the order of 5-10 minutes seeming more reasonable, but I don't\nhave a principled number.\n"},{"id":"546522","messageId":"20260626234312.GA3156205@coredump.intra.peff.net","threadId":"65847","inReplyTo":"aj5ZaZK7xylfs4Xw@pks.im","subject":"Re: [RFH] Why do osx CI jobs so unreliable?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-26T23:43:12Z","receivedAt":"2026-06-26T23:43:14Z","isPatch":false,"body":"On Fri, Jun 26, 2026 at 12:50:17PM +0200, Patrick Steinhardt wrote:\n\n> > Thanks both of you for digging into this. I'm not familiar enough with\n> > Apache's code to pass confident judgement, but your findings certainly\n> > convinced me that this is just an apache bug.\n> \n> The bug manifests both with HTTP/1.1 and HTTP/2 though, so this wouldn't\n> fully fix the flakes we see, right?\n\nIf I understand the situation correctly, there are really two problems.\n\nFirst, the EXPENSIVE tests in t5551 sometimes trigger a timeout with\nApache's stock settings. This presumably became a problem recently due\nto 7a094d68a2 (ci: run expensive tests on push builds to integration\nbranches, 2026-05-08). The same problem exists in t5559, which just\nwraps t5551 but tells us to use http2.\n\nThis timeout will cause test failures in t5551, because we aren't able\nto complete a request we expected to. Obviously bad and annoying.\n\nThe second problem is that when Apache hits the timeout in HTTP/2 mode,\nit hangs forever. And then the CI job hangs for 6 hours until it's\nkilled, which is an even more annoying failure.\n\nSo the root cause is the same (a timeout), but the effect depends on\nHTTP/1.1 vs HTTP/2. I was able to reproduce both cases on my local\nDebian unstable system by dropping the timeout as you suggested. Running\nt5551 with GIT_TEST_LONG yields a failure, whereas running t5559 yields\na hang.\n\nWe can mitigate both cases by bumping the timeout value, since it's\naddressing the root cause.\n\nThere's an open question of whether this is just papering over a problem\nthat real users might experience, and whether Git should be doing more\nto keep the connection alive. I think it's probably OK to ignore this in\npractice. This is an intentionally large request being served by a very\nunderpowered platform. The default apache timeout is 60s. If a\nreal-world server is seeing ls-refs requests take that long then they\nprobably need to reconsider some other decisions, from ref packing to\nbetter hardware to dropping some users. ;) I don't think trying to\ninsert keepalives at the Git layer here is worth the trouble.\n\nTo give a sense of the time options, here are a few timings from my\nlocal machine, timing \"git upload-pack . </dev/null >/dev/null\" in\nt5551's big repo.git (that's a v0 advertisement, but it should be\nroughly the same work as the v2 ls-refs).\n\n  cold-cache, refs not packed:\n  real\t0m9.973s\n  user\t0m0.354s\n  sys\t0m1.364s\n\n  warm cache, refs not packed:\n  real\t0m0.410s\n  user\t0m0.153s\n  sys\t0m0.257s\n\n  cold-cache, refs packed:\n  real\t0m0.149s\n  user\t0m0.086s\n  sys\t0m0.035s\n\n  warm cache, refs packed:\n  real\t0m0.069s\n  user\t0m0.054s\n  sys\t0m0.016s\n\nSo 10s is pretty abysmal (and on an SSD, no less). I would expect the\ncache to be warm (we just wrote these refs!) but I could also believe\nthat CI systems are under heavy I/O and memory pressure, so we sometimes\nend up crossing the 60s mark.\n\nSo bumping Apache's timeout to 600s or something would probably be a\nfine mitigation. That's still not _solving_ the problem, but presumably\nan order of magnitude is enough for it to never come up in practice.\n\nMichael suggested packing the refs as a mitigation. I was lukewarm on\nthat in my previous email, because it wasn't clear to me how close we\nwere on the timeout budget, and if it would just make the race less\nfrequent (rather than never happen). But seeing those cold-cache numbers\nmakes me think it might be worth doing just on principle to make the\ntests more efficient, and any timeout mitigation is a bonus.\n\nOf course the pack-refs process (and the initial ref writes) will still\nhave to touch all of those loose files, so those will still be slow. But\nthey're not on a timeout, and I suspect we read the result many more\ntimes than we write/pack (the test failures we are seeing are not in the\nexpensive tests, but just \"normal\" tests that are stuck with the\ngigantic ref state).\n\n-Peff\n"},{"id":"546583","messageId":"20260628075716.GA3525066@coredump.intra.peff.net","threadId":"65847","inReplyTo":"CAC2QwmLkHUymvtYbjY8aQO9_VogvaSXdbb1_DSZtcBttGfN0tg@mail.gmail.com","subject":"[PATCH 0/3] fixing expensive http test timeouts","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-28T07:57:16Z","receivedAt":"2026-06-28T07:57:24Z","isPatch":true,"body":"On Fri, Jun 26, 2026 at 04:26:28PM -0700, Michael Montalbo wrote:\n\n> I think Peff and Patrick's suggestion to just increase the Apache timeout\n> makes sense. I ran some experiments using a really long timeout with an\n> artificially slowed down CI runner and all the jobs made progress\n> (if slowly) without stalling, and eventually completed successfully:\n> \n> https://github.com/mmontalbo/git/actions/runs/28267019651\n> \n> I haven't spent a lot of time trying to figure out what the right timeout\n> value should be. An hour definitely seems like overkill, with something\n> on the order of 5-10 minutes seeming more reasonable, but I don't\n> have a principled number.\n\nHere are some patches to keep things moving along. I arbitrarily picked\n10 minutes, because multiplying the 1-minute default by 10 felt right. ;)\n\nThe first one just bumps the timeout and should make our problems go\naway. The other two are optimizations, but I'm on the fence on whether\nthe final patch is worth it.\n\nThanks again for all of the digging.\n\n  [1/3]: t/lib-httpd: bump apache timeout\n  [2/3]: t5551: put many-tags case into its own repo\n  [3/3]: t5551: pack refs after creating many tags\n\n t/lib-httpd/apache.conf     |  1 +\n t/t5551-http-fetch-smart.sh | 10 ++++++----\n 2 files changed, 7 insertions(+), 4 deletions(-)\n\n-Peff\n"},{"id":"546584","messageId":"20260628080009.GA107826@coredump.intra.peff.net","threadId":"65847","inReplyTo":"20260628075716.GA3525066@coredump.intra.peff.net","subject":"[PATCH 1/3] t/lib-httpd: bump apache timeout","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-28T08:00:09Z","receivedAt":"2026-06-28T08:00:11Z","isPatch":true,"body":"Since enabling more tests with 7a094d68a2 (ci: run expensive tests on\npush builds to integration branches, 2026-05-08), we sometimes see test\nfailures or timeouts in GitHub CI. The culprit seems to be the \"enormous\nref negotiation\" test in t5551, which creates ~100k tag refs in our http\nserver-side repo.\n\nIterating through the loose refs of this repo to generate a ref\nadvertisement can take a long time, especially on a platform with slow\nI/O. On my otherwise unloaded local machine, a cold cache ref\nadvertisement takes ~10s. On a busy CI machine running tests in\nparallel, it can presumably top 60s, which runs afoul of Apache's\ndefault CGI timeout.\n\nThe result in t5551 is a test failure, where Apache simply hangs up the\nconnection and the client reports an error. But worse, t5559 runs the\nsame test with HTTP/2, and a bug in Apache causes the connection to hang\nindefinitely! We eventually see this as a CI timeout after 6 hours.\n\nLet's bump Apache's timeout to something much larger: 600 seconds. This\ndoesn't eliminate the possibility of a timeout, but it makes it much\nless likely. It should eliminate both the test failures and the CI\ntimeouts in practice, and it protects us from running into similar\nproblems with other tests in the future.\n\nThere are two counter-arguments to consider.\n\nOne, could/should we just make the test faster? Probably yes. The\nbiggest mistake here is having such an absurd number of unpacked refs on\na system which is bottle-necked on I/O. But I think it's worth bumping\nthe timeout so that we can fix this (and possibly other) correctness\nissues, and then consider performance separately (which we'll do in\nsubsequent patches).\n\nAnd two, is this just papering over a problem that users might see in\nthe real world? We could teach Git to handle this case more gracefully\nwith optimizations or keep-alives. But I think it's really an artificial\nsituation. You need a combination of this silly number of loose refs,\nplus a very heavily loaded system. If you were trying to run a real\nserver and it took more than 60s to generate the ref advertisement, I\ndon't think the timeout is your biggest problem. Your crappy service is,\nand you should adjust your resources to match your load. I.e., it is\nprobably reasonable for Git to assume that advertisements happen\nfast-ish and don't need protocol-level keepalives.\n\nThough the patch here is small, tons of work went into analyzing the\nproblem. Many thanks to the contributors credited below.\n\nHelped-by: Michael Montalbo <mmontalbo@gmail.com>\nHelped-by: Patrick Steinhardt <ps@pks.im>\nSigned-off-by: Jeff King <peff@peff.net>\n---\nI didn't reference Michael's bugzilla report directly, because you can't\nread it without a login. :(\n\nMaybe it's worth doing anyway?\n\n t/lib-httpd/apache.conf | 1 +\n 1 file changed, 1 insertion(+)\n\ndiff --git a/t/lib-httpd/apache.conf b/t/lib-httpd/apache.conf\nindex 40a690b0bb..4149fc1078 100644\n--- a/t/lib-httpd/apache.conf\n+++ b/t/lib-httpd/apache.conf\n@@ -4,6 +4,7 @@ DocumentRoot www\n LogFormat \"%h %l %u %t \\\"%r\\\" %>s %b\" common\n CustomLog access.log common\n ErrorLog error.log\n+Timeout 600\n <IfModule !mod_log_config.c>\n \tLoadModule log_config_module modules/mod_log_config.so\n </IfModule>\n-- \n2.55.0.rc2.353.gf769b6597e\n\n"},{"id":"546585","messageId":"20260628080345.GB107826@coredump.intra.peff.net","threadId":"65847","inReplyTo":"20260628075716.GA3525066@coredump.intra.peff.net","subject":"[PATCH 2/3] t5551: put many-tags case into its own repo","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-28T08:03:45Z","receivedAt":"2026-06-28T08:03:47Z","isPatch":true,"body":"Most of the t5551 http fetch tests use a handful of refs. But there are\na few test cases which check our handling of large numbers of refs.\nThese tests use the same server-side repo, so all subsequent tests end\nup having to consider those extra refs, too.\n\nThe result is that the test script is a bit slower than it needs to be.\nIn a normal run, moving the \"2,000 tags\" test into its own repo drops my\nruntime for the whole script from ~2.7s to ~1.9s.\n\nThis is a modest gain, but when we add the \"--long\" flag it gets much\nbigger. There we trigger a test (marked with EXPENSIVE) that adds\n100,000 tags, and the script runtime jumps to ~95s. But if we use the\nsame \"many tags\" repo for that, our runtime drops to just ~37s.\n\nThis is a pretty easy win to drop the cost of the script. It may even be\na larger gain on a heavily loaded system, since one of the main costs\nhere is unpacked refs, which are heavy on system time and I/O costs.\n\nIt's possible we are reducing test coverage, since all of those other\ntests were inadvertently using large ref advertisements (and thus could\nhave uncovered some unexpected interaction). But that seems somewhat\nunlikely; the tests targeted at the large number of refs are doing\nroughly similar things to the other tests.\n\nNote that the real performance culprit is the 100k-tag --long test, not\nthe 2k-tag one. So we could just let the 100k one use its own repo, and\nkeep the 2k tags in the main repo. But since these two tests are\nsomewhat interlinked, it's easier to just move them both (and it does\nprovide a small gain even for the 2000-tag test). I also notice that the\n2000-tag test is gated on the CMDLINE_LIMIT prereq, and without that the\nlater EXPENSIVE test will fail (since we won't have a too-many-refs\nclone). Nobody seems to have noticed or complained after many years, and\nI left it alone for this patch.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n t/t5551-http-fetch-smart.sh | 9 +++++----\n 1 file changed, 5 insertions(+), 4 deletions(-)\n\ndiff --git a/t/t5551-http-fetch-smart.sh b/t/t5551-http-fetch-smart.sh\nindex e236e526f0..cd851f24b8 100755\n--- a/t/t5551-http-fetch-smart.sh\n+++ b/t/t5551-http-fetch-smart.sh\n@@ -397,15 +397,16 @@ create_tags () {\n }\n \n test_expect_success 'create 2,000 tags in the repo' '\n+\tgit init \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n \t(\n-\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/repo.git\" &&\n+\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n \t\tcreate_tags 1 2000\n \t)\n '\n \n test_expect_success CMDLINE_LIMIT \\\n \t'clone the 2,000 tag repo to check OS command line overflow' '\n-\trun_with_limited_cmdline git clone $HTTPD_URL/smart/repo.git too-many-refs &&\n+\trun_with_limited_cmdline git clone $HTTPD_URL/smart/many-tags.git too-many-refs &&\n \t(\n \t\tcd too-many-refs &&\n \t\tgit for-each-ref refs/tags >actual &&\n@@ -483,12 +484,12 @@ test_expect_success 'test allowanysha1inwant with unreachable' '\n test_expect_success EXPENSIVE 'http can handle enormous ref negotiation' '\n \ttest_when_finished \"rm -f tags\" &&\n \t(\n-\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/repo.git\" &&\n+\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n \t\tcreate_tags 2001 50000\n \t) &&\n \tgit -C too-many-refs fetch -q --tags &&\n \t(\n-\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/repo.git\" &&\n+\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n \t\tcreate_tags 50001 100000\n \t) &&\n \tgit -C too-many-refs fetch -q --tags &&\n-- \n2.55.0.rc2.353.gf769b6597e\n\n"},{"id":"546586","messageId":"20260628080710.GC107826@coredump.intra.peff.net","threadId":"65847","inReplyTo":"20260628075716.GA3525066@coredump.intra.peff.net","subject":"[PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-28T08:07:10Z","receivedAt":"2026-06-28T08:07:11Z","isPatch":true,"body":"We have two tests that create 2,000 and 100,000 tags respectively.\nAfter doing so, the resulting state can be a bit slow to work with when\nusing the \"files\" ref backend, as each of those refs is in its own file.\n\nThis isn't a very realistic scenario, as we'd expect most of those refs\nto be packed. If they accrue over time along with objects, they'd get\npacked by maintenance/gc runs. And if you have a process that creates a\nton of refs at once (like a big fast-import), the usual recommendation\nis to run maintenance afterwards.\n\nSo let's follow that recommendation and pack the refs ourselves.\nUnfortunately, this does not seem to produce an improvement to the\nrun-time of the test script! That's because after producing this state,\nwe perform only a few fetches of it. And packing the refs costs at least\nas much as serving a ref advertisement (both have to iterate the refs,\nbut packing additionally must write .lock files as we pack).\n\nMy wall-clock time was slightly improved (but within the noise) with\nthis patch, but my user and system CPU time were slightly worse!\nHowever, on a loaded system with I/O bottlenecks, it may be a net win.\nThat's somewhat of a guess, though.\n\nIt would be nice if we had a way to generate all of these refs without\nwriting so many individual files. But even if we taught the ref code to\nwrite large cases directly to the packed-refs file, we'd still need to\ntake individual locks. The real solution is a backend like reftable,\nwhich shaves ~30% off of the test runtime.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\nI'm iffy on whether this one is worth it.\n\nIf you apply just this patch without patch 2, then the run-time does\nimprove quite a bit. The cost of packing is amortized by the improved\nperformance for all of those subsequent tests (but after patch 2, they\nnever even see the unpacked state).\n\nLikewise, I suspect this would make our timeout problems go away even\nwithout patch 1.\n\nSo the whole series _could_ be reduced to just this one patch. But\nhopefully the reasoning given in the earlier patches makes sense, at\nwhich point this one is kind of superfluous.\n\n t/t5551-http-fetch-smart.sh | 1 +\n 1 file changed, 1 insertion(+)\n\ndiff --git a/t/t5551-http-fetch-smart.sh b/t/t5551-http-fetch-smart.sh\nindex cd851f24b8..e2e729216f 100755\n--- a/t/t5551-http-fetch-smart.sh\n+++ b/t/t5551-http-fetch-smart.sh\n@@ -393,6 +393,7 @@ create_tags () {\n \ttag=$(perl -e \"print \\\"bla\\\" x 30\") &&\n \tsed -e \"s|^:\\([^ ]*\\) \\(.*\\)$|create refs/tags/$tag-\\1 \\2|\" <marks >input &&\n \tgit update-ref --stdin <input &&\n+\tgit pack-refs --all &&\n \trm input\n }\n \n-- \n2.55.0.rc2.353.gf769b6597e\n"},{"id":"546613","messageId":"xmqqy0fy1hnf.fsf@gitster.g","threadId":"65847","inReplyTo":"20260628080710.GC107826@coredump.intra.peff.net","subject":"Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-28T21:25:56Z","receivedAt":"2026-06-28T21:25:59Z","isPatch":true,"body":"\nJeff King <peff@peff.net> writes:\n\n> So let's follow that recommendation and pack the refs ourselves.\n> Unfortunately, this does not seem to produce an improvement to the\n> run-time of the test script! That's because after producing this state,\n> we perform only a few fetches of it. And packing the refs costs at least\n> as much as serving a ref advertisement (both have to iterate the refs,\n> but packing additionally must write .lock files as we pack).\n\nTesting a pathological set-up with too many loose refs may have\nextra value, as long as we are also testing the recommended set-up,\nso ideally we should have both ;-) but if we have to pick only one\nand drop the other, we probably should be testing the packed case.\n\n> I'm iffy on whether this one is worth it.\n\nI am ambivalent, too, about this change for the purpose of the\n\"yeek, apache times out while enumerating refs\" issue.  But see\nabove ;-)\n\n"},{"id":"546614","messageId":"xmqqh5mm1gsf.fsf@gitster.g","threadId":"65847","inReplyTo":"20260628080345.GB107826@coredump.intra.peff.net","subject":"Re: [PATCH 2/3] t5551: put many-tags case into its own repo","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-28T21:44:32Z","receivedAt":"2026-06-28T21:44:34Z","isPatch":true,"body":"Jeff King <peff@peff.net> writes:\n\n> diff --git a/t/t5551-http-fetch-smart.sh b/t/t5551-http-fetch-smart.sh\n> index e236e526f0..cd851f24b8 100755\n> --- a/t/t5551-http-fetch-smart.sh\n> +++ b/t/t5551-http-fetch-smart.sh\n> @@ -397,15 +397,16 @@ create_tags () {\n>  }\n>  \n>  test_expect_success 'create 2,000 tags in the repo' '\n> +\tgit init \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n>  \t(\n> -\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/repo.git\" &&\n> +\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n>  \t\tcreate_tags 1 2000\n>  \t)\n>  '\n\nWhile all the other repositories used in this tests are bare\nrepositories, this new one is a non-bare repository.\n\nIt shouldn't make any difference, but since I noticed it...\n"},{"id":"546618","messageId":"20260629003434.GA1228461@coredump.intra.peff.net","threadId":"65847","inReplyTo":"xmqqh5mm1gsf.fsf@gitster.g","subject":"Re: [PATCH 2/3] t5551: put many-tags case into its own repo","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-29T00:34:34Z","receivedAt":"2026-06-29T00:34:42Z","isPatch":true,"body":"On Sun, Jun 28, 2026 at 02:44:32PM -0700, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > diff --git a/t/t5551-http-fetch-smart.sh b/t/t5551-http-fetch-smart.sh\n> > index e236e526f0..cd851f24b8 100755\n> > --- a/t/t5551-http-fetch-smart.sh\n> > +++ b/t/t5551-http-fetch-smart.sh\n> > @@ -397,15 +397,16 @@ create_tags () {\n> >  }\n> >  \n> >  test_expect_success 'create 2,000 tags in the repo' '\n> > +\tgit init \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n> >  \t(\n> > -\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/repo.git\" &&\n> > +\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n> >  \t\tcreate_tags 1 2000\n> >  \t)\n> >  '\n> \n> While all the other repositories used in this tests are bare\n> repositories, this new one is a non-bare repository.\n> \n> It shouldn't make any difference, but since I noticed it...\n\nAh, yeah. It should work either way, but it is slightly confusing for it\nto be non-bare. I'll wait to re-send (though if nothing else comes up,\nit may be simpler for you to just amend on your side).\n\n-Peff\n"},{"id":"546622","messageId":"akIJQbOUbdBbkTef@pks.im","threadId":"65847","inReplyTo":"20260628080710.GC107826@coredump.intra.peff.net","subject":"Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-29T05:57:21Z","receivedAt":"2026-06-29T05:57:28Z","isPatch":true,"body":"On Sun, Jun 28, 2026 at 04:07:10AM -0400, Jeff King wrote:\n> We have two tests that create 2,000 and 100,000 tags respectively.\n> After doing so, the resulting state can be a bit slow to work with when\n> using the \"files\" ref backend, as each of those refs is in its own file.\n> \n> This isn't a very realistic scenario, as we'd expect most of those refs\n> to be packed. If they accrue over time along with objects, they'd get\n> packed by maintenance/gc runs. And if you have a process that creates a\n> ton of refs at once (like a big fast-import), the usual recommendation\n> is to run maintenance afterwards.\n> \n> So let's follow that recommendation and pack the refs ourselves.\n> Unfortunately, this does not seem to produce an improvement to the\n> run-time of the test script! That's because after producing this state,\n> we perform only a few fetches of it. And packing the refs costs at least\n> as much as serving a ref advertisement (both have to iterate the refs,\n> but packing additionally must write .lock files as we pack).\n\n> My wall-clock time was slightly improved (but within the noise) with\n> this patch, but my user and system CPU time were slightly worse!\n> However, on a loaded system with I/O bottlenecks, it may be a net win.\n> That's somewhat of a guess, though.\n> \n> It would be nice if we had a way to generate all of these refs without\n> writing so many individual files. But even if we taught the ref code to\n> write large cases directly to the packed-refs file, we'd still need to\n> take individual locks. The real solution is a backend like reftable,\n> which shaves ~30% off of the test runtime.\n\nWe kind of already have this with the `REF_TRANSACTION_FLAG_INITIAL`\nflag, but right now it is only used when performing a clone or when\nmigrating references. Also, it requires an empty repository that has no\nreferences yet.\n\nIt raises the question whether we could also extend git-fast-import(1)\nto use it, as it would typically be run on an almost-empty repository.\nIt's the \"almost\" that kills it though, as we already do have at least\nthe HEAD reference. So it could be feasible, but it's not as trivial as\njust setting the flag and then we're magically faster.\n\nAnd besides, in this particular test here we run git-fast-import(1)\nmultiple times in the same repository, so it wouldn't help us.\n\nWe could of course extend all of this so that Git is able to write into\nthe packed-refs directly, even with preexisting refs. But I agree with\nyour sentiment: it doesn't feel worth it as the reftable backend fixes\nscenarios like this anyway.\n\n> Signed-off-by: Jeff King <peff@peff.net>\n> ---\n> I'm iffy on whether this one is worth it.\n> \n> If you apply just this patch without patch 2, then the run-time does\n> improve quite a bit. The cost of packing is amortized by the improved\n> performance for all of those subsequent tests (but after patch 2, they\n> never even see the unpacked state).\n> \n> Likewise, I suspect this would make our timeout problems go away even\n> without patch 1.\n> \n> So the whole series _could_ be reduced to just this one patch. But\n> hopefully the reasoning given in the earlier patches makes sense, at\n> which point this one is kind of superfluous.\n\nAgreed. I'd just merge the first two patches and drop this one here.\n\nThanks!\n\nPatrick\n"},{"id":"546632","messageId":"akIfsaVMB_S6kfJQ@pks.im","threadId":"65847","inReplyTo":"20260628075716.GA3525066@coredump.intra.peff.net","subject":"Re: [PATCH 0/3] fixing expensive http test timeouts","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-29T07:33:05Z","receivedAt":"2026-06-29T07:33:12Z","isPatch":true,"body":"On Sun, Jun 28, 2026 at 03:57:16AM -0400, Jeff King wrote:\n> On Fri, Jun 26, 2026 at 04:26:28PM -0700, Michael Montalbo wrote:\n> \n> > I think Peff and Patrick's suggestion to just increase the Apache timeout\n> > makes sense. I ran some experiments using a really long timeout with an\n> > artificially slowed down CI runner and all the jobs made progress\n> > (if slowly) without stalling, and eventually completed successfully:\n> > \n> > https://github.com/mmontalbo/git/actions/runs/28267019651\n> > \n> > I haven't spent a lot of time trying to figure out what the right timeout\n> > value should be. An hour definitely seems like overkill, with something\n> > on the order of 5-10 minutes seeming more reasonable, but I don't\n> > have a principled number.\n> \n> Here are some patches to keep things moving along. I arbitrarily picked\n> 10 minutes, because multiplying the 1-minute default by 10 felt right. ;)\n> \n> The first one just bumps the timeout and should make our problems go\n> away. The other two are optimizations, but I'm on the fence on whether\n> the final patch is worth it.\n> \n> Thanks again for all of the digging.\n> \n>   [1/3]: t/lib-httpd: bump apache timeout\n>   [2/3]: t5551: put many-tags case into its own repo\n>   [3/3]: t5551: pack refs after creating many tags\n\nBy the way, the only reason why we at GitLab haven't been feeling the\npain is that we only enable GIT_TEST_LONG for GitHub. So I was wondering\nwhether we want to have something like the below patch on top.\n\nPatrick\n\ndiff --git a/ci/lib.sh b/ci/lib.sh\nindex b939110a6e..57801586aa 100755\n--- a/ci/lib.sh\n+++ b/ci/lib.sh\n@@ -215,6 +215,14 @@ then\n \ttest macos != \"$CI_OS_NAME\" || CI_OS_NAME=osx\n \tCI_REPO_SLUG=\"$GITHUB_REPOSITORY\"\n \tCI_JOB_ID=\"$GITHUB_RUN_ID\"\n+\n+\tcase \"$GITHUB_EVENT_NAME\" in\n+\tpull_request)\n+\t\tCI_EVENT=pull_request;;\n+\tpush)\n+\t\tCI_EVENT=push;;\n+\tesac\n+\n \tCC=\"${CC_PACKAGE:-${CC:-gcc}}\"\n \tDONT_SKIP_TAGS=t\n \thandle_failed_tests () {\n@@ -239,6 +247,13 @@ then\n \tCI_BRANCH=\"$CI_COMMIT_REF_NAME\"\n \tCI_COMMIT=\"$CI_COMMIT_SHA\"\n \n+\tcase \"$CI_PIPELINE_SOURCE\" in\n+\tmerge_request_event)\n+\t\tCI_EVENT=pull_request;;\n+\tpush)\n+\t\tCI_EVENT=push;;\n+\tesac\n+\n \tcase \"$OS,$CI_JOB_IMAGE\" in\n \tWindows_NT,*)\n \t\tCI_OS_NAME=windows\n@@ -319,7 +334,7 @@ export SKIP_DASHED_BUILT_INS=YesPlease\n # enable \"expensive\" tests for PR events.\n # In order to catch bugs introduced at integration time by mismerges,\n # enable the long tests for pushes to the integration branches as well.\n-case \"$GITHUB_EVENT_NAME,$CI_BRANCH\" in\n+case \"$CI_EVENT,$CI_BRANCH\" in\n pull_request,*|push,*next*|push,*master*|push,*main*|push,*maint*)\n \texport GIT_TEST_LONG=YesPlease\n \t;;\n\n"},{"id":"546670","messageId":"xmqqldbxz9z4.fsf@gitster.g","threadId":"65847","inReplyTo":"akIfsaVMB_S6kfJQ@pks.im","subject":"Re: [PATCH 0/3] fixing expensive http test timeouts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-29T14:39:59Z","receivedAt":"2026-06-29T14:40:01Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> By the way, the only reason why we at GitLab haven't been feeling the\n> pain is that we only enable GIT_TEST_LONG for GitHub. So I was wondering\n> whether we want to have something like the below patch on top.\n\nIf we can afford the cycles, it would be good to have similarly\nlarger coverage on two different platforms (compared to leaving one\nof them not doing as much as the other when we know it).  On the\nother hand, if we cannot cover _everything_ in one platform, it may\nbe a better use of the resources to have the other platform things\nthat are not covered already.  I see that among different pipeline\nsources, we are doing TEST_LONG for pull requests to any branch, and\npushes only to \"cast in stone\" branches.  If there are other\nbranches that deserve to be tested with TEST_LONG upon other events\nthat the existing GitHub Actions CI does not trigger, it may be good\nto have GitLab CI cover them, perhaps?\n\n\n\n\n\n>\n> Patrick\n>\n> diff --git a/ci/lib.sh b/ci/lib.sh\n> index b939110a6e..57801586aa 100755\n> --- a/ci/lib.sh\n> +++ b/ci/lib.sh\n> @@ -215,6 +215,14 @@ then\n>  \ttest macos != \"$CI_OS_NAME\" || CI_OS_NAME=osx\n>  \tCI_REPO_SLUG=\"$GITHUB_REPOSITORY\"\n>  \tCI_JOB_ID=\"$GITHUB_RUN_ID\"\n> +\n> +\tcase \"$GITHUB_EVENT_NAME\" in\n> +\tpull_request)\n> +\t\tCI_EVENT=pull_request;;\n> +\tpush)\n> +\t\tCI_EVENT=push;;\n> +\tesac\n> +\n>  \tCC=\"${CC_PACKAGE:-${CC:-gcc}}\"\n>  \tDONT_SKIP_TAGS=t\n>  \thandle_failed_tests () {\n> @@ -239,6 +247,13 @@ then\n>  \tCI_BRANCH=\"$CI_COMMIT_REF_NAME\"\n>  \tCI_COMMIT=\"$CI_COMMIT_SHA\"\n>  \n> +\tcase \"$CI_PIPELINE_SOURCE\" in\n> +\tmerge_request_event)\n> +\t\tCI_EVENT=pull_request;;\n> +\tpush)\n> +\t\tCI_EVENT=push;;\n> +\tesac\n> +\n>  \tcase \"$OS,$CI_JOB_IMAGE\" in\n>  \tWindows_NT,*)\n>  \t\tCI_OS_NAME=windows\n> @@ -319,7 +334,7 @@ export SKIP_DASHED_BUILT_INS=YesPlease\n>  # enable \"expensive\" tests for PR events.\n>  # In order to catch bugs introduced at integration time by mismerges,\n>  # enable the long tests for pushes to the integration branches as well.\n> -case \"$GITHUB_EVENT_NAME,$CI_BRANCH\" in\n> +case \"$CI_EVENT,$CI_BRANCH\" in\n>  pull_request,*|push,*next*|push,*master*|push,*main*|push,*maint*)\n>  \texport GIT_TEST_LONG=YesPlease\n>  \t;;\n"},{"id":"546671","messageId":"xmqqh5mlz9uw.fsf@gitster.g","threadId":"65847","inReplyTo":"20260629003434.GA1228461@coredump.intra.peff.net","subject":"Re: [PATCH 2/3] t5551: put many-tags case into its own repo","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-29T14:42:31Z","receivedAt":"2026-06-29T14:42:33Z","isPatch":true,"body":"Jeff King <peff@peff.net> writes:\n\n> On Sun, Jun 28, 2026 at 02:44:32PM -0700, Junio C Hamano wrote:\n>\n>> Jeff King <peff@peff.net> writes:\n>> \n>> > diff --git a/t/t5551-http-fetch-smart.sh b/t/t5551-http-fetch-smart.sh\n>> > index e236e526f0..cd851f24b8 100755\n>> > --- a/t/t5551-http-fetch-smart.sh\n>> > +++ b/t/t5551-http-fetch-smart.sh\n>> > @@ -397,15 +397,16 @@ create_tags () {\n>> >  }\n>> >  \n>> >  test_expect_success 'create 2,000 tags in the repo' '\n>> > +\tgit init \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n>> >  \t(\n>> > -\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/repo.git\" &&\n>> > +\t\tcd \"$HTTPD_DOCUMENT_ROOT_PATH/many-tags.git\" &&\n>> >  \t\tcreate_tags 1 2000\n>> >  \t)\n>> >  '\n>> \n>> While all the other repositories used in this tests are bare\n>> repositories, this new one is a non-bare repository.\n>> \n>> It shouldn't make any difference, but since I noticed it...\n>\n> Ah, yeah. It should work either way, but it is slightly confusing for it\n> to be non-bare. I'll wait to re-send (though if nothing else comes up,\n> it may be simpler for you to just amend on your side).\n\nOK.  It seems both Patrick and you are in favor of using only [1/3]\n& [2/3] but dropping [3/3]?  If that is the concensus I can just\ntweak this one and apply before 2.55 final.\n\nThanks.\n"},{"id":"546679","messageId":"akKYv3nqX0BXcavu@pks.im","threadId":"65847","inReplyTo":"xmqqldbxz9z4.fsf@gitster.g","subject":"Re: [PATCH 0/3] fixing expensive http test timeouts","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-29T16:09:35Z","receivedAt":"2026-06-29T16:09:41Z","isPatch":true,"body":"On Mon, Jun 29, 2026 at 07:39:59AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > By the way, the only reason why we at GitLab haven't been feeling the\n> > pain is that we only enable GIT_TEST_LONG for GitHub. So I was wondering\n> > whether we want to have something like the below patch on top.\n> \n> If we can afford the cycles, it would be good to have similarly\n> larger coverage on two different platforms (compared to leaving one\n> of them not doing as much as the other when we know it).  On the\n> other hand, if we cannot cover _everything_ in one platform, it may\n> be a better use of the resources to have the other platform things\n> that are not covered already.  I see that among different pipeline\n> sources, we are doing TEST_LONG for pull requests to any branch, and\n> pushes only to \"cast in stone\" branches.  If there are other\n> branches that deserve to be tested with TEST_LONG upon other events\n> that the existing GitHub Actions CI does not trigger, it may be good\n> to have GitLab CI cover them, perhaps?\n\nI'm a bit hesitant to do such a split, mostly because the canonical\nsource of truth that the project typically uses is GitHub's CI. So I\nwant us at GitLab to be able to catch the same issues that GitHub would\nflag. And if GitLab's CI stopped detecting everything that GitHub does,\nthen the result would likely be that we often create merge requests on\nboth platforms, which would only result in more wasted resources.\n\nPatrick\n"},{"id":"546681","messageId":"xmqqik71xqtc.fsf@gitster.g","threadId":"65847","inReplyTo":"akKYv3nqX0BXcavu@pks.im","subject":"Re: [PATCH 0/3] fixing expensive http test timeouts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-29T16:19:11Z","receivedAt":"2026-06-29T16:19:14Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n>> pushes only to \"cast in stone\" branches.  If there are other\n>> branches that deserve to be tested with TEST_LONG upon other events\n>> that the existing GitHub Actions CI does not trigger, it may be good\n>> to have GitLab CI cover them, perhaps?\n>\n> I'm a bit hesitant to do such a split, mostly because the canonical\n> source of truth that the project typically uses is GitHub's CI. So I\n> want us at GitLab to be able to catch the same issues that GitHub would\n> flag. And if GitLab's CI stopped detecting everything that GitHub does,\n> then the result would likely be that we often create merge requests on\n> both platforms, which would only result in more wasted resources.\n\nI didn't suggest splitting them into two circles that overlap but\neach with area only it covers, though.  GitLab's coverage can be\nsuperset to GitHub's and that would satify what I suggested.\n\nFWIW, I do not consider GitHub's CI \"the canonical source\" at all.\nIt is a very handy service to use to check how well we are doing,\nbut from time to time it has its own hiccups ;-).\n\nWhat can we do to make the visibility of GitLab's CI more prominent?\n\nI know where the CI jobs that are triggered when I push out the\nintegration branches are found at GitHub's website[*], but I do not\nthink I know the corresponding one at GitLab, for example, and I\nthink that is a shame.\n\n\n[Footnote]\n *1* I just made https://tinyurl.com/github-gitci that points at\n     https://github.com/git/git/actions/workflows/main.yml?query=event%3Apush+actor%3Agitster\n\n"},{"id":"546707","messageId":"20260629203527.GA1895313@coredump.intra.peff.net","threadId":"65847","inReplyTo":"akIJQbOUbdBbkTef@pks.im","subject":"Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-29T20:35:27Z","receivedAt":"2026-06-29T20:35:36Z","isPatch":true,"body":"On Mon, Jun 29, 2026 at 07:57:21AM +0200, Patrick Steinhardt wrote:\n\n> > It would be nice if we had a way to generate all of these refs without\n> > writing so many individual files. But even if we taught the ref code to\n> > write large cases directly to the packed-refs file, we'd still need to\n> > take individual locks. The real solution is a backend like reftable,\n> > which shaves ~30% off of the test runtime.\n> \n> We kind of already have this with the `REF_TRANSACTION_FLAG_INITIAL`\n> flag, but right now it is only used when performing a clone or when\n> migrating references. Also, it requires an empty repository that has no\n> references yet.\n> \n> It raises the question whether we could also extend git-fast-import(1)\n> to use it, as it would typically be run on an almost-empty repository.\n> It's the \"almost\" that kills it though, as we already do have at least\n> the HEAD reference. So it could be feasible, but it's not as trivial as\n> just setting the flag and then we're magically faster.\n> \n> And besides, in this particular test here we run git-fast-import(1)\n> multiple times in the same repository, so it wouldn't help us.\n> \n> We could of course extend all of this so that Git is able to write into\n> the packed-refs directly, even with preexisting refs. But I agree with\n> your sentiment: it doesn't feel worth it as the reftable backend fixes\n> scenarios like this anyway.\n\nYup. In the past I've pondered exposing this via update-ref, but I think\nit's too weird and/or dangerous to do so. Especially because you are\nstill stuck creating all of the .lock files, so the performance is not\neven that much better (though it does save you doing so _twice_ when you\nthen pack the refs).\n\nSo the performance option you really want is \"YOLO, just write some\npacked-refs without locking\". But that is not something I think we want\nto expose to users. ;)\n\nWe could do it ad-hoc within this test like so:\n\ndiff --git a/t/t5551-http-fetch-smart.sh b/t/t5551-http-fetch-smart.sh\nindex dcff0bc7d4..276c7ac002 100755\n--- a/t/t5551-http-fetch-smart.sh\n+++ b/t/t5551-http-fetch-smart.sh\n@@ -389,11 +389,15 @@ create_tags () {\n \t\techo \"from :$1\"\n \tdone | git fast-import --export-marks=marks &&\n \n+\t# should be mostly a noop, but makes sure we have the right header\n+\tgit pack-refs &&\n+\n \t# now assign tags to all the dangling commits we created above\n+\t# It is OK to write directly to the packed-refs file because we know\n+\t# that our entries are sorted by refname, and that they all\n+\t# come after what we wrote earlier.\n \ttag=$(perl -e \"print \\\"bla\\\" x 30\") &&\n-\tsed -e \"s|^:\\([^ ]*\\) \\(.*\\)$|create refs/tags/$tag-\\1 \\2|\" <marks >input &&\n-\tgit update-ref --stdin <input &&\n-\trm input\n+\tsed -e \"s|^:\\([^ ]*\\) \\(.*\\)$|\\2 refs/tags/$tag-\\1|\" <marks >>packed-refs\n }\n \n test_expect_success 'create 2,000 tags in the repo' '\n\nThat gives us the same ~30% speedup that using reftables does, but it\nstill is quite gross and fragile. And it is not even strictly correct,\nbecause we don't zero-pad the numbered tags (so our file is subtly out\nof order).  Plus it would need to be conditional on the ref backend\nbeing used. Yuck.\n\n\nThere's one other thing you might find interesting. While poking at the\ntimings here the other day, I noticed that reftable is very eager to\nstat the tables.list file. Try this:\n\n  git init --ref-format=reftable\n  blob=$(echo foo | git hash-object -w --stdin)\n  seq -f \"create refs/tags/foo-%g $blob\" 2000 |\n  strace -c git update-ref --stdin\n\nWe make 2000 fstat, which strace claims takes 85% of the time. I suspect\nthis is over-emphasized because strace inherently makes syscalls slow,\nbut running with perf also highlights it as a non-trivial cost.\n\nIt has been a long time since I've thought about reftable internals, but\nit feels like we ought to be able to take the lock and then trust that\nthe stack has not been manipulated.\n\nIt may not be worth digging into too much, though. I can make 50,000\nrefs in 150ms on my system, which is probably good enough (especially\ncompared to the files backend).\n\n-Peff\n"},{"id":"546708","messageId":"20260629203608.GB1895313@coredump.intra.peff.net","threadId":"65847","inReplyTo":"xmqqh5mlz9uw.fsf@gitster.g","subject":"Re: [PATCH 2/3] t5551: put many-tags case into its own repo","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-29T20:36:08Z","receivedAt":"2026-06-29T20:36:10Z","isPatch":true,"body":"On Mon, Jun 29, 2026 at 07:42:31AM -0700, Junio C Hamano wrote:\n\n> > Ah, yeah. It should work either way, but it is slightly confusing for it\n> > to be non-bare. I'll wait to re-send (though if nothing else comes up,\n> > it may be simpler for you to just amend on your side).\n> \n> OK.  It seems both Patrick and you are in favor of using only [1/3]\n> & [2/3] but dropping [3/3]?  If that is the concensus I can just\n> tweak this one and apply before 2.55 final.\n\nYep, that sounds great. Looks like it already happened. :)\n\n-Peff\n"},{"id":"546744","messageId":"akOGzAq8Is7ghgIM@pks.im","threadId":"65847","inReplyTo":"xmqqik71xqtc.fsf@gitster.g","subject":"Re: [PATCH 0/3] fixing expensive http test timeouts","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-30T09:05:16Z","receivedAt":"2026-06-30T09:05:23Z","isPatch":true,"body":"On Mon, Jun 29, 2026 at 09:19:11AM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> >> pushes only to \"cast in stone\" branches.  If there are other\n> >> branches that deserve to be tested with TEST_LONG upon other events\n> >> that the existing GitHub Actions CI does not trigger, it may be good\n> >> to have GitLab CI cover them, perhaps?\n> >\n> > I'm a bit hesitant to do such a split, mostly because the canonical\n> > source of truth that the project typically uses is GitHub's CI. So I\n> > want us at GitLab to be able to catch the same issues that GitHub would\n> > flag. And if GitLab's CI stopped detecting everything that GitHub does,\n> > then the result would likely be that we often create merge requests on\n> > both platforms, which would only result in more wasted resources.\n> \n> I didn't suggest splitting them into two circles that overlap but\n> each with area only it covers, though.  GitLab's coverage can be\n> superset to GitHub's and that would satify what I suggested.\n\nFair.\n\n> FWIW, I do not consider GitHub's CI \"the canonical source\" at all.\n> It is a very handy service to use to check how well we are doing,\n> but from time to time it has its own hiccups ;-).\n\nWell, GitLab of course has its own share of hiccups, like for example\nthe Chocolatey issues we've been facing.\n\n> What can we do to make the visibility of GitLab's CI more prominent?\n> \n> I know where the CI jobs that are triggered when I push out the\n> integration branches are found at GitHub's website[*], but I do not\n> think I know the corresponding one at GitLab, for example, and I\n> think that is a shame.\n\nThe pipelines of the official mirror can be found at [1]. We might for\nexample add something like the below patch to our README.md to make it\nmore discoverable.\n\nPatrick\n\n[1]: https://gitlab.com/git-scm/git/-/pipelines\n\ndiff --git a/README.md b/README.md\nindex d87bca1b8c..9ad77fdf7e 100644\n--- a/README.md\n+++ b/README.md\n@@ -1,4 +1,5 @@\n-[![Build status](https://github.com/git/git/workflows/CI/badge.svg)](https://github.com/git/git/actions?query=branch%3Amaster+event%3Apush)\n+[![GitHub build status](https://github.com/git/git/workflows/CI/badge.svg)](https://github.com/git/git/actions?query=branch%3Amaster+event%3Apush)\n+[![GitLab build status](https://gitlab.com/git-scm/git/badges/master/pipeline.svg)](https://gitlab.com/git-scm/git/-/pipelines?ref=master)\n \n Git - fast, scalable, distributed revision control system\n =========================================================\n"},{"id":"546745","messageId":"akOG0oMu2KTqqyW7@pks.im","threadId":"65847","inReplyTo":"20260629203527.GA1895313@coredump.intra.peff.net","subject":"Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-30T09:05:22Z","receivedAt":"2026-06-30T09:05:26Z","isPatch":true,"body":"On Mon, Jun 29, 2026 at 04:35:27PM -0400, Jeff King wrote:\n> There's one other thing you might find interesting. While poking at the\n> timings here the other day, I noticed that reftable is very eager to\n> stat the tables.list file. Try this:\n> \n>   git init --ref-format=reftable\n>   blob=$(echo foo | git hash-object -w --stdin)\n>   seq -f \"create refs/tags/foo-%g $blob\" 2000 |\n>   strace -c git update-ref --stdin\n> \n> We make 2000 fstat, which strace claims takes 85% of the time. I suspect\n> this is over-emphasized because strace inherently makes syscalls slow,\n> but running with perf also highlights it as a non-trivial cost.\n\nYeah, this rings a bell. If I remember correctly, this is because we\ncall `refs_resolve_ref_unsafe()` to verify whether the target already\nexists. And as that function is generic, it wasn't easy to optimize it\nby reusing the already-loaded reftable stack.\n\n> It has been a long time since I've thought about reftable internals, but\n> it feels like we ought to be able to take the lock and then trust that\n> the stack has not been manipulated.\n\nYeah, that would certainly be an option to explore.\n\n> It may not be worth digging into too much, though. I can make 50,000\n> refs in 150ms on my system, which is probably good enough (especially\n> compared to the files backend).\n\nTrue, it's going to be much better compared to the \"files\" backend. But\nthat isn't enough reason to not optimize it even further -- doubly so if\nit actually takes 85% of the time. Sounds like a low-hanging fruit to me\nthat can result in a significant speedup.\n\nI'll probably not get to it anytime soon, but I'll create an issue to\nkeep track of it.\n\nPatrick\n"},{"id":"546805","messageId":"xmqq8q7vsukl.fsf@gitster.g","threadId":"65847","inReplyTo":"akOGzAq8Is7ghgIM@pks.im","subject":"Re: [PATCH 0/3] fixing expensive http test timeouts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-30T19:21:30Z","receivedAt":"2026-06-30T19:21:32Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> The pipelines of the official mirror can be found at [1]. We might for\n> example add something like the below patch to our README.md to make it\n> more discoverable.\n>\n> Patrick\n>\n> [1]: https://gitlab.com/git-scm/git/-/pipelines\n>\n> diff --git a/README.md b/README.md\n> index d87bca1b8c..9ad77fdf7e 100644\n> --- a/README.md\n> +++ b/README.md\n> @@ -1,4 +1,5 @@\n> -[![Build status](https://github.com/git/git/workflows/CI/badge.svg)](https://github.com/git/git/actions?query=branch%3Amaster+event%3Apush)\n> +[![GitHub build status](https://github.com/git/git/workflows/CI/badge.svg)](https://github.com/git/git/actions?query=branch%3Amaster+event%3Apush)\n> +[![GitLab build status](https://gitlab.com/git-scm/git/badges/master/pipeline.svg)](https://gitlab.com/git-scm/git/-/pipelines?ref=master)\n>  \n>  Git - fast, scalable, distributed revision control system\n>  =========================================================\n\nOh, nice.  We of course do not want to be heavily involved in\nadvertising offerings by commercial entities but I think these two\nsites deserve one line each for their continued service to the\ncommunity ;-)\n"},{"id":"546815","messageId":"20260630234702.GA3759976@coredump.intra.peff.net","threadId":"65847","inReplyTo":"akOG0oMu2KTqqyW7@pks.im","subject":"Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-30T23:47:02Z","receivedAt":"2026-06-30T23:47:09Z","isPatch":true,"body":"On Tue, Jun 30, 2026 at 11:05:22AM +0200, Patrick Steinhardt wrote:\n\n> > We make 2000 fstat, which strace claims takes 85% of the time. I suspect\n> > this is over-emphasized because strace inherently makes syscalls slow,\n> > but running with perf also highlights it as a non-trivial cost.\n> \n> Yeah, this rings a bell. If I remember correctly, this is because we\n> call `refs_resolve_ref_unsafe()` to verify whether the target already\n> exists. And as that function is generic, it wasn't easy to optimize it\n> by reusing the already-loaded reftable stack.\n\nI think it's similar to that. I dug down and got some actual performance\nnumbers below for eliminating the fstat() calls, with the conclusion\nthat it's probably not worth spending too much time digging into it. But\ndetails below for completeness.\n\nI was wrong to say fstat() before, it's a regular stat(). The\ninteresting backtrace is:\n\n  #0  __GI___stat64 (file=0x555555ab0bf0 \"/home/peff/tmp/.git/reftable/tables.list\", buf=0x7fffffffd630)\n      at ../sysdeps/unix/sysv/linux/stat64.c:29\n  #1  0x0000555555832ad6 in stack_uptodate (st=0x555555ab0ae0) at reftable/stack.c:575\n  #2  0x0000555555832c5c in reftable_stack_reload (st=0x555555ab0ae0) at reftable/stack.c:625\n  #3  0x000055555581f690 in backend_for (out=0x7fffffffd798, store=0x555555ab0900,\n      refname=0x555555aae178 \"refs/tags/foo-1\", rewritten_ref=0x7fffffffd780, reload=1) at refs/reftable-backend.c:283\n  #4  0x0000555555824957 in reftable_be_reflog_exists (ref_store=0x555555ab0900,\n      refname=0x555555aae178 \"refs/tags/foo-1\") at refs/reftable-backend.c:2266\n  #5  0x000055555581393a in refs_reflog_exists (refs=0x555555ab0900, refname=0x555555aae178 \"refs/tags/foo-1\")\n      at refs.c:2995\n  #6  0x000055555581f730 in should_write_log (refs=0x555555ab0900, refname=0x555555aae178 \"refs/tags/foo-1\")\n      at refs/reftable-backend.c:301\n  #7  0x00005555558224df in write_transaction_table (writer=0x5555569abe30, cb_data=0x5555569ab820)\n      at refs/reftable-backend.c:1511\n  #8  0x000055555583354d in reftable_addition_add (add=0x5555569ab050,\n      write_table=0x5555558220f0 <write_transaction_table>, arg=0x5555569ab820) at reftable/stack.c:902\n  #9  0x0000555555822b2b in reftable_be_transaction_finish (ref_store=0x555555ab0900, transaction=0x555555ab0c80,\n      err=0x7fffffffdbc0) at refs/reftable-backend.c:1633\n  #10 0x0000555555813161 in ref_transaction_commit (transaction=0x555555ab0c80, err=0x7fffffffdbc0) at refs.c:2769\n  #11 0x00005555556956c9 in update_refs_stdin (flags=0) at builtin/update-ref.c:789\n\nSo we are asking about reflogs for each ref under the \"only reflog if a\nlog already exists\" rule. Which means we can easily disable it by\nsetting core.logallrefupdates to \"always\", giving us a way to measure\nthe impact. So we can try:\n\n  git init --ref-format=reftable\n  blob=$(echo foo | git hash-object -w --stdin)\n  seq -f \"create refs/tags/foo-%g $blob\" 50000 >input\n  cp -a .git/reftable reftable.orig\n  hyperfine \\\n  \t-p 'rm -rf .git/reftable; cp -a reftable.orig .git/reftable' \\\n  \t-L config true,always \\\n  \t'git -c core.logallrefupdates={config} update-ref --stdin <input'\n\nwhich yields:\n\n  Benchmark 1: git -c core.logallrefupdates=true update-ref --stdin <input\n    Time (mean ± σ):     128.8 ms ±   1.5 ms    [User: 97.1 ms, System: 31.7 ms]\n    Range (min … max):   126.4 ms … 131.3 ms    23 runs\n  \n  Benchmark 2: git -c core.logallrefupdates=always update-ref --stdin <input\n    Time (mean ± σ):     195.2 ms ±   1.7 ms    [User: 182.1 ms, System: 13.0 ms]\n    Range (min … max):   191.9 ms … 197.6 ms    15 runs\n\nSo we saved all of those stat() calls, but writing the actual reflog\nentries adds much more cost!\n\nSadly there is no \"never\" option for core.logallrefupdates. I hacked one\nin and the result took ~96ms to run. So that tells us the cost of the\nstat() checks: around 25% of the runtime.\n\nWhich is a big-ish percentage, but a small absolute number. It's 0.64us\nper ref. Perhaps not worrying about too much.\n\nThe other interesting thing I noticed is that we seem to write() entries\nindividually with no buffering. Probably we could get an easy speedup\nfor big cases like this by using stdio in the fd_writer abstraction.\nWe're again pinching pennies to some degree, but I expect we may be able\nto drop the 120ms case down to 60ms or so (just guessing based on the\nextra reflog writes adding ~70ms).\n\nThere was one other oddity I didn't quite resolve. You may notice the\ngross reftable.orig stuff in hyperfine. I originally wrote this as:\n\n    git for-each-ref --format=\"delete %(refname)\" | git update-ref --stdin\n\nbut for some reason that causes the subsequent update-ref to loop\ninfinitely on merged_iter_next_entry(). It does so reliably, but I can't\nreproduce it outside of hyperfine. Super weird, and I'm sure I'm missing\nsomething obvious.\n\n-Peff\n"},{"id":"546816","messageId":"20260630235850.GB3759976@coredump.intra.peff.net","threadId":"65847","inReplyTo":"20260630234702.GA3759976@coredump.intra.peff.net","subject":"weird quadratic reftable behavior, was: Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-30T23:58:50Z","receivedAt":"2026-06-30T23:58:52Z","isPatch":true,"body":"On Tue, Jun 30, 2026 at 07:47:02PM -0400, Jeff King wrote:\n\n> There was one other oddity I didn't quite resolve. You may notice the\n> gross reftable.orig stuff in hyperfine. I originally wrote this as:\n> \n>     git for-each-ref --format=\"delete %(refname)\" | git update-ref --stdin\n> \n> but for some reason that causes the subsequent update-ref to loop\n> infinitely on merged_iter_next_entry(). It does so reliably, but I can't\n> reproduce it outside of hyperfine. Super weird, and I'm sure I'm missing\n> something obvious.\n\nAh, maybe not infinite, but probably quadratic. The key is that you have\nto delete a lot of refs and then try to insert them again. So with this\nscript:\n\n  nr=$1; shift\n  rm -rf .git\n  \n  git init --ref-format=reftable\n  blob=$(echo foo | git hash-object -w --stdin)\n  seq -f \"create refs/tags/foo-%g $blob\" $nr >input\n  git update-ref --stdin <input\n  git for-each-ref --format=\"delete %(refname)\" | git update-ref --stdin\n  time git update-ref --stdin <input\n\nI get results like this:\n\n  nr   | runtime\n  ------------\n  1000 | 0.125s\n  2000 | 0.454s\n  4000 | 1.811s\n  8000 | 7.091s\n\nSo for every doubling of the input size, the runtime quadruples. I guess\nit is iterating through some deleted tombstone entries, but I'm not sure\nwhy.\n\nThat's probably a more interesting and productive performance problem to\nwork on than micro-optimizing out the last few microseconds of writing. :)\n\n-Peff\n"},{"id":"546818","messageId":"akSxCUfm2P7ocLJX@pks.im","threadId":"65847","inReplyTo":"20260630235850.GB3759976@coredump.intra.peff.net","subject":"Re: weird quadratic reftable behavior, was: Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-01T06:17:45Z","receivedAt":"2026-07-01T06:17:52Z","isPatch":true,"body":"On Tue, Jun 30, 2026 at 07:58:50PM -0400, Jeff King wrote:\n> On Tue, Jun 30, 2026 at 07:47:02PM -0400, Jeff King wrote:\n> \n> > There was one other oddity I didn't quite resolve. You may notice the\n> > gross reftable.orig stuff in hyperfine. I originally wrote this as:\n> > \n> >     git for-each-ref --format=\"delete %(refname)\" | git update-ref --stdin\n> > \n> > but for some reason that causes the subsequent update-ref to loop\n> > infinitely on merged_iter_next_entry(). It does so reliably, but I can't\n> > reproduce it outside of hyperfine. Super weird, and I'm sure I'm missing\n> > something obvious.\n> \n> Ah, maybe not infinite, but probably quadratic. The key is that you have\n> to delete a lot of refs and then try to insert them again. So with this\n> script:\n> \n>   nr=$1; shift\n>   rm -rf .git\n>   \n>   git init --ref-format=reftable\n>   blob=$(echo foo | git hash-object -w --stdin)\n>   seq -f \"create refs/tags/foo-%g $blob\" $nr >input\n>   git update-ref --stdin <input\n>   git for-each-ref --format=\"delete %(refname)\" | git update-ref --stdin\n>   time git update-ref --stdin <input\n> \n> I get results like this:\n> \n>   nr   | runtime\n>   ------------\n>   1000 | 0.125s\n>   2000 | 0.454s\n>   4000 | 1.811s\n>   8000 | 7.091s\n> \n> So for every doubling of the input size, the runtime quadruples. I guess\n> it is iterating through some deleted tombstone entries, but I'm not sure\n> why.\n> \n> That's probably a more interesting and productive performance problem to\n> work on than micro-optimizing out the last few microseconds of writing. :)\n\nThis is a known issue, I think [1].\n\nThe problem here is the tombstoning: when you delete all references,\nchances are that they are not truly gone but that every reference is\njust tombstoned. The problem with this is that reading refs may now take\nsignifciantly more time as we cannot just say \"this stack is empty\".\nInstead, we need to figure out that it is empty by processing all the\ntombstones, and that takes a lot of time.\n\nI remember that I did some digging back then and improved the status quo\nquite significantly by optimizing `refs_verify_refname_available()`. I'm\nsure there are more opportunities for optimization here though -- I have\na feeling that we for example exhaust the merged iterator until its end\nwhen searching for a specific refname, where we could easily abort once\nthe observed tombstone name sorts lexicographically after the needle.\n\nBut eventually I decided to not care too much about this edge case, as\nit seems very specific to this artificial benchmark scenario. Which of\ncourse doesn't mean that it's not worth doing, I just had bigger fish to\nfry and didn't get around to it yet.\n\nPatrick\n\n[1]: <Z602dzQggtDdcgCX@tapette.crustytoothpaste.net>\n"},{"id":"546826","messageId":"akS6FZGSmwAA8Gdi@pks.im","threadId":"65847","inReplyTo":"xmqq8q7vsukl.fsf@gitster.g","subject":"Re: [PATCH 0/3] fixing expensive http test timeouts","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-01T06:56:21Z","receivedAt":"2026-07-01T06:56:26Z","isPatch":true,"body":"On Tue, Jun 30, 2026 at 12:21:30PM -0700, Junio C Hamano wrote:\n> Patrick Steinhardt <ps@pks.im> writes:\n> \n> > The pipelines of the official mirror can be found at [1]. We might for\n> > example add something like the below patch to our README.md to make it\n> > more discoverable.\n> >\n> > Patrick\n> >\n> > [1]: https://gitlab.com/git-scm/git/-/pipelines\n> >\n> > diff --git a/README.md b/README.md\n> > index d87bca1b8c..9ad77fdf7e 100644\n> > --- a/README.md\n> > +++ b/README.md\n> > @@ -1,4 +1,5 @@\n> > -[![Build status](https://github.com/git/git/workflows/CI/badge.svg)](https://github.com/git/git/actions?query=branch%3Amaster+event%3Apush)\n> > +[![GitHub build status](https://github.com/git/git/workflows/CI/badge.svg)](https://github.com/git/git/actions?query=branch%3Amaster+event%3Apush)\n> > +[![GitLab build status](https://gitlab.com/git-scm/git/badges/master/pipeline.svg)](https://gitlab.com/git-scm/git/-/pipelines?ref=master)\n> >  \n> >  Git - fast, scalable, distributed revision control system\n> >  =========================================================\n> \n> Oh, nice.  We of course do not want to be heavily involved in\n> advertising offerings by commercial entities but I think these two\n> sites deserve one line each for their continued service to the\n> community ;-)\n\nOkay, I'll send a patch then. Thanks!\n\nPatrick\n"},{"id":"546853","messageId":"20260701080014.GA3748390@coredump.intra.peff.net","threadId":"65847","inReplyTo":"akSxCUfm2P7ocLJX@pks.im","subject":"Re: weird quadratic reftable behavior, was: Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-07-01T08:00:14Z","receivedAt":"2026-07-01T08:00:16Z","isPatch":true,"body":"On Wed, Jul 01, 2026 at 08:17:45AM +0200, Patrick Steinhardt wrote:\n\n> This is a known issue, I think [1].\n> \n> The problem here is the tombstoning: when you delete all references,\n> chances are that they are not truly gone but that every reference is\n> just tombstoned. The problem with this is that reading refs may now take\n> signifciantly more time as we cannot just say \"this stack is empty\".\n> Instead, we need to figure out that it is empty by processing all the\n> tombstones, and that takes a lot of time.\n> \n> I remember that I did some digging back then and improved the status quo\n> quite significantly by optimizing `refs_verify_refname_available()`. I'm\n> sure there are more opportunities for optimization here though -- I have\n> a feeling that we for example exhaust the merged iterator until its end\n> when searching for a specific refname, where we could easily abort once\n> the observed tombstone name sorts lexicographically after the needle.\n\nYeah, it is (mostly) the same problem. About half the time is spent in\nrefs_verify_refnames_available().\n\nThe other half is in reftable_be_transaction_prepare(). Looks like it\nmakes individual calls to prepare_single_update(), which reads each ref.\nAnd those reads are expensive because of all of the tombstones. It might\nbe possible to do an iterator merge or similar between the sorted list\nof transaction refs and the reftable contents.\n\n> But eventually I decided to not care too much about this edge case, as\n> it seems very specific to this artificial benchmark scenario. Which of\n> course doesn't mean that it's not worth doing, I just had bigger fish to\n> fry and didn't get around to it yet.\n\nYeah, that's fair. I dug a bit further in case there was anything useful\nto write up, but I don't have much to add beyond what's here and in the\nthread you linked. We can let it live on in the archive for now.\n\n-Peff\n"},{"id":"546864","messageId":"CAL71e4PfXA-ixKR6r7fu_7_QmdzK+rTRs29mOsUYKaq+_a5q5w@mail.gmail.com","threadId":"65847","inReplyTo":"20260701080014.GA3748390@coredump.intra.peff.net","subject":"Re: weird quadratic reftable behavior, was: Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Kristofer Karlsson","fromEmail":"krka@spotify.com","sentAt":"2026-07-01T09:04:45Z","receivedAt":"2026-07-01T09:04:57Z","isPatch":true,"body":"On Wed, 1 Jul 2026 at 10:00, Jeff King <peff@peff.net> wrote:\n\n> Yeah, it is (mostly) the same problem. About half the time is spent in\n> refs_verify_refnames_available().\n>\n> The other half is in reftable_be_transaction_prepare(). Looks like it\n> makes individual calls to prepare_single_update(), which reads each ref.\n> And those reads are expensive because of all of the tombstones. It might\n> be possible to do an iterator merge or similar between the sorted list\n> of transaction refs and the reftable contents.\n\nHi, sorry for jumping in -- I found this interesting and started\npoking at the code. I think both halves may share the same root\ncause.\n\nThe merged iterator's suppress_deletions flag filters out tombstones\ninternally, which means higher-level code with prefix or refname\nbounds never gets a chance to stop iteration early. By letting\ntombstones pass through and filtering them one layer up in the\nreftable backend, the existing bounds checks can kick in before\nwe scan through all the tombstones.\n\nSo instead of doing full scans inside merged_iter_next_void()\nwe can just delegate to merged_iter_next_entry() and instead\nadd a loop to reftable_be_reflog_exists() that skips\ntombstones (but is amortized O(1)).\n\nNow multiple call sites would need to add something like this\nto compensate for returning tombstones:\n\n    if (reftable_log_record_is_deletion(&iter->log))\n        continue;\n\nbut it may be worth it if it reduces cost when there are many refs.\n\nThe key spot is reftable_ref_iterator_advance(), where the deletion\nskip goes right after the existing prefix check -- so a tombstone\npast the prefix stops iteration immediately instead of being\nsilently consumed. The same idea applies to reftable_backend_read_ref()\nand the log iteration paths.\n\nI have a local branch with this attempted fix. Rerunning the\nbenchmark:\n\n  Before:\n    nr=1000  0.306s\n    nr=2000  0.945s\n    nr=4000  3.816s\n    nr=8000  14.93s\n\n  After:\n    nr=1000   0.020s\n    nr=2000   0.044s\n    nr=4000   0.071s\n    nr=8000   0.145s\n    nr=16000  0.258s\n    nr=32000  0.591s\n\nI can send a proper patch if needed/wanted, but I might have missed\nsomething silly here.\n\nThanks,\nKristofer\n"},{"id":"546866","messageId":"akTm7BDohsy85sN8@pks.im","threadId":"65847","inReplyTo":"CAL71e4PfXA-ixKR6r7fu_7_QmdzK+rTRs29mOsUYKaq+_a5q5w@mail.gmail.com","subject":"Re: weird quadratic reftable behavior, was: Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-01T10:07:40Z","receivedAt":"2026-07-01T10:07:47Z","isPatch":true,"body":"On Wed, Jul 01, 2026 at 11:04:45AM +0200, Kristofer Karlsson wrote:\n> On Wed, 1 Jul 2026 at 10:00, Jeff King <peff@peff.net> wrote:\n> \n> > Yeah, it is (mostly) the same problem. About half the time is spent in\n> > refs_verify_refnames_available().\n> >\n> > The other half is in reftable_be_transaction_prepare(). Looks like it\n> > makes individual calls to prepare_single_update(), which reads each ref.\n> > And those reads are expensive because of all of the tombstones. It might\n> > be possible to do an iterator merge or similar between the sorted list\n> > of transaction refs and the reftable contents.\n> \n> Hi, sorry for jumping in -- I found this interesting and started\n> poking at the code. I think both halves may share the same root\n> cause.\n> \n> The merged iterator's suppress_deletions flag filters out tombstones\n> internally, which means higher-level code with prefix or refname\n> bounds never gets a chance to stop iteration early. By letting\n> tombstones pass through and filtering them one layer up in the\n> reftable backend, the existing bounds checks can kick in before\n> we scan through all the tombstones.\n\nAh, interesting. This kind of hits int o the direction that I proposed\nof stopping iteration once we hit the first reference that has a\nlexicographically-larget refname. But I thought about having to solve it\ngenerically in the iterator somehow, and I wasn't really looking forward\nto that.\n\nYour solutions works around that issue indeed. It of course doesn't\nsolve reference iteration, which would still be expensive. But it does\nsolve the issue where we just want to look up a single reference, which\nis what Peff noticed to be slow.\n\n> So instead of doing full scans inside merged_iter_next_void()\n> we can just delegate to merged_iter_next_entry() and instead\n> add a loop to reftable_be_reflog_exists() that skips\n> tombstones (but is amortized O(1)).\n> \n> Now multiple call sites would need to add something like this\n> to compensate for returning tombstones:\n> \n>     if (reftable_log_record_is_deletion(&iter->log))\n>         continue;\n> \n> but it may be worth it if it reduces cost when there are many refs.\n\nI think it certainly might be.\n\n> The key spot is reftable_ref_iterator_advance(), where the deletion\n> skip goes right after the existing prefix check -- so a tombstone\n> past the prefix stops iteration immediately instead of being\n> silently consumed. The same idea applies to reftable_backend_read_ref()\n> and the log iteration paths.\n\nYup.\n\n> I have a local branch with this attempted fix. Rerunning the\n> benchmark:\n> \n>   Before:\n>     nr=1000  0.306s\n>     nr=2000  0.945s\n>     nr=4000  3.816s\n>     nr=8000  14.93s\n> \n>   After:\n>     nr=1000   0.020s\n>     nr=2000   0.044s\n>     nr=4000   0.071s\n>     nr=8000   0.145s\n>     nr=16000  0.258s\n>     nr=32000  0.591s\n> \n> I can send a proper patch if needed/wanted, but I might have missed\n> something silly here.\n\nNice gains. I certainly think it would make sense to polish this a bit\nand then cast it into a patch.\n\nPatrick\n"},{"id":"546952","messageId":"CAC2Qwm+0-O6aL3bEN15+L+8EtVdF3msNARxhysfPCFxxdrBnPQ@mail.gmail.com","threadId":"65847","inReplyTo":"20260628080009.GA107826@coredump.intra.peff.net","subject":"Re: [PATCH 1/3] t/lib-httpd: bump apache timeout","fromName":"Michael Montalbo","fromEmail":"mmontalbo@gmail.com","sentAt":"2026-07-02T03:24:22Z","receivedAt":"2026-07-02T03:24:34Z","isPatch":true,"body":"On Sun, Jun 28, 2026 at 1:00 AM Jeff King <peff@peff.net> wrote:\n>\n> I didn't reference Michael's bugzilla report directly, because you can't\n> read it without a login. :(\n>\n> Maybe it's worth doing anyway?\n>\n\nI also thought the report being behind a login was unfortunate. For the\nhistorical record, I ended up submitting a patch[1] to their public GitHub\nmirror that describes the issue in more detail.\n\n[1] https://github.com/apache/httpd/pull/676\n"},{"id":"547074","messageId":"CAL71e4OavgfXtjN7QxkvmctS3fTpb5MtDsi-iUg=2izZCG5yxg@mail.gmail.com","threadId":"65847","inReplyTo":"akTm7BDohsy85sN8@pks.im","subject":"Re: weird quadratic reftable behavior, was: Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Kristofer Karlsson","fromEmail":"krka@spotify.com","sentAt":"2026-07-03T12:09:45Z","receivedAt":"2026-07-03T12:09:57Z","isPatch":true,"body":"On Wed, 1 Jul 2026 at 12:07, Patrick Steinhardt <ps@pks.im> wrote:\n> >\n> > I can send a proper patch if needed/wanted, but I might have missed\n> > something silly here.\n>\n> Nice gains. I certainly think it would make sense to polish this a bit\n> and then cast it into a patch.\n>\n> Patrick\n\nI have a small draft here https://github.com/gitgitgadget/git/pull/2166\nbut I am honestly not sure if it's worth submitting as a patch - the\nchange is somewhat small, but spread out, and I failed to properly\nreproduce the performance win in any realistic scenario (I had to\ndisable compaction to see the improvement).\n\nI would want to rely on your expertise to know if this change\nwould be valuable to discuss as a patch at all.\n\nThanks,\nKristofer\n"},{"id":"547198","messageId":"aktPP_aRI5Xfo4RA@pks.im","threadId":"65847","inReplyTo":"CAL71e4OavgfXtjN7QxkvmctS3fTpb5MtDsi-iUg=2izZCG5yxg@mail.gmail.com","subject":"Re: weird quadratic reftable behavior, was: Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-06T06:46:23Z","receivedAt":"2026-07-06T06:46:32Z","isPatch":true,"body":"On Fri, Jul 03, 2026 at 02:09:45PM +0200, Kristofer Karlsson wrote:\n> On Wed, 1 Jul 2026 at 12:07, Patrick Steinhardt <ps@pks.im> wrote:\n> > >\n> > > I can send a proper patch if needed/wanted, but I might have missed\n> > > something silly here.\n> >\n> > Nice gains. I certainly think it would make sense to polish this a bit\n> > and then cast it into a patch.\n> >\n> > Patrick\n> \n> I have a small draft here https://github.com/gitgitgadget/git/pull/2166\n> but I am honestly not sure if it's worth submitting as a patch - the\n> change is somewhat small, but spread out, and I failed to properly\n> reproduce the performance win in any realistic scenario (I had to\n> disable compaction to see the improvement).\n\nAn easy scenario where you don't have to disable compaction would be\nwhat Peff posted: you create X references and then delete all of them.\nThat shouldn't result in compaction and directly hits the case that we\ncare about.\n\n> I would want to rely on your expertise to know if this change\n> would be valuable to discuss as a patch at all.\n\nIf we can demonstrate a significant improvement in the above case then\nit would be worth it, I guess.\n\nPatrick\n"},{"id":"547215","messageId":"CAL71e4Pzq176Weu_R7vvCQUgqv5Z=88W82-K-oZdqxoSOQ0=rA@mail.gmail.com","threadId":"65847","inReplyTo":"aktPP_aRI5Xfo4RA@pks.im","subject":"Re: weird quadratic reftable behavior, was: Re: [PATCH 3/3] t5551: pack refs after creating many tags","fromName":"Kristofer Karlsson","fromEmail":"krka@spotify.com","sentAt":"2026-07-06T11:37:36Z","receivedAt":"2026-07-06T11:37:49Z","isPatch":true,"body":"On Mon, 6 Jul 2026 at 08:46, Patrick Steinhardt <ps@pks.im> wrote:\n>\n> An easy scenario where you don't have to disable compaction would be\n> what Peff posted: you create X references and then delete all of them.\n> That shouldn't result in compaction and directly hits the case that we\n> care about.\n\nRight, thanks. The recreate-same-refs case works with compaction\nenabled and shows the expected improvement (~100x for 8000 refs).\n\nI also found a worse case that feels more realistic: delete 8000\n\"old-*\" refs, then create 8000 \"new-*\" refs. Since \"new\" is\nlexicographically after \"old\", every create scans all tombstones.\nThat one goes from 27s to 0.09s after fixing it.\n\n> If we can demonstrate a significant improvement in the above case then\n> it would be worth it, I guess.\n\nI will clean up the patch and submit it shortly.\n\nThanks,\nKristofer\n"}]}