From: Ramsay Jones Date: Thu, 06 Nov 2025 20:28:35 GMT Subject: Re: v2.52.0-rc0 test failure on cygwin Message-ID: In-Reply-To: On 06/11/2025 10:53 am, Patrick Steinhardt wrote: > On Tue, Nov 04, 2025 at 11:49:46PM +0000, Ramsay Jones wrote: >> Just a quick heads up: the rc0 build on cygwin has a flaky test, thus: >> >> $ tail test-out-2-52-rc0 >> Test Summary Report >> ------------------- >> t0610-reftable-basics.sh (Wstat: 256 (exited 1) Tests: 90 Failed: 1) >> Failed test: 29 >> Non-zero exit status: 1 >> Files=1024, Tests=32232, 2703 wallclock secs (23.38 usr 60.53 sys + 7886.88 cusr 10419.88 csys = 18390.67 CPU) >> Result: FAIL >> make[1]: *** [Makefile:78: prove] Error 1 >> make[1]: Leaving directory '/home/ramsay/git/t' >> make: *** [Makefile:3327: test] Error 2 >> $ >> >> Initially, while investigating the failure, I was running the test by hand and it >> didn't fail ... So, I tried a stess test, like so: > > Interesting. My first hunch is that the root cause is auto-maintenance. > git-maintenance(1) spawns `git pack-refs --auto`, and that process will > open the stack so that it can verify whether it needs to be packed or > not. And Windows being Windows, the file being open may mean that it > cannot be written by another process at the same point in time. > > In any case, I was able to reproduce the issue. But disabling auto > maintenance with the following patch does not fix the flake. > > diff --git a/t/t0610-reftable-basics.sh b/t/t0610-reftable-basics.sh > index 3ea5d51532..52bbf4fe57 100755 > --- a/t/t0610-reftable-basics.sh > +++ b/t/t0610-reftable-basics.sh > @@ -204,6 +204,7 @@ test_expect_success 'ref transaction: corrupted tables cause failure' ' > git init repo && > ( > cd repo && > + git config set maintenance.auto false && > test_commit file1 && > for f in .git/reftable/*.ref > do > Thanks for looking into this - yesterday was unexpectedly busy, so I didn't have time to look at this myself. :( I also thought, briefly, about 'git maintenance' since the error seems to happen in parallel heavy workloads. You probably didn't notice that the test finished in 45min, because I ran the test with '-j8'. I have recently replaced my win10 laptop. On my old laptop, the non-parallel test run used to take 6+ hours. With my new laptop it is 4+ hours, so it is still a long time to wait. However, the 'meson test', which by default runs the tests in parallel, was much faster (about 80-90min). So, it was worth a try... At the moment the parallel tests hang about half of the time (prove hangs right at the very end!), so I am still experimenting. [snip] > I wonder whether the issue is surfaced because we use the shell to > truncate the file. If you instead use `file-tool truncate 0` for example > then I cannot reproduce the flake anymore: > > diff --git a/t/t0610-reftable-basics.sh b/t/t0610-reftable-basics.sh > index 3ea5d51532..1058f83993 100755 > --- a/t/t0610-reftable-basics.sh > +++ b/t/t0610-reftable-basics.sh > @@ -207,7 +207,7 @@ test_expect_success 'ref transaction: corrupted tables cause failure' ' > test_commit file1 && > for f in .git/reftable/*.ref > do > - : >"$f" || return 1 > + test-tool truncate "$f" 0 || return 1 > done && > test_must_fail git update-ref refs/heads/main HEAD > ) > > But this may very well just be due to timing again -- spawning the > process will be slower than using shell redirection to trim the file. I tried this patch tonight, letting: $ ./t0610-reftable-basics.sh --run=29 --stress-limit=10 finish, which it did without failure. So that's 32 * 10 successful runs. (I had expected 16 * 10 yesterday, ie 2 * cores * 10, but this laptop has 8 cores 16 threads, so 'getconf _NPROCESSORS_ONLN' returns 16 not 8). > > All of this is quite curious. I don't really have any better idea than > to use something like the above patch. It's ugly, doubly so because I > don't understand either the root cause nor why the patch properly fixes > it. So I'd be grateful if anyone were to enlighten me :) Me too! :) > I have verified that the flake already exists in Git 2.51, so at least > it's not a regression in the current release cycle. OK, that's good to know. Despite the mystery, I think a patch based on the above would be the best solution for now. (Assuming nobody has a better idea). Thanks. ATB, Ramsay Jones