{"thread":{"id":"50506","subject":"RE: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","startedAt":"2019-02-14T15:05:11Z","lastAt":"2019-02-19T18:38:44Z","messageCount":13,"participants":["Randall S. Becker","Junio C Hamano","Johannes Schindelin","SZEDER Gábor","Max Kirillov"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"369312","messageId":"001501d4c476$a94651d0$fbd2f570$@nexbridge.com","threadId":"50506","inReplyTo":null,"subject":"RE: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-14T15:04:56Z","receivedAt":"2019-02-14T15:05:11Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 13, 2019 22:33, Junio C Hamano wrote:\n> A release candidate Git v2.21.0-rc1 is now available for testing at the usual\n> places.  It is comprised of 464 non-merge commits since v2.20.0, contributed\n> by 60 people, 14 of which are new faces.\n\nWe are currently running through a full regression of v2.21.0-rc1 on NonStop. It will take about 30 hours, but preliminary results, relative to breakages found in rc0 are:\n\nt1308 is fixed.\nt1404 is still broken (explainable) - scraping strerror output mismatches reported error on NonStop for EEXIST\nt5318 is fixed.\nt5403 is fixed.\nt5562 still hangs (blocking) - this breaks our CI pipeline since the test hangs and we have no explanation of whether the hang is in git or the tests.\n\nCheers,\nRandall\n\n-- Brief whoami:\n NonStop developer since approximately 211288444200000000\n UNIX developer since approximately 421664400\n-- In my real life, I talk too much.\n\n\n\n\n"},{"id":"369332","messageId":"xmqqk1i287ph.fsf@gitster-ct.c.googlers.com","threadId":"50506","inReplyTo":"001501d4c476$a94651d0$fbd2f570$@nexbridge.com","subject":"Re: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-02-14T19:56:42Z","receivedAt":"2019-02-14T19:56:47Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Randall S. Becker\" <rsbecker@nexbridge.com> writes:\n\n> On February 13, 2019 22:33, Junio C Hamano wrote:\n>> A release candidate Git v2.21.0-rc1 is now available for testing at the usual\n>> places.  It is comprised of 464 non-merge commits since v2.20.0, contributed\n>> by 60 people, 14 of which are new faces.\n>\n> We are currently running through a full regression of v2.21.0-rc1\n> on NonStop. It will take about 30 hours, but preliminary results,\n> relative to breakages found in rc0 are:\n>\n> t1308 is fixed.\n\nNice.\n\n> t1404 is still broken (explainable) - scraping strerror output\n> mismatches reported error on NonStop for EEXIST\n\nIIRC, the consensus was to loosen by not matching for the error\nmessage?  Let me take a look later today.\n\n> t5318 is fixed.\n> t5403 is fixed.\n\nGood.\n\n> t5562 still hangs (blocking) - this breaks our CI pipeline since\n> the test hangs and we have no explanation of whether the hang is\n> in git or the tests.\n\n"},{"id":"369350","messageId":"nycvar.QRO.7.76.6.1902142234070.45@tvgsbejvaqbjf.bet","threadId":"50506","inReplyTo":"001501d4c476$a94651d0$fbd2f570$@nexbridge.com","subject":"RE: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-02-14T21:36:42Z","receivedAt":"2019-02-14T21:37:02Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Randall,\n\nOn Thu, 14 Feb 2019, Randall S. Becker wrote:\n\n> t5562 still hangs (blocking) - this breaks our CI pipeline since the\n> test hangs and we have no explanation of whether the hang is in git or\n> the tests.\n\nI have \"good\" news: it now also hangs on Ubuntu 16.04 in Azure Pipelines'\nLinux agents.\n\nThere is a silver lining with those good news, though: I found a\nworkaround, and it might work for you, too:\n\n\thttps://github.com/gitgitgadget/git/pull/126\n\n(I also submitted this to the Git mailing list, as I really wanted to tag\nGit for Windows' v2.21.0-rc1.windows.1 only with a passing build, and I do\nnot want to keep that patch to the Windows port only.)\n\nCiao,\nJohannes\n"},{"id":"369355","messageId":"005501d4c4b4$39b68f90$ad23aeb0$@nexbridge.com","threadId":"50506","inReplyTo":"nycvar.QRO.7.76.6.1902142234070.45@tvgsbejvaqbjf.bet","subject":"RE: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-14T22:25:38Z","receivedAt":"2019-02-14T22:26:21Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 14, 2019 16:37, Johannes Schindelin wrote:\n> On Thu, 14 Feb 2019, Randall S. Becker wrote:\n> \n> > t5562 still hangs (blocking) - this breaks our CI pipeline since the\n> > test hangs and we have no explanation of whether the hang is in git or\n> > the tests.\n> \n> I have \"good\" news: it now also hangs on Ubuntu 16.04 in Azure Pipelines'\n> Linux agents.\n> \n> There is a silver lining with those good news, though: I found a\nworkaround,\n> and it might work for you, too:\n> \n> \thttps://github.com/gitgitgadget/git/pull/126\n> \n> (I also submitted this to the Git mailing list, as I really wanted to tag\nGit for\n> Windows' v2.21.0-rc1.windows.1 only with a passing build, and I do not\nwant\n> to keep that patch to the Windows port only.)\n\nThanks for trying. It was a good try, but did not fix the hang. See my other\nresponse for the stack trace. I tried debugging once it hung, but the code\nnever exits from the operating system, so I can't get inside. It is hiding\nin waitpid on a process that exists otherwise we would get an error (EINTR,\nECHILD, EFAULT are possible returns). One thing to consider is that we do\nnot have kernel threads, so if that is assumed, that is badness.\n\nRegards,\nRandall\n\n"},{"id":"369381","messageId":"20190215130213.GK1622@szeder.dev","threadId":"50506","inReplyTo":"nycvar.QRO.7.76.6.1902142234070.45@tvgsbejvaqbjf.bet","subject":"Re: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-02-15T13:02:13Z","receivedAt":"2019-02-15T13:02:20Z","isPatch":false,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Thu, Feb 14, 2019 at 10:36:42PM +0100, Johannes Schindelin wrote:\n> On Thu, 14 Feb 2019, Randall S. Becker wrote:\n> \n> > t5562 still hangs (blocking) - this breaks our CI pipeline since the\n> > test hangs and we have no explanation of whether the hang is in git or\n> > the tests.\n> \n> I have \"good\" news: it now also hangs on Ubuntu 16.04 in Azure Pipelines'\n> Linux agents.\n\nI haven't yet seen that hang in the wild and couldn't reproduce it on\npurpose, but there is definitely something fishy with t5562 even on\nLinux and even without that perl generate_zero_bytes helper.\n\n  $ git checkout cc95bc2025^\n  Previous HEAD position was cc95bc2025 t5562: replace /dev/zero with a pipe from generate_zero_bytes\n  HEAD is now at 24b451e77c t5318: replace use of /dev/zero with generate_zero_bytes\n  $ make\n  <snip>\n  $ cd t\n  # take note of the shell's PID\n  $ echo $$\n  15522\n  $ ./t5562-http-backend-content-length.sh --stress |tee LOG\n  OK    3.0\n  OK    1.0\n  OK    6.0\n  OK    0.0\n  <snap>\n\nAnd then in another terminal run this:\n\n  $ pstree -a -p 15522\n\nor, to make it easier noticable what changed and what stayed the same:\n\n  $ watch -d pstree -a -p 15522\n\nThe output will sooner or later will look like this:\n\n  bash,15522\n    └─t5562-http-back,21082 ./t5562-http-backend-content-length.sh --stress\n        ├─t5562-http-back,21089 ./t5562-http-backend-content-length.sh --stress\n        │   └─sh,24906 ./t5562-http-backend-content-length.sh --stress\n        ├─t5562-http-back,21090 ./t5562-http-backend-content-length.sh --stress\n        │   └─sh,26660 ./t5562-http-backend-content-length.sh --stress\n        ├─t5562-http-back,21092 ./t5562-http-backend-content-length.sh --stress\n        │   └─sh,4202 ./t5562-http-backend-content-length.sh --stress\n        │       └─sh,5696 ./t5562-http-backend-content-length.sh --stress\n        │           └─perl,5697 /home/szeder/src/git/t/t5562/invoke-with-content-length.pl push_body.gz.trunc git http-backend\n        │               └─(git,5722)\n        ├─t5562-http-back,21093 ./t5562-http-backend-content-length.sh --stress\n        │   └─sh,25572 ./t5562-http-backend-content-length.sh --stress\n  <snip>\n\nIt won't show most of the processes run in the tests, because they are\njust too fast and short-lived.  However, occasionally it does show a\nstuck git process, which is shown as <defunct> in regular 'ps aux'\noutput:\n\n  szeder   5722  0.0  0.0      0     0 pts/16   Z+   13:36   0:00 [git] <defunct>\n\nNote that this is not a \"proper\" hang, in the sense that this process\nis not stuck forever, but only for about 1 minute, after which it\ndisappears, and the test continues and eventually finishes with\nsuccess.  I've looked into the logs of a couple of such stuck jobs,\nand it seems that it varies in which test that git process happened to\nget stuck.\n\n\n \n"},{"id":"369382","messageId":"000801d4c535$54796f10$fd6c4d30$@nexbridge.com","threadId":"50506","inReplyTo":"20190215130213.GK1622@szeder.dev","subject":"RE: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-15T13:49:48Z","receivedAt":"2019-02-15T13:50:05Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 15, 2019 8:02, SZEDER Gábor wrote:\n> To: Johannes Schindelin <Johannes.Schindelin@gmx.de>\n> Cc: Randall S. Becker <rsbecker@nexbridge.com>; 'Junio C Hamano'\n> <gitster@pobox.com>; git@vger.kernel.org; 'Max Kirillov'\n> <max@max630.net>\n> Subject: Re: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)\n> \n> On Thu, Feb 14, 2019 at 10:36:42PM +0100, Johannes Schindelin wrote:\n> > On Thu, 14 Feb 2019, Randall S. Becker wrote:\n> >\n> > > t5562 still hangs (blocking) - this breaks our CI pipeline since the\n> > > test hangs and we have no explanation of whether the hang is in git\n> > > or the tests.\n> >\n> > I have \"good\" news: it now also hangs on Ubuntu 16.04 in Azure Pipelines'\n> > Linux agents.\n> \n> I haven't yet seen that hang in the wild and couldn't reproduce it on purpose,\n> but there is definitely something fishy with t5562 even on Linux and even\n> without that perl generate_zero_bytes helper.\n> \n>   $ git checkout cc95bc2025^\n>   Previous HEAD position was cc95bc2025 t5562: replace /dev/zero with a\n> pipe from generate_zero_bytes\n>   HEAD is now at 24b451e77c t5318: replace use of /dev/zero with\n> generate_zero_bytes\n>   $ make\n>   <snip>\n>   $ cd t\n>   # take note of the shell's PID\n>   $ echo $$\n>   15522\n>   $ ./t5562-http-backend-content-length.sh --stress |tee LOG\n>   OK    3.0\n>   OK    1.0\n>   OK    6.0\n>   OK    0.0\n>   <snap>\n> \n> And then in another terminal run this:\n> \n>   $ pstree -a -p 15522\n> \n> or, to make it easier noticable what changed and what stayed the same:\n> \n>   $ watch -d pstree -a -p 15522\n> \n> The output will sooner or later will look like this:\n> \n>   bash,15522\n>     └─t5562-http-back,21082 ./t5562-http-backend-content-length.sh --stress\n>         ├─t5562-http-back,21089 ./t5562-http-backend-content-length.sh --\n> stress\n>         │   └─sh,24906 ./t5562-http-backend-content-length.sh --stress\n>         ├─t5562-http-back,21090 ./t5562-http-backend-content-length.sh --\n> stress\n>         │   └─sh,26660 ./t5562-http-backend-content-length.sh --stress\n>         ├─t5562-http-back,21092 ./t5562-http-backend-content-length.sh --\n> stress\n>         │   └─sh,4202 ./t5562-http-backend-content-length.sh --stress\n>         │       └─sh,5696 ./t5562-http-backend-content-length.sh --stress\n>         │           └─perl,5697 /home/szeder/src/git/t/t5562/invoke-with-content-\n> length.pl push_body.gz.trunc git http-backend\n>         │               └─(git,5722)\n>         ├─t5562-http-back,21093 ./t5562-http-backend-content-length.sh --\n> stress\n>         │   └─sh,25572 ./t5562-http-backend-content-length.sh --stress\n>   <snip>\n> \n> It won't show most of the processes run in the tests, because they are just\n> too fast and short-lived.  However, occasionally it does show a stuck git\n> process, which is shown as <defunct> in regular 'ps aux'\n> output:\n> \n>   szeder   5722  0.0  0.0      0     0 pts/16   Z+   13:36   0:00 [git] <defunct>\n> \n> Note that this is not a \"proper\" hang, in the sense that this process is not\n> stuck forever, but only for about 1 minute, after which it disappears, and the\n> test continues and eventually finishes with success.  I've looked into the logs\n> of a couple of such stuck jobs, and it seems that it varies in which test that git\n> process happened to get stuck.\n\nWe see something similar. The 60 seconds is in the support script in the t/t5562 directory. If a SIGCHLD is received, the sleep is interrupted and perl terminates (no hang). If the sleep is not interrupted, NonStop hangs in the close() after coming out of sleep because perl still has output to send somewhere. We are hung in the close call - which is really perplexing considering a close on NonStop in any other product is immediate and rather harsh, but perl's semantics for close() are: \"Closing a pipe also waits for the process executing on the pipe to complete\" (from the Perl spec), which seems to apply on NonStop because the git (5722) is reading but not receiving any data and not terminating - based on your tree above. Or, in other words, perl closing the pipe will not cause git (5722) to terminate because perl is waiting on git (5722) to terminate before completing the close. The only time it would not hang is if git (5722) terminates on its own so that sleep is interrupted without going back for more data to read. I am making a semi-educated guess. From my experience with the NS perl team, they are going to point at that spec and say that perl is exhibiting the correct behaviour and that the hang is expected.\n\nAnother weird observation is that the test generates up to three hangs (subtests 6,8,13) at worst, and one (subtest 13) at best, depending on some unknown factor that might be system load. This is hinting at a race condition. Sadly, we don't have the above cool watch or pstree utilities on platform. What we do have is something called ptrace, which can look at the stack and I/O conditions of all open files and whether there are outstanding I/Os (and how many) on each FD, memory use.\n\n"},{"id":"369417","messageId":"20190215203726.GG3064@jessie.local","threadId":"50506","inReplyTo":"20190215130213.GK1622@szeder.dev","subject":"Re: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","fromName":"Max Kirillov","fromEmail":"max@max630.net","sentAt":"2019-02-15T20:37:26Z","receivedAt":"2019-02-15T20:37:33Z","isPatch":false,"sender":{"key":"max@max630.net","avatar":"https://avatars.githubusercontent.com/u/381560?v=4"},"body":"On Fri, Feb 15, 2019 at 02:02:13PM +0100, SZEDER Gábor wrote:\n> I haven't yet seen that hang in the wild and couldn't reproduce it on\n> purpose, but there is definitely something fishy with t5562 even on\n> Linux and even without that perl generate_zero_bytes helper.\n> \n> It won't show most of the processes run in the tests, because they are\n> just too fast and short-lived.  However, occasionally it does show a\n> stuck git process, which is shown as <defunct> in regular 'ps aux'\n> output:\n> \n>   szeder   5722  0.0  0.0      0     0 pts/16   Z+   13:36   0:00 [git] <defunct>\n> \n> Note that this is not a \"proper\" hang, in the sense that this process\n> is not stuck forever, but only for about 1 minute\n\nThis is probably because of SIGCHILD comes before \"sleep\". I believe this is\nunrelated to the hang issue. The hang issue looks like something is wrong with\ncleanu_children(), or maybe in the child which it tries to kill and wait, not in\ntests.\n\nAs for this zombie issue, could be fixed with, for example, more busy wait like\nthe following. It may with some bigger probability miss SIGCHILD to the first\nsleep because there is a bit more to do before it. But the penalty is only 1\nsecond now, and as it still happens rarely there seems to be no visible\ndegradation.\n\n--- 8< -----------\ndiff --git a/t/t5562/invoke-with-content-length.pl b/t/t5562/invoke-with-content-length.pl\nindex 0943474af2..257e280e3b 100644\n--- a/t/t5562/invoke-with-content-length.pl\n+++ b/t/t5562/invoke-with-content-length.pl\n@@ -29,7 +29,12 @@\n }\n print $out $body_data or die \"Cannot write data: $!\";\n \n-sleep 60; # is interrupted by SIGCHLD\n+my $counter = 0;\n+while (not $exited and $counter < 60) {\n+        sleep 1;\n+        $counter = $counter + 1;\n+}\n+\n if (!$exited) {\n         close($out);\n         die \"Command did not exit after reading whole body\";\n"},{"id":"369418","messageId":"002a01d4c573$4784e670$d68eb350$@nexbridge.com","threadId":"50506","inReplyTo":"20190215203726.GG3064@jessie.local","subject":"RE: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-15T21:13:15Z","receivedAt":"2019-02-15T21:13:33Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 15, 2019 15:37, Max Kirillov wrote:\n> On Fri, Feb 15, 2019 at 02:02:13PM +0100, SZEDER Gábor wrote:\n> > I haven't yet seen that hang in the wild and couldn't reproduce it on\n> > purpose, but there is definitely something fishy with t5562 even on\n> > Linux and even without that perl generate_zero_bytes helper.\n> >\n> > It won't show most of the processes run in the tests, because they are\n> > just too fast and short-lived.  However, occasionally it does show a\n> > stuck git process, which is shown as <defunct> in regular 'ps aux'\n> > output:\n> >\n> >   szeder   5722  0.0  0.0      0     0 pts/16   Z+   13:36   0:00 [git] <defunct>\n> >\n> > Note that this is not a \"proper\" hang, in the sense that this process\n> > is not stuck forever, but only for about 1 minute\n> \n> This is probably because of SIGCHILD comes before \"sleep\". I believe this is\n> unrelated to the hang issue. The hang issue looks like something is wrong\n> with cleanu_children(), or maybe in the child which it tries to kill and wait,\n> not in tests.\n> \n> As for this zombie issue, could be fixed with, for example, more busy wait\n> like the following. It may with some bigger probability miss SIGCHILD to the\n> first sleep because there is a bit more to do before it. But the penalty is only\n> 1 second now, and as it still happens rarely there seems to be no visible\n> degradation.\n> \n> --- 8< -----------\n> diff --git a/t/t5562/invoke-with-content-length.pl b/t/t5562/invoke-with-\n> content-length.pl\n> index 0943474af2..257e280e3b 100644\n> --- a/t/t5562/invoke-with-content-length.pl\n> +++ b/t/t5562/invoke-with-content-length.pl\n> @@ -29,7 +29,12 @@\n>  }\n>  print $out $body_data or die \"Cannot write data: $!\";\n> \n> -sleep 60; # is interrupted by SIGCHLD\n> +my $counter = 0;\n> +while (not $exited and $counter < 60) {\n> +        sleep 1;\n> +        $counter = $counter + 1;\n> +}\n> +\n>  if (!$exited) {\n>          close($out);\n>          die \"Command did not exit after reading whole body\";\n\nFrom the trace I found in perl, we have gone past sleep and are hung at \n          close($out);\n\nCommenting out the close() does nothing because perl still hangs on an implied close resulting from the exception thrown by die(). See my other post on adding GIT_TRACE and the changes resulting from that.\n\nSadly, the fix does not change the results. In fact, it makes the hang far more likely. Subtest 6,7,8 fails here, at close()\n  waitpid + 0x130 (SLr)\n  $n_EnterPriv + 0x280 (Milli)\n  Perl_wait4pid + 0x130 (UCr)\n  Perl_my_pclose + 0x4C0 (UCr)\n  Perl_io_close + 0x180 (UCr)\n  Perl_do_close + 0x620 (UCr)\n  Perl_pp_close + 0xA70 (UCr)\n  Perl_runops_standard + 0xF0 (UCr)\n  S_run_body + 0x870 (UCr)\n  perl_run + 0x2D0 (UCr)\n  main + 0x3D0 (UCr)\n\n\n\n"},{"id":"369433","messageId":"20190216082635.GH3064@jessie.local","threadId":"50506","inReplyTo":"002a01d4c573$4784e670$d68eb350$@nexbridge.com","subject":"Re: [ANNOUNCE] Git v2.21.0-rc1 (NonStop Results)","fromName":"Max Kirillov","fromEmail":"max@max630.net","sentAt":"2019-02-16T08:26:35Z","receivedAt":"2019-02-16T08:29:10Z","isPatch":false,"sender":{"key":"max@max630.net","avatar":"https://avatars.githubusercontent.com/u/381560?v=4"},"body":"On Fri, Feb 15, 2019 at 04:13:15PM -0500, Randall S. Becker wrote:\n> Sadly, the fix does not change the results. In fact, it\n> makes the hang far more likely. Subtest 6,7,8 fails here,\n> at close()\n\nCorrect, I did not expect it to help, it was for the other\nissue.\n\nAs for the hang issue, from your another message it seems to\nme that perl waiting correctly, there are really child\nprocess which do not exit.\n\nWhat you could try is\nhttps://public-inbox.org/git/20181124093719.10705-1-max@max630.net/\n(I'm not sure it would not conflict by now), this would\nremove dependency between tests. If it helps it would be\nvery valuable information.\n"},{"id":"369557","messageId":"20190218205028.32486-1-max@max630.net","threadId":"50506","inReplyTo":"20190215130213.GK1622@szeder.dev","subject":"[PATCH] t5562: chunked sleep to avoid lost SIGCHILD","fromName":"Max Kirillov","fromEmail":"max@max630.net","sentAt":"2019-02-18T20:50:28Z","receivedAt":"2019-02-18T20:50:39Z","isPatch":true,"sender":{"key":"max@max630.net","avatar":"https://avatars.githubusercontent.com/u/381560?v=4"},"body":"If was found during stress-test run that a test may hang by 60 seconds.\nIt supposedly happens because SIGCHILD was received before sleep has\nstarted.\n\nFix by looping by smaller chunks, checking $exited after each of them.\nThen lost SIGCHILD would not cause longer delay than 1 second.\n\nReported-by: SZEDER Gábor <szeder.dev@gmail.com>\nSigned-off-by: Max Kirillov <max@max630.net>\n---\nSubmitting as proper patch. Note: I believe it does not relate to other issues\ndiscussed in this thread.\n t/t5562/invoke-with-content-length.pl | 7 ++++++-\n 1 file changed, 6 insertions(+), 1 deletion(-)\n\ndiff --git a/t/t5562/invoke-with-content-length.pl b/t/t5562/invoke-with-content-length.pl\nindex 0943474af2..257e280e3b 100644\n--- a/t/t5562/invoke-with-content-length.pl\n+++ b/t/t5562/invoke-with-content-length.pl\n@@ -29,7 +29,12 @@\n }\n print $out $body_data or die \"Cannot write data: $!\";\n \n-sleep 60; # is interrupted by SIGCHLD\n+my $counter = 0;\n+while (not $exited and $counter < 60) {\n+        sleep 1;\n+        $counter = $counter + 1;\n+}\n+\n if (!$exited) {\n         close($out);\n         die \"Command did not exit after reading whole body\";\n-- \n2.19.0.1202.g68e1e8f04e\n\n"},{"id":"369558","messageId":"005101d4c7cc$26a3c5b0$73eb5110$@nexbridge.com","threadId":"50506","inReplyTo":"20190218205028.32486-1-max@max630.net","subject":"RE: [PATCH] t5562: chunked sleep to avoid lost SIGCHILD","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-18T20:54:27Z","receivedAt":"2019-02-18T20:54:42Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 18, 2019 15:50, Max Kirillov wrote:\n> To: SZEDER Gábor <szeder.dev@gmail.com>; git@vger.kernel.org\n> Cc: Max Kirillov <max@max630.net>; Johannes Schindelin\n> <Johannes.Schindelin@gmx.de>; Randall S. Becker\n> <rsbecker@nexbridge.com>; 'Junio C Hamano' <gitster@pobox.com>\n> Subject: [PATCH] t5562: chunked sleep to avoid lost SIGCHILD\n> \n> If was found during stress-test run that a test may hang by 60 seconds.\n> It supposedly happens because SIGCHILD was received before sleep has\n> started.\n> \n> Fix by looping by smaller chunks, checking $exited after each of them.\n> Then lost SIGCHILD would not cause longer delay than 1 second.\n> \n> Reported-by: SZEDER Gábor <szeder.dev@gmail.com>\n> Signed-off-by: Max Kirillov <max@max630.net>\n> ---\n> Submitting as proper patch. Note: I believe it does not relate to other issues\n> discussed in this thread.\n>  t/t5562/invoke-with-content-length.pl | 7 ++++++-\n>  1 file changed, 6 insertions(+), 1 deletion(-)\n> \n> diff --git a/t/t5562/invoke-with-content-length.pl b/t/t5562/invoke-with-\n> content-length.pl\n> index 0943474af2..257e280e3b 100644\n> --- a/t/t5562/invoke-with-content-length.pl\n> +++ b/t/t5562/invoke-with-content-length.pl\n> @@ -29,7 +29,12 @@\n>  }\n>  print $out $body_data or die \"Cannot write data: $!\";\n> \n> -sleep 60; # is interrupted by SIGCHLD\n> +my $counter = 0;\n> +while (not $exited and $counter < 60) {\n> +        sleep 1;\n> +        $counter = $counter + 1;\n> +}\n> +\n>  if (!$exited) {\n>          close($out);\n>          die \"Command did not exit after reading whole body\";\n\nI tried this fix and it made no difference to the hang on NonStop. I do not think this fixes the root cause as sleep was never an issue and SIGCHLD was not missed in any test I conducted. Maybe on another platform it is required.\n\n"},{"id":"369561","messageId":"20190218205900.GC3373@jessie.local","threadId":"50506","inReplyTo":"005101d4c7cc$26a3c5b0$73eb5110$@nexbridge.com","subject":"Re: [PATCH] t5562: chunked sleep to avoid lost SIGCHILD","fromName":"Max Kirillov","fromEmail":"max@max630.net","sentAt":"2019-02-18T20:59:00Z","receivedAt":"2019-02-18T20:59:06Z","isPatch":true,"sender":{"key":"max@max630.net","avatar":"https://avatars.githubusercontent.com/u/381560?v=4"},"body":"On Mon, Feb 18, 2019 at 03:54:27PM -0500, Randall S. Becker wrote:\n> On February 18, 2019 15:50, Max Kirillov wrote:\n> > To: SZEDER Gábor <szeder.dev@gmail.com>; git@vger.kernel.org\n> > Cc: Max Kirillov <max@max630.net>; Johannes Schindelin\n> > <Johannes.Schindelin@gmx.de>; Randall S. Becker\n> > <rsbecker@nexbridge.com>; 'Junio C Hamano' <gitster@pobox.com>\n> > Subject: [PATCH] t5562: chunked sleep to avoid lost SIGCHILD\n> > \n> > If was found during stress-test run that a test may hang by 60 seconds.\n> > It supposedly happens because SIGCHILD was received before sleep has\n> > started.\n> > \n> > Fix by looping by smaller chunks, checking $exited after each of them.\n> > Then lost SIGCHILD would not cause longer delay than 1 second.\n> > \n> > Reported-by: SZEDER Gábor <szeder.dev@gmail.com>\n> > Signed-off-by: Max Kirillov <max@max630.net>\n> > ---\n> > Submitting as proper patch. Note: I believe it does not relate to other issues\n> > discussed in this thread.\n> >  t/t5562/invoke-with-content-length.pl | 7 ++++++-\n> >  1 file changed, 6 insertions(+), 1 deletion(-)\n> > \n> > diff --git a/t/t5562/invoke-with-content-length.pl b/t/t5562/invoke-with-\n> > content-length.pl\n> > index 0943474af2..257e280e3b 100644\n> > --- a/t/t5562/invoke-with-content-length.pl\n> > +++ b/t/t5562/invoke-with-content-length.pl\n> > @@ -29,7 +29,12 @@\n> >  }\n> >  print $out $body_data or die \"Cannot write data: $!\";\n> > \n> > -sleep 60; # is interrupted by SIGCHLD\n> > +my $counter = 0;\n> > +while (not $exited and $counter < 60) {\n> > +        sleep 1;\n> > +        $counter = $counter + 1;\n> > +}\n> > +\n> >  if (!$exited) {\n> >          close($out);\n> >          die \"Command did not exit after reading whole body\";\n> \n> I tried this fix and it made no difference to the hang on\n> NonStop. I do not think this fixes the root cause as sleep\n> was never an issue and SIGCHLD was not missed in any test\n> I conducted. Maybe on another platform it is required.\n\nCorrect, as I said it should not be related.\n"},{"id":"369657","messageId":"xmqq1s4339ow.fsf@gitster-ct.c.googlers.com","threadId":"50506","inReplyTo":"20190218205028.32486-1-max@max630.net","subject":"Re: [PATCH] t5562: chunked sleep to avoid lost SIGCHILD","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-02-19T18:38:39Z","receivedAt":"2019-02-19T18:38:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Max Kirillov <max@max630.net> writes:\n\n> If was found during stress-test run that a test may hang by 60 seconds.\n> It supposedly happens because SIGCHILD was received before sleep has\n> started.\n>\n> Fix by looping by smaller chunks, checking $exited after each of them.\n> Then lost SIGCHILD would not cause longer delay than 1 second.\n>\n> Reported-by: SZEDER Gábor <szeder.dev@gmail.com>\n> Signed-off-by: Max Kirillov <max@max630.net>\n> ---\n> Submitting as proper patch. Note: I believe it does not relate to other issues\n> discussed in this thread.\n>  t/t5562/invoke-with-content-length.pl | 7 ++++++-\n>  1 file changed, 6 insertions(+), 1 deletion(-)\n>\n> diff --git a/t/t5562/invoke-with-content-length.pl b/t/t5562/invoke-with-content-length.pl\n> index 0943474af2..257e280e3b 100644\n> --- a/t/t5562/invoke-with-content-length.pl\n> +++ b/t/t5562/invoke-with-content-length.pl\n> @@ -29,7 +29,12 @@\n>  }\n>  print $out $body_data or die \"Cannot write data: $!\";\n>  \n> -sleep 60; # is interrupted by SIGCHLD\n\nAh, of course.  If SIGCHLD interrupts, sets $existed in the handler,\nthen we won't go back to sleep.  But if the signal came before the\nsleep starts, we spend full 60 seconds here before we check $exited.\n\nMakes sense.\n\n> +my $counter = 0;\n> +while (not $exited and $counter < 60) {\n> +        sleep 1;\n> +        $counter = $counter + 1;\n> +}\n> +\n>  if (!$exited) {\n>          close($out);\n>          die \"Command did not exit after reading whole body\";\n"}]}