{"thread":{"id":"50513","subject":"[PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","startedAt":"2019-02-14T21:33:15Z","lastAt":"2019-02-19T14:13:42Z","messageCount":21,"participants":["Johannes Schindelin via GitGitGadget","Junio C Hamano","Randall S. Becker","Max Kirillov","Johannes Schindelin","Ævar Arnfjörð Bjarmason"],"isPatch":true,"patchVersion":1,"patchTotal":1},"messages":[{"id":"369348","messageId":"pull.126.git.gitgitgadget@gmail.com","threadId":"50513","inReplyTo":null,"subject":"[PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-02-14T21:33:11Z","receivedAt":"2019-02-14T21:33:15Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"The last-minute patch to replace /dev/zero with a Perl script snippet broke\nthe Linux part of the CI builds on Azure Pipelines: it timed out. The\nculprit is the rb/no-dev-zero-in-test branch (see the build for this branch \nhere [https://dev.azure.com/gitgitgadget/git/_build/results?buildId=1727]).\n\nAll of master, next, jch and pu are broken that way. You might see it in the\ncommit status of the active branches\n[https://github.com/gitgitgadget/git/branches/active].\n\nTurns out that it is that particular Perl script snippet which for some\nreason hangs the build. If you kill it, t5562.15 succeeds, if you don't kill\nit, it will hang indefinitely (or until killed).\n\nSadly, despite my earnest attempts, I could not figure out why it hangs in\nthose Linux agents (I could not reproduce that hang locally), or for that\nmatter, why it does not hang in the Windows and macOS agents.\n\nLet's avoid that hang. This patch fixes things on Azure Pipelines, and my\nhope is that it also fixes the hang on NonStop.\n\nJohannes Schindelin (1):\n  tests: teach the test-tool to generate NUL bytes and use it\n\n Makefile                               |  1 +\n t/helper/test-genzeros.c               | 22 ++++++++++++++++++++++\n t/helper/test-tool.c                   |  1 +\n t/helper/test-tool.h                   |  1 +\n t/t5562-http-backend-content-length.sh |  2 +-\n t/test-lib-functions.sh                |  8 +-------\n 6 files changed, 27 insertions(+), 8 deletions(-)\n create mode 100644 t/helper/test-genzeros.c\n\n\nbase-commit: 8989e1950a845ceeb186d490321a4f917ca4de47\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-126%2Fdscho%2Ffix-t5562-hang-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-126/dscho/fix-t5562-hang-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/126\n-- \ngitgitgadget\n"},{"id":"369349","messageId":"34cde0f2849a098c17ab83786da5ce06f69cfafa.1550179990.git.gitgitgadget@gmail.com","threadId":"50513","inReplyTo":"pull.126.git.gitgitgadget@gmail.com","subject":"[PATCH 1/1] tests: teach the test-tool to generate NUL bytes and use it","fromName":"Johannes Schindelin via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-02-14T21:33:12Z","receivedAt":"2019-02-14T21:33:17Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"From: Johannes Schindelin <johannes.schindelin@gmx.de>\n\nIn cc95bc2025 (t5562: replace /dev/zero with a pipe from\ngenerate_zero_bytes, 2019-02-09), we replaced usage of /dev/zero (which\nis not available on NonStop, apparently) by a Perl script snippet to\ngenerate NUL bytes.\n\nSadly, it does not seem to work on NonStop, as t5562 reportedly hangs.\n\nWorse, this also hangs in the Ubuntu 16.04 agents of the CI builds on\nAzure Pipelines: for some reason, the Perl script snippet that is run\nvia `generate_zero_bytes` in t5562's 'CONTENT_LENGTH overflow ssite_t'\ntest case tries to write out an infinite amount of NUL bytes unless a\nbroken pipe is encountered, that snippet never encounters the broken\npipe, and keeps going until the build times out.\n\nOddly enough, this does not reproduce on the Windows and macOS agents,\nnor in a local Ubuntu 18.04.\n\nThis developer tried for a day to figure out the exact circumstances\nunder which this hang happens, to no avail, the details remain a\nmystery.\n\nIn the end, though, what counts is that this here change incidentally\nfixes that hang (maybe also on NonStop?). Even more positively, it gets\nrid of yet another unnecessary Perl invocation.\n\nSigned-off-by: Johannes Schindelin <johannes.schindelin@gmx.de>\n---\n Makefile                               |  1 +\n t/helper/test-genzeros.c               | 22 ++++++++++++++++++++++\n t/helper/test-tool.c                   |  1 +\n t/helper/test-tool.h                   |  1 +\n t/t5562-http-backend-content-length.sh |  2 +-\n t/test-lib-functions.sh                |  8 +-------\n 6 files changed, 27 insertions(+), 8 deletions(-)\n create mode 100644 t/helper/test-genzeros.c\n\ndiff --git a/Makefile b/Makefile\nindex f0b2299172..c5240942f2 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -740,6 +740,7 @@ TEST_BUILTINS_OBJS += test-dump-split-index.o\n TEST_BUILTINS_OBJS += test-dump-untracked-cache.o\n TEST_BUILTINS_OBJS += test-example-decorate.o\n TEST_BUILTINS_OBJS += test-genrandom.o\n+TEST_BUILTINS_OBJS += test-genzeros.o\n TEST_BUILTINS_OBJS += test-hash.o\n TEST_BUILTINS_OBJS += test-hashmap.o\n TEST_BUILTINS_OBJS += test-hash-speed.o\ndiff --git a/t/helper/test-genzeros.c b/t/helper/test-genzeros.c\nnew file mode 100644\nindex 0000000000..f607f800a9\n--- /dev/null\n+++ b/t/helper/test-genzeros.c\n@@ -0,0 +1,22 @@\n+#include \"test-tool.h\"\n+#include \"git-compat-util.h\"\n+\n+int cmd__genzeros(int argc, const char **argv)\n+{\n+\tlong count;\n+\n+\tif (argc > 2) {\n+\t\tfprintf(stderr, \"usage: %s [<count>]\\n\", argv[0]);\n+\t\treturn 1;\n+\t}\n+\n+\tcount = argc > 1 ? strtol(argv[1], NULL, 0) : -1L;\n+\n+\twhile (count < 0 || count--) {\n+\t\tif (putchar(0) == EOF)\n+\t\t\treturn -1;\n+\t}\n+\n+\treturn 0;\n+}\n+\ndiff --git a/t/helper/test-tool.c b/t/helper/test-tool.c\nindex 50c55f8b1a..99db7409b8 100644\n--- a/t/helper/test-tool.c\n+++ b/t/helper/test-tool.c\n@@ -19,6 +19,7 @@ static struct test_cmd cmds[] = {\n \t{ \"dump-untracked-cache\", cmd__dump_untracked_cache },\n \t{ \"example-decorate\", cmd__example_decorate },\n \t{ \"genrandom\", cmd__genrandom },\n+\t{ \"genzeros\", cmd__genzeros },\n \t{ \"hashmap\", cmd__hashmap },\n \t{ \"hash-speed\", cmd__hash_speed },\n \t{ \"index-version\", cmd__index_version },\ndiff --git a/t/helper/test-tool.h b/t/helper/test-tool.h\nindex a563df49bf..25abed1cf2 100644\n--- a/t/helper/test-tool.h\n+++ b/t/helper/test-tool.h\n@@ -16,6 +16,7 @@ int cmd__dump_split_index(int argc, const char **argv);\n int cmd__dump_untracked_cache(int argc, const char **argv);\n int cmd__example_decorate(int argc, const char **argv);\n int cmd__genrandom(int argc, const char **argv);\n+int cmd__genzeros(int argc, const char **argv);\n int cmd__hashmap(int argc, const char **argv);\n int cmd__hash_speed(int argc, const char **argv);\n int cmd__index_version(int argc, const char **argv);\ndiff --git a/t/t5562-http-backend-content-length.sh b/t/t5562-http-backend-content-length.sh\nindex bbadde2c6e..c0105423e6 100755\n--- a/t/t5562-http-backend-content-length.sh\n+++ b/t/t5562-http-backend-content-length.sh\n@@ -143,7 +143,7 @@ test_expect_success GZIP 'push gzipped empty' '\n \n test_expect_success 'CONTENT_LENGTH overflow ssite_t' '\n \tNOT_FIT_IN_SSIZE=$(ssize_b100dots) &&\n-\tgenerate_zero_bytes infinity  | env \\\n+\tgenerate_zero_bytes | env \\\n \t\tCONTENT_TYPE=application/x-git-upload-pack-request \\\n \t\tQUERY_STRING=/repo.git/git-upload-pack \\\n \t\tPATH_TRANSLATED=\"$PWD\"/.git/git-upload-pack \\\ndiff --git a/t/test-lib-functions.sh b/t/test-lib-functions.sh\nindex 094c07748a..80402a428f 100644\n--- a/t/test-lib-functions.sh\n+++ b/t/test-lib-functions.sh\n@@ -120,13 +120,7 @@ remove_cr () {\n # If $1 is 'infinity', output forever or until the receiving pipe stops reading,\n # whichever comes first.\n generate_zero_bytes () {\n-\tperl -e 'if ($ARGV[0] == \"infinity\") {\n-\t\twhile (-1) {\n-\t\t\tprint \"\\0\"\n-\t\t}\n-\t} else {\n-\t\tprint \"\\0\" x $ARGV[0]\n-\t}' \"$@\"\n+\ttest-tool genzeros \"$@\"\n }\n \n # In some bourne shell implementations, the \"unset\" builtin returns\n-- \ngitgitgadget\n"},{"id":"369353","messageId":"xmqqimxm6msi.fsf@gitster-ct.c.googlers.com","threadId":"50513","inReplyTo":"34cde0f2849a098c17ab83786da5ce06f69cfafa.1550179990.git.gitgitgadget@gmail.com","subject":"Re: [PATCH 1/1] tests: teach the test-tool to generate NUL bytes and use it","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-02-14T22:13:49Z","receivedAt":"2019-02-14T22:13:54Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\nwrites:\n\n> From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>\n> In cc95bc2025 (t5562: replace /dev/zero with a pipe from\n> generate_zero_bytes, 2019-02-09), we replaced usage of /dev/zero (which\n> is not available on NonStop, apparently) by a Perl script snippet to\n> generate NUL bytes.\n>\n> Sadly, it does not seem to work on NonStop, as t5562 reportedly hangs.\n> ...\n> In the end, though, what counts is that this here change incidentally\n> fixes that hang (maybe also on NonStop?). Even more positively, it gets\n> rid of yet another unnecessary Perl invocation.\n\nThanks for a quick band-aid.\n\nWill apply directly to 'master' so that we won't forget before -rc2.\n\nIn the meantime, perhaps somebody who knows Perl interpreter's\nquirks well can tell us what's different between the obvious and\nsimple C program and an equivalent in Perl to convince us why this\nis a good solution to the problem.\n\n"},{"id":"369354","messageId":"005401d4c4b3$147aa8c0$3d6ffa40$@nexbridge.com","threadId":"50513","inReplyTo":"pull.126.git.gitgitgadget@gmail.com","subject":"RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-14T22:17:26Z","receivedAt":"2019-02-14T22:17:42Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 14, 2019 16:33, Johannes Schindelin wrote:\n> To: git@vger.kernel.org\n> Cc: Randall Becker <rsbecker@nexbridge.com>; Junio C Hamano\n> <gitster@pobox.com>\n> Subject: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1\n> \n> The last-minute patch to replace /dev/zero with a Perl script snippet broke\n> the Linux part of the CI builds on Azure Pipelines: it timed out. The culprit is\n> the rb/no-dev-zero-in-test branch (see the build for this branch here\n> [https://dev.azure.com/gitgitgadget/git/_build/results?buildId=1727]).\n> \n> All of master, next, jch and pu are broken that way. You might see it in the\n> commit status of the active branches\n> [https://github.com/gitgitgadget/git/branches/active].\n> \n> Turns out that it is that particular Perl script snippet which for some reason\n> hangs the build. If you kill it, t5562.15 succeeds, if you don't kill it, it will hang\n> indefinitely (or until killed).\n> \n> Sadly, despite my earnest attempts, I could not figure out why it hangs in\n> those Linux agents (I could not reproduce that hang locally), or for that\n> matter, why it does not hang in the Windows and macOS agents.\n> \n> Let's avoid that hang. This patch fixes things on Azure Pipelines, and my hope\n> is that it also fixes the hang on NonStop.\n> \n> Johannes Schindelin (1):\n>   tests: teach the test-tool to generate NUL bytes and use it\n> \n>  Makefile                               |  1 +\n>  t/helper/test-genzeros.c               | 22 ++++++++++++++++++++++\n>  t/helper/test-tool.c                   |  1 +\n>  t/helper/test-tool.h                   |  1 +\n>  t/t5562-http-backend-content-length.sh |  2 +-\n>  t/test-lib-functions.sh                |  8 +-------\n>  6 files changed, 27 insertions(+), 8 deletions(-)  create mode 100644\n> t/helper/test-genzeros.c\n> \n> \n> base-commit: 8989e1950a845ceeb186d490321a4f917ca4de47\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-\n> 126%2Fdscho%2Ffix-t5562-hang-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-126/dscho/fix-\n> t5562-hang-v1\n> Pull-Request: https://github.com/gitgitgadget/git/pull/126\n\nUnfortunately, subtest 13 still hangs on NonStop, even with this patch, so our Pipeline still hangs. I'm glad it's better on Azure, but I don't think this actually addresses the root cause of the hang. This is now the fourth attempt at fixing this. Is it possible this is not the test that is failing, but actually the git-http-backend? The code is not in a loop, if that helps. It is not consuming any significant cycles. I don't know that part of the code at all, sadly. The code is here:\n\n* in the operating system from here up *\n  cleanup_children + 0x5D0 (UCr)\n  cleanup_children_on_exit + 0x70 (UCr)\n  git_atexit_dispatch + 0x200 (UCr)\n  __process_atexit_functions + 0xA0 (DLL zcredll)\n  CRE_TERMINATOR_ + 0xB50 (DLL zcredll)\n  exit + 0x2A0 (DLL zcrtldll)\n  die_webcgi + 0x240 (UCr)\n  die_errno + 0x360 (UCr)\n  write_or_die + 0x1C0 (UCr)\n  end_headers + 0x1A0 (UCr)\n  die_webcgi + 0x220 (UCr)\n  die + 0x320 (UCr)\n  inflate_request + 0x520 (UCr)\n  run_service + 0xC20 (UCr)\n  service_rpc + 0x530 (UCr)\n  cmd_main + 0xD00 (UCr)\n  main + 0x190 (UCr)\n\nBest guess is that a signal (SIGCHLD?) is possibly getting eaten or neglected somewhere between the test, perl, and git-http-backend.\n\nStuck,\nRandall\n\n"},{"id":"369356","messageId":"xmqqef8a6lnb.fsf@gitster-ct.c.googlers.com","threadId":"50513","inReplyTo":"005401d4c4b3$147aa8c0$3d6ffa40$@nexbridge.com","subject":"Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-02-14T22:38:32Z","receivedAt":"2019-02-14T22:38:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Randall S. Becker\" <rsbecker@nexbridge.com> writes:\n\n> Unfortunately, subtest 13 still hangs on NonStop, even with this\n> patch, so our Pipeline still hangs. I'm glad it's better on Azure,\n> but I don't think this actually addresses the root cause of the\n> hang.\n\nSigh.\n\n> possible this is not the test that is failing, but actually the\n> git-http-backend? The code is not in a loop, if that helps. It is\n> not consuming any significant cycles. I don't know that part of\n> the code at all, sadly. The code is here:\n>\n> * in the operating system from here up *\n>   cleanup_children + 0x5D0 (UCr)\n>   cleanup_children_on_exit + 0x70 (UCr)\n>   git_atexit_dispatch + 0x200 (UCr)\n>   __process_atexit_functions + 0xA0 (DLL zcredll)\n>   CRE_TERMINATOR_ + 0xB50 (DLL zcredll)\n>   exit + 0x2A0 (DLL zcrtldll)\n>   die_webcgi + 0x240 (UCr)\n>   die_errno + 0x360 (UCr)\n>   write_or_die + 0x1C0 (UCr)\n>   end_headers + 0x1A0 (UCr)\n>   die_webcgi + 0x220 (UCr)\n>   die + 0x320 (UCr)\n>   inflate_request + 0x520 (UCr)\n>   run_service + 0xC20 (UCr)\n>   service_rpc + 0x530 (UCr)\n>   cmd_main + 0xD00 (UCr)\n>   main + 0x190 (UCr)\n>\n> Best guess is that a signal (SIGCHLD?) is possibly getting eaten\n> or neglected somewhere between the test, perl, and\n> git-http-backend.\n\nSo we are trying to die(), which actually happens in die_webcgi(),\nand then try to write some message _but_ notice an error inside\nwrite_or_dir() and try to exit because we do not want to recurse\nforever trying to die, giving a message to say how/why we died, and\ndie because failing to give that message, forever.\n\nBut in our attempt to exit(), we try to \"cleanup children\" and that\nis what gets stuck.\n\nOne big difference before and after the /dev/zero change is that the\nprocess is now on a downstream of the pipe.  If we prepare a large\nfile with a finite size full of NULs and replace /dev/null with it,\ninstead of feeding NULs from the pipe, would it change the equation?\n"},{"id":"369357","messageId":"20190214223334.GE3064@jessie.local","threadId":"50513","inReplyTo":"005401d4c4b3$147aa8c0$3d6ffa40$@nexbridge.com","subject":"Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Max Kirillov","fromEmail":"max@max630.net","sentAt":"2019-02-14T22:33:34Z","receivedAt":"2019-02-14T22:40:58Z","isPatch":true,"sender":{"key":"max@max630.net","avatar":"https://avatars.githubusercontent.com/u/381560?v=4"},"body":"On Thu, Feb 14, 2019 at 05:17:26PM -0500, Randall S. Becker wrote:\n> Unfortunately, subtest 13 still hangs on NonStop, even\n> with this patch, so our Pipeline still hangs. I'm glad\n> it's better on Azure, but I don't think this actually\n> addresses the root cause of the hang. This is now the\n> fourth attempt at fixing this. Is it possible this is not\n> the test that is failing, but actually the\n> git-http-backend? The code is not in a loop, if that\n> helps. It is not consuming any significant cycles. I don't\n> know that part of the code at all, sadly. The code is\n> here:\n> \n> * in the operating system from here up *\n>   cleanup_children + 0x5D0 (UCr)\n\n... so does the process which the stack was taken from has\nany children processes still?\n\nI could imagine if a child somehow manages to end up in\nuninterruptible sleep, then probably it would never complete\nthis way, wouldn't it?\n"},{"id":"369359","messageId":"005801d4c4b8$e5669980$b033cc80$@nexbridge.com","threadId":"50513","inReplyTo":"20190214223334.GE3064@jessie.local","subject":"RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-14T22:59:04Z","receivedAt":"2019-02-14T22:59:17Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 14, 2019 17:34, Max Kirillov wrote:\n> On Thu, Feb 14, 2019 at 05:17:26PM -0500, Randall S. Becker wrote:\n> > Unfortunately, subtest 13 still hangs on NonStop, even with this\n> > patch, so our Pipeline still hangs. I'm glad it's better on Azure, but\n> > I don't think this actually addresses the root cause of the hang. This\n> > is now the fourth attempt at fixing this. Is it possible this is not\n> > the test that is failing, but actually the git-http-backend? The code\n> > is not in a loop, if that helps. It is not consuming any significant\n> > cycles. I don't know that part of the code at all, sadly. The code is\n> > here:\n> >\n> > * in the operating system from here up *\n> >   cleanup_children + 0x5D0 (UCr)\n> \n> ... so does the process which the stack was taken from has any children\n> processes still?\n> \n> I could imagine if a child somehow manages to end up in uninterruptible\n> sleep, then probably it would never complete this way, wouldn't it?\n\nFrom what I can tell (previously reported), none of the children are dead.\ngit-http-backend is waiting and the others are in a read state. I can try to\nget full stack traces once the current cycle ends.\n\n"},{"id":"369360","messageId":"005901d4c4b9$3b9a2a60$b2ce7f20$@nexbridge.com","threadId":"50513","inReplyTo":"xmqqef8a6lnb.fsf@gitster-ct.c.googlers.com","subject":"RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-14T23:01:29Z","receivedAt":"2019-02-14T23:01:44Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 14, 2019 17:39, Junio C Hamano wrote:\n> To: Randall S. Becker <rsbecker@nexbridge.com>\n> Cc: 'Johannes Schindelin via GitGitGadget' <gitgitgadget@gmail.com>;\n> git@vger.kernel.org; 'Max Kirillov' <max@max630.net>\n> Subject: Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1\n> \n> \"Randall S. Becker\" <rsbecker@nexbridge.com> writes:\n> \n> > Unfortunately, subtest 13 still hangs on NonStop, even with this\n> > patch, so our Pipeline still hangs. I'm glad it's better on Azure, but\n> > I don't think this actually addresses the root cause of the hang.\n> \n> Sigh.\n> \n> > possible this is not the test that is failing, but actually the\n> > git-http-backend? The code is not in a loop, if that helps. It is not\n> > consuming any significant cycles. I don't know that part of the code\n> > at all, sadly. The code is here:\n> >\n> > * in the operating system from here up *\n> >   cleanup_children + 0x5D0 (UCr)\n> >   cleanup_children_on_exit + 0x70 (UCr)\n> >   git_atexit_dispatch + 0x200 (UCr)\n> >   __process_atexit_functions + 0xA0 (DLL zcredll)\n> >   CRE_TERMINATOR_ + 0xB50 (DLL zcredll)\n> >   exit + 0x2A0 (DLL zcrtldll)\n> >   die_webcgi + 0x240 (UCr)\n> >   die_errno + 0x360 (UCr)\n> >   write_or_die + 0x1C0 (UCr)\n> >   end_headers + 0x1A0 (UCr)\n> >   die_webcgi + 0x220 (UCr)\n> >   die + 0x320 (UCr)\n> >   inflate_request + 0x520 (UCr)\n> >   run_service + 0xC20 (UCr)\n> >   service_rpc + 0x530 (UCr)\n> >   cmd_main + 0xD00 (UCr)\n> >   main + 0x190 (UCr)\n> >\n> > Best guess is that a signal (SIGCHLD?) is possibly getting eaten or\n> > neglected somewhere between the test, perl, and git-http-backend.\n> \n> So we are trying to die(), which actually happens in die_webcgi(), and\nthen try\n> to write some message _but_ notice an error inside\n> write_or_dir() and try to exit because we do not want to recurse forever\n> trying to die, giving a message to say how/why we died, and die because\n> failing to give that message, forever.\n> \n> But in our attempt to exit(), we try to \"cleanup children\" and that is\nwhat gets\n> stuck.\n> \n> One big difference before and after the /dev/zero change is that the\nprocess\n> is now on a downstream of the pipe.  If we prepare a large file with a\nfinite\n> size full of NULs and replace /dev/null with it, instead of feeding NULs\nfrom\n> the pipe, would it change the equation?\n\nDoubtful. The processes are still around, and are waiting on read but not\nactively reading (CPU time is not going up, so we're not reading an infinite\nstream). To me, this is a pipe situation where there is simply nothing\nwaiting on the pipe (maybe a flush missing?). I'm grasping are straws\nwithout knowing the actual process architecture of the test to debug it.\n\n"},{"id":"369361","messageId":"005a01d4c4b9$a9c411e0$fd4c35a0$@nexbridge.com","threadId":"50513","inReplyTo":"20190214223334.GE3064@jessie.local","subject":"RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-14T23:04:34Z","receivedAt":"2019-02-14T23:04:47Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 14, 2019 17:34, Max Kirillov wrote:\n> To: Randall S. Becker <rsbecker@nexbridge.com>\n> Cc: 'Johannes Schindelin via GitGitGadget' <gitgitgadget@gmail.com>;\n> git@vger.kernel.org; 'Junio C Hamano' <gitster@pobox.com>; 'Max Kirillov'\n> <max@max630.net>\n> Subject: Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1\n> \n> On Thu, Feb 14, 2019 at 05:17:26PM -0500, Randall S. Becker wrote:\n> > Unfortunately, subtest 13 still hangs on NonStop, even with this\n> > patch, so our Pipeline still hangs. I'm glad it's better on Azure, but\n> > I don't think this actually addresses the root cause of the hang. This\n> > is now the fourth attempt at fixing this. Is it possible this is not\n> > the test that is failing, but actually the git-http-backend? The code\n> > is not in a loop, if that helps. It is not consuming any significant\n> > cycles. I don't know that part of the code at all, sadly. The code is\n> > here:\n> >\n> > * in the operating system from here up *\n> >   cleanup_children + 0x5D0 (UCr)\n> \n> ... so does the process which the stack was taken from has any children\n> processes still?\n> \n> I could imagine if a child somehow manages to end up in uninterruptible\n> sleep, then probably it would never complete this way, wouldn't it?\n\nYes, this is typical of a hang. Two processes reading on the same pipe, or\none reading on a pipe and the other waiting for something that never shows.\nOr one process attempting both reading and writing on the same pipe (no\nkernel threads here). I did not see anything actually in sleep. perl is in a\nclose call, waiting for its output to be consumed - which never happens,\nmaking me suspect this is a pipe setup issue, but I can't demonstrate that,\nsorry.\n\n"},{"id":"369383","messageId":"nycvar.QRO.7.76.6.1902151558080.45@tvgsbejvaqbjf.bet","threadId":"50513","inReplyTo":"xmqqimxm6msi.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 1/1] tests: teach the test-tool to generate NUL bytes and use it","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-02-15T14:59:12Z","receivedAt":"2019-02-15T14:59:26Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Junio,\n\nOn Thu, 14 Feb 2019, Junio C Hamano wrote:\n\n> \"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\n> writes:\n> \n> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> >\n> > In cc95bc2025 (t5562: replace /dev/zero with a pipe from\n> > generate_zero_bytes, 2019-02-09), we replaced usage of /dev/zero (which\n> > is not available on NonStop, apparently) by a Perl script snippet to\n> > generate NUL bytes.\n> >\n> > Sadly, it does not seem to work on NonStop, as t5562 reportedly hangs.\n> > ...\n> > In the end, though, what counts is that this here change incidentally\n> > fixes that hang (maybe also on NonStop?). Even more positively, it gets\n> > rid of yet another unnecessary Perl invocation.\n> \n> Thanks for a quick band-aid.\n> \n> Will apply directly to 'master' so that we won't forget before -rc2.\n\nThank you, that will be good, as the builds still seem to fail. All of\nthem.\n\n> In the meantime, perhaps somebody who knows Perl interpreter's\n> quirks well can tell us what's different between the obvious and\n> simple C program and an equivalent in Perl to convince us why this\n> is a good solution to the problem.\n\nI would be *really* curious about that, too.\n\nCiao,\nDscho\n"},{"id":"369396","messageId":"xmqq5ztl6jbj.fsf@gitster-ct.c.googlers.com","threadId":"50513","inReplyTo":"nycvar.QRO.7.76.6.1902151558080.45@tvgsbejvaqbjf.bet","subject":"Re: [PATCH 1/1] tests: teach the test-tool to generate NUL bytes and use it","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-02-15T17:41:04Z","receivedAt":"2019-02-15T17:41:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> On Thu, 14 Feb 2019, Junio C Hamano wrote:\n>\n>> \"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\n>> writes:\n>> \n>> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n>> >\n>> > In cc95bc2025 (t5562: replace /dev/zero with a pipe from\n>> > generate_zero_bytes, 2019-02-09), we replaced usage of /dev/zero (which\n>> > is not available on NonStop, apparently) by a Perl script snippet to\n>> > generate NUL bytes.\n>> >\n>> > Sadly, it does not seem to work on NonStop, as t5562 reportedly hangs.\n>> > ...\n>> > In the end, though, what counts is that this here change incidentally\n>> > fixes that hang (maybe also on NonStop?). Even more positively, it gets\n>> > rid of yet another unnecessary Perl invocation.\n>> \n>> Thanks for a quick band-aid.\n>> \n>> Will apply directly to 'master' so that we won't forget before -rc2.\n>\n> Thank you, that will be good, as the builds still seem to fail. All of\n> them.\n\nActually, I am really tempted to instead not apply this, but revert\nthat genzerobytes Perl thing.  This assumes that your Azure thing\ndid not have the breakage before we applied that patchset.  What do\nyou think?\n\nTrying four or more possible band-aids that may or may not work\nwithout knowing what the real cause of the hangs are is not\nsomething I want to see people spend excessive time of theirs on\nthis close to the final.  I'd rather avoid distraction and see\npeople spend their cycles on bugs that matter, instead of trying to\nchase test breakages that have always been present for those without\n/dev/zero.  I am not fundamentally opposed to supporting those\nwithout /dev/zero but I'd prefer to see it happen in 'pu' until we\nidentify and fix the real cause---which may well be a real bug in\nthe http-backend stuff---and the time to do that is not during the\nrc period where we close the tree for new features and non-regression\nfixes.\n\n\n\n"},{"id":"369534","messageId":"nycvar.QRO.7.76.6.1902181652160.45@tvgsbejvaqbjf.bet","threadId":"50513","inReplyTo":"xmqq5ztl6jbj.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 1/1] tests: teach the test-tool to generate NUL bytes and use it","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-02-18T15:55:57Z","receivedAt":"2019-02-18T15:56:16Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Junio,\n\nOn Fri, 15 Feb 2019, Junio C Hamano wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n> \n> > On Thu, 14 Feb 2019, Junio C Hamano wrote:\n> >\n> >> \"Johannes Schindelin via GitGitGadget\" <gitgitgadget@gmail.com>\n> >> writes:\n> >> \n> >> > From: Johannes Schindelin <johannes.schindelin@gmx.de>\n> >> >\n> >> > In cc95bc2025 (t5562: replace /dev/zero with a pipe from\n> >> > generate_zero_bytes, 2019-02-09), we replaced usage of /dev/zero (which\n> >> > is not available on NonStop, apparently) by a Perl script snippet to\n> >> > generate NUL bytes.\n> >> >\n> >> > Sadly, it does not seem to work on NonStop, as t5562 reportedly hangs.\n> >> > ...\n> >> > In the end, though, what counts is that this here change incidentally\n> >> > fixes that hang (maybe also on NonStop?). Even more positively, it gets\n> >> > rid of yet another unnecessary Perl invocation.\n> >> \n> >> Thanks for a quick band-aid.\n> >> \n> >> Will apply directly to 'master' so that we won't forget before -rc2.\n> >\n> > Thank you, that will be good, as the builds still seem to fail. All of\n> > them.\n> \n> Actually, I am really tempted to instead not apply this, but revert\n> that genzerobytes Perl thing.  This assumes that your Azure thing\n> did not have the breakage before we applied that patchset.  What do\n> you think?\n\nHonestly, I don't care, as long as we can stop *all* of the CI builds\nfailing, soon.\n\nSo whether you revert that commit, or apply mine, it is up to you, but\nI'd really rather not see another batch of useless builds.\n\n> Trying four or more possible band-aids that may or may not work\n> without knowing what the real cause of the hangs are is not\n> something I want to see people spend excessive time of theirs on\n> this close to the final.  I'd rather avoid distraction and see\n> people spend their cycles on bugs that matter, instead of trying to\n> chase test breakages that have always been present for those without\n> /dev/zero.  I am not fundamentally opposed to supporting those\n> without /dev/zero but I'd prefer to see it happen in 'pu' until we\n> identify and fix the real cause---which may well be a real bug in\n> the http-backend stuff---and the time to do that is not during the\n> rc period where we close the tree for new features and non-regression\n> fixes.\n\nWhile I am quite positive that my patch helps (even on NonStop, because\nreportedly the hang is in .13 while my fix is about .15, which the commit\nthat caused the regression also touches), I am okay with holding off until\nafter v2.21.0.\n\nBut in the long run, the only sane thing really is to move more and more\nfunctionality in the test suite for which we rely on Perl into test-tool.\nIt is really the only thing that makes sense, from the perspectives of\nperformance, robustness and portability.\n\nCiao,\nDscho\n"},{"id":"369555","messageId":"nycvar.QRO.7.76.6.1902182139490.45@tvgsbejvaqbjf.bet","threadId":"50513","inReplyTo":"005901d4c4b9$3b9a2a60$b2ce7f20$@nexbridge.com","subject":"RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-02-18T20:41:09Z","receivedAt":"2019-02-18T20:41:27Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Randall,\n\nOn Thu, 14 Feb 2019, Randall S. Becker wrote:\n\n> On February 14, 2019 17:39, Junio C Hamano wrote:\n> > To: Randall S. Becker <rsbecker@nexbridge.com>\n> > Cc: 'Johannes Schindelin via GitGitGadget' <gitgitgadget@gmail.com>;\n> > git@vger.kernel.org; 'Max Kirillov' <max@max630.net>\n> > Subject: Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1\n> > \n> > \"Randall S. Becker\" <rsbecker@nexbridge.com> writes:\n> > \n> > > Unfortunately, subtest 13 still hangs on NonStop, even with this\n> > > patch, so our Pipeline still hangs. I'm glad it's better on Azure, but\n> > > I don't think this actually addresses the root cause of the hang.\n> > \n> > Sigh.\n> > \n> > > possible this is not the test that is failing, but actually the\n> > > git-http-backend? The code is not in a loop, if that helps. It is not\n> > > consuming any significant cycles. I don't know that part of the code\n> > > at all, sadly. The code is here:\n> > >\n> > > * in the operating system from here up *\n> > >   cleanup_children + 0x5D0 (UCr)\n> > >   cleanup_children_on_exit + 0x70 (UCr)\n> > >   git_atexit_dispatch + 0x200 (UCr)\n> > >   __process_atexit_functions + 0xA0 (DLL zcredll)\n> > >   CRE_TERMINATOR_ + 0xB50 (DLL zcredll)\n> > >   exit + 0x2A0 (DLL zcrtldll)\n> > >   die_webcgi + 0x240 (UCr)\n> > >   die_errno + 0x360 (UCr)\n> > >   write_or_die + 0x1C0 (UCr)\n> > >   end_headers + 0x1A0 (UCr)\n> > >   die_webcgi + 0x220 (UCr)\n> > >   die + 0x320 (UCr)\n> > >   inflate_request + 0x520 (UCr)\n> > >   run_service + 0xC20 (UCr)\n> > >   service_rpc + 0x530 (UCr)\n> > >   cmd_main + 0xD00 (UCr)\n> > >   main + 0x190 (UCr)\n> > >\n> > > Best guess is that a signal (SIGCHLD?) is possibly getting eaten or\n> > > neglected somewhere between the test, perl, and git-http-backend.\n> > \n> > So we are trying to die(), which actually happens in die_webcgi(), and\n> then try\n> > to write some message _but_ notice an error inside\n> > write_or_dir() and try to exit because we do not want to recurse forever\n> > trying to die, giving a message to say how/why we died, and die because\n> > failing to give that message, forever.\n> > \n> > But in our attempt to exit(), we try to \"cleanup children\" and that is\n> what gets\n> > stuck.\n> > \n> > One big difference before and after the /dev/zero change is that the\n> process\n> > is now on a downstream of the pipe.  If we prepare a large file with a\n> finite\n> > size full of NULs and replace /dev/null with it, instead of feeding NULs\n> from\n> > the pipe, would it change the equation?\n> \n> Doubtful. The processes are still around, and are waiting on read but not\n> actively reading (CPU time is not going up, so we're not reading an infinite\n> stream). To me, this is a pipe situation where there is simply nothing\n> waiting on the pipe (maybe a flush missing?). I'm grasping are straws\n> without knowing the actual process architecture of the test to debug it.\n\nSo could you try with this patch?\n\n-- snipsnap --\ndiff --git a/http-backend.c b/http-backend.c\nindex d5cea0329a..7c1b4a2555 100644\n--- a/http-backend.c\n+++ b/http-backend.c\n@@ -427,6 +427,7 @@ static void inflate_request(const char *prog_name, int out, int buffer_input, ss\n \n done:\n \tgit_inflate_end(&stream);\n+\tclose(0);\n \tclose(out);\n \tfree(full_request);\n }\n\n"},{"id":"369556","messageId":"005001d4c7cb$0cc4ce60$264e6b20$@nexbridge.com","threadId":"50513","inReplyTo":"nycvar.QRO.7.76.6.1902182139490.45@tvgsbejvaqbjf.bet","subject":"RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-18T20:46:34Z","receivedAt":"2019-02-18T20:46:49Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 18, 2019 15:41, Johannes Schindelin wrote:\n> On Thu, 14 Feb 2019, Randall S. Becker wrote:\n> \n> > On February 14, 2019 17:39, Junio C Hamano wrote:\n> > > To: Randall S. Becker <rsbecker@nexbridge.com>\n> > > Cc: 'Johannes Schindelin via GitGitGadget' <gitgitgadget@gmail.com>;\n> > > git@vger.kernel.org; 'Max Kirillov' <max@max630.net>\n> > > Subject: Re: [PATCH 0/1] Fix hang in t5562, introduced in\n> > > v2.21.0-rc1\n> > >\n> > > \"Randall S. Becker\" <rsbecker@nexbridge.com> writes:\n> > >\n> > > > Unfortunately, subtest 13 still hangs on NonStop, even with this\n> > > > patch, so our Pipeline still hangs. I'm glad it's better on Azure,\n> > > > but I don't think this actually addresses the root cause of the\nhang.\n> > >\n> > > Sigh.\n> > >\n> > > > possible this is not the test that is failing, but actually the\n> > > > git-http-backend? The code is not in a loop, if that helps. It is\n> > > > not consuming any significant cycles. I don't know that part of\n> > > > the code at all, sadly. The code is here:\n> > > >\n> > > > * in the operating system from here up *\n> > > >   cleanup_children + 0x5D0 (UCr)\n> > > >   cleanup_children_on_exit + 0x70 (UCr)\n> > > >   git_atexit_dispatch + 0x200 (UCr)\n> > > >   __process_atexit_functions + 0xA0 (DLL zcredll)\n> > > >   CRE_TERMINATOR_ + 0xB50 (DLL zcredll)\n> > > >   exit + 0x2A0 (DLL zcrtldll)\n> > > >   die_webcgi + 0x240 (UCr)\n> > > >   die_errno + 0x360 (UCr)\n> > > >   write_or_die + 0x1C0 (UCr)\n> > > >   end_headers + 0x1A0 (UCr)\n> > > >   die_webcgi + 0x220 (UCr)\n> > > >   die + 0x320 (UCr)\n> > > >   inflate_request + 0x520 (UCr)\n> > > >   run_service + 0xC20 (UCr)\n> > > >   service_rpc + 0x530 (UCr)\n> > > >   cmd_main + 0xD00 (UCr)\n> > > >   main + 0x190 (UCr)\n> > > >\n> > > > Best guess is that a signal (SIGCHLD?) is possibly getting eaten\n> > > > or neglected somewhere between the test, perl, and git-http-backend.\n> > >\n> > > So we are trying to die(), which actually happens in die_webcgi(),\n> > > and\n> > then try\n> > > to write some message _but_ notice an error inside\n> > > write_or_dir() and try to exit because we do not want to recurse\n> > > forever trying to die, giving a message to say how/why we died, and\n> > > die because failing to give that message, forever.\n> > >\n> > > But in our attempt to exit(), we try to \"cleanup children\" and that\n> > > is\n> > what gets\n> > > stuck.\n> > >\n> > > One big difference before and after the /dev/zero change is that the\n> > process\n> > > is now on a downstream of the pipe.  If we prepare a large file with\n> > > a\n> > finite\n> > > size full of NULs and replace /dev/null with it, instead of feeding\n> > > NULs\n> > from\n> > > the pipe, would it change the equation?\n> >\n> > Doubtful. The processes are still around, and are waiting on read but\n> > not actively reading (CPU time is not going up, so we're not reading\n> > an infinite stream). To me, this is a pipe situation where there is\n> > simply nothing waiting on the pipe (maybe a flush missing?). I'm\n> > grasping are straws without knowing the actual process architecture of\nthe\n> test to debug it.\n> \n> So could you try with this patch?\n> \n> -- snipsnap --\n> diff --git a/http-backend.c b/http-backend.c index d5cea0329a..7c1b4a2555\n> 100644\n> --- a/http-backend.c\n> +++ b/http-backend.c\n> @@ -427,6 +427,7 @@ static void inflate_request(const char *prog_name,\n> int out, int buffer_input, ss\n> \n>  done:\n>  \tgit_inflate_end(&stream);\n> +\tclose(0);\n>  \tclose(out);\n>  \tfree(full_request);\n>  }\n\nIn isolation or with the other fixes associated with t5562? Or, which\nbaseline commit should I use? 8989e1950a or d92031209a or some other?\n\n"},{"id":"369559","messageId":"20190218205725.GB3373@jessie.local","threadId":"50513","inReplyTo":"005001d4c7cb$0cc4ce60$264e6b20$@nexbridge.com","subject":"Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Max Kirillov","fromEmail":"max@max630.net","sentAt":"2019-02-18T20:57:25Z","receivedAt":"2019-02-18T20:57:30Z","isPatch":true,"sender":{"key":"max@max630.net","avatar":"https://avatars.githubusercontent.com/u/381560?v=4"},"body":"On Mon, Feb 18, 2019 at 03:46:34PM -0500, Randall S. Becker wrote:\n> On February 18, 2019 15:41, Johannes Schindelin wrote:\n> > So could you try with this patch?\n> > \n> > -- snipsnap --\n> > diff --git a/http-backend.c b/http-backend.c index d5cea0329a..7c1b4a2555\n> > 100644\n> > --- a/http-backend.c\n> > +++ b/http-backend.c\n> > @@ -427,6 +427,7 @@ static void inflate_request(const char *prog_name,\n> > int out, int buffer_input, ss\n> > \n> >  done:\n> >  \tgit_inflate_end(&stream);\n> > +\tclose(0);\n> >  \tclose(out);\n> >  \tfree(full_request);\n> >  }\n> \n> In isolation or with the other fixes associated with t5562? Or, which\n> baseline commit should I use? 8989e1950a or d92031209a or some other?\n\nAs far as I understand, it should be tried instead of \nhttps://public-inbox.org/git/20181124093719.10705-1-max@max630.net/\n"},{"id":"369560","messageId":"005201d4c7cc$90bf29d0$b23d7d70$@nexbridge.com","threadId":"50513","inReplyTo":"005001d4c7cb$0cc4ce60$264e6b20$@nexbridge.com","subject":"RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-18T20:57:26Z","receivedAt":"2019-02-18T20:57:41Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 18, 2019 15:47, I wrote:\n> On February 18, 2019 15:41, Johannes Schindelin wrote:\n> > On Thu, 14 Feb 2019, Randall S. Becker wrote:\n> >\n> > > On February 14, 2019 17:39, Junio C Hamano wrote:\n> > > > To: Randall S. Becker <rsbecker@nexbridge.com>\n> > > > Cc: 'Johannes Schindelin via GitGitGadget'\n> > > > <gitgitgadget@gmail.com>; git@vger.kernel.org; 'Max Kirillov'\n> > > > <max@max630.net>\n> > > > Subject: Re: [PATCH 0/1] Fix hang in t5562, introduced in\n> > > > v2.21.0-rc1\n> > > >\n> > > > \"Randall S. Becker\" <rsbecker@nexbridge.com> writes:\n> > > >\n> > > > > Unfortunately, subtest 13 still hangs on NonStop, even with this\n> > > > > patch, so our Pipeline still hangs. I'm glad it's better on\n> > > > > Azure, but I don't think this actually addresses the root cause\n> > > > > of the\n> hang.\n> > > >\n> > > > Sigh.\n> > > >\n> > > > > possible this is not the test that is failing, but actually the\n> > > > > git-http-backend? The code is not in a loop, if that helps. It\n> > > > > is not consuming any significant cycles. I don't know that part\n> > > > > of the code at all, sadly. The code is here:\n> > > > >\n> > > > > * in the operating system from here up *\n> > > > >   cleanup_children + 0x5D0 (UCr)\n> > > > >   cleanup_children_on_exit + 0x70 (UCr)\n> > > > >   git_atexit_dispatch + 0x200 (UCr)\n> > > > >   __process_atexit_functions + 0xA0 (DLL zcredll)\n> > > > >   CRE_TERMINATOR_ + 0xB50 (DLL zcredll)\n> > > > >   exit + 0x2A0 (DLL zcrtldll)\n> > > > >   die_webcgi + 0x240 (UCr)\n> > > > >   die_errno + 0x360 (UCr)\n> > > > >   write_or_die + 0x1C0 (UCr)\n> > > > >   end_headers + 0x1A0 (UCr)\n> > > > >   die_webcgi + 0x220 (UCr)\n> > > > >   die + 0x320 (UCr)\n> > > > >   inflate_request + 0x520 (UCr)\n> > > > >   run_service + 0xC20 (UCr)\n> > > > >   service_rpc + 0x530 (UCr)\n> > > > >   cmd_main + 0xD00 (UCr)\n> > > > >   main + 0x190 (UCr)\n> > > > >\n> > > > > Best guess is that a signal (SIGCHLD?) is possibly getting eaten\n> > > > > or neglected somewhere between the test, perl, and git-http-\n> backend.\n> > > >\n> > > > So we are trying to die(), which actually happens in die_webcgi(),\n> > > > and\n> > > then try\n> > > > to write some message _but_ notice an error inside\n> > > > write_or_dir() and try to exit because we do not want to recurse\n> > > > forever trying to die, giving a message to say how/why we died,\n> > > > and die because failing to give that message, forever.\n> > > >\n> > > > But in our attempt to exit(), we try to \"cleanup children\" and\n> > > > that is\n> > > what gets\n> > > > stuck.\n> > > >\n> > > > One big difference before and after the /dev/zero change is that\n> > > > the\n> > > process\n> > > > is now on a downstream of the pipe.  If we prepare a large file\n> > > > with a\n> > > finite\n> > > > size full of NULs and replace /dev/null with it, instead of\n> > > > feeding NULs\n> > > from\n> > > > the pipe, would it change the equation?\n> > >\n> > > Doubtful. The processes are still around, and are waiting on read\n> > > but not actively reading (CPU time is not going up, so we're not\n> > > reading an infinite stream). To me, this is a pipe situation where\n> > > there is simply nothing waiting on the pipe (maybe a flush\n> > > missing?). I'm grasping are straws without knowing the actual\n> > > process architecture of\n> the\n> > test to debug it.\n> >\n> > So could you try with this patch?\n> >\n> > -- snipsnap --\n> > diff --git a/http-backend.c b/http-backend.c index\n> > d5cea0329a..7c1b4a2555\n> > 100644\n> > --- a/http-backend.c\n> > +++ b/http-backend.c\n> > @@ -427,6 +427,7 @@ static void inflate_request(const char *prog_name,\n> > int out, int buffer_input, ss\n> >\n> >  done:\n> >  \tgit_inflate_end(&stream);\n> > +\tclose(0);\n> >  \tclose(out);\n> >  \tfree(full_request);\n> >  }\n> \n> In isolation or with the other fixes associated with t5562? Or, which\nbaseline\n> commit should I use? 8989e1950a or d92031209a or some other?\n\nIt works fine based off of d92031209a - the test actually runs a bit faster\nbased on the wall clock.\n\n"},{"id":"369565","messageId":"87pnroolgy.fsf@evledraar.gmail.com","threadId":"50513","inReplyTo":"xmqqef8a6lnb.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-02-18T21:06:21Z","receivedAt":"2019-02-18T21:06:27Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Feb 14 2019, Junio C Hamano wrote:\n\n> \"Randall S. Becker\" <rsbecker@nexbridge.com> writes:\n>\n>> Unfortunately, subtest 13 still hangs on NonStop, even with this\n>> patch, so our Pipeline still hangs. I'm glad it's better on Azure,\n>> but I don't think this actually addresses the root cause of the\n>> hang.\n>\n> Sigh.\n>\n>> possible this is not the test that is failing, but actually the\n>> git-http-backend? The code is not in a loop, if that helps. It is\n>> not consuming any significant cycles. I don't know that part of\n>> the code at all, sadly. The code is here:\n>>\n>> * in the operating system from here up *\n>>   cleanup_children + 0x5D0 (UCr)\n>>   cleanup_children_on_exit + 0x70 (UCr)\n>>   git_atexit_dispatch + 0x200 (UCr)\n>>   __process_atexit_functions + 0xA0 (DLL zcredll)\n>>   CRE_TERMINATOR_ + 0xB50 (DLL zcredll)\n>>   exit + 0x2A0 (DLL zcrtldll)\n>>   die_webcgi + 0x240 (UCr)\n>>   die_errno + 0x360 (UCr)\n>>   write_or_die + 0x1C0 (UCr)\n>>   end_headers + 0x1A0 (UCr)\n>>   die_webcgi + 0x220 (UCr)\n>>   die + 0x320 (UCr)\n>>   inflate_request + 0x520 (UCr)\n>>   run_service + 0xC20 (UCr)\n>>   service_rpc + 0x530 (UCr)\n>>   cmd_main + 0xD00 (UCr)\n>>   main + 0x190 (UCr)\n>>\n>> Best guess is that a signal (SIGCHLD?) is possibly getting eaten\n>> or neglected somewhere between the test, perl, and\n>> git-http-backend.\n>\n> So we are trying to die(), which actually happens in die_webcgi(),\n> and then try to write some message _but_ notice an error inside\n> write_or_dir() and try to exit because we do not want to recurse\n> forever trying to die, giving a message to say how/why we died, and\n> die because failing to give that message, forever.\n>\n> But in our attempt to exit(), we try to \"cleanup children\" and that\n> is what gets stuck.\n\nI have not paid enough attention to this thread to say if this is dumb,\nbut just in case it's useful. For this class of problem where cleanup\nbites you for whatever reason in Perl, you can sometimes use this:\n\n    use POSIX ();\n    POSIX::_exit($code);\n\nThis will call \"exit\" from \"stdlib\" instead of Perl's \"exit\". So go away\n*now* and let the OS deal with the mess. Perl's will run around cleaning\nup stuff, freeing memory, running destructors etc, all of which might\nhave side effects you don't want/care about, and might (as maybe in this\ncase?) cause some hang.\n\n> One big difference before and after the /dev/zero change is that the\n> process is now on a downstream of the pipe.  If we prepare a large\n> file with a finite size full of NULs and replace /dev/null with it,\n> instead of feeding NULs from the pipe, would it change the equation?\n"},{"id":"369567","messageId":"20190218211725.GD3373@jessie.local","threadId":"50513","inReplyTo":"87pnroolgy.fsf@evledraar.gmail.com","subject":"Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Max Kirillov","fromEmail":"max@max630.net","sentAt":"2019-02-18T21:17:25Z","receivedAt":"2019-02-18T21:17:31Z","isPatch":true,"sender":{"key":"max@max630.net","avatar":"https://avatars.githubusercontent.com/u/381560?v=4"},"body":"On Mon, Feb 18, 2019 at 10:06:21PM +0100, Ævar Arnfjörð Bjarmason wrote:\n>> But in our attempt to exit(), we try to \"cleanup children\" and that\n>> is what gets stuck.\n> \n> I have not paid enough attention to this thread to say if this is dumb,\n> but just in case it's useful. For this class of problem where cleanup\n> bites you for whatever reason in Perl, you can sometimes use this:\n> \n>     use POSIX ();\n>     POSIX::_exit($code);\n> \n> This will call \"exit\" from \"stdlib\" instead of Perl's \"exit\". So go away\n> *now* and let the OS deal with the mess. Perl's will run around cleaning\n> up stuff, freeing memory, running destructors etc, all of which might\n> have side effects you don't want/care about, and might (as maybe in this\n> case?) cause some hang.\n\n* Perl is running in foreground, so it cannot outlive test\n  case and spoil the subsequent ones.\n* From the dumps I have an impression that it waits\n  legitimately - there are other processes to wait for.\n  And anyway the waits happen before perl script comes to\n  its exit.\n\nThough I am already convinced that I should have done the\nhelper in C. Let's see when I have time to fix it.\n"},{"id":"369570","messageId":"005c01d4c7d3$e5a3c850$b0eb58f0$@nexbridge.com","threadId":"50513","inReplyTo":"nycvar.QRO.7.76.6.1902182139490.45@tvgsbejvaqbjf.bet","subject":"RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-02-18T21:49:54Z","receivedAt":"2019-02-18T21:50:10Z","isPatch":true,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On February 18, 2019 15:41, Johannes Schindelin wrote:\n> To: Randall S. Becker <rsbecker@nexbridge.com>\n> Cc: 'Junio C Hamano' <gitster@pobox.com>; 'Johannes Schindelin via\n> GitGitGadget' <gitgitgadget@gmail.com>; git@vger.kernel.org; 'Max\nKirillov'\n> <max@max630.net>\n> Subject: RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1\n> \n> Hi Randall,\n> \n> On Thu, 14 Feb 2019, Randall S. Becker wrote:\n> \n> > On February 14, 2019 17:39, Junio C Hamano wrote:\n> > > To: Randall S. Becker <rsbecker@nexbridge.com>\n> > > Cc: 'Johannes Schindelin via GitGitGadget' <gitgitgadget@gmail.com>;\n> > > git@vger.kernel.org; 'Max Kirillov' <max@max630.net>\n> > > Subject: Re: [PATCH 0/1] Fix hang in t5562, introduced in\n> > > v2.21.0-rc1\n> > >\n> > > \"Randall S. Becker\" <rsbecker@nexbridge.com> writes:\n> > >\n> > > > Unfortunately, subtest 13 still hangs on NonStop, even with this\n> > > > patch, so our Pipeline still hangs. I'm glad it's better on Azure,\n> > > > but I don't think this actually addresses the root cause of the\nhang.\n> > >\n> > > Sigh.\n> > >\n> > > > possible this is not the test that is failing, but actually the\n> > > > git-http-backend? The code is not in a loop, if that helps. It is\n> > > > not consuming any significant cycles. I don't know that part of\n> > > > the code at all, sadly. The code is here:\n> > > >\n> > > > * in the operating system from here up *\n> > > >   cleanup_children + 0x5D0 (UCr)\n> > > >   cleanup_children_on_exit + 0x70 (UCr)\n> > > >   git_atexit_dispatch + 0x200 (UCr)\n> > > >   __process_atexit_functions + 0xA0 (DLL zcredll)\n> > > >   CRE_TERMINATOR_ + 0xB50 (DLL zcredll)\n> > > >   exit + 0x2A0 (DLL zcrtldll)\n> > > >   die_webcgi + 0x240 (UCr)\n> > > >   die_errno + 0x360 (UCr)\n> > > >   write_or_die + 0x1C0 (UCr)\n> > > >   end_headers + 0x1A0 (UCr)\n> > > >   die_webcgi + 0x220 (UCr)\n> > > >   die + 0x320 (UCr)\n> > > >   inflate_request + 0x520 (UCr)\n> > > >   run_service + 0xC20 (UCr)\n> > > >   service_rpc + 0x530 (UCr)\n> > > >   cmd_main + 0xD00 (UCr)\n> > > >   main + 0x190 (UCr)\n> > > >\n> > > > Best guess is that a signal (SIGCHLD?) is possibly getting eaten\n> > > > or neglected somewhere between the test, perl, and git-http-backend.\n> > >\n> > > So we are trying to die(), which actually happens in die_webcgi(),\n> > > and\n> > then try\n> > > to write some message _but_ notice an error inside\n> > > write_or_dir() and try to exit because we do not want to recurse\n> > > forever trying to die, giving a message to say how/why we died, and\n> > > die because failing to give that message, forever.\n> > >\n> > > But in our attempt to exit(), we try to \"cleanup children\" and that\n> > > is\n> > what gets\n> > > stuck.\n> > >\n> > > One big difference before and after the /dev/zero change is that the\n> > process\n> > > is now on a downstream of the pipe.  If we prepare a large file with\n> > > a\n> > finite\n> > > size full of NULs and replace /dev/null with it, instead of feeding\n> > > NULs\n> > from\n> > > the pipe, would it change the equation?\n> >\n> > Doubtful. The processes are still around, and are waiting on read but\n> > not actively reading (CPU time is not going up, so we're not reading\n> > an infinite stream). To me, this is a pipe situation where there is\n> > simply nothing waiting on the pipe (maybe a flush missing?). I'm\n> > grasping are straws without knowing the actual process architecture of\nthe\n> test to debug it.\n> \n> So could you try with this patch?\n> \n> -- snipsnap --\n> diff --git a/http-backend.c b/http-backend.c index d5cea0329a..7c1b4a2555\n> 100644\n> --- a/http-backend.c\n> +++ b/http-backend.c\n> @@ -427,6 +427,7 @@ static void inflate_request(const char *prog_name,\n> int out, int buffer_input, ss\n> \n>  done:\n>  \tgit_inflate_end(&stream);\n> +\tclose(0);\n>  \tclose(out);\n>  \tfree(full_request);\n>  }\n\nBased on d62dad7a7d (v2.21.0-rc0) undoing all of the fixes, this change on\nits own makes no difference to the hang situation - it is still there as it\nwas when originally reported. Using POSIX::_exit does not change the outcome\nof the test either on its own or in conjunction with this fix.\n\n"},{"id":"369638","messageId":"nycvar.QRO.7.76.6.1902191507380.41@tvgsbejvaqbjf.bet","threadId":"50513","inReplyTo":"20190218205725.GB3373@jessie.local","subject":"Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-02-19T14:09:15Z","receivedAt":"2019-02-19T14:09:52Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Max & Randall,\n\nOn Mon, 18 Feb 2019, Max Kirillov wrote:\n\n> On Mon, Feb 18, 2019 at 03:46:34PM -0500, Randall S. Becker wrote:\n> > On February 18, 2019 15:41, Johannes Schindelin wrote:\n> > > So could you try with this patch?\n> > > \n> > > -- snipsnap --\n> > > diff --git a/http-backend.c b/http-backend.c index d5cea0329a..7c1b4a2555\n> > > 100644\n> > > --- a/http-backend.c\n> > > +++ b/http-backend.c\n> > > @@ -427,6 +427,7 @@ static void inflate_request(const char *prog_name,\n> > > int out, int buffer_input, ss\n> > > \n> > >  done:\n> > >  \tgit_inflate_end(&stream);\n> > > +\tclose(0);\n> > >  \tclose(out);\n> > >  \tfree(full_request);\n> > >  }\n> > \n> > In isolation or with the other fixes associated with t5562? Or, which\n> > baseline commit should I use? 8989e1950a or d92031209a or some other?\n> \n> As far as I understand, it should be tried instead of \n> https://public-inbox.org/git/20181124093719.10705-1-max@max630.net/\n\nDon't ask me which patches you need to try this with. I was just answering\nto the observation that the hangs happen in the gzip-encoding test cases,\nand this was my guess as to what is going wrong there. I have no idea\nwhether other patches try to address the same thing, are obsoleted by this\ndiff, or whatever, as I have not been able to pay attention to the Git\nmailing list in the past 5 days.\n\nCiao,\nJohannes\n"},{"id":"369639","messageId":"nycvar.QRO.7.76.6.1902191512350.41@tvgsbejvaqbjf.bet","threadId":"50513","inReplyTo":"20190218211725.GD3373@jessie.local","subject":"Re: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2019-02-19T14:13:08Z","receivedAt":"2019-02-19T14:13:42Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Max,\n\nOn Mon, 18 Feb 2019, Max Kirillov wrote:\n\n> On Mon, Feb 18, 2019 at 10:06:21PM +0100, Ævar Arnfjörð Bjarmason wrote:\n> >> But in our attempt to exit(), we try to \"cleanup children\" and that\n> >> is what gets stuck.\n> > \n> > I have not paid enough attention to this thread to say if this is dumb,\n> > but just in case it's useful. For this class of problem where cleanup\n> > bites you for whatever reason in Perl, you can sometimes use this:\n> > \n> >     use POSIX ();\n> >     POSIX::_exit($code);\n> > \n> > This will call \"exit\" from \"stdlib\" instead of Perl's \"exit\". So go away\n> > *now* and let the OS deal with the mess. Perl's will run around cleaning\n> > up stuff, freeing memory, running destructors etc, all of which might\n> > have side effects you don't want/care about, and might (as maybe in this\n> > case?) cause some hang.\n> \n> * Perl is running in foreground, so it cannot outlive test\n>   case and spoil the subsequent ones.\n> * From the dumps I have an impression that it waits\n>   legitimately - there are other processes to wait for.\n>   And anyway the waits happen before perl script comes to\n>   its exit.\n> \n> Though I am already convinced that I should have done the\n> helper in C. Let's see when I have time to fix it.\n\nPerl has this nasty habit of causing unexpected problems, doesn't it?\n\nI look forward to that dependency on Perl going away, thank you so much!\n\nCiao,\nDscho"}]}