git/list[1] front-page[2] threads[3] people[4] search[5] about
 

RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1

From
Randall S. Becker <rsbecker@nexbridge.com>
Date
Feb 18, 2019, 21:49 UTC
Message-ID
<005c01d4c7d3$e5a3c850$b0eb58f0$@nexbridge.com>
In-Reply-To
<nycvar.QRO.7.76.6.1902182139490.45@tvgsbejvaqbjf.bet>
On February 18, 2019 15:41, Johannes Schindelin wrote:
> To: Randall S. Becker <rsbecker@nexbridge.com>
> Cc: 'Junio C Hamano' <gitster@pobox.com>; 'Johannes Schindelin via
> GitGitGadget' <gitgitgadget@gmail.com>; git@vger.kernel.org; 'Max
Kirillov'
Show 19 quoted lines
> <max@max630.net>
> Subject: RE: [PATCH 0/1] Fix hang in t5562, introduced in v2.21.0-rc1
> 
> Hi Randall,
> 
> On Thu, 14 Feb 2019, Randall S. Becker wrote:
> 
> > On February 14, 2019 17:39, Junio C Hamano wrote:
> > > To: Randall S. Becker <rsbecker@nexbridge.com>
> > > Cc: 'Johannes Schindelin via GitGitGadget' <gitgitgadget@gmail.com>;
> > > git@vger.kernel.org; 'Max Kirillov' <max@max630.net>
> > > Subject: Re: [PATCH 0/1] Fix hang in t5562, introduced in
> > > v2.21.0-rc1
> > >
> > > "Randall S. Becker" <rsbecker@nexbridge.com> writes:
> > >
> > > > Unfortunately, subtest 13 still hangs on NonStop, even with this
> > > > patch, so our Pipeline still hangs. I'm glad it's better on Azure,
> > > > but I don't think this actually addresses the root cause of the
hang.
Show 58 quoted lines
> > >
> > > Sigh.
> > >
> > > > possible this is not the test that is failing, but actually the
> > > > git-http-backend? The code is not in a loop, if that helps. It is
> > > > not consuming any significant cycles. I don't know that part of
> > > > the code at all, sadly. The code is here:
> > > >
> > > > * in the operating system from here up *
> > > >   cleanup_children + 0x5D0 (UCr)
> > > >   cleanup_children_on_exit + 0x70 (UCr)
> > > >   git_atexit_dispatch + 0x200 (UCr)
> > > >   __process_atexit_functions + 0xA0 (DLL zcredll)
> > > >   CRE_TERMINATOR_ + 0xB50 (DLL zcredll)
> > > >   exit + 0x2A0 (DLL zcrtldll)
> > > >   die_webcgi + 0x240 (UCr)
> > > >   die_errno + 0x360 (UCr)
> > > >   write_or_die + 0x1C0 (UCr)
> > > >   end_headers + 0x1A0 (UCr)
> > > >   die_webcgi + 0x220 (UCr)
> > > >   die + 0x320 (UCr)
> > > >   inflate_request + 0x520 (UCr)
> > > >   run_service + 0xC20 (UCr)
> > > >   service_rpc + 0x530 (UCr)
> > > >   cmd_main + 0xD00 (UCr)
> > > >   main + 0x190 (UCr)
> > > >
> > > > Best guess is that a signal (SIGCHLD?) is possibly getting eaten
> > > > or neglected somewhere between the test, perl, and git-http-backend.
> > >
> > > So we are trying to die(), which actually happens in die_webcgi(),
> > > and
> > then try
> > > to write some message _but_ notice an error inside
> > > write_or_dir() and try to exit because we do not want to recurse
> > > forever trying to die, giving a message to say how/why we died, and
> > > die because failing to give that message, forever.
> > >
> > > But in our attempt to exit(), we try to "cleanup children" and that
> > > is
> > what gets
> > > stuck.
> > >
> > > One big difference before and after the /dev/zero change is that the
> > process
> > > is now on a downstream of the pipe.  If we prepare a large file with
> > > a
> > finite
> > > size full of NULs and replace /dev/null with it, instead of feeding
> > > NULs
> > from
> > > the pipe, would it change the equation?
> >
> > Doubtful. The processes are still around, and are waiting on read but
> > not actively reading (CPU time is not going up, so we're not reading
> > an infinite stream). To me, this is a pipe situation where there is
> > simply nothing waiting on the pipe (maybe a flush missing?). I'm
> > grasping are straws without knowing the actual process architecture of
the
Show 18 quoted lines
> test to debug it.
> 
> So could you try with this patch?
> 
> -- snipsnap --
> diff --git a/http-backend.c b/http-backend.c index d5cea0329a..7c1b4a2555
> 100644
> --- a/http-backend.c
> +++ b/http-backend.c
> @@ -427,6 +427,7 @@ static void inflate_request(const char *prog_name,
> int out, int buffer_input, ss
> 
>  done:
>  	git_inflate_end(&stream);
> +	close(0);
>  	close(out);
>  	free(full_request);
>  }

Based on d62dad7a7d (v2.21.0-rc0) undoing all of the fixes, this change on its own makes no difference to the hang situation - it is still there as it was when originally reported. Using POSIX::_exit does not change the outcome of the test either on its own or in conjunction with this fix.

Previous: Randall S. BeckerNext: Ævar Arnfjörð Bjarmason
Message 15 of 21 in “Fix hang in t5562, introduced in v2.21.0-rc1”
  1. 0/1 Fix hang in t5562, introduced in v2.21.0-rc1Johannes Schindelin via GitGitGadget, Feb 14, 2019
  2. 1/1 tests: teach the test-tool to generate NUL bytes and use itJohannes Schindelin via GitGitGadget, Feb 14, 2019
  3. Junio C HamanoFeb 14, 2019
  4. Johannes SchindelinFeb 15, 2019
  5. Junio C HamanoFeb 15, 2019
  6. Johannes SchindelinFeb 18, 2019
  7. Randall S. BeckerFeb 14, 2019
  8. Junio C HamanoFeb 14, 2019
  9. Randall S. BeckerFeb 14, 2019
  10. Johannes SchindelinFeb 18, 2019
  11. Randall S. BeckerFeb 18, 2019
  12. Max KirillovFeb 18, 2019
  13. Johannes SchindelinFeb 19, 2019
  14. Randall S. BeckerFeb 18, 2019
  15. Randall S. BeckerFeb 18, 2019
  16. Ævar Arnfjörð BjarmasonFeb 18, 2019
  17. Max KirillovFeb 18, 2019
  18. Johannes SchindelinFeb 19, 2019
  19. Max KirillovFeb 14, 2019
  20. Randall S. BeckerFeb 14, 2019
  21. Randall S. BeckerFeb 14, 2019

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.