git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH] run_processes_parallel(): fix order of sigpipe handling

From
Jeff King <peff@peff.net>
Date
Apr 8, 2026, 17:20 UTC
Message-ID
<20260408172055.GA2293804@coredump.intra.peff.net>
In-Reply-To
<87y0ix1kq4.fsf@collabora.com>
On Wed, Apr 08, 2026 at 07:26:59PM +0300, Adrian Ratiu wrote:
> @Peff
> 
> Please let me know if you wish me to send a patch or if you wish to send
> it yourself, since this investigation is your work & effort. :)
Here it is with a small cleanup and a real commit message.
-- >8 --
Subject: run_processes_parallel(): fix order of sigpipe handling

In commit ec0becacc9 (run-command: add stdin callback for parallelization, 2026-01-28), we taught run_processes_parallel() to ignore SIGPIPE, since we wouldn't want a write() to a broken pipe of one of the children to take down the whole process.

But there's a subtle ordering issue. After we ignore SIGPIPE, we call pp_init(), which installs its own cleanup handler for multiple signals using sigchain_push_common(), which includes SIGPIPE. So if we receive SIGPIPE while writing to a child, we'll trigger that handler first, pop it off the stack, and then re-raise (which is then ignored because of the SIG_IGN we pushed first).

But what does that handler do? It tries to clean up all of the child processes, under the assumption that when we re-raise the signal we'll be exiting the process!

So a hook that exits without reading all of its input will cause us to get SIGPIPE, which will put us in a signal handler that then tries to kill() that same child.

This seems to be mostly harmless on Linux. The process has already exited by this point, and though kill() does not complain (since the process has not been reaped with a wait() call), it does not affect the exit status of the process.

However, this seems not to be true on all platforms. This case is triggered by t5401.13, "pre-receive hook that forgets to read its input". This test fails on NonStop since that hook was converted to the run_processes_parallel() API.

We can fix it by reordering the code a bit. We should run pp_init() first, and then push our SIG_IGN onto the stack afterwards, so that it is truly ignored while feeding the sub-processes.

Note that we also reorder the popping at the end of the function, too. This is not technically necessary, as we are doing two pops either way, but now the pops will correctly match their pushes.

This also fixes a related case that we can't test yet. If we did have more than one process to run, then one child causing SIGPIPE would cause us to kill() all of the children (which might still actually be running). But the hook API is the only user of the new feed_pipe feature, and it does not yet support parallel hook execution. So for now we'll always execute the processes sequentially. Once parallel hook execution exists, we'll be able to add a test which covers this.

Reported-by: Randall S. Becker <rsbecker@nexbridge.com>
Signed-off-by: Jeff King <peff@peff.net>
---
 run-command.c | 11 ++++++++---
 1 file changed, 8 insertions(+), 3 deletions(-)
diff --git a/run-command.c b/run-command.c
index 32c290ee6a..574d5c40f0 100644
--- a/run-command.c
+++ b/run-command.c
@@ -1895,14 +1895,19 @@ void run_processes_parallel(const struct run_process_parallel_opts *opts)
 					   "max:%"PRIuMAX,
 					   (uintmax_t)opts->processes);
 
+	pp_init(&pp, opts, &pp_sig);
+
 	/*
 	 * Child tasks might receive input via stdin, terminating early (or not), so
 	 * ignore the default SIGPIPE which gets handled by each feed_pipe_fn which
 	 * actually writes the data to children stdin fds.
+	 *
+	 * This _must_ come after pp_init(), because it installs its own
+	 * SIGPIPE handler (to cleanup children), and we want to supersede
+	 * that.
 	 */
 	sigchain_push(SIGPIPE, SIG_IGN);
 
-	pp_init(&pp, opts, &pp_sig);
 	while (1) {
 		for (i = 0;
 		    i < spawn_cap && !pp.shutdown &&
@@ -1928,10 +1933,10 @@ void run_processes_parallel(const struct run_process_parallel_opts *opts)
 		}
 	}
 
-	pp_cleanup(&pp, opts);
-
 	sigchain_pop(SIGPIPE);
 
+	pp_cleanup(&pp, opts);
+
 	if (do_trace2)
 		trace2_region_leave(tr2_category, tr2_label, NULL);
 }
-- 
2.54.0.rc1.274.g1f8c576c50
Previous: Adrian RatiuNext: Junio C Hamano
Message 14 of 18 in “Help needed on 2.54.0-rc0 t5301.13 looping.”
  1. rsbecker@nexbridge.comApr 7, 2026
  2. Jeff KingApr 8, 2026
  3. Jeff KingApr 8, 2026
  4. Adrian RatiuApr 8, 2026
  5. rsbecker@nexbridge.comApr 8, 2026
  6. rsbecker@nexbridge.comApr 8, 2026
  7. rsbecker@nexbridge.comApr 8, 2026
  8. Junio C HamanoApr 8, 2026
  9. rsbecker@nexbridge.comApr 8, 2026
  10. Adrian RatiuApr 8, 2026
  11. t5401: test SIGPIPE with parallel hooksJeff King, Apr 8, 2026
  12. Junio C HamanoApr 8, 2026
  13. Adrian RatiuApr 8, 2026
  14. run_processes_parallel(): fix order of sigpipe handlingJeff King, Apr 8, 2026
  15. Junio C HamanoApr 8, 2026
  16. Junio C HamanoApr 8, 2026
  17. Jeff KingApr 8, 2026
  18. Junio C HamanoApr 9, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.