{"thread":{"id":"63710","subject":"[PATCH] daemon: handle EINTR failures from waitpid()","startedAt":"2025-06-30T04:13:34Z","lastAt":"2025-06-30T12:18:27Z","messageCount":3,"participants":["Carlo Marcelo Arenas Belón","Phillip Wood"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"520898","messageId":"20250630041303.93370-1-carenas@gmail.com","threadId":"63710","inReplyTo":null,"subject":"[PATCH] daemon: handle EINTR failures from waitpid()","fromName":"Carlo Marcelo Arenas Belón","fromEmail":"carenas@gmail.com","sentAt":"2025-06-30T04:13:03Z","receivedAt":"2025-06-30T04:13:34Z","isPatch":true,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"Since 695605b508 (git-daemon: Simplify dead-children reaping logic,\n2008-08-14), the logic to check for zombie children was moved out of\nthe SIGCHLD signal handler, but adding checks for a failed waitpid()\nwere missed, with the possibility that a badly timed signal could\nprevent the promptly reaping of those defunct processes.\n\nAfter the refactoring of 30e1560230 (daemon: use run-command api for\nasync serving, 2010-11-04), that reproduced that bug, a single\nprocess could be skipped from reaping, so prevent that by adding the\nmissing error handling, and while at it make sure that ECHILD (or\nother errors) are correctly reported as a BUG().\n\nSigned-off-by: Carlo Marcelo Arenas Belón <carenas@gmail.com>\n---\n daemon.c | 5 ++++-\n 1 file changed, 4 insertions(+), 1 deletion(-)\n\ndiff --git a/daemon.c b/daemon.c\nindex d1be61fd57..16ae66a2da 100644\n--- a/daemon.c\n+++ b/daemon.c\n@@ -864,8 +864,11 @@ static void check_dead_children(void)\n \t\t\tlive_children--;\n \t\t\tchild_process_clear(&blanket->cld);\n \t\t\tfree(blanket);\n-\t\t} else\n+\t\t} else if (!pid)\n \t\t\tcradle = &blanket->next;\n+\t\telse if (errno != EINTR)\n+\t\t\tBUG(\"invalid child '%\" PRIuMAX \"'\",\n+\t\t\t    (uintmax_t)blanket->cld.pid);\n }\n \n static struct strvec cld_argv = STRVEC_INIT;\n-- \n2.50.0.132.g32f443f09a.dirty\n\n"},{"id":"520909","messageId":"16612f65-80ff-4162-a1df-31c2777eb848@gmail.com","threadId":"63710","inReplyTo":"20250630041303.93370-1-carenas@gmail.com","subject":"Re: [PATCH] daemon: handle EINTR failures from waitpid()","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2025-06-30T09:00:09Z","receivedAt":"2025-06-30T09:00:16Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Carlo\n\nOn 30/06/2025 05:13, Carlo Marcelo Arenas Belón wrote:\n> Since 695605b508 (git-daemon: Simplify dead-children reaping logic,\n> 2008-08-14), the logic to check for zombie children was moved out of\n> the SIGCHLD signal handler, but adding checks for a failed waitpid()\n> were missed, with the possibility that a badly timed signal could\n> prevent the promptly reaping of those defunct processes.\n> \n> After the refactoring of 30e1560230 (daemon: use run-command api for\n> async serving, 2010-11-04), that reproduced that bug, a single\n> process could be skipped from reaping, so prevent that by adding the\n> missing error handling, and while at it make sure that ECHILD (or\n> other errors) are correctly reported as a BUG().\n\nI agree with you analysis, I've left a couple of comments on the fix. I \nnoticed this when I was reading the code to see how well it handled \nEINTR and decided it wasn't worth worrying about as we still collect the \nchild the next time we call check_dead_children() but there is no harm \nin checking for EINTR here. It might be worth noting in the commit \nmessage that the linux man page for waitpid() explicitly says that EINTR \ncannot happen when WNOHANG is given though. I wonder if that is the case \non other platforms as well because the calling thread is not suspended \nand EINTR is usually associated with calls that block.\n\n> Signed-off-by: Carlo Marcelo Arenas Belón <carenas@gmail.com>\n> ---\n>   daemon.c | 5 ++++-\n>   1 file changed, 4 insertions(+), 1 deletion(-)\n> \n> diff --git a/daemon.c b/daemon.c\n> index d1be61fd57..16ae66a2da 100644\n> --- a/daemon.c\n> +++ b/daemon.c\n> @@ -864,8 +864,11 @@ static void check_dead_children(void)\n>   \t\t\tlive_children--;\n>   \t\t\tchild_process_clear(&blanket->cld);\n>   \t\t\tfree(blanket);\n> -\t\t} else\n> +\t\t} else if (!pid)\n\nOur style guidelines say that if one clause of an if statement needs \nbraces then all the clauses should be braced.\n\n>   \t\t\tcradle = &blanket->next;\n> +\t\telse if (errno != EINTR)\n> +\t\t\tBUG(\"invalid child '%\" PRIuMAX \"'\",\n> +\t\t\t    (uintmax_t)blanket->cld.pid);\n\nPOSIX says pid_t is signed so I'm not sure about the unsigned cast here. \nDo any of the platforms we support have a pid_t that is wider than a \nlong integer? I wondered if we should be logging an error instead of \ncalling BUG() but I think any error other that EINTR indicates a \nprogramming error so BUG() seems appropriate.\n\nThanks\n\nPhillip\n\n"},{"id":"520914","messageId":"bo6mr2zqf32wh6nxu53aoowy4xxcu3vle2wbepnbmyphe6b2tl@pofvznggzjui","threadId":"63710","inReplyTo":"16612f65-80ff-4162-a1df-31c2777eb848@gmail.com","subject":"Re: [PATCH] daemon: handle EINTR failures from waitpid()","fromName":"Carlo Marcelo Arenas Belón","fromEmail":"carenas@gmail.com","sentAt":"2025-06-30T12:18:25Z","receivedAt":"2025-06-30T12:18:27Z","isPatch":true,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"On Mon, Jun 30, 2025 at 10:00:09AM -0800, Phillip Wood wrote:\n> \n> On 30/06/2025 05:13, Carlo Marcelo Arenas Belón wrote:\n> > Since 695605b508 (git-daemon: Simplify dead-children reaping logic,\n> > 2008-08-14), the logic to check for zombie children was moved out of\n> > the SIGCHLD signal handler, but adding checks for a failed waitpid()\n> > were missed, with the possibility that a badly timed signal could\n> > prevent the promptly reaping of those defunct processes.\n> > \n> > After the refactoring of 30e1560230 (daemon: use run-command api for\n> > async serving, 2010-11-04), that reproduced that bug, a single\n> > process could be skipped from reaping, so prevent that by adding the\n> > missing error handling, and while at it make sure that ECHILD (or\n> > other errors) are correctly reported as a BUG().\n> \n> I agree with you analysis, I've left a couple of comments on the fix. I\n> noticed this when I was reading the code to see how well it handled EINTR\n> and decided it wasn't worth worrying about as we still collect the child the\n> next time we call check_dead_children() but there is no harm in checking for\n> EINTR here. It might be worth noting in the commit message that the linux\n> man page for waitpid() explicitly says that EINTR cannot happen when WNOHANG\n> is given though. I wonder if that is the case on other platforms as well\n> because the calling thread is not suspended and EINTR is usually associated\n> with calls that block.\n\nI wasn't aware of the comment in the Linux man page, and didn't see\nsomething similar in the ones I checked or the POSIX specification.\n\nIf WNOHANG prevents it from returning -1 with errno == EINTR, then my analysis\nis incorrect, and the last refactoring is the only one to blame as it didn't\nadd error handling from ECHILD.\n\nMore importantly, if we consider that regardless of the coment in the Linux\nman page (google found something similar in the one from zVM) that behaviour\nis implementation dependent it might be worth to fix also a similar use case\nin run_command.\n\n> >   \t\t\tcradle = &blanket->next;\n> > +\t\telse if (errno != EINTR)\n> > +\t\t\tBUG(\"invalid child '%\" PRIuMAX \"'\",\n> > +\t\t\t    (uintmax_t)blanket->cld.pid);\n> \n> POSIX says pid_t is signed so I'm not sure about the unsigned cast here.\n\nbut that is only so that a `(pid_t)-1` is valid AFAIK, and all \"real\" pid\nare expected to be positive (even in systems where pid_t is a 8 byte long\nlike Solaris).\n\ncasting them to unsigned to print them and using a uintmax_t for it was\nhow all pid are printed since 85e7283069 (cast pid_t's to uintmax_t to\nimprove portability, 2008-08-31) AFAIK.\n\n> Do\n> any of the platforms we support have a pid_t that is wider than a long\n> integer?\n\nthe ones in AIX are pretty long, but definitely no longer than INT_MAX (with\npid_t being 4 bytes long there).\n\nCarlo\n"}]}