{"thread":{"id":"66046","subject":"[PATCH] fsmonitor: flush pending FSEvents before cookie wait","startedAt":"2026-07-21T21:05:03Z","lastAt":"2026-08-13T09:03:19Z","messageCount":11,"participants":["Tamir Duberstein","Koji Nakamaru","Junio C Hamano","Patrick Steinhardt"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"548743","messageId":"20260721-fsmonitor-darwin-cookie-flush-v1-1-357dc5e32040@gmail.com","threadId":"66046","inReplyTo":null,"subject":"[PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-07-21T21:04:56Z","receivedAt":"2026-07-21T21:05:03Z","isPatch":true,"body":"56cef9cb1a (fsmonitor: use pthread_cond_timedwait for cookie wait,\n2026-04-15) limits the cookie wait to one second so that a filesystem\nwhich never delivers events cannot hang fsmonitor clients. A client that\ntimes out receives a trivial response and scans the entire index.\n\nFSEvents can defer delivery while it batches notifications and does not\nguarantee that its queue is drained in one latency interval. A loaded\nmacOS system can therefore time out even though the event stream is\nworking.\n\nOn an Apple M4 Max (16 cores, 128 GiB RAM) running macOS 26.5.2, two\nworktrees with a 1,001,178-entry index timed out 484 of 545 and 297 of\n365 fsmonitor requests. One status call performed 934,519 lstat() calls\nduring a 47-second preload and took 52 seconds overall.\n\nAsk FSEvents to flush pending notifications after creating the cookie\nand before starting the timed wait. Use the asynchronous form because\nthe client handler holds main_lock, which the listener callback also\nacquires. Keep the timeout and the behavior of the other backends\nunchanged.\n\nSigned-off-by: Tamir Duberstein <tamird@gmail.com>\n---\n builtin/fsmonitor--daemon.c          | 3 +++\n compat/fsmonitor/fsm-darwin-gcc.h    | 1 +\n compat/fsmonitor/fsm-listen-darwin.c | 5 +++++\n compat/fsmonitor/fsm-listen-linux.c  | 4 ++++\n compat/fsmonitor/fsm-listen-win32.c  | 4 ++++\n compat/fsmonitor/fsm-listen.h        | 6 ++++++\n 6 files changed, 23 insertions(+)\n\ndiff --git a/builtin/fsmonitor--daemon.c b/builtin/fsmonitor--daemon.c\nindex 4161dd8282..8e32b5ae5e 100644\n--- a/builtin/fsmonitor--daemon.c\n+++ b/builtin/fsmonitor--daemon.c\n@@ -206,6 +206,9 @@ static enum fsmonitor_cookie_item_result with_lock__wait_for_cookie(\n \tclose(fd);\n \tunlink(cookie_pathname.buf);\n \n+\t/* The listener callback takes main_lock, so this must not block. */\n+\tfsm_listen__flush_async(state);\n+\n \t/*\n \t * Wait for the listener thread to observe the cookie file.\n \t * Time out after a short interval so that the client\ndiff --git a/compat/fsmonitor/fsm-darwin-gcc.h b/compat/fsmonitor/fsm-darwin-gcc.h\nindex 3496e29b3a..c209dc2f68 100644\n--- a/compat/fsmonitor/fsm-darwin-gcc.h\n+++ b/compat/fsmonitor/fsm-darwin-gcc.h\n@@ -82,6 +82,7 @@ CFRunLoopRef CFRunLoopGetCurrent(void);\n extern CFStringRef kCFRunLoopDefaultMode;\n void FSEventStreamSetDispatchQueue(FSEventStreamRef stream, dispatch_queue_t q);\n unsigned char FSEventStreamStart(FSEventStreamRef stream);\n+FSEventStreamEventId FSEventStreamFlushAsync(FSEventStreamRef stream);\n void FSEventStreamStop(FSEventStreamRef stream);\n void FSEventStreamInvalidate(FSEventStreamRef stream);\n void FSEventStreamRelease(FSEventStreamRef stream);\ndiff --git a/compat/fsmonitor/fsm-listen-darwin.c b/compat/fsmonitor/fsm-listen-darwin.c\nindex 43c3a915a0..64bee248d2 100644\n--- a/compat/fsmonitor/fsm-listen-darwin.c\n+++ b/compat/fsmonitor/fsm-listen-darwin.c\n@@ -496,6 +496,11 @@ void fsm_listen__stop_async(struct fsmonitor_daemon_state *state)\n \tpthread_mutex_unlock(&data->dq_lock);\n }\n \n+void fsm_listen__flush_async(struct fsmonitor_daemon_state *state)\n+{\n+\tFSEventStreamFlushAsync(state->listen_data->stream);\n+}\n+\n void fsm_listen__loop(struct fsmonitor_daemon_state *state)\n {\n \tstruct fsm_listen_data *data;\ndiff --git a/compat/fsmonitor/fsm-listen-linux.c b/compat/fsmonitor/fsm-listen-linux.c\nindex e3dca14b62..7aae29ea22 100644\n--- a/compat/fsmonitor/fsm-listen-linux.c\n+++ b/compat/fsmonitor/fsm-listen-linux.c\n@@ -493,6 +493,10 @@ void fsm_listen__stop_async(struct fsmonitor_daemon_state *state)\n \t\tstate->listen_data->shutdown = SHUTDOWN_STOP;\n }\n \n+void fsm_listen__flush_async(struct fsmonitor_daemon_state *state UNUSED)\n+{\n+}\n+\n /*\n  * Process a single inotify event and queue for publication.\n  */\ndiff --git a/compat/fsmonitor/fsm-listen-win32.c b/compat/fsmonitor/fsm-listen-win32.c\nindex 9a6efc9bea..039d797000 100644\n--- a/compat/fsmonitor/fsm-listen-win32.c\n+++ b/compat/fsmonitor/fsm-listen-win32.c\n@@ -290,6 +290,10 @@ void fsm_listen__stop_async(struct fsmonitor_daemon_state *state)\n \tSetEvent(state->listen_data->hListener[LISTENER_SHUTDOWN]);\n }\n \n+void fsm_listen__flush_async(struct fsmonitor_daemon_state *state UNUSED)\n+{\n+}\n+\n static struct one_watch *create_watch(const char *path)\n {\n \tstruct one_watch *watch = NULL;\ndiff --git a/compat/fsmonitor/fsm-listen.h b/compat/fsmonitor/fsm-listen.h\nindex 41650bf897..cfeca1f4b6 100644\n--- a/compat/fsmonitor/fsm-listen.h\n+++ b/compat/fsmonitor/fsm-listen.h\n@@ -38,6 +38,12 @@ void fsm_listen__dtor(struct fsmonitor_daemon_state *state);\n  */\n void fsm_listen__loop(struct fsmonitor_daemon_state *state);\n \n+/*\n+ * Prompt the listener to deliver queued filesystem events, if supported.\n+ * This does not wait for the events to be processed.\n+ */\n+void fsm_listen__flush_async(struct fsmonitor_daemon_state *state);\n+\n /*\n  * Gently request that the fsmonitor listener thread shutdown.\n  * It does not wait for it to stop.  The caller should do a JOIN\n\n---\nbase-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200\nchange-id: 20260721-fsmonitor-darwin-cookie-flush-0f0d6e554a56\n\n"},{"id":"548850","messageId":"CAOTNsDy4pKbPHdK1T688Ax6Mgz15K-qfZR-8fAvTk48z3E43Rg@mail.gmail.com","threadId":"66046","inReplyTo":"20260721-fsmonitor-darwin-cookie-flush-v1-1-357dc5e32040@gmail.com","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Koji Nakamaru","fromEmail":"koji.nakamaru@gree.net","sentAt":"2026-07-24T02:41:08Z","receivedAt":"2026-07-24T02:41:20Z","isPatch":true,"body":"On Wed, Jul 22, 2026 at 6:05 AM Tamir Duberstein <tamird@gmail.com> wrote:\n>\n> 56cef9cb1a (fsmonitor: use pthread_cond_timedwait for cookie wait,\n> 2026-04-15) limits the cookie wait to one second so that a filesystem\n> which never delivers events cannot hang fsmonitor clients. A client that\n> times out receives a trivial response and scans the entire index.\n>\n> FSEvents can defer delivery while it batches notifications and does not\n> guarantee that its queue is drained in one latency interval. A loaded\n> macOS system can therefore time out even though the event stream is\n> working.\n>\n> On an Apple M4 Max (16 cores, 128 GiB RAM) running macOS 26.5.2, two\n> worktrees with a 1,001,178-entry index timed out 484 of 545 and 297 of\n> 365 fsmonitor requests. One status call performed 934,519 lstat() calls\n> during a 47-second preload and took 52 seconds overall.\n>\n> Ask FSEvents to flush pending notifications after creating the cookie\n> and before starting the timed wait. Use the asynchronous form because\n> the client handler holds main_lock, which the listener callback also\n> acquires. Keep the timeout and the behavior of the other backends\n> unchanged.\n>\n> Signed-off-by: Tamir Duberstein <tamird@gmail.com>\n> ---\n>  builtin/fsmonitor--daemon.c          | 3 +++\n>  compat/fsmonitor/fsm-darwin-gcc.h    | 1 +\n>  compat/fsmonitor/fsm-listen-darwin.c | 5 +++++\n>  compat/fsmonitor/fsm-listen-linux.c  | 4 ++++\n>  compat/fsmonitor/fsm-listen-win32.c  | 4 ++++\n>  compat/fsmonitor/fsm-listen.h        | 6 ++++++\n>  6 files changed, 23 insertions(+)\n>\n> diff --git a/builtin/fsmonitor--daemon.c b/builtin/fsmonitor--daemon.c\n> index 4161dd8282..8e32b5ae5e 100644\n> --- a/builtin/fsmonitor--daemon.c\n> +++ b/builtin/fsmonitor--daemon.c\n> @@ -206,6 +206,9 @@ static enum fsmonitor_cookie_item_result with_lock__wait_for_cookie(\n>         close(fd);\n>         unlink(cookie_pathname.buf);\n>\n> +       /* The listener callback takes main_lock, so this must not block. */\n> +       fsm_listen__flush_async(state);\n> +\n>         /*\n>          * Wait for the listener thread to observe the cookie file.\n>          * Time out after a short interval so that the client\n> diff --git a/compat/fsmonitor/fsm-darwin-gcc.h b/compat/fsmonitor/fsm-darwin-gcc.h\n> index 3496e29b3a..c209dc2f68 100644\n> --- a/compat/fsmonitor/fsm-darwin-gcc.h\n> +++ b/compat/fsmonitor/fsm-darwin-gcc.h\n> @@ -82,6 +82,7 @@ CFRunLoopRef CFRunLoopGetCurrent(void);\n>  extern CFStringRef kCFRunLoopDefaultMode;\n>  void FSEventStreamSetDispatchQueue(FSEventStreamRef stream, dispatch_queue_t q);\n>  unsigned char FSEventStreamStart(FSEventStreamRef stream);\n> +FSEventStreamEventId FSEventStreamFlushAsync(FSEventStreamRef stream);\n>  void FSEventStreamStop(FSEventStreamRef stream);\n>  void FSEventStreamInvalidate(FSEventStreamRef stream);\n>  void FSEventStreamRelease(FSEventStreamRef stream);\n> diff --git a/compat/fsmonitor/fsm-listen-darwin.c b/compat/fsmonitor/fsm-listen-darwin.c\n> index 43c3a915a0..64bee248d2 100644\n> --- a/compat/fsmonitor/fsm-listen-darwin.c\n> +++ b/compat/fsmonitor/fsm-listen-darwin.c\n> @@ -496,6 +496,11 @@ void fsm_listen__stop_async(struct fsmonitor_daemon_state *state)\n>         pthread_mutex_unlock(&data->dq_lock);\n>  }\n>\n> +void fsm_listen__flush_async(struct fsmonitor_daemon_state *state)\n> +{\n> +       FSEventStreamFlushAsync(state->listen_data->stream);\n> +}\n> +\n>  void fsm_listen__loop(struct fsmonitor_daemon_state *state)\n>  {\n>         struct fsm_listen_data *data;\n> diff --git a/compat/fsmonitor/fsm-listen-linux.c b/compat/fsmonitor/fsm-listen-linux.c\n> index e3dca14b62..7aae29ea22 100644\n> --- a/compat/fsmonitor/fsm-listen-linux.c\n> +++ b/compat/fsmonitor/fsm-listen-linux.c\n> @@ -493,6 +493,10 @@ void fsm_listen__stop_async(struct fsmonitor_daemon_state *state)\n>                 state->listen_data->shutdown = SHUTDOWN_STOP;\n>  }\n>\n> +void fsm_listen__flush_async(struct fsmonitor_daemon_state *state UNUSED)\n> +{\n> +}\n> +\n>  /*\n>   * Process a single inotify event and queue for publication.\n>   */\n> diff --git a/compat/fsmonitor/fsm-listen-win32.c b/compat/fsmonitor/fsm-listen-win32.c\n> index 9a6efc9bea..039d797000 100644\n> --- a/compat/fsmonitor/fsm-listen-win32.c\n> +++ b/compat/fsmonitor/fsm-listen-win32.c\n> @@ -290,6 +290,10 @@ void fsm_listen__stop_async(struct fsmonitor_daemon_state *state)\n>         SetEvent(state->listen_data->hListener[LISTENER_SHUTDOWN]);\n>  }\n>\n> +void fsm_listen__flush_async(struct fsmonitor_daemon_state *state UNUSED)\n> +{\n> +}\n> +\n>  static struct one_watch *create_watch(const char *path)\n>  {\n>         struct one_watch *watch = NULL;\n> diff --git a/compat/fsmonitor/fsm-listen.h b/compat/fsmonitor/fsm-listen.h\n> index 41650bf897..cfeca1f4b6 100644\n> --- a/compat/fsmonitor/fsm-listen.h\n> +++ b/compat/fsmonitor/fsm-listen.h\n> @@ -38,6 +38,12 @@ void fsm_listen__dtor(struct fsmonitor_daemon_state *state);\n>   */\n>  void fsm_listen__loop(struct fsmonitor_daemon_state *state);\n>\n> +/*\n> + * Prompt the listener to deliver queued filesystem events, if supported.\n> + * This does not wait for the events to be processed.\n> + */\n> +void fsm_listen__flush_async(struct fsmonitor_daemon_state *state);\n> +\n>  /*\n>   * Gently request that the fsmonitor listener thread shutdown.\n>   * It does not wait for it to stop.  The caller should do a JOIN\n>\n> ---\n> base-commit: 5d2e7709234afea1b6ddb25cd4f60d3d5fb3c200\n> change-id: 20260721-fsmonitor-darwin-cookie-flush-0f0d6e554a56\n>\n\nThis patch is carefully designed to minimize any risks. To drain events,\nwe could also call FSEventStreamFlushSync before acquiring main_lock in\ndo_handle_client(), but this patch should be sufficient if it mitigates\nthe issue. The commit message would be much more convincing if you also\nincluded benchmark results showing how many timeouts were reduced.\n\n--\nKoji Nakamaru\n"},{"id":"548917","messageId":"xmqqh5lo5dib.fsf@gitster.g","threadId":"66046","inReplyTo":"CAOTNsDy4pKbPHdK1T688Ax6Mgz15K-qfZR-8fAvTk48z3E43Rg@mail.gmail.com","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-24T20:38:52Z","receivedAt":"2026-07-24T20:38:55Z","isPatch":true,"body":"Koji Nakamaru <koji.nakamaru@gree.net> writes:\n\n> On Wed, Jul 22, 2026 at 6:05 AM Tamir Duberstein <tamird@gmail.com> wrote:\n>>\n>> 56cef9cb1a (fsmonitor: use pthread_cond_timedwait for cookie wait,\n>> 2026-04-15) limits the cookie wait to one second so that a filesystem\n>> which never delivers events cannot hang fsmonitor clients. A client that\n>> times out receives a trivial response and scans the entire index.\n>>\n>> FSEvents can defer delivery while it batches notifications and does not\n>> guarantee that its queue is drained in one latency interval. A loaded\n>> macOS system can therefore time out even though the event stream is\n>> working.\n>>\n>> On an Apple M4 Max (16 cores, 128 GiB RAM) running macOS 26.5.2, two\n>> worktrees with a 1,001,178-entry index timed out 484 of 545 and 297 of\n>> 365 fsmonitor requests. One status call performed 934,519 lstat() calls\n>> during a 47-second preload and took 52 seconds overall.\n>>\n>> Ask FSEvents to flush pending notifications after creating the cookie\n>> and before starting the timed wait. Use the asynchronous form because\n>> the client handler holds main_lock, which the listener callback also\n>> acquires. Keep the timeout and the behavior of the other backends\n>> unchanged.\n>>\n>> Signed-off-by: Tamir Duberstein <tamird@gmail.com>\n>> ---\n>>...\n> This patch is carefully designed to minimize any risks. To drain events,\n> we could also call FSEventStreamFlushSync before acquiring main_lock in\n> do_handle_client(), but this patch should be sufficient if it mitigates\n> the issue. The commit message would be much more convincing if you also\n> included benchmark results showing how many timeouts were reduced.\n\nThanks for a review.\n"},{"id":"549628","messageId":"xmqq4ih9ttyc.fsf@gitster.g","threadId":"66046","inReplyTo":"CAOTNsDy4pKbPHdK1T688Ax6Mgz15K-qfZR-8fAvTk48z3E43Rg@mail.gmail.com","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-04T22:13:47Z","receivedAt":"2026-08-04T22:13:51Z","isPatch":true,"body":"Koji Nakamaru <koji.nakamaru@gree.net> writes:\n\n> On Wed, Jul 22, 2026 at 6:05 AM Tamir Duberstein <tamird@gmail.com> wrote:\n>>\n>> 56cef9cb1a (fsmonitor: use pthread_cond_timedwait for cookie wait,\n>> 2026-04-15) limits the cookie wait to one second so that a filesystem\n>> which never delivers events cannot hang fsmonitor clients. A client that\n>> times out receives a trivial response and scans the entire index.\n>>\n>> FSEvents can defer delivery while it batches notifications and does not\n>> guarantee that its queue is drained in one latency interval. A loaded\n>> macOS system can therefore time out even though the event stream is\n>> working.\n>> ...\n>\n> This patch is carefully designed to minimize any risks. To drain events,\n> we could also call FSEventStreamFlushSync before acquiring main_lock in\n> do_handle_client(), but this patch should be sufficient if it mitigates\n> the issue. The commit message would be much more convincing if you also\n> included benchmark results showing how many timeouts were reduced.\n\nTamir, just to say that it is my understanding that the ball is in\nyour court.  It hasn't been _too_ long since the exchange happened,\nbut we expect people to respond review comments (either positively\nor negatively) and without such discourse a topic would not move\nforward, so ...\n\nThanks.\n"},{"id":"549657","messageId":"anLtSOKqgcCrrNHo@pks.im","threadId":"66046","inReplyTo":"20260721-fsmonitor-darwin-cookie-flush-v1-1-357dc5e32040@gmail.com","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T07:59:15Z","receivedAt":"2026-08-05T07:59:27Z","isPatch":true,"body":"On Tue, Jul 21, 2026 at 05:04:56PM -0400, Tamir Duberstein wrote:\n> 56cef9cb1a (fsmonitor: use pthread_cond_timedwait for cookie wait,\n> 2026-04-15) limits the cookie wait to one second so that a filesystem\n> which never delivers events cannot hang fsmonitor clients. A client that\n> times out receives a trivial response and scans the entire index.\n> \n> FSEvents can defer delivery while it batches notifications and does not\n> guarantee that its queue is drained in one latency interval. A loaded\n> macOS system can therefore time out even though the event stream is\n> working.\n> \n> On an Apple M4 Max (16 cores, 128 GiB RAM) running macOS 26.5.2, two\n> worktrees with a 1,001,178-entry index timed out 484 of 545 and 297 of\n> 365 fsmonitor requests. One status call performed 934,519 lstat() calls\n> during a 47-second preload and took 52 seconds overall.\n> \n> Ask FSEvents to flush pending notifications after creating the cookie\n> and before starting the timed wait. Use the asynchronous form because\n> the client handler holds main_lock, which the listener callback also\n> acquires. Keep the timeout and the behavior of the other backends\n> unchanged.\n\nI cannot really say much about the FSEvent interfaces, but to me it\nfeels quite reasonable to flush the queue when we are waiting for events\nto be delivered. And that's exactly what `FSEventStreamFlushAsync()`\ndoes: it basically overrides the latency we have configured (which is\n1ms) and asks the kernel to flush stuff immediately.\n\n> diff --git a/builtin/fsmonitor--daemon.c b/builtin/fsmonitor--daemon.c\n> index 4161dd8282..8e32b5ae5e 100644\n> --- a/builtin/fsmonitor--daemon.c\n> +++ b/builtin/fsmonitor--daemon.c\n> @@ -206,6 +206,9 @@ static enum fsmonitor_cookie_item_result with_lock__wait_for_cookie(\n>  \tclose(fd);\n>  \tunlink(cookie_pathname.buf);\n>  \n> +\t/* The listener callback takes main_lock, so this must not block. */\n> +\tfsm_listen__flush_async(state);\n> +\n>  \t/*\n>  \t * Wait for the listener thread to observe the cookie file.\n>  \t * Time out after a short interval so that the client\n\nOkay, so we've unlinked the cookie file and the next thing is that we're\nwaiting for all events to have been processed. As said, it feels\nreasonable that we're flushing all events before we start waiting for\nthem.\n\nWhat I find surprising though is that this is supposed to make a\ndifference at all. The latency we pass to `FSEventStreamCreate()` is\n1 millisecond, and we wait up to 1 second for the cookie event. I would\nhave expected that batching events for 1 milliseconds should be totally\nfine when we're waiting for a full second anyway.\n\nSo given that I cannot verify this at all and that I have no clue about\nthe FSEvent interfaces... do you have any explanation why the flush\nseems to help regardless?\n\nI _think_ you're already hinting at this in the commit message, where\nyou say that it's not guaranteed that the queue is drained in a single\nlatency interval. Is there any documentation that tells us what the\nprovided guarantees are?\n\nOther than that the code changes look sensible to me, thanks!\n\nPatrick\n"},{"id":"550289","messageId":"CAJ-ks9=+4rxxx8+7fOF1aLFW67=hdxjhQsHqse1GGBLwZUh2BQ@mail.gmail.com","threadId":"66046","inReplyTo":"anLtSOKqgcCrrNHo@pks.im","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-08-11T15:22:01Z","receivedAt":"2026-08-11T15:22:40Z","isPatch":true,"body":"On Wed, Aug 5, 2026 at 3:59 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Tue, Jul 21, 2026 at 05:04:56PM -0400, Tamir Duberstein wrote:\n> > 56cef9cb1a (fsmonitor: use pthread_cond_timedwait for cookie wait,\n> > 2026-04-15) limits the cookie wait to one second so that a filesystem\n> > which never delivers events cannot hang fsmonitor clients. A client that\n> > times out receives a trivial response and scans the entire index.\n> >\n> > FSEvents can defer delivery while it batches notifications and does not\n> > guarantee that its queue is drained in one latency interval. A loaded\n> > macOS system can therefore time out even though the event stream is\n> > working.\n> >\n> > On an Apple M4 Max (16 cores, 128 GiB RAM) running macOS 26.5.2, two\n> > worktrees with a 1,001,178-entry index timed out 484 of 545 and 297 of\n> > 365 fsmonitor requests. One status call performed 934,519 lstat() calls\n> > during a 47-second preload and took 52 seconds overall.\n> >\n> > Ask FSEvents to flush pending notifications after creating the cookie\n> > and before starting the timed wait. Use the asynchronous form because\n> > the client handler holds main_lock, which the listener callback also\n> > acquires. Keep the timeout and the behavior of the other backends\n> > unchanged.\n>\n> I cannot really say much about the FSEvent interfaces, but to me it\n> feels quite reasonable to flush the queue when we are waiting for events\n> to be delivered. And that's exactly what `FSEventStreamFlushAsync()`\n> does: it basically overrides the latency we have configured (which is\n> 1ms) and asks the kernel to flush stuff immediately.\n>\n> > diff --git a/builtin/fsmonitor--daemon.c b/builtin/fsmonitor--daemon.c\n> > index 4161dd8282..8e32b5ae5e 100644\n> > --- a/builtin/fsmonitor--daemon.c\n> > +++ b/builtin/fsmonitor--daemon.c\n> > @@ -206,6 +206,9 @@ static enum fsmonitor_cookie_item_result with_lock__wait_for_cookie(\n> >       close(fd);\n> >       unlink(cookie_pathname.buf);\n> >\n> > +     /* The listener callback takes main_lock, so this must not block. */\n> > +     fsm_listen__flush_async(state);\n> > +\n> >       /*\n> >        * Wait for the listener thread to observe the cookie file.\n> >        * Time out after a short interval so that the client\n>\n> Okay, so we've unlinked the cookie file and the next thing is that we're\n> waiting for all events to have been processed. As said, it feels\n> reasonable that we're flushing all events before we start waiting for\n> them.\n>\n> What I find surprising though is that this is supposed to make a\n> difference at all. The latency we pass to `FSEventStreamCreate()` is\n> 1 millisecond, and we wait up to 1 second for the cookie event. I would\n> have expected that batching events for 1 milliseconds should be totally\n> fine when we're waiting for a full second anyway.\n>\n> So given that I cannot verify this at all and that I have no clue about\n> the FSEvent interfaces... do you have any explanation why the flush\n> seems to help regardless?\n>\n> I _think_ you're already hinting at this in the commit message, where\n> you say that it's not guaranteed that the queue is drained in a single\n> latency interval. Is there any documentation that tells us what the\n> provided guarantees are?\n>\n> Other than that the code changes look sensible to me, thanks!\n>\n> Patrick\n\nThe following was generated by my coding agent and fact checked and\nedited by me mainly to address you in the second person.\n\nYour question was already answered by the original Git implementation\n- and you yourself predicted this exact regression before it landed.\n\nIn March 2022, Jeff Hostetler introduced Git’s fsmonitor cookie\nprotocol in commit b05880d357. Its commit message explicitly says\nmacOS “does not guarantee that the kernel queue is completely drained”\nafter one FSEvents latency interval. That is precisely why Git\noriginally waited until it actually observed the cookie. Original\ncookie implementation\n(https://github.com/git/git/commit/b05880d357c6dadba8d1d7943f4782fc25e06999)\n\nRegression timeline:\n\n1. February 2026: Paul Tarjan proposed replacing the indefinite cookie\nwait with a one-second timeout to prevent hangs on Linux filesystems\nthat never deliver events. Junio questioned whether one second was\nappropriate and warned about expensive full-scan fallbacks. Junio’s\ninitial concern\n(https://lore.kernel.org/git/xmqqzf4w8r20.fsf@gitster.g/); Junio’s\nfull-scan warning\n(https://lore.kernel.org/git/xmqqfr6mt9uk.fsf@gitster.g/)\n2. Paul’s assumption: He argued that the timeout would trigger only on\nbroken filesystems that never deliver events, while working\nfilesystems would respond promptly. Paul’s explanation\n(https://lore.kernel.org/git/20260227063118.9069-1-github@paulisageek.com/)\n3. March 4: You (Patrick) asked: “Are we sure this is always enough on\na loaded system?” Paul responded that even if a timeout occurred, the\nfallback would simply involve some additional work. Patrick’s earlier\nwarning (https://lore.kernel.org/git/aafifU-befdZW4O0@pks.im/); Paul’s\nresponse (https://lore.kernel.org/git/20260304181745.25673-1-github@paulisageek.com/)\n4. April 15: The one-second timeout landed anyway as 56cef9cb1a.\nAccepted timeout change\n(https://github.com/git/git/commit/56cef9cb1a083c47b12b88548bf2126af8bfb263)\n5. July 21: My (tamird) measurements disproved both assumptions: 781\nof 910 requests timed out on functioning macOS worktrees, and one\nfallback caused 934,519 lstat() calls and a 52-second git status. The\nresult is this patch.\n\nSummary:\n\nThe 1 ms value is not a delivery deadline:\n\n- Apple defines it as the delay the userspace service should apply\nafter it hears about an event from the kernel. It says nothing about\nkernel backlog, service scheduling, callback scheduling, or complete\nqueue drainage. The installed SDK spells this out in\n/Library/Developer/CommandLineTools/SDKs/MacOSX.sdk/System/Library/Frameworks/CoreServices.framework/Frameworks/FSEvents.framework/Headers/FSEvents.h:763.\n\n *    latency:\n *      The number of seconds the service should wait after hearing\n *      about an event from the kernel before passing it along to the\n *      client via its callback. Specifying a larger value may result\n *      in more effective temporal coalescing, resulting in fewer\n *      callbacks and greater overall efficiency.\n\n- Apple explicitly describes notification latency as “inherently\nnon-deterministic.” Apple’s FSEvents programming guide\n(https://developer.apple.com/library/archive/documentation/Darwin/Conceptual/FSEvents_ProgGuide/UsingtheFSEventsFramework/UsingtheFSEventsFramework.html)\n- Apple’s kernel independently implements a 10 ms event-batching\ntimer, demonstrating that the userspace 1 ms parameter is not even the\nonly batching interval. This does not itself explain a one-second\ndelay; it disproves treating 1 ms as an end-to-end guarantee. Apple\nXNU FSEvents implementation\n(https://github.com/apple-oss-distributions/xnu/blob/f6217f891ac0bb64f3d375211650a4c1ff8ca1ea/bsd/vfs/vfs_fsevents.c#L1479-L1525)\n- Git’s original Darwin implementation chose 1 ms because 100 ms\ncaused dropped events in a 100,000-file stress test—not because Apple\nguaranteed delivery within 1 ms. See\ncompat/fsmonitor/fsm-listen-darwin.c:437.\n- FSEventStreamFlushAsync() requests delivery of pending events\nwithout blocking. A synchronous flush at the existing call site would\ndeadlock because the caller already holds the mutex needed by the\ncallback. See builtin/fsmonitor--daemon.c:247.\n\nHope that's helpful.\n"},{"id":"550293","messageId":"antMfAYVSPX9QAk1@pks.im","threadId":"66046","inReplyTo":"CAJ-ks9=+4rxxx8+7fOF1aLFW67=hdxjhQsHqse1GGBLwZUh2BQ@mail.gmail.com","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-11T16:23:24Z","receivedAt":"2026-08-11T16:23:31Z","isPatch":true,"body":"On Tue, Aug 11, 2026 at 11:22:01AM -0400, Tamir Duberstein wrote:\n> On Wed, Aug 5, 2026 at 3:59 AM Patrick Steinhardt <ps@pks.im> wrote:\n> > On Tue, Jul 21, 2026 at 05:04:56PM -0400, Tamir Duberstein wrote:\n> > > 56cef9cb1a (fsmonitor: use pthread_cond_timedwait for cookie wait,\n> > > 2026-04-15) limits the cookie wait to one second so that a filesystem\n> > > which never delivers events cannot hang fsmonitor clients. A client that\n> > > times out receives a trivial response and scans the entire index.\n> > >\n> > > FSEvents can defer delivery while it batches notifications and does not\n> > > guarantee that its queue is drained in one latency interval. A loaded\n> > > macOS system can therefore time out even though the event stream is\n> > > working.\n> > >\n> > > On an Apple M4 Max (16 cores, 128 GiB RAM) running macOS 26.5.2, two\n> > > worktrees with a 1,001,178-entry index timed out 484 of 545 and 297 of\n> > > 365 fsmonitor requests. One status call performed 934,519 lstat() calls\n> > > during a 47-second preload and took 52 seconds overall.\n> > >\n> > > Ask FSEvents to flush pending notifications after creating the cookie\n> > > and before starting the timed wait. Use the asynchronous form because\n> > > the client handler holds main_lock, which the listener callback also\n> > > acquires. Keep the timeout and the behavior of the other backends\n> > > unchanged.\n> >\n> > I cannot really say much about the FSEvent interfaces, but to me it\n> > feels quite reasonable to flush the queue when we are waiting for events\n> > to be delivered. And that's exactly what `FSEventStreamFlushAsync()`\n> > does: it basically overrides the latency we have configured (which is\n> > 1ms) and asks the kernel to flush stuff immediately.\n> >\n> > > diff --git a/builtin/fsmonitor--daemon.c b/builtin/fsmonitor--daemon.c\n> > > index 4161dd8282..8e32b5ae5e 100644\n> > > --- a/builtin/fsmonitor--daemon.c\n> > > +++ b/builtin/fsmonitor--daemon.c\n> > > @@ -206,6 +206,9 @@ static enum fsmonitor_cookie_item_result with_lock__wait_for_cookie(\n> > >       close(fd);\n> > >       unlink(cookie_pathname.buf);\n> > >\n> > > +     /* The listener callback takes main_lock, so this must not block. */\n> > > +     fsm_listen__flush_async(state);\n> > > +\n> > >       /*\n> > >        * Wait for the listener thread to observe the cookie file.\n> > >        * Time out after a short interval so that the client\n> >\n> > Okay, so we've unlinked the cookie file and the next thing is that we're\n> > waiting for all events to have been processed. As said, it feels\n> > reasonable that we're flushing all events before we start waiting for\n> > them.\n> >\n> > What I find surprising though is that this is supposed to make a\n> > difference at all. The latency we pass to `FSEventStreamCreate()` is\n> > 1 millisecond, and we wait up to 1 second for the cookie event. I would\n> > have expected that batching events for 1 milliseconds should be totally\n> > fine when we're waiting for a full second anyway.\n> >\n> > So given that I cannot verify this at all and that I have no clue about\n> > the FSEvent interfaces... do you have any explanation why the flush\n> > seems to help regardless?\n> >\n> > I _think_ you're already hinting at this in the commit message, where\n> > you say that it's not guaranteed that the queue is drained in a single\n> > latency interval. Is there any documentation that tells us what the\n> > provided guarantees are?\n> >\n> > Other than that the code changes look sensible to me, thanks!\n> >\n> > Patrick\n> \n> The following was generated by my coding agent and fact checked and\n> edited by me mainly to address you in the second person.\n> \n[snip]\n> \n> Hope that's helpful.\n\nSorry, but that's not quite helpful. The questions I'm asking are to\nverify whether you understand the consequences and subtleties around the\ncode area that you're proposing to change. If I wanted to only learn\nabout this myself then I could simply ask an agent myself, but that's\nnot really the intent of a code review.\n\nSo what I'm looking for is _your_ explanation, not the explanation of\nAI. Your explanation may of course be informed by AI. But if so it's\nyour responsibility to double-check its assumptions, build your own\nmodel and then share your informed opinion with us.\n\nRight now I don't yet have the feeling that you understand why this\nfixes the underlying issue.\n\nThanks!\n\nPatrick\n"},{"id":"550302","messageId":"CAJ-ks9=oV4SQSjTHNEOGBaQb8Rb4xBqVSp4wYum6yzU-zx3YtQ@mail.gmail.com","threadId":"66046","inReplyTo":"antMfAYVSPX9QAk1@pks.im","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-08-11T16:45:03Z","receivedAt":"2026-08-11T16:45:42Z","isPatch":true,"body":"On Tue, Aug 11, 2026 at 12:23 PM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> On Tue, Aug 11, 2026 at 11:22:01AM -0400, Tamir Duberstein wrote:\n> > On Wed, Aug 5, 2026 at 3:59 AM Patrick Steinhardt <ps@pks.im> wrote:\n> > > On Tue, Jul 21, 2026 at 05:04:56PM -0400, Tamir Duberstein wrote:\n> > > > 56cef9cb1a (fsmonitor: use pthread_cond_timedwait for cookie wait,\n> > > > 2026-04-15) limits the cookie wait to one second so that a filesystem\n> > > > which never delivers events cannot hang fsmonitor clients. A client that\n> > > > times out receives a trivial response and scans the entire index.\n> > > >\n> > > > FSEvents can defer delivery while it batches notifications and does not\n> > > > guarantee that its queue is drained in one latency interval. A loaded\n> > > > macOS system can therefore time out even though the event stream is\n> > > > working.\n> > > >\n> > > > On an Apple M4 Max (16 cores, 128 GiB RAM) running macOS 26.5.2, two\n> > > > worktrees with a 1,001,178-entry index timed out 484 of 545 and 297 of\n> > > > 365 fsmonitor requests. One status call performed 934,519 lstat() calls\n> > > > during a 47-second preload and took 52 seconds overall.\n> > > >\n> > > > Ask FSEvents to flush pending notifications after creating the cookie\n> > > > and before starting the timed wait. Use the asynchronous form because\n> > > > the client handler holds main_lock, which the listener callback also\n> > > > acquires. Keep the timeout and the behavior of the other backends\n> > > > unchanged.\n> > >\n> > > I cannot really say much about the FSEvent interfaces, but to me it\n> > > feels quite reasonable to flush the queue when we are waiting for events\n> > > to be delivered. And that's exactly what `FSEventStreamFlushAsync()`\n> > > does: it basically overrides the latency we have configured (which is\n> > > 1ms) and asks the kernel to flush stuff immediately.\n> > >\n> > > > diff --git a/builtin/fsmonitor--daemon.c b/builtin/fsmonitor--daemon.c\n> > > > index 4161dd8282..8e32b5ae5e 100644\n> > > > --- a/builtin/fsmonitor--daemon.c\n> > > > +++ b/builtin/fsmonitor--daemon.c\n> > > > @@ -206,6 +206,9 @@ static enum fsmonitor_cookie_item_result with_lock__wait_for_cookie(\n> > > >       close(fd);\n> > > >       unlink(cookie_pathname.buf);\n> > > >\n> > > > +     /* The listener callback takes main_lock, so this must not block. */\n> > > > +     fsm_listen__flush_async(state);\n> > > > +\n> > > >       /*\n> > > >        * Wait for the listener thread to observe the cookie file.\n> > > >        * Time out after a short interval so that the client\n> > >\n> > > Okay, so we've unlinked the cookie file and the next thing is that we're\n> > > waiting for all events to have been processed. As said, it feels\n> > > reasonable that we're flushing all events before we start waiting for\n> > > them.\n> > >\n> > > What I find surprising though is that this is supposed to make a\n> > > difference at all. The latency we pass to `FSEventStreamCreate()` is\n> > > 1 millisecond, and we wait up to 1 second for the cookie event. I would\n> > > have expected that batching events for 1 milliseconds should be totally\n> > > fine when we're waiting for a full second anyway.\n> > >\n> > > So given that I cannot verify this at all and that I have no clue about\n> > > the FSEvent interfaces... do you have any explanation why the flush\n> > > seems to help regardless?\n> > >\n> > > I _think_ you're already hinting at this in the commit message, where\n> > > you say that it's not guaranteed that the queue is drained in a single\n> > > latency interval. Is there any documentation that tells us what the\n> > > provided guarantees are?\n> > >\n> > > Other than that the code changes look sensible to me, thanks!\n> > >\n> > > Patrick\n> >\n> > The following was generated by my coding agent and fact checked and\n> > edited by me mainly to address you in the second person.\n> >\n> [snip]\n> >\n> > Hope that's helpful.\n>\n> Sorry, but that's not quite helpful. The questions I'm asking are to\n> verify whether you understand the consequences and subtleties around the\n> code area that you're proposing to change. If I wanted to only learn\n> about this myself then I could simply ask an agent myself, but that's\n> not really the intent of a code review.\n>\n> So what I'm looking for is _your_ explanation, not the explanation of\n> AI. Your explanation may of course be informed by AI. But if so it's\n> your responsibility to double-check its assumptions, build your own\n> model and then share your informed opinion with us.\n>\n> Right now I don't yet have the feeling that you understand why this\n> fixes the underlying issue.\n\nGot it. I agree with you that the flush call feels unnecessary under\nthe interpretation that passing 1ms to FSEventStreamCreate is the\nequivalent of asking it to flush every 1ms. Empirically, though,\nthat's not the case, as described in the commit message.\n\nThere's more precedent for this technique (found by agent, sorry):\nwatchman fixed a similar issue here:\nhttps://github.com/facebook/watchman/commit/d1795de4ecab33672a89802318fe6f0122462194\nand the documented it here:\nhttps://github.com/facebook/watchman/commit/2f80886991ce81585ac0679c2b019fa0e4d9e9dd\n\nI agree this is unsatisfying.\n\nDoes that help?\n"},{"id":"550306","messageId":"xmqqh5l036i8.fsf@gitster.g","threadId":"66046","inReplyTo":"antMfAYVSPX9QAk1@pks.im","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-08-11T17:35:11Z","receivedAt":"2026-08-11T17:35:14Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> On Tue, Aug 11, 2026 at 11:22:01AM -0400, Tamir Duberstein wrote:\n>...\n>> Hope that's helpful.\n>\n> Sorry, but that's not quite helpful. The questions I'm asking are to\n\nThanks for pushing back.\n"},{"id":"550466","messageId":"CAJ-ks9kQR77vH-56eS9tT-iXEnih+Z7SPRMs1gD_wTyg_6gZ_w@mail.gmail.com","threadId":"66046","inReplyTo":"CAJ-ks9=oV4SQSjTHNEOGBaQb8Rb4xBqVSp4wYum6yzU-zx3YtQ@mail.gmail.com","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-08-13T01:19:08Z","receivedAt":"2026-08-13T01:19:47Z","isPatch":true,"body":"On Tue, Aug 11, 2026 at 12:45 PM Tamir Duberstein <tamird@gmail.com> wrote:\n>\n> On Tue, Aug 11, 2026 at 12:23 PM Patrick Steinhardt <ps@pks.im> wrote:\n> >\n> > On Tue, Aug 11, 2026 at 11:22:01AM -0400, Tamir Duberstein wrote:\n> > > On Wed, Aug 5, 2026 at 3:59 AM Patrick Steinhardt <ps@pks.im> wrote:\n> > > > On Tue, Jul 21, 2026 at 05:04:56PM -0400, Tamir Duberstein wrote:\n> > > > > 56cef9cb1a (fsmonitor: use pthread_cond_timedwait for cookie wait,\n> > > > > 2026-04-15) limits the cookie wait to one second so that a filesystem\n> > > > > which never delivers events cannot hang fsmonitor clients. A client that\n> > > > > times out receives a trivial response and scans the entire index.\n> > > > >\n> > > > > FSEvents can defer delivery while it batches notifications and does not\n> > > > > guarantee that its queue is drained in one latency interval. A loaded\n> > > > > macOS system can therefore time out even though the event stream is\n> > > > > working.\n> > > > >\n> > > > > On an Apple M4 Max (16 cores, 128 GiB RAM) running macOS 26.5.2, two\n> > > > > worktrees with a 1,001,178-entry index timed out 484 of 545 and 297 of\n> > > > > 365 fsmonitor requests. One status call performed 934,519 lstat() calls\n> > > > > during a 47-second preload and took 52 seconds overall.\n> > > > >\n> > > > > Ask FSEvents to flush pending notifications after creating the cookie\n> > > > > and before starting the timed wait. Use the asynchronous form because\n> > > > > the client handler holds main_lock, which the listener callback also\n> > > > > acquires. Keep the timeout and the behavior of the other backends\n> > > > > unchanged.\n> > > >\n> > > > I cannot really say much about the FSEvent interfaces, but to me it\n> > > > feels quite reasonable to flush the queue when we are waiting for events\n> > > > to be delivered. And that's exactly what `FSEventStreamFlushAsync()`\n> > > > does: it basically overrides the latency we have configured (which is\n> > > > 1ms) and asks the kernel to flush stuff immediately.\n> > > >\n> > > > > diff --git a/builtin/fsmonitor--daemon.c b/builtin/fsmonitor--daemon.c\n> > > > > index 4161dd8282..8e32b5ae5e 100644\n> > > > > --- a/builtin/fsmonitor--daemon.c\n> > > > > +++ b/builtin/fsmonitor--daemon.c\n> > > > > @@ -206,6 +206,9 @@ static enum fsmonitor_cookie_item_result with_lock__wait_for_cookie(\n> > > > >       close(fd);\n> > > > >       unlink(cookie_pathname.buf);\n> > > > >\n> > > > > +     /* The listener callback takes main_lock, so this must not block. */\n> > > > > +     fsm_listen__flush_async(state);\n> > > > > +\n> > > > >       /*\n> > > > >        * Wait for the listener thread to observe the cookie file.\n> > > > >        * Time out after a short interval so that the client\n> > > >\n> > > > Okay, so we've unlinked the cookie file and the next thing is that we're\n> > > > waiting for all events to have been processed. As said, it feels\n> > > > reasonable that we're flushing all events before we start waiting for\n> > > > them.\n> > > >\n> > > > What I find surprising though is that this is supposed to make a\n> > > > difference at all. The latency we pass to `FSEventStreamCreate()` is\n> > > > 1 millisecond, and we wait up to 1 second for the cookie event. I would\n> > > > have expected that batching events for 1 milliseconds should be totally\n> > > > fine when we're waiting for a full second anyway.\n> > > >\n> > > > So given that I cannot verify this at all and that I have no clue about\n> > > > the FSEvent interfaces... do you have any explanation why the flush\n> > > > seems to help regardless?\n> > > >\n> > > > I _think_ you're already hinting at this in the commit message, where\n> > > > you say that it's not guaranteed that the queue is drained in a single\n> > > > latency interval. Is there any documentation that tells us what the\n> > > > provided guarantees are?\n> > > >\n> > > > Other than that the code changes look sensible to me, thanks!\n> > > >\n> > > > Patrick\n> > >\n> > > The following was generated by my coding agent and fact checked and\n> > > edited by me mainly to address you in the second person.\n> > >\n> > [snip]\n> > >\n> > > Hope that's helpful.\n> >\n> > Sorry, but that's not quite helpful. The questions I'm asking are to\n> > verify whether you understand the consequences and subtleties around the\n> > code area that you're proposing to change. If I wanted to only learn\n> > about this myself then I could simply ask an agent myself, but that's\n> > not really the intent of a code review.\n> >\n> > So what I'm looking for is _your_ explanation, not the explanation of\n> > AI. Your explanation may of course be informed by AI. But if so it's\n> > your responsibility to double-check its assumptions, build your own\n> > model and then share your informed opinion with us.\n> >\n> > Right now I don't yet have the feeling that you understand why this\n> > fixes the underlying issue.\n>\n> Got it. I agree with you that the flush call feels unnecessary under\n> the interpretation that passing 1ms to FSEventStreamCreate is the\n> equivalent of asking it to flush every 1ms. Empirically, though,\n> that's not the case, as described in the commit message.\n>\n> There's more precedent for this technique (found by agent, sorry):\n> watchman fixed a similar issue here:\n> https://github.com/facebook/watchman/commit/d1795de4ecab33672a89802318fe6f0122462194\n> and the documented it here:\n> https://github.com/facebook/watchman/commit/2f80886991ce81585ac0679c2b019fa0e4d9e9dd\n>\n> I agree this is unsatisfying.\n>\n> Does that help?\n\nI did a bunch more digging and I'm withdrawing this patch. I haven't\nsucceeded in proving that this fixes the performance issue. I'll\nresend in case this changes.\n\nThanks all for pushing back!\nTamir\n"},{"id":"550486","messageId":"an2IT78KDS94JqUt@pks.im","threadId":"66046","inReplyTo":"CAJ-ks9kQR77vH-56eS9tT-iXEnih+Z7SPRMs1gD_wTyg_6gZ_w@mail.gmail.com","subject":"Re: [PATCH] fsmonitor: flush pending FSEvents before cookie wait","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-13T09:03:11Z","receivedAt":"2026-08-13T09:03:19Z","isPatch":true,"body":"On Wed, Aug 12, 2026 at 09:19:08PM -0400, Tamir Duberstein wrote:\n> On Tue, Aug 11, 2026 at 12:45 PM Tamir Duberstein <tamird@gmail.com> wrote:\n[snip]\n> > Got it. I agree with you that the flush call feels unnecessary under\n> > the interpretation that passing 1ms to FSEventStreamCreate is the\n> > equivalent of asking it to flush every 1ms. Empirically, though,\n> > that's not the case, as described in the commit message.\n> >\n> > There's more precedent for this technique (found by agent, sorry):\n> > watchman fixed a similar issue here:\n> > https://github.com/facebook/watchman/commit/d1795de4ecab33672a89802318fe6f0122462194\n> > and the documented it here:\n> > https://github.com/facebook/watchman/commit/2f80886991ce81585ac0679c2b019fa0e4d9e9dd\n> >\n> > I agree this is unsatisfying.\n> >\n> > Does that help?\n\nThose links definitely help to provide some more context, thanks!\n\n> I did a bunch more digging and I'm withdrawing this patch. I haven't\n> succeeded in proving that this fixes the performance issue. I'll\n> resend in case this changes.\n\nOne major difference I notice there is that your patch uses\n`FsEventStreamFlushAsync()`, whereas Watchman uses the `Sync()` variant.\nThat could help explain why it works for their use case, as the can now\nguarantee that the cookie was indeed processed once that call finishes.\nBut with our `Async()` variant that's a guarantee that we cannot uphold,\nand consequently we're essentially still racing with the timout.\n\nNow we could of course try to use the synchronous variant ourselves. But\nI'm a bit concerned that this may create new problems that we don't\nreally understand yet. Quite unfortunate indeed :/\n\nThanks!\n\nPatrick\n"}]}