{"thread":{"id":"66048","subject":"[PATCH] regexec: work around macOS TRE memory leak on invalid UTF-8","startedAt":"2026-07-22T05:39:46Z","lastAt":"2026-08-05T09:00:31Z","messageCount":5,"participants":["Chungmin Lee","Junio C Hamano","Patrick Steinhardt"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"548753","messageId":"20260722053127.37244-1-chungmin@chungminlee.com","threadId":"66048","inReplyTo":null,"subject":"[PATCH] regexec: work around macOS TRE memory leak on invalid UTF-8","fromName":"Chungmin Lee","fromEmail":"chungmin@chungminlee.com","sentAt":"2026-07-22T05:31:27Z","receivedAt":"2026-07-22T05:39:46Z","isPatch":true,"body":"On macOS the system regex engine (TRE) leaks the buffer it allocates for\na match (in tre_tnfa_run_parallel()) whenever regexec() encounters an\ninvalid multibyte sequence in a UTF-8 locale: it returns REG_ILLSEQ\nwithout freeing that buffer.  Because regexec_buf() is called once per\nline, grepping a file that mixes text with binary data (for example a PDF\nchecked into a repository) leaks a buffer per line.  The leaked buffer is\nsized in proportion to the line length and the pattern's automaton, so a\nlarge case-insensitive alternation over long binary lines can leak\ngigabytes in a single command; this has been observed to exhaust memory\nand trigger a kernel watchdog panic on the machine running \"git grep\".\n\nA line that contains an invalid multibyte sequence can still contain\nreal matches: when a match lies on either side of the invalid bytes, TRE\nfinds it, returns REG_OK, and frees the buffer normally.  So the fix\ncannot simply reject or truncate such a line -- that would silently drop\nmatches the platform engine itself would report.  Nor can it search the\nwhole line first and fall back to segmentation only on REG_ILLSEQ: by the\ntime regexec() returns REG_ILLSEQ it has already leaked, so the fallback\ncannot prevent it.\n\nAdd a Darwin-only regexec_buf() (compat/regexec.c) that never hands the\nplatform matcher an invalid byte.  It walks the buffer, splits it at each\ninvalid sequence, and searches each maximal run of valid text between\nthem, returning the first match.  Each segment is searched with\nREG_STARTEND (whose offsets are relative to the buffer, so no translation\nis needed) and with REG_NOTBOL / REG_NOTEOL suppressed only when the\nsegment does not reach the true start or end of the buffer, so \"^\" and\n\"$\" keep matching exactly where they would on the whole line.  The split\npoints are chosen with mbrtowc(), so the buffer is cut at exactly the\nbytes the platform's own decoder -- and therefore TRE, which decodes the\nsame way -- rejects.\n\nThe split runs only where a byte can be invalid.  In a single-byte\nlocale (MB_CUR_MAX == 1) nothing is invalid, so the whole buffer is\nsearched as before and behavior is unchanged.  The leak is not specific\nto UTF-8 -- any multibyte locale drives TRE through the same REG_ILLSEQ\npath -- so the guard keys on MB_CUR_MAX rather than the codeset name,\nwhich is also simpler.  MB_CUR_MAX reflects the LC_CTYPE the C library\ninstalled, the same locale mbrtowc() decodes against.\n\nregexec_buf() backs every regex caller, so this affects \"git grep\",\ndiffcore-pickaxe (-G/-S --pickaxe-regex) and diff word splitting alike.\nThe fix does not reproduce the platform's exact output on lines with\ninvalid bytes -- doing so would mean deliberately hiding matches that are\nreally there.  No match the platform engine reported before is lost, and\nevery match reported now is a genuine one: each segment is searched as\nvalid text at its real offset, so nothing spurious is introduced.\nUnanchored matches -- which the engine already found, even next to\ninvalid bytes -- are unchanged.  The only additional matches are\nzero-width ones at a true line boundary next to invalid bytes: a \"^\",\n\"$\", or word boundary that holds there now matches, where the engine used\nto hit the invalid byte, return REG_ILLSEQ, and report nothing even\nthough the match is genuine.\n\nThe workaround is compiled in only when the platform's own regex engine\nis used, across the Makefile, Meson and CMake builds: it is skipped\nwhenever the bundled engine (which has no REG_ILLSEQ and does not leak)\nis selected by NO_REGEX -- as it is under AddressSanitizer in the\nMakefile and Meson builds.  compat/regexec.c guards its body on the same\nmacro, so a build that does not enable the workaround compiles it to\nnothing rather than colliding with the inline regexec_buf().  The\nworkaround avoids the leaking path rather than depending on it, so it\nstays correct if the platform is fixed and can be dropped once a macOS\nwithout the leak is the minimum supported baseline.\n\nA regression test cannot bound a process's memory portably -- setrlimit\nRLIMIT_AS is not enforced on macOS -- so t7810 checks the observable\nbehavior instead: valid text matches on both sides of an invalid byte\nand in a run between two invalid sequences, and \"^\" and \"$\" do not match\nacross an invalid byte in mid-line.  A further check, guarded by the\nMACOS prerequisite because only macOS runs this code, confirms that \"^\"\nand \"$\" match at the true start and end of such a line.  (A line that is\nentirely invalid bytes still does not match an unanchored empty-width\npattern such as \"^\", as before.)\n\nSigned-off-by: Chungmin Lee <chungmin@chungminlee.com>\n---\nThis came out of a real incident: \"git grep -i\" over a repository that\ncontains PDFs exhausted memory on an otherwise idle Mac mini and took the\nmachine down with a kernel watchdog panic (\"no checkins from watchdogd\").\nThe leak is in the system regex engine, not in git, but git is what\ndrives it into the leaking path, once per line.\n\nWhy this belongs in regexec_buf():\n\n  - regexec_buf() already exists as the place to wrap platform regexec()\n    quirks.  It was added in v2.11.0 (backported to v2.10.1) for exactly\n    this kind of reason -- \"a regexec_buf() helper that takes a <ptr,len>\n    pair with REG_STARTEND extension\" (RelNotes/2.11.0).\n\n  - git already treats REG_ILLSEQ on invalid UTF-8 as a cross-platform\n    reality -- \"As FreeBSD is not the only platform whose regexp library\n    reports a REG_ILLSEQ error when fed invalid UTF-8...\"\n    (RelNotes/2.28.0).  This patch keeps returning the matches around the\n    invalid bytes, and just stops leaking on the way there.\n\nOn the behavior change: the fix does not reproduce the platform engine's\nexact output on a line with invalid bytes, because that output is \"leak,\nthen report nothing\".  It never drops a match the engine reported before\nand never invents a spurious one; the only additional matches are\nzero-width assertions at a true line boundary next to the invalid bytes\n(the log describes the exact envelope).\n\nReproducing (macOS, UTF-8 locale):\n\n    perl -e 'print \"\\377\" x 16, \"\\n\" for 1..300000' >binfile\n    git init -q r && mv binfile r && git -C r add binfile &&\n        git -C r commit -qm x\n    P='aa|bb|cc|dd|ee|ff|gg|hh|ii|jj|kk|ll|mm|nn|oo|pp|qq|rr|ss|tt'\n    /usr/bin/time -l git -C r grep -i -E \"$P\" >/dev/null\n\n  Max RSS on this machine (grep is multithreaded, so the stock figure\n  varies run to run):\n    stock  git 2.50.1 (Apple):  ~1-2 GiB\n    patched:                    ~13 MiB\n  The gap grows without bound with the number of invalid-byte lines and\n  the size of the pattern's automaton; the original incident is believed\n  to have reached tens of GiB before the watchdog panic.\n\nI also have a standalone, self-verifying reproducer that measures the\nheap directly (malloc_zone_statistics(), no external tools); it reports\n~3072 bytes leaked per REG_ILLSEQ call, allocated in\ntre_tnfa_run_parallel() (confirmed with MallocStackLogging + leaks(1)).\nI can post it, and I have prepared it to file with Apple as well.\n\nEnvironment:\n    macOS 26.5 (build 25F71), Darwin 25.5.0 arm64, Apple M4\n    Apple clang 21.0.0; /usr/bin/git 2.50.1 (Apple Git-155)\n\nTested: t7810 passes (268/268), including the invalid-byte cases added\nhere.  (t7812 exits clean but its cases are all skipped on this machine\nfor lack of a suitable GETTEXT_LOCALE; it exercises the system-regex path\non platforms that have one.)  Built with the platform regex engine on\nmacOS; NO_REGEX builds, and Makefile/Meson ASAN builds, use the bundled\nengine and skip the workaround.\n\n Makefile                            |   4 ++\n compat/regexec.c                    | 108 ++++++++++++++++++++++++++++\n config.mak.uname                    |   1 +\n contrib/buildsystems/CMakeLists.txt |   5 ++\n git-compat-util.h                   |   5 ++\n meson.build                         |   7 ++\n t/t7810-grep.sh                     |  28 ++++++++\n 7 files changed, 158 insertions(+)\n create mode 100644 compat/regexec.c\n\ndiff --git a/Makefile b/Makefile\nindex 1cec251f4..b568dde52 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -2264,6 +2264,10 @@ ifdef USE_ENHANCED_BASIC_REGULAR_EXPRESSIONS\n \tCOMPAT_CFLAGS += -DUSE_ENHANCED_BASIC_REGULAR_EXPRESSIONS\n \tCOMPAT_OBJS += compat/regcomp_enhanced.o\n endif\n+ifdef DARWIN_TRE_REGEXEC_LEAK_WORKAROUND\n+\tCOMPAT_OBJS += compat/regexec.o\n+\tBASIC_CFLAGS += -DREGEXEC_MAY_LEAK_ON_ILLSEQ\n+endif\n endif\n ifdef NATIVE_CRLF\n \tBASIC_CFLAGS += -DNATIVE_CRLF\ndiff --git a/compat/regexec.c b/compat/regexec.c\nnew file mode 100644\nindex 000000000..0677162a8\n--- /dev/null\n+++ b/compat/regexec.c\n@@ -0,0 +1,108 @@\n+#include \"git-compat-util.h\"\n+\n+#ifdef REGEXEC_MAY_LEAK_ON_ILLSEQ\n+\n+#include <wchar.h>\n+\n+/*\n+ * macOS's libc regex engine (TRE) leaks the buffer it allocates for a\n+ * match whenever regexec() encounters an invalid multibyte sequence in\n+ * a multibyte locale: it returns REG_ILLSEQ without freeing that buffer.\n+ * A single \"git grep\" over a file with binary data can call regexec()\n+ * once per line and leak gigabytes, which has been observed to exhaust\n+ * memory and trigger a kernel watchdog panic.\n+ *\n+ * The leak happens inside regexec() before it returns, so reacting to\n+ * REG_ILLSEQ cannot avoid it: the invalid bytes must never reach the\n+ * matcher.  Split the buffer at each invalid sequence and search the\n+ * surrounding runs of valid text separately.  A match on either side of\n+ * the invalid bytes is still found (the same result the matcher gives on\n+ * valid input), but the leaking REG_ILLSEQ path is never reached.\n+ *\n+ * Use mbrtowc() to decide where to split, so that we split at exactly the\n+ * bytes the platform's own decoder -- and thus the regex engine, which\n+ * decodes the same way -- rejects.  A hand-rolled validator would\n+ * have to guess that boundary; being too lenient reintroduces the leak.\n+ */\n+\n+/*\n+ * Search buf[start, end) for a match.  REG_STARTEND reports offsets\n+ * relative to buf, so a hit needs no translation.  ^ may only match at\n+ * the real start of the buffer and $ only at its real end, so suppress\n+ * them when this segment does not reach those boundaries.\n+ */\n+static int regexec_segment(const regex_t *preg, const char *buf,\n+\t\t\t   size_t start, size_t end, size_t size,\n+\t\t\t   size_t nmatch, regmatch_t pmatch[], int eflags)\n+{\n+\teflags |= REG_STARTEND;\n+\tif (start > 0)\n+\t\teflags |= REG_NOTBOL;\n+\tif (end < size)\n+\t\teflags |= REG_NOTEOL;\n+\tpmatch[0].rm_so = start;\n+\tpmatch[0].rm_eo = end;\n+\treturn regexec(preg, buf, nmatch, pmatch, eflags);\n+}\n+\n+int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n+\t\tsize_t nmatch, regmatch_t pmatch[], int eflags)\n+{\n+\tsize_t seg_start = 0, i = 0;\n+\tmbstate_t mbs;\n+\n+\tassert(nmatch > 0 && pmatch);\n+\n+\t/*\n+\t * Only a multibyte locale drives TRE through the leaking multibyte\n+\t * path.  In a single-byte locale (MB_CUR_MAX == 1) no byte is\n+\t * invalid, so search the whole buffer as before.  MB_CUR_MAX\n+\t * reflects the current LC_CTYPE, the same locale mbrtowc() below\n+\t * decodes against.\n+\t */\n+\tif (MB_CUR_MAX == 1) {\n+\t\tpmatch[0].rm_so = 0;\n+\t\tpmatch[0].rm_eo = size;\n+\t\treturn regexec(preg, buf, nmatch, pmatch, eflags | REG_STARTEND);\n+\t}\n+\n+\tmemset(&mbs, 0, sizeof(mbs));\n+\twhile (i < size) {\n+\t\tunsigned char c = (unsigned char)buf[i];\n+\t\tsize_t n;\n+\n+\t\tif (c < 0x80) {\t\t/* ASCII fast path */\n+\t\t\ti++;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tn = mbrtowc(NULL, buf + i, size - i, &mbs);\n+\t\tif (!n)\t\t\t/* embedded NUL decodes to one byte */\n+\t\t\tn = 1;\n+\t\tif (n != (size_t)-1 && n != (size_t)-2) {\n+\t\t\ti += n;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\t/* buf[i] begins an invalid sequence; search the run before it */\n+\t\tif (i > seg_start) {\n+\t\t\tint ret = regexec_segment(preg, buf, seg_start, i, size,\n+\t\t\t\t\t\t  nmatch, pmatch, eflags);\n+\t\t\tif (ret != REG_NOMATCH)\n+\t\t\t\treturn ret;\n+\t\t}\n+\t\ti++;\t\t\t/* skip the invalid byte and resync */\n+\t\tseg_start = i;\n+\t\tmemset(&mbs, 0, sizeof(mbs));\n+\t}\n+\n+\t/*\n+\t * Search the final run.  Do this even when it is empty (a line that\n+\t * ends in invalid bytes, or an empty buffer) so that \"$\" and\n+\t * empty-matching patterns still match at the true end of the buffer.\n+\t */\n+\treturn regexec_segment(preg, buf, seg_start, size, size,\n+\t\t\t       nmatch, pmatch, eflags);\n+}\n+\n+#endif /* REGEXEC_MAY_LEAK_ON_ILLSEQ */\ndiff --git a/config.mak.uname b/config.mak.uname\nindex 9ebd24037..2402a2449 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -154,6 +154,7 @@ ifeq ($(uname_S),Darwin)\n \tHAVE_DEV_TTY = YesPlease\n \tCOMPAT_OBJS += compat/precompose_utf8.o\n \tBASIC_CFLAGS += -DPRECOMPOSE_UNICODE\n+\tDARWIN_TRE_REGEXEC_LEAK_WORKAROUND = YesPlease\n \tBASIC_CFLAGS += -DPROTECT_HFS_DEFAULT=1\n \tHAVE_BSD_SYSCTL = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex a57c4b464..5cfb48f78 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -519,6 +519,11 @@ if(NOT HAVE_REGEX)\n \tinclude_directories(${CMAKE_SOURCE_DIR}/compat/regex)\n \tlist(APPEND compat_SOURCES compat/regex/regex.c )\n \tadd_compile_definitions(NO_REGEX NO_MBSUPPORT GAWK)\n+elseif(APPLE)\n+\t# macOS's system regex engine (TRE) leaks memory when regexec()\n+\t# hits invalid UTF-8; work around it via compat/regexec.c.\n+\tlist(APPEND compat_SOURCES compat/regexec.c)\n+\tadd_compile_definitions(REGEXEC_MAY_LEAK_ON_ILLSEQ)\n endif()\n \n \ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 880977640..3861c9353 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -992,6 +992,10 @@ static inline int strtol_i(char const *s, int base, int *result)\n #error \"Git requires REG_STARTEND support. Compile with NO_REGEX=NeedsStartEnd\"\n #endif\n \n+#ifdef REGEXEC_MAY_LEAK_ON_ILLSEQ\n+int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n+\t\tsize_t nmatch, regmatch_t pmatch[], int eflags);\n+#else\n static inline int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n \t\t\t      size_t nmatch, regmatch_t pmatch[], int eflags)\n {\n@@ -1000,6 +1004,7 @@ static inline int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n \tpmatch[0].rm_eo = size;\n \treturn regexec(preg, buf, nmatch, pmatch, eflags | REG_STARTEND);\n }\n+#endif\n \n #ifdef USE_ENHANCED_BASIC_REGULAR_EXPRESSIONS\n int git_regcomp(regex_t *preg, const char *pattern, int cflags);\ndiff --git a/meson.build b/meson.build\nindex 3247697f7..2ce37b607 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -1387,6 +1387,13 @@ if not get_option('b_sanitize').contains('address') and get_option('regex').allo\n     libgit_c_args += '-DUSE_ENHANCED_BASIC_REGULAR_EXPRESSIONS'\n     compat_sources += 'compat/regcomp_enhanced.c'\n   endif\n+\n+  # macOS's system regex engine (TRE) leaks memory when regexec()\n+  # hits invalid UTF-8; work around it via compat/regexec.c.\n+  if host_machine.system() == 'darwin'\n+    libgit_c_args += '-DREGEXEC_MAY_LEAK_ON_ILLSEQ'\n+    compat_sources += 'compat/regexec.c'\n+  endif\n elif not get_option('regex').enabled()\n   libgit_c_args += [\n     '-DNO_REGEX',\ndiff --git a/t/t7810-grep.sh b/t/t7810-grep.sh\nindex d61c4a4d7..b325a2e11 100755\n--- a/t/t7810-grep.sh\n+++ b/t/t7810-grep.sh\n@@ -89,6 +89,8 @@ test_expect_success setup '\n \tfunction dummy() {}\n \tEOF\n \tprintf \"\\200\\nASCII\\n\" >invalid-utf8 &&\n+\tprintf \"before\\346world\\n\" >invalid-utf8-embedded &&\n+\tprintf \"a\\346b\\347c\\n\" >invalid-utf8-multi &&\n \tif test_have_prereq FUNNYNAMES\n \tthen\n \t\techo unusual >\"\\\"unusual\\\" pathname\" &&\n@@ -595,6 +597,32 @@ test_expect_success MB_REGEX 'grep two chars in single-char multibyte file' '\n \tLC_ALL=en_US.UTF-8 test_expect_code 1 git grep \"..\" reverse-question-mark\n '\n \n+test_expect_success MB_REGEX 'grep matches valid text on both sides of invalid UTF-8' '\n+\tLC_ALL=en_US.UTF-8 git grep -h before invalid-utf8-embedded >actual &&\n+\ttest_cmp invalid-utf8-embedded actual &&\n+\tLC_ALL=en_US.UTF-8 git grep -h world invalid-utf8-embedded >actual &&\n+\ttest_cmp invalid-utf8-embedded actual\n+'\n+\n+test_expect_success MB_REGEX 'grep matches a run between two invalid sequences' '\n+\tLC_ALL=en_US.UTF-8 git grep -h b invalid-utf8-multi >actual &&\n+\ttest_cmp invalid-utf8-multi actual\n+'\n+\n+test_expect_success MB_REGEX 'grep does not anchor ^ or $ inside an invalid-byte line' '\n+\ttest_expect_code 1 env LC_ALL=en_US.UTF-8 \\\n+\t\tgit grep -h \"^world\" invalid-utf8-embedded &&\n+\ttest_expect_code 1 env LC_ALL=en_US.UTF-8 \\\n+\t\tgit grep -h \"before\\$\" invalid-utf8-embedded\n+'\n+\n+test_expect_success MACOS,MB_REGEX 'grep anchors ^ and $ at true line ends past invalid UTF-8' '\n+\tLC_ALL=en_US.UTF-8 git grep -h \"^before\" invalid-utf8-embedded >actual &&\n+\ttest_cmp invalid-utf8-embedded actual &&\n+\tLC_ALL=en_US.UTF-8 git grep -h \"world\\$\" invalid-utf8-embedded >actual &&\n+\ttest_cmp invalid-utf8-embedded actual\n+'\n+\n cat >expected <<EOF\n file\n EOF\n\nbase-commit: e9019fcafe0040228b8631c30f97ae1adb61bcdc\n-- \n2.55.0\n\n"},{"id":"548859","messageId":"xmqqpl0d9fyq.fsf@gitster.g","threadId":"66048","inReplyTo":"20260722053127.37244-1-chungmin@chungminlee.com","subject":"Re: [PATCH] regexec: work around macOS TRE memory leak on invalid UTF-8","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-24T04:19:57Z","receivedAt":"2026-07-24T04:20:00Z","isPatch":true,"body":"Chungmin Lee <chungmin@chungminlee.com> writes:\n\n> This came out of a real incident: \"git grep -i\" over a repository that\n> contains PDFs exhausted memory on an otherwise idle Mac mini and took the\n> machine down with a kernel watchdog panic (\"no checkins from watchdogd\").\n> The leak is in the system regex engine, not in git, but git is what\n> drives it into the leaking path, once per line.\n>\n> Why this belongs in regexec_buf():\n\nWhy does this belong to Git, not macOS, in the first place?\n\n> diff --git a/compat/regexec.c b/compat/regexec.c\n> new file mode 100644\n> index 000000000..0677162a8\n> --- /dev/null\n> +++ b/compat/regexec.c\n> @@ -0,0 +1,108 @@\n> +#include \"git-compat-util.h\"\n> +\n> +#ifdef REGEXEC_MAY_LEAK_ON_ILLSEQ\n> +\n> +#include <wchar.h>\n> +\n> +/*\n> + * macOS's libc regex engine (TRE) leaks the buffer it allocates for a\n> + * match whenever regexec() encounters an invalid multibyte sequence in\n> + * a multibyte locale: it returns REG_ILLSEQ without freeing that buffer.\n> + * A single \"git grep\" over a file with binary data can call regexec()\n> + * once per line and leak gigabytes, which has been observed to exhaust\n> + * memory and trigger a kernel watchdog panic.\n> + *\n> + * The leak happens inside regexec() before it returns, so reacting to\n> + * REG_ILLSEQ cannot avoid it: the invalid bytes must never reach the\n> + * matcher.  Split the buffer at each invalid sequence and search the\n> + * surrounding runs of valid text separately.  A match on either side of\n> + * the invalid bytes is still found (the same result the matcher gives on\n> + * valid input), but the leaking REG_ILLSEQ path is never reached.\n> + *\n> + * Use mbrtowc() to decide where to split, so that we split at exactly the\n> + * bytes the platform's own decoder -- and thus the regex engine, which\n> + * decodes the same way -- rejects.  A hand-rolled validator would\n> + * have to guess that boundary; being too lenient reintroduces the leak.\n> + */\n\nThis clearly seems to be a workaround for a platform bug.  Do we\nknow how long we will need to keep it?\n\n> +/*\n> + * Search buf[start, end) for a match.  REG_STARTEND reports offsets\n> + * relative to buf, so a hit needs no translation.  ^ may only match at\n> + * the real start of the buffer and $ only at its real end, so suppress\n> + * them when this segment does not reach those boundaries.\n> + */\n> +static int regexec_segment(const regex_t *preg, const char *buf,\n> +\t\t\t   size_t start, size_t end, size_t size,\n> +\t\t\t   size_t nmatch, regmatch_t pmatch[], int eflags)\n> +{\n> +\teflags |= REG_STARTEND;\n> +\tif (start > 0)\n> +\t\teflags |= REG_NOTBOL;\n> +\tif (end < size)\n> +\t\teflags |= REG_NOTEOL;\n> +\tpmatch[0].rm_so = start;\n> +\tpmatch[0].rm_eo = end;\n> +\treturn regexec(preg, buf, nmatch, pmatch, eflags);\n> +}\n> +\n> +int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n> +\t\tsize_t nmatch, regmatch_t pmatch[], int eflags)\n> +{\n> +\tsize_t seg_start = 0, i = 0;\n> +\tmbstate_t mbs;\n> +\n> +\tassert(nmatch > 0 && pmatch);\n> +\n> +\t/*\n> +\t * Only a multibyte locale drives TRE through the leaking multibyte\n> +\t * path.  In a single-byte locale (MB_CUR_MAX == 1) no byte is\n> +\t * invalid, so search the whole buffer as before.  MB_CUR_MAX\n> +\t * reflects the current LC_CTYPE, the same locale mbrtowc() below\n> +\t * decodes against.\n> +\t */\n> +\tif (MB_CUR_MAX == 1) {\n> +\t\tpmatch[0].rm_so = 0;\n> +\t\tpmatch[0].rm_eo = size;\n> +\t\treturn regexec(preg, buf, nmatch, pmatch, eflags | REG_STARTEND);\n> +\t}\n> +\n> +\tmemset(&mbs, 0, sizeof(mbs));\n> +\twhile (i < size) {\n> +\t\tunsigned char c = (unsigned char)buf[i];\n> +\t\tsize_t n;\n> +\n> +\t\tif (c < 0x80) {\t\t/* ASCII fast path */\n> +\t\t\ti++;\n> +\t\t\tcontinue;\n> +\t\t}\n> +\n> +\t\tn = mbrtowc(NULL, buf + i, size - i, &mbs);\n> +\t\tif (!n)\t\t\t/* embedded NUL decodes to one byte */\n> +\t\t\tn = 1;\n> +\t\tif (n != (size_t)-1 && n != (size_t)-2) {\n> +\t\t\ti += n;\n> +\t\t\tcontinue;\n> +\t\t}\n\nOK.  I wonder if we want to document what -1 and -2 signify (in\nother words, why we stop only when the call returns one of these\ntwo values), or is it too obvious for users of mbrtowc()?\n\nIn any case, if control reaches here, we saw either an invalid\nsequence (-1) or not enough bytes to complete a whole multi-byte\ncharacter (-2), i.e., the case where regexec() would have trouble\nmatching starting at offset 'i'.  The bytes before that position\nmake an OK substring.\n\n> +\t\t/* buf[i] begins an invalid sequence; search the run before it */\n> +\t\tif (i > seg_start) {\n> +\t\t\tint ret = regexec_segment(preg, buf, seg_start, i, size,\n> +\t\t\t\t\t\t  nmatch, pmatch, eflags);\n> +\t\t\tif (ret != REG_NOMATCH)\n> +\t\t\t\treturn ret;\n> +\t\t}\n\nNaturally, this \"check the OK prefix string\" approach makes readers\nwonder what happens when the pattern is \"right anchored$\" and the OK\nprefix would match if the string truly ended at 'i' (or, if this is a\nsecond or subsequent segment, the pattern is \"^left anchored\", and\nthe segment would match if the string started at 'seg_start').  The\nuse of 'REG_STARTEND' in regexec_segment() above, combined with\n'REG_NOTBOL'/'REG_NOTEOL', is a clever way to work around it cleanly.\n\n> diff --git a/git-compat-util.h b/git-compat-util.h\n> index 880977640..3861c9353 100644\n> --- a/git-compat-util.h\n> +++ b/git-compat-util.h\n> @@ -992,6 +992,10 @@ static inline int strtol_i(char const *s, int base, int *result)\n>  #error \"Git requires REG_STARTEND support. Compile with NO_REGEX=NeedsStartEnd\"\n>  #endif\n>  \n> +#ifdef REGEXEC_MAY_LEAK_ON_ILLSEQ\n\nHmph, how many different symbols do we need to deal with this?  The\nMakefile has DARWIN_TRE_REGEXEC_LEAK_WORKAROUND and CPP macro is\nREGEXEC_MAY_LEAK_ON_ILLSEQ?\n\n> +int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n> +\t\tsize_t nmatch, regmatch_t pmatch[], int eflags);\n> +#else\n>  static inline int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n>  \t\t\t      size_t nmatch, regmatch_t pmatch[], int eflags)\n>  {\n> @@ -1000,6 +1004,7 @@ static inline int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n>  \tpmatch[0].rm_eo = size;\n>  \treturn regexec(preg, buf, nmatch, pmatch, eflags | REG_STARTEND);\n>  }\n> +#endif\n\nIt is a bit awkward that the next platform needing its own\nimplementation of regexec_buf() to work around a different platform\nbug would have to do:\n\n\t#if defined(REGEXEC_MAY_LEAK_ON_ILLSEQ) || defined(SOME_OTHER_PLATFORM_BUG)\n\tint regexec_buf(.....);\n\t#else\n\tstatic inline int regexec_buf(.....)\n\t... the current definition comes here ...\n\t#endif\n\nI thought it was more common to:\n\n * Have each platform with such a need define an override in its own\n   platform header file:\n\n    int darwin_regexec_buf(.....);\n    #define regexec_buf darwin_regexec_buf\n\n * Have a header file like 'git-compat-util.h' include such a header\n   file (conditionally on relevant platforms, of course); and\n\n * Have the common header file do this:\n\n        #ifndef regexec_buf\n        static inline int regexec_buf(.....)\n        ... the current definition comes here ...\n        #endif\n\nRight now, macOS is the only platform that needs an override, so the\nresult would be about the same amount of code.  However, in the long\nrun, this structure may give us a better organization, no?\n\n> diff --git a/t/t7810-grep.sh b/t/t7810-grep.sh\n> ...\n> +test_expect_success MACOS,MB_REGEX 'grep anchors ^ and $ at true line ends past invalid UTF-8' '\n\nDo we need to allow this test to fail on non macOS hosts?  Why?\n\n> +\tLC_ALL=en_US.UTF-8 git grep -h \"^before\" invalid-utf8-embedded >actual &&\n> +\ttest_cmp invalid-utf8-embedded actual &&\n> +\tLC_ALL=en_US.UTF-8 git grep -h \"world\\$\" invalid-utf8-embedded >actual &&\n> +\ttest_cmp invalid-utf8-embedded actual\n> +'\n> +\n\nThanks.\n"},{"id":"549112","messageId":"20260728052538.12429-1-chungmin@chungminlee.com","threadId":"66048","inReplyTo":"20260722053127.37244-1-chungmin@chungminlee.com","subject":"[PATCH v2] regexec: work around macOS TRE leak on invalid UTF-8","fromName":"Chungmin Lee","fromEmail":"chungmin@chungminlee.com","sentAt":"2026-07-28T05:25:38Z","receivedAt":"2026-07-28T05:26:05Z","isPatch":true,"body":"On macOS, the system regex engine leaks an internal buffer when\nregexec() encounters an invalid multibyte sequence in a UTF-8 locale.\nThe line-by-line path can call regexec_buf() for each pattern on every\nline, so \"git grep\" can leak repeatedly on a file containing invalid\nUTF-8.  The total leak grows with the number of calls, and the per-call\nallocation grows with the pattern's automaton.  In one case, grepping a\nrepository containing PDFs exhausted memory and caused the machine to\nrestart.\n\nce025ae4f61e (grep: disable lookahead on error, 2024-10-20) made \"git\ngrep\" fall back to line-by-line matching when regexec() reports an error\non invalid UTF-8.  That fallback cannot prevent this leak: the allocation\nhas already leaked when regexec() returns REG_ILLSEQ.\n\nAvoid the leaking path by providing a Darwin-specific regexec_buf().\nWalk the input with mbrtowc(), split it at bytes that cannot form a\ncomplete multibyte character, and search each valid segment separately.\nThis preserves matches in valid text on either side of an invalid byte.\n\nSearch each segment with REG_STARTEND so match offsets remain relative to\nthe original buffer.  Set REG_NOTBOL and REG_NOTEOL for internal segment\nboundaries so \"^\" and \"$\" do not match there.  Keep the flags clear at\nthe true beginning and end of the buffer.\n\nUse the normal regexec_buf() path in single-byte locales, where no byte\ncan form an invalid multibyte sequence.  Use the bundled regex\nimplementation unchanged when NO_REGEX is enabled.\n\nDeclare the Darwin override in compat/darwin.h and map regexec_buf() to\ndarwin_regexec_buf().  This follows the platform override pattern used\nby the other compatibility headers and leaves the common inline\nimplementation as the default.\n\nThere is no reliable way to detect a future macOS version in which the\nsystem regex implementation has been fixed.  Even after a fix, Git will\nneed the workaround while it supports affected macOS releases, so treat\nit as an indefinite compatibility workaround.\n\nAdd tests for matches before, after, and between invalid bytes, including\nan offset check after an invalid byte.  Also check incomplete trailing\ninput and anchors at true and internal line boundaries.\n\nSigned-off-by: Chungmin Lee <chungmin@chungminlee.com>\n---\nChanges since v1:\n\n  - Cite ce025ae4f61e, which handles the same macOS regexec() error.\n  - Treat the workaround as indefinite instead of assuming a known\n    removal point.\n  - Document the -1 and -2 returns from mbrtowc().\n  - Use one DARWIN_REGEXEC name in all build systems.\n  - Move the implementation under compat/darwin/ and use a platform\n    header to override regexec_buf().\n  - Add coverage for offsets, incomplete sequences, and internal\n    boundaries.\n  - Use POSIX regex patterns in the tests so they exercise regexec_buf()\n    rather than the fixed-string optimization.\n  - Keep positive invalid-UTF-8 matching tests macOS-only because system\n    regex implementations differ in how they handle invalid multibyte\n    input.  Run the no-false-anchor test on every MB_REGEX platform.\n\nReproduction on macOS in a UTF-8 locale:\n\n    perl -e 'print \"\\377\" x 16, \"\\n\" for 1..300000' >binfile\n    git init -q r &&\n    mv binfile r &&\n    git -C r add binfile &&\n    git -C r commit -qm x\n    P='aa|bb|cc|dd|ee|ff|gg|hh|ii|jj|kk|ll|mm|nn|oo|pp|qq|rr|ss|tt'\n    /usr/bin/time -l git -C r grep -i -E \"$P\" >/dev/null\n\nOn the machine used to reproduce the problem, stock Apple Git 2.50.1\nreached roughly 1--2 GiB maximum RSS.  The patched build used roughly\n13 MiB.  The difference grows with the number of invalid-byte lines and\nthe size of the pattern.\n\nThe leak was also reproduced in a standalone test using\nmalloc_zone_statistics().  It reported about 3072 bytes leaked per\nREG_ILLSEQ call from tre_tnfa_run_parallel().\n\nTested with the native macOS regex engine and with NO_REGEX:\n\n    make\n    make test T=t7810-grep.sh\n    make clean\n    make NO_REGEX=YesPlease\n    make test T=t7810-grep.sh NO_REGEX=YesPlease\n\n Makefile                            |  4 ++\n compat/darwin.h                     |  8 +++\n compat/darwin/regexec.c             | 91 +++++++++++++++++++++++++++++\n config.mak.uname                    |  1 +\n contrib/buildsystems/CMakeLists.txt |  3 +\n git-compat-util.h                   |  5 ++\n meson.build                         |  5 ++\n t/t7810-grep.sh                     | 37 ++++++++++++\n 8 files changed, 154 insertions(+)\n create mode 100644 compat/darwin.h\n create mode 100644 compat/darwin/regexec.c\n\ndiff --git a/Makefile b/Makefile\nindex 1cec251..81075c3 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -2264,6 +2264,10 @@ ifdef USE_ENHANCED_BASIC_REGULAR_EXPRESSIONS\n \tCOMPAT_CFLAGS += -DUSE_ENHANCED_BASIC_REGULAR_EXPRESSIONS\n \tCOMPAT_OBJS += compat/regcomp_enhanced.o\n endif\n+ifdef DARWIN_REGEXEC\n+\tCOMPAT_OBJS += compat/darwin/regexec.o\n+\tBASIC_CFLAGS += -DDARWIN_REGEXEC\n+endif\n endif\n ifdef NATIVE_CRLF\n \tBASIC_CFLAGS += -DNATIVE_CRLF\ndiff --git a/compat/darwin.h b/compat/darwin.h\nnew file mode 100644\nindex 0000000..6fbdc34\n--- /dev/null\n+++ b/compat/darwin.h\n@@ -0,0 +1,8 @@\n+#ifndef COMPAT_DARWIN_H\n+#define COMPAT_DARWIN_H\n+\n+int darwin_regexec_buf(const regex_t *preg, const char *buf, size_t size,\n+\t\t       size_t nmatch, regmatch_t pmatch[], int eflags);\n+#define regexec_buf darwin_regexec_buf\n+\n+#endif\ndiff --git a/compat/darwin/regexec.c b/compat/darwin/regexec.c\nnew file mode 100644\nindex 0000000..13fb7d5\n--- /dev/null\n+++ b/compat/darwin/regexec.c\n@@ -0,0 +1,91 @@\n+#include \"git-compat-util.h\"\n+\n+#include <wchar.h>\n+\n+/*\n+ * Darwin's TRE regex engine leaks an internal buffer when it encounters an\n+ * invalid multibyte sequence.  Since the leak has already happened when\n+ * regexec() reports REG_ILLSEQ, keep invalid bytes out of regexec() by\n+ * searching each valid segment separately.\n+ */\n+\n+/*\n+ * Search buf[start, end), where size is the full size of buf.  REG_STARTEND\n+ * keeps match offsets relative to buf.  Do not let an internal segment create\n+ * a false beginning or end of line.\n+ */\n+static int regexec_segment(const regex_t *preg, const char *buf,\n+\t\t\t   size_t size, size_t start, size_t end,\n+\t\t\t   size_t nmatch, regmatch_t pmatch[], int eflags)\n+{\n+\teflags |= REG_STARTEND;\n+\tif (start > 0)\n+\t\teflags |= REG_NOTBOL;\n+\tif (end < size)\n+\t\teflags |= REG_NOTEOL;\n+\tpmatch[0].rm_so = start;\n+\tpmatch[0].rm_eo = end;\n+\treturn regexec(preg, buf, nmatch, pmatch, eflags);\n+}\n+\n+int darwin_regexec_buf(const regex_t *preg, const char *buf, size_t size,\n+\t\t       size_t nmatch, regmatch_t pmatch[], int eflags)\n+{\n+\tsize_t seg_start = 0, i = 0;\n+\tmbstate_t mbs;\n+\n+\tassert(nmatch > 0 && pmatch);\n+\n+\t/*\n+\t * A single-byte locale cannot contain an invalid multibyte sequence,\n+\t * so use regexec() directly.\n+\t */\n+\tif (MB_CUR_MAX == 1) {\n+\t\tpmatch[0].rm_so = 0;\n+\t\tpmatch[0].rm_eo = size;\n+\t\treturn regexec(preg, buf, nmatch, pmatch, eflags | REG_STARTEND);\n+\t}\n+\n+\tmemset(&mbs, 0, sizeof(mbs));\n+\twhile (i < size) {\n+\t\tunsigned char c = (unsigned char)buf[i];\n+\t\tsize_t n;\n+\n+\t\tif (c < 0x80) {\n+\t\t\ti++;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\tn = mbrtowc(NULL, buf + i, size - i, &mbs);\n+\t\tif (!n)\n+\t\t\tn = 1;\n+\t\tif (n != (size_t)-1 && n != (size_t)-2) {\n+\t\t\ti += n;\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\t/*\n+\t\t * -1 denotes an encoding error; -2 denotes an incomplete\n+\t\t * trailing sequence.  In either case, buf[i] cannot begin a\n+\t\t * complete valid character within this buffer.  Search an\n+\t\t * empty initial segment to preserve zero-width matches at the\n+\t\t * true beginning.\n+\t\t */\n+\t\tif (i > seg_start || i == 0) {\n+\t\t\tint ret = regexec_segment(preg, buf, size, seg_start, i,\n+\t\t\t\t\t\t  nmatch, pmatch, eflags);\n+\t\t\tif (ret != REG_NOMATCH)\n+\t\t\t\treturn ret;\n+\t\t}\n+\t\ti++;\n+\t\tseg_start = i;\n+\t\tmemset(&mbs, 0, sizeof(mbs));\n+\t}\n+\n+\t/*\n+\t * Search the final segment even when it is empty, so an empty buffer\n+\t * or a buffer ending in invalid bytes still has its true end.\n+\t */\n+\treturn regexec_segment(preg, buf, size, seg_start, size,\n+\t\t\t       nmatch, pmatch, eflags);\n+}\ndiff --git a/config.mak.uname b/config.mak.uname\nindex 9ebd240..4660ff3 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -154,6 +154,7 @@ ifeq ($(uname_S),Darwin)\n \tHAVE_DEV_TTY = YesPlease\n \tCOMPAT_OBJS += compat/precompose_utf8.o\n \tBASIC_CFLAGS += -DPRECOMPOSE_UNICODE\n+\tDARWIN_REGEXEC = YesPlease\n \tBASIC_CFLAGS += -DPROTECT_HFS_DEFAULT=1\n \tHAVE_BSD_SYSCTL = YesPlease\n \tFREAD_READS_DIRECTORIES = UnfortunatelyYes\ndiff --git a/contrib/buildsystems/CMakeLists.txt b/contrib/buildsystems/CMakeLists.txt\nindex a57c4b4..83e8b71 100644\n--- a/contrib/buildsystems/CMakeLists.txt\n+++ b/contrib/buildsystems/CMakeLists.txt\n@@ -519,6 +519,9 @@ if(NOT HAVE_REGEX)\n \tinclude_directories(${CMAKE_SOURCE_DIR}/compat/regex)\n \tlist(APPEND compat_SOURCES compat/regex/regex.c )\n \tadd_compile_definitions(NO_REGEX NO_MBSUPPORT GAWK)\n+elseif(APPLE)\n+\tlist(APPEND compat_SOURCES compat/darwin/regexec.c)\n+\tadd_compile_definitions(DARWIN_REGEXEC)\n endif()\n \n \ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 8809776..96995c6 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -162,6 +162,9 @@ static inline int is_xplatform_dir_sep(int c)\n #include \"compat/win32/path-utils.h\"\n #include \"compat/msvc.h\"\n #endif\n+#ifdef DARWIN_REGEXEC\n+#include \"compat/darwin.h\"\n+#endif\n \n /* used on Mac OS X */\n #ifdef PRECOMPOSE_UNICODE\n@@ -992,6 +995,7 @@ static inline int strtol_i(char const *s, int base, int *result)\n #error \"Git requires REG_STARTEND support. Compile with NO_REGEX=NeedsStartEnd\"\n #endif\n \n+#ifndef regexec_buf\n static inline int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n \t\t\t      size_t nmatch, regmatch_t pmatch[], int eflags)\n {\n@@ -1000,6 +1004,7 @@ static inline int regexec_buf(const regex_t *preg, const char *buf, size_t size,\n \tpmatch[0].rm_eo = size;\n \treturn regexec(preg, buf, nmatch, pmatch, eflags | REG_STARTEND);\n }\n+#endif\n \n #ifdef USE_ENHANCED_BASIC_REGULAR_EXPRESSIONS\n int git_regcomp(regex_t *preg, const char *pattern, int cflags);\ndiff --git a/meson.build b/meson.build\nindex 3247697..53c4816 100644\n--- a/meson.build\n+++ b/meson.build\n@@ -1387,6 +1387,11 @@ if not get_option('b_sanitize').contains('address') and get_option('regex').allo\n     libgit_c_args += '-DUSE_ENHANCED_BASIC_REGULAR_EXPRESSIONS'\n     compat_sources += 'compat/regcomp_enhanced.c'\n   endif\n+\n+  if host_machine.system() == 'darwin'\n+    libgit_c_args += '-DDARWIN_REGEXEC'\n+    compat_sources += 'compat/darwin/regexec.c'\n+  endif\n elif not get_option('regex').enabled()\n   libgit_c_args += [\n     '-DNO_REGEX',\ndiff --git a/t/t7810-grep.sh b/t/t7810-grep.sh\nindex d61c4a4..149654e 100755\n--- a/t/t7810-grep.sh\n+++ b/t/t7810-grep.sh\n@@ -89,6 +89,10 @@ test_expect_success setup '\n \tfunction dummy() {}\n \tEOF\n \tprintf \"\\200\\nASCII\\n\" >invalid-utf8 &&\n+\tprintf \"before\\346world\\n\" >invalid-utf8-embedded &&\n+\tprintf \"a\\346b\\347c\\n\" >invalid-utf8-multi &&\n+\tprintf \"\\346world\\n\" >invalid-utf8-leading &&\n+\tprintf \"before\\346\\n\" >invalid-utf8-trailing &&\n \tif test_have_prereq FUNNYNAMES\n \tthen\n \t\techo unusual >\"\\\"unusual\\\" pathname\" &&\n@@ -595,6 +599,39 @@ test_expect_success MB_REGEX 'grep two chars in single-char multibyte file' '\n \tLC_ALL=en_US.UTF-8 test_expect_code 1 git grep \"..\" reverse-question-mark\n '\n \n+test_expect_success MACOS,MB_REGEX 'grep matches valid text on both sides of invalid UTF-8' '\n+\tLC_ALL=en_US.UTF-8 git grep -h \"befo[r]e\" invalid-utf8-embedded >actual &&\n+\ttest_cmp invalid-utf8-embedded actual &&\n+\tLC_ALL=en_US.UTF-8 git grep -h \"worl[d]\" invalid-utf8-embedded >actual &&\n+\ttest_cmp invalid-utf8-embedded actual &&\n+\tLC_ALL=en_US.UTF-8 git grep -h -o \"worl[d]\" invalid-utf8-embedded >actual &&\n+\techo world >expected &&\n+\ttest_cmp expected actual\n+'\n+\n+test_expect_success MACOS,MB_REGEX 'grep matches a run between two invalid sequences' '\n+\tLC_ALL=en_US.UTF-8 git grep -h \"[b]\" invalid-utf8-multi >actual &&\n+\ttest_cmp invalid-utf8-multi actual\n+'\n+\n+test_expect_success MB_REGEX 'grep does not anchor ^ or $ inside an invalid-byte line' '\n+\ttest_expect_code 1 env LC_ALL=en_US.UTF-8 \\\n+\t\tgit grep -h \"^world\" invalid-utf8-embedded &&\n+\ttest_expect_code 1 env LC_ALL=en_US.UTF-8 \\\n+\t\tgit grep -h \"before\\$\" invalid-utf8-embedded\n+'\n+\n+test_expect_success MACOS,MB_REGEX 'grep anchors ^ and $ at true line ends past invalid UTF-8' '\n+\tLC_ALL=en_US.UTF-8 git grep -h \"^before\" invalid-utf8-embedded >actual &&\n+\ttest_cmp invalid-utf8-embedded actual &&\n+\tLC_ALL=en_US.UTF-8 git grep -h \"world\\$\" invalid-utf8-embedded >actual &&\n+\ttest_cmp invalid-utf8-embedded actual &&\n+\tLC_ALL=en_US.UTF-8 git grep -h \"^\" invalid-utf8-leading >actual &&\n+\ttest_cmp invalid-utf8-leading actual &&\n+\tLC_ALL=en_US.UTF-8 git grep -h \"\\$\" invalid-utf8-trailing >actual &&\n+\ttest_cmp invalid-utf8-trailing actual\n+'\n+\n cat >expected <<EOF\n file\n EOF\n\nbase-commit: e9019fcafe0040228b8631c30f97ae1adb61bcdc\n-- \n2.55.0\n"},{"id":"549184","messageId":"xmqqse52bpa9.fsf@gitster.g","threadId":"66048","inReplyTo":"20260728052538.12429-1-chungmin@chungminlee.com","subject":"Re: [PATCH v2] regexec: work around macOS TRE leak on invalid UTF-8","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-29T00:41:34Z","receivedAt":"2026-07-29T00:41:37Z","isPatch":true,"body":"Chungmin Lee <chungmin@chungminlee.com> writes:\n\n> diff --git a/Makefile b/Makefile\n> index 1cec251..81075c3 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -2264,6 +2264,10 @@ ifdef USE_ENHANCED_BASIC_REGULAR_EXPRESSIONS\n>  \tCOMPAT_CFLAGS += -DUSE_ENHANCED_BASIC_REGULAR_EXPRESSIONS\n>  \tCOMPAT_OBJS += compat/regcomp_enhanced.o\n>  endif\n> +ifdef DARWIN_REGEXEC\n> +\tCOMPAT_OBJS += compat/darwin/regexec.o\n> +\tBASIC_CFLAGS += -DDARWIN_REGEXEC\n> +endif\n>  endif\n>  ifdef NATIVE_CRLF\n>  \tBASIC_CFLAGS += -DNATIVE_CRLF\n> diff --git a/compat/darwin.h b/compat/darwin.h\n> new file mode 100644\n> index 0000000..6fbdc34\n> --- /dev/null\n> +++ b/compat/darwin.h\n> @@ -0,0 +1,8 @@\n> +#ifndef COMPAT_DARWIN_H\n> +#define COMPAT_DARWIN_H\n> +\n> +int darwin_regexec_buf(const regex_t *preg, const char *buf, size_t size,\n> +\t\t       size_t nmatch, regmatch_t pmatch[], int eflags);\n> +#define regexec_buf darwin_regexec_buf\n> +\n> +#endif\n\n\nThis iteration looks much easier to grok, at least to me.  Two\nthings:\n\n * The name of the header, <compat/darwin.h>, sounds so nice and\n   central, that those who care a lot more about macOS than I do may\n   want to consolidate other support for the peculiarities macOS has\n   also into it.  I personally do not have a strong opinion.\n\n * We'd need a comment near the beginning of Makefile, like other\n   symbolis like NO_FINK and NO_APPLE_COMMMON_CRYPTO do, to tell the\n   users when to define this new symbol.\n\nThe latter I would feel strong enough, so here is a sample update in\na squashable form.  If you have reasons to send a new iteration, you\nare free to include it.  After waiting for comments from others for\na few days, if you still don't have reasons to send an update, you\ncan just tell me to squash the change on my end (if you agree with\nthe change, that is).\n\n\n\ndiff --git c/Makefile w/Makefile\nindex 81075c38a2..ed2868ce10 100644\n--- c/Makefile\n+++ w/Makefile\n@@ -110,6 +110,9 @@ include shared.mak\n # Define USE_HOMEBREW_LIBICONV to link against libiconv installed by\n # Homebrew, if present.\n #\n+# Define DARWIN_REGEXEC if regexec() in your platform regex library\n+# leaks when fed an invalid UTF-8 sequence.\n+#\n # Define NO_APPLE_COMMON_CRYPTO if you are building on Darwin/Mac OS X\n # and do not want to use Apple's CommonCrypto library.  This allows you\n # to provide your own OpenSSL library, for example from MacPorts.\n\n\n"},{"id":"549661","messageId":"anL7qL2-4h8ZlLcg@pks.im","threadId":"66048","inReplyTo":"20260728052538.12429-1-chungmin@chungminlee.com","subject":"Re: [PATCH v2] regexec: work around macOS TRE leak on invalid UTF-8","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-08-05T09:00:24Z","receivedAt":"2026-08-05T09:00:31Z","isPatch":true,"body":"On Mon, Jul 27, 2026 at 10:25:38PM -0700, Chungmin Lee wrote:\n> On macOS, the system regex engine leaks an internal buffer when\n> regexec() encounters an invalid multibyte sequence in a UTF-8 locale.\n> The line-by-line path can call regexec_buf() for each pattern on every\n> line, so \"git grep\" can leak repeatedly on a file containing invalid\n> UTF-8.  The total leak grows with the number of calls, and the per-call\n> allocation grows with the pattern's automaton.  In one case, grepping a\n> repository containing PDFs exhausted memory and caused the machine to\n> restart.\n> \n> ce025ae4f61e (grep: disable lookahead on error, 2024-10-20) made \"git\n> grep\" fall back to line-by-line matching when regexec() reports an error\n> on invalid UTF-8.  That fallback cannot prevent this leak: the allocation\n> has already leaked when regexec() returns REG_ILLSEQ.\n> \n> Avoid the leaking path by providing a Darwin-specific regexec_buf().\n> Walk the input with mbrtowc(), split it at bytes that cannot form a\n> complete multibyte character, and search each valid segment separately.\n> This preserves matches in valid text on either side of an invalid byte.\n\nHm. I feel like we're adding quite a lot of logic only to fix an\nupstream bug that we expect will be eventually fixed. At the same time\nwe already have a compatibility \"regexec\" implementation that I'd expect\ndoesn't have the bug. So would an alternative be to detect whether the\ngiven platform is susceptible to the bug and, if so, define NO_REGEX and\nthen use our own regex implementation? Or are there good reasons to not\ndo that?\n\nPatrick\n"}]}