{"thread":{"id":"65916","subject":"[PATCH] precompose_utf8: use a flex array for d_name","startedAt":"2026-07-03T02:36:10Z","lastAt":"2026-07-04T23:37:41Z","messageCount":6,"participants":["Ihar Hrachyshka","Torsten Bögershausen","Junio C Hamano","Patrick Steinhardt"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"547019","messageId":"20260703023554.36577-1-ihar.hrachyshka@gmail.com","threadId":"65916","inReplyTo":null,"subject":"[PATCH] precompose_utf8: use a flex array for d_name","fromName":"Ihar Hrachyshka","fromEmail":"ihar.hrachyshka@gmail.com","sentAt":"2026-07-03T02:35:54Z","receivedAt":"2026-07-03T02:36:10Z","isPatch":true,"body":"On macOS, git status may abort while reading a directory entry\nwhose UTF-8 name grows past NAME_MAX bytes:\n\n  __chk_fail_overflow\n  __strlcpy_chk\n  precompose_utf8_readdir\n  read_directory_recursive\n  wt_status_collect\n  cmd_status\n\nThe precompose wrapper already reallocates dirent_prec_psx for\nlong names, but d_name is declared as char[NAME_MAX + 1]. A\nfortified libc can still see that declared object size and reject a\nlarger strlcpy bound, even though the allocation was grown.\n\nMake d_name a FLEX_ARRAY and size allocations from offsetof(). That\nmatches the actual object layout with the dynamic allocation, so the\nfortified copy sees a destination whose size can grow with max_name_len.\n\nAdd a regression test that creates a 261-byte non-ASCII basename and\nruns status with core.precomposeunicode enabled.\n\nSigned-off-by: Ihar Hrachyshka <ihar.hrachyshka@gmail.com>\n---\n compat/precompose_utf8.c     | 12 ++++++++----\n compat/precompose_utf8.h     |  9 +++++----\n t/t3910-mac-os-precompose.sh | 15 +++++++++++++++\n 3 files changed, 28 insertions(+), 8 deletions(-)\n\ndiff --git a/compat/precompose_utf8.c b/compat/precompose_utf8.c\nindex 1711794..8077f62 100644\n--- a/compat/precompose_utf8.c\n+++ b/compat/precompose_utf8.c\n@@ -19,6 +19,11 @@ typedef char *iconv_ibp;\n static const char *repo_encoding = \"UTF-8\";\n static const char *path_encoding = \"UTF-8-MAC\";\n \n+static size_t dirent_prec_psx_size(size_t max_name_len)\n+{\n+\treturn st_add(offsetof(dirent_prec_psx, d_name), max_name_len);\n+}\n+\n static size_t has_non_ascii(const char *s, size_t maxlen, size_t *strlen_c)\n {\n \tconst uint8_t *ptr = (const uint8_t *)s;\n@@ -114,8 +119,8 @@ const char *precompose_argv_prefix(int argc, const char **argv, const char *pref\n PREC_DIR *precompose_utf8_opendir(const char *dirname)\n {\n \tPREC_DIR *prec_dir = xmalloc(sizeof(PREC_DIR));\n-\tprec_dir->dirent_nfc = xmalloc(sizeof(dirent_prec_psx));\n-\tprec_dir->dirent_nfc->max_name_len = sizeof(prec_dir->dirent_nfc->d_name);\n+\tprec_dir->dirent_nfc = xmalloc(dirent_prec_psx_size(NAME_MAX + 1));\n+\tprec_dir->dirent_nfc->max_name_len = NAME_MAX + 1;\n \n \tprec_dir->dirp = opendir(dirname);\n \tif (!prec_dir->dirp) {\n@@ -145,8 +150,7 @@ struct dirent_prec_psx *precompose_utf8_readdir(PREC_DIR *prec_dir)\n \t\tint ret_errno = errno;\n \n \t\tif (new_maxlen > prec_dir->dirent_nfc->max_name_len) {\n-\t\t\tsize_t new_len = sizeof(dirent_prec_psx) + new_maxlen -\n-\t\t\t\tsizeof(prec_dir->dirent_nfc->d_name);\n+\t\t\tsize_t new_len = dirent_prec_psx_size(new_maxlen);\n \n \t\t\tprec_dir->dirent_nfc = xrealloc(prec_dir->dirent_nfc, new_len);\n \t\t\tprec_dir->dirent_nfc->max_name_len = new_maxlen;\ndiff --git a/compat/precompose_utf8.h b/compat/precompose_utf8.h\nindex fea06cf..c7c3cc2 100644\n--- a/compat/precompose_utf8.h\n+++ b/compat/precompose_utf8.h\n@@ -14,11 +14,12 @@ typedef struct dirent_prec_psx {\n \n \t/*\n \t * See http://pubs.opengroup.org/onlinepubs/9699919799/basedefs/dirent.h.html\n-\t * NAME_MAX + 1 should be enough, but some systems have\n-\t * NAME_MAX=255 and strlen(d_name) may return 508 or 510\n-\t * Solution: allocate more when needed, see precompose_utf8_readdir()\n+\t * Start with room for NAME_MAX + 1 bytes, but keep d_name as a\n+\t * flexible array. Some systems have NAME_MAX=255 while strlen(d_name)\n+\t * from readdir() may return 508 or 510 bytes. Grow the allocation as\n+\t * needed in precompose_utf8_readdir().\n \t */\n-\tchar   d_name[NAME_MAX+1];\n+\tchar   d_name[FLEX_ARRAY];\n } dirent_prec_psx;\n \n \ndiff --git a/t/t3910-mac-os-precompose.sh b/t/t3910-mac-os-precompose.sh\nindex 6d5918c..fda4a76 100755\n--- a/t/t3910-mac-os-precompose.sh\n+++ b/t/t3910-mac-os-precompose.sh\n@@ -207,6 +207,21 @@ test_expect_success \"Add long precomposed filename\" '\n \tgit commit -m \"Long filename\"\n '\n \n+test_expect_success \"status with long non-ASCII filename\" '\n+\ttest_when_finished \"rm -rf long-utf8-status\" &&\n+\tgit init long-utf8-status &&\n+\t(\n+\t\tcd long-utf8-status &&\n+\t\ttest \"$(git config --bool core.precomposeunicode)\" = true &&\n+\t\tlong_utf8_name=$(\n+\t\t\tperl -e \"print q(a) x 249, qq(\\342\\200\\224) x 3, q(.md)\"\n+\t\t) &&\n+\t\ttest \"$(printf \"%s\" \"$long_utf8_name\" | wc -c | tr -d \" \")\" = 261 &&\n+\t\tprintf \"content\\n\" >\"$long_utf8_name\" &&\n+\t\tgit status --porcelain=v1 >actual\n+\t)\n+'\n+\n test_expect_failure 'handle existing decomposed filenames' '\n \techo content >\"verbatim.$Adiarnfd\" &&\n \tgit -c core.precomposeunicode=false add \"verbatim.$Adiarnfd\" &&\n\nbase-commit: e9019fcafe0040228b8631c30f97ae1adb61bcdc\n-- \n2.54.0\n\n"},{"id":"547032","messageId":"20260703050800.GA29216@tb-raspi4","threadId":"65916","inReplyTo":"20260703023554.36577-1-ihar.hrachyshka@gmail.com","subject":"Re: [PATCH] precompose_utf8: use a flex array for d_name","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2026-07-03T05:08:00Z","receivedAt":"2026-07-03T05:08:03Z","isPatch":true,"body":"On Thu, Jul 02, 2026 at 10:35:54PM -0400, Ihar Hrachyshka wrote:\n> On macOS, git status may abort while reading a directory entry\n> whose UTF-8 name grows past NAME_MAX bytes:\n> \n>   __chk_fail_overflow\n>   __strlcpy_chk\n>   precompose_utf8_readdir\n>   read_directory_recursive\n>   wt_status_collect\n>   cmd_status\n> \n> The precompose wrapper already reallocates dirent_prec_psx for\n> long names, but d_name is declared as char[NAME_MAX + 1]. A\n> fortified libc can still see that declared object size and reject a\n> larger strlcpy bound, even though the allocation was grown.\n> \n> Make d_name a FLEX_ARRAY and size allocations from offsetof(). That\n> matches the actual object layout with the dynamic allocation, so the\n> fortified copy sees a destination whose size can grow with max_name_len.\n> \n> Add a regression test that creates a 261-byte non-ASCII basename and\n> runs status with core.precomposeunicode enabled.\n> \n> Signed-off-by: Ihar Hrachyshka <ihar.hrachyshka@gmail.com>\n\nNice, thanks for the patch.\nOne minor nit/question:\nDo we need a\ntest_have_prereq PERL\nin t/t3910 ?\n\n[]\n> diff --git a/t/t3910-mac-os-precompose.sh b/t/t3910-mac-os-precompose.sh\n> index 6d5918c..fda4a76 100755\n> --- a/t/t3910-mac-os-precompose.sh\n> +++ b/t/t3910-mac-os-precompose.sh\n> @@ -207,6 +207,21 @@ test_expect_success \"Add long precomposed filename\" '\n>  \tgit commit -m \"Long filename\"\n>  '\n>  \n> +test_expect_success \"status with long non-ASCII filename\" '\n> +\ttest_when_finished \"rm -rf long-utf8-status\" &&\n> +\tgit init long-utf8-status &&\n> +\t(\n> +\t\tcd long-utf8-status &&\n> +\t\ttest \"$(git config --bool core.precomposeunicode)\" = true &&\n> +\t\tlong_utf8_name=$(\n> +\t\t\tperl -e \"print q(a) x 249, qq(\\342\\200\\224) x 3, q(.md)\"\n> +\t\t) &&\n> +\t\ttest \"$(printf \"%s\" \"$long_utf8_name\" | wc -c | tr -d \" \")\" = 261 &&\n> +\t\tprintf \"content\\n\" >\"$long_utf8_name\" &&\n> +\t\tgit status --porcelain=v1 >actual\n> +\t)\n> +'\n> +\n>  test_expect_failure 'handle existing decomposed filenames' '\n>  \techo content >\"verbatim.$Adiarnfd\" &&\n>  \tgit -c core.precomposeunicode=false add \"verbatim.$Adiarnfd\" &&\n> \n"},{"id":"547047","messageId":"xmqq8q7sjwkl.fsf@gitster.g","threadId":"65916","inReplyTo":"20260703050800.GA29216@tb-raspi4","subject":"Re: [PATCH] precompose_utf8: use a flex array for d_name","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-07-03T08:39:54Z","receivedAt":"2026-07-03T08:39:57Z","isPatch":true,"body":"Torsten Bögershausen <tboegi@web.de> writes:\n\n> Nice, thanks for the patch.  One minor nit/question: Do we need a\n> test_have_prereq PERL in t/t3910 ?\n\nGood question.\n\n>> +test_expect_success \"status with long non-ASCII filename\" '\n>> +\ttest_when_finished \"rm -rf long-utf8-status\" &&\n>> +\tgit init long-utf8-status &&\n>> +\t(\n>> +\t\tcd long-utf8-status &&\n>> +\t\ttest \"$(git config --bool core.precomposeunicode)\" = true &&\n>> +\t\tlong_utf8_name=$(\n>> +\t\t\tperl -e \"print q(a) x 249, qq(\\342\\200\\224) x 3, q(.md)\"\n>> +\t\t) &&\n>> +\t\ttest \"$(printf \"%s\" \"$long_utf8_name\" | wc -c | tr -d \" \")\" = 261 &&\n>> +\t\tprintf \"content\\n\" >\"$long_utf8_name\" &&\n\nI would say that if we are going to use this construct as-is, then\nwe do need the prereq.\n\nBut as far as I can see, this is mostly to create a very long\nfilename, which does not require perl at all, with 9 bytes of binary\nwhich could be easily done with printf with the ame backslash\nnotation.\n\nSo, if we can fix the test, that would be preferrable.\n\nThanks.\n\n>> +\t\tgit status --porcelain=v1 >actual\n>> +\t)\n>> +'\n>> +\n>>  test_expect_failure 'handle existing decomposed filenames' '\n>>  \techo content >\"verbatim.$Adiarnfd\" &&\n>>  \tgit -c core.precomposeunicode=false add \"verbatim.$Adiarnfd\" &&\n>> \n"},{"id":"547048","messageId":"akd1m6KoUh7N8yyE@pks.im","threadId":"65916","inReplyTo":"20260703023554.36577-1-ihar.hrachyshka@gmail.com","subject":"Re: [PATCH] precompose_utf8: use a flex array for d_name","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-07-03T08:40:59Z","receivedAt":"2026-07-03T08:41:05Z","isPatch":true,"body":"On Thu, Jul 02, 2026 at 10:35:54PM -0400, Ihar Hrachyshka wrote:\n> On macOS, git status may abort while reading a directory entry\n> whose UTF-8 name grows past NAME_MAX bytes:\n> \n>   __chk_fail_overflow\n>   __strlcpy_chk\n>   precompose_utf8_readdir\n>   read_directory_recursive\n>   wt_status_collect\n>   cmd_status\n> \n> The precompose wrapper already reallocates dirent_prec_psx for\n> long names, but d_name is declared as char[NAME_MAX + 1]. A\n> fortified libc can still see that declared object size and reject a\n> larger strlcpy bound, even though the allocation was grown.\n> \n> Make d_name a FLEX_ARRAY and size allocations from offsetof(). That\n> matches the actual object layout with the dynamic allocation, so the\n> fortified copy sees a destination whose size can grow with max_name_len.\n> \n> Add a regression test that creates a 261-byte non-ASCII basename and\n> runs status with core.precomposeunicode enabled.\n\nHm. Why does macOS even allow you to create a file that has a basename\nlonger than NAME_MAX? Does macOS count unicode characters specially?\n\n> diff --git a/compat/precompose_utf8.c b/compat/precompose_utf8.c\n> index 1711794..8077f62 100644\n> --- a/compat/precompose_utf8.c\n> +++ b/compat/precompose_utf8.c\n> @@ -19,6 +19,11 @@ typedef char *iconv_ibp;\n>  static const char *repo_encoding = \"UTF-8\";\n>  static const char *path_encoding = \"UTF-8-MAC\";\n>  \n> +static size_t dirent_prec_psx_size(size_t max_name_len)\n> +{\n> +\treturn st_add(offsetof(dirent_prec_psx, d_name), max_name_len);\n> +}\n> +\n>  static size_t has_non_ascii(const char *s, size_t maxlen, size_t *strlen_c)\n>  {\n>  \tconst uint8_t *ptr = (const uint8_t *)s;\n> @@ -114,8 +119,8 @@ const char *precompose_argv_prefix(int argc, const char **argv, const char *pref\n>  PREC_DIR *precompose_utf8_opendir(const char *dirname)\n>  {\n>  \tPREC_DIR *prec_dir = xmalloc(sizeof(PREC_DIR));\n> -\tprec_dir->dirent_nfc = xmalloc(sizeof(dirent_prec_psx));\n> -\tprec_dir->dirent_nfc->max_name_len = sizeof(prec_dir->dirent_nfc->d_name);\n> +\tprec_dir->dirent_nfc = xmalloc(dirent_prec_psx_size(NAME_MAX + 1));\n> +\tprec_dir->dirent_nfc->max_name_len = NAME_MAX + 1;\n\nWe have the `FLEX_ALLOC_MEM()` macro that would probably be a better fit\ncompared to introducing `dirent_prec_psx_size()`.\n\nAlso, when converting this to a flex array, can't we do better here and\nallocate the structures with the right size? Otherwise, I expect that we\noverallocate most of the entrise.\n\n> @@ -145,8 +150,7 @@ struct dirent_prec_psx *precompose_utf8_readdir(PREC_DIR *prec_dir)\n>  \t\tint ret_errno = errno;\n>  \n>  \t\tif (new_maxlen > prec_dir->dirent_nfc->max_name_len) {\n> -\t\t\tsize_t new_len = sizeof(dirent_prec_psx) + new_maxlen -\n> -\t\t\t\tsizeof(prec_dir->dirent_nfc->d_name);\n> +\t\t\tsize_t new_len = dirent_prec_psx_size(new_maxlen);\n>  \n>  \t\t\tprec_dir->dirent_nfc = xrealloc(prec_dir->dirent_nfc, new_len);\n>  \t\t\tprec_dir->dirent_nfc->max_name_len = new_maxlen;\n\nOkay, here we indeed have to realloc though, and thus we can't quite\navoid `dirent_prec_psx_size()`. Too bad.\n\nThanks!\n\nPatrick\n"},{"id":"547115","messageId":"f35346ce-056f-4add-b071-2703c2455daa@gmail.com","threadId":"65916","inReplyTo":"akd1m6KoUh7N8yyE@pks.im","subject":"Re: [PATCH] precompose_utf8: use a flex array for d_name","fromName":"Ihar Hrachyshka","fromEmail":"ihar.hrachyshka@gmail.com","sentAt":"2026-07-03T20:20:15Z","receivedAt":"2026-07-03T20:20:17Z","isPatch":true,"body":"On 7/3/26 4:40 AM, Patrick Steinhardt wrote:\n> On Thu, Jul 02, 2026 at 10:35:54PM -0400, Ihar Hrachyshka wrote:\n>> On macOS, git status may abort while reading a directory entry\n>> whose UTF-8 name grows past NAME_MAX bytes:\n>>\n>>    __chk_fail_overflow\n>>    __strlcpy_chk\n>>    precompose_utf8_readdir\n>>    read_directory_recursive\n>>    wt_status_collect\n>>    cmd_status\n>>\n>> The precompose wrapper already reallocates dirent_prec_psx for\n>> long names, but d_name is declared as char[NAME_MAX + 1]. A\n>> fortified libc can still see that declared object size and reject a\n>> larger strlcpy bound, even though the allocation was grown.\n>>\n>> Make d_name a FLEX_ARRAY and size allocations from offsetof(). That\n>> matches the actual object layout with the dynamic allocation, so the\n>> fortified copy sees a destination whose size can grow with max_name_len.\n>>\n>> Add a regression test that creates a 261-byte non-ASCII basename and\n>> runs status with core.precomposeunicode enabled.\n> Hm. Why does macOS even allow you to create a file that has a basename\n> longer than NAME_MAX? Does macOS count unicode characters specially?\n\n\nYes, macOS file names can exceed NAME_MAX bytes because the real dirent \nlimit in system headers is:\n\n#define __DARWIN_MAXPATHLEN 1024\n\n#define __DARWIN_STRUCT_DIRENTRY { \\\nchar d_name[__DARWIN_MAXPATHLEN]; /* entry name (up to MAXPATHLEN bytes) \n*/ \\\n}\n\n(for a very old 32-bit ABI it's 256 but it's not really relevant)\n\nThis in-memory limit may be further capped by file system. For HFS+, \nit's 255 16-bit Unicode characters (as per on-disk format). For APFS, \non-disk theoretically allows up to 1022 UTF-8 bytes, but my testing \nsuggests they still enforce the same 255 character limit somewhere in \nkernel API layer. (Which means that they could later expand the maximum \nfilename length further without changing the on-disk format.)\n\nSo effectively, today on Darwin, the real limit is \"up to 255 2-byte \ncode points\", not bytes. Which is potentially beyond NAME_MAX.\n\n...that said, Linux readdir() doesn't guarantee NAME_MAX limit either. \n From readdir(3):\n\n\"\"\"\n\n         Note that while the call\n\n             fpathconf(fd, _PC_NAME_MAX)\n\n         returns the value 255 for most filesystems, on some filesystems\n         (e.g., CIFS, Windows SMB servers), the null-terminated filename\n         that is (correctly) returned in .d_name can actually exceed this\n         size.  In such cases, the .d_reclen field will contain a value\n         that exceeds the size of the glibc dirent structure shown above.\n\n\"\"\"\n\n\nThe man page also advises against using sizeof() against dirent structs. \n(Which is what we currently do - against our own MacOS helper dirent \nstruct.)\n\n\nAs a side note, it probably means neither Darwin nor Linux readdir() is \nPOSIX compliant, because, as per:\n\nhttps://pubs.opengroup.org/onlinepubs/9799919799/basedefs/dirent.h.html\n\n\n\"The array d_name in each of these structures is of unspecified size, \nbut shall contain a filename of at most {NAME_MAX} bytes followed by a \nterminating null byte.\"\n\n\n>> diff --git a/compat/precompose_utf8.c b/compat/precompose_utf8.c\n>> index 1711794..8077f62 100644\n>> --- a/compat/precompose_utf8.c\n>> +++ b/compat/precompose_utf8.c\n>> @@ -19,6 +19,11 @@ typedef char *iconv_ibp;\n>>   static const char *repo_encoding = \"UTF-8\";\n>>   static const char *path_encoding = \"UTF-8-MAC\";\n>>   \n>> +static size_t dirent_prec_psx_size(size_t max_name_len)\n>> +{\n>> +\treturn st_add(offsetof(dirent_prec_psx, d_name), max_name_len);\n>> +}\n>> +\n>>   static size_t has_non_ascii(const char *s, size_t maxlen, size_t *strlen_c)\n>>   {\n>>   \tconst uint8_t *ptr = (const uint8_t *)s;\n>> @@ -114,8 +119,8 @@ const char *precompose_argv_prefix(int argc, const char **argv, const char *pref\n>>   PREC_DIR *precompose_utf8_opendir(const char *dirname)\n>>   {\n>>   \tPREC_DIR *prec_dir = xmalloc(sizeof(PREC_DIR));\n>> -\tprec_dir->dirent_nfc = xmalloc(sizeof(dirent_prec_psx));\n>> -\tprec_dir->dirent_nfc->max_name_len = sizeof(prec_dir->dirent_nfc->d_name);\n>> +\tprec_dir->dirent_nfc = xmalloc(dirent_prec_psx_size(NAME_MAX + 1));\n>> +\tprec_dir->dirent_nfc->max_name_len = NAME_MAX + 1;\n> We have the `FLEX_ALLOC_MEM()` macro that would probably be a better fit\n> compared to introducing `dirent_prec_psx_size()`.\n>\n> Also, when converting this to a flex array, can't we do better here and\n> allocate the structures with the right size? Otherwise, I expect that we\n> overallocate most of the entrise.\n\n\nAs I understand it, this is a *per-directory* buffer that starts from \nNAME_MAX + 1, then gets expanded as entries with names longer than \nNAME_MAX + 1 are encountered. It is reused for next entries.\n\n\n>> @@ -145,8 +150,7 @@ struct dirent_prec_psx *precompose_utf8_readdir(PREC_DIR *prec_dir)\n>>   \t\tint ret_errno = errno;\n>>   \n>>   \t\tif (new_maxlen > prec_dir->dirent_nfc->max_name_len) {\n>> -\t\t\tsize_t new_len = sizeof(dirent_prec_psx) + new_maxlen -\n>> -\t\t\t\tsizeof(prec_dir->dirent_nfc->d_name);\n>> +\t\t\tsize_t new_len = dirent_prec_psx_size(new_maxlen);\n>>   \n>>   \t\t\tprec_dir->dirent_nfc = xrealloc(prec_dir->dirent_nfc, new_len);\n>>   \t\t\tprec_dir->dirent_nfc->max_name_len = new_maxlen;\n> Okay, here we indeed have to realloc though, and thus we can't quite\n> avoid `dirent_prec_psx_size()`. Too bad.\n>\n> Thanks!\n>\n> Patrick\n\n\n"},{"id":"547141","messageId":"20260704233724.16928-1-ihar.hrachyshka@gmail.com","threadId":"65916","inReplyTo":"20260703023554.36577-1-ihar.hrachyshka@gmail.com","subject":"[PATCH v2] precompose_utf8: use a flex array for d_name","fromName":"Ihar Hrachyshka","fromEmail":"ihar.hrachyshka@gmail.com","sentAt":"2026-07-04T23:37:24Z","receivedAt":"2026-07-04T23:37:41Z","isPatch":true,"body":"On macOS, git status may abort while reading a directory entry\nwhose UTF-8 name grows past NAME_MAX bytes:\n\n  __chk_fail_overflow\n  __strlcpy_chk\n  precompose_utf8_readdir\n  read_directory_recursive\n  wt_status_collect\n  cmd_status\n\nThe precompose wrapper already reallocates dirent_prec_psx for\nlong names, but d_name is declared as char[NAME_MAX + 1]. A\nfortified libc can still see that declared object size and reject a\nlarger strlcpy bound, even though the allocation was grown.\n\nMake d_name a FLEX_ARRAY and size allocations from offsetof(). That\nmatches the actual object layout with the dynamic allocation, so the\nfortified copy sees a destination whose size can grow with max_name_len.\n\nAdd a regression test that creates an over-NAME_MAX non-ASCII basename\nand runs status with core.precomposeunicode enabled.\n\nSigned-off-by: Ihar Hrachyshka <ihar.hrachyshka@gmail.com>\n---\n\nChanges in v2:\n- Drop perl from the regression test and use printf/tr instead.\n- Use the minimal 256-byte filename that reproduces the crash.\n\n compat/precompose_utf8.c     | 12 ++++++++----\n compat/precompose_utf8.h     |  9 +++++----\n t/t3910-mac-os-precompose.sh | 16 ++++++++++++++++\n 3 files changed, 29 insertions(+), 8 deletions(-)\n\ndiff --git a/compat/precompose_utf8.c b/compat/precompose_utf8.c\nindex 1711794..8077f62 100644\n--- a/compat/precompose_utf8.c\n+++ b/compat/precompose_utf8.c\n@@ -19,6 +19,11 @@ typedef char *iconv_ibp;\n static const char *repo_encoding = \"UTF-8\";\n static const char *path_encoding = \"UTF-8-MAC\";\n \n+static size_t dirent_prec_psx_size(size_t max_name_len)\n+{\n+\treturn st_add(offsetof(dirent_prec_psx, d_name), max_name_len);\n+}\n+\n static size_t has_non_ascii(const char *s, size_t maxlen, size_t *strlen_c)\n {\n \tconst uint8_t *ptr = (const uint8_t *)s;\n@@ -114,8 +119,8 @@ const char *precompose_argv_prefix(int argc, const char **argv, const char *pref\n PREC_DIR *precompose_utf8_opendir(const char *dirname)\n {\n \tPREC_DIR *prec_dir = xmalloc(sizeof(PREC_DIR));\n-\tprec_dir->dirent_nfc = xmalloc(sizeof(dirent_prec_psx));\n-\tprec_dir->dirent_nfc->max_name_len = sizeof(prec_dir->dirent_nfc->d_name);\n+\tprec_dir->dirent_nfc = xmalloc(dirent_prec_psx_size(NAME_MAX + 1));\n+\tprec_dir->dirent_nfc->max_name_len = NAME_MAX + 1;\n \n \tprec_dir->dirp = opendir(dirname);\n \tif (!prec_dir->dirp) {\n@@ -145,8 +150,7 @@ struct dirent_prec_psx *precompose_utf8_readdir(PREC_DIR *prec_dir)\n \t\tint ret_errno = errno;\n \n \t\tif (new_maxlen > prec_dir->dirent_nfc->max_name_len) {\n-\t\t\tsize_t new_len = sizeof(dirent_prec_psx) + new_maxlen -\n-\t\t\t\tsizeof(prec_dir->dirent_nfc->d_name);\n+\t\t\tsize_t new_len = dirent_prec_psx_size(new_maxlen);\n \n \t\t\tprec_dir->dirent_nfc = xrealloc(prec_dir->dirent_nfc, new_len);\n \t\t\tprec_dir->dirent_nfc->max_name_len = new_maxlen;\ndiff --git a/compat/precompose_utf8.h b/compat/precompose_utf8.h\nindex fea06cf..c7c3cc2 100644\n--- a/compat/precompose_utf8.h\n+++ b/compat/precompose_utf8.h\n@@ -14,11 +14,12 @@ typedef struct dirent_prec_psx {\n \n \t/*\n \t * See http://pubs.opengroup.org/onlinepubs/9699919799/basedefs/dirent.h.html\n-\t * NAME_MAX + 1 should be enough, but some systems have\n-\t * NAME_MAX=255 and strlen(d_name) may return 508 or 510\n-\t * Solution: allocate more when needed, see precompose_utf8_readdir()\n+\t * Start with room for NAME_MAX + 1 bytes, but keep d_name as a\n+\t * flexible array. Some systems have NAME_MAX=255 while strlen(d_name)\n+\t * from readdir() may return 508 or 510 bytes. Grow the allocation as\n+\t * needed in precompose_utf8_readdir().\n \t */\n-\tchar   d_name[NAME_MAX+1];\n+\tchar   d_name[FLEX_ARRAY];\n } dirent_prec_psx;\n \n \ndiff --git a/t/t3910-mac-os-precompose.sh b/t/t3910-mac-os-precompose.sh\nindex 6d5918c..ea75fb4 100755\n--- a/t/t3910-mac-os-precompose.sh\n+++ b/t/t3910-mac-os-precompose.sh\n@@ -207,6 +207,22 @@ test_expect_success \"Add long precomposed filename\" '\n \tgit commit -m \"Long filename\"\n '\n \n+test_expect_success \"status with long non-ASCII filename\" '\n+\ttest_when_finished \"rm -rf long-utf8-status\" &&\n+\tgit init long-utf8-status &&\n+\t(\n+\t\tcd long-utf8-status &&\n+\t\ttest \"$(git config --bool core.precomposeunicode)\" = true &&\n+\t\tlong_utf8_name=$(\n+\t\t\tprintf \"%253s\\342\\200\\224\" \"\" |\n+\t\t\ttr \" \" a\n+\t\t) &&\n+\t\ttest \"$(printf \"%s\" \"$long_utf8_name\" | wc -c | tr -d \" \")\" = 256 &&\n+\t\tprintf \"content\\n\" >\"$long_utf8_name\" &&\n+\t\tgit status --porcelain=v1 >actual\n+\t)\n+'\n+\n test_expect_failure 'handle existing decomposed filenames' '\n \techo content >\"verbatim.$Adiarnfd\" &&\n \tgit -c core.precomposeunicode=false add \"verbatim.$Adiarnfd\" &&\n-- \n2.54.0\n\n"}]}