{"thread":{"id":"64715","subject":"[PATCH] ignores: handle non UTF-8 exclude files","startedAt":"2026-01-03T22:17:00Z","lastAt":"2026-01-08T01:13:08Z","messageCount":13,"participants":["Matthieu Beauchamp-Boulay via GitGitGadget","Junio C Hamano","Torsten Bögershausen","brian m. carlson","Matthieu Beauchamp","Collin Funk","Phillip Wood"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"532959","messageId":"pull.2157.git.git.1767478617198.gitgitgadget@gmail.com","threadId":"64715","inReplyTo":null,"subject":"[PATCH] ignores: handle non UTF-8 exclude files","fromName":"Matthieu Beauchamp-Boulay via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-03T22:16:57Z","receivedAt":"2026-01-03T22:17:00Z","isPatch":true,"sender":{"key":"name:Matthieu Beauchamp-Boulay","avatar":null},"body":"From: Matthieu Beauchamp-Boulay <matthieu.beauchamp.boulay@gmail.com>\n\nWhen reading exclude files, git assumes it is encoded in UTF-8 and will\nfail to apply patterns if it isn't. This is a silent failure as no warning\nor errors are shown to the users. This is a problem that can take a while\nto diagnose as many users will not think of checking the encoding of their\nfile and may believe their patterns are wrong instead. Users may also\naccidentally commit undesired files.\n\nOn Windows, this happens if a user uses Windows PowerShell to create the\nfile, which results in a UTF-16LE file with a BOM. This issue was discussed\nhere https://github.com/git-for-windows/git/issues/3329. An example of\nwhere a user was confused that his exclude file was not working is cited\nhttps://github.com/git-for-windows/git/issues/3227.\n\nA minimal fix should at least warn the user if git cannot properly decode\nthe exclude file. Ideally, git would handle any given Unicode file.\n\nFirst, check if a BOM is present. If it is, decode the file to UTF-8.\nIf no BOM is detected, then try to parse the file as UTF-8. If that fails,\nattempt to decode the file using the working tree encoding of the file,\nif any. If that fails, print a warning to tell the user that the exclude\nfile could not be decoded and skip the file.\n\nThis raises the issue that if the entire tree is encoded in, for example\nUTF-16BE (no BOM), then even if the encoding is given in .gitattributes,\ngit would not be able to decode it. I believe that this is still\nacceptable since a warning will be emitted for the file (since it has no\nBOM, is not valid UTF-8 and no working tree encoding could be found).\n\nOne case that isn't handled is if a wrong encoding is given in the\nattributes and the exclude file has no BOM and is not UTF-8. Using\niconv to convert an UTF16BE file to UTF-8 while specifying UTF-16LE\nyields gibberish without an error and so this case is a silent failure\nwhere no patterns will match.\n\nSigned-off-by: Matthieu Beauchamp-Boulay <matthieu.beauchamp.boulay@gmail.com>\n---\n    ignores: handle non UTF-8 exclude files\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2157%2FMatthieu-Beauchamp%2Funicode-support-gitignore-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2157/Matthieu-Beauchamp/unicode-support-gitignore-v1\nPull-Request: https://github.com/git/git/pull/2157\n\n dir.c              |  29 +++++++++++++\n t/lib-encoding.sh  |   8 ++++\n t/t0008-ignores.sh | 102 +++++++++++++++++++++++++++++++++++++++++++++\n utf8.c             |  29 +++++++++++++\n utf8.h             |  12 ++++++\n 5 files changed, 180 insertions(+)\n\ndiff --git a/dir.c b/dir.c\nindex b00821f294..895d476253 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -1154,7 +1154,9 @@ static int add_patterns(const char *fname, const char *base, int baselen,\n \tint r;\n \tint fd;\n \tsize_t size = 0;\n+\tsize_t reencoded_size = 0;\n \tchar *buf;\n+\tchar *reencoded = NULL;\n \n \tif (flags & PATTERN_NOFOLLOW)\n \t\tfd = open_nofollow(fname, O_RDONLY);\n@@ -1190,7 +1192,34 @@ static int add_patterns(const char *fname, const char *base, int baselen,\n \t\t\tclose(fd);\n \t\t\treturn -1;\n \t\t}\n+\n+\t\tif (!try_reencode_to_utf8(buf, size, &reencoded, &reencoded_size) && !is_valid_utf8(buf, size)) {\n+\t\t\tstruct conv_attrs ca;\n+\t\t\tconvert_attrs(istate, &ca, fname);\n+\n+\t\t\tif (ca.working_tree_encoding) {\n+\t\t\t\treencoded = reencode_string_len(buf, size, \"UTF-8\", ca.working_tree_encoding, &reencoded_size);\n+\t\t\t\tif (!reencoded) {\n+\t\t\t\t\twarning(_(\"Failed to decode exclude file %s from encoding %s\"), pl->src, ca.working_tree_encoding);\n+\t\t\t\t\tfree(buf);\n+\t\t\t\t\treturn -1;\n+\t\t\t\t}\n+\t\t\t} else {\n+\t\t\t\twarning(_(\"Ignoring exclude file with unknown encoding: %s\"), pl->src);\n+\t\t\t\tfree(buf);\n+\t\t\t\treturn -1;\n+\t\t\t}\n+\t\t}\n+\n+\t\tif (reencoded) {\n+\t\t\tsize = reencoded_size;\n+\t\t\tfree(buf);\n+\t\t\tbuf = xmallocz(size);\n+\t\t\tmemcpy(buf, reencoded, size);\n+\t\t\tfree(reencoded);\n+\t\t}\n \t\tbuf[size++] = '\\n';\n+\n \t\tclose(fd);\n \t\tif (oid_stat) {\n \t\t\tint pos;\ndiff --git a/t/lib-encoding.sh b/t/lib-encoding.sh\nindex 2dabc8c73e..1b1cc357ba 100644\n--- a/t/lib-encoding.sh\n+++ b/t/lib-encoding.sh\n@@ -23,3 +23,11 @@ write_utf32 () {\n \tfi &&\n \ticonv -f UTF-8 -t UTF-32\n }\n+\n+write_encoded () {\n+  iconv -f UTF-8 -t \"$1\"\n+}\n+\n+write_bom () {\n+  echo \"$@\" | perl -pe 's/\\s+//g; $_=pack(\"H*\", $_)'\n+}\n\\ No newline at end of file\ndiff --git a/t/t0008-ignores.sh b/t/t0008-ignores.sh\nindex db8bde280e..d5a3002ffb 100755\n--- a/t/t0008-ignores.sh\n+++ b/t/t0008-ignores.sh\n@@ -4,6 +4,7 @@ test_description=check-ignore\n \n TEST_CREATE_REPO_NO_TEMPLATE=1\n . ./test-lib.sh\n+. \"$TEST_DIRECTORY/lib-encoding.sh\"\n \n init_vars () {\n \tglobal_excludes=\"global-excludes\"\n@@ -963,4 +964,105 @@ test_expect_success EXPENSIVE 'large exclude file ignored in tree' '\n \ttest_cmp expect err\n '\n \n+############################################################################\n+#\n+# test handling of unicode for .gitignore when BOM is preset or worktree encoding is set for the file\n+\n+supports_encoding () {\n+  encoding=\"$1\" &&\n+  d=\"support-bom-$encoding\" &&\n+  shift &&\n+\tmkdir -p \"$d\" &&\n+\ttouch \"$d/file\" \"$d/excluded\" &&\n+\twrite_bom \"$@\" > \"$d/.gitignore\" &&\n+  echo excluded | write_encoded \"$encoding\" >> \"$d/.gitignore\"\n+\tgit check-ignore \"$d/excluded\" > \"$d/actual\" &&\n+\techo \"$d/excluded\" > expect &&\n+\ttest_cmp expect \"$d/actual\"\n+}\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-8 with BOM' '\n+  supports_encoding \"UTF-8\" EF BB BF\n+'\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-16LE when given a BOM' '\n+  supports_encoding \"UTF-16LE\" FF FE\n+'\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-16BE when given a BOM' '\n+  supports_encoding \"UTF-16BE\" FE FF\n+'\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-32LE when given a BOM' '\n+  supports_encoding \"UTF-32LE\" FF FE 00 00\n+'\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-32BE when given a BOM' '\n+  supports_encoding \"UTF-32BE\" 00 00 FE FF\n+'\n+\n+supports_reading_ignore_in_working_tree_encoding () {\n+  encoding=\"$1\" &&\n+  d=\"support-wt-$encoding\" &&\n+\tmkdir -p \"$d\" &&\n+\ttouch \"$d/file\" \"$d/excluded\" &&\n+  echo \".gitignore\t\ttext working-tree-encoding=$encoding\" > \"$d/.gitattributes\" &&\n+  echo excluded | write_encoded \"$encoding\" > \"$d/.gitignore\"\n+\tgit check-ignore \"$d/excluded\" > \"$d/actual\" &&\n+\techo \"$d/excluded\" > expect &&\n+\ttest_cmp expect \"$d/actual\"\n+}\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-8 when it is set as working tree encoding' '\n+  supports_reading_ignore_in_working_tree_encoding \"UTF-8\"\n+'\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-16LE when it is set as working tree encoding' '\n+  supports_reading_ignore_in_working_tree_encoding \"UTF-16LE\"\n+'\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-16BE when it is set as working tree encoding' '\n+  supports_reading_ignore_in_working_tree_encoding \"UTF-16BE\"\n+'\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-32LE when it is set as working tree encoding' '\n+  supports_reading_ignore_in_working_tree_encoding \"UTF-32LE\"\n+'\n+\n+test_expect_success ICONV 'Can read gitignore in UTF-32BE when it is set as working tree encoding' '\n+  supports_reading_ignore_in_working_tree_encoding \"UTF-32BE\"\n+'\n+\n+test_expect_success ICONV 'Issues a warning if encoding cannot be deduced' '\n+  d=\"warn-unknown-encoding\" &&\n+\tmkdir -p \"$d\" &&\n+\ttouch \"$d/file\" \"$d/excluded\" &&\n+\techo excluded | write_encoded \"UTF-16BE\" > \"$d/.gitignore\" &&\n+\ttest_must_fail git check-ignore \"$d/excluded\" > actual 2>&1 &&\n+\techo \"warning: Ignoring exclude file with unknown encoding: $d/.gitignore\" > expect &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_success ICONV 'Warns if the exclude file cannot be decoded due to encoded attributes' '\n+  d=\"warn-cant-decode-attributes\" &&\n+\tmkdir -p \"$d\" &&\n+\ttouch \"$d/file\" \"$d/excluded\" &&\n+\techo excluded | write_encoded \"UTF-16BE\" > \"$d/.gitignore\" &&\n+  echo \".gitignore\t\ttext working-tree-encoding=UTF-16BE\" | write_encoded \"UTF-16BE\" > \"$d/.gitattributes\" &&\n+\ttest_must_fail git check-ignore \"$d/excluded\" > actual 2>&1 &&\n+\techo \"warning: Ignoring exclude file with unknown encoding: $d/.gitignore\" > expect &&\n+\ttest_cmp expect actual\n+'\n+\n+test_expect_failure ICONV 'Issues a warning if the wrong encoding is given' '\n+  d=\"warn-wrong-encoding\" &&\n+\tmkdir -p \"$d\" &&\n+\ttouch \"$d/file\" \"$d/excluded\" &&\n+\techo excluded | write_encoded \"UTF-16BE\" > \"$d/.gitignore\" &&\n+  echo \".gitignore\t\ttext working-tree-encoding=UTF-16LE\" > \"$d/.gitattributes\" &&\n+\ttest_must_fail git check-ignore \"$d/excluded\" > actual 2>&1 &&\n+\techo \"warning: Ignoring exclude file with unknown encoding: $d/.gitignore\" > expect &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\ndiff --git a/utf8.c b/utf8.c\nindex 35a0251939..904cf20f97 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -252,6 +252,17 @@ int is_utf8(const char *text)\n \treturn 1;\n }\n \n+int is_valid_utf8(const char *text, size_t len)\n+{\n+\twhile (text && len > 0) {\n+\t\tucs_char_t ch = pick_one_utf8_char(&text, &len);\n+\t\tif (text && ch == 0 && len > 0)\n+\t\t\treturn 0;\n+\t}\n+\n+\treturn text != NULL;\n+}\n+\n static void strbuf_add_indented_text(struct strbuf *buf, const char *text,\n \t\t\t\t     int indent, int indent2)\n {\n@@ -643,6 +654,24 @@ int is_missing_required_utf_bom(const char *enc, const char *data, size_t len)\n \t);\n }\n \n+int try_reencode_to_utf8(const char *text, size_t len, char **out_text, size_t *out_len)\n+{\n+\tconst char *in_encoding;\n+\tif (has_bom_prefix(text, len, utf32_be_bom, sizeof(utf32_be_bom)))\n+\t\tin_encoding = \"UTF-32BE\";\n+\telse if (has_bom_prefix(text, len, utf32_le_bom, sizeof(utf32_le_bom)))\n+\t\tin_encoding = \"UTF-32LE\";\n+\telse if (has_bom_prefix(text, len, utf16_be_bom, sizeof(utf16_be_bom)))\n+\t\tin_encoding = \"UTF-16BE\";\n+\telse if (has_bom_prefix(text, len, utf16_le_bom, sizeof(utf16_le_bom)))\n+\t\tin_encoding = \"UTF-16LE\";\n+\telse\n+\t\treturn 0;\n+\n+\t*out_text = reencode_string_len(text, len, \"UTF-8\", in_encoding, out_len);\n+\treturn *out_text != NULL;\n+}\n+\n /*\n  * Returns first character length in bytes for multi-byte `text` according to\n  * `encoding`.\ndiff --git a/utf8.h b/utf8.h\nindex cf8ecb0f21..e52e4f7c06 100644\n--- a/utf8.h\n+++ b/utf8.h\n@@ -10,6 +10,12 @@ int utf8_width(const char **start, size_t *remainder_p);\n int utf8_strnwidth(const char *string, size_t len, int skip_ansi);\n int utf8_strwidth(const char *string);\n int is_utf8(const char *text);\n+\n+/*\n+ * Checks that the string is valid UTF-8 that does not contain the null byte\n+ * except at the end of the string\n+ */\n+int is_valid_utf8(const char *text, size_t len);\n int is_encoding_utf8(const char *name);\n int same_encoding(const char *, const char *);\n __attribute__((format (printf, 2, 3)))\n@@ -48,6 +54,12 @@ static inline char *reencode_string(const char *in,\n \t\t\t\t   NULL);\n }\n \n+/*\n+ * Returns true if an unicode BOM is detected and the string can be reencoded to UTF-8.\n+ * In that case the string is reencoded to UTF-8 in *out_text.\n+ */\n+int try_reencode_to_utf8(const char *text, size_t len, char **out_text, size_t *out_len);\n+\n int mbs_chrlen(const char **text, size_t *remainder_p, const char *encoding);\n \n /*\n\nbase-commit: c4a0c8845e2426375ad257b6c221a3a7d92ecfda\n-- \ngitgitgadget\n"},{"id":"532969","messageId":"xmqqsecmnl3d.fsf@gitster.g","threadId":"64715","inReplyTo":"pull.2157.git.git.1767478617198.gitgitgadget@gmail.com","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-01-04T02:54:46Z","receivedAt":"2026-01-04T02:54:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Matthieu Beauchamp-Boulay via GitGitGadget\"\n<gitgitgadget@gmail.com> writes:\n\n> From: Matthieu Beauchamp-Boulay <matthieu.beauchamp.boulay@gmail.com>\n>\n> When reading exclude files, git assumes it is encoded in UTF-8 and will\n> fail to apply patterns if it isn't.\n\nIs it true?  I thought we assume that the exclude patters are\nwritten in such a way to match the encoding of the pathnames,\nwhatever used on the platform that our calls to readdir(3) returns.\nSome platforms may have compat/ code to convert these paths and\nforce use of UTF-8, but please do not write such platform local\nconventions as if it were universal characteristics of our system.\n\n\"ignores\" -> \"exclude\" on the title, as that is the canonical word\nwe use in the codebase to refer to the ignore mechanism.\n\n"},{"id":"532995","messageId":"20260104173524.GA29867@tb-raspi4","threadId":"64715","inReplyTo":"pull.2157.git.git.1767478617198.gitgitgadget@gmail.com","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2026-01-04T17:35:24Z","receivedAt":"2026-01-04T17:35:33Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On Sat, Jan 03, 2026 at 10:16:57PM +0000, Matthieu Beauchamp-Boulay via GitGitGadget wrote:\n> From: Matthieu Beauchamp-Boulay <matthieu.beauchamp.boulay@gmail.com>\nThanks for contributing - some comments inlie\n> \n> When reading exclude files, git assumes it is encoded in UTF-8 and will\nQuestion: The report citet below talks about ignore files.\n\n> fail to apply patterns if it isn't. This is a silent failure as no warning\n> or errors are shown to the users. This is a problem that can take a while\n> to diagnose as many users will not think of checking the encoding of their\n> file and may believe their patterns are wrong instead. Users may also\n> accidentally commit undesired files.\nNote:\ngit status is your friend.\nBlindly commiting without checking what is staged or not may\nlead to unwanted results.\n\n> \n> On Windows, this happens if a user uses Windows PowerShell to create the\n> file, which results in a UTF-16LE file with a BOM.\n>  This issue was discussed\n> here https://github.com/git-for-windows/git/issues/3329. An example of\n> where a user was confused that his exclude file was not working is cited\n> https://github.com/git-for-windows/git/issues/3227.\nA very short research indicates that powershell can be configured\nto use UTF-8. I am not a powershell user, please correct if I am wrong.\n\n> \n> A minimal fix should at least warn the user if git cannot properly decode\n> the exclude file.\nI think that reading an ignore file that contains a '\\0' could/should\nGit to complain. If someone asks my, most users are tempted to ignore\nwarnings for different reasons. Bailing out may feel more unpolite\nbut more clear that somethinh is wrong.\n\n>Ideally, git would handle any given Unicode file.\nThat is debatable.\n\n> \n> First, check if a BOM is present. If it is, decode the file to UTF-8.\n> If no BOM is detected, then try to parse the file as UTF-8. If that fails,\n> attempt to decode the file using the working tree encoding of the file,\n> if any. If that fails, print a warning to tell the user that the exclude\n> file could not be decoded and skip the file.\n> \n> This raises the issue that if the entire tree is encoded in, for example\n> UTF-16BE (no BOM), then even if the encoding is given in .gitattributes,\n> git would not be able to decode it.\n\"able to decode: Yes. But willing to do so: not with the patch, right ?\n> I believe that this is still\n> acceptable since a warning will be emitted for the file (since it has no\n> BOM, is not valid UTF-8 and no working tree encoding could be found).\n> \n> One case that isn't handled is if a wrong encoding is given in the\n> attributes and the exclude file has no BOM and is not UTF-8. Using\n> iconv to convert an UTF16BE file to UTF-8 while specifying UTF-16LE\n> yields gibberish without an error and so this case is a silent failure\n> where no patterns will match.\nOne question is, if we should look at working_tree_encoding at all.\nThe other one is, how much UTF-16 handling of ignore or\nother file should we have have in Git ?\nIt seems that this fix is for a very special case only ?\n\nFrom\nhttps://github.com/git-for-windows/git/issues/3329\nwe read:\n/******/\nif (size > 1 && buf[0] == 0xff && buf[1] == 0xfe) {\n    char *reencoded = reencode_string_len(buf, size, \"UTF-8\", \"UTF16-LE-BOM\", &size);\n    if (!reencoded)\n        die(_(\"could not convert contents of '%s' from UTF-16\"), fname);\n    free(buf);\n    buf = reencoded;\n}\n/******/\n(Which seems a simpler suggestion)\nHowever,  there is no UTF-16-LE-BOM in iconv \n(at least in the majority of implementations), \nso a better approach, totaly untested, may be:\n\nif (size >= 2 && buf[0] == 0xff && buf[1] == 0xfe) {\n    char *reencoded = reencode_string_len(buf+2, size-2, \"UTF-8\", \"UTF16\", &size);\n    if (!reencoded)\n        die(_(\"could not convert contents of '%s' from UTF-16\"), fname);\n    free(buf);\n    buf = reencoded;\n}\n\nThis leads to some free thinking, especially when we look at\nother implementations of Git:\nWould it be better to simply bail out on UTF-16 files ?\nTechically all files with a '\\0'.\n[snip] \n"},{"id":"532999","messageId":"aVrCHr_NRDqNjPn0@fruit.crustytoothpaste.net","threadId":"64715","inReplyTo":"pull.2157.git.git.1767478617198.gitgitgadget@gmail.com","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-01-04T19:40:14Z","receivedAt":"2026-01-04T19:40:23Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2026-01-03 at 22:16:57, Matthieu Beauchamp-Boulay via GitGitGadget wrote:\n> When reading exclude files, git assumes it is encoded in UTF-8 and will\n> fail to apply patterns if it isn't. This is a silent failure as no warning\n> or errors are shown to the users. This is a problem that can take a while\n> to diagnose as many users will not think of checking the encoding of their\n> file and may believe their patterns are wrong instead. Users may also\n> accidentally commit undesired files.\n\nThis isn't actually true.  Git allows arbitrary byte sequences in the\nfile because Git allows filenames to have arbitrary byte sequences, just\nlike Unix.\n\n> On Windows, this happens if a user uses Windows PowerShell to create the\n> file, which results in a UTF-16LE file with a BOM. This issue was discussed\n> here https://github.com/git-for-windows/git/issues/3329. An example of\n> where a user was confused that his exclude file was not working is cited\n> https://github.com/git-for-windows/git/issues/3227.\n\nAh, yes, here's the problem.  UTF-16LE is used on Windows, and on\nWindows, Git stores pathnames as if they were converted into UTF-8, so\nyou do need to write the filenames in UTF-8 in the ignore file.\n\n> A minimal fix should at least warn the user if git cannot properly decode\n> the exclude file. Ideally, git would handle any given Unicode file.\n\nAs I mentioned, the file isn't necessarily in UTF-8 or Unicode.  Here's\nan example shell script to demonstrate (requires a non-macOS Unix):\n\n----\n#!/bin/sh\n\nrm -fr test-repo\ngit init --object-format=sha256 test-repo\ncd test-repo\ntouch abc.txt\ntouch \"$(printf '\\220')\"\nprintf '\\220\\n' >.gitignore\ngit add .\ngit status\ngit ls-files -io --exclude-standard\n----\n\nI'll point out that all of this is also true for things like config\nfiles (which are also used in `.gitmodules`) and `.gitattributes` files.\nIf we wanted to make a change, we would be wise to make it everywhere.\n\nHowever, if we wanted to force `.gitignore` to UTF-8, we'd need to have\nan escape mechanism to write non-UTF-8 sequences, and as far as I know,\nwe don't.\n\n> First, check if a BOM is present. If it is, decode the file to UTF-8.\n> If no BOM is detected, then try to parse the file as UTF-8. If that fails,\n> attempt to decode the file using the working tree encoding of the file,\n> if any. If that fails, print a warning to tell the user that the exclude\n> file could not be decoded and skip the file.\n\nWe do not accept and strip BOMs in UTF-8 files elsewhere (including in\nthings like `git diff` output), so we should not do so here, either.\nFor Unicode files, if there is no BOM, then the standard is that it's\nassumed to automatically be UTF-8, so a BOM is superfluous and not\nrecommended.\n\n> diff --git a/t/lib-encoding.sh b/t/lib-encoding.sh\n> index 2dabc8c73e..1b1cc357ba 100644\n> --- a/t/lib-encoding.sh\n> +++ b/t/lib-encoding.sh\n> @@ -23,3 +23,11 @@ write_utf32 () {\n>  \tfi &&\n>  \ticonv -f UTF-8 -t UTF-32\n>  }\n> +\n> +write_encoded () {\n> +  iconv -f UTF-8 -t \"$1\"\n> +}\n> +\n> +write_bom () {\n> +  echo \"$@\" | perl -pe 's/\\s+//g; $_=pack(\"H*\", $_)'\n> +}\n> \\ No newline at end of file\n\nWe place newlines at the end of our text files unless there's a good\nreason no to.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"533158","messageId":"CALH9GrYpWG2WPM76WDFnK-tFFTFtg3hZFLjg3z27gOYKeXpmxw@mail.gmail.com","threadId":"64715","inReplyTo":"xmqqsecmnl3d.fsf@gitster.g","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"Matthieu Beauchamp","fromEmail":"matthieu.beauchamp.boulay@gmail.com","sentAt":"2026-01-06T19:52:36Z","receivedAt":"2026-01-06T19:52:48Z","isPatch":true,"sender":{"key":"matthieu.beauchamp.boulay@gmail.com","avatar":null},"body":"On Sat, Jan 3, 2026 at 9:54 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Matthieu Beauchamp-Boulay via GitGitGadget\"\n> <gitgitgadget@gmail.com> writes:\n>\n> > From: Matthieu Beauchamp-Boulay <matthieu.beauchamp.boulay@gmail.com>\n> >\n> > When reading exclude files, git assumes it is encoded in UTF-8 and will\n> > fail to apply patterns if it isn't.\n>\n> Is it true?  I thought we assume that the exclude patters are\n> written in such a way to match the encoding of the pathnames,\n> whatever used on the platform that our calls to readdir(3) returns.\n> Some platforms may have compat/ code to convert these paths and\n> force use of UTF-8, but please do not write such platform local\n> conventions as if it were universal characteristics of our system.\n\nI believe you are correct, I wrongly assumed git would always\nmanipulate UTF-8 paths.\nThe revision will need to take the platform into consideration.\n\n> \"ignores\" -> \"exclude\" on the title, as that is the canonical word\n> we use in the codebase to refer to the ignore mechanism.\n\nThank you, I will update in the revision.\n"},{"id":"533159","messageId":"CALH9GrYi0dYo4LJg8ww1cDOETiOT44m0zQgkxLsxqEuMmv_myQ@mail.gmail.com","threadId":"64715","inReplyTo":"20260104173524.GA29867@tb-raspi4","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"Matthieu Beauchamp","fromEmail":"matthieu.beauchamp.boulay@gmail.com","sentAt":"2026-01-06T20:32:50Z","receivedAt":"2026-01-06T20:33:03Z","isPatch":true,"sender":{"key":"matthieu.beauchamp.boulay@gmail.com","avatar":null},"body":"On Sun, Jan 4, 2026 at 12:35 PM Torsten Bögershausen <tboegi@web.de> wrote:\n>\n> On Sat, Jan 03, 2026 at 10:16:57PM +0000, Matthieu Beauchamp-Boulay via GitGitGadget wrote:\n> > From: Matthieu Beauchamp-Boulay <matthieu.beauchamp.boulay@gmail.com>\n> Thanks for contributing - some comments inlie\n> >\n> > When reading exclude files, git assumes it is encoded in UTF-8 and will\n> Question: The report citet below talks about ignore files.\n>\n> > fail to apply patterns if it isn't. This is a silent failure as no warning\n> > or errors are shown to the users. This is a problem that can take a while\n> > to diagnose as many users will not think of checking the encoding of their\n> > file and may believe their patterns are wrong instead. Users may also\n> > accidentally commit undesired files.\n> Note:\n> git status is your friend.\n> Blindly commiting without checking what is staged or not may\n> lead to unwanted results.\n\nYes of course, I'll remove that last line as it is not the problem I'm\nreally trying to fix.\n\n> >\n> > On Windows, this happens if a user uses Windows PowerShell to create the\n> > file, which results in a UTF-16LE file with a BOM.\n> >  This issue was discussed\n> > here https://github.com/git-for-windows/git/issues/3329. An example of\n> > where a user was confused that his exclude file was not working is cited\n> > https://github.com/git-for-windows/git/issues/3227.\n> A very short research indicates that powershell can be configured\n> to use UTF-8. I am not a powershell user, please correct if I am wrong.\n>\n\nYes you are correct, but I want to address the issues for users who may not\nrealize that they used the wrong encoding when creating their exclude file.\nFor that case I don't see how the fact that powershell can be configured to\nUTF-8 helps, aside from preventing repeating the same mistake.\n\n> >\n> > A minimal fix should at least warn the user if git cannot properly decode\n> > the exclude file.\n> I think that reading an ignore file that contains a '\\0' could/should\n> Git to complain. If someone asks my, most users are tempted to ignore\n> warnings for different reasons. Bailing out may feel more unpolite\n> but more clear that somethinh is wrong.\n\nWhile I agree that warnings may be ignored, I feel like a wrongly encoded\nexclude file is not an error that warrants stopping git entirely.\n\nAs other reviewers mentioned, I wrongly assumed that the encoding would\nbe UTF-8. The idea of looking for the null byte in the exclude file\nmay be helpful\nsince any d_name from readdir (3) is null terminated. Checking for a null byte\nbefore the end of the file could be a simple check to detect a bad exclude file.\n\n> >Ideally, git would handle any given Unicode file.\n> That is debatable.\n\nOf course, I'll rephrase that part.\n\n> >\n> > First, check if a BOM is present. If it is, decode the file to UTF-8.\n> > If no BOM is detected, then try to parse the file as UTF-8. If that fails,\n> > attempt to decode the file using the working tree encoding of the file,\n> > if any. If that fails, print a warning to tell the user that the exclude\n> > file could not be decoded and skip the file.\n> >\n> > This raises the issue that if the entire tree is encoded in, for example\n> > UTF-16BE (no BOM), then even if the encoding is given in .gitattributes,\n> > git would not be able to decode it.\n> \"able to decode: Yes. But willing to do so: not with the patch, right ?\n> > I believe that this is still\n> > acceptable since a warning will be emitted for the file (since it has no\n> > BOM, is not valid UTF-8 and no working tree encoding could be found).\n> >\n> > One case that isn't handled is if a wrong encoding is given in the\n> > attributes and the exclude file has no BOM and is not UTF-8. Using\n> > iconv to convert an UTF16BE file to UTF-8 while specifying UTF-16LE\n> > yields gibberish without an error and so this case is a silent failure\n> > where no patterns will match.\n> One question is, if we should look at working_tree_encoding at all.\n> The other one is, how much UTF-16 handling of ignore or\n> other file should we have have in Git ?\n> It seems that this fix is for a very special case only ?\n>\n> From\n> https://github.com/git-for-windows/git/issues/3329\n> we read:\n> /******/\n> if (size > 1 && buf[0] == 0xff && buf[1] == 0xfe) {\n>     char *reencoded = reencode_string_len(buf, size, \"UTF-8\", \"UTF16-LE-BOM\", &size);\n>     if (!reencoded)\n>         die(_(\"could not convert contents of '%s' from UTF-16\"), fname);\n>     free(buf);\n>     buf = reencoded;\n> }\n> /******/\n> (Which seems a simpler suggestion)\n> However,  there is no UTF-16-LE-BOM in iconv\n> (at least in the majority of implementations),\n> so a better approach, totaly untested, may be:\n>\n> if (size >= 2 && buf[0] == 0xff && buf[1] == 0xfe) {\n>     char *reencoded = reencode_string_len(buf+2, size-2, \"UTF-8\", \"UTF16\", &size);\n>     if (!reencoded)\n>         die(_(\"could not convert contents of '%s' from UTF-16\"), fname);\n>     free(buf);\n>     buf = reencoded;\n> }\n>\n> This leads to some free thinking, especially when we look at\n> other implementations of Git:\n> Would it be better to simply bail out on UTF-16 files ?\n> Techically all files with a '\\0'.\n> [snip]\n\nI was trying to cover more possible use cases, but this may not be a desired\nbehavior after all. Other reviewers pointed out that the exclude file may have\nan abitrary encoding that needs to match the encoding of the paths as read\nby git when using readdir (3).\n\nYou are correct, UTF-16-LE-BOM is a 'fictional' encoding handled by git. Git\nhandles the BOM and iconv will be passed the UTF-16LE encoding instead.\n\nI would've liked to be able to handle any wrongly encoded exclude files, but\nit's more complicated than I originally thought. Checking for a null byte\ncould be a simple way to detect some wrong encodings.\n"},{"id":"533161","messageId":"CALH9GrYOjb92gjrtdjwapFH9L73XGg1Kan8uz1aVLpSXNURi+Q@mail.gmail.com","threadId":"64715","inReplyTo":"aVrCHr_NRDqNjPn0@fruit.crustytoothpaste.net","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"Matthieu Beauchamp","fromEmail":"matthieu.beauchamp.boulay@gmail.com","sentAt":"2026-01-06T20:45:56Z","receivedAt":"2026-01-06T20:46:08Z","isPatch":true,"sender":{"key":"matthieu.beauchamp.boulay@gmail.com","avatar":null},"body":"On Sun, Jan 4, 2026 at 2:40 PM brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> On 2026-01-03 at 22:16:57, Matthieu Beauchamp-Boulay via GitGitGadget wrote:\n> > When reading exclude files, git assumes it is encoded in UTF-8 and will\n> > fail to apply patterns if it isn't. This is a silent failure as no warning\n> > or errors are shown to the users. This is a problem that can take a while\n> > to diagnose as many users will not think of checking the encoding of their\n> > file and may believe their patterns are wrong instead. Users may also\n> > accidentally commit undesired files.\n>\n> This isn't actually true.  Git allows arbitrary byte sequences in the\n> file because Git allows filenames to have arbitrary byte sequences, just\n> like Unix.\n\nYes thank you for pointing that out, I had some wrong assumptions about the\nencodings.\n\n> > On Windows, this happens if a user uses Windows PowerShell to create the\n> > file, which results in a UTF-16LE file with a BOM. This issue was discussed\n> > here https://github.com/git-for-windows/git/issues/3329. An example of\n> > where a user was confused that his exclude file was not working is cited\n> > https://github.com/git-for-windows/git/issues/3227.\n>\n> Ah, yes, here's the problem.  UTF-16LE is used on Windows, and on\n> Windows, Git stores pathnames as if they were converted into UTF-8, so\n> you do need to write the filenames in UTF-8 in the ignore file.\n>\n\nYes, the conversion from UTF16-LE to UTF-8 would need to be platform\nspecific.\n\n> > A minimal fix should at least warn the user if git cannot properly decode\n> > the exclude file. Ideally, git would handle any given Unicode file.\n>\n> As I mentioned, the file isn't necessarily in UTF-8 or Unicode.  Here's\n> an example shell script to demonstrate (requires a non-macOS Unix):\n>\n> ----\n> #!/bin/sh\n>\n> rm -fr test-repo\n> git init --object-format=sha256 test-repo\n> cd test-repo\n> touch abc.txt\n> touch \"$(printf '\\220')\"\n> printf '\\220\\n' >.gitignore\n> git add .\n> git status\n> git ls-files -io --exclude-standard\n> ----\n>\n> I'll point out that all of this is also true for things like config\n> files (which are also used in `.gitmodules`) and `.gitattributes` files.\n> If we wanted to make a change, we would be wise to make it everywhere.\n>\n> However, if we wanted to force `.gitignore` to UTF-8, we'd need to have\n> an escape mechanism to write non-UTF-8 sequences, and as far as I know,\n> we don't.\n\nRight, I don't think forcing UTF-8 everywhere is worth it for a relatively\nsimple issue. If I can find a portable way to determine that an encoding\nis incorrect (and possibly reencode it), I could apply it to those other files\nas well.\n\n> > First, check if a BOM is present. If it is, decode the file to UTF-8.\n> > If no BOM is detected, then try to parse the file as UTF-8. If that fails,\n> > attempt to decode the file using the working tree encoding of the file,\n> > if any. If that fails, print a warning to tell the user that the exclude\n> > file could not be decoded and skip the file.\n>\n> We do not accept and strip BOMs in UTF-8 files elsewhere (including in\n> things like `git diff` output), so we should not do so here, either.\n> For Unicode files, if there is no BOM, then the standard is that it's\n> assumed to automatically be UTF-8, so a BOM is superfluous and not\n> recommended.\n\nI meant checking for UTF-16 and UTF-32 BOMs and then converting to UTF-8,\nI will clarify if this part is still in the revision.\n\n> > diff --git a/t/lib-encoding.sh b/t/lib-encoding.sh\n> > index 2dabc8c73e..1b1cc357ba 100644\n> > --- a/t/lib-encoding.sh\n> > +++ b/t/lib-encoding.sh\n> > @@ -23,3 +23,11 @@ write_utf32 () {\n> >       fi &&\n> >       iconv -f UTF-8 -t UTF-32\n> >  }\n> > +\n> > +write_encoded () {\n> > +  iconv -f UTF-8 -t \"$1\"\n> > +}\n> > +\n> > +write_bom () {\n> > +  echo \"$@\" | perl -pe 's/\\s+//g; $_=pack(\"H*\", $_)'\n> > +}\n> > \\ No newline at end of file\n>\n> We place newlines at the end of our text files unless there's a good\n> reason no to.\n> --\n> brian m. carlson (they/them)\n> Toronto, Ontario, CA\n\nI will fix it, I would've assumed that clang-format would fix that.\n"},{"id":"533169","messageId":"aV2ZS1lvLivi8xRH@fruit.crustytoothpaste.net","threadId":"64715","inReplyTo":"CALH9GrYOjb92gjrtdjwapFH9L73XGg1Kan8uz1aVLpSXNURi+Q@mail.gmail.com","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-01-06T23:22:51Z","receivedAt":"2026-01-06T23:23:00Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2026-01-06 at 20:45:56, Matthieu Beauchamp wrote:\n> On Sun, Jan 4, 2026 at 2:40 PM brian m. carlson\n> <sandals@crustytoothpaste.net> wrote:\n> > Ah, yes, here's the problem.  UTF-16LE is used on Windows, and on\n> > Windows, Git stores pathnames as if they were converted into UTF-8, so\n> > you do need to write the filenames in UTF-8 in the ignore file.\n> >\n> \n> Yes, the conversion from UTF16-LE to UTF-8 would need to be platform\n> specific.\n\nWe typically don't want platform-specific behaviour in Git.  Many Git\ncontributors do not work on Windows but we want things to work as much\nas possible identically across all platforms because it makes\ndevelopment easier, as well as making it easier for users to reason\nabout the project.  I, for one, don't have a Windows system (nor do I\nwant one) but I do want my Git code to just work there.\n\nAs an example, we still use a POSIX shell in aliases and other settings\non Windows despite the fact that PowerShell is built into Windows\nbecause it means that aliases and similar functionality just work\ncorrectly regardless of platform and it allows users to write a config\nfile that works everywhere.\n\nInstead of trying to force Git to gracefully handle UTF-16 in its config\nfiles, my strong recommendation is to adjust your PowerShell scripts to\nuse UTF-8 instead[0] or use a POSIX shell.  I'll note that Microsoft's\nnew Edit text editor[1] defaults to UTF-8 (and, except on Windows, LF\nline endings), so I know that Microsoft understands that UTF-8 is the\nproper encoding to use on the Internet today.\n\n[0] https://stackoverflow.com/questions/5596982/using-powershell-to-write-a-file-in-utf-8-without-the-bom\n[1] Available at https://github.com/microsoft/edit and apparently\nshipped with Windows.  I will say that I was impressed at its\nfunctionality for a 231 KiB binary footprint.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"533176","messageId":"87secimchc.fsf@gmail.com","threadId":"64715","inReplyTo":"aV2ZS1lvLivi8xRH@fruit.crustytoothpaste.net","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"Collin Funk","fromEmail":"collin.funk1@gmail.com","sentAt":"2026-01-07T01:35:11Z","receivedAt":"2026-01-07T01:35:14Z","isPatch":true,"sender":{"key":"collin.funk1@gmail.com","avatar":"https://avatars.githubusercontent.com/u/65689063?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> Instead of trying to force Git to gracefully handle UTF-16 in its config\n> files, my strong recommendation is to adjust your PowerShell scripts to\n> use UTF-8 instead[0] or use a POSIX shell.  I'll note that Microsoft's\n> new Edit text editor[1] defaults to UTF-8 (and, except on Windows, LF\n> line endings), so I know that Microsoft understands that UTF-8 is the\n> proper encoding to use on the Internet today.\n>\n> [1] Available at https://github.com/microsoft/edit and apparently\n> shipped with Windows.  I will say that I was impressed at its\n> functionality for a 231 KiB binary footprint.\n\nDoes it handle text that is not UTF-8 encoded?\n\nAn unfortunate trend that I have seen with Rust programs is that they\ncompletely disregard the systems locale. E.g. using\nLC_ALL=en_US.ISO-8859-1 and passing an \"À\" character as an option will\ntypically fail since it is encoded as 0xC0 which is not a valid UTF-8\ncharacter.\n\nI figured it was worth bringing up since Git may wany to think about it\nsome before introducing more Rust. I think it can be worked around by\nusing OsString [1], but I guess many people choose not to.\n\nCollin\n\n[1] https://doc.rust-lang.org/std/ffi/struct.OsString.html\n"},{"id":"533223","messageId":"331595ad-5c6a-4e01-bd0f-1dabb4bc0fcb@gmail.com","threadId":"64715","inReplyTo":"87secimchc.fsf@gmail.com","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-07T14:28:08Z","receivedAt":"2026-01-07T14:28:18Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 07/01/2026 01:35, Collin Funk wrote:\n> \n> An unfortunate trend that I have seen with Rust programs is that they\n> completely disregard the systems locale. E.g. using\n> LC_ALL=en_US.ISO-8859-1 and passing an \"À\" character as an option will\n> typically fail since it is encoded as 0xC0 which is not a valid UTF-8\n> character.\n> \n> I figured it was worth bringing up since Git may wany to think about it\n> some before introducing more Rust. I think it can be worked around by\n> using OsString [1], but I guess many people choose not to.\n\nGit will certainly want to continue to support non-utf8 encodings. \nThat's perfectly possible in rust but in my (rather limited) experience \nit does take a bit more effort than the equivalent code using the \nstandard library's String type. I find it particularly annoying that \n\"cargo run\" refuses to pass non-utf8 arguments to the program being run \nwhen the program has been carefully written to support them.\n\nThanks\n\nPhillip\n"},{"id":"533224","messageId":"072dc5ef-e750-4023-bf6c-30b4b143beca@gmail.com","threadId":"64715","inReplyTo":"CALH9GrYi0dYo4LJg8ww1cDOETiOT44m0zQgkxLsxqEuMmv_myQ@mail.gmail.com","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-01-07T14:36:19Z","receivedAt":"2026-01-07T14:36:22Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 06/01/2026 20:32, Matthieu Beauchamp wrote:\n> \n> Yes you are correct, but I want to address the issues for users who may not\n> realize that they used the wrong encoding when creating their exclude file.\n> For that case I don't see how the fact that powershell can be configured to\n> UTF-8 helps, aside from preventing repeating the same mistake.\n\nMy concern with that is that it ends up hampering collaboration with \npeople using bash on Windows or a native shell on other platforms. If \nthey append to a UTF-16 encoded .gitignore with \"echo path >>.gitignore\" \nyou'll end up with a mix of encodings in the same file. Similarly if you \nuse powershell to append to an existing file that is UTF-8 encoded with \n\"echo hello >>.gitignore\" is the appended text UTF-16 encoded resulting \nin mixed encodings in the same file?\n\nThanks\n\nPhillip\n\n"},{"id":"533256","messageId":"aV7ujZ2FeO7EleT5@fruit.crustytoothpaste.net","threadId":"64715","inReplyTo":"87secimchc.fsf@gmail.com","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2026-01-07T23:38:53Z","receivedAt":"2026-01-07T23:39:01Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2026-01-07 at 01:35:11, Collin Funk wrote:\n> An unfortunate trend that I have seen with Rust programs is that they\n> completely disregard the systems locale. E.g. using\n> LC_ALL=en_US.ISO-8859-1 and passing an \"À\" character as an option will\n> typically fail since it is encoded as 0xC0 which is not a valid UTF-8\n> character.\n\nGit does not usually directly read input and then convert it to other\nencodings unless specifically asked to (e.g., `working-tree-encoding`),\nso I fully expect that nothing will change there.  However, in many\ncases, Git also currently does not honour LC_ALL, such as for commit\nmessages.\n\n> I figured it was worth bringing up since Git may wany to think about it\n> some before introducing more Rust. I think it can be worked around by\n> using OsString [1], but I guess many people choose not to.\n\nThe people who have been working on Rust have been very careful to not\nmake assumptions that all data is UTF-8, and I don't expect that to\nchange.\n\nOsString is slightly problematic because it is effectively UTF-8-ish (on\nWindows, it's actually WTF-8 and on Unix it allows arbitrary bytes) but\nthere is no portable way to get any consistent byte encoding out of it.\n(In versions of Rust too new for us to use, there is a function that\nprovides a byte encoding but it's not guaranteed to be stable across\nversions.)  I have some custom code in one of my branches to handle the\nconversion to and from OsString to a consistent byte encoding using some\ntraits to paper over the operating system differences.\n\nIn general, I expect we will continue to use some C-based interfaces\n(possibly called via Rust wrappers) because Rust also does not expose\nthings like file descriptors on Windows or the full range of stat or\nother information we need.\n\nOne assumption I do think is safe to make is that arbitrary Unicode can\nbe printed to the terminal, such as in error messages.  Considering that\nvirtually everybody sets IUTF8 in Unix terminals and we effectively do\nthat right now with localized text, I think that's okay.\n-- \nbrian m. carlson (they/them)\nToronto, Ontario, CA\n"},{"id":"533258","messageId":"87ldi8aov2.fsf@gmail.com","threadId":"64715","inReplyTo":"aV7ujZ2FeO7EleT5@fruit.crustytoothpaste.net","subject":"Re: [PATCH] ignores: handle non UTF-8 exclude files","fromName":"Collin Funk","fromEmail":"collin.funk1@gmail.com","sentAt":"2026-01-08T01:13:05Z","receivedAt":"2026-01-08T01:13:08Z","isPatch":true,"sender":{"key":"collin.funk1@gmail.com","avatar":"https://avatars.githubusercontent.com/u/65689063?v=4"},"body":"\"brian m. carlson\" <sandals@crustytoothpaste.net> writes:\n\n> On 2026-01-07 at 01:35:11, Collin Funk wrote:\n>> An unfortunate trend that I have seen with Rust programs is that they\n>> completely disregard the systems locale. E.g. using\n>> LC_ALL=en_US.ISO-8859-1 and passing an \"À\" character as an option will\n>> typically fail since it is encoded as 0xC0 which is not a valid UTF-8\n>> character.\n>\n> Git does not usually directly read input and then convert it to other\n> encodings unless specifically asked to (e.g., `working-tree-encoding`),\n> so I fully expect that nothing will change there.  However, in many\n> cases, Git also currently does not honour LC_ALL, such as for commit\n> messages.\n\nThat makes sense.\n\n>> I figured it was worth bringing up since Git may wany to think about it\n>> some before introducing more Rust. I think it can be worked around by\n>> using OsString [1], but I guess many people choose not to.\n>\n> The people who have been working on Rust have been very careful to not\n> make assumptions that all data is UTF-8, and I don't expect that to\n> change.\n\nGreat, glad that it was considered. I guess you have to worry about\ncrates, but I think I recall wide agreement that Git was going to be\ncareful with what it decides to use.\n\n> OsString is slightly problematic because it is effectively UTF-8-ish (on\n> Windows, it's actually WTF-8 and on Unix it allows arbitrary bytes) but\n> there is no portable way to get any consistent byte encoding out of it.\n> (In versions of Rust too new for us to use, there is a function that\n> provides a byte encoding but it's not guaranteed to be stable across\n> versions.)  I have some custom code in one of my branches to handle the\n> conversion to and from OsString to a consistent byte encoding using some\n> traits to paper over the operating system differences.\n\nInteresting, good to know. Thanks.\n\nUnrelated to encoding, but two other things I noticed about Rust. Before\nmain() SIGPIPE is set to SIG_IGN which can be seen with the programs\nbelow:\n\n    $ cat main.rs \n    use std::io::{self, Write};\n    fn main() -> io::Result<()> {\n        io::stdout().write_all(b\"hello world\\n\")?;\n        Ok(())\n    }\n    $ cat main.c\n    #include <unistd.h>\n    #include <stdio.h>\n    #include <stdlib.h>\n    #include <string.h>\n    #include <errno.h>\n    int\n    main (void)\n    {\n      static const char message[] = \"hello world\\n\";\n      if (write (STDOUT_FILENO, message, sizeof message - 1) < 0)\n        {\n          fprintf (stderr, \"%s\\n\", strerror (errno));\n          return EXIT_FAILURE;\n        }\n      return EXIT_SUCCESS;\n    }\n    $ rustc main.rs\n    $ gcc main.c\n    $ ./main | :\n    Error: Os { code: 32, kind: BrokenPipe, message: \"Broken pipe\" }\n    $ echo ${PIPESTATUS[@]}\n    1 0\n    $ ./a.out | :\n    $ echo ${PIPESTATUS[@]}\n    141 0\n\nBefore executing a program using the standard library, SIGPIPE will be\nset to SIG_DFL. That is better than not doing that, but both behaviors\nmean that the typical behavior of inheriting signal actions from the\nparent process is impossible without hacks or an unstable feature that\nhas been unfortunately stagnant for years [1].\n\nBefore main() all standard file descriptors are also opened. While\nreasonable in many cases, is not the desired behavior for all programs.\nUsing the same example programs:\n\n    $ ./main >&-\n    $ echo $?\n    0\n    $ ./a.out >&-\n    Bad file descriptor\n    $ echo $?\n    1\n\nI'm not sure if either of those will affect 'git' at all, assuming it is\nmostly library code that is called from C.\n\nBut it will likely have to be considered if someone wants to write a\nprogram that goes in libexec that is executed by 'git'.\n\nCollin\n\n[1] https://dev-doc.rust-lang.org/beta/unstable-book/language-features/unix-sigpipe.html\n"}]}