From: Matthieu Beauchamp Date: Tue, 06 Jan 2026 20:45:56 GMT Subject: Re: [PATCH] ignores: handle non UTF-8 exclude files Message-ID: In-Reply-To: On Sun, Jan 4, 2026 at 2:40 PM brian m. carlson wrote: > > On 2026-01-03 at 22:16:57, Matthieu Beauchamp-Boulay via GitGitGadget wrote: > > When reading exclude files, git assumes it is encoded in UTF-8 and will > > fail to apply patterns if it isn't. This is a silent failure as no warning > > or errors are shown to the users. This is a problem that can take a while > > to diagnose as many users will not think of checking the encoding of their > > file and may believe their patterns are wrong instead. Users may also > > accidentally commit undesired files. > > This isn't actually true. Git allows arbitrary byte sequences in the > file because Git allows filenames to have arbitrary byte sequences, just > like Unix. Yes thank you for pointing that out, I had some wrong assumptions about the encodings. > > On Windows, this happens if a user uses Windows PowerShell to create the > > file, which results in a UTF-16LE file with a BOM. This issue was discussed > > here https://github.com/git-for-windows/git/issues/3329. An example of > > where a user was confused that his exclude file was not working is cited > > https://github.com/git-for-windows/git/issues/3227. > > Ah, yes, here's the problem. UTF-16LE is used on Windows, and on > Windows, Git stores pathnames as if they were converted into UTF-8, so > you do need to write the filenames in UTF-8 in the ignore file. > Yes, the conversion from UTF16-LE to UTF-8 would need to be platform specific. > > A minimal fix should at least warn the user if git cannot properly decode > > the exclude file. Ideally, git would handle any given Unicode file. > > As I mentioned, the file isn't necessarily in UTF-8 or Unicode. Here's > an example shell script to demonstrate (requires a non-macOS Unix): > > ---- > #!/bin/sh > > rm -fr test-repo > git init --object-format=sha256 test-repo > cd test-repo > touch abc.txt > touch "$(printf '\220')" > printf '\220\n' >.gitignore > git add . > git status > git ls-files -io --exclude-standard > ---- > > I'll point out that all of this is also true for things like config > files (which are also used in `.gitmodules`) and `.gitattributes` files. > If we wanted to make a change, we would be wise to make it everywhere. > > However, if we wanted to force `.gitignore` to UTF-8, we'd need to have > an escape mechanism to write non-UTF-8 sequences, and as far as I know, > we don't. Right, I don't think forcing UTF-8 everywhere is worth it for a relatively simple issue. If I can find a portable way to determine that an encoding is incorrect (and possibly reencode it), I could apply it to those other files as well. > > First, check if a BOM is present. If it is, decode the file to UTF-8. > > If no BOM is detected, then try to parse the file as UTF-8. If that fails, > > attempt to decode the file using the working tree encoding of the file, > > if any. If that fails, print a warning to tell the user that the exclude > > file could not be decoded and skip the file. > > We do not accept and strip BOMs in UTF-8 files elsewhere (including in > things like `git diff` output), so we should not do so here, either. > For Unicode files, if there is no BOM, then the standard is that it's > assumed to automatically be UTF-8, so a BOM is superfluous and not > recommended. I meant checking for UTF-16 and UTF-32 BOMs and then converting to UTF-8, I will clarify if this part is still in the revision. > > diff --git a/t/lib-encoding.sh b/t/lib-encoding.sh > > index 2dabc8c73e..1b1cc357ba 100644 > > --- a/t/lib-encoding.sh > > +++ b/t/lib-encoding.sh > > @@ -23,3 +23,11 @@ write_utf32 () { > > fi && > > iconv -f UTF-8 -t UTF-32 > > } > > + > > +write_encoded () { > > + iconv -f UTF-8 -t "$1" > > +} > > + > > +write_bom () { > > + echo "$@" | perl -pe 's/\s+//g; $_=pack("H*", $_)' > > +} > > \ No newline at end of file > > We place newlines at the end of our text files unless there's a good > reason no to. > -- > brian m. carlson (they/them) > Toronto, Ontario, CA I will fix it, I would've assumed that clang-format would fix that.