Re: [PATCH] ignores: handle non UTF-8 exclude files
- From
- Matthieu Beauchamp <matthieu.beauchamp.boulay@gmail.com>
- Date
- Jan 6, 2026, 20:45 UTC
- Message-ID
- <CALH9GrYOjb92gjrtdjwapFH9L73XGg1Kan8uz1aVLpSXNURi+Q@mail.gmail.com>
- In-Reply-To
- <aVrCHr_NRDqNjPn0@fruit.crustytoothpaste.net>
On Sun, Jan 4, 2026 at 2:40 PM brian m. carlson <sandals@crustytoothpaste.net> wrote:
Show 12 quoted lines
> > On 2026-01-03 at 22:16:57, Matthieu Beauchamp-Boulay via GitGitGadget wrote: > > When reading exclude files, git assumes it is encoded in UTF-8 and will > > fail to apply patterns if it isn't. This is a silent failure as no warning > > or errors are shown to the users. This is a problem that can take a while > > to diagnose as many users will not think of checking the encoding of their > > file and may believe their patterns are wrong instead. Users may also > > accidentally commit undesired files. > > This isn't actually true. Git allows arbitrary byte sequences in the > file because Git allows filenames to have arbitrary byte sequences, just > like Unix.
Yes thank you for pointing that out, I had some wrong assumptions about the encodings.
Show 10 quoted lines
> > On Windows, this happens if a user uses Windows PowerShell to create the > > file, which results in a UTF-16LE file with a BOM. This issue was discussed > > here https://github.com/git-for-windows/git/issues/3329. An example of > > where a user was confused that his exclude file was not working is cited > > https://github.com/git-for-windows/git/issues/3227. > > Ah, yes, here's the problem. UTF-16LE is used on Windows, and on > Windows, Git stores pathnames as if they were converted into UTF-8, so > you do need to write the filenames in UTF-8 in the ignore file. >
Yes, the conversion from UTF16-LE to UTF-8 would need to be platform specific.
Show 27 quoted lines
> > A minimal fix should at least warn the user if git cannot properly decode > > the exclude file. Ideally, git would handle any given Unicode file. > > As I mentioned, the file isn't necessarily in UTF-8 or Unicode. Here's > an example shell script to demonstrate (requires a non-macOS Unix): > > ---- > #!/bin/sh > > rm -fr test-repo > git init --object-format=sha256 test-repo > cd test-repo > touch abc.txt > touch "$(printf '\220')" > printf '\220\n' >.gitignore > git add . > git status > git ls-files -io --exclude-standard > ---- > > I'll point out that all of this is also true for things like config > files (which are also used in `.gitmodules`) and `.gitattributes` files. > If we wanted to make a change, we would be wise to make it everywhere. > > However, if we wanted to force `.gitignore` to UTF-8, we'd need to have > an escape mechanism to write non-UTF-8 sequences, and as far as I know, > we don't.
Right, I don't think forcing UTF-8 everywhere is worth it for a relatively simple issue. If I can find a portable way to determine that an encoding is incorrect (and possibly reencode it), I could apply it to those other files as well.
Show 11 quoted lines
> > First, check if a BOM is present. If it is, decode the file to UTF-8. > > If no BOM is detected, then try to parse the file as UTF-8. If that fails, > > attempt to decode the file using the working tree encoding of the file, > > if any. If that fails, print a warning to tell the user that the exclude > > file could not be decoded and skip the file. > > We do not accept and strip BOMs in UTF-8 files elsewhere (including in > things like `git diff` output), so we should not do so here, either. > For Unicode files, if there is no BOM, then the standard is that it's > assumed to automatically be UTF-8, so a BOM is superfluous and not > recommended.
I meant checking for UTF-16 and UTF-32 BOMs and then converting to UTF-8, I will clarify if this part is still in the revision.
Show 23 quoted lines
> > diff --git a/t/lib-encoding.sh b/t/lib-encoding.sh
> > index 2dabc8c73e..1b1cc357ba 100644
> > --- a/t/lib-encoding.sh
> > +++ b/t/lib-encoding.sh
> > @@ -23,3 +23,11 @@ write_utf32 () {
> > fi &&
> > iconv -f UTF-8 -t UTF-32
> > }
> > +
> > +write_encoded () {
> > + iconv -f UTF-8 -t "$1"
> > +}
> > +
> > +write_bom () {
> > + echo "$@" | perl -pe 's/\s+//g; $_=pack("H*", $_)'
> > +}
> > \ No newline at end of file
>
> We place newlines at the end of our text files unless there's a good
> reason no to.
> --
> brian m. carlson (they/them)
> Toronto, Ontario, CAI will fix it, I would've assumed that clang-format would fix that.