Re: [PATCH] ignores: handle non UTF-8 exclude files
- From
- Matthieu Beauchamp <matthieu.beauchamp.boulay@gmail.com>
- Date
- Jan 6, 2026, 20:32 UTC
- Message-ID
- <CALH9GrYi0dYo4LJg8ww1cDOETiOT44m0zQgkxLsxqEuMmv_myQ@mail.gmail.com>
- In-Reply-To
- <20260104173524.GA29867@tb-raspi4>
On Sun, Jan 4, 2026 at 12:35 PM Torsten Bögershausen <tboegi@web.de> wrote:
Show 17 quoted lines
> > On Sat, Jan 03, 2026 at 10:16:57PM +0000, Matthieu Beauchamp-Boulay via GitGitGadget wrote: > > From: Matthieu Beauchamp-Boulay <matthieu.beauchamp.boulay@gmail.com> > Thanks for contributing - some comments inlie > > > > When reading exclude files, git assumes it is encoded in UTF-8 and will > Question: The report citet below talks about ignore files. > > > fail to apply patterns if it isn't. This is a silent failure as no warning > > or errors are shown to the users. This is a problem that can take a while > > to diagnose as many users will not think of checking the encoding of their > > file and may believe their patterns are wrong instead. Users may also > > accidentally commit undesired files. > Note: > git status is your friend. > Blindly commiting without checking what is staged or not may > lead to unwanted results.
Yes of course, I'll remove that last line as it is not the problem I'm really trying to fix.
Show 10 quoted lines
> > > > On Windows, this happens if a user uses Windows PowerShell to create the > > file, which results in a UTF-16LE file with a BOM. > > This issue was discussed > > here https://github.com/git-for-windows/git/issues/3329. An example of > > where a user was confused that his exclude file was not working is cited > > https://github.com/git-for-windows/git/issues/3227. > A very short research indicates that powershell can be configured > to use UTF-8. I am not a powershell user, please correct if I am wrong. >
Yes you are correct, but I want to address the issues for users who may not realize that they used the wrong encoding when creating their exclude file. For that case I don't see how the fact that powershell can be configured to UTF-8 helps, aside from preventing repeating the same mistake.
Show 7 quoted lines
> > > > A minimal fix should at least warn the user if git cannot properly decode > > the exclude file. > I think that reading an ignore file that contains a '\0' could/should > Git to complain. If someone asks my, most users are tempted to ignore > warnings for different reasons. Bailing out may feel more unpolite > but more clear that somethinh is wrong.
While I agree that warnings may be ignored, I feel like a wrongly encoded exclude file is not an error that warrants stopping git entirely.
As other reviewers mentioned, I wrongly assumed that the encoding would be UTF-8. The idea of looking for the null byte in the exclude file may be helpful since any d_name from readdir (3) is null terminated. Checking for a null byte before the end of the file could be a simple check to detect a bad exclude file.
> >Ideally, git would handle any given Unicode file. > That is debatable.
Of course, I'll rephrase that part.
Show 55 quoted lines
> >
> > First, check if a BOM is present. If it is, decode the file to UTF-8.
> > If no BOM is detected, then try to parse the file as UTF-8. If that fails,
> > attempt to decode the file using the working tree encoding of the file,
> > if any. If that fails, print a warning to tell the user that the exclude
> > file could not be decoded and skip the file.
> >
> > This raises the issue that if the entire tree is encoded in, for example
> > UTF-16BE (no BOM), then even if the encoding is given in .gitattributes,
> > git would not be able to decode it.
> "able to decode: Yes. But willing to do so: not with the patch, right ?
> > I believe that this is still
> > acceptable since a warning will be emitted for the file (since it has no
> > BOM, is not valid UTF-8 and no working tree encoding could be found).
> >
> > One case that isn't handled is if a wrong encoding is given in the
> > attributes and the exclude file has no BOM and is not UTF-8. Using
> > iconv to convert an UTF16BE file to UTF-8 while specifying UTF-16LE
> > yields gibberish without an error and so this case is a silent failure
> > where no patterns will match.
> One question is, if we should look at working_tree_encoding at all.
> The other one is, how much UTF-16 handling of ignore or
> other file should we have have in Git ?
> It seems that this fix is for a very special case only ?
>
> From
> https://github.com/git-for-windows/git/issues/3329
> we read:
> /******/
> if (size > 1 && buf[0] == 0xff && buf[1] == 0xfe) {
> char *reencoded = reencode_string_len(buf, size, "UTF-8", "UTF16-LE-BOM", &size);
> if (!reencoded)
> die(_("could not convert contents of '%s' from UTF-16"), fname);
> free(buf);
> buf = reencoded;
> }
> /******/
> (Which seems a simpler suggestion)
> However, there is no UTF-16-LE-BOM in iconv
> (at least in the majority of implementations),
> so a better approach, totaly untested, may be:
>
> if (size >= 2 && buf[0] == 0xff && buf[1] == 0xfe) {
> char *reencoded = reencode_string_len(buf+2, size-2, "UTF-8", "UTF16", &size);
> if (!reencoded)
> die(_("could not convert contents of '%s' from UTF-16"), fname);
> free(buf);
> buf = reencoded;
> }
>
> This leads to some free thinking, especially when we look at
> other implementations of Git:
> Would it be better to simply bail out on UTF-16 files ?
> Techically all files with a '\0'.
> [snip]I was trying to cover more possible use cases, but this may not be a desired behavior after all. Other reviewers pointed out that the exclude file may have an abitrary encoding that needs to match the encoding of the paths as read by git when using readdir (3).
You are correct, UTF-16-LE-BOM is a 'fictional' encoding handled by git. Git handles the BOM and iconv will be passed the UTF-16LE encoding instead.
I would've liked to be able to handle any wrongly encoded exclude files, but it's more complicated than I originally thought. Checking for a null byte could be a simple way to detect some wrong encodings.