git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] ignores: handle non UTF-8 exclude files

From
Torsten Bögershausen <tboegi@web.de>
Date
Jan 4, 2026, 17:35 UTC
Message-ID
<20260104173524.GA29867@tb-raspi4>
In-Reply-To
<pull.2157.git.git.1767478617198.gitgitgadget@gmail.com>
On Sat, Jan 03, 2026 at 10:16:57PM +0000, Matthieu Beauchamp-Boulay via GitGitGadget wrote:
> From: Matthieu Beauchamp-Boulay <matthieu.beauchamp.boulay@gmail.com>
Thanks for contributing - some comments inlie
> 
> When reading exclude files, git assumes it is encoded in UTF-8 and will
Question: The report citet below talks about ignore files.
Show 5 quoted lines
> fail to apply patterns if it isn't. This is a silent failure as no warning
> or errors are shown to the users. This is a problem that can take a while
> to diagnose as many users will not think of checking the encoding of their
> file and may believe their patterns are wrong instead. Users may also
> accidentally commit undesired files.

Note: git status is your friend. Blindly commiting without checking what is staged or not may lead to unwanted results.

Show 7 quoted lines
> 
> On Windows, this happens if a user uses Windows PowerShell to create the
> file, which results in a UTF-16LE file with a BOM.
>  This issue was discussed
> here https://github.com/git-for-windows/git/issues/3329. An example of
> where a user was confused that his exclude file was not working is cited
> https://github.com/git-for-windows/git/issues/3227.

A very short research indicates that powershell can be configured to use UTF-8. I am not a powershell user, please correct if I am wrong.

> 
> A minimal fix should at least warn the user if git cannot properly decode
> the exclude file.

I think that reading an ignore file that contains a '\0' could/should Git to complain. If someone asks my, most users are tempted to ignore warnings for different reasons. Bailing out may feel more unpolite but more clear that somethinh is wrong.

>Ideally, git would handle any given Unicode file.
That is debatable.
Show 10 quoted lines
> 
> First, check if a BOM is present. If it is, decode the file to UTF-8.
> If no BOM is detected, then try to parse the file as UTF-8. If that fails,
> attempt to decode the file using the working tree encoding of the file,
> if any. If that fails, print a warning to tell the user that the exclude
> file could not be decoded and skip the file.
> 
> This raises the issue that if the entire tree is encoded in, for example
> UTF-16BE (no BOM), then even if the encoding is given in .gitattributes,
> git would not be able to decode it.
"able to decode: Yes. But willing to do so: not with the patch, right ?
Show 9 quoted lines
> I believe that this is still
> acceptable since a warning will be emitted for the file (since it has no
> BOM, is not valid UTF-8 and no working tree encoding could be found).
> 
> One case that isn't handled is if a wrong encoding is given in the
> attributes and the exclude file has no BOM and is not UTF-8. Using
> iconv to convert an UTF16BE file to UTF-8 while specifying UTF-16LE
> yields gibberish without an error and so this case is a silent failure
> where no patterns will match.

One question is, if we should look at working_tree_encoding at all. The other one is, how much UTF-16 handling of ignore or other file should we have have in Git ? It seems that this fix is for a very special case only ?

From
https://github.com/git-for-windows/git/issues/3329
we read:
/******/
if (size > 1 && buf[0] == 0xff && buf[1] == 0xfe) {
    char *reencoded = reencode_string_len(buf, size, "UTF-8", "UTF16-LE-BOM", &size);
    if (!reencoded)
        die(_("could not convert contents of '%s' from UTF-16"), fname);
    free(buf);
    buf = reencoded;
}
/******/
(Which seems a simpler suggestion)
However,  there is no UTF-16-LE-BOM in iconv 
(at least in the majority of implementations), 
so a better approach, totaly untested, may be:
if (size >= 2 && buf[0] == 0xff && buf[1] == 0xfe) {
    char *reencoded = reencode_string_len(buf+2, size-2, "UTF-8", "UTF16", &size);
    if (!reencoded)
        die(_("could not convert contents of '%s' from UTF-16"), fname);
    free(buf);
    buf = reencoded;
}

This leads to some free thinking, especially when we look at other implementations of Git: Would it be better to simply bail out on UTF-16 files ? Techically all files with a '\0'. [snip]

Previous: Matthieu BeauchampNext: Matthieu Beauchamp
Message 4 of 13 in “ignores: handle non UTF-8 exclude files”
  1. ignores: handle non UTF-8 exclude filesMatthieu Beauchamp-Boulay via GitGitGadget, Jan 3, 2026
  2. Junio C HamanoJan 4, 2026
  3. Matthieu BeauchampJan 6, 2026
  4. Torsten BögershausenJan 4, 2026
  5. Matthieu BeauchampJan 6, 2026
  6. Phillip WoodJan 7, 2026
  7. brian m. carlsonJan 4, 2026
  8. Matthieu BeauchampJan 6, 2026
  9. brian m. carlsonJan 6, 2026
  10. Collin FunkJan 7, 2026
  11. Phillip WoodJan 7, 2026
  12. brian m. carlsonJan 7, 2026
  13. Collin FunkJan 8, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.