git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: utf8 BOM

From
Dmitry Potapov <dpotapov@gmail.com>
Date
May 16, 2010, 05:19 UTC
Message-ID
<20100516051927.GA17200@dpotapov.dyndns.org>
In-Reply-To
<61355CFC-EB9E-4B76-9450-F2DF1B2903C0@gmail.com>
On Sat, May 15, 2010 at 10:23:52PM +0200, Eyvind Bernhardsen wrote:
Show 8 quoted lines
> On 14. mai 2010, at 12.16, Dmitry Potapov wrote:
> 
> > Probably, ability of automatic add utf8 BOM on Windows to text files
> > (which are marked as "unicode") can be helpful, but it is just a part
> > of the problem of how to deal with text files in "legacy" encoding,
> > which are still widely used on Windows.
>
> Sounds like something a clean/smudge filter should be able to do.

Yes, it should if you handful files that need such conversion. However, if you want it for every text file, running filters are slow (especially on Windows), and they are not capable to autodetect text.

> (which hopefully works no matter what your code
> page is?  I don't know much about Windows i18n).

Yes, it does. I am not an expert on Windows either, but as far as I know, BOM are used to mark unicode files, which could be either UTF-8 or UTF-16. BTW, UTF-16 are treated by Git as "binary" now, which may not always convenient, because impossible to do "merge" or "diff".

> Adding this to convert.c would be more difficult, at least
> politically, since I assume it would be Windows-specific code.

I don't think it needs any Windows-specific code. We already have some functions to convert text from different charsets, which could be used. But this feature should be developed and tested by people who work on Windows regularly and need this feature, because there is no substitute for testing and experience of how well it works in practice. Currently, I rarely use Windows and can get by clean/smudge filters.

Dmitry
Previous: Eyvind BernhardsenNext: Eyvind Bernhardsen
Message 13 of 27 in “End-of-line normalization, redesigned”
  1. 0/5 End-of-line normalization, redesignedEyvind Bernhardsen, May 12, 2010
  2. 1/5 autocrlf: Make it work also for un-normalized repositoriesEyvind Bernhardsen, May 12, 2010
  3. 2/5 Add tests for per-repository eol normalizationEyvind Bernhardsen, May 12, 2010
  4. 3/5 Add per-repository eol normalizationEyvind Bernhardsen, May 12, 2010
  5. 4/5 Rename "crlf" attribute as "eolconv"Eyvind Bernhardsen, May 12, 2010
  6. Linus TorvaldsMay 13, 2010
  7. Robert BuckMay 13, 2010
  8. Robert BuckMay 13, 2010
  9. Eyvind BernhardsenMay 13, 2010
  10. Robert BuckMay 13, 2010
  11. utf8 BOMDmitry Potapov, May 14, 2010
  12. Eyvind BernhardsenMay 15, 2010
  13. Dmitry PotapovMay 16, 2010
  14. Eyvind BernhardsenMay 16, 2010
  15. TaitMay 16, 2010
  16. Dmitry PotapovMay 16, 2010
  17. Eyvind BernhardsenMay 13, 2010
  18. Linus TorvaldsMay 13, 2010
  19. Robert BuckMay 14, 2010
  20. Jonathan NiederMay 14, 2010
  21. Eyvind BernhardsenMay 14, 2010
  22. Eyvind BernhardsenMay 14, 2010
  23. Eyvind BernhardsenMay 14, 2010
  24. Linus TorvaldsMay 14, 2010
  25. Add "core.eol" variable to control end-of-line conversionEyvind Bernhardsen, May 15, 2010
  26. Robert BuckMay 16, 2010
  27. 5/5 Rename "core.autocrlf" config variable as "core.eolconv"Eyvind Bernhardsen, May 12, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.