git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v11 06/10] convert: add 'working-tree-encoding' attribute

From
Lars Schneider <larsxschneider@gmail.com>
Date
Apr 15, 2018, 16:54 UTC
Message-ID
<8EBD3571-D5FA-471B-BD5F-D8401043D503@gmail.com>
In-Reply-To
<583f9ec3-3aef-d823-9fd6-3cc126ac47f6@web.de>
Show 19 quoted lines
> On 05 Apr 2018, at 18:41, Torsten Bögershausen <tboegi@web.de> wrote:
> 
> On 01.04.18 15:24, Lars Schneider wrote:
>>> TRUE or false are values, but just wrong ones.
>>> If this test is removed, the user will see "failed to encode "TRUE" to "UTF-8",
>>> which should give enough information to fix it.
>> 
>> I see your point. However, I would like to stop the processing right
>> there for these invalid values. How about 
>> 
>>  error(_("true/false are no valid working-tree-encodings"));
>> 
>> I think that is the most straight forward/helpful error message
>> for the enduser (I consider the term "boolean" but dismissed it
>> as potentially confusing to folks not familiar with the term).
>> 
>> OK with you?
> 
> Yes.
Great!
Show 14 quoted lines
> Another thing that came up recently, independent of your series:
> 
> What should happen if a user specifies "UTF-8" and the file
> has an UTF-8 encoded BOM ?
> I ask because I stumbled over such a file coming from a Windows
> which the java compiler under Linux didn't accept.
> 
> And because some tools love to put an UTF-8 encoded BOM
> into text files.
> 
> The clearest thing would be to extend the BOM check in 5/9
> to cover UTF-32, UTF-16 and UTF-8.
> 
> Are there any plans to do so?

If `working-tree-encoding` is not defined or defined as UTF-8, then we would return from encode_to_git() early. That means we would never run validate_encoding() which would check the BOM.

However, adding the UTF-8 BOM would still make sense. This way Git could scream if a user set `working-tree-encoding` to UTF-16 but the file is really UTF-8 encoded.

> And thanks for the work.
Thanks :-)
- Lars
Previous: Torsten BögershausenNext: lars.schneider@autodesk.com
Message 18 of 25 in “convert: add support for different encodings”
  1. 00/10 convert: add support for different encodingslars.schneider@autodesk.com, Mar 9, 2018
  2. 02/10 strbuf: add xstrdup_toupper()lars.schneider@autodesk.com, Mar 9, 2018
  3. 08/10 convert: advise canonical UTF encoding nameslars.schneider@autodesk.com, Mar 9, 2018
  4. Junio C HamanoMar 9, 2018
  5. Lars SchneiderMar 15, 2018
  6. 03/10 strbuf: add a case insensitive starts_with()lars.schneider@autodesk.com, Mar 9, 2018
  7. 09/10 convert: add tracing for 'working-tree-encoding' attributelars.schneider@autodesk.com, Mar 9, 2018
  8. 07/10 convert: check for detectable errors in UTF encodingslars.schneider@autodesk.com, Mar 9, 2018
  9. Junio C HamanoMar 9, 2018
  10. Lars SchneiderMar 9, 2018
  11. Junio C HamanoMar 9, 2018
  12. 06/10 convert: add 'working-tree-encoding' attributelars.schneider@autodesk.com, Mar 9, 2018
  13. Junio C HamanoMar 9, 2018
  14. Lars SchneiderMar 15, 2018
  15. Torsten BögershausenMar 18, 2018
  16. Lars SchneiderApr 1, 2018
  17. Torsten BögershausenApr 5, 2018
  18. Lars SchneiderApr 15, 2018
  19. 10/10 convert: add round trip check based on 'core.checkRoundtripEncoding'lars.schneider@autodesk.com, Mar 9, 2018
  20. Eric SunshineMar 9, 2018
  21. Junio C HamanoMar 9, 2018
  22. Eric SunshineMar 9, 2018
  23. 01/10 strbuf: remove unnecessary NUL assignment in xstrdup_tolower()lars.schneider@autodesk.com, Mar 9, 2018
  24. 05/10 utf8: add function to detect a missing UTF-16/32 BOMlars.schneider@autodesk.com, Mar 9, 2018
  25. 04/10 utf8: add function to detect prohibited UTF-16/32 BOMlars.schneider@autodesk.com, Mar 9, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.