git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [RFC PATCH] Windows: Assume all file names to be UTF-8 encoded.

From
Johannes Sixt <j.sixt@viscovery.net>
Date
Mar 2, 2009, 10:30 UTC
Message-ID
<49ABB529.1080500@viscovery.net>
In-Reply-To
<alpine.DEB.2.00.0903020941120.17877@perkele.intern.softwolves.pp.se>
Peter Krefting schrieb:
> When opening a file through open() or fopen(), the path passed is
> UTF-8 encoded.

I don't think that this assumption is valid. Whenever the Windows API has to convert between Unicode strings and char* strings, it uses the current "ANSI code page". As far as I know, the UTF-8 codepage (65001) cannot be used as the "current ANSI code page". Users will always have some code page set that is not UTF-8.

For example, if the user specifies a file name on the command line, than it will not enter git in UTF-8, but in the current "ANSI" or "OEM code page" encoding. If git prints a file name under the assumption that it is UTF-8 encoded, then it will be displayed incorrectly because the system uses a different encoding.

> Since there is no real file system abstraction beyond using stdio
> (AFAIK), I need to hack it by replacing fopen (and open). Probably
> opendir/readdir as well (might be trickier), and possibly even hack
> around main() to parse the wchar_t command-line instead of the char copy.

I think you are grossly underestimating the venture that you want to undertake here.

Please come up with a plan how you are going to deal with the various issues. File names enter and leave the system through different channels:

- the command line and terminal window
- object database (tree objects)
- opendir/readdir; opening files or directories for reading or writing

And there is probably some more... How do you treat encodings in these channels? What if the file names are not valid UTF-8? Etc.

The biggest obstacle will be that git does not have a notion of "file name encoding" - it simply treats a file name as a stream of bytes. There is no place to write an encoding. If the byte streams are regarded as having an encoding, then you can have ambiguities, mixed encodings, or invalid characters. You would have to deal with this in some way.

> This will lose all chances of Windows 9x compatibility, but I don't know
> if there are any attempts of supporting it anyway?

Windows 9x is already out of the loop. We use GetFileInformationByHandle() that is only available since Windows 2000.

-- Hannes
Previous: Peter KreftingNext: Peter Krefting
Message 2 of 26 in “Windows: Assume all file names to be UTF-8 encoded.”
  1. Windows: Assume all file names to be UTF-8 encoded.Peter Krefting, Mar 2, 2009
  2. Johannes SixtMar 2, 2009
  3. Peter KreftingMar 2, 2009
  4. Johannes SchindelinMar 2, 2009
  5. Peter KreftingMar 2, 2009
  6. Johannes SixtMar 2, 2009
  7. Peter KreftingMar 2, 2009
  8. Robin RosenbergMar 2, 2009
  9. Peter KreftingMar 2, 2009
  10. Robin RosenbergMar 2, 2009
  11. Peter KreftingMar 3, 2009
  12. Dmitry PotapovMar 3, 2009
  13. Peter KreftingMar 3, 2009
  14. Robin RosenbergMar 7, 2009
  15. Peter KreftingMar 2, 2009
  16. Thomas RastMar 2, 2009
  17. Peter KreftingMar 2, 2009
  18. Lars NoschinskiMar 3, 2009
  19. Peter KreftingMar 3, 2009
  20. Lars NoschinskiMar 3, 2009
  21. Robin RosenbergMar 3, 2009
  22. Dmitry PotapovMar 3, 2009
  23. Peter KreftingMar 3, 2009
  24. Dmitry PotapovMar 3, 2009
  25. Peter KreftingMar 4, 2009
  26. Dmitry PotapovMar 4, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.