{"thread":{"id":"39622","subject":"[PATCH] Documentation/i18n.txt: clarify character encoding support","startedAt":"2015-06-13T20:24:01Z","lastAt":"2015-07-02T05:25:07Z","messageCount":6,"participants":["Karsten Blees","Junio C Hamano","Torsten Bögershausen"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"263764","messageId":"557C9161.6020703@gmail.com","threadId":"39622","inReplyTo":null,"subject":"[PATCH] Documentation/i18n.txt: clarify character encoding support","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2015-06-13T20:24:01Z","receivedAt":"2015-06-13T20:24:01Z","isPatch":true,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"As a \"distributed\" VCS, git should better define the encodings of its core\ntextual data structures, in particular those that are part of the network\nprotocol.\n\nThat git is encoding agnostic is only really true for blob objects. E.g.\nthe 'non-NUL bytes' requirement of tree and commit objects excludes\nUTF-16/32, and the special meaning of '/' in the index file as well as\nspace and linefeed in commit objects eliminates EBCDIC and other non-ASCII\nencodings.\n\nGit expects bytes < 0x80 to be pure ASCII, thus CJK encodings that partly\noverlap with the ASCII range are problematic as well. E.g. fmt_ident()\nremoves trailing 0x5C from user names on the assumption that it is ASCII\n'\\'. However, there are over 200 GBK double byte codes that end in 0x5C.\n\nUTF-8 as default encoding on Linux and respective path translations in the\nMac and Windows versions have established UTF-8 NFC as de-facto standard\nfor path names.\n\nUpdate the documentation in i18n.txt to reflect the current status-quo.\n\nSigned-off-by: Karsten Blees <blees@dcon.de>\n---\n Documentation/i18n.txt | 30 ++++++++++++++++++++----------\n 1 file changed, 20 insertions(+), 10 deletions(-)\n\ndiff --git a/Documentation/i18n.txt b/Documentation/i18n.txt\nindex e9a1d5d..e5f6233 100644\n--- a/Documentation/i18n.txt\n+++ b/Documentation/i18n.txt\n@@ -1,18 +1,28 @@\n-At the core level, Git is character encoding agnostic.\n-\n- - The pathnames recorded in the index and in the tree objects\n-   are treated as uninterpreted sequences of non-NUL bytes.\n-   What readdir(2) returns are what are recorded and compared\n-   with the data Git keeps track of, which in turn are expected\n-   to be what lstat(2) and creat(2) accepts.  There is no such\n-   thing as pathname encoding translation.\n+Git is to some extent character encoding agnostic.\n \n  - The contents of the blob objects are uninterpreted sequences\n    of bytes.  There is no encoding translation at the core\n    level.\n \n- - The commit log messages are uninterpreted sequences of non-NUL\n-   bytes.\n+ - Pathnames are encoded in UTF-8 normalization form C. This\n+   applies to tree objects, the index file, ref names and\n+   config files (`.git/config` (see linkgit:git-config[1]),\n+   linkgit:gitignore[5], linkgit:gitattributes[5] and\n+   linkgit:gitmodules[5]).\n+   The Mac and Windows versions automatically translate pathnames\n+   to and from UTF-8 NFC in their readdir(2), lstat(2), creat(2)\n+   etc. APIs. However, there is no such translation on other\n+   platforms. If file system APIs don't use UTF-8 (which may be\n+   file system specific), it is recommended to stick to pure\n+   ASCII file names. While Git technically supports other\n+   extended ASCII encodings at the core level, such repositories\n+   will not be portable.\n+\n+ - Commit log messages are typically encoded in UTF-8, but other\n+   extended ASCII encodings are also supported. This includes\n+   ISO-8859-x, CP125x and many others, but _not_ UTF-16/32,\n+   EBCDIC and CJK multi-byte encodings (GBK, Shift-JIS, Big5,\n+   EUC-x, CP9xx etc.).\n \n Although we encourage that the commit log messages are encoded\n in UTF-8, both the core and Git Porcelain are designed not to\n-- \n2.4.1.windows.1\n"},{"id":"263833","messageId":"xmqqmw01ltid.fsf@gitster.dls.corp.google.com","threadId":"39622","inReplyTo":"557C9161.6020703@gmail.com","subject":"Re: [PATCH] Documentation/i18n.txt: clarify character encoding support","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2015-06-15T00:12:10Z","receivedAt":"2015-06-15T00:12:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Karsten Blees <karsten.blees@gmail.com> writes:\n\n> diff --git a/Documentation/i18n.txt b/Documentation/i18n.txt\n> index e9a1d5d..e5f6233 100644\n> --- a/Documentation/i18n.txt\n> +++ b/Documentation/i18n.txt\n> @@ -1,18 +1,28 @@\n> -At the core level, Git is character encoding agnostic.\n> -\n> - - The pathnames recorded in the index and in the tree objects\n> -   are treated as uninterpreted sequences of non-NUL bytes.\n> -   What readdir(2) returns are what are recorded and compared\n> -   with the data Git keeps track of, which in turn are expected\n> -   to be what lstat(2) and creat(2) accepts.  There is no such\n> -   thing as pathname encoding translation.\n> +Git is to some extent character encoding agnostic.\n\nI do not think the removal of the text makes much sense here unless\nyou add the equivalent to the new text below.\n\n>   - The contents of the blob objects are uninterpreted sequences\n>     of bytes.  There is no encoding translation at the core\n>     level.\n>  \n> - - The commit log messages are uninterpreted sequences of non-NUL\n> -   bytes.\n> + - Pathnames are encoded in UTF-8 normalization form C. This\n\nThat is true only on some systems like OSX (with HFS+) and Windows,\nno?  BSDs in general and Linux do not do any such mangling IIRC.  I\nam OK with mangling described as a notable oddball to warn users,\nthough; i.e. not as a norm as your new text suggests but as an\nexception.\n\n> +   platforms. If file system APIs don't use UTF-8 (which may be\n> +   file system specific), it is recommended to stick to pure\n> +   ASCII file names.\n\nHmph, who endorsed such a recommendation?  It is recommended to\nstick to whatever naming scheme that would not cause troubles to\nproject participants.  If your participants all want to (and can)\nuse ISO-8859-1, we do not discourage them from doing so.\n"},{"id":"263846","messageId":"557EA421.5050706@gmail.com","threadId":"39622","inReplyTo":"xmqqmw01ltid.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH] Documentation/i18n.txt: clarify character encoding support","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2015-06-15T10:08:33Z","receivedAt":"2015-06-15T10:08:33Z","isPatch":true,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"Am 15.06.2015 um 02:12 schrieb Junio C Hamano:\n> Karsten Blees <karsten.blees@gmail.com> writes:\n> \n>> diff --git a/Documentation/i18n.txt b/Documentation/i18n.txt\n>> index e9a1d5d..e5f6233 100644\n>> --- a/Documentation/i18n.txt\n>> +++ b/Documentation/i18n.txt\n>> @@ -1,18 +1,28 @@\n>> -At the core level, Git is character encoding agnostic.\n>> -\n>> - - The pathnames recorded in the index and in the tree objects\n>> -   are treated as uninterpreted sequences of non-NUL bytes.\n>> -   What readdir(2) returns are what are recorded and compared\n>> -   with the data Git keeps track of, which in turn are expected\n>> -   to be what lstat(2) and creat(2) accepts.  There is no such\n>> -   thing as pathname encoding translation.\n>> +Git is to some extent character encoding agnostic.\n> \n> I do not think the removal of the text makes much sense here unless\n> you add the equivalent to the new text below.\n> \n>>   - The contents of the blob objects are uninterpreted sequences\n>>     of bytes.  There is no encoding translation at the core\n>>     level.\n>>  \n>> - - The commit log messages are uninterpreted sequences of non-NUL\n>> -   bytes.\n>> + - Pathnames are encoded in UTF-8 normalization form C. This\n> \n> That is true only on some systems like OSX (with HFS+) and Windows,\n> no?  BSDs in general and Linux do not do any such mangling IIRC.\n\nModern Unices don't need any such mangling because UTF-8 NFC should\nbe the default system encoding. I'm not sure for BSDs, but it has\nbeen the default on all major Linux distros for more than 10 years.\n\n> I\n> am OK with mangling described as a notable oddball to warn users,\n> though; i.e. not as a norm as your new text suggests but as an\n> exception.\n> \n\nI would guess that non-UTF-8 Unices (or file systems) are the oddball\ncase, which is why I described them last. But I could be wrong.\n\n>> +   platforms. If file system APIs don't use UTF-8 (which may be\n>> +   file system specific), it is recommended to stick to pure\n>> +   ASCII file names.\n> \n> Hmph, who endorsed such a recommendation?  It is recommended to\n> stick to whatever naming scheme that would not cause troubles to\n> project participants.  If your participants all want to (and can)\n> use ISO-8859-1, we do not discourage them from doing so.\n> \n\nISO-8859-x file names may be fine if you won't ever need to:\n- use git-web, JGit, gitk, git-gui...\n- exchange repos with \"normal\" (UTF-8) Unices, Mac and Windows systems\n- publish your work on a git hosting service (and expect file and\n  ref names to show up correctly in the web interface)\n- store the repo on Unicode-based file systems (JFS, Joliet, UDF,\n  exFat, NTFS, HFS, CIFS...)\n\nThese restrictions are not that obvious when you start a new git\nproject, and while converting file names after the fact is possible\n(e.g. using the recodetree script we shipped with Git for Windows\n1.7.10), it will destroy history.\n\nThus I think we should strongly discourage users from using anything\nbut UTF-8.\n"},{"id":"264083","messageId":"xmqqr3pa5aix.fsf@gitster.dls.corp.google.com","threadId":"39622","inReplyTo":"557EA421.5050706@gmail.com","subject":"Re: [PATCH] Documentation/i18n.txt: clarify character encoding support","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2015-06-17T20:45:42Z","receivedAt":"2015-06-17T20:45:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Karsten Blees <karsten.blees@gmail.com> writes:\n\n>> I do not think the removal of the text makes much sense here unless\n>> you add the equivalent to the new text below.\n>> \n>>>   - The contents of the blob objects are uninterpreted sequences\n>>>     of bytes.  There is no encoding translation at the core\n>>>     level.\n>>>  \n>>> - - The commit log messages are uninterpreted sequences of non-NUL\n>>> -   bytes.\n>>> + - Pathnames are encoded in UTF-8 normalization form C. This\n>> \n>> That is true only on some systems like OSX (with HFS+) and Windows,\n>> no?  BSDs in general and Linux do not do any such mangling IIRC.\n>\n> Modern Unices don't need any such mangling because UTF-8 NFC should\n> be the default system encoding. I'm not sure for BSDs, but it has\n> been the default on all major Linux distros for more than 10 years.\n\nSo?  All major distros do not have to worry (and do not even need to\nknow).  As I said,...\n\n>> I\n>> am OK with mangling described as a notable oddball to warn users,\n>> though; i.e. not as a norm as your new text suggests but as an\n>> exception.\n\n... I am OK to describe \"pathnames are mangled into UTF-8 NFC on\ncertain filesystems\" as a warning.  I am OK if we encourage the use\nof UTF-8, especially if a project wants to be forward looking\n(i.e. it may currently be a monoculture but may become cross\nplatform in the future).  I just do not want to see us saying \"you\n*must* encode your path in UTF-8 NFC\".\n\n> ISO-8859-x file names may be fine if you won't ever need to:\n> - use git-web, JGit, gitk, git-gui...\n> - exchange repos with \"normal\" (UTF-8) Unices, Mac and Windows systems\n> - publish your work on a git hosting service (and expect file and\n>   ref names to show up correctly in the web interface)\n> - store the repo on Unicode-based file systems (JFS, Joliet, UDF,\n>   exFat, NTFS, HFS, CIFS...)\n\nYes, that is exatly what I said, isn't it?  \"Use whatever works for\nyour project, we do not dictate.\"\n\n> These restrictions are not that obvious when you start a new git\n> project,...\n\nOr any project for that matter, not limited to \"git project\", no?\nPerhaps that is a moot point by now, as everything in the workd\nseems to be a \"git project\" these days.\n"},{"id":"265356","messageId":"55943B37.40101@gmail.com","threadId":"39622","inReplyTo":"xmqqr3pa5aix.fsf@gitster.dls.corp.google.com","subject":"[PATCH v2] Documentation/i18n.txt: clarify character encoding support","fromName":"Karsten Blees","fromEmail":"karsten.blees@gmail.com","sentAt":"2015-07-01T19:10:47Z","receivedAt":"2015-07-01T19:10:47Z","isPatch":true,"sender":{"key":"karsten.blees@gmail.com","avatar":"https://avatars.githubusercontent.com/u/1111200?v=4"},"body":"As a \"distributed\" VCS, git should better define the encodings of its core\ntextual data structures, in particular those that are part of the network\nprotocol.\n\nThat git is encoding agnostic is only really true for blob objects. E.g.\nthe 'non-NUL bytes' requirement of tree and commit objects excludes\nUTF-16/32, and the special meaning of '/' in the index file as well as\nspace and linefeed in commit objects eliminates EBCDIC and other non-ASCII\nencodings.\n\nGit expects bytes < 0x80 to be pure ASCII, thus CJK encodings that partly\noverlap with the ASCII range are problematic as well. E.g. fmt_ident()\nremoves trailing 0x5C from user names on the assumption that it is ASCII\n'\\'. However, there are over 200 GBK double byte codes that end in 0x5C.\n\nUTF-8 as default encoding on Linux and respective path translations in the\nMac and Windows versions have established UTF-8 NFC as de-facto standard\nfor path names.\n\nUpdate the documentation in i18n.txt to reflect the current status-quo.\n\nSigned-off-by: Karsten Blees <blees@dcon.de>\n---\n\nSorry for the delay, got swamped with other stuff...\n\nAm 17.06.2015 um 22:45 schrieb Junio C Hamano:\n> \n> ... I am OK to describe \"pathnames are mangled into UTF-8 NFC on\n> certain filesystems\" as a warning.  I am OK if we encourage the use\n> of UTF-8, especially if a project wants to be forward looking\n> (i.e. it may currently be a monoculture but may become cross\n> platform in the future).  I just do not want to see us saying \"you\n> *must* encode your path in UTF-8 NFC\".\n> \n...\n> Yes, that is exatly what I said, isn't it?  \"Use whatever works for\n> your project, we do not dictate.\"\n\n\nIMO we *have* to clearly specify an encoding. This freedom of choice\nyou're proclaiming just does not work in reality.\n\nE.g. Git for Windows prior to 1.7.10 recorded file names in Windows\nsystem encoding, which was perfectly legitimate according to the\ndocumentation. Yet we had numerous bug reports regarding file name\nencoding problems (you couldn't even share repos across different\nWindows versions, let alone with Linux / Mac / JGit...).\n\nYou cannot simply tell users that this is because of Git's superior,\nflexible design and its their own fault...except of course if you\nwant them to switch to VCSes that *do* properly define their file\nformats and network protocols - such as subversion or bazaar.\n(sorry for the sarcasm, couldn't resist)\n\nI think its important to realize that specifying an encoding is\n*not* a limitation - on the contrary: it *enables* us to do things\nthat would be impossible if file names were just \"uninterpreted\nsequences of non-NUL bytes\". This includes features that are so\nfundamental that we take them for granted, e.g. displaying file\nnames using *real* characters rather than just octal escapes.\n\n\nI've rewritten the path name paragraph to better describe the\nproblems to expect with legacy encodings. I hope you like this\nversion better.\n\nOf course, it would be nice to hear other opinions as well - this\nprobably shouldn't be a discussion between the two of us :-)\n\nKarsten\n\n\n\n Documentation/i18n.txt | 33 +++++++++++++++++++++++----------\n 1 file changed, 23 insertions(+), 10 deletions(-)\n\ndiff --git a/Documentation/i18n.txt b/Documentation/i18n.txt\nindex e9a1d5d..2dd79db 100644\n--- a/Documentation/i18n.txt\n+++ b/Documentation/i18n.txt\n@@ -1,18 +1,31 @@\n-At the core level, Git is character encoding agnostic.\n-\n- - The pathnames recorded in the index and in the tree objects\n-   are treated as uninterpreted sequences of non-NUL bytes.\n-   What readdir(2) returns are what are recorded and compared\n-   with the data Git keeps track of, which in turn are expected\n-   to be what lstat(2) and creat(2) accepts.  There is no such\n-   thing as pathname encoding translation.\n+Git is to some extent character encoding agnostic.\n \n  - The contents of the blob objects are uninterpreted sequences\n    of bytes.  There is no encoding translation at the core\n    level.\n \n- - The commit log messages are uninterpreted sequences of non-NUL\n-   bytes.\n+ - Path names are encoded in UTF-8 normalization form C. This\n+   applies to tree objects, the index file, ref names, as well as\n+   path names in command line arguments, environment variables\n+   and config files (`.git/config` (see linkgit:git-config[1]),\n+   linkgit:gitignore[5], linkgit:gitattributes[5] and\n+   linkgit:gitmodules[5]).\n++\n+Note that Git at the core level treats path names simply as\n+sequences of non-NUL bytes, there are no path name encoding\n+conversions (except on Mac and Windows). Therefore, using\n+non-ASCII path names will mostly work even on platforms and file\n+systems that use legacy extended ASCII encodings. However,\n+repositories created on such systems will not work properly on\n+UTF-8-based systems (e.g. Linux, Mac, Windows) and vice versa.\n+Additionally, many Git-based tools simply assume path names to\n+be UTF-8 and will fail to display other encodings correctly.\n+\n+ - Commit log messages are typically encoded in UTF-8, but other\n+   extended ASCII encodings are also supported. This includes\n+   ISO-8859-x, CP125x and many others, but _not_ UTF-16/32,\n+   EBCDIC and CJK multi-byte encodings (GBK, Shift-JIS, Big5,\n+   EUC-x, CP9xx etc.).\n \n Although we encourage that the commit log messages are encoded\n in UTF-8, both the core and Git Porcelain are designed not to\n-- \n2.4.3.windows.1.1.g87477f9\n"},{"id":"265379","messageId":"5594CB33.7020108@web.de","threadId":"39622","inReplyTo":"55943B37.40101@gmail.com","subject":"Re: [PATCH v2] Documentation/i18n.txt: clarify character encoding support","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2015-07-02T05:25:07Z","receivedAt":"2015-07-02T05:25:07Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 07/01/2015 09:10 PM, Karsten Blees wrote:\n>\n> Of course, it would be nice to hear other opinions as well - this\n> probably shouldn't be a discussion between the two of us :-)\n>\n> Karsten\n>\nI like this paragraf from your previous mail, I think it can go\ninto i18n.txt \"as is\":\n\nISO-8859-x file names may be fine if you won't ever need to:\n- use git-web, JGit, gitk, git-gui...\n- exchange repos with \"normal\" (UTF-8) Unices, Mac and Windows systems\n- publish your work on a git hosting service (and expect file and\n   ref names to show up correctly in the web interface)\n- store the repo on Unicode-based file systems (JFS, Joliet, UDF,\n   exFat, NTFS, HFS+, CIFS...)\n"}]}