{"thread":{"id":"23850","subject":"Git, Mac OS X and German special characters","startedAt":"2010-05-20T07:26:03Z","lastAt":"2010-05-20T18:22:22Z","messageCount":13,"participants":["Matthias Moeller","Ævar Arnfjörð Bjarmason","Michael J Gruber","demerphq","Torsten Bögershausen","Thomas Singer","Jay Soffian"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"141935","messageId":"4BF4E40B.30205@math.tu-dortmund.de","threadId":"23850","inReplyTo":null,"subject":"Git, Mac OS X and German special characters","fromName":"Matthias Moeller","fromEmail":"matthias.moeller@math.tu-dortmund.de","sentAt":"2010-05-20T07:26:03Z","receivedAt":"2010-05-20T07:26:03Z","isPatch":false,"sender":{"key":"matthias.moeller@math.tu-dortmund.de","avatar":null},"body":"Hi,\n\nI have been using git (version 1.7.1) for quite some time to\nbackup/synchronize folders between my two linux workstations at home and\nat work and my Apple laptop. Synchronization works well except for some\nstrange behavior under Mac OS X (10.6).\n\nI have commited some files with German special characters, say,\n\"Übersicht.xls\" under linux and pulled them on the Macbook. However, on\nthe laptop git status says:\n\n# On branch master\n# Untracked files:\n#   (use \"git add <file>...\" to include in what will be committed)\n#\n#       \"U\\314\\210bersicht.xls\"\nnothing added to commit but untracked files present (use \"git add\" to track)\n\nI have been searching the web for help and found lengthy discussions\nwhich state that this is a common problem of the HFS+ filesystem.\nWhat I did not find was a solution to this problem. Is there a solution\nto this problem?\n\nThanks,\nMatthias\n"},{"id":"141938","messageId":"AANLkTimYgkv6q6fTXqNOCq1ZbodxgCZ18Fum_NryyiO8@mail.gmail.com","threadId":"23850","inReplyTo":"4BF4E40B.30205@math.tu-dortmund.de","subject":"Re: Git, Mac OS X and German special characters","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2010-05-20T08:34:47Z","receivedAt":"2010-05-20T08:34:47Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Thu, May 20, 2010 at 07:26, Matthias Moeller\n<matthias.moeller@math.tu-dortmund.de> wrote:\n> I have been searching the web for help and found lengthy discussions\n> which state that this is a common problem of the HFS+ filesystem.\n> What I did not find was a solution to this problem. Is there a solution\n> to this problem?\n\nIs this problem particular to Git, or do you also get it if you\ne.g. rsync from the Linux box to the Mac OS X box?\n\n> #       \"U\\314\\210bersicht.xls\"\n\nYou probably have to configure your shell on OSX to render UTF-8\ncorrectly. It's just showing the raw escaped byte sequence instead of\na character there.\n\nThere isn't anything wrong with OSX in this case, filename encoding on\nany POSIX system is only done by convention. You'll find that you have\nsimilar problems on Linux if you encode filename in Big5 or\nUTF-32.\n\nLinux will happily accept it, but your shell / other applications will\nrender it as unknown goo because they expect UTF-8.\n"},{"id":"141940","messageId":"4BF4F7D7.60002@drmicha.warpmail.net","threadId":"23850","inReplyTo":"AANLkTimYgkv6q6fTXqNOCq1ZbodxgCZ18Fum_NryyiO8@mail.gmail.com","subject":"Re: Git, Mac OS X and German special characters","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2010-05-20T08:50:31Z","receivedAt":"2010-05-20T08:50:31Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Ævar Arnfjörð Bjarmason venit, vidit, dixit 20.05.2010 10:34:\n> On Thu, May 20, 2010 at 07:26, Matthias Moeller\n> <matthias.moeller@math.tu-dortmund.de> wrote:\n>> I have been searching the web for help and found lengthy discussions\n>> which state that this is a common problem of the HFS+ filesystem.\n>> What I did not find was a solution to this problem. Is there a solution\n>> to this problem?\n> \n> Is this problem particular to Git, or do you also get it if you\n> e.g. rsync from the Linux box to the Mac OS X box?\n> \n>> #       \"U\\314\\210bersicht.xls\"\n> \n> You probably have to configure your shell on OSX to render UTF-8\n> correctly. It's just showing the raw escaped byte sequence instead of\n> a character there.\n> \n> There isn't anything wrong with OSX in this case, filename encoding on\n> any POSIX system is only done by convention. You'll find that you have\n> similar problems on Linux if you encode filename in Big5 or\n> UTF-32.\n> \n> Linux will happily accept it, but your shell / other applications will\n> render it as unknown goo because they expect UTF-8.\n\nNo, the problem with git status is not the display. Matthias' problem is\nthat git status reports a tracked file as untracked. The reason is that\non HFS+, you create a file with name A and get a file with name B, where\nA and B are different representations of the same name. There seems to\nbe no way to reliably detect which one HFS+ uses.\n\nMichael\n"},{"id":"141941","messageId":"AANLkTimqn18y8_fB0PCUNY3XTD8nneDsEz4Koimr-SxM@mail.gmail.com","threadId":"23850","inReplyTo":"AANLkTimYgkv6q6fTXqNOCq1ZbodxgCZ18Fum_NryyiO8@mail.gmail.com","subject":"Re: Git, Mac OS X and German special characters","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2010-05-20T08:55:38Z","receivedAt":"2010-05-20T08:55:38Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On 20 May 2010 10:34, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> On Thu, May 20, 2010 at 07:26, Matthias Moeller\n> <matthias.moeller@math.tu-dortmund.de> wrote:\n>> I have been searching the web for help and found lengthy discussions\n>> which state that this is a common problem of the HFS+ filesystem.\n>> What I did not find was a solution to this problem. Is there a solution\n>> to this problem?\n>\n> Is this problem particular to Git, or do you also get it if you\n> e.g. rsync from the Linux box to the Mac OS X box?\n>\n>> #       \"U\\314\\210bersicht.xls\"\n>\n> You probably have to configure your shell on OSX to render UTF-8\n> correctly. It's just showing the raw escaped byte sequence instead of\n> a character there.\n>\n> There isn't anything wrong with OSX in this case, filename encoding on\n> any POSIX system is only done by convention. You'll find that you have\n> similar problems on Linux if you encode filename in Big5 or\n> UTF-32.\n>\n> Linux will happily accept it, but your shell / other applications will\n> render it as unknown goo because they expect UTF-8.\n\nExcept that isnt a normalized utf8 representation of capital U umlaut,\ncode point U+00DC, (utf8 c3,9c), instead presumably it has been\ndecomposed into a captial U followed by a combining character to add\nin the umlaut, which IMO is pretty weird.\n\nAs far as i can tell the filename:\n\n\"Übersicht.xls\"\n\nShould be stored in utf8 as:\n\n\"\\303\\234bersicht.xls\"\n\nAlso minor nit. UTF-32, as it contains nuls for latin-1 chars would be\nmuch much worse than utf8 :-)\n\ncheers,\nYves\n\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"141942","messageId":"AANLkTimKW-eC--JtjOzclfCymhOTizboISDygvTvVbUY@mail.gmail.com","threadId":"23850","inReplyTo":"4BF4F7D7.60002@drmicha.warpmail.net","subject":"Re: Git, Mac OS X and German special characters","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2010-05-20T08:57:59Z","receivedAt":"2010-05-20T08:57:59Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On 20 May 2010 10:50, Michael J Gruber <git@drmicha.warpmail.net> wrote:\n> Ævar Arnfjörð Bjarmason venit, vidit, dixit 20.05.2010 10:34:\n>> On Thu, May 20, 2010 at 07:26, Matthias Moeller\n>> <matthias.moeller@math.tu-dortmund.de> wrote:\n>>> I have been searching the web for help and found lengthy discussions\n>>> which state that this is a common problem of the HFS+ filesystem.\n>>> What I did not find was a solution to this problem. Is there a solution\n>>> to this problem?\n>>\n>> Is this problem particular to Git, or do you also get it if you\n>> e.g. rsync from the Linux box to the Mac OS X box?\n>>\n>>> #       \"U\\314\\210bersicht.xls\"\n>>\n>> You probably have to configure your shell on OSX to render UTF-8\n>> correctly. It's just showing the raw escaped byte sequence instead of\n>> a character there.\n>>\n>> There isn't anything wrong with OSX in this case, filename encoding on\n>> any POSIX system is only done by convention. You'll find that you have\n>> similar problems on Linux if you encode filename in Big5 or\n>> UTF-32.\n>>\n>> Linux will happily accept it, but your shell / other applications will\n>> render it as unknown goo because they expect UTF-8.\n>\n> No, the problem with git status is not the display. Matthias' problem is\n> that git status reports a tracked file as untracked. The reason is that\n> on HFS+, you create a file with name A and get a file with name B, where\n> A and B are different representations of the same name. There seems to\n> be no way to reliably detect which one HFS+ uses.\n\nJudging by the example given the problem is that HFS+ decomposes\nUnicode file names into latin1+combining characters instead of using\nnormalized utf8.\n\nThis implies that if the utf8 is normalized first using canonical\nunicode normalization rules (to eliminate the combining character)\nthat it can then be compared.\n\nYves\n\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"141943","messageId":"4BF4FA89.2040904@gmail.com","threadId":"23850","inReplyTo":"4BF4F7D7.60002@drmicha.warpmail.net","subject":"Re: Git, Mac OS X and German special characters","fromName":"Torsten Bögershausen","fromEmail":"totte.enea@gmail.com","sentAt":"2010-05-20T09:02:01Z","receivedAt":"2010-05-20T09:02:01Z","isPatch":false,"sender":{"key":"totte.enea@gmail.com","avatar":null},"body":"Hej,\nI have the same problem here.\nBelow there is a patch, which may solve the problem.\n(Yes, whitespaces are broken. I'm still fighting with\ngit format-patch -s --cover-letter -M --stdout origin/master | git \nimap-send)\nBut this patch may be a start point for improvements.\nComments welcome\nBR\n/Torsten\n\n\n\nImproved interwork between Mac OS X and linux when umlauts are used\nWhen a git repository containing utf-8 coded umlaut characters\nis cloned onto an Mac OS X machine, the Mac OS system will convert\nall filenames returned by readdir() into denormalized utf-8.\nAs a result of this conversion, git will not find them on disk.\nThis helps by treating the NFD and NFD version of filenames as\nidentical on Mac OS.\n\n\n\n\n\n\nSigned-off-by: Torsten Bögershausen <tboegi@web.de>\n---\nname-hash.c |   40 ++++++++++++++++++++++++++++++++++++++++\nutf8.c      |   55 ++++++++++++++++++++++++++++++++++++++++++++++++-------\nutf8.h      |   11 +++++++++++\n3 files changed, 99 insertions(+), 7 deletions(-)\n\ndiff --git a/name-hash.c b/name-hash.c\nindex 0031d78..e6494e8 100644\n--- a/name-hash.c\n+++ b/name-hash.c\n@@ -7,6 +7,7 @@\n  */\n#define NO_THE_INDEX_COMPATIBILITY_MACROS\n#include \"cache.h\"\n+#include \"utf8.h\"\n\n/*\n  * This removes bit 5 if bit 6 is set.\n@@ -100,6 +101,25 @@ static int same_name(const struct cache_entry *ce, \nconst char *name, int namelen\n     return icase && slow_same_name(name, namelen, ce->name, len);\n}\n\n+#ifdef __APPLE__\n+struct cache_entry *index_name_exists2(struct index_state *istate, \nconst char *name, int icase)\n+{\n+    int namelen = (int)strlen(name);\n+    unsigned int hash = hash_name(name, namelen);\n+    struct cache_entry *ce;\n+\n+    ce = lookup_hash(hash, &istate->name_hash);\n+    while (ce) {\n+        if (!(ce->ce_flags & CE_UNHASHED)) {\n+            if (same_name(ce, name, namelen, icase))\n+                return ce;\n+        }\n+        ce = ce->next;\n+    }\n+    return NULL;\n+}\n+#endif\n+\nstruct cache_entry *index_name_exists(struct index_state *istate, const \nchar *name, int namelen, int icase)\n{\n     unsigned int hash = hash_name(name, namelen);\n@@ -115,5 +135,25 @@ struct cache_entry *index_name_exists(struct \nindex_state *istate, const char *na\n         }\n         ce = ce->next;\n     }\n+#ifdef __APPLE__\n+    {\n+        char *name_nfc_nfd;\n+        name_nfc_nfd = str_nfc2nfd(name);\n+        if (name_nfc_nfd) {\n+            ce = index_name_exists2(istate, name_nfc_nfd, icase);\n+            free(name_nfc_nfd);\n+            if (ce)\n+                return ce;\n+        }\n+        name_nfc_nfd = str_nfd2nfc(name);\n+        if (name_nfc_nfd) {\n+            ce = index_name_exists2(istate, name_nfc_nfd, icase);\n+            free(name_nfc_nfd);\n+            if (ce)\n+                return ce;\n+        }\n+    }\n+#endif\n+\n     return NULL;\n}\ndiff --git a/utf8.c b/utf8.c\nindex 84cfc72..8e794dc 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -2,6 +2,11 @@\n#include \"strbuf.h\"\n#include \"utf8.h\"\n\n+#ifdef __APPLE__\n+static iconv_t my_iconv_nfd2nfc = (iconv_t) -1;\n+static iconv_t my_iconv_nfc2nfd = (iconv_t) -1;\n+#endif\n+\n/* This code is originally from http://www.cl.cam.ac.uk/~mgk25/ucs/ */\n\nstruct interval {\n@@ -424,18 +429,13 @@ int is_encoding_utf8(const char *name)\n#else\n     typedef char * iconv_ibp;\n#endif\n-char *reencode_string(const char *in, const char *out_encoding, const \nchar *in_encoding)\n+\n+char *reencode_string_iconv(const char *in, iconv_t conv)\n{\n-    iconv_t conv;\n     size_t insz, outsz, outalloc;\n     char *out, *outpos;\n     iconv_ibp cp;\n\n-    if (!in_encoding)\n-        return NULL;\n-    conv = iconv_open(out_encoding, in_encoding);\n-    if (conv == (iconv_t) -1)\n-        return NULL;\n     insz = strlen(in);\n     outsz = insz;\n     outalloc = outsz + 1; /* for terminating NUL */\n@@ -469,7 +469,48 @@ char *reencode_string(const char *in, const char \n*out_encoding, const char *in_e\n             break;\n         }\n     }\n+    return out;\n+}\n+\n+char *reencode_string(const char *in, const char *out_encoding, const \nchar *in_encoding)\n+{\n+    iconv_t conv;\n+    char *out;\n+\n+    if (!in_encoding)\n+        return NULL;\n+    conv = iconv_open(out_encoding, in_encoding);\n+    if (conv == (iconv_t) -1)\n+        return NULL;\n+    out = reencode_string_iconv(in, conv);\n     iconv_close(conv);\n     return out;\n}\n+\n+#ifdef __APPLE__\n+char*\n+str_nfc2nfd(const char *in)\n+{\n+    if (my_iconv_nfc2nfd == (iconv_t) -1) {\n+        my_iconv_nfc2nfd = iconv_open(\"utf-8-mac\", \"utf-8\");\n+        if (my_iconv_nfc2nfd == (iconv_t) -1) {\n+            return NULL;\n+        }\n+    }\n+    return reencode_string_iconv(in, my_iconv_nfc2nfd);\n+}\n+\n+char*\n+str_nfd2nfc(const char *in)\n+{\n+    if (my_iconv_nfd2nfc == (iconv_t) -1){\n+        my_iconv_nfd2nfc = iconv_open(\"utf-8\", \"utf-8-mac\");\n+        if (my_iconv_nfd2nfc == (iconv_t) -1) {\n+            return NULL;\n+        }\n+    }\n+    return reencode_string_iconv(in, my_iconv_nfd2nfc);\n+}\n+#endif /* APPLE */\n+\n#endif\ndiff --git a/utf8.h b/utf8.h\nindex ebc4d2f..db29c8a 100644\n--- a/utf8.h\n+++ b/utf8.h\n@@ -13,8 +13,19 @@ int strbuf_add_wrapped_text(struct strbuf *buf,\n\n#ifndef NO_ICONV\nchar *reencode_string(const char *in, const char *out_encoding, const \nchar *in_encoding);\n+char *reencode_string_iconv(const char *in, iconv_t conv);\n+#ifdef __APPLE__\n+char *str_nfc2nfd(const char *in);\n+char *str_nfd2nfc(const char *in);\n+#else\n+#define str_nfc2nfd(in) (NULL)\n+#define str_nfd2nfc(in) (NULL)\n+#endif\n#else\n#define reencode_string(a,b,c) NULL\n+#define reencode_string2(a,b) NULL\n+#define str_nfc2nfd(in) (NULL)\n+#define str_nfd2nfc(in) (NULL)\n#endif\n\n#endif\n-- \n1.7.1.dirty\n\n\n\n\n\n\n\n\n\n\nOn 20.05.10 10:50, Michael J Gruber wrote:\n> Ævar Arnfjörð Bjarmason venit, vidit, dixit 20.05.2010 10:34:\n>    \n>> On Thu, May 20, 2010 at 07:26, Matthias Moeller\n>> <matthias.moeller@math.tu-dortmund.de>  wrote:\n>>      \n>>> I have been searching the web for help and found lengthy discussions\n>>> which state that this is a common problem of the HFS+ filesystem.\n>>> What I did not find was a solution to this problem. Is there a solution\n>>> to this problem?\n>>>        \n>> Is this problem particular to Git, or do you also get it if you\n>> e.g. rsync from the Linux box to the Mac OS X box?\n>>\n>>      \n>>> #       \"U\\314\\210bersicht.xls\"\n>>>        \n>> You probably have to configure your shell on OSX to render UTF-8\n>> correctly. It's just showing the raw escaped byte sequence instead of\n>> a character there.\n>>\n>> There isn't anything wrong with OSX in this case, filename encoding on\n>> any POSIX system is only done by convention. You'll find that you have\n>> similar problems on Linux if you encode filename in Big5 or\n>> UTF-32.\n>>\n>> Linux will happily accept it, but your shell / other applications will\n>> render it as unknown goo because they expect UTF-8.\n>>      \n> No, the problem with git status is not the display. Matthias' problem is\n> that git status reports a tracked file as untracked. The reason is that\n> on HFS+, you create a file with name A and get a file with name B, where\n> A and B are different representations of the same name. There seems to\n> be no way to reliably detect which one HFS+ uses.\n>\n> Michael\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>    \n"},{"id":"141944","messageId":"4BF4FDB4.2010409@drmicha.warpmail.net","threadId":"23850","inReplyTo":"4BF4FA89.2040904@gmail.com","subject":"Re: Git, Mac OS X and German special characters","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2010-05-20T09:15:32Z","receivedAt":"2010-05-20T09:15:32Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Torsten Bögershausen venit, vidit, dixit 20.05.2010 11:02:\n> Hej,\n> I have the same problem here.\n> Below there is a patch, which may solve the problem.\n> (Yes, whitespaces are broken. I'm still fighting with\n> git format-patch -s --cover-letter -M --stdout origin/master | git \n> imap-send)\n> But this patch may be a start point for improvements.\n> Comments welcome\n> BR\n> /Torsten\n> \n> \n> \n> Improved interwork between Mac OS X and linux when umlauts are used\n> When a git repository containing utf-8 coded umlaut characters\n> is cloned onto an Mac OS X machine, the Mac OS system will convert\n> all filenames returned by readdir() into denormalized utf-8.\n> As a result of this conversion, git will not find them on disk.\n> This helps by treating the NFD and NFD version of filenames as\n> identical on Mac OS.\n> \n> \n> \n> \n> \n> \n> Signed-off-by: Torsten Bögershausen <tboegi@web.de>\n\nYou signed off, but is Markus Kuhn's code from UCS GPL2-licensed?\nAlso, a few tests would be nice.\n\nI remember we had threads on this issue in the past. I haven't checked\nyet (Thunderbird pruned my nntp history), but it is worth checking that\nyou addressed any issues mentioned there.\n\nI have no Mac so I can't test, sorry. Would be happy to run Mac OS in a\nvm, but you know...\n\nThanks for looking into this!\n\nMichael\n\n> ---\n> name-hash.c |   40 ++++++++++++++++++++++++++++++++++++++++\n> utf8.c      |   55 ++++++++++++++++++++++++++++++++++++++++++++++++-------\n> utf8.h      |   11 +++++++++++\n> 3 files changed, 99 insertions(+), 7 deletions(-)\n> \n> diff --git a/name-hash.c b/name-hash.c\n> index 0031d78..e6494e8 100644\n> --- a/name-hash.c\n> +++ b/name-hash.c\n> @@ -7,6 +7,7 @@\n>   */\n> #define NO_THE_INDEX_COMPATIBILITY_MACROS\n> #include \"cache.h\"\n> +#include \"utf8.h\"\n> \n> /*\n>   * This removes bit 5 if bit 6 is set.\n> @@ -100,6 +101,25 @@ static int same_name(const struct cache_entry *ce, \n> const char *name, int namelen\n>      return icase && slow_same_name(name, namelen, ce->name, len);\n> }\n> \n> +#ifdef __APPLE__\n> +struct cache_entry *index_name_exists2(struct index_state *istate, \n> const char *name, int icase)\n> +{\n> +    int namelen = (int)strlen(name);\n> +    unsigned int hash = hash_name(name, namelen);\n> +    struct cache_entry *ce;\n> +\n> +    ce = lookup_hash(hash, &istate->name_hash);\n> +    while (ce) {\n> +        if (!(ce->ce_flags & CE_UNHASHED)) {\n> +            if (same_name(ce, name, namelen, icase))\n> +                return ce;\n> +        }\n> +        ce = ce->next;\n> +    }\n> +    return NULL;\n> +}\n> +#endif\n> +\n> struct cache_entry *index_name_exists(struct index_state *istate, const \n> char *name, int namelen, int icase)\n> {\n>      unsigned int hash = hash_name(name, namelen);\n> @@ -115,5 +135,25 @@ struct cache_entry *index_name_exists(struct \n> index_state *istate, const char *na\n>          }\n>          ce = ce->next;\n>      }\n> +#ifdef __APPLE__\n> +    {\n> +        char *name_nfc_nfd;\n> +        name_nfc_nfd = str_nfc2nfd(name);\n> +        if (name_nfc_nfd) {\n> +            ce = index_name_exists2(istate, name_nfc_nfd, icase);\n> +            free(name_nfc_nfd);\n> +            if (ce)\n> +                return ce;\n> +        }\n> +        name_nfc_nfd = str_nfd2nfc(name);\n> +        if (name_nfc_nfd) {\n> +            ce = index_name_exists2(istate, name_nfc_nfd, icase);\n> +            free(name_nfc_nfd);\n> +            if (ce)\n> +                return ce;\n> +        }\n> +    }\n> +#endif\n> +\n>      return NULL;\n> }\n> diff --git a/utf8.c b/utf8.c\n> index 84cfc72..8e794dc 100644\n> --- a/utf8.c\n> +++ b/utf8.c\n> @@ -2,6 +2,11 @@\n> #include \"strbuf.h\"\n> #include \"utf8.h\"\n> \n> +#ifdef __APPLE__\n> +static iconv_t my_iconv_nfd2nfc = (iconv_t) -1;\n> +static iconv_t my_iconv_nfc2nfd = (iconv_t) -1;\n> +#endif\n> +\n> /* This code is originally from http://www.cl.cam.ac.uk/~mgk25/ucs/ */\n> \n> struct interval {\n> @@ -424,18 +429,13 @@ int is_encoding_utf8(const char *name)\n> #else\n>      typedef char * iconv_ibp;\n> #endif\n> -char *reencode_string(const char *in, const char *out_encoding, const \n> char *in_encoding)\n> +\n> +char *reencode_string_iconv(const char *in, iconv_t conv)\n> {\n> -    iconv_t conv;\n>      size_t insz, outsz, outalloc;\n>      char *out, *outpos;\n>      iconv_ibp cp;\n> \n> -    if (!in_encoding)\n> -        return NULL;\n> -    conv = iconv_open(out_encoding, in_encoding);\n> -    if (conv == (iconv_t) -1)\n> -        return NULL;\n>      insz = strlen(in);\n>      outsz = insz;\n>      outalloc = outsz + 1; /* for terminating NUL */\n> @@ -469,7 +469,48 @@ char *reencode_string(const char *in, const char \n> *out_encoding, const char *in_e\n>              break;\n>          }\n>      }\n> +    return out;\n> +}\n> +\n> +char *reencode_string(const char *in, const char *out_encoding, const \n> char *in_encoding)\n> +{\n> +    iconv_t conv;\n> +    char *out;\n> +\n> +    if (!in_encoding)\n> +        return NULL;\n> +    conv = iconv_open(out_encoding, in_encoding);\n> +    if (conv == (iconv_t) -1)\n> +        return NULL;\n> +    out = reencode_string_iconv(in, conv);\n>      iconv_close(conv);\n>      return out;\n> }\n> +\n> +#ifdef __APPLE__\n> +char*\n> +str_nfc2nfd(const char *in)\n> +{\n> +    if (my_iconv_nfc2nfd == (iconv_t) -1) {\n> +        my_iconv_nfc2nfd = iconv_open(\"utf-8-mac\", \"utf-8\");\n> +        if (my_iconv_nfc2nfd == (iconv_t) -1) {\n> +            return NULL;\n> +        }\n> +    }\n> +    return reencode_string_iconv(in, my_iconv_nfc2nfd);\n> +}\n> +\n> +char*\n> +str_nfd2nfc(const char *in)\n> +{\n> +    if (my_iconv_nfd2nfc == (iconv_t) -1){\n> +        my_iconv_nfd2nfc = iconv_open(\"utf-8\", \"utf-8-mac\");\n> +        if (my_iconv_nfd2nfc == (iconv_t) -1) {\n> +            return NULL;\n> +        }\n> +    }\n> +    return reencode_string_iconv(in, my_iconv_nfd2nfc);\n> +}\n> +#endif /* APPLE */\n> +\n> #endif\n> diff --git a/utf8.h b/utf8.h\n> index ebc4d2f..db29c8a 100644\n> --- a/utf8.h\n> +++ b/utf8.h\n> @@ -13,8 +13,19 @@ int strbuf_add_wrapped_text(struct strbuf *buf,\n> \n> #ifndef NO_ICONV\n> char *reencode_string(const char *in, const char *out_encoding, const \n> char *in_encoding);\n> +char *reencode_string_iconv(const char *in, iconv_t conv);\n> +#ifdef __APPLE__\n> +char *str_nfc2nfd(const char *in);\n> +char *str_nfd2nfc(const char *in);\n> +#else\n> +#define str_nfc2nfd(in) (NULL)\n> +#define str_nfd2nfc(in) (NULL)\n> +#endif\n> #else\n> #define reencode_string(a,b,c) NULL\n> +#define reencode_string2(a,b) NULL\n> +#define str_nfc2nfd(in) (NULL)\n> +#define str_nfd2nfc(in) (NULL)\n> #endif\n> \n> #endif\n"},{"id":"141945","messageId":"4BF4FDD9.9090500@math.tu-dortmund.de","threadId":"23850","inReplyTo":"4BF4F7D7.60002@drmicha.warpmail.net","subject":"Re: Git, Mac OS X and German special characters","fromName":"Matthias Moeller","fromEmail":"matthias.moeller@math.tu-dortmund.de","sentAt":"2010-05-20T09:16:09Z","receivedAt":"2010-05-20T09:16:09Z","isPatch":false,"sender":{"key":"matthias.moeller@math.tu-dortmund.de","avatar":null},"body":"On 05/20/2010 10:50 AM, Michael J Gruber wrote:\n>\n>> Is this problem particular to Git, or do you also get it if you\n>> e.g. rsync from the Linux box to the Mac OS X box?\n>>     \n> No, the problem with git status is not the display. Matthias' problem is\n> that git status reports a tracked file as untracked. The reason is that\n> on HFS+, you create a file with name A and get a file with name B, where\n> A and B are different representations of the same name. There seems to\n> be no way to reliably detect which one HFS+ uses.\n>   \n\nYes, the problem is not the display but the filesystem. I had similar\nproblems with unison some time ago.\nBut there was a special fix for utf-8 and Mac OS X in one of the newer\nunison versions.\n\nMatthias\n"},{"id":"141957","messageId":"4BF51133.6040504@syntevo.com","threadId":"23850","inReplyTo":"4BF4F7D7.60002@drmicha.warpmail.net","subject":"Re: Git, Mac OS X and German special characters","fromName":"Thomas Singer","fromEmail":"thomas.singer@syntevo.com","sentAt":"2010-05-20T10:38:43Z","receivedAt":"2010-05-20T10:38:43Z","isPatch":false,"sender":{"key":"thomas.singer@syntevo.com","avatar":null},"body":"On 20.05.2010 10:50, Michael J Gruber wrote:\n> There seems to be no way to reliably detect which one HFS+ uses.\n\nIIRC, HFS+ always stores and reports the file names with decomposed UTF-8,\neven if the file is created using composed UTF-8. IMHO, Git should\nstandardize on the file and text encoding (e.g. commit messages) used in the\nrepository, so such problems can't occur. SVN has standardized on \"UTF-8\" in\nthe repository, but had/s similar problems on OS X with the decomposed\ncharacters:\n\n http://subversion.tigris.org/issues/show_bug.cgi?id=2464\n\nTom\n"},{"id":"141966","messageId":"4BF54744.60602@drmicha.warpmail.net","threadId":"23850","inReplyTo":"4BF5294E.7060206@web.de","subject":"Re: Git, Mac OS X and German special characters","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2010-05-20T14:29:24Z","receivedAt":"2010-05-20T14:29:24Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Torsten Bögershausen venit, vidit, dixit 20.05.2010 14:21:\n> Hej Michael,\n> Thanks for the reply.\n> \n>> You signed off, but is Markus Kuhn's code from UCS GPL2-licensed?\n> Oh, I haven't added any code from Markus here.\n> But if my sign off is a problem, we can remove it ;-)\n> or move the code to another place. (And utf.c will still have code from UCS)\n\nYour sign-off is fine if you can place the code under the terms of the\nproject.\n\nIn your patch there is a line\n\n/* This code is originally from http://www.cl.cam.ac.uk/~mgk25/ucs/ */\n\nbut I missed the missing '+' in front - that comment was there before\nyour patch! Sorry for the confusion.\n\n> \n>> Also, a few tests would be nice.\n> Yes, fully agreed.\n> My feeling is, that at least\n> \"git add\", \"git mv\", \"git rm\" should be tested.\n> I will fix that.\n> But as I become more familiar with the git testsuite,\n> it becomes more and more clear, that testing the new feature will\n> do the same tests as already existing tests.\n>  From that point of view, it seems easier to re-use the existing test\n> cases and run them twice, once with clean ascii, and second time with\n> an internationalized form.\n> As not all platforms support utf-8, the internationalized tests may be\n> either utf-8, 8859-1, or nothing at all.\n>   \n> I feel that at least 50% of the test cases should be \"internationalized\",\n> like \"git merge\", \"git pull\" etc.\n> (And re-writing the tests is a big issue, at least for me as a beginner)\n> Anyway, I will make simple tests.\n\nSimple tests are a good start. More importantly (compared to full\ninternationalisation), we need someone running them on Mac OS ;)\n\nMichael\n"},{"id":"141968","messageId":"AANLkTilvz-n-oBzfCChh0qGo3oWVXaB7CZY4562ineDn@mail.gmail.com","threadId":"23850","inReplyTo":"4BF4FDB4.2010409@drmicha.warpmail.net","subject":"Re: Git, Mac OS X and German special characters","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2010-05-20T15:30:37Z","receivedAt":"2010-05-20T15:30:37Z","isPatch":false,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"2010/5/20 Michael J Gruber <git@drmicha.warpmail.net>:\n> I remember we had threads on this issue in the past. I haven't checked\n> yet (Thunderbird pruned my nntp history), but it is worth checking that\n> you addressed any issues mentioned there.\n\nThis was the monster thread on it:\n\n  http://thread.gmane.org/gmane.comp.version-control.git/70688\n\nLinus added support for the case-insensitivity aliasing issue:\n\n  1102952 (Make git-add behave more sensibly in a case-insensitive environment,\n           2008-03-22)\n  6835550 (When adding files to the index, add support for\ncase-independent matches,\n           2008-03-22)\n\nBut he couldn't care less about HFS+ brain-damage, so he left that for\nothers. And hey, it's only taken 2 years for someone to step up to the\nplate. :-)\n\nj.\n"},{"id":"141969","messageId":"AANLkTil2N2xP1CWj0xxskOn-KCN1JMpJS8d3WpT5Mdg2@mail.gmail.com","threadId":"23850","inReplyTo":"4BF4FA89.2040904@gmail.com","subject":"Re: Git, Mac OS X and German special characters","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2010-05-20T15:50:16Z","receivedAt":"2010-05-20T15:50:16Z","isPatch":false,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"2010/5/20 Torsten Bögershausen <totte.enea@gmail.com>:\n> Improved interwork between Mac OS X and linux when umlauts are used\n> When a git repository containing utf-8 coded umlaut characters\n> is cloned onto an Mac OS X machine, the Mac OS system will convert\n> all filenames returned by readdir() into denormalized utf-8.\n> As a result of this conversion, git will not find them on disk.\n> This helps by treating the NFD and NFD version of filenames as\n> identical on Mac OS.\n\nSo this is an edge case, but what happens if a repo has both the NFC\nand NFD representation of a given name? It should be handled the same\nway as if a repo has both \"File\" and \"file\" and you try to check-out\nonto a case-insensitive filesystem.\n\nAdditionally, note this paragraph from 1102952 (Make git-add behave\nmore sensibly in a case-insensitive environment, 2008-03-22):\n\n    However, if we actually have *both* a file called \"File\" and one called\n    \"file\", and they don't have the same lstat() information (ie we're on a\n    case-sensitive filesystem but have the \"core.ignorecase\" flag set), we\n    will error out if we try to add them both.\n\nTo be consistent, shouldn't we have a core.HFSPlusCompat that can be\nset on non-braindamaged filesystems to prevent filenames which would\nalias on HFS+ from entering the repo?\n\n   http://developer.apple.com/mac/library/technotes/tn/tn1150.html#UnicodeSubtleties\n\nj.\n"},{"id":"141979","messageId":"AANLkTilIqY0760KpZbGrv3jyZLT1e0CbkvKF3lbYB5la@mail.gmail.com","threadId":"23850","inReplyTo":"AANLkTil2N2xP1CWj0xxskOn-KCN1JMpJS8d3WpT5Mdg2@mail.gmail.com","subject":"Re: Git, Mac OS X and German special characters","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2010-05-20T18:22:22Z","receivedAt":"2010-05-20T18:22:22Z","isPatch":false,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"On Thu, May 20, 2010 at 11:50 AM, Jay Soffian <jaysoffian@gmail.com> wrote:\n> To be consistent, shouldn't we have a core.HFSPlusCompat that can be\n> set on non-braindamaged filesystems to prevent filenames which would\n> alias on HFS+ from entering the repo?\n>\n>   http://developer.apple.com/mac/library/technotes/tn/tn1150.html#UnicodeSubtleties\n\nAnd here's an implementation (BSD-style license):\n\n  http://src.chromium.org/svn/trunk/src/base/file_path.cc\n\nSee GetHFSDecomposedForm and HFSFastUnicodeCompare.\n\nC++ so it'll require adapting.\n\nj.\n"}]}