{"thread":{"id":"11690","subject":"[PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","startedAt":"2008-01-21T09:12:09Z","lastAt":"2008-01-22T22:34:58Z","messageCount":22,"participants":["Mark Junker","Junio C Hamano","Johannes Schindelin","H. Peter Anvin","Linus Torvalds","Dmitry Potapov","Nicolas Pitre","Robin Rosenberg"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"66102","messageId":"fn1nl6$ek5$1@ger.gmane.org","threadId":"11690","inReplyTo":null,"subject":"[PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-21T09:12:09Z","receivedAt":"2008-01-21T09:12:09Z","isPatch":true,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8 for readdir \nand get_pathspec.\n\nI had to change get_pathspec too because otherwise git-add wouldn't work \nanymore because it uses the output of get_pathspec as strings to compare \nwith the output of readdir.\n\nI'm quite unsure because this is my first patch for the git project and \nI have several questions:\n\n1. Is FIX_UTF8_MAC the right name for this \"feature\"?\n2. Do I have to introduce a configuration option for this \"feature\"?\n\nSigned-off-by: Mark Junker <mjscod@web.de>\n---\n  Makefile          |    5 +++++\n  compat/readdir.c  |   26 ++++++++++++++++++++++++++\n  git-compat-util.h |    5 +++++\n  setup.c           |   12 ++++++++++++\n  4 files changed, 48 insertions(+), 0 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 5aac0c0..e55914e 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -417,6 +417,7 @@ ifeq ($(uname_S),Darwin)\n  \tendif\n  \tNO_STRLCPY = YesPlease\n  \tNO_MEMMEM = YesPlease\n+\tFIX_UTF8_MAC = YesPlease\n  endif\n  ifeq ($(uname_S),SunOS)\n  \tNEEDS_SOCKET = YesPlease\n@@ -616,6 +617,10 @@ ifdef NO_STRLCPY\n  \tCOMPAT_CFLAGS += -DNO_STRLCPY\n  \tCOMPAT_OBJS += compat/strlcpy.o\n  endif\n+ifdef FIX_UTF8_MAC\n+\tCOMPAT_CFLAGS += -DFIX_UTF8_MAC\n+\tCOMPAT_OBJS += compat/readdir.o\n+endif\n  ifdef NO_STRTOUMAX\n  \tCOMPAT_CFLAGS += -DNO_STRTOUMAX\n  \tCOMPAT_OBJS += compat/strtoumax.o\ndiff --git a/compat/readdir.c b/compat/readdir.c\nnew file mode 100644\nindex 0000000..045cfef\n--- /dev/null\n+++ b/compat/readdir.c\n@@ -0,0 +1,26 @@\n+#include \"../git-compat-util.h\"\n+#include \"../utf8.h\"\n+\n+#undef readdir\n+\n+static struct dirent temp;\n+\n+struct dirent *gitreaddir(DIR *dirp)\n+{\n+\tsize_t utf8_len;\n+\tchar *utf8;\n+\tstruct dirent *result;\n+\tresult = readdir(dirp);\n+\tif (result != NULL) {\n+\t\tmemcpy(&temp, result, sizeof(struct dirent));\n+\t\tutf8 = reencode_string(temp.d_name, \"UTF8\", \"UTF8-MAC\");\n+\t\tif (utf8 != NULL) {\n+\t\t\tutf8_len = strlen(utf8);\n+\t\t\ttemp.d_namlen = (u_int8_t) utf8_len;\n+\t\t\tmemcpy(temp.d_name, utf8, utf8_len + 1);\n+\t\t\tfree(utf8);\n+\t\t\tresult = &temp;\n+\t\t}\n+\t}\n+\treturn result;\n+}\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b6ef544..cd0233d 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -202,6 +202,11 @@ void *gitmemmem(const void *haystack, size_t \nhaystacklen,\n                  const void *needle, size_t needlelen);\n  #endif\n\n+#ifdef FIX_UTF8_MAC\n+#define readdir gitreaddir\n+struct dirent *gitreaddir(DIR *dirp);\n+#endif\n+\n  #ifdef __GLIBC_PREREQ\n  #if __GLIBC_PREREQ(2, 1)\n  #define HAVE_STRCHRNUL\ndiff --git a/setup.c b/setup.c\nindex adede16..4cec28b 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1,5 +1,8 @@\n  #include \"cache.h\"\n  #include \"dir.h\"\n+#ifdef FIX_UTF8_MAC\n+#include \"utf8.h\"\n+#endif\n\n  static int inside_git_dir = -1;\n  static int inside_work_tree = -1;\n@@ -131,6 +134,15 @@ const char **get_pathspec(const char *prefix, const \nchar **pathspec)\n  \tp = pathspec;\n  \tprefixlen = prefix ? strlen(prefix) : 0;\n  \tdo {\n+#ifdef FIX_UTF8_MAC\n+\t\t/* Reencode as UTF8 (composed) to have a counterpart for the\n+\t\t * readdir-replacement on MacOS X.\n+\t\t */\n+\t\tchar *utf8 = reencode_string(entry, \"UTF8\", \"UTF8-MAC\");\n+\t\tif (utf8 != NULL) {\n+\t\t\tentry = utf8;\n+\t\t}\n+#endif\n  \t\t*p = prefix_path(prefix, prefixlen, entry);\n  \t} while ((entry = *++p) != NULL);\n  \treturn (const char **) pathspec;\n-- \n1.5.4.rc3.40.gebe4\n"},{"id":"66103","messageId":"fn1pj9$kkg$1@ger.gmane.org","threadId":"11690","inReplyTo":"fn1nl6$ek5$1@ger.gmane.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-21T09:45:17Z","receivedAt":"2008-01-21T09:45:17Z","isPatch":true,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Sorry, the patch was broken (TABS were converted to spaces).\n\nSigned-off-by: Mark Junker <mjscod@web.de>\n---\n  Makefile          |    5 +++++\n  compat/readdir.c  |   26 ++++++++++++++++++++++++++\n  git-compat-util.h |    5 +++++\n  setup.c           |   12 ++++++++++++\n  4 files changed, 48 insertions(+), 0 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 5aac0c0..e55914e 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -417,6 +417,7 @@ ifeq ($(uname_S),Darwin)\n  \tendif\n  \tNO_STRLCPY = YesPlease\n  \tNO_MEMMEM = YesPlease\n+\tFIX_UTF8_MAC = YesPlease\n  endif\n  ifeq ($(uname_S),SunOS)\n  \tNEEDS_SOCKET = YesPlease\n@@ -616,6 +617,10 @@ ifdef NO_STRLCPY\n  \tCOMPAT_CFLAGS += -DNO_STRLCPY\n  \tCOMPAT_OBJS += compat/strlcpy.o\n  endif\n+ifdef FIX_UTF8_MAC\n+\tCOMPAT_CFLAGS += -DFIX_UTF8_MAC\n+\tCOMPAT_OBJS += compat/readdir.o\n+endif\n  ifdef NO_STRTOUMAX\n  \tCOMPAT_CFLAGS += -DNO_STRTOUMAX\n  \tCOMPAT_OBJS += compat/strtoumax.o\ndiff --git a/compat/readdir.c b/compat/readdir.c\nnew file mode 100644\nindex 0000000..045cfef\n--- /dev/null\n+++ b/compat/readdir.c\n@@ -0,0 +1,26 @@\n+#include \"../git-compat-util.h\"\n+#include \"../utf8.h\"\n+\n+#undef readdir\n+\n+static struct dirent temp;\n+\n+struct dirent *gitreaddir(DIR *dirp)\n+{\n+\tsize_t utf8_len;\n+\tchar *utf8;\n+\tstruct dirent *result;\n+\tresult = readdir(dirp);\n+\tif (result != NULL) {\n+\t\tmemcpy(&temp, result, sizeof(struct dirent));\n+\t\tutf8 = reencode_string(temp.d_name, \"UTF8\", \"UTF8-MAC\");\n+\t\tif (utf8 != NULL) {\n+\t\t\tutf8_len = strlen(utf8);\n+\t\t\ttemp.d_namlen = (u_int8_t) utf8_len;\n+\t\t\tmemcpy(temp.d_name, utf8, utf8_len + 1);\n+\t\t\tfree(utf8);\n+\t\t\tresult = &temp;\n+\t\t}\n+\t}\n+\treturn result;\n+}\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b6ef544..cd0233d 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -202,6 +202,11 @@ void *gitmemmem(const void *haystack, size_t \nhaystacklen,\n                  const void *needle, size_t needlelen);\n  #endif\n\n+#ifdef FIX_UTF8_MAC\n+#define readdir gitreaddir\n+struct dirent *gitreaddir(DIR *dirp);\n+#endif\n+\n  #ifdef __GLIBC_PREREQ\n  #if __GLIBC_PREREQ(2, 1)\n  #define HAVE_STRCHRNUL\ndiff --git a/setup.c b/setup.c\nindex adede16..4cec28b 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1,5 +1,8 @@\n  #include \"cache.h\"\n  #include \"dir.h\"\n+#ifdef FIX_UTF8_MAC\n+#include \"utf8.h\"\n+#endif\n\n  static int inside_git_dir = -1;\n  static int inside_work_tree = -1;\n@@ -131,6 +134,15 @@ const char **get_pathspec(const char *prefix, const \nchar **pathspec)\n  \tp = pathspec;\n  \tprefixlen = prefix ? strlen(prefix) : 0;\n  \tdo {\n+#ifdef FIX_UTF8_MAC\n+\t\t/* Reencode as UTF8 (composed) to have a counterpart for the\n+\t\t * readdir-replacement on MacOS X.\n+\t\t */\n+\t\tchar *utf8 = reencode_string(entry, \"UTF8\", \"UTF8-MAC\");\n+\t\tif (utf8 != NULL) {\n+\t\t\tentry = utf8;\n+\t\t}\n+#endif\n  \t\t*p = prefix_path(prefix, prefixlen, entry);\n  \t} while ((entry = *++p) != NULL);\n  \treturn (const char **) pathspec;\n-- \n1.5.4.rc3.40.gebe4\n"},{"id":"66105","messageId":"fn1ptk$ljj$1@ger.gmane.org","threadId":"11690","inReplyTo":"fn1pj9$kkg$1@ger.gmane.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-21T09:50:47Z","receivedAt":"2008-01-21T09:50:47Z","isPatch":true,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"ARGH! Sorry, I'll try it again later - when I fixed this stupid implicit \nconversion of TAB to SPC ...\n\nRegards,\nMark\n"},{"id":"66107","messageId":"fn1q6b$ljj$2@ger.gmane.org","threadId":"11690","inReplyTo":"fn1ptk$ljj$1@ger.gmane.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-21T09:55:27Z","receivedAt":"2008-01-21T09:55:27Z","isPatch":true,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Hi,\n\nhere's the patch again - sent as an attachment. Sorry for any inconvenience.\n\nRegards,\nMark\n\n\nSigned-off-by: Mark Junker <mjscod@web.de>\n---\n Makefile          |    5 +++++\n compat/readdir.c  |   26 ++++++++++++++++++++++++++\n git-compat-util.h |    5 +++++\n setup.c           |   12 ++++++++++++\n 4 files changed, 48 insertions(+), 0 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 5aac0c0..e55914e 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -417,6 +417,7 @@ ifeq ($(uname_S),Darwin)\n \tendif\n \tNO_STRLCPY = YesPlease\n \tNO_MEMMEM = YesPlease\n+\tFIX_UTF8_MAC = YesPlease\n endif\n ifeq ($(uname_S),SunOS)\n \tNEEDS_SOCKET = YesPlease\n@@ -616,6 +617,10 @@ ifdef NO_STRLCPY\n \tCOMPAT_CFLAGS += -DNO_STRLCPY\n \tCOMPAT_OBJS += compat/strlcpy.o\n endif\n+ifdef FIX_UTF8_MAC\n+\tCOMPAT_CFLAGS += -DFIX_UTF8_MAC\n+\tCOMPAT_OBJS += compat/readdir.o\n+endif\n ifdef NO_STRTOUMAX\n \tCOMPAT_CFLAGS += -DNO_STRTOUMAX\n \tCOMPAT_OBJS += compat/strtoumax.o\ndiff --git a/compat/readdir.c b/compat/readdir.c\nnew file mode 100644\nindex 0000000..045cfef\n--- /dev/null\n+++ b/compat/readdir.c\n@@ -0,0 +1,26 @@\n+#include \"../git-compat-util.h\"\n+#include \"../utf8.h\"\n+\n+#undef readdir\n+\n+static struct dirent temp;\n+\n+struct dirent *gitreaddir(DIR *dirp)\n+{\n+\tsize_t utf8_len;\n+\tchar *utf8;\n+\tstruct dirent *result;\n+\tresult = readdir(dirp);\n+\tif (result != NULL) {\n+\t\tmemcpy(&temp, result, sizeof(struct dirent));\n+\t\tutf8 = reencode_string(temp.d_name, \"UTF8\", \"UTF8-MAC\");\n+\t\tif (utf8 != NULL) {\n+\t\t\tutf8_len = strlen(utf8);\n+\t\t\ttemp.d_namlen = (u_int8_t) utf8_len;\n+\t\t\tmemcpy(temp.d_name, utf8, utf8_len + 1);\n+\t\t\tfree(utf8);\n+\t\t\tresult = &temp;\n+\t\t}\n+\t}\n+\treturn result;\n+}\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b6ef544..cd0233d 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -202,6 +202,11 @@ void *gitmemmem(const void *haystack, size_t haystacklen,\n                 const void *needle, size_t needlelen);\n #endif\n \n+#ifdef FIX_UTF8_MAC\n+#define readdir gitreaddir\n+struct dirent *gitreaddir(DIR *dirp);\n+#endif\n+\n #ifdef __GLIBC_PREREQ\n #if __GLIBC_PREREQ(2, 1)\n #define HAVE_STRCHRNUL\ndiff --git a/setup.c b/setup.c\nindex adede16..4cec28b 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1,5 +1,8 @@\n #include \"cache.h\"\n #include \"dir.h\"\n+#ifdef FIX_UTF8_MAC\n+#include \"utf8.h\"\n+#endif\n \n static int inside_git_dir = -1;\n static int inside_work_tree = -1;\n@@ -131,6 +134,15 @@ const char **get_pathspec(const char *prefix, const char **pathspec)\n \tp = pathspec;\n \tprefixlen = prefix ? strlen(prefix) : 0;\n \tdo {\n+#ifdef FIX_UTF8_MAC\n+\t\t/* Reencode as UTF8 (composed) to have a counterpart for the\n+\t\t * readdir-replacement on MacOS X.\n+\t\t */\n+\t\tchar *utf8 = reencode_string(entry, \"UTF8\", \"UTF8-MAC\");\n+\t\tif (utf8 != NULL) {\n+\t\t\tentry = utf8;\n+\t\t}\n+#endif\n \t\t*p = prefix_path(prefix, prefixlen, entry);\n \t} while ((entry = *++p) != NULL);\n \treturn (const char **) pathspec;\n-- \n1.5.4.rc3.40.gebe4\n\n"},{"id":"66110","messageId":"7vve5nzdqx.fsf@gitster.siamese.dyndns.org","threadId":"11690","inReplyTo":"fn1q6b$ljj$2@ger.gmane.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-21T10:15:02Z","receivedAt":"2008-01-21T10:15:02Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Mark Junker <mjscod@web.de> writes:\n\n> diff --git a/compat/readdir.c b/compat/readdir.c\n> new file mode 100644\n> index 0000000..045cfef\n> --- /dev/null\n> +++ b/compat/readdir.c\n> @@ -0,0 +1,26 @@\n> +#include \"../git-compat-util.h\"\n> +#include \"../utf8.h\"\n> +\n> +#undef readdir\n> +\n> +static struct dirent temp;\n> +\n> +struct dirent *gitreaddir(DIR *dirp)\n> +{\n> +\tsize_t utf8_len;\n> +\tchar *utf8;\n> +\tstruct dirent *result;\n> +\tresult = readdir(dirp);\n> +\tif (result != NULL) {\n> +\t\tmemcpy(&temp, result, sizeof(struct dirent));\n> +\t\tutf8 = reencode_string(temp.d_name, \"UTF8\", \"UTF8-MAC\");\n> +\t\tif (utf8 != NULL) {\n> +\t\t\tutf8_len = strlen(utf8);\n> +\t\t\ttemp.d_namlen = (u_int8_t) utf8_len;\n> +\t\t\tmemcpy(temp.d_name, utf8, utf8_len + 1);\n> +\t\t\tfree(utf8);\n\nI do not know how Macintosh libc implements \"struc dirent\", but\nthis approach does not work in general.  For example, on Linux\nboxes with glibc, \"struct dirent\" is defined like this (pardon\nthe funny indentation --- that is from the original):\n\n        struct dirent\n          {\n        #ifndef __USE_FILE_OFFSET64\n            __ino_t d_ino;\n            __off_t d_off;\n        #else\n            __ino64_t d_ino;\n            __off64_t d_off;\n        #endif\n            unsigned short int d_reclen;\n            unsigned char d_type;\n            char d_name[256];\t\t/* We must not include limits.h! */\n          };\n\nyet you can obtain a path component longer than 256 bytes.\nApparently the library allocates longer d_name[] field than what\nis shown to the user.\n"},{"id":"66114","messageId":"fn1sk4$uh4$1@ger.gmane.org","threadId":"11690","inReplyTo":"7vve5nzdqx.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-21T10:36:55Z","receivedAt":"2008-01-21T10:36:55Z","isPatch":true,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Junio C Hamano schrieb:\n\n> I do not know how Macintosh libc implements \"struc dirent\", but\n> this approach does not work in general.\n\nIMHO there is no need that this approach works in general because this \nis a fix for MacOSX systems only. I also use d_namlen which might not be \navailable on other systems. But on MacOSX this works as expected.\n\n> yet you can obtain a path component longer than 256 bytes.\n> Apparently the library allocates longer d_name[] field than what\n> is shown to the user.\n\nThis is not a problem either because on MacOSX we get decomposed UTF8 \nand we always convert to composed UTF8. This means that the string \nreturned from reencode_string will always be smaller than the original \nfilename that had to be reencoded.\n\nRegards,\nMark\n"},{"id":"66118","messageId":"7vejcbzbge.fsf@gitster.siamese.dyndns.org","threadId":"11690","inReplyTo":"fn1sk4$uh4$1@ger.gmane.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-21T11:04:33Z","receivedAt":"2008-01-21T11:04:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Mark Junker <mjscod@web.de> writes:\n\n> Junio C Hamano schrieb:\n>\n>> I do not know how Macintosh libc implements \"struc dirent\", but\n>> this approach does not work in general.\n>\n> IMHO there is no need that this approach works in general because this\n> is a fix for MacOSX systems only. I also use d_namlen which might not\n> be available on other systems. But on MacOSX this works as expected.\n>\n>> yet you can obtain a path component longer than 256 bytes.\n>> Apparently the library allocates longer d_name[] field than what\n>> is shown to the user.\n>\n> This is not a problem either because on MacOSX we get decomposed UTF8\n> and we always convert to composed UTF8. This means that the string\n> returned from reencode_string will always be smaller than the original\n> filename that had to be reencoded.\n\nIt is not quite enough that this works Ok on MacOS, if you made\nFIX_UTF8_MAC definable in the Makefile.  After all some friendly\nand helpful Linux folks might want to enable it with their build\ntrying to help debugging, right?\n\nIn the short term, as long as it safely runs without overrunning\nthe buffer on MacOS, then that is fine, even though we will need\nsome protection to prevent this code from getting compiled and\nused on Linux with glibc, which does have the issue.\n\nI was specifically talking about this \"static\" thing.\n\n+static struct dirent temp;\n\n\n+struct dirent *gitreaddir(DIR *dirp)\n+{\n+\tsize_t utf8_len;\n+\tchar *utf8;\n+\tstruct dirent *result;\n+\tresult = readdir(dirp);\n+\tif (result != NULL) {\n+\t\tmemcpy(&temp, result, sizeof(struct dirent));\n+\t\tutf8 = reencode_string(temp.d_name, \"UTF8\", \"UTF8-MAC\");\n+\t\tif (utf8 != NULL) {\n+\t\t\tutf8_len = strlen(utf8);\n+\t\t\ttemp.d_namlen = (u_int8_t) utf8_len;\n+\t\t\tmemcpy(temp.d_name, utf8, utf8_len + 1);\n+\t\t\tfree(utf8);\n+\t\t\tresult = &temp;\n+\t\t}\n+\t}\n+\treturn result;\n+}\n\nYou memcpy() what the library gave you in *result to the\nstatically allocated \"temp\".  d_name[] in \"temp\" comes from the\nstructure definition in the user visible include file, which\ncould be much shorter than what the library gave you in *result.\nThe structure definition I showed in my message you are\nresponding to illustrates the issue.  If MacOS uses a similar\ntrick to define d_name[256] and sometimes returns much longer\nname in *result, you are truncating the name by copying only the\nfirst part of the structure and first 256 bytes of d_name[]. \n\nBut you have a Mac, I don't, so as long as you have verified\nthat their header has enough room in statically allocated \"temp\"\nto store longest possible name that can be returned from\nreaddir(), the code is Ok.  I was just being cautious, as I know\nthe above code has a problem on one platform.\n"},{"id":"66122","messageId":"alpine.LSU.1.00.0801211121440.5731@racer.site","threadId":"11690","inReplyTo":"fn1nl6$ek5$1@ger.gmane.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-21T11:24:20Z","receivedAt":"2008-01-21T11:24:20Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 21 Jan 2008, Mark Junker wrote:\n\n> Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8 for readdir \n> and get_pathspec.\n> \n> I had to change get_pathspec too because otherwise git-add wouldn't work \n> anymore because it uses the output of get_pathspec as strings to compare \n> with the output of readdir.\n> \n> I'm quite unsure because this is my first patch for the git project and \n> I have several questions:\n> \n> 1. Is FIX_UTF8_MAC the right name for this \"feature\"?\n> 2. Do I have to introduce a configuration option for this \"feature\"?\n> \n> Signed-off-by: Mark Junker <mjscod@web.de>\n\nI hate three facts about this patch:\n\n- it is too specific to the MacOSX filesystem issues (and better \n  alternatives have _already_ been proposed),\n\n- it is a new feature and not a bug fix, very, _very_ late in the rc \n  cycle,\n\n- it contains questions in the commit message? WTF?  Should it not be \n  marked as PATCH/RFC, possibly without a signoff to make sure that you \n  want to discuss it first?\n\nIt's possible I am grumpy because everybody and her dog seems to work on \nher little projects, while I listen to Junio and try to work with/on \n\"master\" already since a month.\n\nCiao,\nDscho\n"},{"id":"66123","messageId":"7v3asrzab9.fsf@gitster.siamese.dyndns.org","threadId":"11690","inReplyTo":"alpine.LSU.1.00.0801211121440.5731@racer.site","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-21T11:29:14Z","receivedAt":"2008-01-21T11:29:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> - it contains questions in the commit message? WTF?  Should it not be \n>   marked as PATCH/RFC, possibly without a signoff to make sure that you \n>   want to discuss it first?\n\nSign-off is about \"this is kosher, from licensing point of view\"\nand nothing else.  Please do not suggest otherwise.\n\nI do not mind discussions during the feature freeze as long as\nit stays at \"feeler\" level.\n"},{"id":"66129","messageId":"fn20he$c4e$1@ger.gmane.org","threadId":"11690","inReplyTo":"7vejcbzbge.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-21T11:43:45Z","receivedAt":"2008-01-21T11:43:45Z","isPatch":true,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Junio C Hamano schrieb:\n\n> You memcpy() what the library gave you in *result to the\n> statically allocated \"temp\".  d_name[] in \"temp\" comes from the\n> structure definition in the user visible include file, which\n> could be much shorter than what the library gave you in *result.\n> The structure definition I showed in my message you are\n> responding to illustrates the issue.  If MacOS uses a similar\n> trick to define d_name[256] and sometimes returns much longer\n> name in *result, you are truncating the name by copying only the\n> first part of the structure and first 256 bytes of d_name[]. \n\nNow I understand what you mean. Ok, I'll try to change this and make \nthis work on other platforms too.\n\nI didn't know that the readdir function is allowed to return something \nlonger for d_name than the specified length.\n\nRegards,\nMark\n\n\n\nFrom 239de834a4d67b4176a8dbf5504fa2f335989aaa Mon Sep 17 00:00:00 2001\nFrom: Mark Junker <mjscod@web.de>\nDate: Sun, 20 Jan 2008 16:59:32 +0100\nSubject: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8\n\nSigned-off-by: Mark Junker <mjscod@web.de>\n---\n Makefile          |    5 +++++\n compat/readdir.c  |   30 ++++++++++++++++++++++++++++++\n git-compat-util.h |    5 +++++\n setup.c           |   12 ++++++++++++\n 4 files changed, 52 insertions(+), 0 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 5aac0c0..e55914e 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -417,6 +417,7 @@ ifeq ($(uname_S),Darwin)\n \tendif\n \tNO_STRLCPY = YesPlease\n \tNO_MEMMEM = YesPlease\n+\tFIX_UTF8_MAC = YesPlease\n endif\n ifeq ($(uname_S),SunOS)\n \tNEEDS_SOCKET = YesPlease\n@@ -616,6 +617,10 @@ ifdef NO_STRLCPY\n \tCOMPAT_CFLAGS += -DNO_STRLCPY\n \tCOMPAT_OBJS += compat/strlcpy.o\n endif\n+ifdef FIX_UTF8_MAC\n+\tCOMPAT_CFLAGS += -DFIX_UTF8_MAC\n+\tCOMPAT_OBJS += compat/readdir.o\n+endif\n ifdef NO_STRTOUMAX\n \tCOMPAT_CFLAGS += -DNO_STRTOUMAX\n \tCOMPAT_OBJS += compat/strtoumax.o\ndiff --git a/compat/readdir.c b/compat/readdir.c\nnew file mode 100644\nindex 0000000..96a724a\n--- /dev/null\n+++ b/compat/readdir.c\n@@ -0,0 +1,30 @@\n+#include \"../git-compat-util.h\"\n+#include \"../utf8.h\"\n+\n+#undef readdir\n+\n+static struct dirent *temp_dirent = NULL;\n+static size_t temp_dirent_length = 0;\n+\n+struct dirent *gitreaddir(DIR *dirp)\n+{\n+\tstruct dirent *result = readdir(dirp);\n+\tif (result != NULL) {\n+\t\tchar *utf8 = reencode_string(result->d_name, \"UTF8\", \"UTF8-MAC\");\n+\t\tif (utf8 != NULL) {\n+\t\t\tsize_t utf8_len = strlen(utf8);\n+\t\t\t/* Create a copy of the dirent data only if conversion is possible. */\n+\t\t\tif (result->d_reclen > temp_dirent_length) {\n+\t\t\t\t/* Ensure that the buffer is large enough and avoid\n+\t\t\t\t * too much allocations. */\n+\t\t\t\ttemp_dirent_length = result->d_reclen;\n+\t\t\t\ttemp_dirent = realloc(temp_dirent, temp_dirent_length);\n+\t\t\t}\n+\t\t\tmemcpy(temp_dirent, result, result->d_reclen);\n+\t\t\tmemcpy(temp_dirent->d_name, utf8, utf8_len + 1);\n+\t\t\tfree(utf8);\n+\t\t\tresult = temp_dirent;\n+\t\t}\n+\t}\n+\treturn result;\n+}\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex b6ef544..cd0233d 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -202,6 +202,11 @@ void *gitmemmem(const void *haystack, size_t haystacklen,\n                 const void *needle, size_t needlelen);\n #endif\n \n+#ifdef FIX_UTF8_MAC\n+#define readdir gitreaddir\n+struct dirent *gitreaddir(DIR *dirp);\n+#endif\n+\n #ifdef __GLIBC_PREREQ\n #if __GLIBC_PREREQ(2, 1)\n #define HAVE_STRCHRNUL\ndiff --git a/setup.c b/setup.c\nindex adede16..4cec28b 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1,5 +1,8 @@\n #include \"cache.h\"\n #include \"dir.h\"\n+#ifdef FIX_UTF8_MAC\n+#include \"utf8.h\"\n+#endif\n \n static int inside_git_dir = -1;\n static int inside_work_tree = -1;\n@@ -131,6 +134,15 @@ const char **get_pathspec(const char *prefix, const char **pathspec)\n \tp = pathspec;\n \tprefixlen = prefix ? strlen(prefix) : 0;\n \tdo {\n+#ifdef FIX_UTF8_MAC\n+\t\t/* Reencode as UTF8 (composed) to have a counterpart for the\n+\t\t * readdir-replacement on MacOS X.\n+\t\t */\n+\t\tchar *utf8 = reencode_string(entry, \"UTF8\", \"UTF8-MAC\");\n+\t\tif (utf8 != NULL) {\n+\t\t\tentry = utf8;\n+\t\t}\n+#endif\n \t\t*p = prefix_path(prefix, prefixlen, entry);\n \t} while ((entry = *++p) != NULL);\n \treturn (const char **) pathspec;\n-- \n1.5.4.rc3.38.g8166\n\n"},{"id":"66131","messageId":"fn20ra$c4e$2@ger.gmane.org","threadId":"11690","inReplyTo":"alpine.LSU.1.00.0801211121440.5731@racer.site","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-21T11:49:02Z","receivedAt":"2008-01-21T11:49:02Z","isPatch":true,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Johannes Schindelin schrieb:\n\n> - it is too specific to the MacOSX filesystem issues (and better \n>   alternatives have _already_ been proposed),\n\nI know that there were proposed alternatives but I like to use git on \nMacOSX now and not in XY months.\n\n> - it is a new feature and not a bug fix, very, _very_ late in the rc \n>   cycle,\n\nIt was never meant for inclusion now. I know that this is post-1.5.4 stuff.\n\nRegards,\nMark\n"},{"id":"66132","messageId":"alpine.LSU.1.00.0801211209060.5731@racer.site","threadId":"11690","inReplyTo":"fn20ra$c4e$2@ger.gmane.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-21T12:09:31Z","receivedAt":"2008-01-21T12:09:31Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 21 Jan 2008, Mark Junker wrote:\n\n> Johannes Schindelin schrieb:\n> \n> > - it is too specific to the MacOSX filesystem issues (and better \n> > alternatives have _already_ been proposed),\n> \n> I know that there were proposed alternatives but I like to use git on \n> MacOSX now and not in XY months.\n> \n> > - it is a new feature and not a bug fix, very, _very_ late in the rc \n> > cycle,\n> \n> It was never meant for inclusion now. I know that this is post-1.5.4 \n> stuff.\n\nIn this case, I am going to work on my suggestion myself.\n\nNow.\n\nCiao,\nDscho\n"},{"id":"66171","messageId":"alpine.LSU.1.00.0801211913030.5731@racer.site","threadId":"11690","inReplyTo":"alpine.LSU.1.00.0801211209060.5731@racer.site","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-21T19:14:11Z","receivedAt":"2008-01-21T19:14:11Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 21 Jan 2008, Johannes Schindelin wrote:\n\n> On Mon, 21 Jan 2008, Mark Junker wrote:\n> \n> > Johannes Schindelin schrieb:\n> > \n> > > - it is too specific to the MacOSX filesystem issues (and better \n> > > alternatives have _already_ been proposed),\n> > \n> > I know that there were proposed alternatives but I like to use git on \n> > MacOSX now and not in XY months.\n> > \n> > > - it is a new feature and not a bug fix, very, _very_ late in the rc \n> > > cycle,\n> > \n> > It was never meant for inclusion now. I know that this is post-1.5.4 \n> > stuff.\n> \n> In this case, I am going to work on my suggestion myself.\n> \n> Now.\n\nI retract that.  Although I put some work into it, I agree that it is post \n1.5.4.\n\nPlus, I do not want anybody to think that shouting and being a PITA buys \nhim anything.\n\nCiao,\nDscho\n"},{"id":"66271","messageId":"47956C4E.1080903@zytor.com","threadId":"11690","inReplyTo":"fn1sk4$uh4$1@ger.gmane.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-01-22T04:08:46Z","receivedAt":"2008-01-22T04:08:46Z","isPatch":true,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Mark Junker wrote:\n> Junio C Hamano schrieb:\n> \n>> I do not know how Macintosh libc implements \"struc dirent\", but\n>> this approach does not work in general.\n> \n> IMHO there is no need that this approach works in general because this \n> is a fix for MacOSX systems only. I also use d_namlen which might not be \n> available on other systems. But on MacOSX this works as expected.\n> \n>> yet you can obtain a path component longer than 256 bytes.\n>> Apparently the library allocates longer d_name[] field than what\n>> is shown to the user.\n> \n> This is not a problem either because on MacOSX we get decomposed UTF8 \n> and we always convert to composed UTF8. This means that the string \n> returned from reencode_string will always be smaller than the original \n> filename that had to be reencoded.\n> \n\nThat's not true!  There are strings which gets longer when a composing \nnormalization is applied.  Please see section 3.3 of Unicode Techical \nReport 36:\n\n\thttp://www.unicode.org/reports/tr36/\n\n > People assume that NFC always composes, and thus is the same or\n > shorter length than the original source. However, some characters\n > decompose in NFC.\n\n(NFC = Normalization Form Composing.)\n\nU+1D160 MUSICAL SYMBOL EIGHT NOTE is given as an example with a 3x \nexpansion factor when encoded in UTF-8 (I don't know what it expands to; \nseems odd to me.)\n\n\t-hpa\n"},{"id":"66275","messageId":"alpine.LFD.1.00.0801212025050.2957@woody.linux-foundation.org","threadId":"11690","inReplyTo":"7vve5nzdqx.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-22T04:59:56Z","receivedAt":"2008-01-22T04:59:56Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Junio C Hamano wrote:\n> \n> yet you can obtain a path component longer than 256 bytes.\n\nIndividual components are limited to 255 bytes by most filesystems \n(PATH_MAX is the whole path, not any individual component).\n\nThat said, you're right. It's not really a design requirement, and since \nyou never get an array of \"struct dirent\", just a pointer to a single one, \nit would be perfectly normal and natural for \"struct dirent\" to be \ndeclared with a unsized d_name[].\n\nIt's also quite possible that some implementations might even have \nd_name[] not as an array, but as a pointer to somewhere else (POSIX may \nrequire it to be an array, I didn't check).\n\nThat said, I bet that Mark isn't the only one to have written code like \nthat, so I suspect Mark's code probably works in practice pretty much \neverywhere, even if I don't think it's necessarily _required_ to work \ncorrectly.\n\nI do suspect that if you really want to make this portable, and able to \nhandle an expanding d_name[] too, I think you need to make sure you \nallocate a big-enough one. And if you worry about d_name perhaps being a \npointer, that really does mean that you'd need to convert the \nsystem-supplied \"struct dirent\" into a \"git_dirent_t\" that you can \ncontrol.\n\nThat said, I think this patch has a bigger problem, namely just \nfundamentally that\n\n\tchar *utf8 = reencode_string(entry, \"UTF8\", \"UTF8-MAC\");\n\nis just unbelievably slow. That's just not how it should be done.\n\nFirst off, the common case is that the filename likely has everything in \nplain 7-bit ascii. So rather than re-encoding by default, the first thing \nto do is to just see if it even needs re-encoding. Even if it's as simple \nas saying \"does it have any high bits at all\", that's going to be a *huge* \nperformance win.\n\nSo start off with something like\n\n\tint is_usascii(const char *p)\n\t{\n\t\tchar c;\n\n\t\tdo {\n\t\t\tc = *p++;\n\t\t} while (c > 0);\n\t\treturn !c;\n\t}\n\nand now you can do\n\n\tif (is_usascii(entry->d_name))\n\t\treturn entry;\n\nbefore you even *look* at re-encoding it (and this basically works for all \ncases - we really don't care about EBCDIC, do we? So even if this routine \nwas meant to do Latin1<->utf8, the above \"is_usascii()\" test is always the \nright thing to do).\n\nAnyway, even if you do that, our \"reencode_string()\" is really *so* \nexpensive that you really don't want to do it on a filename by filename \nbasis. It literally does a malloc() for each allocation. It might well be \nworth it to find something that is more utf-8-specific (and I could well \nimagine that Mac OS X comes with some UTF libraries, if only because we \ncannot possibly be the only people with this issue).\n\n(Same goes for Latin1<->UTF conversion, for that matter. If somebody wants \nto add that, I suspect it's best done by hand, not using iconv and our \nrather expensive layer around it. That said, latin1->utf8 is actually \nmuch *easier* than utf-8 NFD->NFC).\n\n\t\t\tLinus\n"},{"id":"66281","messageId":"alpine.LFD.1.00.0801212304460.2957@woody.linux-foundation.org","threadId":"11690","inReplyTo":"alpine.LFD.1.00.0801212025050.2957@woody.linux-foundation.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-22T07:16:54Z","receivedAt":"2008-01-22T07:16:54Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Linus Torvalds wrote:\n> \n> I do suspect that if you really want to make this portable, and able to \n> handle an expanding d_name[] too, I think you need to make sure you \n> allocate a big-enough one. And if you worry about d_name perhaps being a \n> pointer, that really does mean that you'd need to convert the \n> system-supplied \"struct dirent\" into a \"git_dirent_t\" that you can \n> control.\n> \n> That said, I think this patch has a bigger problem, namely just \n> fundamentally that\n> \n> \tchar *utf8 = reencode_string(entry, \"UTF8\", \"UTF8-MAC\");\n> \n> is just unbelievably slow. That's just not how it should be done.\n\nHaving thought about this some more, I'm starting to suspect that the \n\"readdir()\" wrapper thing won't work very well.\n\nYes, it will work on OS X, but for all the wrong reasons. It works there \njust because of the stupid normalization that OS X does both on filename \ninput and output, so if we hook into readdir() and munge the name there, \nwe'll still be able to use the munged name for lstat() and open().\n\nHowever, we'll never be able to test it on a sane Unix system, and it \nwon't ever be able to handle the case of a filesystem actually being \nLatin1 but git being asked to try to transparently convert it to utf-8 in \norder to work with others.\n\nBecause most of those readdir() calls will just be fed back not just to \nthe filesystem as lstat() calls later, but also to the recursive directory \ntraversal itself, so if we munge the name, we're also going to screw name \nlookup.\n\nAgain, as an OSX-only workaround it's probably acceptable, and perhaps \nthat's the only thing to look at right now. But it does strike me as a \ndesign mistake to do it at that level.\n\nIt would be conceptually nicer to do it in \"add_file_to_index()\" instead. \nIe anything that creates a \"struct cache_entry\" would do the \nconversion. \n\nSo it would be good if somebody looked at what happens if you do the OSX \nhack in add_file_to_index() instead, and see if it works there..\n\n\t\tLinus\n"},{"id":"66285","messageId":"7vmyqythwa.fsf@gitster.siamese.dyndns.org","threadId":"11690","inReplyTo":"alpine.LFD.1.00.0801212304460.2957@woody.linux-foundation.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-22T07:54:13Z","receivedAt":"2008-01-22T07:54:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Again, as an OSX-only workaround it's probably acceptable, and perhaps \n> that's the only thing to look at right now. But it does strike me as a \n> design mistake to do it at that level.\n\nYes, we would need a reverse conversion when going from index to\nwork tree, including entry.c, in order to be able to emulate\nthis on filesystems that do not take \"equivalent\" but different\nnames on open(), creat() and lstat().\n"},{"id":"66303","messageId":"20080122115741.GI14871@dpotapov.dyndns.org","threadId":"11690","inReplyTo":"alpine.LFD.1.00.0801212025050.2957@woody.linux-foundation.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-22T11:57:41Z","receivedAt":"2008-01-22T11:57:41Z","isPatch":true,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Jan 21, 2008 at 08:59:56PM -0800, Linus Torvalds wrote:\n> \n> Anyway, even if you do that, our \"reencode_string()\" is really *so* \n> expensive that you really don't want to do it on a filename by filename \n> basis. It literally does a malloc() for each allocation. It might well be \n> worth it to find something that is more utf-8-specific (and I could well \n> imagine that Mac OS X comes with some UTF libraries, if only because we \n> cannot possibly be the only people with this issue).\n\nYes, starting with Mac OS X 10.2 there are functions for that.\nhttp://developer.apple.com/qa/qa2001/qa1235.html\n\nAnyway, even if iconv is to be used, I believe it should be possible to\navoid malloc here (I usually allocate 256 on stack and use malloc()/free()\nonly when I need more than that which in practice never happens!). It is\nalso avoidable to call iconv_open/iconv_close for each name by putting the\nallocated descriptor for character set conversion into a static variable.\nThus leaving iconv() alone, which should not be big overhead provided that\nit is done only for non-ASCII names.\n\nDmitry\n"},{"id":"66304","messageId":"20080122122019.GJ14871@dpotapov.dyndns.org","threadId":"11690","inReplyTo":"alpine.LFD.1.00.0801212304460.2957@woody.linux-foundation.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-22T12:20:19Z","receivedAt":"2008-01-22T12:20:19Z","isPatch":true,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Jan 21, 2008 at 11:16:54PM -0800, Linus Torvalds wrote:\n> \n> Yes, it will work on OS X, but for all the wrong reasons. It works there \n> just because of the stupid normalization that OS X does both on filename \n> input and output, so if we hook into readdir() and munge the name there, \n> we'll still be able to use the munged name for lstat() and open().\n\nYes, when I proposed the readdir() wrapper, I meant it to be as OS X\nspecific hack. Just because HFS+ munges names and does that by converting\nthem in the form that is HFS+ specific, we can safely convert then into\nNFC, as we do not lose more information than it is lost already, and\nmore importantly, AFAIK, everything that a user types on Mac is in NFC,\nwhether they are names in the command line or names in .gitatributes.\n\n> However, we'll never be able to test it on a sane Unix system, and it \n> won't ever be able to handle the case of a filesystem actually being \n> Latin1 but git being asked to try to transparently convert it to utf-8 in \n> order to work with others.\n\nYes, but that is a separate issue, which unfortunately is much more\ndifficult to deal with. Basically, there are two approaches -- either\nto wrap all input/output functions, or to find another point where it\nis possible to convert names without re-writing too much code in Git.\nIt seems to me that the first approach may requires wrapping too much\nfunctions, but looking at the code I am not sure that the second will\nbe much easier. There are many places where a filename in the local\nencoding will interact with Git internal encoding used by repo.\n\nIf we spoke about Windows only, I would say that the first approach makes\nmuch more sense, because all i/o functions used on Windows are already\nwrappers over Unicode functions. So, converting UTF-8 <-> UTF-16 makes\nmuch more sense than UTF-8 <-> some-local-encoding(*) <-> UTF-16.\n\n(*) In fact, two different encodings for the same locale setting -- \none for console and the other for non-console programs!\n\n> It would be conceptually nicer to do it in \"add_file_to_index()\" instead. \n> Ie anything that creates a \"struct cache_entry\" would do the \n> conversion. \n\nI don't think it is going to work, without changing a lot of code,\nbecause filenames entered by user and those that are returned by\nreaddir() are different. Also, .gitignore or .gitattributes files will\nhave filenames in the form that differs from returned by readdir().\n\n\nDmitry\n"},{"id":"66310","messageId":"alpine.LFD.1.00.0801220917310.20753@xanadu.home","threadId":"11690","inReplyTo":"alpine.LFD.1.00.0801212025050.2957@woody.linux-foundation.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-01-22T14:21:56Z","receivedAt":"2008-01-22T14:21:56Z","isPatch":true,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 21 Jan 2008, Linus Torvalds wrote:\n\n> First off, the common case is that the filename likely has everything in \n> plain 7-bit ascii. So rather than re-encoding by default, the first thing \n> to do is to just see if it even needs re-encoding. Even if it's as simple \n> as saying \"does it have any high bits at all\", that's going to be a *huge* \n> performance win.\n> \n> So start off with something like\n> \n> \tint is_usascii(const char *p)\n> \t{\n> \t\tchar c;\n> \n> \t\tdo {\n> \t\t\tc = *p++;\n> \t\t} while (c > 0);\n> \t\treturn !c;\n> \t}\n\nYou need to use \"signed char\" here.  On ARM a char is unsigned by \ndefault.  That's the case on some other systems too.\n\n\nNicolas\n"},{"id":"66315","messageId":"alpine.LFD.1.00.0801220757080.2957@woody.linux-foundation.org","threadId":"11690","inReplyTo":"alpine.LFD.1.00.0801220917310.20753@xanadu.home","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-22T15:58:47Z","receivedAt":"2008-01-22T15:58:47Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 22 Jan 2008, Nicolas Pitre wrote:\n> \n> You need to use \"signed char\" here.  On ARM a char is unsigned by \n> default.  That's the case on some other systems too.\n\nCorrect you are. Me bad.\n\n\t\tLinus\n"},{"id":"66333","messageId":"200801222334.58874.robin.rosenberg.lists@dewire.com","threadId":"11690","inReplyTo":"7vmyqythwa.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] Use FIX_UTF8_MAC to enable conversion from UTF8-MAC to UTF8","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2008-01-22T22:34:58Z","receivedAt":"2008-01-22T22:34:58Z","isPatch":true,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"tisdagen den 22 januari 2008 skrev Junio C Hamano:\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > Again, as an OSX-only workaround it's probably acceptable, and perhaps \n> > that's the only thing to look at right now. But it does strike me as a \n> > design mistake to do it at that level.\n> \n> Yes, we would need a reverse conversion when going from index to\n> work tree, including entry.c, in order to be able to emulate\n> this on filesystems that do not take \"equivalent\" but different\n> names on open(), creat() and lstat().\n\nAbout this size:\n\nhttp://rosenberg.homelinux.net/cgi-bin/gitweb/gitweb.cgi?p=GIT.git;a=commitdiff;h=766d84eff841172c3754f67c66363a1d60038de5\n\nAnd messed up expanded tests:\n\nhttp://rosenberg.homelinux.net/cgi-bin/gitweb/gitweb.cgi?p=GIT.git;a=commitdiff;h=5d73e28397f7ec0f85fcb8e31e91326afbcfea19\n\nJunio: @Maybe in five years@, you said.. Four more to go.\n\n-- robin\n"}]}