{"thread":{"id":"23837","subject":"[PATCH] t9129: fix UTF-8 locale detection","startedAt":"2010-05-18T14:41:25Z","lastAt":"2011-01-07T18:49:05Z","messageCount":23,"participants":["Yann Droneaud","Michael J Gruber","Linus Torvalds","Andreas Schwab","Miles Bader","Junio C Hamano"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"141863","messageId":"1274193685-5468-1-git-send-email-yann@droneaud.fr","threadId":"23837","inReplyTo":null,"subject":"[PATCH] t9129: fix UTF-8 locale detection","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2010-05-18T14:41:25Z","receivedAt":"2010-05-18T14:41:25Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Since I don't have en_US.utf8, some tests failed:\n\n  * UTF-8 locale not available, test skipped\n  * skip 10: ISO-8859-1 should match UTF-8 in svn\n  * skip 11: eucJP should match UTF-8 in svn\n  * skip 12: ISO-2022-JP should match UTF-8 in svn\n\nOn my system locale -a reports:\n\n   en_US\n   en_US.ISO-8859-1\n   en_US.UTF-8\n\nAccording to Wikipedia utf8 is not a correct name\nfor the UTF-8 encoding:\nhttp://en.wikipedia.org/wiki/UTF-8#Official_name_and_incorrect_variants\n\nAnd compare_svn_head_with() is explicitly using en_US.UTF-8\nlocale.\n\nSigned-off-by: Yann Droneaud <yann@droneaud.fr>\n---\n t/t9129-git-svn-i18n-commitencoding.sh |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/t/t9129-git-svn-i18n-commitencoding.sh b/t/t9129-git-svn-i18n-commitencoding.sh\nindex b9224bd..ec6ed4f 100755\n--- a/t/t9129-git-svn-i18n-commitencoding.sh\n+++ b/t/t9129-git-svn-i18n-commitencoding.sh\n@@ -69,7 +69,7 @@ do\n \t'\n done\n \n-if locale -a |grep -q en_US.utf8; then\n+if locale -a |grep -q en_US.UTF-8; then\n \ttest_set_prereq UTF8\n else\n \tsay \"UTF-8 locale not available, test skipped\"\n-- \n1.6.4.4\n"},{"id":"141865","messageId":"4BF2BABC.2010405@drmicha.warpmail.net","threadId":"23837","inReplyTo":"1274193685-5468-1-git-send-email-yann@droneaud.fr","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2010-05-18T16:05:16Z","receivedAt":"2010-05-18T16:05:16Z","isPatch":true,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Yann Droneaud venit, vidit, dixit 18.05.2010 16:41:\n> Since I don't have en_US.utf8, some tests failed:\n> \n>   * UTF-8 locale not available, test skipped\n>   * skip 10: ISO-8859-1 should match UTF-8 in svn\n>   * skip 11: eucJP should match UTF-8 in svn\n>   * skip 12: ISO-2022-JP should match UTF-8 in svn\n> \n> On my system locale -a reports:\n> \n>    en_US\n>    en_US.ISO-8859-1\n>    en_US.UTF-8\n> \n\nlocale -a|grep en_US\nen_US\nen_US.iso88591\nen_US.iso885915\nen_US.utf8\n\nThis is on Fedora 13, which is not exactly exotic. What is your system?\n\n> According to Wikipedia utf8 is not a correct name\n> for the UTF-8 encoding:\n> http://en.wikipedia.org/wiki/UTF-8#Official_name_and_incorrect_variants\n> \n> And compare_svn_head_with() is explicitly using en_US.UTF-8\n> locale.\n> \n> Signed-off-by: Yann Droneaud <yann@droneaud.fr>\n> ---\n>  t/t9129-git-svn-i18n-commitencoding.sh |    2 +-\n>  1 files changed, 1 insertions(+), 1 deletions(-)\n> \n> diff --git a/t/t9129-git-svn-i18n-commitencoding.sh b/t/t9129-git-svn-i18n-commitencoding.sh\n> index b9224bd..ec6ed4f 100755\n> --- a/t/t9129-git-svn-i18n-commitencoding.sh\n> +++ b/t/t9129-git-svn-i18n-commitencoding.sh\n> @@ -69,7 +69,7 @@ do\n>  \t'\n>  done\n>  \n> -if locale -a |grep -q en_US.utf8; then\n> +if locale -a |grep -q en_US.UTF-8; then\n>  \ttest_set_prereq UTF8\n>  else\n>  \tsay \"UTF-8 locale not available, test skipped\"\n\nFunny thing is the test succeeds for me, even when run within\nLANG=en_US.iso88591.\nSo I'd suggest to use\n\n-if locale -a |grep -q en_US.utf8; then\n+if locale -a |egrep -q 'en_US.utf8|en_US.UTF-8'; then\n\nand embrace for more variants to appear down the road...\n\nMichael\n"},{"id":"141871","messageId":"1274202486.4228.22.camel@localhost","threadId":"23837","inReplyTo":"4BF2BABC.2010405@drmicha.warpmail.net","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2010-05-18T17:08:06Z","receivedAt":"2010-05-18T17:08:06Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Le mardi 18 mai 2010 à 18:05 +0200, Michael J Gruber a écrit :\n> Yann Droneaud venit, vidit, dixit 18.05.2010 16:41:\n> > Since I don't have en_US.utf8, some tests failed:\n\n> > \n> > On my system locale -a reports:\n> > \n> >    en_US\n> >    en_US.ISO-8859-1\n> >    en_US.UTF-8\n> > \n> \n> locale -a|grep en_US\n> en_US\n> en_US.iso88591\n> en_US.iso885915\n> en_US.utf8\n> \n> This is on Fedora 13, which is not exactly exotic. What is your system?\n> \n\nMandriva Linux 2009.1 and 2010.0, see results of locale -a :\n\nhttp://pastebin.mandriva.com/18557\nhttp://pastebin.mandriva.com/18555\n\nI've double check with Mandriva's developers who have\n\n  en_US\n  en_US.iso88591\n  en_US.utf8\n  en_US.UTF-8\n\n> > According to Wikipedia utf8 is not a correct name\n> > for the UTF-8 encoding:\n> > http://en.wikipedia.org/wiki/UTF-8#Official_name_and_incorrect_variants\n> > \n\nUTF-8 seems to be the correct name.\n\n> >  \n> > -if locale -a |grep -q en_US.utf8; then\n> > +if locale -a |grep -q en_US.UTF-8; then\n> >  \ttest_set_prereq UTF8\n> >  else\n> >  \tsay \"UTF-8 locale not available, test skipped\"\n> \n> Funny thing is the test succeeds for me, even when run within\n> LANG=en_US.iso88591.\n\n> So I'd suggest to use\n> \n> -if locale -a |grep -q en_US.utf8; then\n> +if locale -a |egrep -q 'en_US.utf8|en_US.UTF-8'; then\n> \n> and embrace for more variants to appear down the road...\n> \n\nUsing en_US.UTF-8 seems more accurate when I wrote the patch since, as I\nwrote before, compare_svn_head_with() is using LC_ALL=en_US.UTF-8.\nSo en_US.UTF-8 is an alias for en_US.utf8, whatever the canonical\nversion is.\n\nSo let's go for another version.\n\n-- \nYann Droneaud\n"},{"id":"141872","messageId":"1274203013-1349-1-git-send-email-yann@droneaud.fr","threadId":"23837","inReplyTo":"1274202486.4228.22.camel@localhost","subject":"[PATCH] t9129: fix UTF-8 locale detection","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2010-05-18T17:16:53Z","receivedAt":"2010-05-18T17:16:53Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Since I don't have en_US.utf8, some tests failed:\n\n  * UTF-8 locale not available, test skipped\n  * skip 10: ISO-8859-1 should match UTF-8 in svn\n  * skip 11: eucJP should match UTF-8 in svn\n  * skip 12: ISO-2022-JP should match UTF-8 in svn\n\nOn my system locale -a reports:\n\n   en_US\n   en_US.ISO-8859-1\n   en_US.UTF-8\n\nTests available locales against en_US\\.(utf|UTF)-?8 regexp.\n\nSigned-off-by: Yann Droneaud <yann@droneaud.fr>\n---\n t/t9129-git-svn-i18n-commitencoding.sh |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/t/t9129-git-svn-i18n-commitencoding.sh b/t/t9129-git-svn-i18n-commitencoding.sh\nindex b9224bd..00a273b 100755\n--- a/t/t9129-git-svn-i18n-commitencoding.sh\n+++ b/t/t9129-git-svn-i18n-commitencoding.sh\n@@ -69,7 +69,7 @@ do\n \t'\n done\n \n-if locale -a |grep -q en_US.utf8; then\n+if locale -a |grep -qE '^en_US\\.(utf|UTF)-?8$'; then\n \ttest_set_prereq UTF8\n else\n \tsay \"UTF-8 locale not available, test skipped\"\n-- \n1.6.4.4\n"},{"id":"141873","messageId":"alpine.LFD.2.00.1005181037250.4195@i5.linux-foundation.org","threadId":"23837","inReplyTo":"1274203013-1349-1-git-send-email-yann@droneaud.fr","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2010-05-18T17:45:26Z","receivedAt":"2010-05-18T17:45:26Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 18 May 2010, Yann Droneaud wrote:\n>  \n> -if locale -a |grep -q en_US.utf8; then\n> +if locale -a |grep -qE '^en_US\\.(utf|UTF)-?8$'; then\n\nWhile -E is POSIX, I suspect that it's not universal. iirc, you still have \nsome really crap fileutils tools coming with Solaris, for example. \n\nWouldn't it be easier to just make it ignore case, and do\n\n\tgrep -qi '^en_US\\.utf-?8$'\n\ninstead?\n\nI'm also not entirely sure you want to make that pattern stricter - the \nwhole problem with the old pattern was that it was too exact, so why add \nthe beginning/end requirement?\n\n\t\tLinus\n"},{"id":"141884","messageId":"m24oi5j81q.fsf@igel.home","threadId":"23837","inReplyTo":"alpine.LFD.2.00.1005181037250.4195@i5.linux-foundation.org","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Andreas Schwab","fromEmail":"schwab@linux-m68k.org","sentAt":"2010-05-18T19:58:57Z","receivedAt":"2010-05-18T19:58:57Z","isPatch":true,"sender":{"key":"schwab@linux-m68k.org","avatar":"https://avatars.githubusercontent.com/u/2175493?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Wouldn't it be easier to just make it ignore case, and do\n>\n> \tgrep -qi '^en_US\\.utf-?8$'\n\nYou'll need ERE's for the ? operator.\n\nAndreas.\n\n-- \nAndreas Schwab, schwab@linux-m68k.org\nGPG Key fingerprint = 58CA 54C7 6D53 942B 1756  01D3 44D5 214B 8276 4ED5\n\"And now for something completely different.\"\n"},{"id":"141887","messageId":"alpine.LFD.2.00.1005181300130.7559@i5.linux-foundation.org","threadId":"23837","inReplyTo":"m24oi5j81q.fsf@igel.home","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2010-05-18T20:00:32Z","receivedAt":"2010-05-18T20:00:32Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 18 May 2010, Andreas Schwab wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > Wouldn't it be easier to just make it ignore case, and do\n> >\n> > \tgrep -qi '^en_US\\.utf-?8$'\n> \n> You'll need ERE's for the ? operator.\n\nOh, just replace it with '*' then. It's not like anybody cares.\n\n\t\tLinus\n"},{"id":"141888","messageId":"1274215074.16337.4.camel@localhost","threadId":"23837","inReplyTo":"4BF2BABC.2010405@drmicha.warpmail.net","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2010-05-18T20:37:54Z","receivedAt":"2010-05-18T20:37:54Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Le mardi 18 mai 2010 à 18:05 +0200, Michael J Gruber a écrit :\n> Yann Droneaud venit, vidit, dixit 18.05.2010 16:41:\n\n> This is on Fedora 13, which is not exactly exotic. What is your system?\n\nI've tested on:\n - FreeBSD 8.0\n - NetBSD 5.0.1\n - OpenSolaris 2009.06\n\nand none of these systems report en_US.utf8 in locale -a, they're all\nreporting en_US.UTF-8\n\nRegards\n\n-- \nYann\n"},{"id":"141889","messageId":"1274215797.16337.16.camel@localhost","threadId":"23837","inReplyTo":"alpine.LFD.2.00.1005181037250.4195@i5.linux-foundation.org","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2010-05-18T20:49:57Z","receivedAt":"2010-05-18T20:49:57Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Le mardi 18 mai 2010 à 10:45 -0700, Linus Torvalds a écrit :\n> \n> On Tue, 18 May 2010, Yann Droneaud wrote:\n> >  \n> > -if locale -a |grep -q en_US.utf8; then\n> > +if locale -a |grep -qE '^en_US\\.(utf|UTF)-?8$'; then\n> \n> While -E is POSIX, I suspect that it's not universal. iirc, you still have \n> some really crap fileutils tools coming with Solaris, for example. \n> \n\nYou're right, Solaris's own grep doesn't known about -E nor -e.\n\nAnd it even doesn't know about -q : \n\n $ grep -q                                                     \n grep: illegal option -- q\n Usage: grep -hblcnsviw pattern file . . .\n\nSo the whole test won't work for older Solaris.\n\nSolaris don't support grep -e, but has the non POSIX egrep instead\n(which doesn't support -q too).\n\n[...]\n\n> I'm also not entirely sure you want to make that pattern stricter - the \n> whole problem with the old pattern was that it was too exact, so why add \n> the beginning/end requirement?\n> \n\nJust to be sure it doesn't match \"garbage\". \nInitial regexp was using a straight '.' operator, so while fixing it, I\nthought it would be better to achieve \"perfect match\".\n\nRegards.\n\n-- \nYann Droneaud\n"},{"id":"141897","messageId":"87r5l8mwzl.fsf@catnip.gol.com","threadId":"23837","inReplyTo":"1274215074.16337.4.camel@localhost","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Miles Bader","fromEmail":"miles@gnu.org","sentAt":"2010-05-19T02:44:14Z","receivedAt":"2010-05-19T02:44:14Z","isPatch":true,"sender":{"key":"miles@gnu.org","avatar":"https://gravatar.com/avatar/01069b69593af7bff28e2f97afeb3644ae6fe2f5f56cb3a8cf34c5fb8c36efe5?d=mp&s=160"},"body":"Yann Droneaud <yann@droneaud.fr> writes:\n> I've tested on:\n>  - FreeBSD 8.0\n>  - NetBSD 5.0.1\n>  - OpenSolaris 2009.06\n>\n> and none of these systems report en_US.utf8 in locale -a, they're all\n> reporting en_US.UTF-8\n\nStill, \"utf8\" is common enough, so clearly it should be supported (along\nwith \"utf-8\" and \"UTF-8\" etc).\n\n[On my system::\n\n   $ locale -a\n   C\n   POSIX\n   en_US.utf8\n   ja_JP.utf8\n   ko_KR\n   ko_KR.euckr\n   ko_KR.utf8\n   korean\n   korean.euc\n]\n\n-miles\n\n-- \nIdiot, n. A member of a large and powerful tribe whose influence in human\naffairs has always been dominant and controlling.\n"},{"id":"141909","messageId":"1274282202.4275.68.camel@localhost","threadId":"23837","inReplyTo":"1274215797.16337.16.camel@localhost","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2010-05-19T15:16:42Z","receivedAt":"2010-05-19T15:16:42Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Le mardi 18 mai 2010 à 22:49 +0200, Yann Droneaud a écrit :\n> Le mardi 18 mai 2010 à 10:45 -0700, Linus Torvalds a écrit :\n> > \n> > On Tue, 18 May 2010, Yann Droneaud wrote:\n> > >  \n> > > -if locale -a |grep -q en_US.utf8; then\n> > > +if locale -a |grep -qE '^en_US\\.(utf|UTF)-?8$'; then\n> > \n> > While -E is POSIX, I suspect that it's not universal. iirc, you still have \n> > some really crap fileutils tools coming with Solaris, for example. \n> > \n> \n> You're right, Solaris's own grep doesn't known about -E nor -e.\n> \n> And it even doesn't know about -q : \n> \n>  $ grep -q                                                     \n>  grep: illegal option -- q\n>  Usage: grep -hblcnsviw pattern file . . .\n> \n\nIntegrating Linus's remarks and some autoconf[1][2] hints, here is a\nproposal for a portable test:\n\n   if locale -a |grep -i 'en_US\\.utf-*8' > /dev/null ; then\n\nIt should be as portable as possible without too much work.\n\n\n[1] egrep\n<http://www.gnu.org/software/autoconf/manual/html_node/Limitations-of-Usual-Tools.html#index-g_t_0040command_007begrep_007d-1706>\n\n[2] grep\n<http://www.gnu.org/software/autoconf/manual/html_node/Limitations-of-Usual-Tools.html#index-g_t_0040command_007bgrep_007d-1712>\n\nRegards.\n\n-- \nYann Droneaud\n"},{"id":"142208","messageId":"1274720888.4838.13.camel@localhost","threadId":"23837","inReplyTo":"1274202486.4228.22.camel@localhost","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2010-05-24T17:08:08Z","receivedAt":"2010-05-24T17:08:08Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Le mardi 18 mai 2010 à 19:08 +0200, Yann Droneaud a écrit :\n> Le mardi 18 mai 2010 à 18:05 +0200, Michael J Gruber a écrit :\n> > Yann Droneaud venit, vidit, dixit 18.05.2010 16:41:\n> > > Since I don't have en_US.utf8, some tests failed:\n> \n> > > \n> > > On my system locale -a reports:\n> > > \n> > >    en_US\n> > >    en_US.ISO-8859-1\n> > >    en_US.UTF-8\n> > > \n> > \n> > locale -a|grep en_US\n> > en_US\n> > en_US.iso88591\n> > en_US.iso885915\n> > en_US.utf8\n> > \n> > This is on Fedora 13, which is not exactly exotic. What is your system?\n> > \n> \n\nI've checked carefully multiple system and configuration, and found why\nwe have some little locale problem here.\n\nSince glibc 2.3, a file can hold all locales in file \"locale-archive\"\ninstead of having a tons of directory. To store all the locales in this\nfile, it uses an index based on a \"normalized\" codeset, e.g. it converts\ncodeset to lowercase, removes dash and minus. \nSo when one ask for the locale list, locale first go through the\n\"locale-archive\" content and report normalized codeset (utf8)  instead\nof canonical codeset (UTF-8), then it proceed with the legacy locales\ndirectories, using for them the canonical codeset.\n\nUntil recently, Mandriva Linux doesn't make use of \"locale-archive\", so\nUTF-8 locales were reported. Version in development uses\n\"locale-archive\" + legacy locale directories, hence the mix I've\nreported. Other Linux distributions like Fedora and Ubuntu uses only\n\"locale-archive\" and so, have only \"normalized\" codeset.\nPOSIX doesn't specify the output of locale -a, so it's not really a bug\nto show \"normalized\" codeset name.\n\nBut all others \"POSIX\" system I've found report \"canonical\" codeset,\ne.g. UTF-8 (all but latest cygwin). \n\nHere's the bug report:\nhttp://sourceware.org/bugzilla/show_bug.cgi?id=11629\n\nBTW, I will shortly provided a fix for the testcase, which will handle\nall cases.\n\nRegards.\n\n-- \nYann Droneaud\n"},{"id":"142249","messageId":"4BFB7D60.6090602@drmicha.warpmail.net","threadId":"23837","inReplyTo":"1274720888.4838.13.camel@localhost","subject":"Re: [PATCH] t9129: fix UTF-8 locale detection","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2010-05-25T07:33:52Z","receivedAt":"2010-05-25T07:33:52Z","isPatch":true,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Yann Droneaud venit, vidit, dixit 24.05.2010 19:08:\n> Le mardi 18 mai 2010 à 19:08 +0200, Yann Droneaud a écrit :\n>> Le mardi 18 mai 2010 à 18:05 +0200, Michael J Gruber a écrit :\n>>> Yann Droneaud venit, vidit, dixit 18.05.2010 16:41:\n>>>> Since I don't have en_US.utf8, some tests failed:\n>>\n>>>>\n>>>> On my system locale -a reports:\n>>>>\n>>>>    en_US\n>>>>    en_US.ISO-8859-1\n>>>>    en_US.UTF-8\n>>>>\n>>>\n>>> locale -a|grep en_US\n>>> en_US\n>>> en_US.iso88591\n>>> en_US.iso885915\n>>> en_US.utf8\n>>>\n>>> This is on Fedora 13, which is not exactly exotic. What is your system?\n>>>\n>>\n> \n> I've checked carefully multiple system and configuration, and found why\n> we have some little locale problem here.\n> \n> Since glibc 2.3, a file can hold all locales in file \"locale-archive\"\n> instead of having a tons of directory. To store all the locales in this\n> file, it uses an index based on a \"normalized\" codeset, e.g. it converts\n> codeset to lowercase, removes dash and minus. \n> So when one ask for the locale list, locale first go through the\n> \"locale-archive\" content and report normalized codeset (utf8)  instead\n> of canonical codeset (UTF-8), then it proceed with the legacy locales\n> directories, using for them the canonical codeset.\n> \n> Until recently, Mandriva Linux doesn't make use of \"locale-archive\", so\n> UTF-8 locales were reported. Version in development uses\n> \"locale-archive\" + legacy locale directories, hence the mix I've\n> reported. Other Linux distributions like Fedora and Ubuntu uses only\n> \"locale-archive\" and so, have only \"normalized\" codeset.\n> POSIX doesn't specify the output of locale -a, so it's not really a bug\n> to show \"normalized\" codeset name.\n> \n> But all others \"POSIX\" system I've found report \"canonical\" codeset,\n> e.g. UTF-8 (all but latest cygwin). \n> \n> Here's the bug report:\n> http://sourceware.org/bugzilla/show_bug.cgi?id=11629\n\nThanks a lot for doing the leg work! Is there any way to, say,\nset_local(a) and check whether get_locale() == a up to equivalence?\n\n> \n> BTW, I will shortly provided a fix for the testcase, which will handle\n> all cases.\n> \n> Regards.\n> \n\nThanks,\nMichael\n"},{"id":"142836","messageId":"7vmxvdckmf.fsf_-_@alter.siamese.dyndns.org","threadId":"23837","inReplyTo":"alpine.LFD.2.00.1005181037250.4195@i5.linux-foundation.org","subject":"Re* [PATCH] t9129: fix UTF-8 locale detection","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2010-06-02T19:14:32Z","receivedAt":"2010-06-02T19:14:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Wouldn't it be easier to just make it ignore case, and do\n>\n> \tgrep -qi '^en_US\\.utf-?8$'\n>\n> instead?\n>\n> I'm also not entirely sure you want to make that pattern stricter - the \n> whole problem with the old pattern was that it was too exact, so why add \n> the beginning/end requirement?\n\nSorry for being late to the party...\n\nThe prerequisite test is supposed to protect a real test that does this:\n\n\tLC_ALL=en_US.UTF-8 svn log `git svn info --url` | perl -w -e '...'\n\nand the original patch at least matches what we check with what we\nactually ask for from the system.\n\nI don't know if the above \"svn log\" test would still work if we run it\nunder any locale with UTF-8 (I checked with ja_JP.UTF-8 and it seems to be\nOk), but if it does, then a patch like this might be a better alternative.\n\n t/t9129-git-svn-i18n-commitencoding.sh |   20 +++++++++++++-------\n 1 files changed, 13 insertions(+), 7 deletions(-)\n\ndiff --git a/t/t9129-git-svn-i18n-commitencoding.sh b/t/t9129-git-svn-i18n-commitencoding.sh\nindex b9224bd..1e9a2eb 100755\n--- a/t/t9129-git-svn-i18n-commitencoding.sh\n+++ b/t/t9129-git-svn-i18n-commitencoding.sh\n@@ -14,10 +14,22 @@ compare_git_head_with () {\n \ttest_cmp current \"$1\"\n }\n \n+a_utf8_locale=$(locale -a | sed -n '/\\.[uU][tT][fF]-*8$/{\n+\tp\n+\tq\n+}')\n+\n+if test -n \"$a_utf8_locale\"\n+then\n+\ttest_set_prereq UTF8\n+else\n+\tsay \"UTF-8 locale not available, some tests are skipped\"\n+fi\n+\n compare_svn_head_with () {\n \t# extract just the log message and strip out committer info.\n \t# don't use --limit here since svn 1.1.x doesn't have it,\n-\tLC_ALL=en_US.UTF-8 svn log `git svn info --url` | perl -w -e '\n+\tLC_ALL=\"$a_utf8_locale\" svn log `git svn info --url` | perl -w -e '\n \t\tuse bytes;\n \t\t$/ = (\"-\"x72) . \"\\n\";\n \t\tmy @x = <STDIN>;\n@@ -69,12 +81,6 @@ do\n \t'\n done\n \n-if locale -a |grep -q en_US.utf8; then\n-\ttest_set_prereq UTF8\n-else\n-\tsay \"UTF-8 locale not available, test skipped\"\n-fi\n-\n test_expect_success UTF8 'ISO-8859-1 should match UTF-8 in svn' '\n \t(\n \t\tcd ISO8859-1 &&\n"},{"id":"159027","messageId":"cover.1294312018.git.yann@droneaud.fr","threadId":"23837","inReplyTo":"1274720888.4838.13.camel@localhost","subject":"[PATCH/RFC 0/4] en_US.UTF-8 locale detection","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2011-01-06T14:22:13Z","receivedAt":"2011-01-06T14:22:13Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Following discussions[1] about test t9129 UTF-8 locale detection,\nhere's a generic rewrite of the en_US.UTF-8 locale detection\nto be used with all other tests.\n\n[1] http://thread.gmane.org/gmane.comp.version-control.git/147283\n\nThe proposed mechanism could work for system without \"locale\" command,\nand would detect en_US.UTF-8 locale when named en_US.utf8 or some other\nvariations.\n\nIt must be tested on a wider range of systems (especially non-Linux).\n\nYann Droneaud (4):\n  test: added a library to detect an en_US.UTF-8 locale\n  test-lib.sh: added test_utf8() function\n  test: use test_utf8 and GIT_LC_UTF8 where an en_US.UTF-8 locale is\n    required\n  t9129: use \"$PERL_PATH\" instead of \"perl\"\n\n t/lib-locale.pl                        |  167 ++++++++++++++++++++++++++++++++\n t/t9100-git-svn-basic.sh               |   25 ++---\n t/t9129-git-svn-i18n-commitencoding.sh |   13 +--\n t/test-lib.sh                          |   14 +++\n 4 files changed, 192 insertions(+), 27 deletions(-)\n create mode 100755 t/lib-locale.pl\n\n-- \n1.7.3.4\n"},{"id":"159028","messageId":"4490926da28dcbfedc779cd32c5a59e20a1b55a2.1294312018.git.yann@droneaud.fr","threadId":"23837","inReplyTo":"cover.1294312018.git.yann@droneaud.fr","subject":"[PATCH 1/4] test: add a library to detect an en_US.UTF-8 locale","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2011-01-06T14:22:14Z","receivedAt":"2011-01-06T14:22:14Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Since one can't rely on \"locale\" command availability nor its -a output,\nperl script t/lib-locale.pl first use setlocale() and langinfo(CODESET)\nto search for a working en_US.UTF-8 locale among many name variants.\n\nIf this fail, the script fallback to \"locale\" usage with two steps:\n- try the \"charmap\" keyword, for example LC_ALL=en_US locale charmap\n- then try \"-a\" option and match a pattern looking to the\n  locale names\n\nSigned-off-by: Yann Droneaud <yann@droneaud.fr>\n---\n t/lib-locale.pl |  167 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 files changed, 167 insertions(+), 0 deletions(-)\n create mode 100755 t/lib-locale.pl\n\ndiff --git a/t/lib-locale.pl b/t/lib-locale.pl\nnew file mode 100755\nindex 0000000..33d47a5\n--- /dev/null\n+++ b/t/lib-locale.pl\n@@ -0,0 +1,167 @@\n+#!/usr/bin/env perl\n+\n+#\n+# Copyright (c) 2011 Yann Droneaud <yann@droneaud.fr>\n+#\n+# This program is free software: you can redistribute it and/or modify\n+# it under the terms of the GNU General Public License as published by\n+# the Free Software Foundation, either version 2 of the License, or\n+# (at your option) any later version.\n+#\n+# This program is distributed in the hope that it will be useful,\n+# but WITHOUT ANY WARRANTY; without even the implied warranty of\n+# MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\n+# GNU General Public License for more details.\n+#\n+# You should have received a copy of the GNU General Public License\n+# along with this program.  If not, see http://www.gnu.org/licenses/ .\n+#\n+\n+use strict;\n+if ($] > 5.004) {\n+    use warnings; # unless $] < 5.0004;\n+}\n+use POSIX;\n+\n+# according to manpage: setlocale can be used starting from perl 5.004\n+# starting from May 1997.\n+require 5.004;\n+\n+# it can also be tested with  perl -V:d_setlocale\n+# which must return something like \"d_setlocale='define';\"\n+\n+my $default_locale = setlocale(LC_ALL);\n+\n+# try a locale and return its codeset if available\n+sub locale_get_codeset\n+{\n+    my ($locale) = @_;\n+    my $codeset;\n+    my $ret;\n+\n+    $ret = setlocale(LC_ALL, $locale);\n+    if (!defined($ret)) {\n+\t# print \"can't set locale $locale\\n\";\n+\treturn undef;\n+    }\n+\n+    #\n+    # I18N::Langinfo is not available everywhere\n+    # use the trick from perldoc I18N::Langinfo\n+    #\n+    $codeset = eval {\n+\trequire I18N::Langinfo;\n+\tI18N::Langinfo->import(qw(langinfo CODESET));\n+\tlanginfo(CODESET()); # note the ()\n+    };\n+\n+    if (defined($@) && $@) {\n+\tprint STDERR \"error while testing codeset:\\n $@\";\n+\treturn undef;\n+    }\n+\n+    # restore locale\n+    setlocale(LC_ALL, $default_locale);\n+\n+    return $codeset;\n+}\n+\n+# check that given codeset match UTF-8\n+sub codeset_check\n+{\n+    my ($codeset) = @_;\n+\n+    if (defined($codeset) && $codeset =~ /^UTF[_-]?8$/i) {\n+\treturn 1;\n+    }\n+\n+    return 0;\n+}\n+\n+# try all those locales\n+# require en_US with UTF-8\n+my @locales = ( \"en_US.UTF-8\",\n+\t\t\"en_US.utf8\",\n+\t\t\"en_US.UTF8\",\n+\t\t\"en_US.utf-8\",\n+\t\t\"en_US.UTF_8\",\n+\t\t\"en_US.utf_8\",\n+\t\t\"en_US\" );\n+\n+my $locale;\n+\n+my $codeset;\n+\n+#\n+# Check locale from within perl\n+#\n+foreach $locale (@locales) {\n+\n+    $codeset = locale_get_codeset($locale);\n+\n+    if (codeset_check($codeset)) {\n+\tprint $locale, \"\\n\";\n+\texit 0;\n+    }\n+}\n+\n+#\n+# if 'locale' command is available, test 'locale charmap'\n+#\n+foreach $locale (@locales) {\n+\n+    $codeset = `LC_ALL=$locale locale charmap 2>/dev/null`;\n+\n+    # if any error, skip\n+    if ($!) {\n+\tlast;\n+    }\n+\n+    if (codeset_check($codeset)) {\n+\tprint $locale, \"\\n\";\n+\texit 0;\n+    }\n+}\n+\n+#\n+# try to execute \"locale -a\" command\n+# this command is not always available and\n+# output format is not normalized\n+#\n+use IPC::Open3;\n+use File::Spec;\n+\n+open(NULLR, \"<\", File::Spec->devnull) || die \"Can't open devnull: $!\";\n+open(NULLW, \">\", File::Spec->devnull) || die \"Can't open devnull: $!\";\n+\n+my $pid = open3(\"<&NULLR\", \\*LOCALES, \">&NULLW\" , \"locale\", \"-a\") || die \"Can't launch locale -a: $!\";\n+\n+while(<LOCALES>) {\n+     chomp;\n+     if (/(en_US\\.([\\w-]+))/i) {\n+\t if (codeset_check($2)) {\n+\t     $locale = $1;\n+\t     last;\n+\t }\n+     }\n+}\n+\n+waitpid($pid, 0);\n+\n+my $retcode = $?;\n+\n+close(LOCALES);\n+\n+if ($retcode == 0) {\n+    if (defined $locale) {\n+\tprint \"$locale\\n\";\n+\texit 0;\n+    }\n+}\n+\n+#\n+# Nothing available,\n+# at last, return an error\n+#\n+\n+exit 1;\n-- \n1.7.3.4\n"},{"id":"159029","messageId":"8559d90bff6fca1c18f1cbf3530f2f4cc695f9f4.1294312018.git.yann@droneaud.fr","threadId":"23837","inReplyTo":"cover.1294312018.git.yann@droneaud.fr","subject":"[PATCH 2/4] test-lib.sh: add test_utf8() function","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2011-01-06T14:22:15Z","receivedAt":"2011-01-06T14:22:15Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"test_utf8() checks for en_US.UTF-8 locale availability using lib-locale.pl.\nThe function returns 1 if a locale was not found, otherwise it returns 0,\nset prereq UTF8 and export GIT_LC_UTF8 with the locale name.\n\nIf a test needs to use an en_US.UTF-8 locale, it has to call test_utf8() first.\nThen it can do tests based on prereq UTF8 availability and use LC_ALL=$GIT_LC_UTF8.\n\nSigned-off-by: Yann Droneaud <yann@droneaud.fr>\n---\n t/test-lib.sh |   14 ++++++++++++++\n 1 files changed, 14 insertions(+), 0 deletions(-)\n\ndiff --git a/t/test-lib.sh b/t/test-lib.sh\nindex cb1ca97..3e92360 100644\n--- a/t/test-lib.sh\n+++ b/t/test-lib.sh\n@@ -1066,6 +1066,20 @@ case $(uname -s) in\n \t;;\n esac\n \n+# check UTF-8 locale availability\n+test_utf8 () {\n+    if test_have_prereq PERL ; then\n+\t# output an en_US.UTF-8 locale compatible name\n+\tGIT_LC_UTF8=`$PERL_PATH $GIT_BUILD_DIR/t/lib-locale.pl`\n+    fi\n+    if test -z \"$GIT_LC_UTF8\" ; then\n+\treturn 1\n+    else\n+\ttest_set_prereq UTF8\n+\treturn 0\n+    fi\n+}\n+\n test -z \"$NO_PERL\" && test_set_prereq PERL\n test -z \"$NO_PYTHON\" && test_set_prereq PYTHON\n \n-- \n1.7.3.4\n"},{"id":"159030","messageId":"97423472c08cd83373c769bf1cafdb9b85db37e3.1294312018.git.yann@droneaud.fr","threadId":"23837","inReplyTo":"cover.1294312018.git.yann@droneaud.fr","subject":"[PATCH 3/4] test: use test_utf8 and GIT_LC_UTF8 where an en_US.UTF-8 locale is required","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2011-01-06T14:22:16Z","receivedAt":"2011-01-06T14:22:16Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Signed-off-by: Yann Droneaud <yann@droneaud.fr>\n---\n t/t9100-git-svn-basic.sh               |   25 ++++++++-----------------\n t/t9129-git-svn-i18n-commitencoding.sh |   13 +++----------\n 2 files changed, 11 insertions(+), 27 deletions(-)\n\ndiff --git a/t/t9100-git-svn-basic.sh b/t/t9100-git-svn-basic.sh\nindex b041516..fcdffef 100755\n--- a/t/t9100-git-svn-basic.sh\n+++ b/t/t9100-git-svn-basic.sh\n@@ -4,20 +4,14 @@\n #\n \n test_description='git svn basic tests'\n-GIT_SVN_LC_ALL=${LC_ALL:-$LANG}\n \n . ./lib-git-svn.sh\n \n say 'define NO_SVN_TESTS to skip git svn tests'\n \n-case \"$GIT_SVN_LC_ALL\" in\n-*.UTF-8)\n-\ttest_set_prereq UTF8\n-\t;;\n-*)\n-\tsay \"# UTF-8 locale not set, some tests skipped ($GIT_SVN_LC_ALL)\"\n-\t;;\n-esac\n+if ! test_utf8 ; then\n+\tsay \"# UTF-8 locale not set, some tests skipped\"\n+fi\n \n test_expect_success \\\n     'initialize git svn' '\n@@ -172,15 +166,12 @@ test_expect_success \"$name\" '\n \ttest ! -h \"$SVN_TREE\"/exec-2.sh &&\n \ttest_cmp help \"$SVN_TREE\"/exec-2.sh'\n \n-name=\"commit with UTF-8 message: locale: $GIT_SVN_LC_ALL\"\n-LC_ALL=\"$GIT_SVN_LC_ALL\"\n-export LC_ALL\n+name=\"commit with UTF-8 message: locale: $GIT_LC_UTF8\"\n test_expect_success UTF8 \"$name\" \"\n-\techo '# hello' >> exec-2.sh &&\n-\tgit update-index exec-2.sh &&\n-\tgit commit -m 'éï∏' &&\n-\tgit svn set-tree HEAD\"\n-unset LC_ALL\n+\tLC_ALL=$GIT_LC_UTF8 echo '# hello' >> exec-2.sh &&\n+\tLC_ALL=$GIT_LC_UTF8 git update-index exec-2.sh &&\n+\tLC_ALL=$GIT_LC_UTF8 git commit -m 'éï∏' &&\n+\tLC_ALL=$GIT_LC_UTF8 git svn set-tree HEAD\"\n \n name='test fetch functionality (svn => git) with alternate GIT_SVN_ID'\n GIT_SVN_ID=alt\ndiff --git a/t/t9129-git-svn-i18n-commitencoding.sh b/t/t9129-git-svn-i18n-commitencoding.sh\nindex 8cfdfe7..f3bbde4 100755\n--- a/t/t9129-git-svn-i18n-commitencoding.sh\n+++ b/t/t9129-git-svn-i18n-commitencoding.sh\n@@ -14,22 +14,15 @@ compare_git_head_with () {\n \ttest_cmp current \"$1\"\n }\n \n-a_utf8_locale=$(locale -a | sed -n '/\\.[uU][tT][fF]-*8$/{\n-\tp\n-\tq\n-}')\n-\n-if test -n \"$a_utf8_locale\"\n-then\n-\ttest_set_prereq UTF8\n-else\n+if ! test_utf8 ; then\n \tsay \"# UTF-8 locale not available, some tests are skipped\"\n fi\n \n compare_svn_head_with () {\n \t# extract just the log message and strip out committer info.\n \t# don't use --limit here since svn 1.1.x doesn't have it,\n-\tLC_ALL=\"$a_utf8_locale\" svn log `git svn info --url` | perl -w -e '\n+\tLC_ALL=$GIT_LC_UTF8 svn log `git svn info --url` | \\\n+\t    LC_ALL=$GIT_LC_UTF8 perl -w -e '\n \t\tuse bytes;\n \t\t$/ = (\"-\"x72) . \"\\n\";\n \t\tmy @x = <STDIN>;\n-- \n1.7.3.4\n"},{"id":"159031","messageId":"74b3db2e538abb4c06d7a0792dff9d78636e2758.1294312018.git.yann@droneaud.fr","threadId":"23837","inReplyTo":"cover.1294312018.git.yann@droneaud.fr","subject":"[PATCH 4/4] t9129: use \"$PERL_PATH\" instead of \"perl\"","fromName":"Yann Droneaud","fromEmail":"yann@droneaud.fr","sentAt":"2011-01-06T14:22:17Z","receivedAt":"2011-01-06T14:22:17Z","isPatch":true,"sender":{"key":"yann@droneaud.fr","avatar":null},"body":"Signed-off-by: Yann Droneaud <yann@droneaud.fr>\n---\n t/t9129-git-svn-i18n-commitencoding.sh |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/t/t9129-git-svn-i18n-commitencoding.sh b/t/t9129-git-svn-i18n-commitencoding.sh\nindex f3bbde4..6dda569 100755\n--- a/t/t9129-git-svn-i18n-commitencoding.sh\n+++ b/t/t9129-git-svn-i18n-commitencoding.sh\n@@ -22,7 +22,7 @@ compare_svn_head_with () {\n \t# extract just the log message and strip out committer info.\n \t# don't use --limit here since svn 1.1.x doesn't have it,\n \tLC_ALL=$GIT_LC_UTF8 svn log `git svn info --url` | \\\n-\t    LC_ALL=$GIT_LC_UTF8 perl -w -e '\n+\t    LC_ALL=$GIT_LC_UTF8 $PERL_PATH -w -e '\n \t\tuse bytes;\n \t\t$/ = (\"-\"x72) . \"\\n\";\n \t\tmy @x = <STDIN>;\n-- \n1.7.3.4\n"},{"id":"159129","messageId":"7vy66w7e6j.fsf@alter.siamese.dyndns.org","threadId":"23837","inReplyTo":"4490926da28dcbfedc779cd32c5a59e20a1b55a2.1294312018.git.yann@droneaud.fr","subject":"Re: [PATCH 1/4] test: add a library to detect an en_US.UTF-8 locale","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-07T18:37:40Z","receivedAt":"2011-01-07T18:37:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Yann Droneaud <yann@droneaud.fr> writes:\n\nYann Droneaud <yann@droneaud.fr> writes:\n\n> Since one can't rely on \"locale\" command availability nor its -a output,\n> perl script t/lib-locale.pl first use setlocale() and langinfo(CODESET)\n> to search for a working en_US.UTF-8 locale among many name variants.\n>\n> If this fail, the script fallback to \"locale\" usage with two steps:\n> - try the \"charmap\" keyword, for example LC_ALL=en_US locale charmap\n> - then try \"-a\" option and match a pattern looking to the\n>   locale names\n>\n> Signed-off-by: Yann Droneaud <yann@droneaud.fr>\n> ---\n>  t/lib-locale.pl |  167 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n>  1 files changed, 167 insertions(+), 0 deletions(-)\n>  create mode 100755 t/lib-locale.pl\n\nThe remainder of the series looks more or less sane, but a new 170-line\nscript somehow feels a bit overengineered solution to a minor problem that\nhas already been solved in a much simpler way.\n\n> +# try to execute \"locale -a\" command\n> +# this command is not always available and\n> +# output format is not normalized\n> +#\n> +use IPC::Open3;\n> +use File::Spec;\n> +\n> +open(NULLR, \"<\", File::Spec->devnull) || die \"Can't open devnull: $!\";\n> +open(NULLW, \">\", File::Spec->devnull) || die \"Can't open devnull: $!\";\n> +\n> +my $pid = open3(\"<&NULLR\", \\*LOCALES, \">&NULLW\" , \"locale\", \"-a\") || die \"Can't launch locale -a: $!\";\n> +\n> +while(<LOCALES>) {\n> +     chomp;\n> +     if (/(en_US\\.([\\w-]+))/i) {\n> +\t if (codeset_check($2)) {\n> +\t     $locale = $1;\n> +\t     last;\n> +\t }\n> +     }\n> +}\n> +\n> +waitpid($pid, 0);\n\nYou are trying to buy something with the complexity of Open3 for doing\nthis logic, over a bog-naive \"for (`locale -a`) { ... }\", but is that\nsomething really worth the money?\n\n\"lib-locale.pl\" is a gross misnomer for this script.\n\nIt may be a good helper to be used in 2/4 (test_utf8), but later people\nmay want to have more helper feature in something called \"lib-locale\", not\njust \"pick a single UTF-8 locale randomly from available ones on the\nsystem\" (especially when ab/i18n starts moving again).  I'd suggest either\nto (1) rename it \"pick-utf8-locale.pl\" (i.e. honest naming) or (2) prepare\nthe helper command to be extensible from the beginning, i.e. require a\ncommand word e.g. \"lib-locale.pl pick-utf8-locale\" to trigger the\ncurrently implemented feature (i.e. forward looking naming).\n"},{"id":"159132","messageId":"7vtyhk7du0.fsf@alter.siamese.dyndns.org","threadId":"23837","inReplyTo":"8559d90bff6fca1c18f1cbf3530f2f4cc695f9f4.1294312018.git.yann@droneaud.fr","subject":"Re: [PATCH 2/4] test-lib.sh: add test_utf8() function","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-07T18:45:11Z","receivedAt":"2011-01-07T18:45:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Yann Droneaud <yann@droneaud.fr> writes:\n\n> +# check UTF-8 locale availability\n> +test_utf8 () {\n> +    if test_have_prereq PERL ; then\n> +\t# output an en_US.UTF-8 locale compatible name\n> +\tGIT_LC_UTF8=`$PERL_PATH $GIT_BUILD_DIR/t/lib-locale.pl`\n> +    fi\n> +    if test -z \"$GIT_LC_UTF8\" ; then\n> +\treturn 1\n> +    else\n> +\ttest_set_prereq UTF8\n> +\treturn 0\n> +    fi\n> +}\n\nNice abstraction to have a helper function that picks a locale to be used\nwhen we want to test UTF-8 thingy.  Perhaps pick_utf8_locale might be a\nbetter name, though. \n\nThe comment in the function is not wrong per-se, but it and the\nimplementation in 1/4 may be too restrictive---all it needs to do is to\npick a locale that is UTF-8, and it does not necessarily have to be en_US,\nno?\n"},{"id":"159133","messageId":"7vpqs87dob.fsf@alter.siamese.dyndns.org","threadId":"23837","inReplyTo":"97423472c08cd83373c769bf1cafdb9b85db37e3.1294312018.git.yann@droneaud.fr","subject":"Re: [PATCH 3/4] test: use test_utf8 and GIT_LC_UTF8 where an en_US.UTF-8 locale is required","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-07T18:48:36Z","receivedAt":"2011-01-07T18:48:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Yann Droneaud <yann@droneaud.fr> writes:\n\n> Signed-off-by: Yann Droneaud <yann@droneaud.fr>\n> ---\n>  t/t9100-git-svn-basic.sh               |   25 ++++++++-----------------\n>  t/t9129-git-svn-i18n-commitencoding.sh |   13 +++----------\n>  2 files changed, 11 insertions(+), 27 deletions(-)\n\nBoth are nice changes; this patch shows the earlier abstraction in 2/4 is\nthe right direction to go.\n\n>  compare_svn_head_with () {\n>  \t# extract just the log message and strip out committer info.\n>  \t# don't use --limit here since svn 1.1.x doesn't have it,\n> -\tLC_ALL=\"$a_utf8_locale\" svn log `git svn info --url` | perl -w -e '\n> +\tLC_ALL=$GIT_LC_UTF8 svn log `git svn info --url` | \\\n> +\t    LC_ALL=$GIT_LC_UTF8 perl -w -e '\n\nStyle.\n\n\tLC_ALL=... svn log ... |\n        LC_ALL=... perl -w -e '\n        \t...\n\t'\n\nWhen you end a line with '|', the shell knows that you haven't finished\ntalking to it, so there is no need for the trailing bs-lf there.  Indent\nthe downstream of the pipe to the same level as the upstream.\n"},{"id":"159134","messageId":"7vlj2w7dni.fsf@alter.siamese.dyndns.org","threadId":"23837","inReplyTo":"74b3db2e538abb4c06d7a0792dff9d78636e2758.1294312018.git.yann@droneaud.fr","subject":"Re: [PATCH 4/4] t9129: use \"$PERL_PATH\" instead of \"perl\"","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-01-07T18:49:05Z","receivedAt":"2011-01-07T18:49:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Squash this to 3/4, perhaps saying \"while at it, fix $this\".\n"}]}