git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH] t9129: fix UTF-8 locale detection

From
YDYann Droneaud <yann@droneaud.fr>
Date
May 24, 2010, 17:08 UTC
Message-ID
<1274720888.4838.13.camel@localhost>
In-Reply-To
<1274202486.4228.22.camel@localhost>
Le mardi 18 mai 2010 à 19:08 +0200, Yann Droneaud a écrit :
Show 21 quoted lines
> Le mardi 18 mai 2010 à 18:05 +0200, Michael J Gruber a écrit :
> > Yann Droneaud venit, vidit, dixit 18.05.2010 16:41:
> > > Since I don't have en_US.utf8, some tests failed:
> 
> > > 
> > > On my system locale -a reports:
> > > 
> > >    en_US
> > >    en_US.ISO-8859-1
> > >    en_US.UTF-8
> > > 
> > 
> > locale -a|grep en_US
> > en_US
> > en_US.iso88591
> > en_US.iso885915
> > en_US.utf8
> > 
> > This is on Fedora 13, which is not exactly exotic. What is your system?
> > 
> 

I've checked carefully multiple system and configuration, and found why we have some little locale problem here.

Since glibc 2.3, a file can hold all locales in file "locale-archive" instead of having a tons of directory. To store all the locales in this file, it uses an index based on a "normalized" codeset, e.g. it converts codeset to lowercase, removes dash and minus. So when one ask for the locale list, locale first go through the "locale-archive" content and report normalized codeset (utf8) instead of canonical codeset (UTF-8), then it proceed with the legacy locales directories, using for them the canonical codeset.

Until recently, Mandriva Linux doesn't make use of "locale-archive", so UTF-8 locales were reported. Version in development uses "locale-archive" + legacy locale directories, hence the mix I've reported. Other Linux distributions like Fedora and Ubuntu uses only "locale-archive" and so, have only "normalized" codeset. POSIX doesn't specify the output of locale -a, so it's not really a bug to show "normalized" codeset name.

But all others "POSIX" system I've found report "canonical" codeset, e.g. UTF-8 (all but latest cygwin).

Here's the bug report: http://sourceware.org/bugzilla/show_bug.cgi?id=11629

BTW, I will shortly provided a fix for the testcase, which will handle all cases.

Regards.
-- 
Yann Droneaud
Previous: Junio C HamanoNext: Michael J Gruber
Message 11 of 23 in “t9129: fix UTF-8 locale detection”
  1. t9129: fix UTF-8 locale detectionYann Droneaud, May 18, 2010
  2. Michael J GruberMay 18, 2010
  3. Yann DroneaudMay 18, 2010
  4. t9129: fix UTF-8 locale detectionYann Droneaud, May 18, 2010
  5. Linus TorvaldsMay 18, 2010
  6. Andreas SchwabMay 18, 2010
  7. Linus TorvaldsMay 18, 2010
  8. Yann DroneaudMay 18, 2010
  9. Yann DroneaudMay 19, 2010
  10. Re* [PATCH] t9129: fix UTF-8 locale detectionJunio C Hamano, Jun 2, 2010
  11. Yann DroneaudMay 24, 2010
  12. Michael J GruberMay 25, 2010
  13. 0/4 en_US.UTF-8 locale detectionYann Droneaud, Jan 6, 2011
  14. 1/4 test: add a library to detect an en_US.UTF-8 localeYann Droneaud, Jan 6, 2011
  15. Junio C HamanoJan 7, 2011
  16. 2/4 test-lib.sh: add test_utf8() functionYann Droneaud, Jan 6, 2011
  17. Junio C HamanoJan 7, 2011
  18. 3/4 test: use test_utf8 and GIT_LC_UTF8 where an en_US.UTF-8 locale is requiredYann Droneaud, Jan 6, 2011
  19. Junio C HamanoJan 7, 2011
  20. 4/4 t9129: use "$PERL_PATH" instead of "perl"Yann Droneaud, Jan 6, 2011
  21. Junio C HamanoJan 7, 2011
  22. Yann DroneaudMay 18, 2010
  23. Miles BaderMay 19, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.