git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Fix UTF Encoding issue

From
IDIsmail Dönmez <ismail@pardus.org.tr>
Date
Dec 4, 2007, 08:12 UTC
Message-ID
<200712041012.50935.ismail@pardus.org.tr>
In-Reply-To
<20071204080407.GC31042@auto.tuwien.ac.at>
Tuesday 04 December 2007 10:04:07 Martin Koegler yazmıştı:
Show 52 quoted lines
> On Tue, Dec 04, 2007 at 08:16:24AM +1030, Benjamin Close wrote:
> > Jakub Narebski wrote:
> > >On Mon, 3 Dec 2007, Martin Koegler wrote:
> > >>On Mon, Dec 03, 2007 at 04:06:48AM -0800, Jakub Narebski wrote:
> > >>>Ismail Dönmez <ismail@pardus.org.tr> writes:
> > >>>>Monday 03 December 2007 Tarihinde 12:14:43 yazm??t?:
> > >>>>>Benjamin Close <Benjamin.Close@clearchain.com> writes:
> > >>>>>>-	eval { $res = decode_utf8($str, Encode::FB_CROAK); };
> > >>>>>>-	if (defined $res) {
> > >>>>>>-		return $res;
> > >>>>>>-	} else {
> > >>>>>>-		return decode($fallback_encoding, $str,
> > >>>>>>Encode::FB_DEFAULT);
> > >>>>>>-	}
> > >>>>>>+	eval { return ($res = decode_utf8($str, Encode::FB_CROAK));
> > >>>>>>};
> > >>>>>>+	return decode($fallback_encoding, $str, Encode::FB_DEFAULT);
> > >>>>>> }
> > >>
> > >>This version is broken on Debian sarge and etch. Feeding a UTF-8 and a
> > >>latin1
> > >>encoding of the same character sequence yields to different results.
> >
> > For the record, this was on a debian sid machine.
> >
> > #perl --version
> > This is perl, v5.8.8 built for x86_64-linux-gnu-thread-multi
> >
> > and the result of not using the original patch was:
> >
> > <h1>Software error:</h1>
> > <pre>Cannot decode string with wide characters at
> > /usr/lib/perl/5.8/Encode.pm line 166.
> > </pre>
> >
> >
> > I haven't tried the other solutions tested here.
>
> Debian etch also has v5.8.8.
>
> My main question is, why is the error not catched?
>
> I'm not a perl programmer, but in your patch the first line is a
> NOP. The return in eval seems to only returns from the eval block, so
> any text is decoded as latin1 with the second statement.
>
> In the original version, decode($fallback_encoding, $str,
> Encode::FB_DEFAULT) can not emit an error, else it would in your
> version too.
>
> In your version, eval is able to surpress the error of
> decode_utf8($str, Encode::FB_CROAK);, but not in the original version.
I think just a better method is to use (not tested):
if( is_utf8($str) ) 
{
	return decode_utf8($str);
}
else {
	return decode($str);
}

Regards, ismail

-- 
Never learn by your mistakes, if you do you may never dare to try again.
Previous: Martin KoeglerNext: Martin Koegler
Message 13 of 24 in “Fix UTF Encoding issue”
  1. Benjamin CloseDec 3, 2007
  2. Junio C HamanoDec 3, 2007
  3. Ismail DönmezDec 3, 2007
  4. Jakub NarebskiDec 3, 2007
  5. Martin KoeglerDec 3, 2007
  6. Jakub NarebskiDec 3, 2007
  7. Benjamin CloseDec 3, 2007
  8. Ismail DönmezDec 3, 2007
  9. Benjamin CloseDec 3, 2007
  10. Jakub NarebskiDec 3, 2007
  11. Ismail DönmezDec 4, 2007
  12. Martin KoeglerDec 4, 2007
  13. Ismail DönmezDec 4, 2007
  14. Martin KoeglerDec 4, 2007
  15. Martin KoeglerDec 4, 2007
  16. Ismail DönmezDec 4, 2007
  17. Martin KoeglerDec 4, 2007
  18. Ismail DönmezDec 4, 2007
  19. Ismail DönmezDec 4, 2007
  20. Martin KoeglerDec 4, 2007
  21. Ismail DönmezDec 4, 2007
  22. Ismail DönmezDec 4, 2007
  23. Jakub NarebskiDec 4, 2007
  24. Wincent ColaiutaDec 4, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.