{"thread":{"id":"11118","subject":"Fix UTF Encoding issue","startedAt":"2007-12-03T10:02:01Z","lastAt":"2007-12-04T10:11:04Z","messageCount":24,"participants":["Benjamin Close","Junio C Hamano","Ismail Dönmez","Jakub Narebski","Martin Koegler","Wincent Colaiuta"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"61759","messageId":"4753D419.80503@clearchain.com","threadId":"11118","inReplyTo":null,"subject":"Fix UTF Encoding issue","fromName":"Benjamin Close","fromEmail":"benjamin.close@clearchain.com","sentAt":"2007-12-03T10:02:01Z","receivedAt":"2007-12-03T10:02:01Z","isPatch":false,"sender":{"key":"benjamin.close@clearchain.com","avatar":"https://gravatar.com/avatar/19dca0aa372edfa901b93c556dfda2e78ad4434558fe4d139598e086315d714a?d=mp&s=160"},"body":">From 83042abf3967b455953cddeab43e33c1d59c6f03 Mon Sep 17 00:00:00 2001\nFrom: Benjamin Close <Benjamin.Close@clearchain.com>\nDate: Sun, 2 Dec 2007 15:09:00 -0800\nSubject: [PATCH] Gitweb: Fix encoding to always translate rather than \nsometimes fail\n\nWhen performing the utf translation don't test if $res is defined.\nIt appears that it is defined even when the conversion fails. This causes\nfailures on the writing of the output stream which is expecting UTF.\n\nInstead, immediately return if conversion is successful else force\nthe translation to the fallback encoding\n---\n  gitweb/gitweb.perl |    8 ++------\n  1 files changed, 2 insertions(+), 6 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex 491a3f4..00bbcdf 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -696,12 +696,8 @@ sub validate_refname {\n  sub to_utf8 {\n  \tmy $str = shift;\n  \tmy $res;\n-\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n-\tif (defined $res) {\n-\t\treturn $res;\n-\t} else {\n-\t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n-\t}\n+\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\n+\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n  }\n\n  # quote unsafe chars, but keep the slash, even when it's not\n-- \n1.5.3.6\n"},{"id":"61761","messageId":"7v7ijwjd9o.fsf@gitster.siamese.dyndns.org","threadId":"11118","inReplyTo":"4753D419.80503@clearchain.com","subject":"Re: Fix UTF Encoding issue","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-12-03T10:14:43Z","receivedAt":"2007-12-03T10:14:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Benjamin Close <Benjamin.Close@clearchain.com> writes:\n\n>>From 83042abf3967b455953cddeab43e33c1d59c6f03 Mon Sep 17 00:00:00 2001\n> From: Benjamin Close <Benjamin.Close@clearchain.com>\n> Date: Sun, 2 Dec 2007 15:09:00 -0800\n> Subject: [PATCH] Gitweb: Fix encoding to always translate rather than\n> sometimes fail\n>\n> When performing the utf translation don't test if $res is defined.\n> It appears that it is defined even when the conversion fails. This causes\n> failures on the writing of the output stream which is expecting UTF.\n> @@ -696,12 +696,8 @@ sub validate_refname {\n>  sub to_utf8 {\n>  \tmy $str = shift;\n>  \tmy $res;\n> -\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n> -\tif (defined $res) {\n> -\t\treturn $res;\n> -\t} else {\n> -\t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> -\t}\n> +\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\n> +\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n>  }\n\nThis is funny.\n\nI thought the standard catch ... throw idiom in Perl was to do the above\nlike this:\n\n\tmy $res;\n        eval { $res = decode_utf8($str, Encode::FB_CROAK); };\n        if ($@) {\n        \treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n\t}\n\treturn $res;\n\n(alternatively, you can assign return value of eval {} to $res).\n"},{"id":"61774","messageId":"200712031332.36187.ismail@pardus.org.tr","threadId":"11118","inReplyTo":"7v7ijwjd9o.fsf@gitster.siamese.dyndns.org","subject":"Re: Fix UTF Encoding issue","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-03T11:32:36Z","receivedAt":"2007-12-03T11:32:36Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Monday 03 December 2007 Tarihinde 12:14:43 yazmıştı:\n> Benjamin Close <Benjamin.Close@clearchain.com> writes:\n> >>From 83042abf3967b455953cddeab43e33c1d59c6f03 Mon Sep 17 00:00:00 2001\n> >\n> > From: Benjamin Close <Benjamin.Close@clearchain.com>\n> > Date: Sun, 2 Dec 2007 15:09:00 -0800\n> > Subject: [PATCH] Gitweb: Fix encoding to always translate rather than\n> > sometimes fail\n> >\n> > When performing the utf translation don't test if $res is defined.\n> > It appears that it is defined even when the conversion fails. This causes\n> > failures on the writing of the output stream which is expecting UTF.\n> > @@ -696,12 +696,8 @@ sub validate_refname {\n> >  sub to_utf8 {\n> >  \tmy $str = shift;\n> >  \tmy $res;\n> > -\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n> > -\tif (defined $res) {\n> > -\t\treturn $res;\n> > -\t} else {\n> > -\t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> > -\t}\n> > +\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\n> > +\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> >  }\n>\n> This is funny.\n>\n> I thought the standard catch ... throw idiom in Perl was to do the above\n> like this:\n>\n> \tmy $res;\n>         eval { $res = decode_utf8($str, Encode::FB_CROAK); };\n>         if ($@) {\n>         \treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> \t}\n> \treturn $res;\n\nI think this is correct, but the current code in gitweb doesn't look correct \nsince it checks for $res and not $@.\n\nRegards,\nismail\n\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n"},{"id":"61777","messageId":"m3prxougmx.fsf@roke.D-201","threadId":"11118","inReplyTo":"200712031332.36187.ismail@pardus.org.tr","subject":"Re: Fix UTF Encoding issue","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-12-03T12:06:48Z","receivedAt":"2007-12-03T12:06:48Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Ismail Dönmez <ismail@pardus.org.tr> writes:\n> Monday 03 December 2007 Tarihinde 12:14:43 yazmıştı:\n>> Benjamin Close <Benjamin.Close@clearchain.com> writes:\n\n>>> -\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n>>> -\tif (defined $res) {\n>>> -\t\treturn $res;\n>>> -\t} else {\n>>> -\t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n>>> -\t}\n>>> +\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\n>>> +\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n>>>  }\n>>\n>> I thought the standard catch ... throw idiom in Perl was to do the above\n>> like this:\n>>\n>> \tmy $res;\n>>         eval { $res = decode_utf8($str, Encode::FB_CROAK); };\n>>         if ($@) {\n>>         \treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n>> \t}\n>> \treturn $res;\n> \n> I think this is correct, but the current code in gitweb doesn't look correct \n> since it checks for $res and not $@.\n\nFirst version of the patch was created by Martin Koegler. I have\nparticipated in creating the version which is now in gitweb, but I\nhave to say that I wrote it based on decode_utf8\ndocumentation... which doesn't necessarily agree with facts :-(\n\nI'm all for the \"throw idion\" version. Ack.\n-- \nJakub Narebski\n"},{"id":"61788","messageId":"20071203163856.GA24269@auto.tuwien.ac.at","threadId":"11118","inReplyTo":"m3prxougmx.fsf@roke.D-201","subject":"Re: Fix UTF Encoding issue","fromName":"Martin Koegler","fromEmail":"mkoegler@auto.tuwien.ac.at","sentAt":"2007-12-03T16:38:56Z","receivedAt":"2007-12-03T16:38:56Z","isPatch":false,"sender":{"key":"mkoegler@auto.tuwien.ac.at","avatar":null},"body":"On Mon, Dec 03, 2007 at 04:06:48AM -0800, Jakub Narebski wrote:\n> Ismail Dönmez <ismail@pardus.org.tr> writes:\n> > Monday 03 December 2007 Tarihinde 12:14:43 yazm??t?:\n> >> Benjamin Close <Benjamin.Close@clearchain.com> writes:\n> >>> -\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n> >>> -\tif (defined $res) {\n> >>> -\t\treturn $res;\n> >>> -\t} else {\n> >>> -\t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> >>> -\t}\n> >>> +\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\n> >>> +\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> >>>  }\n\nThis version is broken on Debian sarge and etch. Feeding a UTF-8 and a latin1\nencoding of the same character sequence yields to different results.\n\n> >>\n> >> I thought the standard catch ... throw idiom in Perl was to do the above\n> >> like this:\n> >>\n> >> \tmy $res;\n> >>         eval { $res = decode_utf8($str, Encode::FB_CROAK); };\n> >>         if ($@) {\n> >>         \treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> >> \t}\n> >> \treturn $res;\n> > \n> > I think this is correct, but the current code in gitweb doesn't look correct \n> > since it checks for $res and not $@.\n> \n> First version of the patch was created by Martin Koegler. I have\n> participated in creating the version which is now in gitweb, but I\n> have to say that I wrote it based on decode_utf8\n> documentation... which doesn't necessarily agree with facts :-(\n\neval { $res = decode_utf8(...); }\nif ($@) \n     return decode(...);\nreturn $res\n\nor\n\neval { $res = decode_utf8(...); }\nif (defined $res)\n      return $res;\nelse\n    return decode(...);\n\nshow the same (wrong) behaviour on Debian sarge. They do not always\ndecode non UTF-8 characters correctly, eg.\n#öäü does not work\n#äöüä does work\n\nOn Debian etch, both versions are working.\n\n> I'm all for the \"throw idion\" version. Ack.\n\n\nmfg Martin Kögler\n"},{"id":"61786","messageId":"200712031802.55514.jnareb@gmail.com","threadId":"11118","inReplyTo":"20071203163856.GA24269@auto.tuwien.ac.at","subject":"Re: Fix UTF Encoding issue","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-12-03T17:02:54Z","receivedAt":"2007-12-03T17:02:54Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Mon, 3 Dec 2007, Martin Koegler wrote:\n> On Mon, Dec 03, 2007 at 04:06:48AM -0800, Jakub Narebski wrote:\n>> Ismail Dönmez <ismail@pardus.org.tr> writes:\n>>> Monday 03 December 2007 Tarihinde 12:14:43 yazm??t?:\n>>>> Benjamin Close <Benjamin.Close@clearchain.com> writes:\n>>>>> -\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n>>>>> -\tif (defined $res) {\n>>>>> -\t\treturn $res;\n>>>>> -\t} else {\n>>>>> -\t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n>>>>> -\t}\n>>>>> +\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\n>>>>> +\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n>>>>>  }\n> \n> This version is broken on Debian sarge and etch. Feeding a UTF-8 and a latin1\n> encoding of the same character sequence yields to different results.\n\n[...]\n\n> eval { $res = decode_utf8(...); }\n> if ($@) \n>      return decode(...);\n> return $res\n> \n> or\n> \n> eval { $res = decode_utf8(...); }\n> if (defined $res)\n>       return $res;\n> else\n>     return decode(...);\n> \n> show the same (wrong) behaviour on Debian sarge. They do not always\n> decode non UTF-8 characters correctly, eg.\n> #öäü does not work\n> #äöüä does work\n> \n> On Debian etch, both versions are working.\n\nI don't know enough Perl to decide if it is a bug in gitweb usage\nof decode_utf8, if it is a bug in your version of Encode, or if it\nis bug in Encode.\n\nSend copy of this mail to maintainers of Encode perl module.\n-- \nJakub Narebski\nPoland\n"},{"id":"61825","messageId":"47547930.5070603@clearchain.com","threadId":"11118","inReplyTo":"200712031802.55514.jnareb@gmail.com","subject":"Re: Fix UTF Encoding issue","fromName":"Benjamin Close","fromEmail":"benjamin.close@clearchain.com","sentAt":"2007-12-03T21:46:24Z","receivedAt":"2007-12-03T21:46:24Z","isPatch":false,"sender":{"key":"benjamin.close@clearchain.com","avatar":"https://gravatar.com/avatar/19dca0aa372edfa901b93c556dfda2e78ad4434558fe4d139598e086315d714a?d=mp&s=160"},"body":"Jakub Narebski wrote:\n> On Mon, 3 Dec 2007, Martin Koegler wrote:\n>   \n>> On Mon, Dec 03, 2007 at 04:06:48AM -0800, Jakub Narebski wrote:\n>>     \n>>> Ismail Dönmez <ismail@pardus.org.tr> writes:\n>>>       \n>>>> Monday 03 December 2007 Tarihinde 12:14:43 yazm??t?:\n>>>>         \n>>>>> Benjamin Close <Benjamin.Close@clearchain.com> writes:\n>>>>>           \n>>>>>> -\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n>>>>>> -\tif (defined $res) {\n>>>>>> -\t\treturn $res;\n>>>>>> -\t} else {\n>>>>>> -\t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n>>>>>> -\t}\n>>>>>> +\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\n>>>>>> +\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n>>>>>>  }\n>>>>>>             \n>> This version is broken on Debian sarge and etch. Feeding a UTF-8 and a latin1\n>> encoding of the same character sequence yields to different results.\n>>     \n>\n>   \nFor the record, this was on a debian sid machine.\n\n#perl --version\nThis is perl, v5.8.8 built for x86_64-linux-gnu-thread-multi\n\nand the result of not using the original patch was:\n\n<h1>Software error:</h1>\n<pre>Cannot decode string with wide characters at /usr/lib/perl/5.8/Encode.pm line 166.\n</pre>\n\n\nI haven't tried the other solutions tested here.\n>> eval { $res = decode_utf8(...); }\n>> if ($@) \n>>      return decode(...);\n>> return $res\n>>\n>> or\n>>\n>> eval { $res = decode_utf8(...); }\n>> if (defined $res)\n>>       return $res;\n>> else\n>>     return decode(...);\n>>\n>> show the same (wrong) behaviour on Debian sarge. They do not always\n>> decode non UTF-8 characters correctly, eg.\n>> #öäü does not work\n>> #äöüä does work\n>>\n>> On Debian etch, both versions are working.\n>>     \n>\n> I don't know enough Perl to decide if it is a bug in gitweb usage\n> of decode_utf8, if it is a bug in your version of Encode, or if it\n> is bug in Encode.\n>\n> Send copy of this mail to maintainers of Encode perl module.\n>   \nIsmail do you know if sid was also broken?\n"},{"id":"61828","messageId":"200712040020.26773.ismail@pardus.org.tr","threadId":"11118","inReplyTo":"47547930.5070603@clearchain.com","subject":"Re: Fix UTF Encoding issue","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-03T22:20:26Z","receivedAt":"2007-12-03T22:20:26Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Monday 03 December 2007 Tarihinde 23:46:24 yazmıştı:\n> Jakub Narebski wrote:\n> > On Mon, 3 Dec 2007, Martin Koegler wrote:\n> >> On Mon, Dec 03, 2007 at 04:06:48AM -0800, Jakub Narebski wrote:\n> >>> Ismail Dönmez <ismail@pardus.org.tr> writes:\n> >>>> Monday 03 December 2007 Tarihinde 12:14:43 yazm??t?:\n> >>>>> Benjamin Close <Benjamin.Close@clearchain.com> writes:\n> >>>>>> -\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n> >>>>>> -\tif (defined $res) {\n> >>>>>> -\t\treturn $res;\n> >>>>>> -\t} else {\n> >>>>>> -\t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> >>>>>> -\t}\n> >>>>>> +\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\n> >>>>>> +\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> >>>>>>  }\n> >>\n> >> This version is broken on Debian sarge and etch. Feeding a UTF-8 and a\n> >> latin1 encoding of the same character sequence yields to different\n> >> results.\n>\n> For the record, this was on a debian sid machine.\n>\n> #perl --version\n> This is perl, v5.8.8 built for x86_64-linux-gnu-thread-multi\n>\n> and the result of not using the original patch was:\n>\n> <h1>Software error:</h1>\n> <pre>Cannot decode string with wide characters at\n> /usr/lib/perl/5.8/Encode.pm line 166. </pre>\n\nCan you try the attached patch?\n\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n\n\n--- gitweb/gitweb.perl\t2007-11-28 11:33:14.000000000 +0200\n+++ gitweb/gitweb.perl\t2007-11-28 11:33:42.000000000 +0200\n@@ -2159,7 +2159,7 @@\n \t}\n \tmy $owner = $gcos;\n \t$owner =~ s/[,;].*$//;\n-\treturn to_utf8($owner);\n+\treturn $owner;\n }\n \n ## ......................................................................\n"},{"id":"61840","messageId":"20071203230432.GA1337@wolf.clearchain.com","threadId":"11118","inReplyTo":"200712040020.26773.ismail@pardus.org.tr","subject":"Re: Fix UTF Encoding issue","fromName":"Benjamin Close","fromEmail":"benjamin.close@clearchain.com","sentAt":"2007-12-03T23:04:34Z","receivedAt":"2007-12-03T23:04:34Z","isPatch":false,"sender":{"key":"benjamin.close@clearchain.com","avatar":"https://gravatar.com/avatar/19dca0aa372edfa901b93c556dfda2e78ad4434558fe4d139598e086315d714a?d=mp&s=160"},"body":"On Tue, Dec 04, 2007 at 12:20:26AM +0200, Ismail D??nmez wrote:\n> Monday 03 December 2007 Tarihinde 23:46:24 yazm????t??:\n> > Jakub Narebski wrote:\n> > > On Mon, 3 Dec 2007, Martin Koegler wrote:\n> > >> On Mon, Dec 03, 2007 at 04:06:48AM -0800, Jakub Narebski wrote:\n> > >>> Ismail D??nmez <ismail@pardus.org.tr> writes:\n> > >>>> Monday 03 December 2007 Tarihinde 12:14:43 yazm??t?:\n> > >>>>> Benjamin Close <Benjamin.Close@clearchain.com> writes:\n> > >>>>>> -\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n> > >>>>>> -\tif (defined $res) {\n> > >>>>>> -\t\treturn $res;\n> > >>>>>> -\t} else {\n> > >>>>>> -\t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> > >>>>>> -\t}\n> > >>>>>> +\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\n> > >>>>>> +\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> > >>>>>>  }\n> > >>\n> > >> This version is broken on Debian sarge and etch. Feeding a UTF-8 and a\n> > >> latin1 encoding of the same character sequence yields to different\n> > >> results.\n> >\n> > For the record, this was on a debian sid machine.\n> >\n> > #perl --version\n> > This is perl, v5.8.8 built for x86_64-linux-gnu-thread-multi\n> >\n> > and the result of not using the original patch was:\n> >\n> > <h1>Software error:</h1>\n> > <pre>Cannot decode string with wide characters at\n> > /usr/lib/perl/5.8/Encode.pm line 166. </pre>\n> \n> Can you try the attached patch?\n\nI confirm that the patch corrects the problem.\n\nWithout it I get the Cannot decode string error. With it gitweb displays\ncorrectly.\n\nCheers,\n\tBenjamin\n"},{"id":"61841","messageId":"200712040037.37204.jnareb@gmail.com","threadId":"11118","inReplyTo":"20071203230432.GA1337@wolf.clearchain.com","subject":"Re: Fix UTF Encoding issue","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-12-03T23:37:35Z","receivedAt":"2007-12-03T23:37:35Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, 4 Dec 2007, Benjamin Close wrote:\n> On Tue, Dec 04, 2007 at 12:20:26AM +0200, Ismail Donmez wrote:\n> > \n> > Can you try the attached patch?\n> \n> I confirm that the patch corrects the problem.\n> \n> Without it I get the Cannot decode string error. With it gitweb\n> displays correctly.\n\nBut the patch _avoids_ issue (des not convert owner to utf8), rather\nthan solving it, if I understand it correctly. What if gecos is in \nutf-8?\n\n-- \nJakub Narebski\nPoland\n"},{"id":"61864","messageId":"200712040612.32483.ismail@pardus.org.tr","threadId":"11118","inReplyTo":"200712040037.37204.jnareb@gmail.com","subject":"Re: Fix UTF Encoding issue","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-04T04:12:32Z","receivedAt":"2007-12-04T04:12:32Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Tuesday 04 December 2007 Tarihinde 01:37:35 yazmıştı:\n> On Tue, 4 Dec 2007, Benjamin Close wrote:\n> > On Tue, Dec 04, 2007 at 12:20:26AM +0200, Ismail Donmez wrote:\n> > > Can you try the attached patch?\n> >\n> > I confirm that the patch corrects the problem.\n> >\n> > Without it I get the Cannot decode string error. With it gitweb\n> > displays correctly.\n>\n> But the patch _avoids_ issue (des not convert owner to utf8), rather\n> than solving it, if I understand it correctly. What if gecos is in\n> utf-8?\n\nIndeed its a workaround but UTF-8 username is correctly displayed in gitweb so \nmy understanding was gecos field is already UTF-8.\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n"},{"id":"61875","messageId":"20071204075028.GB31042@auto.tuwien.ac.at","threadId":"11118","inReplyTo":"200712031802.55514.jnareb@gmail.com","subject":"Re: Fix UTF Encoding issue","fromName":"Martin Koegler","fromEmail":"mkoegler@auto.tuwien.ac.at","sentAt":"2007-12-04T07:50:28Z","receivedAt":"2007-12-04T07:50:28Z","isPatch":false,"sender":{"key":"mkoegler@auto.tuwien.ac.at","avatar":null},"body":"On Mon, Dec 03, 2007 at 06:02:54PM +0100, Jakub Narebski wrote:\n> On Mon, 3 Dec 2007, Martin Koegler wrote:\n> > eval { $res = decode_utf8(...); }\n> > if ($@) \n> >      return decode(...);\n> > return $res\n> > \n> > or\n> > \n> > eval { $res = decode_utf8(...); }\n> > if (defined $res)\n> >       return $res;\n> > else\n> >     return decode(...);\n> > \n> > show the same (wrong) behaviour on Debian sarge. They do not always\n> > decode non UTF-8 characters correctly, eg.\n> > #öäü does not work\n> > #äöüä does work\n> > \n> > On Debian etch, both versions are working.\n> \n> I don't know enough Perl to decide if it is a bug in gitweb usage\n> of decode_utf8, if it is a bug in your version of Encode, or if it\n> is bug in Encode.\n> \n> Send copy of this mail to maintainers of Encode perl module.\n\nThe bug affects old versions of perl (Debian sarge = oldstable).\nAs it works on the newer Debian etch, do you really think, that it is\na good idea to report issue? \n\nHow would you handle a bug report, which reports a bug for gitweb in GIT\n1.4, and tells you, that a newer versions works?\n\nAs Debian sarge has reached its end of life, the distribution will\nprobable also issue no update.\n\nmfg Martin Kögler\n"},{"id":"61876","messageId":"200712040955.04655.ismail@pardus.org.tr","threadId":"11118","inReplyTo":"20071204075028.GB31042@auto.tuwien.ac.at","subject":"Re: Fix UTF Encoding issue","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-04T07:55:04Z","receivedAt":"2007-12-04T07:55:04Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Tuesday 04 December 2007 Tarihinde 09:50:28 yazmıştı:\n> The bug affects old versions of perl (Debian sarge = oldstable).\n> As it works on the newer Debian etch, do you really think, that it is\n> a good idea to report issue?\n\nSame problem here with v5.8.8 which is latest stable perl5 release.\n\nRegards,\nismail\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n"},{"id":"61877","messageId":"20071204080407.GC31042@auto.tuwien.ac.at","threadId":"11118","inReplyTo":"47547930.5070603@clearchain.com","subject":"Re: Fix UTF Encoding issue","fromName":"Martin Koegler","fromEmail":"mkoegler@auto.tuwien.ac.at","sentAt":"2007-12-04T08:04:07Z","receivedAt":"2007-12-04T08:04:07Z","isPatch":false,"sender":{"key":"mkoegler@auto.tuwien.ac.at","avatar":null},"body":"On Tue, Dec 04, 2007 at 08:16:24AM +1030, Benjamin Close wrote:\n> Jakub Narebski wrote:\n> >On Mon, 3 Dec 2007, Martin Koegler wrote:\n> >>On Mon, Dec 03, 2007 at 04:06:48AM -0800, Jakub Narebski wrote:\n> >>>Ismail Dönmez <ismail@pardus.org.tr> writes:\n> >>>>Monday 03 December 2007 Tarihinde 12:14:43 yazm??t?:\n> >>>>>Benjamin Close <Benjamin.Close@clearchain.com> writes:\n> >>>>>>-\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n> >>>>>>-\tif (defined $res) {\n> >>>>>>-\t\treturn $res;\n> >>>>>>-\t} else {\n> >>>>>>-\t\treturn decode($fallback_encoding, $str, \n> >>>>>>Encode::FB_DEFAULT);\n> >>>>>>-\t}\n> >>>>>>+\teval { return ($res = decode_utf8($str, Encode::FB_CROAK)); \n> >>>>>>};\n> >>>>>>+\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> >>>>>> }\n> >>>>>>            \n> >>This version is broken on Debian sarge and etch. Feeding a UTF-8 and a \n> >>latin1\n> >>encoding of the same character sequence yields to different results.\n>\n> For the record, this was on a debian sid machine.\n> \n> #perl --version\n> This is perl, v5.8.8 built for x86_64-linux-gnu-thread-multi\n> \n> and the result of not using the original patch was:\n> \n> <h1>Software error:</h1>\n> <pre>Cannot decode string with wide characters at \n> /usr/lib/perl/5.8/Encode.pm line 166.\n> </pre>\n> \n> \n> I haven't tried the other solutions tested here.\n\nDebian etch also has v5.8.8.\n\nMy main question is, why is the error not catched?\n\nI'm not a perl programmer, but in your patch the first line is a\nNOP. The return in eval seems to only returns from the eval block, so\nany text is decoded as latin1 with the second statement.\n\nIn the original version, decode($fallback_encoding, $str,\nEncode::FB_DEFAULT) can not emit an error, else it would in your\nversion too. \n\nIn your version, eval is able to surpress the error of\ndecode_utf8($str, Encode::FB_CROAK);, but not in the original version.\n\nStrange.\n\nmfg Martin Kögler\n"},{"id":"61878","messageId":"200712041012.50935.ismail@pardus.org.tr","threadId":"11118","inReplyTo":"20071204080407.GC31042@auto.tuwien.ac.at","subject":"Re: Fix UTF Encoding issue","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-04T08:12:50Z","receivedAt":"2007-12-04T08:12:50Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Tuesday 04 December 2007 10:04:07 Martin Koegler yazmıştı:\n> On Tue, Dec 04, 2007 at 08:16:24AM +1030, Benjamin Close wrote:\n> > Jakub Narebski wrote:\n> > >On Mon, 3 Dec 2007, Martin Koegler wrote:\n> > >>On Mon, Dec 03, 2007 at 04:06:48AM -0800, Jakub Narebski wrote:\n> > >>>Ismail Dönmez <ismail@pardus.org.tr> writes:\n> > >>>>Monday 03 December 2007 Tarihinde 12:14:43 yazm??t?:\n> > >>>>>Benjamin Close <Benjamin.Close@clearchain.com> writes:\n> > >>>>>>-\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n> > >>>>>>-\tif (defined $res) {\n> > >>>>>>-\t\treturn $res;\n> > >>>>>>-\t} else {\n> > >>>>>>-\t\treturn decode($fallback_encoding, $str,\n> > >>>>>>Encode::FB_DEFAULT);\n> > >>>>>>-\t}\n> > >>>>>>+\teval { return ($res = decode_utf8($str, Encode::FB_CROAK));\n> > >>>>>>};\n> > >>>>>>+\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n> > >>>>>> }\n> > >>\n> > >>This version is broken on Debian sarge and etch. Feeding a UTF-8 and a\n> > >>latin1\n> > >>encoding of the same character sequence yields to different results.\n> >\n> > For the record, this was on a debian sid machine.\n> >\n> > #perl --version\n> > This is perl, v5.8.8 built for x86_64-linux-gnu-thread-multi\n> >\n> > and the result of not using the original patch was:\n> >\n> > <h1>Software error:</h1>\n> > <pre>Cannot decode string with wide characters at\n> > /usr/lib/perl/5.8/Encode.pm line 166.\n> > </pre>\n> >\n> >\n> > I haven't tried the other solutions tested here.\n>\n> Debian etch also has v5.8.8.\n>\n> My main question is, why is the error not catched?\n>\n> I'm not a perl programmer, but in your patch the first line is a\n> NOP. The return in eval seems to only returns from the eval block, so\n> any text is decoded as latin1 with the second statement.\n>\n> In the original version, decode($fallback_encoding, $str,\n> Encode::FB_DEFAULT) can not emit an error, else it would in your\n> version too.\n>\n> In your version, eval is able to surpress the error of\n> decode_utf8($str, Encode::FB_CROAK);, but not in the original version.\n\nI think just a better method is to use (not tested):\n\nif( is_utf8($str) ) \n{\n\treturn decode_utf8($str);\n}\nelse {\n\treturn decode($str);\n}\n\nRegards,\nismail\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n"},{"id":"61879","messageId":"20071204081634.GD31042@auto.tuwien.ac.at","threadId":"11118","inReplyTo":"200712040955.04655.ismail@pardus.org.tr","subject":"Re: Fix UTF Encoding issue","fromName":"Martin Koegler","fromEmail":"mkoegler@auto.tuwien.ac.at","sentAt":"2007-12-04T08:16:34Z","receivedAt":"2007-12-04T08:16:34Z","isPatch":false,"sender":{"key":"mkoegler@auto.tuwien.ac.at","avatar":null},"body":"On Tue, Dec 04, 2007 at 09:55:04AM +0200, Ismail Dönmez wrote:\n> Tuesday 04 December 2007 Tarihinde 09:50:28 yazm????t??:\n> > The bug affects old versions of perl (Debian sarge = oldstable).\n> > As it works on the newer Debian etch, do you really think, that it is\n> > a good idea to report issue?\n> \n> Same problem here with v5.8.8 which is latest stable perl5 release.\n\nI have put together a small perl script, which tests the various ways\nof decoding, which have been posted on the list. The first test is\nwrong by design. A working decoding method should result in\n\"#öäü#äöü\".\n\nDebian sarge:\n#öäü#Ã€Ã¶ÃŒ\n##äöü\n##äöü\n##äöü\n\nDebian etch, OpenSuSE 10.2, Fedora 7:\n#öäü#Ã€Ã¶ÃŒ\n#öäü#äöü\n#öäü#äöü\n#öäü#äöü\n\nmfg Martin Kögler\n\n#!/usr/bin/perl\nuse Encode;\n\nsub t {\nmy $str = shift;\nmy ($res);\neval { return ($res = decode_utf8($str, Encode::FB_CROAK)); };\nreturn decode(\"latin1\", $str, Encode::FB_DEFAULT);\n}\nsub t1 {\nmy $str = shift;\nmy ($res);\neval { ($res = decode_utf8($str, Encode::FB_CROAK)); };\nif ($@) {\nreturn decode(\"latin1\", $str, Encode::FB_DEFAULT); }\nelse\n{ return $res; }\n}\n\nsub t2 {\nmy $str = shift;\nmy ($res);\n\neval { $res = decode_utf8($str, Encode::FB_CROAK); };\n if (defined $res) {\n        return $res;\n} else {\n        return decode(\"latin1\", $str, Encode::FB_DEFAULT);\n}\n}\n\nsub t3 {\n\tmy $str = shift;\n\tmy $res;\n\teval { $res = decode_utf8 ($str, 1); };\n\treturn $res || decode('latin1', $str);\n}\n\nprint t(\"#öäü\");\nprint t(\"#Ã€Ã¶ÃŒ\");\nprint \"\\n\";\nprint t1(\"#öäü\");\nprint t1(\"#Ã€Ã¶ÃŒ\");\nprint \"\\n\";\nprint t2(\"#öäü\");\nprint t2(\"#Ã€Ã¶ÃŒ\");\nprint \"\\n\";\nprint t3(\"#öäü\");\nprint t3(\"#Ã€Ã¶ÃŒ\");\nprint \"\\n\";\n"},{"id":"61880","messageId":"20071204082021.GE31042@auto.tuwien.ac.at","threadId":"11118","inReplyTo":"200712041012.50935.ismail@pardus.org.tr","subject":"Re: Fix UTF Encoding issue","fromName":"Martin Koegler","fromEmail":"mkoegler@auto.tuwien.ac.at","sentAt":"2007-12-04T08:20:21Z","receivedAt":"2007-12-04T08:20:21Z","isPatch":false,"sender":{"key":"mkoegler@auto.tuwien.ac.at","avatar":null},"body":"On Tue, Dec 04, 2007 at 10:12:50AM +0200, Ismail Dönmez wrote:\n> I think just a better method is to use (not tested):\n> \n> if( is_utf8($str) ) \n> {\n> \treturn decode_utf8($str);\n> }\n> else {\n> \treturn decode($str);\n> }\n\nI already tried this function. It does not test, if a string is\nreally UTF-8. It seems to be to intended to check, if perl stores\nthe string internally in a multi byte encoding.\n\nmfg Martin Kögler.\n"},{"id":"61881","messageId":"200712041028.59185.ismail@pardus.org.tr","threadId":"11118","inReplyTo":"20071204081634.GD31042@auto.tuwien.ac.at","subject":"Re: Fix UTF Encoding issue","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-04T08:28:59Z","receivedAt":"2007-12-04T08:28:59Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Tuesday 04 December 2007 10:16:34 Martin Koegler yazmıştı:\n[...]\n> print t(\"#öäü\");\n> print t(\"#Ã€Ã¶ÃŒ\");\n> print \"\\n\";\n\nHow about this one, doesn't even use Encode, uses just built-in utf8 \nfunction :\n\n[~]> cat test.pl\nbinmode STDOUT, ':utf8';\n\nmy $str = \"#öäü\";\n\nif (utf8::valid($str))\n{\n    utf8::decode($str);\n}\n\nprint $str.\"\\n\";\n\n[~]> perl test.pl\n#öäü\n\nRegards,\nismail\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n"},{"id":"61883","messageId":"200712041033.39579.ismail@pardus.org.tr","threadId":"11118","inReplyTo":"200712041028.59185.ismail@pardus.org.tr","subject":"Re: Fix UTF Encoding issue","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-04T08:33:39Z","receivedAt":"2007-12-04T08:33:39Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Tuesday 04 December 2007 10:28:59 Ismail Dönmez yazmıştı:\n> Tuesday 04 December 2007 10:16:34 Martin Koegler yazmıştı:\n> [...]\n>\n> > print t(\"#öäü\");\n> > print t(\"#Ã€Ã¶ÃŒ\");\n> > print \"\\n\";\n>\n> How about this one, doesn't even use Encode, uses just built-in utf8\n> function :\n>\n> [~]> cat test.pl\n> binmode STDOUT, ':utf8';\n>\n> my $str = \"#öäü\";\n>\n> if (utf8::valid($str))\n> {\n>     utf8::decode($str);\n> }\n>\n> print $str.\"\\n\";\n>\n> [~]> perl test.pl\n> #öäü\n\nFollowing to_utf8 function works for me :\n\nsub to_utf8 {\n·   my $str = shift;\n\n    if(utf8::valid($str))\n    {\n        utf8::decode($str);\n    }\n·\n    return $str;\n}\n\nRegards,\nismail\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n"},{"id":"61887","messageId":"20071204084412.GA19597@auto.tuwien.ac.at","threadId":"11118","inReplyTo":"200712041033.39579.ismail@pardus.org.tr","subject":"Re: Fix UTF Encoding issue","fromName":"Martin Koegler","fromEmail":"mkoegler@auto.tuwien.ac.at","sentAt":"2007-12-04T08:44:12Z","receivedAt":"2007-12-04T08:44:12Z","isPatch":false,"sender":{"key":"mkoegler@auto.tuwien.ac.at","avatar":null},"body":"On Tue, Dec 04, 2007 at 10:33:39AM +0200, Ismail Dönmez wrote:\n> Following to_utf8 function works for me :\n\nFor me too (Debian sarge+etch).\n\n> sub to_utf8 {\n> ·   my $str = shift;\n> \n>     if(utf8::valid($str))\n>     {\n>         utf8::decode($str);\n>     }\n> ·\n>     return $str;\n\nIn the original thread, there was some discussion, that some people\nmight want a different fallback endcoding. So mayme you should \nkeep the second call to decode for the fallback encoding.\n\n> }\n\nmfg Martin Kögler\n"},{"id":"61888","messageId":"200712041047.39340.ismail@pardus.org.tr","threadId":"11118","inReplyTo":"20071204084412.GA19597@auto.tuwien.ac.at","subject":"Re: Fix UTF Encoding issue","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-04T08:47:39Z","receivedAt":"2007-12-04T08:47:39Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Tuesday 04 December 2007 10:44:12 Martin Koegler yazmıştı:\n> On Tue, Dec 04, 2007 at 10:33:39AM +0200, Ismail Dönmez wrote:\n> > Following to_utf8 function works for me :\n>\n> For me too (Debian sarge+etch).\n\nThanks for testing.\n\n> > sub to_utf8 {\n> > ·   my $str = shift;\n> >\n> >     if(utf8::valid($str))\n> >     {\n> >         utf8::decode($str);\n> >     }\n> > ·\n> >     return $str;\n>\n> In the original thread, there was some discussion, that some people\n> might want a different fallback endcoding. So mayme you should\n> keep the second call to decode for the fallback encoding.\n\nProbably, I just wanted to fix this damn UTF-8 bug surfacing over and over =)\n\nRegards,\nismail\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n"},{"id":"61890","messageId":"200712041055.41593.ismail@pardus.org.tr","threadId":"11118","inReplyTo":"200712041047.39340.ismail@pardus.org.tr","subject":"Re: Fix UTF Encoding issue","fromName":"Ismail Dönmez","fromEmail":"ismail@pardus.org.tr","sentAt":"2007-12-04T08:55:41Z","receivedAt":"2007-12-04T08:55:41Z","isPatch":false,"sender":{"key":"ismail@pardus.org.tr","avatar":null},"body":"Tuesday 04 December 2007 10:47:39 Ismail Dönmez yazmıştı:\n> Tuesday 04 December 2007 10:44:12 Martin Koegler yazmıştı:\n> > On Tue, Dec 04, 2007 at 10:33:39AM +0200, Ismail Dönmez wrote:\n> > > Following to_utf8 function works for me :\n> >\n> > For me too (Debian sarge+etch).\n>\n> Thanks for testing.\n\nUse Perl built-in utf8 function for UTF-8 decoding.\n\nSigned-off-by: İsmail Dönmez <ismail@pardus.org.tr>\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex ff5daa7..db255c1 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -695,10 +695,9 @@ sub validate_refname {\n # in utf-8 thanks to \"binmode STDOUT, ':utf8'\" at beginning\n sub to_utf8 {\n \tmy $str = shift;\n-\tmy $res;\n-\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n-\tif (defined $res) {\n-\t\treturn $res;\n+        if (utf8::valid($str)) {\n+                utf8::decode($str);\n+                return $str;\n \t} else {\n \t\treturn decode($fallback_encoding, $str, Encode::FB_DEFAULT);\n \t}\n\n\n\n-- \nNever learn by your mistakes, if you do you may never dare to try again.\n"},{"id":"61891","messageId":"200712041007.44525.jnareb@gmail.com","threadId":"11118","inReplyTo":"200712041055.41593.ismail@pardus.org.tr","subject":"Re: Fix UTF Encoding issue","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-12-04T09:07:43Z","receivedAt":"2007-12-04T09:07:43Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, 4 Dec 2007, Ismail Dönmez wrote:\n\n> Use Perl built-in utf8 function for UTF-8 decoding.\n> \n> Signed-off-by: İsmail Dönmez <ismail@pardus.org.tr>\n \nLooks nice. I have not tested it, but if it works: Ack.\n-- \nJakub Narebski\nPoland\n"},{"id":"61896","messageId":"8BED0A14-31F0-4288-B7A2-59B2CEE6DF97@wincent.com","threadId":"11118","inReplyTo":"200712041055.41593.ismail@pardus.org.tr","subject":"Re: Fix UTF Encoding issue","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2007-12-04T10:11:04Z","receivedAt":"2007-12-04T10:11:04Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 4/12/2007, a las 9:55, Ismail Dönmez escribió:\n\n> Tuesday 04 December 2007 10:47:39 Ismail Dönmez yazmıştı:\n>> Tuesday 04 December 2007 10:44:12 Martin Koegler yazmıştı:\n>>> On Tue, Dec 04, 2007 at 10:33:39AM +0200, Ismail Dönmez wrote:\n>>>> Following to_utf8 function works for me :\n>>>\n>>> For me too (Debian sarge+etch).\n>>\n>> Thanks for testing.\n>\n> Use Perl built-in utf8 function for UTF-8 decoding.\n>\n> Signed-off-by: İsmail Dönmez <ismail@pardus.org.tr>\n>\n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index ff5daa7..db255c1 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -695,10 +695,9 @@ sub validate_refname {\n> # in utf-8 thanks to \"binmode STDOUT, ':utf8'\" at beginning\n> sub to_utf8 {\n> \tmy $str = shift;\n> -\tmy $res;\n> -\teval { $res = decode_utf8($str, Encode::FB_CROAK); };\n> -\tif (defined $res) {\n> -\t\treturn $res;\n> +        if (utf8::valid($str)) {\n> +                utf8::decode($str);\n> +                return $str;\n\nThis is good as it fixes another problem which some may have  \nencountered. On at least one distro that I use (Red Hat Enterprise  \nLinux 3) the Encode module is very old (it's 1.83; the latest release  \nis 2.23), and so gitweb won't even run, dying during compilation with  \nthis:\n\n\tToo many arguments for Encode::decode_utf8 at gitweb.cgi line 686,  \nnear \"Encode::FB_CROAK)\"\n\nOf course, the workaround is to install a newer version of the module,  \nbut this patch eliminates that dependency which is IMO a good thing.\n\nCheers,\nWincent\n"}]}