{"thread":{"id":"29509","subject":"[RFC PATCH] gitweb: use CGI with -utf8","startedAt":"2012-02-01T22:50:53Z","lastAt":"2012-02-03T21:09:12Z","messageCount":14,"participants":["Michał Kiedrowicz","Jakub Narebski","Michal Kiedrowicz","Junio C Hamano"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"183511","messageId":"1328136653-20559-1-git-send-email-michal.kiedrowicz@gmail.com","threadId":"29509","inReplyTo":null,"subject":"[RFC PATCH] gitweb: use CGI with -utf8","fromName":"Michał Kiedrowicz","fromEmail":"michal.kiedrowicz@gmail.com","sentAt":"2012-02-01T22:50:53Z","receivedAt":"2012-02-01T22:50:53Z","isPatch":true,"sender":{"key":"michal.kiedrowicz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/14072847?v=4"},"body":"I noticed that gitweb tries a lot to properly process UTF-8 data, for\nexample it prints my name correctly in log and commit information, but\nit echos junk in the search field. It looks like:\n\n\tMichaÅ Kiedrowicz\n\nI don't know CGI well and I never touched gitewb code, but I found this\non http://www.lemoda.net/cgi/perl-unicode/index.html:\n\n\tuse CGI '-utf8';\n\tmy $value = params ('input');\n\nI tried it and that fixed my problem. I'm not sure about the\nconsequences, maybe someone more experienced in CGI might help?\n---\n gitweb/gitweb.perl |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex abb5a79..74d45b1 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -10,7 +10,7 @@\n use 5.008;\n use strict;\n use warnings;\n-use CGI qw(:standard :escapeHTML -nosticky);\n+use CGI qw(:standard :escapeHTML -nosticky -utf8);\n use CGI::Util qw(unescape);\n use CGI::Carp qw(fatalsToBrowser set_message);\n use Encode;\n-- \n1.7.3.4\n"},{"id":"183645","messageId":"m37h05c8c1.fsf@localhost.localdomain","threadId":"29509","inReplyTo":"1328136653-20559-1-git-send-email-michal.kiedrowicz@gmail.com","subject":"Re: [RFC PATCH] gitweb: use CGI with -utf8","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2012-02-02T20:01:41Z","receivedAt":"2012-02-02T20:01:41Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Michał Kiedrowicz <michal.kiedrowicz@gmail.com> writes:\n\n> I noticed that gitweb tries a lot to properly process UTF-8 data, for\n> example it prints my name correctly in log and commit information, but\n> it echos junk in the search field. It looks like:\n> \n> \tMichaÅ Kiedrowicz\n> \n> I don't know CGI well and I never touched gitewb code, but I found this\n> on http://www.lemoda.net/cgi/perl-unicode/index.html:\n> \n> \tuse CGI '-utf8';\n> \tmy $value = params ('input');\n> \n> I tried it and that fixed my problem. I'm not sure about the\n> consequences, maybe someone more experienced in CGI might help?\n\nI have reworded this to form a proper commit message (see\nDocumentation/SubmittingPatches) and I'll resend this as a reply to\nthis email.\n\n> ---\n>  gitweb/gitweb.perl |    2 +-\n>  1 files changed, 1 insertions(+), 1 deletions(-)\n> \n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index abb5a79..74d45b1 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -10,7 +10,7 @@\n>  use 5.008;\n>  use strict;\n>  use warnings;\n> -use CGI qw(:standard :escapeHTML -nosticky);\n> +use CGI qw(:standard :escapeHTML -nosticky -utf8);\n>  use CGI::Util qw(unescape);\n>  use CGI::Carp qw(fatalsToBrowser set_message);\n>  use Encode;\n> -- \n\nDoes this actually work for you?  Because it doesn't work for me\n(perhaps I have too old CGI module: what CGI.pm and what Perl version\ndo you use?).\n\nSee other solution to this in other reply to this email.\n\n-- \nJakub Narebski\n"},{"id":"183649","messageId":"201202022108.51353.jnareb@gmail.com","threadId":"29509","inReplyTo":"m37h05c8c1.fsf@localhost.localdomain","subject":"[PATCH/RFC (version A)] gitweb: use CGI with -utf8 to process Unicode query parameters correctly","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2012-02-02T20:08:50Z","receivedAt":"2012-02-02T20:08:50Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Gitweb tries hard to properly process UTF-8 data, by marking output\nfrom git commands and contents of files as UTF-8 with to_utf8()\nsubroutine.  This ensures that gitweb would print correctly UTF-8\ne.g. in 'log' and 'commit' views.\n\nUnfortunately it misses another source of potentially Unicode input,\nnamely query parameters.  The result is that one cannot search for a\nstring containing characters outside US-ASCII.  For example searching\nfor \"Michał Kiedrowicz\" (containing letter 'ł' - LATIN SMALL LETTER L\nWITH STROKE, with Unicode codepoint U+0142, represented with 0xc5 0x82\nbytes in UTF-8 and percent-encoded as %C5%81) result in the following\nincorrect data in search field\n\n\tMichaÅ Kiedrowicz\n\nThis is caused by CGI by default treating '0xc5 0x82' bytes as two\ncharacters in Perl legacy encoding latin-1 (iso-8859-1), because 's'\nquery parameter is not processed explicitly as UTF-8 encoded string.\n\nAccording to \"Using Unicode in a Perl CGI script\" article on\nhttp://www.lemoda.net/cgi/perl-unicode/index.html the simplest\nsolution is to just import '-utf8' pragma for CGI module:\n\n\tuse CGI '-utf8';\n\tmy $value = params('input');\n\nAccording to CGI module documentation, the '-utf8' pragma may cause\nproblems with POST requests containing binary files... but gitweb\ncurrently do not use POST requests at all, so this should be not a\nproblem now.\n\nAlternate solution would be to explicity decode query parameters when\nstoring them in %input_params (and perhaps also path_info).\n\n[jn: reworded / rewritten commit message]\n\nSigned-off-by: Michał Kiedrowicz <michal.kiedrowicz@gmail.com>\nSigned-off-by: Jakub Narębski <jnareb@gmail.com>\n---\n gitweb/gitweb.perl |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex 9cf7e71..a7441ef 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -10,7 +10,7 @@\n use 5.008;\n use strict;\n use warnings;\n-use CGI qw(:standard :escapeHTML -nosticky);\n+use CGI qw(:standard :escapeHTML -nosticky -utf8);\n use CGI::Util qw(unescape);\n use CGI::Carp qw(fatalsToBrowser set_message);\n use Encode;\n-- \n1.7.6\n"},{"id":"183651","messageId":"201202022110.07127.jnareb@gmail.com","threadId":"29509","inReplyTo":"m37h05c8c1.fsf@localhost.localdomain","subject":"[PATCH/RFC (version B)] gitweb: Allow UTF-8 encoded CGI query parameters and path_info","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2012-02-02T20:10:06Z","receivedAt":"2012-02-02T20:10:06Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Gitweb tries hard to properly process UTF-8 data, by marking output\nfrom git commands and contents of files as UTF-8 with to_utf8()\nsubroutine.  This ensures that gitweb would print correctly UTF-8\ne.g. in 'log' and 'commit' views.\n\nUnfortunately it misses another source of potentially Unicode input,\nnamely query parameters.  The result is that one cannot search for a\nstring containing characters outside US-ASCII.  For example searching\nfor \"Michał Kiedrowicz\" (containing letter 'ł' - LATIN SMALL LETTER L\nWITH STROKE, with Unicode codepoint U+0142, represented with 0xc5 0x82\nbytes in UTF-8 and percent-encoded as %C5%81) result in the following\nincorrect data in search field\n\n\tMichaÅ Kiedrowicz\n\nThis is caused by CGI by default treating '0xc5 0x82' bytes as two\ncharacters in Perl legacy encoding latin-1 (iso-8859-1), because 's'\nquery parameter is not processed explicitly as UTF-8 encoded string.\n\nThe solution used here follows \"Using Unicode in a Perl CGI script\"\narticle on http://www.lemoda.net/cgi/perl-unicode/index.html:\n\n\tuse CGI;\n\tuse Encode 'decode_utf8;\n\tmy $value = params('input');\n\t$value = decode_utf8($value);\n\nThis is done when filling %input_params hash; this required to move\nfrom explicit $cgi->param(<label>) to $input_params{<name>} in a few\nplaces.\n\nAlternate solution would be to simply use the '-utf8' pragma (via\n\"use CGI '-utf8';\"), but according to CGI.pm documentation it may\ncause problems with POST requests containing binary files... and\nit doesn't work with old CGI.pm version 3.10 from Perl v5.8.6.\n\nNoticed-by: Michał Kiedrowicz <michal.kiedrowicz@gmail.com>\nSigned-off-by: Jakub Narębski <jnareb@gmail.com>\n---\n gitweb/gitweb.perl |   12 ++++++------\n 1 files changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex 9cf7e71..55b2c24 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -52,7 +52,7 @@ sub evaluate_uri {\n \t# as base URL.\n \t# Therefore, if we needed to strip PATH_INFO, then we know that we have\n \t# to build the base URL ourselves:\n-\tour $path_info = $ENV{\"PATH_INFO\"};\n+\tour $path_info = decode_utf8($ENV{\"PATH_INFO\"});\n \tif ($path_info) {\n \t\tif ($my_url =~ s,\\Q$path_info\\E$,, &&\n \t\t    $my_uri =~ s,\\Q$path_info\\E$,, &&\n@@ -816,9 +816,9 @@ sub evaluate_query_params {\n \n \twhile (my ($name, $symbol) = each %cgi_param_mapping) {\n \t\tif ($symbol eq 'opt') {\n-\t\t\t$input_params{$name} = [ $cgi->param($symbol) ];\n+\t\t\t$input_params{$name} = [ map { decode_utf8($_) } $cgi->param($symbol) ];\n \t\t} else {\n-\t\t\t$input_params{$name} = $cgi->param($symbol);\n+\t\t\t$input_params{$name} = decode_utf8($cgi->param($symbol));\n \t\t}\n \t}\n }\n@@ -2767,7 +2767,7 @@ sub git_populate_project_tagcloud {\n \t}\n \n \tmy $cloud;\n-\tmy $matched = $cgi->param('by_tag');\n+\tmy $matched = $input_params{'ctag'};\n \tif (eval { require HTML::TagCloud; 1; }) {\n \t\t$cloud = HTML::TagCloud->new;\n \t\tforeach my $ctag (sort keys %ctags_lc) {\n@@ -5282,7 +5282,7 @@ sub git_project_list_body {\n \n \tmy $check_forks = gitweb_check_feature('forks');\n \tmy $show_ctags  = gitweb_check_feature('ctags');\n-\tmy $tagfilter = $show_ctags ? $cgi->param('by_tag') : undef;\n+\tmy $tagfilter = $show_ctags ? $input_params{'ctag'} : undef;\n \t$check_forks = undef\n \t\tif ($tagfilter || $searchtext);\n \n@@ -6197,7 +6197,7 @@ sub git_tag {\n \n sub git_blame_common {\n \tmy $format = shift || 'porcelain';\n-\tif ($format eq 'porcelain' && $cgi->param('js')) {\n+\tif ($format eq 'porcelain' && $input_params{'javascript'}) {\n \t\t$format = 'incremental';\n \t\t$action = 'blame_incremental'; # for page title etc\n \t}\n-- \n1.7.6\n"},{"id":"183652","messageId":"201202022111.03336.jnareb@gmail.com","threadId":"29509","inReplyTo":"201202022108.51353.jnareb@gmail.com","subject":"Re: [PATCH/RFC (version A)] gitweb: use CGI with -utf8 to process Unicode query parameters correctly","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2012-02-02T20:11:02Z","receivedAt":"2012-02-02T20:11:02Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"> According to \"Using Unicode in a Perl CGI script\" article on\n> http://www.lemoda.net/cgi/perl-unicode/index.html the simplest\n> solution is to just import '-utf8' pragma for CGI module:\n> \n> \tuse CGI '-utf8';\n> \tmy $value = params('input');\n[...]\n> ---\n\nExcept it doesn't work for me...\n\n>  gitweb/gitweb.perl |    2 +-\n>  1 files changed, 1 insertions(+), 1 deletions(-)\n> \n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index 9cf7e71..a7441ef 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -10,7 +10,7 @@\n>  use 5.008;\n>  use strict;\n>  use warnings;\n> -use CGI qw(:standard :escapeHTML -nosticky);\n> +use CGI qw(:standard :escapeHTML -nosticky -utf8);\n>  use CGI::Util qw(unescape);\n>  use CGI::Carp qw(fatalsToBrowser set_message);\n>  use Encode;\n> -- \n> 1.7.6\n> \n> \n\n-- \nJakub Narebski\nPoland\n"},{"id":"183659","messageId":"20120202213816.1eabe031@gmail.com","threadId":"29509","inReplyTo":"m37h05c8c1.fsf@localhost.localdomain","subject":"Re: [RFC PATCH] gitweb: use CGI with -utf8","fromName":"Michał Kiedrowicz","fromEmail":"michal.kiedrowicz@gmail.com","sentAt":"2012-02-02T20:38:16Z","receivedAt":"2012-02-02T20:38:16Z","isPatch":true,"sender":{"key":"michal.kiedrowicz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/14072847?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> wrote:\n\n> Michał Kiedrowicz <michal.kiedrowicz@gmail.com> writes:\n> \n> > I noticed that gitweb tries a lot to properly process UTF-8 data, for\n> > example it prints my name correctly in log and commit information, but\n> > it echos junk in the search field. It looks like:\n> > \n> > \tMichaÅ Kiedrowicz\n> > \n> > I don't know CGI well and I never touched gitewb code, but I found this\n> > on http://www.lemoda.net/cgi/perl-unicode/index.html:\n> > \n> > \tuse CGI '-utf8';\n> > \tmy $value = params ('input');\n> > \n> > I tried it and that fixed my problem. I'm not sure about the\n> > consequences, maybe someone more experienced in CGI might help?\n> \n> I have reworded this to form a proper commit message (see\n> Documentation/SubmittingPatches) and I'll resend this as a reply to\n> this email.\n\nThanks, your message is much better.\n\n> \n> > ---\n> >  gitweb/gitweb.perl |    2 +-\n> >  1 files changed, 1 insertions(+), 1 deletions(-)\n> > \n> > diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> > index abb5a79..74d45b1 100755\n> > --- a/gitweb/gitweb.perl\n> > +++ b/gitweb/gitweb.perl\n> > @@ -10,7 +10,7 @@\n> >  use 5.008;\n> >  use strict;\n> >  use warnings;\n> > -use CGI qw(:standard :escapeHTML -nosticky);\n> > +use CGI qw(:standard :escapeHTML -nosticky -utf8);\n> >  use CGI::Util qw(unescape);\n> >  use CGI::Carp qw(fatalsToBrowser set_message);\n> >  use Encode;\n> > -- \n> \n> Does this actually work for you? \n\nYes. It correctly displays \"ł\" in the search form.\n\n>  Because it doesn't work for me\n> (perhaps I have too old CGI module: what CGI.pm and what Perl version\n> do you use?).\n> \n\n$ perl --version\n\nThis is perl 5, version 12, subversion 4 (v5.12.4) built for x86_64-linux\n(with 12 registered patches, see perl -V for more detail)\n\n$ eix -e CGI -c\n[I] perl-core/CGI (3.510@01.02.2012): Simple Common Gateway Interface Class\n\n> See other solution to this in other reply to this email.\n> \n"},{"id":"183660","messageId":"20120202214336.0c9daf9f@gmail.com","threadId":"29509","inReplyTo":"201202022108.51353.jnareb@gmail.com","subject":"Re: [PATCH/RFC (version A)] gitweb: use CGI with -utf8 to process Unicode query parameters correctly","fromName":"Michał Kiedrowicz","fromEmail":"michal.kiedrowicz@gmail.com","sentAt":"2012-02-02T20:43:36Z","receivedAt":"2012-02-02T20:43:36Z","isPatch":true,"sender":{"key":"michal.kiedrowicz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/14072847?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> wrote:\n\n> Gitweb tries hard to properly process UTF-8 data, by marking output\n> from git commands and contents of files as UTF-8 with to_utf8()\n> subroutine.  This ensures that gitweb would print correctly UTF-8\n> e.g. in 'log' and 'commit' views.\n> \n> Unfortunately it misses another source of potentially Unicode input,\n> namely query parameters.  The result is that one cannot search for a\n> string containing characters outside US-ASCII.  For example searching\n> for \"Michał Kiedrowicz\" (containing letter 'ł' - LATIN SMALL LETTER L\n> WITH STROKE, with Unicode codepoint U+0142, represented with 0xc5 0x82\n> bytes in UTF-8 and percent-encoded as %C5%81) result in the following\n> incorrect data in search field\n> \n> \tMichaÅ Kiedrowicz\n> \n> This is caused by CGI by default treating '0xc5 0x82' bytes as two\n> characters in Perl legacy encoding latin-1 (iso-8859-1), because 's'\n> query parameter is not processed explicitly as UTF-8 encoded string.\n> \n> According to \"Using Unicode in a Perl CGI script\" article on\n> http://www.lemoda.net/cgi/perl-unicode/index.html the simplest\n> solution is to just import '-utf8' pragma for CGI module:\n> \n> \tuse CGI '-utf8';\n> \tmy $value = params('input');\n> \n> According to CGI module documentation, the '-utf8' pragma may cause\n> problems with POST requests containing binary files... but gitweb\n> currently do not use POST requests at all, so this should be not a\n> problem now.\n\nThis was exactly my feeling  when I sent this patch :).\n\n> \n> Alternate solution would be to explicity decode query parameters when\n> storing them in %input_params (and perhaps also path_info).\n> \n> [jn: reworded / rewritten commit message]\n> \n> Signed-off-by: Michał Kiedrowicz <michal.kiedrowicz@gmail.com>\n\nThanks, I forgot about that.\n\n> Signed-off-by: Jakub Narębski <jnareb@gmail.com>\n> ---\n>  gitweb/gitweb.perl |    2 +-\n>  1 files changed, 1 insertions(+), 1 deletions(-)\n> \n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index 9cf7e71..a7441ef 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -10,7 +10,7 @@\n>  use 5.008;\n>  use strict;\n>  use warnings;\n> -use CGI qw(:standard :escapeHTML -nosticky);\n> +use CGI qw(:standard :escapeHTML -nosticky -utf8);\n>  use CGI::Util qw(unescape);\n>  use CGI::Carp qw(fatalsToBrowser set_message);\n>  use Encode;\n"},{"id":"183664","messageId":"20120202214646.1b84f23e@gmail.com","threadId":"29509","inReplyTo":"201202022110.07127.jnareb@gmail.com","subject":"Re: [PATCH/RFC (version B)] gitweb: Allow UTF-8 encoded CGI query parameters and path_info","fromName":"Michał Kiedrowicz","fromEmail":"michal.kiedrowicz@gmail.com","sentAt":"2012-02-02T20:46:46Z","receivedAt":"2012-02-02T20:46:46Z","isPatch":true,"sender":{"key":"michal.kiedrowicz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/14072847?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> wrote:\n\n> Gitweb tries hard to properly process UTF-8 data, by marking output\n> from git commands and contents of files as UTF-8 with to_utf8()\n> subroutine.  This ensures that gitweb would print correctly UTF-8\n> e.g. in 'log' and 'commit' views.\n> \n> Unfortunately it misses another source of potentially Unicode input,\n> namely query parameters.  The result is that one cannot search for a\n> string containing characters outside US-ASCII.  For example searching\n> for \"Michał Kiedrowicz\" (containing letter 'ł' - LATIN SMALL LETTER L\n> WITH STROKE, with Unicode codepoint U+0142, represented with 0xc5 0x82\n> bytes in UTF-8 and percent-encoded as %C5%81) result in the following\n> incorrect data in search field\n> \n> \tMichaÅ Kiedrowicz\n> \n> This is caused by CGI by default treating '0xc5 0x82' bytes as two\n> characters in Perl legacy encoding latin-1 (iso-8859-1), because 's'\n> query parameter is not processed explicitly as UTF-8 encoded string.\n> \n> The solution used here follows \"Using Unicode in a Perl CGI script\"\n> article on http://www.lemoda.net/cgi/perl-unicode/index.html:\n> \n> \tuse CGI;\n> \tuse Encode 'decode_utf8;\n> \tmy $value = params('input');\n> \t$value = decode_utf8($value);\n> \n> This is done when filling %input_params hash; this required to move\n> from explicit $cgi->param(<label>) to $input_params{<name>} in a few\n> places.\n\nI'm sorry but this doesn't work for me. I would be happy to help if you\nhave some questions about it.\n\n> \n> Alternate solution would be to simply use the '-utf8' pragma (via\n> \"use CGI '-utf8';\"), but according to CGI.pm documentation it may\n> cause problems with POST requests containing binary files... and\n> it doesn't work with old CGI.pm version 3.10 from Perl v5.8.6.\n> \n> Noticed-by: Michał Kiedrowicz <michal.kiedrowicz@gmail.com>\n> Signed-off-by: Jakub Narębski <jnareb@gmail.com>\n> ---\n>  gitweb/gitweb.perl |   12 ++++++------\n>  1 files changed, 6 insertions(+), 6 deletions(-)\n> \n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index 9cf7e71..55b2c24 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -52,7 +52,7 @@ sub evaluate_uri {\n>  \t# as base URL.\n>  \t# Therefore, if we needed to strip PATH_INFO, then we know that we have\n>  \t# to build the base URL ourselves:\n> -\tour $path_info = $ENV{\"PATH_INFO\"};\n> +\tour $path_info = decode_utf8($ENV{\"PATH_INFO\"});\n>  \tif ($path_info) {\n>  \t\tif ($my_url =~ s,\\Q$path_info\\E$,, &&\n>  \t\t    $my_uri =~ s,\\Q$path_info\\E$,, &&\n> @@ -816,9 +816,9 @@ sub evaluate_query_params {\n>  \n>  \twhile (my ($name, $symbol) = each %cgi_param_mapping) {\n>  \t\tif ($symbol eq 'opt') {\n> -\t\t\t$input_params{$name} = [ $cgi->param($symbol) ];\n> +\t\t\t$input_params{$name} = [ map { decode_utf8($_) } $cgi->param($symbol) ];\n>  \t\t} else {\n> -\t\t\t$input_params{$name} = $cgi->param($symbol);\n> +\t\t\t$input_params{$name} = decode_utf8($cgi->param($symbol));\n>  \t\t}\n>  \t}\n>  }\n> @@ -2767,7 +2767,7 @@ sub git_populate_project_tagcloud {\n>  \t}\n>  \n>  \tmy $cloud;\n> -\tmy $matched = $cgi->param('by_tag');\n> +\tmy $matched = $input_params{'ctag'};\n>  \tif (eval { require HTML::TagCloud; 1; }) {\n>  \t\t$cloud = HTML::TagCloud->new;\n>  \t\tforeach my $ctag (sort keys %ctags_lc) {\n> @@ -5282,7 +5282,7 @@ sub git_project_list_body {\n>  \n>  \tmy $check_forks = gitweb_check_feature('forks');\n>  \tmy $show_ctags  = gitweb_check_feature('ctags');\n> -\tmy $tagfilter = $show_ctags ? $cgi->param('by_tag') : undef;\n> +\tmy $tagfilter = $show_ctags ? $input_params{'ctag'} : undef;\n>  \t$check_forks = undef\n>  \t\tif ($tagfilter || $searchtext);\n>  \n> @@ -6197,7 +6197,7 @@ sub git_tag {\n>  \n>  sub git_blame_common {\n>  \tmy $format = shift || 'porcelain';\n> -\tif ($format eq 'porcelain' && $cgi->param('js')) {\n> +\tif ($format eq 'porcelain' && $input_params{'javascript'}) {\n>  \t\t$format = 'incremental';\n>  \t\t$action = 'blame_incremental'; # for page title etc\n>  \t}\n"},{"id":"183666","messageId":"201202022207.52220.jnareb@gmail.com","threadId":"29509","inReplyTo":"20120202214646.1b84f23e@gmail.com","subject":"Re: [PATCH/RFC (version B)] gitweb: Allow UTF-8 encoded CGI query parameters and path_info","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2012-02-02T21:07:51Z","receivedAt":"2012-02-02T21:07:51Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Thu, 2 Jan 2012, Michał Kiedrowicz wrote:\n> Jakub Narebski <jnareb@gmail.com> wrote:\n> \n> > Gitweb tries hard to properly process UTF-8 data, by marking output\n> > from git commands and contents of files as UTF-8 with to_utf8()\n> > subroutine.  This ensures that gitweb would print correctly UTF-8\n> > e.g. in 'log' and 'commit' views.\n> > \n> > Unfortunately it misses another source of potentially Unicode input,\n> > namely query parameters.  The result is that one cannot search for a\n> > string containing characters outside US-ASCII.  For example searching\n> > for \"Michał Kiedrowicz\" (containing letter 'ł' - LATIN SMALL LETTER L\n> > WITH STROKE, with Unicode codepoint U+0142, represented with 0xc5 0x82\n> > bytes in UTF-8 and percent-encoded as %C5%81) result in the following\n> > incorrect data in search field\n> > \n> > \tMichaÅ Kiedrowicz\n> > \n> > This is caused by CGI by default treating '0xc5 0x82' bytes as two\n> > characters in Perl legacy encoding latin-1 (iso-8859-1), because 's'\n> > query parameter is not processed explicitly as UTF-8 encoded string.\n> > \n> > The solution used here follows \"Using Unicode in a Perl CGI script\"\n> > article on http://www.lemoda.net/cgi/perl-unicode/index.html:\n> > \n> > \tuse CGI;\n> > \tuse Encode 'decode_utf8;\n> > \tmy $value = params('input');\n> > \t$value = decode_utf8($value);\n> > \n> > This is done when filling %input_params hash; this required to move\n> > from explicit $cgi->param(<label>) to $input_params{<name>} in a few\n> > places.\n> \n> I'm sorry but this doesn't work for me. I would be happy to help if you\n> have some questions about it.\n\nStrange.  http://www.lemoda.net/cgi/perl-unicode/index.html says that\nthose two approaches should be equivalent.  The -utf8 pragma version\ndoesn't work for me at all, while this one works in that if finds what\nit is supposed to, but shows garbage in search form.\n\nWill investigate.\n \n> > Alternate solution would be to simply use the '-utf8' pragma (via\n> > \"use CGI '-utf8';\"), but according to CGI.pm documentation it may\n> > cause problems with POST requests containing binary files... and\n> > it doesn't work with old CGI.pm version 3.10 from Perl v5.8.6.\n\n[...]\n> > @@ -816,9 +816,9 @@ sub evaluate_query_params {\n> >  \n> >  \twhile (my ($name, $symbol) = each %cgi_param_mapping) {\n> >  \t\tif ($symbol eq 'opt') {\n> > -\t\t\t$input_params{$name} = [ $cgi->param($symbol) ];\n> > +\t\t\t$input_params{$name} = [ map { decode_utf8($_) } $cgi->param($symbol) ];\n> >  \t\t} else {\n> > -\t\t\t$input_params{$name} = $cgi->param($symbol);\n> > +\t\t\t$input_params{$name} = decode_utf8($cgi->param($symbol));\n> >  \t\t}\n> >  \t}\n> >  }\n-- \nJakub Narebski\nPoland\n"},{"id":"183678","messageId":"201202022357.29569.jnareb@gmail.com","threadId":"29509","inReplyTo":"201202022207.52220.jnareb@gmail.com","subject":"Re: [PATCH/RFC (version B)] gitweb: Allow UTF-8 encoded CGI query parameters and path_info","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2012-02-02T22:57:29Z","receivedAt":"2012-02-02T22:57:29Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Thu, 2 Feb 2012, Jakub Narebski wrote:\n> On Thu, 2 Feb 2012, Michał Kiedrowicz wrote:\n> > Jakub Narebski <jnareb@gmail.com> wrote:\n> > \n> > > Gitweb tries hard to properly process UTF-8 data, by marking output\n> > > from git commands and contents of files as UTF-8 with to_utf8()\n> > > subroutine.  This ensures that gitweb would print correctly UTF-8\n> > > e.g. in 'log' and 'commit' views.\n> > > \n> > > Unfortunately it misses another source of potentially Unicode input,\n> > > namely query parameters.  The result is that one cannot search for a\n> > > string containing characters outside US-ASCII.  For example searching\n> > > for \"Michał Kiedrowicz\" (containing letter 'ł' - LATIN SMALL LETTER L\n> > > WITH STROKE, with Unicode codepoint U+0142, represented with 0xc5 0x82\n> > > bytes in UTF-8 and percent-encoded as %C5%81) result in the following\n> > > incorrect data in search field\n> > > \n> > > \tMichaÅ Kiedrowicz\n> > > \n> > > This is caused by CGI by default treating '0xc5 0x82' bytes as two\n> > > characters in Perl legacy encoding latin-1 (iso-8859-1), because 's'\n> > > query parameter is not processed explicitly as UTF-8 encoded string.\n> > > \n> > > The solution used here follows \"Using Unicode in a Perl CGI script\"\n> > > article on http://www.lemoda.net/cgi/perl-unicode/index.html:\n> > > \n> > > \tuse CGI;\n> > > \tuse Encode 'decode_utf8;\n> > > \tmy $value = params('input');\n> > > \t$value = decode_utf8($value);\n> > > \n> > > This is done when filling %input_params hash; this required to move\n> > > from explicit $cgi->param(<label>) to $input_params{<name>} in a few\n> > > places.\n> > \n> > I'm sorry but this doesn't work for me. I would be happy to help if you\n> > have some questions about it.\n> \n> Strange.  http://www.lemoda.net/cgi/perl-unicode/index.html says that\n> those two approaches should be equivalent.  The -utf8 pragma version\n> doesn't work for me at all, while this one works in that if finds what\n> it is supposed to, but shows garbage in search form.\n\nIs it what you mean by \"this doesn't work for me\", i.e. working search,\ngarbage in search field?\n \n> Will investigate.\n\nDamn.  If we use $cgi->textfield(-name => \"s\", -value => $searchtext) like\nin gitweb, CGI.pm would read $cgi->param(\"s\") by itself - without decoding.\nTo skip this we need to pass -force=>1  or  -override=>1 (i.e. further\nchanges to gitweb).\n\n-utf8 pragma works with more modern CGI.pm, but does not with 3.10.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"183700","messageId":"20120203083935.5d9d4b18@mkiedrowicz.ivo.pl","threadId":"29509","inReplyTo":"201202022357.29569.jnareb@gmail.com","subject":"Re: [PATCH/RFC (version B)] gitweb: Allow UTF-8 encoded CGI query parameters and path_info","fromName":"Michal Kiedrowicz","fromEmail":"michal.kiedrowicz@gmail.com","sentAt":"2012-02-03T07:39:35Z","receivedAt":"2012-02-03T07:39:35Z","isPatch":true,"sender":{"key":"michal.kiedrowicz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/14072847?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> wrote:\n\n> On Thu, 2 Feb 2012, Jakub Narebski wrote:\n> > On Thu, 2 Feb 2012, Michał Kiedrowicz wrote:\n> > > Jakub Narebski <jnareb@gmail.com> wrote:\n> > > \n> > > > Gitweb tries hard to properly process UTF-8 data, by marking\n> > > > output from git commands and contents of files as UTF-8 with\n> > > > to_utf8() subroutine.  This ensures that gitweb would print\n> > > > correctly UTF-8 e.g. in 'log' and 'commit' views.\n> > > > \n> > > > Unfortunately it misses another source of potentially Unicode\n> > > > input, namely query parameters.  The result is that one cannot\n> > > > search for a string containing characters outside US-ASCII.\n> > > > For example searching for \"Michał Kiedrowicz\" (containing\n> > > > letter 'ł' - LATIN SMALL LETTER L WITH STROKE, with Unicode\n> > > > codepoint U+0142, represented with 0xc5 0x82 bytes in UTF-8 and\n> > > > percent-encoded as %C5%81) result in the following incorrect\n> > > > data in search field\n> > > > \n> > > > \tMichaÅ Kiedrowicz\n> > > > \n> > > > This is caused by CGI by default treating '0xc5 0x82' bytes as\n> > > > two characters in Perl legacy encoding latin-1 (iso-8859-1),\n> > > > because 's' query parameter is not processed explicitly as\n> > > > UTF-8 encoded string.\n> > > > \n> > > > The solution used here follows \"Using Unicode in a Perl CGI\n> > > > script\" article on\n> > > > http://www.lemoda.net/cgi/perl-unicode/index.html:\n> > > > \n> > > > \tuse CGI;\n> > > > \tuse Encode 'decode_utf8;\n> > > > \tmy $value = params('input');\n> > > > \t$value = decode_utf8($value);\n> > > > \n> > > > This is done when filling %input_params hash; this required to\n> > > > move from explicit $cgi->param(<label>) to\n> > > > $input_params{<name>} in a few places.\n> > > \n> > > I'm sorry but this doesn't work for me. I would be happy to help\n> > > if you have some questions about it.\n> > \n> > Strange.  http://www.lemoda.net/cgi/perl-unicode/index.html says\n> > that those two approaches should be equivalent.  The -utf8 pragma\n> > version doesn't work for me at all, while this one works in that if\n> > finds what it is supposed to, but shows garbage in search form.\n> \n> Is it what you mean by \"this doesn't work for me\", i.e. working\n> search, garbage in search field?\n\nI mean \"garbage in search field\". Search works even without the patch\n(at least on Debian with git-1.7.7.3, perl-5.10.1 and CGI-3.43; I\ndon't have my notebook nearby at the moment to check).\n\n>  \n> > Will investigate.\n\nThanks for your time spending on this. I wouldn't call this problem\n\"production critial\" but it seems wrong to support UTF-8 everywhere\nproperly except for one place.\n\n> \n> Damn.  If we use $cgi->textfield(-name => \"s\", -value => $searchtext)\n> like in gitweb, CGI.pm would read $cgi->param(\"s\") by itself -\n> without decoding. \n\nMakes sense. When I tried calling to_utf8() in the line that defines\ntextfield (this was my first approach to this problem), it haven't\nchanged anything.\n\n> To skip this we need to pass -force=>1  or\n> -override=>1 (i.e. further changes to gitweb).\n> \n> -utf8 pragma works with more modern CGI.pm, but does not with 3.10.\n> \n"},{"id":"183713","messageId":"201202031344.55750.jnareb@gmail.com","threadId":"29509","inReplyTo":"20120203083935.5d9d4b18@mkiedrowicz.ivo.pl","subject":"[PATCH/RFCv2 (version B)] gitweb: Allow UTF-8 encoded CGI query parameters and path_info","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2012-02-03T12:44:54Z","receivedAt":"2012-02-03T12:44:54Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Gitweb tries hard to properly process UTF-8 data, by marking output\nfrom git commands and contents of files as UTF-8 with to_utf8()\nsubroutine.  This ensures that gitweb would print correctly UTF-8\ne.g. in 'log' and 'commit' views.\n\nUnfortunately it misses another source of potentially Unicode input,\nnamely query parameters.  The result is that one cannot search for a\nstring containing characters outside US-ASCII.  For example searching\nfor \"Michał Kiedrowicz\" (containing letter 'ł' - LATIN SMALL LETTER L\nWITH STROKE, with Unicode codepoint U+0142, represented with 0xc5 0x82\nbytes in UTF-8 and percent-encoded as %C5%81) result in the following\nincorrect data in search field\n\n\tMichaÅ Kiedrowicz\n\nThis is caused by CGI by default treating '0xc5 0x82' bytes as two\ncharacters in Perl legacy encoding latin-1 (iso-8859-1), because 's'\nquery parameter is not processed explicitly as UTF-8 encoded string.\n\nThe solution used here follows \"Using Unicode in a Perl CGI script\"\narticle on http://www.lemoda.net/cgi/perl-unicode/index.html:\n\n\tuse CGI;\n\tuse Encode 'decode_utf8;\n\tmy $value = params('input');\n\t$value = decode_utf8($value);\n\nDecoding UTF-8 is done when filling %input_params hash and $path_info\nvariable; the former required to move from explicit $cgi->param(<label>)\nto $input_params{<name>} in a few places, which is a good idea anyway.\n\nAnother required change was to add -override=>1 parameter to\n$cgi->textfield() invocation (in search form).  Otherwise CGI would\nuse values from query string if it is present, filling value from\n$cgi->param... without decode_utf8().  As we are using value of\nappropriate parameter anyway, -override=>1 doesn't change the\nsituation but makes gitweb fill search field correctly.\n\nAlternate solution would be to simply use the '-utf8' pragma (via\n\"use CGI '-utf8';\"), but according to CGI.pm documentation it may\ncause problems with POST requests containing binary files... and\nit requires CGI 3.31 (I think), released with perl v5.8.9.\n\nNoticed-by: Michał Kiedrowicz <michal.kiedrowicz@gmail.com>\nSigned-off-by: Jakub Narębski <jnareb@gmail.com>\n---\nOn Fri, 3 Feb 2012, Michal Kiedrowicz wrote:\n> Jakub Narebski <jnareb@gmail.com> wrote:\n\n> > Is it what you mean by \"this doesn't work for me\", i.e. working\n> > search, garbage in search field?\n> \n> I mean \"garbage in search field\". Search works even without the patch\n> (at least on Debian with git-1.7.7.3, perl-5.10.1 and CGI-3.43; I\n> don't have my notebook nearby at the moment to check).\n[...]\n\n> > Damn.  If we use $cgi->textfield(-name => \"s\", -value => $searchtext)\n> > like in gitweb, CGI.pm would read $cgi->param(\"s\") by itself -\n> > without decoding. \n> \n> Makes sense. When I tried calling to_utf8() in the line that defines\n> textfield (this was my first approach to this problem), it haven't\n> changed anything.\n\nYes, and it doesn't makes sense in gitweb case - we use value of \n$cgi->param(\"s\") as default value of text field anyway, but in\nUnicode-aware way.\n \n> > To skip this we need to pass -force=>1  or\n> > -override=>1 (i.e. further changes to gitweb).\n\nThis patch does this.  \n\nDoes it make work for you?\n\n> > -utf8 pragma works with more modern CGI.pm, but does not with 3.10.\n\n-utf8 pragma was added in CVS revision 1.238 of CGI.pm, which I think\nis present in CGI 3.31, released with perl v5.8.9.  Theoretically gitweb\nmaintains backward compatibility with perl v5.8.3 or something \n(\"use 5.008;\" but IIRC 5.8.3 is needed for correct Unicde handling anyway).\n\n gitweb/gitweb.perl |   16 ++++++++--------\n 1 files changed, 8 insertions(+), 8 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex 9cf7e71..bd5fff9 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -52,7 +52,7 @@ sub evaluate_uri {\n \t# as base URL.\n \t# Therefore, if we needed to strip PATH_INFO, then we know that we have\n \t# to build the base URL ourselves:\n-\tour $path_info = $ENV{\"PATH_INFO\"};\n+\tour $path_info = decode_utf8($ENV{\"PATH_INFO\"});\n \tif ($path_info) {\n \t\tif ($my_url =~ s,\\Q$path_info\\E$,, &&\n \t\t    $my_uri =~ s,\\Q$path_info\\E$,, &&\n@@ -816,9 +816,9 @@ sub evaluate_query_params {\n \n \twhile (my ($name, $symbol) = each %cgi_param_mapping) {\n \t\tif ($symbol eq 'opt') {\n-\t\t\t$input_params{$name} = [ $cgi->param($symbol) ];\n+\t\t\t$input_params{$name} = [ map { decode_utf8($_) } $cgi->param($symbol) ];\n \t\t} else {\n-\t\t\t$input_params{$name} = $cgi->param($symbol);\n+\t\t\t$input_params{$name} = decode_utf8($cgi->param($symbol));\n \t\t}\n \t}\n }\n@@ -2767,7 +2767,7 @@ sub git_populate_project_tagcloud {\n \t}\n \n \tmy $cloud;\n-\tmy $matched = $cgi->param('by_tag');\n+\tmy $matched = $input_params{'ctag'};\n \tif (eval { require HTML::TagCloud; 1; }) {\n \t\t$cloud = HTML::TagCloud->new;\n \t\tforeach my $ctag (sort keys %ctags_lc) {\n@@ -3873,7 +3873,7 @@ sub print_search_form {\n \t                       -values => ['commit', 'grep', 'author', 'committer', 'pickaxe']) .\n \t      $cgi->sup($cgi->a({-href => href(action=>\"search_help\")}, \"?\")) .\n \t      \" search:\\n\",\n-\t      $cgi->textfield(-name => \"s\", -value => $searchtext) . \"\\n\" .\n+\t      $cgi->textfield(-name => \"s\", -value => $searchtext, -override => 1) . \"\\n\" .\n \t      \"<span title=\\\"Extended regular expression\\\">\" .\n \t      $cgi->checkbox(-name => 'sr', -value => 1, -label => 're',\n \t                     -checked => $search_use_regexp) .\n@@ -5282,7 +5282,7 @@ sub git_project_list_body {\n \n \tmy $check_forks = gitweb_check_feature('forks');\n \tmy $show_ctags  = gitweb_check_feature('ctags');\n-\tmy $tagfilter = $show_ctags ? $cgi->param('by_tag') : undef;\n+\tmy $tagfilter = $show_ctags ? $input_params{'ctag'} : undef;\n \t$check_forks = undef\n \t\tif ($tagfilter || $searchtext);\n \n@@ -5994,7 +5994,7 @@ sub git_project_list {\n \t}\n \tprint $cgi->startform(-method => \"get\") .\n \t      \"<p class=\\\"projsearch\\\">Search:\\n\" .\n-\t      $cgi->textfield(-name => \"s\", -value => $searchtext) . \"\\n\" .\n+\t      $cgi->textfield(-name => \"s\", -value => $searchtext, -override => 1) . \"\\n\" .\n \t      \"</p>\" .\n \t      $cgi->end_form() . \"\\n\";\n \tgit_project_list_body(\\@list, $order);\n@@ -6197,7 +6197,7 @@ sub git_tag {\n \n sub git_blame_common {\n \tmy $format = shift || 'porcelain';\n-\tif ($format eq 'porcelain' && $cgi->param('js')) {\n+\tif ($format eq 'porcelain' && $input_params{'javascript'}) {\n \t\t$format = 'incremental';\n \t\t$action = 'blame_incremental'; # for page title etc\n \t}\n-- \n1.7.6\n"},{"id":"183745","messageId":"20120203184557.59042dec@gmail.com","threadId":"29509","inReplyTo":"201202031344.55750.jnareb@gmail.com","subject":"Re: [PATCH/RFCv2 (version B)] gitweb: Allow UTF-8 encoded CGI query parameters and path_info","fromName":"Michał Kiedrowicz","fromEmail":"michal.kiedrowicz@gmail.com","sentAt":"2012-02-03T17:45:57Z","receivedAt":"2012-02-03T17:45:57Z","isPatch":true,"sender":{"key":"michal.kiedrowicz@gmail.com","avatar":"https://avatars.githubusercontent.com/u/14072847?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> wrote:\n\n> Gitweb tries hard to properly process UTF-8 data, by marking output\n> from git commands and contents of files as UTF-8 with to_utf8()\n> subroutine.  This ensures that gitweb would print correctly UTF-8\n> e.g. in 'log' and 'commit' views.\n> \n> Unfortunately it misses another source of potentially Unicode input,\n> namely query parameters.  The result is that one cannot search for a\n> string containing characters outside US-ASCII.  For example searching\n> for \"Michał Kiedrowicz\" (containing letter 'ł' - LATIN SMALL LETTER L\n> WITH STROKE, with Unicode codepoint U+0142, represented with 0xc5 0x82\n> bytes in UTF-8 and percent-encoded as %C5%81) result in the following\n> incorrect data in search field\n> \n> \tMichaÅ Kiedrowicz\n> \n> This is caused by CGI by default treating '0xc5 0x82' bytes as two\n> characters in Perl legacy encoding latin-1 (iso-8859-1), because 's'\n> query parameter is not processed explicitly as UTF-8 encoded string.\n> \n> The solution used here follows \"Using Unicode in a Perl CGI script\"\n> article on http://www.lemoda.net/cgi/perl-unicode/index.html:\n> \n> \tuse CGI;\n> \tuse Encode 'decode_utf8;\n> \tmy $value = params('input');\n> \t$value = decode_utf8($value);\n> \n> Decoding UTF-8 is done when filling %input_params hash and $path_info\n> variable; the former required to move from explicit $cgi->param(<label>)\n> to $input_params{<name>} in a few places, which is a good idea anyway.\n> \n> Another required change was to add -override=>1 parameter to\n> $cgi->textfield() invocation (in search form).  Otherwise CGI would\n> use values from query string if it is present, filling value from\n> $cgi->param... without decode_utf8().  As we are using value of\n> appropriate parameter anyway, -override=>1 doesn't change the\n> situation but makes gitweb fill search field correctly.\n> \n> Alternate solution would be to simply use the '-utf8' pragma (via\n> \"use CGI '-utf8';\"), but according to CGI.pm documentation it may\n> cause problems with POST requests containing binary files... and\n> it requires CGI 3.31 (I think), released with perl v5.8.9.\n> \n> Noticed-by: Michał Kiedrowicz <michal.kiedrowicz@gmail.com>\n> Signed-off-by: Jakub Narębski <jnareb@gmail.com>\n> ---\n> On Fri, 3 Feb 2012, Michal Kiedrowicz wrote:\n> > Jakub Narebski <jnareb@gmail.com> wrote:\n> \n> > > Is it what you mean by \"this doesn't work for me\", i.e. working\n> > > search, garbage in search field?\n> > \n> > I mean \"garbage in search field\". Search works even without the patch\n> > (at least on Debian with git-1.7.7.3, perl-5.10.1 and CGI-3.43; I\n> > don't have my notebook nearby at the moment to check).\n> [...]\n> \n> > > Damn.  If we use $cgi->textfield(-name => \"s\", -value => $searchtext)\n> > > like in gitweb, CGI.pm would read $cgi->param(\"s\") by itself -\n> > > without decoding. \n> > \n> > Makes sense. When I tried calling to_utf8() in the line that defines\n> > textfield (this was my first approach to this problem), it haven't\n> > changed anything.\n> \n> Yes, and it doesn't makes sense in gitweb case - we use value of \n> $cgi->param(\"s\") as default value of text field anyway, but in\n> Unicode-aware way.\n>  \n> > > To skip this we need to pass -force=>1  or\n> > > -override=>1 (i.e. further changes to gitweb).\n> \n> This patch does this.  \n> \n> Does it make work for you?\n> \n\nYes, it works for me. Search form properly displays \"ł\". Thanks!\n"},{"id":"183760","messageId":"7vzkczmxon.fsf@alter.siamese.dyndns.org","threadId":"29509","inReplyTo":"201202031344.55750.jnareb@gmail.com","subject":"Re: [PATCH/RFCv2 (version B)] gitweb: Allow UTF-8 encoded CGI query parameters and path_info","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-02-03T21:09:12Z","receivedAt":"2012-02-03T21:09:12Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> Gitweb tries hard to properly process UTF-8 data, by marking output\n> from git commands and contents of files as UTF-8 with to_utf8()\n> subroutine.  This ensures that gitweb would print correctly UTF-8\n> e.g. in 'log' and 'commit' views.\n>\n> Unfortunately it misses another source of potentially Unicode input,\n> namely query parameters.  The result is that one cannot search for a\n\nI think two lines should suffice instead of the above two paragraphs.\n\n        Gitweb forgot to turn query parameters into UTF-8. This results in\n        a bug that one cannot search for a\n\n> string containing characters outside US-ASCII.  For example searching\n> ...\n\nThe remainder explains the problem and the solution very well modulo minor\ntypos.\n\n> Noticed-by: Michał Kiedrowicz <michal.kiedrowicz@gmail.com>\n> Signed-off-by: Jakub Narębski <jnareb@gmail.com>\n> ---\n\nWe can add \"Tested-by:\" to this now.  Will queue.\n\nThanks.\n"}]}