threads / rfc / 36657

RFC patchGitweb: Convert UTF-8 encoded file names

Subject: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

## tl;dr

21 messages between May 14, 2014 and Jun 4, 2014. Diffs are folded; open one to read it.

replies: 20people: 5as markdown or json

Michael Wagner· May 14, 2014, 18:41 UTC · lore

Perl has an internal encoding used to store text strings. Currently, trying to view files with UTF-8 encoded names results in an error (either "404 - Cannot find file" [blob_plain] or "XML Parsing Error" [blob]). Converting these UTF-8 encoded file names into Perl's internal format resolves these errors.

Signed-off-by: Michael Wagner <accounts@mwagner.org>
---
 gitweb/gitweb.perl | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)
Show changes to gitweb/gitweb.perl +1 −1
diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
index a9f57d6..6046977 100755
--- a/gitweb/gitweb.perl
+++ b/gitweb/gitweb.perl
@@ -1056,7 +1056,7 @@ sub evaluate_and_validate_params {
 		}
 	}
 
-	our $file_name = $input_params{'file_name'};
+	our $file_name = decode("utf-8", $input_params{'file_name'});
 	if (defined $file_name) {
 		if (!is_valid_pathname($file_name)) {
 			die_error(400, "Invalid file parameter");
-- 
1.9.0
Junio C Hamano· May 14, 2014, 21:57 UTC · re: Michael Wagner · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

Michael Wagner <accounts@mwagner.org> writes:
Show 7 quoted lines
> Perl has an internal encoding used to store text strings. Currently, trying to
> view files with UTF-8 encoded names results in an error (either "404 - Cannot
> find file" [blob_plain] or "XML Parsing Error" [blob]). Converting these UTF-8
> encoded file names into Perl's internal format resolves these errors.
>
> Signed-off-by: Michael Wagner <accounts@mwagner.org>
> ---
Cc'ing Jakub, who have been the area maintainer, for comments.

One thing I wonder is that, if there are some additional calls to encode() necessary before we embed $file_name (which are now decoded to the internal string form, not a byte-sequence that happens to be in utf-8) in the generated pages, if we were to do this change.

Show 16 quoted lines
>  gitweb/gitweb.perl | 2 +-
>  1 file changed, 1 insertion(+), 1 deletion(-)
>
> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
> index a9f57d6..6046977 100755
> --- a/gitweb/gitweb.perl
> +++ b/gitweb/gitweb.perl
> @@ -1056,7 +1056,7 @@ sub evaluate_and_validate_params {
>  		}
>  	}
>  
> -	our $file_name = $input_params{'file_name'};
> +	our $file_name = decode("utf-8", $input_params{'file_name'});
>  	if (defined $file_name) {
>  		if (!is_valid_pathname($file_name)) {
>  			die_error(400, "Invalid file parameter");
Jakub Narębski· May 14, 2014, 22:25 UTC · re: Junio C Hamano · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

On Wed, May 14, 2014 at 11:57 PM, Junio C Hamano <gitster@pobox.com> wrote:
Show 6 quoted lines
> Michael Wagner <accounts@mwagner.org> writes:
>
>> Perl has an internal encoding used to store text strings. Currently, trying to
>> view files with UTF-8 encoded names results in an error (either "404 - Cannot
>> find file" [blob_plain] or "XML Parsing Error" [blob]). Converting these UTF-8
>> encoded file names into Perl's internal format resolves these errors.

Could you give us an example? What is important is whether filename is passed via path_info or via query string.

Because in evaluate_uri() there is
     our $path_info = decode_utf8($ENV{"PATH_INFO"});
and in evaluate_query_params() there is
    $input_params{$name} = decode_utf8($cgi->param($symbol));
Show 9 quoted lines
>> Signed-off-by: Michael Wagner <accounts@mwagner.org>
>> ---
>
> Cc'ing Jakub, who have been the area maintainer, for comments.
>
> One thing I wonder is that, if there are some additional calls to
> encode() necessary before we embed $file_name (which are now decoded
> to the internal string form, not a byte-sequence that happens to be
> in utf-8) in the generated pages, if we were to do this change.

There should be no problem with output encoding. esc_path(), which should be used for filenames, includes to_utf8, which in turn uses decode($fallback_encoding, $str, Encode::FB_DEFAULT);

Show 16 quoted lines
>>  gitweb/gitweb.perl | 2 +-
>>  1 file changed, 1 insertion(+), 1 deletion(-)
>>
>> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
>> index a9f57d6..6046977 100755
>> --- a/gitweb/gitweb.perl
>> +++ b/gitweb/gitweb.perl
>> @@ -1056,7 +1056,7 @@ sub evaluate_and_validate_params {
>>               }
>>       }
>>
>> -     our $file_name = $input_params{'file_name'};
>> +     our $file_name = decode("utf-8", $input_params{'file_name'});
>>       if (defined $file_name) {
>>               if (!is_valid_pathname($file_name)) {
>>                       die_error(400, "Invalid file parameter");

Hmm... all %input_params should have been properly decoded already, how it was missed?

Also, branchname (hash_base etc.), search query, filename in file_parent, project name can be UTF-8 too, so it is at best partial fix.

-- 
Jakub Narębski
Michael Wagner· May 15, 2014, 05:08 UTC · re: Jakub Narębski · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

On Thu, May 15, 2014 at 12:25:45AM +0200, Jakub Narębski wrote:
Show 11 quoted lines
> On Wed, May 14, 2014 at 11:57 PM, Junio C Hamano <gitster@pobox.com> wrote:
> > Michael Wagner <accounts@mwagner.org> writes:
> >
> >> Perl has an internal encoding used to store text strings. Currently, trying to
> >> view files with UTF-8 encoded names results in an error (either "404 - Cannot
> >> find file" [blob_plain] or "XML Parsing Error" [blob]). Converting these UTF-8
> >> encoded file names into Perl's internal format resolves these errors.
> 
> Could you give us an example?  What is important is whether filename
> is passed via path_info or via query string.
> 

There is a file named "Gütekriterien.txt" in my repository. Trying to view this file as "blob_plain" produces an 404 error (displaying the file name with an additional print statement):

$ REQUEST_METHOD=GET QUERY_STRING='p=notes.git;a=blob_plain;f=work/G%C3%83%C2%BCtekriterien.txt;hb=HEAD' ./gitweb.cgi
work/Gütekriterien.txt
Status: 404 Not Found

Decoding the UTF-8 encoded file name (again with an additional print statement):

$ REQUEST_METHOD=GET QUERY_STRING='p=notes.git;a=blob_plain;f=work/G%C3%83%C2%BCtekriterien.txt;hb=HEAD' ./gitweb.cgi
work/Gütekriterien.txt
Content-disposition: inline; filename="work/Gütekriterien.txt"
Show 17 quoted lines
> Because in evaluate_uri() there is
> 
>      our $path_info = decode_utf8($ENV{"PATH_INFO"});
> 
> and in evaluate_query_params() there is
> 
>     $input_params{$name} = decode_utf8($cgi->param($symbol));
> 
> >> Signed-off-by: Michael Wagner <accounts@mwagner.org>
> >> ---
> >
> > Cc'ing Jakub, who have been the area maintainer, for comments.
> >
> > One thing I wonder is that, if there are some additional calls to
> > encode() necessary before we embed $file_name (which are now decoded
> > to the internal string form, not a byte-sequence that happens to be
> > in utf-8) in the generated pages, if we were to do this change.
The generated pages show the correct file names. 
Show 30 quoted lines
> 
> There should be no problem with output encoding.  esc_path(), which
> should be used for filenames, includes to_utf8, which in turn uses
> decode($fallback_encoding, $str, Encode::FB_DEFAULT);
> 
> >>  gitweb/gitweb.perl | 2 +-
> >>  1 file changed, 1 insertion(+), 1 deletion(-)
> >>
> >> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
> >> index a9f57d6..6046977 100755
> >> --- a/gitweb/gitweb.perl
> >> +++ b/gitweb/gitweb.perl
> >> @@ -1056,7 +1056,7 @@ sub evaluate_and_validate_params {
> >>               }
> >>       }
> >>
> >> -     our $file_name = $input_params{'file_name'};
> >> +     our $file_name = decode("utf-8", $input_params{'file_name'});
> >>       if (defined $file_name) {
> >>               if (!is_valid_pathname($file_name)) {
> >>                       die_error(400, "Invalid file parameter");
> 
> Hmm... all %input_params should have been properly decoded
> already, how it was missed?
> 
> Also, branchname (hash_base etc.), search query, filename in file_parent,
> project name can be UTF-8 too, so it is at best partial fix.
> 
> -- 
> Jakub Narębski
Peter Krefting· May 15, 2014, 09:04 UTC · re: Michael Wagner · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

Michael Wagner:
Show 7 quoted lines
> Decoding the UTF-8 encoded file name (again with an additional print
> statement):
>
> $ REQUEST_METHOD=GET QUERY_STRING='p=notes.git;a=blob_plain;f=work/G%C3%83%C2%BCtekriterien.txt;hb=HEAD' ./gitweb.cgi
>
> work/Gütekriterien.txt
> Content-disposition: inline; filename="work/Gütekriterien.txt"

You should fix the code path that created that URI, though, as it is not what you expected.

%C3%83 decodes to U+00C3 Latin Capital Letter A With Tilde %C2%BC decodes to U+00BC Vulgar Graction One Quarter

The proper UTF-8 encoding for ü (U+00FC) is, as you can probably guess from looking at which two characters the sequence above yielded, C3 BC, which in a URI is represented as %C3%BC.

Your QUERY_STRING should thus be
   p=notes.git;a=blob_plain;f=work/G%C3%BCtekriterien.txt;hb=HEAD
which probably works as expected.

What is happening is that whatever is generating the URI us UTF-8-encoding the string twice (i.e., it generates a string with the proper C3 BC in it, and then interprets it as iso-8859-1 data and runs that through a UTF-8 encoder again, yielding the C3 83 C2 BC sequence you see above).

-- 
\\// Peter - http://www.softwolves.pp.se/
Jakub Narębski· May 15, 2014, 12:32 UTC · re: Michael Wagner · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

On Thu, May 15, 2014 at 7:08 AM, Michael Wagner <accounts@mwagner.org> wrote:
Show 21 quoted lines
> On Thu, May 15, 2014 at 12:25:45AM +0200, Jakub Narębski wrote:
>> On Wed, May 14, 2014 at 11:57 PM, Junio C Hamano <gitster@pobox.com> wrote:
>>> Michael Wagner <accounts@mwagner.org> writes:
>>>
>>>> Perl has an internal encoding used to store text strings. Currently, trying to
>>>> view files with UTF-8 encoded names results in an error (either "404 - Cannot
>>>> find file" [blob_plain] or "XML Parsing Error" [blob]). Converting these UTF-8
>>>> encoded file names into Perl's internal format resolves these errors.
>>
>> Could you give us an example?  What is important is whether filename
>> is passed via path_info or via query string.
>>
>
> There is a file named "Gütekriterien.txt" in my repository. Trying to
> view this file as "blob_plain" produces an 404 error (displaying the
> file name with an additional print statement):
>
> $ REQUEST_METHOD=GET QUERY_STRING='p=notes.git;a=blob_plain;f=work/G%C3%83%C2%BCtekriterien.txt;hb=HEAD' ./gitweb.cgi
>
> work/Gütekriterien.txt
> Status: 404 Not Found

You have URI encoding of "ü" wrong! "ü" encodes as %C3%BC, not as %C3%83%C2%BC (4 bytes?)

  http://www.url-encode-decode.com/
You tested with wrong input.

BTW. there probably should be test for UTF-8 encoding, similar to the one for XSS in t9502-gitweb-standalone-parse-output

-- 
Jakub Narębski
Junio C Hamano· May 15, 2014, 17:24 UTC · re: Peter Krefting · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

Peter Krefting <peter@softwolves.pp.se> writes:
Show 5 quoted lines
> What is happening is that whatever is generating the URI us
> UTF-8-encoding the string twice (i.e., it generates a string with the
> proper C3 BC in it, and then interprets it as iso-8859-1 data and runs
> that through a UTF-8 encoder again, yielding the C3 83 C2 BC sequence
> you see above).

Thanks for a quick response. If the input was unnecessarily encoded one extra time, it is no wonder it needed one unnecessary extra decoding.

Michael Wagner· May 15, 2014, 18:48 UTC · re: Peter Krefting · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

On Thu, May 15, 2014 at 10:04:24AM +0100, Peter Krefting wrote:
Show 25 quoted lines
> Michael Wagner:
> 
> >Decoding the UTF-8 encoded file name (again with an additional print
> >statement):
> >
> >$ REQUEST_METHOD=GET QUERY_STRING='p=notes.git;a=blob_plain;f=work/G%C3%83%C2%BCtekriterien.txt;hb=HEAD' ./gitweb.cgi
> >
> >work/Gütekriterien.txt
> >Content-disposition: inline; filename="work/Gütekriterien.txt"
> 
> You should fix the code path that created that URI, though, as it is not
> what you expected.
> 
> %C3%83 decodes to U+00C3 Latin Capital Letter A With Tilde
> %C2%BC decodes to U+00BC Vulgar Graction One Quarter
> 
> The proper UTF-8 encoding for ü (U+00FC) is, as you can probably guess from
> looking at which two characters the sequence above yielded, C3 BC, which in
> a URI is represented as %C3%BC.
> 
> Your QUERY_STRING should thus be
> 
>   p=notes.git;a=blob_plain;f=work/G%C3%BCtekriterien.txt;hb=HEAD
> 
> which probably works as expected.
Obviously, you are right, thanks.
Show 6 quoted lines
> 
> What is happening is that whatever is generating the URI us UTF-8-encoding
> the string twice (i.e., it generates a string with the proper C3 BC in it,
> and then interprets it as iso-8859-1 data and runs that through a UTF-8
> encoder again, yielding the C3 83 C2 BC sequence you see above).
> 

The subroutine "git tree" generates the tree view. It stores the output of "git ls-tree -z ..." in an array named "@entries". Printing the content of this array yields the following result:

00644 blob 6419cd06a9461c38d4f94d9705d97eaaa887156a     520 Gütekriterien.txt

This leads to the "doubled" encoding. Declaring the encoding in the call to open yields the following result:

100644 blob 6419cd06a9461c38d4f94d9705d97eaaa887156a     520 Gütekriterien.txt
---
Show changes to gitweb/gitweb.perl +1 −1
diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
index a9f57d6..f1414e1 100755
--- a/gitweb/gitweb.perl
+++ b/gitweb/gitweb.perl
@@ -7138,7 +7138,7 @@ sub git_tree {
        my @entries = ();
        {
                local $/ = "\0";
-               open my $fd, "-|", git_cmd(), "ls-tree", '-z',
+               open my $fd, "-|encoding(UTF-8)", git_cmd(), "ls-tree", '-z',
                        ($show_sizes ? '-l' : ()), @extra_options, $hash
                        or die_error(500, "Open git-ls-tree failed");
                @entries = map { chomp; $_ } <$fd>;

> -- 
> \\// Peter - http://www.softwolves.pp.se/
Jakub Narębski· May 15, 2014, 19:28 UTC · re: Michael Wagner · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

On Thu, May 15, 2014 at 8:48 PM, Michael Wagner <accounts@mwagner.org> wrote:
Show 42 quoted lines
> On Thu, May 15, 2014 at 10:04:24AM +0100, Peter Krefting wrote:
>> Michael Wagner:
>>
>>>Decoding the UTF-8 encoded file name (again with an additional print
>>>statement):
>>>
>>>$ REQUEST_METHOD=GET QUERY_STRING='p=notes.git;a=blob_plain;f=work/G%C3%83%C2%BCtekriterien.txt;hb=HEAD' ./gitweb.cgi
>>>
>>>work/Gütekriterien.txt
>>>Content-disposition: inline; filename="work/Gütekriterien.txt"
>>
>> You should fix the code path that created that URI, though, as it is not
>> what you expected.
>>
>> %C3%83 decodes to U+00C3 Latin Capital Letter A With Tilde
>> %C2%BC decodes to U+00BC Vulgar Graction One Quarter
>>
>> The proper UTF-8 encoding for ü (U+00FC) is, as you can probably guess from
>> looking at which two characters the sequence above yielded, C3 BC, which in
>> a URI is represented as %C3%BC.
>>
>> Your QUERY_STRING should thus be
>>
>>   p=notes.git;a=blob_plain;f=work/G%C3%BCtekriterien.txt;hb=HEAD
>>
>> which probably works as expected.
>>
>> What is happening is that whatever is generating the URI us UTF-8-encoding
>> the string twice (i.e., it generates a string with the proper C3 BC in it,
>> and then interprets it as iso-8859-1 data and runs that through a UTF-8
>> encoder again, yielding the C3 83 C2 BC sequence you see above).
>
> The subroutine "git tree" generates the tree view. It stores the output
> of "git ls-tree -z ..." in an array named "@entries". Printing the content
> of this array yields the following result:
>
> 00644 blob 6419cd06a9461c38d4f94d9705d97eaaa887156a     520 Gütekriterien.txt
>
> This leads to the "doubled" encoding. Declaring the encoding in the call
> to open yields the following result:
>
> 100644 blob 6419cd06a9461c38d4f94d9705d97eaaa887156a     520 Gütekriterien.txt
Good catch.

Writing test for this would not be easy, and require some HTML parser (WWW::Mechanize, Web::Scraper, HTML::Query, pQuery, ... or low level HTML::TreeBuilder, or other low level parser).

Show 14 quoted lines
> ---
>
> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
> index a9f57d6..f1414e1 100755
> --- a/gitweb/gitweb.perl
> +++ b/gitweb/gitweb.perl
> @@ -7138,7 +7138,7 @@ sub git_tree {
>         my @entries = ();
>         {
>                 local $/ = "\0";
> -               open my $fd, "-|", git_cmd(), "ls-tree", '-z',
> +               open my $fd, "-|encoding(UTF-8)", git_cmd(), "ls-tree", '-z',
>                         ($show_sizes ? '-l' : ()), @extra_options, $hash
>                         or die_error(500, "Open git-ls-tree failed");
Or put
                   binmode $fd, ':utf8';
like in the rest of the code.
>                 @entries = map { chomp; $_ } <$fd>;
>
Even better solution would be to use
    use open IN => ':encoding(utf-8)';
at the beginning of gitweb.perl, once and for all.

Unfortunately the output equivalent requires creating Perl module for gitweb, to be able to use

    use open OUT => ':encoding(utf-8-with-fallback)';
-- 
Jakub Narebski
Jakub Narębski· May 15, 2014, 19:37 UTC · re: Jakub Narębski · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

On Thu, May 15, 2014 at 9:28 PM, Jakub Narębski <jnareb@gmail.com> wrote:
> On Thu, May 15, 2014 at 8:48 PM, Michael Wagner <accounts@mwagner.org> wrote:
[...]
Show 10 quoted lines
>> The subroutine "git tree" generates the tree view. It stores the output
>> of "git ls-tree -z ..." in an array named "@entries". Printing the content
>> of this array yields the following result:
>>
>> 00644 blob 6419cd06a9461c38d4f94d9705d97eaaa887156a     520 Gütekriterien.txt
>>
>> This leads to the "doubled" encoding. Declaring the encoding in the call
>> to open yields the following result:
>>
>> 100644 blob 6419cd06a9461c38d4f94d9705d97eaaa887156a     520 Gütekriterien.txt
Show 20 quoted lines
>> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
>> index a9f57d6..f1414e1 100755
>> --- a/gitweb/gitweb.perl
>> +++ b/gitweb/gitweb.perl
>> @@ -7138,7 +7138,7 @@ sub git_tree {
>>         my @entries = ();
>>         {
>>                 local $/ = "\0";
>> -               open my $fd, "-|", git_cmd(), "ls-tree", '-z',
>> +               open my $fd, "-|encoding(UTF-8)", git_cmd(), "ls-tree", '-z',
>>                         ($show_sizes ? '-l' : ()), @extra_options, $hash
>>                         or die_error(500, "Open git-ls-tree failed");
>
> Or put
>
>                    binmode $fd, ':utf8';
>
> like in the rest of the code.
>
>>                 @entries = map { chomp; $_ } <$fd>;

Though to be exact there isn't any mechanism that ensures that filenames in tree objects use utf-8 encoding, so perhaps a safer solution would be to use

   to_utf8($file_name)
(which respects $fallback_encoding) in appropriate places.
-- 
Jakub Narębski
Junio C Hamano· May 15, 2014, 19:38 UTC · re: Jakub Narębski · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

Jakub Narębski <jnareb@gmail.com> writes:
> Writing test for this would not be easy, and require some HTML
> parser (WWW::Mechanize, Web::Scraper, HTML::Query, pQuery,
> ... or low level HTML::TreeBuilder, or other low level parser).

Hmph. Is it more than just looking for a specific run of %xx we would expect to see in the output of the tree view for a repository in which there is one tree with non-ASCII name?

Show 18 quoted lines
>> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
>> index a9f57d6..f1414e1 100755
>> --- a/gitweb/gitweb.perl
>> +++ b/gitweb/gitweb.perl
>> @@ -7138,7 +7138,7 @@ sub git_tree {
>>         my @entries = ();
>>         {
>>                 local $/ = "\0";
>> -               open my $fd, "-|", git_cmd(), "ls-tree", '-z',
>> +               open my $fd, "-|encoding(UTF-8)", git_cmd(), "ls-tree", '-z',
>>                         ($show_sizes ? '-l' : ()), @extra_options, $hash
>>                         or die_error(500, "Open git-ls-tree failed");
>
> Or put
>
>                    binmode $fd, ':utf8';
>
> like in the rest of the code.

I expect a patch to do so and can forget about this thread myself, then, OK?

Thanks all for digging this to the root.
Jakub Narębski· May 15, 2014, 20:45 UTC · re: Junio C Hamano · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

On Thu, May 15, 2014 at 9:38 PM, Junio C Hamano <gitster@pobox.com> wrote:
Show 9 quoted lines
> Jakub Narębski <jnareb@gmail.com> writes:
>
>> Writing test for this would not be easy, and require some HTML
>> parser (WWW::Mechanize, Web::Scraper, HTML::Query, pQuery,
>> ... or low level HTML::TreeBuilder, or other low level parser).
>
> Hmph.  Is it more than just looking for a specific run of %xx we
> would expect to see in the output of the tree view for a repository
> in which there is one tree with non-ASCII name?

There is if we want to check (in non-fragile way) that said specific run is in 'href' *attribute* of 'a' element (link target).

-- 
Jakub Narebski
Junio C Hamano· May 16, 2014, 01:26 UTC · re: Jakub Narębski · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

Jakub Narębski <jnareb@gmail.com> writes:
Show 13 quoted lines
> On Thu, May 15, 2014 at 9:38 PM, Junio C Hamano <gitster@pobox.com> wrote:
>> Jakub Narębski <jnareb@gmail.com> writes:
>>
>>> Writing test for this would not be easy, and require some HTML
>>> parser (WWW::Mechanize, Web::Scraper, HTML::Query, pQuery,
>>> ... or low level HTML::TreeBuilder, or other low level parser).
>>
>> Hmph.  Is it more than just looking for a specific run of %xx we
>> would expect to see in the output of the tree view for a repository
>> in which there is one tree with non-ASCII name?
>
> There is if we want to check (in non-fragile way) that said
> specific run is in 'href' *attribute* of 'a' element (link target).

Correct, but is "where does it appear" the question we are primarily interested in, wrt this breakage and its fix?

If gitweb output has some volatile parts that do not depend on the contents of the Git test repository (e.g. showing contents of /etc/motd, date/time of when the test was run, or the full pathname leading to the trash directory), then preparing a tree whose name is äéìõû and making sure that the properly encoded version of äéìõû appears anywhere in the output may not be sufficient to validate that we got the encoding right, as that string may appear in the parts that are totally unrelated to the contents being shown and not under our control. But is that really the case?

Also we may introduce a bug and misspell the attr name and produce an anchor element with hpef attribute with the properly encoded URL in it, and your "parse HTML properly" approach would catch it, but is that the kind of breakage under discussion? You hinted at new tests for UTF-8 encoding in the other message in the thread earlier, and I've been assuming that we were talking about the encoding test, not a test to catch s/href/hpef/ kind of breakage.

Jakub Narębski· May 16, 2014, 07:54 UTC · re: Junio C Hamano · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

On Fri, May 16, 2014 at 3:26 AM, Junio C Hamano <gitster@pobox.com> wrote:
Show 17 quoted lines
> Jakub Narębski <jnareb@gmail.com> writes:
>> On Thu, May 15, 2014 at 9:38 PM, Junio C Hamano <gitster@pobox.com> wrote:
>>> Jakub Narębski <jnareb@gmail.com> writes:
>>>
>>>> Writing test for this would not be easy, and require some HTML
>>>> parser (WWW::Mechanize, Web::Scraper, HTML::Query, pQuery,
>>>> ... or low level HTML::TreeBuilder, or other low level parser).
>>>
>>> Hmph.  Is it more than just looking for a specific run of %xx we
>>> would expect to see in the output of the tree view for a repository
>>> in which there is one tree with non-ASCII name?
>>
>> There is if we want to check (in non-fragile way) that said
>> specific run is in 'href' *attribute* of 'a' element (link target).
>
> Correct, but is "where does it appear" the question we are
> primarily interested in, wrt this breakage and its fix?

That of course depends on how we want to test gitweb output. The simplest solution, comparing with known output with perhaps fragile / variable elements masked out could be done quickly... but changes in output (even if they don't change functionality, or don't change visible output) require regenerating test cases (expected output) to test against - which might be source of errors in test suite.

Another simple solution, grepping for expected strings, also easy to create, has the disadvantage of being only positive test - you cannot [easily] test that there are no *wrong* output, only that right string exists somewhere.

Show 9 quoted lines
> If gitweb output has some volatile parts that do not depend on the
> contents of the Git test repository (e.g. showing contents of
> /etc/motd, date/time of when the test was run, or the full pathname
> leading to the trash directory), then preparing a tree whose name is
> äéìõû and making sure that the properly encoded version of äéìõû
> appears anywhere in the output may not be sufficient to validate
> that we got the encoding right, as that string may appear in the
> parts that are totally unrelated to the contents being shown and not
> under our control.  But is that really the case?

Well, I guess that any test is better than no test (though OTOH Heartbleed and "goto fail" bugs shows the importance of negative tests).

Show 7 quoted lines
> Also we may introduce a bug and misspell the attr name and produce
> an anchor element with hpef attribute with the properly encoded URL
> in it, and your "parse HTML properly" approach would catch it, but
> is that the kind of breakage under discussion?  You hinted at new
> tests for UTF-8 encoding in the other message in the thread earlier,
> and I've been assuming that we were talking about the encoding test,
> not a test to catch s/href/hpef/ kind of breakage.

One of tests possible with HTML parser (e.g. WWW::Mechanize::CGI) is to check that all [internal] links leads to 200-OK pages, which accidentally would also be a test against this breakage.

-- 
Jakub Narebski
Junio C Hamano· May 16, 2014, 17:05 UTC · re: Jakub Narębski · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

Jakub Narębski <jnareb@gmail.com> writes:
Show 10 quoted lines
>> Correct, but is "where does it appear" the question we are
>> primarily interested in, wrt this breakage and its fix?
>
> That of course depends on how we want to test gitweb output.
> The simplest solution, comparing with known output with perhaps
> fragile / variable elements masked out could be done quickly...
> but changes in output (even if they don't change functionality,
> or don't change visible output) require regenerating test cases
> (expected output) to test against - which might be source of
> errors in test suite.

I agree with your "to test it fully, we need extra dependencies", but my point is that it does not have to be a full "HTML-validating, picking the expected attribute via XPATH matching" kind of test if what we want is only to add a new test to protect this particular fix from future breakages.

For example, I think it is sufficient to grep for 'href="...%xx%xx"' in the output after preparing a sample tree with one entry to show. The expected substring either exists (in which case we got it right), or it doesn't (in which case we are showing garbage). Of course that depends on the assumption that its output is not too heavily contaminated with volatile parts outside our control, as I already mentioned in the message you are responding to.

But it all depends on "if" we wanted to add a new test ;-)
Junio C Hamano· May 16, 2014, 18:17 UTC · re: Jakub Narębski · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

(sorry if you receive a dup; pobox.com seems to be constipated right now)
Jakub Narębski <jnareb@gmail.com> writes:
Show 10 quoted lines
>> Correct, but is "where does it appear" the question we are
>> primarily interested in, wrt this breakage and its fix?
>
> That of course depends on how we want to test gitweb output.
> The simplest solution, comparing with known output with perhaps
> fragile / variable elements masked out could be done quickly...
> but changes in output (even if they don't change functionality,
> or don't change visible output) require regenerating test cases
> (expected output) to test against - which might be source of
> errors in test suite.

I agree with your "to test it fully, we need extra dependencies", but my point is that it does not have to be a full "HTML-validating, picking the expected attribute via XPATH matching" kind of test if what we want is only to add a new test to protect this particular fix from future breakages.

For example, I think it is sufficient to grep for 'href="...%xx%xx"' in the output after preparing a sample tree with one entry to show. The expected substring either exists (in which case we got it right), or it doesn't (in which case we are showing garbage). Of course that depends on the assumption that its output is not too heavily contaminated with volatile parts outside our control, as I already mentioned in the message you are responding to.

But it all depends on "if" we wanted to add a new test ;-)
Jakub Narębski· May 27, 2014, 14:18 UTC · re: Junio C Hamano · lore

Re: [PATCH/RFC] Gitweb: Convert UTF-8 encoded file names

W dniu 2014-05-16 19:05, Junio C Hamano pisze:
Show 28 quoted lines
> Jakub Narębski <jnareb@gmail.com> writes:
> 
>>> Correct, but is "where does it appear" the question we are
>>> primarily interested in, wrt this breakage and its fix?
>>
>> That of course depends on how we want to test gitweb output.
>> The simplest solution, comparing with known output with perhaps
>> fragile / variable elements masked out could be done quickly...
>> but changes in output (even if they don't change functionality,
>> or don't change visible output) require regenerating test cases
>> (expected output) to test against - which might be source of
>> errors in test suite.
> 
> I agree with your "to test it fully, we need extra dependencies",
> but my point is that it does not have to be a full "HTML-validating,
> picking the expected attribute via XPATH matching" kind of test if
> what we want is only to add a new test to protect this particular
> fix from future breakages.
> 
> For example, I think it is sufficient to grep for 'href="...%xx%xx"'
> in the output after preparing a sample tree with one entry to show.
> The expected substring either exists (in which case we got it
> right), or it doesn't (in which case we are showing garbage).  Of
> course that depends on the assumption that its output is not too
> heavily contaminated with volatile parts outside our control, as I
> already mentioned in the message you are responding to.
> 
> But it all depends on "if" we wanted to add a new test ;-)

I tried to add such simple test to t9502, but instead of tests failing with current version, the test setup fails but succeeds (i.e. test library says that it failed, but manual examination shows that everything is O.K.).

-- >8 --
From: Jakub Narebski <jnareb@gmail.com>
Subject: [PATCH/RFC] gitweb test: Test proper encoding of non US-ASCII filenames in output (WIP)

This t9502 test is intended to test for proper encoding of non US-ASCII filenames (i.e. UTF-8 filenames) in generated links (which need some form of URI encoding) and in generated HTML (which needs HTML encoding / escaping).

For now it tests only 'tree' view (though incidentally it also tests UTF-8 in commit subject), as this was the action where reportedly there was bug in link encoding: $t{'name'} coming from the "git ls-tree -z ..." command via @ntries array was not marked as UTF-8, making Perl assume that it is in internal Perl format i.e. iso-8859-1 encoding and URI-escaping it as if it was in iso-8859-1 encoding (e.g. "Gütekriterien.txt" in UTF-8 is "Gütekriterien.txt" if treated as iso-8859-1, and it then encodes to "G%C3%83%C2%BCtekriterien.txt" instead of correct "G%C3%BCtekriterien.txt").

UNFORTUNATELY test does not fail as it should, even though the issue was not fixed... OTOH it fails in setup though it is successful.

Reported-by: Michael Wagner <accounts@mwagner.org>
Signed-off-by: Jakub Narębski <jnareb@gmail.com>
---
 t/t9502-gitweb-standalone-parse-output.sh |   34 +++++++++++++++++++++++++++++
 1 files changed, 34 insertions(+), 0 deletions(-)
Show changes to t/t9502-gitweb-standalone-parse-output.sh +34 −0
diff --git a/t/t9502-gitweb-standalone-parse-output.sh b/t/t9502-gitweb-standalone-parse-output.sh
index 86dfee2..37246a3 100755
--- a/t/t9502-gitweb-standalone-parse-output.sh
+++ b/t/t9502-gitweb-standalone-parse-output.sh
@@ -201,4 +201,38 @@ test_expect_success 'xss checks' '
 	xss "a=rss&p=foo.git&f=$TAG"
 '
 
+link_check () {
+	grep -F   "%3C__%C2%A3%C3%A5%C3%AB%C3%AE%C3%B1%C3%B2%C3%BB%C3%BD%C2%B6" \
+		gitweb.body &&
+	! grep -F "%3C__%A3%E5%EB%EE%F1%F2%FB%FD%B6" \
+		gitweb.body
+}
+
+test_expect_success 'prepare UTF-8 output tests' '
+	FILENAME="<__£åëîñòûý¶  +;?&__>" &&
+	test_commit "Adding $FILENAME" "$FILENAME" "$FILENAME contents"
+'
+
+test_expect_success 'check URI-escaped UTF-8 filename in query-params link' '
+	cat >>gitweb_config.perl <<-\EOF &&
+	$feature{"pathinfo"}{"default"} = [0];
+	EOF
+	gitweb_run "p=.git;a=tree" &&
+	link_check
+'
+
+test_expect_success 'check URI-escaped UTF-8 filename in path_info link' '
+	cat >>gitweb_config.perl <<-\EOF &&
+	$feature{"pathinfo"}{"default"} = [1];
+	EOF
+	gitweb_run "" "/.git/tree" &&
+	link_check
+'
+
+test_expect_success 'check HTML-escaped UTF-8 filename in body' '
+	gitweb_run "p=.git;a=tree" &&
+	grep -F "&lt;__£åëîñòûý¶  +;?&amp;__&gt;" gitweb.body &&
+	! grep -F  "<__£åëîñòûý¶  +;?&__>" gitweb.body
+'
+
 test_done
-- 
1.7.1


 
Jakub Narębski· May 27, 2014, 14:22 UTC · re: Jakub Narębski · lore

[PATCH] gitweb: Harden UTF-8 handling in generated links

W dniu 2014-05-15 21:28, Jakub Narębski pisze:
> On Thu, May 15, 2014 at 8:48 PM, Michael Wagner <accounts@mwagner.org> wrote:
>> On Thu, May 15, 2014 at 10:04:24AM +0100, Peter Krefting wrote:
>>> Michael Wagner:
Show 27 quoted lines
>> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
>> index a9f57d6..f1414e1 100755
>> --- a/gitweb/gitweb.perl
>> +++ b/gitweb/gitweb.perl
>> @@ -7138,7 +7138,7 @@ sub git_tree {
>>          my @entries = ();
>>          {
>>                  local $/ = "\0";
>> -               open my $fd, "-|", git_cmd(), "ls-tree", '-z',
>> +               open my $fd, "-|encoding(UTF-8)", git_cmd(), "ls-tree", '-z',
>>                          ($show_sizes ? '-l' : ()), @extra_options, $hash
>>                          or die_error(500, "Open git-ls-tree failed");
> 
> Or put
> 
>                     binmode $fd, ':utf8';
> 
> like in the rest of the code.
> 
>>                  @entries = map { chomp; $_ } <$fd>;
>>
> 
> Even better solution would be to use
> 
>      use open IN => ':encoding(utf-8)';
> 
> at the beginning of gitweb.perl, once and for all.

Or harden esc_param / esc_path_info the same way esc_html is hardened against missing ':utf8' flag.

-- >8 -- 
Subject: [PATCH] gitweb: Harden UTF-8 handling in generated links

esc_html() ensures that its input is properly UTF-8 encoded and marked as UTF-8 with to_utf8(). Make esc_param() (used for query parameters in generated URLs), esc_path_info() (for escaping path_info components) and esc_url() use it too.

This hardens gitweb against errors in UTF-8 handling; because to_utf8() is idempotent it won't change correct output.

Reported-by: Michael Wagner <accounts@mwagner.org>
Signed-off-by: Jakub Narębski <jnareb@gmail.com>
---
 gitweb/gitweb.perl |    7 +++++++
 1 files changed, 7 insertions(+), 0 deletions(-)
Show changes to gitweb/gitweb.perl +7 −0
diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
index a9f57d6..77e1312 100755
--- a/gitweb/gitweb.perl
+++ b/gitweb/gitweb.perl
@@ -1548,8 +1548,11 @@ sub to_utf8 {
 sub esc_param {
 	my $str = shift;
 	return undef unless defined $str;
+
+	$str = to_utf8($str);
 	$str =~ s/([^A-Za-z0-9\-_.~()\/:@ ]+)/CGI::escape($1)/eg;
 	$str =~ s/ /\+/g;
+
 	return $str;
 }
 
@@ -1558,6 +1561,7 @@ sub esc_path_info {
 	my $str = shift;
 	return undef unless defined $str;
 
+	$str = to_utf8($str);
 	# path_info doesn't treat '+' as space (specially), but '?' must be escaped
 	$str =~ s/([^A-Za-z0-9\-_.~();\/;:@&= +]+)/CGI::escape($1)/eg;
 
@@ -1568,8 +1572,11 @@ sub esc_path_info {
 sub esc_url {
 	my $str = shift;
 	return undef unless defined $str;
+
+	$str = to_utf8($str);
 	$str =~ s/([^A-Za-z0-9\-_.~();\/;?:@&= ]+)/CGI::escape($1)/eg;
 	$str =~ s/ /\+/g;
+
 	return $str;
 }
 
-- 
1.7.1
Michael Wagner· Jun 4, 2014, 15:41 UTC · re: Jakub Narębski · lore

Re: [PATCH] gitweb: Harden UTF-8 handling in generated links

On Tue, May 27, 2014 at 04:22:42PM +0200, Jakub Narębski wrote:
Show 93 quoted lines
> W dniu 2014-05-15 21:28, Jakub Narębski pisze:
> > On Thu, May 15, 2014 at 8:48 PM, Michael Wagner <accounts@mwagner.org> wrote:
> >> On Thu, May 15, 2014 at 10:04:24AM +0100, Peter Krefting wrote:
> >>> Michael Wagner:
> 
> >> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
> >> index a9f57d6..f1414e1 100755
> >> --- a/gitweb/gitweb.perl
> >> +++ b/gitweb/gitweb.perl
> >> @@ -7138,7 +7138,7 @@ sub git_tree {
> >>          my @entries = ();
> >>          {
> >>                  local $/ = "\0";
> >> -               open my $fd, "-|", git_cmd(), "ls-tree", '-z',
> >> +               open my $fd, "-|encoding(UTF-8)", git_cmd(), "ls-tree", '-z',
> >>                          ($show_sizes ? '-l' : ()), @extra_options, $hash
> >>                          or die_error(500, "Open git-ls-tree failed");
> > 
> > Or put
> > 
> >                     binmode $fd, ':utf8';
> > 
> > like in the rest of the code.
> > 
> >>                  @entries = map { chomp; $_ } <$fd>;
> >>
> > 
> > Even better solution would be to use
> > 
> >      use open IN => ':encoding(utf-8)';
> > 
> > at the beginning of gitweb.perl, once and for all.
> 
> Or harden esc_param / esc_path_info the same way esc_html
> is hardened against missing ':utf8' flag.
> 
> -- >8 -- 
> Subject: [PATCH] gitweb: Harden UTF-8 handling in generated links
> 
> esc_html() ensures that its input is properly UTF-8 encoded and marked
> as UTF-8 with to_utf8().  Make esc_param() (used for query parameters
> in generated URLs), esc_path_info() (for escaping path_info
> components) and esc_url() use it too.
> 
> This hardens gitweb against errors in UTF-8 handling; because
> to_utf8() is idempotent it won't change correct output.
> 
> Reported-by: Michael Wagner <accounts@mwagner.org>
> Signed-off-by: Jakub Narębski <jnareb@gmail.com>
> ---
>  gitweb/gitweb.perl |    7 +++++++
>  1 files changed, 7 insertions(+), 0 deletions(-)
> 
> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
> index a9f57d6..77e1312 100755
> --- a/gitweb/gitweb.perl
> +++ b/gitweb/gitweb.perl
> @@ -1548,8 +1548,11 @@ sub to_utf8 {
>  sub esc_param {
>  	my $str = shift;
>  	return undef unless defined $str;
> +
> +	$str = to_utf8($str);
>  	$str =~ s/([^A-Za-z0-9\-_.~()\/:@ ]+)/CGI::escape($1)/eg;
>  	$str =~ s/ /\+/g;
> +
>  	return $str;
>  }
>  
> @@ -1558,6 +1561,7 @@ sub esc_path_info {
>  	my $str = shift;
>  	return undef unless defined $str;
>  
> +	$str = to_utf8($str);
>  	# path_info doesn't treat '+' as space (specially), but '?' must be escaped
>  	$str =~ s/([^A-Za-z0-9\-_.~();\/;:@&= +]+)/CGI::escape($1)/eg;
>  
> @@ -1568,8 +1572,11 @@ sub esc_path_info {
>  sub esc_url {
>  	my $str = shift;
>  	return undef unless defined $str;
> +
> +	$str = to_utf8($str);
>  	$str =~ s/([^A-Za-z0-9\-_.~();\/;?:@&= ]+)/CGI::escape($1)/eg;
>  	$str =~ s/ /\+/g;
> +
>  	return $str;
>  }
>  
> -- 
> 1.7.1
> 
> 

While trying to view a "blob_plain" of "Gütekritierien.txt", a 404 error occured. "git_get_hash_by_path" tries to resolve the hash with the wrong filename (git ls-tree -z HEAD -- Gütekriterien.txt) and fails.

The filename needs the correct encoding. Something like this is probably
needed for all filenames and should be done at a prior stage:
---
 gitweb/gitweb.perl |    2 +-
 1 files changed, 1 insertions(+), 1 deletions(-)
Show changes to gitweb/gitweb.perl +1 −1
diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
index 77e1312..e4a50e7 100755
--- a/gitweb/gitweb.perl
+++ b/gitweb/gitweb.perl
@@ -4725,7 +4725,7 @@ sub git_print_tree_entry {
                }
                print " | " .
                        $cgi->a({-href => href(action=>"blob_plain", hash_base=>$hash_base,
-                                              file_name=>"$basedir$t->{'name'}")},
+                                              file_name=>"$basedir" . to_utf8($t->{'name'}))}, 
                                "raw");
                print "</td>\n";
-- 
1.7.1
Jakub Narębski· Jun 4, 2014, 18:47 UTC · re: Michael Wagner · lore

Re: [PATCH] gitweb: Harden UTF-8 handling in generated links

Michael Wagner wrote:
> On Tue, May 27, 2014 at 04:22:42PM +0200, Jakub Narębski wrote:
Show 9 quoted lines
>> Subject: [PATCH] gitweb: Harden UTF-8 handling in generated links
>>
>> esc_html() ensures that its input is properly UTF-8 encoded and marked
>> as UTF-8 with to_utf8().  Make esc_param() (used for query parameters
>> in generated URLs), esc_path_info() (for escaping path_info
>> components) and esc_url() use it too.
>>
>> This hardens gitweb against errors in UTF-8 handling; because
>> to_utf8() is idempotent it won't change correct output.
[...]
Show 10 quoted lines
>>   sub esc_param {
>>   	my $str = shift;
>>   	return undef unless defined $str;
>> +
>> +	$str = to_utf8($str);
>>   	$str =~ s/([^A-Za-z0-9\-_.~()\/:@ ]+)/CGI::escape($1)/eg;
>>   	$str =~ s/ /\+/g;
>> +
>>   	return $str;
>>   }   
Show 6 quoted lines
> While trying to view a "blob_plain" of "Gütekritierien.txt", a 404 error
> occured. "git_get_hash_by_path" tries to resolve the hash with the wrong
> filename (git ls-tree -z HEAD -- Gütekriterien.txt) and fails.
> 
> The filename needs the correct encoding. Something like this is probably
> needed for all filenames and should be done at a prior stage:
True.

First, I wonder why the tests I did for this situation didn't show any errors even before the "harden href()" patch. What is different in your config that you see those errors?

Show 14 quoted lines
> ---
>   gitweb/gitweb.perl |    2 +-
>   1 files changed, 1 insertions(+), 1 deletions(-)
> 
> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
> index 77e1312..e4a50e7 100755
> --- a/gitweb/gitweb.perl
> +++ b/gitweb/gitweb.perl
> @@ -4725,7 +4725,7 @@ sub git_print_tree_entry {
>                  }
>                  print " | " .
>                          $cgi->a({-href => href(action=>"blob_plain", hash_base=>$hash_base,
> -                                              file_name=>"$basedir$t->{'name'}")},
> +                                              file_name=>"$basedir" . to_utf8($t->{'name'}))},

Second, my "harder href()" patch does not work for this because concatenation of non-UFT8 with UTF8 string screws up Perl knowledge what is and isn't UTF8. So to_utf8() after concat doesn't help.

>                                  "raw");
>                  print "</td>\n";
> 
Michael Wagner· Jun 4, 2014, 20:47 UTC · re: Jakub Narębski · lore

Re: [PATCH] gitweb: Harden UTF-8 handling in generated links

On Wed, Jun 04, 2014 at 08:47:54PM +0200, Jakub Narębski wrote:
Show 37 quoted lines
> Michael Wagner wrote:
> > On Tue, May 27, 2014 at 04:22:42PM +0200, Jakub Narębski wrote:
> 
> >> Subject: [PATCH] gitweb: Harden UTF-8 handling in generated links
> >>
> >> esc_html() ensures that its input is properly UTF-8 encoded and marked
> >> as UTF-8 with to_utf8().  Make esc_param() (used for query parameters
> >> in generated URLs), esc_path_info() (for escaping path_info
> >> components) and esc_url() use it too.
> >>
> >> This hardens gitweb against errors in UTF-8 handling; because
> >> to_utf8() is idempotent it won't change correct output.
> [...]
> >>   sub esc_param {
> >>   	my $str = shift;
> >>   	return undef unless defined $str;
> >> +
> >> +	$str = to_utf8($str);
> >>   	$str =~ s/([^A-Za-z0-9\-_.~()\/:@ ]+)/CGI::escape($1)/eg;
> >>   	$str =~ s/ /\+/g;
> >> +
> >>   	return $str;
> >>   }   
>  
> > While trying to view a "blob_plain" of "Gütekritierien.txt", a 404 error
> > occured. "git_get_hash_by_path" tries to resolve the hash with the wrong
> > filename (git ls-tree -z HEAD -- Gütekriterien.txt) and fails.
> > 
> > The filename needs the correct encoding. Something like this is probably
> > needed for all filenames and should be done at a prior stage:
> 
> True.
> 
> First, I wonder why the tests I did for this situation didn't
> show any errors even before the "harden href()" patch. What
> is different in your config that you see those errors?
> 
Nothing special. It is reproducible with git 1.9.3 (Fedora 20), git
instaweb (lighttpd) and LANG=de_DE.UTF-8.  
 

← back to recent threads