{"thread":{"id":"18920","subject":"[PATCH] gitweb: filter escapes from longer commit titles that break firefox","startedAt":"2009-04-17T16:24:33Z","lastAt":"2009-04-25T09:04:42Z","messageCount":7,"participants":["Paul Gortmaker","Jakub Narebski"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"111503","messageId":"1239985473-666-1-git-send-email-paul.gortmaker@windriver.com","threadId":"18920","inReplyTo":null,"subject":"[PATCH] gitweb: filter escapes from longer commit titles that break firefox","fromName":"Paul Gortmaker","fromEmail":"paul.gortmaker@windriver.com","sentAt":"2009-04-17T16:24:33Z","receivedAt":"2009-04-17T16:24:33Z","isPatch":true,"sender":{"key":"paul.gortmaker@windriver.com","avatar":null},"body":"If there is a commit that ends in ^X and is longer in length than\nwhat will fit in title_short, then it doesn't get fed through\nesc_html() and so the ^X will appear as-is in the page source.\n\nWhen Firefox comes across this, it will fail to display the page,\nand only display a couple lines of error messages that read like:\n\n   XML Parsing Error: not well-formed\n   Location: http://git ....\n\nSigned-off-by: Paul Gortmaker <paul.gortmaker@windriver.com>\n---\n gitweb/gitweb.perl |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex 33ef190..e686e82 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -2470,7 +2470,7 @@ sub parse_commit_text {\n \tforeach my $title (@commit_lines) {\n \t\t$title =~ s/^    //;\n \t\tif ($title ne \"\") {\n-\t\t\t$co{'title'} = chop_str($title, 80, 5);\n+\t\t\t$co{'title'} = chop_and_escape_str($title, 80, 5);\n \t\t\t# remove leading stuff of merges to make the interesting part visible\n \t\t\tif (length($title) > 50) {\n \t\t\t\t$title =~ s/^Automatic //;\n-- \n1.6.2.3\n"},{"id":"111700","messageId":"m3r5znpt5g.fsf@localhost.localdomain","threadId":"18920","inReplyTo":"1239985473-666-1-git-send-email-paul.gortmaker@windriver.com","subject":"Re: [PATCH] gitweb: filter escapes from longer commit titles that break firefox","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-04-20T09:32:35Z","receivedAt":"2009-04-20T09:32:35Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Paul Gortmaker <paul.gortmaker@windriver.com> writes:\n\n> If there is a commit that ends in ^X and is longer in length than\n> what will fit in title_short, then it doesn't get fed through\n> esc_html() and so the ^X will appear as-is in the page source.\n> \n> When Firefox comes across this, it will fail to display the page,\n> and only display a couple lines of error messages that read like:\n> \n>    XML Parsing Error: not well-formed\n>    Location: http://git ....\n> \n> Signed-off-by: Paul Gortmaker <paul.gortmaker@windriver.com>\n\nThis is an issue for when project doesn't follow sanity (control\ncharacters in commit message) nor commit message conventions of git\n(limiting length of first line of commit message to 60-70 characters).\n\nBut I do not think that the solution presented here is good solution\nfor this problem.  chop_and_escape_str is meant as _output_ filter,\nbecause it generates (can generate) fragment of HTML.  It is not a\ngood solution to use it for shortening in intermediate representation\nof %co{'title'}.\n\nAnd I think that issue might be a bug elsewhere in gitweb if we have\ntext output which is not passed through esc_html... or bug in CGI.pm\nif the error is in not escaping of -title _attribute_ (attribute\nescaping has slightly different rules than escaping HTML, and should\nbe done automatically by CGI.pm).\n\n\nSo thanks for noticing the issue, but NAK on the solution.\n\n> ---\n>  gitweb/gitweb.perl |    2 +-\n>  1 files changed, 1 insertions(+), 1 deletions(-)\n> \n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index 33ef190..e686e82 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -2470,7 +2470,7 @@ sub parse_commit_text {\n>  \tforeach my $title (@commit_lines) {\n>  \t\t$title =~ s/^    //;\n>  \t\tif ($title ne \"\") {\n> -\t\t\t$co{'title'} = chop_str($title, 80, 5);\n> +\t\t\t$co{'title'} = chop_and_escape_str($title, 80, 5);\n>  \t\t\t# remove leading stuff of merges to make the interesting part visible\n>  \t\t\tif (length($title) > 50) {\n>  \t\t\t\t$title =~ s/^Automatic //;\n> -- \n> 1.6.2.3\n> \n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"111730","messageId":"49EC78AB.6020009@windriver.com","threadId":"18920","inReplyTo":"m3r5znpt5g.fsf@localhost.localdomain","subject":"Re: [PATCH] gitweb: filter escapes from longer commit titles that break firefox","fromName":"Paul Gortmaker","fromEmail":"paul.gortmaker@windriver.com","sentAt":"2009-04-20T13:29:15Z","receivedAt":"2009-04-20T13:29:15Z","isPatch":true,"sender":{"key":"paul.gortmaker@windriver.com","avatar":null},"body":"Jakub Narebski wrote:\n> Paul Gortmaker <paul.gortmaker@windriver.com> writes:\n>\n>   \n>> If there is a commit that ends in ^X and is longer in length than\n>> what will fit in title_short, then it doesn't get fed through\n>> esc_html() and so the ^X will appear as-is in the page source.\n>>\n>> When Firefox comes across this, it will fail to display the page,\n>> and only display a couple lines of error messages that read like:\n>>\n>>    XML Parsing Error: not well-formed\n>>    Location: http://git ....\n>>\n>> Signed-off-by: Paul Gortmaker <paul.gortmaker@windriver.com>\n>>     \n>\n> This is an issue for when project doesn't follow sanity (control\n> characters in commit message) nor commit message conventions of git\n> (limiting length of first line of commit message to 60-70 characters).\n>   \n\nI agree - the situation should be that it doesn't happen, but it can \nhappen (and it did\nhappen) that a novice, or a simple mistake ends up with such a commit. \n\n> But I do not think that the solution presented here is good solution\n> for this problem.  chop_and_escape_str is meant as _output_ filter,\n> because it generates (can generate) fragment of HTML.  It is not a\n> good solution to use it for shortening in intermediate representation\n> of %co{'title'}.\n>\n> And I think that issue might be a bug elsewhere in gitweb if we have\n> text output which is not passed through esc_html... or bug in CGI.pm\n> if the error is in not escaping of -title _attribute_ (attribute\n> escaping has slightly different rules than escaping HTML, and should\n> be done automatically by CGI.pm).\n>\n>\n> So thanks for noticing the issue, but NAK on the solution.\n>   \n\nFair enough -- I wasn't familiar with the code in there, and there \nwasn't really any indication that it was for output only.  I can easily \nbelieve that there is a better place for it -- I just didn't see where \nany global esc_html filtering was taking place...\n\nPaul.\n\n>   \n>> ---\n>>  gitweb/gitweb.perl |    2 +-\n>>  1 files changed, 1 insertions(+), 1 deletions(-)\n>>\n>> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n>> index 33ef190..e686e82 100755\n>> --- a/gitweb/gitweb.perl\n>> +++ b/gitweb/gitweb.perl\n>> @@ -2470,7 +2470,7 @@ sub parse_commit_text {\n>>  \tforeach my $title (@commit_lines) {\n>>  \t\t$title =~ s/^    //;\n>>  \t\tif ($title ne \"\") {\n>> -\t\t\t$co{'title'} = chop_str($title, 80, 5);\n>> +\t\t\t$co{'title'} = chop_and_escape_str($title, 80, 5);\n>>  \t\t\t# remove leading stuff of merges to make the interesting part visible\n>>  \t\t\tif (length($title) > 50) {\n>>  \t\t\t\t$title =~ s/^Automatic //;\n>> -- \n>> 1.6.2.3\n>>\n>>     \n>\n>   \n"},{"id":"112222","messageId":"200904241953.37187.jnareb@gmail.com","threadId":"18920","inReplyTo":"49EC78AB.6020009@windriver.com","subject":"Re: [PATCH] gitweb: filter escapes from longer commit titles that break firefox","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-04-24T17:53:35Z","receivedAt":"2009-04-24T17:53:35Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Mon, 20 April 2009, Paul Gortmaker wrote:\n> Jakub Narebski wrote:\n>> Paul Gortmaker <paul.gortmaker@windriver.com> writes:\n>>\n>>   \n>>> If there is a commit that ends in ^X and is longer in length than\n>>> what will fit in title_short, then it doesn't get fed through\n>>> esc_html() and so the ^X will appear as-is in the page source.\n>>>\n>>> When Firefox comes across this, it will fail to display the page,\n>>> and only display a couple lines of error messages that read like:\n>>>\n>>>    XML Parsing Error: not well-formed\n>>>    Location: http://git ....\n>>>\n>>> Signed-off-by: Paul Gortmaker <paul.gortmaker@windriver.com>\n\n>> But I do not think that the solution presented here is good solution\n>> for this problem.  chop_and_escape_str is meant as _output_ filter,\n>> because it generates (can generate) fragment of HTML.  It is not a\n>> good solution to use it for shortening in intermediate representation\n>> of %co{'title'}.\n>>\n>> And I think that issue might be a bug elsewhere in gitweb if we have\n>> text output which is not passed through esc_html... or bug in CGI.pm\n>> if the error is in not escaping of -title _attribute_ (attribute\n>> escaping has slightly different rules than escaping HTML, and should\n>> be done automatically by CGI.pm).\n>>\n>>\n>> So thanks for noticing the issue, but NAK on the solution.\n> \n> Fair enough -- I wasn't familiar with the code in there, and there \n> wasn't really any indication that it was for output only.  I can easily \n> believe that there is a better place for it -- I just didn't see where \n> any global esc_html filtering was taking place...\n\nThe name chop_and_escape_str for this subroutine is not a very good\nname; it rather should follow format_* as a naming convention for this\nsubroutine.\n\nWhat more important is: can you find out in more detail _where_\nan error (unescaped control character) occurs: is it tag contents or\n'title' attribute for some tag, what tag is it (name and class), in\nwhat view or views this bug is present, and in which part this occur?\nWithout those details it would b much harder to diagnose this bug...\n-- \nJakub Narebski\nPoland\n"},{"id":"112230","messageId":"20090424194822.GA15846@windriver.com","threadId":"18920","inReplyTo":"200904241953.37187.jnareb@gmail.com","subject":"Re: [PATCH] gitweb: filter escapes from longer commit titles that break firefox","fromName":"Paul Gortmaker","fromEmail":"paul.gortmaker@windriver.com","sentAt":"2009-04-24T19:48:22Z","receivedAt":"2009-04-24T19:48:22Z","isPatch":true,"sender":{"key":"paul.gortmaker@windriver.com","avatar":null},"body":"[Re: [PATCH] gitweb: filter escapes from longer commit titles that break firefox] On 24/04/2009 (Fri 19:53) Jakub Narebski wrote:\n\n> On Mon, 20 April 2009, Paul Gortmaker wrote:\n> > Jakub Narebski wrote:\n> >> Paul Gortmaker <paul.gortmaker@windriver.com> writes:\n> >>\n> >>   \n> >>> If there is a commit that ends in ^X and is longer in length than\n> >>> what will fit in title_short, then it doesn't get fed through\n> >>> esc_html() and so the ^X will appear as-is in the page source.\n> >>>\n> >>> When Firefox comes across this, it will fail to display the page,\n> >>> and only display a couple lines of error messages that read like:\n> >>>\n> >>>    XML Parsing Error: not well-formed\n> >>>    Location: http://git ....\n> >>>\n> >>> Signed-off-by: Paul Gortmaker <paul.gortmaker@windriver.com>\n> \n> >> But I do not think that the solution presented here is good solution\n> >> for this problem.  chop_and_escape_str is meant as _output_ filter,\n> >> because it generates (can generate) fragment of HTML.  It is not a\n> >> good solution to use it for shortening in intermediate representation\n> >> of %co{'title'}.\n> >>\n> >> And I think that issue might be a bug elsewhere in gitweb if we have\n> >> text output which is not passed through esc_html... or bug in CGI.pm\n> >> if the error is in not escaping of -title _attribute_ (attribute\n> >> escaping has slightly different rules than escaping HTML, and should\n> >> be done automatically by CGI.pm).\n> >>\n> >>\n> >> So thanks for noticing the issue, but NAK on the solution.\n> > \n> > Fair enough -- I wasn't familiar with the code in there, and there \n> > wasn't really any indication that it was for output only.  I can easily \n> > believe that there is a better place for it -- I just didn't see where \n> > any global esc_html filtering was taking place...\n> \n> The name chop_and_escape_str for this subroutine is not a very good\n> name; it rather should follow format_* as a naming convention for this\n> subroutine.\n> \n> What more important is: can you find out in more detail _where_\n> an error (unescaped control character) occurs: is it tag contents or\n> 'title' attribute for some tag, what tag is it (name and class), in\n> what view or views this bug is present, and in which part this occur?\n> Without those details it would b much harder to diagnose this bug...\n\nNo problem -- It appears to be in the title attribute, and it appears\nstraight away when I go to the toplevel view of the repo, assuming that\nthe commit is within the top 10 recent commits that are shown on the\nsummary page.  I've put more details below on how I can reproduce it\nand the page source deltas -- hopefully this will help.  If there is\nsomething else I can provide that would help, don't hesitate to ask.\n\nPaul.\n\n-------\n\nSetup:\n\nyow-d4:test$mkdir bad_commit\nyow-d4:test$cd bad_commit/\nyow-d4:bad_commit$git init\nInitialized empty Git repository in /home/pgortmak/test/bad_commit/.git/\nyow-d4:bad_commit$echo bbb > bbb\nyow-d4:bad_commit$git add bbb\nyow-d4:bad_commit$git commit -m 'some string that is longer than roughly\n50chars, with a ^X embedded at the end^X'\n[master (root-commit) 8735814] some string that is longer than roughly\n50chars, with a ^X embedded at the end\n 1 files changed, 1 insertions(+), 0 deletions(-)\n create mode 100644 bbb\nyow-d4:bad_commit$\n--------\n\nI've used ^V^X to embed the ^X at the end of the commit message above;\nthe other \"^X\" is just literally a ^ followed by a X.\n\nThen I load it with firefox (default shipping with Ubuntu Jaunty).\nWith my workaround patch sent previously, it renders OK, and I save\nthe source to \"source-ok\".  Then I take out my hack patch, and it fails\nto render, instead giving:\n\n------\nXML Parsing Error: not well-formed\nLocation:\nhttp://yow-somehost.com/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=summary\nLine Number 54, Column 114:<td><a class=\"list subject\" title=\"some\nstring that is longer than roughly 50chars, with a ^X embedded at the\nend\"\nhref=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=commit;h=8735814a15cf930c48fd33563f5922a103b6b4ea\">some string that is longer than roughly 50chars, with...  <span class=\"refs\"> <span class=\"head\" title=\"heads/master\"><a href=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=shortlog;h=refs/heads/master\">master</a></span></span></a></td>\n-----------------------------------------------------------------------------------------------------------------^\n\nYou probably can't tell in mail, but the line of dashes and the caret\nare pointing at firefox's rendering of the ^X at EOL, which is displayed\nas a little square box with a 00 above an 18 inside it.\n\nIf I save this page to source-bad, and then diff the two, I get:\n\n--- source-ok\t2009-04-24 15:29:11.000000000 -0400\n+++ source-bad\t2009-04-24 15:30:20.000000000 -0400\n@@ -54,9 +54,9 @@\n </div>\n <table class=\"shortlog\">\n <tr class=\"dark\">\n-<td title=\"2009-04-24\"><i>2 min ago</i></td>\n+<td title=\"2009-04-24\"><i>6 min ago</i></td>\n <td><i>Paul Gortmaker</i></td>\n-<td><a class=\"list subject\" title=\"some string that is longer than roughly 50chars, with a ^X embedded at the end&lt;span class=&quot;cntrl&quot;&gt;\\18&lt;/span&gt;\" href=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=commit;h=8735814a15cf930c48fd33563f5922a103b6b4ea\">some string that is longer than roughly 50chars, with...  <span class=\"refs\"> <span class=\"head\" title=\"heads/master\"><a href=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=shortlog;h=refs/heads/master\">master</a></span></span></a></td>\n+<td><a class=\"list subject\" title=\"some string that is longer than roughly 50chars, with a ^X embedded at the end\" href=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=commit;h=8735814a15cf930c48fd33563f5922a103b6b4ea\">some string that is longer than roughly 50chars, with...  <span class=\"refs\"> <span class=\"head\" title=\"heads/master\"><a href=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=shortlog;h=refs/heads/master\">master</a></span></span></a></td>\n <td class=\"link\"><a href=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=commit;h=8735814a15cf930c48fd33563f5922a103b6b4ea\">commit</a> | <a href=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=commitdiff;h=8735814a15cf930c48fd33563f5922a103b6b4ea\">commitdiff</a> | <a href=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=tree;h=8735814a15cf930c48fd33563f5922a103b6b4ea;hb=8735814a15cf930c48fd33563f5922a103b6b4ea\">tree</a> | <a title=\"in format: tar.gz\" href=\"/gitweb/gitweb.cgi?p=local/pgortmak/test/bad_commit/.git;a=snapshot;h=8735814a15cf930c48fd33563f5922a103b6b4ea;sf=tgz\">snapshot</a></td>\n \n </tr>\n"},{"id":"112246","messageId":"200904250010.46299.jnareb@gmail.com","threadId":"18920","inReplyTo":"20090424194822.GA15846@windriver.com","subject":"Re: [PATCH] gitweb: filter escapes from longer commit titles that break firefox","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-04-24T22:10:44Z","receivedAt":"2009-04-24T22:10:44Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Fri, 24 April 2009, Paul Gortmaker wrote:\n> [Re: [PATCH] gitweb: filter escapes from longer commit titles that break firefox]\n> On 24/04/2009 (Fri 19:53) Jakub Narebski wrote:  \n>> On Mon, 20 April 2009, Paul Gortmaker wrote:\n>>> Jakub Narebski wrote:\n>>>> Paul Gortmaker <paul.gortmaker@windriver.com> writes:\n>>>>\n>>>>   \n>>>>> If there is a commit that ends in ^X and is longer in length than\n>>>>> what will fit in title_short, then it doesn't get fed through\n>>>>> esc_html() and so the ^X will appear as-is in the page source.\n>>>>>\n>>>>> When Firefox comes across this, it will fail to display the page,\n>>>>> and only display a couple lines of error messages that read like:\n>>>>>\n>>>>>    XML Parsing Error: not well-formed\n>>>>>    Location: http://git ....\n>>>>>\n>>>>> Signed-off-by: Paul Gortmaker <paul.gortmaker@windriver.com>\n\n>>>> And I think that issue might be a bug elsewhere in gitweb if we have\n>>>> text output which is not passed through esc_html... or bug in CGI.pm\n>>>> if the error is in not escaping of -title _attribute_ (attribute\n>>>> escaping has slightly different rules than escaping HTML, and should\n>>>> be done automatically by CGI.pm).\n\n>> What more important is: can you find out in more detail _where_\n>> an error (unescaped control character) occurs: is it tag contents or\n>> 'title' attribute for some tag, what tag is it (name and class), in\n>> what view or views this bug is present, and in which part this occur?\n>> Without those details it would b much harder to diagnose this bug...\n> \n> No problem -- It appears to be in the title attribute, and it appears\n> straight away when I go to the toplevel view of the repo, assuming that\n> the commit is within the top 10 recent commits that are shown on the\n> summary page.  I've put more details below on how I can reproduce it\n> and the page source deltas -- hopefully this will help.  If there is\n> something else I can provide that would help, don't hesitate to ask.\n\nAhh... that is what I thought.\n\nThe problem that we have to solve to fix this bug is twofold:\n\n * CGI.pm does by default slight escaping (simple_escape from CGI::Util)\n   of _attribute_ values, but for obvious reasons it cannot do\n   unconditional escaping of tag _contents_ (because it can be HTML\n   itself).\n\n   This escaping, at least in CGI.pm version 3.10 (most current version\n   at CPAN is 3.43), is minimal: only '\"', '&', '<' and '>' are escaped\n   using named HTML entity references (&quot;, &amp;, &lt; and &gt;\n   respectively).  simple_escape does not do escaping of control\n   characters such as ^X which are invalid in XHTML (in strict mode).\n   Note that IIRC escaping '<' and '>' in attributes is not strictly\n   necessary.\n\n   Gitweb relies on the fact that CGI.pm does escaping of attribute\n   values.  We cannot escape attributes (e.g. \"title\" attribute with\n   (almost) full commit subject) as it is now, because it would lead\n   to double escaping.  Fortunately it is possible to turn off\n   autoescaping by using $cgi->autoEscape(undef); note however that\n   we would have to do attribute escaping by ourself in the scope of\n   this declaration.\n\n * Rules for escaping attribute values are slightly different for rules\n   for escaping HTML.  For attribute values we have to escape '\"'\n   because it is attribute delimiter, and '&' because it is escape\n   character; escaping '<' and '>' is not strictly necessary.  For\n   escaping HTML we need to escape '<' and '>' because they introduce\n   tags, and '&' because it is escape character; escaping '\"' is not\n   strictly necessary.  It does not make sense to replace spaces by\n   &nbsp; in attribute values, although it shouldn't harm.  OTOH we\n   should perhaps escape newlines in attribute values.\n\n   For esc_html and esc_path we replace (currently) control characters\n   by character escape codes (e.g. \"\\f\" for form-feed, \"\\0\" for NUL,\n   hexadecimal escapes for 'other' control characters).  But it is not\n   the only possible solution.  We can use Unicode printable\n   representation of control characters instead (0x2400 sheet).  Or we\n   can use control key sequence / caret notation e.g. ^X for \\0x18,\n   or ^L for \"\\f\" there.  We probably should discus this in more detail.\n\nSo it is not that simple...\n\n\nP.S. The subject (one line summary of this change) should be also\nchanged to for example \"gitweb: escape control characters in attributes\"\nand in commit message itself you should explain that control characters\nbreak rendering in Firefox in strict XML compliance mode... or something\nlike that.\n-- \nJakub Narebski\nPoland\n"},{"id":"112285","messageId":"200904251104.45236.jnareb@gmail.com","threadId":"18920","inReplyTo":"200904250010.46299.jnareb@gmail.com","subject":"Re: [PATCH] gitweb: filter escapes from longer commit titles that break firefox","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-04-25T09:04:42Z","receivedAt":"2009-04-25T09:04:42Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, 25 April 2009, Jakub Narebski wrote:\n\n> So it is not that simple...\n\nThat said, here is simple patch which should fix the bug you found.\nIt always creates sensible short and long values, contrary to your patch\n(take a look at gitweb output after your patch, including tooltips on\nmouseover).\n\nBut it is NOT TESTED if it works correctly, and if it covers all\noccurrences.  And it might be not necessary in all its complication:\nwe could simply replace control characters by '?' like in\nchop_and_escape_str subroutine (which would also make gitweb more\nconsistent).  It also lacks commit message.\n\nNevertheless it might be good bandaid for your problem:\n\n-- >8 --\ndiff --git c/gitweb/gitweb.perl w/gitweb/gitweb.perl\nindex 3f99361..8575d5f 100755\n--- c/gitweb/gitweb.perl\n+++ w/gitweb/gitweb.perl\n@@ -1035,6 +1035,24 @@ sub esc_url {\n \treturn $str;\n }\n \n+# quote and escape tag attribute values; autoEscape has to be turned off\n+sub esc_attr {\n+\tmy $str = shift;\n+\treturn $str unless defined $str;\n+\n+\tmy %ent = ( # named HTML entities\n+\t\t'\"' => '&quot;',\n+\t\t'&' => '&amp;',\n+\t\t'<' => '&lt;',\n+\t\t'>' => '&gt;',\n+\t);\n+\t$str = to_utf8($str);\n+\t$str =~ s|([\\\"&<>])|$ent{$1}|eg;\n+\t$str =~ s|([[:cntrl:]])|(($1 ne \"\\t\") ? quot_upr($1) : $1)|eg;\n+\n+\treturn $str;\n+}\n+\n # replace invalid utf8 character with SUBSTITUTION sequence\n sub esc_html ($;%) {\n \tmy $str = shift;\n@@ -1457,14 +1475,19 @@ sub format_subject_html {\n \tmy ($long, $short, $href, $extra) = @_;\n \t$extra = '' unless defined($extra);\n \n+\tmy $ret = '';\n \tif (length($short) < length($long)) {\n-\t\treturn $cgi->a({-href => $href, -class => \"list subject\",\n-\t\t                -title => to_utf8($long)},\n+\t\tmy $autoescape = $cgi->autoEscape(undef);\n+\t\t# or just replace s/([[:cntrl:]])/?/g in -title\n+\t\t$ret = $cgi->a({-href => $href, -class => \"list subject\",\n+\t\t                -title => esc_attr($long)},\n \t\t       esc_html($short) . $extra);\n+\t\t$cgi->autoEscape($autoescape); # restore original value\n \t} else {\n-\t\treturn $cgi->a({-href => $href, -class => \"list subject\"},\n+\t\t$ret = $cgi->a({-href => $href, -class => \"list subject\"},\n \t\t       esc_html($long)  . $extra);\n \t}\n+\treturn $ret;\n }\n \n # format git diff header line, i.e. \"diff --(git|combined|cc) ...\"\n"}]}