{"thread":{"id":"12075","subject":"[PATCH] Do not chop HTML tags in commit search result","startedAt":"2008-02-13T17:37:24Z","lastAt":"2008-02-25T20:18:54Z","messageCount":16,"participants":["Jean-Baptiste Quenot","Jakub Narebski","Junio C Hamano","Karl Hasselström"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"68660","messageId":"ae63f8b50802130937mddf9df9re2a95bee44661ee3@mail.gmail.com","threadId":"12075","inReplyTo":null,"subject":"[PATCH] Do not chop HTML tags in commit search result","fromName":"Jean-Baptiste Quenot","fromEmail":"jbq@caraldi.com","sentAt":"2008-02-13T17:37:24Z","receivedAt":"2008-02-13T17:37:24Z","isPatch":true,"sender":{"key":"jbq@caraldi.com","avatar":null},"body":"Hi there,\n\nThanks for Git! It's a great program.  I encountered an annoying bug\nwith gitweb 1.5.4.1, when searching for commits, if the search string\nis too long, the generated HTML is munged leading to an ill-formed\nXHTML document.\n\nHere is the patch, hope it helps:\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex ae2d057..2c0b990 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -3780,7 +3780,10 @@ sub git_search_grep_body {\n                                my $trail = esc_html($3) || \"\";\n                                $trail = chop_str($trail, 30, 10);\n                                my $text = \"$lead<span\nclass=\\\"match\\\">$match</span>$trail\";\n-                               print chop_str($text, 80, 5) . \"<br/>\\n\";\n+                               # Do not chop $text as match can be\nlong, and we don't want to\n+                               # munge HTML tags!\n+                               #print chop_str($text, 80, 5) . \"<br/>\\n\";\n+                               print $text . \"<br/>\\n\";\n                        }\n                }\n                print \"</td>\\n\" .\n-- \n1.5.4.1-dirty\n"},{"id":"68670","messageId":"m3bq6kd5vq.fsf@localhost.localdomain","threadId":"12075","inReplyTo":"ae63f8b50802130937mddf9df9re2a95bee44661ee3@mail.gmail.com","subject":"Re: [PATCH] Do not chop HTML tags in commit search result","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-13T19:16:15Z","receivedAt":"2008-02-13T19:16:15Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Jean-Baptiste Quenot\" <jbq@caraldi.com> writes:\n\n> Thanks for Git! It's a great program.  I encountered an annoying bug\n> with gitweb 1.5.4.1, when searching for commits, if the search string\n> is too long, the generated HTML is munged leading to an ill-formed\n> XHTML document.\n> \n> Here is the patch, hope it helps:\n\nIt would be better if the patch wasn't wrapped and whitespace\ncorrupted.\n\n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index ae2d057..2c0b990 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -3780,7 +3780,10 @@ sub git_search_grep_body {\n>                                 my $trail = esc_html($3) || \"\";\n>                                 $trail = chop_str($trail, 30, 10);\n>                                 my $text = \"$lead<span class=\\\"match\\\">$match</span>$trail\";\n> -                               print chop_str($text, 80, 5) . \"<br/>\\n\";\n> +                               # Do not chop $text as match can be long, and we don't want to\n> +                               # munge HTML tags!\n> +                               #print chop_str($text, 80, 5) . \"<br/>\\n\";\n> +                               print $text . \"<br/>\\n\";\n>                         }\n>                 }\n>                 print \"</td>\\n\" .\n\nWhile this might be good bigfix, I think it is not a good solution. If\nwe select to show only neighbourhood of the match, we should probably\nchop trailing text, or both leading (cutting at beginning) and\ntrailing, perhaps even match if it is overly long.\n\nOr we could alternatively show all lines of commit message, and only\nmark matching fragment... but that I think would have to wait for\nrefactoring of generation of log-like views (log, shorlog, history,\nrss) in gitweb.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"68671","messageId":"7vbq6kprql.fsf@gitster.siamese.dyndns.org","threadId":"12075","inReplyTo":"ae63f8b50802130937mddf9df9re2a95bee44661ee3@mail.gmail.com","subject":"Re: [PATCH] Do not chop HTML tags in commit search result","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-02-13T19:43:14Z","receivedAt":"2008-02-13T19:43:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Jean-Baptiste Quenot\" <jbq@caraldi.com> writes:\n\n> ... I encountered an annoying bug\n> with gitweb 1.5.4.1, when searching for commits, if the search string\n> is too long, the generated HTML is munged leading to an ill-formed\n> XHTML document.\n\n> Here is the patch, hope it helps:\n>\n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index ae2d057..2c0b990 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -3780,7 +3780,10 @@ sub git_search_grep_body {\n>                                 my $trail = esc_html($3) || \"\";\n>                                 $trail = chop_str($trail, 30, 10);\n> ...\n>                                 my $text = \"$lead<span\n> class=\\\"match\\\">$match</span>$trail\";\n> -                               print chop_str($text, 80, 5) . \"<br/>\\n\";\n\nI think esc_html() and chop_str() are backwards here.  If $3 is\noverlong it is cut in the middle of some markup.  Even though\nchop_str() claims to be \"HTML aware\", I do not think it is.  It\nseems to know about \"&entities;\" but not mark-ups.\n\nThere are quite a many instances of esc_html() first then chop_str()\nin that function, and I think they all deserve to be fixed.\n\n\tmy $comment = $co{'comment'};\n\tforeach my $line (@$comment) {\n\t\tif ($line =~ m/^(.*)($search_regexp)(.*)$/i) {\n\t\t\tmy $lead = esc_html($1) || \"\";\n\t\t\t$lead = chop_str($lead, 30, 10);\n\t\t\tmy $match = esc_html($2) || \"\";\n\t\t\tmy $trail = esc_html($3) || \"\";\n\t\t\t$trail = chop_str($trail, 30, 10);\n\t\t\tmy $text = \"$lead<span class=\\\"match\\\">$match</span>$trail\";\n\t\t\tprint chop_str($text, 80, 5) . \"<br/>\\n\";\n\t\t}\n\t}\n\nI think this is trying to fit the result on a line, showing the\nmatch sandwitched by not-too-long part taken from leading and\ntrailing context ($lead and $trail can be chomped aggressively\nthan $match).  But $lead and $trail are escaped then chomped\nwhich is already wrong.\n\nI think the body of that if() would be better written like this:\n\n\tmy ($lead, $match, $trail) = ($1, $2, $3);\n\t$match = chop_str($match, 70, $slop); # in case it is very long...\n\t$contextlen = (80 - len($match)) / 2; # and the remainder...\n        if ($contextlen > 30) { $contextlen = 30 }; # but not too much\n        $trail = chop_str($trail, $contextlen, $slop);\n        $lead = chop_str($lead, $contextlen, $slop);\n\n\t$lead = esc_html($lead);\n\t$match = esc_html($match);\n\t$trail = esc_html($trail);\n\n        print \"$lead<span ...>$match</span>$trail\";\n"},{"id":"69590","messageId":"20080222163035.5942.93410.stgit@localhost.localdomain","threadId":"12075","inReplyTo":"7vbq6kprql.fsf@gitster.siamese.dyndns.org","subject":"[PATCH] gitweb: Better chopping in commit search results","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-22T16:33:47Z","receivedAt":"2008-02-22T16:33:47Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\nFrom: Junio C Hamano <gitster@pobox.com>\nSubject: [PATCH] gitweb: Better chopping in commit search results\n\nWhen searching commit messages (commit search), if matched string is\ntoo long, the generated HTML was munged leading to an ill-formed XHTML\ndocument.\n\nNow gitweb chop leading, trailing and matched parts, HTML escapes\nthose parts, then composes and marks up match info.  HTML output is\nnever chopped.  Limiting matched info to 80 columns (with slop) is now\ndone by dividing remaining characters after chopping match equally to\nleading and trailing part, not by chopping composed and HTML marked\noutput.\n\nNoticed-by: Jean-Baptiste Quenot <jbq@caraldi.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Jakub Narebski <jnareb@gmail.com>\n---\nThis is just slightly reworked Junio's patch; probably should be \nmarked as from Junio, so I'm trying to send it as it.\n\nStrange that StGit always sends patches (stg mail) as if repo owner\nwas their author, regardless of path/commit author (I think; unless\n\"stg edit\" cannot change authorship).\n\n gitweb/gitweb.perl |   24 +++++++++++++++---------\n 1 files changed, 15 insertions(+), 9 deletions(-)\n\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex 8ed6d04..326e27c 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -3784,18 +3784,24 @@ sub git_search_grep_body {\n \t\tprint \"<td title=\\\"$co{'age_string_age'}\\\"><i>$co{'age_string_date'}</i></td>\\n\" .\n \t\t      \"<td><i>\" . $author . \"</i></td>\\n\" .\n \t\t      \"<td>\" .\n-\t\t      $cgi->a({-href => href(action=>\"commit\", hash=>$co{'id'}), -class => \"list subject\"},\n-\t\t\t       chop_and_escape_str($co{'title'}, 50) . \"<br/>\");\n+\t\t      $cgi->a({-href => href(action=>\"commit\", hash=>$co{'id'}),\n+\t\t               -class => \"list subject\"},\n+\t\t              chop_and_escape_str($co{'title'}, 50) . \"<br/>\");\n \t\tmy $comment = $co{'comment'};\n \t\tforeach my $line (@$comment) {\n \t\t\tif ($line =~ m/^(.*)($search_regexp)(.*)$/i) {\n-\t\t\t\tmy $lead = esc_html($1) || \"\";\n-\t\t\t\t$lead = chop_str($lead, 30, 10);\n-\t\t\t\tmy $match = esc_html($2) || \"\";\n-\t\t\t\tmy $trail = esc_html($3) || \"\";\n-\t\t\t\t$trail = chop_str($trail, 30, 10);\n-\t\t\t\tmy $text = \"$lead<span class=\\\"match\\\">$match</span>$trail\";\n-\t\t\t\tprint chop_str($text, 80, 5) . \"<br/>\\n\";\n+\t\t\t\tmy ($lead, $match, $trail) = ($1, $2, $3);\n+\t\t\t\t$match = chop_str($match, 70, 5);       # in case match is very long\n+\t\t\t\tmy $contextlen = (80 - len($match))/2;  # is left for the remainder\n+\t\t\t\t$contextlen = 30 if ($contextlen > 30); # but not too much\n+\t\t\t\t$lead  = chop_str($lead,  $contextlen, 10);\n+\t\t\t\t$trail = chop_str($trail, $contextlen, 10);\n+\n+\t\t\t\t$lead  = esc_html($lead);\n+\t\t\t\t$match = esc_html($match);\n+\t\t\t\t$trail = esc_html($trail);\n+\n+\t\t\t\tprint \"$lead<span class=\\\"match\\\">$match</span>$trail<br />\";\n \t\t\t}\n \t\t}\n \t\tprint \"</td>\\n\" .\n"},{"id":"69593","messageId":"7voda8ap6r.fsf@gitster.siamese.dyndns.org","threadId":"12075","inReplyTo":"20080222163035.5942.93410.stgit@localhost.localdomain","subject":"Re: [PATCH] gitweb: Better chopping in commit search results","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-02-22T17:14:36Z","receivedAt":"2008-02-22T17:14:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> From: Junio C Hamano <gitster@pobox.com>\n> Subject: [PATCH] gitweb: Better chopping in commit search results\n>\n> When searching commit messages (commit search), if matched string is\n> too long, the generated HTML was munged leading to an ill-formed XHTML\n> document.\n>\n> Now gitweb chop leading, trailing and matched parts, HTML escapes\n> those parts, then composes and marks up match info.  HTML output is\n> never chopped.  Limiting matched info to 80 columns (with slop) is now\n> done by dividing remaining characters after chopping match equally to\n> leading and trailing part, not by chopping composed and HTML marked\n> output.\n\nCould somebody test this with very long search string, as that\nwas how the issue initially came up, to see (1) if it really\nfixes the \"mark-up chopped in the middle\" issue, (2) and how the\nactual output looks like?\n\nRegarding the latter, I have a slight suspicion that chopping\nthe tail of the middle part and showing very little context may\nnot produce a very useful output.\n\nFor example, if you are looking for \"very long ... and how\"\nin the first paragraph of message (if it were all on a single\nline), wouldn't you want to see:\n\n    ...st this with <<very long ... and how>> the actual out...\n\nrather than:\n\n    Could som... <<very long search stri...>> the actual out...\n\nin the result?\n\nThat is, it any chopping ever needs to happen, I suspect a more\nuseful way to shorten the output would be to:\n\n - divide the available space to give enough space to give\n   context for head and tail part.\n\n - chop head from the left, if needed, with leading ellipsis;\n\n - chop tail from the right, if needed, with trailing ellipsis;\n\n - chop search string from both ends, if needed, with leading\n   and trailing ellipses.\n\nHmm?\n"},{"id":"69599","messageId":"200802221849.44054.jnareb@gmail.com","threadId":"12075","inReplyTo":"7voda8ap6r.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] gitweb: Better chopping in commit search results","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-22T17:49:43Z","receivedAt":"2008-02-22T17:49:43Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Fri, 22 Feb 2008, Junio C Hamano wrote:\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n>> From: Junio C Hamano <gitster@pobox.com>\n>> Subject: [PATCH] gitweb: Better chopping in commit search results\n>>\n>> When searching commit messages (commit search), if matched string is\n>> too long, the generated HTML was munged leading to an ill-formed XHTML\n>> document.\n>>\n>> Now gitweb chop leading, trailing and matched parts, HTML escapes\n>> those parts, then composes and marks up match info.  HTML output is\n>> never chopped.  Limiting matched info to 80 columns (with slop) is now\n>> done by dividing remaining characters after chopping match equally to\n>> leading and trailing part, not by chopping composed and HTML marked\n>> output.\n> \n> Could somebody test this with very long search string, as that\n> was how the issue initially came up, to see (1) if it really\n> fixes the \"mark-up chopped in the middle\" issue, (2) [...]\n\nThe bug in question was cause by the chop _after_ doing HTML\nmarkup. Now gitweb chops, then HTML escapes, and chops no more.\nThere is no way this bug can happen now.\n\nBTW if commit messages follows \"wrap at 76 column\" convention\nit is not easy to test this condition... :-) \n\n\nBut you are right that output should be improved...\n \n> For example, if you are looking for \"very long ... and how\"\n> in the first paragraph of message (if it were all on a single\n> line), wouldn't you want to see:\n> \n>     ...st this with <<very long ... and how>> the actual out...\n> \n> rather than:\n> \n>     Could som... <<very long search stri...>> the actual out...\n> \n> in the result?\n\n...but I think it is better left for another patch.\n\nP.S. When testing this commit I have noticed that currently, probably\ndue to some misquoting, or interaction between escapemeta and quoting,\nsearching for messages which contain \"'\" (apostrophe), e.g. \"don't\"\ncurrently doesn't work. Will investigate...\n-- \nJakub Narebski\nPoland\n"},{"id":"69608","messageId":"200802222014.13205.jnareb@gmail.com","threadId":"12075","inReplyTo":"200802221849.44054.jnareb@gmail.com","subject":"Re: [PATCH] gitweb: Better chopping in commit search results","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-22T19:14:12Z","receivedAt":"2008-02-22T19:14:12Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jakub Narebski wrote:\n> On Fri, 22 Feb 2008, Junio C Hamano wrote:\n\n>> For example, if you are looking for \"very long ... and how\"\n>> in the first paragraph of message (if it were all on a single\n>> line), wouldn't you want to see:\n>> \n>>     ...st this with <<very long ... and how>> the actual out...\n>> \n>> rather than:\n>> \n>>     Could som... <<very long search stri...>> the actual out...\n>> \n>> in the result?\n> \n> ...but I think it is better left for another patch.\n\nEnd here is proposed improved chop_str which can do chopping at\nbeginning, in the middle, and (as it used to do) at the end.\n\nSome questions about the code:\n * should we divide slop in two also when chopping in the middle?\n * what should extra option be named, and what should be names of\n   posible values of this option (the option deciding where to chop)\n * $add_len has default value if not provided, or if 0 (!), or if '';\n   you have to use chop_str($str, 20, undef, -pos=>'center') trick\n   to use it with extra options.\n * can the code be improved? I'm not Perl expert.\n\n\n-- >8 --\nsub chop_str {\n\tmy $str = shift;\n\tmy $len = shift;\n\tmy $add_len = shift || 10;\n\t# supported opts:\n\t# * -pos => 'left' | 'center' | 'right', defaults to 'right'\n\t#   denotes where (which part) to chop\n\tmy %opts = @_;\n\n\t# allow only $len chars, but don't cut a word if it would fit in $add_len\n\t# if it doesn't fit, cut it if it's still longer than the dots we would add\n\t# remove chopped character entities entirely\n\n\t# when chopping in the middle, distribute $len into left and right part\n\tif (defined $opts{'-pos'} && $opts{'-pos'} eq 'center') {\n\t\t$len = int($len/2);\n\t}\n\n\t# regexps: ending and beginning with word part up to $add_len\n\tmy $endre = qr/.{0,$len}[^ \\/\\-_:\\.@]{0,$add_len}/;\n\tmy $begre = qr/[^ \\/\\-_:\\.@]{0,$add_len}.{0,$len}/;\n\n\tif (defined $opts{'-pos'} && $opts{'-pos'} eq 'left') {\n\t\t$str =~ m/^(.*?)($begre)$/;\n\t\tmy ($lead, $body) = ($1, $2);\n\t\tif (length($lead) > 4) {\n\t\t\tif ($lead =~ m/&[^;]*$/) {\n\t\t\t\t$body =~ s/^[^;]*;//; \n\t\t\t}\n\t\t\t$lead = \"... \";\n\t\t}\n\t\treturn \"$lead$body\";\n\n\t} elsif (defined $opts{'-pos'} && $opts{'-pos'} eq 'center') {\n\t\t$str =~ m/^($endre)(.*)$/;\n\t\tmy ($left, $str)  = ($1, $2);\n\t\t$str =~ m/^(.*?)($begre)$/;\n\t\tmy ($mid, $right) = ($1, $2);\n\t\tif (length($mid) > 5) {\n\t\t\t$left =~ s/&[^;]*$//;\n\t\t\tif ($mid =~ m/&[^;]*$/) {\n\t\t\t\t$right =~ s/^[^;]*;//;\n\t\t\t}\n\t\t\t$mid = \" ... \";\n\t\t}\n\t\treturn \"$left$mid$right\";\n\n\t} else {\n\t\t$str =~ m/^($endre)(.*)$/;\n\t\tmy $body = $1;\n\t\tmy $tail = $2;\n\t\tif (length($tail) > 4) {\n\t\t\t$body =~ s/&[^;]*$//;\n\t\t\t$tail = \" ...\";\n\t\t}\n\t\treturn \"$body$tail\";\n\t}\n}\n-- >8 --\n\nExample usage:\n  chop_str($str, 15, 5, -pos=>'center')\n\n-- \nJakub Narebski\nPoland\n"},{"id":"69669","messageId":"20080223092722.GB1390@diana.vm.bytemark.co.uk","threadId":"12075","inReplyTo":"20080222163035.5942.93410.stgit@localhost.localdomain","subject":"Re: [PATCH] gitweb: Better chopping in commit search results","fromName":"Karl Hasselström","fromEmail":"kha@treskal.com","sentAt":"2008-02-23T09:27:22Z","receivedAt":"2008-02-23T09:27:22Z","isPatch":true,"sender":{"key":"kha@treskal.com","avatar":"https://gravatar.com/avatar/f0120c734b5279b345075a28521e1ac66acb20c9913ffe9bf6ae97e53f7f3f13?d=mp&s=160"},"body":"On 2008-02-22 17:33:47 +0100, Jakub Narebski wrote:\n\n> Strange that StGit always sends patches (stg mail) as if repo owner\n> was their author, regardless of path/commit author (I think; unless\n> \"stg edit\" cannot change authorship).\n\nIt'll always put the sender in the From: header in the mail,\nobviously. But the patch mail template contains a \"%(fromauth)s\"\nstring, which will be replaced by \"From: Real Author\n<ra@example.com>\\n\", or the empty string if the patch author is the\nsame as the sender.\n\nYour problem might be that you're using an old template.\n\n-- \nKarl Hasselström, kha@treskal.com\n      www.treskal.com/kalle\n"},{"id":"69670","messageId":"200802231120.19448.jnareb@gmail.com","threadId":"12075","inReplyTo":"20080223092722.GB1390@diana.vm.bytemark.co.uk","subject":"Re: [PATCH] gitweb: Better chopping in commit search results","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-23T10:20:18Z","receivedAt":"2008-02-23T10:20:18Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, 23 Feb 2008, Karl Hasselström wrote:\n> On 2008-02-22 17:33:47 +0100, Jakub Narebski wrote:\n> \n> > Strange that StGit always sends patches (stg mail) as if repo owner\n> > was their author, regardless of path/commit author (I think; unless\n> > \"stg edit\" cannot change authorship).\n> \n> It'll always put the sender in the From: header in the mail,\n> obviously. But the patch mail template contains a \"%(fromauth)s\"\n> string, which will be replaced by \"From: Real Author\n> <ra@example.com>\\n\", or the empty string if the patch author is the\n> same as the sender.\n> \n> Your problem might be that you're using an old template.\n\nThanks a lot. That was it.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"69716","messageId":"20080223214226.16470.29333.stgit@localhost.localdomain","threadId":"12075","inReplyTo":"200802222014.13205.jnareb@gmail.com","subject":"[RFC/PATCH] gitweb: Option to chop at beginning and in the middle in chop_str","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-23T21:44:24Z","receivedAt":"2008-02-23T21:44:24Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\nAdd support for '-cut' option to chop_str subroutine, to cut at the\nbeginning (from the left side of the string), in the middle (center of\nthe string), or at the end (from the right side of the string) which\nis the default:\n  chop_str(somestring, len, slop, -pos=>'left')    ->  ... string\n  chop_str(somestring, len, slop, -pos=>'center')  ->  som ... ing\n  chop_str(somestring, len, slop, -pos=>'right')   ->  somestr ...\n\nSimplify passing all arguments to chop_str in chop_and_escape_str\nsubroutine. This was needed to pass additional options to chop_str.\n\n\nMake use of new feature of chop_str to better cut matched string and\nits context in match info for searching commit messages (commit\nsearch), as proposed by Junio C Hamano.  For example, if you are\nlooking for \"very long ... and how\" in the first paragraph of message\n(if it were all on a single line), you would now see:\n\n    ...st this with <<very long ... and how>> the actual out...\n\ninstead of:\n\n    Could som... <<very long search stri...>> the actual out...\n\n(where <<something>> denotes emphasized / colored fragment).\n\nSigned-off-by: Jakub Narebski <jnareb@gmail.com>\n---\nAnd here it is as a patch to gitweb to play with. I agree with Junio\nthat the output is better, but I'm not sure if it is worth the\ncomplication in code; perhaps is.\n\n gitweb/gitweb.perl |   74 +++++++++++++++++++++++++++++++++++++++++-----------\n 1 files changed, 58 insertions(+), 16 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex e8226b1..59d44b3 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -848,32 +848,74 @@ sub project_in_list {\n ## ----------------------------------------------------------------------\n ## HTML aware string manipulation\n \n+# cut (chop) string to given length, with additional slop,\n+# optionally from left and in the middle.\n sub chop_str {\n \tmy $str = shift;\n \tmy $len = shift;\n \tmy $add_len = shift || 10;\n+\t# supported opts:\n+\t# * -cut => 'left' | 'center' | 'right', defaults to 'right'\n+\t#   denotes where (which part) to chop\n+\tmy %opts = @_;\n \n \t# allow only $len chars, but don't cut a word if it would fit in $add_len\n \t# if it doesn't fit, cut it if it's still longer than the dots we would add\n-\t$str =~ m/^(.{0,$len}[^ \\/\\-_:\\.@]{0,$add_len})(.*)/;\n-\tmy $body = $1;\n-\tmy $tail = $2;\n-\tif (length($tail) > 4) {\n-\t\t$tail = \" ...\";\n-\t\t$body =~ s/&[^;]*$//; # remove chopped character entities\n+\t# remove chopped character entities entirely\n+\n+\t# when chopping in the middle, distribute $len into left and right part\n+\tif (defined $opts{'-cut'} && $opts{'-cut'} eq 'center') {\n+\t\t$len = int($len/2);\n+\t}\n+\n+\t# regexps: ending and beginning with word part up to $add_len\n+\tmy $endre = qr/.{0,$len}[^ \\/\\-_:\\.@]{0,$add_len}/;\n+\tmy $begre = qr/[^ \\/\\-_:\\.@]{0,$add_len}.{0,$len}/;\n+\n+\tif (defined $opts{'-cut'} && $opts{'-cut'} eq 'left') {\n+\t\t$str =~ m/^(.*?)($begre)$/;\n+\t\tmy ($lead, $body) = ($1, $2);\n+\t\tif (length($lead) > 4) {\n+\t\t\tif ($lead =~ m/&[^;]*$/) {\n+\t\t\t\t$body =~ s/^[^;]*;//;\n+\t\t\t}\n+\t\t\t$lead = \"... \";\n+\t\t}\n+\t\treturn \"$lead$body\";\n+\n+\t} elsif (defined $opts{'-cut'} && $opts{'-cut'} eq 'center') {\n+\t\t$str =~ m/^($endre)(.*)$/;\n+\t\tmy ($left, $str)  = ($1, $2);\n+\t\t$str =~ m/^(.*?)($begre)$/;\n+\t\tmy ($mid, $right) = ($1, $2);\n+\t\tif (length($mid) > 5) {\n+\t\t\t$left =~ s/&[^;]*$//;\n+\t\t\tif ($mid =~ m/&[^;]*$/) {\n+\t\t\t\t$right =~ s/^[^;]*;//;\n+\t\t\t}\n+\t\t\t$mid = \" ... \";\n+\t\t}\n+\t\treturn \"$left$mid$right\";\n+\n+\t} else {\n+\t\t$str =~ m/^($endre)(.*)$/;\n+\t\tmy $body = $1;\n+\t\tmy $tail = $2;\n+\t\tif (length($tail) > 4) {\n+\t\t\t$body =~ s/&[^;]*$//;\n+\t\t\t$tail = \" ...\";\n+\t\t}\n+\t\treturn \"$body$tail\";\n \t}\n-\treturn \"$body$tail\";\n }\n \n # takes the same arguments as chop_str, but also wraps a <span> around the\n # result with a title attribute if it does get chopped. Additionally, the\n # string is HTML-escaped.\n sub chop_and_escape_str {\n-\tmy $str = shift;\n-\tmy $len = shift;\n-\tmy $add_len = shift || 10;\n+\tmy ($str) = @_;\n \n-\tmy $chopped = chop_str($str, $len, $add_len);\n+\tmy $chopped = chop_str(@_);\n \tif ($chopped eq $str) {\n \t\treturn esc_html($chopped);\n \t} else {\n@@ -3791,11 +3833,11 @@ sub git_search_grep_body {\n \t\tforeach my $line (@$comment) {\n \t\t\tif ($line =~ m/^(.*)($search_regexp)(.*)$/i) {\n \t\t\t\tmy ($lead, $match, $trail) = ($1, $2, $3);\n-\t\t\t\t$match = chop_str($match, 70, 5);       # in case match is very long\n-\t\t\t\tmy $contextlen = int((80 - length($match))/2); # for the remainder\n-\t\t\t\t$contextlen = 30 if ($contextlen > 30); # but not too much\n-\t\t\t\t$lead  = chop_str($lead,  $contextlen, 10);\n-\t\t\t\t$trail = chop_str($trail, $contextlen, 10);\n+\t\t\t\t$match = chop_str($match, 70, 5, -cut=>'center');\n+\t\t\t\tmy $contextlen = int((80 - length($match))/2);\n+\t\t\t\t$contextlen = 30 if ($contextlen > 30);\n+\t\t\t\t$lead  = chop_str($lead,  $contextlen, 10, -cut=>'left');\n+\t\t\t\t$trail = chop_str($trail, $contextlen, 10, -cut=>'right');\n \n \t\t\t\t$lead  = esc_html($lead);\n \t\t\t\t$match = esc_html($match);\n"},{"id":"69718","messageId":"7vd4qn1ga2.fsf@gitster.siamese.dyndns.org","threadId":"12075","inReplyTo":"200802222014.13205.jnareb@gmail.com","subject":"Re: [PATCH] gitweb: Better chopping in commit search results","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-02-23T22:04:05Z","receivedAt":"2008-02-23T22:04:05Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> \t# regexps: ending and beginning with word part up to $add_len\n> \tmy $endre = qr/.{0,$len}[^ \\/\\-_:\\.@]{0,$add_len}/;\n> \tmy $begre = qr/[^ \\/\\-_:\\.@]{0,$add_len}.{0,$len}/;\n\nI have no idea what these line noise characters inside [] are.\nDid you mean something like \"\\w\"?\n\nI have a suspicion that it may be easier to read and could be\neven more efficient to split an overlong line at word boundaries\nand to remove elements from the end you are removing from until\nit fits.\n\nsub chop_whence {\n\tmy ($line, $max, $slop, $where) = @_;\n\n\tmy $len = length($line);\n\tif ($len < $max + $slop) {\n\t\treturn $line;\n\t}\n\n\t# Cut at word boundaries\n\tmy @split = split(/\\b/, $line);\n\tmy $filler = \"...\";\n\n\twhile ((2 < @split)) {\n\t\tmy $removed;\n\t\tmy $splice_at;\n\t\tif ($where eq 'left') {\n\t\t\t$removed = shift @split;\n\t\t} elsif ($where eq 'right') {\n\t\t\t$removed = pop @split;\n\t\t} else {\n\t\t\tmy $splice_at = int($#split / 2);\n\t\t\t$removed = splice(@split, $splice_at, 1);\n\t\t}\n\t\t$len -= length($removed) + length($filler);\n\t\tif ($len < $max + $slop) {\n\t\t\tif ($where eq 'left') {\n\t\t\t\tunshift @split, $filler;\n\t\t\t} elsif ($where eq 'right') {\n\t\t\t\tpush @split, $filler;\n\t\t\t} else {\n\t\t\t\tmy $splice_at = int($#split / 2);\n\t\t\t\tsplice(@split, $splice_at, 0, $filler);\n\t\t\t}\n\t\t\treturn join('', @split);\n\t\t}\n\t}\n\t# give up\n\treturn $line;\n}\n"},{"id":"69739","messageId":"200802240036.02406.jnareb@gmail.com","threadId":"12075","inReplyTo":"7vd4qn1ga2.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] gitweb: Better chopping in commit search results","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-23T23:36:01Z","receivedAt":"2008-02-23T23:36:01Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, 23 Feb 2008, Junio C Hamano wrote:\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n> > \t# regexps: ending and beginning with word part up to $add_len\n> > \tmy $endre = qr/.{0,$len}[^ \\/\\-_:\\.@]{0,$add_len}/;\n> > \tmy $begre = qr/[^ \\/\\-_:\\.@]{0,$add_len}.{0,$len}/;\n> \n> I have no idea what these line noise characters inside [] are.\n> Did you mean something like \"\\w\"?\n\nThey were in original chop_str, written by Kay Sievers if I have\nchecked correctly, I have only repeated it. But changing it to\n\"\\w\" might be a good idea.\n\n> I have a suspicion that it may be easier to read and could be\n> even more efficient to split an overlong line at word boundaries\n> and to remove elements from the end you are removing from until\n> it fits.\n> \n> sub chop_whence {\n> \tmy ($line, $max, $slop, $where) = @_;\n\nIMHO it is neither easier to read, nor more efficient. And changes\nsemantic a bit.\n\n\nOriginal chop_str used to mean:\n\n chop_str($str, $len, $add_len) means: Try to chop on a word boundary \n between position $len and $len+$add_len. If there is no word boundary\n between $len and $len+$add_len, chop at $add_len, and replace chopped\n part by \" ...\". Do not chop if to be replaced part is shorter than\n ellipsis, i.e. \" ...\" replacement.\n\n(I think I add this description to gitweb). With this description the \ncode is I think quite obvious.\n-- \nJakub Narebski\nPoland\n"},{"id":"69785","messageId":"20080224125920.24448.2179.stgit@localhost.localdomain","threadId":"12075","inReplyTo":"200802222014.13205.jnareb@gmail.com","subject":"[RFC/PATCH v2] gitweb: Option to chop at beginning and in the middle in chop_str","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-24T13:01:17Z","receivedAt":"2008-02-24T13:01:17Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Add support for '-cut' option to chop_str subroutine, to cut at the\nbeginning (from the left side of the string), in the middle (center of\nthe string), or at the end (from the right side of the string) which\nis the default:\n  chop_str(somestring, len, slop, -cut=>'left')    ->  ' ...string'\n  chop_str(somestring, len, slop, -cut=>'center')  ->  'som ... ing'\n  chop_str(somestring, len, slop, -cut=>'right')   ->  'somestr... '\n\nWhile at it return from chop_str early if given string is so short\nthat chop_str couldn't shorten it, and simplify regexp used. Make\nellipsis (ots) stick to shorthened fragment for cutting at ends.\n\nSimplify passing all arguments to chop_str in chop_and_escape_str\nsubroutine. This was needed to pass additional options to chop_str.\n\n\nMake use of new feature of chop_str to better cut matched string and\nits context in match info for searching commit messages (commit\nsearch), as proposed by Junio C Hamano.  For example, if you are\nlooking for \"very long ... and how\" in the first paragraph of message\n(if it were all on a single line), you would now see:\n\n    ...st this with <<very long ... and how>> the actual out...\n\ninstead of:\n\n    Could som... <<very long search stri...>> the actual out...\n\n(where <<something>> denotes emphasized / colored fragment).\n\nSigned-off-by: Jakub Narebski <jnareb@gmail.com>\n---\nThis is second version of the patch, with regexp simplified,\nand early return. Rudimentarly tested.\n\n gitweb/gitweb.perl |   77 +++++++++++++++++++++++++++++++++++++++++-----------\n 1 files changed, 61 insertions(+), 16 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex e8226b1..9a8e0a6 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -848,32 +848,77 @@ sub project_in_list {\n ## ----------------------------------------------------------------------\n ## HTML aware string manipulation\n \n+# Try to chop given string on a word boundary between position\n+# $len and $len+$add_len. If there is no word boundary there,\n+# chop at $len+$add_len. Do not chop if chopped part plus ellipsis\n+# (marking chopped part) would be longer than given string.\n sub chop_str {\n \tmy $str = shift;\n \tmy $len = shift;\n \tmy $add_len = shift || 10;\n+\tmy $where = shift || 'right' # 'left' | 'center' | 'right'\n \n \t# allow only $len chars, but don't cut a word if it would fit in $add_len\n \t# if it doesn't fit, cut it if it's still longer than the dots we would add\n-\t$str =~ m/^(.{0,$len}[^ \\/\\-_:\\.@]{0,$add_len})(.*)/;\n-\tmy $body = $1;\n-\tmy $tail = $2;\n-\tif (length($tail) > 4) {\n-\t\t$tail = \" ...\";\n-\t\t$body =~ s/&[^;]*$//; # remove chopped character entities\n+\t# remove chopped character entities entirely\n+\n+\t# when chopping in the middle, distribute $len into left and right part\n+\t# return early if chupping wouldn't make string shorter\n+\tif ($where eq 'center') {\n+\t\treturn $str if ($len + 5 >= length($str)); # filler is length 5\n+\t\t$len = int($len/2);\n+\t} else {\n+\t\treturn $str if ($len + 4 >= length($str)); # filler is length 4\n+\t}\n+\n+\t# regexps: ending and beginning with word part up to $add_len\n+\tmy $endre = qr/.{$len}\\w{0,$add_len}/;\n+\tmy $begre = qr/\\w{0,$add_len}.{$len}/;\n+\n+\tif ($where eq 'left') {\n+\t\t$str =~ m/^(.*?)($begre)$/;\n+\t\tmy ($lead, $body) = ($1, $2);\n+\t\tif (length($lead) > 4) {\n+\t\t\tif ($lead =~ m/&[^;]*$/) {\n+\t\t\t\t$body =~ s/^[^;]*;//;\n+\t\t\t}\n+\t\t\t$lead = \" ...\";\n+\t\t}\n+\t\treturn \"$lead$body\";\n+\n+\t} elsif ($where eq 'center') {\n+\t\t$str =~ m/^($endre)(.*)$/;\n+\t\tmy ($left, $str)  = ($1, $2);\n+\t\t$str =~ m/^(.*?)($begre)$/;\n+\t\tmy ($mid, $right) = ($1, $2);\n+\t\tif (length($mid) > 5) {\n+\t\t\t$left =~ s/&[^;]*$//;\n+\t\t\tif ($mid =~ m/&[^;]*$/) {\n+\t\t\t\t$right =~ s/^[^;]*;//;\n+\t\t\t}\n+\t\t\t$mid = \" ... \";\n+\t\t}\n+\t\treturn \"$left$mid$right\";\n+\n+\t} else {\n+\t\t$str =~ m/^($endre)(.*)$/;\n+\t\tmy $body = $1;\n+\t\tmy $tail = $2;\n+\t\tif (length($tail) > 4) {\n+\t\t\t$body =~ s/&[^;]*$//;\n+\t\t\t$tail = \"... \";\n+\t\t}\n+\t\treturn \"$body$tail\";\n \t}\n-\treturn \"$body$tail\";\n }\n \n # takes the same arguments as chop_str, but also wraps a <span> around the\n # result with a title attribute if it does get chopped. Additionally, the\n # string is HTML-escaped.\n sub chop_and_escape_str {\n-\tmy $str = shift;\n-\tmy $len = shift;\n-\tmy $add_len = shift || 10;\n+\tmy ($str) = @_;\n \n-\tmy $chopped = chop_str($str, $len, $add_len);\n+\tmy $chopped = chop_str(@_);\n \tif ($chopped eq $str) {\n \t\treturn esc_html($chopped);\n \t} else {\n@@ -3791,11 +3836,11 @@ sub git_search_grep_body {\n \t\tforeach my $line (@$comment) {\n \t\t\tif ($line =~ m/^(.*)($search_regexp)(.*)$/i) {\n \t\t\t\tmy ($lead, $match, $trail) = ($1, $2, $3);\n-\t\t\t\t$match = chop_str($match, 70, 5);       # in case match is very long\n-\t\t\t\tmy $contextlen = int((80 - length($match))/2); # for the remainder\n-\t\t\t\t$contextlen = 30 if ($contextlen > 30); # but not too much\n-\t\t\t\t$lead  = chop_str($lead,  $contextlen, 10);\n-\t\t\t\t$trail = chop_str($trail, $contextlen, 10);\n+\t\t\t\t$match = chop_str($match, 70, 5, 'center');\n+\t\t\t\tmy $contextlen = int((80 - length($match))/2);\n+\t\t\t\t$contextlen = 30 if ($contextlen > 30);\n+\t\t\t\t$lead  = chop_str($lead,  $contextlen, 10, 'left');\n+\t\t\t\t$trail = chop_str($trail, $contextlen, 10, 'right');\n \n \t\t\t\t$lead  = esc_html($lead);\n \t\t\t\t$match = esc_html($match);\n\n\n-- \nStacked GIT 0.14.1\ngit version 1.5.4.2\n"},{"id":"69838","messageId":"7v8x19st7x.fsf@gitster.siamese.dyndns.org","threadId":"12075","inReplyTo":"20080224125920.24448.2179.stgit@localhost.localdomain","subject":"Re: [RFC/PATCH v2] gitweb: Option to chop at beginning and in the middle in chop_str","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-02-25T01:46:58Z","receivedAt":"2008-02-25T01:46:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> Make use of new feature of chop_str to better cut matched string and\n> its context in match info for searching commit messages (commit\n> search), as proposed by Junio C Hamano.  For example, if you are\n> looking for \"very long ... and how\" in the first paragraph of this message\n> (if it were all on a single line), you would now see:\n>\n>     ...st this with <<very long ... and how>> the actual out...\n>\n> instead of:\n>\n>     Could som... <<very long search stri...>> the actual out...\n>\n> (where <<something>> denotes emphasized / colored fragment).\n\nThis part needs rewritten; the first paragraph of what message is that?  \n\nAlso I think the subject is wrong.  Yes, it is adding an option\nto an internal subroutine but who cares?  The net effect the\n\"gitweb\" users see is that the way the grep result is shown\ndifferently, hopefully in a more understandable way, and that\nchange is not _optional_ at all.\n\nThe code looks easier to read than before, but I may be partial\n;-)\n"},{"id":"69913","messageId":"200802252107.59366.jnareb@gmail.com","threadId":"12075","inReplyTo":"7v8x19st7x.fsf@gitster.siamese.dyndns.org","subject":"[RFC/PATCH v3] gitweb: Better cutting matched string and its context","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-02-25T20:07:57Z","receivedAt":"2008-02-25T20:07:57Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Mon, 25 Feb 2008, Junio C Hamano wrote:\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n>> Make use of new feature of chop_str to better cut matched string and\n>> its context in match info for searching commit messages (commit\n>> search), as proposed by Junio C Hamano.  For example, if you are\n>> looking for \"very long ... and how\" in the first paragraph of this message\n>> (if it were all on a single line), you would now see:\n>>\n>>     ...st this with <<very long ... and how>> the actual out...\n>>\n>> instead of:\n>>\n>>     Could som... <<very long search stri...>> the actual out...\n>>\n>> (where <<something>> denotes emphasized / colored fragment).\n> \n> This part needs rewritten; the first paragraph of what message is that?  \n\nOoops. Sorry about that. I have forgot to transport context when\ndoing copy'n'paste.\n\nBTW. the example was subtly wrong: you can search _lines_, like grep,\nyou cannot search multiline.\n \n> Also I think the subject is wrong.  Yes, it is adding an option\n> to an internal subroutine but who cares?  The net effect the\n> \"gitweb\" users see is that the way the grep result is shown\n> differently, hopefully in a more understandable way, and that\n> change is not _optional_ at all.\n\nBelow there is corrected commit message (rewritten to match code).\n\n> The code looks easier to read than before, but I may be partial\n> ;-)\n\nFurther improved, IMVHO.\n\n\nBTW. I have queued 3-patch series further improving gitweb search,\nto be sent soon...\n\n.....................................................................\n-- >8 --\nFrom: Jakub Narebski <jnareb@gmail.com>\nSubject: [PATCH] gitweb: Better cutting matched string and its context\n\nImprove look of commit search output ('search' view) by better cutting\nof matched string and its context in match info, as proposed by Junio\nHamano.  For example, if you are looking for \"very long search string\"\nin the following line:\n\n    Could somebody test this with very long search string, and see how\n\nyou would now see:\n\n    ...this with <<very long ... string>>, and see...\n\ninstead of:\n\n    Could som... <<very long search...>>, and see...\n\n(where <<something>> denotes emphasized / colored fragment; matched\nfragment to be more exact).\n\n\nFor this feature support for fourth [optional] parameter to chop_str\nsubroutine was added.  This fourth parameter is used to denote where\nto cut string to make it shorter.  chop_str can now cut at the\nbeginning (from the _left_ side of the string), in the middle\n(_center_ of the string), or at the end (from the _right_ side of\nthe string); cutting from right is the default:\n\n  chop_str(somestring, len, slop, 'left')    ->  ' ...string'\n  chop_str(somestring, len, slop, 'center')  ->  'som ... ing'\n  chop_str(somestring, len, slop, 'right')   ->  'somestr... '\n\nIf you want to use default slop (default additional length), use undef\nas value for third parameter to chop_str.\n\nWhile at it, return from chop_str early if given string is so short\nthat chop_str couldn't shorten it.  Simplify also regexp used by\nchop_str.  Make ellipsis (dots) stick to shortened fragment for\ncutting at ends, to better see which part got shortened.\n\nSimplify passing all arguments to chop_str in chop_and_escape_str\nsubroutine. This was needed to pass additional options to chop_str.\n\nSigned-off-by: Jakub Narebski <jnareb@gmail.com>\n---\n gitweb/gitweb.perl |   73 ++++++++++++++++++++++++++++++++++++++++-----------\n 1 files changed, 57 insertions(+), 16 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex e8226b1..fc95e2c 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -848,32 +848,73 @@ sub project_in_list {\n ## ----------------------------------------------------------------------\n ## HTML aware string manipulation\n \n+# Try to chop given string on a word boundary between position\n+# $len and $len+$add_len. If there is no word boundary there,\n+# chop at $len+$add_len. Do not chop if chopped part plus ellipsis\n+# (marking chopped part) would be longer than given string.\n sub chop_str {\n \tmy $str = shift;\n \tmy $len = shift;\n \tmy $add_len = shift || 10;\n+\tmy $where = shift || 'right'; # 'left' | 'center' | 'right'\n \n \t# allow only $len chars, but don't cut a word if it would fit in $add_len\n \t# if it doesn't fit, cut it if it's still longer than the dots we would add\n-\t$str =~ m/^(.{0,$len}[^ \\/\\-_:\\.@]{0,$add_len})(.*)/;\n-\tmy $body = $1;\n-\tmy $tail = $2;\n-\tif (length($tail) > 4) {\n-\t\t$tail = \" ...\";\n-\t\t$body =~ s/&[^;]*$//; # remove chopped character entities\n+\t# remove chopped character entities entirely\n+\n+\t# when chopping in the middle, distribute $len into left and right part\n+\t# return early if chopping wouldn't make string shorter\n+\tif ($where eq 'center') {\n+\t\treturn $str if ($len + 5 >= length($str)); # filler is length 5\n+\t\t$len = int($len/2);\n+\t} else {\n+\t\treturn $str if ($len + 4 >= length($str)); # filler is length 4\n+\t}\n+\n+\t# regexps: ending and beginning with word part up to $add_len\n+\tmy $endre = qr/.{$len}\\w{0,$add_len}/;\n+\tmy $begre = qr/\\w{0,$add_len}.{$len}/;\n+\n+\tif ($where eq 'left') {\n+\t\t$str =~ m/^(.*?)($begre)$/;\n+\t\tmy ($lead, $body) = ($1, $2);\n+\t\tif (length($lead) > 4) {\n+\t\t\t$body =~ s/^[^;]*;// if ($lead =~ m/&[^;]*$/);\n+\t\t\t$lead = \" ...\";\n+\t\t}\n+\t\treturn \"$lead$body\";\n+\n+\t} elsif ($where eq 'center') {\n+\t\t$str =~ m/^($endre)(.*)$/;\n+\t\tmy ($left, $str)  = ($1, $2);\n+\t\t$str =~ m/^(.*?)($begre)$/;\n+\t\tmy ($mid, $right) = ($1, $2);\n+\t\tif (length($mid) > 5) {\n+\t\t\t$left  =~ s/&[^;]*$//;\n+\t\t\t$right =~ s/^[^;]*;// if ($mid =~ m/&[^;]*$/);\n+\t\t\t$mid = \" ... \";\n+\t\t}\n+\t\treturn \"$left$mid$right\";\n+\n+\t} else {\n+\t\t$str =~ m/^($endre)(.*)$/;\n+\t\tmy $body = $1;\n+\t\tmy $tail = $2;\n+\t\tif (length($tail) > 4) {\n+\t\t\t$body =~ s/&[^;]*$//;\n+\t\t\t$tail = \"... \";\n+\t\t}\n+\t\treturn \"$body$tail\";\n \t}\n-\treturn \"$body$tail\";\n }\n \n # takes the same arguments as chop_str, but also wraps a <span> around the\n # result with a title attribute if it does get chopped. Additionally, the\n # string is HTML-escaped.\n sub chop_and_escape_str {\n-\tmy $str = shift;\n-\tmy $len = shift;\n-\tmy $add_len = shift || 10;\n+\tmy ($str) = @_;\n \n-\tmy $chopped = chop_str($str, $len, $add_len);\n+\tmy $chopped = chop_str(@_);\n \tif ($chopped eq $str) {\n \t\treturn esc_html($chopped);\n \t} else {\n@@ -3791,11 +3832,11 @@ sub git_search_grep_body {\n \t\tforeach my $line (@$comment) {\n \t\t\tif ($line =~ m/^(.*)($search_regexp)(.*)$/i) {\n \t\t\t\tmy ($lead, $match, $trail) = ($1, $2, $3);\n-\t\t\t\t$match = chop_str($match, 70, 5);       # in case match is very long\n-\t\t\t\tmy $contextlen = int((80 - length($match))/2); # for the remainder\n-\t\t\t\t$contextlen = 30 if ($contextlen > 30); # but not too much\n-\t\t\t\t$lead  = chop_str($lead,  $contextlen, 10);\n-\t\t\t\t$trail = chop_str($trail, $contextlen, 10);\n+\t\t\t\t$match = chop_str($match, 70, 5, 'center');\n+\t\t\t\tmy $contextlen = int((80 - length($match))/2);\n+\t\t\t\t$contextlen = 30 if ($contextlen > 30);\n+\t\t\t\t$lead  = chop_str($lead,  $contextlen, 10, 'left');\n+\t\t\t\t$trail = chop_str($trail, $contextlen, 10, 'right');\n \n \t\t\t\t$lead  = esc_html($lead);\n \t\t\t\t$match = esc_html($match);\n-- \n1.5.4.2\n"},{"id":"69918","messageId":"7vskzghjrl.fsf@gitster.siamese.dyndns.org","threadId":"12075","inReplyTo":"200802252107.59366.jnareb@gmail.com","subject":"Re: [RFC/PATCH v3] gitweb: Better cutting matched string and its context","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-02-25T20:18:54Z","receivedAt":"2008-02-25T20:18:54Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> Ooops. Sorry about that. I have forgot to transport context when\n> doing copy'n'paste.\n> ...\n> BTW. the example was subtly wrong: you can search _lines_, like grep,\n> you cannot search multiline.\n\nYeah, the original had:\n\n        For example, if you are looking for \"very long ... and how\"\n        in the first paragraph of message (if it were all on a single\n        line), wouldn't you want to see:\n\nand you dropped \"if it were all on a single line\" part, as you\nforgot to transport context when doing copy-n-paste ;-)\n\n> Below there is corrected commit message (rewritten to match code).\n\nThanks.  Will take a look later; I am at day-job today.\n"}]}