git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[RFC/PATCH] gitweb: Option to chop at beginning and in the middle in chop_str

From
Jakub Narebski <jnareb@gmail.com>
Date
Feb 23, 2008, 21:44 UTC
Message-ID
<20080223214226.16470.29333.stgit@localhost.localdomain>
In-Reply-To
<200802222014.13205.jnareb@gmail.com>
Add support for '-cut' option to chop_str subroutine, to cut at the
beginning (from the left side of the string), in the middle (center of
the string), or at the end (from the right side of the string) which
is the default:
  chop_str(somestring, len, slop, -pos=>'left')    ->  ... string
  chop_str(somestring, len, slop, -pos=>'center')  ->  som ... ing
  chop_str(somestring, len, slop, -pos=>'right')   ->  somestr ...

Simplify passing all arguments to chop_str in chop_and_escape_str subroutine. This was needed to pass additional options to chop_str.

Make use of new feature of chop_str to better cut matched string and its context in match info for searching commit messages (commit search), as proposed by Junio C Hamano. For example, if you are looking for "very long ... and how" in the first paragraph of message (if it were all on a single line), you would now see:

    ...st this with <<very long ... and how>> the actual out...
instead of:
    Could som... <<very long search stri...>> the actual out...
(where <<something>> denotes emphasized / colored fragment).
Signed-off-by: Jakub Narebski <jnareb@gmail.com>
---
And here it is as a patch to gitweb to play with. I agree with Junio
that the output is better, but I'm not sure if it is worth the
complication in code; perhaps is.
 gitweb/gitweb.perl |   74 +++++++++++++++++++++++++++++++++++++++++-----------
 1 files changed, 58 insertions(+), 16 deletions(-)
diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl
index e8226b1..59d44b3 100755
--- a/gitweb/gitweb.perl
+++ b/gitweb/gitweb.perl
@@ -848,32 +848,74 @@ sub project_in_list {
 ## ----------------------------------------------------------------------
 ## HTML aware string manipulation
 
+# cut (chop) string to given length, with additional slop,
+# optionally from left and in the middle.
 sub chop_str {
 	my $str = shift;
 	my $len = shift;
 	my $add_len = shift || 10;
+	# supported opts:
+	# * -cut => 'left' | 'center' | 'right', defaults to 'right'
+	#   denotes where (which part) to chop
+	my %opts = @_;
 
 	# allow only $len chars, but don't cut a word if it would fit in $add_len
 	# if it doesn't fit, cut it if it's still longer than the dots we would add
-	$str =~ m/^(.{0,$len}[^ \/\-_:\.@]{0,$add_len})(.*)/;
-	my $body = $1;
-	my $tail = $2;
-	if (length($tail) > 4) {
-		$tail = " ...";
-		$body =~ s/&[^;]*$//; # remove chopped character entities
+	# remove chopped character entities entirely
+
+	# when chopping in the middle, distribute $len into left and right part
+	if (defined $opts{'-cut'} && $opts{'-cut'} eq 'center') {
+		$len = int($len/2);
+	}
+
+	# regexps: ending and beginning with word part up to $add_len
+	my $endre = qr/.{0,$len}[^ \/\-_:\.@]{0,$add_len}/;
+	my $begre = qr/[^ \/\-_:\.@]{0,$add_len}.{0,$len}/;
+
+	if (defined $opts{'-cut'} && $opts{'-cut'} eq 'left') {
+		$str =~ m/^(.*?)($begre)$/;
+		my ($lead, $body) = ($1, $2);
+		if (length($lead) > 4) {
+			if ($lead =~ m/&[^;]*$/) {
+				$body =~ s/^[^;]*;//;
+			}
+			$lead = "... ";
+		}
+		return "$lead$body";
+
+	} elsif (defined $opts{'-cut'} && $opts{'-cut'} eq 'center') {
+		$str =~ m/^($endre)(.*)$/;
+		my ($left, $str)  = ($1, $2);
+		$str =~ m/^(.*?)($begre)$/;
+		my ($mid, $right) = ($1, $2);
+		if (length($mid) > 5) {
+			$left =~ s/&[^;]*$//;
+			if ($mid =~ m/&[^;]*$/) {
+				$right =~ s/^[^;]*;//;
+			}
+			$mid = " ... ";
+		}
+		return "$left$mid$right";
+
+	} else {
+		$str =~ m/^($endre)(.*)$/;
+		my $body = $1;
+		my $tail = $2;
+		if (length($tail) > 4) {
+			$body =~ s/&[^;]*$//;
+			$tail = " ...";
+		}
+		return "$body$tail";
 	}
-	return "$body$tail";
 }
 
 # takes the same arguments as chop_str, but also wraps a <span> around the
 # result with a title attribute if it does get chopped. Additionally, the
 # string is HTML-escaped.
 sub chop_and_escape_str {
-	my $str = shift;
-	my $len = shift;
-	my $add_len = shift || 10;
+	my ($str) = @_;
 
-	my $chopped = chop_str($str, $len, $add_len);
+	my $chopped = chop_str(@_);
 	if ($chopped eq $str) {
 		return esc_html($chopped);
 	} else {
@@ -3791,11 +3833,11 @@ sub git_search_grep_body {
 		foreach my $line (@$comment) {
 			if ($line =~ m/^(.*)($search_regexp)(.*)$/i) {
 				my ($lead, $match, $trail) = ($1, $2, $3);
-				$match = chop_str($match, 70, 5);       # in case match is very long
-				my $contextlen = int((80 - length($match))/2); # for the remainder
-				$contextlen = 30 if ($contextlen > 30); # but not too much
-				$lead  = chop_str($lead,  $contextlen, 10);
-				$trail = chop_str($trail, $contextlen, 10);
+				$match = chop_str($match, 70, 5, -cut=>'center');
+				my $contextlen = int((80 - length($match))/2);
+				$contextlen = 30 if ($contextlen > 30);
+				$lead  = chop_str($lead,  $contextlen, 10, -cut=>'left');
+				$trail = chop_str($trail, $contextlen, 10, -cut=>'right');
 
 				$lead  = esc_html($lead);
 				$match = esc_html($match);
Previous: Jakub NarebskiNext: Junio C Hamano
Message 8 of 16 in “Do not chop HTML tags in commit search result”
  1. Do not chop HTML tags in commit search resultJean-Baptiste Quenot, Feb 13, 2008
  2. Jakub NarebskiFeb 13, 2008
  3. Junio C HamanoFeb 13, 2008
  4. gitweb: Better chopping in commit search resultsJakub Narebski, Feb 22, 2008
  5. Junio C HamanoFeb 22, 2008
  6. Jakub NarebskiFeb 22, 2008
  7. Jakub NarebskiFeb 22, 2008
  8. gitweb: Option to chop at beginning and in the middle in chop_strJakub Narebski, Feb 23, 2008
  9. Junio C HamanoFeb 23, 2008
  10. Jakub NarebskiFeb 23, 2008
  11. gitweb: Option to chop at beginning and in the middle in chop_strJakub Narebski, Feb 24, 2008
  12. Junio C HamanoFeb 25, 2008
  13. gitweb: Better cutting matched string and its contextJakub Narebski, Feb 25, 2008
  14. Junio C HamanoFeb 25, 2008
  15. Karl HasselströmFeb 23, 2008
  16. Jakub NarebskiFeb 23, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.