git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH 02/18] chainlint.pl: add POSIX shell lexical analyzer

From
Eric Sunshine via GitGitGadget <gitgitgadget@gmail.com>
Date
Sep 1, 2022, 00:29 UTC
Message-ID
<c1042b9bcd94b9ecb0bf73dfbd4334b9f30ba99a.1661992197.git.gitgitgadget@gmail.com>
In-Reply-To
<pull.1322.git.git.1661992197.gitgitgadget@gmail.com>
From: Eric Sunshine <sunshine@sunshineco.com>

Begin fleshing out chainlint.pl by adding a lexical analyzer for the POSIX shell command language. The sole entry point Lexer::scan_token() returns the next token from the input. It will be called by the upcoming shell language parser.

Signed-off-by: Eric Sunshine <sunshine@sunshineco.com>
---
 t/chainlint.pl | 177 +++++++++++++++++++++++++++++++++++++++++++++++++
 1 file changed, 177 insertions(+)
diff --git a/t/chainlint.pl b/t/chainlint.pl
index e8ab95c7858..81ffbf28bf3 100755
--- a/t/chainlint.pl
+++ b/t/chainlint.pl
@@ -21,6 +21,183 @@ use Getopt::Long;
 my $show_stats;
 my $emit_all;
 
+# Lexer tokenizes POSIX shell scripts. It is roughly modeled after section 2.3
+# "Token Recognition" of POSIX chapter 2 "Shell Command Language". Although
+# similar to lexical analyzers for other languages, this one differs in a few
+# substantial ways due to quirks of the shell command language.
+#
+# For instance, in many languages, newline is just whitespace like space or
+# TAB, but in shell a newline is a command separator, thus a distinct lexical
+# token. A newline is significant and returned as a distinct token even at the
+# end of a shell comment.
+#
+# In other languages, `1+2` would typically be scanned as three tokens
+# (`1`, `+`, and `2`), but in shell it is a single token. However, the similar
+# `1 + 2`, which embeds whitepace, is scanned as three token in shell, as well.
+# In shell, several characters with special meaning lose that meaning when not
+# surrounded by whitespace. For instance, the negation operator `!` is special
+# when standing alone surrounded by whitespace; whereas in `foo!uucp` it is
+# just a plain character in the longer token "foo!uucp". In many other
+# languages, `"string"/foo:'string'` might be scanned as five tokens ("string",
+# `/`, `foo`, `:`, and 'string'), but in shell, it is just a single token.
+#
+# The lexical analyzer for the shell command language is also somewhat unusual
+# in that it recursively invokes the parser to handle the body of `$(...)`
+# expressions which can contain arbitrary shell code. Such expressions may be
+# encountered both inside and outside of double-quoted strings.
+#
+# The lexical analyzer is responsible for consuming shell here-doc bodies which
+# extend from the line following a `<<TAG` operator until a line consisting
+# solely of `TAG`. Here-doc consumption begins when a newline is encountered.
+# It is legal for multiple here-doc `<<TAG` operators to be present on a single
+# line, in which case their bodies must be present one following the next, and
+# are consumed in the (left-to-right) order the `<<TAG` operators appear on the
+# line. A special complication is that the bodies of all here-docs must be
+# consumed when the newline is encountered even if the parse context depth has
+# changed. For instance, in `cat <<A && x=$(cat <<B &&\n`, bodies of here-docs
+# "A" and "B" must be consumed even though "A" was introduced outside the
+# recursive parse context in which "B" was introduced and in which the newline
+# is encountered.
+package Lexer;
+
+sub new {
+	my ($class, $parser, $s) = @_;
+	bless {
+		parser => $parser,
+		buff => $s,
+		heretags => []
+	} => $class;
+}
+
+sub scan_heredoc_tag {
+	my $self = shift @_;
+	${$self->{buff}} =~ /\G(-?)/gc;
+	my $indented = $1;
+	my $tag = $self->scan_token();
+	$tag =~ s/['"\\]//g;
+	push(@{$self->{heretags}}, $indented ? "\t$tag" : "$tag");
+	return "<<$indented$tag";
+}
+
+sub scan_op {
+	my ($self, $c) = @_;
+	my $b = $self->{buff};
+	return $c unless $$b =~ /\G(.)/sgc;
+	my $cc = $c . $1;
+	return scan_heredoc_tag($self) if $cc eq '<<';
+	return $cc if $cc =~ /^(?:&&|\|\||>>|;;|<&|>&|<>|>\|)$/;
+	pos($$b)--;
+	return $c;
+}
+
+sub scan_sqstring {
+	my $self = shift @_;
+	${$self->{buff}} =~ /\G([^']*'|.*\z)/sgc;
+	return "'" . $1;
+}
+
+sub scan_dqstring {
+	my $self = shift @_;
+	my $b = $self->{buff};
+	my $s = '"';
+	while (1) {
+		# slurp up non-special characters
+		$s .= $1 if $$b =~ /\G([^"\$\\]+)/gc;
+		# handle special characters
+		last unless $$b =~ /\G(.)/sgc;
+		my $c = $1;
+		$s .= '"', last if $c eq '"';
+		$s .= '$' . $self->scan_dollar(), next if $c eq '$';
+		if ($c eq '\\') {
+			$s .= '\\', last unless $$b =~ /\G(.)/sgc;
+			$c = $1;
+			next if $c eq "\n"; # line splice
+			# backslash escapes only $, `, ", \ in dq-string
+			$s .= '\\' unless $c =~ /^[\$`"\\]$/;
+			$s .= $c;
+			next;
+		}
+		die("internal error scanning dq-string '$c'\n");
+	}
+	return $s;
+}
+
+sub scan_balanced {
+	my ($self, $c1, $c2) = @_;
+	my $b = $self->{buff};
+	my $depth = 1;
+	my $s = $c1;
+	while ($$b =~ /\G([^\Q$c1$c2\E]*(?:[\Q$c1$c2\E]|\z))/gc) {
+		$s .= $1;
+		$depth++, next if $s =~ /\Q$c1\E$/;
+		$depth--;
+		last if $depth == 0;
+	}
+	return $s;
+}
+
+sub scan_subst {
+	my $self = shift @_;
+	my @tokens = $self->{parser}->parse(qr/^\)$/);
+	$self->{parser}->next_token(); # closing ")"
+	return @tokens;
+}
+
+sub scan_dollar {
+	my $self = shift @_;
+	my $b = $self->{buff};
+	return $self->scan_balanced('(', ')') if $$b =~ /\G\((?=\()/gc; # $((...))
+	return '(' . join(' ', $self->scan_subst()) . ')' if $$b =~ /\G\(/gc; # $(...)
+	return $self->scan_balanced('{', '}') if $$b =~ /\G\{/gc; # ${...}
+	return $1 if $$b =~ /\G(\w+)/gc; # $var
+	return $1 if $$b =~ /\G([@*#?$!0-9-])/gc; # $*, $1, $$, etc.
+	return '';
+}
+
+sub swallow_heredocs {
+	my $self = shift @_;
+	my $b = $self->{buff};
+	my $tags = $self->{heretags};
+	while (my $tag = shift @$tags) {
+		my $indent = $tag =~ s/^\t// ? '\\s*' : '';
+		$$b =~ /(?:\G|\n)$indent\Q$tag\E(?:\n|\z)/gc;
+	}
+}
+
+sub scan_token {
+	my $self = shift @_;
+	my $b = $self->{buff};
+	my $token = '';
+RESTART:
+	$$b =~ /\G[ \t]+/gc; # skip whitespace (but not newline)
+	return "\n" if $$b =~ /\G#[^\n]*(?:\n|\z)/gc; # comment
+	while (1) {
+		# slurp up non-special characters
+		$token .= $1 if $$b =~ /\G([^\\;&|<>(){}'"\$\s]+)/gc;
+		# handle special characters
+		last unless $$b =~ /\G(.)/sgc;
+		my $c = $1;
+		last if $c =~ /^[ \t]$/; # whitespace ends token
+		pos($$b)--, last if length($token) && $c =~ /^[;&|<>(){}\n]$/;
+		$token .= $self->scan_sqstring(), next if $c eq "'";
+		$token .= $self->scan_dqstring(), next if $c eq '"';
+		$token .= $c . $self->scan_dollar(), next if $c eq '$';
+		$self->swallow_heredocs(), $token = $c, last if $c eq "\n";
+		$token = $self->scan_op($c), last if $c =~ /^[;&|<>]$/;
+		$token = $c, last if $c =~ /^[(){}]$/;
+		if ($c eq '\\') {
+			$token .= '\\', last unless $$b =~ /\G(.)/sgc;
+			$c = $1;
+			next if $c eq "\n" && length($token); # line splice
+			goto RESTART if $c eq "\n"; # line splice
+			$token .= '\\' . $c;
+			next;
+		}
+		die("internal error scanning character '$c'\n");
+	}
+	return length($token) ? $token : undef;
+}
+
 package ScriptParser;
 
 sub new {
-- 
gitgitgadget
Previous: Eric SunshineNext: Ævar Arnfjörð Bjarmason
Message 5 of 51 in “make test "linting" more comprehensive”
  1. 00/18 make test "linting" more comprehensiveEric Sunshine via GitGitGadget, Sep 1, 2022
  2. 01/18 t: add skeleton chainlint.plEric Sunshine via GitGitGadget, Sep 1, 2022
  3. Ævar Arnfjörð BjarmasonSep 1, 2022
  4. Eric SunshineSep 2, 2022
  5. 02/18 chainlint.pl: add POSIX shell lexical analyzerEric Sunshine via GitGitGadget, Sep 1, 2022
  6. Ævar Arnfjörð BjarmasonSep 1, 2022
  7. Eric SunshineSep 3, 2022
  8. 04/18 chainlint.pl: add parser to validate testsEric Sunshine via GitGitGadget, Sep 1, 2022
  9. 03/18 chainlint.pl: add POSIX shell parserEric Sunshine via GitGitGadget, Sep 1, 2022
  10. 06/18 chainlint.pl: validate test scripts in parallelEric Sunshine via GitGitGadget, Sep 1, 2022
  11. Ævar Arnfjörð BjarmasonSep 1, 2022
  12. Eric SunshineSep 3, 2022
  13. Eric WongSep 6, 2022
  14. Eric SunshineSep 6, 2022
  15. Jeff KingSep 6, 2022
  16. Eric SunshineNov 21, 2022
  17. Ævar Arnfjörð BjarmasonNov 21, 2022
  18. Eric SunshineNov 21, 2022
  19. Ævar Arnfjörð BjarmasonNov 21, 2022
  20. Eric SunshineNov 21, 2022
  21. Jeff KingNov 21, 2022
  22. Eric SunshineNov 21, 2022
  23. Eric SunshineNov 21, 2022
  24. Jeff KingNov 21, 2022
  25. Eric SunshineNov 21, 2022
  26. Jeff KingNov 21, 2022
  27. Ævar Arnfjörð BjarmasonNov 22, 2022
  28. 05/18 chainlint.pl: add parser to identify test definitionsEric Sunshine via GitGitGadget, Sep 1, 2022
  29. 07/18 chainlint.pl: don't require `return|exit|continue` to end with `&&`Eric Sunshine via GitGitGadget, Sep 1, 2022
  30. 10/18 chainlint.pl: don't flag broken &&-chain if `$?` handled explicitlyEric Sunshine via GitGitGadget, Sep 1, 2022
  31. 12/18 chainlint.pl: complain about loops lacking explicit failure handlingEric Sunshine via GitGitGadget, Sep 1, 2022
  32. 09/18 chainlint.pl: don't require `&` background command to end with `&&`Eric Sunshine via GitGitGadget, Sep 1, 2022
  33. 08/18 t/Makefile: apply chainlint.pl to existing self-testsEric Sunshine via GitGitGadget, Sep 1, 2022
  34. 11/18 chainlint.pl: don't flag broken &&-chain if failure indicated explicitlyEric Sunshine via GitGitGadget, Sep 1, 2022
  35. 13/18 chainlint.pl: allow `|| echo` to signal failure upstream of a pipeEric Sunshine via GitGitGadget, Sep 1, 2022
  36. 14/18 t/chainlint: add more chainlint.pl self-testsEric Sunshine via GitGitGadget, Sep 1, 2022
  37. 15/18 test-lib: retire "lint harder" optimization hackEric Sunshine via GitGitGadget, Sep 1, 2022
  38. 16/18 test-lib: replace chainlint.sed with chainlint.plEric Sunshine via GitGitGadget, Sep 1, 2022
  39. Elijah NewrenSep 3, 2022
  40. Eric SunshineSep 3, 2022
  41. 18/18 t: retire unused chainlint.sedEric Sunshine via GitGitGadget, Sep 1, 2022
  42. Johannes SchindelinSep 2, 2022
  43. Eric SunshineSep 2, 2022
  44. Jeff KingSep 2, 2022
  45. Junio C HamanoSep 2, 2022
  46. 17/18 t/Makefile: teach `make test` and `make prove` to run chainlint.plEric Sunshine via GitGitGadget, Sep 1, 2022
  47. Jeff KingSep 11, 2022
  48. Eric SunshineSep 11, 2022
  49. Jeff KingSep 11, 2022
  50. Eric SunshineSep 12, 2022
  51. Jeff KingSep 13, 2022

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.