{"thread":{"id":"55399","subject":"[GSOC][PATCH] userdiff: add support for Scheme","startedAt":"2021-03-27T17:41:32Z","lastAt":"2021-04-12T23:04:10Z","messageCount":35,"participants":["Atharva Raykar","Junio C Hamano","Johannes Sixt","Ævar Arnfjörð Bjarmason","Phillip Wood"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"420302","messageId":"20210327173938.59391-1-raykar.ath@gmail.com","threadId":"55399","inReplyTo":null,"subject":"[GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-27T17:39:38Z","receivedAt":"2021-03-27T17:41:32Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"Add a diff driver for Scheme (R5RS and R6RS) which\nrecognizes top level and local `define` forms,\nwhether it is a function definition, binding, syntax\ndefinition or a user-defined `define-xyzzy` form.\n\nThe rationale for picking `define` forms for the\nhunk headers is because it is usually the only\nsignificant form for defining the structure of the\nprogram, and it is a common pattern for schemers to\nhave local function definitions to hide their\nvisibility, so it is not only the top level\n`define`'s that are of interest. Schemers also\nextend the language with macros to provide their\nown define forms (for example, something like a\n`define-test-suite`) which is also captured in the\nhunk header.\n\nThe word regex is a best-effort attempt to conform\nto R6RS[1] valid identifiers, symbols and numbers.\n\n[1] http://www.r6rs.org/final/html/r6rs/r6rs-Z-H-7.html#node_chap_4\n\nSigned-off-by: Atharva Raykar <raykar.ath@gmail.com>\n---\n\nHi, first-time contributor here, I wanted to have a go at this as\na microproject.\n\nA few things I had to consider:\n - Going through the mailing list, there have already been two other\n   patches that are for lispy languages that have taken slightly\n   different approaches: Elisp[1] and Clojure[2]. Would it make any\n   sense to have a single userdiff driver for lisp that just captures\n   all top level forms in the hunk? I personally felt it's better to\n   differentiate the drivers for each language, as they have different\n   constructs.\n\n - It was hard to decide exactly which forms should appear on the hunk\n   headers. Having programmed in a Scheme before, I went with the forms I\n   would have liked to see when looking at git diffs, which would be the\n   nearest `define` along with `define-syntax` and other define forms that\n   are created as user-defined macros. I am willing to ask around in\n   certain active scheme communities for some kind of consensus, but there\n   is no single large consolidated group of schemers (the closest is\n   probably comp.lang.scheme?).\n\n - By best-effort attempt at the wordregex, I mean that it is a little\n   more permissive than it has to be, as it accepts a few words that are\n   technically invalid in Scheme.\n   Making it handle all cases like numbers and identifiers with separate\n   regexen would be greatly complicated (Eg: #x#e10.2f3 is a valid number\n   but #x#f10.2e3 is not; 10t1 is a valid identifier, but 10s1 is a number\n   -- my wordregex just clubs all of these into a generic 'word match' which\n   trades of granularity for simplicity, and it usually does the right thing).\n\n[1] http://public-inbox.org/git/20210213192447.6114-1-git@adamspiers.org/\n[2] http://public-inbox.org/git/pull.902.git.1615667191368.gitgitgadget@gmail.com/\n\n Documentation/gitattributes.txt    | 2 ++\n t/t4018-diff-funcname.sh           | 1 +\n t/t4018/scheme-define-syntax       | 8 ++++++++\n t/t4018/scheme-local-define        | 4 ++++\n t/t4018/scheme-top-level-define    | 4 ++++\n t/t4018/scheme-user-defined-define | 6 ++++++\n t/t4034-diff-words.sh              | 1 +\n t/t4034/scheme/expect              | 9 +++++++++\n t/t4034/scheme/post                | 4 ++++\n t/t4034/scheme/pre                 | 4 ++++\n userdiff.c                         | 8 ++++++++\n 11 files changed, 51 insertions(+)\n create mode 100644 t/t4018/scheme-define-syntax\n create mode 100644 t/t4018/scheme-local-define\n create mode 100644 t/t4018/scheme-top-level-define\n create mode 100644 t/t4018/scheme-user-defined-define\n create mode 100644 t/t4034/scheme/expect\n create mode 100644 t/t4034/scheme/post\n create mode 100644 t/t4034/scheme/pre\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 0a60472bb5..cfcfa800c2 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -845,6 +845,8 @@ patterns are available:\n \n - `rust` suitable for source code in the Rust language.\n \n+- `scheme` suitable for source code in the Scheme language.\n+\n - `tex` suitable for source code for LaTeX documents.\n \n \ndiff --git a/t/t4018-diff-funcname.sh b/t/t4018-diff-funcname.sh\nindex 9675bc17db..823ea96acb 100755\n--- a/t/t4018-diff-funcname.sh\n+++ b/t/t4018-diff-funcname.sh\n@@ -48,6 +48,7 @@ diffpatterns=\"\n \tpython\n \truby\n \trust\n+\tscheme\n \ttex\n \tcustom1\n \tcustom2\ndiff --git a/t/t4018/scheme-define-syntax b/t/t4018/scheme-define-syntax\nnew file mode 100644\nindex 0000000000..603b99cea4\n--- /dev/null\n+++ b/t/t4018/scheme-define-syntax\n@@ -0,0 +1,8 @@\n+(define-syntax define-test-suite RIGHT\n+  (syntax-rules ()\n+    ((_ suite-name (name test) ChangeMe ...)\n+     (define suite-name\n+       (let ((tests\n+              `((name . ,test) ...)))\n+         (lambda ()\n+           (ChangeMe 'suite-name tests)))))))\n\\ No newline at end of file\ndiff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\nnew file mode 100644\nindex 0000000000..90e75dcce8\n--- /dev/null\n+++ b/t/t4018/scheme-local-define\n@@ -0,0 +1,4 @@\n+(define (higher-order)\n+  (define local-function RIGHT\n+    (lambda (x)\n+     (car \"this is\" \"ChangeMe\"))))\n\\ No newline at end of file\ndiff --git a/t/t4018/scheme-top-level-define b/t/t4018/scheme-top-level-define\nnew file mode 100644\nindex 0000000000..03acdc628d\n--- /dev/null\n+++ b/t/t4018/scheme-top-level-define\n@@ -0,0 +1,4 @@\n+(define (some-func x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\n\\ No newline at end of file\ndiff --git a/t/t4018/scheme-user-defined-define b/t/t4018/scheme-user-defined-define\nnew file mode 100644\nindex 0000000000..401093bac3\n--- /dev/null\n+++ b/t/t4018/scheme-user-defined-define\n@@ -0,0 +1,6 @@\n+(define-test-suite record-case-tests RIGHT\n+  (record-case-1 (lambda (fail)\n+                   (let ((a (make-foo 1 2)))\n+                     (record-case a\n+                       ((bar x) (ChangeMe))\n+                       ((foo a b) (+ a b)))))))\n\\ No newline at end of file\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 56f1e62a97..ee7721ab91 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -325,6 +325,7 @@ test_language_driver perl\n test_language_driver php\n test_language_driver python\n test_language_driver ruby\n+test_language_driver scheme\n test_language_driver tex\n \n test_expect_success 'word-diff with diff.sbe' '\ndiff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\nnew file mode 100644\nindex 0000000000..eed21e803c\n--- /dev/null\n+++ b/t/t4034/scheme/expect\n@@ -0,0 +1,9 @@\n+<BOLD>diff --git a/pre b/post<RESET>\n+<BOLD>index 6a5efba..7c4a6b4 100644<RESET>\n+<BOLD>--- a/pre<RESET>\n+<BOLD>+++ b/post<RESET>\n+<CYAN>@@ -1,4 +1,4 @@<RESET>\n+(define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n+  ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n+  (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n+    (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\ndiff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\nnew file mode 100644\nindex 0000000000..7c4a6b4f3d\n--- /dev/null\n+++ b/t/t4034/scheme/post\n@@ -0,0 +1,4 @@\n+(define (my-func first second)\n+  ; This is a (moderately) cool function.\n+  (let ((c (add1 first)))\n+    (format \"one more than the total is %d\" (+ c second))))\n\\ No newline at end of file\ndiff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\nnew file mode 100644\nindex 0000000000..6a5efbae61\n--- /dev/null\n+++ b/t/t4034/scheme/pre\n@@ -0,0 +1,4 @@\n+(define (myfunc a b)\n+  ; This is a really cool function.\n+  (let ((c (+ a b)))\n+    (format \"one more than the total is %d\" (add1 c))))\n\\ No newline at end of file\ndiff --git a/userdiff.c b/userdiff.c\nindex 3f81a2261c..c51a8c98ba 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -191,6 +191,14 @@ PATTERNS(\"rust\",\n \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n+PATTERNS(\"scheme\",\n+         \"^[\\t ]*(\\\\(define-?.*)$\",\n+         /* \n+          * Scheme allows symbol names to have any character,\n+          * as long as it is not a form of a parenthesis.\n+          * The spaces must be escaped.\n+          */\n+         \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n-- \n2.31.0\n\n"},{"id":"420313","messageId":"xmqq5z1cqki7.fsf@gitster.g","threadId":"55399","inReplyTo":"20210327173938.59391-1-raykar.ath@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-03-27T22:50:24Z","receivedAt":"2021-03-27T22:51:39Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Atharva Raykar <raykar.ath@gmail.com> writes:\n\n> diff --git a/t/t4018/scheme-define-syntax b/t/t4018/scheme-define-syntax\n> new file mode 100644\n> index 0000000000..603b99cea4\n> --- /dev/null\n> +++ b/t/t4018/scheme-define-syntax\n> @@ -0,0 +1,8 @@\n> +(define-syntax define-test-suite RIGHT\n> +  (syntax-rules ()\n> +    ((_ suite-name (name test) ChangeMe ...)\n> +     (define suite-name\n> +       (let ((tests\n> +              `((name . ,test) ...)))\n> +         (lambda ()\n> +           (ChangeMe 'suite-name tests)))))))\n> \\ No newline at end of file\n\nIs there a good reason to leave the final line incomplete?  If there\nisn't, complete it (applies to other newly-created files in the patch).\n\n> diff --git a/userdiff.c b/userdiff.c\n> index 3f81a2261c..c51a8c98ba 100644\n> --- a/userdiff.c\n> +++ b/userdiff.c\n> @@ -191,6 +191,14 @@ PATTERNS(\"rust\",\n>  \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n>  \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>  \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n> +PATTERNS(\"scheme\",\n> +         \"^[\\t ]*(\\\\(define-?.*)$\",\n\nDidn't \"git diff HEAD\" before committing (or \"git show\") highlighted\nthese whitespace errors?\n\n.git/rebase-apply/patch:183: indent with spaces.\n         \"^[\\t ]*(\\\\(define-?.*)$\",\n.git/rebase-apply/patch:184: trailing whitespace, indent with spaces.\n         /* \n.git/rebase-apply/patch:185: indent with spaces.\n          * Scheme allows symbol names to have any character,\n.git/rebase-apply/patch:186: indent with spaces.\n          * as long as it is not a form of a parenthesis.\n.git/rebase-apply/patch:187: indent with spaces.\n          * The spaces must be escaped.\nwarning: squelched 2 whitespace errors\nwarning: 7 lines applied after fixing whitespace errors.\n\n\n> +         /* \n> +          * Scheme allows symbol names to have any character,\n> +          * as long as it is not a form of a parenthesis.\n> +          * The spaces must be escaped.\n> +          */\n> +         \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n\nOne or more \"dot or anything other than SP or parentheses\"?  But\na dot \".\" is neither a space or any {bra-ce} letter, so would the\nabove be equivalent to\n\n\t\"[^][()\\\\{\\\\} \\t]+\"\n\nI wonder...\n\nI am also trying to figure out what you wanted to achieve by\nmentioning \"The spaces must be escaped.\".  Did you mean something\nlike (string->symbol \"a symbol with SP in it\") is a symbol?  Even\nso, I cannot quite guess the significance of that fact wrt the\nregexp you added here?\n\nAs we are trying to catch program identifiers (symbols in scheme)\nand numeric literals, treating any group of non-whitespace letters\nthat is delimited by one or more whitespaces as a \"word\" would be a\ngood first-order approximation, but in addition, as can be seen in\nan example like (a(b(c))), parentheses can also serve as such \"word\ndelimiters\" in addition to whitespaces.  So the regexp given above\nmakes sense to me from that angle, especially if you do not limit\nthe whitespace to only SP, but include HT (\\t) as well.  But was\nthat how you came up with the regexp?\n\nThanks.\n\n>  PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n>  \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n>  PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n"},{"id":"420319","messageId":"xmqq1rc0qjn1.fsf@gitster.g","threadId":"55399","inReplyTo":"xmqq5z1cqki7.fsf@gitster.g","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-03-27T23:09:06Z","receivedAt":"2021-03-27T23:09:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Atharva Raykar <raykar.ath@gmail.com> writes:\n> ...\n>> +           (ChangeMe 'suite-name tests)))))))\n>> \\ No newline at end of file\n>\n> Is there a good reason to leave the final line incomplete?  ...\n> I am also trying to figure out what you wanted to achieve ...\n\nTaking all of them together, here is what I hope you may agree as\nits improved version.  The only differences from what you posted are\ncorrections to all the \"\\ No newline at end of file\" and the simplification\nof the pattern (remove \"a dot\" from the alternative and add \\t next\nto SP).  Without changes, the new tests still pass so ... ;-)\n\n    diff --git c/userdiff.c w/userdiff.c\n    index 5fd0eb31ec..685fe712aa 100644\n    --- c/userdiff.c\n    +++ w/userdiff.c\n    @@ -193,12 +193,8 @@ PATTERNS(\"rust\",\n             \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n     PATTERNS(\"scheme\",\n             \"^[\\t ]*(\\\\(define-?.*)$\",\n    -\t /*\n    -\t  * Scheme allows symbol names to have any character,\n    -\t  * as long as it is not a form of a parenthesis.\n    -\t  * The spaces must be escaped.\n    -\t  */\n    -\t \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n    +\t /* whitespace separated tokens, but parentheses also can delimit words */\n    +\t \"([^][)(\\\\}\\\\{ \\t])+\"),\n     PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n             \"[={}\\\"]|[^={}\\\" \\t]+\"),\n     PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n\n----- >8 ---------- >8 ---------- >8 ---------- >8 ---------- >8 -----\nFrom: Atharva Raykar <raykar.ath@gmail.com>\nDate: Sat, 27 Mar 2021 23:09:38 +0530\nSubject: [PATCH] userdiff: add support for Scheme\n\nAdd a diff driver for Scheme (R5RS and R6RS) which\nrecognizes top level and local `define` forms,\nwhether it is a function definition, binding, syntax\ndefinition or a user-defined `define-xyzzy` form.\n\nThe rationale for picking `define` forms for the\nhunk headers is because it is usually the only\nsignificant form for defining the structure of the\nprogram, and it is a common pattern for schemers to\nhave local function definitions to hide their\nvisibility, so it is not only the top level\n`define`'s that are of interest. Schemers also\nextend the language with macros to provide their\nown define forms (for example, something like a\n`define-test-suite`) which is also captured in the\nhunk header.\n\nSince the identifier syntax is quite forgiving, we start\nour word regexp from \"words delimited by whitespaces\" and\nthen loosen to include various forms of parentheses characters\nto word-delimiters.\n\nSigned-off-by: Atharva Raykar <raykar.ath@gmail.com>\n[jc: simplified word regex and its explanation; fixed whitespace errors]\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n Documentation/gitattributes.txt    | 2 ++\n t/t4018-diff-funcname.sh           | 1 +\n t/t4018/scheme-define-syntax       | 8 ++++++++\n t/t4018/scheme-local-define        | 4 ++++\n t/t4018/scheme-top-level-define    | 4 ++++\n t/t4018/scheme-user-defined-define | 6 ++++++\n t/t4034-diff-words.sh              | 1 +\n t/t4034/scheme/expect              | 9 +++++++++\n t/t4034/scheme/post                | 4 ++++\n t/t4034/scheme/pre                 | 4 ++++\n userdiff.c                         | 4 ++++\n 11 files changed, 47 insertions(+)\n create mode 100644 t/t4018/scheme-define-syntax\n create mode 100644 t/t4018/scheme-local-define\n create mode 100644 t/t4018/scheme-top-level-define\n create mode 100644 t/t4018/scheme-user-defined-define\n create mode 100644 t/t4034/scheme/expect\n create mode 100644 t/t4034/scheme/post\n create mode 100644 t/t4034/scheme/pre\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 0a60472bb5..cfcfa800c2 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -845,6 +845,8 @@ patterns are available:\n \n - `rust` suitable for source code in the Rust language.\n \n+- `scheme` suitable for source code in the Scheme language.\n+\n - `tex` suitable for source code for LaTeX documents.\n \n \ndiff --git a/t/t4018-diff-funcname.sh b/t/t4018-diff-funcname.sh\nindex 9675bc17db..823ea96acb 100755\n--- a/t/t4018-diff-funcname.sh\n+++ b/t/t4018-diff-funcname.sh\n@@ -48,6 +48,7 @@ diffpatterns=\"\n \tpython\n \truby\n \trust\n+\tscheme\n \ttex\n \tcustom1\n \tcustom2\ndiff --git a/t/t4018/scheme-define-syntax b/t/t4018/scheme-define-syntax\nnew file mode 100644\nindex 0000000000..33fa50c844\n--- /dev/null\n+++ b/t/t4018/scheme-define-syntax\n@@ -0,0 +1,8 @@\n+(define-syntax define-test-suite RIGHT\n+  (syntax-rules ()\n+    ((_ suite-name (name test) ChangeMe ...)\n+     (define suite-name\n+       (let ((tests\n+              `((name . ,test) ...)))\n+         (lambda ()\n+           (ChangeMe 'suite-name tests)))))))\ndiff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\nnew file mode 100644\nindex 0000000000..bc6d8aebbe\n--- /dev/null\n+++ b/t/t4018/scheme-local-define\n@@ -0,0 +1,4 @@\n+(define (higher-order)\n+  (define local-function RIGHT\n+    (lambda (x)\n+     (car \"this is\" \"ChangeMe\"))))\ndiff --git a/t/t4018/scheme-top-level-define b/t/t4018/scheme-top-level-define\nnew file mode 100644\nindex 0000000000..624743c22b\n--- /dev/null\n+++ b/t/t4018/scheme-top-level-define\n@@ -0,0 +1,4 @@\n+(define (some-func x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\ndiff --git a/t/t4018/scheme-user-defined-define b/t/t4018/scheme-user-defined-define\nnew file mode 100644\nindex 0000000000..70e403c5e2\n--- /dev/null\n+++ b/t/t4018/scheme-user-defined-define\n@@ -0,0 +1,6 @@\n+(define-test-suite record-case-tests RIGHT\n+  (record-case-1 (lambda (fail)\n+                   (let ((a (make-foo 1 2)))\n+                     (record-case a\n+                       ((bar x) (ChangeMe))\n+                       ((foo a b) (+ a b)))))))\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 56f1e62a97..ee7721ab91 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -325,6 +325,7 @@ test_language_driver perl\n test_language_driver php\n test_language_driver python\n test_language_driver ruby\n+test_language_driver scheme\n test_language_driver tex\n \n test_expect_success 'word-diff with diff.sbe' '\ndiff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\nnew file mode 100644\nindex 0000000000..eed21e803c\n--- /dev/null\n+++ b/t/t4034/scheme/expect\n@@ -0,0 +1,9 @@\n+<BOLD>diff --git a/pre b/post<RESET>\n+<BOLD>index 6a5efba..7c4a6b4 100644<RESET>\n+<BOLD>--- a/pre<RESET>\n+<BOLD>+++ b/post<RESET>\n+<CYAN>@@ -1,4 +1,4 @@<RESET>\n+(define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n+  ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n+  (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n+    (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\ndiff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\nnew file mode 100644\nindex 0000000000..28f59c6584\n--- /dev/null\n+++ b/t/t4034/scheme/post\n@@ -0,0 +1,4 @@\n+(define (my-func first second)\n+  ; This is a (moderately) cool function.\n+  (let ((c (add1 first)))\n+    (format \"one more than the total is %d\" (+ c second))))\ndiff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\nnew file mode 100644\nindex 0000000000..4bd0069493\n--- /dev/null\n+++ b/t/t4034/scheme/pre\n@@ -0,0 +1,4 @@\n+(define (myfunc a b)\n+  ; This is a really cool function.\n+  (let ((c (+ a b)))\n+    (format \"one more than the total is %d\" (add1 c))))\ndiff --git a/userdiff.c b/userdiff.c\nindex 3f81a2261c..685fe712aa 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -191,6 +191,10 @@ PATTERNS(\"rust\",\n \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n+PATTERNS(\"scheme\",\n+\t \"^[\\t ]*(\\\\(define-?.*)$\",\n+\t /* whitespace separated tokens, but parentheses also can delimit words */\n+\t \"([^][)(\\\\}\\\\{ \\t])+\"),\n PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n-- \n2.31.1-255-g3df2b433e7\n\n"},{"id":"420320","messageId":"3def82fd-71a7-3ad9-0fa2-48598bfd3313@kdbg.org","threadId":"55399","inReplyTo":"20210327173938.59391-1-raykar.ath@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2021-03-27T23:46:11Z","receivedAt":"2021-03-27T23:59:57Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 27.03.21 um 18:39 schrieb Atharva Raykar:\n>  - By best-effort attempt at the wordregex, I mean that it is a little\n>    more permissive than it has to be, as it accepts a few words that are\n>    technically invalid in Scheme.\n>    Making it handle all cases like numbers and identifiers with separate\n>    regexen would be greatly complicated (Eg: #x#e10.2f3 is a valid number\n>    but #x#f10.2e3 is not; 10t1 is a valid identifier, but 10s1 is a number\n>    -- my wordregex just clubs all of these into a generic 'word match' which\n>    trades of granularity for simplicity, and it usually does the right thing).\n\nIt is ok to have regex that capture tokens that are not valid. A\nuserdiff driver can assume that it operates only text that is valid in\nthe language.\n\n> diff --git a/t/t4018/scheme-define-syntax b/t/t4018/scheme-define-syntax\n> new file mode 100644\n> index 0000000000..603b99cea4\n> --- /dev/null\n> +++ b/t/t4018/scheme-define-syntax\n> @@ -0,0 +1,8 @@\n> +(define-syntax define-test-suite RIGHT\n> +  (syntax-rules ()\n> +    ((_ suite-name (name test) ChangeMe ...)\n> +     (define suite-name\n\nThis test is suspicious. Notice the \"ChangeMe\" above? That is sufficient\nto let the test case succeed. The \"ChangeMe\" in the last line below\nshould be the only one.\n\nBut then there is this indented '(define' that is not marked as RIGHT,\nand I wonder how is it different from...\n\n> +       (let ((tests\n> +              `((name . ,test) ...)))\n> +         (lambda ()\n> +           (ChangeMe 'suite-name tests)))))))\n> \\ No newline at end of file\n> diff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\n> new file mode 100644\n> index 0000000000..90e75dcce8\n> --- /dev/null\n> +++ b/t/t4018/scheme-local-define\n> @@ -0,0 +1,4 @@\n> +(define (higher-order)\n> +  (define local-function RIGHT\n\n... this one, which is also indented and *is* marked as RIGHT.\n\nBTW, it's good to see test cases for both indented and not-indented\ntrigger lines.\n\n> +    (lambda (x)\n> +     (car \"this is\" \"ChangeMe\"))))\n> \\ No newline at end of file\n\n> diff --git a/userdiff.c b/userdiff.c\n> index 3f81a2261c..c51a8c98ba 100644\n> --- a/userdiff.c\n> +++ b/userdiff.c\n> @@ -191,6 +191,14 @@ PATTERNS(\"rust\",\n>  \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n>  \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>  \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n> +PATTERNS(\"scheme\",\n> +         \"^[\\t ]*(\\\\(define-?.*)$\",\n\nThis \"optional hyphen followed by anything\" in the regex is strange.\nWouldn't that also capture a line that looks like, e.g.,\n\n    (defined-foo bar)\n\nPerhaps we want \"define[- \\t].*\" in the regex?\n\n> +         /* \n> +          * Scheme allows symbol names to have any character,\n> +          * as long as it is not a form of a parenthesis.\n> +          * The spaces must be escaped.\n> +          */\n> +         \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n>  PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n>  \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n>  PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n> \n\n-- Hannes\n"},{"id":"420356","messageId":"87blb4nf2n.fsf@evledraar.gmail.com","threadId":"55399","inReplyTo":"xmqq1rc0qjn1.fsf@gitster.g","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-03-28T03:16:00Z","receivedAt":"2021-03-28T03:26:06Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sun, Mar 28 2021, Junio C Hamano wrote:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> Atharva Raykar <raykar.ath@gmail.com> writes:\n>> ...\n>>> +           (ChangeMe 'suite-name tests)))))))\n>>> \\ No newline at end of file\n>>\n>> Is there a good reason to leave the final line incomplete?  ...\n>> I am also trying to figure out what you wanted to achieve ...\n>\n> Taking all of them together, here is what I hope you may agree as\n> its improved version.  The only differences from what you posted are\n> corrections to all the \"\\ No newline at end of file\" and the simplification\n> of the pattern (remove \"a dot\" from the alternative and add \\t next\n> to SP).  Without changes, the new tests still pass so ... ;-)\n>\n>     diff --git c/userdiff.c w/userdiff.c\n>     index 5fd0eb31ec..685fe712aa 100644\n>     --- c/userdiff.c\n>     +++ w/userdiff.c\n>     @@ -193,12 +193,8 @@ PATTERNS(\"rust\",\n>              \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n>      PATTERNS(\"scheme\",\n>              \"^[\\t ]*(\\\\(define-?.*)$\",\n>     -\t /*\n>     -\t  * Scheme allows symbol names to have any character,\n>     -\t  * as long as it is not a form of a parenthesis.\n>     -\t  * The spaces must be escaped.\n>     -\t  */\n>     -\t \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n>     +\t /* whitespace separated tokens, but parentheses also can delimit words */\n>     +\t \"([^][)(\\\\}\\\\{ \\t])+\"),\n>      PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n>              \"[={}\\\"]|[^={}\\\" \\t]+\"),\n>      PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n>\n> ----- >8 ---------- >8 ---------- >8 ---------- >8 ---------- >8 -----\n> From: Atharva Raykar <raykar.ath@gmail.com>\n> Date: Sat, 27 Mar 2021 23:09:38 +0530\n> Subject: [PATCH] userdiff: add support for Scheme\n>\n> Add a diff driver for Scheme (R5RS and R6RS) which\n> recognizes top level and local `define` forms,\n> whether it is a function definition, binding, syntax\n> definition or a user-defined `define-xyzzy` form.\n>\n> The rationale for picking `define` forms for the\n> hunk headers is because it is usually the only\n> significant form for defining the structure of the\n> program, and it is a common pattern for schemers to\n> have local function definitions to hide their\n> visibility, so it is not only the top level\n> `define`'s that are of interest. Schemers also\n> extend the language with macros to provide their\n> own define forms (for example, something like a\n> `define-test-suite`) which is also captured in the\n> hunk header.\n>\n> Since the identifier syntax is quite forgiving, we start\n> our word regexp from \"words delimited by whitespaces\" and\n> then loosen to include various forms of parentheses characters\n> to word-delimiters.\n>\n> Signed-off-by: Atharva Raykar <raykar.ath@gmail.com>\n> [jc: simplified word regex and its explanation; fixed whitespace errors]\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>  Documentation/gitattributes.txt    | 2 ++\n>  t/t4018-diff-funcname.sh           | 1 +\n>  t/t4018/scheme-define-syntax       | 8 ++++++++\n>  t/t4018/scheme-local-define        | 4 ++++\n>  t/t4018/scheme-top-level-define    | 4 ++++\n>  t/t4018/scheme-user-defined-define | 6 ++++++\n>  t/t4034-diff-words.sh              | 1 +\n>  t/t4034/scheme/expect              | 9 +++++++++\n>  t/t4034/scheme/post                | 4 ++++\n>  t/t4034/scheme/pre                 | 4 ++++\n>  userdiff.c                         | 4 ++++\n>  11 files changed, 47 insertions(+)\n>  create mode 100644 t/t4018/scheme-define-syntax\n>  create mode 100644 t/t4018/scheme-local-define\n>  create mode 100644 t/t4018/scheme-top-level-define\n>  create mode 100644 t/t4018/scheme-user-defined-define\n>  create mode 100644 t/t4034/scheme/expect\n>  create mode 100644 t/t4034/scheme/post\n>  create mode 100644 t/t4034/scheme/pre\n>\n> diff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\n> index 0a60472bb5..cfcfa800c2 100644\n> --- a/Documentation/gitattributes.txt\n> +++ b/Documentation/gitattributes.txt\n> @@ -845,6 +845,8 @@ patterns are available:\n>  \n>  - `rust` suitable for source code in the Rust language.\n>  \n> +- `scheme` suitable for source code in the Scheme language.\n> +\n>  - `tex` suitable for source code for LaTeX documents.\n>  \n>  \n> diff --git a/t/t4018-diff-funcname.sh b/t/t4018-diff-funcname.sh\n> index 9675bc17db..823ea96acb 100755\n> --- a/t/t4018-diff-funcname.sh\n> +++ b/t/t4018-diff-funcname.sh\n> @@ -48,6 +48,7 @@ diffpatterns=\"\n>  \tpython\n>  \truby\n>  \trust\n> +\tscheme\n>  \ttex\n>  \tcustom1\n>  \tcustom2\n> diff --git a/t/t4018/scheme-define-syntax b/t/t4018/scheme-define-syntax\n> new file mode 100644\n> index 0000000000..33fa50c844\n> --- /dev/null\n> +++ b/t/t4018/scheme-define-syntax\n> @@ -0,0 +1,8 @@\n> +(define-syntax define-test-suite RIGHT\n> +  (syntax-rules ()\n> +    ((_ suite-name (name test) ChangeMe ...)\n> +     (define suite-name\n> +       (let ((tests\n> +              `((name . ,test) ...)))\n> +         (lambda ()\n> +           (ChangeMe 'suite-name tests)))))))\n> diff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\n> new file mode 100644\n> index 0000000000..bc6d8aebbe\n> --- /dev/null\n> +++ b/t/t4018/scheme-local-define\n> @@ -0,0 +1,4 @@\n> +(define (higher-order)\n> +  (define local-function RIGHT\n> +    (lambda (x)\n> +     (car \"this is\" \"ChangeMe\"))))\n> diff --git a/t/t4018/scheme-top-level-define b/t/t4018/scheme-top-level-define\n> new file mode 100644\n> index 0000000000..624743c22b\n> --- /dev/null\n> +++ b/t/t4018/scheme-top-level-define\n> @@ -0,0 +1,4 @@\n> +(define (some-func x y z) RIGHT\n> +  (let ((a x)\n> +        (b y))\n> +        (ChangeMe a b)))\n> diff --git a/t/t4018/scheme-user-defined-define b/t/t4018/scheme-user-defined-define\n> new file mode 100644\n> index 0000000000..70e403c5e2\n> --- /dev/null\n> +++ b/t/t4018/scheme-user-defined-define\n> @@ -0,0 +1,6 @@\n> +(define-test-suite record-case-tests RIGHT\n> +  (record-case-1 (lambda (fail)\n> +                   (let ((a (make-foo 1 2)))\n> +                     (record-case a\n> +                       ((bar x) (ChangeMe))\n> +                       ((foo a b) (+ a b)))))))\n> diff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\n> index 56f1e62a97..ee7721ab91 100755\n> --- a/t/t4034-diff-words.sh\n> +++ b/t/t4034-diff-words.sh\n> @@ -325,6 +325,7 @@ test_language_driver perl\n>  test_language_driver php\n>  test_language_driver python\n>  test_language_driver ruby\n> +test_language_driver scheme\n>  test_language_driver tex\n>  \n>  test_expect_success 'word-diff with diff.sbe' '\n> diff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\n> new file mode 100644\n> index 0000000000..eed21e803c\n> --- /dev/null\n> +++ b/t/t4034/scheme/expect\n> @@ -0,0 +1,9 @@\n> +<BOLD>diff --git a/pre b/post<RESET>\n> +<BOLD>index 6a5efba..7c4a6b4 100644<RESET>\n> +<BOLD>--- a/pre<RESET>\n> +<BOLD>+++ b/post<RESET>\n> +<CYAN>@@ -1,4 +1,4 @@<RESET>\n> +(define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n> +  ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n> +  (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n> +    (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\n> diff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\n> new file mode 100644\n> index 0000000000..28f59c6584\n> --- /dev/null\n> +++ b/t/t4034/scheme/post\n> @@ -0,0 +1,4 @@\n> +(define (my-func first second)\n> +  ; This is a (moderately) cool function.\n> +  (let ((c (add1 first)))\n> +    (format \"one more than the total is %d\" (+ c second))))\n> diff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\n> new file mode 100644\n> index 0000000000..4bd0069493\n> --- /dev/null\n> +++ b/t/t4034/scheme/pre\n> @@ -0,0 +1,4 @@\n> +(define (myfunc a b)\n> +  ; This is a really cool function.\n> +  (let ((c (+ a b)))\n> +    (format \"one more than the total is %d\" (add1 c))))\n> diff --git a/userdiff.c b/userdiff.c\n> index 3f81a2261c..685fe712aa 100644\n> --- a/userdiff.c\n> +++ b/userdiff.c\n> @@ -191,6 +191,10 @@ PATTERNS(\"rust\",\n>  \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n>  \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>  \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n> +PATTERNS(\"scheme\",\n> +\t \"^[\\t ]*(\\\\(define-?.*)$\",\n\nThe \"define-?.*\" can be simplified to just \"define.*\", but looking at\nthe tests is that the intent? From the tests it looks like \"define[- ]\"\nis what the author wants, unless this is meant to also match\n\"(definements\".\n\nHas this been tested on some real-world scheme code? E.g. I have guile\ninstalled locally, and it has really large top-level eval-when\nblocks. These rules would jump over those to whatever the function above\nthem is.\n\n> +\t /* whitespace separated tokens, but parentheses also can delimit words */\n> +\t \"([^][)(\\\\}\\\\{ \\t])+\"),\n>  PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n>  \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n>  PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n\n"},{"id":"420359","messageId":"xmqqtuovon3f.fsf@gitster.g","threadId":"55399","inReplyTo":"87blb4nf2n.fsf@evledraar.gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-03-28T05:37:24Z","receivedAt":"2021-03-28T05:49:02Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n[jc: redirecting the question to patch author---I am just a messenger]\n\n>> diff --git a/userdiff.c b/userdiff.c\n>> index 3f81a2261c..685fe712aa 100644\n>> --- a/userdiff.c\n>> +++ b/userdiff.c\n>> @@ -191,6 +191,10 @@ PATTERNS(\"rust\",\n>>  \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n>>  \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>>  \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n>> +PATTERNS(\"scheme\",\n>> +\t \"^[\\t ]*(\\\\(define-?.*)$\",\n>\n> The \"define-?.*\" can be simplified to just \"define.*\", but looking at\n> the tests is that the intent? From the tests it looks like \"define[- ]\"\n> is what the author wants, unless this is meant to also match\n> \"(definements\".\n>\n> Has this been tested on some real-world scheme code? E.g. I have guile\n> installed locally, and it has really large top-level eval-when\n> blocks. These rules would jump over those to whatever the function above\n> them is.\n>\n>> +\t /* whitespace separated tokens, but parentheses also can delimit words */\n>> +\t \"([^][)(\\\\}\\\\{ \\t])+\"),\n>>  PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n>>  \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n>>  PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n"},{"id":"420383","messageId":"EBC020E6-BE8B-4332-8225-A988CB7CFA69@gmail.com","threadId":"55399","inReplyTo":"xmqq5z1cqki7.fsf@gitster.g","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-28T11:51:16Z","receivedAt":"2021-03-28T11:51:55Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"On 28-Mar-2021, at 04:20, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> Atharva Raykar <raykar.ath@gmail.com> writes:\n> \n>> diff --git a/t/t4018/scheme-define-syntax b/t/t4018/scheme-define-syntax\n>> new file mode 100644\n>> index 0000000000..603b99cea4\n>> --- /dev/null\n>> +++ b/t/t4018/scheme-define-syntax\n>> @@ -0,0 +1,8 @@\n>> +(define-syntax define-test-suite RIGHT\n>> +  (syntax-rules ()\n>> +    ((_ suite-name (name test) ChangeMe ...)\n>> +     (define suite-name\n>> +       (let ((tests\n>> +              `((name . ,test) ...)))\n>> +         (lambda ()\n>> +           (ChangeMe 'suite-name tests)))))))\n>> \\ No newline at end of file\n> \n> Is there a good reason to leave the final line incomplete?  If there\n> isn't, complete it (applies to other newly-created files in the patch).\n\nWill do.\n\n>> diff --git a/userdiff.c b/userdiff.c\n>> index 3f81a2261c..c51a8c98ba 100644\n>> --- a/userdiff.c\n>> +++ b/userdiff.c\n>> @@ -191,6 +191,14 @@ PATTERNS(\"rust\",\n>> \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n>> \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>> \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n>> +PATTERNS(\"scheme\",\n>> +         \"^[\\t ]*(\\\\(define-?.*)$\",\n> \n> Didn't \"git diff HEAD\" before committing (or \"git show\") highlighted\n> these whitespace errors?\n> \n> .git/rebase-apply/patch:183: indent with spaces.\n>         \"^[\\t ]*(\\\\(define-?.*)$\",\n> .git/rebase-apply/patch:184: trailing whitespace, indent with spaces.\n>         /* \n> .git/rebase-apply/patch:185: indent with spaces.\n>          * Scheme allows symbol names to have any character,\n> .git/rebase-apply/patch:186: indent with spaces.\n>          * as long as it is not a form of a parenthesis.\n> .git/rebase-apply/patch:187: indent with spaces.\n>          * The spaces must be escaped.\n> warning: squelched 2 whitespace errors\n> warning: 7 lines applied after fixing whitespace errors.\n\nIt did highlight the spaces (which I accidentally overlooked), but I\ndidn’t receive these warnings. It shows up with the --check flag though.\nI'll recheck my configuration. Thanks for pointing this out.\n\n> \n>> +         /* \n>> +          * Scheme allows symbol names to have any character,\n>> +          * as long as it is not a form of a parenthesis.\n>> +          * The spaces must be escaped.\n>> +          */\n>> +         \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n> \n> One or more \"dot or anything other than SP or parentheses\"?  But\n> a dot \".\" is neither a space or any {bra-ce} letter, so would the\n> above be equivalent to\n> \n> \t\"[^][()\\\\{\\\\} \\t]+\"\n> \n> I wonder...\n\nA backslash is allowed in scheme identifiers, and I erroneously thought that\nthe first part handles the case for identifiers such as `component\\new` or \n`\\\"id-with-quotes\\\"`. (I tested it with a regex engine that behaves differently\nthan the one git is using, my bad.)\n\n> I am also trying to figure out what you wanted to achieve by\n> mentioning \"The spaces must be escaped.\".  Did you mean something\n> like (string->symbol \"a symbol with SP in it\") is a symbol?  Even\n> so, I cannot quite guess the significance of that fact wrt the\n> regexp you added here?\n\nI initially tried using identifiers like `space\\ separated` and they\nseemed to work in my REPL, but turns out space separated identifiers in\nscheme do not work with backslashes, and it was working because of the way\nmy terminal handled escaping. Space separated identifiers are declared like\n`|space separated|` and this too only seems to work with Racket, not\nthe other Scheme implementations. So I stand corrected here, and it's better\nto drop this feature altogether.\n\nBut somehow, the regexp you suggested, ie:\n\n\t\"[^][()\\\\{\\\\} \\t]+\"\n\ndoes not handle the case of make\\foo -> make\\bar (it will only diff on foo).\nI am not too sure why it treats backslashes as delimiters.\n\nThis seems to actually do what I was going for:\n\n\t\"(\\\\\\\\|[^][)(\\\\}\\\\{ ])+\"\n\n> As we are trying to catch program identifiers (symbols in scheme)\n> and numeric literals, treating any group of non-whitespace letters\n> that is delimited by one or more whitespaces as a \"word\" would be a\n> good first-order approximation, but in addition, as can be seen in\n> an example like (a(b(c))), parentheses can also serve as such \"word\n> delimiters\" in addition to whitespaces.  So the regexp given above\n> makes sense to me from that angle, especially if you do not limit\n> the whitespace to only SP, but include HT (\\t) as well.  But was\n> that how you came up with the regexp?\n\nYes, this is exactly what I was trying to express. All words should be\ndelimited by either whitespace or a parenthesis, and all other special\ncharacters should be accepted as part of the word."},{"id":"420384","messageId":"5BA00FC6-9810-49AB-8DE2-D4F4010E2F82@gmail.com","threadId":"55399","inReplyTo":"3def82fd-71a7-3ad9-0fa2-48598bfd3313@kdbg.org","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-28T12:23:21Z","receivedAt":"2021-03-28T12:24:33Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"On 28-Mar-2021, at 05:16, Johannes Sixt <j6t@kdbg.org> wrote:\n>> diff --git a/t/t4018/scheme-define-syntax b/t/t4018/scheme-define-syntax\n>> new file mode 100644\n>> index 0000000000..603b99cea4\n>> --- /dev/null\n>> +++ b/t/t4018/scheme-define-syntax\n>> @@ -0,0 +1,8 @@\n>> +(define-syntax define-test-suite RIGHT\n>> +  (syntax-rules ()\n>> +    ((_ suite-name (name test) ChangeMe ...)\n>> +     (define suite-name\n> \n> This test is suspicious. Notice the \"ChangeMe\" above? That is sufficient\n> to let the test case succeed. The \"ChangeMe\" in the last line below\n> should be the only one.\n\nThanks for pointing this out. The second \"ChangeMe\" was not supposed to be\nthere.\n\nWhat I wanted to test was the hunk header showing the line for\n'(define-syntax ...' and not the internal '(define ...' below it. Thus the\nChangeMe should be located above the internal define so that the hunk header\nwould show define-syntax and not the local define.\n\n> But then there is this indented '(define' that is not marked as RIGHT,\n> and I wonder how is it different from...\n> \n>> +       (let ((tests\n>> +              `((name . ,test) ...)))\n>> +         (lambda ()\n>> +           (ChangeMe 'suite-name tests)))))))\n>> \\ No newline at end of file\n>> diff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\n>> new file mode 100644\n>> index 0000000000..90e75dcce8\n>> --- /dev/null\n>> +++ b/t/t4018/scheme-local-define\n>> @@ -0,0 +1,4 @@\n>> +(define (higher-order)\n>> +  (define local-function RIGHT\n> \n> ... this one, which is also indented and *is* marked as RIGHT.\n\nIn this test case, I was explicitly testing for an indented '(define'\nwhereas in the former, I was testing for the top-level '(define-syntax',\nwhich happened to have an internal define (which will inevitably show up\nin a lot of scheme code).\n\n>> +    (lambda (x)\n>> +     (car \"this is\" \"ChangeMe\"))))\n>> \\ No newline at end of file\n> \n>> diff --git a/userdiff.c b/userdiff.c\n>> index 3f81a2261c..c51a8c98ba 100644\n>> --- a/userdiff.c\n>> +++ b/userdiff.c\n>> @@ -191,6 +191,14 @@ PATTERNS(\"rust\",\n>> \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n>> \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>> \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n>> +PATTERNS(\"scheme\",\n>> +         \"^[\\t ]*(\\\\(define-?.*)$\",\n> \n> This \"optional hyphen followed by anything\" in the regex is strange.\n> Wouldn't that also capture a line that looks like, e.g.,\n> \n>    (defined-foo bar)\n> \n> Perhaps we want \"define[- \\t].*\" in the regex?\n\nYes, this is what I intended to do, thanks for correcting it.\n\n"},{"id":"420385","messageId":"578FC14B-CB72-41CA-A8FD-1480EBCCB968@gmail.com","threadId":"55399","inReplyTo":"87blb4nf2n.fsf@evledraar.gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-28T12:40:56Z","receivedAt":"2021-03-28T12:42:03Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"On 28-Mar-2021, at 08:46, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> The \"define-?.*\" can be simplified to just \"define.*\", but looking at\n> the tests is that the intent? From the tests it looks like \"define[- ]\"\n> is what the author wants, unless this is meant to also match\n> \"(definements\".\n\nYes, you captured my intent correctly. Will fix it.\n\n> Has this been tested on some real-world scheme code? E.g. I have guile\n> installed locally, and it has really large top-level eval-when\n> blocks. These rules would jump over those to whatever the function above\n> them is.\n\nI do not have a large scheme codebase on my own, I usually use Racket,\nwhich is a much larger language with many more forms. Other Schemes like\nGuile also extend the language a lot, like in your example, eval-when is\nan extension provided by Guile (and Chicken and Chez), but not a part of\nthe R6RS document when I searched its index.\n\nSo the 'define' forms are the only one that I know would reliably be present\nacross all schemes. But one can also make a case where some of these non-standard\nforms may be common enough that they are worth adding in. In that case which\nforms to include? Should we consider everything in the SRFI's[1]? Should the\nvarious module definitions of Racket be included? It's a little tricky to know\nwhere to stop.\n\nThat being said, I will try to run this through more Scheme codebases that I can\nfind and see if there are any forms that seem to show up commonly enough that they\nare worth including.\n\n[1] https://en.wikipedia.org/wiki/Scheme_Requests_for_Implementation"},{"id":"420386","messageId":"166B3835-32C3-4CE6-9799-C187284E5756@gmail.com","threadId":"55399","inReplyTo":"xmqq1rc0qjn1.fsf@gitster.g","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-28T12:45:06Z","receivedAt":"2021-03-28T12:46:25Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"On 28-Mar-2021, at 04:39, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> Junio C Hamano <gitster@pobox.com> writes:\n> \n>> Atharva Raykar <raykar.ath@gmail.com> writes:\n>> ...\n>>> +           (ChangeMe 'suite-name tests)))))))\n>>> \\ No newline at end of file\n>> \n>> Is there a good reason to leave the final line incomplete?  ...\n>> I am also trying to figure out what you wanted to achieve ...\n> \n> Taking all of them together, here is what I hope you may agree as\n> its improved version.  The only differences from what you posted are\n> corrections to all the \"\\ No newline at end of file\" and the simplification\n> of the pattern (remove \"a dot\" from the alternative and add \\t next\n> to SP).  Without changes, the new tests still pass so ... ;-)\n> \n>    diff --git c/userdiff.c w/userdiff.c\n>    index 5fd0eb31ec..685fe712aa 100644\n>    --- c/userdiff.c\n>    +++ w/userdiff.c\n>    @@ -193,12 +193,8 @@ PATTERNS(\"rust\",\n>             \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n>     PATTERNS(\"scheme\",\n>             \"^[\\t ]*(\\\\(define-?.*)$\",\n>    -\t /*\n>    -\t  * Scheme allows symbol names to have any character,\n>    -\t  * as long as it is not a form of a parenthesis.\n>    -\t  * The spaces must be escaped.\n>    -\t  */\n>    -\t \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n>    +\t /* whitespace separated tokens, but parentheses also can delimit words */\n>    +\t \"([^][)(\\\\}\\\\{ \\t])+\"),\n>     PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n>             \"[={}\\\"]|[^={}\\\" \\t]+\"),\n>     PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n\nThanks for these. I will eventually send another patch with the whitespaces corrected,\nand try to see if there is a better way to handle backslashes, other than the regex I\nsuggested. I will also be writing another test case to check that case properly.\n\nI will also incorporate the other changes suggested by Johannes and Ævar as well,\nMy regex was not supposed to capture forms like `defined-thing`. And there are a\nfew rough edges with some of my test cases, which I will correct as well in the next\npatch. It is also worth spending some more time and see if there is any other form other\nthan definitions that a Scheme programmer other than myself may be interested in. I will\nconsult a few Scheme communities and mailing lists and see what more experienced\nprogrammers have to say."},{"id":"420417","messageId":"xmqqft0fm9uu.fsf@gitster.g","threadId":"55399","inReplyTo":"EBC020E6-BE8B-4332-8225-A988CB7CFA69@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-03-28T18:06:17Z","receivedAt":"2021-03-28T18:07:28Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Atharva Raykar <raykar.ath@gmail.com> writes:\n\n>>> +         \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n>> \n>> One or more \"dot or anything other than SP or parentheses\"?  But\n>> a dot \".\" is neither a space or any {bra-ce} letter, so would the\n>> above be equivalent to\n>> \n>> \t\"[^][()\\\\{\\\\} \\t]+\"\n>> \n>> I wonder...\n>\n> A backslash is allowed in scheme identifiers, and I erroneously thought that\n> the first part handles the case for identifiers such as `component\\new` or \n> `\\\"id-with-quotes\\\"`. (I tested it with a regex engine that behaves differently\n> than the one git is using, my bad.)\n\nAh, perhaps you didn't have enough backslashes.  A half of the\ndoubled one before the dot is eaten by the C compiler, so the regexp\nengine is seeing only a single backslash before the dot, which means\n\"literally a single dot\".  If you meant \"literally a single\nbackslash, followed by any single char\", you probably would write 4\nbackslashes and a dot---half of the backslashes would be eaten by\nthe compiler, so you'd be passing two backslashes and a dot, which\nis probably what you meant.\n\nHaving said that, two further points.\n\n - the \"anything but whitespaces and various forms of parentheses\"\n   set would include backslash, so 'component\\new' would be taken as\n   a single word with \"[^][()\\\\{\\\\} \\t]+\", wouldn't it?\n\n - how common is the use of backslashes in identifiers?  I am trying\n   to see if the additional complexity needed to support them is\n   worth the benefit.\n\n> But somehow, the regexp you suggested, ie:\n>\n> \t\"[^][()\\\\{\\\\} \\t]+\"\n>\n> does not handle the case of make\\foo -> make\\bar (it will only diff on foo).\n> I am not too sure why it treats backslashes as delimiters.\n\nPerhaps because you have included two backslashes inside [] to say\n\"backslash is not a word character\" in the original, and I blindly\ncopied that?  IOW, do you need to quote {} inside []?\n\n> Yes, this is exactly what I was trying to express. All words should be\n> delimited by either whitespace or a parenthesis, and all other special\n> characters should be accepted as part of the word.\n\nThat sentence after \"All words should be...\" would be a good comment\nto replace what you wrote in the original, then ;-).\n"},{"id":"420434","messageId":"562DCDA0-EAE6-408F-97D7-127689DE5559@gmail.com","threadId":"55399","inReplyTo":"xmqqft0fm9uu.fsf@gitster.g","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-29T08:12:22Z","receivedAt":"2021-03-29T08:13:36Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"On 28-Mar-2021, at 23:36, Junio C Hamano <gitster@pobox.com> wrote:\n> \n> Atharva Raykar <raykar.ath@gmail.com> writes:\n> \n>>>> +         \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n>>> \n>>> One or more \"dot or anything other than SP or parentheses\"?  But\n>>> a dot \".\" is neither a space or any {bra-ce} letter, so would the\n>>> above be equivalent to\n>>> \n>>> \t\"[^][()\\\\{\\\\} \\t]+\"\n>>> \n>>> I wonder...\n>> \n>> A backslash is allowed in scheme identifiers, and I erroneously thought that\n>> the first part handles the case for identifiers such as `component\\new` or \n>> `\\\"id-with-quotes\\\"`. (I tested it with a regex engine that behaves differently\n>> than the one git is using, my bad.)\n> \n> Ah, perhaps you didn't have enough backslashes.  A half of the\n> doubled one before the dot is eaten by the C compiler, so the regexp\n> engine is seeing only a single backslash before the dot, which means\n> \"literally a single dot\".  If you meant \"literally a single\n> backslash, followed by any single char\", you probably would write 4\n> backslashes and a dot---half of the backslashes would be eaten by\n> the compiler, so you'd be passing two backslashes and a dot, which\n> is probably what you meant.\n> \n> Having said that, two further points.\n> \n> - the \"anything but whitespaces and various forms of parentheses\"\n>   set would include backslash, so 'component\\new' would be taken as\n>   a single word with \"[^][()\\\\{\\\\} \\t]+\", wouldn't it?\n> \n> - how common is the use of backslashes in identifiers?  I am trying\n>   to see if the additional complexity needed to support them is\n>   worth the benefit.\n\nI have refined the regex, and now it is much simpler and does all of what\nI want it to:\n\n\t\"([^][)(}{[:space:]])+\"\n\nI did not have to escape the various parentheses, so I avoided the need to\nhandle backslashes separately. The \"\\\\t\" was causing problems as well because\nit took it as a '\\' followed by a 't' (Thanks to j416 on #git-devel for\nhelping me out on this).\n\n>> Yes, this is exactly what I was trying to express. All words should be\n>> delimited by either whitespace or a parenthesis, and all other special\n>> characters should be accepted as part of the word.\n> \n> That sentence after \"All words should be...\" would be a good comment\n> to replace what you wrote in the original, then ;-).\n\nYes, that should make it a lot more clear.\n\n"},{"id":"420439","messageId":"62695830-2f9e-c3b5-856c-01b97eb2c3af@gmail.com","threadId":"55399","inReplyTo":"578FC14B-CB72-41CA-A8FD-1480EBCCB968@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2021-03-29T10:08:57Z","receivedAt":"2021-03-29T10:10:05Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Atharva\n\nOn 28/03/2021 13:40, Atharva Raykar wrote:\n> On 28-Mar-2021, at 08:46, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>> The \"define-?.*\" can be simplified to just \"define.*\", but looking at\n>> the tests is that the intent? From the tests it looks like \"define[- ]\"\n>> is what the author wants, unless this is meant to also match\n>> \"(definements\".\n> \n> Yes, you captured my intent correctly. Will fix it.\n> \n>> Has this been tested on some real-world scheme code? E.g. I have guile\n>> installed locally, and it has really large top-level eval-when\n>> blocks. These rules would jump over those to whatever the function above\n>> them is.\n> \n> I do not have a large scheme codebase on my own, I usually use Racket,\n> which is a much larger language with many more forms. Other Schemes like\n> Guile also extend the language a lot, like in your example, eval-when is\n> an extension provided by Guile (and Chicken and Chez), but not a part of\n> the R6RS document when I searched its index.\n> \n> So the 'define' forms are the only one that I know would reliably be present\n> across all schemes. But one can also make a case where some of these non-standard\n> forms may be common enough that they are worth adding in. In that case which\n> forms to include? Should we consider everything in the SRFI's[1]? Should the\n> various module definitions of Racket be included? It's a little tricky to know\n> where to stop.\n\nIf there are some common forms such as eval-when then it would be good \nto include them, otherwise we end up needing a different rule for each \nscheme implementation as they all seem to tweak something. Gerbil uses \n'def...' e.g def, defsyntax, defstruct, defrules rather than define, \ndefine-syntax, define-record etc. I'm not user if we want to accommodate \nthat or not.\n\nBest Wishes\n\nPhillip\n\n\n> That being said, I will try to run this through more Scheme codebases that I can\n> find and see if there are any forms that seem to show up commonly enough that they\n> are worth including.\n> \n> [1] https://en.wikipedia.org/wiki/Scheme_Requests_for_Implementation\n> \n\n"},{"id":"420440","messageId":"cfd75dc0-7828-f36c-fd7b-c9f5a2e8d4cc@gmail.com","threadId":"55399","inReplyTo":"EBC020E6-BE8B-4332-8225-A988CB7CFA69@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2021-03-29T10:12:33Z","receivedAt":"2021-03-29T10:13:17Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Atharva\n\nOn 28/03/2021 12:51, Atharva Raykar wrote:\n> On 28-Mar-2021, at 04:20, Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> Atharva Raykar <raykar.ath@gmail.com> writes:\n>>\n>>> +         /*\n>>> +          * Scheme allows symbol names to have any character,\n>>> +          * as long as it is not a form of a parenthesis.\n>>> +          * The spaces must be escaped.\n>>> +          */\n>>> +         \"(\\\\.|[^][)(\\\\}\\\\{ ])+\"),\n>>\n>> One or more \"dot or anything other than SP or parentheses\"?  But\n>> a dot \".\" is neither a space or any {bra-ce} letter, so would the\n>> above be equivalent to\n>>\n>> \t\"[^][()\\\\{\\\\} \\t]+\"\n>>\n>> I wonder...\n> \n> A backslash is allowed in scheme identifiers, and I erroneously thought that\n> the first part handles the case for identifiers such as `component\\new` or\n> `\\\"id-with-quotes\\\"`. (I tested it with a regex engine that behaves differently\n> than the one git is using, my bad.)\n> \n>> I am also trying to figure out what you wanted to achieve by\n>> mentioning \"The spaces must be escaped.\".  Did you mean something\n>> like (string->symbol \"a symbol with SP in it\") is a symbol?  Even\n>> so, I cannot quite guess the significance of that fact wrt the\n>> regexp you added here?\n> \n> I initially tried using identifiers like `space\\ separated` and they\n> seemed to work in my REPL, but turns out space separated identifiers in\n> scheme do not work with backslashes, and it was working because of the way\n> my terminal handled escaping. Space separated identifiers are declared like\n> `|space separated|` and this too only seems to work with Racket, not\n> the other Scheme implementations.\n\nI think the bar notation works with some other such as gambit and \npossibly guile (it's a while since I used the latter)\n\nBest wishes\n\nPhillip\n\n  So I stand corrected here, and it's better\n> to drop this feature altogether.\n> \n> But somehow, the regexp you suggested, ie:\n> \n> \t\"[^][()\\\\{\\\\} \\t]+\"\n> \n> does not handle the case of make\\foo -> make\\bar (it will only diff on foo).\n> I am not too sure why it treats backslashes as delimiters.\n> \n> This seems to actually do what I was going for:\n> \n> \t\"(\\\\\\\\|[^][)(\\\\}\\\\{ ])+\"\n> \n>> As we are trying to catch program identifiers (symbols in scheme)\n>> and numeric literals, treating any group of non-whitespace letters\n>> that is delimited by one or more whitespaces as a \"word\" would be a\n>> good first-order approximation, but in addition, as can be seen in\n>> an example like (a(b(c))), parentheses can also serve as such \"word\n>> delimiters\" in addition to whitespaces.  So the regexp given above\n>> makes sense to me from that angle, especially if you do not limit\n>> the whitespace to only SP, but include HT (\\t) as well.  But was\n>> that how you came up with the regexp?\n> \n> Yes, this is exactly what I was trying to express. All words should be\n> delimited by either whitespace or a parenthesis, and all other special\n> characters should be accepted as part of the word.\n> \n\n"},{"id":"420441","messageId":"09678471-b2a2-8504-2293-e2b34a3a96e8@gmail.com","threadId":"55399","inReplyTo":"5BA00FC6-9810-49AB-8DE2-D4F4010E2F82@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2021-03-29T10:18:49Z","receivedAt":"2021-03-29T10:19:42Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Atharva\n\nOn 28/03/2021 13:23, Atharva Raykar wrote:\n> On 28-Mar-2021, at 05:16, Johannes Sixt <j6t@kdbg.org> wrote:\n > [...]\n>>> diff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\n>>> new file mode 100644\n>>> index 0000000000..90e75dcce8\n>>> --- /dev/null\n>>> +++ b/t/t4018/scheme-local-define\n>>> @@ -0,0 +1,4 @@\n>>> +(define (higher-order)\n>>> +  (define local-function RIGHT\n>>\n>> ... this one, which is also indented and *is* marked as RIGHT.\n> \n> In this test case, I was explicitly testing for an indented '(define'\n> whereas in the former, I was testing for the top-level '(define-syntax',\n> which happened to have an internal define (which will inevitably show up\n> in a lot of scheme code).\n\nIt would be nice to include indented define forms but including them \nmeans that any change to the body of a function is attributed to the \nlast internal definition rather than the actual function. For example\n\n(define (f arg)\n   (define (g x)\n     (+ 1 x))\n\n   (some-func ...)\n   ;;any change here will have '(define (g x)' in the hunk header, not \n'(define (f arg)'\n\nI don't think this can be avoided as we rely on regexs rather than \nparsing the source so it is probably best to only match toplevel defines.\n\nBest Wishes\n\nPhillip\n\n> \n>>> +    (lambda (x)\n>>> +     (car \"this is\" \"ChangeMe\"))))\n>>> \\ No newline at end of file\n>>\n>>> diff --git a/userdiff.c b/userdiff.c\n>>> index 3f81a2261c..c51a8c98ba 100644\n>>> --- a/userdiff.c\n>>> +++ b/userdiff.c\n>>> @@ -191,6 +191,14 @@ PATTERNS(\"rust\",\n>>> \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n>>> \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>>> \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n>>> +PATTERNS(\"scheme\",\n>>> +         \"^[\\t ]*(\\\\(define-?.*)$\",\n>>\n>> This \"optional hyphen followed by anything\" in the regex is strange.\n>> Wouldn't that also capture a line that looks like, e.g.,\n>>\n>>     (defined-foo bar)\n>>\n>> Perhaps we want \"define[- \\t].*\" in the regex?\n> \n> Yes, this is what I intended to do, thanks for correcting it.\n> \n\n"},{"id":"420445","messageId":"71c34328-9814-2777-3a9d-f908602dd36f@kdbg.org","threadId":"55399","inReplyTo":"09678471-b2a2-8504-2293-e2b34a3a96e8@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2021-03-29T10:48:20Z","receivedAt":"2021-03-29T10:49:05Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 29.03.21 um 12:18 schrieb Phillip Wood:\n> It would be nice to include indented define forms but including them\n> means that any change to the body of a function is attributed to the\n> last internal definition rather than the actual function. For example\n> \n> (define (f arg)\n>   (define (g x)\n>     (+ 1 x))\n> \n>   (some-func ...)\n>   ;;any change here will have '(define (g x)' in the hunk header, not\n> '(define (f arg)'\n> \n> I don't think this can be avoided as we rely on regexs rather than\n> parsing the source so it is probably best to only match toplevel defines.\n\nThere can be two rules, one that matches '(define-' that is indented,\nand another one that matches all non-indented forms of definitions. If\nthat is what you mean.\n\n-- Hannes\n"},{"id":"420465","messageId":"87wntqm7dj.fsf@evledraar.gmail.com","threadId":"55399","inReplyTo":"71c34328-9814-2777-3a9d-f908602dd36f@kdbg.org","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-03-29T13:12:08Z","receivedAt":"2021-03-29T13:12:52Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Mar 29 2021, Johannes Sixt wrote:\n\n> Am 29.03.21 um 12:18 schrieb Phillip Wood:\n>> It would be nice to include indented define forms but including them\n>> means that any change to the body of a function is attributed to the\n>> last internal definition rather than the actual function. For example\n>> \n>> (define (f arg)\n>>   (define (g x)\n>>     (+ 1 x))\n>> \n>>   (some-func ...)\n>>   ;;any change here will have '(define (g x)' in the hunk header, not\n>> '(define (f arg)'\n>> \n>> I don't think this can be avoided as we rely on regexs rather than\n>> parsing the source so it is probably best to only match toplevel defines.\n>\n> There can be two rules, one that matches '(define-' that is indented,\n> and another one that matches all non-indented forms of definitions. If\n> that is what you mean.\n\nYes, but that doesn't help in these sorts of cases because what a rule\nlike that really wants is some version of \"don't match this line, but\nonly if you can reasonably match this other rule\".\n\nWe can only do rule precedence on a per-line basis via the inverted\nmatches.\n\nSo for languages like cl/elisp/scheme and others where it's common to\nhave nested function definitions (then -W would like the top-level) *OR*\nsimilarly looking nested function definitions, but the top-level isn't a\nfunction but a (setq) or whatever we're basically stuck with picking one\nor the other.\n\nI've pondered how to get around this problem in my userdiff.c hacking\nwithout resorting to supporting some general-purpose Turing machine, and\nhave so far come up with nothing.\n\nYou can see lots of prior art by grepping Emacs's source code for\nbeginning-of-defun, it solves this problem by exposing a Turing machine\n:)\n"},{"id":"420476","messageId":"7edaee06-2149-a547-4fa9-c91b241ff966@gmail.com","threadId":"55399","inReplyTo":"87wntqm7dj.fsf@evledraar.gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2021-03-29T14:06:27Z","receivedAt":"2021-03-29T14:07:26Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"On 29/03/2021 14:12, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Mon, Mar 29 2021, Johannes Sixt wrote:\n> \n>> Am 29.03.21 um 12:18 schrieb Phillip Wood:\n>>> It would be nice to include indented define forms but including them\n>>> means that any change to the body of a function is attributed to the\n>>> last internal definition rather than the actual function. For example\n>>>\n>>> (define (f arg)\n>>>    (define (g x)\n>>>      (+ 1 x))\n>>>\n>>>    (some-func ...)\n>>>    ;;any change here will have '(define (g x)' in the hunk header, not\n>>> '(define (f arg)'\n>>>\n>>> I don't think this can be avoided as we rely on regexs rather than\n>>> parsing the source so it is probably best to only match toplevel defines.\n>>\n>> There can be two rules, one that matches '(define-' that is indented,\n>> and another one that matches all non-indented forms of definitions. If\n>> that is what you mean.\n> \n> Yes, but that doesn't help in these sorts of cases because what a rule\n> like that really wants is some version of \"don't match this line, but\n> only if you can reasonably match this other rule\".\n> \n> We can only do rule precedence on a per-line basis via the inverted\n> matches.\n> \n> So for languages like cl/elisp/scheme and others where it's common to\n> have nested function definitions (then -W would like the top-level) *OR*\n> similarly looking nested function definitions, but the top-level isn't a\n> function but a (setq) or whatever we're basically stuck with picking one\n> or the other.\n\nExactly\n\n> I've pondered how to get around this problem in my userdiff.c hacking\n> without resorting to supporting some general-purpose Turing machine, and\n> have so far come up with nothing.\n\nI think using an indentation heuristic would probably work quite well \nfor most languages - see \nhttps://public-inbox.org/git/20200923215859.102981-1-rtzoeller@rtzoeller.com/ \nfor a discussion from last year (from memory there were some problems \nwith the approach in those patches but I think there are some suggestion \nfrom Peff and me later in the thread on how they could be overcome)\n\nBest Wishes\n\nPhillip\n\n\n> You can see lots of prior art by grepping Emacs's source code for\n> beginning-of-defun, it solves this problem by exposing a Turing machine\n> :)\n> \n"},{"id":"420509","messageId":"xmqq1rbxk7qi.fsf@gitster.g","threadId":"55399","inReplyTo":"562DCDA0-EAE6-408F-97D7-127689DE5559@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-03-29T20:47:17Z","receivedAt":"2021-03-29T20:48:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Atharva Raykar <raykar.ath@gmail.com> writes:\n\n>> Having said that, two further points.\n>> \n>> - the \"anything but whitespaces and various forms of parentheses\"\n>>   set would include backslash, so 'component\\new' would be taken as\n>>   a single word with \"[^][()\\\\{\\\\} \\t]+\", wouldn't it?\n>> \n>> - how common is the use of backslashes in identifiers?  I am trying\n>>   to see if the additional complexity needed to support them is\n>>   worth the benefit.\n>\n> I have refined the regex, and now it is much simpler and does all of what\n> I want it to:\n>\n> \t\"([^][)(}{[:space:]])+\"\n\nOK, [:space:] is already used elsewhere, so it would be OK.\n\nIn practice, the only difference from \"[ \\t]\" (which is used in many\nother patterns in the same file) is that [:space:] class includes\nform-feed (\\Ctrl-L); nobody would write vertical-tab in the code,\nand the matching is done one line at a time, so the fact that LF (or\nCRLF) is in the [:space:] class does not make a difference anyway.\n\n> I did not have to escape the various parentheses, so I avoided the need to\n> handle backslashes separately. The \"\\\\t\" was causing problems as well because\n\nIf you spelled \"\\\\t\" that would have caused a problem of your own\nmaking ;-)\n\nI think what I gave in the message you are responding to was a\nsingle backslash followed by a 't', to let the compiler turn them\ninto a single HT character, and that wouldn't have had such a\nproblem---in fact \"[ \\t]\" is used in many other existing rules in\nthe same file.\n\nThanks.\n\n"},{"id":"420556","messageId":"59DFC82F-A3EA-4637-94AE-4042697448FF@gmail.com","threadId":"55399","inReplyTo":"62695830-2f9e-c3b5-856c-01b97eb2c3af@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-30T06:41:31Z","receivedAt":"2021-03-30T06:42:22Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"\nOn 29-Mar-2021, at 15:38, Phillip Wood <phillip.wood123@gmail.com> wrote:\n> On 28/03/2021 13:40, Atharva Raykar wrote:\n>> On 28-Mar-2021, at 08:46, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>>> The \"define-?.*\" can be simplified to just \"define.*\", but looking at\n>>> the tests is that the intent? From the tests it looks like \"define[- ]\"\n>>> is what the author wants, unless this is meant to also match\n>>> \"(definements\".\n>> Yes, you captured my intent correctly. Will fix it.\n>>> Has this been tested on some real-world scheme code? E.g. I have guile\n>>> installed locally, and it has really large top-level eval-when\n>>> blocks. These rules would jump over those to whatever the function above\n>>> them is.\n>> I do not have a large scheme codebase on my own, I usually use Racket,\n>> which is a much larger language with many more forms. Other Schemes like\n>> Guile also extend the language a lot, like in your example, eval-when is\n>> an extension provided by Guile (and Chicken and Chez), but not a part of\n>> the R6RS document when I searched its index.\n>> So the 'define' forms are the only one that I know would reliably be present\n>> across all schemes. But one can also make a case where some of these non-standard\n>> forms may be common enough that they are worth adding in. In that case which\n>> forms to include? Should we consider everything in the SRFI's[1]? Should the\n>> various module definitions of Racket be included? It's a little tricky to know\n>> where to stop.\n> \n> If there are some common forms such as eval-when then it would be good to include them, otherwise we end up needing a different rule for each scheme implementation as they all seem to tweak something. Gerbil uses 'def...' e.g def, defsyntax, defstruct, defrules rather than define, define-syntax, define-record etc. I'm not user if we want to accommodate that or not.\n\nYes, this is the part that is hard for me to figure out. I am going by\ntwo heuristics: what Scheme communities in other places would generally\nprefer, and what patterns I see happen more often in scheme code.\n\nThe former is tricky to do. I posted to a few mailing lists about this,\nbut they don't seem active enough to garner any responses.\n\nThe latter is a little easier to measure quickly. I did a GitHub search,\nwhere I filtered results to only consider Scheme files (language:scheme).\n\nSome armchair stats, just for a broad understanding:\n\n  Total number of scheme files: 529,339\n  No. of times a construct is used in those files:\n    define and its variants : 431,090 (81.4%)\n    def and its variants    :  18,466 ( 3.5%)\n    eval-when               :   3,375 ( 0.6%)\n\nThere was no way for me to quickly know which of these uses are at the top\nlevel, but either way of the more structural forms that do show up in Scheme\ncode, define and its variants seem like a clear winner. I am not sure if\nit's worth adding more rules to check for def and its variants, given that\nthey are not nearly as common."},{"id":"420558","messageId":"D8256AFA-898E-4388-8FCC-7D3D340C001E@gmail.com","threadId":"55399","inReplyTo":"09678471-b2a2-8504-2293-e2b34a3a96e8@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-30T07:04:38Z","receivedAt":"2021-03-30T07:05:27Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"\n\n> On 29-Mar-2021, at 15:48, Phillip Wood <phillip.wood123@gmail.com> wrote:\n> \n> Hi Atharva\n> \n> On 28/03/2021 13:23, Atharva Raykar wrote:\n>> On 28-Mar-2021, at 05:16, Johannes Sixt <j6t@kdbg.org> wrote:\n> > [...]\n>>>> diff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\n>>>> new file mode 100644\n>>>> index 0000000000..90e75dcce8\n>>>> --- /dev/null\n>>>> +++ b/t/t4018/scheme-local-define\n>>>> @@ -0,0 +1,4 @@\n>>>> +(define (higher-order)\n>>>> +  (define local-function RIGHT\n>>> \n>>> ... this one, which is also indented and *is* marked as RIGHT.\n>> In this test case, I was explicitly testing for an indented '(define'\n>> whereas in the former, I was testing for the top-level '(define-syntax',\n>> which happened to have an internal define (which will inevitably show up\n>> in a lot of scheme code).\n> \n> It would be nice to include indented define forms but including them means that any change to the body of a function is attributed to the last internal definition rather than the actual function. For example\n> \n> (define (f arg)\n>  (define (g x)\n>    (+ 1 x))\n> \n>  (some-func ...)\n>  ;;any change here will have '(define (g x)' in the hunk header, not '(define (f arg)'\n\nThe reason I went for this over the top level forms, is because\nI felt it was useful to see the nearest definition for internal\nfunctions that often have a lot of the actual business logic of\nthe program (at least a lot of SICP seems to follow this pattern).\nThe disadvantage is as you said, it might also catch trivial inner\nfunctions and the developer might lose context.\n\nAnother problem is it may match more trivial bindings, like:\n\n(define (some-func things)\n  ...\n  (define items '(eggs\n                  ham\n                  peanut-butter))\n  ...)\n\nWhat I have noticed *anecdotally* is that this is not common enough\nto be too much of a problem, and local define bindings seem to be more\nfavoured in Racket than other Schemes, that use 'let' more often.\n\n> I don't think this can be avoided as we rely on regexs rather than parsing the source so it is probably best to only match toplevel defines.\n\nThe other issue with only matching top level defines is that a\nlot of scheme programs are library definitions, something like\n\n(library\n    (foo bar)\n  (export ...)\n  (define ...)\n  (define ...)\n  ;; and a bunch of other definitions...\n)\n\nOnly matching top level defines will completely ignore matching all\nthe definitions in these files."},{"id":"420569","messageId":"A3C3DD12-3C00-49ED-B427-37AAB4211C2A@gmail.com","threadId":"55399","inReplyTo":"D8256AFA-898E-4388-8FCC-7D3D340C001E@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-30T10:22:51Z","receivedAt":"2021-03-30T10:23:54Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"\n\n> On 30-Mar-2021, at 12:34, Atharva Raykar <raykar.ath@gmail.com> wrote:\n> \n> \n> \n>> On 29-Mar-2021, at 15:48, Phillip Wood <phillip.wood123@gmail.com> wrote:\n>> \n>> Hi Atharva\n>> \n>> On 28/03/2021 13:23, Atharva Raykar wrote:\n>>> On 28-Mar-2021, at 05:16, Johannes Sixt <j6t@kdbg.org> wrote:\n>>> [...]\n>>>>> diff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\n>>>>> new file mode 100644\n>>>>> index 0000000000..90e75dcce8\n>>>>> --- /dev/null\n>>>>> +++ b/t/t4018/scheme-local-define\n>>>>> @@ -0,0 +1,4 @@\n>>>>> +(define (higher-order)\n>>>>> +  (define local-function RIGHT\n>>>> \n>>>> ... this one, which is also indented and *is* marked as RIGHT.\n>>> In this test case, I was explicitly testing for an indented '(define'\n>>> whereas in the former, I was testing for the top-level '(define-syntax',\n>>> which happened to have an internal define (which will inevitably show up\n>>> in a lot of scheme code).\n>> \n>> It would be nice to include indented define forms but including them means that any change to the body of a function is attributed to the last internal definition rather than the actual function. For example\n>> \n>> (define (f arg)\n>> (define (g x)\n>>   (+ 1 x))\n>> \n>> (some-func ...)\n>> ;;any change here will have '(define (g x)' in the hunk header, not '(define (f arg)'\n> \n> The reason I went for this over the top level forms, is because\n> I felt it was useful to see the nearest definition for internal\n> functions that often have a lot of the actual business logic of\n> the program (at least a lot of SICP seems to follow this pattern).\n> The disadvantage is as you said, it might also catch trivial inner\n> functions and the developer might lose context.\n\nNever mind this message, I had misunderstood the problem you were trying to\ndemonstrate. I wholeheartedly agree with what you are trying to say, and\nthe indentation heuristic discussed does look interesting. I shall have a\nglance at the RFC you linked in the other reply.\n\n> The disadvantage is as you said, it might also catch trivial inner\n> functions and the developer might lose context.\n\nFeel free to disregard me misquoting you here. You did not say that (:\n\n> Another problem is it may match more trivial bindings, like:\n> \n> (define (some-func things)\n>  ...\n>  (define items '(eggs\n>                  ham\n>                  peanut-butter))\n>  ...)\n> \n> What I have noticed *anecdotally* is that this is not common enough\n> to be too much of a problem, and local define bindings seem to be more\n> favoured in Racket than other Schemes, that use 'let' more often.\n> \n>> I don't think this can be avoided as we rely on regexs rather than parsing the source so it is probably best to only match toplevel defines.\n> \n> The other issue with only matching top level defines is that a\n> lot of scheme programs are library definitions, something like\n> \n> (library\n>    (foo bar)\n>  (export ...)\n>  (define ...)\n>  (define ...)\n>  ;; and a bunch of other definitions...\n> )\n> \n> Only matching top level defines will completely ignore matching all\n> the definitions in these files.\n\nThat said, I still stand by the fact that only catching top level defines\nwill lead to a lot of definitions being ignored. Maybe the occasional\nmismatch may be worth the gain in the number of function contexts being\ndetected?\n\n\n"},{"id":"420575","messageId":"874kgsn6kb.fsf@evledraar.gmail.com","threadId":"55399","inReplyTo":"59DFC82F-A3EA-4637-94AE-4042697448FF@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-03-30T12:56:36Z","receivedAt":"2021-03-30T12:57:43Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Mar 30 2021, Atharva Raykar wrote:\n\n> On 29-Mar-2021, at 15:38, Phillip Wood <phillip.wood123@gmail.com> wrote:\n>> On 28/03/2021 13:40, Atharva Raykar wrote:\n>>> On 28-Mar-2021, at 08:46, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>>>> The \"define-?.*\" can be simplified to just \"define.*\", but looking at\n>>>> the tests is that the intent? From the tests it looks like \"define[- ]\"\n>>>> is what the author wants, unless this is meant to also match\n>>>> \"(definements\".\n>>> Yes, you captured my intent correctly. Will fix it.\n>>>> Has this been tested on some real-world scheme code? E.g. I have guile\n>>>> installed locally, and it has really large top-level eval-when\n>>>> blocks. These rules would jump over those to whatever the function above\n>>>> them is.\n>>> I do not have a large scheme codebase on my own, I usually use Racket,\n>>> which is a much larger language with many more forms. Other Schemes like\n>>> Guile also extend the language a lot, like in your example, eval-when is\n>>> an extension provided by Guile (and Chicken and Chez), but not a part of\n>>> the R6RS document when I searched its index.\n>>> So the 'define' forms are the only one that I know would reliably be present\n>>> across all schemes. But one can also make a case where some of these non-standard\n>>> forms may be common enough that they are worth adding in. In that case which\n>>> forms to include? Should we consider everything in the SRFI's[1]? Should the\n>>> various module definitions of Racket be included? It's a little tricky to know\n>>> where to stop.\n>> \n>> If there are some common forms such as eval-when then it would be good to include them, otherwise we end up needing a different rule for each scheme implementation as they all seem to tweak something. Gerbil uses 'def...' e.g def, defsyntax, defstruct, defrules rather than define, define-syntax, define-record etc. I'm not user if we want to accommodate that or not.\n>\n> Yes, this is the part that is hard for me to figure out. I am going by\n> two heuristics: what Scheme communities in other places would generally\n> prefer, and what patterns I see happen more often in scheme code.\n>\n> The former is tricky to do. I posted to a few mailing lists about this,\n> but they don't seem active enough to garner any responses.\n>\n> The latter is a little easier to measure quickly. I did a GitHub search,\n> where I filtered results to only consider Scheme files (language:scheme).\n>\n> Some armchair stats, just for a broad understanding:\n>\n>   Total number of scheme files: 529,339\n>   No. of times a construct is used in those files:\n>     define and its variants : 431,090 (81.4%)\n>     def and its variants    :  18,466 ( 3.5%)\n>     eval-when               :   3,375 ( 0.6%)\n>\n> There was no way for me to quickly know which of these uses are at the top\n> level, but either way of the more structural forms that do show up in Scheme\n> code, define and its variants seem like a clear winner. I am not sure if\n> it's worth adding more rules to check for def and its variants, given that\n> they are not nearly as common.\n\nIn those cases we should veer on the side of inclusion. The only problem\nwe'll have is if \"eval-when\" is a \"setq\"-like function top-level form in\nsome other scheme dialect, so we'll have a conflict.\n\nOtherwise it's fine, programs that only use \"define\" won't be bothered\nby an eval-when rule.\n"},{"id":"420601","messageId":"88DF47D5-F104-4677-A2E4-7B23FEFCF022@gmail.com","threadId":"55399","inReplyTo":"874kgsn6kb.fsf@evledraar.gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-03-30T13:48:04Z","receivedAt":"2021-03-30T13:48:55Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"On 30-Mar-2021, at 18:26, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n> \n> \n> On Tue, Mar 30 2021, Atharva Raykar wrote:\n> \n>> On 29-Mar-2021, at 15:38, Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>> On 28/03/2021 13:40, Atharva Raykar wrote:\n>>>> On 28-Mar-2021, at 08:46, Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>>>>> The \"define-?.*\" can be simplified to just \"define.*\", but looking at\n>>>>> the tests is that the intent? From the tests it looks like \"define[- ]\"\n>>>>> is what the author wants, unless this is meant to also match\n>>>>> \"(definements\".\n>>>> Yes, you captured my intent correctly. Will fix it.\n>>>>> Has this been tested on some real-world scheme code? E.g. I have guile\n>>>>> installed locally, and it has really large top-level eval-when\n>>>>> blocks. These rules would jump over those to whatever the function above\n>>>>> them is.\n>>>> I do not have a large scheme codebase on my own, I usually use Racket,\n>>>> which is a much larger language with many more forms. Other Schemes like\n>>>> Guile also extend the language a lot, like in your example, eval-when is\n>>>> an extension provided by Guile (and Chicken and Chez), but not a part of\n>>>> the R6RS document when I searched its index.\n>>>> So the 'define' forms are the only one that I know would reliably be present\n>>>> across all schemes. But one can also make a case where some of these non-standard\n>>>> forms may be common enough that they are worth adding in. In that case which\n>>>> forms to include? Should we consider everything in the SRFI's[1]? Should the\n>>>> various module definitions of Racket be included? It's a little tricky to know\n>>>> where to stop.\n>>> \n>>> If there are some common forms such as eval-when then it would be good to include them, otherwise we end up needing a different rule for each scheme implementation as they all seem to tweak something. Gerbil uses 'def...' e.g def, defsyntax, defstruct, defrules rather than define, define-syntax, define-record etc. I'm not user if we want to accommodate that or not.\n>> \n>> Yes, this is the part that is hard for me to figure out. I am going by\n>> two heuristics: what Scheme communities in other places would generally\n>> prefer, and what patterns I see happen more often in scheme code.\n>> \n>> The former is tricky to do. I posted to a few mailing lists about this,\n>> but they don't seem active enough to garner any responses.\n>> \n>> The latter is a little easier to measure quickly. I did a GitHub search,\n>> where I filtered results to only consider Scheme files (language:scheme).\n>> \n>> Some armchair stats, just for a broad understanding:\n>> \n>>  Total number of scheme files: 529,339\n>>  No. of times a construct is used in those files:\n>>    define and its variants : 431,090 (81.4%)\n>>    def and its variants    :  18,466 ( 3.5%)\n>>    eval-when               :   3,375 ( 0.6%)\n>> \n>> There was no way for me to quickly know which of these uses are at the top\n>> level, but either way of the more structural forms that do show up in Scheme\n>> code, define and its variants seem like a clear winner. I am not sure if\n>> it's worth adding more rules to check for def and its variants, given that\n>> they are not nearly as common.\n> \n> In those cases we should veer on the side of inclusion. The only problem\n> we'll have is if \"eval-when\" is a \"setq\"-like function top-level form in\n> some other scheme dialect, so we'll have a conflict.\n> \n> Otherwise it's fine, programs that only use \"define\" won't be bothered\n> by an eval-when rule.\n\nI would like some clarification, since my knowledge of Common Lisp's setq\nand Guile's/Other's eval-when is pretty surface level.\n\n> The only problem we'll have is if \"eval-when\" is a \"setq\"-like function\n> top-level form in some other scheme dialect, so we'll have a conflict.\n\nI am not sure what you mean when you say if \"eval-when\" is a \"setq\"-like\ntop level form, and exactly what kind of problem it may cause.\n\nI also realized from my understanding of the Guile Documentation[1],\nthat \"eval-when\" is used to tell the compiler which expressions should be\nmade available during the expansion phase.\n\nIt does not seem to have anything that may help identify the location of the\nhunk, which I understand is the primary purpose of these hunk headers.\nAll uses of \"eval-when\" would be some variation of:\n\n\t(eval-when (expand load eval) ; no identifier in this form\n\t  ...)\n\nunlike a \"define\" which will always name the nearest function, which helps as\na landmark.\n\nWould that be a valid reason to exclude \"eval-when\"?\n\n"},{"id":"420932","messageId":"20210403131612.97194-1-raykar.ath@gmail.com","threadId":"55399","inReplyTo":"20210327173938.59391-1-raykar.ath@gmail.com","subject":"[GSoC][PATCH v2 0/1] userdiff: add support for scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-04-03T13:16:11Z","receivedAt":"2021-04-03T13:17:14Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"Hello all,\n\nThis is v2 of the patch I sent to add userdiff support to Scheme. I have\nmodified my approach to be more inclusive of as many Schemes as possible and\nthus added many more forms. Since Ævar suggested we veer on the side of\ninclusion when talking about the Gerbil scheme syntax, I felt it made sense to\nextend the driver to include common, but non-standard extensions of Scheme,\nincluding the forms of Racket and Guile.\n\nForms added since last patch:\n\n - Variants of define such as def, defsyntax etc as suggested by Phillip and the\n   other reviewers. - The library form of R6RS, as well as module definitions of\n   Racket, Gerbil and Guile. * These forms were added on the recommendation of\n   Göran Weinholt, creator of the Akku scheme package manager and the Loko\n   Scheme implementation:\n   https://groups.google.com/g/comp.lang.scheme/c/Aczn0TNEr5g/m/Jq3AlKvZBgAJ -\n   The Racket forms for defining structs and classes.\n\nI have restricted the \"def\" forms to only certain keywords, so that it does not\nover-match to words like \"deflate\", \"deform\", \"defer\" etc.\n\nI have also allowed the use of \"/\" after a define or def form, as some scheme\ncode uses it as a convention for defines in a certain context, such as Racket's\n\"define/public\".\n\nI have also fixed some a test case which had a redundant \"ChangeMe\".\n\nFinally in the word regex, which has been simplified a lot, while retaining the\nsame functioning after taking into account Junio's suggestions.\n\nAtharva Raykar (1):\n  userdiff: add support for Scheme\n\n Documentation/gitattributes.txt    |  2 ++\n t/t4018-diff-funcname.sh           |  1 +\n t/t4018/scheme-class               |  7 +++++++\n t/t4018/scheme-def                 |  4 ++++\n t/t4018/scheme-def-variant         |  4 ++++\n t/t4018/scheme-define-slash-public |  7 +++++++\n t/t4018/scheme-define-syntax       |  8 ++++++++\n t/t4018/scheme-define-variant      |  4 ++++\n t/t4018/scheme-library             | 11 +++++++++++\n t/t4018/scheme-local-define        |  4 ++++\n t/t4018/scheme-module              |  6 ++++++\n t/t4018/scheme-top-level-define    |  4 ++++\n t/t4018/scheme-user-defined-define |  6 ++++++\n t/t4034-diff-words.sh              |  1 +\n t/t4034/scheme/expect              | 10 ++++++++++\n t/t4034/scheme/post                |  5 +++++\n t/t4034/scheme/pre                 |  5 +++++\n userdiff.c                         |  4 ++++\n 18 files changed, 93 insertions(+)\n create mode 100644 t/t4018/scheme-class\n create mode 100644 t/t4018/scheme-def\n create mode 100644 t/t4018/scheme-def-variant\n create mode 100644 t/t4018/scheme-define-slash-public\n create mode 100644 t/t4018/scheme-define-syntax\n create mode 100644 t/t4018/scheme-define-variant\n create mode 100644 t/t4018/scheme-library\n create mode 100644 t/t4018/scheme-local-define\n create mode 100644 t/t4018/scheme-module\n create mode 100644 t/t4018/scheme-top-level-define\n create mode 100644 t/t4018/scheme-user-defined-define\n create mode 100644 t/t4034/scheme/expect\n create mode 100644 t/t4034/scheme/post\n create mode 100644 t/t4034/scheme/pre\n\n-- \n2.31.1\n\n"},{"id":"420933","messageId":"20210403131612.97194-2-raykar.ath@gmail.com","threadId":"55399","inReplyTo":"20210403131612.97194-1-raykar.ath@gmail.com","subject":"[GSoC][PATCH v2 1/1] userdiff: add support for scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-04-03T13:16:12Z","receivedAt":"2021-04-03T13:17:36Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"Add a diff driver for Scheme-like languages which recognizes top level\nand local `define` forms, whether it is a function definition, binding,\nsyntax definition or a user-defined `define-xyzzy` form.\n\nAlso supports R6RS `library` forms, `module` forms along with class and\nstruct declarations used in Racket (PLT Scheme).\n\nAlternate \"def\" syntax such as those in Gerbil Scheme are also\nsupported, like defstruct, defsyntax and so on.\n\nThe rationale for picking `define` forms for the hunk headers is because\nit is usually the only significant form for defining the structure of\nthe program, and it is a common pattern for schemers to have local\nfunction definitions to hide their visibility, so it is not only the top\nlevel `define`'s that are of interest. Schemers also extend the language\nwith macros to provide their own define forms (for example, something\nlike a `define-test-suite`) which is also captured in the hunk header.\n\nSince it is common practice to extend syntax with variants of a form\nlike `module+`, `class*` etc, those have been supported as well.\n\nThe word regex is a best-effort attempt to conform to R6RS[1] valid\nidentifiers, symbols and numbers.\n\n[1] http://www.r6rs.org/final/html/r6rs/r6rs-Z-H-7.html#node_chap_4\n\nSigned-off-by: Atharva Raykar <raykar.ath@gmail.com>\n---\n Documentation/gitattributes.txt    |  2 ++\n t/t4018-diff-funcname.sh           |  1 +\n t/t4018/scheme-class               |  7 +++++++\n t/t4018/scheme-def                 |  4 ++++\n t/t4018/scheme-def-variant         |  4 ++++\n t/t4018/scheme-define-slash-public |  7 +++++++\n t/t4018/scheme-define-syntax       |  8 ++++++++\n t/t4018/scheme-define-variant      |  4 ++++\n t/t4018/scheme-library             | 11 +++++++++++\n t/t4018/scheme-local-define        |  4 ++++\n t/t4018/scheme-module              |  6 ++++++\n t/t4018/scheme-top-level-define    |  4 ++++\n t/t4018/scheme-user-defined-define |  6 ++++++\n t/t4034-diff-words.sh              |  1 +\n t/t4034/scheme/expect              | 10 ++++++++++\n t/t4034/scheme/post                |  5 +++++\n t/t4034/scheme/pre                 |  5 +++++\n userdiff.c                         |  4 ++++\n 18 files changed, 93 insertions(+)\n create mode 100644 t/t4018/scheme-class\n create mode 100644 t/t4018/scheme-def\n create mode 100644 t/t4018/scheme-def-variant\n create mode 100644 t/t4018/scheme-define-slash-public\n create mode 100644 t/t4018/scheme-define-syntax\n create mode 100644 t/t4018/scheme-define-variant\n create mode 100644 t/t4018/scheme-library\n create mode 100644 t/t4018/scheme-local-define\n create mode 100644 t/t4018/scheme-module\n create mode 100644 t/t4018/scheme-top-level-define\n create mode 100644 t/t4018/scheme-user-defined-define\n create mode 100644 t/t4034/scheme/expect\n create mode 100644 t/t4034/scheme/post\n create mode 100644 t/t4034/scheme/pre\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 0a60472bb5..cfcfa800c2 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -845,6 +845,8 @@ patterns are available:\n \n - `rust` suitable for source code in the Rust language.\n \n+- `scheme` suitable for source code in the Scheme language.\n+\n - `tex` suitable for source code for LaTeX documents.\n \n \ndiff --git a/t/t4018-diff-funcname.sh b/t/t4018-diff-funcname.sh\nindex 9675bc17db..823ea96acb 100755\n--- a/t/t4018-diff-funcname.sh\n+++ b/t/t4018-diff-funcname.sh\n@@ -48,6 +48,7 @@ diffpatterns=\"\n \tpython\n \truby\n \trust\n+\tscheme\n \ttex\n \tcustom1\n \tcustom2\ndiff --git a/t/t4018/scheme-class b/t/t4018/scheme-class\nnew file mode 100644\nindex 0000000000..e5e07b43fb\n--- /dev/null\n+++ b/t/t4018/scheme-class\n@@ -0,0 +1,7 @@\n+(define book-class%\n+  (class* () object% RIGHT\n+    (field (pages 5))\n+    (field (ChangeMe 5))\n+    (define/public (letters)\n+      (* pages 500))\n+    (super-new)))\ndiff --git a/t/t4018/scheme-def b/t/t4018/scheme-def\nnew file mode 100644\nindex 0000000000..1e2673da96\n--- /dev/null\n+++ b/t/t4018/scheme-def\n@@ -0,0 +1,4 @@\n+(def (some-func x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\ndiff --git a/t/t4018/scheme-def-variant b/t/t4018/scheme-def-variant\nnew file mode 100644\nindex 0000000000..d857a61d64\n--- /dev/null\n+++ b/t/t4018/scheme-def-variant\n@@ -0,0 +1,4 @@\n+(defmethod {print point} RIGHT\n+  (lambda (self)\n+    (with ((point x y) self)\n+      (printf \"{ChangeMe x:~a y:~a}~n\" x y))))\ndiff --git a/t/t4018/scheme-define-slash-public b/t/t4018/scheme-define-slash-public\nnew file mode 100644\nindex 0000000000..39a93a1600\n--- /dev/null\n+++ b/t/t4018/scheme-define-slash-public\n@@ -0,0 +1,7 @@\n+(define bar-class%\n+  (class object%\n+    (field (info 5))\n+    (define/public (foo) RIGHT\n+      (+ info 42)\n+      (* info ChangeMe))\n+    (super-new)))\ndiff --git a/t/t4018/scheme-define-syntax b/t/t4018/scheme-define-syntax\nnew file mode 100644\nindex 0000000000..7d5e99e0fc\n--- /dev/null\n+++ b/t/t4018/scheme-define-syntax\n@@ -0,0 +1,8 @@\n+(define-syntax define-test-suite RIGHT\n+  (syntax-rules ()\n+    ((_ suite-name (name test) ChangeMe ...)\n+     (define suite-name\n+       (let ((tests\n+              `((name . ,test) ...)))\n+         (lambda ()\n+           (run-suite 'suite-name tests)))))))\ndiff --git a/t/t4018/scheme-define-variant b/t/t4018/scheme-define-variant\nnew file mode 100644\nindex 0000000000..911708854d\n--- /dev/null\n+++ b/t/t4018/scheme-define-variant\n@@ -0,0 +1,4 @@\n+(define* (some-func x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\ndiff --git a/t/t4018/scheme-library b/t/t4018/scheme-library\nnew file mode 100644\nindex 0000000000..82ea3df510\n--- /dev/null\n+++ b/t/t4018/scheme-library\n@@ -0,0 +1,11 @@\n+(library (my-helpers id-stuff) RIGHT\n+  (export find-dup)\n+  (import (ChangeMe))\n+  (define (find-dup l)\n+    (and (pair? l)\n+         (let loop ((rest (cdr l)))\n+           (cond\n+            [(null? rest) (find-dup (cdr l))]\n+            [(bound-identifier=? (car l) (car rest))\n+             (car rest)]\n+            [else (loop (cdr rest))])))))\ndiff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\nnew file mode 100644\nindex 0000000000..bc6d8aebbe\n--- /dev/null\n+++ b/t/t4018/scheme-local-define\n@@ -0,0 +1,4 @@\n+(define (higher-order)\n+  (define local-function RIGHT\n+    (lambda (x)\n+     (car \"this is\" \"ChangeMe\"))))\ndiff --git a/t/t4018/scheme-module b/t/t4018/scheme-module\nnew file mode 100644\nindex 0000000000..edfae0ebf7\n--- /dev/null\n+++ b/t/t4018/scheme-module\n@@ -0,0 +1,6 @@\n+(module A RIGHT\n+  (export with-display-exception)\n+  (extern (display-exception display-exception ChangeMe))\n+  (def (with-display-exception thunk)\n+    (with-catch (lambda (e) (display-exception e (current-error-port)) e)\n+      thunk)))\ndiff --git a/t/t4018/scheme-top-level-define b/t/t4018/scheme-top-level-define\nnew file mode 100644\nindex 0000000000..624743c22b\n--- /dev/null\n+++ b/t/t4018/scheme-top-level-define\n@@ -0,0 +1,4 @@\n+(define (some-func x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\ndiff --git a/t/t4018/scheme-user-defined-define b/t/t4018/scheme-user-defined-define\nnew file mode 100644\nindex 0000000000..35fe7cc9bf\n--- /dev/null\n+++ b/t/t4018/scheme-user-defined-define\n@@ -0,0 +1,6 @@\n+(define-test-suite record\\ case-tests RIGHT\n+  (record-case-1 (lambda (fail)\n+                   (let ((a (make-foo 1 2)))\n+                     (record-case a\n+                       ((bar x) (ChangeMe))\n+                       ((foo a b) (+ a b)))))))\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 56f1e62a97..ee7721ab91 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -325,6 +325,7 @@ test_language_driver perl\n test_language_driver php\n test_language_driver python\n test_language_driver ruby\n+test_language_driver scheme\n test_language_driver tex\n \n test_expect_success 'word-diff with diff.sbe' '\ndiff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\nnew file mode 100644\nindex 0000000000..d9bb82225b\n--- /dev/null\n+++ b/t/t4034/scheme/expect\n@@ -0,0 +1,10 @@\n+<BOLD>diff --git a/pre b/post<RESET>\n+<BOLD>index cb56df5..09d9506 100644<RESET>\n+<BOLD>--- a/pre<RESET>\n+<BOLD>+++ b/post<RESET>\n+<CYAN>@@ -1,5 +1,5 @@<RESET>\n+(define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n+  ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n+  (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n+  (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n+    (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\ndiff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\nnew file mode 100644\nindex 0000000000..06bf0b34f9\n--- /dev/null\n+++ b/t/t4034/scheme/post\n@@ -0,0 +1,5 @@\n+(define (my-func first second)\n+  ; This is a (moderately) cool function.\n+  (that\\place (+ 3 4))\n+  (let ((c (add1 first)))\n+    (format \"one more than the total is %d\" (+ c second))))\ndiff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\nnew file mode 100644\nindex 0000000000..270a2d0cc5\n--- /dev/null\n+++ b/t/t4034/scheme/pre\n@@ -0,0 +1,5 @@\n+(define (myfunc a b)\n+  ; This is a really cool function.\n+  (this\\place (+ 3 4))\n+  (let ((c (+ a b)))\n+    (format \"one more than the total is %d\" (add1 c))))\ndiff --git a/userdiff.c b/userdiff.c\nindex 3f81a2261c..ac1999bbc5 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -191,6 +191,10 @@ PATTERNS(\"rust\",\n \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n+PATTERNS(\"scheme\",\n+\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\",\n+\t /* All words should be delimited by spaces or parentheses */\n+\t \"([^][)(}{[ \\t])+\"),\n PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n-- \n2.31.1\n\n"},{"id":"420977","messageId":"61622cda-3ce5-7cd9-acd6-54906297500c@gmail.com","threadId":"55399","inReplyTo":"A3C3DD12-3C00-49ED-B427-37AAB4211C2A@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2021-04-05T10:04:39Z","receivedAt":"2021-04-05T10:04:47Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Atharva\n\nOn 30/03/2021 11:22, Atharva Raykar wrote:\n> \n> \n>> On 30-Mar-2021, at 12:34, Atharva Raykar <raykar.ath@gmail.com> wrote:\n>>\n>>\n>>\n>>> On 29-Mar-2021, at 15:48, Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>>\n>>> Hi Atharva\n>>>\n>>> On 28/03/2021 13:23, Atharva Raykar wrote:\n>>>> On 28-Mar-2021, at 05:16, Johannes Sixt <j6t@kdbg.org> wrote:\n>>>> [...]\n>>>>>> diff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\n>>>>>> new file mode 100644\n>>>>>> index 0000000000..90e75dcce8\n>>>>>> --- /dev/null\n>>>>>> +++ b/t/t4018/scheme-local-define\n>>>>>> @@ -0,0 +1,4 @@\n>>>>>> +(define (higher-order)\n>>>>>> +  (define local-function RIGHT\n>>>>>\n>>>>> ... this one, which is also indented and *is* marked as RIGHT.\n>>>> In this test case, I was explicitly testing for an indented '(define'\n>>>> whereas in the former, I was testing for the top-level '(define-syntax',\n>>>> which happened to have an internal define (which will inevitably show up\n>>>> in a lot of scheme code).\n>>>\n>>> It would be nice to include indented define forms but including them means that any change to the body of a function is attributed to the last internal definition rather than the actual function. For example\n>>>\n>>> (define (f arg)\n>>> (define (g x)\n>>>    (+ 1 x))\n>>>\n>>> (some-func ...)\n>>> ;;any change here will have '(define (g x)' in the hunk header, not '(define (f arg)'\n>>\n>> The reason I went for this over the top level forms, is because\n>> I felt it was useful to see the nearest definition for internal\n>> functions that often have a lot of the actual business logic of\n>> the program (at least a lot of SICP seems to follow this pattern).\n>> The disadvantage is as you said, it might also catch trivial inner\n>> functions and the developer might lose context.\n> \n> Never mind this message, I had misunderstood the problem you were trying to\n> demonstrate. I wholeheartedly agree with what you are trying to say, and\n> the indentation heuristic discussed does look interesting. I shall have a\n> glance at the RFC you linked in the other reply.\n> \n>> The disadvantage is as you said, it might also catch trivial inner\n>> functions and the developer might lose context.\n> \n> Feel free to disregard me misquoting you here. You did not say that (:\n> \n>> Another problem is it may match more trivial bindings, like:\n>>\n>> (define (some-func things)\n>>   ...\n>>   (define items '(eggs\n>>                   ham\n>>                   peanut-butter))\n>>   ...)\n>>\n>> What I have noticed *anecdotally* is that this is not common enough\n>> to be too much of a problem, and local define bindings seem to be more\n>> favoured in Racket than other Schemes, that use 'let' more often.\n>>\n>>> I don't think this can be avoided as we rely on regexs rather than parsing the source so it is probably best to only match toplevel defines.\n>>\n>> The other issue with only matching top level defines is that a\n>> lot of scheme programs are library definitions, something like\n>>\n>> (library\n>>     (foo bar)\n>>   (export ...)\n>>   (define ...)\n>>   (define ...)\n>>   ;; and a bunch of other definitions...\n>> )\n>>\n>> Only matching top level defines will completely ignore matching all\n>> the definitions in these files.\n> \n> That said, I still stand by the fact that only catching top level defines\n> will lead to a lot of definitions being ignored. Maybe the occasional\n> mismatch may be worth the gain in the number of function contexts being\n> detected?\n\nI'm not sure that the mismatches will be occasional - every time you \nhave an internal definition in a function the hunk header will be wrong \nwhen you change the main body of the function. This will affect grep \n--function-context and diff -W as well as the normal hunk headers. The \nproblem is there is no way to avoid that and provide something useful in \nthe library example you have above. It would be useful to find some code \nbases and diff the output of 'git log --patch' with and without the \nleading whitespace match in the function pattern to see how often this \nis a problem (i.e. when the funcnames do not match see which one is \ncorrect).\n\nBest Wishes\n\nPhillip\n\n\n"},{"id":"420978","messageId":"4a9bdf0c-dc0f-a0fa-5c13-2b4732d21ca8@gmail.com","threadId":"55399","inReplyTo":"20210403131612.97194-2-raykar.ath@gmail.com","subject":"Re: [GSoC][PATCH v2 1/1] userdiff: add support for scheme","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2021-04-05T10:21:24Z","receivedAt":"2021-04-05T10:21:29Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Atharva\nOn 03/04/2021 14:16, Atharva Raykar wrote:\n> Add a diff driver for Scheme-like languages which recognizes top level\n> and local `define` forms, whether it is a function definition, binding,\n> syntax definition or a user-defined `define-xyzzy` form.\n> \n> Also supports R6RS `library` forms, `module` forms along with class and\n> struct declarations used in Racket (PLT Scheme).\n> \n> Alternate \"def\" syntax such as those in Gerbil Scheme are also\n> supported, like defstruct, defsyntax and so on.\n> \n> The rationale for picking `define` forms for the hunk headers is because\n> it is usually the only significant form for defining the structure of\n> the program, and it is a common pattern for schemers to have local\n> function definitions to hide their visibility, so it is not only the top\n> level `define`'s that are of interest. Schemers also extend the language\n> with macros to provide their own define forms (for example, something\n> like a `define-test-suite`) which is also captured in the hunk header.\n> \n> Since it is common practice to extend syntax with variants of a form\n> like `module+`, `class*` etc, those have been supported as well.\n> \n> The word regex is a best-effort attempt to conform to R6RS[1] valid\n> identifiers, symbols and numbers.\n> \n> [1] http://www.r6rs.org/final/html/r6rs/r6rs-Z-H-7.html#node_chap_4\n> \n> Signed-off-by: Atharva Raykar <raykar.ath@gmail.com>\n>[...]\n> diff --git a/userdiff.c b/userdiff.c\n> index 3f81a2261c..ac1999bbc5 100644\n> --- a/userdiff.c\n> +++ b/userdiff.c\n> @@ -191,6 +191,10 @@ PATTERNS(\"rust\",\n>   \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n>   \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>   \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n> +PATTERNS(\"scheme\",\n> +\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\",\n> +\t /* All words should be delimited by spaces or parentheses */\n> +\t \"([^][)(}{[ \\t])+\"),\n\nI think it would be nice to match single '(' and '[' to highlight when \nthey have been added or deleted - I find this useful when I get a syntax \nerror. Also it would be nice to handle r7rs identifiers like | this is a \nsymbol |. Maybe something like\n\"(\\\\|([^\\\\\\\\|]*(\\\\\\\\|)*)*\\\\||[^][}{)( \\t]|[][(){}])\"\n\nBest Wishes\n\nPhillip\n\n>   PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n>   \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n>   PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n> \n\n"},{"id":"421001","messageId":"01f17458-3d99-d4f5-aee8-0f77f73063d2@kdbg.org","threadId":"55399","inReplyTo":"61622cda-3ce5-7cd9-acd6-54906297500c@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2021-04-05T17:58:28Z","receivedAt":"2021-04-05T17:58:34Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 05.04.21 um 12:04 schrieb Phillip Wood:\n> Hi Atharva\n> \n> On 30/03/2021 11:22, Atharva Raykar wrote:\n>>\n>>\n>>> On 30-Mar-2021, at 12:34, Atharva Raykar <raykar.ath@gmail.com> wrote:\n>>>\n>>>\n>>>\n>>>> On 29-Mar-2021, at 15:48, Phillip Wood <phillip.wood123@gmail.com>\n>>>> wrote:\n>>>>\n>>>> Hi Atharva\n>>>>\n>>>> On 28/03/2021 13:23, Atharva Raykar wrote:\n>>>>> On 28-Mar-2021, at 05:16, Johannes Sixt <j6t@kdbg.org> wrote:\n>>>>> [...]\n>>>>>>> diff --git a/t/t4018/scheme-local-define\n>>>>>>> b/t/t4018/scheme-local-define\n>>>>>>> new file mode 100644\n>>>>>>> index 0000000000..90e75dcce8\n>>>>>>> --- /dev/null\n>>>>>>> +++ b/t/t4018/scheme-local-define\n>>>>>>> @@ -0,0 +1,4 @@\n>>>>>>> +(define (higher-order)\n>>>>>>> +  (define local-function RIGHT\n>>>>>>\n>>>>>> ... this one, which is also indented and *is* marked as RIGHT.\n>>>>> In this test case, I was explicitly testing for an indented '(define'\n>>>>> whereas in the former, I was testing for the top-level\n>>>>> '(define-syntax',\n>>>>> which happened to have an internal define (which will inevitably\n>>>>> show up\n>>>>> in a lot of scheme code).\n>>>>\n>>>> It would be nice to include indented define forms but including them\n>>>> means that any change to the body of a function is attributed to the\n>>>> last internal definition rather than the actual function. For example\n>>>>\n>>>> (define (f arg)\n>>>> (define (g x)\n>>>>    (+ 1 x))\n>>>>\n>>>> (some-func ...)\n>>>> ;;any change here will have '(define (g x)' in the hunk header, not\n>>>> '(define (f arg)'\n>>>\n>>> The reason I went for this over the top level forms, is because\n>>> I felt it was useful to see the nearest definition for internal\n>>> functions that often have a lot of the actual business logic of\n>>> the program (at least a lot of SICP seems to follow this pattern).\n>>> The disadvantage is as you said, it might also catch trivial inner\n>>> functions and the developer might lose context.\n>>\n>> Never mind this message, I had misunderstood the problem you were\n>> trying to\n>> demonstrate. I wholeheartedly agree with what you are trying to say, and\n>> the indentation heuristic discussed does look interesting. I shall have a\n>> glance at the RFC you linked in the other reply.\n>>\n>>> The disadvantage is as you said, it might also catch trivial inner\n>>> functions and the developer might lose context.\n>>\n>> Feel free to disregard me misquoting you here. You did not say that (:\n>>\n>>> Another problem is it may match more trivial bindings, like:\n>>>\n>>> (define (some-func things)\n>>>   ...\n>>>   (define items '(eggs\n>>>                   ham\n>>>                   peanut-butter))\n>>>   ...)\n>>>\n>>> What I have noticed *anecdotally* is that this is not common enough\n>>> to be too much of a problem, and local define bindings seem to be more\n>>> favoured in Racket than other Schemes, that use 'let' more often.\n>>>\n>>>> I don't think this can be avoided as we rely on regexs rather than\n>>>> parsing the source so it is probably best to only match toplevel\n>>>> defines.\n>>>\n>>> The other issue with only matching top level defines is that a\n>>> lot of scheme programs are library definitions, something like\n>>>\n>>> (library\n>>>     (foo bar)\n>>>   (export ...)\n>>>   (define ...)\n>>>   (define ...)\n>>>   ;; and a bunch of other definitions...\n>>> )\n>>>\n>>> Only matching top level defines will completely ignore matching all\n>>> the definitions in these files.\n>>\n>> That said, I still stand by the fact that only catching top level defines\n>> will lead to a lot of definitions being ignored. Maybe the occasional\n>> mismatch may be worth the gain in the number of function contexts being\n>> detected?\n> \n> I'm not sure that the mismatches will be occasional - every time you\n> have an internal definition in a function the hunk header will be wrong\n> when you change the main body of the function. This will affect grep\n> --function-context and diff -W as well as the normal hunk headers. The\n> problem is there is no way to avoid that and provide something useful in\n> the library example you have above. It would be useful to find some code\n> bases and diff the output of 'git log --patch' with and without the\n> leading whitespace match in the function pattern to see how often this\n> is a problem (i.e. when the funcnames do not match see which one is\n> correct).\n\n--function-context is just one application of the function matcher. To\nwork properly with nested function definitions, it would have to\nunderstand the nesting. But it does not; there is nothing that we can do\nabout it without a proper language parser. Therefore, the argument that\nthe matcher does not work well with --function-context for nested\nfunctions is of little relevance.\n\nIMO, the primary concern should be whether the matcher decorates hunk\ncontexts sufficiently well.\n\n-- Hannes\n"},{"id":"421052","messageId":"86EBA5B6-2DAE-4D4C-BCCD-46279595FA75@gmail.com","threadId":"55399","inReplyTo":"4a9bdf0c-dc0f-a0fa-5c13-2b4732d21ca8@gmail.com","subject":"Re: [GSoC][PATCH v2 1/1] userdiff: add support for scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-04-06T10:32:54Z","receivedAt":"2021-04-06T10:33:04Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"On 05-Apr-2021, at 15:51, Phillip Wood <phillip.wood123@gmail.com> wrote:\n> \n> Hi Atharva\n> On 03/04/2021 14:16, Atharva Raykar wrote:\n>> Add a diff driver for Scheme-like languages which recognizes top level\n>> and local `define` forms, whether it is a function definition, binding,\n>> syntax definition or a user-defined `define-xyzzy` form.\n>> Also supports R6RS `library` forms, `module` forms along with class and\n>> struct declarations used in Racket (PLT Scheme).\n>> Alternate \"def\" syntax such as those in Gerbil Scheme are also\n>> supported, like defstruct, defsyntax and so on.\n>> The rationale for picking `define` forms for the hunk headers is because\n>> it is usually the only significant form for defining the structure of\n>> the program, and it is a common pattern for schemers to have local\n>> function definitions to hide their visibility, so it is not only the top\n>> level `define`'s that are of interest. Schemers also extend the language\n>> with macros to provide their own define forms (for example, something\n>> like a `define-test-suite`) which is also captured in the hunk header.\n>> Since it is common practice to extend syntax with variants of a form\n>> like `module+`, `class*` etc, those have been supported as well.\n>> The word regex is a best-effort attempt to conform to R6RS[1] valid\n>> identifiers, symbols and numbers.\n>> [1] http://www.r6rs.org/final/html/r6rs/r6rs-Z-H-7.html#node_chap_4\n>> Signed-off-by: Atharva Raykar <raykar.ath@gmail.com>\n>> [...]\n>> diff --git a/userdiff.c b/userdiff.c\n>> index 3f81a2261c..ac1999bbc5 100644\n>> --- a/userdiff.c\n>> +++ b/userdiff.c\n>> @@ -191,6 +191,10 @@ PATTERNS(\"rust\",\n>>  \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n>>  \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>>  \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n>> +PATTERNS(\"scheme\",\n>> +\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\",\n>> +\t /* All words should be delimited by spaces or parentheses */\n>> +\t \"([^][)(}{[ \\t])+\"),\n> \n> I think it would be nice to match single '(' and '[' to highlight when they have been added or deleted - I find this useful when I get a syntax error. Also it would be nice to handle r7rs identifiers like | this is a symbol |. Maybe something like\n> \"(\\\\|([^\\\\\\\\|]*(\\\\\\\\|)*)*\\\\||[^][}{)( \\t]|[][(){}])\"\n\nMy patch seems to detect additions and removals of singular parentheses\nalready -- I am not sure why it works, but my suspicion is that the\nuserdiff code seems to fall back to some default rules for additions and\nremovals that do not match the current word regex? Either way that seems\nto work.\n\nAs for the R7RS identifiers, I can definitely add that, thanks for\npointing that out!\n\n> Best Wishes\n> \n> Phillip\n> \n>>  PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n>>  \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n>>  PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n\n"},{"id":"421059","messageId":"60D8CE48-926B-4A09-9355-4331C14F6753@gmail.com","threadId":"55399","inReplyTo":"61622cda-3ce5-7cd9-acd6-54906297500c@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-04-06T12:29:43Z","receivedAt":"2021-04-06T12:29:52Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"On 05-Apr-2021, at 15:34, Phillip Wood <phillip.wood123@gmail.com> wrote:\n> \n> Hi Atharva\n> \n> On 30/03/2021 11:22, Atharva Raykar wrote:\n>>> On 30-Mar-2021, at 12:34, Atharva Raykar <raykar.ath@gmail.com> wrote:\n>>> \n>>> \n>>> \n>>>> On 29-Mar-2021, at 15:48, Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>>> \n>>>> Hi Atharva\n>>>> \n>>>> On 28/03/2021 13:23, Atharva Raykar wrote:\n>>>>> On 28-Mar-2021, at 05:16, Johannes Sixt <j6t@kdbg.org> wrote:\n>>>>> [...]\n>>>>>>> diff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\n>>>>>>> new file mode 100644\n>>>>>>> index 0000000000..90e75dcce8\n>>>>>>> --- /dev/null\n>>>>>>> +++ b/t/t4018/scheme-local-define\n>>>>>>> @@ -0,0 +1,4 @@\n>>>>>>> +(define (higher-order)\n>>>>>>> +  (define local-function RIGHT\n>>>>>> \n>>>>>> ... this one, which is also indented and *is* marked as RIGHT.\n>>>>> In this test case, I was explicitly testing for an indented '(define'\n>>>>> whereas in the former, I was testing for the top-level '(define-syntax',\n>>>>> which happened to have an internal define (which will inevitably show up\n>>>>> in a lot of scheme code).\n>>>> \n>>>> It would be nice to include indented define forms but including them means that any change to the body of a function is attributed to the last internal definition rather than the actual function. For example\n>>>> \n>>>> (define (f arg)\n>>>> (define (g x)\n>>>>   (+ 1 x))\n>>>> \n>>>> (some-func ...)\n>>>> ;;any change here will have '(define (g x)' in the hunk header, not '(define (f arg)'\n>>> \n>>> The reason I went for this over the top level forms, is because\n>>> I felt it was useful to see the nearest definition for internal\n>>> functions that often have a lot of the actual business logic of\n>>> the program (at least a lot of SICP seems to follow this pattern).\n>>> The disadvantage is as you said, it might also catch trivial inner\n>>> functions and the developer might lose context.\n>> Never mind this message, I had misunderstood the problem you were trying to\n>> demonstrate. I wholeheartedly agree with what you are trying to say, and\n>> the indentation heuristic discussed does look interesting. I shall have a\n>> glance at the RFC you linked in the other reply.\n>>> The disadvantage is as you said, it might also catch trivial inner\n>>> functions and the developer might lose context.\n>> Feel free to disregard me misquoting you here. You did not say that (:\n>>> Another problem is it may match more trivial bindings, like:\n>>> \n>>> (define (some-func things)\n>>>  ...\n>>>  (define items '(eggs\n>>>                  ham\n>>>                  peanut-butter))\n>>>  ...)\n>>> \n>>> What I have noticed *anecdotally* is that this is not common enough\n>>> to be too much of a problem, and local define bindings seem to be more\n>>> favoured in Racket than other Schemes, that use 'let' more often.\n>>> \n>>>> I don't think this can be avoided as we rely on regexs rather than parsing the source so it is probably best to only match toplevel defines.\n>>> \n>>> The other issue with only matching top level defines is that a\n>>> lot of scheme programs are library definitions, something like\n>>> \n>>> (library\n>>>    (foo bar)\n>>>  (export ...)\n>>>  (define ...)\n>>>  (define ...)\n>>>  ;; and a bunch of other definitions...\n>>> )\n>>> \n>>> Only matching top level defines will completely ignore matching all\n>>> the definitions in these files.\n>> That said, I still stand by the fact that only catching top level defines\n>> will lead to a lot of definitions being ignored. Maybe the occasional\n>> mismatch may be worth the gain in the number of function contexts being\n>> detected?\n> \n> I'm not sure that the mismatches will be occasional - every time you have an internal definition in a function the hunk header will be wrong when you change the main body of the function. This will affect grep --function-context and diff -W as well as the normal hunk headers. The problem is there is no way to avoid that and provide something useful in the library example you have above. It would be useful to find some code bases and diff the output of 'git log --patch' with and without the leading whitespace match in the function pattern to see how often this is a problem (i.e. when the funcnames do not match see which one is correct).\n\nYou are right -- on trying out the function on a two other scheme\ncodebases, I noticed that there are a lot more wrongly matched functions\nthan I initially thought. About half of them identify the wrong function\nin one of the repositories I tried. However, removing the leading\nwhitespace in the pattern did not lead to better matching; it just led\nto a lot of the hunk headers going blank. I am not sure what causes this\nbehaviour, but my guess is that the function contexts are shown only if\nit is within a certain distance from the function definition?\n\nEven if it did match only the top level defines correctly, the functions\nmatched would still often be technically wrong -- it will show the outer\nfunction as the context when the user has edited an internal function\n(and in Scheme, there is heavy usage of internal functions).\n\nAfter running 'git grep --function-context' with the leading whitespace\nremoved, it seems to match too aggressively, as it captures a huge\nregion to match all the way upto the top level. Especially for files\nwhere all the definitions are in a 'library'.\n\nOverall, I personally felt that there were more downsides to matching\nonly at the top level. I'd rather the hunk header have the nearest\nfunction to provide the context, than have no function displayed at all.\nEven when the match is wrong, it at least helps me locate where the\nchange was made more easily.\n\n\n> Best Wishes\n> \n> Phillip\n> \n> \n\n"},{"id":"421088","messageId":"570c172a-0bc5-8a3d-eeca-c5dd81c84206@gmail.com","threadId":"55399","inReplyTo":"60D8CE48-926B-4A09-9355-4331C14F6753@gmail.com","subject":"Re: [GSOC][PATCH] userdiff: add support for Scheme","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2021-04-06T19:10:57Z","receivedAt":"2021-04-06T19:11:06Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"Hi Atharva\n\nOn 06/04/2021 13:29, Atharva Raykar wrote:\n> On 05-Apr-2021, at 15:34, Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>\n>> Hi Atharva\n>>\n>> On 30/03/2021 11:22, Atharva Raykar wrote:\n>>>> On 30-Mar-2021, at 12:34, Atharva Raykar <raykar.ath@gmail.com> wrote:\n>>>>\n>>>>\n>>>>\n>>>>> On 29-Mar-2021, at 15:48, Phillip Wood <phillip.wood123@gmail.com> wrote:\n>>>>>\n>>>>> Hi Atharva\n>>>>>\n>>>>> On 28/03/2021 13:23, Atharva Raykar wrote:\n>>>>>> On 28-Mar-2021, at 05:16, Johannes Sixt <j6t@kdbg.org> wrote:\n>>>>>> [...]\n>>>>>>>> diff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\n>>>>>>>> new file mode 100644\n>>>>>>>> index 0000000000..90e75dcce8\n>>>>>>>> --- /dev/null\n>>>>>>>> +++ b/t/t4018/scheme-local-define\n>>>>>>>> @@ -0,0 +1,4 @@\n>>>>>>>> +(define (higher-order)\n>>>>>>>> +  (define local-function RIGHT\n>>>>>>>\n>>>>>>> ... this one, which is also indented and *is* marked as RIGHT.\n>>>>>> In this test case, I was explicitly testing for an indented '(define'\n>>>>>> whereas in the former, I was testing for the top-level '(define-syntax',\n>>>>>> which happened to have an internal define (which will inevitably show up\n>>>>>> in a lot of scheme code).\n>>>>>\n>>>>> It would be nice to include indented define forms but including them means that any change to the body of a function is attributed to the last internal definition rather than the actual function. For example\n>>>>>\n>>>>> (define (f arg)\n>>>>> (define (g x)\n>>>>>    (+ 1 x))\n>>>>>\n>>>>> (some-func ...)\n>>>>> ;;any change here will have '(define (g x)' in the hunk header, not '(define (f arg)'\n>>>>\n>>>> The reason I went for this over the top level forms, is because\n>>>> I felt it was useful to see the nearest definition for internal\n>>>> functions that often have a lot of the actual business logic of\n>>>> the program (at least a lot of SICP seems to follow this pattern).\n>>>> The disadvantage is as you said, it might also catch trivial inner\n>>>> functions and the developer might lose context.\n>>> Never mind this message, I had misunderstood the problem you were trying to\n>>> demonstrate. I wholeheartedly agree with what you are trying to say, and\n>>> the indentation heuristic discussed does look interesting. I shall have a\n>>> glance at the RFC you linked in the other reply.\n>>>> The disadvantage is as you said, it might also catch trivial inner\n>>>> functions and the developer might lose context.\n>>> Feel free to disregard me misquoting you here. You did not say that (:\n>>>> Another problem is it may match more trivial bindings, like:\n>>>>\n>>>> (define (some-func things)\n>>>>   ...\n>>>>   (define items '(eggs\n>>>>                   ham\n>>>>                   peanut-butter))\n>>>>   ...)\n>>>>\n>>>> What I have noticed *anecdotally* is that this is not common enough\n>>>> to be too much of a problem, and local define bindings seem to be more\n>>>> favoured in Racket than other Schemes, that use 'let' more often.\n>>>>\n>>>>> I don't think this can be avoided as we rely on regexs rather than parsing the source so it is probably best to only match toplevel defines.\n>>>>\n>>>> The other issue with only matching top level defines is that a\n>>>> lot of scheme programs are library definitions, something like\n>>>>\n>>>> (library\n>>>>     (foo bar)\n>>>>   (export ...)\n>>>>   (define ...)\n>>>>   (define ...)\n>>>>   ;; and a bunch of other definitions...\n>>>> )\n>>>>\n>>>> Only matching top level defines will completely ignore matching all\n>>>> the definitions in these files.\n>>> That said, I still stand by the fact that only catching top level defines\n>>> will lead to a lot of definitions being ignored. Maybe the occasional\n>>> mismatch may be worth the gain in the number of function contexts being\n>>> detected?\n>>\n>> I'm not sure that the mismatches will be occasional - every time you have an internal definition in a function the hunk header will be wrong when you change the main body of the function. This will affect grep --function-context and diff -W as well as the normal hunk headers. The problem is there is no way to avoid that and provide something useful in the library example you have above. It would be useful to find some code bases and diff the output of 'git log --patch' with and without the leading whitespace match in the function pattern to see how often this is a problem (i.e. when the funcnames do not match see which one is correct).\n> \n> You are right -- on trying out the function on a two other scheme\n> codebases, I noticed that there are a lot more wrongly matched functions\n> than I initially thought. About half of them identify the wrong function\n> in one of the repositories I tried. However, removing the leading\n> whitespace in the pattern did not lead to better matching; it just led\n> to a lot of the hunk headers going blank. I am not sure what causes this\n> behaviour, but my guess is that the function contexts are shown only if\n> it is within a certain distance from the function definition?\n> \n> Even if it did match only the top level defines correctly, the functions\n> matched would still often be technically wrong -- it will show the outer\n> function as the context when the user has edited an internal function\n> (and in Scheme, there is heavy usage of internal functions).\n> \n> After running 'git grep --function-context' with the leading whitespace\n> removed, it seems to match too aggressively, as it captures a huge\n> region to match all the way upto the top level. Especially for files\n> where all the definitions are in a 'library'.\n> \n> Overall, I personally felt that there were more downsides to matching\n> only at the top level. I'd rather the hunk header have the nearest\n> function to provide the context, than have no function displayed at all.\n> Even when the match is wrong, it at least helps me locate where the\n> change was made more easily.\n\nThanks for taking the time to check the differences between the two \napproaches, as there is no perfect solution I'm happy to go with the one \nthat seemed to be best in your investigations\n\nBest Wishes\n\nPhillip\n\n>> Best Wishes\n>>\n>> Phillip\n>>\n>>\n> \n"},{"id":"421199","messageId":"20210408091442.22740-1-raykar.ath@gmail.com","threadId":"55399","inReplyTo":"20210403131612.97194-1-raykar.ath@gmail.com","subject":"[GSoC][PATCH v3 0/1] userdiff: add support for scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-04-08T09:14:42Z","receivedAt":"2021-04-08T09:15:51Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"Changes since v2:\n\n - Add word regex to support for R7RS identifiers that allow for special\n   characters between vertical lines, such as |foo bar baz|.\n   - Change t4034 to include the above case.\n\nAtharva Raykar (1):\n  userdiff: add support for Scheme\n\n Documentation/gitattributes.txt    |  2 ++\n t/t4018-diff-funcname.sh           |  1 +\n t/t4018/scheme-class               |  7 +++++++\n t/t4018/scheme-def                 |  4 ++++\n t/t4018/scheme-def-variant         |  4 ++++\n t/t4018/scheme-define-slash-public |  7 +++++++\n t/t4018/scheme-define-syntax       |  8 ++++++++\n t/t4018/scheme-define-variant      |  4 ++++\n t/t4018/scheme-library             | 11 +++++++++++\n t/t4018/scheme-local-define        |  4 ++++\n t/t4018/scheme-module              |  6 ++++++\n t/t4018/scheme-top-level-define    |  4 ++++\n t/t4018/scheme-user-defined-define |  6 ++++++\n t/t4034-diff-words.sh              |  1 +\n t/t4034/scheme/expect              | 11 +++++++++++\n t/t4034/scheme/post                |  6 ++++++\n t/t4034/scheme/pre                 |  6 ++++++\n userdiff.c                         |  9 +++++++++\n 18 files changed, 101 insertions(+)\n create mode 100644 t/t4018/scheme-class\n create mode 100644 t/t4018/scheme-def\n create mode 100644 t/t4018/scheme-def-variant\n create mode 100644 t/t4018/scheme-define-slash-public\n create mode 100644 t/t4018/scheme-define-syntax\n create mode 100644 t/t4018/scheme-define-variant\n create mode 100644 t/t4018/scheme-library\n create mode 100644 t/t4018/scheme-local-define\n create mode 100644 t/t4018/scheme-module\n create mode 100644 t/t4018/scheme-top-level-define\n create mode 100644 t/t4018/scheme-user-defined-define\n create mode 100644 t/t4034/scheme/expect\n create mode 100644 t/t4034/scheme/post\n create mode 100644 t/t4034/scheme/pre\n\n-- \n2.31.1\n\n"},{"id":"421200","messageId":"20210408091442.22740-2-raykar.ath@gmail.com","threadId":"55399","inReplyTo":"20210403131612.97194-1-raykar.ath@gmail.com","subject":"[GSoC][PATCH v3 1/1] userdiff: add support for Scheme","fromName":"Atharva Raykar","fromEmail":"raykar.ath@gmail.com","sentAt":"2021-04-08T09:14:43Z","receivedAt":"2021-04-08T09:16:09Z","isPatch":true,"sender":{"key":"raykar.ath@gmail.com","avatar":"https://avatars.githubusercontent.com/u/24277692?v=4"},"body":"Add a diff driver for Scheme-like languages which recognizes top level\nand local `define` forms, whether it is a function definition, binding,\nsyntax definition or a user-defined `define-xyzzy` form.\n\nAlso supports R6RS `library` forms, `module` forms along with class and\nstruct declarations used in Racket (PLT Scheme).\n\nAlternate \"def\" syntax such as those in Gerbil Scheme are also\nsupported, like defstruct, defsyntax and so on.\n\nThe rationale for picking `define` forms for the hunk headers is because\nit is usually the only significant form for defining the structure of\nthe program, and it is a common pattern for schemers to have local\nfunction definitions to hide their visibility, so it is not only the top\nlevel `define`'s that are of interest. Schemers also extend the language\nwith macros to provide their own define forms (for example, something\nlike a `define-test-suite`) which is also captured in the hunk header.\n\nSince it is common practice to extend syntax with variants of a form\nlike `module+`, `class*` etc, those have been supported as well.\n\nThe word regex is a best-effort attempt to conform to R7RS[1] valid\nidentifiers, symbols and numbers.\n\n[1] https://small.r7rs.org/attachment/r7rs.pdf (section 2.1)\n\nSigned-off-by: Atharva Raykar <raykar.ath@gmail.com>\n---\n Documentation/gitattributes.txt    |  2 ++\n t/t4018-diff-funcname.sh           |  1 +\n t/t4018/scheme-class               |  7 +++++++\n t/t4018/scheme-def                 |  4 ++++\n t/t4018/scheme-def-variant         |  4 ++++\n t/t4018/scheme-define-slash-public |  7 +++++++\n t/t4018/scheme-define-syntax       |  8 ++++++++\n t/t4018/scheme-define-variant      |  4 ++++\n t/t4018/scheme-library             | 11 +++++++++++\n t/t4018/scheme-local-define        |  4 ++++\n t/t4018/scheme-module              |  6 ++++++\n t/t4018/scheme-top-level-define    |  4 ++++\n t/t4018/scheme-user-defined-define |  6 ++++++\n t/t4034-diff-words.sh              |  1 +\n t/t4034/scheme/expect              | 11 +++++++++++\n t/t4034/scheme/post                |  6 ++++++\n t/t4034/scheme/pre                 |  6 ++++++\n userdiff.c                         |  9 +++++++++\n 18 files changed, 101 insertions(+)\n create mode 100644 t/t4018/scheme-class\n create mode 100644 t/t4018/scheme-def\n create mode 100644 t/t4018/scheme-def-variant\n create mode 100644 t/t4018/scheme-define-slash-public\n create mode 100644 t/t4018/scheme-define-syntax\n create mode 100644 t/t4018/scheme-define-variant\n create mode 100644 t/t4018/scheme-library\n create mode 100644 t/t4018/scheme-local-define\n create mode 100644 t/t4018/scheme-module\n create mode 100644 t/t4018/scheme-top-level-define\n create mode 100644 t/t4018/scheme-user-defined-define\n create mode 100644 t/t4034/scheme/expect\n create mode 100644 t/t4034/scheme/post\n create mode 100644 t/t4034/scheme/pre\n\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex 0a60472bb5..cfcfa800c2 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -845,6 +845,8 @@ patterns are available:\n \n - `rust` suitable for source code in the Rust language.\n \n+- `scheme` suitable for source code in the Scheme language.\n+\n - `tex` suitable for source code for LaTeX documents.\n \n \ndiff --git a/t/t4018-diff-funcname.sh b/t/t4018-diff-funcname.sh\nindex 9675bc17db..823ea96acb 100755\n--- a/t/t4018-diff-funcname.sh\n+++ b/t/t4018-diff-funcname.sh\n@@ -48,6 +48,7 @@ diffpatterns=\"\n \tpython\n \truby\n \trust\n+\tscheme\n \ttex\n \tcustom1\n \tcustom2\ndiff --git a/t/t4018/scheme-class b/t/t4018/scheme-class\nnew file mode 100644\nindex 0000000000..e5e07b43fb\n--- /dev/null\n+++ b/t/t4018/scheme-class\n@@ -0,0 +1,7 @@\n+(define book-class%\n+  (class* () object% RIGHT\n+    (field (pages 5))\n+    (field (ChangeMe 5))\n+    (define/public (letters)\n+      (* pages 500))\n+    (super-new)))\ndiff --git a/t/t4018/scheme-def b/t/t4018/scheme-def\nnew file mode 100644\nindex 0000000000..1e2673da96\n--- /dev/null\n+++ b/t/t4018/scheme-def\n@@ -0,0 +1,4 @@\n+(def (some-func x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\ndiff --git a/t/t4018/scheme-def-variant b/t/t4018/scheme-def-variant\nnew file mode 100644\nindex 0000000000..d857a61d64\n--- /dev/null\n+++ b/t/t4018/scheme-def-variant\n@@ -0,0 +1,4 @@\n+(defmethod {print point} RIGHT\n+  (lambda (self)\n+    (with ((point x y) self)\n+      (printf \"{ChangeMe x:~a y:~a}~n\" x y))))\ndiff --git a/t/t4018/scheme-define-slash-public b/t/t4018/scheme-define-slash-public\nnew file mode 100644\nindex 0000000000..39a93a1600\n--- /dev/null\n+++ b/t/t4018/scheme-define-slash-public\n@@ -0,0 +1,7 @@\n+(define bar-class%\n+  (class object%\n+    (field (info 5))\n+    (define/public (foo) RIGHT\n+      (+ info 42)\n+      (* info ChangeMe))\n+    (super-new)))\ndiff --git a/t/t4018/scheme-define-syntax b/t/t4018/scheme-define-syntax\nnew file mode 100644\nindex 0000000000..7d5e99e0fc\n--- /dev/null\n+++ b/t/t4018/scheme-define-syntax\n@@ -0,0 +1,8 @@\n+(define-syntax define-test-suite RIGHT\n+  (syntax-rules ()\n+    ((_ suite-name (name test) ChangeMe ...)\n+     (define suite-name\n+       (let ((tests\n+              `((name . ,test) ...)))\n+         (lambda ()\n+           (run-suite 'suite-name tests)))))))\ndiff --git a/t/t4018/scheme-define-variant b/t/t4018/scheme-define-variant\nnew file mode 100644\nindex 0000000000..911708854d\n--- /dev/null\n+++ b/t/t4018/scheme-define-variant\n@@ -0,0 +1,4 @@\n+(define* (some-func x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\ndiff --git a/t/t4018/scheme-library b/t/t4018/scheme-library\nnew file mode 100644\nindex 0000000000..82ea3df510\n--- /dev/null\n+++ b/t/t4018/scheme-library\n@@ -0,0 +1,11 @@\n+(library (my-helpers id-stuff) RIGHT\n+  (export find-dup)\n+  (import (ChangeMe))\n+  (define (find-dup l)\n+    (and (pair? l)\n+         (let loop ((rest (cdr l)))\n+           (cond\n+            [(null? rest) (find-dup (cdr l))]\n+            [(bound-identifier=? (car l) (car rest))\n+             (car rest)]\n+            [else (loop (cdr rest))])))))\ndiff --git a/t/t4018/scheme-local-define b/t/t4018/scheme-local-define\nnew file mode 100644\nindex 0000000000..bc6d8aebbe\n--- /dev/null\n+++ b/t/t4018/scheme-local-define\n@@ -0,0 +1,4 @@\n+(define (higher-order)\n+  (define local-function RIGHT\n+    (lambda (x)\n+     (car \"this is\" \"ChangeMe\"))))\ndiff --git a/t/t4018/scheme-module b/t/t4018/scheme-module\nnew file mode 100644\nindex 0000000000..edfae0ebf7\n--- /dev/null\n+++ b/t/t4018/scheme-module\n@@ -0,0 +1,6 @@\n+(module A RIGHT\n+  (export with-display-exception)\n+  (extern (display-exception display-exception ChangeMe))\n+  (def (with-display-exception thunk)\n+    (with-catch (lambda (e) (display-exception e (current-error-port)) e)\n+      thunk)))\ndiff --git a/t/t4018/scheme-top-level-define b/t/t4018/scheme-top-level-define\nnew file mode 100644\nindex 0000000000..624743c22b\n--- /dev/null\n+++ b/t/t4018/scheme-top-level-define\n@@ -0,0 +1,4 @@\n+(define (some-func x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\ndiff --git a/t/t4018/scheme-user-defined-define b/t/t4018/scheme-user-defined-define\nnew file mode 100644\nindex 0000000000..35fe7cc9bf\n--- /dev/null\n+++ b/t/t4018/scheme-user-defined-define\n@@ -0,0 +1,6 @@\n+(define-test-suite record\\ case-tests RIGHT\n+  (record-case-1 (lambda (fail)\n+                   (let ((a (make-foo 1 2)))\n+                     (record-case a\n+                       ((bar x) (ChangeMe))\n+                       ((foo a b) (+ a b)))))))\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 56f1e62a97..ee7721ab91 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -325,6 +325,7 @@ test_language_driver perl\n test_language_driver php\n test_language_driver python\n test_language_driver ruby\n+test_language_driver scheme\n test_language_driver tex\n \n test_expect_success 'word-diff with diff.sbe' '\ndiff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\nnew file mode 100644\nindex 0000000000..496cd5de8c\n--- /dev/null\n+++ b/t/t4034/scheme/expect\n@@ -0,0 +1,11 @@\n+<BOLD>diff --git a/pre b/post<RESET>\n+<BOLD>index 74b6605..63b6ac4 100644<RESET>\n+<BOLD>--- a/pre<RESET>\n+<BOLD>+++ b/post<RESET>\n+<CYAN>@@ -1,6 +1,6 @@<RESET>\n+(define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n+  ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n+  (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n+  (define <RED>some-text<RESET><GREEN>|a greeting|<RESET> \"hello\")\n+  (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n+    (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\ndiff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\nnew file mode 100644\nindex 0000000000..63b6ac4f87\n--- /dev/null\n+++ b/t/t4034/scheme/post\n@@ -0,0 +1,6 @@\n+(define (my-func first second)\n+  ; This is a (moderately) cool function.\n+  (that\\place (+ 3 4))\n+  (define |a greeting| \"hello\")\n+  (let ((c (add1 first)))\n+    (format \"one more than the total is %d\" (+ c second))))\ndiff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\nnew file mode 100644\nindex 0000000000..74b6605357\n--- /dev/null\n+++ b/t/t4034/scheme/pre\n@@ -0,0 +1,6 @@\n+(define (myfunc a b)\n+  ; This is a really cool function.\n+  (this\\place (+ 3 4))\n+  (define some-text \"hello\")\n+  (let ((c (+ a b)))\n+    (format \"one more than the total is %d\" (add1 c))))\ndiff --git a/userdiff.c b/userdiff.c\nindex 3f81a2261c..3897317aff 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -191,6 +191,15 @@ PATTERNS(\"rust\",\n \t \"[a-zA-Z_][a-zA-Z0-9_]*\"\n \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n+PATTERNS(\"scheme\",\n+\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\",\n+\t /*\n+\t  * R7RS valid identifiers include any sequence enclosed\n+\t  * within vertical lines having no backslashes\n+\t  */\n+\t \"\\\\|([^\\\\\\\\]*)\\\\|\"\n+\t /* All other words should be delimited by spaces or parentheses */\n+\t \"|([^][)(}{[ \\t])+\"),\n PATTERNS(\"bibtex\", \"(@[a-zA-Z]{1,}[ \\t]*\\\\{{0,1}[ \\t]*[^ \\t\\\"@',\\\\#}{~%]*).*$\",\n \t \"[={}\\\"]|[^={}\\\" \\t]+\"),\n PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n-- \n2.31.1\n\n"},{"id":"421820","messageId":"xmqqzgy39kbf.fsf@gitster.g","threadId":"55399","inReplyTo":"20210408091442.22740-2-raykar.ath@gmail.com","subject":"Re: [GSoC][PATCH v3 1/1] userdiff: add support for Scheme","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-04-12T23:04:04Z","receivedAt":"2021-04-12T23:04:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Atharva Raykar <raykar.ath@gmail.com> writes:\n\n> Add a diff driver for Scheme-like languages which recognizes top level\n> and local `define` forms, whether it is a function definition, binding,\n> syntax definition or a user-defined `define-xyzzy` form.\n>\n> Also supports R6RS `library` forms, `module` forms along with class and\n> struct declarations used in Racket (PLT Scheme).\n>\n> Alternate \"def\" syntax such as those in Gerbil Scheme are also\n> supported, like defstruct, defsyntax and so on.\n>\n> The rationale for picking `define` forms for the hunk headers is because\n> it is usually the only significant form for defining the structure of\n> the program, and it is a common pattern for schemers to have local\n> function definitions to hide their visibility, so it is not only the top\n> level `define`'s that are of interest. Schemers also extend the language\n> with macros to provide their own define forms (for example, something\n> like a `define-test-suite`) which is also captured in the hunk header.\n>\n> Since it is common practice to extend syntax with variants of a form\n> like `module+`, `class*` etc, those have been supported as well.\n>\n> The word regex is a best-effort attempt to conform to R7RS[1] valid\n> identifiers, symbols and numbers.\n>\n> [1] https://small.r7rs.org/attachment/r7rs.pdf (section 2.1)\n>\n> Signed-off-by: Atharva Raykar <raykar.ath@gmail.com>\n> ---\n>  Documentation/gitattributes.txt    |  2 ++\n>  t/t4018-diff-funcname.sh           |  1 +\n>  t/t4018/scheme-class               |  7 +++++++\n>  t/t4018/scheme-def                 |  4 ++++\n>  t/t4018/scheme-def-variant         |  4 ++++\n>  t/t4018/scheme-define-slash-public |  7 +++++++\n>  t/t4018/scheme-define-syntax       |  8 ++++++++\n>  t/t4018/scheme-define-variant      |  4 ++++\n>  t/t4018/scheme-library             | 11 +++++++++++\n>  t/t4018/scheme-local-define        |  4 ++++\n>  t/t4018/scheme-module              |  6 ++++++\n>  t/t4018/scheme-top-level-define    |  4 ++++\n>  t/t4018/scheme-user-defined-define |  6 ++++++\n>  t/t4034-diff-words.sh              |  1 +\n>  t/t4034/scheme/expect              | 11 +++++++++++\n>  t/t4034/scheme/post                |  6 ++++++\n>  t/t4034/scheme/pre                 |  6 ++++++\n>  userdiff.c                         |  9 +++++++++\n>  18 files changed, 101 insertions(+)\n>  create mode 100644 t/t4018/scheme-class\n>  create mode 100644 t/t4018/scheme-def\n>  create mode 100644 t/t4018/scheme-def-variant\n>  create mode 100644 t/t4018/scheme-define-slash-public\n>  create mode 100644 t/t4018/scheme-define-syntax\n>  create mode 100644 t/t4018/scheme-define-variant\n>  create mode 100644 t/t4018/scheme-library\n>  create mode 100644 t/t4018/scheme-local-define\n>  create mode 100644 t/t4018/scheme-module\n>  create mode 100644 t/t4018/scheme-top-level-define\n>  create mode 100644 t/t4018/scheme-user-defined-define\n>  create mode 100644 t/t4034/scheme/expect\n>  create mode 100644 t/t4034/scheme/post\n>  create mode 100644 t/t4034/scheme/pre\n\nWe have seen reviews on previous rounds, and haven't heard anything\non this round yet.\n\nIf I do not hear from anybody in a few days, let's merge it to\n'next'.\n"}]}