{"thread":{"id":"64483","subject":"[PATCH] diff: \"lisp\" userdiff_driver","startedAt":"2025-11-15T10:17:47Z","lastAt":"2026-04-15T06:54:32Z","messageCount":27,"participants":["Scott L. Burson via GitGitGadget","Johannes Sixt","Scott L. Burson","Junio C Hamano","D. Ben Knoble","Johannes Sixt via GitGitGadget"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"530732","messageId":"pull.2000.git.1763201865025.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":null,"subject":"[PATCH] diff: \"lisp\" userdiff_driver","fromName":"Scott L. Burson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-15T10:17:45Z","receivedAt":"2025-11-15T10:17:47Z","isPatch":true,"body":"From: \"Scott L. Burson\" <Scott@sympoiesis.com>\n\nThe \"scheme\" driver doesn't quite work for Common Lisp.  This driver\nis very generic and should work for almost any dialect of Lisp,\nincluding Common Lisp.\n\nSigned-off-by: Scott L. Burson <Scott@sympoiesis.com>\n---\n    diff: \"lisp\" userdiff_driver\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2000%2Fslburson%2Flisp-userdiff_driver-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2000/slburson/lisp-userdiff_driver-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/2000\n\n userdiff.c | 8 ++++++++\n 1 file changed, 8 insertions(+)\n\ndiff --git a/userdiff.c b/userdiff.c\nindex fe710a68bf..e127b4a1f1 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -249,6 +249,14 @@ PATTERNS(\"kotlin\",\n \t \"|[.][0-9][0-9_]*([Ee][-+]?[0-9]+)?[fFlLuU]?\"\n \t /* unary and binary operators */\n \t \"|[-+*/<>%&^|=!]==?|--|\\\\+\\\\+|<<=|>>=|&&|\\\\|\\\\||->|\\\\.\\\\*|!!|[?:.][.:]\"),\n+PATTERNS(\"lisp\",\n+\t /* Either an unindented left paren, or a slightly indented line\n+\t  * starting with \"(def\" */\n+\t \"^((\\\\(|:space:{1,2}\\\\(def).*)$\",\n+\t /* Common Lisp symbol syntax allows arbitrary strings between vertical bars */\n+\t \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n+\t /* All other words are delimited by spaces or parentheses/brackets/braces */\n+\t \"|([^][(){} \\t])+\"),\n PATTERNS(\"markdown\",\n \t \"^ {0,3}#{1,6}[ \\t].*\",\n \t /* -- */\n\nbase-commit: fd372d9b1a69a01a676398882bbe3840bf51fe72\n-- \ngitgitgadget\n"},{"id":"530750","messageId":"773d3233-c890-4df9-8f7e-32ff8a48651e@kdbg.org","threadId":"64483","inReplyTo":"pull.2000.git.1763201865025.gitgitgadget@gmail.com","subject":"Re: [PATCH] diff: \"lisp\" userdiff_driver","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2025-11-15T17:06:44Z","receivedAt":"2025-11-15T17:40:59Z","isPatch":true,"body":"[Cc the author of the Scheme driver]\n\nAm 15.11.25 um 11:17 schrieb Scott L. Burson via GitGitGadget:\n> From: \"Scott L. Burson\" <Scott@sympoiesis.com>\n> \n> The \"scheme\" driver doesn't quite work for Common Lisp.  This driver\n> is very generic and should work for almost any dialect of Lisp,\n> including Common Lisp.\n\n\"Doesn't quite work\" is unfortunately a description of the problem that\ndoes not help a reviewer decide whether this change is justified. Please\nadd a lot more details why the Scheme driver is unsuitable for Lisp and\nwhy a new driver is needed.\n\nIt is customary to mark changes to the drivers in the subject line with\n\"userdiff:\". Have a look at `git log userdiff.c`. It would be\nappreciated to stay away from nerdy tokens like \"userdiff_driver\" when\nthe change can be summarized in plain English language.\n\n> \n> Signed-off-by: Scott L. Burson <Scott@sympoiesis.com>\n> ---\n>     diff: \"lisp\" userdiff_driver\n> \n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-2000%2Fslburson%2Flisp-userdiff_driver-v1\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2000/slburson/lisp-userdiff_driver-v1\n> Pull-Request: https://github.com/gitgitgadget/git/pull/2000\n> \n>  userdiff.c | 8 ++++++++\n>  1 file changed, 8 insertions(+)\n> \n> diff --git a/userdiff.c b/userdiff.c\n> index fe710a68bf..e127b4a1f1 100644\n> --- a/userdiff.c\n> +++ b/userdiff.c\n> @@ -249,6 +249,14 @@ PATTERNS(\"kotlin\",\n>  \t \"|[.][0-9][0-9_]*([Ee][-+]?[0-9]+)?[fFlLuU]?\"\n>  \t /* unary and binary operators */\n>  \t \"|[-+*/<>%&^|=!]==?|--|\\\\+\\\\+|<<=|>>=|&&|\\\\|\\\\||->|\\\\.\\\\*|!!|[?:.][.:]\"),\n> +PATTERNS(\"lisp\",\n> +\t /* Either an unindented left paren, or a slightly indented line\n> +\t  * starting with \"(def\" */\n> +\t \"^((\\\\(|:space:{1,2}\\\\(def).*)$\",\n\nCompared to the Scheme driver, this regular expression is\n\n- more restrictive because it does not permit arbitrary indentation;\n\n- less restrictive because it permits everything that begins with \"(def\".\n\nWhat would happen if this regular expression were added to the Scheme\ndriver? Would it pick up additional and unwanted hunk headers is typical\nScheme code?\n\nJust so you know: The string literal for hunk headers can contain \"\\n\"\nto separate regular expressions. For a given line, they are attempted to\nmatch in order until the first match is found.\n\n> +\t /* Common Lisp symbol syntax allows arbitrary strings between vertical bars */\n> +\t \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n\nThe Scheme driver has an similar description of this word token, but it\nhas only half as many backslashes. Is the difference necessary? Isn't\nactually one or the other incorrect? (I did not try to understand what\nthis version here does.)\n\n> +\t /* All other words are delimited by spaces or parentheses/brackets/braces */\n> +\t \"|([^][(){} \\t])+\"),\n\nThis one looks identical to the same line in Scheme with a slightly\naltered comment.\n\nIn summary, I feel like this could be unified with the Scheme driver,\nexcept when awkward false positive hunk headers in existing Scheme code\nare to be expected. That's a question that I cannot answer.\n\n-- Hannes\n\n"},{"id":"530765","messageId":"CAF5LJ4D4q2S2VFhvEgVOe1Ar0e6cu=H3e_o_98VwHN7wYHh+DQ@mail.gmail.com","threadId":"64483","inReplyTo":"773d3233-c890-4df9-8f7e-32ff8a48651e@kdbg.org","subject":"Re: [PATCH] diff: \"lisp\" userdiff_driver","fromName":"Scott L. Burson","fromEmail":"scott@sympoiesis.com","sentAt":"2025-11-15T23:32:54Z","receivedAt":"2025-11-15T23:33:33Z","isPatch":true,"body":"On Sat, Nov 15, 2025 at 9:06 AM Johannes Sixt <j6t@kdbg.org> wrote:\n>\n> [Cc the author of the Scheme driver]\n>\n> Am 15.11.25 um 11:17 schrieb Scott L. Burson via GitGitGadget:\n> > From: \"Scott L. Burson\" <Scott@sympoiesis.com>\n>\n> Please\n> add a lot more details why the Scheme driver is unsuitable for Lisp and\n> why a new driver is needed.\n\nHere is text I propose for the commit message:\n\n----\nCommon Lisp has top-level forms 'defun' and 'deftype' that are not\nmatched by the current Scheme pattern.  Also, it is more common when\ndefining user macros intended as top-level forms to prefix their names\nwith \"def\" instead of \"define\"; such forms are also not matched.  And\nsome such forms don't even begin with \"def\".\n\nOn the other hand, it is an established formatting convention in the\nLisp community that only top-level forms start at the left margin.  So\nmatching any unindented line starting with an open parenthesis is an\nacceptable heuristic; false positives will be rare.\n\nHowever, there are also cases where notionally top-level forms are\ngrouped together within some containing form.  At least in the Common\nLisp community, it is conventional to indent these by two spaces, or\nsometimes one.  But matching just an open parenthesis indented by two\nspaces would be too broad; so the pattern added by this commit\nrequires an indented form to start with \"(def\".  It is believed that\nthis strikes a good balance between potential false positives and\nfalse negatives.\n----\n\nI discussed the pattern with some other experienced Common Lisp\ndevelopers on a mailing list, and this is what I settled on after\nincorporating their feedback.\n\n> It is customary to mark changes to the drivers in the subject line with\n> \"userdiff:\". Have a look at `git log userdiff.c`. It would be\n> appreciated to stay away from nerdy tokens like \"userdiff_driver\" when\n> the change can be summarized in plain English language.\n\nWill do.\n\n> >\n> > Signed-off-by: Scott L. Burson <Scott@sympoiesis.com>\n> > ---\n> >     diff: \"lisp\" userdiff_driver\n> >\n> > Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-2000%2Fslburson%2Flisp-userdiff_driver-v1\n> > Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2000/slburson/lisp-userdiff_driver-v1\n> > Pull-Request: https://github.com/gitgitgadget/git/pull/2000\n> >\n> >  userdiff.c | 8 ++++++++\n> >  1 file changed, 8 insertions(+)\n> >\n> > diff --git a/userdiff.c b/userdiff.c\n> > index fe710a68bf..e127b4a1f1 100644\n> > --- a/userdiff.c\n> > +++ b/userdiff.c\n> > @@ -249,6 +249,14 @@ PATTERNS(\"kotlin\",\n> >        \"|[.][0-9][0-9_]*([Ee][-+]?[0-9]+)?[fFlLuU]?\"\n> >        /* unary and binary operators */\n> >        \"|[-+*/<>%&^|=!]==?|--|\\\\+\\\\+|<<=|>>=|&&|\\\\|\\\\||->|\\\\.\\\\*|!!|[?:.][.:]\"),\n> > +PATTERNS(\"lisp\",\n> > +      /* Either an unindented left paren, or a slightly indented line\n> > +       * starting with \"(def\" */\n> > +      \"^((\\\\(|:space:{1,2}\\\\(def).*)$\",\n>\n> Compared to the Scheme driver, this regular expression is\n>\n> - more restrictive because it does not permit arbitrary indentation;\n>\n> - less restrictive because it permits everything that begins with \"(def\".\n>\n> What would happen if this regular expression were added to the Scheme\n> driver? Would it pick up additional and unwanted hunk headers is typical\n> Scheme code?\n\nThat is a good question.  I don't think so, but I don't work in Scheme.\nI see that you have CC'ed Atharva Raykar; let's see whether he would\nhave any objection.\n\nI would point out that Scheme is a dialect of Lisp, not the other way\naround.  (Lisp is unusual in being a family of languages, rather than a\nsingle language.)  And having a separate \"lisp\" driver might aid\ndiscoverability.\n\nBut I understand: Scheme got their driver in first, and you have to fight\nagainst the tendency of the driver list to grow unboundedly.\n\nOoh, that reminds me: if we do decide to add a \"lisp\" driver, I'll also need\nto add it to 'Documentation/gitattributes.adoc'.\n\n> The string literal for hunk headers can contain \"\\n\"\n\nNoted.\n\n> > +      /* Common Lisp symbol syntax allows arbitrary strings between vertical bars */\n> > +      \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n>\n> The Scheme driver has an similar description of this word token, but it\n> has only half as many backslashes. Is the difference necessary? Isn't\n> actually one or the other incorrect? (I did not try to understand what\n> this version here does.)\n\nIt's not important, but technically, Common Lisp allows an escaped\nbackslash between vertical bars, but the R7RS formal grammar does not.\nHowever, I just tried Chicken Scheme, which claims to be at least\npartially R7RS compliant, and it does accept the escaped backslash.  I\nam left to conclude that Scheme implementors think that the omission\nof the escaped backslash from the R7RS formal grammar is an oversight\n(I think so too).\n\nOf course, no one would actually write a symbol name with an escaped\nbackslash in it unless they were submitting to an obfuscated Lisp\ncontest.  So we are really being pedantic here.  Still, may as well allow\nit.\n\nAtharva, any comments?\n\n-- Scott\n"},{"id":"530767","messageId":"xmqqbjl2ee8t.fsf@gitster.g","threadId":"64483","inReplyTo":"773d3233-c890-4df9-8f7e-32ff8a48651e@kdbg.org","subject":"Re: [PATCH] diff: \"lisp\" userdiff_driver","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-16T05:30:10Z","receivedAt":"2025-11-16T05:30:13Z","isPatch":true,"body":"Johannes Sixt <j6t@kdbg.org> writes:\n\n>> +\t /* Either an unindented left paren, or a slightly indented line\n>> +\t  * starting with \"(def\" */\n>> +\t \"^((\\\\(|:space:{1,2}\\\\(def).*)$\",\n>\n> Compared to the Scheme driver, this regular expression is\n>\n> - more restrictive because it does not permit arbitrary indentation;\n>\n> - less restrictive because it permits everything that begins with \"(def\".\n>\n> What would happen if this regular expression were added to the Scheme\n> driver? Would it pick up additional and unwanted hunk headers is typical\n> Scheme code?\n\nAs we generally assume that the file being edited is syntactically\nsound, even if one lisp variant understands \"(deffoo\" and others do\nnot, it should be generally fine for the pattern to say something\nlike \"at the beginning of the line, optionally following a few\nspaces, four-letter sequence '(def' is likely to be the beginning of\na function definition\", as long as there is some convention that\nuser defined functions and macros, unless they are to behave\nsimilarly to \"(defun\", would not be named so confusingly to start\nwith d-e-f.\n\nIt would be nice if a single set of rules can cover what existing\nscheme patterns cover, Emacs lisp, and Common lisp.\n\n"},{"id":"530863","messageId":"CAF5LJ4CMtEaJgDYRHXvCTUm9Pjpv2GAsMQN9D-DL-Ric3ADMXQ@mail.gmail.com","threadId":"64483","inReplyTo":"xmqqbjl2ee8t.fsf@gitster.g","subject":"Re: [PATCH] diff: \"lisp\" userdiff_driver","fromName":"Scott L. Burson","fromEmail":"scott@sympoiesis.com","sentAt":"2025-11-17T23:23:19Z","receivedAt":"2025-11-17T23:23:57Z","isPatch":true,"body":"On Sat, Nov 15, 2025 at 9:30 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Johannes Sixt <j6t@kdbg.org> writes:\n>\n> >> +     /* Either an unindented left paren, or a slightly indented line\n> >> +      * starting with \"(def\" */\n> >> +     \"^((\\\\(|:space:{1,2}\\\\(def).*)$\",\n> >\n> > Compared to the Scheme driver, this regular expression is\n> >\n> > - more restrictive because it does not permit arbitrary indentation;\n> >\n> > - less restrictive because it permits everything that begins with \"(def\".\n> >\n> > What would happen if this regular expression were added to the Scheme\n> > driver? Would it pick up additional and unwanted hunk headers is typical\n> > Scheme code?\n\nHmm, we haven't heard from Atharva.  I'll try asking around in the\nScheme community.\n\nThe regex I proposed has a bug.  The use of the Posix character class\nis incorrect, because that class includes tabs.  I will replace it\nwith a literal space.  Also, many Lisps, including Common Lisp in its\ndefault configuration, are case-insensitive, and at least in the\n1970s, it wasn't completely unheard-of to write Lisp code in\nuppercase; I'll change the entry to use 'IPATTERN'.\n\n> As we generally assume that the file being edited is syntactically\n> sound, even if one lisp variant understands \"(deffoo\" and others do\n> not, it should be generally fine for the pattern to say something\n> like \"at the beginning of the line, optionally following a few\n> spaces, four-letter sequence '(def' is likely to be the beginning of\n> a function definition\", as long as there is some convention that\n> user defined functions and macros, unless they are to behave\n> similarly to \"(defun\", would not be named so confusingly to start\n> with d-e-f.\n\nAgreed, but this is not the most important point.  The greater\npotential for false positives comes from the rule (in my proposal)\nthat a left parenthesis in column 0 is taken as indicating a top-level\ndefinition, without even looking at the following characters.\nAlthough Lisp dialects certainly vary, I have not seen one in which\nstandard indentation practice does not indent internal expressions;\ncertainly, Lisp mode in Emacs indents them.  And, I think the rule\nreally does need to be that broad, because top-level forms don't\nalways begin with \"def\"; indeed, one can put any executable expression\nat top level in a source file to perform load-time initializations.\n\nIt's only when there is some indentation that I think the regex needs\nto require a word beginning with \"def\".\n\n> It would be nice if a single set of rules can cover what existing\n> scheme patterns cover, Emacs lisp, and Common lisp.\n\nAgreed.  I do think it would be a little better for non-Scheme users\nif the single driver were named \"lisp\" instead of \"scheme\".  Renaming\nthe driver out from under the Scheme community, though, seems like it\nwould be unfriendly, even after a deprecation period.\n\nOne solution would be to add an aliasing mechanism to the\ndriver table.  Perhaps there would be other use cases for it.  If you\nwould consider a patch along these lines, I can code it up.\n"},{"id":"530871","messageId":"xmqqldk4ez05.fsf@gitster.g","threadId":"64483","inReplyTo":"CAF5LJ4CMtEaJgDYRHXvCTUm9Pjpv2GAsMQN9D-DL-Ric3ADMXQ@mail.gmail.com","subject":"Re: [PATCH] diff: \"lisp\" userdiff_driver","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-11-18T04:38:34Z","receivedAt":"2025-11-18T04:38:37Z","isPatch":true,"body":"\"Scott L. Burson\" <Scott@sympoiesis.com> writes:\n\n> ...  The greater\n> potential for false positives comes from the rule (in my proposal)\n> that a left parenthesis in column 0 is taken as indicating a top-level\n> definition, without even looking at the following characters.\n\nI didn't respond to that part as I didn't know if you were serious\nor joking ;-).\n\n> Although Lisp dialects certainly vary, I have not seen one in which\n> standard indentation practice does not indent internal expressions;\n> certainly, Lisp mode in Emacs indents them.  And, I think the rule\n> really does need to be that broad, because top-level forms don't\n> always begin with \"def\"; indeed, one can put any executable expression\n> at top level in a source file to perform load-time initializations.\n\nExactly, but the more important question is are they considered as\nthe beginning of an important, and sematically distinct, block, just\nlike the beginning of a function is.  I am somewhat negative to the\n\"anything not indented is a beginning of a significant group\", as I\ndo not know how well it meshes with the \"(defXX is a beginning of a\nfunction\", when they are used together.\n\n> One solution would be to add an aliasing mechanism to the\n> driver table.  Perhaps there would be other use cases for it.  If you\n> would consider a patch along these lines, I can code it up.\n\nIt is not a particularly interesting part of the problem, simply\nbecause as the first approximation, we can just advertise \"you can\nmark your lisp files as 'scheme'\".  A more interesting issue is if\nwe can indeed come up with such a superset of patterns that can\ncover all Lisp variants that matter.\n"},{"id":"531067","messageId":"CALnO6CBFKjewrkPeEUh7Q-A2dZ7Fknjy4DszG8xCKu-NvGETfQ@mail.gmail.com","threadId":"64483","inReplyTo":"CAF5LJ4D4q2S2VFhvEgVOe1Ar0e6cu=H3e_o_98VwHN7wYHh+DQ@mail.gmail.com","subject":"Re: [PATCH] diff: \"lisp\" userdiff_driver","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-11-20T16:47:02Z","receivedAt":"2025-11-20T16:47:15Z","isPatch":true,"body":"On Sat, Nov 15, 2025 at 6:33 PM Scott L. Burson <Scott@sympoiesis.com> wrote:\n> > >\n> > > Signed-off-by: Scott L. Burson <Scott@sympoiesis.com>\n> > > ---\n> > >     diff: \"lisp\" userdiff_driver\n> > >\n> > > Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-2000%2Fslburson%2Flisp-userdiff_driver-v1\n> > > Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2000/slburson/lisp-userdiff_driver-v1\n> > > Pull-Request: https://github.com/gitgitgadget/git/pull/2000\n> > >\n> > >  userdiff.c | 8 ++++++++\n> > >  1 file changed, 8 insertions(+)\n> > >\n> > > diff --git a/userdiff.c b/userdiff.c\n> > > index fe710a68bf..e127b4a1f1 100644\n> > > --- a/userdiff.c\n> > > +++ b/userdiff.c\n> > > @@ -249,6 +249,14 @@ PATTERNS(\"kotlin\",\n> > >        \"|[.][0-9][0-9_]*([Ee][-+]?[0-9]+)?[fFlLuU]?\"\n> > >        /* unary and binary operators */\n> > >        \"|[-+*/<>%&^|=!]==?|--|\\\\+\\\\+|<<=|>>=|&&|\\\\|\\\\||->|\\\\.\\\\*|!!|[?:.][.:]\"),\n> > > +PATTERNS(\"lisp\",\n> > > +      /* Either an unindented left paren, or a slightly indented line\n> > > +       * starting with \"(def\" */\n> > > +      \"^((\\\\(|:space:{1,2}\\\\(def).*)$\",\n> >\n> > Compared to the Scheme driver, this regular expression is\n> >\n> > - more restrictive because it does not permit arbitrary indentation;\n> >\n> > - less restrictive because it permits everything that begins with \"(def\".\n> >\n> > What would happen if this regular expression were added to the Scheme\n> > driver? Would it pick up additional and unwanted hunk headers is typical\n> > Scheme code?\n>\n> That is a good question.  I don't think so, but I don't work in Scheme.\n> I see that you have CC'ed Atharva Raykar; let's see whether he would\n> have any objection.\n>\n> I would point out that Scheme is a dialect of Lisp, not the other way\n> around.  (Lisp is unusual in being a family of languages, rather than a\n> single language.)  And having a separate \"lisp\" driver might aid\n> discoverability.\n\nWithout \"going there,\" I think there are enough differences to warrant\na different driver. (OTOH, I have sometimes wanted to teach the Scheme\ndriver that most \"def\" things are probably definitions.) Our\nindentation is less rigid in that indented forms may be more deeply\nnested than only one or two spaces (and we of course have more\ndefinition forms than \"only things starting with def\"), and I don't\nunderstand the downthread desire to not permit tabs.\n\nAs for Scheme community, I'll suggest asking on the Racket channels\n(Discourse is probably best if you want a mailing-list-like\ndiscussion?)\n\n\n-- \nD. Ben Knoble\n"},{"id":"531363","messageId":"CAF5LJ4AgJvMHej1uRm5Q0v_BTavmm+aXPBC-nGmBp9URX16Gkw@mail.gmail.com","threadId":"64483","inReplyTo":"CALnO6CBFKjewrkPeEUh7Q-A2dZ7Fknjy4DszG8xCKu-NvGETfQ@mail.gmail.com","subject":"Re: [PATCH] diff: \"lisp\" userdiff_driver","fromName":"Scott L. Burson","fromEmail":"scott@sympoiesis.com","sentAt":"2025-11-27T02:10:00Z","receivedAt":"2025-11-27T02:10:38Z","isPatch":true,"body":"On Thu, Nov 20, 2025 at 8:47 AM D. Ben Knoble <ben.knoble@gmail.com> wrote:\n> Without \"going there,\" I think there are enough differences to warrant\n> a different driver. (OTOH, I have sometimes wanted to teach the Scheme\n> driver that most \"def\" things are probably definitions.) Our\n> indentation is less rigid in that indented forms may be more deeply\n> nested than only one or two spaces (and we of course have more\n> definition forms than \"only things starting with def\"), and I don't\n> understand the downthread desire to not permit tabs.\n\nThe proposal on the table is not to replace the current Scheme regexp\nwith the Lisp one, but rather to disjoin them, so as to match any line\nthat either one would match.  So you don't need to worry about false\nnegatives.\n\nThe thing about tabs was that, the way I had initially written the\nregexp, it would have matched a line starting with one or two tabs, or\na tab and a space -- but no other whitespace -- followed by \"(def\".\nThis was not my intention.  I was just trying to match lines starting\nwith a space or two and \"(def\".  Again, this is in addition to the\nlines currently being matched by the Scheme regexp.\n\n> As for Scheme community, I'll suggest asking on the Racket channels\n> (Discourse is probably best if you want a mailing-list-like\n> discussion?)\n\nI have asked on Reddit under /r/scheme, and got only the same\nobjection that you offered (\"define\" etc. forms more deeply nested),\nwith the same resolution (disjoining the new pattern to the existing\none).\n\nReddit says my post got some 4300 views.  That seems like an adequate\nsample to me, but I can try elsewhere if you think it's important.\n\nI have an updated version of the patch ready; I will submit it\nshortly.\n"},{"id":"531365","messageId":"pull.2000.v2.git.1764211096.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":"pull.2000.git.1763201865025.gitgitgadget@gmail.com","subject":"[PATCH v2 0/2] userdiff: extend Scheme support to cover other Lisp dialects","fromName":"Scott L. Burson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-27T02:38:14Z","receivedAt":"2025-11-27T02:38:18Z","isPatch":true,"body":"Common Lisp, Emacs Lisp, and other dialects have some top-level forms, most\nimportantly 'defun', that are not matched by the current Scheme pattern.\nAlso, it is common in these dialects, when defining user macros intended as\ntop-level forms, to prefix their names with \"def\" instead of \"define\"; such\nforms are also not currently matched. Some such forms don't even begin with\n\"def\".\n\nOn the other hand, it is an established formatting convention in the Lisp\ncommunity that only top-level forms start at the left margin. So matching\nany unindented line starting with an open parenthesis is an acceptable\nheuristic; false positives will be rare.\n\nHowever, there are also cases where notionally top-level forms are grouped\ntogether within some containing form. At least in the Common Lisp community,\nit is conventional to indent these by two spaces, or sometimes one. But\nmatching just an open parenthesis indented by two spaces would be too broad;\nso the pattern added by this commit requires an indented form to start with\n\"(def\". It is believed that this strikes a good balance between potential\nfalse positives and false negatives.\n\nThis commit disjoins a regexp employing these heuristics to the existing\nScheme regexp, so it will still match everything that it did previously.\n\nChanges since v1:\n\n * unified with Scheme driver\n * fixed whitespace bug (tabs were allowed incorrectly)\n * made \"(def\" case-insensitive\n * improved commit summary line\n * improved commit description\n\nScott L. Burson (2):\n  diff: \"lisp\" userdiff_driver\n  merge with Scheme regexp; fix bugs\n\n userdiff.c | 17 ++++++++++++-----\n 1 file changed, 12 insertions(+), 5 deletions(-)\n\n\nbase-commit: fd372d9b1a69a01a676398882bbe3840bf51fe72\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2000%2Fslburson%2Flisp-userdiff_driver-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2000/slburson/lisp-userdiff_driver-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/2000\n\nRange-diff vs v1:\n\n 1:  da99bb0bcd = 1:  da99bb0bcd diff: \"lisp\" userdiff_driver\n -:  ---------- > 2:  86315aa3e3 merge with Scheme regexp; fix bugs\n\n-- \ngitgitgadget\n"},{"id":"531366","messageId":"da99bb0bcd8c92e0d6de8b929b67095fae251f88.1764211096.git.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":"pull.2000.v2.git.1764211096.gitgitgadget@gmail.com","subject":"[PATCH v2 1/2] diff: \"lisp\" userdiff_driver","fromName":"Scott L. Burson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-27T02:38:15Z","receivedAt":"2025-11-27T02:38:19Z","isPatch":true,"body":"From: \"Scott L. Burson\" <Scott@sympoiesis.com>\n\nThe \"scheme\" driver doesn't quite work for Common Lisp.  This driver\nis very generic and should work for almost any dialect of Lisp,\nincluding Common Lisp.\n\nSigned-off-by: Scott L. Burson <Scott@sympoiesis.com>\n---\n userdiff.c | 8 ++++++++\n 1 file changed, 8 insertions(+)\n\ndiff --git a/userdiff.c b/userdiff.c\nindex fe710a68bf..e127b4a1f1 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -249,6 +249,14 @@ PATTERNS(\"kotlin\",\n \t \"|[.][0-9][0-9_]*([Ee][-+]?[0-9]+)?[fFlLuU]?\"\n \t /* unary and binary operators */\n \t \"|[-+*/<>%&^|=!]==?|--|\\\\+\\\\+|<<=|>>=|&&|\\\\|\\\\||->|\\\\.\\\\*|!!|[?:.][.:]\"),\n+PATTERNS(\"lisp\",\n+\t /* Either an unindented left paren, or a slightly indented line\n+\t  * starting with \"(def\" */\n+\t \"^((\\\\(|:space:{1,2}\\\\(def).*)$\",\n+\t /* Common Lisp symbol syntax allows arbitrary strings between vertical bars */\n+\t \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n+\t /* All other words are delimited by spaces or parentheses/brackets/braces */\n+\t \"|([^][(){} \\t])+\"),\n PATTERNS(\"markdown\",\n \t \"^ {0,3}#{1,6}[ \\t].*\",\n \t /* -- */\n-- \ngitgitgadget\n\n"},{"id":"531367","messageId":"86315aa3e36afa1ee741a2c9b9e95a71ca569302.1764211096.git.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":"pull.2000.v2.git.1764211096.gitgitgadget@gmail.com","subject":"[PATCH v2 2/2] merge with Scheme regexp; fix bugs","fromName":"Scott L. Burson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2025-11-27T02:38:16Z","receivedAt":"2025-11-27T02:38:21Z","isPatch":true,"body":"From: \"Scott L. Burson\" <Scott@sympoiesis.com>\n\nThis commit merges (by disjoining) the new generic Lisp regexp into\nthe existing Scheme regexp.  It also fixes two bugs: the new regexp\nwas unintentionally allowing tabs, and the matching of \"(def\" should\nbe case-insensitive.\n\nSigned-off-by: Scott L. Burson <Scott@sympoiesis.com>\n---\n userdiff.c | 25 ++++++++++++-------------\n 1 file changed, 12 insertions(+), 13 deletions(-)\n\ndiff --git a/userdiff.c b/userdiff.c\nindex e127b4a1f1..b67dfddbef 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -249,14 +249,6 @@ PATTERNS(\"kotlin\",\n \t \"|[.][0-9][0-9_]*([Ee][-+]?[0-9]+)?[fFlLuU]?\"\n \t /* unary and binary operators */\n \t \"|[-+*/<>%&^|=!]==?|--|\\\\+\\\\+|<<=|>>=|&&|\\\\|\\\\||->|\\\\.\\\\*|!!|[?:.][.:]\"),\n-PATTERNS(\"lisp\",\n-\t /* Either an unindented left paren, or a slightly indented line\n-\t  * starting with \"(def\" */\n-\t \"^((\\\\(|:space:{1,2}\\\\(def).*)$\",\n-\t /* Common Lisp symbol syntax allows arbitrary strings between vertical bars */\n-\t \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n-\t /* All other words are delimited by spaces or parentheses/brackets/braces */\n-\t \"|([^][(){} \\t])+\"),\n PATTERNS(\"markdown\",\n \t \"^ {0,3}#{1,6}[ \\t].*\",\n \t /* -- */\n@@ -352,14 +344,21 @@ PATTERNS(\"rust\",\n \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n PATTERNS(\"scheme\",\n-\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\",\n+\t /* A possibly indented left paren followed by a Scheme keyword. */\n+\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\\n\"\n+\t /*\n+\t  * For other Lisp dialects: either an unindented left paren, or a\n+\t  * slightly indented line starting with \"(def\".\n+\t  */\n+\t \"^((\\\\(| {1,2}\\\\([Dd][Ee][Ff]).*)$\",\n \t /*\n-\t  * R7RS valid identifiers include any sequence enclosed\n-\t  * within vertical lines having no backslashes\n+\t  * The union of R7RS and Common Lisp symbol syntax: allows arbitrary\n+\t  * strings between vertical bars, including escaped backslashes and\n+\t  * vertical bars.\n \t  */\n-\t \"\\\\|([^\\\\\\\\]*)\\\\|\"\n+\t \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n \t /* All other words should be delimited by spaces or parentheses */\n-\t \"|([^][)(}{[ \\t])+\"),\n+\t \"|([^][)(}{ \\t])+\"),\n PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n \t \"\\\\\\\\[a-zA-Z@]+|\\\\\\\\.|([a-zA-Z0-9]|[^\\x01-\\x7f])+\"),\n { .name = \"default\", .binary = -1 },\n-- \ngitgitgadget\n"},{"id":"531370","messageId":"CAF5LJ4B2PeLPZi5gD6Htqdwhj5T-5U9Od_NhDe-8kXTN1-v6_Q@mail.gmail.com","threadId":"64483","inReplyTo":"da99bb0bcd8c92e0d6de8b929b67095fae251f88.1764211096.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 1/2] diff: \"lisp\" userdiff_driver","fromName":"Scott L. Burson","fromEmail":"scott@sympoiesis.com","sentAt":"2025-11-27T10:32:07Z","receivedAt":"2025-11-27T10:32:48Z","isPatch":true,"body":"On Wed, Nov 26, 2025, 6:38 PM Scott L. Burson via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n[duplicate of first message]\n\nHmm.  Somehow I have screwed things up so that GitGitGadget sends out\nonly the first commit, not including the changes in the second, nor\nusing the PR title and the text in the description.  Sorry for the\nnoise.  Is there an easy fix, or do I need to make a new PR?\n\n-- Scott\n"},{"id":"531371","messageId":"8800e796-77f4-4613-8c68-25bfa091d424@kdbg.org","threadId":"64483","inReplyTo":"CAF5LJ4B2PeLPZi5gD6Htqdwhj5T-5U9Od_NhDe-8kXTN1-v6_Q@mail.gmail.com","subject":"Re: [PATCH v2 1/2] diff: \"lisp\" userdiff_driver","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2025-11-27T10:51:23Z","receivedAt":"2025-11-27T10:51:27Z","isPatch":true,"body":"Am 27.11.25 um 11:32 schrieb Scott L. Burson:\n> On Wed, Nov 26, 2025, 6:38 PM Scott L. Burson via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n> [duplicate of first message]\n> \n> Hmm.  Somehow I have screwed things up so that GitGitGadget sends out\n> only the first commit, not including the changes in the second, nor\n> using the PR title and the text in the description.  Sorry for the\n> noise.  Is there an easy fix, or do I need to make a new PR?\nI think all went as expected as far as GGG is concerend. You made a new\ncommit on top of the one in the earlier round, and then asked GGG to\nsubmit the PR to the mailing list. GGG made a patch series with a cover\nletter and the two patches. That's expected, because the first patch\nhasn't been integrated in upstream Git, yet.\n\nIf you intended to send just the second patch (as a fixup of the first\none), then GGG cannot help you, not even if you make another pull request.\n\nBut you shouldn't have sent a fixup commit anyway, but that is a matter\nI'll address in a separate message.\n\n-- Hannes\n\n"},{"id":"531376","messageId":"b6656e6d-d1e8-4ebe-821f-9211643a71ab@kdbg.org","threadId":"64483","inReplyTo":"86315aa3e36afa1ee741a2c9b9e95a71ca569302.1764211096.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v2 2/2] merge with Scheme regexp; fix bugs","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2025-11-27T16:09:32Z","receivedAt":"2025-11-27T16:09:43Z","isPatch":true,"body":"Am 27.11.25 um 03:38 schrieb Scott L. Burson via GitGitGadget:\n> From: \"Scott L. Burson\" <Scott@sympoiesis.com>\n> \n> This commit merges (by disjoining) the new generic Lisp regexp into\n> the existing Scheme regexp.  It also fixes two bugs: the new regexp\n> was unintentionally allowing tabs, and the matching of \"(def\" should\n> be case-insensitive.\n> \n> Signed-off-by: Scott L. Burson <Scott@sympoiesis.com>\n> ---\n>  userdiff.c | 25 ++++++++++++-------------\n>  1 file changed, 12 insertions(+), 13 deletions(-)\n> \n> diff --git a/userdiff.c b/userdiff.c\n> index e127b4a1f1..b67dfddbef 100644\n> --- a/userdiff.c\n> +++ b/userdiff.c\n> @@ -249,14 +249,6 @@ PATTERNS(\"kotlin\",\n>  \t \"|[.][0-9][0-9_]*([Ee][-+]?[0-9]+)?[fFlLuU]?\"\n>  \t /* unary and binary operators */\n>  \t \"|[-+*/<>%&^|=!]==?|--|\\\\+\\\\+|<<=|>>=|&&|\\\\|\\\\||->|\\\\.\\\\*|!!|[?:.][.:]\"),\n> -PATTERNS(\"lisp\",\n> -\t /* Either an unindented left paren, or a slightly indented line\n> -\t  * starting with \"(def\" */\n> -\t \"^((\\\\(|:space:{1,2}\\\\(def).*)$\",\n> -\t /* Common Lisp symbol syntax allows arbitrary strings between vertical bars */\n> -\t \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n> -\t /* All other words are delimited by spaces or parentheses/brackets/braces */\n> -\t \"|([^][(){} \\t])+\"),\n>  PATTERNS(\"markdown\",\n>  \t \"^ {0,3}#{1,6}[ \\t].*\",\n>  \t /* -- */\n\nYou made this commit a fixup commit of the commit from the first round.\nThis isn't desirable as long as the earlier patch has not been\nintegrated in \"next\", yet.\n\nYou should have squashed the commits into one. The cover letter gives a\nreally good justification for this change and should be the commit's\nmessage (with its subject line, ie, the PR title). However, don't write\n\"This commit does X\", but write \"Do X\" instead: you give someone an\norder to change the code. (Also, after squashing there is no bug to fix\nanymore, of course.)\n\n> @@ -352,14 +344,21 @@ PATTERNS(\"rust\",\n>  \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>  \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n>  PATTERNS(\"scheme\",\n> -\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\",\n> +\t /* A possibly indented left paren followed by a Scheme keyword. */\n> +\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\\n\"\n\nMental note how this RE is nested:\n\n\t[\\t ]*(\n\t\t\\((\n\t\t\t(\n\t\t\t\tdefine|def(\n\t\t\t\t\tstruct|syntax|class|method\n\t\t\t\t\t|rules|record|proto|alias\n\t\t\t\t)?\n\t\t\t)[-*/ \\t]\n\t\t\t|\n\t\t\t(\n\t\t\t\tlibrary|module|struct|class\n\t\t\t)[*+ \\t]\n\t\t).*\n\t)$\n\n> +\t /*\n> +\t  * For other Lisp dialects: either an unindented left paren, or a\n> +\t  * slightly indented line starting with \"(def\".\n> +\t  */\n> +\t \"^((\\\\(| {1,2}\\\\([Dd][Ee][Ff]).*)$\",\n\nHere you are adding a very generous new pattern, the opening parenthesis\nwithout indentation. This will not only apply to \"other Lisp dialects\",\nas the comment says, but also Scheme code and will produce new matches.\nIt does not change the test cases in t/t4018/scheme-*, because all have\nadditional matches later.\n\nAs such it would possibly be more honest to extract it out into its own\n(first) pattern and marked as applying to all dialects:\n\n\t/*\n\t * An unindented opening parenthesis identifies a top-level\n         * structure in all Lisp dialects.\n\t */\n\t\"^(\\\\(.*)$\\n\",\n\nNote that the Scheme pattern excludes the indentation from the capture.\nYou may want to do so here, too (and simplify \"one or two spaces\" like\nthis):\n\n\t\"^  ?(\\\\([Dd][Ee][Ff].*)$\",\n\nWould it be possible to have test cases of Lisp code that is not covered\nby the Scheme pattern?\n\n>  \t /*\n> -\t  * R7RS valid identifiers include any sequence enclosed\n> -\t  * within vertical lines having no backslashes\n> +\t  * The union of R7RS and Common Lisp symbol syntax: allows arbitrary\n> +\t  * strings between vertical bars, including escaped backslashes and\n> +\t  * vertical bars.\n>  \t  */\n> -\t \"\\\\|([^\\\\\\\\]*)\\\\|\"\n> +\t \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n\nWithout the C quoting we have\n\n\t\\|([^\\\\]|\\\\\\\\|\\\\\\|)*\\|\n\nSo, this is everthing from | up to the next |, except that \\| does not\nstop scanning and \\\\ is also considered so that \\\\| is not regarded as \\\nfollowed by \\|. Good.\n\n>  \t /* All other words should be delimited by spaces or parentheses */\n> -\t \"|([^][)(}{[ \\t])+\"),\n> +\t \"|([^][)(}{ \\t])+\"),\n\nHere we have a single bracket expression. The removed opening [ does not\nbegin a new one, but is a duplicated character. Good.\n\n-- Hannes\n\n"},{"id":"531546","messageId":"7c642644-09a5-4a50-931b-a630d459932d@kdbg.org","threadId":"64483","inReplyTo":"b6656e6d-d1e8-4ebe-821f-9211643a71ab@kdbg.org","subject":"Re: [PATCH v2 2/2] merge with Scheme regexp; fix bugs","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2025-12-02T10:27:38Z","receivedAt":"2025-12-02T10:27:48Z","isPatch":true,"body":"Am 27.11.25 um 17:09 schrieb Johannes Sixt:\n> Am 27.11.25 um 03:38 schrieb Scott L. Burson via GitGitGadget:\n>>  \t /*\n>> -\t  * R7RS valid identifiers include any sequence enclosed\n>> -\t  * within vertical lines having no backslashes\n>> +\t  * The union of R7RS and Common Lisp symbol syntax: allows arbitrary\n>> +\t  * strings between vertical bars, including escaped backslashes and\n>> +\t  * vertical bars.\n>>  \t  */\n>> -\t \"\\\\|([^\\\\\\\\]*)\\\\|\"\n>> +\t \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n> \n> Without the C quoting we have\n> \n> \t\\|([^\\\\]|\\\\\\\\|\\\\\\|)*\\|\n> \n> So, this is everthing from | up to the next |, except that \\| does not\n> stop scanning and \\\\ is also considered so that \\\\| is not regarded as \\\n> followed by \\|. Good.\n\nActually, no. Regular expressions don't choose the first match if a\ndifferent alternative gives a longer match in total. For example, for\nthe change\n\n-  (let ((|one two| |three four|)))\n+  (let ((|1 two| |three four|)))\n\nwe get to see the word diff\n\n  (let (([-|one two| |three four|-]{+|1 two| |three four|+})))\n\nbut the desired result is\n\n  (let (([-|one two|-]{+|1 two|+} |three four|)))\n\nThe problem isn't new with the proposed change, but if we change the RE,\nwe could fix this at the same time. I think it helps to include | in the\nbracket expression. It may be worth its own patch that also adds a test\nin t/t4034/scheme/.\n\nThe worddiff test case is a bit too sloppy. I've tightened it in the\npatch below. You may want to make it the first of your series. (If you\ndo, don't forget to apply your sign-off when you cherry-pick the\ncommit.) It is also available here:\nhttps://github.com/j6t/git/tree/userdiff-scheme\n(commit 8f6cb42a02cc).\n\n----- 8< -----\nuserdiff: tighten word-diff test case of the scheme driver\n\nThe scheme driver separates identifiers only at parentheses of all\nsorts and whitespace, except that vertical bars act as brackets that\nenclose an identifier.\n\nThe test case attempts to demonstrate the vertical bars with a change\nfrom 'some-text' to '|a greeting|'. However, this misses the goal\nbecause the same word coloring would be applied if '|a greeting|'\nwere parsed as two words.\n\nHave an identifier between vertical bars with a space in both the pre-\nand the post-image and change only one side of the space to show that\nthe single word exists between the vertical bars.\n\nAlso add cases that change parentheses of all kinds in a sequence of\nparentheses to show that they are their own word each.\n\nSigned-off-by: Johannes Sixt <j6t@kdbg.org>\n---\n t/t4034/scheme/expect | 5 +++--\n t/t4034/scheme/post   | 1 +\n t/t4034/scheme/pre    | 3 ++-\n 3 files changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\nindex 496cd5de8c..138abe9f56 100644\n--- a/t/t4034/scheme/expect\n+++ b/t/t4034/scheme/expect\n@@ -2,10 +2,11 @@\n <BOLD>index 74b6605..63b6ac4 100644<RESET>\n <BOLD>--- a/pre<RESET>\n <BOLD>+++ b/post<RESET>\n-<CYAN>@@ -1,6 +1,6 @@<RESET>\n+<CYAN>@@ -1,7 +1,7 @@<RESET>\n (define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n   ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n   (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n-  (define <RED>some-text<RESET><GREEN>|a greeting|<RESET> \"hello\")\n+  (define <RED>|the greeting|<RESET><GREEN>|a greeting|<RESET> \"hello\")\n+  ({<RED>}<RESET>(([<RED>]<RESET>(func-n)<RED>[<RESET>]))<RED>{<RESET>})\n   (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n     (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\ndiff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\nindex 63b6ac4f87..0e3bab101d 100644\n--- a/t/t4034/scheme/post\n+++ b/t/t4034/scheme/post\n@@ -2,5 +2,6 @@\n   ; This is a (moderately) cool function.\n   (that\\place (+ 3 4))\n   (define |a greeting| \"hello\")\n+  ({(([(func-n)]))})\n   (let ((c (add1 first)))\n     (format \"one more than the total is %d\" (+ c second))))\ndiff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\nindex 74b6605357..03d77c7c43 100644\n--- a/t/t4034/scheme/pre\n+++ b/t/t4034/scheme/pre\n@@ -1,6 +1,7 @@\n (define (myfunc a b)\n   ; This is a really cool function.\n   (this\\place (+ 3 4))\n-  (define some-text \"hello\")\n+  (define |the greeting| \"hello\")\n+  ({}(([](func-n)[])){})\n   (let ((c (+ a b)))\n     (format \"one more than the total is %d\" (add1 c))))\n-- \n2.52.0.rc0.206.g6c0125c11f\n\n"},{"id":"533794","messageId":"CAF5LJ4DrKkJpCfOkkEsYvDH7qF1Bx-v75GryxUbr6UgmJq05cw@mail.gmail.com","threadId":"64483","inReplyTo":"7c642644-09a5-4a50-931b-a630d459932d@kdbg.org","subject":"Re: [PATCH v2 2/2] merge with Scheme regexp; fix bugs","fromName":"Scott L. Burson","fromEmail":"scott@sympoiesis.com","sentAt":"2026-01-14T06:18:01Z","receivedAt":"2026-01-14T06:18:42Z","isPatch":true,"body":"On Tue, Dec 2, 2025 at 2:27 AM Johannes Sixt <j6t@kdbg.org> wrote:\n>\n> Am 27.11.25 um 17:09 schrieb Johannes Sixt:\n> > Am 27.11.25 um 03:38 schrieb Scott L. Burson via GitGitGadget:\n> >>       /*\n> >> -      * R7RS valid identifiers include any sequence enclosed\n> >> -      * within vertical lines having no backslashes\n> >> +      * The union of R7RS and Common Lisp symbol syntax: allows arbitrary\n> >> +      * strings between vertical bars, including escaped backslashes and\n> >> +      * vertical bars.\n> >>        */\n> >> -     \"\\\\|([^\\\\\\\\]*)\\\\|\"\n> >> +     \"\\\\|([^\\\\\\\\]|\\\\\\\\\\\\\\\\|\\\\\\\\\\\\|)*\\\\|\"\n> >\n> > Without the C quoting we have\n> >\n> >       \\|([^\\\\]|\\\\\\\\|\\\\\\|)*\\|\n> >\n> > So, this is everthing from | up to the next |, except that \\| does not\n> > stop scanning and \\\\ is also considered so that \\\\| is not regarded as \\\n> > followed by \\|. Good.\n>\n> Actually, no. Regular expressions don't choose the first match if a\n> different alternative gives a longer match in total.\n\nAh, good catch.\n\nI noticed another bug.  At least in Common Lisp, and I expect also in\nScheme, while backslash and vertical bar are the only characters that\nmust be escaped to be included, in fact any character _may_ be\nescaped.  (This came to my attention when Emacs Paredit escaped\na double-quote for me, between vertical bars, unnecessarily.  Of\ncourse, in a string, double-quote would need to be escaped.)\n\nSo the correct regexp, with both of these bugs fixed, is\n\n    \"\\\\|([^|\\\\\\\\]|\\\\\\\\.)*\\\\|\"\n\nOr, without the C quoting:\n\n    \\|([^|\\\\]|\\\\.)*\\|\n\nFor example, for\n> the change\n>\n> -  (let ((|one two| |three four|)))\n> +  (let ((|1 two| |three four|)))\n>\n> we get to see the word diff\n>\n>   (let (([-|one two| |three four|-]{+|1 two| |three four|+})))\n>\n> but the desired result is\n>\n>   (let (([-|one two|-]{+|1 two|+} |three four|)))\n>\n> I think it helps to include | in the bracket expression.\n\nDone.\n\n> It may be worth its own patch that also adds a test in t/t4034/scheme/.\n\nI updated the existing test to check for both of these bugs (and verified\nthat it did so by reintroducing them).  It's all in this one patch.\n\nThe branch now contains two commits, yours and mine.  Do I just\ndo /submit at this point, or do I need to submit them separately?\n"},{"id":"533808","messageId":"e0e82af6-7577-43ba-beef-944715d7743b@kdbg.org","threadId":"64483","inReplyTo":"CAF5LJ4DrKkJpCfOkkEsYvDH7qF1Bx-v75GryxUbr6UgmJq05cw@mail.gmail.com","subject":"Re: [PATCH v2 2/2] merge with Scheme regexp; fix bugs","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2026-01-14T08:40:34Z","receivedAt":"2026-01-14T08:40:53Z","isPatch":true,"body":"Am 14.01.26 um 07:18 schrieb Scott L. Burson:\n> The branch now contains two commits, yours and mine.  Do I just\n> do /submit at this point, or do I need to submit them separately?\n\nYou should submit both commits with a single /submit command.\n\nBefore you do so, though, you may want to update the authorship of my\npatch (`git commit --author=\"Johannes Sixt <j6t@kdbg.org>\" --amend`).\nThat you have appended your own sign-off to the commit message is very\nmuch appreciated.\n\n-- Hannes\n\n"},{"id":"534001","messageId":"pull.2000.v3.git.1768519120.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":"pull.2000.v2.git.1764211096.gitgitgadget@gmail.com","subject":"[PATCH v3 0/2] userdiff: extend Scheme support to cover other Lisp dialects","fromName":"Scott L. Burson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-15T23:18:38Z","receivedAt":"2026-01-15T23:18:43Z","isPatch":true,"body":"Common Lisp, Emacs Lisp, and other dialects have some top-level forms, most\nimportantly 'defun', that are not matched by the current Scheme pattern.\nAlso, it is common in these dialects, when defining user macros intended as\ntop-level forms, to prefix their names with \"def\" instead of \"define\"; such\nforms are also not currently matched. Some such forms don't even begin with\n\"def\".\n\nOn the other hand, it is an established formatting convention in the Lisp\ncommunity that only top-level forms start at the left margin. So matching\nany unindented line starting with an open parenthesis is an acceptable\nheuristic; false positives will be rare.\n\nHowever, there are also cases where notionally top-level forms are grouped\ntogether within some containing form. At least in the Common Lisp community,\nit is conventional to indent these by two spaces, or sometimes one. But\nmatching just an open parenthesis indented by two spaces would be too broad;\nso the pattern added by this commit requires an indented form to start with\n\"(def\". It is believed that this strikes a good balance between potential\nfalse positives and false negatives.\n\nThis commit disjoins a regexp employing these heuristics to the existing\nScheme regexp, so it will still match everything that it did previously.\n\nJohannes Sixt (1):\n  userdiff: tighten word-diff test case of the scheme driver\n\nScott L. Burson (1):\n  userdiff: extend Scheme support to cover other Lisp dialects\n\n Documentation/gitattributes.adoc           |  1 +\n t/t4018/scheme-lisp-defun-a                |  4 ++++\n t/t4018/scheme-lisp-defun-b                |  4 ++++\n t/t4018/scheme-lisp-eval-when              |  4 ++++\n t/t4018/{scheme-module => scheme-module-a} |  0\n t/t4018/scheme-module-b                    |  6 ++++++\n t/t4034/scheme/expect                      |  5 +++--\n t/t4034/scheme/post                        |  3 ++-\n t/t4034/scheme/pre                         |  3 ++-\n userdiff.c                                 | 22 ++++++++++++++++------\n 10 files changed, 42 insertions(+), 10 deletions(-)\n create mode 100644 t/t4018/scheme-lisp-defun-a\n create mode 100644 t/t4018/scheme-lisp-defun-b\n create mode 100644 t/t4018/scheme-lisp-eval-when\n rename t/t4018/{scheme-module => scheme-module-a} (100%)\n create mode 100644 t/t4018/scheme-module-b\n\n\nbase-commit: 8745eae506f700657882b9e32b2aa00f234a6fb6\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2000%2Fslburson%2Flisp-userdiff_driver-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2000/slburson/lisp-userdiff_driver-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/2000\n\nRange-diff vs v2:\n\n 1:  da99bb0bcd < -:  ---------- diff: \"lisp\" userdiff_driver\n 2:  86315aa3e3 < -:  ---------- merge with Scheme regexp; fix bugs\n -:  ---------- > 1:  e20ac5b6a6 userdiff: tighten word-diff test case of the scheme driver\n -:  ---------- > 2:  fb4c8dc5d4 userdiff: extend Scheme support to cover other Lisp dialects\n\n-- \ngitgitgadget\n"},{"id":"534002","messageId":"e20ac5b6a6257e909fc676f6472230540268146b.1768519120.git.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":"pull.2000.v3.git.1768519120.gitgitgadget@gmail.com","subject":"[PATCH v3 1/2] userdiff: tighten word-diff test case of the scheme driver","fromName":"Johannes Sixt via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-15T23:18:39Z","receivedAt":"2026-01-15T23:18:44Z","isPatch":true,"body":"From: Johannes Sixt <j6t@kdbg.org>\n\nThe scheme driver separates identifiers only at parentheses of all\nsorts and whitespace, except that vertical bars act as brackets that\nenclose an identifier.\n\nThe test case attempts to demonstrate the vertical bars with a change\nfrom 'some-text' to '|a greeting|'. However, this misses the goal\nbecause the same word coloring would be applied if '|a greeting|'\nwere parsed as two words.\n\nHave an identifier between vertical bars with a space in both the pre-\nand the post-image and change only one side of the space to show that\nthe single word exists between the vertical bars.\n\nAlso add cases that change parentheses of all kinds in a sequence of\nparentheses to show that they are their own word each.\n\nSigned-off-by: Johannes Sixt <j6t@kdbg.org>\nSigned-off-by: Scott L. Burson <Scott@sympoiesis.com>\n---\n t/t4034/scheme/expect | 5 +++--\n t/t4034/scheme/post   | 1 +\n t/t4034/scheme/pre    | 3 ++-\n 3 files changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\nindex 496cd5de8c..138abe9f56 100644\n--- a/t/t4034/scheme/expect\n+++ b/t/t4034/scheme/expect\n@@ -2,10 +2,11 @@\n <BOLD>index 74b6605..63b6ac4 100644<RESET>\n <BOLD>--- a/pre<RESET>\n <BOLD>+++ b/post<RESET>\n-<CYAN>@@ -1,6 +1,6 @@<RESET>\n+<CYAN>@@ -1,7 +1,7 @@<RESET>\n (define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n   ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n   (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n-  (define <RED>some-text<RESET><GREEN>|a greeting|<RESET> \"hello\")\n+  (define <RED>|the greeting|<RESET><GREEN>|a greeting|<RESET> \"hello\")\n+  ({<RED>}<RESET>(([<RED>]<RESET>(func-n)<RED>[<RESET>]))<RED>{<RESET>})\n   (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n     (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\ndiff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\nindex 63b6ac4f87..0e3bab101d 100644\n--- a/t/t4034/scheme/post\n+++ b/t/t4034/scheme/post\n@@ -2,5 +2,6 @@\n   ; This is a (moderately) cool function.\n   (that\\place (+ 3 4))\n   (define |a greeting| \"hello\")\n+  ({(([(func-n)]))})\n   (let ((c (add1 first)))\n     (format \"one more than the total is %d\" (+ c second))))\ndiff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\nindex 74b6605357..03d77c7c43 100644\n--- a/t/t4034/scheme/pre\n+++ b/t/t4034/scheme/pre\n@@ -1,6 +1,7 @@\n (define (myfunc a b)\n   ; This is a really cool function.\n   (this\\place (+ 3 4))\n-  (define some-text \"hello\")\n+  (define |the greeting| \"hello\")\n+  ({}(([](func-n)[])){})\n   (let ((c (+ a b)))\n     (format \"one more than the total is %d\" (add1 c))))\n-- \ngitgitgadget\n\n"},{"id":"534003","messageId":"fb4c8dc5d4434deab9c8f1872f309a79351dc799.1768519120.git.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":"pull.2000.v3.git.1768519120.gitgitgadget@gmail.com","subject":"[PATCH v3 2/2] userdiff: extend Scheme support to cover other Lisp dialects","fromName":"Scott L. Burson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-01-15T23:18:40Z","receivedAt":"2026-01-15T23:18:45Z","isPatch":true,"body":"From: \"Scott L. Burson\" <Scott@sympoiesis.com>\n\nCommon Lisp has top-level forms, such as 'defun' and 'defmacro', that\nare not matched by the current Scheme pattern.  Also, it is more\ncommon in CL, when defining user macros intended as top-level forms,\nto prefix their names with \"def\" instead of \"define\"; such forms are\nalso not matched.  And some top-level forms don't even begin with\n\"def\".\n\nOn the other hand, it is an established formatting convention in the\nLisp community that only top-level forms start at the left margin.  So\nmatching any unindented line starting with an open parenthesis is an\nacceptable heuristic; false positives will be rare.\n\nHowever, there are also cases where notionally top-level forms are\ngrouped together within some containing form.  At least in the Common\nLisp community, it is conventional to indent these by two spaces, or\nsometimes one.  But matching just an open parenthesis indented by two\nspaces would be too broad; so the pattern added by this commit\nrequires an indented form to start with \"(def\".  It is believed that\nthis strikes a good balance between potential false positives and\nfalse negatives.\n\nSigned-off-by: Scott L. Burson <Scott@sympoiesis.com>\n---\n Documentation/gitattributes.adoc           |  1 +\n t/t4018/scheme-lisp-defun-a                |  4 ++++\n t/t4018/scheme-lisp-defun-b                |  4 ++++\n t/t4018/scheme-lisp-eval-when              |  4 ++++\n t/t4018/{scheme-module => scheme-module-a} |  0\n t/t4018/scheme-module-b                    |  6 ++++++\n t/t4034/scheme/expect                      |  2 +-\n t/t4034/scheme/post                        |  2 +-\n t/t4034/scheme/pre                         |  2 +-\n userdiff.c                                 | 22 ++++++++++++++++------\n 10 files changed, 38 insertions(+), 9 deletions(-)\n create mode 100644 t/t4018/scheme-lisp-defun-a\n create mode 100644 t/t4018/scheme-lisp-defun-b\n create mode 100644 t/t4018/scheme-lisp-eval-when\n rename t/t4018/{scheme-module => scheme-module-a} (100%)\n create mode 100644 t/t4018/scheme-module-b\n\ndiff --git a/Documentation/gitattributes.adoc b/Documentation/gitattributes.adoc\nindex f20041a323..a9ce5adef9 100644\n--- a/Documentation/gitattributes.adoc\n+++ b/Documentation/gitattributes.adoc\n@@ -912,6 +912,7 @@ patterns are available:\n - `rust` suitable for source code in the Rust language.\n \n - `scheme` suitable for source code in the Scheme language.\n+Also handles Emacs Lisp, Common Lisp, and most other dialects.\n \n - `tex` suitable for source code for LaTeX documents.\n \ndiff --git a/t/t4018/scheme-lisp-defun-a b/t/t4018/scheme-lisp-defun-a\nnew file mode 100644\nindex 0000000000..c3c750f76d\n--- /dev/null\n+++ b/t/t4018/scheme-lisp-defun-a\n@@ -0,0 +1,4 @@\n+(defun some-func (x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\ndiff --git a/t/t4018/scheme-lisp-defun-b b/t/t4018/scheme-lisp-defun-b\nnew file mode 100644\nindex 0000000000..21be305968\n--- /dev/null\n+++ b/t/t4018/scheme-lisp-defun-b\n@@ -0,0 +1,4 @@\n+(macrolet ((foo (x) `(bar ,x)))\n+  (defun mumble (x) ; RIGHT\n+    (when (> x 0)\n+      (foo x)))) ; ChangeMe\ndiff --git a/t/t4018/scheme-lisp-eval-when b/t/t4018/scheme-lisp-eval-when\nnew file mode 100644\nindex 0000000000..5d941d7e0e\n--- /dev/null\n+++ b/t/t4018/scheme-lisp-eval-when\n@@ -0,0 +1,4 @@\n+(eval-when (:compile-toplevel :load-toplevel :execute)  ; RIGHT\n+  (set-macro-character #\\?\n+\t\t       (lambda (stream char)\n+\t\t\t `(make-pattern-variable ,(read stream)))))  ; ChangeMe\ndiff --git a/t/t4018/scheme-module b/t/t4018/scheme-module-a\nsimilarity index 100%\nrename from t/t4018/scheme-module\nrename to t/t4018/scheme-module-a\ndiff --git a/t/t4018/scheme-module-b b/t/t4018/scheme-module-b\nnew file mode 100644\nindex 0000000000..77bc0c5eff\n--- /dev/null\n+++ b/t/t4018/scheme-module-b\n@@ -0,0 +1,6 @@\n+(module A\n+  (export with-display-exception)\n+  (extern (display-exception display-exception))\n+  (def (with-display-exception thunk) RIGHT\n+    (with-catch (lambda (e) (display-exception e (current-error-port)) e)\n+      thunk ChangeMe)))\ndiff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\nindex 138abe9f56..72592665f1 100644\n--- a/t/t4034/scheme/expect\n+++ b/t/t4034/scheme/expect\n@@ -6,7 +6,7 @@\n (define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n   ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n   (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n-  (define <RED>|the greeting|<RESET><GREEN>|a greeting|<RESET> \"hello\")\n+  (define <RED>|the \\greeting|<RESET><GREEN>|a \\greeting|<RESET> |hello there|)\n   ({<RED>}<RESET>(([<RED>]<RESET>(func-n)<RED>[<RESET>]))<RED>{<RESET>})\n   (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n     (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\ndiff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\nindex 0e3bab101d..450cc234f7 100644\n--- a/t/t4034/scheme/post\n+++ b/t/t4034/scheme/post\n@@ -1,7 +1,7 @@\n (define (my-func first second)\n   ; This is a (moderately) cool function.\n   (that\\place (+ 3 4))\n-  (define |a greeting| \"hello\")\n+  (define |a \\greeting| |hello there|)\n   ({(([(func-n)]))})\n   (let ((c (add1 first)))\n     (format \"one more than the total is %d\" (+ c second))))\ndiff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\nindex 03d77c7c43..ba8b8ac0a4 100644\n--- a/t/t4034/scheme/pre\n+++ b/t/t4034/scheme/pre\n@@ -1,7 +1,7 @@\n (define (myfunc a b)\n   ; This is a really cool function.\n   (this\\place (+ 3 4))\n-  (define |the greeting| \"hello\")\n+  (define |the \\greeting| |hello there|)\n   ({}(([](func-n)[])){})\n   (let ((c (+ a b)))\n     (format \"one more than the total is %d\" (add1 c))))\ndiff --git a/userdiff.c b/userdiff.c\nindex fe710a68bf..b5412e6bc3 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -344,14 +344,24 @@ PATTERNS(\"rust\",\n \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n PATTERNS(\"scheme\",\n-\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\",\n \t /*\n-\t  * R7RS valid identifiers include any sequence enclosed\n-\t  * within vertical lines having no backslashes\n+\t  * An unindented opening parenthesis identifies a top-level\n+\t  * expression in all Lisp dialects.\n \t  */\n-\t \"\\\\|([^\\\\\\\\]*)\\\\|\"\n-\t /* All other words should be delimited by spaces or parentheses */\n-\t \"|([^][)(}{[ \\t])+\"),\n+\t \"^(\\\\(.*)$\\n\"\n+\t /* For Scheme: a possibly indented left paren followed by a keyword. */\n+\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\\n\"\n+\t /*\n+\t  * For all Lisp dialects: a slightly indented line starting with \"(def\".\n+\t  */\n+\t \"^  ?(\\\\([Dd][Ee][Ff].*)$\",\n+\t /*\n+\t  * The union of R7RS and Common Lisp symbol syntax: allows arbitrary\n+\t  * strings between vertical bars, including any escaped characters.\n+\t  */\n+\t \"\\\\|([^|\\\\\\\\]|\\\\\\\\.)*\\\\|\"\n+\t /* All other words should be delimited by spaces or parentheses. */\n+\t \"|([^][)(}{ \\t])+\"),\n PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n \t \"\\\\\\\\[a-zA-Z@]+|\\\\\\\\.|([a-zA-Z0-9]|[^\\x01-\\x7f])+\"),\n { .name = \"default\", .binary = -1 },\n-- \ngitgitgadget\n"},{"id":"534023","messageId":"3243b63b-b0c1-42d5-beeb-df42b891f09e@kdbg.org","threadId":"64483","inReplyTo":"fb4c8dc5d4434deab9c8f1872f309a79351dc799.1768519120.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 2/2] userdiff: extend Scheme support to cover other Lisp dialects","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2026-01-16T08:49:54Z","receivedAt":"2026-01-16T08:50:12Z","isPatch":true,"body":"Am 16.01.26 um 00:18 schrieb Scott L. Burson via GitGitGadget:\n> From: \"Scott L. Burson\" <Scott@sympoiesis.com>\n> \n> Common Lisp has top-level forms, such as 'defun' and 'defmacro', that\n> are not matched by the current Scheme pattern.  Also, it is more\n> common in CL, when defining user macros intended as top-level forms,\n> to prefix their names with \"def\" instead of \"define\"; such forms are\n> also not matched.  And some top-level forms don't even begin with\n> \"def\".\n> \n> On the other hand, it is an established formatting convention in the\n> Lisp community that only top-level forms start at the left margin.  So\n> matching any unindented line starting with an open parenthesis is an\n> acceptable heuristic; false positives will be rare.\n> \n> However, there are also cases where notionally top-level forms are\n> grouped together within some containing form.  At least in the Common\n> Lisp community, it is conventional to indent these by two spaces, or\n> sometimes one.  But matching just an open parenthesis indented by two\n> spaces would be too broad; so the pattern added by this commit\n> requires an indented form to start with \"(def\".  It is believed that\n> this strikes a good balance between potential false positives and\n> false negatives.\n\nThe commit message doesn't mention the changes regarding the word-diff\npattern.  I would have prefered to have them in their own patch; it\nwould make the patch text less obscure about what it actually changes.\n\n> \n> Signed-off-by: Scott L. Burson <Scott@sympoiesis.com>\n\n> diff --git a/Documentation/gitattributes.adoc b/Documentation/gitattributes.adoc\n> index f20041a323..a9ce5adef9 100644\n> --- a/Documentation/gitattributes.adoc\n> +++ b/Documentation/gitattributes.adoc\n> @@ -912,6 +912,7 @@ patterns are available:\n>  - `rust` suitable for source code in the Rust language.\n>  \n>  - `scheme` suitable for source code in the Scheme language.\n> +Also handles Emacs Lisp, Common Lisp, and most other dialects.\n\nSaying \"most dialects\" immediately begs the questions \"which dialects\nare not covered\" and \"is the dialect that I'm using covered\". Let's\nwrite it this way:\n\n- `scheme` suitable for source code in the Lisp dialects including\n  Scheme, Emacs Lisp, Common Lisp.\n\nNote the indentation of the continuation line (see the 'bash' entry, for\nexample).\n\n>  \n>  - `tex` suitable for source code for LaTeX documents.\n>  \n> diff --git a/t/t4018/scheme-lisp-defun-a b/t/t4018/scheme-lisp-defun-a\n> new file mode 100644\n> index 0000000000..c3c750f76d\n> --- /dev/null\n> +++ b/t/t4018/scheme-lisp-defun-a\n> @@ -0,0 +1,4 @@\n> +(defun some-func (x y z) RIGHT\n> +  (let ((a x)\n> +        (b y))\n> +        (ChangeMe a b)))\n\nThis also demonstrates that \"(let\" isn't picked up. Good.\n\n> diff --git a/t/t4018/scheme-lisp-defun-b b/t/t4018/scheme-lisp-defun-b\n> new file mode 100644\n> index 0000000000..21be305968\n> --- /dev/null\n> +++ b/t/t4018/scheme-lisp-defun-b\n> @@ -0,0 +1,4 @@\n> +(macrolet ((foo (x) `(bar ,x)))\n> +  (defun mumble (x) ; RIGHT\n> +    (when (> x 0)\n> +      (foo x)))) ; ChangeMe\n\nIndented \"(defun\" overrides the earlier structure that begins in the\nfirst column. Good.\n\n> diff --git a/t/t4018/scheme-lisp-eval-when b/t/t4018/scheme-lisp-eval-when\n> new file mode 100644\n> index 0000000000..5d941d7e0e\n> --- /dev/null\n> +++ b/t/t4018/scheme-lisp-eval-when\n> @@ -0,0 +1,4 @@\n> +(eval-when (:compile-toplevel :load-toplevel :execute)  ; RIGHT\n> +  (set-macro-character #\\?\n> +\t\t       (lambda (stream char)\n> +\t\t\t `(make-pattern-variable ,(read stream)))))  ; ChangeMe\n\nAny structure beginning in the first column is picked up. Good.\n\n> diff --git a/t/t4018/scheme-module b/t/t4018/scheme-module-a\n> similarity index 100%\n> rename from t/t4018/scheme-module\n> rename to t/t4018/scheme-module-a\n> diff --git a/t/t4018/scheme-module-b b/t/t4018/scheme-module-b\n> new file mode 100644\n> index 0000000000..77bc0c5eff\n> --- /dev/null\n> +++ b/t/t4018/scheme-module-b\n> @@ -0,0 +1,6 @@\n> +(module A\n> +  (export with-display-exception)\n> +  (extern (display-exception display-exception))\n> +  (def (with-display-exception thunk) RIGHT\n> +    (with-catch (lambda (e) (display-exception e (current-error-port)) e)\n> +      thunk ChangeMe)))\n\nmodule-a and module-b are basically the same text. module-b changes the\nlast line, so that \"(def\" is picked up, while module-a changes the\n\"(extern\" line, so that \"(module\" is picked up. Good.\n\n> diff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\n> index 138abe9f56..72592665f1 100644\n> --- a/t/t4034/scheme/expect\n> +++ b/t/t4034/scheme/expect\n> @@ -6,7 +6,7 @@\n>  (define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n>    ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n>    (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n> -  (define <RED>|the greeting|<RESET><GREEN>|a greeting|<RESET> \"hello\")\n> +  (define <RED>|the \\greeting|<RESET><GREEN>|a \\greeting|<RESET> |hello there|)\n>    ({<RED>}<RESET>(([<RED>]<RESET>(func-n)<RED>[<RESET>]))<RED>{<RESET>})\n>    (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n>      (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\n\nThis tests backslash between vertical bars and non-greediness of the\npattern. Good.\n\nUsing the identifier \"|the \\| greeting|\" could make the test even more\ncomplete, I think.\n\n> diff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\n> index 0e3bab101d..450cc234f7 100644\n> --- a/t/t4034/scheme/post\n> +++ b/t/t4034/scheme/post\n> @@ -1,7 +1,7 @@\n>  (define (my-func first second)\n>    ; This is a (moderately) cool function.\n>    (that\\place (+ 3 4))\n> -  (define |a greeting| \"hello\")\n> +  (define |a \\greeting| |hello there|)\n>    ({(([(func-n)]))})\n>    (let ((c (add1 first)))\n>      (format \"one more than the total is %d\" (+ c second))))\n> diff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\n> index 03d77c7c43..ba8b8ac0a4 100644\n> --- a/t/t4034/scheme/pre\n> +++ b/t/t4034/scheme/pre\n> @@ -1,7 +1,7 @@\n>  (define (myfunc a b)\n>    ; This is a really cool function.\n>    (this\\place (+ 3 4))\n> -  (define |the greeting| \"hello\")\n> +  (define |the \\greeting| |hello there|)\n>    ({}(([](func-n)[])){})\n>    (let ((c (+ a b)))\n>      (format \"one more than the total is %d\" (add1 c))))\n> diff --git a/userdiff.c b/userdiff.c\n> index fe710a68bf..b5412e6bc3 100644\n> --- a/userdiff.c\n> +++ b/userdiff.c\n> @@ -344,14 +344,24 @@ PATTERNS(\"rust\",\n>  \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n>  \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n>  PATTERNS(\"scheme\",\n> -\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\",\n>  \t /*\n> -\t  * R7RS valid identifiers include any sequence enclosed\n> -\t  * within vertical lines having no backslashes\n> +\t  * An unindented opening parenthesis identifies a top-level\n> +\t  * expression in all Lisp dialects.\n>  \t  */\n> -\t \"\\\\|([^\\\\\\\\]*)\\\\|\"\n> -\t /* All other words should be delimited by spaces or parentheses */\n> -\t \"|([^][)(}{[ \\t])+\"),\n> +\t \"^(\\\\(.*)$\\n\"\n> +\t /* For Scheme: a possibly indented left paren followed by a keyword. */\n> +\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\\n\"\n> +\t /*\n> +\t  * For all Lisp dialects: a slightly indented line starting with \"(def\".\n> +\t  */\n> +\t \"^  ?(\\\\([Dd][Ee][Ff].*)$\",\n> +\t /*\n> +\t  * The union of R7RS and Common Lisp symbol syntax: allows arbitrary\n> +\t  * strings between vertical bars, including any escaped characters.\n> +\t  */\n> +\t \"\\\\|([^|\\\\\\\\]|\\\\\\\\.)*\\\\|\"\n> +\t /* All other words should be delimited by spaces or parentheses. */\n> +\t \"|([^][)(}{ \\t])+\"),\n>  PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n>  \t \"\\\\\\\\[a-zA-Z@]+|\\\\\\\\.|([a-zA-Z0-9]|[^\\x01-\\x7f])+\"),\n>  { .name = \"default\", .binary = -1 },\n\nThis change matches my expectations.\n\n-- Hannes\n\n"},{"id":"534107","messageId":"CAF5LJ4Du-x9ND-EHCe0Npz9GaE7kinEYNnpP_416cKpZuxc9hg@mail.gmail.com","threadId":"64483","inReplyTo":"3243b63b-b0c1-42d5-beeb-df42b891f09e@kdbg.org","subject":"Re: [PATCH v3 2/2] userdiff: extend Scheme support to cover other Lisp dialects","fromName":"Scott L. Burson","fromEmail":"scott@sympoiesis.com","sentAt":"2026-01-17T02:09:17Z","receivedAt":"2026-01-17T02:09:55Z","isPatch":true,"body":"On Fri, Jan 16, 2026 at 12:49 AM Johannes Sixt <j6t@kdbg.org> wrote:\n>\n> The commit message doesn't mention the changes regarding the word-diff\n> pattern.  I would have prefered to have them in their own patch; it\n> would make the patch text less obscure about what it actually changes.\n\nOkay, will do.\n\n> > diff --git a/Documentation/gitattributes.adoc b/Documentation/gitattributes.adoc\n> > index f20041a323..a9ce5adef9 100644\n> > --- a/Documentation/gitattributes.adoc\n> > +++ b/Documentation/gitattributes.adoc\n> > @@ -912,6 +912,7 @@ patterns are available:\n> >\n> >  - `scheme` suitable for source code in the Scheme language.\n> > +Also handles Emacs Lisp, Common Lisp, and most other dialects.\n>\n> Saying \"most dialects\" immediately begs the questions \"which dialects\n> are not covered\" and \"is the dialect that I'm using covered\". Let's\n> write it this way:\n>\n> - `scheme` suitable for source code in the Lisp dialects including\n>   Scheme, Emacs Lisp, Common Lisp.\n>\n> Note the indentation of the continuation line\n\nOf course I will fix the indentation, but I don't agree with your\nproposal for the text.  There are many Lisp dialects in use, indeed\nprobably thousands; lots of people write their own.  As previously\nnoted, matching an unindented open parenthesis is a very general\nheuristic that is likely to work for the vast majority of dialects.\nWhile we can't answer the question \"is the dialect I'm using covered?\"\nfor everyone, I think the text should encourage them to give the\ndriver a try.\n\nSo how about this:\n\n- `scheme` suitable for source code in most Lisp dialects,\n  including Scheme, Emacs Lisp, Common Lisp, and Clojure.\n\nI've looked at some Clojure and I believe the proposed regexp will\nwork for it.  I think it's a good idea to mention it explicitly,\nbecause people might search the text for it.\n\n> Using the identifier \"|the \\| greeting|\" could make the test even more\n> complete, I think.\n\nAgreed, will do.\n\n-- Scott\n"},{"id":"534109","messageId":"fcfaaab9-d719-406b-bf18-0e4a0ce6a1db@kdbg.org","threadId":"64483","inReplyTo":"CAF5LJ4Du-x9ND-EHCe0Npz9GaE7kinEYNnpP_416cKpZuxc9hg@mail.gmail.com","subject":"Re: [PATCH v3 2/2] userdiff: extend Scheme support to cover other Lisp dialects","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2026-01-17T08:15:36Z","receivedAt":"2026-01-17T08:15:48Z","isPatch":true,"body":"Am 17.01.26 um 03:09 schrieb Scott L. Burson:\n> So how about this:\n> \n> - `scheme` suitable for source code in most Lisp dialects,\n>   including Scheme, Emacs Lisp, Common Lisp, and Clojure.\n\nOK, let's leave it at that.\n\n-- Hannes\n\n"},{"id":"541609","messageId":"pull.2000.v4.git.1776220063.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":"pull.2000.v3.git.1768519120.gitgitgadget@gmail.com","subject":"[PATCH v4 0/2] userdiff: extend Scheme support to cover other Lisp dialects","fromName":"Scott L. Burson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-15T02:27:41Z","receivedAt":"2026-04-15T02:27:46Z","isPatch":true,"body":"Common Lisp, Emacs Lisp, and other dialects have some top-level forms, most\nimportantly 'defun', that are not matched by the current Scheme pattern.\nAlso, it is common in these dialects, when defining user macros intended as\ntop-level forms, to prefix their names with \"def\" instead of \"define\"; such\nforms are also not currently matched. Some such forms don't even begin with\n\"def\".\n\nOn the other hand, it is an established formatting convention in the Lisp\ncommunity that only top-level forms start at the left margin. So matching\nany unindented line starting with an open parenthesis is an acceptable\nheuristic; false positives will be rare.\n\nHowever, there are also cases where notionally top-level forms are grouped\ntogether within some containing form. At least in the Common Lisp community,\nit is conventional to indent these by two spaces, or sometimes one. But\nmatching just an open parenthesis indented by two spaces would be too broad;\nso the pattern added by this commit requires an indented form to start with\n\"(def\". It is believed that this strikes a good balance between potential\nfalse positives and false negatives.\n\nThis commit disjoins a regexp employing these heuristics to the existing\nScheme regexp, so it will still match everything that it did previously.\n\nJohannes Sixt (1):\n  userdiff: tighten word-diff test case of the scheme driver\n\nScott L. Burson (1):\n  userdiff: extend Scheme support to cover other Lisp dialects\n\n Documentation/gitattributes.adoc           |  3 ++-\n t/t4018/scheme-lisp-defun-a                |  4 ++++\n t/t4018/scheme-lisp-defun-b                |  4 ++++\n t/t4018/scheme-lisp-eval-when              |  4 ++++\n t/t4018/{scheme-module => scheme-module-a} |  0\n t/t4018/scheme-module-b                    |  6 ++++++\n t/t4034/scheme/expect                      |  5 +++--\n t/t4034/scheme/post                        |  3 ++-\n t/t4034/scheme/pre                         |  3 ++-\n userdiff.c                                 | 22 ++++++++++++++++------\n 10 files changed, 43 insertions(+), 11 deletions(-)\n create mode 100644 t/t4018/scheme-lisp-defun-a\n create mode 100644 t/t4018/scheme-lisp-defun-b\n create mode 100644 t/t4018/scheme-lisp-eval-when\n rename t/t4018/{scheme-module => scheme-module-a} (100%)\n create mode 100644 t/t4018/scheme-module-b\n\n\nbase-commit: 9e8f4e9c04e3efa494e78b710e0c5f6cc77a0a5e\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-2000%2Fslburson%2Flisp-userdiff_driver-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-2000/slburson/lisp-userdiff_driver-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/2000\n\nRange-diff vs v3:\n\n 1:  e20ac5b6a6 = 1:  8e0b1e3d01 userdiff: tighten word-diff test case of the scheme driver\n 2:  fb4c8dc5d4 ! 2:  0bd51e02ba userdiff: extend Scheme support to cover other Lisp dialects\n     @@ Commit message\n      \n       ## Documentation/gitattributes.adoc ##\n      @@ Documentation/gitattributes.adoc: patterns are available:\n     + \n       - `rust` suitable for source code in the Rust language.\n       \n     - - `scheme` suitable for source code in the Scheme language.\n     -+Also handles Emacs Lisp, Common Lisp, and most other dialects.\n     +-- `scheme` suitable for source code in the Scheme language.\n     ++- `scheme` suitable for source code in most Lisp dialects,\n     ++  including Scheme, Emacs Lisp, Common Lisp, and Clojure.\n       \n       - `tex` suitable for source code for LaTeX documents.\n       \n     @@ t/t4034/scheme/expect\n         ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n         (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n      -  (define <RED>|the greeting|<RESET><GREEN>|a greeting|<RESET> \"hello\")\n     -+  (define <RED>|the \\greeting|<RESET><GREEN>|a \\greeting|<RESET> |hello there|)\n     ++  (define <RED>|the \\| \\greeting|<RESET><GREEN>|a \\greeting|<RESET> |hello there|)\n         ({<RED>}<RESET>(([<RED>]<RESET>(func-n)<RED>[<RESET>]))<RED>{<RESET>})\n         (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n           (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\n     @@ t/t4034/scheme/pre\n         ; This is a really cool function.\n         (this\\place (+ 3 4))\n      -  (define |the greeting| \"hello\")\n     -+  (define |the \\greeting| |hello there|)\n     ++  (define |the \\| \\greeting| |hello there|)\n         ({}(([](func-n)[])){})\n         (let ((c (+ a b)))\n           (format \"one more than the total is %d\" (add1 c))))\n\n-- \ngitgitgadget\n"},{"id":"541610","messageId":"8e0b1e3d013ee379335bc89801525a861047929d.1776220063.git.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":"pull.2000.v4.git.1776220063.gitgitgadget@gmail.com","subject":"[PATCH v4 1/2] userdiff: tighten word-diff test case of the scheme driver","fromName":"Johannes Sixt via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-15T02:27:42Z","receivedAt":"2026-04-15T02:27:48Z","isPatch":true,"body":"From: Johannes Sixt <j6t@kdbg.org>\n\nThe scheme driver separates identifiers only at parentheses of all\nsorts and whitespace, except that vertical bars act as brackets that\nenclose an identifier.\n\nThe test case attempts to demonstrate the vertical bars with a change\nfrom 'some-text' to '|a greeting|'. However, this misses the goal\nbecause the same word coloring would be applied if '|a greeting|'\nwere parsed as two words.\n\nHave an identifier between vertical bars with a space in both the pre-\nand the post-image and change only one side of the space to show that\nthe single word exists between the vertical bars.\n\nAlso add cases that change parentheses of all kinds in a sequence of\nparentheses to show that they are their own word each.\n\nSigned-off-by: Johannes Sixt <j6t@kdbg.org>\nSigned-off-by: Scott L. Burson <Scott@sympoiesis.com>\n---\n t/t4034/scheme/expect | 5 +++--\n t/t4034/scheme/post   | 1 +\n t/t4034/scheme/pre    | 3 ++-\n 3 files changed, 6 insertions(+), 3 deletions(-)\n\ndiff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\nindex 496cd5de8c..138abe9f56 100644\n--- a/t/t4034/scheme/expect\n+++ b/t/t4034/scheme/expect\n@@ -2,10 +2,11 @@\n <BOLD>index 74b6605..63b6ac4 100644<RESET>\n <BOLD>--- a/pre<RESET>\n <BOLD>+++ b/post<RESET>\n-<CYAN>@@ -1,6 +1,6 @@<RESET>\n+<CYAN>@@ -1,7 +1,7 @@<RESET>\n (define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n   ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n   (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n-  (define <RED>some-text<RESET><GREEN>|a greeting|<RESET> \"hello\")\n+  (define <RED>|the greeting|<RESET><GREEN>|a greeting|<RESET> \"hello\")\n+  ({<RED>}<RESET>(([<RED>]<RESET>(func-n)<RED>[<RESET>]))<RED>{<RESET>})\n   (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n     (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\ndiff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\nindex 63b6ac4f87..0e3bab101d 100644\n--- a/t/t4034/scheme/post\n+++ b/t/t4034/scheme/post\n@@ -2,5 +2,6 @@\n   ; This is a (moderately) cool function.\n   (that\\place (+ 3 4))\n   (define |a greeting| \"hello\")\n+  ({(([(func-n)]))})\n   (let ((c (add1 first)))\n     (format \"one more than the total is %d\" (+ c second))))\ndiff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\nindex 74b6605357..03d77c7c43 100644\n--- a/t/t4034/scheme/pre\n+++ b/t/t4034/scheme/pre\n@@ -1,6 +1,7 @@\n (define (myfunc a b)\n   ; This is a really cool function.\n   (this\\place (+ 3 4))\n-  (define some-text \"hello\")\n+  (define |the greeting| \"hello\")\n+  ({}(([](func-n)[])){})\n   (let ((c (+ a b)))\n     (format \"one more than the total is %d\" (add1 c))))\n-- \ngitgitgadget\n\n"},{"id":"541611","messageId":"0bd51e02ba1aec92f2149a3c870af2dd1fc200b4.1776220063.git.gitgitgadget@gmail.com","threadId":"64483","inReplyTo":"pull.2000.v4.git.1776220063.gitgitgadget@gmail.com","subject":"[PATCH v4 2/2] userdiff: extend Scheme support to cover other Lisp dialects","fromName":"Scott L. Burson via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-15T02:27:43Z","receivedAt":"2026-04-15T02:27:50Z","isPatch":true,"body":"From: \"Scott L. Burson\" <Scott@sympoiesis.com>\n\nCommon Lisp has top-level forms, such as 'defun' and 'defmacro', that\nare not matched by the current Scheme pattern.  Also, it is more\ncommon in CL, when defining user macros intended as top-level forms,\nto prefix their names with \"def\" instead of \"define\"; such forms are\nalso not matched.  And some top-level forms don't even begin with\n\"def\".\n\nOn the other hand, it is an established formatting convention in the\nLisp community that only top-level forms start at the left margin.  So\nmatching any unindented line starting with an open parenthesis is an\nacceptable heuristic; false positives will be rare.\n\nHowever, there are also cases where notionally top-level forms are\ngrouped together within some containing form.  At least in the Common\nLisp community, it is conventional to indent these by two spaces, or\nsometimes one.  But matching just an open parenthesis indented by two\nspaces would be too broad; so the pattern added by this commit\nrequires an indented form to start with \"(def\".  It is believed that\nthis strikes a good balance between potential false positives and\nfalse negatives.\n\nSigned-off-by: Scott L. Burson <Scott@sympoiesis.com>\n---\n Documentation/gitattributes.adoc           |  3 ++-\n t/t4018/scheme-lisp-defun-a                |  4 ++++\n t/t4018/scheme-lisp-defun-b                |  4 ++++\n t/t4018/scheme-lisp-eval-when              |  4 ++++\n t/t4018/{scheme-module => scheme-module-a} |  0\n t/t4018/scheme-module-b                    |  6 ++++++\n t/t4034/scheme/expect                      |  2 +-\n t/t4034/scheme/post                        |  2 +-\n t/t4034/scheme/pre                         |  2 +-\n userdiff.c                                 | 22 ++++++++++++++++------\n 10 files changed, 39 insertions(+), 10 deletions(-)\n create mode 100644 t/t4018/scheme-lisp-defun-a\n create mode 100644 t/t4018/scheme-lisp-defun-b\n create mode 100644 t/t4018/scheme-lisp-eval-when\n rename t/t4018/{scheme-module => scheme-module-a} (100%)\n create mode 100644 t/t4018/scheme-module-b\n\ndiff --git a/Documentation/gitattributes.adoc b/Documentation/gitattributes.adoc\nindex f20041a323..bd76167a45 100644\n--- a/Documentation/gitattributes.adoc\n+++ b/Documentation/gitattributes.adoc\n@@ -911,7 +911,8 @@ patterns are available:\n \n - `rust` suitable for source code in the Rust language.\n \n-- `scheme` suitable for source code in the Scheme language.\n+- `scheme` suitable for source code in most Lisp dialects,\n+  including Scheme, Emacs Lisp, Common Lisp, and Clojure.\n \n - `tex` suitable for source code for LaTeX documents.\n \ndiff --git a/t/t4018/scheme-lisp-defun-a b/t/t4018/scheme-lisp-defun-a\nnew file mode 100644\nindex 0000000000..c3c750f76d\n--- /dev/null\n+++ b/t/t4018/scheme-lisp-defun-a\n@@ -0,0 +1,4 @@\n+(defun some-func (x y z) RIGHT\n+  (let ((a x)\n+        (b y))\n+        (ChangeMe a b)))\ndiff --git a/t/t4018/scheme-lisp-defun-b b/t/t4018/scheme-lisp-defun-b\nnew file mode 100644\nindex 0000000000..21be305968\n--- /dev/null\n+++ b/t/t4018/scheme-lisp-defun-b\n@@ -0,0 +1,4 @@\n+(macrolet ((foo (x) `(bar ,x)))\n+  (defun mumble (x) ; RIGHT\n+    (when (> x 0)\n+      (foo x)))) ; ChangeMe\ndiff --git a/t/t4018/scheme-lisp-eval-when b/t/t4018/scheme-lisp-eval-when\nnew file mode 100644\nindex 0000000000..5d941d7e0e\n--- /dev/null\n+++ b/t/t4018/scheme-lisp-eval-when\n@@ -0,0 +1,4 @@\n+(eval-when (:compile-toplevel :load-toplevel :execute)  ; RIGHT\n+  (set-macro-character #\\?\n+\t\t       (lambda (stream char)\n+\t\t\t `(make-pattern-variable ,(read stream)))))  ; ChangeMe\ndiff --git a/t/t4018/scheme-module b/t/t4018/scheme-module-a\nsimilarity index 100%\nrename from t/t4018/scheme-module\nrename to t/t4018/scheme-module-a\ndiff --git a/t/t4018/scheme-module-b b/t/t4018/scheme-module-b\nnew file mode 100644\nindex 0000000000..77bc0c5eff\n--- /dev/null\n+++ b/t/t4018/scheme-module-b\n@@ -0,0 +1,6 @@\n+(module A\n+  (export with-display-exception)\n+  (extern (display-exception display-exception))\n+  (def (with-display-exception thunk) RIGHT\n+    (with-catch (lambda (e) (display-exception e (current-error-port)) e)\n+      thunk ChangeMe)))\ndiff --git a/t/t4034/scheme/expect b/t/t4034/scheme/expect\nindex 138abe9f56..fb7f2616fe 100644\n--- a/t/t4034/scheme/expect\n+++ b/t/t4034/scheme/expect\n@@ -6,7 +6,7 @@\n (define (<RED>myfunc a b<RESET><GREEN>my-func first second<RESET>)\n   ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n   (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n-  (define <RED>|the greeting|<RESET><GREEN>|a greeting|<RESET> \"hello\")\n+  (define <RED>|the \\| \\greeting|<RESET><GREEN>|a \\greeting|<RESET> |hello there|)\n   ({<RED>}<RESET>(([<RED>]<RESET>(func-n)<RED>[<RESET>]))<RED>{<RESET>})\n   (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n     (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\ndiff --git a/t/t4034/scheme/post b/t/t4034/scheme/post\nindex 0e3bab101d..450cc234f7 100644\n--- a/t/t4034/scheme/post\n+++ b/t/t4034/scheme/post\n@@ -1,7 +1,7 @@\n (define (my-func first second)\n   ; This is a (moderately) cool function.\n   (that\\place (+ 3 4))\n-  (define |a greeting| \"hello\")\n+  (define |a \\greeting| |hello there|)\n   ({(([(func-n)]))})\n   (let ((c (add1 first)))\n     (format \"one more than the total is %d\" (+ c second))))\ndiff --git a/t/t4034/scheme/pre b/t/t4034/scheme/pre\nindex 03d77c7c43..e16ee75849 100644\n--- a/t/t4034/scheme/pre\n+++ b/t/t4034/scheme/pre\n@@ -1,7 +1,7 @@\n (define (myfunc a b)\n   ; This is a really cool function.\n   (this\\place (+ 3 4))\n-  (define |the greeting| \"hello\")\n+  (define |the \\| \\greeting| |hello there|)\n   ({}(([](func-n)[])){})\n   (let ((c (+ a b)))\n     (format \"one more than the total is %d\" (add1 c))))\ndiff --git a/userdiff.c b/userdiff.c\nindex fe710a68bf..b5412e6bc3 100644\n--- a/userdiff.c\n+++ b/userdiff.c\n@@ -344,14 +344,24 @@ PATTERNS(\"rust\",\n \t \"|[0-9][0-9_a-fA-Fiosuxz]*(\\\\.([0-9]*[eE][+-]?)?[0-9_fF]*)?\"\n \t \"|[-+*\\\\/<>%&^|=!:]=|<<=?|>>=?|&&|\\\\|\\\\||->|=>|\\\\.{2}=|\\\\.{3}|::\"),\n PATTERNS(\"scheme\",\n-\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\",\n \t /*\n-\t  * R7RS valid identifiers include any sequence enclosed\n-\t  * within vertical lines having no backslashes\n+\t  * An unindented opening parenthesis identifies a top-level\n+\t  * expression in all Lisp dialects.\n \t  */\n-\t \"\\\\|([^\\\\\\\\]*)\\\\|\"\n-\t /* All other words should be delimited by spaces or parentheses */\n-\t \"|([^][)(}{[ \\t])+\"),\n+\t \"^(\\\\(.*)$\\n\"\n+\t /* For Scheme: a possibly indented left paren followed by a keyword. */\n+\t \"^[\\t ]*(\\\\(((define|def(struct|syntax|class|method|rules|record|proto|alias)?)[-*/ \\t]|(library|module|struct|class)[*+ \\t]).*)$\\n\"\n+\t /*\n+\t  * For all Lisp dialects: a slightly indented line starting with \"(def\".\n+\t  */\n+\t \"^  ?(\\\\([Dd][Ee][Ff].*)$\",\n+\t /*\n+\t  * The union of R7RS and Common Lisp symbol syntax: allows arbitrary\n+\t  * strings between vertical bars, including any escaped characters.\n+\t  */\n+\t \"\\\\|([^|\\\\\\\\]|\\\\\\\\.)*\\\\|\"\n+\t /* All other words should be delimited by spaces or parentheses. */\n+\t \"|([^][)(}{ \\t])+\"),\n PATTERNS(\"tex\", \"^(\\\\\\\\((sub)*section|chapter|part)\\\\*{0,1}\\\\{.*)$\",\n \t \"\\\\\\\\[a-zA-Z@]+|\\\\\\\\.|([a-zA-Z0-9]|[^\\x01-\\x7f])+\"),\n { .name = \"default\", .binary = -1 },\n-- \ngitgitgadget\n"},{"id":"541613","messageId":"5880e69f-a617-426f-ab1b-12a27c763fcd@kdbg.org","threadId":"64483","inReplyTo":"pull.2000.v4.git.1776220063.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 0/2] userdiff: extend Scheme support to cover other Lisp dialects","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2026-04-15T06:54:21Z","receivedAt":"2026-04-15T06:54:32Z","isPatch":true,"body":"Am 15.04.26 um 04:27 schrieb Scott L. Burson via GitGitGadget:\n> Range-diff vs v3:\n> \n>  1:  e20ac5b6a6 = 1:  8e0b1e3d01 userdiff: tighten word-diff test case of the scheme driver\n>  2:  fb4c8dc5d4 ! 2:  0bd51e02ba userdiff: extend Scheme support to cover other Lisp dialects\n>      @@ Commit message\n>       \n>        ## Documentation/gitattributes.adoc ##\n>       @@ Documentation/gitattributes.adoc: patterns are available:\n>      + \n>        - `rust` suitable for source code in the Rust language.\n>        \n>      - - `scheme` suitable for source code in the Scheme language.\n>      -+Also handles Emacs Lisp, Common Lisp, and most other dialects.\n>      +-- `scheme` suitable for source code in the Scheme language.\n>      ++- `scheme` suitable for source code in most Lisp dialects,\n>      ++  including Scheme, Emacs Lisp, Common Lisp, and Clojure.\n>        \n>        - `tex` suitable for source code for LaTeX documents.\n>        \n>      @@ t/t4034/scheme/expect\n>          ; This is a <RED>really<RESET><GREEN>(moderately)<RESET> cool function.\n>          (<RED>this\\place<RESET><GREEN>that\\place<RESET> (+ 3 4))\n>       -  (define <RED>|the greeting|<RESET><GREEN>|a greeting|<RESET> \"hello\")\n>      -+  (define <RED>|the \\greeting|<RESET><GREEN>|a \\greeting|<RESET> |hello there|)\n>      ++  (define <RED>|the \\| \\greeting|<RESET><GREEN>|a \\greeting|<RESET> |hello there|)\n>          ({<RED>}<RESET>(([<RED>]<RESET>(func-n)<RED>[<RESET>]))<RED>{<RESET>})\n>          (let ((c (<RED>+ a b<RESET><GREEN>add1 first<RESET>)))\n>            (format \"one more than the total is %d\" (<RED>add1<RESET><GREEN>+<RESET> c <GREEN>second<RESET>))))\n>      @@ t/t4034/scheme/pre\n>          ; This is a really cool function.\n>          (this\\place (+ 3 4))\n>       -  (define |the greeting| \"hello\")\n>      -+  (define |the \\greeting| |hello there|)\n>      ++  (define |the \\| \\greeting| |hello there|)\n>          ({}(([](func-n)[])){})\n>          (let ((c (+ a b)))\n>            (format \"one more than the total is %d\" (add1 c))))\n> \n\nThis implements the conclusion of our discussion from three months ago.\n\nAcked-by: Johannes Sixt <j6t@kdbg.org>\n\n-- Hannes\n\n"}]}