git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: grep: fix multibyte regex handling under macOS (1819ad327b7a1f19540a819813b70a0e8a7f798f)

From
D. Ben Knoble <ben.knoble@gmail.com>
Date
Feb 7, 2023, 22:27 UTC
Message-ID
<CALnO6CCzp0brmCjbNmOKe9E1bcxLzGNrqpfK_JrU=+LXt-DUyQ@mail.gmail.com>
In-Reply-To
<Y+KXC5b0qUg2/nxt@coredump.intra.peff.net>
CC'ing Jonathan Nieder
On Tue, Feb 7, 2023 at 1:23 PM Jeff King <peff@peff.net> wrote:
Show 12 quoted lines
>
> On Sun, Feb 05, 2023 at 02:51:05PM -0500, D. Ben Knoble wrote:
>
> > Any thoughts on some sort of stop-gap measure to fix --word-diff while
> > Git decides how to handle the regex engine incompatibilities? How
> > important is the sequence of bytes at the end of --word-diff regexes
> > in userdiff.c?
>
> It comes from 664d44ee7f (userdiff: simplify word-diff safeguard,
> 2011-01-11). So presumably we'd want to figure out a way to accomplish
> the same thing in a portable way. I'm not sure that's possible, though,
> without making assumptions about the regex engine.

If "use the safeguard portably" implies "make assumptions about the regex engine," that sounds like an argument for Git to ship its own engine with exactly the necessary features. If that implementation includes proper locale and UTF-8 support alongside support for the high-byte character classes, I think we would be all set…

OTOH, perhaps there is a way to express the safeguard character classes portably?

Jonathan, can you provide more context for the safeguard? I've read this message several times

Show 20 quoted lines
> git's diff-words support has a detail that can be a little dangerous:
> any text not matched by a given language's tokenization pattern is
> treated as whitespace and changes in such text would go unnoticed.
> Therefore each of the built-in regexes allows a special token type
> consisting of a single non-whitespace character [^[:space:]].
>
> To make sure UTF-8 sequences remain human readable, the builtin
> regexes also have a special token type for runs of bytes with the high
> bit set.  In English, non-ASCII characters are usually isolated so
> this is analogous to the [^[:space:]] pattern, except it matches a
> single _multibyte_ character despite use of the C locale.
>
> Unfortunately it is easy to make typos or forget entirely to include
> these catch-all token types when adding support for new languages (see
> v1.7.3.5~16, userdiff: fix typo in ruby and python word regexes,
> 2010-12-18).  Avoid this by including them automatically within the
> PATTERNS and IPATTERN macros.
>
> While at it, change the UTF-8 sequence token type to match exactly one
> non-ASCII multi-byte character, rather than an arbitrary run of them.
and I can hardly make heads or tails of it.
-- 
D. Ben Knoble
Previous: Jeff KingNext: Jeff King
Message 18 of 22 in “RE: grep: fix multibyte regex handling under macOS (1819ad327b7a1f19540a819813b70a0e8a7f798f)”
  1. D. Ben KnobleFeb 1, 2023
  2. demerphqFeb 1, 2023
  3. D. Ben KnobleFeb 1, 2023
  4. demerphqFeb 1, 2023
  5. Junio C HamanoFeb 1, 2023
  6. D. Ben KnobleFeb 1, 2023
  7. D. Ben KnobleFeb 1, 2023
  8. Junio C HamanoFeb 1, 2023
  9. Jeff KingFeb 1, 2023
  10. demerphqFeb 2, 2023
  11. D. Ben KnobleFeb 2, 2023
  12. Jeff KingFeb 3, 2023
  13. Ævar Arnfjörð BjarmasonFeb 3, 2023
  14. Jeff KingFeb 4, 2023
  15. demerphqFeb 4, 2023
  16. D. Ben KnobleFeb 5, 2023
  17. Jeff KingFeb 7, 2023
  18. D. Ben KnobleFeb 7, 2023
  19. Jeff KingFeb 7, 2023
  20. D. Ben KnobleFeb 2, 2023
  21. Jeff KingFeb 3, 2023
  22. D. Ben KnobleFeb 3, 2023

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.