git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: bug#60690: [PATCH v2] grep: correctly identify utf-8 characters with \{b, w} in -P

From
Paul Eggert <eggert@cs.ucla.edu>
Date
Jan 9, 2023, 23:12 UTC
Message-ID
<80b42740-c85b-cef2-622c-c5b2450e264c@cs.ucla.edu>
In-Reply-To
<230109.865ydf1mdu.gmgdl@evledraar.gmail.com>
On 1/9/23 11:51, Ævar Arnfjörð Bjarmason wrote:
Show 15 quoted lines
> 	/b:
> 	155781
> 	(*UCP)/b:
> 	46035
> 	/s:
> 	0
> 	(*UCP)/s:
> 	0
> 	/w:
> 	142468
> 	(*UCP)/w:
> 	9706
> 
> So the output still differs, and some of those differences may or may
> not be wanted.

I took a look at the output, and by and large I'd want the differences; that is, I'd want the UCP version, which generates less output. This is because several Emacs source files are not UTF-8, and \b has nonsense matches when searching text files encoded via Shift-JIS or Big 5 or whatever. For this sort of thing, the fewer matches the better.

> If all you're doing is matching either ASCII or Japanese text and you
> want "locale-aware numbers" it might do the wrong thing.

I'm not seeing much of a problem here. When searching Japanese text, I would expect \d and [0-90-9] (using both ASCII and full-width digits) to be equivalent so (assuming UCP) it's not a big deal as to which regex you use, since Japanese text won't contain Bengali (or whatever) digits. And when searching binary data, I'd expect a bunch of garbage no matter how \d is interpreted.

Here I'm assuming [0-9] (using full-width digits) has the expected meaning in PCRE2, i.e., that PCRE2 didn't make the same mistake that POSIX made.

Previous: Ævar Arnfjörð BjarmasonNext: Carlo Arenas
Message 7 of 17 in “grep: correctly identify utf-8 characters with \{b,w} in -P”
  1. grep: correctly identify utf-8 characters with \{b,w} in -PCarlo Marcelo Arenas Belón, Jan 8, 2023
  2. Junio C HamanoJan 8, 2023
  3. grep: correctly identify utf-8 characters with \{b,w} in -PCarlo Marcelo Arenas Belón, Jan 8, 2023
  4. Ævar Arnfjörð BjarmasonJan 9, 2023
  5. Paul EggertJan 9, 2023
  6. Ævar Arnfjörð BjarmasonJan 9, 2023
  7. Paul EggertJan 9, 2023
  8. Carlo ArenasJan 10, 2023
  9. Junio C HamanoJan 16, 2023
  10. grep: correctly identify utf-8 characters with \{b,w} in -PCarlo Marcelo Arenas Belón, Jan 17, 2023
  11. Ævar Arnfjörð BjarmasonJan 17, 2023
  12. Junio C HamanoJan 17, 2023
  13. Carlo ArenasJan 18, 2023
  14. Ævar Arnfjörð BjarmasonJan 18, 2023
  15. Junio C HamanoJan 18, 2023
  16. Ævar Arnfjörð BjarmasonJan 18, 2023
  17. Junio C HamanoJan 18, 2023

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.