git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: bug#60690: -P '\d' in GNU and git grep

From
Paul Eggert <eggert@cs.ucla.edu>
Date
Apr 5, 2023, 18:32 UTC
Message-ID
<6d86214a-1b80-eb88-1efb-36e61fd3203e@cs.ucla.edu>
In-Reply-To
<xmqqttxvzbo8.fsf@gitster.g>
On 2023-04-04 12:31, Junio C Hamano wrote:
> My personal inclination is to let Perl folks decide
> and follow them (even though I am skeptical about the wisdom of
> letting '\d' match anything other than [0-9])

I looked into what pcre2grep does. It has always done only 8-bit processing unless you use the -u or --utf option, so plain "pcre2grep '\d'" matches only ASCII digits.

Although this causes pcre2grep to mishandle Unicode characters:
   $ echo 'Ævar' | pcre2grep '[Ssß]'
   Ævar
it mimics Perl 5.36:
   $ echo 'Ævar' | perl -ne 'print $_ if /[Ssß]/'
   Ævar
so this seems to be what Perl users expect, despite its infelicities.

For better Unicode handling one can use pcre2grep's -u or --utf option, which causes pcre2grep to behave more like GNU grep -P and git grep -P: "echo 'Ævar' | pcre2grep -u '[Ssß]'" outputs nothing, which I think is what most people would expect (unless they're Perl users :-).

Neither git grep -P nor the current release of pcre2grep -u have \d matching non-ASCII digits, because they do not use PCRE2_UCP. However, in a February 8 commit[1], Philip Hazel changed pcre2grep to use PCRE2_UCP, so this will mean 10.43 pcre2grep -u will behave like 3.9 GNU grep -P did (though 3.10 has changed this).

That February commit also added a --no-ucp option, to disable PCRE2_UCP. So as I understand it, if you're in a UTF-8 locale:

* 10.43 pcre2grep -u will behave like 3.9 GNU grep -P.
* 10.43 pcre2grep -u --no-ucp will behave like git grep -P.
* Current GNU grep -P is different from everybody else.
This incompatibility is not good.

Here are two ways forward to fix this incompatibility (there are other possibilities of course):

(A) GNU grep adds a --no-ucp option that acts like 10.43 pcre2grep --no-ucp, and git grep -P follows suit. That is, both GNU and git grep act like 10.43 pcre2grep -u, in that they enable PCRE2_UTF, and also enable PCRE2_UCP unless --no-ucp is given. This would cause \d to match non-ASCII digits unless --no-ucp is given.

(B) GNU grep -P and git grep -P mimic pcre2grep in both -u and --no-ucp. That is, they would both do 8-bit-only by default, and use PCRE2_UTF only when -u or --utf is given, and use PCRE2_UCP only when --no-ucp is absent. This would cause \d to match non-ASCII digits only when -u is given but --no-ucp is not.

Under either (A) or (B), future pcre2grep -u, GNU grep -P, and git grep -P would be consistent.

I mildly prefer (B) but (A) would also work. (One advantage of (B) is that it should be faster....)

[1]: https://github.com/PCRE2Project/pcre2/commit/8385df8c97b6f8069a48e600c7e4e94cc3e3ebd9ht

Previous: Junio C HamanoNext: Paul Eggert
Message 8 of 19 in “-P '\d' in GNU and git grep”
  1. Paul EggertApr 3, 2023
  2. Jim MeyeringApr 4, 2023
  3. Paul EggertApr 4, 2023
  4. Jim MeyeringApr 4, 2023
  5. Carlo ArenasApr 4, 2023
  6. Paul EggertApr 4, 2023
  7. Junio C HamanoApr 4, 2023
  8. Paul EggertApr 5, 2023
  9. Paul EggertApr 5, 2023
  10. Junio C HamanoApr 5, 2023
  11. Jim MeyeringApr 5, 2023
  12. Paul EggertApr 5, 2023
  13. Carlo ArenasApr 5, 2023
  14. demerphqApr 6, 2023
  15. Paul EggertApr 7, 2023
  16. demerphqApr 6, 2023
  17. Paul EggertApr 7, 2023
  18. Carlo ArenasApr 8, 2023
  19. Paul EggertApr 8, 2023

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.