git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: bug#60690: -P '\d' in GNU and git grep

From
Paul Eggert <eggert@cs.ucla.edu>
Date
Apr 7, 2023, 19:00 UTC
Message-ID
<065bcdcb-5770-5384-5afe-4a4d29272274@cs.ucla.edu>
In-Reply-To
<CANgJU+U+xXsh9psd0z5Xjr+Se5QgdKkjQ7LUQ-PdUULSN3n4+g@mail.gmail.com>
On 2023-04-06 06:39, demerphq wrote:
> Unicode specifies that \d match any digit
> in any script that it supports.

"Specifies" is too strong. The Unicode Regular Expressions technical standard (UTS#18) mentions \d only in Annex C[1], next to the word "digit" in a column labeled "Property" (even though \d is really syntax not a property). This is at best an informal recommendation, not a requirement, as UTS#18 0.2[2] says that UTS#18's syntax is only for illustration and that although it's similar to Perl's, the two syntax forms may not be exactly the same. So we can't look to UTS#18 for a definitive way out of the \d mess, as the Unicode folks specifically delegated matters to us.

Even ignoring the \d issue the digit situation is messy. UTS#18 Annex C says "\p{gc=Decimal_Number}" is the standard recommended syntax assignment for digits. However, PCRE2 does not support this syntax; it supports another variant \p{Nd} that UTS#18 also recommends. So it appears that PCRE2 already does not implement every recommended aspect of UTS#18 syntax. PCRE2 also doesn't match Perl, which does support "\p{gc=Decimal_Number}".

Anyway, since grep -P '\p{Nd}' implements Unicode's decimal digit class, that's clearly enough for grep -P to conform to UTS#18 with respect to digits.

> A) how do you tell the regular expression
> engine what semantics you want and B) how does the regular expression
> library identify the encoding in the file, and how does it handle
> malformed content in that file.
Here's how GNU grep does it:
* RE semantics are specified via command-line options like -P.
* Text encoding is specified by locale, e.g., LC_ALL='en_US.utf8'.
* REs do not match encoding errors.
> on *nix there is no tradition of using BOM's to
> distinguish the 6 different possible encodings of Unicode (UTF-8,
> UTF-EBCDIC, UTF-16LE, UTF-16BE, UTF-32LE, UTF-32BE)

Yes, GNU/Linux never really experienced the joys of UTF-EBCDIC, Oracle UTFE, UTF-16LE vs UTF-16BE etc. If you're running legacy IBM mainframe or MS-Windows code these legacy encodings are obviously a big deal. However, there seems little reason to force their nontrivial hassles onto every GNU/Linux program that processes text. A few specialized apps like 'iconv' deal with offbeat encodings, and that is probably a better approach all around.

> there seems
> to be some level of desire of matching with unicode semantics against
> files that are not uniformly encoded in one of these formats.
That is a use case, yes. It's what 'strings' and 'grep' do.

[1]: https://unicode.org/reports/tr18/#Compatibility_Properties [2]: https://unicode.org/reports/tr18/#Conformance

Previous: demerphqNext: Carlo Arenas
Message 17 of 19 in “-P '\d' in GNU and git grep”
  1. Paul EggertApr 3, 2023
  2. Jim MeyeringApr 4, 2023
  3. Paul EggertApr 4, 2023
  4. Jim MeyeringApr 4, 2023
  5. Carlo ArenasApr 4, 2023
  6. Paul EggertApr 4, 2023
  7. Junio C HamanoApr 4, 2023
  8. Paul EggertApr 5, 2023
  9. Paul EggertApr 5, 2023
  10. Junio C HamanoApr 5, 2023
  11. Jim MeyeringApr 5, 2023
  12. Paul EggertApr 5, 2023
  13. Carlo ArenasApr 5, 2023
  14. demerphqApr 6, 2023
  15. Paul EggertApr 7, 2023
  16. demerphqApr 6, 2023
  17. Paul EggertApr 7, 2023
  18. Carlo ArenasApr 8, 2023
  19. Paul EggertApr 8, 2023

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.