{"thread":{"id":"59531","subject":"-P '\\d' in GNU and git grep","startedAt":"2023-04-03T21:47:46Z","lastAt":"2023-04-08T22:45:26Z","messageCount":19,"participants":["Paul Eggert","Jim Meyering","Carlo Arenas","Junio C Hamano","demerphq"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"474716","messageId":"2554712d-e386-3bab-bc6c-1f0e85d999db@cs.ucla.edu","threadId":"59531","inReplyTo":null,"subject":"-P '\\d' in GNU and git grep","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2023-04-03T21:38:42Z","receivedAt":"2023-04-03T21:47:46Z","isPatch":false,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"I've recently done some bug-report maintenance about a set of GNU grep \nbug reports related to whether whether \"grep -P '\\d'\" should match \nnon-ASCII digits, and have some thoughts about coordinating GNU grep \nwith git grep in this department.\n\nGNU Bug#62605[1] \"`[\\d]` does not work with PCRE\" has been fixed on \nSavannah's copy of GNU grep, and some sort of fix should appear in the \nnext grep release. However, I'm leaving the GNU grep bug report open for \nnow because it's related to Bug#60690[2] \"[PATCH v2] grep: correctly \nidentify utf-8 characters with \\{b,w} in -P\" and to Bug#62552[3] \"Bug \nfound in latest stable release v3.10 of grep\". I merged these related \nbug reports, and the oldest one, Bug#60690, is now the representative \ndisplayed in the GNU grep bug list[4].\n\nFor this set of grep bug reports there's still a pending issue discussed \nin my recent email[5], which proposes a patch so I've tagged Bug#60690 \nwith \"patch\". The proposal is that GNU grep -P '\\d' should revert to the \ngrep 3.9 behavior, i.e., that in a UTF-8 locale, \\d should also match \nnon-ASCII decimal digits.\n\nIn researching this a bit further, I found that on March 23 Git disabled \nthe use of PCRE2_UCP in PCRE2 10.34 or earlier[6], due to a PCRE2 bug \nthat can cause a crash when PCRE2_UCP is used[7]. A bug fix[8] should \nappear in the next PCRE2 release.\n\nWhen PCRE2 10.35 comes out, it appears that 'git grep -P' will behave \nlike 'grep -P' only if GNU grep adopts something like the solution \nproposed in [5].\n\n[1]: https://bugs.gnu.org/62605\n[2]: https://bugs.gnu.org/60690\n[3]: https://bugs.gnu.org/62552\n[4]: https://debbugs.gnu.org/cgi/pkgreport.cgi?package=grep\n[5]: https://lists.gnu.org/archive/html/grep-devel/2023-04/msg00004.html\n[6]: \nhttps://github.com/git/git/commit/14b9a044798ebb3858a1f1a1377309a3d6054ac8\n[7]: \nhttps://lore.kernel.org/git/7E83DAA1-F9A9-4151-8D07-D80EA6D59EEA@clumio.com/\n[8]: \nhttps://github.com/git/git/commit/14b9a044798ebb3858a1f1a1377309a3d6054ac8\n"},{"id":"474761","messageId":"CA+8g5KHuE-kQqmi9cVjeJbpyt54v9m9omh9A9we1zmR0+aTDHg@mail.gmail.com","threadId":"59531","inReplyTo":"2554712d-e386-3bab-bc6c-1f0e85d999db@cs.ucla.edu","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Jim Meyering","fromEmail":"jim@meyering.net","sentAt":"2023-04-04T03:30:02Z","receivedAt":"2023-04-04T03:30:26Z","isPatch":false,"sender":{"key":"jim@meyering.net","avatar":"https://avatars.githubusercontent.com/u/710630?v=4"},"body":"On Mon, Apr 3, 2023 at 2:39 PM Paul Eggert <eggert@cs.ucla.edu> wrote:\n> I've recently done some bug-report maintenance about a set of GNU grep\n> bug reports related to whether whether \"grep -P '\\d'\" should match\n> non-ASCII digits, and have some thoughts about coordinating GNU grep\n> with git grep in this department.\n>\n> GNU Bug#62605[1] \"`[\\d]` does not work with PCRE\" has been fixed on\n> Savannah's copy of GNU grep, and some sort of fix should appear in the\n> next grep release. However, I'm leaving the GNU grep bug report open for\n> now because it's related to Bug#60690[2] \"[PATCH v2] grep: correctly\n> identify utf-8 characters with \\{b,w} in -P\" and to Bug#62552[3] \"Bug\n> found in latest stable release v3.10 of grep\". I merged these related\n> bug reports, and the oldest one, Bug#60690, is now the representative\n> displayed in the GNU grep bug list[4].\n>\n> For this set of grep bug reports there's still a pending issue discussed\n> in my recent email[5], which proposes a patch so I've tagged Bug#60690\n> with \"patch\". The proposal is that GNU grep -P '\\d' should revert to the\n> grep 3.9 behavior, i.e., that in a UTF-8 locale, \\d should also match\n> non-ASCII decimal digits.\n>\n> In researching this a bit further, I found that on March 23 Git disabled\n> the use of PCRE2_UCP in PCRE2 10.34 or earlier[6], due to a PCRE2 bug\n> that can cause a crash when PCRE2_UCP is used[7]. A bug fix[8] should\n> appear in the next PCRE2 release.\n>\n> When PCRE2 10.35 comes out,\n\nThanks for finding that.\nIt's clearly a good idea to disable PCRE2_UCP for those using those\nolder, known-buggy versions of pcre2.\n\nThe latest is 10.42, per https://github.com/PCRE2Project/pcre2/releases\n\n> it appears that 'git grep -P' will behave\n> like 'grep -P' only if GNU grep adopts something like the solution\n> proposed in [5].\n>\n> [1]: https://bugs.gnu.org/62605\n> [2]: https://bugs.gnu.org/60690\n> [3]: https://bugs.gnu.org/62552\n> [4]: https://debbugs.gnu.org/cgi/pkgreport.cgi?package=grep\n> [5]: https://lists.gnu.org/archive/html/grep-devel/2023-04/msg00004.html\n> [6]:\n> https://github.com/git/git/commit/14b9a044798ebb3858a1f1a1377309a3d6054ac8\n> [7]:\n> https://lore.kernel.org/git/7E83DAA1-F9A9-4151-8D07-D80EA6D59EEA@clumio.com/\n> [8]:\n> https://github.com/git/git/commit/14b9a044798ebb3858a1f1a1377309a3d6054ac8\n\nThanks for all of the links. However, have you seen justification\n(other than for compatibility with some other tool or language) for\nallowing \\d to match non-ASCII by default, in spite of the risks?\nIMHO, we have an obligation to retain compatibility with how grep -P\n'\\d' has worked since -P was added. I'd be happy to see an option to\nenable the match-multibyte-digits behavior, but making it the default\nseems too likely to introduce unwarranted risk.\n"},{"id":"474763","messageId":"920dcc8d-9e45-a03e-af06-6b420c6e0f81@cs.ucla.edu","threadId":"59531","inReplyTo":"CA+8g5KHuE-kQqmi9cVjeJbpyt54v9m9omh9A9we1zmR0+aTDHg@mail.gmail.com","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2023-04-04T06:46:57Z","receivedAt":"2023-04-04T06:47:05Z","isPatch":false,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"On 2023-04-03 20:30, Jim Meyering wrote:\n> have you seen justification\n> (other than for compatibility with some other tool or language) for\n> allowing \\d to match non-ASCII by default, in spite of the risks?\n\nIn the example Ævar supplied in <https://bugs.gnu.org/60690>, my \nimpression was that it was better when \\d matched non-ASCII digits. That \nis, in a UTF-8 locale it's better when \\d finds matches in these lines:\n\n>> \t> git-gui/po/ja.po:\"- 第１行: 何をしたか、を１行で要約。\\n\"\n>> \t> git-gui/po/ja.po:\"- 第２行: 空白\\n\"\n\nbecause they contain the Japanese digits \"１\" and \"２\". This was the only \nexample I recall being given.\n\nAlso, I find it odd that grep -P '^[\\w\\d]*$' matches lines containing \nany sort of Arabic word characters, but it rejects lines containing \nArabic digits like \"٣\" that are perfectly reasonable in Arabic-language \ntext. I also find it odd that [\\d] and [[:digit:]] mean different things.\n\nThere are arguments on the other side, otherwise we wouldn't be having \nthis discussion. And it's true that grep -P '\\d' formerly rejected \nArabic digits (though it's also true that grep -P '\\w' formerly rejected \nArabic letters...). Still, the cure's oddness and incompatibility with \nGit, Perl, etc. appears to me to be worse than the disease of dealing \nwith grep -P invocations that need to use [0-9] or LC_ALL=\"C\" anyway if \nthey want to be portable to any program other than GNU grep.\n"},{"id":"474764","messageId":"CAPUEspj1m6F0_XgOFUVaq3Aq_Ah3PzCUs7YUyFH9_Zz-MOYTTA@mail.gmail.com","threadId":"59531","inReplyTo":"2554712d-e386-3bab-bc6c-1f0e85d999db@cs.ucla.edu","subject":"Re: -P '\\d' in GNU and git grep","fromName":"Carlo Arenas","fromEmail":"carenas@gmail.com","sentAt":"2023-04-04T06:56:54Z","receivedAt":"2023-04-04T06:57:11Z","isPatch":false,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"On Mon, Apr 3, 2023 at 2:38 PM Paul Eggert <eggert@cs.ucla.edu> wrote:\n>\n> In researching this a bit further, I found that on March 23 Git disabled\n> the use of PCRE2_UCP in PCRE2 10.34 or earlier[6], due to a PCRE2 bug\n> that can cause a crash when PCRE2_UCP is used[7]. A bug fix[8] should\n> appear in the next PCRE2 release.\n\nPresume PCRE2 is a typo and should have been \"git\" here?\n\nFWIW the PCRE2 fix[1] has been released already with 10.35 and\nbackporting to the Ubuntu 20.04 package that crashed in the original\nreport would also solve the crash with 10.34.\n\nCarlo\n\n[1] https://github.com/PCRE2Project/pcre2/commit/c21bd977547d\n"},{"id":"474780","messageId":"CA+8g5KGGnsf0xMCXO28R1m8-z76=kG_AiYRh6=OgRL+x5C1yqQ@mail.gmail.com","threadId":"59531","inReplyTo":"920dcc8d-9e45-a03e-af06-6b420c6e0f81@cs.ucla.edu","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Jim Meyering","fromEmail":"jim@meyering.net","sentAt":"2023-04-04T15:31:44Z","receivedAt":"2023-04-04T15:32:02Z","isPatch":false,"sender":{"key":"jim@meyering.net","avatar":"https://avatars.githubusercontent.com/u/710630?v=4"},"body":"On Mon, Apr 3, 2023 at 11:47 PM Paul Eggert <eggert@cs.ucla.edu> wrote:\n> On 2023-04-03 20:30, Jim Meyering wrote:\n> > have you seen justification\n> > (other than for compatibility with some other tool or language) for\n> > allowing \\d to match non-ASCII by default, in spite of the risks?\n>\n> In the example Ævar supplied in <https://bugs.gnu.org/60690>, my\n> impression was that it was better when \\d matched non-ASCII digits. That\n> is, in a UTF-8 locale it's better when \\d finds matches in these lines:\n>\n> >>      > git-gui/po/ja.po:\"- 第１行: 何をしたか、を１行で要約。\\n\"\n> >>      > git-gui/po/ja.po:\"- 第２行: 空白\\n\"\n>\n> because they contain the Japanese digits \"１\" and \"２\". This was the only\n> example I recall being given.\n\nBefore it was unintentionally enabled in grep-3.9, lines like that have\nnever been matched by grep -P's '\\d'. By relaxing \\d, we'd weaken\nany application that uses say grep -P '^\\d+$' to perform input\nvalidation intending to ensure that some input is all ASCII digits.\nIt's not a big stretch to imagine that some downstream processor\nof that \"verified\" data is not prepared to deal with multi-byte digits.\n\n> Also, I find it odd that grep -P '^[\\w\\d]*$' matches lines containing\n> any sort of Arabic word characters, but it rejects lines containing\n> Arabic digits like \"٣\" that are perfectly reasonable in Arabic-language\n> text. I also find it odd that [\\d] and [[:digit:]] mean different things.\n>\n> There are arguments on the other side, otherwise we wouldn't be having\n> this discussion. And it's true that grep -P '\\d' formerly rejected\n> Arabic digits (though it's also true that grep -P '\\w' formerly rejected\n> Arabic letters...). Still, the cure's oddness and incompatibility with\n> Git, Perl, etc. appears to me to be worse than the disease of dealing\n> with grep -P invocations that need to use [0-9] or LC_ALL=\"C\" anyway if\n> they want to be portable to any program other than GNU grep.\n\nI'm primarily concerned about not introducing a persistent regression in\nhow GNU grep's -P '\\d' works in multibyte locales. The corner cases you\nmention do matter, of course, but are far less likely to matter in practice.\n"},{"id":"474795","messageId":"96358c4e-7200-e5a5-869e-5da9d0de3503@cs.ucla.edu","threadId":"59531","inReplyTo":"CAPUEspj1m6F0_XgOFUVaq3Aq_Ah3PzCUs7YUyFH9_Zz-MOYTTA@mail.gmail.com","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2023-04-04T18:25:59Z","receivedAt":"2023-04-04T18:29:33Z","isPatch":false,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"On 4/3/23 23:56, Carlo Arenas wrote:\n> On Mon, Apr 3, 2023 at 2:38 PM Paul Eggert <eggert@cs.ucla.edu> wrote:\n>>\n>> on March 23 Git disabled\n>> the use of PCRE2_UCP in PCRE2 10.34 or earlier[6], due to a PCRE2 bug\n>> that can cause a crash when PCRE2_UCP is used[7]. A bug fix[8] should\n>> appear in the next PCRE2 release.\n> \n> Presume PCRE2 is a typo and should have been \"git\" here?\n\nNo, I was talking about what options Git uses when it calls PCRE2 \nfunctions. In other words, this is about whether GNU 'grep -P' should be \ncompatible with 'git grep -P' (as well as with Perl and with pcregrep), \nwhen interpreting \\d and similar constructs.\n\nThis is an evolving area. Git master is fiddling with flags and options, \nand so is GNU grep master, and so is PCRE2, and there are bugs. If \nyou're running bleeding-edge versions of this code you'll get different \nbehavior than if you're running grep 3.8, pcregrep 8.45, Perl 5.36, and \ngit 2.39.2 (which is what Fedora 37 has).\n\nWhat I'm fearing is that we may evolve into mutually incompatible \ninterpretations of how Perl regular expressions deal with UTF-8 text. \nThat'd be a recipe for confusion down the road.\n"},{"id":"474799","messageId":"xmqqttxvzbo8.fsf@gitster.g","threadId":"59531","inReplyTo":"96358c4e-7200-e5a5-869e-5da9d0de3503@cs.ucla.edu","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-04-04T19:31:51Z","receivedAt":"2023-04-04T19:31:59Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Paul Eggert <eggert@cs.ucla.edu> writes:\n\n> This is an evolving area. Git master is fiddling with flags and\n> options, and so is GNU grep master, and so is PCRE2, and there are\n> bugs. If you're running bleeding-edge versions of this code you'll get\n> different behavior than if you're running grep 3.8, pcregrep 8.45,\n> Perl 5.36, and git 2.39.2 (which is what Fedora 37 has).\n>\n> What I'm fearing is that we may evolve into mutually incompatible\n> interpretations of how Perl regular expressions deal with UTF-8\n> text. That'd be a recipe for confusion down the road.\n\nNicely said.  My personal inclination is to let Perl folks decide\nand follow them (even though I am skeptical about the wisdom of\nletting '\\d' match anything other than [0-9]), but even in Git\ncircle there would be different opinions, so I am glad that the\ndiscussion is visible on the list to those who are intrested.\n\n"},{"id":"474873","messageId":"6d86214a-1b80-eb88-1efb-36e61fd3203e@cs.ucla.edu","threadId":"59531","inReplyTo":"xmqqttxvzbo8.fsf@gitster.g","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2023-04-05T18:32:38Z","receivedAt":"2023-04-05T18:32:46Z","isPatch":false,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"On 2023-04-04 12:31, Junio C Hamano wrote:\n\n> My personal inclination is to let Perl folks decide\n> and follow them (even though I am skeptical about the wisdom of\n> letting '\\d' match anything other than [0-9])\n\nI looked into what pcre2grep does. It has always done only 8-bit \nprocessing unless you use the -u or --utf option, so plain \"pcre2grep \n'\\d'\" matches only ASCII digits.\n\nAlthough this causes pcre2grep to mishandle Unicode characters:\n\n   $ echo 'Ævar' | pcre2grep '[Ssß]'\n   Ævar\n\nit mimics Perl 5.36:\n\n   $ echo 'Ævar' | perl -ne 'print $_ if /[Ssß]/'\n   Ævar\n\nso this seems to be what Perl users expect, despite its infelicities.\n\nFor better Unicode handling one can use pcre2grep's -u or --utf option, \nwhich causes pcre2grep to behave more like GNU grep -P and git grep -P: \n\"echo 'Ævar' | pcre2grep -u '[Ssß]'\" outputs nothing, which I think is \nwhat most people would expect (unless they're Perl users :-).\n\nNeither git grep -P nor the current release of pcre2grep -u have \\d \nmatching non-ASCII digits, because they do not use PCRE2_UCP. However, \nin a February 8 commit[1], Philip Hazel changed pcre2grep to use \nPCRE2_UCP, so this will mean 10.43 pcre2grep -u will behave like 3.9 GNU \ngrep -P did (though 3.10 has changed this).\n\nThat February commit also added a --no-ucp option, to disable PCRE2_UCP. \nSo as I understand it, if you're in a UTF-8 locale:\n\n* 10.43 pcre2grep -u will behave like 3.9 GNU grep -P.\n\n* 10.43 pcre2grep -u --no-ucp will behave like git grep -P.\n\n* Current GNU grep -P is different from everybody else.\n\nThis incompatibility is not good.\n\nHere are two ways forward to fix this incompatibility (there are other \npossibilities of course):\n\n(A) GNU grep adds a --no-ucp option that acts like 10.43 pcre2grep \n--no-ucp, and git grep -P follows suit. That is, both GNU and git grep \nact like 10.43 pcre2grep -u, in that they enable PCRE2_UTF, and also \nenable PCRE2_UCP unless --no-ucp is given. This would cause \\d to match \nnon-ASCII digits unless --no-ucp is given.\n\n(B) GNU grep -P and git grep -P mimic pcre2grep in both -u and --no-ucp. \nThat is, they would both do 8-bit-only by default, and use PCRE2_UTF \nonly when -u or --utf is given, and use PCRE2_UCP only when --no-ucp is \nabsent. This would cause \\d to match non-ASCII digits only when -u is \ngiven but --no-ucp is not.\n\nUnder either (A) or (B), future pcre2grep -u, GNU grep -P, and git grep \n-P would be consistent.\n\nI mildly prefer (B) but (A) would also work. (One advantage of (B) is \nthat it should be faster....)\n\n[1]: \nhttps://github.com/PCRE2Project/pcre2/commit/8385df8c97b6f8069a48e600c7e4e94cc3e3ebd9ht\n"},{"id":"474875","messageId":"33b3eb15-73e2-8004-9f06-19e5ec5c5877@cs.ucla.edu","threadId":"59531","inReplyTo":"6d86214a-1b80-eb88-1efb-36e61fd3203e@cs.ucla.edu","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2023-04-05T19:04:28Z","receivedAt":"2023-04-05T19:04:46Z","isPatch":false,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"On 2023-04-05 11:32, Paul Eggert wrote:\n\n> in a February 8 commit[1], Philip Hazel changed pcre2grep to use \n> PCRE2_UCP, so this will mean 10.43 pcre2grep -u will behave like 3.9 GNU \n> grep -P did (though 3.10 has changed this).\n\nSorry, due to fumblefingers I gave the wrong URL for [1]. Here's a \ncorrected URL:\n\nhttps://github.com/PCRE2Project/pcre2/commit/8385df8c97b6f8069a48e600c7e4e94cc3e3ebd9\n\nIt also mentions a new --case-restrict option, intended for 10.43 \npcre2grep. Given Perl's and PCRE2's plethora of options I suppose one \ncould imagine several other options of that ilk.\n"},{"id":"474879","messageId":"xmqqlej6unle.fsf@gitster.g","threadId":"59531","inReplyTo":"6d86214a-1b80-eb88-1efb-36e61fd3203e@cs.ucla.edu","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2023-04-05T19:37:49Z","receivedAt":"2023-04-05T19:38:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Paul Eggert <eggert@cs.ucla.edu> writes:\n\n> Here are two ways forward to fix this incompatibility (there are other\n> possibilities of course):\n>\n> (A) GNU grep adds a --no-ucp option that acts like 10.43 pcre2grep\n> --no-ucp, and git grep -P follows suit. That is, both GNU and git grep\n> act like 10.43 pcre2grep -u, in that they enable PCRE2_UTF, and also\n> enable PCRE2_UCP unless --no-ucp is given. This would cause \\d to\n> match non-ASCII digits unless --no-ucp is given.\n>\n> (B) GNU grep -P and git grep -P mimic pcre2grep in both -u and\n> --no-ucp. That is, they would both do 8-bit-only by default, and use\n> PCRE2_UTF only when -u or --utf is given, and use PCRE2_UCP only when\n> --no-ucp is absent. This would cause \\d to match non-ASCII digits only\n> when -u is given but --no-ucp is not.\n>\n> Under either (A) or (B), future pcre2grep -u, GNU grep -P, and git\n> grep -P would be consistent.\n>\n> I mildly prefer (B) but (A) would also work. (One advantage of (B) is\n> that it should be faster....)\n\nFor \"git grep -P\", I would like to hear from Carlo and Ævar; I agree\nboth (A) and (B) would be workable solutions, and have a slight\npreference on a solution that does not add more options that take\nonly in effect when -P is given, simply because these options are\ncumbersome to document and explain, but that is a very minor point.\n\nThanks.\n"},{"id":"474880","messageId":"CA+8g5KHYqgAZPpTOXWekDpWv-mvj-rBkGu+4MXy4OB1VDeS4Lw@mail.gmail.com","threadId":"59531","inReplyTo":"6d86214a-1b80-eb88-1efb-36e61fd3203e@cs.ucla.edu","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Jim Meyering","fromEmail":"jim@meyering.net","sentAt":"2023-04-05T19:40:18Z","receivedAt":"2023-04-05T19:41:26Z","isPatch":false,"sender":{"key":"jim@meyering.net","avatar":"https://avatars.githubusercontent.com/u/710630?v=4"},"body":"On Wed, Apr 5, 2023 at 11:33 AM Paul Eggert <eggert@cs.ucla.edu> wrote:\n> On 2023-04-04 12:31, Junio C Hamano wrote:\n> > My personal inclination is to let Perl folks decide\n> > and follow them (even though I am skeptical about the wisdom of\n> > letting '\\d' match anything other than [0-9])\n>\n> I looked into what pcre2grep does. It has always done only 8-bit\n> processing unless you use the -u or --utf option, so plain \"pcre2grep\n> '\\d'\" matches only ASCII digits.\n>\n> Although this causes pcre2grep to mishandle Unicode characters:\n>\n>    $ echo 'Ævar' | pcre2grep '[Ssß]'\n>    Ævar\n>\n> it mimics Perl 5.36:\n>\n>    $ echo 'Ævar' | perl -ne 'print $_ if /[Ssß]/'\n>    Ævar\n>\n> so this seems to be what Perl users expect, despite its infelicities.\n>\n> For better Unicode handling one can use pcre2grep's -u or --utf option,\n> which causes pcre2grep to behave more like GNU grep -P and git grep -P:\n> \"echo 'Ævar' | pcre2grep -u '[Ssß]'\" outputs nothing, which I think is\n> what most people would expect (unless they're Perl users :-).\n\nGood argument for making PCRE2_UCP the default.\n\n> Neither git grep -P nor the current release of pcre2grep -u have \\d\n> matching non-ASCII digits, because they do not use PCRE2_UCP. However,\n> in a February 8 commit[1], Philip Hazel changed pcre2grep to use\n> PCRE2_UCP, so this will mean 10.43 pcre2grep -u will behave like 3.9 GNU\n> grep -P did (though 3.10 has changed this).\n>\n> That February commit also added a --no-ucp option, to disable PCRE2_UCP.\n> So as I understand it, if you're in a UTF-8 locale:\n>\n> * 10.43 pcre2grep -u will behave like 3.9 GNU grep -P.\n>\n> * 10.43 pcre2grep -u --no-ucp will behave like git grep -P.\n>\n> * Current GNU grep -P is different from everybody else.\n>\n> This incompatibility is not good.\n>\n> Here are two ways forward to fix this incompatibility (there are other\n> possibilities of course):\n>\n> (A) GNU grep adds a --no-ucp option that acts like 10.43 pcre2grep\n> --no-ucp, and git grep -P follows suit. That is, both GNU and git grep\n> act like 10.43 pcre2grep -u, in that they enable PCRE2_UTF, and also\n> enable PCRE2_UCP unless --no-ucp is given. This would cause \\d to match\n> non-ASCII digits unless --no-ucp is given.\n>\n> (B) GNU grep -P and git grep -P mimic pcre2grep in both -u and --no-ucp.\n> That is, they would both do 8-bit-only by default, and use PCRE2_UTF\n> only when -u or --utf is given, and use PCRE2_UCP only when --no-ucp is\n> absent. This would cause \\d to match non-ASCII digits only when -u is\n> given but --no-ucp is not.\n\nChanging grep -P's \\d to match multibyte digits by default would break\nan important contract. Avoiding that feels like it must outweigh any\ncross-tool portability concern.\n\n(C)  preserve grep -P's tradition of \\d matching only 0..9, and once\ngrep uses 10.43 or newer, \\b and \\w will also work as desired.\n\n> Under either (A) or (B), future pcre2grep -u, GNU grep -P, and git grep\n> -P would be consistent.\n\nI hope git grep -P's \\d will also stick to ASCII-only by default.\nThose rare few who desire multibyte matches can always specify \\p{Nd}\ninstead of \\d, or (with new enough PCRE2), use (?-aD) and (?aD) to\ntoggle the digit-matching mode.\n"},{"id":"474884","messageId":"ed237a07-2f77-74eb-2f52-49b9b8f08873@cs.ucla.edu","threadId":"59531","inReplyTo":"CA+8g5KHYqgAZPpTOXWekDpWv-mvj-rBkGu+4MXy4OB1VDeS4Lw@mail.gmail.com","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2023-04-05T20:03:51Z","receivedAt":"2023-04-05T20:03:58Z","isPatch":false,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"On 2023-04-05 12:40, Jim Meyering wrote:\n> (C)  preserve grep -P's tradition of \\d matching only 0..9, and once\n> grep uses 10.43 or newer, \\b and \\w will also work as desired.\n\nIf I understand you correctly, (C) would mean that GNU grep -P, git grep \n-P, and pcre2grep -u would all use PCRE2_UTF | PCRE2_UCP, and would also \nuse the extra option PCRE2_EXTRA_ASCII_BSD that is planned for 10.43 PCRE2.\n\nThis would require changes to bleeding-edge pcre2grep -u (since it would \nneed to add PCRE2_EXTRA_ASCII_BSD unless --no-ucp is also given), and to \ngit grep -P (which would need to add PCRE2_UCP and \nPCRE2_EXTRA_ASCII_BSD, when libpcre2 is new enough to #define \nPCRE2_EXTRA_ASCII_BSD).\n\nThis option works for me as well. In fact it's the least work for me \nsince I already implemented it in bleeding-edge GNU grep (so it works \nthis way already :-).\n\n"},{"id":"474887","messageId":"CAPUEspjM6PtsY9LiK9Lqb2+H2UrWEfPziWVrOPwZGVpVbx7aJQ@mail.gmail.com","threadId":"59531","inReplyTo":"CA+8g5KHYqgAZPpTOXWekDpWv-mvj-rBkGu+4MXy4OB1VDeS4Lw@mail.gmail.com","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Carlo Arenas","fromEmail":"carenas@gmail.com","sentAt":"2023-04-05T21:20:51Z","receivedAt":"2023-04-05T21:21:08Z","isPatch":false,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"On Wed, Apr 5, 2023 at 12:40 PM Jim Meyering <jim@meyering.net> wrote:\n>\n> Changing grep -P's \\d to match multibyte digits by default would break\n> an important contract.\n\nWhile I tend to agree[1] (and indeed that is why PCRE2_EXTRA_ASCII_BSD\nwas invented), it would be also important to note that it goes against\nthe Unicode recommendation[2] and it is actually not true already[3]\nfor Python, .NET or Rust (which means ripgrep behaves like GNU grep -P\n3.9).\n\nFWIW I also agree that (at least `git grep -P`) should use\nPCRE2_EXTRA_ASCII_BSD by default as that is what makes more sense in\nthe context of matching source code and using instead `\\P{Nd}` if you\nreally want all Unicode digits is not much of a burden, but I am also\nnot sure if that makes sense in other contexts, specially considering\nthat I am obviously biased since the languages I mostly interact with\nONLY use arabic numerals and therefore `\\d` meaning `[0-9]` seems\n\"normal\".\n\nCarlo\n\nCC: changed to the real email address for PCRE2 development, for full\ncontext on this thread use [4]\n\n[1] https://github.com/PCRE2Project/pcre2/pull/186\n[2] https://unicode.org/reports/tr18/\n[3] https://regex101.com/r/S5RW4c/1\n[4] https://lore.kernel.org/git/230109.86v8lf297g.gmgdl@evledraar.gmail.com/T/\n"},{"id":"474933","messageId":"CANgJU+U+xXsh9psd0z5Xjr+Se5QgdKkjQ7LUQ-PdUULSN3n4+g@mail.gmail.com","threadId":"59531","inReplyTo":"xmqqttxvzbo8.fsf@gitster.g","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2023-04-06T13:39:31Z","receivedAt":"2023-04-06T13:39:58Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On Tue, 4 Apr 2023 at 21:31, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Paul Eggert <eggert@cs.ucla.edu> writes:\n>\n> > This is an evolving area. Git master is fiddling with flags and\n> > options, and so is GNU grep master, and so is PCRE2, and there are\n> > bugs. If you're running bleeding-edge versions of this code you'll get\n> > different behavior than if you're running grep 3.8, pcregrep 8.45,\n> > Perl 5.36, and git 2.39.2 (which is what Fedora 37 has).\n> >\n> > What I'm fearing is that we may evolve into mutually incompatible\n> > interpretations of how Perl regular expressions deal with UTF-8\n> > text. That'd be a recipe for confusion down the road.\n>\n> Nicely said.  My personal inclination is to let Perl folks decide\n> and follow them (even though I am skeptical about the wisdom of\n> letting '\\d' match anything other than [0-9]), but even in Git\n> circle there would be different opinions, so I am glad that the\n> discussion is visible on the list to those who are intrested.\n\n\nPerl matches Unicode text according to the rules specified by the\nUnicode consortium. It is the reference implementation for Unicode\nregular expression matching. Unicode specifies that \\d match any digit\nin any script that it supports. Thus \\d matches far more codepoints\nthan \\p{PosixDigit} or [0-9] would.  Be aware that Unicode contains\nand separates numbers and digits, eg, \\x{1EC9E} represents a Lakh,\nwhich is used in many Indian languages for 100,000, but which is not\nconsidered a *digit* for obvious reasons.\n\nFWIW, someone mentioned [[:digit:]] which matches the same as \\d does\non Unicode strings and under the /u matching flag for regexes in Perl.\nArguably this was a mistake, [[:digit:]] is a POSIX character class,\nand POSIX doesn't support Unicode so it should have matched [0-9] or\n\\p{PosixDigit}. But historically \\d and [[:digit:]] in Perl were the\nsame and when \\d was extended to meet the Unicode specification\n[[:digit:]] came along for the ride likely inadvertently, thus\n\\p{PosixDigit} is equivalent to [0-9], but \\p{XPosixDigit} is\nequivalent to \\d and [[:digit:]].\n\nI notice that other posts in this thread have moved the conversation\non, and covered most of the points I wanted to make here. However I\nwanted to say that there seem to be two different issues here. The\nfirst is \"what semantics do i expect from my regular expressions\",\nUnicode or legacy-ASCII, mostly this relates to case-insensitive\nmatching, but things like \\d also surface discrepancies. The second is\n\"what encodings does the regular expression engine understand\".\nUnfortunately on *nix there is no tradition of using BOM's to\ndistinguish the 6 different possible encodings of Unicode (UTF-8,\nUTF-EBCDIC, UTF-16LE, UTF-16BE, UTF-32LE, UTF-32BE), and there seems\nto be some level of desire of matching with unicode semantics against\nfiles that are not uniformly encoded in one of these formats.\n\nSo the question comes up, A) how do you tell the regular expression\nengine what semantics you want and B) how does the regular expression\nlibrary identify the encoding in the file, and how does it handle\nmalformed content in that file. For instance if I have a file which\ncontains snippets of UTF8 encoded data, *and* snippets of data that is\nillegal in UTF8, what should the regular expression engine do if it is\nasked to do a case insensitive match against that file.\n\ncheers,\nyves\n"},{"id":"474936","messageId":"CANgJU+XoyptS8NU+f6uMLrKjQakv=iN2c4DQydVaBVH3dK3s-w@mail.gmail.com","threadId":"59531","inReplyTo":"6d86214a-1b80-eb88-1efb-36e61fd3203e@cs.ucla.edu","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2023-04-06T15:45:09Z","receivedAt":"2023-04-06T15:45:25Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On Wed, 5 Apr 2023 at 20:32, Paul Eggert <eggert@cs.ucla.edu> wrote:\n>\n> On 2023-04-04 12:31, Junio C Hamano wrote:\n>\n> > My personal inclination is to let Perl folks decide\n> > and follow them (even though I am skeptical about the wisdom of\n> > letting '\\d' match anything other than [0-9])\n>\n> I looked into what pcre2grep does. It has always done only 8-bit\n> processing unless you use the -u or --utf option, so plain \"pcre2grep\n> '\\d'\" matches only ASCII digits.\n>\n> Although this causes pcre2grep to mishandle Unicode characters:\n>\n>    $ echo 'Ævar' | pcre2grep '[Ssß]'\n>    Ævar\n>\n> it mimics Perl 5.36:\n>\n>    $ echo 'Ævar' | perl -ne 'print $_ if /[Ssß]/'\n>    Ævar\n>\n> so this seems to be what Perl users expect, despite its infelicities.\n\nActually no, I think you have misunderstood what is happening at the\ndifferent layers involved here.\n\nYour terminal is rendering ß as a glyph. But it is almost certainly\nactually the octets C3 9F (which is the UTF8 canonical representation\nof the codepoint U+DF). So the code you provided to perl is close to\nthe equivalent of\n\necho 'Ævar' | perl -ne 'print $_ if /[Ss\\x{C3}\\x{9F}]/'\n\nAnd if you check, you will see that U+C6 \"Æ\" in utf8 is represented as\nthe octets C3 86.\n\nSo what you have done is the equivalent of:\n\nperl -le'print \"\\x{C3}\\x{86}\"' | perl -ne'print $_ if /[Ss\\x{C3}\\x{9F}]/'\n\nwhich of course matches. \\x{C3} matches \\x{C3} always and everywhere.\n\nWhat you should have done is something like this:\n\n$ echo 'Ævar' | perl -ne 'utf8::decode($_); print $_ if /[Ss\\x{DF}]/u'\n$ echo 'baß' | perl -MEncode -ne 'utf8::decode($_); print\nencode_utf8($_) if /[Ss\\x{DF}]/u'\nbaß\n$ echo 'Ævar' | perl -MEncode -ne 'utf8::decode($_); print\nencode_utf8($_) if /[Ss\\x{C6}]/u'\nÆvar\n$ echo 'Ævar' | perl -MEncode -ne 'utf8::decode($_); print\nencode_utf8($_) if /[Ss\\x{e6}]/ui'\nÆvar\n\nThe \"utf8::decode($_)\" tells perl to decode the input string as though\nit contained utf8 (which in this case it does). THe /u suffix tells\nthe regex engine that you want Unicode semantics.\n\nI believe that the same thing is true of your pcre2grep example. You\nsimply aren't checking what you think you are checking. You terminal\nrenders UTF8 as glyphs, but the programs you are feeding those glyphs\nto aren't seeing glyphs, they are seeing UTF8 sequences as distinct\noctets, and are not decoding their input back as codepoints.\n\nYou could have checked your assumptions by using the -Mre=debug option to perl:\n\n$ echo 'Ævar' | perl -Mre=debug -ne 'print $_ if /[Ssß]/'\nCompiling REx \"[Ss%x{c3}%x{9f}]\"\nFinal program:\n   1: ANYOF[Ss\\x9F\\xC3] (11)\n  11: END (0)\nstclass ANYOF[Ss\\x9F\\xC3] minlen 1\nMatching REx \"[Ss%x{c3}%x{9f}]\" against \"%x{c3}%x{86}var%n\"\nMatching stclass ANYOF[Ss\\x9F\\xC3] against \"%x{c3}%x{86}var%n\" (6 bytes)\n   0 <> <%x{c3}>             |   0| 1:ANYOF[Ss\\x9F\\xC3](11)\n   1 <%x{c3}> <%x{86}var>    |   0| 11:END(0)\nMatch successful!\nÆvar\nFreeing REx: \"[Ss%x{c3}%x{9f}]\"\n\nThe line: Matching REx \"[Ss%x{c3}%x{9f}]\" against \"%x{c3}%x{86}var%n\"\n\nbasically says it all. Perl has not decoded the UTF8 into U+C6, and it\nhas not decoded the UTF8 for U+DF either. Instead you have asked it if\nthe UTF8 sequence that represents U+C6 contains any of the same octets\nas the UTF8 representation of U+53, U+73 and U+DF would. Which gives\nthe common octet of \\x{c3}.\n\n> For better Unicode handling one can use pcre2grep's -u or --utf option,\n> which causes pcre2grep to behave more like GNU grep -P and git grep -P:\n> \"echo 'Ævar' | pcre2grep -u '[Ssß]'\" outputs nothing, which I think is\n> what most people would expect (unless they're Perl users :-).\n\nIt is what Perl users would expect also, assuming you actually wrote\nthe character class [Ss\\x{DF}] and asked for unicode semantics. \\x{DF}\nis the Latin1 codepoint range, so perl will assume that you meant\nASCII semantics unless you tell it otherwise.\n\nBasically these tests you have quoted here are just examples of\ngarbage in garbage out.\n\nPerl has been working together with the Unicode consortium for over 20\nyears. Afaik we were and are the reference implementation for the spec\non regular expression matching in Unicode and we have a long history\nof working together with the Unicode consortium to refine and\nimplement the spec. You should assume that if Perl seems to have made\na gross error in how it does Unicode matching that you are simply\nusing it wrong, we take a great deal of pride in having the best\nUnicode support there is.\n\nhttps://unicode.org/reports/tr18/\n\nFWIW, i think this email nicely illustrates the issues with git and\nregular expressions. To do regular expressions properly you need to\nknow a) what semantics do you expect, b) how to decode the text you\nare matching against. If you want unicode semantics you need to have a\nway to ask for it. If you want to match against Unicode data then you\nneed a way to determine which of the 6 possible encodings[1] of\nUnicode data you are using. If you get either wrong you will not get\nthe results you expect.  You may even want to deal with cases where\nyou want Unicode semantics, but to match against non-unicode data. For\ninstance Latin-1. In Latin-1 the codepoint U+DF is the *octet* 0xDF.\nMaybe you want that octet to match \"ss\" case-insensitively, as a\nGerman speaker would expect and as Unicode specifies is correct.  Or\nvice versa, maybe you are like some of the posters to this thread who\nseem to expect that \\d should not match U+16B51 (as a Hmong speaker\nmight expect). Perl resolves these problems at the pattern level by\nsupporting the suffixes /a and /u (for ascii and unicode), and at the\nstring level it supports two type of string, unicode strings, and\nbinary/ASCII strings. By default input is the latter but there are a\nvariety of ways of saying that a file handle should decode to Unicode\ninstead.\n\ncheers,\nYves\n[1] UTF-EBCDIC, UTF-8, UTF-16LE, UTF-16BE, UTF-32LE, UTF-32BE.\n\n\n\n\n--\nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"475010","messageId":"767d3617-e35a-a693-6ec8-f65421c68e5f@cs.ucla.edu","threadId":"59531","inReplyTo":"CANgJU+XoyptS8NU+f6uMLrKjQakv=iN2c4DQydVaBVH3dK3s-w@mail.gmail.com","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2023-04-07T16:48:40Z","receivedAt":"2023-04-07T16:48:51Z","isPatch":false,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"On 2023-04-06 08:45, demerphq wrote:\n>> Although this causes pcre2grep to mishandle Unicode characters:\n>>\n>>     $ echo 'Ævar' | pcre2grep '[Ssß]'\n>>     Ævar\n>>\n>> it mimics Perl 5.36:\n>>\n>>     $ echo 'Ævar' | perl -ne 'print $_ if /[Ssß]/'\n>>     Ævar\n>>\n>> so this seems to be what Perl users expect, despite its infelicities.\n> Actually no, I think you have misunderstood what is happening at the\n> different layers involved here.\n\nNo, I understood what was going on. My point was that Perl users seem to \nhave accepted this behavior, even though it does not match what people \nwould ordinarily expect.\n\n\n> What you should have done is something like this:\n\nNo, for two reasons. First, I'm no Perl expert and so I don't know (and \ndon't particularly want to learn) its complicated Unicode options and \ncalls. Second, /[Ss\\x{DF}]/u is hard to read. If I want the S letters of \ntraditional German, I'll write them in the obvious way, as [Ssß]. No \ndoubt Perl will let me do this somehow - but it is telling that none of \nyour examples do it in such a straightforward way.\n\n> $ echo 'Ævar' | perl -ne 'utf8::decode($_); print $_ if /[Ss\\x{DF}]/u'\n> $ echo 'baß' | perl -MEncode -ne 'utf8::decode($_); print\n> encode_utf8($_) if /[Ss\\x{DF}]/u'\n> baß\n> $ echo 'Ævar' | perl -MEncode -ne 'utf8::decode($_); print\n> encode_utf8($_) if /[Ss\\x{C6}]/u'\n> Ævar\n> $ echo 'Ævar' | perl -MEncode -ne 'utf8::decode($_); print\n> encode_utf8($_) if /[Ss\\x{e6}]/ui'\n> Ævar\n\n\n"},{"id":"475016","messageId":"065bcdcb-5770-5384-5afe-4a4d29272274@cs.ucla.edu","threadId":"59531","inReplyTo":"CANgJU+U+xXsh9psd0z5Xjr+Se5QgdKkjQ7LUQ-PdUULSN3n4+g@mail.gmail.com","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2023-04-07T19:00:16Z","receivedAt":"2023-04-07T19:02:09Z","isPatch":false,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"On 2023-04-06 06:39, demerphq wrote:\n\n> Unicode specifies that \\d match any digit\n> in any script that it supports.\n\n\"Specifies\" is too strong. The Unicode Regular Expressions technical \nstandard (UTS#18) mentions \\d only in Annex C[1], next to the word \n\"digit\" in a column labeled \"Property\" (even though \\d is really syntax \nnot a property). This is at best an informal recommendation, not a \nrequirement, as UTS#18 0.2[2] says that UTS#18's syntax is only for \nillustration and that although it's similar to Perl's, the two syntax \nforms may not be exactly the same. So we can't look to UTS#18 for a \ndefinitive way out of the \\d mess, as the Unicode folks specifically \ndelegated matters to us.\n\nEven ignoring the \\d issue the digit situation is messy. UTS#18 Annex C \nsays \"\\p{gc=Decimal_Number}\" is the standard recommended syntax \nassignment for digits. However, PCRE2 does not support this syntax; it \nsupports another variant \\p{Nd} that UTS#18 also recommends. So it \nappears that PCRE2 already does not implement every recommended aspect \nof UTS#18 syntax. PCRE2 also doesn't match Perl, which does support \n\"\\p{gc=Decimal_Number}\".\n\nAnyway, since grep -P '\\p{Nd}' implements Unicode's decimal digit class, \nthat's clearly enough for grep -P to conform to UTS#18 with respect to \ndigits.\n\n\n> A) how do you tell the regular expression\n> engine what semantics you want and B) how does the regular expression\n> library identify the encoding in the file, and how does it handle\n> malformed content in that file.\n\nHere's how GNU grep does it:\n\n* RE semantics are specified via command-line options like -P.\n\n* Text encoding is specified by locale, e.g., LC_ALL='en_US.utf8'.\n\n* REs do not match encoding errors.\n\n\n> on *nix there is no tradition of using BOM's to\n> distinguish the 6 different possible encodings of Unicode (UTF-8,\n> UTF-EBCDIC, UTF-16LE, UTF-16BE, UTF-32LE, UTF-32BE)\n\nYes, GNU/Linux never really experienced the joys of UTF-EBCDIC, Oracle \nUTFE, UTF-16LE vs UTF-16BE etc. If you're running legacy IBM mainframe \nor MS-Windows code these legacy encodings are obviously a big deal. \nHowever, there seems little reason to force their nontrivial hassles \nonto every GNU/Linux program that processes text. A few specialized apps \nlike 'iconv' deal with offbeat encodings, and that is probably a better \napproach all around.\n\n\n> there seems\n> to be some level of desire of matching with unicode semantics against\n> files that are not uniformly encoded in one of these formats.\n\nThat is a use case, yes. It's what 'strings' and 'grep' do.\n\n\n[1]: https://unicode.org/reports/tr18/#Compatibility_Properties\n[2]: https://unicode.org/reports/tr18/#Conformance\n\n"},{"id":"475031","messageId":"CAPUEspjtN-cwm=Nn=hMCcbOcOgPaVHsBfLW9TXn1HZrxtRR3BQ@mail.gmail.com","threadId":"59531","inReplyTo":"065bcdcb-5770-5384-5afe-4a4d29272274@cs.ucla.edu","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Carlo Arenas","fromEmail":"carenas@gmail.com","sentAt":"2023-04-08T05:01:14Z","receivedAt":"2023-04-08T05:01:38Z","isPatch":false,"sender":{"key":"carenas@gmail.com","avatar":"https://avatars.githubusercontent.com/u/76036?v=4"},"body":"On Fri, Apr 7, 2023 at 12:00 PM Paul Eggert <eggert@cs.ucla.edu> wrote:\n>\n> On 2023-04-06 06:39, demerphq wrote:\n>\n> > Unicode specifies that \\d match any digit\n> > in any script that it supports.\n>\n> \"Specifies\" is too strong. The Unicode Regular Expressions technical\n> standard (UTS#18) mentions \\d only in Annex C[1], next to the word\n> \"digit\" in a column labeled \"Property\" (even though \\d is really syntax\n> not a property). This is at best an informal recommendation, not a\n> requirement, as UTS#18 0.2[2] says that UTS#18's syntax is only for\n> illustration and that although it's similar to Perl's, the two syntax\n> forms may not be exactly the same. So we can't look to UTS#18 for a\n> definitive way out of the \\d mess, as the Unicode folks specifically\n> delegated matters to us.\n>\n> Even ignoring the \\d issue the digit situation is messy. UTS#18 Annex C\n> says \"\\p{gc=Decimal_Number}\" is the standard recommended syntax\n> assignment for digits. However, PCRE2 does not support this syntax; it\n> supports another variant \\p{Nd} that UTS#18 also recommends. So it\n> appears that PCRE2 already does not implement every recommended aspect\n> of UTS#18 syntax. PCRE2 also doesn't match Perl, which does support\n> \"\\p{gc=Decimal_Number}\".\n\nNot sure I follow the whole logic here, but PCRE2[3] (search for\n\"general category\" which is what the \"gc\" above stands for) only\nsupports the abbreviated form of the unicode classes and `Nd` is\nindeed the one that corresponds to `Decimal_Number`.\n\nCarlo\n\n[1]: https://unicode.org/reports/tr18/#Compatibility_Properties\n[2]: https://unicode.org/reports/tr18/#Conformance\n[3]: https://pcre2project.github.io/pcre2/doc/html/pcre2pattern.html\n"},{"id":"475040","messageId":"43d04851-2463-2922-44e3-075080129ec3@cs.ucla.edu","threadId":"59531","inReplyTo":"CAPUEspjtN-cwm=Nn=hMCcbOcOgPaVHsBfLW9TXn1HZrxtRR3BQ@mail.gmail.com","subject":"Re: bug#60690: -P '\\d' in GNU and git grep","fromName":"Paul Eggert","fromEmail":"eggert@cs.ucla.edu","sentAt":"2023-04-08T22:45:20Z","receivedAt":"2023-04-08T22:45:26Z","isPatch":false,"sender":{"key":"eggert@cs.ucla.edu","avatar":"https://avatars.githubusercontent.com/u/572024?v=4"},"body":"On 2023-04-07 22:01, Carlo Arenas wrote:\n\n> Not sure I follow the whole logic here, but PCRE2[3] (search for\n> \"general category\" which is what the \"gc\" above stands for) only\n> supports the abbreviated form of the unicode classes and `Nd` is\n> indeed the one that corresponds to `Decimal_Number`.\n\nThat's fine: all that UTS#18[1] requires is that PCRE2 provide syntax \nfor a regular expression that matches the Decimal Number class. Which \nPCRE2 does, via \\p{Nd}.\n\nThe logic is that UTF#18 does not require that \\d must behave like \n\\p{Nd}, or even that \\p{gc=Decimal_Number} must behave like \\p{Nd}. It \nmerely requires that there be some syntax for matching Decimal Number, \nand it says the choice of syntax is up to the implementer. This is why \nUTF#18 doesn't require that \\d must also match non-ASCII digits (which \nis what I think Yves was saying).\n\n[1]: https://unicode.org/reports/tr18/\n"}]}