{"thread":{"id":"33957","subject":"[PATCH] wildmatch: properly fold case everywhere","startedAt":"2013-05-28T12:32:41Z","lastAt":"2013-06-02T23:42:51Z","messageCount":18,"participants":["Anthony Ramine","Duy Nguyen","Eric Sunshine","Junio C Hamano"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"218644","messageId":"1369744361-44918-1-git-send-email-n.oxyde@gmail.com","threadId":"33957","inReplyTo":null,"subject":"[PATCH] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-05-28T12:32:41Z","receivedAt":"2013-05-28T12:32:41Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"Case folding is not done correctly when matching against the [:upper:]\ncharacter class and uppercased character ranges (e.g. A-Z).\nSpecifically, an uppercase letter fails to match against any of them\nwhen case folding is requested because plain characters in the pattern\nand the whole string and preemptively lowercased to handle the base case\nfast.\n\nThat optimization is kept and ISLOWER() is used in the [:upper:] case\nwhen case folding is requested, while matching against a character range\nis retried with toupper() if the character was lowercase.\n\nSigned-off-by: Anthony Ramine <n.oxyde@gmail.com>\n---\n t/t3070-wildmatch.sh | 43 +++++++++++++++++++++++++++++++++++++------\n wildmatch.c          |  7 +++++++\n 2 files changed, 44 insertions(+), 6 deletions(-)\n\ndiff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nindex 4c37057..17315aa 100755\n--- a/t/t3070-wildmatch.sh\n+++ b/t/t3070-wildmatch.sh\n@@ -6,20 +6,20 @@ test_description='wildmatch tests'\n \n match() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"wildmatch:    match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:     match '$3' '$4'\" \"\n \t    test-wildmatch wildmatch '$3' '$4'\n \t\"\n     else\n-\ttest_expect_success \"wildmatch: no match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:  no match '$3' '$4'\" \"\n \t    ! test-wildmatch wildmatch '$3' '$4'\n \t\"\n     fi\n     if [ $2 = 1 ]; then\n-\ttest_expect_success \"fnmatch:      match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:       match '$3' '$4'\" \"\n \t    test-wildmatch fnmatch '$3' '$4'\n \t\"\n     elif [ $2 = 0 ]; then\n-\ttest_expect_success \"fnmatch:   no match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:    no match '$3' '$4'\" \"\n \t    ! test-wildmatch fnmatch '$3' '$4'\n \t\"\n #    else\n@@ -29,13 +29,25 @@ match() {\n     fi\n }\n \n+imatch() {\n+    if [ $1 = 1 ]; then\n+\ttest_expect_success \"iwildmatch:    match '$2' '$3'\" \"\n+\t    test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    else\n+\ttest_expect_success \"iwildmatch: no match '$2' '$3'\" \"\n+\t    ! test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    fi\n+}\n+\n pathmatch() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"pathmatch:    match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:     match '$2' '$3'\" \"\n \t    test-wildmatch pathmatch '$2' '$3'\n \t\"\n     else\n-\ttest_expect_success \"pathmatch: no match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:  no match '$2' '$3'\" \"\n \t    ! test-wildmatch pathmatch '$2' '$3'\n \t\"\n     fi\n@@ -235,4 +247,23 @@ pathmatch 1 abcXdefXghi '*X*i'\n pathmatch 1 ab/cXd/efXg/hi '*/*X*/*/*i'\n pathmatch 1 ab/cXd/efXg/hi '*Xg*i'\n \n+# Case-sensitivy features\n+match 0 x 'a' '[A-Z]'\n+match 1 x 'A' '[A-Z]'\n+match 0 x 'A' '[a-z]'\n+match 1 x 'a' '[a-z]'\n+match 0 x 'a' '[[:upper:]]'\n+match 1 x 'A' '[[:upper:]]'\n+match 0 x 'A' '[[:lower:]]'\n+match 1 x 'a' '[[:lower:]]'\n+\n+imatch 1 'a' '[A-Z]'\n+imatch 1 'A' '[A-Z]'\n+imatch 1 'A' '[a-z]'\n+imatch 1 'a' '[a-z]'\n+imatch 1 'a' '[[:upper:]]'\n+imatch 1 'A' '[[:upper:]]'\n+imatch 1 'A' '[[:lower:]]'\n+imatch 1 'a' '[[:lower:]]'\n+\n test_done\ndiff --git a/wildmatch.c b/wildmatch.c\nindex 7192bdc..ea318d3 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t}\n \t\t\t\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n \t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n+\t\t\t\t\t\tt_ch = toupper(t_ch);\n+\t\t\t\t\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n+\t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t}\n \t\t\t\t\tp_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n \t\t\t\t} else if (p_ch == '[' && p[1] == ':') {\n \t\t\t\t\tconst uchar *s;\n@@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t} else if (CC_EQ(s,i, \"upper\")) {\n \t\t\t\t\t\tif (ISUPPER(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n+\t\t\t\t\t\t\tmatched = 1;\n \t\t\t\t\t} else if (CC_EQ(s,i, \"xdigit\")) {\n \t\t\t\t\t\tif (ISXDIGIT(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n-- \n1.8.3\n"},{"id":"218645","messageId":"CACsJy8Bu2XvapGHjcKZJuATqB90MXSCoNHkke9CBeiwSvzpH8A@mail.gmail.com","threadId":"33957","inReplyTo":"1369744361-44918-1-git-send-email-n.oxyde@gmail.com","subject":"Re: [PATCH] wildmatch: properly fold case everywhere","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-05-28T12:53:47Z","receivedAt":"2013-05-28T12:53:47Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, May 28, 2013 at 7:32 PM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n> @@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>                                         }\n>                                         if (t_ch <= p_ch && t_ch >= prev_ch)\n>                                                 matched = 1;\n> +                                       else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n> +                                               t_ch = toupper(t_ch);\n\nThis happens in a while loop where t_ch may be used again. Should we\nmake a local copy of toupper(t_ch) and leave t_ch untouched?\n\n> +                                               if (t_ch <= p_ch && t_ch >= prev_ch)\n> +                                                       matched = 1;\n> +                                       }\n>                                         p_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n>                                 } else if (p_ch == '[' && p[1] == ':') {\n>                                         const uchar *s;\n--\nDuy\n"},{"id":"218657","messageId":"5A688100-5F54-4945-85BB-643B69C05F85@gmail.com","threadId":"33957","inReplyTo":"CACsJy8Bu2XvapGHjcKZJuATqB90MXSCoNHkke9CBeiwSvzpH8A@mail.gmail.com","subject":"Re: [PATCH] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-05-28T13:01:51Z","receivedAt":"2013-05-28T13:01:51Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"You're right, I will amend my patch. How do I make git-send-email reply to that thread?\n\n-- \nAnthony Ramine\n\nLe 28 mai 2013 à 14:53, Duy Nguyen a écrit :\n\n> On Tue, May 28, 2013 at 7:32 PM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n>> @@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>>                                        }\n>>                                        if (t_ch <= p_ch && t_ch >= prev_ch)\n>>                                                matched = 1;\n>> +                                       else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n>> +                                               t_ch = toupper(t_ch);\n> \n> This happens in a while loop where t_ch may be used again. Should we\n> make a local copy of toupper(t_ch) and leave t_ch untouched?\n> \n>> +                                               if (t_ch <= p_ch && t_ch >= prev_ch)\n>> +                                                       matched = 1;\n>> +                                       }\n>>                                        p_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n>>                                } else if (p_ch == '[' && p[1] == ':') {\n>>                                        const uchar *s;\n> --\n> Duy\n"},{"id":"218658","messageId":"1369746650-53869-1-git-send-email-n.oxyde@gmail.com","threadId":"33957","inReplyTo":"1369744361-44918-1-git-send-email-n.oxyde@gmail.com","subject":"[PATCH v2] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-05-28T13:10:50Z","receivedAt":"2013-05-28T13:10:50Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"Case folding is not done correctly when matching against the [:upper:]\ncharacter class and uppercased character ranges (e.g. A-Z).\nSpecifically, an uppercase letter fails to match against any of them\nwhen case folding is requested because plain characters in the pattern\nand the whole string and preemptively lowercased to handle the base case\nfast.\n\nThat optimization is kept and ISLOWER() is used in the [:upper:] case\nwhen case folding is requested, while matching against a character range\nis retried with toupper() if the character was lowercase.\n\nSigned-off-by: Anthony Ramine <n.oxyde@gmail.com>\n---\n t/t3070-wildmatch.sh | 47 +++++++++++++++++++++++++++++++++++++++++------\n wildmatch.c          |  7 +++++++\n 2 files changed, 48 insertions(+), 6 deletions(-)\n\ndiff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nindex 4c37057..e1b45e6 100755\n--- a/t/t3070-wildmatch.sh\n+++ b/t/t3070-wildmatch.sh\n@@ -6,20 +6,20 @@ test_description='wildmatch tests'\n \n match() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"wildmatch:    match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:     match '$3' '$4'\" \"\n \t    test-wildmatch wildmatch '$3' '$4'\n \t\"\n     else\n-\ttest_expect_success \"wildmatch: no match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:  no match '$3' '$4'\" \"\n \t    ! test-wildmatch wildmatch '$3' '$4'\n \t\"\n     fi\n     if [ $2 = 1 ]; then\n-\ttest_expect_success \"fnmatch:      match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:       match '$3' '$4'\" \"\n \t    test-wildmatch fnmatch '$3' '$4'\n \t\"\n     elif [ $2 = 0 ]; then\n-\ttest_expect_success \"fnmatch:   no match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:    no match '$3' '$4'\" \"\n \t    ! test-wildmatch fnmatch '$3' '$4'\n \t\"\n #    else\n@@ -29,13 +29,25 @@ match() {\n     fi\n }\n \n+imatch() {\n+    if [ $1 = 1 ]; then\n+\ttest_expect_success \"iwildmatch:    match '$2' '$3'\" \"\n+\t    test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    else\n+\ttest_expect_success \"iwildmatch: no match '$2' '$3'\" \"\n+\t    ! test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    fi\n+}\n+\n pathmatch() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"pathmatch:    match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:     match '$2' '$3'\" \"\n \t    test-wildmatch pathmatch '$2' '$3'\n \t\"\n     else\n-\ttest_expect_success \"pathmatch: no match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:  no match '$2' '$3'\" \"\n \t    ! test-wildmatch pathmatch '$2' '$3'\n \t\"\n     fi\n@@ -235,4 +247,27 @@ pathmatch 1 abcXdefXghi '*X*i'\n pathmatch 1 ab/cXd/efXg/hi '*/*X*/*/*i'\n pathmatch 1 ab/cXd/efXg/hi '*Xg*i'\n \n+# Case-sensitivy features\n+match 0 x 'a' '[A-Z]'\n+match 1 x 'A' '[A-Z]'\n+match 0 x 'A' '[a-z]'\n+match 1 x 'a' '[a-z]'\n+match 0 x 'a' '[[:upper:]]'\n+match 1 x 'A' '[[:upper:]]'\n+match 0 x 'A' '[[:lower:]]'\n+match 1 x 'a' '[[:lower:]]'\n+match 0 x 'A' '[B-Za]'\n+match 1 x 'a' '[B-Za]'\n+\n+imatch 1 'a' '[A-Z]'\n+imatch 1 'A' '[A-Z]'\n+imatch 1 'A' '[a-z]'\n+imatch 1 'a' '[a-z]'\n+imatch 1 'a' '[[:upper:]]'\n+imatch 1 'A' '[[:upper:]]'\n+imatch 1 'A' '[[:lower:]]'\n+imatch 1 'a' '[[:lower:]]'\n+imatch 1 'A' '[B-Za]'\n+imatch 1 'a' '[B-Za]'\n+\n test_done\ndiff --git a/wildmatch.c b/wildmatch.c\nindex 7192bdc..ea318d3 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t}\n \t\t\t\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n \t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n+\t\t\t\t\t\tt_ch = toupper(t_ch);\n+\t\t\t\t\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n+\t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t}\n \t\t\t\t\tp_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n \t\t\t\t} else if (p_ch == '[' && p[1] == ':') {\n \t\t\t\t\tconst uchar *s;\n@@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t} else if (CC_EQ(s,i, \"upper\")) {\n \t\t\t\t\t\tif (ISUPPER(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n+\t\t\t\t\t\t\tmatched = 1;\n \t\t\t\t\t} else if (CC_EQ(s,i, \"xdigit\")) {\n \t\t\t\t\t\tif (ISXDIGIT(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n-- \n1.8.3\n"},{"id":"218666","messageId":"1369749497-55610-1-git-send-email-n.oxyde@gmail.com","threadId":"33957","inReplyTo":"1369744361-44918-1-git-send-email-n.oxyde@gmail.com","subject":"[PATCH v3] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-05-28T13:58:17Z","receivedAt":"2013-05-28T13:58:17Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"Case folding is not done correctly when matching against the [:upper:]\ncharacter class and uppercased character ranges (e.g. A-Z).\nSpecifically, an uppercase letter fails to match against any of them\nwhen case folding is requested because plain characters in the pattern\nand the whole string and preemptively lowercased to handle the base case\nfast.\n\nThat optimization is kept and ISLOWER() is used in the [:upper:] case\nwhen case folding is requested, while matching against a character range\nis retried with toupper() if the character was lowercase.\n\nSigned-off-by: Anthony Ramine <n.oxyde@gmail.com>\n---\n t/t3070-wildmatch.sh | 47 +++++++++++++++++++++++++++++++++++++++++------\n wildmatch.c          |  7 +++++++\n 2 files changed, 48 insertions(+), 6 deletions(-)\n\nPlease disregard PATCH v2, it is identical to the first one.\n\ndiff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nindex 4c37057..e1b45e6 100755\n--- a/t/t3070-wildmatch.sh\n+++ b/t/t3070-wildmatch.sh\n@@ -6,20 +6,20 @@ test_description='wildmatch tests'\n \n match() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"wildmatch:    match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:     match '$3' '$4'\" \"\n \t    test-wildmatch wildmatch '$3' '$4'\n \t\"\n     else\n-\ttest_expect_success \"wildmatch: no match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:  no match '$3' '$4'\" \"\n \t    ! test-wildmatch wildmatch '$3' '$4'\n \t\"\n     fi\n     if [ $2 = 1 ]; then\n-\ttest_expect_success \"fnmatch:      match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:       match '$3' '$4'\" \"\n \t    test-wildmatch fnmatch '$3' '$4'\n \t\"\n     elif [ $2 = 0 ]; then\n-\ttest_expect_success \"fnmatch:   no match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:    no match '$3' '$4'\" \"\n \t    ! test-wildmatch fnmatch '$3' '$4'\n \t\"\n #    else\n@@ -29,13 +29,25 @@ match() {\n     fi\n }\n \n+imatch() {\n+    if [ $1 = 1 ]; then\n+\ttest_expect_success \"iwildmatch:    match '$2' '$3'\" \"\n+\t    test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    else\n+\ttest_expect_success \"iwildmatch: no match '$2' '$3'\" \"\n+\t    ! test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    fi\n+}\n+\n pathmatch() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"pathmatch:    match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:     match '$2' '$3'\" \"\n \t    test-wildmatch pathmatch '$2' '$3'\n \t\"\n     else\n-\ttest_expect_success \"pathmatch: no match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:  no match '$2' '$3'\" \"\n \t    ! test-wildmatch pathmatch '$2' '$3'\n \t\"\n     fi\n@@ -235,4 +247,27 @@ pathmatch 1 abcXdefXghi '*X*i'\n pathmatch 1 ab/cXd/efXg/hi '*/*X*/*/*i'\n pathmatch 1 ab/cXd/efXg/hi '*Xg*i'\n \n+# Case-sensitivy features\n+match 0 x 'a' '[A-Z]'\n+match 1 x 'A' '[A-Z]'\n+match 0 x 'A' '[a-z]'\n+match 1 x 'a' '[a-z]'\n+match 0 x 'a' '[[:upper:]]'\n+match 1 x 'A' '[[:upper:]]'\n+match 0 x 'A' '[[:lower:]]'\n+match 1 x 'a' '[[:lower:]]'\n+match 0 x 'A' '[B-Za]'\n+match 1 x 'a' '[B-Za]'\n+\n+imatch 1 'a' '[A-Z]'\n+imatch 1 'A' '[A-Z]'\n+imatch 1 'A' '[a-z]'\n+imatch 1 'a' '[a-z]'\n+imatch 1 'a' '[[:upper:]]'\n+imatch 1 'A' '[[:upper:]]'\n+imatch 1 'A' '[[:lower:]]'\n+imatch 1 'a' '[[:lower:]]'\n+imatch 1 'A' '[B-Za]'\n+imatch 1 'a' '[B-Za]'\n+\n test_done\ndiff --git a/wildmatch.c b/wildmatch.c\nindex 7192bdc..f91ba99 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t}\n \t\t\t\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n \t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n+\t\t\t\t\t\tuchar t_ch_upper = toupper(t_ch);\n+\t\t\t\t\t\tif (t_ch_upper <= p_ch && t_ch_upper >= prev_ch)\n+\t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t}\n \t\t\t\t\tp_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n \t\t\t\t} else if (p_ch == '[' && p[1] == ':') {\n \t\t\t\t\tconst uchar *s;\n@@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t} else if (CC_EQ(s,i, \"upper\")) {\n \t\t\t\t\t\tif (ISUPPER(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n+\t\t\t\t\t\t\tmatched = 1;\n \t\t\t\t\t} else if (CC_EQ(s,i, \"xdigit\")) {\n \t\t\t\t\t\tif (ISXDIGIT(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n-- \n1.8.3\n"},{"id":"218824","messageId":"CACsJy8CY_T44ymUnLWv4FpF3zpL3WKSysJ1wBhfxGHNPJ6kSmg@mail.gmail.com","threadId":"33957","inReplyTo":"1369749497-55610-1-git-send-email-n.oxyde@gmail.com","subject":"Re: [PATCH v3] wildmatch: properly fold case everywhere","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-05-29T13:22:11Z","receivedAt":"2013-05-29T13:22:11Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Tue, May 28, 2013 at 8:58 PM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n> Case folding is not done correctly when matching against the [:upper:]\n> character class and uppercased character ranges (e.g. A-Z).\n> Specifically, an uppercase letter fails to match against any of them\n> when case folding is requested because plain characters in the pattern\n> and the whole string and preemptively lowercased to handle the base case\n> fast.\n\nI did a little test with glibc fnmatch and also checked the source\ncode. I don't think 'a' matches [:upper:]. So I'm not sure if that's a\ncorrect behavior or a bug in glibc. The spec is not clear (I think) on\nthis. I guess we should just assume that 'a' should match '[:upper:]'?\n\n> @@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>                                         }\n>                                         if (t_ch <= p_ch && t_ch >= prev_ch)\n>                                                 matched = 1;\n> +                                       else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n> +                                               uchar t_ch_upper = toupper(t_ch);\n> +                                               if (t_ch_upper <= p_ch && t_ch_upper >= prev_ch)\n> +                                                       matched = 1;\n> +                                       }\n\nOr we could stick with to tolower. Something like this\n\nif ((t_ch <= p_ch && t_ch >= prev_ch) ||\n   ((flags & WM_CASEFOLD) &&\n      t_ch <= tolower(p_ch) && t_ch >= tolower(prev_ch)))\n   match = 1;\n\nI think it's easier to read if we either downcase all, or upcase all, not both.\n\n>                                         p_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n>                                 } else if (p_ch == '[' && p[1] == ':') {\n>                                         const uchar *s;\n> @@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>                                         } else if (CC_EQ(s,i, \"upper\")) {\n>                                                 if (ISUPPER(t_ch))\n>                                                         matched = 1;\n> +                                               else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n> +                                                       matched = 1;\n>                                         } else if (CC_EQ(s,i, \"xdigit\")) {\n>                                                 if (ISXDIGIT(t_ch))\n>                                                         matched = 1;\n\nIf WM_CASEFOLD is set, maybe isalpha(t_ch) is enough then?\n--\nDuy\n"},{"id":"218831","messageId":"4E816EBA-A22D-4507-BED0-0DE55D2E619C@gmail.com","threadId":"33957","inReplyTo":"CACsJy8CY_T44ymUnLWv4FpF3zpL3WKSysJ1wBhfxGHNPJ6kSmg@mail.gmail.com","subject":"Re: [PATCH v3] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-05-29T13:37:10Z","receivedAt":"2013-05-29T13:37:10Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"Replied inline.\n\nRegards,\n\n-- \nAnthony Ramine\n\nLe 29 mai 2013 à 15:22, Duy Nguyen a écrit :\n\n> On Tue, May 28, 2013 at 8:58 PM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n>> Case folding is not done correctly when matching against the [:upper:]\n>> character class and uppercased character ranges (e.g. A-Z).\n>> Specifically, an uppercase letter fails to match against any of them\n>> when case folding is requested because plain characters in the pattern\n>> and the whole string and preemptively lowercased to handle the base case\n>> fast.\n> \n> I did a little test with glibc fnmatch and also checked the source\n> code. I don't think 'a' matches [:upper:]. So I'm not sure if that's a\n> correct behavior or a bug in glibc. The spec is not clear (I think) on\n> this. I guess we should just assume that 'a' should match '[:upper:]'?\n\nI don't know, in my opinion if case folding is enabled we should say [:upper:], [:lower:] and [:alpha:] are equivalent.\n\nThis opinion is shared by GNU Flex [1]:\n\n> \t• If your scanner is case-insensitive (the ‘-i’ flag), then ‘[:upper:]’ and ‘[:lower:]’ are equivalent to ‘[:alpha:]’.\n\n[1] http://flex.sourceforge.net/manual/Patterns.html\n\n>> @@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>>                                        }\n>>                                        if (t_ch <= p_ch && t_ch >= prev_ch)\n>>                                                matched = 1;\n>> +                                       else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n>> +                                               uchar t_ch_upper = toupper(t_ch);\n>> +                                               if (t_ch_upper <= p_ch && t_ch_upper >= prev_ch)\n>> +                                                       matched = 1;\n>> +                                       }\n> \n> Or we could stick with to tolower. Something like this\n> \n> if ((t_ch <= p_ch && t_ch >= prev_ch) ||\n>   ((flags & WM_CASEFOLD) &&\n>      t_ch <= tolower(p_ch) && t_ch >= tolower(prev_ch)))\n>   match = 1;\n> \n> I think it's easier to read if we either downcase all, or upcase all, not both.\n\nIf the range to match against is [A-_], it will become [a-_] which is an empty range, ord('a') > ord('_'). I think it is simpler to reuse toupper() after the fact as I did.\n\nAnyway maybe I should add a test for that corner case?\n\n>>                                        p_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n>>                                } else if (p_ch == '[' && p[1] == ':') {\n>>                                        const uchar *s;\n>> @@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>>                                        } else if (CC_EQ(s,i, \"upper\")) {\n>>                                                if (ISUPPER(t_ch))\n>>                                                        matched = 1;\n>> +                                               else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n>> +                                                       matched = 1;\n>>                                        } else if (CC_EQ(s,i, \"xdigit\")) {\n>>                                                if (ISXDIGIT(t_ch))\n>>                                                        matched = 1;\n> \n> If WM_CASEFOLD is set, maybe isalpha(t_ch) is enough then?\n\nYes isalpha() is enought but I wanted to keep the two cases separated, I can amend that if you want.\n"},{"id":"218838","messageId":"CACsJy8A61nYu9a-BhUiBhBEv-e6_CtYyZE3sG9iCiau+3EKVdw@mail.gmail.com","threadId":"33957","inReplyTo":"4E816EBA-A22D-4507-BED0-0DE55D2E619C@gmail.com","subject":"Re: [PATCH v3] wildmatch: properly fold case everywhere","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-05-29T13:52:07Z","receivedAt":"2013-05-29T13:52:07Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, May 29, 2013 at 8:37 PM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n> Le 29 mai 2013 à 15:22, Duy Nguyen a écrit :\n>\n>> On Tue, May 28, 2013 at 8:58 PM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n>>> Case folding is not done correctly when matching against the [:upper:]\n>>> character class and uppercased character ranges (e.g. A-Z).\n>>> Specifically, an uppercase letter fails to match against any of them\n>>> when case folding is requested because plain characters in the pattern\n>>> and the whole string and preemptively lowercased to handle the base case\n>>> fast.\n>>\n>> I did a little test with glibc fnmatch and also checked the source\n>> code. I don't think 'a' matches [:upper:]. So I'm not sure if that's a\n>> correct behavior or a bug in glibc. The spec is not clear (I think) on\n>> this. I guess we should just assume that 'a' should match '[:upper:]'?\n>\n> I don't know, in my opinion if case folding is enabled we should say [:upper:], [:lower:] and [:alpha:] are equivalent.\n>\n> This opinion is shared by GNU Flex [1]:\n>\n>>       • If your scanner is case-insensitive (the ‘-i’ flag), then ‘[:upper:]’ and ‘[:lower:]’ are equivalent to ‘[:alpha:]’.\n>\n> [1] http://flex.sourceforge.net/manual/Patterns.html\n\nThen we should do it too because of this precedent, I think.\n\n>>> @@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>>>                                        }\n>>>                                        if (t_ch <= p_ch && t_ch >= prev_ch)\n>>>                                                matched = 1;\n>>> +                                       else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n>>> +                                               uchar t_ch_upper = toupper(t_ch);\n>>> +                                               if (t_ch_upper <= p_ch && t_ch_upper >= prev_ch)\n>>> +                                                       matched = 1;\n>>> +                                       }\n>>\n>> Or we could stick with to tolower. Something like this\n>>\n>> if ((t_ch <= p_ch && t_ch >= prev_ch) ||\n>>   ((flags & WM_CASEFOLD) &&\n>>      t_ch <= tolower(p_ch) && t_ch >= tolower(prev_ch)))\n>>   match = 1;\n>>\n>> I think it's easier to read if we either downcase all, or upcase all, not both.\n>\n> If the range to match against is [A-_], it will become [a-_] which is an empty range, ord('a') > ord('_'). I think it is simpler to reuse toupper() after the fact as I did.\n>\n> Anyway maybe I should add a test for that corner case?\n\nYeah I was thinking about such a case, but I saw glibc do it... I\nguess we just found another bug, at least in compat/fnmatch.c. Yes a\ntest for it would be great, in case I change my mind 2 years from now\nand decide to turn it the other way ;)\n\n>\n>>>                                        p_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n>>>                                } else if (p_ch == '[' && p[1] == ':') {\n>>>                                        const uchar *s;\n>>> @@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>>>                                        } else if (CC_EQ(s,i, \"upper\")) {\n>>>                                                if (ISUPPER(t_ch))\n>>>                                                        matched = 1;\n>>> +                                               else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n>>> +                                                       matched = 1;\n>>>                                        } else if (CC_EQ(s,i, \"xdigit\")) {\n>>>                                                if (ISXDIGIT(t_ch))\n>>>                                                        matched = 1;\n>>\n>> If WM_CASEFOLD is set, maybe isalpha(t_ch) is enough then?\n>\n> Yes isalpha() is enought but I wanted to keep the two cases separated, I can amend that if you want.\n\nEither way is fine. I don't think this code is performance critical. Your call.\n--\nDuy\n"},{"id":"218866","messageId":"BAB62C57-FE7D-476A-ACA7-5831BAF3E558@gmail.com","threadId":"33957","inReplyTo":"CACsJy8A61nYu9a-BhUiBhBEv-e6_CtYyZE3sG9iCiau+3EKVdw@mail.gmail.com","subject":"Re: [PATCH v3] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-05-29T17:57:44Z","receivedAt":"2013-05-29T17:57:44Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"Replied inline.\n\n-- \nAnthony Ramine\n\nLe 29 mai 2013 à 15:52, Duy Nguyen a écrit :\n\n> On Wed, May 29, 2013 at 8:37 PM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n>> Le 29 mai 2013 à 15:22, Duy Nguyen a écrit :\n>> \n>>> On Tue, May 28, 2013 at 8:58 PM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n>>>> Case folding is not done correctly when matching against the [:upper:]\n>>>> character class and uppercased character ranges (e.g. A-Z).\n>>>> Specifically, an uppercase letter fails to match against any of them\n>>>> when case folding is requested because plain characters in the pattern\n>>>> and the whole string and preemptively lowercased to handle the base case\n>>>> fast.\n>>> \n>>> I did a little test with glibc fnmatch and also checked the source\n>>> code. I don't think 'a' matches [:upper:]. So I'm not sure if that's a\n>>> correct behavior or a bug in glibc. The spec is not clear (I think) on\n>>> this. I guess we should just assume that 'a' should match '[:upper:]'?\n>> \n>> I don't know, in my opinion if case folding is enabled we should say [:upper:], [:lower:] and [:alpha:] are equivalent.\n>> \n>> This opinion is shared by GNU Flex [1]:\n>> \n>>>      • If your scanner is case-insensitive (the ‘-i’ flag), then ‘[:upper:]’ and ‘[:lower:]’ are equivalent to ‘[:alpha:]’.\n>> \n>> [1] http://flex.sourceforge.net/manual/Patterns.html\n> \n> Then we should do it too because of this precedent, I think.\n> \n>>>> @@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>>>>                                       }\n>>>>                                       if (t_ch <= p_ch && t_ch >= prev_ch)\n>>>>                                               matched = 1;\n>>>> +                                       else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n>>>> +                                               uchar t_ch_upper = toupper(t_ch);\n>>>> +                                               if (t_ch_upper <= p_ch && t_ch_upper >= prev_ch)\n>>>> +                                                       matched = 1;\n>>>> +                                       }\n>>> \n>>> Or we could stick with to tolower. Something like this\n>>> \n>>> if ((t_ch <= p_ch && t_ch >= prev_ch) ||\n>>>  ((flags & WM_CASEFOLD) &&\n>>>     t_ch <= tolower(p_ch) && t_ch >= tolower(prev_ch)))\n>>>  match = 1;\n>>> \n>>> I think it's easier to read if we either downcase all, or upcase all, not both.\n>> \n>> If the range to match against is [A-_], it will become [a-_] which is an empty range, ord('a') > ord('_'). I think it is simpler to reuse toupper() after the fact as I did.\n>> \n>> Anyway maybe I should add a test for that corner case?\n> \n> Yeah I was thinking about such a case, but I saw glibc do it... I\n> guess we just found another bug, at least in compat/fnmatch.c. Yes a\n> test for it would be great, in case I change my mind 2 years from now\n> and decide to turn it the other way ;)\n\nShould I patch compat/fnmatch.c too? That would make it different from the glibc's one.\n\n>> \n>>>>                                       p_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n>>>>                               } else if (p_ch == '[' && p[1] == ':') {\n>>>>                                       const uchar *s;\n>>>> @@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>>>>                                       } else if (CC_EQ(s,i, \"upper\")) {\n>>>>                                               if (ISUPPER(t_ch))\n>>>>                                                       matched = 1;\n>>>> +                                               else if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n>>>> +                                                       matched = 1;\n>>>>                                       } else if (CC_EQ(s,i, \"xdigit\")) {\n>>>>                                               if (ISXDIGIT(t_ch))\n>>>>                                                       matched = 1;\n>>> \n>>> If WM_CASEFOLD is set, maybe isalpha(t_ch) is enough then?\n>> \n>> Yes isalpha() is enought but I wanted to keep the two cases separated, I can amend that if you want.\n> \n> Either way is fine. I don't think this code is performance critical. Your call.\n> --\n> Duy\n"},{"id":"218914","messageId":"CACsJy8CuaowyZJGKh7X+43qRwYAdUCDbVo8P5CpEtukBzRiReg@mail.gmail.com","threadId":"33957","inReplyTo":"BAB62C57-FE7D-476A-ACA7-5831BAF3E558@gmail.com","subject":"Re: [PATCH v3] wildmatch: properly fold case everywhere","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-05-30T00:04:30Z","receivedAt":"2013-05-30T00:04:30Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, May 30, 2013 at 12:57 AM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n>>> If the range to match against is [A-_], it will become [a-_] which is an empty range, ord('a') > ord('_'). I think it is simpler to reuse toupper() after the fact as I did.\n>>>\n>>> Anyway maybe I should add a test for that corner case?\n>>\n>> Yeah I was thinking about such a case, but I saw glibc do it... I\n>> guess we just found another bug, at least in compat/fnmatch.c. Yes a\n>> test for it would be great, in case I change my mind 2 years from now\n>> and decide to turn it the other way ;)\n>\n> Should I patch compat/fnmatch.c too? That would make it different from the glibc's one.\n\nNo. I plan to remove compat/fnmatch and always use wildmatch, even\nignoring system's fnmatch. That would keep the matching behavior\nconsistent across platforms.\n--\nDuy\n"},{"id":"218954","messageId":"1369903506-72731-1-git-send-email-n.oxyde@gmail.com","threadId":"33957","inReplyTo":"CACsJy8CuaowyZJGKh7X+43qRwYAdUCDbVo8P5CpEtukBzRiReg@mail.gmail.com","subject":"[PATCH] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-05-30T08:45:06Z","receivedAt":"2013-05-30T08:45:06Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"Case folding is not done correctly when matching against the [:upper:]\ncharacter class and uppercased character ranges (e.g. A-Z).\nSpecifically, an uppercase letter fails to match against any of them\nwhen case folding is requested because plain characters in the pattern\nand the whole string and preemptively lowercased to handle the base case\nfast.\n\nThat optimization is kept and ISLOWER() is used in the [:upper:] case\nwhen case folding is requested, while matching against a character range\nis retried with toupper() if the character was lowercase, as the bounds\nof the range itself cannot be modified (in a case-insensitive context,\n[A-_] is not equivalent to [a-_]).\n\nSigned-off-by: Anthony Ramine <n.oxyde@gmail.com>\n---\n t/t3070-wildmatch.sh | 55 ++++++++++++++++++++++++++++++++++++++++++++++------\n wildmatch.c          |  7 +++++++\n 2 files changed, 56 insertions(+), 6 deletions(-)\n\nI added four tests for the [A-_] range case and a note about it in the\ncommit message.\n\ndiff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nindex 4c37057..38446a0 100755\n--- a/t/t3070-wildmatch.sh\n+++ b/t/t3070-wildmatch.sh\n@@ -6,20 +6,20 @@ test_description='wildmatch tests'\n \n match() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"wildmatch:    match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:     match '$3' '$4'\" \"\n \t    test-wildmatch wildmatch '$3' '$4'\n \t\"\n     else\n-\ttest_expect_success \"wildmatch: no match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:  no match '$3' '$4'\" \"\n \t    ! test-wildmatch wildmatch '$3' '$4'\n \t\"\n     fi\n     if [ $2 = 1 ]; then\n-\ttest_expect_success \"fnmatch:      match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:       match '$3' '$4'\" \"\n \t    test-wildmatch fnmatch '$3' '$4'\n \t\"\n     elif [ $2 = 0 ]; then\n-\ttest_expect_success \"fnmatch:   no match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:    no match '$3' '$4'\" \"\n \t    ! test-wildmatch fnmatch '$3' '$4'\n \t\"\n #    else\n@@ -29,13 +29,25 @@ match() {\n     fi\n }\n \n+imatch() {\n+    if [ $1 = 1 ]; then\n+\ttest_expect_success \"iwildmatch:    match '$2' '$3'\" \"\n+\t    test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    else\n+\ttest_expect_success \"iwildmatch: no match '$2' '$3'\" \"\n+\t    ! test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    fi\n+}\n+\n pathmatch() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"pathmatch:    match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:     match '$2' '$3'\" \"\n \t    test-wildmatch pathmatch '$2' '$3'\n \t\"\n     else\n-\ttest_expect_success \"pathmatch: no match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:  no match '$2' '$3'\" \"\n \t    ! test-wildmatch pathmatch '$2' '$3'\n \t\"\n     fi\n@@ -235,4 +247,35 @@ pathmatch 1 abcXdefXghi '*X*i'\n pathmatch 1 ab/cXd/efXg/hi '*/*X*/*/*i'\n pathmatch 1 ab/cXd/efXg/hi '*Xg*i'\n \n+# Case-sensitivy features\n+match 0 x 'a' '[A-Z]'\n+match 1 x 'A' '[A-Z]'\n+match 0 x 'A' '[a-z]'\n+match 1 x 'a' '[a-z]'\n+match 0 x 'a' '[[:upper:]]'\n+match 1 x 'A' '[[:upper:]]'\n+match 0 x 'A' '[[:lower:]]'\n+match 1 x 'a' '[[:lower:]]'\n+match 0 x 'A' '[B-Za]'\n+match 1 x 'a' '[B-Za]'\n+match 0 x 'A' '[B-a]'\n+match 1 x 'a' '[B-a]'\n+match 0 x 'z' '[Z-y]'\n+match 1 x 'Z' '[Z-y]'\n+\n+imatch 1 'a' '[A-Z]'\n+imatch 1 'A' '[A-Z]'\n+imatch 1 'A' '[a-z]'\n+imatch 1 'a' '[a-z]'\n+imatch 1 'a' '[[:upper:]]'\n+imatch 1 'A' '[[:upper:]]'\n+imatch 1 'A' '[[:lower:]]'\n+imatch 1 'a' '[[:lower:]]'\n+imatch 1 'A' '[B-Za]'\n+imatch 1 'a' '[B-Za]'\n+imatch 1 'A' '[B-a]'\n+imatch 1 'a' '[B-a]'\n+imatch 1 'z' '[Z-y]'\n+imatch 1 'Z' '[Z-y]'\n+\n test_done\ndiff --git a/wildmatch.c b/wildmatch.c\nindex 7192bdc..f91ba99 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t}\n \t\t\t\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n \t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n+\t\t\t\t\t\tuchar t_ch_upper = toupper(t_ch);\n+\t\t\t\t\t\tif (t_ch_upper <= p_ch && t_ch_upper >= prev_ch)\n+\t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t}\n \t\t\t\t\tp_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n \t\t\t\t} else if (p_ch == '[' && p[1] == ':') {\n \t\t\t\t\tconst uchar *s;\n@@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t} else if (CC_EQ(s,i, \"upper\")) {\n \t\t\t\t\t\tif (ISUPPER(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n+\t\t\t\t\t\t\tmatched = 1;\n \t\t\t\t\t} else if (CC_EQ(s,i, \"xdigit\")) {\n \t\t\t\t\t\tif (ISXDIGIT(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n-- \n1.8.3\n"},{"id":"218955","messageId":"CACsJy8DpJKPgc2h0yV3ZHOV6mhEGs=j1NZJ2WQBWG7hAk_iuBw@mail.gmail.com","threadId":"33957","inReplyTo":"1369903506-72731-1-git-send-email-n.oxyde@gmail.com","subject":"Re: [PATCH] wildmatch: properly fold case everywhere","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2013-05-30T08:52:09Z","receivedAt":"2013-05-30T08:52:09Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, May 30, 2013 at 3:45 PM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n> Case folding is not done correctly when matching against the [:upper:]\n> character class and uppercased character ranges (e.g. A-Z).\n> Specifically, an uppercase letter fails to match against any of them\n> when case folding is requested because plain characters in the pattern\n> and the whole string and preemptively lowercased to handle the base case\n> fast.\n>\n> That optimization is kept and ISLOWER() is used in the [:upper:] case\n> when case folding is requested, while matching against a character range\n> is retried with toupper() if the character was lowercase, as the bounds\n> of the range itself cannot be modified (in a case-insensitive context,\n> [A-_] is not equivalent to [a-_]).\n>\n> Signed-off-by: Anthony Ramine <n.oxyde@gmail.com>\n\nReviewed-by: Duy Nguyen <pclouds@gmail.com>\n\nIf you have time, you may want to send a similar patch to rsync, which\ncontains original wildmatch implementation. It does not look much\ndifferent from this one, except that (flags & WM_CASEFOLD) is replaced\nwith force_lower_case. Thanks.\n--\nDuy\n"},{"id":"218957","messageId":"CAPig+cTfaj3e_sRZhHLQUDWYinFVsNieFFA027zJSfdSty1x1g@mail.gmail.com","threadId":"33957","inReplyTo":"1369903506-72731-1-git-send-email-n.oxyde@gmail.com","subject":"Re: [PATCH] wildmatch: properly fold case everywhere","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2013-05-30T09:07:26Z","receivedAt":"2013-05-30T09:07:26Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Thu, May 30, 2013 at 4:45 AM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n> Case folding is not done correctly when matching against the [:upper:]\n> character class and uppercased character ranges (e.g. A-Z).\n> Specifically, an uppercase letter fails to match against any of them\n> when case folding is requested because plain characters in the pattern\n> and the whole string and preemptively lowercased to handle the base case\n\nDid you mean s/and preemptively/are preemptively/ ?\n\n> fast.\n>\n> That optimization is kept and ISLOWER() is used in the [:upper:] case\n> when case folding is requested, while matching against a character range\n> is retried with toupper() if the character was lowercase, as the bounds\n> of the range itself cannot be modified (in a case-insensitive context,\n> [A-_] is not equivalent to [a-_]).\n>\n> Signed-off-by: Anthony Ramine <n.oxyde@gmail.com>\n"},{"id":"218958","messageId":"E670228E-B029-422C-B048-5F28E3AEB731@gmail.com","threadId":"33957","inReplyTo":"CAPig+cTfaj3e_sRZhHLQUDWYinFVsNieFFA027zJSfdSty1x1g@mail.gmail.com","subject":"Re: [PATCH] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-05-30T09:29:25Z","receivedAt":"2013-05-30T09:29:25Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"Yes indeed. Will amend. Should I add your name in Reviewed-by as well?\n\n-- \nAnthony Ramine\n\nLe 30 mai 2013 à 11:07, Eric Sunshine a écrit :\n\n> On Thu, May 30, 2013 at 4:45 AM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n>> Case folding is not done correctly when matching against the [:upper:]\n>> character class and uppercased character ranges (e.g. A-Z).\n>> Specifically, an uppercase letter fails to match against any of them\n>> when case folding is requested because plain characters in the pattern\n>> and the whole string and preemptively lowercased to handle the base case\n> \n> Did you mean s/and preemptively/are preemptively/ ?\n> \n>> fast.\n>> \n>> That optimization is kept and ISLOWER() is used in the [:upper:] case\n>> when case folding is requested, while matching against a character range\n>> is retried with toupper() if the character was lowercase, as the bounds\n>> of the range itself cannot be modified (in a case-insensitive context,\n>> [A-_] is not equivalent to [a-_]).\n>> \n>> Signed-off-by: Anthony Ramine <n.oxyde@gmail.com>\n"},{"id":"218961","messageId":"CAPig+cQZd4Y5gwJkf4NT73xv5hYmNQeZ_kccrzF9_12WbzONJw@mail.gmail.com","threadId":"33957","inReplyTo":"E670228E-B029-422C-B048-5F28E3AEB731@gmail.com","subject":"Re: [PATCH] wildmatch: properly fold case everywhere","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2013-05-30T10:09:42Z","receivedAt":"2013-05-30T10:09:42Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Thu, May 30, 2013 at 5:29 AM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n> Yes indeed. Will amend. Should I add your name in Reviewed-by as well?\n\nNo. I merely spotted a minor typographical error.\n\n> --\n> Anthony Ramine\n>\n> Le 30 mai 2013 à 11:07, Eric Sunshine a écrit :\n>\n>> On Thu, May 30, 2013 at 4:45 AM, Anthony Ramine <n.oxyde@gmail.com> wrote:\n>>> Case folding is not done correctly when matching against the [:upper:]\n>>> character class and uppercased character ranges (e.g. A-Z).\n>>> Specifically, an uppercase letter fails to match against any of them\n>>> when case folding is requested because plain characters in the pattern\n>>> and the whole string and preemptively lowercased to handle the base case\n>>\n>> Did you mean s/and preemptively/are preemptively/ ?\n>>\n>>> fast.\n>>>\n>>> That optimization is kept and ISLOWER() is used in the [:upper:] case\n>>> when case folding is requested, while matching against a character range\n>>> is retried with toupper() if the character was lowercase, as the bounds\n>>> of the range itself cannot be modified (in a case-insensitive context,\n>>> [A-_] is not equivalent to [a-_]).\n>>>\n>>> Signed-off-by: Anthony Ramine <n.oxyde@gmail.com>\n>\n"},{"id":"218963","messageId":"1369909150-73114-1-git-send-email-n.oxyde@gmail.com","threadId":"33957","inReplyTo":"1369903506-72731-1-git-send-email-n.oxyde@gmail.com","subject":"[PATCH v5] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-05-30T10:19:10Z","receivedAt":"2013-05-30T10:19:10Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"Case folding is not done correctly when matching against the [:upper:]\ncharacter class and uppercased character ranges (e.g. A-Z).\nSpecifically, an uppercase letter fails to match against any of them\nwhen case folding is requested because plain characters in the pattern\nand the whole string are preemptively lowercased to handle the base case\nfast.\n\nThat optimization is kept and ISLOWER() is used in the [:upper:] case\nwhen case folding is requested, while matching against a character range\nis retried with toupper() if the character was lowercase, as the bounds\nof the range itself cannot be modified (in a case-insensitive context,\n[A-_] is not equivalent to [a-_]).\n\nSigned-off-by: Anthony Ramine <n.oxyde@gmail.com>\nReviewed-by: Duy Nguyen <pclouds@gmail.com>\n---\n t/t3070-wildmatch.sh | 55 ++++++++++++++++++++++++++++++++++++++++++++++------\n wildmatch.c          |  7 +++++++\n 2 files changed, 56 insertions(+), 6 deletions(-)\n\nI added Duy as reviewer and fixed a typo in the commit message reported by\nEric Sunshine.\n\ndiff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nindex 4c37057..38446a0 100755\n--- a/t/t3070-wildmatch.sh\n+++ b/t/t3070-wildmatch.sh\n@@ -6,20 +6,20 @@ test_description='wildmatch tests'\n \n match() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"wildmatch:    match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:     match '$3' '$4'\" \"\n \t    test-wildmatch wildmatch '$3' '$4'\n \t\"\n     else\n-\ttest_expect_success \"wildmatch: no match '$3' '$4'\" \"\n+\ttest_expect_success \"wildmatch:  no match '$3' '$4'\" \"\n \t    ! test-wildmatch wildmatch '$3' '$4'\n \t\"\n     fi\n     if [ $2 = 1 ]; then\n-\ttest_expect_success \"fnmatch:      match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:       match '$3' '$4'\" \"\n \t    test-wildmatch fnmatch '$3' '$4'\n \t\"\n     elif [ $2 = 0 ]; then\n-\ttest_expect_success \"fnmatch:   no match '$3' '$4'\" \"\n+\ttest_expect_success \"fnmatch:    no match '$3' '$4'\" \"\n \t    ! test-wildmatch fnmatch '$3' '$4'\n \t\"\n #    else\n@@ -29,13 +29,25 @@ match() {\n     fi\n }\n \n+imatch() {\n+    if [ $1 = 1 ]; then\n+\ttest_expect_success \"iwildmatch:    match '$2' '$3'\" \"\n+\t    test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    else\n+\ttest_expect_success \"iwildmatch: no match '$2' '$3'\" \"\n+\t    ! test-wildmatch iwildmatch '$2' '$3'\n+\t\"\n+    fi\n+}\n+\n pathmatch() {\n     if [ $1 = 1 ]; then\n-\ttest_expect_success \"pathmatch:    match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:     match '$2' '$3'\" \"\n \t    test-wildmatch pathmatch '$2' '$3'\n \t\"\n     else\n-\ttest_expect_success \"pathmatch: no match '$2' '$3'\" \"\n+\ttest_expect_success \"pathmatch:  no match '$2' '$3'\" \"\n \t    ! test-wildmatch pathmatch '$2' '$3'\n \t\"\n     fi\n@@ -235,4 +247,35 @@ pathmatch 1 abcXdefXghi '*X*i'\n pathmatch 1 ab/cXd/efXg/hi '*/*X*/*/*i'\n pathmatch 1 ab/cXd/efXg/hi '*Xg*i'\n \n+# Case-sensitivy features\n+match 0 x 'a' '[A-Z]'\n+match 1 x 'A' '[A-Z]'\n+match 0 x 'A' '[a-z]'\n+match 1 x 'a' '[a-z]'\n+match 0 x 'a' '[[:upper:]]'\n+match 1 x 'A' '[[:upper:]]'\n+match 0 x 'A' '[[:lower:]]'\n+match 1 x 'a' '[[:lower:]]'\n+match 0 x 'A' '[B-Za]'\n+match 1 x 'a' '[B-Za]'\n+match 0 x 'A' '[B-a]'\n+match 1 x 'a' '[B-a]'\n+match 0 x 'z' '[Z-y]'\n+match 1 x 'Z' '[Z-y]'\n+\n+imatch 1 'a' '[A-Z]'\n+imatch 1 'A' '[A-Z]'\n+imatch 1 'A' '[a-z]'\n+imatch 1 'a' '[a-z]'\n+imatch 1 'a' '[[:upper:]]'\n+imatch 1 'A' '[[:upper:]]'\n+imatch 1 'A' '[[:lower:]]'\n+imatch 1 'a' '[[:lower:]]'\n+imatch 1 'A' '[B-Za]'\n+imatch 1 'a' '[B-Za]'\n+imatch 1 'A' '[B-a]'\n+imatch 1 'a' '[B-a]'\n+imatch 1 'z' '[Z-y]'\n+imatch 1 'Z' '[Z-y]'\n+\n test_done\ndiff --git a/wildmatch.c b/wildmatch.c\nindex 7192bdc..f91ba99 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t}\n \t\t\t\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n \t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n+\t\t\t\t\t\tuchar t_ch_upper = toupper(t_ch);\n+\t\t\t\t\t\tif (t_ch_upper <= p_ch && t_ch_upper >= prev_ch)\n+\t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t}\n \t\t\t\t\tp_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n \t\t\t\t} else if (p_ch == '[' && p[1] == ':') {\n \t\t\t\t\tconst uchar *s;\n@@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n \t\t\t\t\t} else if (CC_EQ(s,i, \"upper\")) {\n \t\t\t\t\t\tif (ISUPPER(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n+\t\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n+\t\t\t\t\t\t\tmatched = 1;\n \t\t\t\t\t} else if (CC_EQ(s,i, \"xdigit\")) {\n \t\t\t\t\t\tif (ISXDIGIT(t_ch))\n \t\t\t\t\t\t\tmatched = 1;\n-- \n1.8.3\n"},{"id":"219193","messageId":"7vppw4mb77.fsf@alter.siamese.dyndns.org","threadId":"33957","inReplyTo":"1369909150-73114-1-git-send-email-n.oxyde@gmail.com","subject":"Re: [PATCH v5] wildmatch: properly fold case everywhere","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-06-02T21:53:16Z","receivedAt":"2013-06-02T21:53:16Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Anthony Ramine <n.oxyde@gmail.com> writes:\n\n> ase folding is not done correctly when matching against the [:upper:]\n> character class and uppercased character ranges (e.g. A-Z).\n> Specifically, an uppercase letter fails to match against any of them\n> when case folding is requested because plain characters in the pattern\n> and the whole string are preemptively lowercased to handle the base case\n> fast.\n>\n> That optimization is kept and ISLOWER() is used in the [:upper:] case\n> when case folding is requested, while matching against a character range\n> is retried with toupper() if the character was lowercase, as the bounds\n> of the range itself cannot be modified (in a case-insensitive context,\n> [A-_] is not equivalent to [a-_]).\n>\n> Signed-off-by: Anthony Ramine <n.oxyde@gmail.com>\n> Reviewed-by: Duy Nguyen <pclouds@gmail.com>\n> ---\n\nThanks.\n\n>  t/t3070-wildmatch.sh | 55 ++++++++++++++++++++++++++++++++++++++++++++++------\n>  wildmatch.c          |  7 +++++++\n>  2 files changed, 56 insertions(+), 6 deletions(-)\n>\n> I added Duy as reviewer and fixed a typo in the commit message reported by\n> Eric Sunshine.\n>\n> diff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\n> index 4c37057..38446a0 100755\n> --- a/t/t3070-wildmatch.sh\n> +++ b/t/t3070-wildmatch.sh\n> @@ -6,20 +6,20 @@ test_description='wildmatch tests'\n>  \n>  match() {\n>      if [ $1 = 1 ]; then\n> -\ttest_expect_success \"wildmatch:    match '$3' '$4'\" \"\n> +\ttest_expect_success \"wildmatch:     match '$3' '$4'\" \"\n>  \t    test-wildmatch wildmatch '$3' '$4'\n>  \t\"\n>      else\n> -\ttest_expect_success \"wildmatch: no match '$3' '$4'\" \"\n> +\ttest_expect_success \"wildmatch:  no match '$3' '$4'\" \"\n>  \t    ! test-wildmatch wildmatch '$3' '$4'\n>  \t\"\n>      fi\n>      if [ $2 = 1 ]; then\n> -\ttest_expect_success \"fnmatch:      match '$3' '$4'\" \"\n> +\ttest_expect_success \"fnmatch:       match '$3' '$4'\" \"\n>  \t    test-wildmatch fnmatch '$3' '$4'\n>  \t\"\n>      elif [ $2 = 0 ]; then\n> -\ttest_expect_success \"fnmatch:   no match '$3' '$4'\" \"\n> +\ttest_expect_success \"fnmatch:    no match '$3' '$4'\" \"\n>  \t    ! test-wildmatch fnmatch '$3' '$4'\n>  \t\"\n>  #    else\n\nIs the above about aligning $3/$4 across checks of different types\n(i.e. purely cosmetic)?  I am not complaining; just making sure if\nthere is nothing deeper going on.\n\nIt is outside the scope of this change, but the shell style of this\nscript (most notably use of [] instead of test) needs to be fixed\nsomeday, preferrably soon, including the commented out else clause\nat the end of the hunk.\n\n> @@ -235,4 +247,35 @@ pathmatch 1 abcXdefXghi '*X*i'\n>  pathmatch 1 ab/cXd/efXg/hi '*/*X*/*/*i'\n>  pathmatch 1 ab/cXd/efXg/hi '*Xg*i'\n>  \n> +# Case-sensitivy features\n> +match 0 x 'a' '[A-Z]'\n> +match 1 x 'A' '[A-Z]'\n> +match 0 x 'A' '[a-z]'\n> +match 1 x 'a' '[a-z]'\n> +match 0 x 'a' '[[:upper:]]'\n> +match 1 x 'A' '[[:upper:]]'\n> +match 0 x 'A' '[[:lower:]]'\n> +match 1 x 'a' '[[:lower:]]'\n> +match 0 x 'A' '[B-Za]'\n> +match 1 x 'a' '[B-Za]'\n> +match 0 x 'A' '[B-a]'\n> +match 1 x 'a' '[B-a]'\n> +match 0 x 'z' '[Z-y]'\n> +match 1 x 'Z' '[Z-y]'\n> +\n> +imatch 1 'a' '[A-Z]'\n\nDo we want \"# Case-insensitivity features\" commment here as well?\n\n> +imatch 1 'A' '[A-Z]'\n> +imatch 1 'A' '[a-z]'\n> +imatch 1 'a' '[a-z]'\n> +imatch 1 'a' '[[:upper:]]'\n> +imatch 1 'A' '[[:upper:]]'\n> +imatch 1 'A' '[[:lower:]]'\n> +imatch 1 'a' '[[:lower:]]'\n> +imatch 1 'A' '[B-Za]'\n> +imatch 1 'a' '[B-Za]'\n> +imatch 1 'A' '[B-a]'\n> +imatch 1 'a' '[B-a]'\n> +imatch 1 'z' '[Z-y]'\n> +imatch 1 'Z' '[Z-y]'\n\n> +\n>  test_done\n> diff --git a/wildmatch.c b/wildmatch.c\n> index 7192bdc..f91ba99 100644\n> --- a/wildmatch.c\n> +++ b/wildmatch.c\n> @@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>  \t\t\t\t\t}\n>  \t\t\t\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n>  \t\t\t\t\t\tmatched = 1;\n> +\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n> +\t\t\t\t\t\tuchar t_ch_upper = toupper(t_ch);\n> +\t\t\t\t\t\tif (t_ch_upper <= p_ch && t_ch_upper >= prev_ch)\n> +\t\t\t\t\t\t\tmatched = 1;\n> +\t\t\t\t\t}\n>  \t\t\t\t\tp_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n\nHmm, this looks somewhat strange.\n\n * At the beginning of the outermost \"per characters in the text\"\n   loop, we seem to downcase t_ch when WM_CASEFOLD is set.\n * Also at the same place, we also seem to downcase p_ch under the\n   same condition.\n\nwhich makes me wonder why the fix is not like this:\n\n\t+\tif (flags & WM_CASEFOLD) {\n        +\t\tif (ISUPPER(p_ch))\n        +\t\t\tp_ch = tolower(p_ch);\n        +\t\tif (prev_ch && ISUPPER(prev_ch))\n        +\t\t\tprev_ch = tolower(prev_ch);\n\t+\t}\n\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n\t\t\tmatched = 1;\n\t\tp_ch = 0; /* This sets \"prev_ch\" to 0 */\n\n\nAhh, OK, the \"seemingly strange\" construct is about handling a range\nlike \"[Z-y]\"; we do not want to upcase or downcase the p_ch/prev_ch\nmake the range \"[z-y]\" (empty) which would exclude bytes like \"^\",\n\"_\" or even \"Z\".\n\nAnd it is also OK to downcase p_ch in a single-letter case, not the\ncharacters in a range, at the beginning of the outermost loop,\nbecause we always compare for equality against t_ch (which is\ndowncased) in that case.\n\n> @@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>  \t\t\t\t\t} else if (CC_EQ(s,i, \"upper\")) {\n>  \t\t\t\t\t\tif (ISUPPER(t_ch))\n>  \t\t\t\t\t\t\tmatched = 1;\n> +\t\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n> +\t\t\t\t\t\t\tmatched = 1;\n\nThis also looks somewhat strange but correct in that t_ch is already\ndowncased so we do not need a corresponding change for CC_EQ(\"lower\")\ncodepath.\n\nInteresting.  Will apply.\n\nThanks.\n"},{"id":"219203","messageId":"586F64C2-0F44-4DAB-B91A-DB88C624FEC2@gmail.com","threadId":"33957","inReplyTo":"7vppw4mb77.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH v5] wildmatch: properly fold case everywhere","fromName":"Anthony Ramine","fromEmail":"n.oxyde@gmail.com","sentAt":"2013-06-02T23:42:51Z","receivedAt":"2013-06-02T23:42:51Z","isPatch":true,"sender":{"key":"n.oxyde@gmail.com","avatar":"https://avatars.githubusercontent.com/u/123095?v=4"},"body":"Hello Junio,\n\nReplied inline.\n\nRegards,\n\n-- \nAnthony Ramine\n\nLe 2 juin 2013 à 23:53, Junio C Hamano a écrit :\n\n> Anthony Ramine <n.oxyde@gmail.com> writes:\n> \n>> ase folding is not done correctly when matching against the [:upper:]\n>> character class and uppercased character ranges (e.g. A-Z).\n>> Specifically, an uppercase letter fails to match against any of them\n>> when case folding is requested because plain characters in the pattern\n>> and the whole string are preemptively lowercased to handle the base case\n>> fast.\n>> \n>> That optimization is kept and ISLOWER() is used in the [:upper:] case\n>> when case folding is requested, while matching against a character range\n>> is retried with toupper() if the character was lowercase, as the bounds\n>> of the range itself cannot be modified (in a case-insensitive context,\n>> [A-_] is not equivalent to [a-_]).\n>> \n>> Signed-off-by: Anthony Ramine <n.oxyde@gmail.com>\n>> Reviewed-by: Duy Nguyen <pclouds@gmail.com>\n>> ---\n> \n> Thanks.\n> \n>> t/t3070-wildmatch.sh | 55 ++++++++++++++++++++++++++++++++++++++++++++++------\n>> wildmatch.c          |  7 +++++++\n>> 2 files changed, 56 insertions(+), 6 deletions(-)\n>> \n>> I added Duy as reviewer and fixed a typo in the commit message reported by\n>> Eric Sunshine.\n>> \n>> diff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\n>> index 4c37057..38446a0 100755\n>> --- a/t/t3070-wildmatch.sh\n>> +++ b/t/t3070-wildmatch.sh\n>> @@ -6,20 +6,20 @@ test_description='wildmatch tests'\n>> \n>> match() {\n>>     if [ $1 = 1 ]; then\n>> -\ttest_expect_success \"wildmatch:    match '$3' '$4'\" \"\n>> +\ttest_expect_success \"wildmatch:     match '$3' '$4'\" \"\n>> \t    test-wildmatch wildmatch '$3' '$4'\n>> \t\"\n>>     else\n>> -\ttest_expect_success \"wildmatch: no match '$3' '$4'\" \"\n>> +\ttest_expect_success \"wildmatch:  no match '$3' '$4'\" \"\n>> \t    ! test-wildmatch wildmatch '$3' '$4'\n>> \t\"\n>>     fi\n>>     if [ $2 = 1 ]; then\n>> -\ttest_expect_success \"fnmatch:      match '$3' '$4'\" \"\n>> +\ttest_expect_success \"fnmatch:       match '$3' '$4'\" \"\n>> \t    test-wildmatch fnmatch '$3' '$4'\n>> \t\"\n>>     elif [ $2 = 0 ]; then\n>> -\ttest_expect_success \"fnmatch:   no match '$3' '$4'\" \"\n>> +\ttest_expect_success \"fnmatch:    no match '$3' '$4'\" \"\n>> \t    ! test-wildmatch fnmatch '$3' '$4'\n>> \t\"\n>> #    else\n> \n> Is the above about aligning $3/$4 across checks of different types\n> (i.e. purely cosmetic)?  I am not complaining; just making sure if\n> there is nothing deeper going on.\n\nYes purely cosmetic.\n\n> It is outside the scope of this change, but the shell style of this\n> script (most notably use of [] instead of test) needs to be fixed\n> someday, preferrably soon, including the commented out else clause\n> at the end of the hunk.\n> \n>> @@ -235,4 +247,35 @@ pathmatch 1 abcXdefXghi '*X*i'\n>> pathmatch 1 ab/cXd/efXg/hi '*/*X*/*/*i'\n>> pathmatch 1 ab/cXd/efXg/hi '*Xg*i'\n>> \n>> +# Case-sensitivy features\n>> +match 0 x 'a' '[A-Z]'\n>> +match 1 x 'A' '[A-Z]'\n>> +match 0 x 'A' '[a-z]'\n>> +match 1 x 'a' '[a-z]'\n>> +match 0 x 'a' '[[:upper:]]'\n>> +match 1 x 'A' '[[:upper:]]'\n>> +match 0 x 'A' '[[:lower:]]'\n>> +match 1 x 'a' '[[:lower:]]'\n>> +match 0 x 'A' '[B-Za]'\n>> +match 1 x 'a' '[B-Za]'\n>> +match 0 x 'A' '[B-a]'\n>> +match 1 x 'a' '[B-a]'\n>> +match 0 x 'z' '[Z-y]'\n>> +match 1 x 'Z' '[Z-y]'\n>> +\n>> +imatch 1 'a' '[A-Z]'\n> \n> Do we want \"# Case-insensitivity features\" commment here as well?\n\nI don't particularly care, some sections have titles, some don't.\n\n>> +imatch 1 'A' '[A-Z]'\n>> +imatch 1 'A' '[a-z]'\n>> +imatch 1 'a' '[a-z]'\n>> +imatch 1 'a' '[[:upper:]]'\n>> +imatch 1 'A' '[[:upper:]]'\n>> +imatch 1 'A' '[[:lower:]]'\n>> +imatch 1 'a' '[[:lower:]]'\n>> +imatch 1 'A' '[B-Za]'\n>> +imatch 1 'a' '[B-Za]'\n>> +imatch 1 'A' '[B-a]'\n>> +imatch 1 'a' '[B-a]'\n>> +imatch 1 'z' '[Z-y]'\n>> +imatch 1 'Z' '[Z-y]'\n> \n>> +\n>> test_done\n>> diff --git a/wildmatch.c b/wildmatch.c\n>> index 7192bdc..f91ba99 100644\n>> --- a/wildmatch.c\n>> +++ b/wildmatch.c\n>> @@ -196,6 +196,11 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>> \t\t\t\t\t}\n>> \t\t\t\t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n>> \t\t\t\t\t\tmatched = 1;\n>> +\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch)) {\n>> +\t\t\t\t\t\tuchar t_ch_upper = toupper(t_ch);\n>> +\t\t\t\t\t\tif (t_ch_upper <= p_ch && t_ch_upper >= prev_ch)\n>> +\t\t\t\t\t\t\tmatched = 1;\n>> +\t\t\t\t\t}\n>> \t\t\t\t\tp_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n> \n> Hmm, this looks somewhat strange.\n> \n> * At the beginning of the outermost \"per characters in the text\"\n>   loop, we seem to downcase t_ch when WM_CASEFOLD is set.\n> * Also at the same place, we also seem to downcase p_ch under the\n>   same condition.\n> \n> which makes me wonder why the fix is not like this:\n> \n> \t+\tif (flags & WM_CASEFOLD) {\n>        +\t\tif (ISUPPER(p_ch))\n>        +\t\t\tp_ch = tolower(p_ch);\n>        +\t\tif (prev_ch && ISUPPER(prev_ch))\n>        +\t\t\tprev_ch = tolower(prev_ch);\n> \t+\t}\n> \t\tif (t_ch <= p_ch && t_ch >= prev_ch)\n> \t\t\tmatched = 1;\n> \t\tp_ch = 0; /* This sets \"prev_ch\" to 0 */\n> \n> \n> Ahh, OK, the \"seemingly strange\" construct is about handling a range\n> like \"[Z-y]\"; we do not want to upcase or downcase the p_ch/prev_ch\n> make the range \"[z-y]\" (empty) which would exclude bytes like \"^\",\n> \"_\" or even \"Z\".\n> \n> And it is also OK to downcase p_ch in a single-letter case, not the\n> characters in a range, at the beginning of the outermost loop,\n> because we always compare for equality against t_ch (which is\n> downcased) in that case.\n\nYeah I tried to explain that in the commit message but it is surely too dense.\n\n>> @@ -245,6 +250,8 @@ static int dowild(const uchar *p, const uchar *text, unsigned int flags)\n>> \t\t\t\t\t} else if (CC_EQ(s,i, \"upper\")) {\n>> \t\t\t\t\t\tif (ISUPPER(t_ch))\n>> \t\t\t\t\t\t\tmatched = 1;\n>> +\t\t\t\t\t\telse if ((flags & WM_CASEFOLD) && ISLOWER(t_ch))\n>> +\t\t\t\t\t\t\tmatched = 1;\n> \n> This also looks somewhat strange but correct in that t_ch is already\n> downcased so we do not need a corresponding change for CC_EQ(\"lower\")\n> codepath.\n\nI chose to still lowercase everything first, to keep the common path as is.\n\n> Interesting.  Will apply.\n> \n> Thanks.\n\nGreat. You're welcome.\n"}]}