{"thread":{"id":"31810","subject":"[PATCH v5 00/12] nd/wildmatch","startedAt":"2012-10-14T02:34:58Z","lastAt":"2012-11-16T04:19:32Z","messageCount":37,"participants":["Nguyễn Thái Ngọc Duy","Junio C Hamano","Nguyen Thai Ngoc Duy","Torsten Bögershausen","René Scharfe","Jan H. Schönherr","Johannes Sixt","Linus Torvalds"],"isPatch":true,"patchVersion":5,"patchTotal":12},"messages":[{"id":"201100","messageId":"1350182110-25936-1-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":null,"subject":"[PATCH v5 00/12] nd/wildmatch","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:34:58Z","receivedAt":"2012-10-14T02:34:58Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"This version splits fnmatch/wildmatch tests separately in t3070 and\ndisables a lot more fnmatch tests. It also fixes the \"cd ..\" in t0003\ntest and a comment in front of dowild(). No functional changes.\n\nNguyễn Thái Ngọc Duy (12):\n  ctype: make sane_ctype[] const array\n  ctype: support iscntrl, ispunct, isxdigit and isprint\n  Import wildmatch from rsync\n  wildmatch: remove unnecessary functions\n  Integrate wildmatch to git\n  t3070: disable unreliable fnmatch tests\n  wildmatch: make wildmatch's return value compatible with fnmatch\n  wildmatch: remove static variable force_lower_case\n  wildmatch: fix case-insensitive matching\n  wildmatch: adjust \"**\" behavior\n  wildmatch: make /**/ match zero or more directories\n  Support \"**\" wildcard in .gitignore and .gitattributes\n\n .gitignore                         |   1 +\n Documentation/gitignore.txt        |  19 +++\n Makefile                           |   3 +\n attr.c                             |   4 +-\n ctype.c                            |  20 +++-\n dir.c                              |   4 +-\n git-compat-util.h                  |  15 ++-\n t/t0003-attributes.sh              |  37 ++++++\n t/t3001-ls-files-others-exclude.sh |  19 +++\n t/t3070-wildmatch.sh               | 195 ++++++++++++++++++++++++++++++\n test-wildmatch.c                   |  14 +++\n wildmatch.c                        | 239 +++++++++++++++++++++++++++++++++++++\n wildmatch.h                        |   9 ++\n 13 files changed, 575 insertions(+), 4 deletions(-)\n create mode 100755 t/t3070-wildmatch.sh\n create mode 100644 test-wildmatch.c\n create mode 100644 wildmatch.c\n create mode 100644 wildmatch.h\n\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201101","messageId":"1350182110-25936-2-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 01/12] ctype: make sane_ctype[] const array","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:34:59Z","receivedAt":"2012-10-14T02:34:59Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n ctype.c           | 2 +-\n git-compat-util.h | 2 +-\n 2 files changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/ctype.c b/ctype.c\nindex 9353271..faeaf34 100644\n--- a/ctype.c\n+++ b/ctype.c\n@@ -14,7 +14,7 @@ enum {\n \tP = GIT_PATHSPEC_MAGIC  /* other non-alnum, except for ] and } */\n };\n \n-unsigned char sane_ctype[256] = {\n+const unsigned char sane_ctype[256] = {\n \t0, 0, 0, 0, 0, 0, 0, 0, 0, S, S, 0, 0, S, 0, 0,\t\t/*   0.. 15 */\n \t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\t\t/*  16.. 31 */\n \tS, P, P, P, R, P, P, P, R, R, G, R, P, P, R, P,\t\t/*  32.. 47 */\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 2fbf1fd..f8b859c 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -510,7 +510,7 @@ extern const char tolower_trans_tbl[256];\n #undef isupper\n #undef tolower\n #undef toupper\n-extern unsigned char sane_ctype[256];\n+extern const unsigned char sane_ctype[256];\n #define GIT_SPACE 0x01\n #define GIT_DIGIT 0x02\n #define GIT_ALPHA 0x04\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201102","messageId":"1350182110-25936-3-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 02/12] ctype: support iscntrl, ispunct, isxdigit and isprint","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:00Z","receivedAt":"2012-10-14T02:35:00Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n ctype.c           | 18 ++++++++++++++++++\n git-compat-util.h | 13 +++++++++++++\n 2 files changed, 31 insertions(+)\n\ndiff --git a/ctype.c b/ctype.c\nindex faeaf34..b4bf48a 100644\n--- a/ctype.c\n+++ b/ctype.c\n@@ -26,6 +26,24 @@ const unsigned char sane_ctype[256] = {\n \t/* Nothing in the 128.. range */\n };\n \n+enum {\n+\tCN = GIT_CNTRL,\n+\tPU = GIT_PUNCT,\n+\tXD = GIT_XDIGIT,\n+};\n+\n+const unsigned char sane_ctype2[256] = {\n+\tCN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, /*    0..15 */\n+\tCN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, /*   16..31 */\n+\t0,  PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, /*   32..47 */\n+\tXD, XD, XD, XD, XD, XD, XD, XD, XD, XD, PU, PU, PU, PU, PU, PU, /*   48..63 */\n+\tPU, 0,\tXD, 0,\tXD, 0,\tXD, 0,\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t/*   64..79 */\n+\t0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  PU, PU, PU, PU, PU, /*   80..95 */\n+\tPU, 0,\tXD, 0,\tXD, 0,\tXD, 0,\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t/*  96..111 */\n+\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t0,  PU, PU, PU, PU, CN, /* 112..127 */\n+\t/* Nothing in the 128.. range */\n+};\n+\n /* For case-insensitive kwset */\n const char tolower_trans_tbl[256] = {\n \t0x00, 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07,\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex f8b859c..ea11694 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -510,14 +510,23 @@ extern const char tolower_trans_tbl[256];\n #undef isupper\n #undef tolower\n #undef toupper\n+#undef iscntrl\n+#undef ispunct\n+#undef isxdigit\n+#undef isprint\n extern const unsigned char sane_ctype[256];\n+extern const unsigned char sane_ctype2[256];\n #define GIT_SPACE 0x01\n #define GIT_DIGIT 0x02\n #define GIT_ALPHA 0x04\n #define GIT_GLOB_SPECIAL 0x08\n #define GIT_REGEX_SPECIAL 0x10\n #define GIT_PATHSPEC_MAGIC 0x20\n+#define GIT_CNTRL 0x01\n+#define GIT_PUNCT 0x02\n+#define GIT_XDIGIT 0x04\n #define sane_istest(x,mask) ((sane_ctype[(unsigned char)(x)] & (mask)) != 0)\n+#define sane_istest2(x,mask) ((sane_ctype2[(unsigned char)(x)] & (mask)) != 0)\n #define isascii(x) (((x) & ~0x7f) == 0)\n #define isspace(x) sane_istest(x,GIT_SPACE)\n #define isdigit(x) sane_istest(x,GIT_DIGIT)\n@@ -527,6 +536,10 @@ extern const unsigned char sane_ctype[256];\n #define isupper(x) sane_iscase(x, 0)\n #define is_glob_special(x) sane_istest(x,GIT_GLOB_SPECIAL)\n #define is_regex_special(x) sane_istest(x,GIT_GLOB_SPECIAL | GIT_REGEX_SPECIAL)\n+#define iscntrl(x) sane_istest2(x, GIT_CNTRL)\n+#define ispunct(x) sane_istest2(x, GIT_PUNCT)\n+#define isxdigit(x) sane_istest2(x, GIT_XDIGIT)\n+#define isprint(x) (isalnum(x) || isspace(x) || ispunct(x))\n #define tolower(x) sane_case((unsigned char)(x), 0x20)\n #define toupper(x) sane_case((unsigned char)(x), 0)\n #define is_pathspec_magic(x) sane_istest(x,GIT_PATHSPEC_MAGIC)\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201103","messageId":"1350182110-25936-4-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 03/12] Import wildmatch from rsync","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:01Z","receivedAt":"2012-10-14T02:35:01Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"These files are from rsync.git commit\nf92f5b166e3019db42bc7fe1aa2f1a9178cd215d, which was the last commit\nbefore rsync turned GPL-3. All files are imported as-is and\nno-op. Adaptation is done in a separate patch.\n\nrsync.git           ->  git.git\nlib/wildmatch.[ch]      wildmatch.[ch]\nwildtest.txt            t/t3070/wildtest.txt\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n t/t3070/wildtest.txt | 165 +++++++++++++++++++++++\n wildmatch.c          | 368 +++++++++++++++++++++++++++++++++++++++++++++++++++\n wildmatch.h          |   6 +\n 3 files changed, 539 insertions(+)\n create mode 100644 t/t3070/wildtest.txt\n create mode 100644 wildmatch.c\n create mode 100644 wildmatch.h\n\ndiff --git a/t/t3070/wildtest.txt b/t/t3070/wildtest.txt\nnew file mode 100644\nindex 0000000..42c1678\n--- /dev/null\n+++ b/t/t3070/wildtest.txt\n@@ -0,0 +1,165 @@\n+# Input is in the following format (all items white-space separated):\n+#\n+# The first two items are 1 or 0 indicating if the wildmat call is expected to\n+# succeed and if fnmatch works the same way as wildmat, respectively.  After\n+# that is a text string for the match, and a pattern string.  Strings can be\n+# quoted (if desired) in either double or single quotes, as well as backticks.\n+#\n+# MATCH FNMATCH_SAME \"text to match\" 'pattern to use'\n+\n+# Basic wildmat features\n+1 1 foo\t\t\tfoo\n+0 1 foo\t\t\tbar\n+1 1 ''\t\t\t\"\"\n+1 1 foo\t\t\t???\n+0 1 foo\t\t\t??\n+1 1 foo\t\t\t*\n+1 1 foo\t\t\tf*\n+0 1 foo\t\t\t*f\n+1 1 foo\t\t\t*foo*\n+1 1 foobar\t\t*ob*a*r*\n+1 1 aaaaaaabababab\t*ab\n+1 1 foo*\t\tfoo\\*\n+0 1 foobar\t\tfoo\\*bar\n+1 1 f\\oo\t\tf\\\\oo\n+1 1 ball\t\t*[al]?\n+0 1 ten\t\t\t[ten]\n+1 1 ten\t\t\t**[!te]\n+0 1 ten\t\t\t**[!ten]\n+1 1 ten\t\t\tt[a-g]n\n+0 1 ten\t\t\tt[!a-g]n\n+1 1 ton\t\t\tt[!a-g]n\n+1 1 ton\t\t\tt[^a-g]n\n+1 1 a]b\t\t\ta[]]b\n+1 1 a-b\t\t\ta[]-]b\n+1 1 a]b\t\t\ta[]-]b\n+0 1 aab\t\t\ta[]-]b\n+1 1 aab\t\t\ta[]a-]b\n+1 1 ]\t\t\t]\n+\n+# Extended slash-matching features\n+0 1 foo/baz/bar\t\tfoo*bar\n+1 1 foo/baz/bar\t\tfoo**bar\n+0 1 foo/bar\t\tfoo?bar\n+0 1 foo/bar\t\tfoo[/]bar\n+0 1 foo/bar\t\tf[^eiu][^eiu][^eiu][^eiu][^eiu]r\n+1 1 foo-bar\t\tf[^eiu][^eiu][^eiu][^eiu][^eiu]r\n+0 1 foo\t\t\t**/foo\n+1 1 /foo\t\t**/foo\n+1 1 bar/baz/foo\t\t**/foo\n+0 1 bar/baz/foo\t\t*/foo\n+0 0 foo/bar/baz\t\t**/bar*\n+1 1 deep/foo/bar/baz\t**/bar/*\n+0 1 deep/foo/bar/baz/\t**/bar/*\n+1 1 deep/foo/bar/baz/\t**/bar/**\n+0 1 deep/foo/bar\t**/bar/*\n+1 1 deep/foo/bar/\t**/bar/**\n+1 1 foo/bar/baz\t\t**/bar**\n+1 1 foo/bar/baz/x\t*/bar/**\n+0 0 deep/foo/bar/baz/x\t*/bar/**\n+1 1 deep/foo/bar/baz/x\t**/bar/*/*\n+\n+# Various additional tests\n+0 1 acrt\t\ta[c-c]st\n+1 1 acrt\t\ta[c-c]rt\n+0 1 ]\t\t\t[!]-]\n+1 1 a\t\t\t[!]-]\n+0 1 ''\t\t\t\\\n+0 1 \\\t\t\t\\\n+0 1 /\\\t\t\t*/\\\n+1 1 /\\\t\t\t*/\\\\\n+1 1 foo\t\t\tfoo\n+1 1 @foo\t\t@foo\n+0 1 foo\t\t\t@foo\n+1 1 [ab]\t\t\\[ab]\n+1 1 [ab]\t\t[[]ab]\n+1 1 [ab]\t\t[[:]ab]\n+0 1 [ab]\t\t[[::]ab]\n+1 1 [ab]\t\t[[:digit]ab]\n+1 1 [ab]\t\t[\\[:]ab]\n+1 1 ?a?b\t\t\\??\\?b\n+1 1 abc\t\t\t\\a\\b\\c\n+0 1 foo\t\t\t''\n+1 1 foo/bar/baz/to\t**/t[o]\n+\n+# Character class tests\n+1 1 a1B\t\t[[:alpha:]][[:digit:]][[:upper:]]\n+0 1 a\t\t[[:digit:][:upper:][:space:]]\n+1 1 A\t\t[[:digit:][:upper:][:space:]]\n+1 1 1\t\t[[:digit:][:upper:][:space:]]\n+0 1 1\t\t[[:digit:][:upper:][:spaci:]]\n+1 1 ' '\t\t[[:digit:][:upper:][:space:]]\n+0 1 .\t\t[[:digit:][:upper:][:space:]]\n+1 1 .\t\t[[:digit:][:punct:][:space:]]\n+1 1 5\t\t[[:xdigit:]]\n+1 1 f\t\t[[:xdigit:]]\n+1 1 D\t\t[[:xdigit:]]\n+1 1 _\t\t[[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]\n+#1 1 �\t\t[^[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]\n+1 1 \t\t[^[:alnum:][:alpha:][:blank:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]\n+1 1 .\t\t[^[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:lower:][:space:][:upper:][:xdigit:]]\n+1 1 5\t\t[a-c[:digit:]x-z]\n+1 1 b\t\t[a-c[:digit:]x-z]\n+1 1 y\t\t[a-c[:digit:]x-z]\n+0 1 q\t\t[a-c[:digit:]x-z]\n+\n+# Additional tests, including some malformed wildmats\n+1 1 ]\t\t[\\\\-^]\n+0 1 [\t\t[\\\\-^]\n+1 1 -\t\t[\\-_]\n+1 1 ]\t\t[\\]]\n+0 1 \\]\t\t[\\]]\n+0 1 \\\t\t[\\]]\n+0 1 ab\t\ta[]b\n+0 1 a[]b\ta[]b\n+0 1 ab[\t\tab[\n+0 1 ab\t\t[!\n+0 1 ab\t\t[-\n+1 1 -\t\t[-]\n+0 1 -\t\t[a-\n+0 1 -\t\t[!a-\n+1 1 -\t\t[--A]\n+1 1 5\t\t[--A]\n+1 1 ' '\t\t'[ --]'\n+1 1 $\t\t'[ --]'\n+1 1 -\t\t'[ --]'\n+0 1 0\t\t'[ --]'\n+1 1 -\t\t[---]\n+1 1 -\t\t[------]\n+0 1 j\t\t[a-e-n]\n+1 1 -\t\t[a-e-n]\n+1 1 a\t\t[!------]\n+0 1 [\t\t[]-a]\n+1 1 ^\t\t[]-a]\n+0 1 ^\t\t[!]-a]\n+1 1 [\t\t[!]-a]\n+1 1 ^\t\t[a^bc]\n+1 1 -b]\t\t[a-]b]\n+0 1 \\\t\t[\\]\n+1 1 \\\t\t[\\\\]\n+0 1 \\\t\t[!\\\\]\n+1 1 G\t\t[A-\\\\]\n+0 1 aaabbb\tb*a\n+0 1 aabcaa\t*ba*\n+1 1 ,\t\t[,]\n+1 1 ,\t\t[\\\\,]\n+1 1 \\\t\t[\\\\,]\n+1 1 -\t\t[,-.]\n+0 1 +\t\t[,-.]\n+0 1 -.]\t\t[,-.]\n+1 1 2\t\t[\\1-\\3]\n+1 1 3\t\t[\\1-\\3]\n+0 1 4\t\t[\\1-\\3]\n+1 1 \\\t\t[[-\\]]\n+1 1 [\t\t[[-\\]]\n+1 1 ]\t\t[[-\\]]\n+0 1 -\t\t[[-\\]]\n+\n+# Test recursion and the abort code (use \"wildtest -i\" to see iteration counts)\n+1 1 -adobe-courier-bold-o-normal--12-120-75-75-m-70-iso8859-1\t-*-*-*-*-*-*-12-*-*-*-m-*-*-*\n+0 1 -adobe-courier-bold-o-normal--12-120-75-75-X-70-iso8859-1\t-*-*-*-*-*-*-12-*-*-*-m-*-*-*\n+0 1 -adobe-courier-bold-o-normal--12-120-75-75-/-70-iso8859-1\t-*-*-*-*-*-*-12-*-*-*-m-*-*-*\n+1 1 /adobe/courier/bold/o/normal//12/120/75/75/m/70/iso8859/1\t/*/*/*/*/*/*/12/*/*/*/m/*/*/*\n+0 1 /adobe/courier/bold/o/normal//12/120/75/75/X/70/iso8859/1\t/*/*/*/*/*/*/12/*/*/*/m/*/*/*\n+1 1 abcd/abcdefg/abcdefghijk/abcdefghijklmnop.txt\t\t**/*a*b*g*n*t\n+0 1 abcd/abcdefg/abcdefghijk/abcdefghijklmnop.txtz\t\t**/*a*b*g*n*t\ndiff --git a/wildmatch.c b/wildmatch.c\nnew file mode 100644\nindex 0000000..f3a1731\n--- /dev/null\n+++ b/wildmatch.c\n@@ -0,0 +1,368 @@\n+/*\n+**  Do shell-style pattern matching for ?, \\, [], and * characters.\n+**  It is 8bit clean.\n+**\n+**  Written by Rich $alz, mirror!rs, Wed Nov 26 19:03:17 EST 1986.\n+**  Rich $alz is now <rsalz@bbn.com>.\n+**\n+**  Modified by Wayne Davison to special-case '/' matching, to make '**'\n+**  work differently than '*', and to fix the character-class code.\n+*/\n+\n+#include \"rsync.h\"\n+\n+/* What character marks an inverted character class? */\n+#define NEGATE_CLASS\t'!'\n+#define NEGATE_CLASS2\t'^'\n+\n+#define FALSE 0\n+#define TRUE 1\n+#define ABORT_ALL -1\n+#define ABORT_TO_STARSTAR -2\n+\n+#define CC_EQ(class, len, litmatch) ((len) == sizeof (litmatch)-1 \\\n+\t\t\t\t    && *(class) == *(litmatch) \\\n+\t\t\t\t    && strncmp((char*)class, litmatch, len) == 0)\n+\n+#if defined STDC_HEADERS || !defined isascii\n+# define ISASCII(c) 1\n+#else\n+# define ISASCII(c) isascii(c)\n+#endif\n+\n+#ifdef isblank\n+# define ISBLANK(c) (ISASCII(c) && isblank(c))\n+#else\n+# define ISBLANK(c) ((c) == ' ' || (c) == '\\t')\n+#endif\n+\n+#ifdef isgraph\n+# define ISGRAPH(c) (ISASCII(c) && isgraph(c))\n+#else\n+# define ISGRAPH(c) (ISASCII(c) && isprint(c) && !isspace(c))\n+#endif\n+\n+#define ISPRINT(c) (ISASCII(c) && isprint(c))\n+#define ISDIGIT(c) (ISASCII(c) && isdigit(c))\n+#define ISALNUM(c) (ISASCII(c) && isalnum(c))\n+#define ISALPHA(c) (ISASCII(c) && isalpha(c))\n+#define ISCNTRL(c) (ISASCII(c) && iscntrl(c))\n+#define ISLOWER(c) (ISASCII(c) && islower(c))\n+#define ISPUNCT(c) (ISASCII(c) && ispunct(c))\n+#define ISSPACE(c) (ISASCII(c) && isspace(c))\n+#define ISUPPER(c) (ISASCII(c) && isupper(c))\n+#define ISXDIGIT(c) (ISASCII(c) && isxdigit(c))\n+\n+#ifdef WILD_TEST_ITERATIONS\n+int wildmatch_iteration_count;\n+#endif\n+\n+static int force_lower_case = 0;\n+\n+/* Match pattern \"p\" against the a virtually-joined string consisting\n+ * of \"text\" and any strings in array \"a\". */\n+static int dowild(const uchar *p, const uchar *text, const uchar*const *a)\n+{\n+    uchar p_ch;\n+\n+#ifdef WILD_TEST_ITERATIONS\n+    wildmatch_iteration_count++;\n+#endif\n+\n+    for ( ; (p_ch = *p) != '\\0'; text++, p++) {\n+\tint matched, special;\n+\tuchar t_ch, prev_ch;\n+\twhile ((t_ch = *text) == '\\0') {\n+\t    if (*a == NULL) {\n+\t\tif (p_ch != '*')\n+\t\t    return ABORT_ALL;\n+\t\tbreak;\n+\t    }\n+\t    text = *a++;\n+\t}\n+\tif (force_lower_case && ISUPPER(t_ch))\n+\t    t_ch = tolower(t_ch);\n+\tswitch (p_ch) {\n+\t  case '\\\\':\n+\t    /* Literal match with following character.  Note that the test\n+\t     * in \"default\" handles the p[1] == '\\0' failure case. */\n+\t    p_ch = *++p;\n+\t    /* FALLTHROUGH */\n+\t  default:\n+\t    if (t_ch != p_ch)\n+\t\treturn FALSE;\n+\t    continue;\n+\t  case '?':\n+\t    /* Match anything but '/'. */\n+\t    if (t_ch == '/')\n+\t\treturn FALSE;\n+\t    continue;\n+\t  case '*':\n+\t    if (*++p == '*') {\n+\t\twhile (*++p == '*') {}\n+\t\tspecial = TRUE;\n+\t    } else\n+\t\tspecial = FALSE;\n+\t    if (*p == '\\0') {\n+\t\t/* Trailing \"**\" matches everything.  Trailing \"*\" matches\n+\t\t * only if there are no more slash characters. */\n+\t\tif (!special) {\n+\t\t    do {\n+\t\t\tif (strchr((char*)text, '/') != NULL)\n+\t\t\t    return FALSE;\n+\t\t    } while ((text = *a++) != NULL);\n+\t\t}\n+\t\treturn TRUE;\n+\t    }\n+\t    while (1) {\n+\t\tif (t_ch == '\\0') {\n+\t\t    if ((text = *a++) == NULL)\n+\t\t\tbreak;\n+\t\t    t_ch = *text;\n+\t\t    continue;\n+\t\t}\n+\t\tif ((matched = dowild(p, text, a)) != FALSE) {\n+\t\t    if (!special || matched != ABORT_TO_STARSTAR)\n+\t\t\treturn matched;\n+\t\t} else if (!special && t_ch == '/')\n+\t\t    return ABORT_TO_STARSTAR;\n+\t\tt_ch = *++text;\n+\t    }\n+\t    return ABORT_ALL;\n+\t  case '[':\n+\t    p_ch = *++p;\n+#ifdef NEGATE_CLASS2\n+\t    if (p_ch == NEGATE_CLASS2)\n+\t\tp_ch = NEGATE_CLASS;\n+#endif\n+\t    /* Assign literal TRUE/FALSE because of \"matched\" comparison. */\n+\t    special = p_ch == NEGATE_CLASS? TRUE : FALSE;\n+\t    if (special) {\n+\t\t/* Inverted character class. */\n+\t\tp_ch = *++p;\n+\t    }\n+\t    prev_ch = 0;\n+\t    matched = FALSE;\n+\t    do {\n+\t\tif (!p_ch)\n+\t\t    return ABORT_ALL;\n+\t\tif (p_ch == '\\\\') {\n+\t\t    p_ch = *++p;\n+\t\t    if (!p_ch)\n+\t\t\treturn ABORT_ALL;\n+\t\t    if (t_ch == p_ch)\n+\t\t\tmatched = TRUE;\n+\t\t} else if (p_ch == '-' && prev_ch && p[1] && p[1] != ']') {\n+\t\t    p_ch = *++p;\n+\t\t    if (p_ch == '\\\\') {\n+\t\t\tp_ch = *++p;\n+\t\t\tif (!p_ch)\n+\t\t\t    return ABORT_ALL;\n+\t\t    }\n+\t\t    if (t_ch <= p_ch && t_ch >= prev_ch)\n+\t\t\tmatched = TRUE;\n+\t\t    p_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n+\t\t} else if (p_ch == '[' && p[1] == ':') {\n+\t\t    const uchar *s;\n+\t\t    int i;\n+\t\t    for (s = p += 2; (p_ch = *p) && p_ch != ']'; p++) {} /*SHARED ITERATOR*/\n+\t\t    if (!p_ch)\n+\t\t\treturn ABORT_ALL;\n+\t\t    i = p - s - 1;\n+\t\t    if (i < 0 || p[-1] != ':') {\n+\t\t\t/* Didn't find \":]\", so treat like a normal set. */\n+\t\t\tp = s - 2;\n+\t\t\tp_ch = '[';\n+\t\t\tif (t_ch == p_ch)\n+\t\t\t    matched = TRUE;\n+\t\t\tcontinue;\n+\t\t    }\n+\t\t    if (CC_EQ(s,i, \"alnum\")) {\n+\t\t\tif (ISALNUM(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"alpha\")) {\n+\t\t\tif (ISALPHA(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"blank\")) {\n+\t\t\tif (ISBLANK(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"cntrl\")) {\n+\t\t\tif (ISCNTRL(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"digit\")) {\n+\t\t\tif (ISDIGIT(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"graph\")) {\n+\t\t\tif (ISGRAPH(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"lower\")) {\n+\t\t\tif (ISLOWER(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"print\")) {\n+\t\t\tif (ISPRINT(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"punct\")) {\n+\t\t\tif (ISPUNCT(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"space\")) {\n+\t\t\tif (ISSPACE(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"upper\")) {\n+\t\t\tif (ISUPPER(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else if (CC_EQ(s,i, \"xdigit\")) {\n+\t\t\tif (ISXDIGIT(t_ch))\n+\t\t\t    matched = TRUE;\n+\t\t    } else /* malformed [:class:] string */\n+\t\t\treturn ABORT_ALL;\n+\t\t    p_ch = 0; /* This makes \"prev_ch\" get set to 0. */\n+\t\t} else if (t_ch == p_ch)\n+\t\t    matched = TRUE;\n+\t    } while (prev_ch = p_ch, (p_ch = *++p) != ']');\n+\t    if (matched == special || t_ch == '/')\n+\t\treturn FALSE;\n+\t    continue;\n+\t}\n+    }\n+\n+    do {\n+\tif (*text)\n+\t    return FALSE;\n+    } while ((text = *a++) != NULL);\n+\n+    return TRUE;\n+}\n+\n+/* Match literal string \"s\" against the a virtually-joined string consisting\n+ * of \"text\" and any strings in array \"a\". */\n+static int doliteral(const uchar *s, const uchar *text, const uchar*const *a)\n+{\n+    for ( ; *s != '\\0'; text++, s++) {\n+\twhile (*text == '\\0') {\n+\t    if ((text = *a++) == NULL)\n+\t\treturn FALSE;\n+\t}\n+\tif (*text != *s)\n+\t    return FALSE;\n+    }\n+\n+    do {\n+\tif (*text)\n+\t    return FALSE;\n+    } while ((text = *a++) != NULL);\n+\n+    return TRUE;\n+}\n+\n+/* Return the last \"count\" path elements from the concatenated string.\n+ * We return a string pointer to the start of the string, and update the\n+ * array pointer-pointer to point to any remaining string elements. */\n+static const uchar *trailing_N_elements(const uchar*const **a_ptr, int count)\n+{\n+    const uchar*const *a = *a_ptr;\n+    const uchar*const *first_a = a;\n+\n+    while (*a)\n+\t    a++;\n+\n+    while (a != first_a) {\n+\tconst uchar *s = *--a;\n+\ts += strlen((char*)s);\n+\twhile (--s >= *a) {\n+\t    if (*s == '/' && !--count) {\n+\t\t*a_ptr = a+1;\n+\t\treturn s+1;\n+\t    }\n+\t}\n+    }\n+\n+    if (count == 1) {\n+\t*a_ptr = a+1;\n+\treturn *a;\n+    }\n+\n+    return NULL;\n+}\n+\n+/* Match the \"pattern\" against the \"text\" string. */\n+int wildmatch(const char *pattern, const char *text)\n+{\n+    static const uchar *nomore[1]; /* A NULL pointer. */\n+#ifdef WILD_TEST_ITERATIONS\n+    wildmatch_iteration_count = 0;\n+#endif\n+    return dowild((const uchar*)pattern, (const uchar*)text, nomore) == TRUE;\n+}\n+\n+/* Match the \"pattern\" against the forced-to-lower-case \"text\" string. */\n+int iwildmatch(const char *pattern, const char *text)\n+{\n+    static const uchar *nomore[1]; /* A NULL pointer. */\n+    int ret;\n+#ifdef WILD_TEST_ITERATIONS\n+    wildmatch_iteration_count = 0;\n+#endif\n+    force_lower_case = 1;\n+    ret = dowild((const uchar*)pattern, (const uchar*)text, nomore) == TRUE;\n+    force_lower_case = 0;\n+    return ret;\n+}\n+\n+/* Match pattern \"p\" against the a virtually-joined string consisting\n+ * of all the pointers in array \"texts\" (which has a NULL pointer at the\n+ * end).  The int \"where\" can be 0 (normal matching), > 0 (match only\n+ * the trailing N slash-separated filename components of \"texts\"), or < 0\n+ * (match the \"pattern\" at the start or after any slash in \"texts\"). */\n+int wildmatch_array(const char *pattern, const char*const *texts, int where)\n+{\n+    const uchar *p = (const uchar*)pattern;\n+    const uchar*const *a = (const uchar*const*)texts;\n+    const uchar *text;\n+    int matched;\n+\n+#ifdef WILD_TEST_ITERATIONS\n+    wildmatch_iteration_count = 0;\n+#endif\n+\n+    if (where > 0)\n+\ttext = trailing_N_elements(&a, where);\n+    else\n+\ttext = *a++;\n+    if (!text)\n+\treturn FALSE;\n+\n+    if ((matched = dowild(p, text, a)) != TRUE && where < 0\n+     && matched != ABORT_ALL) {\n+\twhile (1) {\n+\t    if (*text == '\\0') {\n+\t\tif ((text = (uchar*)*a++) == NULL)\n+\t\t    return FALSE;\n+\t\tcontinue;\n+\t    }\n+\t    if (*text++ == '/' && (matched = dowild(p, text, a)) != FALSE\n+\t     && matched != ABORT_TO_STARSTAR)\n+\t\tbreak;\n+\t}\n+    }\n+    return matched == TRUE;\n+}\n+\n+/* Match literal string \"s\" against the a virtually-joined string consisting\n+ * of all the pointers in array \"texts\" (which has a NULL pointer at the\n+ * end).  The int \"where\" can be 0 (normal matching), or > 0 (match\n+ * only the trailing N slash-separated filename components of \"texts\"). */\n+int litmatch_array(const char *string, const char*const *texts, int where)\n+{\n+    const uchar *s = (const uchar*)string;\n+    const uchar*const *a = (const uchar* const*)texts;\n+    const uchar *text;\n+\n+    if (where > 0)\n+\ttext = trailing_N_elements(&a, where);\n+    else\n+\ttext = *a++;\n+    if (!text)\n+\treturn FALSE;\n+\n+    return doliteral(s, text, a) == TRUE;\n+}\ndiff --git a/wildmatch.h b/wildmatch.h\nnew file mode 100644\nindex 0000000..e7f1a35\n--- /dev/null\n+++ b/wildmatch.h\n@@ -0,0 +1,6 @@\n+/* wildmatch.h */\n+\n+int wildmatch(const char *pattern, const char *text);\n+int iwildmatch(const char *pattern, const char *text);\n+int wildmatch_array(const char *pattern, const char*const *texts, int where);\n+int litmatch_array(const char *string, const char*const *texts, int where);\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201104","messageId":"1350182110-25936-5-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 04/12] wildmatch: remove unnecessary functions","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:02Z","receivedAt":"2012-10-14T02:35:02Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n wildmatch.c | 164 ++++--------------------------------------------------------\n wildmatch.h |   2 -\n 2 files changed, 10 insertions(+), 156 deletions(-)\n\ndiff --git a/wildmatch.c b/wildmatch.c\nindex f3a1731..fae7397 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -53,33 +53,18 @@\n #define ISUPPER(c) (ISASCII(c) && isupper(c))\n #define ISXDIGIT(c) (ISASCII(c) && isxdigit(c))\n \n-#ifdef WILD_TEST_ITERATIONS\n-int wildmatch_iteration_count;\n-#endif\n-\n static int force_lower_case = 0;\n \n-/* Match pattern \"p\" against the a virtually-joined string consisting\n- * of \"text\" and any strings in array \"a\". */\n-static int dowild(const uchar *p, const uchar *text, const uchar*const *a)\n+/* Match pattern \"p\" against \"text\" */\n+static int dowild(const uchar *p, const uchar *text)\n {\n     uchar p_ch;\n \n-#ifdef WILD_TEST_ITERATIONS\n-    wildmatch_iteration_count++;\n-#endif\n-\n     for ( ; (p_ch = *p) != '\\0'; text++, p++) {\n \tint matched, special;\n \tuchar t_ch, prev_ch;\n-\twhile ((t_ch = *text) == '\\0') {\n-\t    if (*a == NULL) {\n-\t\tif (p_ch != '*')\n-\t\t    return ABORT_ALL;\n-\t\tbreak;\n-\t    }\n-\t    text = *a++;\n-\t}\n+\tif ((t_ch = *text) == '\\0' && p_ch != '*')\n+\t\treturn ABORT_ALL;\n \tif (force_lower_case && ISUPPER(t_ch))\n \t    t_ch = tolower(t_ch);\n \tswitch (p_ch) {\n@@ -107,21 +92,15 @@ static int dowild(const uchar *p, const uchar *text, const uchar*const *a)\n \t\t/* Trailing \"**\" matches everything.  Trailing \"*\" matches\n \t\t * only if there are no more slash characters. */\n \t\tif (!special) {\n-\t\t    do {\n \t\t\tif (strchr((char*)text, '/') != NULL)\n \t\t\t    return FALSE;\n-\t\t    } while ((text = *a++) != NULL);\n \t\t}\n \t\treturn TRUE;\n \t    }\n \t    while (1) {\n-\t\tif (t_ch == '\\0') {\n-\t\t    if ((text = *a++) == NULL)\n-\t\t\tbreak;\n-\t\t    t_ch = *text;\n-\t\t    continue;\n-\t\t}\n-\t\tif ((matched = dowild(p, text, a)) != FALSE) {\n+\t\tif (t_ch == '\\0')\n+\t\t    break;\n+\t\tif ((matched = dowild(p, text)) != FALSE) {\n \t\t    if (!special || matched != ABORT_TO_STARSTAR)\n \t\t\treturn matched;\n \t\t} else if (!special && t_ch == '/')\n@@ -225,144 +204,21 @@ static int dowild(const uchar *p, const uchar *text, const uchar*const *a)\n \t}\n     }\n \n-    do {\n-\tif (*text)\n-\t    return FALSE;\n-    } while ((text = *a++) != NULL);\n-\n-    return TRUE;\n-}\n-\n-/* Match literal string \"s\" against the a virtually-joined string consisting\n- * of \"text\" and any strings in array \"a\". */\n-static int doliteral(const uchar *s, const uchar *text, const uchar*const *a)\n-{\n-    for ( ; *s != '\\0'; text++, s++) {\n-\twhile (*text == '\\0') {\n-\t    if ((text = *a++) == NULL)\n-\t\treturn FALSE;\n-\t}\n-\tif (*text != *s)\n-\t    return FALSE;\n-    }\n-\n-    do {\n-\tif (*text)\n-\t    return FALSE;\n-    } while ((text = *a++) != NULL);\n-\n-    return TRUE;\n-}\n-\n-/* Return the last \"count\" path elements from the concatenated string.\n- * We return a string pointer to the start of the string, and update the\n- * array pointer-pointer to point to any remaining string elements. */\n-static const uchar *trailing_N_elements(const uchar*const **a_ptr, int count)\n-{\n-    const uchar*const *a = *a_ptr;\n-    const uchar*const *first_a = a;\n-\n-    while (*a)\n-\t    a++;\n-\n-    while (a != first_a) {\n-\tconst uchar *s = *--a;\n-\ts += strlen((char*)s);\n-\twhile (--s >= *a) {\n-\t    if (*s == '/' && !--count) {\n-\t\t*a_ptr = a+1;\n-\t\treturn s+1;\n-\t    }\n-\t}\n-    }\n-\n-    if (count == 1) {\n-\t*a_ptr = a+1;\n-\treturn *a;\n-    }\n-\n-    return NULL;\n+    return *text ? FALSE : TRUE;\n }\n \n /* Match the \"pattern\" against the \"text\" string. */\n int wildmatch(const char *pattern, const char *text)\n {\n-    static const uchar *nomore[1]; /* A NULL pointer. */\n-#ifdef WILD_TEST_ITERATIONS\n-    wildmatch_iteration_count = 0;\n-#endif\n-    return dowild((const uchar*)pattern, (const uchar*)text, nomore) == TRUE;\n+    return dowild((const uchar*)pattern, (const uchar*)text) == TRUE;\n }\n \n /* Match the \"pattern\" against the forced-to-lower-case \"text\" string. */\n int iwildmatch(const char *pattern, const char *text)\n {\n-    static const uchar *nomore[1]; /* A NULL pointer. */\n     int ret;\n-#ifdef WILD_TEST_ITERATIONS\n-    wildmatch_iteration_count = 0;\n-#endif\n     force_lower_case = 1;\n-    ret = dowild((const uchar*)pattern, (const uchar*)text, nomore) == TRUE;\n+    ret = dowild((const uchar*)pattern, (const uchar*)text) == TRUE;\n     force_lower_case = 0;\n     return ret;\n }\n-\n-/* Match pattern \"p\" against the a virtually-joined string consisting\n- * of all the pointers in array \"texts\" (which has a NULL pointer at the\n- * end).  The int \"where\" can be 0 (normal matching), > 0 (match only\n- * the trailing N slash-separated filename components of \"texts\"), or < 0\n- * (match the \"pattern\" at the start or after any slash in \"texts\"). */\n-int wildmatch_array(const char *pattern, const char*const *texts, int where)\n-{\n-    const uchar *p = (const uchar*)pattern;\n-    const uchar*const *a = (const uchar*const*)texts;\n-    const uchar *text;\n-    int matched;\n-\n-#ifdef WILD_TEST_ITERATIONS\n-    wildmatch_iteration_count = 0;\n-#endif\n-\n-    if (where > 0)\n-\ttext = trailing_N_elements(&a, where);\n-    else\n-\ttext = *a++;\n-    if (!text)\n-\treturn FALSE;\n-\n-    if ((matched = dowild(p, text, a)) != TRUE && where < 0\n-     && matched != ABORT_ALL) {\n-\twhile (1) {\n-\t    if (*text == '\\0') {\n-\t\tif ((text = (uchar*)*a++) == NULL)\n-\t\t    return FALSE;\n-\t\tcontinue;\n-\t    }\n-\t    if (*text++ == '/' && (matched = dowild(p, text, a)) != FALSE\n-\t     && matched != ABORT_TO_STARSTAR)\n-\t\tbreak;\n-\t}\n-    }\n-    return matched == TRUE;\n-}\n-\n-/* Match literal string \"s\" against the a virtually-joined string consisting\n- * of all the pointers in array \"texts\" (which has a NULL pointer at the\n- * end).  The int \"where\" can be 0 (normal matching), or > 0 (match\n- * only the trailing N slash-separated filename components of \"texts\"). */\n-int litmatch_array(const char *string, const char*const *texts, int where)\n-{\n-    const uchar *s = (const uchar*)string;\n-    const uchar*const *a = (const uchar* const*)texts;\n-    const uchar *text;\n-\n-    if (where > 0)\n-\ttext = trailing_N_elements(&a, where);\n-    else\n-\ttext = *a++;\n-    if (!text)\n-\treturn FALSE;\n-\n-    return doliteral(s, text, a) == TRUE;\n-}\ndiff --git a/wildmatch.h b/wildmatch.h\nindex e7f1a35..562faa3 100644\n--- a/wildmatch.h\n+++ b/wildmatch.h\n@@ -2,5 +2,3 @@\n \n int wildmatch(const char *pattern, const char *text);\n int iwildmatch(const char *pattern, const char *text);\n-int wildmatch_array(const char *pattern, const char*const *texts, int where);\n-int litmatch_array(const char *string, const char*const *texts, int where);\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201105","messageId":"1350182110-25936-6-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 05/12] Integrate wildmatch to git","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:03Z","receivedAt":"2012-10-14T02:35:03Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n .gitignore           |   1 +\n Makefile             |   3 +\n t/t3070-wildmatch.sh | 188 +++++++++++++++++++++++++++++++++++++++++++++++++++\n t/t3070/wildtest.txt | 165 --------------------------------------------\n test-wildmatch.c     |  14 ++++\n wildmatch.c          |   5 +-\n 6 files changed, 210 insertions(+), 166 deletions(-)\n create mode 100755 t/t3070-wildmatch.sh\n delete mode 100644 t/t3070/wildtest.txt\n create mode 100644 test-wildmatch.c\n\ndiff --git a/.gitignore b/.gitignore\nindex a188a82..37c3507 100644\n--- a/.gitignore\n+++ b/.gitignore\n@@ -197,6 +197,7 @@\n /test-string-list\n /test-subprocess\n /test-svn-fe\n+/test-wildmatch\n /common-cmds.h\n *.tar.gz\n *.dsc\ndiff --git a/Makefile b/Makefile\nindex f69979e..c752673 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -523,6 +523,7 @@ TEST_PROGRAMS_NEED_X += test-sigchain\n TEST_PROGRAMS_NEED_X += test-string-list\n TEST_PROGRAMS_NEED_X += test-subprocess\n TEST_PROGRAMS_NEED_X += test-svn-fe\n+TEST_PROGRAMS_NEED_X += test-wildmatch\n \n TEST_PROGRAMS = $(patsubst %,%$X,$(TEST_PROGRAMS_NEED_X))\n \n@@ -695,6 +696,7 @@ LIB_H += userdiff.h\n LIB_H += utf8.h\n LIB_H += varint.h\n LIB_H += walker.h\n+LIB_H += wildmatch.h\n LIB_H += wt-status.h\n LIB_H += xdiff-interface.h\n LIB_H += xdiff/xdiff.h\n@@ -826,6 +828,7 @@ LIB_OBJS += utf8.o\n LIB_OBJS += varint.o\n LIB_OBJS += version.o\n LIB_OBJS += walker.o\n+LIB_OBJS += wildmatch.o\n LIB_OBJS += wrapper.o\n LIB_OBJS += write_or_die.o\n LIB_OBJS += ws.o\ndiff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nnew file mode 100755\nindex 0000000..dbd3c8b\n--- /dev/null\n+++ b/t/t3070-wildmatch.sh\n@@ -0,0 +1,188 @@\n+#!/bin/sh\n+\n+test_description='wildmatch tests'\n+\n+. ./test-lib.sh\n+\n+match() {\n+    if [ $1 = 1 ]; then\n+\ttest_expect_success \"wildmatch:    match '$3' '$4'\" \"\n+\t    test-wildmatch wildmatch '$3' '$4'\n+\t\"\n+    else\n+\ttest_expect_success \"wildmatch: no match '$3' '$4'\" \"\n+\t    ! test-wildmatch wildmatch '$3' '$4'\n+\t\"\n+    fi\n+    if [ $2 = 1 ]; then\n+\ttest_expect_success \"fnmatch:      match '$3' '$4'\" \"\n+\t    test-wildmatch fnmatch '$3' '$4'\n+\t\"\n+    elif [ $2 = 0 ]; then\n+\ttest_expect_success \"fnmatch:   no match '$3' '$4'\" \"\n+\t    ! test-wildmatch fnmatch '$3' '$4'\n+\t\"\n+#    else\n+#\ttest_expect_success BROKEN_FNMATCH \"fnmatch:       '$3' '$4'\" \"\n+#\t    ! test-wildmatch fnmatch '$3' '$4'\n+#\t\"\n+    fi\n+}\n+\n+# Basic wildmat features\n+match 1 1 foo foo\n+match 0 0 foo bar\n+match 1 1 '' \"\"\n+match 1 1 foo '???'\n+match 0 0 foo '??'\n+match 1 1 foo '*'\n+match 1 1 foo 'f*'\n+match 0 0 foo '*f'\n+match 1 1 foo '*foo*'\n+match 1 1 foobar '*ob*a*r*'\n+match 1 1 aaaaaaabababab '*ab'\n+match 1 1 'foo*' 'foo\\*'\n+match 0 0 foobar 'foo\\*bar'\n+match 1 1 'f\\oo' 'f\\\\oo'\n+match 1 1 ball '*[al]?'\n+match 0 0 ten '[ten]'\n+match 1 1 ten '**[!te]'\n+match 0 0 ten '**[!ten]'\n+match 1 1 ten 't[a-g]n'\n+match 0 0 ten 't[!a-g]n'\n+match 1 1 ton 't[!a-g]n'\n+match 1 1 ton 't[^a-g]n'\n+match 1 1 'a]b' 'a[]]b'\n+match 1 1 a-b 'a[]-]b'\n+match 1 1 'a]b' 'a[]-]b'\n+match 0 0 aab 'a[]-]b'\n+match 1 1 aab 'a[]a-]b'\n+match 1 1 ']' ']'\n+\n+# Extended slash-matching features\n+match 0 0 'foo/baz/bar' 'foo*bar'\n+match 1 0 'foo/baz/bar' 'foo**bar'\n+match 0 0 'foo/bar' 'foo?bar'\n+match 0 0 'foo/bar' 'foo[/]bar'\n+match 0 0 'foo/bar' 'f[^eiu][^eiu][^eiu][^eiu][^eiu]r'\n+match 1 1 'foo-bar' 'f[^eiu][^eiu][^eiu][^eiu][^eiu]r'\n+match 0 0 'foo' '**/foo'\n+match 1 1 '/foo' '**/foo'\n+match 1 0 'bar/baz/foo' '**/foo'\n+match 0 0 'bar/baz/foo' '*/foo'\n+match 0 0 'foo/bar/baz' '**/bar*'\n+match 1 0 'deep/foo/bar/baz' '**/bar/*'\n+match 0 0 'deep/foo/bar/baz/' '**/bar/*'\n+match 1 0 'deep/foo/bar/baz/' '**/bar/**'\n+match 0 0 'deep/foo/bar' '**/bar/*'\n+match 1 0 'deep/foo/bar/' '**/bar/**'\n+match 1 0 'foo/bar/baz' '**/bar**'\n+match 1 0 'foo/bar/baz/x' '*/bar/**'\n+match 0 0 'deep/foo/bar/baz/x' '*/bar/**'\n+match 1 0 'deep/foo/bar/baz/x' '**/bar/*/*'\n+\n+# Various additional tests\n+match 0 0 'acrt' 'a[c-c]st'\n+match 1 1 'acrt' 'a[c-c]rt'\n+match 0 0 ']' '[!]-]'\n+match 1 1 'a' '[!]-]'\n+match 0 0 '' '\\'\n+match 0 0 '\\' '\\'\n+match 0 0 '/\\' '*/\\'\n+match 1 1 '/\\' '*/\\\\'\n+match 1 1 'foo' 'foo'\n+match 1 1 '@foo' '@foo'\n+match 0 0 'foo' '@foo'\n+match 1 1 '[ab]' '\\[ab]'\n+match 1 1 '[ab]' '[[]ab]'\n+match 1 1 '[ab]' '[[:]ab]'\n+match 0 0 '[ab]' '[[::]ab]'\n+match 1 1 '[ab]' '[[:digit]ab]'\n+match 1 1 '[ab]' '[\\[:]ab]'\n+match 1 1 '?a?b' '\\??\\?b'\n+match 1 1 'abc' '\\a\\b\\c'\n+match 0 0 'foo' ''\n+match 1 0 'foo/bar/baz/to' '**/t[o]'\n+\n+# Character class tests\n+match 1 1 'a1B' '[[:alpha:]][[:digit:]][[:upper:]]'\n+match 0 0 'a' '[[:digit:][:upper:][:space:]]'\n+match 1 1 'A' '[[:digit:][:upper:][:space:]]'\n+match 1 0 '1' '[[:digit:][:upper:][:space:]]'\n+match 0 0 '1' '[[:digit:][:upper:][:spaci:]]'\n+match 1 1 ' ' '[[:digit:][:upper:][:space:]]'\n+match 0 0 '.' '[[:digit:][:upper:][:space:]]'\n+match 1 1 '.' '[[:digit:][:punct:][:space:]]'\n+match 1 1 '5' '[[:xdigit:]]'\n+match 1 1 'f' '[[:xdigit:]]'\n+match 1 1 'D' '[[:xdigit:]]'\n+match 1 0 '_' '[[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]'\n+match 1 0 '_' '[[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]'\n+match 1 1 '.' '[^[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:lower:][:space:][:upper:][:xdigit:]]'\n+match 1 1 '5' '[a-c[:digit:]x-z]'\n+match 1 1 'b' '[a-c[:digit:]x-z]'\n+match 1 1 'y' '[a-c[:digit:]x-z]'\n+match 0 0 'q' '[a-c[:digit:]x-z]'\n+\n+# Additional tests, including some malformed wildmats\n+match 1 1 ']' '[\\\\-^]'\n+match 0 0 '[' '[\\\\-^]'\n+match 1 1 '-' '[\\-_]'\n+match 1 1 ']' '[\\]]'\n+match 0 0 '\\]' '[\\]]'\n+match 0 0 '\\' '[\\]]'\n+match 0 0 'ab' 'a[]b'\n+match 0 1 'a[]b' 'a[]b'\n+match 0 1 'ab[' 'ab['\n+match 0 0 'ab' '[!'\n+match 0 0 'ab' '[-'\n+match 1 1 '-' '[-]'\n+match 0 0 '-' '[a-'\n+match 0 0 '-' '[!a-'\n+match 1 1 '-' '[--A]'\n+match 1 1 '5' '[--A]'\n+match 1 1 ' ' '[ --]'\n+match 1 1 '$' '[ --]'\n+match 1 1 '-' '[ --]'\n+match 0 0 '0' '[ --]'\n+match 1 1 '-' '[---]'\n+match 1 1 '-' '[------]'\n+match 0 0 'j' '[a-e-n]'\n+match 1 1 '-' '[a-e-n]'\n+match 1 1 'a' '[!------]'\n+match 0 0 '[' '[]-a]'\n+match 1 1 '^' '[]-a]'\n+match 0 0 '^' '[!]-a]'\n+match 1 1 '[' '[!]-a]'\n+match 1 1 '^' '[a^bc]'\n+match 1 1 '-b]' '[a-]b]'\n+match 0 0 '\\' '[\\]'\n+match 1 1 '\\' '[\\\\]'\n+match 0 0 '\\' '[!\\\\]'\n+match 1 1 'G' '[A-\\\\]'\n+match 0 0 'aaabbb' 'b*a'\n+match 0 0 'aabcaa' '*ba*'\n+match 1 1 ',' '[,]'\n+match 1 1 ',' '[\\\\,]'\n+match 1 1 '\\' '[\\\\,]'\n+match 1 1 '-' '[,-.]'\n+match 0 0 '+' '[,-.]'\n+match 0 0 '-.]' '[,-.]'\n+match 1 1 '2' '[\\1-\\3]'\n+match 1 1 '3' '[\\1-\\3]'\n+match 0 0 '4' '[\\1-\\3]'\n+match 1 1 '\\' '[[-\\]]'\n+match 1 1 '[' '[[-\\]]'\n+match 1 1 ']' '[[-\\]]'\n+match 0 0 '-' '[[-\\]]'\n+\n+# Test recursion and the abort code (use \"wildtest -i\" to see iteration counts)\n+match 1 1 '-adobe-courier-bold-o-normal--12-120-75-75-m-70-iso8859-1' '-*-*-*-*-*-*-12-*-*-*-m-*-*-*'\n+match 0 0 '-adobe-courier-bold-o-normal--12-120-75-75-X-70-iso8859-1' '-*-*-*-*-*-*-12-*-*-*-m-*-*-*'\n+match 0 0 '-adobe-courier-bold-o-normal--12-120-75-75-/-70-iso8859-1' '-*-*-*-*-*-*-12-*-*-*-m-*-*-*'\n+match 1 1 '/adobe/courier/bold/o/normal//12/120/75/75/m/70/iso8859/1' '/*/*/*/*/*/*/12/*/*/*/m/*/*/*'\n+match 0 0 '/adobe/courier/bold/o/normal//12/120/75/75/X/70/iso8859/1' '/*/*/*/*/*/*/12/*/*/*/m/*/*/*'\n+match 1 0 'abcd/abcdefg/abcdefghijk/abcdefghijklmnop.txt' '**/*a*b*g*n*t'\n+match 0 0 'abcd/abcdefg/abcdefghijk/abcdefghijklmnop.txtz' '**/*a*b*g*n*t'\n+\n+test_done\ndiff --git a/t/t3070/wildtest.txt b/t/t3070/wildtest.txt\ndeleted file mode 100644\nindex 42c1678..0000000\n--- a/t/t3070/wildtest.txt\n+++ /dev/null\n@@ -1,165 +0,0 @@\n-# Input is in the following format (all items white-space separated):\n-#\n-# The first two items are 1 or 0 indicating if the wildmat call is expected to\n-# succeed and if fnmatch works the same way as wildmat, respectively.  After\n-# that is a text string for the match, and a pattern string.  Strings can be\n-# quoted (if desired) in either double or single quotes, as well as backticks.\n-#\n-# MATCH FNMATCH_SAME \"text to match\" 'pattern to use'\n-\n-# Basic wildmat features\n-1 1 foo\t\t\tfoo\n-0 1 foo\t\t\tbar\n-1 1 ''\t\t\t\"\"\n-1 1 foo\t\t\t???\n-0 1 foo\t\t\t??\n-1 1 foo\t\t\t*\n-1 1 foo\t\t\tf*\n-0 1 foo\t\t\t*f\n-1 1 foo\t\t\t*foo*\n-1 1 foobar\t\t*ob*a*r*\n-1 1 aaaaaaabababab\t*ab\n-1 1 foo*\t\tfoo\\*\n-0 1 foobar\t\tfoo\\*bar\n-1 1 f\\oo\t\tf\\\\oo\n-1 1 ball\t\t*[al]?\n-0 1 ten\t\t\t[ten]\n-1 1 ten\t\t\t**[!te]\n-0 1 ten\t\t\t**[!ten]\n-1 1 ten\t\t\tt[a-g]n\n-0 1 ten\t\t\tt[!a-g]n\n-1 1 ton\t\t\tt[!a-g]n\n-1 1 ton\t\t\tt[^a-g]n\n-1 1 a]b\t\t\ta[]]b\n-1 1 a-b\t\t\ta[]-]b\n-1 1 a]b\t\t\ta[]-]b\n-0 1 aab\t\t\ta[]-]b\n-1 1 aab\t\t\ta[]a-]b\n-1 1 ]\t\t\t]\n-\n-# Extended slash-matching features\n-0 1 foo/baz/bar\t\tfoo*bar\n-1 1 foo/baz/bar\t\tfoo**bar\n-0 1 foo/bar\t\tfoo?bar\n-0 1 foo/bar\t\tfoo[/]bar\n-0 1 foo/bar\t\tf[^eiu][^eiu][^eiu][^eiu][^eiu]r\n-1 1 foo-bar\t\tf[^eiu][^eiu][^eiu][^eiu][^eiu]r\n-0 1 foo\t\t\t**/foo\n-1 1 /foo\t\t**/foo\n-1 1 bar/baz/foo\t\t**/foo\n-0 1 bar/baz/foo\t\t*/foo\n-0 0 foo/bar/baz\t\t**/bar*\n-1 1 deep/foo/bar/baz\t**/bar/*\n-0 1 deep/foo/bar/baz/\t**/bar/*\n-1 1 deep/foo/bar/baz/\t**/bar/**\n-0 1 deep/foo/bar\t**/bar/*\n-1 1 deep/foo/bar/\t**/bar/**\n-1 1 foo/bar/baz\t\t**/bar**\n-1 1 foo/bar/baz/x\t*/bar/**\n-0 0 deep/foo/bar/baz/x\t*/bar/**\n-1 1 deep/foo/bar/baz/x\t**/bar/*/*\n-\n-# Various additional tests\n-0 1 acrt\t\ta[c-c]st\n-1 1 acrt\t\ta[c-c]rt\n-0 1 ]\t\t\t[!]-]\n-1 1 a\t\t\t[!]-]\n-0 1 ''\t\t\t\\\n-0 1 \\\t\t\t\\\n-0 1 /\\\t\t\t*/\\\n-1 1 /\\\t\t\t*/\\\\\n-1 1 foo\t\t\tfoo\n-1 1 @foo\t\t@foo\n-0 1 foo\t\t\t@foo\n-1 1 [ab]\t\t\\[ab]\n-1 1 [ab]\t\t[[]ab]\n-1 1 [ab]\t\t[[:]ab]\n-0 1 [ab]\t\t[[::]ab]\n-1 1 [ab]\t\t[[:digit]ab]\n-1 1 [ab]\t\t[\\[:]ab]\n-1 1 ?a?b\t\t\\??\\?b\n-1 1 abc\t\t\t\\a\\b\\c\n-0 1 foo\t\t\t''\n-1 1 foo/bar/baz/to\t**/t[o]\n-\n-# Character class tests\n-1 1 a1B\t\t[[:alpha:]][[:digit:]][[:upper:]]\n-0 1 a\t\t[[:digit:][:upper:][:space:]]\n-1 1 A\t\t[[:digit:][:upper:][:space:]]\n-1 1 1\t\t[[:digit:][:upper:][:space:]]\n-0 1 1\t\t[[:digit:][:upper:][:spaci:]]\n-1 1 ' '\t\t[[:digit:][:upper:][:space:]]\n-0 1 .\t\t[[:digit:][:upper:][:space:]]\n-1 1 .\t\t[[:digit:][:punct:][:space:]]\n-1 1 5\t\t[[:xdigit:]]\n-1 1 f\t\t[[:xdigit:]]\n-1 1 D\t\t[[:xdigit:]]\n-1 1 _\t\t[[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]\n-#1 1 �\t\t[^[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]\n-1 1 \t\t[^[:alnum:][:alpha:][:blank:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]\n-1 1 .\t\t[^[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:lower:][:space:][:upper:][:xdigit:]]\n-1 1 5\t\t[a-c[:digit:]x-z]\n-1 1 b\t\t[a-c[:digit:]x-z]\n-1 1 y\t\t[a-c[:digit:]x-z]\n-0 1 q\t\t[a-c[:digit:]x-z]\n-\n-# Additional tests, including some malformed wildmats\n-1 1 ]\t\t[\\\\-^]\n-0 1 [\t\t[\\\\-^]\n-1 1 -\t\t[\\-_]\n-1 1 ]\t\t[\\]]\n-0 1 \\]\t\t[\\]]\n-0 1 \\\t\t[\\]]\n-0 1 ab\t\ta[]b\n-0 1 a[]b\ta[]b\n-0 1 ab[\t\tab[\n-0 1 ab\t\t[!\n-0 1 ab\t\t[-\n-1 1 -\t\t[-]\n-0 1 -\t\t[a-\n-0 1 -\t\t[!a-\n-1 1 -\t\t[--A]\n-1 1 5\t\t[--A]\n-1 1 ' '\t\t'[ --]'\n-1 1 $\t\t'[ --]'\n-1 1 -\t\t'[ --]'\n-0 1 0\t\t'[ --]'\n-1 1 -\t\t[---]\n-1 1 -\t\t[------]\n-0 1 j\t\t[a-e-n]\n-1 1 -\t\t[a-e-n]\n-1 1 a\t\t[!------]\n-0 1 [\t\t[]-a]\n-1 1 ^\t\t[]-a]\n-0 1 ^\t\t[!]-a]\n-1 1 [\t\t[!]-a]\n-1 1 ^\t\t[a^bc]\n-1 1 -b]\t\t[a-]b]\n-0 1 \\\t\t[\\]\n-1 1 \\\t\t[\\\\]\n-0 1 \\\t\t[!\\\\]\n-1 1 G\t\t[A-\\\\]\n-0 1 aaabbb\tb*a\n-0 1 aabcaa\t*ba*\n-1 1 ,\t\t[,]\n-1 1 ,\t\t[\\\\,]\n-1 1 \\\t\t[\\\\,]\n-1 1 -\t\t[,-.]\n-0 1 +\t\t[,-.]\n-0 1 -.]\t\t[,-.]\n-1 1 2\t\t[\\1-\\3]\n-1 1 3\t\t[\\1-\\3]\n-0 1 4\t\t[\\1-\\3]\n-1 1 \\\t\t[[-\\]]\n-1 1 [\t\t[[-\\]]\n-1 1 ]\t\t[[-\\]]\n-0 1 -\t\t[[-\\]]\n-\n-# Test recursion and the abort code (use \"wildtest -i\" to see iteration counts)\n-1 1 -adobe-courier-bold-o-normal--12-120-75-75-m-70-iso8859-1\t-*-*-*-*-*-*-12-*-*-*-m-*-*-*\n-0 1 -adobe-courier-bold-o-normal--12-120-75-75-X-70-iso8859-1\t-*-*-*-*-*-*-12-*-*-*-m-*-*-*\n-0 1 -adobe-courier-bold-o-normal--12-120-75-75-/-70-iso8859-1\t-*-*-*-*-*-*-12-*-*-*-m-*-*-*\n-1 1 /adobe/courier/bold/o/normal//12/120/75/75/m/70/iso8859/1\t/*/*/*/*/*/*/12/*/*/*/m/*/*/*\n-0 1 /adobe/courier/bold/o/normal//12/120/75/75/X/70/iso8859/1\t/*/*/*/*/*/*/12/*/*/*/m/*/*/*\n-1 1 abcd/abcdefg/abcdefghijk/abcdefghijklmnop.txt\t\t**/*a*b*g*n*t\n-0 1 abcd/abcdefg/abcdefghijk/abcdefghijklmnop.txtz\t\t**/*a*b*g*n*t\ndiff --git a/test-wildmatch.c b/test-wildmatch.c\nnew file mode 100644\nindex 0000000..ac56420\n--- /dev/null\n+++ b/test-wildmatch.c\n@@ -0,0 +1,14 @@\n+#include \"cache.h\"\n+#include \"wildmatch.h\"\n+\n+int main(int argc, char **argv)\n+{\n+\tif (!strcmp(argv[1], \"wildmatch\"))\n+\t\treturn wildmatch(argv[3], argv[2]) ? 0 : 1;\n+\telse if (!strcmp(argv[1], \"iwildmatch\"))\n+\t\treturn iwildmatch(argv[3], argv[2]) ? 0 : 1;\n+\telse if (!strcmp(argv[1], \"fnmatch\"))\n+\t\treturn !!fnmatch(argv[3], argv[2], FNM_PATHNAME);\n+\telse\n+\t\treturn 1;\n+}\ndiff --git a/wildmatch.c b/wildmatch.c\nindex fae7397..d0b906a 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -9,7 +9,10 @@\n **  work differently than '*', and to fix the character-class code.\n */\n \n-#include \"rsync.h\"\n+#include \"cache.h\"\n+#include \"wildmatch.h\"\n+\n+typedef unsigned char uchar;\n \n /* What character marks an inverted character class? */\n #define NEGATE_CLASS\t'!'\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201106","messageId":"1350182110-25936-7-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 06/12] t3070: disable unreliable fnmatch tests","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:04Z","receivedAt":"2012-10-14T02:35:04Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"These tests show different results on different fnmatch() versions. We\ndon't want to test fnmatch here. We want to make sure wildmatch\nbehavior matches fnmatch and that only makes sense in cases when\nfnmatch() behaves consistently.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n t/t3070-wildmatch.sh | 86 ++++++++++++++++++++++++++--------------------------\n 1 file changed, 43 insertions(+), 43 deletions(-)\n\ndiff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nindex dbd3c8b..dd95b00 100755\n--- a/t/t3070-wildmatch.sh\n+++ b/t/t3070-wildmatch.sh\n@@ -52,11 +52,11 @@ match 1 1 ten 't[a-g]n'\n match 0 0 ten 't[!a-g]n'\n match 1 1 ton 't[!a-g]n'\n match 1 1 ton 't[^a-g]n'\n-match 1 1 'a]b' 'a[]]b'\n-match 1 1 a-b 'a[]-]b'\n-match 1 1 'a]b' 'a[]-]b'\n-match 0 0 aab 'a[]-]b'\n-match 1 1 aab 'a[]a-]b'\n+match 1 x 'a]b' 'a[]]b'\n+match 1 x a-b 'a[]-]b'\n+match 1 x 'a]b' 'a[]-]b'\n+match 0 x aab 'a[]-]b'\n+match 1 x aab 'a[]a-]b'\n match 1 1 ']' ']'\n \n # Extended slash-matching features\n@@ -67,7 +67,7 @@ match 0 0 'foo/bar' 'foo[/]bar'\n match 0 0 'foo/bar' 'f[^eiu][^eiu][^eiu][^eiu][^eiu]r'\n match 1 1 'foo-bar' 'f[^eiu][^eiu][^eiu][^eiu][^eiu]r'\n match 0 0 'foo' '**/foo'\n-match 1 1 '/foo' '**/foo'\n+match 1 x '/foo' '**/foo'\n match 1 0 'bar/baz/foo' '**/foo'\n match 0 0 'bar/baz/foo' '*/foo'\n match 0 0 'foo/bar/baz' '**/bar*'\n@@ -85,77 +85,77 @@ match 1 0 'deep/foo/bar/baz/x' '**/bar/*/*'\n match 0 0 'acrt' 'a[c-c]st'\n match 1 1 'acrt' 'a[c-c]rt'\n match 0 0 ']' '[!]-]'\n-match 1 1 'a' '[!]-]'\n+match 1 x 'a' '[!]-]'\n match 0 0 '' '\\'\n-match 0 0 '\\' '\\'\n-match 0 0 '/\\' '*/\\'\n-match 1 1 '/\\' '*/\\\\'\n+match 0 x '\\' '\\'\n+match 0 x '/\\' '*/\\'\n+match 1 x '/\\' '*/\\\\'\n match 1 1 'foo' 'foo'\n match 1 1 '@foo' '@foo'\n match 0 0 'foo' '@foo'\n match 1 1 '[ab]' '\\[ab]'\n match 1 1 '[ab]' '[[]ab]'\n-match 1 1 '[ab]' '[[:]ab]'\n-match 0 0 '[ab]' '[[::]ab]'\n-match 1 1 '[ab]' '[[:digit]ab]'\n-match 1 1 '[ab]' '[\\[:]ab]'\n+match 1 x '[ab]' '[[:]ab]'\n+match 0 x '[ab]' '[[::]ab]'\n+match 1 x '[ab]' '[[:digit]ab]'\n+match 1 x '[ab]' '[\\[:]ab]'\n match 1 1 '?a?b' '\\??\\?b'\n match 1 1 'abc' '\\a\\b\\c'\n match 0 0 'foo' ''\n match 1 0 'foo/bar/baz/to' '**/t[o]'\n \n # Character class tests\n-match 1 1 'a1B' '[[:alpha:]][[:digit:]][[:upper:]]'\n-match 0 0 'a' '[[:digit:][:upper:][:space:]]'\n-match 1 1 'A' '[[:digit:][:upper:][:space:]]'\n-match 1 0 '1' '[[:digit:][:upper:][:space:]]'\n-match 0 0 '1' '[[:digit:][:upper:][:spaci:]]'\n-match 1 1 ' ' '[[:digit:][:upper:][:space:]]'\n-match 0 0 '.' '[[:digit:][:upper:][:space:]]'\n-match 1 1 '.' '[[:digit:][:punct:][:space:]]'\n+match 1 x 'a1B' '[[:alpha:]][[:digit:]][[:upper:]]'\n+match 0 x 'a' '[[:digit:][:upper:][:space:]]'\n+match 1 x 'A' '[[:digit:][:upper:][:space:]]'\n+match 1 x '1' '[[:digit:][:upper:][:space:]]'\n+match 0 x '1' '[[:digit:][:upper:][:spaci:]]'\n+match 1 x ' ' '[[:digit:][:upper:][:space:]]'\n+match 0 x '.' '[[:digit:][:upper:][:space:]]'\n+match 1 x '.' '[[:digit:][:punct:][:space:]]'\n match 1 1 '5' '[[:xdigit:]]'\n match 1 1 'f' '[[:xdigit:]]'\n match 1 1 'D' '[[:xdigit:]]'\n-match 1 0 '_' '[[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]'\n-match 1 0 '_' '[[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]'\n-match 1 1 '.' '[^[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:lower:][:space:][:upper:][:xdigit:]]'\n-match 1 1 '5' '[a-c[:digit:]x-z]'\n-match 1 1 'b' '[a-c[:digit:]x-z]'\n-match 1 1 'y' '[a-c[:digit:]x-z]'\n-match 0 0 'q' '[a-c[:digit:]x-z]'\n+match 1 x '_' '[[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]'\n+match 1 x '_' '[[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:graph:][:lower:][:print:][:punct:][:space:][:upper:][:xdigit:]]'\n+match 1 x '.' '[^[:alnum:][:alpha:][:blank:][:cntrl:][:digit:][:lower:][:space:][:upper:][:xdigit:]]'\n+match 1 x '5' '[a-c[:digit:]x-z]'\n+match 1 x 'b' '[a-c[:digit:]x-z]'\n+match 1 x 'y' '[a-c[:digit:]x-z]'\n+match 0 x 'q' '[a-c[:digit:]x-z]'\n \n # Additional tests, including some malformed wildmats\n-match 1 1 ']' '[\\\\-^]'\n+match 1 x ']' '[\\\\-^]'\n match 0 0 '[' '[\\\\-^]'\n-match 1 1 '-' '[\\-_]'\n-match 1 1 ']' '[\\]]'\n+match 1 x '-' '[\\-_]'\n+match 1 x ']' '[\\]]'\n match 0 0 '\\]' '[\\]]'\n match 0 0 '\\' '[\\]]'\n match 0 0 'ab' 'a[]b'\n-match 0 1 'a[]b' 'a[]b'\n-match 0 1 'ab[' 'ab['\n+match 0 x 'a[]b' 'a[]b'\n+match 0 x 'ab[' 'ab['\n match 0 0 'ab' '[!'\n match 0 0 'ab' '[-'\n match 1 1 '-' '[-]'\n match 0 0 '-' '[a-'\n match 0 0 '-' '[!a-'\n-match 1 1 '-' '[--A]'\n-match 1 1 '5' '[--A]'\n+match 1 x '-' '[--A]'\n+match 1 x '5' '[--A]'\n match 1 1 ' ' '[ --]'\n match 1 1 '$' '[ --]'\n match 1 1 '-' '[ --]'\n match 0 0 '0' '[ --]'\n-match 1 1 '-' '[---]'\n-match 1 1 '-' '[------]'\n+match 1 x '-' '[---]'\n+match 1 x '-' '[------]'\n match 0 0 'j' '[a-e-n]'\n-match 1 1 '-' '[a-e-n]'\n-match 1 1 'a' '[!------]'\n+match 1 x '-' '[a-e-n]'\n+match 1 x 'a' '[!------]'\n match 0 0 '[' '[]-a]'\n-match 1 1 '^' '[]-a]'\n+match 1 x '^' '[]-a]'\n match 0 0 '^' '[!]-a]'\n-match 1 1 '[' '[!]-a]'\n+match 1 x '[' '[!]-a]'\n match 1 1 '^' '[a^bc]'\n-match 1 1 '-b]' '[a-]b]'\n+match 1 x '-b]' '[a-]b]'\n match 0 0 '\\' '[\\]'\n match 1 1 '\\' '[\\\\]'\n match 0 0 '\\' '[!\\\\]'\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201107","messageId":"1350182110-25936-8-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 07/12] wildmatch: make wildmatch's return value compatible with fnmatch","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:05Z","receivedAt":"2012-10-14T02:35:05Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"wildmatch returns non-zero if matched, zero otherwise. This patch\nmakes it return zero if matches, non-zero otherwise, like fnmatch().\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n test-wildmatch.c |  4 ++--\n wildmatch.c      | 21 ++++++++++++---------\n 2 files changed, 14 insertions(+), 11 deletions(-)\n\ndiff --git a/test-wildmatch.c b/test-wildmatch.c\nindex ac56420..77014e9 100644\n--- a/test-wildmatch.c\n+++ b/test-wildmatch.c\n@@ -4,9 +4,9 @@\n int main(int argc, char **argv)\n {\n \tif (!strcmp(argv[1], \"wildmatch\"))\n-\t\treturn wildmatch(argv[3], argv[2]) ? 0 : 1;\n+\t\treturn !!wildmatch(argv[3], argv[2]);\n \telse if (!strcmp(argv[1], \"iwildmatch\"))\n-\t\treturn iwildmatch(argv[3], argv[2]) ? 0 : 1;\n+\t\treturn !!iwildmatch(argv[3], argv[2]);\n \telse if (!strcmp(argv[1], \"fnmatch\"))\n \t\treturn !!fnmatch(argv[3], argv[2], FNM_PATHNAME);\n \telse\ndiff --git a/wildmatch.c b/wildmatch.c\nindex d0b906a..e3ac6cc 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -20,6 +20,9 @@ typedef unsigned char uchar;\n \n #define FALSE 0\n #define TRUE 1\n+\n+#define NOMATCH 1\n+#define MATCH 0\n #define ABORT_ALL -1\n #define ABORT_TO_STARSTAR -2\n \n@@ -78,12 +81,12 @@ static int dowild(const uchar *p, const uchar *text)\n \t    /* FALLTHROUGH */\n \t  default:\n \t    if (t_ch != p_ch)\n-\t\treturn FALSE;\n+\t\treturn NOMATCH;\n \t    continue;\n \t  case '?':\n \t    /* Match anything but '/'. */\n \t    if (t_ch == '/')\n-\t\treturn FALSE;\n+\t\treturn NOMATCH;\n \t    continue;\n \t  case '*':\n \t    if (*++p == '*') {\n@@ -96,14 +99,14 @@ static int dowild(const uchar *p, const uchar *text)\n \t\t * only if there are no more slash characters. */\n \t\tif (!special) {\n \t\t\tif (strchr((char*)text, '/') != NULL)\n-\t\t\t    return FALSE;\n+\t\t\t    return NOMATCH;\n \t\t}\n-\t\treturn TRUE;\n+\t\treturn MATCH;\n \t    }\n \t    while (1) {\n \t\tif (t_ch == '\\0')\n \t\t    break;\n-\t\tif ((matched = dowild(p, text)) != FALSE) {\n+\t\tif ((matched = dowild(p, text)) != NOMATCH) {\n \t\t    if (!special || matched != ABORT_TO_STARSTAR)\n \t\t\treturn matched;\n \t\t} else if (!special && t_ch == '/')\n@@ -202,18 +205,18 @@ static int dowild(const uchar *p, const uchar *text)\n \t\t    matched = TRUE;\n \t    } while (prev_ch = p_ch, (p_ch = *++p) != ']');\n \t    if (matched == special || t_ch == '/')\n-\t\treturn FALSE;\n+\t\treturn NOMATCH;\n \t    continue;\n \t}\n     }\n \n-    return *text ? FALSE : TRUE;\n+    return *text ? NOMATCH : MATCH;\n }\n \n /* Match the \"pattern\" against the \"text\" string. */\n int wildmatch(const char *pattern, const char *text)\n {\n-    return dowild((const uchar*)pattern, (const uchar*)text) == TRUE;\n+    return dowild((const uchar*)pattern, (const uchar*)text);\n }\n \n /* Match the \"pattern\" against the forced-to-lower-case \"text\" string. */\n@@ -221,7 +224,7 @@ int iwildmatch(const char *pattern, const char *text)\n {\n     int ret;\n     force_lower_case = 1;\n-    ret = dowild((const uchar*)pattern, (const uchar*)text) == TRUE;\n+    ret = dowild((const uchar*)pattern, (const uchar*)text);\n     force_lower_case = 0;\n     return ret;\n }\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201108","messageId":"1350182110-25936-9-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 08/12] wildmatch: remove static variable force_lower_case","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:06Z","receivedAt":"2012-10-14T02:35:06Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"One place less to worry about thread safety. Also combine wildmatch\nand iwildmatch into one.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n test-wildmatch.c |  4 ++--\n wildmatch.c      | 21 +++++----------------\n wildmatch.h      |  3 +--\n 3 files changed, 8 insertions(+), 20 deletions(-)\n\ndiff --git a/test-wildmatch.c b/test-wildmatch.c\nindex 77014e9..74c0864 100644\n--- a/test-wildmatch.c\n+++ b/test-wildmatch.c\n@@ -4,9 +4,9 @@\n int main(int argc, char **argv)\n {\n \tif (!strcmp(argv[1], \"wildmatch\"))\n-\t\treturn !!wildmatch(argv[3], argv[2]);\n+\t\treturn !!wildmatch(argv[3], argv[2], 0);\n \telse if (!strcmp(argv[1], \"iwildmatch\"))\n-\t\treturn !!iwildmatch(argv[3], argv[2]);\n+\t\treturn !!wildmatch(argv[3], argv[2], FNM_CASEFOLD);\n \telse if (!strcmp(argv[1], \"fnmatch\"))\n \t\treturn !!fnmatch(argv[3], argv[2], FNM_PATHNAME);\n \telse\ndiff --git a/wildmatch.c b/wildmatch.c\nindex e3ac6cc..20c5ef6 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -59,10 +59,8 @@ typedef unsigned char uchar;\n #define ISUPPER(c) (ISASCII(c) && isupper(c))\n #define ISXDIGIT(c) (ISASCII(c) && isxdigit(c))\n \n-static int force_lower_case = 0;\n-\n /* Match pattern \"p\" against \"text\" */\n-static int dowild(const uchar *p, const uchar *text)\n+static int dowild(const uchar *p, const uchar *text, int force_lower_case)\n {\n     uchar p_ch;\n \n@@ -106,7 +104,7 @@ static int dowild(const uchar *p, const uchar *text)\n \t    while (1) {\n \t\tif (t_ch == '\\0')\n \t\t    break;\n-\t\tif ((matched = dowild(p, text)) != NOMATCH) {\n+\t\tif ((matched = dowild(p, text, force_lower_case)) != NOMATCH) {\n \t\t    if (!special || matched != ABORT_TO_STARSTAR)\n \t\t\treturn matched;\n \t\t} else if (!special && t_ch == '/')\n@@ -214,17 +212,8 @@ static int dowild(const uchar *p, const uchar *text)\n }\n \n /* Match the \"pattern\" against the \"text\" string. */\n-int wildmatch(const char *pattern, const char *text)\n-{\n-    return dowild((const uchar*)pattern, (const uchar*)text);\n-}\n-\n-/* Match the \"pattern\" against the forced-to-lower-case \"text\" string. */\n-int iwildmatch(const char *pattern, const char *text)\n+int wildmatch(const char *pattern, const char *text, int flags)\n {\n-    int ret;\n-    force_lower_case = 1;\n-    ret = dowild((const uchar*)pattern, (const uchar*)text);\n-    force_lower_case = 0;\n-    return ret;\n+    return dowild((const uchar*)pattern, (const uchar*)text,\n+\t\t  flags & FNM_CASEFOLD ? 1 : 0);\n }\ndiff --git a/wildmatch.h b/wildmatch.h\nindex 562faa3..e974f9a 100644\n--- a/wildmatch.h\n+++ b/wildmatch.h\n@@ -1,4 +1,3 @@\n /* wildmatch.h */\n \n-int wildmatch(const char *pattern, const char *text);\n-int iwildmatch(const char *pattern, const char *text);\n+int wildmatch(const char *pattern, const char *text, int flags);\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201109","messageId":"1350182110-25936-10-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 09/12] wildmatch: fix case-insensitive matching","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:07Z","receivedAt":"2012-10-14T02:35:07Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"dowild() does case insensitive matching by lower-casing the text. That\nmeans lower case letters in patterns imply case-insensitive matching,\nbut upper case means exact matching.\n\nWe do not want that subtlety. Lower case pattern too so iwildmatch()\nalways does what we expect it to do.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n wildmatch.c | 2 ++\n 1 file changed, 2 insertions(+)\n\ndiff --git a/wildmatch.c b/wildmatch.c\nindex 20c5ef6..6542524 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -71,6 +71,8 @@ static int dowild(const uchar *p, const uchar *text, int force_lower_case)\n \t\treturn ABORT_ALL;\n \tif (force_lower_case && ISUPPER(t_ch))\n \t    t_ch = tolower(t_ch);\n+\tif (force_lower_case && ISUPPER(p_ch))\n+\t    p_ch = tolower(p_ch);\n \tswitch (p_ch) {\n \t  case '\\\\':\n \t    /* Literal match with following character.  Note that the test\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201110","messageId":"1350182110-25936-11-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 10/12] wildmatch: adjust \"**\" behavior","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:08Z","receivedAt":"2012-10-14T02:35:08Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Standard wildmatch() sees consecutive asterisks as \"*\" that can also\nmatch slashes. But that may be hard to explain to users as\n\"abc/**/def\" can match \"abcdef\", \"abcxyzdef\", \"abc/def\", \"abc/x/def\",\n\"abc/x/y/def\"...\n\nThis patch changes wildmatch so that users can do\n\n- \"**/def\" -> all paths ending with file/directory 'def'\n- \"abc/**\" - equivalent to \"/abc/\"\n- \"abc/**/def\" -> \"abc/x/def\", \"abc/x/y/def\"...\n- otherwise consider the pattern malformed if \"**\" is found\n\nBasically the magic of \"**\" only remains if it's wrapped around by\nslashes.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n t/t3070-wildmatch.sh |  5 +++--\n wildmatch.c          | 13 +++++++------\n wildmatch.h          |  6 ++++++\n 3 files changed, 16 insertions(+), 8 deletions(-)\n\ndiff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nindex dd95b00..15848d5 100755\n--- a/t/t3070-wildmatch.sh\n+++ b/t/t3070-wildmatch.sh\n@@ -46,7 +46,7 @@ match 0 0 foobar 'foo\\*bar'\n match 1 1 'f\\oo' 'f\\\\oo'\n match 1 1 ball '*[al]?'\n match 0 0 ten '[ten]'\n-match 1 1 ten '**[!te]'\n+match 0 1 ten '**[!te]'\n match 0 0 ten '**[!ten]'\n match 1 1 ten 't[a-g]n'\n match 0 0 ten 't[!a-g]n'\n@@ -61,7 +61,8 @@ match 1 1 ']' ']'\n \n # Extended slash-matching features\n match 0 0 'foo/baz/bar' 'foo*bar'\n-match 1 0 'foo/baz/bar' 'foo**bar'\n+match 0 0 'foo/baz/bar' 'foo**bar'\n+match 0 1 'foobazbar' 'foo**bar'\n match 0 0 'foo/bar' 'foo?bar'\n match 0 0 'foo/bar' 'foo[/]bar'\n match 0 0 'foo/bar' 'f[^eiu][^eiu][^eiu][^eiu][^eiu]r'\ndiff --git a/wildmatch.c b/wildmatch.c\nindex 6542524..7209f26 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -21,11 +21,6 @@ typedef unsigned char uchar;\n #define FALSE 0\n #define TRUE 1\n \n-#define NOMATCH 1\n-#define MATCH 0\n-#define ABORT_ALL -1\n-#define ABORT_TO_STARSTAR -2\n-\n #define CC_EQ(class, len, litmatch) ((len) == sizeof (litmatch)-1 \\\n \t\t\t\t    && *(class) == *(litmatch) \\\n \t\t\t\t    && strncmp((char*)class, litmatch, len) == 0)\n@@ -90,8 +85,14 @@ static int dowild(const uchar *p, const uchar *text, int force_lower_case)\n \t    continue;\n \t  case '*':\n \t    if (*++p == '*') {\n+\t\tconst uchar *prev_p = p - 2;\n \t\twhile (*++p == '*') {}\n-\t\tspecial = TRUE;\n+\t\tif ((prev_p == text || *prev_p == '/') ||\n+\t\t    (*p == '\\0' || *p == '/' ||\n+\t\t     (p[0] == '\\\\' && p[1] == '/'))) {\n+\t\t    special = TRUE;\n+\t\t} else\n+\t\t    return ABORT_MALFORMED;\n \t    } else\n \t\tspecial = FALSE;\n \t    if (*p == '\\0') {\ndiff --git a/wildmatch.h b/wildmatch.h\nindex e974f9a..984a38c 100644\n--- a/wildmatch.h\n+++ b/wildmatch.h\n@@ -1,3 +1,9 @@\n /* wildmatch.h */\n \n+#define ABORT_MALFORMED 2\n+#define NOMATCH 1\n+#define MATCH 0\n+#define ABORT_ALL -1\n+#define ABORT_TO_STARSTAR -2\n+\n int wildmatch(const char *pattern, const char *text, int flags);\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201111","messageId":"1350182110-25936-12-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 11/12] wildmatch: make /**/ match zero or more directories","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:09Z","receivedAt":"2012-10-14T02:35:09Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\"foo/**/bar\" matches \"foo/x/bar\", \"foo/x/y/bar\"... but not\n\"foo/bar\". We make a special case, when foo/**/ is detected (and\n\"foo/\" part is already matched), try matching \"bar\" with the rest of\nthe string.\n\n\"Match one or more directories\" semantics can be easily achieved using\n\"foo/*/**/bar\".\n\nThis also makes \"**/foo\" match \"foo\" in addition to \"x/foo\",\n\"x/y/foo\"..\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n t/t3070-wildmatch.sh |  8 +++++++-\n wildmatch.c          | 17 +++++++++++++++++\n 2 files changed, 24 insertions(+), 1 deletion(-)\n\ndiff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nindex 15848d5..e6ad6f4 100755\n--- a/t/t3070-wildmatch.sh\n+++ b/t/t3070-wildmatch.sh\n@@ -63,11 +63,17 @@ match 1 1 ']' ']'\n match 0 0 'foo/baz/bar' 'foo*bar'\n match 0 0 'foo/baz/bar' 'foo**bar'\n match 0 1 'foobazbar' 'foo**bar'\n+match 1 1 'foo/baz/bar' 'foo/**/bar'\n+match 1 0 'foo/baz/bar' 'foo/**/**/bar'\n+match 1 0 'foo/b/a/z/bar' 'foo/**/bar'\n+match 1 0 'foo/b/a/z/bar' 'foo/**/**/bar'\n+match 1 0 'foo/bar' 'foo/**/bar'\n+match 1 0 'foo/bar' 'foo/**/**/bar'\n match 0 0 'foo/bar' 'foo?bar'\n match 0 0 'foo/bar' 'foo[/]bar'\n match 0 0 'foo/bar' 'f[^eiu][^eiu][^eiu][^eiu][^eiu]r'\n match 1 1 'foo-bar' 'f[^eiu][^eiu][^eiu][^eiu][^eiu]r'\n-match 0 0 'foo' '**/foo'\n+match 1 0 'foo' '**/foo'\n match 1 x '/foo' '**/foo'\n match 1 0 'bar/baz/foo' '**/foo'\n match 0 0 'bar/baz/foo' '*/foo'\ndiff --git a/wildmatch.c b/wildmatch.c\nindex 7209f26..35c34ac 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -90,6 +90,23 @@ static int dowild(const uchar *p, const uchar *text, int force_lower_case)\n \t\tif ((prev_p == text || *prev_p == '/') ||\n \t\t    (*p == '\\0' || *p == '/' ||\n \t\t     (p[0] == '\\\\' && p[1] == '/'))) {\n+\t\t\t/*\n+\t\t\t * Assuming we already match 'foo/' and are at\n+\t\t\t * <star star slash>, just assume it matches\n+\t\t\t * nothing and go ahead match the rest of the\n+\t\t\t * pattern with the remaining string. This\n+\t\t\t * helps make foo/<*><*>/bar (<> because\n+\t\t\t * otherwise it breaks C comment syntax) match\n+\t\t\t * both foo/bar and foo/a/bar.\n+\t\t\t *\n+\t\t\t * Crazy patterns like /<*><*>/<*><*>/ are\n+\t\t\t * treated like /<*><*>/. But undefined\n+\t\t\t * behavior is even appropriate for people\n+\t\t\t * writing such a pattern.\n+\t\t\t */\n+\t\t\tif (p[0] == '/' &&\n+\t\t\t    dowild(p + 1, text, force_lower_case) == MATCH)\n+\t\t\t\treturn MATCH;\n \t\t    special = TRUE;\n \t\t} else\n \t\t    return ABORT_MALFORMED;\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201112","messageId":"1350182110-25936-13-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1350182110-25936-1-git-send-email-pclouds@gmail.com","subject":"[PATCH v5 12/12] Support \"**\" wildcard in .gitignore and .gitattributes","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T02:35:10Z","receivedAt":"2012-10-14T02:35:10Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Documentation/gitignore.txt        | 19 +++++++++++++++++++\n attr.c                             |  4 +++-\n dir.c                              |  4 +++-\n t/t0003-attributes.sh              | 37 +++++++++++++++++++++++++++++++++++++\n t/t3001-ls-files-others-exclude.sh | 19 +++++++++++++++++++\n 5 files changed, 81 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/gitignore.txt b/Documentation/gitignore.txt\nindex 1b82fe1..91a6438 100644\n--- a/Documentation/gitignore.txt\n+++ b/Documentation/gitignore.txt\n@@ -108,6 +108,25 @@ PATTERN FORMAT\n    For example, \"/{asterisk}.c\" matches \"cat-file.c\" but not\n    \"mozilla-sha1/sha1.c\".\n \n+Two consecutive asterisks (\"`**`\") in patterns matched against\n+full pathname may have special meaning:\n+\n+ - A leading \"`**`\" followed by a slash means match in all\n+   directories. For example, \"`**/foo`\" matches file or directory\n+   \"`foo`\" anywhere, the same as pattern \"`foo`\". \"**/foo/bar\"\n+   matches file or directory \"`bar`\" anywhere that is directly\n+   under directory \"`foo`\".\n+\n+ - A trailing \"/**\" matches everything inside. For example,\n+   \"abc/**\" matches all files inside directory \"abc\", relative\n+   to the location of the `.gitignore` file, with infinite depth.\n+\n+ - A slash followed by two consecutive asterisks then a slash\n+   matches zero or more directories. For example, \"`a/**/b`\"\n+   matches \"`a/b`\", \"`a/x/b`\", \"`a/x/y/b`\" and so on.\n+\n+ - Other consecutive asterisks are considered invalid.\n+\n NOTES\n -----\n \ndiff --git a/attr.c b/attr.c\nindex 887a9ae..8010429 100644\n--- a/attr.c\n+++ b/attr.c\n@@ -12,6 +12,7 @@\n #include \"exec_cmd.h\"\n #include \"attr.h\"\n #include \"dir.h\"\n+#include \"wildmatch.h\"\n \n const char git_attr__true[] = \"(builtin)true\";\n const char git_attr__false[] = \"\\0(builtin)false\";\n@@ -666,7 +667,8 @@ static int path_matches(const char *pathname, int pathlen,\n \t\treturn 0;\n \tif (baselen != 0)\n \t\tbaselen++;\n-\treturn fnmatch_icase(pattern, pathname + baselen, FNM_PATHNAME) == 0;\n+\treturn wildmatch(pattern, pathname + baselen,\n+\t\t\t ignore_case ? FNM_CASEFOLD : 0) == 0;\n }\n \n static int macroexpand_one(int attr_nr, int rem);\ndiff --git a/dir.c b/dir.c\nindex 4868339..442db1c 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -8,6 +8,7 @@\n #include \"cache.h\"\n #include \"dir.h\"\n #include \"refs.h\"\n+#include \"wildmatch.h\"\n \n struct path_simplify {\n \tint len;\n@@ -575,7 +576,8 @@ int excluded_from_list(const char *pathname,\n \t\t\tnamelen -= prefix;\n \t\t}\n \n-\t\tif (!namelen || !fnmatch_icase(exclude, name, FNM_PATHNAME))\n+\t\tif (!namelen ||\n+\t\t    wildmatch(exclude, name, ignore_case ? FNM_CASEFOLD : 0) == 0)\n \t\t\treturn to_exclude;\n \t}\n \treturn -1; /* undecided */\ndiff --git a/t/t0003-attributes.sh b/t/t0003-attributes.sh\nindex febc45c..d5a6946 100755\n--- a/t/t0003-attributes.sh\n+++ b/t/t0003-attributes.sh\n@@ -196,6 +196,43 @@ test_expect_success 'root subdir attribute test' '\n \tattr_check subdir/a/i unspecified\n '\n \n+test_expect_success '\"**\" test' '\n+\techo \"**/f foo=bar\" >.gitattributes &&\n+\tcat <<\\EOF >expect &&\n+f: foo: bar\n+a/f: foo: bar\n+a/b/f: foo: bar\n+a/b/c/f: foo: bar\n+EOF\n+\tgit check-attr foo -- \"f\" >actual 2>err &&\n+\tgit check-attr foo -- \"a/f\" >>actual 2>>err &&\n+\tgit check-attr foo -- \"a/b/f\" >>actual 2>>err &&\n+\tgit check-attr foo -- \"a/b/c/f\" >>actual 2>>err &&\n+\ttest_cmp expect actual &&\n+\ttest_line_count = 0 err\n+'\n+\n+test_expect_success '\"**\" with no slashes test' '\n+\techo \"a**f foo=bar\" >.gitattributes &&\n+\tgit check-attr foo -- \"f\" >actual &&\n+\tcat <<\\EOF >expect &&\n+f: foo: unspecified\n+af: foo: bar\n+axf: foo: bar\n+a/f: foo: unspecified\n+a/b/f: foo: unspecified\n+a/b/c/f: foo: unspecified\n+EOF\n+\tgit check-attr foo -- \"f\" >actual 2>err &&\n+\tgit check-attr foo -- \"af\" >>actual 2>err &&\n+\tgit check-attr foo -- \"axf\" >>actual 2>err &&\n+\tgit check-attr foo -- \"a/f\" >>actual 2>>err &&\n+\tgit check-attr foo -- \"a/b/f\" >>actual 2>>err &&\n+\tgit check-attr foo -- \"a/b/c/f\" >>actual 2>>err &&\n+\ttest_cmp expect actual &&\n+\ttest_line_count = 0 err\n+'\n+\n test_expect_success 'setup bare' '\n \tgit clone --bare . bare.git &&\n \tcd bare.git\ndiff --git a/t/t3001-ls-files-others-exclude.sh b/t/t3001-ls-files-others-exclude.sh\nindex c8fe978..278315d 100755\n--- a/t/t3001-ls-files-others-exclude.sh\n+++ b/t/t3001-ls-files-others-exclude.sh\n@@ -214,4 +214,23 @@ test_expect_success 'subdirectory ignore (l1)' '\n \ttest_cmp expect actual\n '\n \n+\n+test_expect_success 'ls-files with \"**\" patterns' '\n+\tcat <<\\EOF >expect &&\n+a.1\n+one/a.1\n+one/two/a.1\n+three/a.1\n+EOF\n+\tgit ls-files -o -i --exclude \"**/a.1\" >actual\n+\ttest_cmp expect actual\n+'\n+\n+\n+test_expect_success 'ls-files with \"**\" patterns and no slashes' '\n+\t: >expect &&\n+\tgit ls-files -o -i --exclude \"one**a.1\" >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \n1.8.0.rc2.11.g2b79d01\n"},{"id":"201121","messageId":"7va9vp1wa8.fsf@alter.siamese.dyndns.org","threadId":"31810","inReplyTo":"1350182110-25936-3-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v5 02/12] ctype: support iscntrl, ispunct, isxdigit and isprint","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-14T05:02:07Z","receivedAt":"2012-10-14T05:02:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n\nThe description to justify why it is ctype2[] seems to have been\nlost.  Intended?\n\n>  ctype.c           | 18 ++++++++++++++++++\n>  git-compat-util.h | 13 +++++++++++++\n>  2 files changed, 31 insertions(+)\n>\n> diff --git a/ctype.c b/ctype.c\n> index faeaf34..b4bf48a 100644\n> --- a/ctype.c\n> +++ b/ctype.c\n> @@ -26,6 +26,24 @@ const unsigned char sane_ctype[256] = {\n>  \t/* Nothing in the 128.. range */\n>  };\n>  \n> +enum {\n> +\tCN = GIT_CNTRL,\n> +\tPU = GIT_PUNCT,\n> +\tXD = GIT_XDIGIT,\n> +};\n> +\n> +const unsigned char sane_ctype2[256] = {\n> +\tCN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, /*    0..15 */\n> +\tCN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, /*   16..31 */\n> +\t0,  PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, /*   32..47 */\n> +\tXD, XD, XD, XD, XD, XD, XD, XD, XD, XD, PU, PU, PU, PU, PU, PU, /*   48..63 */\n> +\tPU, 0,\tXD, 0,\tXD, 0,\tXD, 0,\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t/*   64..79 */\n> +\t0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  PU, PU, PU, PU, PU, /*   80..95 */\n> +\tPU, 0,\tXD, 0,\tXD, 0,\tXD, 0,\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t/*  96..111 */\n> +\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t0,  PU, PU, PU, PU, CN, /* 112..127 */\n> +\t/* Nothing in the 128.. range */\n> +};\n> +\n>  /* For case-insensitive kwset */\n>  const char tolower_trans_tbl[256] = {\n>  \t0x00, 0x01, 0x02, 0x03, 0x04, 0x05, 0x06, 0x07,\n> diff --git a/git-compat-util.h b/git-compat-util.h\n> index f8b859c..ea11694 100644\n> --- a/git-compat-util.h\n> +++ b/git-compat-util.h\n> @@ -510,14 +510,23 @@ extern const char tolower_trans_tbl[256];\n>  #undef isupper\n>  #undef tolower\n>  #undef toupper\n> +#undef iscntrl\n> +#undef ispunct\n> +#undef isxdigit\n> +#undef isprint\n>  extern const unsigned char sane_ctype[256];\n> +extern const unsigned char sane_ctype2[256];\n>  #define GIT_SPACE 0x01\n>  #define GIT_DIGIT 0x02\n>  #define GIT_ALPHA 0x04\n>  #define GIT_GLOB_SPECIAL 0x08\n>  #define GIT_REGEX_SPECIAL 0x10\n>  #define GIT_PATHSPEC_MAGIC 0x20\n> +#define GIT_CNTRL 0x01\n> +#define GIT_PUNCT 0x02\n> +#define GIT_XDIGIT 0x04\n>  #define sane_istest(x,mask) ((sane_ctype[(unsigned char)(x)] & (mask)) != 0)\n> +#define sane_istest2(x,mask) ((sane_ctype2[(unsigned char)(x)] & (mask)) != 0)\n>  #define isascii(x) (((x) & ~0x7f) == 0)\n>  #define isspace(x) sane_istest(x,GIT_SPACE)\n>  #define isdigit(x) sane_istest(x,GIT_DIGIT)\n> @@ -527,6 +536,10 @@ extern const unsigned char sane_ctype[256];\n>  #define isupper(x) sane_iscase(x, 0)\n>  #define is_glob_special(x) sane_istest(x,GIT_GLOB_SPECIAL)\n>  #define is_regex_special(x) sane_istest(x,GIT_GLOB_SPECIAL | GIT_REGEX_SPECIAL)\n> +#define iscntrl(x) sane_istest2(x, GIT_CNTRL)\n> +#define ispunct(x) sane_istest2(x, GIT_PUNCT)\n> +#define isxdigit(x) sane_istest2(x, GIT_XDIGIT)\n> +#define isprint(x) (isalnum(x) || isspace(x) || ispunct(x))\n>  #define tolower(x) sane_case((unsigned char)(x), 0x20)\n>  #define toupper(x) sane_case((unsigned char)(x), 0)\n>  #define is_pathspec_magic(x) sane_istest(x,GIT_PATHSPEC_MAGIC)\n"},{"id":"201124","messageId":"7vr4p1zl2s.fsf@alter.siamese.dyndns.org","threadId":"31810","inReplyTo":"1350182110-25936-5-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v5 04/12] wildmatch: remove unnecessary functions","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-14T05:04:00Z","receivedAt":"2012-10-14T05:04:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n\nThe comment-fix seems to be new but otherwise this is unchanged,\nright?\n\n\n>  wildmatch.c | 164 ++++--------------------------------------------------------\n>  wildmatch.h |   2 -\n>  2 files changed, 10 insertions(+), 156 deletions(-)\n>\n> diff --git a/wildmatch.c b/wildmatch.c\n> index f3a1731..fae7397 100644\n> --- a/wildmatch.c\n> +++ b/wildmatch.c\n> @@ -53,33 +53,18 @@\n>  #define ISUPPER(c) (ISASCII(c) && isupper(c))\n>  #define ISXDIGIT(c) (ISASCII(c) && isxdigit(c))\n>  \n> -#ifdef WILD_TEST_ITERATIONS\n> -int wildmatch_iteration_count;\n> -#endif\n> -\n>  static int force_lower_case = 0;\n>  \n> -/* Match pattern \"p\" against the a virtually-joined string consisting\n> - * of \"text\" and any strings in array \"a\". */\n> -static int dowild(const uchar *p, const uchar *text, const uchar*const *a)\n> +/* Match pattern \"p\" against \"text\" */\n> +static int dowild(const uchar *p, const uchar *text)\n>  {\n>      uchar p_ch;\n>  \n> -#ifdef WILD_TEST_ITERATIONS\n> -    wildmatch_iteration_count++;\n> -#endif\n> -\n>      for ( ; (p_ch = *p) != '\\0'; text++, p++) {\n>  \tint matched, special;\n>  \tuchar t_ch, prev_ch;\n> -\twhile ((t_ch = *text) == '\\0') {\n> -\t    if (*a == NULL) {\n> -\t\tif (p_ch != '*')\n> -\t\t    return ABORT_ALL;\n> -\t\tbreak;\n> -\t    }\n> -\t    text = *a++;\n> -\t}\n> +\tif ((t_ch = *text) == '\\0' && p_ch != '*')\n> +\t\treturn ABORT_ALL;\n>  \tif (force_lower_case && ISUPPER(t_ch))\n>  \t    t_ch = tolower(t_ch);\n>  \tswitch (p_ch) {\n> @@ -107,21 +92,15 @@ static int dowild(const uchar *p, const uchar *text, const uchar*const *a)\n>  \t\t/* Trailing \"**\" matches everything.  Trailing \"*\" matches\n>  \t\t * only if there are no more slash characters. */\n>  \t\tif (!special) {\n> -\t\t    do {\n>  \t\t\tif (strchr((char*)text, '/') != NULL)\n>  \t\t\t    return FALSE;\n> -\t\t    } while ((text = *a++) != NULL);\n>  \t\t}\n>  \t\treturn TRUE;\n>  \t    }\n>  \t    while (1) {\n> -\t\tif (t_ch == '\\0') {\n> -\t\t    if ((text = *a++) == NULL)\n> -\t\t\tbreak;\n> -\t\t    t_ch = *text;\n> -\t\t    continue;\n> -\t\t}\n> -\t\tif ((matched = dowild(p, text, a)) != FALSE) {\n> +\t\tif (t_ch == '\\0')\n> +\t\t    break;\n> +\t\tif ((matched = dowild(p, text)) != FALSE) {\n>  \t\t    if (!special || matched != ABORT_TO_STARSTAR)\n>  \t\t\treturn matched;\n>  \t\t} else if (!special && t_ch == '/')\n> @@ -225,144 +204,21 @@ static int dowild(const uchar *p, const uchar *text, const uchar*const *a)\n>  \t}\n>      }\n>  \n> -    do {\n> -\tif (*text)\n> -\t    return FALSE;\n> -    } while ((text = *a++) != NULL);\n> -\n> -    return TRUE;\n> -}\n> -\n> -/* Match literal string \"s\" against the a virtually-joined string consisting\n> - * of \"text\" and any strings in array \"a\". */\n> -static int doliteral(const uchar *s, const uchar *text, const uchar*const *a)\n> -{\n> -    for ( ; *s != '\\0'; text++, s++) {\n> -\twhile (*text == '\\0') {\n> -\t    if ((text = *a++) == NULL)\n> -\t\treturn FALSE;\n> -\t}\n> -\tif (*text != *s)\n> -\t    return FALSE;\n> -    }\n> -\n> -    do {\n> -\tif (*text)\n> -\t    return FALSE;\n> -    } while ((text = *a++) != NULL);\n> -\n> -    return TRUE;\n> -}\n> -\n> -/* Return the last \"count\" path elements from the concatenated string.\n> - * We return a string pointer to the start of the string, and update the\n> - * array pointer-pointer to point to any remaining string elements. */\n> -static const uchar *trailing_N_elements(const uchar*const **a_ptr, int count)\n> -{\n> -    const uchar*const *a = *a_ptr;\n> -    const uchar*const *first_a = a;\n> -\n> -    while (*a)\n> -\t    a++;\n> -\n> -    while (a != first_a) {\n> -\tconst uchar *s = *--a;\n> -\ts += strlen((char*)s);\n> -\twhile (--s >= *a) {\n> -\t    if (*s == '/' && !--count) {\n> -\t\t*a_ptr = a+1;\n> -\t\treturn s+1;\n> -\t    }\n> -\t}\n> -    }\n> -\n> -    if (count == 1) {\n> -\t*a_ptr = a+1;\n> -\treturn *a;\n> -    }\n> -\n> -    return NULL;\n> +    return *text ? FALSE : TRUE;\n>  }\n>  \n>  /* Match the \"pattern\" against the \"text\" string. */\n>  int wildmatch(const char *pattern, const char *text)\n>  {\n> -    static const uchar *nomore[1]; /* A NULL pointer. */\n> -#ifdef WILD_TEST_ITERATIONS\n> -    wildmatch_iteration_count = 0;\n> -#endif\n> -    return dowild((const uchar*)pattern, (const uchar*)text, nomore) == TRUE;\n> +    return dowild((const uchar*)pattern, (const uchar*)text) == TRUE;\n>  }\n>  \n>  /* Match the \"pattern\" against the forced-to-lower-case \"text\" string. */\n>  int iwildmatch(const char *pattern, const char *text)\n>  {\n> -    static const uchar *nomore[1]; /* A NULL pointer. */\n>      int ret;\n> -#ifdef WILD_TEST_ITERATIONS\n> -    wildmatch_iteration_count = 0;\n> -#endif\n>      force_lower_case = 1;\n> -    ret = dowild((const uchar*)pattern, (const uchar*)text, nomore) == TRUE;\n> +    ret = dowild((const uchar*)pattern, (const uchar*)text) == TRUE;\n>      force_lower_case = 0;\n>      return ret;\n>  }\n> -\n> -/* Match pattern \"p\" against the a virtually-joined string consisting\n> - * of all the pointers in array \"texts\" (which has a NULL pointer at the\n> - * end).  The int \"where\" can be 0 (normal matching), > 0 (match only\n> - * the trailing N slash-separated filename components of \"texts\"), or < 0\n> - * (match the \"pattern\" at the start or after any slash in \"texts\"). */\n> -int wildmatch_array(const char *pattern, const char*const *texts, int where)\n> -{\n> -    const uchar *p = (const uchar*)pattern;\n> -    const uchar*const *a = (const uchar*const*)texts;\n> -    const uchar *text;\n> -    int matched;\n> -\n> -#ifdef WILD_TEST_ITERATIONS\n> -    wildmatch_iteration_count = 0;\n> -#endif\n> -\n> -    if (where > 0)\n> -\ttext = trailing_N_elements(&a, where);\n> -    else\n> -\ttext = *a++;\n> -    if (!text)\n> -\treturn FALSE;\n> -\n> -    if ((matched = dowild(p, text, a)) != TRUE && where < 0\n> -     && matched != ABORT_ALL) {\n> -\twhile (1) {\n> -\t    if (*text == '\\0') {\n> -\t\tif ((text = (uchar*)*a++) == NULL)\n> -\t\t    return FALSE;\n> -\t\tcontinue;\n> -\t    }\n> -\t    if (*text++ == '/' && (matched = dowild(p, text, a)) != FALSE\n> -\t     && matched != ABORT_TO_STARSTAR)\n> -\t\tbreak;\n> -\t}\n> -    }\n> -    return matched == TRUE;\n> -}\n> -\n> -/* Match literal string \"s\" against the a virtually-joined string consisting\n> - * of all the pointers in array \"texts\" (which has a NULL pointer at the\n> - * end).  The int \"where\" can be 0 (normal matching), or > 0 (match\n> - * only the trailing N slash-separated filename components of \"texts\"). */\n> -int litmatch_array(const char *string, const char*const *texts, int where)\n> -{\n> -    const uchar *s = (const uchar*)string;\n> -    const uchar*const *a = (const uchar* const*)texts;\n> -    const uchar *text;\n> -\n> -    if (where > 0)\n> -\ttext = trailing_N_elements(&a, where);\n> -    else\n> -\ttext = *a++;\n> -    if (!text)\n> -\treturn FALSE;\n> -\n> -    return doliteral(s, text, a) == TRUE;\n> -}\n> diff --git a/wildmatch.h b/wildmatch.h\n> index e7f1a35..562faa3 100644\n> --- a/wildmatch.h\n> +++ b/wildmatch.h\n> @@ -2,5 +2,3 @@\n>  \n>  int wildmatch(const char *pattern, const char *text);\n>  int iwildmatch(const char *pattern, const char *text);\n> -int wildmatch_array(const char *pattern, const char*const *texts, int where);\n> -int litmatch_array(const char *string, const char*const *texts, int where);\n"},{"id":"201125","messageId":"7vlif9zl2n.fsf@alter.siamese.dyndns.org","threadId":"31810","inReplyTo":"1350182110-25936-6-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v5 05/12] Integrate wildmatch to git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-14T05:06:22Z","receivedAt":"2012-10-14T05:06:22Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> +++ b/t/t3070-wildmatch.sh\n> @@ -0,0 +1,188 @@\n> +#!/bin/sh\n> +\n> +test_description='wildmatch tests'\n> +\n> +. ./test-lib.sh\n> +\n> +match() {\n> +    if [ $1 = 1 ]; then\n> +\ttest_expect_success \"wildmatch:    match '$3' '$4'\" \"\n> +\t    test-wildmatch wildmatch '$3' '$4'\n> +\t\"\n> +    else\n> +\ttest_expect_success \"wildmatch: no match '$3' '$4'\" \"\n> +\t    ! test-wildmatch wildmatch '$3' '$4'\n> +\t\"\n> +    fi\n> +    if [ $2 = 1 ]; then\n> +\ttest_expect_success \"fnmatch:      match '$3' '$4'\" \"\n> +\t    test-wildmatch fnmatch '$3' '$4'\n> +\t\"\n> +    elif [ $2 = 0 ]; then\n> +\ttest_expect_success \"fnmatch:   no match '$3' '$4'\" \"\n> +\t    ! test-wildmatch fnmatch '$3' '$4'\n> +\t\"\n> +#    else\n> +#\ttest_expect_success BROKEN_FNMATCH \"fnmatch:       '$3' '$4'\" \"\n> +#\t    ! test-wildmatch fnmatch '$3' '$4'\n> +#\t\"\n> +    fi\n\nHeh, broken can be two-way.  Either it may succeed matching what it\nshouldn't, or it may not match what it should.\n"},{"id":"201122","messageId":"CACsJy8B6Yo5zZsf8aBF=fPgUfmAfNEnnDNkh8+WF28GcGuyidw@mail.gmail.com","threadId":"31810","inReplyTo":"7va9vp1wa8.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH v5 02/12] ctype: support iscntrl, ispunct, isxdigit and isprint","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T05:07:11Z","receivedAt":"2012-10-14T05:07:11Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Oct 14, 2012 at 12:02 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n>\n>> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n>> ---\n>\n> The description to justify why it is ctype2[] seems to have been\n> lost.  Intended?\n\nNope. I added the description after generating patches and forgot to\nupdate the same to my branch. Thanks for catching.\n-- \nDuy\n"},{"id":"201123","messageId":"7vwqytzlkk.fsf@alter.siamese.dyndns.org","threadId":"31810","inReplyTo":"1350182110-25936-8-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v5 07/12] wildmatch: make wildmatch's return value compatible with fnmatch","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-10-14T05:09:31Z","receivedAt":"2012-10-14T05:09:31Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n\n> wildmatch returns non-zero if matched, zero otherwise. This patch\n> makes it return zero if matches, non-zero otherwise, like fnmatch().\n>\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n\nOK, so ABORT cases where the patterns are either broken or impossible\nto match are also taken as not-matching, which sounds like the right\nthing to do.\n\n>  test-wildmatch.c |  4 ++--\n>  wildmatch.c      | 21 ++++++++++++---------\n>  2 files changed, 14 insertions(+), 11 deletions(-)\n>\n> diff --git a/test-wildmatch.c b/test-wildmatch.c\n> index ac56420..77014e9 100644\n> --- a/test-wildmatch.c\n> +++ b/test-wildmatch.c\n> @@ -4,9 +4,9 @@\n>  int main(int argc, char **argv)\n>  {\n>  \tif (!strcmp(argv[1], \"wildmatch\"))\n> -\t\treturn wildmatch(argv[3], argv[2]) ? 0 : 1;\n> +\t\treturn !!wildmatch(argv[3], argv[2]);\n>  \telse if (!strcmp(argv[1], \"iwildmatch\"))\n> -\t\treturn iwildmatch(argv[3], argv[2]) ? 0 : 1;\n> +\t\treturn !!iwildmatch(argv[3], argv[2]);\n>  \telse if (!strcmp(argv[1], \"fnmatch\"))\n>  \t\treturn !!fnmatch(argv[3], argv[2], FNM_PATHNAME);\n>  \telse\n> diff --git a/wildmatch.c b/wildmatch.c\n> index d0b906a..e3ac6cc 100644\n> --- a/wildmatch.c\n> +++ b/wildmatch.c\n> @@ -20,6 +20,9 @@ typedef unsigned char uchar;\n>  \n>  #define FALSE 0\n>  #define TRUE 1\n> +\n> +#define NOMATCH 1\n> +#define MATCH 0\n>  #define ABORT_ALL -1\n>  #define ABORT_TO_STARSTAR -2\n>  \n> @@ -78,12 +81,12 @@ static int dowild(const uchar *p, const uchar *text)\n>  \t    /* FALLTHROUGH */\n>  \t  default:\n>  \t    if (t_ch != p_ch)\n> -\t\treturn FALSE;\n> +\t\treturn NOMATCH;\n>  \t    continue;\n>  \t  case '?':\n>  \t    /* Match anything but '/'. */\n>  \t    if (t_ch == '/')\n> -\t\treturn FALSE;\n> +\t\treturn NOMATCH;\n>  \t    continue;\n>  \t  case '*':\n>  \t    if (*++p == '*') {\n> @@ -96,14 +99,14 @@ static int dowild(const uchar *p, const uchar *text)\n>  \t\t * only if there are no more slash characters. */\n>  \t\tif (!special) {\n>  \t\t\tif (strchr((char*)text, '/') != NULL)\n> -\t\t\t    return FALSE;\n> +\t\t\t    return NOMATCH;\n>  \t\t}\n> -\t\treturn TRUE;\n> +\t\treturn MATCH;\n>  \t    }\n>  \t    while (1) {\n>  \t\tif (t_ch == '\\0')\n>  \t\t    break;\n> -\t\tif ((matched = dowild(p, text)) != FALSE) {\n> +\t\tif ((matched = dowild(p, text)) != NOMATCH) {\n>  \t\t    if (!special || matched != ABORT_TO_STARSTAR)\n>  \t\t\treturn matched;\n>  \t\t} else if (!special && t_ch == '/')\n> @@ -202,18 +205,18 @@ static int dowild(const uchar *p, const uchar *text)\n>  \t\t    matched = TRUE;\n>  \t    } while (prev_ch = p_ch, (p_ch = *++p) != ']');\n>  \t    if (matched == special || t_ch == '/')\n> -\t\treturn FALSE;\n> +\t\treturn NOMATCH;\n>  \t    continue;\n>  \t}\n>      }\n>  \n> -    return *text ? FALSE : TRUE;\n> +    return *text ? NOMATCH : MATCH;\n>  }\n>  \n>  /* Match the \"pattern\" against the \"text\" string. */\n>  int wildmatch(const char *pattern, const char *text)\n>  {\n> -    return dowild((const uchar*)pattern, (const uchar*)text) == TRUE;\n> +    return dowild((const uchar*)pattern, (const uchar*)text);\n>  }\n>  \n>  /* Match the \"pattern\" against the forced-to-lower-case \"text\" string. */\n> @@ -221,7 +224,7 @@ int iwildmatch(const char *pattern, const char *text)\n>  {\n>      int ret;\n>      force_lower_case = 1;\n> -    ret = dowild((const uchar*)pattern, (const uchar*)text) == TRUE;\n> +    ret = dowild((const uchar*)pattern, (const uchar*)text);\n>      force_lower_case = 0;\n>      return ret;\n>  }\n"},{"id":"201151","messageId":"CACsJy8DsH9OfkdJw3MnkvUidjgdUFP_ODdjLnj25jgKbB9yr7w@mail.gmail.com","threadId":"31810","inReplyTo":"7vr4p1zl2s.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH v5 04/12] wildmatch: remove unnecessary functions","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T06:29:36Z","receivedAt":"2012-10-14T06:29:36Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Oct 14, 2012 at 12:04 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Nguyễn Thái Ngọc Duy  <pclouds@gmail.com> writes:\n>\n>> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n>> ---\n>\n> The comment-fix seems to be new but otherwise this is unchanged,\n> right?\n\nRight.--\nDuy\n"},{"id":"201160","messageId":"507A9CF8.1040603@web.de","threadId":"31810","inReplyTo":"1350182110-25936-6-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v5 05/12] Integrate wildmatch to git","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2012-10-14T11:07:36Z","receivedAt":"2012-10-14T11:07:36Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"diff --git a/t/t3070-wildmatch.sh b/t/t3070-wildmatch.sh\nnew file mode 100755\nindex 0000000..dbd3c8b\n--- /dev/null\n+++ b/t/t3070-wildmatch.sh\n@@ -0,0 +1,188 @@\n+#!/bin/sh\n+#    else\n+#    test_expect_success BROKEN_FNMATCH \"fnmatch:       '$3' '$4'\" \"\n+#        ! test-wildmatch fnmatch '$3' '$4'\n+#    \"\n+    fi\n+}\n+\n\nThanks:\nOn my Mac OS X box:\n# passed all 259 test(s)\n\nAnd a quick test on cygwin:\n$ ./t3070-wildmatch.sh  2>&1 | grep \"not ok\"\nnot ok - 148 fnmatch:      match '5' '[[:xdigit:]]'\nnot ok - 150 fnmatch:      match 'f' '[[:xdigit:]]'\nnot ok - 152 fnmatch:      match 'D' '[[:xdigit:]]'\n\n\nAnd 2 micronits:\na) Commented out code\nb) Whithespace damage\n( 4 spaces used for an indent of 1, TAB for indent of 2)\n\n/Torsten\n"},{"id":"201170","messageId":"507AB73D.8010406@lsrfire.ath.cx","threadId":"31810","inReplyTo":"1350182110-25936-3-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH v5 02/12] ctype: support iscntrl, ispunct, isxdigit and isprint","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2012-10-14T12:59:41Z","receivedAt":"2012-10-14T12:59:41Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 14.10.2012 04:35, schrieb Nguyễn Thái Ngọc Duy:\n>\n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>   ctype.c           | 18 ++++++++++++++++++\n>   git-compat-util.h | 13 +++++++++++++\n>   2 files changed, 31 insertions(+)\n>\n> diff --git a/ctype.c b/ctype.c\n> index faeaf34..b4bf48a 100644\n> --- a/ctype.c\n> +++ b/ctype.c\n> @@ -26,6 +26,24 @@ const unsigned char sane_ctype[256] = {\n>   \t/* Nothing in the 128.. range */\n>   };\n>\n> +enum {\n> +\tCN = GIT_CNTRL,\n> +\tPU = GIT_PUNCT,\n> +\tXD = GIT_XDIGIT,\n> +};\n> +\n> +const unsigned char sane_ctype2[256] = {\n> +\tCN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, /*    0..15 */\n> +\tCN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, /*   16..31 */\n> +\t0,  PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, /*   32..47 */\n> +\tXD, XD, XD, XD, XD, XD, XD, XD, XD, XD, PU, PU, PU, PU, PU, PU, /*   48..63 */\n> +\tPU, 0,\tXD, 0,\tXD, 0,\tXD, 0,\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t/*   64..79 */\n> +\t0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  PU, PU, PU, PU, PU, /*   80..95 */\n> +\tPU, 0,\tXD, 0,\tXD, 0,\tXD, 0,\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t/*  96..111 */\n> +\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t0,  0,\t0,  PU, PU, PU, PU, CN, /* 112..127 */\n\nShouldn't [ace] (65, 67, 69) and [ACE] (97, 99, 101) be xdigits as well?\n\nBut how about using the existing hexval_table instead, like this:\n\n\t#define isxdigit(x) (hexval_table[(x)] != -1)\n\nWith that, couldn't you squeeze the other two classes into the existing \nsane_type?\n\nBy the way, I'm working on a patch series for implementing a lot more \ncharacter classes with table lookups.  It grew out of a desire to make \nbad_ref_char() faster but perhaps got a bit out of hand by now; it's at \n24 patches and still not finished.  I'm curious how long we have until \nit escapes. ;-)\n\n>  #define is_regex_special(x) sane_istest(x,GIT_GLOB_SPECIAL | GIT_REGEX_SPECIAL)\n> +#define iscntrl(x) sane_istest2(x, GIT_CNTRL)\n> +#define ispunct(x) sane_istest2(x, GIT_PUNCT)\n> +#define isxdigit(x) sane_istest2(x, GIT_XDIGIT)\n> +#define isprint(x) (isalnum(x) || isspace(x) || ispunct(x))\n\nIf a single table is used, you can do with a single table lookup by \nadding the bits for the component classes, like isalnum and \nis_regex_special do.\n\nRené\n"},{"id":"201171","messageId":"CACsJy8B+6OPkP6ijMDzm+n0eHnDZ4Pj8UO_KasdfEP4wF+_hww@mail.gmail.com","threadId":"31810","inReplyTo":"507AB73D.8010406@lsrfire.ath.cx","subject":"Re: [PATCH v5 02/12] ctype: support iscntrl, ispunct, isxdigit and isprint","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T13:25:10Z","receivedAt":"2012-10-14T13:25:10Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Oct 14, 2012 at 7:59 PM, René Scharfe\n<rene.scharfe@lsrfire.ath.cx> wrote:\n>> +const unsigned char sane_ctype2[256] = {\n>> +       CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, /*\n>> 0..15 */\n>> +       CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, CN, /*\n>> 16..31 */\n>> +       0,  PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, PU, /*\n>> 32..47 */\n>> +       XD, XD, XD, XD, XD, XD, XD, XD, XD, XD, PU, PU, PU, PU, PU, PU, /*\n>> 48..63 */\n>> +       PU, 0,  XD, 0,  XD, 0,  XD, 0,  0,  0,  0,  0,  0,  0,  0,  0,  /*\n>> 64..79 */\n>> +       0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  PU, PU, PU, PU, PU, /*\n>> 80..95 */\n>> +       PU, 0,  XD, 0,  XD, 0,  XD, 0,  0,  0,  0,  0,  0,  0,  0,  0,  /*\n>> 96..111 */\n>> +       0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  0,  PU, PU, PU, PU, CN, /*\n>> 112..127 */\n>\n>\n> Shouldn't [ace] (65, 67, 69) and [ACE] (97, 99, 101) be xdigits as well?\n\nHmm.. I generated it from LANG=C. I wonder where I got it wrong..\n\n> But how about using the existing hexval_table instead, like this:\n>\n>         #define isxdigit(x) (hexval_table[(x)] != -1)\n>\n> With that, couldn't you squeeze the other two classes into the existing\n> sane_type?\n\nNo there are still conflicts: 9, 10 and 13 as spaces (vs controls) and\n123, 124 and 126 as regex/pathspec special (vs punctuation).\n\n> By the way, I'm working on a patch series for implementing a lot more\n> character classes with table lookups.  It grew out of a desire to make\n> bad_ref_char() faster but perhaps got a bit out of hand by now; it's at 24\n> patches and still not finished.  I'm curious how long we have until it\n> escapes. ;-)\n\nI don't think the series is going to graduate any time soon :)\n-- \nDuy\n"},{"id":"201174","messageId":"507AC543.2020402@lsrfire.ath.cx","threadId":"31810","inReplyTo":"CACsJy8B+6OPkP6ijMDzm+n0eHnDZ4Pj8UO_KasdfEP4wF+_hww@mail.gmail.com","subject":"Re: [PATCH v5 02/12] ctype: support iscntrl, ispunct, isxdigit and isprint","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2012-10-14T13:59:31Z","receivedAt":"2012-10-14T13:59:31Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 14.10.2012 15:25, schrieb Nguyen Thai Ngoc Duy:\n> On Sun, Oct 14, 2012 at 7:59 PM, René Scharfe\n> <rene.scharfe@lsrfire.ath.cx> wrote:\n>> With that, couldn't you squeeze the other two classes into the existing\n>> sane_type?\n>\n> No there are still conflicts: 9, 10 and 13 as spaces (vs controls) and\n> 123, 124 and 126 as regex/pathspec special (vs punctuation).\n\nThat's not a problem, an entry in the table can have more than one bit \nset -- just OR them together in ctype.c.  It may not look as nice, but \nthat's OK.  You could also define a character for GIT_SPACE | GIT_CNTRL \netc. for cosmetic reasons.\n\nRené\n"},{"id":"201175","messageId":"20121014142624.GA992@do","threadId":"31810","inReplyTo":"507AC543.2020402@lsrfire.ath.cx","subject":"Re: [PATCH v5 02/12] ctype: support iscntrl, ispunct, isxdigit and isprint","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-14T14:26:24Z","receivedAt":"2012-10-14T14:26:24Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sun, Oct 14, 2012 at 03:59:31PM +0200, René Scharfe wrote:\n> Am 14.10.2012 15:25, schrieb Nguyen Thai Ngoc Duy:\n> > On Sun, Oct 14, 2012 at 7:59 PM, René Scharfe\n> > <rene.scharfe@lsrfire.ath.cx> wrote:\n> >> With that, couldn't you squeeze the other two classes into the existing\n> >> sane_type?\n> >\n> > No there are still conflicts: 9, 10 and 13 as spaces (vs controls) and\n> > 123, 124 and 126 as regex/pathspec special (vs punctuation).\n> \n> That's not a problem, an entry in the table can have more than one bit \n> set -- just OR them together in ctype.c.  It may not look as nice, but \n> that's OK.  You could also define a character for GIT_SPACE | GIT_CNTRL \n> etc. for cosmetic reasons.\n\nOnly space chars is not a subset of control chars, which needs a new\ncombination. So the result does not look as bad as I thought:\n\n-- 8< --\ndiff --git a/ctype.c b/ctype.c\nindex faeaf34..0bfebb4 100644\n--- a/ctype.c\n+++ b/ctype.c\n@@ -11,18 +11,21 @@ enum {\n \tD = GIT_DIGIT,\n \tG = GIT_GLOB_SPECIAL,\t/* *, ?, [, \\\\ */\n \tR = GIT_REGEX_SPECIAL,\t/* $, (, ), +, ., ^, {, | */\n-\tP = GIT_PATHSPEC_MAGIC  /* other non-alnum, except for ] and } */\n+\tP = GIT_PATHSPEC_MAGIC, /* other non-alnum, except for ] and } */\n+\tX = GIT_CNTRL,\n+\tU = GIT_PUNCT,\n+\tZ = GIT_CNTRL | GIT_SPACE\n };\n \n const unsigned char sane_ctype[256] = {\n-\t0, 0, 0, 0, 0, 0, 0, 0, 0, S, S, 0, 0, S, 0, 0,\t\t/*   0.. 15 */\n-\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\t\t/*  16.. 31 */\n+\tX, X, X, X, X, X, X, X, X, Z, Z, X, X, Z, X, X,\t\t/*   0.. 15 */\n+\tX, X, X, X, X, X, X, X, X, X, X, X, X, X, X, X,\t\t/*  16.. 31 */\n \tS, P, P, P, R, P, P, P, R, R, G, R, P, P, R, P,\t\t/*  32.. 47 */\n \tD, D, D, D, D, D, D, D, D, D, P, P, P, P, P, G,\t\t/*  48.. 63 */\n \tP, A, A, A, A, A, A, A, A, A, A, A, A, A, A, A,\t\t/*  64.. 79 */\n-\tA, A, A, A, A, A, A, A, A, A, A, G, G, 0, R, P,\t\t/*  80.. 95 */\n+\tA, A, A, A, A, A, A, A, A, A, A, G, G, U, R, P,\t\t/*  80.. 95 */\n \tP, A, A, A, A, A, A, A, A, A, A, A, A, A, A, A,\t\t/*  96..111 */\n-\tA, A, A, A, A, A, A, A, A, A, A, R, R, 0, P, 0,\t\t/* 112..127 */\n+\tA, A, A, A, A, A, A, A, A, A, A, R, R, U, P, X,\t\t/* 112..127 */\n \t/* Nothing in the 128.. range */\n };\n \ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex f8b859c..db77f3e 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -510,6 +510,10 @@ extern const char tolower_trans_tbl[256];\n #undef isupper\n #undef tolower\n #undef toupper\n+#undef iscntrl\n+#undef ispunct\n+#undef isxdigit\n+#undef isprint\n extern const unsigned char sane_ctype[256];\n #define GIT_SPACE 0x01\n #define GIT_DIGIT 0x02\n@@ -517,6 +521,8 @@ extern const unsigned char sane_ctype[256];\n #define GIT_GLOB_SPECIAL 0x08\n #define GIT_REGEX_SPECIAL 0x10\n #define GIT_PATHSPEC_MAGIC 0x20\n+#define GIT_CNTRL 0x40\n+#define GIT_PUNCT 0x80\n #define sane_istest(x,mask) ((sane_ctype[(unsigned char)(x)] & (mask)) != 0)\n #define isascii(x) (((x) & ~0x7f) == 0)\n #define isspace(x) sane_istest(x,GIT_SPACE)\n@@ -527,6 +533,13 @@ extern const unsigned char sane_ctype[256];\n #define isupper(x) sane_iscase(x, 0)\n #define is_glob_special(x) sane_istest(x,GIT_GLOB_SPECIAL)\n #define is_regex_special(x) sane_istest(x,GIT_GLOB_SPECIAL | GIT_REGEX_SPECIAL)\n+#define iscntrl(x) (sane_istest(x,GIT_CNTRL))\n+#define ispunct(x) sane_istest(x, GIT_PUNCT | GIT_REGEX_SPECIAL | \\\n+\t\tGIT_GLOB_SPECIAL | GIT_PATHSPEC_MAGIC)\n+#define isxdigit(x) (hexval_table[x] != -1)\n+#define isprint(x) (sane_istest(x, GIT_ALPHA | GIT_DIGIT | GIT_SPACE | \\\n+\t\tGIT_PUNCT | GIT_REGEX_SPECIAL | GIT_GLOB_SPECIAL | \\\n+\t\tGIT_PATHSPEC_MAGIC))\n #define tolower(x) sane_case((unsigned char)(x), 0x20)\n #define toupper(x) sane_case((unsigned char)(x), 0)\n #define is_pathspec_magic(x) sane_istest(x,GIT_PATHSPEC_MAGIC)\n-- 8< --\n\n-- \nDuy\n"},{"id":"201420","messageId":"507E9FDE.7080706@cs.tu-berlin.de","threadId":"31810","inReplyTo":"20121014142624.GA992@do","subject":"Re: [PATCH v5 02/12] ctype: support iscntrl, ispunct, isxdigit and isprint","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-10-17T12:09:02Z","receivedAt":"2012-10-17T12:09:02Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"Hi Nguyen.\n\nI just had a need for isprint() myself, and then I found\nyour code here.\n\nI had a look at the POSIX locale as describe here:\n\nhttp://sourceware.org/git/?p=glibc.git;a=blob;f=localedata/locales/POSIX\n\nSome remarks below.\n\nAm 14.10.2012 16:26, schrieb Nguyen Thai Ngoc Duy:\n> -- 8< --\n> diff --git a/ctype.c b/ctype.c\n> index faeaf34..0bfebb4 100644\n> --- a/ctype.c\n> +++ b/ctype.c\n> @@ -11,18 +11,21 @@ enum {\n>  \tD = GIT_DIGIT,\n>  \tG = GIT_GLOB_SPECIAL,\t/* *, ?, [, \\\\ */\n>  \tR = GIT_REGEX_SPECIAL,\t/* $, (, ), +, ., ^, {, | */\n> -\tP = GIT_PATHSPEC_MAGIC  /* other non-alnum, except for ] and } */\n> +\tP = GIT_PATHSPEC_MAGIC, /* other non-alnum, except for ] and } */\n> +\tX = GIT_CNTRL,\n> +\tU = GIT_PUNCT,\n> +\tZ = GIT_CNTRL | GIT_SPACE\n>  };\n>  \n>  const unsigned char sane_ctype[256] = {\n> -\t0, 0, 0, 0, 0, 0, 0, 0, 0, S, S, 0, 0, S, 0, 0,\t\t/*   0.. 15 */\n> -\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\t\t/*  16.. 31 */\n> +\tX, X, X, X, X, X, X, X, X, Z, Z, X, X, Z, X, X,\t\t/*   0.. 15 */\n> +\tX, X, X, X, X, X, X, X, X, X, X, X, X, X, X, X,\t\t/*  16.. 31 */\n\n\"Normal\" isspace() also includes vertical tab (11) and form-feed (12) as\nwhite-space characters. Is there a reason, why they are not included here?\n\n>  \tS, P, P, P, R, P, P, P, R, R, G, R, P, P, R, P,\t\t/*  32.. 47 */\n>  \tD, D, D, D, D, D, D, D, D, D, P, P, P, P, P, G,\t\t/*  48.. 63 */\n>  \tP, A, A, A, A, A, A, A, A, A, A, A, A, A, A, A,\t\t/*  64.. 79 */\n> -\tA, A, A, A, A, A, A, A, A, A, A, G, G, 0, R, P,\t\t/*  80.. 95 */\n> +\tA, A, A, A, A, A, A, A, A, A, A, G, G, U, R, P,\t\t/*  80.. 95 */\n>  \tP, A, A, A, A, A, A, A, A, A, A, A, A, A, A, A,\t\t/*  96..111 */\n> -\tA, A, A, A, A, A, A, A, A, A, A, R, R, 0, P, 0,\t\t/* 112..127 */\n> +\tA, A, A, A, A, A, A, A, A, A, A, R, R, U, P, X,\t\t/* 112..127 */\n>  \t/* Nothing in the 128.. range */\n>  };\n>  \n> diff --git a/git-compat-util.h b/git-compat-util.h\n> index f8b859c..db77f3e 100644\n> --- a/git-compat-util.h\n> +++ b/git-compat-util.h\n[...]\n> @@ -527,6 +533,13 @@ extern const unsigned char sane_ctype[256];\n>  #define isupper(x) sane_iscase(x, 0)\n>  #define is_glob_special(x) sane_istest(x,GIT_GLOB_SPECIAL)\n>  #define is_regex_special(x) sane_istest(x,GIT_GLOB_SPECIAL | GIT_REGEX_SPECIAL)\n> +#define iscntrl(x) (sane_istest(x,GIT_CNTRL))\n> +#define ispunct(x) sane_istest(x, GIT_PUNCT | GIT_REGEX_SPECIAL | \\\n> +\t\tGIT_GLOB_SPECIAL | GIT_PATHSPEC_MAGIC)\n> +#define isxdigit(x) (hexval_table[x] != -1)\n> +#define isprint(x) (sane_istest(x, GIT_ALPHA | GIT_DIGIT | GIT_SPACE | \\\n> +\t\tGIT_PUNCT | GIT_REGEX_SPECIAL | GIT_GLOB_SPECIAL | \\\n> +\t\tGIT_PATHSPEC_MAGIC))\n\n\"Normal\" isprint() only includes space (32) from the white-space characters.\nThe other white-space characters are not considered printable.\n\nDo we want to stay close to the \"original\", or not?\n\nRegards\nJan\n"},{"id":"201423","messageId":"CACsJy8D3WteqsQN_1UxYyE8ADZom6T4Udo8W=hCqiAn+W4K8vQ@mail.gmail.com","threadId":"31810","inReplyTo":"507E9FDE.7080706@cs.tu-berlin.de","subject":"Re: [PATCH v5 02/12] ctype: support iscntrl, ispunct, isxdigit and isprint","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-10-17T12:26:57Z","receivedAt":"2012-10-17T12:26:57Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Oct 17, 2012 at 7:09 PM, \"Jan H. Schönherr\"\n<schnhrr@cs.tu-berlin.de> wrote:\n>>  const unsigned char sane_ctype[256] = {\n>> -     0, 0, 0, 0, 0, 0, 0, 0, 0, S, S, 0, 0, S, 0, 0,         /*   0.. 15 */\n>> -     0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,         /*  16.. 31 */\n>> +     X, X, X, X, X, X, X, X, X, Z, Z, X, X, Z, X, X,         /*   0.. 15 */\n>> +     X, X, X, X, X, X, X, X, X, X, X, X, X, X, X, X,         /*  16.. 31 */\n>\n> \"Normal\" isspace() also includes vertical tab (11) and form-feed (12) as\n> white-space characters. Is there a reason, why they are not included here?\n\nI'm not sure. They were not classified as spaces in the very first\nversion in 4546738 (Unlocalized isspace and friends - 2005-10-13).\nMaybe Linus had a reason to do so.\n\n>> +#define isprint(x) (sane_istest(x, GIT_ALPHA | GIT_DIGIT | GIT_SPACE | \\\n>> +             GIT_PUNCT | GIT_REGEX_SPECIAL | GIT_GLOB_SPECIAL | \\\n>> +             GIT_PATHSPEC_MAGIC))\n>\n> \"Normal\" isprint() only includes space (32) from the white-space characters.\n> The other white-space characters are not considered printable.\n>\n> Do we want to stay close to the \"original\", or not?\n\nWe do. I followed [1] but obvious missed the last sentence in \"print\"\ndescription: \"No characters specified for the keyword cntrl shall be\nspecified\". Thanks for catching. I'll fix it soon.\n\n[1] http://pubs.opengroup.org/onlinepubs/009695399/basedefs/xbd_chap07.html\n-- \nDuy\n"},{"id":"203070","messageId":"1352803572-14547-1-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"507E9FDE.7080706@cs.tu-berlin.de","subject":"[PATCH nd/wildmatch] Correct Git's version of isprint and isspace","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-11-13T10:46:12Z","receivedAt":"2012-11-13T10:46:12Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Git's ispace does not include 11 and 12. Git's isprint includes\ncontrol space characters (10-13). According to glibc-2.14.1 on C\nlocale on Linux, this is wrong. This patch fixes it.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n I wrote a small C program to compare the result of all is* functions\n that Git replaces against the libc version. These are the only ones that\n differ. Which matches what Jan Schönherr commented.\n\n ctype.c           |  6 +++---\n git-compat-util.h | 11 ++++++-----\n 2 files changed, 9 insertions(+), 8 deletions(-)\n\ndiff --git a/ctype.c b/ctype.c\nindex 0bfebb4..71311a3 100644\n--- a/ctype.c\n+++ b/ctype.c\n@@ -14,11 +14,11 @@ enum {\n \tP = GIT_PATHSPEC_MAGIC, /* other non-alnum, except for ] and } */\n \tX = GIT_CNTRL,\n \tU = GIT_PUNCT,\n-\tZ = GIT_CNTRL | GIT_SPACE\n+\tZ = GIT_CNTRL_SPACE\n };\n \n-const unsigned char sane_ctype[256] = {\n-\tX, X, X, X, X, X, X, X, X, Z, Z, X, X, Z, X, X,\t\t/*   0.. 15 */\n+const unsigned int sane_ctype[256] = {\n+\tX, X, X, X, X, X, X, X, X, Z, Z, Z, Z, Z, X, X,\t\t/*   0.. 15 */\n \tX, X, X, X, X, X, X, X, X, X, X, X, X, X, X, X,\t\t/*  16.. 31 */\n \tS, P, P, P, R, P, P, P, R, R, G, R, P, P, R, P,\t\t/*  32.. 47 */\n \tD, D, D, D, D, D, D, D, D, D, P, P, P, P, P, G,\t\t/*  48.. 63 */\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 02f48f6..4ed3f94 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -474,8 +474,8 @@ extern const char tolower_trans_tbl[256];\n #undef ispunct\n #undef isxdigit\n #undef isprint\n-extern const unsigned char sane_ctype[256];\n-#define GIT_SPACE 0x01\n+extern const unsigned int sane_ctype[256];\n+#define GIT_CNTRL_SPACE 0x01\n #define GIT_DIGIT 0x02\n #define GIT_ALPHA 0x04\n #define GIT_GLOB_SPECIAL 0x08\n@@ -483,9 +483,10 @@ extern const unsigned char sane_ctype[256];\n #define GIT_PATHSPEC_MAGIC 0x20\n #define GIT_CNTRL 0x40\n #define GIT_PUNCT 0x80\n-#define sane_istest(x,mask) ((sane_ctype[(unsigned char)(x)] & (mask)) != 0)\n+#define GIT_SPACE 0x100\n+#define sane_istest(x,mask) ((sane_ctype[(unsigned int)(x)] & (mask)) != 0)\n #define isascii(x) (((x) & ~0x7f) == 0)\n-#define isspace(x) sane_istest(x,GIT_SPACE)\n+#define isspace(x) sane_istest(x,GIT_SPACE | GIT_CNTRL_SPACE)\n #define isdigit(x) sane_istest(x,GIT_DIGIT)\n #define isalpha(x) sane_istest(x,GIT_ALPHA)\n #define isalnum(x) sane_istest(x,GIT_ALPHA | GIT_DIGIT)\n@@ -493,7 +494,7 @@ extern const unsigned char sane_ctype[256];\n #define isupper(x) sane_iscase(x, 0)\n #define is_glob_special(x) sane_istest(x,GIT_GLOB_SPECIAL)\n #define is_regex_special(x) sane_istest(x,GIT_GLOB_SPECIAL | GIT_REGEX_SPECIAL)\n-#define iscntrl(x) (sane_istest(x,GIT_CNTRL))\n+#define iscntrl(x) (sane_istest(x,GIT_CNTRL | GIT_CNTRL_SPACE))\n #define ispunct(x) sane_istest(x, GIT_PUNCT | GIT_REGEX_SPECIAL | \\\n \t\tGIT_GLOB_SPECIAL | GIT_PATHSPEC_MAGIC)\n #define isxdigit(x) (hexval_table[x] != -1)\n-- \n1.8.0.rc2.23.g1fb49df\n"},{"id":"203131","messageId":"50A29866.1070700@cs.tu-berlin.de","threadId":"31810","inReplyTo":"1352803572-14547-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH nd/wildmatch] Correct Git's version of isprint and isspace","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-11-13T18:58:46Z","receivedAt":"2012-11-13T18:58:46Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"Hi.\n\nAm 13.11.2012 11:46, schrieb Nguyễn Thái Ngọc Duy:\n> Git's ispace does not include 11 and 12. Git's isprint includes\n> control space characters (10-13). According to glibc-2.14.1 on C\n> locale on Linux, this is wrong. This patch fixes it.\n> \n> Signed-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n> ---\n>  I wrote a small C program to compare the result of all is* functions\n>  that Git replaces against the libc version. These are the only ones that\n>  differ. Which matches what Jan Schönherr commented.\n> \n>  ctype.c           |  6 +++---\n>  git-compat-util.h | 11 ++++++-----\n>  2 files changed, 9 insertions(+), 8 deletions(-)\n> \n> diff --git a/ctype.c b/ctype.c\n> index 0bfebb4..71311a3 100644\n> --- a/ctype.c\n> +++ b/ctype.c\n> @@ -14,11 +14,11 @@ enum {\n>  \tP = GIT_PATHSPEC_MAGIC, /* other non-alnum, except for ] and } */\n>  \tX = GIT_CNTRL,\n>  \tU = GIT_PUNCT,\n> -\tZ = GIT_CNTRL | GIT_SPACE\n> +\tZ = GIT_CNTRL_SPACE\n>  };\n>  \n> -const unsigned char sane_ctype[256] = {\n> -\tX, X, X, X, X, X, X, X, X, Z, Z, X, X, Z, X, X,\t\t/*   0.. 15 */\n> +const unsigned int sane_ctype[256] = {\n> +\tX, X, X, X, X, X, X, X, X, Z, Z, Z, Z, Z, X, X,\t\t/*   0.. 15 */\n>  \tX, X, X, X, X, X, X, X, X, X, X, X, X, X, X, X,\t\t/*  16.. 31 */\n>  \tS, P, P, P, R, P, P, P, R, R, G, R, P, P, R, P,\t\t/*  32.. 47 */\n>  \tD, D, D, D, D, D, D, D, D, D, P, P, P, P, P, G,\t\t/*  48.. 63 */\n\nAn alternative to switching from 1-byte to 4-byte values (don't we have\na 2-byte datatype?), would be to free up GIT_CNTRL and simply do:\n\n#define iscntrl(x) ((x) < 0x20)\n\n\n> diff --git a/git-compat-util.h b/git-compat-util.h\n> index 02f48f6..4ed3f94 100644\n> --- a/git-compat-util.h\n> +++ b/git-compat-util.h\n[...]\n> @@ -483,9 +483,10 @@ extern const unsigned char sane_ctype[256];\n>  #define GIT_PATHSPEC_MAGIC 0x20\n>  #define GIT_CNTRL 0x40\n>  #define GIT_PUNCT 0x80\n> -#define sane_istest(x,mask) ((sane_ctype[(unsigned char)(x)] & (mask)) != 0)\n> +#define GIT_SPACE 0x100\n> +#define sane_istest(x,mask) ((sane_ctype[(unsigned int)(x)] & (mask)) != 0)\n\nThat should better be left \"(unsigned char)\"? We might access values after the\narray otherwise.\n\n(That said, it wasn't really correct before either, when there really is a\npossibility that x >= 0x100.)\n\nRegards\nJan\n\nPS: It looks like my isprint() version was given precedence over your\nisprint() version during the merge into next. That should also be sorted out,\nbut I've no idea which one is actually better: two comparisons versus one\ncache lookup and a bitop... (though my guess is that comparisons are cheaper,\nbut then we should also convert isdigit()...)\n"},{"id":"203132","messageId":"50A29C1F.6000006@lsrfire.ath.cx","threadId":"31810","inReplyTo":"1352803572-14547-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH nd/wildmatch] Correct Git's version of isprint and isspace","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2012-11-13T19:14:39Z","receivedAt":"2012-11-13T19:14:39Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 13.11.2012 11:46, schrieb Nguyễn Thái Ngọc Duy:\n> Git's isprint includes\n> control space characters (10-13). According to glibc-2.14.1 on C\n> locale on Linux, this is wrong. This patch fixes it.\n\nisprint() is not in master, yet.  Can we perhaps still introduce it in \nsuch a way that we never have an incorrect version in master's history?\n\nAnd could you please update test-ctype.c to match the change to \nisspace()?  The tests there just documented the status quo before I made \nchanges to ctype.c long ago, so it's definitions are just as correct (or \nwrong) as the original implementation.\n\nThanks,\nRené\n"},{"id":"203133","messageId":"50A29C3A.1070000@lsrfire.ath.cx","threadId":"31810","inReplyTo":"1352803572-14547-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH nd/wildmatch] Correct Git's version of isprint and isspace","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2012-11-13T19:15:06Z","receivedAt":"2012-11-13T19:15:06Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 13.11.2012 11:46, schrieb Nguyễn Thái Ngọc Duy:\n> Git's ispace does not include 11 and 12.  [...]\n > According to glibc-2.14.1 on C locale on Linux, this is wrong.\n\n11 and 12 being vertical tab (\\v) and form-feed (\\f).  This lack goes \nback to the introduction of git's own character classifier macros seven \nyears ago in 4546738b (Unlocalized isspace and friends).\n\nLinus, do you remember if you left them out on purpose?\n\nThanks,\nRené\n"},{"id":"203136","messageId":"CA+55aFwsjpOop=4mVkx4e=zw5LH41sD9x-b_WMo4Hvo7ygjEtQ@mail.gmail.com","threadId":"31810","inReplyTo":"50A29C3A.1070000@lsrfire.ath.cx","subject":"Re: [PATCH nd/wildmatch] Correct Git's version of isprint and isspace","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2012-11-13T19:40:58Z","receivedAt":"2012-11-13T19:40:58Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Tue, Nov 13, 2012 at 11:15 AM, René Scharfe\n<rene.scharfe@lsrfire.ath.cx> wrote:\n>\n> Linus, do you remember if you left them out on purpose?\n\nUmm, no.\n\nI have to wonder why you care? As far as I'm concerned, the only valid\nspace is space, TAB and CR/LF.\n\nAnything else is *noise*, not space. What's the reason for even caring?\n\n                  Linus\n"},{"id":"203135","messageId":"50A2A254.9030908@kdbg.org","threadId":"31810","inReplyTo":"1352803572-14547-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH nd/wildmatch] Correct Git's version of isprint and isspace","fromName":"Johannes Sixt","fromEmail":"j6t@kdbg.org","sentAt":"2012-11-13T19:41:08Z","receivedAt":"2012-11-13T19:41:08Z","isPatch":true,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Am 13.11.2012 11:46, schrieb Nguyễn Thái Ngọc Duy:\n> @@ -14,11 +14,11 @@ enum {\n>  \tP = GIT_PATHSPEC_MAGIC, /* other non-alnum, except for ] and } */\n>  \tX = GIT_CNTRL,\n>  \tU = GIT_PUNCT,\n> -\tZ = GIT_CNTRL | GIT_SPACE\n> +\tZ = GIT_CNTRL_SPACE\n>  };\n>  \n> -const unsigned char sane_ctype[256] = {\n> -\tX, X, X, X, X, X, X, X, X, Z, Z, X, X, Z, X, X,\t\t/*   0.. 15 */\n> +const unsigned int sane_ctype[256] = {\n> +\tX, X, X, X, X, X, X, X, X, Z, Z, Z, Z, Z, X, X,\t\t/*   0.. 15 */\n>  \tX, X, X, X, X, X, X, X, X, X, X, X, X, X, X, X,\t\t/*  16.. 31 */\n>  \tS, P, P, P, R, P, P, P, R, R, G, R, P, P, R, P,\t\t/*  32.. 47 */\n>  \tD, D, D, D, D, D, D, D, D, D, P, P, P, P, P, G,\t\t/*  48.. 63 */\n> diff --git a/git-compat-util.h b/git-compat-util.h\n> index 02f48f6..4ed3f94 100644\n> --- a/git-compat-util.h\n> +++ b/git-compat-util.h\n> @@ -474,8 +474,8 @@ extern const char tolower_trans_tbl[256];\n>  #undef ispunct\n>  #undef isxdigit\n>  #undef isprint\n> -extern const unsigned char sane_ctype[256];\n> -#define GIT_SPACE 0x01\n> +extern const unsigned int sane_ctype[256];\n> +#define GIT_CNTRL_SPACE 0x01\n>  #define GIT_DIGIT 0x02\n>  #define GIT_ALPHA 0x04\n>  #define GIT_GLOB_SPECIAL 0x08\n> @@ -483,9 +483,10 @@ extern const unsigned char sane_ctype[256];\n>  #define GIT_PATHSPEC_MAGIC 0x20\n>  #define GIT_CNTRL 0x40\n>  #define GIT_PUNCT 0x80\n> -#define sane_istest(x,mask) ((sane_ctype[(unsigned char)(x)] & (mask)) != 0)\n> +#define GIT_SPACE 0x100\n> +#define sane_istest(x,mask) ((sane_ctype[(unsigned int)(x)] & (mask)) != 0)\n>  #define isascii(x) (((x) & ~0x7f) == 0)\n> -#define isspace(x) sane_istest(x,GIT_SPACE)\n> +#define isspace(x) sane_istest(x,GIT_SPACE | GIT_CNTRL_SPACE)\n>  #define isdigit(x) sane_istest(x,GIT_DIGIT)\n>  #define isalpha(x) sane_istest(x,GIT_ALPHA)\n>  #define isalnum(x) sane_istest(x,GIT_ALPHA | GIT_DIGIT)\n> @@ -493,7 +494,7 @@ extern const unsigned char sane_ctype[256];\n>  #define isupper(x) sane_iscase(x, 0)\n>  #define is_glob_special(x) sane_istest(x,GIT_GLOB_SPECIAL)\n>  #define is_regex_special(x) sane_istest(x,GIT_GLOB_SPECIAL | GIT_REGEX_SPECIAL)\n> -#define iscntrl(x) (sane_istest(x,GIT_CNTRL))\n> +#define iscntrl(x) (sane_istest(x,GIT_CNTRL | GIT_CNTRL_SPACE))\n>  #define ispunct(x) sane_istest(x, GIT_PUNCT | GIT_REGEX_SPECIAL | \\\n>  \t\tGIT_GLOB_SPECIAL | GIT_PATHSPEC_MAGIC)\n>  #define isxdigit(x) (hexval_table[x] != -1)\n\nSo we have two properties that overlap:\n\n      SSSSSSSSSS\n   CCCCCCCC\n\nYou seem to generate partions:\n\n   XXXYYYYYZZZZZ\n\nthen assign individual bits to each partition. Now each entry in the\nlookup table has only one bit set. Then you define isxxx() to check for\none of the two possible bits:\n\n   iscntrl is X or Y\n   isspace is Y or Z\n\nBut shouldn't you just assign one bit for S and another one for C, have\nentries in the lookup table with more than one bit set, and check for\nonly one bit in the isxxx macro?\n\nThat way you don't run out of bits as easily as you do with this patch.\n\n-- Hannes\n"},{"id":"203137","messageId":"CA+55aFynRG-CbSp-aLoo1iZTvfBWMgt6kwrPiQjSZJ0ZzraDKQ@mail.gmail.com","threadId":"31810","inReplyTo":"CA+55aFwsjpOop=4mVkx4e=zw5LH41sD9x-b_WMo4Hvo7ygjEtQ@mail.gmail.com","subject":"Re: [PATCH nd/wildmatch] Correct Git's version of isprint and isspace","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2012-11-13T19:50:15Z","receivedAt":"2012-11-13T19:50:15Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Tue, Nov 13, 2012 at 11:40 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> I have to wonder why you care? As far as I'm concerned, the only valid\n> space is space, TAB and CR/LF.\n>\n> Anything else is *noise*, not space. What's the reason for even caring?\n\nBtw, expanding the whitespace selection may actually be very\ncounter-productive. It is used primarily for things like removing\nextraneous space at the end of lines etc, and for that, the current\nselection of SPACE, TAB and LF/CR is the right thing to do.\n\nAdding things like FF etc - that are *technically* whitespace, but\naren't the normal kind of silent whitespace - is potentially going to\nchange things too much. People might *want* a form-feed in their\nmessages, for all we know.\n\nSo I really object to changing things \"just because\". There's a reason\nwe do our own ctype.c: it avoids the crazy crap. It avoids the idiotic\nlocalization issues, and it avoids the ambiguous cases.\n\nSo just let it be, unless you have some major real reason to actually\ncare about a real-world case. And if you do, please explain it. Don't\nchange things just because.\n\n               Linus\n"},{"id":"203213","messageId":"50A3F15C.7000200@lsrfire.ath.cx","threadId":"31810","inReplyTo":"CA+55aFynRG-CbSp-aLoo1iZTvfBWMgt6kwrPiQjSZJ0ZzraDKQ@mail.gmail.com","subject":"Re: [PATCH nd/wildmatch] Correct Git's version of isprint and isspace","fromName":"René Scharfe","fromEmail":"rene.scharfe@lsrfire.ath.cx","sentAt":"2012-11-14T19:30:36Z","receivedAt":"2012-11-14T19:30:36Z","isPatch":true,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 13.11.2012 20:50, schrieb Linus Torvalds:\n> On Tue, Nov 13, 2012 at 11:40 AM, Linus Torvalds\n> <torvalds@linux-foundation.org> wrote:\n>>\n>> I have to wonder why you care? As far as I'm concerned, the only valid\n>> space is space, TAB and CR/LF.\n>>\n>> Anything else is *noise*, not space. What's the reason for even caring?\n>\n> Btw, expanding the whitespace selection may actually be very\n> counter-productive. It is used primarily for things like removing\n> extraneous space at the end of lines etc, and for that, the current\n> selection of SPACE, TAB and LF/CR is the right thing to do.\n>\n> Adding things like FF etc - that are *technically* whitespace, but\n> aren't the normal kind of silent whitespace - is potentially going to\n> change things too much. People might *want* a form-feed in their\n> messages, for all we know.\n\nThe patch was motivated by the integration of the wildmatch library, \nwhich exposes named character classes to users.  It replaces a call of \nfnmatch in match_pathname.  Users probably expect [:space:] to mean the \nsame in git as in other programs.\n\nI never saw a vertical tab and I can't imagine what it's used for.  I'd \nexpect form-feeds to be matched as space, though.  Didn't see them very \noften, admittedly.\n\nNevertheless, it's unfortunate that we have an isspace() that *almost* \ndoes what the widely known thing of the same name does.  I'd shy away \nfrom changing git's version directly, because it's used more than a \nhundred times in the code, and estimating the impact of adding \\v and \\f \nto it.  Perhaps renaming it to isgitspace() is a good first step, \nfollowed by adding a \"standard\" version of isspace() for wildmatch?\n\nRené\n"},{"id":"203286","messageId":"1352981983-22005-1-git-send-email-pclouds@gmail.com","threadId":"31810","inReplyTo":"1352803572-14547-1-git-send-email-pclouds@gmail.com","subject":"[PATCH] wildmatch: correct isprint and isspace","fromName":"Nguyễn Thái Ngọc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-11-15T12:19:43Z","receivedAt":"2012-11-15T12:19:43Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"Current isprint() incorrectly includes control characters 9, 10 and\n13, which is fixed by this patch.\n\nCurrent isspace() lacks 11 and 12. But Git's isspace() has been\ndesigned this way since the beginning and has over 100 call sites\nrelying on this. Instead of updating isspace() behavior (which could be\ntricky as patches from other topics may come in parallel that assume\nthe old isspace()), a new isspace_posix() is introduced and used by\nwildmatch.c. Other part of Git can be converted to use this new\nfunction if it seems appropriate.\n\nSigned-off-by: Nguyễn Thái Ngọc Duy <pclouds@gmail.com>\n---\n Sorry for the late response. I'll reply to everybody in one mail.\n\n On Wed, Nov 14, 2012 at 1:58 AM, \"Jan H. Schönherr\" <schnhrr@cs.tu-berlin.de> wrote:\n > An alternative to switching from 1-byte to 4-byte values (don't we have\n > a 2-byte datatype?), would be to free up GIT_CNTRL and simply do:\n >\n > #define iscntrl(x) ((x) < 0x20)\n \n No. 127 is also a control character.\n\n On Wed, Nov 14, 2012 at 2:41 AM, Johannes Sixt <j6t@kdbg.org> wrote:\n > So we have two properties that overlap:\n >\n >       SSSSSSSSSS\n >    CCCCCCCC\n >\n > You seem to generate partions:\n >\n >    XXXYYYYYZZZZZ\n >\n > then assign individual bits to each partition. Now each entry in the\n > lookup table has only one bit set. Then you define isxxx() to check for\n > one of the two possible bits:\n >\n >    iscntrl is X or Y\n >    isspace is Y or Z\n >\n > But shouldn't you just assign one bit for S and another one for C, have\n > entries in the lookup table with more than one bit set, and check for\n > only one bit in the isxxx macro?\n >\n > That way you don't run out of bits as easily as you do with this patch.\n\n I need three sets of characters actually: control, spaces and\n printable (which contains non-control spaces). Making it\n (isspace(x) && (x) >= 32) is simpler and because isprint() is only used in\n wildmatch, I don't need to think about performance penalty (yet).\n\n On Thu, Nov 15, 2012 at 2:30 AM, René Scharfe <rene.scharfe@lsrfire.ath.cx> wrote:\n > Nevertheless, it's unfortunate that we have an isspace() that *almost* does\n > what the widely known thing of the same name does.  I'd shy away from\n > changing git's version directly, because it's used more than a hundred times\n > in the code, and estimating the impact of adding \\v and \\f to it.\n > Perhaps renaming it to isgitspace() is a good first step, followed by\n > adding a \"standard\" version of isspace() for wildmatch?\n\n There are just too many call sites of isspace() and there is a risk\n of new call sites coming in independently. So I think keeping isspace()\n as-is and using a different name for the standard version is probably\n a better choice.\n\n As the new isspace_posix() is only used by wildmatch, its performance\n as of now is not critical and a simple macro like in this patch is\n probably enough. We can optimize it later if we need to.\n\n git-compat-util.h | 4 +++-\n wildmatch.c       | 8 ++------\n 2 files changed, 5 insertions(+), 7 deletions(-)\n\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex 02f48f6..d4c3fda 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -486,6 +486,7 @@ extern const unsigned char sane_ctype[256];\n #define sane_istest(x,mask) ((sane_ctype[(unsigned char)(x)] & (mask)) != 0)\n #define isascii(x) (((x) & ~0x7f) == 0)\n #define isspace(x) sane_istest(x,GIT_SPACE)\n+#define isspace_posix(x) (((x) >= 9 && (x) <= 13) || (x) == 32)\n #define isdigit(x) sane_istest(x,GIT_DIGIT)\n #define isalpha(x) sane_istest(x,GIT_ALPHA)\n #define isalnum(x) sane_istest(x,GIT_ALPHA | GIT_DIGIT)\n@@ -499,7 +500,8 @@ extern const unsigned char sane_ctype[256];\n #define isxdigit(x) (hexval_table[x] != -1)\n #define isprint(x) (sane_istest(x, GIT_ALPHA | GIT_DIGIT | GIT_SPACE | \\\n \t\tGIT_PUNCT | GIT_REGEX_SPECIAL | GIT_GLOB_SPECIAL | \\\n-\t\tGIT_PATHSPEC_MAGIC))\n+\t\tGIT_PATHSPEC_MAGIC) && \\\n+\t\t(x) >= 32)\n #define tolower(x) sane_case((unsigned char)(x), 0x20)\n #define toupper(x) sane_case((unsigned char)(x), 0)\n #define is_pathspec_magic(x) sane_istest(x,GIT_PATHSPEC_MAGIC)\ndiff --git a/wildmatch.c b/wildmatch.c\nindex 3972e26..fd74efd 100644\n--- a/wildmatch.c\n+++ b/wildmatch.c\n@@ -37,11 +37,7 @@ typedef unsigned char uchar;\n # define ISBLANK(c) ((c) == ' ' || (c) == '\\t')\n #endif\n \n-#ifdef isgraph\n-# define ISGRAPH(c) (ISASCII(c) && isgraph(c))\n-#else\n-# define ISGRAPH(c) (ISASCII(c) && isprint(c) && !isspace(c))\n-#endif\n+#define ISGRAPH(c) (ISASCII(c) && isprint(c) && !isspace_posix(c))\n \n #define ISPRINT(c) (ISASCII(c) && isprint(c))\n #define ISDIGIT(c) (ISASCII(c) && isdigit(c))\n@@ -50,7 +46,7 @@ typedef unsigned char uchar;\n #define ISCNTRL(c) (ISASCII(c) && iscntrl(c))\n #define ISLOWER(c) (ISASCII(c) && islower(c))\n #define ISPUNCT(c) (ISASCII(c) && ispunct(c))\n-#define ISSPACE(c) (ISASCII(c) && isspace(c))\n+#define ISSPACE(c) (ISASCII(c) && isspace_posix(c))\n #define ISUPPER(c) (ISASCII(c) && isupper(c))\n #define ISXDIGIT(c) (ISASCII(c) && isxdigit(c))\n \n-- \n1.8.0.rc2.23.g1fb49df\n"},{"id":"203301","messageId":"50A522B5.7080206@cs.tu-berlin.de","threadId":"31810","inReplyTo":"1352981983-22005-1-git-send-email-pclouds@gmail.com","subject":"Re: [PATCH] wildmatch: correct isprint and isspace","fromName":"Jan H. Schönherr","fromEmail":"schnhrr@cs.tu-berlin.de","sentAt":"2012-11-15T17:13:25Z","receivedAt":"2012-11-15T17:13:25Z","isPatch":true,"sender":{"key":"schnhrr@cs.tu-berlin.de","avatar":null},"body":"Am 15.11.2012 13:19, schrieb Nguyễn Thái Ngọc Duy:\n>  On Thu, Nov 15, 2012 at 2:30 AM, René Scharfe <rene.scharfe@lsrfire.ath.cx> wrote:\n>  > Nevertheless, it's unfortunate that we have an isspace() that *almost* does\n>  > what the widely known thing of the same name does.  I'd shy away from\n>  > changing git's version directly, because it's used more than a hundred times\n>  > in the code, and estimating the impact of adding \\v and \\f to it.\n>  > Perhaps renaming it to isgitspace() is a good first step, followed by\n>  > adding a \"standard\" version of isspace() for wildmatch?\n> \n>  There are just too many call sites of isspace() and there is a risk\n>  of new call sites coming in independently. So I think keeping isspace()\n>  as-is and using a different name for the standard version is probably\n>  a better choice.\n\nAfter having a closer look, where wildmatch is actually used -- matching\nfilenames -- and I've not yet seen \\v or \\f in a filename, it's possibly\nunnecessary to do anything about isspace() right now.\n\n(It's probably more an issue that filenames can be localized, and we only\nsupport unlocalized character classes.)\n\n> diff --git a/git-compat-util.h b/git-compat-util.h\n> index 02f48f6..d4c3fda 100644\n> --- a/git-compat-util.h\n> +++ b/git-compat-util.h\n> @@ -486,6 +486,7 @@ extern const unsigned char sane_ctype[256];\n>  #define sane_istest(x,mask) ((sane_ctype[(unsigned char)(x)] & (mask)) != 0)\n>  #define isascii(x) (((x) & ~0x7f) == 0)\n>  #define isspace(x) sane_istest(x,GIT_SPACE)\n> +#define isspace_posix(x) (((x) >= 9 && (x) <= 13) || (x) == 32)\n>  #define isdigit(x) sane_istest(x,GIT_DIGIT)\n>  #define isalpha(x) sane_istest(x,GIT_ALPHA)\n>  #define isalnum(x) sane_istest(x,GIT_ALPHA | GIT_DIGIT)\n> @@ -499,7 +500,8 @@ extern const unsigned char sane_ctype[256];\n>  #define isxdigit(x) (hexval_table[x] != -1)\n\nThis was from a previous patch, but maybe: \"hexval_table[(unsigned char)x]\"\n\n>  #define isprint(x) (sane_istest(x, GIT_ALPHA | GIT_DIGIT | GIT_SPACE | \\\n>  \t\tGIT_PUNCT | GIT_REGEX_SPECIAL | GIT_GLOB_SPECIAL | \\\n> -\t\tGIT_PATHSPEC_MAGIC))\n> +\t\tGIT_PATHSPEC_MAGIC) && \\\n> +\t\t(x) >= 32)\n\nMay I suggest the current is_print() implementation in master:\n\n#define isprint(x) ((x) >= 0x20 && (x) <= 0x7e)\n\n\nTo summarize my opinion:\n\nI no longer see a reason to correct isspace() (unless somebody with an actual\nuse case complains), and a more POSIXly isprint() is already in master.\n\n=> Nothing to do. :)\n\nRegards\nJan\n"},{"id":"203328","messageId":"CACsJy8AnghB1F8pwHbZpZJKoCZqbQ-TWpV9x_+49ZEJdE0bW2g@mail.gmail.com","threadId":"31810","inReplyTo":"50A522B5.7080206@cs.tu-berlin.de","subject":"Re: [PATCH] wildmatch: correct isprint and isspace","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2012-11-16T04:19:32Z","receivedAt":"2012-11-16T04:19:32Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Fri, Nov 16, 2012 at 12:13 AM, \"Jan H. Schönherr\"\n<schnhrr@cs.tu-berlin.de> wrote:\n>>  #define isprint(x) (sane_istest(x, GIT_ALPHA | GIT_DIGIT | GIT_SPACE | \\\n>>               GIT_PUNCT | GIT_REGEX_SPECIAL | GIT_GLOB_SPECIAL | \\\n>> -             GIT_PATHSPEC_MAGIC))\n>> +             GIT_PATHSPEC_MAGIC) && \\\n>> +             (x) >= 32)\n>\n> May I suggest the current is_print() implementation in master:\n>\n> #define isprint(x) ((x) >= 0x20 && (x) <= 0x7e)\n>\n>\n> To summarize my opinion:\n>\n> I no longer see a reason to correct isspace() (unless somebody with an actual\n> use case complains), and a more POSIXly isprint() is already in master.\n>\n> => Nothing to do. :)\n\n\nYeah. I remember to remind myself to check \"the implementation in\nmaster\" you mentioned but I probably failed at that. Just checked that\nisprint() is already in master, and your comment about isspace() use\nin wildmatch.c makes sense too. So I'm all for doing nothing.\n-- \nDuy\n"}]}