{"thread":{"id":"34345","subject":"[PATCH v2 0/2] commit: improve UTF-8 validation","startedAt":"2013-07-04T17:17:36Z","lastAt":"2013-08-06T07:03:51Z","messageCount":11,"participants":["brian m. carlson","Torsten Bögershausen","Peter Krefting","Junio C Hamano"],"isPatch":true,"patchVersion":2,"patchTotal":2},"messages":[{"id":"222569","messageId":"cover.1372957719.git.sandals@crustytoothpaste.net","threadId":"34345","inReplyTo":null,"subject":"[PATCH v2 0/2] commit: improve UTF-8 validation","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2013-07-04T17:17:36Z","receivedAt":"2013-07-04T17:17:36Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"This series contains a pair of patches that improve the validation of\nthe UTF-8 used in commit messages.  Invalid codepoints, such as\nsurrogates and guaranteed non-characters, are rejected, along with\noverlong UTF-8 sequences.\n\nChanges from v1:\n\n* Improved comments to aid those less familiar with Unicode.\n* Generated test files using printf as part of the test.\n* Removed FIXME comments for things that have been fixed.\n* Use a shorter form for detecting surrogate pairs.\n\nbrian m. carlson (2):\n  commit: reject invalid UTF-8 codepoints\n  commit: reject overlong UTF-8 sequences\n\n commit.c               | 34 ++++++++++++++++++++++++++++------\n t/t3900-i18n-commit.sh | 23 +++++++++++++++++++++++\n 2 files changed, 51 insertions(+), 6 deletions(-)\n\n-- \n1.8.3.1\n"},{"id":"222571","messageId":"20130704171943.GA267700@vauxhall.crustytoothpaste.net","threadId":"34345","inReplyTo":"cover.1372957719.git.sandals@crustytoothpaste.net","subject":"[PATCH v2 1/2] commit: reject invalid UTF-8 codepoints","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2013-07-04T17:19:43Z","receivedAt":"2013-07-04T17:19:43Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"The commit code already contains code for validating UTF-8, but it does not\ncheck for invalid values, such as guaranteed non-characters and surrogates.  Fix\nthis by explicitly checking for and rejecting such characters.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n commit.c               | 27 ++++++++++++++++++++++-----\n t/t3900-i18n-commit.sh | 12 ++++++++++++\n 2 files changed, 34 insertions(+), 5 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex 888e02a..2264106 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1244,6 +1244,7 @@ static int find_invalid_utf8(const char *buf, int len)\n \twhile (len) {\n \t\tunsigned char c = *buf++;\n \t\tint bytes, bad_offset;\n+\t\tunsigned int codepoint;\n \n \t\tlen--;\n \t\toffset++;\n@@ -1264,24 +1265,40 @@ static int find_invalid_utf8(const char *buf, int len)\n \t\t\tbytes++;\n \t\t}\n \n-\t\t/* Must be between 1 and 5 more bytes */\n-\t\tif (bytes < 1 || bytes > 5)\n+\t\t/*\n+\t\t * Must be between 1 and 3 more bytes.  Longer sequences result in\n+\t\t * codepoints beyond U+10FFFF, which are guaranteed never to exist.\n+\t\t */\n+\t\tif (bytes < 1 || bytes > 3)\n \t\t\treturn bad_offset;\n \n \t\t/* Do we *have* that many bytes? */\n \t\tif (len < bytes)\n \t\t\treturn bad_offset;\n \n+\t\t/* Place the encoded bits at the bottom of the value. */\n+\t\tcodepoint = (c & 0x7f) >> bytes;\n+\n \t\toffset += bytes;\n \t\tlen -= bytes;\n \n \t\t/* And verify that they are good continuation bytes */\n \t\tdo {\n+\t\t\tcodepoint <<= 6;\n+\t\t\tcodepoint |= *buf & 0x3f;\n \t\t\tif ((*buf++ & 0xc0) != 0x80)\n \t\t\t\treturn bad_offset;\n \t\t} while (--bytes);\n \n-\t\t/* We could/should check the value and length here too */\n+\t\t/* No codepoints can ever be allocated beyond U+10FFFF. */\n+\t\tif (codepoint > 0x10ffff)\n+\t\t\treturn bad_offset;\n+\t\t/* Surrogates are only for UTF-16 and cannot be encoded in UTF-8. */\n+\t\tif ((codepoint & 0x1ff800) == 0xd800)\n+\t\t\treturn bad_offset;\n+\t\t/* U+FFFE and U+FFFF are guaranteed non-characters. */\n+\t\tif ((codepoint & 0x1ffffe) == 0xfffe)\n+\t\t\treturn bad_offset;\n \t}\n \treturn -1;\n }\n@@ -1292,8 +1309,8 @@ static int find_invalid_utf8(const char *buf, int len)\n  * If it isn't, it assumes any non-utf8 characters are Latin1,\n  * and does the conversion.\n  *\n- * Fixme: we should probably also disallow overlong forms and\n- * invalid characters. But we don't do that currently.\n+ * Fixme: we should probably also disallow overlong forms.\n+ * But we don't do that currently.\n  */\n static int verify_utf8(struct strbuf *buf)\n {\ndiff --git a/t/t3900-i18n-commit.sh b/t/t3900-i18n-commit.sh\nindex 37ddabb..ee8ba6c 100755\n--- a/t/t3900-i18n-commit.sh\n+++ b/t/t3900-i18n-commit.sh\n@@ -39,6 +39,18 @@ test_expect_failure 'UTF-16 refused because of NULs' '\n \tgit commit -a -F \"$TEST_DIRECTORY\"/t3900/UTF-16.txt\n '\n \n+test_expect_success 'UTF-8 invalid characters refused' '\n+\ttest_when_finished \"rm -f $HOME/stderr $HOME/invalid\" && \n+\trm -f \"$HOME/stderr\" &&\n+\techo \"UTF-8 characters\" >F &&\n+\tprintf \"Commit message\\n\\nInvalid surrogate:\\355\\240\\200\\n\" \\\n+\t\t>\"$HOME/invalid\" &&\n+\tgit commit -a -F \"$HOME/invalid\" \\\n+\t\t2>\"$HOME\"/stderr &&\n+\tgrep \"did not conform\" \"$HOME\"/stderr\n+'\n+\n+rm -f \"$HOME/stderr\"\n \n for H in ISO8859-1 eucJP ISO-2022-JP\n do\n-- \n1.8.3.1\n"},{"id":"222572","messageId":"20130704172034.GB267700@vauxhall.crustytoothpaste.net","threadId":"34345","inReplyTo":"cover.1372957719.git.sandals@crustytoothpaste.net","subject":"[PATCH v2 2/2] commit: reject overlong UTF-8 sequences","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2013-07-04T17:20:34Z","receivedAt":"2013-07-04T17:20:34Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"The commit code accepts pseudo-UTF-8 sequences that encode a character with more\nbytes than necessary.  Reject such sequences, since they are not valid UTF-8.\n\nSigned-off-by: brian m. carlson <sandals@crustytoothpaste.net>\n---\n commit.c               | 17 +++++++++++------\n t/t3900-i18n-commit.sh | 11 +++++++++++\n 2 files changed, 22 insertions(+), 6 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex 2264106..b59c187 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1240,11 +1240,15 @@ int commit_tree(const struct strbuf *msg, unsigned char *tree,\n static int find_invalid_utf8(const char *buf, int len)\n {\n \tint offset = 0;\n+\tstatic const unsigned int max_codepoint[] = {\n+\t\t0x7f, 0x7ff, 0xffff, 0x10ffff\n+\t};\n \n \twhile (len) {\n \t\tunsigned char c = *buf++;\n \t\tint bytes, bad_offset;\n \t\tunsigned int codepoint;\n+\t\tunsigned int min_val, max_val;\n \n \t\tlen--;\n \t\toffset++;\n@@ -1276,8 +1280,12 @@ static int find_invalid_utf8(const char *buf, int len)\n \t\tif (len < bytes)\n \t\t\treturn bad_offset;\n \n-\t\t/* Place the encoded bits at the bottom of the value. */\n+\t\t/* Place the encoded bits at the bottom of the value and compute the\n+\t\t * valid range.\n+\t\t */\n \t\tcodepoint = (c & 0x7f) >> bytes;\n+\t\tmin_val = max_codepoint[bytes-1] + 1;\n+\t\tmax_val = max_codepoint[bytes];\n \n \t\toffset += bytes;\n \t\tlen -= bytes;\n@@ -1290,8 +1298,8 @@ static int find_invalid_utf8(const char *buf, int len)\n \t\t\t\treturn bad_offset;\n \t\t} while (--bytes);\n \n-\t\t/* No codepoints can ever be allocated beyond U+10FFFF. */\n-\t\tif (codepoint > 0x10ffff)\n+\t\t/* Reject codepoints that are out of range for the sequence length. */\n+\t\tif (codepoint < min_val || codepoint > max_val)\n \t\t\treturn bad_offset;\n \t\t/* Surrogates are only for UTF-16 and cannot be encoded in UTF-8. */\n \t\tif ((codepoint & 0x1ff800) == 0xd800)\n@@ -1308,9 +1316,6 @@ static int find_invalid_utf8(const char *buf, int len)\n  *\n  * If it isn't, it assumes any non-utf8 characters are Latin1,\n  * and does the conversion.\n- *\n- * Fixme: we should probably also disallow overlong forms.\n- * But we don't do that currently.\n  */\n static int verify_utf8(struct strbuf *buf)\n {\ndiff --git a/t/t3900-i18n-commit.sh b/t/t3900-i18n-commit.sh\nindex ee8ba6c..94fa1e8 100755\n--- a/t/t3900-i18n-commit.sh\n+++ b/t/t3900-i18n-commit.sh\n@@ -50,6 +50,17 @@ test_expect_success 'UTF-8 invalid characters refused' '\n \tgrep \"did not conform\" \"$HOME\"/stderr\n '\n \n+test_expect_success 'UTF-8 overlong sequences rejected' '\n+\ttest_when_finished \"rm -f $HOME/stderr $HOME/invalid\" &&\n+\trm -f \"$HOME/stderr\" \"$HOME/invalid\" &&\n+\techo \"UTF-8 overlong\" >F &&\n+\tprintf \"\\340\\202\\251ommit message\\n\\nThis is not a space:\\300\\240\\n\" \\\n+\t\t>\"$HOME/invalid\" &&\n+\tgit commit -a -F \"$HOME/invalid\" \\\n+\t\t2>\"$HOME\"/stderr &&\n+\tgrep \"did not conform\" \"$HOME\"/stderr\n+'\n+\n rm -f \"$HOME/stderr\"\n \n for H in ISO8859-1 eucJP ISO-2022-JP\n-- \n1.8.3.1\n"},{"id":"222582","messageId":"51D5D3D0.3030102@web.de","threadId":"34345","inReplyTo":"20130704171943.GA267700@vauxhall.crustytoothpaste.net","subject":"Re: [PATCH v2 1/2] commit: reject invalid UTF-8 codepoints","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2013-07-04T19:58:08Z","receivedAt":"2013-07-04T19:58:08Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 2013-07-04 19.19, brian m. carlson wrote:\n> The commit code already contains code for validating UTF-8, but it does not\n> check for invalid values, such as guaranteed non-characters and surrogates.  Fix\ns/guaranteed non-characters/code points out of range/\n> this by explicitly checking for and rejecting such characters.\nDo we really reject them, or do we (only) warn about them ? \n\nOther question:\nNow that we have a check for codepoints out of range, beyond U+10FFFF,\ndo we want to have an additional testcase ?\n\n\n> +test_expect_success 'UTF-8 invalid characters refused' '\nMay be:\n test_expect_success 'UTF-8 invalid surrogate' '\n\n\n> +\ttest_when_finished \"rm -f $HOME/stderr $HOME/invalid\" && \n> +\trm -f \"$HOME/stderr\" &&\n> +\techo \"UTF-8 characters\" >F &&\n> +\tprintf \"Commit message\\n\\nInvalid surrogate:\\355\\240\\200\\n\" \\\n> +\t\t>\"$HOME/invalid\" &&\ngood\n> +\tgit commit -a -F \"$HOME/invalid\" \\\n> +\t\t2>\"$HOME\"/stderr &&\n> +\tgrep \"did not conform\" \"$HOME\"/stderr\n> +'\n> +\n> +rm -f \"$HOME/stderr\"\nDoes it make sense to \"grep on the fly\", like this:\ngit commit -a -F \"$HOME/invalid\" 2>&1  | grep \"did not conform\"\n"},{"id":"222590","messageId":"20130704203921.GR862789@vauxhall.crustytoothpaste.net","threadId":"34345","inReplyTo":"51D5D3D0.3030102@web.de","subject":"Re: [PATCH v2 1/2] commit: reject invalid UTF-8 codepoints","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2013-07-04T20:39:21Z","receivedAt":"2013-07-04T20:39:21Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Thu, Jul 04, 2013 at 09:58:08PM +0200, Torsten Bögershausen wrote:\n> On 2013-07-04 19.19, brian m. carlson wrote:\n> > The commit code already contains code for validating UTF-8, but it does not\n> > check for invalid values, such as guaranteed non-characters and surrogates.  Fix\n> s/guaranteed non-characters/code points out of range/\n\nThe \"such as\" is meant to be illustrative, not all-inclusive, and my\npatch does check for U+FFFE and U+FFFF, which are guaranteed\nnon-characters.\n\n> > this by explicitly checking for and rejecting such characters.\n> Do we really reject them, or do we (only) warn about them ? \n\nWell, find_invalid_utf8 rejects them as invalid, and verify_utf8 fixes\nthem up as if they were Latin-1, and commit_tree_extended warns about\nthem.  My interpretation was from the point of view of the code that I\ntouched (find_invalid_utf8), not the binary.  It would be nice if the\nbinary actually rejected it, too, but that isn't within the scope of\nthis patch.\n\n> Other question:\n> Now that we have a check for codepoints out of range, beyond U+10FFFF,\n> do we want to have an additional testcase ?\n\nSure, why not?\n\n> > +test_expect_success 'UTF-8 invalid characters refused' '\n> May be:\n>  test_expect_success 'UTF-8 invalid surrogate' '\n\nSince I'll be adding at least one more unit test, as you requested, I'll\nchange the name.  I suppose I might as well add a test for the\nnon-characters as well.\n\n> Does it make sense to \"grep on the fly\", like this:\n> git commit -a -F \"$HOME/invalid\" 2>&1  | grep \"did not conform\"\n\nI am interested in making sure that git commit succeeds, and using a\npipe will cause any failure of git commit to be ignored.\n\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | http://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: RSA v4 4096b: 88AC E9B2 9196 305B A994 7552 F1BA 225C 0223 B187\n"},{"id":"222635","messageId":"alpine.DEB.2.00.1307051345260.11814@ds9.cixit.se","threadId":"34345","inReplyTo":"20130704171943.GA267700@vauxhall.crustytoothpaste.net","subject":"Re: [PATCH v2 1/2] commit: reject invalid UTF-8 codepoints","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2013-07-05T12:51:03Z","receivedAt":"2013-07-05T12:51:03Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"brian m. carlson:\n\n> +\t\t/* U+FFFE and U+FFFF are guaranteed non-characters. */\n> +\t\tif ((codepoint & 0x1ffffe) == 0xfffe)\n> +\t\t\treturn bad_offset;\n\nI missed this the first time around: All Unicode characters whose \nlower 16-bits are FFFE or FFFF are non-characters, so you can re-write \nthat to:\n\n   /* U+xxFFFE and U+xxFFFF are guaranteed non-characters. */\n   if ((codepoint & 0xfffe) == 0xfffe)\n    return bad_offset;\n\nAlso, the range U+FDD0--U+FDEF are also non-characters, if you wish to \nbe really pedantic.\n\n$ grep '^[0-9A-F].*<not a' NamesList.txt\nFDD0\t<not a character>\nFDD1\t<not a character>\nFDD2\t<not a character>\nFDD3\t<not a character>\nFDD4\t<not a character>\nFDD5\t<not a character>\nFDD6\t<not a character>\nFDD7\t<not a character>\nFDD8\t<not a character>\nFDD9\t<not a character>\nFDDA\t<not a character>\nFDDB\t<not a character>\nFDDC\t<not a character>\nFDDD\t<not a character>\nFDDE\t<not a character>\nFDDF\t<not a character>\nFDE0\t<not a character>\nFDE1\t<not a character>\nFDE2\t<not a character>\nFDE3\t<not a character>\nFDE4\t<not a character>\nFDE5\t<not a character>\nFDE6\t<not a character>\nFDE7\t<not a character>\nFDE8\t<not a character>\nFDE9\t<not a character>\nFDEA\t<not a character>\nFDEB\t<not a character>\nFDEC\t<not a character>\nFDED\t<not a character>\nFDEE\t<not a character>\nFDEF\t<not a character>\nFFFE\t<not a character>\nFFFF\t<not a character>\n1FFFE\t<not a character>\n1FFFF\t<not a character>\n2FFFE\t<not a character>\n2FFFF\t<not a character>\n3FFFE\t<not a character>\n3FFFF\t<not a character>\n4FFFE\t<not a character>\n4FFFF\t<not a character>\n5FFFE\t<not a character>\n5FFFF\t<not a character>\n6FFFE\t<not a character>\n6FFFF\t<not a character>\n7FFFE\t<not a character>\n7FFFF\t<not a character>\n8FFFE\t<not a character>\n8FFFF\t<not a character>\n9FFFE\t<not a character>\n9FFFF\t<not a character>\nAFFFE\t<not a character>\nAFFFF\t<not a character>\nBFFFE\t<not a character>\nBFFFF\t<not a character>\nCFFFE\t<not a character>\nCFFFF\t<not a character>\nDFFFE\t<not a character>\nDFFFF\t<not a character>\nEFFFE\t<not a character>\nEFFFF\t<not a character>\nFFFFE\t<not a character>\nFFFFF\t<not a character>\n10FFFE\t<not a character>\n10FFFF\t<not a character>\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"222857","messageId":"7vfvvozvx4.fsf@alter.siamese.dyndns.org","threadId":"34345","inReplyTo":"alpine.DEB.2.00.1307051345260.11814@ds9.cixit.se","subject":"Re: [PATCH v2 1/2] commit: reject invalid UTF-8 codepoints","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-07-08T19:36:07Z","receivedAt":"2013-07-08T19:36:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Peter Krefting <peter@softwolves.pp.se> writes:\n\n> brian m. carlson:\n>\n>> +\t\t/* U+FFFE and U+FFFF are guaranteed non-characters. */\n>> +\t\tif ((codepoint & 0x1ffffe) == 0xfffe)\n>> +\t\t\treturn bad_offset;\n>\n> I missed this the first time around: All Unicode characters whose\n> lower 16-bits are FFFE or FFFF are non-characters, so you can re-write\n> that to:\n>\n>   /* U+xxFFFE and U+xxFFFF are guaranteed non-characters. */\n>   if ((codepoint & 0xfffe) == 0xfffe)\n>    return bad_offset;\n>\n> Also, the range U+FDD0--U+FDEF are also non-characters, if you wish to\n> be really pedantic.\n\nYeah, while we are at it, doing this may not hurt.  I think Brian's\ntwo patches are in fairly good shape otherwise, so perhaps you can\ndo this as a follow-up patch on top of the tip of the topic,\ne82bd6cc (commit: reject overlong UTF-8 sequences, 2013-07-04)?\n"},{"id":"222915","messageId":"alpine.DEB.2.00.1307091213090.2313@ds9.cixit.se","threadId":"34345","inReplyTo":"7vfvvozvx4.fsf@alter.siamese.dyndns.org","subject":"[PATCH] commit: reject non-characters","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2013-07-09T11:16:33Z","receivedAt":"2013-07-09T11:16:33Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Unicode clause D14 defines all characters U+nFFFE and U+nFFFF (where\n0 <= n <= 10h) as well as the range U+FDD0..U+FDEF as non-characters,\nreserved for internal use only.  Disallow these characters in commit\nmessages as they are normally not recommended for interchange.\n\nSigned-off-by: Peter Krefting <peter@softwolves.pp.se>\n---\nJunio C Hamano:\n\n> Yeah, while we are at it, doing this may not hurt.  I think Brian's\n> two patches are in fairly good shape otherwise, so perhaps you can\n> do this as a follow-up patch on top of the tip of the topic,\n> e82bd6cc (commit: reject overlong UTF-8 sequences, 2013-07-04)?\n\nOK, here you are. Enjoy :)\n\n  commit.c               |  7 +++++--\n  t/t3900-i18n-commit.sh | 18 ++++++++++++++++++\n  2 files changed, 23 insertions(+), 2 deletions(-)\n\ndiff --git a/commit.c b/commit.c\nindex 5097dba..0587732 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1305,8 +1305,11 @@ static int find_invalid_utf8(const char *buf, int len)\n  \t\t/* Surrogates are only for UTF-16 and cannot be encoded in UTF-8. */\n  \t\tif ((codepoint & 0x1ff800) == 0xd800)\n  \t\t\treturn bad_offset;\n-\t\t/* U+FFFE and U+FFFF are guaranteed non-characters. */\n-\t\tif ((codepoint & 0x1ffffe) == 0xfffe)\n+\t\t/* U+xxFFFE and U+xxFFFF are guaranteed non-characters. */\n+\t\tif ((codepoint & 0xffffe) == 0xfffe)\n+\t\t\treturn bad_offset;\n+\t\t/* So are anything in the range U+FDD0..U+FDEF. */\n+\t\tif (codepoint >= 0xfdd0 && codepoint <= 0xfdef)\n  \t\t\treturn bad_offset;\n  \t}\n  \treturn -1;\ndiff --git a/t/t3900-i18n-commit.sh b/t/t3900-i18n-commit.sh\nindex 051ea9d..38b00c3 100755\n--- a/t/t3900-i18n-commit.sh\n+++ b/t/t3900-i18n-commit.sh\n@@ -58,6 +58,24 @@ test_expect_success 'UTF-8 overlong sequences rejected' '\n  \tgrep \"did not conform\" \"$HOME\"/stderr\n  '\n\n+test_expect_success 'UTF-8 non-characters refused' '\n+\ttest_when_finished \"rm -f $HOME/stderr $HOME/invalid\" &&\n+\techo \"UTF-8 non-character 1\" >F &&\n+\tprintf \"Commit message\\n\\nNon-character:\\364\\217\\277\\276\\n\" \\\n+\t\t>\"$HOME/invalid\" &&\n+\tgit commit -a -F \"$HOME/invalid\" 2>\"$HOME\"/stderr &&\n+\tgrep \"did not conform\" \"$HOME\"/stderr\n+'\n+\n+test_expect_success 'UTF-8 non-characters refused' '\n+\ttest_when_finished \"rm -f $HOME/stderr $HOME/invalid\" &&\n+\techo \"UTF-8 non-character 2.\" >F &&\n+\tprintf \"Commit message\\n\\nNon-character:\\357\\267\\220\\n\" \\\n+\t\t>\"$HOME/invalid\" &&\n+\tgit commit -a -F \"$HOME/invalid\" 2>\"$HOME\"/stderr &&\n+\tgrep \"did not conform\" \"$HOME\"/stderr\n+'\n+\n  for H in ISO8859-1 eucJP ISO-2022-JP\n  do\n  \ttest_expect_success \"$H setup\" '\n-- \n1.8.3.1\n"},{"id":"224575","messageId":"alpine.DEB.2.00.1308051346530.3657@ds9.cixit.se","threadId":"34345","inReplyTo":"alpine.DEB.2.00.1307091213090.2313@ds9.cixit.se","subject":"Re: [PATCH] commit: reject non-characters","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2013-08-05T12:48:28Z","receivedAt":"2013-08-05T12:48:28Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Peter Krefting:\n\n> -\t\t/* U+FFFE and U+FFFF are guaranteed non-characters. */\n> -\t\tif ((codepoint & 0x1ffffe) == 0xfffe)\n> +\t\t/* U+xxFFFE and U+xxFFFF are guaranteed non-characters. */\n> +\t\tif ((codepoint & 0xffffe) == 0xfffe)\n> +\t\t\treturn bad_offset;\n\nDrats, there is an F too many in the bitmask, it should be:\n\n  +\t\tif ((codepoint & 0xfffe) == 0xfffe)\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"224604","messageId":"7vpptskrhd.fsf@alter.siamese.dyndns.org","threadId":"34345","inReplyTo":"alpine.DEB.2.00.1308051346530.3657@ds9.cixit.se","subject":"Re: [PATCH] commit: reject non-characters","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-08-05T16:54:54Z","receivedAt":"2013-08-05T16:54:54Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Peter Krefting <peter@softwolves.pp.se> writes:\n\n> Peter Krefting:\n>\n>> -\t\t/* U+FFFE and U+FFFF are guaranteed non-characters. */\n>> -\t\tif ((codepoint & 0x1ffffe) == 0xfffe)\n>> +\t\t/* U+xxFFFE and U+xxFFFF are guaranteed non-characters. */\n>> +\t\tif ((codepoint & 0xffffe) == 0xfffe)\n>> +\t\t\treturn bad_offset;\n>\n> Drats, there is an F too many in the bitmask, it should be:\n>\n>  +\t\tif ((codepoint & 0xfffe) == 0xfffe)\n\nIndeed.\n\n-- >8 --\nSubject: [PATCH] commit: typofix for xxFFF[EF] check\n\nWe wanted to catch all codepoints that ends with FFFE and FFFF,\nnot with 0FFFE and 0FFFF.\n\nNoticed and corrected by Peter Krefting.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n commit.c | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/commit.c b/commit.c\nindex 7dcfeea..38d8979 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -1306,7 +1306,7 @@ static int find_invalid_utf8(const char *buf, int len)\n \t\tif ((codepoint & 0x1ff800) == 0xd800)\n \t\t\treturn bad_offset;\n \t\t/* U+xxFFFE and U+xxFFFF are guaranteed non-characters. */\n-\t\tif ((codepoint & 0xffffe) == 0xfffe)\n+\t\tif ((codepoint & 0xfffe) == 0xfffe)\n \t\t\treturn bad_offset;\n \t\t/* So are anything in the range U+FDD0..U+FDEF. */\n \t\tif (codepoint >= 0xfdd0 && codepoint <= 0xfdef)\n-- \n1.8.4-rc1-129-g1f3472b\n"},{"id":"224655","messageId":"alpine.DEB.2.00.1308060802570.9867@ds9.cixit.se","threadId":"34345","inReplyTo":"7vpptskrhd.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] commit: reject non-characters","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2013-08-06T07:03:51Z","receivedAt":"2013-08-06T07:03:51Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Junio C Hamano:\n\n> Indeed.\n\nThanks. Testcases are good, but not if they don't actually catch the \nbug one has just introduced :-)\n\n> -- >8 --\n> Subject: [PATCH] commit: typofix for xxFFF[EF] check\n>\n> We wanted to catch all codepoints that ends with FFFE and FFFF,\n> not with 0FFFE and 0FFFF.\n>\n> Noticed and corrected by Peter Krefting.\n>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n\nSigned-off-by: Peter Krefting <peter@softwolves.pp.se>\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"}]}