{"thread":{"id":"41365","subject":"Test failures with GNU grep 2.23","startedAt":"2016-02-07T16:25:40Z","lastAt":"2016-02-24T10:24:12Z","messageCount":27,"participants":["John Keeping","Jeff King","Eric Sunshine","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"277697","messageId":"20160207162540.GK29880@serenity.lan","threadId":"41365","inReplyTo":null,"subject":"Test failures with GNU grep 2.23","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2016-02-07T16:25:40Z","receivedAt":"2016-02-07T16:25:40Z","isPatch":false,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"It seems that binary file detection has changed in GNU grep 2.23 as a\nresult of commit 40ed879 (grep: fix bug with with invalid unibyte\nsequence).\n\nThis causes a couple of test failures in t8005 and t9200 (the t9200 case\nis less obvious so I'm only including t8005 here):\n\n-- >8 --\n$ ./t8005-blame-i18n.sh -v -i\n[snip]\nexpecting success: \n        git blame --incremental file | \\\n                egrep \"^(author|summary) \" > actual &&\n        test_cmp actual expected\n\n--- actual      2016-02-07 16:14:55.372510307 +0000\n+++ expected    2016-02-07 16:14:55.359510341 +0000\n@@ -1 +1,6 @@\n-Binary file (standard input) matches\n+author �R�c ���Y\n+summary �u���[���̃e�X�g�ł��B\n+author �R�c ���Y\n+summary �u���[���̃e�X�g�ł��B\n+author �R�c ���Y\n+summary �u���[���̃e�X�g�ł��B\nnot ok 2 - blame respects i18n.commitencoding\n#\n#               git blame --incremental file | \\\n#                       egrep \"^(author|summary) \" > actual &&\n#               test_cmp actual expected\n#\n-- 8< --\n\nThe following patch fixes the tests for me, but I wonder if \"-a\" is\nsupported on all target platforms (it's not in POSIX, which specifies\nthat the \"input files shall be text files\") or whether we should do\nsomething more comprehensive to provide sane_{e,f,}grep which guarantee\nto treat input as text.\n\nI also tried setting POSIXLY_CORRECT but that doesn't affect the\ntext/binary decision.\n\n-- >8 --\ndiff --git a/t/t8005-blame-i18n.sh b/t/t8005-blame-i18n.sh\nindex 847d098..3b6e697 100755\n--- a/t/t8005-blame-i18n.sh\n+++ b/t/t8005-blame-i18n.sh\n@@ -36,7 +36,7 @@ EOF\n test_expect_success !MINGW \\\n \t'blame respects i18n.commitencoding' '\n \tgit blame --incremental file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\t\tegrep -a \"^(author|summary) \" > actual &&\n \ttest_cmp actual expected\n '\n \n@@ -53,7 +53,7 @@ test_expect_success !MINGW \\\n \t'blame respects i18n.logoutputencoding' '\n \tgit config i18n.logoutputencoding eucJP &&\n \tgit blame --incremental file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\t\tegrep -a \"^(author|summary) \" > actual &&\n \ttest_cmp actual expected\n '\n \n@@ -69,7 +69,7 @@ EOF\n test_expect_success !MINGW \\\n \t'blame respects --encoding=UTF-8' '\n \tgit blame --incremental --encoding=UTF-8 file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\t\tegrep -a \"^(author|summary) \" > actual &&\n \ttest_cmp actual expected\n '\n \n@@ -85,7 +85,7 @@ EOF\n test_expect_success !MINGW \\\n \t'blame respects --encoding=none' '\n \tgit blame --incremental --encoding=none file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\t\tegrep -a \"^(author|summary) \" > actual &&\n \ttest_cmp actual expected\n '\n \ndiff --git a/t/t9200-git-cvsexportcommit.sh b/t/t9200-git-cvsexportcommit.sh\nindex 5cfb9cf..f05578a 100755\n--- a/t/t9200-git-cvsexportcommit.sh\n+++ b/t/t9200-git-cvsexportcommit.sh\n@@ -35,7 +35,7 @@ exit 1\n \n check_entries () {\n \t# $1 == directory, $2 == expected\n-\tgrep '^/' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n+\tgrep -a '^/' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n \tif test -z \"$2\"\n \tthen\n \t\t>expected\n"},{"id":"278662","messageId":"20160219115928.GA10204@sigill.intra.peff.net","threadId":"41365","inReplyTo":"20160207162540.GK29880@serenity.lan","subject":"Re: Test failures with GNU grep 2.23","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-02-19T11:59:29Z","receivedAt":"2016-02-19T11:59:29Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 07, 2016 at 04:25:40PM +0000, John Keeping wrote:\n\n> It seems that binary file detection has changed in GNU grep 2.23 as a\n> result of commit 40ed879 (grep: fix bug with with invalid unibyte\n> sequence).\n\nI read this bug report a while ago when you posted it, but happily\nignored it until today, when my debian unstable system pulled in the new\nversion of grep. :)\n\n> This causes a couple of test failures in t8005 and t9200 (the t9200 case\n> is less obvious so I'm only including t8005 here):\n> \n> -- >8 --\n> $ ./t8005-blame-i18n.sh -v -i\n> [snip]\n> expecting success: \n>         git blame --incremental file | \\\n>                 egrep \"^(author|summary) \" > actual &&\n>         test_cmp actual expected\n\nJust a side note while we are touching these tests:\n\n - we probably should not pipe, so we check the exist code from\n   git-blame\n\n - we usually flip the test_cmp file order, to show the difference from\n   expectation when there is a failure\n\n - no space after \">\" redirection :)\n\n> The following patch fixes the tests for me, but I wonder if \"-a\" is\n> supported on all target platforms (it's not in POSIX, which specifies\n> that the \"input files shall be text files\") or whether we should do\n> something more comprehensive to provide sane_{e,f,}grep which guarantee\n> to treat input as text.\n> \n> I also tried setting POSIXLY_CORRECT but that doesn't affect the\n> text/binary decision.\n\nYeah, I'd worry that \"-a\" is not portable. OTOH, BSD grep seems to have\nit, so between that and GNU, I think most systems are covered. We could\ndo:\n\n  test_lazy_prereq GREP_A '\n\techo foo | grep -a foo\n  '\n\nand mark these tests with it. I'd also be happy to skip that step and\njust do it if and when somebody actually complains about a system\nwithout it (I wouldn't be surprised if most people on antique systems\nend up installing GNU grep anyway).\n\nAnother option might be using \"sed -ne '/^author/p'\" or similar. But\nthat may very well just be trading one portability problem for another.\n\nI also wondered whether we could get away without grepping at all here.\nBut the blame output has a bunch of cruft we don't care about; I think\nthe readability of the tests would suffer if we tried to match the whole\nthing in a test_cmp.\n\n-Peff\n"},{"id":"278673","messageId":"CAPig+cQ0v4edA58=W3YdGUSDc8MeDC3H3Y=-s26ND4=hXj2bpg@mail.gmail.com","threadId":"41365","inReplyTo":"20160219115928.GA10204@sigill.intra.peff.net","subject":"Re: Test failures with GNU grep 2.23","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2016-02-19T17:27:32Z","receivedAt":"2016-02-19T17:27:32Z","isPatch":false,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Fri, Feb 19, 2016 at 6:59 AM, Jeff King <peff@peff.net> wrote:\n> On Sun, Feb 07, 2016 at 04:25:40PM +0000, John Keeping wrote:\n>> The following patch fixes the tests for me, but I wonder if \"-a\" is\n>> supported on all target platforms (it's not in POSIX, which specifies\n>> that the \"input files shall be text files\") or whether we should do\n>> something more comprehensive to provide sane_{e,f,}grep which guarantee\n>> to treat input as text.\n>>\n>> I also tried setting POSIXLY_CORRECT but that doesn't affect the\n>> text/binary decision.\n>\n> Yeah, I'd worry that \"-a\" is not portable. OTOH, BSD grep seems to have\n> it, so between that and GNU, I think most systems are covered.\n\nMac OS X grep seems to support -a and tests in t8005 still pass with\n-a added to the egrep invocations.\n\n> We could\n> do:\n>\n>   test_lazy_prereq GREP_A '\n>         echo foo | grep -a foo\n>   '\n>\n> and mark these tests with it. I'd also be happy to skip that step and\n> just do it if and when somebody actually complains about a system\n> without it (I wouldn't be surprised if most people on antique systems\n> end up installing GNU grep anyway).\n>\n> Another option might be using \"sed -ne '/^author/p'\" or similar. But\n> that may very well just be trading one portability problem for another.\n>\n> I also wondered whether we could get away without grepping at all here.\n> But the blame output has a bunch of cruft we don't care about; I think\n> the readability of the tests would suffer if we tried to match the whole\n> thing in a test_cmp.\n>\n> -Peff\n"},{"id":"278675","messageId":"xmqqmvqwd2ie.fsf@gitster.mtv.corp.google.com","threadId":"41365","inReplyTo":"20160219115928.GA10204@sigill.intra.peff.net","subject":"Re: Test failures with GNU grep 2.23","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-02-19T17:38:17Z","receivedAt":"2016-02-19T17:38:17Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Yeah, I'd worry that \"-a\" is not portable. OTOH, BSD grep seems to have\n> it, so between that and GNU, I think most systems are covered. We could\n> do:\n>\n>   test_lazy_prereq GREP_A '\n> \techo foo | grep -a foo\n>   '\n>\n> and mark these tests with it. I'd also be happy to skip that step and\n> just do it if and when somebody actually complains about a system\n> without it (I wouldn't be surprised if most people on antique systems\n> end up installing GNU grep anyway).\n>\n> Another option might be using \"sed -ne '/^author/p'\" or similar. But\n> that may very well just be trading one portability problem for another.\n\nWould $PERL help, I wonder?\n"},{"id":"278695","messageId":"20160219191125.GB777@sigill.intra.peff.net","threadId":"41365","inReplyTo":"xmqqmvqwd2ie.fsf@gitster.mtv.corp.google.com","subject":"Re: Test failures with GNU grep 2.23","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-02-19T19:11:25Z","receivedAt":"2016-02-19T19:11:25Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 19, 2016 at 09:38:17AM -0800, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > Yeah, I'd worry that \"-a\" is not portable. OTOH, BSD grep seems to have\n> > it, so between that and GNU, I think most systems are covered. We could\n> > do:\n> >\n> >   test_lazy_prereq GREP_A '\n> > \techo foo | grep -a foo\n> >   '\n> >\n> > and mark these tests with it. I'd also be happy to skip that step and\n> > just do it if and when somebody actually complains about a system\n> > without it (I wouldn't be surprised if most people on antique systems\n> > end up installing GNU grep anyway).\n> >\n> > Another option might be using \"sed -ne '/^author/p'\" or similar. But\n> > that may very well just be trading one portability problem for another.\n> \n> Would $PERL help, I wonder?\n\nIt would, though I think you would need to call `binmode` to make it\nreliable. I was hesitant to suggest it, because I seem to recall some\nresistance to more perl dependencies in the test suite, but I think we\nmay be past the point of no return there, anyway.\n\n-Peff\n"},{"id":"278697","messageId":"20160219192311.GB1766@serenity.lan","threadId":"41365","inReplyTo":"xmqqmvqwd2ie.fsf@gitster.mtv.corp.google.com","subject":"Re: Test failures with GNU grep 2.23","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2016-02-19T19:23:11Z","receivedAt":"2016-02-19T19:23:11Z","isPatch":false,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Fri, Feb 19, 2016 at 09:38:17AM -0800, Junio C Hamano wrote:\n> Jeff King <peff@peff.net> writes:\n> \n> > Yeah, I'd worry that \"-a\" is not portable. OTOH, BSD grep seems to have\n> > it, so between that and GNU, I think most systems are covered. We could\n> > do:\n> >\n> >   test_lazy_prereq GREP_A '\n> > \techo foo | grep -a foo\n> >   '\n> >\n> > and mark these tests with it. I'd also be happy to skip that step and\n> > just do it if and when somebody actually complains about a system\n> > without it (I wouldn't be surprised if most people on antique systems\n> > end up installing GNU grep anyway).\n> >\n> > Another option might be using \"sed -ne '/^author/p'\" or similar. But\n> > that may very well just be trading one portability problem for another.\n> \n> Would $PERL help, I wonder?\n\nI suspect that any grep that lacks \"-a\" also lacks binary file handling\nthat will break these tests.  I found a Solaris grep that doesn't\nsupport \"-a\" and it treats these files as text.\n\n>From that perspective, it would be better to have a central place that\ndeals with figuring out how to get grep to work for us.  Perhaps we need\ntest_grep to get this right.  We already have test_cmp_bin() as a thin\nwrapper around cmp so I don't think this is completely unprecedented.\n"},{"id":"278698","messageId":"20160219193310.GA1299@sigill.intra.peff.net","threadId":"41365","inReplyTo":"20160219192311.GB1766@serenity.lan","subject":"Re: Test failures with GNU grep 2.23","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-02-19T19:33:10Z","receivedAt":"2016-02-19T19:33:10Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 19, 2016 at 07:23:11PM +0000, John Keeping wrote:\n\n> I suspect that any grep that lacks \"-a\" also lacks binary file handling\n> that will break these tests.  I found a Solaris grep that doesn't\n> support \"-a\" and it treats these files as text.\n> \n> From that perspective, it would be better to have a central place that\n> deals with figuring out how to get grep to work for us.  Perhaps we need\n> test_grep to get this right.  We already have test_cmp_bin() as a thin\n> wrapper around cmp so I don't think this is completely unprecedented.\n\nI think 99% of the time we are using grep for ascii text. As evidenced\nby the number of test failures we see with the new grep, it is a small\nminority that feed binary gibberish. I'd prefer if \"-a\" handling didn't\nneed to pollute anything outside of this narrow range of tests (and as\nwith my prereq suggestion, I am even find just skipping this narrow\nrange of tests on platforms with no \"-a\", though falling back to running\nwithout \"-a\" is fine if it works).\n\n-Peff\n"},{"id":"278756","messageId":"cover.1456075680.git.john@keeping.me.uk","threadId":"41365","inReplyTo":"20160219193310.GA1299@sigill.intra.peff.net","subject":"[PATCH 0/2] Fix test failures with GNU grep 2.23","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2016-02-21T17:32:20Z","receivedAt":"2016-02-21T17:32:20Z","isPatch":true,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Fri, Feb 19, 2016 at 02:33:10PM -0500, Jeff King wrote:\n> On Fri, Feb 19, 2016 at 07:23:11PM +0000, John Keeping wrote:\n> \n> > I suspect that any grep that lacks \"-a\" also lacks binary file handling\n> > that will break these tests.  I found a Solaris grep that doesn't\n> > support \"-a\" and it treats these files as text.\n> > \n> > From that perspective, it would be better to have a central place that\n> > deals with figuring out how to get grep to work for us.  Perhaps we need\n> > test_grep to get this right.  We already have test_cmp_bin() as a thin\n> > wrapper around cmp so I don't think this is completely unprecedented.\n> \n> I think 99% of the time we are using grep for ascii text. As evidenced\n> by the number of test failures we see with the new grep, it is a small\n> minority that feed binary gibberish. I'd prefer if \"-a\" handling didn't\n> need to pollute anything outside of this narrow range of tests (and as\n> with my prereq suggestion, I am even find just skipping this narrow\n> range of tests on platforms with no \"-a\", though falling back to running\n> without \"-a\" is fine if it works).\n\nI went with using sed in this series because it seems to be the simplest\nand most compatible way to extract lines from the input.  We don't need\nany special casing to figure out if an implementation needs \"-a\" or if\nit doesn't support that option and all the implementation I tested\nsupport the constructs used here.\n\nJohn Keeping (2):\n  t8005: avoid grep on non-ASCII data\n  t9200: avoid grep on non-ASCII data\n\n t/t8005-blame-i18n.sh          | 16 ++++++++--------\n t/t9200-git-cvsexportcommit.sh |  2 +-\n 2 files changed, 9 insertions(+), 9 deletions(-)\n\n-- \n2.7.1.503.g3cfa3ac\n"},{"id":"278757","messageId":"81ec83acd004ef050a4c8df62fb158b41f0a0a80.1456075680.git.john@keeping.me.uk","threadId":"41365","inReplyTo":"cover.1456075680.git.john@keeping.me.uk","subject":"[PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2016-02-21T17:32:21Z","receivedAt":"2016-02-21T17:32:21Z","isPatch":true,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"GNU grep 2.23 detects the input used in this test as binary data so it\ndoes not work for extracting lines from a file.  We could add the \"-a\"\noption to force grep to treat the input as text, but not all\nimplementations support that.  Instead, use sed to extract the desired\nlines since it will always treat its input as text.\n\nWhile touching these lines, modernize the test style to avoid hiding the\nexit status of \"git blame\" and remove a space following a redirection\noperator.\n\nSigned-off-by: John Keeping <john@keeping.me.uk>\n---\n t/t8005-blame-i18n.sh | 16 ++++++++--------\n 1 file changed, 8 insertions(+), 8 deletions(-)\n\ndiff --git a/t/t8005-blame-i18n.sh b/t/t8005-blame-i18n.sh\nindex 847d098..0a86c72 100755\n--- a/t/t8005-blame-i18n.sh\n+++ b/t/t8005-blame-i18n.sh\n@@ -35,8 +35,8 @@ EOF\n \n test_expect_success !MINGW \\\n \t'blame respects i18n.commitencoding' '\n-\tgit blame --incremental file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\tgit blame --incremental file >output &&\n+\tsed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n \ttest_cmp actual expected\n '\n \n@@ -52,8 +52,8 @@ EOF\n test_expect_success !MINGW \\\n \t'blame respects i18n.logoutputencoding' '\n \tgit config i18n.logoutputencoding eucJP &&\n-\tgit blame --incremental file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\tgit blame --incremental file >output &&\n+\tsed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n \ttest_cmp actual expected\n '\n \n@@ -68,8 +68,8 @@ EOF\n \n test_expect_success !MINGW \\\n \t'blame respects --encoding=UTF-8' '\n-\tgit blame --incremental --encoding=UTF-8 file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\tgit blame --incremental --encoding=UTF-8 file >output &&\n+\tsed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n \ttest_cmp actual expected\n '\n \n@@ -84,8 +84,8 @@ EOF\n \n test_expect_success !MINGW \\\n \t'blame respects --encoding=none' '\n-\tgit blame --incremental --encoding=none file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\tgit blame --incremental --encoding=none file >output &&\n+\tsed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n \ttest_cmp actual expected\n '\n \n-- \n2.7.1.503.g3cfa3ac\n"},{"id":"278758","messageId":"42c95c23bffcbb526aaae302f80667867d164876.1456075680.git.john@keeping.me.uk","threadId":"41365","inReplyTo":"cover.1456075680.git.john@keeping.me.uk","subject":"[PATCH 2/2] t9200: avoid grep on non-ASCII data","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2016-02-21T17:32:22Z","receivedAt":"2016-02-21T17:32:22Z","isPatch":true,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"GNU grep 2.23 detects the input used in this test as binary data so it\ndoes not work for extracting lines from a file.  We could add the \"-a\"\noption to force grep to treat the input as text, but not all\nimplementations support that.  Instead, use sed to extract the desired\nlines since it will always treat its input as text.\n\nSigned-off-by: John Keeping <john@keeping.me.uk>\n---\n t/t9200-git-cvsexportcommit.sh | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/t/t9200-git-cvsexportcommit.sh b/t/t9200-git-cvsexportcommit.sh\nindex 812c9cd..0765d52 100755\n--- a/t/t9200-git-cvsexportcommit.sh\n+++ b/t/t9200-git-cvsexportcommit.sh\n@@ -35,7 +35,7 @@ exit 1\n \n check_entries () {\n \t# $1 == directory, $2 == expected\n-\tgrep '^/' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n+\tsed -ne '\\!^/!p' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n \tif test -z \"$2\"\n \tthen\n \t\t>expected\n-- \n2.7.1.503.g3cfa3ac\n"},{"id":"278781","messageId":"CAPig+cQ9n4Eg73Uyeg_g_4wzebuwn8=0R-LMb8F9QLFxanwVVg@mail.gmail.com","threadId":"41365","inReplyTo":"81ec83acd004ef050a4c8df62fb158b41f0a0a80.1456075680.git.john@keeping.me.uk","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2016-02-21T21:01:27Z","receivedAt":"2016-02-21T21:01:27Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Sun, Feb 21, 2016 at 12:32 PM, John Keeping <john@keeping.me.uk> wrote:\n> GNU grep 2.23 detects the input used in this test as binary data so it\n> does not work for extracting lines from a file.  We could add the \"-a\"\n> option to force grep to treat the input as text, but not all\n> implementations support that.  Instead, use sed to extract the desired\n> lines since it will always treat its input as text.\n>\n> While touching these lines, modernize the test style to avoid hiding the\n> exit status of \"git blame\" and remove a space following a redirection\n> operator.\n>\n> Signed-off-by: John Keeping <john@keeping.me.uk>\n> ---\n> diff --git a/t/t8005-blame-i18n.sh b/t/t8005-blame-i18n.sh\n> @@ -35,8 +35,8 @@ EOF\n>  test_expect_success !MINGW \\\n>         'blame respects i18n.commitencoding' '\n> -       git blame --incremental file | \\\n> -               egrep \"^(author|summary) \" > actual &&\n> +       git blame --incremental file >output &&\n> +       sed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n\nThese tests all crash and burn with BSD sed (including Mac OS X) since\nyou're not restricting yourself to BRE (basic regular expressions).\nYou _could_ request extended regular expressions, which do work on\nthose platforms, as well as with GNU sed:\n\n    sed -nEe \"/^(author|summary) /p\" ...\n\n>         test_cmp actual expected\n>  '\n>\n> @@ -52,8 +52,8 @@ EOF\n>  test_expect_success !MINGW \\\n>         'blame respects i18n.logoutputencoding' '\n>         git config i18n.logoutputencoding eucJP &&\n> -       git blame --incremental file | \\\n> -               egrep \"^(author|summary) \" > actual &&\n> +       git blame --incremental file >output &&\n> +       sed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n>         test_cmp actual expected\n>  '\n>\n> @@ -68,8 +68,8 @@ EOF\n>\n>  test_expect_success !MINGW \\\n>         'blame respects --encoding=UTF-8' '\n> -       git blame --incremental --encoding=UTF-8 file | \\\n> -               egrep \"^(author|summary) \" > actual &&\n> +       git blame --incremental --encoding=UTF-8 file >output &&\n> +       sed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n>         test_cmp actual expected\n>  '\n>\n> @@ -84,8 +84,8 @@ EOF\n>\n>  test_expect_success !MINGW \\\n>         'blame respects --encoding=none' '\n> -       git blame --incremental --encoding=none file | \\\n> -               egrep \"^(author|summary) \" > actual &&\n> +       git blame --incremental --encoding=none file >output &&\n> +       sed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n>         test_cmp actual expected\n>  '\n"},{"id":"278782","messageId":"CAPig+cQkcUPD5+0rUPkKCcJSzRC0NkuRYKHmW54eZ041PqaqmQ@mail.gmail.com","threadId":"41365","inReplyTo":"42c95c23bffcbb526aaae302f80667867d164876.1456075680.git.john@keeping.me.uk","subject":"Re: [PATCH 2/2] t9200: avoid grep on non-ASCII data","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2016-02-21T21:15:31Z","receivedAt":"2016-02-21T21:15:31Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Sun, Feb 21, 2016 at 12:32 PM, John Keeping <john@keeping.me.uk> wrote:\n> GNU grep 2.23 detects the input used in this test as binary data so it\n> does not work for extracting lines from a file.  We could add the \"-a\"\n> option to force grep to treat the input as text, but not all\n> implementations support that.  Instead, use sed to extract the desired\n> lines since it will always treat its input as text.\n>\n> Signed-off-by: John Keeping <john@keeping.me.uk>\n> ---\n> diff --git a/t/t9200-git-cvsexportcommit.sh b/t/t9200-git-cvsexportcommit.sh\n> @@ -35,7 +35,7 @@ exit 1\n>  check_entries () {\n>         # $1 == directory, $2 == expected\n> -       grep '^/' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n> +       sed -ne '\\!^/!p' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n\nThis works with BSD sed, but double negatives are confusing. Have you\nconsidered this instead?\n\n    sed -ne '/^\\//p' ...\n\n>         if test -z \"$2\"\n>         then\n>                 >expected\n"},{"id":"278791","messageId":"20160221231913.GA4094@sigill.intra.peff.net","threadId":"41365","inReplyTo":"CAPig+cQ9n4Eg73Uyeg_g_4wzebuwn8=0R-LMb8F9QLFxanwVVg@mail.gmail.com","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-02-21T23:19:14Z","receivedAt":"2016-02-21T23:19:14Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 21, 2016 at 04:01:27PM -0500, Eric Sunshine wrote:\n\n> On Sun, Feb 21, 2016 at 12:32 PM, John Keeping <john@keeping.me.uk> wrote:\n> > GNU grep 2.23 detects the input used in this test as binary data so it\n> > does not work for extracting lines from a file.  We could add the \"-a\"\n> > option to force grep to treat the input as text, but not all\n> > implementations support that.  Instead, use sed to extract the desired\n> > lines since it will always treat its input as text.\n> >\n> > While touching these lines, modernize the test style to avoid hiding the\n> > exit status of \"git blame\" and remove a space following a redirection\n> > operator.\n> >\n> > Signed-off-by: John Keeping <john@keeping.me.uk>\n> > ---\n> > diff --git a/t/t8005-blame-i18n.sh b/t/t8005-blame-i18n.sh\n> > @@ -35,8 +35,8 @@ EOF\n> >  test_expect_success !MINGW \\\n> >         'blame respects i18n.commitencoding' '\n> > -       git blame --incremental file | \\\n> > -               egrep \"^(author|summary) \" > actual &&\n> > +       git blame --incremental file >output &&\n> > +       sed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n> \n> These tests all crash and burn with BSD sed (including Mac OS X) since\n> you're not restricting yourself to BRE (basic regular expressions).\n> You _could_ request extended regular expressions, which do work on\n> those platforms, as well as with GNU sed:\n> \n>     sed -nEe \"/^(author|summary) /p\" ...\n\nAt that point, I think we may as well use grep, because obscure\nplatforms are probably broken either way.\n\nI'm tempted to just go the perl route. We already depend on at least a\nbaisc version of perl5 being installed for many of the other tests, so\nit's not really introducing a new dependency.\n\nSomething like the patch below works for me. I think we could make it\nshorter by using $PERLIO to get the raw behavior, but using binmode will\nwork even on ancient versions of perl.\n\nJohn, if you agree on the direction, feel free to combine it with your\npatch.\n\ndiff --git a/t/t8005-blame-i18n.sh b/t/t8005-blame-i18n.sh\nindex 847d098..f7a02d8 100755\n--- a/t/t8005-blame-i18n.sh\n+++ b/t/t8005-blame-i18n.sh\n@@ -33,10 +33,20 @@ author $SJIS_NAME\n summary $SJIS_MSG\n EOF\n \n+filter_blame () {\n+\tperl -e '\n+\t\tbinmode STDIN;\n+\t\tbinmode STDOUT;\n+\t\twhile (<>) {\n+\t\t\tprint if /^(author|summary) /;\n+\t\t}\n+\t'\n+}\n+\n test_expect_success !MINGW \\\n \t'blame respects i18n.commitencoding' '\n \tgit blame --incremental file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\t\tfilter_blame >actual &&\n \ttest_cmp actual expected\n '\n \n@@ -53,7 +63,7 @@ test_expect_success !MINGW \\\n \t'blame respects i18n.logoutputencoding' '\n \tgit config i18n.logoutputencoding eucJP &&\n \tgit blame --incremental file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\t\tfilter_blame > actual &&\n \ttest_cmp actual expected\n '\n \n@@ -69,7 +79,7 @@ EOF\n test_expect_success !MINGW \\\n \t'blame respects --encoding=UTF-8' '\n \tgit blame --incremental --encoding=UTF-8 file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\t\tfilter_blame >actual &&\n \ttest_cmp actual expected\n '\n \n@@ -85,7 +95,7 @@ EOF\n test_expect_success !MINGW \\\n \t'blame respects --encoding=none' '\n \tgit blame --incremental --encoding=none file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n+\t\tfilter_blame >actual &&\n \ttest_cmp actual expected\n '\n \n"},{"id":"278796","messageId":"CAPig+cSXsk4Pp9adi4KvYjdCwaw4R0Jrv2vwC0JTCyzomWxaww@mail.gmail.com","threadId":"41365","inReplyTo":"20160221231913.GA4094@sigill.intra.peff.net","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2016-02-21T23:31:08Z","receivedAt":"2016-02-21T23:31:08Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Sun, Feb 21, 2016 at 6:19 PM, Jeff King <peff@peff.net> wrote:\n> On Sun, Feb 21, 2016 at 04:01:27PM -0500, Eric Sunshine wrote:\n>> On Sun, Feb 21, 2016 at 12:32 PM, John Keeping <john@keeping.me.uk> wrote:\n>> > -       git blame --incremental file | \\\n>> > -               egrep \"^(author|summary) \" > actual &&\n>> > +       git blame --incremental file >output &&\n>> > +       sed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n>>\n>> These tests all crash and burn with BSD sed (including Mac OS X) since\n>> you're not restricting yourself to BRE (basic regular expressions).\n>> You _could_ request extended regular expressions, which do work on\n>> those platforms, as well as with GNU sed:\n>>\n>>     sed -nEe \"/^(author|summary) /p\" ...\n>\n> At that point, I think we may as well use grep, because obscure\n> platforms are probably broken either way.\n\nI came to the same conclusion but forgot to say so at the end of my message.\n\n> I'm tempted to just go the perl route. We already depend on at least a\n> baisc version of perl5 being installed for many of the other tests, so\n> it's not really introducing a new dependency.\n>\n> Something like the patch below works for me. I think we could make it\n> shorter by using $PERLIO to get the raw behavior, but using binmode will\n> work even on ancient versions of perl.\n>\n> +filter_blame () {\n> +       perl -e '\n> +               binmode STDIN;\n> +               binmode STDOUT;\n\nI was worried about binmode() due to some vague recollection from\nyears and years ago of it being problematic on Windows, but I see\nthese tests are all protected by !MINGW anyhow...\n\n> +               while (<>) {\n> +                       print if /^(author|summary) /;\n> +               }\n> +       '\n> +}\n> +\n>  test_expect_success !MINGW \\\n>         'blame respects i18n.commitencoding' '\n>         git blame --incremental file | \\\n> -               egrep \"^(author|summary) \" > actual &&\n> +               filter_blame >actual &&\n>         test_cmp actual expected\n>  '\n>\n> @@ -53,7 +63,7 @@ test_expect_success !MINGW \\\n>         'blame respects i18n.logoutputencoding' '\n>         git config i18n.logoutputencoding eucJP &&\n>         git blame --incremental file | \\\n> -               egrep \"^(author|summary) \" > actual &&\n> +               filter_blame > actual &&\n>         test_cmp actual expected\n>  '\n>\n> @@ -69,7 +79,7 @@ EOF\n>  test_expect_success !MINGW \\\n>         'blame respects --encoding=UTF-8' '\n>         git blame --incremental --encoding=UTF-8 file | \\\n> -               egrep \"^(author|summary) \" > actual &&\n> +               filter_blame >actual &&\n>         test_cmp actual expected\n>  '\n>\n> @@ -85,7 +95,7 @@ EOF\n>  test_expect_success !MINGW \\\n>         'blame respects --encoding=none' '\n>         git blame --incremental --encoding=none file | \\\n> -               egrep \"^(author|summary) \" > actual &&\n> +               filter_blame >actual &&\n>         test_cmp actual expected\n>  '\n"},{"id":"278797","messageId":"xmqqsi0l8wt8.fsf@gitster.mtv.corp.google.com","threadId":"41365","inReplyTo":"CAPig+cQ9n4Eg73Uyeg_g_4wzebuwn8=0R-LMb8F9QLFxanwVVg@mail.gmail.com","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-02-21T23:31:47Z","receivedAt":"2016-02-21T23:31:47Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Eric Sunshine <sunshine@sunshineco.com> writes:\n\n> These tests all crash and burn with BSD sed (including Mac OS X) since\n> you're not restricting yourself to BRE (basic regular expressions).\n> You _could_ request extended regular expressions, which do work on\n> those platforms, as well as with GNU sed:\n>\n>     sed -nEe \"/^(author|summary) /p\" ...\n\nAn obvious way to avoid any RE is to write it as two separate\nstatements.  As there are repeated invocations of this filtering\nin this script, perhaps a helper function can hide this ugliness?\n"},{"id":"278799","messageId":"20160221233533.GD4094@sigill.intra.peff.net","threadId":"41365","inReplyTo":"CAPig+cSXsk4Pp9adi4KvYjdCwaw4R0Jrv2vwC0JTCyzomWxaww@mail.gmail.com","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-02-21T23:35:33Z","receivedAt":"2016-02-21T23:35:33Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 21, 2016 at 06:31:08PM -0500, Eric Sunshine wrote:\n\n> > Something like the patch below works for me. I think we could make it\n> > shorter by using $PERLIO to get the raw behavior, but using binmode will\n> > work even on ancient versions of perl.\n> >\n> > +filter_blame () {\n> > +       perl -e '\n> > +               binmode STDIN;\n> > +               binmode STDOUT;\n> \n> I was worried about binmode() due to some vague recollection from\n> years and years ago of it being problematic on Windows, but I see\n> these tests are all protected by !MINGW anyhow...\n\nThanks for mentioning that. I meant to put a note on that at the end of\n_my_ message, but forgot. :)\n\nIt does mean we won't do CRLF processing. We could get around that with\nsome explicit `chomp`-ing, I think. Or just leave it as-is and assume\nthese will lose the !MINGW prereq.\n\nI see Junio just mentioned elsewhere that we can simply avoid the\nextended regular expressions by using two sed commands. That would be\nfine with me, too.\n\n-Peff\n"},{"id":"278800","messageId":"CAPig+cQQWAsd9MB4yUKaFePdVoJqrqWEZCFakQxTiWKJWW+4cw@mail.gmail.com","threadId":"41365","inReplyTo":"xmqqsi0l8wt8.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2016-02-21T23:40:04Z","receivedAt":"2016-02-21T23:40:04Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Sun, Feb 21, 2016 at 6:31 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Eric Sunshine <sunshine@sunshineco.com> writes:\n>\n>> These tests all crash and burn with BSD sed (including Mac OS X) since\n>> you're not restricting yourself to BRE (basic regular expressions).\n>> You _could_ request extended regular expressions, which do work on\n>> those platforms, as well as with GNU sed:\n>>\n>>     sed -nEe \"/^(author|summary) /p\" ...\n>\n> An obvious way to avoid any RE is to write it as two separate\n> statements.\n\nYes, that's even better; now I feel stupid for not thinking of it.\n"},{"id":"278801","messageId":"20160221234135.GA14382@river.lan","threadId":"41365","inReplyTo":"20160221231913.GA4094@sigill.intra.peff.net","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2016-02-21T23:41:35Z","receivedAt":"2016-02-21T23:41:35Z","isPatch":true,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Sun, Feb 21, 2016 at 06:19:14PM -0500, Jeff King wrote:\n> On Sun, Feb 21, 2016 at 04:01:27PM -0500, Eric Sunshine wrote:\n> \n> > On Sun, Feb 21, 2016 at 12:32 PM, John Keeping <john@keeping.me.uk> wrote:\n> > > GNU grep 2.23 detects the input used in this test as binary data so it\n> > > does not work for extracting lines from a file.  We could add the \"-a\"\n> > > option to force grep to treat the input as text, but not all\n> > > implementations support that.  Instead, use sed to extract the desired\n> > > lines since it will always treat its input as text.\n> > >\n> > > While touching these lines, modernize the test style to avoid hiding the\n> > > exit status of \"git blame\" and remove a space following a redirection\n> > > operator.\n> > >\n> > > Signed-off-by: John Keeping <john@keeping.me.uk>\n> > > ---\n> > > diff --git a/t/t8005-blame-i18n.sh b/t/t8005-blame-i18n.sh\n> > > @@ -35,8 +35,8 @@ EOF\n> > >  test_expect_success !MINGW \\\n> > >         'blame respects i18n.commitencoding' '\n> > > -       git blame --incremental file | \\\n> > > -               egrep \"^(author|summary) \" > actual &&\n> > > +       git blame --incremental file >output &&\n> > > +       sed -ne \"/^\\(author\\|summary\\) /p\" output >actual &&\n> > \n> > These tests all crash and burn with BSD sed (including Mac OS X) since\n> > you're not restricting yourself to BRE (basic regular expressions).\n> > You _could_ request extended regular expressions, which do work on\n> > those platforms, as well as with GNU sed:\n> > \n> >     sed -nEe \"/^(author|summary) /p\" ...\n> \n> At that point, I think we may as well use grep, because obscure\n> platforms are probably broken either way.\n\nAlso GNU sed doesn't understand \"-E\", it uses \"-r\" for --regexp-extended.\n\n> I'm tempted to just go the perl route. We already depend on at least a\n> baisc version of perl5 being installed for many of the other tests, so\n> it's not really introducing a new dependency.\n> \n> Something like the patch below works for me. I think we could make it\n> shorter by using $PERLIO to get the raw behavior, but using binmode will\n> work even on ancient versions of perl.\n> \n> John, if you agree on the direction, feel free to combine it with your\n> patch.\n\nMy original sed version was:\n\n\tsed -ne \"/^author /p\" -e \"/^summary /p\"\n\nwhich I think will work on all platforms (we already use it in\nt0000-basic.sh) but then I decided to be too clever :-(\n\nI still think sed is simpler than introducing a new function to wrap a\nperl script.\n"},{"id":"278802","messageId":"20160221234345.GB14382@river.lan","threadId":"41365","inReplyTo":"CAPig+cQkcUPD5+0rUPkKCcJSzRC0NkuRYKHmW54eZ041PqaqmQ@mail.gmail.com","subject":"Re: [PATCH 2/2] t9200: avoid grep on non-ASCII data","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2016-02-21T23:43:45Z","receivedAt":"2016-02-21T23:43:45Z","isPatch":true,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Sun, Feb 21, 2016 at 04:15:31PM -0500, Eric Sunshine wrote:\n> On Sun, Feb 21, 2016 at 12:32 PM, John Keeping <john@keeping.me.uk> wrote:\n> > GNU grep 2.23 detects the input used in this test as binary data so it\n> > does not work for extracting lines from a file.  We could add the \"-a\"\n> > option to force grep to treat the input as text, but not all\n> > implementations support that.  Instead, use sed to extract the desired\n> > lines since it will always treat its input as text.\n> >\n> > Signed-off-by: John Keeping <john@keeping.me.uk>\n> > ---\n> > diff --git a/t/t9200-git-cvsexportcommit.sh b/t/t9200-git-cvsexportcommit.sh\n> > @@ -35,7 +35,7 @@ exit 1\n> >  check_entries () {\n> >         # $1 == directory, $2 == expected\n> > -       grep '^/' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n> > +       sed -ne '\\!^/!p' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n> \n> This works with BSD sed, but double negatives are confusing. Have you\n> considered this instead?\n> \n>     sed -ne '/^\\//p' ...\n\nWhat do you mean double negatives?  Do you mean using \"!\" as an\nalternative delimiter?  I find changing delimters is normally simpler\nthan following multiple levels of quoting for escaping slashes, although\nin this case it's simple enough that it doesn't make much difference.\n"},{"id":"278803","messageId":"CAPig+cQvuc+Gn1jsnYr0_wMe+fggtHyJPXx4dV=h5sMGWpGRuQ@mail.gmail.com","threadId":"41365","inReplyTo":"20160221234135.GA14382@river.lan","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2016-02-21T23:50:15Z","receivedAt":"2016-02-21T23:50:15Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Sun, Feb 21, 2016 at 6:41 PM, John Keeping <john@keeping.me.uk> wrote:\n> On Sun, Feb 21, 2016 at 06:19:14PM -0500, Jeff King wrote:\n>> On Sun, Feb 21, 2016 at 04:01:27PM -0500, Eric Sunshine wrote:\n>> > These tests all crash and burn with BSD sed (including Mac OS X) since\n>> > you're not restricting yourself to BRE (basic regular expressions).\n>> > You _could_ request extended regular expressions, which do work on\n>> > those platforms, as well as with GNU sed:\n>> >\n>> >     sed -nEe \"/^(author|summary) /p\" ...\n>>\n>> At that point, I think we may as well use grep, because obscure\n>> platforms are probably broken either way.\n>\n> Also GNU sed doesn't understand \"-E\", it uses \"-r\" for --regexp-extended.\n\nIt actually does recognize -E in all the versions I've tested,\nhowever, apparently it's undocumented (thus probably should be\navoided).\n\n> My original sed version was:\n>\n>         sed -ne \"/^author /p\" -e \"/^summary /p\"\n>\n> which I think will work on all platforms (we already use it in\n> t0000-basic.sh) but then I decided to be too clever :-(\n\nThe unclever version seems fine.\n"},{"id":"278805","messageId":"CAPig+cSQkjjLKAQvd1HR+1p9Gni_V7YXB-MK7qsh4yP1USzyLg@mail.gmail.com","threadId":"41365","inReplyTo":"20160221234345.GB14382@river.lan","subject":"Re: [PATCH 2/2] t9200: avoid grep on non-ASCII data","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2016-02-22T00:04:22Z","receivedAt":"2016-02-22T00:04:22Z","isPatch":true,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Sun, Feb 21, 2016 at 6:43 PM, John Keeping <john@keeping.me.uk> wrote:\n> On Sun, Feb 21, 2016 at 04:15:31PM -0500, Eric Sunshine wrote:\n>> On Sun, Feb 21, 2016 at 12:32 PM, John Keeping <john@keeping.me.uk> wrote:\n>> > GNU grep 2.23 detects the input used in this test as binary data so it\n>> > does not work for extracting lines from a file.  We could add the \"-a\"\n>> > option to force grep to treat the input as text, but not all\n>> > implementations support that.  Instead, use sed to extract the desired\n>> > lines since it will always treat its input as text.\n>> >\n>> > Signed-off-by: John Keeping <john@keeping.me.uk>\n>> > ---\n>> > diff --git a/t/t9200-git-cvsexportcommit.sh b/t/t9200-git-cvsexportcommit.sh\n>> > @@ -35,7 +35,7 @@ exit 1\n>> >  check_entries () {\n>> >         # $1 == directory, $2 == expected\n>> > -       grep '^/' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n>> > +       sed -ne '\\!^/!p' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n>>\n>> This works with BSD sed, but double negatives are confusing. Have you\n>> considered this instead?\n>>\n>>     sed -ne '/^\\//p' ...\n>\n> What do you mean double negatives?  Do you mean using \"!\" as an\n> alternative delimiter?  I find changing delimters is normally simpler\n> than following multiple levels of quoting for escaping slashes, although\n> in this case it's simple enough that it doesn't make much difference.\n\nNice, I learned something new today. If I recall correctly, historic\nsed did not allow the delimiter to be changed (or it wasn't documented\nor I simply forgot about the capability). So, feel free to ignore me.\n"},{"id":"278911","messageId":"20160222221811.GC18522@sigill.intra.peff.net","threadId":"41365","inReplyTo":"20160221234135.GA14382@river.lan","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-02-22T22:18:11Z","receivedAt":"2016-02-22T22:18:11Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 21, 2016 at 11:41:35PM +0000, John Keeping wrote:\n\n> My original sed version was:\n> \n> \tsed -ne \"/^author /p\" -e \"/^summary /p\"\n> \n> which I think will work on all platforms (we already use it in\n> t0000-basic.sh) but then I decided to be too clever :-(\n> \n> I still think sed is simpler than introducing a new function to wrap a\n> perl script.\n\nYeah, I think that is good (personally I'd use a function anyway, but I\nthink it is short enough that we could go either way).\n\n-Peff\n"},{"id":"278914","messageId":"20160222222503.GD18522@sigill.intra.peff.net","threadId":"41365","inReplyTo":"20160221234345.GB14382@river.lan","subject":"Re: [PATCH 2/2] t9200: avoid grep on non-ASCII data","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-02-22T22:25:04Z","receivedAt":"2016-02-22T22:25:04Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 21, 2016 at 11:43:45PM +0000, John Keeping wrote:\n\n> On Sun, Feb 21, 2016 at 04:15:31PM -0500, Eric Sunshine wrote:\n> > On Sun, Feb 21, 2016 at 12:32 PM, John Keeping <john@keeping.me.uk> wrote:\n> > > GNU grep 2.23 detects the input used in this test as binary data so it\n> > > does not work for extracting lines from a file.  We could add the \"-a\"\n> > > option to force grep to treat the input as text, but not all\n> > > implementations support that.  Instead, use sed to extract the desired\n> > > lines since it will always treat its input as text.\n> > >\n> > > Signed-off-by: John Keeping <john@keeping.me.uk>\n> > > ---\n> > > diff --git a/t/t9200-git-cvsexportcommit.sh b/t/t9200-git-cvsexportcommit.sh\n> > > @@ -35,7 +35,7 @@ exit 1\n> > >  check_entries () {\n> > >         # $1 == directory, $2 == expected\n> > > -       grep '^/' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n> > > +       sed -ne '\\!^/!p' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n> > \n> > This works with BSD sed, but double negatives are confusing. Have you\n> > considered this instead?\n> > \n> >     sed -ne '/^\\//p' ...\n> \n> What do you mean double negatives?  Do you mean using \"!\" as an\n> alternative delimiter?  I find changing delimters is normally simpler\n> than following multiple levels of quoting for escaping slashes, although\n> in this case it's simple enough that it doesn't make much difference.\n\nI agree that changing delimiters is much nicer than backslashes. But I\nwonder if using \"!\" is more confusing than it needs to be, given its\nother meanings.\n\nI dunno. I admit that the backslash threw me off, too (since it needs\nescaped in interactive shells, I first assumed that's what was going\non). Using backslash to select the delimiter was new to me. I've usually\nseen:\n\n  s!/foo/!/bar/!\n\nwhich is arguably a little more clear. Too bad we cannot do:\n\n  m!/foo!\n\nwhich I think reads better. Oh well. Maybe:\n\n  sed -ne '\\#^/#p'\n\nwould be more readable, but I'm just bikeshedding at this point.  The\ngrep invocation really was the most clear. :-/\n\n-Peff\n"},{"id":"278915","messageId":"xmqq1t845qmg.fsf@gitster.mtv.corp.google.com","threadId":"41365","inReplyTo":"20160222221811.GC18522@sigill.intra.peff.net","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-02-22T22:25:59Z","receivedAt":"2016-02-22T22:25:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Sun, Feb 21, 2016 at 11:41:35PM +0000, John Keeping wrote:\n>\n>> My original sed version was:\n>> \n>> \tsed -ne \"/^author /p\" -e \"/^summary /p\"\n>> \n>> which I think will work on all platforms (we already use it in\n>> t0000-basic.sh) but then I decided to be too clever :-(\n>> \n>> I still think sed is simpler than introducing a new function to wrap a\n>> perl script.\n>\n> Yeah, I think that is good (personally I'd use a function anyway, but I\n> think it is short enough that we could go either way).\n\nAgreed, and because there are repeated invocation of the same sed\nscript in this file, it would be sensible to hide it behind a helper\nfunction.\n\nThanks.\n"},{"id":"279086","messageId":"xmqqio1fxcj4.fsf@gitster.mtv.corp.google.com","threadId":"41365","inReplyTo":"20160222222503.GD18522@sigill.intra.peff.net","subject":"Re: [PATCH 2/2] t9200: avoid grep on non-ASCII data","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-02-23T22:55:11Z","receivedAt":"2016-02-23T22:55:11Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Sun, Feb 21, 2016 at 11:43:45PM +0000, John Keeping wrote:\n>\n>> On Sun, Feb 21, 2016 at 04:15:31PM -0500, Eric Sunshine wrote:\n>> > On Sun, Feb 21, 2016 at 12:32 PM, John Keeping <john@keeping.me.uk> wrote:\n>> > > GNU grep 2.23 detects the input used in this test as binary data so it\n>> > > does not work for extracting lines from a file.  We could add the \"-a\"\n>> > > option to force grep to treat the input as text, but not all\n>> > > implementations support that.  Instead, use sed to extract the desired\n>> > > lines since it will always treat its input as text.\n>> > >\n>> > > Signed-off-by: John Keeping <john@keeping.me.uk>\n>> > > ---\n>> > > diff --git a/t/t9200-git-cvsexportcommit.sh b/t/t9200-git-cvsexportcommit.sh\n>> > > @@ -35,7 +35,7 @@ exit 1\n>> > >  check_entries () {\n>> > >         # $1 == directory, $2 == expected\n>> > > -       grep '^/' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n>> > > +       sed -ne '\\!^/!p' \"$1/CVS/Entries\" | sort | cut -d/ -f2,3,5 >actual\n>> > \n>> > This works with BSD sed, but double negatives are confusing. Have you\n>> > considered this instead?\n>> > \n>> >     sed -ne '/^\\//p' ...\n>> \n>> What do you mean double negatives?  Do you mean using \"!\" as an\n>> alternative delimiter?  I find changing delimters is normally simpler\n>> than following multiple levels of quoting for escaping slashes, although\n>> in this case it's simple enough that it doesn't make much difference.\n>\n> I agree that changing delimiters is much nicer than backslashes. But I\n> wonder if using \"!\" is more confusing than it needs to be, given its\n> other meanings.\n>\n> I dunno. I admit that the backslash threw me off, too (since it needs\n> escaped in interactive shells, I first assumed that's what was going\n> on). Using backslash to select the delimiter was new to me. I've usually\n> seen:\n>\n>   s!/foo/!/bar/!\n>\n> which is arguably a little more clear. Too bad we cannot do:\n>\n>   m!/foo!\n>\n> which I think reads better. Oh well. Maybe:\n>\n>   sed -ne '\\#^/#p'\n>\n> would be more readable, but I'm just bikeshedding at this point.  The\n> grep invocation really was the most clear. :-/\n\nEric's '/^\\//' was the most straight-forward and easiest to see what\nis going on, I would think.\n"},{"id":"279088","messageId":"xmqqegc3xc88.fsf@gitster.mtv.corp.google.com","threadId":"41365","inReplyTo":"20160221234135.GA14382@river.lan","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-02-23T23:01:43Z","receivedAt":"2016-02-23T23:01:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"John Keeping <john@keeping.me.uk> writes:\n\n> My original sed version was:\n>\n> \tsed -ne \"/^author /p\" -e \"/^summary /p\"\n>\n> which I think will work on all platforms (we already use it in\n> t0000-basic.sh) but then I decided to be too clever :-(\n>\n> I still think sed is simpler than introducing a new function to wrap a\n> perl script.\n\nLet's do this, before everybody forgets what we discussed.\n\n-- >8 --\nFrom: John Keeping <john@keeping.me.uk>\nDate: Sun, 21 Feb 2016 17:32:21 +0000\nSubject: [PATCH] t8005: avoid grep on non-ASCII data\n\nGNU grep 2.23 detects the input used in this test as binary data so it\ndoes not work for extracting lines from a file.  We could add the \"-a\"\noption to force grep to treat the input as text, but not all\nimplementations support that.  Instead, use sed to extract the desired\nlines since it will always treat its input as text.\n\nWhile touching these lines, modernize the test style to avoid hiding the\nexit status of \"git blame\" and remove a space following a redirection\noperator.  Also swap the order of the expected and actual output\nfiles given to test_cmp; we compare expect and actual to show how\nactual output differs from what is expected.\n\nSigned-off-by: John Keeping <john@keeping.me.uk>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n t/t8005-blame-i18n.sh | 28 ++++++++++++++++------------\n 1 file changed, 16 insertions(+), 12 deletions(-)\n\ndiff --git a/t/t8005-blame-i18n.sh b/t/t8005-blame-i18n.sh\nindex 847d098..75da219 100755\n--- a/t/t8005-blame-i18n.sh\n+++ b/t/t8005-blame-i18n.sh\n@@ -33,11 +33,15 @@ author $SJIS_NAME\n summary $SJIS_MSG\n EOF\n \n+filter_author_summary () {\n+\tsed -n -e '/^author /p' -e '/^summary /p' \"$@\"\n+}\n+\n test_expect_success !MINGW \\\n \t'blame respects i18n.commitencoding' '\n-\tgit blame --incremental file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n-\ttest_cmp actual expected\n+\tgit blame --incremental file >output &&\n+\tfilter_author_summary output >actual &&\n+\ttest_cmp expected actual\n '\n \n cat >expected <<EOF\n@@ -52,9 +56,9 @@ EOF\n test_expect_success !MINGW \\\n \t'blame respects i18n.logoutputencoding' '\n \tgit config i18n.logoutputencoding eucJP &&\n-\tgit blame --incremental file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n-\ttest_cmp actual expected\n+\tgit blame --incremental file >output &&\n+\tfilter_author_summary output >actual &&\n+\ttest_cmp expected actual\n '\n \n cat >expected <<EOF\n@@ -68,9 +72,9 @@ EOF\n \n test_expect_success !MINGW \\\n \t'blame respects --encoding=UTF-8' '\n-\tgit blame --incremental --encoding=UTF-8 file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n-\ttest_cmp actual expected\n+\tgit blame --incremental --encoding=UTF-8 file >output &&\n+\tfilter_author_summary output >actual &&\n+\ttest_cmp expected actual\n '\n \n cat >expected <<EOF\n@@ -84,9 +88,9 @@ EOF\n \n test_expect_success !MINGW \\\n \t'blame respects --encoding=none' '\n-\tgit blame --incremental --encoding=none file | \\\n-\t\tegrep \"^(author|summary) \" > actual &&\n-\ttest_cmp actual expected\n+\tgit blame --incremental --encoding=none file >output &&\n+\tfilter_author_summary output >actual &&\n+\ttest_cmp expected actual\n '\n \n test_done\n-- \n2.7.2-532-g79873b4\n"},{"id":"279140","messageId":"20160224102411.GO1766@serenity.lan","threadId":"41365","inReplyTo":"xmqqegc3xc88.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH 1/2] t8005: avoid grep on non-ASCII data","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2016-02-24T10:24:12Z","receivedAt":"2016-02-24T10:24:12Z","isPatch":true,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Tue, Feb 23, 2016 at 03:01:43PM -0800, Junio C Hamano wrote:\n> John Keeping <john@keeping.me.uk> writes:\n> \n> > My original sed version was:\n> >\n> > \tsed -ne \"/^author /p\" -e \"/^summary /p\"\n> >\n> > which I think will work on all platforms (we already use it in\n> > t0000-basic.sh) but then I decided to be too clever :-(\n> >\n> > I still think sed is simpler than introducing a new function to wrap a\n> > perl script.\n> \n> Let's do this, before everybody forgets what we discussed.\n\nThanks, this looks good to me.\n\n> -- >8 --\n> From: John Keeping <john@keeping.me.uk>\n> Date: Sun, 21 Feb 2016 17:32:21 +0000\n> Subject: [PATCH] t8005: avoid grep on non-ASCII data\n> \n> GNU grep 2.23 detects the input used in this test as binary data so it\n> does not work for extracting lines from a file.  We could add the \"-a\"\n> option to force grep to treat the input as text, but not all\n> implementations support that.  Instead, use sed to extract the desired\n> lines since it will always treat its input as text.\n> \n> While touching these lines, modernize the test style to avoid hiding the\n> exit status of \"git blame\" and remove a space following a redirection\n> operator.  Also swap the order of the expected and actual output\n> files given to test_cmp; we compare expect and actual to show how\n> actual output differs from what is expected.\n> \n> Signed-off-by: John Keeping <john@keeping.me.uk>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>  t/t8005-blame-i18n.sh | 28 ++++++++++++++++------------\n>  1 file changed, 16 insertions(+), 12 deletions(-)\n> \n> diff --git a/t/t8005-blame-i18n.sh b/t/t8005-blame-i18n.sh\n> index 847d098..75da219 100755\n> --- a/t/t8005-blame-i18n.sh\n> +++ b/t/t8005-blame-i18n.sh\n> @@ -33,11 +33,15 @@ author $SJIS_NAME\n>  summary $SJIS_MSG\n>  EOF\n>  \n> +filter_author_summary () {\n> +\tsed -n -e '/^author /p' -e '/^summary /p' \"$@\"\n> +}\n> +\n>  test_expect_success !MINGW \\\n>  \t'blame respects i18n.commitencoding' '\n> -\tgit blame --incremental file | \\\n> -\t\tegrep \"^(author|summary) \" > actual &&\n> -\ttest_cmp actual expected\n> +\tgit blame --incremental file >output &&\n> +\tfilter_author_summary output >actual &&\n> +\ttest_cmp expected actual\n>  '\n>  \n>  cat >expected <<EOF\n> @@ -52,9 +56,9 @@ EOF\n>  test_expect_success !MINGW \\\n>  \t'blame respects i18n.logoutputencoding' '\n>  \tgit config i18n.logoutputencoding eucJP &&\n> -\tgit blame --incremental file | \\\n> -\t\tegrep \"^(author|summary) \" > actual &&\n> -\ttest_cmp actual expected\n> +\tgit blame --incremental file >output &&\n> +\tfilter_author_summary output >actual &&\n> +\ttest_cmp expected actual\n>  '\n>  \n>  cat >expected <<EOF\n> @@ -68,9 +72,9 @@ EOF\n>  \n>  test_expect_success !MINGW \\\n>  \t'blame respects --encoding=UTF-8' '\n> -\tgit blame --incremental --encoding=UTF-8 file | \\\n> -\t\tegrep \"^(author|summary) \" > actual &&\n> -\ttest_cmp actual expected\n> +\tgit blame --incremental --encoding=UTF-8 file >output &&\n> +\tfilter_author_summary output >actual &&\n> +\ttest_cmp expected actual\n>  '\n>  \n>  cat >expected <<EOF\n> @@ -84,9 +88,9 @@ EOF\n>  \n>  test_expect_success !MINGW \\\n>  \t'blame respects --encoding=none' '\n> -\tgit blame --incremental --encoding=none file | \\\n> -\t\tegrep \"^(author|summary) \" > actual &&\n> -\ttest_cmp actual expected\n> +\tgit blame --incremental --encoding=none file >output &&\n> +\tfilter_author_summary output >actual &&\n> +\ttest_cmp expected actual\n>  '\n>  \n>  test_done\n> -- \n> 2.7.2-532-g79873b4\n"}]}