{"thread":{"id":"38226","subject":"[PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","startedAt":"2014-12-22T17:52:24Z","lastAt":"2014-12-23T18:26:48Z","messageCount":10,"participants":["Ben Walton","Junio C Hamano","Jonathan Nieder"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"253941","messageId":"1419270744-1408-1-git-send-email-bdwalton@gmail.com","threadId":"38226","inReplyTo":null,"subject":"[PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Ben Walton","fromEmail":"bdwalton@gmail.com","sentAt":"2014-12-22T17:52:24Z","receivedAt":"2014-12-22T17:52:24Z","isPatch":true,"sender":{"key":"bdwalton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/396061?v=4"},"body":"The awk statements previously used in this test weren't compatible\nwith the native versions of awk on Solaris:\n\necho \"dir\" | /bin/awk -v c=0 '$1 {++c} END {print c}'\nawk: syntax error near line 1\nawk: bailing out near line 1\n\necho \"dir\" | /usr/xpg4/bin/awk -v c=0 '$1 {++c} END {print c}'\n0\n\nAnd with GNU awk for comparison:\necho \"dir\" | /opt/csw/gnu/awk -v c=0 '$1 {++c} END {print c}'\n1\n\nInstead of modifying the awk code to work, use wc -w instead as that\nis both adequate and simpler.\n\nSigned-off-by: Ben Walton <bdwalton@gmail.com>\n---\n t/t0090-cache-tree.sh | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/t/t0090-cache-tree.sh b/t/t0090-cache-tree.sh\nindex 067f4c6..f2b1c9c 100755\n--- a/t/t0090-cache-tree.sh\n+++ b/t/t0090-cache-tree.sh\n@@ -22,7 +22,7 @@ generate_expected_cache_tree_rec () {\n \t# ls-files might have foo/bar, foo/bar/baz, and foo/bar/quux\n \t# We want to count only foo because it's the only direct child\n \tsubtrees=$(git ls-files|grep /|cut -d / -f 1|uniq) &&\n-\tsubtree_count=$(echo \"$subtrees\"|awk -v c=0 '$1 {++c} END {print c}') &&\n+\tsubtree_count=$(echo \"$subtrees\"|wc -w) &&\n \tentries=$(git ls-files|wc -l) &&\n \tprintf \"SHA $dir (%d entries, %d subtrees)\\n\" \"$entries\" \"$subtree_count\" &&\n \tfor subtree in $subtrees\n-- \n1.9.1\n"},{"id":"253949","messageId":"xmqqk31j8ik9.fsf@gitster.dls.corp.google.com","threadId":"38226","inReplyTo":"1419270744-1408-1-git-send-email-bdwalton@gmail.com","subject":"Re: [PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-12-22T21:45:42Z","receivedAt":"2014-12-22T21:45:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Walton <bdwalton@gmail.com> writes:\n\n> The awk statements previously used in this test weren't compatible\n> with the native versions of awk on Solaris:\n>\n> echo \"dir\" | /bin/awk -v c=0 '$1 {++c} END {print c}'\n> awk: syntax error near line 1\n> awk: bailing out near line 1\n>\n> echo \"dir\" | /usr/xpg4/bin/awk -v c=0 '$1 {++c} END {print c}'\n> 0\n>\n> And with GNU awk for comparison:\n> echo \"dir\" | /opt/csw/gnu/awk -v c=0 '$1 {++c} END {print c}'\n> 1\n>\n> Instead of modifying the awk code to work, use wc -w instead as that\n> is both adequate and simpler.\n\nHmm, why \"wc -w\" not \"wc -l\", though?  Is somebody squashing a\none-elem-per-line output from ls-files onto a single line?\n"},{"id":"253954","messageId":"20141222220117.GS29365@google.com","threadId":"38226","inReplyTo":"1419270744-1408-1-git-send-email-bdwalton@gmail.com","subject":"Re: [PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2014-12-22T22:01:17Z","receivedAt":"2014-12-22T22:01:17Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Ben Walton wrote:\n\n> echo \"dir\" | /usr/xpg4/bin/awk -v c=0 '$1 {++c} END {print c}'\n> 0\n\nThanks.  Weird.  Does\n\n\tawk -v c=0 '$1 != \"\" {++c} END {print c}'\n\nwork better?\n\n[...]\n> --- a/t/t0090-cache-tree.sh\n> +++ b/t/t0090-cache-tree.sh\n> @@ -22,7 +22,7 @@ generate_expected_cache_tree_rec () {\n>  \t# ls-files might have foo/bar, foo/bar/baz, and foo/bar/quux\n>  \t# We want to count only foo because it's the only direct child\n>  \tsubtrees=$(git ls-files|grep /|cut -d / -f 1|uniq) &&\n> -\tsubtree_count=$(echo \"$subtrees\"|awk -v c=0 '$1 {++c} END {print c}') &&\n> +\tsubtree_count=$(echo \"$subtrees\"|wc -w) &&\n>  \tentries=$(git ls-files|wc -l) &&\n>  \tprintf \"SHA $dir (%d entries, %d subtrees)\\n\" \"$entries\" \"$subtree_count\" &&\n\nSome implementations of wc add a trailing space, causing\n\n\tprintf: 1 : invalid number\n\nUsing\n\n\tprintf \"SHA $dir (%d entries, %d subtrees)\\n\" \"$entries\" $subtree_count &&\n\n(with no quotes around $subtree_count) would avoid trouble, though\nthat's a little subtle.\n\nHope that helps,\nJonathan\n"},{"id":"253955","messageId":"20141222220209.GT29365@google.com","threadId":"38226","inReplyTo":"xmqqk31j8ik9.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2014-12-22T22:02:09Z","receivedAt":"2014-12-22T22:02:09Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n> Ben Walton <bdwalton@gmail.com> writes:\n\n>> echo \"dir\" | /usr/xpg4/bin/awk -v c=0 '$1 {++c} END {print c}'\n>> 0\n>>\n>> And with GNU awk for comparison:\n>> echo \"dir\" | /opt/csw/gnu/awk -v c=0 '$1 {++c} END {print c}'\n>> 1\n>>\n>> Instead of modifying the awk code to work, use wc -w instead as that\n>> is both adequate and simpler.\n>\n> Hmm, why \"wc -w\" not \"wc -l\", though?  Is somebody squashing a\n> one-elem-per-line output from ls-files onto a single line?\n\nThe old code was trying to skip empty lines.\n"},{"id":"253960","messageId":"xmqq7fxj8gp3.fsf@gitster.dls.corp.google.com","threadId":"38226","inReplyTo":"20141222220209.GT29365@google.com","subject":"Re: [PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-12-22T22:26:00Z","receivedAt":"2014-12-22T22:26:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> Junio C Hamano wrote:\n>> Ben Walton <bdwalton@gmail.com> writes:\n>\n>>> echo \"dir\" | /usr/xpg4/bin/awk -v c=0 '$1 {++c} END {print c}'\n>>> 0\n>>>\n>>> And with GNU awk for comparison:\n>>> echo \"dir\" | /opt/csw/gnu/awk -v c=0 '$1 {++c} END {print c}'\n>>> 1\n>>>\n>>> Instead of modifying the awk code to work, use wc -w instead as that\n>>> is both adequate and simpler.\n>>\n>> Hmm, why \"wc -w\" not \"wc -l\", though?  Is somebody squashing a\n>> one-elem-per-line output from ls-files onto a single line?\n>\n> The old code was trying to skip empty lines.\n\nAhh, I misread the original.\n\nYour suggestion to explicitly check $1 != \"\" makes sense to me now.\n\nTo be blunt, I do not have much sympathy to those who insist using\n/usr/bin versions of various tools on Solaris that are overriden by\nxpg variants, but it is somewhat disturbing that the one from xpg4\ndoes not work.\n"},{"id":"253974","messageId":"xmqqd27b6zd3.fsf@gitster.dls.corp.google.com","threadId":"38226","inReplyTo":"1419270744-1408-1-git-send-email-bdwalton@gmail.com","subject":"Re: [PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-12-22T23:25:44Z","receivedAt":"2014-12-22T23:25:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"From: Ben Walton <bdwalton@gmail.com>\n\nThe awk statements previously used in this test weren't compatible\nwith the native versions of awk on Solaris:\n\n    echo \"dir\" | /bin/awk -v c=0 '$1 {++c} END {print c}'\n    awk: syntax error near line 1\n    awk: bailing out near line 1\n\n    echo \"dir\" | /usr/xpg4/bin/awk -v c=0 '$1 {++c} END {print c}'\n    0\n\nAnd with GNU awk for comparison:\n\n    echo \"dir\" | /opt/csw/gnu/awk -v c=0 '$1 {++c} END {print c}'\n    1\n\nWork it around by using $1 != \"\" to state more explicitly that we\nare skipping empty lines.\n\nHelped-by: Jonathan Nieder <jrnieder@gmail.com>\nSigned-off-by: Ben Walton <bdwalton@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n * Then let's queue this, perhaps?\n\n t/t0090-cache-tree.sh | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/t/t0090-cache-tree.sh b/t/t0090-cache-tree.sh\nindex 067f4c6..601d02d 100755\n--- a/t/t0090-cache-tree.sh\n+++ b/t/t0090-cache-tree.sh\n@@ -22,7 +22,7 @@ generate_expected_cache_tree_rec () {\n \t# ls-files might have foo/bar, foo/bar/baz, and foo/bar/quux\n \t# We want to count only foo because it's the only direct child\n \tsubtrees=$(git ls-files|grep /|cut -d / -f 1|uniq) &&\n-\tsubtree_count=$(echo \"$subtrees\"|awk -v c=0 '$1 {++c} END {print c}') &&\n+\tsubtree_count=$(echo \"$subtrees\"|awk -v c=0 '$1 != \"\" {++c} END {print c}') &&\n \tentries=$(git ls-files|wc -l) &&\n \tprintf \"SHA $dir (%d entries, %d subtrees)\\n\" \"$entries\" \"$subtree_count\" &&\n \tfor subtree in $subtrees\n"},{"id":"253975","messageId":"xmqq8uhz6za0.fsf@gitster.dls.corp.google.com","threadId":"38226","inReplyTo":"xmqqd27b6zd3.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-12-22T23:27:35Z","receivedAt":"2014-12-22T23:27:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> From: Ben Walton <bdwalton@gmail.com>\n>\n> The awk statements previously used in this test weren't compatible\n> with the native versions of awk on Solaris:\n>\n>     echo \"dir\" | /bin/awk -v c=0 '$1 {++c} END {print c}'\n>     awk: syntax error near line 1\n>     awk: bailing out near line 1\n>\n>     echo \"dir\" | /usr/xpg4/bin/awk -v c=0 '$1 {++c} END {print c}'\n>     0\n>\n> And with GNU awk for comparison:\n>\n>     echo \"dir\" | /opt/csw/gnu/awk -v c=0 '$1 {++c} END {print c}'\n>     1\n>\n> Work it around by using $1 != \"\" to state more explicitly that we\n> are skipping empty lines.\n>\n> Helped-by: Jonathan Nieder <jrnieder@gmail.com>\n> Signed-off-by: Ben Walton <bdwalton@gmail.com>\n> Signed-off-by: Junio C Hamano <gitster@pobox.com>\n> ---\n>\n>  * Then let's queue this, perhaps?\n\nheh, not like that without updating the subject, perhaps like this:\n\nSubject: t0090: tweak awk statement for Solaris /usr/xpg4/bin/awk\n\nSorry for the noise.\n\n\n\n>  t/t0090-cache-tree.sh | 2 +-\n>  1 file changed, 1 insertion(+), 1 deletion(-)\n>\n> diff --git a/t/t0090-cache-tree.sh b/t/t0090-cache-tree.sh\n> index 067f4c6..601d02d 100755\n> --- a/t/t0090-cache-tree.sh\n> +++ b/t/t0090-cache-tree.sh\n> @@ -22,7 +22,7 @@ generate_expected_cache_tree_rec () {\n>  \t# ls-files might have foo/bar, foo/bar/baz, and foo/bar/quux\n>  \t# We want to count only foo because it's the only direct child\n>  \tsubtrees=$(git ls-files|grep /|cut -d / -f 1|uniq) &&\n> -\tsubtree_count=$(echo \"$subtrees\"|awk -v c=0 '$1 {++c} END {print c}') &&\n> +\tsubtree_count=$(echo \"$subtrees\"|awk -v c=0 '$1 != \"\" {++c} END {print c}') &&\n>  \tentries=$(git ls-files|wc -l) &&\n>  \tprintf \"SHA $dir (%d entries, %d subtrees)\\n\" \"$entries\" \"$subtree_count\" &&\n>  \tfor subtree in $subtrees\n"},{"id":"253976","messageId":"xmqq4msn6yyl.fsf@gitster.dls.corp.google.com","threadId":"38226","inReplyTo":"xmqqd27b6zd3.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-12-22T23:34:26Z","receivedAt":"2014-12-22T23:34:26Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> From: Ben Walton <bdwalton@gmail.com>\n>\n> The awk statements previously used in this test weren't compatible\n> with the native versions of awk on Solaris:\n>\n>     echo \"dir\" | /bin/awk -v c=0 '$1 {++c} END {print c}'\n>     awk: syntax error near line 1\n>     awk: bailing out near line 1\n>\n>     echo \"dir\" | /usr/xpg4/bin/awk -v c=0 '$1 {++c} END {print c}'\n>     0\n>\n> And with GNU awk for comparison:\n>\n>     echo \"dir\" | /opt/csw/gnu/awk -v c=0 '$1 {++c} END {print c}'\n>     1\n>\n> Work it around by using $1 != \"\" to state more explicitly that we\n> are skipping empty lines.\n\nBy the way, I was hoping (eh, what kind of hope is that???) that $1\nalone is not a kosher POSIX way but a GNUism, but that does not seem\nto be the case.  POSIX has this [*1*]\n\n    When an expression is used in a Boolean context, if it has a\n    numeric value, a value of zero shall be treated as false and any\n    other value shall be treated as true. Otherwise, a string value\n    of the null string shall be treated as false and any other value\n    shall be treated as true. A Boolean context shall be one of the\n    following:\n\nand among the \"Boolean context\" listed is:\n\n    * An expression used as a pattern (as in Overall Program Structure)\n\nSo the example with /usr/xpg4/bin/awk does not seem to be a\nbehaviour from a conformant implementationd, and it seems to be\ncorrect to label this as \"work it around by ...\" (not \"avoid using\nGNUism\").\n\nWe learn new things every day (not that I really wanted to learn\nglitches in various implementations of awk) ;-).\n\nThanks.\n\n\n[Reference]\n\n*1* http://pubs.opengroup.org/onlinepubs/9699919799/utilities/awk.html\n"},{"id":"253977","messageId":"20141222233807.GU29365@google.com","threadId":"38226","inReplyTo":"xmqq8uhz6za0.fsf@gitster.dls.corp.google.com","subject":"Re: [PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2014-12-22T23:38:07Z","receivedAt":"2014-12-22T23:38:07Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n\n>> From: Ben Walton <bdwalton@gmail.com>\n>>\n>> The awk statements previously used in this test weren't compatible\n>> with the native versions of awk on Solaris:\n>>\n>>     echo \"dir\" | /bin/awk -v c=0 '$1 {++c} END {print c}'\n>>     awk: syntax error near line 1\n>>     awk: bailing out near line 1\n\nIf I were doing it, I'd leave the above four lines out --- they are\ndescribing an unrelated problem.  I wonder if we should make the test\nharness respect SANE_TOOL_PATH to avoid that kind of problem in the\nfuture.\n\n[...]\n> heh, not like that without updating the subject, perhaps like this:\n>\n> Subject: t0090: tweak awk statement for Solaris /usr/xpg4/bin/awk\n\nWith the updated subject,\n\nReviewed-by: Jonathan Nieder <jrnieder@gmail.com>\n\nThanks.\n"},{"id":"254034","messageId":"xmqqk31i2pef.fsf@gitster.dls.corp.google.com","threadId":"38226","inReplyTo":"20141222233807.GU29365@google.com","subject":"Re: [PATCH] Use wc instead of awk to count subtrees in t0090-cache-tree","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-12-23T18:26:48Z","receivedAt":"2014-12-23T18:26:48Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> With the updated subject,\n>\n> Reviewed-by: Jonathan Nieder <jrnieder@gmail.com>\n\nThanks.  Here is what I tentatively queued for today's pushout.\n\n-- >8 --\nFrom: Ben Walton <bdwalton@gmail.com>\nDate: Mon, 22 Dec 2014 15:25:44 -0800\nSubject: [PATCH] t0090: tweak awk statement for Solaris /usr/xpg4/bin/awk\n\nThe awk statements previously used in this test weren't compatible\nwith the native versions of awk on Solaris:\n\n    echo \"dir\" | /bin/awk -v c=0 '$1 {++c} END {print c}'\n    awk: syntax error near line 1\n    awk: bailing out near line 1\n\n    echo \"dir\" | /usr/xpg4/bin/awk -v c=0 '$1 {++c} END {print c}'\n    0\n\nEven though we do not cater to tools in /usr/bin on Solaris that are\noverridden by corresponding ones in /usr/xpg?/bin, in this case,\neven the XPG version does not work correctly.\n\nWith GNU awk for comparison:\n\n    echo \"dir\" | /opt/csw/gnu/awk -v c=0 '$1 {++c} END {print c}'\n    1\n\nwhich is what this test expects (and is in line with POSIX; non-empty\nstring is true and an empty string is false).\n\nWork this issue around by using $1 != \"\" to state more explicitly\nthat we are skipping empty lines.\n\nHelped-by: Jonathan Nieder <jrnieder@gmail.com>\nSigned-off-by: Ben Walton <bdwalton@gmail.com>\nReviewed-by: Jonathan Nieder <jrnieder@gmail.com>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n t/t0090-cache-tree.sh | 2 +-\n 1 file changed, 1 insertion(+), 1 deletion(-)\n\ndiff --git a/t/t0090-cache-tree.sh b/t/t0090-cache-tree.sh\nindex 067f4c6..601d02d 100755\n--- a/t/t0090-cache-tree.sh\n+++ b/t/t0090-cache-tree.sh\n@@ -22,7 +22,7 @@ generate_expected_cache_tree_rec () {\n \t# ls-files might have foo/bar, foo/bar/baz, and foo/bar/quux\n \t# We want to count only foo because it's the only direct child\n \tsubtrees=$(git ls-files|grep /|cut -d / -f 1|uniq) &&\n-\tsubtree_count=$(echo \"$subtrees\"|awk -v c=0 '$1 {++c} END {print c}') &&\n+\tsubtree_count=$(echo \"$subtrees\"|awk -v c=0 '$1 != \"\" {++c} END {print c}') &&\n \tentries=$(git ls-files|wc -l) &&\n \tprintf \"SHA $dir (%d entries, %d subtrees)\\n\" \"$entries\" \"$subtree_count\" &&\n \tfor subtree in $subtrees\n-- \n2.2.1-321-gd161b79\n"}]}