{"thread":{"id":"6723","subject":"git log filtering","startedAt":"2007-02-07T16:41:41Z","lastAt":"2007-03-07T18:03:32Z","messageCount":34,"participants":["Don Zickus","Uwe Kleine-König","Linus Torvalds","Jakub Narebski","Johannes Schindelin","Junio C Hamano","Horst H. von Brand","Jeff King","Shawn O. Pearce","Sergey Vlasov","Paolo Bonzini"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"33819","messageId":"68948ca0702070841m76817d9el7ce2ec69835c50e@mail.gmail.com","threadId":"6723","inReplyTo":null,"subject":"git log filtering","fromName":"Don Zickus","fromEmail":"dzickus@gmail.com","sentAt":"2007-02-07T16:41:41Z","receivedAt":"2007-02-07T16:41:41Z","isPatch":false,"sender":{"key":"dzickus@gmail.com","avatar":"https://gravatar.com/avatar/fbc96d0d5584c05dec11867b861650fe9f5d7d0ddec2655a1f90542fe07d9769?d=mp&s=160"},"body":"I was curious to know what is the easiest way to filter info inside a\ncommit message.\n\nFor example say I wanted to find out what patches Joe User has\nsubmitted to the git project.\nI know I can do something like ' git log |grep -B2 \"^Author: Joe User\"\n' and it will output the matches and the commit id.  However, if I\nwanted to filter on something like \"Signed-off-by: Joe User\", then it\nis a little harder to dig for the commit id.\n\nIs there a better way of doing this?  Or should I accept the fact that\ngit wasn't designed to filter info like this very quickly?\n\nI guess what I was looking to do was embed some metadata inside the\ncommit message and parse through it at a later time (ie like a\nbugzilla number or something).\n\nAny thoughts/tips/tricks would be helpful.\n\nCheers,\nDon\n"},{"id":"33824","messageId":"eqd09b$4hg$1@sea.gmane.org","threadId":"6723","inReplyTo":"68948ca0702070841m76817d9el7ce2ec69835c50e@mail.gmail.com","subject":"Re: git log filtering","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-02-07T16:55:07Z","receivedAt":"2007-02-07T16:55:07Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[Cc: git@vger.kernel.org]\n\nDon Zickus wrote:\n\n> I was curious to know what is the easiest way to filter info inside a\n> commit message.\n> \n> For example say I wanted to find out what patches Joe User has\n> submitted to the git project.\n>\n> I know I can do something like ' git log |grep -B2 \"^Author: Joe User\"\n> ' and it will output the matches and the commit id.  However, if I\n> wanted to filter on something like \"Signed-off-by: Joe User\", then it\n> is a little harder to dig for the commit id.\n> \n> Is there a better way of doing this?  Or should I accept the fact that\n> git wasn't designed to filter info like this very quickly?\n\nYou can use \"git log --grep=<pattern>\" for that, instead. This greps\nraw commit message. You can use --author and --comitter to grep those\nheaders.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"33820","messageId":"20070207170122.GB18704@informatik.uni-freiburg.de","threadId":"6723","inReplyTo":"68948ca0702070841m76817d9el7ce2ec69835c50e@mail.gmail.com","subject":"Re: git log filtering","fromName":"Uwe Kleine-König","fromEmail":"ukleinek@informatik.uni-freiburg.de","sentAt":"2007-02-07T17:01:22Z","receivedAt":"2007-02-07T17:01:22Z","isPatch":false,"sender":{"key":"u.kleine-koenig@pengutronix.de","avatar":"https://gravatar.com/avatar/354b5e3ceb2806a2f1e1e382ac29ddbdad18288654da62b61eb13583a857eee7?d=mp&s=160"},"body":"Don Zickus wrote:\n> I was curious to know what is the easiest way to filter info inside a\n> commit message.\n> \n> For example say I wanted to find out what patches Joe User has\n> submitted to the git project.\n> I know I can do something like ' git log |grep -B2 \"^Author: Joe User\"\nWhat about\n\n\tgit log --author=\"Joe User\"\n\n> ' and it will output the matches and the commit id.  However, if I\n> wanted to filter on something like \"Signed-off-by: Joe User\", then it\n> is a little harder to dig for the commit id.\n> \n> Is there a better way of doing this?  Or should I accept the fact that\n> git wasn't designed to filter info like this very quickly?\n> \n> I guess what I was looking to do was embed some metadata inside the\n> commit message and parse through it at a later time (ie like a\n> bugzilla number or something).\n> \n> Any thoughts/tips/tricks would be helpful.\n\nMaybe:\n\n\tgit log | awk -v sob=\"Joe User\" '$1 == \"commit\" {commit = $2} /Signed-off-by:/ {if (match($0, sob)) print commit}'\n\nBest regards\nUwe\n\n-- \nUwe Kleine-König\n\nhttp://www.google.com/search?q=2004+in+roman+numerals\n"},{"id":"33825","messageId":"Pine.LNX.4.63.0702071811060.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6723","inReplyTo":"20070207170122.GB18704@informatik.uni-freiburg.de","subject":"Re: git log filtering","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-07T17:12:12Z","receivedAt":"2007-02-07T17:12:12Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 7 Feb 2007, Uwe Kleine-König wrote:\n\n> Don Zickus wrote:\n> > I was curious to know what is the easiest way to filter info inside a\n> > commit message.\n> > \n> > For example say I wanted to find out what patches Joe User has\n> > submitted to the git project.\n> > I know I can do something like ' git log |grep -B2 \"^Author: Joe User\"\n> What about\n> \n> \tgit log --author=\"Joe User\"\n> \n> > ' and it will output the matches and the commit id.  However, if I\n> > wanted to filter on something like \"Signed-off-by: Joe User\", then it\n> > is a little harder to dig for the commit id.\n> > \n> > Is there a better way of doing this?  Or should I accept the fact that\n> > git wasn't designed to filter info like this very quickly?\n> > \n> > I guess what I was looking to do was embed some metadata inside the\n> > commit message and parse through it at a later time (ie like a\n> > bugzilla number or something).\n> > \n> > Any thoughts/tips/tricks would be helpful.\n> \n> Maybe:\n> \n> \tgit log | awk -v sob=\"Joe User\" '$1 == \"commit\" {commit = $2} /Signed-off-by:/ {if (match($0, sob)) print commit}'\n\n*grin* Why do you know --author, but not --grep?\n\ngit log --grep=Signed-off-by:\\ Joe\\ User\n\nCiao,\nDscho\n"},{"id":"33823","messageId":"Pine.LNX.4.64.0702070856190.8424@woody.linux-foundation.org","threadId":"6723","inReplyTo":"68948ca0702070841m76817d9el7ce2ec69835c50e@mail.gmail.com","subject":"Re: git log filtering","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-07T17:12:38Z","receivedAt":"2007-02-07T17:12:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 7 Feb 2007, Don Zickus wrote:\n>\n> I was curious to know what is the easiest way to filter info inside a\n> commit message.\n> \n> For example say I wanted to find out what patches Joe User has\n> submitted to the git project.\n> I know I can do something like ' git log |grep -B2 \"^Author: Joe User\"\n> ' and it will output the matches and the commit id.  However, if I\n> wanted to filter on something like \"Signed-off-by: Joe User\", then it\n> is a little harder to dig for the commit id.\n\nThere are two ways:\n\n - \"git log\" can itself do a lot of filtering. Both on date, on revisions, \n   on \"modifies files/directories X, Y and Z\" _and_ on strings.\n\n   See \"man git-rev-list\" for more (it doesn't apply to just \"git log\", it \n   applies to just about any revision listing, including gitk etc)\n\n   For example,\n\n\tgit log [--author=pattern] [--committer=pattern] [--grep=pattern]\n\n   will likely do exactly what you want. You can do\n\n\tgit log --grep=\"Signed-off-by:.*akpm\"\n\n   on the kernel archive to see which ones were signed off by Andrew.\n\nSo the above works, and catches *most* uses. But it has problems if you \nwant to do something fancier (and I think that includes something as \nsimple as doing a case-insensitive grep). So the other approach is:\n\n - The hacky way: use \"git log --pretty -z\", and GNU grep -z:\n\n\tgit log --pretty -z |\n\t\tgrep -i -z Signed-off-by:.*junkio |\n\t\ttr '\\0' '\\n'\n\n   which allows you to do anything you want with grep (or other unix tools \n   that take zero-terminated output).\n\n> Is there a better way of doing this?  Or should I accept the fact that\n> git wasn't designed to filter info like this very quickly?\n\nGit definitely was designed to do it. The \"-z\" option in particular is \nvery much designed for any generic UNIX scripting, but the *easy* cases \ngit does internally.\n\n\t\tLinus\n"},{"id":"33826","messageId":"Pine.LNX.4.63.0702071822430.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702070856190.8424@woody.linux-foundation.org","subject":"Re: git log filtering","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-07T17:25:04Z","receivedAt":"2007-02-07T17:25:04Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 7 Feb 2007, Linus Torvalds wrote:\n\n> You can do\n> \n> \tgit log --grep=\"Signed-off-by:.*akpm\"\n> \n>    on the kernel archive to see which ones were signed off by Andrew.\n> \n> So the above works, and catches *most* uses. But it has problems if you \n> want to do something fancier (and I think that includes something as \n> simple as doing a case-insensitive grep).\n\n[TIC PATCH] revision.c: accept \"-i\" to make --grep case insensitive\n\nWhen calling\n\n\tgit log --grep=blabla -i --grep=blublu\n\nthe expression \"blabla\" is greppend case _sensitively_, but \"blublu\"\ncase _insensitively_.\n\nSigned-off-by: Johannes Schindelin <Johannes.Schindelin@gmx.de>\n\n---\n\n revision.c |    9 +++++++++\n 1 files changed, 9 insertions(+), 0 deletions(-)\n\ndiff --git a/revision.c b/revision.c\nindex 42ba310..843aa8e 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -9,6 +9,8 @@\n #include \"grep.h\"\n #include \"reflog-walk.h\"\n \n+static int case_insensitive_grep = 0;\n+\n static char *path_name(struct name_path *path, const char *name)\n {\n \tstruct name_path *p;\n@@ -742,6 +744,8 @@ static void add_grep(struct rev_info *revs, const char *ptn, enum grep_pat_token\n \t\topt->status_only = 1;\n \t\topt->pattern_tail = &(opt->pattern_list);\n \t\topt->regflags = REG_NEWLINE;\n+\t\tif (case_insensitive_grep)\n+\t\t\topt->regflags |= REG_ICASE;\n \t\trevs->grep_filter = opt;\n \t}\n \tappend_grep_pattern(revs->grep_filter, ptn,\n@@ -1042,6 +1046,11 @@ int setup_revisions(int argc, const char **argv, struct rev_info *revs, const ch\n \t\t\t\tadd_header_grep(revs, \"committer\", arg+12);\n \t\t\t\tcontinue;\n \t\t\t}\n+\t\t\tif (!strcmp(arg, \"-i\") ||\n+\t\t\t\t\t!strcmp(arg, \"--case-insensitive\")) {\n+\t\t\t\tcase_insensitive_grep = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n \t\t\tif (!strncmp(arg, \"--grep=\", 7)) {\n \t\t\t\tadd_message_grep(revs, arg+7);\n \t\t\t\tcontinue;\n"},{"id":"33832","messageId":"Pine.LNX.4.64.0702070919320.8424@woody.linux-foundation.org","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702070856190.8424@woody.linux-foundation.org","subject":"Re: git log filtering","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-07T18:16:15Z","receivedAt":"2007-02-07T18:16:15Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 7 Feb 2007, Linus Torvalds wrote:\n> \n> \tgit log --pretty -z |\n\nGaah. If all you want is normal logs, you don't need the \"--pretty\", \nof course, since that's the default. Just \"git log -z\" will give you \nzero-terminated logs. \n\nBut if you want to grep on committer, you'd need to use \"--pretty=full\" or \nsomething, of course, so the \"--pretty=xyz\" thing is indeed often \napplicable for things like this.\n\nAlso, I just checked, and we have a bug. Merges do not have the ending \nzero in \"git log -z\" output. It seems to be connected to the fact that we \nhandle the \"always_show_header\" commits differently (the ones that we \nwouldn't normally show because they have no diffs associated with them).\n\nThe obvious fix for that failed. I'll look at it some more.\n\n\t\tLinus\n"},{"id":"33833","messageId":"68948ca0702071019i5704f24fnf2b9a7d6dfa74d86@mail.gmail.com","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702070856190.8424@woody.linux-foundation.org","subject":"Re: git log filtering","fromName":"Don Zickus","fromEmail":"dzickus@gmail.com","sentAt":"2007-02-07T18:19:03Z","receivedAt":"2007-02-07T18:19:03Z","isPatch":false,"sender":{"key":"dzickus@gmail.com","avatar":"https://gravatar.com/avatar/fbc96d0d5584c05dec11867b861650fe9f5d7d0ddec2655a1f90542fe07d9769?d=mp&s=160"},"body":">  - \"git log\" can itself do a lot of filtering. Both on date, on revisions,\n>    on \"modifies files/directories X, Y and Z\" _and_ on strings.\n>\n>    See \"man git-rev-list\" for more (it doesn't apply to just \"git log\", it\n>    applies to just about any revision listing, including gitk etc)\n>\n>    For example,\n>\n>         git log [--author=pattern] [--committer=pattern] [--grep=pattern]\n>\n>    will likely do exactly what you want. You can do\n>\n>         git log --grep=\"Signed-off-by:.*akpm\"\n>\n>    on the kernel archive to see which ones were signed off by Andrew.\n\nCool.  The hidden little options.  :-)  This is exactly what I was\nlooking for.  Thanks.\n\nI didn't see these options in the man pages.  Might be worth putting in there??\n\nCheers,\nDon\n"},{"id":"33834","messageId":"Pine.LNX.4.64.0702071025100.8424@woody.linux-foundation.org","threadId":"6723","inReplyTo":"68948ca0702071019i5704f24fnf2b9a7d6dfa74d86@mail.gmail.com","subject":"Re: git log filtering","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-07T18:27:17Z","receivedAt":"2007-02-07T18:27:17Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 7 Feb 2007, Don Zickus wrote:\n>\n> I didn't see these options in the man pages.  Might be worth putting in\n> there??\n\nWell, they really _are_ there, indirectly:\n\n\tThe command takes options applicable to the git-rev-list(1) command \n\tto control what is shown and how, and options applicable to the \n\tgit-diff-tree(1) commands to control how the change each commit \n\tintroduces are shown.\n\nso you have to look at both git-rev-list and git-diff-tree to get all the \noptions.\n\nIt then goes on to say:\n\n\tThis manual page describes only the most frequently used options.\n\t                           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\n\nso technically it's complete and true.\n\nBut yeah, maybe we could include all the options there.\n\n\t\tLinus\n"},{"id":"33836","messageId":"Pine.LNX.4.64.0702071139090.8424@woody.linux-foundation.org","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702070919320.8424@woody.linux-foundation.org","subject":"Fix \"git log -z\" behaviour","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-07T19:49:56Z","receivedAt":"2007-02-07T19:49:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nFor commit messages, we should really put the \"line_termination\" when we \noutput the character in between different commits, *not* between the \ncommit and the diff. The diff goes hand-in-hand with the commit, it \nshouldn't be separated from it with the termination character.\n\nSo this:\n - uses the termination character for true inter-commit spacing\n - uses a regular newline between the commit log and the diff\n\nWe had it the other way around.\n\nFor the normal case where the termination character is '\\n', this \nobviously doesn't change anything at all, since we just switched two \nidentical characters around. So it's very safe - it doesn't change any \nnormal usage, but it definitely fixes \"git log -z\".\n\nBy fixing \"git log -z\", you can now also do insane things like\n\n\tgit log -p -z |\n\t\tgrep -z \"some patch expression\" |\n\t\ttr '\\0' '\\n' |\n\t\tless -S\n\nand you will see only those commits that have the \"some patch expression\" \nin their commit message _or_ their patches.\n\n(This is slightly different from 'git log -S\"some patch expression\"', \nsince the latter requires the expression to literally *change* in the \npatch, while the \"git log -p -z | grep ..\" approach will see it if it's \njust an unchanged _part_ of the patch context)\n\nOf course, if you actually do something like the above, you're probably \ninsane, but hey, it works!\n\nTry the above command line for a demonstration (of course, you need to \nchange the \"some patch expression\" to be something relevant). The old \nbehaviour of \"git log -p -z\" was useless (and got things completely wrong \nfor log entries without patches).\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n\nOn Wed, 7 Feb 2007, Linus Torvalds wrote:\n> \n> Also, I just checked, and we have a bug. Merges do not have the ending \n> zero in \"git log -z\" output. It seems to be connected to the fact that we \n> handle the \"always_show_header\" commits differently (the ones that we \n> wouldn't normally show because they have no diffs associated with them).\n> \n> The obvious fix for that failed. I'll look at it some more.\n\nActually, the obvious fix was right, I just did the *wrong* obvious fix at \nfirst ;)\n\n log-tree.c |    7 +++----\n 1 files changed, 3 insertions(+), 4 deletions(-)\n\ndiff --git a/log-tree.c b/log-tree.c\nindex d8ca36b..85acd66 100644\n--- a/log-tree.c\n+++ b/log-tree.c\n@@ -143,7 +143,7 @@ void show_log(struct rev_info *opt, const char *sep)\n \tif (*sep != '\\n' && opt->commit_format == CMIT_FMT_ONELINE)\n \t\textra = \"\\n\";\n \tif (opt->shown_one && opt->commit_format != CMIT_FMT_ONELINE)\n-\t\tputchar('\\n');\n+\t\tputchar(opt->diffopt.line_termination);\n \topt->shown_one = 1;\n \n \t/*\n@@ -270,9 +270,8 @@ int log_tree_diff_flush(struct rev_info *opt)\n \t\t    opt->commit_format != CMIT_FMT_ONELINE) {\n \t\t\tint pch = DIFF_FORMAT_DIFFSTAT | DIFF_FORMAT_PATCH;\n \t\t\tif ((pch & opt->diffopt.output_format) == pch)\n-\t\t\t\tprintf(\"---%c\", opt->diffopt.line_termination);\n-\t\t\telse\n-\t\t\t\tputchar(opt->diffopt.line_termination);\n+\t\t\t\tprintf(\"---\");\n+\t\t\tputchar('\\n');\n \t\t}\n \t}\n \tdiff_flush(&opt->diffopt);\n"},{"id":"33837","messageId":"7vmz3p7neb.fsf@assigned-by-dhcp.cox.net","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702071139090.8424@woody.linux-foundation.org","subject":"Re: Fix \"git log -z\" behaviour","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-07T19:55:56Z","receivedAt":"2007-02-07T19:55:56Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ah, I was looking at other minor issues and then came up with\nthis one liner.  But obviously \"termination should be the true\ninter-commit spacing\" is the right direction, so I'll chuck this\none.\n\ndiff --git a/log-tree.c b/log-tree.c\nindex d8ca36b..410f90f 100644\n--- a/log-tree.c\n+++ b/log-tree.c\n@@ -354,6 +354,8 @@ int log_tree_commit(struct rev_info *opt, struct commit *commit)\n \tif (!shown && opt->loginfo && opt->always_show_header) {\n \t\tlog.parent = NULL;\n \t\tshow_log(opt, \"\");\n+\t\tif (!opt->diffopt.line_termination)\n+\t\t\tputchar(0);\n \t\tshown = 1;\n \t}\n \topt->loginfo = NULL;\n"},{"id":"33847","messageId":"Pine.LNX.4.64.0702071257490.8424@woody.linux-foundation.org","threadId":"6723","inReplyTo":"7v64ad7l12.fsf@assigned-by-dhcp.cox.net","subject":"Re: git log filtering","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-07T21:03:05Z","receivedAt":"2007-02-07T21:03:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 7 Feb 2007, Junio C Hamano wrote:\n> \n> This is very tempting but, ... hmmmm...\n\nI would actually prefer to have it be some marker on the expression \nitself.\n\nWe already do that '^' handling by hand for \"author\"/\"committer\" things. \nWe could do other things like that.\n\nAlthough I guess the downside of not doing standard regexps would be too \nbig.\n\n\t\tLinus\n"},{"id":"33849","messageId":"7vps8l65fh.fsf@assigned-by-dhcp.cox.net","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702071257490.8424@woody.linux-foundation.org","subject":"Re: git log filtering","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-07T21:09:22Z","receivedAt":"2007-02-07T21:09:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Wed, 7 Feb 2007, Junio C Hamano wrote:\n>> \n>> This is very tempting but, ... hmmmm...\n>\n> I would actually prefer to have it be some marker on the expression \n> itself.\n>\n> We already do that '^' handling by hand for \"author\"/\"committer\" things. \n> We could do other things like that.\n>\n> Although I guess the downside of not doing standard regexps would be too \n> big.\n>\n> \t\tLinus\n\nWe could go pcre and let you say \"(?i)\".  That would all be post\n1.5.0, though.\n"},{"id":"33859","messageId":"Pine.LNX.4.64.0702071334060.8424@woody.linux-foundation.org","threadId":"6723","inReplyTo":"7vps8l65fh.fsf@assigned-by-dhcp.cox.net","subject":"Re: git log filtering","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-07T21:53:18Z","receivedAt":"2007-02-07T21:53:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 7 Feb 2007, Junio C Hamano wrote:\n> \n> We could go pcre and let you say \"(?i)\".  That would all be post\n> 1.5.0, though.\n\nHmm. PCRE is probably wide-spread enough that it could be an option. \n\nWhat's PCRE performance like? I'd hate to make \"git grep\" slower, and it \nwould be stupid and confusing to use two different regex libraries..\n\nMaybe somebody could test - afaik, PCRE has a regex-compatible (from a API \nstandpoint, not from a regex standpoint!) wrapper thing, and it might be \ninteresting to hear if doing \"git grep\" is slower or faster..\n\n(I realize that the performance thing depends heavily on the patterns and \nthe working set they are used on, but I guess _I_ personally only care \nabout fairly simple patterns on the kernel ;)\n\n\t\tLinus\n"},{"id":"33868","messageId":"68948ca0702071453i3c4d1b66hcf173fc17919acd6@mail.gmail.com","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702071139090.8424@woody.linux-foundation.org","subject":"Re: Fix \"git log -z\" behaviour","fromName":"Don Zickus","fromEmail":"dzickus@gmail.com","sentAt":"2007-02-07T22:53:22Z","receivedAt":"2007-02-07T22:53:22Z","isPatch":false,"sender":{"key":"dzickus@gmail.com","avatar":"https://gravatar.com/avatar/fbc96d0d5584c05dec11867b861650fe9f5d7d0ddec2655a1f90542fe07d9769?d=mp&s=160"},"body":">\n> For commit messages, we should really put the \"line_termination\" when we\n> output the character in between different commits, *not* between the\n> commit and the diff. The diff goes hand-in-hand with the commit, it\n> shouldn't be separated from it with the termination character.\n>\n> So this:\n>  - uses the termination character for true inter-commit spacing\n>  - uses a regular newline between the commit log and the diff\n>\n> We had it the other way around.\n>\n> For the normal case where the termination character is '\\n', this\n> obviously doesn't change anything at all, since we just switched two\n> identical characters around. So it's very safe - it doesn't change any\n> normal usage, but it definitely fixes \"git log -z\".\n>\n> By fixing \"git log -z\", you can now also do insane things like\n>\n>         git log -p -z |\n>                 grep -z \"some patch expression\" |\n>                 tr '\\0' '\\n' |\n>                 less -S\n>\n> and you will see only those commits that have the \"some patch expression\"\n> in their commit message _or_ their patches.\n>\n> (This is slightly different from 'git log -S\"some patch expression\"',\n> since the latter requires the expression to literally *change* in the\n> patch, while the \"git log -p -z | grep ..\" approach will see it if it's\n> just an unchanged _part_ of the patch context)\n>\n> Of course, if you actually do something like the above, you're probably\n> insane, but hey, it works!\n>\n> Try the above command line for a demonstration (of course, you need to\n> change the \"some patch expression\" to be something relevant). The old\n> behaviour of \"git log -p -z\" was useless (and got things completely wrong\n> for log entries without patches).\n>\n> Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n> ---\n>\n> On Wed, 7 Feb 2007, Linus Torvalds wrote:\n> >\n> > Also, I just checked, and we have a bug. Merges do not have the ending\n> > zero in \"git log -z\" output. It seems to be connected to the fact that we\n> > handle the \"always_show_header\" commits differently (the ones that we\n> > wouldn't normally show because they have no diffs associated with them).\n> >\n> > The obvious fix for that failed. I'll look at it some more.\n>\n> Actually, the obvious fix was right, I just did the *wrong* obvious fix at\n> first ;)\n\nWorks for me.  :)\nAnd I thought I had a handle on a lot of the Unix commands.  That -z\nstuff just threw me for a loop.  It's pretty neat to be able to grep\ncommits and have the output display the whole commit and diff.\n\nCheers,\nDon\n"},{"id":"33877","messageId":"Pine.LNX.4.64.0702071503100.8424@woody.linux-foundation.org","threadId":"6723","inReplyTo":"68948ca0702071453i3c4d1b66hcf173fc17919acd6@mail.gmail.com","subject":"Re: Fix \"git log -z\" behaviour","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-07T23:05:11Z","receivedAt":"2007-02-07T23:05:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 7 Feb 2007, Don Zickus wrote:\n> \n> And I thought I had a handle on a lot of the Unix commands.  That -z\n> stuff just threw me for a loop.  It's pretty neat to be able to grep\n> commits and have the output display the whole commit and diff.\n\nThe whole \"-z\" flag to grep is a GNU extension, as far as I know. I don't \nthink it's portable. \n\nEven for GNU grep, it's not mentioned in the man-page. Whether that is \njust due to the normal inane FSF rules (\"man-pages are evil, you should \nuse those idiotic info pages\") or whether it is a conscious effort to not \ndocument nonstandard features, I don't know.\n\n\t\tLinus\n"},{"id":"33908","messageId":"200702080159.l181xiK3021514@laptop13.inf.utfsm.cl","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702071257490.8424@woody.linux-foundation.org","subject":"Re: git log filtering","fromName":"Horst H. von Brand","fromEmail":"vonbrand@inf.utfsm.cl","sentAt":"2007-02-08T01:59:44Z","receivedAt":"2007-02-08T01:59:44Z","isPatch":false,"sender":{"key":"vonbrand@inf.utfsm.cl","avatar":"https://avatars.githubusercontent.com/u/211384?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> On Wed, 7 Feb 2007, Junio C Hamano wrote:\n> > This is very tempting but, ... hmmmm...\n> \n> I would actually prefer to have it be some marker on the expression \n> itself.\n> \n> We already do that '^' handling by hand for \"author\"/\"committer\" things. \n> We could do other things like that.\n> \n> Although I guess the downside of not doing standard regexps would be too \n> big.\n\nUse Perl's regexps? the pcre library packs them, and they have all sorts of\ngoodies like markers in the expression itself. \n-- \nDr. Horst H. von Brand                   User #22616 counter.li.org\nDepartamento de Informatica                    Fono: +56 32 2654431\nUniversidad Tecnica Federico Santa Maria             +56 32 2654239\nCasilla 110-V, Valparaiso, Chile               Fax:  +56 32 2797513\n"},{"id":"33913","messageId":"20070208061654.GA8813@coredump.intra.peff.net","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702071334060.8424@woody.linux-foundation.org","subject":"Re: git log filtering","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-02-08T06:16:54Z","receivedAt":"2007-02-08T06:16:54Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Feb 07, 2007 at 01:53:18PM -0800, Linus Torvalds wrote:\n\n> What's PCRE performance like? I'd hate to make \"git grep\" slower, and it \n> would be stupid and confusing to use two different regex libraries..\n>\n> Maybe somebody could test - afaik, PCRE has a regex-compatible (from a API \n> standpoint, not from a regex standpoint!) wrapper thing, and it might be \n> interesting to hear if doing \"git grep\" is slower or faster..\n\nThe patch is delightfully simple (though a real patch would probably be\nconditional):\n\ndiff --git a/Makefile b/Makefile\nindex aca96c8..cf391dc 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -323,7 +323,7 @@ BUILTIN_OBJS = \\\n \tbuiltin-pack-refs.o\n \n GITLIBS = $(LIB_FILE) $(XDIFF_LIB)\n-EXTLIBS = -lz\n+EXTLIBS = -lz -lpcreposix -lpcre\n \n #\n # Platform specific tweaks\ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex c1bcb00..a6c77f9 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -40,7 +40,7 @@\n #include <sys/poll.h>\n #include <sys/socket.h>\n #include <assert.h>\n-#include <regex.h>\n+#include <pcreposix.h>\n #include <netinet/in.h>\n #include <netinet/tcp.h>\n #include <arpa/inet.h>\n\n\nA few numbers, all from a fully packed kernel repository:\n\n# glibc, trivial regex\n$ /usr/bin/time git grep --cached foo >/dev/null\n10.07user 0.15system 0:10.23elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+36617minor)pagefaults 0swaps\n\n# glibc, complex regex\n$ /usr/bin/time git grep --cached '[a-z][0-9][a-z][0-9][a-z]'  >/dev/null\n24.42user 0.15system 0:24.60elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+36210minor)pagefaults 0swaps\n\n# pcre, trivial regex\n$ /usr/bin/time git grep --cached foo >/dev/null\n7.82user 0.12system 0:08.00elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+36571minor)pagefaults 0swaps\n\n# pcre, complex regex\n$ /usr/bin/time git grep --cached '[a-z][0-9][a-z][0-9][a-z]'  >/dev/null\n36.51user 0.13system 0:36.65elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+36583minor)pagefaults 0swaps\n\n\nSo the winner seems to vary based on the complexity of the pattern.\nThere are some less rudimentary but non-git performance tests here:\n\n  http://www.boost.org/libs/regex/doc/gcc-performance.html\n\nIn every case there, pcre has either comparable performance, or simply\nblows away glibc.\n\nOne final note that caused some confusion during my testing: git-grep\nstill uses external grep for working tree greps (i.e., 'git grep foo').\nThis meant that 'git grep' and 'git grep --cached' produced wildly\ndifferent results once I was using pcre internally. Something to look\nout for if we switch to pcre (or any other library which doesn't exactly\nmatch external grep behavior!).\n\n-Peff\n"},{"id":"33949","messageId":"Pine.LNX.4.63.0702081905570.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6723","inReplyTo":"20070208061654.GA8813@coredump.intra.peff.net","subject":"Re: git log filtering","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-08T18:06:25Z","receivedAt":"2007-02-08T18:06:25Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 8 Feb 2007, Jeff King wrote:\n\n> On Wed, Feb 07, 2007 at 01:53:18PM -0800, Linus Torvalds wrote:\n> \n> > What's PCRE performance like? I'd hate to make \"git grep\" slower, and it \n> > would be stupid and confusing to use two different regex libraries..\n> >\n> > Maybe somebody could test - afaik, PCRE has a regex-compatible (from a API \n> > standpoint, not from a regex standpoint!) wrapper thing, and it might be \n> > interesting to hear if doing \"git grep\" is slower or faster..\n> \n> The patch is delightfully simple (though a real patch would probably be\n> conditional):\n>\n> [...]\n\nMay I register a complaint? This is yet _another_ dependency.\n\nCiao,\nDscho\n"},{"id":"33992","messageId":"20070208223336.GA9422@coredump.intra.peff.net","threadId":"6723","inReplyTo":"Pine.LNX.4.63.0702081905570.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: git log filtering","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-02-08T22:33:37Z","receivedAt":"2007-02-08T22:33:37Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 08, 2007 at 07:06:25PM +0100, Johannes Schindelin wrote:\n\n> May I register a complaint? This is yet _another_ dependency.\n\nUnlike other dependencies, I think it's quite natural to make it a\nconditional dependency. If you have pcre, you get more featureful\nregular expressions. If you don't, you get posix regular expressions.\nDo you object to a few extra lines in the Makefile?\n\n-Peff\n"},{"id":"33994","messageId":"7v7iusz3c2.fsf@assigned-by-dhcp.cox.net","threadId":"6723","inReplyTo":"Pine.LNX.4.64.0702071139090.8424@woody.linux-foundation.org","subject":"Re: Fix \"git log -z\" behaviour","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-08T22:34:05Z","receivedAt":"2007-02-08T22:34:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> For the normal case where the termination character is '\\n', this \n> obviously doesn't change anything at all, since we just switched two \n> identical characters around. So it's very safe - it doesn't change any \n> normal usage, but it definitely fixes \"git log -z\".\n\nGaah.\n\nI have already applied this but I think this has fallout for\nexisting users of \"-z --raw\".  Nothing in-tree uses \"git log\" as\nthe upstream of a pipe as far as I know because in-tree stuff\ntend to stick to plumbing when it comes to scripting, but I\nthink your patch would affect the plumbing level as well.\n\nScripts that read from \"-z --raw\" have been expecting to get a\nrecord whose first 7 bytes are \"commit \" to be a log, which is\nfollowed by an arbitrary number of records whose first byte is\n\":\" (and then it needs variable number of records to complete\none diff record).  This patch removes the separator NUL between\nthe log message and the first diff record.\n"},{"id":"34017","messageId":"Pine.LNX.4.63.0702090115180.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6723","inReplyTo":"20070208223336.GA9422@coredump.intra.peff.net","subject":"Re: git log filtering","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-09T00:18:01Z","receivedAt":"2007-02-09T00:18:01Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 8 Feb 2007, Jeff King wrote:\n\n> On Thu, Feb 08, 2007 at 07:06:25PM +0100, Johannes Schindelin wrote:\n> \n> > May I register a complaint? This is yet _another_ dependency.\n> \n> Unlike other dependencies, I think it's quite natural to make it a\n> conditional dependency. If you have pcre, you get more featureful\n> regular expressions. If you don't, you get posix regular expressions.\n> Do you object to a few extra lines in the Makefile?\n\nYes, I do. Not because of the extra lines, but because of the inconsistent \ninterface.\n\nWe included libxdiff _exactly_ to ensure consistency between different git \ninstallations (remember, diff behaves quite differently on different \nplatforms, and even GNU diff behaves differently depending on which \nversion you use).\n\nSo no, I do not like the idea of using git on some random box, only to \nrealize that what I have grown used to does not work.\n\nCiao,\nDscho\n"},{"id":"34020","messageId":"20070209002344.GF1556@spearce.org","threadId":"6723","inReplyTo":"Pine.LNX.4.63.0702090115180.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: git log filtering","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-02-09T00:23:44Z","receivedAt":"2007-02-09T00:23:44Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> We included libxdiff _exactly_ to ensure consistency between different git \n> installations (remember, diff behaves quite differently on different \n> platforms, and even GNU diff behaves differently depending on which \n> version you use).\n\npcre is covered by the BSD license.  Can we ship it with git, like\nwe ship libxdiff?  I want to say Apache ships with pcre, but they\nuse the Apache License so it might be easier for them to do so.\n\n-- \nShawn.\n"},{"id":"34024","messageId":"Pine.LNX.4.63.0702090144310.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6723","inReplyTo":"20070209002344.GF1556@spearce.org","subject":"Re: git log filtering","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-09T00:45:09Z","receivedAt":"2007-02-09T00:45:09Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 8 Feb 2007, Shawn O. Pearce wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > We included libxdiff _exactly_ to ensure consistency between different \n> > git installations (remember, diff behaves quite differently on \n> > different platforms, and even GNU diff behaves differently depending \n> > on which version you use).\n> \n> pcre is covered by the BSD license.  Can we ship it with git, like we \n> ship libxdiff?  I want to say Apache ships with pcre, but they use the \n> Apache License so it might be easier for them to do so.\n\nIf we bundle it like we do with libxdiff, I do not have any objections. It \nwould also help MinGW.\n\nCiao,\nDscho\n"},{"id":"34029","messageId":"20070209015925.GD10574@coredump.intra.peff.net","threadId":"6723","inReplyTo":"Pine.LNX.4.63.0702090115180.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: git log filtering","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-02-09T01:59:25Z","receivedAt":"2007-02-09T01:59:25Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 09, 2007 at 01:18:01AM +0100, Johannes Schindelin wrote:\n\n> Yes, I do. Not because of the extra lines, but because of the inconsistent \n> interface.\n\nOK, so we may either:\n  1. always use the lowest common denominator (i.e., no pcre support)\n  2. force a dependency for new features (i.e., require pcre)\n  3. have inconsistency between builds (i.e., conditional dependency)\n  4. include all dependencies, or re-write them natively\n\nI agree that 4 can make some sense in limited situations, but I worry\nthat it will eventually cease to be scalable (we don't get improvements\nor bugfixes automatically from other packages, we potentially re-invent\nthe wheel). We already have '3' for other things: openssl, curl, expat,\neven perl.\n\n-Peff\n"},{"id":"34053","messageId":"20070209131522.2df5d0b2.vsu@altlinux.ru","threadId":"6723","inReplyTo":"20070209002344.GF1556@spearce.org","subject":"Re: git log filtering","fromName":"Sergey Vlasov","fromEmail":"vsu@altlinux.ru","sentAt":"2007-02-09T10:15:22Z","receivedAt":"2007-02-09T10:15:22Z","isPatch":false,"sender":{"key":"vsu@altlinux.ru","avatar":"https://avatars.githubusercontent.com/u/616082?v=4"},"body":"On Thu, 8 Feb 2007 19:23:44 -0500 Shawn O. Pearce wrote:\n\n> pcre is covered by the BSD license.  Can we ship it with git, like\n> we ship libxdiff?  I want to say Apache ships with pcre, but they\n> use the Apache License so it might be easier for them to do so.\n\nIf you do this, please do not forget to add a way to use the system\ncopy of libpcre instead of the bundled version.\n"},{"id":"34059","messageId":"Pine.LNX.4.63.0702091410230.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6723","inReplyTo":"20070209015925.GD10574@coredump.intra.peff.net","subject":"Re: git log filtering","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-09T13:13:18Z","receivedAt":"2007-02-09T13:13:18Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 8 Feb 2007, Jeff King wrote:\n\n> On Fri, Feb 09, 2007 at 01:18:01AM +0100, Johannes Schindelin wrote:\n> \n> > Yes, I do. Not because of the extra lines, but because of the inconsistent \n> > interface.\n> \n> OK, so we may either:\n>   1. always use the lowest common denominator (i.e., no pcre support)\n>   2. force a dependency for new features (i.e., require pcre)\n>   3. have inconsistency between builds (i.e., conditional dependency)\n>   4. include all dependencies, or re-write them natively\n> \n> I agree that 4 can make some sense in limited situations, but I worry\n> that it will eventually cease to be scalable (we don't get improvements\n> or bugfixes automatically from other packages, we potentially re-invent\n> the wheel). We already have '3' for other things: openssl, curl, expat,\n> even perl.\n\nThe difference, of course, is that with the \"other things\", we either have \nno alternative (if you do not have curl, you cannot use HTTP transport), \nor we have workalikes (if you don't use openssl, the (possibly slower) \nSHA1 replacements take effect).\n\nWe _used_ to rely on external \"diff\" and \"merge\", but have them as inbuilt \ncomponents, exactly to avoid \"if you have a slightly differing setup, \ngit behaves differently\".\n\nCiao,\nDscho\n"},{"id":"34060","messageId":"20070209132239.GA727@coredump.intra.peff.net","threadId":"6723","inReplyTo":"Pine.LNX.4.63.0702091410230.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: git log filtering","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2007-02-09T13:22:39Z","receivedAt":"2007-02-09T13:22:39Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 09, 2007 at 02:13:18PM +0100, Johannes Schindelin wrote:\n\n> The difference, of course, is that with the \"other things\", we either have \n> no alternative (if you do not have curl, you cannot use HTTP transport), \n> or we have workalikes (if you don't use openssl, the (possibly slower) \n> SHA1 replacements take effect).\n\nI'm not a pcre expert, but I thought most of the additions to posix\nextended regular expressions were expressed through constructs that\nwould otherwise be invalid patterns. For example, '(?i)' doesn't make\nany sense as a pattern. Thus you would only see different behavior when\ninputting nonsense. Of course, we're not currently using extended\nregexps, but that could be made the default without additional\ndependencies.\n\n> We _used_ to rely on external \"diff\" and \"merge\", but have them as inbuilt \n> components, exactly to avoid \"if you have a slightly differing setup, \n> git behaves differently\".\n\nBut you're OK with \"if you didn't built against curl, http transport\njust doesn't work.\" So what if there is a '--pcre' option and a\ncorresponding config option? Thus you get the same results always,\nunless you use --pcre and it's not built, in which case git dies. That\nseems to be the moral equivalent of the curl situation.\n\n\nAt any rate, you didn't address my original point, which is _all_ of\nthose options have drawbacks. I think the drawbacks of re-writing or\nre-packaging a regular expression library outweigh those of adding the\ndependency (or even having slightly irregular behavior).\n\n-Peff\n"},{"id":"34062","messageId":"Pine.LNX.4.63.0702091557420.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6723","inReplyTo":"20070209132239.GA727@coredump.intra.peff.net","subject":"Re: git log filtering","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-02-09T15:02:24Z","receivedAt":"2007-02-09T15:02:24Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 9 Feb 2007, Jeff King wrote:\n\n> On Fri, Feb 09, 2007 at 02:13:18PM +0100, Johannes Schindelin wrote:\n> \n> > The difference, of course, is that with the \"other things\", we either have \n> > no alternative (if you do not have curl, you cannot use HTTP transport), \n> > or we have workalikes (if you don't use openssl, the (possibly slower) \n> > SHA1 replacements take effect).\n> \n> I'm not a pcre expert, but I thought most of the additions to posix\n> extended regular expressions were expressed through constructs that\n> would otherwise be invalid patterns.\n\nSo, once pcre is used, you can use these constructs. Even in scripts. \nWhich just so happen to break on platforms where git is not compiled with \npcre support.\n\nOr do you suggest checking (in git!) if the pattern is a pcre special or \nnot? That would be insane.\n\n> > We _used_ to rely on external \"diff\" and \"merge\", but have them as \n> > inbuilt components, exactly to avoid \"if you have a slightly differing \n> > setup, git behaves differently\".\n> \n> But you're OK with \"if you didn't built against curl, http transport \n> just doesn't work.\"\n\nYes, I am. Since HTTP is itself only a second-class citizen.\n\n> So what if there is a '--pcre' option and a corresponding config option? \n> Thus you get the same results always, unless you use --pcre and it's not \n> built, in which case git dies. That seems to be the moral equivalent of \n> the curl situation.\n\nI might be wrong, but most of git does not depend on HTTP.\n\n> At any rate, you didn't address my original point, which is _all_ of \n> those options have drawbacks. I think the drawbacks of re-writing or \n> re-packaging a regular expression library outweigh those of adding the \n> dependency (or even having slightly irregular behavior).\n\nThis is only because you do not really have problems with dependencies. \nYou just install, or compile, the dependent thing, which happens to be no \nhassle, since you use Linux. And you can compile & install things.\n\nOnce everybody runs Linux, and is allowed to compile & install things, I \nwill no longer complain about trillions of dependencies.\n\nCiao,\nDscho\n"},{"id":"34088","messageId":"7vtzxumps5.fsf@assigned-by-dhcp.cox.net","threadId":"6723","inReplyTo":"7v7iusz3c2.fsf@assigned-by-dhcp.cox.net","subject":"Re: Fix \"git log -z\" behaviour","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-10T07:32:10Z","receivedAt":"2007-02-10T07:32:10Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>\n>> For the normal case where the termination character is '\\n', this \n>> obviously doesn't change anything at all, since we just switched two \n>> identical characters around. So it's very safe - it doesn't change any \n>> normal usage, but it definitely fixes \"git log -z\".\n>\n> Gaah.\n>\n> I have already applied this but I think this has fallout for\n> existing users of \"-z --raw\".  Nothing in-tree uses \"git log\" as\n> the upstream of a pipe as far as I know because in-tree stuff\n> tend to stick to plumbing when it comes to scripting, but I\n> think your patch would affect the plumbing level as well.\n\nI think the new semantics for -z (\"inter-record termination is\nNUL\") makes a lot more sense for \"-p -z\" format that shows\ncommit log message and the patch text.  It makes filtering the\noutput with \"grep -z\" feel much more natural.\n\nThe new semantics is however quite inconsistent with the other\nformats: --raw, --name-only and --name-status.  These already\nuse NUL for separating pathnames and fields when -z is given, in\norder to allow scripts sensibly deal with pathname that contain\nfunny characters (e.g. LF and HT).  Nobody is likely to feed\ntheir output to \"grep -z\", but one problematic case I see is to\nuse this:\n\n\tgit log -z --raw -r --pretty=raw $commit\n\nor its equivalent:\n\n\tgit rev-list $commit |\n        git diff-tree --stdin --raw -r --pretty=raw\n\nto prepare data to feed something like fast-import.\n\nBut such newly written scripts can read from non -z and unwrap\npaths themselves just as easily (the pathname safety with NUL\nwas invented before we started using c-quote consistently), so\nit might be Ok to leave them (slightly) broken.\n\nSo, I give up.\n"},{"id":"34093","messageId":"7vlkj6mk0q.fsf@assigned-by-dhcp.cox.net","threadId":"6723","inReplyTo":"7vtzxumps5.fsf@assigned-by-dhcp.cox.net","subject":"Re: Fix \"git log -z\" behaviour","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-02-10T09:36:37Z","receivedAt":"2007-02-10T09:36:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> Junio C Hamano <junkio@cox.net> writes:\n>\n>> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>>\n>>> For the normal case where the termination character is '\\n', this \n>>> obviously doesn't change anything at all, since we just switched two \n>>> identical characters around. So it's very safe - it doesn't change any \n>>> normal usage, but it definitely fixes \"git log -z\".\n>>\n>> Gaah.\n>>\n>> I have already applied this but I think this has fallout for\n>> existing users of \"-z --raw\".  Nothing in-tree uses \"git log\" as\n>> the upstream of a pipe as far as I know because in-tree stuff\n>> tend to stick to plumbing when it comes to scripting, but I\n>> think your patch would affect the plumbing level as well.\n>\n> I think the new semantics for -z (\"inter-record termination is\n> NUL\") makes a lot more sense for \"-p -z\" format that shows\n> commit log message and the patch text.  It makes filtering the\n> output with \"grep -z\" feel much more natural.\n>\n> The new semantics is however quite inconsistent with the other\n> formats: --raw, --name-only and --name-status.  These already\n> use NUL for separating pathnames and fields when -z is given, in\n> order to allow scripts sensibly deal with pathname that contain\n> funny characters (e.g. LF and HT).  Nobody is likely to feed\n> their output to \"grep -z\", but one problematic case I see is to\n> use this:\n>\n> \tgit log -z --raw -r --pretty=raw $commit\n>\n> or its equivalent:\n>\n> \tgit rev-list $commit |\n>         git diff-tree --stdin --raw -r --pretty=raw\n>\n> to prepare data to feed something like fast-import.\n>\n> But such newly written scripts can read from non -z and unwrap\n> paths themselves just as easily (the pathname safety with NUL\n> was invented before we started using c-quote consistently), so\n> it might be Ok to leave them (slightly) broken.\n>\n> So, I give up.\n\n... well, it just occured to me that it might make sense not to\nlet this new \"use NUL as inter-commit separator for grep -z\"\nsemantics hijack existing -z option, but introduce another\noption, say, -Z.  Then you could even do something like:\n\n\tgit log -Z -r --numstat |\n        grep -z -e '^[1-9][0-9][0-9][0-9]*\t'\n\nto find commits that has more than 100 lines of additions to a\nfile.  (or use --stat and grep for '| *[1-9][0-9][0-9][0-9]* ' to\nlook for sum of addition+deletion ).\n\nHmmmm.\n"},{"id":"34125","messageId":"Pine.LNX.4.64.0702100902250.8424@woody.linux-foundation.org","threadId":"6723","inReplyTo":"7vlkj6mk0q.fsf@assigned-by-dhcp.cox.net","subject":"Re: Fix \"git log -z\" behaviour","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-02-10T17:09:44Z","receivedAt":"2007-02-10T17:09:44Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 10 Feb 2007, Junio C Hamano wrote:\n> \n> ... well, it just occured to me that it might make sense not to\n> let this new \"use NUL as inter-commit separator for grep -z\"\n> semantics hijack existing -z option, but introduce another\n> option, say, -Z.\n\nI don't think I disagree, but I do suspect it's not worth it.\n\nYes, we really do have two \"line_termination\" characters: the one between \ncommits, and the one we use within raw diffs. However, I don't think the \n*combination* ever makes sense any more (*), so using the same flag \ndoesn't seem to really be a problem.\n\nAnd the -z \"line_termination\" already got hijacked a long time ago for \ninter-commit messages too, so while adding a \"-Z\" would perhaps avoid a \ncertain ambiguity, it would actually potentially break stuff that just did\n\n\tgit-rev-list -z --pretty .. | ...\n\nwhich is actually _more_ likely than the \"multiple commit messages _and_ \nraw outpu _and_ '-z'\" combination.\n\nSo I would suggest leaving it as-is, especially since I don't think \nanybody has actually even noticed (ie nobody probably used that \ncombination), and the new semantics in many ways are both more useful and \nmore logical.\n\n\t\tLinus\n\n(*) It may well have made sense a year and a half ago, I don't think it \nmakes much sense any more.\n"},{"id":"36561","messageId":"Pine.LNX.4.63.0703071807250.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6723","inReplyTo":"20070208061654.GA8813@coredump.intra.peff.net","subject":"pcre performance, was Re: git log filtering","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-03-07T17:37:23Z","receivedAt":"2007-03-07T17:37:23Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 8 Feb 2007, Jeff King wrote:\n\n> In every case there, pcre has either comparable performance, or simply \n> blows away glibc.\n\nSo I tested this against external grep. For completeness' sake, I tested \nthese against each other: GNU regex-0.12, Git _without_ external grep \n(relies on glibc's regex), Git _with_ external grep (\"original\"), pcre, \nand for good measure, pcre with NO_MMAP=1 (to test if disk access is the \nproblem).\n\nHere are the numbers:\n\ngrep-gnu-regex:\n\n21.41user 1.08system 0:22.52elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7210minor)pagefaults 0swaps\n21.40user 1.06system 0:22.47elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7209minor)pagefaults 0swaps\n21.61user 1.06system 0:22.68elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7209minor)pagefaults 0swaps\n21.30user 1.10system 0:22.48elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7210minor)pagefaults 0swaps\n21.30user 1.08system 0:22.43elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7209minor)pagefaults 0swaps\n\ngrep-no-external-grep:\n\n6.98user 1.17system 0:08.16elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7120minor)pagefaults 0swaps\n7.07user 1.16system 0:08.27elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7121minor)pagefaults 0swaps\n6.98user 1.12system 0:08.11elapsed 100%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7121minor)pagefaults 0swaps\n7.00user 1.18system 0:08.20elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7121minor)pagefaults 0swaps\n\ngrep-original:\n\n0.82user 1.15system 0:01.97elapsed 100%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7090minor)pagefaults 0swaps\n0.94user 1.03system 0:01.97elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7099minor)pagefaults 0swaps\n0.89user 1.07system 0:01.96elapsed 100%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7092minor)pagefaults 0swaps\n0.81user 1.15system 0:01.97elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7092minor)pagefaults 0swaps\n\ngrep-pcre:\n\n4.04user 1.18system 0:05.24elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7205minor)pagefaults 0swaps\n4.16user 1.08system 0:05.25elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7206minor)pagefaults 0swaps\n4.24user 0.98system 0:05.23elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7206minor)pagefaults 0swaps\n4.08user 1.14system 0:05.23elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7206minor)pagefaults 0swaps\n\ngrep-pcre-no-mmap:\n\n4.15user 1.07system 0:05.22elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7210minor)pagefaults 0swaps\n4.01user 1.14system 0:05.17elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7209minor)pagefaults 0swaps\n3.94user 1.18system 0:05.14elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7210minor)pagefaults 0swaps\n4.11user 1.06system 0:05.18elapsed 99%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+7210minor)pagefaults 0swaps\n\nBTW this was \"git grep Lin.*valds\" on linux-2.6, just updated.\n\nThe first test was run 5 times instead of 4 to make sure it is hot cache. \nThis is on a dual 1.2GHz 2GB machine.\n\nI cannot really say anything about the pagefaults, so I'll leave that to \nthe wizards.\n\nResult: external grep wins hands-down. GNU regex loses hands-down. pcre \nseems to be better than glibc's regex engine, and gains ever so slightly \nwhen using NO_MMAP.\n\nI ran the same test on a 1GHz 256MB machine which is overloaded, and in \nthat case, GNU regex is still worst (~55 sec), while glibc and pcre are \nequal (glibc slightly slower with ~35 sec, pcre ~34 sec), and external \ngrep wins (~29 sec). Of course, this is io-bound, but it shows that pcre \nuses more memory than glibc.\n\nCiao,\nDscho\n"},{"id":"36564","messageId":"45EEFE74.1090309@lu.unisi.ch","threadId":"6723","inReplyTo":"Pine.LNX.4.63.0703071807250.22628@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: pcre performance, was Re: git log filtering","fromName":"Paolo Bonzini","fromEmail":"paolo.bonzini@lu.unisi.ch","sentAt":"2007-03-07T18:03:32Z","receivedAt":"2007-03-07T18:03:32Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n> Result: external grep wins hands-down. GNU regex loses hands-down. pcre \n> seems to be better than glibc's regex engine, and gains ever so slightly \n> when using NO_MMAP.\n\nIndeed GNU regex 0.12 loses, and that's why it was rewritten for (IIRC)\nglibc 2.3.  Older glibc's use code derived from GNU regex 0.12; but the\nold GNU regex code is dead in general (maybe it survives in Emacs -- but\nI don't remember), and the glibc regex code can be used by external\nprograms via gnulib.\n\nglibc is slower than PCRE mostly because it is internationalized.  So\nfor example it supports things like stra[.ss.]e matching both strasse\nand straße in a German locale, or [[=a=]] matching aàáäâ and possibly\nmore variations.  In theory.  In practice I couldn't make it work\nwhile writing this message...\n\nExternal grep wins hands-down because it's a DFA engine.  If the regex\nuses backreferences (or the above esoteric constructs), however, external\ngrep will not be able to give a definite answer using the fast engine,\nand will fall back to glibc regex.\n\nPaolo\n"}]}