{"thread":{"id":"17516","subject":"webgit highlightes mem adresses as git versions","startedAt":"2009-02-02T21:04:46Z","lastAt":"2009-02-07T14:01:31Z","messageCount":22,"participants":["Toralf Förster","Jakub Narebski","Johannes Schindelin","Rafael Garcia-Suarez","Jay Soffian","Junio C Hamano","demerphq"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"102897","messageId":"200902022204.46651.toralf.foerster@gmx.de","threadId":"17516","inReplyTo":null,"subject":"webgit highlightes mem adresses as git versions","fromName":"Toralf Förster","fromEmail":"toralf.foerster@gmx.de","sentAt":"2009-02-02T21:04:46Z","receivedAt":"2009-02-02T21:04:46Z","isPatch":false,"sender":{"key":"toralf.foerster@gmx.de","avatar":null},"body":"As seen here \nhttp://git.kernel.org/?p=linux/kernel/git/stable/linux-2.6.27.y.git;a=commit;h=8ca2918f99b5861359de1805f27b08023c82abd2 \nthe strings [<c043d0f3>] and firends shouldn't be recognized as git hashes, \nisn't it ?\n\n-- \nMfG/Sincerely\n\nToralf Förster\npgp finger print: 7B1A 07F4 EC82 0F90 D4C2 8936 872A E508 7DB6 9DA3\n"},{"id":"102906","messageId":"m3ljsowisv.fsf@localhost.localdomain","threadId":"17516","inReplyTo":"200902022204.46651.toralf.foerster@gmx.de","subject":"Re: webgit highlightes mem adresses as git versions","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-02T22:54:20Z","receivedAt":"2009-02-02T22:54:20Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Toralf Förster <toralf.foerster@gmx.de> writes:\n\n> As seen here \n> http://git.kernel.org/?p=linux/kernel/git/stable/linux-2.6.27.y.git;a=commit;h=8ca2918f99b5861359de1805f27b08023c82abd2 \n> the strings [<c043d0f3>] and firends shouldn't be recognized as git hashes, \n> isn't it ?\n\nGitweb, not webgit.  And gitweb considers ([0-9a-fA-F]{8,40}) i.e.\nfrom 8 to 40 hexadecimal characters to be (shortened) SHA-1.  It\nsimply cannot afford checking if such object exists when displaying\ncommit message (for example in 'log' view).\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"102950","messageId":"200902031204.21435.toralf.foerster@gmx.de","threadId":"17516","inReplyTo":"m3ljsowisv.fsf@localhost.localdomain","subject":"Re: webgit highlightes mem adresses as git versions","fromName":"Toralf Förster","fromEmail":"toralf.foerster@gmx.de","sentAt":"2009-02-03T11:04:21Z","receivedAt":"2009-02-03T11:04:21Z","isPatch":false,"sender":{"key":"toralf.foerster@gmx.de","avatar":null},"body":"At Monday 02 February 2009 23:54:20 Jakub Narebski wrote :\n> Toralf Förster <toralf.foerster@gmx.de> writes:\n> > As seen here\n> > http://git.kernel.org/?p=linux/kernel/git/stable/linux-2.6.27.y.git;a=com\n> >mit;h=8ca2918f99b5861359de1805f27b08023c82abd2 the strings [<c043d0f3>]\n> > and firends shouldn't be recognized as git hashes, isn't it ?\n>\n> Gitweb, not webgit.  And gitweb considers ([0-9a-fA-F]{8,40}) i.e.\n> from 8 to 40 hexadecimal characters to be (shortened) SHA-1.  It\n> simply cannot afford checking if such object exists when displaying\n> commit message (for example in 'log' view).\n\nAh - ok, what's about expecting spaces around such SHA-1 keys ?\n\n-- \nMfG/Sincerely\n\nToralf Förster\npgp finger print: 7B1A 07F4 EC82 0F90 D4C2 8936 872A E508 7DB6 9DA3\n"},{"id":"102953","messageId":"alpine.DEB.1.00.0902031327340.6573@intel-tinevez-2-302","threadId":"17516","inReplyTo":"200902031204.21435.toralf.foerster@gmx.de","subject":"Re: webgit highlightes mem adresses as git versions","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-02-03T12:31:44Z","receivedAt":"2009-02-03T12:31:44Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 3 Feb 2009, Toralf Förster wrote:\n\n> At Monday 02 February 2009 23:54:20 Jakub Narebski wrote :\n> > Toralf Förster <toralf.foerster@gmx.de> writes:\n> > > As seen here\n> > > http://git.kernel.org/?p=linux/kernel/git/stable/linux-2.6.27.y.git;a=com\n> > >mit;h=8ca2918f99b5861359de1805f27b08023c82abd2 the strings [<c043d0f3>]\n> > > and firends shouldn't be recognized as git hashes, isn't it ?\n> >\n> > Gitweb, not webgit.  And gitweb considers ([0-9a-fA-F]{8,40}) i.e.\n> > from 8 to 40 hexadecimal characters to be (shortened) SHA-1.  It\n> > simply cannot afford checking if such object exists when displaying\n> > commit message (for example in 'log' view).\n> \n> Ah - ok, what's about expecting spaces around such SHA-1 keys ?\n\nWon't fly: there was a recommendation at some point that you should refer \nto commits in such a form:\n\n\t2819075(Merge branch 'maint-1.6.0' into maint)\n\nHowever, gitweb being written in Perl, I think a lookbehind like (?<!0x), \ni.e. that a 0x in front of the hexadecimal characters means it is no \nSHA-1.\n\nEven better would be using word boundaries, I guess, but all that fails \nwhen you have a hexdump in the commit message.\n\nCiao,\nDscho\n"},{"id":"103435","messageId":"200902061012.42943.jnareb@gmail.com","threadId":"17516","inReplyTo":"alpine.DEB.1.00.0902031327340.6573@intel-tinevez-2-302","subject":"[PATCH] gitweb: Better regexp for SHA-1 committag match","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-06T09:12:41Z","receivedAt":"2009-02-06T09:12:41Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, 3 Feb 2009, Johannes Schindelin wrote:\n> On Tue, 3 Feb 2009, Toralf Förster wrote:\n>> At Monday 02 February 2009 23:54:20 Jakub Narebski wrote :\n>>> Toralf Förster <toralf.foerster@gmx.de> writes:\n\n>>>> As seen here\n>>>> http://git.kernel.org/?p=linux/kernel/git/stable/linux-2.6.27.y.git;a=commit;h=8ca2918f99b5861359de1805f27b08023c82abd2 the strings [<c043d0f3>]\n>>>> and firends shouldn't be recognized as git hashes, isn't it ?\n>>>\n>>> Gitweb, not webgit.  And gitweb considers ([0-9a-fA-F]{8,40}) i.e.\n>>> from 8 to 40 hexadecimal characters to be (shortened) SHA-1.  It\n>>> simply cannot afford checking if such object exists when displaying\n>>> commit message (for example in 'log' view).\n>> \n>> Ah - ok, what's about expecting spaces around such SHA-1 keys ?\n> \n> Won't fly: there was a recommendation at some point that you should refer \n> to commits in such a form:\n> \n> \t2819075(Merge branch 'maint-1.6.0' into maint)\n> \n> However, gitweb being written in Perl, I think a lookbehind like (?<!0x), \n> i.e. that a 0x in front of the hexadecimal characters means it is no \n> SHA-1.\n> \n> Even better would be using word boundaries, I guess, but all that fails \n> when you have a hexdump in the commit message.\n\nHere you have it: anchoring SHA-1 regexp to word boundary. It would\nhelp eliminate _some_ of false matches.\n\n-- >8 --\nSubject: [PATCH] gitweb: Better regexp for SHA-1 committag match\n\nMake SHA-1 regexp to be turned into hyperlink (the SHA-1 committag)\nto match word boundary at the beginning and the end.  This way we\nreduce number of false matches, for example we now don't match\n0x74a5cd01 which is hex decimal (for example memory address),\nbut is not SHA-1.\n\nSuggested-by: Johannes Schindelin <Johannes.Schindelin@gmx.de>\nSigned-off-by: Jakub Narebski <jnareb@gmail.com>\n---\n gitweb/gitweb.perl |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex f27dbb6..bec1af6 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -1364,7 +1364,7 @@ sub format_log_line_html {\n \tmy $line = shift;\n \n \t$line = esc_html($line, -nbsp=>1);\n-\tif ($line =~ m/([0-9a-fA-F]{8,40})/) {\n+\tif ($line =~ m/\\b([0-9a-fA-F]{8,40})\\b/) {\n \t\tmy $hash_text = $1;\n \t\tmy $link =\n \t\t\t$cgi->a({-href => href(action=>\"object\", hash=>$hash_text),\n-- \n1.6.1\n"},{"id":"103440","messageId":"b77c1dce0902060149j25e76250q76c2368bd3ca5857@mail.gmail.com","threadId":"17516","inReplyTo":"200902061012.42943.jnareb@gmail.com","subject":"Re: [PATCH] gitweb: Better regexp for SHA-1 committag match","fromName":"Rafael Garcia-Suarez","fromEmail":"rgarciasuarez@gmail.com","sentAt":"2009-02-06T09:49:57Z","receivedAt":"2009-02-06T09:49:57Z","isPatch":true,"sender":{"key":"rgarciasuarez@gmail.com","avatar":null},"body":"2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n> Make SHA-1 regexp to be turned into hyperlink (the SHA-1 committag)\n> to match word boundary at the beginning and the end.  This way we\n> reduce number of false matches, for example we now don't match\n> 0x74a5cd01 which is hex decimal (for example memory address),\n> but is not SHA-1.\n\nFurther suggestion: you could also turn the final \\b into (\\b|\\@), so\nit skips stuff that might look like a message-id.\nHere's an example :\nhttp://perl5.git.perl.org/perl.git/commit/f57255841c18e91c7a719a2400645e39398f3947\n(We get loads of this in the Perl repository)\n\n> Suggested-by: Johannes Schindelin <Johannes.Schindelin@gmx.de>\n> Signed-off-by: Jakub Narebski <jnareb@gmail.com>\n> ---\n>  gitweb/gitweb.perl |    2 +-\n>  1 files changed, 1 insertions(+), 1 deletions(-)\n>\n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index f27dbb6..bec1af6 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -1364,7 +1364,7 @@ sub format_log_line_html {\n>        my $line = shift;\n>\n>        $line = esc_html($line, -nbsp=>1);\n> -       if ($line =~ m/([0-9a-fA-F]{8,40})/) {\n> +       if ($line =~ m/\\b([0-9a-fA-F]{8,40})\\b/) {\n>                my $hash_text = $1;\n>                my $link =\n>                        $cgi->a({-href => href(action=>\"object\", hash=>$hash_text),\n> --\n> 1.6.1\n>\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n\n-- \n\"You don't mean odds and ends, you mean des curieux et des bouts\",\ncorrected the manager.\n-- Terry Pratchett, Hogfather\n"},{"id":"103442","messageId":"200902061126.18418.jnareb@gmail.com","threadId":"17516","inReplyTo":"b77c1dce0902060149j25e76250q76c2368bd3ca5857@mail.gmail.com","subject":"Re: [PATCH] gitweb: Better regexp for SHA-1 committag match","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-06T10:26:16Z","receivedAt":"2009-02-06T10:26:16Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Dnia piątek 6. lutego 2009 10:49, Rafael Garcia-Suarez napisał:\n> 2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n\n> > Make SHA-1 regexp to be turned into hyperlink (the SHA-1 committag)\n> > to match word boundary at the beginning and the end.  This way we\n> > reduce number of false matches, for example we now don't match\n> > 0x74a5cd01 which is hex decimal (for example memory address),\n> > but is not SHA-1.\n> \n> Further suggestion: you could also turn the final \\b into (\\b|\\@),\n\nYou meant \\b -> \\b(?!\\@), didn't you?  Word boundary _not_ followed\nby '@', and not word boundary _OR_ '@' as you wrote...\n\n> so it skips stuff that might look like a message-id.\n> Here's an example :\n> http://perl5.git.perl.org/perl.git/commit/f57255841c18e91c7a719a2400645e39398f3947\n\nFor those who do not want to open browser, it is:\n\nMessage-ID: <46A0F33545E63740BC7563DE59CA9C6D0939A0@exchsvr2.npl.ad.local>\n\n> (We get loads of this in the Perl repository)\n\n> > -       if ($line =~ m/([0-9a-fA-F]{8,40})/) {\n> > +       if ($line =~ m/\\b([0-9a-fA-F]{8,40})\\b/) {\n\n-- \nJakub Narebski\nPoland\n"},{"id":"103444","messageId":"b77c1dce0902060231u358587d5o940eb322fde52a68@mail.gmail.com","threadId":"17516","inReplyTo":"200902061126.18418.jnareb@gmail.com","subject":"Re: [PATCH] gitweb: Better regexp for SHA-1 committag match","fromName":"Rafael Garcia-Suarez","fromEmail":"rgarciasuarez@gmail.com","sentAt":"2009-02-06T10:31:39Z","receivedAt":"2009-02-06T10:31:39Z","isPatch":true,"sender":{"key":"rgarciasuarez@gmail.com","avatar":null},"body":"2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n> Dnia piątek 6. lutego 2009 10:49, Rafael Garcia-Suarez napisał:\n>> 2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n>\n>> > Make SHA-1 regexp to be turned into hyperlink (the SHA-1 committag)\n>> > to match word boundary at the beginning and the end.  This way we\n>> > reduce number of false matches, for example we now don't match\n>> > 0x74a5cd01 which is hex decimal (for example memory address),\n>> > but is not SHA-1.\n>>\n>> Further suggestion: you could also turn the final \\b into (\\b|\\@),\n>\n> You meant \\b -> \\b(?!\\@), didn't you?  Word boundary _not_ followed\n> by '@', and not word boundary _OR_ '@' as you wrote...\n\nAh right, shame on me.\n"},{"id":"103446","messageId":"200902061149.16210.jnareb@gmail.com","threadId":"17516","inReplyTo":"b77c1dce0902060231u358587d5o940eb322fde52a68@mail.gmail.com","subject":"[PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-06T10:49:14Z","receivedAt":"2009-02-06T10:49:14Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Make SHA-1 regexp to be turned into hyperlink (SHA-1 committag)\nto match word boundary at beginning and end.  This way we limit\nfalse matches, for example 0x74a5cd01 which is hex decimal (for\nexample memory address) but not SHA-1.\n\nAlso make sure that it is not Message-ID, which fragment just\nlooks like SHA-1 (e.g. \"Message-ID: <46A0F335@example.com>\"),\nby using zero-width negative look-ahead assertion to _not_\nmatch '@' after.\n\nSuggested-by: Johannes Schindelin <Johannes.Schindelin@gmx.de>\nSuggested-by: Rafael Garcia-Suarez <rgarciasuarez@gmail.com>\nSigned-off-by: Jakub Narebski <jnareb@gmail.com>\n---\nv2: Added protection against matching Message-IDs fragments.\n\n gitweb/gitweb.perl |    2 +-\n 1 files changed, 1 insertions(+), 1 deletions(-)\n\ndiff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\nindex f27dbb6..5dcc108 100755\n--- a/gitweb/gitweb.perl\n+++ b/gitweb/gitweb.perl\n@@ -1364,7 +1364,7 @@ sub format_log_line_html {\n \tmy $line = shift;\n \n \t$line = esc_html($line, -nbsp=>1);\n-\tif ($line =~ m/([0-9a-fA-F]{8,40})/) {\n+\tif ($line =~ m/\\b([0-9a-fA-F]{8,40})\\b(!?\\@)/) {\n \t\tmy $hash_text = $1;\n \t\tmy $link =\n \t\t\t$cgi->a({-href => href(action=>\"object\", hash=>$hash_text),\n-- \n1.6.1\n"},{"id":"103468","messageId":"alpine.DEB.1.00.0902061403130.7377@intel-tinevez-2-302","threadId":"17516","inReplyTo":"200902061149.16210.jnareb@gmail.com","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-02-06T13:03:39Z","receivedAt":"2009-02-06T13:03:39Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 6 Feb 2009, Jakub Narebski wrote:\n\n> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n> index f27dbb6..5dcc108 100755\n> --- a/gitweb/gitweb.perl\n> +++ b/gitweb/gitweb.perl\n> @@ -1364,7 +1364,7 @@ sub format_log_line_html {\n>  \tmy $line = shift;\n>  \n>  \t$line = esc_html($line, -nbsp=>1);\n> -\tif ($line =~ m/([0-9a-fA-F]{8,40})/) {\n> +\tif ($line =~ m/\\b([0-9a-fA-F]{8,40})\\b(!?\\@)/) {\n\nLooks good to me!\n\nThanks,\nDscho\n"},{"id":"103561","messageId":"76718490902061347h5bc35e7et9e1b66bf9dd2c93a@mail.gmail.com","threadId":"17516","inReplyTo":"alpine.DEB.1.00.0902061403130.7377@intel-tinevez-2-302","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2009-02-06T21:47:57Z","receivedAt":"2009-02-06T21:47:57Z","isPatch":false,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"On Fri, Feb 6, 2009 at 8:03 AM, Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Fri, 6 Feb 2009, Jakub Narebski wrote:\n>\n>> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n>> index f27dbb6..5dcc108 100755\n>> --- a/gitweb/gitweb.perl\n>> +++ b/gitweb/gitweb.perl\n>> @@ -1364,7 +1364,7 @@ sub format_log_line_html {\n>>       my $line = shift;\n>>\n>>       $line = esc_html($line, -nbsp=>1);\n>> -     if ($line =~ m/([0-9a-fA-F]{8,40})/) {\n>> +     if ($line =~ m/\\b([0-9a-fA-F]{8,40})\\b(!?\\@)/) {\n>\n> Looks good to me!\n\nI wonder if just matching lower-case a-f would be sufficient as well?\n\nj.\n"},{"id":"103562","messageId":"200902062300.16798.jnareb@gmail.com","threadId":"17516","inReplyTo":"76718490902061347h5bc35e7et9e1b66bf9dd2c93a@mail.gmail.com","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-06T22:00:13Z","receivedAt":"2009-02-06T22:00:13Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Jay Soffian wrote:\n> On Fri, Feb 6, 2009 at 8:03 AM, Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n>> On Fri, 6 Feb 2009, Jakub Narebski wrote:\n>>\n>>> diff --git a/gitweb/gitweb.perl b/gitweb/gitweb.perl\n>>> index f27dbb6..5dcc108 100755\n>>> --- a/gitweb/gitweb.perl\n>>> +++ b/gitweb/gitweb.perl\n>>> @@ -1364,7 +1364,7 @@ sub format_log_line_html {\n>>>       my $line = shift;\n>>>\n>>>       $line = esc_html($line, -nbsp=>1);\n>>> -     if ($line =~ m/([0-9a-fA-F]{8,40})/) {\n>>> +     if ($line =~ m/\\b([0-9a-fA-F]{8,40})\\b(!?\\@)/) {\n>>\n>> Looks good to me!\n> \n> I wonder if just matching lower-case a-f would be sufficient as well?\n\nWell... \n\nOn one hand side git generates always lower-case a-f for SHA-1.\nOn the other hand git _accepts_ upper-case A-F for SHA-1 of object.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"103593","messageId":"7vd4duzo07.fsf@gitster.siamese.dyndns.org","threadId":"17516","inReplyTo":"200902061149.16210.jnareb@gmail.com","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-02-07T07:48:08Z","receivedAt":"2009-02-07T07:48:08Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> Make SHA-1 regexp to be turned into hyperlink (SHA-1 committag)\n> to match word boundary at beginning and end.  This way we limit\n> false matches, for example 0x74a5cd01 which is hex decimal (for\n> example memory address) but not SHA-1.\n>\n> Also make sure that it is not Message-ID, which fragment just\n> looks like SHA-1 (e.g. \"Message-ID: <46A0F335@example.com>\"),\n> by using zero-width negative look-ahead assertion to _not_\n> match '@' after.\n\nYour message I am responding to is:\n\n    Message-ID: <200902061149.16210.jnareb@gmail.com>\n\nDoes your description mean that \"200902061149\" would match, because the\nLAA will say \"Ah, dot is not an at-sign\"?\n"},{"id":"103599","messageId":"200902070934.50555.jnareb@gmail.com","threadId":"17516","inReplyTo":"7vd4duzo07.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-07T08:34:48Z","receivedAt":"2009-02-07T08:34:48Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n> > Make SHA-1 regexp to be turned into hyperlink (SHA-1 committag)\n> > to match word boundary at beginning and end.  This way we limit\n> > false matches, for example 0x74a5cd01 which is hex decimal (for\n> > example memory address) but not SHA-1.\n> >\n> > Also make sure that it is not Message-ID, which fragment just\n> > looks like SHA-1 (e.g. \"Message-ID: <46A0F335@example.com>\"),\n> > by using zero-width negative look-ahead assertion to _not_\n> > match '@' after.\n> \n> Your message I am responding to is:\n> \n>     Message-ID: <200902061149.16210.jnareb@gmail.com>\n> \n> Does your description mean that \"200902061149\" would match, because the\n> LAA will say \"Ah, dot is not an at-sign\"?\n\nIt would unfortunately falsely match... but we cannot eliminate this\ncase (well, at least not checking if hexnumber is followed by dot),\nbecause of totally legitimate\n\n   ... at commit 8457bb9e.\n\nSo even with that we would have still false matches...\n-- \nJakub Narebski\nPoland\n"},{"id":"103602","messageId":"7v7i42y6ms.fsf@gitster.siamese.dyndns.org","threadId":"17516","inReplyTo":"200902070934.50555.jnareb@gmail.com","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-02-07T08:48:43Z","receivedAt":"2009-02-07T08:48:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> writes:\n\n> It would unfortunately falsely match... but we cannot eliminate this\n> case (well, at least not checking if hexnumber is followed by dot),\n> because of totally legitimate\n>\n>    ... at commit 8457bb9e.\n>\n> So even with that we would have still false matches...\n\nYeah, so what's the value in v2 over v1?  It is still wrong but it is less\nwrong than it used to be?  I think the word-boundary one made a good\nsense.  I do not see the @lookahead adding much value at all.\n"},{"id":"103603","messageId":"9b18b3110902070122r3397888aqcaebfcf3e6d40d51@mail.gmail.com","threadId":"17516","inReplyTo":"200902061126.18418.jnareb@gmail.com","subject":"Re: [PATCH] gitweb: Better regexp for SHA-1 committag match","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2009-02-07T09:22:15Z","receivedAt":"2009-02-07T09:22:15Z","isPatch":true,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n> Dnia piątek 6. lutego 2009 10:49, Rafael Garcia-Suarez napisał:\n>> 2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n>\n>> > Make SHA-1 regexp to be turned into hyperlink (the SHA-1 committag)\n>> > to match word boundary at the beginning and the end.  This way we\n>> > reduce number of false matches, for example we now don't match\n>> > 0x74a5cd01 which is hex decimal (for example memory address),\n>> > but is not SHA-1.\n>>\n>> Further suggestion: you could also turn the final \\b into (\\b|\\@),\n>\n> You meant \\b -> \\b(?!\\@), didn't you?  Word boundary _not_ followed\n> by '@', and not word boundary _OR_ '@' as you wrote...\n\nSince \\b(?!\\@) is effectively two zero width negative assertions in a\nrow you could simplify by saying:\n\n  (?![^\\w\\@])\n\nand that way you can easily add the '.' case as well.\n\nYves\n\n\n\n\n\n\n\n\n\n\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"103604","messageId":"200902071025.02491.jnareb@gmail.com","threadId":"17516","inReplyTo":"7v7i42y6ms.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-07T09:25:01Z","receivedAt":"2009-02-07T09:25:01Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Junio C Hamano wrote:\n> Jakub Narebski <jnareb@gmail.com> writes:\n> \n> > It would unfortunately falsely match... but we cannot eliminate this\n> > case (well, at least not checking if hexnumber is followed by dot),\n> > because of totally legitimate\n> >\n> >    ... at commit 8457bb9e.\n> >\n> > So even with that we would have still false matches...\n> \n> Yeah, so what's the value in v2 over v1?  It is still wrong but it is less\n> wrong than it used to be?  I think the word-boundary one made a good\n> sense.  I do not see the @lookahead adding much value at all.\n\nRight. So v2 is less useful that I thought it to be; and adding further\n\"exceptions\" doesn't seem like a good idea.  The 'msgid' committag\nwhen/if it gets implemented would help there...\n\nSo please take v1, as it is sane improvement and generic enough.\n-- \nJakub Narebski\nPoland\n"},{"id":"103605","messageId":"9b18b3110902070132h2401a2f1w7abefa1c9906a567@mail.gmail.com","threadId":"17516","inReplyTo":"200902071025.02491.jnareb@gmail.com","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2009-02-07T09:32:37Z","receivedAt":"2009-02-07T09:32:37Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"2009/2/7 Jakub Narebski <jnareb@gmail.com>:\n> Junio C Hamano wrote:\n>> Jakub Narebski <jnareb@gmail.com> writes:\n>>\n>> > It would unfortunately falsely match... but we cannot eliminate this\n>> > case (well, at least not checking if hexnumber is followed by dot),\n>> > because of totally legitimate\n>> >\n>> >    ... at commit 8457bb9e.\n>> >\n>> > So even with that we would have still false matches...\n>>\n>> Yeah, so what's the value in v2 over v1?  It is still wrong but it is less\n>> wrong than it used to be?  I think the word-boundary one made a good\n>> sense.  I do not see the @lookahead adding much value at all.\n>\n> Right. So v2 is less useful that I thought it to be; and adding further\n> \"exceptions\" doesn't seem like a good idea.  The 'msgid' committag\n> when/if it gets implemented would help there...\n>\n> So please take v1, as it is sane improvement and generic enough.\n\nIf you make it configurable then everybody can be happy right?\n\nYves\n\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"103607","messageId":"200902071107.33428.jnareb@gmail.com","threadId":"17516","inReplyTo":"9b18b3110902070122r3397888aqcaebfcf3e6d40d51@mail.gmail.com","subject":"Re: [PATCH] gitweb: Better regexp for SHA-1 committag match","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-07T10:07:32Z","receivedAt":"2009-02-07T10:07:32Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, 7 Feb 2009, demerphq wrote:\n> 2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n>> Dnia piątek 6. lutego 2009 10:49, Rafael Garcia-Suarez napisał:\n>>> 2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n\n>>>> Make SHA-1 regexp to be turned into hyperlink (the SHA-1 committag)\n>>>> to match word boundary at the beginning and the end.  This way we\n>>>> reduce number of false matches, for example we now don't match\n>>>> 0x74a5cd01 which is hex decimal (for example memory address),\n>>>> but is not SHA-1.\n>>>\n>>> Further suggestion: you could also turn the final \\b into (\\b|\\@),\n>>\n>> You meant \\b -> \\b(?!\\@), didn't you?  Word boundary _not_ followed\n>> by '@', and not word boundary _OR_ '@' as you wrote...\n> \n> Since \\b(?!\\@) is effectively two zero width negative assertions in a\n> row you could simplify by saying:\n> \n>   (?![^\\w\\@])\n\nI don't know if \"sth\\b\" is effectively \"sth(!?[^\\w])\"... perhaps it is.\n\n> \n> and that way you can easily add the '.' case as well.\n\nWe cannot add '.' case, because it there can be legitimate SHA-1 match\nending sentence, e.g.\n\n     ... at commit 8457bb9e.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"103606","messageId":"200902071109.27894.jnareb@gmail.com","threadId":"17516","inReplyTo":"9b18b3110902070132h2401a2f1w7abefa1c9906a567@mail.gmail.com","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-07T10:09:27Z","receivedAt":"2009-02-07T10:09:27Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Sat, 7 Feb 2009, demerphq wrote:\n> 2009/2/7 Jakub Narebski <jnareb@gmail.com>:\n>> Junio C Hamano wrote:\n>>> Jakub Narebski <jnareb@gmail.com> writes:\n>>>\n>>>> It would unfortunately falsely match... but we cannot eliminate this\n>>>> case (well, at least not checking if hexnumber is followed by dot),\n>>>> because of totally legitimate\n>>>>\n>>>>    ... at commit 8457bb9e.\n>>>>\n>>>> So even with that we would have still false matches...\n>>>\n>>> Yeah, so what's the value in v2 over v1?  It is still wrong but it is less\n>>> wrong than it used to be?  I think the word-boundary one made a good\n>>> sense.  I do not see the @lookahead adding much value at all.\n>>\n>> Right. So v2 is less useful that I thought it to be; and adding further\n>> \"exceptions\" doesn't seem like a good idea.  The 'msgid' committag\n>> when/if it gets implemented would help there...\n>>\n>> So please take v1, as it is sane improvement and generic enough.\n> \n> If you make it configurable then everybody can be happy right?\n\nThat are the long term plans, to implement generic 'committags' support\n(which would include current SHA-1 and signoff committags).\n\nBTW. it is the 'configurable' part that makes it difficult... ;-)\n-- \nJakub Narebski\nPoland\n"},{"id":"103617","messageId":"9b18b3110902070530s70c93813se529ee7ab69b1f7e@mail.gmail.com","threadId":"17516","inReplyTo":"200902071107.33428.jnareb@gmail.com","subject":"Re: [PATCH] gitweb: Better regexp for SHA-1 committag match","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2009-02-07T13:30:22Z","receivedAt":"2009-02-07T13:30:22Z","isPatch":true,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"2009/2/7 Jakub Narebski <jnareb@gmail.com>:\n> On Sat, 7 Feb 2009, demerphq wrote:\n>> 2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n>>> Dnia piątek 6. lutego 2009 10:49, Rafael Garcia-Suarez napisał:\n>>>> 2009/2/6 Jakub Narebski <jnareb@gmail.com>:\n>\n>>>>> Make SHA-1 regexp to be turned into hyperlink (the SHA-1 committag)\n>>>>> to match word boundary at the beginning and the end.  This way we\n>>>>> reduce number of false matches, for example we now don't match\n>>>>> 0x74a5cd01 which is hex decimal (for example memory address),\n>>>>> but is not SHA-1.\n>>>>\n>>>> Further suggestion: you could also turn the final \\b into (\\b|\\@),\n>>>\n>>> You meant \\b -> \\b(?!\\@), didn't you?  Word boundary _not_ followed\n>>> by '@', and not word boundary _OR_ '@' as you wrote...\n>>\n>> Since \\b(?!\\@) is effectively two zero width negative assertions in a\n>> row you could simplify by saying:\n>>\n>>   (?![^\\w\\@])\n>\n> I don't know if \"sth\\b\" is effectively \"sth(!?[^\\w])\"... perhaps it is.\n\nSorry, my bad, that is double negation, I meant (?![\\w\\@])\n\nOn of the ways you can express \\b is as:\n(?:(?<=\\w)(?!\\w)|(?<=\\W)(?!\\W)|\\A)\n\nBut the point here is you are looking for the end of a hex sequence,\nso you can just use the \"end of string\" bit of the alternation which\nis: (?!\\w).\n\n>>\n>> and that way you can easily add the '.' case as well.\n>\n> We cannot add '.' case, because it there can be legitimate SHA-1 match\n> ending sentence, e.g.\n>\n>     ... at commit 8457bb9e.\n\n/(?<!\\w)([a-fA-F0-9]+)(?!(?:\\.\\w|[\\w@]))/\n\n:-)\n\nYves\n\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"103622","messageId":"200902071501.34312.jnareb@gmail.com","threadId":"17516","inReplyTo":"200902061149.16210.jnareb@gmail.com","subject":"Re: [PATCHv2] gitweb: Better regexp for SHA-1 committag match","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-02-07T14:01:31Z","receivedAt":"2009-02-07T14:01:31Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Dnia piątek 6. lutego 2009 11:49, Jakub Narebski napisał:\n\n> +\tif ($line =~ m/\\b([0-9a-fA-F]{8,40})\\b(!?\\@)/) {\n\n+\tif ($line =~ m/\\b([0-9a-fA-F]{8,40})\\b(?!\\@)/) {\n\nNot that it matters, because adding such single-case exceptions\nis not a good idea, so it is v1 which would be (I hope) in git.\n\n-- \nJakub Narebski\nPoland\n"}]}