{"thread":{"id":"9147","subject":"Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","startedAt":"2007-07-22T18:48:49Z","lastAt":"2007-07-24T13:08:23Z","messageCount":24,"participants":["Jon Smirl","Johannes Schindelin","Linus Torvalds","David Kastrup","Jan Engelhardt","Simon 'corecode' Schubert","Jakub Narebski"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"48161","messageId":"9e4733910707221148g69d7600bk632abb7452ce9c7c@mail.gmail.com","threadId":"9147","inReplyTo":null,"subject":"Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-22T18:48:49Z","receivedAt":"2007-07-22T18:48:49Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"It would really be useful if git diff had an option for suppressing\ndiffs caused by CVS keyword expansion. I run into this problem over\nand over when trying to recover stuff out of old kernel sources that\npeople checked into CVS and then posted CVS diffs to fulfill their GPL\nobligations. I sometimes wonder if vendors are doing this on purpose\nto make it more difficult to recover the changes they made to the\ncode.\n\nTo prevent this in the future, I'd love to see a patch removing all\nCVS keywords from the kernel sources. Quick grep of kernel shows 1,535\n$Id, 197 $Log, 441 $Revision, 144 $Date. I guestimate that this would\nremove about 5,000 lines form the kernel source and touch 1,700 files.\nWould this be accept or do those expansions contain useful info?\n\nExample diff caused by $Log expansion:\n\ndiff --git a/arch/cris/kernel/ptrace.c b/arch/cris/kernel/ptrace.c\nindex b3d10fb..32e47ca 100644\n--- a/arch/cris/kernel/ptrace.c\n+++ b/arch/cris/kernel/ptrace.c\n@@ -8,6 +8,9 @@\n  * Authors:   Bjorn Wesen\n  *\n  * $Log: ptrace.c,v $\n+ * Revision 1.2  2003/06/26 21:08:16  vangool\n+ * Commit kernel changes.\n+ *\n  * Revision 1.8  2001/11/12 18:26:21  pkj\n  * Fixed compiler warnings.\n  *\n\nOr in this case the person edited out all of the lines in the kernel\ncontaining $Id\n\ndiff --git a/arch/i386/boot/tools/build.c b/arch/i386/boot/tools/build.c\nindex 2edd0a4..792aeb1 100644\n--- a/arch/i386/boot/tools/build.c\n+++ b/arch/i386/boot/tools/build.c\n@@ -1,5 +1,4 @@\n /*\n- *  $Id: build.c,v 1.5 1997/05/19 12:29:58 mj Exp $\n  *\n  *  Copyright (C) 1991, 1992  Linus Torvalds\n  *  Copyright (C) 1997 Martin Mares\n\nA variation on this where the Id was updated..\n- *  $Id: build.c,v 1.5 1997/05/19 12:29:58 mj Exp $\n+ *  $Id: build.c,v 1.5 2002/09/29 12:29:58 mj Exp $\n\nThere are a more CVS keywords that can cause expansion, probably around twenty.\nWhen removing these with a simple grep it can get fooled and mess up.\nCVS keyword expansion adds about 60,000 lines of noise to a typical\ndiff.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"48166","messageId":"Pine.LNX.4.64.0707221959100.14781@racer.site","threadId":"9147","inReplyTo":"9e4733910707221148g69d7600bk632abb7452ce9c7c@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-22T19:00:36Z","receivedAt":"2007-07-22T19:00:36Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[culling lkml list, since this is really a Git problem]\n\nOn Sun, 22 Jul 2007, Jon Smirl wrote:\n\n> It would really be useful if git diff had an option for suppressing\n> diffs caused by CVS keyword expansion.\n\nLooks to me slightly a bit like an XY problem.\n\nHow about using git-filter-branch to get rid of the expansions?  Or even \nbetter, if you have access to the CVS server, use the -k option to \ncvsimport?\n\nHth,\nDscho\n"},{"id":"48168","messageId":"alpine.LFD.0.999.0707221205080.3607@woody.linux-foundation.org","threadId":"9147","inReplyTo":"9e4733910707221148g69d7600bk632abb7452ce9c7c@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-22T19:09:01Z","receivedAt":"2007-07-22T19:09:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 22 Jul 2007, Jon Smirl wrote:\n>\n> It would really be useful if git diff had an option for suppressing\n> diffs caused by CVS keyword expansion.\n\nI really think it's not a \"git diff\" issue, but it might be a \"import\" \nissue.\n\nIOW, I think you'd be a *lot* better off just not importing those things \nin the first place (which is what CVS does internally), or possibly \nimporting them as two trees (ie you'd have the \"non-log\" version and the \n\"log expansion\" version, so that you can track and compare both).\n\nDoing the thing at \"diff\" time is certainly possible, but this is simply \nmuch better done as a totally independent preprocessing phase. The diff \nhandling is already some of the more complex parts (and very central), it \nwould be much simpler and efficient to not try to make that thing fancier, \nand instead solve the problem at the front-end.\n\n\t\t\tLinus\n"},{"id":"48169","messageId":"9e4733910707221210t2b2896b5ob4ce7bf95d4a707a@mail.gmail.com","threadId":"9147","inReplyTo":"Pine.LNX.4.64.0707221959100.14781@racer.site","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-22T19:10:04Z","receivedAt":"2007-07-22T19:10:04Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> [culling lkml list, since this is really a Git problem]\n>\n> On Sun, 22 Jul 2007, Jon Smirl wrote:\n>\n> > It would really be useful if git diff had an option for suppressing\n> > diffs caused by CVS keyword expansion.\n>\n> Looks to me slightly a bit like an XY problem.\n>\n> How about using git-filter-branch to get rid of the expansions?  Or even\n> better, if you have access to the CVS server, use the -k option to\n> cvsimport?\n\nNo access to the server the people handing over the diffs and not\nmaking things easy. You get diffs like this when asking companies to\nturn over their changes for GPL compliance purposes. A big source is\nfrom companies building embedded systems., Usually the diffs are\npretty worthless but every once in a while they contain very useful\nnuggets (like inadvertently documenting secret hardware features). It\nis hard to spot the nuggets in all of the noise.\n\n\n>\n> Hth,\n> Dscho\n>\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"48170","messageId":"9e4733910707221212v2f6cc1c4kf7a35e84f351e4cd@mail.gmail.com","threadId":"9147","inReplyTo":"alpine.LFD.0.999.0707221205080.3607@woody.linux-foundation.org","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-22T19:12:03Z","receivedAt":"2007-07-22T19:12:03Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 7/22/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n>\n>\n> On Sun, 22 Jul 2007, Jon Smirl wrote:\n> >\n> > It would really be useful if git diff had an option for suppressing\n> > diffs caused by CVS keyword expansion.\n>\n> I really think it's not a \"git diff\" issue, but it might be a \"import\"\n> issue.\n>\n> IOW, I think you'd be a *lot* better off just not importing those things\n> in the first place (which is what CVS does internally), or possibly\n> importing them as two trees (ie you'd have the \"non-log\" version and the\n> \"log expansion\" version, so that you can track and compare both).\n>\n> Doing the thing at \"diff\" time is certainly possible, but this is simply\n> much better done as a totally independent preprocessing phase. The diff\n> handling is already some of the more complex parts (and very central), it\n> would be much simpler and efficient to not try to make that thing fancier,\n> and instead solve the problem at the front-end.\n\nThese diffs are coming from companies doing GPL compliance without\nreally wanting to comply. CVS servers are not made available.\n\n\n>\n>                         Linus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"48171","messageId":"Pine.LNX.4.64.0707222013200.14781@racer.site","threadId":"9147","inReplyTo":"9e4733910707221210t2b2896b5ob4ce7bf95d4a707a@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-22T19:15:29Z","receivedAt":"2007-07-22T19:15:29Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 22 Jul 2007, Jon Smirl wrote:\n\n> On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>\n> > On Sun, 22 Jul 2007, Jon Smirl wrote:\n> > \n> > > It would really be useful if git diff had an option for suppressing \n> > > diffs caused by CVS keyword expansion.\n> > \n> > Looks to me slightly a bit like an XY problem.\n> > \n> > How about using git-filter-branch to get rid of the expansions?  Or \n> > even better, if you have access to the CVS server, use the -k option \n> > to cvsimport?\n> \n> No access to the server the people handing over the diffs and not\n> making things easy.\n\nAh yes, that's what I feared.  Darn.\n\nBut still, I think that it would be much better not to put this into Git.  \nWe do have diff gitattributes now, so that you can roll your own diff for \nspecific files, but I still think that this is more a task for a \nstandalone perl script.  Possibly being called from filter-branch to be \ndone with the conversion once and for all times.\n\nCiao,\nDscho\n"},{"id":"48174","messageId":"85644ci6yv.fsf@lola.goethe.zz","threadId":"9147","inReplyTo":"alpine.LFD.0.999.0707221205080.3607@woody.linux-foundation.org","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2007-07-22T19:17:12Z","receivedAt":"2007-07-22T19:17:12Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sun, 22 Jul 2007, Jon Smirl wrote:\n>>\n>> It would really be useful if git diff had an option for suppressing\n>> diffs caused by CVS keyword expansion.\n>\n> I really think it's not a \"git diff\" issue, but it might be a \"import\" \n> issue.\n>\n> IOW, I think you'd be a *lot* better off just not importing those\n> things in the first place (which is what CVS does internally), or\n> possibly importing them as two trees (ie you'd have the \"non-log\"\n> version and the \"log expansion\" version, so that you can track and\n> compare both).\n\nOne problem is that those strings more often than not are involved in\nsome magic computation, like placing version and date information into\nextracted output.  While an import of the unexpanded version is really\nthe sanest option, it might render the resulting code inoperable.\n\nCVS diff itself reports those differences, too (and it makes for a\ngood quota of merge problems even in CVS), so it is not like we are in\nbad company here.\n\nUh, strike that.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"48177","messageId":"alpine.LFD.0.999.0707221234250.3607@woody.linux-foundation.org","threadId":"9147","inReplyTo":"9e4733910707221212v2f6cc1c4kf7a35e84f351e4cd@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-07-22T19:37:22Z","receivedAt":"2007-07-22T19:37:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 22 Jul 2007, Jon Smirl wrote:\n\n> On 7/22/07, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> > \n> > \n> > On Sun, 22 Jul 2007, Jon Smirl wrote:\n> > >\n> > > It would really be useful if git diff had an option for suppressing\n> > > diffs caused by CVS keyword expansion.\n> > \n> > I really think it's not a \"git diff\" issue, but it might be a \"import\"\n> > issue.\n> > \n> > IOW, I think you'd be a *lot* better off just not importing those things\n> > in the first place (which is what CVS does internally), or possibly\n> > importing them as two trees (ie you'd have the \"non-log\" version and the\n> > \"log expansion\" version, so that you can track and compare both).\n> > \n> > Doing the thing at \"diff\" time is certainly possible, but this is simply\n> > much better done as a totally independent preprocessing phase. The diff\n> > handling is already some of the more complex parts (and very central), it\n> > would be much simpler and efficient to not try to make that thing fancier,\n> > and instead solve the problem at the front-end.\n> \n> These diffs are coming from companies doing GPL compliance without\n> really wanting to comply. CVS servers are not made available.\n\nThat wasn't what I said. \n\nYou want to supporess the CVS keyword expansion.\n\nI'm telling you that you should just do so.\n\nBut I'm *also* telling you that this has nothing to do with \"git diff\".\n\nThe way to get \"git diff\" to not show the CVS expansion is to *remove* the \nCVS expansion. By simply running some pre-processing on the patches \n*before* you put them into git in the first place (or, like Dscho \nsuggested: you could do it later too, by using git filter-branch).\n\nOnce you have the version that doesn't have the CVS log, \"git diff\" will \nautomatically do the rigth thing.\n\nIn other words, I'm just saying that you're trying to solve the wrong \nproblem. Once you solve the *right* problem, the wrong problem just goes \naway.\n\n\t\t\tLinus\n"},{"id":"48179","messageId":"Pine.LNX.4.64.0707222139400.15705@fbirervta.pbzchgretzou.qr","threadId":"9147","inReplyTo":"9e4733910707221148g69d7600bk632abb7452ce9c7c@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jan Engelhardt","fromEmail":"jengelh@computergmbh.de","sentAt":"2007-07-22T19:47:02Z","receivedAt":"2007-07-22T19:47:02Z","isPatch":false,"sender":{"key":"jengelh@computergmbh.de","avatar":null},"body":"\nOn Jul 22 2007 14:48, Jon Smirl wrote:\n>\n> It would really be useful if git diff had an option for suppressing\n> diffs caused by CVS keyword expansion. I run into this problem over\n> and over when trying to recover stuff out of old kernel sources that\n> people checked into CVS and then posted CVS diffs to fulfill their GPL\n> obligations. I sometimes wonder if vendors are doing this on purpose\n> to make it more difficult to recover the changes they made to the\n> code.\n>\n> To prevent this in the future, I'd love to see a patch removing all\n> CVS keywords from the kernel sources.\n\nI say: yes please. Even svn (which Linus dislikes very much) has cvs\nkeywords disabled by default. And at the same time, a patch to\nCodingStyle not to include CVS tags, and a patch to checkpatch.pl\nto enforce it :)\n\n> Quick grep of kernel shows 1,535 $Id, 197 $Log, 441 $Revision, 144 $Date. I\n> guestimate that this would remove about 5,000 lines form the kernel source\n> and touch 1,700 files. Would this be accept or do those expansions contain\n> useful info?\n\nCommon argument for $Id$ tags has been, that upon a problem, a user\ncan look into dmesg (or whatever log is appropraite - it's not\nlimited to the kernel) for the version string instead of figuring it\nout himself (by means or rpm -q, or whatever hides it).\n\nFor the Linux kernel however, I do not think it is important, since:\n\n- the code of a module changes faster than you can think (e.g.\n  a change in the network code touched a network filesystem). Result:\n  same $Id$ tag in the vendor's version AND the mainline version.\n  So the tag is useless.\n\n- you know what version you are running, mingetty usually prints it\n  for you; uname -a also shows it.\n\n- if you have a problem, people kindly ask you to bisect anyway, in\n  which case $Id$ numbers become irrelevant anyway, because you got\n  the git SHA1 ids.\n\nYep, that's pretty much it from my side. Conclusion: Remove CVS\ntags from kernel.\n\n\n\n\tJan\n-- \n"},{"id":"48180","messageId":"9e4733910707221248q45fb3aaala9c79afd4b09830e@mail.gmail.com","threadId":"9147","inReplyTo":"Pine.LNX.4.64.0707222013200.14781@racer.site","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-22T19:48:53Z","receivedAt":"2007-07-22T19:48:53Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> But still, I think that it would be much better not to put this into Git.\n> We do have diff gitattributes now, so that you can roll your own diff for\n> specific files, but I still think that this is more a task for a\n> standalone perl script.  Possibly being called from filter-branch to be\n> done with the conversion once and for all times.\n\nI can provide sample diffs if you want something to play with.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"48197","messageId":"Pine.LNX.4.64.0707222238180.14781@racer.site","threadId":"9147","inReplyTo":"9e4733910707221248q45fb3aaala9c79afd4b09830e@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-22T21:39:50Z","receivedAt":"2007-07-22T21:39:50Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 22 Jul 2007, Jon Smirl wrote:\n\n> On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > But still, I think that it would be much better not to put this into Git.\n> > We do have diff gitattributes now, so that you can roll your own diff for\n> > specific files, but I still think that this is more a task for a\n> > standalone perl script.  Possibly being called from filter-branch to be\n> > done with the conversion once and for all times.\n> \n> I can provide sample diffs if you want something to play with.\n\nAs you already did, this is my attempt at a perl script...  Feel free to \nbash my Perl capabilities, or to correct it...\n\nCiao,\nDscho\n\n-- snipsnap --\n#!/usr/bin/perl\n\nsub init_hunk {\n\t$_ = $_[0];\n\t$current_hunk = \"\";\n\t$current_hunk_header = $_;\n\t($start_minus, $dummy, $start_plus, $dummy) =\n\t\t/^\\@\\@ -(\\d+)(,\\d+|) \\+(\\d+)(,\\d+|) \\@\\@/;\n\t$plus = $minus = $space = 0;\n\t$skip_logs = 0;\n}\n\nsub flush_hunk {\n\tif ($plus > 0 || $minus > 0) {\n\t\tif ($current_file ne \"\") {\n\t\t\tprint $current_file;\n\t\t\t$current_file = \"\";\n\t\t}\n\t\t$minus += $space;\n\t\t$plus += $space;\n\t\tprint \"\\@\\@ -$start_minus,$minus \"\n\t\t\t. \"+$start_plus,$plus \\@\\@\\n\";\n\t\tprint $current_hunk;\n\t}\n}\n\nsub check_file {\n\t$_ = $_[0];\n\t$current_file = $_;\n\twhile (<>) {\n\t\tif (/^\\@\\@/) {\n\t\t\tlast;\n\t\t}\n\t\t$current_file .= $_;\n\t}\n\n\tinit_hunk $_;\n\n\t# check hunks\n\twhile (<>) {\n\t\tif ($skip_logs && /^\\+ *\\*/) {\n\t\t\t# do nothing\n\t\t} elsif (/^\\@\\@.*/) {\n\t\t\tflush_hunk;\n\t\t\tinit_hunk $_;\n\t\t} elsif (/^diff/) {\n\t\t\tflush_hunk;\n\t\t\treturn;\n\t\t} elsif (/^-.*\\$(Id|Revision|Author|Date).*\\$/) {\n\t\t\t$key = $1;\n\t\t\ts/^-/ /;\n\t\t\t$current_hunk .= $_;\n\t\t\t$space++;\n\t\t\t$_ = <>;\n\t\t\tif (!/\\+.*\\$Id.*\\$/) {\n\t\t\t\tdie \"Expected some changed \\$$key line: $_\";\n\t\t\t}\n\t\t\t$skip_logs = 0;\n\t\t} elsif (/^ .*\\$Log.*\\$/) {\n\t\t\t$current_hunk .= $_;\n\t\t\t$space++;\n\t\t\t$skip_logs++;\n\t\t} elsif (/^ /) {\n\t\t\t$current_hunk .= $_;\n\t\t\t$space++;\n\t\t\t$skip_logs = 0;\n\t\t} elsif (/^\\+/) {\n\t\t\t$current_hunk .= $_;\n\t\t\t$plus++;\n\t\t} elsif (/^-/) {\n\t\t\t$current_hunk .= $_;\n\t\t\t$minus++;\n\t\t\t$skip_logs = 0;\n\t\t} else {\n\t\t\tdie \"Unexpected line: $_\";\n\t\t}\n\t}\n}\n\nwhile (<>) {\n\tif (/^diff/) {\n\t\tdo {\n\t\t\tcheck_file $_;\n\t\t} while(/^diff/);\n\t} else {\n\t\tprintf $_;\n\t}\n}\n"},{"id":"48242","messageId":"9e4733910707221645x21d74e70y3c43bc8c02a9d4ca@mail.gmail.com","threadId":"9147","inReplyTo":"Pine.LNX.4.64.0707222238180.14781@racer.site","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-22T23:45:58Z","receivedAt":"2007-07-22T23:45:58Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> As you already did, this is my attempt at a perl script...  Feel free to\n> bash my Perl capabilities, or to correct it...\n\nI don't really know Perl but Perl is probably a better language for\nthis. I was doing it in C.\nThis doesn't run on the two full kernel samples I sent you. That's the\nproblem I was having, I can catch 95% of the expanded keywords with my\nprogram but I have to touch up 5% by hand. I'll stare at this and see\nif I can increase my understanding of Perl.\n\n>\n> Ciao,\n> Dscho\n>\n> -- snipsnap --\n> #!/usr/bin/perl\n>\n> sub init_hunk {\n>         $_ = $_[0];\n>         $current_hunk = \"\";\n>         $current_hunk_header = $_;\n>         ($start_minus, $dummy, $start_plus, $dummy) =\n>                 /^\\@\\@ -(\\d+)(,\\d+|) \\+(\\d+)(,\\d+|) \\@\\@/;\n>         $plus = $minus = $space = 0;\n>         $skip_logs = 0;\n> }\n>\n> sub flush_hunk {\n>         if ($plus > 0 || $minus > 0) {\n>                 if ($current_file ne \"\") {\n>                         print $current_file;\n>                         $current_file = \"\";\n>                 }\n>                 $minus += $space;\n>                 $plus += $space;\n>                 print \"\\@\\@ -$start_minus,$minus \"\n>                         . \"+$start_plus,$plus \\@\\@\\n\";\n>                 print $current_hunk;\n>         }\n> }\n>\n> sub check_file {\n>         $_ = $_[0];\n>         $current_file = $_;\n>         while (<>) {\n>                 if (/^\\@\\@/) {\n>                         last;\n>                 }\n>                 $current_file .= $_;\n>         }\n>\n>         init_hunk $_;\n>\n>         # check hunks\n>         while (<>) {\n>                 if ($skip_logs && /^\\+ *\\*/) {\n>                         # do nothing\n>                 } elsif (/^\\@\\@.*/) {\n>                         flush_hunk;\n>                         init_hunk $_;\n>                 } elsif (/^diff/) {\n>                         flush_hunk;\n>                         return;\n>                 } elsif (/^-.*\\$(Id|Revision|Author|Date).*\\$/) {\n>                         $key = $1;\n>                         s/^-/ /;\n>                         $current_hunk .= $_;\n>                         $space++;\n>                         $_ = <>;\n>                         if (!/\\+.*\\$Id.*\\$/) {\n>                                 die \"Expected some changed \\$$key line: $_\";\n>                         }\n>                         $skip_logs = 0;\n>                 } elsif (/^ .*\\$Log.*\\$/) {\n>                         $current_hunk .= $_;\n>                         $space++;\n>                         $skip_logs++;\n>                 } elsif (/^ /) {\n>                         $current_hunk .= $_;\n>                         $space++;\n>                         $skip_logs = 0;\n>                 } elsif (/^\\+/) {\n>                         $current_hunk .= $_;\n>                         $plus++;\n>                 } elsif (/^-/) {\n>                         $current_hunk .= $_;\n>                         $minus++;\n>                         $skip_logs = 0;\n>                 } else {\n>                         die \"Unexpected line: $_\";\n>                 }\n>         }\n> }\n>\n> while (<>) {\n>         if (/^diff/) {\n>                 do {\n>                         check_file $_;\n>                 } while(/^diff/);\n>         } else {\n>                 printf $_;\n>         }\n> }\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"48243","messageId":"Pine.LNX.4.64.0707230048570.14781@racer.site","threadId":"9147","inReplyTo":"9e4733910707221645x21d74e70y3c43bc8c02a9d4ca@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-22T23:50:38Z","receivedAt":"2007-07-22T23:50:38Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 22 Jul 2007, Jon Smirl wrote:\n\n> On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > As you already did, this is my attempt at a perl script...  Feel free to\n> > bash my Perl capabilities, or to correct it...\n> \n> I don't really know Perl but Perl is probably a better language for\n> this. I was doing it in C.\n\nYes, Perl is much much better for an application like this.  You do not \nhave to cater for newbies, but can keep it to the technical level.\n\n> This doesn't run on the two full kernel samples I sent you.\n\nYou mean my script?  Or your C program?\n\n> That's the problem I was having, I can catch 95% of the expanded \n> keywords with my program but I have to touch up 5% by hand. I'll stare \n> at this and see if I can increase my understanding of Perl.\n\nIf it is my script, please tell me what is not working.  I'll fix it.  In \nthe process, I could even add some comments...\n\nCiao,\nDscho\n"},{"id":"48246","messageId":"9e4733910707221711u6e965e6cr29e06fa8fb09165@mail.gmail.com","threadId":"9147","inReplyTo":"Pine.LNX.4.64.0707230048570.14781@racer.site","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-23T00:11:23Z","receivedAt":"2007-07-23T00:11:23Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > This doesn't run on the two full kernel samples I sent you.\n>\n> You mean my script?  Or your C program?\n\nYour script, I sent two big samples directly to you as attachments\nthree hours ago.\nScript processes them for a few pages of output and then bails out.\nI have more diffs,  but those two are representative of the sample\n\nphytec patch has support for the NXP 3180 CPU which is not in the\nkernel. I have a project interested in using this CPU so we are\nlooking at extracting the essence of 3180 support from the patch and\ntrying to apply it to the current kernel.\n\n>From the phytec patch:\ndiff -uarN linux-2.6.10/arch/arm/vfp/entry.S\nlinux-2.6.10-lpc3180/arch/arm/vfp/entry.S\n--- linux-2.6.10/arch/arm/vfp/entry.S   2004-12-25 05:35:27.000000000 +0800\n+++ linux-2.6.10-lpc3180/arch/arm/vfp/entry.S   2006-11-20\n15:49:30.000000000 +0800\n@@ -19,6 +19,7 @@\n #include <linux/init.h>\n #include <asm/thread_info.h>\n #include <asm/vfpmacros.h>\n+#include <asm/constants.h>\n\n        .globl  do_vfp\n do_vfp:\nExpected some changed $Revision line: +#define DRIVER_VERSION \"$Revision: 1.2 $\"\njonsmirl@jonsmirl:/extra/diffs$\n\n>From the sonos one:\ndiff --git a/Documentation/cachetlb.txt b/Documentation/cachetlb.txt\nindex 127a661..02dec92 100644\n--- a/Documentation/cachetlb.txt\n+++ b/Documentation/cachetlb.txt\n@@ -222,7 +222,7 @@\n this value.\n\n NOTE: This does not fix shared mmaps, check out the sparc64 port for\n-one way to solve this (in particular SPARC_FLAG_MMAPSHARED).\n+one way to solve this (in particular arch_get_unmapped_area).\n\n Next, you have two methods to solve the D-cache aliasing issue for all\n other cases.  Please keep in mind that fact that, for a given page\nExpected some changed $Id line:            Readme-File\n/usr/src/Documentation/cdrom/aztcd\njonsmirl@jonsmirl:/extra/diffs$\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"48249","messageId":"Pine.LNX.4.64.0707230136360.14781@racer.site","threadId":"9147","inReplyTo":"9e4733910707221711u6e965e6cr29e06fa8fb09165@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-23T00:37:25Z","receivedAt":"2007-07-23T00:37:25Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 22 Jul 2007, Jon Smirl wrote:\n\n> On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > > This doesn't run on the two full kernel samples I sent you.\n> > \n> > You mean my script?  Or your C program?\n> \n> Your script, I sent two big samples directly to you as attachments\n> three hours ago.\n\nOkay, I did not really test thoroughly, it seems.  Sorry.  Next try.\n\nCiao,\nDscho\n\n-- snipsnap --\n#!/usr/bin/perl\n\nsub init_hunk {\n\t$_ = $_[0];\n\t$current_hunk = \"\";\n\t$current_hunk_header = $_;\n\t($start_minus, $dummy, $start_plus, $dummy) =\n\t\t/^\\@\\@ -(\\d+)(,\\d+|) \\+(\\d+)(,\\d+|) \\@\\@/;\n\t$plus = $minus = $space = 0;\n\t$skip_logs = 0;\n\t$key = 0;\n}\n\nsub flush_hunk {\n\tif ($plus > 0 || $minus > 0) {\n\t\tif ($current_file ne \"\") {\n\t\t\tprint $current_file;\n\t\t\t$current_file = \"\";\n\t\t}\n\t\t$minus += $space;\n\t\t$plus += $space;\n\t\tprint \"\\@\\@ -$start_minus,$minus \"\n\t\t\t. \"+$start_plus,$plus \\@\\@\\n\";\n\t\tprint $current_hunk;\n\t}\n}\n\nsub check_file {\n\t$_ = $_[0];\n\t$current_file = $_;\n\twhile (<>) {\n\t\tif (/^\\@\\@/) {\n\t\t\tlast;\n\t\t}\n\t\t$current_file .= $_;\n\t}\n\n\tinit_hunk $_;\n\n\t# check hunks\n\twhile (<>) {\n\t\tif ($skip_logs && /^\\+ *\\*/) {\n\t\t\t# do nothing\n\t\t} elsif ($key && /^\\+.*\\$$key.*\\$/) {\n\t\t\t# do nothing\n\t\t} elsif (/^\\@\\@.*/) {\n\t\t\tflush_hunk;\n\t\t\tinit_hunk $_;\n\t\t} elsif (/^diff/) {\n\t\t\tflush_hunk;\n\t\t\treturn;\n\t\t} elsif (/^-.*\\$(Id|Revision|Author|Date).*\\$/) {\n\t\t\t$key = $1;\n\t\t\ts/^-/ /;\n\t\t\t$current_hunk .= $_;\n\t\t\t$space++;\n\t\t\t$skip_logs = 0;\n\t\t\t$key = 0;\n\t\t} elsif (/^ .*\\$Log.*\\$/) {\n\t\t\t$current_hunk .= $_;\n\t\t\t$space++;\n\t\t\t$skip_logs = 1;\n\t\t\t$key = 0;\n\t\t} elsif (/^ /) {\n\t\t\t$current_hunk .= $_;\n\t\t\t$space++;\n\t\t\t$skip_logs = 0;\n\t\t\t$key = 0;\n\t\t} elsif (/^\\+/) {\n\t\t\t$current_hunk .= $_;\n\t\t\t$plus++;\n\t\t\t$key = 0;\n\t\t} elsif (/^-/) {\n\t\t\t$current_hunk .= $_;\n\t\t\t$minus++;\n\t\t\t$skip_logs = 0;\n\t\t\t$key = 0;\n\t\t} elsif (/^\\\\/) {\n\t\t\tprint $_;\n\t\t} else {\n\t\t\tdie \"Unexpected line: $_\";\n\t\t}\n\t}\n}\n\nwhile (<>) {\n\tif (/^diff/) {\n\t\tdo {\n\t\t\tcheck_file $_;\n\t\t} while(/^diff/);\n\t} else {\n\t\tprintf $_;\n\t}\n}\n"},{"id":"48310","messageId":"9e4733910707230744u2d3a0a31t9f65d5c9e68c9805@mail.gmail.com","threadId":"9147","inReplyTo":"Pine.LNX.4.64.0707230136360.14781@racer.site","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-23T14:44:00Z","receivedAt":"2007-07-23T14:44:00Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> Okay, I did not really test thoroughly, it seems.  Sorry.  Next try.\n\nIt's working a lot better. People deal with $Id two different ways.\nOne way they delete all of the $Id lines, that is the case of the\nSonos patch. In the Phytec patch they left all of the $Id lines in\nplace which caused them to get modified. In both cases you just want\nthe lines with $Id to disappear in the patch.\n\nIt doesn't catch the $Id case from the Phytec patch.\n\ndiff -uarN linux-2.6.10/arch/cris/arch-v10/boot/rescue/head.S\nlinux-2.6.10-lpc3180/arch/cris/arch-v10/boot/rescue/head.S\n--- linux-2.6.10/arch/cris/arch-v10/boot/rescue/head.S  2004-12-25\n05:35:24.000000000 +0800\n+++ linux-2.6.10-lpc3180/arch/cris/arch-v10/boot/rescue/head.S\n2006-11-20 15:49:30.000000000 +0800\n@@ -1,4 +1,5 @@\n /* $Id: head.S,v 1.6 2003/04/09 08:12:43 pkj Exp $\n+/* $Id: head.S,v 1.2 2005/02/18 13:06:31 mike Exp $\n  *\n  * Rescue code, made to reside at the beginning of the\n  * flash-memory. when it starts, it checks a partition\n\nIt's not catching all of the $Revision and $Date deltas.\n\nThe output diff shouldn't contain any CVS keywords. It is somewhat\ntricky to catch all of the cases and fix up the diffs. This filter\nshould get written and debugged once and then made part of something\nlike git so that it doesn't get written over and over again. Perl is\nway better for this I had 1000 lines of C in my program and it was\nstill missing 10% of the cases.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"48312","messageId":"46A4C51D.2060202@fs.ei.tum.de","threadId":"9147","inReplyTo":"9e4733910707230744u2d3a0a31t9f65d5c9e68c9805@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-07-23T15:11:25Z","receivedAt":"2007-07-23T15:11:25Z","isPatch":false,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"Jon Smirl wrote:\n> On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>> Okay, I did not really test thoroughly, it seems.  Sorry.  Next try.\n> \n> It's working a lot better. People deal with $Id two different ways.\n> One way they delete all of the $Id lines, that is the case of the\n> Sonos patch. In the Phytec patch they left all of the $Id lines in\n> place which caused them to get modified. In both cases you just want\n> the lines with $Id to disappear in the patch.\n> \n> It doesn't catch the $Id case from the Phytec patch.\n> \n> diff -uarN linux-2.6.10/arch/cris/arch-v10/boot/rescue/head.S\n> linux-2.6.10-lpc3180/arch/cris/arch-v10/boot/rescue/head.S\n> --- linux-2.6.10/arch/cris/arch-v10/boot/rescue/head.S  2004-12-25\n> 05:35:24.000000000 +0800\n> +++ linux-2.6.10-lpc3180/arch/cris/arch-v10/boot/rescue/head.S\n> 2006-11-20 15:49:30.000000000 +0800\n> @@ -1,4 +1,5 @@\n> /* $Id: head.S,v 1.6 2003/04/09 08:12:43 pkj Exp $\n> +/* $Id: head.S,v 1.2 2005/02/18 13:06:31 mike Exp $\n>  *\n>  * Rescue code, made to reside at the beginning of the\n>  * flash-memory. when it starts, it checks a partition\n> \n> It's not catching all of the $Revision and $Date deltas.\n> \n> The output diff shouldn't contain any CVS keywords. It is somewhat\n> tricky to catch all of the cases and fix up the diffs. This filter\n> should get written and debugged once and then made part of something\n> like git so that it doesn't get written over and over again. Perl is\n> way better for this I had 1000 lines of C in my program and it was\n> still missing 10% of the cases.\n\nMaybe I am missing something, but apart from $Log$, what's so hard about collapsing the CVS keywords?  That's something like\n\ns/\\$(Id|Date|Header|CVSHeader|Author|Revision):[^$]*\\$/$\\1$/\n\nor not?\n\nI see that people want to modify the patch text to remove the churn, but that seems wrong.  It would be much easier to just collapse the keywords before applying the patch, and then work on a newly created diff (from the git repo).\n\ncheers\n  simon\n"},{"id":"48340","messageId":"Pine.LNX.4.64.0707231933030.14781@racer.site","threadId":"9147","inReplyTo":"9e4733910707230744u2d3a0a31t9f65d5c9e68c9805@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-23T20:11:18Z","receivedAt":"2007-07-23T20:11:18Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 23 Jul 2007, Jon Smirl wrote:\n\n> On 7/22/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > Okay, I did not really test thoroughly, it seems.  Sorry.  Next try.\n> \n> It's working a lot better. People deal with $Id two different ways.\n> One way they delete all of the $Id lines, that is the case of the\n> Sonos patch. In the Phytec patch they left all of the $Id lines in\n> place which caused them to get modified. In both cases you just want\n> the lines with $Id to disappear in the patch.\n> \n> It doesn't catch the $Id case from the Phytec patch.\n> \n> diff -uarN linux-2.6.10/arch/cris/arch-v10/boot/rescue/head.S\n> linux-2.6.10-lpc3180/arch/cris/arch-v10/boot/rescue/head.S\n> --- linux-2.6.10/arch/cris/arch-v10/boot/rescue/head.S  2004-12-25\n> 05:35:24.000000000 +0800\n> +++ linux-2.6.10-lpc3180/arch/cris/arch-v10/boot/rescue/head.S\n> 2006-11-20 15:49:30.000000000 +0800\n> @@ -1,4 +1,5 @@\n> /* $Id: head.S,v 1.6 2003/04/09 08:12:43 pkj Exp $\n> +/* $Id: head.S,v 1.2 2005/02/18 13:06:31 mike Exp $\n>  *\n>  * Rescue code, made to reside at the beginning of the\n>  * flash-memory. when it starts, it checks a partition\n> \n> It's not catching all of the $Revision and $Date deltas.\n> \n> The output diff shouldn't contain any CVS keywords. It is somewhat\n> tricky to catch all of the cases and fix up the diffs. This filter\n> should get written and debugged once and then made part of something\n> like git so that it doesn't get written over and over again. Perl is\n> way better for this I had 1000 lines of C in my program and it was\n> still missing 10% of the cases.\n\nNext try.  It includes documentation, skips $Source:$ and $Header:$ as \nwell, and the output diff does not contain any CVS keywords.\n\nAh yes, and if a file was actually created, it substitutes the --- file \nname with /dev/null.  In that case, the keywords are skipped, too!  \nHowever, when a file was deleted, they are _not_ skipped.\n\nStrictly speaking, it is wrong to just cut out the lines containing CVS \nkeywords, since it is possible that there is an important change on that \nline, but we can cross that bridge once companies play such cute games \nwith us.\n\nBut at least I tested the resulting patch with your phytec3180 patch this \ntime...\n\nCiao,\nDscho\n\n-- snipsnap --\n#!/usr/bin/perl\n\n# This is a simple state machine.\n#\n# There is the state of the current file; its header is stored\n# in $current_file to avoid outputting it when all hunks were\n# culled.  It is only printed before the first hunk, and then\n# set to \"\" to avoid outputting it twice.\n#\n# There are the states of the current hunk, stored in\n# * $current_hunk (possibly modified hunk)\n# * $start_minus, $start_plus (from the original)\n# * $plus, $minus, $space (the current count of the respective lines)\n# a hunk is only printed (in flush_hunk) if any '+' or '-' lines\n# are left after filtering.\n#\n# For $Log..$, there is the state $skip_logs, which is set to 1\n# after seeing such a line, and set to 0 when the first line\n# was seen which does not begin with '+'.\n#\n# A particularly nasty special case is when a single \"*\" was\n# misattributed by the diff to be _inserted before_ a $Log, instead\n# of _appended after_ a $Log.\n# This is the purpose of the $before_log and $after_log variables:\n# if not empty, the state machine expects the next line to begin\n# with '+' or '-', respectively, and followed by a $Log.  If this\n# expectation is not met, the variable is output.\n#\n# The variable $plus_minus_adjust contains the number of lines which\n# were skipped from the \"+\" side, so that the correct offset is shown.\n\n\n# This function gets a hunk header.\n#\n# It initializes the state variables described above\n\nsub init_hunk {\n\t$_ = $_[0];\n\t$current_hunk = \"\";\n\t$current_hunk_header = $_;\n\t($start_minus, $dummy, $start_plus, $dummy) =\n\t\t/^\\@\\@ -(\\d+)(,\\d+|) \\+(\\d+)(,\\d+|) \\@\\@/;\n\t$plus = $minus = $space = 0;\n\t$skip_logs = 0;\n\t$before_log = '';\n\t$after_log = '';\n\n\t# we prefer /dev/null as original file name when a file is new\n\tif ($start_minus eq 0) {\n\t\t$current_file =~ s/\\n--- .*\\n/\\n--- \\/dev\\/null\\n/;\n\t} elsif ($start_plus eq 0) {\n\t\t$current_file =~ s/\\n\\+\\+\\+ .*\\n/\\n+++ \\/dev\\/null\\n/;\n\t}\n}\n\n# This function is called whenever there is possibly a hunk to print.\n# Nothing is printed if no '+' or '-' lines are left.\n# Otherwise, if the file header was not yet shown, it does so now.\n\nsub flush_hunk {\n\tif ($plus > 0 || $minus > 0) {\n\t\tif ($current_file ne \"\") {\n\t\t\tprint $current_file;\n\t\t\t$current_file = \"\";\n\t\t}\n\t\t$minus += $space;\n\t\t$plus += $space;\n\t\tprint \"\\@\\@ -$start_minus,$minus \"\n\t\t\t. \"+\" . ($start_plus - $start_plus_adjust)\n\t\t\t. \",$plus \\@\\@\\n\";\n\t\tprint $current_hunk;\n\t}\n}\n\n# This adds a line to the current hunk and updates $space, $plus or $minus\n\nsub add_line {\n\tmy $line = $_[0];\n\t$current_hunk .= $line;\n\tif ($line =~ /^ /) {\n\t\t$space++;\n\t} elsif ($line =~ /^\\+/) {\n\t\t$plus++;\n\t} elsif ($line =~ /^-/) {\n\t\t$minus++;\n\t} elsif ($line =~ /^\\\\/) {\n\t\t# do nothing\n\t} else {\n\t\tdie \"Unexpected line: $line\";\n\t}\n}\n\n# This function splits the current hunk into the part before the current\n# line, and the part after the current line.\n\nsub skip_line {\n\tmy $line = $_[0];\n\n\tif ($start_minus == 0) {\n\t\t# This patch adds a new file, just ignore that line\n\t\treturn;\n\t} elsif ($start_plus == 0) {\n\t\t# This patch removes a file, so include the line nevertheless\n\t\tadd_line $_;\n\t\treturn;\n\t}\n\n\tflush_hunk;\n\tif ($line =~ /^-/) {\n\t\t$minus++;\n\t} elsif ($line =~ /^\\+/) {\n\t\t$plus++;\n\t\t$start_plus_adjust++;\n\t}\n\tinit_hunk \"@@ -\" . ($start_minus + $minus + $space)\n\t\t. \" +\" . ($start_plus + $plus + $space)\n\t\t. \" @@\\n\";\n}\n\n$simple_keyword = \"Id|Revision|Author|Date|Source|Header\";\n\n# This is the main loop\n\nsub check_file {\n\t$_ = $_[0];\n\t$current_file = $_;\n\t$start_plus_adjust = 0;\n\twhile (<>) {\n\t\tif (/^\\@\\@/) {\n\t\t\tlast;\n\t\t}\n\t\t$current_file .= $_;\n\t}\n\n\tinit_hunk $_;\n\n\t# check hunks\n\twhile (<>) {\n\t\tif ($before_log) {\n\t\t\tif (!/\\+.*\\$Log.*\\$/) {\n\t\t\t\tadd_line $before_log;\n\t\t\t} else {\n\t\t\t\tskip_line $before_log;\n\t\t\t}\n\t\t\t$before_log = '';\n\t\t}\n\n\t\tif ($after_log) {\n\t\t\tif (!/-.*\\$Log.*\\$/) {\n\t\t\t\tadd_line $after_log;\n\t\t\t} else {\n\t\t\t\tskip_line $after_log;\n\t\t\t}\n\t\t\t$after_log = '';\n\t\t}\n\n\t\tif ($skip_logs) {\n\t\t\tif (/^\\+/) {\n\t\t\t\tskip_line $_;\n\t\t\t\t$skip_logs = 1;\n\t\t\t} else {\n\t\t\t\t$skip_logs = 0;\n\t\t\t\tif (/^ *\\*$/) {\n\t\t\t\t\t$after_log = $_;\n\t\t\t\t}\n\t\t\t}\n\t\t} elsif (/^\\+.*\\$($simple_keyword).*\\$/) {\n\t\t\tskip_line $_;\n\t\t} elsif (/^\\@\\@.*/) {\n\t\t\tflush_hunk;\n\t\t\tinit_hunk $_;\n\t\t} elsif (/^diff/) {\n\t\t\tflush_hunk;\n\t\t\treturn;\n\t\t} elsif (/^-.*\\$($simple_keyword).*\\$/) {\n\t\t\t# fake new hunk\n\t\t\tskip_line $_;\n\t\t} elsif (/^\\+ *\\*$/) {\n\t\t\t$before_log = $_;\n\t\t} elsif (/^([- \\+]).*\\$Log.*\\$/) {\n\t\t\tskip_line $_;\n\t\t\t$skip_logs = 1;\n\t\t} else {\n\t\t\tadd_line $_;\n\t\t}\n\t}\n}\n\n# This loop just shows everything before the first diff, and then hands\n# over to check_file whenever it sees a line beginning with \"diff\".\n\nwhile (<>) {\n\tif (/^diff/) {\n\t\tdo {\n\t\t\tcheck_file $_;\n\t\t} while(/^diff/);\n\t\tflush_hunk;\n\t} else {\n\t\tprintf $_;\n\t}\n}\n"},{"id":"48388","messageId":"9e4733910707231743w759afabfvd43045ad2e2eba5a@mail.gmail.com","threadId":"9147","inReplyTo":"Pine.LNX.4.64.0707231933030.14781@racer.site","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-24T00:43:31Z","receivedAt":"2007-07-24T00:43:31Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"Thanks for working on this. I'd like to see it added to git toolkit.\nThis diff is way easier to read now. I can see that it has SPI support\nbackported from some future version.\n\nBut... it still has some problems.\nFor the phytec patch it's not getting the $Log changes in the qlogic\nfiles right.\n\nI'm checking the output diff line by line for problems. It's down to\n11,528 lines from 88,787. That's a lot of junk removed, I'll have to\nmake sure it isn't removing too much.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"48389","messageId":"9e4733910707231806r40d9d427t25853ac65ac756f4@mail.gmail.com","threadId":"9147","inReplyTo":"9e4733910707231743w759afabfvd43045ad2e2eba5a@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-24T01:06:12Z","receivedAt":"2007-07-24T01:06:12Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 7/23/07, Jon Smirl <jonsmirl@gmail.com> wrote:\n> Thanks for working on this. I'd like to see it added to git toolkit.\n> This diff is way easier to read now. I can see that it has SPI support\n> backported from some future version.\n>\n> But... it still has some problems.\n> For the phytec patch it's not getting the $Log changes in the qlogic\n> files right.\n>\n> I'm checking the output diff line by line for problems. It's down to\n> 11,528 lines from 88,787. That's a lot of junk removed, I'll have to\n> make sure it isn't removing too much.\n\nI forgot about newly added files, looks like only 20,000 lines of CVS noise.\n\nThe phytec patch has several ia64 files added to it, those are obvious\nnot part of the support and an ARM9 processor. Never noticed those\nbefore.\n\n\n>\n> --\n> Jon Smirl\n> jonsmirl@gmail.com\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"48390","messageId":"Pine.LNX.4.64.0707240214500.14781@racer.site","threadId":"9147","inReplyTo":"9e4733910707231743w759afabfvd43045ad2e2eba5a@mail.gmail.com","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-24T01:16:21Z","receivedAt":"2007-07-24T01:16:21Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 23 Jul 2007, Jon Smirl wrote:\n\n> Thanks for working on this. I'd like to see it added to git toolkit. \n\nI plan to submit it to patchutils instead, since this is not really \ndependent on git.\n\n> This diff is way easier to read now. I can see that it has SPI support \n> backported from some future version.\n> \n> But... it still has some problems. For the phytec patch it's not getting \n> the $Log changes in the qlogic files right.\n\nYou are correct.  This slipped by my eyes.\n\n> I'm checking the output diff line by line for problems. It's down to\n> 11,528 lines from 88,787. That's a lot of junk removed, I'll have to\n> make sure it isn't removing too much.\n\nMay I ask you to test this version?\n\nThanks,\nDscho\n\n-- snipsnap --\n#!/usr/bin/perl\n\n# This is a simple state machine.\n#\n# There is the state of the current file; its header is stored\n# in $current_file to avoid outputting it when all hunks were\n# culled.  It is only printed before the first hunk, and then\n# set to \"\" to avoid outputting it twice.\n#\n# There are the states of the current hunk, stored in\n# * $current_hunk (possibly modified hunk)\n# * $start_minus, $start_plus (from the original)\n# * $plus, $minus, $space (the current count of the respective lines)\n# a hunk is only printed (in flush_hunk) if any '+' or '-' lines\n# are left after filtering.\n#\n# For $Log..$, there is the state $skip_logs, which is set to 1\n# after seeing such a line, and set to 0 when the first line\n# was seen which does not begin with '+'.\n#\n# A particularly nasty special case is when a single \"*\" was\n# misattributed by the diff to be _inserted before_ a $Log, instead\n# of _appended after_ a $Log.\n# This is the purpose of the $before_log and $after_log variables:\n# if not empty, the state machine expects the next line to begin\n# with '+' or '-', respectively, and followed by a $Log.  If this\n# expectation is not met, the variable is output.\n#\n# The variable $plus_minus_adjust contains the number of lines which\n# were skipped from the \"+\" side, so that the correct offset is shown.\n\n\n# This function gets a hunk header.\n#\n# It initializes the state variables described above\n\nsub init_hunk {\n\tmy $line = $_[0];\n\t$current_hunk = \"\";\n\t($start_minus, $dummy, $start_plus, $dummy) =\n\t\t($line =~ /^\\@\\@ -(\\d+)(,\\d+|) \\+(\\d+)(,\\d+|) \\@\\@/);\n\t$plus = $minus = $space = 0;\n\t$skip_logs = 0;\n\t$before_log = '';\n\t$after_log = '';\n\n\t# we prefer /dev/null as original file name when a file is new\n\tif ($start_minus eq 0) {\n\t\t$current_file =~ s/\\n--- .*\\n/\\n--- \\/dev\\/null\\n/;\n\t} elsif ($start_plus eq 0) {\n\t\t$current_file =~ s/\\n\\+\\+\\+ .*\\n/\\n+++ \\/dev\\/null\\n/;\n\t}\n}\n\n# This function is called whenever there is possibly a hunk to print.\n# Nothing is printed if no '+' or '-' lines are left.\n# Otherwise, if the file header was not yet shown, it does so now.\n\nsub flush_hunk {\n\tif (($plus > 0 || $minus > 0) && $current_hunk ne '') {\n\t\tif ($current_file ne \"\") {\n\t\t\tprint $current_file;\n\t\t\t$current_file = \"\";\n\t\t}\n\t\t$minus += $space;\n\t\t$plus += $space;\n\t\tprint \"\\@\\@ -$start_minus,$minus \"\n\t\t\t. \"+\" . ($start_plus - $start_plus_adjust)\n\t\t\t. \",$plus \\@\\@\\n\";\n\t\tprint $current_hunk;\n\t\t$current_hunk = '';\n\t}\n}\n\n# This adds a line to the current hunk and updates $space, $plus or $minus\n\nsub add_line {\n\tmy $line = $_[0];\n\t$current_hunk .= $line;\n\tif ($line =~ /^ /) {\n\t\t$space++;\n\t} elsif ($line =~ /^\\+/) {\n\t\t$plus++;\n\t} elsif ($line =~ /^-/) {\n\t\t$minus++;\n\t} elsif ($line =~ /^\\\\/) {\n\t\t# do nothing\n\t} else {\n\t\tdie \"Unexpected line: $line\";\n\t}\n}\n\n# This function splits the current hunk into the part before the current\n# line, and the part after the current line.\n\nsub skip_line {\n\tmy $line = $_[0];\n\n\tif ($start_minus == 0) {\n\t\t# This patch adds a new file, just ignore that line\n\t\treturn;\n\t} elsif ($start_plus == 0) {\n\t\t# This patch removes a file, so include the line nevertheless\n\t\tadd_line $_;\n\t\treturn;\n\t}\n\n\tflush_hunk;\n\tif ($line =~ /^-/) {\n\t\t$minus++;\n\t} elsif ($line =~ /^\\+/) {\n\t\t$plus++;\n\t\t$start_plus_adjust++;\n\t}\n\tinit_hunk \"@@ -\" . ($start_minus + $minus + $space)\n\t\t. \" +\" . ($start_plus + $plus + $space)\n\t\t. \" @@\\n\";\n}\n\n$simple_keyword = \"Id|Revision|Author|Date|Source|Header\";\n\n# This is the main loop\n\nsub check_file {\n\t$_ = $_[0];\n\t$current_file = $_;\n\t$start_plus_adjust = 0;\n\twhile (<>) {\n\t\tif (/^\\@\\@/) {\n\t\t\tlast;\n\t\t}\n\t\t$current_file .= $_;\n\t}\n\n\tinit_hunk $_;\n\n\t# check hunks\n\twhile (<>) {\n\t\tif ($before_log) {\n\t\t\tif (!/\\+.*\\$Log.*\\$/) {\n\t\t\t\tadd_line $before_log;\n\t\t\t} else {\n\t\t\t\tskip_line $before_log;\n\t\t\t}\n\t\t\t$before_log = '';\n\t\t}\n\n\t\tif ($after_log) {\n\t\t\tif (!/-.*\\$Log.*\\$/) {\n\t\t\t\tadd_line $after_log;\n\t\t\t} else {\n\t\t\t\tskip_line $after_log;\n\t\t\t}\n\t\t\t$after_log = '';\n\t\t}\n\n\t\tif ($skip_logs) {\n\t\t\tif (/^\\+/) {\n\t\t\t\tskip_line $_;\n\t\t\t\t$skip_logs = 1;\n\t\t\t} else {\n\t\t\t\t$skip_logs = 0;\n\t\t\t\tif (/^ *\\*$/) {\n\t\t\t\t\t$after_log = $_;\n\t\t\t\t}\n\t\t\t}\n\t\t} elsif (/^\\+.*\\$($simple_keyword).*\\$/) {\n\t\t\tskip_line $_;\n\t\t} elsif (/^\\@\\@.*/) {\n\t\t\tflush_hunk;\n\t\t\tinit_hunk $_;\n\t\t} elsif (/^diff/) {\n\t\t\tflush_hunk;\n\t\t\treturn;\n\t\t} elsif (/^-.*\\$($simple_keyword).*\\$/) {\n\t\t\t# fake new hunk\n\t\t\tskip_line $_;\n\t\t} elsif (/^\\+ *\\*$/) {\n\t\t\t$before_log = $_;\n\t\t} elsif (/^([- \\+]).*\\$Log.*\\$/) {\n\t\t\tskip_line $_;\n\t\t\t$skip_logs = 1;\n\t\t} else {\n\t\t\tadd_line $_;\n\t\t}\n\t}\n}\n\n# This loop just shows everything before the first diff, and then hands\n# over to check_file whenever it sees a line beginning with \"diff\".\n\nwhile (<>) {\n\tif (/^diff/) {\n\t\tdo {\n\t\t\tcheck_file $_;\n\t\t} while(/^diff/);\n\t} else {\n\t\tprintf $_;\n\t}\n}\nflush_hunk;\n"},{"id":"48426","messageId":"f84jh8$e27$2@sea.gmane.org","threadId":"9147","inReplyTo":"Pine.LNX.4.64.0707240214500.14781@racer.site","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-07-24T10:16:09Z","receivedAt":"2007-07-24T10:16:09Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"[Cc: Johannes Schindelin <Johannes.Schindelin@gmx.de>, git@vger.kernel.org]\n\nJohannes Schindelin wrote:\n\n> On Mon, 23 Jul 2007, Jon Smirl wrote:\n> \n>> Thanks for working on this. I'd like to see it added to git toolkit. \n> \n> I plan to submit it to patchutils instead, since this is not really \n> dependent on git.\n\nCould you also add it to contrib/ area, please?\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"48444","messageId":"Pine.LNX.4.64.0707241223440.14781@racer.site","threadId":"9147","inReplyTo":"f84jh8$e27$2@sea.gmane.org","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-07-24T11:26:02Z","receivedAt":"2007-07-24T11:26:02Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 24 Jul 2007, Jakub Narebski wrote:\n\n> [Cc: Johannes Schindelin <Johannes.Schindelin@gmx.de>, git@vger.kernel.org]\n\nThanks.\n\n> Johannes Schindelin wrote:\n> \n> > On Mon, 23 Jul 2007, Jon Smirl wrote:\n> > \n> >> Thanks for working on this. I'd like to see it added to git toolkit. \n> > \n> > I plan to submit it to patchutils instead, since this is not really \n> > dependent on git.\n> \n> Could you also add it to contrib/ area, please?\n\nSure.  Once it is really all fleshed out.  But I really think that \npatchutils is a better place.  That way, the script also helps the poor \nsouls stuck with Mercurial or Darcs :-)\n\nThe really tedious part is the testing, and the verifying.  Fortunately, \nJon made up for my incapability in both areas, with an incredible \npatience.\n\nCiao,\nDscho\n"},{"id":"48455","messageId":"9e4733910707240608gd201d5bv3d7002d9a9c186f6@mail.gmail.com","threadId":"9147","inReplyTo":"Pine.LNX.4.64.0707241223440.14781@racer.site","subject":"Re: Git help for kernel archeology, suppress diffs caused by CVS keyword expansion","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-07-24T13:08:23Z","receivedAt":"2007-07-24T13:08:23Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 7/24/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> The really tedious part is the testing, and the verifying.  Fortunately,\n> Jon made up for my incapability in both areas, with an incredible\n> patience.\n\nI haven't finished checking every last line yet, if anything is left\nit is small. We lost power here all evening last night. That's not\nsupposed to happen in urban Boston. It was black over a mile from me\nin all directions.\n\n>\n> Ciao,\n> Dscho\n>\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"}]}