{"thread":{"id":"19958","subject":"git mailinfo strips important context from patch subjects","startedAt":"2009-06-28T19:38:58Z","lastAt":"2009-09-23T00:26:36Z","messageCount":21,"participants":["Roger Leigh","Jeff King","Paolo Bonzini","Junio C Hamano","Andreas Ericsson","Jakub Narebski","Neil Roberts","Jason Holden"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"117139","messageId":"20090628193858.GA29467@codelibre.net","threadId":"19958","inReplyTo":null,"subject":"git mailinfo strips important context from patch subjects","fromName":"Roger Leigh","fromEmail":"rleigh@codelibre.net","sentAt":"2009-06-28T19:38:58Z","receivedAt":"2009-06-28T19:38:58Z","isPatch":false,"sender":{"key":"rleigh@codelibre.net","avatar":null},"body":"[I'm not currently subscribed to the list; I'd appreciate a CC\non any replies, thanks!]\n\nHi,\n\nIn most of the projects I work on, the git commit message has\nthe affected subsystem or component in square brackets, such as\n\n  [foo] change bar to baz\n\nFor example, with a single patch from a series produced by\ngit format-patch:\n\n% head -n4 /tmp/patches/0005-sbuild-chroot_mountable-Don-t-derive-from-chroot.patch\nFrom f01579584f1e7d77cf1e9c3306601a4cccff8c55 Mon Sep 17 00:00:00 2001\nFrom: Roger Leigh <rleigh@debian.org>\nDate: Fri, 10 Apr 2009 19:43:15 +0100\nSubject: [PATCH 05/15] [sbuild] chroot_mountable: Don't derive from chroot\n\n% git mailinfo </tmp/patches/0005-sbuild-chroot_mountable-Don-t-derive-from-chroot.patch /dev/null /dev/null\nAuthor: Roger Leigh\nEmail: rleigh@debian.org\nSubject: chroot_mountable: Don't derive from chroot\nDate: Fri, 10 Apr 2009 19:43:15 +0100\n\nThe [sbuild] prefix has been dropped from the Subject, so an\nimportant bit of context about the patch has been lost.\n\nIt's a bit of a bug that you can't round trip from a git-format-patch\nto import with git-am and then not be able to produce the exact same\npatch set with git-format-patch again (assuming preparing and applying\nto the same point, of course).\n\nWould it be possible to change the git-mailinfo logic to use a less\ngreedy pattern match so it leaves everything after\n([PATCH( [0-9/])+])+ in the subject?  AFAICT this is cleanup_subject in\nbuiltin-mailinfo.c?  Could this rather complex function not just do a\nsimple regex match which can also take care of stripping ([Rr]e:) ?\n\n\nThanks,\nRoger\n\n-- \n  .''`.  Roger Leigh\n : :' :  Debian GNU/Linux             http://people.debian.org/~rleigh/\n `. `'   Printing on GNU/Linux?       http://gutenprint.sourceforge.net/\n   `-    GPG Public Key: 0x25BFB848   Please GPG sign your mail.\n"},{"id":"117140","messageId":"20090628200259.GB8828@sigio.peff.net","threadId":"19958","inReplyTo":"20090628193858.GA29467@codelibre.net","subject":"Re: git mailinfo strips important context from patch subjects","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-06-28T20:02:59Z","receivedAt":"2009-06-28T20:02:59Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Jun 28, 2009 at 08:38:58PM +0100, Roger Leigh wrote:\n\n> In most of the projects I work on, the git commit message has\n> the affected subsystem or component in square brackets, such as\n> \n>   [foo] change bar to baz\n>\n> [...]\n>\n> The [sbuild] prefix has been dropped from the Subject, so an\n> important bit of context about the patch has been lost.\n> \n> It's a bit of a bug that you can't round trip from a git-format-patch\n> to import with git-am and then not be able to produce the exact same\n> patch set with git-format-patch again (assuming preparing and applying\n> to the same point, of course).\n\nAs an immediate solution, you probably want to use \"-k\" when generating\nthe patch (not to add the [PATCH] munging) and \"-k\" when reading the\npatch via \"git am\" (which will avoid trying to strip any munging).\n\nHowever:\n\n> Would it be possible to change the git-mailinfo logic to use a less\n> greedy pattern match so it leaves everything after\n> ([PATCH( [0-9/])+])+ in the subject?  AFAICT this is cleanup_subject in\n> builtin-mailinfo.c?  Could this rather complex function not just do a\n> simple regex match which can also take care of stripping ([Rr]e:) ?\n\nYes, I think in the long run it makes sense to strip just the _first_\nset of brackets. I don't think we want to be more specific than that in\nthe match, because we allow arbitrary cruft inside the brackets (like\n\"[RFC/PATCH]\", etc). But if format-patch always puts exactly one set of\nbrackets, and am strips exactly one set, then that should retain your\nsubject in practice, even if it starts with [foo].\n\n-Peff\n"},{"id":"117142","messageId":"1246219664-11000-1-git-send-email-bonzini@gnu.org","threadId":"19958","inReplyTo":"20090628193858.GA29467@codelibre.net","subject":"[PATCH] git mailinfo strips important context from patch subjects","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2009-06-28T20:07:44Z","receivedAt":"2009-06-28T20:07:44Z","isPatch":true,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"> Would it be possible to change the git-mailinfo logic to use a less\n> greedy pattern match?\n\nLike this?  (I also simplified the first part of the if condition since I\nwas at it).  Anyone, feel free to resubmit it as a proper patch.\n\nAlmost-Signed-off-by: Paolo Bonzini <bonzini@gnu.org>\n---\n builtin-mailinfo.c |    3 ++-\n 1 files changed, 2 insertions(+), 1 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 92637ac..d340ae6 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -237,7 +237,8 @@ static void cleanup_subject(struct strbuf *subject)\n \t\tcase '[':\n \t\t\tif ((pos = strchr(subject->buf, ']'))) {\n \t\t\t\tremove = pos - subject->buf;\n-\t\t\t\tif (remove <= (subject->len - remove) * 2) {\n+\t\t\t\tif (remove <= subject->len * 2 / 3\n+\t\t\t\t    && memmem(subject->buf, remove, 'PATCH', 5)) {\n \t\t\t\t\tstrbuf_remove(subject, 0, remove + 1);\n \t\t\t\t\tcontinue;\n \t\t\t\t}\n-- \n1.6.0.3\n"},{"id":"117148","messageId":"7vfxdkez96.fsf@alter.siamese.dyndns.org","threadId":"19958","inReplyTo":"20090628200259.GB8828@sigio.peff.net","subject":"Re: git mailinfo strips important context from patch subjects","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-06-28T23:04:37Z","receivedAt":"2009-06-28T23:04:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Sun, Jun 28, 2009 at 08:38:58PM +0100, Roger Leigh wrote:\n>\n>> In most of the projects I work on, the git commit message has\n>> the affected subsystem or component in square brackets, such as\n>> \n>>   [foo] change bar to baz\n>>\n>> [...]\n>>\n>> The [sbuild] prefix has been dropped from the Subject, so an\n>> important bit of context about the patch has been lost.\n>> \n>> It's a bit of a bug that you can't round trip from a git-format-patch\n>> to import with git-am and then not be able to produce the exact same\n>> patch set with git-format-patch again (assuming preparing and applying\n>> to the same point, of course).\n>\n> As an immediate solution, you probably want to use \"-k\" when generating\n> the patch (not to add the [PATCH] munging) and \"-k\" when reading the\n> patch via \"git am\" (which will avoid trying to strip any munging).\n>\n> However:\n>\n>> Would it be possible to change the git-mailinfo logic to use a less\n>> greedy pattern match so it leaves everything after\n>> ([PATCH( [0-9/])+])+ in the subject?  AFAICT this is cleanup_subject in\n>> builtin-mailinfo.c?  Could this rather complex function not just do a\n>> simple regex match which can also take care of stripping ([Rr]e:) ?\n>\n> Yes, I think in the long run it makes sense to strip just the _first_\n> set of brackets. I don't think we want to be more specific than that in\n> the match, because we allow arbitrary cruft inside the brackets (like\n> \"[RFC/PATCH]\", etc). But if format-patch always puts exactly one set of\n> brackets, and am strips exactly one set, then that should retain your\n> subject in practice, even if it starts with [foo].\n\nI think it may still make sense to insist that PATCH appears somewhere in\nthe first set of brackets, but I have stop and wonder if it is even\nnecessary.\n\nBecause git removes [sbuild] at the beginning, Roger is unhappy.\n\n * Is he happy that git removes [PATCH]?  In E-mail based workflow it is\n   a good practice to mark messages that are patches clearly so that they\n   can be quickly found among the discussions that lead to them, and it is\n   plausible that his project accepted that as an established practice\n   supported well by git.\n\n * Is he happy that git treats the first paragraph of the commit message\n   specially from the rest of the message?  In a project with many\n   commits, it is essential that people write good commit summaries that\n   fits on a single line so that tools like shortlog and gitweb can be\n   used to get a bird-eye view of what happened recently.  Perhaps his\n   project picked it up as the best current practice supported well by\n   git.\n\n * Is he happy that git takes \"---\" as the end of message marker, so that\n   any other commentary can be added to the message to facilitate the\n   communication without adding noise to the commits?  Perhaps he is and\n   his project picked it up as a good practice supported well by git.\n\nThere are many other conventions in git that does not have anything to do\nwith what the underlying git datastructure supports, but conventions can\nalways be seen as \"don't do that, instead do it this way\", limitations,\nand to some of them Roger may not be happy.  Where would we draw a line?\n\n_An_ established (note that I did not say _the_ nor _best current_)\npractice supported well by git to note the area being affected in a\nproject of nontrivial size is to prefix the single line summary with the\nname of the area followed by a colon.  There is no difference between\n\"[sbuild] foo\" and \"sbuild: foo\" at the information content point-of-view,\nbut the latter has an advantage of being one letter shorter and less\ndistracting in MUA.  He does not have a very strong reason to choose\nsomething different only to make his life harder, does he?\n\nUsers can take advantage of this established practice when running\nshortlog with \"--grep=^area:\" to limit the birds-eye-view to a specific\narea.  If this turns out to be useful, we could even add an option to \"git\nlog --area=name\" that limits this kind of match to the first paragraph of\nthe commit log message, for example.\n\nSupporting a slightly different convention may seem to be accomodating and\nnice, but if there is no real technical difference between the two (and\nagain, \"area:\" is one letter shorter ;-), letting people run with\ndifferent convention longer, when they can switch easily to another\nconvention that is already well supported, may actually hurt them in the\nlong run.  \"[sbuild]\" will not match \"--area=sbuild\" that will internally\nbecome \"--grep-only-first-line=sbuild:\" so either he will miss out\nbenefiting from the new feature, or the implementation of the new feature\nunnecessarily needs more code.\n\nIt is not about discouraging a wrong workflow or practice, because there\nis nothing _wrong_ per-se in [sbuild] prefix.  It is just that it makes\nthings harder in the long run.  In this particular case, it is only very\nslightly harder, but these things tend to add up from different fronts.\n"},{"id":"117151","messageId":"4A48870B.5050802@op5.se","threadId":"19958","inReplyTo":"1246219664-11000-1-git-send-email-bonzini@gnu.org","subject":"Re: [PATCH] git mailinfo strips important context from patch subjects","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2009-06-29T09:19:07Z","receivedAt":"2009-06-29T09:19:07Z","isPatch":true,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Paolo Bonzini wrote:\n>> Would it be possible to change the git-mailinfo logic to use a less\n>> greedy pattern match?\n> \n> Like this?  (I also simplified the first part of the if condition since I\n> was at it).  Anyone, feel free to resubmit it as a proper patch.\n> \n> Almost-Signed-off-by: Paolo Bonzini <bonzini@gnu.org>\n> ---\n>  builtin-mailinfo.c |    3 ++-\n>  1 files changed, 2 insertions(+), 1 deletions(-)\n> \n> diff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\n> index 92637ac..d340ae6 100644\n> --- a/builtin-mailinfo.c\n> +++ b/builtin-mailinfo.c\n> @@ -237,7 +237,8 @@ static void cleanup_subject(struct strbuf *subject)\n>  \t\tcase '[':\n>  \t\t\tif ((pos = strchr(subject->buf, ']'))) {\n>  \t\t\t\tremove = pos - subject->buf;\n> -\t\t\t\tif (remove <= (subject->len - remove) * 2) {\n> +\t\t\t\tif (remove <= subject->len * 2 / 3\n> +\t\t\t\t    && memmem(subject->buf, remove, 'PATCH', 5)) {\n>  \t\t\t\t\tstrbuf_remove(subject, 0, remove + 1);\n>  \t\t\t\t\tcontinue;\n>  \t\t\t\t}\n\n\nPardon my ignorance, but wouldn't this still remove not only\n\"[PATCH 4/5]\", but all of [PATCH 4/5] [sbuild]\" anyway? The\nparameters to strbuf_remove() seem unchanged.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n\nConsidering the successes of the wars on alcohol, poverty, drugs and\nterror, I think we should give some serious thought to declaring war\non peace.\n"},{"id":"117153","messageId":"4A488F07.10002@op5.se","threadId":"19958","inReplyTo":"7vfxdkez96.fsf@alter.siamese.dyndns.org","subject":"Re: git mailinfo strips important context from patch subjects","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2009-06-29T09:53:11Z","receivedAt":"2009-06-29T09:53:11Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> Jeff King <peff@peff.net> writes:\n> \n>> On Sun, Jun 28, 2009 at 08:38:58PM +0100, Roger Leigh wrote:\n>>\n>>> In most of the projects I work on, the git commit message has\n>>> the affected subsystem or component in square brackets, such as\n>>>\n>>>   [foo] change bar to baz\n>>>\n>>> [...]\n>>>\n>>> The [sbuild] prefix has been dropped from the Subject, so an\n>>> important bit of context about the patch has been lost.\n>>>\n>>> It's a bit of a bug that you can't round trip from a git-format-patch\n>>> to import with git-am and then not be able to produce the exact same\n>>> patch set with git-format-patch again (assuming preparing and applying\n>>> to the same point, of course).\n>> As an immediate solution, you probably want to use \"-k\" when generating\n>> the patch (not to add the [PATCH] munging) and \"-k\" when reading the\n>> patch via \"git am\" (which will avoid trying to strip any munging).\n>>\n>> However:\n>>\n>>> Would it be possible to change the git-mailinfo logic to use a less\n>>> greedy pattern match so it leaves everything after\n>>> ([PATCH( [0-9/])+])+ in the subject?  AFAICT this is cleanup_subject in\n>>> builtin-mailinfo.c?  Could this rather complex function not just do a\n>>> simple regex match which can also take care of stripping ([Rr]e:) ?\n>> Yes, I think in the long run it makes sense to strip just the _first_\n>> set of brackets. I don't think we want to be more specific than that in\n>> the match, because we allow arbitrary cruft inside the brackets (like\n>> \"[RFC/PATCH]\", etc). But if format-patch always puts exactly one set of\n>> brackets, and am strips exactly one set, then that should retain your\n>> subject in practice, even if it starts with [foo].\n> \n> I think it may still make sense to insist that PATCH appears somewhere in\n> the first set of brackets, but I have stop and wonder if it is even\n> necessary.\n> \n> Because git removes [sbuild] at the beginning, Roger is unhappy.\n> \n\n[ and a lot more ]\n\n> \n> _An_ established (note that I did not say _the_ nor _best current_)\n> practice supported well by git to note the area being affected in a\n> project of nontrivial size is to prefix the single line summary with the\n> name of the area followed by a colon.  There is no difference between\n> \"[sbuild] foo\" and \"sbuild: foo\" at the information content point-of-view,\n> but the latter has an advantage of being one letter shorter and less\n> distracting in MUA.  He does not have a very strong reason to choose\n> something different only to make his life harder, does he?\n> \n\nTrue, but it seems wrong to have am remove more of the subject than\nformat-patch prepends. Imagine a commit subject looking like this:\n  \"Allow [ and ] in the blurble.foostuff table\".\n\nShould am strip the subject all the way up to the last ']'? I think\nnot, and I'd be very vexed if it did.\n\n> Users can take advantage of this established practice when running\n> shortlog with \"--grep=^area:\" to limit the birds-eye-view to a specific\n> area.  If this turns out to be useful, we could even add an option to \"git\n> log --area=name\" that limits this kind of match to the first paragraph of\n> the commit log message, for example.\n> \n> Supporting a slightly different convention may seem to be accomodating and\n> nice, but if there is no real technical difference between the two (and\n> again, \"area:\" is one letter shorter ;-), letting people run with\n> different convention longer, when they can switch easily to another\n> convention that is already well supported, may actually hurt them in the\n> long run.  \"[sbuild]\" will not match \"--area=sbuild\" that will internally\n> become \"--grep-only-first-line=sbuild:\" so either he will miss out\n> benefiting from the new feature, or the implementation of the new feature\n> unnecessarily needs more code.\n> \n> It is not about discouraging a wrong workflow or practice, because there\n> is nothing _wrong_ per-se in [sbuild] prefix.  It is just that it makes\n> things harder in the long run.  In this particular case, it is only very\n> slightly harder, but these things tend to add up from different fronts.\n\nAgreed, but there are valid use-cases orthogonal to subsystem naming to\nplace [] in the patch subject. I still feel that since format-patch only\nadds one set, am (mailinfo) should really only remove one set, too. It's\nwhat makes sense, really.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n\nConsidering the successes of the wars on alcohol, poverty, drugs and\nterror, I think we should give some serious thought to declaring war\non peace.\n"},{"id":"117154","messageId":"1246269351-26929-1-git-send-email-ae@op5.se","threadId":"19958","inReplyTo":"4A488F07.10002@op5.se","subject":"[PATCH] mailinfo: Remove only one set of square brackets","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2009-06-29T09:55:51Z","receivedAt":"2009-06-29T09:55:51Z","isPatch":true,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"git-format-patch prepends patches with a [PATCH x/n] prefix, but\nmailinfo used to remove any number of square-bracket pairs and\nthe content between them. This prevents one from using a commit\nsubject like this:\n\n  [ and ] must be allowed as input\n\nRemoving the square bracket pair from this rather clumsily\nconstructed subject line loses important information, so we must\ntake care not to.\n\nThis patch causes the subject stripping to stop after it has\nencountered one pair of square brackets.\n\nOne possible downside of this patch is that the patch-handling\nprograms will now fail at removing author-added square-brackets\nto be removed, such as\n\n  [RFC][PATCH x/n]\n\nHowever, since format-patch only adds one set of square brackets,\nthis behaviour is quite easily undesrstood and defended while the\nprevious behaviour is not.\n\nSigned-off-by: Andreas Ericsson <ae@op5.se>\n---\n builtin-mailinfo.c |    7 +++++++\n 1 files changed, 7 insertions(+), 0 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 92637ac..fb5ad70 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -221,6 +221,8 @@ static void cleanup_subject(struct strbuf *subject)\n {\n \tchar *pos;\n \tsize_t remove;\n+\tint brackets_removed = 0;\n+\n \twhile (subject->len) {\n \t\tswitch (*subject->buf) {\n \t\tcase 'r': case 'R':\n@@ -235,10 +237,15 @@ static void cleanup_subject(struct strbuf *subject)\n \t\t\tstrbuf_remove(subject, 0, 1);\n \t\t\tcontinue;\n \t\tcase '[':\n+\t\t\t/* remove only one set of square brackets */\n+\t\t\tif (brackets_removed)\n+\t\t\t\tbreak;\n+\n \t\t\tif ((pos = strchr(subject->buf, ']'))) {\n \t\t\t\tremove = pos - subject->buf;\n \t\t\t\tif (remove <= (subject->len - remove) * 2) {\n \t\t\t\t\tstrbuf_remove(subject, 0, remove + 1);\n+\t\t\t\t\tbrackets_removed = 1;\n \t\t\t\t\tcontinue;\n \t\t\t\t}\n \t\t\t} else\n-- \n1.6.3.3.354.gfb24\n"},{"id":"117155","messageId":"4A48959A.3060404@gmail.com","threadId":"19958","inReplyTo":"4A48870B.5050802@op5.se","subject":"Re: [PATCH] git mailinfo strips important context from patch subjects","fromName":"Paolo Bonzini","fromEmail":"paolo.bonzini@gmail.com","sentAt":"2009-06-29T10:21:14Z","receivedAt":"2009-06-29T10:21:14Z","isPatch":true,"sender":{"key":"paolo.bonzini@gmail.com","avatar":"https://gravatar.com/avatar/7817ef2e168b4ef0570c5bb5bdc1d4b44f34d3075fe32b871710dd942d0a89f5?d=mp&s=160"},"body":"\n>> case '[':\n>> if ((pos = strchr(subject->buf, ']'))) {\n>> remove = pos - subject->buf;\n>> - if (remove <= (subject->len - remove) * 2) {\n>> + if (remove <= subject->len * 2 / 3\n>> + && memmem(subject->buf, remove, 'PATCH', 5)) {\n>> strbuf_remove(subject, 0, remove + 1);\n>> continue;\n>> }\n>\n>\n> Pardon my ignorance, but wouldn't this still remove not only\n> \"[PATCH 4/5]\", but all of [PATCH 4/5] [sbuild]\" anyway? The\n> parameters to strbuf_remove() seem unchanged.\n\nI don't exclude I've screwed up, but note that pos is computed with \nstrchr, not strrchr.  Since the second memmem does not find [PATCH], it \ndoes not remove anything.\n\n(BTW, cairo uses the [...] convention).\n\nPaolo\n"},{"id":"117156","messageId":"4A489D5C.2000406@op5.se","threadId":"19958","inReplyTo":"4A48959A.3060404@gmail.com","subject":"Re: [PATCH] git mailinfo strips important context from patch subjects","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2009-06-29T10:54:20Z","receivedAt":"2009-06-29T10:54:20Z","isPatch":true,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Paolo Bonzini wrote:\n> \n>>> case '[':\n>>> if ((pos = strchr(subject->buf, ']'))) {\n>>> remove = pos - subject->buf;\n>>> - if (remove <= (subject->len - remove) * 2) {\n>>> + if (remove <= subject->len * 2 / 3\n>>> + && memmem(subject->buf, remove, 'PATCH', 5)) {\n>>> strbuf_remove(subject, 0, remove + 1);\n>>> continue;\n>>> }\n>>\n>>\n>> Pardon my ignorance, but wouldn't this still remove not only\n>> \"[PATCH 4/5]\", but all of [PATCH 4/5] [sbuild]\" anyway? The\n>> parameters to strbuf_remove() seem unchanged.\n> \n> I don't exclude I've screwed up, but note that pos is computed with \n> strchr, not strrchr.  Since the second memmem does not find [PATCH], it \n> does not remove anything.\n> \n\nIt removes one character, which means the subject still gets mangled. If\nit *doesn't* remove one character and also doesn't break out of the loop,\nit'll loop indefinitely, since *subject->buf will never change.\n\nThere's something else wrong with your patch though, as mailinfo dumps\ncore with it for a patch starting with \"[PATCH] [git]\". It happens in\nmemmem(). Here's the backtrace:\n\n(gdb) bt\n#0  0x00c67c76 in memmem (haystack_start=0x8a4bae0, haystack_len=6, \n    needle_start=0x41544348, needle_len=5) at memmem.c:66\n#1  0x0807a1e5 in cleanup_subject () at builtin-mailinfo.c:240\n#2  handle_info () at builtin-mailinfo.c:878\n#3  mailinfo () at builtin-mailinfo.c:929\n#4  cmd_mailinfo (argc=4, argv=<value optimized out>, prefix=0x0)\n    at builtin-mailinfo.c:966\n#5  0x0804b0f7 in run_builtin () at git.c:247\n#6  handle_internal_command (argc=4, argv=0xbfb00f58) at git.c:393\n#7  0x0804b2e2 in run_argv () at git.c:439\n#8  main (argc=4, argv=0xbfb00f58) at git.c:510\n\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n\nConsidering the successes of the wars on alcohol, poverty, drugs and\nterror, I think we should give some serious thought to declaring war\non peace.\n"},{"id":"117165","messageId":"7v8wjbc98d.fsf@alter.siamese.dyndns.org","threadId":"19958","inReplyTo":"1246269351-26929-1-git-send-email-ae@op5.se","subject":"Re: [PATCH] mailinfo: Remove only one set of square brackets","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-06-29T16:09:38Z","receivedAt":"2009-06-29T16:09:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Andreas Ericsson <ae@op5.se> writes:\n\n> git-format-patch prepends patches with a [PATCH x/n] prefix, but\n> mailinfo used to remove any number of square-bracket pairs and\n> the content between them. This prevents one from using a commit\n> subject like this:\n>\n>   [ and ] must be allowed as input\n>\n> Removing the square bracket pair from this rather clumsily\n> constructed subject line loses important information, so we must\n> take care not to.\n>\n> This patch causes the subject stripping to stop after it has\n> encountered one pair of square brackets.\n>\n> One possible downside of this patch is that the patch-handling\n> programs will now fail at removing author-added square-brackets\n> to be removed, such as\n>\n>   [RFC][PATCH x/n]\n>\n> However, since format-patch only adds one set of square brackets,\n> this behaviour is quite easily undesrstood and defended while the\n> previous behaviour is not.\n>\n> Signed-off-by: Andreas Ericsson <ae@op5.se>\n> ---\n\nAll good points, and I like this one, including its Subject: line.\n\n>  builtin-mailinfo.c |    7 +++++++\n>  1 files changed, 7 insertions(+), 0 deletions(-)\n>\n> diff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\n> index 92637ac..fb5ad70 100644\n> --- a/builtin-mailinfo.c\n> +++ b/builtin-mailinfo.c\n> @@ -221,6 +221,8 @@ static void cleanup_subject(struct strbuf *subject)\n>  {\n>  \tchar *pos;\n>  \tsize_t remove;\n> +\tint brackets_removed = 0;\n> +\n>  \twhile (subject->len) {\n>  \t\tswitch (*subject->buf) {\n>  \t\tcase 'r': case 'R':\n> @@ -235,10 +237,15 @@ static void cleanup_subject(struct strbuf *subject)\n>  \t\t\tstrbuf_remove(subject, 0, 1);\n>  \t\t\tcontinue;\n>  \t\tcase '[':\n> +\t\t\t/* remove only one set of square brackets */\n> +\t\t\tif (brackets_removed)\n> +\t\t\t\tbreak;\n> +\n>  \t\t\tif ((pos = strchr(subject->buf, ']'))) {\n>  \t\t\t\tremove = pos - subject->buf;\n>  \t\t\t\tif (remove <= (subject->len - remove) * 2) {\n>  \t\t\t\t\tstrbuf_remove(subject, 0, remove + 1);\n> +\t\t\t\t\tbrackets_removed = 1;\n>  \t\t\t\t\tcontinue;\n>  \t\t\t\t}\n>  \t\t\t} else\n> -- \n> 1.6.3.3.354.gfb24\n"},{"id":"117178","messageId":"1246310220-16909-1-git-send-email-rleigh@debian.org","threadId":"19958","inReplyTo":"7vfxdkez96.fsf@alter.siamese.dyndns.org","subject":"[PATCH] builtin-mailinfo.c: Trim only first pair of square brackets in subject","fromName":"Roger Leigh","fromEmail":"rleigh@debian.org","sentAt":"2009-06-29T21:17:00Z","receivedAt":"2009-06-29T21:17:00Z","isPatch":true,"sender":{"key":"rleigh@debian.org","avatar":null},"body":"Use a regular expression to match text after \"Re:\" or any text in the\nfirst pair of square brackets such as \"[PATCH n/m]\".  This replaces\nthe complex hairy string munging with a simple single  pattern match.\n\nSigned-off-by: Roger Leigh <rleigh@debian.org>\n---\n builtin-mailinfo.c |   61 +++++++++++++++++++++++++++++-----------------------\n 1 files changed, 34 insertions(+), 27 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 92637ac..6d19046 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -219,35 +219,42 @@ static int is_multipart_boundary(const struct strbuf *line)\n \n static void cleanup_subject(struct strbuf *subject)\n {\n-\tchar *pos;\n-\tsize_t remove;\n-\twhile (subject->len) {\n-\t\tswitch (*subject->buf) {\n-\t\tcase 'r': case 'R':\n-\t\t\tif (subject->len <= 3)\n-\t\t\t\tbreak;\n-\t\t\tif (!memcmp(subject->buf + 1, \"e:\", 2)) {\n-\t\t\t\tstrbuf_remove(subject, 0, 3);\n-\t\t\t\tcontinue;\n-\t\t\t}\n-\t\t\tbreak;\n-\t\tcase ' ': case '\\t': case ':':\n-\t\t\tstrbuf_remove(subject, 0, 1);\n-\t\t\tcontinue;\n-\t\tcase '[':\n-\t\t\tif ((pos = strchr(subject->buf, ']'))) {\n-\t\t\t\tremove = pos - subject->buf;\n-\t\t\t\tif (remove <= (subject->len - remove) * 2) {\n-\t\t\t\t\tstrbuf_remove(subject, 0, remove + 1);\n-\t\t\t\t\tcontinue;\n-\t\t\t\t}\n-\t\t\t} else\n-\t\t\t\tstrbuf_remove(subject, 0, 1);\n-\t\t\tbreak;\n-\t\t}\n+\tint status;\n+\tregex_t regex;\n+\tregmatch_t match[4];\n+\n+\t/* Strip off 'Re:' and/or the first text in square brackets, such as\n+\t   '[PATCH]' at the start of the mail Subject. */\n+\tstatus = regcomp(&regex,\n+\t\t\t \"^([Rr]e:)?([^]]*\\\\[[^]]+\\\\])(.*)$\",\n+\t\t\t REG_EXTENDED);\n+\n+\tif (status) {\n+\t\t/* Compiling the regex failed.  Find out why and tell\n+\t\t   the user.  This is always a bug in the code. */\n+\t\tint esize = regerror(status, &regex, NULL, 0);\n+\t\tstruct strbuf etext = STRBUF_INIT;\n+\n+\t\tstrbuf_grow(&etext, esize);\n+\t\tregerror(status, &regex, etext.buf, esize);\n+\t\tfprintf (stderr,\n+\t\t\t \"Error compiling regular expression: %s\\n\",\n+\t\t\t etext.buf);\n+\t\tstrbuf_release(&etext);\n+\t\texit(1);\n+\t}\n+\n+\t/* Store any matches in match. */\n+\tstatus = regexec(&regex, subject->buf, 4, match, 0);\n+\n+\t/* If there was a match for \\3 in the regex, trim the subject\n+\t   to this match. */\n+\tif (!status && match[3].rm_so > 0) {\n+\t\tstrbuf_remove(subject, 0, match[3].rm_so);\n \t\tstrbuf_trim(subject);\n-\t\treturn;\n \t}\n+\n+\treturn;\n }\n \n static void cleanup_space(struct strbuf *sb)\n-- \n1.6.3.3\n"},{"id":"117179","messageId":"m3ljnawx3h.fsf@localhost.localdomain","threadId":"19958","inReplyTo":"1246310220-16909-1-git-send-email-rleigh@debian.org","subject":"Re: [PATCH] builtin-mailinfo.c: Trim only first pair of square brackets in subject","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-06-29T21:26:45Z","receivedAt":"2009-06-29T21:26:45Z","isPatch":true,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Roger Leigh <rleigh@debian.org> writes:\n\n> Use a regular expression to match text after \"Re:\" or any text in the\n> first pair of square brackets such as \"[PATCH n/m]\".  This replaces\n> the complex hairy string munging with a simple single  pattern match.\n\n[...]\n> +\t/* Strip off 'Re:' and/or the first text in square brackets, such as\n> +\t   '[PATCH]' at the start of the mail Subject. */\n> +\tstatus = regcomp(&regex,\n> +\t\t\t \"^([Rr]e:)?([^]]*\\\\[[^]]+\\\\])(.*)$\",\n> +\t\t\t REG_EXTENDED);\n\nSidenote: it probably didn't worked before either, but there are some\nbroken mail readers in the wold (*cough* MS Outlook *cough*), that\nmisinterpret RFCs and use translated form of \"Re:\" e.g. \"Odp:\" (Polish),\nor not strip \"Re:\" when replying resulting in string of \"Re: Re: Re: ...\",\nor use capitalized form of \"Re:\", i.e. \"RE:\", or use yet another form \ne.g. compact form of repeated \"Re: Re: Re: ...\" in form of \"Re(3):\".\n\nBut I guess it didn't worked before either.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"117182","messageId":"1246311269-18120-1-git-send-email-rleigh@debian.org","threadId":"19958","inReplyTo":"7vfxdkez96.fsf@alter.siamese.dyndns.org","subject":"[PATCH 2/2] builtin-mailinfo.c: Free regular expression after use","fromName":"Roger Leigh","fromEmail":"rleigh@debian.org","sentAt":"2009-06-29T21:34:29Z","receivedAt":"2009-06-29T21:34:29Z","isPatch":true,"sender":{"key":"rleigh@debian.org","avatar":null},"body":"Signed-off-by: Roger Leigh <rleigh@debian.org>\n---\n builtin-mailinfo.c |    2 ++\n 1 files changed, 2 insertions(+), 0 deletions(-)\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 6d19046..6559c37 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -254,6 +254,8 @@ static void cleanup_subject(struct strbuf *subject)\n \t\tstrbuf_trim(subject);\n \t}\n \n+\tregfree(&regex);\n+\n \treturn;\n }\n \n-- \n1.6.3.3\n"},{"id":"117183","messageId":"20090629213625.GA5397@codelibre.net","threadId":"19958","inReplyTo":"7vfxdkez96.fsf@alter.siamese.dyndns.org","subject":"Re: git mailinfo strips important context from patch subjects","fromName":"Roger Leigh","fromEmail":"rleigh@codelibre.net","sentAt":"2009-06-29T21:36:26Z","receivedAt":"2009-06-29T21:36:26Z","isPatch":false,"sender":{"key":"rleigh@codelibre.net","avatar":null},"body":"On Sun, Jun 28, 2009 at 04:04:37PM -0700, Junio C Hamano wrote:\n> Jeff King <peff@peff.net> writes:\n> \n> > On Sun, Jun 28, 2009 at 08:38:58PM +0100, Roger Leigh wrote:\n> >\n> >> In most of the projects I work on, the git commit message has\n> >> the affected subsystem or component in square brackets, such as\n> >> \n> >>   [foo] change bar to baz\n> >>\n> >> [...]\n> >>\n> >> The [sbuild] prefix has been dropped from the Subject, so an\n> >> important bit of context about the patch has been lost.\n> >> \n> >> It's a bit of a bug that you can't round trip from a git-format-patch\n> >> to import with git-am and then not be able to produce the exact same\n> >> patch set with git-format-patch again (assuming preparing and applying\n> >> to the same point, of course).\n> >\n> > As an immediate solution, you probably want to use \"-k\" when generating\n> > the patch (not to add the [PATCH] munging) and \"-k\" when reading the\n> > patch via \"git am\" (which will avoid trying to strip any munging).\n> >\n> > However:\n> >\n> >> Would it be possible to change the git-mailinfo logic to use a less\n> >> greedy pattern match so it leaves everything after\n> >> ([PATCH( [0-9/])+])+ in the subject?  AFAICT this is cleanup_subject in\n> >> builtin-mailinfo.c?  Could this rather complex function not just do a\n> >> simple regex match which can also take care of stripping ([Rr]e:) ?\n> >\n> > Yes, I think in the long run it makes sense to strip just the _first_\n> > set of brackets. I don't think we want to be more specific than that in\n> > the match, because we allow arbitrary cruft inside the brackets (like\n> > \"[RFC/PATCH]\", etc). But if format-patch always puts exactly one set of\n> > brackets, and am strips exactly one set, then that should retain your\n> > subject in practice, even if it starts with [foo].\n> \n> I think it may still make sense to insist that PATCH appears somewhere in\n> the first set of brackets, but I have stop and wonder if it is even\n> necessary.\n\nI imagine not.  I've submitted a patch separately which implements\nthis behaviour (more on that below).\n\n> Because git removes [sbuild] at the beginning, Roger is unhappy.\n> \n>  * Is he happy that git removes [PATCH]?  In E-mail based workflow it is\n>    a good practice to mark messages that are patches clearly so that they\n>    can be quickly found among the discussions that lead to them, and it is\n>    plausible that his project accepted that as an established practice\n>    supported well by git.\n\nI'm perfectly happy that [PATCH] is removed.  My requirement is that\nthe commit created by \"git am\" is identical to the commit represented\nin the patch created with \"git format-patch\".  The removal of this\nis IMO correct, and I agree that it's presence is useful in an email-\nbased workflow.\n\n>  * Is he happy that git treats the first paragraph of the commit message\n>    specially from the rest of the message?  In a project with many\n>    commits, it is essential that people write good commit summaries that\n>    fits on a single line so that tools like shortlog and gitweb can be\n>    used to get a bird-eye view of what happened recently.  Perhaps his\n>    project picked it up as the best current practice supported well by\n>    git.\n\nI'm also happy with this, and make use of it.  As for the previous\nparagraph, I would like the commit message to be preserved correctly\nso that the message committed by \"git am\" matches the original\ncommit message exactly.\n\n>  * Is he happy that git takes \"---\" as the end of message marker, so that\n>    any other commentary can be added to the message to facilitate the\n>    communication without adding noise to the commits?  Perhaps he is and\n>    his project picked it up as a good practice supported well by git.\n\nThis sounds just fine, though I have not yet had the need to use it.\n\n> _An_ established (note that I did not say _the_ nor _best current_)\n> practice supported well by git to note the area being affected in a\n> project of nontrivial size is to prefix the single line summary with the\n> name of the area followed by a colon.  There is no difference between\n> \"[sbuild] foo\" and \"sbuild: foo\" at the information content point-of-view,\n> but the latter has an advantage of being one letter shorter and less\n> distracting in MUA.  He does not have a very strong reason to choose\n> something different only to make his life harder, does he?\n\nWell, I sometimes use the format\n\n  [foo] bar: baz\n\nbut my more general point was not my specific usage but that the\nexisting behaviour was causing loss of information.  I think it\nwould be preferable to guarantee that data from the original\ncommit is not lost and is preserved exactly if at all possible.\n\n> Supporting a slightly different convention may seem to be accomodating and\n> nice, but if there is no real technical difference between the two (and\n> again, \"area:\" is one letter shorter ;-), letting people run with\n> different convention longer, when they can switch easily to another\n> convention that is already well supported, may actually hurt them in the\n> long run.  \"[sbuild]\" will not match \"--area=sbuild\" that will internally\n> become \"--grep-only-first-line=sbuild:\" so either he will miss out\n> benefiting from the new feature, or the implementation of the new feature\n> unnecessarily needs more code.\n\nThis is a nice feature I wasn't aware of, so thanks for pointing it\nout.  It might be useful to alter my workflow to allow it to be used,\nor alternatively customisation to allow a custom regex stored e.g.\nin .git/config would allow me to match both forms?\n\nThe patch I sent to the list separately replaces the existing\ncleanup_subject string munging (which is rather complex and\nhairy), with a single regular expression to match the bits of\nthe string we don't want such as '^Re:' and the first set of\nsquare brackets.  We then just keep the remainder.  I initally\nwent with the following extended regex:\n\n  ^([Rr]e: )?(.*PATCH[^]]*\\\\] )(.*)$\n\nbut as per your comments above about removing the first set of\nbrackets whatever the contents, chose the following more\ngeneral expression:\n\n  ^([Rr]e:)?([^]]*\\\\[[^]]+\\\\])(.*)$\n\nThis should be rather more maintainable and flexible than the\nexisting code, because one can just tweak the regex rather than\nfiddling with hairy string offsets.  This preserves the\nexisting behaviour with the exception of matching the first []\npair only rather than being \"greedy\" and removing everything up\nto the last \"]\".\n\n\nRegards,\nRoger\n\n-- \n  .''`.  Roger Leigh\n : :' :  Debian GNU/Linux             http://people.debian.org/~rleigh/\n `. `'   Printing on GNU/Linux?       http://gutenprint.sourceforge.net/\n   `-    GPG Public Key: 0x25BFB848   Please GPG sign your mail.\n"},{"id":"117186","messageId":"20090629214919.GB5397@codelibre.net","threadId":"19958","inReplyTo":"m3ljnawx3h.fsf@localhost.localdomain","subject":"Re: [PATCH] builtin-mailinfo.c: Trim only first pair of square brackets in subject","fromName":"Roger Leigh","fromEmail":"rleigh@codelibre.net","sentAt":"2009-06-29T21:49:20Z","receivedAt":"2009-06-29T21:49:20Z","isPatch":true,"sender":{"key":"rleigh@codelibre.net","avatar":null},"body":"On Mon, Jun 29, 2009 at 02:26:45PM -0700, Jakub Narebski wrote:\n> Roger Leigh <rleigh@debian.org> writes:\n> \n> > Use a regular expression to match text after \"Re:\" or any text in the\n> > first pair of square brackets such as \"[PATCH n/m]\".  This replaces\n> > the complex hairy string munging with a simple single  pattern match.\n> \n> [...]\n> > +\t/* Strip off 'Re:' and/or the first text in square brackets, such as\n> > +\t   '[PATCH]' at the start of the mail Subject. */\n> > +\tstatus = regcomp(&regex,\n> > +\t\t\t \"^([Rr]e:)?([^]]*\\\\[[^]]+\\\\])(.*)$\",\n> > +\t\t\t REG_EXTENDED);\n> \n> Sidenote: it probably didn't worked before either, but there are some\n> broken mail readers in the wold (*cough* MS Outlook *cough*), that\n> misinterpret RFCs and use translated form of \"Re:\" e.g. \"Odp:\" (Polish),\n> or not strip \"Re:\" when replying resulting in string of \"Re: Re: Re: ...\",\n> or use capitalized form of \"Re:\", i.e. \"RE:\", or use yet another form \n> e.g. compact form of repeated \"Re: Re: Re: ...\" in form of \"Re(3):\".\n> \n> But I guess it didn't worked before either.\n\nOne could update the regex to cope with that easily enough such as\n\n  \"^([Rr]e:[[:space:]]*)*([^]]*\\\\[[^]]+\\\\])(.*)$\"\n\nfor the \"Re: Re: Re:\" case, though I can't say I've seen anything\nexcept \"Re:\" for years.  Maybe I just don't get mail and patches\nfrom Outlook users ;-)\n\n\nRegards,\nRoger\n\n-- \n  .''`.  Roger Leigh\n : :' :  Debian GNU/Linux             http://people.debian.org/~rleigh/\n `. `'   Printing on GNU/Linux?       http://gutenprint.sourceforge.net/\n   `-    GPG Public Key: 0x25BFB848   Please GPG sign your mail.\n"},{"id":"117209","messageId":"20090630053333.GD29643@sigio.peff.net","threadId":"19958","inReplyTo":"1246269351-26929-1-git-send-email-ae@op5.se","subject":"Re: [PATCH] mailinfo: Remove only one set of square brackets","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2009-06-30T05:33:34Z","receivedAt":"2009-06-30T05:33:34Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jun 29, 2009 at 11:55:51AM +0200, Andreas Ericsson wrote:\n\n> git-format-patch prepends patches with a [PATCH x/n] prefix, but\n> mailinfo used to remove any number of square-bracket pairs and\n> the content between them. This prevents one from using a commit\n> subject like this:\n> \n>   [ and ] must be allowed as input\n> \n> Removing the square bracket pair from this rather clumsily\n> constructed subject line loses important information, so we must\n> take care not to.\n> \n> This patch causes the subject stripping to stop after it has\n> encountered one pair of square brackets.\n\nI think this is a definite improvement, though I would be much more\nconvinced that the does the right thing if there were some tests. :)\n\n> One possible downside of this patch is that the patch-handling\n> programs will now fail at removing author-added square-brackets\n> to be removed, such as\n> \n>   [RFC][PATCH x/n]\n> \n> However, since format-patch only adds one set of square brackets,\n> this behaviour is quite easily undesrstood and defended while the\n> previous behaviour is not.\n\nAgreed. And I think Junio raised a good point elsewhere: there are\ncertain formatting conventions that are part of format-patch output. So\nI think we do need to address \"this subject munging is totally idiot\nproof and will always reproduce the input patch text exactly\". But\nrather \"is this a sane and useful way to do the munging?\". And I think\nit is a useful convention.\n\nThis is a user-visible change that might impact people's workflows (if\nonly slightly), though, so it should probably get a good mention in the\nrelease notes.\n\n-Peff\n"},{"id":"123622","messageId":"87hbuv5km2.fsf@janet.wally","threadId":"19958","inReplyTo":"1246310220-16909-1-git-send-email-rleigh@debian.org","subject":"Re: [PATCH] builtin-mailinfo.c: Trim only first pair of square brackets in subject","fromName":"Neil Roberts","fromEmail":"bpeeluk@yahoo.co.uk","sentAt":"2009-09-22T10:39:33Z","receivedAt":"2009-09-22T10:39:33Z","isPatch":true,"sender":{"key":"bpeeluk@yahoo.co.uk","avatar":"https://gravatar.com/avatar/91a7a1ae011c2431594cf617b8c5e8af439673b724caa86836f22c51254deb93?d=mp&s=160"},"body":"Roger Leigh <rleigh@debian.org> writes:\n\n> Use a regular expression to match text after \"Re:\" or any text in the\n> first pair of square brackets such as \"[PATCH n/m]\".  This replaces\n> the complex hairy string munging with a simple single pattern match.\n\nIs this patch going to get applied? We like to use the '[topic]' format\nin Clutter¹ because it looks so much nicer than 'topic:' and it would be\nreally nice not to have to manually fix the commit when a contributor\nsends a patch in the same format.\n\n- Neil\n\n[1] http://git.clutter-project.org/cgit.cgi?url=clutter/log/\n"},{"id":"123626","messageId":"87pr9juohc.fsf_-_@janet.wally","threadId":"19958","inReplyTo":"87hbuv5km2.fsf@janet.wally","subject":"Re: [PATCH] builtin-mailinfo.c: Improve the regexp for cleaning up the subject","fromName":"Neil Roberts","fromEmail":"bpeeluk@yahoo.co.uk","sentAt":"2009-09-22T12:56:47Z","receivedAt":"2009-09-22T12:56:47Z","isPatch":true,"sender":{"key":"bpeeluk@yahoo.co.uk","avatar":"https://gravatar.com/avatar/91a7a1ae011c2431594cf617b8c5e8af439673b724caa86836f22c51254deb93?d=mp&s=160"},"body":"Previously the regular expression would remove the first set of square\nbrackets regardless of what came before it. If a patch with a summary\nsuch as 'Added a[0] to a line' was passed through git-format-patch\nwith the -k option then the summary would be cropped to 'to a line'\nwhen applied with git-am.\n\nThe new regular expression also matches any number of 're:' prefixes\nwhich apparently can be generated by some old mail clients.\n\nThe old regexp required that there be at least one set of square\nbrackets before it would remove the 're:' and this is now fixed.\n---\n builtin-mailinfo.c |    9 ++++-----\n 1 files changed, 4 insertions(+), 5 deletions(-)\n\nThis patch is meant to apply on top of the two previous patches by\nRoger Leigh which are available here:\n\nhttp://marc.info/?l=git&m=124839483217718&w=2\nhttp://marc.info/?l=git&m=124839483317722&w=2\n\nIt fixes some small problems as described above.\n\ndiff --git a/builtin-mailinfo.c b/builtin-mailinfo.c\nindex 7098c90..f5799f1 100644\n--- a/builtin-mailinfo.c\n+++ b/builtin-mailinfo.c\n@@ -227,7 +227,7 @@ static void cleanup_subject(struct strbuf *subject)\n \t/* Strip off 'Re:' and/or the first text in square brackets, such as\n \t   '[PATCH]' at the start of the mail Subject. */\n \tstatus = regcomp(&regex,\n-\t\t\t \"^([Rr]e:)?([^]]*\\\\[[^]]+\\\\])(.*)$\",\n+\t\t\t \"^([Rr]e:[ \\t]*)*(\\\\[[^]]+\\\\][ \\t]*)?\",\n \t\t\t REG_EXTENDED);\n \n \tif (status) {\n@@ -248,10 +248,9 @@ static void cleanup_subject(struct strbuf *subject)\n \t/* Store any matches in match. */\n \tstatus = regexec(&regex, subject->buf, 4, match, 0);\n \n-\t/* If there was a match for \\3 in the regex, trim the subject\n-\t   to this match. */\n-\tif (!status && match[3].rm_so > 0) {\n-\t\tstrbuf_remove(subject, 0, match[3].rm_so);\n+\t/* If there was a match, remove it */\n+\tif (!status && match[0].rm_so >= 0) {\n+\t\tstrbuf_remove(subject, 0, match[0].rm_eo);\n \t\tstrbuf_trim(subject);\n \t}\n \n-- \n1.6.0.4\n"},{"id":"123638","messageId":"7vocp3t0oz.fsf@alter.siamese.dyndns.org","threadId":"19958","inReplyTo":"87hbuv5km2.fsf@janet.wally","subject":"Re: [PATCH] builtin-mailinfo.c: Trim only first pair of square brackets in subject","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-09-22T16:15:56Z","receivedAt":"2009-09-22T16:15:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Neil Roberts <bpeeluk@yahoo.co.uk> writes:\n\n> Roger Leigh <rleigh@debian.org> writes:\n>\n>> Use a regular expression to match text after \"Re:\" or any text in the\n>> first pair of square brackets such as \"[PATCH n/m]\".  This replaces\n>> the complex hairy string munging with a simple single pattern match.\n>\n> Is this patch going to get applied?\n\nI do not think it is likely to happen for a patch without much comments\nnor progress after this long blank period, without a refresher discussion.\n\nIt definitely won't be applied silently in its original form, especially\nbecause the final comment in the old discussion on the patch in question\nbegan with \"One could _update_ ...\" from the author of the patch, and then\nnothing happened.\n\n    http://thread.gmane.org/gmane.comp.version-control.git/122418/focus=122466\n\nI actually liked the much simpler one by Andreas in the original thread,\nbut if you really want to use a regexp (which we didn't have to) we should\nmake it configurable.  See the neighbouring discussion here as well.\n\n    http://thread.gmane.org/gmane.comp.version-control.git/123322\n\nI think we all agree that the behaviour should be improved, but I think\nneither Roger's patch nor Andreas's one was the solution..  People who\ncare need to carry discussions and proposed patches forward to help us\nagree on an acceptable solution.\n"},{"id":"123641","messageId":"87k4zqvs70.fsf@janet.wally","threadId":"19958","inReplyTo":"7vocp3t0oz.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] builtin-mailinfo.c: Trim only first pair of square brackets in subject","fromName":"Neil Roberts","fromEmail":"bpeeluk@yahoo.co.uk","sentAt":"2009-09-22T16:51:15Z","receivedAt":"2009-09-22T16:51:15Z","isPatch":true,"sender":{"key":"bpeeluk@yahoo.co.uk","avatar":"https://gravatar.com/avatar/91a7a1ae011c2431594cf617b8c5e8af439673b724caa86836f22c51254deb93?d=mp&s=160"},"body":"> Neil Roberts <bpeeluk@yahoo.co.uk> writes:\n>\n>> Is this patch going to get applied?\n\nJunio C Hamano <gitster@pobox.com> writes:\n\n> I do not think it is likely to happen for a patch without much\n> comments nor progress after this long blank period, without a\n> refresher discussion.\n>\n> It definitely won't be applied silently in its original form,\n> especially because the final comment in the old discussion on the\n> patch in question began with \"One could _update_ ...\" from the author\n> of the patch, and then nothing happened.\n>\n>     http://thread.gmane.org/gmane.comp.version-control.git/122418/focus=122466\n\nOk, fair enough. I submitted another patch to mailing list earlier which\nat least addresses the issue mentioned by the original author when he\nsays \"One could _update_ ...\".\n\n> I actually liked the much simpler one by Andreas in the original\n> thread, but if you really want to use a regexp (which we didn't have\n> to) we should make it configurable.  See the neighbouring discussion\n> here as well.\n>\n>     http://thread.gmane.org/gmane.comp.version-control.git/123322\n\nOh I didn't see that thread, sorry. It's quite tricky to track the issue\nwhen it is spread across multiple threads in a mailing list.\n\nI'm not particularly set on the idea of it being a regular expression so\nI'd be happy with an improved version of the existing loop. I'd\ncertainly be happy with it being an option as in your patch here:\n\nhttp://article.gmane.org/gmane.comp.version-control.git/123340\n\nIf it is an option as in that patch surely it's quite safe as it can't\naffect anyone's existing workflow? It might also be nice if it was\npossible to change it in .git/config so you could enable it by default\nfor projects that use the '[topic]' syntax (such as Cairo and Clutter).\n\n> I think we all agree that the behaviour should be improved, but I\n> think neither Roger's patch nor Andreas's one was the solution..\n> People who care need to carry discussions and proposed patches forward\n> to help us agree on an acceptable solution.\n\nOk, well I do care about this issue and it annoys me regularly so I\nwould love to reopen the discussion. What are the issues with the last\npatch mentioned above?\n\n- Neil\n"},{"id":"123657","messageId":"4AB96B3C.1040107@gmail.com","threadId":"19958","inReplyTo":"7vocp3t0oz.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] builtin-mailinfo.c: Trim only first pair of square brackets in subject","fromName":"Jason Holden","fromEmail":"jason.k.holden@gmail.com","sentAt":"2009-09-23T00:26:36Z","receivedAt":"2009-09-23T00:26:36Z","isPatch":true,"sender":{"key":"jason.k.holden@gmail.com","avatar":null},"body":"Junio C Hamano wrote:\n> \n> I think we all agree that the behaviour should be improved, but I think\n> neither Roger's patch nor Andreas's one was the solution..  People who\n> care need to carry discussions and proposed patches forward to help us\n> agree on an acceptable solution.\n\n\nAn additional use case for this is that at $dayjob, we use GForge\nAdvanced Server.  With GForge, commits are tied to the bug-tracker\nby including the bug-id in the commit message with the syntax\n[#NNN], where NNN is a unique id for each submitted bug.\n\nSo the typical first line of a commit message looks something like:\n[#100] Fix bug in foo.c\n\nsent using git-send-email, this becomes\n[PATCH] [#100] Fix bug in foo.c\n\nBut of course, both [PATCH] and [#100] get stripped off when applied\nwith git-am, forcing a manual edit.\n\nThe reg-expression stuff isn't necessary for my particular use-case.\n Stripping off brackets that have any variant of \"PATCH\" in them, or\njust stripping off the first set of brackets would work for me.\n\n-- \nRegards,\nJason Holden\n"}]}