{"thread":{"id":"26465","subject":"Using Origin hashes to improve rebase behavior","startedAt":"2011-02-10T21:13:10Z","lastAt":"2011-02-21T23:49:37Z","messageCount":14,"participants":["John Wiegley","Johan Herland","Jeff King","skillzero@gmail.com","Junio C Hamano","Thomas Rast","Enrico Weigelt","Dave Abrahams"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"160838","messageId":"m21v3fvbix.fsf@hermes.luannocracy.com","threadId":"26465","inReplyTo":null,"subject":"Using Origin hashes to improve rebase behavior","fromName":"John Wiegley","fromEmail":"johnw@boostpro.com","sentAt":"2011-02-10T21:13:10Z","receivedAt":"2011-02-10T21:13:10Z","isPatch":false,"sender":{"key":"johnw@boostpro.com","avatar":null},"body":"The following proposal is a check to see if this approach would be sane and\nwhether someone is already doing similar work.  If not, I offer to implement\nthis solution.\n\nTHE PROBLEM\n\nSay I have a master from which I have branched locally, and that this private\nbranch has four commits:\n\n    a   b   c\n    o---o---o\n             \\\n              o---o---o---o\n              1   2   3   4\n\nI then decide to cherry pick commit 3 onto master.  Please believe that my\nsituation is such that I cannot immediately rebase the private branch to drop\nthe now-duplicated change.  I end up with this:\n\n    a   b   c   3'\n    o---o---o---o\n             \\\n              o---o---o---o\n              1   2   3   4\n\nLater, there is work on master which changes the same lines of code that 3'\nhas changed.  The commit which changes 3' is e*\n\n    a   b   c   3'  d   e*  f\n    o---o---o---o---o---o---o\n             \\\n              o---o---o---o\n              1   2   3   4\n\nAt a later date, I want to rebase the private branch onto master.  What will\nhappen is that the changes in 3 will conflict with the rewritten changes in\ne*.  However, I'd like Git to know that 3 was already incorporated at some\nearlier time, and *not consider it during the rebase*, since it doesn't need\nto.\n\nTHE SOLUTION\n\nFor the purposes of this discussion, I'd like to define the term \"aggregate\nidentity\" (insert better name here) as a set including: a commit's sha, and\nzero or more shas stored in a new field named \"Origin-Ids\".\n\nIf, when cherry-picking, the originating's commit id is stored in the\nOrigin-Ids field of the cherry-picked commit, then rebase could know whether a\ngiven commit's changes had already been applied.  The logic would look like\nthis:\n\n  1. When rebasing a branch A onto B, find the common ancestor of A and B.\n  2. Examine every commit on B since that common ancestor, collecting a\n     set of their aggregate identities.\n  3. For each commit on A, ignore it if its aggregate identity occurs in\n     that set.\n\nThis would cause commit 3 to be ignored during the rebase above, since 3'\nwould have an origin id referring to 3.\n\nIMPLEMENTATION\n\nA few things need to be done:\n\n - Extend commit objects to have an Origin field, which can be zero, one or a\n   list of hashes.\n\n - Add an option to git commit so that one or more origin ids can be specified\n   at the time any commit is made.  There may be occasions when it's useful to\n   explicitly state that a new commit should somehow 'override' the contents\n   of another during a rebase.\n\n - git cherry-pick and git am should add this Origin field, showing the commit\n   their contents originated from.\n\n - git merge --squash would store the commit ids, and the origin ids, of every\n   commit involved in the merge into the resulting commit's Origin field.\n\n   Note that nothing can be done about rebasing a squashed merge commit onto\n   another squashed merge commit, even though it could be detected that they\n   had common changes.  I don't believe it would even be useful to warn about\n   this, the user would just have to resolve the conflicts manually.\n\n - git log could be extended to show the \"parentage\" (really, the aunt/uncle)\n   of commits with origin info, assuming those origin commits are not dangling\n   (which is OK, and likely to occur after the originating branch is deleted,\n   or if the originating branch is in another repository).\n\n   Where there are multiple Origin ids, a search could be done to find the set\n   of most descendent commits, so that history could be usefully shown after\n   an octopus squash, for example.\n\nQUESTIONS\n\nIs it allowable to add new metadata fields to a commit, and would this require\nbumping the repository version number?  Or should this be implemented by\nappending a Header-style textual field at the end of the commit message?\n\n--\n  John Wiegley\n  BoostPro Computing\n  http://www.boostpro.com\n"},{"id":"160843","messageId":"201102102316.23628.johan@herland.net","threadId":"26465","inReplyTo":"m21v3fvbix.fsf@hermes.luannocracy.com","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2011-02-10T22:16:23Z","receivedAt":"2011-02-10T22:16:23Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Thursday 10 February 2011, John Wiegley wrote:\n> Is it allowable to add new metadata fields to a commit, and would this\n> require bumping the repository version number?  Or should this be\n> implemented by appending a Header-style textual field at the end of the\n> commit message?\n\nMany have tried before you to add such fields to the commit objects \n(including, literally, storing the origin of cherry-picks to help with \nrebases; search the archives for several examples). They have not succeeded. \nWith good reason. This information does not belong in the commit object \nheader section (see earlier discussions for a more complete rationale). \nPutting them at the end of the commit message is your best bet. Or even \nbetter: as a note object stored in a special-purpose notes ref (e.g. \nrefs/notes/cherry-picks). The note approach also allows you to retroactively \nadd this field to previous cherry-picks. AND it allows you to remove Origin-\nIDs that refer to no-longer-existing commits. AND it pretty much solves the \n\"git log should show this info\" for you as well. In short, this is exactly \nthe thing that notes were created to do.\n\nAlso, don't forget that the existing -x option to cherry-pick pretty much \ndoes exactly what you want to add to the commit object.\n\nAs for making use of this information in other git commands (e.g. rebase, \nlog, etc.), you should show the list that the feature works well and solves \nReal Problems(tm) in the real world. If you can do so (with patches), and \nare willing to work with the list to address issues raised, and improve your \npatches, I'll guess you have a pretty good shot at getting this accepted.\n\nAFAIK, nobody else is working in this area right now, although I don't read \nthe mailing list religuously, so I may have missed things. As I said, others \nhave previously proposed similar features, so you'll want to search the \narchive for those discussions to make sure you don't repeat the same \nmistakes.\n\n\nHave fun! :)\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"160849","messageId":"20110210225428.GA21335@sigill.intra.peff.net","threadId":"26465","inReplyTo":"m21v3fvbix.fsf@hermes.luannocracy.com","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-10T22:54:28Z","receivedAt":"2011-02-10T22:54:28Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 10, 2011 at 04:13:10PM -0500, John Wiegley wrote:\n\n> For the purposes of this discussion, I'd like to define the term \"aggregate\n> identity\" (insert better name here) as a set including: a commit's sha, and\n> zero or more shas stored in a new field named \"Origin-Ids\".\n> \n> If, when cherry-picking, the originating's commit id is stored in the\n> Origin-Ids field of the cherry-picked commit, then rebase could know whether a\n> given commit's changes had already been applied.  The logic would look like\n> this:\n> \n>   1. When rebasing a branch A onto B, find the common ancestor of A and B.\n>   2. Examine every commit on B since that common ancestor, collecting a\n>      set of their aggregate identities.\n>   3. For each commit on A, ignore it if its aggregate identity occurs in\n>      that set.\n> \n> This would cause commit 3 to be ignored during the rebase above, since 3'\n> would have an origin id referring to 3.\n\nThis can work in some cases, but there are other cases where it might\nnot. For example, consider:\n\n   1. I cherry-pick commit X from some branch \"topic\" onto master as X'.\n      We record \"Origin-ID: X\" in X'.\n\n   2. I rebase \"topic\" (either onto some other branch, or perhaps I use\n      rebase -i to rewrite some earlier commit). X now becomes some X''.\n\n   3. I now rebase \"topic\" onto master. But we fail to note that X''\n      matches X, so we try to rebase it.\n\nNow, in step 2 we could record \"X\" as an origin ID of X'' and during the\nrebase in step 3, calculate the intersection of the origins of X'' and\nX', and see that they are both just X. And I think maybe you already\nrealize that, since you talk about Origin-IDs as sets.\n\nBut now you have an interesting question: during which operations does a\ncommit retain its Origin-ID on a source? I think it is pretty clear that\na cherry-pick that cleanly applies is probably a good candidate. But\nwhat if there is a conflict, and I fix up the conflict? Should I still\nskip the original commit during the rebase? Maybe, but there are cases\nwhere you wouldn't want to. For example, consider this sequence:\n\n  1. On my master branch, I have a function foo() which takes one\n     argument.\n\n  2. I branch a topic from master. The first commit adds a new caller to\n     foo().\n\n  3. The second commit changes foo() to take two arguments. I fix up the\n     function itself, any old callers, and the new caller.\n\n  4. I cherry-pick the second commit onto master. There is a conflict,\n     since one of the callers it updates doesn't exist in master. So I\n     drop that part of the patch.\n\n  5. On master I update the implementation of foo().\n\n  6. Now I try to rebase my topic on top of master. We could get a\n     conflict because the second commit from (3) will conflict with the\n     updated implementation in (5). This is more or less the case you\n     described in your initial email, and we'd like git to automatically\n     realize that the conflict is uninteresting.\n\n     So let's imagine we recorded Origin-IDs as you describe, and we\n     skip it. But that means we are _also_ skipping the part where we\n     update the new caller from the commit in (2), the part that was\n     dropped during conflict resolution. So our end result is broken,\n     because the new caller is still calling with one argument.\n\nAnd there are lots of other cases. What about \"git cherry-pick -n\"? What\nabout rebasing? If there are no conflicts, is it OK to copy the origin\nfield? How about if there are conflicts? How about in a \"git rebase -i\",\nwhere we may stop and the user can arbitrarily split, amend, or add new\ncommits. How do the old commits map to the new ones with respect to\norigin fields?\n\nSo there are lots of corner cases where it won't work, because git is\nmore than happy to give you lots of ways to tweak tree state and\nhistory, and it fundamentally doesn't care as much about process as it\ndoes about the end states that you reach. That's part of what makes git\nso flexible, but it also makes niceties like \"did I already apply this\ncommit on this branch\" much harder to make sense of.\n\nNow, I don't want to discourage you from working on this. Because while\nthere are lots of cases where it won't work, there are plenty of cases\nwhere it _will_, and it will save rebasers time and effort. So it is\nworth pursuing, but I think it is also worth keeping things simple and\nconservative, and not affecting the people who have cases where this\nwon't help.\n\n>  - Extend commit objects to have an Origin field, which can be zero, one or a\n>    list of hashes.\n\nIt probably shouldn't be a new header field, but rather a text-style\npseudo-header at the end of the commit.\n\nBut consider for a moment whether you actually want this field in the\nresulting commit at all, or whether it should be an external annotation.\nFor example, let's say I cherry-pick from a private branch that is going\nto end up rebased anyway. Now the history for all time will have a\ncommit that refers to some totally useless sha1 that nobody even knows\nabout.\n\nWe already went through this with cherry-pick. It used to always put\n\"cherry-picked from X...\" in the commit message. And then we realized\nthat in many cases, that information is not interesting, because X is\nnot something people actually know about. So now we don't do it by\ndefault, but for cases where you are cherry-picking from one\nlong-running branch to another, you can use \"cherry-pick -x\".\n\nSo consider instead putting this information into a commit-note for the\nnew commit. Possibly even reversing the direction of the mapping (so\nthat the old commit says \"I was cherry-picked to X\"). And then when the\nold, rebased commit goes away, the note will automagically get pruned by\nthe notes-pruning mechanism.\n\nThere may be reasons why that isn't a good idea, and I haven't thought\nit through. But I think you should consider it as an alternate\nimplementation and tell me why I'm dumb in that case. ;)\n\n>  - Add an option to git commit so that one or more origin ids can be specified\n>    at the time any commit is made.  There may be occasions when it's useful to\n>    explicitly state that a new commit should somehow 'override' the contents\n>    of another during a rebase.\n> \n>  - git cherry-pick and git am should add this Origin field, showing the commit\n>    their contents originated from.\n\nWe already have this to some degree, in the form of \"cherry-pick -x\".\nYou could do it with \"git am\", but you would need \"git format-patch\" to\nactually generate the information (well, technically speaking it is in\nthe mbox \"From \" header, but that usually doesn't make it through mail\ntransports for obvious reasons).\n\nSo I wonder if your proposal can be restructured as:\n\n  1. Change rebase to look for cherry-picked-from headers on the --onto\n     side, and skip source commits that appear to exist already. That\n     will start helping people immediately using existing history.\n\n     You can also deal with uncertainty by leaving this decision to the\n     last minute, or even leaving it up to the user. The usual patch-id\n     detection works in a lot of cases. Let it work when it does. When\n     it fails, check if the conflicted commit exists in a\n     cherry-picked-from line. If it does, either do the skip then, or\n     when we barf with the \"there was a conflict; fix it up and rebase\n     --continue\" message, mention the cherry-picked-from line and let\n     the user inspect the commits themselves and make a decision.\n\n     They can always do \"git rebase --skip\" even now, so all we are\n     really doing is saying \"By the way, you might want the extra\n     information that this was cherry-picked earlier\". And that makes\n     this a very low-risk change, since we are just giving the user\n     extra information for a decision they are already making.\n\n  2. (Optional) Start adding the \"Cherry picked from\" message in a more\n     machine-readable format, like an \"Origin-ID: ...\" header. This has\n     already been discussed before. People were generally positive, but\n     it didn't seem especially useful. This is a use.\n\n     And obviously make the corresponding change in rebase to also parse\n     these kinds of headers (but don't drop parsing the original format,\n     obviously, for compatibility).\n\n  3. For people who don't want the \"cherry picked from\" (or \"origin-id\")\n     in their commit, because they are cherry-picking from a private\n     source, start recording \"cherry picked from\" in a git-note. You\n     could even do this by default, since you are not impacting the\n     commits themselves in any way.\n\n     And then make the corresponding change in rebase to start using\n     these notes as a source.\n\n  4. We already have some functionality to copy notes about commit A to\n     commit B during certain operations (like rebasing and\n     cherry-picking). Check out how these interact with the notes\n     introduced in (3) to see if transitive stuff works (like\n     cherry-picking A to A', and then A' to A''; you should still be\n     able to figure out that A'' came from A).\n\nAnd I think at that point we have more or less the functionality you\nwere asking for, though we arrived in several non-controversial steps.\nAnd there are lots of enhancements you could add on top, like skipping\nwithout bothering the user about it, or better heuristics for when to\nrecord an origin-id or not to. But we can do those once we see how the\nbasic dumb part performs. I.e., how useful it is in practice, and how\noften it is wrong about when to skip.\n\n>  - git merge --squash would store the commit ids, and the origin ids, of every\n>    commit involved in the merge into the resulting commit's Origin field.\n\nI hadn't thought about merge --squash as a commit copying operation, but\nI think it is. I wonder if squash merges (or squash rebases) should also\nbe copying notes (or if they do already, I haven't checked).\n\n>  - git log could be extended to show the \"parentage\" (really, the aunt/uncle)\n>    of commits with origin info, assuming those origin commits are not dangling\n>    (which is OK, and likely to occur after the originating branch is deleted,\n>    or if the originating branch is in another repository).\n\nIf you do it with a combination of text in the commit message and\ngit-notes, then this is all done for you. The commit message you\nobviously see by default, and you could explicitly ask for it to show\nthe refs/notes/origin-id notes tree.\n\n\nWhew, that turned out long. I hope it's helpful. I think the problem\nyou're trying to solve is a real one, and I think your approach is the\nright direction. I just think we can leverage existing git features to\ndo most of it, and because it is sort of a heuristic, we should be\nconservative in how it's introduced.\n\n-Peff\n"},{"id":"160874","messageId":"m2oc6jtg8o.fsf@hermes.luannocracy.com","threadId":"26465","inReplyTo":"20110210225428.GA21335@sigill.intra.peff.net","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"John Wiegley","fromEmail":"johnw@boostpro.com","sentAt":"2011-02-11T03:14:15Z","receivedAt":"2011-02-11T03:14:15Z","isPatch":false,"sender":{"key":"johnw@boostpro.com","avatar":null},"body":"Jeff King <peff@peff.net> writes:\n\n> Now, in step 2 we could record \"X\" as an origin ID of X'' and during the\n> rebase in step 3, calculate the intersection of the origins of X'' and X',\n> and see that they are both just X. And I think maybe you already realize\n> that, since you talk about Origin-IDs as sets.\n\nRight.\n\n> And there are lots of other cases. What about \"git cherry-pick -n\"? What\n> about rebasing? If there are no conflicts, is it OK to copy the origin\n> field? How about if there are conflicts? How about in a \"git rebase -i\",\n> where we may stop and the user can arbitrarily split, amend, or add new\n> commits. How do the old commits map to the new ones with respect to\n> origin fields?\n\nDuring rebasing, any commits which can be rebased without conflict have their\norigin transferred (and each time it would cause the origin id list to grow by\none), but any commits which are squashed or edited would not transfer.\n\nFor cherry-pick -n, if the index is empty at the time the cherry-pick is done\n(is this required?), then a file is created under .git/ with the SHA of the\nchanges placed in the index, so that when git-commit is later run and the\nindex has not been changed, then the Origin-Id for that originating commit\ngets placed at the bottom of the commit message.\n\n> So there are lots of corner cases where it won't work, because git is\n> more than happy to give you lots of ways to tweak tree state and\n> history, and it fundamentally doesn't care as much about process as it\n> does about the end states that you reach. That's part of what makes git\n> so flexible, but it also makes niceties like \"did I already apply this\n> commit on this branch\" much harder to make sense of.\n\nI think we'd want to restrict this system to those commits which were\nautomatically rewritten without conflicts.  Any user intervention in the\nprocess would invalidate the meaning of the Origin-Id.\n\n> It probably shouldn't be a new header field, but rather a text-style\n> pseudo-header at the end of the commit.\n\nI understand.\n\n> But consider for a moment whether you actually want this field in the\n> resulting commit at all, or whether it should be an external annotation.\n> For example, let's say I cherry-pick from a private branch that is going\n> to end up rebased anyway. Now the history for all time will have a\n> commit that refers to some totally useless sha1 that nobody even knows\n> about.\n\nThe problem with an external annotation is that if developers are sharing\nfeature branches, as a branch maintainer I want to know whether commits coming\nfrom those feature branches are already in the branch I'm maintaining.\n\n> There may be reasons why that isn't a good idea, and I haven't thought it\n> through. But I think you should consider it as an alternate implementation\n> and tell me why I'm dumb in that case. ;)\n\nI'll give it a bit more thought as I consider the implementation of this.\n\n> Whew, that turned out long. I hope it's helpful. I think the problem\n> you're trying to solve is a real one, and I think your approach is the\n> right direction. I just think we can leverage existing git features to\n> do most of it, and because it is sort of a heuristic, we should be\n> conservative in how it's introduced.\n\nThat's all extremely helpful, thank you!  You've brought up several use cases\nI hadn't thought of, and perhaps this feature will indeed never cover\neverything, but if it can reliably ease maintenance 80% of the time, I think\nit's a relatively simple addition.\n\nJohn\n"},{"id":"160875","messageId":"20110211044539.GA2071@sigill.intra.peff.net","threadId":"26465","inReplyTo":"m2oc6jtg8o.fsf@hermes.luannocracy.com","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-11T04:45:40Z","receivedAt":"2011-02-11T04:45:40Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 10, 2011 at 10:14:15PM -0500, John Wiegley wrote:\n\n> > And there are lots of other cases. What about \"git cherry-pick -n\"? What\n> > about rebasing? If there are no conflicts, is it OK to copy the origin\n> > field? How about if there are conflicts? How about in a \"git rebase -i\",\n> > where we may stop and the user can arbitrarily split, amend, or add new\n> > commits. How do the old commits map to the new ones with respect to\n> > origin fields?\n> \n> During rebasing, any commits which can be rebased without conflict have their\n> origin transferred (and each time it would cause the origin id list to grow by\n> one), but any commits which are squashed or edited would not transfer.\n\nOK. That's certainly the conservative answer, and where we should start.\nBut I wonder in practice how many times we'll hit all the criteria just\nright for this feature to kick in (i.e., a cherry pick or rebase with no\nconflicts, followed by one that would cause a conflict). But I think\nthere's nothing to do but implement and see how it works.\n\nAfter thinking about this a bit more, the whole idea of \"is this\ncherry-picked/rebased/whatever commit the same as the one before\" is\nreally the same as the notes-rewriting case (i.e., copying notes on\ncommit A when it is rebased into A'). Which makes me excited about using\nnotes for this, because the rules that you do figure out to work in\npractice will be good rules for notes rewriting in general.\n\n> The problem with an external annotation is that if developers are sharing\n> feature branches, as a branch maintainer I want to know whether commits coming\n> from those feature branches are already in the branch I'm maintaining.\n\nIn that case, I would suggest putting it in git-notes and sharing the\nnotes with each other. The notes code should happily merge them all\ntogether, and then everyone gets to see everybody else's\ncherry-pick/rebase annotations.\n\n-Peff\n"},{"id":"160876","messageId":"m2fwrvta4t.fsf@hermes.luannocracy.com","threadId":"26465","inReplyTo":"20110211044539.GA2071@sigill.intra.peff.net","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"John Wiegley","fromEmail":"johnw@boostpro.com","sentAt":"2011-02-11T05:26:10Z","receivedAt":"2011-02-11T05:26:10Z","isPatch":false,"sender":{"key":"johnw@boostpro.com","avatar":null},"body":"Jeff King <peff@peff.net> writes:\n\n> OK. That's certainly the conservative answer, and where we should start.\n> But I wonder in practice how many times we'll hit all the criteria just\n> right for this feature to kick in (i.e., a cherry pick or rebase with no\n> conflicts, followed by one that would cause a conflict). But I think\n> there's nothing to do but implement and see how it works.\n>\n> After thinking about this a bit more, the whole idea of \"is this\n> cherry-picked/rebased/whatever commit the same as the one before\" is\n> really the same as the notes-rewriting case (i.e., copying notes on\n> commit A when it is rebased into A'). Which makes me excited about using\n> notes for this, because the rules that you do figure out to work in\n> practice will be good rules for notes rewriting in general.\n>\n>> The problem with an external annotation is that if developers are sharing\n>> feature branches, as a branch maintainer I want to know whether commits coming\n>> from those feature branches are already in the branch I'm maintaining.\n>\n> In that case, I would suggest putting it in git-notes and sharing the\n> notes with each other. The notes code should happily merge them all\n> together, and then everyone gets to see everybody else's\n> cherry-pick/rebase annotations.\n\nThe more I've talked this over with my friend, the more we discover how\ndifficult this is to get right in certain situations, and also how rare the\nactual use cases that require storage within the commit message are -- but at\nthe same time, how valuable that information is when those cases occur!\n\nThis may be a bit more than I can chew right now, so thank you for bringing to\nmy attention the depth of this problem.  That's exactly why I posted here\nbefore beginning to punch out code that might solve just the naive cases. :)\n\nThanks, John\n"},{"id":"160886","messageId":"AANLkTikrVPCr92XHirn1u=73eM--T190V-7nbE6fo8ng@mail.gmail.com","threadId":"26465","inReplyTo":"m21v3fvbix.fsf@hermes.luannocracy.com","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"","fromEmail":"skillzero@gmail.com","sentAt":"2011-02-11T10:02:52Z","receivedAt":"2011-02-11T10:02:52Z","isPatch":false,"sender":{"key":"skillzero@gmail.com","avatar":null},"body":"On Thu, Feb 10, 2011 at 1:13 PM, John Wiegley <johnw@boostpro.com> wrote:\n> The following proposal is a check to see if this approach would be sane and\n> whether someone is already doing similar work.  If not, I offer to implement\n> this solution.\n>\n> THE PROBLEM\n>\n> Say I have a master from which I have branched locally, and that this private\n> branch has four commits:\n>\n>    a   b   c\n>    o---o---o\n>             \\\n>              o---o---o---o\n>              1   2   3   4\n>\n> I then decide to cherry pick commit 3 onto master.  Please believe that my\n> situation is such that I cannot immediately rebase the private branch to drop\n> the now-duplicated change.  I end up with this:\n>\n>    a   b   c   3'\n>    o---o---o---o\n>             \\\n>              o---o---o---o\n>              1   2   3   4\n>\n> Later, there is work on master which changes the same lines of code that 3'\n> has changed.  The commit which changes 3' is e*\n>\n>    a   b   c   3'  d   e*  f\n>    o---o---o---o---o---o---o\n>             \\\n>              o---o---o---o\n>              1   2   3   4\n>\n> At a later date, I want to rebase the private branch onto master.  What will\n> happen is that the changes in 3 will conflict with the rewritten changes in\n> e*.  However, I'd like Git to know that 3 was already incorporated at some\n> earlier time, and *not consider it during the rebase*, since it doesn't need\n> to.\n\nI don't know very much about how git really works so what I'm saying\nmay be dumb, but rather than record where a commit came from, would it\nbe reasonable for rebase to look at the patch-id for each change on\nthe topic branch after the merge base and automatically remove topic\nbranch commits that match that patch-id? So in your example, rebase\nwould check each topic branch commit against 3', d, e*, and f and see\nthat the 3' patch-id is the same as the topic branch 3 and remove\ntopic branch 3 before it gets to e*?\n"},{"id":"160891","messageId":"201102111240.29746.johan@herland.net","threadId":"26465","inReplyTo":"AANLkTikrVPCr92XHirn1u=73eM--T190V-7nbE6fo8ng@mail.gmail.com","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2011-02-11T11:40:29Z","receivedAt":"2011-02-11T11:40:29Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Friday 11 February 2011, skillzero@gmail.com wrote:\n> On Thu, Feb 10, 2011 at 1:13 PM, John Wiegley <johnw@boostpro.com> wrote:\n> >    a   b   c   3'  d   e*  f\n> >    o---o---o---o---o---o---o\n> >             \\\n> >              o---o---o---o\n> >              1   2   3   4\n> > \n> > At a later date, I want to rebase the private branch onto master.  What\n> > will happen is that the changes in 3 will conflict with the rewritten\n> > changes in e*.  However, I'd like Git to know that 3 was already\n> > incorporated at some earlier time, and *not consider it during the\n> > rebase*, since it doesn't need to.\n> \n> I don't know very much about how git really works so what I'm saying\n> may be dumb, but rather than record where a commit came from, would it\n> be reasonable for rebase to look at the patch-id for each change on\n> the topic branch after the merge base and automatically remove topic\n> branch commits that match that patch-id? So in your example, rebase\n> would check each topic branch commit against 3', d, e*, and f and see\n> that the 3' patch-id is the same as the topic branch 3 and remove\n> topic branch 3 before it gets to e*?\n\nI believe \"git rebase\" already does exactly what you describe [1].\n\nHowever, comparing patch-ids stops working when the cherry-pick (3 -> 3') \nhas conflicts. IINM, it is the conflicting cases that John is interested in \nsolving...\n\n\n...Johan\n\n[1]: I tested the above scenario, and got no conflicts:\n\n$ git init\n$ FOO=a && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ FOO=b && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ FOO=c && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ git checkout -b topic\n$ FOO=1 && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ FOO=2 && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ FOO=3 && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ FOO=4 && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ git checkout master\n$ git cherry-pick topic^\n$ FOO=d && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ echo e >> 3 && git add 3\n$ FOO=e && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ FOO=f && echo $FOO > $FOO && git add $FOO && git commit -m $FOO\n$ git checkout topic\n$ git rebase master\nFirst, rewinding head to replay your work on top of it...\nApplying: 1\nApplying: 2\nApplying: 4\n$ # Look, no conflicts.\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"160907","messageId":"20110211190326.GB29203@sigill.intra.peff.net","threadId":"26465","inReplyTo":"201102111240.29746.johan@herland.net","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-11T19:03:26Z","receivedAt":"2011-02-11T19:03:26Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 11, 2011 at 12:40:29PM +0100, Johan Herland wrote:\n\n> > I don't know very much about how git really works so what I'm saying\n> > may be dumb, but rather than record where a commit came from, would it\n> > be reasonable for rebase to look at the patch-id for each change on\n> > the topic branch after the merge base and automatically remove topic\n> > branch commits that match that patch-id? So in your example, rebase\n> > would check each topic branch commit against 3', d, e*, and f and see\n> > that the 3' patch-id is the same as the topic branch 3 and remove\n> > topic branch 3 before it gets to e*?\n> \n> I believe \"git rebase\" already does exactly what you describe [1].\n\nYep. It uses format-patch's \"--ignore-if-in-upstream\", which computes\npatch-ids (you can get the same list with \"git cherry\").\n\n> However, comparing patch-ids stops working when the cherry-pick (3 -> 3') \n> has conflicts. IINM, it is the conflicting cases that John is interested in \n> solving...\n\nExactly. One other possible solution to this problem would be to somehow\nmake patch-ids handle fuzzy situations better. I doubt it is possible to\ndo that without introducing a lot of false positives, though.\n\n-Peff\n"},{"id":"160908","messageId":"7v4o8a1i6k.fsf@alter.siamese.dyndns.org","threadId":"26465","inReplyTo":"20110211190326.GB29203@sigill.intra.peff.net","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2011-02-11T19:32:03Z","receivedAt":"2011-02-11T19:32:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Exactly. One other possible solution to this problem would be to somehow\n> make patch-ids handle fuzzy situations better. I doubt it is possible to\n> do that without introducing a lot of false positives, though.\n\nWe need to remember that we would want to tolerate _no_ false positive.\n\nWe try hard to err on the safer side and leave the hard case to the users\nfor a reason.  A tool that records correct results 99.9% of the time but\nproduces wrong results for the rest of the time _silently_ is a tool that\ncannot be trusted, and forces the user to inspect its output carefully to\nmake sure it is correct, not just for the 0.1% cases but for all of them.\n\nAmong the many automation support facilities we have gained over time, the\nthree-way merge, recursive merge to come up with a synthetic merge base\ntree, detecting change similarity with patch-id, and detecting renames by\ncontent inspection all proved themselves to be reasonably trustworthy\nwithout false positives, even though they sometimes fail with false\nnegatives and they do so rather loudly by failing.  I find the heuristics\nin rerere is trustable most of the time but I still do not completely\ntrust it myself.\n\nPatching with fuzz and a user declaration that \"this change came from\nthat\", especially if the user can declare the correspondence even when\nconflict resolution is involved during the porting of changes from totally\ndifferent context, fall into a different, a lot less trustworthy, basket.\nIt needs to start from totally trivial cases and punt _loudly_ when there\nis any doubt.\n"},{"id":"160909","messageId":"20110211194541.GA32023@sigill.intra.peff.net","threadId":"26465","inReplyTo":"7v4o8a1i6k.fsf@alter.siamese.dyndns.org","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2011-02-11T19:45:41Z","receivedAt":"2011-02-11T19:45:41Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 11, 2011 at 11:32:03AM -0800, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > Exactly. One other possible solution to this problem would be to somehow\n> > make patch-ids handle fuzzy situations better. I doubt it is possible to\n> > do that without introducing a lot of false positives, though.\n> \n> We need to remember that we would want to tolerate _no_ false positive.\n\nYeah, I agree with everything you say here. My original message should\nhave been s/a lot of// in the last line.\n\n-Peff\n"},{"id":"160951","messageId":"201102121536.25789.trast@student.ethz.ch","threadId":"26465","inReplyTo":"20110210225428.GA21335@sigill.intra.peff.net","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Thomas Rast","fromEmail":"trast@student.ethz.ch","sentAt":"2011-02-12T14:36:25Z","receivedAt":"2011-02-12T14:36:25Z","isPatch":false,"sender":{"key":"tr@thomasrast.ch","avatar":"https://avatars.githubusercontent.com/u/153510?v=4"},"body":"[I skipped most of the thread, so here's just one minor point.]\n\nJeff King wrote:\n> I hadn't thought about merge --squash as a commit copying operation, but\n> I think it is. I wonder if squash merges (or squash rebases) should also\n> be copying notes (or if they do already, I haven't checked).\n\nSquash rebases do but squash merges don't.  Doing it \"elegantly\" would\nprobably involve caching the list of commits since the merge-bases,\nwhich would be a bit of work.\n\n-- \nThomas Rast\ntrast@{inf,student}.ethz.ch\n"},{"id":"161761","messageId":"20110220174914.GA23366@nibiru.local","threadId":"26465","inReplyTo":"m21v3fvbix.fsf@hermes.luannocracy.com","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Enrico Weigelt","fromEmail":"weigelt@metux.de","sentAt":"2011-02-20T17:49:14Z","receivedAt":"2011-02-20T17:49:14Z","isPatch":false,"sender":{"key":"weigelt@metux.de","avatar":null},"body":"* John Wiegley <johnw@boostpro.com> wrote:\n\n<snip>\n\n> Later, there is work on master which changes the same lines of code that 3'\n> has changed.  The commit which changes 3' is e*\n> \n>     a   b   c   3'  d   e*  f\n>     o---o---o---o---o---o---o\n>              \\\n>               o---o---o---o\n>               1   2   3   4\n> \n> At a later date, I want to rebase the private branch onto master.  What will\n> happen is that the changes in 3 will conflict with the rewritten changes in\n> e*.  However, I'd like Git to know that 3 was already incorporated at some\n> earlier time, and *not consider it during the rebase*, since it doesn't need\n> to.\n\nI'm solving these situations by incremental rebase (rebasing onto earlier\ncommits than the head, iteratively). A command for that would be nice.\n\n\ncu\n-- \n----------------------------------------------------------------------\n Enrico Weigelt, metux IT service -- http://www.metux.de/\n\n phone:  +49 36207 519931  email: weigelt@metux.de\n mobile: +49 151 27565287  icq:   210169427         skype: nekrad666\n----------------------------------------------------------------------\n Embedded-Linux / Portierung / Opensource-QM / Verteilte Systeme\n----------------------------------------------------------------------\n"},{"id":"161846","messageId":"loom.20110222T004926-60@post.gmane.org","threadId":"26465","inReplyTo":"m21v3fvbix.fsf@hermes.luannocracy.com","subject":"Re: Using Origin hashes to improve rebase behavior","fromName":"Dave Abrahams","fromEmail":"dave@boostpro.com","sentAt":"2011-02-21T23:49:37Z","receivedAt":"2011-02-21T23:49:37Z","isPatch":false,"sender":{"key":"dave@boostpro.com","avatar":"https://gravatar.com/avatar/df0921f05114687777894565de21c052fb137ba7c303a399528b43d08833f065?d=mp&s=160"},"body":"Johan Herland <johan <at> herland.net> writes:\n\n> \n> On Friday 11 February 2011, skillzero <at> gmail.com wrote:\n> > On Thu, Feb 10, 2011 at 1:13 PM, John Wiegley\n> > <johnw <at> boostpro.com> wrote:\n>\n> > I don't know very much about how git really works so what I'm saying\n> > may be dumb, but rather than record where a commit came from, would it\n> > be reasonable for rebase to look at the patch-id for each change on\n> > the topic branch after the merge base and automatically remove topic\n> > branch commits that match that patch-id? So in your example, rebase\n> > would check each topic branch commit against 3', d, e*, and f and see\n> > that the 3' patch-id is the same as the topic branch 3 and remove\n> > topic branch 3 before it gets to e*?\n> \n> I believe \"git rebase\" already does exactly what you describe [1].\n\nI can imagine that we could make merges do something similar:\n\ngit merge <sources> :=\n\n   Attempt the merge as it works today\n   If there are conflicts\n       for s in <sources>\n            rebase s onto HEAD\n       if there are no conflicts\n            use the current tree as the result of\n                the merge (with the merge's heritage)\n            commit\n       else\n            reset to the conflicted merge state\n"}]}