{"thread":{"id":"20581","subject":"rebase-with-history -- a technique for rebasing without trashing your repo history","startedAt":"2009-08-13T12:46:07Z","lastAt":"2009-08-15T03:36:23Z","messageCount":10,"participants":["Michael Haggerty","Björn Steinbrink","Bryan O'Sullivan","Abderrahim Kitouni","Sitaram Chamarty","Nanako Shiraishi"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"120526","messageId":"4A840B0F.9060003@alum.mit.edu","threadId":"20581","inReplyTo":null,"subject":"rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2009-08-13T12:46:07Z","receivedAt":"2009-08-13T12:46:07Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Sorry to cross-post, but I think this might be interesting to all three\nprojects...\n\nI've been thinking a lot about the problems of tracking upstream changes\nwhile developing a feature branch.  As I think everybody knows, both\nrebasing and merging have serious disadvantages for this use case.\nRebasing discards history and makes it difficult to share\nwork-in-progress with others, whereas merging makes it difficult to\nprepare a clean patch series that is suitable for submission upstream.\n\nI've written some articles describing another possibility, which\ncombines the advantages of both methods.  The key idea is to retain\nrebase history correctly, on a patch-by-patch level.  The resulting DAG\nretains enough history to prevent problems with merge conflicts\ndownstream, while also allowing the patch series to be kept tidy.\n\n(Please note that this technique only works for the typical \"tracking\nupstream\" type of rebase; it doesn't help with rebases whose goals are\nchanging the order of commits, moving only part of a branch, rewriting\ncommits, etc.)\n\nFor more information, please see the full articles:\n\n* A truce in the merge vs. rebase war? [1]\n* Upstream rebase Just Works™ if history is retained [2]\n* Rebase with history -- implementation ideas [3]\n\nI'd appreciate feedback!\n\nMichael\n\n[1]\nhttp://softwareswirl.blogspot.com/2009/04/truce-in-merge-vs-rebase-war.html\n[2]\nhttp://softwareswirl.blogspot.com/2009/08/upstream-rebase-just-works-if-history.html\n[3]\nhttp://softwareswirl.blogspot.com/2009/08/rebase-with-history-implementation.html\n"},{"id":"120535","messageId":"20090813161256.GA8292@atjola.homenet","threadId":"20581","inReplyTo":"4A840B0F.9060003@alum.mit.edu","subject":"Re: rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-08-13T16:12:56Z","receivedAt":"2009-08-13T16:12:56Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.08.13 14:46:07 +0200, Michael Haggerty wrote:\n> Sorry to cross-post, but I think this might be interesting to all three\n> projects...\n> \n> I've been thinking a lot about the problems of tracking upstream changes\n> while developing a feature branch.  As I think everybody knows, both\n> rebasing and merging have serious disadvantages for this use case.\n> Rebasing discards history and makes it difficult to share\n> work-in-progress with others, whereas merging makes it difficult to\n> prepare a clean patch series that is suitable for submission upstream.\n> \n> I've written some articles describing another possibility, which\n> combines the advantages of both methods.  The key idea is to retain\n> rebase history correctly, on a patch-by-patch level.  The resulting DAG\n> retains enough history to prevent problems with merge conflicts\n> downstream, while also allowing the patch series to be kept tidy.\n> \n> (Please note that this technique only works for the typical \"tracking\n> upstream\" type of rebase; it doesn't help with rebases whose goals are\n> changing the order of commits, moving only part of a branch, rewriting\n> commits, etc.)\n\nHm, so that pretty much doesn't work at all for creating a clean patch\nseries, which usually involves rewriting commits, squasing bug fixes\ninto the original commits that introduced the bug etc.\n\nAnd even for just continously forward porting a series of commits, a\ncommon case might be that upstream applied some patches, but not all.\nCan you deal with that?\n\nExample:\n\nA---B---C (upstream)\n     \\\n      H---I---J---K (yours)\n\nUpstream takes some changes:\n\nA---B---C---I'--K'--D (upstream)\n     \\\n      H---I---J---K (yours)\n\nrebase leads to:\n\nA---B---C---I'--K'--D (upstream)\n                     \\\n                      H'--J' (yours)\n\nWhat would your approach generate in that case?\n\n> For more information, please see the full articles:\n> \n> * Upstream rebase Just Works™ if history is retained [2]\n\nIn this one you have two DAGs:\n(I fixed the second one to also have the merge commit in \"subsystem\"\ninstead of \"topic\", so they only differ WRT to the rebased stuff)\n\nA)\nm---N---m---m---m---m---m---M  (master)\n     \\                       \\\n      o---o---O---o---o       o'--o'--o'--o'--o'--S  (subsystem)\n                       \\                         /\n                        *---*---*-..........-*--T (topic)\n\n\nB)\nm---N---m---m---m---m---m---M  (master)\n     \\                       \\\n      \\                       o'--o'--o'--o'--o'----------S  (subsystem)\n       \\                     /   /   /   /   /           /\n        --------------------o---o---O---o---o---*---*---T (topic)\n\n\nAnd you say that the former creates problems when you want to merge\nagain. How so?\n\nMerging \"master\" to \"subsystem\" is a no-op in both cases.\nMerging \"subsystem\" to \"master\" is a fast-forward in both cases.\nMerging \"subsystem\" to \"topic\" is a fast-forward in both cases.\nMerging \"topic\" to \"subsystem\" is a no-op in both cases.\n\nMerging \"topic\" and \"master\" (in either direction) has merge base N in\nboth cases.\n\nLet's assume that there's another dev, having his own history based on\nthe old \"O\" commit. So:\n\nA)\nm---N---m---m---m---m---m---M  (master)\n     \\                       \\\n      o---o---O---o---o       o'--o'--o'--o'--o'--S  (subsystem)\n               \\       \\                         /\n                \\       *---*---*-..........-*--T (topic)\n                 \\\n                  X---Y---Z (outsider)\n\nB)\nm---N---m---m---m---m---m---M  (master)\n     \\                       \\\n      \\                       o'--o'--o'--o'--o'----------S  (subsystem)\n       \\                     /   /   /   /   /           /\n        --------------------o---o---O---o---o---*---*---T (topic)\n                                     \\\n                                      X---Y---Z (outsider)\n\nMerging \"master\" and \"outsider\" has merge base N in both cases.\nMerging \"subsystem\" and \"outsider\" has merge base O in both cases.\nMerging \"topic\" and \"outsider\" has merge base O in both cases.\n\nThe only thing that really makes a difference is when you have another\ndev having based his history upon one of the o' commits. If that history\nis merged with \"topic\", then you get merge base N in A) and one of the\no's in B).\n\nBut, in that case, merging \"topic\" is basically just a complicated way\nof merging \"subsystem\". Both contain the full series of \"o\" commits, and\nall the \"topic\" commits (due to the merge). So you could just trivially\nmerge \"subsystem\" instead, which leads the o' commit as the merge base\nin both cases.\n\nAnd, that merge of \"topic\" to \"subsystem\" was wrong to begin with. If\nyou rewrite history, that has to trickle down. So A) should really have\nbeen:\n\nm---m---m---m-...--m---m (master)\n                        \\\n                         o'--o'-...-o'--*'--*'--*' (topic) (subsystem)\n\n(topic was rebased, and subsystem fast-forwarded)\n\nSo AFAICT, what your system achieves WRT ease of rebasing is that it\nobsoletes the need to use \"--onto\" with git's rebase.\n\nInstead of \"git rebase --onto subsystem old_subsystem\", you can just say\n\"git rebase subsystem\", at the cost of a very complicated DAG.\n\nBjörn\n"},{"id":"120553","messageId":"c290c4f20908131039s576b5cdycbefa9d2fed99b49@mail.gmail.com","threadId":"20581","inReplyTo":"4A840B0F.9060003@alum.mit.edu","subject":"Re: rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Bryan O'Sullivan","fromEmail":"bos@serpentine.com","sentAt":"2009-08-13T17:39:21Z","receivedAt":"2009-08-13T17:39:21Z","isPatch":false,"sender":{"key":"bos@serpentine.com","avatar":"https://gravatar.com/avatar/e7371b984ae8671c92ccac9e2c686faafaf40ca062d9b682f5afdc0f8fc999c5?d=mp&s=160"},"body":"On Thu, Aug 13, 2009 at 5:46 AM, Michael Haggerty <mhagger@alum.mit.edu>wrote:\n\n> Sorry to cross-post, but I think this might be interesting to all three\n> projects...\n>\n\nPlease do not cross post, no matter how interesting you think the topic\nmight be.\n\n\n"},{"id":"120576","messageId":"3d6b0edb0908131331l5a3177d1t3d1e1858fc139b2e@mail.gmail.com","threadId":"20581","inReplyTo":"4A840B0F.9060003@alum.mit.edu","subject":"Re: rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Abderrahim Kitouni","fromEmail":"a.kitouni@gmail.com","sentAt":"2009-08-13T20:31:04Z","receivedAt":"2009-08-13T20:31:04Z","isPatch":false,"sender":{"key":"a.kitouni@gmail.com","avatar":null},"body":"2009/8/13 Michael Haggerty <mhagger@alum.mit.edu>:\n> I've been thinking a lot about the problems of tracking upstream changes\n> while developing a feature branch.\nIsn't this the purpose of pbranch [1] (and I beleive bzr's loom[2] and\ngit's topgit[3])?\n\nPeace,\nAbderrahim\n\n[1] http://arrenbrecht.ch/mercurial/pbranch/\n[2] https://launchpad.net/bzr-loom\n[3]http://repo.or.cz/w/topgit.git\n"},{"id":"120597","messageId":"4A849634.1020609@alum.mit.edu","threadId":"20581","inReplyTo":"20090813161256.GA8292@atjola.homenet","subject":"Re: rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2009-08-13T22:39:48Z","receivedAt":"2009-08-13T22:39:48Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Björn Steinbrink wrote:\n> On 2009.08.13 14:46:07 +0200, Michael Haggerty wrote:\n>> (Please note that this technique only works for the typical \"tracking\n>> upstream\" type of rebase; it doesn't help with rebases whose goals are\n>> changing the order of commits, moving only part of a branch, rewriting\n>> commits, etc.)\n> \n> Hm, so that pretty much doesn't work at all for creating a clean patch\n> series, which usually involves rewriting commits, squasing bug fixes\n> into the original commits that introduced the bug etc.\n\nNow that you mention it, there are some other uses of rebase whose\nhistory could be recorded correctly, or at least better, in the DAG.  I\nam not ready to advocate any of these changes, but I think they are\nworth discussing.\n\n\nA squash of two adjacent commits currently transforms this:\n\nA---B1---B2---C\n\nto this\n\nA---B12---C'\n\n(C' has the same contents and commit message as C but a different\nhistory and therefore a different SHA1.)\n\nBut B12 includes both B1 and B2, so the correct history is this:\n\nA---B1---B2----C\n \\         \\    \\\n  ---------B12---C'\n\nThe fact that B12 has B1 and B2 as ancestors tells git that it\nincorporates both of their changes, which correctly describes reality.\n\n\nSplitting a commit (not really an elementary rebase operation, but\nachievable with \"edit\") transforms this:\n\nA---B12---C\n\nto this:\n\nA---B1---B2---C'\n\nIt is not possible to represent this correctly in the DAG, because there\nis no way to express that \"B1\" includes part of \"B12\" as a parent.  But\nthe following would be more accurate than discarding all history:\n\nA---B12---C\n \\     \\   \\\n  B1---B2---C'\n\nIt would of course be difficult for the user-interface layer to be\nconfident that the changes in B1 and B2 are really equivalent to B12\nunless the content of B12 and B2 are identical.\n\n\nInserting a new commit into the history (for example as part of\nreordering later commits) transforms this:\n\nA---C\n\nto this:\n\nA---B---C'\n\nIn this case C' includes everything that is in C, so the correct history is\n\nA---C\n \\   \\\n  B---C'\n\n\nHowever, deleting a commit, which transforms this:\n\nA---B---C\n\nto this:\n\nA---C'\n\ncannot be represented in the DAG, because there is no way to express\nthat C' includes the changes from C without also implying that it\nincludes the changes from B.\n\n\nRewriting a single commit, under the assumption that the rewritten\ncommit is the logical equivalent of the original, transforms this:\n\nA---B1---C\n\nto this:\n\nA---B2---C'\n\nwhere B2 is a hand-rewritten version of B1, and C' is the version of C\nproduced by the rebase.  In this case, the history could be recorded as:\n\nA---B1---C\n \\    \\   \\\n  ----B2---C'\n\nbut again, it is impossible for the user-interface layer to ascertain\nthat B2 is equivalent to B1 without help from the user.\n\n\nAll of this extra history would currently create far more clutter than\nit is worth, but if there were a way to suppress the display of rebased\ncommits (as discussed in the third article I quoted), then the extra\ninformation would be there to help git without overwhelming users.\n\n> And even for just continously forward porting a series of commits, a\n> common case might be that upstream applied some patches, but not all.\n> Can you deal with that?\n> \n> Example:\n> \n> A---B---C (upstream)\n>      \\\n>       H---I---J---K (yours)\n> \n> Upstream takes some changes:\n> \n> A---B---C---I'--K'--D (upstream)\n>      \\\n>       H---I---J---K (yours)\n> \n> rebase leads to:\n> \n> A---B---C---I'--K'--D (upstream)\n>                      \\\n>                       H'--J' (yours)\n> \n> What would your approach generate in that case?\n\nThere *is no way* to represent this history in a DAG, and therefore the\nhistory of this operation will necessarily be lost.  (Well, of course it\ncould be recorded in metadata supplemental to the DAG, but since the\nhistory would not affect future merges it would be pointless.)  The\nproblem is that there is no way to claim that I' is derived from I\nwithout also implying that I' includes the change in H (which it\ndoesn't).  I discuss this sort of thing in another article [1].\n\n>> For more information, please see the full articles: [...]\n> \n> In this one you have two DAGs:\n> (I fixed the second one to also have the merge commit in \"subsystem\"\n> instead of \"topic\", so they only differ WRT to the rebased stuff)\n> \n> A)\n> m---N---m---m---m---m---m---M  (master)\n>      \\                       \\\n>       o---o---O---o---o       o'--o'--o'--o'--o'--S  (subsystem)\n>                        \\                         /\n>                         *---*---*-..........-*--T (topic)\n> \n> \n> B)\n> m---N---m---m---m---m---m---M  (master)\n>      \\                       \\\n>       \\                       o'--o'--o'--o'--o'----------S  (subsystem)\n>        \\                     /   /   /   /   /           /\n>         --------------------o---o---O---o---o---*---*---T (topic)\n> \n> \n> And you say that the former creates problems when you want to merge\n> again. How so?\n\nAs you very clearly showed (thanks!), the merge problems that I claimed\nonly occur in some obscure edge cases.\n\nWhat I *should* have emphasized is that the merge S itself is much more\nprone to conflicts in case A) (with merge base N) than in case B) (with\nthe last \"o\" as merge base).  That is the first advantage of\nrebase-with-history.\n\nAnd please note that I really advocate C), not B):\n\nC)\nm---N---m---m---m---m---m---M  master\n     \\                       \\\n      \\                       o'--o'--o'--o'--o'  subsystem\n       \\                     /   /   /   /   / \\\n        --------------------o---o---O---o---o   \\\n                                             \\   \\\n                                              \\   *'--*'--*'  topic\n                                               \\ /   /   /\n                                                *---*---*\n\nwhere the topic branch is not merged into the subsystem branch but\nrather rebased-with-history.  C) has the significant advantage over A)\nor B) that the topic branch can be converted to a series of patches (the\n*' patches) that apply cleanly to the rebased subsystem branch and can\ntherefore be submitted upstream.  In the case of A) or B), the only\navailable patch that applies cleanly to the rebased subsystem branch is\nS, which is a single commit that squashes together the entire topic\nbranch and is therefore difficult to review.\n\nSo rebasing in a public repository makes it difficult for downstream\ndevelopers to apply their work to the rebased branch (because they have\nto repeat the conflict resolution that was done in the upstream rebase),\nand merging in a topic branch makes it more difficult to create an\neasily-reviewable patch series.  rebase-with-history has neither of\nthese problems.\n\nMichael\n\n[1] \"Git, Mercurial, and Bazaar—simplicity through inflexibility\",\nhttp://softwareswirl.blogspot.com/2009/08/git-mercurial-and-bazaarsimplicity.html\n"},{"id":"120599","messageId":"20090813233027.GA19833@atjola.homenet","threadId":"20581","inReplyTo":"4A849634.1020609@alum.mit.edu","subject":"Re: rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-08-13T23:30:34Z","receivedAt":"2009-08-13T23:30:34Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.08.14 00:39:48 +0200, Michael Haggerty wrote:\n> Björn Steinbrink wrote:\n> > On 2009.08.13 14:46:07 +0200, Michael Haggerty wrote:\n> > And even for just continously forward porting a series of commits, a\n> > common case might be that upstream applied some patches, but not all.\n> > Can you deal with that?\n> > \n> > Example:\n> > \n> > A---B---C (upstream)\n> >      \\\n> >       H---I---J---K (yours)\n> > \n> > Upstream takes some changes:\n> > \n> > A---B---C---I'--K'--D (upstream)\n> >      \\\n> >       H---I---J---K (yours)\n> > \n> > rebase leads to:\n> > \n> > A---B---C---I'--K'--D (upstream)\n> >                      \\\n> >                       H'--J' (yours)\n> > \n> > What would your approach generate in that case?\n> \n> There *is no way* to represent this history in a DAG, and therefore the\n> history of this operation will necessarily be lost.  (Well, of course it\n> could be recorded in metadata supplemental to the DAG, but since the\n> history would not affect future merges it would be pointless.)  The\n> problem is that there is no way to claim that I' is derived from I\n> without also implying that I' includes the change in H (which it\n> doesn't).  I discuss this sort of thing in another article [1].\n\nWell, I' isn't even interesting here. I' is _upstream's_ commit, likely\ncreated from an email. Upstream didn't have \"I\" when I' was created to\nbegin with. The interesting commits are H' and J', which were actually\nrebased. And the situation is even \"worse\" than for I' or K'. H' does\ninclude H, I, and K, but not J and J' contains H, I, J and K.\n\nSo possibly you'd have:\n\nA---B---C---I'--K'--D (upstream)\n     \\               \\\n      \\       --------H'--J'\n       \\     /           /\n        H---I---J---K----\n\nWhich only ignores that K is \"contained\" in H'.\n\nBut you only get to know that I is in H' (and that K is in J') _after_\nH' (or J') have been created. So you'd need a preprocessing run.\n\nThe non-preprocessing result would be something like:\n\nA---B---C---I'--K'--D (upstream)\n     \\               \\\n      \\  -------------H'--J'\n       \\/                /\n        H---I---J--------\n                 \\\n                  K\n\nBut that's obviously total crap.\n\n> >> For more information, please see the full articles: [...]\n> > \n> > In this one you have two DAGs:\n> > (I fixed the second one to also have the merge commit in \"subsystem\"\n> > instead of \"topic\", so they only differ WRT to the rebased stuff)\n> > \n> > A)\n> > m---N---m---m---m---m---m---M  (master)\n> >      \\                       \\\n> >       o---o---O---o---o       o'--o'--o'--o'--o'--S  (subsystem)\n> >                        \\                         /\n> >                         *---*---*-..........-*--T (topic)\n> > \n> > \n> > B)\n> > m---N---m---m---m---m---m---M  (master)\n> >      \\                       \\\n> >       \\                       o'--o'--o'--o'--o'----------S  (subsystem)\n> >        \\                     /   /   /   /   /           /\n> >         --------------------o---o---O---o---o---*---*---T (topic)\n> > \n> > \n> > And you say that the former creates problems when you want to merge\n> > again. How so?\n> \n> As you very clearly showed (thanks!), the merge problems that I claimed\n> only occur in some obscure edge cases.\n> \n> What I *should* have emphasized is that the merge S itself is much more\n> prone to conflicts in case A) (with merge base N) than in case B) (with\n> the last \"o\" as merge base).  That is the first advantage of\n> rebase-with-history.\n\nYeah, but as I said, actually, topic should have been rebased, not\nmerged, and then that ends up the same, as you'd do:\ngit rebase --onto subsystem $last_o topic\n\nTaking the last o commit as upstream.\n\n> And please note that I really advocate C), not B):\n> \n> C)\n> m---N---m---m---m---m---m---M  master\n>      \\                       \\\n>       \\                       o'--o'--o'--o'--o'  subsystem\n>        \\                     /   /   /   /   / \\\n>         --------------------o---o---O---o---o   \\\n>                                              \\   \\\n>                                               \\   *'--*'--*'  topic\n>                                                \\ /   /   /\n>                                                 *---*---*\n> \n\nFine, but this should then be compared to the result from the above\nrebase command, which is:\n\nD)\nm--...--m (master)\n         \\\n          o'-...-o' (subsystem)\n                  \\\n                  *'-...-*' (topic)\n\n> where the topic branch is not merged into the subsystem branch but\n> rather rebased-with-history.  C) has the significant advantage over A)\n> or B) that the topic branch can be converted to a series of patches (the\n> *' patches) that apply cleanly to the rebased subsystem branch and can\n> therefore be submitted upstream.  In the case of A) or B), the only\n> available patch that applies cleanly to the rebased subsystem branch is\n> S, which is a single commit that squashes together the entire topic\n> branch and is therefore difficult to review.\n\nSame for D)\n\n> So rebasing in a public repository makes it difficult for downstream\n> developers to apply their work to the rebased branch (because they have\n> to repeat the conflict resolution that was done in the upstream rebase),\n\nThere's no difference between C) and D) there, except for the fact that\nD) requires you to use --onto, because that needs to differ from\n<upstream>.\n\nLet's take the pre-topic-rebase history:\n\nm---m---m (master)\n     \\   \\\n      \\   o'--o'--O' (subsystem)\n       \\\n        o---o---O---*---*---* (topic)\n\nDoing a plain \"git rebase subsystem topic\" would of course also try to\nrebase the \"o\" commits, so that problematic. Instead, you do:\n\ngit rebase --onto subsystem O topic\n\nThat turns O..topic (the * commits) into patches, and applies them on\ntop of O'. So the \"o\" commits aren't to be rebased.\n\nAnd that's exactly what your rebase-with-history would do as well. Just\nthat O is naturally a common ancestor of subsystem and topic, and so\njust using \"git rebase-w-h subsystem topic\" would be enough. Conflicts\netc. should be 100% the same.\n\nIf you know that your upstream is going to rebase/rewrite history, you\ncan tag (or otherwise mark) the current branching point of your branch,\nso you can easily specify it for the --onto rebase. IOW: This is\nprimarily a social problem (tell your downstream that you rebase this or\nthat branch), but having built-in support to store the branching point\nfor rebasing _might_ be worth a thought.\n\n> and merging in a topic branch makes it more difficult to create an\n> easily-reviewable patch series.  rebase-with-history has neither of\n> these problems.\n\nSure, merging is a no-go if you submit patches by email (or other,\nsimilar means). But you compared that to an \"enhanced\" rebase approach,\ninstead of comparing your rebase approach to the currently available\none.\n\nSo, as I see it, your approach does:\n * Save the need to use --onto, allowing to just specify <upstream>, as\n   if <upstream> was not rewritten.\n * Allows to keep older versions of commits more easily accessible for\n   inspection, e.g. creating interdiffs.\n\nThe latter is (to me) of limited use, and the former could be done by\ntracking the branching point, not sure how well it work out, but maybe\nworth investigating.\n\nBjörn\n"},{"id":"120600","messageId":"2e24e5b90908132017i1c6be9abt9b08219acc1cb600@mail.gmail.com","threadId":"20581","inReplyTo":"4A849634.1020609@alum.mit.edu","subject":"Re: rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2009-08-14T03:17:02Z","receivedAt":"2009-08-14T03:17:02Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"Hi,\n\nI'm one of those wannabe experts who thinks he knows enough about git\nto teach people in his workplace but obviously pales in this group,\nbut with that caveat, let me say:\n\n2009/8/14 Michael Haggerty <mhagger@alum.mit.edu>:\n\n> Now that you mention it, there are some other uses of rebase whose\n> history could be recorded correctly, or at least better, in the DAG.  I\n> am not ready to advocate any of these changes, but I think they are\n\nI see you've made your own caveat :-)\n\n> worth discussing.\n\n[snip]\n\n> A---B1---B2----C\n>  \\         \\    \\\n>  ---------B12---C'\n\n> A---B12---C\n>  \\     \\   \\\n>  B1---B2---C'\n\n[snip]\n\n> A---C\n>  \\   \\\n>  B---C'\n\n[etc etc... many such snipped]\n\nTo me, the ability to *forget* the mistakes I made (for whatever\ndefinition of \"mistake\" you wish) as long as it's private to my repo,\nis one of the main attractions of git.  I'm one of those guys who\nsaves early, and saves often, when editing files.  This translates to\ncommit early, commit often, in the git world.\n\nI see no earthly reason why I would ever *want* those commits\npreserved, so I hope that, if this sort of thing ever gets into the\ncode, it is definitely *not* the default :-)  It is not sufficient for\nme that the GUI knows how to suppress their display, it is necessary\nthat they *disappear completely*.\n\nAnd that reminds me.  You often hear people on #git ask how to get rid\nof some files (maybe containing passwords etc) that inadvertently got\ninto the repo, and the answer, a lot of the time, is filter-branch,\nbecause the \"bad\" commit is pretty old.  I suspect that for every\nperson who asks that question on the list because he already pushed,\nthere are 4 who discovered such an error much earlier, (when the file\nwent into only a couple of commits at the top maybe), did a rebase -i\nwith \"edit\" or whatever, and got rid of the evidence, err I mean\npassword :-)  If this sort of thing were to be the default, they'd\nhave to use a filter-branch even for such simple cases.\n\nFinally, speaking as someone who teaches git, this adds enormous\ncomplexity to the basic concepts.  Complexity is good when the\nbenefits are obvious, but to me they are not obvious [see *my* caveat\nat the top before you react to this statement]\n"},{"id":"120663","messageId":"4A85D53D.9050805@alum.mit.edu","threadId":"20581","inReplyTo":"20090813233027.GA19833@atjola.homenet","subject":"Re: rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2009-08-14T21:21:01Z","receivedAt":"2009-08-14T21:21:01Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Björn Steinbrink wrote:\n> On 2009.08.14 00:39:48 +0200, Michael Haggerty wrote:\n>> Björn Steinbrink wrote:\n>>> On 2009.08.13 14:46:07 +0200, Michael Haggerty wrote:\n>>> And even for just continously forward porting a series of commits, a\n>>> common case might be that upstream applied some patches, but not all.\n>>> Can you deal with that?\n\n[A discussion of various unsatisfactory approaches omitted...]\n\n> But that's obviously total crap.\n\nSo I think we agree that it is not possible to retain history for a case\nlike this (which is essentially a general cherry-pick).\n\n> [...]\n> Doing a plain \"git rebase subsystem topic\" would of course also try to\n> rebase the \"o\" commits, so that problematic. Instead, you do:\n> \n> git rebase --onto subsystem O topic\n> \n> That turns O..topic (the * commits) into patches, and applies them on\n> top of O'. So the \"o\" commits aren't to be rebased.\n> \n> And that's exactly what your rebase-with-history would do as well. Just\n> that O is naturally a common ancestor of subsystem and topic, and so\n> just using \"git rebase-w-h subsystem topic\" would be enough. Conflicts\n> etc. should be 100% the same.\n> \n> If you know that your upstream is going to rebase/rewrite history, you\n> can tag (or otherwise mark) the current branching point of your branch,\n> so you can easily specify it for the --onto rebase. IOW: This is\n> primarily a social problem (tell your downstream that you rebase this or\n> that branch), but having built-in support to store the branching point\n> for rebasing _might_ be worth a thought.\n\nRecording branch points manually, coordinating merges via email -- OMG\nyou are giving me flashbacks of CVS ;-)\n\n*Of course* you can get around all of these problems if you put the\nburden of bookkeeping on the user.  The whole point of\nrebase-with-history is to have the VCS handle it automatically!\n\n>> and merging in a topic branch makes it more difficult to create an\n>> easily-reviewable patch series.  rebase-with-history has neither of\n>> these problems.\n> \n> Sure, merging is a no-go if you submit patches by email (or other,\n> similar means). But you compared that to an \"enhanced\" rebase approach,\n> instead of comparing your rebase approach to the currently available\n> one.\n\nIn [1] I compared rebase-with-history with both of the\ncurrently-available options (rebase and merge).  Rebase and merge can\neach deal with some of the issues that come up, but each one falls flat\non others.  I believe that rebase-with-history has the advantages of both.\n\nThe example in [2] was taken straight from the git-rebase man page [3];\nI did not want to claim that current practice would use merging in this\nsituation, but rather just to show that rebase-with-history removes the\npain from this well-known example.\n\nI think we are mostly in agreement.  Rebase-with-history is obviously\nnot an earth-shattering revolution in DVCS technology, but my hope is\nthat it could unobtrusively assist with a few minor pain points.\n\nMichael\n\n[1]\nhttp://softwareswirl.blogspot.com/2009/04/truce-in-merge-vs-rebase-war.html\n[2]\nhttp://softwareswirl.blogspot.com/2009/08/upstream-rebase-just-works-if-history.html\n[3] http://www.kernel.org/pub/software/scm/git/docs/git-rebase.html\n"},{"id":"120666","messageId":"20090815064001.6117@nanako3.lavabit.com","threadId":"20581","inReplyTo":"4A85D53D.9050805@alum.mit.edu","subject":"Re: rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Nanako Shiraishi","fromEmail":"nanako3@lavabit.com","sentAt":"2009-08-14T21:40:01Z","receivedAt":"2009-08-14T21:40:01Z","isPatch":false,"sender":{"key":"nanako3@lavabit.com","avatar":"https://gravatar.com/avatar/3777b9e201c5883a62b1a6fdf7c53f2d712d1d80989146063ea861e33aad72a8?d=mp&s=160"},"body":"Quoting Michael Haggerty <mhagger@alum.mit.edu>\n\n> In [1] I compared rebase-with-history with both of the\n> currently-available options (rebase and merge).  Rebase and merge can\n> each deal with some of the issues that come up, but each one falls flat\n> on others.  I believe that rebase-with-history has the advantages of both.\n> ....  Rebase-with-history is obviously\n> not an earth-shattering revolution in DVCS technology, but my hope is\n> that it could unobtrusively assist with a few minor pain points.\n\nThe saddest part is that your [1] works only in a case a user can easily handle manually, and doesn't help cases more complex than the most trivial ones, such as reordering and squashing commits, where the user may benefit if an automated support from VCS were available.\n\n-- \nNanako Shiraishi\nhttp://ivory.ap.teacup.com/nanako3/\n"},{"id":"120687","messageId":"20090815033623.GB19833@atjola.homenet","threadId":"20581","inReplyTo":"4A85D53D.9050805@alum.mit.edu","subject":"Re: rebase-with-history -- a technique for rebasing without trashing your repo history","fromName":"Björn Steinbrink","fromEmail":"b.steinbrink@gmx.de","sentAt":"2009-08-15T03:36:23Z","receivedAt":"2009-08-15T03:36:23Z","isPatch":false,"sender":{"key":"b.steinbrink@gmx.de","avatar":"https://avatars.githubusercontent.com/u/230962?v=4"},"body":"On 2009.08.14 23:21:01 +0200, Michael Haggerty wrote:\n> Björn Steinbrink wrote:\n> > On 2009.08.14 00:39:48 +0200, Michael Haggerty wrote:\n> >> Björn Steinbrink wrote:\n> >>> On 2009.08.13 14:46:07 +0200, Michael Haggerty wrote:\n> > [...]\n> > Doing a plain \"git rebase subsystem topic\" would of course also try to\n> > rebase the \"o\" commits, so that problematic. Instead, you do:\n> > \n> > git rebase --onto subsystem O topic\n> > \n> > That turns O..topic (the * commits) into patches, and applies them on\n> > top of O'. So the \"o\" commits aren't to be rebased.\n> > \n> > And that's exactly what your rebase-with-history would do as well. Just\n> > that O is naturally a common ancestor of subsystem and topic, and so\n> > just using \"git rebase-w-h subsystem topic\" would be enough. Conflicts\n> > etc. should be 100% the same.\n> > \n> > If you know that your upstream is going to rebase/rewrite history, you\n> > can tag (or otherwise mark) the current branching point of your branch,\n> > so you can easily specify it for the --onto rebase. IOW: This is\n> > primarily a social problem (tell your downstream that you rebase this or\n> > that branch), but having built-in support to store the branching point\n> > for rebasing _might_ be worth a thought.\n> \n> Recording branch points manually, coordinating merges via email -- OMG\n> you are giving me flashbacks of CVS ;-)\n\nNot merging, but rewriting history. One of the primary purposes of\nrebasing is to forget the old history, the new version overrides it. And\ntelling someone to forget something is a social problem. You can help\nthe user to forget the history by tracking the branching points and I\nsaid that git could maybe learn to do that, so the user doesn't have\nto do so. Quick idea:\n\nOn branch creation, create refs/bases/<branchname> (let's call that\n<base>) referencing the commit the branch initially references.\n\nOn rebase, check if <branchname>..<onto> is not empty. If so, update\nrefs/bases/<branchname> to reference <base>.\n\nOn reset, check if the commit the branch head is being reset to is\nreachable through the commit the branch head currently references. If\nnot, update <base> to reference the commit we're resetting to.\n\nFind some sane syntax for rebase that implicitly uses <base> as the\n<upstream> argument, e.g. just \"git rebase --onto <whatever>\" could work\nas \"git rebase --onto <whatever> <base>\".\n\nMost likely, I missed a bunch of corner cases though...\n\n> *Of course* you can get around all of these problems if you put the\n> burden of bookkeeping on the user.  The whole point of\n> rebase-with-history is to have the VCS handle it automatically!\n\nWhat your approach does, is simply moving the \"just forget the\nhistory\" part. Instead of forgetting it at rebase time, you have to\nforget it when you want to submit patches. It's obviously a bit easier\nthough, as you can just say \"--first-parent <upstream>\", assuming that\nyou teach format-patch to use a special first-parent diff mode for the\nmerge commits (see below).\n\n> >> and merging in a topic branch makes it more difficult to create an\n> >> easily-reviewable patch series.  rebase-with-history has neither of\n> >> these problems.\n> > \n> > Sure, merging is a no-go if you submit patches by email (or other,\n> > similar means). But you compared that to an \"enhanced\" rebase approach,\n> > instead of comparing your rebase approach to the currently available\n> > one.\n> \n> In [1] I compared rebase-with-history with both of the\n> currently-available options (rebase and merge).  Rebase and merge can\n> each deal with some of the issues that come up, but each one falls flat\n> on others.  I believe that rebase-with-history has the advantages of both.\n\nAnd some disadvantages.\n\n1) Cluttered history, which needs to be rewritten again when the emailed\npatches are just for review, but the maintainer will actually merge from\nyou later.\n\nTaking the old master, subsystem, topic example, you get (for example):\n\n          o2--o2 (subsystem)\n         /     \\\nm---m---m---m---m (master)\n     \\   \\\n      \\   o'--o'\n       \\ /   / \\\n        o---o   *'--*' (topic)\n             \\ /   /\n              *---*\n\nNow the user that maintains \"topic\" is back at the hard case. He now\nneeds to rebase onto master, using the last o' as <upstream>. The\nDAG doesn't help here, the base-tracking would handle that.\n\n\n2) Merge commits, which are usually displayed in a special format. So\nfor \"git show\" or \"git log -p\" to give useful output for those special\nmerges, you'd have to introduce a new \"diff only against first-parent\"\nmode, and mark those merge in a special way, so that diff mode is used\nfor them, but not for real merges. And users of old git versions would\nhave to deal with the basic -m merge diff mode, ignoring the useless\ndiff for the second parent and the fact that the real merges also get\nshown in that format. The base tracking doesn't have this problem\neither.\n\n\n> The example in [2] was taken straight from the git-rebase man page [3];\n> I did not want to claim that current practice would use merging in this\n> situation, but rather just to show that rebase-with-history removes the\n> pain from this well-known example.\n\nWell, the man pages says: Don't merge, rebase needs to trickle down, but\nyou'll likely need to use \"git rebase --onto subsystem subsystem@{1}\".\nSo the rebase-with-history really just saves that \"use --onto and the\nright <upstream>\" from the hard case. The plain base-tracking does the\nsame.\n\nAnother way to reach the same goal would be just to explictly override\nthe old history.\n\nm---m---m (master)\n     \\\n      o---o (subsystem)\n           \\\n            *---* (topic)\n\n(Hypothetical): git rebase --override master subsystem\n\nLeads to:\n\nm---m---m---- (master)\n     \\       \\\n      o---o---O---o'--o' (subsystem)\n           \\\n            *---* (topic)\n\nWhere O is an --ours merge, that just marks the old o commits as merged,\nbut has the same tree as the last m commit.\n\nNow topic can be rebase using: git rebase --override subsystem topic\n\nm---m---m---- (master)\n     \\       \\\n      o---o---O---o'--o' (subsystem)\n           \\           \\\n            *---*-------X---*'--*' (topic)\n\nAgain, X being an ours merge.\n\nAs the O and X commits have the last o and * commits as their second\nparents, this even doesn't break things like \"git show\" and \"git log\n-p\", as the interesting commits aren't merge commits. So \"git\nlog -p --first-parent subsystem..topic\" would do the right thing\n(optionally with --no-merges to avoid the merge commit, but seeing that\ndoesn't hurt that much I guess).\n\nThis also trivially supports the reorder, squash, edit whatever stuff,\nas it doesn't rely on 1:1 commit counterparts to exist. But it also\nfalls flat on its face as soon as subsystem gets \"really\" rewritting, so\nthat the old history is no longer reachable from the new history.\n\nBjörn\n"}]}