{"thread":{"id":"13940","subject":"RFC: rebase without pain","startedAt":"2008-06-14T01:38:44Z","lastAt":"2008-06-15T19:36:34Z","messageCount":3,"participants":["Luke Lu","Dmitry Potapov"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"79800","messageId":"5B4BD573-8C89-4E27-8ADB-F870EA503D00@vicaya.com","threadId":"13940","inReplyTo":null,"subject":"RFC: rebase without pain","fromName":"Luke Lu","fromEmail":"git@vicaya.com","sentAt":"2008-06-14T01:38:44Z","receivedAt":"2008-06-14T01:38:44Z","isPatch":false,"sender":{"key":"git@vicaya.com","avatar":null},"body":"This may have been discussed before, but I could not find it. If so, I  \napologize for the noise and hope somebody is working on the issue.\n\nBased on my observation, rebase is the single most interesting and  \nmisunderstood feature in git compared with other VCS. Once I  \ndiscovered rebase -i, I can't stop using it, because I'd like to keep  \nmy history clean for readability and maintenance purpose. The downside  \nis that once you publish your branch, further rebasing could cause a  \nlot of pain for people who have already rebased and merged this  \nbranch, as rebasing/merging against the branch by them would cause a  \nlot conflicts that have to be resolved manually due to some loss of  \ncommon history (one or more SHA1s were rewritten).\n\nOf these lost SHA1s, the last SHA1 of the last rebase is most  \ninteresting, as it's usually the parent of other people's rebase or  \nmerge. The same applies to a series of patches that are being modified  \nover time. If we preserve a  \n<last_sha1_of_the_last_rebase_or_original_head> record, say in reflog,  \nfor every rebase, we can have more intelligent merges of rebased  \nhistories with much less conflicts (specifically, squashing,  \nreordering and comment editing will result zero conflicts for  \ndownstream users). We probably need to extend git protocol to send  \nalong these rebase records. We can even modify format-patch to  \ngenerate these rebase records for each rebase of a branch. We also  \nneed to modify gc code to reserve these records in reflog (along with  \nstashes as discussed in another thread). I think this approach can be  \nmade backward compatible.\n\nOne obvious problem of this approach is that people can branch in the  \nmiddle of a patch series after they rebased/merged it, which means  \nthey'll have to resolve conflicts manually even with this approach. I  \nargue that this is a rare case and that the above approach works for  \nmost common cases.\n\nIn summary, the proposal tries to solve the rebase problems by  \npreserving and propagating rebase history and let the merge algorithm  \ntake advantage of the information.\n\nWhat do you guys think?\n\nDisclaimer: I don't know much about git internal yet. My understanding  \nof git is only at conceptual level (commit DAG and reflog)\n\n__Luke\n"},{"id":"79851","messageId":"20080614111758.GB5737@dpotapov.dyndns.org","threadId":"13940","inReplyTo":"5B4BD573-8C89-4E27-8ADB-F870EA503D00@vicaya.com","subject":"Re: RFC: rebase without pain","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-06-14T11:17:58Z","receivedAt":"2008-06-14T11:17:58Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Fri, Jun 13, 2008 at 06:38:44PM -0700, Luke Lu wrote:\n> This may have been discussed before, but I could not find it. If so, I  \n> apologize for the noise and hope somebody is working on the issue.\n\nI think we have had a somewhat similar discussion not so long ago.\nIt was called \"inexplicable failure to merge recursively across\ncherry-picks\". I think you can find it here:\nhttp://kerneltrap.org/mailarchive/git/2007/10/9/333729\n\nPlease, read carefully this Linus' posts:\nhttp://kerneltrap.org/mailarchive/git/2007/10/10/334129\nhttp://kerneltrap.org/mailarchive/git/2007/10/11/335451\n\n> \n> Based on my observation, rebase is the single most interesting and  \n> misunderstood feature in git compared with other VCS. Once I  \n> discovered rebase -i, I can't stop using it, because I'd like to keep  \n> my history clean for readability and maintenance purpose.\n\nThe downside of rebase is that you are *re-writing* branch history. It\nis okay when you do that in your private branch, but when you publish\nsomething there is no way back. It is like when you prepare on some\narticle, you can make a lot of drafts but when you publish it then\nit is published. Any attempt to falsify history will cause a lot of\nconfusion. Also, please, notice that even if a branch was rebased\nwithout a single conflict, it does not mean that it will work. So,\nyou can break things just by rebasing and it will be impossible to\nfind later who caused the breakage. Sometimes, even if the final\nstate after rebase is working, the intermediate commits may not work\nor even not compile.\n\nSo, I don't think that rebasing published history is a good idea.\n\nDmitry\n"},{"id":"79949","messageId":"C8DA824B-97D6-46FB-8CD4-F66458FBC7B2@vicaya.com","threadId":"13940","inReplyTo":"20080614111758.GB5737@dpotapov.dyndns.org","subject":"Re: RFC: rebase without pain","fromName":"Luke Lu","fromEmail":"git@vicaya.com","sentAt":"2008-06-15T19:36:34Z","receivedAt":"2008-06-15T19:36:34Z","isPatch":false,"sender":{"key":"git@vicaya.com","avatar":null},"body":"On Jun 14, 2008, at 4:17 AM, Dmitry Potapov wrote:\n> On Fri, Jun 13, 2008 at 06:38:44PM -0700, Luke Lu wrote:\n>> This may have been discussed before, but I could not find it. If  \n>> so, I\n>> apologize for the noise and hope somebody is working on the issue.\n>\n> I think we have had a somewhat similar discussion not so long ago.\n> It was called \"inexplicable failure to merge recursively across\n> cherry-picks\". I think you can find it here:\n> http://kerneltrap.org/mailarchive/git/2007/10/9/333729\n>\n> Please, read carefully this Linus' posts:\n> http://kerneltrap.org/mailarchive/git/2007/10/10/334129\n> http://kerneltrap.org/mailarchive/git/2007/10/11/335451\n\nThanks for the pointers, Dmitry. They are indeed illuminating. I  \nactually agree with Linus on all accounts here.\n\n>> Based on my observation, rebase is the single most interesting and\n>> misunderstood feature in git compared with other VCS. Once I\n>> discovered rebase -i, I can't stop using it, because I'd like to keep\n>> my history clean for readability and maintenance purpose.\n>\n> The downside of rebase is that you are *re-writing* branch history.\n\nI think the word \"history\" is too overloaded in VCS world. Sometimes  \nit really means the actually chronological steps developer performed  \nto solve a problem. Sometimes it just means a portion of commit/patch  \nDAG, that is, a series of transformations (or functions if you prefer)  \nthat given an input (e.g., a parent of a tree) will produce a  \ndeterministic outcome. The two semantics are actually orthogonal,  \ndespite the fact they're often identical in reality. I think it's a  \nmisuse of the word \"history\" here, when you actually mean a set of  \npatches. You can never rewrite history by definition of the original  \nmeaning of history because you can't turn back time. The real history  \nof git is kept in the reflog.\n\nRebasing is rewriting a set of patches. It's a form of  \n(meta)programing to compose and decompose features. The problem of  \nrebasing (in addition to merge) is that we don't have proper tools to  \ntrack rebase itself. Even though the history is in the reflog, the  \ninformation is only kept temporarily and not propagated and used  \nthrough fetch/merge.\n\n> It\n> is okay when you do that in your private branch, but when you publish\n> something there is no way back. It is like when you prepare on some\n> article, you can make a lot of drafts but when you publish it then\n> it is published. Any attempt to falsify history will cause a lot of\n> confusion.\n\nI think the confusion mainly comes from the conflation of meanings of  \nthe word \"history\" and merge conflicts that result from lack of tool  \nsupport.\n\n> Also, please, notice that even if a branch was rebased\n> without a single conflict, it does not mean that it will work.\n\nThe same applies to usual merges as well. The advantage of merge is  \nthat it maintains the original ancestry of the commits, so it's  \nrelatively easy to visualize and debug merge problems.\n\n> So, you can break things just by rebasing\n\nSo can you by just straight merging.\n\n> and it will be impossible to\n> find later who caused the breakage.\n\nYes, that's the real problem. But it's mainly caused by lack of tools  \nto track rebase. I think we should probably put a 'rebase' node in the  \ncommit just like merge. The rebase commit will contain enough  \ninformation to track the rebase. git log can display who did the  \nrebase and gitk can even visualize the graph transformation.\n\n> Sometimes, even if the final state after rebase is working, the  \n> intermediate commits may not work\n> or even not compile.\n\nYes, that's another problem, but with rebase -i, we can hopefully fix  \nthem :)\n\n> So, I don't think that rebasing published history is a good idea.\n\nYes, rewriting published patch set (again you can't rewrite history,  \npublic or not) without proper tool to track them is definitely not a  \ngood idea. But don't you think we need to develop tools to track  \nrebase properly?\n\nOne common use case would be maintaining a patch set against a release  \npoint. It's already a common practice with or without VCS support:  \nPeople release some software version n, then release a giant patch to  \ncover several bugs; later on, they realize that they need to split the  \npatch for each bug and vice versa. That's rebase right there. But  \npeople are not really confused because they know they need to reapply  \nthe patch set against an official release.\n\nYou can't do that easily with git by simply pulling from upstream, yet.\n\n__Luke\n"}]}