{"thread":{"id":"47487","subject":"Bring together merge and rebase","startedAt":"2017-12-23T06:14:45Z","lastAt":"2018-01-06T21:38:59Z","messageCount":44,"participants":["Carl Baldwin","Ævar Arnfjörð Bjarmason","Randall S. Becker","Johannes Schindelin","Alexei Lozovsky","Theodore Ts'o","Jacob Keller","Paul Smith","Mike Hommey","Igor Djordjevic","Martin Fick","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"335218","messageId":"CALiLy7pBvyqA+NjTZHOK9t0AFGYbwqwRVD3sZjUg0ZLx5y1h3A@mail.gmail.com","threadId":"47487","inReplyTo":null,"subject":"Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-23T06:10:19Z","receivedAt":"2017-12-23T06:14:45Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"The big contention among git users is whether to rebase or to merge\nchanges [2][3] while iterating. I used to firmly believe that merging\nwas the way to go and rebase was harmful. More recently, I have worked\nin some environments where I saw rebase used very effectively while\niterating on changes and I relaxed my stance a lot. Now, I'm on the\nfence. I appreciate the strengths and weaknesses of both approaches. I\nwaffle between the two depending on the situation, the tools being\nused, and I guess, to some extent, my mood.\n\nI think what git needs is something brand new that brings the two\ntogether and has all of the advantages of both approaches. Let me\nexplain what I've got in mind...\n\nI've been calling this proposal `git replay` or `git replace` but I'd\nlike to hear other suggestions for what to name it. It works like\nrebase except with one very important difference. Instead of orphaning\nthe original commit, it keeps a pointer to it in the commit just like\na `parent` entry but calls it `replaces` instead to distinguish it\nfrom regular history. In the resulting commit history, following\n`parent` pointers shows exactly the same history as if the commit had\nbeen rebased. Meanwhile, the history of iterating on the change itself\nis available by following `replaces` pointers. The new commit replaces\nthe old one but keeps it around to record how the change evolved.\n\nThe git history now has two dimensions. The first shows a cleaned up\nhistory where fix ups and code review feedback have been rolled into\nthe original changes and changes can possibly be ordered in a nice\nlinear progression that is much easier to understand. The second\ndrills into the history of a change. There is no loss and you don't\nchange history in a way that will cause problems for others who have\nthe older commits.\n\nReplay handles collaboration between multiple authors on a single\nchange. This is difficult and prone to accidental loss when using\nrebase and it results in a complex history when done with merge. With\nreplay, collaborators could merge while collaborating on a single\nchange and a record of each one's contributions can be preserved.\nAttempting this level of collaboration caused me many headaches when I\nworked with the gerrit workflow (which in many ways, I like a lot).\n\nI blogged about this proposal earlier this year when I first thought\nof it [1]. I got busy and didn't think about it for a while. Now with\na little time off of work, I've come back to revisit it. The blog\nentry has a few examples showing how it works and how the history will\nlook in a few examples. Take a look.\n\nVarious git commands will have to learn how to handle this kind of\nhistory. For example, things like fetch, push, gc, and others that\nmove history around and clean out orphaned history should treat\nanything reachable through `replaces` pointers as precious. Log and\nrelated history commands may need new switches to traverse the history\ndifferently in different situations. Bisect is a interesting one. I\ntend to think that bisect should prefer the regular commit history but\nhave the ability to drill into the change history if necessary.\n\nIn my opinion, this proposal would bring together rebase and merge in\na powerful way and could end the contention. Thanks for your\nconsideration.\n\nCarl Baldwin\n\n[1] http://blog.episodicgenius.com/post/merge-or-rebase--neither/\n[2] https://git-scm.com/book/en/v2/Git-Branching-Rebasing\n[3] http://changelog.complete.org/archives/586-rebase-considered-harmful\n"},{"id":"335231","messageId":"877etds220.fsf@evledraar.gmail.com","threadId":"47487","inReplyTo":"CALiLy7pBvyqA+NjTZHOK9t0AFGYbwqwRVD3sZjUg0ZLx5y1h3A@mail.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2017-12-23T18:59:35Z","receivedAt":"2017-12-23T18:59:47Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, Dec 23 2017, Carl Baldwin jotted:\n\n> The big contention among git users is whether to rebase or to merge\n> changes [2][3] while iterating. I used to firmly believe that merging\n> was the way to go and rebase was harmful. More recently, I have worked\n> in some environments where I saw rebase used very effectively while\n> iterating on changes and I relaxed my stance a lot. Now, I'm on the\n> fence. I appreciate the strengths and weaknesses of both approaches. I\n> waffle between the two depending on the situation, the tools being\n> used, and I guess, to some extent, my mood.\n>\n> I think what git needs is something brand new that brings the two\n> together and has all of the advantages of both approaches. Let me\n> explain what I've got in mind...\n>\n> I've been calling this proposal `git replay` or `git replace` but I'd\n> like to hear other suggestions for what to name it. It works like\n> rebase except with one very important difference. Instead of orphaning\n> the original commit, it keeps a pointer to it in the commit just like\n> a `parent` entry but calls it `replaces` instead to distinguish it\n> from regular history. In the resulting commit history, following\n> `parent` pointers shows exactly the same history as if the commit had\n> been rebased. Meanwhile, the history of iterating on the change itself\n> is available by following `replaces` pointers. The new commit replaces\n> the old one but keeps it around to record how the change evolved.\n>\n> The git history now has two dimensions. The first shows a cleaned up\n> history where fix ups and code review feedback have been rolled into\n> the original changes and changes can possibly be ordered in a nice\n> linear progression that is much easier to understand. The second\n> drills into the history of a change. There is no loss and you don't\n> change history in a way that will cause problems for others who have\n> the older commits.\n>\n> Replay handles collaboration between multiple authors on a single\n> change. This is difficult and prone to accidental loss when using\n> rebase and it results in a complex history when done with merge. With\n> replay, collaborators could merge while collaborating on a single\n> change and a record of each one's contributions can be preserved.\n> Attempting this level of collaboration caused me many headaches when I\n> worked with the gerrit workflow (which in many ways, I like a lot).\n>\n> I blogged about this proposal earlier this year when I first thought\n> of it [1]. I got busy and didn't think about it for a while. Now with\n> a little time off of work, I've come back to revisit it. The blog\n> entry has a few examples showing how it works and how the history will\n> look in a few examples. Take a look.\n>\n> Various git commands will have to learn how to handle this kind of\n> history. For example, things like fetch, push, gc, and others that\n> move history around and clean out orphaned history should treat\n> anything reachable through `replaces` pointers as precious. Log and\n> related history commands may need new switches to traverse the history\n> differently in different situations. Bisect is a interesting one. I\n> tend to think that bisect should prefer the regular commit history but\n> have the ability to drill into the change history if necessary.\n>\n> In my opinion, this proposal would bring together rebase and merge in\n> a powerful way and could end the contention. Thanks for your\n> consideration.\n>\n> Carl Baldwin\n>\n> [1] http://blog.episodicgenius.com/post/merge-or-rebase--neither/\n> [2] https://git-scm.com/book/en/v2/Git-Branching-Rebasing\n> [3] http://changelog.complete.org/archives/586-rebase-considered-harmful\n\nI think this is a worthwhile thing to implement, there are certainly\nuse-cases where you'd like to have your cake & eat it too as it were,\ni.e. have a nice rebased history in \"git log\", but also have the \"raw\"\nhistory for all the reasons the fossil people like to talk about, or for\nsome compliance reasons.\n\nBut I don't see why you think this needs a new \"replaces\" parent pointer\northagonal to parent pointers, i.e. something that would need to be a\nnew field in the commit object (I may have misread the proposal, it's\nnot heavy on technical details).\n\nConsider a merge use case like this:\n\n          A---B---C topic\n         /         \\\n    D---E---F---G---H master\n\nHere we worked on a topic with commits A,B & C, maybe we regret not\nsquashing B into A, but it gives us the \"raw\" history. Instead we might\nrebase it like this:\n\n          A+B---C topic\n         /\n    G---H master\n\nNow we can push \"topic\" to master, but as you've noted this loses the\nraw history, but now consider doing this instead:\n\n          A---B---C   A2+B2---C2 topic\n         /         \\ /\n    D---E---F---G---G master\n\nI.e. you could have started working on commit A/B/C, now you \"git\nreplace\" them (which would be some fancy rebase alias), and what it'll\ndo is create a merge commit that entirely resolves the conflict so that\nhte tree is equivalent to what \"master\" was already at. Then you rewrite\nthem and re-apply them on top.\n\nIf you run \"git log\" it will already ignore A,B,C unless you specify\n--full-history, so git already knows to ignore these sort of side\nhistories that result in no changes on the branch they got merged\ninto. I don't know about bisect, but if it's not doing something similar\nalready it would be easy to make it do so.\n\nYou could even add a new field to the commit object of A2+B2 & C2 which\nwould be one or more of \"replaces <sha1 of A/B/C>\", commit objects\nsupport adding arbitrary new fields without anything breaking.\n\nBut most importantly, while I think this gives you the same things from\na UX level, it doesn't need any changes to fetch, push, gc or whatever,\nsince it's all stuff we support today, someone just needs to hack\n\"rebase\" to create this sort of no-op merge commit to take advantage of\nit.\n"},{"id":"335233","messageId":"20171223210141.GA24715@hpz.ecbaldwin.net","threadId":"47487","inReplyTo":"877etds220.fsf@evledraar.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-23T21:01:42Z","receivedAt":"2017-12-23T21:01:51Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Sat, Dec 23, 2017 at 07:59:35PM +0100, Ævar Arnfjörð Bjarmason wrote:\n> I think this is a worthwhile thing to implement, there are certainly\n> use-cases where you'd like to have your cake & eat it too as it were,\n> i.e. have a nice rebased history in \"git log\", but also have the \"raw\"\n> history for all the reasons the fossil people like to talk about, or for\n> some compliance reasons.\n\nThank you kindly for your reply. I do think we can have the cake and eat\nit too in this case. At a high level, what you describe above is what\nI'm after. I'm sorry if I left something out or was unclear. I hoped to\nkeep my original post brief. Maybe it was too brief to be useful.\nHowever, I'd like to follow up and be understood.\n\n> But I don't see why you think this needs a new \"replaces\" parent pointer\n> orthagonal to parent pointers, i.e. something that would need to be a\n> new field in the commit object (I may have misread the proposal, it's\n> not heavy on technical details).\n\nJust to clarify, I am proposing a new \"replaces\" pointer in the commit\nobject. Imagine starting with rebase exactly as it works today. This new\nfield would be inserted into any new commit created by a rebase command\nto reference the original commit on which it was based. Though, I'm not\nsure if it would be better to change the behavior of the existing rebase\ncommand, provide a switch or config option to turn it on, or provide a\nnew command entirely (e.g. git replay or git replace) to avoid\ncompatibility issues with the existing rebase.\n\nI imagine that a \"git commit --amend\" would also insert a \"replaces\"\nreference to the original commit but I failed to mention that in my\noriginal post. The amend use case is similar to adding a fixup commit\nand then doing a squash in interactive mode.\n\n> Consider a merge use case like this:\n> \n>           A---B---C topic\n>          /         \\\n>     D---E---F---G---H master\n\nThis is a bit different than the use cases that I've had in mind. You\nshow that the topic has already merged to master. I have imagined this\nproposal being useful before the topic becomes a part of the master\nbranch. I'm thinking in the context of something like a github pull\nrequest under active development and review or a gerrit review. So, at\nthis point, we still look like this:\n\n          A---B---C topic\n         /\n    D---E---F---G\n\n> Here we worked on a topic with commits A,B & C, maybe we regret not\n> squashing B into A, but it gives us the \"raw\" history. Instead we might\n> rebase it like this:\n> \n>           A+B---C topic\n>          /\n>     G---H master\n\nSince H already merged the topic. I'm not sure what the A+B and C\ncommits are doing.\n\nAt the point where I have C and G above, let's say I regret not having\nsquashed A and B as you suggested. My proposal would end up as I draw\nbelow where the primes are the new versions of the commits (A' is A+B).\nBare with me, I'm not sure the best way to draw this in ascii. It has\nthat orthogoal dimension that makes the ascii drawings a little more\ncomplex: (I left out the parent of A' which is still E)\n\n       A--B---C\n        \\ |    \\                    <- \"replaces\" rather than \"parent\"\n         -A'----C' topic\n         /\n    D---E---F---G master\n\nWe can continue by actually changing the base. All of these commits are\nkept, I just drop them from the drawings to avoid getting too complex.\n\n                A'--C'\n                 \\   \\              <- \"replaces\" rather than \"parent\"\n                  A\"--C\" topic\n                 /\n    D---E---F---G master\n\nNormal git log operations would ignore them by default. When finally\nmerging to master, it ends up very simple (by default) but the history\nis still there to support archealogic operations.\n\n    D---E---F---G---A\"--C\" master\n\n> Now we can push \"topic\" to master, but as you've noted this loses the\n> raw history, but now consider doing this instead:\n> \n>           A---B---C   A2+B2---C2 topic\n>          /         \\ /\n>     D---E---F---G---G master\n\nThere are two Gs in this drawing. Should the second be H? Sorry, I'm\njust trying to understanding the use case you're describing and I don't\nunderstand it yet which makes it difficult to comment on the rest of\nyour reply.\n\n> I.e. you could have started working on commit A/B/C, now you \"git\n> replace\" them (which would be some fancy rebase alias), and what it'll\n> do is create a merge commit that entirely resolves the conflict so that\n> hte tree is equivalent to what \"master\" was already at. Then you rewrite\n> them and re-apply them on top.\n> \n> If you run \"git log\" it will already ignore A,B,C unless you specify\n> --full-history, so git already knows to ignore these sort of side\n> histories that result in no changes on the branch they got merged\n> into. I don't know about bisect, but if it's not doing something similar\n> already it would be easy to make it do so.\n\nI haven't had the need to use --full-history much. Let me see if I can\nplay around with it to see if I can figure out how to use it in a way\nthat gives me what I'm after.\n\n> You could even add a new field to the commit object of A2+B2 & C2 which\n> would be one or more of \"replaces <sha1 of A/B/C>\", commit objects\n> support adding arbitrary new fields without anything breaking.\n> \n> But most importantly, while I think this gives you the same things from\n> a UX level, it doesn't need any changes to fetch, push, gc or whatever,\n> since it's all stuff we support today, someone just needs to hack\n> \"rebase\" to create this sort of no-op merge commit to take advantage of\n> it.\n\nAvoiding changes would be very nice. I'm not convinced yet that it can\nbe done but maybe when I understand your counter proposal, it will\nbecome clearer.\n\nThank you,\nCarl Baldwin\n"},{"id":"335241","messageId":"87608xrt8o.fsf@evledraar.gmail.com","threadId":"47487","inReplyTo":"20171223210141.GA24715@hpz.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2017-12-23T22:09:59Z","receivedAt":"2017-12-23T22:12:51Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, Dec 23 2017, Carl Baldwin jotted:\n\n> On Sat, Dec 23, 2017 at 07:59:35PM +0100, Ævar Arnfjörð Bjarmason wrote:\n>> I think this is a worthwhile thing to implement, there are certainly\n>> use-cases where you'd like to have your cake & eat it too as it were,\n>> i.e. have a nice rebased history in \"git log\", but also have the \"raw\"\n>> history for all the reasons the fossil people like to talk about, or for\n>> some compliance reasons.\n>\n> Thank you kindly for your reply. I do think we can have the cake and eat\n> it too in this case. At a high level, what you describe above is what\n> I'm after. I'm sorry if I left something out or was unclear. I hoped to\n> keep my original post brief. Maybe it was too brief to be useful.\n> However, I'd like to follow up and be understood.\n>\n>> But I don't see why you think this needs a new \"replaces\" parent pointer\n>> orthagonal to parent pointers, i.e. something that would need to be a\n>> new field in the commit object (I may have misread the proposal, it's\n>> not heavy on technical details).\n>\n> Just to clarify, I am proposing a new \"replaces\" pointer in the commit\n> object. Imagine starting with rebase exactly as it works today. This new\n> field would be inserted into any new commit created by a rebase command\n> to reference the original commit on which it was based. Though, I'm not\n> sure if it would be better to change the behavior of the existing rebase\n> command, provide a switch or config option to turn it on, or provide a\n> new command entirely (e.g. git replay or git replace) to avoid\n> compatibility issues with the existing rebase.\n\nYeah that sounds fine, I thought you meant that this \"replaces\" field\nwould replace the \"parent\" field, which would require some rather deep\nincompatible changes to all git clients.\n\nBut then I don't get why you think fetch/pull/gc would need to be\naltered, if it's because you thought that adding arbitrary *new* fields\nto the commit object would require changes to those that's not the case.\n\n> I imagine that a \"git commit --amend\" would also insert a \"replaces\"\n> reference to the original commit but I failed to mention that in my\n> original post. The amend use case is similar to adding a fixup commit\n> and then doing a squash in interactive mode.\n>\n>> Consider a merge use case like this:\n>>\n>>           A---B---C topic\n>>          /         \\\n>>     D---E---F---G---H master\n>\n> This is a bit different than the use cases that I've had in mind. You\n> show that the topic has already merged to master. I have imagined this\n> proposal being useful before the topic becomes a part of the master\n> branch. I'm thinking in the context of something like a github pull\n> request under active development and review or a gerrit review. So, at\n> this point, we still look like this:\n>\n>           A---B---C topic\n>          /\n>     D---E---F---G\n\nRight, I'm just mentioning this for context, i.e. \"if you only used\ngit-merge\".\n\n>> Here we worked on a topic with commits A,B & C, maybe we regret not\n>> squashing B into A, but it gives us the \"raw\" history. Instead we might\n>> rebase it like this:\n>>\n>>           A+B---C topic\n>>          /\n>>     G---H master\n>\n> Since H already merged the topic. I'm not sure what the A+B and C\n> commits are doing.\n\nThis means that master is at commit H, but your newly rebased topic is\nat C, i.e. master has no new commits so you could `git push origin\nC:master` without -f.\n\n> At the point where I have C and G above, let's say I regret not having\n> squashed A and B as you suggested. My proposal would end up as I draw\n> below where the primes are the new versions of the commits (A' is A+B).\n> Bare with me, I'm not sure the best way to draw this in ascii. It has\n> that orthogoal dimension that makes the ascii drawings a little more\n> complex: (I left out the parent of A' which is still E)\n>\n>        A--B---C\n>         \\ |    \\                    <- \"replaces\" rather than \"parent\"\n>          -A'----C' topic\n>          /\n>     D---E---F---G master\n>\n> We can continue by actually changing the base. All of these commits are\n> kept, I just drop them from the drawings to avoid getting too complex.\n>\n>                 A'--C'\n>                  \\   \\              <- \"replaces\" rather than \"parent\"\n>                   A\"--C\" topic\n>                  /\n>     D---E---F---G master\n>\n> Normal git log operations would ignore them by default. When finally\n> merging to master, it ends up very simple (by default) but the history\n> is still there to support archealogic operations.\n>\n>     D---E---F---G---A\"--C\" master\n>\n>> Now we can push \"topic\" to master, but as you've noted this loses the\n>> raw history, but now consider doing this instead:\n>>\n>>           A---B---C   A2+B2---C2 topic\n>>          /         \\ /\n>>     D---E---F---G---G master\n>\n> There are two Gs in this drawing. Should the second be H? Sorry, I'm\n> just trying to understanding the use case you're describing and I don't\n> understand it yet which makes it difficult to comment on the rest of\n> your reply.\n\nYes this is very confusing, sorry for not clarifying this.\n\nWhat the letters in *this* diagram actually mean is they're all unique\nids for commits that parse to the same value given;\n\n    git rev-parse $commit^{tree}\n\nI.e. you'd merge C into the G commit, and you'd end up with a commit\nthat would give you the exact same tree, see \"ours\" under \"MERGE\nSTRATEGIES\" in git-commit(1).\n\nYou can try to create one of these with:\n\n    (\n        rm -rf /tmp/testgit &&\n        git clone git@github.com:antirez/rax.git /tmp/testgit &&\n        cd /tmp/testgit &&\n        git checkout -b wip-rebase master &&\n        for f in foo bar baz; do\n            echo $f >$f &&\n            git add $f &&\n            git commit -m\"$f\"\n        done &&\n        git checkout master &&\n        git merge --no-edit -s ours wip-rebase &&\n        git rev-parse origin/master^{tree} &&\n        git rev-parse HEAD^{tree}\n    )\n\nNote that the output of the two rev-parse commands is the same,\ni.e. I've created a bunch of content on a side branch and merged it in,\nbut due to \"-s ours\" the end result is exactly the same as if it had\nnever been merged as far as the content of the tree at HEAD goes.\n\nBut I see now that I was wrong/misremembering about --full-history. In\nthis case if you just run \"git log\" you'd get those foo/bar/baz changes,\nhowever if you run;\n\n    git log -- foo\n\nYou get nothing, but run:\n\n    git log --full-history -- foo\n\nAnd you get that no-op merge.\n\nBut in any case, regardless of what the history simplification does\n*now* I was trying to point out, with the assumption (see my comment\nabout pull/fetch/gc above) that you were suggesting some deep changes in\nhow git's object model works.\n\nInstead, if I understand what you're actually trying to do, it could\nalso be done as:\n\n 1) Just add a new replaces <sha1> field to new commit objects\n\n 2) Make git-rebase know how to write those, e.g. add two of those\n    pointing to A & B when it squashes them into AB.\n\n 3) Write a history traversal mechanism similar to --full-history\n    that'll ignore any commits on branches that yield no changes, or\n    only those whose commits are referenced by this \"replaces\" field.\n\nYou'd then end up with:\n\n A) A way to \"stash\" these commits in the permanent history\n\n B) ... that wouldn't be visble in \"git log\" by default\n\n C) Would require no underlying changes to the commit model, i.e. it\n    would work with all past & future git clients, if they didn't know\n    about the \"replaces\" field they'd just show more verbose history.\n\n>> I.e. you could have started working on commit A/B/C, now you \"git\n>> replace\" them (which would be some fancy rebase alias), and what it'll\n>> do is create a merge commit that entirely resolves the conflict so that\n>> hte tree is equivalent to what \"master\" was already at. Then you rewrite\n>> them and re-apply them on top.\n>>\n>> If you run \"git log\" it will already ignore A,B,C unless you specify\n>> --full-history, so git already knows to ignore these sort of side\n>> histories that result in no changes on the branch they got merged\n>> into. I don't know about bisect, but if it's not doing something similar\n>> already it would be easy to make it do so.\n>\n> I haven't had the need to use --full-history much. Let me see if I can\n> play around with it to see if I can figure out how to use it in a way\n> that gives me what I'm after.\n>\n>> You could even add a new field to the commit object of A2+B2 & C2 which\n>> would be one or more of \"replaces <sha1 of A/B/C>\", commit objects\n>> support adding arbitrary new fields without anything breaking.\n>>\n>> But most importantly, while I think this gives you the same things from\n>> a UX level, it doesn't need any changes to fetch, push, gc or whatever,\n>> since it's all stuff we support today, someone just needs to hack\n>> \"rebase\" to create this sort of no-op merge commit to take advantage of\n>> it.\n>\n> Avoiding changes would be very nice. I'm not convinced yet that it can\n> be done but maybe when I understand your counter proposal, it will\n> become clearer.\n>\n> Thank you,\n> Carl Baldwin\n"},{"id":"335242","messageId":"000d01d37c3c$207d7050$617850f0$@nexbridge.com","threadId":"47487","inReplyTo":"20171223210141.GA24715@hpz.ecbaldwin.net","subject":"RE: Bring together merge and rebase","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2017-12-23T22:19:35Z","receivedAt":"2017-12-23T22:19:48Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On December 23, 2017 4:02 PM, Carl Baldwin wrote:\n> On Sat, Dec 23, 2017 at 07:59:35PM +0100, Ævar Arnfjörð Bjarmason wrote:\n> > I think this is a worthwhile thing to implement, there are certainly\n> > use-cases where you'd like to have your cake & eat it too as it were,\n> > i.e. have a nice rebased history in \"git log\", but also have the \"raw\"\n> > history for all the reasons the fossil people like to talk about, or\n> > for some compliance reasons.\n> \n> Thank you kindly for your reply. I do think we can have the cake and eat it\n> too in this case. At a high level, what you describe above is what I'm after.\n> I'm sorry if I left something out or was unclear. I hoped to keep my original\n> post brief. Maybe it was too brief to be useful.\n> However, I'd like to follow up and be understood.\n> \n> > But I don't see why you think this needs a new \"replaces\" parent\n> > pointer orthagonal to parent pointers, i.e. something that would need\n> > to be a new field in the commit object (I may have misread the\n> > proposal, it's not heavy on technical details).\n> \n> Just to clarify, I am proposing a new \"replaces\" pointer in the commit object.\n> Imagine starting with rebase exactly as it works today. This new field would\n> be inserted into any new commit created by a rebase command to reference\n> the original commit on which it was based. Though, I'm not sure if it would\n> be better to change the behavior of the existing rebase command, provide a\n> switch or config option to turn it on, or provide a new command entirely (e.g.\n> git replay or git replace) to avoid compatibility issues with the existing rebase.\n> \n> I imagine that a \"git commit --amend\" would also insert a \"replaces\"\n> reference to the original commit but I failed to mention that in my original\n> post. The amend use case is similar to adding a fixup commit and then doing\n> a squash in interactive mode.\n> \n> > Consider a merge use case like this:\n> >\n> >           A---B---C topic\n> >          /         \\\n> >     D---E---F---G---H master\n> \n> This is a bit different than the use cases that I've had in mind. You show that\n> the topic has already merged to master. I have imagined this proposal being\n> useful before the topic becomes a part of the master branch. I'm thinking in\n> the context of something like a github pull request under active development\n> and review or a gerrit review. So, at this point, we still look like this:\n> \n>           A---B---C topic\n>          /\n>     D---E---F---G\n> \n> > Here we worked on a topic with commits A,B & C, maybe we regret not\n> > squashing B into A, but it gives us the \"raw\" history. Instead we\n> > might rebase it like this:\n> >\n> >           A+B---C topic\n> >          /\n> >     G---H master\n> \n> Since H already merged the topic. I'm not sure what the A+B and C commits\n> are doing.\n> \n> At the point where I have C and G above, let's say I regret not having\n> squashed A and B as you suggested. My proposal would end up as I draw\n> below where the primes are the new versions of the commits (A' is A+B).\n> Bare with me, I'm not sure the best way to draw this in ascii. It has that\n> orthogoal dimension that makes the ascii drawings a little more\n> complex: (I left out the parent of A' which is still E)\n> \n>        A--B---C\n>         \\ |    \\                    <- \"replaces\" rather than \"parent\"\n>          -A'----C' topic\n>          /\n>     D---E---F---G master\n> \n> We can continue by actually changing the base. All of these commits are\n> kept, I just drop them from the drawings to avoid getting too complex.\n> \n>                 A'--C'\n>                  \\   \\              <- \"replaces\" rather than \"parent\"\n>                   A\"--C\" topic\n>                  /\n>     D---E---F---G master\n> \n> Normal git log operations would ignore them by default. When finally\n> merging to master, it ends up very simple (by default) but the history is still\n> there to support archealogic operations.\n> \n>     D---E---F---G---A\"--C\" master\n> \n> > Now we can push \"topic\" to master, but as you've noted this loses the\n> > raw history, but now consider doing this instead:\n> >\n> >           A---B---C   A2+B2---C2 topic\n> >          /         \\ /\n> >     D---E---F---G---G master\n> \n> There are two Gs in this drawing. Should the second be H? Sorry, I'm just\n> trying to understanding the use case you're describing and I don't\n> understand it yet which makes it difficult to comment on the rest of your\n> reply.\n> \n> > I.e. you could have started working on commit A/B/C, now you \"git\n> > replace\" them (which would be some fancy rebase alias), and what it'll\n> > do is create a merge commit that entirely resolves the conflict so\n> > that hte tree is equivalent to what \"master\" was already at. Then you\n> > rewrite them and re-apply them on top.\n> >\n> > If you run \"git log\" it will already ignore A,B,C unless you specify\n> > --full-history, so git already knows to ignore these sort of side\n> > histories that result in no changes on the branch they got merged\n> > into. I don't know about bisect, but if it's not doing something\n> > similar already it would be easy to make it do so.\n> \n> I haven't had the need to use --full-history much. Let me see if I can play\n> around with it to see if I can figure out how to use it in a way that gives me\n> what I'm after.\n> \n> > You could even add a new field to the commit object of A2+B2 & C2\n> > which would be one or more of \"replaces <sha1 of A/B/C>\", commit\n> > objects support adding arbitrary new fields without anything breaking.\n> >\n> > But most importantly, while I think this gives you the same things\n> > from a UX level, it doesn't need any changes to fetch, push, gc or\n> > whatever, since it's all stuff we support today, someone just needs to\n> > hack \"rebase\" to create this sort of no-op merge commit to take\n> > advantage of it.\n> \n> Avoiding changes would be very nice. I'm not convinced yet that it can be\n> done but maybe when I understand your counter proposal, it will become\n> clearer.\n\nNo matter how this plays out, let's please make very sure to provide sufficient user documentation so that those of us who have to explain the differences to users have a decent reference. Even now, explaining rebase vs. merge is difficult enough for people new to git to choose which to use when (sometimes pummeling is involved to get the point across 😉 ), even though it should be intuitive to most of us. I am predicting that adding this capability is going to further confuse the *new* user community a little. Entirely out of enlighted self-interest, I am offering to help document (edits/contribution//whatever) this once we get to that point in development.\n\nSomething else to consider is how (or if) this capability is going to be presented in front-ends and in Cloud services. GitK is a given, of course. I'm still impatiently waiting for worktree support from some other front-ends.\n\nCheers,\nRandall\n\n-- Brief whoami: NonStop&UNIX developer since approximately UNIX(421664400)/NonStop(211288444200000000)\n-- In my real life, I talk too much.\n\n\n\n"},{"id":"335243","messageId":"alpine.DEB.2.21.1.1712232248410.406@MININT-6BKU6QN.europe.corp.microsoft.com","threadId":"47487","inReplyTo":"877etds220.fsf@evledraar.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-12-23T22:30:33Z","receivedAt":"2017-12-23T22:30:44Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Ævar,\n\nOn Sat, 23 Dec 2017, Ævar Arnfjörð Bjarmason wrote:\n\n> On Sat, Dec 23 2017, Carl Baldwin jotted:\n> \n> > The big contention among git users is whether to rebase or to merge\n> > changes [2][3] while iterating. I used to firmly believe that merging\n> > was the way to go and rebase was harmful. More recently, I have worked\n> > in some environments where I saw rebase used very effectively while\n> > iterating on changes and I relaxed my stance a lot. Now, I'm on the\n> > fence. I appreciate the strengths and weaknesses of both approaches. I\n> > waffle between the two depending on the situation, the tools being\n> > used, and I guess, to some extent, my mood.\n> >\n> > I think what git needs is something brand new that brings the two\n> > together and has all of the advantages of both approaches. Let me\n> > explain what I've got in mind...\n> >\n> > I've been calling this proposal `git replay` or `git replace` but I'd\n> > like to hear other suggestions for what to name it. It works like\n> > rebase except with one very important difference. Instead of orphaning\n> > the original commit, it keeps a pointer to it in the commit just like\n> > a `parent` entry but calls it `replaces` instead to distinguish it\n> > from regular history. In the resulting commit history, following\n> > `parent` pointers shows exactly the same history as if the commit had\n> > been rebased. Meanwhile, the history of iterating on the change itself\n> > is available by following `replaces` pointers. The new commit replaces\n> > the old one but keeps it around to record how the change evolved.\n> >\n> > The git history now has two dimensions. The first shows a cleaned up\n> > history where fix ups and code review feedback have been rolled into\n> > the original changes and changes can possibly be ordered in a nice\n> > linear progression that is much easier to understand. The second\n> > drills into the history of a change. There is no loss and you don't\n> > change history in a way that will cause problems for others who have\n> > the older commits.\n> >\n> > Replay handles collaboration between multiple authors on a single\n> > change. This is difficult and prone to accidental loss when using\n> > rebase and it results in a complex history when done with merge. With\n> > replay, collaborators could merge while collaborating on a single\n> > change and a record of each one's contributions can be preserved.\n> > Attempting this level of collaboration caused me many headaches when I\n> > worked with the gerrit workflow (which in many ways, I like a lot).\n> >\n> > I blogged about this proposal earlier this year when I first thought\n> > of it [1]. I got busy and didn't think about it for a while. Now with\n> > a little time off of work, I've come back to revisit it. The blog\n> > entry has a few examples showing how it works and how the history will\n> > look in a few examples. Take a look.\n> >\n> > Various git commands will have to learn how to handle this kind of\n> > history. For example, things like fetch, push, gc, and others that\n> > move history around and clean out orphaned history should treat\n> > anything reachable through `replaces` pointers as precious. Log and\n> > related history commands may need new switches to traverse the history\n> > differently in different situations. Bisect is a interesting one. I\n> > tend to think that bisect should prefer the regular commit history but\n> > have the ability to drill into the change history if necessary.\n> >\n> > In my opinion, this proposal would bring together rebase and merge in\n> > a powerful way and could end the contention. Thanks for your\n> > consideration.\n> >\n> > Carl Baldwin\n> >\n> > [1] http://blog.episodicgenius.com/post/merge-or-rebase--neither/ [2]\n> > https://git-scm.com/book/en/v2/Git-Branching-Rebasing [3]\n> > http://changelog.complete.org/archives/586-rebase-considered-harmful\n> \n> I think this is a worthwhile thing to implement, there are certainly\n> use-cases where you'd like to have your cake & eat it too as it were,\n> i.e. have a nice rebased history in \"git log\", but also have the \"raw\"\n> history for all the reasons the fossil people like to talk about, or for\n> some compliance reasons.\n> \n> But I don't see why you think this needs a new \"replaces\" parent pointer\n> orthagonal to parent pointers, i.e. something that would need to be a\n> new field in the commit object (I may have misread the proposal, it's\n> not heavy on technical details).\n> \n> Consider a merge use case like this:\n> \n>           A---B---C topic\n>          /         \\\n>     D---E---F---G---H master\n> \n> Here we worked on a topic with commits A,B & C, maybe we regret not\n> squashing B into A, but it gives us the \"raw\" history. Instead we might\n> rebase it like this:\n> \n>           A+B---C topic\n>          /\n>     G---H master\n> \n> Now we can push \"topic\" to master, but as you've noted this loses the\n> raw history, but now consider doing this instead:\n> \n>           A---B---C   A2+B2---C2 topic\n>          /         \\ /\n>     D---E---F---G---G master\n> \n> I.e. you could have started working on commit A/B/C, now you \"git\n> replace\" them (which would be some fancy rebase alias), and what it'll\n> do is create a merge commit that entirely resolves the conflict so that\n> hte tree is equivalent to what \"master\" was already at. Then you rewrite\n> them and re-apply them on top.\n\n1) you just described the \"merging rebase\" I use in Git for Windows for\n*quite* a while (five years or so):\n\nhttps://github.com/git-for-windows/build-extra/blob/af9cff5005/shears.sh#L12-L18\n\nJust look for commits in https://github.com/git-for-windows/git/commits\nwhose oneline begins with \"Start the merging-rebase\".\n\n2) you do not resolve merge conflicts here, as there may not be any.\nInstead, you use the \"ours\" merge strategy.\n\n> If you run \"git log\" it will already ignore A,B,C unless you specify\n> --full-history, so git already knows to ignore these sort of side\n> histories that result in no changes on the branch they got merged\n> into. I don't know about bisect, but if it's not doing something similar\n> already it would be easy to make it do so.\n\nSadly, it is not as easy as that. When you call \"git log\", you often want\nto know *when* a change was introduced originally. In this case, you would\n*not* want A, B nor C ignored, but you would really want to dig into that\nhistory that is ignored by default.\n\nIn general, the technique you described (and that I described years before\nyou, and employ for years, too, so I actually already have experience with\nits pros and cons) works, but leaves quite a bit to be desired.\n\nFor example, when anybody asks me \"when was XYZ fixed in Git for Windows?\"\nit is not enough to run `git blame` and then `git name-rev` on the commit\nidentified by `git blame`: this would be your A2, and there could be any\nnumber of previous iterations *of the same patch*. What I do in this case\nis to search for the matching oneline in the full history, from the end.\nThis is a costly, and not very automatable operation (as there have been\nrewordings at times, in particular when the oneline contained a typo).\n\n> You could even add a new field to the commit object of A2+B2 & C2 which\n> would be one or more of \"replaces <sha1 of A/B/C>\", commit objects\n> support adding arbitrary new fields without anything breaking.\n\nThis is a very fragile way of doing things because you cannot fix rebases\ndone in the past. Those commits won't have that header, and you cannot put\nit there after the fact, not even manually identifying the mapping.\n\nBesides, there are plenty of scenarios when you do not actually want a\nmerging-rebase, e.g. when you develop a patch series and have to iterate\nit over a dozen times until it is finally accepted into core Git. In this\ninstance, you may want to retain the iterations' commit histories during\nthe time of the development, but you probably won't need it any longer\nafter the patches have been integrated into a released version.\n\nBaking those names into the commit object would kind of cause broken links\nin such a scenario.\n\nBTW it gets a lot more complicated when you think about\n\n1) fixups and squashes, and\n\n2) the often much more interesting question: *with what commit* was this\none replaced?\n\n> But most importantly, while I think this gives you the same things from\n> a UX level, it doesn't need any changes to fetch, push, gc or whatever,\n> since it's all stuff we support today, someone just needs to hack\n> \"rebase\" to create this sort of no-op merge commit to take advantage of\n> it.\n\nI already did that. You can use the above-linked shears.sh script to\nperform such a merging-rebase.\n\nHowever, my experience is that the lack of UX in Git's tools do hurt at\ntimes, it is incorrect to say that log or blame already support this use\ncase (see the discussion above).\n\nThe most important part would be to record the mapping between old and new\ncommits; the `post-rewrite` hook would pose a natural way to implement\nthis: it gets called after a successful rebase, receiving a stream of\nlines via stdin of the form `<old-commit> <new-commit>` (listing multiple\nlines for the same <new-commit> if there were fixups and/or squashes).\n\nThis points to another command that absolutely would need patching: `git\npull` (or more correctly: `git merge`). Why, you ask? If you use that\ntechnique (and we do, as I pointed out earlier, in Git for Windows),\ncontributors *will* start to send you contributions *based on previous\niterations*. If you integrate them (\"pull\"), the merge may very well\nsucceed. But the next time you rebase, those commits will not be ordered\ncorrectly, as they are considered older than the patches on which they are\nbased (which have been rebased in the meantime already).\n\nIn Git for Windows, I do have these problems, and they are even worse\nthere because instead of a linearizing rebase, I recreate the branch\nstructure instead (see https://github.com/git/git/pull/447 for my upcoming\nattempt to implement this functionality directly in core Git, in proper,\nperformant and portable C). So the base commits are *really* important\ninformation. And I had to work harder in the past to find out what the\nnewer iteration of the base commit really is.\n\nBTW I would *strongly* suggest to use notes instead of a new commit\nheader. You absolutely do not want to impose this on any user who does not\nwant it. And notes can be added for already-completed rebases. Even\nmanually, if need be.\n\nCiao,\nJohannes"},{"id":"335245","messageId":"alpine.DEB.2.21.1.1712232353390.406@MININT-6BKU6QN.europe.corp.microsoft.com","threadId":"47487","inReplyTo":"20171223210141.GA24715@hpz.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-12-23T23:01:38Z","receivedAt":"2017-12-23T23:04:12Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Carl,\n\nOn Sat, 23 Dec 2017, Carl Baldwin wrote:\n\n> I imagine that a \"git commit --amend\" would also insert a \"replaces\"\n> reference to the original commit but I failed to mention that in my\n> original post.\n\nAnd cherry-pick, too, of course.\n\nBoth of these examples hint at a rather huge urge of some users to turn\nthis feature off because the referenced commits may very well be\nthrow-away commits in their case, making the newly-recorded information\ncompletely undesired.\n\nExample: I am working on a topic branch. In the middle, I see a typo. I\ncommit a fix, continue to work on the topic branch. Later, I cherry-pick\nthat commit to a separate topic branch because I really don't think that\nthose two topics are related. Now I definitely do not want a reference of\nthe cherry-picked commit to the original one: the latter will never be\npushed to a public repository, and gc'ed in a few weeks.\n\nOf course, that is only my wish, other users in similar situations may\nwant that information. Demonstrating that you would be better served with\nan opt-in feature that uses notes rather than a baked-in commit header.\n\nCiao,\nJohannes\n"},{"id":"335256","messageId":"16725929-1BD2-44D3-8E71-E97C4A2C4034@gmail.com","threadId":"47487","inReplyTo":"alpine.DEB.2.21.1.1712232353390.406@MININT-6BKU6QN.europe.corp.microsoft.com","subject":"Re: Bring together merge and rebase","fromName":"Alexei Lozovsky","fromEmail":"a.lozovsky@gmail.com","sentAt":"2017-12-24T14:13:23Z","receivedAt":"2017-12-24T14:16:10Z","isPatch":false,"sender":{"key":"a.lozovsky@gmail.com","avatar":"https://gravatar.com/avatar/8bb8ff5ec366dd64bd8e08f768082934513da367ae047bcb0039292e4ed6bda5?d=mp&s=160"},"body":"On Dec 24, 2017, at 01:01, Johannes Schindelin wrote:\n> \n> Hi Carl,\n> \n> On Sat, 23 Dec 2017, Carl Baldwin wrote:\n> \n>> I imagine that a \"git commit --amend\" would also insert a \"replaces\"\n>> reference to the original commit but I failed to mention that in my\n>> original post.\n> \n> And cherry-pick, too, of course.\n\nWhy would it? In my mind, cherry-picking does not 'replace' or 'refine'\ncommits, it copies them into other, unrelated branches (usually something\nlike stable branches maintained separately from the mainline). If anything,\ncherry-pick could add a separate \"cherry-picked from\" reference which may\nbe useful, I guess, for conflict resolution if two branches with the same\ncommit are merged.\n\n> Of course, that is only my wish, other users in similar situations may\n> want that information. Demonstrating that you would be better served with\n> an opt-in feature that uses notes rather than a baked-in commit header.\n\nUsing notes also allows to test and evaluate this new feature without\nany changes to core git, using it as an extension at first.\n"},{"id":"335276","messageId":"20171225035215.GC1257@thunk.org","threadId":"47487","inReplyTo":"CALiLy7pBvyqA+NjTZHOK9t0AFGYbwqwRVD3sZjUg0ZLx5y1h3A@mail.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2017-12-25T03:52:15Z","receivedAt":"2017-12-25T03:52:26Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Fri, Dec 22, 2017 at 11:10:19PM -0700, Carl Baldwin wrote:\n> I've been calling this proposal `git replay` or `git replace` but I'd\n> like to hear other suggestions for what to name it. It works like\n> rebase except with one very important difference. Instead of orphaning\n> the original commit, it keeps a pointer to it in the commit just like\n> a `parent` entry but calls it `replaces` instead to distinguish it\n> from regular history. In the resulting commit history, following\n> `parent` pointers shows exactly the same history as if the commit had\n> been rebased. Meanwhile, the history of iterating on the change itself\n> is available by following `replaces` pointers. The new commit replaces\n> the old one but keeps it around to record how the change evolved.\n\nAs a suggestion, before diving into the technical details of your\nproposal, it might be useful consider the usage scenario you are\ntargetting.  Things like \"git rebase\" and \"git merge\" and your\nproposed \"git replace/replay\" are *mechanisms*.\n\nBut how they fit into a particular workflow is much more important\nfrom a design perspective, and given that there are many different git\nworkflows which are used by different projects, and by different\ndevelopers within a particular project.\n\nFor example, rebase gets used in many different ways, and many of the\ndebates when people talk about \"git rebase\" being evil generally\npresuppose a particular workflow that that the advocate has in mind.\nIf someone is using git rebase or git commit --amend before git\ncommits have ever been pushed out to a public repository, or to anyone\nelse, that's a very different case where it has been visible\nelsewhere.  Even the the most strident, \"you must never rewrite a\ncommit and all history must be preserved\" generally don't insist that\nevery single edit must be preserved on the theory that \"all history is\nvaluable\".\n\n> The git history now has two dimensions. The first shows a cleaned up\n> history where fix ups and code review feedback have been rolled into\n> the original changes and changes can possibly be ordered in a nice\n> linear progression that is much easier to understand. The second\n> drills into the history of a change. There is no loss and you don't\n> change history in a way that will cause problems for others who have\n> the older commits.\n\nIf your goal is to preserve the history of the change, one of the\nproblems with any git-centric solution is that you generally lose the\ncode review feedback and the discussions that are involved with a\ncommit.  Just simply preserving the different versions of the commits\nis going to lose a huge amount of the context that makes the history\nvaluable.\n\nSo for example, I would claim that if *that* is your goal, a better\nsolution is to use Gerrit, so that all of the different versions of\nthe commits are preserved along with the line-by-line comments and\ndiscussions that were part of the code review.  In that model, each\ncommit has something like this in the commit trailer:\n\nChange-Id: I8d89b33683274451bcd6bfbaf75bce98\n\nYou can then cut and paste the Change-Id into the Gerrit user\ninterface, and see the different commits, more important, the\ndiscussion surrounding each change.\n\n\nIf the complaint about Gerrit is that it's not a core part of Git, the\nchallenge is (a) how to carry the code review comments in the git\nrepository, and (b) do so in a while that it doesn't bloat the core\nrepository, since most of the time, you *don't* want or need to keep a\nlocal copy of all of the code review comments going back since the\nbeginning of the project.\n\n-------------\n\nHere's another potential use case.  The stable kernels (e.g., 3.18.y,\n4.4.y, 4.9.y, etc.) have cherry picks from the the upstream kernel,\nand this is handled by putting in the commit body something like this:\n\n    [ Upstream commit 3a4b77cd47bb837b8557595ec7425f281f2ca1fe ]\n\n----\n\nAnd here's yet another use case.  For internal Google kernel\ndevelopment, we maintain a kernel that has a large number of patches\non top of a kernel version.  When we backport an upstream fix (say,\none that first appeared in the 4.12 version of the upstream kernel),\nwe include a line in the commit body that looks like this:\n\nUpstream-4.12-SHA1: 5649645d725c73df4302428ee4e02c869248b4c5\n\nThis is useful, because when we switch to use a newer upstream kernel,\nwe need make sure we can account for all patches that were built on\ntop of the 3xx kernel (which might have been using 4.10, for the sake\nof argument), to the 4xx kernel series (which might be using 4.15 ---\nthe version numbers have been changed to protect the innocent).  This\nmeans going through each and every patch that was on top of the 3xx\nkernel, and if it has a line such as \"Upstream 4.12-SHA1\", we know\nthat it will already be included in a 4.15 based kernel, so we don't\nneed to worry about carrying that patch forward.\n\nIn other cases, we might decide that the patch is no longer needed.\nIt could be because the patch has already be included upstream, in\nwhich case we might check in a commit with an empty patch body, but\nwhose header contains something like this in the 4xx kernel:\n\nOrigin-3xx-SHA1: fe546bdfc46a92255ebbaa908dc3a942bc422faa\nUpstream-Dropped-4.11-SHA1: d90dc0ae7c264735bfc5ac354c44ce2e\n\nOr we could decide that the commit is no longer no longer needed ---\nperhaps because the relevant subsystem was completely rewritten and\nthe functionality was added in a different way.  Then we might have\njust have an empty commit with an explanation of why the commit is no\nlonger needed and the commit body would have the metadata:\n\nOrigin-Dropped-3xx-SHA1: 26f49fcbb45e4bc18ad5b52dc93c3afe\n\nOr perhaps the commit is still needed, and for various reasons the\ncommit was never upstreamed; perhaps because it's only useful for\nGoogle-specific hardware, or the patch was rejected upstream.  The we\nwill have a cherry-pick that would include in the body:\n\nOrigin-3xx-SHA1: 8f3b6df74b9b4ec3ab615effb984c1b5\n\n\n(Note: all commits that are added in the rebase workflow, even the\nempty commits that just have the Origin-Dropped-3xx-SHA1 or\nUpstream-Droped-4.11-SHA1 headers, are patch reviewed through Gerrit,\nso we have an audited, second-engineer review to make sure each commit\nin the 3xx kernel that Google had been carrying had the correct\ndisposition when rebasing to the 4xx kernel.)\n\nThe point is that for this much more complex, real-world workflow, we\nneed much *more* metadata than a simple \"Replaces\" metadata.  (And we\nalso have other metadata --- for example, we have a \"Tested: \" trailer\nthat explains how to test the commit, or which unit test can and\nshould be used to test this commit, combined with a link to the test\nlog in our automated unit tester that has the test run, and a\n\"Rebase-Tested-4xx: \" trailer that might just have the URL to the test\nlog when the commit was rebased since the testing instructions in the\nTested: trailer is still relevant.)\n\nAnd since this metadata is not really needed by the core git\nmachinery, we just use text trailers in the commit body; it's not hard\nto write code which parses this out of the git commit.\n\n> Various git commands will have to learn how to handle this kind of\n> history. For example, things like fetch, push, gc, and others that\n> move history around and clean out orphaned history should treat\n> anything reachable through `replaces` pointers as precious. Log and\n> related history commands may need new switches to traverse the history\n> differently in different situations.\n\nI'd encourage you to think very hard about how exactly \"git log\" and\n\"gitk\" might actually deal with these links.  In the Google kernel\ndevelopment use cases, we use different repos for the 3xx and 4xx\nkernels.  It would be possible to make hot links for the\nOriginal-3xx-SHA1: trailers, but you couldn't do it using gitk.  It\nwould actually have to be a completely new tool.  (And we do have new\ntools, most especially a dashboard so we can keep track of how many\ncommits in the 3xx kernel still have to be rebased to the 4xx kernel,\nor can be confirmed to be in the upstream kernel, or can be confirmed\nto be dropped.  We have a *large* number of patches that we carry, so\nit's a multi-month effort involving a large number of engineers\nworking together to do a kernel rebase operation from a 4.x upstream\nkernel to a 4.y upstream kernel.  So having a dashboard is useful\nbecause we can see whether a particular subsystem team is ahead or\nbehind the curve in terms of handling those commits which are their\nresponsibility.)\n\n\nMy experience, from seeing these much more complex use cases ---\nstarting with something as simple as the Linux Kernel Stable Kernel\nSeries, and extending to something much more complex such as the\nworkflow that is used to support a Google Kernel Rebase, is that using\njust a simple extra \"Replaces\" pointer in the commit header is not\nnearly expressive enough.  And, if you make it a core part of the\ncommit data structure, there are all sorts of compatibility headaches\nwith older versions of git that wouldn't know about it.  And if it\nthen turns out it's not sufficient more the more complex workflows\n*anyway*, maybe adding a new \"replace\" pointer in the core git data\nstructures isn't worth it.  It might be that just keeping such things\nas trailers in the commit body might be the better way to go.\n\nCheers,\n\n\t\t\t\t\t\t- Ted\n"},{"id":"335301","messageId":"20171225200509.GA24104@hpz.ecbaldwin.net","threadId":"47487","inReplyTo":"000d01d37c3c$207d7050$617850f0$@nexbridge.com","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-25T20:05:11Z","receivedAt":"2017-12-25T20:09:39Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Sat, Dec 23, 2017 at 05:19:35PM -0500, Randall S. Becker wrote:\n> No matter how this plays out, let's please make very sure to provide\n> sufficient user documentation so that those of us who have to explain\n> the differences to users have a decent reference. Even now, explaining\n> rebase vs. merge is difficult enough for people new to git to choose\n> which to use when (sometimes pummeling is involved to get the point\n> across 😉 ), even though it should be intuitive to most of us. I am\n> predicting that adding this capability is going to further confuse the\n> *new* user community a little. Entirely out of enlighted\n> self-interest, I am offering to help document\n> (edits/contribution//whatever) this once we get to that point in\n> development.\n\nI agree. I have a feeling that it may take a while for this to play out.\nThis has been on my mind for a while and think there will be some more\ndiscussion before anything gets started.\n\nCarl\n\n> Something else to consider is how (or if) this capability is going to\n> be presented in front-ends and in Cloud services. GitK is a given, of\n> course. I'm still impatiently waiting for worktree support from some\n> other front-ends.\n\nIt all takes time. :)\n\n> Cheers,\n> Randall\n> \n> -- Brief whoami: NonStop&UNIX developer since approximately UNIX(421664400)/NonStop(211288444200000000)\n> -- In my real life, I talk too much.\n"},{"id":"335303","messageId":"20171225234334.GB24104@hpz.ecbaldwin.net","threadId":"47487","inReplyTo":"alpine.DEB.2.21.1.1712232353390.406@MININT-6BKU6QN.europe.corp.microsoft.com","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-25T23:43:35Z","receivedAt":"2017-12-26T00:06:35Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Sun, Dec 24, 2017 at 12:01:38AM +0100, Johannes Schindelin wrote:\n> Hi Carl,\n> \n> On Sat, 23 Dec 2017, Carl Baldwin wrote:\n> \n> > I imagine that a \"git commit --amend\" would also insert a \"replaces\"\n> > reference to the original commit but I failed to mention that in my\n> > original post.\n> \n> And cherry-pick, too, of course.\n\nThis brings up a good point. I do think this can be applied to\ncherry-pick, but as someone else pointed out, the name \"replaces\"\ndoesn't seem right in the context of a cherry-pick. So, maybe \"replaces\"\nis not the right name. I'm open to suggestions.\n\nIt occurs to me now that the reason that I want a separate, orthogonal\nhistory dimension is that a \"replaces\" reference does not imply that the\nreferenced commit is pulled in with all of its history like a \"parent\"\nreference does. It isn't creating a merge commit. It means that the\nreferenced commit is derived from the other one and, at least in the\ncontext of this branch's main history, renders it obsolete. Given this\ndefinition, I think it applies to a cherry-pick.\n\n> Both of these examples hint at a rather huge urge of some users to turn\n> this feature off because the referenced commits may very well be\n> throw-away commits in their case, making the newly-recorded information\n> completely undesired.\n\nI certainly don't want to make it difficult to get rid of throw-away\ncommits.\n\nThe workflows I'm interested in are mostly around iterating on what will\nend up looking like a single commit in the final history. I'm imagining\nposting a change, (or changes) somewhere to be reviewed by others.\nOthers submit feedback and I continue iterating given the feedback. If\ncertain intermediate throw-away commits have only been seen locally by\nthe author, they could be squashed into a single minimal new update.\n\nI'm diving deeper into these workflows in my reply to Theodore. To avoid\nfragmenting my ideas too much, I'll take the details over to that reply.\nI hope to finished that soon.\n\nCarl\n"},{"id":"335305","messageId":"004e01d37ddc$a6683280$f3389780$@nexbridge.com","threadId":"47487","inReplyTo":"20171225234334.GB24104@hpz.ecbaldwin.net","subject":"RE: Bring together merge and rebase","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2017-12-26T00:01:10Z","receivedAt":"2017-12-26T00:06:39Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On December 25, 2017 6:44 PM Carl Baldwin wrote:\n> On Sun, Dec 24, 2017 at 12:01:38AM +0100, Johannes Schindelin wrote:\n> > On Sat, 23 Dec 2017, Carl Baldwin wrote:\n> > > I imagine that a \"git commit --amend\" would also insert a \"replaces\"\n> > > reference to the original commit but I failed to mention that in my\n> > > original post.\n> >\n> > And cherry-pick, too, of course.\n> \n> This brings up a good point. I do think this can be applied to cherry-pick, but\n> as someone else pointed out, the name \"replaces\"\n> doesn't seem right in the context of a cherry-pick. So, maybe \"replaces\"\n> is not the right name. I'm open to suggestions.\n\nJust an off the wall suggestion: what about \"stitch\" or \"suture\" since this is now way beyond a band-aid solution (sorry 😉 , but only a little). I was thinking along the lines of \"blend\" but that seems less graphic and doesn't apply to cherry-picking.\n\nHoliday Cheers,\nRandall\n\n-- Brief whoami: NonStop&UNIX developer since approximately UNIX(421664400)/NonStop(211288444200000000)\n-- In my real life, I talk too much.\n\n\n\n"},{"id":"335306","messageId":"20171226001622.GA16219@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"87608xrt8o.fsf@evledraar.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-26T00:16:22Z","receivedAt":"2017-12-26T00:16:30Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Sat, Dec 23, 2017 at 11:09:59PM +0100, Ævar Arnfjörð Bjarmason wrote:\n> >> But I don't see why you think this needs a new \"replaces\" parent\n> >> pointer orthagonal to parent pointers, i.e. something that would\n> >> need to be a new field in the commit object (I may have misread the\n> >> proposal, it's not heavy on technical details).\n> >\n> > Just to clarify, I am proposing a new \"replaces\" pointer in the commit\n> > object. Imagine starting with rebase exactly as it works today. This new\n> > field would be inserted into any new commit created by a rebase command\n> > to reference the original commit on which it was based. Though, I'm not\n> > sure if it would be better to change the behavior of the existing rebase\n> > command, provide a switch or config option to turn it on, or provide a\n> > new command entirely (e.g. git replay or git replace) to avoid\n> > compatibility issues with the existing rebase.\n> \n> Yeah that sounds fine, I thought you meant that this \"replaces\" field\n> would replace the \"parent\" field, which would require some rather deep\n> incompatible changes to all git clients.\n> \n> But then I don't get why you think fetch/pull/gc would need to be\n> altered, if it's because you thought that adding arbitrary *new* fields\n> to the commit object would require changes to those that's not the case.\n\nThank you again for your reply. Following is the kind of commit that I\nwould like to create.\n\n    tree fcce2f309177c7da9c795448a3e392a137434cf1\n    parent b3758d9223b63ebbfbc16c9b23205e42272cd4b9\n    replaces e8aa79baf6aef573da930a385e4db915187d5187\n    author Carl Baldwin <carl@ecbaldwin.net> 1514057225 -0700\n    committer Carl Baldwin <carl@ecbaldwin.net> 1514058444 -0700\n\nWhat will happen if I create this today? I assumed git would just choke\non it but I'm not certain. It has been a long time since I attempted to\nget into the internals of git.\n\nEven if core git code does not simply choke on it, I would like push and\npull to follow these pointers and transfer the history behind them. I\nassumed that git would not do this today. I would also like gc to\npreserve e8aa79baf6 as if it were referenced by a parent pointer so that\nit doesn't purge it from the history.\n\nI'm currently thinking of an example of the workflow that I'm after in\nresponse to Theodore Ts'o's message from yesterday. Stay tuned, I hope\nit makes it clearer why I want it this way.\n\n[snip]\n\n> Instead, if I understand what you're actually trying to do, it could\n> also be done as:\n> \n>  1) Just add a new replaces <sha1> field to new commit objects\n> \n>  2) Make git-rebase know how to write those, e.g. add two of those\n>     pointing to A & B when it squashes them into AB.\n> \n>  3) Write a history traversal mechanism similar to --full-history\n>     that'll ignore any commits on branches that yield no changes, or\n>     only those whose commits are referenced by this \"replaces\" field.\n> \n> You'd then end up with:\n> \n>  A) A way to \"stash\" these commits in the permanent history\n> \n>  B) ... that wouldn't be visble in \"git log\" by default\n> \n>  C) Would require no underlying changes to the commit model, i.e. it\n>     would work with all past & future git clients, if they didn't know\n>     about the \"replaces\" field they'd just show more verbose history.\n\nI get this point. I don't underestimate how difficult making such a\nchange to the core model is. I know there are older clients which cannot\nsimply be updated. There are also alternate implementations (e.g. jgit)\nthat also need to be considered. This is the thing I worry about the\nmost. I think at the very least, this new feature will have to be an\nopt-in feature for teams who can easily ensure a minimum version of git\nwill be used. Maybe the core.repositoryformatversion config or something\nlike that would have to play into it. There may also be some minimal\namount that could be backported to older clients to at least avoid\nchoking on new repos (I know this doesn't guarantee older clients will\nbe updated). Just throwing a few ideas out.\n\nI want to be sure that the implications have been explored before giving\nup and doing something external to git.\n\nCarl\n"},{"id":"335307","messageId":"20171226011638.GA16552@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"20171225035215.GC1257@thunk.org","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-26T01:16:40Z","receivedAt":"2017-12-26T01:24:39Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Sun, Dec 24, 2017 at 10:52:15PM -0500, Theodore Ts'o wrote:\n> As a suggestion, before diving into the technical details of your\n> proposal, it might be useful consider the usage scenario you are\n> targetting.  Things like \"git rebase\" and \"git merge\" and your\n> proposed \"git replace/replay\" are *mechanisms*.\n> \n> But how they fit into a particular workflow is much more important\n> from a design perspective, and given that there are many different git\n> workflows which are used by different projects, and by different\n> developers within a particular project.\n> \n> For example, rebase gets used in many different ways, and many of the\n> debates when people talk about \"git rebase\" being evil generally\n> presuppose a particular workflow that that the advocate has in mind.\n> If someone is using git rebase or git commit --amend before git\n> commits have ever been pushed out to a public repository, or to anyone\n> else, that's a very different case where it has been visible\n> elsewhere.  Even the the most strident, \"you must never rewrite a\n> commit and all history must be preserved\" generally don't insist that\n> every single edit must be preserved on the theory that \"all history is\n> valuable\".\n> \n> > The git history now has two dimensions. The first shows a cleaned up\n> > history where fix ups and code review feedback have been rolled into\n> > the original changes and changes can possibly be ordered in a nice\n> > linear progression that is much easier to understand. The second\n> > drills into the history of a change. There is no loss and you don't\n> > change history in a way that will cause problems for others who have\n> > the older commits.\n> \n> If your goal is to preserve the history of the change, one of the\n> problems with any git-centric solution is that you generally lose the\n> code review feedback and the discussions that are involved with a\n> commit.  Just simply preserving the different versions of the commits\n> is going to lose a huge amount of the context that makes the history\n> valuable.\n> \n> So for example, I would claim that if *that* is your goal, a better\n> solution is to use Gerrit, so that all of the different versions of\n> the commits are preserved along with the line-by-line comments and\n> discussions that were part of the code review.  In that model, each\n> commit has something like this in the commit trailer:\n> \n> Change-Id: I8d89b33683274451bcd6bfbaf75bce98\n\nThank you for your reply. I agree that discussing the workflows is very\nvaluable and I certainly haven't done that justice yet.\n\nGerrit is the tool that got me thinking about my proposal in the first\nplace. I spent a few years developing and doing a significant number of\ncode reviews using it. I've since changed to an environment where I no\nlonger have it. It turns out that \"a better solution is to use Gerrit\"\nis not helpful to me now because it isn't up to me. Gerrit is not nearly\nas ubiquitous as git itself.\n\nIn my opinion, Gerrit has shown us the power of the \"change\". As you\npoint out, it introduced the change-id embedded into the commit message\nand uses it to track a change's progress as a \"review.\" I think these\nare powerful concepts and Gerrit did a nice job with them. I guess one\nof my goals with my proposal here is to formalize the \"change\" idea so\nthat any git-based tool understands it and can interoperate. This is why\nI want it in the core git commit object and I want push, pull, gc, and\nother commands to understand it.\n\nAt this point, you might wonder why I'm not proposing to simply add a\n\"change-id\" to the commit object. The short answer is that the\n\"change-id\" Gerrit uses in the commit messages cannot stand on its own.\nIt depends on data stored on the server which maintains a relationship\nof commits to a review number and a linear ordering of commits within\nthe review (hopefully I'm not over simplifying this). The \"replaces\"\nreference is an attempt to make something which can stand on its own. I\ndon't think we need to solve the problem of where to keep comments at\nthis point.\n\nAn unbroken chain of \"replaces\" references obviates the need for the\nchange id in the commit message. From any given commit in the chain, we\ncan follow the references to the first commit which started the review.\nHowever, the chain is even more useful because it is not limited to a\nlinear progression of revisions. Let me try to explain how this can\nsolve some of the most common issues I ran into with the rebase type\nworkflow.\n\nLook at what happens in a rebase type workflow in any of the following\nscenarios. All of these came up regularly in my time with Gerrit.\n\n    1. Make a quick edit through the web UI then later work on the\n       change again in your local clone. It is easy to forget to pull\n       down the change made through the UI before starting to work on it\n       again. If that happens, the change made through the UI will\n       almost certainly be clobbered.\n\n    2. You or someone else creates a second change that is dependent on\n       yours and works on it while yours is still evolving. If the\n       second change gets rebased with an older copy of the base change\n       and then posted back up for review, newer work in the base change\n       has just been clobbered.\n\n    3. As a reviewer, you decide the best way to explain how you'd like\n       to see something done differently is to make the quick change\n       yourself and push it up. If the author fails to fetch what you\n       pushed before continuing onto something else, it gets clobbered.\n\n    4. You want to collaborate on a single change with someone else in\n       any way and for whatever reason. As soon as that change starts\n       hitting multiple work spaces, there are synchronization issues\n       that currently take careful manual intervention.\n\nThese kinds of scenarios usually end up being used as arguments against\na rebased based workflow. On the other hand, with a chain of \"replaces\"\nreferences, these scenarios end up branching the change. This is where\nit will be useful for my local git command to be able to pull down the\nupstream state, understand what branched, help me merge, and then push\nthe result upstream. I really think this brings the power and benefits\nof branching and merging to the rebase workflow.\n\nAnyway, now I am compelled to use github which is also a fine tool and I\nappreciate all of the work that has gone into it. About 80% of the time,\nI rebase and force push to my branch to update a pull request. I've come\nto like the end product of the rebase workflow. However, github doesn't\nexcel at this approach. For one, it doesn't preserve older revisions\nwhich were already reviewed which makes it is difficult for reviewers to\npick up where they left off the last time. If it preserved them, as\ngerrit does, the reviewer can compare a new revision with the most\nrecent older revision they reviewed to see just what has been addressed\nsince then.\n\nI think it would be great if git standardized the way that revisions of\n\"changes\" are preserved in the repository as commits. Preserving the\ncomments attached to each commit could be left up to each of the tools,\nin my opinion. Or, at least left to a later discussion.\n\nThe other 20% of the time, I revert to branching and merging only. This\nhelps reviewers with the problem of picking up where they left off. It\nis also an absolute necessity if another developer is going to be\ncollaborating with me on the pr or basing a dependent pr on it. However,\nit leaves the history in a more complex state in the end. Github offers\n\"squash merge\" and \"rebase merge\" to help out with this but these don't\ngive me the control I want over what goes into each change when I want\nto end up with more than one. They also cause problems if there is a\ndependent pr.\n\nAfter a pull request in github has progressed, it can be difficult for a\nnew reviewer to jump on. They may want to review a large pr commit by\ncommit from the bottom up instead of trying to tackle the entire thing\nat once. However, if fixups have been made in later commits, they will\nbe looking at the old stuff first, seeing all the bugs and issues that\nhave already been addressed.\n\n> You can then cut and paste the Change-Id into the Gerrit user\n> interface, and see the different commits, more important, the\n> discussion surrounding each change.\n> \n> \n> If the complaint about Gerrit is that it's not a core part of Git, the\n> challenge is (a) how to carry the code review comments in the git\n> repository, and (b) do so in a while that it doesn't bloat the core\n> repository, since most of the time, you *don't* want or need to keep a\n> local copy of all of the code review comments going back since the\n> beginning of the project.\n> \n> -------------\n\nI need to give your kernel patching use cases some thought. I once\ndesigned a process to do similar patching against a different project so\nI think I know what you're getting at. I just need a little more time to\nthink about it. Hopefully, I'll have a little more time to post another\nreply.\n\nCarl\n"},{"id":"335308","messageId":"CA+P7+xp9v8adrbF7JUYa3X+PvurHiW1QNTnodJt6-vyB3_dWAQ@mail.gmail.com","threadId":"47487","inReplyTo":"20171226001622.GA16219@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2017-12-26T01:28:10Z","receivedAt":"2017-12-26T01:28:36Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Mon, Dec 25, 2017 at 4:16 PM, Carl Baldwin <carl@ecbaldwin.net> wrote:\n> On Sat, Dec 23, 2017 at 11:09:59PM +0100, Ęvar Arnfjörš Bjarmason wrote:\n>> >> But I don't see why you think this needs a new \"replaces\" parent\n>> >> pointer orthagonal to parent pointers, i.e. something that would\n>> >> need to be a new field in the commit object (I may have misread the\n>> >> proposal, it's not heavy on technical details).\n>> >\n>> > Just to clarify, I am proposing a new \"replaces\" pointer in the commit\n>> > object. Imagine starting with rebase exactly as it works today. This new\n>> > field would be inserted into any new commit created by a rebase command\n>> > to reference the original commit on which it was based. Though, I'm not\n>> > sure if it would be better to change the behavior of the existing rebase\n>> > command, provide a switch or config option to turn it on, or provide a\n>> > new command entirely (e.g. git replay or git replace) to avoid\n>> > compatibility issues with the existing rebase.\n>>\n>> Yeah that sounds fine, I thought you meant that this \"replaces\" field\n>> would replace the \"parent\" field, which would require some rather deep\n>> incompatible changes to all git clients.\n>>\n>> But then I don't get why you think fetch/pull/gc would need to be\n>> altered, if it's because you thought that adding arbitrary *new* fields\n>> to the commit object would require changes to those that's not the case.\n>\n> Thank you again for your reply. Following is the kind of commit that I\n> would like to create.\n>\n>     tree fcce2f309177c7da9c795448a3e392a137434cf1\n>     parent b3758d9223b63ebbfbc16c9b23205e42272cd4b9\n>     replaces e8aa79baf6aef573da930a385e4db915187d5187\n>     author Carl Baldwin <carl@ecbaldwin.net> 1514057225 -0700\n>     committer Carl Baldwin <carl@ecbaldwin.net> 1514058444 -0700\n>\n> What will happen if I create this today? I assumed git would just choke\n> on it but I'm not certain. It has been a long time since I attempted to\n> get into the internals of git.\n>\n> Even if core git code does not simply choke on it, I would like push and\n> pull to follow these pointers and transfer the history behind them. I\n> assumed that git would not do this today. I would also like gc to\n> preserve e8aa79baf6 as if it were referenced by a parent pointer so that\n> it doesn't purge it from the history.\n>\n> I'm currently thinking of an example of the workflow that I'm after in\n> response to Theodore Ts'o's message from yesterday. Stay tuned, I hope\n> it makes it clearer why I want it this way.\n>\n> [snip]\n>\n>> Instead, if I understand what you're actually trying to do, it could\n>> also be done as:\n>>\n>>  1) Just add a new replaces <sha1> field to new commit objects\n>>\n>>  2) Make git-rebase know how to write those, e.g. add two of those\n>>     pointing to A & B when it squashes them into AB.\n>>\n>>  3) Write a history traversal mechanism similar to --full-history\n>>     that'll ignore any commits on branches that yield no changes, or\n>>     only those whose commits are referenced by this \"replaces\" field.\n>>\n>> You'd then end up with:\n>>\n>>  A) A way to \"stash\" these commits in the permanent history\n>>\n>>  B) ... that wouldn't be visble in \"git log\" by default\n>>\n>>  C) Would require no underlying changes to the commit model, i.e. it\n>>     would work with all past & future git clients, if they didn't know\n>>     about the \"replaces\" field they'd just show more verbose history.\n>\n> I get this point. I don't underestimate how difficult making such a\n> change to the core model is. I know there are older clients which cannot\n> simply be updated. There are also alternate implementations (e.g. jgit)\n> that also need to be considered. This is the thing I worry about the\n> most. I think at the very least, this new feature will have to be an\n> opt-in feature for teams who can easily ensure a minimum version of git\n> will be used. Maybe the core.repositoryformatversion config or something\n> like that would have to play into it. There may also be some minimal\n> amount that could be backported to older clients to at least avoid\n> choking on new repos (I know this doesn't guarantee older clients will\n> be updated). Just throwing a few ideas out.\n>\n> I want to be sure that the implications have been explored before giving\n> up and doing something external to git.\n>\n> Carl\n\nWhat about some way to take the reflog and turn it into a commit-based\nlinkage and export that? Rather than tying it into the individual\ncommit history, keep track of it outside the commit, possibly via\nsomething like notes, or some other mechanism.\n\nThis also ties into work done by Josh Triplett on git series [1] and\nsome previous mail discussions that I remember. He had some mechanism\nfor tracking series history which works ok, but can cause problems you\nmentioned when simply adding a second parent commit.\n\nI tend to think some mechanism to store both patch/commit history and\nreview based comments would be a very useful thing to integrate so\nthat multiple platforms had a more generic way of sharing things such\nas line-based commentary, and so on. It could even be some sort of\nunformatted method to at least get the mechanism of \"how to share this\namong clients\" to be stable across tools, even if each review tool\nmade its own format (thus we don't lock the *type* of review comments\nin stone).\n\nI definitely think that storing just the history of commits isn't as\nvaluable without storing the comments made in the review process.\n\n-Jake\n\n[1] https://github.com/git-series/git-series\n"},{"id":"335309","messageId":"CA+P7+xpj4o+N3uy2ea7DM-Y0oY_scayUARZMWP5QCwJEG02bZg@mail.gmail.com","threadId":"47487","inReplyTo":"20171226011638.GA16552@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2017-12-26T01:47:55Z","receivedAt":"2017-12-26T01:48:21Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Mon, Dec 25, 2017 at 5:16 PM, Carl Baldwin <carl@ecbaldwin.net> wrote:\n> Anyway, now I am compelled to use github which is also a fine tool and I\n> appreciate all of the work that has gone into it. About 80% of the time,\n> I rebase and force push to my branch to update a pull request. I've come\n> to like the end product of the rebase workflow. However, github doesn't\n> excel at this approach. For one, it doesn't preserve older revisions\n> which were already reviewed which makes it is difficult for reviewers to\n> pick up where they left off the last time. If it preserved them, as\n> gerrit does, the reviewer can compare a new revision with the most\n> recent older revision they reviewed to see just what has been addressed\n> since then.\n\nA bit of a tangent here, but a thought I didn't wanna lose: In the\ngeneral case where a patch was rebased and the original parent pointer\nwas changed, it is actually quite hard to show a diff of what changed\nbetween versions.\n\nThe best I've found is to do something like a 4-way --cc merge diff,\nwhich mostly works, but has a few awkward cases, and ends up usually\nshowing double ++ and -- notation.\n\nJust something I've thought about a fair bit, trying to come up with\nsome good way to show \"what changed between A1 and A2, but ignore all\nchanges between parent P1 and P2 which you don't care that much about\nin this context.\n\nThanks,\nJake\n"},{"id":"335312","messageId":"20171226060229.GB18783@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"CA+P7+xpj4o+N3uy2ea7DM-Y0oY_scayUARZMWP5QCwJEG02bZg@mail.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-26T06:02:30Z","receivedAt":"2017-12-26T06:02:42Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Mon, Dec 25, 2017 at 05:47:55PM -0800, Jacob Keller wrote:\n> On Mon, Dec 25, 2017 at 5:16 PM, Carl Baldwin <carl@ecbaldwin.net> wrote:\n> > Anyway, now I am compelled to use github which is also a fine tool and I\n> > appreciate all of the work that has gone into it. About 80% of the time,\n> > I rebase and force push to my branch to update a pull request. I've come\n> > to like the end product of the rebase workflow. However, github doesn't\n> > excel at this approach. For one, it doesn't preserve older revisions\n> > which were already reviewed which makes it is difficult for reviewers to\n> > pick up where they left off the last time. If it preserved them, as\n> > gerrit does, the reviewer can compare a new revision with the most\n> > recent older revision they reviewed to see just what has been addressed\n> > since then.\n> \n> A bit of a tangent here, but a thought I didn't wanna lose: In the\n> general case where a patch was rebased and the original parent pointer\n> was changed, it is actually quite hard to show a diff of what changed\n> between versions.\n>\n> The best I've found is to do something like a 4-way --cc merge diff,\n> which mostly works, but has a few awkward cases, and ends up usually\n> showing double ++ and -- notation.\n>\n> Just something I've thought about a fair bit, trying to come up with\n> some good way to show \"what changed between A1 and A2, but ignore all\n> changes between parent P1 and P2 which you don't care that much about\n> in this context.\n\nI ran into this all the time with gerrit. I wrote a script that you'd\nrun on a working copy (with no local changes). I'd fetch and checkout\nthe latest patchset that I want to review(say, for example, its patchset\n5) from gerrit. Then, say I wanted to compare it with patch set 3 which\nhas a different parent. I'd run this from the top level of my working\ncopy.\n\n    compare-to-previous-patchset 3\n\nIt would fetch patch set 3 from gerrit, rebase it to the same parent as\nthe current patch set on a detached HEAD and then git diff it with the\ncurrent patch set. If there were conflicts, it would just commit the\nconflict markers to the commit. There is no attempt to resolve the\nconflicts. The script was crude but it helped me out many times and it\nwas nice to be able to review how conflicts were resolved when those\ncame up.\n\nCarl\n\nPS In case you're curious, here's my script...\n\n#!/bin/bash\n\nremote=gerrit\nprevious_patchset=$1; shift\n\n# Assumes we're sitting on the latest patch set.\nnew_patch_set_id=$(git rev-parse HEAD)\n\nbranch=$(git branch | awk '/^\\*/ {print$2}')\n[ \"$branch\" = \"(no\" ] && branch=\n\n# set user, host, port, and project from git config\neval $(echo \"$(git config remote.$remote.url)\" |\n       sed 's,ssh://\\(.*\\)@\\(.*\\):\\([[:digit:]]*\\)/\\(.*\\).git,user=\\1 host=\\2 p<\n\ngerrit() {\n    ssh $user@$host -p $port gerrit ${1+\"$@\"}\n}\n\n# Grabs a bunch of information from gerrit about the current patch\neval $(gerrit query --current-patch-set $new_patch_set_id |\n    awk '\n        BEGIN {mode=\"main\"}\n        / currentPatchSet:/ { mode=\"currentPatchSet\" }\n        / ref:/ { printf \"new_patch_ref=%s\\n\", $2 }\n        / number:/ {\n            if (mode==\"main\") {\n                printf \"review_num=%s\\n\", $2\n            }\n            if (mode==\"currentPatchSet\") {\n                printf \"new_patchset=%s\\n\", $2\n            }\n        }\n    ')\n\n# Fetch the old patch set\nold_patch_ref=${new_patch_ref%$new_patchset}$previous_patchset\ngit fetch $remote $old_patch_ref && git checkout FETCH_HEAD\n\n# Rebase the old patch set to the parent of the new patch set.\nif ! git rebase HEAD^ --onto ${new_patch_set_id}^\nthen\n    git diff --name-only --diff-filter=U -z | xargs -0 git add\n    git rebase --continue\nfi\n\nprevious_patchset_rebased=$(git rev-parse HEAD)\n\n# Go back to the new patch set and diff it against the rebased old one.\nif [ \"$branch\" ]\nthen\n    git checkout $branch\nelse\n    git checkout $new_patch_set_id\nfi\ngit diff $previous_patchset_rebased\n"},{"id":"335313","messageId":"CA+P7+xpvuCjdnjyQxQg3B5iMwbnx-CerQMAP+bDQHR_-ALJOkQ@mail.gmail.com","threadId":"47487","inReplyTo":"20171226060229.GB18783@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2017-12-26T08:40:26Z","receivedAt":"2017-12-26T08:40:54Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Mon, Dec 25, 2017 at 10:02 PM, Carl Baldwin <carl@ecbaldwin.net> wrote:\n> On Mon, Dec 25, 2017 at 05:47:55PM -0800, Jacob Keller wrote:\n>> On Mon, Dec 25, 2017 at 5:16 PM, Carl Baldwin <carl@ecbaldwin.net> wrote:\n>> > Anyway, now I am compelled to use github which is also a fine tool and I\n>> > appreciate all of the work that has gone into it. About 80% of the time,\n>> > I rebase and force push to my branch to update a pull request. I've come\n>> > to like the end product of the rebase workflow. However, github doesn't\n>> > excel at this approach. For one, it doesn't preserve older revisions\n>> > which were already reviewed which makes it is difficult for reviewers to\n>> > pick up where they left off the last time. If it preserved them, as\n>> > gerrit does, the reviewer can compare a new revision with the most\n>> > recent older revision they reviewed to see just what has been addressed\n>> > since then.\n>>\n>> A bit of a tangent here, but a thought I didn't wanna lose: In the\n>> general case where a patch was rebased and the original parent pointer\n>> was changed, it is actually quite hard to show a diff of what changed\n>> between versions.\n>>\n>> The best I've found is to do something like a 4-way --cc merge diff,\n>> which mostly works, but has a few awkward cases, and ends up usually\n>> showing double ++ and -- notation.\n>>\n>> Just something I've thought about a fair bit, trying to come up with\n>> some good way to show \"what changed between A1 and A2, but ignore all\n>> changes between parent P1 and P2 which you don't care that much about\n>> in this context.\n>\n> I ran into this all the time with gerrit. I wrote a script that you'd\n> run on a working copy (with no local changes). I'd fetch and checkout\n> the latest patchset that I want to review(say, for example, its patchset\n> 5) from gerrit. Then, say I wanted to compare it with patch set 3 which\n> has a different parent. I'd run this from the top level of my working\n> copy.\n>\n>     compare-to-previous-patchset 3\n>\n> It would fetch patch set 3 from gerrit, rebase it to the same parent as\n> the current patch set on a detached HEAD and then git diff it with the\n> current patch set. If there were conflicts, it would just commit the\n> conflict markers to the commit. There is no attempt to resolve the\n> conflicts. The script was crude but it helped me out many times and it\n> was nice to be able to review how conflicts were resolved when those\n> came up.\n>\n> Carl\n>\n\nInteresting. That could work fairly well. I usually do something along\nthe lines of:\n\ngit diff patch-new patch-old patch-base-new patch-base-old --cc, which\nproduces a combined diff format patch which usually works ok.\n\nMy biggest gripes are that the gerrit web interface doesn't itself do\nsomething like this (and jgit does not appear to be able to generate\ncombined diffs at all!)\n\n> PS In case you're curious, here's my script...\n>\n> #!/bin/bash\n>\n> remote=gerrit\n> previous_patchset=$1; shift\n>\n> # Assumes we're sitting on the latest patch set.\n> new_patch_set_id=$(git rev-parse HEAD)\n>\n> branch=$(git branch | awk '/^\\*/ {print$2}')\n> [ \"$branch\" = \"(no\" ] && branch=\n>\n> # set user, host, port, and project from git config\n> eval $(echo \"$(git config remote.$remote.url)\" |\n>        sed 's,ssh://\\(.*\\)@\\(.*\\):\\([[:digit:]]*\\)/\\(.*\\).git,user=\\1 host=\\2 p<\n>\n> gerrit() {\n>     ssh $user@$host -p $port gerrit ${1+\"$@\"}\n> }\n>\n> # Grabs a bunch of information from gerrit about the current patch\n> eval $(gerrit query --current-patch-set $new_patch_set_id |\n>     awk '\n>         BEGIN {mode=\"main\"}\n>         / currentPatchSet:/ { mode=\"currentPatchSet\" }\n>         / ref:/ { printf \"new_patch_ref=%s\\n\", $2 }\n>         / number:/ {\n>             if (mode==\"main\") {\n>                 printf \"review_num=%s\\n\", $2\n>             }\n>             if (mode==\"currentPatchSet\") {\n>                 printf \"new_patchset=%s\\n\", $2\n>             }\n>         }\n>     ')\n>\n> # Fetch the old patch set\n> old_patch_ref=${new_patch_ref%$new_patchset}$previous_patchset\n> git fetch $remote $old_patch_ref && git checkout FETCH_HEAD\n>\n> # Rebase the old patch set to the parent of the new patch set.\n> if ! git rebase HEAD^ --onto ${new_patch_set_id}^\n> then\n>     git diff --name-only --diff-filter=U -z | xargs -0 git add\n>     git rebase --continue\n> fi\n>\n> previous_patchset_rebased=$(git rev-parse HEAD)\n>\n> # Go back to the new patch set and diff it against the rebased old one.\n> if [ \"$branch\" ]\n> then\n>     git checkout $branch\n> else\n>     git checkout $new_patch_set_id\n> fi\n> git diff $previous_patchset_rebased\n\nOne thing you might do is have it create a temporary worktree in order\nto avoid problems with being in the local checkout.\n\nThanks,\nJake\n"},{"id":"335323","messageId":"87vagtqszf.fsf@evledraar.gmail.com","threadId":"47487","inReplyTo":"20171226001622.GA16219@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2017-12-26T17:49:56Z","receivedAt":"2017-12-26T17:50:05Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Tue, Dec 26 2017, Carl Baldwin jotted:\n\n> On Sat, Dec 23, 2017 at 11:09:59PM +0100, Ævar Arnfjörð Bjarmason wrote:\n>> >> But I don't see why you think this needs a new \"replaces\" parent\n>> >> pointer orthagonal to parent pointers, i.e. something that would\n>> >> need to be a new field in the commit object (I may have misread the\n>> >> proposal, it's not heavy on technical details).\n>> >\n>> > Just to clarify, I am proposing a new \"replaces\" pointer in the commit\n>> > object. Imagine starting with rebase exactly as it works today. This new\n>> > field would be inserted into any new commit created by a rebase command\n>> > to reference the original commit on which it was based. Though, I'm not\n>> > sure if it would be better to change the behavior of the existing rebase\n>> > command, provide a switch or config option to turn it on, or provide a\n>> > new command entirely (e.g. git replay or git replace) to avoid\n>> > compatibility issues with the existing rebase.\n>>\n>> Yeah that sounds fine, I thought you meant that this \"replaces\" field\n>> would replace the \"parent\" field, which would require some rather deep\n>> incompatible changes to all git clients.\n>>\n>> But then I don't get why you think fetch/pull/gc would need to be\n>> altered, if it's because you thought that adding arbitrary *new* fields\n>> to the commit object would require changes to those that's not the case.\n>\n> Thank you again for your reply. Following is the kind of commit that I\n> would like to create.\n>\n>     tree fcce2f309177c7da9c795448a3e392a137434cf1\n>     parent b3758d9223b63ebbfbc16c9b23205e42272cd4b9\n>     replaces e8aa79baf6aef573da930a385e4db915187d5187\n>     author Carl Baldwin <carl@ecbaldwin.net> 1514057225 -0700\n>     committer Carl Baldwin <carl@ecbaldwin.net> 1514058444 -0700\n>\n> What will happen if I create this today? I assumed git would just choke\n> on it but I'm not certain. It has been a long time since I attempted to\n> get into the internals of git.\n\nNew headers should be added after existing headers, but other than that\nit won't choke on it. See 4b2bced559 when the encoding header was added,\nthis also passes most tests:\n\n    diff --git a/commit.c b/commit.c\n    index cab8d4455b..cd2bafbaa0 100644\n    --- a/commit.c\n    +++ b/commit.c\n    @@ -1565,6 +1565,8 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n            if (!encoding_is_utf8)\n                    strbuf_addf(&buffer, \"encoding %s\\n\", git_commit_encoding);\n\n    +       strbuf_addf(&buffer, \"replaces 0000000000000000000000000000000000000000\\n\");\n    +\n            while (extra) {\n                    add_extra_header(&buffer, extra);\n                    extra = extra->next;\n\nOnly \"most\" since of course this changes the sha1 of every commit git\ncreates from what you get now.\n\n> Even if core git code does not simply choke on it, I would like push and\n> pull to follow these pointers and transfer the history behind them. I\n> assumed that git would not do this today. I would also like gc to\n> preserve e8aa79baf6 as if it were referenced by a parent pointer so that\n> it doesn't purge it from the history.\n\nIt won't pay any attention to them if \"replaces\" is something entirely\nnew, what I was pointing out in my earlier reply is that you can simply\n*also* create the parent pointers to these no-op merge commits that hide\naway the previous history the \"replaces\" headers will be referencing.\n\nThe reason to do that is 100% backwards compatibility, and and only\nneeding to make minor UI changes to have this feature (to e.g. history\nwalking), as opposed to needing to hack everything that now follows\n\"parent\" or constructs a commit graph.\n\n> I'm currently thinking of an example of the workflow that I'm after in\n> response to Theodore Ts'o's message from yesterday. Stay tuned, I hope\n> it makes it clearer why I want it this way.\n>\n> [snip]\n>\n>> Instead, if I understand what you're actually trying to do, it could\n>> also be done as:\n>>\n>>  1) Just add a new replaces <sha1> field to new commit objects\n>>\n>>  2) Make git-rebase know how to write those, e.g. add two of those\n>>     pointing to A & B when it squashes them into AB.\n>>\n>>  3) Write a history traversal mechanism similar to --full-history\n>>     that'll ignore any commits on branches that yield no changes, or\n>>     only those whose commits are referenced by this \"replaces\" field.\n>>\n>> You'd then end up with:\n>>\n>>  A) A way to \"stash\" these commits in the permanent history\n>>\n>>  B) ... that wouldn't be visble in \"git log\" by default\n>>\n>>  C) Would require no underlying changes to the commit model, i.e. it\n>>     would work with all past & future git clients, if they didn't know\n>>     about the \"replaces\" field they'd just show more verbose history.\n>\n> I get this point. I don't underestimate how difficult making such a\n> change to the core model is. I know there are older clients which cannot\n> simply be updated. There are also alternate implementations (e.g. jgit)\n> that also need to be considered. This is the thing I worry about the\n> most. I think at the very least, this new feature will have to be an\n> opt-in feature for teams who can easily ensure a minimum version of git\n> will be used. Maybe the core.repositoryformatversion config or something\n> like that would have to play into it. There may also be some minimal\n> amount that could be backported to older clients to at least avoid\n> choking on new repos (I know this doesn't guarantee older clients will\n> be updated). Just throwing a few ideas out.\n\nSure, it could be opt in, be a new format etc. But you haven't explained\nwhy you think a feature like this would need to rely on an entirely new\nparent structure and side-DAG, as opposed to just the more minor changes\nI'm pointing out above, and which I think will give you what you need\nfrom a UX level.\n\n> I want to be sure that the implications have been explored before giving\n> up and doing something external to git.\n"},{"id":"335325","messageId":"20171226180436.GA28565@thunk.org","threadId":"47487","inReplyTo":"20171226011638.GA16552@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2017-12-26T18:04:36Z","receivedAt":"2017-12-26T18:04:45Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Dec 25, 2017 at 06:16:40PM -0700, Carl Baldwin wrote:\n> At this point, you might wonder why I'm not proposing to simply add a\n> \"change-id\" to the commit object. The short answer is that the\n> \"change-id\" Gerrit uses in the commit messages cannot stand on its own.\n> It depends on data stored on the server which maintains a relationship\n> of commits to a review number and a linear ordering of commits within\n> the review (hopefully I'm not over simplifying this). The \"replaces\"\n> reference is an attempt to make something which can stand on its own. I\n> don't think we need to solve the problem of where to keep comments at\n> this point.\n\nI strongly disagree, and one way to see that is by doing a real-life\nexperiment.  If you take a look at a gerrit change that, which in my\nexperience can have up to ten or twelve revisions, and strip out the\ncomments, so all you get to look at it is half-dozen or more\nrevisions.  How useful is it *really*?  How does it get used in\npractice?  What development problem does it help to solve?\n\nAnd when you say that it is a bug that the Gerrit Change-Id does not\nstand alone, consider that it can also be a *feature*.  If you keep\nall of this in the main repo, the number of commits can easily grow by\nan order of magnitude.  And these are commits that you have to keep\nforever, which means it slows down every subsequent git clone, git gc\noperation, git tag --contains search, etc.\n\nSo what are the benefits, and what are the costs?  If the benefits\nwere huge, then perhaps it would be worthwhile.  But if you lose a\nhuge amount of the value because you are missing the *why* between the\nhalf-dozen to dozen past revisions of the commit, then is it really\nworth it to adopt that particular workflow?\n\nIt seems to me your argument is contrasting a \"replaces\" pointer\nversus the github PR.  But compared to the Gerrit solution, I don't\nthink the \"replaces\" pointer proposal is as robust or as featureful.\nAlso, please keep in mind that just because it's in core git doesn't\nguarantee that Github will support it.  As far as I know github has\nzero support notes, for example.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"335327","messageId":"20171226194408.GA22855@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"87vagtqszf.fsf@evledraar.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-26T19:44:09Z","receivedAt":"2017-12-26T19:44:17Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Tue, Dec 26, 2017 at 06:49:56PM +0100, Ævar Arnfjörð Bjarmason wrote:\n> New headers should be added after existing headers, but other than\n> that it won't choke on it. See 4b2bced559 when the encoding header was\n> added, this also passes most tests:\n> \n>     diff --git a/commit.c b/commit.c\n>     index cab8d4455b..cd2bafbaa0 100644\n>     --- a/commit.c\n>     +++ b/commit.c\n>     @@ -1565,6 +1565,8 @@ int commit_tree_extended(const char *msg, size_t msg_len,\n>             if (!encoding_is_utf8)\n>                     strbuf_addf(&buffer, \"encoding %s\\n\", git_commit_encoding);\n> \n>     +       strbuf_addf(&buffer, \"replaces 0000000000000000000000000000000000000000\\n\");\n>     +\n>             while (extra) {\n>                     add_extra_header(&buffer, extra);\n>                     extra = extra->next;\n> \n> Only \"most\" since of course this changes the sha1 of every commit git\n> creates from what you get now.\n> \n> > Even if core git code does not simply choke on it, I would like push and\n> > pull to follow these pointers and transfer the history behind them. I\n> > assumed that git would not do this today. I would also like gc to\n> > preserve e8aa79baf6 as if it were referenced by a parent pointer so that\n> > it doesn't purge it from the history.\n> \n> It won't pay any attention to them if \"replaces\" is something entirely\n> new, what I was pointing out in my earlier reply is that you can simply\n> *also* create the parent pointers to these no-op merge commits that hide\n> away the previous history the \"replaces\" headers will be referencing.\n> \n> The reason to do that is 100% backwards compatibility, and and only\n> needing to make minor UI changes to have this feature (to e.g. history\n> walking), as opposed to needing to hack everything that now follows\n> \"parent\" or constructs a commit graph.\n\nThank you for clarifying this. I have learned something.\n\n> Sure, it could be opt in, be a new format etc. But you haven't\n> explained why you think a feature like this would need to rely on an\n> entirely new parent structure and side-DAG, as opposed to just the\n> more minor changes I'm pointing out above, and which I think will give\n> you what you need from a UX level.\n\nI have not wrapped my head around it enough to convince myself that it\ngives what I'm after. Let me spend a little more time with it to get a\nfeel for it.\n\nCarl\n"},{"id":"335328","messageId":"20171226203153.GA21429@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"20171226180436.GA28565@thunk.org","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-26T20:31:55Z","receivedAt":"2017-12-26T20:32:03Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Tue, Dec 26, 2017 at 01:04:36PM -0500, Theodore Ts'o wrote:\n> On Mon, Dec 25, 2017 at 06:16:40PM -0700, Carl Baldwin wrote:\n> > At this point, you might wonder why I'm not proposing to simply add a\n> > \"change-id\" to the commit object. The short answer is that the\n> > \"change-id\" Gerrit uses in the commit messages cannot stand on its own.\n> > It depends on data stored on the server which maintains a relationship\n> > of commits to a review number and a linear ordering of commits within\n> > the review (hopefully I'm not over simplifying this). The \"replaces\"\n> > reference is an attempt to make something which can stand on its own. I\n> > don't think we need to solve the problem of where to keep comments at\n> > this point.\n> \n> I strongly disagree, and one way to see that is by doing a real-life\n> experiment.  If you take a look at a gerrit change that, which in my\n> experience can have up to ten or twelve revisions, and strip out the\n> comments, so all you get to look at it is half-dozen or more\n> revisions.  How useful is it *really*?  How does it get used in\n> practice?  What development problem does it help to solve?\n\nI didn't mean to imply that we need to get along without the comments. I\nwas only pointing out that gerrit, github, other code review UIs have\nalready figured out how to store comments archored to specific revisions\nof files in the repository. I'm suggesting that we let them continue to\ndo that part while we take the first step of specifying how the\nintermediate revisions are kept.\n\nIf the various code review servers adopted this then we'd have a client\nside which could push up revisions for review to any of them. In\naddition, they'd all get the collaborative functionality that I\ndescribed in my reply to your previous message.\n\nWhat we get with this proposal is if I push up a review and that review\nis changed by someone (maybe even me) outside of my original workspace,\nmy client gives me the tools to detect it and merge with it. If I try to\npush over (clobber) that work then I get an error that the remote cannot\nbe fast-forwarded and I'm forced to fetch it and merge it.\n\nI get this while using the rebase methodology I've grown to enjoy having\nsince using gerrit and I end up with a mainline history that looks\nexactly the way I want it to.\n\n> And when you say that it is a bug that the Gerrit Change-Id does not\n> stand alone, consider that it can also be a *feature*.  If you keep\n> all of this in the main repo, the number of commits can easily grow by\n> an order of magnitude.  And these are commits that you have to keep\n> forever, which means it slows down every subsequent git clone, git gc\n> operation, git tag --contains search, etc.\n\nI didn't say it was a bug; just that it is at odds with what I'm hoping\nto do.\n\nI agree that the number of commits in the repository will go up.\nHowever, I think there will be ways to mitigate the costs.\n\nThe commits are not in the mainline history. So, I wouldn't expect a git\ntag --contains or most other commands that traverse history to consider\nthem at all.\n\nIt could be possible to make the default git clone skip them all and\nonly fetch them on demand for specific changes.\n\n> So what are the benefits, and what are the costs?  If the benefits\n> were huge, then perhaps it would be worthwhile.  But if you lose a\n> huge amount of the value because you are missing the *why* between the\n> half-dozen to dozen past revisions of the commit, then is it really\n> worth it to adopt that particular workflow?\n> \n> It seems to me your argument is contrasting a \"replaces\" pointer\n> versus the github PR.  But compared to the Gerrit solution, I don't\n> think the \"replaces\" pointer proposal is as robust or as featureful.\n> Also, please keep in mind that just because it's in core git doesn't\n> guarantee that Github will support it.  As far as I know github has\n> zero support notes, for example.\n\nWhat I propose is that gerrit and github could end up more robust,\nfeatureful, and interoperable if they had this feature to build from.\n\nWith gerrit specifically, adopting this feature would make the \"change\"\nconcept richer than it is now because it could supersede the change-id\nin the commit message and allow a change to evolve in a distributed\nnon-linear way with protection against clobbering work.\n\nI have no intention to disparage either tool. I love them both. They've\nboth made my career better in different ways. I know there is no\nguarantee that github, gerrit, or any other tool will do anything to\nadopt this. But, I'm hoping they are reading this thread and that they\nrecognize how this feature can make them a little bit better and jump in\nand help. I know it is a lot to hope for but I think it could be great\nif it happened.\n\nCarl\n"},{"id":"335329","messageId":"1514319542.2717.406.camel@mad-scientist.net","threadId":"47487","inReplyTo":"20171226194408.GA22855@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Paul Smith","fromEmail":"paul@mad-scientist.net","sentAt":"2017-12-26T20:19:02Z","receivedAt":"2017-12-26T20:42:13Z","isPatch":false,"sender":{"key":"paul@mad-scientist.net","avatar":"https://avatars.githubusercontent.com/u/109636?v=4"},"body":"On Tue, 2017-12-26 at 12:44 -0700, Carl Baldwin wrote:\n> > Sure, it could be opt in, be a new format etc. But you haven't\n> > explained why you think a feature like this would need to rely on\n> > an entirely new parent structure and side-DAG, as opposed to just\n> > the more minor changes I'm pointing out above, and which I think\n> > will give you what you need from a UX level.\n> \n> I have not wrapped my head around it enough to convince myself that\n> it gives what I'm after. Let me spend a little more time with it to\n> get a feel for it.\n\nAs someone working in an environment where we do a lot of rebasing and\nvery little merging, I read these proposals with interest.  I'm not\nconvinced that we would switch to using a \"replaces\"-type feature, but\nI'm pretty sure that the \"null-merge and rebase\" trick described\npreviously would not be something we're interested in using.\n\nAlthough \"git log\" doesn't follow these merges (unless requested), all\nthe graphical tools that are used to display history WOULD show all\nthose branches.  In a \"replaces\"-type environment I think the point is\nthat we would not want to see them (certainly not by default) as they\nwould be used mainly for deeper spelunking, but since they just seem\nlike normal merges I don't see any way to turn them off.\n\nIf \"replaces\" was a separate capability then it could be treated\ndifferently by history browsing tools, and shown or not shown as\ndesired.  For example, a commit that had a \"replaces\" element could be\nselected somehow and you could expand that set of commits that were\nreplaced, or something like that.\n"},{"id":"335330","messageId":"20171226210727.GB22855@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"1514319542.2717.406.camel@mad-scientist.net","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-26T21:07:29Z","receivedAt":"2017-12-26T21:07:36Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Tue, Dec 26, 2017 at 03:19:02PM -0500, Paul Smith wrote:\n> As someone working in an environment where we do a lot of rebasing and\n> very little merging, I read these proposals with interest.  I'm not\n> convinced that we would switch to using a \"replaces\"-type feature, but\n> I'm pretty sure that the \"null-merge and rebase\" trick described\n> previously would not be something we're interested in using.\n\nIn the near term, maybe. I'm still working with it to be sure I\nunderstand it right.\n\n> Although \"git log\" doesn't follow these merges (unless requested), all\n> the graphical tools that are used to display history WOULD show all\n> those branches.  In a \"replaces\"-type environment I think the point is\n> that we would not want to see them (certainly not by default) as they\n> would be used mainly for deeper spelunking, but since they just seem\n> like normal merges I don't see any way to turn them off.\n\nYou've touched on some of my concerns with the null-merge approach. I\nwant the end result to be as clean as possible which I think is what\nlures many to the rebase methodology in the first place.\n\n> If \"replaces\" was a separate capability then it could be treated\n> differently by history browsing tools, and shown or not shown as\n> desired.  For example, a commit that had a \"replaces\" element could be\n> selected somehow and you could expand that set of commits that were\n> replaced, or something like that.\n\nExactly!\n\nCarl\n"},{"id":"335331","messageId":"20171226040843.h7o6txkrp6zlv7u5@glandium.org","threadId":"47487","inReplyTo":"CALiLy7pBvyqA+NjTZHOK9t0AFGYbwqwRVD3sZjUg0ZLx5y1h3A@mail.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2017-12-26T04:08:45Z","receivedAt":"2017-12-26T21:34:04Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Fri, Dec 22, 2017 at 11:10:19PM -0700, Carl Baldwin wrote:\n> The big contention among git users is whether to rebase or to merge\n> changes [2][3] while iterating. I used to firmly believe that merging\n> was the way to go and rebase was harmful. More recently, I have worked\n> in some environments where I saw rebase used very effectively while\n> iterating on changes and I relaxed my stance a lot. Now, I'm on the\n> fence. I appreciate the strengths and weaknesses of both approaches. I\n> waffle between the two depending on the situation, the tools being\n> used, and I guess, to some extent, my mood.\n> \n> I think what git needs is something brand new that brings the two\n> together and has all of the advantages of both approaches. Let me\n> explain what I've got in mind...\n> \n> I've been calling this proposal `git replay` or `git replace` but I'd\n> like to hear other suggestions for what to name it. It works like\n> rebase except with one very important difference. Instead of orphaning\n> the original commit, it keeps a pointer to it in the commit just like\n> a `parent` entry but calls it `replaces` instead to distinguish it\n> from regular history. In the resulting commit history, following\n> `parent` pointers shows exactly the same history as if the commit had\n> been rebased. Meanwhile, the history of iterating on the change itself\n> is available by following `replaces` pointers. The new commit replaces\n> the old one but keeps it around to record how the change evolved.\n> \n> The git history now has two dimensions. The first shows a cleaned up\n> history where fix ups and code review feedback have been rolled into\n> the original changes and changes can possibly be ordered in a nice\n> linear progression that is much easier to understand. The second\n> drills into the history of a change. There is no loss and you don't\n> change history in a way that will cause problems for others who have\n> the older commits.\n> \n> Replay handles collaboration between multiple authors on a single\n> change. This is difficult and prone to accidental loss when using\n> rebase and it results in a complex history when done with merge. With\n> replay, collaborators could merge while collaborating on a single\n> change and a record of each one's contributions can be preserved.\n> Attempting this level of collaboration caused me many headaches when I\n> worked with the gerrit workflow (which in many ways, I like a lot).\n> \n> I blogged about this proposal earlier this year when I first thought\n> of it [1]. I got busy and didn't think about it for a while. Now with\n> a little time off of work, I've come back to revisit it. The blog\n> entry has a few examples showing how it works and how the history will\n> look in a few examples. Take a look.\n> \n> Various git commands will have to learn how to handle this kind of\n> history. For example, things like fetch, push, gc, and others that\n> move history around and clean out orphaned history should treat\n> anything reachable through `replaces` pointers as precious. Log and\n> related history commands may need new switches to traverse the history\n> differently in different situations. Bisect is a interesting one. I\n> tend to think that bisect should prefer the regular commit history but\n> have the ability to drill into the change history if necessary.\n> \n> In my opinion, this proposal would bring together rebase and merge in\n> a powerful way and could end the contention. Thanks for your\n> consideration.\n\nFWIW, your proposal has a lot in common (but is not quite equivalent) to\nmercurial's obsolescence markers and changeset evolution features.\n\nMike\n"},{"id":"335342","messageId":"c03c67f2-7d2c-da94-08f8-48c41c2a55ed@gmail.com","threadId":"47487","inReplyTo":"CA+P7+xp9v8adrbF7JUYa3X+PvurHiW1QNTnodJt6-vyB3_dWAQ@mail.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Igor Djordjevic","fromEmail":"igor.d.djordjevic@gmail.com","sentAt":"2017-12-26T23:30:57Z","receivedAt":"2017-12-26T23:31:14Z","isPatch":false,"sender":{"key":"igor.d.djordjevic@gmail.com","avatar":null},"body":"Very interesting topic, just this one part I wanted to comment on:\n\nOn 26/12/2017 02:28, Jacob Keller wrote:\n> \n> What about some way to take the reflog and turn it into a commit-based\n> linkage and export that? Rather than tying it into the individual\n> commit history, keep track of it outside the commit, possibly via\n> something like notes, or some other mechanism.\n\nThis seems like the most useful approach, might be not touching reflog \nper se, but having some kind of \"cherry-picked commits source\" log \n(where rebasing is a subset of cherry-picking). What Johannes \nmentioned, a mapping between \"old\" and \"new\" commits. Might be notes \ncould fit in nicely, but I`m not competent to comment on that at the \nmoment.\n\nFor me, the most interesting use case is not even tied to code review \n(thus no review comments to think about), but a situation where one \nmight be rebasing a set of downstream patches on top of updating \nupstream - it might be possible for a bug to slip through due to some \nupstream changes, even where there are no conflicts and test suite is \nexecuted regularly (might be test reveling the bug is yet to be added).\n\nIn that situation, instead of just going back in \"regular\" history \n(single dimension) and eventually finding the offending (rebased) \ncommit (its N-th rebased version, that is), it might be great to \nactually keep drilling down the \"rebase history\" now (second \ndimension), finding the exact rebase iteration / rebased commit \nversion where the error first appeared.\n\nCarl, you described this well in your document[1], and Johannes \nprovided a valuable first-hand experience[2] from working around the \nvery same native Git limitation for years, mentioning using (fragile, \ncostly and not very automatible) rebased commits message search to \ndrill down the second dimension (rebase iterations), which seems to \nbe the only possible approach at the moment, with \"vanilla\" Git, at \nleast.\n\nSo this might be much more interesting case, if code review one is \nless appropriate because of review comments being also relevant to \ncommit rebase iterations (which should be then stored somewhere, too, \nrelating to corresponding commits, not to lose context).\n\nRegards, Buga\n\np.s. \"Merging rebase\" and \"shears.sh\" script[3] seem to be orthogonal \nto this - really great on their own in improving rebase itself and \nmaking it smarter and much more powerful and useful, where I guess \nthey would benefit from native Git \"cherry-picked (rebased) commits \niterations tracking\" (old/source <> new/destination commit mapping), \ntoo, as would other Git tools.\n\n[1] http://blog.episodicgenius.com/post/merge-or-rebase--neither/\n[2] https://public-inbox.org/git/20171226040843.h7o6txkrp6zlv7u5@glandium.org/T/#m2e5079488bed2968d4ea52a10051a06c06ff61e0\n[3] https://github.com/git-for-windows/build-extra/blob/af9cff5005/shears.sh#L12-L18\n"},{"id":"335345","messageId":"20171227024413.GA26579@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"20171226040843.h7o6txkrp6zlv7u5@glandium.org","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-27T02:44:15Z","receivedAt":"2017-12-27T02:44:24Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Tue, Dec 26, 2017 at 01:08:45PM +0900, Mike Hommey wrote:\n> FWIW, your proposal has a lot in common (but is not quite equivalent)\n> to mercurial's obsolescence markers and changeset evolution features.\n\nI've had experience with mercurial but not since about 2009. After\nreading up a little bit on this changeset evolution feature, it looks\nvery much like what I'm proposing. Obsolescence markers look a lot like\nreplaces references except, as illustrated by this blog [1], they point\nthe other way! Hence, the illustrations confused me for a moment. It\nseems more natural to embed the reference in the new commit pointing at\nthe old. That said, the illustrated direction of the arrows doesn't\nreally affect the usefulness of the idea.\n\nHis third example (#3-working-with-other-people), appears to be the kind\nof collaboration that I'm trying to describe here. To quote the blog:\n\n  In git or vanilla (no extension) mercurial, you would have to figure\n  out that b’ and b” are two new versions of b and merge them. Changeset\n  evolution detects that situation, marks b’ and b” as being divergent.\n  It then suggests automatic resolution with a merge and preserves\n  history.\n\nThis is the kind of thing that I had to deal with manually in gerrit. I\nhadn't seen this feature in mercurial but I'm glad to know now there is\na precedent for it.\n\nCarl\n\n[1] https://blog.laurentcharignon.com/post/2016-02-02-changeset-evolution/\n"},{"id":"335346","messageId":"20171227043544.GB26579@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"20171225035215.GC1257@thunk.org","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-27T04:35:46Z","receivedAt":"2017-12-27T04:36:03Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Sun, Dec 24, 2017 at 10:52:15PM -0500, Theodore Ts'o wrote:\n> Here's another potential use case.  The stable kernels (e.g., 3.18.y,\n> 4.4.y, 4.9.y, etc.) have cherry picks from the the upstream kernel,\n> and this is handled by putting in the commit body something like this:\n> \n>     [ Upstream commit 3a4b77cd47bb837b8557595ec7425f281f2ca1fe ]\n\nI think replaces could apply to cherry picks like this too. The more I\nthink about it, I actually think that replaces isn't a bad name for it\nin the cherry pick context. When you cherry pick a commit, you create a\nnew commit that is derived from it and stands in for or replaces it in\nthe new context. It is a stretch but I don't think it is that bad.\n\nYou can tell that it is a cherry pick because the referenced commit's\nhistory is not reachable in the current context.\n\nThough we could consider some different names like \"derivedfrom\",\n\"obsoletes\", \"succeeds\", \"supersedes\", \"supplants\"\n\n> ----\n> \n> And here's yet another use case.  For internal Google kernel\n> development, we maintain a kernel that has a large number of patches\n> on top of a kernel version.  When we backport an upstream fix (say,\n> one that first appeared in the 4.12 version of the upstream kernel),\n> we include a line in the commit body that looks like this:\n> \n> Upstream-4.12-SHA1: 5649645d725c73df4302428ee4e02c869248b4c5\n> \n> This is useful, because when we switch to use a newer upstream kernel,\n> we need make sure we can account for all patches that were built on\n> top of the 3xx kernel (which might have been using 4.10, for the sake\n> of argument), to the 4xx kernel series (which might be using 4.15 ---\n> the version numbers have been changed to protect the innocent).  This\n> means going through each and every patch that was on top of the 3xx\n> kernel, and if it has a line such as \"Upstream 4.12-SHA1\", we know\n> that it will already be included in a 4.15 based kernel, so we don't\n> need to worry about carrying that patch forward.\n\nAre 3xx and 4xx internal version numbers? If I understand correctly, in\nyour example, 3xx is the heavily patched internal kernel based on 4.10\nand 4xx is the internal patched version of 4.15. I think I'm following\nso far.\n\nLet's say that you used a \"replaces\" reference instead of your\n\"Upstream-4.12-SHA1\" reference. The only piece of metadata that is\nmissing is the \"4.12\" of your string. However, you could replicate this\nwith some set arithmetic. If the sha1 referred to by \"replaces\" exists\nin the set of commits reachable from 4.15 then you've answered the same\nquestion.\n\n> In other cases, we might decide that the patch is no longer needed.\n> It could be because the patch has already be included upstream, in\n> which case we might check in a commit with an empty patch body, but\n> whose header contains something like this in the 4xx kernel:\n> \n> Origin-3xx-SHA1: fe546bdfc46a92255ebbaa908dc3a942bc422faa\n> Upstream-Dropped-4.11-SHA1: d90dc0ae7c264735bfc5ac354c44ce2e\n\nSo, the first reference is the old commit that patched the 3xx series?\nWhat is the second reference? What is \"4.11\" indicating? Is that the\npatch that was included in the upstream kernel that obsoleted your 3xx\npatch?\n\nIf I understood that correctly. You could use a \"replaces\" reference for\nthe first line and the second line would still have to be included as a\nseparate header in your commit message? Does this mean \"replaces\" is not\nuseful in your case? I don't think so.\n\n> Or we could decide that the commit is no longer no longer needed ---\n\nno longer no longer needed? Is this a double negative indicating that it\nis needed again? Or, is it a mistake?\n\n> perhaps because the relevant subsystem was completely rewritten and\n> the functionality was added in a different way.  Then we might have\n> just have an empty commit with an explanation of why the commit is no\n> longer needed and the commit body would have the metadata:\n> \n> Origin-Dropped-3xx-SHA1: 26f49fcbb45e4bc18ad5b52dc93c3afe\n\nThe metadata in this reference indicates that it was dropped since 3xx.\nDoesn't the empty body (and maybe a commit message saying dropping a\npatch) indicate this if a \"references\" pointer were used instead? The\n3xx part of the metadata could be derived again by set arithmetic.\n\n> Or perhaps the commit is still needed, and for various reasons the\n> commit was never upstreamed; perhaps because it's only useful for\n> Google-specific hardware, or the patch was rejected upstream.  The we\n> will have a cherry-pick that would include in the body:\n> \n> Origin-3xx-SHA1: 8f3b6df74b9b4ec3ab615effb984c1b5\n\nReplaces reference and set arithmetic.\n\n> (Note: all commits that are added in the rebase workflow, even the\n> empty commits that just have the Origin-Dropped-3xx-SHA1 or\n> Upstream-Droped-4.11-SHA1 headers, are patch reviewed through Gerrit,\n> so we have an audited, second-engineer review to make sure each commit\n> in the 3xx kernel that Google had been carrying had the correct\n> disposition when rebasing to the 4xx kernel.)\n\nThis is great! I designed a strikingly similar workflow for local\npatches to Openstack Neutron about four years ago. Each time we moved\nforward to a new version of upstream, we went through a very similar\nprocess. I don't have access to those scripts any longer but here is\nwhat I recall.\n\nWith each now upstream version, we'd generate a list of commits we\ncreated locally using git. I recall it being a fairly simple set\ndifference between the upstream tag and our downstream tag. I wrote\nscripts that would take each of them and proposed them as new gerrit\nreviews against the new upstream. I made sure to keep the same gerrit\nchange-id as the old one in the new review. Gerrit allowed this because\nit was a new branch for each now revision and gerrit allows the same\nchange-id in different reviews as long as the branch is different. I'm\nnot sure if you kept the same change-id or not but it was very useful to\nme. I could see the history of the patch applied to various upstream\nversions with the click of a link in gerrit.\n\nMy automated script would cherry pick but wouldn't attempt to resolve\nany conflicts. Instead, it would commit the conflict markers exactly as\nthey occurred and flag the review as conflicted. A human would have to\ncome along and recreate the cherry pick in their own workspace to\nresolve the conflicts. Then they'd post the result to the same gerrit\nreview as the second patch set. This way, the resolution of the\nconflicts could easily be reviewed in the tool.\n\nWe'd run our CI against each and every one also to be sure that we\nweren't breaking it.\n\nSometimes, patches were rendered obsolete by something upstream. In this\ncase, we would close the gerrit review without merging indicating that\nthe patch was dropped and the reason.\n\nFor patches that we proposed upstream, we'd use the same gerrit id in\nthe upstream review. This helped us tie them together and identify them\nas equivalent patches.\n\nWhen all of the gerrit reviews for all of the patches were merged, we'd\nmerge the result to master. I recall doing something special for this\nfinal merge but I don't recall exactly what it was. Maybe it was to use\nthe \"theirs\" strategy or something like that.\n\nI remember having a script called \"delinearize\" which would actually\nfind the minimum chain of preceding patches that had to come before a\ngiven one in order for it to rebase cleanly to the upstream base. Most\nof the time, the patches rebased cleanly to the upstream. This meant\nthat they really didn't depend on any of the patches that came before\nthem. This was very useful because it allowed us humans to switch\nbetween a bunch of mostly independent gerrit reviews for the newly\ncherry-picked patches and do a lot of things in parallel. We could merge\nthem in any order as long as they went through the entire review\nprocess.\n\n> The point is that for this much more complex, real-world workflow, we\n> need much *more* metadata than a simple \"Replaces\" metadata.  (And we\n> also have other metadata --- for example, we have a \"Tested: \" trailer\n> that explains how to test the commit, or which unit test can and\n> should be used to test this commit, combined with a link to the test\n> log in our automated unit tester that has the test run, and a\n> \"Rebase-Tested-4xx: \" trailer that might just have the URL to the test\n> log when the commit was rebased since the testing instructions in the\n> Tested: trailer is still relevant.)\n\nSo far, I haven't seen that it is that much more complex. I've actually\nhad experience doing practically the same thing. Yes, it was a complex\nprocess but we didn't need much more than the gerrit-id in the reviews\nand an external dashboard listing the patches out with a link to each\nreview. Even the dashboard was pretty much obsolete once the whole\nprocess was declared done for a given upstream revision.\n\nYour \"Tested\" trailer sounds completely orthogonal. I'm not sure that\nshowing a need for other orthogonal metadata is necessarily a good\nargument against my proposal. It doesn't seem relevant.\n\n> And since this metadata is not really needed by the core git\n> machinery, we just use text trailers in the commit body; it's not hard\n> to write code which parses this out of the git commit.\n> \n> > Various git commands will have to learn how to handle this kind of\n> > history. For example, things like fetch, push, gc, and others that\n> > move history around and clean out orphaned history should treat\n> > anything reachable through `replaces` pointers as precious. Log and\n> > related history commands may need new switches to traverse the history\n> > differently in different situations.\n> \n> I'd encourage you to think very hard about how exactly \"git log\" and\n> \"gitk\" might actually deal with these links.  In the Google kernel\n> development use cases, we use different repos for the 3xx and 4xx\n> kernels.  It would be possible to make hot links for the\n> Original-3xx-SHA1: trailers, but you couldn't do it using gitk.  It\n> would actually have to be a completely new tool.  (And we do have new\n> tools, most especially a dashboard so we can keep track of how many\n> commits in the 3xx kernel still have to be rebased to the 4xx kernel,\n> or can be confirmed to be in the upstream kernel, or can be confirmed\n> to be dropped.  We have a *large* number of patches that we carry, so\n> it's a multi-month effort involving a large number of engineers\n> working together to do a kernel rebase operation from a 4.x upstream\n> kernel to a 4.y upstream kernel.  So having a dashboard is useful\n> because we can see whether a particular subsystem team is ahead or\n> behind the curve in terms of handling those commits which are their\n> responsibility.)\n> \n> My experience, from seeing these much more complex use cases ---\n> starting with something as simple as the Linux Kernel Stable Kernel\n> Series, and extending to something much more complex such as the\n> workflow that is used to support a Google Kernel Rebase, is that using\n> just a simple extra \"Replaces\" pointer in the commit header is not\n> nearly expressive enough.  And, if you make it a core part of the\n> commit data structure, there are all sorts of compatibility headaches\n> with older versions of git that wouldn't know about it.  And if it\n\nThe more I think about this, the less I worry. Be sure that you're using \n\n> then turns out it's not sufficient more the more complex workflows\n> *anyway*, maybe adding a new \"replace\" pointer in the core git data\n> structures isn't worth it.  It might be that just keeping such things\n> as trailers in the commit body might be the better way to go.\n\nIt doesn't need to be everything to everyone to be useful. I hope to\nshow in this thread that it is useful enough to be a compelling addition\nto git. I think I've also shown that it could be used as a part of your\nmore complex workflow. Maybe even a bigger part of it than you had\nrealized.\n\nCarl\n"},{"id":"335356","messageId":"C82A30ED-D608-4F79-B824-C23DDB078DD9@gmail.com","threadId":"47487","inReplyTo":"20171227043544.GB26579@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Alexei Lozovsky","fromEmail":"a.lozovsky@gmail.com","sentAt":"2017-12-27T13:35:58Z","receivedAt":"2017-12-27T13:36:10Z","isPatch":false,"sender":{"key":"a.lozovsky@gmail.com","avatar":"https://gravatar.com/avatar/8bb8ff5ec366dd64bd8e08f768082934513da367ae047bcb0039292e4ed6bda5?d=mp&s=160"},"body":"On Dec 27, 2017, at 06:35, Carl Baldwin <carl@ecbaldwin.net> wrote:\n> \n> On Sun, Dec 24, 2017 at 10:52:15PM -0500, Theodore Ts'o wrote:\n>> \n>> My experience, from seeing these much more complex use cases ---\n>> starting with something as simple as the Linux Kernel Stable Kernel\n>> Series, and extending to something much more complex such as the\n>> workflow that is used to support a Google Kernel Rebase, is that using\n>> just a simple extra \"Replaces\" pointer in the commit header is not\n>> nearly expressive enough.  And, if you make it a core part of the\n>> commit data structure, there are all sorts of compatibility headaches\n>> with older versions of git that wouldn't know about it.  And if it\n> \n> The more I think about this, the less I worry. Be sure that you're using \n> \n>> then turns out it's not sufficient more the more complex workflows\n>> *anyway*, maybe adding a new \"replace\" pointer in the core git data\n>> structures isn't worth it.  It might be that just keeping such things\n>> as trailers in the commit body might be the better way to go.\n> \n> It doesn't need to be everything to everyone to be useful. I hope to\n> show in this thread that it is useful enough to be a compelling addition\n> to git. I think I've also shown that it could be used as a part of your\n> more complex workflow. Maybe even a bigger part of it than you had\n> realized.\n\nI think the reasoning behind Theo's words is that it would be better to\nfirst implement the commit relationship tracking as an add-in which uses\ncommit messages for data storage, then evaluate its usefulness when it's\nactually available (including extensions to gitk and stuff to support the\nnew metadata), and then it could be moved into core git data structures,\nwhen it has proven itself useful. It's not a trivial feature which warrants\nimmediate addition to git and its design can change when faced with real-\nworld use-cases, so it would be bad for compatibility to rush its addition.\nStorage location for metadata seems to be an implementation detail which\ncould be technically changed more or less easily. But it's much easier to\nignore a trailer in commit message in the favor of a commit header field\nthan to replace a deprecated commit header field with a better one, which\ncould cause massive headache for all git repositories in the world.\n"},{"id":"335456","messageId":"20171228052303.GA33027@Carl-MBP","threadId":"47487","inReplyTo":"C82A30ED-D608-4F79-B824-C23DDB078DD9@gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2017-12-28T05:23:08Z","receivedAt":"2017-12-28T05:23:19Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Wed, Dec 27, 2017 at 03:35:58PM +0200, Alexei Lozovsky wrote:\n> I think the reasoning behind Theo's words is that it would be better\n> to first implement the commit relationship tracking as an add-in which\n> uses commit messages for data storage, then evaluate its usefulness\n> when it's actually available (including extensions to gitk and stuff\n> to support the new metadata), and then it could be moved into core git\n> data structures, when it has proven itself useful. It's not a trivial\n> feature which warrants immediate addition to git and its design can\n> change when faced with real- world use-cases, so it would be bad for\n> compatibility to rush its addition. Storage location for metadata\n> seems to be an implementation detail which could be technically\n> changed more or less easily. But it's much easier to ignore a trailer\n> in commit message in the favor of a commit header field than to\n> replace a deprecated commit header field with a better one, which\n> could cause massive headache for all git repositories in the world.\n\nYeah, this is a point that everyone is eager to make instead of really\ntrying to understand what I'm trying to do and offering constructive\nsuggestions. It's not that I'm not listening. I'm not really concerned\nabout headers vs trailers or the asthetics of the whole thing as much as\nI'm concerned about how the server / client interaction will be. I worry\nthat anything that I come up with that isn't implemented in the regular\ngit core push and fetch will end up being awkward or end up needing to\nreimplement a lot of what's already in git. But, maybe it just needs a\nlittle more thought. Let me try to think through it...\n\nImagine John posts a new change up for review to a review server. The\ncurrent master points at commit A and so he grabs it and drafts his\nfirst proposal, B1.\n\n    digraph history {\n        B1 -> A\n    }\n\nSoon after posting, he notices a couple of simple errors and uses the\nweb UI to correct them. This creates B2. (Dashed edges are replaces\nreferences).\n\n    digraph history {\n        B1 -> A\n        B2 -> A\n        B2 -> B1 [ style=\"dashed\"; ]\n    }\n\nAnna reviews B2 and finds a small nit. She asks John if she can just fix\nit and push up a new review. He agrees. She pushes up B3.\n\n    digraph history {\n        B1 -> A\n        B2 -> A\n        B3 -> A\n        B2 -> B1 [ style=\"dashed\"; ]\n        B3 -> B2 [ style=\"dashed\"; ]\n    }\n\nJohn goes back to his workspace and does a little more work on B. He\ncreates the fourth revision, B4 but since he didn't update his workspace\nwith the other two most recent revisions, his new revision is derived\nfrom B1.\n\n    digraph history {\n        B1 -> A\n        B2 -> A\n        B3 -> A\n        B4 -> A\n        B2 -> B1 [ style=\"dashed\"; ]\n        B3 -> B2 [ style=\"dashed\"; ]\n        B4 -> B1 [ style=\"dashed\"; ]\n    }\n\nJohn then pushes to the server. I imagined that would be a command\nsimilar to what gerrit does.\n\n    git push codereview refs/for/master\n\nAt this point, I want a couple of things to happen. First, the server\nshould be able to match the new revision to the change by following the\nreplaces references to the commits it already has. Then it should\nrecognize that this is not a fast forward update to the change and\nreject it on those grounds.\n\nAfter that, John needs to be able to fetch B2 and B3 so that his local\nclient can perform a merge. I guess John needs to know what change he's\ntrying to fetch. In this case, he needs to fetch both B2 and B3 in order\nget the full history graph of the change. The problem I see here is that\ntoday's git fetch would see B2 and B3 as unrelated branches. There could\nbe any number of them to fetch. So, how does he ask for everything\nrelated to the change? Does he do a wild card or something?\n\n    git fetch codereview refs/changes/123/*\n\nOr does he just fetch all refs (this could be many on a busy review\nserver)? Or do we need to do something out of band to discover the list\nof references that need to be fetched?\n\nI've been thinking out loud a bit. I guess this could be a path forward.\nI guess to make gc happy, I've got to keep around a ref pointing at each\nnew revision so that it doesn't get garbage collected.\n\nCarl\n"},{"id":"335820","messageId":"alpine.DEB.2.21.1.1801041643400.31@MININT-6BKU6QN.europe.corp.microsoft.com","threadId":"47487","inReplyTo":"16725929-1BD2-44D3-8E71-E97C4A2C4034@gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2018-01-04T15:44:32Z","receivedAt":"2018-01-04T15:44:44Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 24 Dec 2017, Alexei Lozovsky wrote:\n\n> On Dec 24, 2017, at 01:01, Johannes Schindelin wrote:\n> > \n> > Hi Carl,\n> > \n> > On Sat, 23 Dec 2017, Carl Baldwin wrote:\n> > \n> >> I imagine that a \"git commit --amend\" would also insert a \"replaces\"\n> >> reference to the original commit but I failed to mention that in my\n> >> original post.\n> > \n> > And cherry-pick, too, of course.\n> \n> Why would it?\n\nBecause that's the command you use if you perform an interactive rebase\n\"manually\". Or if you need to split a topic branch into two.\n\nCiao,\nJohannes\n"},{"id":"335841","messageId":"2180514.YZ24uruNv3@mfick-lnx","threadId":"47487","inReplyTo":"CA+P7+xpvuCjdnjyQxQg3B5iMwbnx-CerQMAP+bDQHR_-ALJOkQ@mail.gmail.com","subject":"Re: Bring together merge and rebase","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-01-04T19:19:34Z","receivedAt":"2018-01-04T19:19:42Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Tuesday, December 26, 2017 12:40:26 AM Jacob Keller \nwrote:\n> On Mon, Dec 25, 2017 at 10:02 PM, Carl Baldwin \n<carl@ecbaldwin.net> wrote:\n> >> On Mon, Dec 25, 2017 at 5:16 PM, Carl Baldwin \n<carl@ecbaldwin.net> wrote:\n> >> A bit of a tangent here, but a thought I didn't wanna\n> >> lose: In the general case where a patch was rebased\n> >> and the original parent pointer was changed, it is\n> >> actually quite hard to show a diff of what changed\n> >> between versions.\n> \n> My biggest gripes are that the gerrit web interface\n> doesn't itself do something like this (and jgit does not\n> appear to be able to generate combined diffs at all!)\n\nI believe it now does, a presentation was given at the \nGerrit User summit in London describing this work.  It would \nindeed be great if git could do this also!\n\n-Martin \n\n\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"335854","messageId":"2363617.2KWG4LOxlS@mfick-lnx","threadId":"47487","inReplyTo":"alpine.DEB.2.21.1.1712232353390.406@MININT-6BKU6QN.europe.corp.microsoft.com","subject":"Re: Bring together merge and rebase","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-01-04T19:49:24Z","receivedAt":"2018-01-04T19:49:36Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Sunday, December 24, 2017 12:01:38 AM Johannes Schindelin \nwrote:\n> Hi Carl,\n> \n> On Sat, 23 Dec 2017, Carl Baldwin wrote:\n> > I imagine that a \"git commit --amend\" would also insert\n> > a \"replaces\" reference to the original commit but I\n> > failed to mention that in my original post.\n> \n> And cherry-pick, too, of course.\n> \n> Both of these examples hint at a rather huge urge of some\n> users to turn this feature off because the referenced\n> commits may very well be throw-away commits in their\n> case, making the newly-recorded information completely\n> undesired.\n> \n> Example: I am working on a topic branch. In the middle, I\n> see a typo. I commit a fix, continue to work on the topic\n> branch. Later, I cherry-pick that commit to a separate\n> topic branch because I really don't think that those two\n> topics are related. Now I definitely do not want a\n> reference of the cherry-picked commit to the original\n> one: the latter will never be pushed to a public\n> repository, and gc'ed in a few weeks.\n> \n> Of course, that is only my wish, other users in similar\n> situations may want that information. Demonstrating that\n> you would be better served with an opt-in feature that\n> uses notes rather than a baked-in commit header.\n\nI think what you are highlighting is not when to track this, \nbut rather when to share this tracking.  In my local repo, I \nwould definitely want to know that I cherry-picked this from \nelsewhere, it helps me understand what I have done later \nwhen I look back at old commits and branches that need to \npotentially be thrown away.  But I agree you may not want to \nshare these publicly.\n\nI am not sure what the right formula is, for when to share \nthese pointers publicly, but it seems like it might be that \nwhenever you push something, it should push along any \nreferences to amended commits that were publicly available \nalready.  I am not sure how to track that, but I suspect it \nis a subset of the union of commits you have fetched, and \ncommits you have pushed (i.e. you got it from elsewhere, or \nyou created it and already shared it with others)?  Maybe it \nshould also include any commits reachable by advertisements \nto places you are pushing to (in case it got shared some \nother way)?\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"335855","messageId":"3447055.jsE6nH3DQt@mfick-lnx","threadId":"47487","inReplyTo":"20171226011638.GA16552@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-01-04T19:54:00Z","receivedAt":"2018-01-04T19:54:06Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Monday, December 25, 2017 06:16:40 PM Carl Baldwin wrote:\n> On Sun, Dec 24, 2017 at 10:52:15PM -0500, Theodore Ts'o \nwrote:\n> Look at what happens in a rebase type workflow in any of\n> the following scenarios. All of these came up regularly\n> in my time with Gerrit.\n> \n>     1. Make a quick edit through the web UI then later\n> work on the change again in your local clone. It is easy\n> to forget to pull down the change made through the UI\n> before starting to work on it again. If that happens, the\n> change made through the UI will almost certainly be\n> clobbered.\n> \n>     2. You or someone else creates a second change that is\n> dependent on yours and works on it while yours is still\n> evolving. If the second change gets rebased with an older\n> copy of the base change and then posted back up for\n> review, newer work in the base change has just been\n> clobbered.\n> \n>     3. As a reviewer, you decide the best way to explain\n> how you'd like to see something done differently is to\n> make the quick change yourself and push it up. If the\n> author fails to fetch what you pushed before continuing\n> onto something else, it gets clobbered.\n> \n>     4. You want to collaborate on a single change with\n> someone else in any way and for whatever reason. As soon\n> as that change starts hitting multiple work spaces, there\n> are synchronization issues that currently take careful\n> manual intervention.\n\nThese scenarios seem to come up most for me at Gerrit hack-\na-thons where we collaborate a lot in short time spans on \nchanges.  We (the Gerrit maintainers) too have wanted and \nsometimes discussed ways to track the relation of \"amended\" \ncommits (which is generally what Gerrit patchsets are).  We \nalso concluded that some sort of parent commit pointer was \nneeded, although parent is somewhat the wrong term since \nthat already means something in git.  Rather, maybe some \n\"predecessor\" type of term would be better, maybe \n\"antecedent\", but \"amended-commit\" pointer might be best?\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"335858","messageId":"45915667.ye9M3CbBo5@mfick-lnx","threadId":"47487","inReplyTo":"20171226203153.GA21429@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-01-04T20:06:27Z","receivedAt":"2018-01-04T20:06:34Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"On Tuesday, December 26, 2017 01:31:55 PM Carl Baldwin \nwrote:\n...\n> What I propose is that gerrit and github could end up more\n> robust, featureful, and interoperable if they had this\n> feature to build from.\n\nI agree (assuming we come up with a well defined feature)\n\n> With gerrit specifically, adopting this feature would make\n> the \"change\" concept richer than it is now because it\n> could supersede the change-id in the commit message and\n> allow a change to evolve in a distributed non-linear way\n> with protection against clobbering work.\n\nWe (the Gerrit maintainers) would like changes to be able to \nevolve non-linearly so that we can eventually support \ndistributed Gerrit reviews, and the amended-commit pointer \nis one way I have thought to resolve this.\n\n> I have no intention to disparage either tool. I love them\n> both. They've both made my career better in different\n> ways. I know there is no guarantee that github, gerrit,\n> or any other tool will do anything to adopt this. But,\n> I'm hoping they are reading this thread and that they\n> recognize how this feature can make them a little bit\n> better and jump in and help. I know it is a lot to hope\n> for but I think it could be great if it happened.\n\nWe (the Gerrit maintainers) do recognize it, and I am glad \nthat someone is pushing for solutions in this space.  I am \nnot sure what the right solution is, and how to modify \nworkflows to deal better with this.  I do think that starting \nby making your local repo track pointers to amended-commits, \nlikely with various git hooks and notes (as also proposed by \nJohannes Schindelin), would be a good start.   With that in \nplace, then you can attack various specific workflows.\n\nIf you want to then attack the Gerrit workflow, it would be \ngood if you could prevent pushing new patchests that are \namended versions of patchsets that are out of date.  While \nit would be great if Gerrit could reject such pushes, I \nwonder if to start, git could detect and it prevent the push \nin this situation?  Could a git push hook analyze the ref \nadvertisements and figure this out (all the patchsets are in \nthe advertisement)?  Can a git hook look at the ref \nadvertisement?\n\n-Martin\n\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"335899","messageId":"5636043.i54Fe5jZ1b@mfick-lnx","threadId":"47487","inReplyTo":"2180514.YZ24uruNv3@mfick-lnx","subject":"Re: Bring together merge and rebase","fromName":"Martin Fick","fromEmail":"mfick@codeaurora.org","sentAt":"2018-01-05T00:31:39Z","receivedAt":"2018-01-05T00:31:49Z","isPatch":false,"sender":{"key":"mfick@codeaurora.org","avatar":null},"body":"> On Jan 4, 2018 11:19 AM, \"Martin Fick\" \n<mfick@codeaurora.org> wrote:\n> > On Tuesday, December 26, 2017 12:40:26 AM Jacob Keller\n> > \n> > wrote:\n> > > On Mon, Dec 25, 2017 at 10:02 PM, Carl Baldwin\n> > \n> > <carl@ecbaldwin.net> wrote:\n> > > >> On Mon, Dec 25, 2017 at 5:16 PM, Carl Baldwin\n> > \n> > <carl@ecbaldwin.net> wrote:\n> > > >> A bit of a tangent here, but a thought I didn't\n> > > >> wanna\n> > > >> lose: In the general case where a patch was rebased\n> > > >> and the original parent pointer was changed, it is\n> > > >> actually quite hard to show a diff of what changed\n> > > >> between versions.\n> > > \n> > > My biggest gripes are that the gerrit web interface\n> > > doesn't itself do something like this (and jgit does\n> > > not\n> > > appear to be able to generate combined diffs at all!)\n> > \n> > I believe it now does, a presentation was given at the\n> > Gerrit User summit in London describing this work.  It\n> > would indeed be great if git could do this also!\n\n\nOn Thursday, January 04, 2018 04:02:40 PM Jacob Keller \nwrote:\n> Any chance slides or a recording was posted anywhere? I'm\n> quite interested in this topic.\n\nSlides and video + transcript here:\n\nhttps://gerrit.googlesource.com/summit/2017/+/master/sessions/new-in-2.15.md\n\nWatch the part after the backend improvements,\n\n-Martin\n\n-- \nThe Qualcomm Innovation Center, Inc. is a member of Code \nAurora Forum, hosted by The Linux Foundation\n\n"},{"id":"335903","messageId":"20180105040837.GA12861@Carl-MBP.local","threadId":"47487","inReplyTo":"3447055.jsE6nH3DQt@mfick-lnx","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2018-01-05T04:08:38Z","receivedAt":"2018-01-05T04:08:43Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Thu, Jan 04, 2018 at 12:54:00PM -0700, Martin Fick wrote:\n> On Monday, December 25, 2017 06:16:40 PM Carl Baldwin wrote:\n> > On Sun, Dec 24, 2017 at 10:52:15PM -0500, Theodore Ts'o \n> wrote:\n> > Look at what happens in a rebase type workflow in any of\n> > the following scenarios. All of these came up regularly\n> > in my time with Gerrit.\n> > \n> >     1. Make a quick edit through the web UI then later\n> > work on the change again in your local clone. It is easy\n> > to forget to pull down the change made through the UI\n> > before starting to work on it again. If that happens, the\n> > change made through the UI will almost certainly be\n> > clobbered.\n> > \n> >     2. You or someone else creates a second change that is\n> > dependent on yours and works on it while yours is still\n> > evolving. If the second change gets rebased with an older\n> > copy of the base change and then posted back up for\n> > review, newer work in the base change has just been\n> > clobbered.\n> > \n> >     3. As a reviewer, you decide the best way to explain\n> > how you'd like to see something done differently is to\n> > make the quick change yourself and push it up. If the\n> > author fails to fetch what you pushed before continuing\n> > onto something else, it gets clobbered.\n> > \n> >     4. You want to collaborate on a single change with\n> > someone else in any way and for whatever reason. As soon\n> > as that change starts hitting multiple work spaces, there\n> > are synchronization issues that currently take careful\n> > manual intervention.\n> \n> These scenarios seem to come up most for me at Gerrit hack-\n> a-thons where we collaborate a lot in short time spans on \n> changes.  We (the Gerrit maintainers) too have wanted and \n> sometimes discussed ways to track the relation of \"amended\" \n> commits (which is generally what Gerrit patchsets are).  We \n> also concluded that some sort of parent commit pointer was \n> needed, although parent is somewhat the wrong term since \n> that already means something in git.  Rather, maybe some \n> \"predecessor\" type of term would be better, maybe \n> \"antecedent\", but \"amended-commit\" pointer might be best?\n\nI like \"replaces\" as I have proposed or \"supersedes\". \"predecessor\" also\nseems pretty good. I may add that to my list of favorites.\n\nCarl\n"},{"id":"335905","messageId":"20180105050647.GB12861@Carl-MBP.local","threadId":"47487","inReplyTo":"45915667.ye9M3CbBo5@mfick-lnx","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2018-01-05T05:06:50Z","receivedAt":"2018-01-05T05:06:58Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Thu, Jan 04, 2018 at 01:06:27PM -0700, Martin Fick wrote:\n> On Tuesday, December 26, 2017 01:31:55 PM Carl Baldwin \n> wrote:\n> ...\n> > What I propose is that gerrit and github could end up more\n> > robust, featureful, and interoperable if they had this\n> > feature to build from.\n> \n> I agree (assuming we come up with a well defined feature)\n> \n> > With gerrit specifically, adopting this feature would make\n> > the \"change\" concept richer than it is now because it\n> > could supersede the change-id in the commit message and\n> > allow a change to evolve in a distributed non-linear way\n> > with protection against clobbering work.\n> \n> We (the Gerrit maintainers) would like changes to be able to \n> evolve non-linearly so that we can eventually support \n> distributed Gerrit reviews, and the amended-commit pointer \n> is one way I have thought to resolve this.\n\nI really think that keeping these references is the key to doing this.\n\n> > I have no intention to disparage either tool. I love them\n> > both. They've both made my career better in different\n> > ways. I know there is no guarantee that github, gerrit,\n> > or any other tool will do anything to adopt this. But,\n> > I'm hoping they are reading this thread and that they\n> > recognize how this feature can make them a little bit\n> > better and jump in and help. I know it is a lot to hope\n> > for but I think it could be great if it happened.\n> \n> We (the Gerrit maintainers) do recognize it, and I am glad \n> that someone is pushing for solutions in this space.  I am \n> not sure what the right solution is, and how to modify \n> workflows to deal better with this.  I do think that starting \n> by making your local repo track pointers to amended-commits, \n> likely with various git hooks and notes (as also proposed by \n> Johannes Schindelin), would be a good start.   With that in \n> place, then you can attack various specific workflows.\n\nI have started a prototype that I will use to demonstrate this. I hope\nto have something in a couple of weeks. I do have a day job also, so it\nwill be slow going. One idea that I had was to put my own server with\nspecial hooks in it in front of gerrit to illustrate how collaboration\non a gerrit change, or even a chain of them, can be made safe. It would\nact as a middle man between my client and the gerrit server. I'd just\nhave to change remote reference on my client to demonstrate.\n\n> If you want to then attack the Gerrit workflow, it would be \n> good if you could prevent pushing new patchests that are \n> amended versions of patchsets that are out of date.  While \n> it would be great if Gerrit could reject such pushes, I \n> wonder if to start, git could detect and it prevent the push \n> in this situation?  Could a git push hook analyze the ref \n> advertisements and figure this out (all the patchsets are in \n> the advertisement)?  Can a git hook look at the ref \n> advertisement?\n\nI'll think about this. At the least, the hook would have to look at the\nserver to see if there are new revisions. It would be difficult to close\nrace conditions that occur because the client will always be using\npotentially out of date information even if it just went and pulled down\nthe latest stuff. I think I still like my middle man idea better as a\nshort term proof of concept.\n\nPreventing pushing amended/rebased versions of out of date changes is\nsimple. Follow the \"predecessor\" references until you hit a known\ncommit. If that commit is the latest revision of the change then it is\nup to date. If that commit not the latest revision, then it is out of\ndate. Reject it. This is what I plan to illustrate in my middle man\nserver.\n\nIf you traverse the entire graph of predecessors without finding a known\ncommit, then you have a new change. (In fact, the changeset id in the\ncommit message in a gerrit change seems unnecessary at this point). It\ngets a little more complicated when you think about combining/squashing\nchanges (resulting in two or more \"predecessor\" references from a single\ncommit) or dividing a change into multiple but it works.\n\nThe harder part is the push/pull interaction between client and server.\nWhen you go to push your amended update to a patchset, you want git to\nsend along any other new commits to complete the predecessor graph on\nthe server side. For example, you might rebase your commit and then\namend it to fix something. Personally, I'd like the rebase and the amend\nto both be kept separately.\n\nSimilarly, when you've just had a push rejected because you're out of\ndate, you want to be able to easily pull down the commits you're missing\nso that you can merge locally and try to push again.\n\nYou also don't want gc to garbage collect the intermediate commits. I\nthink gerrit uses many references internally in the git repo to \"pin\"\nolder revisions in the repository so that they don't appear orphaned. I\nthink I'm going to have to do something similar in my prototype.\n\nIf you think about it, this is all very much like what git already does\nwith its commit history and branches. If you stick to a strict\nbranch/merge model and don't rewrite commits, then it is unnecessary.\nHowever, for those that do rewrite commits (such as anyone using the\ngerrit workflow), this is a way to bring that power to them.\n\nI'd like to point out the RECOVERING FROM UPSTREAM REBASE section of the\ngit-rebase man page. If we have the graph of \"predecessor\" references\nfor a change, it could be used to automatically recover from the cases\ndescribed in this section much like regular branching and merging.\nRewriting changes would no longer be something to consider \"a bad idea\"\nfor these reasons.\n\nCarl\n"},{"id":"335906","messageId":"20180105050919.GA14525@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"2180514.YZ24uruNv3@mfick-lnx","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2018-01-05T05:09:20Z","receivedAt":"2018-01-05T05:09:28Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Thu, Jan 04, 2018 at 12:19:34PM -0700, Martin Fick wrote:\n> On Tuesday, December 26, 2017 12:40:26 AM Jacob Keller \n> wrote:\n> > On Mon, Dec 25, 2017 at 10:02 PM, Carl Baldwin \n> <carl@ecbaldwin.net> wrote:\n> > >> On Mon, Dec 25, 2017 at 5:16 PM, Carl Baldwin \n> <carl@ecbaldwin.net> wrote:\n> > >> A bit of a tangent here, but a thought I didn't wanna\n> > >> lose: In the general case where a patch was rebased\n> > >> and the original parent pointer was changed, it is\n> > >> actually quite hard to show a diff of what changed\n> > >> between versions.\n> > \n> > My biggest gripes are that the gerrit web interface\n> > doesn't itself do something like this (and jgit does not\n> > appear to be able to generate combined diffs at all!)\n> \n> I believe it now does, a presentation was given at the \n> Gerrit User summit in London describing this work.  It would \n> indeed be great if git could do this also!\n\nThis would be very cool. I've wanted to tackle this for a long time. I\nthink I even filed an issue with gerrit about this years ago.\n\nCarl\n"},{"id":"335908","messageId":"20180105052043.GB14525@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"20180105050919.GA14525@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2018-01-05T05:20:44Z","receivedAt":"2018-01-05T05:20:52Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Thu, Jan 04, 2018 at 10:09:19PM -0700, Carl Baldwin wrote:\n> This would be very cool. I've wanted to tackle this for a long time. I\n> think I even filed an issue with gerrit about this years ago.\n\nYep, it turned out that it was a duplicate but I described what I did to\nwork around it.\n\nhttps://bugs.chromium.org/p/gerrit/issues/detail?id=2375\n"},{"id":"335959","messageId":"xmqq4lo0cbbv.fsf@gitster.mtv.corp.google.com","threadId":"47487","inReplyTo":"3447055.jsE6nH3DQt@mfick-lnx","subject":"Re: Bring together merge and rebase","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-01-05T20:14:28Z","receivedAt":"2018-01-05T20:14:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Martin Fick <mfick@codeaurora.org> writes:\n\n> These scenarios seem to come up most for me at Gerrit hack-\n> a-thons where we collaborate a lot in short time spans on \n> changes.  We (the Gerrit maintainers) too have wanted and \n> sometimes discussed ways to track the relation of \"amended\" \n> commits (which is generally what Gerrit patchsets are).  We \n> also concluded that some sort of parent commit pointer was \n> needed, although parent is somewhat the wrong term since \n> that already means something in git.  Rather, maybe some \n> \"predecessor\" type of term would be better, maybe \n> \"antecedent\", but \"amended-commit\" pointer might be best?\n\nIn general, I agree that you would want richer set of \"relationship\"\nthan mere \"predecessor\" or \"related\", but I do not think \"amended\"\nis sufficient.  I certainly do not think a \"pointer\" embedded in a\ncommit object is a good idea, either (a new commit object header is\nout of question, but I doubt it is a good idea to make a pointer\nback to an existing commit as a part of the log message).\n\nYou may used to have a set of n-patches A1, A2, ..., An, that turned\ninto m-patches X1, X2, ..., Xm, after refactoring.  During the work,\nit may turned out that some things the original tried to do are not\nsensible and dropped, while some other things are added in the final.\nseries.  \n\nWhen n==m==1, \"amended\" pointer from X1 to A1 may allow you to\nanswer \"Is this the first attempt?  If this is refined, what did the\nearlier one look like?\" when given X1, but you would also want to\nanswer a related question \"This was a good start, but did the effort\nresult in a refined patch, and if so what is it?\" when given A1, and\n\"amended\" pointer won't help at all.  Needless to say, the \"pointer\"\napproach breaks down when !(n==m==1).\n\n"},{"id":"336057","messageId":"20180106172919.GA17272@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"xmqq4lo0cbbv.fsf@gitster.mtv.corp.google.com","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2018-01-06T17:29:21Z","receivedAt":"2018-01-06T17:29:31Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Fri, Jan 05, 2018 at 12:14:28PM -0800, Junio C Hamano wrote:\n> Martin Fick <mfick@codeaurora.org> writes:\n> \n> > These scenarios seem to come up most for me at Gerrit hack-\n> > a-thons where we collaborate a lot in short time spans on \n> > changes.  We (the Gerrit maintainers) too have wanted and \n> > sometimes discussed ways to track the relation of \"amended\" \n> > commits (which is generally what Gerrit patchsets are).  We \n> > also concluded that some sort of parent commit pointer was \n> > needed, although parent is somewhat the wrong term since \n> > that already means something in git.  Rather, maybe some \n> > \"predecessor\" type of term would be better, maybe \n> > \"antecedent\", but \"amended-commit\" pointer might be best?\n> \n> In general, I agree that you would want richer set of \"relationship\"\n> than mere \"predecessor\" or \"related\", but I do not think \"amended\"\n> is sufficient.  I certainly do not think a \"pointer\" embedded in a\n> commit object is a good idea, either (a new commit object header is\n\nTo me, this is roughly equivalent to saying that parent pointers\nembedded in a commit object is a good idea because we want a richer\nrelationship than mere \"parent\". Look how much we've done with this\nsimple relationship. Similarly, the new relationship that I'm proposing\nhandles much more than the simple m==n==1 case. Read below for more\ndetail.\n\n> out of question, but I doubt it is a good idea to make a pointer\n> back to an existing commit as a part of the log message).\n> \n> You may used to have a set of n-patches A1, A2, ..., An, that turned\n> into m-patches X1, X2, ..., Xm, after refactoring.  During the work,\n> it may turned out that some things the original tried to do are not\n> sensible and dropped, while some other things are added in the final.\n> series.  \n> \n> When n==m==1, \"amended\" pointer from X1 to A1 may allow you to\n> answer \"Is this the first attempt?  If this is refined, what did the\n> earlier one look like?\" when given X1, but you would also want to\n> answer a related question \"This was a good start, but did the effort\n> result in a refined patch, and if so what is it?\" when given A1, and\n> \"amended\" pointer won't help at all.  Needless to say, the \"pointer\"\n> approach breaks down when !(n==m==1).\n\nIt doesn't break down. It merely presents more sophisticated situations\nthat may be more work for the tool to help out with. This is where I\nthink a prototype will help see these situations and develop the tool to\nmanage them.\n\nWhen each of n commits is amended or rebased trivially into m==n new\ncommits then each change is represented by a distinct graph of\npredecessors that can be followed independently of others. With rebase,\nthis is accomplished by using only \"pick\" in interactive mode or not\nusing interactive mode at all (and no autosquash).\n\nThe more sophisticated cases can be broken down into two operations that\nchange the number of resulting commits.\n\n  1. Squashing two commits together (\"fixup\", \"squash\"). In this case,\n     the resulting commit will have two or more pointers. This clearly\n     shows that multiple changes converged into one at this point.\n\n  2. Splitting a single commit into multiple new commits (\"edit\"). In\n     this case, the graph shows multiple new commits pointing to the\n     same predecessor. In my experience, this is less common. It also is\n     a little more challenging to think about the tool managing\n     divergent work but I think it is possible.\n\nThe end result is m commits where m can be any positive number (even,\ncoincidentally, n). However, the graph of amended commits still tells\nthe story quite well. Even if commits are reordered, the graphs can\nstill be useful. The predecessor graph is independent of the parent\ngraph which makes up normal git commit history so it isn't inherently\nbad that the order of commits was changed.\n\nWe can dream up some very interesting graphs. Sure, as we do\nincreasingly more complicated history rewriting, it is going to be\nincreasingly more difficult for the tool to help out. I'm not really\ndeterred by this at this point. I want to experiment and work it out\nwith a prototype.\n\nMy primary objective personally is to detect where work on a single\nchange has diverged by working on it from more than one workspace\nwhether its multiple people chipping in or just me. Merely having the\nability to reject an update that clobbers divergent work is a big win.\nNo more silent corruption of work.\n\nMy secondary objective is to develop a tool to help get the divergent\nwork back on track. I believe that in the majority of common cases, this\ntool can be successful in either finding an automatic way to bring the\ndivergent work back into a new revision of the change or present the\nuser with conflicts to resolve that end up being much easier than what\nI've had to do in past experience with rebase workflows.\n\nCarl\n"},{"id":"336058","messageId":"20180106173225.GB17272@Carl-MBP.ecbaldwin.net","threadId":"47487","inReplyTo":"20180106172919.GA17272@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Carl Baldwin","fromEmail":"carl@ecbaldwin.net","sentAt":"2018-01-06T17:32:26Z","receivedAt":"2018-01-06T17:32:34Z","isPatch":false,"sender":{"key":"carl@ecbaldwin.net","avatar":null},"body":"On Sat, Jan 06, 2018 at 10:29:19AM -0700, Carl Baldwin wrote:\n> To me, this is roughly equivalent to saying that parent pointers\n> embedded in a commit object is a good idea because we want a richer\n> relationship than mere \"parent\". Look how much we've done with this\n> simple relationship. Similarly, the new relationship that I'm\n> proposing handles much more than the simple m==n==1 case. Read below\n> for more detail.\n\nOf course, I meant to say \"is not a good idea\" in the above paragraph.\nPlease pardon my error.\n"},{"id":"336069","messageId":"20180106213845.GD2404@thunk.org","threadId":"47487","inReplyTo":"20180106172919.GA17272@Carl-MBP.ecbaldwin.net","subject":"Re: Bring together merge and rebase","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2018-01-06T21:38:45Z","receivedAt":"2018-01-06T21:38:59Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Sat, Jan 06, 2018 at 10:29:21AM -0700, Carl Baldwin wrote:\n> > When n==m==1, \"amended\" pointer from X1 to A1 may allow you to\n> > answer \"Is this the first attempt?  If this is refined, what did the\n> > earlier one look like?\" when given X1, but you would also want to\n> > answer a related question \"This was a good start, but did the effort\n> > result in a refined patch, and if so what is it?\" when given A1, and\n> > \"amended\" pointer won't help at all.  Needless to say, the \"pointer\"\n> > approach breaks down when !(n==m==1).\n> \n> It doesn't break down. It merely presents more sophisticated situations\n> that may be more work for the tool to help out with. This is where I\n> think a prototype will help see these situations and develop the tool to\n> manage them.\n\nThat's another way of saying \"break down\".\n\nAnd if the goal is a prototype, may I gently suggest that the way\nforward is trailers in the commit body, ala:\n\n\tChange-Id: I0b793feac9664bcc8935d8ec04ca16d5\n\nor\n\n\tUpstream-4.15-SHA1: 73875fc2b3934e45b4b9a94eb57ca8cd\n\nMaking changes in the commit header is complex, and has all *sorts* of\nforward and backwards compatibility challenges, especially when it's\nnot clear what the proper data model should be.\n\nCheers,\n\n\t\t\t\t\t\t -Ted\n"}]}