{"thread":{"id":"12311","subject":"\"Contributors never merge\" and preserving history","startedAt":"2008-02-25T15:59:45Z","lastAt":"2008-02-26T16:41:10Z","messageCount":6,"participants":["John Goerzen","Linus Torvalds","Asheesh Laroia"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"69891","messageId":"slrnfs5pfh.lkc.jgoerzen@katherina.lan.complete.org","threadId":"12311","inReplyTo":null,"subject":"\"Contributors never merge\" and preserving history","fromName":"John Goerzen","fromEmail":"jgoerzen@complete.org","sentAt":"2008-02-25T15:59:45Z","receivedAt":"2008-02-25T15:59:45Z","isPatch":false,"sender":{"key":"jgoerzen@complete.org","avatar":null},"body":"Hi folks,\n\nI have a question about git philosophy.  Yesterday I was on #git\n(thanks to those of you that were there and very helpful).  This\ncomment was made: [1] \n\n<vmiklos> patch series are created by contributors while\n          contributors never merge\n\nNow, here's my question, posed in IRC and over at [2]:\n\n  Say we started from a common base where line 10 of file X said \"hi\", I\n  locally changed it to \"foo\", upstream changed it to \"bar\", and at\n  merge time I decide that we were both wrong and change it to \"baz\". I\n  don't want to lose the fact that I once had it at \"foo\", in case it\n  turns out later that really was the right decision.\n\nI understand that git projects that use a kernel-like development\nmodel -- which seems to be common -- do not want to know that I once\nhad it at \"foo\".  That is fine to me.  But I want a local branch -- or\na local *something* -- to store that, and make it convenient to\naccess.\n\nThat can be done easily enough if I use git pull instead of git rebase\nto pull in upstream updates.  I'll be committing a merge patch on\nevery upstream update.  I suppose I'm OK with that, though it's not ideal.\n\nThe canonical answer from #git seems to be \"never pull\", always use\nfetch and rebase when submitting patches upstream using\ngit-format-patch.\n\nI have tried various ways to somehow make rebase get along with\nmerging, but have failed on that, mainly because merging confuses\nrebase's algorithm that figures out which patches have already been\napplied.\n\nSo, the question is: what is the best way to be able to keep a full\nlocal history, while still using format-patch to interact easily with\nupstream?  I'm thinking the answer may involve a submission branch\nthat uses squashing, but I'm not entirely certain.\n\nAnd the corrolary: As upstream, how can I facilitate users that may\nwish to do this?\n\n[1] http://colabti.org/irclogger/irclogger_log/git?date=2008-02-25,Mon&sel=66#l99\n[2] http://changelog.complete.org/posts/690-Git-looks-really-nice,-until.....html\n"},{"id":"69920","messageId":"alpine.LFD.1.00.0802251202380.14934@woody.linux-foundation.org","threadId":"12311","inReplyTo":"slrnfs5pfh.lkc.jgoerzen@katherina.lan.complete.org","subject":"Re: \"Contributors never merge\" and preserving history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-02-25T20:31:26Z","receivedAt":"2008-02-25T20:31:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Feb 2008, John Goerzen wrote:\n> \n> The canonical answer from #git seems to be \"never pull\", always use\n> fetch and rebase when submitting patches upstream using\n> git-format-patch.\n\nIf you're going to submit them as patches, that is the correct answer, \nbecause what you want to do is basically keep your patch-queue up-to-date \n(which is exactly what \"git rebase\" does).\n\nAnd no, you must never mix merging and rebasing - not because it's \ntechnically impossible (it *can* make sense in some circumstances), but \nbecause the two flows really are very different. Either you keep history \nand submit it as such (git-to-git merges up and down the chain) or you \nwork with a \"set of patches\" model. Mixing the two in the same tree just \nleads to insanity (although mixing the two int he same project among \ndifferent repositories can be a very good way to handle things).\n\nBut \"never pull\" isn't quite true either. \n\nBasically, the way to think about development is to try to keep things in \n\"topic branches\", which git is really good at. And the rule really \nshouldn' be \"never pull\" as much as \"try to keep those topic-branches \nseparate\".\n\nSo pulling generally by definition mixes different branches (that's the \nmerge part) in a way that a rebase does not. That's *especially* true \nabout pulling from \"upstream\", because - pretty much by definition - that \nupstream is generally not even a well-defined topic branch that you want \nto merge, but simply the sum of all the *other* topic branches that have \nbeen merged upstream.\n\nSo the reason you should generally pull from downstream rather than \nupstream is that it keeps your development branch \"focused\" or \"on target\" \nor whatever you want to call it. And that's always a good idea, because \nnow anybody who works together with you knows what he is getting.\n\nSo think of it as a cleanliness issue - it may not matter all that much if \nthe only person you expect to ever pull your tree is always that same \nupstream (so even if you pull from your upstream, your upstream isn't \nreally getting any mixed-up new code when he in turn pulls from you), but \none thign I personally hoped for as a \"design\" was that what git really \nallows you to do (and _should_ encourage) is to have a less than perfectly \nhierarchical development stream.\n\nAnd yes, in practice, pretty much every project ends up being pretty \nhierarchical after all, and that may be because of how people work (they \nwant clarity, they want a simple \"which tree is in charge\" kind of model), \nbut I still suspect it's at least partly - and perhaps mostly - simply due \nto historical patterns that it's just really hard to break.\n\nSo if you think of different git repositories as different branches (and \nthat's what they really are!), then the \"avoid pulling from upstream\" is \nreally about that \"keep the topic branch focused and clean!\".\n\nAnd quite frankly, as \"the upstream\" for the kernel, I really appreciate \npeople who ask me to pull, and that keep their histories clean, so that \nwhen I do a \"gitk ORIG_HEAD..\" after a pull, I get something that just \nlooks real and not too messy. IOW, I like seeing myself pulling clear and \nwell-defined topic branches (even if the \"topics\" I pull tend to be pretty \nbig-picture topics, ie they may encompass \"everything networking\" or \nsimilar).\n\nBUT!\n\nThere's always a but. In some cases, I will literally _ask_ a downstream \nperson to pull from me. Havign the downstream doing a merge makes sense if \nthe merge is non-trivial, and then the conflicting changes in a topic \nbranch should generally be resolved by the side that has (a) the knowledge \nto do so (obviously) but also (b) the one who has the more specific \nchanges (ie the side that has less work, and more targeted knowledge, of \nthe things that conflict).\n\nIOW, in the case of non-trivial conflicts, suddenly downstream is usually \nthe one that has more knowledge, and now they should do the merge. And \nquite often, downstream may well know that ahead of time and be proactive, \nand just do the pull the \"wrong way\" and when asking me as an upstream \nmember to merge, they'll let me know that they've already resolved the \nconflicts with what was in my tree.\n\nThe latter case is also something where a really long-lived topic branch \nmay simply be doing those pulls over time every once in a while to just \nmake sure that the topic branch never gets *too* far out of sync. However, \nif that's the reason for doing a merge, I'd almost suggest not just doing \na \"git pull\" from upstream, but fetching and then merging at \nwell-specified points (ie releases or release candidates).\n\nThat way you can also make a better and more useful merge message: not \njust \"merged with upstream\", but actually make it be \"Synchronized with \nrelease v1.7.9\". Which now makes a whole lot of conceptual sense at a \nhigher level.\n\nTo recap:\n\n - from a purely technical sense it doesn't make any difference \n   what-so-ever who pulls and who doesn't, although you don't want it to \n   be *too* rare so that the different branches diverge so far as to make \n   it technically hard to synchronize later!\n\n - from a cleanliness angle - and *especially* if you want to work not \n   just in strict \"upstream\" and \"downstream\" patterns, but expect to \n   maybe have multiple upstreams (think \"stable branch\" or \"vendor X wants \n   to pull this too\"), a clean and clear topic banch is a really really \n   good idea.\n\n   For example, let's say that you're developing a driver. If you start at \n   some specific kernel version (say, 2.6.24) and you do *not* generally \n   merge from my development tree, now suddenly other people can happily \n   pull from your tree to get the driver, even if they are stable kernels \n   or vendor kernels that don't want all the development crud that is in \n   my tree!\n\n   See? Keeping a clean history actually makes your tree more useful!\n\n - But there are cases where pulling from up-stream really makes sense. \n   There may be specific points at upstream that you simply want to \n   synchronize with, or there may be conflicts that you want to resolve \n   simply because others aren't as knowledgeable about your topic branch\n   etc.\n\nSo don't believe in \"never pull from upstream\". But *do* believe in \"try \nto keep your branches on topic, because it will make everybody happier, \nand you'll be more easily able to read your own history too!\".\n\n\t\t\tLinus\n"},{"id":"69927","messageId":"alpine.DEB.1.00.0802251330340.28694@dell.linuxdev.us.dell.com","threadId":"12311","inReplyTo":"alpine.LFD.1.00.0802251202380.14934@woody.linux-foundation.org","subject":"Re: \"Contributors never merge\" and preserving history","fromName":"Asheesh Laroia","fromEmail":"asheesh@asheesh.org","sentAt":"2008-02-25T21:35:36Z","receivedAt":"2008-02-25T21:35:36Z","isPatch":false,"sender":{"key":"asheesh@asheesh.org","avatar":"https://avatars.githubusercontent.com/u/25457?v=4"},"body":"On Mon, 25 Feb 2008, Linus Torvalds wrote:\n\n>   For example, let's say that you're developing a driver. If you start at\n>   some specific kernel version (say, 2.6.24) and you do *not* generally\n>   merge from my development tree, now suddenly other people can happily\n>   pull from your tree to get the driver, even if they are stable kernels\n>   or vendor kernels that don't want all the development crud that is in\n>   my tree!\n>\n>   See? Keeping a clean history actually makes your tree more useful!\n\nI'm going to chime in on this thread as a relative newcomer to git.  If \nI'm developing a driver or other feature branch, and then a new upstream \nrelease comes along, I can't rebase and push - that would make the \"is not \na strict subset of local ref\" complaint.\n\nIs the right workflow, then, to rebase against 2.6.25 in a new local \nbranch, and push that to a new remote branch for others (like you say, \nvendor kernel maintainers) to pull from?\n\nThanks!\n\n-- Asheesh.\n\n-- \nWho will take care of the world after you're gone?\n"},{"id":"69939","messageId":"alpine.LFD.1.00.0802251347530.14934@woody.linux-foundation.org","threadId":"12311","inReplyTo":"alpine.DEB.1.00.0802251330340.28694@dell.linuxdev.us.dell.com","subject":"Re: \"Contributors never merge\" and preserving history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-02-25T22:02:11Z","receivedAt":"2008-02-25T22:02:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 25 Feb 2008, Asheesh Laroia wrote:\n>\n> On Mon, 25 Feb 2008, Linus Torvalds wrote:\n> > \n> >   For example, let's say that you're developing a driver. If you start at\n> >   some specific kernel version (say, 2.6.24) and you do *not* generally\n> >   merge from my development tree, now suddenly other people can happily\n> >   pull from your tree to get the driver, even if they are stable kernels\n> >   or vendor kernels that don't want all the development crud that is in\n> >   my tree!\n> > \n> >   See? Keeping a clean history actually makes your tree more useful!\n> \n> I'm going to chime in on this thread as a relative newcomer to git.  If I'm\n> developing a driver or other feature branch, and then a new upstream release\n> comes along, I can't rebase and push - that would make the \"is not a strict\n> subset of local ref\" complaint.\n> \n> Is the right workflow, then, to rebase against 2.6.25 in a new local branch,\n> and push that to a new remote branch for others (like you say, vendor kernel\n> maintainers) to pull from?\n\nAlmost always, the right workflow is to *neither* rebase *nor* pull.\n\nQuite frankly, if you're working on some new feature like a driver, then \nin most projects you shouldn't need to care all that deeply about what is \ngoing on in other drivers etc. Merging or rebasing is just going to be a \ndistraction, and open you up to new bugs that aren't even in your code!\n\nOf course, this kind of situation can certainly be taken too far. The \ninfrastructure may be changing, and you may want to rebase for that \nreason, but in most projects that kind of churn is (a) generally kept to a \nminimum and (b) shouldn't really necessarily be your headache anyway (ie \nI'm happy to handle merge conflicts even if I can't always _test_ them, \nand if other maintainers make changes to infrastructure they are also \nsupposed to end up helping fix up the fallout!).\n\nSo I would not in general suggest that a driver writer maintain multiple \nbranches based on different versions. That's just not worth your time, I \nthink. You'd be better off staying back on whatever version you're \ncomfortable with, and then perhaps rebasing or merging very occasionally \nwhen you start thinking that your base is simply too old to be relevant.\n\nIOW, I'd suggest not merging or rebasing more than maybe once a month, if \neven that, unless you happen to be very bleeding edge (ie the \ninfrastructure you depend on may itself be developing quickly, so a \nwireless driver in the kernel would generally see more need for being kept \nup-to-date than a random other driver).\n\nBut hey, it's also a matter of your personal taste, and how you work. \nThere really are different models:\n\n - the \"rebase\" model is more amenable to a daily \"fetch+rebase\" kind of \n   ritual, and if you're the kind of person who really wants to feel that \n   you're always on the bleeding edge, maybe that's the right model for \n   you, even though I actually don't think it's really a logically very \n   good model (ie if you are actively doing development, you really don't \n   need the distraction!)\n\n - if you're working with somebody else (or a group), pulling from \n   *each*other* may be a great thing to do, and may well be the right \n   approach. But if you start doing that, then everybody involved should \n   avoid rebasing or pulling from upstream, because otherwise you'll just \n   get either tons of duplicate history (which wil *really* mess up \n   debugging: things like \"git bisect\" will work much worse if you have \n   the same bug introduced in multiple places etc)\n\n - the optimal strategy if you're just buffered enough is likely to not \n   rebase and not merge, and just work on your own thing, and then when \n   you feel ready, just say \"please pull\" to upstream.\n\nThe last one may feel a bit boring and staid, but I think it's the best \none when it works (and \"when it works\" is a lot about _your_ psychology \ntoo: some people like that kind of insulation where they don't have to \nworry about what everybody else is doing, while others hate feeling like \nthey're working on a tree that is a week old).\n\n\t\t\tLinus\n"},{"id":"69994","messageId":"slrnfs8749.prc.jgoerzen@katherina.lan.complete.org","threadId":"12311","inReplyTo":"alpine.LFD.1.00.0802251202380.14934@woody.linux-foundation.org","subject":"Re: \"Contributors never merge\" and preserving history","fromName":"John Goerzen","fromEmail":"jgoerzen@complete.org","sentAt":"2008-02-26T14:04:57Z","receivedAt":"2008-02-26T14:04:57Z","isPatch":false,"sender":{"key":"jgoerzen@complete.org","avatar":null},"body":"On 2008-02-25, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n\n[ snip ]\n\n> So the reason you should generally pull from downstream rather than \n> upstream is that it keeps your development branch \"focused\" or \"on target\" \n> or whatever you want to call it. And that's always a good idea, because \n> now anybody who works together with you knows what he is getting.\n\nHi Linus,\n\nThank you very much for these two informative messages.  I think that\nthere were a lot of shades of gray to the kernel workflow that I failed\nto appreciate before, for whatever reason.\n\nI do have a question about the point you make above though.  I'm not\nquite understanding what you're saying here.  Technically speaking,\nthe end result of a merge where you pulled from me would be identical\nto a merge where I pulled from you.  Moreover, say I'm pretty far down\non the seniority list, kernel-wise.  Do you expect subsystem\nmaintainers to honor a request from me to pull from my tree, even if\nthey've never heard of me before, or would you think they'd only want\ngit format-patch output?\n\nI ask because let's say I follow that advice above, and there are some\n\"downstreams\" to me.  I pull from them, which involves some merging,\nand then I want to format-patch.  It seems format-patch doesn't work\nso well with merging.  What would I then do in this situation?  Should\nI just use rebase to merge unless I know for sure that upstream will\nhonor a pull request?  But then again, we get into trouble if one of\nmy downstreams did a merge.\n\n-- John\n"},{"id":"70006","messageId":"alpine.LFD.1.00.0802260831220.14934@woody.linux-foundation.org","threadId":"12311","inReplyTo":"slrnfs8749.prc.jgoerzen@katherina.lan.complete.org","subject":"Re: \"Contributors never merge\" and preserving history","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-02-26T16:41:10Z","receivedAt":"2008-02-26T16:41:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 26 Feb 2008, John Goerzen wrote:\n> \n> I do have a question about the point you make above though.  I'm not\n> quite understanding what you're saying here.  Technically speaking,\n> the end result of a merge where you pulled from me would be identical\n> to a merge where I pulled from you.\n\nYes. Except for where the end result is!\n\nThat's kind of the point. If you're developing a driver, your tree is the \n\"driver tree\". But if you keep pulling from me, now it's no longer a \ndriver tree, it's a \"driver and Linus' code tree\".\n\n> Moreover, say I'm pretty far down on the seniority list, kernel-wise.  \n> Do you expect subsystem maintainers to honor a request from me to pull \n> from my tree, even if they've never heard of me before, or would you \n> think they'd only want git format-patch output?\n\nIt probably depends on the submaintainer. But you're absolutely right that \nat least early on, most of them will want just emailed patches. And for \nthat, the \"fetch + rebase\" model is the better one.\n\nHOWEVER.\n\nWhat happens with me is that I personally prefer patches from people if\n\n - they are \"single\" patches at a time (not necessarily just one, but at \n   most a couple at a time)\n\n - I've really never worked with you before\n\nbut if you have a real patch-series with more than (say) 4-5 patches, and \nI've seen patches from you before, _and_ you have a clean git tree, at \nthat point I'd more likely actually already prefer a git pull if you can \njust describe your patches well enough in the email.\n\n[ Of course, when it comes to me personally, another big requirement is \n  that I don't feel like you're going past some subsystem maintainer.\n\n  It's not that I am a strict hierarchical person, it's that when it comes \n  to most subsystems I often don't feel competent enough to make the \n  decision, so I want things to go through submaintainers simply because \n  it's an extra layer of \"filters\".\n\n  So the things I'd take through git are things that are either really \n  obvious or things I'd take anyway for other reasons. Those things are \n  seldom \"patch series\", but it happens.. ]\n\nSo at least judging by my own preferences, I don't think the barrier to \ndoing a git merge is actually all that high. The *biggest* barrier may \nindeed be that I can happily do a \"git pull\", but if it doesn't look like \na clean topic branch, I'd probably undo it.\n\n\t\t\tLinus\n"}]}