threads / discuss / 12311

"Contributors never merge" and preserving history

Subject: "Contributors never merge" and preserving history

## tl;dr

6 messages between Feb 25, 2008 and Feb 26, 2008.

replies: 5people: 3as markdown or json

John Goerzen· Feb 25, 2008, 15:59 UTC · lore
Hi folks,

I have a question about git philosophy. Yesterday I was on #git (thanks to those of you that were there and very helpful). This comment was made: [1]

<vmiklos> patch series are created by contributors while
          contributors never merge
Now, here's my question, posed in IRC and over at [2]:
  Say we started from a common base where line 10 of file X said "hi", I
  locally changed it to "foo", upstream changed it to "bar", and at
  merge time I decide that we were both wrong and change it to "baz". I
  don't want to lose the fact that I once had it at "foo", in case it
  turns out later that really was the right decision.

I understand that git projects that use a kernel-like development model -- which seems to be common -- do not want to know that I once had it at "foo". That is fine to me. But I want a local branch -- or a local *something* -- to store that, and make it convenient to access.

That can be done easily enough if I use git pull instead of git rebase to pull in upstream updates. I'll be committing a merge patch on every upstream update. I suppose I'm OK with that, though it's not ideal.

The canonical answer from #git seems to be "never pull", always use fetch and rebase when submitting patches upstream using git-format-patch.

I have tried various ways to somehow make rebase get along with merging, but have failed on that, mainly because merging confuses rebase's algorithm that figures out which patches have already been applied.

So, the question is: what is the best way to be able to keep a full local history, while still using format-patch to interact easily with upstream? I'm thinking the answer may involve a submission branch that uses squashing, but I'm not entirely certain.

And the corrolary: As upstream, how can I facilitate users that may wish to do this?

[1] http://colabti.org/irclogger/irclogger_log/git?date=2008-02-25,Mon&sel=66#l99 [2] http://changelog.complete.org/posts/690-Git-looks-really-nice,-until.....html

Linus Torvalds· Feb 25, 2008, 20:31 UTC · re: John Goerzen · lore

Re: "Contributors never merge" and preserving history

On Mon, 25 Feb 2008, John Goerzen wrote:
> 
> The canonical answer from #git seems to be "never pull", always use
> fetch and rebase when submitting patches upstream using
> git-format-patch.

If you're going to submit them as patches, that is the correct answer, because what you want to do is basically keep your patch-queue up-to-date (which is exactly what "git rebase" does).

And no, you must never mix merging and rebasing - not because it's technically impossible (it *can* make sense in some circumstances), but because the two flows really are very different. Either you keep history and submit it as such (git-to-git merges up and down the chain) or you work with a "set of patches" model. Mixing the two in the same tree just leads to insanity (although mixing the two int he same project among different repositories can be a very good way to handle things).

But "never pull" isn't quite true either. 

Basically, the way to think about development is to try to keep things in "topic branches", which git is really good at. And the rule really shouldn' be "never pull" as much as "try to keep those topic-branches separate".

So pulling generally by definition mixes different branches (that's the merge part) in a way that a rebase does not. That's *especially* true about pulling from "upstream", because - pretty much by definition - that upstream is generally not even a well-defined topic branch that you want to merge, but simply the sum of all the *other* topic branches that have been merged upstream.

So the reason you should generally pull from downstream rather than upstream is that it keeps your development branch "focused" or "on target" or whatever you want to call it. And that's always a good idea, because now anybody who works together with you knows what he is getting.

So think of it as a cleanliness issue - it may not matter all that much if the only person you expect to ever pull your tree is always that same upstream (so even if you pull from your upstream, your upstream isn't really getting any mixed-up new code when he in turn pulls from you), but one thign I personally hoped for as a "design" was that what git really allows you to do (and _should_ encourage) is to have a less than perfectly hierarchical development stream.

And yes, in practice, pretty much every project ends up being pretty hierarchical after all, and that may be because of how people work (they want clarity, they want a simple "which tree is in charge" kind of model), but I still suspect it's at least partly - and perhaps mostly - simply due to historical patterns that it's just really hard to break.

So if you think of different git repositories as different branches (and that's what they really are!), then the "avoid pulling from upstream" is really about that "keep the topic branch focused and clean!".

And quite frankly, as "the upstream" for the kernel, I really appreciate people who ask me to pull, and that keep their histories clean, so that when I do a "gitk ORIG_HEAD.." after a pull, I get something that just looks real and not too messy. IOW, I like seeing myself pulling clear and well-defined topic branches (even if the "topics" I pull tend to be pretty big-picture topics, ie they may encompass "everything networking" or similar).

BUT!

There's always a but. In some cases, I will literally _ask_ a downstream person to pull from me. Havign the downstream doing a merge makes sense if the merge is non-trivial, and then the conflicting changes in a topic branch should generally be resolved by the side that has (a) the knowledge to do so (obviously) but also (b) the one who has the more specific changes (ie the side that has less work, and more targeted knowledge, of the things that conflict).

IOW, in the case of non-trivial conflicts, suddenly downstream is usually the one that has more knowledge, and now they should do the merge. And quite often, downstream may well know that ahead of time and be proactive, and just do the pull the "wrong way" and when asking me as an upstream member to merge, they'll let me know that they've already resolved the conflicts with what was in my tree.

The latter case is also something where a really long-lived topic branch may simply be doing those pulls over time every once in a while to just make sure that the topic branch never gets *too* far out of sync. However, if that's the reason for doing a merge, I'd almost suggest not just doing a "git pull" from upstream, but fetching and then merging at well-specified points (ie releases or release candidates).

That way you can also make a better and more useful merge message: not just "merged with upstream", but actually make it be "Synchronized with release v1.7.9". Which now makes a whole lot of conceptual sense at a higher level.

To recap:
 - from a purely technical sense it doesn't make any difference 
   what-so-ever who pulls and who doesn't, although you don't want it to 
   be *too* rare so that the different branches diverge so far as to make 
   it technically hard to synchronize later!
 - from a cleanliness angle - and *especially* if you want to work not 
   just in strict "upstream" and "downstream" patterns, but expect to 
   maybe have multiple upstreams (think "stable branch" or "vendor X wants 
   to pull this too"), a clean and clear topic banch is a really really 
   good idea.
   For example, let's say that you're developing a driver. If you start at 
   some specific kernel version (say, 2.6.24) and you do *not* generally 
   merge from my development tree, now suddenly other people can happily 
   pull from your tree to get the driver, even if they are stable kernels 
   or vendor kernels that don't want all the development crud that is in 
   my tree!
   See? Keeping a clean history actually makes your tree more useful!
 - But there are cases where pulling from up-stream really makes sense. 
   There may be specific points at upstream that you simply want to 
   synchronize with, or there may be conflicts that you want to resolve 
   simply because others aren't as knowledgeable about your topic branch
   etc.

So don't believe in "never pull from upstream". But *do* believe in "try to keep your branches on topic, because it will make everybody happier, and you'll be more easily able to read your own history too!".

			Linus
Asheesh Laroia· Feb 25, 2008, 21:35 UTC · re: Linus Torvalds · lore

Re: "Contributors never merge" and preserving history

On Mon, 25 Feb 2008, Linus Torvalds wrote:
Show 8 quoted lines
>   For example, let's say that you're developing a driver. If you start at
>   some specific kernel version (say, 2.6.24) and you do *not* generally
>   merge from my development tree, now suddenly other people can happily
>   pull from your tree to get the driver, even if they are stable kernels
>   or vendor kernels that don't want all the development crud that is in
>   my tree!
>
>   See? Keeping a clean history actually makes your tree more useful!

I'm going to chime in on this thread as a relative newcomer to git. If I'm developing a driver or other feature branch, and then a new upstream release comes along, I can't rebase and push - that would make the "is not a strict subset of local ref" complaint.

Is the right workflow, then, to rebase against 2.6.25 in a new local branch, and push that to a new remote branch for others (like you say, vendor kernel maintainers) to pull from?

Thanks!
-- Asheesh.
-- 
Who will take care of the world after you're gone?
Linus Torvalds· Feb 25, 2008, 22:02 UTC · re: Asheesh Laroia · lore

Re: "Contributors never merge" and preserving history

On Mon, 25 Feb 2008, Asheesh Laroia wrote:
Show 20 quoted lines
>
> On Mon, 25 Feb 2008, Linus Torvalds wrote:
> > 
> >   For example, let's say that you're developing a driver. If you start at
> >   some specific kernel version (say, 2.6.24) and you do *not* generally
> >   merge from my development tree, now suddenly other people can happily
> >   pull from your tree to get the driver, even if they are stable kernels
> >   or vendor kernels that don't want all the development crud that is in
> >   my tree!
> > 
> >   See? Keeping a clean history actually makes your tree more useful!
> 
> I'm going to chime in on this thread as a relative newcomer to git.  If I'm
> developing a driver or other feature branch, and then a new upstream release
> comes along, I can't rebase and push - that would make the "is not a strict
> subset of local ref" complaint.
> 
> Is the right workflow, then, to rebase against 2.6.25 in a new local branch,
> and push that to a new remote branch for others (like you say, vendor kernel
> maintainers) to pull from?
Almost always, the right workflow is to *neither* rebase *nor* pull.

Quite frankly, if you're working on some new feature like a driver, then in most projects you shouldn't need to care all that deeply about what is going on in other drivers etc. Merging or rebasing is just going to be a distraction, and open you up to new bugs that aren't even in your code!

Of course, this kind of situation can certainly be taken too far. The infrastructure may be changing, and you may want to rebase for that reason, but in most projects that kind of churn is (a) generally kept to a minimum and (b) shouldn't really necessarily be your headache anyway (ie I'm happy to handle merge conflicts even if I can't always _test_ them, and if other maintainers make changes to infrastructure they are also supposed to end up helping fix up the fallout!).

So I would not in general suggest that a driver writer maintain multiple branches based on different versions. That's just not worth your time, I think. You'd be better off staying back on whatever version you're comfortable with, and then perhaps rebasing or merging very occasionally when you start thinking that your base is simply too old to be relevant.

IOW, I'd suggest not merging or rebasing more than maybe once a month, if even that, unless you happen to be very bleeding edge (ie the infrastructure you depend on may itself be developing quickly, so a wireless driver in the kernel would generally see more need for being kept up-to-date than a random other driver).

But hey, it's also a matter of your personal taste, and how you work. There really are different models:

 - the "rebase" model is more amenable to a daily "fetch+rebase" kind of 
   ritual, and if you're the kind of person who really wants to feel that 
   you're always on the bleeding edge, maybe that's the right model for 
   you, even though I actually don't think it's really a logically very 
   good model (ie if you are actively doing development, you really don't 
   need the distraction!)
 - if you're working with somebody else (or a group), pulling from 
   *each*other* may be a great thing to do, and may well be the right 
   approach. But if you start doing that, then everybody involved should 
   avoid rebasing or pulling from upstream, because otherwise you'll just 
   get either tons of duplicate history (which wil *really* mess up 
   debugging: things like "git bisect" will work much worse if you have 
   the same bug introduced in multiple places etc)
 - the optimal strategy if you're just buffered enough is likely to not 
   rebase and not merge, and just work on your own thing, and then when 
   you feel ready, just say "please pull" to upstream.

The last one may feel a bit boring and staid, but I think it's the best one when it works (and "when it works" is a lot about _your_ psychology too: some people like that kind of insulation where they don't have to worry about what everybody else is doing, while others hate feeling like they're working on a tree that is a week old).

			Linus
John Goerzen· Feb 26, 2008, 14:04 UTC · re: Linus Torvalds · lore

Re: "Contributors never merge" and preserving history

On 2008-02-25, Linus Torvalds <torvalds@linux-foundation.org> wrote:
[ snip ]
> So the reason you should generally pull from downstream rather than 
> upstream is that it keeps your development branch "focused" or "on target" 
> or whatever you want to call it. And that's always a good idea, because 
> now anybody who works together with you knows what he is getting.
Hi Linus,

Thank you very much for these two informative messages. I think that there were a lot of shades of gray to the kernel workflow that I failed to appreciate before, for whatever reason.

I do have a question about the point you make above though. I'm not quite understanding what you're saying here. Technically speaking, the end result of a merge where you pulled from me would be identical to a merge where I pulled from you. Moreover, say I'm pretty far down on the seniority list, kernel-wise. Do you expect subsystem maintainers to honor a request from me to pull from my tree, even if they've never heard of me before, or would you think they'd only want git format-patch output?

I ask because let's say I follow that advice above, and there are some "downstreams" to me. I pull from them, which involves some merging, and then I want to format-patch. It seems format-patch doesn't work so well with merging. What would I then do in this situation? Should I just use rebase to merge unless I know for sure that upstream will honor a pull request? But then again, we get into trouble if one of my downstreams did a merge.

-- John
Linus Torvalds· Feb 26, 2008, 16:41 UTC · re: John Goerzen · lore

Re: "Contributors never merge" and preserving history

On Tue, 26 Feb 2008, John Goerzen wrote:
Show 5 quoted lines
> 
> I do have a question about the point you make above though.  I'm not
> quite understanding what you're saying here.  Technically speaking,
> the end result of a merge where you pulled from me would be identical
> to a merge where I pulled from you.
Yes. Except for where the end result is!

That's kind of the point. If you're developing a driver, your tree is the "driver tree". But if you keep pulling from me, now it's no longer a driver tree, it's a "driver and Linus' code tree".

> Moreover, say I'm pretty far down on the seniority list, kernel-wise.  
> Do you expect subsystem maintainers to honor a request from me to pull 
> from my tree, even if they've never heard of me before, or would you 
> think they'd only want git format-patch output?

It probably depends on the submaintainer. But you're absolutely right that at least early on, most of them will want just emailed patches. And for that, the "fetch + rebase" model is the better one.

HOWEVER.
What happens with me is that I personally prefer patches from people if
 - they are "single" patches at a time (not necessarily just one, but at 
   most a couple at a time)
 - I've really never worked with you before

but if you have a real patch-series with more than (say) 4-5 patches, and I've seen patches from you before, _and_ you have a clean git tree, at that point I'd more likely actually already prefer a git pull if you can just describe your patches well enough in the email.

[ Of course, when it comes to me personally, another big requirement is 
  that I don't feel like you're going past some subsystem maintainer.
  It's not that I am a strict hierarchical person, it's that when it comes 
  to most subsystems I often don't feel competent enough to make the 
  decision, so I want things to go through submaintainers simply because 
  it's an extra layer of "filters".
  So the things I'd take through git are things that are either really 
  obvious or things I'd take anyway for other reasons. Those things are 
  seldom "patch series", but it happens.. ]

So at least judging by my own preferences, I don't think the barrier to doing a git merge is actually all that high. The *biggest* barrier may indeed be that I can happily do a "git pull", but if it doesn't look like a clean topic branch, I'd probably undo it.

			Linus

← back to recent threads