# "Contributors never merge" and preserving history

6 messages from 2008-02-25 to 2008-02-26. Participants: John Goerzen, Linus Torvalds, Asheesh Laroia.
Thread: https://gitlist.dev/t/12311

## John Goerzen, 2008-02-25 15:59

Subject: "Contributors never merge" and preserving history
Message-ID: <slrnfs5pfh.lkc.jgoerzen@katherina.lan.complete.org>
URL: https://gitlist.dev/e/slrnfs5pfh.lkc.jgoerzen%40katherina.lan.complete.org

```
Hi folks,

I have a question about git philosophy.  Yesterday I was on #git
(thanks to those of you that were there and very helpful).  This
comment was made: [1] 

<vmiklos> patch series are created by contributors while
          contributors never merge

Now, here's my question, posed in IRC and over at [2]:

  Say we started from a common base where line 10 of file X said "hi", I
  locally changed it to "foo", upstream changed it to "bar", and at
  merge time I decide that we were both wrong and change it to "baz". I
  don't want to lose the fact that I once had it at "foo", in case it
  turns out later that really was the right decision.

I understand that git projects that use a kernel-like development
model -- which seems to be common -- do not want to know that I once
had it at "foo".  That is fine to me.  But I want a local branch -- or
a local *something* -- to store that, and make it convenient to
access.

That can be done easily enough if I use git pull instead of git rebase
to pull in upstream updates.  I'll be committing a merge patch on
every upstream update.  I suppose I'm OK with that, though it's not ideal.

The canonical answer from #git seems to be "never pull", always use
fetch and rebase when submitting patches upstream using
git-format-patch.

I have tried various ways to somehow make rebase get along with
merging, but have failed on that, mainly because merging confuses
rebase's algorithm that figures out which patches have already been
applied.

So, the question is: what is the best way to be able to keep a full
local history, while still using format-patch to interact easily with
upstream?  I'm thinking the answer may involve a submission branch
that uses squashing, but I'm not entirely certain.

And the corrolary: As upstream, how can I facilitate users that may
wish to do this?

[1] http://colabti.org/irclogger/irclogger_log/git?date=2008-02-25,Mon&sel=66#l99
[2] http://changelog.complete.org/posts/690-Git-looks-really-nice,-until.....html

```

## Linus Torvalds, 2008-02-25 20:31

Subject: Re: "Contributors never merge" and preserving history
Message-ID: <alpine.LFD.1.00.0802251202380.14934@woody.linux-foundation.org>
URL: https://gitlist.dev/e/alpine.LFD.1.00.0802251202380.14934%40woody.linux-foundation.org
In-Reply-To: <slrnfs5pfh.lkc.jgoerzen@katherina.lan.complete.org>

```


On Mon, 25 Feb 2008, John Goerzen wrote:
> 
> The canonical answer from #git seems to be "never pull", always use
> fetch and rebase when submitting patches upstream using
> git-format-patch.

If you're going to submit them as patches, that is the correct answer, 
because what you want to do is basically keep your patch-queue up-to-date 
(which is exactly what "git rebase" does).

And no, you must never mix merging and rebasing - not because it's 
technically impossible (it *can* make sense in some circumstances), but 
because the two flows really are very different. Either you keep history 
and submit it as such (git-to-git merges up and down the chain) or you 
work with a "set of patches" model. Mixing the two in the same tree just 
leads to insanity (although mixing the two int he same project among 
different repositories can be a very good way to handle things).

But "never pull" isn't quite true either. 

Basically, the way to think about development is to try to keep things in 
"topic branches", which git is really good at. And the rule really 
shouldn' be "never pull" as much as "try to keep those topic-branches 
separate".

So pulling generally by definition mixes different branches (that's the 
merge part) in a way that a rebase does not. That's *especially* true 
about pulling from "upstream", because - pretty much by definition - that 
upstream is generally not even a well-defined topic branch that you want 
to merge, but simply the sum of all the *other* topic branches that have 
been merged upstream.

So the reason you should generally pull from downstream rather than 
upstream is that it keeps your development branch "focused" or "on target" 
or whatever you want to call it. And that's always a good idea, because 
now anybody who works together with you knows what he is getting.

So think of it as a cleanliness issue - it may not matter all that much if 
the only person you expect to ever pull your tree is always that same 
upstream (so even if you pull from your upstream, your upstream isn't 
really getting any mixed-up new code when he in turn pulls from you), but 
one thign I personally hoped for as a "design" was that what git really 
allows you to do (and _should_ encourage) is to have a less than perfectly 
hierarchical development stream.

And yes, in practice, pretty much every project ends up being pretty 
hierarchical after all, and that may be because of how people work (they 
want clarity, they want a simple "which tree is in charge" kind of model), 
but I still suspect it's at least partly - and perhaps mostly - simply due 
to historical patterns that it's just really hard to break.

So if you think of different git repositories as different branches (and 
that's what they really are!), then the "avoid pulling from upstream" is 
really about that "keep the topic branch focused and clean!".

And quite frankly, as "the upstream" for the kernel, I really appreciate 
people who ask me to pull, and that keep their histories clean, so that 
when I do a "gitk ORIG_HEAD.." after a pull, I get something that just 
looks real and not too messy. IOW, I like seeing myself pulling clear and 
well-defined topic branches (even if the "topics" I pull tend to be pretty 
big-picture topics, ie they may encompass "everything networking" or 
similar).

BUT!

There's always a but. In some cases, I will literally _ask_ a downstream 
person to pull from me. Havign the downstream doing a merge makes sense if 
the merge is non-trivial, and then the conflicting changes in a topic 
branch should generally be resolved by the side that has (a) the knowledge 
to do so (obviously) but also (b) the one who has the more specific 
changes (ie the side that has less work, and more targeted knowledge, of 
the things that conflict).

IOW, in the case of non-trivial conflicts, suddenly downstream is usually 
the one that has more knowledge, and now they should do the merge. And 
quite often, downstream may well know that ahead of time and be proactive, 
and just do the pull the "wrong way" and when asking me as an upstream 
member to merge, they'll let me know that they've already resolved the 
conflicts with what was in my tree.

The latter case is also something where a really long-lived topic branch 
may simply be doing those pulls over time every once in a while to just 
make sure that the topic branch never gets *too* far out of sync. However, 
if that's the reason for doing a merge, I'd almost suggest not just doing 
a "git pull" from upstream, but fetching and then merging at 
well-specified points (ie releases or release candidates).

That way you can also make a better and more useful merge message: not 
just "merged with upstream", but actually make it be "Synchronized with 
release v1.7.9". Which now makes a whole lot of conceptual sense at a 
higher level.

To recap:

 - from a purely technical sense it doesn't make any difference 
   what-so-ever who pulls and who doesn't, although you don't want it to 
   be *too* rare so that the different branches diverge so far as to make 
   it technically hard to synchronize later!

 - from a cleanliness angle - and *especially* if you want to work not 
   just in strict "upstream" and "downstream" patterns, but expect to 
   maybe have multiple upstreams (think "stable branch" or "vendor X wants 
   to pull this too"), a clean and clear topic banch is a really really 
   good idea.

   For example, let's say that you're developing a driver. If you start at 
   some specific kernel version (say, 2.6.24) and you do *not* generally 
   merge from my development tree, now suddenly other people can happily 
   pull from your tree to get the driver, even if they are stable kernels 
   or vendor kernels that don't want all the development crud that is in 
   my tree!

   See? Keeping a clean history actually makes your tree more useful!

 - But there are cases where pulling from up-stream really makes sense. 
   There may be specific points at upstream that you simply want to 
   synchronize with, or there may be conflicts that you want to resolve 
   simply because others aren't as knowledgeable about your topic branch
   etc.

So don't believe in "never pull from upstream". But *do* believe in "try 
to keep your branches on topic, because it will make everybody happier, 
and you'll be more easily able to read your own history too!".

			Linus

```

## Asheesh Laroia, 2008-02-25 21:35

Subject: Re: "Contributors never merge" and preserving history
Message-ID: <alpine.DEB.1.00.0802251330340.28694@dell.linuxdev.us.dell.com>
URL: https://gitlist.dev/e/alpine.DEB.1.00.0802251330340.28694%40dell.linuxdev.us.dell.com
In-Reply-To: <alpine.LFD.1.00.0802251202380.14934@woody.linux-foundation.org>

```
On Mon, 25 Feb 2008, Linus Torvalds wrote:

>   For example, let's say that you're developing a driver. If you start at
>   some specific kernel version (say, 2.6.24) and you do *not* generally
>   merge from my development tree, now suddenly other people can happily
>   pull from your tree to get the driver, even if they are stable kernels
>   or vendor kernels that don't want all the development crud that is in
>   my tree!
>
>   See? Keeping a clean history actually makes your tree more useful!

I'm going to chime in on this thread as a relative newcomer to git.  If 
I'm developing a driver or other feature branch, and then a new upstream 
release comes along, I can't rebase and push - that would make the "is not 
a strict subset of local ref" complaint.

Is the right workflow, then, to rebase against 2.6.25 in a new local 
branch, and push that to a new remote branch for others (like you say, 
vendor kernel maintainers) to pull from?

Thanks!

-- Asheesh.

-- 
Who will take care of the world after you're gone?

```

## Linus Torvalds, 2008-02-25 22:02

Subject: Re: "Contributors never merge" and preserving history
Message-ID: <alpine.LFD.1.00.0802251347530.14934@woody.linux-foundation.org>
URL: https://gitlist.dev/e/alpine.LFD.1.00.0802251347530.14934%40woody.linux-foundation.org
In-Reply-To: <alpine.DEB.1.00.0802251330340.28694@dell.linuxdev.us.dell.com>

```


On Mon, 25 Feb 2008, Asheesh Laroia wrote:
>
> On Mon, 25 Feb 2008, Linus Torvalds wrote:
> > 
> >   For example, let's say that you're developing a driver. If you start at
> >   some specific kernel version (say, 2.6.24) and you do *not* generally
> >   merge from my development tree, now suddenly other people can happily
> >   pull from your tree to get the driver, even if they are stable kernels
> >   or vendor kernels that don't want all the development crud that is in
> >   my tree!
> > 
> >   See? Keeping a clean history actually makes your tree more useful!
> 
> I'm going to chime in on this thread as a relative newcomer to git.  If I'm
> developing a driver or other feature branch, and then a new upstream release
> comes along, I can't rebase and push - that would make the "is not a strict
> subset of local ref" complaint.
> 
> Is the right workflow, then, to rebase against 2.6.25 in a new local branch,
> and push that to a new remote branch for others (like you say, vendor kernel
> maintainers) to pull from?

Almost always, the right workflow is to *neither* rebase *nor* pull.

Quite frankly, if you're working on some new feature like a driver, then 
in most projects you shouldn't need to care all that deeply about what is 
going on in other drivers etc. Merging or rebasing is just going to be a 
distraction, and open you up to new bugs that aren't even in your code!

Of course, this kind of situation can certainly be taken too far. The 
infrastructure may be changing, and you may want to rebase for that 
reason, but in most projects that kind of churn is (a) generally kept to a 
minimum and (b) shouldn't really necessarily be your headache anyway (ie 
I'm happy to handle merge conflicts even if I can't always _test_ them, 
and if other maintainers make changes to infrastructure they are also 
supposed to end up helping fix up the fallout!).

So I would not in general suggest that a driver writer maintain multiple 
branches based on different versions. That's just not worth your time, I 
think. You'd be better off staying back on whatever version you're 
comfortable with, and then perhaps rebasing or merging very occasionally 
when you start thinking that your base is simply too old to be relevant.

IOW, I'd suggest not merging or rebasing more than maybe once a month, if 
even that, unless you happen to be very bleeding edge (ie the 
infrastructure you depend on may itself be developing quickly, so a 
wireless driver in the kernel would generally see more need for being kept 
up-to-date than a random other driver).

But hey, it's also a matter of your personal taste, and how you work. 
There really are different models:

 - the "rebase" model is more amenable to a daily "fetch+rebase" kind of 
   ritual, and if you're the kind of person who really wants to feel that 
   you're always on the bleeding edge, maybe that's the right model for 
   you, even though I actually don't think it's really a logically very 
   good model (ie if you are actively doing development, you really don't 
   need the distraction!)

 - if you're working with somebody else (or a group), pulling from 
   *each*other* may be a great thing to do, and may well be the right 
   approach. But if you start doing that, then everybody involved should 
   avoid rebasing or pulling from upstream, because otherwise you'll just 
   get either tons of duplicate history (which wil *really* mess up 
   debugging: things like "git bisect" will work much worse if you have 
   the same bug introduced in multiple places etc)

 - the optimal strategy if you're just buffered enough is likely to not 
   rebase and not merge, and just work on your own thing, and then when 
   you feel ready, just say "please pull" to upstream.

The last one may feel a bit boring and staid, but I think it's the best 
one when it works (and "when it works" is a lot about _your_ psychology 
too: some people like that kind of insulation where they don't have to 
worry about what everybody else is doing, while others hate feeling like 
they're working on a tree that is a week old).

			Linus

```

## John Goerzen, 2008-02-26 14:04

Subject: Re: "Contributors never merge" and preserving history
Message-ID: <slrnfs8749.prc.jgoerzen@katherina.lan.complete.org>
URL: https://gitlist.dev/e/slrnfs8749.prc.jgoerzen%40katherina.lan.complete.org
In-Reply-To: <alpine.LFD.1.00.0802251202380.14934@woody.linux-foundation.org>

```
On 2008-02-25, Linus Torvalds <torvalds@linux-foundation.org> wrote:

[ snip ]

> So the reason you should generally pull from downstream rather than 
> upstream is that it keeps your development branch "focused" or "on target" 
> or whatever you want to call it. And that's always a good idea, because 
> now anybody who works together with you knows what he is getting.

Hi Linus,

Thank you very much for these two informative messages.  I think that
there were a lot of shades of gray to the kernel workflow that I failed
to appreciate before, for whatever reason.

I do have a question about the point you make above though.  I'm not
quite understanding what you're saying here.  Technically speaking,
the end result of a merge where you pulled from me would be identical
to a merge where I pulled from you.  Moreover, say I'm pretty far down
on the seniority list, kernel-wise.  Do you expect subsystem
maintainers to honor a request from me to pull from my tree, even if
they've never heard of me before, or would you think they'd only want
git format-patch output?

I ask because let's say I follow that advice above, and there are some
"downstreams" to me.  I pull from them, which involves some merging,
and then I want to format-patch.  It seems format-patch doesn't work
so well with merging.  What would I then do in this situation?  Should
I just use rebase to merge unless I know for sure that upstream will
honor a pull request?  But then again, we get into trouble if one of
my downstreams did a merge.

-- John

```

## Linus Torvalds, 2008-02-26 16:41

Subject: Re: "Contributors never merge" and preserving history
Message-ID: <alpine.LFD.1.00.0802260831220.14934@woody.linux-foundation.org>
URL: https://gitlist.dev/e/alpine.LFD.1.00.0802260831220.14934%40woody.linux-foundation.org
In-Reply-To: <slrnfs8749.prc.jgoerzen@katherina.lan.complete.org>

```


On Tue, 26 Feb 2008, John Goerzen wrote:
> 
> I do have a question about the point you make above though.  I'm not
> quite understanding what you're saying here.  Technically speaking,
> the end result of a merge where you pulled from me would be identical
> to a merge where I pulled from you.

Yes. Except for where the end result is!

That's kind of the point. If you're developing a driver, your tree is the 
"driver tree". But if you keep pulling from me, now it's no longer a 
driver tree, it's a "driver and Linus' code tree".

> Moreover, say I'm pretty far down on the seniority list, kernel-wise.  
> Do you expect subsystem maintainers to honor a request from me to pull 
> from my tree, even if they've never heard of me before, or would you 
> think they'd only want git format-patch output?

It probably depends on the submaintainer. But you're absolutely right that 
at least early on, most of them will want just emailed patches. And for 
that, the "fetch + rebase" model is the better one.

HOWEVER.

What happens with me is that I personally prefer patches from people if

 - they are "single" patches at a time (not necessarily just one, but at 
   most a couple at a time)

 - I've really never worked with you before

but if you have a real patch-series with more than (say) 4-5 patches, and 
I've seen patches from you before, _and_ you have a clean git tree, at 
that point I'd more likely actually already prefer a git pull if you can 
just describe your patches well enough in the email.

[ Of course, when it comes to me personally, another big requirement is 
  that I don't feel like you're going past some subsystem maintainer.

  It's not that I am a strict hierarchical person, it's that when it comes 
  to most subsystems I often don't feel competent enough to make the 
  decision, so I want things to go through submaintainers simply because 
  it's an extra layer of "filters".

  So the things I'd take through git are things that are either really 
  obvious or things I'd take anyway for other reasons. Those things are 
  seldom "patch series", but it happens.. ]

So at least judging by my own preferences, I don't think the barrier to 
doing a git merge is actually all that high. The *biggest* barrier may 
indeed be that I can happily do a "git pull", but if it doesn't look like 
a clean topic branch, I'd probably undo it.

			Linus

```
