threads / discuss / 8424

Git Vs. Svn for a project which *must* distribute binaries too.

Subject: Git Vs. Svn for a project which *must* distribute binaries too.

## tl;dr

21 messages between Jun 4, 2007 and Jun 6, 2007.

replies: 20people: 12as markdown or json

Bryan Childs· Jun 4, 2007, 11:48 UTC · lore
Hello git users / maintainers / fans,

My fellow projecteers and I watched a presentation given by Linus Torvalds on the advantages of git given at a google questions session sometime recently.

Our project, www.rockbox.org, an open source firmware replacement project for digital audio players currently makes use of subversion for it's source code management system, but Linus's eloquent (though sometimes rather blunt) speech has made us question whether git is perhaps a better solution for us.

On the whole, we like a lot of the features it offers but, we have a couple of issues which we've discussed, and so far have failed to come up with a decent resolution for them.

1) Due to the nature of our project, with multiple architectures
supported, we strive to provide a binary build of our software with
every commit to the subversion repository. This is so that we can
provide a working firmware for the majority of our users that don't
have the necessary know-how for cross-compiling and so forth.
2) Unlike the Linux Kernel, which Linus uses as a prime example of
something git is very useful for, the Rockbox project has no central
figurehead for anyone to consider as owning the "master" repository
from which to build the "current" version of the Rockbox firmware for
any given target.
3) With a central repository, for which we have a limited number of
individuals having commit access, it's easy for us to automate a build
based on each commit the repository receives.

Given these three points, we wonder how we'd best achieve the same using git. As far as we can make out we'd need to appoint someone as a maintainer for a master repository whose job it is to co-ordinate pulls from people based on when they've made changes we wish to include in the latest version of our software. This sounds like a time consuming role for a project which is only staffed by volunteers.

Can anyone offer any insights for us here?
Bryan
Julian Phillips· Jun 4, 2007, 11:56 UTC · re: Bryan Childs · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, 4 Jun 2007, Bryan Childs wrote:
Show 16 quoted lines
> 2) Unlike the Linux Kernel, which Linus uses as a prime example of
> something git is very useful for, the Rockbox project has no central
> figurehead for anyone to consider as owning the "master" repository
> from which to build the "current" version of the Rockbox firmware for
> any given target.
>
> 3) With a central repository, for which we have a limited number of
> individuals having commit access, it's easy for us to automate a build
> based on each commit the repository receives.
>
> Given these three points, we wonder how we'd best achieve the same
> using git. As far as we can make out we'd need to appoint someone as a
> maintainer for a master repository whose job it is to co-ordinate
> pulls from people based on when they've made changes we wish to
> include in the latest version of our software. This sounds like a time
> consuming role for a project which is only staffed by volunteers.
You can setup git to work in a centralised style if you wish.
See http://www.kernel.org/pub/software/scm/git/docs/cvs-migration.html
-- 
Julian

  ---
If reporters don't know that truth is plural, they ought to be lawyers.
 		-- Tom Wicker
Theodore Tso· Jun 4, 2007, 13:18 UTC · re: Bryan Childs · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, Jun 04, 2007 at 12:48:17PM +0100, Bryan Childs wrote:
Show 5 quoted lines
> 2) Unlike the Linux Kernel, which Linus uses as a prime example of
> something git is very useful for, the Rockbox project has no central
> figurehead for anyone to consider as owning the "master" repository
> from which to build the "current" version of the Rockbox firmware for
> any given target.
> 3) With a central repository, for which we have a limited number of
> individuals having commit access, it's easy for us to automate a build
> based on each commit the repository receives.

You might want to take a look at http://repo.or.cz for an example of how you can have a limited number of trusted inidividuals with commit access. As has been said before, <SCM> is not a substitute for communication, and if you have multiple people who can commit into a repository, you had better make sure those trusted individuals with commit access are talking to each other.

There are some folks who have created hooks to do more fine-grained access control systems, if you want to replicate SVN's ability to control who can commit to which branch.

Regards,
						- Ted
Johannes Schindelin· Jun 4, 2007, 14:58 UTC · re: Bryan Childs · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

Hi,
On Mon, 4 Jun 2007, Bryan Childs wrote:
> 1) Due to the nature of our project, with multiple architectures
> supported, we strive to provide a binary build of our software with
> every commit to the subversion repository.

Git has no problems with binaries. Actually, one could argue that it has less problems with binary files than with text files, since it only recently acquired the capability (disabled by default) to transcribe certain files into the CR/LF line ending some Windows programs still insist on.

As for checking in binaries, you even could set up a post-commit hook, which builds the binary, and checks it into a separate branch...

Ciao, Dscho

Linus Torvalds· Jun 4, 2007, 15:20 UTC · re: Bryan Childs · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, 4 Jun 2007, Bryan Childs wrote:
Show 6 quoted lines
> 
> 1) Due to the nature of our project, with multiple architectures
> supported, we strive to provide a binary build of our software with
> every commit to the subversion repository. This is so that we can
> provide a working firmware for the majority of our users that don't
> have the necessary know-how for cross-compiling and so forth.

Git has no problems with binaries, but I _really_ I hope that you don't actually want to check these binaries into the repository? You could do that, and the git delta algorithm might even be able to compress the binaries against each other, but it could still be pretty nasty.

And by "pretty nasty" I don't mean that git won't be able to handle it: I suspect it's no worse from a disk size perspective than SVN. But since git is distributed, it means that everybody who fetches it will get the whole archive with whole history - it means that cloning the result is going to be really painful with tons of old binaries that nobody really cares about beign pushed around.

So I *hope* that you want to just have automated build machinery that builds the binaries to a *separate* location? You could use git to archive them, and you can obviously (and easily) name the resulting binary blobs by the versions in the source tree, but I'm just saying that trying to track the binaries from within the same git repository as the source code is less than optimal.

Show 5 quoted lines
> 2) Unlike the Linux Kernel, which Linus uses as a prime example of
> something git is very useful for, the Rockbox project has no central
> figurehead for anyone to consider as owning the "master" repository
> from which to build the "current" version of the Rockbox firmware for
> any given target.

The kernel is really kind of odd in that it has just a single maintainer. That's usually the case only for much smaller projects.

And no, git is not at all exclusively *designed* for that situation, although it is arguably one situation that git works really well for.

There is nothing to say that you cannot have shared repositories that are writably by multiple users. Anything that works for a single person works equally well for a "group of people" that all write to the same central git repo. It ends up not being how the kernel does things (not because of git, but because it's not how I've ever worked), but the kernel situation really _is_ pretty unusual.

So git makes everybody have their own repository in order to commit, but you can (and some people do) just view that as your "CVS working tree", and every time you commit, you end up pushing to some central repository that is writable by the "core group" that has commit access.

In *practice*, I suspect that once you get used to the git model, you'd actually end up with a hybrid scheme, where you might have a *smaller* core group with commit access to the central repository (in git, it wouldn't be "commit access", it would really be "ability to push", but that's a technical difference ratehr than anything conceptually huge), and members in that core group end up pulling from others.

But that would literally be once you have gotten used to the git model, and you can start out just totally emulating the old CVS/SVN model with a single central repository.

> 3) With a central repository, for which we have a limited number of
> individuals having commit access, it's easy for us to automate a build
> based on each commit the repository receives.

.. and that's exactly how you'd do it with git too. You wouldn't have a "commit trigger", but you'd have a "receive trigger", which triggers whenever somebody pushes to the central repository.

And that does mean that a developer might do a series of _five_ commits locally on his own machine, and they are totally invisible to everybody until he pushes to the central repository: and then the build will build just the top-most end result commit. So you'd not necessarily have a binary for _each_ commit, but:

 - you could (if you really wanted to) actually force people to always 
   send just one commit at a time. You could even enforce that in the 
   pre-receive triggers, so that people *cannot* push multiple commits at 
   a time.
   Quite frankly, I really don't think you want to go this way. I think 
   you want to perhaps _encourage_ people to send just one commit at a 
   time, but the much better model is the other choice:
 - realize that the git model tends to encourage many small commits 
   (because you *can* make commits without impacting others), so when you 
   fix something, or add a new feature, with git, you can do it as many 
   small steps, and then only "push" when it's ready.
   IOW, if you encourage people to do small step-wise changes, you 
   probably don't even *want* a build for each commit, you really want a 
   build for the case where "my feature is now ready, I'll push". So you'd 
   effectively get one build not per commit, but per "publication point".

But anyway, it really boils down to: you *can* use a distributed development model to emulate a totally centralized situation (put another way: "centralized" is just one very trivial special case of "distributed"), but I suspect that while you might want to start out trying to change as little as possible in your development model, I equally strongly suspect that you'll find out that the distributed nature makes _some_ changes to the model very natural, and you'll end up with more of a hybrid setup: aspects of a centralized model, but with distributed elements.

		Linus
Bryan Childs· Jun 4, 2007, 15:38 UTC · re: Linus Torvalds · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On 6/4/07, Linus Torvalds < [send email to torvalds@linux-foundation.org via gmail] torvalds@linux-foundation.org> wrote:

Show 6 quoted lines
> So I *hope* that you want to just have automated build machinery that
> builds the binaries to a *separate* location? You could use git to archive
> them, and you can obviously (and easily) name the resulting binary blobs
> by the versions in the source tree, but I'm just saying that trying to
> track the binaries from within the same git repository as the source code
> is less than optimal.

Oh lord no - I never meant to imply that we'd be checking those binaries in, I just meant to hi-light that we need a central repository to build those binaries from - otherwise we'd end up with a selection of binaries for our users to download which contain a bunch of different features if they were built from a combination of repositories. I know you think everyone else is a moron, but we're not quite dumb enough to think maintaining binaries in a repository is a good idea :)

Show 6 quoted lines
> In *practice*, I suspect that once you get used to the git model, you'd
> actually end up with a hybrid scheme, where you might have a *smaller*
> core group with commit access to the central repository (in git, it
> wouldn't be "commit access", it would really be "ability to push", but
> that's a technical difference rather than anything conceptually huge), and
> members in that core group end up pulling from others.

This sounds like what we eventually came up with. I'm not sure how soon we'll make a switch to a git repository, but when we do, this seems to be the best model for the conversion in the short term, and perhaps in the long term too.

> .. and that's exactly how you'd do it with git too. You wouldn't have a
> "commit trigger", but you'd have a "receive trigger", which triggers
> whenever somebody pushes to the central repository.

Yes, after I'd sent my email this morning I found you could do pushes as well as pulls. That'll teach me to RTFM properly next time.

>  - realize that the git model tends to encourage many small commits
>    (because you *can* make commits without impacting others), so when you
>    fix something, or add a new feature, with git, you can do it as many
>    small steps, and then only "push" when it's ready.

This is what I personally was trying to advocate in our discussion - but I'm not sure everyone quite understood it. Hopefully your explanation will do a better job :)

>    IOW, if you encourage people to do small step-wise changes, you
>    probably don't even *want* a build for each commit, you really want a
>    build for the case where "my feature is now ready, I'll push". So you'd
>    effectively get one build not per commit, but per "publication point".
Absolutely.
>                 Linus

Thanks for your time (and everyone else who replied) - it's very much appreciated!

Bryan
Linus Torvalds· Jun 4, 2007, 16:23 UTC · re: Bryan Childs · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, 4 Jun 2007, Bryan Childs wrote:
> 
> Oh lord no - I never meant to imply that we'd be checking those
> binaries in, I just meant to hi-light that we need a central
> repository to build those binaries from

Heh. I get worried (and judging from other responses, I wasn't the only one) when people start talking about generated binaries and SCM's.

Because people _have_ traditionally done things like commit the generated files too.

But if it's just an automated build server, everything is good. That's trivial to do.

Show 11 quoted lines
> > In *practice*, I suspect that once you get used to the git model, you'd
> > actually end up with a hybrid scheme, where you might have a *smaller*
> > core group with commit access to the central repository (in git, it
> > wouldn't be "commit access", it would really be "ability to push", but
> > that's a technical difference rather than anything conceptually huge), and
> > members in that core group end up pulling from others.
> 
> This sounds like what we eventually came up with. I'm not sure how
> soon we'll make a switch to a git repository, but when we do, this
> seems to be the best model for the conversion in the short term, and
> perhaps in the long term too.

Yes. As mentioned, the kernel model of having just one person push is actually fairly rare.

When you have multiple people pushing, you have issues that I never have, but that you've already seen with CVS/SVN, for all the same reasons: you may need to merge the changes that others have done while you were working on yours.

However, the git "push" model is *different* from the CVS/SVN "commit" model.

In CVS/SVN, if you want to commit, and somebody else has done updates to the central repository, the "cvs commit" phase will obviously tell you that you're not up-to-date, and you cannot commit at all. So you end up doing a "cvs update -d" equivalent to first update your tree, then you have to resolve any conflicts, and then you can try to commit again.

In git, this is technically very different, yet similar. Since you can always commit to your *local* repository, when you do a "git commit", you'll never have any conflicts at all, because there is no conflicting work!

But the conflicts happen when you then do a "git push" to send out your commit(s) to the central repository. If nobody else has done any changes, at that point, you'll get exactly the same kind of situation as when you do a CVS commit, and the server will tell you that you're not up-to-date, and will refuse to take your push.

(The message is different: git will tell you that you try to push a commit that is not a "strict superset" of what the central repository has).

So when that happens with git, you actually have two different options:
 - you can do "git pull" to merge the central changes, and in that case 
   you get the exact same kinds of conflict markers for any conflicting 
   code that you would have gotten for "cvs update"
   This is how most people would probably use it, and it's the simplest 
   one, where you get very traditional commit conflict markers, fix it up, 
   and commit the merge. 
   However, it does end up making the history explicitly showing the 
   parallelism that happened, and while that is *correct* and can be very 
   useful, sometimes it means that especially if you've done just trivial 
   changes, you might want to take an alternate approach that "linearizes" 
   the history and makes it appear linear instead of parallel:
 - instead of doing a "git pull" that merges the two branches (your work, 
   and the work that happened by somebody else in the central repo while 
   you did it), you *may* also just want to do a "git fetch" to fetch the 
   changes from the central repo, and then do "git rebase origin" to 
   linearize the work you did on _top_ of those central repo one (so that 
   it no longer looks like a branch, and looks linear)
   In the "git rebase" case, you'll effectively merge your commits one at 
   a time, and you may thus have to fix up *multiple* conflicts. So it's 
   potentially more work, but it results in a simpler history if you want 
   it.

Regardless of how you ended up sorting out the fact that you had parallel development, once you've resolved it, you do a "git push" again, and now the stuff you're pushing is a proper superset of what the central repository had, so it will happily push it out.

(Of course, the exact same thing that can happen with CVS central repositories can happen with git ones too: by the time you've resolved all the differences and are ready to push them to the central one, somebody else might have pushed *more*, and you may need to do another "update" ;)

> Yes, after I'd sent my email this morning I found you could do pushes
> as well as pulls. That'll teach me to RTFM properly next time.

I think we talk a lot more about pulls, because we have had more people ask about them, and because more people tend to pull than to push.

The pull is also somewhat easier to explain. The pushing thing always has to talk about resolving differences when different people have pushed, so teaching people to push by necessity involves first teaching them about merging (ie pull or rebase).

Also, "push" is also a bit more interesting to explain, because a "push" won't update the working tree on the other end, so when you explain pushing, you should also explain about "bare" repositories (which I didn't do)), ie about having git repositories without any working tree associated with them.

So there is a bit of a learning experience involved, but espeically if some of the developers have seen git used in other environments (perhaps not as developers, just as users), it shouldn't be *that* hard to pick up. But there does seem to be a pretty big mental leap from the "centralized" thing to the "distributed" thing - I just moved over so long ago that I even have trouble understanding why people sometimes don't seem to find the distributed model the only natural and sane thing to do.

(It really does seem to be one of those "aha!" moments. People think distributed just adds a lot of complexity, and it takes a "Oh, *THAT* is how it works" kind of enlightenment to just switch your brain over, and I guarantee that once that moment on enlightenment hits, you'll never go back, but I cannot guarantee that that moment will happen for all developers ;)

		Linus
Thomas Glanzmann· Jun 4, 2007, 17:57 UTC · re: Linus Torvalds · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

Hello,
Show 6 quoted lines
>  - instead of doing a "git pull" that merges the two branches (your work, 
>    and the work that happened by somebody else in the central repo while 
>    you did it), you *may* also just want to do a "git fetch" to fetch the 
>    changes from the central repo, and then do "git rebase origin" to 
>    linearize the work you did on _top_ of those central repo one (so that 
>    it no longer looks like a branch, and looks linear)
>    In the "git rebase" case, you'll effectively merge your commits one at 
>    a time, and you may thus have to fix up *multiple* conflicts. So it's 
>    potentially more work, but it results in a simpler history if you want 
>    it.
Thank you a lot. I finally understood what "git rebase" is all about!
        Thomas
Linus Torvalds· Jun 4, 2007, 20:45 UTC · re: Thomas Glanzmann · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, 4 Jun 2007, Thomas Glanzmann wrote:
Show 14 quoted lines
> 
> >  - instead of doing a "git pull" that merges the two branches (your work, 
> >    and the work that happened by somebody else in the central repo while 
> >    you did it), you *may* also just want to do a "git fetch" to fetch the 
> >    changes from the central repo, and then do "git rebase origin" to 
> >    linearize the work you did on _top_ of those central repo one (so that 
> >    it no longer looks like a branch, and looks linear)
> > 
> >    In the "git rebase" case, you'll effectively merge your commits one at 
> >    a time, and you may thus have to fix up *multiple* conflicts. So it's 
> >    potentially more work, but it results in a simpler history if you want 
> >    it.
> 
> Thank you a lot. I finally understood what "git rebase" is all about!
I'd like to point out some more upsides and downsides of "git rebase".
Downsides:
 - you're rewriting history, so you MUST NOT have made your pre-rebase 
   changes available publicly anywhere else (or you are in a world of pain 
   with duplicate history and tons of confusion)
 - you can only rebase "simple" commits. If you don't just have a linear 
   history of your own commits, but have merged from others, rebasing 
   isn't a sane alternative (yeah, we could make it do something half-way 
   sane, but really, it's not worth even contemplating)
Upsides:
 - while there may be more conflicts you have to sort out, they may be 
   individually  simpler, so you *might* actually prefer to do it that 
   way.
 - if the reason for the conflicts is that upstream did some nice cleanup 
   in the same area, and you decide that you would actually want to re-do 
   your development based on that nice cleanup, then "git rebase" can 
   actually be used as a way to help you do exactly that. IOW, you can 
   take _advantage_ of the conflicts as a way to re-apply the patches but 
   also then fix them up by hand to work in the new (better) world order.

And finally, the upside that is probably the most common case for using "git rebase", and has nothing to do with resolving conflicts before pushing them out with "git push":

 - if you actually want to send your changes upstream as emailed *patches* 
   rather than by pushing them out (or asking somebody else to pull them),
   rebasing is an excellent way to keep the set of patches "fresh" on top 
   of the current development tree.
   People who send their patches out as emails are also unlikely to have 
   the downsides (ie they normally send them as patches exactly *because* 
   they don't want to make their git trees public, and they probably just 
   have a small set of simple patches in their tree anyway)

So I have to say, I'm still very ambivalent about rebasing. It's definitely a very useful thing to do, but at the same time I think "git pull" in many ways is often the more honest and correct way to do things.

		Linus
Olivier Galibert· Jun 4, 2007, 21:21 UTC · re: Linus Torvalds · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, Jun 04, 2007 at 01:45:26PM -0700, Linus Torvalds wrote:
Show 7 quoted lines
> I'd like to point out some more upsides and downsides of "git rebase".
> 
> Downsides:
> 
>  - you're rewriting history, so you MUST NOT have made your pre-rebase 
>    changes available publicly anywhere else (or you are in a world of pain 
>    with duplicate history and tons of confusion)

Wouldn't it be possible to register the rebase somewhere (weak parent? some kind of note not influencing the sha1 ?) that pull/merge could follow? Rebases and cherry-picking are a special kind of merge, so maybe it can be handled like one where it counts...

  OG.
Linus Torvalds· Jun 4, 2007, 21:33 UTC · re: Olivier Galibert · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, 4 Jun 2007, Olivier Galibert wrote:
Show 13 quoted lines
> On Mon, Jun 04, 2007 at 01:45:26PM -0700, Linus Torvalds wrote:
> > I'd like to point out some more upsides and downsides of "git rebase".
> > 
> > Downsides:
> > 
> >  - you're rewriting history, so you MUST NOT have made your pre-rebase 
> >    changes available publicly anywhere else (or you are in a world of pain 
> >    with duplicate history and tons of confusion)
> 
> Wouldn't it be possible to register the rebase somewhere (weak parent?
> some kind of note not influencing the sha1 ?) that pull/merge could
> follow?  Rebases and cherry-picking are a special kind of merge, so
> maybe it can be handled like one where it counts...

Well, it's not like duplicate history is a disaster from a *technical* angle. It might be a small space-waster etc, but that's really not the real issue.

The problem with duplicate history is that it just makes things much harder to look at. IOW, it's *messy*. So the "tons of confusion" part is basically purely about humans, not about git itself. Git won't really care, and there's no reason to "handle" it specially in that sense.

So I would strongly discourage people from ever making rebased history available, but that's not because of any particular git technical issues as just because of it being a good way to confuse all the _humans_ involved.

(That said, gits own 'pu' branch ends up jumping around, and it hasn't caused all that much confusion, so maybe I'm overstating even that human confusion)

			Linus
Joel Becker· Jun 4, 2007, 22:30 UTC · re: Linus Torvalds · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, Jun 04, 2007 at 02:33:18PM -0700, Linus Torvalds wrote:
> (That said, gits own 'pu' branch ends up jumping around, and it hasn't 
> caused all that much confusion, so maybe I'm overstating even that human 
> confusion)
	It survives because it is well-known.  Everyone expects it to
break.  ocfs2 has an "ALL" branch that is everything we have working,
sort of a "test this bleeding edge" thing.  It gets rebased all the
time, and everyone knows that they can't trust it to update linearly.
Other developers have similar things in their repositories.
Joel
-- 
"What no boss of a programmer can ever understand is that a programmer
 is working when he's staring out of the window"
	- With apologies to Burton Rascoe

Joel Becker
Principal Software Developer
Oracle
E-mail: joel.becker@oracle.com
Phone: (650) 506-8127
Theodore Tso· Jun 5, 2007, 11:19 UTC · re: Joel Becker · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, Jun 04, 2007 at 03:30:03PM -0700, Joel Becker wrote:
Show 5 quoted lines
> 	It survives because it is well-known.  Everyone expects it to
> break.  ocfs2 has an "ALL" branch that is everything we have working,
> sort of a "test this bleeding edge" thing.  It gets rebased all the
> time, and everyone knows that they can't trust it to update linearly.
> Other developers have similar things in their repositories.

I wonder if it would be useful to be able to be able to flag a branches as "jumping around a lot", where this flag would be downloaded from another repository when it is cloned, so that a naive user could get some kind of warning before committing a patch on top of one of these branches that is known jump around.

	"This branch gets rebased all the time and is really meant for
	testing.  If you really want to commit this changeset, please
	configure yourself for expert mode or use the --force."
Or maybe just a warning, ala what we do with detached heads.
						- Ted
Johannes Schindelin· Jun 5, 2007, 02:56 UTC · re: Olivier Galibert · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

Hi,
On Mon, 4 Jun 2007, Olivier Galibert wrote:
Show 13 quoted lines
> On Mon, Jun 04, 2007 at 01:45:26PM -0700, Linus Torvalds wrote:
>
> > I'd like to point out some more upsides and downsides of "git rebase".
> > 
> > Downsides:
> > 
> >  - you're rewriting history, so you MUST NOT have made your pre-rebase 
> >    changes available publicly anywhere else (or you are in a world of 
> >    pain with duplicate history and tons of confusion)
> 
> Wouldn't it be possible to register the rebase somewhere (weak parent? 
> some kind of note not influencing the sha1 ?) that pull/merge could 
> follow?

Actually, with reflogs (if you did not explicitely disable them), you should have the information already.

> Rebases and cherry-picking are a special kind of merge, so maybe it can 
> be handled like one where it counts...
There is something I have to add as a real disadvantage in rebase:

Usually you are expected to test your commits. So, say that you work on some patch series, and produce 3 well tested patches. Then you fetch upstream and realize it advanced by some commits, and rebase your three patches.

However, _none_ of your patches is well tested, because there is a quite real chance that your patches interact _badly_ with the patches you just fetched.

And if that is the case, git-bisect can very well attribute it to a wrong patch, either because more than one patch is bad, or because the last patch in your series _exposes_ the bug (but does not _introduce_ it).

Ciao, Dscho

Martin Langhoff· Jun 4, 2007, 22:29 UTC · re: Bryan Childs · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On 6/5/07, Bryan Childs <godeater@gmail.com> wrote:
> Oh lord no - I never meant to imply that we'd be checking those
> binaries in, I just meant to hi-light that we need a central
> repository to build those binaries from - otherwise we'd end up with a

If your infrastructure to build the binaries is automated, you can easily script the build for new incoming commits. The output of git-describe is really useful for this if you are going to name your builds `git describe`-<arch>.tar.gz.

OTOH, commit is different from push (vs SVN where both are one op), and that means that when using git you can present a large change as a better-explained patch-series. That's actually a good practice for new development, and it might not make sense to have literally one-build-per-commit.

Maybe I'd enable auto-builds for maintenance/bugfixes branches, and on other (experimental/devel) branches only auto-build commits selected explicitly (tagged?).

cheers,
martin
Daniel Barkalow· Jun 4, 2007, 23:48 UTC · re: Bryan Childs · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, 4 Jun 2007, Bryan Childs wrote:
Show 18 quoted lines
> On 6/4/07, Linus Torvalds < [send email to
> torvalds@linux-foundation.org via gmail]
> torvalds@linux-foundation.org> wrote:
> > So I *hope* that you want to just have automated build machinery that
> > builds the binaries to a *separate* location? You could use git to archive
> > them, and you can obviously (and easily) name the resulting binary blobs
> > by the versions in the source tree, but I'm just saying that trying to
> > track the binaries from within the same git repository as the source code
> > is less than optimal.
> 
> Oh lord no - I never meant to imply that we'd be checking those
> binaries in, I just meant to hi-light that we need a central
> repository to build those binaries from - otherwise we'd end up with a
> selection of binaries for our users to download which contain a bunch
> of different features if they were built from a combination of
> repositories. I know you think everyone else is a moron, but we're not
> quite dumb enough to think maintaining binaries in a repository is a
> good idea :)

Actually, I've been playing with using git's data-distribution mechanism to distribute generated binaries. You can do tags for arbitrary binary content (not in a tree or commit), and, if you have some way of finding the right tag name, you can fetch that and extract it.

I came up with this at my job when we were trying to decide what to do with firmware images that we'd shipped, so that we'd be able to examine them again even if we lose the compiler version we used at the time. We needed an immutable data store with a mapping of tags to objects, and I realized that we already had something with these exact characteristics.

	-Daniel
*This .sig left intentionally blank*
Linus Torvalds· Jun 5, 2007, 00:21 UTC · re: Daniel Barkalow · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, 4 Jun 2007, Daniel Barkalow wrote:
Show 5 quoted lines
> 
> Actually, I've been playing with using git's data-distribution mechanism 
> to distribute generated binaries. You can do tags for arbitrary binary 
> content (not in a tree or commit), and, if you have some way of finding 
> the right tag name, you can fetch that and extract it.

Yes, I think git should be very nice for doing binary stuff like firmware images too, my only worry is literally about "mixing it in" with other stuff.

Putting lots of binary blobs into a git archive should work fine: but if you would then start tying them together (with a commit chain), it just means that even if you only really want _one_ of them, you end up getting them all, which sounds like a potential disaster.

On the other hand, if you actually want a way to really *archive* the dang things, that may well be what you actually want. In that case, having a separate branch that only contains the binary stuff might actually be what you want to do (and depending on the kind of binary data you have, the delta algorithm might even be good at finding common data sequences and compressing it).

Show 5 quoted lines
> I came up with this at my job when we were trying to decide what to do 
> with firmware images that we'd shipped, so that we'd be able to examine 
> them again even if we lose the compiler version we used at the time. We 
> needed an immutable data store with a mapping of tags to objects, and I 
> realized that we already had something with these exact characteristics.

Yeah, if you just tag individual blobs, git will keep track of them, but won't link them together, so you can easily just look up and fetch a single one from such an archive. Sounds sane enough.

		Linus
david@lang.hm· Jun 5, 2007, 01:42 UTC · re: Linus Torvalds · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, 4 Jun 2007, Linus Torvalds wrote:
Show 15 quoted lines
> On Mon, 4 Jun 2007, Daniel Barkalow wrote:
>>
>> Actually, I've been playing with using git's data-distribution mechanism
>> to distribute generated binaries. You can do tags for arbitrary binary
>> content (not in a tree or commit), and, if you have some way of finding
>> the right tag name, you can fetch that and extract it.
>
> Yes, I think git should be very nice for doing binary stuff like firmware
> images too, my only worry is literally about "mixing it in" with other
> stuff.
>
> Putting lots of binary blobs into a git archive should work fine: but
> if you would then start tying them together (with a commit chain), it just
> means that even if you only really want _one_ of them, you end up getting
> them all, which sounds like a potential disaster.

if you put the binaries in a seperate repository and do shallow clones to avoid getting all the old stuff wouldn't that work well?

David Lang
Show 23 quoted lines
> On the other hand, if you actually want a way to really *archive* the dang
> things, that may well be what you actually want. In that case, having a
> separate branch that only contains the binary stuff might actually be what
> you want to do (and depending on the kind of binary data you have, the
> delta algorithm might even be good at finding common data sequences and
> compressing it).
>
>> I came up with this at my job when we were trying to decide what to do
>> with firmware images that we'd shipped, so that we'd be able to examine
>> them again even if we lose the compiler version we used at the time. We
>> needed an immutable data store with a mapping of tags to objects, and I
>> realized that we already had something with these exact characteristics.
>
> Yeah, if you just tag individual blobs, git will keep track of them, but
> won't link them together, so you can easily just look up and fetch a
> single one from such an archive. Sounds sane enough.
>
> 		Linus
> -
> To unsubscribe from this list: send the line "unsubscribe git" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>
Linus Torvalds· Jun 5, 2007, 03:58 UTC · re: david@lang.hm · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

On Mon, 4 Jun 2007, david@lang.hm wrote:
> 
> if you put the binaries in a seperate repository and do shallow clones to
> avoid getting all the old stuff wouldn't that work well?

Yes. I'm not a huge fan of shallow clones, and I suspect they've not gotten all that much testing, but that would certainly solve the problem of getting unnecessarily much data..

		Linus
Jakub Narebski· Jun 4, 2007, 23:46 UTC · re: Bryan Childs · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

Bryan Childs wrote:
> 3) With a central repository, for which we have a limited number of
> individuals having commit access, it's easy for us to automate a build
> based on each commit the repository receives.

Check out contrib/continuous/ scripts in git repository: you would have to enable it only on one machine, of course.

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Jakub Narebski· Jun 6, 2007, 22:34 UTC · re: Jakub Narebski · lore

Re: Git Vs. Svn for a project which *must* distribute binaries too.

Jakub Narebski wrote:
Show 8 quoted lines
> Bryan Childs wrote:
> 
>> 3) With a central repository, for which we have a limited number of
>> individuals having commit access, it's easy for us to automate a build
>> based on each commit the repository receives.
> 
> Check out contrib/continuous/ scripts in git repository: you would have
> to enable it only on one machine, of course.

You can also use something similar to dodoc.sh script in 'todo' branch of git repository, which script makes some build results and saves them in _separate_ branch of repository.

-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git

← back to recent threads