threads / discuss / 14144

Re: policy and mechanism for less-connected clients

Subject: Re: policy and mechanism for less-connected clients

## tl;dr

5 messages between Jun 25, 2008 and Jun 25, 2008.

replies: 4people: 4as markdown or json

Theodore Tso· Jun 25, 2008, 13:34 UTC · lore
On Wed, Jun 25, 2008 at 05:20:49AM -0000, David Jeske wrote:
Show 6 quoted lines
> The other big one is ACLs in 'well named' repositories, so multiple
> people can safely be allowed to add changes to them, without giving
> them ability to blow away the repository. I can see this isn't the
> way all git users work, but at least a few users working this way
> now with shared push repositories. This is just making it
> 'safer'. Also seems pretty easy to do.

So this isn't true security, since someone determined (or an ingenious enough fool) can always blow away repository if you allow them to add changes; they could just add a change which rm's all of the files, yes? You just want to prevent something stupid.

Well, as long as they don't do non-fast forward updates (i.e., they never do something like: "git push publish +head:head", or any other incantation involving a leading '+' in the refspec), they should be pretty safe. I don't see how they would do any damage just due to user confusion. So I think git is pretty safe as-is.

Show 9 quoted lines
> > This is also easy; you just establish remote tracking branches. I
> > have a single shell scripted command, git-get-all, which pulls from
> > all of the repositories I am interested in into various remote
> > tracking branches so while I am disconnected, I can see what other
> > folks have done on their trees.
> 
> Yes, so I'd have the same thing, except instead of a remote
> repository, it would be a pattern of the branch namespace, such as
> /origin/users/jeske/*.

And the advantage of using branch namespaces instead of separate remote repositories is.... ? I don't see any....

Show 6 quoted lines
> Think about using CVS. user does "cvs up; hack hack hack; cvs commit
> (to server)". In git, this workflow is "git pull; hack; commit;
> hack; commit; git push (to server)". I want those interum "commits"
> to share the changes with the server. I want to change this to "git
> pull; hack; commit-and-share; hack; commit-and-share; git-push (to
> shared branch tag)"

OK, so *why* is it a good idea to ask people to share their in-progress work? What's the upside? Maybe if the idea is as backup if people are working from their laptops, and they're about to travel internationally or some such, but in general, sharing in-progress work is highly overrated.

The other thing is in your design assumption is that remote repositories are somehow expensive, when in fact they are very cheap; use either repo.or.cz or github; they support repo sharing so there isn't major cost to letting each developer having their own repository to push to.

So the way I would do things is to simply encourage people to do start their work by branching off of an up-to-date master branch, but *not* do any git pulls or git pushes. They can use git commit as necessary to save interim work, and they do all of this work on a private branch. When they are done doing their work, they should review the git commit points and make sure they make sense; in some cases they may be better off squashing the commits down to a single commit, or possibly refactoring their work so that each individual commit is free-standing, so that their series of commits is git-bisectable (i.e., after each commit the tree will fully compile and fully pass the project regression test suite).

Once they have done *that*, they make sure the master branch has been fully updated, and then do a git-rebase on their feature branch so that it is up-to-date with respect to master, and then they do a full build and regression test. Then they switch back to the master branch, and do a "git push publish" --- where <publish> is defined in .git/config to be something like this:

[remote "publish"]
	url = ssh://master.kernel.org/pub/scm/linux/kernel/git/tytso/ext4.git
	push = refs/heads/master:refs/heads/master

This will *only* push the master branch (and not any of the feature branches), and it will not allow non-fast forward merges. Hence, if the user screwed up and accidentally made changes to the master branch (say, an accidental git-rebase while on the master branch, or something else bone-headed), the git push will fail. This gives you the safety you desire about not accidentally screwing up the master branch.

And you're done. The only reason why you need a per-user repository if you want some safety in terms of backups in case the work being done on the laptop gets destroyed, but you can get that pretty much for free via git.or.cz or github. I really don't buy the sharing argument, because if you are in the middle of implementing a feature, it's generally not useful for others to look at your in-progress work.

> I know that all of what I wrote above seems strange if you don't buy into the
> design assumptions. That it's critical to share a single server-repository,
> that it's critical to have a shared 'well known' branch that only trusts
> clients to add new changes to, etc.. However, these are important.

Yep. And you still haven't justified why it's critical to share a single server repository. ***Why*** is that important?

And when you have shared push repositories, as long as users don't use the '+', in practice they can only add new changes. And if you don't trust them not to use the '+' character in refspecs, are you really going to trust them not to introduce either bone-headed mistakes into the code? Or to "git rm" the wrong files, git commit them, and then merge that into the repository? If all you care about is avoiding the accidentally stupid user mistakes, then putting in a convenience default so that "git push publish" always does what you want should be good enough.

So fundamentally, yeah, I think your primary problem is with the design assumptions, which haven't been justified at all.

       		    	  	       		     - Ted
Junio C Hamano· Jun 25, 2008, 17:34 UTC · re: Theodore Tso · lore
Theodore Tso <tytso@mit.edu> writes:
Show 5 quoted lines
> And when you have shared push repositories, as long as users don't use
> the '+', in practice they can only add new changes.  And if you don't
> trust them not to use the '+' character in refspecs, are you really
> going to trust them not to introduce either bone-headed mistakes into
> the code?

Well, if you do not trust them, just set receive.denynonfastforwards and they won't be able to.

David Jeske· Jun 25, 2008, 20:38 UTC · re: Theodore Tso · lore

Thanks for the info about shared object storage for shared repositories. That's great, and looks like a good implementation method.

Previously I was thinking in terms of making a different server to change behavior. However, I think the comments I've read are shifting my mindset towards making a client-wrapper. I want to provide a system [wrapper] without the user-burden of thinking about three repositories (local, my-public, shared-public). Doing this as a wrapper has other benefits, like the fact that users can treat services like repo.or.cz as the "networked filesystem of their version control system", so I like it.

I have a model for the operations of this wrapper below.
-- Theodore Tso wrote:
> [snip] sharing in-progress work is highly overrated.

_Seeing_ unfinished changes is overrated. However, so is managing multiple repositories and managing which data is shared.

I think my new wrapper approach below eliminates this overly-aggressive sharing while still reducing complexity for the average user.

> So the way I would do things is to simply encourage people to do start
> their work by branching off of an up-to-date master branch, but *not*
> do any git pulls or git pushes.

You confused me here. If their repo.or.cz private repository is their only way of sharing (because their home directory is inaccessible and emailing patches is cumbersome), how do they exchange their own changes without pushing? Even in a short time on git mailing list I see mini-unfinished-patches being posted.

> [ description of commit rewriting, rebase, push ]

The method you describe is burdening all users with learning a bunch of new concepts to do things that are unnecessary micromanagement for their needs. I'd prefer to give my users many of the benefits of DVCS/git with a command/argument set 1/20th the size and a much simpler mental model.

Most of the software we're all using was developed while working with centralized source control, where people just hack and commit and those commits are not even known-working. They don't bother with patch/commit rewriting and management, and it works out just fine. I can see how that finer granularity may be valuable for linux kernel coordinators. However, most projects don't need to bother with all that, and even in the ones that do, most of their contributors don't.

Despite the success of centralized revision control, distributed source control revision models have some very attractive features which can add efficiency to a shared-central-repo model without straying far from the familiar (cvs up; hack; hack; cvs up; cvs commit;) workflow. I read some commentary from Linus that compared git to a 'filesystem', and that's what I see.. a really awesome underlying set of mechanisms for implementing SCM.. I'm trying to understand how to layer an easy to use SCM system on top if it.

Some 'git' users might say the right thing to do is do a different project, but I think, just like with the filesystem-analogy, there is significant benefit to sharing a single repository model so a simple source control system can then be used in powerful ways by powerful users. This is similar to the direction "eg" (easy git) is heading, but more extreme and extending to the server.

In fact, it seems like we might be better off if all of these source control user-interfaces (cvs, perforce, git, eg, mercurial, etc. etc.) could be written on top of a version-control-api that they shared. Witness the similar implementation strategies of this modern rash of DVCS systems.

--------------------------------------------------------------

I'll try to explain my wrapper model in terms of an example... Imagine I'm going to deliver a "cvs drop in replacement", ncvs, that mostly keeps the cvs mental model, but is implemented underneath using git and just works better than cvs (yet is simpler than git). I'll use the exact cvs command parameters for illustration, but I wouldn't plan to do this. Notice how each ncvs command uses many git commands. It's possible these things should be done in terms of plumbing instead of porcelain to reduce dependence on git changes, but it's more concise to express them as porcelain.

>From the earlier feedback, there are now two repositories, one is considered
the "shared-root" while the other is the "user" repository.

(1) make "cvs update" safe, make it easy to see granular comments for things you have not pushed

CVS users do potentially destructive merges all the time. Despite the way we use terminology, working files ARE a branch, and "cvs up" IS a merge. That merge can require edits to resolve, and after those edits are complete, the previous state is NOT recoverable. There is no reason for this. We can easily save the delta by just making "cvs up" equal "git commit; git pull;", or alternately, "git stash; git pull; git apply;".

: "ncvs up" -> : : git stash; git pull; git apply; : git diff --stat <baseof:current branch> - un-pushed filenames : git-show-branch <current branch> - un-pushed comments

Question: when I say "baseof:current branch", I mean "the common-ancestor
between my local-repo tracking branch and the remote-repo branch it's
tracking". How do I find that out?

Adding "git diff --stat <baseof:current branch>" helps keep us aware of what changes are in our local repo. Any files not pushed up to the branch head on the server are seen. Likewise with "git-show-branch <current branch>" (which somehow is not the same as git-show-branch --current).

(2) make "planned ahead of time" branches cheap to make

"cvs up" is the easiest merge in cvs, therefore, separate sets of checked out working files become the most common form of branching in cvs. They are basically personal work branches that you can't commit on, and can't collaborate on. I've seen developers with cvs working directories weeks or months old because that's an easier way to work on different ideas than creating a branch and checking them in. DVCS fixes this, by making branches cheap to make, and by making all branch merges closer to the simplicity of cvs's easy branch merge "cvs up". However, I don't need to burden the user with the extra complexity and workload of the default being local branches, which they then need to do more work to share. I want branches to be shared by default.

: "ncvs tag -b --shared $branch" -> : : [ create a branch on the "shared root" repo, pointing : to where I am in my local tree, if I have permission ] : git branch --track $branch origin/$branch

: "ncvs tag -b mybranch : : [ create a branch on my "user" repo, pointing to where I am in my local tree, if I have permission ] : git branch --track $mybranch my-origin/$branch

Question: I'm not sure what commands to use above. How do I create a branch on
a remote repo when I'm on my local machine, without sshing to it?

The advantages of git's repository over cvs's repository in this use-case are not created because the branch is on the local machine. In fact, we also created it on the server. The benefit comes from the git revision storage model being faster and BETTER.

Then to switch our working pointer to this branch, we might do:

: "ncvs up -r mybranch" -> : : git stash; git checkout mybranch; git pull; : git stash show --relevant --recent;

Our "safe update" automatically saved away any local directory changes before switching off to the branch (if there were any). Our "stash show" is there always to show us if any stashes hang off a recent parent of the tree we just switched to, but it only shows them if they are hanging off this tree, and only if they are recent. If there is, we might want to look at or grab it, or we might just ignore it and not care.

(3) allow users to commit their 'final' changes to others (only on the branch they are on)

: "ncvs commit" -> "git commit; git push <only this branch>;"
Question: how do I only push the branch I'm on? "eg" says it does this, but
from a quick look at the code, it wasn't obvious to me how.

Developers who are plenty happy with their existing model of never saving local changes, can continue doing what they are doing. This makes the ability to save local changes an added benefit to the users like me that want to do it, instead of an extra burden to the other users. It also simplifies the issue of which changes are pushed to the server and which are not, because pushing is managed by "git push <only this branch>", not by creating and managing local and remote branch names separately. (easy git took the same approach with push)

(4) Allow users to save interim changes, without ahead of time planning, ahead of time nameing, and hopefully, without naming at all.

Saving interim changes in a cvs working tree before merging with head is not cheap. Making my own branch tag isn't too hard, but it takes a long time on a big tree. Ironically, perforce made branching mechanism faster while making the cognitive load of branch hing much higher.

: "ncvs save" -> "git commit -a" : : "ncvs stash [$name]" -> : : $currentbranch = `git branch` : $base-ish = '<baseof: current branch>' : git stash; : git branch -m $currentbranch $name; : git checkout $baseish; : git branch $currentbranch

This "ncvs stash" is acknowledging the value of the "git stash" idea, while also recognizing that when I'm using "git commit" regularly, I don't have anything in the working set! I really want to stash the changes made since "origin/<branchname>" and return there with my local <branchname>. This is really after the fact branch creation. If no $name is supplied, then it can auto-generate one like stash does.

(5) make it obvious there is a difference between local and remote changes, but make it easy to diff against remote before "ncvs commit;"

: "ncvs diff" -> : : echo -n "since commit(-C): " \ : `git diff --shortstat <baseof:current branch>`; \ : echo : echo -n "since save(-S): " \ : `git diff`; echo : : "ncvs diff -S" -> "git diff" : "ncvs diff -C" -> "git diff <baseof:current branch> --------------------------------------------------------------

I'm primarily trying to understand how to map my model to git. Continued thanks for the discussion and help.

Jakub Narebski· Jun 25, 2008, 20:52 UTC · re: David Jeske · lore
<opublikowany i wysłany>
David Jeske wrote:
> Question: when I say "baseof:current branch", I mean "the common-ancestor
> between my local-repo tracking branch and the remote-repo branch it's
> tracking". How do I find that out?
git-merge-base
-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git
Jakub Narebski· Jun 25, 2008, 20:54 UTC · re: David Jeske · lore
David Jeske wrote:
> Question: how do I only push the branch I'm on? "eg" says it does this, but
> from a quick look at the code, it wasn't obvious to me how.
git push <remote> HEAD   # with current enough git
-- 
Jakub Narebski
Warsaw, Poland
ShadeHawk on #git

← back to recent threads