threads / discuss / 14133

Re: policy and mechanism for less-connected clients

Subject: Re: policy and mechanism for less-connected clients

## tl;dr

6 messages between Jun 25, 2008 and Aug 14, 2016.

replies: 5people: 5as markdown or json

Theodore Tso· Jun 25, 2008, 02:33 UTC · lore
On Wed, Jun 25, 2008 at 12:36:03AM -0000, David Jeske wrote:
Show 5 quoted lines
> The purpose of this mechanism is to host a distributed source
> repository in a world where most most developer contributors are
> behind firewalls and do not have access to, or do not want to
> configure a unix server, ftp, or ssh to possibly contribute to a
> project. 
Show 8 quoted lines
> design assumptions:
> 
> - all developers are firewalled and can not be "pulled" from directly.
> - there can be one or more well-connected servers which all users can access.
> - .. but which they cannot have ssh, ftp, or other dangerous access to
> - .. and whose protocol should be layered on http(s)
> - there is a shared namespace for branches, and tags
> - .. users are not-trusted to change the branches or tags of other users

Up to here, you can do this all with repo.or.cz, and/or github; you just give each developer their own repository, which they are allowed to push to, and no once else. Within their own repository they can make changes to their branches, so that all works just fine.

Show 6 quoted lines
> (a) safely "share" every DAG, branch, and tag data in their
> repository to a well-connected server, into an established
> namespace, while only changing branches and tags in their
> namespace. This will allow all users to see the changes of other
> users, without needing direct access to their trees (which are
> inaccessible behind firewalls). [1]
Right, so thats github and/or git.or.cz.  Each user gets his/her own
> repository, but thats a very minor change.  Not a big deal.
> (b) fetch selected DAG, branch, and tag data of others to their tree, to see
> the changes of others (whether merged with head or not) while disconnected or
> remote.

This is also easy; you just establish remote tracking branches. I have a single shell scripted command, git-get-all, which pulls from all of the repositories I am interested in into various remote tracking branches so while I am disconnected, I can see what other folks have done on their trees.

> (c) grant and enforce permission for certain users to submit _merges
> only_ onto certain sub-portions of the "well-named branches"

This is the wierd one. *** Why ***? There is nothing magical about merges; all a merge is a commit that contains more than one parent. You can put anything into a merge, and in theory the result of a merge could have nothing to do with either parent. It would be a very perverse merge, but it's certainly possible. So what's the point of trying to enforce rules about "merges only"?

					- Ted
David Jeske· Jun 25, 2008, 05:30 UTC · re: Theodore Tso · lore
-- Theodore Tso wrote:
> Up to here, you can do this all with repo.or.cz, and/or github; you
> just give each developer their own repository, which they are allowed
> to push to, and no once else. Within their own repository they can
> make changes to their branches, so that all works just fine.

Yup. That's one of the reasons git is so attractive. There is some good stuff under "here" though....

Show 9 quoted lines
> > (a) safely "share" every DAG, branch, and tag data in their
> > repository to a well-connected server, into an established
> > namespace, while only changing branches and tags in their
> > namespace. This will allow all users to see the changes of other
> > users, without needing direct access to their trees (which are
> > inaccessible behind firewalls). [1]
>
> Right, so thats github and/or git.or.cz. Each user gets his/her own
> repository, but thats a very minor change. Not a big deal.

...most notably, all their DAGs in a single repository to save space is important. Thousands of copies of thousands of repositories adds up. Especially when most of the users who want to commit something probably commit <1-10k of unique stuff. Seems pretty easy to change though. git.or.cz and github will both be wanting this eventually.

The other big one is ACLs in 'well named' repositories, so multiple people can safely be allowed to add changes to them, without giving them ability to blow away the repository. I can see this isn't the way all git users work, but at least a few users working this way now with shared push repositories. This is just making it 'safer'. Also seems pretty easy to do.

> > (b) fetch selected DAG, branch, and tag data of others to their tree, to
see
> > the changes of others (whether merged with head or not) while disconnected
or
Show 7 quoted lines
> > remote.
>
> This is also easy; you just establish remote tracking branches. I
> have a single shell scripted command, git-get-all, which pulls from
> all of the repositories I am interested in into various remote
> tracking branches so while I am disconnected, I can see what other
> folks have done on their trees.

Yes, so I'd have the same thing, except instead of a remote repository, it would be a pattern of the branch namespace, such as /origin/users/jeske/*. It doesn't seem like the current remote tracking branch stuff can do this, but it would be easy to provide a client wrapper that would. Users who tracked the whole repository would just get everything, which is also fine. Maybe a client patch to make this better would be accepted.

Show 9 quoted lines
> > (c) grant and enforce permission for certain users to submit _merges
> > only_ onto certain sub-portions of the "well-named branches"
>
> This is the wierd one. *** Why ***? There is nothing magical about
> merges; all a merge is a commit that contains more than one parent.
> You can put anything into a merge, and in theory the result of a merge
> could have nothing to do with either parent. It would be a very
> perverse merge, but it's certainly possible. So what's the point of
> trying to enforce rules about "merges only"?

I'll explain why I wrote this, but I admit it's a strange roundabout way to get what I was hoping for. I hope there is a better way. One better way is to just change the client, but I was hoping not to have to do that. let me explain..

Think about using CVS. user does "cvs up; hack hack hack; cvs commit (to server)". In git, this workflow is "git pull; hack; commit; hack; commit; git push (to server)". I want those interum "commits" to share the changes with the server. I want to change this to "git pull; hack; commit-and-share; hack; commit-and-share; git-push (to shared branch tag)"

It would be nice if "commit-and-share" could just use "git-push". However, because users are going to do this habitually every commit, probably through a script or merged command, I didn't want users who are accidentally working directly in the master to accidentally fast-forward origin/master. (everyone seems to discourage working on master anyhow). I was hoping to enforce this only with server policy, so any git client works. That leaves me with the challenge of figuring out which commits on origin/master are actually intended to move the pointer, and which are accidents because someone forgot to branch before hacking in their client. One simple way to do this is to require any origin/master commit to have two children, one on the master, one somewhere else. If you have a commit that is directly hanging off of master in this design, you are doing the wrong thing. The server would tell you to "git checkout master; git branch -b mymaster; git reset origin/master; git push". This would put their local changes onto their private branch where they should be. When they wanted to do the equivilant of "cvs commit;" or current "git push;", they would do a merge to the master, and push again. The server would allow it, because it sees the merge.

I recognize this is a bit strange. I'd love to have a better solution, but this is the solution I can think of which only involves server enforcement. Other solutions I thought of would all require client changes that would change everyone's behavior. The candidate I liked best was: disallowing changes to tracking branches, including master, probably by implicitly creating a branch on commit to a tracking branch... However, I don't get the impression this will fit into current git very well, because users would need to turn their current "git push", into a "git merge master;git push"

I'm interested in other ideas to address this.

I know that all of what I wrote above seems strange if you don't buy into the design assumptions. That it's critical to share a single server-repository, that it's critical to have a shared 'well known' branch that only trusts clients to add new changes to, etc.. However, these are important.

Jakub Narebski· Jun 25, 2008, 09:30 UTC · re: David Jeske · lore
"David Jeske" <jeske@willowmail.com> writes:
Show 18 quoted lines
> -- Theodore Tso wrote:
> > ???
> > > 
> > > (a) safely "share" every DAG, branch, and tag data in their
> > > repository to a well-connected server, into an established
> > > namespace, while only changing branches and tags in their
> > > namespace. This will allow all users to see the changes of other
> > > users, without needing direct access to their trees (which are
> > > inaccessible behind firewalls). [1]
> >
> > Right, so thats github and/or git.or.cz. Each user gets his/her own
> > repository, but thats a very minor change. Not a big deal.
> 
> ...most notably, all their DAGs in a single repository to save space
> is important. Thousands of copies of thousands of repositories adds
> up. Especially when most of the users who want to commit something
> probably commit <1-10k of unique stuff. Seems pretty easy to change
> though. git.or.cz and github will both be wanting this eventually.

repo.or.cz has support for forks, i.e. sharing object database (for old objects) via alternates, although it is not "common object database" (as in, for example, $GIT_DIR/objects symlinked to single common parent repository)

GitHub has also some support for "forks", but as it is closed source I don't think anybody knows how it is done.

-- 
Jakub Narebski
Poland
ShadeHawk on #git
David Jeske· Aug 14, 2016, 00:43 UTC · re: Theodore Tso · lore
-- Theodore Tso wrote:
> Up to here, you can do this all with repo.or.cz, and/or github; you
> just give each developer their own repository, which they are allowed
> to push to, and no once else. Within their own repository they can
> make changes to their branches, so that all works just fine.

Yup. That's one of the reasons git is so attractive. There is some good stuff under "here" though....

Show 9 quoted lines
> > (a) safely "share" every DAG, branch, and tag data in their
> > repository to a well-connected server, into an established
> > namespace, while only changing branches and tags in their
> > namespace. This will allow all users to see the changes of other
> > users, without needing direct access to their trees (which are
> > inaccessible behind firewalls). [1]
>
> Right, so thats github and/or git.or.cz. Each user gets his/her own
> repository, but thats a very minor change. Not a big deal.

...most notably, all their DAGs in a single repository to save space is important. Thousands of copies of thousands of repositories adds up. Especially when most of the users who want to commit something probably commit <1-10k of unique stuff. Seems pretty easy to change though. git.or.cz and github will both be wanting this eventually.

The other big one is ACLs in 'well named' repositories, so multiple people can safely be allowed to add changes to them, without giving them ability to blow away the repository. I can see this isn't the way all git users work, but at least a few users working this way now with shared push repositories. This is just making it 'safer'. Also seems pretty easy to do.

> > (b) fetch selected DAG, branch, and tag data of others to their tree, to
see
> > the changes of others (whether merged with head or not) while disconnected
or
Show 7 quoted lines
> > remote.
>
> This is also easy; you just establish remote tracking branches. I
> have a single shell scripted command, git-get-all, which pulls from
> all of the repositories I am interested in into various remote
> tracking branches so while I am disconnected, I can see what other
> folks have done on their trees.

Yes, so I'd have the same thing, except instead of a remote repository, it would be a pattern of the branch namespace, such as /origin/users/jeske/*. It doesn't seem like the current remote tracking branch stuff can do this, but it would be easy to provide a client wrapper that would. Users who tracked the whole repository would just get everything, which is also fine. Maybe a client patch to make this better would be accepted.

Show 9 quoted lines
> > (c) grant and enforce permission for certain users to submit _merges
> > only_ onto certain sub-portions of the "well-named branches"
>
> This is the wierd one. *** Why ***? There is nothing magical about
> merges; all a merge is a commit that contains more than one parent.
> You can put anything into a merge, and in theory the result of a merge
> could have nothing to do with either parent. It would be a very
> perverse merge, but it's certainly possible. So what's the point of
> trying to enforce rules about "merges only"?

I'll explain why I wrote this, but I admit it's a strange roundabout way to get what I was hoping for. I hope there is a better way. One better way is to just change the client, but I was hoping not to have to do that. let me explain..

Think about using CVS. user does "cvs up; hack hack hack; cvs commit (to server)". In git, this workflow is "git pull; hack; commit; hack; commit; git push (to server)". I want those interum "commits" to share the changes with the server. I want to change this to "git pull; hack; commit-and-share; hack; commit-and-share; git-push (to shared branch tag)"

It would be nice if "commit-and-share" could just use "git-push". However, because users are going to do this habitually every commit, probably through a script or merged command, I didn't want users who are accidentally working directly in the master to accidentally fast-forward origin/master. (everyone seems to discourage working on master anyhow). I was hoping to enforce this only with server policy, so any git client works. That leaves me with the challenge of figuring out which commits on origin/master are actually intended to move the pointer, and which are accidents because someone forgot to branch before hacking in their client. One simple way to do this is to require any origin/master commit to have two children, one on the master, one somewhere else. If you have a commit that is directly hanging off of master in this design, you are doing the wrong thing. The server would tell you to "git checkout master; git branch -b mymaster; git reset origin/master; git push". This would put their local changes onto their private branch where they should be. When they wanted to do the equivilant of "cvs commit;" or current "git push;", they would do a merge to the master, and push again. The server would allow it, because it sees the merge.

I recognize this is a bit strange. I'd love to have a better solution, but this is the solution I can think of which only involves server enforcement. Other solutions I thought of would all require client changes that would change everyone's behavior. The candidate I liked best was: disallowing changes to tracking branches, including master, probably by implicitly creating a branch on commit to a tracking branch... However, I don't get the impression this will fit into current git very well, because users would need to turn their current "git push", into a "git merge master;git push"

I'm interested in other ideas to address this.

I know that all of what I wrote above seems strange if you don't buy into the design assumptions. That it's critical to share a single server-repository, that it's critical to have a shared 'well known' branch that only trusts clients to add new changes to, etc.. However, these are important.

Daniel Barkalow· Jun 25, 2008, 19:17 UTC · re: David Jeske · lore
On Wed, 25 Jun 2008, David Jeske wrote:
Show 6 quoted lines
> Yes, so I'd have the same thing, except instead of a remote repository, it
> would be a pattern of the branch namespace, such as /origin/users/jeske/*. It
> doesn't seem like the current remote tracking branch stuff can do this, but it
> would be easy to provide a client wrapper that would. Users who tracked the
> whole repository would just get everything, which is also fine. Maybe a client
> patch to make this better would be accepted.

Git actually has good support for large numbers of repositories sharing the same object storage. It's actually more efficient (in terms of server load) to have thousands of repositories with the same contents than one repository with thousands of branches.

Show 37 quoted lines
> > > (c) grant and enforce permission for certain users to submit _merges
> > > only_ onto certain sub-portions of the "well-named branches"
> >
> > This is the wierd one. *** Why ***? There is nothing magical about
> > merges; all a merge is a commit that contains more than one parent.
> > You can put anything into a merge, and in theory the result of a merge
> > could have nothing to do with either parent. It would be a very
> > perverse merge, but it's certainly possible. So what's the point of
> > trying to enforce rules about "merges only"?
> 
> I'll explain why I wrote this, but I admit it's a strange roundabout way to get
> what I was hoping for. I hope there is a better way. One better way is to just
> change the client, but I was hoping not to have to do that. let me explain..
> 
> Think about using CVS. user does "cvs up; hack hack hack; cvs commit (to
> server)". In git, this workflow is "git pull; hack; commit; hack; commit; git
> push (to server)". I want those interum "commits" to share the changes with the
> server. I want to change this to "git pull; hack; commit-and-share; hack;
> commit-and-share; git-push (to shared branch tag)"
> 
> It would be nice if "commit-and-share" could just use "git-push". However,
> because users are going to do this habitually every commit, probably through a
> script or merged command, I didn't want users who are accidentally working
> directly in the master to accidentally fast-forward origin/master. (everyone
> seems to discourage working on master anyhow). I was hoping to enforce this
> only with server policy, so any git client works. That leaves me with the
> challenge of figuring out which commits on origin/master are actually intended
> to move the pointer, and which are accidents because someone forgot to branch
> before hacking in their client. One simple way to do this is to require any
> origin/master commit to have two children, one on the master, one somewhere
> else. If you have a commit that is directly hanging off of master in this
> design, you are doing the wrong thing. The server would tell you to "git
> checkout master; git branch -b mymaster; git reset origin/master; git push".
> This would put their local changes onto their private branch where they should
> be. When they wanted to do the equivilant of "cvs commit;" or current "git
> push;", they would do a merge to the master, and push again. The server would
> allow it, because it sees the merge.

You have a fundamental misconception about git's data model. A commit doesn't have a particular branch it is on. There is only the DAG, where each node is a commit that is structured identically to all of the other commits. Branches pick out particular nodes in the DAG at particular times.

You can even think of there being a single theoretical universal DAG, independant of the actual development that gets done, and developers work to find the interesting portions, which are ones that contain trees that contain working code and useful messages and history that is informative. And they use branches to hold references to worthwhile parts of the DAG, and not (as in systems like SVN) to partition the DAG, which makes no reference to branches.

It therefore doesn't make any sense to ask if a commit is directly hanging off of master. If your local branch is up to date, and you commit, your commit's parent is the current master. If you now check out master and merge your local branch, master gets the same (non-merge) commit.

> I recognize this is a bit strange. I'd love to have a better solution, but this
> is the solution I can think of which only involves server enforcement.

You fundamentally can't do what you want with only server enforcement, because git doesn't provide the history of what local operations were used to prepare to ask the server to change something. It fundamentally can't, because there's no room in its data model of changes to hold that, and because its design is to allow flexibility in this preparation.

Show 7 quoted lines
> Other solutions I thought of would all require client changes that would 
> change everyone's behavior. The candidate I liked best was: disallowing 
> changes to tracking branches, including master, probably by implicitly 
> creating a branch on commit to a tracking branch... However, I don't get 
> the impression this will fit into current git very well, because users 
> would need to turn their current "git push", into a "git merge 
> master;git push"

Git prevents you from committing to tracking branches at all. Any branch you can commit to is inherently a local branch, because that's what it means for a branch to be local. The "push" operation updates a remote branch from a local branch.

Now, what might be good would be to introduce a type of ref that you can update with "merge" but not with "commit". Of course, this has to be client-side, because the final state doesn't depend on whether you commit in a temporary branch and merge into a publishing branch or commit directory in the publishing branch.

	-Daniel
*This .sig left intentionally blank*
Raimund Bauer· Jun 25, 2008, 20:12 UTC · re: Daniel Barkalow · lore
On Wed, 2008-06-25 at 15:17 -0400, Daniel Barkalow wrote:
Show 5 quoted lines
> You have a fundamental misconception about git's data model. A commit 
> doesn't have a particular branch it is on. There is only the DAG, where 
> each node is a commit that is structured identically to all of the other 
> commits. Branches pick out particular nodes in the DAG at particular 
> times.

But a branch in repository also has a local history. The ref-log. And git could use that to produce a distributed branch-history.

<wishful thinking>

A developer prepares a series of commits in a local branch to push to the server. On the server the ref-log of a branch gets updated with a new entry for each push, and other developers pulling from the server get the servers ref-log as ref-log of their remote tracking branch and can see the push-points there.

Those push-points seem to be somehow more important than other commits - there was a reason for the first developer to push right this branch tip, right? Seems like valuable (optional) information to me.

</wishful thinking>
> It therefore doesn't make any sense to ask if a commit is directly hanging 
> off of master. If your local branch is up to date, and you commit, your 
> commit's parent is the current master. If you now check out master and 
> merge your local branch, master gets the same (non-merge) commit.
Check if the commit is in master's ref-log?

regards, Ray

← back to recent threads