{"thread":{"id":"14133","subject":"Re: policy and mechanism for less-connected clients","startedAt":"2008-06-25T02:33:53Z","lastAt":"2016-08-14T00:43:11Z","messageCount":6,"participants":["Theodore Tso","David Jeske","Jakub Narebski","Daniel Barkalow","Raimund Bauer"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"81036","messageId":"20080625023352.GC20361@mit.edu","threadId":"14133","inReplyTo":null,"subject":"Re: policy and mechanism for less-connected clients","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-06-25T02:33:53Z","receivedAt":"2008-06-25T02:33:53Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Jun 25, 2008 at 12:36:03AM -0000, David Jeske wrote:\n> The purpose of this mechanism is to host a distributed source\n> repository in a world where most most developer contributors are\n> behind firewalls and do not have access to, or do not want to\n> configure a unix server, ftp, or ssh to possibly contribute to a\n> project. \n\n> design assumptions:\n> \n> - all developers are firewalled and can not be \"pulled\" from directly.\n> - there can be one or more well-connected servers which all users can access.\n> - .. but which they cannot have ssh, ftp, or other dangerous access to\n> - .. and whose protocol should be layered on http(s)\n> - there is a shared namespace for branches, and tags\n> - .. users are not-trusted to change the branches or tags of other users\n\nUp to here, you can do this all with repo.or.cz, and/or github; you\njust give each developer their own repository, which they are allowed\nto push to, and no once else.  Within their own repository they can\nmake changes to their branches, so that all works just fine.\n\n> (a) safely \"share\" every DAG, branch, and tag data in their\n> repository to a well-connected server, into an established\n> namespace, while only changing branches and tags in their\n> namespace. This will allow all users to see the changes of other\n> users, without needing direct access to their trees (which are\n> inaccessible behind firewalls). [1]\n\nRight, so thats github and/or git.or.cz.  Each user gets his/her own\n> repository, but thats a very minor change.  Not a big deal.\n\n> (b) fetch selected DAG, branch, and tag data of others to their tree, to see\n> the changes of others (whether merged with head or not) while disconnected or\n> remote.\n\nThis is also easy; you just establish remote tracking branches.  I\nhave a single shell scripted command, git-get-all, which pulls from\nall of the repositories I am interested in into various remote\ntracking branches so while I am disconnected, I can see what other\nfolks have done on their trees.\n\n> (c) grant and enforce permission for certain users to submit _merges\n> only_ onto certain sub-portions of the \"well-named branches\"\n\nThis is the wierd one.  *** Why ***?  There is nothing magical about\nmerges; all a merge is a commit that contains more than one parent.\nYou can put anything into a merge, and in theory the result of a merge\ncould have nothing to do with either parent.  It would be a very\nperverse merge, but it's certainly possible.  So what's the point of\ntrying to enforce rules about \"merges only\"?\n\n\t\t\t\t\t- Ted\n"},{"id":"81071","messageId":"10634.0258535512$1214372002@news.gmane.org","threadId":"14133","inReplyTo":"20080625023352.GC20361@mit.edu","subject":"Re: policy and mechanism for less-connected clients","fromName":"David Jeske","fromEmail":"jeske@willowmail.com","sentAt":null,"receivedAt":"2008-06-25T05:30:55Z","isPatch":false,"sender":{"key":"jeske@willowmail.com","avatar":null},"body":"-- Theodore Tso wrote:\n> Up to here, you can do this all with repo.or.cz, and/or github; you\n> just give each developer their own repository, which they are allowed\n> to push to, and no once else. Within their own repository they can\n> make changes to their branches, so that all works just fine.\n\nYup. That's one of the reasons git is so attractive. There is some good stuff\nunder \"here\" though....\n\n> > (a) safely \"share\" every DAG, branch, and tag data in their\n> > repository to a well-connected server, into an established\n> > namespace, while only changing branches and tags in their\n> > namespace. This will allow all users to see the changes of other\n> > users, without needing direct access to their trees (which are\n> > inaccessible behind firewalls). [1]\n>\n> Right, so thats github and/or git.or.cz. Each user gets his/her own\n> repository, but thats a very minor change. Not a big deal.\n\n...most notably, all their DAGs in a single repository to save space is\nimportant. Thousands of copies of thousands of repositories adds up. Especially\nwhen most of the users who want to commit something probably commit <1-10k of\nunique stuff. Seems pretty easy to change though. git.or.cz and github will\nboth be wanting this eventually.\n\nThe other big one is ACLs in 'well named' repositories, so multiple people can\nsafely be allowed to add changes to them, without giving them ability to blow\naway the repository. I can see this isn't the way all git users work, but at\nleast a few users working this way now with shared push repositories. This is\njust making it 'safer'. Also seems pretty easy to do.\n\n> > (b) fetch selected DAG, branch, and tag data of others to their tree, to\nsee\n> > the changes of others (whether merged with head or not) while disconnected\nor\n> > remote.\n>\n> This is also easy; you just establish remote tracking branches. I\n> have a single shell scripted command, git-get-all, which pulls from\n> all of the repositories I am interested in into various remote\n> tracking branches so while I am disconnected, I can see what other\n> folks have done on their trees.\n\nYes, so I'd have the same thing, except instead of a remote repository, it\nwould be a pattern of the branch namespace, such as /origin/users/jeske/*. It\ndoesn't seem like the current remote tracking branch stuff can do this, but it\nwould be easy to provide a client wrapper that would. Users who tracked the\nwhole repository would just get everything, which is also fine. Maybe a client\npatch to make this better would be accepted.\n\n> > (c) grant and enforce permission for certain users to submit _merges\n> > only_ onto certain sub-portions of the \"well-named branches\"\n>\n> This is the wierd one. *** Why ***? There is nothing magical about\n> merges; all a merge is a commit that contains more than one parent.\n> You can put anything into a merge, and in theory the result of a merge\n> could have nothing to do with either parent. It would be a very\n> perverse merge, but it's certainly possible. So what's the point of\n> trying to enforce rules about \"merges only\"?\n\nI'll explain why I wrote this, but I admit it's a strange roundabout way to get\nwhat I was hoping for. I hope there is a better way. One better way is to just\nchange the client, but I was hoping not to have to do that. let me explain..\n\nThink about using CVS. user does \"cvs up; hack hack hack; cvs commit (to\nserver)\". In git, this workflow is \"git pull; hack; commit; hack; commit; git\npush (to server)\". I want those interum \"commits\" to share the changes with the\nserver. I want to change this to \"git pull; hack; commit-and-share; hack;\ncommit-and-share; git-push (to shared branch tag)\"\n\nIt would be nice if \"commit-and-share\" could just use \"git-push\". However,\nbecause users are going to do this habitually every commit, probably through a\nscript or merged command, I didn't want users who are accidentally working\ndirectly in the master to accidentally fast-forward origin/master. (everyone\nseems to discourage working on master anyhow). I was hoping to enforce this\nonly with server policy, so any git client works. That leaves me with the\nchallenge of figuring out which commits on origin/master are actually intended\nto move the pointer, and which are accidents because someone forgot to branch\nbefore hacking in their client. One simple way to do this is to require any\norigin/master commit to have two children, one on the master, one somewhere\nelse. If you have a commit that is directly hanging off of master in this\ndesign, you are doing the wrong thing. The server would tell you to \"git\ncheckout master; git branch -b mymaster; git reset origin/master; git push\".\nThis would put their local changes onto their private branch where they should\nbe. When they wanted to do the equivilant of \"cvs commit;\" or current \"git\npush;\", they would do a merge to the master, and push again. The server would\nallow it, because it sees the merge.\n\nI recognize this is a bit strange. I'd love to have a better solution, but this\nis the solution I can think of which only involves server enforcement. Other\nsolutions I thought of would all require client changes that would change\neveryone's behavior. The candidate I liked best was: disallowing changes to\ntracking branches, including master, probably by implicitly creating a branch\non commit to a tracking branch... However, I don't get the impression this will\nfit into current git very well, because users would need to turn their current\n\"git push\", into a \"git merge master;git push\"\n\nI'm interested in other ideas to address this.\n\nI know that all of what I wrote above seems strange if you don't buy into the\ndesign assumptions. That it's critical to share a single server-repository,\nthat it's critical to have a shared 'well known' branch that only trusts\nclients to add new changes to, etc.. However, these are important.\n"},{"id":"81101","messageId":"m37icdkgkl.fsf@localhost.localdomain","threadId":"14133","inReplyTo":"10634.0258535512$1214372002@news.gmane.org","subject":"Re: policy and mechanism for less-connected clients","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-06-25T09:30:09Z","receivedAt":"2008-06-25T09:30:09Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"David Jeske\" <jeske@willowmail.com> writes:\n> -- Theodore Tso wrote:\n> > ???\n> > > \n> > > (a) safely \"share\" every DAG, branch, and tag data in their\n> > > repository to a well-connected server, into an established\n> > > namespace, while only changing branches and tags in their\n> > > namespace. This will allow all users to see the changes of other\n> > > users, without needing direct access to their trees (which are\n> > > inaccessible behind firewalls). [1]\n> >\n> > Right, so thats github and/or git.or.cz. Each user gets his/her own\n> > repository, but thats a very minor change. Not a big deal.\n> \n> ...most notably, all their DAGs in a single repository to save space\n> is important. Thousands of copies of thousands of repositories adds\n> up. Especially when most of the users who want to commit something\n> probably commit <1-10k of unique stuff. Seems pretty easy to change\n> though. git.or.cz and github will both be wanting this eventually.\n\nrepo.or.cz has support for forks, i.e. sharing object database (for\nold objects) via alternates, although it is not \"common object\ndatabase\" (as in, for example, $GIT_DIR/objects symlinked to single\ncommon parent repository)\n\nGitHub has also some support for \"forks\", but as it is closed source I\ndon't think anybody knows how it is done.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"81164","messageId":"alpine.LNX.1.00.0806251421520.19665@iabervon.org","threadId":"14133","inReplyTo":"willow-jeske-01l6Cy0dFEDjCVqc","subject":"Re: policy and mechanism for less-connected clients","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-06-25T19:17:29Z","receivedAt":"2008-06-25T19:17:29Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Wed, 25 Jun 2008, David Jeske wrote:\n\n> Yes, so I'd have the same thing, except instead of a remote repository, it\n> would be a pattern of the branch namespace, such as /origin/users/jeske/*. It\n> doesn't seem like the current remote tracking branch stuff can do this, but it\n> would be easy to provide a client wrapper that would. Users who tracked the\n> whole repository would just get everything, which is also fine. Maybe a client\n> patch to make this better would be accepted.\n\nGit actually has good support for large numbers of repositories sharing \nthe same object storage. It's actually more efficient (in terms of \nserver load) to have thousands of repositories with the same contents than \none repository with thousands of branches.\n\n> > > (c) grant and enforce permission for certain users to submit _merges\n> > > only_ onto certain sub-portions of the \"well-named branches\"\n> >\n> > This is the wierd one. *** Why ***? There is nothing magical about\n> > merges; all a merge is a commit that contains more than one parent.\n> > You can put anything into a merge, and in theory the result of a merge\n> > could have nothing to do with either parent. It would be a very\n> > perverse merge, but it's certainly possible. So what's the point of\n> > trying to enforce rules about \"merges only\"?\n> \n> I'll explain why I wrote this, but I admit it's a strange roundabout way to get\n> what I was hoping for. I hope there is a better way. One better way is to just\n> change the client, but I was hoping not to have to do that. let me explain..\n> \n> Think about using CVS. user does \"cvs up; hack hack hack; cvs commit (to\n> server)\". In git, this workflow is \"git pull; hack; commit; hack; commit; git\n> push (to server)\". I want those interum \"commits\" to share the changes with the\n> server. I want to change this to \"git pull; hack; commit-and-share; hack;\n> commit-and-share; git-push (to shared branch tag)\"\n> \n> It would be nice if \"commit-and-share\" could just use \"git-push\". However,\n> because users are going to do this habitually every commit, probably through a\n> script or merged command, I didn't want users who are accidentally working\n> directly in the master to accidentally fast-forward origin/master. (everyone\n> seems to discourage working on master anyhow). I was hoping to enforce this\n> only with server policy, so any git client works. That leaves me with the\n> challenge of figuring out which commits on origin/master are actually intended\n> to move the pointer, and which are accidents because someone forgot to branch\n> before hacking in their client. One simple way to do this is to require any\n> origin/master commit to have two children, one on the master, one somewhere\n> else. If you have a commit that is directly hanging off of master in this\n> design, you are doing the wrong thing. The server would tell you to \"git\n> checkout master; git branch -b mymaster; git reset origin/master; git push\".\n> This would put their local changes onto their private branch where they should\n> be. When they wanted to do the equivilant of \"cvs commit;\" or current \"git\n> push;\", they would do a merge to the master, and push again. The server would\n> allow it, because it sees the merge.\n\nYou have a fundamental misconception about git's data model. A commit \ndoesn't have a particular branch it is on. There is only the DAG, where \neach node is a commit that is structured identically to all of the other \ncommits. Branches pick out particular nodes in the DAG at particular \ntimes.\n\nYou can even think of there being a single theoretical universal DAG, \nindependant of the actual development that gets done, and developers work \nto find the interesting portions, which are ones that contain trees that \ncontain working code and useful messages and history that is informative. \nAnd they use branches to hold references to worthwhile parts of the DAG, \nand not (as in systems like SVN) to partition the DAG, which makes no \nreference to branches.\n\nIt therefore doesn't make any sense to ask if a commit is directly hanging \noff of master. If your local branch is up to date, and you commit, your \ncommit's parent is the current master. If you now check out master and \nmerge your local branch, master gets the same (non-merge) commit.\n\n> I recognize this is a bit strange. I'd love to have a better solution, but this\n> is the solution I can think of which only involves server enforcement.\n\nYou fundamentally can't do what you want with only server enforcement, \nbecause git doesn't provide the history of what local operations were used \nto prepare to ask the server to change something. It fundamentally can't, \nbecause there's no room in its data model of changes to hold that, and \nbecause its design is to allow flexibility in this preparation.\n\n> Other solutions I thought of would all require client changes that would \n> change everyone's behavior. The candidate I liked best was: disallowing \n> changes to tracking branches, including master, probably by implicitly \n> creating a branch on commit to a tracking branch... However, I don't get \n> the impression this will fit into current git very well, because users \n> would need to turn their current \"git push\", into a \"git merge \n> master;git push\"\n\nGit prevents you from committing to tracking branches at all. Any branch \nyou can commit to is inherently a local branch, because that's what it \nmeans for a branch to be local. The \"push\" operation updates a remote \nbranch from a local branch.\n\nNow, what might be good would be to introduce a type of ref that you can \nupdate with \"merge\" but not with \"commit\". Of course, this has to be \nclient-side, because the final state doesn't depend on whether you commit \nin a temporary branch and merge into a publishing branch or commit \ndirectory in the publishing branch.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"81177","messageId":"1214424774.6570.21.camel@doriath","threadId":"14133","inReplyTo":"alpine.LNX.1.00.0806251421520.19665@iabervon.org","subject":"Re: policy and mechanism for less-connected clients","fromName":"Raimund Bauer","fromEmail":"ray007@gmx.net","sentAt":"2008-06-25T20:12:54Z","receivedAt":"2008-06-25T20:12:54Z","isPatch":false,"sender":{"key":"ray007@gmx.net","avatar":null},"body":"On Wed, 2008-06-25 at 15:17 -0400, Daniel Barkalow wrote:\n\n> You have a fundamental misconception about git's data model. A commit \n> doesn't have a particular branch it is on. There is only the DAG, where \n> each node is a commit that is structured identically to all of the other \n> commits. Branches pick out particular nodes in the DAG at particular \n> times.\n\nBut a branch in repository also has a local history. The ref-log.\nAnd git could use that to produce a distributed branch-history.\n\n<wishful thinking>\n\nA developer prepares a series of commits in a local branch to push to\nthe server.\nOn the server the ref-log of a branch gets updated with a new entry for\neach push, and other developers pulling from the server get the servers\nref-log as ref-log of their remote tracking branch and can see the\npush-points there.\n\nThose push-points seem to be somehow more important than other commits -\nthere was a reason for the first developer to push right this branch\ntip, right?\nSeems like valuable (optional) information to me.\n\n﻿</wishful thinking>\n\n> It therefore doesn't make any sense to ask if a commit is directly hanging \n> off of master. If your local branch is up to date, and you commit, your \n> commit's parent is the current master. If you now check out master and \n> merge your local branch, master gets the same (non-merge) commit.\n\nCheck if the commit is in master's ref-log?\n\nregards,\nRay\n"},{"id":"299232","messageId":"willow-jeske-01l6Cy0dFEDjCVqc","threadId":"14133","inReplyTo":"20080625023352.GC20361@mit.edu","subject":"Re: policy and mechanism for less-connected clients","fromName":"David Jeske","fromEmail":"jeske@willowmail.com","sentAt":null,"receivedAt":"2016-08-14T00:43:11Z","isPatch":false,"sender":{"key":"jeske@willowmail.com","avatar":null},"body":"-- Theodore Tso wrote:\n> Up to here, you can do this all with repo.or.cz, and/or github; you\n> just give each developer their own repository, which they are allowed\n> to push to, and no once else. Within their own repository they can\n> make changes to their branches, so that all works just fine.\n\nYup. That's one of the reasons git is so attractive. There is some good stuff\nunder \"here\" though....\n\n> > (a) safely \"share\" every DAG, branch, and tag data in their\n> > repository to a well-connected server, into an established\n> > namespace, while only changing branches and tags in their\n> > namespace. This will allow all users to see the changes of other\n> > users, without needing direct access to their trees (which are\n> > inaccessible behind firewalls). [1]\n>\n> Right, so thats github and/or git.or.cz. Each user gets his/her own\n> repository, but thats a very minor change. Not a big deal.\n\n...most notably, all their DAGs in a single repository to save space is\nimportant. Thousands of copies of thousands of repositories adds up. Especially\nwhen most of the users who want to commit something probably commit <1-10k of\nunique stuff. Seems pretty easy to change though. git.or.cz and github will\nboth be wanting this eventually.\n\nThe other big one is ACLs in 'well named' repositories, so multiple people can\nsafely be allowed to add changes to them, without giving them ability to blow\naway the repository. I can see this isn't the way all git users work, but at\nleast a few users working this way now with shared push repositories. This is\njust making it 'safer'. Also seems pretty easy to do.\n\n> > (b) fetch selected DAG, branch, and tag data of others to their tree, to\nsee\n> > the changes of others (whether merged with head or not) while disconnected\nor\n> > remote.\n>\n> This is also easy; you just establish remote tracking branches. I\n> have a single shell scripted command, git-get-all, which pulls from\n> all of the repositories I am interested in into various remote\n> tracking branches so while I am disconnected, I can see what other\n> folks have done on their trees.\n\nYes, so I'd have the same thing, except instead of a remote repository, it\nwould be a pattern of the branch namespace, such as /origin/users/jeske/*. It\ndoesn't seem like the current remote tracking branch stuff can do this, but it\nwould be easy to provide a client wrapper that would. Users who tracked the\nwhole repository would just get everything, which is also fine. Maybe a client\npatch to make this better would be accepted.\n\n> > (c) grant and enforce permission for certain users to submit _merges\n> > only_ onto certain sub-portions of the \"well-named branches\"\n>\n> This is the wierd one. *** Why ***? There is nothing magical about\n> merges; all a merge is a commit that contains more than one parent.\n> You can put anything into a merge, and in theory the result of a merge\n> could have nothing to do with either parent. It would be a very\n> perverse merge, but it's certainly possible. So what's the point of\n> trying to enforce rules about \"merges only\"?\n\nI'll explain why I wrote this, but I admit it's a strange roundabout way to get\nwhat I was hoping for. I hope there is a better way. One better way is to just\nchange the client, but I was hoping not to have to do that. let me explain..\n\nThink about using CVS. user does \"cvs up; hack hack hack; cvs commit (to\nserver)\". In git, this workflow is \"git pull; hack; commit; hack; commit; git\npush (to server)\". I want those interum \"commits\" to share the changes with the\nserver. I want to change this to \"git pull; hack; commit-and-share; hack;\ncommit-and-share; git-push (to shared branch tag)\"\n\nIt would be nice if \"commit-and-share\" could just use \"git-push\". However,\nbecause users are going to do this habitually every commit, probably through a\nscript or merged command, I didn't want users who are accidentally working\ndirectly in the master to accidentally fast-forward origin/master. (everyone\nseems to discourage working on master anyhow). I was hoping to enforce this\nonly with server policy, so any git client works. That leaves me with the\nchallenge of figuring out which commits on origin/master are actually intended\nto move the pointer, and which are accidents because someone forgot to branch\nbefore hacking in their client. One simple way to do this is to require any\norigin/master commit to have two children, one on the master, one somewhere\nelse. If you have a commit that is directly hanging off of master in this\ndesign, you are doing the wrong thing. The server would tell you to \"git\ncheckout master; git branch -b mymaster; git reset origin/master; git push\".\nThis would put their local changes onto their private branch where they should\nbe. When they wanted to do the equivilant of \"cvs commit;\" or current \"git\npush;\", they would do a merge to the master, and push again. The server would\nallow it, because it sees the merge.\n\nI recognize this is a bit strange. I'd love to have a better solution, but this\nis the solution I can think of which only involves server enforcement. Other\nsolutions I thought of would all require client changes that would change\neveryone's behavior. The candidate I liked best was: disallowing changes to\ntracking branches, including master, probably by implicitly creating a branch\non commit to a tracking branch... However, I don't get the impression this will\nfit into current git very well, because users would need to turn their current\n\"git push\", into a \"git merge master;git push\"\n\nI'm interested in other ideas to address this.\n\nI know that all of what I wrote above seems strange if you don't buy into the\ndesign assumptions. That it's critical to share a single server-repository,\nthat it's critical to have a shared 'well known' branch that only trusts\nclients to add new changes to, etc.. However, these are important.\n"}]}