{"thread":{"id":"14144","subject":"Re: policy and mechanism for less-connected clients","startedAt":"2008-06-25T13:34:58Z","lastAt":"2008-06-25T20:54:12Z","messageCount":5,"participants":["Theodore Tso","Junio C Hamano","David Jeske","Jakub Narebski"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"81117","messageId":"20080625133458.GE20361@mit.edu","threadId":"14144","inReplyTo":null,"subject":"Re: policy and mechanism for less-connected clients","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-06-25T13:34:58Z","receivedAt":"2008-06-25T13:34:58Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Jun 25, 2008 at 05:20:49AM -0000, David Jeske wrote:\n> The other big one is ACLs in 'well named' repositories, so multiple\n> people can safely be allowed to add changes to them, without giving\n> them ability to blow away the repository. I can see this isn't the\n> way all git users work, but at least a few users working this way\n> now with shared push repositories. This is just making it\n> 'safer'. Also seems pretty easy to do.\n\nSo this isn't true security, since someone determined (or an ingenious\nenough fool) can always blow away repository if you allow them to add\nchanges; they could just add a change which rm's all of the files,\nyes?  You just want to prevent something stupid.\n\nWell, as long as they don't do non-fast forward updates (i.e., they\nnever do something like: \"git push publish +head:head\", or any other\nincantation involving a leading '+' in the refspec), they should be\npretty safe.  I don't see how they would do any damage just due to\nuser confusion.  So I think git is pretty safe as-is.\n\n> > This is also easy; you just establish remote tracking branches. I\n> > have a single shell scripted command, git-get-all, which pulls from\n> > all of the repositories I am interested in into various remote\n> > tracking branches so while I am disconnected, I can see what other\n> > folks have done on their trees.\n> \n> Yes, so I'd have the same thing, except instead of a remote\n> repository, it would be a pattern of the branch namespace, such as\n> /origin/users/jeske/*.\n\nAnd the advantage of using branch namespaces instead of separate\nremote repositories is.... ?  I don't see any....\n\n> Think about using CVS. user does \"cvs up; hack hack hack; cvs commit\n> (to server)\". In git, this workflow is \"git pull; hack; commit;\n> hack; commit; git push (to server)\". I want those interum \"commits\"\n> to share the changes with the server. I want to change this to \"git\n> pull; hack; commit-and-share; hack; commit-and-share; git-push (to\n> shared branch tag)\"\n\nOK, so *why* is it a good idea to ask people to share their\nin-progress work?  What's the upside?  Maybe if the idea is as backup\nif people are working from their laptops, and they're about to travel\ninternationally or some such, but in general, sharing in-progress work\nis highly overrated.\n\nThe other thing is in your design assumption is that remote\nrepositories are somehow expensive, when in fact they are very cheap;\nuse either repo.or.cz or github; they support repo sharing so there\nisn't major cost to letting each developer having their own repository\nto push to.\n\nSo the way I would do things is to simply encourage people to do start\ntheir work by branching off of an up-to-date master branch, but *not*\ndo any git pulls or git pushes.  They can use git commit as necessary\nto save interim work, and they do all of this work on a private\nbranch.  When they are done doing their work, they should review the\ngit commit points and make sure they make sense; in some cases they\nmay be better off squashing the commits down to a single commit, or\npossibly refactoring their work so that each individual commit is\nfree-standing, so that their series of commits is git-bisectable\n(i.e., after each commit the tree will fully compile and fully pass\nthe project regression test suite).\n\nOnce they have done *that*, they make sure the master branch has been\nfully updated, and then do a git-rebase on their feature branch so\nthat it is up-to-date with respect to master, and then they do a full\nbuild and regression test.  Then they switch back to the master\nbranch, and do a \"git push publish\" --- where <publish> is defined in\n.git/config to be something like this:\n\n[remote \"publish\"]\n\turl = ssh://master.kernel.org/pub/scm/linux/kernel/git/tytso/ext4.git\n\tpush = refs/heads/master:refs/heads/master\n\nThis will *only* push the master branch (and not any of the feature\nbranches), and it will not allow non-fast forward merges.  Hence, if\nthe user screwed up and accidentally made changes to the master branch\n(say, an accidental git-rebase while on the master branch, or\nsomething else bone-headed), the git push will fail.  This gives you\nthe safety you desire about not accidentally screwing up the master branch.\n\n\nAnd you're done.  The only reason why you need a per-user repository\nif you want some safety in terms of backups in case the work being\ndone on the laptop gets destroyed, but you can get that pretty much\nfor free via git.or.cz or github.  I really don't buy the sharing\nargument, because if you are in the middle of implementing a feature,\nit's generally not useful for others to look at your in-progress work.\n\n> I know that all of what I wrote above seems strange if you don't buy into the\n> design assumptions. That it's critical to share a single server-repository,\n> that it's critical to have a shared 'well known' branch that only trusts\n> clients to add new changes to, etc.. However, these are important.\n\nYep.  And you still haven't justified why it's critical to share a\nsingle server repository.  ***Why*** is that important?\n\nAnd when you have shared push repositories, as long as users don't use\nthe '+', in practice they can only add new changes.  And if you don't\ntrust them not to use the '+' character in refspecs, are you really\ngoing to trust them not to introduce either bone-headed mistakes into\nthe code?  Or to \"git rm\" the wrong files, git commit them, and then\nmerge that into the repository?  If all you care about is avoiding the\naccidentally stupid user mistakes, then putting in a convenience\ndefault so that \"git push publish\" always does what you want should be\ngood enough.\n\nSo fundamentally, yeah, I think your primary problem is with the\ndesign assumptions, which haven't been justified at all.\n\n       \t\t    \t  \t       \t\t     - Ted\n"},{"id":"81145","messageId":"7vwskd1kr7.fsf@gitster.siamese.dyndns.org","threadId":"14144","inReplyTo":"20080625133458.GE20361@mit.edu","subject":"Re: policy and mechanism for less-connected clients","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-06-25T17:34:36Z","receivedAt":"2008-06-25T17:34:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Theodore Tso <tytso@mit.edu> writes:\n\n> And when you have shared push repositories, as long as users don't use\n> the '+', in practice they can only add new changes.  And if you don't\n> trust them not to use the '+' character in refspecs, are you really\n> going to trust them not to introduce either bone-headed mistakes into\n> the code?\n\nWell, if you do not trust them, just set receive.denynonfastforwards\nand they won't be able to.\n"},{"id":"81182","messageId":"43260.7826347978$1214426654@news.gmane.org","threadId":"14144","inReplyTo":"20080625133458.GE20361@mit.edu","subject":"Re: policy and mechanism for less-connected clients","fromName":"David Jeske","fromEmail":"jeske@willowmail.com","sentAt":null,"receivedAt":"2008-06-25T20:38:22Z","isPatch":false,"sender":{"key":"jeske@willowmail.com","avatar":null},"body":"Thanks for the info about shared object storage for shared repositories. That's\ngreat, and looks like a good implementation method.\n\nPreviously I was thinking in terms of making a different server to change\nbehavior. However, I think the comments I've read are shifting my mindset\ntowards making a client-wrapper. I want to provide a system [wrapper] without\nthe user-burden of thinking about three repositories (local, my-public,\nshared-public). Doing this as a wrapper has other benefits, like the fact that\nusers can treat services like repo.or.cz as the \"networked filesystem of their\nversion control system\", so I like it.\n\nI have a model for the operations of this wrapper below.\n\n-- Theodore Tso wrote:\n> [snip] sharing in-progress work is highly overrated.\n\n_Seeing_ unfinished changes is overrated. However, so is managing multiple\nrepositories and managing which data is shared.\n\nI think my new wrapper approach below eliminates this overly-aggressive sharing\nwhile still reducing complexity for the average user.\n\n> So the way I would do things is to simply encourage people to do start\n> their work by branching off of an up-to-date master branch, but *not*\n> do any git pulls or git pushes.\n\nYou confused me here. If their repo.or.cz private repository is their only way\nof sharing (because their home directory is inaccessible and emailing patches\nis cumbersome), how do they exchange their own changes without pushing? Even in\na short time on git mailing list I see mini-unfinished-patches being posted.\n\n> [ description of commit rewriting, rebase, push ]\n\nThe method you describe is burdening all users with learning a bunch of new\nconcepts to do things that are unnecessary micromanagement for their needs. I'd\nprefer to give my users many of the benefits of DVCS/git with a\ncommand/argument set 1/20th the size and a much simpler mental model.\n\nMost of the software we're all using was developed while working with\ncentralized source control, where people just hack and commit and those commits\nare not even known-working. They don't bother with patch/commit rewriting and\nmanagement, and it works out just fine. I can see how that finer granularity\nmay be valuable for linux kernel coordinators. However, most projects don't\nneed to bother with all that, and even in the ones that do, most of their\ncontributors don't.\n\nDespite the success of centralized revision control, distributed source control\nrevision models have some very attractive features which can add efficiency to\na shared-central-repo model without straying far from the familiar (cvs up;\nhack; hack; cvs up; cvs commit;) workflow. I read some commentary from Linus\nthat compared git to a 'filesystem', and that's what I see.. a really awesome\nunderlying set of mechanisms for implementing SCM.. I'm trying to understand\nhow to layer an easy to use SCM system on top if it.\n\nSome 'git' users might say the right thing to do is do a different project, but\nI think, just like with the filesystem-analogy, there is significant benefit to\nsharing a single repository model so a simple source control system can then be\nused in powerful ways by powerful users. This is similar to the direction \"eg\"\n(easy git) is heading, but more extreme and extending to the server.\n\nIn fact, it seems like we might be better off if all of these source control\nuser-interfaces (cvs, perforce, git, eg, mercurial, etc. etc.) could be written\non top of a version-control-api that they shared. Witness the similar\nimplementation strategies of this modern rash of DVCS systems.\n\n--------------------------------------------------------------\n\nI'll try to explain my wrapper model in terms of an example... Imagine I'm\ngoing to deliver a \"cvs drop in replacement\", ncvs, that mostly keeps the cvs\nmental model, but is implemented underneath using git and just works better\nthan cvs (yet is simpler than git). I'll use the exact cvs command parameters\nfor illustration, but I wouldn't plan to do this. Notice how each ncvs command\nuses many git commands. It's possible these things should be done in terms of\nplumbing instead of porcelain to reduce dependence on git changes, but it's\nmore concise to express them as porcelain.\n\n>From the earlier feedback, there are now two repositories, one is considered\nthe \"shared-root\" while the other is the \"user\" repository.\n\n(1) make \"cvs update\" safe, make it easy to see granular comments for things\nyou have not pushed\n\nCVS users do potentially destructive merges all the time. Despite the way we\nuse terminology, working files ARE a branch, and \"cvs up\" IS a merge. That\nmerge can require edits to resolve, and after those edits are complete, the\nprevious state is NOT recoverable. There is no reason for this. We can easily\nsave the delta by just making \"cvs up\" equal \"git commit; git pull;\", or\nalternately, \"git stash; git pull; git apply;\".\n\n: \"ncvs up\" ->\n:\n: git stash; git pull; git apply;\n: git diff --stat <baseof:current branch> - un-pushed filenames\n: git-show-branch <current branch> - un-pushed comments\n\nQuestion: when I say \"baseof:current branch\", I mean \"the common-ancestor\nbetween my local-repo tracking branch and the remote-repo branch it's\ntracking\". How do I find that out?\n\nAdding \"git diff --stat <baseof:current branch>\" helps keep us aware of what\nchanges are in our local repo. Any files not pushed up to the branch head on\nthe server are seen. Likewise with \"git-show-branch <current branch>\" (which\nsomehow is not the same as git-show-branch --current).\n\n(2) make \"planned ahead of time\" branches cheap to make\n\n\"cvs up\" is the easiest merge in cvs, therefore, separate sets of checked out\nworking files become the most common form of branching in cvs. They are\nbasically personal work branches that you can't commit on, and can't\ncollaborate on. I've seen developers with cvs working directories weeks or\nmonths old because that's an easier way to work on different ideas than\ncreating a branch and checking them in. DVCS fixes this, by making branches\ncheap to make, and by making all branch merges closer to the simplicity of\ncvs's easy branch merge \"cvs up\". However, I don't need to burden the user with\nthe extra complexity and workload of the default being local branches, which\nthey then need to do more work to share. I want branches to be shared by\ndefault.\n\n: \"ncvs tag -b --shared $branch\" ->\n:\n: [ create a branch on the \"shared root\" repo, pointing\n:   to where I am in my local tree, if I have permission ]\n:  git branch --track $branch origin/$branch\n\n: \"ncvs tag -b mybranch\n:\n: [ create a branch on my \"user\" repo, pointing to where I am\nin my local tree, if I have permission ]\n: git branch --track $mybranch my-origin/$branch\n\nQuestion: I'm not sure what commands to use above. How do I create a branch on\na remote repo when I'm on my local machine, without sshing to it?\n\nThe advantages of git's repository over cvs's repository in this use-case are\nnot created because the branch is on the local machine. In fact, we also\ncreated it on the server. The benefit comes from the git revision storage model\nbeing faster and BETTER.\n\nThen to switch our working pointer to this branch, we might do:\n\n: \"ncvs up -r mybranch\" ->\n:\n: git stash; git checkout mybranch; git pull;\n: git stash show --relevant --recent;\n\nOur \"safe update\" automatically saved away any local directory changes before\nswitching off to the branch (if there were any). Our \"stash show\" is there\nalways to show us if any stashes hang off a recent parent of the tree we just\nswitched to, but it only shows them if they are hanging off this tree, and only\nif they are recent. If there is, we might want to look at or grab it, or we\nmight just ignore it and not care.\n\n(3) allow users to commit their 'final' changes to others (only on the branch\nthey are on)\n\n: \"ncvs commit\" -> \"git commit; git push <only this branch>;\"\n\nQuestion: how do I only push the branch I'm on? \"eg\" says it does this, but\nfrom a quick look at the code, it wasn't obvious to me how.\n\nDevelopers who are plenty happy with their existing model of never saving local\nchanges, can continue doing what they are doing. This makes the ability to save\nlocal changes an added benefit to the users like me that want to do it, instead\nof an extra burden to the other users. It also simplifies the issue of which\nchanges are pushed to the server and which are not, because pushing is managed\nby \"git push <only this branch>\", not by creating and managing local and remote\nbranch names separately. (easy git took the same approach with push)\n\n(4) Allow users to save interim changes, without ahead of time planning, ahead\nof time nameing, and hopefully, without naming at all.\n\nSaving interim changes in a cvs working tree before merging with head is not\ncheap. Making my own branch tag isn't too hard, but it takes a long time on a\nbig tree. Ironically, perforce made branching mechanism faster while making the\ncognitive load of branch hing much higher.\n\n: \"ncvs save\" -> \"git commit -a\"\n:\n: \"ncvs stash [$name]\" ->\n:\n: $currentbranch = `git branch`\n: $base-ish = '<baseof: current branch>'\n: git stash;\n: git branch -m $currentbranch $name;\n: git checkout $baseish;\n: git branch $currentbranch\n\nThis \"ncvs stash\" is acknowledging the value of the \"git stash\" idea, while\nalso recognizing that when I'm using \"git commit\" regularly, I don't have\nanything in the working set! I really want to stash the changes made since\n\"origin/<branchname>\" and return there with my local <branchname>. This is\nreally after the fact branch creation. If no $name is supplied, then it can\nauto-generate one like stash does.\n\n(5) make it obvious there is a difference between local and remote changes, but\nmake it easy to diff against remote before \"ncvs commit;\"\n\n: \"ncvs diff\" ->\n:\n: echo -n \"since commit(-C): \"  \\\n:   `git diff --shortstat <baseof:current branch>`; \\\n:   echo\n: echo -n \"since save(-S): \" \\\n:   `git diff`; echo\n:\n: \"ncvs diff -S\" -> \"git diff\"\n: \"ncvs diff -C\" -> \"git diff <baseof:current branch>\n--------------------------------------------------------------\n\nI'm primarily trying to understand how to map my model to git.\nContinued thanks for the discussion and help.\n"},{"id":"81186","messageId":"g3ub6v$7cu$1@ger.gmane.org","threadId":"14144","inReplyTo":"43260.7826347978$1214426654@news.gmane.org","subject":"Re: policy and mechanism for less-connected clients","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-06-25T20:52:47Z","receivedAt":"2008-06-25T20:52:47Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"<opublikowany i wysłany>\n\nDavid Jeske wrote:\n\n> Question: when I say \"baseof:current branch\", I mean \"the common-ancestor\n> between my local-repo tracking branch and the remote-repo branch it's\n> tracking\". How do I find that out?\n\ngit-merge-base\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"81187","messageId":"g3ub9l$b71$1@ger.gmane.org","threadId":"14144","inReplyTo":"43260.7826347978$1214426654@news.gmane.org","subject":"Re: policy and mechanism for less-connected clients","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-06-25T20:54:12Z","receivedAt":"2008-06-25T20:54:12Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"David Jeske wrote:\n\n> Question: how do I only push the branch I'm on? \"eg\" says it does this, but\n> from a quick look at the code, it wasn't obvious to me how.\n\ngit push <remote> HEAD   # with current enough git\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"}]}