{"thread":{"id":"36631","subject":"optimising a push by fetching objects from nearby repos","startedAt":"2014-05-10T13:39:37Z","lastAt":"2014-05-12T01:50:59Z","messageCount":13,"participants":["Sitaram Chamarty","Duy Nguyen","brian m. carlson","milki","Junio C Hamano","Storm-Olsen, Marius"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"241244","messageId":"536E2C19.3000202@gmail.com","threadId":"36631","inReplyTo":null,"subject":"optimising a push by fetching objects from nearby repos","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2014-05-10T13:39:37Z","receivedAt":"2014-05-10T13:39:37Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"Hi,\n\nIs there a trick to optimising a push by telling the receiver to pick up\nmissing objects from some other repo on its own server, to cut down even\nmore on network traffic?\n\nSo, hypothetically,\n\n     git push user@host:repo1 --look-for-objects-in=repo2\n\nI'm aware of the alternates mechanism, but that makes the dependency on\nthe other repo sort-of permanent.  I'm looking for a temporary\ndependence, just for the duration of the push.  Naturally, the objects\nshould be brought into the target repo for that to happen, except that\nthis would be doing more from disk and less from the network.\n\nMy gut says this isn't possible, and I've searched enough to almost be\nsure, but before I give up, I wanted to ask.\n\nthanks\nsitaram\n\nMilki: I'm sure you won't mind the cc, since you know the context :-)\n"},{"id":"241246","messageId":"CACsJy8DJq2e149wH6cQb4p1fQx9hubKQvxToMU8LLO+UmP+XmA@mail.gmail.com","threadId":"36631","inReplyTo":"536E2C19.3000202@gmail.com","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2014-05-10T13:54:53Z","receivedAt":"2014-05-10T13:54:53Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Sat, May 10, 2014 at 8:39 PM, Sitaram Chamarty <sitaramc@gmail.com> wrote:\n> Hi,\n>\n> Is there a trick to optimising a push by telling the receiver to pick up\n> missing objects from some other repo on its own server, to cut down even\n> more on network traffic?\n>\n> So, hypothetically,\n>\n>     git push user@host:repo1 --look-for-objects-in=repo2\n>\n> I'm aware of the alternates mechanism, but that makes the dependency on\n> the other repo sort-of permanent.  I'm looking for a temporary\n> dependence, just for the duration of the push.  Naturally, the objects\n> should be brought into the target repo for that to happen, except that\n> this would be doing more from disk and less from the network.\n>\n> My gut says this isn't possible, and I've searched enough to almost be\n> sure, but before I give up, I wanted to ask.\n\nMy feeling is it is possible, assuming that the target sees and reuses\nobjects from repo2 already. Injecting an alternate repo at runtime\nshould be possible. We exclude objects from sending at commit level,\nnot object level. So after the initial exclusion, we may need to run\nthe to-be-sent objects against the alternate repo to skip some more,\nbut that should not cost much if repo2 is fully packed. The receiver\nalways does the connectivity test. So if you make a mistake and\nspecify repo3 instead, the receiver will reject the push and the\ntarget repo won't be corrupted.\n-- \nDuy\n"},{"id":"241248","messageId":"20140510172338.GB45511@vauxhall.crustytoothpaste.net","threadId":"36631","inReplyTo":"536E2C19.3000202@gmail.com","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2014-05-10T17:23:39Z","receivedAt":"2014-05-10T17:23:39Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sat, May 10, 2014 at 07:09:37PM +0530, Sitaram Chamarty wrote:\n> Hi,\n> \n> Is there a trick to optimising a push by telling the receiver to pick up\n> missing objects from some other repo on its own server, to cut down even\n> more on network traffic?\n> \n> So, hypothetically,\n> \n>     git push user@host:repo1 --look-for-objects-in=repo2\n> \n> I'm aware of the alternates mechanism, but that makes the dependency on\n> the other repo sort-of permanent.  I'm looking for a temporary\n> dependence, just for the duration of the push.  Naturally, the objects\n> should be brought into the target repo for that to happen, except that\n> this would be doing more from disk and less from the network.\n> \n> My gut says this isn't possible, and I've searched enough to almost be\n> sure, but before I give up, I wanted to ask.\n\nI don't believe this is possible.  There has been some discussion on\nrelated matters at least fairly recently, though.\n\nPart of the reason nobody has implemented this is because it exposes\nadditional security concerns.  If I create a commit that references an\nobject I don't own, but is in someone else's repository, this feature\ncould allow me to gain access to objects which I shouldn't have access\nto unless the authentication and permissions layer is very, very\ncareful.  This would make many very simple HTTPS and SSH setups much\nmore complex.  Alternates don't have this problem because they're done\nserver-side.\n\nI definitely understand the desire for this, though.  I would probably\nuse it myself if it were available.\n\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | http://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: RSA v4 4096b: 88AC E9B2 9196 305B A994 7552 F1BA 225C 0223 B187\n"},{"id":"241249","messageId":"20140510173226.GA27483@hal.rescomp.berkeley.edu","threadId":"36631","inReplyTo":"20140510172338.GB45511@vauxhall.crustytoothpaste.net","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"milki","fromEmail":"milki@rescomp.berkeley.edu","sentAt":"2014-05-10T17:32:26Z","receivedAt":"2014-05-10T17:32:26Z","isPatch":false,"sender":{"key":"milki@rescomp.berkeley.edu","avatar":"https://gravatar.com/avatar/dbe8b44fae13d8846b110e1bd18a9f61f55089b728372f02a424bee1f825308f?d=mp&s=160"},"body":"On 17:23 Sat 10 May     , brian m. carlson wrote:\n> I don't believe this is possible.  There has been some discussion on\n> related matters at least fairly recently, though.\n> \n> Part of the reason nobody has implemented this is because it exposes\n> additional security concerns.  If I create a commit that references an\n> object I don't own, but is in someone else's repository, this feature\n> could allow me to gain access to objects which I shouldn't have access\n> to unless the authentication and permissions layer is very, very\n> careful.  This would make many very simple HTTPS and SSH setups much\n> more complex.  Alternates don't have this problem because they're done\n> server-side.\n\nIf this were implemented service side and specified with, say, a config\noption, would this security concern go away?\n\n-- \nmilki\n"},{"id":"241251","messageId":"20140510200459.GC45511@vauxhall.crustytoothpaste.net","threadId":"36631","inReplyTo":"20140510173226.GA27483@hal.rescomp.berkeley.edu","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2014-05-10T20:04:59Z","receivedAt":"2014-05-10T20:04:59Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sat, May 10, 2014 at 10:32:26AM -0700, milki wrote:\n> On 17:23 Sat 10 May     , brian m. carlson wrote:\n> > I don't believe this is possible.  There has been some discussion on\n> > related matters at least fairly recently, though.\n> > \n> > Part of the reason nobody has implemented this is because it exposes\n> > additional security concerns.  If I create a commit that references an\n> > object I don't own, but is in someone else's repository, this feature\n> > could allow me to gain access to objects which I shouldn't have access\n> > to unless the authentication and permissions layer is very, very\n> > careful.  This would make many very simple HTTPS and SSH setups much\n> > more complex.  Alternates don't have this problem because they're done\n> > server-side.\n> \n> If this were implemented service side and specified with, say, a config\n> option, would this security concern go away?\n\nIt would probably be fine if it were a config option.  I'd prefer it be\noff by default, though, to prevent surprises.\n\nThe attack scenario I'm thinking of is where you have several different\nusers, but the web server runs as one system user.  So /git/bmc/foo.git\nis owned by bmc, and /git/alice/bar.git is owned by alice.  The web\nserver will check authentication based on the path, and approve or deny\nit.  If it's approved, it will invoke the git daemon as a CGI script.\n\nBut the git daemon itself only knows that it was authenticated as a\ngiven user, and knows nothing about what the permissions scheme is.  So\nit will blithely let me refer to any other repository and import its\ndata if the option is enabled.  The web server only considered the path\nit was fed, so it couldn't have blocked this.\n\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | http://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: RSA v4 4096b: 88AC E9B2 9196 305B A994 7552 F1BA 225C 0223 B187\n"},{"id":"241253","messageId":"xmqqtx8xuz3b.fsf@gitster.dls.corp.google.com","threadId":"36631","inReplyTo":"536E2C19.3000202@gmail.com","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-05-10T21:02:16Z","receivedAt":"2014-05-10T21:02:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sitaram Chamarty <sitaramc@gmail.com> writes:\n\n> Is there a trick to optimising a push by telling the receiver to pick up\n> missing objects from some other repo on its own server, to cut down even\n> more on network traffic?\n>\n> So, hypothetically,\n>\n>     git push user@host:repo1 --look-for-objects-in=repo2\n>\n> I'm aware of the alternates mechanism, but that makes the dependency on\n> the other repo sort-of permanent.\n\nIn the direction of fetching, this may be give a good starting point.\n\n    http://thread.gmane.org/gmane.comp.version-control.git/243918/focus=245397\n\nIn the direction of pushing, theoretically you could:\n\n - define a new capability \"look-for-objects-in\" to pass the name of\n   the repository from \"git push\" to the \"receive-pack\";\n\n - have \"receive-pack\" temporarily borrow from the named repository\n   (if the policy on the server side allows it), and accept the push;\n\n - repack in order to dissociate the receiving repository from the\n   other repository it temporarily borrowed from.\n\nwhich would be the natural inverse of the approach suggested in the\n\"Can I borrow just temporarily while cloning?\" thread.\n\nBut I haven't thought things through with respect to what else need\nto be modified to make sure this does not have adverse interaction\nwith simultaneous pushes into the same repository, which would make\nit harder to solve for \"receive-pack\" than for \"clone/fetch\".\n"},{"id":"241255","messageId":"536ECC93.1070102@gmail.com","threadId":"36631","inReplyTo":"xmqqtx8xuz3b.fsf@gitster.dls.corp.google.com","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2014-05-11T01:04:19Z","receivedAt":"2014-05-11T01:04:19Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"On 05/11/2014 02:32 AM, Junio C Hamano wrote:\n> Sitaram Chamarty <sitaramc@gmail.com> writes:\n>\n>> Is there a trick to optimising a push by telling the receiver to pick up\n>> missing objects from some other repo on its own server, to cut down even\n>> more on network traffic?\n>>\n>> So, hypothetically,\n>>\n>>      git push user@host:repo1 --look-for-objects-in=repo2\n>>\n>> I'm aware of the alternates mechanism, but that makes the dependency on\n>> the other repo sort-of permanent.\n>\n> In the direction of fetching, this may be give a good starting point.\n>\n>      http://thread.gmane.org/gmane.comp.version-control.git/243918/focus=245397\n\nThat's an interesting thread and it's recent too.  However, it's about\nclone (though the intro email mentions other commands also).\n\nI'm specifically interested in push efficiency right now.  When you\n\"fork\" someone's repo to your own space, and you push your fork to the\nsame server, it ought to be able to get most of the common objects from\ndisk (specifically, from the repo you forked), and only what extra you\ndid from the network.\n\nClones do have a workaround (clone with --reference, then repack, as you\nsaid in that thread), but no such workaround exists for push.\n\n> In the direction of pushing, theoretically you could:\n>\n>   - define a new capability \"look-for-objects-in\" to pass the name of\n>     the repository from \"git push\" to the \"receive-pack\";\n>\n>   - have \"receive-pack\" temporarily borrow from the named repository\n>     (if the policy on the server side allows it), and accept the push;\n>\n>   - repack in order to dissociate the receiving repository from the\n>     other repository it temporarily borrowed from.\n>\n> which would be the natural inverse of the approach suggested in the\n> \"Can I borrow just temporarily while cloning?\" thread.\n>\n> But I haven't thought things through with respect to what else need\n> to be modified to make sure this does not have adverse interaction\n> with simultaneous pushes into the same repository, which would make\n> it harder to solve for \"receive-pack\" than for \"clone/fetch\".\n\nI'll leave it in your capable hands :-)  My C coding days are long gone!\n\nI do have a way to do this in gitolite (haven't coded it yet; just\nthinking).  Gitolite lets you specify something to do before git-*-pack\nruns, and I was planning something like this:\n\nterminology: borrow, borrower repo, reference repo\n\n\"borrow = relaxed\" mode\n\n     1.  check if the user has read access to the reference repo; skip\n         the rest of this if he doesn't\n\n     2.  from reference repo's \"objects\", find all directories and\n         \"mkdir\" them into borrower's objects directory, then find all\n         files and \"ln\" (hardlink) them. This is presumably what \"clone\n         -l\" does.\n\n     This method is close to constant time since we're not copying\n     objects.\n\n     It has the potential issue that if an object existed in the\n     reference repo that was subsequently *deleted* (say, a commit that\n     contained a password, which was quickly overwritten when\n     discovered), and the attacker knows the SHA, he can get the commit\n     out by sending an commit that depends on it, then fetching it back.\n\n     (He could do that to the reference repo directly if he had write\n     access, but we'll assume he doesn't, so this *is* a possible\n     attack).\n\n\"borrow = strict\" mode\n\n     1.  (same as for \"relaxed\" mode)\n\n     2.  actually *fetch* all refs from the reference repo to the\n         borrower (into, say, 'refs/borrowed'), then delete all those\n         refs so you just have the objects now.\n\n     Unlike the previous method, this takes time proportional to the\n     delta between borrower and reference, and may load the system a bit,\n     but unless the reference repo is highly volatile, this will settle\n     down. The point is that it cannot be used to get anything that the\n     user doesn't already have access to anyway.\n\nI still have to try it, but it sounds like both these would work.\n\nI'd appreciate any comments though...\n\nregards\nsitaram\n"},{"id":"241256","messageId":"1399772049733.13154@student.bi.no","threadId":"36631","inReplyTo":"536ECC93.1070102@gmail.com","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"Storm-Olsen, Marius","fromEmail":"marius.storm-olsen@student.bi.no","sentAt":"2014-05-11T01:34:10Z","receivedAt":"2014-05-11T01:34:10Z","isPatch":false,"sender":{"key":"marius.storm-olsen@student.bi.no","avatar":null},"body":"On 5/10/2014 8:04 PM, Sitaram Chamarty wrote:\n> On 05/11/2014 02:32 AM, Junio C Hamano wrote: That's an interesting\n> thread and it's recent too.  However, it's about clone (though the\n> intro email mentions other commands also).\n>\n> I'm specifically interested in push efficiency right now.  When you\n> \"fork\" someone's repo to your own space, and you push your fork to\n> the same server, it ought to be able to get most of the common\n> objects from disk (specifically, from the repo you forked), and only\n> what extra you did from the network.\n>\n...\n>\n> I do have a way to do this in gitolite (haven't coded it yet; just\n> thinking).  Gitolite lets you specify something to do before\n> git-*-pack runs, and I was planning something like this:\n\nAnd here you're poking the stick at the real solution to your problem.\n\nMany of the Git repo managers will neatly set up a server-side repo \nclone for you, with alternates into the original repo saving both \nnetwork and disk I/O.\n\nSo your work flow would instead be:\n   1. Fork repo on server\n   2. Remotely clone your own forked repo\n\nI think it's more appropriate to handle this higher level operation \nwithin the security context of a git repo manager, rather than directly \nin git.\n\n-- \n.marius\n"},{"id":"241257","messageId":"536EDC1C.5040101@gmail.com","threadId":"36631","inReplyTo":"1399772049733.13154@student.bi.no","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2014-05-11T02:10:36Z","receivedAt":"2014-05-11T02:10:36Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"On 05/11/2014 07:04 AM, Storm-Olsen, Marius wrote:\n> On 5/10/2014 8:04 PM, Sitaram Chamarty wrote:\n>> On 05/11/2014 02:32 AM, Junio C Hamano wrote: That's an interesting\n>> thread and it's recent too.  However, it's about clone (though the\n>> intro email mentions other commands also).\n>>\n>> I'm specifically interested in push efficiency right now.  When you\n>> \"fork\" someone's repo to your own space, and you push your fork to\n>> the same server, it ought to be able to get most of the common\n>> objects from disk (specifically, from the repo you forked), and only\n>> what extra you did from the network.\n>>\n> ...\n>>\n>> I do have a way to do this in gitolite (haven't coded it yet; just\n>> thinking).  Gitolite lets you specify something to do before\n>> git-*-pack runs, and I was planning something like this:\n>\n> And here you're poking the stick at the real solution to your problem.\n>\n> Many of the Git repo managers will neatly set up a server-side repo\n> clone for you, with alternates into the original repo saving both\n> network and disk I/O.\n\nGitolite already has a \"fork\" command that does that (though it uses\n\"-l\", not alternates).  I specifically don't want to use alternates, and\nI also specifically am looking for something that activates on a push --\nin the situations I am looking to optimise, the clone already happened.\n\n> So your work flow would instead be:\n>     1. Fork repo on server\n>     2. Remotely clone your own forked repo\n>\n> I think it's more appropriate to handle this higher level operation\n> within the security context of a git repo manager, rather than directly\n> in git.\n\nYes, because of the \"read access\" check in my suggested procedure to\nhandle this.  (Otherwise this is as valid as the plan suggested for\nclone in Junior's email in [1]).\n\n[1]: http://thread.gmane.org/gmane.comp.version-control.git/243918/focus=245397\n\nI will certainly be doing this in gitolite.  The point of my post was to\nvalidate the flow with the *git* experts in case they catch something I\nmissed, not to say \"this should be done *in* git\".\n"},{"id":"241258","messageId":"1399777917522.41294@student.bi.no","threadId":"36631","inReplyTo":"536EDC1C.5040101@gmail.com","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"Storm-Olsen, Marius","fromEmail":"marius.storm-olsen@student.bi.no","sentAt":"2014-05-11T03:11:57Z","receivedAt":"2014-05-11T03:11:57Z","isPatch":false,"sender":{"key":"marius.storm-olsen@student.bi.no","avatar":null},"body":"On 5/10/2014 9:10 PM, Sitaram Chamarty wrote:\n> On 05/11/2014 07:04 AM, Storm-Olsen, Marius wrote:\n>> On 5/10/2014 8:04 PM, Sitaram Chamarty wrote: Many of the Git repo\n>> managers will neatly set up a server-side repo clone for you, with\n>> alternates into the original repo saving both network and disk\n>> I/O.\n>\n> Gitolite already has a \"fork\" command that does that (though it uses\n> \"-l\", not alternates).  I specifically don't want to use alternates,\n> and I also specifically am looking for something that activates on a\n> push -- in the situations I am looking to optimise, the clone already\n> happened.\n\nYou can probably get the managers to do a fork without alternates too.\n\nAlso, it doesn't matter if you have already cloned from the original \nrepo remotely. If you use the git manager to clone the original repo on \nthe server, and you push to your new repo, only your changes will go \nback over the wire. The git protocol will figure out only which objects \nare missing to complete the new HEAD, and send those.\n\nSo\n    1. Clone remote repo\n    2. Hack hack hack\n    3. Fork repo on server\n    4. Push changes to your own remote repo\nis equally efficient.\n\n\n>> So your work flow would instead be:\n>>     1. Fork repo on server\n>>     2. Remotely clone your own forked repo\n>>\n>> I think it's more appropriate to handle this higher level operation\n>> within the security context of a git repo manager, rather than directly\n>> in git.\n>\n> Yes, because of the \"read access\" check in my suggested procedure to\n> handle this.  (Otherwise this is as valid as the plan suggested for\n> clone in Junior's email in [1]).\n\nIt's similar, but security issues come into play due to the swapped \ndirection, which is why I think it's wrong to place it in the push \ncommand. Now, having the 'borrow' complement to 'reference' in Git seems \nlike a good idea, and should work for your case too, but IMO should be \nconfigured with in the security context of the repo manager, and not on \nan individual push. *shrug*\n\n\n> [1]:\n> http://thread.gmane.org/gmane.comp.version-control.git/243918/focus=245397\n>\n> I will certainly be doing this in gitolite.  The point of my post was to\n> validate the flow with the *git* experts in case they catch something I\n> missed, not to say \"this should be done *in* git\".\n\nAbsolutely, and I think that's how everyone perceived it :) It's a good \nidea, with some tweaks, I think.\n\n\n-- \n.marius\n"},{"id":"241259","messageId":"536F08C5.3010705@gmail.com","threadId":"36631","inReplyTo":"1399777917522.41294@student.bi.no","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2014-05-11T05:21:09Z","receivedAt":"2014-05-11T05:21:09Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"On 05/11/2014 08:41 AM, Storm-Olsen, Marius wrote:\n> On 5/10/2014 9:10 PM, Sitaram Chamarty wrote:\n\n>      1. Clone remote repo\n>      2. Hack hack hack\n>      3. Fork repo on server\n>      4. Push changes to your own remote repo\n> is equally efficient.\n\nYour suggestions are good for a manual setup where the target repo\ndoesn't already exist.\n\nBut what I was looking for was validation from git.git folks of the idea\nof replicating what \"git clone -l\" does, for an *existing* repo.\n\nFor example, I'm assuming that bringing in only the objects -- without\nany of the refs pointing to them, making them all dangling objects --\nwill still allow the optimisation to occur (i.e., git will still say \"oh\nyeah I have these objects, even if they're dangling so I won't ask for\nthem from the pusher\" and not \"oh these are dangling objects; so I don't\nrecognise them from this perspective -- you'll have to send me those\nagain\").\n\n[1]: for any gitolite-aware folks reading this: this involves mirroring,\nbringing a new mirror into play, normal repos, wild repos, and on and\non...\n"},{"id":"241265","messageId":"xmqqbnv4ur7t.fsf@gitster.dls.corp.google.com","threadId":"36631","inReplyTo":"536F08C5.3010705@gmail.com","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2014-05-11T18:04:38Z","receivedAt":"2014-05-11T18:04:38Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sitaram Chamarty <sitaramc@gmail.com> writes:\n\n> But what I was looking for was validation from git.git folks of the idea\n> of replicating what \"git clone -l\" does, for an *existing* repo.\n>\n> For example, I'm assuming that bringing in only the objects -- without\n> any of the refs pointing to them, making them all dangling objects --\n> will still allow the optimisation to occur (i.e., git will still say \"oh\n> yeah I have these objects, even if they're dangling so I won't ask for\n> them from the pusher\" and not \"oh these are dangling objects; so I don't\n> recognise them from this perspective -- you'll have to send me those\n> again\").\n\nSo here is an educated guess by a git.git folk.  I haven't read the\ncodepath for some time, so I may be missing some details:\n\n - The set of objects sent over the wire in \"push\" direction is\n   determined by the receiving end listing what it has to the\n   sending end, and then the sending end excluding what the\n   receiving end told that it already has.\n\n - The receiving end tells the sending end what it has by showing\n   the names of its refs and their values.\n\nHaving otherwise dangling objects in your object store alone will\nnot make them reachable from the refs shown to the sending end.  But\nthere is another trick the receiving end employes.\n\n - The receiving end also includes the refs and their values that\n   appear in the repository it borrows objects from its alternate\n   repositories, when it tells what objects it already has to the\n   sending end.\n\nSo what you \"assumed\" is not entirely correct---bringing in only the\nobjects will not give you any optimization.\n\nBut because we infer from the location of the object store\n(i.e. \"objects\" directory) where the refs that point at these\nborrowed objects exist (i.e. in \"../refs\" relative to that \"objects\"\ndirectory) in order to make sure that we do not have to say \"oh\nthese are dangling but we know their history is not broken\", we\nstill get the same optimisation.\n\nAt least, that is the theory ;-)\n"},{"id":"241273","messageId":"53702903.20904@gmail.com","threadId":"36631","inReplyTo":"xmqqbnv4ur7t.fsf@gitster.dls.corp.google.com","subject":"Re: optimising a push by fetching objects from nearby repos","fromName":"Sitaram Chamarty","fromEmail":"sitaramc@gmail.com","sentAt":"2014-05-12T01:50:59Z","receivedAt":"2014-05-12T01:50:59Z","isPatch":false,"sender":{"key":"sitaramc@gmail.com","avatar":"https://avatars.githubusercontent.com/u/43316?v=4"},"body":"On 05/11/2014 11:34 PM, Junio C Hamano wrote:\n> Sitaram Chamarty <sitaramc@gmail.com> writes:\n>\n>> But what I was looking for was validation from git.git folks of the idea\n>> of replicating what \"git clone -l\" does, for an *existing* repo.\n>>\n>> For example, I'm assuming that bringing in only the objects -- without\n>> any of the refs pointing to them, making them all dangling objects --\n>> will still allow the optimisation to occur (i.e., git will still say \"oh\n>> yeah I have these objects, even if they're dangling so I won't ask for\n>> them from the pusher\" and not \"oh these are dangling objects; so I don't\n>> recognise them from this perspective -- you'll have to send me those\n>> again\").\n>\n> So here is an educated guess by a git.git folk.  I haven't read the\n> codepath for some time, so I may be missing some details:\n>\n>   - The set of objects sent over the wire in \"push\" direction is\n>     determined by the receiving end listing what it has to the\n>     sending end, and then the sending end excluding what the\n>     receiving end told that it already has.\n>\n>   - The receiving end tells the sending end what it has by showing\n>     the names of its refs and their values.\n>\n> Having otherwise dangling objects in your object store alone will\n> not make them reachable from the refs shown to the sending end.  But\n> there is another trick the receiving end employes.\n>\n>   - The receiving end also includes the refs and their values that\n>     appear in the repository it borrows objects from its alternate\n>     repositories, when it tells what objects it already has to the\n>     sending end.\n>\n> So what you \"assumed\" is not entirely correct---bringing in only the\n> objects will not give you any optimization.\n>\n> But because we infer from the location of the object store\n> (i.e. \"objects\" directory) where the refs that point at these\n> borrowed objects exist (i.e. in \"../refs\" relative to that \"objects\"\n> directory) in order to make sure that we do not have to say \"oh\n> these are dangling but we know their history is not broken\", we\n> still get the same optimisation.\n\nThanks!\n\nEverything makes sense.  However, I'm not using the alternates\nmechanism.\n\nSince gitolite has the advantage of allowing me to do something before\nand something after the git-receive-pack, I'm fetching all the refs into\na temporary namespace before, and deleting all of them after.  So, just\nfor the duration of the push, the refs do exist, and optimisation (of\nnetwork traffic) therefore happens.\n\nIn addition, since I check that the user has read access to the lender\nrepo (and don't do this optimisation if he does not), there is -- by\ndefinition -- no security issue, in the sense that he cannot get\nanything from the lender repo that he could not have got directly.\n\nThanks for all your help again, especially the very clear explanation!\n\nregards\nsitaram\n"}]}