{"thread":{"id":"8170","subject":"Smart fetch via HTTP?","startedAt":"2007-05-15T20:10:06Z","lastAt":"2007-05-20T10:30:10Z","messageCount":47,"participants":["Jan Hudec","A Large Angry SCM","Shawn O. Pearce","Junio C Hamano","Martin Langhoff","Johannes Schindelin","Jakub Narebski","david@lang.hm","Nicolas Pitre","Matthieu Moy","Theodore Tso","Petr Baudis","Linus Torvalds","alan","Joel Becker"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"42247","messageId":"20070515201006.GD3653@efreet.light.src","threadId":"8170","inReplyTo":null,"subject":"Smart fetch via HTTP?","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2007-05-15T20:10:06Z","receivedAt":"2007-05-15T20:10:06Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"Hello,\n\nDid anyone already think about fetching over HTTP working similarly to the\nnative git protocol?\n\nThat is rather than reading the raw content of the repository, there would be\na CGI script (could be integrated to gitweb), that would negotiate what the\nclient needs and then generate and send a single pack with it.\n\nMercurial and bzr both have this option. It would IMO have three benefits:\n - Fast access for people behind paranoid firewalls, that only let http and\n   https (you can tunel anything through, but only to port 443) through.\n - Can be run on shared machine. If you have web space on machine shared\n   by many people, you can set up your own gitweb, but cannot/are not allowed\n   to start your own network server for git native protocol.\n - Less things to set up. If you are setting up gitweb anyway, you'd not need\n   to set up additional thing for providing fetch access.\n\nThan a question is how to implement it. The current protocol is stateful on\nboth sides, but the stateless nature of HTTP more or less requires the\nprotocol to be stateless on the server.\n\nI think it would be possible to use basically the same protocol as now, but\nmake it stateless for server. That is server first sends it's heads and than\nclient repeatedly sends all it's wants and some haves until the server acks\nall of them and sends the pack.\n\nAlternatively I am thinking about using Bloom filters (somebody came with\nsuch idea on the bzr list when I still followed it). It might be useful, as\nover HTTP we need to send as many haves as possible in one go.\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"42256","messageId":"464A3471.9070007@gmail.com","threadId":"8170","inReplyTo":"20070515201006.GD3653@efreet.light.src","subject":"Re: Smart fetch via HTTP?","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2007-05-15T22:30:09Z","receivedAt":"2007-05-15T22:30:09Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"Jan Hudec wrote:\n> Hello,\n> \n> Did anyone already think about fetching over HTTP working similarly to the\n> native git protocol?\n> \n> That is rather than reading the raw content of the repository, there would be\n> a CGI script (could be integrated to gitweb), that would negotiate what the\n> client needs and then generate and send a single pack with it.\n> \n> Mercurial and bzr both have this option. It would IMO have three benefits:\n>  - Fast access for people behind paranoid firewalls, that only let http and\n>    https (you can tunel anything through, but only to port 443) through.\n>  - Can be run on shared machine. If you have web space on machine shared\n>    by many people, you can set up your own gitweb, but cannot/are not allowed\n>    to start your own network server for git native protocol.\n>  - Less things to set up. If you are setting up gitweb anyway, you'd not need\n>    to set up additional thing for providing fetch access.\n> \n> Than a question is how to implement it. The current protocol is stateful on\n> both sides, but the stateless nature of HTTP more or less requires the\n> protocol to be stateless on the server.\n> \n> I think it would be possible to use basically the same protocol as now, but\n> make it stateless for server. That is server first sends it's heads and than\n> client repeatedly sends all it's wants and some haves until the server acks\n> all of them and sends the pack.\n> \n> Alternatively I am thinking about using Bloom filters (somebody came with\n> such idea on the bzr list when I still followed it). It might be useful, as\n> over HTTP we need to send as many haves as possible in one go.\n> \n\nBundles?\n\nClient POSTs it's ref set; server uses the ref set to generate and \nreturn the bundle.\n\nPush over http(s) could work the same...\n"},{"id":"42265","messageId":"20070515232946.GR3141@spearce.org","threadId":"8170","inReplyTo":"20070515201006.GD3653@efreet.light.src","subject":"Re: Smart fetch via HTTP?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-05-15T23:29:47Z","receivedAt":"2007-05-15T23:29:47Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Jan Hudec <bulb@ucw.cz> wrote:\n> Did anyone already think about fetching over HTTP working similarly to the\n> native git protocol?\n\nNo work has been done on this (that I know of) but I've discussed\nit to some extent with Simon 'corecode' Schubert on #git, and I\nthink he also brought it up on the mailing list not too long after.\n\nI've certainly thought about adding some sort of pack-objects\nfrontend into gitweb.cgi for this exact purpose.  It is really\nquite easy, except for the negotation of what the client has.  ;-)\n \n> Than a question is how to implement it. The current protocol is stateful on\n> both sides, but the stateless nature of HTTP more or less requires the\n> protocol to be stateless on the server.\n> \n> I think it would be possible to use basically the same protocol as now, but\n> make it stateless for server. That is server first sends it's heads and than\n> client repeatedly sends all it's wants and some haves until the server acks\n> all of them and sends the pack.\n\nI think Simon was talking about doubling the number of haves the\nclient sends in each request.  So the client POSTs initially all\nof its current refs; then current refs and their parents; then 4\ncommits back, then 8, etc.  The server replies to each POST request\nwith either a \"send more please\" or the packfile.\n\n-- \nShawn.\n"},{"id":"42273","messageId":"7vlkfphab9.fsf@assigned-by-dhcp.cox.net","threadId":"8170","inReplyTo":"20070515232946.GR3141@spearce.org","subject":"Re: Smart fetch via HTTP?","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-05-16T00:38:34Z","receivedAt":"2007-05-16T00:38:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> Jan Hudec <bulb@ucw.cz> wrote:\n>> Did anyone already think about fetching over HTTP working similarly to the\n>> native git protocol?\n>\n> No work has been done on this (that I know of) but I've discussed\n> it to some extent with Simon 'corecode' Schubert on #git, and I\n> think he also brought it up on the mailing list not too long after.\n>\n> I've certainly thought about adding some sort of pack-objects\n> frontend into gitweb.cgi for this exact purpose.  It is really\n> quite easy, except for the negotation of what the client has.  ;-)\n>  \n>> Than a question is how to implement it. The current protocol is stateful on\n>> both sides, but the stateless nature of HTTP more or less requires the\n>> protocol to be stateless on the server.\n>> \n>> I think it would be possible to use basically the same protocol as now, but\n>> make it stateless for server. That is server first sends it's heads and than\n>> client repeatedly sends all it's wants and some haves until the server acks\n>> all of them and sends the pack.\n>\n> I think Simon was talking about doubling the number of haves the\n> client sends in each request.  So the client POSTs initially all\n> of its current refs; then current refs and their parents; then 4\n> commits back, then 8, etc.  The server replies to each POST request\n> with either a \"send more please\" or the packfile.\n\nI kinda' like the bundle suggestion ;-)\n"},{"id":"42285","messageId":"46a038f90705152225y529c9db3x8615822e876c25a8@mail.gmail.com","threadId":"8170","inReplyTo":"20070515201006.GD3653@efreet.light.src","subject":"Re: Smart fetch via HTTP?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-05-16T05:25:29Z","receivedAt":"2007-05-16T05:25:29Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/16/07, Jan Hudec <bulb@ucw.cz> wrote:\n> Did anyone already think about fetching over HTTP working similarly to the\n> native git protocol?\n\nDo the indexes have enough info to use them with http ranges? It'd be\nchunkier than a smart protocol, but it'd still work with dumb servers.\n\ncheers,\n\n\nm\n"},{"id":"42310","messageId":"Pine.LNX.4.64.0705161232120.6410@racer.site","threadId":"8170","inReplyTo":"46a038f90705152225y529c9db3x8615822e876c25a8@mail.gmail.com","subject":"Re: Smart fetch via HTTP?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-16T11:33:18Z","receivedAt":"2007-05-16T11:33:18Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 16 May 2007, Martin Langhoff wrote:\n\n> On 5/16/07, Jan Hudec <bulb@ucw.cz> wrote:\n> > Did anyone already think about fetching over HTTP working similarly to the\n> > native git protocol?\n> \n> Do the indexes have enough info to use them with http ranges? It'd be\n> chunkier than a smart protocol, but it'd still work with dumb servers.\n\nIt would not be really performant, would it? Besides, not all Web servers \nspeak HTTP/1.1...\n\nCiao,\nDscho\n"},{"id":"42332","messageId":"46a038f90705161426n3b928086t2d3e68749557f866@mail.gmail.com","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705161232120.6410@racer.site","subject":"Re: Smart fetch via HTTP?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-05-16T21:26:44Z","receivedAt":"2007-05-16T21:26:44Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/16/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> On Wed, 16 May 2007, Martin Langhoff wrote:\n> > Do the indexes have enough info to use them with http ranges? It'd be\n> > chunkier than a smart protocol, but it'd still work with dumb servers.\n> It would not be really performant, would it? Besides, not all Web servers\n> speak HTTP/1.1...\n\nPerformant compared to downloading a huge packfile to get 10% of it?\nSure! It'd probably take a few trips, and you'd end up fetching 20% of\nthe file, still better than 100%.\n\n> Besides, not all Web servers speak HTTP/1.1...\n\nAre there any interesting webservers out there that don't? Hand-rolled\npurpose-built webservers often don't but those don't serve files, they\nserve web apps. When it comes to serving files, any webserver that is\nsupported (security-wise) these days is HTTP/1.1.\n\nAnd for services like SF.net it'd be a safe low-cpu way of serving git\nfiles. 'cause the git protocol is quite expensive server-side (io+cpu)\nas we've seen with kernel.org. Being really smart with a cgi is\nprobably going to be expensive too.\n\ncheers,\n\n\nm\n"},{"id":"42335","messageId":"f2fua2$e2s$1@sea.gmane.org","threadId":"8170","inReplyTo":"46a038f90705161426n3b928086t2d3e68749557f866@mail.gmail.com","subject":"Re: Smart fetch via HTTP?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-05-16T21:54:42Z","receivedAt":"2007-05-16T21:54:42Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Martin Langhoff wrote:\n\n> On 5/16/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>> On Wed, 16 May 2007, Martin Langhoff wrote:\n>> > Do the indexes have enough info to use them with http ranges? It'd be\n>> > chunkier than a smart protocol, but it'd still work with dumb servers.\n>> It would not be really performant, would it? Besides, not all Web servers\n>> speak HTTP/1.1...\n> \n> Performant compared to downloading a huge packfile to get 10% of it?\n> Sure! It'd probably take a few trips, and you'd end up fetching 20% of\n> the file, still better than 100%.\n\nThat's why you should have something akin to backup policy for pack files,\nlike daily packs, weekly packs, ..., and the rest, just for the dumb\nprotocols.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"42348","messageId":"Pine.LNX.4.64.0705170152470.6410@racer.site","threadId":"8170","inReplyTo":"46a038f90705161426n3b928086t2d3e68749557f866@mail.gmail.com","subject":"Re: Smart fetch via HTTP?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-17T00:52:55Z","receivedAt":"2007-05-17T00:52:55Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 May 2007, Martin Langhoff wrote:\n\n> On 5/16/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > On Wed, 16 May 2007, Martin Langhoff wrote:\n> > > Do the indexes have enough info to use them with http ranges? It'd be\n> > > chunkier than a smart protocol, but it'd still work with dumb servers.\n> > It would not be really performant, would it? Besides, not all Web servers\n> > speak HTTP/1.1...\n> \n> Performant compared to downloading a huge packfile to get 10% of it?\n> Sure! It'd probably take a few trips, and you'd end up fetching 20% of\n> the file, still better than 100%.\n\nDon't forget that those 10% probably do not do you the favour to be in \nlarge chunks. Chances are that _every_ _single_ wanted object is separate \nfrom the others.\n\n> > Besides, not all Web servers speak HTTP/1.1...\n> \n> Are there any interesting webservers out there that don't? Hand-rolled \n> purpose-built webservers often don't but those don't serve files, they \n> serve web apps. When it comes to serving files, any webserver that is \n> supported (security-wise) these days is HTTP/1.1.\n> \n> And for services like SF.net it'd be a safe low-cpu way of serving git\n> files. 'cause the git protocol is quite expensive server-side (io+cpu)\n> as we've seen with kernel.org. Being really smart with a cgi is\n> probably going to be expensive too.\n\nIt's probably better and faster than relying on a feature which does not \nexactly help.\n\nCiao,\nDscho\n"},{"id":"42349","messageId":"20070517010335.GU3141@spearce.org","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705170152470.6410@racer.site","subject":"Re: Smart fetch via HTTP?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-05-17T01:03:35Z","receivedAt":"2007-05-17T01:03:35Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> Don't forget that those 10% probably do not do you the favour to be in \n> large chunks. Chances are that _every_ _single_ wanted object is separate \n> from the others.\n\nThat's completely possible.  Assuming the objects even are packed\nin the first place.  Its very unlikely that you would be able to\nfetch very large of a range from an existing packfile, you would be\nsubmitting most of your range requests for very very small sections.\n \n> > And for services like SF.net it'd be a safe low-cpu way of serving git\n> > files. 'cause the git protocol is quite expensive server-side (io+cpu)\n> > as we've seen with kernel.org. Being really smart with a cgi is\n> > probably going to be expensive too.\n> \n> It's probably better and faster than relying on a feature which does not \n> exactly help.\n\nYes.  Packing more often and pack v4 may help a lot there.\n\nThe other thing is kernel.org should really try to encourage the\nfolks with repositories there to try and share against one master\nrepository, so the poor OS has a better chance at holding the bulk\nof linux-2.6.git in buffer cache.\n\nI'm not suggesting they share specifically against Linus' repository;\nmaybe hpa and the other admins can host one seperately from Linus and\nenourage users to use that repository when on a system they maintain.\n\nIn an SF.net type case this doesn't help however.  Most of SF.net\nis tiny projects with very few, if any, developers.  Hence most\nof that is going to be unsharable, infrequently accessed, and uh,\nnot needed to be stored in buffer cache.  For the few projects that\nare hosted there that have a large developer base they could use\na shared repository approach as I just suggested for kernel.org.\n\naka the \"forks\" thing in gitweb, and on repo.or.cz...\n\n-- \nShawn.\n"},{"id":"42350","messageId":"Pine.LNX.4.64.0705161803580.1280@asgard.lang.hm","threadId":"8170","inReplyTo":"20070517010335.GU3141@spearce.org","subject":"Re: Smart fetch via HTTP?","fromName":"","fromEmail":"david@lang.hm","sentAt":"2007-05-17T01:04:55Z","receivedAt":"2007-05-17T01:04:55Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Wed, 16 May 2007, Shawn O. Pearce wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>\n>>> And for services like SF.net it'd be a safe low-cpu way of serving git\n>>> files. 'cause the git protocol is quite expensive server-side (io+cpu)\n>>> as we've seen with kernel.org. Being really smart with a cgi is\n>>> probably going to be expensive too.\n>>\n>> It's probably better and faster than relying on a feature which does not\n>> exactly help.\n>\n> Yes.  Packing more often and pack v4 may help a lot there.\n>\n> The other thing is kernel.org should really try to encourage the\n> folks with repositories there to try and share against one master\n> repository, so the poor OS has a better chance at holding the bulk\n> of linux-2.6.git in buffer cache.\n\ndo you mean more precisely share against one object store or do you really \nmean repository?\n\nDavid Lang\n\n> I'm not suggesting they share specifically against Linus' repository;\n> maybe hpa and the other admins can host one seperately from Linus and\n> enourage users to use that repository when on a system they maintain.\n>\n> In an SF.net type case this doesn't help however.  Most of SF.net\n> is tiny projects with very few, if any, developers.  Hence most\n> of that is going to be unsharable, infrequently accessed, and uh,\n> not needed to be stored in buffer cache.  For the few projects that\n> are hosted there that have a large developer base they could use\n> a shared repository approach as I just suggested for kernel.org.\n>\n> aka the \"forks\" thing in gitweb, and on repo.or.cz...\n>\n>\n"},{"id":"42351","messageId":"20070517012602.GV3141@spearce.org","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705161803580.1280@asgard.lang.hm","subject":"Re: Smart fetch via HTTP?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-05-17T01:26:03Z","receivedAt":"2007-05-17T01:26:03Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"david@lang.hm wrote:\n> On Wed, 16 May 2007, Shawn O. Pearce wrote:\n> >\n> >The other thing is kernel.org should really try to encourage the\n> >folks with repositories there to try and share against one master\n> >repository, so the poor OS has a better chance at holding the bulk\n> >of linux-2.6.git in buffer cache.\n> \n> do you mean more precisely share against one object store or do you really \n> mean repository?\n\nSorry, I did mean \"object store\".  ;-)\n\nRepository is insanity, as the refs and tags namespaces are suddenly\nshared.  What a nightmare that would become.\n\n-- \nShawn.\n"},{"id":"42353","messageId":"20070517014542.GW3141@spearce.org","threadId":"8170","inReplyTo":"20070517012602.GV3141@spearce.org","subject":"Re: Smart fetch via HTTP?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-05-17T01:45:42Z","receivedAt":"2007-05-17T01:45:42Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> wrote:\n> david@lang.hm wrote:\n> > On Wed, 16 May 2007, Shawn O. Pearce wrote:\n> > >\n> > >The other thing is kernel.org should really try to encourage the\n> > >folks with repositories there to try and share against one master\n> > >repository, so the poor OS has a better chance at holding the bulk\n> > >of linux-2.6.git in buffer cache.\n> > \n> > do you mean more precisely share against one object store or do you really \n> > mean repository?\n> \n> Sorry, I did mean \"object store\".  ;-)\n\nAnd even there, I don't mean symlink objects to a shared database,\nI mean use the objects/info/alternates file to point to the shared,\nread-only location.\n\nIts not perfect.  The hotter parts of the object database is almost\nalways the recent stuff, as that's what people are actively trying\nto fetch, or are using as a base when they are trying to fetch from\nsomeone else.  The hotter parts are also probably too new to be\nin the shared store offered by kernel.org admins, which means you\ncannot get good IO buffering.  Back to the current set of problems.\n\nA single shared object directory that everyone can write new files\ninto, but cannot modify or delete from, would help that problem quite\na bit.  But it opens up huge problems about pruning, as there is no\nway to perform garbage collection on that database without scanning\nevery ref on the system, and that's just not simply possible on a\nbusy system like kernel.org.\n\n-- \nShawn.\n"},{"id":"42367","messageId":"alpine.LFD.0.99.0705162309310.24220@xanadu.home","threadId":"8170","inReplyTo":"20070517010335.GU3141@spearce.org","subject":"Re: Smart fetch via HTTP?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-05-17T03:45:30Z","receivedAt":"2007-05-17T03:45:30Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 16 May 2007, Shawn O. Pearce wrote:\n\n> Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > Don't forget that those 10% probably do not do you the favour to be in \n> > large chunks. Chances are that _every_ _single_ wanted object is separate \n> > from the others.\n> \n> That's completely possible.  Assuming the objects even are packed\n> in the first place.  Its very unlikely that you would be able to\n> fetch very large of a range from an existing packfile, you would be\n> submitting most of your range requests for very very small sections.\n\nWell, in the commit objects case you're likely to have a bunch of them \nall contigous.\n\nFor tree and blob objects it is less likely.\n\nAnd of course there is the question of deltas for which you might or \nmight not have the base object locally already.\n\nStill... I wonder if this could be actually workable.  A typical daily \nupdate on the Linux kernel repository might consist of a couple hundreds \nor a few tousands objects.  This could still be faster to fetch parts of \na pack than the whole pack if the size difference is above a certain \ntreshold.  It is certainly not worse than fetching loose objects.\n\nThings would be pretty horrid if you think of fetching a commit object, \nparsing it to find out what tree object to fetch, then parse that tree \nobject to find out what other objects to fetch, and so on.\n\nBut if you only take the approach of fetching the pack index files, \nfinding out about the objects that the remote has that are not available \nlocally, and then fetching all those objects from within pack files \nwithout even looking at them (except for deltas), then it should be \npossible to issue a couple requests in parallel and possibly have decent \nperformances.  And if it turns out that more than, say, 70% of a \nparticular pack is to be fetched (you can determine that up front), then \nit might be decided to fetch the whole pack.\n\nThere is no way to sensibly keep those objects packed on the receiving \nend of course, but storing them as loose objects and repacking them \nafterwards should be just fine.\n\nOf course you'll get objects from branches in the remote repository you \nmight not be interested in, but that's a price to pay for such a hack.  \nOn average the overhead shouldn't be that big anyway if branches within \na repository are somewhat related.\n\nI think this is something worth experimenting.\n\n\nNicolas\n"},{"id":"42384","messageId":"Pine.LNX.4.64.0705171143350.6410@racer.site","threadId":"8170","inReplyTo":"alpine.LFD.0.99.0705162309310.24220@xanadu.home","subject":"Re: Smart fetch via HTTP?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-17T10:48:47Z","receivedAt":"2007-05-17T10:48:47Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 16 May 2007, Nicolas Pitre wrote:\n\n> Still... I wonder if this could be actually workable.  A typical daily \n> update on the Linux kernel repository might consist of a couple hundreds \n> or a few tousands objects.  This could still be faster to fetch parts of \n> a pack than the whole pack if the size difference is above a certain \n> treshold.  It is certainly not worse than fetching loose objects.\n> \n> Things would be pretty horrid if you think of fetching a commit object, \n> parsing it to find out what tree object to fetch, then parse that tree \n> object to find out what other objects to fetch, and so on.\n> \n> But if you only take the approach of fetching the pack index files, \n> finding out about the objects that the remote has that are not available \n> locally, and then fetching all those objects from within pack files \n> without even looking at them (except for deltas), then it should be \n> possible to issue a couple requests in parallel and possibly have decent \n> performances.  And if it turns out that more than, say, 70% of a \n> particular pack is to be fetched (you can determine that up front), then \n> it might be decided to fetch the whole pack.\n> \n> There is no way to sensibly keep those objects packed on the receiving \n> end of course, but storing them as loose objects and repacking them \n> afterwards should be just fine.\n> \n> Of course you'll get objects from branches in the remote repository you \n> might not be interested in, but that's a price to pay for such a hack.  \n> On average the overhead shouldn't be that big anyway if branches within \n> a repository are somewhat related.\n> \n> I think this is something worth experimenting.\n\nI am a bit wary about that, because it is so complex. IMHO a cgi which \ngets, say, up to a hundred refs (maybe something like ref~0, ref~1, ref~2, \nref~4, ref~8, ref~16, ... for the refs), and then makes a bundle for that \ncase on the fly, is easier to do.\n\nOf course, as with all cgi scripts, you have to make sure that DOS attacks \nhave a low probability of success.\n\nCiao,\nDscho\n"},{"id":"42387","messageId":"vpq8xbnlmdv.fsf@bauges.imag.fr","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705170152470.6410@racer.site","subject":"Re: Smart fetch via HTTP?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-05-17T11:28:44Z","receivedAt":"2007-05-17T11:28:44Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Hi,\n>\n> On Thu, 17 May 2007, Martin Langhoff wrote:\n>\n>> On 5/16/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>> > On Wed, 16 May 2007, Martin Langhoff wrote:\n>> > > Do the indexes have enough info to use them with http ranges? It'd be\n>> > > chunkier than a smart protocol, but it'd still work with dumb servers.\n>> > It would not be really performant, would it? Besides, not all Web servers\n>> > speak HTTP/1.1...\n>> \n>> Performant compared to downloading a huge packfile to get 10% of it?\n>> Sure! It'd probably take a few trips, and you'd end up fetching 20% of\n>> the file, still better than 100%.\n>\n> Don't forget that those 10% probably do not do you the favour to be in \n> large chunks. Chances are that _every_ _single_ wanted object is separate \n> from the others.\n\nFYI, bzr uses HTTP range requests, and the introduction of this\nfeature lead to significant performance improvement for them (bzr is\nmore dumb-protocol oriented than git is, so that's really important\nthere). They have this \"index file+data file\" system too, so you\ndownload the full index file, and then send an HTTP range request to\nget only the relevant parts of the data file.\n\nThe thing is, AAUI, they don't send N range requests to get N chunks,\nbut one HTTP request, requesting the N ranges at a time, and get the N\nchunks a a whole (IIRC, a kind of MIME-encoded response from the\nserver). So, you pay the price of a longer HTTP request, but not the\nprice of N networks round-trips.\n\nThat's surely not as efficient as anything smart on the server, but\nmight really help for the cases where the server is /not/ smart.\n\n-- \nMatthieu\n"},{"id":"42389","messageId":"20070517123637.GA9514@thunk.org","threadId":"8170","inReplyTo":"20070517014542.GW3141@spearce.org","subject":"Re: Smart fetch via HTTP?","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-05-17T12:36:37Z","receivedAt":"2007-05-17T12:36:37Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, May 16, 2007 at 09:45:42PM -0400, Shawn O. Pearce wrote:\n> Its not perfect.  The hotter parts of the object database is almost\n> always the recent stuff, as that's what people are actively trying\n> to fetch, or are using as a base when they are trying to fetch from\n> someone else.  The hotter parts are also probably too new to be\n> in the shared store offered by kernel.org admins, which means you\n> cannot get good IO buffering.  Back to the current set of problems.\n\nActually, as long as objects/info/alternates is pointing at Linus's\nkernel.org tree, I would think that it should work relatively well,\nsince everyone is normally basing their work on top of his tree as a\nstarting point.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"42390","messageId":"20070517124006.GO4489@pasky.or.cz","threadId":"8170","inReplyTo":"20070515201006.GD3653@efreet.light.src","subject":"Re: Smart fetch via HTTP?","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2007-05-17T12:40:06Z","receivedAt":"2007-05-17T12:40:06Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"  Hi,\n\nOn Tue, May 15, 2007 at 10:10:06PM CEST, Jan Hudec wrote:\n> Did anyone already think about fetching over HTTP working similarly to the\n> That is rather than reading the raw content of the repository, there would be\n> a CGI script (could be integrated to gitweb), that would negotiate what the\n> client needs and then generate and send a single pack with it.\n\n  frankly, I'm not that excited. I'm not disputing that this would be\nuseful, but I have my doubts on just how *much* useful it would be - I'm\nnot so sure the set of users affected is really all that large. So I'm\njust cooling people down here. ;-))\n\n> Mercurial and bzr both have this option. It would IMO have three benefits:\n>  - Fast access for people behind paranoid firewalls, that only let http and\n>    https (you can tunel anything through, but only to port 443) through.\n\n  How many users really have this problem? I'm not so sure. There are\ncertainly some, but enough for this to be a viable argument?\n\n>  - Can be run on shared machine. If you have web space on machine shared\n>    by many people, you can set up your own gitweb, but cannot/are not allowed\n>    to start your own network server for git native protocol.\n\n  You need to have CGI-enabled hosting, set up the CGI script etc. -\noverally, the setup is similarly complicated as git-daemon setup, so\nit's not \"zero-setup\" solution anymore.\n\n  Again, I'm not sure just how many people are in the situation that\nthey can run real CGI (not just PHP) but not git-daemon.\n\n>  - Less things to set up. If you are setting up gitweb anyway, you'd not need\n>    to set up additional thing for providing fetch access.\n\n  Except, well, how do you \"set it up\"? You need to make sure\ngit-update-server-info is run, yes, but that shouldn't be a problem (I'm\nnot so sure if git does this for you automagically - Cogito would...).\n\n  I think 95% of people don't set up gitweb.cgi either for their small\nHTTP repositories. :-)\n\n  Then again, it's not that it would be really technically complicated -\nadding \"give me a bundle\" support to gitweb should be pretty easy.\nHowever, this support has some \"social\" costs as well: no compatibility\nwith older git versions, support cost, confusion between dumb HTTP and\ngitweb HTTP transports, more lack of motivation for improving dumb HTTP\ntransport...\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nEver try. Ever fail. No matter. // Try again. Fail again. Fail better.\n\t\t-- Samuel Beckett\n"},{"id":"42393","messageId":"vpqlkfnipjl.fsf@bauges.imag.fr","threadId":"8170","inReplyTo":"20070517124006.GO4489@pasky.or.cz","subject":"Re: Smart fetch via HTTP?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-05-17T12:48:46Z","receivedAt":"2007-05-17T12:48:46Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Petr Baudis <pasky@suse.cz> writes:\n\n>> Mercurial and bzr both have this option. It would IMO have three benefits:\n>>  - Fast access for people behind paranoid firewalls, that only let http and\n>>    https (you can tunel anything through, but only to port 443) through.\n>\n>   How many users really have this problem? I'm not so sure.\n\nMany (if not most?) of the people working in a big company, I'd say.\nYear, it sucks, but people having used a paranoid firewall with a\nnot-less-paranoid and broken proxy understand what I mean.\n\n>>  - Can be run on shared machine. If you have web space on machine shared\n>>    by many people, you can set up your own gitweb, but cannot/are not allowed\n>>    to start your own network server for git native protocol.\n>\n>   You need to have CGI-enabled hosting, set up the CGI script etc. -\n> overally, the setup is similarly complicated as git-daemon setup, so\n> it's not \"zero-setup\" solution anymore.\n>\n>   Again, I'm not sure just how many people are in the situation that\n> they can run real CGI (not just PHP) but not git-daemon.\n\nAny volunteer to write a full-PHP version of git? ;-)\n\n-- \nMatthieu\n"},{"id":"42394","messageId":"46a038f90705170610mf9c9b0eu7b40af709469a601@mail.gmail.com","threadId":"8170","inReplyTo":"vpq8xbnlmdv.fsf@bauges.imag.fr","subject":"Re: Smart fetch via HTTP?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-05-17T13:10:35Z","receivedAt":"2007-05-17T13:10:35Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/17/07, Matthieu Moy <Matthieu.Moy@imag.fr> wrote:\n> FYI, bzr uses HTTP range requests, and the introduction of this\n> feature lead to significant performance improvement for them (bzr is\n> more dumb-protocol oriented than git is, so that's really important\n> there). They have this \"index file+data file\" system too, so you\n> download the full index file, and then send an HTTP range request to\n> get only the relevant parts of the data file.\n\nThat's the kind of thing I was imagining. Between the index and an\nadditional \"index-supplement-for-dumb-protocols\" maintained by\nupdate-server-info, http ranges can be bent to our evil purposes.\n\nOf course it won't be as network-efficient as the git proto, or even\nas the git-over-cgi proto, but it'll surely be server-cpu-and-memory\nefficient. And people will benefit from it without having to do any\nadditional setup.\n\nIt might be hard to come up with a usable approach to http ranges. But\nI do think it's worth considering carefully.\n\ncheers,\n\n\n\nm\n"},{"id":"42397","messageId":"Pine.LNX.4.64.0705171445100.6410@racer.site","threadId":"8170","inReplyTo":"46a038f90705170610mf9c9b0eu7b40af709469a601@mail.gmail.com","subject":"Re: Smart fetch via HTTP?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-17T13:47:59Z","receivedAt":"2007-05-17T13:47:59Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[I missed this mail, because Matthieu culled the Cc list again]\n\nOn Fri, 18 May 2007, Martin Langhoff wrote:\n\n> On 5/17/07, Matthieu Moy <Matthieu.Moy@imag.fr> wrote:\n>\n> > FYI, bzr uses HTTP range requests, and the introduction of this\n> > feature lead to significant performance improvement for them (bzr is\n> > more dumb-protocol oriented than git is, so that's really important\n> > there). They have this \"index file+data file\" system too, so you\n> > download the full index file, and then send an HTTP range request to\n> > get only the relevant parts of the data file.\n> \n> That's the kind of thing I was imagining. Between the index and an\n> additional \"index-supplement-for-dumb-protocols\" maintained by\n> update-server-info, http ranges can be bent to our evil purposes.\n> \n> Of course it won't be as network-efficient as the git proto, or even\n> as the git-over-cgi proto, but it'll surely be server-cpu-and-memory\n> efficient. And people will benefit from it without having to do any\n> additional setup.\n\nOf course, the problem is that only the server can know beforehand which \nobjects are needed. Imagine this:\n\nX - Y - Z\n  \\\n    A\n\n\nClient has \"X\", wants \"Z\", but not \"A\". Client needs \"Y\" and \"Z\". But \nclient cannot know that it needs \"Y\" before getting \"Z\", except if the \nserver says so.\n\nIf you have a solution for that problem, please enlighten me: I don't.\n\nCiao,\nDscho\n"},{"id":"42399","messageId":"vpqhcqbim0i.fsf@bauges.imag.fr","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705171445100.6410@racer.site","subject":"Re: Smart fetch via HTTP?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-05-17T14:05:01Z","receivedAt":"2007-05-17T14:05:01Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> Hi,\n>\n> [I missed this mail, because Matthieu culled the Cc list again]\n\nSorry about that, miss-configuration of my mailer. I didn't find time\nto solve it before.\n\nOTOH, since most people actually complain when you Cc them on a\nmailing list, the choice \"To Cc or not to Cc\" has no universal\nsolution ;-).\n\n-- \nMatthieu\n"},{"id":"42400","messageId":"46a038f90705170709j7eb23d4fy6811fc2985dd888d@mail.gmail.com","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705171445100.6410@racer.site","subject":"Re: Smart fetch via HTTP?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-05-17T14:09:29Z","receivedAt":"2007-05-17T14:09:29Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/18/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> If you have a solution for that problem, please enlighten me: I don't.\n\nOk - worst case scenario - have a minimal hints file that tells me the\nranges to fetch all commits and all trees. To reduce that Add to the\nhints file data to name the hashes (or even better - offsets) for the\ndelta chains that contain commits+trees relevant to all the heads -\nminus 10, 20, 30, 40 commits and 1,2,4,8 and 16 days.\n\nSo there's a good chance the client can get the commits+trees needed\nefficiently. For blobs, all you need is the index to mark the delta\nchains you need.\n\ncheers,\n\n\nm\n"},{"id":"42402","messageId":"alpine.LFD.0.99.0705170954200.24220@xanadu.home","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705171143350.6410@racer.site","subject":"Re: Smart fetch via HTTP?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-05-17T14:41:37Z","receivedAt":"2007-05-17T14:41:37Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 17 May 2007, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Wed, 16 May 2007, Nicolas Pitre wrote:\n> \n> > Still... I wonder if this could be actually workable.  A typical daily \n> > update on the Linux kernel repository might consist of a couple hundreds \n> > or a few tousands objects.  This could still be faster to fetch parts of \n> > a pack than the whole pack if the size difference is above a certain \n> > treshold.  It is certainly not worse than fetching loose objects.\n> > \n> > Things would be pretty horrid if you think of fetching a commit object, \n> > parsing it to find out what tree object to fetch, then parse that tree \n> > object to find out what other objects to fetch, and so on.\n> > \n> > But if you only take the approach of fetching the pack index files, \n> > finding out about the objects that the remote has that are not available \n> > locally, and then fetching all those objects from within pack files \n> > without even looking at them (except for deltas), then it should be \n> > possible to issue a couple requests in parallel and possibly have decent \n> > performances.  And if it turns out that more than, say, 70% of a \n> > particular pack is to be fetched (you can determine that up front), then \n> > it might be decided to fetch the whole pack.\n> > \n> > There is no way to sensibly keep those objects packed on the receiving \n> > end of course, but storing them as loose objects and repacking them \n> > afterwards should be just fine.\n> > \n> > Of course you'll get objects from branches in the remote repository you \n> > might not be interested in, but that's a price to pay for such a hack.  \n> > On average the overhead shouldn't be that big anyway if branches within \n> > a repository are somewhat related.\n> > \n> > I think this is something worth experimenting.\n> \n> I am a bit wary about that, because it is so complex. IMHO a cgi which \n> gets, say, up to a hundred refs (maybe something like ref~0, ref~1, ref~2, \n> ref~4, ref~8, ref~16, ... for the refs), and then makes a bundle for that \n> case on the fly, is easier to do.\n\nAnd if you have 1) the permission and 2) the CPU power to execute such a \ncgi on the server and obviously 3) the knowledge to set it up properly, \nthen why aren't you running the Git daemon in the first place?  After \nall, they both boil down to running git-pack-objects and sending out the \nresult.  I don't think such a solution really buys much.\n\nOn the other hand, if the client does all the work and provides the \nserver with a list of ranges within a pack it wants to be sent, then you \nsimply have zero special setup to perform on the hosting server and you \nkeep the server load down due to not running pack-objects there.  That, \nat least, is different enough from the Git daemon to be worth \nconsidering.  Not only does it provide an advantage to those who cannot \ndo anything but http out of their segregated network, but it also \nprovide many advantages on the server side too while the cgi approach \ndoesn't.\n\nAnd actually finding out the list of objects the remote has that you \ndon't have is not that complex.  It could go as follows:\n\n1) Fetch every .idx files the remote has.\n\n2) From those .idx files, keep only a list of objects that are unknown \n   locally.  A good starting point for doing this really efficiently is \n   the code for git-pack-redundant.\n\n3) From the .idx files we got in (1), create a reverse index to get each \n   object's size in the remote pack.  The code to do this already exists \n   in builtin-pack-objects.c.\n\n4) With the list of missing objects from (2) along with their offset and \n   size within a given pack file, fetch those objects from the remote \n   server.  Either perform multiple requests in parallel, or as someone \n   mentioned already, provide the server with a list of ranges you want \n   to be sent.\n\n5) Store the received objects as loose objects locally.  If a given \n   object is a delta, verify if its base is available locally, or if it \n   is listed amongst those objects to be fetched from the server.  If \n   not, add it to the list.  In most cases, delta base objects will be \n   objects already listed to be fetched anyway.  To greatly simplify \n   things, the loose delta object type from 2 years ago could be revived \n   (commit 91d7b8afc2) since a repack will get rid of them.\n\n6 Repeat (4) and (5) until everything has been fetched.\n\n7) Run git-pack-objects with the list of fetched objects.\n\nEt voilà.  Oh, and of course update your local refs from the remote's.\n\nActually there is nothing really complex in the above operations. And \nwith this the server side remains really simple with no special setup \nnor extra load beyond the simple serving of file content.\n\n\nNicolas\n"},{"id":"42403","messageId":"alpine.LFD.0.99.0705171043070.24220@xanadu.home","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705171445100.6410@racer.site","subject":"Re: Smart fetch via HTTP?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-05-17T14:50:37Z","receivedAt":"2007-05-17T14:50:37Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 17 May 2007, Johannes Schindelin wrote:\n\n> Hi,\n> \n> [I missed this mail, because Matthieu culled the Cc list again]\n> \n> On Fri, 18 May 2007, Martin Langhoff wrote:\n> \n> > On 5/17/07, Matthieu Moy <Matthieu.Moy@imag.fr> wrote:\n> >\n> > > FYI, bzr uses HTTP range requests, and the introduction of this\n> > > feature lead to significant performance improvement for them (bzr is\n> > > more dumb-protocol oriented than git is, so that's really important\n> > > there). They have this \"index file+data file\" system too, so you\n> > > download the full index file, and then send an HTTP range request to\n> > > get only the relevant parts of the data file.\n> > \n> > That's the kind of thing I was imagining. Between the index and an\n> > additional \"index-supplement-for-dumb-protocols\" maintained by\n> > update-server-info, http ranges can be bent to our evil purposes.\n> > \n> > Of course it won't be as network-efficient as the git proto, or even\n> > as the git-over-cgi proto, but it'll surely be server-cpu-and-memory\n> > efficient. And people will benefit from it without having to do any\n> > additional setup.\n> \n> Of course, the problem is that only the server can know beforehand which \n> objects are needed.\n\nBut the whole idea is that we don't care.\n\n> Imagine this:\n> \n> X - Y - Z\n>   \\\n>     A\n> \n> \n> Client has \"X\", wants \"Z\", but not \"A\". Client needs \"Y\" and \"Z\". But \n> client cannot know that it needs \"Y\" before getting \"Z\", except if the \n> server says so.\n> \n> If you have a solution for that problem, please enlighten me: I don't.\n\nWe're talking about a _dumb_ protocol here.  If you want something \nfancy, just use the Git daemon.\n\nOtherwise, you'll simply get everything the remote has that you don't \nhave, including A.\n\nIn practice this shouldn't be a problem because people tend to have \nclean repositories on machines they want their stuff to be published, \nmeaning that those public repos are usually the result of pushes, hence \nthey contain only the minimum set of needed objects.  Of course you get \nevery branches and not only a particular one, but that's the price to \npay with a dumb protocol.\n\n\nNicolas\n"},{"id":"42405","messageId":"alpine.LFD.0.99.0705171054450.24220@xanadu.home","threadId":"8170","inReplyTo":"46a038f90705170709j7eb23d4fy6811fc2985dd888d@mail.gmail.com","subject":"Re: Smart fetch via HTTP?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-05-17T15:01:02Z","receivedAt":"2007-05-17T15:01:02Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 18 May 2007, Martin Langhoff wrote:\n\n> On 5/18/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> > If you have a solution for that problem, please enlighten me: I don't.\n> \n> Ok - worst case scenario - have a minimal hints file that tells me the\n> ranges to fetch all commits and all trees. To reduce that Add to the\n> hints file data to name the hashes (or even better - offsets) for the\n> delta chains that contain commits+trees relevant to all the heads -\n> minus 10, 20, 30, 40 commits and 1,2,4,8 and 16 days.\n\nNO !\n\nThis is unreliable, unnecessary, and actually kills the beauty of \nthe solution's simplicity.\n\nYou get updates for every branches the remote has, period.\n\nNo server side extra files, no guesses, no arbitrary ranges, no backward \ncompatibility issues, no crap!\n\n\nNicolas\n"},{"id":"42406","messageId":"46a038f90705170824g4ef8c800w826ada3964b711a@mail.gmail.com","threadId":"8170","inReplyTo":"alpine.LFD.0.99.0705170954200.24220@xanadu.home","subject":"Re: Smart fetch via HTTP?","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-05-17T15:24:54Z","receivedAt":"2007-05-17T15:24:54Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 5/18/07, Nicolas Pitre <nico@cam.org> wrote:\n> And if you have 1) the permission and 2) the CPU power to execute such a\n> cgi on the server and obviously 3) the knowledge to set it up properly,\n> then why aren't you running the Git daemon in the first place?\n\nAnd you probably _are_ running git daemon. But some clients may be on\nshitty connections that only allow http. That's one of the scenarios\nwe're discussing.\n\ncheers,\n\n\nm\n"},{"id":"42407","messageId":"alpine.LFD.0.99.0705171130030.24220@xanadu.home","threadId":"8170","inReplyTo":"46a038f90705170824g4ef8c800w826ada3964b711a@mail.gmail.com","subject":"Re: Smart fetch via HTTP?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-05-17T15:34:25Z","receivedAt":"2007-05-17T15:34:25Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 18 May 2007, Martin Langhoff wrote:\n\n> On 5/18/07, Nicolas Pitre <nico@cam.org> wrote:\n> > And if you have 1) the permission and 2) the CPU power to execute such a\n> > cgi on the server and obviously 3) the knowledge to set it up properly,\n> > then why aren't you running the Git daemon in the first place?\n> \n> And you probably _are_ running git daemon. But some clients may be on\n> shitty connections that only allow http. That's one of the scenarios\n> we're discussing.\n\nThat's not what I'm disputing at all.\n\nI'm disputing the vertue of an HTTP solution involving a cgi with Git \nbundles vs an HTTP solution involving static file range serving.  The \nclients on shitty connections don't care either ways.\n\n\nNicolas\n"},{"id":"42427","messageId":"20070517200431.GA3079@efreet.light.src","threadId":"8170","inReplyTo":"alpine.LFD.0.99.0705170954200.24220@xanadu.home","subject":"Re: Smart fetch via HTTP?","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2007-05-17T20:04:31Z","receivedAt":"2007-05-17T20:04:31Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Thu, May 17, 2007 at 10:41:37 -0400, Nicolas Pitre wrote:\n> On Thu, 17 May 2007, Johannes Schindelin wrote:\n> > On Wed, 16 May 2007, Nicolas Pitre wrote:\n> And if you have 1) the permission and 2) the CPU power to execute such a \n> cgi on the server and obviously 3) the knowledge to set it up properly, \n> then why aren't you running the Git daemon in the first place?  After \n> all, they both boil down to running git-pack-objects and sending out the \n> result.  I don't think such a solution really buys much.\n\nYes, it does. I had 2 accounts where I could run CGI, but not separate\nserver, at university while I studied and now I can get the same on friend's\nserver. Neither of them would probably be ok for serving larger busy git\nrepository, but something smaller accessed by several people is OK. I think\nthis is quite common for university students.\n\nOf course your suggestion which moves the logic to client-side is a good one,\nbut even the cgi with logic on server side would help in some situations.\n\n> On the other hand, if the client does all the work and provides the \n> server with a list of ranges within a pack it wants to be sent, then you \n> simply have zero special setup to perform on the hosting server and you \n> keep the server load down due to not running pack-objects there.  That, \n> at least, is different enough from the Git daemon to be worth \n> considering.  Not only does it provide an advantage to those who cannot \n> do anything but http out of their segregated network, but it also \n> provide many advantages on the server side too while the cgi approach \n> doesn't.\n> \n> And actually finding out the list of objects the remote has that you \n> don't have is not that complex.  It could go as follows:\n> \n> 1) Fetch every .idx files the remote has.\n\n... for git it's 1.2 MiB. And that definitely isn't a huge source tree.\nOf course the local side could remember which indices it already saw during\nprevious fetch from that location and not re-fetch them.\n\nA slight problem is, that git-repack normally recombines everything to\na single pack, so the index would have to be re-fetched again anyway.\n\n> 2) From those .idx files, keep only a list of objects that are unknown \n>    locally.  A good starting point for doing this really efficiently is \n>    the code for git-pack-redundant.\n> \n> 3) From the .idx files we got in (1), create a reverse index to get each \n>    object's size in the remote pack.  The code to do this already exists \n>    in builtin-pack-objects.c.\n> \n> 4) With the list of missing objects from (2) along with their offset and \n>    size within a given pack file, fetch those objects from the remote \n>    server.  Either perform multiple requests in parallel, or as someone \n>    mentioned already, provide the server with a list of ranges you want \n>    to be sent.\n\nDoes the git server really have to do so much beyond that? I didn't look at\nthe algorithm that finds what deltas should be based on, but depending on\nthat it might (or might not) be possible to proof the client has everything to\nunderstand if the server sends the objects as it currently has them.\n\n> 5) Store the received objects as loose objects locally.  If a given \n>    object is a delta, verify if its base is available locally, or if it \n>    is listed amongst those objects to be fetched from the server.  If \n>    not, add it to the list.  In most cases, delta base objects will be \n>    objects already listed to be fetched anyway.  To greatly simplify \n>    things, the loose delta object type from 2 years ago could be revived \n>    (commit 91d7b8afc2) since a repack will get rid of them.\n> \n> 6 Repeat (4) and (5) until everything has been fetched.\n\nUnless I am really seriously missing something, there is no point in\nrepeating. For each pack you need to unpack a delta either:\n - you have it => ok.\n - you don't have it, but the server does =>\n    but than it's already in the fetch set calculated in 2.\n - you don't have it and nor does server =>\n    the repository at server is corrupted and you can't fix it.\n\n> 7) Run git-pack-objects with the list of fetched objects.\n> \n> Et voilà.  Oh, and of course update your local refs from the remote's.\n> \n> Actually there is nothing really complex in the above operations. And \n> with this the server side remains really simple with no special setup \n> nor extra load beyond the simple serving of file content.\n\nOn the other hand the amount of data transfered is larger, than with the git\nserver approach, because at least the indices have to be transfered in\nentirety. So each approach has it's own advantages.\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"42428","messageId":"20070517202655.GB3079@efreet.light.src","threadId":"8170","inReplyTo":"20070517124006.GO4489@pasky.or.cz","subject":"Re: Smart fetch via HTTP?","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2007-05-17T20:26:55Z","receivedAt":"2007-05-17T20:26:55Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Thu, May 17, 2007 at 14:40:06 +0200, Petr Baudis wrote:\n> On Tue, May 15, 2007 at 10:10:06PM CEST, Jan Hudec wrote:\n> >  - Can be run on shared machine. If you have web space on machine shared\n> >    by many people, you can set up your own gitweb, but cannot/are not allowed\n> >    to start your own network server for git native protocol.\n> \n>   You need to have CGI-enabled hosting, set up the CGI script etc. -\n> overally, the setup is similarly complicated as git-daemon setup, so\n> it's not \"zero-setup\" solution anymore.\n> \n>   Again, I'm not sure just how many people are in the situation that\n> they can run real CGI (not just PHP) but not git-daemon.\n\nA particular case would be a group of students wanting to publish their\nsoftware project (I mean the PRG023 or equivalent). Private computers in the\nhostel are not allowed to serve anything, so they'd use some of the lab\nservers (eg. artax, ss1000...). All of them allow full CGI, but running\ndaemons is forbiden.\n\n> >  - Less things to set up. If you are setting up gitweb anyway, you'd not need\n> >    to set up additional thing for providing fetch access.\n> \n>   Except, well, how do you \"set it up\"? You need to make sure\n> git-update-server-info is run, yes, but that shouldn't be a problem (I'm\n> not so sure if git does this for you automagically - Cogito would...).\n\nNo. If it worked similar to git-upload-pack, only over http, it would work\nwithout update-server-info, no?\n\n>   I think 95% of people don't set up gitweb.cgi either for their small\n> HTTP repositories. :-)\n> \n>   Then again, it's not that it would be really technically complicated -\n> adding \"give me a bundle\" support to gitweb should be pretty easy.\n> However, this support has some \"social\" costs as well: no compatibility\n> with older git versions, support cost, confusion between dumb HTTP and\n> gitweb HTTP transports, more lack of motivation for improving dumb HTTP\n> transport...\n\nThe dumb transport is definitely useful. Extending it to use ranges if\npossible would be useful as well (and maybe more than upload-pack-over-http).\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"42429","messageId":"alpine.LFD.0.99.0705171618410.24220@xanadu.home","threadId":"8170","inReplyTo":"20070517200431.GA3079@efreet.light.src","subject":"Re: Smart fetch via HTTP?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-05-17T20:31:50Z","receivedAt":"2007-05-17T20:31:50Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 17 May 2007, Jan Hudec wrote:\n\n> On Thu, May 17, 2007 at 10:41:37 -0400, Nicolas Pitre wrote:\n> > On Thu, 17 May 2007, Johannes Schindelin wrote:\n> > > On Wed, 16 May 2007, Nicolas Pitre wrote:\n> > And if you have 1) the permission and 2) the CPU power to execute such a \n> > cgi on the server and obviously 3) the knowledge to set it up properly, \n> > then why aren't you running the Git daemon in the first place?  After \n> > all, they both boil down to running git-pack-objects and sending out the \n> > result.  I don't think such a solution really buys much.\n> \n> Yes, it does. I had 2 accounts where I could run CGI, but not separate\n> server, at university while I studied and now I can get the same on friend's\n> server. Neither of them would probably be ok for serving larger busy git\n> repository, but something smaller accessed by several people is OK. I think\n> this is quite common for university students.\n> \n> Of course your suggestion which moves the logic to client-side is a good one,\n> but even the cgi with logic on server side would help in some situations.\n\nYou could simply wrap git-bundle within a cgi.  That is certainly easy \nenough.\n\n> > On the other hand, if the client does all the work and provides the \n> > server with a list of ranges within a pack it wants to be sent, then you \n> > simply have zero special setup to perform on the hosting server and you \n> > keep the server load down due to not running pack-objects there.  That, \n> > at least, is different enough from the Git daemon to be worth \n> > considering.  Not only does it provide an advantage to those who cannot \n> > do anything but http out of their segregated network, but it also \n> > provide many advantages on the server side too while the cgi approach \n> > doesn't.\n> > \n> > And actually finding out the list of objects the remote has that you \n> > don't have is not that complex.  It could go as follows:\n> > \n> > 1) Fetch every .idx files the remote has.\n> \n> ... for git it's 1.2 MiB. And that definitely isn't a huge source tree.\n> Of course the local side could remember which indices it already saw during\n> previous fetch from that location and not re-fetch them.\n\nRight.  The name of the pack/index plus its time stamp can be cached.  \nIf the remote doesn't repack too often then the overhead would be \nminimal.\n\n> > 2) From those .idx files, keep only a list of objects that are unknown \n> >    locally.  A good starting point for doing this really efficiently is \n> >    the code for git-pack-redundant.\n> > \n> > 3) From the .idx files we got in (1), create a reverse index to get each \n> >    object's size in the remote pack.  The code to do this already exists \n> >    in builtin-pack-objects.c.\n> > \n> > 4) With the list of missing objects from (2) along with their offset and \n> >    size within a given pack file, fetch those objects from the remote \n> >    server.  Either perform multiple requests in parallel, or as someone \n> >    mentioned already, provide the server with a list of ranges you want \n> >    to be sent.\n> \n> Does the git server really have to do so much beyond that?\n\nYes it does.  The real thing perform a full object reachability walk and \nonly the objects that are needed for the wanted branch(es) are sent in a \ncustom pack meaning that the data transfer is really optimal.\n\n> > 5) Store the received objects as loose objects locally.  If a given \n> >    object is a delta, verify if its base is available locally, or if it \n> >    is listed amongst those objects to be fetched from the server.  If \n> >    not, add it to the list.  In most cases, delta base objects will be \n> >    objects already listed to be fetched anyway.  To greatly simplify \n> >    things, the loose delta object type from 2 years ago could be revived \n> >    (commit 91d7b8afc2) since a repack will get rid of them.\n> > \n> > 6 Repeat (4) and (5) until everything has been fetched.\n> \n> Unless I am really seriously missing something, there is no point in\n> repeating. For each pack you need to unpack a delta either:\n>  - you have it => ok.\n>  - you don't have it, but the server does =>\n>     but than it's already in the fetch set calculated in 2.\n>  - you don't have it and nor does server =>\n>     the repository at server is corrupted and you can't fix it.\n\nYou're right of course.\n\n\nNicolas\n"},{"id":"42430","messageId":"alpine.LFD.0.99.0705171633440.24220@xanadu.home","threadId":"8170","inReplyTo":"20070517202655.GB3079@efreet.light.src","subject":"Re: Smart fetch via HTTP?","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-05-17T20:38:41Z","receivedAt":"2007-05-17T20:38:41Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 17 May 2007, Jan Hudec wrote:\n\n> A particular case would be a group of students wanting to publish their\n> software project (I mean the PRG023 or equivalent). Private computers in the\n> hostel are not allowed to serve anything, so they'd use some of the lab\n> servers (eg. artax, ss1000...). All of them allow full CGI, but running\n> daemons is forbiden.\n\nAnd wouldn't the admin authority for those lab servers be amenable to \ninstall a Git daemon service?  That'd be a much better solution to me.\n\n\nNicolas\n"},{"id":"42433","messageId":"Pine.LNX.4.64.0705171358070.16479@asgard.lang.hm","threadId":"8170","inReplyTo":"alpine.LFD.0.99.0705171618410.24220@xanadu.home","subject":"Re: Smart fetch via HTTP?","fromName":"","fromEmail":"david@lang.hm","sentAt":"2007-05-17T21:00:44Z","receivedAt":"2007-05-17T21:00:44Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 17 May 2007, Nicolas Pitre wrote:\n\n> On Thu, 17 May 2007, Jan Hudec wrote:\n>\n>> On Thu, May 17, 2007 at 10:41:37 -0400, Nicolas Pitre wrote:\n>>> On Thu, 17 May 2007, Johannes Schindelin wrote:\n>>>> On Wed, 16 May 2007, Nicolas Pitre wrote:\n>>> And if you have 1) the permission and 2) the CPU power to execute such a\n>>> cgi on the server and obviously 3) the knowledge to set it up properly,\n>>> then why aren't you running the Git daemon in the first place?  After\n>>> all, they both boil down to running git-pack-objects and sending out the\n>>> result.  I don't think such a solution really buys much.\n>>\n>> Yes, it does. I had 2 accounts where I could run CGI, but not separate\n>> server, at university while I studied and now I can get the same on friend's\n>> server. Neither of them would probably be ok for serving larger busy git\n>> repository, but something smaller accessed by several people is OK. I think\n>> this is quite common for university students.\n>>\n>> Of course your suggestion which moves the logic to client-side is a good one,\n>> but even the cgi with logic on server side would help in some situations.\n>\n> You could simply wrap git-bundle within a cgi.  That is certainly easy\n> enough.\n\nisn't this (or something very similar) exactly what we want for a smalrt \nfetch via http?\n\nafter all, we're completely in control of the client software, and the \nuseual reason for HTTP-only access is on the client side rather then the \nserver side. so http access that wraps the git protocol in http would make \nlife much cleaner for lots of people\n\nthere are a few cases where all you have is static web space, but I don't \nthink it's worth trying to optimize that too much as you still have the \nsafety issues to worry about\n\nDavid Lang\n"},{"id":"42439","messageId":"f2inc6$hl5$1@sea.gmane.org","threadId":"8170","inReplyTo":"46a038f90705170709j7eb23d4fy6811fc2985dd888d@mail.gmail.com","subject":"Re: Smart fetch via HTTP?","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-05-17T23:14:49Z","receivedAt":"2007-05-17T23:14:49Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Martin Langhoff wrote:\n\n> On 5/18/07, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n>> If you have a solution for that problem, please enlighten me: I don't.\n> \n> Ok - worst case scenario - have a minimal hints file that tells me the\n> ranges to fetch all commits and all trees. To reduce that Add to the\n> hints file data to name the hashes (or even better - offsets) for the\n> delta chains that contain commits+trees relevant to all the heads -\n> minus 10, 20, 30, 40 commits and 1,2,4,8 and 16 days.\n> \n> So there's a good chance the client can get the commits+trees needed\n> efficiently. For blobs, all you need is the index to mark the delta\n> chains you need.\n\nBy the way, I think we always should get the whole delta chain, unless we\nare absolutely sure that we have base object(s) in repo.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"42458","messageId":"Pine.LNX.4.64.0705180958390.6410@racer.site","threadId":"8170","inReplyTo":"20070517200431.GA3079@efreet.light.src","subject":"Re: Smart fetch via HTTP?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-05-18T09:01:52Z","receivedAt":"2007-05-18T09:01:52Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 May 2007, Jan Hudec wrote:\n\n> On Thu, May 17, 2007 at 10:41:37 -0400, Nicolas Pitre wrote:\n>\n> > And if you have 1) the permission and 2) the CPU power to execute such \n> > a cgi on the server and obviously 3) the knowledge to set it up \n> > properly, then why aren't you running the Git daemon in the first \n> > place?  After all, they both boil down to running git-pack-objects and \n> > sending out the result.  I don't think such a solution really buys \n> > much.\n> \n> Yes, it does. I had 2 accounts where I could run CGI, but not separate \n> server, at university while I studied and now I can get the same on \n> friend's server. Neither of them would probably be ok for serving larger \n> busy git repository, but something smaller accessed by several people is \n> OK. I think this is quite common for university students.\n\n1) This has nothing to do with the way the repo is served, but how much \nyou advertise it. The load will not be lower, just because you use a CGI \nscript.\n\n2) you say yourself that git-daemon would have less impact on the load:\n\n> > [...]\n> >\n> > Et voilà.  Oh, and of course update your local refs from the \n> > remote's.\n> > \n> > Actually there is nothing really complex in the above operations. And \n> > with this the server side remains really simple with no special setup \n> > nor extra load beyond the simple serving of file content.\n> \n> On the other hand the amount of data transfered is larger, than with the \n> git server approach, because at least the indices have to be transfered \n> in entirety.\n\nCiao,\nDscho\n"},{"id":"42497","messageId":"20070518173527.GA3327@efreet.light.src","threadId":"8170","inReplyTo":"alpine.LFD.0.99.0705171633440.24220@xanadu.home","subject":"Re: Smart fetch via HTTP?","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2007-05-18T17:35:27Z","receivedAt":"2007-05-18T17:35:27Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Thu, May 17, 2007 at 16:38:41 -0400, Nicolas Pitre wrote:\n> On Thu, 17 May 2007, Jan Hudec wrote:\n> \n> > A particular case would be a group of students wanting to publish their\n> > software project (I mean the PRG023 or equivalent). Private computers in the\n> > hostel are not allowed to serve anything, so they'd use some of the lab\n> > servers (eg. artax, ss1000...). All of them allow full CGI, but running\n> > daemons is forbiden.\n> \n> And wouldn't the admin authority for those lab servers be amenable to \n> install a Git daemon service?  That'd be a much better solution to me.\n\nIt would. But it would really depend on the administrator goodwill.\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"42498","messageId":"20070518175155.GB3327@efreet.light.src","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705180958390.6410@racer.site","subject":"Re: Smart fetch via HTTP?","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2007-05-18T17:51:55Z","receivedAt":"2007-05-18T17:51:55Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Fri, May 18, 2007 at 10:01:52 +0100, Johannes Schindelin wrote:\n> Hi,\n> \n> On Thu, 17 May 2007, Jan Hudec wrote:\n> \n> > On Thu, May 17, 2007 at 10:41:37 -0400, Nicolas Pitre wrote:\n> >\n> > > And if you have 1) the permission and 2) the CPU power to execute such \n> > > a cgi on the server and obviously 3) the knowledge to set it up \n> > > properly, then why aren't you running the Git daemon in the first \n> > > place?  After all, they both boil down to running git-pack-objects and \n> > > sending out the result.  I don't think such a solution really buys \n> > > much.\n> > \n> > Yes, it does. I had 2 accounts where I could run CGI, but not separate \n> > server, at university while I studied and now I can get the same on \n> > friend's server. Neither of them would probably be ok for serving larger \n> > busy git repository, but something smaller accessed by several people is \n> > OK. I think this is quite common for university students.\n> \n> 1) This has nothing to do with the way the repo is served, but how much \n> you advertise it. The load will not be lower, just because you use a CGI \n> script.\n\nThat won't. But that was never the purpose of \"smart cgi\". The purpose was to\nminimize the bandwidth usage (and connectivity is still not so cheap that\nyou'd not care) while still working over http either because the users need\nto access it from behind firewall or because administrator is not willing to\nset up git-daemon for you, while CGI you can run yourself.\n\n> 2) you say yourself that git-daemon would have less impact on the load:\n\nNO, I didn't -- at least not in the paragraph below.\n\nIn the below paragraph I said, that *network* use will never be as good with\n*dumb* solution, as it can be with smart solution, no matter whether it is\nover special protocol or HTTP.\n\n---\n\nOf course it would be less efficient in both CPU and network load, because\nthere is the overhead of the web server and overhead of the http headers.\n\nActually I like the ranges solution. If accompanied with repack stategy that\ndoes not pack everything together, but instead creates packs of limited\nnumber of objects -- so that the indices don't exceed configurable size, say\n64kB -- could not so much less efficient for the network and have the\nadvantage of working without ability to execute CGI.\n\n> > > [...]\n> > >\n> > > Et voilà.  Oh, and of course update your local refs from the \n> > > remote's.\n> > > \n> > > Actually there is nothing really complex in the above operations. And \n> > > with this the server side remains really simple with no special setup \n> > > nor extra load beyond the simple serving of file content.\n> > \n> > On the other hand the amount of data transfered is larger, than with the \n> > git server approach, because at least the indices have to be transfered \n> > in entirety.\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"},{"id":"42499","messageId":"alpine.LFD.0.98.0705181123590.3890@woody.linux-foundation.org","threadId":"8170","inReplyTo":"vpqlkfnipjl.fsf@bauges.imag.fr","subject":"Re: Smart fetch via HTTP?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-05-18T18:27:22Z","receivedAt":"2007-05-18T18:27:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 17 May 2007, Matthieu Moy wrote:\n> \n> Many (if not most?) of the people working in a big company, I'd say.\n> Year, it sucks, but people having used a paranoid firewall with a\n> not-less-paranoid and broken proxy understand what I mean.\n\nWell, we could try to support the git protocol over port 80..\n\nIOW, it's probably easier to try to get people to use\n\n\tgit clone git://some.host:80/project\n\nand just run git-daemon on port 80, than it is to try to set of magic cgi \nscripts etc.\n\nDoing that with virtual hosts etc should be pretty trivial. Much more so \nthan trying to make a git-cgi script.\n\nAnd yes, I do realize that in theory you can have http-aware firewalls \nthat expect to see the normal http sequences in the first few packets in \norder to pass things through, but I seriously doubt it's very common.\n\n\t\t\tLinus\n"},{"id":"42506","messageId":"Pine.LNX.4.64.0705181130030.13214@blackbox.fnordora.org","threadId":"8170","inReplyTo":"alpine.LFD.0.98.0705181123590.3890@woody.linux-foundation.org","subject":"Re: Smart fetch via HTTP?","fromName":"alan","fromEmail":"alan@clueserver.org","sentAt":"2007-05-18T18:33:02Z","receivedAt":"2007-05-18T18:33:02Z","isPatch":false,"sender":{"key":"alan@clueserver.org","avatar":null},"body":"On Fri, 18 May 2007, Linus Torvalds wrote:\n\n>\n>\n> On Thu, 17 May 2007, Matthieu Moy wrote:\n>>\n>> Many (if not most?) of the people working in a big company, I'd say.\n>> Year, it sucks, but people having used a paranoid firewall with a\n>> not-less-paranoid and broken proxy understand what I mean.\n>\n> Well, we could try to support the git protocol over port 80..\n>\n> IOW, it's probably easier to try to get people to use\n>\n> \tgit clone git://some.host:80/project\n>\n> and just run git-daemon on port 80, than it is to try to set of magic cgi\n> scripts etc.\n\nExcept some filtering firewalls try and strip content from data (like \nActiveX controls.)\n\nRunning git on port 53 will bypass pretty much every firewall out there.\n\n(If you want to learn how to bypass an overactive firewall, talk to a \nbunch of teenagers at a school with an agressive porn filter.)\n\n-- \n\"ANSI C says access to the padding fields of a struct is undefined.\nANSI C also says that struct assignment is a memcpy. Therefore struct\nassignment in ANSI C is a violation of ANSI C...\"\n                                   - Alan Cox\n"},{"id":"42507","messageId":"20070518190159.GS24644@ca-server1.us.oracle.com","threadId":"8170","inReplyTo":"alpine.LFD.0.98.0705181123590.3890@woody.linux-foundation.org","subject":"Re: Smart fetch via HTTP?","fromName":"Joel Becker","fromEmail":"joel.becker@oracle.com","sentAt":"2007-05-18T19:01:59Z","receivedAt":"2007-05-18T19:01:59Z","isPatch":false,"sender":{"key":"joel.becker@oracle.com","avatar":null},"body":"On Fri, May 18, 2007 at 11:27:22AM -0700, Linus Torvalds wrote:\n> Well, we could try to support the git protocol over port 80..\n> \n> IOW, it's probably easier to try to get people to use\n> \n> \tgit clone git://some.host:80/project\n> \n> and just run git-daemon on port 80, than it is to try to set of magic cgi \n> scripts etc.\n\n\tCan we tech the git-daemon to parse the HTTP headers\n(specifically, the URL) and return the appropriate HTTP response?\n\n> And yes, I do realize that in theory you can have http-aware firewalls \n> that expect to see the normal http sequences in the first few packets in \n> order to pass things through, but I seriously doubt it's very common.\n\n\tIt's not about packet scanning, it's about GET vs CONNECT.  If\nthe proxy allows GET but not CONNECT, it's going to forward the HTTP\nprotocol to the server, and git-daemon is going to see \"GET /project\nHTTP/1.1\" as its first input.  Now, perhaps we can cook that up behind\nsome apache so that apache handles vhosting the URL, then calls\ngit-daemon which can take the stdin.  So we'd be doing POST, not GET.\n\tOn the other hand, if the proxy allows CONNECT, there is no\nscanning for HTTP sequences done by the proxy.  It just allows all raw\ndata (as it figures you're doing SSL).\n\tA normal company needs to have their firewall allow CONNECT to\n9418.  Then git proxying over HTTP is possible to a standard git-daemon.\n\nJoel\n\n-- \n\n\"The first requisite of a good citizen in this republic of ours\n is that he shall be able and willing to pull his weight.\"\n\t- Theodore Roosevelt\n\nJoel Becker\nPrincipal Software Developer\nOracle\nE-mail: joel.becker@oracle.com\nPhone: (650) 506-8127\n"},{"id":"42531","messageId":"vpqbqgh99s5.fsf@bauges.imag.fr","threadId":"8170","inReplyTo":"20070518190159.GS24644@ca-server1.us.oracle.com","subject":"Re: Smart fetch via HTTP?","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2007-05-18T20:06:18Z","receivedAt":"2007-05-18T20:06:18Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"Joel Becker <Joel.Becker@oracle.com> writes:\n\n> \tA normal company needs to have their firewall allow CONNECT to\n> 9418.  Then git proxying over HTTP is possible to a standard\n> git-daemon.\n\n443 should work too (that's HTTPS, and the proxy can't filter it,\nsince this would be a man-in-the-middle attack).\n\n-- \nMatthieu\n"},{"id":"42532","messageId":"alpine.LFD.0.98.0705181312060.3890@woody.linux-foundation.org","threadId":"8170","inReplyTo":"20070518190159.GS24644@ca-server1.us.oracle.com","subject":"Re: Smart fetch via HTTP?","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-05-18T20:13:36Z","receivedAt":"2007-05-18T20:13:36Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 18 May 2007, Joel Becker wrote:\n> \n> \tIt's not about packet scanning, it's about GET vs CONNECT.  If\n> the proxy allows GET but not CONNECT, it's going to forward the HTTP\n> protocol to the server, and git-daemon is going to see \"GET /project\n> HTTP/1.1\" as its first input.  Now, perhaps we can cook that up behind\n> some apache so that apache handles vhosting the URL, then calls\n> git-daemon which can take the stdin.  So we'd be doing POST, not GET.\n\nIf it's _just_ the initial GET/CONNECT strings, yeah, we could probably \neasily make the git-daemon just ignore them. That shouldn't be a problem.\n\nBut if there's anything *else* required, it gets uglier much more quickly.\n\n\t\tLinus\n"},{"id":"42537","messageId":"20070518215607.GT24644@ca-server1.us.oracle.com","threadId":"8170","inReplyTo":"alpine.LFD.0.98.0705181312060.3890@woody.linux-foundation.org","subject":"Re: Smart fetch via HTTP?","fromName":"Joel Becker","fromEmail":"joel.becker@oracle.com","sentAt":"2007-05-18T21:56:07Z","receivedAt":"2007-05-18T21:56:07Z","isPatch":false,"sender":{"key":"joel.becker@oracle.com","avatar":null},"body":"On Fri, May 18, 2007 at 01:13:36PM -0700, Linus Torvalds wrote:\n> If it's _just_ the initial GET/CONNECT strings, yeah, we could probably \n> easily make the git-daemon just ignore them. That shouldn't be a problem.\n> \n> But if there's anything *else* required, it gets uglier much more quickly.\n\n\tWith CONNECT, there isn't anything.  That is, your\nGIT_PROXY_COMMAND handles talking to the proxy, then gives git itself a\nraw data pipe.  My proxy allows CONNECT to 9418, and that's how I use it\ntoday.\n\tIf you tried to make POST work (It'd be POST, not GET, as you\nneed to connect up the sending side), either apache would have to front\nit for us, or \"git-daemon --http\" would have to accept the HTTP headers\non before the input, and output a proper HTTP response before sending\noutput.  Seeing the headers would allow for us to vhost, even.\n\tHmm, but the proxy may not allow two-way communication.  Does\nthe git protocol have more than one round-trip?  That is:\n\nClient:\n    POST http://server.git.host:80/projects/thisproject HTTP/1.1\n    Host: server.git.host\n\n    fetch-pack <sha1>\n    EOF\n\nServer:\n    200 OK HTTP/1.1\n    \n    <data>\n    EOF\n\nshould work, I'd think.\n\nJoel\n\n\n-- \n\n\"Ninety feet between bases is perhaps as close as man has ever come\n to perfection.\"\n\t- Red Smith\n\nJoel Becker\nPrincipal Software Developer\nOracle\nE-mail: joel.becker@oracle.com\nPhone: (650) 506-8127\n"},{"id":"42575","messageId":"Pine.LNX.4.64.0705181742470.20116@asgard.lang.hm","threadId":"8170","inReplyTo":"alpine.LFD.0.98.0705181123590.3890@woody.linux-foundation.org","subject":"Re: Smart fetch via HTTP?","fromName":"","fromEmail":"david@lang.hm","sentAt":"2007-05-19T00:50:08Z","receivedAt":"2007-05-19T00:50:08Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 18 May 2007, Linus Torvalds wrote:\n\n> On Thu, 17 May 2007, Matthieu Moy wrote:\n>>\n>> Many (if not most?) of the people working in a big company, I'd say.\n>> Year, it sucks, but people having used a paranoid firewall with a\n>> not-less-paranoid and broken proxy understand what I mean.\n>\n> Well, we could try to support the git protocol over port 80..\n>\n> IOW, it's probably easier to try to get people to use\n>\n> \tgit clone git://some.host:80/project\n>\n> and just run git-daemon on port 80, than it is to try to set of magic cgi\n> scripts etc.\n>\n> Doing that with virtual hosts etc should be pretty trivial. Much more so\n> than trying to make a git-cgi script.\n>\n> And yes, I do realize that in theory you can have http-aware firewalls\n> that expect to see the normal http sequences in the first few packets in\n> order to pass things through, but I seriously doubt it's very common.\n\nthey are actually more common than you think, and getting even more common \nthanks to IE\n\nwhen a person browsing a hostile website will allow that website to take \nover the machine the demand is created for 'malware filters' for http, to \ndo this the firewalls need to decode the http, and in the process limit \nyou to only doing legitimate http.\n\nit's also the case that the companies that have firewalls paranoid enough \nto not let you get to the git port are highly likely to be paranoid enough \nto have a malware filtering http firewall.\n\nDavid Lang\n"},{"id":"42582","messageId":"20070519035856.GB3141@spearce.org","threadId":"8170","inReplyTo":"Pine.LNX.4.64.0705181742470.20116@asgard.lang.hm","subject":"Re: Smart fetch via HTTP?","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-05-19T03:58:56Z","receivedAt":"2007-05-19T03:58:56Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"david@lang.hm wrote:\n> when a person browsing a hostile website will allow that website to take \n> over the machine the demand is created for 'malware filters' for http, to \n> do this the firewalls need to decode the http, and in the process limit \n> you to only doing legitimate http.\n> \n> it's also the case that the companies that have firewalls paranoid enough \n> to not let you get to the git port are highly likely to be paranoid enough \n> to have a malware filtering http firewall.\n\nI'm behind such a filter, and fetch git.git via HTTP just to keep\nmy work system current with Junio.  ;-)\n\nOf course we're really really really paranoid about our firewall,\nbut are also so paranoid that any other web browser *except*\nMicrosoft Internet Explorer is thought to be a security risk and\nis more-or-less banned from the network.\n\nThe kicker is some of our developers create public websites, where\ntesting your local webpage with Firefox and Safari is pretty much\nrequired...  but those browsers still aren't as trusted as IE and\nrequire special clearances.  *shakes head*\n\nWe're pretty much limited to:\n\n *) Running the native Git protocol SSL, where the remote system\n is answering to port 443.  It may not need to be HTTP at all,\n but it probably has to smell enough like SSL to get it through\n the malware filter.  Oh, what's that?  The filter cannot actually\n filter the SSL data?  Funny!  ;-)\n\n *) Using a single POST upload followed by response from server,\n formatted with minimal HTTP headers.  The real problem as people\n have pointed out is not the HTTP headers, but it is the single\n exchange.\n\nOne might think you could use HTTP pipelining to try and get a\nbi-directional channel with the remote system, but I'm sure proxy\nservers are not required to reuse the same TCP connection to the\nremote HTTP server when the inside client piplines a new request.\nSo any sort of hack on pipelining won't work.\n\nIf you really want a stateful exchange you have to treat HTTP as\nthough it were IP, but with reliable (and much more expensive)\npacket delivery, and make the Git daemon keep track of the protocol\nstate with the client.  Yes, that means that when the client suddenly\ngoes away and doesn't tell you he went away you also have to garbage\ncollect your state.  No nice messages from your local kernel.  :-(\n\n-- \nShawn.\n"},{"id":"42588","messageId":"Pine.LNX.4.64.0705182154540.6938@asgard.lang.hm","threadId":"8170","inReplyTo":"20070519035856.GB3141@spearce.org","subject":"Re: Smart fetch via HTTP?","fromName":"","fromEmail":"david@lang.hm","sentAt":"2007-05-19T04:58:57Z","receivedAt":"2007-05-19T04:58:57Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 18 May 2007, Shawn O. Pearce wrote:\n\n> david@lang.hm wrote:\n>> when a person browsing a hostile website will allow that website to take\n>> over the machine the demand is created for 'malware filters' for http, to\n>> do this the firewalls need to decode the http, and in the process limit\n>> you to only doing legitimate http.\n>>\n>> it's also the case that the companies that have firewalls paranoid enough\n>> to not let you get to the git port are highly likely to be paranoid enough\n>> to have a malware filtering http firewall.\n>\n> I'm behind such a filter, and fetch git.git via HTTP just to keep\n> my work system current with Junio.  ;-)\n>\n> Of course we're really really really paranoid about our firewall,\n> but are also so paranoid that any other web browser *except*\n> Microsoft Internet Explorer is thought to be a security risk and\n> is more-or-less banned from the network.\n>\n> The kicker is some of our developers create public websites, where\n> testing your local webpage with Firefox and Safari is pretty much\n> required...  but those browsers still aren't as trusted as IE and\n> require special clearances.  *shakes head*\n\nthis isn't paranoia, this is just bullheadedness\n\n> We're pretty much limited to:\n>\n> *) Running the native Git protocol SSL, where the remote system\n> is answering to port 443.  It may not need to be HTTP at all,\n> but it probably has to smell enough like SSL to get it through\n> the malware filter.  Oh, what's that?  The filter cannot actually\n> filter the SSL data?  Funny!  ;-)\n\nwe're actually paranoid enough to have devices that do man-in-the-middle \ndecryption for some sites, and are given copies of the encryption keys \nthat other sites (and browsers) use so that it can decrypt the SSL and \ncheck it. I admit that this is far more paranoid then almost all sites \nthough :-)\n\n> *) Using a single POST upload followed by response from server,\n> formatted with minimal HTTP headers.  The real problem as people\n> have pointed out is not the HTTP headers, but it is the single\n> exchange.\n\n> If you really want a stateful exchange you have to treat HTTP as\n> though it were IP, but with reliable (and much more expensive)\n> packet delivery, and make the Git daemon keep track of the protocol\n> state with the client.  Yes, that means that when the client suddenly\n> goes away and doesn't tell you he went away you also have to garbage\n> collect your state.  No nice messages from your local kernel.  :-(\n\nunfortunantly you are right about this.\n\nDavid Lang\n"},{"id":"42694","messageId":"20070520103010.GA27087@efreet.light.src","threadId":"8170","inReplyTo":"20070518215607.GT24644@ca-server1.us.oracle.com","subject":"Re: Smart fetch via HTTP?","fromName":"Jan Hudec","fromEmail":"bulb@ucw.cz","sentAt":"2007-05-20T10:30:10Z","receivedAt":"2007-05-20T10:30:10Z","isPatch":false,"sender":{"key":"bulb@ucw.cz","avatar":null},"body":"On Fri, May 18, 2007 at 14:56:07 -0700, Joel Becker wrote:\n> On Fri, May 18, 2007 at 01:13:36PM -0700, Linus Torvalds wrote:\n> > If it's _just_ the initial GET/CONNECT strings, yeah, we could probably \n> > easily make the git-daemon just ignore them. That shouldn't be a problem.\n> > \n> > But if there's anything *else* required, it gets uglier much more quickly.\n> \n> \tWith CONNECT, there isn't anything.  That is, your\n> GIT_PROXY_COMMAND handles talking to the proxy, then gives git itself a\n> raw data pipe.  My proxy allows CONNECT to 9418, and that's how I use it\n> today.\n\nYes. Connect is easy. However many companies only allow CONNECT to 443\n(not that it's much more secure than allowing it anywhere, but at least it\nhas to block CONNECT to 25 to block sending spam).\n\n> \tIf you tried to make POST work (It'd be POST, not GET, as you\n> need to connect up the sending side), either apache would have to front\n> it for us, or \"git-daemon --http\" would have to accept the HTTP headers\n> on before the input, and output a proper HTTP response before sending\n> output.  Seeing the headers would allow for us to vhost, even.\n> \tHmm, but the proxy may not allow two-way communication.  Does\n> the git protocol have more than one round-trip?  That is:\n> \n> Client:\n>     POST http://server.git.host:80/projects/thisproject HTTP/1.1\n>     Host: server.git.host\n> \n>     fetch-pack <sha1>\n>     EOF\n> \n> Server:\n>     200 OK HTTP/1.1\n>     \n>     <data>\n>     EOF\n> \n> should work, I'd think.\n\nWell, that does not require git at all -- apache can handle this all right.\nBut it's not network-efficient. To be network-efficient, it is necessary to\nnegotiate the list of objects that need to be send. And that requires more\nthan one round-trip. Additionally, the current git protocol is streaming --\nthe client sends data without waiting for the server. So it would require\nslightly different protocol over HTTP.\n\n-- \n\t\t\t\t\t\t Jan 'Bulb' Hudec <bulb@ucw.cz>\n"}]}