{"thread":{"id":"14798","subject":"More on git over HTTP POST","startedAt":"2008-08-01T21:50:49Z","lastAt":"2008-08-13T02:37:59Z","messageCount":40,"participants":["H. Peter Anvin","Shawn O. Pearce","Daniel Stenberg","Petr Baudis","Junio C Hamano","Mike Hommey","david@lang.hm","Rogan Dawes","Johannes Schindelin"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"85937","messageId":"48938539.9060003@zytor.com","threadId":"14798","inReplyTo":null,"subject":"More on git over HTTP POST","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-01T21:50:49Z","receivedAt":"2008-08-01T21:50:49Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Hi all,\n\nI have investigated a bit what it would take to support git protocol \n(smart transport) over HTTP POST transactions.\n\nThe current proxy system is broken, for a very simple reason: it doesn't \nconvey information about when the channel should be turned around.\n\nHTTP POST -- or, for that matter, any RPC-style transport, is a half \nduplex transport: only one direction can be active at a time, after \nwhich the channel has to be explicitly turned around.  The \"turning \naround\" consists of posting the queued transaction and listening for the \nreply.\n\nUltimately, it comes down to the following: the transactor needs to be \ngiven explicit information when the git protocol goes from writing to \nreading (the opposite direction information is obvious.)  I was hoping \nthat it would be possible to get this information from snooping the \nprotocol, but it doesn't seem to be so lucky.\n\nI started to hack on a variant which would embed a VFS-style interface \nin git itself, looking something like:\n\nstruct transactor;\n\nstruct transact_ops {\n\tssize_t (*read)(struct transactor *, void *, size_t);\n\tssize_t (*write)(struct transactor *, const void *, size_t);\n\tint (*close)(struct transactor *);\n};\n\nstruct transactor {\n\tunion {\n\t\tvoid *p;\n\t\tintptr_t i;\n\t} u;\n\tconst struct transact_ops *ops;\n};\n\nReplacing the usual fd operations with this interface would allow a \ndifferent transactor to see the phase changes explicitly; the \nreplacement to use xread() and xwrite() is obvious.\n\nOf course, I started hacking on it and found myself with zero time to \ncontinue, but I thought I'd post what I had come up with.\n\n\t-hpa\n"},{"id":"86009","messageId":"20080802205702.GA24723@spearce.org","threadId":"14798","inReplyTo":"48938539.9060003@zytor.com","subject":"Re: More on git over HTTP POST","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-02T20:57:02Z","receivedAt":"2008-08-02T20:57:02Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> I have investigated a bit what it would take to support git protocol  \n> (smart transport) over HTTP POST transactions.\n\nI have started to think about this more myself, not just for POST\nput also for some form of GET that can return an efficient pack,\nrather than making the client walk the object chains itself.\n\nHave you looked at the Mecurial wire protocol?  It runs over HTTP\nand uses a relatively efficient means of deciding where to cut the\ntransfer at.\n\n  http://www.selenic.com/mercurial/wiki/index.cgi/WireProtocol\n\nMost of their smarts are in the branches() and between() operations.\n\nUnfortunately this documentation isn't very complete and/or there\nare some simplifications that the Mecurial team took due to their\nrepository format not initially supporting multiple branches like\nthe Git format does.\n\n> The current proxy system is broken, for a very simple reason: it doesn't  \n> convey information about when the channel should be turned around.\n\nWell, over git:// (or any protocol that wraps git:// like ssh)\nwe assume a full-duplex channel.  Some proxy systems are able to\ndo such a channel.  HTTP however does not offer it.\n\n> I started to hack on a variant which would embed a VFS-style interface  \n> in git itself, looking something like:\n>\n> struct transactor;\n>\n> struct transact_ops {\n> \tssize_t (*read)(struct transactor *, void *, size_t);\n> \tssize_t (*write)(struct transactor *, const void *, size_t);\n> \tint (*close)(struct transactor *);\n> };\n\nNo, the git:// protocol implementation in fetch-pack/upload-pack\nruns more efficient than that by keeping a sliding window of stuff\nthat is in-flight.  Its I guess two async RPCs running in parallel,\nbut from the client and server perspective both RPCs go into the\nsame computation.\n\nHTTP POST is actually trivial if you don't want to support the new\ntell-me-more extension that was added to git-push.  Hell, I could\nwrite the CGI in a few minutes I think.  Its really just a small\nwrapper around git-receive-pack.\n\nWhat's a bitch is the efficient fetch, and getting tell-me-more to\nwork on push.\n\n-- \nShawn.\n"},{"id":"86012","messageId":"alpine.LRH.1.10.0808022257470.25900@yvahk3.pbagnpgbe.fr","threadId":"14798","inReplyTo":"20080802205702.GA24723@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"Daniel Stenberg","fromEmail":"daniel@haxx.se","sentAt":"2008-08-02T21:00:41Z","receivedAt":"2008-08-02T21:00:41Z","isPatch":false,"sender":{"key":"daniel@haxx.se","avatar":"https://gravatar.com/avatar/69fdca87edd17cee21ca2e79fc2ff671d644603c3dc27167430f3cd3dbab7ba8?d=mp&s=160"},"body":"On Sat, 2 Aug 2008, Shawn O. Pearce wrote:\n\n> Well, over git:// (or any protocol that wraps git:// like ssh) we assume a \n> full-duplex channel.  Some proxy systems are able to do such a channel. \n> HTTP however does not offer it.\n\nYes it does. The CONNECT method is used to get a full-duplex channel to a \nremote site through a HTTP proxy. The downside with that is of course that \nmost proxies are setup to disallow CONNECT to other ports than 443 (the https \ndefault port).\n\n-- \n\n  / daniel.haxx.se\n"},{"id":"86014","messageId":"20080802210828.GE24723@spearce.org","threadId":"14798","inReplyTo":"alpine.LRH.1.10.0808022257470.25900@yvahk3.pbagnpgbe.fr","subject":"Re: More on git over HTTP POST","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-02T21:08:28Z","receivedAt":"2008-08-02T21:08:28Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Daniel Stenberg <daniel@haxx.se> wrote:\n> On Sat, 2 Aug 2008, Shawn O. Pearce wrote:\n>\n>> Well, over git:// (or any protocol that wraps git:// like ssh) we \n>> assume a full-duplex channel.  Some proxy systems are able to do such a \n>> channel. HTTP however does not offer it.\n>\n> Yes it does. The CONNECT method is used to get a full-duplex channel to a \n> remote site through a HTTP proxy. The downside with that is of course \n> that most proxies are setup to disallow CONNECT to other ports than 443 \n> (the https default port).\n\nAh, yes.  CONNECT.  Very few servers wind up supporting it I think.\n\nI know one very big company who cannot use or support Git because\nGit over HTTP is too slow to be useful.  They support other tools\nlike Subversion instead.  :-|\n\nReally we just need smart protocol support in half-duplex RPC like\nhpa was going after.  Then it doesn't matter what we serialize it\ninto, almost any RPC system will be useful.  Of course the only\none that probably matters in practice is HTTP.\n\n-- \nShawn.\n"},{"id":"86018","messageId":"20080802212357.GU32184@machine.or.cz","threadId":"14798","inReplyTo":"20080802210828.GE24723@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2008-08-02T21:23:58Z","receivedAt":"2008-08-02T21:23:58Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Sat, Aug 02, 2008 at 02:08:28PM -0700, Shawn O. Pearce wrote:\n> I know one very big company who cannot use or support Git because\n> Git over HTTP is too slow to be useful.  They support other tools\n> like Subversion instead.  :-|\n\nOn what projects? I'm currently using Git over HTTP (read-only) a lot\nand it doesn't seem really all that impractical to me. Maybe just using\na more dumb-friendly packing scheme could help a lot?\n\n\t\t\t\tPetr \"Pasky\" Baudis\n"},{"id":"86020","messageId":"20080802213248.GH24723@spearce.org","threadId":"14798","inReplyTo":"20080802212357.GU32184@machine.or.cz","subject":"Re: More on git over HTTP POST","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-02T21:32:48Z","receivedAt":"2008-08-02T21:32:48Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Petr Baudis <pasky@suse.cz> wrote:\n> On Sat, Aug 02, 2008 at 02:08:28PM -0700, Shawn O. Pearce wrote:\n> > I know one very big company who cannot use or support Git because\n> > Git over HTTP is too slow to be useful.  They support other tools\n> > like Subversion instead.  :-|\n> \n> On what projects? I'm currently using Git over HTTP (read-only) a lot\n> and it doesn't seem really all that impractical to me. Maybe just using\n> a more dumb-friendly packing scheme could help a lot?\n\nThey tested by taking the SVN source code and importing it into\nboth Git and Hg, then cloned them both over a WAN link.  Git was\n22x slower.  I suspect they didn't pack the Git repository at all,\nso Git had to issue thousands of HTTP GET requests for the loose\nobjects.  But I also suspect there was bias in the testing so they\ndidn't realize they needed to repack, and didn't care to find out.\n\nI've probably already said too much.  I'm under NDAs.\n\nBut anyway.  The point I was trying to make was that there are\nnot just some proxy servers, but also some server platforms, that\ncannot handle bidirectional communiction.  E.g. servers that are\nbehind reverse proxies, where the reverse proxy is acting as a sort\nof firewall or content cache accelerator.\n\n-- \nShawn.\n"},{"id":"86029","messageId":"20080803025602.GB27465@spearce.org","threadId":"14798","inReplyTo":"20080802205702.GA24723@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-03T02:56:02Z","receivedAt":"2008-08-03T02:56:02Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> wrote:\n> \"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> > I have investigated a bit what it would take to support git protocol  \n> > (smart transport) over HTTP POST transactions.\n> \n> I have started to think about this more myself, not just for POST\n> put also for some form of GET that can return an efficient pack,\n> rather than making the client walk the object chains itself.\n...\n> HTTP POST is actually trivial if you don't want to support the new\n> tell-me-more extension that was added to git-push.  Hell, I could\n> write the CGI in a few minutes I think.  Its really just a small\n> wrapper around git-receive-pack.\n\nSo I have this draft of how smart push might work.  Its slated\nfor the Documentation/technical directory.  Thus far I have only\nwritten about push support, but Ilari on #git has some ideas about\nhow to do a smart fetch protocol.\n\nImplementation wise in C git I think this is just a new C\nprogram (git-http-backend?) that turns around and proxies\ninto git-receive-pack, at least for the push support.\n\nWhat I don't know is how we could configure URI translation from\n/path/to/repository.git received out of the $PATH_INFO in the\nCGI environment to a physical directory.  Should we rely on the\nserver's $PATH_TRANSLATED?\n\n\nSmart HTTP transfer protocols\n=============================\n\nGit supports two HTTP based transfer protocols.  A \"dumb\" protocol\nwhich requires only a standard HTTP server on the server end of the\nconnection, and a \"smart\" protocol which requires a Git aware CGI\n(or server module).  This document describes the \"smart\" protocol.\n\nAuthentication\n--------------\n\nStandard HTTP authentication is used, and must be configured and\nenforced by the HTTP server software.\n\nChunked Transfer Encoding\n-------------------------\n\nFor performance reasons the HTTP/1.1 chunked transfer encoding is\nused frequently to transfer variable length objects.  This avoids\nneeding to produce large results in memory to compute the proper\ncontent-length.\n\nDetecting Smart Servers\n-----------------------\n\nHTTP clients can detect a smart Git-aware server by sending the\nshow-ref request (below) to the server.  If the response has a\nstatus of 200 and the magic x-application/git-refs content type\nthen the server can be assumed to be a smart Git-aware server.\n\nIf any other response is received the client must assume dumb\nprotocol support, as the server did not correctly response to\nthe request.\n\n\nShow Refs\n---------\n\nObtains the available refs from the remote repository.  The response\nis a sequence of git \"packet lines\", one per ref, and a final flush\npacket line to indicate the end of stream.\n\n\tC: GET /path/to/repository.git?show-ref HTTP/1.0\n\n\tS: HTTP/1.1 200 OK\n\tS: Content-Type: x-application/git-refs\n\tS: Transfer-Encoding: chunked\n\tS:\n\tS: 62\n\tS: 003e95dcfa3633004da0049d3d0fa03f80589cbcaf31 refs/heads/maint\n\tS: \n\tS: 63\n\tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n\tS: \n\tS: 59\n\tS: 003b2cb58b79488a98d2721cea644875a8dd0026b115 refs/heads/pu\n\tS: \n\tS: 4\n\tS: 0000\n\tS: 0\n\nPush Pack\n---------\n\nUploads a pack and updates refs.  The start of the stream is the\ncommands to update the refs and the remainder of the stream is the\npack file itself.  See git-receive-pack and its network protocol\nin pack-protocol.txt, as this is essentially the same.\n\n\tC: POST /path/to/repository.git?receive-pack HTTP/1.0\n\tC: Content-Type: x-application/git-receive-pack\n\tC: Transfer-Encoding: chunked\n\tC:\n\tC: 103\n\tC: 006395dcfa3633004da0049d3d0fa03f80589cbcaf31 d049f6c27a2244e12041955e262a404c7faba355 refs/heads/maint\n\tC: 4\n\tC: 0000\n\tC: 12\n\tC: PACK\n\t...\n\tC: 0\n\n\tS: HTTP/1.0 200 OK\n\tS: Content-type: x-application/git-status\n\tS: Transfer-Encoding: chunked\n\tS:\n\tS: ...<output of receive-pack>...\n\n\n\n-- \nShawn.\n"},{"id":"86031","messageId":"7v63qiydzg.fsf@gitster.siamese.dyndns.org","threadId":"14798","inReplyTo":"20080803025602.GB27465@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-08-03T03:27:47Z","receivedAt":"2008-08-03T03:27:47Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> Show Refs\n> ---------\n>\n> Obtains the available refs from the remote repository.  The response\n> is a sequence of git \"packet lines\", one per ref, and a final flush\n> packet line to indicate the end of stream.\n\nAs the initial protocol exchange request, I suspect that you would regret\nif you do not leave room for some \"capability advertisement\" in this\nexchange.\n\nWith the git native protocol, we luckily found space to do so after the\nref payload (because pkt-line is \"length + payload\" format but the code\nthat reads payload happened to ignore anything after NUL).  You would want\nto define how these are given by the server to the client over HTTP\nchannel.  For example, putting them on extra HTTP headers is probably Ok.\n"},{"id":"86032","messageId":"20080803033155.GC27465@spearce.org","threadId":"14798","inReplyTo":"7v63qiydzg.fsf@gitster.siamese.dyndns.org","subject":"Re: More on git over HTTP POST","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-03T03:31:55Z","receivedAt":"2008-08-03T03:31:55Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n> \n> > Show Refs\n> > ---------\n> >\n> > Obtains the available refs from the remote repository.  The response\n> > is a sequence of git \"packet lines\", one per ref, and a final flush\n> > packet line to indicate the end of stream.\n> \n> As the initial protocol exchange request, I suspect that you would regret\n> if you do not leave room for some \"capability advertisement\" in this\n> exchange.\n> \n> With the git native protocol, we luckily found space to do so after the\n> ref payload (because pkt-line is \"length + payload\" format but the code\n> that reads payload happened to ignore anything after NUL).  You would want\n> to define how these are given by the server to the client over HTTP\n> channel.  For example, putting them on extra HTTP headers is probably Ok.\n\nYea, I thought that the HTTP headers would be more than enough\nspace to add capability advertisements.  Most client libraries\nwill happily parse and store these for the application, and won't\nmake a fuss if the application doesn't read them.\n\nHence there's more than enough room in the protocol to extend it\nin the future with additional capabilities.\n\nWe do have to be careful though.  Any cachable resource must only\nrely upon the URI and the standard headers which compute into the\ncache key for a request.  There aren't many, though I think the\nContent-Type header may be among them.\n\n-- \nShawn.\n"},{"id":"86033","messageId":"48952A62.6050709@zytor.com","threadId":"14798","inReplyTo":"7v63qiydzg.fsf@gitster.siamese.dyndns.org","subject":"Re: More on git over HTTP POST","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-03T03:47:46Z","receivedAt":"2008-08-03T03:47:46Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Junio C Hamano wrote:\n> With the git native protocol, we luckily found space to do so after the\n> ref payload (because pkt-line is \"length + payload\" format but the code\n> that reads payload happened to ignore anything after NUL).  You would want\n> to define how these are given by the server to the client over HTTP\n> channel.  For example, putting them on extra HTTP headers is probably Ok.\n\nI think that would be a mistake, just because it's one more thing for \nproxies to screw up on.  It's better to have negotiation information in \nthe payload, before the \"real\" data.\n\nObviously one thing that needs to be included in each transaction is a \ntransaction ID that will be reported back on the next transaction, since \nyou can't rely on a persistent connection.\n\n\t-hpa\n"},{"id":"86034","messageId":"48952B2E.3030209@zytor.com","threadId":"14798","inReplyTo":"20080803025602.GB27465@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-03T03:51:10Z","receivedAt":"2008-08-03T03:51:10Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> Chunked Transfer Encoding\n> -------------------------\n> \n> For performance reasons the HTTP/1.1 chunked transfer encoding is\n> used frequently to transfer variable length objects.  This avoids\n> needing to produce large results in memory to compute the proper\n> content-length.\n\nNote: you cannot rely on HTTP/1.1 being supported by an intermediate \nproxy; you might have to handle HTTP/1.0, where the data is terminated \nby connection close.\n\n\t-hpa\n"},{"id":"86035","messageId":"48952D87.2070707@zytor.com","threadId":"14798","inReplyTo":"20080803025602.GB27465@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-03T04:01:11Z","receivedAt":"2008-08-03T04:01:11Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> Chunked Transfer Encoding\n> -------------------------\n> \n> For performance reasons the HTTP/1.1 chunked transfer encoding is\n> used frequently to transfer variable length objects.  This avoids\n> needing to produce large results in memory to compute the proper\n> content-length.\n\nOne more thing about chunked transfer encodings: you cannot assume that \na proxy will maintain chunk boundaries, any more than you can assume \nthat a firewall will maintain TCP packet boundaries.\n\n> Detecting Smart Servers\n> -----------------------\n> \n> HTTP clients can detect a smart Git-aware server by sending the\n> show-ref request (below) to the server.  If the response has a\n> status of 200 and the magic x-application/git-refs content type\n> then the server can be assumed to be a smart Git-aware server.\n> \n> If any other response is received the client must assume dumb\n> protocol support, as the server did not correctly response to\n> the request.\n\nI think it should be application/x-git-refs, but that's splitting hairs.\n\n> Obtains the available refs from the remote repository.  The response\n> is a sequence of git \"packet lines\", one per ref, and a final flush\n> packet line to indicate the end of stream.\n> \n> \tC: GET /path/to/repository.git?show-ref HTTP/1.0\n> \n\nI really think it would make more sense to use POST requests for \neverything, and have the command part of the POSTed payload.  Putting \nstuff in the URL just complicates the namespace to the detriment of the \nadmin.\n\n> \tS: HTTP/1.1 200 OK\n> \tS: Content-Type: x-application/git-refs\n> \tS: Transfer-Encoding: chunked\n\nTransfer-encoding: chunked is illegal with a HTTP/1.0 client.\n\n\t-hpa\n"},{"id":"86036","messageId":"20080803041014.GD27465@spearce.org","threadId":"14798","inReplyTo":"48952A62.6050709@zytor.com","subject":"Re: More on git over HTTP POST","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-03T04:10:14Z","receivedAt":"2008-08-03T04:10:14Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Junio C Hamano wrote:\n>>  For example, putting them [capabilities] on extra HTTP headers is probably Ok.\n>\n> I think that would be a mistake, just because it's one more thing for  \n> proxies to screw up on.\n\nI didn't realize we were in an era of proxies that are that\nbrain-damaged that they cannot relay the other headers.  The Amazon\nS3 service relies heavily upon their own extended headers to make\ntheir REST API work.  If proxies stripped that stuff out then the\nclient wouldn't work at all.\n\nIOW I had thought we were past this dark age of the Internet.\n\n> It's better to have negotiation information in  \n> the payload, before the \"real\" data.\n\nI guess I could do that.  At least for the really complex stuff.\n\n> Obviously one thing that needs to be included in each transaction is a  \n> transaction ID that will be reported back on the next transaction, since  \n> you can't rely on a persistent connection.\n\nNo.  That requires the server to maintain state.  We don't want to\ndo that if we can avoid it.  I would much rather have the clients\nhandle the state management as it simplifies the server side,\nespecially when you start talking about reverse proxies and/or\nload-balancers running in front of the server farm.\n\n-- \nShawn.\n"},{"id":"86037","messageId":"20080803041258.GE27465@spearce.org","threadId":"14798","inReplyTo":"48952B2E.3030209@zytor.com","subject":"Re: More on git over HTTP POST","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-03T04:12:58Z","receivedAt":"2008-08-03T04:12:58Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>> Chunked Transfer Encoding\n>> -------------------------\n>>\n>> For performance reasons the HTTP/1.1 chunked transfer encoding is\n>> used frequently to transfer variable length objects.  This avoids\n>> needing to produce large results in memory to compute the proper\n>> content-length.\n>\n> Note: you cannot rely on HTTP/1.1 being supported by an intermediate  \n> proxy; you might have to handle HTTP/1.0, where the data is terminated  \n> by connection close.\n\nWell, that proxy is going to be crying when we upload a 120M pack\nduring a push to it, and it buffers the damn thing to figure out\nthe proper Content-Length so it can convert an HTTP/1.1 client\nrequest into an HTTP/1.0 request to forward to the server.  That's\njust _stupid_.\n\nBut from the client side perspective the chunked transfer encoding\nis used only to avoid generating in advance and producing the\ncontent-length header.  I fully expect the encoding to disappear\n(e.g. in a proxy, or in the HTTP client library) before any sort\nof Git code gets its fingers on the data.\n\nHence to your other remark, I _do not_ rely upon the encoding\nboundaries to remain intact.  That is why there is Git pkt-line\nencodings inside of the HTTP data stream.  We can rely on the\npkt-line encoding being present, even if the HTTP chunks were\nmoved around (or removed entirely) by a proxy.\n\n-- \nShawn.\n"},{"id":"86042","messageId":"20080803064338.GA3686@glandium.org","threadId":"14798","inReplyTo":"20080803025602.GB27465@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-08-03T06:43:38Z","receivedAt":"2008-08-03T06:43:38Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Sat, Aug 02, 2008 at 07:56:02PM -0700, Shawn O. Pearce wrote:\n> Smart HTTP transfer protocols\n> =============================\n> \n> Git supports two HTTP based transfer protocols.  A \"dumb\" protocol\n> which requires only a standard HTTP server on the server end of the\n> connection, and a \"smart\" protocol which requires a Git aware CGI\n> (or server module).  This document describes the \"smart\" protocol.\n\nIf you want, I have a patch series that introduces a small API to make\nHTTP requests easier to make.\n\nMike\n"},{"id":"86043","messageId":"1217748317-70096-1-git-send-email-spearce@spearce.org","threadId":"14798","inReplyTo":"20080803025602.GB27465@spearce.org","subject":"[RFC 1/2] Add backdoor options to receive-pack for use in Git-aware CGI","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-03T07:25:16Z","receivedAt":"2008-08-03T07:25:16Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"The new --report-status flag forces the status report feature of\nthe push protocol to be enabled.  This can be useful in a CGI\nprogram that implements the server side of a \"smart\" Git-aware\nHTTP transport.  The CGI code can perform the selection of the\nfeature and ask receive-pack to enable it automatically.\n\nThe new --no-advertise-heads causes receive-pack to bypass its usual\ndisplay of known refs to the client, and instead immediately start\nreading the commands and pack from stdin.  This is useful in a CGI\nsituation where we want to hand off all input to receive-pack.\n\nSigned-off-by: Shawn O. Pearce <spearce@spearce.org>\n---\n receive-pack.c |   19 ++++++++++++++-----\n 1 files changed, 14 insertions(+), 5 deletions(-)\n\ndiff --git a/receive-pack.c b/receive-pack.c\nindex d44c19e..512eae6 100644\n--- a/receive-pack.c\n+++ b/receive-pack.c\n@@ -464,6 +464,7 @@ static int delete_only(struct command *cmd)\n \n int main(int argc, char **argv)\n {\n+\tint advertise_heads = 1;\n \tint i;\n \tchar *dir = NULL;\n \n@@ -472,7 +473,15 @@ int main(int argc, char **argv)\n \t\tchar *arg = *argv++;\n \n \t\tif (*arg == '-') {\n-\t\t\t/* Do flag handling here */\n+\t\t\tif (!strcmp(arg, \"--report-status\")) {\n+\t\t\t\treport_status = 1;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tif (!strcmp(arg, \"--no-advertise-heads\")) {\n+\t\t\t\tadvertise_heads = 0;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\n \t\t\tusage(receive_pack_usage);\n \t\t}\n \t\tif (dir)\n@@ -497,10 +506,10 @@ int main(int argc, char **argv)\n \telse if (0 <= receive_unpack_limit)\n \t\tunpack_limit = receive_unpack_limit;\n \n-\twrite_head_info();\n-\n-\t/* EOF */\n-\tpacket_flush(1);\n+\tif (advertise_heads) {\n+\t\twrite_head_info();\n+\t\tpacket_flush(1);\n+\t}\n \n \tread_head_info();\n \tif (commands) {\n-- \n1.6.0.rc1.221.g9ae23\n"},{"id":"86044","messageId":"1217748317-70096-2-git-send-email-spearce@spearce.org","threadId":"14798","inReplyTo":"1217748317-70096-1-git-send-email-spearce@spearce.org","subject":"[RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-03T07:25:17Z","receivedAt":"2008-08-03T07:25:17Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"This CGI can be loaded into an Apache server using ScriptAlias,\nsuch as with the following configuration:\n\n  LoadModule cgi_module /usr/libexec/apache2/mod_cgi.so\n  LoadModule alias_module /usr/libexec/apache2/mod_alias.so\n  ScriptAlias /git/ /usr/libexec/git-core/git-http-backend/\n\nRepositories are accessed via the translated PATH_INFO.\n\nThe CGI is backwards compatible with the dumb client, allowing the\nclient to detect the server's smarts by looking at the content-type\nreturned from \"GET /repo.git/info/refs\".  If the returned content\ntype is the magic application/x-git-refs type then the client can\nassume the server is Git-aware.\n\nSigned-off-by: Shawn O. Pearce <spearce@spearce.org>\n---\n .gitignore                                |    1 +\n Documentation/technical/http-protocol.txt |   88 +++++++++\n Makefile                                  |    1 +\n http-backend.c                            |  302 +++++++++++++++++++++++++++++\n 4 files changed, 392 insertions(+), 0 deletions(-)\n create mode 100644 Documentation/technical/http-protocol.txt\n create mode 100644 http-backend.c\n\ndiff --git a/.gitignore b/.gitignore\nindex a213e8e..02eaf3a 100644\n--- a/.gitignore\n+++ b/.gitignore\n@@ -51,6 +51,7 @@ git-gc\n git-get-tar-commit-id\n git-grep\n git-hash-object\n+git-http-backend\n git-http-fetch\n git-http-push\n git-imap-send\ndiff --git a/Documentation/technical/http-protocol.txt b/Documentation/technical/http-protocol.txt\nnew file mode 100644\nindex 0000000..6cb96f3\n--- /dev/null\n+++ b/Documentation/technical/http-protocol.txt\n@@ -0,0 +1,88 @@\n+Smart HTTP transfer protocols\n+=============================\n+\n+Git supports two HTTP based transfer protocols.  A \"dumb\" protocol\n+which requires only a standard HTTP server on the server end of the\n+connection, and a \"smart\" protocol which requires a Git aware CGI\n+(or server module).  This document describes the \"smart\" protocol.\n+\n+As a design feature smart servers automatically degrade to the\n+dumb protocol when speaking with a dumb client.  This may cause\n+more load to be placed on the server as the file GET requests are\n+handled by a CGI rather than the server itself.\n+\n+\n+Authentication\n+--------------\n+\n+Standard HTTP authentication is used, and must be configured and\n+enforced by the HTTP server software.\n+\n+Chunked Transfer Encoding\n+-------------------------\n+\n+For performance reasons the HTTP/1.1 chunked transfer encoding is\n+used frequently to transfer variable length objects.  This avoids\n+needing to produce large results in memory to compute the proper\n+content-length.\n+\n+Detecting Smart Servers\n+-----------------------\n+\n+HTTP clients can detect a smart Git-aware server by sending the\n+/info/refs request (below) to the server.  If the response has a\n+status of 200 and the magic application/x-git-refs content type\n+then the server can be assumed to be a smart Git-aware server.\n+\n+\n+Show Refs\n+---------\n+\n+Obtains the available refs from the remote repository.  The response\n+is a sequence of refs, one per line.  The actual format matches that\n+of the $GIT_DIR/info/refs file normally used by a \"dumb\" protocol.\n+\n+\tC: GET /path/to/repository.git/info/refs HTTP/1.0\n+\n+\tS: HTTP/1.1 200 OK\n+\tS: Content-Type: application/x-git-refs\n+\tS: Transfer-Encoding: chunked\n+\tS:\n+\tS: 62\n+\tS: 95dcfa3633004da0049d3d0fa03f80589cbcaf31 refs/heads/maint\n+\tS:\n+\tS: 63\n+\tS: d049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n+\tS:\n+\tS: 59\n+\tS: 2cb58b79488a98d2721cea644875a8dd0026b115 refs/heads/pu\n+\tS:\n+\n+Push Pack\n+---------\n+\n+Uploads a pack and updates refs.  The start of the stream is the\n+commands to update the refs and the remainder of the stream is the\n+pack file itself.  See git-receive-pack and its network protocol\n+in pack-protocol.txt, as this is essentially the same.\n+\n+\tC: POST /path/to/repository.git/receive-pack HTTP/1.0\n+\tC: Content-Type: application/x-git-receive-pack\n+\tC: Transfer-Encoding: chunked\n+\tC:\n+\tC: 103\n+\tC: 006395dcfa3633004da0049d3d0fa03f80589cbcaf31 d049f6c27a2244e12041955e262a404c7faba355 refs/heads/maint\n+\tC: 4\n+\tC: 0000\n+\tC: 12\n+\tC: PACK\n+\t...\n+\tC: 0\n+\n+\tS: HTTP/1.0 200 OK\n+\tS: Content-type: application/x-git-receive-pack-status\n+\tS: Transfer-Encoding: chunked\n+\tS:\n+\tS: ...<output of receive-pack>...\n+\n+\ndiff --git a/Makefile b/Makefile\nindex 52c67c1..3a93bf6 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -298,6 +298,7 @@ PROGRAMS += git-unpack-file$X\n PROGRAMS += git-update-server-info$X\n PROGRAMS += git-upload-pack$X\n PROGRAMS += git-var$X\n+PROGRAMS += git-http-backend$X\n \n # List built-in command $C whose implementation cmd_$C() is not in\n # builtin-$C.o but is linked in as part of some other command.\ndiff --git a/http-backend.c b/http-backend.c\nnew file mode 100644\nindex 0000000..a498f89\n--- /dev/null\n+++ b/http-backend.c\n@@ -0,0 +1,302 @@\n+#include \"cache.h\"\n+#include \"refs.h\"\n+#include \"pkt-line.h\"\n+#include \"object.h\"\n+#include \"tag.h\"\n+#include \"exec_cmd.h\"\n+#include \"run-command.h\"\n+\n+static const char content_type[] = \"Content-Type\";\n+static const char content_length[] = \"Content-Length\";\n+\n+static int can_chunk;\n+static char buffer[1000];\n+\n+static void send_status(unsigned code, const char *msg)\n+{\n+\tsize_t n;\n+\n+\tn = snprintf(buffer, sizeof(buffer), \"Status: %u %s\\r\\n\", code, msg);\n+\tif (n >= sizeof(buffer))\n+\t\tdie(\"protocol error: impossibly long header\");\n+\tsafe_write(1, buffer, n);\n+}\n+\n+static void send_header(const char *name, const char *value)\n+{\n+\tsize_t n;\n+\n+\tn = snprintf(buffer, sizeof(buffer), \"%s: %s\\r\\n\", name, value);\n+\tif (n >= sizeof(buffer))\n+\t\tdie(\"protocol error: impossibly long header\");\n+\tsafe_write(1, buffer, n);\n+}\n+\n+static void end_headers(void)\n+{\n+\tsafe_write(1, \"\\r\\n\", 2);\n+}\n+\n+static void send_nocaching(void)\n+{\n+\tconst char *proto = getenv(\"SERVER_PROTOCOL\");\n+\tif (!proto || !strcmp(proto, \"HTTP/1.0\"))\n+\t\tsend_header(\"Expires\", \"Mon, 17 Sep 2001 00:00:00 GMT\");\n+\telse\n+\t\tsend_header(\"Cache-Control\", \"no-cache\");\n+}\n+\n+static void send_connection_close(void)\n+{\n+\tsend_header(\"Connection\", \"close\");\n+}\n+\n+static void enable_chunking(void)\n+{\n+\tconst char *proto = getenv(\"SERVER_PROTOCOL\");\n+\n+\tcan_chunk = proto && strcmp(proto, \"HTTP/1.0\");\n+\tif (can_chunk)\n+\t\tsend_header(\"Transfer-Encoding\", \"chunked\");\n+\telse\n+\t\tsend_connection_close();\n+}\n+\n+#define hex(a) (hexchar[(a) & 15])\n+static void chunked_write(const char *fmt, ...)\n+{\n+\tstatic const char hexchar[] = \"0123456789abcdef\";\n+\tva_list args;\n+\tunsigned n;\n+\n+\tva_start(args, fmt);\n+\tn = vsnprintf(buffer + 6, sizeof(buffer) - 8, fmt, args);\n+\tva_end(args);\n+\tif (n >= sizeof(buffer) - 8)\n+\t\tdie(\"protocol error: impossibly long line\");\n+\n+\tif (can_chunk) {\n+\t\tunsigned len = n + 4, b = 4;\n+\n+\t\tbuffer[4] = '\\r';\n+\t\tbuffer[5] = '\\n';\n+\t\tbuffer[n + 6] = '\\r';\n+\t\tbuffer[n + 7] = '\\n';\n+\n+\t\twhile (n > 0) {\n+\t\t\tbuffer[--b] = hex(n);\n+\t\t\tn >>= 4;\n+\t\t\tlen++;\n+\t\t}\n+\n+\t\tsafe_write(1, buffer + b, len);\n+\t} else\n+\t\tsafe_write(1, buffer + 6, n);\n+}\n+\n+static void end_chunking(void)\n+{\n+\tstatic const char flush_chunk[] = \"0\\r\\n\\r\\n\";\n+\tif (can_chunk)\n+\t\tsafe_write(1, flush_chunk, strlen(flush_chunk));\n+}\n+\n+static void NORETURN invalid_request(const char *msg)\n+{\n+\tstatic const char header[] = \"error: \";\n+\n+\tsend_status(400, \"Bad Request\");\n+\tsend_header(content_type, \"text/plain\");\n+\tend_headers();\n+\n+\tsafe_write(1, header, strlen(header));\n+\tsafe_write(1, msg, strlen(msg));\n+\tsafe_write(1, \"\\n\", 1);\n+\n+\texit(0);\n+}\n+\n+static void not_found(void)\n+{\n+\tsend_status(404, \"Not Found\");\n+\tend_headers();\n+}\n+\n+static void server_error(void)\n+{\n+\tsend_status(500, \"Internal Error\");\n+\tend_headers();\n+}\n+\n+static void require_content_type(const char *need_type)\n+{\n+\tconst char *input_type = getenv(\"CONTENT_TYPE\");\n+\tif (!input_type || strcmp(input_type, need_type))\n+\t\tinvalid_request(\"Unsupported content-type\");\n+}\n+\n+static void do_GET_any_file(char *name)\n+{\n+\tconst char *p = git_path(\"%s\", name);\n+\tstruct stat sb;\n+\tuintmax_t remaining;\n+\tsize_t n;\n+\tint fd = open(p, O_RDONLY);\n+\n+\tif (fd < 0) {\n+\t\tnot_found();\n+\t\treturn;\n+\t}\n+\tif (fstat(fd, &sb) < 0) {\n+\t\tclose(fd);\n+\t\tserver_error();\n+\t\tdie(\"fstat on plain file failed\");\n+\t}\n+\tremaining = (uintmax_t)sb.st_size;\n+\n+\tn = snprintf(buffer, sizeof(buffer),\n+\t\t\"Content-Length: %\" PRIuMAX \"\\r\\n\", remaining);\n+\tif (n >= sizeof(buffer))\n+\t\tdie(\"protocol error: impossibly long header\");\n+\tsafe_write(1, buffer, n);\n+\tsend_header(content_type, \"application/octet-stream\");\n+\tend_headers();\n+\n+\twhile (remaining) {\n+\t\tn = xread(fd, buffer, sizeof(buffer));\n+\t\tif (n < 0)\n+\t\t\tdie(\"error reading from %s\", p);\n+\t\tn = safe_write(1, buffer, n);\n+\t\tif (n <= 0)\n+\t\t\tbreak;\n+\t}\n+\tclose(fd);\n+}\n+\n+static int show_one_ref(const char *name, const unsigned char *sha1,\n+\tint flag, void *cb_data)\n+{\n+\tstruct object *o = parse_object(sha1);\n+\tif (!o)\n+\t\treturn 0;\n+\n+\tchunked_write(\"%s\\t%s\\n\", sha1_to_hex(sha1), name);\n+\tif (o->type == OBJ_TAG) {\n+\t\to = deref_tag(o, name, 0);\n+\t\tif (!o)\n+\t\t\treturn 0;\n+\t\tchunked_write(\"%s\\t%s^{}\\n\", sha1_to_hex(o->sha1), name);\n+\t}\n+\n+\treturn 0;\n+}\n+\n+static void do_GET_info_refs(char *arg)\n+{\n+\tsend_header(content_type, \"application/x-git-refs\");\n+\tsend_nocaching();\n+\tenable_chunking();\n+\tend_headers();\n+\n+\tfor_each_ref(show_one_ref, NULL);\n+\tend_chunking();\n+}\n+\n+static void do_GET_info_packs(char *arg)\n+{\n+\tsize_t objdirlen = strlen(get_object_directory());\n+\tstruct packed_git *p;\n+\n+\tsend_nocaching();\n+\tenable_chunking();\n+\tend_headers();\n+\n+\tprepare_packed_git();\n+\tfor (p = packed_git; p; p = p->next) {\n+\t\tif (!p->pack_local)\n+\t\t\tcontinue;\n+\t\tchunked_write(\"P %s\\n\", p->pack_name + objdirlen + 6);\n+\t}\n+\tchunked_write(\"\\n\");\n+\tend_chunking();\n+}\n+\n+static void do_POST_receive_pack(char *arg)\n+{\n+\trequire_content_type(\"application/x-git-receive-pack\");\n+\tsend_header(content_type, \"application/x-git-receive-pack-status\");\n+\tsend_nocaching();\n+\tsend_connection_close();\n+\tend_headers();\n+\n+\texecl_git_cmd(\"receive-pack\",\n+\t\t\"--report-status\",\n+\t\t\"--no-advertise-heads\",\n+\t\t\".\",\n+\t\tNULL);\n+\tdie(\"Failed to start receive-pack\");\n+}\n+\n+static struct service_cmd {\n+\tconst char *method;\n+\tconst char *pattern;\n+\tvoid (*imp)(char *);\n+} services[] = {\n+\t{\"GET\", \"/info/refs$\", do_GET_info_refs},\n+\t{\"GET\", \"/objects/info/packs\", do_GET_info_packs},\n+\n+\t{\"GET\", \"/HEAD$\", do_GET_any_file},\n+\t{\"GET\", \"/objects/../.{38}$\", do_GET_any_file},\n+\t{\"GET\", \"/objects/pack/pack-[^/]*$\", do_GET_any_file},\n+\t{\"GET\", \"/objects/info/[^/]*$\", do_GET_any_file},\n+\n+\t{\"POST\", \"/receive-pack\", do_POST_receive_pack}\n+};\n+\n+int main(int argc, char **argv)\n+{\n+\tchar *input_method = getenv(\"REQUEST_METHOD\");\n+\tchar *dir = getenv(\"PATH_TRANSLATED\");\n+\tstruct service_cmd *cmd = NULL;\n+\tchar *cmd_arg = NULL;\n+\tint i;\n+\n+\tif (!input_method)\n+\t\tdie(\"No REQUEST_METHOD from server\");\n+\tif (!strcmp(input_method, \"HEAD\"))\n+\t\tinput_method = \"GET\";\n+\n+\tif (!dir)\n+\t\tdie(\"No PATH_TRANSLATED from server\");\n+\n+\tfor (i = 0; i < ARRAY_SIZE(services); i++) {\n+\t\tstruct service_cmd *c = &services[i];\n+\t\tregex_t re;\n+\t\tregmatch_t out[1];\n+\n+\t\tif (strcmp(input_method, c->method))\n+\t\t\tcontinue;\n+\t\tif (regcomp(&re, c->pattern, REG_EXTENDED))\n+\t\t\tdie(\"Bogus re in service table: %s\", c->pattern);\n+\t\tif (!regexec(&re, dir, 2, out, 0)) {\n+\t\t\tsize_t n = out[0].rm_eo - out[0].rm_so;\n+\t\t\tcmd = c;\n+\t\t\tcmd_arg = xmalloc(n);\n+\t\t\tstrncpy(cmd_arg, dir + out[0].rm_so + 1, n);\n+\t\t\tcmd_arg[n] = 0;\n+\t\t\tdir[out[0].rm_so] = 0;\n+\t\t\tbreak;\n+\t\t}\n+\t\tregfree(&re);\n+\t}\n+\n+\tif (!cmd)\n+\t\tinvalid_request(\"Unsupported query request\");\n+\n+\tsetup_path();\n+\tif (!enter_repo(dir, 0))\n+\t\tinvalid_request(\"Not a Git repository\");\n+\n+\tcmd->imp(cmd_arg);\n+\treturn 0;\n+}\n-- \n1.6.0.rc1.221.g9ae23\n"},{"id":"86048","messageId":"alpine.DEB.1.10.0808030105510.22058@asgard.lang.hm","threadId":"14798","inReplyTo":"20080803041014.GD27465@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-03T08:10:37Z","receivedAt":"2008-08-03T08:10:37Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Sat, 2 Aug 2008, Shawn O. Pearce wrote:\n\n> \n> \"H. Peter Anvin\" <hpa@zytor.com> wrote:\n>> Junio C Hamano wrote:\n>>>  For example, putting them [capabilities] on extra HTTP headers is probably Ok.\n>>\n>> I think that would be a mistake, just because it's one more thing for\n>> proxies to screw up on.\n>\n> I didn't realize we were in an era of proxies that are that\n> brain-damaged that they cannot relay the other headers.  The Amazon\n> S3 service relies heavily upon their own extended headers to make\n> their REST API work.  If proxies stripped that stuff out then the\n> client wouldn't work at all.\n>\n> IOW I had thought we were past this dark age of the Internet.\n\nactually, it's not just a matter of not getting 'past this dark age of the \nInternet', it's an issue that so many people are tunneling _everyting_ \nover http (including the bad guys tunneling malware) that proxies are \ngetting more aggressive then they have ever been before in pulling apart \nthe payload and analysing it before letting it get through to the far \nside.\n\nDavid Lang\n\n>> It's better to have negotiation information in\n>> the payload, before the \"real\" data.\n>\n> I guess I could do that.  At least for the really complex stuff.\n>\n>> Obviously one thing that needs to be included in each transaction is a\n>> transaction ID that will be reported back on the next transaction, since\n>> you can't rely on a persistent connection.\n>\n> No.  That requires the server to maintain state.  We don't want to\n> do that if we can avoid it.  I would much rather have the clients\n> handle the state management as it simplifies the server side,\n> especially when you start talking about reverse proxies and/or\n> load-balancers running in front of the server farm.\n>\n>\n"},{"id":"86061","messageId":"489596A2.9030704@zytor.com","threadId":"14798","inReplyTo":"20080803041014.GD27465@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-03T11:29:38Z","receivedAt":"2008-08-03T11:29:38Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> IOW I had thought we were past this dark age of the Internet.\n> \n\nIf we were, there wouldn't be a need for this project at all.  The whole \npurpose of it is to deal with corporate proxies that try to prevent \nactual communication because of \"security\", and it's really hard to \npredict what utterly arbitrary heuristics they have applied.\n\n\t-hpa\n"},{"id":"86060","messageId":"48959701.5050705@zytor.com","threadId":"14798","inReplyTo":"20080803041258.GE27465@spearce.org","subject":"Re: More on git over HTTP POST","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-03T11:31:13Z","receivedAt":"2008-08-03T11:31:13Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> But from the client side perspective the chunked transfer encoding\n> is used only to avoid generating in advance and producing the\n> content-length header.  I fully expect the encoding to disappear\n> (e.g. in a proxy, or in the HTTP client library) before any sort\n> of Git code gets its fingers on the data.\n> \n> Hence to your other remark, I _do not_ rely upon the encoding\n> boundaries to remain intact.  That is why there is Git pkt-line\n> encodings inside of the HTTP data stream.  We can rely on the\n> pkt-line encoding being present, even if the HTTP chunks were\n> moved around (or removed entirely) by a proxy.\n> \n\nExcellent.  I did not mean that as criticism, obviously, I just wanted \nthat to be clear.\n\nHTTP/1.1 does chunked encoding, and HTTP/1.0 does terminate on \nconnection close; both serve the same purpose.\n\n\t-hpa\n"},{"id":"86062","messageId":"489598C5.6060508@zytor.com","threadId":"14798","inReplyTo":"1217748317-70096-2-git-send-email-spearce@spearce.org","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-03T11:38:45Z","receivedAt":"2008-08-03T11:38:45Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> +#define hex(a) (hexchar[(a) & 15])\n> +static void chunked_write(const char *fmt, ...)\n> +{\n> +\tstatic const char hexchar[] = \"0123456789abcdef\";\n> +\tva_list args;\n> +\tunsigned n;\n> +\n> +\tva_start(args, fmt);\n> +\tn = vsnprintf(buffer + 6, sizeof(buffer) - 8, fmt, args);\n> +\tva_end(args);\n> +\tif (n >= sizeof(buffer) - 8)\n> +\t\tdie(\"protocol error: impossibly long line\");\n> +\n> +\tif (can_chunk) {\n> +\t\tunsigned len = n + 4, b = 4;\n> +\n> +\t\tbuffer[4] = '\\r';\n> +\t\tbuffer[5] = '\\n';\n> +\t\tbuffer[n + 6] = '\\r';\n> +\t\tbuffer[n + 7] = '\\n';\n> +\n> +\t\twhile (n > 0) {\n> +\t\t\tbuffer[--b] = hex(n);\n> +\t\t\tn >>= 4;\n> +\t\t\tlen++;\n> +\t\t}\n> +\n> +\t\tsafe_write(1, buffer + b, len);\n> +\t} else\n> +\t\tsafe_write(1, buffer + 6, n);\n> +}\n\nMaybe I am slightly confused, but I thought handling HTTP chunking for \nHTTP/1.1+ clients was usually done by Apache above the level of the CGI \nscript?\n\n\t-hpa\n"},{"id":"86064","messageId":"489599BE.3050000@zytor.com","threadId":"14798","inReplyTo":"alpine.DEB.1.10.0808030105510.22058@asgard.lang.hm","subject":"Re: More on git over HTTP POST","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-03T11:42:54Z","receivedAt":"2008-08-03T11:42:54Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"david@lang.hm wrote:\n> \n> actually, it's not just a matter of not getting 'past this dark age of \n> the Internet', it's an issue that so many people are tunneling \n> _everyting_ over http (including the bad guys tunneling malware) that \n> proxies are getting more aggressive then they have ever been before in \n> pulling apart the payload and analysing it before letting it get through \n> to the far side.\n> \n\n... which is of course because of said proxies that this is happening, too.\n\nThere are too many idiots out there building \"security software\" and \nrunning IT departments, that's really the bottom line.\n\nBy the way, I want to say *thank you* to Shawn for tackling this \nproject: this has been a major issue for kernel.org, and getting \nsomething like this deployed would be incredibly helpful.\n\n\t-hpa\n"},{"id":"86108","messageId":"20080803212516.GB31762@spearce.org","threadId":"14798","inReplyTo":"489598C5.6060508@zytor.com","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-03T21:25:16Z","receivedAt":"2008-08-03T21:25:16Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>> +#define hex(a) (hexchar[(a) & 15])\n>> +static void chunked_write(const char *fmt, ...)\n>> +{\n>\n> Maybe I am slightly confused, but I thought handling HTTP chunking for  \n> HTTP/1.1+ clients was usually done by Apache above the level of the CGI  \n> script?\n\nYou may be right.  Apache undoes the chunking during a POST before\nfeeding the data to the CGI script.  If we can omit this mess of\ncode from git-http-backend that's a good thing.\n\nThanks for the sanity check.\n\n-- \nShawn.\n"},{"id":"86114","messageId":"7vwsix7nhw.fsf@gitster.siamese.dyndns.org","threadId":"14798","inReplyTo":"1217748317-70096-2-git-send-email-spearce@spearce.org","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-08-03T22:16:43Z","receivedAt":"2008-08-03T22:16:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"I very much like it.\n\nBut could you be a bit more explicit than application/x-git-refs magic?  I\nsuspect very strongly that clueless server operators would advertise the\ntype on repositories statically hosted there, and would defeat the point\nof your patch.\n\nWe are not changing update-server-info so if we can find a place we can\nuse to hide the \"magic\", it would be a much more robust.\n\nPerhaps \"#\" comment line in info/refs that is ignored on the reading side\nbut update-server-info never generates on its own?\n\nOr perhaps sort the output differently from how update-server-info\nproduces its output, so that older client would not care but the magic\naware client can notice?\n"},{"id":"86145","messageId":"20080804035921.GB2963@spearce.org","threadId":"14798","inReplyTo":"7vwsix7nhw.fsf@gitster.siamese.dyndns.org","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-04T03:59:21Z","receivedAt":"2008-08-04T03:59:21Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> But could you be a bit more explicit than application/x-git-refs magic?  I\n> suspect very strongly that clueless server operators would advertise the\n> type on repositories statically hosted there, and would defeat the point\n> of your patch.\n\nThis is a very valid concern.  I started to worry about it myself\nlast night, but decided it was late enough and just wanted to start\nthe discussion on the list, extending JH's thread even further.\n \n> Perhaps \"#\" comment line in info/refs that is ignored on the reading side\n> but update-server-info never generates on its own?\n\nThis is a good idea.  I think anyone who consumes info/refs does\nso with the understanding that \"#\" comment lines exist, and should\nbe skipped, but this is not something that has been heavily tested\nin the wild yet.\n\nMy concern here goes back to the remark you made above. What if a\nserver owner mirrors a smart server by a non-Git aware device like\nwget?  They will now have a copy of the info/refs content which will\nsuggest we have Git smarts on the backend, but really it isn't there.\n\nPerhaps the smart server detection is something like:\n\n\tSmart Server Detection\n\t----------------------\n\n\tTo detect a smart (Git-aware) server a client sends an\n\tempty POST request to info/refs; if a 200 OK response is\n\treceived with the proper content type then the server can\n\tbe assumed to be Git-aware, and the result contains the\n\tcurrent info/refs data for that repository.\n\n\t\tC: POST /repository.git/info/refs HTTP/1.0\n\t\tC: Content-Length: 0\n\n\t\tS: HTTP/1.0 200 OK\n\t\tS: Content-Type: application/x-git-refs\n\t\tS:\n\t\tS: 95dcfa3633004da0049d3d0fa03f80589cbcaf31\trefs/heads/maint\n\nThen clients should just attempt this POST first before issuing\na GET info/refs.  Non Git-aware servers will issue an error code,\nand the client can retry with a standard GET request, and assume\nthe server isn't a newer style.\n\n-- \nShawn.\n"},{"id":"86167","messageId":"4896D19C.6040704@dawes.za.net","threadId":"14798","inReplyTo":"20080804035921.GB2963@spearce.org","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2008-08-04T09:53:32Z","receivedAt":"2008-08-04T09:53:32Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Shawn O. Pearce wrote:\n\n> Perhaps the smart server detection is something like:\n> \n> \tSmart Server Detection\n> \t----------------------\n> \n> \tTo detect a smart (Git-aware) server a client sends an\n> \tempty POST request to info/refs; if a 200 OK response is\n> \treceived with the proper content type then the server can\n> \tbe assumed to be Git-aware, and the result contains the\n> \tcurrent info/refs data for that repository.\n> \n> \t\tC: POST /repository.git/info/refs HTTP/1.0\n> \t\tC: Content-Length: 0\n> \n> \t\tS: HTTP/1.0 200 OK\n> \t\tS: Content-Type: application/x-git-refs\n> \t\tS:\n> \t\tS: 95dcfa3633004da0049d3d0fa03f80589cbcaf31\trefs/heads/maint\n> \n> Then clients should just attempt this POST first before issuing\n> a GET info/refs.  Non Git-aware servers will issue an error code,\n> and the client can retry with a standard GET request, and assume\n> the server isn't a newer style.\n> \n\nI don't understand why you would want to keep the commands in the URL \nwhen you are doing a POST?\n\nHow about something like:\n\n\tC: POST /repository.git/ HTTP/1.0\n\tC: Content-Length: <calculated>\n         C:\n         C: <whatever command you want>\n\nA dumb server will respond with:\n\n\tS: HTTP/1.1 405 Method not allowed\n\n(expected according to the RFC)\n\nOr\n\n\tS: HTTP/1.1 404 Not Found\n\n(resulting from testing against my own repo :-) )\n\nWhile a smart server will respond with a \"200 Ok\" and the results of the \ncommand.\n\nAlso, if everything is done via POST, you don't have to worry about a \nwget-cloned server appearing to be \"smart\", since no \"smarts\" will ever \nbe returned in response to a GET request (and to the best of my \nknowledge, wget can't mirror using POST).\n\nRogan\n"},{"id":"86168","messageId":"alpine.DEB.1.00.0808041208060.9611@pacific.mpi-cbg.de.mpi-cbg.de","threadId":"14798","inReplyTo":"4896D19C.6040704@dawes.za.net","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-08-04T10:08:29Z","receivedAt":"2008-08-04T10:08:29Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 4 Aug 2008, Rogan Dawes wrote:\n\n> Shawn O. Pearce wrote:\n> \n> > Perhaps the smart server detection is something like:\n> > \n> >  Smart Server Detection\n> >  ----------------------\n> > \n> >  To detect a smart (Git-aware) server a client sends an\n> >  empty POST request to info/refs; if a 200 OK response is\n> >  received with the proper content type then the server can\n> >  be assumed to be Git-aware, and the result contains the\n> >  current info/refs data for that repository.\n> > \n> > C: POST /repository.git/info/refs HTTP/1.0\n> > C: Content-Length: 0\n> > \n> > S: HTTP/1.0 200 OK\n> > S: Content-Type: application/x-git-refs\n> > S:\n> > S:95dcfa3633004da0049d3d0fa03f80589cbcaf31\trefs/heads/maint\n> > \n> > Then clients should just attempt this POST first before issuing\n> > a GET info/refs.  Non Git-aware servers will issue an error code,\n> > and the client can retry with a standard GET request, and assume\n> > the server isn't a newer style.\n> > \n> \n> I don't understand why you would want to keep the commands in the URL \n> when you are doing a POST?\n\nCaching.\n\nHth,\nDscho\n"},{"id":"86171","messageId":"4896D669.30402@dawes.za.net","threadId":"14798","inReplyTo":"alpine.DEB.1.00.0808041208060.9611@pacific.mpi-cbg.de.mpi-cbg.de","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2008-08-04T10:14:01Z","receivedAt":"2008-08-04T10:14:01Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Johannes Schindelin wrote:\n> Hi,\n> \n> On Mon, 4 Aug 2008, Rogan Dawes wrote:\n> \n>> I don't understand why you would want to keep the commands in the URL \n>> when you are doing a POST?\n> \n> Caching.\n> \n> Hth,\n> Dscho\n> \n\nIf you are expecting something to be cacheable, then should you not be \nusing a GET anyway?\n\nAnyway, from RFC 2616:\n\n> 13.10 Invalidation After Updates or Deletions\n> \n> ...\n> \n> Some HTTP methods MUST cause a cache to invalidate an entity. This is\n> either the entity referred to by the Request-URI, or by the Location\n > or Content-Location headers (if present). These methods are:\n> \n>       - PUT\n>       - DELETE\n>       - POST\n\nThis doesn't seem negotiable to me.\n\nUnless I am misunderstanding your \"Caching\" comment to mean \"To enable \ncaching\", as opposed to \"To prevent caching\"?\n\nRogan\n"},{"id":"86173","messageId":"alpine.DEB.1.00.0808041225050.9611@pacific.mpi-cbg.de.mpi-cbg.de","threadId":"14798","inReplyTo":"4896D669.30402@dawes.za.net","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-08-04T10:26:01Z","receivedAt":"2008-08-04T10:26:01Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 4 Aug 2008, Rogan Dawes wrote:\n\n> Johannes Schindelin wrote:\n> \n> > On Mon, 4 Aug 2008, Rogan Dawes wrote:\n> > \n> > > I don't understand why you would want to keep the commands in the \n> > > URL when you are doing a POST?\n> > \n> > Caching.\n> \n> If you are expecting something to be cacheable, then should you not be \n> using a GET anyway?\n\nYes.\n\nAnd I think the wget thing is not an issue: we should not try to prevent \nevery single idiocy.\n\nCiao,\nDscho\n"},{"id":"86196","messageId":"20080804144824.GB27666@spearce.org","threadId":"14798","inReplyTo":"4896D19C.6040704@dawes.za.net","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-04T14:48:24Z","receivedAt":"2008-08-04T14:48:24Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Rogan Dawes <lists@dawes.za.net> wrote:\n> Shawn O. Pearce wrote:\n>>\n>> \tSmart Server Detection\n>> \t----------------------\n>>\n>> \tTo detect a smart (Git-aware) server a client sends an\n>> \tempty POST request to info/refs; [...]\n>>\n>> \t\tC: POST /repository.git/info/refs HTTP/1.0\n>> \t\tC: Content-Length: 0\n>\n> I don't understand why you would want to keep the commands in the URL  \n> when you are doing a POST?\n\nWell, as Dscho pointed out this partly has to do with caching and\nthe transparent dumb server functionality.  By using the command in\nthe URL, and having the command match that of the dumb server file,\nits easier to emulate a dumb server and also to permit caching.\n\nCurrently git-http-backend requests no caching for info/refs, but\nI could see us tweaking that to permit several minutes of caching,\nespecially on big public sites like kernel.org.  Having info/refs\nreport stale by 5 minutes is not an issue when writes to there\nalready have a lag due to the master-slave mirroring system in use.\n\nBecause git-http-backend emulates a dumb server there is a command\ndispatch table based upon the URL submitted.  Thus we already have\nthe command dispatch behavior implemented in the URL and doing it\nin the POST body would only complicate the code further.\n\n> Also, if everything is done via POST, you don't have to worry about a  \n> wget-cloned server appearing to be \"smart\", since no \"smarts\" will ever  \n> be returned in response to a GET request (and to the best of my  \n> knowledge, wget can't mirror using POST).\n\nI think we fixed the wget-cloned server issue by requesting\nthat clients use POST /info/refs to identify a smart server.\nA wget-cloned repository will fail on this, and the client can\nfallback to GET /info/refs and assume it must use the object\nwalker to fetch (or WebDAV to push).  A smart server would\nrespond to the POST /info/refs request correctly and the\nclient would know its smart.\n\n-- \nShawn.\n"},{"id":"86209","messageId":"48972437.5050008@dawes.za.net","threadId":"14798","inReplyTo":"20080804144824.GB27666@spearce.org","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2008-08-04T15:45:59Z","receivedAt":"2008-08-04T15:45:59Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Shawn O. Pearce wrote:\n> Rogan Dawes <lists@dawes.za.net> wrote:\n>> Shawn O. Pearce wrote:\n>>> \tSmart Server Detection\n>>> \t----------------------\n>>>\n>>> \tTo detect a smart (Git-aware) server a client sends an\n>>> \tempty POST request to info/refs; [...]\n>>>\n>>> \t\tC: POST /repository.git/info/refs HTTP/1.0\n>>> \t\tC: Content-Length: 0\n>> I don't understand why you would want to keep the commands in the URL  \n>> when you are doing a POST?\n> \n> Well, as Dscho pointed out this partly has to do with caching and\n> the transparent dumb server functionality.  By using the command in\n> the URL, and having the command match that of the dumb server file,\n> its easier to emulate a dumb server and also to permit caching.\n> \n> Currently git-http-backend requests no caching for info/refs, but\n> I could see us tweaking that to permit several minutes of caching,\n> especially on big public sites like kernel.org.  Having info/refs\n> report stale by 5 minutes is not an issue when writes to there\n> already have a lag due to the master-slave mirroring system in use.\n\nFair enough, but what about the quote from RFC2616 that I posted in \nrebuttal to Dscho?\n\n > 13.10 Invalidation After Updates or Deletions\n >\n > ...\n >\n > Some HTTP methods MUST cause a cache to invalidate an entity. This is\n > either the entity referred to by the Request-URI, or by the Location\n > or Content-Location headers (if present). These methods are:\n >\n >       - PUT\n >       - DELETE\n >       - POST\n\nThis doesn't seem negotiable to me.\n\nFor those resources that are expected to be cacheable, the request \nshould be made using a GET.\n\n> Because git-http-backend emulates a dumb server there is a command\n> dispatch table based upon the URL submitted.  Thus we already have\n> the command dispatch behavior implemented in the URL and doing it\n> in the POST body would only complicate the code further.\n\nNot by a huge amount, surely?\n\nif (method == \"GET\") command = ...\nelse if (method == \"POST\") command = ...\ndispatch(command);\n\nRogan\n"},{"id":"86210","messageId":"20080804155956.GF27666@spearce.org","threadId":"14798","inReplyTo":"48972437.5050008@dawes.za.net","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-04T15:59:56Z","receivedAt":"2008-08-04T15:59:56Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Rogan Dawes <lists@dawes.za.net> wrote:\n> Shawn O. Pearce wrote:\n>> Currently git-http-backend requests no caching for info/refs [...]\n>\n> Fair enough, but what about the quote from RFC2616 that I posted in  \n> rebuttal to Dscho?\n>\n> > 13.10 Invalidation After Updates or Deletions\n> >\n> > ...\n> >\n> > Some HTTP methods MUST cause a cache to invalidate an entity. This is\n> > either the entity referred to by the Request-URI, or by the Location\n> > or Content-Location headers (if present). These methods are:\n> >\n> >       - PUT\n> >       - DELETE\n> >       - POST\n>\n> This doesn't seem negotiable to me.\n\nIts not negotiable.  POST requires no caching.  End of discussion.\n\n> For those resources that are expected to be cacheable, the request  \n> should be made using a GET.\n\nThat's exactly what we are doing.  Where caching is reasonable we are\nusing a GET request.  Where caching cannot be performed as the server\nstate is changing (e.g. actually updating refs) we are using POST.\nThat is entirely within the guidelines of the RFC.\n\nHowever we are \"abusing\" POST for \"POST /info/refs\" to detect a\nGit-aware HTTP server.  Sending POST to a static resource should\nalways fail.\n\n>> Because git-http-backend emulates a dumb server there is a command\n>> dispatch table based upon the URL submitted.  Thus we already have\n>> the command dispatch behavior implemented in the URL and doing it\n>> in the POST body would only complicate the code further.\n>\n> Not by a huge amount, surely?\n>\n> if (method == \"GET\") command = ...\n> else if (method == \"POST\") command = ...\n> dispatch(command);\n\nWell, true, we could do that.  But then we have to break the\ncommand name out of the input stream.  In some cases we may just be\nexec'ing another Git process and letting it handle the input stream.\nShoving the command name into the start of it just makes it that\nmuch harder to parse out.\n\nWe already have to handle splitting PATH_TRANSLATED into a pair of\n(GIT_DIR, command) so we can handle that for a GET.  We might as\nwell just use that very same code for POST to select the command.\n\nBesides, by placing the command name into the URL server admins can\nuse regex filters in their configurations to control access.  If we\nshove the command name into the body of a POST they cannot do this.\n\nI can see sites wanting to offer anonymous smart fetch, but require\npassword protected smart push on the same repository URL.  Slapping\na directive like:\n\n\t<Location ~ ^/git/.*/receive-pack$>\n\t\trequire valid-user\n\t\t...\n\t</Location>\n\nWould easily make Apache implement this for us.  Most modern HTTP\nservers should be able to be configured like this.\n\nOne of the problems with these RPC-in-HTTP systems is always the\nfact that the true nature of the action isn't visible in the method\nand URL, causing servers and proxies to have to parse the stream to\nimplement firewall rules.  Or to provide access control.  I'm trying\nto reuse as much of the access control support as possible from the\nHTTP server and put as little of it as possible into the backend CGI.\n\nSince the backend CGI is based upon git-receive-pack itself admins\ncan use the standard pre-receive/update hook pair to manage branch\nlevel security in a repository, while gross-level read/write can\nbe done in the server.\n\n-- \nShawn.\n"},{"id":"86212","messageId":"48972BEC.1060105@dawes.za.net","threadId":"14798","inReplyTo":"20080804155956.GF27666@spearce.org","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Rogan Dawes","fromEmail":"lists@dawes.za.net","sentAt":"2008-08-04T16:18:52Z","receivedAt":"2008-08-04T16:18:52Z","isPatch":false,"sender":{"key":"lists@dawes.za.net","avatar":null},"body":"Shawn O. Pearce wrote:\n> Rogan Dawes <lists@dawes.za.net> wrote:\n>> Shawn O. Pearce wrote:\n>>> Currently git-http-backend requests no caching for info/refs [...]\n>> Fair enough, but what about the quote from RFC2616 that I posted in  \n>> rebuttal to Dscho?\n>>\n>>> 13.10 Invalidation After Updates or Deletions\n>>>\n>>> ...\n>>>\n>>> Some HTTP methods MUST cause a cache to invalidate an entity. This is\n>>> either the entity referred to by the Request-URI, or by the Location\n>>> or Content-Location headers (if present). These methods are:\n>>>\n>>>       - PUT\n>>>       - DELETE\n>>>       - POST\n>> This doesn't seem negotiable to me.\n> \n> Its not negotiable.  POST requires no caching.  End of discussion.\n\nAha. So now I see the objective. I had misunderstood the intention to be \nto *allow* caching of POST'ed resources.\n\n>> For those resources that are expected to be cacheable, the request  \n>> should be made using a GET.\n> \n> That's exactly what we are doing.  Where caching is reasonable we are\n> using a GET request.  Where caching cannot be performed as the server\n> state is changing (e.g. actually updating refs) we are using POST.\n> That is entirely within the guidelines of the RFC.\n> \n> However we are \"abusing\" POST for \"POST /info/refs\" to detect a\n> Git-aware HTTP server.  Sending POST to a static resource should\n> always fail.\n\nRight. Either with a \"405 Method not supported\", or a \"404 Not found\". \nas I discovered.\n\n>>> Because git-http-backend emulates a dumb server there is a command\n>>> dispatch table based upon the URL submitted.  Thus we already have\n>>> the command dispatch behavior implemented in the URL and doing it\n>>> in the POST body would only complicate the code further.\n>> Not by a huge amount, surely?\n>>\n>> if (method == \"GET\") command = ...\n>> else if (method == \"POST\") command = ...\n>> dispatch(command);\n> \n> Well, true, we could do that.  But then we have to break the\n> command name out of the input stream.  In some cases we may just be\n> exec'ing another Git process and letting it handle the input stream.\n> Shoving the command name into the start of it just makes it that\n> much harder to parse out.\n\nFair enough. I had not thought about other uses for the input stream.\n\n> One of the problems with these RPC-in-HTTP systems is always the\n> fact that the true nature of the action isn't visible in the method\n> and URL, causing servers and proxies to have to parse the stream to\n> implement firewall rules.  Or to provide access control.  I'm trying\n> to reuse as much of the access control support as possible from the\n> HTTP server and put as little of it as possible into the backend CGI.\n> \n> Since the backend CGI is based upon git-receive-pack itself admins\n> can use the standard pre-receive/update hook pair to manage branch\n> level security in a repository, while gross-level read/write can\n> be done in the server.\n\nWorks for me!\n\nThanks for doing all the hard thinking for this feature :-)\n\nRogan\n"},{"id":"86228","messageId":"4897A6E4.3070508@zytor.com","threadId":"14798","inReplyTo":"20080804144824.GB27666@spearce.org","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-05T01:03:32Z","receivedAt":"2008-08-05T01:03:32Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> Currently git-http-backend requests no caching for info/refs, but\n> I could see us tweaking that to permit several minutes of caching,\n> especially on big public sites like kernel.org.  Having info/refs\n> report stale by 5 minutes is not an issue when writes to there\n> already have a lag due to the master-slave mirroring system in use.\n> \n> Because git-http-backend emulates a dumb server there is a command\n> dispatch table based upon the URL submitted.  Thus we already have\n> the command dispatch behavior implemented in the URL and doing it\n> in the POST body would only complicate the code further.\n> \n\nLet's put it this way: we're not seeing a huge amount of load from git \nprotocol requests, and I'm going to assume \"git+http\" protocol to be \nused only by sites behind braindamaged firewalls (everyone else would \nuse git protocol), so I'm not really all that worried about it.\n\nI'm not sure if \"emulating a dumb server\" is desirable at all; it seems \nlike it would at least in part defeat the purpose of minimizing the \ntransaction count and otherwise be as much of a \"smart\" server as the \nmedium permits.\n\n\t-hpa\n"},{"id":"86231","messageId":"20080805012459.GC32543@spearce.org","threadId":"14798","inReplyTo":"4897A6E4.3070508@zytor.com","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-05T01:24:59Z","receivedAt":"2008-08-05T01:24:59Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>>\n>> Currently git-http-backend requests no caching for info/refs [...]\n>\n> Let's put it this way: we're not seeing a huge amount of load from git  \n> protocol requests, and I'm going to assume \"git+http\" protocol to be  \n> used only by sites behind braindamaged firewalls (everyone else would  \n> use git protocol), so I'm not really all that worried about it.\n\nAgreed.  There's another application I want git+http for, but that\nmay never materialize.  Or maybe it will someday.  I just have to\nadopt a wait and see approach there.\n\n> I'm not sure if \"emulating a dumb server\" is desirable at all; it seems  \n> like it would at least in part defeat the purpose of minimizing the  \n> transaction count and otherwise be as much of a \"smart\" server as the  \n> medium permits.\n\nI think it is a really good idea.  Then clients don't have to worry\nabout which HTTP URL is the \"correct\" one for them to be using.\nEnd users will just magically get the smart git+http variant if\nboth sides support it and they need to use HTTP due to firewalls.\nClients will fall back onto the dumb protocol if the server doesn't\nsupport smart clones.  Older clients (pre git+http) will still be\nable to talk to a smart server, just slower.  This is nice for the\nend user.  No thinking is required.\n\nNever ask a human to do what a machine can do in less time.\n\nI think its just 1 extra HTTP hit per fetch/push done against\na dumb server.  On a smart server that first hit will also give\nus what we need to begin the conversation (the info/refs data).\nOn a dumb server its a wasted hit, but a dumb server is already\ndoing to suck.  One extra HTTP request against a dumb server is a\ndrop in the bucket.  Its also a pretty small request (an empty POST).\n\n-- \nShawn.\n"},{"id":"86233","messageId":"4897AE53.4030107@zytor.com","threadId":"14798","inReplyTo":"20080805012459.GC32543@spearce.org","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-05T01:35:15Z","receivedAt":"2008-08-05T01:35:15Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n>> I'm not sure if \"emulating a dumb server\" is desirable at all; it seems  \n>> like it would at least in part defeat the purpose of minimizing the  \n>> transaction count and otherwise be as much of a \"smart\" server as the  \n>> medium permits.\n> \n> I think it is a really good idea.  Then clients don't have to worry\n> about which HTTP URL is the \"correct\" one for them to be using.\n> End users will just magically get the smart git+http variant if\n> both sides support it and they need to use HTTP due to firewalls.\n> Clients will fall back onto the dumb protocol if the server doesn't\n> support smart clones.  Older clients (pre git+http) will still be\n> able to talk to a smart server, just slower.  This is nice for the\n> end user.  No thinking is required.\n> \n> Never ask a human to do what a machine can do in less time.\n> \n> I think its just 1 extra HTTP hit per fetch/push done against\n> a dumb server.  On a smart server that first hit will also give\n> us what we need to begin the conversation (the info/refs data).\n> On a dumb server its a wasted hit, but a dumb server is already\n> doing to suck.  One extra HTTP request against a dumb server is a\n> drop in the bucket.  Its also a pretty small request (an empty POST).\n> \n\nNot arguing that URL compatibility isn't a good thing, but there are \nother ways to accomplish it, too.  After detecting either a smart or \ndumb server, we can use a redirect to point them to a different URL, as \nappropriate.\n\nFurthermore, in the case of round-robin sites like kernel.org, this is \nactually *mandatory* in the case of a stateful server (we need a \nredirect to a server-specific URL), and highly recommended in the case \nof a stateless server (because of potential skew.)\n\n\t-hpa\n"},{"id":"86237","messageId":"20080805015717.GB383@spearce.org","threadId":"14798","inReplyTo":"4897AE53.4030107@zytor.com","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-05T01:57:17Z","receivedAt":"2008-08-05T01:57:17Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>>\n>> I think it is a really good idea.  Then clients don't have to worry\n>> about which HTTP URL is the \"correct\" one for them to be using.\n>\n> Not arguing that URL compatibility isn't a good thing, but there are  \n> other ways to accomplish it, too.  After detecting either a smart or  \n> dumb server, we can use a redirect to point them to a different URL, as  \n> appropriate.\n\nI'm not sure this is necessary.\n\nOf course it all comes down to \"how does an admin map Git repositories\ninto the URL space of the server\"?\n\nI thought it would be simple if the admin was able to map\nrepositories using a ScriptAlias and allow the server to perform\npath info translation to give us the filesystem location of the\nrepository.  Then we don't have to configure our own map of the\navailable Git repositories.\n\nOnce you do that though you now have the URL space associated with\nthat repository served by a CGI.  For older clients we need to\neither serve them the file, or issue a redirect to serve the file.\nThe redirect is messy because we need some configuration to explain\nwhere the files are available in the server's URL space.\n\nOr you go the other way, and have newer git+http clients try to\nfind the git aware server by a redirect.  Again we have to explain\nwhere that git aware server is in the URL space of the server.\n\n*sigh*\n\n> Furthermore, in the case of round-robin sites like kernel.org, this is  \n> actually *mandatory* in the case of a stateful server (we need a  \n> redirect to a server-specific URL), and highly recommended in the case  \n> of a stateless server (because of potential skew.)\n\nWell, the git+http protocol will hold all state in the client, making\neach RPC a stateless RPC operation.  The only issue is then dealing with\nskew in a server farm.\n\nI guess we need to ask client implementations to honor a redirect\non the first request and reuse that new base URL for all subsequent\nrequests that are part of the same \"operation\".  Then server farms\ncan issue a redirect to a server-specific hostname if a client\ncomes in with a round-robin DNS hostname, thus ensuring that for\nthis current operation there isn't skew.\n\n-- \nShawn.\n"},{"id":"86238","messageId":"4897B4A5.4030700@zytor.com","threadId":"14798","inReplyTo":"20080805015717.GB383@spearce.org","subject":"Re: [RFC 2/2] Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-05T02:02:13Z","receivedAt":"2008-08-05T02:02:13Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> I guess we need to ask client implementations to honor a redirect\n> on the first request and reuse that new base URL for all subsequent\n> requests that are part of the same \"operation\".  Then server farms\n> can issue a redirect to a server-specific hostname if a client\n> comes in with a round-robin DNS hostname, thus ensuring that for\n> this current operation there isn't skew.\n> \n\nEither that, or you can pass a \"chase URL\" in the payload of the \nrequest... it's more or less the same concept.\n\n\t-hpa\n"},{"id":"86996","messageId":"48A23F56.7000607@zytor.com","threadId":"14798","inReplyTo":"4897B4A5.4030700@zytor.com","subject":"Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-13T01:56:38Z","receivedAt":"2008-08-13T01:56:38Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Anything we can do to keep this moving forward?  I was extremely \nencouraged with the fast progress on this; this would be great to get to \nthe point where we (kernel.org) can deploy it at least for testing.\n\n\t-hpa\n"},{"id":"86997","messageId":"20080813023759.GA5760@spearce.org","threadId":"14798","inReplyTo":"48A23F56.7000607@zytor.com","subject":"Re: Add Git-aware CGI for Git-aware smart HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-13T02:37:59Z","receivedAt":"2008-08-13T02:37:59Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Anything we can do to keep this moving forward?  I was extremely  \n> encouraged with the fast progress on this; this would be great to get to  \n> the point where we (kernel.org) can deploy it at least for testing.\n\nSorry, I dropped it with my egit work.  I'll pick it up again and\ntry to continue it further.  I left off trying to implement the\npush client and saying \"damn, jgit is better structured to make this\nsort of change than C git\" and decided it was too late at night to\ncontinue it more.  That was like a week ago.\n\n-- \nShawn.\n"}]}