{"thread":{"id":"15201","subject":"Git-aware HTTP transport","startedAt":"2008-08-26T01:26:43Z","lastAt":"2013-02-13T15:29:03Z","messageCount":57,"participants":["Shawn O. Pearce","H. Peter Anvin","david@lang.hm","Imran M Yousuf","Nicolas Pitre","Junio C Hamano","Daniel Stenberg","Mike Hommey","Tarmigan","Scott Chacon"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"88531","messageId":"20080826012643.GD26523@spearce.org","threadId":"15201","inReplyTo":null,"subject":"Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-26T01:26:43Z","receivedAt":"2008-08-26T01:26:43Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"I spent some time on Friday kicking around how the fetch part of\nan HTTP protocol could be implemented.  What I seem to have settled\non at this point in time is a more condensed version of the native\ngit protocol, batched into 256 commit blocks.\n\n--8<--\nSmart HTTP transfer protocols\n=============================\n\nGit supports two HTTP based transfer protocols.  A \"dumb\" protocol\nwhich requires only a standard HTTP server on the server end of the\nconnection, and a \"smart\" protocol which requires a Git aware CGI\n(or server module).  This document describes the \"smart\" protocol.\n\nAs a design feature smart clients can automatically translate and\nupgrade \"dumb\" protocol URLs.  This permits all users to have the\nsame published URL, with the peers automatically choosing to use\nthe most efficient transport available to them.\n\nAuthentication\n--------------\n\nStandard HTTP authentication is used if authentication is required\nto access a repository, and must be configured and enforced by the\nHTTP server software itself.\n\nStateless\n---------\n\nThe protocol, much like its underlying HTTP, is stateless, from the\nperspective of the HTTP server side.  All state must be retained and\nmanaged by the client.  This permits round-robin load-balancing on\nthe server side, among many other implementation details.\n\nHTTP/1.1 Preference\n-------------------\n\nFor performance reasons the HTTP/1.1 chunked transfer encoding\nis used whenever possible to transfer variable length objects.\nThis avoids needing to produce large results in memory to compute\nthe proper content-length.\n\nDetecting Smart Servers\n-----------------------\n\nHTTP clients can detect a smart Git-aware server by HEADing\n$repo/backend.git-http and looking for a 302 redirect to the\nrepository's smart service URL:\n\n\tC: HEAD /path/to/repository.git/backend.git-http HTTP/1.1\n\n\tS: HTTP/1.1 302 Found\n\tS: Location: /git/path/to/repository.git\n\nA dumb server would respond with a 304 Not Found (or 200 OK).\n\nSmart servers may send a redirect to any URL that does not\ncontain query args (e.g. \"foo?repo=path.git\" is invalid).\nThe URL must be sufficient to provide the location of the\nrepository to the smart service code.\n\nA valid redirect can be to yourself, for example:\n\n\tC: HEAD /path/to/repository.git/backend.git-http HTTP/1.1\n\n\tS: HTTP/1.1 302 Found\n\tS: Location: /path/to/repository.git/backend.git-http/.\n\nAll subsequent communcation for this transaction is done through\nthe smart service URL ($ssurl), not the original URL.\n\nGET $ssurl/refs\n---------------\n\nObtains the available refs from the remote repository.  The response\nis a sequence of refs, one per Git packet line.  The final packet\nline has a length of 0 to indicate the end.  This is basically\nthe same protocol that is used by the git-upload-pack service to\nadvertise the available refs.\n\n\tC: GET $ssurl/refs HTTP/1.1\n\n\tS: HTTP/1.1 200 OK\n\tS: Content-Type: application/x-git-refs\n\tS:\n\tS: 003295dcfa3633004da0049d3d0fa03f80589cbcaf31 HEAD\n\tS: 003e95dcfa3633004da0049d3d0fa03f80589cbcaf31 refs/heads/maint\n\tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n\tS: 003b2cb58b79488a98d2721cea644875a8dd0026b115 refs/heads/pu\n\tS: 0000\n\nPOST $ssurl/upload-pack\n-----------------------\n\nPrepares an estimated minimal pack to transfer new objects to the\nclient.\n\nThe computation to select the minimal pack proceeds as follows\n(c = client, s = server):\n\n init step:\n (c) Use /refs to obtain the advertised refs.\n (c) Place any object seen in /refs into set ADVERTISED.\n\n (c) Build a set, WANT, of the objects from ADVERTISED the client\n     wants to fetch, based on what it saw from /refs.\n\n (c) Start a queue, C_PENDING, ordered by commit time (popping newest\n     first).  Add all client refs.  When a commit is popped from the\n     queue its parents should be automatically inserted back.  Commits\n     should only enter the queue once.\n\n one compute step:\n (c) Send a /upload-pack request:\n\n\tC: POST $ssurl/upload-pack HTTP/1.1\n\tC: Content-Type: application/x-git-uploadpack\n\tC: Content-Length: ...\n\tC:\n\tC: 0009want\n\tC: 0xxx<WANT list>\n\tC: 000bcommon\n\tC: 0xxx<COMMON list>\n\tC: 0009have\n\tC: 0xxx<HAVE list>\n\tC: 0000\n\n     The stream is organized into \"sections\", where each section is\n     composed of two git pkt-lines.  The first pkt-line provides the\n     name of the section (\"want\", \"have\", \"common\").  The second\n     pkt-line has the binary SHA-1 ids which compose that section.\n\n     The \"want\" section is required.  The other sections (\"have\",\n     \"common\") are optional.  A missing \"want\" section should be\n     answered with a \"400 Bad Request\".\n\n     Sections must appear in the following order, if they appear\n     at all in the request stream:\n\n       * want\n       * common\n       * have\n\n     Each section may appear multiple times.  Client implementions\n     are encouraged to use as few sections as possible, however the\n     limit of 64k per pkt-line limits the number of ids to 3,276 per\n     section entry.\n\n     The stream is terminated by a pkt-line flush (\"0000\").\n\n     The HAVE list is created by popping the first 256 commits\n     from C_PENDING.  Less can be supplied if C_PENDING empties.\n\n  (s) Parse the /upload-pack request.\n\n      Verify all objects in WANT are reachable from refs.  As\n      this may require walking backwards through history to\n      the very beginning on invalid requests the server may\n      use a reasonable limit of commits (e.g. 1000) walked\n      beyond any ref tip before giving up.\n\n      If any WANT object is not reachable, send a 409 error:\n\n\tS: HTTP/1.1 409 Conflict\n\tS: Content-Type: application/x-git-error\n\tS:\n\tS: %s not reachable\n\n     Create an empty list, S_COMMON.\n\n     If 'common' was sent:\n\n     Load all objects into S_COMMON.\n\n     If 'have' was sent:\n\n     Loop through the objects in the order supplied by the client.\n     For each object, if the server has the object reachable from\n     a ref, add it to S_COMMON.  If a commit is added to S_COMMON,\n     do not add any ancestors, even if they also appear in HAVE.\n\n  (s) Send the /upload-pack response:\n\n\tS: HTTP/1.1 200 OK\n\tS: Content-Type: application/x-git-uploadpack\n\n\tS: 000bcommon\n\tS: 0xxx<S_COMMON list>\n\tS: 0000\n\n     The stream formatting rules are the same as the request.\n\n     The section \"common\" details the contents of S_COMMON,\n     that is all objects from HAVE that the server also has.\n\n     If the server has found a closed set of objects to pack,\n     it replies with the pack and not x-git-uploadpack response.\n\n\tS: HTTP/1.1 200 OK\n\tS: Content-Type: application/x-git-pack\n\n\tS: 000c.PACK...\n\n     The returned stream is the side-band-64k protocol supported\n     by the git-upload-pack service, and the pack is embedded into\n     stream 1.  Progress messages from the server side may appear\n     in stream 2.\n\n  (c) Parse the /upload-pack response:\n\n      If the Content-Type is application/x-git-uploadpack:\n\n      Reset COMMON to the items in S_COMMON.  The new S_COMMON\n      should be a superset of the existing COMMON set.\n\n      Remove all items in S_COMMON, and all of their ancestors,\n      from PENDING.\n\n      Do another /compute-common step.\n\n      If the Content-Type is application/x-git-pack:\n\n      Process the pack stream and update the local refs.\n\n\nPOST $ssurl/receive-pack\n------------------------\n\nTBD: Still a work in progress.\n\nUploads a pack and updates refs.  The start of the stream is the\ncommands to update the refs and the remainder of the stream is the\npack file itself.  See git-receive-pack and its network protocol\nin pack-protocol.txt, as this is essentially the same.\n\n\tC: POST /path/to/repository.git/receive-pack HTTP/1.0\n\tC: Content-Type: application/x-git-receivepack\n\tC: Transfer-Encoding: chunked\n\tC:\n\tC: 103\n\tC: 006395dcfa3633004da0049d3d0fa03f80589cbcaf31 d049f6c27a2244e12041955e262a404c7faba355 refs/heads/maint\n\tC: 4\n\tC: 0000\n\tC: 12\n\tC: PACK\n\t...\n\tC: 0\n\n\tS: HTTP/1.0 200 OK\n\tS: Content-type: application/x-git-receive-pack-status\n\tS: Transfer-Encoding: chunked\n\tS:\n\tS: ...<output of receive-pack>...\n\n\n-- \nShawn.\n"},{"id":"88533","messageId":"48B36BCA.8060103@zytor.com","threadId":"15201","inReplyTo":"20080826012643.GD26523@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-26T02:34:50Z","receivedAt":"2008-08-26T02:34:50Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"This is a bit more detailed review than I have done in the past (I'm \nactually in town...) so please pardon me for commenting on things that \nhas been dealt with in the past.\n\nOverall, I really keeping in mind that you're doing a layer on top of \nHTTP, and resist the temptation to delve into details how things are to \nbe implemented at the HTTP level.\n\nThe HTTP layer, furthermore, has a couple of important properties:\n\n- GET requests may be cached, even if you tell it not to.\n   (Some proxies, transparent or not, ignore caching directives.)\n- POST requests generally will not.\n\nSo don't implement things as GET requests unless you genuinely can deal \nwith the request being cached.  Using POST requests throughout seems \nlike a safer bet to me; on the other hand, since the only use of GET is \nobtaining a list of refs the worst thing that can happen, I presume, is \nadditional latency for the user behind the proxy.\n\nAgain, please don't take this as anything other than technical review \ntype criticism.  I'm obviously really happy about the project and want \nto see it happen.\n\nI do have one, very specific question: would the load on the server be \nlower if it was using a stateful protocol (like the standard git \nprotocol)?  If there is value in the server maintaining state, then I \nwould like to suggest a slightly different protocol.\n\n\t-hpa\n\n\nShawn O. Pearce wrote:\n> \n> HTTP/1.1 Preference\n> -------------------\n> \n> For performance reasons the HTTP/1.1 chunked transfer encoding\n> is used whenever possible to transfer variable length objects.\n> This avoids needing to produce large results in memory to compute\n> the proper content-length.\n> \n\nThis piece is unnecessary; it's a detail of the underlying HTTP layer.\n\n> Detecting Smart Servers\n> -----------------------\n> \n> HTTP clients can detect a smart Git-aware server by HEADing\n> $repo/backend.git-http and looking for a 302 redirect to the\n> repository's smart service URL:\n> \n> \tC: HEAD /path/to/repository.git/backend.git-http HTTP/1.1\n> \n> \tS: HTTP/1.1 302 Found\n> \tS: Location: /git/path/to/repository.git\n> \n> A dumb server would respond with a 304 Not Found (or 200 OK).\n> \n> Smart servers may send a redirect to any URL that does not\n> contain query args (e.g. \"foo?repo=path.git\" is invalid).\n> The URL must be sufficient to provide the location of the\n> repository to the smart service code.\n> \n> A valid redirect can be to yourself, for example:\n> \n> \tC: HEAD /path/to/repository.git/backend.git-http HTTP/1.1\n> \n> \tS: HTTP/1.1 302 Found\n> \tS: Location: /path/to/repository.git/backend.git-http/.\n> \n> All subsequent communcation for this transaction is done through\n> the smart service URL ($ssurl), not the original URL.\n\nI actually suggest embedding the forwarding URL into an ordinary \npayload.  Instead of a HEAD request here, then do a GET (or, even \nbetter, POST) and get the redirected URL in return.\n\nWhy?  Because it's common enough to redirect entire trees, and use of \nHTTP-layer redirections here is an unnecessary layering violation.\n\nIf you insist on using a HTTP status code, I would claim that 303 is a \nbetter status code.\n\n> GET $ssurl/refs\n> ---------------\n> \n> Obtains the available refs from the remote repository.  The response\n> is a sequence of refs, one per Git packet line.  The final packet\n> line has a length of 0 to indicate the end.  This is basically\n> the same protocol that is used by the git-upload-pack service to\n> advertise the available refs.\n> \n> \tC: GET $ssurl/refs HTTP/1.1\n> \n> \tS: HTTP/1.1 200 OK\n> \tS: Content-Type: application/x-git-refs\n> \tS:\n> \tS: 003295dcfa3633004da0049d3d0fa03f80589cbcaf31 HEAD\n> \tS: 003e95dcfa3633004da0049d3d0fa03f80589cbcaf31 refs/heads/maint\n> \tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n> \tS: 003b2cb58b79488a98d2721cea644875a8dd0026b115 refs/heads/pu\n> \tS: 0000\n> \n> POST $ssurl/upload-pack\n> -----------------------\n> \n> Prepares an estimated minimal pack to transfer new objects to the\n> client.\n> \n> The computation to select the minimal pack proceeds as follows\n> (c = client, s = server):\n> \n>  init step:\n>  (c) Use /refs to obtain the advertised refs.\n>  (c) Place any object seen in /refs into set ADVERTISED.\n> \n>  (c) Build a set, WANT, of the objects from ADVERTISED the client\n>      wants to fetch, based on what it saw from /refs.\n> \n>  (c) Start a queue, C_PENDING, ordered by commit time (popping newest\n>      first).  Add all client refs.  When a commit is popped from the\n>      queue its parents should be automatically inserted back.  Commits\n>      should only enter the queue once.\n> \n>  one compute step:\n>  (c) Send a /upload-pack request:\n> \n> \tC: POST $ssurl/upload-pack HTTP/1.1\n> \tC: Content-Type: application/x-git-uploadpack\n\nInstead of \"application/x-git-blah\" I would suggest using \n\"application/x-git; action=blah\"; that way we can probably even register \napplication/git with IANA.\n\n> \tC: Content-Length: ...\n> \tC:\n> \tC: 0009want\n> \tC: 0xxx<WANT list>\n> \tC: 000bcommon\n> \tC: 0xxx<COMMON list>\n> \tC: 0009have\n> \tC: 0xxx<HAVE list>\n> \tC: 0000\n> \n>      The stream is organized into \"sections\", where each section is\n>      composed of two git pkt-lines.  The first pkt-line provides the\n>      name of the section (\"want\", \"have\", \"common\").  The second\n>      pkt-line has the binary SHA-1 ids which compose that section.\n> \n>      The \"want\" section is required.  The other sections (\"have\",\n>      \"common\") are optional.  A missing \"want\" section should be\n>      answered with a \"400 Bad Request\".\n> \n>      Sections must appear in the following order, if they appear\n>      at all in the request stream:\n> \n>        * want\n>        * common\n>        * have\n> \n>      Each section may appear multiple times.  Client implementions\n>      are encouraged to use as few sections as possible, however the\n>      limit of 64k per pkt-line limits the number of ids to 3,276 per\n>      section entry.\n> \n>      The stream is terminated by a pkt-line flush (\"0000\").\n> \n>      The HAVE list is created by popping the first 256 commits\n>      from C_PENDING.  Less can be supplied if C_PENDING empties.\n> \n>   (s) Parse the /upload-pack request.\n> \n>       Verify all objects in WANT are reachable from refs.  As\n>       this may require walking backwards through history to\n>       the very beginning on invalid requests the server may\n>       use a reasonable limit of commits (e.g. 1000) walked\n>       beyond any ref tip before giving up.\n> \n>       If any WANT object is not reachable, send a 409 error:\n\nAgain, I think the 409 error code here is an unnecessary layering \nviolation.  It's simply Yet Another Thing that an HTTP proxy can screw \nup.  Having the HTTP server return a normal 200 reply (meaning that the \n*transport* succeeded) and have the error embedded in a lower layer \nshould avoid that class of problems.\n\n> \tS: HTTP/1.1 409 Conflict\n> \tS: Content-Type: application/x-git-error\n> \tS:\n> \tS: %s not reachable\n> \n>      Create an empty list, S_COMMON.\n> \n>      If 'common' was sent:\n> \n>      Load all objects into S_COMMON.\n> \n>      If 'have' was sent:\n> \n>      Loop through the objects in the order supplied by the client.\n>      For each object, if the server has the object reachable from\n>      a ref, add it to S_COMMON.  If a commit is added to S_COMMON,\n>      do not add any ancestors, even if they also appear in HAVE.\n> \n>   (s) Send the /upload-pack response:\n> \n> \tS: HTTP/1.1 200 OK\n> \tS: Content-Type: application/x-git-uploadpack\n> \n> \tS: 000bcommon\n> \tS: 0xxx<S_COMMON list>\n> \tS: 0000\n> \n>      The stream formatting rules are the same as the request.\n> \n>      The section \"common\" details the contents of S_COMMON,\n>      that is all objects from HAVE that the server also has.\n> \n>      If the server has found a closed set of objects to pack,\n>      it replies with the pack and not x-git-uploadpack response.\n> \n> \tS: HTTP/1.1 200 OK\n> \tS: Content-Type: application/x-git-pack\n> \n> \tS: 000c.PACK...\n> \n>      The returned stream is the side-band-64k protocol supported\n>      by the git-upload-pack service, and the pack is embedded into\n>      stream 1.  Progress messages from the server side may appear\n>      in stream 2.\n> \n>   (c) Parse the /upload-pack response:\n> \n>       If the Content-Type is application/x-git-uploadpack:\n> \n>       Reset COMMON to the items in S_COMMON.  The new S_COMMON\n>       should be a superset of the existing COMMON set.\n> \n>       Remove all items in S_COMMON, and all of their ancestors,\n>       from PENDING.\n> \n>       Do another /compute-common step.\n> \n>       If the Content-Type is application/x-git-pack:\n> \n>       Process the pack stream and update the local refs.\n> \n> \n> POST $ssurl/receive-pack\n> ------------------------\n> \n> TBD: Still a work in progress.\n> \n> Uploads a pack and updates refs.  The start of the stream is the\n> commands to update the refs and the remainder of the stream is the\n> pack file itself.  See git-receive-pack and its network protocol\n> in pack-protocol.txt, as this is essentially the same.\n> \n> \tC: POST /path/to/repository.git/receive-pack HTTP/1.0\n> \tC: Content-Type: application/x-git-receivepack\n> \tC: Transfer-Encoding: chunked\n> \tC:\n> \tC: 103\n> \tC: 006395dcfa3633004da0049d3d0fa03f80589cbcaf31 d049f6c27a2244e12041955e262a404c7faba355 refs/heads/maint\n> \tC: 4\n> \tC: 0000\n> \tC: 12\n> \tC: PACK\n> \t...\n> \tC: 0\n> \n> \tS: HTTP/1.0 200 OK\n> \tS: Content-type: application/x-git-receive-pack-status\n> \tS: Transfer-Encoding: chunked\n> \tS:\n> \tS: ...<output of receive-pack>...\n> \n> \n"},{"id":"88536","messageId":"20080826034544.GA32334@spearce.org","threadId":"15201","inReplyTo":"48B36BCA.8060103@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-26T03:45:45Z","receivedAt":"2008-08-26T03:45:45Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> So don't implement things as GET requests unless you genuinely can deal  \n> with the request being cached.  Using POST requests throughout seems  \n> like a safer bet to me; on the other hand, since the only use of GET is  \n> obtaining a list of refs the worst thing that can happen, I presume, is  \n> additional latency for the user behind the proxy.\n\nThis is a good point.  There is probably not any reason to cache the\nrefs content if we don't also support caching the pack files.  So in\nthis latest draft I have moved the ref listing to also be a POST.\n\n> I do have one, very specific question: would the load on the server be  \n> lower if it was using a stateful protocol (like the standard git  \n> protocol)?  If there is value in the server maintaining state, then I  \n> would like to suggest a slightly different protocol.\n\nIts possible the load would be lower, but it would complicate the\nserver implementation considerably.  Looking at the algorithm used\nto compute upload-pack the server has relatively little to do in\nany request.\n\nValidation of WANT should be just matching the requested objects\nagainst the refs; most clients will be asking for the current tips.\nA client that started the process just before a fast-forward push\nmay incur at most a few hundred commit walk during its last couple\nof computation round-trips.\n\nMarking commits COMMON is just a matter of looking them up in the\ndatabase and setting their flags.\n\nEvaluation of the HAVE list avoids duplicates in a well-behaved\nclient.  So the server sees each candidate commit from a client\nonly once, even if it spans multiple upload-pack requests.\n\nSo I think the cost may actually break even with a stateful protocol\nif we imagine that the server is actually a farm of systems and\nsimple round-robin load-balancing is being done in front of the\nGit-aware server.\n\nI'd really like to keep the protocol stateless on the server side, as\nthis makes it easier to embed into certain commerical server farms.\n\n> Shawn O. Pearce wrote:\n>>\n>> HTTP/1.1 Preference\n>\n> This piece is unnecessary; it's a detail of the underlying HTTP layer.\n\nGone from the latest draft.\n\n>> Detecting Smart Servers\n...\n> I actually suggest embedding the forwarding URL into an ordinary  \n> payload.  Instead of a HEAD request here, then do a GET (or, even  \n> better, POST) and get the redirected URL in return.\n>\n> Why?  Because it's common enough to redirect entire trees, and use of  \n> HTTP-layer redirections here is an unnecessary layering violation.\n\nThis has been completely rewritten to not use URL redirection at all.\n\n--8<--\nSmart HTTP transfer protocols\n=============================\n\nGit supports two HTTP based transfer protocols.  A \"dumb\" protocol\nwhich requires only a standard HTTP server on the server end of the\nconnection, and a \"smart\" protocol which requires a Git aware CGI\n(or server module).  This document describes the \"smart\" protocol.\n\nAs a design feature smart clients can automatically translate and\nupgrade \"dumb\" protocol URLs.  This permits all users to have the\nsame published URL, with the peers automatically choosing to use\nthe most efficient transport available to them.\n\nHTTP Transport\n--------------\n\nAll requests are encoded as HTTP POST requests to the smart service\nURL, \"$url/backend.git-http/$service\".\n\nAll responses are encoded as 200 Ok responses, even if the server\nside has \"failed\" the request.  Service specific success/failure\ncodes are embedded in the content.\n\nAuthentication\n--------------\n\nStandard HTTP authentication is used if authentication is required\nto access a repository, and must be configured and enforced by the\nHTTP server software itself.\n\nStateless\n---------\n\nThe protocol, much like its underlying HTTP, is stateless, from the\nperspective of the HTTP server side.  All state must be retained and\nmanaged by the client.  This permits round-robin load-balancing on\nthe server side, among many other implementation details.\n\nContent Type\n------------\n\nAll requests/responses use \"application/x-git\" as the content type.\nAction specific subtypes are specified by the parameter \"service\",\ne.g. \"application/x-git; service=upload-pack\".\n\nDetecting Smart Servers\n-----------------------\n\nHTTP clients can detect a smart Git-aware server by sending\na request to service \"show-ref\".\n\nA Git-aware server will respond with a valid response (see below).\nA dumb server should respond with an error message. \n\nService show-ref\n----------------\n\nObtains the available refs from the remote repository.\n\nURL: $url/backend.git-http/show-ref\nContent-Type: application/x-git; service=show-ref\n\nThe request is an empty body.\n\nThe response is a sequence of refs, one per Git packet line.\nThe final packet line has a length of 0 to indicate the end.\n\n\tS: 003295dcfa3633004da0049d3d0fa03f80589cbcaf31 HEAD\n\tS: 003e95dcfa3633004da0049d3d0fa03f80589cbcaf31 refs/heads/maint\n\tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n\tS: 003b2cb58b79488a98d2721cea644875a8dd0026b115 refs/heads/pu\n\tS: 0000\n\nService upload-pack\n-------------------\n\nPrepares an estimated minimal pack to transfer new objects to the\nclient.\n\nURL: $url/backend.git-http/upload-pack\nContent-Type: application/x-git; service=upload-pack\n\nThe computation to select the minimal pack proceeds as follows\n(c = client, s = server):\n\n init step:\n (c) Use show-ref to obtain the advertised refs.\n (c) Place any object seen in show-ref into set ADVERTISED.\n\n (c) Build a set, WANT, of the objects from ADVERTISED the client\n     wants to fetch, based on what it saw from show-ref.\n\n (c) Start a queue, C_PENDING, ordered by commit time (popping newest\n     first).  Add all client refs.  When a commit is popped from the\n     queue its parents should be automatically inserted back.  Commits\n     should only enter the queue once.\n\n one compute step:\n (c) Send an upload-pack request:\n\n\tC: 0009want\n\tC: 0xxx<WANT list>\n\tC: 000bcommon\n\tC: 0xxx<COMMON list>\n\tC: 0009have\n\tC: 0xxx<HAVE list>\n\tC: 0000\n\n     The stream is organized into \"sections\", where each section is\n     composed of two git pkt-lines.  The first pkt-line provides the\n     name of the section (\"want\", \"have\", \"common\").  The second\n     pkt-line has the binary SHA-1 ids which compose that section.\n\n     The \"want\" section is required.  The other sections (\"have\",\n     \"common\") are optional.  A missing \"want\" section should be\n     answered with an error.\n\n     Sections must appear in the following order, if they appear\n     at all in the request stream:\n\n       * want\n       * common\n       * have\n\n     Each section may appear multiple times.  Client implementions\n     are encouraged to use as few sections as possible, however the\n     limit of 64k per pkt-line limits the number of ids to 3,276 per\n     section entry.\n\n     The stream is terminated by a pkt-line flush (\"0000\").\n\n     The HAVE list is created by popping the first 256 commits\n     from C_PENDING.  Less can be supplied if C_PENDING empties.\n\n  (s) Parse the upload-pack request:\n\n      Verify all objects in WANT are reachable from refs.  As\n      this may require walking backwards through history to\n      the very beginning on invalid requests the server may\n      use a reasonable limit of commits (e.g. 1000) walked\n      beyond any ref tip before giving up.\n\n      If no WANT objects are received, send an error:\n\n\tS: 0019status error no want\n\n      If any WANT object is not reachable, send an error:\n\n\tS: 001estatus error invalid want\n\n     Create an empty list, S_COMMON.\n\n     If 'common' was sent:\n\n     Load all objects into S_COMMON.\n\n     If 'have' was sent:\n\n     Loop through the objects in the order supplied by the client.\n     For each object, if the server has the object reachable from\n     a ref, add it to S_COMMON.  If a commit is added to S_COMMON,\n     do not add any ancestors, even if they also appear in HAVE.\n\n  (s) Send the upload-pack response:\n\n     If the server has found a closed set of objects to pack,\n     it replies with the pack.\n\n\tS: 0010status pack\n\tS: 000c.PACK...\n\n     The returned stream is the side-band-64k protocol supported\n     by the git-upload-pack service, and the pack is embedded into\n     stream 1.  Progress messages from the server side may appear\n     in stream 2.\n\n     If the server wants more information, it replies with a\n     status continue response:\n\n\tS: 0014status continue\n\tS: 000bcommon\n\tS: 0xxx<S_COMMON list>\n\tS: 0000\n\n     The stream formatting rules are the same as the request.\n\n     The section \"common\" details the contents of S_COMMON,\n     that is all objects from HAVE that the server also has.\n\n  (c) Parse the upload-pack response:\n\n      If the status pkt-line is \"status pack:\"\n\n      Process the pack stream and update the local refs.\n\n      If the status pkt-line is \"status continue\":\n\n      Reset COMMON to the items in S_COMMON.  The new S_COMMON\n      should be a superset of the existing COMMON set.\n\n      Remove all items in S_COMMON, and all of their ancestors,\n      from PENDING.\n\n      Do another compute step.\n\n\nService receive-pack\n--------------------\n\nUploads a pack and updates refs.\n\nURL: $url/backend.git-http/receive-pack\nContent-Type: application/x-git; service=receive-pack\n\nThe start of the stream is the commands to update the refs and\nthe remainder of the stream is the pack file itself.  See\ngit-receive-pack and its network protocol in pack-protocol.txt,\nas this is essentially the same.\n\n\tC: 006395dcfa3633004da0049d3d0fa03f80589cbcaf31 d049f6c27a2244e12041955e262a404c7faba355 refs/heads/maint\n\tC: 0000\n\tC: PACK...\n\n\tS: ...<output of receive-pack>...\n\n\n-- \nShawn.\n"},{"id":"88537","messageId":"alpine.DEB.1.10.0808252052350.29665@asgard.lang.hm","threadId":"15201","inReplyTo":"20080826034544.GA32334@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-26T03:59:33Z","receivedAt":"2008-08-26T03:59:33Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Mon, 25 Aug 2008, Shawn O. Pearce wrote:\n\n> \"H. Peter Anvin\" <hpa@zytor.com> wrote:\n>> So don't implement things as GET requests unless you genuinely can deal\n>> with the request being cached.  Using POST requests throughout seems\n>> like a safer bet to me; on the other hand, since the only use of GET is\n>> obtaining a list of refs the worst thing that can happen, I presume, is\n>> additional latency for the user behind the proxy.\n>\n> This is a good point.  There is probably not any reason to cache the\n> refs content if we don't also support caching the pack files.  So in\n> this latest draft I have moved the ref listing to also be a POST.\n\non the other hand, it would be a good thing if pack files could be cached.\n\nin a peer-peer git environment the cache would not be used very much, but \nwhen you have a large number of people tracking a central repository (or \neven a pseudo-central one like the kernel) you have a lot of people \nupgrading from one point to the next point.\n\nand for cloneing (and especially thing like linux-next where you \nessentially re-clone daily) letting the pack get cached is probably a very \ngood thing.\n\nI know it would be another round-trip, but how painful would it be to \ncompute what the contents of a pack would be (what objects would be in it, \nnot calculating the deltas nessasary for a full pack file), and return \nthat to the client so that the client could do a GET for the pack itself.\n\nif that exact pack happens to be in the cache, great, if not the server \ntakes the data from the client and creates a pack file with those objects \nin it.\n\nDavid Lang\n"},{"id":"88538","messageId":"48B38340.8050504@zytor.com","threadId":"15201","inReplyTo":"20080826034544.GA32334@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-26T04:14:56Z","receivedAt":"2008-08-26T04:14:56Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> So I think the cost may actually break even with a stateful protocol\n> if we imagine that the server is actually a farm of systems and\n> simple round-robin load-balancing is being done in front of the\n> Git-aware server.\n> \n> I'd really like to keep the protocol stateless on the server side, as\n> this makes it easier to embed into certain commerical server farms.\n> \n\nIndeed.  It was a question, not a statement of any sort.  I was curious \nabout the answer.\n\nI really like the new draft, with the one consideration below.\n\n> \n> HTTP Transport\n> --------------\n> \n> All requests are encoded as HTTP POST requests to the smart service\n> URL, \"$url/backend.git-http/$service\".\n> \n> All responses are encoded as 200 Ok responses, even if the server\n> side has \"failed\" the request.  Service specific success/failure\n> codes are embedded in the content.\n> \n\nI still would like to have an indirection step at the start, in order to \nkeep a single client on a server in the case of skew.  I suggest simply \ndo it as HTTP POST $url/backend.git-http, empty body, and return a URL \nprefix to use for the remainder of the session.  That way a server who \nwants a stateful setup can return a URL which contains a session cookie; \nothers can return a URL containing a target server, and finally others \ncan simply return the requesting URL.\n\n\t-hpa\n"},{"id":"88539","messageId":"48B38377.3050901@zytor.com","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808252052350.29665@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-26T04:15:51Z","receivedAt":"2008-08-26T04:15:51Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"david@lang.hm wrote:\n> \n> on the other hand, it would be a good thing if pack files could be cached.\n> \n> in a peer-peer git environment the cache would not be used very much, \n> but when you have a large number of people tracking a central repository \n> (or even a pseudo-central one like the kernel) you have a lot of people \n> upgrading from one point to the next point.\n> \n\nWorth noting that this also applies to the raw git protocol.\n\n\t-hpa\n"},{"id":"88540","messageId":"alpine.DEB.1.10.0808252121510.30743@asgard.lang.hm","threadId":"15201","inReplyTo":"48B38377.3050901@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-26T04:25:06Z","receivedAt":"2008-08-26T04:25:06Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Mon, 25 Aug 2008, H. Peter Anvin wrote:\n\n> david@lang.hm wrote:\n>> \n>> on the other hand, it would be a good thing if pack files could be cached.\n>> \n>> in a peer-peer git environment the cache would not be used very much, but \n>> when you have a large number of people tracking a central repository (or \n>> even a pseudo-central one like the kernel) you have a lot of people \n>> upgrading from one point to the next point.\n>> \n>\n> Worth noting that this also applies to the raw git protocol.\n\nIIRC the native git server will use existing packs when it can.\n\nit would be interesting to modify git to record what packs it generates \nand then see how much a big server (like kernel.org) would re-use a pack \nunder different caching strategies.\n\nDavid Lang\n"},{"id":"88541","messageId":"48B389BD.9050606@zytor.com","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808252121510.30743@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-26T04:42:37Z","receivedAt":"2008-08-26T04:42:37Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"david@lang.hm wrote:\n> On Mon, 25 Aug 2008, H. Peter Anvin wrote:\n> \n>> david@lang.hm wrote:\n>>>\n>>> on the other hand, it would be a good thing if pack files could be \n>>> cached.\n>>>\n>>> in a peer-peer git environment the cache would not be used very much, \n>>> but when you have a large number of people tracking a central \n>>> repository (or even a pseudo-central one like the kernel) you have a \n>>> lot of people upgrading from one point to the next point.\n>>>\n>>\n>> Worth noting that this also applies to the raw git protocol.\n> \n> IIRC the native git server will use existing packs when it can.\n> \n\nYes (and the smart http server should, too).  However, neither of them \ncan currently generate new packfiles and save them for future use in a \nseparate directory from the repository tree.\n\n\t-hpa\n"},{"id":"88542","messageId":"7bfdc29a0808252145k40c41993h4e3504a6aff66e12@mail.gmail.com","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808252121510.30743@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"Imran M Yousuf","fromEmail":"imyousuf@gmail.com","sentAt":"2008-08-26T04:45:41Z","receivedAt":"2008-08-26T04:45:41Z","isPatch":false,"sender":{"key":"imyousuf@gmail.com","avatar":"https://gravatar.com/avatar/fda3c870262849d03c7b9c4d288842e128d6d80769fa7bc2d22731b7597928be?d=mp&s=160"},"body":"On Tue, Aug 26, 2008 at 10:25 AM,  <david@lang.hm> wrote:\n> On Mon, 25 Aug 2008, H. Peter Anvin wrote:\n>\n>> david@lang.hm wrote:\n>>>\n>>> on the other hand, it would be a good thing if pack files could be\n>>> cached.\n>>>\n>>> in a peer-peer git environment the cache would not be used very much, but\n>>> when you have a large number of people tracking a central repository (or\n>>> even a pseudo-central one like the kernel) you have a lot of people\n>>> upgrading from one point to the next point.\n>>>\n>>\n>> Worth noting that this also applies to the raw git protocol.\n>\n> IIRC the native git server will use existing packs when it can.\n>\n> it would be interesting to modify git to record what packs it generates and\n> then see how much a big server (like kernel.org) would re-use a pack under\n> different caching strategies.\n\nI fully agree with the caching logic as well. In this regard I was\nthinking whether the protocol could be modified a bit to accommodate\nit or not. From initial proposal GET was dropped because there will be\ncaching, which I also agree :), and we need GET in order to achieve\ncache - so I would have done something such as - initial request would\nbe POST and if there is no change and cache can be used I would\nredirect it to a equivalen GET URL and if cache is invalid (which the\nserver can track by pinging the GET URL) serve directly through the\nPOST method untill either the GET is out of the cache or is updated.\n\n- Imran\n\n>\n> David Lang\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n\n-- \nImran M Yousuf\nEmail: imran@smartitengineering.com\nBlog: http://imyousuf-tech.blogs.smartitengineering.com/\nMobile: +880-1711402557\n"},{"id":"88582","messageId":"20080826145857.GF26523@spearce.org","threadId":"15201","inReplyTo":"48B36BCA.8060103@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-26T14:58:57Z","receivedAt":"2008-08-26T14:58:57Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n>> Detecting Smart Servers\n>> -----------------------\n>>\n>> HTTP clients can detect a smart Git-aware server by HEADing\n>> $repo/backend.git-http and looking for a 302 redirect to the\n>> repository's smart service URL:\n...\n>> All subsequent communcation for this transaction is done through\n>> the smart service URL ($ssurl), not the original URL.\n>\n> I actually suggest embedding the forwarding URL into an ordinary  \n> payload.  Instead of a HEAD request here, then do a GET (or, even  \n> better, POST) and get the redirected URL in return.\n>\n> Why?  Because it's common enough to redirect entire trees, and use of  \n> HTTP-layer redirections here is an unnecessary layering violation.\n\nHmm.  I'm actually thinking the exact opposite here.  My rationale\nfor putting the response as a standard HTTP 302/303 style redirect\nis to permit hardware load balancers or Apache mod_rewrite rules\nto implement simple load balancing with a HTTP redirect.\n\nIf we embed the redirect URL into the payload then configuring that\nwill become a lot more complex.  At the minimum you may have to\nmake up a dummy file for each server (holding the response payload)\nthen then let mod_rewrite rewrite the request internally to make\nApache serve that file.  Ugly.\n\n> If you insist on using a HTTP status code, I would claim that 303 is a  \n> better status code.\n\nOk.\n\n-- \nShawn.\n"},{"id":"88586","messageId":"20080826161425.GG26523@spearce.org","threadId":"15201","inReplyTo":"20080826145857.GF26523@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-26T16:14:25Z","receivedAt":"2008-08-26T16:14:25Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> wrote:\n> \"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> >\n> > I actually suggest embedding the forwarding URL into an ordinary  \n> > payload.  Instead of a HEAD request here, then do a GET (or, even  \n> > better, POST) and get the redirected URL in return.\n> \n> Hmm.  I'm actually thinking the exact opposite here.\n\nHere's the delta from the last draft I emailed.  Its basically just\nabout this redirect stuff.\n\ndiff --git a/Documentation/technical/http-protocol.txt b/Documentation/technical/http-protocol.txt\nindex 99d7623..a3f7379 100644\n--- a/Documentation/technical/http-protocol.txt\n+++ b/Documentation/technical/http-protocol.txt\n@@ -43,14 +43,40 @@ All requests/responses use \"application/x-git\" as the content type.\n Action specific subtypes are specified by the parameter \"service\",\n e.g. \"application/x-git; service=upload-pack\".\n \n+Redirects\n+---------\n+\n+If a POST request results in an HTTP 302 or 303 redirect response\n+clients should retry the request by updating the URL and POSTing\n+the request to the new location.\n+\n+If the new request is successful clients should trim off the\n+trailing \"/backend.git/$service\" portion of the new loaction\n+and use the remainder as the base URL for future requests in\n+the same transaction.\n+\n+This redirection permits Apache's mod_rewrite (and many other\n+servers) to implement a form of round-robin load balancing by\n+redirecting all requests to a generic host to a specific host.\n+\n Detecting Smart Servers\n -----------------------\n \n HTTP clients can detect a smart Git-aware server by sending\n a request to service \"show-ref\".\n \n-A Git-aware server will respond with a valid response (see below).\n-A dumb server should respond with an error message. \n+A Git-aware server will respond with a valid response.  Clients\n+must check the following properties to prevent being fooled by\n+misconfigured servers:\n+\n+  * HTTP status code is 200.\n+  * Content-Type is \"application/x-git; service=show-ref\"\n+  * The body can be parsed without errors.  The length of\n+    each pkt-line must be 4 valid hex digits.\n+\n+A dumb server will respond with a non-200 HTTP status code.\n+A misconfigured server may respond with a normal 200 status\n+code, but an incorrect content type.\n \n Service show-ref\n ----------------\n\n-- \nShawn.\n"},{"id":"88590","messageId":"48B4303C.3080409@zytor.com","threadId":"15201","inReplyTo":"20080826145857.GF26523@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-26T16:33:00Z","receivedAt":"2008-08-26T16:33:00Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> Hmm.  I'm actually thinking the exact opposite here.  My rationale\n> for putting the response as a standard HTTP 302/303 style redirect\n> is to permit hardware load balancers or Apache mod_rewrite rules\n> to implement simple load balancing with a HTTP redirect.\n> \n> If we embed the redirect URL into the payload then configuring that\n> will become a lot more complex.  At the minimum you may have to\n> make up a dummy file for each server (holding the response payload)\n> then then let mod_rewrite rewrite the request internally to make\n> Apache serve that file.  Ugly.\n> \n\nNo, you're thinking backwards.  What you want is the standard HTTP \nredirect load balancing to take effect *before* the initial request is \nserviced.  The front-end load balancer will take effect on the initial \nrequest, and then redirect the request to a node (via a 302 reply.)  The \ntarget node then sends a self-referencing URL to keep the service local, \nif that is desired -- otherwise it doesn't.\n\nAgain, the 300-class redirect is treated as a part of the HTTP transport \nin this case; it doesn't have to be visible to the RPC layer.  However, \nin order to maintain the integrity of an interchange, we do need an \nadditional level of redirection visible to the RPC layer.\n\n> If we embed the redirect URL into the payload then configuring that\n> will become a lot more complex.  At the minimum you may have to\n> make up a dummy file for each server (holding the response payload)\n> then then let mod_rewrite rewrite the request internally to make\n> Apache serve that file.  Ugly.\n\nA very simple CGI/PHP script will do this, and it's really very very \ntrivial to set up.\n\nPlease keep in mind I'm not talking hypotheticals at all.  What you have \nproposed is actually a lot uglier for kernel.org to implement, simply \nbecause we try to stay with strict IP-based vhosting\n\n\t-hpa\n"},{"id":"88594","messageId":"alpine.LFD.1.10.0808261255570.1624@xanadu.home","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808252052350.29665@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-08-26T17:01:20Z","receivedAt":"2008-08-26T17:01:20Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 25 Aug 2008, david@lang.hm wrote:\n\n> and for cloneing (and especially thing like linux-next where you essentially\n> re-clone daily) letting the pack get cached is probably a very good thing.\n\nI hope that people recloning linux-next daily are very few.  This is an \nincredible waste of bandwidth, regardless of the protocol used, dumb or \nnot.  A standard fetch with a remote tracking branch (with -f or with a \nplus sign on the \"fetch\" line in your config file) should be all that's \nneeded to significantly reduce the amount of data needed to transfer.\n\n\nNicolas\n"},{"id":"88595","messageId":"20080826170315.GH26523@spearce.org","threadId":"15201","inReplyTo":"alpine.LFD.1.10.0808261255570.1624@xanadu.home","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-26T17:03:15Z","receivedAt":"2008-08-26T17:03:15Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Nicolas Pitre <nico@cam.org> wrote:\n> On Mon, 25 Aug 2008, david@lang.hm wrote:\n> \n> > and for cloneing (and especially thing like linux-next where you essentially\n> > re-clone daily) letting the pack get cached is probably a very good thing.\n> \n> I hope that people recloning linux-next daily are very few.  This is an \n> incredible waste of bandwidth, regardless of the protocol used, dumb or \n> not.  A standard fetch with a remote tracking branch (with -f or with a \n> plus sign on the \"fetch\" line in your config file) should be all that's \n> needed to significantly reduce the amount of data needed to transfer.\n\nOr at least clone with --reference.  You get about the same benefit if\nyour local reference repository is fairly current, say with a stable\nupstream like Linus' own tree.\n\n-- \nShawn.\n"},{"id":"88607","messageId":"20080826172648.GK26523@spearce.org","threadId":"15201","inReplyTo":"48B4303C.3080409@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-26T17:26:48Z","receivedAt":"2008-08-26T17:26:48Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>>\n>> Hmm.  I'm actually thinking the exact opposite here.  My rationale\n>> for putting the response as a standard HTTP 302/303 style redirect\n>> is to permit hardware load balancers [...]\n>> to implement simple load balancing with a HTTP redirect.\n>\n> No, you're thinking backwards.  What you want is the standard HTTP  \n> redirect load balancing to take effect *before* the initial request is  \n> serviced.\n...\n> Please keep in mind I'm not talking hypotheticals at all.  What you have  \n> proposed is actually a lot uglier for kernel.org to implement, simply  \n> because we try to stay with strict IP-based vhosting\n\nDiscard my prior patch from today.\n\nThis is a patch to last night's full document edition\n(http://article.gmane.org/gmane.comp.version-control.git/93704)\nand addresses only the issue of redirects.\n\n--8<--\ndiff --git a/Documentation/technical/http-protocol.txt b/Documentation/technical/http-protocol.txt\nindex 99d7623..99dc88d 100644\n--- a/Documentation/technical/http-protocol.txt\n+++ b/Documentation/technical/http-protocol.txt\n@@ -43,14 +43,34 @@ All requests/responses use \"application/x-git\" as the content type.\n Action specific subtypes are specified by the parameter \"service\",\n e.g. \"application/x-git; service=upload-pack\".\n \n+HTTP Redirects\n+--------------\n+\n+If a POST request results in an HTTP 302 or 303 redirect response\n+clients should retry the request by updating the URL and POSTing\n+the same request to the new location.  Subsequent requests should\n+still be sent to the original URL.\n+\n Detecting Smart Servers\n -----------------------\n \n HTTP clients can detect a smart Git-aware server by sending\n a request to service \"show-ref\".\n \n-A Git-aware server will respond with a valid response (see below).\n-A dumb server should respond with an error message. \n+A Git-aware server will respond with a valid response.  Clients\n+must check the following properties to prevent being fooled by\n+misconfigured servers:\n+\n+  * HTTP status code is 200.\n+  * Content-Type is \"application/x-git; service=show-ref\"\n+  * The body can be parsed without errors.  The length of\n+    each pkt-line must be 4 valid hex digits.\n+\n+A dumb server will respond with a non-200 HTTP status code.\n+A misconfigured server may respond with a normal 200 status\n+code, but an incorrect content type, or an invalid leading\n+4 byte sequence for a pkt-line (e.g. \"<htm\" or \"<!DO\" are\n+not valid lengths).\n \n Service show-ref\n ----------------\n@@ -62,15 +82,46 @@ Content-Type: application/x-git; service=show-ref\n \n The request is an empty body.\n \n-The response is a sequence of refs, one per Git packet line.\n-The final packet line has a length of 0 to indicate the end.\n+The response is a pkt-line with \"refs\", followed by zero\n+or more ref pkt-lines (\"$id $name\"), and a final pkt-line\n+with a length of 0:\n \n+\tS: 0009refs\n \tS: 003295dcfa3633004da0049d3d0fa03f80589cbcaf31 HEAD\n \tS: 003e95dcfa3633004da0049d3d0fa03f80589cbcaf31 refs/heads/maint\n \tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n \tS: 003b2cb58b79488a98d2721cea644875a8dd0026b115 refs/heads/pu\n \tS: 0000\n \n+The response may begin with an optional redirect to a new service\n+URL for the repository:\n+\n+\tS: 0028redirect http://s1.example.com/git/\n+\tS: 0009refs\n+\tS: 003295dcfa3633004da0049d3d0fa03f80589cbcaf31 HEAD\n+\tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n+\tS: 0000\n+\n+or be composed of only a redirect:\n+\n+\tS: 0028redirect http://s1.example.com/git/\n+\tS: 0000\n+\n+If a redirect is returned the client should update itself\n+to use the new URL as the location for future requests.\n+A server may use the redirect to request that the client\n+\"pin\" itself to a particular server for the remainder of\n+the current transaction.\n+\n+The URL listed in any redirect should be the base URL\n+without any query args.  The client will automatically\n+append \"/backend.git-http/$service\" as it makes each\n+future request.\n+\n+If no \"refs\" line was received in the response, but\n+a \"redirect\" was received, the client should retry\n+its request at the new location before giving up.\n+\n Service upload-pack\n -------------------\n \n-- \nShawn.\n"},{"id":"88660","messageId":"48B485F8.5030109@zytor.com","threadId":"15201","inReplyTo":"20080826172648.GK26523@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-26T22:38:48Z","receivedAt":"2008-08-26T22:38:48Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> Discard my prior patch from today.\n> \n> This is a patch to last night's full document edition\n> (http://article.gmane.org/gmane.comp.version-control.git/93704)\n> and addresses only the issue of redirects.\n> \n\nLooks great to me.\n\n\t-hpa\n"},{"id":"88678","messageId":"7bfdc29a0808261951t59af433bkefa77ef8c370ecbd@mail.gmail.com","threadId":"15201","inReplyTo":"48B485F8.5030109@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Imran M Yousuf","fromEmail":"imyousuf@gmail.com","sentAt":"2008-08-27T02:51:20Z","receivedAt":"2008-08-27T02:51:20Z","isPatch":false,"sender":{"key":"imyousuf@gmail.com","avatar":"https://gravatar.com/avatar/fda3c870262849d03c7b9c4d288842e128d6d80769fa7bc2d22731b7597928be?d=mp&s=160"},"body":"On Wed, Aug 27, 2008 at 4:38 AM, H. Peter Anvin <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>>\n>> Discard my prior patch from today.\n>>\n>> This is a patch to last night's full document edition\n>> (http://article.gmane.org/gmane.comp.version-control.git/93704)\n>> and addresses only the issue of redirects.\n>>\n>\n> Looks great to me.\n\nThis looks really good! The redirect idea just seems cool!\n\n- Imran\n\n>\n>        -hpa\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n-- \nImran M Yousuf\nEmail: imran@smartitengineering.com\nBlog: http://imyousuf-tech.blogs.smartitengineering.com/\nMobile: +880-1711402557\n"},{"id":"88864","messageId":"20080828035018.GA10010@spearce.org","threadId":"15201","inReplyTo":"48B485F8.5030109@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-28T03:50:18Z","receivedAt":"2008-08-28T03:50:18Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n>\n> Looks great to me.\n\nSo this is what may be the final draft of the HTTP protocol.\nI've added stuff about capability selection between the peers for\nfuture expansion support.  The upload-pack service has a better\nuse of it than receive-pack.  Otherwise it is what I think you are\nagreeing to above.  ;-)\n\nI'm hoping to start implementating a prototype of this on Friday.\nI may do it in JGit first; the transport infrastructure there is\na lot more modular so experimentation should be quicker.  I would\nobviously also implement it in C Git, unless someone else comes\nalong and beats me to it.  This project is only a fraction of my\ntotal Git time in any given week.  :-|\n\n--8<--\nSmart HTTP transfer protocols\n=============================\n\nGit supports two HTTP based transfer protocols.  A \"dumb\" protocol\nwhich requires only a standard HTTP server on the server end of the\nconnection, and a \"smart\" protocol which requires a Git aware CGI\n(or server module).  This document describes the \"smart\" protocol.\n\nAs a design feature smart clients can automatically translate and\nupgrade \"dumb\" protocol URLs.  This permits all users to have the\nsame published URL, with the peers automatically choosing to use\nthe most efficient transport available to them.\n\nHTTP Transport\n--------------\n\nAll requests are encoded as HTTP POST requests to the smart service\nURL, \"$url/backend.git-http/$service\".\n\nAll responses are encoded as 200 Ok responses, even if the server\nside has \"failed\" the request.  Service specific success/failure\ncodes are embedded in the content.\n\nAuthentication\n--------------\n\nStandard HTTP authentication is used if authentication is required\nto access a repository, and must be configured and enforced by the\nHTTP server software itself.\n\nStateless\n---------\n\nThe protocol, much like its underlying HTTP, is stateless, from the\nperspective of the HTTP server side.  All state must be retained and\nmanaged by the client.  This permits round-robin load-balancing on\nthe server side, among many other implementation details.\n\nContent Type\n------------\n\nAll requests/responses use \"application/x-git\" as the content type.\nAction specific subtypes are specified by the parameter \"service\",\ne.g. \"application/x-git; service=upload-pack\".\n\nHTTP Redirects\n--------------\n\nIf a POST request results in an HTTP 302 or 303 redirect response\nclients should retry the request by updating the URL and POSTing\nthe same request to the new location.  Subsequent requests should\nstill be sent to the original URL.\n\nDetecting Smart Servers\n-----------------------\n\nHTTP clients can detect a smart Git-aware server by sending\na request to service \"show-ref\".\n\nA Git-aware server will respond with a valid response.  Clients\nmust check the following properties to prevent being fooled by\nmisconfigured servers:\n\n  * HTTP status code is 200.\n  * Content-Type is \"application/x-git; service=show-ref\"\n  * The body can be parsed without errors.  The length of\n    each pkt-line must be 4 valid hex digits.\n\nA dumb server will respond with a non-200 HTTP status code.\nA misconfigured server may respond with a normal 200 status\ncode, but an incorrect content type, or an invalid leading\n4 byte sequence for a pkt-line (e.g. \"<htm\" or \"<!DO\" are\nnot valid lengths).\n\nService show-ref\n----------------\n\nObtains the available refs from the remote repository.\n\nURL: $url/backend.git-http/show-ref\nContent-Type: application/x-git; service=show-ref\n\nThe request is an empty body.\n\nThe response is a pkt-line with \"refs\", followed by zero\nor more ref pkt-lines (\"$id $name\"), and a final pkt-line\nwith a length of 0:\n\n\tS: 0009refs\n\tS: 003295dcfa3633004da0049d3d0fa03f80589cbcaf31 HEAD\n\tS: 003e95dcfa3633004da0049d3d0fa03f80589cbcaf31 refs/heads/maint\n\tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n\tS: 003b2cb58b79488a98d2721cea644875a8dd0026b115 refs/heads/pu\n\tS: 0000\n\nThe response may begin with an optional redirect to a new service\nURL for the repository:\n\n\tS: 0028redirect http://s1.example.com/git/\n\tS: 0009refs\n\tS: 003295dcfa3633004da0049d3d0fa03f80589cbcaf31 HEAD\n\tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n\tS: 0000\n\nor be composed of only a redirect:\n\n\tS: 0028redirect http://s1.example.com/git/\n\tS: 0000\n\nIf a redirect is returned the client should update itself\nto use the new URL as the location for future requests.\nA server may use the redirect to request that the client\n\"pin\" itself to a particular server for the remainder of\nthe current transaction.\n\nThe URL listed in any redirect should be the base URL\nwithout any query args.  The client will automatically\nappend \"/backend.git-http/$service\" as it makes each\nfuture request.\n\nIf no \"refs\" line was received in the response, but\na \"redirect\" was received, the client should retry\nits request at the new location before giving up.\n\nService upload-pack\n-------------------\n\nPrepares an estimated minimal pack to transfer new objects to the\nclient.\n\nURL: $url/backend.git-http/upload-pack\nContent-Type: application/x-git; service=upload-pack\n\nThe computation to select the minimal pack proceeds as follows\n(c = client, s = server):\n\n init step:\n (c) Use show-ref to obtain the advertised refs.\n (c) Place any object seen in show-ref into set ADVERTISED.\n\n (c) Build a set, WANT, of the objects from ADVERTISED the client\n     wants to fetch, based on what it saw from show-ref.\n\n (c) Start a queue, C_PENDING, ordered by commit time (popping newest\n     first).  Add all client refs.  When a commit is popped from the\n     queue its parents should be automatically inserted back.  Commits\n     should only enter the queue once.\n\n one compute step:\n (c) Send an upload-pack request:\n\n\tC: 0011capabilities\n\tC: 0024thin-pack include-tag ofs-delta\n\tC: 0009want\n\tC: 0xxx<WANT list>\n\tC: 000bcommon\n\tC: 0xxx<COMMON list>\n\tC: 0009have\n\tC: 0xxx<HAVE list>\n\tC: 0000\n\n     The stream is organized into \"sections\", where each section is\n     composed of two git pkt-lines.  The first pkt-line provides the\n     name of the section (\"capabilities\", \"want\", \"have\", \"common\").\n     The second pkt-line has the binary SHA-1 ids which compose that\n     section.\n\n     The \"want\" section is required.  The other sections (\"have\",\n     \"common\") are optional.  A missing \"want\" section should be\n     answered with an error.\n\n     Sections must appear in the following order, if they appear\n     at all in the request stream:\n\n       * capabilities\n       * want\n       * common\n       * have\n\n     Each section may appear multiple times.  Client implementions\n     are encouraged to use as few sections as possible, however the\n     limit of 64k per pkt-line limits the number of ids to 3,276 per\n     section entry.\n\n     The stream is terminated by a pkt-line flush (\"0000\").\n\n     The HAVE list is created by popping the first 256 commits\n     from C_PENDING.  Less can be supplied if C_PENDING empties.\n\n  (s) Parse the upload-pack request:\n\n      Verify all objects in WANT are reachable from refs.  As\n      this may require walking backwards through history to\n      the very beginning on invalid requests the server may\n      use a reasonable limit of commits (e.g. 1000) walked\n      beyond any ref tip before giving up.\n\n      If no WANT objects are received, send an error:\n\n\tS: 0019status error no want\n\n      If any WANT object is not reachable, send an error:\n\n\tS: 001estatus error invalid want\n\n     Create an empty list, S_COMMON.\n\n     If 'common' was sent:\n\n     Load all objects into S_COMMON.\n\n     If 'have' was sent:\n\n     Loop through the objects in the order supplied by the client.\n     For each object, if the server has the object reachable from\n     a ref, add it to S_COMMON.  If a commit is added to S_COMMON,\n     do not add any ancestors, even if they also appear in HAVE.\n\n  (s) Send the upload-pack response:\n\n     If the server has found a closed set of objects to pack, it\n     replies with the pack and the enabled capabilities.  The set\n     of enabled capabilities is limited to the intersection of\n     what the client requested and what the server supports.\n\n\tS: 0010status pack\n\tS: 0011capabilities\n\tS: 0024thin-pack include-tag ofs-delta\n\tS: 000c.PACK...\n\n     The returned stream is the side-band-64k protocol supported\n     by the git-upload-pack service, and the pack is embedded into\n     stream 1.  Progress messages from the server side may appear\n     in stream 2.\n\n     If the server wants more information, it replies with a\n     status continue response:\n\n\tS: 0014status continue\n\tS: 000bcommon\n\tS: 0xxx<S_COMMON list>\n\tS: 0000\n\n     The stream formatting rules are the same as the request.\n\n     The section \"common\" details the contents of S_COMMON,\n     that is all objects from HAVE that the server also has.\n\n  (c) Parse the upload-pack response:\n\n      If the status pkt-line is \"status pack:\"\n\n      Process the pack stream and update the local refs.\n\n      If the status pkt-line is \"status continue\":\n\n      Reset COMMON to the items in S_COMMON.  The new S_COMMON\n      should be a superset of the existing COMMON set.\n\n      Remove all items in S_COMMON, and all of their ancestors,\n      from PENDING.\n\n      Do another compute step.\n\n\nService receive-pack\n--------------------\n\nUploads a pack and updates refs.\n\nURL: $url/backend.git-http/receive-pack\nContent-Type: application/x-git; service=receive-pack\n\nThe start of the stream is the commands to update the refs and\nthe remainder of the stream is the pack file itself.  See\ngit-receive-pack and its network protocol in pack-protocol.txt,\nas this is essentially the same.\n\n\tC: 0011capabilities\n\tC: 0005\n\tC: 006395dcfa3633004da0049d3d0fa03f80589cbcaf31 d049f6c27a2244e12041955e262a404c7faba355 refs/heads/maint\n\tC: 0000\n\tC: PACK...\n\n\tS: 0011capabilities\n\tS: 0005\n\tS: ...<output of receive-pack>...\n\nThe capabilities are handled exactly as in the fetch protocol,\nhowever the server may reject a pack and its associated commands\nif an invalid capability request is made by the client, or the\nclient has assumed a pack capability that the server does not\nhave support for.  In the latter case the server must still send\nthe capabilities key in the response so the client can correct\nitself and try again.\n\n-- \nShawn.\n"},{"id":"88866","messageId":"48B62B6F.7010103@zytor.com","threadId":"15201","inReplyTo":"20080828035018.GA10010@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-28T04:37:03Z","receivedAt":"2008-08-28T04:37:03Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> So this is what may be the final draft of the HTTP protocol.\n> I've added stuff about capability selection between the peers for\n> future expansion support.  The upload-pack service has a better\n> use of it than receive-pack.  Otherwise it is what I think you are\n> agreeing to above.  ;-)\n> \n\nIt looks good to me.  I *really* like the option of combining a redirect \nwith a refs list in one reply; this will make things substantially \neasier do deploy on kernel.org, and saves a round trip to boot.\n\nJust an implementation detail for the server, however: for an *empty* \nrepository (one which has no refs at all), the server needs to *not* \ntransmit the redirect, or there will be a loop :)  It is unnecessary, \nanyway, since there is inherently nothing to do.\n\n\t-hpa\n"},{"id":"88867","messageId":"20080828044205.GB10238@spearce.org","threadId":"15201","inReplyTo":"48B62B6F.7010103@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-28T04:42:05Z","receivedAt":"2008-08-28T04:42:05Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>>\n>> So this is what may be the final draft of the HTTP protocol.\n>\n> It looks good to me.  I *really* like the option of combining a redirect  \n> with a refs list in one reply; this will make things substantially  \n> easier do deploy on kernel.org, and saves a round trip to boot.\n\nYea, I had a draft that didn't combine these and I realized how\nstupid that was.  So I allowed them to appear together if the\nserver operator wants to do that.\n\n> Just an implementation detail for the server, however: for an *empty*  \n> repository (one which has no refs at all), the server needs to *not*  \n> transmit the redirect, or there will be a loop :)  It is unnecessary,  \n> anyway, since there is inherently nothing to do.\n\nActually that's not true.  A correct client won't loop.\n\nAn empty repository is required to send \"refs\" section header.\nSo the client will see the \"refs\" header and know that the complete\nset of refs is following.  Only nothing follows, so it knows the\ncomplete set is the empty set.\n\nA redirect with no ref data won't have the \"refs\" section header.\nSo the client knows that it cannot conclude anything from that\nexchange and must follow the redirect.\n\nAn empty repository sending a redirect will send both \"redirect\"\nand \"refs\", but no refs follow the \"refs\" section header.  So the\nclient knows that it is empty and it does not need to follow the\nredirect it received.\n\nNow if the server is stupid and keeps sending a redirect with no\nrefs header, yea, the client can loop.  So the clients should have\na maximum recursion limit configured into them, just like a good\nbrowser would, so you can't get stuck in an A->B->C->A loop.\n\n-- \nShawn.\n"},{"id":"88868","messageId":"7vhc95iwcs.fsf@gitster.siamese.dyndns.org","threadId":"15201","inReplyTo":"20080828035018.GA10010@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-08-28T04:42:27Z","receivedAt":"2008-08-28T04:42:27Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> HTTP Redirects\n> --------------\n>\n> If a POST request results in an HTTP 302 or 303 redirect response\n> clients should retry the request by updating the URL and POSTing\n> the same request to the new location.  Subsequent requests should\n> still be sent to the original URL.\n\nAt the first reading I was confused because this seemed to contradict with\nthe server pinning that is done by the payload level redirect.\n\n> Service upload-pack\n> -------------------\n>\n> Prepares an estimated minimal pack to transfer new objects to the\n> client.\n>\n> URL: $url/backend.git-http/upload-pack\n> Content-Type: application/x-git; service=upload-pack\n>\n> The computation to select the minimal pack proceeds as follows\n> (c = client, s = server):\n>\n>  init step:\n>  (c) Use show-ref to obtain the advertised refs.\n>  (c) Place any object seen in show-ref into set ADVERTISED.\n>\n>  (c) Build a set, WANT, of the objects from ADVERTISED the client\n>      wants to fetch, based on what it saw from show-ref.\n>\n>  (c) Start a queue, C_PENDING, ordered by commit time (popping newest\n>      first).  Add all client refs.  When a commit is popped from the\n>      queue its parents should be automatically inserted back.  Commits\n>      should only enter the queue once.\n>\n>  one compute step:\n>  (c) Send an upload-pack request:\n>\n> \tC: 0011capabilities\n> \tC: 0024thin-pack include-tag ofs-delta\n> \tC: 0009want\n> \tC: 0xxx<WANT list>\n> \tC: 000bcommon\n> \tC: 0xxx<COMMON list>\n> \tC: 0009have\n> \tC: 0xxx<HAVE list>\n> \tC: 0000\n>\n>      The stream is organized into \"sections\", where each section is\n>      composed of two git pkt-lines.  The first pkt-line provides the\n>      name of the section (\"capabilities\", \"want\", \"have\", \"common\").\n>      The second pkt-line has the binary SHA-1 ids which compose that\n>      section.\n\nIt appears that you really meant \"Binary\", as opposed to \"Hexadecimal\"\nthat show-ref example illustrate, judging from the later 3,276 number.\nI'd prefer hexadecimal here.\n\nAs a protocol specification, you'd eventually need to describe the\npkt-line format, namely, (1) four hexadecimal digits that represents the\nlength of the line (including that four bytes), followed by that many\nnumber of bytes as the line's payload, or (2) \"0000\" which is \"flush\".\nAlso typically the text based line payload is LF terminated (hence the\nfour-hexdigit length counts the terminating LF).  Also \"capabilities\" need\nto be defined.\n\n>   (s) Parse the upload-pack request:\n>\n>       Verify all objects in WANT are reachable from refs.  As\n>       this may require walking backwards through history to\n>       the very beginning on invalid requests the server may\n>       use a reasonable limit of commits (e.g. 1000) walked\n>       beyond any ref tip before giving up.\n\nI suspect moving as much work to the client side by erroring out and\nhaving the client restart from show-ref might be a better tradeoff (also\nthis has been advertised as a security feature on the native protocol\nside).\n\n>       If any WANT object is not reachable, send an error:\n>\n> \tS: 001estatus error invalid want\n>\n>      Create an empty list, S_COMMON.\n>\n>      If 'common' was sent:\n>\n>      Load all objects into S_COMMON.\n\nSecurity?  Error out if some of them do not exist on the server end, at\nleast.\n\n>   (s) Send the upload-pack response:\n>\n>      If the server has found a closed set of objects to pack, it\n>      replies with the pack and the enabled capabilities.  The set\n>      of enabled capabilities is limited to the intersection of\n>      what the client requested and what the server supports.\n\nDefine \"closed set\".\n\n>      The stream formatting rules are the same as the request.\n>\n>      The section \"common\" details the contents of S_COMMON,\n>      that is all objects from HAVE that the server also has.\n\nAn object in HAVE that exists on the server end can be a descendant of\nmany other HAVEs. Answering with that youngest one alone is enough,\nwithout the other HAVEs the server end also has as its ancestors, as they\nare redundant information.\n\n>   (c) Parse the upload-pack response:\n>\n>       If the status pkt-line is \"status pack:\"\n>\n>       Process the pack stream and update the local refs.\n>\n>       If the status pkt-line is \"status continue\":\n>\n>       Reset COMMON to the items in S_COMMON.  The new S_COMMON\n>       should be a superset of the existing COMMON set.\n\nIs there a way to detect bad clients that does not obey this rule without\nserver side states?\n"},{"id":"88871","messageId":"48B6305D.1080004@zytor.com","threadId":"15201","inReplyTo":"20080828044205.GB10238@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-28T04:58:05Z","receivedAt":"2008-08-28T04:58:05Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n>> Just an implementation detail for the server, however: for an *empty*  \n>> repository (one which has no refs at all), the server needs to *not*  \n>> transmit the redirect, or there will be a loop :)  It is unnecessary,  \n>> anyway, since there is inherently nothing to do.\n> \n> Actually that's not true.  A correct client won't loop.\n> \n> An empty repository is required to send \"refs\" section header.\n> So the client will see the \"refs\" header and know that the complete\n> set of refs is following.  Only nothing follows, so it knows the\n> complete set is the empty set.\n> \n> A redirect with no ref data won't have the \"refs\" section header.\n> So the client knows that it cannot conclude anything from that\n> exchange and must follow the redirect.\n> \n\nAh, good point.\n\n\t-hpa\n"},{"id":"88976","messageId":"7bfdc29a0808272340kdc2f3b0x250eef32b25dcdcb@mail.gmail.com","threadId":"15201","inReplyTo":"48B62B6F.7010103@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Imran M Yousuf","fromEmail":"imyousuf@gmail.com","sentAt":"2008-08-28T06:40:07Z","receivedAt":"2008-08-28T06:40:07Z","isPatch":false,"sender":{"key":"imyousuf@gmail.com","avatar":"https://gravatar.com/avatar/fda3c870262849d03c7b9c4d288842e128d6d80769fa7bc2d22731b7597928be?d=mp&s=160"},"body":"On Thu, Aug 28, 2008 at 10:37 AM, H. Peter Anvin <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>>\n>> So this is what may be the final draft of the HTTP protocol.\n>> I've added stuff about capability selection between the peers for\n>> future expansion support.  The upload-pack service has a better\n>> use of it than receive-pack.  Otherwise it is what I think you are\n>> agreeing to above.  ;-)\n>>\n>\n> It looks good to me.  I *really* like the option of combining a redirect\n> with a refs list in one reply; this will make things substantially easier do\n> deploy on kernel.org, and saves a round trip to boot.\n\nI agree, this is a very cool feature of the protocol...\n\n- Imran\n\n>\n> Just an implementation detail for the server, however: for an *empty*\n> repository (one which has no refs at all), the server needs to *not*\n> transmit the redirect, or there will be a loop :)  It is unnecessary,\n> anyway, since there is inherently nothing to do.\n>\n>        -hpa\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n\n-- \nImran M Yousuf\nEntrepreneur & Software Engineer\nSmart IT Engineering\nDhaka, Bangladesh\nEmail: imran@smartitengineering.com\nBlog: http://imyousuf-tech.blogs.smartitengineering.com/\nMobile: +880-1711402557\n"},{"id":"88897","messageId":"20080828145706.GB21072@spearce.org","threadId":"15201","inReplyTo":"7vhc95iwcs.fsf@gitster.siamese.dyndns.org","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-28T14:57:06Z","receivedAt":"2008-08-28T14:57:06Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n> \n> > HTTP Redirects\n> > --------------\n> >\n> > If a POST request results in an HTTP 302 or 303 redirect response\n> > clients should retry the request by updating the URL and POSTing\n> > the same request to the new location.  Subsequent requests should\n> > still be sent to the original URL.\n> \n> At the first reading I was confused because this seemed to contradict with\n> the server pinning that is done by the payload level redirect.\n\nThis is meant to help load balancing initially target to a server.\nI think its also reasonable to honor a transport level redirect,\nmuch as we honor whatever route IP gives us (not that we have a\nlot of choice - or even want one at that level).\n \n> > Service upload-pack\n> > -------------------\n> >  one compute step:\n> >  (c) Send an upload-pack request:\n> >\n> > \tC: 0011capabilities\n> > \tC: 0024thin-pack include-tag ofs-delta\n> > \tC: 0009want\n> > \tC: 0xxx<WANT list>\n> > \tC: 000bcommon\n> > \tC: 0xxx<COMMON list>\n> > \tC: 0009have\n> > \tC: 0xxx<HAVE list>\n> > \tC: 0000\n> >\n> >      The stream is organized into \"sections\", where each section is\n> >      composed of two git pkt-lines.  The first pkt-line provides the\n> >      name of the section (\"capabilities\", \"want\", \"have\", \"common\").\n> >      The second pkt-line has the binary SHA-1 ids which compose that\n> >      section.\n> \n> It appears that you really meant \"Binary\", as opposed to \"Hexadecimal\"\n> that show-ref example illustrate, judging from the later 3,276 number.\n> I'd prefer hexadecimal here.\n\nYes, I really did mean for this part of the protocol to be in binary.\nWe have to exchange a bunch of commits to figure out what is common.\nThe binary form is 1/2 the size of the hexadecimal form, resulting\nin fewer TCP packets for the same request.\n\nReading/writing the SHA-1s in binary is usually faster than doing\nit in hex; you don't have to go through the formatting routines.\nSo there's a few less CPU cycles on the server end.\n\nBut the rest of the protocol is in hex and ASCII, so I guess it\ndoes make sense to make this part be in hex too. I can change it\nin the next draft.\n \n> As a protocol specification, you'd eventually need to describe the\n> pkt-line format, namely, (1) four hexadecimal digits that represents the\n> length of the line (including that four bytes), followed by that many\n> number of bytes as the line's payload, or (2) \"0000\" which is \"flush\".\n> Also typically the text based line payload is LF terminated (hence the\n> four-hexdigit length counts the terminating LF).\n\nYes.  I'll add that into the next draft.\n\n> Also \"capabilities\" need\n> to be defined.\n\nWell, currently its just room for expansion.  But I'll try to define\nit out better.  My initial thought is to do something like we have\nin the native protocol where there are capability names \"hidden\"\non the end of the first pkt-line.  Only I'm making it explicit.\n \n> >   (s) Parse the upload-pack request:\n> >\n> >       Verify all objects in WANT are reachable from refs.  As\n> >       this may require walking backwards through history to\n> >       the very beginning on invalid requests the server may\n> >       use a reasonable limit of commits (e.g. 1000) walked\n> >       beyond any ref tip before giving up.\n> \n> I suspect moving as much work to the client side by erroring out and\n> having the client restart from show-ref might be a better tradeoff (also\n> this has been advertised as a security feature on the native protocol\n> side).\n\nI'm concerned about livelock.  If the client sees something in\nshow-ref, starts upload-pack and gets 2 round-trips into the common\nexchange and then someone updates a ref the client wants, the client\nhas to go back to the beginning and start all over.\n\nBut if the object they want is still reachable (within a reasonable\ndistance) from the current refs, what is the harm in letting the\nclient see the stale view?  Especially since grabbing the most\ncurrent refs would still make that object available to the client?\n\nRemember that is how the native protocol behaves.  You get a single\nupload-pack process which has grabbed a snapshot of the refs.\nIf they change during the want-have-ack-nack exchange the client\ndoesn't get kicked out and asked to start all over again.  Same idea.\n \n> >       If any WANT object is not reachable, send an error:\n> >\n> > \tS: 001estatus error invalid want\n> >\n> >      Create an empty list, S_COMMON.\n> >\n> >      If 'common' was sent:\n> >\n> >      Load all objects into S_COMMON.\n> \n> Security?  Error out if some of them do not exist on the server end, at\n> least.\n\nI think I can add something saying its a protocol error if that\nhappens.  Its not a security risk, remember the S_COMMON set\neventually turns into the\n\n  git rev-list --objects-edge $WANT --not $S_COMMON \\\n  | git pack-objects --stdout\n\nSending an S_COMMON the server doesn't have just causes it to fail.\n\nIf the server silently prunes out ones it doesn't know its not a\nconcern.  If the server has it, but it isn't advertised in a ref,\nits also not a security risk.  No data from those objects is sent\nback to the client.\n \n> >   (s) Send the upload-pack response:\n> >\n> >      If the server has found a closed set of objects to pack, it\n> >      replies with the pack and the enabled capabilities.  The set\n> >      of enabled capabilities is limited to the intersection of\n> >      what the client requested and what the server supports.\n> \n> Define \"closed set\".\n\nYea, not only that I don't describe how the client can give up and\njust ask for everything that is left.  Like say on an initial clone.\n \n> >      The stream formatting rules are the same as the request.\n> >\n> >      The section \"common\" details the contents of S_COMMON,\n> >      that is all objects from HAVE that the server also has.\n> \n> An object in HAVE that exists on the server end can be a descendant of\n> many other HAVEs. Answering with that youngest one alone is enough,\n> without the other HAVEs the server end also has as its ancestors, as they\n> are redundant information.\n\nYes, obviously.  I must not have made that clear here.  I'll try\nto improve the language.\n \n> >   (c) Parse the upload-pack response:\n> >\n> >       If the status pkt-line is \"status pack:\"\n> >\n> >       Process the pack stream and update the local refs.\n> >\n> >       If the status pkt-line is \"status continue\":\n> >\n> >       Reset COMMON to the items in S_COMMON.  The new S_COMMON\n> >       should be a superset of the existing COMMON set.\n> \n> Is there a way to detect bad clients that does not obey this rule without\n> server side states?\n\nNo.  Is that really a concern though?\n\nThe worst a bad client can do here is cause itself to receive\nmore data than it wants by refusing to put things into COMMON.\nEventually it gives up and just clones the entire repository.  How is\nthat any different from a well behaved client doing an initial clone?\n\nA bad client could also stick random things into COMMON.  If the\nserver doesn't have the object we error out (as you suggest above)\nduring the next call.  So the client has only DOS'd the server.\nIt can DOS the server easier other ways.\n\nA bad client could stick only part of what S_COMMON into COMMON.\nThat may cause it to get a bigger pack file than it asked for as the\nrev-list call won't be as limited.  How is that any different from\na well behaved client that is really behind and has a lot to fetch?\n\n-- \nShawn.\n"},{"id":"88917","messageId":"48B6DABD.7090800@zytor.com","threadId":"15201","inReplyTo":"7vhc95iwcs.fsf@gitster.siamese.dyndns.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-28T17:05:01Z","receivedAt":"2008-08-28T17:05:01Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Junio C Hamano wrote:\n\n> \n> It appears that you really meant \"Binary\", as opposed to \"Hexadecimal\"\n> that show-ref example illustrate, judging from the later 3,276 number.\n> I'd prefer hexadecimal here.\n> \n\nI *think* the \"native\" git protocol uses binary here.  It makes sense to \nbe consistent, to allow them to share code?\n\n\t-hpa\n"},{"id":"88918","messageId":"20080828171052.GC21072@spearce.org","threadId":"15201","inReplyTo":"48B6DABD.7090800@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-28T17:10:52Z","receivedAt":"2008-08-28T17:10:52Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Junio C Hamano wrote:\n>\n>>\n>> It appears that you really meant \"Binary\", as opposed to \"Hexadecimal\"\n>> that show-ref example illustrate, judging from the later 3,276 number.\n>> I'd prefer hexadecimal here.\n>>\n>\n> I *think* the \"native\" git protocol uses binary here.  It makes sense to  \n> be consistent, to allow them to share code?\n\nNo, the native protocol is horribly verbose here:\n\n\t0032want ac3abe10ed54d512fbbaeb7cef19972eedd8e4a8\n\t0032want 404c3bbec34f5c65c5024c856eed4dbbfc27831e\n\t0032want 9bcc7aff6095549c1425aef6ca0034c47189705d\n\t0032have 471287a3c311e486206d3c6ff94faf3dfffc736c\n\t0032have 48f27055a4fa5f4da8234f44808f0b0c70629218\n\t0032have d4cc612f218b3dd3b831e3b976bf85165cd4f3d4\n\t...\n\nso its doing it in hex, and its using 10 bytes of \"framing\" for\nevery SHA-1 it sends as each is sent in its own pkt-line with the\nhave/want header.\n\n-- \nShawn.\n"},{"id":"88921","messageId":"48B6DE7A.1020207@zytor.com","threadId":"15201","inReplyTo":"20080828171052.GC21072@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-28T17:20:58Z","receivedAt":"2008-08-28T17:20:58Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \"H. Peter Anvin\" <hpa@zytor.com> wrote:\n>> Junio C Hamano wrote:\n>>\n>>> It appears that you really meant \"Binary\", as opposed to \"Hexadecimal\"\n>>> that show-ref example illustrate, judging from the later 3,276 number.\n>>> I'd prefer hexadecimal here.\n>>>\n>> I *think* the \"native\" git protocol uses binary here.  It makes sense to  \n>> be consistent, to allow them to share code?\n> \n> No, the native protocol is horribly verbose here:\n> \n> \t0032want ac3abe10ed54d512fbbaeb7cef19972eedd8e4a8\n> \t0032want 404c3bbec34f5c65c5024c856eed4dbbfc27831e\n> \t0032want 9bcc7aff6095549c1425aef6ca0034c47189705d\n> \t0032have 471287a3c311e486206d3c6ff94faf3dfffc736c\n> \t0032have 48f27055a4fa5f4da8234f44808f0b0c70629218\n> \t0032have d4cc612f218b3dd3b831e3b976bf85165cd4f3d4\n> \t...\n> \n> so its doing it in hex, and its using 10 bytes of \"framing\" for\n> every SHA-1 it sends as each is sent in its own pkt-line with the\n> have/want header.\n> \n\nHm.  It's probably not enough data to worry significantly about.\n\n\t-hpa\n"},{"id":"88927","messageId":"alpine.DEB.1.10.0808281023200.2713@asgard.lang.hm","threadId":"15201","inReplyTo":"20080828145706.GB21072@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-28T17:26:15Z","receivedAt":"2008-08-28T17:26:15Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 28 Aug 2008, Shawn O. Pearce wrote:\n\n> Junio C Hamano <gitster@pobox.com> wrote:\n>> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n>>\n>> It appears that you really meant \"Binary\", as opposed to \"Hexadecimal\"\n>> that show-ref example illustrate, judging from the later 3,276 number.\n>> I'd prefer hexadecimal here.\n>\n> Yes, I really did mean for this part of the protocol to be in binary.\n> We have to exchange a bunch of commits to figure out what is common.\n> The binary form is 1/2 the size of the hexadecimal form, resulting\n> in fewer TCP packets for the same request.\n\nexcept that HTTP cannot transport binary data, if you feed it binary data \nit then encodes it into 7-bit safe forms for transport.\n\nwhile it's true that it can pack it more efficiantly than hex, it's not \ndouble the density.\n\n> Reading/writing the SHA-1s in binary is usually faster than doing\n> it in hex; you don't have to go through the formatting routines.\n> So there's a few less CPU cycles on the server end.\n\nexcept you then need to go through the formatting routines to send it via \nHTTP.\n\nDavid Lang\n"},{"id":"88928","messageId":"20080828172623.GD21072@spearce.org","threadId":"15201","inReplyTo":"48B6DE7A.1020207@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-28T17:26:23Z","receivedAt":"2008-08-28T17:26:23Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>> \"H. Peter Anvin\" <hpa@zytor.com> wrote:\n>>>\n>>> I *think* the \"native\" git protocol uses binary here.  It makes sense \n>>> to  be consistent, to allow them to share code?\n>>\n>> No, the native protocol is horribly verbose here:\n>>\n>> \t0032want ac3abe10ed54d512fbbaeb7cef19972eedd8e4a8\n>> \t...\n>>\n>> so its doing it in hex, and its using 10 bytes of \"framing\" for\n>> every SHA-1 it sends as each is sent in its own pkt-line with the\n>> have/want header.\n>\n> Hm.  It's probably not enough data to worry significantly about.\n\nShould I change the HTTP protocol then to use the same format,\nso they have a better chance at sharing code between them?\n\n-- \nShawn.\n"},{"id":"88929","messageId":"20080828172853.GE21072@spearce.org","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808281023200.2713@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-28T17:28:53Z","receivedAt":"2008-08-28T17:28:53Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"david@lang.hm wrote:\n> On Thu, 28 Aug 2008, Shawn O. Pearce wrote:\n>>> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n>>\n>> Yes, I really did mean for this part of the protocol to be in binary.\n>\n> except that HTTP cannot transport binary data, if you feed it binary data \n> it then encodes it into 7-bit safe forms for transport.\n\nSo then how does it transport a GIF file to my browser?  uuencoded?\nLast time I read the RFCs I was pretty certain HTTP is 8-bit clean\nin both directions.\n\nOf course this may all be moot.  I think we're moving in a direction\nof matching the git native protocol more exactly.\n\n-- \nShawn.\n"},{"id":"88930","messageId":"alpine.DEB.1.10.0808281033070.2713@asgard.lang.hm","threadId":"15201","inReplyTo":"20080828172853.GE21072@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-28T17:37:04Z","receivedAt":"2008-08-28T17:37:04Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 28 Aug 2008, Shawn O. Pearce wrote:\n\n> david@lang.hm wrote:\n>> On Thu, 28 Aug 2008, Shawn O. Pearce wrote:\n>>>> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n>>>\n>>> Yes, I really did mean for this part of the protocol to be in binary.\n>>\n>> except that HTTP cannot transport binary data, if you feed it binary data\n>> it then encodes it into 7-bit safe forms for transport.\n>\n> So then how does it transport a GIF file to my browser?  uuencoded?\n\nsomething like that. it uses the mimetype mechanisms to identify the \nvarious pieces and encodes each piece (if nothing else it needs to make \nsure that the mimetype seperators don't appear in the data) uuencode is \none of the available mechanisms.\n\n> Last time I read the RFCs I was pretty certain HTTP is 8-bit clean\n> in both directions.\n\nI could be wrong, but I'm pretty sure I'm not. to test this yourself find \na webserver with an image file and retrieve it via telnet (telnet hostname \n80<enter>GET /path/to/file HTTP/1.0<enter><enter>) and what will come back \nwill be text.\n\n> Of course this may all be moot.  I think we're moving in a direction\n> of matching the git native protocol more exactly.\n\ntrue, but it's never a waste of time to learn something (whichever one of \nus is right :-)\n\nDavid Lang\n"},{"id":"88932","messageId":"alpine.LRH.1.10.0808281937580.10660@yvahk3.pbagnpgbe.fr","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808281033070.2713@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"Daniel Stenberg","fromEmail":"daniel@haxx.se","sentAt":"2008-08-28T17:38:59Z","receivedAt":"2008-08-28T17:38:59Z","isPatch":false,"sender":{"key":"daniel@haxx.se","avatar":"https://gravatar.com/avatar/69fdca87edd17cee21ca2e79fc2ff671d644603c3dc27167430f3cd3dbab7ba8?d=mp&s=160"},"body":"On Thu, 28 Aug 2008, david@lang.hm wrote:\n\n>>> except that HTTP cannot transport binary data, if you feed it binary data\n>>> it then encodes it into 7-bit safe forms for transport.\n>> \n>> So then how does it transport a GIF file to my browser?  uuencoded?\n>\n> something like that. it uses the mimetype mechanisms to identify the various \n> pieces and encodes each piece (if nothing else it needs to make sure that \n> the mimetype seperators don't appear in the data) uuencode is one of the \n> available mechanisms.\n\nNo. HTTP is 8bit clean and sends and receives binary just fine. You seem to \nthink of SMTP or something.\n\n-- \n\n  / daniel.haxx.se\n"},{"id":"88933","messageId":"48B6E3C0.4010608@zytor.com","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808281023200.2713@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-28T17:43:28Z","receivedAt":"2008-08-28T17:43:28Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"david@lang.hm wrote:\n> except that HTTP cannot transport binary data, if you feed it binary \n> data it then encodes it into 7-bit safe forms for transport.\n\nTotal utter bunk.  You're thinking of email and news, which had to deal \nwith broken legacy code.\n\nHTTP has *always* been binary clean.  It does not encode anything into \n7-bit safe anything.  The only \"encoding\" that it ever does is \nHTTP/1.1's chunked encoding, which is a way to deal with the fact that \nit might not always know the total length of the data before it starts \nthe transfer; it sends the data in arbitrary-sized \"chunks\" prefixed by \na byte count.  It does this to support connection caching in HTTP/1.1; \nHTTP/1.0 would simply close the connection to indicate end of (binary) data.\n\n\t-hpa\n"},{"id":"88931","messageId":"20080828174334.GF21072@spearce.org","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808281033070.2713@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-28T17:43:34Z","receivedAt":"2008-08-28T17:43:34Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"david@lang.hm wrote:\n> On Thu, 28 Aug 2008, Shawn O. Pearce wrote:\n>> david@lang.hm wrote:\n>>> On Thu, 28 Aug 2008, Shawn O. Pearce wrote:\n>>>>> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n>>>>\n>>>> Yes, I really did mean for this part of the protocol to be in binary.\n>>>\n>>> except that HTTP cannot transport binary data, if you feed it binary data\n>>> it then encodes it into 7-bit safe forms for transport.\n>>\n>> So then how does it transport a GIF file to my browser?  uuencoded?\n...\n> I could be wrong, but I'm pretty sure I'm not. to test this yourself find \n> a webserver with an image file and retrieve it via telnet (telnet \n> hostname 80<enter>GET /path/to/file HTTP/1.0<enter><enter>) and what will \n> come back will be text.\n\n  $ telnet www.google.com 80\n  Trying 74.125.19.104...\n  Connected to www.google.com (74.125.19.104).\n  Escape character is '^]'.\n  GET /intl/en_ALL/images/logo.gif HTTP/1.0\n  \n  HTTP/1.0 200 OK\n  Content-Type: image/gif\n  Last-Modified: Wed, 07 Jun 2006 19:38:24 GMT\n  Expires: Sun, 17 Jan 2038 19:14:07 GMT\n  Cache-Control: public\n  Date: Thu, 28 Aug 2008 17:40:44 GMT\n  Server: gws\n  Content-Length: 8558\n  X-Google-Backends: /bns/pq/borg/pq/bns/gws-prod/staticweb.staticfrontend.gws/16:9836,dauf30:80\n  X-Google-Service: static\n  X-Google-GFE-Request-Trace: dauf30:80,/bns/pq/borg/pq/bns/gws-prod/staticweb.staticfrontend.gws/16:9836,dauf30:80\n  Connection: Close\n  \n  GIF89a\t\u0001n���������������έ\t���\u0018E�\u0018I�\u00104�\u0010<��\u0018�������\u0010ƾ����\u0018M������������ｾ�����$c!Y�����\u0018QΜ�������e֮c��1e�J}<�������s�����9q�k\n\nVery funny.  It trashed my tty.  Even reset won't restore the\nsettings.  Anyway.\n\nI chose the Google logo on the Google homepage because I know we\ntry really hard to conform to standards, so we can have the biggest\npossible user base.  Micro$oft or Yahoo! probably would have come\nout the same way.  Or some image on kernel.org.\n\nAnyway, I didn't send any browser data, so the server had to assume\nthe dumbest f'ing browser on the planet, and I got back binary data.\n\n-- \nShawn.\n"},{"id":"88934","messageId":"48B6E3F3.1070903@zytor.com","threadId":"15201","inReplyTo":"20080828172623.GD21072@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-28T17:44:19Z","receivedAt":"2008-08-28T17:44:19Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> Should I change the HTTP protocol then to use the same format,\n> so they have a better chance at sharing code between them?\n> \n\nI leave that up to you and Junio.  My feel would that it's not worth \noptimizing the HTTP protocol separately.\n\n\t-hpa\n"},{"id":"88936","messageId":"20080828174642.GG21072@spearce.org","threadId":"15201","inReplyTo":"48B6E3F3.1070903@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-28T17:46:42Z","receivedAt":"2008-08-28T17:46:42Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>>\n>> Should I change the HTTP protocol then to use the same format,\n>> so they have a better chance at sharing code between them?\n>>\n>\n> I leave that up to you and Junio.  My feel would that it's not worth  \n> optimizing the HTTP protocol separately.\n\nYea, I'm leaning towards just keeping them the same.  I may be\nable to reuse a lot of code in JGit that way.  In C Git its going\nto take some refactoring to disentagle the IO parts of fetch-pack\nfrom the protocol, but I should be able to reuse a lot there too.\n\n-- \nShawn.\n"},{"id":"88937","messageId":"48B6E4A5.7030605@zytor.com","threadId":"15201","inReplyTo":"20080828174334.GF21072@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-28T17:47:17Z","receivedAt":"2008-08-28T17:47:17Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n>   GET /intl/en_ALL/images/logo.gif HTTP/1.0\n\n> Anyway, I didn't send any browser data, so the server had to assume\n> the dumbest f'ing browser on the planet, and I got back binary data.\n\nSlight nitpick: you did (HTTP/1.0).  If you really want to show the \nbottom-of-the barrel behaviour, drop that off which gets you HTTP 0.x \nbehaviour -- not exactly commonly encountered today :)  HTTP 0.x had no \nprovisions for a request header, so you should not need to send a blank \nline after the GET.\n\n\t-hpa\n"},{"id":"88938","messageId":"20080828180452.GB14901@glandium.org","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808281033070.2713@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-08-28T18:04:52Z","receivedAt":"2008-08-28T18:04:52Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Thu, Aug 28, 2008 at 10:37:04AM -0700, david@lang.hm wrote:\n> On Thu, 28 Aug 2008, Shawn O. Pearce wrote:\n>\n>> david@lang.hm wrote:\n>>> On Thu, 28 Aug 2008, Shawn O. Pearce wrote:\n>>>>> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n>>>>\n>>>> Yes, I really did mean for this part of the protocol to be in binary.\n>>>\n>>> except that HTTP cannot transport binary data, if you feed it binary data\n>>> it then encodes it into 7-bit safe forms for transport.\n>>\n>> So then how does it transport a GIF file to my browser?  uuencoded?\n>\n> something like that. it uses the mimetype mechanisms to identify the  \n> various pieces and encodes each piece (if nothing else it needs to make  \n> sure that the mimetype seperators don't appear in the data) uuencode is  \n> one of the available mechanisms.\n>\n>> Last time I read the RFCs I was pretty certain HTTP is 8-bit clean\n>> in both directions.\n>\n> I could be wrong, but I'm pretty sure I'm not. to test this yourself find \n> a webserver with an image file and retrieve it via telnet (telnet \n> hostname 80<enter>GET /path/to/file HTTP/1.0<enter><enter>) and what will \n> come back will be text.\n\nNo it won't. Try it *yourself*.\n\n$ nc www.google.com 80 | sed '1,/^\\r$/d' > /tmp/logo.gif\nGET /intl/en_ALL/images/logo.gif HTTP/1.1\nHost: www.google.com\nConnection: close\n\n$ file /tmp/logo.gif \n/tmp/logo.gif: GIF image data, version 89a, 276 x 110\n\nMike\n\nPS: sed only removes the HTTP response headers.\n"},{"id":"88939","messageId":"alpine.DEB.1.10.0808281111310.2713@asgard.lang.hm","threadId":"15201","inReplyTo":"48B6E3C0.4010608@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-28T18:12:13Z","receivedAt":"2008-08-28T18:12:13Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 28 Aug 2008, H. Peter Anvin wrote:\n\n> david@lang.hm wrote:\n>> except that HTTP cannot transport binary data, if you feed it binary data \n>> it then encodes it into 7-bit safe forms for transport.\n>\n> Total utter bunk.  You're thinking of email and news, which had to deal with \n> broken legacy code.\n>\n> HTTP has *always* been binary clean.  It does not encode anything into 7-bit \n> safe anything.  The only \"encoding\" that it ever does is HTTP/1.1's chunked \n> encoding, which is a way to deal with the fact that it might not always know \n> the total length of the data before it starts the transfer; it sends the data \n> in arbitrary-sized \"chunks\" prefixed by a byte count.  It does this to \n> support connection caching in HTTP/1.1; HTTP/1.0 would simply close the \n> connection to indicate end of (binary) data.\n\nOk, I was wrong, thanks to everyone for correcting me. I now know this.\n\nDavid Lang\n"},{"id":"88940","messageId":"48B6EAF5.4030704@zytor.com","threadId":"15201","inReplyTo":"alpine.DEB.1.10.0808281111310.2713@asgard.lang.hm","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-28T18:14:13Z","receivedAt":"2008-08-28T18:14:13Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"david@lang.hm wrote:\n> \n> Ok, I was wrong, thanks to everyone for correcting me. I now know this.\n> \n\nIt's an easy mistake to make given HTTP's apparent connections with \nMIME.  However, all it uses from MIME is the type tags, it doesn't use \nMIME encapsulation format at all.\n\n\t-hpa\n"},{"id":"88941","messageId":"alpine.DEB.1.10.0808281117410.2713@asgard.lang.hm","threadId":"15201","inReplyTo":"48B6EAF5.4030704@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-08-28T18:18:13Z","receivedAt":"2008-08-28T18:18:13Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 28 Aug 2008, H. Peter Anvin wrote:\n\n> david@lang.hm wrote:\n>> \n>> Ok, I was wrong, thanks to everyone for correcting me. I now know this.\n>> \n>\n> It's an easy mistake to make given HTTP's apparent connections with MIME. \n> However, all it uses from MIME is the type tags, it doesn't use MIME \n> encapsulation format at all.\n\nas I said earlier, it's never a waste of time to learn things, no matter \nwho is right. ;-)\n\nDavid Lang\n"},{"id":"88942","messageId":"alpine.LFD.1.10.0808281432240.1624@xanadu.home","threadId":"15201","inReplyTo":"20080828172623.GD21072@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-08-28T18:40:21Z","receivedAt":"2008-08-28T18:40:21Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 28 Aug 2008, Shawn O. Pearce wrote:\n\n> \"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> > Shawn O. Pearce wrote:\n> >> \"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> >>>\n> >>> I *think* the \"native\" git protocol uses binary here.  It makes sense \n> >>> to  be consistent, to allow them to share code?\n> >>\n> >> No, the native protocol is horribly verbose here:\n> >>\n> >> \t0032want ac3abe10ed54d512fbbaeb7cef19972eedd8e4a8\n> >> \t...\n> >>\n> >> so its doing it in hex, and its using 10 bytes of \"framing\" for\n> >> every SHA-1 it sends as each is sent in its own pkt-line with the\n> >> have/want header.\n> >\n> > Hm.  It's probably not enough data to worry significantly about.\n> \n> Should I change the HTTP protocol then to use the same format,\n> so they have a better chance at sharing code between them?\n\nGiven that the ref exchange happens on multiple lines (one ref per line) \nin the native protocol, and that your proposal is using one line for \nmultiple refs, I don't see this as a big factor wrt code reuse.  Since \nyou'll have separate \"output\" code anyway, why not simply going with \nrefs in straight binary for the HTTP protocol?  Even the debugability of \nrefs exchange in plain text is dubious especially with all refs on the \nsame line (that'll be a pain to split refs out of a long stream of hex \nby hand).\n\n\nNicolas\n"},{"id":"88944","messageId":"48B6F2C7.1010801@zytor.com","threadId":"15201","inReplyTo":"alpine.LFD.1.10.0808281432240.1624@xanadu.home","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-28T18:47:35Z","receivedAt":"2008-08-28T18:47:35Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Nicolas Pitre wrote:\n>> Should I change the HTTP protocol then to use the same format,\n>> so they have a better chance at sharing code between them?\n> \n> Given that the ref exchange happens on multiple lines (one ref per line) \n> in the native protocol, and that your proposal is using one line for \n> multiple refs, I don't see this as a big factor wrt code reuse.  Since \n> you'll have separate \"output\" code anyway, why not simply going with \n> refs in straight binary for the HTTP protocol?  Even the debugability of \n> refs exchange in plain text is dubious especially with all refs on the \n> same line (that'll be a pain to split refs out of a long stream of hex \n> by hand).\n\nWell, I think the real question was to go to multiple lines, \nnative-protocol style, or go to binary.\n\n\t-hpa\n"},{"id":"89079","messageId":"7vwsi0a2op.fsf@gitster.siamese.dyndns.org","threadId":"15201","inReplyTo":"20080828145706.GB21072@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-08-29T04:02:46Z","receivedAt":"2008-08-29T04:02:46Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> writes:\n\n> Junio C Hamano <gitster@pobox.com> wrote:\n>> \"Shawn O. Pearce\" <spearce@spearce.org> writes:\n>> \n>> > HTTP Redirects\n>> > --------------\n>> >\n>> > If a POST request results in an HTTP 302 or 303 redirect response\n>> > clients should retry the request by updating the URL and POSTing\n>> > the same request to the new location.  Subsequent requests should\n>> > still be sent to the original URL.\n>> \n>> At the first reading I was confused because this seemed to contradict with\n>> the server pinning that is done by the payload level redirect.\n>\n> This is meant to help load balancing initially target to a server.\n> I think its also reasonable to honor a transport level redirect,\n> much as we honor whatever route IP gives us (not that we have a\n> lot of choice - or even want one at that level).\n\nYeah, I think understood what you are trying to achieve here after reading\nthe document twice.\n\nI was just pointing out that the language (or presentation order) was\nconfusing to me and I needed to read these two sections twice to see the\ndifference between the two redirects.\n\n>> Is there a way to detect bad clients that does not obey this rule without\n>> server side states?\n>\n> No.  Is that really a concern though?\n\nI was more concerned about a bad/broken client not giving up forever, and\nnot giving the server enough cue to give up, saying \"I've conversed with\nthis guy long enough but haven't reached the conclusion yet --- there is\nsomething wrong\".  Even without server side states, if we were to trust\nclients, we can add \"this is Nth round\" to the protocol to help the server\ndetect \"long enough\" part, but that somehow does not feel right.\n"},{"id":"89084","messageId":"48B784FD.3080005@zytor.com","threadId":"15201","inReplyTo":"7vwsi0a2op.fsf@gitster.siamese.dyndns.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-08-29T05:11:25Z","receivedAt":"2008-08-29T05:11:25Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Junio C Hamano wrote:\n> \n>>> Is there a way to detect bad clients that does not obey this rule without\n>>> server side states?\n>> No.  Is that really a concern though?\n> \n> I was more concerned about a bad/broken client not giving up forever, and\n> not giving the server enough cue to give up, saying \"I've conversed with\n> this guy long enough but haven't reached the conclusion yet --- there is\n> something wrong\".  Even without server side states, if we were to trust\n> clients, we can add \"this is Nth round\" to the protocol to help the server\n> detect \"long enough\" part, but that somehow does not feel right.\n> \n\nWe should be able to detect either inconsistency, or lack of forward \nprogress, but as long as there is forward progress made there doesn't \nseem to be a strong need to terminate.\n\n\t-hpa\n"},{"id":"89090","messageId":"7vej488gcu.fsf@gitster.siamese.dyndns.org","threadId":"15201","inReplyTo":"48B784FD.3080005@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-08-29T06:50:25Z","receivedAt":"2008-08-29T06:50:25Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> We should be able to detect either inconsistency, or lack of forward\n> progress, but as long as there is forward progress made there doesn't\n> seem to be a strong need to terminate.\n\nYeah, what I wanted to say was that it would be tricky to detect (any/lack\nof) forward progress without having server side state and without trusting\nthe client.\n"},{"id":"89145","messageId":"20080829173954.GG7403@spearce.org","threadId":"15201","inReplyTo":"7vej488gcu.fsf@gitster.siamese.dyndns.org","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-08-29T17:39:54Z","receivedAt":"2008-08-29T17:39:54Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Yet another draft follows.  I believe that I have covered all\ncomments with this draft.  But I welcome any additional ones,\nas thus far it has been a very constructive process.\n\nThe updated protocol looks more like the current native protocol\ndoes.  This should make it easier to reuse code between the two\nprotocol implementations.\n\n--8<--\nSmart HTTP transfer protocols\n=============================\n\nGit supports two HTTP based transfer protocols.  A \"dumb\" protocol\nwhich requires only a standard HTTP server on the server end of the\nconnection, and a \"smart\" protocol which requires a Git aware CGI\n(or server module).  This document describes the \"smart\" protocol.\n\nAs a design feature smart clients can automatically translate and\nupgrade \"dumb\" protocol URLs.  This permits all users to have the\nsame published URL, with the peers automatically choosing to use\nthe most efficient transport available to them.\n\nHTTP Transport\n--------------\n\nAll requests are encoded as HTTP POST requests to the smart service\nURL, \"$url/backend.git-http/$service\".\n\nAll responses are encoded as 200 Ok responses, even if the server\nside has \"failed\" the request.  Service specific success/failure\ncodes are embedded in the content.\n\nAuthentication\n--------------\n\nStandard HTTP authentication is used if authentication is required\nto access a repository, and must be configured and enforced by the\nHTTP server software itself.\n\nStateless\n---------\n\nThe protocol, much like its underlying HTTP, is stateless, from the\nperspective of the HTTP server side.  All state must be retained and\nmanaged by the client.  This permits round-robin load-balancing on\nthe server side, among many other implementation details.\n\nContent Type\n------------\n\nAll requests/responses use \"application/x-git\" as the content type.\nAction specific subtypes are specified by the parameter \"service\",\ne.g. \"application/x-git; service=upload-pack\".\n\nHTTP Redirects\n--------------\n\nIf a POST request results in an HTTP 302 or 303 redirect response\nclients should retry the request by updating the URL and POSTing\nthe same request to the new location.  Subsequent requests should\nstill be sent to the original URL.\n\nThis redirect behavior is unrelated to the in-payload redirect\nthat is described below in \"Service show-ref\".\n\nDetecting Smart Servers\n-----------------------\n\nHTTP clients can detect a smart Git-aware server by sending\na request to service \"show-ref\".\n\nA Git-aware server will respond with a valid response.  Clients\nmust check the following properties to prevent being fooled by\nmisconfigured servers:\n\n  * HTTP status code is 200.\n  * Content-Type is \"application/x-git; service=show-ref\"\n  * The body can be parsed without errors.  The length of\n    each pkt-line must be 4 valid hex digits.\n\nA dumb server will respond with a non-200 HTTP status code.\nA misconfigured server may respond with a normal 200 status\ncode, but an incorrect content type, or an invalid leading\n4 byte sequence for a pkt-line (e.g. \"<htm\" or \"<!DO\" are\nnot valid lengths).\n\npkt-line Format\n---------------\n\nMuch of the payload is described around pkt-lines.\n\nA pkt-line is a variable length binary string.  The first four bytes\nof the line indicates the total length of the line, in hexadecimal.\nThe total length includes the 4 bytes used to denote the length.  A\nline is usually terminated by an LF, which must be included in the\ntotal length if present.\n\nBinary data is permitted within a pkt-line so implementors should\nensure their pkt-line parsing/formatting routines are 8-bit clean.\nThe maximum length of a pkt-line's data is 65532 bytes (65536 - 4).\n\nExamples (as C-style strings):\n\n  pkt-line          actual value\n  ---------------------------------\n  \"0006a\\n\"         \"a\\n\"\n  \"0005a\"           \"a\"\n  \"000bfoobar\\n\"    \"foobar\\n\"\n  \"0004\"            \"\"\n\nA pkt-line with a length of 0 (\"0000\") is a special case and is\ntreated as a break or terminator in the payload.\n\nService show-ref\n----------------\n\nObtains the available refs from the remote repository.\n\nURL: $url/backend.git-http/show-ref\nContent-Type: application/x-git; service=show-ref\n\nThe request is an empty body.\n\nThe response is a pkt-line with \"refs\", followed by zero\nor more ref pkt-lines (\"$id $name\"), and a final pkt-line\nwith a length of 0:\n\n\tS: 0009refs\n\tS: 003295dcfa3633004da0049d3d0fa03f80589cbcaf31 HEAD\n\tS: 003e95dcfa3633004da0049d3d0fa03f80589cbcaf31 refs/heads/maint\n\tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n\tS: 003b2cb58b79488a98d2721cea644875a8dd0026b115 refs/heads/pu\n\tS: 0000\n\nThe response may begin with an optional redirect to a new service\nURL for the repository:\n\n\tS: 0028redirect http://s1.example.com/git/\n\tS: 0009refs\n\tS: 003295dcfa3633004da0049d3d0fa03f80589cbcaf31 HEAD\n\tS: 003fd049f6c27a2244e12041955e262a404c7faba355 refs/heads/master\n\tS: 0000\n\nor be composed of only a redirect:\n\n\tS: 0028redirect http://s1.example.com/git/\n\tS: 0000\n\nIf a redirect is returned the client should update itself\nto use the new URL as the location for future requests.\nA server may use the redirect to request that the client\n\"pin\" itself to a particular server for the remainder of\nthe current transaction.\n\nThe URL listed in any redirect should be the base URL\nwithout any query args.  The client will automatically\nappend \"/backend.git-http/$service\" as it makes each\nfuture request.\n\nIf no \"refs\" line was received in the response, but\na \"redirect\" was received, the client should retry\nits request at the new location before giving up.\n\nService upload-pack\n-------------------\n\nPrepares an estimated minimal pack to transfer new objects to the\nclient.\n\nURL: $url/backend.git-http/upload-pack\nContent-Type: application/x-git; service=upload-pack\n\nThe computation to select the minimal pack proceeds as follows\n(c = client, s = server):\n\n init step:\n (c) Use show-ref to obtain the advertised refs.\n (c) Place any object seen in show-ref into set ADVERTISED.\n\n (c) Build a set, WANT, of the objects from ADVERTISED the client\n     wants to fetch, based on what it saw from show-ref.\n (c) Build an empty set, COMMON, to hold the objects that are later\n     determined to be on both ends.\n\n (c) Start a queue, C_PENDING, ordered by commit time (popping newest\n     first).  Add all client refs.  When a commit is popped from the\n     queue its parents should be automatically inserted back.  Commits\n     should only enter the queue once.\n\n one compute step:\n (c) Send an upload-pack request:\n\n\tC: 001bcapability include-tag\n\tC: 0019capability thin-pack\n\t....\n\tC: 0032want <WANT #1>...............................\n\tC: 0032want <WANT #2>...............................\n\t....\n\tC: 0034common <COMMON #1>.............................\n\tC: 0034common <COMMON #2>.............................\n\t....\n\tC: 0032have <HAVE #1>...............................\n\tC: 0032have <HAVE #2>...............................\n\t....\n\tC: 0000\n\n     The stream is organized into \"commands\", with each command\n     appearing by itself in a pkt-line.  Within a command line\n     the text leading up to the first space is the command name,\n     and the remainder of the line to the first LF is the value.\n     Command lines are terminated with an LF as the last byte of\n     the pkt-line value.\n\n     Servers must ignore commands which they do not recognize.\n     This permits newer clients to transmit additional data to\n     an unknown server, in case the server is new enough to use\n     the additional information.\n\n     Commands must appear in the following order, if they appear\n     at all in the request stream:\n\n       * capability\n       * want\n       * common\n       * have\n       * give-up\n\n     The stream is terminated by a pkt-line flush (\"0000\").\n\n     The \"capability\" command requests a single protocol feature\n     to be enabled by the server.  Typically these are used to\n     describe aspects of the pack that will be returned.  See\n     below for more details on the current capabilities.\n\n     A single \"want\", \"common\", or \"have\" command has one hex\n     formatted SHA-1 as its value.  Multiple SHA-1s can be sent\n     by sending multiple commands.\n\n     The HAVE list is created by popping the first 64 commits\n     from C_PENDING.  Less can be supplied if C_PENDING empties.\n\n     If the client has sent 256 HAVE commits and has not yet\n     received one of those back from S_COMMON, or the client\n     has emptied C_PENDING it should include a \"give-up\"\n     command to let the server know it won't proceed:\n\n\tC: 000cgive-up\n\n  (s) Parse the upload-pack request:\n\n      Verify all objects in WANT are reachable from refs.  As\n      this may require walking backwards through history to\n      the very beginning on invalid requests the server may\n      use a reasonable limit of commits (e.g. 1000) walked\n      beyond any ref tip before giving up.\n\n      If no WANT objects are received, send an error:\n\n\tS: 0019status error no want\n\n      If any WANT object is not reachable, send an error:\n\n\tS: 001estatus error invalid want\n\n     Create an empty list, S_COMMON.\n\n     If 'common' was sent:\n\n     Load all objects into S_COMMON.  If an object appears in\n     'common' but the server does not have the object locally\n     an error should be returned:\n\n\tS: 001estatus error invalid common\n\n     If 'have' was sent:\n\n     Loop through the objects in the order supplied by the client.\n     For each object, if the server has the object reachable from\n     a ref, add it to S_COMMON.  If a commit is added to S_COMMON,\n     do not add any ancestors, even if they also appear in HAVE.\n\n  (s) Send the upload-pack response:\n\n     If the server has found a closed set of objects to pack or the\n     request contains \"give-up\", it replies with the pack and the\n     enabled capabilities.  The set of enabled capabilities is limited\n     to the intersection of what the client requested and what the\n     server supports.\n\n\tS: 0010status pack\n\tC: 001bcapability include-tag\n\tC: 0019capability thin-pack\n\tS: 000c.PACK...\n\n     The returned stream is the side-band-64k protocol supported\n     by the git-upload-pack service, and the pack is embedded into\n     stream 1.  Progress messages from the server side may appear\n     in stream 2.\n\n     Here a \"closed set of objects\" is defined to have at least\n     one path from every WANT to at least one COMMON object.\n\n     If the server needs more information, it replies with a\n     status continue response:\n\n\tS: 0014status continue\n\tS: 0034common <S_COMMON #1>...........................\n\tS: 0034common <S_COMMON #2>...........................\n\t...\n\tS: 0000\n\n     The stream formatting rules are the same as the request.\n\n     The \"common\" command details the contents of S_COMMON,\n     that is all objects from HAVE that the server also has.\n\n  (c) Parse the upload-pack response:\n\n      If the status pkt-line is \"status pack:\"\n\n      Process the pack stream and update the local refs.\n\n      If the status pkt-line is \"status continue\":\n\n      Reset COMMON to the items in S_COMMON.  The new S_COMMON\n      should be a superset of the existing COMMON set.\n\n      Remove all items in S_COMMON, and all of their ancestors,\n      from PENDING.\n\n      Do another compute step.\n\nCapability include-tag\n~~~~~~~~~~~~~~~~~~~~~~\n\nWhen packing an object that an annotated tag points at, include\nthe tag object too.  Clients can request this if they want to\nfetch tags, but don't know which tags they will need until after\nthey receive the branch data.  By enabling include-tag an entire\ncall to upload-pack can be avoided.\n\nCapability thin-pack\n~~~~~~~~~~~~~~~~~~~~\n\nWhen packing a deltified object the base is not included if the\nbase is reachable from an object listed in the COMMON set by the\nclient.  This reduces the bandwidth required to transfer, but it\ndoes slightly increase processing time for the client to save the\npack to disk.\n\nService receive-pack\n--------------------\n\nUploads a pack and updates refs.\n\nURL: $url/backend.git-http/receive-pack\nContent-Type: application/x-git; service=receive-pack\n\nThe start of the stream is the commands to update the refs and\nthe remainder of the stream is the pack file itself.  See\ngit-receive-pack and its network protocol in pack-protocol.txt,\nas this is essentially the same.\n\n\tC: 006395dcfa3633004da0049d3d0fa03f80589cbcaf31 d049f6c27a2244e12041955e262a404c7faba355 refs/heads/maint\n\tC: 0000\n\tC: PACK...\n\n\tS: 0005\n\tS: ...<output of receive-pack>...\n\nThe capabilities are handled exactly as in the fetch protocol,\nhowever the server may reject a pack and its associated commands\nif an invalid capability request is made by the client, or the\nclient has assumed a pack capability that the server does not\nhave support for.  In the latter case the server must still send\nthe capabilities key in the response so the client can correct\nitself and try again.\n\n-- \nShawn.\n"},{"id":"89152","messageId":"alpine.LFD.1.10.0808291546570.1624@xanadu.home","threadId":"15201","inReplyTo":"20080829173954.GG7403@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-08-29T19:55:44Z","receivedAt":"2008-08-29T19:55:44Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Fri, 29 Aug 2008, Shawn O. Pearce wrote:\n\n> Yet another draft follows.  I believe that I have covered all\n> comments with this draft.  But I welcome any additional ones,\n> as thus far it has been a very constructive process.\n> \n> The updated protocol looks more like the current native protocol\n> does.  This should make it easier to reuse code between the two\n> protocol implementations.\n[...]\n\n> pkt-line Format\n> ---------------\n> \n> Much of the payload is described around pkt-lines.\n> \n> A pkt-line is a variable length binary string.  The first four bytes\n> of the line indicates the total length of the line, in hexadecimal.\n> The total length includes the 4 bytes used to denote the length.  A\n> line is usually terminated by an LF, which must be included in the\n> total length if present.\n> \n> Binary data is permitted within a pkt-line so implementors should\n> ensure their pkt-line parsing/formatting routines are 8-bit clean.\n> The maximum length of a pkt-line's data is 65532 bytes (65536 - 4).\n\nShouldn't that be 65531, since you cannot represent 65536 with 4 hex \ndigits?\n\n> \tC: 001bcapability include-tag\n> \tC: 0019capability thin-pack\n> \t....\n[...]\n>      The \"capability\" command requests a single protocol feature\n>      to be enabled by the server.  Typically these are used to\n>      describe aspects of the pack that will be returned.  See\n>      below for more details on the current capabilities.\n\nWhy not having all capabilities listed at once on a single line instead?  \nThat's more or less what the current protocol does already.\n\n\nNicolas\n"},{"id":"89424","messageId":"905315640809010905w20f4ceeo43e7b0a14abd48a3@mail.gmail.com","threadId":"15201","inReplyTo":"20080829173954.GG7403@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"Tarmigan","fromEmail":"tarmigan+git@gmail.com","sentAt":"2008-09-01T16:05:16Z","receivedAt":"2008-09-01T16:05:16Z","isPatch":false,"sender":{"key":"tarmigan+git@gmail.com","avatar":null},"body":"On Fri, Aug 29, 2008 at 10:39 AM, Shawn O. Pearce <spearce@spearce.org> wrote:\n> Yet another draft follows.  I believe that I have covered all\n> comments with this draft.  But I welcome any additional ones,\n> as thus far it has been a very constructive process.\n\nSorry I'm jumping into this a bit late, but something just occurred to me.\n\n>\n> The updated protocol looks more like the current native protocol\n> does.  This should make it easier to reuse code between the two\n> protocol implementations.\n>\n> --8<--\n> Smart HTTP transfer protocols\n\n[...]\n\n> HTTP Redirects\n> --------------\n>\n> If a POST request results in an HTTP 302 or 303 redirect response\n> clients should retry the request by updating the URL and POSTing\n> the same request to the new location.  Subsequent requests should\n> still be sent to the original URL.\n>\n> This redirect behavior is unrelated to the in-payload redirect\n> that is described below in \"Service show-ref\".\n\nI just want to see smart http could support a new feature (please yell\nif git:// already supports this and I am not aware of it).   The idea\nis from http://lkml.org/lkml/2008/8/21/347, the relevant portion\nbeing:\n\nGreg KH wrote:\n>David Vrabel wrote:\n>> Or you can pull the changes from the uwb branch of\n>>\n>> git://pear.davidvrabel.org.uk/git/uwb.git\n>>\n>> (Please don't clone the entire tree from here as I have very limited\n>> bandwidth.)\n>\n> If this is an issue, I think you can use the --reference option to\n> git-clone when creating the tree to reference an external tree (like\n> Linus's).  That way you don't have the whole tree on your server for\n> stuff like this.\n\nI do not believe that the server (either git:// or http://) can\ncurrently be setup with --reference to redirect to another server for\ncertain refs, but perhaps with smart http and the POST 302/303\nredirect responses, this would now be possible as a way to reduce\nbandwidth for people's home servers?  I have also seen similar\nrequests before (\"don't pull the whole kernel from me, just add my\nrepo as a remote after you've cloned linus-2.6\"), so for larger\nprojects, it might be a nice feature.  Would that be something\ndesirable to support?\n\nWould the current proposal be able to support this kind of partial\nredirect?  I don't quite see how it would, but it seems very close.\nPerhaps if the show-ref redirect could appear partway through the\nshow-ref response and then the client could go off, fetch the some\nrefs from that server and then return to the original server for the\nremainder?  Or maybe in the upload-pack negotiations, there could be a\nspecial redirect command as part of the \"status continue\" response\nthat told the client to run off and look for a specific sha at another\nurl?  Something like\n\nstatus continue\n\n S: 0014status continue\n       S: 0034common <S_COMMON #1>...........................\n       S: 0034common <S_COMMON #2>...........................\n       ...\n\n\nOtherwise, it looks very cool, but I have a few more minor questions\nto help my general understanding...\n\n>     If the client has sent 256 HAVE commits and has not yet\n>     received one of those back from S_COMMON, or the client\n>     has emptied C_PENDING it should include a \"give-up\"\n>     command to let the server know it won't proceed:\n>\n>        C: 000cgive-up\n\nWhat does the server do after a 000cgive-up ?  Does the server send\nback a complete pack (like a new clone) or if not, how does clone work\nover smart http?  Does that mean that if I fall more than 256 commits\nbehind, I have to redownload the whole repo?  Or am I missing\nsomething about the the C_PENDING commits being sparse and doing some\nkind of smart back-off (I'm not at all familiar with the existing\nreceive-pack/upload-pack)?\n\n>  (s) Parse the upload-pack request:\n>\n>      Verify all objects in WANT are reachable from refs.  As\n>      this may require walking backwards through history to\n>      the very beginning on invalid requests the server may\n>      use a reasonable limit of commits (e.g. 1000) walked\n>      beyond any ref tip before giving up.\n>\n>      If no WANT objects are received, send an error:\n>\n>        S: 0019status error no want\n>\n>      If any WANT object is not reachable, send an error:\n>\n>        S: 001estatus error invalid want\n\nSo again, if the client falls more than 1000 commits behind (not hard\nto do for example during the linux merge window), and then the client\nWANTs HEAD^1001, what happens?  Does the get nothing from the server,\nor does the client essentially reclone, or I am missing something?\n\n>  (s) Send the upload-pack response:\n>\n>     If the server has found a closed set of objects to pack or the\n>     request contains \"give-up\", it replies with the pack and the\n>     enabled capabilities.  The set of enabled capabilities is limited\n>     to the intersection of what the client requested and what the\n>     server supports.\n>\n>        S: 0010status pack\n>        C: 001bcapability include-tag\n>        C: 0019capability thin-pack\n>        S: 000c.PACK...\n\nShould these be all S: ... ?\n\nThanks,\nTarmigan\n"},{"id":"89425","messageId":"905315640809010913v56961d41nc397f71f71077e39@mail.gmail.com","threadId":"15201","inReplyTo":"905315640809010905w20f4ceeo43e7b0a14abd48a3@mail.gmail.com","subject":"Re: Git-aware HTTP transport","fromName":"Tarmigan","fromEmail":"tarmigan+git@gmail.com","sentAt":"2008-09-01T16:13:42Z","receivedAt":"2008-09-01T16:13:42Z","isPatch":false,"sender":{"key":"tarmigan+git@gmail.com","avatar":null},"body":"(Oops, hit send too early by mistake, so some of my thoughts were incomplete)\n\nOn Mon, Sep 1, 2008 at 9:05 AM, Tarmigan <tarmigan+git@gmail.com> wrote:\n> On Fri, Aug 29, 2008 at 10:39 AM, Shawn O. Pearce <spearce@spearce.org> wrote:\n>> Yet another draft follows.  I believe that I have covered all\n>> comments with this draft.  But I welcome any additional ones,\n>> as thus far it has been a very constructive process.\n>\n> Sorry I'm jumping into this a bit late, but something just occurred to me.\n>\n>>\n>> The updated protocol looks more like the current native protocol\n>> does.  This should make it easier to reuse code between the two\n>> protocol implementations.\n>>\n>> --8<--\n>> Smart HTTP transfer protocols\n>\n> [...]\n>\n>> HTTP Redirects\n>> --------------\n>>\n>> If a POST request results in an HTTP 302 or 303 redirect response\n>> clients should retry the request by updating the URL and POSTing\n>> the same request to the new location.  Subsequent requests should\n>> still be sent to the original URL.\n>>\n>> This redirect behavior is unrelated to the in-payload redirect\n>> that is described below in \"Service show-ref\".\n>\n> I just want to see smart http could support a new feature (please yell\n> if git:// already supports this and I am not aware of it).   The idea\n> is from http://lkml.org/lkml/2008/8/21/347, the relevant portion\n> being:\n>\n> Greg KH wrote:\n>>David Vrabel wrote:\n>>> Or you can pull the changes from the uwb branch of\n>>>\n>>> git://pear.davidvrabel.org.uk/git/uwb.git\n>>>\n>>> (Please don't clone the entire tree from here as I have very limited\n>>> bandwidth.)\n>>\n>> If this is an issue, I think you can use the --reference option to\n>> git-clone when creating the tree to reference an external tree (like\n>> Linus's).  That way you don't have the whole tree on your server for\n>> stuff like this.\n>\n> I do not believe that the server (either git:// or http://) can\n> currently be setup with --reference to redirect to another server for\n> certain refs, but perhaps with smart http and the POST 302/303\n> redirect responses, this would now be possible as a way to reduce\n> bandwidth for people's home servers?  I have also seen similar\n> requests before (\"don't pull the whole kernel from me, just add my\n> repo as a remote after you've cloned linus-2.6\"), so for larger\n> projects, it might be a nice feature.  Would that be something\n> desirable to support?\n>\n> Would the current proposal be able to support this kind of partial\n> redirect?  I don't quite see how it would, but it seems very close.\n> Perhaps if the show-ref redirect could appear partway through the\n> show-ref response and then the client could go off, fetch the some\n> refs from that server and then return to the original server for the\n> remainder?  Or maybe in the upload-pack negotiations, there could be a\n> special redirect command as part of the \"status continue\" response\n> that told the client to run off and look for a specific sha at another\n> url?  Something like\n>\n> status continue\n>\n>  S: 0014status continue\n>       S: 0034common <S_COMMON #1>...........................\n>       S: 0034common <S_COMMON #2>...........................\n>       ...\n\nI meant to write:\n\n         S: 0014status continue\n         S: 0034common <S_COMMON #1>...........................\n         S: 0034common <S_COMMON #2>...........................\n         S: 00xxredirect <WILL_BE_COMMON> <REMOTE_URL>\n\nand then the client could go try the remote url, fetch that SHA and\nancestors, and then resume the upload pack negotiations with\n<WILL_BE_COMMON> among the <COMMON> commits.   Obviously it's still\nsomewhat of a half baked idea, and would probably need some kind of\nfallback, but does that seem like a reasonable thing to do and a\nreasonable way to do it?\n\n>\n> Otherwise, it looks very cool, but I have a few more minor questions\n> to help my general understanding...\n>\n>>     If the client has sent 256 HAVE commits and has not yet\n>>     received one of those back from S_COMMON, or the client\n>>     has emptied C_PENDING it should include a \"give-up\"\n>>     command to let the server know it won't proceed:\n>>\n>>        C: 000cgive-up\n>\n> What does the server do after a 000cgive-up ?  Does the server send\n> back a complete pack (like a new clone) or if not, how does clone work\n> over smart http?  Does that mean that if I fall more than 256 commits\n> behind, I have to redownload the whole repo?  Or am I missing\n> something about the the C_PENDING commits being sparse and doing some\n> kind of smart back-off (I'm not at all familiar with the existing\n> receive-pack/upload-pack)?\n>\n>>  (s) Parse the upload-pack request:\n>>\n>>      Verify all objects in WANT are reachable from refs.  As\n>>      this may require walking backwards through history to\n>>      the very beginning on invalid requests the server may\n>>      use a reasonable limit of commits (e.g. 1000) walked\n>>      beyond any ref tip before giving up.\n>>\n>>      If no WANT objects are received, send an error:\n>>\n>>        S: 0019status error no want\n>>\n>>      If any WANT object is not reachable, send an error:\n>>\n>>        S: 001estatus error invalid want\n>\n> So again, if the client falls more than 1000 commits behind (not hard\n> to do for example during the linux merge window), and then the client\n> WANTs HEAD^1001, what happens?  Does the get nothing from the server,\n> or does the client essentially reclone, or I am missing something?\n>\n>>  (s) Send the upload-pack response:\n>>\n>>     If the server has found a closed set of objects to pack or the\n>>     request contains \"give-up\", it replies with the pack and the\n>>     enabled capabilities.  The set of enabled capabilities is limited\n>>     to the intersection of what the client requested and what the\n>>     server supports.\n>>\n>>        S: 0010status pack\n>>        C: 001bcapability include-tag\n>>        C: 0019capability thin-pack\n>>        S: 000c.PACK...\n>\n> Should these be all S: ... ?\n\nThanks,\nTarmigan\n"},{"id":"89480","messageId":"20080902060608.GG13248@spearce.org","threadId":"15201","inReplyTo":"905315640809010905w20f4ceeo43e7b0a14abd48a3@mail.gmail.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-09-02T06:06:08Z","receivedAt":"2008-09-02T06:06:08Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Tarmigan <tarmigan+git@gmail.com> wrote:\n> On Fri, Aug 29, 2008 at 10:39 AM, Shawn O. Pearce <spearce@spearce.org> wrote:\n> \n> I just want to see smart http could support a new feature (please yell\n> if git:// already supports this and I am not aware of it).   The idea\n> is from http://lkml.org/lkml/2008/8/21/347, the relevant portion\n> being:\n> \n> Greg KH wrote:\n> >David Vrabel wrote:\n> >> Or you can pull the changes from the uwb branch of\n> >>\n> >> git://pear.davidvrabel.org.uk/git/uwb.git\n> >>\n> >> (Please don't clone the entire tree from here as I have very limited\n> >> bandwidth.)\n> >\n> > If this is an issue, I think you can use the --reference option to\n> > git-clone when creating the tree to reference an external tree (like\n> > Linus's).  That way you don't have the whole tree on your server for\n> > stuff like this.\n> \n> I do not believe that the server (either git:// or http://) can\n> currently be setup with --reference to redirect to another server for\n> certain refs,\n\nCorrect.  Today _none_ of the transport protocols allow the server\nto force the client to use some sort of reference repository for an\ninitial clone.  There are likely two reasons for this:\n\n *) Its a lot simpler to program to just get everything from\n    one location.\n\n *) If you really are forking an open source project then in\n    some cases you may need to distribute the full source,\n\tnot your delta.  You may just as well distribute the full\n\tsource and call it a day.\n\nThe dumb http:// currently supports getting packs from a remote HTTP\nserver via its objects/info/http-alternates.  But the native and\nrsync protocols don't support that.  The logic behind http-alternates\nisn't to allow moving load onto a different server, but to make\nthe locally available alternate repository available through the\nsame web server.  The path on the UNIX filesystem that is used in\nobjects/info/alternates may not be the same path used in the web\nserver's namespace.\n\n> but perhaps with smart http and the POST 302/303\n> redirect responses, this would now be possible as a way to reduce\n> bandwidth for people's home servers?  I have also seen similar\n> requests before (\"don't pull the whole kernel from me, just add my\n> repo as a remote after you've cloned linus-2.6\"), so for larger\n> projects, it might be a nice feature.  Would that be something\n> desirable to support?\n\nI think this isn't a bad idea, but I'd rather have the server say\n\"In order to talk to me you need at least these objects in common\nwith me: ...\".  If you don't have those then the user should go\nfind it on their own, rather than forcing them to a particular URL\nand automatically following it.\n\nI'm a little concerned about a US user putting a US mirror site\nof kernel.org into the server and forcing a user in India to do a\nfull clone over the Atlantic links when they could have just used\na more local mirror for that initial \"linus-2.6\" clone.\n \n> Otherwise, it looks very cool, but I have a few more minor questions\n> to help my general understanding...\n> \n> >     If the client has sent 256 HAVE commits and has not yet\n> >     received one of those back from S_COMMON, or the client\n> >     has emptied C_PENDING it should include a \"give-up\"\n> >     command to let the server know it won't proceed:\n> >\n> >        C: 000cgive-up\n> \n> What does the server do after a 000cgive-up ?  Does the server send\n> back a complete pack (like a new clone) or if not, how does clone work\n> over smart http?\n\nWhen the server receives a \"give-up\" it needs to create a pack\nbased on \"git rev-list --objects-boundary $WANT --not $COMMON\".\nIf the set $COMMON is non-empty then its a partial pack; if $COMMON\nis empty then its a full clone.  This is what the native protocol\ndoes when the client gives up.\n\n> Does that mean that if I fall more than 256 commits\n> behind, I have to redownload the whole repo?\n\nYou are thinking the wrong way.  If you have more than 256 commits\nthat the other side doesn't have you may give up too early.\nFor that to be true you need to create 256 commits locally that\naren't on the remote peer and whose timestamps are all ahead of\nthe commits you last fetched from the remote peer.\n\nYes, it can happen.  But its less likely than you think because\nwe're talking about you doing 256 commits worth of development and\nnot picking up any new commits from remote peers in the middle of\nthat time period.  Get just one and it resets the counter back to\n0 and allows it to try another 256 commits before giving up.\n\nI should amend this section to talk about what giving up here\nreally means.  If we have nothing sent in common yet or maybe\nvery little sent in common we may have existing remote refs tied\nto this URL in .git/config that can send, and we may have one or\nmore annotated tags that we know for a fact are in common as both\npeers have the same tag name pointing to the same tag object.\n\nA smart(er) client might try to toss some recently dated annotated\ntags at the server before throwing a give-up if it would otherwise\nthrow a give-up.  Its likely to narrow the result set, and doesn't\nhurt if it doesn't.\n\n> >  (s) Parse the upload-pack request:\n> >\n> >      Verify all objects in WANT are reachable from refs.  As\n> >      this may require walking backwards through history to\n> >      the very beginning on invalid requests the server may\n> >      use a reasonable limit of commits (e.g. 1000) walked\n> >      beyond any ref tip before giving up.\n> >\n> >      If no WANT objects are received, send an error:\n> >\n> >        S: 0019status error no want\n> >\n> >      If any WANT object is not reachable, send an error:\n> >\n> >        S: 001estatus error invalid want\n> \n> So again, if the client falls more than 1000 commits behind (not hard\n> to do for example during the linux merge window), and then the client\n> WANTs HEAD^1001, what happens?  Does the get nothing from the server,\n> or does the client essentially reclone, or I am missing something?\n\nOh, this is a live-lock condition.  If the client grabs the list of\nrefs from the server, then has to wait 100 ms to get back to the\nserver and start upload-pack (due to latency) and in that 100ms\nwindow Linus shoves a 1001 commit merge into his tree then yes,\nthe server may abort and tell the client \"error invalid want\".\n\nAt which point the client may try to restart from the beginning,\nor just plain give up and tell the end user try again later.\n\nThis condition of 1000 is just some aribtrary limit to allow the\nclient to still continue with an in-progress download if right in\nthe middle of the client's RPCs the remote was modified by its owner.\n \n> >  (s) Send the upload-pack response:\n> >\n> >     If the server has found a closed set of objects to pack or the\n> >     request contains \"give-up\", it replies with the pack and the\n> >     enabled capabilities.  The set of enabled capabilities is limited\n> >     to the intersection of what the client requested and what the\n> >     server supports.\n> >\n> >        S: 0010status pack\n> >        C: 001bcapability include-tag\n> >        C: 0019capability thin-pack\n> >        S: 000c.PACK...\n> \n> Should these be all S: ... ?\n\nYes, thanks.  I will make the correction.  Damn copy and paste.\n\n-- \nShawn.\n"},{"id":"89481","messageId":"48BCD899.3010805@zytor.com","threadId":"15201","inReplyTo":"20080902060608.GG13248@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-09-02T06:09:29Z","receivedAt":"2008-09-02T06:09:29Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> Correct.  Today _none_ of the transport protocols allow the server\n> to force the client to use some sort of reference repository for an\n> initial clone.  There are likely two reasons for this:\n> \n>  *) Its a lot simpler to program to just get everything from\n>     one location.\n> \n>  *) If you really are forking an open source project then in\n>     some cases you may need to distribute the full source,\n> \tnot your delta.  You may just as well distribute the full\n> \tsource and call it a day.\n> \n\n3) it encourages single points of failure.\n\n\t-hpa\n"},{"id":"89482","messageId":"20080902061342.GI13248@spearce.org","threadId":"15201","inReplyTo":"48BCD899.3010805@zytor.com","subject":"Re: Git-aware HTTP transport","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-09-02T06:13:42Z","receivedAt":"2008-09-02T06:13:42Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> wrote:\n> Shawn O. Pearce wrote:\n>>\n>> Correct.  Today _none_ of the transport protocols allow the server\n>> to force the client to use some sort of reference repository for an\n>> initial clone.  There are likely two reasons for this:\n>>\n>>  *) Its a lot simpler to program to just get everything from\n>>     one location.\n>>\n>>  *) If you really are forking an open source project then in\n>>     some cases you may need to distribute the full source,\n>> \tnot your delta.  You may just as well distribute the full\n>> \tsource and call it a day.\n>>\n>\n> 3) it encourages single points of failure.\n\nOr bad network usage, as I pointed out later about an India user\nunknowingly being forced into a US based mirror when another was\ncloser to them.\n\nI didn't make it clear in my response but I'm really against our\nprotocol having this sort of explicit redirect.  I'd rather put a\nrequirement in that says \"Unless you have X,Y,Z in common with me\n(directly or indirectly) I'm just not going to give you a pack\".\n\nFWIW that fixes an issue for me at day-job that people will be\ncursing about later this year in public.  Not my fault.  We would\nall rather just publish the entire repository.  Instead we have\nto publish something that requires the user to clone it from\nanother source first, and use fetch or \"clone --reference\" to get\nour updates.  *sigh*\n\n-- \nShawn.\n"},{"id":"89521","messageId":"905315640809021120j13ee5f5t21e1d2618b63568c@mail.gmail.com","threadId":"15201","inReplyTo":"20080902060608.GG13248@spearce.org","subject":"Re: Git-aware HTTP transport","fromName":"Tarmigan","fromEmail":"tarmigan+git@gmail.com","sentAt":"2008-09-02T18:20:31Z","receivedAt":"2008-09-02T18:20:31Z","isPatch":false,"sender":{"key":"tarmigan+git@gmail.com","avatar":null},"body":"On Mon, Sep 1, 2008 at 11:06 PM, Shawn O. Pearce <spearce@spearce.org> wrote:\n>> What does the server do after a 000cgive-up ?  Does the server send\n>> back a complete pack (like a new clone) or if not, how does clone work\n>> over smart http?\n>\n> When the server receives a \"give-up\" it needs to create a pack\n> based on \"git rev-list --objects-boundary $WANT --not $COMMON\".\n> If the set $COMMON is non-empty then its a partial pack; if $COMMON\n> is empty then its a full clone.  This is what the native protocol\n> does when the client gives up.\n\nOK, that makes sense now.\n\n>> Does that mean that if I fall more than 256 commits\n>> behind, I have to redownload the whole repo?\n>\n> You are thinking the wrong way.  If you have more than 256 commits\n> that the other side doesn't have you may give up too early.\n> For that to be true you need to create 256 commits locally that\n> aren't on the remote peer and whose timestamps are all ahead of\n> the commits you last fetched from the remote peer.\n>\n> Yes, it can happen.  But its less likely than you think because\n> we're talking about you doing 256 commits worth of development and\n> not picking up any new commits from remote peers in the middle of\n> that time period.  Get just one and it resets the counter back to\n> 0 and allows it to try another 256 commits before giving up.\n>\n> I should amend this section to talk about what giving up here\n> really means.  If we have nothing sent in common yet or maybe\n> very little sent in common we may have existing remote refs tied\n> to this URL in .git/config that can send, and we may have one or\n> more annotated tags that we know for a fact are in common as both\n> peers have the same tag name pointing to the same tag object.\n>\n> A smart(er) client might try to toss some recently dated annotated\n> tags at the server before throwing a give-up if it would otherwise\n> throw a give-up.  Its likely to narrow the result set, and doesn't\n> hurt if it doesn't.\n\nYes, throwing in tags and remotes as a last resort sounds like a good idea.\n\n>> So again, if the client falls more than 1000 commits behind (not hard\n>> to do for example during the linux merge window), and then the client\n>> WANTs HEAD^1001, what happens?  Does the get nothing from the server,\n>> or does the client essentially reclone, or I am missing something?\n>\n> Oh, this is a live-lock condition.  If the client grabs the list of\n> refs from the server, then has to wait 100 ms to get back to the\n> server and start upload-pack (due to latency) and in that 100ms\n> window Linus shoves a 1001 commit merge into his tree then yes,\n> the server may abort and tell the client \"error invalid want\".\n\nAhh, now I get it.  Somehow I forgot that the WANTs were only boundary\ncommits and not a list of all the commits that the client wants.\n\nOn Mon, Sep 1, 2008 at 11:13 PM, Shawn O. Pearce <spearce@spearce.org> wrote:\n> \"H. Peter Anvin\" <hpa@zytor.com> wrote:\n>> Shawn O. Pearce wrote:\n>>>\n>>> Correct.  Today _none_ of the transport protocols allow the server\n>>> to force the client to use some sort of reference repository for an\n>>> initial clone.  There are likely two reasons for this:\n>>>\n>>>  *) Its a lot simpler to program to just get everything from\n>>>     one location.\n>>>\n>>>  *) If you really are forking an open source project then in\n>>>     some cases you may need to distribute the full source,\n>>>      not your delta.  You may just as well distribute the full\n>>>      source and call it a day.\n>>>\n>>\n>> 3) it encourages single points of failure.\n>\n> Or bad network usage, as I pointed out later about an India user\n> unknowingly being forced into a US based mirror when another was\n> closer to them.\n>\n> I didn't make it clear in my response but I'm really against our\n> protocol having this sort of explicit redirect.  I'd rather put a\n> requirement in that says \"Unless you have X,Y,Z in common with me\n> (directly or indirectly) I'm just not going to give you a pack\".\n\nOK, this all makes sense. http:// and git:// are probably the wrong\nprotocols to reduce bandwidth for the server for new clones.  Long\nterm, maybe gittorrent will be the right solution...\n\nThanks,\nTarmigan\n"},{"id":"209448","messageId":"511AED98.5070809@zytor.com","threadId":"15201","inReplyTo":"20080826012643.GD26523@spearce.org","subject":"Re: Git-aware HTTP transport docs","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2013-02-13T01:34:16Z","receivedAt":"2013-02-13T01:34:16Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Hi Shawn,\n\nYou wrote a really great protocol spec for the smart HTTP protocol back\nin the day.  It would be really great if it could be checked into the\ngit repository (updated if need be).  Someone mentioned today trying to\nreverse-engineer the protocol because of a lack of specs, and I was a\nbit surprised to day the least.\n\n\t-hpa\n"},{"id":"209449","messageId":"CAP2yMaLz=vpOVgpxG0CwVwWD_sq+T9px3w0KXE7doUFhKqNZWQ@mail.gmail.com","threadId":"15201","inReplyTo":"511AED98.5070809@zytor.com","subject":"Re: Git-aware HTTP transport docs","fromName":"Scott Chacon","fromEmail":"schacon@gmail.com","sentAt":"2013-02-13T02:23:02Z","receivedAt":"2013-02-13T02:23:02Z","isPatch":false,"sender":{"key":"schacon@gmail.com","avatar":"https://gravatar.com/avatar/9b13a8a078e1dcf8588c4eea9554445d51ebed6c41b51f56f4d96738130b05c6?d=mp&s=160"},"body":"I don't believe it was ever merged into the Git docs.  I have a copy of it here:\n\nhttps://www.dropbox.com/s/pwawp8kmwgyc3w2/http-protocol.txt\n\nScott\n\nOn Tue, Feb 12, 2013 at 5:34 PM, H. Peter Anvin <hpa@zytor.com> wrote:\n> Hi Shawn,\n>\n> You wrote a really great protocol spec for the smart HTTP protocol back\n> in the day.  It would be really great if it could be checked into the\n> git repository (updated if need be).  Someone mentioned today trying to\n> reverse-engineer the protocol because of a lack of specs, and I was a\n> bit surprised to day the least.\n>\n>         -hpa\n>\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"209473","messageId":"7vliasb5c0.fsf@alter.siamese.dyndns.org","threadId":"15201","inReplyTo":"CAP2yMaLz=vpOVgpxG0CwVwWD_sq+T9px3w0KXE7doUFhKqNZWQ@mail.gmail.com","subject":"Re: Git-aware HTTP transport docs","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2013-02-13T15:29:03Z","receivedAt":"2013-02-13T15:29:03Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Scott Chacon <schacon@gmail.com> writes:\n\n> I don't believe it was ever merged into the Git docs.  I have a copy of it here:\n>\n> https://www.dropbox.com/s/pwawp8kmwgyc3w2/http-protocol.txt\n\nThanks for a pointer.  It seems that it wasn't in a shape ready to\nbe \"merged\" yet.\n\nDoes somebody want to pick it up and polish it further?\n"}]}