{"thread":{"id":"13056","subject":"Corporate firewall braindamage","startedAt":"2008-04-10T21:11:19Z","lastAt":"2008-04-11T08:25:08Z","messageCount":6,"participants":["H. Peter Anvin","Junio C Hamano","Shawn O. Pearce","david@lang.hm"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"74097","messageId":"47FE8277.8070503@zytor.com","threadId":"13056","inReplyTo":null,"subject":"Corporate firewall braindamage","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-04-10T21:11:19Z","receivedAt":"2008-04-10T21:11:19Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"The apparent commonality of corporate firewall braindamage, and the \nresulting \"need\" of people to pull over dumb (http) transport, is an \nongoing problem on kernel.org.\n\nI have thought some about what can be done to improve the situation, and \nI have come up with the following list of possibilities, pretty much \nlisted in order from easiest and least generic to hardest and most generic.\n\nIt would be very interesting if people who have familiarity with this \nparticular class of braindamaged firewalls could comment on how many \nusers would be helped by which ones of these solutions.\n\n\n1. git protocol via CONNECT http proxy\n\n    Connect to http proxy, and use a CONNECT method to establish a link\n    to the git server, using the normal git protocol.\n\n    Minor change to TCP connection setup, but no other changes needed.\n    No changes on the server side.\n\n\n2. git protocol over SSL via CONNECT http proxy\n\n    Same as #1, but encapsulate the data stream in an SSL connection.\n    If the git server is run on port 443, then the fact that the data\n    on the SSL connection isn't actually HTTP should be invisible to the\n    proxy, and thus this *should* work anywhere which allows https://\n    traffic.\n\n    Requires the git server to speak SSL.\n\n\n3. git protocol encapsulated in HTTP POST transaction\n\n    git protocol is already fundamentally a RPC protocol, where the\n    client sends a query and the server responds.  Furthermore, it\n    tries to minimize the number of round trips (RPC calls), which is\n    of course desirable.\n\n    Each such RPC transaction could be formulated as an HTTP POST\n    transaction.\n\n    This requires modifications to both the client and the server;\n    furthermore, the server can no longer rely on the invariant \"one TCP\n    connection == one session\"; a proxy might break a single session\n    into arbitrarily many TCP connections.\n\nThoughts?\n\n\t-hpa\n"},{"id":"74102","messageId":"7v7if5wbdd.fsf@gitster.siamese.dyndns.org","threadId":"13056","inReplyTo":"47FE8277.8070503@zytor.com","subject":"Re: Corporate firewall braindamage","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-04-10T23:14:54Z","receivedAt":"2008-04-10T23:14:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> 1. git protocol via CONNECT http proxy\n>\n>    Connect to http proxy, and use a CONNECT method to establish a link\n>    to the git server, using the normal git protocol.\n>\n>    Minor change to TCP connection setup, but no other changes needed.\n>    No changes on the server side.\n\nMany firewalls will detect that CONNECT will not going to 443 and block\nyou, and even if you run git:// daemon on 443, they will detect that you\nare not talking SSL initial exchange and shut you off.\n\n> 2. git protocol over SSL via CONNECT http proxy\n>\n>    Same as #1, but encapsulate the data stream in an SSL connection.\n>    If the git server is run on port 443, then the fact that the data\n>    on the SSL connection isn't actually HTTP should be invisible to the\n>    proxy, and thus this *should* work anywhere which allows https://\n>    traffic.\n>\n>    Requires the git server to speak SSL.\n\nYes, perhaps putting it behind an independent ssl relay would give you a\nsolution without any code change.\n\n> 3. git protocol encapsulated in HTTP POST transaction\n>\n>    git protocol is already fundamentally a RPC protocol, where the\n>    client sends a query and the server responds.  Furthermore, it\n>    tries to minimize the number of round trips (RPC calls), which is\n>    of course desirable.\n>\n>    Each such RPC transaction could be formulated as an HTTP POST\n>    transaction.\n>\n>    This requires modifications to both the client and the server;\n>    furthermore, the server can no longer rely on the invariant \"one TCP\n>    connection == one session\"; a proxy might break a single session\n>    into arbitrarily many TCP connections.\n\nIt would probably be a one-CS/EE-student-half-a-summer sized project to\ncreate such a server-side support with a specialized client.\n"},{"id":"74104","messageId":"20080410233328.GQ10274@spearce.org","threadId":"13056","inReplyTo":"7v7if5wbdd.fsf@gitster.siamese.dyndns.org","subject":"Re: Corporate firewall braindamage","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2008-04-10T23:33:28Z","receivedAt":"2008-04-10T23:33:28Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Junio C Hamano <gitster@pobox.com> wrote:\n> \"H. Peter Anvin\" <hpa@zytor.com> writes:\n> > 3. git protocol encapsulated in HTTP POST transaction\n> >\n> >    git protocol is already fundamentally a RPC protocol, where the\n> >    client sends a query and the server responds.  Furthermore, it\n> >    tries to minimize the number of round trips (RPC calls), which is\n> >    of course desirable.\n> >\n> >    Each such RPC transaction could be formulated as an HTTP POST\n> >    transaction.\n> >\n> >    This requires modifications to both the client and the server;\n> >    furthermore, the server can no longer rely on the invariant \"one TCP\n> >    connection == one session\"; a proxy might break a single session\n> >    into arbitrarily many TCP connections.\n> \n> It would probably be a one-CS/EE-student-half-a-summer sized project to\n> create such a server-side support with a specialized client.\n\nFunny you say that.  This was a GSoC 2008 project idea.  We even\nreceived an application from a student for it.\n\nThe hard part is either making the server side stateful, so it can\nremember what the last RCP call had said it wants/haves, or doing a\nstateless protocol where the client uses an exponential expansion\n(or some such behavior) of its have list until the server replies\nwith the pack data.\n\n-- \nShawn.\n"},{"id":"74105","messageId":"47FEA7AE.1050403@zytor.com","threadId":"13056","inReplyTo":"20080410233328.GQ10274@spearce.org","subject":"Re: Corporate firewall braindamage","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2008-04-10T23:50:06Z","receivedAt":"2008-04-10T23:50:06Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Shawn O. Pearce wrote:\n> \n> Funny you say that.  This was a GSoC 2008 project idea.  We even\n> received an application from a student for it.\n> \n> The hard part is either making the server side stateful, so it can\n> remember what the last RCP call had said it wants/haves, or doing a\n> stateless protocol where the client uses an exponential expansion\n> (or some such behavior) of its have list until the server replies\n> with the pack data.\n> \n\nOne easy way of doing the former is to have a session reassociator in \nthe flow; pretty much a multiplexer which receives the HTTP request, and \npasses it onto a work slave (which can be an ordinary process, in fact, \ncan be the ordinary git daemon) based on a session and sequence ID.\n\n\t-hpa\n"},{"id":"74109","messageId":"47FEB8E4.3070306@kernel.org","threadId":"13056","inReplyTo":"47FEA7AE.1050403@zytor.com","subject":"Re: Corporate firewall braindamage","fromName":"H. Peter Anvin","fromEmail":"hpa@kernel.org","sentAt":"2008-04-11T01:03:32Z","receivedAt":"2008-04-11T01:03:32Z","isPatch":false,"sender":{"key":"hpa@kernel.org","avatar":null},"body":"H. Peter Anvin wrote:\n> Shawn O. Pearce wrote:\n>>\n>> Funny you say that.  This was a GSoC 2008 project idea.  We even\n>> received an application from a student for it.\n>>\n>> The hard part is either making the server side stateful, so it can\n>> remember what the last RCP call had said it wants/haves, or doing a\n>> stateless protocol where the client uses an exponential expansion\n>> (or some such behavior) of its have list until the server replies\n>> with the pack data.\n>>\n> \n> One easy way of doing the former is to have a session reassociator in \n> the flow; pretty much a multiplexer which receives the HTTP request, and \n> passes it onto a work slave (which can be an ordinary process, in fact, \n> can be the ordinary git daemon) based on a session and sequence ID.\n> \n\ns/multiplexer/demultiplexer/\n\nThe best might be to turn the demultiplexer either into an Apache module \nor some scripting language which can run inside Apache (e.g. mod_perl) \nto avoid Apache spawning a CGI program which is only used to talk to the \ngit daemon backend.\n\n\t-hpa\n"},{"id":"74116","messageId":"alpine.DEB.1.10.0804110123030.4615@asgard","threadId":"13056","inReplyTo":"7v7if5wbdd.fsf@gitster.siamese.dyndns.org","subject":"Re: Corporate firewall braindamage","fromName":"","fromEmail":"david@lang.hm","sentAt":"2008-04-11T08:25:08Z","receivedAt":"2008-04-11T08:25:08Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 10 Apr 2008, Junio C Hamano wrote:\n\n> \"H. Peter Anvin\" <hpa@zytor.com> writes:\n>\n>> 1. git protocol via CONNECT http proxy\n>>\n>>    Connect to http proxy, and use a CONNECT method to establish a link\n>>    to the git server, using the normal git protocol.\n>>\n>>    Minor change to TCP connection setup, but no other changes needed.\n>>    No changes on the server side.\n>\n> Many firewalls will detect that CONNECT will not going to 443 and block\n> you, and even if you run git:// daemon on 443, they will detect that you\n> are not talking SSL initial exchange and shut you off.\n>\n>> 2. git protocol over SSL via CONNECT http proxy\n>>\n>>    Same as #1, but encapsulate the data stream in an SSL connection.\n>>    If the git server is run on port 443, then the fact that the data\n>>    on the SSL connection isn't actually HTTP should be invisible to the\n>>    proxy, and thus this *should* work anywhere which allows https://\n>>    traffic.\n>>\n>>    Requires the git server to speak SSL.\n>\n> Yes, perhaps putting it behind an independent ssl relay would give you a\n> solution without any code change.\n\nin more pananoid locations they are putting client certs on desktops and \ngiving those to the IDS systems so that they can decrypt the SSL traffic, \nso if it doesn't look like HTTP inside the SSL they will block it.\n\nthis isn't very common now, but the firewalls that are blocking #1 weren't \nvery common a year or so ago either.\n\n>> 3. git protocol encapsulated in HTTP POST transaction\n>>\n>>    git protocol is already fundamentally a RPC protocol, where the\n>>    client sends a query and the server responds.  Furthermore, it\n>>    tries to minimize the number of round trips (RPC calls), which is\n>>    of course desirable.\n>>\n>>    Each such RPC transaction could be formulated as an HTTP POST\n>>    transaction.\n>>\n>>    This requires modifications to both the client and the server;\n>>    furthermore, the server can no longer rely on the invariant \"one TCP\n>>    connection == one session\"; a proxy might break a single session\n>>    into arbitrarily many TCP connections.\n>\n> It would probably be a one-CS/EE-student-half-a-summer sized project to\n> create such a server-side support with a specialized client.\n\nthis is probably the best long-term option.\n\nDavid Lang\n"}]}