{"thread":{"id":"3692","subject":"Re: Cloning from sites with 404 overridden","startedAt":"2006-03-22T02:59:21Z","lastAt":"2006-03-24T18:40:32Z","messageCount":18,"participants":["linux@horizon.com","Shawn Pearce","Linus Torvalds","Marco Costalba","Junio C Hamano","Andreas Ericsson","Nick Hengeveld","Radoslaw Szkodzinski","Mark Wooding","Morten Welinder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"17750","messageId":"20060322025921.1722.qmail@science.horizon.com","threadId":"3692","inReplyTo":null,"subject":"Re: Cloning from sites with 404 overridden","fromName":"","fromEmail":"linux@horizon.com","sentAt":"2006-03-22T02:59:21Z","receivedAt":"2006-03-22T02:59:21Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"If someone feels ambitious, you can detect this condition automatically\nby searching for a file that you know won't be there and seeing if you\nget a 404 response to that.\n\nTo avoid punishing good servers, it would be nice to defer the test\nuntil reciving the first corrupted object.\n\nI'm not sure what the best \"object that's not supposed to be there\" is.\nIt could just be a random hash, or would a malformed object file name\nbe better?  Any fixed name has a finite chance of being created by\nsomeone somewhere, but generating 160-bit random numbers is a PITA on\nnon-freenix platforms.\n\n\n(As an aside, I suspect this is all caused by Microsoft's \"friendly HTML\nerror messages\" invention.)\n"},{"id":"17751","messageId":"20060322031200.GB17954@spearce.org","threadId":"3692","inReplyTo":"20060322025921.1722.qmail@science.horizon.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-03-22T03:12:00Z","receivedAt":"2006-03-22T03:12:00Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"'0' x 40.  :-) There's some places already in the GIT source\nwhich would have ``issues'' if they got an object with this hash.\nNot sure if it is actually an entirely impossible hash or just one\nthat is highly improbable.\n\nMy own website has this problem and its because I'm using WordPress\nto handle all URLs on the site; I haven't yet found a way to\nconfigure WordPress to return a proper 404 when the URL can't be\nmapped to something on the server.  Note that 404 status codes can\nin fact return pretty HTML content for the user, and many websites\ndo this and many browsers display that pretty HTML.  But a bot can\nthen also recognize the status code and DTRT.\n\nThe webservers are just plain broken, mine included.  I think the\nbest option is to delay corrupt object reporting to the end of\nthe download process if you get only one corrupt object and that\ncorrupt object was actually attainable from a pack.  And in this\ncase its just a minor warning:\n\n\tWarning: The server appears to not return proper HTTP status\n\tcodes on missing files.  The files were found in one or\n\tmore packs so the download is OK, but the server administrator\n\tshould really fix their server.  If you know the server\n\tadministrator you might want to prod them to do so.\n\nBut that's already been suggested and I thought someone worked up\na patch based on that idea?  If not I could try to do so since my\nown damn server has the problem.  :-)\n\nlinux@horizon.com wrote:\n> If someone feels ambitious, you can detect this condition automatically\n> by searching for a file that you know won't be there and seeing if you\n> get a 404 response to that.\n> \n> To avoid punishing good servers, it would be nice to defer the test\n> until reciving the first corrupted object.\n> \n> I'm not sure what the best \"object that's not supposed to be there\" is.\n> It could just be a random hash, or would a malformed object file name\n> be better?  Any fixed name has a finite chance of being created by\n> someone somewhere, but generating 160-bit random numbers is a PITA on\n> non-freenix platforms.\n> \n> \n> (As an aside, I suspect this is all caused by Microsoft's \"friendly HTML\n> error messages\" invention.)\n\n-- \nShawn.\n"},{"id":"17755","messageId":"Pine.LNX.4.64.0603212011430.26286@g5.osdl.org","threadId":"3692","inReplyTo":"20060322031200.GB17954@spearce.org","subject":"Re: Cloning from sites with 404 overridden","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-03-22T04:13:21Z","receivedAt":"2006-03-22T04:13:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 21 Mar 2006, Shawn Pearce wrote:\n>\n> '0' x 40.  :-) There's some places already in the GIT source\n> which would have ``issues'' if they got an object with this hash.\n> Not sure if it is actually an entirely impossible hash or just one\n> that is highly improbable.\n\nThe all-zeroes hash is as improbable as any other one, and finding a \n\"collision\" (ie a \"real object\") with that hash is as improbable as any \nother collision, ie we can (and do) depend on it beign a unique identifier \nfor \"does not exist\".\n\n\t\t\tLinus\n"},{"id":"17760","messageId":"e5bfff550603212206k1924d352xcbdb0e5a11b88a50@mail.gmail.com","threadId":"3692","inReplyTo":"20060322025921.1722.qmail@science.horizon.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-03-22T06:06:13Z","receivedAt":"2006-03-22T06:06:13Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 21 Mar 2006 21:59:21 -0500, linux@horizon.com <linux@horizon.com> wrote:\n> If someone feels ambitious, you can detect this condition automatically\n> by searching for a file that you know won't be there and seeing if you\n> get a 404 response to that.\n>\n\nPerhaps I am proposing a total idiocy, I don't know git-fetch\ninternals, but wouldn't be better to avoid trying to download a non\nexisting object? So to fix the problem at the origin?\n\nI don't know if it is possible to list contents before try to download\nso to avoid asking for a non existing object.\n\nMarco\n"},{"id":"17761","messageId":"7v1wwvascr.fsf@assigned-by-dhcp.cox.net","threadId":"3692","inReplyTo":"e5bfff550603212206k1924d352xcbdb0e5a11b88a50@mail.gmail.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-22T06:47:16Z","receivedAt":"2006-03-22T06:47:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Marco Costalba\" <mcostalba@gmail.com> writes:\n\n> Perhaps I am proposing a total idiocy, I don't know git-fetch\n> internals, but wouldn't be better to avoid trying to download a non\n> existing object? So to fix the problem at the origin?\n>\n> I don't know if it is possible to list contents before try to download\n> so to avoid asking for a non existing object.\n\nThere is no way for the downloader to know if the upstream\nrepository has packed which object.  What is happening is that\nthe commit walker asks for loose object first because it does\nnot know.  Upon getting a \"no such file\" (or in the case of\nmisconfigured HTTP server that does not say 404, \"corrupt\nobject\"), it then checks if the object appears in the pack by\ndownloading the pack index.  It can tell what objects are in the\npacks by looking at the pack index and downloads the pack that\ncontains needed object.\n"},{"id":"17768","messageId":"442152E0.4020604@op5.se","threadId":"3692","inReplyTo":"20060322025921.1722.qmail@science.horizon.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-03-22T13:36:32Z","receivedAt":"2006-03-22T13:36:32Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"linux@horizon.com wrote:\n> If someone feels ambitious, you can detect this condition automatically\n> by searching for a file that you know won't be there and seeing if you\n> get a 404 response to that.\n> \n> To avoid punishing good servers, it would be nice to defer the test\n> until reciving the first corrupted object.\n> \n> I'm not sure what the best \"object that's not supposed to be there\" is.\n\n.git/objects/00/hoping-for-a-404-or-webadmin-should-fix\n\nIt has the right number of chars so it should fit in wherever a real \nobject name does but is obviously bogus anyways.\n\n\n> It could just be a random hash, or would a malformed object file name\n> be better?\n\nA malformed object name is infinitely better. Otherwise we'd end up with \na wild guess that hits home some day, to much surprise and a bug-report \nI wouldn't want to track. Not to mention the embarrassment when \nexplaining why that object-name was chosen.\n\n> \n> (As an aside, I suspect this is all caused by Microsoft's \"friendly HTML\n> error messages\" invention.)\n\nThe body of the 404-page has absolutely nothing to do with it.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"17770","messageId":"20060322172227.GO3997@reactrix.com","threadId":"3692","inReplyTo":"20060322025921.1722.qmail@science.horizon.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2006-03-22T17:22:27Z","receivedAt":"2006-03-22T17:22:27Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Tue, Mar 21, 2006 at 09:59:21PM -0500, linux@horizon.com wrote:\n\n> If someone feels ambitious, you can detect this condition automatically\n> by searching for a file that you know won't be there and seeing if you\n> get a 404 response to that.\n\nIt might be feasible to detect this condition using the Content-Type:\nheader in the server response.  So far, all the GIT repositories I've\ntried return text/plain for loose objects and a special 404 page will\nlikely be text/html.\n\n-- \nFor a successful technology, reality must take precedence over public\nrelations, for nature cannot be fooled.\n"},{"id":"17771","messageId":"20060322183621.GP3997@reactrix.com","threadId":"3692","inReplyTo":"20060322172227.GO3997@reactrix.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2006-03-22T18:36:21Z","receivedAt":"2006-03-22T18:36:21Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Wed, Mar 22, 2006 at 09:22:27AM -0800, Nick Hengeveld wrote:\n\n> It might be feasible to detect this condition using the Content-Type:\n> header in the server response.  So far, all the GIT repositories I've\n> tried return text/plain for loose objects and a special 404 page will\n> likely be text/html.\n\nSomething like this:\n\nhttp_fetch: report text/html responses for loose objects\n\nSome HTTP server environments return a 200 status and text/html error\ndocument or a redirect to one rather than a 404 status if a loose\nobject does not exist.  This patch detects and reports this condition\nto differentiate between a misconfigured server and an actual corrupt\nobject on the server.\n\nSigned-off-by: Nick Hengeveld <nickh@reactrix.com>\n\n\n---\n\n http-fetch.c |   19 ++++++++++++++++++-\n 1 files changed, 18 insertions(+), 1 deletions(-)\n\n61069cc348640fef2b8c503b8b8f00f689872cab\ndiff --git a/http-fetch.c b/http-fetch.c\nindex dc67218..ee5b585 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -41,6 +41,7 @@ struct object_request\n \tCURLcode curl_result;\n \tchar errorstr[CURL_ERROR_SIZE];\n \tlong http_code;\n+\tchar *content_type;\n \tunsigned char real_sha1[20];\n \tSHA_CTX c;\n \tz_stream stream;\n@@ -258,9 +259,15 @@ static void finish_object_request(struct\n \n static void process_object_response(void *callback_data)\n {\n+\tchar *content_type;\n \tstruct object_request *obj_req =\n \t\t(struct object_request *)callback_data;\n \n+\tcurl_easy_getinfo(obj_req->slot->curl, CURLINFO_CONTENT_TYPE,\n+\t\t\t  &content_type);\n+\tif (content_type)\n+\t\tobj_req->content_type = strdup(content_type);\n+\n \tobj_req->curl_result = obj_req->slot->curl_result;\n \tobj_req->http_code = obj_req->slot->http_code;\n \tobj_req->slot = NULL;\n@@ -298,6 +305,8 @@ static void release_object_request(struc\n \t\t\tentry->next = entry->next->next;\n \t}\n \n+\tif (obj_req->content_type)\n+\t\tfree(obj_req->content_type);\n \tfree(obj_req->url);\n \tfree(obj_req);\n }\n@@ -340,6 +349,7 @@ void prefetch(unsigned char *sha1)\n \tmemcpy(newreq->sha1, sha1, 20);\n \tnewreq->repo = alt;\n \tnewreq->url = NULL;\n+\tnewreq->content_type = NULL;\n \tnewreq->local = -1;\n \tnewreq->state = WAITING;\n \tsnprintf(newreq->filename, sizeof(newreq->filename), \"%s\", filename);\n@@ -836,7 +846,14 @@ static int fetch_object(struct alt_base \n \t\t\t\t    obj_req->http_code, hex);\n \t} else if (obj_req->zret != Z_STREAM_END) {\n \t\tcorrupt_object_found++;\n-\t\tret = error(\"File %s (%s) corrupt\", hex, obj_req->url);\n+\t\tif (obj_req->content_type &&\n+\t\t    !strcmp(obj_req->content_type, \"text/html\")) {\n+\t\t\tret = error(\"text/html response for file %s (%s)\",\n+\t\t\t\t    sha1_to_hex(obj_req->sha1), obj_req->url);\n+\t\t} else {\n+\t\t\tret = error(\"File %s (%s) corrupt\",\n+\t\t\t\t    sha1_to_hex(obj_req->sha1), obj_req->url);\n+\t\t}\n \t} else if (memcmp(obj_req->sha1, obj_req->real_sha1, 20)) {\n \t\tret = error(\"File %s has bad hash\", hex);\n \t} else if (obj_req->rename < 0) {\n-- \n1.2.4.gb1bc1d-dirty\n"},{"id":"17772","messageId":"7vslpa8fld.fsf@assigned-by-dhcp.cox.net","threadId":"3692","inReplyTo":"20060322183621.GP3997@reactrix.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-22T19:05:50Z","receivedAt":"2006-03-22T19:05:50Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nick Hengeveld <nickh@reactrix.com> writes:\n\n> Some HTTP server environments return a 200 status and text/html error\n> document or a redirect to one rather than a 404 status if a loose\n> object does not exist.  This patch detects and reports this condition\n> to differentiate between a misconfigured server and an actual corrupt\n> object on the server.\n\n> 61069cc348640fef2b8c503b8b8f00f689872cab\n> diff --git a/http-fetch.c b/http-fetch.c\n> index dc67218..ee5b585 100644\n> --- a/http-fetch.c\n> +++ b/http-fetch.c\n> @@ -41,6 +41,7 @@ struct object_request\n>  \tCURLcode curl_result;\n>...\n> +\tchar *content_type;\n>  \tunsigned char real_sha1[20];\n>...\n\nYou probably need only one bit here,...\n\n> @@ -258,9 +259,15 @@ static void finish_object_request(struct\n>  \n>  static void process_object_response(void *callback_data)\n>...  \n> +\tcurl_easy_getinfo(obj_req->slot->curl, CURLINFO_CONTENT_TYPE,\n> +\t\t\t  &content_type);\n> +\tif (content_type)\n> +\t\tobj_req->content_type = strdup(content_type);\n> +\n\n... and note if that is an HTML document or not.\n\nWe do bend backwards to support ISP HTTP servers, but this might\nbe going a bit too far.  Also I wonder if ISP runs a really\ndumb-friendly configured server that defaults to text/html\nunless the mimemap says otherwise.  Loose object files do not\nhave suffixes and I am expecting these servers would give\nwhatever the server default is.\n"},{"id":"17776","messageId":"7vacbi8eu1.fsf@assigned-by-dhcp.cox.net","threadId":"3692","inReplyTo":"7vslpa8fld.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cloning from sites with 404 overridden","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-22T19:22:14Z","receivedAt":"2006-03-22T19:22:14Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <junkio@cox.net> writes:\n\n> We do bend backwards to support ISP HTTP servers, but this might\n> be going a bit too far.  Also I wonder if ISP runs a really\n> dumb-friendly configured server that defaults to text/html\n> unless the mimemap says otherwise.  Loose object files do not\n> have suffixes and I am expecting these servers would give\n> whatever the server default is.\n\nClarification.  Even if a server configured as such existed and\nsent an otherwise valid loose object with text/html, your code\ndoes the right thing.\n\nHowever the patch would not help when such a server also did a\n\"Sorry, did you mistype the URL?\" HTML response, and I was\nwondering how typical that would be.\n"},{"id":"17783","messageId":"200603222224.08061.astralstorm@o2.pl","threadId":"3692","inReplyTo":"7vslpa8fld.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cloning from sites with 404 overridden","fromName":"Radoslaw Szkodzinski","fromEmail":"astralstorm@o2.pl","sentAt":"2006-03-22T21:24:02Z","receivedAt":"2006-03-22T21:24:02Z","isPatch":false,"sender":{"key":"astralstorm@o2.pl","avatar":null},"body":"On Wednesday 22 March 2006 20:05, Junio C Hamano wrote yet:\n>\n> .. and note if that is an HTML document or not.\n>\n\nBetter yet, see first if the object is corrupt. If it is and its Content-Type \nis text/html, error out.\n\n> We do bend backwards to support ISP HTTP servers, but this might\n> be going a bit too far.  Also I wonder if ISP runs a really\n> dumb-friendly configured server that defaults to text/html\n> unless the mimemap says otherwise.  Loose object files do not\n> have suffixes and I am expecting these servers would give\n> whatever the server default is.\n\nThat server would break a *lot* of file types. That admin should be hanged, \nshot, then burned.\n\nI think of only one reason for doing that: to restrict file types posted on \nthe server to, say, zip and html.\n\n-- \nGPG Key id:  0xD1F10BA2\nFingerprint: 96E2 304A B9C4 949A 10A0  9105 9543 0453 D1F1 0BA2\n\nAstralStorm\n"},{"id":"17817","messageId":"20060323184351.GA3892@reactrix.com","threadId":"3692","inReplyTo":"7vacbi8eu1.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cloning from sites with 404 overridden","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2006-03-23T18:43:51Z","receivedAt":"2006-03-23T18:43:51Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Wed, Mar 22, 2006 at 11:22:14AM -0800, Junio C Hamano wrote:\n\n> You probably need only one bit here,...\n> ... and note if that is an HTML document or not.\n\n/me smacks self...\n\n> However the patch would not help when such a server also did a\n> \"Sorry, did you mistype the URL?\" HTML response, and I was\n> wondering how typical that would be.\n\nSeems like there are three cases to worry about:\n\n1) the server returns a 200 status and a text/html response instead of a\n   404, and the server's default content type is not text/html\n2) the server returns a 200 status and a text/html response instead of a\n   404, and the server's default content type is text/html\n3) the server returns a corrupt object from the repository\n\nI don't think there's a way to distinguish between #2 and #3, so all we\ncan really do is display as helpful an error message as possible.\n\nWe can detect #1 if there has been a previous successful loose object\ntransfer by tracking whether the repo's default content type is\ntext/html.  In such a case should http-fetch behave as if the server\nreturned 404?  If there have been no successful loose object transfers,\nwe'd have to respond as with #2.  This approach could potentially break\nif requests are load-balanced to servers with different\nmisconfigurations - but I think trying to detect that is bending\nbackwards a little too far.\n\nOn a related note, I noticed that http-fetch will continue to try\ninflating/sha1_updating the response after an inflate error has been\ndetected.  It's probably not a huge deal, but we could just error out\nimmediately at that point or at least stop the unnecessary processing.\n\nSomething like this?  Tested by cloning\nhttp://digilander.libero.it/mcostalba/scm/qgit.git\n\n\n[PATCH] http-fetch: try to detect 404s from misconfigured servers\n\nSome HTTP server environments return a 200 status and text/html error\ndocument or a redirect to one rather than a 404 status if a loose\nobject does not exist.  This patch tries to detect such a response\nand treat it as a 404.\n\nSigned-off-by: Nick Hengeveld <nickh@reactrix.com>\n\n\n---\n\n http-fetch.c |   24 ++++++++++++++++++++++--\n 1 files changed, 22 insertions(+), 2 deletions(-)\n\nab97429c5b0a4b4466ee0072f75706399e42b675\ndiff --git a/http-fetch.c b/http-fetch.c\nindex dc67218..bb75050 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -16,6 +16,7 @@ struct alt_base\n {\n \tchar *base;\n \tint got_indices;\n+\tint default_html_content_type;\n \tstruct packed_git *packs;\n \tstruct alt_base *next;\n };\n@@ -41,6 +42,7 @@ struct object_request\n \tCURLcode curl_result;\n \tchar errorstr[CURL_ERROR_SIZE];\n \tlong http_code;\n+\tchar html_content_type;\n \tunsigned char real_sha1[20];\n \tSHA_CTX c;\n \tz_stream stream;\n@@ -249,6 +251,9 @@ static void finish_object_request(struct\n \t\tunlink(obj_req->tmpfile);\n \t\treturn;\n \t}\n+\tif (obj_req->repo->default_html_content_type == -1)\n+\t\tobj_req->repo->default_html_content_type =\n+\t\t\tobj_req->html_content_type;\n \tobj_req->rename =\n \t\tmove_temp_to_file(obj_req->tmpfile, obj_req->filename);\n \n@@ -258,9 +263,15 @@ static void finish_object_request(struct\n \n static void process_object_response(void *callback_data)\n {\n+\tchar *content_type;\n \tstruct object_request *obj_req =\n \t\t(struct object_request *)callback_data;\n \n+\tcurl_easy_getinfo(obj_req->slot->curl, CURLINFO_CONTENT_TYPE,\n+\t\t\t  &content_type);\n+\tif (content_type && !strcmp(content_type, \"text/html\"))\n+\t\tobj_req->html_content_type = 1;\n+\n \tobj_req->curl_result = obj_req->slot->curl_result;\n \tobj_req->http_code = obj_req->slot->http_code;\n \tobj_req->slot = NULL;\n@@ -340,6 +351,7 @@ void prefetch(unsigned char *sha1)\n \tmemcpy(newreq->sha1, sha1, 20);\n \tnewreq->repo = alt;\n \tnewreq->url = NULL;\n+\tnewreq->html_content_type = 0;\n \tnewreq->local = -1;\n \tnewreq->state = WAITING;\n \tsnprintf(newreq->filename, sizeof(newreq->filename), \"%s\", filename);\n@@ -539,6 +551,7 @@ static void process_alternates_response(\n \t\t\t\tnewalt->next = NULL;\n \t\t\t\tnewalt->base = target;\n \t\t\t\tnewalt->got_indices = 0;\n+\t\t\t\tnewalt->default_html_content_type = -1;\n \t\t\t\tnewalt->packs = NULL;\n \t\t\t\twhile (tail->next != NULL)\n \t\t\t\t\ttail = tail->next;\n@@ -835,8 +848,14 @@ static int fetch_object(struct alt_base \n \t\t\t\t    obj_req->errorstr, obj_req->curl_result,\n \t\t\t\t    obj_req->http_code, hex);\n \t} else if (obj_req->zret != Z_STREAM_END) {\n-\t\tcorrupt_object_found++;\n-\t\tret = error(\"File %s (%s) corrupt\", hex, obj_req->url);\n+\t\tif (obj_req->html_content_type &&\n+\t\t    !obj_req->repo->default_html_content_type)\n+\t\t\tret = -1; /* Be silent, looks like a 404 */\n+\t\telse {\n+\t\t\tcorrupt_object_found++;\n+\t\t\tret = error(\"File %s (%s) corrupt\",\n+\t\t\t\t    sha1_to_hex(obj_req->sha1), obj_req->url);\n+\t\t}\n \t} else if (memcmp(obj_req->sha1, obj_req->real_sha1, 20)) {\n \t\tret = error(\"File %s has bad hash\", hex);\n \t} else if (obj_req->rename < 0) {\n@@ -985,6 +1004,7 @@ int main(int argc, char **argv)\n \talt = xmalloc(sizeof(*alt));\n \talt->base = url;\n \talt->got_indices = 0;\n+\talt->default_html_content_type = -1;\n \talt->packs = NULL;\n \talt->next = NULL;\n \n-- \n1.2.4.gb1bc1d-dirty\n"},{"id":"17821","messageId":"7vwtek51r4.fsf@assigned-by-dhcp.cox.net","threadId":"3692","inReplyTo":"20060323184351.GA3892@reactrix.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-23T20:45:19Z","receivedAt":"2006-03-23T20:45:19Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nick Hengeveld <nickh@reactrix.com> writes:\n\n> Seems like there are three cases to worry about:\n>\n> 1) the server returns a 200 status and a text/html response instead of a\n>    404, and the server's default content type is not text/html\n> 2) the server returns a 200 status and a text/html response instead of a\n>    404, and the server's default content type is text/html\n> 3) the server returns a corrupt object from the repository\n\n> I don't think there's a way to distinguish between #2 and #3, so all we\n> can really do is display as helpful an error message as possible.\n\nThe code behaves correctly the same way whether the server says\n404 or 200 with human readable \"No such object\", and this is\njust for formatting error messages, and to be honest I do not\nreally care at this point.  I think the existing error message\nat the end of transfer we added recently should be sufficient.\n\n> On a related note, I noticed that http-fetch will continue to try\n> inflating/sha1_updating the response after an inflate error has been\n> detected.  It's probably not a huge deal, but we could just error out\n> immediately at that point or at least stop the unnecessary processing.\n\nThat would probably be more helpful.\n"},{"id":"17881","messageId":"slrne28b3a.cp6.mdw@metalzone.distorted.org.uk","threadId":"3692","inReplyTo":"442152E0.4020604@op5.se","subject":"Re: Cloning from sites with 404 overridden","fromName":"Mark Wooding","fromEmail":"mdw@distorted.org.uk","sentAt":"2006-03-24T17:29:14Z","receivedAt":"2006-03-24T17:29:14Z","isPatch":false,"sender":{"key":"mdw@distorted.org.uk","avatar":null},"body":"Andreas Ericsson <ae@op5.se> wrote:\n\n>> I'm not sure what the best \"object that's not supposed to be there\" is.\n>\n> .git/objects/00/hoping-for-a-404-or-webadmin-should-fix\n\nIf .git/objects/00/00000000000000000000000000000000000000 exists, the\nrepository has big problems already.\n\n(Aside: `C-u 38 0' doesn't work because Emacs hears `C-u 380' and waits\nfor a key.  `M-: (insert-char ?0 38) RET' does the right thing, but is\nugly.  Any better suggestions?)\n\n-- [mdw]\n"},{"id":"17882","messageId":"7vu09nzq50.fsf@assigned-by-dhcp.cox.net","threadId":"3692","inReplyTo":"slrne28b3a.cp6.mdw@metalzone.distorted.org.uk","subject":"Re: Cloning from sites with 404 overridden","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-24T17:52:43Z","receivedAt":"2006-03-24T17:52:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Mark Wooding <mdw@distorted.org.uk> writes:\n\n> (Aside: `C-u 38 0' doesn't work because Emacs hears `C-u 380' and waits\n> for a key.  `M-: (insert-char ?0 38) RET' does the right thing, but is\n> ugly.  Any better suggestions?)\n\nC-u 38 C-u 0\n"},{"id":"17883","messageId":"Pine.LNX.4.64.0603240945470.26286@g5.osdl.org","threadId":"3692","inReplyTo":"slrne28b3a.cp6.mdw@metalzone.distorted.org.uk","subject":"Re: Cloning from sites with 404 overridden","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-03-24T17:53:59Z","receivedAt":"2006-03-24T17:53:59Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nOn Fri, 24 Mar 2006, Mark Wooding wrote:\n> \n> (Aside: `C-u 38 0' doesn't work because Emacs hears `C-u 380' and waits\n> for a key.  `M-: (insert-char ?0 38) RET' does the right thing, but is\n> ugly.  Any better suggestions?)\n\nI don't do GNU emacs, but the way to do it in some other editors that do\nrepeats somewhat similarly is to do the action that starts with a number\nas a macro, and do that macro 37 more times. \n\nOn uemacs: ^X '(' '0' ^X ')' ESC '3' '7' ^X 'E'\n\n(Of course, the easier way is to just do '0' LEFT ^K to put the 0 in the\nbuffer, and than ESC '3' '8' ^Y to yank it 38 times, but the macro trick\nis generic, even if it's a few more keystrokes). \n\n\t\tLinus \"teaching people the one true editor\" Torvalds\n"},{"id":"17884","messageId":"118833cc0603241016h46124f8cr9634ee9f121f2f15@mail.gmail.com","threadId":"3692","inReplyTo":"slrne28b3a.cp6.mdw@metalzone.distorted.org.uk","subject":"Re: Cloning from sites with 404 overridden","fromName":"Morten Welinder","fromEmail":"mwelinder@gmail.com","sentAt":"2006-03-24T18:16:19Z","receivedAt":"2006-03-24T18:16:19Z","isPatch":false,"sender":{"key":"mwelinder@gmail.com","avatar":null},"body":"> (Aside: `C-u 38 0' doesn't work because Emacs hears `C-u 380' and waits\n> for a key.  `M-: (insert-char ?0 38) RET' does the right thing, but is\n> ugly.  Any better suggestions?)\n\nThere's a million ways to skin that cat.\n\nESC 38 C-q 60 RET\n\n[Octal 060 == '0']\n\nM.\n"},{"id":"17887","messageId":"44243D20.9060309@op5.se","threadId":"3692","inReplyTo":"slrne28b3a.cp6.mdw@metalzone.distorted.org.uk","subject":"Re: Cloning from sites with 404 overridden","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2006-03-24T18:40:32Z","receivedAt":"2006-03-24T18:40:32Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Mark Wooding wrote:\n> Andreas Ericsson <ae@op5.se> wrote:\n> \n> \n>>>I'm not sure what the best \"object that's not supposed to be there\" is.\n>>\n>>.git/objects/00/hoping-for-a-404-or-webadmin-should-fix\n> \n> \n> If .git/objects/00/00000000000000000000000000000000000000 exists, the\n> repository has big problems already.\n> \n\nIndeed. I'm off sobriety again, it being friday and all, but I'm \nassuming there are more than 18 zeroes there, yes? The \"feature\" of the \nabove line is that it will fit in any buffer that already exists, and \nwill match any third argument to send(2) that already exists.\n\n\n> (Aside: `C-u 38 0' doesn't work because Emacs hears `C-u 380' and waits\n> for a key.  `M-: (insert-char ?0 38) RET' does the right thing, but is\n> ugly.  Any better suggestions?)\n> \n\nThis I happily don't understand at all. I'm also happy ignorant of what \nit has to do with the issue at hand.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"}]}