{"thread":{"id":"3676","subject":"Cloning from sites with 404 overridden","startedAt":"2006-03-19T10:52:05Z","lastAt":"2006-03-20T19:54:09Z","messageCount":17,"participants":["Marco Costalba","Paolo Ciarrocchi","Junio C Hamano","Petr Baudis","Randal L. Schwartz","Lukas Sandström","Nick Hengeveld"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"17666","messageId":"e5bfff550603190252n7e3e1cbbp94e3f15c92f12d07@mail.gmail.com","threadId":"3676","inReplyTo":null,"subject":"Cloning from sites with 404 overridden","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-03-19T10:52:05Z","receivedAt":"2006-03-19T10:52:05Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"Hi all,\n\n    I have set a git repository on a hosted public site:\nhttp://digilander.libero.it/mcostalba/scm/qgit.git\n\nI cannot run any process (read git-daemon) on that site, so git-clone uses\na 'dumb server' type protocol and this is what I got.\n\n$ git clone http://digilander.libero.it/mcostalba/scm/qgit.git\nerror: File 8dea03519e75f47da91108330dde3043defddd60\n(http://digilander.libero.it/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60)\ncorrupt\nGetting pack list for http://digilander.libero.it/mcostalba/scm/qgit.git/\nGetting index for pack fe1f3586b38e70e963de47f31379ef170adc5ca9\nGetting pack fe1f3586b38e70e963de47f31379ef170adc5ca9\n which contains 8dea03519e75f47da91108330dde3043defddd60\nwalk 8dea03519e75f47da91108330dde3043defddd60\nwalk ec47dab590fb838ba2be7af5bf9aa46d9f2e502d\n\n-------------- cut ------------------------\n\nwalk 907d47e836f4f174386d02d21e38aeafc1e79626\nwalk 5d3454248bbb3aaba080057dc9666a3c3aaeca1f\n$\n\nThe above mentioned error belongs to git requests a non existing object\n(8dea03519e75f47da91108330dde3043defddd60) _and_  the site answers with\na pre-canned 'page not found' html page instead of reporting 404 error.\n\nAfter some research I found it is quite common for public hosting\nsites to use a pre-canned\n'Sorry, no page here' html stuff instead of 404.\n\nSo my request is if it is possible for git to _learn_ this and to\navoid been fooled by\nthese kind of public sites.\n\nThanks\nMarco\n"},{"id":"17667","messageId":"4d8e3fd30603190525o5a01fba8w5bcdedd064c213ec@mail.gmail.com","threadId":"3676","inReplyTo":"e5bfff550603190252n7e3e1cbbp94e3f15c92f12d07@mail.gmail.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Paolo Ciarrocchi","fromEmail":"paolo.ciarrocchi@gmail.com","sentAt":"2006-03-19T13:25:21Z","receivedAt":"2006-03-19T13:25:21Z","isPatch":false,"sender":{"key":"paolo.ciarrocchi@gmail.com","avatar":null},"body":"On 3/19/06, Marco Costalba <mcostalba@gmail.com> wrote:\n> Hi all,\n\nCiao Marco,\n\n>     I have set a git repository on a hosted public site:\n> http://digilander.libero.it/mcostalba/scm/qgit.git\n>\n> I cannot run any process (read git-daemon) on that site, so git-clone uses\n> a 'dumb server' type protocol and this is what I got.\n>\n> $ git clone http://digilander.libero.it/mcostalba/scm/qgit.git\n> error: File 8dea03519e75f47da91108330dde3043defddd60\n> (http://digilander.libero.it/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60)\n> corrupt\n> Getting pack list for http://digilander.libero.it/mcostalba/scm/qgit.git/\n> Getting index for pack fe1f3586b38e70e963de47f31379ef170adc5ca9\n> Getting pack fe1f3586b38e70e963de47f31379ef170adc5ca9\n>  which contains 8dea03519e75f47da91108330dde3043defddd60\n> walk 8dea03519e75f47da91108330dde3043defddd60\n> walk ec47dab590fb838ba2be7af5bf9aa46d9f2e502d\n>\n> -------------- cut ------------------------\n>\n> walk 907d47e836f4f174386d02d21e38aeafc1e79626\n> walk 5d3454248bbb3aaba080057dc9666a3c3aaeca1f\n> $\n>\n> The above mentioned error belongs to git requests a non existing object\n> (8dea03519e75f47da91108330dde3043defddd60) _and_  the site answers with\n> a pre-canned 'page not found' html page instead of reporting 404 error.\n>\n> After some research I found it is quite common for public hosting\n> sites to use a pre-canned\n> 'Sorry, no page here' html stuff instead of 404.\n>\n> So my request is if it is possible for git to _learn_ this and to\n> avoid been fooled by\n> these kind of public sites.\n>\n\nHow about getting an account on kernel.org?\n\nAnyway, here is what I did:\npaolo@Italia:~$ cg-clone\nhttp://digilander.libero.it/mcostalba/scm/qgit.git qgit defaulting to\nlocal storage area\nFetching head...\nFetching objects...\nerror: File 8dea03519e75f47da91108330dde3043defddd60\n(http://digilander.libero.i\nt/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60)\ncorr upt\n\nGetting pack list for http://digilander.libero.it/mcostalba/scm/qgit.git/\nGetting index for pack fe1f3586b38e70e963de47f31379ef170adc5ca9\nGetting pack fe1f3586b38e70e963de47f31379ef170adc5ca9\n which contains 8dea03519e75f47da91108330dde3043defddd60\nFetching tags...\nMissing tag qgit-0.93... retrieved\nMissing tag qgit-0.94... retrieved\nMissing tag qgit-0.94.1... retrieved\nMissing tag qgit-0.95.1... retrieved\nMissing tag qgit-0.96... retrieved\nMissing tag qgit-0.96.1... retrieved\nMissing tag qgit-0.97... retrieved\nMissing tag qgit-0.97.1... retrieved\nMissing tag qgit-0.97.2... retrieved\nMissing tag qgit-1.0... retrieved\nMissing tag qgit-1.1rc1... retrieved\nMissing tag qgit-1.1rc3... retrieved\nNew branch: 8dea03519e75f47da91108330dde3043defddd60\nCloned to qgit/ (origin\nhttp://digilander.libero.it/mcostalba/scm/qgit.git available as branch\n\"origin\")\n\n\nWhy am I getting this error?\nerror: File 8dea03519e75f47da91108330dde3043defddd60\n(http://digilander.libero.i\nt/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60)\ncorr upt\n\n\n--\nPaolo\nhttp://paolociarrocchi.googlepages.com\n"},{"id":"17668","messageId":"e5bfff550603190604ne4364f3o6a862d25267a2dce@mail.gmail.com","threadId":"3676","inReplyTo":"4d8e3fd30603190525o5a01fba8w5bcdedd064c213ec@mail.gmail.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-03-19T14:04:43Z","receivedAt":"2006-03-19T14:04:43Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/19/06, Paolo Ciarrocchi <paolo.ciarrocchi@gmail.com> wrote:\n> On 3/19/06, Marco Costalba <mcostalba@gmail.com> wrote:\n> >\n>\n> How about getting an account on kernel.org?\n>\n\nI don't think I have the credentials to ask for ;-)\n\n> Anyway, here is what I did:\n> paolo@Italia:~$ cg-clone\n> http://digilander.libero.it/mcostalba/scm/qgit.git qgit defaulting to\n>\n> Why am I getting this error?\n> error: File 8dea03519e75f47da91108330dde3043defddd60\n> (http://digilander.libero.i\n> t/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60)\n> corr upt\n>\n\nBecause http server of digilander.libero.it instead of responding with\n404 code (page not\nfound) sends a not standard html page as answer. To see the page just point\nyour browser to:\nhttp://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d\n\nGit does not understand object is missing and thinks what site sends\n_is_ the requested\nobject and then founds that is (of course) corrupted.\n\n\nMarco\n"},{"id":"17672","messageId":"7vk6aqql9e.fsf@assigned-by-dhcp.cox.net","threadId":"3676","inReplyTo":"e5bfff550603190604ne4364f3o6a862d25267a2dce@mail.gmail.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-19T19:37:01Z","receivedAt":"2006-03-19T19:37:01Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Marco Costalba\" <mcostalba@gmail.com> writes:\n\n> http://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d\n>\n> Git does not understand object is missing and thinks what site sends\n> _is_ the requested\n> object and then founds that is (of course) corrupted.\n\nTo be fair, the site is _not_ missing anything from HTTP\nprotocol perspective, because when git asks 8d/ea0351... file,\nthe server responds with a regular \"HTTP/1.0 200 OK\" response.\nSo it is _your_ repository that is corrupt -- instead of\ncorrectly _lacking_ the file you should have removed with\nprune-packed, it has a garbage file.\n\nHaving said that, I agree that it would be nicer if we support\nsuch a site, in the same spirit that we already bend backwards\nto support really dumb hosted http servers that do not give\ndirectory index by using objects/info/packs and info/refs.\n\nI think it wouldn't be too much a hassle to add logic to\nhttp-fetch.c (perhaps with an additional \"--no-404\" option or\nsomesuch) to fall back on pack transfer upon seeing a corrupt\nloose object.  We do the falling back when getting 404 error to\na request for a loose object, so the new code would essentially\ndo the same and you might be OK.\n"},{"id":"17674","messageId":"7v7j6qqks6.fsf@assigned-by-dhcp.cox.net","threadId":"3676","inReplyTo":"e5bfff550603190604ne4364f3o6a862d25267a2dce@mail.gmail.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-19T19:47:21Z","receivedAt":"2006-03-19T19:47:21Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Marco Costalba\" <mcostalba@gmail.com> writes:\n\n> On 3/19/06, Paolo Ciarrocchi <paolo.ciarrocchi@gmail.com> wrote:\n>>\n>> How about getting an account on kernel.org?\n>\n> I don't think I have the credentials to ask for ;-)\n\nHeh, it has a striking resemblance to the first thing I said\nwhen Linus asked me if I want to take over git.git: \"It would\nbe embarrassing to be the first person to have an account there\nwithout having a single line of code in the kernel\" ;-).\n\nWell, you won't be the first (in fact it appears I wasn't\neither), and it would never hurt to ask.\n"},{"id":"17676","messageId":"20060319213125.GE18185@pasky.or.cz","threadId":"3676","inReplyTo":"7v7j6qqks6.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cloning from sites with 404 overridden","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-03-19T21:31:25Z","receivedAt":"2006-03-19T21:31:25Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sun, Mar 19, 2006 at 08:47:21PM CET, I got a letter\nwhere Junio C Hamano <junkio@cox.net> said that...\n> \"Marco Costalba\" <mcostalba@gmail.com> writes:\n> \n> > On 3/19/06, Paolo Ciarrocchi <paolo.ciarrocchi@gmail.com> wrote:\n> >>\n> >> How about getting an account on kernel.org?\n> >\n> > I don't think I have the credentials to ask for ;-)\n> \n> Heh, it has a striking resemblance to the first thing I said\n> when Linus asked me if I want to take over git.git: \"It would\n> be embarrassing to be the first person to have an account there\n> without having a single line of code in the kernel\" ;-).\n> \n> Well, you won't be the first (in fact it appears I wasn't\n> either), and it would never hurt to ask.\n\nYeah, I think I was there before you... ;-)\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"17679","messageId":"e5bfff550603191340u466d3551t8a95c3808eb977c1@mail.gmail.com","threadId":"3676","inReplyTo":"7vk6aqql9e.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cloning from sites with 404 overridden","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-03-19T21:40:41Z","receivedAt":"2006-03-19T21:40:41Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/19/06, Junio C Hamano <junkio@cox.net> wrote:\n> \"Marco Costalba\" <mcostalba@gmail.com> writes:\n>\n> > http://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d\n> >\n> > Git does not understand object is missing and thinks what site sends\n> > _is_ the requested\n> > object and then founds that is (of course) corrupted.\n>\n> To be fair, the site is _not_ missing anything from HTTP\n> protocol perspective, because when git asks 8d/ea0351... file,\n> the server responds with a regular \"HTTP/1.0 200 OK\" response.\n> So it is _your_ repository that is corrupt -- instead of\n> correctly _lacking_ the file you should have removed with\n> prune-packed, it has a garbage file.\n>\n\nCurrently my git repo layout is as follow\n$ pwd\n<local master copy>/qgit.git/.git\n$ ls\nbranches/  description  HEAD    index  objects/   refs/\nconfig     FETCH_HEAD   hooks/  info/  ORIG_HEAD  remotes/\n$ ls objects\n2c/  32/  53/  5c/  6a/  info/  pack/\n\nThe host copy should be the exact mirror of the local copy (I use\nsitecopy to sync\nhost). I have also verified this directly accessing the host with ftp.\n\nSo the 8d/ea0351... file is really not existent. BTW I have run git\nprune and git-prune-packed\nalso.\n\nFinally accessing the missing object with a browser\n\nhttp://digilander.libero.it/mcostalba/\nscm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60\n\ngives a pre-canned (in italian) 'Sorry page not found' stuff.\n\nSo I really think the site \"HTTP/1.0 200 OK\" response it's a fake.\nPerhaps security related to avoid sniffing (just a guess because I have\nabsolutely zero competence in security related things).\n\n\nMarco\n"},{"id":"17680","messageId":"20060319214307.GG18185@pasky.or.cz","threadId":"3676","inReplyTo":"20060319213125.GE18185@pasky.or.cz","subject":"Re: Cloning from sites with 404 overridden","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-03-19T21:43:07Z","receivedAt":"2006-03-19T21:43:07Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Sun, Mar 19, 2006 at 10:31:25PM CET, I got a letter\nwhere Petr Baudis <pasky@suse.cz> said that...\n> Dear diary, on Sun, Mar 19, 2006 at 08:47:21PM CET, I got a letter\n> where Junio C Hamano <junkio@cox.net> said that...\n> > Heh, it has a striking resemblance to the first thing I said\n> > when Linus asked me if I want to take over git.git: \"It would\n> > be embarrassing to be the first person to have an account there\n> > without having a single line of code in the kernel\" ;-).\n> > \n> > Well, you won't be the first (in fact it appears I wasn't\n> > either), and it would never hurt to ask.\n> \n> Yeah, I think I was there before you... ;-)\n\nSilly me, on a second thought I've realized that I already had some\nstuff in the kernel by then. Sorry for the noise.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"17681","messageId":"e5bfff550603191345m5c784604yb3e63ab7f3ae5efd@mail.gmail.com","threadId":"3676","inReplyTo":"20060319213125.GE18185@pasky.or.cz","subject":"Re: Cloning from sites with 404 overridden","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-03-19T21:45:24Z","receivedAt":"2006-03-19T21:45:24Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/19/06, Petr Baudis <pasky@suse.cz> wrote:\n> Dear diary, on Sun, Mar 19, 2006 at 08:47:21PM CET, I got a letter\n> where Junio C Hamano <junkio@cox.net> said that...\n> > \"Marco Costalba\" <mcostalba@gmail.com> writes:\n> >\n> > > On 3/19/06, Paolo Ciarrocchi <paolo.ciarrocchi@gmail.com> wrote:\n> > >>\n> > >> How about getting an account on kernel.org?\n> > >\n> > > I don't think I have the credentials to ask for ;-)\n> >\n> > Heh, it has a striking resemblance to the first thing I said\n> > when Linus asked me if I want to take over git.git: \"It would\n> > be embarrassing to be the first person to have an account there\n> > without having a single line of code in the kernel\" ;-).\n> >\n> > Well, you won't be the first (in fact it appears I wasn't\n> > either), and it would never hurt to ask.\n>\n> Yeah, I think I was there before you... ;-)\n>\n> --\n\nPlease could someone tell me what door I should knock at?\n\nThanks\nMarco\n"},{"id":"17688","messageId":"7vmzfmm35t.fsf@assigned-by-dhcp.cox.net","threadId":"3676","inReplyTo":"e5bfff550603191340u466d3551t8a95c3808eb977c1@mail.gmail.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-19T23:21:34Z","receivedAt":"2006-03-19T23:21:34Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Marco Costalba\" <mcostalba@gmail.com> writes:\n\n> Finally accessing the missing object with a browser\n>\n> http://digilander.libero.it/mcostalba/\n> scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60\n>\n> gives a pre-canned (in italian) 'Sorry page not found' stuff.\n>\n> So I really think the site \"HTTP/1.0 200 OK\" response it's a fake.\n> Perhaps security related to avoid sniffing (just a guess because I have\n> absolutely zero competence in security related things).\n\nI think you are just rephrasing what I said.  From the HTTP\nprotocol perspective, you _do_ have that 8d/3a0351 thing on that\nserver, because you do not correctly say \"No we donot have it\"\nusing 404 response.\n\nYour inability to produce 404 is a different matter -- often the\nhosting server is not under your control.  But that does not\nchange the fact that the repository observed by your clients is\n\"broken\".  That is why a workaround flag like I suggested may be\nneeded for such a setup.\n\nThis is totally untested, but maybe something like this?\n\n---\ndiff --git a/http-fetch.c b/http-fetch.c\nindex 7de818b..d523798 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -8,6 +8,7 @@\n #define RANGE_HEADER_SIZE 30\n \n static int got_alternates = -1;\n+static int unreliable_404 = 0;\n \n static struct curl_slist *no_pragma_header;\n \n@@ -822,12 +823,18 @@ static int fetch_object(struct alt_base \n \t\tclose(obj_req->local); obj_req->local = -1;\n \t}\n \n+\t\n+\n \tif (obj_req->state == ABORTED) {\n \t\tret = error(\"Request for %s aborted\", hex);\n-\t} else if (obj_req->curl_result != CURLE_OK &&\n-\t\t   obj_req->http_code != 416) {\n+\t} else if ((obj_req->curl_result != CURLE_OK &&\n+\t\t    obj_req->http_code != 416)  ||\n+\t\t   (unreliable_404 &&\n+\t\t    obj_req->curl_result == CURLE_OK &&\n+\t\t    obj_req->zret != Z_STREAM_END)) {\n \t\tif (obj_req->http_code == 404 ||\n-\t\t    obj_req->curl_result == CURLE_FILE_COULDNT_READ_FILE)\n+\t\t    obj_req->curl_result == CURLE_FILE_COULDNT_READ_FILE ||\n+\t\t    unreliable_404)\n \t\t\tret = -1; /* Be silent, it is probably in a pack. */\n \t\telse\n \t\t\tret = error(\"%s (curl_result = %d, http_code = %ld, sha1 = %s)\",\n@@ -966,6 +973,8 @@ int main(int argc, char **argv)\n \t\t\targ++;\n \t\t} else if (!strcmp(argv[arg], \"--recover\")) {\n \t\t\tget_recover = 1;\n+\t\t} else if (!strcmp(argv[arg], \"--unreliable-404\")) {\n+\t\t\tunreliable_404 = 1;\n \t\t}\n \t\targ++;\n \t}\n"},{"id":"17694","messageId":"863bhdvir4.fsf@blue.stonehenge.com","threadId":"3676","inReplyTo":"7v7j6qqks6.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cloning from sites with 404 overridden","fromName":"Randal L. Schwartz","fromEmail":"merlyn@stonehenge.com","sentAt":"2006-03-20T04:32:15Z","receivedAt":"2006-03-20T04:32:15Z","isPatch":false,"sender":{"key":"merlyn@stonehenge.com","avatar":"https://gravatar.com/avatar/dc528d210743ff0333e6213f9ee7b33b23f1b7bc1f3c5a8c2d819074ecd7ab19?d=mp&s=160"},"body":">>>>> \"Junio\" == Junio C Hamano <junkio@cox.net> writes:\n\nJunio> Heh, it has a striking resemblance to the first thing I said\nJunio> when Linus asked me if I want to take over git.git: \"It would\nJunio> be embarrassing to be the first person to have an account there\nJunio> without having a single line of code in the kernel\" ;-).\n\nJunio> Well, you won't be the first (in fact it appears I wasn't\nJunio> either), and it would never hurt to ask.\n\nWow.  That would perhaps completely rule out people who have never owned\nanything that can execute the x86 instruction set except in emulation. :)\n\n-- \nRandal L. Schwartz - Stonehenge Consulting Services, Inc. - +1 503 777 0095\n<merlyn@stonehenge.com> <URL:http://www.stonehenge.com/merlyn/>\nPerl/Unix/security consulting, Technical writing, Comedy, etc. etc.\nSee PerlTraining.Stonehenge.com for onsite and open-enrollment Perl training!\n"},{"id":"17695","messageId":"e5bfff550603192231k7843a741xbf14394bc5e4c57@mail.gmail.com","threadId":"3676","inReplyTo":"7vmzfmm35t.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cloning from sites with 404 overridden","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-03-20T06:31:03Z","receivedAt":"2006-03-20T06:31:03Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/20/06, Junio C Hamano <junkio@cox.net> wrote:\n>\n> Your inability to produce 404 is a different matter -- often the\n> hosting server is not under your control.  But that does not\n> change the fact that the repository observed by your clients is\n> \"broken\".  That is why a workaround flag like I suggested may be\n> needed for such a setup.\n>\n> This is totally untested, but maybe something like this?\n>\n\nIt works for me. Just some trailing white space warning when applying.\n\nI didn't found a way to pass '--unreliable-404' flag from git-clone,\nperhaps my bad,\nI have tested forcing the flag in sources.\n\n\nMarco\n"},{"id":"17696","messageId":"7v8xr5ld38.fsf@assigned-by-dhcp.cox.net","threadId":"3676","inReplyTo":"e5bfff550603192231k7843a741xbf14394bc5e4c57@mail.gmail.com","subject":"Re: Cloning from sites with 404 overridden","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-03-20T08:44:43Z","receivedAt":"2006-03-20T08:44:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Marco Costalba\" <mcostalba@gmail.com> writes:\n\n>> This is totally untested, but maybe something like this?\n>\n> It works for me. Just some trailing white space warning when applying.\n\nThe change only removes the error message without changing any\nother logic, so if that works for you, I wonder if leaving\nthings as they are is a better option than doing anything short\nof implementing an AI that tries to pattern-match the \"allegedly\ncorrupt file\" with \"sorry no such page found\" in many natural\nlanguages.\n\nMy test patch makes it impossible to track down the real\nbreakage when an HTTP-reachable repository _does_ have a corrupt\nobject.\n\nSo how about doing this instead?\n\n-- >8 --\ndiff --git a/http-fetch.c b/http-fetch.c\nindex 8fd9de0..1405c1f 100644\n--- a/http-fetch.c\n+++ b/http-fetch.c\n@@ -8,6 +8,7 @@\n #define RANGE_HEADER_SIZE 30\n \n static int got_alternates = -1;\n+static int corrupt_object_found = 0;\n \n static struct curl_slist *no_pragma_header;\n \n@@ -830,6 +831,7 @@ static int fetch_object(struct alt_base \n \t\t\t\t    obj_req->errorstr, obj_req->curl_result,\n \t\t\t\t    obj_req->http_code, hex);\n \t} else if (obj_req->zret != Z_STREAM_END) {\n+\t\tcorrupt_object_found++;\n \t\tret = error(\"File %s (%s) corrupt\", hex, obj_req->url);\n \t} else if (memcmp(obj_req->sha1, obj_req->real_sha1, 20)) {\n \t\tret = error(\"File %s has bad hash\", hex);\n@@ -989,5 +991,11 @@ int main(int argc, char **argv)\n \n \thttp_cleanup();\n \n+\tif (corrupt_object_found) {\n+\t\tfprintf(stderr,\n+\"Some loose object were found to be corrupt, but they might be just\\n\"\n+\"a false '404 Not Found' error message sent with incorrect HTTP\\n\"\n+\"status code.  Suggest running git fsck-objects.\\n\");\n+\t}\n \treturn rc;\n }\n"},{"id":"17704","messageId":"e5bfff550603200417h71c083f9saebbf1fe6b21076c@mail.gmail.com","threadId":"3676","inReplyTo":"7v8xr5ld38.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cloning from sites with 404 overridden","fromName":"Marco Costalba","fromEmail":"mcostalba@gmail.com","sentAt":"2006-03-20T12:17:55Z","receivedAt":"2006-03-20T12:17:55Z","isPatch":false,"sender":{"key":"mcostalba@gmail.com","avatar":null},"body":"On 3/20/06, Junio C Hamano <junkio@cox.net> wrote:\n> \"Marco Costalba\" <mcostalba@gmail.com> writes:\n>\n> >> This is totally untested, but maybe something like this?\n> >\n> > It works for me. Just some trailing white space warning when applying.\n>\n> The change only removes the error message without changing any\n> other logic, so if that works for you, I wonder if leaving\n> things as they are is a better option than doing anything short\n> of implementing an AI that tries to pattern-match the \"allegedly\n> corrupt file\" with \"sorry no such page found\" in many natural\n> languages.\n>\n> My test patch makes it impossible to track down the real\n> breakage when an HTTP-reachable repository _does_ have a corrupt\n> object.\n>\n> So how about doing this instead?\n>\n> -- >8 --\n\n> +               fprintf(stderr,\n> +\"Some loose object were found to be corrupt, but they might be just\\n\"\n> +\"a false '404 Not Found' error message sent with incorrect HTTP\\n\"\n> +\"status code.  Suggest running git fsck-objects.\\n\");\n> +       }\n>         return rc;\n>  }\n>\n\nI think it's better, read more correct.\n\nCould be a real corrupted file or just a false 404, so better a\nwarning then an error message and also better a warning then nothing.\n\nMarco\n"},{"id":"17712","messageId":"441EF46E.5050907@etek.chalmers.se","threadId":"3676","inReplyTo":"7vk6aqql9e.fsf@assigned-by-dhcp.cox.net","subject":"Re: Cloning from sites with 404 overridden","fromName":"Lukas Sandström","fromEmail":"lukass@etek.chalmers.se","sentAt":"2006-03-20T18:29:02Z","receivedAt":"2006-03-20T18:29:02Z","isPatch":false,"sender":{"key":"luksan@gmail.com","avatar":"https://avatars.githubusercontent.com/u/152281?v=4"},"body":"Junio C Hamano wrote:\n> \"Marco Costalba\" <mcostalba@gmail.com> writes:\n>>http://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d\n> \n> To be fair, the site is _not_ missing anything from HTTP\n> protocol perspective, because when git asks 8d/ea0351... file,\n> the server responds with a regular \"HTTP/1.0 200 OK\" response.\n> So it is _your_ repository that is corrupt -- instead of\n> correctly _lacking_ the file you should have removed with\n> prune-packed, it has a garbage file.\n\nActually, it sends a 302 redirect. \n\nPerhaps a repository config option to treat a 302 as a 404?\n\n/Lukas Sandström\n"},{"id":"17715","messageId":"20060320194316.GO18185@pasky.or.cz","threadId":"3676","inReplyTo":"441EF46E.5050907@etek.chalmers.se","subject":"Re: Cloning from sites with 404 overridden","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-03-20T19:43:16Z","receivedAt":"2006-03-20T19:43:16Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Mon, Mar 20, 2006 at 07:29:02PM CET, I got a letter\nwhere Lukas Sandström <lukass@etek.chalmers.se> said that...\n> Junio C Hamano wrote:\n> > \"Marco Costalba\" <mcostalba@gmail.com> writes:\n> >>http://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d\n> > \n> > To be fair, the site is _not_ missing anything from HTTP\n> > protocol perspective, because when git asks 8d/ea0351... file,\n> > the server responds with a regular \"HTTP/1.0 200 OK\" response.\n> > So it is _your_ repository that is corrupt -- instead of\n> > correctly _lacking_ the file you should have removed with\n> > prune-packed, it has a garbage file.\n> \n> Actually, it sends a 302 redirect. \n> \n> Perhaps a repository config option to treat a 302 as a 404?\n\nI think that would be too ugly _and_ specific a workaround for the\nparticular site. It's reasonable to keep it generalized for all the\nbroken repositories when already doing it.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nRight now I am having amnesia and deja-vu at the same time.  I think\nI have forgotten this before.\n"},{"id":"17717","messageId":"20060320195409.GN3997@reactrix.com","threadId":"3676","inReplyTo":"441EF46E.5050907@etek.chalmers.se","subject":"Re: Cloning from sites with 404 overridden","fromName":"Nick Hengeveld","fromEmail":"nickh@reactrix.com","sentAt":"2006-03-20T19:54:09Z","receivedAt":"2006-03-20T19:54:09Z","isPatch":false,"sender":{"key":"nickh@reactrix.com","avatar":null},"body":"On Mon, Mar 20, 2006 at 07:29:02PM +0100, Lukas Sandström wrote:\n\n> Perhaps a repository config option to treat a 302 as a 404?\n\nFWIW, it used to work that way and was modified to follow redirects back at\ncommit 66c9ec25553ce7332c46e2017b9c4d7c26310fff.\n\n-- \nFor a successful technology, reality must take precedence over public\nrelations, for nature cannot be fooled.\n"}]}