threads / discuss / 3676

Cloning from sites with 404 overridden

Subject: Cloning from sites with 404 overridden

## tl;dr

17 messages between Mar 19, 2006 and Mar 20, 2006.

replies: 16people: 7as markdown or json

Marco Costalba· Mar 19, 2006, 10:52 UTC · lore
Hi all,
    I have set a git repository on a hosted public site:
http://digilander.libero.it/mcostalba/scm/qgit.git

I cannot run any process (read git-daemon) on that site, so git-clone uses a 'dumb server' type protocol and this is what I got.

$ git clone http://digilander.libero.it/mcostalba/scm/qgit.git
error: File 8dea03519e75f47da91108330dde3043defddd60
(http://digilander.libero.it/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60)
corrupt
Getting pack list for http://digilander.libero.it/mcostalba/scm/qgit.git/
Getting index for pack fe1f3586b38e70e963de47f31379ef170adc5ca9
Getting pack fe1f3586b38e70e963de47f31379ef170adc5ca9
 which contains 8dea03519e75f47da91108330dde3043defddd60
walk 8dea03519e75f47da91108330dde3043defddd60
walk ec47dab590fb838ba2be7af5bf9aa46d9f2e502d
-------------- cut ------------------------

walk 907d47e836f4f174386d02d21e38aeafc1e79626 walk 5d3454248bbb3aaba080057dc9666a3c3aaeca1f $

The above mentioned error belongs to git requests a non existing object (8dea03519e75f47da91108330dde3043defddd60) _and_ the site answers with a pre-canned 'page not found' html page instead of reporting 404 error.

After some research I found it is quite common for public hosting sites to use a pre-canned 'Sorry, no page here' html stuff instead of 404.

So my request is if it is possible for git to _learn_ this and to avoid been fooled by these kind of public sites.

Thanks Marco

Paolo Ciarrocchi· Mar 19, 2006, 13:25 UTC · re: Marco Costalba · lore

Re: Cloning from sites with 404 overridden

On 3/19/06, Marco Costalba <mcostalba@gmail.com> wrote:
> Hi all,
Ciao Marco,
Show 35 quoted lines
>     I have set a git repository on a hosted public site:
> http://digilander.libero.it/mcostalba/scm/qgit.git
>
> I cannot run any process (read git-daemon) on that site, so git-clone uses
> a 'dumb server' type protocol and this is what I got.
>
> $ git clone http://digilander.libero.it/mcostalba/scm/qgit.git
> error: File 8dea03519e75f47da91108330dde3043defddd60
> (http://digilander.libero.it/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60)
> corrupt
> Getting pack list for http://digilander.libero.it/mcostalba/scm/qgit.git/
> Getting index for pack fe1f3586b38e70e963de47f31379ef170adc5ca9
> Getting pack fe1f3586b38e70e963de47f31379ef170adc5ca9
>  which contains 8dea03519e75f47da91108330dde3043defddd60
> walk 8dea03519e75f47da91108330dde3043defddd60
> walk ec47dab590fb838ba2be7af5bf9aa46d9f2e502d
>
> -------------- cut ------------------------
>
> walk 907d47e836f4f174386d02d21e38aeafc1e79626
> walk 5d3454248bbb3aaba080057dc9666a3c3aaeca1f
> $
>
> The above mentioned error belongs to git requests a non existing object
> (8dea03519e75f47da91108330dde3043defddd60) _and_  the site answers with
> a pre-canned 'page not found' html page instead of reporting 404 error.
>
> After some research I found it is quite common for public hosting
> sites to use a pre-canned
> 'Sorry, no page here' html stuff instead of 404.
>
> So my request is if it is possible for git to _learn_ this and to
> avoid been fooled by
> these kind of public sites.
>
How about getting an account on kernel.org?

Anyway, here is what I did: paolo@Italia:~$ cg-clone http://digilander.libero.it/mcostalba/scm/qgit.git qgit defaulting to local storage area Fetching head... Fetching objects... error: File 8dea03519e75f47da91108330dde3043defddd60 (http://digilander.libero.i t/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60) corr upt

Getting pack list for http://digilander.libero.it/mcostalba/scm/qgit.git/
Getting index for pack fe1f3586b38e70e963de47f31379ef170adc5ca9
Getting pack fe1f3586b38e70e963de47f31379ef170adc5ca9
 which contains 8dea03519e75f47da91108330dde3043defddd60
Fetching tags...
Missing tag qgit-0.93... retrieved
Missing tag qgit-0.94... retrieved
Missing tag qgit-0.94.1... retrieved
Missing tag qgit-0.95.1... retrieved
Missing tag qgit-0.96... retrieved
Missing tag qgit-0.96.1... retrieved
Missing tag qgit-0.97... retrieved
Missing tag qgit-0.97.1... retrieved
Missing tag qgit-0.97.2... retrieved
Missing tag qgit-1.0... retrieved
Missing tag qgit-1.1rc1... retrieved
Missing tag qgit-1.1rc3... retrieved
New branch: 8dea03519e75f47da91108330dde3043defddd60
Cloned to qgit/ (origin
http://digilander.libero.it/mcostalba/scm/qgit.git available as branch
"origin")

Why am I getting this error? error: File 8dea03519e75f47da91108330dde3043defddd60 (http://digilander.libero.i t/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60) corr upt

-- Paolo http://paolociarrocchi.googlepages.com

Marco Costalba· Mar 19, 2006, 14:04 UTC · re: Paolo Ciarrocchi · lore

Re: Cloning from sites with 404 overridden

On 3/19/06, Paolo Ciarrocchi <paolo.ciarrocchi@gmail.com> wrote:
Show 5 quoted lines
> On 3/19/06, Marco Costalba <mcostalba@gmail.com> wrote:
> >
>
> How about getting an account on kernel.org?
>
I don't think I have the credentials to ask for ;-)
Show 10 quoted lines
> Anyway, here is what I did:
> paolo@Italia:~$ cg-clone
> http://digilander.libero.it/mcostalba/scm/qgit.git qgit defaulting to
>
> Why am I getting this error?
> error: File 8dea03519e75f47da91108330dde3043defddd60
> (http://digilander.libero.i
> t/mcostalba/scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60)
> corr upt
>

Because http server of digilander.libero.it instead of responding with 404 code (page not found) sends a not standard html page as answer. To see the page just point your browser to: http://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d

Git does not understand object is missing and thinks what site sends _is_ the requested object and then founds that is (of course) corrupted.

Marco
Junio C Hamano· Mar 19, 2006, 19:37 UTC · re: Marco Costalba · lore

Re: Cloning from sites with 404 overridden

"Marco Costalba" <mcostalba@gmail.com> writes:
Show 5 quoted lines
> http://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d
>
> Git does not understand object is missing and thinks what site sends
> _is_ the requested
> object and then founds that is (of course) corrupted.

To be fair, the site is _not_ missing anything from HTTP protocol perspective, because when git asks 8d/ea0351... file, the server responds with a regular "HTTP/1.0 200 OK" response. So it is _your_ repository that is corrupt -- instead of correctly _lacking_ the file you should have removed with prune-packed, it has a garbage file.

Having said that, I agree that it would be nicer if we support such a site, in the same spirit that we already bend backwards to support really dumb hosted http servers that do not give directory index by using objects/info/packs and info/refs.

I think it wouldn't be too much a hassle to add logic to http-fetch.c (perhaps with an additional "--no-404" option or somesuch) to fall back on pack transfer upon seeing a corrupt loose object. We do the falling back when getting 404 error to a request for a loose object, so the new code would essentially do the same and you might be OK.

Marco Costalba· Mar 19, 2006, 21:40 UTC · re: Junio C Hamano · lore

Re: Cloning from sites with 404 overridden

On 3/19/06, Junio C Hamano <junkio@cox.net> wrote:
Show 15 quoted lines
> "Marco Costalba" <mcostalba@gmail.com> writes:
>
> > http://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d
> >
> > Git does not understand object is missing and thinks what site sends
> > _is_ the requested
> > object and then founds that is (of course) corrupted.
>
> To be fair, the site is _not_ missing anything from HTTP
> protocol perspective, because when git asks 8d/ea0351... file,
> the server responds with a regular "HTTP/1.0 200 OK" response.
> So it is _your_ repository that is corrupt -- instead of
> correctly _lacking_ the file you should have removed with
> prune-packed, it has a garbage file.
>

Currently my git repo layout is as follow $ pwd <local master copy>/qgit.git/.git $ ls branches/ description HEAD index objects/ refs/ config FETCH_HEAD hooks/ info/ ORIG_HEAD remotes/ $ ls objects 2c/ 32/ 53/ 5c/ 6a/ info/ pack/

The host copy should be the exact mirror of the local copy (I use sitecopy to sync host). I have also verified this directly accessing the host with ftp.

So the 8d/ea0351... file is really not existent. BTW I have run git prune and git-prune-packed also.

Finally accessing the missing object with a browser

http://digilander.libero.it/mcostalba/ scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60

gives a pre-canned (in italian) 'Sorry page not found' stuff.

So I really think the site "HTTP/1.0 200 OK" response it's a fake. Perhaps security related to avoid sniffing (just a guess because I have absolutely zero competence in security related things).

Marco
Junio C Hamano· Mar 19, 2006, 23:21 UTC · re: Marco Costalba · lore

Re: Cloning from sites with 404 overridden

"Marco Costalba" <mcostalba@gmail.com> writes:
Show 10 quoted lines
> Finally accessing the missing object with a browser
>
> http://digilander.libero.it/mcostalba/
> scm/qgit.git/objects/8d/ea03519e75f47da91108330dde3043defddd60
>
> gives a pre-canned (in italian) 'Sorry page not found' stuff.
>
> So I really think the site "HTTP/1.0 200 OK" response it's a fake.
> Perhaps security related to avoid sniffing (just a guess because I have
> absolutely zero competence in security related things).

I think you are just rephrasing what I said. From the HTTP protocol perspective, you _do_ have that 8d/3a0351 thing on that server, because you do not correctly say "No we donot have it" using 404 response.

Your inability to produce 404 is a different matter -- often the hosting server is not under your control. But that does not change the fact that the repository observed by your clients is "broken". That is why a workaround flag like I suggested may be needed for such a setup.

This is totally untested, but maybe something like this?
---
diff --git a/http-fetch.c b/http-fetch.c
index 7de818b..d523798 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -8,6 +8,7 @@
 #define RANGE_HEADER_SIZE 30
 
 static int got_alternates = -1;
+static int unreliable_404 = 0;
 
 static struct curl_slist *no_pragma_header;
 
@@ -822,12 +823,18 @@ static int fetch_object(struct alt_base 
 		close(obj_req->local); obj_req->local = -1;
 	}
 
+	
+
 	if (obj_req->state == ABORTED) {
 		ret = error("Request for %s aborted", hex);
-	} else if (obj_req->curl_result != CURLE_OK &&
-		   obj_req->http_code != 416) {
+	} else if ((obj_req->curl_result != CURLE_OK &&
+		    obj_req->http_code != 416)  ||
+		   (unreliable_404 &&
+		    obj_req->curl_result == CURLE_OK &&
+		    obj_req->zret != Z_STREAM_END)) {
 		if (obj_req->http_code == 404 ||
-		    obj_req->curl_result == CURLE_FILE_COULDNT_READ_FILE)
+		    obj_req->curl_result == CURLE_FILE_COULDNT_READ_FILE ||
+		    unreliable_404)
 			ret = -1; /* Be silent, it is probably in a pack. */
 		else
 			ret = error("%s (curl_result = %d, http_code = %ld, sha1 = %s)",
@@ -966,6 +973,8 @@ int main(int argc, char **argv)
 			arg++;
 		} else if (!strcmp(argv[arg], "--recover")) {
 			get_recover = 1;
+		} else if (!strcmp(argv[arg], "--unreliable-404")) {
+			unreliable_404 = 1;
 		}
 		arg++;
 	}
Marco Costalba· Mar 20, 2006, 06:31 UTC · re: Junio C Hamano · lore

Re: Cloning from sites with 404 overridden

On 3/20/06, Junio C Hamano <junkio@cox.net> wrote:
Show 9 quoted lines
>
> Your inability to produce 404 is a different matter -- often the
> hosting server is not under your control.  But that does not
> change the fact that the repository observed by your clients is
> "broken".  That is why a workaround flag like I suggested may be
> needed for such a setup.
>
> This is totally untested, but maybe something like this?
>
It works for me. Just some trailing white space warning when applying.

I didn't found a way to pass '--unreliable-404' flag from git-clone, perhaps my bad, I have tested forcing the flag in sources.

Marco
Junio C Hamano· Mar 20, 2006, 08:44 UTC · re: Marco Costalba · lore

Re: Cloning from sites with 404 overridden

"Marco Costalba" <mcostalba@gmail.com> writes:
>> This is totally untested, but maybe something like this?
>
> It works for me. Just some trailing white space warning when applying.

The change only removes the error message without changing any other logic, so if that works for you, I wonder if leaving things as they are is a better option than doing anything short of implementing an AI that tries to pattern-match the "allegedly corrupt file" with "sorry no such page found" in many natural languages.

My test patch makes it impossible to track down the real breakage when an HTTP-reachable repository _does_ have a corrupt object.

So how about doing this instead?
-- >8 --
diff --git a/http-fetch.c b/http-fetch.c
index 8fd9de0..1405c1f 100644
--- a/http-fetch.c
+++ b/http-fetch.c
@@ -8,6 +8,7 @@
 #define RANGE_HEADER_SIZE 30
 
 static int got_alternates = -1;
+static int corrupt_object_found = 0;
 
 static struct curl_slist *no_pragma_header;
 
@@ -830,6 +831,7 @@ static int fetch_object(struct alt_base 
 				    obj_req->errorstr, obj_req->curl_result,
 				    obj_req->http_code, hex);
 	} else if (obj_req->zret != Z_STREAM_END) {
+		corrupt_object_found++;
 		ret = error("File %s (%s) corrupt", hex, obj_req->url);
 	} else if (memcmp(obj_req->sha1, obj_req->real_sha1, 20)) {
 		ret = error("File %s has bad hash", hex);
@@ -989,5 +991,11 @@ int main(int argc, char **argv)
 
 	http_cleanup();
 
+	if (corrupt_object_found) {
+		fprintf(stderr,
+"Some loose object were found to be corrupt, but they might be just\n"
+"a false '404 Not Found' error message sent with incorrect HTTP\n"
+"status code.  Suggest running git fsck-objects.\n");
+	}
 	return rc;
 }
Marco Costalba· Mar 20, 2006, 12:17 UTC · re: Junio C Hamano · lore

Re: Cloning from sites with 404 overridden

On 3/20/06, Junio C Hamano <junkio@cox.net> wrote:
Show 20 quoted lines
> "Marco Costalba" <mcostalba@gmail.com> writes:
>
> >> This is totally untested, but maybe something like this?
> >
> > It works for me. Just some trailing white space warning when applying.
>
> The change only removes the error message without changing any
> other logic, so if that works for you, I wonder if leaving
> things as they are is a better option than doing anything short
> of implementing an AI that tries to pattern-match the "allegedly
> corrupt file" with "sorry no such page found" in many natural
> languages.
>
> My test patch makes it impossible to track down the real
> breakage when an HTTP-reachable repository _does_ have a corrupt
> object.
>
> So how about doing this instead?
>
> -- >8 --
Show 8 quoted lines
> +               fprintf(stderr,
> +"Some loose object were found to be corrupt, but they might be just\n"
> +"a false '404 Not Found' error message sent with incorrect HTTP\n"
> +"status code.  Suggest running git fsck-objects.\n");
> +       }
>         return rc;
>  }
>
I think it's better, read more correct.

Could be a real corrupted file or just a false 404, so better a warning then an error message and also better a warning then nothing.

Marco
Lukas Sandström· Mar 20, 2006, 18:29 UTC · re: Junio C Hamano · lore

Re: Cloning from sites with 404 overridden

Junio C Hamano wrote:
Show 9 quoted lines
> "Marco Costalba" <mcostalba@gmail.com> writes:
>>http://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d
> 
> To be fair, the site is _not_ missing anything from HTTP
> protocol perspective, because when git asks 8d/ea0351... file,
> the server responds with a regular "HTTP/1.0 200 OK" response.
> So it is _your_ repository that is corrupt -- instead of
> correctly _lacking_ the file you should have removed with
> prune-packed, it has a garbage file.
Actually, it sends a 302 redirect. 
Perhaps a repository config option to treat a 302 as a 404?
/Lukas Sandström
Petr Baudis· Mar 20, 2006, 19:43 UTC · re: Lukas Sandström · lore

Re: Cloning from sites with 404 overridden

Dear diary, on Mon, Mar 20, 2006 at 07:29:02PM CET, I got a letter where Lukas Sandström <lukass@etek.chalmers.se> said that...

Show 14 quoted lines
> Junio C Hamano wrote:
> > "Marco Costalba" <mcostalba@gmail.com> writes:
> >>http://digilander.libero.it /mcostalba/scm/qgit.git/objects/8d/ea03519e75f47d
> > 
> > To be fair, the site is _not_ missing anything from HTTP
> > protocol perspective, because when git asks 8d/ea0351... file,
> > the server responds with a regular "HTTP/1.0 200 OK" response.
> > So it is _your_ repository that is corrupt -- instead of
> > correctly _lacking_ the file you should have removed with
> > prune-packed, it has a garbage file.
> 
> Actually, it sends a 302 redirect. 
> 
> Perhaps a repository config option to treat a 302 as a 404?

I think that would be too ugly _and_ specific a workaround for the particular site. It's reasonable to keep it generalized for all the broken repositories when already doing it.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Nick Hengeveld· Mar 20, 2006, 19:54 UTC · re: Lukas Sandström · lore

Re: Cloning from sites with 404 overridden

On Mon, Mar 20, 2006 at 07:29:02PM +0100, Lukas Sandström wrote:
> Perhaps a repository config option to treat a 302 as a 404?

FWIW, it used to work that way and was modified to follow redirects back at commit 66c9ec25553ce7332c46e2017b9c4d7c26310fff.

-- 
For a successful technology, reality must take precedence over public
relations, for nature cannot be fooled.
Junio C Hamano· Mar 19, 2006, 19:47 UTC · re: Marco Costalba · lore

Re: Cloning from sites with 404 overridden

"Marco Costalba" <mcostalba@gmail.com> writes:
Show 5 quoted lines
> On 3/19/06, Paolo Ciarrocchi <paolo.ciarrocchi@gmail.com> wrote:
>>
>> How about getting an account on kernel.org?
>
> I don't think I have the credentials to ask for ;-)

Heh, it has a striking resemblance to the first thing I said when Linus asked me if I want to take over git.git: "It would be embarrassing to be the first person to have an account there without having a single line of code in the kernel" ;-).

Well, you won't be the first (in fact it appears I wasn't either), and it would never hurt to ask.

Petr Baudis· Mar 19, 2006, 21:31 UTC · re: Junio C Hamano · lore

Re: Cloning from sites with 404 overridden

Dear diary, on Sun, Mar 19, 2006 at 08:47:21PM CET, I got a letter where Junio C Hamano <junkio@cox.net> said that...

Show 15 quoted lines
> "Marco Costalba" <mcostalba@gmail.com> writes:
> 
> > On 3/19/06, Paolo Ciarrocchi <paolo.ciarrocchi@gmail.com> wrote:
> >>
> >> How about getting an account on kernel.org?
> >
> > I don't think I have the credentials to ask for ;-)
> 
> Heh, it has a striking resemblance to the first thing I said
> when Linus asked me if I want to take over git.git: "It would
> be embarrassing to be the first person to have an account there
> without having a single line of code in the kernel" ;-).
> 
> Well, you won't be the first (in fact it appears I wasn't
> either), and it would never hurt to ask.
Yeah, I think I was there before you... ;-)
-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Petr Baudis· Mar 19, 2006, 21:43 UTC · re: Petr Baudis · lore

Re: Cloning from sites with 404 overridden

Dear diary, on Sun, Mar 19, 2006 at 10:31:25PM CET, I got a letter where Petr Baudis <pasky@suse.cz> said that...

Show 11 quoted lines
> Dear diary, on Sun, Mar 19, 2006 at 08:47:21PM CET, I got a letter
> where Junio C Hamano <junkio@cox.net> said that...
> > Heh, it has a striking resemblance to the first thing I said
> > when Linus asked me if I want to take over git.git: "It would
> > be embarrassing to be the first person to have an account there
> > without having a single line of code in the kernel" ;-).
> > 
> > Well, you won't be the first (in fact it appears I wasn't
> > either), and it would never hurt to ask.
> 
> Yeah, I think I was there before you... ;-)

Silly me, on a second thought I've realized that I already had some stuff in the kernel by then. Sorry for the noise.

-- 
				Petr "Pasky" Baudis
Stuff: http://pasky.or.cz/
Right now I am having amnesia and deja-vu at the same time.  I think
I have forgotten this before.
Marco Costalba· Mar 19, 2006, 21:45 UTC · re: Petr Baudis · lore

Re: Cloning from sites with 404 overridden

On 3/19/06, Petr Baudis <pasky@suse.cz> wrote:
Show 21 quoted lines
> Dear diary, on Sun, Mar 19, 2006 at 08:47:21PM CET, I got a letter
> where Junio C Hamano <junkio@cox.net> said that...
> > "Marco Costalba" <mcostalba@gmail.com> writes:
> >
> > > On 3/19/06, Paolo Ciarrocchi <paolo.ciarrocchi@gmail.com> wrote:
> > >>
> > >> How about getting an account on kernel.org?
> > >
> > > I don't think I have the credentials to ask for ;-)
> >
> > Heh, it has a striking resemblance to the first thing I said
> > when Linus asked me if I want to take over git.git: "It would
> > be embarrassing to be the first person to have an account there
> > without having a single line of code in the kernel" ;-).
> >
> > Well, you won't be the first (in fact it appears I wasn't
> > either), and it would never hurt to ask.
>
> Yeah, I think I was there before you... ;-)
>
> --
Please could someone tell me what door I should knock at?

Thanks Marco

Randal L. Schwartz· Mar 20, 2006, 04:32 UTC · re: Junio C Hamano · lore

Re: Cloning from sites with 404 overridden

>>>>> "Junio" == Junio C Hamano <junkio@cox.net> writes:

Junio> Heh, it has a striking resemblance to the first thing I said Junio> when Linus asked me if I want to take over git.git: "It would Junio> be embarrassing to be the first person to have an account there Junio> without having a single line of code in the kernel" ;-).

Junio> Well, you won't be the first (in fact it appears I wasn't Junio> either), and it would never hurt to ask.

Wow. That would perhaps completely rule out people who have never owned anything that can execute the x86 instruction set except in emulation. :)

-- 
Randal L. Schwartz - Stonehenge Consulting Services, Inc. - +1 503 777 0095
<merlyn@stonehenge.com> <URL:http://www.stonehenge.com/merlyn/>
Perl/Unix/security consulting, Technical writing, Comedy, etc. etc.
See PerlTraining.Stonehenge.com for onsite and open-enrollment Perl training!

← back to recent threads