git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Cloning from sites with 404 overridden

From
Shawn Pearce <spearce@spearce.org>
Date
Mar 22, 2006, 03:12 UTC
Message-ID
<20060322031200.GB17954@spearce.org>
In-Reply-To
<20060322025921.1722.qmail@science.horizon.com>

'0' x 40. :-) There's some places already in the GIT source which would have ``issues'' if they got an object with this hash. Not sure if it is actually an entirely impossible hash or just one that is highly improbable.

My own website has this problem and its because I'm using WordPress to handle all URLs on the site; I haven't yet found a way to configure WordPress to return a proper 404 when the URL can't be mapped to something on the server. Note that 404 status codes can in fact return pretty HTML content for the user, and many websites do this and many browsers display that pretty HTML. But a bot can then also recognize the status code and DTRT.

The webservers are just plain broken, mine included. I think the best option is to delay corrupt object reporting to the end of the download process if you get only one corrupt object and that corrupt object was actually attainable from a pack. And in this case its just a minor warning:

	Warning: The server appears to not return proper HTTP status
	codes on missing files.  The files were found in one or
	more packs so the download is OK, but the server administrator
	should really fix their server.  If you know the server
	administrator you might want to prod them to do so.

But that's already been suggested and I thought someone worked up a patch based on that idea? If not I could try to do so since my own damn server has the problem. :-)

linux@horizon.com wrote:
Show 16 quoted lines
> If someone feels ambitious, you can detect this condition automatically
> by searching for a file that you know won't be there and seeing if you
> get a 404 response to that.
> 
> To avoid punishing good servers, it would be nice to defer the test
> until reciving the first corrupted object.
> 
> I'm not sure what the best "object that's not supposed to be there" is.
> It could just be a random hash, or would a malformed object file name
> be better?  Any fixed name has a finite chance of being created by
> someone somewhere, but generating 160-bit random numbers is a PITA on
> non-freenix platforms.
> 
> 
> (As an aside, I suspect this is all caused by Microsoft's "friendly HTML
> error messages" invention.)
-- 
Shawn.
Previous: linux@horizon.comNext: Linus Torvalds
Message 2 of 18 in “Re: Cloning from sites with 404 overridden”
  1. linux@horizon.comMar 22, 2006
  2. Shawn PearceMar 22, 2006
  3. Linus TorvaldsMar 22, 2006
  4. Marco CostalbaMar 22, 2006
  5. Junio C HamanoMar 22, 2006
  6. Andreas EricssonMar 22, 2006
  7. Mark WoodingMar 24, 2006
  8. Junio C HamanoMar 24, 2006
  9. Linus TorvaldsMar 24, 2006
  10. Morten WelinderMar 24, 2006
  11. Andreas EricssonMar 24, 2006
  12. Nick HengeveldMar 22, 2006
  13. Nick HengeveldMar 22, 2006
  14. Junio C HamanoMar 22, 2006
  15. Junio C HamanoMar 22, 2006
  16. Nick HengeveldMar 23, 2006
  17. Junio C HamanoMar 23, 2006
  18. Radoslaw SzkodzinskiMar 22, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.