git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: I'm a total push-over..

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Jan 23, 2008, 16:06 UTC
Message-ID
<alpine.LFD.1.00.0801230754180.1741@woody.linux-foundation.org>
In-Reply-To
<4796FBB6.9080609@op5.se>
On Wed, 23 Jan 2008, Andreas Ericsson wrote:
> 
> Insofar as hashes go, it's not that shabby for hashing filenames.

Hashing filenames is pretty easy. You can do a reasonable job with any "multiply by an odd number, add in value". Picking the odd number is a random choice, some are better than others (and it depends on whether you end up always having a power-of-two bucket size etc), but it generally won't be horrible.

And considering that we generally won't have tons and tons of pathnames (ie we'd generally have thousands, not millions), and that the underlying hash not only resizes itself but actually uses the ful 32 bits as the lookup key, I don't worry too much. I suspect my random choice is fine.

So no, I didn't think my hash would _suck_, although I also didn't research which odd numbers to pick.

No, it's a joke because it doesn't really give an example of how to do the _expected_ hash collissions.

Here's another hash that is actually going to collide *much* more (and on purpose!), but is actually showing an example of something that actually hashes UTF strings so that upper-case and lower-case (and normalization) will all still hash to the same value:

	static unsigned int hash_name(const char *name, int namelen)
	{
	        unsigned int hash = 0x123;
	        do {
	                unsigned char c = *name++;
			if (c & 0x80)
				c = 0;
			c &= ~0x20;
	                hash = hash*101 + c;
	        } while (--namelen);
	        return hash;
	}

but the above does so by making the hash much much worse (although probably still acceptable for "normal source code name distributions" that don't have very many same-name-in-different-cases and high bit characters anyway).

The above is still fairly fast, but obviously at a serious cost in hash goodness, to the point of being totally unusable for anybody who uses Asian characters in their filenames.

To actually be really useful, you'd have to teach it about the character system and do a lookup into a case/normalization table.

So *that* is mainly why it's a joke. But it should work fine for the case it is used for now (exact match).

Picking a better-researched constant might still be a good idea.
		Linus
Previous: Andreas Ericsson
Message 51 of 51 in “I'm a total push-over..”
  1. Linus TorvaldsJan 22, 2008
  2. Kevin BallardJan 23, 2008
  3. Junio C HamanoJan 23, 2008
  4. Junio C HamanoJan 23, 2008
  5. Johannes SchindelinJan 23, 2008
  6. David KastrupJan 23, 2008
  7. Theodore TsoJan 23, 2008
  8. Linus TorvaldsJan 23, 2008
  9. Linus TorvaldsJan 23, 2008
  10. Junio C HamanoJan 25, 2008
  11. Linus TorvaldsJan 25, 2008
  12. Junio C HamanoJan 23, 2008
  13. Johannes SchindelinJan 23, 2008
  14. Linus TorvaldsJan 23, 2008
  15. Johannes SchindelinJan 23, 2008
  16. Linus TorvaldsJan 23, 2008
  17. Linus TorvaldsJan 23, 2008
  18. Jeremy Maitin-ShepardJan 25, 2008
  19. Johannes SchindelinJan 25, 2008
  20. Jeremy Maitin-ShepardJan 25, 2008
  21. Johannes SchindelinJan 25, 2008
  22. Junio C HamanoJan 25, 2008
  23. Andreas EricssonJan 23, 2008
  24. Dmitry PotapovJan 23, 2008
  25. Andreas EricssonJan 23, 2008
  26. Marko KreenJan 23, 2008
  27. Andreas EricssonJan 23, 2008
  28. Luke LuJan 24, 2008
  29. Andreas EricssonJan 24, 2008
  30. Marko KreenJan 24, 2008
  31. Andreas EricssonJan 24, 2008
  32. Marko KreenJan 24, 2008
  33. Dmitry PotapovJan 24, 2008
  34. Linus TorvaldsJan 24, 2008
  35. Dmitry PotapovJan 24, 2008
  36. Linus TorvaldsJan 24, 2008
  37. Marko KreenJan 25, 2008
  38. Linus TorvaldsJan 25, 2008
  39. Linus TorvaldsJan 25, 2008
  40. Marko KreenJan 26, 2008
  41. Linus TorvaldsJan 27, 2008
  42. Dmitry PotapovJan 27, 2008
  43. Johannes SchindelinJan 27, 2008
  44. Dmitry PotapovJan 27, 2008
  45. Marko KreenJan 27, 2008
  46. Dmitry PotapovJan 27, 2008
  47. Marko KreenJan 26, 2008
  48. Marko KreenJan 25, 2008
  49. Dmitry PotapovJan 23, 2008
  50. Andreas EricssonJan 24, 2008
  51. Linus TorvaldsJan 23, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.