git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: I'm a total push-over..

From
Andreas Ericsson <ae@op5.se>
Date
Jan 24, 2008, 10:24 UTC
Message-ID
<47986775.7010603@op5.se>
In-Reply-To
<F23CA352-416C-49EC-8132-688784CF3C18@vicaya.com>
Luke Lu wrote:
Show 63 quoted lines
> On Jan 23, 2008, at 6:39 AM, Andreas Ericsson wrote:
>> Marko Kreen wrote:
>>> On 1/23/08, Andreas Ericsson <ae@op5.se> wrote:
>>>> Dmitry Potapov wrote:
>>>>> On Wed, Jan 23, 2008 at 09:32:54AM +0100, Andreas Ericsson wrote:
>>>>>> The FNV hash would be better (pasted below), but I doubt
>>>>>> anyone will ever care, and there will be larger differences
>>>>>> between architectures with this one than the lt_git hash (well,
>>>>>> a function's gotta have a name).
>>>>> Actually, Bob Jenkins' lookup3 hash is twice faster in my tests
>>>>> than FNV, and also it is much less likely to have any collision.
>>>>>
>>>> >From http://burtleburtle.net/bob/hash/doobs.html
>>>> ---
>>>> FNV Hash
>>>>
>>>> I need to fill this in. Search the web for FNV hash. It's faster 
>>>> than my hash on Intel (because Intel has fast multiplication), but 
>>>> slower on most other platforms. Preliminary tests suggested it has 
>>>> decent distributions.
>>> I suspect that this paragraph was about comparison with lookup2
>>
>>
>> It might be. It's from the link Dmitry posted in his reply to my original
>> message. (something/something/doobs.html).
>>
>>> (not lookup3) because lookup3 beat easily all the "simple" hashes
>>
>> By how much? FNV beat Linus' hash by 0.01 microseconds / insertion,
>> and 0.1 microsecons / lookup. We're talking about a case here where
>> there will never be more lookups than insertions (unless I'm much
>> mistaken).
>>
>>> If you don't mind few percent speed penalty compared to Jenkings
>>> own optimized version, you can use my simplified version:
>>>   
>>> http://repo.or.cz/w/pgbouncer.git?a=blob;f=src/hash.c;h=5c9a73639ad098c296c0be562c34573189f3e083;hb=HEAD 
>>>
>>
>> I don't, but I don't care that deeply either. On the one hand,
>> it would be nifty to have an excellent hash-function in git.
>> On the other hand, it would look stupid with something that's
>> quite clearly over-kill.
>>
>>> It works always with "native" endianess, unlike Jenkins fixed-endian
>>> hashlittle() / hashbig().  It may or may not matter if you plan
>>> to write values on disk.
>>> Speed-wise it may be 10-30% slower worst case (in my case sparc-classic
>>> with unaligned data), but on x86, lucky gcc version and maybe
>>> also memcpy() hack seen in system.h, it tends to be ~10% faster,
>>> especially as it does always 4byte read in main loop.
>>
>> It would have to be a significant improvement in wall-clock time
>> on a test-case of hashing 30k strings to warrant going from 6 to 80
>> lines of code, imo. I still believe the original dumb hash Linus
>> wrote is "good enough".
>>
>> On a side-note, it was very interesting reading, and I shall have
>> to add jenkins3_mkreen() to my test-suite (although the "keep
>> copyright note" license thing bugs me a bit).
> 
> Would you, for completeness' sake, please add Tcl and STL hashes to your 
> test suite?

I could do that. Or I just publish the entire ugly thing and let someone else add them ;-)

> The numbers are quite interesting. Is your test suite 
> available somewhere, so we can test with our own data and hardware as 
> well.

Not yet, no. I usually munge it up quite a lot when I want to test hashes for a specific input, so it's not what anyone would call "pretty".

Show 8 quoted lines
> Both Tcl hash and STL (from SGI probably HP days, still the 
> current default with g++) string hashes are extremely simple (excluding 
> the loop constructs):
> 
> Tcl: h += (h<<3) + c;     // essentially *9+c (but work better on 
> non-late-intels)
> STL: h = h * 5 + c;    // worse than above for most of my data
> 

They sure do look simple enough. As for loop constructs, I've tried to use the same looping mechanics for everything, so as to let the algorithm be the only difference. Otherwise it gets tricky to do comparisons. The exceptions are ofcourse hashes relying on Duff's device or similar alignment trickery.

-- 
Andreas Ericsson                   andreas.ericsson@op5.se
OP5 AB                             www.op5.se
Tel: +46 8-230225                  Fax: +46 8-230231
Previous: Luke LuNext: Marko Kreen
Message 29 of 51 in “I'm a total push-over..”
  1. Linus TorvaldsJan 22, 2008
  2. Kevin BallardJan 23, 2008
  3. Junio C HamanoJan 23, 2008
  4. Junio C HamanoJan 23, 2008
  5. Johannes SchindelinJan 23, 2008
  6. David KastrupJan 23, 2008
  7. Theodore TsoJan 23, 2008
  8. Linus TorvaldsJan 23, 2008
  9. Linus TorvaldsJan 23, 2008
  10. Junio C HamanoJan 25, 2008
  11. Linus TorvaldsJan 25, 2008
  12. Junio C HamanoJan 23, 2008
  13. Johannes SchindelinJan 23, 2008
  14. Linus TorvaldsJan 23, 2008
  15. Johannes SchindelinJan 23, 2008
  16. Linus TorvaldsJan 23, 2008
  17. Linus TorvaldsJan 23, 2008
  18. Jeremy Maitin-ShepardJan 25, 2008
  19. Johannes SchindelinJan 25, 2008
  20. Jeremy Maitin-ShepardJan 25, 2008
  21. Johannes SchindelinJan 25, 2008
  22. Junio C HamanoJan 25, 2008
  23. Andreas EricssonJan 23, 2008
  24. Dmitry PotapovJan 23, 2008
  25. Andreas EricssonJan 23, 2008
  26. Marko KreenJan 23, 2008
  27. Andreas EricssonJan 23, 2008
  28. Luke LuJan 24, 2008
  29. Andreas EricssonJan 24, 2008
  30. Marko KreenJan 24, 2008
  31. Andreas EricssonJan 24, 2008
  32. Marko KreenJan 24, 2008
  33. Dmitry PotapovJan 24, 2008
  34. Linus TorvaldsJan 24, 2008
  35. Dmitry PotapovJan 24, 2008
  36. Linus TorvaldsJan 24, 2008
  37. Marko KreenJan 25, 2008
  38. Linus TorvaldsJan 25, 2008
  39. Linus TorvaldsJan 25, 2008
  40. Marko KreenJan 26, 2008
  41. Linus TorvaldsJan 27, 2008
  42. Dmitry PotapovJan 27, 2008
  43. Johannes SchindelinJan 27, 2008
  44. Dmitry PotapovJan 27, 2008
  45. Marko KreenJan 27, 2008
  46. Dmitry PotapovJan 27, 2008
  47. Marko KreenJan 26, 2008
  48. Marko KreenJan 25, 2008
  49. Dmitry PotapovJan 23, 2008
  50. Andreas EricssonJan 24, 2008
  51. Linus TorvaldsJan 23, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.