git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: I'm a total push-over..

From
MKMarko Kreen <markokr@gmail.com>
Date
Jan 26, 2008, 12:16 UTC
Message-ID
<e51f66da0801260416p5f5ffb98w16fe832fe62dc7c9@mail.gmail.com>
In-Reply-To
<alpine.LFD.1.00.0801251407010.5056@hp.linux-foundation.org>
On 1/26/08, Linus Torvalds <torvalds@linux-foundation.org> wrote:
Show 23 quoted lines
> On Fri, 25 Jan 2008, Marko Kreen wrote:
> > Well, although this is very clever approach, I suggest against it.
> > You'll end up with complex code that gives out substandard results.
>
> Actually, *your* operation is the one that gives substandard results.
>
> > I think its better to have separate case-folding function (or several),
> > that copies string to temp buffer and then run proper optimized hash
> > function on that buffer.
>
> I'm sorry, but you just cannot do that efficiently and portably.
>
> I can write a hash function that reliably does 8 bytes at a time for the
> common case on a 64-bit architecture, exactly because it's easy to do
> "test high bits in parallel" with a simple bitwise 'and', and we can do
> the same with "approximate lower-to-uppercase 8 bytes at a time" for a
> hash by just clearing bit 5.
>
> In contrast, trying to do the same thing in half-way portable C, but being
> limited to having to get the case-folding *exactly* right (which you need
> for the comparison function) is much much harder. It's basically
> impossible in portable C (it's doable with architecture-specific features,
> ie vector extensions that have per-byte compares etc).
Here you misunderstood me, I was proposing following:
int hash_folded(const char *str, int len)
{
   char buf[512];
   do_folding(buf, str, len);
   return do_hash(buf, len);
}
That is - the folded string should stay internal to hash function.

Only difference from combined foling+hashing would be that you can code each part separately.

Show 15 quoted lines
> And hashing is performance-critical, much more so than the compares (ie
> you're likely to have to hash tens of thousands of files, while you will
> only compare a couple). So it really is worth optimizing for.
>
> And the thing is, "performance" isn't a secondary feature. It's also not
> something you can add later by optimizing.
>
> It's also a mindset issue. Quite frankly, people who do this by "convert
> to some folded/normalized form, then do the operation" will generally make
> much more fundamental mistakes. Once you get into the mindset of "let's
> pass a corrupted strign around", you are in trouble. You start thinking
> that the corrupted string isn't really "corrupt", it's in an "optimized
> format".
>
> And it's all downhill from there. Don't do it.

Againg, you seem to keep HFS+ behaviour in mind, but that was not what I did suggest. Probably my mistake, sorry.

-- 
marko
Previous: Linus TorvaldsNext: Linus Torvalds
Message 40 of 51 in “I'm a total push-over..”
  1. Linus TorvaldsJan 22, 2008
  2. Kevin BallardJan 23, 2008
  3. Junio C HamanoJan 23, 2008
  4. Junio C HamanoJan 23, 2008
  5. Johannes SchindelinJan 23, 2008
  6. David KastrupJan 23, 2008
  7. Theodore TsoJan 23, 2008
  8. Linus TorvaldsJan 23, 2008
  9. Linus TorvaldsJan 23, 2008
  10. Junio C HamanoJan 25, 2008
  11. Linus TorvaldsJan 25, 2008
  12. Junio C HamanoJan 23, 2008
  13. Johannes SchindelinJan 23, 2008
  14. Linus TorvaldsJan 23, 2008
  15. Johannes SchindelinJan 23, 2008
  16. Linus TorvaldsJan 23, 2008
  17. Linus TorvaldsJan 23, 2008
  18. Jeremy Maitin-ShepardJan 25, 2008
  19. Johannes SchindelinJan 25, 2008
  20. Jeremy Maitin-ShepardJan 25, 2008
  21. Johannes SchindelinJan 25, 2008
  22. Junio C HamanoJan 25, 2008
  23. Andreas EricssonJan 23, 2008
  24. Dmitry PotapovJan 23, 2008
  25. Andreas EricssonJan 23, 2008
  26. Marko KreenJan 23, 2008
  27. Andreas EricssonJan 23, 2008
  28. Luke LuJan 24, 2008
  29. Andreas EricssonJan 24, 2008
  30. Marko KreenJan 24, 2008
  31. Andreas EricssonJan 24, 2008
  32. Marko KreenJan 24, 2008
  33. Dmitry PotapovJan 24, 2008
  34. Linus TorvaldsJan 24, 2008
  35. Dmitry PotapovJan 24, 2008
  36. Linus TorvaldsJan 24, 2008
  37. Marko KreenJan 25, 2008
  38. Linus TorvaldsJan 25, 2008
  39. Linus TorvaldsJan 25, 2008
  40. Marko KreenJan 26, 2008
  41. Linus TorvaldsJan 27, 2008
  42. Dmitry PotapovJan 27, 2008
  43. Johannes SchindelinJan 27, 2008
  44. Dmitry PotapovJan 27, 2008
  45. Marko KreenJan 27, 2008
  46. Dmitry PotapovJan 27, 2008
  47. Marko KreenJan 26, 2008
  48. Marko KreenJan 25, 2008
  49. Dmitry PotapovJan 23, 2008
  50. Andreas EricssonJan 24, 2008
  51. Linus TorvaldsJan 23, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.