git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: I'm a total push-over..

From
Johannes Schindelin <johannes.schindelin@gmx.de>
Date
Jan 23, 2008, 16:34 UTC
Message-ID
<alpine.LSU.1.00.0801231630480.5731@racer.site>
In-Reply-To
<alpine.LFD.1.00.0801230817390.1741@woody.linux-foundation.org>
Hi,
On Wed, 23 Jan 2008, Linus Torvalds wrote:
Show 16 quoted lines
> On Wed, 23 Jan 2008, Johannes Schindelin wrote:
> > 
> > I fully expect it to be noticable with that UTF-8 "normalisation".  
> > But then, the infrastructure is there, and whoever has an itch to 
> > scratch...
> 
> Actually, it's going to be totally invisible even with UTF-8 
> normalization, because we're going to do it sanely.
> 
> And by "sanely" I mean just having the code test the high bit, and using 
> US-ASCII as-is (possibly with that " & ~0x20 " thing to ignore case in 
> it).
> 
> End result: practically all projects will never notice anything at all for 
> 99.9% of all files. One extra well-predicted branch, and a few more hash 
> collissions for cases where you have both "Makefile" and "makefile" etc.

Well, that's the point, to avoid having both "Makefile" and "makefile" in your repository when you are on case-challenged filesystems, right?

Show 15 quoted lines
> Doing names with *lots* of UTF-8 characters will be rather slower. It's 
> still not horrible to do if you do it the smart way, though. In fact, 
> it's pretty simple, just a few table lookups (one to find the NFD form, 
> one to do the upcasing).
> 
> And yes, for hashing, it makes sense to turn things into NFD because 
> it's generally simpler, but the point is that you really don't actually 
> modify the name itself at all, you just hash things (or compare things) 
> character by expanded character.
> 
> IOW, only a total *moron* does Unicode name comparisons with
> 
> 	strcmp(convert_to_nfd(a), convert_to_nfd(b));
> 
> which is essentially what Apple does.

Heh, indeed that is what I would have done as an initial step (out of laziness).

Show 6 quoted lines
> It's quite possible to do
> 
> 	utf8_nfd_strcmp(a,b)
> 
> and (a) do it tons and tons faster and (b) never have to modify the 
> strings themselves. Same goes (even more) for hashing.
Okay.  Point taken.

But I really hope that you are not proposing to use the case-ignoring hash when we are _not_ on a case-challenged filesystem...

Ciao, Dscho

Previous: Linus TorvaldsNext: Linus Torvalds
Message 15 of 51 in “I'm a total push-over..”
  1. Linus TorvaldsJan 22, 2008
  2. Kevin BallardJan 23, 2008
  3. Junio C HamanoJan 23, 2008
  4. Junio C HamanoJan 23, 2008
  5. Johannes SchindelinJan 23, 2008
  6. David KastrupJan 23, 2008
  7. Theodore TsoJan 23, 2008
  8. Linus TorvaldsJan 23, 2008
  9. Linus TorvaldsJan 23, 2008
  10. Junio C HamanoJan 25, 2008
  11. Linus TorvaldsJan 25, 2008
  12. Junio C HamanoJan 23, 2008
  13. Johannes SchindelinJan 23, 2008
  14. Linus TorvaldsJan 23, 2008
  15. Johannes SchindelinJan 23, 2008
  16. Linus TorvaldsJan 23, 2008
  17. Linus TorvaldsJan 23, 2008
  18. Jeremy Maitin-ShepardJan 25, 2008
  19. Johannes SchindelinJan 25, 2008
  20. Jeremy Maitin-ShepardJan 25, 2008
  21. Johannes SchindelinJan 25, 2008
  22. Junio C HamanoJan 25, 2008
  23. Andreas EricssonJan 23, 2008
  24. Dmitry PotapovJan 23, 2008
  25. Andreas EricssonJan 23, 2008
  26. Marko KreenJan 23, 2008
  27. Andreas EricssonJan 23, 2008
  28. Luke LuJan 24, 2008
  29. Andreas EricssonJan 24, 2008
  30. Marko KreenJan 24, 2008
  31. Andreas EricssonJan 24, 2008
  32. Marko KreenJan 24, 2008
  33. Dmitry PotapovJan 24, 2008
  34. Linus TorvaldsJan 24, 2008
  35. Dmitry PotapovJan 24, 2008
  36. Linus TorvaldsJan 24, 2008
  37. Marko KreenJan 25, 2008
  38. Linus TorvaldsJan 25, 2008
  39. Linus TorvaldsJan 25, 2008
  40. Marko KreenJan 26, 2008
  41. Linus TorvaldsJan 27, 2008
  42. Dmitry PotapovJan 27, 2008
  43. Johannes SchindelinJan 27, 2008
  44. Dmitry PotapovJan 27, 2008
  45. Marko KreenJan 27, 2008
  46. Dmitry PotapovJan 27, 2008
  47. Marko KreenJan 26, 2008
  48. Marko KreenJan 25, 2008
  49. Dmitry PotapovJan 23, 2008
  50. Andreas EricssonJan 24, 2008
  51. Linus TorvaldsJan 23, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.