git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: git on MacOSX and files with decomposed utf-8 file names

From
Kevin Ballard <kevin@sb.org>
Date
Jan 21, 2008, 19:05 UTC
Message-ID
<C6C0E6A1-053B-48CE-90B3-8FFB44061C3B@sb.org>
In-Reply-To
<alpine.LFD.1.00.0801210934400.2957@woody.linux-foundation.org>
On Jan 21, 2008, at 1:12 PM, Linus Torvalds wrote:
Show 20 quoted lines
> On Mon, 21 Jan 2008, Kevin Ballard wrote:
>> On Jan 21, 2008, at 9:14 AM, Peter Karlsson wrote:
>>>
>>> I happen to prefer the text-as-string-of-characters (or code points,
>>> since you use the other meaning of characters in your posts),  
>>> since I
>>> come from the text world, having worked a lot on Unicode text
>>> processing.
>>>
>>> You apparently prefer the text-as-sequence-of-octets, which I tend  
>>> to
>>> dislike because I would have thought computer engineers would have
>>> evolved beyond this when we left the 1900s.
>>
>> I agree. Every single problem that I can recall Linus bringing up  
>> as a
>> consequence of HFS+ treating filenames as strings [..]
>
> You say "I agree", BUT YOU DON'T EVEN SEEM TO UNDERSTAND WHAT IS  
> GOING ON.
I could say the same thing about you.
Show 30 quoted lines
> The fact is, text-as-string-of-codepoints (let's make the "codepoints"
> obvious, so that there is no ambiguity, but I'd also like to make it  
> clear
> that a codepoint *is* how a Unicode character is defined, and a  
> Unicode
> "string" is actually *defined* to be a sequence of codepoints, and  
> totally
> independent of normalization!) is fine.
>
> That was never the issue at all. Unicode codepoints are wonderful.
>
> Now, git _also_ heavily depends on the actual encoding of those
> codepoints, since we create hashes etc, so in fact, as far ass git is
> concerned, names have to be in some particular encoding to be  
> hashed, and
> UTF-8 is the only sane encoding for Unicode. People can blather about
> UCS-2 and UTF-16 and UTF-32 all they want, but the fact is, UTF-8 is
> simply technically superior in so many ways that I don't even  
> understand
> why anybody ever uses anything else.
>
> So I would not disagree with using UTF-8 at all.
>
> But that is *entirely* a separate issue from "normalization".
>
> Kevin, you seem to think that normalization is somehow forced on you  
> by
> the "text-as-codepoints" decision, and that is SIMPLY NOT TRUE.
> Normalization is a totally separate decision, and it's a STUPID one,
> because it breaks so many of the _nice_ properties of using UTF-8.

I'm not saying it's forced on you, I'm saying when you treat filenames as text, it DOESN'T MATTER if the string gets normalized. As long as the string remains equivalent, YOU DON'T CARE about the underlying byte stream.

Show 12 quoted lines
> And THAT is where we differ. It has nothing to do with "octets". It  
> has
> nothing to do with not liking Unicode. It has nothing to do with
> "strings".
>
> In short:
>
> - normalization is by no means required or even a good feature. It's
>   something you do when you want to know if two strings are  
> equivalent,
>   but that doesn't actually mean that you should keep the strings
>   normalized all the time!

Alright, fine. I'm not saying HFS+ is right in storing the normalized version, but I do believe the authors of HFS+ must have had a reason to do that, and I also believe that it shouldn't make any difference to me since it remains equivalent.

> - normalization has *nothing* to do with "treating text as octets".
>   That's entirely an encoding issue.

Sure it does. Normalizing a string produces an equivalent string, and so unless I look at the octets the two strings are, for all intents and purposes, the same.

Show 5 quoted lines
> - of *course* git has to treat things as a binary stream at some  
> point,
>   since you need that to even compute a SHA1 in the first place, but  
> that
>   has *nothing* to do with normalization or the lack of it.

You're right, but it doesn't have to treat it as a binary stream at the level I care about. I mean, no matter what you do at some level the string is evaluated as a binary stream. For our purposes, just redefine the hashing algorithm to hash all equivalent strings the same, and you can implement that by using SHA1 on a particular encoding of the string.

> Got it? Forced normalization is stupid, because it changes the data  
> and
> removes information, and unless you know that change is safe, it's the
> wrong thing to do.

Decomposing and recomposing shouldn't lose any information we care about - when treating filenames as text, a<COMBINING DIARESIS> and <A WITH DIARESIS> are equivalent, and thus no distinction is made between them. I'm not sure what other information you might be considering lost in this case.

Show 11 quoted lines
> One reason _not_ to do normalization is that if you don't, you can  
> still
> interact with no ambiguity with other non-Unicode locales. You can  
> do the
> 1:1 Latin1<->Unicode translation, and you *never* get into trouble. In
> cotnrast, if you normalize, it's no longer a 1:1 translation any  
> more, and
> you can get into a situation where the translation from Latin1 to  
> Unicode
> and back results in a *different* filename than the one you started  
> with!
I don't believe you. See below.
Show 9 quoted lines
> See? That's a *serious*problem*. A system that forces normalization BY
> DEFINITION cannot work with people who use a Latin1 filesystem,  
> because it
> will corrupt the filenames!
>
> But you are apparently too damn stupid to understand that "data
> corruption" == "bad", and too damn stupid to see that "Unicode" does  
> not
> mean "Forced normalization".
When have I ever said that Unicode meant Forced normalization?
Show 14 quoted lines
> But I'll try one more time. Let's say that I work on a project where  
> there
> are some people who use Latin1, and some people who use UTF-8, and  
> we use
> special characters. It should all work, as long as we use only the  
> common
> subset, and we teach git to convert to UTF-8 as a common base. Right?
>
> In your *idiotic* world, where you have to normalize and corrupting
> filenames is ok, that doesn't work! It works wonderfully well if you  
> do
> the obvious 1:1 translation and you do *not* normalize, but the  
> moment you
> start normalizing, you actually corrupt the filenames!
Wrong.
Show 15 quoted lines
> And yes, the character sequence 'a¨' is exactly one such sequence.  
> It's
> perfectly representable in both Latin1 and in UTF-8: in latin1 it is a
> two-character '\x61\xa8', and when doing a Latin1->UTF-8 conversion,  
> it
> becomes '\x61\xc2\xa8', and you can convert back and forth between  
> those
> two forms an infinite amount of times, and you never corrupt it.
>
> But the moment you add normalization to the mix, you start screwing  
> up.
> Suddenly, the sequence '\x61\xa8' in Latin1 becomes (assuming NFD)
> '\xc3\xa4' in UTF-8, and when converted back to Latin1, it is now  
> '\xe4',
> ie that filename hass been corrupted!

Wrong. '\x61\x18' in Latin1, when converted to UTF-8 (NFD) is still '\x61\xc2\xa8'. You're mixing up DIARESIS (U+00A8) and COMBINING DIARESIS (U+0308).

I suspect this is why you've been yelling so much - you have a fundamental misunderstanding about what normalization is actually doing.

Show 27 quoted lines
> See? Normalization in the face of working together with others is a  
> total
> and utter mistake, and yes, it really *does* corrupt data. It makes it
> fundamentally impossible to reliably work together with other  
> encodings -
> even when you do converstion between the two!
>
> [ And that's the really sad part. Non-normalized Unicode can pretty  
> much
>  be used as a "generic encoding" for just about all locales - if you  
> know
>  the locale you convert from and to, you can generally use UTF-8 as an
>  internal format, knowing that you can always get the same result  
> back in
>  the original encoding. Normalization literally breaks that wonderful
>  generic capability of Unicode.
>
>  And the fact that Unicode is such a "generic replacement" for any  
> locale
>  is exactly what makes it so wonderful, and allows you to fairly
>  seamlessly convert piece-meal from some particular locale to Unicode:
>  even if you have some programs that still work in the original  
> locale,
>  you know that you can convert back to it without loss of information.
>
>  Except if you normalize. In that case, you *do* lose information, and
>  suddenly one of the best things about Unicode simply disappears.

See above as to why you're not losing the information you so fervently believe you are.

>  As a result, people who force-normalize are idiots. But they seem to
>  also be stupid enough that they don't understand that they are  
> idiots.
>  Sad.

People who insult others run the risk of looking like a fool when shown to be wrong.

Show 12 quoted lines
>  It's a bit like whitespace. Whitespace "doesn't matter" in text (==  
> is
>  equivalent), but an email client that force-normalizes whitespace in
>  text is a really *broken* email client, because it turns out that
>  sometimes even the "equivalent" forms simply do matter. Patches are
>  text, but whitespace is meaningful there.
>
>  Same exact deal: it's good to have the *ability* to normalize
>  whitespace (in email, we call this "text=flowed" or similar), and in
>  some ceses you might even want to make it the default action, but
>  *forcing* normalization is total idiocy and actually makes the system
>  less useful! ]

Sure, it all depends on what level you need to evaluate text. If we're talking about english paragraphs, then whitespace can be messed with. When we're talking about unicode strings, then specific encoding can be messed with. When talking about byte sequence, nothing can be messed with.

In our case, when working on an HFS+ filesystem all you have to care about is the unicode string level. The specific encoding can be messed with, and the client shouldn't care. Problems only arise when attempting to interoperate with filesystems that work at the byte sequence level.

The only information you lose when doing canonical normalization is what the original byte sequence was. Sure, this is a problem when working on a filesystem that cares about byte sequence, but it's not a problem when working on a filesystem that cares about the unicode string.

-Kevin Ballard
-- 
Kevin Ballard
http://kevin.sb.org
kevin@sb.org
http://www.tildesoft.com
Previous: Linus TorvaldsNext: Linus Torvalds
Message 120 of 260 in “git on MacOSX and files with decomposed utf-8 file names”
  1. Mark JunkerJan 16, 2008
  2. Johannes SchindelinJan 16, 2008
  3. Kevin BallardJan 16, 2008
  4. Johannes SchindelinJan 16, 2008
  5. Jakub NarebskiJan 16, 2008
  6. Kevin BallardJan 16, 2008
  7. Jakub NarebskiJan 16, 2008
  8. Kevin BallardJan 16, 2008
  9. Johannes SchindelinJan 16, 2008
  10. Kevin BallardJan 16, 2008
  11. Linus TorvaldsJan 16, 2008
  12. Linus TorvaldsJan 16, 2008
  13. Kevin BallardJan 16, 2008
  14. Linus TorvaldsJan 16, 2008
  15. Pedro MeloJan 16, 2008
  16. Linus TorvaldsJan 17, 2008
  17. Pedro MeloJan 17, 2008
  18. David KastrupJan 17, 2008
  19. Pedro MeloJan 17, 2008
  20. Wincent ColaiutaJan 17, 2008
  21. Johannes SchindelinJan 17, 2008
  22. Linus TorvaldsJan 17, 2008
  23. Kevin BallardJan 17, 2008
  24. Johannes SchindelinJan 17, 2008
  25. Pedro MeloJan 17, 2008
  26. Peter KarlssonJan 18, 2008
  27. Jakub NarebskiJan 18, 2008
  28. David KastrupJan 16, 2008
  29. Linus TorvaldsJan 17, 2008
  30. Kevin BallardJan 17, 2008
  31. Linus TorvaldsJan 17, 2008
  32. Johannes SchindelinJan 17, 2008
  33. Pedro MeloJan 17, 2008
  34. Johannes SchindelinJan 17, 2008
  35. Linus TorvaldsJan 17, 2008
  36. Linus TorvaldsJan 17, 2008
  37. Kevin BallardJan 17, 2008
  38. Linus TorvaldsJan 17, 2008
  39. Kevin BallardJan 17, 2008
  40. Martin LanghoffJan 17, 2008
  41. Kevin BallardJan 17, 2008
  42. Geert BoschJan 17, 2008
  43. Mitch TishmackJan 17, 2008
  44. Wincent ColaiutaJan 17, 2008
  45. Kevin BallardJan 17, 2008
  46. Johannes SchindelinJan 17, 2008
  47. Kevin BallardJan 17, 2008
  48. Robin RosenbergJan 18, 2008
  49. Andrew HeybeyJan 17, 2008
  50. Kevin BallardJan 17, 2008
  51. Kyle MoffettJan 19, 2008
  52. Kevin BallardJan 19, 2008
  53. Wincent ColaiutaJan 17, 2008
  54. Linus TorvaldsJan 17, 2008
  55. Mark JunkerJan 17, 2008
  56. Pedro MeloJan 17, 2008
  57. Johannes SchindelinJan 17, 2008
  58. Mark JunkerJan 17, 2008
  59. Pedro MeloJan 17, 2008
  60. Linus TorvaldsJan 17, 2008
  61. Pedro MeloJan 17, 2008
  62. Linus TorvaldsJan 17, 2008
  63. Mark JunkerJan 17, 2008
  64. Pedro MeloJan 17, 2008
  65. Theodore TsoJan 17, 2008
  66. Linus TorvaldsJan 17, 2008
  67. Kevin BallardJan 18, 2008
  68. Linus TorvaldsJan 18, 2008
  69. Robin RosenbergJan 18, 2008
  70. Linus TorvaldsJan 18, 2008
  71. Brian DessentJan 18, 2008
  72. Dmitry PotapovJan 18, 2008
  73. Robin RosenbergJan 18, 2008
  74. Dmitry PotapovJan 18, 2008
  75. Peter KarlssonJan 18, 2008
  76. Jakub NarebskiJan 18, 2008
  77. Peter KarlssonJan 18, 2008
  78. Dmitry PotapovJan 18, 2008
  79. Peter KarlssonJan 18, 2008
  80. Linus TorvaldsJan 18, 2008
  81. Kevin BallardJan 18, 2008
  82. Dmitry PotapovJan 19, 2008
  83. Kevin BallardJan 19, 2008
  84. Dmitry PotapovJan 19, 2008
  85. Linus TorvaldsJan 19, 2008
  86. Mark JunkerJan 19, 2008
  87. Johannes SchindelinJan 19, 2008
  88. Dmitry PotapovJan 20, 2008
  89. Linus TorvaldsJan 20, 2008
  90. Johannes SchindelinJan 20, 2008
  91. Wincent ColaiutaJan 20, 2008
  92. Linus TorvaldsJan 20, 2008
  93. Mike HommeyJan 20, 2008
  94. Linus TorvaldsJan 20, 2008
  95. Mike HommeyJan 20, 2008
  96. Linus TorvaldsJan 20, 2008
  97. Dmitry PotapovJan 20, 2008
  98. Dmitry PotapovJan 20, 2008
  99. Wincent ColaiutaJan 20, 2008
  100. Junio C HamanoJan 18, 2008
  101. Johannes SchindelinJan 18, 2008
  102. Eric W. BiedermanJan 23, 2008
  103. Junio C HamanoJan 23, 2008
  104. Nicolas PitreJan 23, 2008
  105. Junio C HamanoJan 23, 2008
  106. Peter KarlssonJan 21, 2008
  107. Kevin BallardJan 21, 2008
  108. David KastrupJan 21, 2008
  109. Kevin BallardJan 21, 2008
  110. Dmitry PotapovJan 21, 2008
  111. Kevin BallardJan 21, 2008
  112. David KastrupJan 21, 2008
  113. Dmitry PotapovJan 21, 2008
  114. Jeff KingJan 21, 2008
  115. Nicolas PitreJan 21, 2008
  116. Kevin BallardJan 21, 2008
  117. David KastrupJan 21, 2008
  118. David KastrupJan 21, 2008
  119. Linus TorvaldsJan 21, 2008
  120. Kevin BallardJan 21, 2008
  121. Linus TorvaldsJan 21, 2008
  122. Kevin BallardJan 21, 2008
  123. Linus TorvaldsJan 21, 2008
  124. Kevin BallardJan 21, 2008
  125. David KastrupJan 21, 2008
  126. Martin LanghoffJan 21, 2008
  127. Kevin BallardJan 21, 2008
  128. Martin LanghoffJan 21, 2008
  129. Linus TorvaldsJan 21, 2008
  130. Kevin BallardJan 21, 2008
  131. Linus TorvaldsJan 21, 2008
  132. Kevin BallardJan 21, 2008
  133. Martin LanghoffJan 21, 2008
  134. Theodore TsoJan 21, 2008
  135. Kevin BallardJan 21, 2008
  136. Linus TorvaldsJan 21, 2008
  137. Kevin BallardJan 22, 2008
  138. Linus TorvaldsJan 22, 2008
  139. Linus TorvaldsJan 22, 2008
  140. Kevin BallardJan 22, 2008
  141. Linus TorvaldsJan 22, 2008
  142. Kevin BallardJan 22, 2008
  143. Linus TorvaldsJan 22, 2008
  144. Martin LanghoffJan 22, 2008
  145. Kevin BallardJan 22, 2008
  146. Theodore TsoJan 21, 2008
  147. Kevin BallardJan 21, 2008
  148. Theodore TsoJan 21, 2008
  149. Kevin BallardJan 21, 2008
  150. Theodore TsoJan 21, 2008
  151. Kevin BallardJan 21, 2008
  152. Dmitry PotapovJan 21, 2008
  153. Kevin BallardJan 21, 2008
  154. Dmitry PotapovJan 21, 2008
  155. Kevin BallardJan 21, 2008
  156. Dmitry PotapovJan 21, 2008
  157. Mike HommeyJan 21, 2008
  158. Dmitry PotapovJan 21, 2008
  159. Martin LanghoffJan 21, 2008
  160. David KastrupJan 21, 2008
  161. Linus TorvaldsJan 21, 2008
  162. Martin LanghoffJan 21, 2008
  163. Dmitry PotapovJan 21, 2008
  164. Linus TorvaldsJan 21, 2008
  165. Dmitry PotapovJan 17, 2008
  166. JM IbanezJan 17, 2008
  167. Johannes SchindelinJan 17, 2008
  168. Robin RosenbergJan 18, 2008
  169. Linus TorvaldsJan 17, 2008
  170. Dmitry PotapovJan 17, 2008
  171. Dmitry PotapovJan 16, 2008
  172. Eyvind BernhardsenJan 16, 2008
  173. Wincent ColaiutaJan 16, 2008
  174. Miles BaderJan 17, 2008
  175. Jay SoffianJan 17, 2008
  176. Jay SoffianJan 17, 2008
  177. Junio C HamanoJan 17, 2008
  178. Wincent ColaiutaJan 17, 2008
  179. Johannes SchindelinJan 17, 2008
  180. Pedro MeloJan 17, 2008
  181. Wincent ColaiutaJan 17, 2008
  182. Johannes SchindelinJan 17, 2008
  183. Wincent ColaiutaJan 17, 2008
  184. Junio C HamanoJan 17, 2008
  185. Johan HerlandJan 17, 2008
  186. Johannes SchindelinJan 17, 2008
  187. Wincent ColaiutaJan 17, 2008
  188. Linus TorvaldsJan 17, 2008
  189. Theodore TsoJan 21, 2008
  190. Kevin BallardJan 21, 2008
  191. Martin LanghoffJan 21, 2008
  192. Kevin BallardJan 21, 2008
  193. Johannes SchindelinJan 22, 2008
  194. Kevin BallardJan 22, 2008
  195. David KastrupJan 22, 2008
  196. Martin LanghoffJan 22, 2008
  197. Johannes SchindelinJan 22, 2008
  198. Martin LanghoffJan 22, 2008
  199. Johannes SchindelinJan 22, 2008
  200. David KastrupJan 21, 2008
  201. Kevin BallardJan 22, 2008
  202. David KastrupJan 22, 2008
  203. Kevin BallardJan 21, 2008
  204. Martin LanghoffJan 21, 2008
  205. Kevin BallardJan 22, 2008
  206. Theodore TsoJan 23, 2008
  207. Kevin BallardJan 23, 2008
  208. Martin LanghoffJan 23, 2008
  209. Theodore TsoJan 23, 2008
  210. David KastrupJan 23, 2008
  211. Linus TorvaldsJan 23, 2008
  212. Martin LanghoffJan 23, 2008
  213. Kevin BallardJan 23, 2008
  214. Martin LanghoffJan 23, 2008
  215. Theodore TsoJan 23, 2008
  216. Linus TorvaldsJan 23, 2008
  217. Kevin BallardJan 23, 2008
  218. Mike HommeyJan 23, 2008
  219. Kevin BallardJan 23, 2008
  220. Dmitry PotapovJan 23, 2008
  221. Jonathan del StrotherJan 23, 2008
  222. Dmitry PotapovJan 23, 2008
  223. Mike HommeyJan 23, 2008
  224. Dmitry PotapovJan 23, 2008
  225. Mike HommeyJan 23, 2008
  226. Theodore TsoJan 23, 2008
  227. Linus TorvaldsJan 23, 2008
  228. Theodore TsoJan 23, 2008
  229. Kevin BallardJan 23, 2008
  230. Linus TorvaldsJan 23, 2008
  231. On pathnamesJunio C Hamano, Jan 24, 2008
  232. Nicolas PitreJan 24, 2008
  233. Martin LanghoffJan 25, 2008
  234. Junio C HamanoJan 25, 2008
  235. Junio C HamanoJan 25, 2008
  236. Pedro MeloJan 25, 2008
  237. Johannes SchindelinJan 25, 2008
  238. David KastrupJan 25, 2008
  239. Wincent ColaiutaJan 25, 2008
  240. SeanJan 24, 2008
  241. Johannes SchindelinJan 25, 2008
  242. Daniel BarkalowJan 25, 2008
  243. Junio C HamanoJan 25, 2008
  244. Johannes SchindelinJan 25, 2008
  245. Daniel BarkalowJan 25, 2008
  246. Johannes SchindelinJan 25, 2008
  247. Jeff KingJan 25, 2008
  248. Jay SoffianJan 23, 2008
  249. Martin LanghoffJan 23, 2008
  250. Kevin BallardJan 23, 2008
  251. Dmitry PotapovJan 23, 2008
  252. Kevin BallardJan 23, 2008
  253. Kevin BallardJan 24, 2008
  254. Junio C HamanoJan 24, 2008
  255. Martin LanghoffJan 24, 2008
  256. Kevin BallardJan 24, 2008
  257. Steffen ProhaskaJan 24, 2008
  258. Mitch TishmackJan 24, 2008
  259. Mitch TishmackJan 24, 2008
  260. Kevin BallardJan 24, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.