{"thread":{"id":"11645","subject":"git on MacOSX and files with decomposed utf-8 file names","startedAt":"2008-01-16T15:17:01Z","lastAt":"2008-01-25T17:34:47Z","messageCount":260,"participants":["Mark Junker","Johannes Schindelin","Kevin Ballard","Jakub Narebski","Linus Torvalds","Wincent Colaiuta","Eyvind Bernhardsen","Dmitry Potapov","Pedro Melo","David Kastrup","Jay Soffian","Martin Langhoff","Junio C Hamano","Geert Bosch","Mitch Tishmack","Miles Bader","Johan Herland","Theodore Tso","JM Ibanez","Robin Rosenberg","Brian Dessent","Andrew Heybey","Peter Karlsson","Kyle Moffett","Mike Hommey","Jeff King","Nicolas Pitre","Eric W. Biederman","Jonathan del Strother","Steffen Prohaska","Sean","Daniel Barkalow"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"65546","messageId":"478E1FED.5010801@web.de","threadId":"11645","inReplyTo":null,"subject":"git on MacOSX and files with decomposed utf-8 file names","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-16T15:17:01Z","receivedAt":"2008-01-16T15:17:01Z","isPatch":false,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Hi,\n\nI have some files like \"Lüftung.txt\" in my repository. The strange thing \nis that I can pull / add / commit / push those files without problem but \ngit-status always complains that thoes files are untraced (but not \nmissing). My assumption is that it's a problem with the way MacOSX \nstores the file names (decomposed UTF-8). So something like \n\"Lüftung.txt\" becomes \"Lüftung.txt\".\n\nIt seems that git-status does two things:\n1. Find files under version control (i.e. search for missing files)\n2. Find files not under version control (i.e. search for untracked files)\n\nI guess that the first look-up succeeds because MacOS X converts \ncomposed UTF-8 to decomposed UTF-8 when searching for a file. But it \nseems that the second look-up takes the file names as-is (decomposed) \nwithout converting them to composed UTF-8.\n\nIs there an easy way to fix this behaviour? It's really annoying to see \nall those \"untracked\" files that are already under version control when \nexecuting a git-status.\n\nRegards,\nMark\n"},{"id":"65548","messageId":"alpine.LSU.1.00.0801161531030.17650@racer.site","threadId":"11645","inReplyTo":"478E1FED.5010801@web.de","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-16T15:34:18Z","receivedAt":"2008-01-16T15:34:18Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 16 Jan 2008, Mark Junker wrote:\n\n> I have some files like \"Lüftung.txt\" in my repository. The strange thing is\n> that I can pull / add / commit / push those files without problem but\n> git-status always complains that thoes files are untraced (but not missing).\n\nThis is a known problem.  Unfortunately, noone has implemented a fix, \nalthough if you're serious about it, I can point you to threads where it \nhas been hinted how to solve the issue.\n\nFWIW the issue is that Mac OS X decides that it knows better how to encode \nyour filename than you could yourself.\n\nCiao,\nDscho"},{"id":"65550","messageId":"427BE4FD-6534-4CB2-91F8-F9014DC82B54@sb.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801161531030.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-16T15:43:24Z","receivedAt":"2008-01-16T15:43:24Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 10:34 AM, Johannes Schindelin wrote:\n\n> On Wed, 16 Jan 2008, Mark Junker wrote:\n>\n>> I have some files like \"Lüftung.txt\" in my repository. The strange  \n>> thing is\n>> that I can pull / add / commit / push those files without problem but\n>> git-status always complains that thoes files are untraced (but not  \n>> missing).\n>\n> This is a known problem.  Unfortunately, noone has implemented a fix,\n> although if you're serious about it, I can point you to threads  \n> where it\n> has been hinted how to solve the issue.\n>\n> FWIW the issue is that Mac OS X decides that it knows better how to  \n> encode\n> your filename than you could yourself.\n\n\nMore like, Mac OS X has standardized on Unicode and the rest of the  \nworld hasn't caught up yet. Git is the only tool I've ever heard of  \nthat has a problem with OS X using Unicode.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65556","messageId":"alpine.LSU.1.00.0801161629580.17650@racer.site","threadId":"11645","inReplyTo":"427BE4FD-6534-4CB2-91F8-F9014DC82B54@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-16T16:32:32Z","receivedAt":"2008-01-16T16:32:32Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 16 Jan 2008, Kevin Ballard wrote:\n\n> On Jan 16, 2008, at 10:34 AM, Johannes Schindelin wrote:\n> \n> > On Wed, 16 Jan 2008, Mark Junker wrote:\n> > \n> > > I have some files like \"Lüftung.txt\" in my repository. The strange \n> > > thing is that I can pull / add / commit / push those files without \n> > > problem but git-status always complains that thoes files are \n> > > untraced (but not missing).\n> > \n> > This is a known problem.  Unfortunately, noone has implemented a fix, \n> > although if you're serious about it, I can point you to threads where \n> > it has been hinted how to solve the issue.\n> > \n> > FWIW the issue is that Mac OS X decides that it knows better how to \n> > encode your filename than you could yourself.\n> \n> More like, Mac OS X has standardized on Unicode and the rest of the \n> world hasn't caught up yet. Git is the only tool I've ever heard of that \n> has a problem with OS X using Unicode.\n\nNo.  That's not at all the problem.  Mac OS X insists on storing _another_ \nencoding of your filename.  Both are UTF-8.  Both encode the _same_ \nstring.  Yet they are different, bytewise.  For no good reason.\n\nStop spreading FUD.  Git can handle Unicode just fine.  In fact, Git does \nnot _care_ how the filename is encoded, it _respects_ the user's choice, \nnot only of the encoding _type_, but the _encoding_, too.\n\nOkay?\n\nHth,\nDscho\n"},{"id":"65559","messageId":"m33asxn2gt.fsf@roke.D-201","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801161629580.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-01-16T16:46:29Z","receivedAt":"2008-01-16T16:46:29Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> On Wed, 16 Jan 2008, Kevin Ballard wrote:\n> \n> > On Jan 16, 2008, at 10:34 AM, Johannes Schindelin wrote:\n> > \n> > > On Wed, 16 Jan 2008, Mark Junker wrote:\n> > > \n> > > > I have some files like \"Lüftung.txt\" in my repository. The strange \n> > > > thing is that I can pull / add / commit / push those files without \n> > > > problem but git-status always complains that thoes files are \n> > > > untraced (but not missing).\n> > > \n> > > This is a known problem.  Unfortunately, noone has implemented a fix, \n> > > although if you're serious about it, I can point you to threads where \n> > > it has been hinted how to solve the issue.\n> > > \n> > > FWIW the issue is that Mac OS X decides that it knows better how to \n> > > encode your filename than you could yourself.\n> > \n> > More like, Mac OS X has standardized on Unicode and the rest of the \n> > world hasn't caught up yet. Git is the only tool I've ever heard of that \n> > has a problem with OS X using Unicode.\n> \n> No.  That's not at all the problem.  Mac OS X insists on storing _another_ \n> encoding of your filename.  Both are UTF-8.  Both encode the _same_ \n> string.  Yet they are different, bytewise.  For no good reason.\n\nTo be more exact encoding used to _create_ file differs from encoding\nreturned when _reading directory_... \n \n> Stop spreading FUD.  Git can handle Unicode just fine.  In fact, Git does \n> not _care_ how the filename is encoded, it _respects_ the user's choice, \n> not only of the encoding _type_, but the _encoding_, too.\n\n...which means that sequence of bytes differ. And Git by design is\n(both for filenames and for blob contents) encoding agnostic.\n\nHFS+ is just _stupid_. And unfortunately Git doesn't support stupid\nfilesystems (e.g. case insensitive filesystems) well.\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"65602","messageId":"65026F2B-5CE8-4238-A9AB-D3545D336B41@sb.org","threadId":"11645","inReplyTo":"m33asxn2gt.fsf@roke.D-201","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-16T20:39:36Z","receivedAt":"2008-01-16T20:39:36Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 11:46 AM, Jakub Narebski wrote:\n\n>>> More like, Mac OS X has standardized on Unicode and the rest of the\n>>> world hasn't caught up yet. Git is the only tool I've ever heard  \n>>> of that\n>>> has a problem with OS X using Unicode.\n>>\n>> No.  That's not at all the problem.  Mac OS X insists on storing  \n>> _another_\n>> encoding of your filename.  Both are UTF-8.  Both encode the _same_\n>> string.  Yet they are different, bytewise.  For no good reason.\n>\n> To be more exact encoding used to _create_ file differs from encoding\n> returned when _reading directory_...\n>\n>> Stop spreading FUD.  Git can handle Unicode just fine.  In fact,  \n>> Git does\n>> not _care_ how the filename is encoded, it _respects_ the user's  \n>> choice,\n>> not only of the encoding _type_, but the _encoding_, too.\n>\n> ...which means that sequence of bytes differ. And Git by design is\n> (both for filenames and for blob contents) encoding agnostic.\n>\n> HFS+ is just _stupid_. And unfortunately Git doesn't support stupid\n> filesystems (e.g. case insensitive filesystems) well.\n\nThere's two different ways to do filesystem encodings. One is to have  \nthe fs simply not care about encoding, which is what the linux world  \nseems to prefer. Sure, this is great in that what you create the file  \nwith is what you get back, but on the other hand, given an arbitrary  \nnon-ASCII file on disk, you have absolutely no idea what the encoding  \nshould be and you can't display it without making assumptions (yes you  \ncan use heuristics, but you're still making assumptions). Filesystems  \nlike HFS+ that standardize the encoding, on the other hand, make it  \nsuch that you always know what the encoding of a file should be, so  \nyou can always display and use the filename intelligently. It also  \nmeans it plays much nicer in a non-ASCII world, since you don't have  \nto worry about different normalizations of a given string referring to  \ndifferent files (it's one thing to be case-sensitive, but claiming  \nthat \"föo\" and \"föo\" are different files just because one uses a  \ncomposed character and the other doesn't is extremely user- \nunfriendly). On the other hand, what you create the file with may not  \nbe what you read back later, since the name has been standardized.  \nIt's hard to say one is better than the other, they're just different  \nways of doing it. However, I have noticed that everybody who's voiced  \nan opinion on this list in favor of the encoding-agnostic approach  \nseem to be unwilling to accept that any other approach might have  \nvalidity, to the extent of calling an OS/filesystem that does things  \ndifferent stupid or insane. This strikes me as extremely elitist and  \nrisks alienating what I expect to be a fast-growing group of users  \n(i.e. OS X users).\n\nI'm willing to give Linus a free pass on calling other OS's stupid and  \ninsane, as I don't think Linux would exist as it does today without  \nhis strong opinions, but I don't think this should give carte blanche  \nto the rest of the community for this inflammatory behavior.\n\nI should note that I'm only taking the time to discuss this because,  \ndespite the fact that I'm new to git, I really like it and I want it  \nto work better. And one area that it has a problem with is the de- \nfacto filesystem on my OS of choice. However, attempts to discuss the  \nproblem invariable end up with multiple people calling my OS stupid  \nand insane simply because it differs in a particular design decision.  \nThis is not a good way to build a community or to build a better  \nproduct, and I hope it can be improved.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65608","messageId":"200801162251.54219.jnareb@gmail.com","threadId":"11645","inReplyTo":"65026F2B-5CE8-4238-A9AB-D3545D336B41@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-01-16T21:51:53Z","receivedAt":"2008-01-16T21:51:53Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Wed, 16 Jan 2008, Kevin Ballard wrote:\n> On Jan 16, 2008, at 11:46 AM, Jakub Narebski wrote:\n>>>> More like, Mac OS X has standardized on Unicode and the rest of the\n>>>> world hasn't caught up yet. Git is the only tool I've ever heard  \n>>>> which has a problem with OS X using Unicode.\n>>>\n>>> No.  That's not at all the problem.  Mac OS X insists on storing  \n>>> _another_  encoding of your filename.  Both are UTF-8.  Both encode\n>>> the _same_ string.  Yet they are different, bytewise.  For no good\n>>> reason. \n>>\n>> To be more exact encoding used to _create_ file differs from encoding\n>> returned when _reading directory_...\n>>\n>>> Stop spreading FUD.  Git can handle Unicode just fine.  In fact,  \n>>> Git does not _care_ how the filename is encoded, it _respects_ the\n>>> user's choice, not only of the encoding _type_, but the _encoding_,\n>>> too. \n>>\n>> ...which means that sequence of bytes differ. And Git by design is\n>> (both for filenames and for blob contents) encoding agnostic.\n>>\n>> HFS+ is just _stupid_. And unfortunately Git doesn't support stupid\n>> filesystems (e.g. case insensitive filesystems) well.\n\nBy the way, calling HFS+ stupid, or rather calling at least two \ndifferent normalizations of UTF-8 (two different encodings) used for \nwriting and reading filenames stupid is wrong _for me_. I have quoted \nLinus here, when I think I should use other description.\n \n> There's two different ways to do filesystem encodings. One is to have  \n> the fs simply not care about encoding, which is what the linux world  \n> seems to prefer. Sure, this is great in that what you create the file  \n> with is what you get back, but on the other hand, given an arbitrary  \n> non-ASCII file on disk, you have absolutely no idea what the encoding  \n> should be and you can't display it without making assumptions (yes you  \n> can use heuristics, but you're still making assumptions). Filesystems  \n> like HFS+ that standardize the encoding, on the other hand, make it  \n> such that you always know what the encoding of a file should be, so  \n> you can always display and use the filename intelligently. It also  \n> means it plays much nicer in a non-ASCII world, since you don't have  \n> to worry about different normalizations of a given string referring to  \n> different files (it's one thing to be case-sensitive, but claiming  \n> that \"föo\" and \"föo\" are different files just because one uses a  \n> composed character and the other doesn't is extremely user- \n> unfriendly).\n\nFor me it looks like a layering violation... but my knowledge about \nfilesystem is cluse to nil. IMHO it is VFS and libc which should do the \ntranslating.\n\n> On the other hand, what you create the file with may not   \n> be what you read back later, since the name has been standardized.  \n> It's hard to say one is better than the other, they're just different  \n> ways of doing it.\n\nBut using one encoding to create file, and another when reding filenames \nis strange. It is IMHO better to simply refuse creating filenames which \nare outside chosen encoding / normalization. But having different \nencodings used for reading and writing on the level of filesystem \naccess (not on level of UI) is strange.\n\n> However, I have noticed that everybody who's voiced   \n> an opinion on this list in favor of the encoding-agnostic approach  \n> seem to be unwilling to accept that any other approach might have  \n> validity, to the extent of calling an OS/filesystem that does things  \n> different stupid or insane. This strikes me as extremely elitist and  \n> risks alienating what I expect to be a fast-growing group of users  \n> (i.e. OS X users).\n\nFirst, it is Git philosophy and very core of design to be encoding \nagnostic (to be \"content tracker\"). Second, using the same sequence of \nbytes on filesystem, in the index, and in 'tree' objects ensures good \nperformance... this is something to think about if you want to add \npatches which would deal with HFS+ API/UI quirks.\n\n[cut]\n-- \nJakub Narebski\nPoland\n"},{"id":"65611","messageId":"1574A90A-8C45-46AD-9402-34AE6F582B3F@sb.org","threadId":"11645","inReplyTo":"200801162251.54219.jnareb@gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-16T22:06:05Z","receivedAt":"2008-01-16T22:06:05Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 4:51 PM, Jakub Narebski wrote:\n\n>> On the other hand, what you create the file with may not\n>> be what you read back later, since the name has been standardized.\n>> It's hard to say one is better than the other, they're just different\n>> ways of doing it.\n>\n> But using one encoding to create file, and another when reding  \n> filenames\n> is strange. It is IMHO better to simply refuse creating filenames  \n> which\n> are outside chosen encoding / normalization. But having different\n> encodings used for reading and writing on the level of filesystem\n> access (not on level of UI) is strange.\n\nIt's not using different encodings, it's all Unicode. However, it  \naccepts different normalization variants of Unicode, since it can read  \nthem all and it would be folly to require everybody to conform to its  \nown special internal variant. But it does have to normalize them,  \notherwise how would it detect the same filename using different  \nnormalizations? Also, it may seem strange to have different names  \nbetween reading and writing, but that's only if you think of the name  \nas a sequence of bytes - when treated as a sequence of characters, you  \nget the same result. In other words, you're used to filenames as  \nbytes, HFS+ treats filenames as strings.\n\n>> However, I have noticed that everybody who's voiced\n>> an opinion on this list in favor of the encoding-agnostic approach\n>> seem to be unwilling to accept that any other approach might have\n>> validity, to the extent of calling an OS/filesystem that does things\n>> different stupid or insane. This strikes me as extremely elitist and\n>> risks alienating what I expect to be a fast-growing group of users\n>> (i.e. OS X users).\n>\n> First, it is Git philosophy and very core of design to be encoding\n> agnostic (to be \"content tracker\"). Second, using the same sequence of\n> bytes on filesystem, in the index, and in 'tree' objects ensures good\n> performance... this is something to think about if you want to add\n> patches which would deal with HFS+ API/UI quirks.\n\nSure, it makes sense from a performance perspective, but it causes  \nproblems with HFS+ and any other filesystem that behaves the same way.  \nIn the previous discussion about case-sensitivity, somebody suggested  \nusing a lookup table to map between git's internal representation and  \nthe name the filesystem returns, which seems like a decent idea and  \none that could be enabled with a config parameter to avoid penalizing  \nrepos on other filesystems. But I don't know enough about the  \ninternals of git to even think of trying to implement it myself.\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65612","messageId":"alpine.LSU.1.00.0801162222180.17650@racer.site","threadId":"11645","inReplyTo":"1574A90A-8C45-46AD-9402-34AE6F582B3F@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-16T22:23:17Z","receivedAt":"2008-01-16T22:23:17Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 16 Jan 2008, Kevin Ballard wrote:\n\n> It's not using different encodings, it's all Unicode.\n\nBut that's the _point_!  It _is_ Unicode, yet it uses _different_ \nencodings of the _same_ string.\n\nNow, this discussion gets really annoying.  The real question is: will you \ndo something about it, or reply with another 500-line email?\n\nCiao,\nDscho\n"},{"id":"65613","messageId":"alpine.LFD.1.00.0801161424040.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"1574A90A-8C45-46AD-9402-34AE6F582B3F@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-16T22:32:38Z","receivedAt":"2008-01-16T22:32:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 16 Jan 2008, Kevin Ballard wrote:\n> \n> It's not using different encodings, it's all Unicode. However, it accepts\n> different normalization variants of Unicode, since it can read them all and it\n> would be folly to require everybody to conform to its own special internal\n> variant. But it does have to normalize them, otherwise how would it detect the\n> same filename using different normalizations?\n\nThat's a singularly *stupid* argument.\n\nHere, let me rephrase that same idiotic argument:\n\n  \"But it does have to uppercase them, otherwise how would it detect the \n   same filename using different cases?\"\n\n..and if you don't see how that's *exactly* the same argument, you really \nare stupid.\n\nThe fact is, normalization is wrong.\n\nIt's wrong when you normalize upper/lower case (no, the word \"Polish\" is \nnot the same as \"polish\"), and it's equally wrong when you normalize for \n\"looks similar\".\n\n> In other words, you're used to filenames as bytes, HFS+ treats filenames \n> as strings.\n\nNo. HFS+ treats users as idiots and thinks that it should \"fix\" the \nfilename for them. And it causes problems.\n\nIt causes problems for exactly the same reasons case-independence causes \nproblems, because it's EXACTLY THE SAME ISSUE. People may think that \"but \nthey are the same\", but they aren't. Case matters. And so does \"single \ncharacter\" vs \"two character overlay\". \n\nDoes it always matter? Hell no. But the problem with a filesystem that \nthinks it knows better is that when it *sometimes* matters, the filesystem \nsimply DOES THE WRONG THING.\n\nCan't you understand that?\n\n\t\t\tLinus\n"},{"id":"65623","messageId":"24A6A8D6-EED9-4E30-AE4C-125F98A7F6A3@orakel.ntnu.no","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801161629580.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Eyvind Bernhardsen","fromEmail":"eyvind-git@orakel.ntnu.no","sentAt":"2008-01-16T22:37:58Z","receivedAt":"2008-01-16T22:37:58Z","isPatch":false,"sender":{"key":"eyvind.bernhardsen@gmail.com","avatar":"https://avatars.githubusercontent.com/u/106762?v=4"},"body":"On 16. jan.. 2008, at 17.32, Johannes Schindelin wrote:\n\n>>> FWIW the issue is that Mac OS X decides that it knows better how to\n>>> encode your filename than you could yourself.\n>>\n>> More like, Mac OS X has standardized on Unicode and the rest of the\n>> world hasn't caught up yet. Git is the only tool I've ever heard of  \n>> that\n>> has a problem with OS X using Unicode.\n>\n> No.  That's not at all the problem.  Mac OS X insists on storing  \n> _another_\n> encoding of your filename.  Both are UTF-8.  Both encode the _same_\n> string.  Yet they are different, bytewise.  For no good reason.\n>\n> Stop spreading FUD.  Git can handle Unicode just fine.  In fact, Git  \n> does\n> not _care_ how the filename is encoded, it _respects_ the user's  \n> choice,\n> not only of the encoding _type_, but the _encoding_, too.\n\n\"FUD\" is a bit strong, don't you think?  HFS+ is the way it is and it  \nwould be nice if Git could deal with it.\n\nThe problem is that HFS+ normalizes filenames to avoid multiple files  \nthat appear to have the same name (eg \"M<A WITH UMLAUT>rchen\" vs  \n\"Ma<UMLAUT MODIFIER>rchen\", in gitweb/test).  This is sort of like  \ncase sensitivity, but filenames are normalized when a file is  \n_created_.  Git, not unreasonably, expects a file to keep the name it  \nwas created with.\n\nAs far as I can tell, as long as you add all your internationally  \nbecharactered files to git from an HFS+ file system using a gui or  \ncommand-line completion, you'll be okay; trouble starts when you check  \nin a file with the composed form of a character, by typing the name on  \nthe command line (I'm not sure about this one) or committing on  \nanother OS.  Git will store the filename in composed form, but the  \nMac's filesystem will decompose the filename when you check the file  \nout.\n\nThe result looks like this:\n\nvredefort:[git]% git status\n# On branch master\n# Untracked files:\n#   (use \"git add <file>...\" to include in what will be committed)\n#\n#\tgitweb/test/Märchen\nnothing added to commit but untracked files present (use \"git add\" to  \ntrack)\n\n(this is directly after checking out git.git @ v1.5.4-rc3)\n\nThere are two things to note here.  One is that Git thinks that there  \nis a new file called \"gitweb/test/Märchen\" (decomposed) when it's  \n\"really\" just the same \"gitweb/test/Märchen\" (precomposed) that's in  \nthe repository.  The other is that git _thinks_ that the \"gitweb/test/ \nMärchen\" (precomposed) it's expecting is still there, because the  \nfilesystem, when asked for \"gitweb/test/Märchen\" in any form will  \nreturn the file \"gitweb/test/Märchen\" (decomposed).\n\nTrying to check out the \"next\" branch at this point is a pain since  \nnext's \"Märchen\" would overwrite the untracked \"Märchen\".\n\nI can't provide links to any previous discussions about this, but  \nhere's Apple's Technical Q&A on the subject:\n\nhttp://developer.apple.com/qa/qa2001/qa1235.html\n\nFinding a sane way of allowing git to handle this behaviour is left as  \nan exercise for the reader.\n\nEyvind Bernhardsen\n"},{"id":"65618","messageId":"alpine.LFD.1.00.0801161433380.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161424040.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-16T22:52:32Z","receivedAt":"2008-01-16T22:52:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 16 Jan 2008, Linus Torvalds wrote:\n> \n> Does it always matter? Hell no. But the problem with a filesystem that \n> thinks it knows better is that when it *sometimes* matters, the filesystem \n> simply DOES THE WRONG THING.\n> \n> Can't you understand that?\n\nSide note: there are ways to do it right.\n\nYou can:\n\n - not do conversion at all (which is always right). Not corrupting the \n   user data means that the user never gets something back that he didn't \n   put in\n\n   (And, btw, the \"security\" argument is total BS. The fact that two \n   characters look the same does not mean that they should act the same, \n   and it is *not* a security feature. Quite the reverse. Having programs \n   that get different results back from what they actually wrote, *that* \n   tends to be a security issue, because now you have a confused program, \n   and I guarantee that there are more bugs in unexpected cases than in \n   the expected ones)\n\n - Not accept data in formats that you don't like. This is also always \n   right, but can be rather impolite.\n\n - Not accept data in formats that you don't like, and give people \n   explicit conversion and comparison routines so that they can then make \n   their own decisions and they are *aware* of the conversion (so that \n   they don't come back to the problem of being confused)\n\nSo there are certainly many ways to handle things like this.\n\nThe one thing you shouldn't do is to silently convert data behind the \nprograms back, without even giving any way to disable it (and that disable \nhas to be on a use-by-use casis, not some \"disable/enable for all users of \nthis filesystem\", because you can - and do - have different programs that \nhave different expectations).\n\nAnd finally: all of the above is true at *all* levels. It doesn't matter \none whit whether the automatic conversion conversion is in the kernel or \nin a library. Doing it on a library level has advantages (namely the whole \n\"disable/enable\" thing tends to get *much* easier to do, and applications \ncan decide to link against a particular version to get the behaviour \n*they* want, for example).\n\nSo doing it inside the kernel is just about the worst possible case, \nexactly because it makes it really hard to do a \"on a case-by-case\" basis. \n\nYes, Linux does it too, but it does it only for filesystems that are \n*defined* to be insane. OS X really should have known better. Especially \nsince they already fixed the applications (ie they do allow for \ncase-sensitive filesystems).\n\nI can understand normalization when it's about case-insensitivity (there \nare lots of _technical_ reasons to do it there), but once you let the \ncase-insensitivity go, there just isn't any excuse any more.\n\n\t\t\tLinus\n"},{"id":"65622","messageId":"BB900023-A5D7-4ECA-95A6-8334F7976224@wincent.com","threadId":"11645","inReplyTo":"427BE4FD-6534-4CB2-91F8-F9014DC82B54@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-16T23:03:51Z","receivedAt":"2008-01-16T23:03:51Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 16/1/2008, a las 16:43, Kevin Ballard escribió:\n\n> On Jan 16, 2008, at 10:34 AM, Johannes Schindelin wrote:\n>\n>> On Wed, 16 Jan 2008, Mark Junker wrote:\n>>\n>>> I have some files like \"Lüftung.txt\" in my repository. The strange  \n>>> thing is\n>>> that I can pull / add / commit / push those files without problem  \n>>> but\n>>> git-status always complains that thoes files are untraced (but not  \n>>> missing).\n>>\n>> This is a known problem.  Unfortunately, noone has implemented a fix,\n>> although if you're serious about it, I can point you to threads  \n>> where it\n>> has been hinted how to solve the issue.\n>>\n>> FWIW the issue is that Mac OS X decides that it knows better how to  \n>> encode\n>> your filename than you could yourself.\n>\n> More like, Mac OS X has standardized on Unicode and the rest of the  \n> world hasn't caught up yet. Git is the only tool I've ever heard of  \n> that has a problem with OS X using Unicode.\n\nAs far as I know, Subversion has basically exactly the same problem,  \nand any time you consume/produce files on Mac OS X that are be  \nconsumed/produced on other platforms you will run into this kind of  \nissue, with any software.\n\nTell Mac OS X to write a file with \"ó\" in the file name (\"\\xc3\\xb3\" in  \nUTF-8), and it will \"normalize\" it prior to writing by converting it  \ninto a decomposed form (that is, ASCII \"o\" followed by \"\\xcc\\x81\", or  \n\"combining acute accent\"). So they're both valid Unicode, both valid  \nUTF-8, and they encode exactly the same characters but the byte stream  \nis different.\n\nIf you only work on Mac OS X then this will never be a problem because  \nall the files you create and therefore all the files you add to your  \nGit repository will have their names in decomposed UTF-8. But when you  \nstart cloning repositories containing files added on other systems,  \nsystems which might use precomposed rather than decomposed UTF-8 then  \nyou'll run into exactly this kind of problem. The git.git repo has one  \nsuch file itself (gitweb/test/Märchen, if I remember correctly, which  \nGit reports as untracked).\n\nNow, Mac OS X's behaviour is not entirely \"insane\" as some would  \nclaim; there is indeed a rationale behind it even if you don't agree  \nwith it, but it *does* produce some unfortunate teething problems for  \npeople wanting to use Mac OS X in a cross-platform environment.\n\nHere are some Apple docs on the subject:\n\nhttp://developer.apple.com/qa/qa2001/qa1173.html\n\nhttp://developer.apple.com/qa/qa2001/qa1235.html\n\nI personally wish that UTF-8 didn't allow different normalization  \nforms; then this kind of problem wouldn't arise. But it has arisen and  \nwe have to live with it. Some workarounds have been proposed for Git,  \nbut I haven't seen any convincing proposals yet.\n\nCheers,\nWincent\n"},{"id":"65625","messageId":"7652B11D-9B9F-45EA-9465-8294B701FE7C@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161424040.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-16T23:11:31Z","receivedAt":"2008-01-16T23:11:31Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 5:32 PM, Linus Torvalds wrote:\n\n> On Wed, 16 Jan 2008, Kevin Ballard wrote:\n>>\n>> It's not using different encodings, it's all Unicode. However, it  \n>> accepts\n>> different normalization variants of Unicode, since it can read them  \n>> all and it\n>> would be folly to require everybody to conform to its own special  \n>> internal\n>> variant. But it does have to normalize them, otherwise how would it  \n>> detect the\n>> same filename using different normalizations?\n>\n> That's a singularly *stupid* argument.\n>\n> Here, let me rephrase that same idiotic argument:\n>\n>  \"But it does have to uppercase them, otherwise how would it detect  \n> the\n>   same filename using different cases?\"\n>\n> ..and if you don't see how that's *exactly* the same argument, you  \n> really\n> are stupid.\n\nYou're right, it doesn't actually have to store the normalized form.  \nAnd yes, it's possible to compare without normalizing them.  \nAdmittedly, I don't know much about the implementation details of  \nunicode, but I would assume that the easiest way to compare two  \nstrings is to normalize them first. But in the case of the filesystem,  \nnormalization actually is important if you're thinking about filenames  \nin terms of characters rather than bytes. When I feed the filesystem a  \ngiven unicode string, it has to find the file I'm talking about -  \nshould it do a relatively expensive unicode-sensitive comparison of  \nall the filenames with the one I gave it, or should it just normalize  \nall names and do the much cheaper lookup that way? I don't know about  \nyou, but I'd prefer to let my filesystem normalize the name and run  \nfaster.\n\n> The fact is, normalization is wrong.\n>\n> It's wrong when you normalize upper/lower case (no, the word  \n> \"Polish\" is\n> not the same as \"polish\"), and it's equally wrong when you normalize  \n> for\n> \"looks similar\".\n\nThere's a difference between \"looks similar\" as in \"Polish\" vs  \n\"polish\", and actually is the same string as in \"Ma<UMLAUT  \nMODIFIER>rchen\" vs \"M<A WITH UMLAUT>rchen\". Capitalization has a valid  \nsemantic meaning, normalization doesn't. The only way to argue that  \nnormalization is wrong is by providing a good reason to preserve the  \nexact byte sequence, and so far the only reason I've seen is to help  \ngit. Applications in general don't care one whit about the byte  \nsequence of the filename, they care about the underlying file the name  \nrepresents. Additionally, it would be a terrible experience for a user  \nto enter \"Märchen\" and have the application say \"sorry, I can't find  \nthis file\" simply because the application used decomposed characters  \nand the filename used composed characters. Unless the user is  \nknowledgeable about the OS, filesystems, and unicode, they wouldn't  \nhave a hope of figuring out what the problem was.\n\n>\n>> In other words, you're used to filenames as bytes, HFS+ treats  \n>> filenames\n>> as strings.\n>\n> No. HFS+ treats users as idiots and thinks that it should \"fix\" the\n> filename for them. And it causes problems.\n\nHow do you figure? When I type \"Märchen\", I'm typing a string, not a  \nbyte sequence. I have no control over the normalization of the  \ncharacters. Therefore, depending on what program I'm typing the name  \nin, I might use the same normalization as the filename, or I might  \nmiss. It's completely out of my control. This is why the filesystem  \nhas to step in and say \"You composed that character differently, but I  \nknow you were trying to specify this file\".\n\n> It causes problems for exactly the same reasons case-independence  \n> causes\n> problems, because it's EXACTLY THE SAME ISSUE. People may think that  \n> \"but\n> they are the same\", but they aren't. Case matters. And so does \"single\n> character\" vs \"two character overlay\".\n\nThere are valid reasons for case to matter, but what reason is there  \nfor \"single character\" vs\" two character overlay\" to matter in  \nfilenames? They're different representations of the exact same string,  \nand that's what a filename is - a string.\n\nIt seems like your arguments stem from the assumption that the user  \ncares about the byte sequence that represents the filename, which is  \nwrong. The user has no idea what the byte sequence is - the user cares  \nabout the string. Normalization is meant to help computers, not users,  \nand claiming that different normalizations of the same string produces  \ndifferent meaningful strings is complete bunk.\n\nIf you were to have two different files on your system, both of them  \ncalled \"Märchen\", but one precomposed and one decomposed, how would  \nyou specify which one you wanted? Unless Linux has a special text  \ninput system which gives the user control over the normalization of  \ntheir typed characters, you'd have to write out the UTF-8 bytes  \nmanually.\n\nI just don't understand this insistence on treating the specific byte  \nsequence that makes up the filename as significant.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65627","messageId":"F1428C56-9BAE-4903-A1A7-27019CBF4487@sb.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801162222180.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-16T23:16:35Z","receivedAt":"2008-01-16T23:16:35Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 5:23 PM, Johannes Schindelin wrote:\n\n> On Wed, 16 Jan 2008, Kevin Ballard wrote:\n>\n>> It's not using different encodings, it's all Unicode.\n>\n> But that's the _point_!  It _is_ Unicode, yet it uses _different_\n> encodings of the _same_ string.\n>\n> Now, this discussion gets really annoying.  The real question is:  \n> will you\n> do something about it, or reply with another 500-line email?\n\nI wish I could do something about it. But right now I'm a full-time  \nstudent trying to do contracting jobs on the side, and I don't believe  \nI have the time to learn enough about the guts of git to try and make  \nany changes to something as core as index filename handling. I just  \nwant people here to recognize that this is a valid problem instead of  \nsimply dismissing it as \"HFS+ is insane, lets just ignore this issue\".\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65633","messageId":"alpine.LFD.1.00.0801161522160.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"7652B11D-9B9F-45EA-9465-8294B701FE7C@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-16T23:38:35Z","receivedAt":"2008-01-16T23:38:35Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 16 Jan 2008, Kevin Ballard wrote:\n> \n> There's a difference between \"looks similar\" as in \"Polish\" vs \"polish\", and\n> actually is the same string as in \"Ma<UMLAUT MODIFIER>rchen\" vs \"M<A WITH\n> UMLAUT>rchen\". Capitalization has a valid semantic meaning, normalization\n> doesn't. \n\nThat simply isn't true.\n\nNormalization actually has real semantic meaning. If it didn't, there \nwould never ever be a reason why you'd use the non-normalized form in the \nfirst place.\n\nOthers have argued the exact same thing for capitalization. \"A\" is the \nsame letter as \"a\". Except there is a distinction.\n\nThe same is true of \"a<UMLAUT MODIFIER>\" and \"<a WITH UMLAUT>\". Yes, it's \nthe same \"chacter\" in either case. Except when there is a distinction.\n\nAnd there *are* cases where there are distinctions. Especially inside \ncomputers. For one thing, you may not be talking about \"characters on \nscreen\", but you may be talking about \"key sequences\". And suddenly \n\"a<UMLAUT MODIFIER>\" is a two-key sequence, and \"<a WITH UMLAUT>\" is a \nsingle-key sequence, and THEY ARE DIFFERENT.\n\nSee?\n\n\"a\" and \"A\" are the same letter. But sometimes case matters.\n\nMulti-character UTF-8 sequences may be the same character. But sometimes \nthe sequence matters.\n\nSame exact thing.\n\n>\tThe only way to argue that normalization is wrong is by providing a\n> good reason to preserve the exact byte sequence, and so far the only reason\n> I've seen is to help git.\n\nGit doesn't care. Just use the *same* sequence everywhere. Make sure \nsomething doesn't change it. Because if something changes it, git will \ntrack it.\n\n> How do you figure? When I type \"Märchen\", I'm typing a string, not a byte\n> sequence. I have no control over the normalization of the characters.\n> Therefore, depending on what program I'm typing the name in, I might use the\n> same normalization as the filename, or I might miss. It's completely out of my\n> control. This is why the filesystem has to step in and say \"You composed that\n> character differently, but I know you were trying to specify this file\".\n\nPure and utter garbage.\n\nWhat you are describing is an *input method* issue, not a filesystem \nissue.\n\nThe fact that you think this has anything what-so-ever to do with \nfilesystems, I cannot understand.\n\nHere's an example: I can type Märchen two different ways on my keyboard: I \ncan press the 'ä' key (yes, I have one, I have a Swedish keyboard), or I \ncould press the '¨' key and the 'a' key.\n\nSee: I get 'ä' and 'ä' respectively.\n\nAnd as I send this email off, those characters never *ever* got written as \nfilenames to any filesystem. But they *did* get written as part of \ntext-files to the disk using \"write()\", yes.\n\nAnd according to your *insane* logic, that write() call should have \nconverted them to the same representation, no?\n\nHell no! That conversion has absolutely nothing to do with the filesystem. \nIt's done at a totally different layer that actually knows what it is \ndoing, and turned them both into \\xc3\\xa4 (and then, the email client \nprobably will turn this into Latin1, and send it out as a single-byte \n'\\xe4' character).\n\nSee? Putting the conversion in the filesystem IS INSANE. You wouldn't make \nthe filesystem convert the characters in the data stream (because it would \ncause strange data conversion issues) AND FOR EXACTLY THE SAME REASON it \nshouldn't do it for filenames either!\n\nAnd your claim that \"you have no control over the normalization of \ncharacters\" is simply insane. Of course you have. It's just not supposed \nto be at the filesystem level - whether it's a write() call or a creat() \ncall!\n\n\t\t\tLinus\n"},{"id":"65635","messageId":"20080116235257.GA2901@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"65026F2B-5CE8-4238-A9AB-D3545D336B41@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-16T23:52:58Z","receivedAt":"2008-01-16T23:52:58Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"\nOn Wed, Jan 16, 2008 at 03:39:36PM -0500, Kevin Ballard wrote:\n> On Jan 16, 2008, at 11:46 AM, Jakub Narebski wrote:\n> \n> >\n> >HFS+ is just _stupid_. And unfortunately Git doesn't support stupid\n> >filesystems (e.g. case insensitive filesystems) well.\n> \n> There's two different ways to do filesystem encodings. One is to have  \n> the fs simply not care about encoding, which is what the linux world  \n> seems to prefer. \n\nThere is no technical reason for *kernel* to care about file name\nencoding. It is something that can be and should be dealt with in\nthe user space (except some special cases like smbfs).\n\n> Sure, this is great in that what you create the file  \n> with is what you get back,\n\nAnd also because a user space program can deal with it much more\ngracefully...\n\n> but on the other hand, given an arbitrary  \n> non-ASCII file on disk, you have absolutely no idea what the encoding  \n> should be and you can't display it without making assumptions (yes you  \n> can use heuristics, but you're still making assumptions).\n\nWrong. If you have a policy that all file names are stored in UTF-8\nencoding then there is no problem here. It should not be a kernel\nproblem to care about encoding, besides you cannot fully solve it\nin the kernel space anyway...\n\n> Filesystems  \n> like HFS+ that standardize the encoding,\n\nYeah, right... Like Microsoft likes to \"standardize\" everything, which\nin practice means forcing on others something fundamentally broken and\nthat does not follow any existing standard precisely:\n\n===\nIMPORTANT:\nThe terms used in this Q&A, decomposed and precomposed, roughly\ncorrespond to Unicode Normal Forms D and C, respectively. However, most\nvolume formats do not follow the exact specification for these normal\nforms.\n===\nhttp://developer.apple.com/qa/qa2001/qa1173.html\n\nNot to mention that the use of decomposed Unicode as the standard is\noutright silly -- no sane person writes in \"decomposed\" Unicode...\n\n> on the other hand, make it  \n> such that you always know what the encoding of a file should be, so  \n> you can always display and use the filename intelligently.\n\nSomehow I have no problem with displaying non-ASCII names on Linux.\nI can see both Unicode Normal Forms C and D encoded symbols without\nany problem, though the kernel is completely unaware about them.\n\n> It also  \n> means it plays much nicer in a non-ASCII world, since you don't have  \n> to worry about different normalizations of a given string referring to  \n> different files (it's one thing to be case-sensitive, but claiming  \n> that \"föo\" and \"föo\" are different files\n\nAs you typed them, they both are exactly the same, and both of them are\nin the Normal Forms C (which Mac calls as precomposed). So why do you\nuse one encoding in your writings and the other in your file names?\n\n> just because one uses a  \n> composed character and the other doesn't is extremely user- \n> unfriendly). On the other hand, what you create the file with may not  \n> be what you read back later, since the name has been standardized.  \n> It's hard to say one is better than the other, they're just different  \n> ways of doing it. However, I have noticed that everybody who's voiced  \n> an opinion on this list in favor of the encoding-agnostic approach  \n> seem to be unwilling to accept that any other approach might have  \n> validity, to the extent of calling an OS/filesystem that does things  \n> different stupid or insane. This strikes me as extremely elitist and  \n> risks alienating what I expect to be a fast-growing group of users  \n> (i.e. OS X users).\n\nI am sure everyone here is scared to death... I mean we have used to\nhear such threats from some MS salespeople, but from a Mac guy? It is\nreally scare....\n\nWake up, and stop shooting this nonsense at us. If you have technical\nreasons why your solution is better, let us know. So far, you do not\nsound very convincing here. Why do think that the issue of encoding can\nnot be dealt with in the user space? Why does Mac OS X uses so-called\ndecomposed Unicode, which even does not follow any standard precisely?\nWhy does Mac OS X chose to decompose characters while it does not\nsolve any real issue?\n\n> And one area that it has a problem with is the de- \n> facto filesystem on my OS of choice.\n\nI suppose it would be much better a subject for discussion...\nAt least, it would be more likely to result in that Git working\nbetter on your OS.\n\n> However, attempts to discuss the  \n> problem invariable end up with multiple people calling my OS stupid  \n> and insane simply because it differs in a particular design decision.  \n\nFirst, no one called Mac OS X insane, but case insensitive filesystems,\nand there are good reasons to think so, because no one has demonstrated\nso far any advantage of that approach, but disadvantages are quite \nobvious to anyone -- comparison of a stored file list with readdir()\nis much more problematic, and you cannot say that you have solved the\nproblem with encoding if you force other people to *duplicate* some\nlogic that Mac OS X does in its kernel just to get things working...\nSo, no one thinks it is insane because it is different, but because it\nrequires much more efforts to do the same thing -- compare two file\nlists, and this operation is important for Git to work properly...\n\n\nDmitry\n"},{"id":"65636","messageId":"BA518A23-FBF8-49BB-BEFB-D9A6BA1E302C@simplicidade.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161522160.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-16T23:57:21Z","receivedAt":"2008-01-16T23:57:21Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"\nOn Jan 16, 2008, at 11:38 PM, Linus Torvalds wrote:\n> On Wed, 16 Jan 2008, Kevin Ballard wrote:\n>> \tThe only way to argue that normalization is wrong is by providing a\n>> good reason to preserve the exact byte sequence, and so far the  \n>> only reason\n>> I've seen is to help git.\n>\n> Git doesn't care. Just use the *same* sequence everywhere. Make sure\n> something doesn't change it. Because if something changes it, git will\n> track it.\n\nThe problem is that you don't control the sequence that everybody uses.\n\nSee this example:\n\nmelo@speed(~)$ uname -a\nLinux speed.simplicidade.org 2.6.9-55.ELsmp #1 SMP Wed May 2 14:28:44  \nEDT 2007 i686 i686 i386 GNU/Linux\nmelo@speed(~)$ set | grep LANG\nLANG=en_US.UTF-8\nmelo@speed(~)$ mkdir t\nmelo@speed(~)$ cd t\nmelo@speed(~/t)$ git init\nInitialized empty Git repository in .git/\nmelo@speed(~/t)$ touch á\nmelo@speed(~/t)$ git-add á\nmelo@speed(~/t)$ git-commit -m \"added a in utf8\"\nCreated initial commit 7a473a2: added a in utf8\n  0 files changed, 0 insertions(+), 0 deletions(-)\n  create mode 100644 \"\\303\\241\"\nmelo@speed(~/t)$ export LANG=en_US\nmelo@speed(~/t)$ touch á\nmelo@speed(~/t)$ ls -la\ntotal 12\ndrwxrwxr-x   3 melo melo 4096 Jan 16 23:44 .\ndrwx--x--x  31 melo melo 4096 Jan 16 23:43 ..\n-rw-rw-r--   1 melo melo    0 Jan 16 23:44 á\n-rw-rw-r--   1 melo melo    0 Jan 16 23:43 Ã¡\ndrwxrwxr-x   8 melo melo 4096 Jan 16 23:43 .git\nmelo@speed(~/t)$ git-add á\nmelo@speed(~/t)$ git-commit -m \"added a in iso-latin-1\"\nCreated commit 4282fca: OlÃ¡x!\n  0 files changed, 0 insertions(+), 0 deletions(-)\n  create mode 100644 \"\\341\"\n\nSo two (simulated in this test) users who use different LANG settings  \nwill be in trouble in no time.\n\nWhat I take from this conversation is that I have to specify, for  \neach project I work on, which encoding we should use, across all  \nusers, before they start using git with files with accented chars.\n\nThe difference I see between us is that if I tell my filesystem that  \nI want to name my file with a particular string encoded in X, users  \nusing encoding Y will be able to read it correctly. I  like my  \nfilesystem to make that work for me.\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65637","messageId":"85ir1tpbk8.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161522160.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-16T23:58:47Z","receivedAt":"2008-01-16T23:58:47Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Wed, 16 Jan 2008, Kevin Ballard wrote:\n>> \n>> There's a difference between \"looks similar\" as in \"Polish\" vs \"polish\", and\n>> actually is the same string as in \"Ma<UMLAUT MODIFIER>rchen\" vs \"M<A WITH\n>> UMLAUT>rchen\". Capitalization has a valid semantic meaning, normalization\n>> doesn't. \n>\n> That simply isn't true.\n>\n> Normalization actually has real semantic meaning. If it didn't, there\n> would never ever be a reason why you'd use the non-normalized form in\n> the first place.\n\nActually, there is no good reason for non-normalized forms (deficient\nsoftware not able to deal with some of the normalized forms is not a\ngood reason: such software should be fixed).\n\nIt is just that the file system is a rather quirky place for enforcing\nthe normalization.  One should not be able to get unnormalized forms\ncreated easily in the first place, be it command line or script.\n\n> And there *are* cases where there are distinctions. Especially inside\n> computers. For one thing, you may not be talking about \"characters on\n> screen\", but you may be talking about \"key sequences\". And suddenly\n> \"a<UMLAUT MODIFIER>\" is a two-key sequence, and \"<a WITH UMLAUT>\" is a\n> single-key sequence, and THEY ARE DIFFERENT.\n>\n> See?\n\nNo.  Input methods are not the same as their resulting string.  I can\neven produce some ASCII characters on my keyboard in more than one way\nand would not expect them to lead to different codes.\n\n>> How do you figure? When I type \"Märchen\", I'm typing a string, not a\n>> byte sequence. I have no control over the normalization of the\n>> characters.  Therefore, depending on what program I'm typing the name\n>> in, I might use the same normalization as the filename, or I might\n>> miss. It's completely out of my control. This is why the filesystem\n>> has to step in and say \"You composed that character differently, but\n>> I know you were trying to specify this file\".\n>\n> Pure and utter garbage.\n>\n> What you are describing is an *input method* issue, not a filesystem\n> issue.\n>\n> The fact that you think this has anything what-so-ever to do with\n> filesystems, I cannot understand.\n\nHow nice.  We are actually in agreement here.\n\n> See? Putting the conversion in the filesystem IS INSANE. You wouldn't\n> make the filesystem convert the characters in the data stream (because\n> it would cause strange data conversion issues) AND FOR EXACTLY THE\n> SAME REASON it shouldn't do it for filenames either!\n\nYup.  But that does not mean that normalization is a bad idea.  It is\njust that the filesystem is not the right place for it.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"65638","messageId":"B45968C6-3029-48B6-BED2-E7D5A88747F7@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161522160.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-17T00:09:59Z","receivedAt":"2008-01-17T00:09:59Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 6:38 PM, Linus Torvalds wrote:\n\n> On Wed, 16 Jan 2008, Kevin Ballard wrote:\n>>\n>> There's a difference between \"looks similar\" as in \"Polish\" vs  \n>> \"polish\", and\n>> actually is the same string as in \"Ma<UMLAUT MODIFIER>rchen\" vs  \n>> \"M<A WITH\n>> UMLAUT>rchen\". Capitalization has a valid semantic meaning,  \n>> normalization\n>> doesn't.\n>\n> That simply isn't true.\n>\n> Normalization actually has real semantic meaning. If it didn't, there\n> would never ever be a reason why you'd use the non-normalized form  \n> in the\n> first place.\n\nMy understanding is that normalization is there to help the computer.  \nThat doesn't give it any semantic meaning, because all normal forms of  \na given string still represent the exact same string to the user.\n\n> Others have argued the exact same thing for capitalization. \"A\" is the\n> same letter as \"a\". Except there is a distinction.\n\nThe argument for case insensitivity is different than the argument for  \nnormalization. I certainly hope you understand why they are different  \narguments, or there's really no point in going further.\n\n> The same is true of \"a<UMLAUT MODIFIER>\" and \"<a WITH UMLAUT>\". Yes,  \n> it's\n> the same \"chacter\" in either case. Except when there is a distinction.\n>\n> And there *are* cases where there are distinctions. Especially inside\n> computers. For one thing, you may not be talking about \"characters on\n> screen\", but you may be talking about \"key sequences\". And suddenly\n> \"a<UMLAUT MODIFIER>\" is a two-key sequence, and \"<a WITH UMLAUT>\" is a\n> single-key sequence, and THEY ARE DIFFERENT.\n>\n> See?\n>\n> \"a\" and \"A\" are the same letter. But sometimes case matters.\n>\n> Multi-character UTF-8 sequences may be the same character. But  \n> sometimes\n> the sequence matters.\n>\n> Same exact thing.\n\nYou're right, sometimes the sequence matters. As in key sequences. But  \nwe're not talking about key sequences, we're talking about strings.  \nJust because it matters sometimes doesn't mean it matters all the time.\n\n\n>> \tThe only way to argue that normalization is wrong is by providing a\n>> good reason to preserve the exact byte sequence, and so far the  \n>> only reason\n>> I've seen is to help git.\n>\n> Git doesn't care. Just use the *same* sequence everywhere. Make sure\n> something doesn't change it. Because if something changes it, git will\n> track it.\n\nAnd how am I supposed to use the same sequence everywhere? When I type  \n\"Märchen\", I don't know which form I'm typing, nor should I. It's not  \nsomething that I, as a user, should have to know. Especially if I pass  \nthis name through various other utilities before using it - I have no  \nidea if another utility is going to end up normalizing the name, and  \nit shouldn't matter, as they are equivalent strings.\n\n>> How do you figure? When I type \"Märchen\", I'm typing a string, not  \n>> a byte\n>> sequence. I have no control over the normalization of the characters.\n>> Therefore, depending on what program I'm typing the name in, I  \n>> might use the\n>> same normalization as the filename, or I might miss. It's  \n>> completely out of my\n>> control. This is why the filesystem has to step in and say \"You  \n>> composed that\n>> character differently, but I know you were trying to specify this  \n>> file\".\n>\n> Pure and utter garbage.\n>\n> What you are describing is an *input method* issue, not a filesystem\n> issue.\n>\n> The fact that you think this has anything what-so-ever to do with\n> filesystems, I cannot understand.\n>\n> Here's an example: I can type Märchen two different ways on my  \n> keyboard: I\n> can press the 'ä' key (yes, I have one, I have a Swedish keyboard),  \n> or I\n> could press the '¨' key and the 'a' key.\n>\n> See: I get 'ä' and 'ä' respectively.\n\nOn a US keyboard I only have one way of typing ä, and I have no idea  \nwhether it ends up precomposed or decomposed in the resulting byte  \nstream. And I don't care. Because I'm typing characters, not bytes. I  \ncould be typing in a file in ISO-Latin-1 and I still wouldn't care,  \nbecause it looks the same to me. If my filesystem did make a  \ndistinction between the normal forms, and I see that I have a file  \nnamed \"Märchen\", how am I supposed to type that at my keyboard? I  \ndon't know which normal form it's using.\n\nThe fact that you think the normalization of the string matters, I  \ndon't understand.\n\n> And as I send this email off, those characters never *ever* got  \n> written as\n> filenames to any filesystem. But they *did* get written as part of\n> text-files to the disk using \"write()\", yes.\n>\n> And according to your *insane* logic, that write() call should have\n> converted them to the same representation, no?\n>\n>\n> Hell no! That conversion has absolutely nothing to do with the  \n> filesystem.\n> It's done at a totally different layer that actually knows what it is\n> doing, and turned them both into \\xc3\\xa4 (and then, the email client\n> probably will turn this into Latin1, and send it out as a single-byte\n> '\\xe4' character).\n>\n> See? Putting the conversion in the filesystem IS INSANE. You  \n> wouldn't make\n> the filesystem convert the characters in the data stream (because it  \n> would\n> cause strange data conversion issues) AND FOR EXACTLY THE SAME  \n> REASON it\n> shouldn't do it for filenames either!\n\nWhat a fabulous straw man argument you just put together. I hope you  \ndon't need me to point out why this argument is fundamentally flawed.\n\n> And your claim that \"you have no control over the normalization of\n> characters\" is simply insane. Of course you have. It's just not  \n> supposed\n> to be at the filesystem level - whether it's a write() call or a  \n> creat()\n> call!\n\nI'm speaking as a user, and as such, I shouldn't even have to know  \nthat it's possible to write the same character in multiple different  \nways. As a user, HFS+ behaves exactly the way I want it to. You were  \ntalking earlier about not messing with the \"user data\", but what is  \nthe \"user data\"? It's the string, not the byte sequence. That's all I  \ncare about - the string. That's all the OS cares about, that's all any  \napplication I use cares about, and that's all git should care about.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65640","messageId":"alpine.LFD.1.00.0801161615330.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"BA518A23-FBF8-49BB-BEFB-D9A6BA1E302C@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T00:16:58Z","receivedAt":"2008-01-17T00:16:58Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 16 Jan 2008, Pedro Melo wrote:\n> \n> The difference I see between us is that if I tell my filesystem that I want to\n> name my file with a particular string encoded in X, users using encoding Y\n> will be able to read it correctly. I  like my filesystem to make that work for\n> me.\n\nThe difference I see between us is that when I tell you that this is \nexactly the same thing as your file *contents*, you don't seem to get it.\n\nAn OS that silently changes the contents of your files is *crap*.\n\nGet it?\n\nAn OS that silently changes the contents of your directories is *crap*.\n\nGet it now?\n\n\t\tLinus\n"},{"id":"65643","messageId":"alpine.LFD.1.00.0801161617070.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"85ir1tpbk8.fsf@lola.goethe.zz","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T00:19:30Z","receivedAt":"2008-01-17T00:19:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 17 Jan 2008, David Kastrup wrote:\n> \n> Actually, there is no good reason for non-normalized forms (deficient\n> software not able to deal with some of the normalized forms is not a\n> good reason: such software should be fixed).\n\nI'd actually agree, and it then boils down to the second sane choice I \ngave earlier:\n\n - don't accept data you don't like\n\nif you don't like non-normalized names, don't create them. That's fine.\n\nBut don't go normalizing them behind the users back.\n\n> Yup.  But that does not mean that normalization is a bad idea.  It is\n> just that the filesystem is not the right place for it.\n\nOh, absolutely. You can - and often should - normalize in the application \n(or have libraries to do it for you). \n\nNot silently and behind peoples backs.\n\n\t\tLinus\n"},{"id":"65641","messageId":"alpine.LFD.1.00.0801161619420.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"B45968C6-3029-48B6-BED2-E7D5A88747F7@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T00:25:52Z","receivedAt":"2008-01-17T00:25:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 16 Jan 2008, Kevin Ballard wrote:\n> \n> My understanding is that normalization is there to help the computer. That\n> doesn't give it any semantic meaning, because all normal forms of a given\n> string still represent the exact same string to the user.\n\nTHAT IS NOT TRUE!\n\nHow the hell does the computer know what the string means?\n\nHint: it does not.\n\nThe fact is, the user may use a non-normalized string on purpose. It's not \nyour place to say that the user is wrong. Your \"undestanding\" is simply \nwrong. Two strings are *different* if they are [un]normalized differently.\n\nReally.\n\nThe exact same way the word Polish and polish are different, just because \nthey are capitalized differently.\n\n> The argument for case insensitivity is different than the argument for\n> normalization. I certainly hope you understand why they are different\n> arguments, or there's really no point in going further.\n\nYou do not understand.\n\nIn *order* to do case-insensitivity, you generally need to normalize (and \ndo other things too - normalization is just *one* of the things you need \nto do).\n\nSo if you are a case-insensitive filesystem, then normalization is sane.\n\nBut if you aren't, then there is no reason to normalize.\n\n> You're right, sometimes the sequence matters. As in key sequences. But we're\n> not talking about key sequences, we're talking about strings.\n\nYou define \"string\" to be something totally made-up.\n\nIn your world \"string\" means \"normalized\". BUT IT'S NOT TRUE!\n\nYou define normalization to be a property of strings, without any actual \nbacking for why that would be.\n\nThe fact is, *looks the same* is very very different from *is the same*.\n\nBut you seem to be too stupid to undestand the differce.\n\n\t\tLinus\n"},{"id":"65642","messageId":"E8E76634-FFEC-426B-B04D-3C2CD3790D5E@simplicidade.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161615330.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-17T00:27:52Z","receivedAt":"2008-01-17T00:27:52Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"Hi,\n\nOn Jan 17, 2008, at 12:16 AM, Linus Torvalds wrote:\n> On Wed, 16 Jan 2008, Pedro Melo wrote:\n>>\n>> The difference I see between us is that if I tell my filesystem  \n>> that I want to\n>> name my file with a particular string encoded in X, users using  \n>> encoding Y\n>> will be able to read it correctly. I  like my filesystem to make  \n>> that work for\n>> me.\n>\n> The difference I see between us is that when I tell you that this is\n> exactly the same thing as your file *contents*, you don't seem to  \n> get it.\n\nI get that you think its the same thing.\n\nWhat I don't get is why a user should be forced to know what type of  \nencoding he and the other users are using on all the layers going  \ndown to the filesystem. If two users on different systems or in  \ndifferent configurations, choose the same unicode string as the name,  \nwhy do we need to make it harder for things to just work out?\n\nThe content of the file is sacred, we both agree on that. We disagree  \non the filename, because for me it's more important that equal  \nstrings, even if encoded to different byte sequences, should be  \ntreated as the same file.\n\n> An OS that silently changes the contents of your files is *crap*.\n>\n> Get it?\n\nI was not talking about content of files, those are sacred. I was  \ntalking about filenames. Those *for me* are not, but are for you. No  \nproblem, we just have different values: I want my computer to work  \nfor me, not me working for the computer. I'm willing to accept a file  \nsystem or other layer that normalizes encoding of filenames if that  \nmakes the end-user life easier, specially in a tool distributed by  \nnature.\n\n> An OS that silently changes the contents of your directories is  \n> *crap*.\n>\n> Get it now?\n\nAs I said before, we disagree on file meta-data, not on file  \ncontents. For you, byte in must be the same byte out. For me string  \nin must be the same string out.\n\nAnd as I said in the previous email, what I learned today is that in  \na distributed project using git, and if you need to use accented  \ncharacters, I need to tell all the users to use the same LANG settings.\n\nIt's important information, at least for me.\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65644","messageId":"85zlv5nvge.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"E8E76634-FFEC-426B-B04D-3C2CD3790D5E@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-17T00:32:01Z","receivedAt":"2008-01-17T00:32:01Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Pedro Melo <melo@simplicidade.org> writes:\n\n> On Jan 17, 2008, at 12:16 AM, Linus Torvalds wrote:\n>> On Wed, 16 Jan 2008, Pedro Melo wrote:\n>>>\n>>> The difference I see between us is that if I tell my filesystem that\n>>> I want to name my file with a particular string encoded in X, users\n>>> using encoding Y will be able to read it correctly. I like my\n>>> filesystem to make that work for me.\n>>\n>> The difference I see between us is that when I tell you that this is\n>> exactly the same thing as your file *contents*, you don't seem to get\n>> it.\n>\n> I get that you think its the same thing.\n>\n> What I don't get is why a user should be forced to know what type of\n> encoding he and the other users are using on all the layers going down\n> to the filesystem. If two users on different systems or in different\n> configurations, choose the same unicode string as the name, why do we\n> need to make it harder for things to just work out?\n\nIf you do the normalization in the right place, things will just work\nout.  The file system is not the right place.\n\n> I'm willing to accept a file system or other layer that normalizes\n> encoding of filenames if that makes the end-user life easier,\n> specially in a tool distributed by nature.\n\nWell, as the issue shows it does not make life for the end-user easier.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"65645","messageId":"alpine.LSU.1.00.0801170032230.17650@racer.site","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161619420.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T00:33:38Z","receivedAt":"2008-01-17T00:33:38Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 16 Jan 2008, Linus Torvalds wrote:\n\n> So if you are a case-insensitive filesystem, then normalization is sane.\n\nActually, no.  Even an case-challenged filesystem should keep the \n_original_ name around, if only for the exact same argument you used \nearlier: if the user chooses to capitalise some letters, but not others, \nit is not the filesystem's place to \"correct\" that.\n\nCiao,\nDscho\n"},{"id":"65646","messageId":"alpine.LSU.1.00.0801170034300.17650@racer.site","threadId":"11645","inReplyTo":"E8E76634-FFEC-426B-B04D-3C2CD3790D5E@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T00:35:05Z","receivedAt":"2008-01-17T00:35:05Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 Jan 2008, Pedro Melo wrote:\n\n> The content of the file is sacred, we both agree on that. We disagree on \n> the filename, because for me it's more important that equal strings, \n> even if encoded to different byte sequences, should be treated as the \n> same file.\n\nWhy should the filename be _stored_ normalised?  I agree on the lookup, \nyes, but not the storage.\n\nHth,\nDscho\n"},{"id":"65647","messageId":"B2E52451-5153-4EFD-ADBE-AACDCEF6169E@simplicidade.org","threadId":"11645","inReplyTo":"85zlv5nvge.fsf@lola.goethe.zz","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-17T00:40:55Z","receivedAt":"2008-01-17T00:40:55Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"Hello,\n\nOn Jan 17, 2008, at 12:32 AM, David Kastrup wrote:\n> Pedro Melo <melo@simplicidade.org> writes:\n>> On Jan 17, 2008, at 12:16 AM, Linus Torvalds wrote:\n>>> On Wed, 16 Jan 2008, Pedro Melo wrote:\n>>>>\n>>>> The difference I see between us is that if I tell my filesystem  \n>>>> that\n>>>> I want to name my file with a particular string encoded in X, users\n>>>> using encoding Y will be able to read it correctly. I like my\n>>>> filesystem to make that work for me.\n>>>\n>>> The difference I see between us is that when I tell you that this is\n>>> exactly the same thing as your file *contents*, you don't seem to  \n>>> get\n>>> it.\n>>\n>> I get that you think its the same thing.\n>>\n>> What I don't get is why a user should be forced to know what type of\n>> encoding he and the other users are using on all the layers going  \n>> down\n>> to the filesystem. If two users on different systems or in different\n>> configurations, choose the same unicode string as the name, why do we\n>> need to make it harder for things to just work out?\n>\n> If you do the normalization in the right place, things will just work\n> out.  The file system is not the right place.\n\nNo problem, but don't you think that git should to it?\n\nDon't you think its important in a distributed tool that no matter  \nwhat system they use, be it linux or solaris, they are able to talk  \nabout a file with non-ascii chars and be the same file to both of them?\n\nThat's the point I'm making. The fact that I need to set LANG across  \nall users of a project is insane...\n\n>> I'm willing to accept a file system or other layer that normalizes\n>> encoding of filenames if that makes the end-user life easier,\n>> specially in a tool distributed by nature.\n>\n> Well, as the issue shows it does not make life for the end-user  \n> easier.\n\nI'm assuming you are talking about HFS+ and the strange normalization  \nit does.\n\nI'm sorry but that was not the problem I sent. I sent a scenario, in  \nwhich two users, using the same linux system but with different LANG  \nsettings cannot use git reliably.\n\nAlthough this thread started because of HFS+ \"choices\", the problem  \nis not really related to HFS+ given that you can have the same issues  \neven on the same physical <insert flavor here> POSIX system.\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65648","messageId":"16C3D0BC-3A70-441C-89D9-71B5F0DA0790@simplicidade.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801170032230.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-17T00:43:30Z","receivedAt":"2008-01-17T00:43:30Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"Hi,\n\nOn Jan 17, 2008, at 12:33 AM, Johannes Schindelin wrote:\n\n> On Wed, 16 Jan 2008, Linus Torvalds wrote:\n>\n>> So if you are a case-insensitive filesystem, then normalization is  \n>> sane.\n>\n> Actually, no.  Even an case-challenged filesystem should keep the\n> _original_ name around, if only for the exact same argument you used\n> earlier: if the user chooses to capitalise some letters, but not  \n> others,\n> it is not the filesystem's place to \"correct\" that.\n\nFor the record, HFS+ is case-insensitive but case-preserving so I  \nbelieve they keep the original filename around. I don't have the spec  \nin front of me, but from memory I believe that this is what they do.\n\nBut I think that focusing on HFS+ is loosing sight of the real  \nproblem. It's not about encoding at the filesystem, but encoding  \ninside the git structures.\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65649","messageId":"A61EFF37-A235-49C3-8F1F-64B4FB26D10B@simplicidade.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801170034300.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-17T00:45:50Z","receivedAt":"2008-01-17T00:45:50Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"\nOn Jan 17, 2008, at 12:35 AM, Johannes Schindelin wrote:\n> On Thu, 17 Jan 2008, Pedro Melo wrote:\n>\n>> The content of the file is sacred, we both agree on that. We  \n>> disagree on\n>> the filename, because for me it's more important that equal strings,\n>> even if encoded to different byte sequences, should be treated as the\n>> same file.\n>\n> Why should the filename be _stored_ normalised?  I agree on the  \n> lookup,\n> yes, but not the storage.\n\nPersonally I don't care how you store it. It's an implementation  \ndetail, and you should choose the best one for your use cases. If  \nthat means that you store the original version and a normalized  \nversion just for lookups, fine.\n\nWhat I think its important is that if two users use different  \nencodings for the same string in a filename, git should treat that as  \nthe same file.\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65650","messageId":"D32FF2AF-EA90-4737-8320-836B52AF4612@wincent.com","threadId":"11645","inReplyTo":"B2E52451-5153-4EFD-ADBE-AACDCEF6169E@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-17T00:54:05Z","receivedAt":"2008-01-17T00:54:05Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 17/1/2008, a las 1:40, Pedro Melo escribió:\n\n> That's the point I'm making. The fact that I need to set LANG across  \n> all users of a project is insane...\n\nI don't think I'd call that \"insane\" (in fact, I think these  \ndiscussions would be much less irritating for all involved if we  \ndidn't use that word so often, even when it's not called for). It's  \nnot that different than the whole LF/CRLF line-ending thing.\n\nThe real problem is that setting LANG won't help you on Mac OS X; set  \nLANG to whatever you want and there is *nothing* that you can do to  \nstop your filenames being normalized into decomposed UTF-8, short of  \ndropping HFS+. You can use an alternative filesystem, but support for  \nbasically everything except HFS+ is suboptimal in Mac OS X at the  \nmoment.\n\nCheers,\nWincent\n"},{"id":"65651","messageId":"alpine.LSU.1.00.0801170054280.17650@racer.site","threadId":"11645","inReplyTo":"16C3D0BC-3A70-441C-89D9-71B5F0DA0790@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T00:57:20Z","receivedAt":"2008-01-17T00:57:20Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 Jan 2008, Pedro Melo wrote:\n\n> On Jan 17, 2008, at 12:33 AM, Johannes Schindelin wrote:\n> \n> > On Wed, 16 Jan 2008, Linus Torvalds wrote:\n> > \n> > > So if you are a case-insensitive filesystem, then normalization is \n> > > sane.\n> > \n> > Actually, no.  Even an case-challenged filesystem should keep the \n> > _original_ name around, if only for the exact same argument you used \n> > earlier: if the user chooses to capitalise some letters, but not \n> > others, it is not the filesystem's place to \"correct\" that.\n> \n> For the record, HFS+ is case-insensitive but case-preserving so I \n> believe they keep the original filename around.\n\nFor the record, that's only the default setting.  AFAIK you can configure \nit to care about case, too.\n\nAlso for the record, the whole thread was about HFS+ _not_ keeping the \noriginal filename around, but _only_ a normalised version of it.\n\n> But I think that focusing on HFS+ is loosing sight of the real problem. \n> It's not about encoding at the filesystem, but encoding inside the git \n> structures.\n\nSo far I have not seen anyone talking _seriously_ about this issue.  Only \na few shouts \"you should support\", and a few shouts back \"I don't care \nabout insane filesystems\".\n\nTherefore, I fully agree with you that we're losing sight of the real \nproblem.\n\nCiao,\nDscho\n"},{"id":"65653","messageId":"alpine.LFD.1.00.0801161705520.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801170032230.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T01:06:33Z","receivedAt":"2008-01-17T01:06:33Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 17 Jan 2008, Johannes Schindelin wrote:\n> \n> On Wed, 16 Jan 2008, Linus Torvalds wrote:\n> \n> > So if you are a case-insensitive filesystem, then normalization is sane.\n> \n> Actually, no.  Even an case-challenged filesystem should keep the \n> _original_ name around\n\nYou're right. The normalization only really needs to happen as part of the \nname comparison itself.\n\n\t\tLinus\n"},{"id":"65654","messageId":"alpine.LSU.1.00.0801170106390.17650@racer.site","threadId":"11645","inReplyTo":"D32FF2AF-EA90-4737-8320-836B52AF4612@wincent.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T01:08:15Z","receivedAt":"2008-01-17T01:08:15Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n\n> El 17/1/2008, a las 1:40, Pedro Melo escribió:\n> \n> > That's the point I'm making. The fact that I need to set LANG across \n> > all users of a project is insane...\n\nFWIW if you use another filesystem, such as reiserfs or ext[2-4], the \nfilenames will be _unaffected_ by your particular setting of LANG.  They \nwill be stored byte-wise exactly like asked for.  That's why I call them \n\"sane\".\n\nHth,\nDscho"},{"id":"65659","messageId":"alpine.LFD.1.00.0801161707150.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"B45968C6-3029-48B6-BED2-E7D5A88747F7@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T01:16:45Z","receivedAt":"2008-01-17T01:16:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 16 Jan 2008, Kevin Ballard wrote:\n> \n> I'm speaking as a user, and as such, I shouldn't even have to know that it's\n> possible to write the same character in multiple different ways.\n\nThe thing is, you seem to argue that what OS X does helps you as the user.\n\nBut you are arguing based on incorrect assumptions.\n\nFirst off, we've had years and years and years of usage of non-corrupting \nfilesystems (pretty much every UNIX OS around since day 1, and many other \nOS's too), and it's simply not true that it's a problem. You see the \nfilename in the file dialog, and you open it, and you're done. OS X isn't \nany \"easier\" in this regard.\n\nIn fact, this whole thread comes from the fact that the OS X choice that \nyou *think* is easier, is in fact not easier at all. It's not easier for \nthe user, it's not easier for the application programmer, and the really \nsad part is that it's very much *not* easier for OS X itself either (ie \nthey had to literally write extra code with nasty tables to do it, and it \nreally does hurt them in performance and complexity).\n\nAnd _that_ is why the OS X situation is so sad. Apple literally added \nextra code to make things slower and more complex *and* harder to use \nreliably.\n\nDoes it show up in normal behaviour? Of course not. You'd probably never \nsee it in real life outside of test-suites. People simply don't even tend \nto use filenames outside of US-ASCII, and when they do use them, input \nmethods really *do* tend to do the normalization for you.\n\nBut when it comes to automation (which is what computers are all about), \nthe OS X choice is literally the wrong one. And there's no _upside_. It's \nall downside. Which is why it's so stupid.\n\nI bet it only exists because OS X engineers didn't really even think about \nit, and they just assumed that \"normalization is helpful\". They took your \nstance - thinking it was worth it, without ever really thinking it \nthrough.\n\n\t\t\tLinus\n"},{"id":"65663","messageId":"alpine.LFD.1.00.0801161717160.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801170106390.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T01:41:20Z","receivedAt":"2008-01-17T01:41:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 17 Jan 2008, Johannes Schindelin wrote:\n> On Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n> \n> > El 17/1/2008, a las 1:40, Pedro Melo escribió:\n> > \n> > > That's the point I'm making. The fact that I need to set LANG across \n> > > all users of a project is insane...\n> \n> FWIW if you use another filesystem, such as reiserfs or ext[2-4], the \n> filenames will be _unaffected_ by your particular setting of LANG.  They \n> will be stored byte-wise exactly like asked for.  That's why I call them \n> \"sane\".\n\nOne of the advantages (the biggest one, in fact, apart from the obvious \nUS-ASCII down-compatibility and the fact that you can do C-compatible \nNUL-terminated strings) of UTF-8 is that it's locale-independent, and \ndoesn't care about LANG, because it's valid in all languages.\n\nAnd that's really important. It's important for a very simple reason: \nthere is almost never such a thing as \"a locale\" except for US-ASCII. Once \nyou move away from US-ASCII, it actually tends to be much more common that \nyou have a *mixture* of locales - often in the same \"document\" - than to \nhave one single locale.\n\nIt very much happens even in filenames - people \"mix\" locales in trivial \nways even within a single pathname component (non-US-ASCII filename, but \nwith a regular file extension), but much more interestingly they do so \nwithin a directory tree (ie you have have translation subdirectories where \nthe filenames themselves are in another language, and you can have full \npathnames where different components are in different languages, for \nexample).\n\nAnd UTF-8 is _wonderful_ for this, because LANG doesn't matter, and \ncannot matter, and thus mixing isn't a problem.\n\nOf course, you can screw it up. Locales still can change things like sort \norder and capitalization etc, so even if you use UTF-8, you sure can get \ninto trouble with LANG and thinking that a per-session locale makes sense.\n\nSo choosing UTF-8 for the filesystem isn't wrong per se. It's a fine \nchoice, and has no issues with LANG in itself. Limiting it to strictly \nvalid UTF-8 encodings is also fine. Limiting it (further) to only \ncharacter normalized UTF-8 is also fine.\n\nMost Linux filesystems don't limit it in any way, so you can make \nfilenames that aren't valid UTF-8 at all, much less normalizing \nmulti-character sequences.\n\nI personally think that's the best option, but I probably do so mostly \nbecause I know some people still use Latin1 as their only locale (and I \nsuspect Asia will take decades before it has converted to UTF-8 and will \nalso have cases where they use other non-UTF locales).\n\nBut enforcing clean UTF-8 is not a bad idea per se. Not allowing byte \nsequences that aren't a valid UTF-8 encoding (eg \\xc0\\xc0 is not a valid \nUTF-8 character) is fine.\n\nI wouldn't call people crazy for doing that, although it does mean that \nyou cannot, for example, decide to write a Latin1 filename (which is not \nnecessarily a *good* idea in this day and age, but I think there's a \ndifference between \"that's not a good idea\" and \"you cannot do that\").\n\nAnd even limiting the UTF-8 charset further to only the minimal \nrepresentation of one particular glyph (ie not allowing multi-character \nsequences that can be represented more simply) may be even *more* \nbig-brother, but would at least not cause the technical aliasing issues. I \npersonally think that's so controlling as to be stupid (and has no real \nadvantage), but hey, at least it doesn't *corrupt* anything silently.\n\nSo I think that using UTF-8 as a character encoding is a *good* thing to \ndo, and that automatically means that LANG shouldn't matter for filenames, \nbut within that choice of UTF-8 there are still mistakes that you can \nmake. Notably multi-character normalization and case-insensitivity.\n\n\t\t\tLinus\n"},{"id":"65678","messageId":"8AC4CC86-A711-483D-9F9C-5F8497006A1D@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161707150.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-17T03:52:19Z","receivedAt":"2008-01-17T03:52:19Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 8:16 PM, Linus Torvalds <torvalds@linux-foundation.org \n > wrote:\n\n> On Wed, 16 Jan 2008, Kevin Ballard wrote:\n>>\n>> I'm speaking as a user, and as such, I shouldn't even have to know  \n>> that it's\n>> possible to write the same character in multiple different ways.\n>\n> The thing is, you seem to argue that what OS X does helps you as the  \n> user.\n>\n> But you are arguing based on incorrect assumptions.\n>\n> First off, we've had years and years and years of usage of non- \n> corrupting\n> filesystems (pretty much every UNIX OS around since day 1, and many  \n> other\n> OS's too), and it's simply not true that it's a problem. You see the\n> filename in the file dialog, and you open it, and you're done. OS X  \n> isn't\n> any \"easier\" in this regard.\n>\n> In fact, this whole thread comes from the fact that the OS X choice  \n> that\n> you *think* is easier, is in fact not easier at all. It's not easier  \n> for\n> the user, it's not easier for the application programmer, and the  \n> really\n> sad part is that it's very much *not* easier for OS X itself either  \n> (ie\n> they had to literally write extra code with nasty tables to do it,  \n> and it\n> really does hurt them in performance and complexity).\n>\n> And _that_ is why the OS X situation is so sad. Apple literally added\n> extra code to make things slower and more complex *and* harder to use\n> reliably.\n>\n> Does it show up in normal behaviour? Of course not. You'd probably  \n> never\n> see it in real life outside of test-suites. People simply don't even  \n> tend\n> to use filenames outside of US-ASCII, and when they do use them, input\n> methods really *do* tend to do the normalization for you.\n>\n> But when it comes to automation (which is what computers are all  \n> about),\n> the OS X choice is literally the wrong one. And there's no _upside_.  \n> It's\n> all downside. Which is why it's so stupid.\n>\n> I bet it only exists because OS X engineers didn't really even think  \n> about\n> it, and they just assumed that \"normalization is helpful\". They took  \n> your\n> stance - thinking it was worth it, without ever really thinking it\n> through.\n>\n>            Linus\n>\n\nI believe it exists because HFS+ was created at a time when the Mac  \nwas moving from a multi-encoding world (which was a nightmare) to a  \nUnicode world and they wanted to remove ambiguity in filenames. But I  \nwasn't around when they made this decision so this is just a guess.\n\n-Kevin Ballard\n"},{"id":"65679","messageId":"A4CE8450-E470-4F32-BCBD-05BF9A458D87@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161717160.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-17T04:07:25Z","receivedAt":"2008-01-17T04:07:25Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 8:41 PM, Linus Torvalds wrote:\n\n> On Thu, 17 Jan 2008, Johannes Schindelin wrote:\n>> On Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n>>\n>>> El 17/1/2008, a las 1:40, Pedro Melo escribió:\n>>>\n>>>> That's the point I'm making. The fact that I need to set LANG  \n>>>> across\n>>>> all users of a project is insane...\n>>\n>> FWIW if you use another filesystem, such as reiserfs or ext[2-4], the\n>> filenames will be _unaffected_ by your particular setting of LANG.   \n>> They\n>> will be stored byte-wise exactly like asked for.  That's why I call  \n>> them\n>> \"sane\".\n>\n> One of the advantages (the biggest one, in fact, apart from the  \n> obvious\n> US-ASCII down-compatibility and the fact that you can do C-compatible\n> NUL-terminated strings) of UTF-8 is that it's locale-independent, and\n> doesn't care about LANG, because it's valid in all languages.\n>\n> And that's really important. It's important for a very simple reason:\n> there is almost never such a thing as \"a locale\" except for US- \n> ASCII. Once\n> you move away from US-ASCII, it actually tends to be much more  \n> common that\n> you have a *mixture* of locales - often in the same \"document\" -  \n> than to\n> have one single locale.\n>\n> It very much happens even in filenames - people \"mix\" locales in  \n> trivial\n> ways even within a single pathname component (non-US-ASCII filename,  \n> but\n> with a regular file extension), but much more interestingly they do so\n> within a directory tree (ie you have have translation subdirectories  \n> where\n> the filenames themselves are in another language, and you can have  \n> full\n> pathnames where different components are in different languages, for\n> example).\n>\n> And UTF-8 is _wonderful_ for this, because LANG doesn't matter, and\n> cannot matter, and thus mixing isn't a problem.\n>\n> Of course, you can screw it up. Locales still can change things like  \n> sort\n> order and capitalization etc, so even if you use UTF-8, you sure can  \n> get\n> into trouble with LANG and thinking that a per-session locale makes  \n> sense.\n>\n> So choosing UTF-8 for the filesystem isn't wrong per se. It's a fine\n> choice, and has no issues with LANG in itself. Limiting it to strictly\n> valid UTF-8 encodings is also fine. Limiting it (further) to only\n> character normalized UTF-8 is also fine.\n>\n> Most Linux filesystems don't limit it in any way, so you can make\n> filenames that aren't valid UTF-8 at all, much less normalizing\n> multi-character sequences.\n>\n> I personally think that's the best option, but I probably do so mostly\n> because I know some people still use Latin1 as their only locale  \n> (and I\n> suspect Asia will take decades before it has converted to UTF-8 and  \n> will\n> also have cases where they use other non-UTF locales).\n>\n> But enforcing clean UTF-8 is not a bad idea per se. Not allowing byte\n> sequences that aren't a valid UTF-8 encoding (eg \\xc0\\xc0 is not a  \n> valid\n> UTF-8 character) is fine.\n>\n> I wouldn't call people crazy for doing that, although it does mean  \n> that\n> you cannot, for example, decide to write a Latin1 filename (which is  \n> not\n> necessarily a *good* idea in this day and age, but I think there's a\n> difference between \"that's not a good idea\" and \"you cannot do that\").\n>\n> And even limiting the UTF-8 charset further to only the minimal\n> representation of one particular glyph (ie not allowing multi- \n> character\n> sequences that can be represented more simply) may be even *more*\n> big-brother, but would at least not cause the technical aliasing  \n> issues. I\n> personally think that's so controlling as to be stupid (and has no  \n> real\n> advantage), but hey, at least it doesn't *corrupt* anything silently.\n>\n> So I think that using UTF-8 as a character encoding is a *good*  \n> thing to\n> do, and that automatically means that LANG shouldn't matter for  \n> filenames,\n> but within that choice of UTF-8 there are still mistakes that you can\n> make. Notably multi-character normalization and case-insensitivity.\n>\n> \t\t\tLinus\n\nAlright, you've made your point, and I'm willing to concede at least  \nsome of what you've said. So perhaps we can now move onto the more  \nrelevant and practical issue of: HFS+, despite how stupid it may or  \nmay not be, normalizes filenames (and is case-insensitive, which is a  \nrelated issue). This causes a problem with git. How can this be solved?\n\nI'm more than willing to do work to solve it, my biggest issue is I  \ndon't believe I actually have the free time to learn the git internals  \nwell enough to actually do proper work on what I would assume is a  \nfairly performance-critical section of git's code. However, I would be  \nhappy to work with others who are perhaps more knowledgeable in this  \narea.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65680","messageId":"alpine.LFD.1.00.0801161959210.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"8AC4CC86-A711-483D-9F9C-5F8497006A1D@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T04:08:48Z","receivedAt":"2008-01-17T04:08:48Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 16 Jan 2008, Kevin Ballard wrote:\n> \n> I believe it exists because HFS+ was created at a time when the Mac was moving\n> from a multi-encoding world (which was a nightmare) to a Unicode world and\n> they wanted to remove ambiguity in filenames. But I wasn't around when they\n> made this decision so this is just a guess.\n\nI do agree. And I think starting out case-insensitive (something they must \nreally hate by now) also made it less of an issue. When you're \ncase-insensitive, the issues with any UTF-8 normalization are simply \nswamped by all the issues of case, so you probably don't even think about \nit very much.\n\nThe big problem with any name rewriting is that I can open file 'xyz', and \nI literally have a very hard time knowing whether that file I know I \nopened and created has anything to do with the file 'Xyz' that I see when \nI do a readdir().\n\nAre they the same? Maybe. But it's literally hard to tell on OS X. I can \ndo an fstat() on my file descriptor and on the directory entry, and if \nthey get the same d_ino they *probably are the same entry, but even then \nit actually could have been a hardlink (and my 'xyz' is really *another* \nname for it entirely, and the filesystem is actually case-sensitive and \n'Xyz' was a *different* name that somebody else did!).\n\nSee? If you're creating a content tracker, these kinds of issues are not \n\"idle chatter\". It's really *really* important. Was that file the one I \nwas told to track? Or was it a temporary file that was just hardlinked? \n\nThis is why case-insensitivity is so hard: you have a very real \"aliasing\" \non the filesystem level, where all those really *different* pathnames end \nup being the same thing.\n\nAnd all the same issues show up with utf-8 rewriting, so if you normalize \nutf-8 names, you actually end up having almost all the same problems that \na case-insensitive filesystem has. They're just much rarer in practice, so \nyou just won't hit them as often - but when you do, they are equally \npainful!\n\n(In fact, they can be a whole lot *more* painful, because now they are \nreally rare, and really confusing when they happen!)\n\nBut if you come from a case-insensitive background, all the UTF-8 \nrewriting really looks like such a small problem compared to all the \nhorrid problems that you had with different locales and cases, so I \nsuspect they didn't even realize what a big mistake they did!\n\n\t\t\tLinus\n"},{"id":"65681","messageId":"4C21C1AF-40B0-48C7-8F0E-2DAF3C5FAB29@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161959210.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-17T04:30:01Z","receivedAt":"2008-01-17T04:30:01Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 11:08 PM, Linus Torvalds wrote:\n\n> On Wed, 16 Jan 2008, Kevin Ballard wrote:\n>>\n>> I believe it exists because HFS+ was created at a time when the Mac  \n>> was moving\n>> from a multi-encoding world (which was a nightmare) to a Unicode  \n>> world and\n>> they wanted to remove ambiguity in filenames. But I wasn't around  \n>> when they\n>> made this decision so this is just a guess.\n>\n> I do agree. And I think starting out case-insensitive (something  \n> they must\n> really hate by now) also made it less of an issue. When you're\n> case-insensitive, the issues with any UTF-8 normalization are simply\n> swamped by all the issues of case, so you probably don't even think  \n> about\n> it very much.\n\nThose of us who grew up on a case-insensitive filesystem don't find  \nthere to be any problem with it. I can count on one hand the number of  \ntimes I've run into a problem caused by a case-insensitive filesystem.  \nThat number is 1. And that 1 time is when git screwed up trying to  \ntrack CS4536 and cs4536 in the same directory (see earlier thread).\n\n> The big problem with any name rewriting is that I can open file  \n> 'xyz', and\n> I literally have a very hard time knowing whether that file I know I\n> opened and created has anything to do with the file 'Xyz' that I see  \n> when\n> I do a readdir().\n\nThat's only true if you don't know what type of filesystem you're on.  \nAnd, in the vast majority of cases (in fact, a content tracker is the  \nonly exception I can think of), it doesn't matter. If the user said  \n'xyz' and you can stat() it, great, that's what the user wanted! Just  \nbecause it's really called 'Xyz' on the filesystem doesn't make any  \ndifference.\n\n> Are they the same? Maybe. But it's literally hard to tell on OS X. I  \n> can\n> do an fstat() on my file descriptor and on the directory entry, and if\n> they get the same d_ino they *probably are the same entry, but even  \n> then\n> it actually could have been a hardlink (and my 'xyz' is really  \n> *another*\n> name for it entirely, and the filesystem is actually case-sensitive  \n> and\n> 'Xyz' was a *different* name that somebody else did!).\n>\n> See? If you're creating a content tracker, these kinds of issues are  \n> not\n> \"idle chatter\". It's really *really* important. Was that file the  \n> one I\n> was told to track? Or was it a temporary file that was just  \n> hardlinked?\n\nBut git is a content tracker, so even if it's really a different  \nhardlink that shouldn't matter, it's still referencing the same  \ncontent. Go ahead and track whatever name the user specified  \noriginally, as long as it maps to a file on disk with the expected  \ncontent you're set. If the file is really called 'foo' and I told git  \nto track 'Foo', I'm perfectly happy with it continuing to think 'foo'  \nis 'Foo' until I use 'git mv Foo foo'.\n\n> This is why case-insensitivity is so hard: you have a very real  \n> \"aliasing\"\n> on the filesystem level, where all those really *different*  \n> pathnames end\n> up being the same thing.\n\nI don't see that as being a problem. Think of it, if you will, as if  \nevery single file simply had an implicit hardlink for every possible  \ncase or normalization variant. The whole point of the filename is that  \nit is meta-information, used as an identifier and not as actual  \ncontent, and thus it is perfectly fine for it to be a real string,  \nsubject to interpretation, rather than treated as a sacred binary blob  \nlike content is. The whole purpose of the name is to identify the  \ninode in question, and case and normalization aren't particularly  \nrelevant here. As long as we can identify the file, we're happy.\n\n> And all the same issues show up with utf-8 rewriting, so if you  \n> normalize\n> utf-8 names, you actually end up having almost all the same problems  \n> that\n> a case-insensitive filesystem has. They're just much rarer in  \n> practice, so\n> you just won't hit them as often - but when you do, they are equally\n> painful!\n>\n> (In fact, they can be a whole lot *more* painful, because now they are\n> really rare, and really confusing when they happen!)\n>\n> But if you come from a case-insensitive background, all the UTF-8\n> rewriting really looks like such a small problem compared to all the\n> horrid problems that you had with different locales and cases, so I\n> suspect they didn't even realize what a big mistake they did!\n\nAgain, as someone who grew up in a case-insensitive world, there's no  \nproblems here. I wish I could tell you that it causes problems, I wish  \nI could agree with you, but I can't.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65682","messageId":"76718490801162043w3884435ex435f38b9de837540@mail.gmail.com","threadId":"11645","inReplyTo":"478E1FED.5010801@web.de","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Jay Soffian","fromEmail":"jaysoffian+git@gmail.com","sentAt":"2008-01-17T04:43:14Z","receivedAt":"2008-01-17T04:43:14Z","isPatch":false,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"FWIW, here's Sun's take on the issue of filesystems and i18n:\n\nhttp://developers.sun.com/global/products_platforms/solaris/reference/presentations/IUC29-FileSystems.pdf\n\nj.\n"},{"id":"65683","messageId":"46a038f90801162051s5ce40abcm623599269943a24@mail.gmail.com","threadId":"11645","inReplyTo":"4C21C1AF-40B0-48C7-8F0E-2DAF3C5FAB29@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-17T04:51:40Z","receivedAt":"2008-01-17T04:51:40Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 17, 2008 5:30 PM, Kevin Ballard <kevin@sb.org> wrote:\n> Those of us who grew up on a case-insensitive filesystem don't find\n> there to be any problem with it. I can count on one hand the number of\n\nI guess you haven't used unix tools much. The ever-popular HEAD perl\nutility (which does an HTTP HEAD against a URL), when installed,\nsilently overwrites the head shell utility, which is used for all\nsorts of things, some even in startup scripts. Ooops! I've been hit by\nthis more than once - and if you google for it, it hurt a lot of\npeople.\n\n> That's only true if you don't know what type of filesystem you're on.\n> And, in the vast majority of cases (in fact, a content tracker is the\n> only exception I can think of), it doesn't matter. If the user said\n\nHmmm. Many important tools - that I wouldn't want to ever fail! - have\nsimilar needs to git. Backup/restore and file replication tools for\nexample.\n\n> > This is why case-insensitivity is so hard: you have a very real\n> > \"aliasing\"\n> > on the filesystem level, where all those really *different*\n> > pathnames end\n> > up being the same thing.\n>\n> I don't see that as being a problem. Think of it, if you will, as if\n> every single file simply had an implicit hardlink for every possible\n> case or normalization variant. The whole point of the filename is that\n\nOk - but how do you track the directory then (in git's terms, the\ntree). There's no way to tell what the user wants. Does the user want\na copy of the file with different capitalization, or is the OS playing\ngames?\n> it is meta-information, used as an identifier and not as actual\n> content, and thus it is perfectly fine for it to be a real string,\n> subject to interpretation,\n\nI don't think you *actually* want it subject to interpretation.\n\n> Again, as someone who grew up in a case-insensitive world, there's no\n> problems here. I wish I could tell you that it causes problems, I wish\n> I could agree with you, but I can't.\n\nProbably because you have been surrounded by tools that have a lot of\nextra code to cope with the case insensitive way of life, and learned\nto not do things that are completely valid, just to avoid trouble.\nWhich is ok, but I don't think it makes the OS design decision\ndefensible.\n\ncheers,\n\n\nm\n"},{"id":"65684","messageId":"76718490801162059i2472cd82va34010caa3130b7e@mail.gmail.com","threadId":"11645","inReplyTo":"76718490801162043w3884435ex435f38b9de837540@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Jay Soffian","fromEmail":"jaysoffian+git@gmail.com","sentAt":"2008-01-17T04:59:53Z","receivedAt":"2008-01-17T04:59:53Z","isPatch":false,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"So here's what I can see as being useful additions to git:\n\n* Allowing a repo to be *optionally* configured to disallow two files\nin a directory that can cause aliasing problems, with options for\nunicode normalization aliasing and/or case-insensitivity aliasing. Can\nthis already be done via hooks and someone just needs to write the\nappropriate hooks?\n\n* Having git warn during checkout if there are files which alias in\nthe working copy filesystem. I guess it might be interesting if there\nwere a mechanism in this situation for telling git which of the\naliases you want checked out, though that doesn't seem like a very\ngood feature.\n\nThoughts (besides \"patches welcomed\")?\n\nj.\n"},{"id":"65685","messageId":"alpine.LFD.1.00.0801162059460.2806@woody.linux-foundation.org","threadId":"11645","inReplyTo":"76718490801162043w3884435ex435f38b9de837540@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T05:11:20Z","receivedAt":"2008-01-17T05:11:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 16 Jan 2008, Jay Soffian wrote:\n>\n> FWIW, here's Sun's take on the issue of filesystems and i18n:\n\nPretty sane, from a quick read-through, although most of it seems to not \nbe about general issue, as about \"let's emulate others correctly on their \nfilesystems\" (ie the rules are different for NTFS and HFS+, little enough \ndiscussion about \"native\" preferred logic).\n\nHowever, while they don't consider normalization on file creates to be the \n\"preferred solution\", they *do* consider filename comparison with \ncanonical equivalence to be that. Which means that you can get the same \nodd problems:\n\n\tfd = open(filename, O_CREAT);\n\t+\n\treaddir()\n\ncan actually return a *different* filename than the one we just created, \nif it already existed in the directory under the different normalization.\n\nSo it's basically \"normalization-preserving, but normalization-ignoring\" \n(the same way many filesystems are case-preserving, but case-ignoring). I \ndon't much like it either, but as with case, the \"preserving\" behaviour is \nprobably the nicer one.\n\nI'd guess the problems are harder to trigger in practice, but you can \nstill get some pretty hairy cases. It's just painful when readdir() and \nyour own file creation doesn't have any obvious 1:1 relationship.\n\n\t\t\tLinus\n"},{"id":"65686","messageId":"7vejchkp6o.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"76718490801162059i2472cd82va34010caa3130b7e@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-17T05:15:43Z","receivedAt":"2008-01-17T05:15:43Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Jay Soffian\" <jaysoffian+git@gmail.com> writes:\n\n> So here's what I can see as being useful additions to git:\n> ...\n> Thoughts (besides \"patches welcomed\")?\n\nI think we already discussed a plan to store normalization\nmapping in the index extension section and use it to avoid\ngetting confused by readdir(3) that lies to us.  Is there any\nmore thing that need to be discussed?\n\nI would presume that we would still add _new_ paths using the\npathname we receive from the user (there is no need for us to be\nsimilarly insane as broken \"normalizing\" filesystems), but when\ndeciding if a path is new or we already have it in the index\nwould be done by seeing if an entry already exists in the index\nwhose \"normalized\" form is the same as the \"normalized\" form of\nthe given path --- that way we would not add two paths to the\nindex that would \"normalize\" to the same string.\n"},{"id":"65687","messageId":"ACDB98F4-178C-43C3-99C4-A1D03DD6A8F5@sb.org","threadId":"11645","inReplyTo":"46a038f90801162051s5ce40abcm623599269943a24@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-17T05:23:48Z","receivedAt":"2008-01-17T05:23:48Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 16, 2008, at 11:51 PM, Martin Langhoff wrote:\n\n> On Jan 17, 2008 5:30 PM, Kevin Ballard <kevin@sb.org> wrote:\n>> Those of us who grew up on a case-insensitive filesystem don't find\n>> there to be any problem with it. I can count on one hand the number  \n>> of\n>\n> I guess you haven't used unix tools much. The ever-popular HEAD perl\n> utility (which does an HTTP HEAD against a URL), when installed,\n> silently overwrites the head shell utility, which is used for all\n> sorts of things, some even in startup scripts. Ooops! I've been hit by\n> this more than once - and if you google for it, it hurt a lot of\n> people.\n\nI can imagine. However, I've never been hit by such a situation. This  \ndoesn't mean a case-insensitive filesystem is a problem per se, it  \nmeans interactions between a case-insensitive and a case-sensitive  \nfilesystem can be a problem. That doesn't mean either way is \"correct\"  \nit just means both don't work well together.\n\nI like ice cream, and I like steak, but I sure don't think a mixture  \nof steak and ice cream would go well together. Do you?\n\n>> That's only true if you don't know what type of filesystem you're on.\n>> And, in the vast majority of cases (in fact, a content tracker is the\n>> only exception I can think of), it doesn't matter. If the user said\n>\n> Hmmm. Many important tools - that I wouldn't want to ever fail! - have\n> similar needs to git. Backup/restore and file replication tools for\n> example.\n\nBoth of which would be replicating the directory contents, not a  \nlisting of files specified by the user. If, as a user, I were to say  \n\"please replicate file FOO\" and the file was really called \"foo\", I  \nwouldn't be in the least surprised to see the tool take me at my word  \nand produce a file called \"FOO\" with the contents of \"foo\". But in  \ngeneral, things like this operate on the filesystem, not on the user  \nargs.\n\n>>> This is why case-insensitivity is so hard: you have a very real\n>>> \"aliasing\"\n>>> on the filesystem level, where all those really *different*\n>>> pathnames end\n>>> up being the same thing.\n>>\n>> I don't see that as being a problem. Think of it, if you will, as if\n>> every single file simply had an implicit hardlink for every possible\n>> case or normalization variant. The whole point of the filename is  \n>> that\n>\n> Ok - but how do you track the directory then (in git's terms, the\n> tree). There's no way to tell what the user wants. Does the user want\n> a copy of the file with different capitalization, or is the OS playing\n> games?\n\nIf I say \"track FOO\", I probably mean it. So go ahead and track \"FOO\",  \neven if you end up tracking the contents of file \"foo\". I certainly  \nwon't blame the tool for doing what I told it.\n\n>> it is meta-information, used as an identifier and not as actual\n>> content, and thus it is perfectly fine for it to be a real string,\n>> subject to interpretation,\n>\n> I don't think you *actually* want it subject to interpretation.\n\nSure I do. I find it  very convenient, for example, to say \"cd  \ndocuments/school\" when I really want to go to \"Documents/School\".  \nSimilarly, if I'm trying to reference gitweb/tests/Märchen, I'm quite  \nhappy to not have to figure out what normalization the filename is  \nusing and attempt to replicate that (especially as I have no idea  \nwhich normalization my input mechanism uses - unlike Linus, I don't  \nhave a key dedicated to ä, and even if I did I wouldn't necessarily  \nexpect it to use precomposed vs decomposed). I can't think of a single  \nreason why I'd want to be able to have 2 different files named  \n\"Märchen\" on my disk. On the other hand, treating unicode  \nnormalization as significant can pose security risks - how am I to  \nknow that the file that is named \"foo.txt\" is really the same file  \n\"foo.txt\" that I last saw? Someone I know on IRC sent me this  \nimage[1], which shows 6 files all apparently named \"foo.txt\" on a disk  \nimage. This is possible because on a case-sensitive HFS+ volume, the  \nfile system doesn't ignore ignorables when comparing filenames (it  \ndoes on a case-insensitive HFS+ system), and so all of those filenames  \nlook identical up until you actually pipe their names through xxd and  \nlook at the byte sequence. When this sort of tomfoolery is possible, I  \nsimply cannot trust the names of any of my files anymore.\n\n[1]: http://sailor月.com/imgs/ignorable.png\n\n>> Again, as someone who grew up in a case-insensitive world, there's no\n>> problems here. I wish I could tell you that it causes problems, I  \n>> wish\n>> I could agree with you, but I can't.\n>\n> Probably because you have been surrounded by tools that have a lot of\n> extra code to cope with the case insensitive way of life, and learned\n> to not do things that are completely valid, just to avoid trouble.\n> Which is ok, but I don't think it makes the OS design decision\n\nExtra code? I don't think so. The only reason I'd need extra code is  \nif I were attempting to explicitly detect the \"real\" filename for a  \nuser-supplied argument, by scanning the directory contents until I  \nfound a file that was equivalent to the given argument. But there's no  \nreason to do that. None of the code I've ever written, or any of the  \ncode I've ever seen, has had to do any extra work because it was on a  \ncase-insensitive filesystem. I contribute to a packaging system for  \nthe Mac called MacPorts, and I've never seen any patches on any of the  \n4000+ ports to handle case insensitivity (granted, I haven't looked at  \nevery port, but I've looked at a significant fraction). It's a  \ncomplete non-issue.\n\nThe content of files is sacred. The filename is only there to provide  \na handle to locate the contents. I don't see any problem with  \nexpanding the equivalency scope of the filename to accept multiple  \nencodings and cases. The only arguments I can see that have any  \nvalidity at all are the ones that sound like \"we use case-sensitive  \nfilesystems, and your case-insensitivity and normalization are causing  \nproblems with our tools! Conform to our world!\". As I said above, this  \nisn't a problem of case-insensitivity or normalization, it's a problem  \nof interaction between two incompatible viewpoints. All I want to do  \nis make git play nicer in an HFS+ world, and this would be far easier  \nif you guys were willing to admit this is a problem that should be  \nsolved in the tool rather than a problem with the system.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65690","messageId":"A915BECA-A486-477B-A07D-D1033E44DCBD@adacore.com","threadId":"11645","inReplyTo":"ACDB98F4-178C-43C3-99C4-A1D03DD6A8F5@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Geert Bosch","fromEmail":"bosch@adacore.com","sentAt":"2008-01-17T06:13:26Z","receivedAt":"2008-01-17T06:13:26Z","isPatch":false,"sender":{"key":"bosch@adacore.com","avatar":null},"body":"For those on Mac OS X: it is possible to create a case-sensitive HFS+  \npartition and\nuse it with git. You even can just create a disk image and mount it.  \nHowever,\nI wouldn't quite try to use it as startup filesystem...\n\n   -Geert\n\nPS. I'm working on a proposal/patch for addressing the UFS/case  \nsensitivity issues.\n     Will try to mail something later this week.\n"},{"id":"65693","messageId":"AD012876-3B4A-41EE-8CCB-F60D5C812903@gmail.com","threadId":"11645","inReplyTo":"A915BECA-A486-477B-A07D-D1033E44DCBD@adacore.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mitch Tishmack","fromEmail":"mitcht.git@gmail.com","sentAt":"2008-01-17T07:11:11Z","receivedAt":"2008-01-17T07:11:11Z","isPatch":false,"sender":{"key":"mitcht.git@gmail.com","avatar":null},"body":"I was going to post this earlier, but wanted to search the archives  \nfirst. Here are the commands assuming you don't want to or can't  \npartition a drive and format as ufs (I don't care for HFS+ much). I  \ncan't believe I didn't find the command in the git list archives, so  \nvoilà:\n\n$ hdiutil create -size 300m -fs UFS foo.dmg\n...............................................................................\ncreated: /Users/mitch/foo.dmg\n$ hdiutil attach foo.dmg\n/dev/disk2          \tGUID_partition_scheme          \t\n/dev/disk2s1        \tApple_UFS                      \t/Volumes/untitled\n$ cd /Volumes/untitled && git clone git://git.kernel.org/pub/scm/git/ \ngit.git\n... snipped ...\n$ cd git && git status\n# On branch master\nnothing to commit (working directory clean)\n\nAfter git clone in HFS+ land...\n$ git status\n# On branch master\n# Untracked files:\n#   (use \"git add <file>...\" to include in what will be committed)\n#\n#\tgitweb/test/MaÌrchen\nnothing added to commit but untracked files present (use \"git add\" to  \ntrack)\n\nShould I just add this to the wiki? Then we can all go back to  \nignoring the insane filesystems.\n\nMitch\n\n\nOn Jan 17, 2008, at 12:13 AM, Geert Bosch wrote:\n\n> For those on Mac OS X: it is possible to create a case-sensitive HFS \n> + partition and\n> use it with git. You even can just create a disk image and mount it.  \n> However,\n> I wouldn't quite try to use it as startup filesystem...\n>\n>  -Geert\n>\n> PS. I'm working on a proposal/patch for addressing the UFS/case  \n> sensitivity issues.\n>    Will try to mail something later this week.\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"65694","messageId":"buo1w8gnc4w.fsf@dhapc248.dev.necel.com","threadId":"11645","inReplyTo":"427BE4FD-6534-4CB2-91F8-F9014DC82B54@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Miles Bader","fromEmail":"miles.bader@necel.com","sentAt":"2008-01-17T07:29:19Z","receivedAt":"2008-01-17T07:29:19Z","isPatch":false,"sender":{"key":"miles.bader@necel.com","avatar":"https://gravatar.com/avatar/be062d4050eb88e04229cbdb60f803e1bd647923a015996c2439e76f23e336a7?d=mp&s=160"},"body":"Kevin Ballard <kevin@sb.org> writes:\n> More like, Mac OS X has standardized on Unicode and the rest of the\n> world hasn't caught up yet. Git is the only tool I've ever heard of\n> that has a problem with OS X using Unicode.\n\nApple's decision[*] to use _decomposed_ unicode causes all sorts of\nlittle problems because other tools aren't expecting to see strings\nchanged behind their backs.\n\nI know little about the gritty details, but I see the bug reports...\n\n-Miles\n\n-- \nAny man who is a triangle, has thee right, when in Cartesian Space, to\nhave angles, which when summed, come to know more, nor no less, than\nnine score degrees, should he so wish.  [TEMPLE OV THEE LEMUR]\n.\n"},{"id":"65698","messageId":"B719D4A2-0D05-4C55-95FC-AB880D58E1AC@wincent.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161959210.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-17T10:08:41Z","receivedAt":"2008-01-17T10:08:41Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 17/1/2008, a las 5:08, Linus Torvalds escribió:\n\n> On Wed, 16 Jan 2008, Kevin Ballard wrote:\n>>\n>> I believe it exists because HFS+ was created at a time when the Mac  \n>> was moving\n>> from a multi-encoding world (which was a nightmare) to a Unicode  \n>> world and\n>> they wanted to remove ambiguity in filenames. But I wasn't around  \n>> when they\n>> made this decision so this is just a guess.\n>\n> I do agree. And I think starting out case-insensitive (something  \n> they must\n> really hate by now) also made it less of an issue.\n\nI hope you're right (about them hating it), but we'll see. They've  \njust opened the source for the ZFS port they're working on. By the  \ntime it goes final and becomes the default FS, replacing HFS+,  \nprobably within a couple of years, we'll see if they make the same two  \ndesign decisions which cause the kinds of problems being discussed  \nhere (case-insensitivity, and ubiquitous FS-level UTF-8 normalization).\n\nI've done a dumb search in the ZFS source code for \"CASE\" and see that  \nit can in theory support case-insensitivity as an optional feature.  \nThe potential is there for Apple to use this. I personally hope that  \nthey don't, because as has already been pointed out, these little  \ntricks tend to make life more difficult for users rather than helping  \nthem (the day I have two files in the same directory called \"Märchen\"  \nand want to specify one of them on the command line I'll worry about  \nthat when I come to it).\n\nhttp://fuzzy.wordpress.com/2007/06/09/zfsandfilesystemoptions/\n\nCheers,\nWincent\n"},{"id":"65699","messageId":"17846BF5-1215-4C28-8BBC-2C745A053156@wincent.com","threadId":"11645","inReplyTo":"AD012876-3B4A-41EE-8CCB-F60D5C812903@gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-17T10:22:08Z","receivedAt":"2008-01-17T10:22:08Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 17/1/2008, a las 8:11, Mitch Tishmack escribió:\n\n> I was going to post this earlier, but wanted to search the archives  \n> first. Here are the commands assuming you don't want to or can't  \n> partition a drive and format as ufs (I don't care for HFS+ much). I  \n> can't believe I didn't find the command in the git list archives, so  \n> voilà:\n>\n> $ hdiutil create -size 300m -fs UFS foo.dmg\n> ...............................................................................\n> created: /Users/mitch/foo.dmg\n> $ hdiutil attach foo.dmg\n> /dev/disk2          \tGUID_partition_scheme          \t\n> /dev/disk2s1        \tApple_UFS                      \t/Volumes/untitled\n> $ cd /Volumes/untitled && git clone git://git.kernel.org/pub/scm/git/ \n> git.git\n> ... snipped ...\n> $ cd git && git status\n> # On branch master\n> nothing to commit (working directory clean)\n>\n> After git clone in HFS+ land...\n> $ git status\n> # On branch master\n> # Untracked files:\n> #   (use \"git add <file>...\" to include in what will be committed)\n> #\n> #\tgitweb/test/MaÌˆrchen\n> nothing added to commit but untracked files present (use \"git add\"  \n> to track)\n>\n> Should I just add this to the wiki?\n\nDefinitely.\n\n> Then we can all go back to ignoring the insane filesystems.\n\nWhile it's a nice workaround, it really is just that (a workaround)  \nbecause performance will be suboptimal in a repository running on a  \ndisk image (and many of switched to Git because of its speed).\n\nCheers,\nWincent\n"},{"id":"65700","messageId":"32DB7E53-1062-4F7C-A42D-6EC5945A70A3@wincent.com","threadId":"11645","inReplyTo":"7vejchkp6o.fsf@gitster.siamese.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-17T10:28:52Z","receivedAt":"2008-01-17T10:28:52Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 17/1/2008, a las 6:15, Junio C Hamano escribió:\n\n> \"Jay Soffian\" <jaysoffian+git@gmail.com> writes:\n>\n>> So here's what I can see as being useful additions to git:\n>> ...\n>> Thoughts (besides \"patches welcomed\")?\n>\n> I think we already discussed a plan to store normalization\n> mapping in the index extension section and use it to avoid\n> getting confused by readdir(3) that lies to us.  Is there any\n> more thing that need to be discussed?\n>\n> I would presume that we would still add _new_ paths using the\n> pathname we receive from the user (there is no need for us to be\n> similarly insane as broken \"normalizing\" filesystems), but when\n> deciding if a path is new or we already have it in the index\n> would be done by seeing if an entry already exists in the index\n> whose \"normalized\" form is the same as the \"normalized\" form of\n> the given path --- that way we would not add two paths to the\n> index that would \"normalize\" to the same string.\n\nAnd what do we do when asked to check out a tree which has two  \ndifferent files in it whose normalized forms are the same (ie. a clone  \nof a repo created on a non-HFS+ filesystem)?\n\nWe either have to fail catastrophically, preventing the user from  \nworking with that tree on HFS+, or arbitrarily pick one of the files  \nas the \"winner\" which gets written out into the work tree. None of the  \noptions is particularly attractive, although luckily this exact  \nsituation is unlikely to come up in practice.\n\nCheers,\nWincent\n"},{"id":"65706","messageId":"alpine.LSU.1.00.0801171106510.17650@racer.site","threadId":"11645","inReplyTo":"32DB7E53-1062-4F7C-A42D-6EC5945A70A3@wincent.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T11:10:19Z","receivedAt":"2008-01-17T11:10:19Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\n[Jay, don't cull Cc: lists on vger.kernel.org.  I consider it rude.]\n\nOn Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n\n> El 17/1/2008, a las 6:15, Junio C Hamano escribió:\n> \n> > \"Jay Soffian\" <jaysoffian+git@gmail.com> writes:\n> > \n> > > So here's what I can see as being useful additions to git:\n> > > ...\n> > > Thoughts (besides \"patches welcomed\")?\n> > \n> > I think we already discussed a plan to store normalization mapping in \n> > the index extension section and use it to avoid getting confused by \n> > readdir(3) that lies to us.  Is there any more thing that need to be \n> > discussed?\n\nYes, and I think that a lot of time would have more wisely spent on \nreading that, and trying to implement it, than writing a number of long \nmails, repeating the _same_ (refuted) points over and over again.\n\n> > I would presume that we would still add _new_ paths using the pathname \n> > we receive from the user (there is no need for us to be similarly \n> > insane as broken \"normalizing\" filesystems), but when deciding if a \n> > path is new or we already have it in the index would be done by seeing \n> > if an entry already exists in the index whose \"normalized\" form is the \n> > same as the \"normalized\" form of the given path --- that way we would \n> > not add two paths to the index that would \"normalize\" to the same \n> > string.\n\nAgree.\n\n> And what do we do when asked to check out a tree which has two different \n> files in it whose normalized forms are the same (ie. a clone of a repo \n> created on a non-HFS+ filesystem)?\n> \n> We either have to fail catastrophically, preventing the user from \n> working with that tree on HFS+, or arbitrarily pick one of the files as \n> the \"winner\" which gets written out into the work tree. None of the \n> options is particularly attractive, although luckily this exact \n> situation is unlikely to come up in practice.\n\nAnything else but failure would be Not What You Want.  You might want a \nspecial mode where you use a _different_ name on-disk (something like the \ninfamous short names on FAT), but that _must_ be turned off by default: \nthink of Martin's HEAD example.  Sometimes, it's just not possible to \ncheck such a tree out on a less-than-nice system.\n\nCiao,\nDscho\n"},{"id":"65707","messageId":"C7439732-3B79-4F2B-9D0C-679C1EC8EA0E@simplicidade.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801171106510.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-17T11:23:14Z","receivedAt":"2008-01-17T11:23:14Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"Hi,\n\nOn Jan 17, 2008, at 11:10 AM, Johannes Schindelin wrote:\n> [Jay, don't cull Cc: lists on vger.kernel.org.  I consider it rude.]\n>\n> On Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n>\n>> El 17/1/2008, a las 6:15, Junio C Hamano escribió:\n>>\n>>> \"Jay Soffian\" <jaysoffian+git@gmail.com> writes:\n>>>\n>>>> So here's what I can see as being useful additions to git:\n>>>> ...\n>>>> Thoughts (besides \"patches welcomed\")?\n>>>\n>>> I think we already discussed a plan to store normalization  \n>>> mapping in\n>>> the index extension section and use it to avoid getting confused by\n>>> readdir(3) that lies to us.  Is there any more thing that need to be\n>>> discussed?\n>\n> Yes, and I think that a lot of time would have more wisely spent on\n> reading that, and trying to implement it, than writing a number of  \n> long\n> mails, repeating the _same_ (refuted) points over and over again.\n\nI searched the archives for the posts about normalization and I could  \nnot find them, sorry.\n\nIs stringprep (RFC 3454) being proposed as an optional normalization  \nstep before lookups in the index?\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65708","messageId":"7C01ADFC-6D1D-403F-A917-DBF289AF754C@wincent.com","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801171106510.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-17T11:46:44Z","receivedAt":"2008-01-17T11:46:44Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 17/1/2008, a las 12:10, Johannes Schindelin escribió:\n\n> On Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n>\n>> And what do we do when asked to check out a tree which has two  \n>> different\n>> files in it whose normalized forms are the same (ie. a clone of a  \n>> repo\n>> created on a non-HFS+ filesystem)?\n>>\n>> We either have to fail catastrophically, preventing the user from\n>> working with that tree on HFS+, or arbitrarily pick one of the  \n>> files as\n>> the \"winner\" which gets written out into the work tree. None of the\n>> options is particularly attractive, although luckily this exact\n>> situation is unlikely to come up in practice.\n>\n> Anything else but failure would be Not What You Want.  You might  \n> want a\n> special mode where you use a _different_ name on-disk (something  \n> like the\n> infamous short names on FAT), but that _must_ be turned off by  \n> default:\n> think of Martin's HEAD example.  Sometimes, it's just not possible to\n> check such a tree out on a less-than-nice system.\n\nSuch a special mode would be mostly useless in most contexts, where  \nGit is used to track source code. It would enable you to check out the  \ntree for inspection, but you probably couldn't build anything from it  \nseeing as at least one of the filenames specified in your Makefile  \nwouldn't be present in the work tree.\n\nAs such, in that kind of situation I'd rather see a big red warning  \nprinted out that the checkout failed because a particular file  \ncouldn't be written out, and perhaps an instruction to the user that  \nthey can use \"git show\" if they want to see the blob/s which wasn't/ \nweren't written.\n\nCheers,\nWincent\n"},{"id":"65709","messageId":"16D4755D-EAEC-4F4A-B6B4-F262A6841F66@wincent.com","threadId":"11645","inReplyTo":"C7439732-3B79-4F2B-9D0C-679C1EC8EA0E@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-17T11:51:12Z","receivedAt":"2008-01-17T11:51:12Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 17/1/2008, a las 12:23, Pedro Melo escribió:\n\n> On Jan 17, 2008, at 11:10 AM, Johannes Schindelin wrote:\n>> [Jay, don't cull Cc: lists on vger.kernel.org.  I consider it rude.]\n>>\n>> On Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n>>\n>>> El 17/1/2008, a las 6:15, Junio C Hamano escribió:\n>>>\n>>>> \"Jay Soffian\" <jaysoffian+git@gmail.com> writes:\n>>>>\n>>>>> So here's what I can see as being useful additions to git:\n>>>>> ...\n>>>>> Thoughts (besides \"patches welcomed\")?\n>>>>\n>>>> I think we already discussed a plan to store normalization  \n>>>> mapping in\n>>>> the index extension section and use it to avoid getting confused by\n>>>> readdir(3) that lies to us.  Is there any more thing that need to  \n>>>> be\n>>>> discussed?\n>>\n>> Yes, and I think that a lot of time would have more wisely spent on\n>> reading that, and trying to implement it, than writing a number of  \n>> long\n>> mails, repeating the _same_ (refuted) points over and over again.\n>\n> I searched the archives for the posts about normalization and I  \n> could not find them, sorry.\n>\n> Is stringprep (RFC 3454) being proposed as an optional normalization  \n> step before lookups in the index?\n\nIf this is really just a platform-specific hack, can we use platform- \nspecific code to do the normalization?\n\nOn Mac OS X we have (unfortunately only 10.4 and up):\n\nCFStringCreateWithFileSystemRepresentation()\nCFStringGetFileSystemRepresentation()\nCFStringGetMaximumSizeOfFileSystemRepresentation()\n\nIf we were to use those you'd at least know that you're getting the  \ntrue normalized form as the system defines it.\n\n> The terms used in this Q&A, decomposed and precomposed, roughly  \n> correspond to Unicode Normal Forms D and C, respectively. However,  \n> most volume formats do not follow the exact specification for these  \n> normal forms. For example, HFS Plus uses a variant of Normal Form D  \n> in which U+2000 through U+2FFF, U+F900 through U+FAFF, and U+2F800  \n> through U+2FAFF are not decomposed (this avoids problems with round  \n> trip conversions from old Mac text encodings). It's likely that your  \n> volume format has similar oddities.\n\n\nhttp://developer.apple.com/qa/qa2001/qa1173.html\n\nCheers,\nWincent\n"},{"id":"65711","messageId":"alpine.LSU.1.00.0801171252540.17650@racer.site","threadId":"11645","inReplyTo":"16D4755D-EAEC-4F4A-B6B4-F262A6841F66@wincent.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T12:53:29Z","receivedAt":"2008-01-17T12:53:29Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n\n> On Mac OS X we have (unfortunately only 10.4 and up):\n\nThat remark about the version raises my eyebrows.  Where I live, 10.2.8 is \n_still_ quite common.\n\nCiao,\nDscho\n"},{"id":"65714","messageId":"alpine.LSU.1.00.0801171302410.17650@racer.site","threadId":"11645","inReplyTo":"C7439732-3B79-4F2B-9D0C-679C1EC8EA0E@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T13:05:05Z","receivedAt":"2008-01-17T13:05:05Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 Jan 2008, Pedro Melo wrote:\n\n> On Jan 17, 2008, at 11:10 AM, Johannes Schindelin wrote:\n> > [Jay, don't cull Cc: lists on vger.kernel.org.  I consider it rude.]\n> > \n> > On Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n> > \n> > > El 17/1/2008, a las 6:15, Junio C Hamano escribió:\n> > > \n> > > > \"Jay Soffian\" <jaysoffian+git@gmail.com> writes:\n> > > > \n> > > > > So here's what I can see as being useful additions to git:\n> > > > > ...\n> > > > > Thoughts (besides \"patches welcomed\")?\n> > > > \n> > > > I think we already discussed a plan to store normalization mapping \n> > > > in the index extension section and use it to avoid getting \n> > > > confused by readdir(3) that lies to us.  Is there any more thing \n> > > > that need to be discussed?\n> > \n> > Yes, and I think that a lot of time would have more wisely spent on \n> > reading that, and trying to implement it, than writing a number of \n> > long mails, repeating the _same_ (refuted) points over and over again.\n> \n> I searched the archives for the posts about normalization and I could \n> not find them, sorry.\n\nHere's my pointer:\n\nhttp://thread.gmane.org/gmane.comp.gnu.make.devel/387/focus=61073\n\nFWIW I searched by the term \"readdir\", and then browsed the thread to find \na more interesting post than the first hit.\n\n> Is stringprep (RFC 3454) being proposed as an optional normalization \n> step before lookups in the index?\n\nI don't know.  I'd probably prefer something using iconv (which we use \nalready if it's available), so that the same system can be used for \ncase-insensitivity, UTF-8 normalisation, but also other transformations \nyou might wish to perform.\n\nCiao,\nDscho\n"},{"id":"65719","messageId":"526D0C73-A004-4A2A-A177-274D2602D8DD@wincent.com","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801171252540.17650@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-17T13:40:50Z","receivedAt":"2008-01-17T13:40:50Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 17/1/2008, a las 13:53, Johannes Schindelin escribió:\n\n> Hi,\n>\n> On Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n>\n>> On Mac OS X we have (unfortunately only 10.4 and up):\n>\n> That remark about the version raises my eyebrows.  Where I live,  \n> 10.2.8 is\n> _still_ quite common.\n\nThere may be alternatives that I don't know about.\n\nAll the way back to 10.0 you have -[NSString  \nfileSystemRepresentation], which does the same thing but that's  \nObjective-C. I wouldn't be surprised if that's just a wrapper for the  \nCF functions; that's often the way it is on Mac OS X. And often, the  \nCF functions *are* present on older systems, but they're just not  \ndeclared in public headers. I wouldn't actually recommend using a  \nprivate SPI, but they are often there.\n\nCheers,\nWincent\n"},{"id":"65720","messageId":"15C01F6E-052B-401C-B189-833CBAB20787@sb.org","threadId":"11645","inReplyTo":"17846BF5-1215-4C28-8BBC-2C745A053156@wincent.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-17T13:44:37Z","receivedAt":"2008-01-17T13:44:37Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 17, 2008, at 5:22 AM, Wincent Colaiuta wrote:\n\n> El 17/1/2008, a las 8:11, Mitch Tishmack escribió:\n>\n>> I was going to post this earlier, but wanted to search the archives  \n>> first. Here are the commands assuming you don't want to or can't  \n>> partition a drive and format as ufs (I don't care for HFS+ much). I  \n>> can't believe I didn't find the command in the git list archives,  \n>> so voilà:\n>>\n>> $ hdiutil create -size 300m -fs UFS foo.dmg\n>> ...............................................................................\n>> created: /Users/mitch/foo.dmg\n>> $ hdiutil attach foo.dmg\n>> /dev/disk2          \tGUID_partition_scheme          \t\n>> /dev/disk2s1        \tApple_UFS                      \t/Volumes/ \n>> untitled\n>> $ cd /Volumes/untitled && git clone git://git.kernel.org/pub/scm/ \n>> git/git.git\n>> ... snipped ...\n>> $ cd git && git status\n>> # On branch master\n>> nothing to commit (working directory clean)\n>>\n>> After git clone in HFS+ land...\n>> $ git status\n>> # On branch master\n>> # Untracked files:\n>> #   (use \"git add <file>...\" to include in what will be committed)\n>> #\n>> #\tgitweb/test/MaÌˆrchen\n>> nothing added to commit but untracked files present (use \"git add\"  \n>> to track)\n>>\n>> Should I just add this to the wiki?\n>\n> Definitely.\n>\n>> Then we can all go back to ignoring the insane filesystems.\n>\n> While it's a nice workaround, it really is just that (a workaround)  \n> because performance will be suboptimal in a repository running on a  \n> disk image (and many of switched to Git because of its speed).\n\nNot only is it suboptimal, it's also not acceptable, plain and simple.  \nIf an individual wants to do that, sure, but it's simply not an  \nappropriate solution in general for this problem. I certainly don't  \nwant to have to attach a disk image every time I want access to  \nanything I keep in a git repo, nor do I want to be restricted to  \nkeeping everything within a certain filesystem on disk. Additionally,  \nwhile I'm not certain it's impossible, it's certainly very difficult  \nto attach a disk image without anybody logged into the system at the  \nGUI, as diskarbitrationd won't be running.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65831","messageId":"85zlv435z4.fsf@stiegl.mj.niksun.com","threadId":"11645","inReplyTo":"A915BECA-A486-477B-A07D-D1033E44DCBD@adacore.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Andrew Heybey","fromEmail":"ath@niksun.com","sentAt":"2008-01-17T14:02:39Z","receivedAt":"2008-01-17T14:02:39Z","isPatch":false,"sender":{"key":"ath@niksun.com","avatar":null},"body":"Geert Bosch <bosch@adacore.com> writes:\n\n> For those on Mac OS X: it is possible to create a case-sensitive HFS+\n> partition and\n> use it with git. You even can just create a disk image and mount it.\n> However,\n> I wouldn't quite try to use it as startup filesystem...\n\nThis is starting to stray far afield, but the first thing I did when I\ngot a Macbook was to reinstall it with case-sensitive HFS as the boot\nfile system.  Works fine, including with git.  The only problem I have\nhad is that FileVault does not work.  There are rumored to be some\nthird-part apps that do not work but I do not use that many of those\nanyway.\n\nandrew\n"},{"id":"65723","messageId":"3D987338-EC1E-422F-850D-D8C52345A6A7@sb.org","threadId":"11645","inReplyTo":"85zlv435z4.fsf@stiegl.mj.niksun.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-17T15:04:34Z","receivedAt":"2008-01-17T15:04:34Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 17, 2008, at 9:02 AM, Andrew Heybey wrote:\n\n> Geert Bosch <bosch@adacore.com> writes:\n>\n>> For those on Mac OS X: it is possible to create a case-sensitive HFS+\n>> partition and\n>> use it with git. You even can just create a disk image and mount it.\n>> However,\n>> I wouldn't quite try to use it as startup filesystem...\n>\n> This is starting to stray far afield, but the first thing I did when I\n> got a Macbook was to reinstall it with case-sensitive HFS as the boot\n> file system.  Works fine, including with git.  The only problem I have\n> had is that FileVault does not work.  There are rumored to be some\n> third-part apps that do not work but I do not use that many of those\n> anyway.\n>\n> andrew\n\nThe main problem with this approach is you know for certain that using  \nHFSX as the boot partition is barely tested by Apple, and certainly  \nuntested by third-party apps. This means the potential for breakage is  \nextremely high.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65728","messageId":"alpine.LSU.1.00.0801171556170.5731@racer.site","threadId":"11645","inReplyTo":"15C01F6E-052B-401C-B189-833CBAB20787@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T15:57:20Z","receivedAt":"2008-01-17T15:57:20Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 Jan 2008, Kevin Ballard wrote:\n\n> On Jan 17, 2008, at 5:22 AM, Wincent Colaiuta wrote:\n> \n> > While it's a nice workaround, it really is just that (a workaround) \n> > because performance will be suboptimal in a repository running on a \n> > disk image (and many of switched to Git because of its speed).\n> \n> Not only is it suboptimal, it's also not acceptable, plain and simple.\n\nIf it's not acceptable, do something about it (and I don't mean writing 50 \nemails). If you don't want to do something about it, I have to assume that \nyou accept it as-is.\n\nCiao,\nDscho\n"},{"id":"65729","messageId":"alpine.LFD.1.00.0801170842280.14959@woody.linux-foundation.org","threadId":"11645","inReplyTo":"B719D4A2-0D05-4C55-95FC-AB880D58E1AC@wincent.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T16:43:25Z","receivedAt":"2008-01-17T16:43:25Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nOn Thu, 17 Jan 2008, Wincent Colaiuta wrote:\n>\n> (the day I have two files in the same directory called \"Märchen\" and \n> want to specify one of them on the command line I'll worry about that \n> when I come to it).\n\nSide note: the thing is, the reason people shouldn't worry about it is \nthat this is a *trivial* thing to handle. You really don't even need to \nknow what you're doing. And you can test it today, easily.\n\nHaving two (differently encoded) files like that is really no different \nfrom the traditional UNIX FAQ of \"how do I remove a file starting with \n'-'\" or even more closely \"how do I remove a file that has a character in \nit that I cannot get at the keyboard\".\n\nIn other words, on a bog-standard UNIX (and yes, in this case, I bet OS X \nworks fine too for this test), just try this\n\n\tfilename1=$(echo -e \"hello\\002there\")\n\tfilename2=$(echo -e \"hello\\003there\")\n\techo Odd file > \"$filename1\"\n\techo Another odd file > \"$filename2\"\n\nand now you have a filename that is actually rather hard to type on the \ncommand line. In fact, for me they even *look* the same:\n\n\t[torvalds@woody ~]$ ll hello*\n\t-rw-rw-r-- 1 torvalds torvalds  9 2008-01-17 08:23 hello?there\n\t-rw-rw-r-- 1 torvalds torvalds 17 2008-01-17 08:23 hello?there\n\nSee?\n\nEven in my graphical browser, those two filenames look 100% *identical*. I \ncould give you a screen-shot, but I'm lazy. Just take my word for it, or \njust fire up konqueror on Linux (but it may well depend on the particular \nfont you're using).\n\n[ And yes, for other browsers, you might have something that shows them as \n  different characters - depending on the font, it might show up as a \n  small box with [00 02] vs [00 03] in it, for example. But that's also \n  actually 100% true of the two different encodings of 'ä' - you could \n  easily have a file broswer that shows the multi-character as a \n  multi-character, exactly to distinguish them and show that one of them \n  isn't \"normalized\"!\n\n  The point is, once the filesystem doesn't corrupt the data, it's always \n  easy to get at, and there is never any ambiguity. ]\n\nHow is this different from \"Märchen\" spelled with two different encodings \nfor that \"ä\"?\n\nI'll tell you: it's not at all different. It's 100% the exact same issue.\n\nAnd does that make you perhaps go \"Hunh? How do I remove it, or open it?\"\n\nAnd the fact is, those \"idential looking\" filenames (and thus they must be \nthe same, and something should have normalized them to the same thing, \nno?) are obviously two different files, and they are *really*easy* to edit \nand look at.\n\nFire up that graphical browser again, and it doesn't even matter whether \nthe filename looks identical or not, it shows up as two different files, \nand you can drag them around independently, rename them there, and at \nleast my file browser shows clearly which is which, because I get a small \nicon with a preview in it, so I directly see which one is the \"Odd file\" \nand which one is the \"Another odd file\".\n\nSo the whole \"but they _look_ the same\" argument is just total BS. In just \nabout all character encodings there has always been unique and different \n\"characters\" that _look_ the same on screen, and it has never really made \nthem actually *be* the same, and it has never been a valid argument for \nthem being considered the same.\n\nBecause even when they *look* the same, that file browser that didn't show \nthe difference in names visually, still showed them correctly as two \nseparate files, and I could still just rename them by hand by \nright-clicking on them and picking \"rename\". \n\nSo \"look the same\" is really not a new thing, nor is it even a really hard \nthing. Yes, people can get confused by it, but hey, people can get \nconfused by *anything*. People get confused by filenames starting with a \n\"-\", yet nobody sane really says that filenames cannot start with a dash.\n\n\t\t\tLinus\n"},{"id":"65730","messageId":"2010BC03-E5AE-4333-96CA-4A9B700AD720@sb.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801171556170.5731@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-17T16:53:56Z","receivedAt":"2008-01-17T16:53:56Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 17, 2008, at 10:57 AM, Johannes Schindelin wrote:\n\n> On Thu, 17 Jan 2008, Kevin Ballard wrote:\n>\n>> On Jan 17, 2008, at 5:22 AM, Wincent Colaiuta wrote:\n>>\n>>> While it's a nice workaround, it really is just that (a workaround)\n>>> because performance will be suboptimal in a repository running on a\n>>> disk image (and many of switched to Git because of its speed).\n>>\n>> Not only is it suboptimal, it's also not acceptable, plain and  \n>> simple.\n>\n> If it's not acceptable, do something about it (and I don't mean  \n> writing 50\n> emails). If you don't want to do something about it, I have to  \n> assume that\n> you accept it as-is.\n\nI never said I don't want to do anything about it. However, I do  \nbelieve that it will take a significant investment of time and energy  \nto learn all the gooey details of how git handles filenames and how  \nthe index works and all that jazz, which is knowledge that other  \npeople already have. I believe that, for me to solve this problem  \nindependently, it may require so much time that it never gets done  \n(after all, I am fairly busy). However, if other people who already  \nhave this knowledge are willing to help, that would make this task far  \neasier, especially given that if nobody else even acknowledges that  \nthis is a problem I don't have much hope of getting a patch accepted.\n\nSo again, I'm certainly going to try, but working by myself it simply  \nmay never get done.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65743","messageId":"7vfxwwjpv3.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"16D4755D-EAEC-4F4A-B6B4-F262A6841F66@wincent.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-17T17:58:40Z","receivedAt":"2008-01-17T17:58:40Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Wincent Colaiuta <win@wincent.com> writes:\n\n> If this is really just a platform-specific hack, can we use platform- \n> specific code to do the normalization?\n\nUnfortunately, I do not think this can be a platform-specific\nhack.\n\nIf a project wants to be usable on both sane and insane\nfilesystems, people on platforms whose filesystems treat \"foo\"\nand \"Foo\" as two distinct pathnames (and \"Ma<UMLAUT>rchen\" and\n\"M<A-with-UMLAUT>rchen\" as two distinct ones) need to be\nprevented from creating both in their tree objects at the same\ntime.  Once you create two pathnames xt_connmark.c and\nxt_CONNMARK.c in the same tree object in your project, people on\ncase insensitive filesystems cannot work with your project (you\ncannot check out the kernel source tree and work on it on vfat).\n\nThis is exactly the same logic as making autocrlf=safe (or at\nleast 'input') the default for projects that people need to work\nboth on UNIX and Windows, which Steffen Prohaska has been\nadovocating in another thread.\n"},{"id":"65745","messageId":"478F99E7.1050503@web.de","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801170842280.14959@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-17T18:09:43Z","receivedAt":"2008-01-17T18:09:43Z","isPatch":false,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Linus Torvalds schrieb:\n\n> In other words, on a bog-standard UNIX (and yes, in this case, I bet OS X \n> works fine too for this test), just try this\n> \n> \tfilename1=$(echo -e \"hello\\002there\")\n> \tfilename2=$(echo -e \"hello\\003there\")\n> \techo Odd file > \"$filename1\"\n> \techo Another odd file > \"$filename2\"\n> \n> and now you have a filename that is actually rather hard to type on the \n> command line. In fact, for me they even *look* the same:\n> \n> \t[torvalds@woody ~]$ ll hello*\n> \t-rw-rw-r-- 1 torvalds torvalds  9 2008-01-17 08:23 hello?there\n> \t-rw-rw-r-- 1 torvalds torvalds 17 2008-01-17 08:23 hello?there\n> \n> See?\n\nSorry, but you're using different characters that look the same. But \nKevins point was that it's a different thing if you use two characters \nthat look the same or the same character with different encodings. This \nmakes this HFS-specific problem different from the \"look the same\"- or \nthe \"case-insensitivity\"-issues.\n\nBTW: I also read about your argument that you wouldn't convert file data \nto normalized UTF-8 (I agree with you that this would be nonsense) and \ntherefore filenames shouldn't be converted too. This is something where \nI have to disagree because a filename (like ctime, mtime, atime, ...) \nare meta data (while file contents isn't) and - until now - I would've \nguessed that you agree on this point because git doesn't care about \nfilenames but contents.\n\nIMHO it would be the best solution when git stores all string meta data \nin UTF-8 and converts it to the target systems file system encoding. \nThat would fix all those problems with different locales and file system \nencodings ...\n\nHowever, I have to agree that the enforced character set conversion \ncauses more problems than it solves.\n\nRegards,\nMark\n"},{"id":"65747","messageId":"6E1A0E9A-34D7-4D85-BD4B-CF56CE3927CA@simplicidade.org","threadId":"11645","inReplyTo":"478F99E7.1050503@web.de","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-17T18:12:53Z","receivedAt":"2008-01-17T18:12:53Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"Hi,\n\nOn Jan 17, 2008, at 6:09 PM, Mark Junker wrote:\n> Linus Torvalds schrieb:\n> IMHO it would be the best solution when git stores all string meta  \n> data in UTF-8 and converts it to the target systems file system  \n> encoding. That would fix all those problems with different locales  \n> and file system encodings ...\n\n+1.\n\nAnd I would suggest the use of RFC 3454 as the guidelines for UTF-8  \nnormalization.\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65749","messageId":"alpine.LSU.1.00.0801171817340.5731@racer.site","threadId":"11645","inReplyTo":"6E1A0E9A-34D7-4D85-BD4B-CF56CE3927CA@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T18:18:22Z","receivedAt":"2008-01-17T18:18:22Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 17 Jan 2008, Pedro Melo wrote:\n\n> On Jan 17, 2008, at 6:09 PM, Mark Junker wrote:\n>\n> > IMHO it would be the best solution when git stores all string meta \n> > data in UTF-8 and converts it to the target systems file system \n> > encoding. That would fix all those problems with different locales and \n> > file system encodings ...\n> \n> +1.\n\n-1.\n\nIt's just too arrogant to force your particular preferences down the \nthroat of every git user.\n\nCiao,\nDscho\n"},{"id":"65756","messageId":"200801171922.48343.johan@herland.net","threadId":"11645","inReplyTo":"7vfxwwjpv3.fsf@gitster.siamese.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2008-01-17T18:22:47Z","receivedAt":"2008-01-17T18:22:47Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Thursday 17 January 2008, Junio C Hamano wrote:\n> Wincent Colaiuta <win@wincent.com> writes:\n> \n> > If this is really just a platform-specific hack, can we use platform- \n> > specific code to do the normalization?\n> \n> Unfortunately, I do not think this can be a platform-specific\n> hack.\n> \n> If a project wants to be usable on both sane and insane\n> filesystems, people on platforms whose filesystems treat \"foo\"\n> and \"Foo\" as two distinct pathnames (and \"Ma<UMLAUT>rchen\" and\n> \"M<A-with-UMLAUT>rchen\" as two distinct ones) need to be\n> prevented from creating both in their tree objects at the same\n> time.  Once you create two pathnames xt_connmark.c and\n> xt_CONNMARK.c in the same tree object in your project, people on\n> case insensitive filesystems cannot work with your project (you\n> cannot check out the kernel source tree and work on it on vfat).\n\nIMHO, support for insane filesystems should be split into two parts:\n\n1. A git config setting (probably in .gitattributes) that is enabled\n   by the project to prevent anyone from committing files that would\n   cause problems on insane filesystems. This setting must be enabled\n   for everybody in the project (which is why it cannot easily be\n   solved by the current hooks infrastructure which is per-repo only).\n\n2. A platform-specific hack that detects whenever you're about to\n   check out a problematic filename on an insane filesystem.\n   The hack should either warn or (probably better) FAIL to check out\n   the problematic file(s) (with an appropriate error message\n   pointing at the setting in (1)).\n\nAFAICS, _both_ are needed in order to solve this problem properly.\n\n\nHave fun!\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"65753","messageId":"478FA030.9010807@web.de","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801171817340.5731@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-17T18:36:32Z","receivedAt":"2008-01-17T18:36:32Z","isPatch":false,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Johannes Schindelin schrieb:\n\n> It's just too arrogant to force your particular preferences down the \n> throat of every git user.\n\nIt's not arrogant to make a suggestion. Where is your alternative solution?\n\nHowever, what about storing an additional information like the file \nsystem encoding (for every file)? This would result in the same \nbehaviour (and speed) as today as long as the file system encoding is \nthe same. Conversion will only be done when the targets file system \nencoding is different.\n\nBTW: This reminds me of the code page switching stuff back in the times \nof MS-DOS 4/5. This really wasn't funny.\n\nRegards,\nMark\n"},{"id":"65755","messageId":"10F7A0B4-AF3C-456A-BC2A-7687FF264E31@simplicidade.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801171817340.5731@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-17T18:38:47Z","receivedAt":"2008-01-17T18:38:47Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"Hi,\n\nOn Jan 17, 2008, at 6:18 PM, Johannes Schindelin wrote:\n> On Thu, 17 Jan 2008, Pedro Melo wrote:\n>> On Jan 17, 2008, at 6:09 PM, Mark Junker wrote:\n>>\n>>> IMHO it would be the best solution when git stores all string meta\n>>> data in UTF-8 and converts it to the target systems file system\n>>> encoding. That would fix all those problems with different  \n>>> locales and\n>>> file system encodings ...\n>>\n>> +1.\n>\n> -1.\n>\n> It's just too arrogant to force your particular preferences down the\n> throat of every git user.\n\nDo you agree that you need to store or at least calculate a  \nnormalized version of each filename to see if you are already  \ntracking the file, to take in account all the the filesystems out  \nthere who are not case-preserving, case-sensitive?\n\nIf so, do you think those rules should be an option? Or a preference?\n\nShould I specify in my config file that I want my filenames to be  \nnormalized?\n\nIgnoring encoding, and case-sensitive issues in the git index creates  \nproblems for those people who want/need to use non-ascii chars in  \ntheir filenames, and have some change of being able to collaborate  \nwith other users on different operating systems.\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65757","messageId":"alpine.LFD.1.00.0801171017460.14959@woody.linux-foundation.org","threadId":"11645","inReplyTo":"478F99E7.1050503@web.de","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T18:42:11Z","receivedAt":"2008-01-17T18:42:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 17 Jan 2008, Mark Junker wrote:\n> \n> Sorry, but you're using different characters that look the same. But Kevins\n> point was that it's a different thing if you use two characters that look the\n> same or the same character with different encodings.\n\nBut that's exactly the case he gave - 'ä' vs 'a¨' are exactly that: \ndifferent strings (not even characters: the second is actually a \nmulti-character) that just look the same.\n\nYou try to twist the argument by just claiming that they are the same \n\"character\". They aren't, unless you *define* character to be the same as \n\"glyph\". Of course, if you claim that, then you can always support your \nargument, but I claim that is a bogus and incorrect axiom to start with!\n\nToo many people confuse \"character\" and \"glyph\". They are different.\n\nSee, for example\n\n\thttp://en.wikipedia.org/wiki/Unicode\n\nand notice the *many* places where they try to make that distinction \nbetween \"character\" and \"glyph\" clear (and also \"code values\", which are \nthe actual bytes that encode a character).\n\nSee also\n\n\thttp://en.wikipedia.org/wiki/Unicode_normalization\n\nand realize that a Unicode sequence is a sequence of *characters* even if \nit is not normalized! Those things are still characters, when they are the \n\"simpler\" non-combined characters.\n\nYou are trying to make a totally BOGUS argument, and you base it on the \nINCORRECT basis that the TWO characters 'a'+'¨' somehow aren't independent \ncharacters. They *are*. They are *different* characters from 'ä', even \nthough they may be \"Canonically equivalent\" as a sequence.\n\nThe fact is that \"equivalent\" does not mean \"same\". Why cannot people \naccept that?\n\n\t\t\tLinus\n"},{"id":"65758","messageId":"alpine.LFD.1.00.0801171042330.14959@woody.linux-foundation.org","threadId":"11645","inReplyTo":"6E1A0E9A-34D7-4D85-BD4B-CF56CE3927CA@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T18:44:43Z","receivedAt":"2008-01-17T18:44:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 17 Jan 2008, Pedro Melo wrote:\n> \n> On Jan 17, 2008, at 6:09 PM, Mark Junker wrote:\n>\n> > IMHO it would be the best solution when git stores all string meta data in\n> > UTF-8 and converts it to the target systems file system encoding. That would\n> > fix all those problems with different locales and file system encodings ...\n> \n> +1.\n> \n> And I would suggest the use of RFC 3454 as the guidelines for UTF-8\n> normalization.\n\nThe problem is that there is no way to know what the \"target system \nencoding\" is.\n\nAnd it wouldn't actually solve the bigger problem on OS X anyway: as long \nas you are case-insensitive, you'll have all the same problems (ie the \ninsane OS X filesystem presumably thinks that \"MÄRCHEN\" and \"Märchen\" are \nalso identical, because they are \"equivalent\" names).\n\n\t\t\tLinus\n"},{"id":"65759","messageId":"478FA373.20904@web.de","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801171017460.14959@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-17T18:50:27Z","receivedAt":"2008-01-17T18:50:27Z","isPatch":false,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Linus Torvalds schrieb:\n\n> You try to twist the argument by just claiming that they are the same \n> \"character\". They aren't, unless you *define* character to be the same as \n> \"glyph\". Of course, if you claim that, then you can always support your \n> argument, but I claim that is a bogus and incorrect axiom to start with!\n\nAhhhh ... now I understand.\n\nRegards,\nMark\n"},{"id":"65761","messageId":"F666FFD2-9777-47EA-BEF4-C78906CA8901@simplicidade.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801171017460.14959@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-17T18:52:57Z","receivedAt":"2008-01-17T18:52:57Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"Hi,\n\nOn Jan 17, 2008, at 6:42 PM, Linus Torvalds wrote:\n> Too many people confuse \"character\" and \"glyph\". They are different.\n\nThis is very true.\n\n\n> The fact is that \"equivalent\" does not mean \"same\". Why cannot people\n> accept that?\n\nI'll shut up now if you can answer me one question,  because it  \nreally is a problem for my team.\n\nWe have people using windows, people using Macs, and people using  \nseveral flavors of Linux desktops. They all have different settings  \nand if I add a file like áéióú that happens to be UTF-8 encoded, it  \nwill reach a iso-latin-1 user as visual garbage. git will track the  \nfile perfectly, we know that, because the sequence of bytes that my  \nsystem used to create the file will be the same on all \"sane\"  \nsystems, but the file will look \"funny\" to some users, and we get  \ncomplaints for some less enlightened ones.\n\nThe answer is that users should not create filenames with non-ascii  \ncharacters if they want a consistent experience, right?\n\nThis is just so that I can write a best practices document to them...\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65762","messageId":"20080117190105.GB5547@mit.edu","threadId":"11645","inReplyTo":"F666FFD2-9777-47EA-BEF4-C78906CA8901@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-17T19:01:05Z","receivedAt":"2008-01-17T19:01:05Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Jan 17, 2008 at 06:52:57PM +0000, Pedro Melo wrote:\n> The answer is that users should not create filenames with non-ascii \n> characters if they want a consistent experience, right?\n>\n> This is just so that I can write a best practices document to them...\n\nThat's the easist thing to do if you want to assure that things will\nmostly work across multiple different OS's, with different levels of\nsanity.  You might also want to include that it's a bad idea to create\ntwo filenames that are identical on case-insensitive filesystems,\ni.e., \"makefile\" and \"Makefile\", or \"foo.H\" and \"foo.h\" which even\nthough it works Just Fine on Linux, will likely cause problems on\nWindows and MacOS filesystems, and other systems that are insane with\nrespect to case insensitivity.\n\n\t\t\t\t\t\t\t- Ted\n"},{"id":"65763","messageId":"B4BACE1E-C408-4229-A9AB-FBBAA0200019@simplicidade.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801171042330.14959@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-17T19:02:00Z","receivedAt":"2008-01-17T19:02:00Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"Hi,\n\nOn Jan 17, 2008, at 6:44 PM, Linus Torvalds wrote:\n> On Thu, 17 Jan 2008, Pedro Melo wrote:\n>>\n>> On Jan 17, 2008, at 6:09 PM, Mark Junker wrote:\n>>\n>>> IMHO it would be the best solution when git stores all string  \n>>> meta data in\n>>> UTF-8 and converts it to the target systems file system encoding.  \n>>> That would\n>>> fix all those problems with different locales and file system  \n>>> encodings ...\n>>\n>> +1.\n>>\n>> And I would suggest the use of RFC 3454 as the guidelines for UTF-8\n>> normalization.\n>\n> The problem is that there is no way to know what the \"target system\n> encoding\" is.\n\nCorrect. Storing or using a normalized version of the filename is  \nonly part of the problem.\n\nThe full problem is:\n\nUser A <-> filesystem A <-#-> git < ...... > git <-#-> filesystem B <- \n > user B.\n\nYou have to encode/decode/normalize on all the <-#-> and there is no  \nmagic bullet. Each user would have to tell git \"Hey I'm using utf-8\"  \nor \"Hey, I'm a masochist using HFS+\".\n\nBut I think its important for git to store the filenames in something  \nthat at least permits this kind of scenario.\n\nAll encoding/decoding/normalization is of course optional, and for  \ngit, it still is a sequence of bytes.\n\n> And it wouldn't actually solve the bigger problem on OS X anyway:  \n> as long\n> as you are case-insensitive, you'll have all the same problems (ie the\n> insane OS X filesystem presumably thinks that \"MÄRCHEN\" and  \n> \"Märchen\" are\n> also identical, because they are \"equivalent\" names).\n\nCorrect. HFS+ has bigger problems. I'm not sure if this is enough to  \nsolve it.\n\nBut it would solve two linux users using different encodings.\n\nAnd given that the filtering layers are optional, you have to  \nconfigure them, it wont bite nobody.\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"65765","messageId":"alpine.LFD.1.00.0801171100330.14959@woody.linux-foundation.org","threadId":"11645","inReplyTo":"F666FFD2-9777-47EA-BEF4-C78906CA8901@simplicidade.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T19:11:27Z","receivedAt":"2008-01-17T19:11:27Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 17 Jan 2008, Pedro Melo wrote:\n>\n> We have people using windows, people using Macs, and people using several\n> flavors of Linux desktops. They all have different settings and if I add a\n> file like áéióú that happens to be UTF-8 encoded, it will reach a iso-latin-1\n> user as visual garbage.\n\nYes.\n\n> git will track the file perfectly, we know that, because the sequence of \n> bytes that my system used to create the file will be the same on all \n> \"sane\" systems, but the file will look \"funny\" to some users, and we get \n> complaints for some less enlightened ones.\n\nI can't really suggest anything else than trying to make everybody use \nUTF-8.\n\n[ Not just for filenames, by the way - this is one of the reasons I think\n  it is so *important* to not corrupt filenames, exactly because this is \n  in no way filename-specific at all, and filenames are generally \"textual \n  data\" exactly the same way a text-file is.\n\n  But only totally insane people think that you should force-normalize \n  text-files, even though all the issues are obviously all the same \n  regardless of whether it's a filename or a word in textfile. ]\n\nAnd yes, I also realize that it's not going to be realistic. We're \nprobably *closer* to that than we used to be, but I don't think you can \neven make Windows think FAT is UTF-8.\n\nI don't know how NTFS works (I know it is Unicode-aware, and I think it \nencodes filenames in UCS-2 or possibly UTF-16, but there is an obvious 1:1 \ntranslation to UTF-8, and since we use C strings, I'd assume/hope Windows \nactually uses that unambiguous translation for any filenames).\n\nUnder modern Linux and OS X, UTF-8 is basically the only way (older Linux \ndistros may be set up for Latin1, but at least the newer ones seem to all \ndefault to a UTF-8 locale).\n\n> The answer is that users should not create filenames with non-ascii characters\n> if they want a consistent experience, right?\n\nOh, absolutely. That takes care of 99.9% of all source projects. Even then \nyou can have problems with case insensitivity (the Linux kernel sources \nare all US-ASCII filenames, for example, but *literally* has many files \nthat are identical if you ignore case, and that's not unheard of).\n\nSo yes, to a first approximation, the answer is to simply avoid using \nanything but US-ASCII. It's seldom a big limitation when talking about \nfilenames.\n\n\t\t\tLinus\n"},{"id":"65787","messageId":"20080117212700.GB14088@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"478F99E7.1050503@web.de","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-17T21:27:00Z","receivedAt":"2008-01-17T21:27:00Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Thu, Jan 17, 2008 at 07:09:43PM +0100, Mark Junker wrote:\n> \n> Sorry, but you're using different characters that look the same. But \n> Kevins point was that it's a different thing if you use two characters \n> that look the same or the same character with different encodings.\n\nNo, the encoding was the same -- UTF-8. MacOSX converts one sequence of\nUnicode characters to *another* sequence, which are canonical equivalent,\nbut being canonical equivalent does not mean they are the same characters.\nIn the same way, as being compatible equivalent does not mean being the\nsame. As well as, being case-insensitive equivalent does not mean being\nthe same... Do you remember DOS? It stored all filenames in upper-case,\nso they original and stored names are case-insensitive equivalent, but\nthey are not the same!\n\nDmitry\n"},{"id":"65789","messageId":"87odbkyuvq.fsf@adler.orangeandbronze.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801170842280.14959@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"JM Ibanez","fromEmail":"jm@orangeandbronze.com","sentAt":"2008-01-17T22:01:13Z","receivedAt":"2008-01-17T22:01:13Z","isPatch":false,"sender":{"key":"jm@orangeandbronze.com","avatar":null},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n> So the whole \"but they _look_ the same\" argument is just total BS. In just \n> about all character encodings there has always been unique and different \n> \"characters\" that _look_ the same on screen, and it has never really made \n> them actually *be* the same, and it has never been a valid argument for \n> them being considered the same.\n\nWith the exception of Unicode. If you check the standard, two Unicode\ncodepoints (i.e. the numeric value that gets stored on disk) *can* map\nto the same character, hence they are the same. They don't just look the\nsame, they are the same character -- even if the codepoints are\ndifferent (i.e. precomposed vs. decomposed characters). In fact, part of\nthe Unicode standard deals with that. (Technically, Unicode calls it\nequivalence, but what the hey).\n\nIn other words, Unicode treats e.g. both U+0065 and U+00E9 as\nfundamentally the same character. This comes even more into play in such\nalphabets as Hangul (Korean) and the Japanese Kana.\n\n\n-- \nJM Ibanez\nSoftware Architect\nOrange & Bronze Software Labs, Ltd. Co.\n\njm@orangeandbronze.com\nhttp://software.orangeandbronze.com/\n"},{"id":"65791","messageId":"alpine.LSU.1.00.0801172209080.5731@racer.site","threadId":"11645","inReplyTo":"87odbkyuvq.fsf@adler.orangeandbronze.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-17T22:09:40Z","receivedAt":"2008-01-17T22:09:40Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 18 Jan 2008, JM Ibanez wrote:\n\n> If you check the standard, two Unicode codepoints (i.e. the numeric \n> value that gets stored on disk) *can* map to the same character, hence \n> they are the same.\n\nAs Linus _already_ pointed out, you are confusing characters with glyphs.\n\nHth,\nDscho\n"},{"id":"65794","messageId":"alpine.LFD.1.00.0801171412420.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"87odbkyuvq.fsf@adler.orangeandbronze.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-17T23:05:23Z","receivedAt":"2008-01-17T23:05:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 18 Jan 2008, JM Ibanez wrote:\n>\n> With the exception of Unicode. If you check the standard, two Unicode\n> codepoints (i.e. the numeric value that gets stored on disk) *can* map\n> to the same character, hence they are the same.\n\nBut if you want to make it clear, you can use \"encoded character\" or yes, \n\"code point\". \n\nBut the thing is, even the unicode standard tends to just say \"character\", \nand a unicode string (for example) is defined to be a sequence of \"code \nunits\" which in turn is about those *encoded* characters, which is all \nabout the code points.\n\nSo you'll find that they are very careful in some technical definition \nparts to talk about \"code points\", but then in other sequences they talk \nabout \"character\" even though they are referring to the actual code point \n(ie the figure literally has the unicode number in it!)\n\nIn fact, they sometimes even talk about \"characters\" in the totally \nnon-encoding meaning of \"glyph\".\n\nSo yes, \"character\" is often ambiguous. It would be good to never use the \nword at all, and only talk about \"code point\" and \"glyph\" and one of the \nwell-defined special terms like \"combining character\" or \"replacement \ncharacter\".\n\nBut to take a representative example from The Unicode Standard, Chapter 2: \n\"Unicode Design Principles\":\n\n  Characters are represented by code points that reside only in a memory \n  representation, as strings in memory, on disk, or in data transmission. \n  The Unicode Standard deals only with character codes.\n\n(any speling mistakes mine). In other words, from the very beginning of \nthe standard, very basic design principles chapter, it starts talking \nabout characters being represented by code points and explicitly says that \nit really only deals with CHARACTER CODES.\n\nYes, I'm sure you can argue ad infinitum that all the \"equivalences\" and \nother crap means that a \"character\" can sometimes mean just about \nanything, but I'd say that it's pretty damn reasonable to equate \"unicode \ncharacter\" with \"code point\" or \"character code\".\n\n\t\t\tLinus\n"},{"id":"65795","messageId":"20080117231006.GA14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"87odbkyuvq.fsf@adler.orangeandbronze.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-17T23:10:06Z","receivedAt":"2008-01-17T23:10:06Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Fri, Jan 18, 2008 at 06:01:13AM +0800, JM Ibanez wrote:\n> \n> With the exception of Unicode.\n\nNice exception...\n\n> If you check the standard,\n\nThe standard of what? Could you provide the exact reference?\n\n> two Unicode\n> codepoints (i.e. the numeric value that gets stored on disk)\n\nDoes the standard say something about disk storage?\n\n> *can* map to the same character, \n\nSo what?\n\n> hence they are the same.\n\nnon-sequitor.\n\n> They don't just look the\n> same, they are the same character\n\nBecause?\n\n> -- even if the codepoints are\n> different (i.e. precomposed vs. decomposed characters).\n\nAnd where exactly does the standard says so?\n\n> In fact, part of\n> the Unicode standard deals with that. (Technically, Unicode calls it\n> equivalence, but what the hey).\n\nSo they are not the same after all? It is just you don't care\nabout what it actually says, right? How about this: Unicode\nprovides a unique number for every character. So, if numbers\nare not the same then by definition of the Unicode standard\nthose characters are different.\n\n> \n> In other words, Unicode treats e.g. both U+0065 and U+00E9 as\n> fundamentally the same character.\n\nThere is no notion \"fundamentally the same character\" in the Unicode\nstandard as far as I know, and the characters you mentioned are very\ndifferent in Unicode:\nhttp://www.fileformat.info/info/unicode/char/0065/index.htm\nhttp://www.fileformat.info/info/unicode/char/00e9/index.htm\nThere have different names, they have different glyphs, and they\nare functional different.\n\nDmitry\n"},{"id":"65798","messageId":"9BA6785F-A715-4F5C-B192-96471243FE20@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801171100330.14959@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-18T00:18:17Z","receivedAt":"2008-01-18T00:18:17Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 17, 2008, at 2:11 PM, Linus Torvalds wrote:\n\n> [ Not just for filenames, by the way - this is one of the reasons I  \n> think\n>  it is so *important* to not corrupt filenames, exactly because this  \n> is\n>  in no way filename-specific at all, and filenames are generally  \n> \"textual\n>  data\" exactly the same way a text-file is.\n\nI just don't understand why you insist that the filename is data, when  \nit is clearly metadata. The filename has two purposes: the identify  \nthe file to the user, and to provide a handle with which to reference  \nthe file contents. The specific byte sequence is in no way sacred.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65799","messageId":"alpine.LFD.1.00.0801171626470.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"9BA6785F-A715-4F5C-B192-96471243FE20@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-18T00:35:19Z","receivedAt":"2008-01-18T00:35:19Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 17 Jan 2008, Kevin Ballard wrote:\n> \n> I just don't understand why you insist that the filename is data, when it is\n> clearly metadata.\n\nUhh. And exactly how do you know the difference, and why should it matter?\n\nA lot of data is metadata. Look at the git index file. It's *all* \nmetadata. Does that mean that the OS has the right to corrupt it?\n\nIOW, why do you seem to argue that metadata something you can corrupt, but \nnot then \"regular\" data?\n\nWhy is it ok to change a filename, when that same filename may *also* be \nencoded by the user in a regular data file (think about MD5SUM files, for \nexample, that include the pathname, but now the pathname is part of the \nfile data, not on a filesystem). \n\nSo filenames are data, they're metadata, they're whatever. None of that \nmeans that it's acceptable to corrupt them, or gives the OS any reason to \nsay that it \"knows better\" than the user in how users use them. It's still \nthe *users* metadata, not the filesystems own metadata!\n\nIn many cases, users use filenames *as* data, ie the filename actually has \na meaning in itself, not just as a handle to get the file contents.\n\nIf this was truly metadata that isn't visible to the user, and not under \nthe users control (ie indirect block numbers etc), then you'd have a good \npoint. At that point, it's obviously entirely up to the filesystem how the \nheck it encodes it.\n\nBut that's not what filenames are. Filenames are an index specified by the \nuser, not by the computer. \n\n\t\tLinus\n"},{"id":"65800","messageId":"200801180144.06253.robin.rosenberg.lists@dewire.com","threadId":"11645","inReplyTo":"2010BC03-E5AE-4333-96CA-4A9B700AD720@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2008-01-18T00:44:04Z","receivedAt":"2008-01-18T00:44:04Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"torsdagen den 17 januari 2008 skrev Kevin Ballard:\n> On Jan 17, 2008, at 10:57 AM, Johannes Schindelin wrote:\n> \n> > On Thu, 17 Jan 2008, Kevin Ballard wrote:\n> >\n> >> On Jan 17, 2008, at 5:22 AM, Wincent Colaiuta wrote:\n> >>\n> >>> While it's a nice workaround, it really is just that (a workaround)\n> >>> because performance will be suboptimal in a repository running on a\n> >>> disk image (and many of switched to Git because of its speed).\n> >>\n> >> Not only is it suboptimal, it's also not acceptable, plain and  \n> >> simple.\n> >\n> > If it's not acceptable, do something about it (and I don't mean  \n> > writing 50\n> > emails). If you don't want to do something about it, I have to  \n> > assume that\n> > you accept it as-is.\n> \n> I never said I don't want to do anything about it. However, I do  \n> believe that it will take a significant investment of time and energy  \n> to learn all the gooey details of how git handles filenames and how  \n> the index works and all that jazz, which is knowledge that other  \n> people already have. I believe that, for me to solve this problem  \n> independently, it may require so much time that it never gets done  \n> (after all, I am fairly busy). However, if other people who already  \n> have this knowledge are willing to help, that would make this task far  \n> easier, especially given that if nobody else even acknowledges that  \n> this is a problem I don't have much hope of getting a patch accepted.\n> \n> So again, I'm certainly going to try, but working by myself it simply  \n> may never get done.\n\n(This is only for those that think the problem should be solved somehow. The\nrest can move on - nothing to see here)\n\nYou may look at http://rosenberg.homelinux.net/cgi-bin/gitweb/gitweb.cgi?p=GIT.git;a=log;h=i18n\nfor inspiration. It's pretty obsolete by now and only a \"proof of concept\", i.e.\nit can be done, not that it necessarily should be done exactly this way.\n\nBasically it intercepts the user's access to git, i.e. certain commands\nand how files are named (since those names represent a user interface). Then\nit assumes the internal encoding is UTF-8 (or garbage) converting to and\nfrom the user's local encoding. The heuristics is based on the assumption that\na string (even random onesthat looks like UTF-8, with a very high probablity\nactually is UTF-8 encoded.\n\nThe test cases might be usable almost as is.\n\n-- robin\n"},{"id":"65802","messageId":"200801180205.28742.robin.rosenberg.lists@dewire.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801171100330.14959@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2008-01-18T01:05:28Z","receivedAt":"2008-01-18T01:05:28Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"torsdagen den 17 januari 2008 skrev Linus Torvalds:\n> And yes, I also realize that it's not going to be realistic. We're \n> probably *closer* to that than we used to be, but I don't think you can \n> even make Windows think FAT is UTF-8.\nIt's UTF-16 (when needed). I think it's all in the Linux kernel for you\nto see.\n\n> I don't know how NTFS works (I know it is Unicode-aware, and I think it \n> encodes filenames in UCS-2 or possibly UTF-16, but there is an obvious 1:1 \nUTF-16 (was UCS-2 until MS did a s/UCS-2/UTF-16/ on the documentation).\n\n> translation to UTF-8, and since we use C strings, I'd assume/hope Windows \n> actually uses that unambiguous translation for any filenames).\n\nIt uses the local 8-bit codepage, which is not UTF-8, often some latin-inspired\nthingy, but in Asia multi-byte encodings are used. In western Europe it is\nWindows-1252, which is almost, but not exactly iso-8859-1. Oh, and then we\nhave the cmd prompt which has another encoding in 8-bit mode.\n\nI think there is a cygwin patch that converts to and from UTF-8. An application\ncan choose to use the \"A\" or \"W\" interfaces. The W-API's are the real ones and \nthe others' are just wrappers that convert to and from UTF-16 before anything\nhappens (i.e. CreateFileA is slower than CreateFileW and so on). \n\n-- robin\n"},{"id":"65803","messageId":"alpine.LFD.1.00.0801171716310.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"200801180205.28742.robin.rosenberg.lists@dewire.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-18T01:24:01Z","receivedAt":"2008-01-18T01:24:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 18 Jan 2008, Robin Rosenberg wrote:\n\n> torsdagen den 17 januari 2008 skrev Linus Torvalds:\n> > And yes, I also realize that it's not going to be realistic. We're \n> > probably *closer* to that than we used to be, but I don't think you can \n> > even make Windows think FAT is UTF-8.\n>\n> It's UTF-16 (when needed). I think it's all in the Linux kernel for you\n> to see.\n\n.. well, FAT certainly wasn't. But yes, VFAT probably is.  Not that I want \nto look at it ;)\n\n> > translation to UTF-8, and since we use C strings, I'd assume/hope Windows \n> > actually uses that unambiguous translation for any filenames).\n> \n> It uses the local 8-bit codepage, which is not UTF-8, often some latin-inspired\n> thingy, but in Asia multi-byte encodings are used. In western Europe it is\n> Windows-1252, which is almost, but not exactly iso-8859-1. Oh, and then we\n> have the cmd prompt which has another encoding in 8-bit mode.\n\nWell, if it uses a 8-bit codepage, then that means that as far as the \nPOSIX filename interface is concerned, it has nothing what-so-ever to do \nwith Unicode (ie unicode is just a totally invisible internal encoding \nissue, not externally visible).\n\nI assume you have to use some insane Windows-only UCS-2 filename function \nto actually see any Unicode behaviour.\n\nSad. Because there really is no reason to use a local 8-bit codepage when \nyou could just use UTF-8.\n\n> I think there is a cygwin patch that converts to and from UTF-8. An application\n> can choose to use the \"A\" or \"W\" interfaces. The W-API's are the real ones and \n> the others' are just wrappers that convert to and from UTF-16 before anything\n> happens (i.e. CreateFileA is slower than CreateFileW and so on). \n\nSo the CreateFileW() is the \"native UTF-16 interface\", and CreateFileA() \nis the 8-bit codepage one that has nothing to do with Unicode and is \npurely some local thing.\n\nBut for a UNIX interface layer, the most logical thing would probably be \nto map \"open()\" and friends not to CreateFileA(), but to \nCreateFileW(utf8_to_utf16(filename)). \n\nOnce you do that, then it sounds like Windows would basically be Unicode, \nand hopefully without any crazy normalization (but presumably all the \ncrazy case-insensitivity cannot be fixed ;^).\n\nSo it probably really only depends on whether you choose to use the insane \n8-bit code page translation or whether you just use a sane and trivial \nUTF8<->UTF16 conversion.\n\nAnybody know which one cygwin/mingw does?\n\n\t\t\tLinus\n"},{"id":"65804","messageId":"200801180227.19242.robin.rosenberg.lists@dewire.com","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801172209080.5731@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2008-01-18T01:27:18Z","receivedAt":"2008-01-18T01:27:18Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"torsdagen den 17 januari 2008 skrev Johannes Schindelin:\n> Hi,\n> \n> On Fri, 18 Jan 2008, JM Ibanez wrote:\n> \n> > If you check the standard, two Unicode codepoints (i.e. the numeric \n> > value that gets stored on disk) *can* map to the same character, hence \n> > they are the same.\n> \n> As Linus _already_ pointed out, you are confusing characters with glyphs.\n> \nSomeone is. \n\nHe is refering to the unicode definition of an (abstract) character.\n\nCh3.4 D11 - \"A single abstract character may also be represented by a sequence\nof code points—for example, latin capital letter g with acute may be represented\nby the sequence <U+0047 latin capital letter g, U+0301 combining acute accent>, \nrather than being mapped to a single code point.\n\n\n-- robin\n"},{"id":"65823","messageId":"47902653.31E38914@dessent.net","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801171716310.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Brian Dessent","fromEmail":"brian@dessent.net","sentAt":"2008-01-18T04:08:51Z","receivedAt":"2008-01-18T04:08:51Z","isPatch":false,"sender":{"key":"brian@dessent.net","avatar":null},"body":"Linus Torvalds wrote:\n\n> But for a UNIX interface layer, the most logical thing would probably be\n> to map \"open()\" and friends not to CreateFileA(), but to\n> CreateFileW(utf8_to_utf16(filename)).\n> \n> Once you do that, then it sounds like Windows would basically be Unicode,\n> and hopefully without any crazy normalization (but presumably all the\n> crazy case-insensitivity cannot be fixed ;^).\n> \n> So it probably really only depends on whether you choose to use the insane\n> 8-bit code page translation or whether you just use a sane and trivial\n> UTF8<->UTF16 conversion.\n> \n> Anybody know which one cygwin/mingw does?\n\nCygwin does not yet support doing the smart thing.  At the moment you\ncan only open() files in the current 8 bit codepage.  There is a patch\nfloating around to allow using UTF-8, but it was rejected for inclusion\nbecause it was considered too hackish.  Instead work has been ongoing\nfor some time to replumb the internal representation of Windows\nfilenames to use UTF-16 instead of plain chars, so that conversion\noverhead can be held at a minimum.  In conjuction with dropping Win9x/ME\nsupport this also means the Native APIs like NtCreateFile() can be used\ndirectly, as they are more low level than the Win32 -A and -W functions\nand expose more flexibility, such as the ability to implement the\nopenat() family of functions natively (no pun intended) without\nemulation.  These two items (unicode and dropping non-NT windows) are\nthe big features for 1.7.\n\nOf course since a lot of what Cygwin does is translate paths in\nsometimes unobvious and complicated ways, there's a lot of path handling\ncode to adapt, so it's taking a while.\n\nIncidently, the ridiculously short MAX_PATH of 260 on Windows comes from\nthe Win32 -A version of the functions.  The -W API and the Native API\ncan cope with paths of up to 32k wide chars, so a side benefit of this\nshould be the ability to finally stop running into length limits.  Of\ncourse there's always a catch: when using long filenames with the Win32\n-W API or the Native API you can only use absolute paths, so either you\nhave to live with the 260 limitation for relative paths or you keep\ntrack of the current directory and always do a rel->abs conversion.  Or\nbetter, if you stick to the Native API you can do a directory handle\nrelative openat-type thing which I suppose starts to sound relatively\nsane.  However, there's another catch here: For some time Cygwin has\nmaintained a separate and private value of CWD behind Windows' back, and\nonly synced the two when spawning a non-Cygwin binary.  This allows\nWindows to happly think the process' CWD is always C:\\ or whatever, and\nnot hold an open handle to the actual CWD.  In turn Cygwin uses this to\nallow POSIX filesystem behavior of being able to unlink the current dir,\nwhich some programs or build systems assume they can do but is not\npossible in straight Win32.  This is a roundabout way of saying that\ngoing back to actually having to keep a handle to CWD open again in\norder to do relative paths might be complicated.\n\nBrian\n"},{"id":"65834","messageId":"Pine.LNX.4.64.0801180902470.817@ds9.cixit.se","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801161615330.2806@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Peter Karlsson","fromEmail":"peter@softwolves.pp.se","sentAt":"2008-01-18T08:29:36Z","receivedAt":"2008-01-18T08:29:36Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Linus Torvalds:\n\n> The difference I see between us is that when I tell you that this is\n> exactly the same thing as your file *contents*,\n\nThis is the same issue as the CRLF issue I posted on earlier, and it\nall stems from that git also sees file names as a stream of bytes, not\na string of characters, just as it does text.\n\n> An OS that silently changes the contents of your files is *crap*.\n> Get it?\n\nA program that silently ignores the conventions of the platform it runs\non is *crap*, no matter if the conventions are not the same as for\nother platforms.\n\n> An OS that silently changes the contents of your directories is *crap*.\n> Get it now?\n\nA program that silently ignores the conventions of the file system it\ntries to store its files on is *crap* :-)\n\n\nIn my perfect world, file names would be stored as a string of characters,\nso if I save a file with an å in it, that å would be preserved no\nmatter if I run Linux on ext2 with my locale is set to latin-1 (which\nstores it as byte 0xE5), on Windows with NTFS (which stores it as the\nUTF-16 code 0x00E5), on Windows/DOS with FAT (which stores it as the\nbyte 0x86) or on Mac OS X which stores it as decomposed UTF-8 (whose\nbyte sequence I don't know at the top of my head). If that was just\nstored as U+00E5 in whatever encoding in the filename index, the local\nimplementation of git can just check it out in the form needed.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"65838","messageId":"20080118084912.GC14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801171716310.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-18T08:49:12Z","receivedAt":"2008-01-18T08:49:12Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Thu, Jan 17, 2008 at 05:24:01PM -0800, Linus Torvalds wrote:\n> \n> On Fri, 18 Jan 2008, Robin Rosenberg wrote:\n> \n> > It uses the local 8-bit codepage, which is not UTF-8, often some latin-inspired\n> > thingy, but in Asia multi-byte encodings are used. In western Europe it is\n> > Windows-1252, which is almost, but not exactly iso-8859-1. Oh, and then we\n> > have the cmd prompt which has another encoding in 8-bit mode.\n\nYes, the default code page for the command prompt uses so-called OEM\nencoding, and GUI programs uses another one, which MS calls as \"ANSI\"\nencoding. However, if you use Cygwin, then you have ANSI encoding in\nthe command prompt. So, in the same command prompt window, you can have\nCygwin programs using one encoding and other window console programs\nusing a different encoding.\n\n> \n> Well, if it uses a 8-bit codepage, then that means that as far as the \n> POSIX filename interface is concerned, it has nothing what-so-ever to do \n> with Unicode (ie unicode is just a totally invisible internal encoding \n> issue, not externally visible).\n\nSome people tried to set the current code page to 65001, which is\nthe Microsoft code page for UTF-8. However, it seems that does not\nwork very well.\n\nhttp://support.microsoft.com/kb/175392\nhttp://blogs.msdn.com/michkap/archive/2006/03/13/550191.aspx\n\nIt seems to me that Win32 API functions work correctly with\nUTF-8 (after all, they are just wrappers over UTF-16 functions),\nbut Microsoft's C library cannot handle UTF-8 (or any other\nencoding that requires more than two bytes per character).\n\n> Anybody know which one cygwin/mingw does?\n\nThere is a patch for Cygwin that adds UTF-8 support for it, however,\nCygwin maintainers do not like it, so it is not integrated. I think\nCygwin 1.7 will support UTF-8, but I have no idea how soon it will be\nreleased.\n\nI don't know much about mingw, but if I am not mistaken, mingw relies\non Microsoft's C library, so I suppose it uses an \"OEM\" code page for\nconsole programs by default.\n\n\nDmitry\n"},{"id":"65842","messageId":"200801181042.37391.robin.rosenberg.lists@dewire.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801171716310.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2008-01-18T09:42:36Z","receivedAt":"2008-01-18T09:42:36Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"fredagen den 18 januari 2008 skrev Linus Torvalds:\n> > > translation to UTF-8, and since we use C strings, I'd assume/hope Windows \n> > > actually uses that unambiguous translation for any filenames).\n> > \n> > It uses the local 8-bit codepage, which is not UTF-8, often some latin-inspired\n> > thingy, but in Asia multi-byte encodings are used. In western Europe it is\n> > Windows-1252, which is almost, but not exactly iso-8859-1. Oh, and then we\n> > have the cmd prompt which has another encoding in 8-bit mode.\n> \n> Well, if it uses a 8-bit codepage, then that means that as far as the \n> POSIX filename interface is concerned, it has nothing what-so-ever to do \n> with Unicode (ie unicode is just a totally invisible internal encoding \n> issue, not externally visible).\n\nI just had to investigate this a bit, so on a Vista machine I started a cmd\nprompt and typed mode con: cp select=65001, selected the lucida font and then\necho å >x.txt and opened it in notepad and it was UTF-8 encoded. So there might\nbe some hope after all. I don't know how to change the encoding for non-console\napps. I leave that as an excercise for the list.\n\n-- robin\n"},{"id":"65848","messageId":"Pine.LNX.4.64.0801181114430.817@ds9.cixit.se","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801171100330.14959@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Peter Karlsson","fromEmail":"peter@softwolves.pp.se","sentAt":"2008-01-18T10:19:21Z","receivedAt":"2008-01-18T10:19:21Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Linus Torvalds:\n\n> But that's exactly the case he gave - 'ä' vs 'a¨' are exactly that: \n> different strings (not even characters: the second is actually a \n> multi-character) that just look the same.\n\nBut they are not different strings, they are canonically equivalent as\nfar as Unicode is concerned. They're even supposed to map to the same\nglyph (if the font has an \"ä\", it should display it in both cases, if\nit has an \"a\" and a combining diaeresis, it should make up one).\n\nYou cannot do a binary comparison of text to see if two strings are\nequivalent.\n\n> You try to twist the argument by just claiming that they are the same\n> \"character\". They aren't, unless you *define* character to be the\n> same as \"glyph\".\n\nWhereas you are confusing characters and code points.\n\n\"ä\" and \"a¨\" use different code points, but they encode the same\ncharacter, and from the user's perspective it is the *character* that\nis interesting (although he might confuse it with the glyph).\n\n\n> I don't know how NTFS works (I know it is Unicode-aware, and I think\n> it encodes filenames in UCS-2 or possibly UTF-16,\n\nActually, NTFS is a bit broken. It sees file names as a string of\n16-bit words. It doesn't check that it is valid UTF-16, or even valid\nUCS-2, it allows almost anything.\n\n\nApple made Mac OS X handle filenames properly, by seeing that file\nnames are a string of characters, not code points, so they use a\ncanonical form for all characters (personally, I would have preferred\nthe pre-composed form, though).\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"65850","messageId":"20080118103036.GD14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"200801181042.37391.robin.rosenberg.lists@dewire.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-18T10:30:36Z","receivedAt":"2008-01-18T10:30:36Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Fri, Jan 18, 2008 at 10:42:36AM +0100, Robin Rosenberg wrote:\n> \n> I just had to investigate this a bit, so on a Vista machine I started a cmd\n> prompt and typed mode con: cp select=65001, selected the lucida font and then\n> echo å >x.txt and opened it in notepad and it was UTF-8 encoded. \n\nYes, but have you tried to run any batch file? At least, on WinXP\nall batch files silently stopped working after choosing 65001, and\nI don't know what else gets broken, because Microsoft C library\ndoes not work with encoding that requires more than two bytes per\ncharacter.\n\n> So there might\n> be some hope after all. I don't know how to change the encoding for non-console\n> apps. I leave that as an excercise for the list.\n\nIt is not difficult to change the current encoding in any Windows\napplication, the real issue is that neither Microsoft C library nor\nCygwin library does not work correctly with UTF-8. There is a patch\nfor Cygwin though...\n\nDmitry\n"},{"id":"65853","messageId":"20080118105040.GE14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"Pine.LNX.4.64.0801181114430.817@ds9.cixit.se","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-18T10:50:40Z","receivedAt":"2008-01-18T10:50:40Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Fri, Jan 18, 2008 at 11:19:21AM +0100, Peter Karlsson wrote:\n> Linus Torvalds:\n> \n> > But that's exactly the case he gave - 'ä' vs 'a¨' are exactly that: \n> > different strings (not even characters: the second is actually a \n> > multi-character) that just look the same.\n> \n> But they are not different strings, they are canonically equivalent as\n> far as Unicode is concerned.\n\nThere are canonically equivalent, but they are different sequences\nof characters as Unicode is concerned. In one case, we have one\ncharacter in the other case, we have two characters that canonically\nequivalent to the first one.\n\n> They're even supposed to map to the same\n> glyph (if the font has an \"ä\", it should display it in both cases, if\n> it has an \"a\" and a combining diaeresis, it should make up one).\n\nBy defition, sequences of characters that are canonically equivalent\nare both visual and functional equivalent...\n\n> You cannot do a binary comparison of text to see if two strings are\n> equivalent.\n\nOf course, you can't. Who argues otherwise?\n\n> > You try to twist the argument by just claiming that they are the same\n> > \"character\". They aren't, unless you *define* character to be the\n> > same as \"glyph\".\n> \n> Whereas you are confusing characters and code points.\n\nI am afraid it is you who confuses \"characters\" with \"abstract\ncharacters\", there is no place in the standard saying that\n\"characters\" are \"abstract characters\" only. On contrary, the\nterm \"characters\" is used to refer non abstract characters.\n\nDmitry\n"},{"id":"65858","messageId":"200801181217.01198.jnareb@gmail.com","threadId":"11645","inReplyTo":"Pine.LNX.4.64.0801180902470.817@ds9.cixit.se","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-01-18T11:16:58Z","receivedAt":"2008-01-18T11:16:58Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Peter Karlsson wrote:\n> Linus Torvalds wrote:\n> \n> > The difference I see between us is that when I tell you that this is\n> > exactly the same thing as your file *contents*,\n> \n> This is the same issue as the CRLF issue I posted on earlier, and it\n> all stems from that git also sees file names as a stream of bytes, not\n> a string of characters, just as it does text.\n\nYou have to be careful about CRLF conversion, lest you corrupt your\nbinary files. CRLF conversion is off by default.\n\n> > An OS that silently changes the contents of your files is *crap*.\n> > Get it?\n> \n> A program that silently ignores the conventions of the platform it runs\n> on is *crap*, no matter if the conventions are not the same as for\n> other platforms.\n> \n> > An OS that silently changes the contents of your directories is *crap*.\n> > Get it now?\n> \n> A program that silently ignores the conventions of the file system it\n> tries to store its files on is *crap* :-)\n\nGit philosophy to see the contents of files and \"contents\" of directories\n(filenames) as stream of bytes, i.e. to use 'native' encoding works\nperfectly well and _fast_ if all developers work in the same environment.\nTroubles start if you are working across operating systems, and across\nfilesystems.\n\n> In my perfect world, file names would be stored as a string of characters,\n> so if I save a file with an å in it, that å would be preserved no\n> matter if I run Linux on ext2 with my locale is set to latin-1 (which\n> stores it as byte 0xE5), on Windows with NTFS (which stores it as the\n> UTF-16 code 0x00E5), on Windows/DOS with FAT (which stores it as the\n> byte 0x86) or on Mac OS X which stores it as decomposed UTF-8 (whose\n> byte sequence I don't know at the top of my head). If that was just\n> stored as U+00E5 in whatever encoding in the filename index, the local\n> implementation of git can just check it out in the form needed.\n\nGit has for a long time i18n.commitEncoding, and from some time it\nsaves it in 'encoding' header in commit object (if different from\n'uft-8') and has also i18n.logOutputEncoding.\n\nFor dealing with different filesystem encodings you would also have\nto have both: encoding used in 'tree' objects (by repository) for\nfilenames saved somewhere in repository, either in tree object (argh!)\nor in some kind of .gitconfig file; encoding used by filesystem in\nrepository config as i18n.filesystemEncoding or something like that.\nAnd think what to put in the on disk index, and in memory index.\n\n\nNOTE, NOTE, NOTE! If filename is used somewherein the file contents\n(manifest-like file, include-like statement), and this filename uses\ncharacters which are differently encoded in different encoding you\nare screwed with this fancy system, badly, anyway.\n\n-- \nJakub Narebski\nPoland\n"},{"id":"65885","messageId":"Pine.LNX.4.64.0801181626060.817@ds9.cixit.se","threadId":"11645","inReplyTo":"20080118105040.GE14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Peter Karlsson","fromEmail":"peter@softwolves.pp.se","sentAt":"2008-01-18T15:30:24Z","receivedAt":"2008-01-18T15:30:24Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Dmitry Potapov:\n\n> I am afraid it is you who confuses \"characters\" with \"abstract\n> characters\", there is no place in the standard saying that\n> \"characters\" are \"abstract characters\" only. On contrary, the term\n> \"characters\" is used to refer non abstract characters.\n\nPerhaps it's just a case of confusion about naming conventions. I tend\nto use \"character\" as a \"grapheme cluster\", i.e a \"user character\" (to\nthe end user, \"ä\" and \"a\"+diaeresis is the same character, no matter if\nthey would display as different glyphs), whereas some people use\n\"character\" as a \"code point\", which would be more of a \"programmer\ncharacter\". And then there are some people that still use \"character\"\ninterchangibly for \"bytes\" or \"code units\" (for UTF-16; a pair of\nsurrogate code units is still only one \"code point\").\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"65886","messageId":"Pine.LNX.4.64.0801181631150.817@ds9.cixit.se","threadId":"11645","inReplyTo":"20080118103036.GD14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Peter Karlsson","fromEmail":"peter@softwolves.pp.se","sentAt":"2008-01-18T15:37:17Z","receivedAt":"2008-01-18T15:37:17Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Dmitry Potapov:\n\n> because Microsoft C library does not work with encoding that requires\n> more than two bytes per character.\n\nIndeed. On Windows, you should avoid using UTF-8 and instead use UTF-16\neverywhere. That usually works better, and if you run on an NT-based\nsystem it will convert all the data to WinAPI to UTF-16 anyway.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"65897","messageId":"alpine.LFD.1.00.0801180909000.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"Pine.LNX.4.64.0801181114430.817@ds9.cixit.se","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-18T17:11:47Z","receivedAt":"2008-01-18T17:11:47Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 18 Jan 2008, Peter Karlsson wrote:\n> \n> But they are not different strings, they are canonically equivalent as\n> far as Unicode is concerned.\n\nFuck me with a spoon.\n\nWhy the hell cannot people see that \"equivalent\" and \"same\" are two \ntotally different meanings.\n\n> You cannot do a binary comparison of text to see if two strings are\n> equivalent.\n\n.. and this is relevant how? They are different strings. Not the same.\n\nEquivalence doesn't matter. Equivalence is *evil*. Equivalence is what \ngives us case-insensitive filesystems (\"because the names are \nequivalent\").\n\nFilesystems don't *want* equivalence. They want a much stronger exactness \nguarantee. Exactly because sometimes the differences matter.\n\n\t\tLinus\n"},{"id":"65899","messageId":"fmqncq$5sf$1@ger.gmane.org","threadId":"11645","inReplyTo":"Pine.LNX.4.64.0801181631150.817@ds9.cixit.se","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-01-18T17:24:44Z","receivedAt":"2008-01-18T17:24:44Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Peter Karlsson wrote:\n\n> Dmitry Potapov:\n> \n>> because Microsoft C library does not work with encoding that requires\n>> more than two bytes per character.\n> \n> Indeed. On Windows, you should avoid using UTF-8 and instead use UTF-16\n> everywhere. That usually works better, and if you run on an NT-based\n> system it will convert all the data to WinAPI to UTF-16 anyway.\n\nErrr... doesn't UTF-16 (as compared to USC-2) sometimes (for some exotic\ncharacters) require more than two bytes per character?\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"65915","messageId":"2E6F57FC-3E78-4DD2-9B5B-CF75975D6A60@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801180909000.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-18T20:24:34Z","receivedAt":"2008-01-18T20:24:34Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"As far as I can tell, the only time you ever run into the problems  \nyou've described on a filesystem which treats filenames as unicode  \nstrings (and therefore is free to normalize), are when you're trying  \nto interact with a filesystem that treats filenames as sequences of  \nbytes.\n\nThis doesn't mean treating filenames as unicode strings is wrong, it  \njust means that the world would be much better if every filesystem had  \nthe same behaviour here. It's kinda like the endian issue, except  \nthere's no simple solution here.\n\n-Kevin Ballard\n\nOn Jan 18, 2008, at 12:11 PM, Linus Torvalds wrote:\n\n> On Fri, 18 Jan 2008, Peter Karlsson wrote:\n>>\n>> But they are not different strings, they are canonically equivalent  \n>> as\n>> far as Unicode is concerned.\n>\n> Fuck me with a spoon.\n>\n> Why the hell cannot people see that \"equivalent\" and \"same\" are two\n> totally different meanings.\n>\n>> You cannot do a binary comparison of text to see if two strings are\n>> equivalent.\n>\n> .. and this is relevant how? They are different strings. Not the same.\n>\n> Equivalence doesn't matter. Equivalence is *evil*. Equivalence is what\n> gives us case-insensitive filesystems (\"because the names are\n> equivalent\").\n>\n> Filesystems don't *want* equivalence. They want a much stronger  \n> exactness\n> guarantee. Exactly because sometimes the differences matter.\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65916","messageId":"7vr6gedgk9.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801180909000.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-18T20:28:22Z","receivedAt":"2008-01-18T20:28:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Fri, 18 Jan 2008, Peter Karlsson wrote:\n>> \n>> But they are not different strings, they are canonically equivalent as\n>> far as Unicode is concerned.\n>\n> Fuck me with a spoon.\n>\n> Why the hell cannot people see that \"equivalent\" and \"same\" are two \n> totally different meanings.\n\nCould people _please_ stop this already?\n\nI think the sane people see the difference between equivalence\nand sameness, and we established that a filesystem that mangles\nthe filenames behind user's back is a bad design.  Anybody who\nfollowed the thread and still does not agree with you is, eh,\n\"ugly-and-stupid\", as you might say ;-).  You cannot educate\nthem all.\n\nThe thing is, even if you mange to educate them all, that broken\nfilesystem, and other filesystems with similar brokenness, do\nnot go away.  If your ultimate objective is to declare that it\nis the right thing for git not to support such broken\nfilesystems, and to make everybody agree to it, that is fine.\nPlease keep pouring fuel to the fire.  But if that is not the\ncase, we would need to devise a way to help lives easier for the\nunfortunate people who are stuck on such filesystems.  They may\nnot even realize that they are unfortunate now, and I agree that\nsome education is justified, but this thread has raged on long\nenough to salvage any salvageable lost souls (the remaining ones\nmay be beyond salvation but let's not waste time on them).\n\nI'd rather see our mental bandwidth spent on coming up with a\nworkable workaround for such broken filesystems, while not\nhurting use of git on sane platforms.\n\nI fear it might have to end up to be very messy and slow,\nthough.\n"},{"id":"65919","messageId":"alpine.LSU.1.00.0801182042360.5731@racer.site","threadId":"11645","inReplyTo":"7vr6gedgk9.fsf@gitster.siamese.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-18T20:50:35Z","receivedAt":"2008-01-18T20:50:35Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 18 Jan 2008, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > On Fri, 18 Jan 2008, Peter Karlsson wrote:\n> >> \n> >> But they are not different strings, they are canonically equivalent as\n> >> far as Unicode is concerned.\n> >\n> > Fuck me with a spoon.\n> >\n> > Why the hell cannot people see that \"equivalent\" and \"same\" are two \n> > totally different meanings.\n> \n> Could people _please_ stop this already?\n\nWelcome, voice of reason.\n\n> I think the sane people see the difference between equivalence\n> and sameness, and we established that a filesystem that mangles\n> the filenames behind user's back is a bad design.  Anybody who\n> followed the thread and still does not agree with you is, eh,\n> \"ugly-and-stupid\", as you might say ;-).  You cannot educate\n> them all.\n\nActually, I see some value in calling them names, see \nhttp://video.google.nl/videoplay?docid=-4216011961522818645 for why.\n\n> The thing is, even if you mange to educate them all, that broken \n> filesystem, and other filesystems with similar brokenness, do not go \n> away.\n\nI was almost starting with hacking on this, but then the discussion \nannoyed me too much, and I asked myself for who I think I'd do this.\n\nIMHO those people should ask \"how could I begin to work on this\".\n\nInstead, they started a useless flamewar.\n\nNow, back to the issue: Robin posted a link to his UTF-8 work.  While it \nis way too intrusive, and not limited to filenames at all, I think it has \na few good pointers.\n\nCiao,\nDscho \"who needs to calm down now\"\n"},{"id":"65947","messageId":"20080119084814.GH14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"2E6F57FC-3E78-4DD2-9B5B-CF75975D6A60@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-19T08:48:14Z","receivedAt":"2008-01-19T08:48:14Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"Hi,\n\n[please do not top post. Just delete everything you do not reply to]\n\nOn Fri, Jan 18, 2008 at 03:24:34PM -0500, Kevin Ballard wrote:\n> As far as I can tell, the only time you ever run into the problems  \n> you've described on a filesystem which treats filenames as unicode  \n> strings (and therefore is free to normalize), are when you're trying  \n> to interact with a filesystem that treats filenames as sequences of  \n> bytes.\n\nIf you read the first message in this thread, you would probably know\nthat the problem exists on Mac even without any other filesystem being\ninvolved. My understanding of it is that is caused by HFS+ converting\none sequence of Unicode characters (generated by Mac keyboard driver)\nto another sequence using \"fast decomposed\" conversion.\n\n[And please stop calling by normalization what is not. Mac does NOT\nnormalize Unicode strings, it uses some sub-standard conversion,\nwhich neither produce a normalized string nor is guaranteed to be\nstable across versions of Unicode.]\n\n> This doesn't mean treating filenames as unicode strings is wrong, it  \n> just means that the world would be much better if every filesystem had  \n> the same behaviour here. It's kinda like the endian issue, except  \n> there's no simple solution here.\n\nActually, there is, if you care to do something. You can write a wrapper\naround readdir(3) that will recodes filenames in Unicode Normal Forms C.\nThis does not require much knowledge of Git -- what it requires the\ndesire to do something to solve the problem. Of course, this step alone\nis not a complete solution (it does not solve case-insensitive issue),\nbut the first step in the right direction...\n\nBTW, Git is far from being only software that ran into this problem with\nMac. But not being first, we can benefit from other people experiences:\nhttp://osdir.com/ml/network.gnutella.limewire.core.devel/2003-01/msg00000.html\n\n\nDmitry\n"},{"id":"65966","messageId":"FD3512A5-1CC6-4F02-8C56-4CAD6F50981B@sb.org","threadId":"11645","inReplyTo":"20080119084814.GH14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-19T14:55:45Z","receivedAt":"2008-01-19T14:55:45Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 19, 2008, at 3:48 AM, Dmitry Potapov wrote:\n\n> On Fri, Jan 18, 2008 at 03:24:34PM -0500, Kevin Ballard wrote:\n>> As far as I can tell, the only time you ever run into the problems\n>> you've described on a filesystem which treats filenames as unicode\n>> strings (and therefore is free to normalize), are when you're trying\n>> to interact with a filesystem that treats filenames as sequences of\n>> bytes.\n>\n> If you read the first message in this thread, you would probably know\n> that the problem exists on Mac even without any other filesystem being\n> involved. My understanding of it is that is caused by HFS+ converting\n> one sequence of Unicode characters (generated by Mac keyboard driver)\n> to another sequence using \"fast decomposed\" conversion.\n\nIn this case, git's index counts as a filesystem that treats filenames  \nas sequences of bytes. But yes, it is possible, though somewhat  \ndifficult, to produce this problem on just HFS+. It's far more common  \nwhen the file was originally added on a different filesystem\n\n> [And please stop calling by normalization what is not. Mac does NOT\n> normalize Unicode strings, it uses some sub-standard conversion,\n> which neither produce a normalized string nor is guaranteed to be\n> stable across versions of Unicode.]\n\n From what the HFS+ technote says, it produces a variant of Normal  \nForm D. This variant, while not guaranteed to be stable across  \nversions of HFS+, but in practice it is stable.\n\nWhat would you prefer I call it?\n\n>> This doesn't mean treating filenames as unicode strings is wrong, it\n>> just means that the world would be much better if every filesystem  \n>> had\n>> the same behaviour here. It's kinda like the endian issue, except\n>> there's no simple solution here.\n>\n> Actually, there is, if you care to do something. You can write a  \n> wrapper\n> around readdir(3) that will recodes filenames in Unicode Normal  \n> Forms C.\n> This does not require much knowledge of Git -- what it requires the\n> desire to do something to solve the problem. Of course, this step  \n> alone\n> is not a complete solution (it does not solve case-insensitive issue),\n> but the first step in the right direction...\n\nI'm not sure how that would solve anything. Sure, it would provide a  \nstable, known encoding for git to compare filenames against, but that  \nwould only work if the filename is known to be Unicode, and as it has  \nbeen pointed out on other filesystems the filename can be whatever  \nencoding the user chooses (which, IMHO, is a flaw).\n\n> BTW, Git is far from being only software that ran into this problem  \n> with\n> Mac. But not being first, we can benefit from other people  \n> experiences:\n> http://osdir.com/ml/network.gnutella.limewire.core.devel/2003-01/msg00000.html\n\nIt looks like their problem was binary compatibility with strings from  \nother clients that were using Normal Form C instead of Normal Form D.  \ngit's problem is that it's only even using a known encoding on HFS+.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65983","messageId":"alpine.LFD.1.00.0801191026500.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"20080119084814.GH14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-19T18:58:44Z","receivedAt":"2008-01-19T18:58:44Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 19 Jan 2008, Dmitry Potapov wrote:\n> \n> Actually, there is, if you care to do something. You can write a wrapper\n> around readdir(3) that will recodes filenames in Unicode Normal Forms C.\n\nIf somebody wants to do this, then readdir() isn't the only place, but \nyes, readdir() is one of the places.\n\nI suspect that if we were to just do the \"turn into NFC on readdir() on OS \nX\", that might actually be good enough to hide most of the problems. The \nissue isn't just that OS X mangles the filenames, it's that it picks a \nparticularly *stupid* way to mangle them (the decomposed forms), which \nmeans that OS X will actually not just corrupt \"odd cases\" of Unicode, but \nwill corrupt the obvious and *common* Latin1 translations of Unicode.\n\nI don't know if NFC is better for other locales, but I doubt it. Usually \npeople want to do the *composite* forms, not the *de*composed forms.\n\nA trivial example of this for some cross-OS issue:\n\n - let's say that you have a file \"Märchen\" on just about *any* other OS \n   than OS X. It could be Latin1 or it could be Unicode, but even if it is \n   Unicode, I can almost guarantee that the 'ä' is going to be the \n   *single* Unicode character U+00e4 (utf-8: \"\\xc3\\xa4\", latin1: \"\\xe4\")\n\n   So from a cross-OS standpoint, that's the *common* representation, and \n   yes, you can create the file that way (I don't know what happens if you \n   actually create it with the Latin1 encoding, but I would not be \n   surprised if OS X notices that it's not a valid UTF sequence and \n   assumes it's Latin1 and converts it to Unicode)\n\n - But on OS X, because of Apples *insane* choice of normal form, it will \n   then be turned into \"a¨\". I doubt *anybody* else does that. If you have \n   to normalize it, NFD is just about the *worst* choice.\n\nSo yeah, even just re-coding it as NFC on readdir() would at least mean \nthat any OS X git client would be MORE LIKELY to pick the same \nrepresentation as git clients on other OS's.\n\nIt wouldn't solve all problems (and it would almost certainly create a few \nnew ones), but it would likely at least increase compatibility between \nsystems.\n\nSo doing the NFC conversion on readdir() on OS X is probably a good idea, \nand probably is the simplest way to make it interact better with other \nOS's. And it's definitely safe on OS X, since OS X _already_ corrupted the \nname, so we're not losing any information (in contrast, on other systems, \ndoing a NFC conversion would possibly lose encoding detail _and_ might be \nincorrect simply because they might not use Unicode in the first place).\n\nAnybody want to creat a compat layer around \"readdir()\" that does that NFC \nconversion on OS X but not elsewhere?\n\n\t\tLinus\n"},{"id":"65984","messageId":"0D595B77-FAD6-4899-B475-8C1190C28E0F@mac.com","threadId":"11645","inReplyTo":"3D987338-EC1E-422F-850D-D8C52345A6A7@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kyle Moffett","fromEmail":"mrlinuxman@mac.com","sentAt":"2008-01-19T19:29:49Z","receivedAt":"2008-01-19T19:29:49Z","isPatch":false,"sender":{"key":"mrlinuxman@mac.com","avatar":null},"body":"On Jan 17, 2008, at 10:04, Kevin Ballard wrote:\n> The main problem with this approach is you know for certain that  \n> using HFSX as the boot partition is barely tested by Apple, and  \n> certainly untested by third-party apps. This means the potential for  \n> breakage is extremely high.\n\nNo, actually, HFSX boot partitions are fairly well tested by Apple and  \nmost 3rd-party programs.  I had one for a while and the only problems  \nI encountered were with programs ported from Windows without Mac  \nversions, such as \"Microsoft Office for Mac\" and \"World of Warcraft\".   \n\"Quake 4\" has a few quirks which are easily worked around.\n\nCheers,\nKyle Moffett\n"},{"id":"65987","messageId":"EC5ED13B-88CD-4D7C-A4DF-30185BAEE7A7@sb.org","threadId":"11645","inReplyTo":"0D595B77-FAD6-4899-B475-8C1190C28E0F@mac.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-19T19:57:57Z","receivedAt":"2008-01-19T19:57:57Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 19, 2008, at 2:29 PM, Kyle Moffett wrote:\n\n> On Jan 17, 2008, at 10:04, Kevin Ballard wrote:\n>> The main problem with this approach is you know for certain that  \n>> using HFSX as the boot partition is barely tested by Apple, and  \n>> certainly untested by third-party apps. This means the potential  \n>> for breakage is extremely high.\n>\n> No, actually, HFSX boot partitions are fairly well tested by Apple  \n> and most 3rd-party programs.  I had one for a while and the only  \n> problems I encountered were with programs ported from Windows  \n> without Mac versions, such as \"Microsoft Office for Mac\" and \"World  \n> of Warcraft\".  \"Quake 4\" has a few quirks which are easily worked  \n> around.\n\n\nPerhaps the big name companies might do some testing on HFSX, but I  \ncan guarantee most third-party programs will not be tested under HFSX.\n\nAlso, World of Warcraft isn't a ported program. It was developed for  \nthe Mac concurrently with the Windows version. Same with MS Office -  \nit's an entirely different team (the Mac BU) developing MS Office for  \nMac independently of the Windows version, not a porting job. However,  \nif you're saying these two big-name programs had problems, I wouldn't  \nbe surprised to see many more problems on various other third-party  \napps from smaller companies.\n\nIn any case, \"just use HFSX\" is still not an appropriate solution to  \nthe problem, especially since that will only take care of case  \nsensitivity and not the utf-8 stuff.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"65989","messageId":"47925FE6.8020704@web.de","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801191026500.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-19T20:39:02Z","receivedAt":"2008-01-19T20:39:02Z","isPatch":false,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Linus Torvalds schrieb:\n\n>  - let's say that you have a file \"Märchen\" on just about *any* other OS \n>    than OS X. It could be Latin1 or it could be Unicode, but even if it is \n>    Unicode, I can almost guarantee that the 'ä' is going to be the \n>    *single* Unicode character U+00e4 (utf-8: \"\\xc3\\xa4\", latin1: \"\\xe4\")\n> \n>    So from a cross-OS standpoint, that's the *common* representation, and \n>    yes, you can create the file that way (I don't know what happens if you \n>    actually create it with the Latin1 encoding, but I would not be \n>    surprised if OS X notices that it's not a valid UTF sequence and \n>    assumes it's Latin1 and converts it to Unicode)\n\nFWIW: I just made a test and it seems that MacOS X refuses the creation \nof a file with this invalid name.\n\n> Anybody want to creat a compat layer around \"readdir()\" that does that NFC \n> conversion on OS X but not elsewhere?\n\nMaybe I'll try it.\n\nRegards,\nMark\n"},{"id":"65990","messageId":"20080119211706.GK14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"FD3512A5-1CC6-4F02-8C56-4CAD6F50981B@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-19T21:17:06Z","receivedAt":"2008-01-19T21:17:06Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Sat, Jan 19, 2008 at 09:55:45AM -0500, Kevin Ballard wrote:\n> \n> >[And please stop calling by normalization what is not. Mac does NOT\n> >normalize Unicode strings, it uses some sub-standard conversion,\n> >which neither produce a normalized string nor is guaranteed to be\n> >stable across versions of Unicode.]\n> \n> From what the HFS+ technote says, it produces a variant of Normal  \n> Form D. \n\nThere is no such thing in the standard as a variant of NFD. Moreover,\neven if this conversion were described in the standard, it would never\ncalled as normalization, because normalization means conversion that\nmakes all equivalent strings having identical binary representations.\nHFS+ conversion does not met this criterion, so it not normalization.\n\n> This variant, while not guaranteed to be stable across  \n> versions of HFS+, but in practice it is stable.\n> \n> What would you prefer I call it?\n\nApple calls it as decomposition, which is correct even if it is not full\ndecomposition as stated in the technote.\n\n> \n> >>This doesn't mean treating filenames as unicode strings is wrong, it\n> >>just means that the world would be much better if every filesystem  \n> >>had\n> >>the same behaviour here. It's kinda like the endian issue, except\n> >>there's no simple solution here.\n> >\n> >Actually, there is, if you care to do something. You can write a  \n> >wrapper\n> >around readdir(3) that will recodes filenames in Unicode Normal  \n> >Forms C.\n> >This does not require much knowledge of Git -- what it requires the\n> >desire to do something to solve the problem. Of course, this step  \n> >alone\n> >is not a complete solution (it does not solve case-insensitive issue),\n> >but the first step in the right direction...\n> \n> I'm not sure how that would solve anything. Sure, it would provide a  \n> stable, known encoding for git to compare filenames against, but that  \n> would only work if the filename is known to be Unicode, and as it has  \n> been pointed out on other filesystems the filename can be whatever  \n> encoding the user chooses (which, IMHO, is a flaw).\n\nI believe that Git internally should use only UTF-8 for encoding file\nnames, commit messages, etc. The problem with some other filesystems\nshould be addressed separately (by those who work on those systems or\nat least have access to them). Regardless interoperability with other\nsystems, this change alone should solve the issue that was described\nin the first message of this thread.\n\nDmitry\n"},{"id":"65996","messageId":"alpine.LSU.1.00.0801192256480.5731@racer.site","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801191026500.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-19T22:58:08Z","receivedAt":"2008-01-19T22:58:08Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sat, 19 Jan 2008, Linus Torvalds wrote:\n\n> On Sat, 19 Jan 2008, Dmitry Potapov wrote:\n> > \n> > Actually, there is, if you care to do something. You can write a \n> > wrapper around readdir(3) that will recodes filenames in Unicode \n> > Normal Forms C.\n> \n> If somebody wants to do this, then readdir() isn't the only place, but \n> yes, readdir() is one of the places.\n> \n> I suspect that if we were to just do the \"turn into NFC on readdir() on \n> OS X\", that might actually be good enough to hide most of the problems.\n\nI think a better approach would be to try to match the name to what we \nhave in the index.  Then we could implement case-insensitivity and MacOSX \nworkaround at the same time.\n\nCiao,\nDscho\n"},{"id":"66003","messageId":"B4FDA32F-16C9-497A-AAD8-27A8D510C4CB@wincent.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801191026500.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-20T00:11:57Z","receivedAt":"2008-01-20T00:11:57Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 19/1/2008, a las 19:58, Linus Torvalds escribió:\n\n> I suspect that if we were to just do the \"turn into NFC on readdir()  \n> on OS\n> X\", that might actually be good enough to hide most of the problems.  \n> The\n> issue isn't just that OS X mangles the filenames, it's that it picks a\n> particularly *stupid* way to mangle them (the decomposed forms), which\n> means that OS X will actually not just corrupt \"odd cases\" of  \n> Unicode, but\n> will corrupt the obvious and *common* Latin1 translations of Unicode.\n\n\nFor what it's worth, their choice wasn't entirely \"insane\" ie. it did  \nhave an element of rationality: that decomposed forms are a little bit  \nsimpler to sort.\n\nOf course, this doesn't excuse them for creating a file system that  \ninteracts so horridly with basically everything else out there.\n\nCheers,\nWincent\n"},{"id":"66004","messageId":"alpine.LFD.1.00.0801191659350.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"B4FDA32F-16C9-497A-AAD8-27A8D510C4CB@wincent.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-20T01:04:09Z","receivedAt":"2008-01-20T01:04:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Jan 2008, Wincent Colaiuta wrote:\n> \n> For what it's worth, their choice wasn't entirely \"insane\" ie. it did have an\n> element of rationality: that decomposed forms are a little bit simpler to\n> sort.\n\nNo they are *not*.\n\nIn many languages, 'ä' does *not* sort like 'a' at all, and if you think \nit does, you'll sort at least Finnish and Swedish totally wrong (åäö are \nreal letters, and they sort at the *end* of the alphabet, they have \nnothing what-so-ever to do with the letters 'a' or 'o').\n\nThe fact that in *some* languages the decomposed forms sort as the base \nletter is immaterial. It's only true in some cases.\n\nSo no, sort order is not it. To sort right, you need to use the a real \nUnicode sort (and the decomposed form is *not* going to help you one bit, \nquite the reverse).\n\nIt may be that a case compare is easier in NFD (ie you basically only do \nthe case-compare on the base letter).\n\n\t\tLinus\n"},{"id":"66013","messageId":"20080120052735.GA18581@glandium.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801191659350.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-01-20T05:27:35Z","receivedAt":"2008-01-20T05:27:35Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Sat, Jan 19, 2008 at 05:04:09PM -0800, Linus Torvalds wrote:\n> \n> \n> On Sun, 20 Jan 2008, Wincent Colaiuta wrote:\n> > \n> > For what it's worth, their choice wasn't entirely \"insane\" ie. it did have an\n> > element of rationality: that decomposed forms are a little bit simpler to\n> > sort.\n> \n> No they are *not*.\n> \n> In many languages, 'ä' does *not* sort like 'a' at all, and if you think \n> it does, you'll sort at least Finnish and Swedish totally wrong (åäö are \n> real letters, and they sort at the *end* of the alphabet, they have \n> nothing what-so-ever to do with the letters 'a' or 'o').\n\nBut there is no way to know whether 'ä' in a document is the Finnish 'ä'\nor a 'ä' from, say, German, that sorts after 'a'.\n\n> The fact that in *some* languages the decomposed forms sort as the base \n> letter is immaterial. It's only true in some cases.\n> \n> So no, sort order is not it. To sort right, you need to use the a real \n> Unicode sort (and the decomposed form is *not* going to help you one bit, \n> quite the reverse).\n\nUnicode sort is not enough, there is no language indicator in an Unicode\ndocument, which is why Unicode, while solving a bunch of problems, has\nits very own, cf. the infamous CJK problem.\n\nBut that's all very OT.\n\nMike\n"},{"id":"66014","messageId":"alpine.LFD.1.00.0801192130180.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"20080120052735.GA18581@glandium.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-20T05:45:40Z","receivedAt":"2008-01-20T05:45:40Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Jan 2008, Mike Hommey wrote:\n> \n> But there is no way to know whether 'ä' in a document is the Finnish 'ä'\n> or a 'ä' from, say, German, that sorts after 'a'.\n\n... without knowing the locale. Correct.\n\nThat's why sorting is locale-dependent, even in Unicode. And why you \nshould always sort using the *combined* character, not think that you can \nsort by decompsed sequence.\n\nThat said, even then you get the wrong thing. Some things cannot be sorted \ncharacter by character at all, and have semantical sorting at a higher \nlevel entirely. I think most European family names are traditionally \nsorted by effectively using the prefixes (ie d', von, etc) as a secondary \nsort key (so even though they are in front, they sort as if they were \nat the _end_ of the name).\n\nSo unicode doesn't help with sorting, and you shouldn't even try to find \nsort rules in the Unicode spec or tech reports. But in general, \ndecomposing the characters just makes things worse, not better. To sort \nwell, you tend to need the bigger picture, not the details.\n\nOf course, for something like git, we sort by binary value, because we \nalso require the sort to be not just well-defined, but *stable*. A sort \nbased on any kind of unicode rule is rather likely to change over time.\n\n\t\t\tLinus\n"},{"id":"66015","messageId":"20080120061419.GL14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801192256480.5731@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-20T06:14:19Z","receivedAt":"2008-01-20T06:14:19Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Sat, Jan 19, 2008 at 10:58:08PM +0000, Johannes Schindelin wrote:\n> \n> I think a better approach would be to try to match the name to what we \n> have in the index.  Then we could implement case-insensitivity and MacOSX \n> workaround at the same time.\n\nI thought about that, but the problem is that HFS+ _already_ mangled\nnames from what the user entered (and what is used by anyone else)\nto some sub-standard form, which no one outside of Mac likes or uses.\nThus, bringing filenames back to the NFC form (which is what almost\nanyone uses) is the only sane thing do, because no one outside of Mac\nreally needs to know about this HFS+ specific craziness.\n\nSo I really dislike the idea that due to some HFS+ specific conversion,\nwe may end up having some strangely encoded names in a Git repository.\nSane people enter names only in NFC, so why should they suffer because\nof some insane conversation made by filesystem behind everyone's back?\nAnd I am not entertaining the idea of having this Mac OS/X specific\nworkaround outside of Mac OS/X.\n\nBesides, writing a wrapper around readdir() is not difficult. We\nalready have git-compat-util.h, which redefines some functions for\nsome platforms, so I don't see any problem with writing a wrapper\naround readdir().\n\nDmitry\n"},{"id":"66016","messageId":"alpine.LFD.1.00.0801192250260.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"20080120061419.GL14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-20T06:53:19Z","receivedAt":"2008-01-20T06:53:19Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Jan 2008, Dmitry Potapov wrote:\n>\n> On Sat, Jan 19, 2008 at 10:58:08PM +0000, Johannes Schindelin wrote:\n> > \n> > I think a better approach would be to try to match the name to what we \n> > have in the index.  Then we could implement case-insensitivity and MacOSX \n> > workaround at the same time.\n> \n> I thought about that, but the problem is that HFS+ _already_ mangled\n> names from what the user entered (and what is used by anyone else)\n> to some sub-standard form, which no one outside of Mac likes or uses.\n\nWell, more importantly, most of the important cases actually don't have an \nindex entry yet.\n\nFor example, what about \"git add\"? That's when it really matters that you \nadd things in a sane format, and by definition, you don't have an index \nentry to try to match to. \n\nSo once you aim for NFC in \"git add\", now the index will generally be in \nNFC anyway (since I agree that that's what you'd normally get on non-OSX \nsystems), so there is little point in then matching the index.\n\nBut no, it won't fix all problems. I do suspect it would make them less \nobvious in practice, though.\n\n\t\tLinus\n"},{"id":"66018","messageId":"20080120070018.GA11015@glandium.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801192130180.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-01-20T07:00:18Z","receivedAt":"2008-01-20T07:00:18Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Sat, Jan 19, 2008 at 09:45:40PM -0800, Linus Torvalds wrote:\n> \n> \n> On Sun, 20 Jan 2008, Mike Hommey wrote:\n> > \n> > But there is no way to know whether 'ä' in a document is the Finnish 'ä'\n> > or a 'ä' from, say, German, that sorts after 'a'.\n> \n> ... without knowing the locale. Correct.\n> \n> That's why sorting is locale-dependent, even in Unicode. And why you \n> should always sort using the *combined* character, not think that you can \n> sort by decompsed sequence.\n\nThat said, the locale doesn't necessarily express the language in which\nthe document is written. It's easy enough to read documents that are not\nwritten in your native language on the net. That's already what we are both\ndoing right now. Fortunately, HTTP and HTML have ways to indicate the\nlanguage in which a document is written in, but that leaves out plain\nmail, for instance. \n\nThat said, the \"decomposed\" version of UTF-8 has nice side effects on\nOSX, with UTF-8 encoded RockRidge ISO-9660 volumes (with or without\nJoliet ; OSX will use RockRidge by default when it's there), for instance.\n\nMike\n"},{"id":"66020","messageId":"alpine.LFD.1.00.0801192311190.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"20080120070018.GA11015@glandium.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-20T07:26:23Z","receivedAt":"2008-01-20T07:26:23Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Jan 2008, Mike Hommey wrote:\n> \n> That said, the locale doesn't necessarily express the language in which\n> the document is written.\n\n.. and quite commonly, there are multiple languages per document.\n\nThe good news is that sorting is almost never relevant or done over \ngeneral documents. You sort almost only well-behaved data, and quite often \nthe exact order is less than important: and when it is, you have very \nspecific rules (which probably seldom have anything what-so-ever to do \nwith general unicode ;).\n\n> It's easy enough to read documents that are not\n> written in your native language on the net. That's already what we are both\n> doing right now. Fortunately, HTTP and HTML have ways to indicate the\n> language in which a document is written in, but that leaves out plain\n> mail, for instance. \n\nWell, Unicode already handles the \"reading\" part, just not the sorting.\n\n> That said, the \"decomposed\" version of UTF-8 has nice side effects on\n> OSX, with UTF-8 encoded RockRidge ISO-9660 volumes (with or without\n> Joliet ; OSX will use RockRidge by default when it's there), for instance.\n\nI think Unicode in general (and UTF-8 in particular) is a great thing. I \ndo not argue against Unicode at all.  It's what I use myself.\n\nThe thing I argue against is that they force normalization (and then, as a \nsecondary complaint, their insane choice of target format).\n\nLinux is generally UTF-8 too, and does all of this much better. No forced \nnormalization, and it uses UTF-8 everywhere as the encoding model. Joliet \nand RR works beautifully.\n\n(I don't think RR is NFD, btw. It's the standard microsoft UTF-16 without \nnormalization, afaik. I think you can happily generate a Rock Ridge disk \nthat has two _different_ filenames that OS X cannot tell apart, but that \nboth Linux and Windows can see peoperly)\n\n\t\tLinus\n"},{"id":"66021","messageId":"20080120080056.GO14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"20080120070018.GA11015@glandium.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-20T08:00:56Z","receivedAt":"2008-01-20T08:00:56Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Sun, Jan 20, 2008 at 08:00:18AM +0100, Mike Hommey wrote:\n> \n> That said, the \"decomposed\" version of UTF-8 has nice side effects on\n> OSX, with UTF-8 encoded RockRidge ISO-9660 volumes (with or without\n> Joliet ; OSX will use RockRidge by default when it's there), for instance.\n\nAFAIK, the RockRidge standard prescribes to use the portable character\nset, and it has nothing to do with Unicode. Basically, it is a subset of\nASCII.\n\nhttp://www.opengroup.org/onlinepubs/009695399/basedefs/xbd_chap06.html\n\nSo, I don't think UTF-8 encoded filenames are valid regardless whether\nthey are decomposed or not.\n\nDmitry\n"},{"id":"66022","messageId":"20080120081236.GP14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"20080120080056.GO14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-20T08:12:36Z","receivedAt":"2008-01-20T08:12:36Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Sun, Jan 20, 2008 at 11:00:56AM +0300, Dmitry Potapov wrote:\n> On Sun, Jan 20, 2008 at 08:00:18AM +0100, Mike Hommey wrote:\n> > \n> > That said, the \"decomposed\" version of UTF-8 has nice side effects on\n> > OSX, with UTF-8 encoded RockRidge ISO-9660 volumes (with or without\n> > Joliet ; OSX will use RockRidge by default when it's there), for instance.\n> \n> AFAIK, the RockRidge standard prescribes to use the portable character\n> set, \n\nActually, it prescribes to use the portable *filename* character set,\nwhich is even more restrictive than just portable character set.\n\nhttp://www.opengroup.org/onlinepubs/009695399/basedefs/xbd_chap03.html#tag_03_276\n\nAnyway, there is no place for UTF-8 in it.\n\nDmitry\n"},{"id":"66025","messageId":"B2D68DE5-4A97-4E89-8C1F-A889380FE32A@wincent.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801191659350.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-20T09:34:54Z","receivedAt":"2008-01-20T09:34:54Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 20/1/2008, a las 2:04, Linus Torvalds escribió:\n\n> On Sun, 20 Jan 2008, Wincent Colaiuta wrote:\n>>\n>> For what it's worth, their choice wasn't entirely \"insane\" ie. it  \n>> did have an\n>> element of rationality: that decomposed forms are a little bit  \n>> simpler to\n>> sort.\n>\n> No they are *not*.\n>\n> In many languages, 'ä' does *not* sort like 'a' at all, and if you  \n> think\n> it does, you'll sort at least Finnish and Swedish totally wrong (åäö  \n> are\n> real letters, and they sort at the *end* of the alphabet, they have\n> nothing what-so-ever to do with the letters 'a' or 'o').\n>\n> The fact that in *some* languages the decomposed forms sort as the  \n> base\n> letter is immaterial. It's only true in some cases.\n>\n> So no, sort order is not it. To sort right, you need to use the a real\n> Unicode sort (and the decomposed form is *not* going to help you one  \n> bit,\n> quite the reverse).\n\nThat's what I get for believing Wikipedia (\"This makes sorting far  \nsimpler\"):\n\nhttp://en.wikipedia.org/wiki/UTF-8#Mac_OS_X\n\nCheers,\nWincent\n"},{"id":"66032","messageId":"alpine.LSU.1.00.0801201300190.5731@racer.site","threadId":"11645","inReplyTo":"20080120061419.GL14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-20T13:15:48Z","receivedAt":"2008-01-20T13:15:48Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 20 Jan 2008, Dmitry Potapov wrote:\n\n> On Sat, Jan 19, 2008 at 10:58:08PM +0000, Johannes Schindelin wrote:\n> > \n> > I think a better approach would be to try to match the name to what we \n> > have in the index.  Then we could implement case-insensitivity and \n> > MacOSX workaround at the same time.\n> \n> I thought about that, but the problem is that HFS+ _already_ mangled \n> names from what the user entered (and what is used by anyone else) to \n> some sub-standard form, which no one outside of Mac likes or uses.\n\nSo?  That's why I said \"match\", not \"compare for identity\".\n\nTo be a little bit more precise: I think a viable plan would be to\n\n- have a config switch which determines what type of filename mangling we\n  allow the host OS to perform (Unicode \"normalisation\", case mongering),\n  and leave _everybody_ alone who left that switch unset,\n\n- \"overload\" readdir() (by the famous git_X(); #define X git_X trick),\n\n- have the overloaded readdir() _know_ which is the current prefix, and\n  load the index if it has not yet been loaded (but probably into a static\n  variable to avoid reloading, and to avoid interfering with the global\n  \"cache\" instance).\n\nIt _could_ be wise to store the \"normalised\" forms at one stage (instead \nof the index) to speed up comparison -- the prefix has to be normalised \nfor readdir()s purposes, too, then.\n\nThis is possible with the HFS+ problem, since we know exactly how HFS+ \ntries to \"help\", and for case insensitivity too, I think.  But it may be \nrestricting ourselves for other filename \"equivalences\" we might want to \nhandle one day.\n\nBTW: I cannot think of anything else than readdir() which should have the \n\"problem\" of reading back a name that the user did not specify.  What am I \nmissing?\n\n> Thus, bringing filenames back to the NFC form (which is what almost \n> anyone uses) is the only sane thing do, because no one outside of Mac \n> really needs to know about this HFS+ specific craziness.\n\nNo.  I think that would be a serious mistake.  If you add a file on MacOSX \n(with a _mangled_ filename, think of \"git add .\"), git should not try to \nbe as clever as HFS+ and \"remangle\" it.\n\n> So I really dislike the idea that due to some HFS+ specific conversion, \n> we may end up having some strangely encoded names in a Git repository.\n\nIt _is_ UTF-8, so what's the problem?\n\nAs for the HFS+ specfic conversion: like the CRLF issue, I am opposed to \nhave a \"solution\" affecting other people than those on broken system.  So \nI very much _want_ it to be an HFS+ specific conversion.\n\n> Besides, writing a wrapper around readdir() is not difficult. We already \n> have git-compat-util.h, which redefines some functions for some \n> platforms, so I don't see any problem with writing a wrapper around \n> readdir().\n\nExactly.\n\nCiao,\nDscho\n"},{"id":"66141","messageId":"Pine.LNX.4.64.0801211509490.17095@ds9.cixit.se","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801180909000.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Peter Karlsson","fromEmail":"peter@softwolves.pp.se","sentAt":"2008-01-21T14:14:16Z","receivedAt":"2008-01-21T14:14:16Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Linus Torvalds:\n\n> Fuck me with a spoon.\n\nI'd prefer not to.\n\n> .. and this is relevant how? They are different strings. Not the same.\n\nIt is relevant because the Mac OS file system stores file names as a\nsequence of Unicode code points, in a (apparently slightly modified)\nnormalized form, whereas Git prefers to see file systems that store\nfile names as a sequence of octets, which may, or may not, actually map\nto something that the user would call characters.\n\nI happen to prefer the text-as-string-of-characters (or code points,\nsince you use the other meaning of characters in your posts), since I\ncome from the text world, having worked a lot on Unicode text\nprocessing.\n\nYou apparently prefer the text-as-sequence-of-octets, which I tend to\ndislike because I would have thought computer engineers would have\nevolved beyond this when we left the 1900s.\n\nBut the real issue is that Git cannot use it's filenames as string of\noctets on Mac OS X, since the file system doesn't handle it. So Git\nneeds to do something sensible. That's part of porting. Preferrably\nthat would involve supporting real Unicode file names, which would also\nwork on Windows (through it's UTF-16 file APIs), and in part on other\nsystems (through conversion to the systems' locale encoding).\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"66150","messageId":"440E4426-BFB5-4836-93DF-05C99EF204E6@sb.org","threadId":"11645","inReplyTo":"Pine.LNX.4.64.0801211509490.17095@ds9.cixit.se","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T16:43:54Z","receivedAt":"2008-01-21T16:43:54Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 9:14 AM, Peter Karlsson wrote:\n\n> I happen to prefer the text-as-string-of-characters (or code points,\n> since you use the other meaning of characters in your posts), since I\n> come from the text world, having worked a lot on Unicode text\n> processing.\n>\n> You apparently prefer the text-as-sequence-of-octets, which I tend to\n> dislike because I would have thought computer engineers would have\n> evolved beyond this when we left the 1900s.\n\nI agree. Every single problem that I can recall Linus bringing up as a  \nconsequence of HFS+ treating filenames as strings is in fact only a  \nproblem if you then think of the filename as octets at some point. If  \nyou stick with UTF-8 equivalence comparison the entire time, then  \neverything just works.\n\nGranted, this is a problem when you have to operate on a filesystem  \nthat thinks of filenames as octets, but as I said before, this doesn't  \nmean the HFS+ approach is wrong, it just means it's incompatible with  \nLinus's approach.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66151","messageId":"86r6gbi0pr.fsf@lola.quinscape.zz","threadId":"11645","inReplyTo":"440E4426-BFB5-4836-93DF-05C99EF204E6@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-21T16:48:32Z","receivedAt":"2008-01-21T16:48:32Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Kevin Ballard <kevin@sb.org> writes:\n\n> On Jan 21, 2008, at 9:14 AM, Peter Karlsson wrote:\n>\n>> I happen to prefer the text-as-string-of-characters (or code points,\n>> since you use the other meaning of characters in your posts), since I\n>> come from the text world, having worked a lot on Unicode text\n>> processing.\n>>\n>> You apparently prefer the text-as-sequence-of-octets, which I tend to\n>> dislike because I would have thought computer engineers would have\n>> evolved beyond this when we left the 1900s.\n>\n> I agree. Every single problem that I can recall Linus bringing up as a\n> consequence of HFS+ treating filenames as strings is in fact only a\n> problem if you then think of the filename as octets at some point. If\n> you stick with UTF-8 equivalence comparison the entire time, then\n> everything just works.\n\ngit calculates hashes over filenames and sorts them.  This is not a mere\nquestion of \"UTF-8 equivalence comparison\".\n\n> Granted, this is a problem when you have to operate on a filesystem\n> that thinks of filenames as octets,\n\nIt also is a problem when operating on a filesystem that considers \"ä\" a\nsingle utf-8 character instead of decomposing it.\n\n> but as I said before, this doesn't mean the HFS+ approach is wrong, it\n> just means it's incompatible with Linus's approach.\n\nIt is not the business of a file system to juggle with filename\nrepresentations.\n\n-- \nDavid Kastrup\n"},{"id":"66152","messageId":"20080121165343.GA12308@sigill.intra.peff.net","threadId":"11645","inReplyTo":"440E4426-BFB5-4836-93DF-05C99EF204E6@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-01-21T16:53:44Z","receivedAt":"2008-01-21T16:53:44Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jan 21, 2008 at 11:43:54AM -0500, Kevin Ballard wrote:\n\n> I agree. Every single problem that I can recall Linus bringing up as a  \n> consequence of HFS+ treating filenames as strings is in fact only a  \n> problem if you then think of the filename as octets at some point. If you \n> stick with UTF-8 equivalence comparison the entire time, then everything \n> just works.\n\nGit's data model relies on SHA-1 hashing of data, including filenames.\nSo at some level, git _has_ to treat data as octets, and \"equivalent\"\nstrings must be the same at the octet level (or else you lose all of the\nuseful properties that the hashing data model provides). You can argue\nabout where in the program conversion and normalization occur, but I\ndon't think you can get around the fact that you're going to need\nto think of the \"filename as octets at some point.\"\n\n-Peff\n"},{"id":"66153","messageId":"CFF9E74C-4A4C-4E5F-8DA3-662D80095503@sb.org","threadId":"11645","inReplyTo":"86r6gbi0pr.fsf@lola.quinscape.zz","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T16:59:24Z","receivedAt":"2008-01-21T16:59:24Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 11:48 AM, David Kastrup wrote:\n\n> Kevin Ballard <kevin@sb.org> writes:\n>\n>> On Jan 21, 2008, at 9:14 AM, Peter Karlsson wrote:\n>>\n>>> I happen to prefer the text-as-string-of-characters (or code points,\n>>> since you use the other meaning of characters in your posts),  \n>>> since I\n>>> come from the text world, having worked a lot on Unicode text\n>>> processing.\n>>>\n>>> You apparently prefer the text-as-sequence-of-octets, which I tend  \n>>> to\n>>> dislike because I would have thought computer engineers would have\n>>> evolved beyond this when we left the 1900s.\n>>\n>> I agree. Every single problem that I can recall Linus bringing up  \n>> as a\n>> consequence of HFS+ treating filenames as strings is in fact only a\n>> problem if you then think of the filename as octets at some point. If\n>> you stick with UTF-8 equivalence comparison the entire time, then\n>> everything just works.\n>\n> git calculates hashes over filenames and sorts them.  This is not a  \n> mere\n> question of \"UTF-8 equivalence comparison\".\n\nNo, it's a question of hashing algorithm. And it's one that's fairly  \neasily solved simply by picking a specific nonambiguous UTF-8 encoding  \nbefore hashing.\n\n>> Granted, this is a problem when you have to operate on a filesystem\n>> that thinks of filenames as octets,\n>\n> It also is a problem when operating on a filesystem that considers  \n> \"ä\" a\n> single utf-8 character instead of decomposing it.\n\nWhat makes you say that?\n\n>> but as I said before, this doesn't mean the HFS+ approach is wrong,  \n>> it\n>> just means it's incompatible with Linus's approach.\n>\n> It is not the business of a file system to juggle with filename\n> representations.\n\nYou're right, that probably belongs in the VFS layer, but the behavior  \nis the same either way. You can't leave it up to user-space libraries  \nto enforce a filesystem encoding, because you can't rely on all  \nclients to behave properly.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66154","messageId":"alpine.LFD.1.00.0801211151180.20753@xanadu.home","threadId":"11645","inReplyTo":"440E4426-BFB5-4836-93DF-05C99EF204E6@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-01-21T17:08:29Z","receivedAt":"2008-01-21T17:08:29Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 21 Jan 2008, Kevin Ballard wrote:\n\n> On Jan 21, 2008, at 9:14 AM, Peter Karlsson wrote:\n> \n> > I happen to prefer the text-as-string-of-characters (or code points,\n> > since you use the other meaning of characters in your posts), since I\n> > come from the text world, having worked a lot on Unicode text\n> > processing.\n> > \n> > You apparently prefer the text-as-sequence-of-octets, which I tend to\n> > dislike because I would have thought computer engineers would have\n> > evolved beyond this when we left the 1900s.\n> \n> I agree. Every single problem that I can recall Linus bringing up as a\n> consequence of HFS+ treating filenames as strings is in fact only a problem if\n> you then think of the filename as octets at some point. If you stick with\n> UTF-8 equivalence comparison the entire time, then everything just works.\n> \n> Granted, this is a problem when you have to operate on a filesystem that\n> thinks of filenames as octets, but as I said before, this doesn't mean the\n> HFS+ approach is wrong, it just means it's incompatible with Linus's approach.\n\nLinus' approach is _FAST_.\n\nWhy do you think Git has now acquired a reputation of kicking asses all \naround the SCM scene?\n\nThe HFS+ approach might be fine if you think of it in terms of \"the user \nwill be awfully confused if two file names are shown identically in the \nFile Open dialog box\".  But it otherwise sucks big time when it comes to \nhigh performance applications needing to deal with a huge amount of file \nnames at once.\n\nNormalization will always hurt performances.  This is an overhead.  \nSometimes that overhead might be insignificant and not be perceptible, \nbut sometimes it is.  And Git is clearly in the later case. Performances \nwill be hurt big time the day it is made aware of that normalization. \nThis is why there is so much resistance about it, especially when the \nbenefits of normalizing file names are not shown to be worth their cost \nin performance and complexity, as other systems do rather fine without \nit.\n\n\nNicolas\n"},{"id":"66155","messageId":"AE99FDAA-F8D3-49F7-A0B9-CDFCC4903824@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211151180.20753@xanadu.home","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T17:25:52Z","receivedAt":"2008-01-21T17:25:52Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 12:08 PM, Nicolas Pitre wrote:\n\n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n>\n>> On Jan 21, 2008, at 9:14 AM, Peter Karlsson wrote:\n>>\n>>> I happen to prefer the text-as-string-of-characters (or code points,\n>>> since you use the other meaning of characters in your posts),  \n>>> since I\n>>> come from the text world, having worked a lot on Unicode text\n>>> processing.\n>>>\n>>> You apparently prefer the text-as-sequence-of-octets, which I tend  \n>>> to\n>>> dislike because I would have thought computer engineers would have\n>>> evolved beyond this when we left the 1900s.\n>>\n>> I agree. Every single problem that I can recall Linus bringing up  \n>> as a\n>> consequence of HFS+ treating filenames as strings is in fact only a  \n>> problem if\n>> you then think of the filename as octets at some point. If you  \n>> stick with\n>> UTF-8 equivalence comparison the entire time, then everything just  \n>> works.\n>>\n>> Granted, this is a problem when you have to operate on a filesystem  \n>> that\n>> thinks of filenames as octets, but as I said before, this doesn't  \n>> mean the\n>> HFS+ approach is wrong, it just means it's incompatible with  \n>> Linus's approach.\n>\n> Linus' approach is _FAST_.\n>\n> Why do you think Git has now acquired a reputation of kicking asses  \n> all\n> around the SCM scene?\n>\n> The HFS+ approach might be fine if you think of it in terms of \"the  \n> user\n> will be awfully confused if two file names are shown identically in  \n> the\n> File Open dialog box\".  But it otherwise sucks big time when it  \n> comes to\n> high performance applications needing to deal with a huge amount of  \n> file\n> names at once.\n>\n> Normalization will always hurt performances.  This is an overhead.\n> Sometimes that overhead might be insignificant and not be perceptible,\n> but sometimes it is.  And Git is clearly in the later case.  \n> Performances\n> will be hurt big time the day it is made aware of that normalization.\n> This is why there is so much resistance about it, especially when the\n> benefits of normalizing file names are not shown to be worth their  \n> cost\n> in performance and complexity, as other systems do rather fine without\n> it.\n\nI agree, Linus's approach is indeed fast. And if speed is more  \nimportant than treating filenames as text instead of octets, then so  \nbe it. This is a trade-off. But a trade-off doesn't mean one approach  \nis \"wrong\", it just means the authors of HFS+ thought it was an  \nacceptable trade-off. HFS+ wasn't designed to be a high-performance  \nfilesystem that deals with lots of files, it was designed to be a  \nfilesystem used by regular people on the Mac, and I believe treating  \nfilenames as text is a good choice in this scenario. Unfortunately,  \nthis does mean git has to do extra work to behave correctly on this  \nsystem.\n\nNow, to move on to actually coming up with a solution. Unfortunately I  \ndon't know enough about the internals of git to really evaluate the  \nproposed ideas myself, or to write a patch. Hopefully I'll come up  \nwith the time to acquire the necessary knowledge, but until then I can  \nonly participate in these higher-level discussions.\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66159","messageId":"alpine.LFD.1.00.0801210934400.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"440E4426-BFB5-4836-93DF-05C99EF204E6@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-21T18:12:01Z","receivedAt":"2008-01-21T18:12:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n> On Jan 21, 2008, at 9:14 AM, Peter Karlsson wrote:\n> > \n> > I happen to prefer the text-as-string-of-characters (or code points,\n> > since you use the other meaning of characters in your posts), since I\n> > come from the text world, having worked a lot on Unicode text\n> > processing.\n> > \n> > You apparently prefer the text-as-sequence-of-octets, which I tend to\n> > dislike because I would have thought computer engineers would have\n> > evolved beyond this when we left the 1900s.\n> \n> I agree. Every single problem that I can recall Linus bringing up as a\n> consequence of HFS+ treating filenames as strings [..]\n\nYou say \"I agree\", BUT YOU DON'T EVEN SEEM TO UNDERSTAND WHAT IS GOING ON.\n\nThe fact is, text-as-string-of-codepoints (let's make the \"codepoints\" \nobvious, so that there is no ambiguity, but I'd also like to make it clear \nthat a codepoint *is* how a Unicode character is defined, and a Unicode \n\"string\" is actually *defined* to be a sequence of codepoints, and totally \nindependent of normalization!) is fine.\n\nThat was never the issue at all. Unicode codepoints are wonderful.\n\nNow, git _also_ heavily depends on the actual encoding of those \ncodepoints, since we create hashes etc, so in fact, as far ass git is \nconcerned, names have to be in some particular encoding to be hashed, and \nUTF-8 is the only sane encoding for Unicode. People can blather about \nUCS-2 and UTF-16 and UTF-32 all they want, but the fact is, UTF-8 is \nsimply technically superior in so many ways that I don't even understand \nwhy anybody ever uses anything else.\n\nSo I would not disagree with using UTF-8 at all.\n\nBut that is *entirely* a separate issue from \"normalization\". \n\nKevin, you seem to think that normalization is somehow forced on you by \nthe \"text-as-codepoints\" decision, and that is SIMPLY NOT TRUE. \nNormalization is a totally separate decision, and it's a STUPID one, \nbecause it breaks so many of the _nice_ properties of using UTF-8.\n\nAnd THAT is where we differ. It has nothing to do with \"octets\". It has \nnothing to do with not liking Unicode. It has nothing to do with \n\"strings\". \n\nIn short:\n\n - normalization is by no means required or even a good feature. It's \n   something you do when you want to know if two strings are equivalent, \n   but that doesn't actually mean that you should keep the strings \n   normalized all the time!\n\n - normalization has *nothing* to do with \"treating text as octets\". \n   That's entirely an encoding issue.\n\n - of *course* git has to treat things as a binary stream at some point, \n   since you need that to even compute a SHA1 in the first place, but that \n   has *nothing* to do with normalization or the lack of it.\n\nGot it? Forced normalization is stupid, because it changes the data and \nremoves information, and unless you know that change is safe, it's the \nwrong thing to do.\n\nOne reason _not_ to do normalization is that if you don't, you can still \ninteract with no ambiguity with other non-Unicode locales. You can do the \n1:1 Latin1<->Unicode translation, and you *never* get into trouble. In \ncotnrast, if you normalize, it's no longer a 1:1 translation any more, and \nyou can get into a situation where the translation from Latin1 to Unicode \nand back results in a *different* filename than the one you started with!\n\nSee? That's a *serious*problem*. A system that forces normalization BY \nDEFINITION cannot work with people who use a Latin1 filesystem, because it \nwill corrupt the filenames!\n\nBut you are apparently too damn stupid to understand that \"data \ncorruption\" == \"bad\", and too damn stupid to see that \"Unicode\" does not \nmean \"Forced normalization\".\n\nBut I'll try one more time. Let's say that I work on a project where there \nare some people who use Latin1, and some people who use UTF-8, and we use \nspecial characters. It should all work, as long as we use only the common \nsubset, and we teach git to convert to UTF-8 as a common base. Right?\n\nIn your *idiotic* world, where you have to normalize and corrupting \nfilenames is ok, that doesn't work! It works wonderfully well if you do \nthe obvious 1:1 translation and you do *not* normalize, but the moment you \nstart normalizing, you actually corrupt the filenames!\n\nAnd yes, the character sequence 'a¨' is exactly one such sequence. It's \nperfectly representable in both Latin1 and in UTF-8: in latin1 it is a \ntwo-character '\\x61\\xa8', and when doing a Latin1->UTF-8 conversion, it \nbecomes '\\x61\\xc2\\xa8', and you can convert back and forth between those \ntwo forms an infinite amount of times, and you never corrupt it.\n\nBut the moment you add normalization to the mix, you start screwing up. \nSuddenly, the sequence '\\x61\\xa8' in Latin1 becomes (assuming NFD) \n'\\xc3\\xa4' in UTF-8, and when converted back to Latin1, it is now '\\xe4', \nie that filename hass been corrupted!\n\nSee? Normalization in the face of working together with others is a total \nand utter mistake, and yes, it really *does* corrupt data. It makes it \nfundamentally impossible to reliably work together with other encodings - \neven when you do converstion between the two!\n\n[ And that's the really sad part. Non-normalized Unicode can pretty much \n  be used as a \"generic encoding\" for just about all locales - if you know \n  the locale you convert from and to, you can generally use UTF-8 as an \n  internal format, knowing that you can always get the same result back in \n  the original encoding. Normalization literally breaks that wonderful \n  generic capability of Unicode.\n\n  And the fact that Unicode is such a \"generic replacement\" for any locale \n  is exactly what makes it so wonderful, and allows you to fairly \n  seamlessly convert piece-meal from some particular locale to Unicode: \n  even if you have some programs that still work in the original locale, \n  you know that you can convert back to it without loss of information.\n\n  Except if you normalize. In that case, you *do* lose information, and \n  suddenly one of the best things about Unicode simply disappears.\n\n  As a result, people who force-normalize are idiots. But they seem to \n  also be stupid enough that they don't understand that they are idiots.\n  Sad. \n\n  It's a bit like whitespace. Whitespace \"doesn't matter\" in text (== is \n  equivalent), but an email client that force-normalizes whitespace in \n  text is a really *broken* email client, because it turns out that \n  sometimes even the \"equivalent\" forms simply do matter. Patches are \n  text, but whitespace is meaningful there. \n\n  Same exact deal: it's good to have the *ability* to normalize \n  whitespace (in email, we call this \"text=flowed\" or similar), and in \n  some ceses you might even want to make it the default action, but \n  *forcing* normalization is total idiocy and actually makes the system \n  less useful! ]\n\n\t\tLinus\n"},{"id":"66160","messageId":"alpine.LFD.1.00.0801211013270.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"Pine.LNX.4.64.0801211509490.17095@ds9.cixit.se","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-21T18:16:10Z","receivedAt":"2008-01-21T18:16:10Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Peter Karlsson wrote:\n> \n> It is relevant because the Mac OS file system stores file names as a\n> sequence of Unicode code points, in a (apparently slightly modified)\n> normalized form, whereas Git prefers to see file systems that store\n> file names as a sequence of octets, which may, or may not, actually map\n> to something that the user would call characters.\n\nNo. The *only* issue is that git doesn't normalize.\n\nYou can think of git as a UTF-8 namespace all you want, and it will work \ntogether wonderfully with OS X. \n\nGit just doesn't force-normalize the names.\n\n> You apparently prefer the text-as-sequence-of-octets, which I tend to\n> dislike because I would have thought computer engineers would have\n> evolved beyond this when we left the 1900s.\n\nSome of us just know what we're doing, and have been working with UTF-8 \nfor a long time. It's not about sequence-of-octets, it's about not \ncorrupting the data.\n\nYou think data should be changed behind peoples backs, potentially causing \ncorruption due to unintended conversions. And I don't.\n\nYou can call me \"left behind in the 1900s\", but that's apparently because \nyou don't understand the issues. Data corruption wasn't something that \nmagically became ok just because we switched into a new century.\n\n\t\t\tLinus\n"},{"id":"66166","messageId":"C6C0E6A1-053B-48CE-90B3-8FFB44061C3B@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801210934400.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T19:05:51Z","receivedAt":"2008-01-21T19:05:51Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 1:12 PM, Linus Torvalds wrote:\n\n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n>> On Jan 21, 2008, at 9:14 AM, Peter Karlsson wrote:\n>>>\n>>> I happen to prefer the text-as-string-of-characters (or code points,\n>>> since you use the other meaning of characters in your posts),  \n>>> since I\n>>> come from the text world, having worked a lot on Unicode text\n>>> processing.\n>>>\n>>> You apparently prefer the text-as-sequence-of-octets, which I tend  \n>>> to\n>>> dislike because I would have thought computer engineers would have\n>>> evolved beyond this when we left the 1900s.\n>>\n>> I agree. Every single problem that I can recall Linus bringing up  \n>> as a\n>> consequence of HFS+ treating filenames as strings [..]\n>\n> You say \"I agree\", BUT YOU DON'T EVEN SEEM TO UNDERSTAND WHAT IS  \n> GOING ON.\n\nI could say the same thing about you.\n\n> The fact is, text-as-string-of-codepoints (let's make the \"codepoints\"\n> obvious, so that there is no ambiguity, but I'd also like to make it  \n> clear\n> that a codepoint *is* how a Unicode character is defined, and a  \n> Unicode\n> \"string\" is actually *defined* to be a sequence of codepoints, and  \n> totally\n> independent of normalization!) is fine.\n>\n> That was never the issue at all. Unicode codepoints are wonderful.\n>\n> Now, git _also_ heavily depends on the actual encoding of those\n> codepoints, since we create hashes etc, so in fact, as far ass git is\n> concerned, names have to be in some particular encoding to be  \n> hashed, and\n> UTF-8 is the only sane encoding for Unicode. People can blather about\n> UCS-2 and UTF-16 and UTF-32 all they want, but the fact is, UTF-8 is\n> simply technically superior in so many ways that I don't even  \n> understand\n> why anybody ever uses anything else.\n>\n> So I would not disagree with using UTF-8 at all.\n>\n> But that is *entirely* a separate issue from \"normalization\".\n>\n> Kevin, you seem to think that normalization is somehow forced on you  \n> by\n> the \"text-as-codepoints\" decision, and that is SIMPLY NOT TRUE.\n> Normalization is a totally separate decision, and it's a STUPID one,\n> because it breaks so many of the _nice_ properties of using UTF-8.\n\nI'm not saying it's forced on you, I'm saying when you treat filenames  \nas text, it DOESN'T MATTER if the string gets normalized. As long as  \nthe string remains equivalent, YOU DON'T CARE about the underlying  \nbyte stream.\n\n> And THAT is where we differ. It has nothing to do with \"octets\". It  \n> has\n> nothing to do with not liking Unicode. It has nothing to do with\n> \"strings\".\n>\n> In short:\n>\n> - normalization is by no means required or even a good feature. It's\n>   something you do when you want to know if two strings are  \n> equivalent,\n>   but that doesn't actually mean that you should keep the strings\n>   normalized all the time!\n\nAlright, fine. I'm not saying HFS+ is right in storing the normalized  \nversion, but I do believe the authors of HFS+ must have had a reason  \nto do that, and I also believe that it shouldn't make any difference  \nto me since it remains equivalent.\n\n> - normalization has *nothing* to do with \"treating text as octets\".\n>   That's entirely an encoding issue.\n\nSure it does. Normalizing a string produces an equivalent string, and  \nso unless I look at the octets the two strings are, for all intents  \nand purposes, the same.\n\n> - of *course* git has to treat things as a binary stream at some  \n> point,\n>   since you need that to even compute a SHA1 in the first place, but  \n> that\n>   has *nothing* to do with normalization or the lack of it.\n\nYou're right, but it doesn't have to treat it as a binary stream at  \nthe level I care about. I mean, no matter what you do at some level  \nthe string is evaluated as a binary stream. For our purposes, just  \nredefine the hashing algorithm to hash all equivalent strings the  \nsame, and you can implement that by using SHA1 on a particular  \nencoding of the string.\n\n> Got it? Forced normalization is stupid, because it changes the data  \n> and\n> removes information, and unless you know that change is safe, it's the\n> wrong thing to do.\n\nDecomposing and recomposing shouldn't lose any information we care  \nabout - when treating filenames as text, a<COMBINING DIARESIS> and <A  \nWITH DIARESIS> are equivalent, and thus no distinction is made between  \nthem. I'm not sure what other information you might be considering  \nlost in this case.\n\n> One reason _not_ to do normalization is that if you don't, you can  \n> still\n> interact with no ambiguity with other non-Unicode locales. You can  \n> do the\n> 1:1 Latin1<->Unicode translation, and you *never* get into trouble. In\n> cotnrast, if you normalize, it's no longer a 1:1 translation any  \n> more, and\n> you can get into a situation where the translation from Latin1 to  \n> Unicode\n> and back results in a *different* filename than the one you started  \n> with!\n\nI don't believe you. See below.\n\n> See? That's a *serious*problem*. A system that forces normalization BY\n> DEFINITION cannot work with people who use a Latin1 filesystem,  \n> because it\n> will corrupt the filenames!\n>\n> But you are apparently too damn stupid to understand that \"data\n> corruption\" == \"bad\", and too damn stupid to see that \"Unicode\" does  \n> not\n> mean \"Forced normalization\".\n\nWhen have I ever said that Unicode meant Forced normalization?\n\n> But I'll try one more time. Let's say that I work on a project where  \n> there\n> are some people who use Latin1, and some people who use UTF-8, and  \n> we use\n> special characters. It should all work, as long as we use only the  \n> common\n> subset, and we teach git to convert to UTF-8 as a common base. Right?\n>\n> In your *idiotic* world, where you have to normalize and corrupting\n> filenames is ok, that doesn't work! It works wonderfully well if you  \n> do\n> the obvious 1:1 translation and you do *not* normalize, but the  \n> moment you\n> start normalizing, you actually corrupt the filenames!\n\nWrong.\n\n> And yes, the character sequence 'a¨' is exactly one such sequence.  \n> It's\n> perfectly representable in both Latin1 and in UTF-8: in latin1 it is a\n> two-character '\\x61\\xa8', and when doing a Latin1->UTF-8 conversion,  \n> it\n> becomes '\\x61\\xc2\\xa8', and you can convert back and forth between  \n> those\n> two forms an infinite amount of times, and you never corrupt it.\n>\n> But the moment you add normalization to the mix, you start screwing  \n> up.\n> Suddenly, the sequence '\\x61\\xa8' in Latin1 becomes (assuming NFD)\n> '\\xc3\\xa4' in UTF-8, and when converted back to Latin1, it is now  \n> '\\xe4',\n> ie that filename hass been corrupted!\n\nWrong. '\\x61\\x18' in Latin1, when converted to UTF-8 (NFD) is still  \n'\\x61\\xc2\\xa8'. You're mixing up DIARESIS (U+00A8) and COMBINING  \nDIARESIS (U+0308).\n\nI suspect this is why you've been yelling so much - you have a  \nfundamental misunderstanding about what normalization is actually doing.\n\n> See? Normalization in the face of working together with others is a  \n> total\n> and utter mistake, and yes, it really *does* corrupt data. It makes it\n> fundamentally impossible to reliably work together with other  \n> encodings -\n> even when you do converstion between the two!\n>\n> [ And that's the really sad part. Non-normalized Unicode can pretty  \n> much\n>  be used as a \"generic encoding\" for just about all locales - if you  \n> know\n>  the locale you convert from and to, you can generally use UTF-8 as an\n>  internal format, knowing that you can always get the same result  \n> back in\n>  the original encoding. Normalization literally breaks that wonderful\n>  generic capability of Unicode.\n>\n>  And the fact that Unicode is such a \"generic replacement\" for any  \n> locale\n>  is exactly what makes it so wonderful, and allows you to fairly\n>  seamlessly convert piece-meal from some particular locale to Unicode:\n>  even if you have some programs that still work in the original  \n> locale,\n>  you know that you can convert back to it without loss of information.\n>\n>  Except if you normalize. In that case, you *do* lose information, and\n>  suddenly one of the best things about Unicode simply disappears.\n\nSee above as to why you're not losing the information you so fervently  \nbelieve you are.\n\n>  As a result, people who force-normalize are idiots. But they seem to\n>  also be stupid enough that they don't understand that they are  \n> idiots.\n>  Sad.\n\nPeople who insult others run the risk of looking like a fool when  \nshown to be wrong.\n\n>  It's a bit like whitespace. Whitespace \"doesn't matter\" in text (==  \n> is\n>  equivalent), but an email client that force-normalizes whitespace in\n>  text is a really *broken* email client, because it turns out that\n>  sometimes even the \"equivalent\" forms simply do matter. Patches are\n>  text, but whitespace is meaningful there.\n>\n>  Same exact deal: it's good to have the *ability* to normalize\n>  whitespace (in email, we call this \"text=flowed\" or similar), and in\n>  some ceses you might even want to make it the default action, but\n>  *forcing* normalization is total idiocy and actually makes the system\n>  less useful! ]\n\nSure, it all depends on what level you need to evaluate text. If we're  \ntalking about english paragraphs, then whitespace can be messed with.  \nWhen we're talking about unicode strings, then specific encoding can  \nbe messed with. When talking about byte sequence, nothing can be  \nmessed with.\n\nIn our case, when working on an HFS+ filesystem all you have to care  \nabout is the unicode string level. The specific encoding can be messed  \nwith, and the client shouldn't care. Problems only arise when  \nattempting to interoperate with filesystems that work at the byte  \nsequence level.\n\nThe only information you lose when doing canonical normalization is  \nwhat the original byte sequence was. Sure, this is a problem when  \nworking on a filesystem that cares about byte sequence, but it's not a  \nproblem when working on a filesystem that cares about the unicode  \nstring.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66174","messageId":"alpine.LFD.1.00.0801211129130.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"C6C0E6A1-053B-48CE-90B3-8FFB44061C3B@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-21T19:41:08Z","receivedAt":"2008-01-21T19:41:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n> \n> I'm not saying it's forced on you, I'm saying when you treat filenames as\n> text, it DOESN'T MATTER if the string gets normalized. As long as the string\n> remains equivalent, YOU DON'T CARE about the underlying byte stream.\n\nSure I do, because it matters a lot for things like - wait for it - things \nlike checksumming it.\n\n> Alright, fine. I'm not saying HFS+ is right in storing the normalized version,\n> but I do believe the authors of HFS+ must have had a reason to do that, and I\n> also believe that it shouldn't make any difference to me since it remains\n> equivalent.\n\nI've already told you the reason: they did the mistake of wanting to be \ncase-independent, and a (bad) case compare is easier in NFD.\n\nOnce you give strings semantic meaning (and \"case independent\" implies \nthat semantic meaning), suddenly normalization looks like a good idea, and \nsince you're going to corrupt the data *anyway*, who cares? You just \ncreated a file like \"Hello\", and readdir() returns \"hello\" (because there \nwas an old file under that name), and it's a lot more obviously corrupt \nthan just due to normalization.\n\n> Sure it does. Normalizing a string produces an equivalent string, and so\n> unless I look at the octets the two strings are, for all intents and purposes,\n> the same.\n\n.. but you *have* to look at the octets at some point. They're kind of \nwhat the string is built up of. They never went away, even if you chose to \nignore them. The encoding is really quite important, and is visible both \nin memory and on disk.\n\nIt's what shows up when you sha1sum, but it's also as simple as what shows \nup when you do an \"ls -l\" and look at a file size.\n\nIt doesn't matter if the text is \"equivalent\", when you then see the \ndifferences in all these small details.\n\nYou can shut your eyes as much as you want, and say that you don't care, \nbut the differences are real, and they are visible.\n\n> Decomposing and recomposing shouldn't lose any information we care about -\n> when treating filenames as text, a<COMBINING DIARESIS> and <A WITH DIARESIS>\n> are equivalent, and thus no distinction is made between them. I'm not sure\n> what other information you might be considering lost in this case.\n\nYou're right, I messed up. I used a non-combining diaeresis, and you're \nright, it doesn't get corrupted. And I think that means that if Apple had \nused NFC, we'd not have this problem with Latin1 systems (because then the \nUTF-8 representation would be the same).\n\nSo I still think that normalization is totally idiotic, but the thing that \nactually causes most problems for people on OS X is that they chose the \nreally inconvenient one.\n\n\t\t\tLinus\n"},{"id":"66175","messageId":"20080121194408.GC11763@glandium.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801210934400.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-01-21T19:44:08Z","receivedAt":"2008-01-21T19:44:08Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Mon, Jan 21, 2008 at 10:12:01AM -0800, Linus Torvalds wrote:\n> \n> \n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n> > On Jan 21, 2008, at 9:14 AM, Peter Karlsson wrote:\n> > > \n> > > I happen to prefer the text-as-string-of-characters (or code points,\n> > > since you use the other meaning of characters in your posts), since I\n> > > come from the text world, having worked a lot on Unicode text\n> > > processing.\n> > > \n> > > You apparently prefer the text-as-sequence-of-octets, which I tend to\n> > > dislike because I would have thought computer engineers would have\n> > > evolved beyond this when we left the 1900s.\n> > \n> > I agree. Every single problem that I can recall Linus bringing up as a\n> > consequence of HFS+ treating filenames as strings [..]\n> \n> You say \"I agree\", BUT YOU DON'T EVEN SEEM TO UNDERSTAND WHAT IS GOING ON.\n> \n> The fact is, text-as-string-of-codepoints (let's make the \"codepoints\" \n> obvious, so that there is no ambiguity, but I'd also like to make it clear \n> that a codepoint *is* how a Unicode character is defined, and a Unicode \n> \"string\" is actually *defined* to be a sequence of codepoints, and totally \n> independent of normalization!) is fine.\n> \n> That was never the issue at all. Unicode codepoints are wonderful.\n> \n> Now, git _also_ heavily depends on the actual encoding of those \n> codepoints, since we create hashes etc, so in fact, as far ass git is \n> concerned, names have to be in some particular encoding to be hashed, and \n> UTF-8 is the only sane encoding for Unicode. People can blather about \n> UCS-2 and UTF-16 and UTF-32 all they want, but the fact is, UTF-8 is \n> simply technically superior in so many ways that I don't even understand \n> why anybody ever uses anything else.\n\nMaybe because it's 1.5 times bigger for any text in chinese, japanese or\nkorean ?\n\nMike\n"},{"id":"66181","messageId":"20080121195703.GE29792@mit.edu","threadId":"11645","inReplyTo":"C6C0E6A1-053B-48CE-90B3-8FFB44061C3B@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-21T19:57:03Z","receivedAt":"2008-01-21T19:57:03Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jan 21, 2008 at 02:05:51PM -0500, Kevin Ballard wrote:\n> You're right, but it doesn't have to treat it as a binary stream at the \n> level I care about. I mean, no matter what you do at some level the string \n> is evaluated as a binary stream. For our purposes, just redefine the \n> hashing algorithm to hash all equivalent strings the same, and you can \n> implement that by using SHA1 on a particular encoding of the string.\n\nThat's horribly broken, for a couple of reasons.  First of all,\nchanging the hash algorithm breaks compatibility with existing\nrepositories; sure, you can try to guess what will least likely break\nexisting repository (which won't be the native MacOSX normalization\nalgorithm, since it's more likely the combined character will likely\nbe used on other environments), but there's still no guarantee there\naren't filenames that use some other form of byte-string for the\nfilename.\n\nSecondly, the hash algorithm would not be stable.  Unicode is not\nstatic, and new characters can get added that may be composable, and\nthus would be normalized differently.  This is one of the reasons why\nUnicode is so horribly broken as a standard.  It was originally\ncreated by representatives from the printing world that were horribly\nclueless about what was needed with respect to canonicalization\nrepresentation, so they compromised allowed both forms, not realizing\nwhat a massive f*ckup this would cause later on.  So people have over\nthe years piled kludges on top of kludges in order to make Unicode\n\"work\".  \n\nSo we can't blame all of the craziness on the MacOS designers,\nalthough they have seen to have been very creative about how to take a\nbad situation and make it worse....\n\n\t\t\t\t\t- Ted\n"},{"id":"66180","messageId":"373E260A-6786-4932-956A-68706AA7C469@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211129130.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T19:58:29Z","receivedAt":"2008-01-21T19:58:29Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 2:41 PM, Linus Torvalds wrote:\n\n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n>>\n>> I'm not saying it's forced on you, I'm saying when you treat  \n>> filenames as\n>> text, it DOESN'T MATTER if the string gets normalized. As long as  \n>> the string\n>> remains equivalent, YOU DON'T CARE about the underlying byte stream.\n>\n> Sure I do, because it matters a lot for things like - wait for it -  \n> things\n> like checksumming it.\n\nI believe I already responded to the issue of hashing. In summary,  \njust re-define your hash function to convert the string to a specific  \nencoding. Sure, you'll lose some speed, but we're already assuming  \nthat it's worth taking a speed hit in order to treat filenames as  \nstrings (please don't argue this point, it's an opinion, not a factual  \nstatement, and I'm not necessarily saying I agree with it, I'm just  \nsaying it's valid).\n\n>> Alright, fine. I'm not saying HFS+ is right in storing the  \n>> normalized version,\n>> but I do believe the authors of HFS+ must have had a reason to do  \n>> that, and I\n>> also believe that it shouldn't make any difference to me since it  \n>> remains\n>> equivalent.\n>\n> I've already told you the reason: they did the mistake of wanting to  \n> be\n> case-independent, and a (bad) case compare is easier in NFD.\n>\n> Once you give strings semantic meaning (and \"case independent\" implies\n> that semantic meaning), suddenly normalization looks like a good  \n> idea, and\n> since you're going to corrupt the data *anyway*, who cares? You just\n> created a file like \"Hello\", and readdir() returns \"hello\" (because  \n> there\n> was an old file under that name), and it's a lot more obviously  \n> corrupt\n> than just due to normalization.\n\nPerhaps that is the reason, I don't know (neither do you, you're just  \nguessing). However, my point still stands - as long as the string  \nstays canonically equivalent, it doesn't matter to me if the  \nfilesystem changes the encoding, since I'm working at the string level.\n\n>> Sure it does. Normalizing a string produces an equivalent string,  \n>> and so\n>> unless I look at the octets the two strings are, for all intents  \n>> and purposes,\n>> the same.\n>\n> .. but you *have* to look at the octets at some point. They're kind of\n> what the string is built up of. They never went away, even if you  \n> chose to\n> ignore them. The encoding is really quite important, and is visible  \n> both\n> in memory and on disk.\n\nSomeone has to look at the octets, but it doesn't have to be me. As  \nlong as I use unicode-aware libraries and such, I can let the  \nunderlying system care about the byte order and my code will be clean.\n\n> It's what shows up when you sha1sum, but it's also as simple as what  \n> shows\n> up when you do an \"ls -l\" and look at a file size.\n\nIt does? Why on earth should it do that? Filename doesn't contribute  \nto the listed filesize on OS X.\n\nkevin@KBLAPTOP:~> echo foo > foo; echo foo > foobar\nkevin@KBLAPTOP:~> ls -l foo*\n-rw-r--r--  1 kevin  kevin  4 Jan 21 14:50 foo\n-rw-r--r--  1 kevin  kevin  4 Jan 21 14:50 foobar\n\nIt would be singularly stupid for the filesize to reflect the  \nfilename, especially since this means you would report different  \nfilesizes for hardlinks.\n\n> It doesn't matter if the text is \"equivalent\", when you then see the\n> differences in all these small details.\n>\n> You can shut your eyes as much as you want, and say that you don't  \n> care,\n> but the differences are real, and they are visible.\n\nVisible at some level, sure, but not visible at the level my code  \nworks on. And thus, I don't have to care about it.\n\n>> Decomposing and recomposing shouldn't lose any information we care  \n>> about -\n>> when treating filenames as text, a<COMBINING DIARESIS> and <A WITH  \n>> DIARESIS>\n>> are equivalent, and thus no distinction is made between them. I'm  \n>> not sure\n>> what other information you might be considering lost in this case.\n>\n> You're right, I messed up. I used a non-combining diaeresis, and  \n> you're\n> right, it doesn't get corrupted. And I think that means that if  \n> Apple had\n> used NFC, we'd not have this problem with Latin1 systems (because  \n> then the\n> UTF-8 representation would be the same).\n\nI'm not sure what you mean. The byte sequence is different from Latin1  \nto UTF-8 even if you use NFC, so I don't think, in this case, it makes  \nany difference whether you use NFC or NFD. Yes, the codepoints are the  \nsame in Latin1 and UTF-8 if you use NFC, but that's hardly relevant.  \nPlease correct me if I'm wrong, but I believe Latin1->UTF-8->Latin1  \nconversion will always produce the same Latin1 text whether you use  \nNFC or NFD.\n\n> So I still think that normalization is totally idiotic, but the  \n> thing that\n> actually causes most problems for people on OS X is that they chose  \n> the\n> really inconvenient one.\n\nThe only reason it's particularly inconvenient is because it's  \ndifferent from what most other systems picked. And if you want to  \nblame someone for that, blame Unicode for having so many different  \nnormalization forms.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66182","messageId":"998717B0-0165-4383-AAB8-33BD2A49954E@sb.org","threadId":"11645","inReplyTo":"20080121195703.GE29792@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T20:01:43Z","receivedAt":"2008-01-21T20:01:43Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 2:57 PM, Theodore Tso wrote:\n\n> On Mon, Jan 21, 2008 at 02:05:51PM -0500, Kevin Ballard wrote:\n>> You're right, but it doesn't have to treat it as a binary stream at  \n>> the\n>> level I care about. I mean, no matter what you do at some level the  \n>> string\n>> is evaluated as a binary stream. For our purposes, just redefine the\n>> hashing algorithm to hash all equivalent strings the same, and you  \n>> can\n>> implement that by using SHA1 on a particular encoding of the string.\n>\n> That's horribly broken, for a couple of reasons.  First of all,\n> changing the hash algorithm breaks compatibility with existing\n> repositories; sure, you can try to guess what will least likely break\n> existing repository (which won't be the native MacOSX normalization\n> algorithm, since it's more likely the combined character will likely\n> be used on other environments), but there's still no guarantee there\n> aren't filenames that use some other form of byte-string for the\n> filename.\n>\n> Secondly, the hash algorithm would not be stable.  Unicode is not\n> static, and new characters can get added that may be composable, and\n> thus would be normalized differently.  This is one of the reasons why\n> Unicode is so horribly broken as a standard.  It was originally\n> created by representatives from the printing world that were horribly\n> clueless about what was needed with respect to canonicalization\n> representation, so they compromised allowed both forms, not realizing\n> what a massive f*ckup this would cause later on.  So people have over\n> the years piled kludges on top of kludges in order to make Unicode\n> \"work\".\n>\n> So we can't blame all of the craziness on the MacOS designers,\n> although they have seen to have been very creative about how to take a\n> bad situation and make it worse..\n\nYou seem to be under the impression that I'm advocating that git treat  \nall filenames as unicode strings, and thus change its hashing  \nalgorithm as described. I am not. I am saying that, if git only had to  \ndeal with HFS+, then it could treat all filenames as strings, etc.  \nHowever, since git does not only have to deal with HFS+, this will not  \nwork. What I am describing is an ideal, not a practicality.\n\nIn other words, what I'm saying is that treating filenames as strings  \nworks perfectly fine, *provided you can do that 100% of the time*. git  \ncannot do that 100% of the time, therefore it's not appropriate here.  \nThe purpose of this argument is to illustrate that treating filenames  \nas strings isn't wrong, it's simply incompatible with treating  \nfilenames as byte sequences.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66183","messageId":"20080121201530.GF29792@mit.edu","threadId":"11645","inReplyTo":"998717B0-0165-4383-AAB8-33BD2A49954E@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-21T20:15:30Z","receivedAt":"2008-01-21T20:15:30Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jan 21, 2008 at 03:01:43PM -0500, Kevin Ballard wrote:\n>\n> You seem to be under the impression that I'm advocating that git treat all \n> filenames as unicode strings, and thus change its hashing algorithm as \n> described. I am not. I am saying that, if git only had to deal with HFS+, \n> then it could treat all filenames as strings, etc. However, since git does \n> not only have to deal with HFS+, this will not work. What I am describing \n> is an ideal, not a practicality.\n\nWell, why are you arguing on the git list about precisely that (when\nyou reponsed to Linus), then?\n\n> In other words, what I'm saying is that treating filenames as strings works \n> perfectly fine, *provided you can do that 100% of the time*. git cannot do \n> that 100% of the time, therefore it's not appropriate here. The purpose of \n> this argument is to illustrate that treating filenames as strings isn't \n> wrong, it's simply incompatible with treating filenames as byte sequences.\n\nNo, it's still broken, because of the Unicode-is-not-static problem.\nWhat happens when you start adding more composable characters, which\nsome future version of HFS+ will start breaking apart? \n\nPresumably the whole *reason* why HFS+ was corrupting strings was so\nthat \"stupid applications\" that only did byte comparisons would work\ncorrectly.  But when you upgrade from Mac OS 10.5 to 10.6, and it adds\nsupport for new composable characters, and you now take a USB hard\ndrive that was hooked up to a MacBook Air, running one version of\nMacOS, and hook it up to another Macintosh, running another version of\nMacOS, the normalization algorithm will be different, so the byte\ncomparisons won't work.  \n\nSo all of this extra work which MacOS put in to corrupt filenames\nbehind our back doesn't actually do any good; applications still need\nto be smart, or there will be rare, hard to reproduce bugs\nnevertheless.  So if MacOS wants to supply Unicode libraries that\ncompare strings keeping in mind Unicode \"equivalences\" it can be our\nguest (although how they deal with different versions of Unicode with\ndifferent equivalence classes will be their cross to bear).  BUT MacOS\nX SHOULD NOT BE CORRUPTING FILENAMES.  TO DO SO IS BROKEN.\n\nEven Microsoft got this right; its filesystem is case-preserving, but\nit has case-insensitive lookups.  Hence, it is not corrupting\nfilenames behind the application's back, unlike MacOS.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"66187","messageId":"20080121203043.GV14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"440E4426-BFB5-4836-93DF-05C99EF204E6@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-21T20:30:43Z","receivedAt":"2008-01-21T20:30:43Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Jan 21, 2008 at 11:43:54AM -0500, Kevin Ballard wrote:\n> \n> I agree. Every single problem that I can recall Linus bringing up as a  \n> consequence of HFS+ treating filenames as strings is in fact only a  \n> problem if you then think of the filename as octets at some point.\n\nAt *some* point everything stored in computers is a sequence of octets.\nIn fact, the whole point of the Unicode standard is to define characters\nand how to map each character to a unique number (code points) and then\nhow to encode this number into sequence of octets.\n\n> If  \n> you stick with UTF-8 equivalence comparison the entire time, then  \n> everything just works.\n\nThere are more than one equivalence comparison. The unicode standard\ndefines at least two, and for some other purpose you may want to use\nsome others, but for some reason you are trying to present that to\nwork with text means to follow only one type of equivalence the entire\ntime...\n\nDmitry\n"},{"id":"66186","messageId":"8F85366A-C990-47B1-BF60-936185B9E438@sb.org","threadId":"11645","inReplyTo":"20080121201530.GF29792@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T20:31:02Z","receivedAt":"2008-01-21T20:31:02Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 3:15 PM, Theodore Tso wrote:\n\n> On Mon, Jan 21, 2008 at 03:01:43PM -0500, Kevin Ballard wrote:\n>>\n>> You seem to be under the impression that I'm advocating that git  \n>> treat all\n>> filenames as unicode strings, and thus change its hashing algorithm  \n>> as\n>> described. I am not. I am saying that, if git only had to deal with  \n>> HFS+,\n>> then it could treat all filenames as strings, etc. However, since  \n>> git does\n>> not only have to deal with HFS+, this will not work. What I am  \n>> describing\n>> is an ideal, not a practicality.\n>\n> Well, why are you arguing on the git list about precisely that (when\n> you reponsed to Linus), then?\n\nBecause of the way in which an argument evolves. This started out as  \n\"HFS+ is stupid because it normalizes\", and I was arguing that said  \nnormalization wasn't stupid. This turned into an argument as to why HFS \n+ wasn't stupid for normalization, which is basically this argument of  \nthe ideal. Yes, I realize that it's not producing any practical  \nresults, but I'm stubborn (as, apparently, are most of you), and I  \nbelieve that if the official stance of the git project is \"HFS+ is  \nstupid\" then there's a lower chance of a patch being accepted then if  \npeople accept that \"HFS+ is different in an incompatible fashion\".\n\n>> In other words, what I'm saying is that treating filenames as  \n>> strings works\n>> perfectly fine, *provided you can do that 100% of the time*. git  \n>> cannot do\n>> that 100% of the time, therefore it's not appropriate here. The  \n>> purpose of\n>> this argument is to illustrate that treating filenames as strings  \n>> isn't\n>> wrong, it's simply incompatible with treating filenames as byte  \n>> sequences.\n>\n> No, it's still broken, because of the Unicode-is-not-static problem.\n> What happens when you start adding more composable characters, which\n> some future version of HFS+ will start breaking apart?\n\nIf you need a static representation, you normalize to a specific form.  \nAnd in fact, adding new composable characters doesn't matter, since if  \nthey didn't exist before, you couldn't have possibly used them. Unless  \nyou mean adding new composed forms of existing simpler characters, at  \nwhich point you seem to be arguing for NFD instead of NFC.\n\n> Presumably the whole *reason* why HFS+ was corrupting strings was so\n> that \"stupid applications\" that only did byte comparisons would work\n> correctly.  But when you upgrade from Mac OS 10.5 to 10.6, and it adds\n> support for new composable characters, and you now take a USB hard\n> drive that was hooked up to a MacBook Air, running one version of\n> MacOS, and hook it up to another Macintosh, running another version of\n> MacOS, the normalization algorithm will be different, so the byte\n> comparisons won't work.\n\nI doubt that HFS+ normalized so that \"stupid applications\" could do  \nbyte comparisons. But even if that were the case, see previous  \nparagraph.\n\n> So all of this extra work which MacOS put in to corrupt filenames\n> behind our back doesn't actually do any good; applications still need\n> to be smart, or there will be rare, hard to reproduce bugs\n> nevertheless.  So if MacOS wants to supply Unicode libraries that\n> compare strings keeping in mind Unicode \"equivalences\" it can be our\n> guest (although how they deal with different versions of Unicode with\n> different equivalence classes will be their cross to bear).  BUT MacOS\n> X SHOULD NOT BE CORRUPTING FILENAMES.  TO DO SO IS BROKEN.\n\nYour entire argument is based on the assumption that HFS+ \"corrupts\"  \nfilenames in order to allow dumb clients to do byte comparisons, and I  \ndon't believe that to be the case. In fact, it's only considered a  \ncorruption if you care about the byte sequence of filenames, and my  \nargument is that, on HFS+, you aren't supposed to care about the byte  \nsequence.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66248","messageId":"85tzl66ht6.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211151180.20753@xanadu.home","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-21T20:32:21Z","receivedAt":"2008-01-21T20:32:21Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> Normalization will always hurt performances.  This is an overhead.\n> Sometimes that overhead might be insignificant and not be perceptible,\n> but sometimes it is.  And Git is clearly in the later\n> case. Performances will be hurt big time the day it is made aware of\n> that normalization.  This is why there is so much resistance about it,\n> especially when the benefits of normalizing file names are not shown\n> to be worth their cost in performance and complexity, as other systems\n> do rather fine without it.\n\nNormalization is cheap if you normalize user input.  The user will\nalways be quite slower than any reasonable normalization algorithm.  But\nin the filesystem, one is normalizing the same stuff over and over.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66188","messageId":"alpine.LFD.1.00.0801211210270.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"373E260A-6786-4932-956A-68706AA7C469@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-21T20:33:47Z","receivedAt":"2008-01-21T20:33:47Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n> \n> > It's what shows up when you sha1sum, but it's also as simple as what shows\n> > up when you do an \"ls -l\" and look at a file size.\n> \n> It does? Why on earth should it do that? Filename doesn't contribute to the\n> listed filesize on OS X.\n\nUmm. What's this inability to see that data is data is data?\n\nWhy do you think Unicode has anything in particular to do with filenames?\n\nThose same unicode strings are often part of the file data itself, and \nthen that encoding damn well is visible in \"ls -l\".\n\nDoing\n\n\techo ä > file\n\tls -l file\n\nsure shows that \"underlying octet\" thing that you wanted to avoid so much. \nMy point was that those underlying octets are always there, and they do \nmatter. The fact that the differences may not be visible when you compare \nthe normalized forms doesn't make it any less true.\n\nYou can choose to put blinders on and try to claim that normalization is \ninvisible, but it's only invisible TO THOSE THINGS THAT DON'T WANT TO SEE \nIT.\n\nBut that doesn't change the fact that a lot of things *do* see it. There \nare very few things that are \"Unicode specific\", and a *lot* of tools that \nare just \"general data tools\".\n\nAnd git tries to be a general data tool, not a Unicode-specific one.\n\n> I'm not sure what you mean. The byte sequence is different from Latin1 to\n> UTF-8 even if you use NFC, so I don't think, in this case, it makes any\n> difference whether you use NFC or NFD.\n>\n> Yes, the codepoints are the same in Latin1 and UTF-8 if you use NFC, but \n> that's hardly relevant. Please correct me if I'm wrong, but I believe \n> Latin1->UTF-8->Latin1 conversion will always produce the same Latin1 \n> text whether you use NFC or NFD.\n\nThe problem is that the UTF-8 form is different, so if you save things in \nUTF-8 (which we hopefully agree is a sane thing to do), then you should \ntry to use a representation that people agree on.\n\nAnd NFC is the more common normalization form by far, so by normalizing to \nsomething else, you actually de-normalize as far as those other people are \nconcerned.\n\nSo if you have to normalize, at least use the normal form!\n\n> The only reason it's particularly inconvenient is because it's different from\n> what most other systems picked. And if you want to blame someone for that,\n> blame Unicode for having so many different normalization forms.\n\nI blame them for encouraging normalization at all.\n\nIt's stupid.\n\nYou don't need it.\n\nThe people who care about \"are these strings equivalent\" shouldn't do a \n\"memcmp()\" on them in the first place. And if you don't do a memcmp() on \nthings, then you don't need to normalize. \n\nSo you have two cases:\n (a) the cases that care about *identity*. They don't want normalization\n (b) the cases that care about *equivalence*. And they shouldn't do \n      octet-by-octet comparison.\n\nSee? Either you want to see equivalence, or you don't. And in neither case \nis normalization the right thing to do (except as *possibly* an internal \npart of the comparison, but there are actually better ways to check for \nequivalence than the brute-force \"normalize both and compare results \nbitwise\").\n\n\t\t\tLinus\n"},{"id":"66242","messageId":"85prvu6hnl.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"AE99FDAA-F8D3-49F7-A0B9-CDFCC4903824@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-21T20:35:42Z","receivedAt":"2008-01-21T20:35:42Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Kevin Ballard <kevin@sb.org> writes:\n\n> I agree, Linus's approach is indeed fast. And if speed is more\n> important than treating filenames as text instead of octets, then so\n> be it. This is a trade-off. But a trade-off doesn't mean one approach\n> is \"wrong\", it just means the authors of HFS+ thought it was an\n> acceptable trade-off. HFS+ wasn't designed to be a high-performance\n> filesystem that deals with lots of files, it was designed to be a\n> filesystem used by regular people on the Mac, and I believe treating\n> filenames as text is a good choice in this scenario.\n\nRegular people have brains, not filesystems.  HFS+ is employed by\ncomputers, and computers can produce or query or process lots of data in\nvery short time spans, in their own pace.  And if Mac users did not want\nto make use of that, they would still be using Mac classics.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66199","messageId":"20080121203629.GW14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801210934400.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-21T20:36:29Z","receivedAt":"2008-01-21T20:36:29Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Jan 21, 2008 at 10:12:01AM -0800, Linus Torvalds wrote:\n> \n> The fact is, text-as-string-of-codepoints (let's make the \"codepoints\" \n> obvious, so that there is no ambiguity, but I'd also like to make it clear \n> that a codepoint *is* how a Unicode character is defined, and a Unicode \n> \"string\" is actually *defined* to be a sequence of codepoints, and totally \n> independent of normalization!) is fine.\n\nCode point is a unique numerical value assigned to every Unicode character.\nAlso, every Unicode character has a uniqie name assigned to it. There are\nsome other non-unique properties that every Unicode has. So, to say that\na Unicode character is just a code point is not exactly correct, because\nthe code point is one of properties of a unicode character. But, yes, any\nUnicode character can be identified by its code point. So, it is one to\none relation.\n\nDmitry\n"},{"id":"66190","messageId":"20080121204330.GX14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"CFF9E74C-4A4C-4E5F-8DA3-662D80095503@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-21T20:43:30Z","receivedAt":"2008-01-21T20:43:30Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Jan 21, 2008 at 11:59:24AM -0500, Kevin Ballard wrote:\n> \n> No, it's a question of hashing algorithm. And it's one that's fairly  \n> easily solved simply by picking a specific nonambiguous UTF-8 encoding  \n> before hashing.\n\nUTF-8 is a *single* encoding, and it maps every Unicode character to\na unique binary representation. So, it is completely nonambiguous.\n\nDmitry\n"},{"id":"66191","messageId":"20080121204614.GG29792@mit.edu","threadId":"11645","inReplyTo":"8F85366A-C990-47B1-BF60-936185B9E438@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-21T20:46:14Z","receivedAt":"2008-01-21T20:46:14Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jan 21, 2008 at 03:31:02PM -0500, Kevin Ballard wrote:\n>> No, it's still broken, because of the Unicode-is-not-static problem.\n>> What happens when you start adding more composable characters, which\n>> some future version of HFS+ will start breaking apart?\n>\n> If you need a static representation, you normalize to a specific form. And \n> in fact, adding new composable characters doesn't matter, since if they \n> didn't exist before, you couldn't have possibly used them. \n\nSure you can.  Suppose you unpack the same tar file or zip file that\ncontains one of these new-fangled characters, one on a MacOS 10.5\nsystem, and one on a MacOS 10.9 system.  How HFS+ will corrupt that\nfilename will differ depending which version of MacOS you are running.\nHence, normalizing the filename when you store it is stupid and\nbroken.  MacOS and its applications and libraries want to do\nnormalization in the privacy of its own address space, that's it's\nbusiness.  It can pursue any fetish it wants, among consenting adults.\nSafe, sane and consensual, and all that... well, consensual, anyway.\nI'm not sure about \"safe\" and \"sane\"....\n\nMy arguement is basically is that there is absolutely no value in what\nHFS+ is doing, by corrupting filenames --- if you want to call it\n\"normalizing\" them, fine, but since Unicode is not static, so you\ncan't even call it a \"canonical\" form.  It's just some random\ncorruption of what was passed in at open(2) time, that can and will\nchange depending on what version of MacOS you are running.\n\nIf you want to play the insane Unicode game of \"equivalent\"\ncharacters, you have to do it at comparison time, so there's no point\ntrying to \"normalize\" them when you store them.  It doesn't buy you\nanything, and it causes all sorts of pain.\n\n> Your entire argument is based on the assumption that HFS+ \"corrupts\" \n> filenames in order to allow dumb clients to do byte comparisons, and I \n> don't believe that to be the case. \n\nOK, what's your reason for why HFS+ corrupts filenames?  What do you\nthink is its excuse?  What problem does it solve?  If the answer is\n\"no reason at all, but because it *can*\", according to the Great God\nUnicode, then that's really not very impressive....\n\n\t\t\t\t\t\t- Ted\n"},{"id":"66193","messageId":"A101AC54-F0CF-4C33-8986-CC6B8DD7F17F@sb.org","threadId":"11645","inReplyTo":"20080121204330.GX14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T20:53:10Z","receivedAt":"2008-01-21T20:53:10Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 3:43 PM, Dmitry Potapov wrote:\n\n> On Mon, Jan 21, 2008 at 11:59:24AM -0500, Kevin Ballard wrote:\n>>\n>> No, it's a question of hashing algorithm. And it's one that's fairly\n>> easily solved simply by picking a specific nonambiguous UTF-8  \n>> encoding\n>> before hashing.\n>\n> UTF-8 is a *single* encoding, and it maps every Unicode character to\n> a unique binary representation. So, it is completely nonambiguous.\n\nIn this case, encoding refers to normalization form, as other people  \nhave used it in the conversation besides me.\n\nI suggest you stop trying to find inconsequential stuff to argue  \nabout, especially when a tiny bit of critical thinking would reveal  \nthe answer.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66194","messageId":"7EB98659-4036-45DA-BD50-42CB23ED517A@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211210270.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T20:53:25Z","receivedAt":"2008-01-21T20:53:25Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 3:33 PM, Linus Torvalds wrote:\n\n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n>>\n>>> It's what shows up when you sha1sum, but it's also as simple as  \n>>> what shows\n>>> up when you do an \"ls -l\" and look at a file size.\n>>\n>> It does? Why on earth should it do that? Filename doesn't  \n>> contribute to the\n>> listed filesize on OS X.\n>\n> Umm. What's this inability to see that data is data is data?\n\nI'm not sure what you mean. I stated a fact - at least on OS X, the  \nfilename does not contribute to the listed filesize, so changing the  \nencoding of the filename doesn't change the filesize. This isn't a  \nphilosophical point, it's a factual statement.\n\n> Why do you think Unicode has anything in particular to do with  \n> filenames?\n\nI don't, but I do think this discussion revolves around filenames,  \ntherefore it should not surprise you when I talk about filenames.\n\n> Those same unicode strings are often part of the file data itself, and\n> then that encoding damn well is visible in \"ls -l\".\n>\n> Doing\n>\n> \techo ä > file\n> \tls -l file\n>\n> sure shows that \"underlying octet\" thing that you wanted to avoid so  \n> much.\n> My point was that those underlying octets are always there, and they  \n> do\n> matter. The fact that the differences may not be visible when you  \n> compare\n> the normalized forms doesn't make it any less true.\n\nYes, I am well aware that the encoding of the *file contents* affects  \nfilesize. But when did I suggest changing the encoding of filenames  \ninside file contents? If you treat filenames as strings, there's no  \nrequirement to change the encoding of filenames inside file contents.  \nI'm talking specifically about the filenames, not about file contents,  \nso stop trying to argue against that which is irrelevant.\n\n> You can choose to put blinders on and try to claim that  \n> normalization is\n> invisible, but it's only invisible TO THOSE THINGS THAT DON'T WANT  \n> TO SEE\n> IT.\n\nDon't want to, or don't need to? It's not a matter of ignoring  \nencoding because I don't want to deal with it, it's ignoring encoding  \nbecause it's simply not relevant if I treat filenames as strings.\n\n> But that doesn't change the fact that a lot of things *do* see it.  \n> There\n> are very few things that are \"Unicode specific\", and a *lot* of  \n> tools that\n> are just \"general data tools\".\n>\n> And git tries to be a general data tool, not a Unicode-specific one.\n\nYes, I realize that. See my previous message about discussing ideal vs  \npracticality.\n\n>> I'm not sure what you mean. The byte sequence is different from  \n>> Latin1 to\n>> UTF-8 even if you use NFC, so I don't think, in this case, it makes  \n>> any\n>> difference whether you use NFC or NFD.\n>>\n>> Yes, the codepoints are the same in Latin1 and UTF-8 if you use  \n>> NFC, but\n>> that's hardly relevant. Please correct me if I'm wrong, but I believe\n>> Latin1->UTF-8->Latin1 conversion will always produce the same Latin1\n>> text whether you use NFC or NFD.\n>\n> The problem is that the UTF-8 form is different, so if you save  \n> things in\n> UTF-8 (which we hopefully agree is a sane thing to do), then you  \n> should\n> try to use a representation that people agree on.\n>\n> And NFC is the more common normalization form by far, so by  \n> normalizing to\n> something else, you actually de-normalize as far as those other  \n> people are\n> concerned.\n>\n> So if you have to normalize, at least use the normal form!\n\nWas NFC the common normalization form back in 1998? My understanding  \nis Unicode was still in the process of being adopted back then, so  \nthere was no one common standard that was obvious for everyone to use.\n\n>> The only reason it's particularly inconvenient is because it's  \n>> different from\n>> what most other systems picked. And if you want to blame someone  \n>> for that,\n>> blame Unicode for having so many different normalization forms.\n>\n> I blame them for encouraging normalization at all.\n>\n> It's stupid.\n>\n> You don't need it.\n>\n> The people who care about \"are these strings equivalent\" shouldn't  \n> do a\n> \"memcmp()\" on them in the first place. And if you don't do a  \n> memcmp() on\n> things, then you don't need to normalize.\n>\n> So you have two cases:\n> (a) the cases that care about *identity*. They don't want  \n> normalization\n> (b) the cases that care about *equivalence*. And they shouldn't do\n>      octet-by-octet comparison.\n>\n> See? Either you want to see equivalence, or you don't. And in  \n> neither case\n> is normalization the right thing to do (except as *possibly* an  \n> internal\n> part of the comparison, but there are actually better ways to check  \n> for\n> equivalence than the brute-force \"normalize both and compare results\n> bitwise\").\n\nI could argue against this, but frankly, I'm really tired of arguing  \nthis same point. I suggest we simply agree to disagree, and move on to  \nactually fixing the problem.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66195","messageId":"20080121205615.GY14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"C6C0E6A1-053B-48CE-90B3-8FFB44061C3B@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-21T20:56:15Z","receivedAt":"2008-01-21T20:56:15Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Jan 21, 2008 at 02:05:51PM -0500, Kevin Ballard wrote:\n> >\n> >But that is *entirely* a separate issue from \"normalization\".\n> >\n> >Kevin, you seem to think that normalization is somehow forced on you  \n> >by\n> >the \"text-as-codepoints\" decision, and that is SIMPLY NOT TRUE.\n> >Normalization is a totally separate decision, and it's a STUPID one,\n> >because it breaks so many of the _nice_ properties of using UTF-8.\n> \n> I'm not saying it's forced on you, I'm saying when you treat filenames  \n> as text,\n\nto treat as text could mean different for different people. Some\nmay prefer to fi and fi_ligature to be treated as same in some\ncontext.\n\n> it DOESN'T MATTER if the string gets normalized. As long as  \n> the string remains equivalent,\n\nAs matter of fact it does, otherwise characters would be the\nsame and we would not have this conversation at all. String\ncan be equivalent and not equivalent at the time, because there\nare different equivalent relations. Finally, what HFS+ does\nis even not normalization. In the technote, Apple explains\nthat they decompose some characters but not others for better\ncompatibility. So, you see, there is a PROBLEM here.\n\n> YOU DON'T CARE about the underlying  \n> byte stream.\n\nIt is not about byte stream. After all, if it were UTF-16 instead\nof UTF-8, it would be one to one conversion for each character.\nSo, what gets corrupted by HFS+ are Unicode *characters*.\n\n> \n> Alright, fine. I'm not saying HFS+ is right in storing the normalized  \n> version, but I do believe the authors of HFS+ must have had a reason  \n> to do that,\n\nI don't say they do that without *any* reason, but I suppose all\nApple developers in the Copland project had some reasons for they\ndid, but the outcome was not very good...\n\n> The only information you lose when doing canonical normalization is  \n> what the original byte sequence was. \n\nNot true. You lose the original sequence of *characters*.\n\nDmitry\n"},{"id":"66244","messageId":"85ejca6glx.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"7EB98659-4036-45DA-BD50-42CB23ED517A@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-21T20:58:18Z","receivedAt":"2008-01-21T20:58:18Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Kevin Ballard <kevin@sb.org> writes:\n\n> On Jan 21, 2008, at 3:33 PM, Linus Torvalds wrote:\n>\n>> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n\n>>> It does? Why on earth should it do that? Filename doesn't\n>>> contribute to the\n>>> listed filesize on OS X.\n>>\n>> Umm. What's this inability to see that data is data is data?\n>\n> I'm not sure what you mean. I stated a fact - at least on OS X, the\n> filename does not contribute to the listed filesize, so changing the\n> encoding of the filename doesn't change the filesize. This isn't a\n> philosophical point, it's a factual statement.\n\nChanging the encoding of the file name most certainly changes the\nfile size of the directory.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66196","messageId":"20863708-2F5A-4E9A-BA67-A6C29D324BD8@sb.org","threadId":"11645","inReplyTo":"20080121204614.GG29792@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T20:59:51Z","receivedAt":"2008-01-21T20:59:51Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"Note: resent to list due to bounce.\nOriginal CC list: tytso@MIT.EDU, torvalds@linux-foundation.org, peter@softwolves.pp.se \n, mjscod@web.de, melo@simplicidade.org\n\nOn Jan 21, 2008, at 3:46 PM, Theodore Tso wrote:\n\n> On Mon, Jan 21, 2008 at 03:31:02PM -0500, Kevin Ballard wrote:\n>>> No, it's still broken, because of the Unicode-is-not-static problem.\n>>> What happens when you start adding more composable characters, which\n>>> some future version of HFS+ will start breaking apart?\n>>\n>> If you need a static representation, you normalize to a specific  \n>> form. And\n>> in fact, adding new composable characters doesn't matter, since if  \n>> they\n>> didn't exist before, you couldn't have possibly used them.\n>\n> Sure you can.  Suppose you unpack the same tar file or zip file that\n> contains one of these new-fangled characters, one on a MacOS 10.5\n> system, and one on a MacOS 10.9 system.  How HFS+ will corrupt that\n> filename will differ depending which version of MacOS you are running.\n> Hence, normalizing the filename when you store it is stupid and\n> broken.  MacOS and its applications and libraries want to do\n> normalization in the privacy of its own address space, that's it's\n> business.  It can pursue any fetish it wants, among consenting adults.\n> Safe, sane and consensual, and all that... well, consensual, anyway.\n> I'm not sure about \"safe\" and \"sane\"....\n\nYou're making the huge assumption that the HFS+ normalization  \nalgorithms will change. As the technote states:\n\n\"Platform algorithms tend to evolve with the Unicode standard. The HFS  \nPlus algorithms cannot evolve because such evolution would invalidate  \nexisting HFS Plus volumes.\"\n\n> My arguement is basically is that there is absolutely no value in what\n> HFS+ is doing, by corrupting filenames --- if you want to call it\n> \"normalizing\" them, fine, but since Unicode is not static, so you\n> can't even call it a \"canonical\" form.  It's just some random\n> corruption of what was passed in at open(2) time, that can and will\n> change depending on what version of MacOS you are running.\n\nAgain with the huge assumptions.\n\n> If you want to play the insane Unicode game of \"equivalent\"\n> characters, you have to do it at comparison time, so there's no point\n> trying to \"normalize\" them when you store them.  It doesn't buy you\n> anything, and it causes all sorts of pain.\n\nIt must have bought somebody something, or they never would have done  \nit.\n\n>> Your entire argument is based on the assumption that HFS+ \"corrupts\"\n>> filenames in order to allow dumb clients to do byte comparisons,  \n>> and I\n>> don't believe that to be the case.\n>\n> OK, what's your reason for why HFS+ corrupts filenames?  What do you\n> think is its excuse?  What problem does it solve?  If the answer is\n> \"no reason at all, but because it *can*\", according to the Great God\n> Unicode, then that's really not very impressive....\n\nI have no idea why HFS+ stores filenames in a normalized form, and  \nfurther I am smart enough to know that speculating is completely  \npointless. I assume the authors had a good reason (which should be a  \nsafe assumption, filesystem authors are a smart bunch). The reason may  \nnot be valid anymore, but if it was valid back in 1998, then I can  \naccept it without complaining.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66245","messageId":"85abmy6g9c.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"A101AC54-F0CF-4C33-8986-CC6B8DD7F17F@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-21T21:05:51Z","receivedAt":"2008-01-21T21:05:51Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Kevin Ballard <kevin@sb.org> writes:\n\n> On Jan 21, 2008, at 3:43 PM, Dmitry Potapov wrote:\n>\n>> On Mon, Jan 21, 2008 at 11:59:24AM -0500, Kevin Ballard wrote:\n>>>\n>>> No, it's a question of hashing algorithm. And it's one that's fairly\n>>> easily solved simply by picking a specific nonambiguous UTF-8\n>>> encoding before hashing.\n>>\n>> UTF-8 is a *single* encoding, and it maps every Unicode character to\n>> a unique binary representation. So, it is completely nonambiguous.\n>\n> In this case, encoding refers to normalization form, as other people\n> have used it in the conversation besides me.\n\nThere exists more than one \"normalization form\" (even across MacOS\nplatforms), and git is cross-platform.  And people can't be made to\nagree on normalization forms, anyway.  You are aware that Unicode code\npoints are shared between some Chinese and Japanese signs, and that\nstroked forms might be composed differently in different languages?  We\ndon't need to go to the Far East, anyway: in Turkish, İ and i are\nequivalent, as are I and ı, whereas in other European languages, I is\ninstead equivalent to i.  In the Netherlands, ÿ is IIRC equivalent to\nij.  And so on.\n\n> I suggest you stop trying to find inconsequential stuff to argue\n> about, especially when a tiny bit of critical thinking would reveal\n> the answer.\n\nNow that you have established that you are the only person on the list\ncapable of critical thinking, how about going elsewhere where you can\nfind similarly sharp critical thinkers?\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66197","messageId":"46a038f90801211306g3dd9a167wb74d06e444b18b93@mail.gmail.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801210934400.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-21T21:06:05Z","receivedAt":"2008-01-21T21:06:05Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 7:12 AM, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> Now, git _also_ heavily depends on the actual encoding of those\n> codepoints, since we create hashes etc, so in fact, as far ass git is\n> concerned, names have to be in some particular encoding to be hashed, and\n> UTF-8 is the only sane encoding for Unicode. People can blather about\n> UCS-2 and UTF-16 and UTF-32 all they want, but the fact is, UTF-8 is\n> simply technically superior in so many ways that I don't even understand\n> why anybody ever uses anything else.\n>\n> So I would not disagree with using UTF-8 at all.\n\nLinus,\n\n(slightly offtopic) are you praising UTF-8 as storage format (for disk\nand network) or in general? UTF-8-aware string ops like counting\ncharacters seem to me a horrendous thing at the ASM level.\n\nMore on topic, I suspect Kevin's experience is more on end-user apps,\nwhere input sanitization and even canonicalisation are common\npractice. From a kernel and filesystems POV, a filename is data as\nsacred as file data. On the webapp world, we \"corrupt\" user input\nliberally to avoid XSS attacks and the like. In some cases, these\npractices are stupid and can be replaced with escaping data properly,\nbut in other cases, the web platform is so broken that there's no\noption.\n\nAt least in Moodle we store *exactly*  what the user POSTed and\ncleanup^Wcorrupt it when displaying it, so that if it does happen that\nthe cleanup was buggy, we never corrupted the data.\n\nSo no point in calling eachother stupid this much. Once is enough ;-)\nAnd no point in arguing that something that is ok for an end-user app\nis a good design decision for an OS.\n\n\nmartin\n"},{"id":"66198","messageId":"DEC058ED-EBF0-4A1E-BF7B-448B16DBBD6E@sb.org","threadId":"11645","inReplyTo":"20080121205615.GY14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T21:07:27Z","receivedAt":"2008-01-21T21:07:27Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 3:56 PM, Dmitry Potapov wrote:\n\n> On Mon, Jan 21, 2008 at 02:05:51PM -0500, Kevin Ballard wrote:\n>>>\n>>> But that is *entirely* a separate issue from \"normalization\".\n>>>\n>>> Kevin, you seem to think that normalization is somehow forced on you\n>>> by\n>>> the \"text-as-codepoints\" decision, and that is SIMPLY NOT TRUE.\n>>> Normalization is a totally separate decision, and it's a STUPID one,\n>>> because it breaks so many of the _nice_ properties of using UTF-8.\n>>\n>> I'm not saying it's forced on you, I'm saying when you treat  \n>> filenames\n>> as text,\n>\n> to treat as text could mean different for different people. Some\n> may prefer to fi and fi_ligature to be treated as same in some\n> context.\n\nThose people can use NFKC/NFKD (compatibility equivalence). As I've  \nsaid before, I'm talking about canonical equivalence, because that  \ndoesn't lose information like compatibility equivalence does (ex. the  \nfi ligature gets turned into fi in compatibility equivalence, but not  \ncanonical equivalence).\n\n>> it DOESN'T MATTER if the string gets normalized. As long as\n>> the string remains equivalent,\n>\n> As matter of fact it does, otherwise characters would be the\n> same and we would not have this conversation at all. String\n> can be equivalent and not equivalent at the time, because there\n> are different equivalent relations. Finally, what HFS+ does\n> is even not normalization. In the technote, Apple explains\n> that they decompose some characters but not others for better\n> compatibility. So, you see, there is a PROBLEM here.\n\nAgain, I've specified many times that I'm talking about canonical  \nequivalence.\n\nAnd yes, HFS+ does normalization, it just doesn't use NFD. It uses a  \ncustom variant. I fail to see how this is a problem.\n\n>> Alright, fine. I'm not saying HFS+ is right in storing the normalized\n>> version, but I do believe the authors of HFS+ must have had a reason\n>> to do that,\n>\n> I don't say they do that without *any* reason, but I suppose all\n> Apple developers in the Copland project had some reasons for they\n> did, but the outcome was not very good...\n\nStupid engineers don't get to work on developing new filesystems. And  \nCopland didn't fail because of stupid engineers anyway. If I had to  \nblame someone, I'd blame management.\n\n>> The only information you lose when doing canonical normalization is\n>> what the original byte sequence was.\n>\n> Not true. You lose the original sequence of *characters*.\n\nWhich is only a problem if you care about the byte sequence, which is  \nkinda the whole point of my argument.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66247","messageId":"8563xm6g38.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"46a038f90801211306g3dd9a167wb74d06e444b18b93@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-21T21:09:31Z","receivedAt":"2008-01-21T21:09:31Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n\n> (slightly offtopic) are you praising UTF-8 as storage format (for disk\n> and network) or in general? UTF-8-aware string ops like counting\n> characters seem to me a horrendous thing at the ASM level.\n\nHuh?  Why?  Just count all characters in the range 00-bf.  That's the\nexact character count of utf-8 characters.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66201","messageId":"46a038f90801211317v4902ffd3ic8ccc35f8df72bd9@mail.gmail.com","threadId":"11645","inReplyTo":"7EB98659-4036-45DA-BD50-42CB23ED517A@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-21T21:17:45Z","receivedAt":"2008-01-21T21:17:45Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 9:53 AM, Kevin Ballard <kevin@sb.org> wrote:\n> On Jan 21, 2008, at 3:33 PM, Linus Torvalds wrote:\n> > Umm. What's this inability to see that data is data is data?\n>\n> I'm not sure what you mean. I stated a fact - at least on OS X, the\n> filename does not contribute to the listed filesize, so changing the\n> encoding of the filename doesn't change the filesize. This isn't a\n> philosophical point, it's a factual statement.\n\nKevin,\n\nas you might know, Linus' \"other hobby\" is to write kernels ;-) From\ntaht POV, a filename is as much data as the data in the file. Doing\nodd things like sorting it, searching through it, etc, is all work for\ncode higher in the stack that is free to mangle the data in any way it\nwants, including creating nice case-insensitive indexes, and\nwho-knows-what for ideogram-based languages. In contrast, the core OS\ntreats user data a sacred stuff, and I'm thankful it does.\n\nAnd from a kernel/filesystem POV, a directory is also a file. So if a\nfilename has a different number of octets, the directory will be\ndifferent.\n\nFor all the searching and matching, it really makes sense to have\nsomething like locate or SpotLight or whatever to index user files\nthat should be easy to find and match, because all the locale rules\nfor matching are hideously expensive to apply. Even today, most UTF-8\naware (and supposedly collation-smart) applications have trouble\nmatching MARTÍN when asked for martín in a case-insensitive search.\nThat pesky latin í trips them up everytime.\n\ncheers,\n\n\nmartin\n"},{"id":"66200","messageId":"20080121211802.GH29792@mit.edu","threadId":"11645","inReplyTo":"6E303071-82A4-4D69-AA0C-EC41168B9AFE@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-21T21:18:02Z","receivedAt":"2008-01-21T21:18:02Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jan 21, 2008 at 03:58:03PM -0500, Kevin Ballard wrote:\n> You're making the huge assumption that the HFS+ normalization algorithms \n> will change. As the technote states:\n>\n> \"Platform algorithms tend to evolve with the Unicode standard. The HFS Plus \n> algorithms cannot evolve because such evolution would invalidate existing \n> HFS Plus volumes.\"\n\nGreat, so even worse.  Does the tech note then specify exactly what\nversion of Unicode HFS+ is using to do its \"normalization\"?  Or\nexactly what characters it will normalize?  After all, Unicode has\nadded all sorts of characters since 1998, and I'm sure some of them\nwere combining characters.\n\nAnd you *really* want to continue argue that a sane thing for a\ncross-platform system to do is to pervert its hash algorithm to take\ninto account *one* particular OS that happened to freeze a\nnormalization algorithm at some arbitrary point in time, approximately\nnine years ago?  Talk about the tail wagging the dog!!  Especially\nwhen you can't even justify why it was done nine years ago!\n\n> It must have bought somebody something, or they never would have done it.\n\nYour faith in the HFS+ designers is touching.\n\n> I have no idea why HFS+ stores filenames in a normalized form, and further \n> I am smart enough to know that speculating is completely pointless. I \n> assume the authors had a good reason (which should be a safe assumption, \n> filesystem authors are a smart bunch). The reason may not be valid anymore, \n> but if it was valid back in 1998, then I can accept it without complaining.\n\nWell, I *AM* a filesystem designer (ext2/ext3/ext4), and well before\n1998, I knew that trying to do anything with Unicode normalization was\na fool's errand.  So if you're going to blindly trust filesystme\ndesigners (not something I would recommend, actually :-), trust me.\nWhat HFS+ is doing is dumb, dumb, dumb.\n\nAnd even if *you* can accept it, why should the git designers pervert\nany core part of git's design to support this behaviour?  Especially\nif it's legacy behaviour which will hopefully be going away, say when\nMacOS adopts ZFS --- there's an opportunity for them to start afresh,\nand not make the same mistakes they made nine years ago!\n\nSo why don't you suggest some kind of sane fix in the Mac specific\ncode that doesn't impact any core part of git, such as its hash\nalgorithm?  It would be far more productive than trying to defend a\nbad design decision made nine years ago....   :-}\n\n\t\t\t\t\t- Ted\n"},{"id":"66204","messageId":"E46A8ED2-7E95-41CF-A5A7-2CB9BC18F95D@sb.org","threadId":"11645","inReplyTo":"46a038f90801211317v4902ffd3ic8ccc35f8df72bd9@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T21:28:41Z","receivedAt":"2008-01-21T21:28:41Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 4:17 PM, Martin Langhoff wrote:\n\n> On Jan 22, 2008 9:53 AM, Kevin Ballard <kevin@sb.org> wrote:\n>> On Jan 21, 2008, at 3:33 PM, Linus Torvalds wrote:\n>>> Umm. What's this inability to see that data is data is data?\n>>\n>> I'm not sure what you mean. I stated a fact - at least on OS X, the\n>> filename does not contribute to the listed filesize, so changing the\n>> encoding of the filename doesn't change the filesize. This isn't a\n>> philosophical point, it's a factual statement.\n>\n> Kevin,\n>\n> as you might know, Linus' \"other hobby\" is to write kernels ;-) From\n> taht POV, a filename is as much data as the data in the file. Doing\n> odd things like sorting it, searching through it, etc, is all work for\n> code higher in the stack that is free to mangle the data in any way it\n> wants, including creating nice case-insensitive indexes, and\n> who-knows-what for ideogram-based languages. In contrast, the core OS\n> treats user data a sacred stuff, and I'm thankful it does.\n\nThat's certainly a reasonable POV. However, it's not the only one. As  \nevidenced by the Mac, treating filenames as strings rather than bytes  \nis a viable alternative POV - you can't argue that it doesn't work,  \nbecause OS X proves it does.\n\nHowever, it is a trade-off.\n\n> And from a kernel/filesystem POV, a directory is also a file. So if a\n> filename has a different number of octets, the directory will be\n> different.\n\nSure, that makes sense. That's why, if you are going to mangle  \nfilenames, you need to pick a stable form to always use, which HFS+  \ndoes.\n\n> For all the searching and matching, it really makes sense to have\n> something like locate or SpotLight or whatever to index user files\n> that should be easy to find and match, because all the locale rules\n> for matching are hideously expensive to apply. Even today, most UTF-8\n> aware (and supposedly collation-smart) applications have trouble\n> matching MARTÍN when asked for martín in a case-insensitive search.\n> That pesky latin í trips them up everytime.\n\n\nPerhaps you should try OS X. Every single Cocoa app should do the  \nsearch properly. In fact, I just checked using 3 different text  \nengines (WebKit, Cocoa's text engine, and ATSUI) and all 3 did the  \ncase-insensitive search properly. That said, this isn't particularly  \nrelevant.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66206","messageId":"alpine.LFD.1.00.0801211323120.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"7EB98659-4036-45DA-BD50-42CB23ED517A@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-21T21:33:52Z","receivedAt":"2008-01-21T21:33:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n> \n> I'm not sure what you mean. I stated a fact - at least on OS X, the filename\n> does not contribute to the listed filesize, so changing the encoding of the\n> filename doesn't change the filesize. This isn't a philosophical point, it's a\n> factual statement.\n\nAnd my point was that your *whole* argument boils down to \"normalization \nis invisible\".\n\nWhen it isn't. It's not invisible for filenames, it's not invisible for \nfile contents.\n\nYou're trying to claim that normalization cannot matter. I'm just pointing \nout that it sure as hell can. Exactly because lots of things don't \nactually look at data other than as just a Unicode string. They do look at \nthe raw format.\n\nAnd that's true both of file contents and file names.\n\n> I don't, but I do think this discussion revolves around filenames, therefore\n> it should not surprise you when I talk about filenames.\n\nI'm surprised that you make generalized sweeping statements about how it's \nok to normalize because normalization is \"invisible\", and then when I \npoint out that that isn't true, you try to limit it.\n\nAnd no, that normalization is not invisible EVEN IN FILENAMES. If it was, \ngit wouldn't ever have noticed it, would it?\n\n> > And git tries to be a general data tool, not a Unicode-specific one.\n> \n> Yes, I realize that. See my previous message about discussing ideal vs\n> practicality.\n\nI don't know which argument you're talking about. Git (and, btw, Linux) \ndoes the \"ideal\" thing (don't screw up peoples data), and it turns out to \nbe the \"practical\" thing too (it can handle a wider range of cases than OS \nX can).\n\nSo no, this is not \"ideal\" vs \"practical\". They aren't in any conflict \nhere.\n\n> I could argue against this, but frankly, I'm really tired of arguing this same\n> point. I suggest we simply agree to disagree, and move on to actually fixing\n> the problem.\n\n.. and people have even suggested how. Hide the idiotic OS X choices by \nmaking a OS X-specific wrapper around readdir() that turns it into NFC.\n\nThat's just about the best we can do. We can't *fix* the thing that OS X \nloses information, but a least we can then show the lost information in \nthe same form it _probably_ was in originally.\n\nBut no, it won't \"fix\" git on OS X. \n\n\t\t\tLinus\n"},{"id":"66207","messageId":"alpine.LFD.1.00.0801211334540.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"46a038f90801211306g3dd9a167wb74d06e444b18b93@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-21T21:42:47Z","receivedAt":"2008-01-21T21:42:47Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 22 Jan 2008, Martin Langhoff wrote:\n> \n> (slightly offtopic) are you praising UTF-8 as storage format (for disk\n> and network) or in general? UTF-8-aware string ops like counting\n> characters seem to me a horrendous thing at the ASM level.\n\nI'm praising UTF-8 (without normalization) as a wonderful format where you \ncan do 99.9% of everything without ever caring about all the expensive \nstuff.\n\nBut in order to do that, you really need to avoid normalization, and you \nalso need to accept mis-formed UTF-8 strings (because even if it is real \nUTF-8, the string may actually be just a fragment of some larger string).\n\nOnce you do that (and _only_ if you do that), then UTF-8 is actually a \nwonderful thing. You can consider it to be a traditional \"everything is a \nstream of bytes\", and everything that only cares about a stream of byte \nwill work wonderfully well.\n\nAnd then, the (actually relatively few) things that want to do things like \nshow things on the screen, or check for equivalence, or worry about width \nof the characters, *those* can still do so. \n\nSo the beauty of UTF-8 is that you can switch between thinking of it like \njust a binary blob and thinking of it like text, and everythign works \n(including the traditional C null-termination).\n\nAnd yes, that was obviously the explicit design goal. It's a good thing.\n\n> More on topic, I suspect Kevin's experience is more on end-user apps,\n> where input sanitization and even canonicalisation are common\n> practice.\n\nSure. And I'm not arguing against them. Knowing the rules for combining \ncharacters is really important for input and output. \n\n> At least in Moodle we store *exactly*  what the user POSTed and\n> cleanup^Wcorrupt it when displaying it, so that if it does happen that\n> the cleanup was buggy, we never corrupted the data.\n\nAbsolutely. It's what the kernel does, and I think that's what perl does \ntoo for their \"strings\". It works really well. It also allows you to \nhandle binary data (ie data that *really* isn't text) with shared routines \netc etc.\n\nAnd that's the beauty of non-normalized (and possibly badly formed) UTF-8.\n\n\t\tLinus\n"},{"id":"66208","messageId":"46a038f90801211343q4a36195dp319beb3020e57600@mail.gmail.com","threadId":"11645","inReplyTo":"E46A8ED2-7E95-41CF-A5A7-2CB9BC18F95D@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-21T21:43:16Z","receivedAt":"2008-01-21T21:43:16Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 10:28 AM, Kevin Ballard <kevin@sb.org> wrote:\n> That's certainly a reasonable POV. However, it's not the only one. As\n> evidenced by the Mac, treating filenames as strings rather than bytes\n> is a viable alternative POV - you can't argue that it doesn't work,\n> because OS X proves it does.\n\nWith its own slew of bugs. See Ted's reply earlier for a mouthful of\nwoe in HFS+ that is not easy to workaround.\n\n> Perhaps you should try OS X. Every single Cocoa app should do the\n\nOSX has given me enough grief with other filesystem and general OS\nproblems that I have definitely abandoned it after 2 years trying to\nuse it part-time. It has been back to linux for me ;-)\n\ncheers,\n\n\n\nmartin\n"},{"id":"66209","messageId":"45C7CC4A-155F-4FE4-B741-8EE6CF7D3700@sb.org","threadId":"11645","inReplyTo":"20080121211802.GH29792@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T21:43:26Z","receivedAt":"2008-01-21T21:43:26Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 4:18 PM, Theodore Tso wrote:\n\n> On Mon, Jan 21, 2008 at 03:58:03PM -0500, Kevin Ballard wrote:\n>> You're making the huge assumption that the HFS+ normalization  \n>> algorithms\n>> will change. As the technote states:\n>>\n>> \"Platform algorithms tend to evolve with the Unicode standard. The  \n>> HFS Plus\n>> algorithms cannot evolve because such evolution would invalidate  \n>> existing\n>> HFS Plus volumes.\"\n>\n> Great, so even worse.  Does the tech note then specify exactly what\n> version of Unicode HFS+ is using to do its \"normalization\"?  Or\n> exactly what characters it will normalize?  After all, Unicode has\n> added all sorts of characters since 1998, and I'm sure some of them\n> were combining characters.\n>\n> And you *really* want to continue argue that a sane thing for a\n> cross-platform system to do is to pervert its hash algorithm to take\n> into account *one* particular OS that happened to freeze a\n> normalization algorithm at some arbitrary point in time, approximately\n> nine years ago?  Talk about the tail wagging the dog!!  Especially\n> when you can't even justify why it was done nine years ago!\n\nI suggest you go back and read the emails where I specifically stated  \nthat I'm *not* suggesting this.\n\n>> It must have bought somebody something, or they never would have  \n>> done it.\n>\n> Your faith in the HFS+ designers is touching.\n\nAnd your arrogance is troubling. Do you really believe you are so  \nsmart you can claim the HFS+ designers had no reason for this decision?\n\n>> I have no idea why HFS+ stores filenames in a normalized form, and  \n>> further\n>> I am smart enough to know that speculating is completely pointless. I\n>> assume the authors had a good reason (which should be a safe  \n>> assumption,\n>> filesystem authors are a smart bunch). The reason may not be valid  \n>> anymore,\n>> but if it was valid back in 1998, then I can accept it without  \n>> complaining.\n>\n> Well, I *AM* a filesystem designer (ext2/ext3/ext4), and well before\n> 1998, I knew that trying to do anything with Unicode normalization was\n> a fool's errand.  So if you're going to blindly trust filesystme\n> designers (not something I would recommend, actually :-), trust me.\n> What HFS+ is doing is dumb, dumb, dumb.\n\nAgain, I'm not saying that they necessarily did the \"correct\" thing,  \nas I can't evaluate that without knowing their reason. I'm just saying  \nthere must have been a reason.\n\n> And even if *you* can accept it, why should the git designers pervert\n> any core part of git's design to support this behaviour?  Especially\n> if it's legacy behaviour which will hopefully be going away, say when\n> MacOS adopts ZFS --- there's an opportunity for them to start afresh,\n> and not make the same mistakes they made nine years ago!\n\nAnd why do you believe MacOS is going to adopt ZFS? Sure, they might,  \nbut assuming stuff about the future is just as bad as assuming stuff  \nabout the past. And git should \"pervert\" itself because of the simple  \nfact that git has a problem on HFS+. Keeping your code \"pure\" is all  \nwell and good, except it's not particularly practical. If the git  \nproject has any interest in being a viable system on OS X, it really  \nshould behave properly. I'm sure you have various \"perversions\" for  \nother cases.\n\n> So why don't you suggest some kind of sane fix in the Mac specific\n> code that doesn't impact any core part of git, such as its hash\n> algorithm?  It would be far more productive than trying to defend a\n> bad design decision made nine years ago....   :-}\n\nHow many times must I say I never suggested actually changing git's  \nhashing algorithm? And if you want me to suggest a fix to git that  \nworks, first you have to wait for me to learn how git's internals  \nwork, and frankly, I have too much work on my plate right now to  \ndevote the time necessary to learning git's internals well enough to  \nfix this problem.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66211","messageId":"46a038f90801211349mfb57c0an9416832c2967c172@mail.gmail.com","threadId":"11645","inReplyTo":"45C7CC4A-155F-4FE4-B741-8EE6CF7D3700@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-21T21:49:14Z","receivedAt":"2008-01-21T21:49:14Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 10:43 AM, Kevin Ballard <kevin@sb.org> wrote:\n> How many times must I say I never suggested actually changing git's\n> hashing algorithm? And if you want me to suggest a fix to git that\n> works, first you have to wait for me to learn how git's internals\n> work, and frankly, I have too much work on my plate right now to\n> devote the time necessary to learning git's internals well enough to\n> fix this problem.\n\nLOL! Spare us the flamefesting and you will have plenty of time for\nlearning git internals. You might even learn something.\n\ncheers,\n\n\nmartin\n"},{"id":"66212","messageId":"C373E12A-2AC4-4961-833A-7D51584143C9@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211323120.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T21:49:46Z","receivedAt":"2008-01-21T21:49:46Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 4:33 PM, Linus Torvalds wrote:\n\n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n>>\n>> I'm not sure what you mean. I stated a fact - at least on OS X, the  \n>> filename\n>> does not contribute to the listed filesize, so changing the  \n>> encoding of the\n>> filename doesn't change the filesize. This isn't a philosophical  \n>> point, it's a\n>> factual statement.\n>\n> And my point was that your *whole* argument boils down to  \n> \"normalization\n> is invisible\".\n>\n> When it isn't. It's not invisible for filenames, it's not invisible  \n> for\n> file contents.\n>\n> You're trying to claim that normalization cannot matter. I'm just  \n> pointing\n> out that it sure as hell can. Exactly because lots of things don't\n> actually look at data other than as just a Unicode string. They do  \n> look at\n> the raw format.\n>\n> And that's true both of file contents and file names.\n>\n>> I don't, but I do think this discussion revolves around filenames,  \n>> therefore\n>> it should not surprise you when I talk about filenames.\n>\n> I'm surprised that you make generalized sweeping statements about  \n> how it's\n> ok to normalize because normalization is \"invisible\", and then when I\n> point out that that isn't true, you try to limit it.\n>\n> And no, that normalization is not invisible EVEN IN FILENAMES. If it  \n> was,\n> git wouldn't ever have noticed it, would it?\n\nI'm really surprised that, after all of this, you're still horribly  \nmisunderstanding my argument. I never said it was invisible. NEVER.\n\nI'm also surprised that you seem to care more about this argument then  \nmy offer to stop arguing and work towards fixing the problem.\n\n>>> And git tries to be a general data tool, not a Unicode-specific one.\n>>\n>> Yes, I realize that. See my previous message about discussing ideal  \n>> vs\n>> practicality.\n>\n> I don't know which argument you're talking about. Git (and, btw,  \n> Linux)\n> does the \"ideal\" thing (don't screw up peoples data), and it turns  \n> out to\n> be the \"practical\" thing too (it can handle a wider range of cases  \n> than OS\n> X can).\n>\n> So no, this is not \"ideal\" vs \"practical\". They aren't in any conflict\n> here.\n\nYou misunderstand my point. In a previous email I specifically used  \nthe words \"ideal\" and \"practical\" to describe arguments, which is what  \nI was referring to here.\n\n>> I could argue against this, but frankly, I'm really tired of  \n>> arguing this same\n>> point. I suggest we simply agree to disagree, and move on to  \n>> actually fixing\n>> the problem.\n>\n> .. and people have even suggested how. Hide the idiotic OS X choices  \n> by\n> making a OS X-specific wrapper around readdir() that turns it into  \n> NFC.\n\nAnd I've responded to that suggestion, multiple times, saying that  \nthis doesn't actually fix the problem, it only hides it.\n\n> That's just about the best we can do. We can't *fix* the thing that  \n> OS X\n> loses information, but a least we can then show the lost information  \n> in\n> the same form it _probably_ was in originally.\n>\n> But no, it won't \"fix\" git on OS X.\n\nQuite a while ago it was suggested that git uses a table that maps the  \noriginal byte sequence as seen in the index to the form returned by  \nreaddir(). So far this has sounded like the best solution, but as I've  \nsaid before I don't know git's internals enough (or, really, at all)  \nto be able to work on this myself.\n\nThis solution should only \"lose\" information in the case where the  \nindex has 2 filenames that HFS+ treats as a single filename.\n\nIs there some reason this won't work?\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66213","messageId":"BA2377B5-BE7E-40F2-9C3C-679663A966A4@sb.org","threadId":"11645","inReplyTo":"46a038f90801211349mfb57c0an9416832c2967c172@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T21:57:08Z","receivedAt":"2008-01-21T21:57:08Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"Again with the bouncing. I gotta figure out how to fix this.\nOriginal CC list: martin.langhoff@gmail.com, tytso@mit.edu, torvalds@linux-foundation.org \n, peter@softwolves.pp.se, mjscod@web.de, melo@simplicidade.org\n\nOn Jan 21, 2008, at 4:49 PM, Martin Langhoff wrote:\n\n> On Jan 22, 2008 10:43 AM, Kevin Ballard <kevin@sb.org> wrote:\n>> How many times must I say I never suggested actually changing git's\n>> hashing algorithm? And if you want me to suggest a fix to git that\n>> works, first you have to wait for me to learn how git's internals\n>> work, and frankly, I have too much work on my plate right now to\n>> devote the time necessary to learning git's internals well enough to\n>> fix this problem.\n>\n> LOL! Spare us the flamefesting and you will have plenty of time for\n> learning git internals. You might even learn something.\n\nAh, so I'm flaming while you are providing a well-reasoned and  \narticulate argument? Glad to know the difference.\n\nIn any case, you should be very familiar with the fact that writing  \nemails and learning code are two vastly different activities that  \nrequire vastly different amounts of concentration and time. I'm  \nresponding to email while doing other things - were I to replace the  \ntime spent writing email with learning git's internals, I would be  \npulled away so frequently that I would end up not having learned  \nanything. Therefore, I write emails now, and I leave learning git's  \ninternals until later, when I have the undisturbed time to devote.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66217","messageId":"alpine.LFD.1.00.0801211407130.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"C373E12A-2AC4-4961-833A-7D51584143C9@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-21T22:34:32Z","receivedAt":"2008-01-21T22:34:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n> \n> I'm really surprised that, after all of this, you're still horribly\n> misunderstanding my argument. I never said it was invisible. NEVER.\n\nYou said it was invisible when you treat things \"as text\". Here's the \nquote:\n\n\t.. when you treat filenames as text, it DOESN'T MATTER if the \n \tstring gets normalized ..\n\nWithout ever apparently realizing that \"as text\" is part of the problem in \nitself. What is \"text\" to one person is gibberish to another.\n\nIn particular, the biggest reason to not normalize is that you don't know \nit's text or Unicode in the first place. Which is why git doesn't do it.\n\nAnd no, even with filenames you don't know that they are \"text\". People \nencode stuff in them. And people don't always use UTF-8. \n\nOf course, you could ask everybody to create OS X-only programs that know \nthat under OS X, you only have a subset of filenames. If so, you're \ncomplaining about the wrong tool. Especially when the whole point of the \ntool was to be distributed (not to mention coming from an environment that \nsimply doesn't have the same silly limitations OS X has).\n\nSo here's a few clues:\n\n - \"as text\" isn't \"as unicode\": it may well be Latin1 or EUC-JP or\n   something. Yes, it's still used. Git doesn't care, and very consciously \n   has avoided forcing character sets, even if the *default* (and notice \n   how it's overridable) commit message encoding may be utf-8.\n\n - In fact, even in unicode, the difference between \"identical\" and \n   \"equivalent\" strings exists, and even in the standard, unicode \n   strings are very much defined to be arbitrary codepoint sequences, not \n   normalized.\n\nSo even for the very specific case of unicode text, it's simply not true \nthat \"it doesn't matter if the string gets normalized\". The unicode spec \nitself talks about cases where even canonical normalization makes a \ndifference.\n\nSearch for this quote:\n\n  \"Not all processes are required to respect canonical equivalence. For \n   example:\n\n    * A function that collects a set of the General_Category values \n      present in a string will and should produce a different value for \n      <angstrom sign, semicolon> than for <A, combining ring above, greek \n      question mark>, even though they are canonically equivalent.\n    * A function that does a binary comparison of strings will also find \n      these two sequences different.\"\n\nand notice that first case. Even things that are *very*much* aware of \nUnicode text do actually have cases where canonical equivalence doesn't \nmean crud.\n\n\t\tLinus\n"},{"id":"66246","messageId":"85zluy4xf0.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"45C7CC4A-155F-4FE4-B741-8EE6CF7D3700@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-21T22:38:11Z","receivedAt":"2008-01-21T22:38:11Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Kevin Ballard <kevin@sb.org> writes:\n\n> On Jan 21, 2008, at 4:18 PM, Theodore Tso wrote:\n>\n>> Your faith in the HFS+ designers is touching.\n>\n> And your arrogance is troubling. Do you really believe you are so\n> smart you can claim the HFS+ designers had no reason for this\n> decision?\n\nNo reason?  Don't say where he says that.  No sane reason?  Certainly.\nIf the visibility of the upsides is not in the same order of magnitude\nas that of the downsides (and your \"I trust they must have had good\nreason\" is implicating exactly that), then yes, this appears like a\nmisdesign, however well-intended.  Because its cleverness hinges on an\nwhat amounts to an arbitrary historic point of stability with only\nfleeting convenience.\n\nIt reminds me of the self-defeatingproblem haunting the MIPS\n(microprocessor without interlocked pipeline stages) architecture: for\npipelined processors, one has to add logic that prevents one command\nfrom working before the results from other commands arrive.  Now the\ningenious idea of the MIPS architecture was to move that logic into the\ncompiler instead of the hardware.  But then the implications of that\nidea got intermingled with binary compatibility and the result was that\nthe advantages lasted for one processor generation, and afterwards, the\nm-stage pipelines needed logic that simulated the n-stage pipeline of\nthe first MIPS processor rather then a comparatively simple 1-stage\npipeline of a conceptual sequential processor.  Rendering the whole\noriginal idea completely absurd and requiring rather more complicated\nrather than simpler hardware as originally envisioned.\n\nThe road to hell is paved with good intentions.\n\n>>> The reason may not be valid anymore, but if it was valid back in\n>>> 1998, then I can accept it without complaining.\n\nThere is no shortage in short-sighted decisions to repeat.  Some\npolitical parties bank on it.\n\n> Again, I'm not saying that they necessarily did the \"correct\" thing,\n> as I can't evaluate that without knowing their reason. I'm just saying\n> there must have been a reason.\n\nJumping to blind faith-based conclusions is never a good move.  You\ndon't end up improving the work of your predecessors that way.\n\n> And why do you believe MacOS is going to adopt ZFS? Sure, they might,\n> but assuming stuff about the future is just as bad as assuming stuff\n> about the past. And git should \"pervert\" itself because of the simple\n> fact that git has a problem on HFS+. Keeping your code \"pure\" is all\n> well and good, except it's not particularly practical. If the git\n> project has any interest in being a viable system on OS X, it really\n> should behave properly.\n\nIf OS X has any interest in being a viable system, perios, it really\nshould behave properly.\n\n> How many times must I say I never suggested actually changing git's\n> hashing algorithm? And if you want me to suggest a fix to git that\n> works, first you have to wait for me to learn how git's internals\n> work, and frankly, I have too much work on my plate right now to\n> devote the time necessary to learning git's internals well enough to\n> fix this problem.\n\nThen please understand that you have too much work on your plate right\nnow to devote the time necessary to provide any constructive criticism.\nA smart person in this situation would shut up until he has the time.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66218","messageId":"20080121224131.GD14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"DEC058ED-EBF0-4A1E-BF7B-448B16DBBD6E@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-21T22:41:31Z","receivedAt":"2008-01-21T22:41:31Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Jan 21, 2008 at 04:07:27PM -0500, Kevin Ballard wrote:\n> \n> Again, I've specified many times that I'm talking about canonical  \n> equivalence.\n> \n> And yes, HFS+ does normalization, it just doesn't use NFD. It uses a  \n> custom variant. I fail to see how this is a problem.\n\nIf you think that HFS+ does normalization then you apparently have no\nidea of what the term \"normalization\" means. Have you? But if you\ndon't know what is \"normalization\" then you cannot really know what\ncanonical equivalence means.\n\n> >\n> >I don't say they do that without *any* reason, but I suppose all\n> >Apple developers in the Copland project had some reasons for they\n> >did, but the outcome was not very good...\n> \n> Stupid engineers don't get to work on developing new filesystems.\n\nAssigning someone to work on a new filesystem does not make him\nsuddenly smart. As to that stupid engineers don't get to work,\nit is like saying there is no stupid engineers at all. There are\nplenty evidence to contrary. And when management is disastrous\nthen most idiots with big mouth and little capacity to produce\nany useful does get assignment to develop new features, while\nthose who can actually solve problems are assigned to fix the\nnext build, because the only thing that this management worries\nabout how to survive another year or another months...\n\n> And  \n> Copland didn't fail because of stupid engineers anyway. If I had to  \n> blame someone, I'd blame management.\n\nBut if the code was so good then why was most of that code thrown away\nlater when management was changed? Still bad management?\n\n> \n> >>The only information you lose when doing canonical normalization is\n> >>what the original byte sequence was.\n> >\n> >Not true. You lose the original sequence of *characters*.\n> \n> Which is only a problem if you care about the byte sequence, which is  \n> kinda the whole point of my argument.\n\nByte sequences are not an issue here. If the filesystem used UTF-16 to\nstore filenames, that would NOT cause this problem, because characters\nwould be the same even though bytes stored on the disk were different.\nSo, what you actually lose here is the original sequence of *characters*.\n\nDmitry\n"},{"id":"66220","messageId":"46a038f90801211445g2fa65bd6ida8b59882d60bd78@mail.gmail.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211334540.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-21T22:45:17Z","receivedAt":"2008-01-21T22:45:17Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 10:42 AM, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> I'm praising UTF-8 (without normalization) as a wonderful format where you\n> can do 99.9% of everything without ever caring about all the expensive\n> stuff.\n\n*thanks* for these notes. Very useful, and...\n\n...\n> And then, the (actually relatively few) things that want to do things like\n> show things on the screen, or check for equivalence, or worry about width\n> of the characters, *those* can still do so.\n\nI find the above amusing -- different worlds we live in. Programming\nwebapps means that 90% of the code deals with a bit of metaprogramming\n(with lots of string manipulation) to talk SQL to a backend, and then\ndoing lots of string manipulation on the data the DB returns, which\nends up in humongous strings of goop otherwise known as HTML+CSS+JS.\nAfter waiting for the DB to return data, over 50% of cpu time is spent\nin regexes, concatenations, counting words, array ops, etc. So it is\npretty significant.\n\nSo now I have to worry about cost and correctness of stuff that I took\nfor granted in the pre-unicode days - strtolower() can be quite\nexpensive and... buggy! But that's mainly due to Unicode, not UTF8. I\nthink the only slowdown I can pin on UTF-8 is in counting chars, and\nprobably slower regexes. Not that I deal with the C implementation of\nany of this stuff -- and so happy about it! ;-)\n\n</offtopic>\n\n(...)\n\n> And that's the beauty of non-normalized (and possibly badly formed) UTF-8.\n\nI had a few issues with Perl v5.6's utf-8 handling that wasn't binary\nsafe (fread() to a fixed-length buffer would break the input if a\nunicode char landed across the boundary - ouch!) -- made me think that\nyou couldn't do this in binary safe ways. So I tend to tell Perl to\ntreatfiles as binary, and switch to utf-8 in specially chosen spots. I\nsuspect that 5.8 is a bit saner about this, but I'm not taking\nchances.\n\ncheers,\n\n\nmartin\n"},{"id":"66221","messageId":"0CA4DF3F-1B64-4F62-8794-6F82C21BD068@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211407130.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T22:46:27Z","receivedAt":"2008-01-21T22:46:27Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 5:34 PM, Linus Torvalds wrote:\n\n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n>>\n>> I'm really surprised that, after all of this, you're still horribly\n>> misunderstanding my argument. I never said it was invisible. NEVER.\n>\n> You said it was invisible when you treat things \"as text\". Here's the\n> quote:\n>\n> \t.. when you treat filenames as text, it DOESN'T MATTER if the\n> \tstring gets normalized ..\n>\n> Without ever apparently realizing that \"as text\" is part of the  \n> problem in\n> itself. What is \"text\" to one person is gibberish to another.\n\nWhich is actually a good argument as to why filenames should be  \nenforced as UTF-8.\n\n> In particular, the biggest reason to not normalize is that you don't  \n> know\n> it's text or Unicode in the first place. Which is why git doesn't do  \n> it.\n\nSure, I understand why git doesn't do it. I'm saying in a system which  \nuses unicode top-to-bottom, which you can create if you're using HFS+  \nonly, can do it. On HFS+ you know the filename is unicode.\n\n> And no, even with filenames you don't know that they are \"text\".  \n> People\n> encode stuff in them. And people don't always use UTF-8.\n\nAgain, I was talking about a system that used unicode top-to-bottom.  \nOn HFS+ you have to use UTF-8 for your filename or it simply won't work.\n\n> Of course, you could ask everybody to create OS X-only programs that  \n> know\n> that under OS X, you only have a subset of filenames. If so, you're\n> complaining about the wrong tool. Especially when the whole point of  \n> the\n> tool was to be distributed (not to mention coming from an  \n> environment that\n> simply doesn't have the same silly limitations OS X has).\n>\n> So here's a few clues:\n>\n> - \"as text\" isn't \"as unicode\": it may well be Latin1 or EUC-JP or\n>   something. Yes, it's still used. Git doesn't care, and very  \n> consciously\n>   has avoided forcing character sets, even if the *default* (and  \n> notice\n>   how it's overridable) commit message encoding may be utf-8.\n>\n> - In fact, even in unicode, the difference between \"identical\" and\n>   \"equivalent\" strings exists, and even in the standard, unicode\n>   strings are very much defined to be arbitrary codepoint sequences,  \n> not\n>   normalized.\n>\n> So even for the very specific case of unicode text, it's simply not  \n> true\n> that \"it doesn't matter if the string gets normalized\". The unicode  \n> spec\n> itself talks about cases where even canonical normalization makes a\n> difference.\n>\n> Search for this quote:\n>\n>  \"Not all processes are required to respect canonical equivalence. For\n>   example:\n>\n>    * A function that collects a set of the General_Category values\n>      present in a string will and should produce a different value for\n>      <angstrom sign, semicolon> than for <A, combining ring above,  \n> greek\n>      question mark>, even though they are canonically equivalent.\n>    * A function that does a binary comparison of strings will also  \n> find\n>      these two sequences different.\"\n>\n> and notice that first case. Even things that are *very*much* aware of\n> Unicode text do actually have cases where canonical equivalence  \n> doesn't\n> mean crud.\n\nI find it amusing that you keep arguing against having git treat  \nfilenames as unicode when, if you had actually taken my advice and  \nread my previous email talking about \"ideal\" vs \"practical\", you'd  \nrealize that I was not suggesting git should. I was simply describing  \nwhy having the filesystem specifically treat filenames as utf-8 isn't  \na problem when the entire system is unicode-aware, and thus showing  \nhow the problems that are cropping up in git aren't because the  \nfilesystem treats filenames as unicode, but rather because the  \nfilesystem treats filenames differently than other filesystems. In  \nother words, I was trying to illustrate that HFS+ isn't wrong, it's  \njust different, and the difference is causing the problem.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66223","messageId":"DC78D5CB-18FF-4504-BD8B-985D8B202817@sb.org","threadId":"11645","inReplyTo":"20080121224131.GD14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T22:53:50Z","receivedAt":"2008-01-21T22:53:50Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 5:41 PM, Dmitry Potapov wrote:\n\n> On Mon, Jan 21, 2008 at 04:07:27PM -0500, Kevin Ballard wrote:\n>>\n>> Again, I've specified many times that I'm talking about canonical\n>> equivalence.\n>>\n>> And yes, HFS+ does normalization, it just doesn't use NFD. It uses a\n>> custom variant. I fail to see how this is a problem.\n>\n> If you think that HFS+ does normalization then you apparently have no\n> idea of what the term \"normalization\" means. Have you? But if you\n> don't know what is \"normalization\" then you cannot really know what\n> canonical equivalence means.\n\nI would go look up specifics to back me up, but my DNS is screwing up  \nright now so I can't access most of the internet. In any case, there  \nare 4 standard normalization forms - NFC, NFD, NFKC, NFKD. If there  \nare others, they aren't notable enough to be listed in the resource I  \nwas reading. HFS+ uses a variant on NFD - it's a well-defined variant,  \nand thus can safely be called its own normalization form. I fail to  \nsee how this means it's not \"normalization\".\n\n>>> I don't say they do that without *any* reason, but I suppose all\n>>> Apple developers in the Copland project had some reasons for they\n>>> did, but the outcome was not very good...\n>>\n>> Stupid engineers don't get to work on developing new filesystems.\n>\n> Assigning someone to work on a new filesystem does not make him\n> suddenly smart. As to that stupid engineers don't get to work,\n> it is like saying there is no stupid engineers at all. There are\n> plenty evidence to contrary. And when management is disastrous\n> then most idiots with big mouth and little capacity to produce\n> any useful does get assignment to develop new features, while\n> those who can actually solve problems are assigned to fix the\n> next build, because the only thing that this management worries\n> about how to survive another year or another months...\n\nI'm not talking about assigning engineers, I'm saying developing a new  \nfilesystem, especially one that's proven itself to be usable and  \nextendable for the last decade, is something that only smart engineers  \nwould be capable of doing.\n\n>> And\n>> Copland didn't fail because of stupid engineers anyway. If I had to\n>> blame someone, I'd blame management.\n>\n> But if the code was so good then why was most of that code thrown away\n> later when management was changed? Still bad management?\n\nYes. Even the best of engineers will produce crap code when overworked  \nand required to implement new features instead of fixing bugs and  \nstabilizing the system. Copland is well-known to have suffered from  \nfeaturitis, to the extent that it was practically impossible to test  \nin any sane fashion. Bad management can kill any project regardless of  \nhow good the engineers are.\n\n>>>> The only information you lose when doing canonical normalization is\n>>>> what the original byte sequence was.\n>>>\n>>> Not true. You lose the original sequence of *characters*.\n>>\n>> Which is only a problem if you care about the byte sequence, which is\n>> kinda the whole point of my argument.\n>\n> Byte sequences are not an issue here. If the filesystem used UTF-16 to\n> store filenames, that would NOT cause this problem, because characters\n> would be the same even though bytes stored on the disk were different.\n> So, what you actually lose here is the original sequence of  \n> *characters*.\n\nI've already talked about that, but you are apparently incapable of  \nunderstanding.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66224","messageId":"46a038f90801211456h4b16ff2cl4378df88023bbc34@mail.gmail.com","threadId":"11645","inReplyTo":"0CA4DF3F-1B64-4F62-8794-6F82C21BD068@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-21T22:56:24Z","receivedAt":"2008-01-21T22:56:24Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 11:46 AM, Kevin Ballard <kevin@sb.org> wrote:\n> Again, I was talking about a system that used unicode top-to-bottom.\n> On HFS+ you have to use UTF-8 for your filename or it simply won't work.\n\nHmmm. I m pretty sure HFS+ has a lot of problems if you run OSX as an\nNFS server with clients in different encodings. It would never work in\nreal life. The \"envelope\" OSs have to work in is hugely varied -- much\nmore so than any other apps. You should try writing one someday ;-)\n\n> other words, I was trying to illustrate that HFS+ isn't wrong, it's\n> just different, and the difference is causing the problem.\n\nDid you spot the rather nasty issues that Ted mentioned earlier in the\nthread? I would say HFS+ is a bit \"special\" rather than \"different\".\n\ncheers,\n\n\n\nm\n"},{"id":"66226","messageId":"20080121230053.GA317@mit.edu","threadId":"11645","inReplyTo":"0CA4DF3F-1B64-4F62-8794-6F82C21BD068@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-21T23:00:53Z","receivedAt":"2008-01-21T23:00:53Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jan 21, 2008 at 05:46:27PM -0500, Kevin Ballard wrote:\n> I find it amusing that you keep arguing against having git treat filenames \n> as unicode when, if you had actually taken my advice and read my previous \n> email talking about \"ideal\" vs \"practical\"...\n\nIf by \"ideal\" you mean a world where 100% of all computers were\ndesigned by Steve Jobs, you might have a point.  But trying to argue\nfor such a state of idealism seems to be stupid, and certainly a\ncomplete waste of everyone's time on the git mailing list.  It's\nsimply not reality.  It's like with the infamous resource forks, which\nwould have worked fine if all the world were MacOS, but which had a\ntendency to get stripped off whenver you used a program that wasn't\nresource fork aware, like zip, or a protocol that wasn't resource fork\naware, like FTP.  And so people had to put in all sorts of kludges\nlike BinHex to work around MacOS's \"if only the entire world was like\n*me*, no one would get hurt\" attitude.  In some ways, the MacOS\ndesigners are even worse than Microsoft in terms of having the \"the\nworld revolves around us\" attitude.\n\n> In other words, I was trying to illustrate that \n> HFS+ isn't wrong, it's just different, and the difference is causing the \n> problem.\n\nAnd if you want to interoperate with the rest of the world, where at\nleast count over 92% of computers are NOT running HFS+, then \"Thinking\nDifferent\" is indeed causing the problem, yes.  And whose fault is that?\n\nThe whole point of interoperability is that when we communicate, we\nhave to do so in a uniform and predictable way.  If we can't, the next\nbest thing is to have protocol translators; but in order to do that,\nwe must avoid lossy transformations, such as HFS+'s\npseudo-normalization.  (Why, by the way, will not result in a \"normal\"\nform for any glyph which can be encoded with and without a combining\ncharacter if said glyph was introduced into Unicode after 1988.  So\nyou can't even call it a \"normalization\" algorithm, but just a\npseudo-normalization transformation which is lossy and which DESTROYS\nfilename information in an irrecoverable way.)\n\n\t      \t     \t     \t \t  \t    - Ted\n"},{"id":"66225","messageId":"20080121230118.GF14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"A101AC54-F0CF-4C33-8986-CC6B8DD7F17F@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-21T23:01:18Z","receivedAt":"2008-01-21T23:01:18Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Jan 21, 2008 at 03:53:10PM -0500, Kevin Ballard wrote:\n> On Jan 21, 2008, at 3:43 PM, Dmitry Potapov wrote:\n> \n> >On Mon, Jan 21, 2008 at 11:59:24AM -0500, Kevin Ballard wrote:\n> >>\n> >>No, it's a question of hashing algorithm. And it's one that's fairly\n> >>easily solved simply by picking a specific nonambiguous UTF-8  \n> >>encoding\n> >>before hashing.\n> >\n> >UTF-8 is a *single* encoding, and it maps every Unicode character to\n> >a unique binary representation. So, it is completely nonambiguous.\n> \n> In this case, encoding refers to normalization form,\n\nI thought we spoke about HFS+, and it does not use any normalization\nform, because normalization should produce binary identitical strings\nfor equivalent strings and HFS+ conversion does not. So, it looks\nlike you redefine both words \"encoding\" and \"normalization\" here.\n\n> as other people  \n> have used it in the conversation besides me.\n\nAll your arguments based on confusion and the fact that some other\npeople were probably confused does not make your arguments any more\nvalid.\n\n> I suggest you stop trying to find inconsequential stuff to argue  \n> about, especially when a tiny bit of critical thinking would reveal  \n> the answer.\n\nIMHO, most of your arguments are inconsequential stuff, so I am not\nsure what I am supposed to do about your writings. Probably, it does\nnot make sense to respond your mails anymore...\n\nAs to critical thinking, it definitely reveals that Apple's choice\nwas far from being. Is it so difficult to accept?\n\nAnyway, if you think that you know better than other how to properly\ndeal with the problem, why don't you try to actually *do* something\nand write some code that works as your propose.\n\nDmitry\n"},{"id":"66227","messageId":"CE44F12F-C056-46A6-9795-7F8E280AE9FC@sb.org","threadId":"11645","inReplyTo":"53C76BEA-2232-4940-8776-9DF1880089A4@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T23:05:56Z","receivedAt":"2008-01-21T23:05:56Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"Yet another bounce.\nOriginal CC: martin.langhoff@gmail.com, torvalds@linux-foundation.org, peter@softwolves.pp.se \n, mjscod@web.de, melo@simplicidade.org\n\nOn Jan 21, 2008, at 6:02 PM, Kevin Ballard wrote:\n\n> On Jan 21, 2008, at 5:56 PM, Martin Langhoff wrote:\n>\n>> On Jan 22, 2008 11:46 AM, Kevin Ballard <kevin@sb.org> wrote:\n>>> Again, I was talking about a system that used unicode top-to-bottom.\n>>> On HFS+ you have to use UTF-8 for your filename or it simply won't  \n>>> work.\n>>\n>> Hmmm. I m pretty sure HFS+ has a lot of problems if you run OSX as an\n>> NFS server with clients in different encodings. It would never work  \n>> in\n>> real life. The \"envelope\" OSs have to work in is hugely varied --  \n>> much\n>> more so than any other apps. You should try writing one someday ;-)\n>\n> I'd imagine writing an OS to be a horrifically complicated task. And  \n> yes, I can certainly imagine HFS+ might have issues when used to  \n> back an NFS server with other clients, but that still leads back to  \n> the original point, which is that all these problems stem from the  \n> differences between HFS+ and other filesystems, not any inherent  \n> problem with HFS+ itself.\n>\n>>> other words, I was trying to illustrate that HFS+ isn't wrong, it's\n>>> just different, and the difference is causing the problem.\n>>\n>> Did you spot the rather nasty issues that Ted mentioned earlier in  \n>> the\n>> thread? I would say HFS+ is a bit \"special\" rather than \"different\".\n>\n> IIRC, the biggest problem he talked about was the changing unicode  \n> standard, but since the technote appears to state that HFS+ will not  \n> be changing its normalization algorithms to preserve backwards  \n> compatibility with existing volumes, that doesn't appear to be a  \n> nasty issue after all. Is there another issue I've failed to address  \n> in this thread?\n>\n> -Kevin Ballard\n>\n> -- \n> Kevin Ballard\n> http://kevin.sb.org\n> kevin@sb.org\n> http://www.tildesoft.com\n>\n>\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66228","messageId":"C331E3F9-E177-4C10-B497-528ED08D1E1C@sb.org","threadId":"11645","inReplyTo":"20080121230053.GA317@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-21T23:09:22Z","receivedAt":"2008-01-21T23:09:22Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 6:00 PM, Theodore Tso wrote:\n\n> On Mon, Jan 21, 2008 at 05:46:27PM -0500, Kevin Ballard wrote:\n>> I find it amusing that you keep arguing against having git treat  \n>> filenames\n>> as unicode when, if you had actually taken my advice and read my  \n>> previous\n>> email talking about \"ideal\" vs \"practical\"...\n>\n> If by \"ideal\" you mean a world where 100% of all computers were\n> designed by Steve Jobs, you might have a point.\n\nNO NO NO NO NO. READ MY EMAIL. STOP MAKING ASSUMPTIONS ABOUT WHAT I'M  \nTALKING ABOUT.\n\nThe most frustrating thing about this thread is everybody keeps  \narguing about what they *assume* I'm talking about without actually  \nbothering to read what I'm saying.\n\n>> In other words, I was trying to illustrate that\n>> HFS+ isn't wrong, it's just different, and the difference is  \n>> causing the\n>> problem.\n>\n> And if you want to interoperate with the rest of the world, where at\n> least count over 92% of computers are NOT running HFS+, then \"Thinking\n> Different\" is indeed causing the problem, yes.  And whose fault is  \n> that?\n\nAnd if you want to interoperate with the rest of the world, where at  \nleast count over 92% of computers are running Windows, then using  \nanother OS is stupid, right? Right? I mean, if everyone else is doing  \nit, we should too, shouldn't we?\n\n> The whole point of interoperability is that when we communicate, we\n> have to do so in a uniform and predictable way.  If we can't, the next\n> best thing is to have protocol translators; but in order to do that,\n> we must avoid lossy transformations, such as HFS+'s\n> pseudo-normalization.  (Why, by the way, will not result in a \"normal\"\n> form for any glyph which can be encoded with and without a combining\n> character if said glyph was introduced into Unicode after 1988.  So\n> you can't even call it a \"normalization\" algorithm, but just a\n> pseudo-normalization transformation which is lossy and which DESTROYS\n> filename information in an irrecoverable way.)\n\nSure it's normalization, it's just not using one of the standard  \nforms. But the form is well-defined.\n\nAnd yes, protocol translators are a good idea. That's why I thought  \nthe original suggestion of using a table to map index filenames <-> HFS \n+ filenames sounded like it could work. The only time that should fail  \nis if the index contains multiple filenames that HFS+ will treat as a  \nsingle filename. Is there a problem with this approach?\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66230","messageId":"46a038f90801211516j1f4e10c3q364b446d1293542b@mail.gmail.com","threadId":"11645","inReplyTo":"53C76BEA-2232-4940-8776-9DF1880089A4@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-21T23:16:16Z","receivedAt":"2008-01-21T23:16:16Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 12:02 PM, Kevin Ballard <kevin@sb.org> wrote:\n> I'd imagine writing an OS to be a horrifically complicated task. And yes, I\n> can certainly imagine HFS+ might have issues when used to back an NFS server\n> with other clients, but that still leads back to the original point, which\n> is that all these problems stem from the differences between HFS+ and other\n> filesystems, not any inherent problem with HFS+ itself.\n\nRight. If you are defining the requirements for a new FS on a new OS,\nwould you not include a requirement that says \"must not add any funny\nrule that prevents clean interoperation with other filesystems or\nOSs\"? Forgetting that requirement is... a big one! And if someone asks\n\"how do we do nice user-friendly filename matching with these\ntechnical differences that users mostly don't care about\"... wouldn't\nyou say \"do it in the GUI facilities, changing the FS to handle this\nis wrong because it will break the OS as a server, as a reliable file\nstorage\"?\n\nFSs have pretty hard requirements these days -- all the modern FS\nyou've heard about respect the requirement above, and a ton more that\nyou have to be in the FS business to be aware of. Mostly anyway,\nwherever they don't, users have all sorts of trouble.\n\n> IIRC, the biggest problem he talked about was the changing unicode standard,\n> but since the technote appears to state that HFS+ will not be changing its\n> normalization algorithms to preserve backwards compatibility with existing\n> volumes, that doesn't appear to be a nasty issue after all. Is there another\n> issue I've failed to address in this thread?\n\nWell, Ted answered that part, noting that then the \"normalisation\" is\npatchy, and everyone else is left to guess what chars are normalised\nand what chars aren't so being HFS+ compatible becomes a very weird\ngame indeed. You didn't reply to his explanation -- you called him\narrogant instead. Did you manage to read to the end if his email?\n\nThe HFS+ designers mucked it up -- and then papered over it with the\nOSX libraries. But a good chunk of the world does not use them, they\nforgot about the little \"interop\" requirement.\n\ncheers,\n\n\nm\n"},{"id":"66231","messageId":"20080121232102.GG14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"DC78D5CB-18FF-4504-BD8B-985D8B202817@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-21T23:21:02Z","receivedAt":"2008-01-21T23:21:02Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Jan 21, 2008 at 05:53:50PM -0500, Kevin Ballard wrote:\n> \n> I would go look up specifics to back me up, but my DNS is screwing up  \n> right now so I can't access most of the internet.\n\nThen you are lucky that your mails reach this ML without problem.\n\n> In any case, there  \n> are 4 standard normalization forms - NFC, NFD, NFKC, NFKD. If there  \n> are others, they aren't notable enough to be listed in the resource I  \n> was reading. HFS+ uses a variant on NFD - it's a well-defined variant,  \n> and thus can safely be called its own normalization form. I fail to  \n> see how this means it's not \"normalization\".\n\nThe defining property of normalization is producing binary identitical\nstrings for equivalent strings, IOW, normalization allows you to tell\nwhat strings are equivalent and what are not just by binary comparision.\nHFS+ decomposition lacks that property, because strings are not fully\ndecomposed thus being comparision of equivalent strings may give false\nresult.\n\n> \n> I'm not talking about assigning engineers, I'm saying developing a new  \n> filesystem, especially one that's proven itself to be usable and  \n> extendable for the last decade, is something that only smart engineers  \n> would be capable of doing.\n\nYou know, many people still use FAT, but somehow I don't think that\nFAT is good despite of it being extendable for more than a decade...\nApparently, HFS+ was not worst part of the Copland project, but I\nsee no evidence to think that it was developed by the best engineers.\n\n\n> >>And\n> >>Copland didn't fail because of stupid engineers anyway. If I had to\n> >>blame someone, I'd blame management.\n> >\n> >But if the code was so good then why was most of that code thrown away\n> >later when management was changed? Still bad management?\n> \n> Yes. Even the best of engineers will produce crap code when overworked  \n> and required to implement new features instead of fixing bugs and  \n> stabilizing the system. \n\nI don't think that anyone asked them to implement so much new features.\nAFAIK, it was very difficult (nearly impossible) to get anyone to work\non stabilizing existing software and fixing existing bugs in it.\n\n> Copland is well-known to have suffered from  \n> featuritis, to the extent that it was practically impossible to test  \n> in any sane fashion.\n\nExactly. IMHO, both management and developers are equally responsible\nfor that feature-mania.\n\n> Bad management can kill any project regardless of  \n> how good the engineers are.\n\nSure.\n\n> >Byte sequences are not an issue here. If the filesystem used UTF-16 to\n> >store filenames, that would NOT cause this problem, because characters\n> >would be the same even though bytes stored on the disk were different.\n> >So, what you actually lose here is the original sequence of  \n> >*characters*.\n> \n> I've already talked about that, but you are apparently incapable of  \n> understanding.\n\nWell, it is *you* who is incapable of understanding anything, even\nbasic terms as encoding and normalization...\n\nDmitry\n"},{"id":"66232","messageId":"alpine.LFD.1.00.0801211538590.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"0CA4DF3F-1B64-4F62-8794-6F82C21BD068@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-21T23:44:49Z","receivedAt":"2008-01-21T23:44:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n> \n> I find it amusing that you keep arguing against having git treat filenames as\n> unicode when\n\nNO I DO NOT!\n\nDammit, stop this idiocy.\n\nI think it's fine having git treat filenames \"as unicode\", as long as you \ndon't do any munging on it.\n\nWhy? Because if it's utf-8, then treating them \"as unicode\" means exactly \nthe same as treating them \"as a user-specified string\".\n\nSo stop lying about this whole thing. I have never *ever* argued against \nunicode per se.\n\nAll my complaints - every single one of them - comes down to making the \nidiotic choice of trying to munge those strings (not even strictly \n\"normalize\") into something they are not.\n\nAnd what you don't seem to understand is that once you accept _unmodified_ \nraw UTF-8 as a good unicode transport mechanism, suddenly other encodings \nare possible. I'm not out to force my world-view on users. If they are \nusing legacy encodings (whether in filenames *or* in commit texts or in \ntheir file contents), that's *their* choice.\n\nI actually personally happen to use UTF-8-encoded unicode.\n\nI'm just not stupid enough to think that (a) corrupting it is a good idea, \n*or* (b) that I should force every Asian installation of git to also force \npeople to use unicode (or even having all the conversion libraries and \noverheads!)\n\nSo stop this idiotic \"unicode == normalization\" crap. \n\nI'm a huge fan of UTF-8. But that does not mean that I think normalization \nis a good idea.\n\n\t\tLinus\n"},{"id":"66234","messageId":"A8C7EBCD-8E7F-44AA-89F2-89F6026980B0@sb.org","threadId":"11645","inReplyTo":"46a038f90801211516j1f4e10c3q364b446d1293542b@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-22T00:30:11Z","receivedAt":"2008-01-22T00:30:11Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"\nOn Jan 21, 2008, at 6:16 PM, Martin Langhoff wrote:\n\n> On Jan 22, 2008 12:02 PM, Kevin Ballard <kevin@sb.org> wrote:\n>> I'd imagine writing an OS to be a horrifically complicated task.  \n>> And yes, I\n>> can certainly imagine HFS+ might have issues when used to back an  \n>> NFS server\n>> with other clients, but that still leads back to the original  \n>> point, which\n>> is that all these problems stem from the differences between HFS+  \n>> and other\n>> filesystems, not any inherent problem with HFS+ itself.\n>\n> Right. If you are defining the requirements for a new FS on a new OS,\n> would you not include a requirement that says \"must not add any funny\n> rule that prevents clean interoperation with other filesystems or\n> OSs\"? Forgetting that requirement is... a big one! And if someone asks\n> \"how do we do nice user-friendly filename matching with these\n> technical differences that users mostly don't care about\"... wouldn't\n> you say \"do it in the GUI facilities, changing the FS to handle this\n> is wrong because it will break the OS as a server, as a reliable file\n> storage\"?\n\nSure, but you have to remember, HFS+ was developed back for Mac OS 8,  \nwhich really wasn't a very good server machine.\n\n> FSs have pretty hard requirements these days -- all the modern FS\n> you've heard about respect the requirement above, and a ton more that\n> you have to be in the FS business to be aware of. Mostly anyway,\n> wherever they don't, users have all sorts of trouble.\n>\n>> IIRC, the biggest problem he talked about was the changing unicode  \n>> standard,\n>> but since the technote appears to state that HFS+ will not be  \n>> changing its\n>> normalization algorithms to preserve backwards compatibility with  \n>> existing\n>> volumes, that doesn't appear to be a nasty issue after all. Is  \n>> there another\n>> issue I've failed to address in this thread?\n>\n> Well, Ted answered that part, noting that then the \"normalisation\" is\n> patchy, and everyone else is left to guess what chars are normalised\n> and what chars aren't so being HFS+ compatible becomes a very weird\n> game indeed. You didn't reply to his explanation -- you called him\n> arrogant instead. Did you manage to read to the end if his email?\n\nI've read every single email in this thread, all the way through. Ted  \nwas arguing against calling it \"normalization\". If you want to argue  \nthat it's using a non-standard normal form, go ahead, but surely you  \ncan figure out that can you simply re-normalize it to whatever form  \nyou want.\n\n> The HFS+ designers mucked it up -- and then papered over it with the\n> OSX libraries. But a good chunk of the world does not use them, they\n> forgot about the little \"interop\" requirement.\n\nSure, maybe they did forget about interop. Or maybe they developed  \nthis back on Mac OS 8 where the only real competitor was Windows, and  \nthey didn't have to worry about the Mac being used as an NFS server,  \nand thus interop wasn't even a requirement.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66236","messageId":"alpine.LSU.1.00.0801220031490.5731@racer.site","threadId":"11645","inReplyTo":"BA2377B5-BE7E-40F2-9C3C-679663A966A4@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-22T00:36:42Z","receivedAt":"2008-01-22T00:36:42Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n\n> On Jan 21, 2008, at 4:49 PM, Martin Langhoff wrote:\n> \n> > LOL! Spare us the flamefesting and you will have plenty of time for \n> > learning git internals. You might even learn something.\n> \n> Ah, so I'm flaming while you are providing a well-reasoned and \n> articulate argument? Glad to know the difference.\n\nENOUGH ALREADY!\n\nYes, you are flaming.  You sent easily over 30 totally useless mails in \nthis thread.  Over a couple of days.\n\nAnd Martin is right, in that same amount of time, you could have learnt \nthe internals of git _easily_.  Especially since I provided an own chapter \nin the manual for people like you.\n\nSo instead of _PESTERING_ us with things that we do _NOT CARE_ about, you \ncould _DO SOMETHING USEFUL_ instead.\n\nFor example, read up that chapter, and not type any _POINTLESS_ mails in \nthat _UTTERLY POINTLESS_ thread anymore.\n\nI hoped that subtle _HINTS_ would give you an _IDEA_, but _EVIDENTLY_ I \nhave to use _ALL-CAPS_ on you.\n\nSo _EITHER_ read up _OR_ go away, but _DO NOT BOTHER_ to post _ANY_ \nresponse without patches in this _DARNED_ thread.\n\nSheesh,\nDscho\n"},{"id":"66238","messageId":"32CA7B5D-FED0-4C70-A7DE-14E24B74FD1F@sb.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801220031490.5731@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-22T00:42:44Z","receivedAt":"2008-01-22T00:42:44Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 7:36 PM, Johannes Schindelin wrote:\n\n> Hi,\n>\n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n>\n>> On Jan 21, 2008, at 4:49 PM, Martin Langhoff wrote:\n>>\n>>> LOL! Spare us the flamefesting and you will have plenty of time for\n>>> learning git internals. You might even learn something.\n>>\n>> Ah, so I'm flaming while you are providing a well-reasoned and\n>> articulate argument? Glad to know the difference.\n>\n> ENOUGH ALREADY!\n>\n> Yes, you are flaming.  You sent easily over 30 totally useless mails  \n> in\n> this thread.  Over a couple of days.\n\nAnd so has EVERYONE ELSE. You cannot hold me to a standard which you  \nyourself do not apply.\n\n> And Martin is right, in that same amount of time, you could have  \n> learnt\n> the internals of git _easily_.  Especially since I provided an own  \n> chapter\n> in the manual for people like you.\n\nAs I said before, I've been responding to emails in the midst of doing  \nother things. So no, I can't learn an entirely new system in my off- \ntime between other tasks, but I can respond to emails.\n\n> So instead of _PESTERING_ us with things that we do _NOT CARE_  \n> about, you\n> could _DO SOMETHING USEFUL_ instead.\n\nYou sure argue a lot for something you don't care about.\n\nI'm especially annoyed since I have made MANY offerings to stop this  \nargument and work towards a solution, but NOBODY ELSE seems to care  \nenough to actually accept that offer.\n\nAnd treating me like a moron (using all caps?) just shows how bad you  \nare at actually reading my emails. I've been giving you and everyone  \nelse the courtesy of actually reading your emails, glad to know nobody  \nelse has bothered to read to the end of mine.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66239","messageId":"F663E088-BCAD-4C5D-89D5-EAF97A29C1DE@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211538590.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-22T00:47:49Z","receivedAt":"2008-01-22T00:47:49Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"Please read to the bottom of this email. As near as I can figure out,  \nyou haven't done that on any of my previous emails.\n\nOn Jan 21, 2008, at 6:44 PM, Linus Torvalds wrote:\n\n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n>>\n>> I find it amusing that you keep arguing against having git treat  \n>> filenames as\n>> unicode when\n>\n> NO I DO NOT!\n>\n> Dammit, stop this idiocy.\n>\n> I think it's fine having git treat filenames \"as unicode\", as long  \n> as you\n> don't do any munging on it.\n\nWhen I say \"treat filenames as unicode\" I'm implying the equivalence  \ncomparisons and everything else that we've been talking about.\n\n> Why? Because if it's utf-8, then treating them \"as unicode\" means  \n> exactly\n> the same as treating them \"as a user-specified string\".\n\nIf that's what \"as unicode\" meant, then the phrase \"as unicode\" has  \nzero meaning.\n\n> So stop lying about this whole thing. I have never *ever* argued  \n> against\n> unicode per se.\n\nNo, you've argued against unicode equivalency in filenames. Can't you  \nfigure out, when the entire time I've been talking about equivalency,  \nthat I'm *still* talking about equivalency?\n\n> All my complaints - every single one of them - comes down to making  \n> the\n> idiotic choice of trying to munge those strings (not even strictly\n> \"normalize\") into something they are not.\n\nYes, I understand quite well that you are against munging strings.\n\n> And what you don't seem to understand is that once you accept  \n> _unmodified_\n> raw UTF-8 as a good unicode transport mechanism, suddenly other  \n> encodings\n> are possible. I'm not out to force my world-view on users. If they are\n> using legacy encodings (whether in filenames *or* in commit texts or  \n> in\n> their file contents), that's *their* choice.\n\nYou're not using raw UTF-8, you're just using raw bytes. Calling it  \nUTF-8 doesn't mean anything, since you don't actually know that's what  \nit is. But this is fairly irrelevant.\n\n> I actually personally happen to use UTF-8-encoded unicode.\n>\n> I'm just not stupid enough to think that (a) corrupting it is a good  \n> idea,\n> *or* (b) that I should force every Asian installation of git to also  \n> force\n> people to use unicode (or even having all the conversion libraries and\n> overheads!)\n>\n> So stop this idiotic \"unicode == normalization\" crap.\n>\n> I'm a huge fan of UTF-8. But that does not mean that I think  \n> normalization\n> is a good idea.\n\nHow many times must I say the same thing over and over? I'm not  \narguing that forced normalization is a good thing. I'm arguing that,  \nin a system which is unicode-aware top to bottom, forced normalization  \nis irrelevant to the user, since they don't care about the exact byte  \nsequence. And I'm also arguing that git should have some solution to  \nthis problem. I find it interesting that you're perfectly happy to  \nrant and rail against your misperception of my argument, and yet you  \nconsistently and repeatedly ignore my offers to stop this argument and  \nwork towards a solution, as well as my comments on existing proposed  \nsolutions.\n\nAre you even reading to the end of my emails?\n\n- Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66243","messageId":"85myqy4re8.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"32CA7B5D-FED0-4C70-A7DE-14E24B74FD1F@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-22T00:48:15Z","receivedAt":"2008-01-22T00:48:15Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Kevin Ballard <kevin@sb.org> writes:\n\n> On Jan 21, 2008, at 7:36 PM, Johannes Schindelin wrote:\n>\n>> And Martin is right, in that same amount of time, you could have\n>> learnt the internals of git _easily_.  Especially since I provided an\n>> own chapter in the manual for people like you.\n>\n> As I said before, I've been responding to emails in the midst of doing\n> other things. So no, I can't learn an entirely new system in my off-\n> time between other tasks, but I can respond to emails.\n\nBut there is no point in responding to emails as long as you don't have\na clue what you are actually talking about.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66241","messageId":"alpine.LFD.1.00.0801211656130.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"F663E088-BCAD-4C5D-89D5-EAF97A29C1DE@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-22T01:01:21Z","receivedAt":"2008-01-22T01:01:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n\n> > I think it's fine having git treat filenames \"as unicode\", as long as you\n> > don't do any munging on it.\n> \n> When I say \"treat filenames as unicode\" I'm implying the equivalence\n> comparisons and everything else that we've been talking about.\n\nYes, because you're an idiot.\n\nI've told you over and over again that equivalence is stupid.\n\nIt's stupid when it's \"equivalent except for case\", and it's stupid when \nit's \"canonically equivalent\".\n\n> No, you've argued against unicode equivalency in filenames. Can't you figure\n> out, when the entire time I've been talking about equivalency, that I'm\n> *still* talking about equivalency?\n\nI agree: normalization and equivalency is idiotic.\n\nBut the two actually go hand in hand:\n\n> > All my complaints - every single one of them - comes down to making the\n> > idiotic choice of trying to munge those strings (not even strictly\n> > \"normalize\") into something they are not.\n> \n> Yes, I understand quite well that you are against munging strings.\n\nYou don't seem to.\n\nThe thing is, the two are inexorably intertwined. Any filename equivalence \n(except for the trivial \"identity\" equivalence) INVARIABLY means that \nfilenames get munged.\n\nWhy?\n\nThink about the file name \"Abc\", and think about what happens when you \ncreate it.\n\nNow, think about what happens if that filename is considered equivalent in \ncase..\n\nSee? The filesystem has to *corrupt* the filename.\n\nCan you not UNDERSTAND this? Equivalence and normalization is STUPID. It's \njust two sides of the exact same coin. They both INVARIABLY cause the \nfilename to be munged.\n\nAnd changing user data is not acceptable.\n\nDo you get it now?\n\n\t\t\tLinus \"probably not\" Torvalds\n"},{"id":"66249","messageId":"46a038f90801211706p1d428509ne2c74a647b40b1c2@mail.gmail.com","threadId":"11645","inReplyTo":"32CA7B5D-FED0-4C70-A7DE-14E24B74FD1F@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-22T01:06:12Z","receivedAt":"2008-01-22T01:06:12Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 1:42 PM, Kevin Ballard <kevin@sb.org> wrote:\n> And so has EVERYONE ELSE. You cannot hold me to a standard which you\n> yourself do not apply.\n\nHi Kevin,\n\nnot sure if you are just joking, but perhaps you have not noticed that\nin technical lists like these, you get karma to voice strong opinions\nonce you've contributed lots of good code. *That* is the standard -\nand Johannes has earned his karma points, as Ted and Linus have, by\nlearning the slow way and writing tons of code. Alas, you don't seem\nto have time for such a thing!\n\nStrong opinions without working code tend to not get much respect. So\nthat is the standard that is applied to all of us, including you.\n\ncheers,\n\n\nmartin\n"},{"id":"66250","messageId":"alpine.LFD.1.00.0801211702050.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211656130.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-22T01:13:48Z","receivedAt":"2008-01-22T01:13:48Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Linus Torvalds wrote:\n> \n> Think about the file name \"Abc\", and think about what happens when you \n> create it.\n> \n> Now, think about what happens if that filename is considered equivalent in \n> case..\n> \n> See? The filesystem has to *corrupt* the filename.\n\nLet me make this really clear, because I'm afraid that you won't get it \nwhen I leave out any steps of the way.\n\nLet us say that there is a filename \"xyz\" that is equivalent to a filename \n\"abc\" in *any* way. It does not matter if xyz/abc is Hello/hello, or \nwhether it's two canonically equivalent strings.\n\nSo now, do\n\n\tclose(open(xyz, O_WRONLY | O_CREAT, 0666));\n\tclose(open(abc, O_WRONLY | O_CREAT, 0666));\n\nand then look at the directory contents afterwards.\n\nThere are two, and only two, choices here (*):\n - the filesystem created both files, and they show up as created\n - the filesystem decided they were equivalent, and munged one (or both) \n   of them\n\nNow, let's go back to my claim:\n - munging user data is unacceptable\nand realize that equivalence BY DEFINITION must do it.\n\nSo no, you do *not* get to have your cake and eat it too. You simply \nfundamentally *cannot* have both filename equivalence and a non-munging \nfilesystem. See above why.\n\n\t\tLinus\n\n(*) Actually, there is third choice above, which is:\n\n - the filesystem created the first file, and errored out on the second \n   because it noticed it was equivalent - but not identical - to one it \n   already had\n\n   This one is actually a perfectly fine choice, but it's not \"your\" kind \n   of equivalence, since it actually makes a difference between two \n   equivalent but non-identical names. So the filenames aren't actually \n   interchangable, and this case is really more of a \"the filesystem has \n   some very specific limitations on what it allows\".\n"},{"id":"66252","messageId":"alpine.LSU.1.00.0801220128030.5731@racer.site","threadId":"11645","inReplyTo":"32CA7B5D-FED0-4C70-A7DE-14E24B74FD1F@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-22T01:34:06Z","receivedAt":"2008-01-22T01:34:06Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n\n> On Jan 21, 2008, at 7:36 PM, Johannes Schindelin wrote:\n> \n> > On Mon, 21 Jan 2008, Kevin Ballard wrote:\n> > \n> > > On Jan 21, 2008, at 4:49 PM, Martin Langhoff wrote:\n> > > \n> > > > LOL! Spare us the flamefesting and you will have plenty of time \n> > > > for learning git internals. You might even learn something.\n> > > \n> > > Ah, so I'm flaming while you are providing a well-reasoned and \n> > > articulate argument? Glad to know the difference.\n> > \n> > ENOUGH ALREADY!\n> > \n> > Yes, you are flaming.  You sent easily over 30 totally useless mails \n> > in this thread.  Over a couple of days.\n> \n> And so has EVERYONE ELSE. You cannot hold me to a standard which you \n> yourself do not apply.\n\nWhile I was playing chess, unsuspectingly, there were 40 mails in this \nuseless thread.  You sent 15 of them.\n\nNow, just making a _conservative_ guess, git@vger.kernel.org has about \n50000 subscribers (I know for a fact this figure is too low, so it is \nconservative).\n\nI estimate all except for 10 of them, let's be conservative, 100, did not \ncare about your lognorrhea.  They needed at least 0.5 seconds to delete \nthose mails _you_ are responsible for.\n\nCongratulations, you cost at least 500 seconds today.  And many nerves.\n\nThe sad thing: I read about two people the other day, Singh and McKinstry, \nwho I'd rather see writing mails here.\n\nI don't care if you honour my contributions to git.  I don't care about \nyou.\n\nYou must get a kick out of annoying people.  Why don't you try to play \nchicken with a train?\n\nCiao,\nDscho\n"},{"id":"66256","messageId":"46a038f90801211753p3baebb61y5ed5f389d5151168@mail.gmail.com","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801220128030.5731@racer.site","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-22T01:53:41Z","receivedAt":"2008-01-22T01:53:41Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 2:34 PM, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n\nHey guys. Let's stop right here -- Kevin has perhaps been annoying but\nthis is a *technical* argument, so let's go back to working code.\n\nAnyone send a patch to clear the air? *please*?\n\n\n\n\n\nm\n"},{"id":"66257","messageId":"alpine.LSU.1.00.0801220202030.5731@racer.site","threadId":"11645","inReplyTo":"46a038f90801211753p3baebb61y5ed5f389d5151168@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-22T02:03:30Z","receivedAt":"2008-01-22T02:03:30Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 22 Jan 2008, Martin Langhoff wrote:\n\n> On Jan 22, 2008 2:34 PM, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> \n> Hey guys. Let's stop right here -- Kevin has perhaps been annoying but \n> this is a *technical* argument, so let's go back to working code.\n> \n> Anyone send a patch to clear the air? *please*?\n\nNot me.  This thread was not _my_ fault.\n\nAnd I'll be _damned_ if I encourage people to be a pain in the rear end, \nin order to get other people to write code/patches for them.\n\nI think it is clear who has the obligation of contributing some _code_, \nfor a change.\n\nCiao,\nDscho\n"},{"id":"66260","messageId":"34103945-2078-4983-B409-2D01EF071A8B@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211702050.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-22T02:33:12Z","receivedAt":"2008-01-22T02:33:12Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"Linus, have you even bothered to read my arguments, or do you just get  \na kick out of building these straw man arguments? You have  \nconsistently failed to actually address what I'm talking about, and  \ninstead persist in explaining stuff I already know, as if that was the  \nanswer to anything I've been talking about. You are clearly incapable  \nof understanding my basic point, no matter how simple I break it down.  \nI suspect it's because you've been working low-level so long you can't  \nthink high-level, and so you manage to misinterpret my high-level  \narguments as boneheaded low-level mistakes.\n\nAnyway, please see my countless former emails where I ask to work  \ntowards a solution instead of just arguing.\n\n-Kevin Ballard\n\nOn Jan 21, 2008, at 8:13 PM, Linus Torvalds wrote:\n\n>\n>\n> On Mon, 21 Jan 2008, Linus Torvalds wrote:\n>>\n>> Think about the file name \"Abc\", and think about what happens when  \n>> you\n>> create it.\n>>\n>> Now, think about what happens if that filename is considered  \n>> equivalent in\n>> case..\n>>\n>> See? The filesystem has to *corrupt* the filename.\n>\n> Let me make this really clear, because I'm afraid that you won't get  \n> it\n> when I leave out any steps of the way.\n>\n> Let us say that there is a filename \"xyz\" that is equivalent to a  \n> filename\n> \"abc\" in *any* way. It does not matter if xyz/abc is Hello/hello, or\n> whether it's two canonically equivalent strings.\n>\n> So now, do\n>\n> \tclose(open(xyz, O_WRONLY | O_CREAT, 0666));\n> \tclose(open(abc, O_WRONLY | O_CREAT, 0666));\n>\n> and then look at the directory contents afterwards.\n>\n> There are two, and only two, choices here (*):\n> - the filesystem created both files, and they show up as created\n> - the filesystem decided they were equivalent, and munged one (or  \n> both)\n>   of them\n>\n> Now, let's go back to my claim:\n> - munging user data is unacceptable\n> and realize that equivalence BY DEFINITION must do it.\n>\n> So no, you do *not* get to have your cake and eat it too. You simply\n> fundamentally *cannot* have both filename equivalence and a non- \n> munging\n> filesystem. See above why.\n>\n> \t\tLinus\n>\n> (*) Actually, there is third choice above, which is:\n>\n> - the filesystem created the first file, and errored out on the second\n>   because it noticed it was equivalent - but not identical - to one it\n>   already had\n>\n>   This one is actually a perfectly fine choice, but it's not \"your\"  \n> kind\n>   of equivalence, since it actually makes a difference between two\n>   equivalent but non-identical names. So the filenames aren't actually\n>   interchangable, and this case is really more of a \"the filesystem  \n> has\n>   some very specific limitations on what it allows\".\n>\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66261","messageId":"31F9ADDC-008D-4F06-97E6-CF1D16238DF9@sb.org","threadId":"11645","inReplyTo":"85zluy4xf0.fsf@lola.goethe.zz","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-22T02:34:49Z","receivedAt":"2008-01-22T02:34:49Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 5:38 PM, David Kastrup wrote:\n\n>> How many times must I say I never suggested actually changing git's\n>> hashing algorithm? And if you want me to suggest a fix to git that\n>> works, first you have to wait for me to learn how git's internals\n>> work, and frankly, I have too much work on my plate right now to\n>> devote the time necessary to learning git's internals well enough to\n>> fix this problem.\n>\n> Then please understand that you have too much work on your plate right\n> now to devote the time necessary to provide any constructive  \n> criticism.\n> A smart person in this situation would shut up until he has the time.\n\nA smart person would not join the conversation late and respond to  \npoints that have already been exhausted ages ago.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66262","messageId":"alpine.LFD.1.00.0801211846010.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"34103945-2078-4983-B409-2D01EF071A8B@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-22T02:50:24Z","receivedAt":"2008-01-22T02:50:24Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n> \n> Anyway, please see my countless former emails where I ask to work towards a\n> solution instead of just arguing.\n\nWe know what the solution is:\n\n - The OS X filesystem _is_ crap (and you seem to have almost admitted as \n   much by your comment that the HFS+ designers did it back in the dark \n   ages and didn't mean for it to ever be a server filesystem anyway)\n\n - But we can at least make a wrapper around readdir() return the NFC form \n   on OS X, and effectively hide much of the fallout from the crap.\n\nThere is no way around it. Your \"solutions\" all seem to boil down to \nasking git to do the same idiotic crap that OS X does, taking all the \nsame performance hits, and just generally doing crap just to work around \ncrap in your favourite OS.\n\nAnd no, making git be stupid just to suit a stupid filesystem simply isn't \ngoing to happen.\n\nSo how about you see _my_ point instead: OS X may have an inferior \nfilesystem, but we don't have to make git inferior just for that. The fact \nthat OS X does case independence is *its* problem, not git's.\n\n\t\tLinus\n"},{"id":"66263","messageId":"E3E4F5B3-1740-47E4-A432-C881830E2037@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801211846010.2957@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-22T03:04:07Z","receivedAt":"2008-01-22T03:04:07Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 9:50 PM, Linus Torvalds wrote:\n\n> On Mon, 21 Jan 2008, Kevin Ballard wrote:\n>>\n>> Anyway, please see my countless former emails where I ask to work  \n>> towards a\n>> solution instead of just arguing.\n>\n> We know what the solution is:\n>\n> - The OS X filesystem _is_ crap (and you seem to have almost  \n> admitted as\n>   much by your comment that the HFS+ designers did it back in the dark\n>   ages and didn't mean for it to ever be a server filesystem anyway)\n\nI agree that HFS+ isn't well suited for tasks which it is being asked  \nto do. I was never arguing that it was the perfect filesystem. But  \nthat hardly matters now, I know nobody's going to bother understanding  \nmy argument so I may as well just stop trying.\n\n> - But we can at least make a wrapper around readdir() return the NFC  \n> form\n>   on OS X, and effectively hide much of the fallout from the crap.\n\nAgain, I don't think that's the correct solution. What about the  \ntranslation table that was suggested back at the beginning of the  \nthread? That would solve the case insensitivity issue as well, whereas  \nthis NFC \"solution\" does nothing for that.\n\n> There is no way around it. Your \"solutions\" all seem to boil down to\n> asking git to do the same idiotic crap that OS X does, taking all the\n> same performance hits, and just generally doing crap just to work  \n> around\n> crap in your favourite OS.\n\nNo, I am not asking git to do the same thing HFS+ does. You just  \npersist in misinterpreting my arguments, no matter how many times I  \nprotest that this is not what I am saying.\n\n> And no, making git be stupid just to suit a stupid filesystem simply  \n> isn't\n> going to happen.\n>\n> So how about you see _my_ point instead: OS X may have an inferior\n> filesystem, but we don't have to make git inferior just for that.  \n> The fact\n> that OS X does case independence is *its* problem, not git's.\n\nSo, what, you're saying git shouldn't do any work at all to try and  \nbehave nicer on OS X? Because OS X sure as hell can't change to suit  \ngit.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66264","messageId":"alpine.LFD.1.00.0801211914500.2957@woody.linux-foundation.org","threadId":"11645","inReplyTo":"E3E4F5B3-1740-47E4-A432-C881830E2037@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-22T03:17:51Z","receivedAt":"2008-01-22T03:17:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Jan 2008, Kevin Ballard wrote:\n> \n> No, I am not asking git to do the same thing HFS+ does. You just persist in\n> misinterpreting my arguments, no matter how many times I protest that this is\n> not what I am saying.\n\nSure you do. You continue to say that unicode is the only choice, and you \ncontinue to say that unicode requires that equivalent names be considered \nthe same.\n\nWhat part of that was I mis-interpreting?\n\n> So, what, you're saying git shouldn't do any work at all to try and behave\n> nicer on OS X? Because OS X sure as hell can't change to suit git.\n\nUmm. Git works perfectly fine on OS X, and it's not like we can do a whole \nlot more about it, exactly because we cannot fix the real problem. We can \nhide some of the fallout (idiotic choice of normalization), but the bigger \nissues we can hardly even do anything about (case independence).\n\nAnd quite frankly, you've also made sure that I have absolutely zero \ninterest in even trying to help people with it.\n\n\t\t\tLinus\n"},{"id":"66266","messageId":"46a038f90801211921j777748cfjf080220e98b01fb6@mail.gmail.com","threadId":"11645","inReplyTo":"E3E4F5B3-1740-47E4-A432-C881830E2037@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-22T03:21:06Z","receivedAt":"2008-01-22T03:21:06Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 22, 2008 4:04 PM, Kevin Ballard <kevin@sb.org> wrote:\n> Again, I don't think that's the correct solution. What about the\n> translation table that was suggested back at the beginning of the\n> thread? That would solve the case insensitivity issue as well, whereas\n> this NFC \"solution\" does nothing for that.\n\nKevin,\n\nyou seem to know the problem fairly well. Could you write up a set of\ntestcases that show the bug? See the \"t\" directory in the git sources\n-- you don't need to learn much about git internals, they are just\nshell scripts (mostly, I think there's some perl there too). That\ncould lead to a good contribution to the project.\n\n... and keep you from telling everyone else that you know better how\nto hack a project that you know nothing about ;-)\n\n(...)\n> So, what, you're saying git shouldn't do any work at all to try and\n> behave nicer on OS X?\n\nKevin - for your edification, that question is usually referred to as\n\"trolling\" in this place we call the internet. Linus outlined what his\ntechnical plan is, so git will probably do something designed by\nsomeone who knows a thing or two about git's internals. So when you\npretend that he is saying the opposite of what he is saying... well,\npeople do get upset.\n\ncheers,\n\n\nm\n"},{"id":"66273","messageId":"2D7A24EC-06DF-4F73-8638-ABE437F33311@sb.org","threadId":"11645","inReplyTo":"46a038f90801211921j777748cfjf080220e98b01fb6@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-22T04:22:23Z","receivedAt":"2008-01-22T04:22:23Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 21, 2008, at 10:21 PM, Martin Langhoff wrote:\n\n> On Jan 22, 2008 4:04 PM, Kevin Ballard <kevin@sb.org> wrote:\n>> Again, I don't think that's the correct solution. What about the\n>> translation table that was suggested back at the beginning of the\n>> thread? That would solve the case insensitivity issue as well,  \n>> whereas\n>> this NFC \"solution\" does nothing for that.\n>\n> Kevin,\n>\n> you seem to know the problem fairly well. Could you write up a set of\n> testcases that show the bug? See the \"t\" directory in the git sources\n> -- you don't need to learn much about git internals, they are just\n> shell scripts (mostly, I think there's some perl there too). That\n> could lead to a good contribution to the project.\n\nSee now this is actually a very good suggestion. I probably should  \nhave done this long ago. Thank you very much for actually responding  \nabout the problem. You are the first person to do so.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66284","messageId":"85abmy47so.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"31F9ADDC-008D-4F06-97E6-CF1D16238DF9@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-22T07:51:35Z","receivedAt":"2008-01-22T07:51:35Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Kevin Ballard <kevin@sb.org> writes:\n\n> On Jan 21, 2008, at 5:38 PM, David Kastrup wrote:\n>\n>>> How many times must I say I never suggested actually changing git's\n>>> hashing algorithm? And if you want me to suggest a fix to git that\n>>> works, first you have to wait for me to learn how git's internals\n>>> work, and frankly, I have too much work on my plate right now to\n>>> devote the time necessary to learning git's internals well enough to\n>>> fix this problem.\n>>\n>> Then please understand that you have too much work on your plate right\n>> now to devote the time necessary to provide any constructive\n>> criticism.\n>> A smart person in this situation would shut up until he has the time.\n>\n> A smart person would not join the conversation late and respond to\n> points that have already been exhausted ages ago.\n\nFind somebody willing to explain to you the difference between Email and\nIRC, and how to read \"Date:\" headers.  I have no doubt you'll be able to\ngrasp the basic involved principles in less than a week.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66338","messageId":"20080123000841.GA22704@mit.edu","threadId":"11645","inReplyTo":"20080122133427.GB17804@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-23T00:08:41Z","receivedAt":"2008-01-23T00:08:41Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Jan 22, 2008 at 08:34:27AM -0500, Theodore Tso wrote:\n> \t* Documenting HFS+'s current pseudo-normalization algorithm.\n> \t  It's not enough to say that you need to decompose all\n> \t  Unicode characters, since you've claimed that HFS+ doesn't\n> \t  decompose Unicode characters after some magic date,\n> \t  presumably roughly 9 years ago.\n\nI did some research on this point, since if we really are going to be\ncompatible with MacOS X's crappy HFS+ system, we need to know what the\ndecomposition algorithm actually is.  Turns out, there are *two* of\nthem.  Kevin didn't know what he was talking about.  In fact,\ndifferent versions of Mac OS X use different normalization algorithms.\n\nMac OS X 8.1 through 10.2.x used decompositions based on Unicode 2.1.\nMac OS X 10.3 and later use decompositions based on Unicode 3.2.[1]\n\nAs I correctly predicted, Apple is changing their normalization\nalgorithm in different versions of Mac OS X.  It is not static, which\nmeands there will be compatibility problems when moving hard drives\nbetween Mac OS X versions.  I don't know if they try to fix this in\ntheir fsck or not, when upgrading from 10.2 to 10.3, but if not,\ncertain files could disappear as part of the Mac OS X upgrade.  Fun\nfun fun.\n\nAnd clearly Kevin didn't read the tech note very carefully, since it\nclearly admits why they did it.  The Mac OS X developers were being\ncheasy with how they implemented their HFS B-tree algorithms, and took\nthe cheap, easy way out.  So yeah, \"crappy\" is the only word that can\nbe used for what Mac OS X perpetuated on the world.  Because of that,\na quick Google search shows it causes problems all over the stack, for\nmany different programs beyond just git, including limewire and\ngnutella[2][3], Slim[4], and no doubt others.\n\n[1] http://developer.apple.com/technotes/tn/tn1150.html#UnicodeSubtleties\n[2] http://lists.limewire.org/pipermail/gui-dev/2003-January/001110.html\n[3] http://osdir.com/ml/network.gnutella.limewire.core.devel/2003-01/msg00000.html\n[4] http://forums.slimdevices.com/showthread.php?t=40582\n\nIn any case, it seems pretty clear that by now everyone except Kevin\nhas realized that HFS+ is crappy and causes Internet-wide\ninteroperability problems.  So I'll justify sending this note by\npointing out the specific table of Mac OS's filesystem corruption\nalgorithm can be found here:\n\n\t  http://developer.apple.com/technotes/tn/tn1150table.html\n\nI'd also recommend that the Mac OS X code try to either figure out\nwhether it is running on an HFS+ partition, or let the HFS+ workaround\ncode be something that can be controlled via .git/config.  It\nshouldn't be on unconditionally even on a Mac OS X system, since if\nthe git repository is on a ZFS or NFS filesystem, there's no reason to\npay the overhead of working around the HFS+ bugs.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"66339","messageId":"E6F76F93-24C9-4D10-813C-770A9C3A9828@sb.org","threadId":"11645","inReplyTo":"20080123000841.GA22704@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-23T00:38:04Z","receivedAt":"2008-01-23T00:38:04Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 22, 2008, at 7:08 PM, Theodore Tso wrote:\n\n> On Tue, Jan 22, 2008 at 08:34:27AM -0500, Theodore Tso wrote:\n>> \t* Documenting HFS+'s current pseudo-normalization algorithm.\n>> \t  It's not enough to say that you need to decompose all\n>> \t  Unicode characters, since you've claimed that HFS+ doesn't\n>> \t  decompose Unicode characters after some magic date,\n>> \t  presumably roughly 9 years ago.\n>\n> I did some research on this point, since if we really are going to be\n> compatible with MacOS X's crappy HFS+ system, we need to know what the\n> decomposition algorithm actually is.  Turns out, there are *two* of\n> them.  Kevin didn't know what he was talking about.  In fact,\n> different versions of Mac OS X use different normalization algorithms.\n>\n> Mac OS X 8.1 through 10.2.x used decompositions based on Unicode 2.1.\n> Mac OS X 10.3 and later use decompositions based on Unicode 3.2.[1]\n>\n> As I correctly predicted, Apple is changing their normalization\n> algorithm in different versions of Mac OS X.  It is not static, which\n> meands there will be compatibility problems when moving hard drives\n> between Mac OS X versions.  I don't know if they try to fix this in\n> their fsck or not, when upgrading from 10.2 to 10.3, but if not,\n> certain files could disappear as part of the Mac OS X upgrade.  Fun\n> fun fun.\n>\n> And clearly Kevin didn't read the tech note very carefully, since it\n> clearly admits why they did it.  The Mac OS X developers were being\n> cheasy with how they implemented their HFS B-tree algorithms, and took\n> the cheap, easy way out.  So yeah, \"crappy\" is the only word that can\n> be used for what Mac OS X perpetuated on the world.  Because of that,\n> a quick Google search shows it causes problems all over the stack, for\n> many different programs beyond just git, including limewire and\n> gnutella[2][3], Slim[4], and no doubt others.\n>\n> [1] http://developer.apple.com/technotes/tn/tn1150.html#UnicodeSubtleties\n> [2] http://lists.limewire.org/pipermail/gui-dev/2003-January/001110.html\n> [3] http://osdir.com/ml/network.gnutella.limewire.core.devel/2003-01/msg00000.html\n> [4] http://forums.slimdevices.com/showthread.php?t=40582\n>\n> In any case, it seems pretty clear that by now everyone except Kevin\n> has realized that HFS+ is crappy and causes Internet-wide\n> interoperability problems.  So I'll justify sending this note by\n> pointing out the specific table of Mac OS's filesystem corruption\n> algorithm can be found here:\n>\n> \t  http://developer.apple.com/technotes/tn/tn1150table.html\n>\n> I'd also recommend that the Mac OS X code try to either figure out\n> whether it is running on an HFS+ partition, or let the HFS+ workaround\n> code be something that can be controlled via .git/config.  It\n> shouldn't be on unconditionally even on a Mac OS X system, since if\n> the git repository is on a ZFS or NFS filesystem, there's no reason to\n> pay the overhead of working around the HFS+ bugs.\n\nI just finished talking to one of the HFS+ developers, so I suspect I  \nknow a lot more on this subject now than you do. Here's some of the  \nrelevant information:\n\n* Any new characters added to Unicode will only have one form  \n(decomposed), so HFS+ will always accept new characters as they will  \nbe NFD. The only exception is case-sensitivity, as the case-folding  \ntables in HFS+ are static, so new characters with case variants will  \nbe treated in a case-sensitive manner. However, as they are already  \ndecomposed, the NFD algorithm will not change their encoding. This  \nmeans that no, there are zero problems moving HFS+ drives between  \nversions of OS X.\n\n* At the time HFS+ was developed, there was no one common standard for  \nnormalization. The HFS+ developers picked NFD because they thought it  \nwas \"a more flexible, future-looking form\", but Microsoft ended up  \npicking the opposite just a short time later. Interestingly, NFC is a  \nweird hybrid form which only has composed forms for pre-existing  \ncharacters, and decomposed forms for all new characters (as they only  \nhave one form). So in a sense NFD is more sane then NFC.\n\n* The core issue here, which is why you think HFS+ is so stupid, is  \nthat you guys see no problem with having 2 files \"Märchen\" (NFC) and  \n\"Märchen\" (NFD), whereas the HFS+ developers don't consider it  \nacceptable to have 2 visually identical names as independent files.  \nUnfortunately, the only way to do this matching is to store the  \nnormalized form in the filesystem, because it would be a performance  \nnightmare to try and do this matching any other way. The HFS+  \ndevelopers considered it an acceptable trade-off, and as an  \napplication developer I tend to agree with them.\n\nAs I have stated in the past, this isn't a case of HFS+ being stupid  \nand causing problems, it's a case of HFS+ being *different* and  \ncausing problems. But this difference is just as much your fault as it  \nis HFS+'s fault.\n\n* For detecting case-sensitive filesystems you can use pathconf(2):  \n_PC_CASE_SENSITIVE (if unsupported, you can assume the filesystem is  \ncase-sensitive). There is also the getattrlist(2) attribute:  \nVOL_CAP_FMT_CASE_SENSITIVE.\n\nThere appears to be no API for determining if normalization will be  \napplied. However, any filesystem that uses UTF-8 explicitly as storage  \n(unlike the Linux filesystems, which you claim use UTF-8 but is  \nobviously you really use nothing at all) is pretty much guaranteed to  \nhave to normalize or it will have abysmal performance.\n\nI must say it is shocking that someone as smart as you is still more  \ninterested in finding ways to prove me wrong then to actually address  \nthe problem. It's obvious that the only research you did was intended  \nto find ways to call me stupid.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66340","messageId":"alpine.LFD.1.00.0801221625510.1741@woody.linux-foundation.org","threadId":"11645","inReplyTo":"20080123000841.GA22704@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-23T00:38:37Z","receivedAt":"2008-01-23T00:38:37Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 22 Jan 2008, Theodore Tso wrote:\n> \n> I'd also recommend that the Mac OS X code try to either figure out\n> whether it is running on an HFS+ partition, or let the HFS+ workaround\n> code be something that can be controlled via .git/config.  It\n> shouldn't be on unconditionally even on a Mac OS X system, since if\n> the git repository is on a ZFS or NFS filesystem, there's no reason to\n> pay the overhead of working around the HFS+ bugs.\n\nOne thing I'd like somebody to check: what _does_ happen with OS X and NFS \n(OS X as a client, not server)? In particular:\n\n - Is it suddenly sane and case-sensitive?\n\n - Does the NFS client do any unicode conversion?\n\nI tried to google for it, but didn't find the right keywords to get \nanything useful out of that modern-day internet oracle.\n\n\t\tLinus\n"},{"id":"66341","messageId":"46a038f90801221714x1c31ac8du7944980d47335263@mail.gmail.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801221625510.1741@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-23T01:14:55Z","receivedAt":"2008-01-23T01:14:55Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 23, 2008 1:38 PM, Linus Torvalds <torvalds@linux-foundation.org> wrote:\n> One thing I'd like somebody to check: what _does_ happen with OS X and NFS\n> (OS X as a client, not server)? In particular:\n>\n>  - Is it suddenly sane and case-sensitive?\n\nYes. Similarlty with UFS partitions. After much grief with\ncase-insensitivity on OSX I reinstalled the OS on a UFS partition,\nonly to find that most 3rd party apps can't cope with case-sensitive\nFSs (this was a while ago, I hope it's gotten better).\n\n>  - Does the NFS client do any unicode conversion?\n\nDon't know, unfortunately. I suspect both bits of mangling happen in\nthe fs code.\n\n\nmartin\n"},{"id":"66342","messageId":"98315FA6-CFEF-4BB1-997B-0B10BDBBE37B@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801221625510.1741@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-23T01:16:04Z","receivedAt":"2008-01-23T01:16:04Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 22, 2008, at 7:38 PM, Linus Torvalds wrote:\n\n> On Tue, 22 Jan 2008, Theodore Tso wrote:\n>>\n>> I'd also recommend that the Mac OS X code try to either figure out\n>> whether it is running on an HFS+ partition, or let the HFS+  \n>> workaround\n>> code be something that can be controlled via .git/config.  It\n>> shouldn't be on unconditionally even on a Mac OS X system, since if\n>> the git repository is on a ZFS or NFS filesystem, there's no reason  \n>> to\n>> pay the overhead of working around the HFS+ bugs.\n>\n> One thing I'd like somebody to check: what _does_ happen with OS X  \n> and NFS\n> (OS X as a client, not server)? In particular:\n>\n> - Is it suddenly sane and case-sensitive?\n>\n> - Does the NFS client do any unicode conversion?\n>\n> I tried to google for it, but didn't find the right keywords to get\n> anything useful out of that modern-day internet oracle.\n\nStraight from the horse's mouth, so to speak:\n\n>> Here's one further question: How does OS X behave as an NFS client?  \n>> Does it do any unicode normalization? Is it case-sensitive?\n>>\n> No conversions are done.. the MacOS X nfs client just sends\n> whatever string it was passed to the server.  If I connect\n> to a MacOS X server exporting an HFS file system, I can\n> \"touch FOO\" and then \"rm foo\" and the rm will work.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66343","messageId":"46a038f90801221727t188b33a9t86ddd2747473b274@mail.gmail.com","threadId":"11645","inReplyTo":"98315FA6-CFEF-4BB1-997B-0B10BDBBE37B@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-23T01:27:55Z","receivedAt":"2008-01-23T01:27:55Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 23, 2008 2:16 PM, Kevin Ballard <kevin@sb.org> wrote:\n> > If I connect\n> > to a MacOS X server exporting an HFS file system, I can\n> > \"touch FOO\" and then \"rm foo\" and the rm will work.\n\nSo this bit of insanity can affect users on other OSs too, if they use\ngit on an NFS mountpoint hosted on OSX/HFS+.\n\nIIRC Apple does recommend UFS for servers though. I wonder how XServe\nmachines ship by default.\n\n\n\nm\n"},{"id":"66344","messageId":"20080123013325.GB1320@mit.edu","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801221625510.1741@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-23T01:33:25Z","receivedAt":"2008-01-23T01:33:25Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Jan 22, 2008 at 04:38:37PM -0800, Linus Torvalds wrote:\n> One thing I'd like somebody to check: what _does_ happen with OS X and NFS \n> (OS X as a client, not server)? In particular:\n> \n>  - Is it suddenly sane and case-sensitive?\n\nUsing a Linux server, and a OS X client, over NFS, it is in\ncase-sensitive.  This is not unexpected, since you can mount UFS\npartitions on Mac OS X, or reformat HFS+ filesystems and make them be\ncase-sensitive.\n\n>  - Does the NFS client do any unicode conversion?\n\nNope:\n\n# perl -CO -e 'print pack(\"U\",0x00C4).\"\\n\"'  | xargs touch\n# ls -l | cat -v\ntotal 0\n0 -rw-r--r--   1 nobody  nobody  0 Jan 22 20:30 M-CM-^D\n\nIt's pretty clear the Unicode conversion is being done in HFS+, not in\nthe VFS layer of Mac OS X.\n\nSo presumably if and when Mac OS adopts ZFS, they will be able to be\nfree of this mess, at least if they care about being compatible with\nSolaris.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"66346","messageId":"46a038f90801221747m53975f39y7e2f6b48c5a7f6aa@mail.gmail.com","threadId":"11645","inReplyTo":"E6F76F93-24C9-4D10-813C-770A9C3A9828@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-23T01:47:48Z","receivedAt":"2008-01-23T01:47:48Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 23, 2008 1:38 PM, Kevin Ballard <kevin@sb.org> wrote:\n> I must say it is shocking\n\nDon't ruin it. You were silent for 12hs and lots of patches and\nresearch on the problem started flowing. If you keep making a nuisance\nof yourself, people will turn from helping you to beating you up for\nbeing so annoying.\n\nPerhaps help prepare those tests you said it was a good idea to work\non. If you manage to stay silent a bit, we'll need them soon ;-)\n\n\nm\n"},{"id":"66347","messageId":"alpine.LFD.1.00.0801221743470.1741@woody.linux-foundation.org","threadId":"11645","inReplyTo":"20080123013325.GB1320@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-23T01:56:54Z","receivedAt":"2008-01-23T01:56:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 22 Jan 2008, Theodore Tso wrote:\n> \n> It's pretty clear the Unicode conversion is being done in HFS+, not in\n> the VFS layer of Mac OS X.\n\nOk. That's going to make it both easier and harder for them in the future. \nIn particular, it probably means that their VFS layer really has no notion \nof this at all, and it's going to be fairly hard to support any kind of \ngeneric \"backwards compatibility\" layer on top of other filesystems.\n\n> So presumably if and when Mac OS adopts ZFS, they will be able to be\n> free of this mess, at least if they care about being compatible with\n> Solaris.\n\nI wouldn't hold my breadth on ZFS, considering the memory requirements. \nZFS apparently wants *lots* of memory:\n\n\thttp://www.solarisinternals.com/wiki/index.php/ZFS_Best_Practices_Guide#ZFS_Administration_Considerations\n\thttp://wiki.freebsd.org/ZFSTuningGuide\n\nin fact it seems that the FreeBSD people basically recomment against using \nZFS on 32-bit kernels because of the memory use issues.\n\nYes, it could be BSD-specific, but considering Solaris has the same \nrecommendation, it sure seems like ZFS isn't ready for prime time on any \nlow-end (read: consumer) hardware.\n\nOf course, in a year or two, 2GB will be the norm. Right now it's still \nfairly unusual on Mac hardware outside of the Mac Pro line (which, I \nthink, comes with a *minimum* of 2GB), and the people who get it want it \nnot for the filesystem caches, but for big photo editing jobs..\n\n\t\t\tLinus\n"},{"id":"66348","messageId":"A3E36E38-5BBC-4E71-8220-FAB498A59EB5@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801221743470.1741@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-23T02:02:35Z","receivedAt":"2008-01-23T02:02:35Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 22, 2008, at 8:56 PM, Linus Torvalds wrote:\n\n> On Tue, 22 Jan 2008, Theodore Tso wrote:\n>>\n>> It's pretty clear the Unicode conversion is being done in HFS+, not  \n>> in\n>> the VFS layer of Mac OS X.\n>\n> Ok. That's going to make it both easier and harder for them in the  \n> future.\n> In particular, it probably means that their VFS layer really has no  \n> notion\n> of this at all, and it's going to be fairly hard to support any kind  \n> of\n> generic \"backwards compatibility\" layer on top of other filesystems.\n\nHFS+ was developed on Mac OS 8, which I believe didn't have the notion  \nof a VFS, or at least not one that would have been in any way capable  \nof doing the case-insensitivity and normalization necessary. However,  \nI'm not sure what you mean by a \"backwards compatibility\" layer on  \nother filesystems - if you mean treating another filesystem like HFS+,  \nwell, if you're using a filesystem that doesn't do normalization then  \nthe VFS really shouldn't do it for you.\n\n>> So presumably if and when Mac OS adopts ZFS, they will be able to be\n>> free of this mess, at least if they care about being compatible with\n>> Solaris.\n>\n> I wouldn't hold my breadth on ZFS, considering the memory  \n> requirements.\n> ZFS apparently wants *lots* of memory:\n>\n> \thttp://www.solarisinternals.com/wiki/index.php/ZFS_Best_Practices_Guide#ZFS_Administration_Considerations\n> \thttp://wiki.freebsd.org/ZFSTuningGuide\n>\n> in fact it seems that the FreeBSD people basically recomment against  \n> using\n> ZFS on 32-bit kernels because of the memory use issues.\n>\n> Yes, it could be BSD-specific, but considering Solaris has the same\n> recommendation, it sure seems like ZFS isn't ready for prime time on  \n> any\n> low-end (read: consumer) hardware.\n>\n> Of course, in a year or two, 2GB will be the norm. Right now it's  \n> still\n> fairly unusual on Mac hardware outside of the Mac Pro line (which, I\n> think, comes with a *minimum* of 2GB), and the people who get it  \n> want it\n> not for the filesystem caches, but for big photo editing jobs..\n\nActually, interestingly the new MacBook Air comes with 2GB stock (I'm  \nassuming it's soldered onto the motherboard, though, so it makes sense  \nthat Apple's giving customers 2GB as they can't upgrade themselves).\n\nIn any case, everybody's making a big fuss about ZFS, but it really  \ndoesn't make a lot of sense to use for a consumer system, it seems  \nmore geared for a server.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66349","messageId":"20080123020607.GC1320@mit.edu","threadId":"11645","inReplyTo":"E6F76F93-24C9-4D10-813C-770A9C3A9828@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-23T02:06:07Z","receivedAt":"2008-01-23T02:06:07Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Jan 22, 2008 at 07:38:04PM -0500, Kevin Ballard wrote:\n> * Any new characters added to Unicode will only have one form (decomposed), \n> so HFS+ will always accept new characters as they will be NFD. The only \n> exception is case-sensitivity, as the case-folding tables in HFS+ are \n> static, so new characters with case variants will be treated in a \n> case-sensitive manner. However, as they are already decomposed, the NFD \n> algorithm will not change their encoding. This means that no, there are \n> zero problems moving HFS+ drives between versions of OS X.\n\nExcept there *are* problems, because this promise doesn't apply to\nUnicode 2.1 (Mac OS 10.2 and before) and Unicode 3.2 (Mac OS 10.3 and\nabove).  And there were changes between the normalization algorithm\nbetween Unicode 3.2 and the Unicode version 4.1.  So taking a hard\ndrive between Mac OS X 10.2 and 10.3 *will* cause problems.  The\nguarantees of Unicode stability didn't come until well past Unicode\n2.1.\n\nAlso, I know of no guarantee that there will be no more new\ncompositions.  According to Unicode Stnadard Annex #15\n(http://unicode.org/reports/tr15/), new characters that can be\ndecomposed are strongly discouraged, but \"It would be possible to add\nmore compositions in a future version of Unicode\".  Got a reference to\nback up your claim that there will never be any more?\n\n> * At the time HFS+ was developed, there was no one common standard for \n> normalization. The HFS+ developers picked NFD because they thought it was \n> \"a more flexible, future-looking form\", but Microsoft ended up picking the \n> opposite just a short time later. Interestingly, NFC is a weird hybrid form \n> which only has composed forms for pre-existing characters, and decomposed \n> forms for all new characters (as they only have one form). So in a sense \n> NFD is more sane then NFC.\n\nNFC is better if you care about compatibility with existing legacy\ncharacter sets, where you want round-trip conversions to be\nidempotent.  On the other hand, given that Mac OS has historically\nnever cared about being compatible with the rest of the world, it\nmakes sense that it would choose NFD.\n\n> * The core issue here, which is why you think HFS+ is so stupid, is that \n> you guys see no problem with having 2 files \"Märchen\" (NFC) and \"Märchen\" \n> (NFD), whereas the HFS+ developers don't consider it acceptable to have 2 \n> visually identical names as independent files.\n\nYep.  No problems to do that.  You seem to think that supporting\nUnicode requires imposing this constraint, but that's simply not true,\nexcept maybe in some kind of religious sense.\n\n> Unfortunately, the only way \n> to do this matching is to store the normalized form in the filesystem, \n> because it would be a performance nightmare to try and do this matching any \n> other way.\n\nNope.  They were just not clever enough.  If they use a hashed key for\ntheir b-tree and used a hash which had the property that two strings\nthat were equivalent in the Unicode sense have the same hash value,\nit's quite possible to do Unicode-equivalence lookups quickly.  Yeah,\ncalculating the hash algorithm takes a bit amount of time, but it gets\ncalled no more than the normalization routine, and its performance\noverhead is no worse than the normalizing a string.\n\nI know how to do it in a Linux filesystem; it's just an insane thing\nto do, and so I choose not to do it.  But it is doable; if you must\npersue the course of filesystem insanity, it's possible to do it in a\nperformant way, without normalization; it's the same way that you can\nuse b-tree lookups in a case insensitive way.\n\n> I must say it is shocking that someone as smart as you is still more \n> interested in finding ways to prove me wrong then to actually address the \n> problem. It's obvious that the only research you did was intended to find \n> ways to call me stupid.\n\nNo, I did the research to try to find the HFS-specific filename\nmangling algorithm.  And given that's based on an back-level, old\nversion of Unicode, you can't just use NFD algorithm from the latest\nUnicode spec.  As I did that research, I came across the evidence that\nclaims you had made (i.e., that HFS had never changed the Unicode\nversion for its Normalization algorithm), was directly contradicted by\nthe Apple TechNote.\n\n    \t  \t\t\t\t\t- Ted\n"},{"id":"66355","messageId":"m17ii1fecy.fsf@ebiederm.dsl.xmission.com","threadId":"11645","inReplyTo":"7vr6gedgk9.fsf@gitster.siamese.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Eric W. Biederman","fromEmail":"ebiederm@xmission.com","sentAt":"2008-01-23T02:46:37Z","receivedAt":"2008-01-23T02:46:37Z","isPatch":false,"sender":{"key":"ebiederm@xmission.com","avatar":"https://avatars.githubusercontent.com/u/7477136?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> I'd rather see our mental bandwidth spent on coming up with a\n> workable workaround for such broken filesystems, while not\n> hurting use of git on sane platforms.\n>\n> I fear it might have to end up to be very messy and slow,\n> though.\n\nRandom thought.  Would it make sense to implement a git paranoid\nmode to autodetect name mangling.\n\nI.e.  After opening or creating a file by name we do a readdir in the\nsame directory to make certain we can find that same name/inode\ncombination.  Then on name-mangling systems we can autodetect they\nexist and limit ourselves to just what they don't mangle with no\nprior knowledge.  By refusing to process names that actively\nget mangled.   For small directories that you frequently see in\ndevelopment it shouldn't even be that slow.\n\nEric\n"},{"id":"66358","messageId":"7vsl0pp7u5.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"m17ii1fecy.fsf@ebiederm.dsl.xmission.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-23T02:57:22Z","receivedAt":"2008-01-23T02:57:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"ebiederm@xmission.com (Eric W. Biederman) writes:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n>\n>> I'd rather see our mental bandwidth spent on coming up with a\n>> workable workaround for such broken filesystems, while not\n>> hurting use of git on sane platforms.\n>>\n>> I fear it might have to end up to be very messy and slow,\n>> though.\n>\n> Random thought.  Would it make sense to implement a git paranoid\n> mode to autodetect name mangling.\n>\n> I.e.  After opening or creating a file by name we do a readdir in the\n> same directory to make certain we can find that same name/inode\n> combination.  Then on name-mangling systems we can autodetect they\n> exist and limit ourselves to just what they don't mangle with no\n> prior knowledge.  By refusing to process names that actively\n> get mangled.   For small directories that you frequently see in\n> development it shouldn't even be that slow.\n\nInside init-db where we already check how the filesystem\nbehaves, we could have an autodetection. A rough equivalent of\nwhat I had in mind is:\n\n\tmkdir -p \"Märchen/Märchen\"\n\tif test \"$(cd Märchen && echo M*)\" = \"Märchen\"\n        then\n        \t: not mangling\n\telse\n        \tgit config core.namemangle true\n\tfi\n\n(of course we do that in C not in shell).\n"},{"id":"66371","messageId":"20080123064139.GC16297@glandium.org","threadId":"11645","inReplyTo":"20080123013325.GB1320@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-01-23T06:41:39Z","receivedAt":"2008-01-23T06:41:39Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Tue, Jan 22, 2008 at 08:33:25PM -0500, Theodore Tso wrote:\n> On Tue, Jan 22, 2008 at 04:38:37PM -0800, Linus Torvalds wrote:\n> > One thing I'd like somebody to check: what _does_ happen with OS X and NFS \n> > (OS X as a client, not server)? In particular:\n> > \n> >  - Is it suddenly sane and case-sensitive?\n> \n> Using a Linux server, and a OS X client, over NFS, it is in\n> case-sensitive.  This is not unexpected, since you can mount UFS\n> partitions on Mac OS X, or reformat HFS+ filesystems and make them be\n> case-sensitive.\n> \n> >  - Does the NFS client do any unicode conversion?\n> \n> Nope:\n> \n> # perl -CO -e 'print pack(\"U\",0x00C4).\"\\n\"'  | xargs touch\n> # ls -l | cat -v\n> total 0\n> 0 -rw-r--r--   1 nobody  nobody  0 Jan 22 20:30 M-CM-^D\n> \n> It's pretty clear the Unicode conversion is being done in HFS+, not in\n> the VFS layer of Mac OS X.\n\nThere must be something at the VFS layer, or some other layer:\n- IIRC, Joliet iso9660 volumes end up being mounted with files names in\n  NFS when the real file names are NFC on the disk.\n- Likewise for Samba shares.\n- When I had my problems with iso9660 rockridge volumes using NFC (you\n  can create that just fine with mkisofs), the volume is mounted without\n  normalisation, i.e. if you get to a shell and want to access files,\n  you must use NFC, but at least the Finder does transliteration at some\n  stage, because going into the mount point and opening some files fail\n  because it's trying to open the file with the name transliterated to\n  NFD. I just hope the same doesn't happen with other filesystems.\n\nAlso, OSX using NFD widely, a file created from non Unix applications\nmay end up being named in NFD on any file system. File contents, too,\nmay end up being transliterated whenever a file is modified with non\nUnix applications, introducing unwanted changes.\nTyping file names in the Terminal might also make them encoded in NFD,\ntoo.\n\nMike\n"},{"id":"66375","messageId":"4697E0BA-7243-4C35-A384-0BD261EC21AF@sb.org","threadId":"11645","inReplyTo":"20080123064139.GC16297@glandium.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-23T08:15:02Z","receivedAt":"2008-01-23T08:15:02Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 23, 2008, at 1:41 AM, Mike Hommey wrote:\n\n> On Tue, Jan 22, 2008 at 08:33:25PM -0500, Theodore Tso wrote:\n>> On Tue, Jan 22, 2008 at 04:38:37PM -0800, Linus Torvalds wrote:\n>>> One thing I'd like somebody to check: what _does_ happen with OS X  \n>>> and NFS\n>>> (OS X as a client, not server)? In particular:\n>>>\n>>> - Is it suddenly sane and case-sensitive?\n>>\n>> Using a Linux server, and a OS X client, over NFS, it is in\n>> case-sensitive.  This is not unexpected, since you can mount UFS\n>> partitions on Mac OS X, or reformat HFS+ filesystems and make them be\n>> case-sensitive.\n>>\n>>> - Does the NFS client do any unicode conversion?\n>>\n>> Nope:\n>>\n>> # perl -CO -e 'print pack(\"U\",0x00C4).\"\\n\"'  | xargs touch\n>> # ls -l | cat -v\n>> total 0\n>> 0 -rw-r--r--   1 nobody  nobody  0 Jan 22 20:30 M-CM-^D\n>>\n>> It's pretty clear the Unicode conversion is being done in HFS+, not  \n>> in\n>> the VFS layer of Mac OS X.\n>\n> There must be something at the VFS layer, or some other layer:\n> - IIRC, Joliet iso9660 volumes end up being mounted with files names  \n> in\n>  NFS when the real file names are NFC on the disk.\n\nI assume you mean NFD, not NFS, but here's what one of the HFS+  \nengineers had to say:\n\n\"In Mac OS X,  SMB, MSDOS, UDF, ISO 9660 (Joliet), NTFS and ZFS file  \nsystems all store in one form -- NFC.  We store in NFC since that what  \nis expected for these files systems.\"\n\n> - Likewise for Samba shares.\n\nSee above.\n\n> - When I had my problems with iso9660 rockridge volumes using NFC (you\n>  can create that just fine with mkisofs), the volume is mounted  \n> without\n>  normalisation, i.e. if you get to a shell and want to access files,\n>  you must use NFC, but at least the Finder does transliteration at  \n> some\n>  stage, because going into the mount point and opening some files fail\n>  because it's trying to open the file with the name transliterated to\n>  NFD. I just hope the same doesn't happen with other filesystems.\n\nCan you produce a reproducible set of steps for this? Because the  \nFinder shouldn't be doing any of this work on its own, all the  \nnormalization stuff happens directly in HFS+.\n\n> Also, OSX using NFD widely, a file created from non Unix applications\n> may end up being named in NFD on any file system. File contents, too,\n> may end up being transliterated whenever a file is modified with non\n> Unix applications, introducing unwanted changes.\n> Typing file names in the Terminal might also make them encoded in NFD,\n> too.\n\nEntirely possible, though renormalizing file contents seems a bit less  \nlikely. I will point out that the text input system in OS X seems to  \ndefault to producing NFC (at least, typing `echo 'Märchen' | xxd` in  \nthe Terminal shows that the input string there is NFC). So user input  \nwill most likely produce NFC, the only way you're probably going to  \nend up with NFD is if you move a file from HFS+ to another filesystem.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66380","messageId":"20080123084345.GN14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"4697E0BA-7243-4C35-A384-0BD261EC21AF@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-23T08:43:45Z","receivedAt":"2008-01-23T08:43:45Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Wed, Jan 23, 2008 at 03:15:02AM -0500, Kevin Ballard wrote:\n> \n> Entirely possible, though renormalizing file contents seems a bit less  \n> likely. I will point out that the text input system in OS X seems to  \n> default to producing NFC (at least, typing `echo 'Märchen' | xxd` in  \n> the Terminal shows that the input string there is NFC).\n\nI wonder what happens if you do this:\n\ntouch 'Märchen'\necho M*rchen | xxd -g1\n\nWill that produce NFC or NFD?\n\nDmitry\n"},{"id":"66381","messageId":"85tzl5udzp.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"E6F76F93-24C9-4D10-813C-770A9C3A9828@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-23T08:45:30Z","receivedAt":"2008-01-23T08:45:30Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Kevin Ballard <kevin@sb.org> writes:\n\n> I just finished talking to one of the HFS+ developers, so I suspect I\n> know a lot more on this subject now than you do.\n\nUh, Ted is a filesystem developer.  I can't count the hours I spent\ntalking with my father, a theoretical physicist, but that does not make\nme qualified to consider myself a better authority on physics than a\nsub-average actual grad student of the matter.\n\nIf you don't manage to check your arrogance eventually, you'll be\ncausing more damage to your cause than if you just shut up.  You make\nabundantly clear that you don't understand the _implications_ of the\ndetails you may or not may happen to find out.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66382","messageId":"57518fd10801230102n4f430219p2701c7561f184569@mail.gmail.com","threadId":"11645","inReplyTo":"20080123084345.GN14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Jonathan del Strother","fromEmail":"maillist@steelskies.com","sentAt":"2008-01-23T09:02:43Z","receivedAt":"2008-01-23T09:02:43Z","isPatch":false,"sender":{"key":"jon.delstrother@bestbefore.tv","avatar":"https://gravatar.com/avatar/754e21ab701c00e2d21fc261187254c34b2a1c0b959d9ee5be1a295990be3081?d=mp&s=160"},"body":"On Jan 23, 2008 8:43 AM, Dmitry Potapov <dpotapov@gmail.com> wrote:\n> On Wed, Jan 23, 2008 at 03:15:02AM -0500, Kevin Ballard wrote:\n> >\n> > Entirely possible, though renormalizing file contents seems a bit less\n> > likely. I will point out that the text input system in OS X seems to\n> > default to producing NFC (at least, typing `echo 'Märchen' | xxd` in\n> > the Terminal shows that the input string there is NFC).\n>\n> I wonder what happens if you do this:\n>\n> touch 'Märchen'\n> echo M*rchen | xxd -g1\n>\n> Will that produce NFC or NFD?\n>\n\n0000000: 4d 61 cc 88 72 63 68 65 6e 0a                    Ma..rchen.\n"},{"id":"66383","messageId":"20080123091240.GO14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"57518fd10801230102n4f430219p2701c7561f184569@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-23T09:12:40Z","receivedAt":"2008-01-23T09:12:40Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Wed, Jan 23, 2008 at 09:02:43AM +0000, Jonathan del Strother wrote:\n> On Jan 23, 2008 8:43 AM, Dmitry Potapov <dpotapov@gmail.com> wrote:\n> > On Wed, Jan 23, 2008 at 03:15:02AM -0500, Kevin Ballard wrote:\n> > >\n> > > Entirely possible, though renormalizing file contents seems a bit less\n> > > likely. I will point out that the text input system in OS X seems to\n> > > default to producing NFC (at least, typing `echo 'Märchen' | xxd` in\n> > > the Terminal shows that the input string there is NFC).\n> >\n> > I wonder what happens if you do this:\n> >\n> > touch 'Märchen'\n> > echo M*rchen | xxd -g1\n> >\n> > Will that produce NFC or NFD?\n> >\n> \n> 0000000: 4d 61 cc 88 72 63 68 65 6e 0a                    Ma..rchen.\n\nThis is NFC! Did you do that on HFS+?\n\nIf so, it means that shell on Mac also converts filenames to NFC when\nit reads them from the disk.\n\nDmitry\n"},{"id":"66385","messageId":"20080123091958.GA6969@glandium.org","threadId":"11645","inReplyTo":"20080123091240.GO14871@dpotapov.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-01-23T09:19:59Z","receivedAt":"2008-01-23T09:19:59Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Wed, Jan 23, 2008 at 12:12:40PM +0300, Dmitry Potapov <dpotapov@gmail.com> wrote:\n> On Wed, Jan 23, 2008 at 09:02:43AM +0000, Jonathan del Strother wrote:\n> > On Jan 23, 2008 8:43 AM, Dmitry Potapov <dpotapov@gmail.com> wrote:\n> > > On Wed, Jan 23, 2008 at 03:15:02AM -0500, Kevin Ballard wrote:\n> > > >\n> > > > Entirely possible, though renormalizing file contents seems a bit less\n> > > > likely. I will point out that the text input system in OS X seems to\n> > > > default to producing NFC (at least, typing `echo 'Märchen' | xxd` in\n> > > > the Terminal shows that the input string there is NFC).\n> > >\n> > > I wonder what happens if you do this:\n> > >\n> > > touch 'Märchen'\n> > > echo M*rchen | xxd -g1\n> > >\n> > > Will that produce NFC or NFD?\n> > >\n> > \n> > 0000000: 4d 61 cc 88 72 63 68 65 6e 0a                    Ma..rchen.\n> \n> This is NFC! Did you do that on HFS+?\n\nNFD, you mean ?\n\nMike\n"},{"id":"66386","messageId":"20080123093242.GQ14871@dpotapov.dyndns.org","threadId":"11645","inReplyTo":"20080123091958.GA6969@glandium.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-23T09:32:42Z","receivedAt":"2008-01-23T09:32:42Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Wed, Jan 23, 2008 at 10:19:59AM +0100, Mike Hommey wrote:\n> On Wed, Jan 23, 2008 at 12:12:40PM +0300, Dmitry Potapov <dpotapov@gmail.com> wrote:\n> > On Wed, Jan 23, 2008 at 09:02:43AM +0000, Jonathan del Strother wrote:\n> > > On Jan 23, 2008 8:43 AM, Dmitry Potapov <dpotapov@gmail.com> wrote:\n> > > > On Wed, Jan 23, 2008 at 03:15:02AM -0500, Kevin Ballard wrote:\n> > > > >\n> > > > > Entirely possible, though renormalizing file contents seems a bit less\n> > > > > likely. I will point out that the text input system in OS X seems to\n> > > > > default to producing NFC (at least, typing `echo 'Märchen' | xxd` in\n> > > > > the Terminal shows that the input string there is NFC).\n> > > >\n> > > > I wonder what happens if you do this:\n> > > >\n> > > > touch 'Märchen'\n> > > > echo M*rchen | xxd -g1\n> > > >\n> > > > Will that produce NFC or NFD?\n> > > >\n> > > \n> > > 0000000: 4d 61 cc 88 72 63 68 65 6e 0a                    Ma..rchen.\n> > \n> > This is NFC! Did you do that on HFS+?\n> \n> NFD, you mean ?\n\nOops, you are right.\n\nDmitry\n"},{"id":"66387","messageId":"20080123094052.GB6969@glandium.org","threadId":"11645","inReplyTo":"4697E0BA-7243-4C35-A384-0BD261EC21AF@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2008-01-23T09:40:52Z","receivedAt":"2008-01-23T09:40:52Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Wed, Jan 23, 2008 at 03:15:02AM -0500, Kevin Ballard <kevin@sb.org> wrote:\n> \"In Mac OS X,  SMB, MSDOS, UDF, ISO 9660 (Joliet), NTFS and ZFS file  \n> systems all store in one form -- NFC.  We store in NFC since that what  \n> is expected for these files systems.\"\n\nThat's the point. It's stored in NFC, but what applications see is NFD.\n\n> >- Likewise for Samba shares.\n> \n> See above.\n> \n> >- When I had my problems with iso9660 rockridge volumes using NFC (you\n> > can create that just fine with mkisofs), the volume is mounted  \n> >without\n> > normalisation, i.e. if you get to a shell and want to access files,\n> > you must use NFC, but at least the Finder does transliteration at  \n> >some\n> > stage, because going into the mount point and opening some files fail\n> > because it's trying to open the file with the name transliterated to\n> > NFD. I just hope the same doesn't happen with other filesystems.\n> \n> Can you produce a reproducible set of steps for this? Because the  \n> Finder shouldn't be doing any of this work on its own, all the  \n> normalization stuff happens directly in HFS+.\n\nSimple : on a Linux host, create files with NFC names, and create an iso\nimage with mkisofs, with rockridge but no joliet. Burn this to a disc, and\ninsert the disc in your OSX host, and try to open files from the finder.\nInterestingly, IIRC, Finder is able to copy the files, though.\n\nAs a bonus, try the same with an iso volume name in NFC, it's even better :\nthe created mount point is NFD, but it tries to mount on the name in NFC and\nfails. And then you just can't eject the CD anymore.\n\nMike\n"},{"id":"66397","messageId":"20080123133802.GC7415@mit.edu","threadId":"11645","inReplyTo":"20080123094052.GB6969@glandium.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-23T13:38:02Z","receivedAt":"2008-01-23T13:38:02Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"Here's a reliable test case to test filename normalization on Mac OS.\n\n------ cut here -------\ncat > test.pl << EOF\n#!/usr/bin/perl -CO\nprint \"M\".pack(\"U\",0x00E4).\"rchen\\n\";\nprint \"Ma\".pack(\"U\",0x0308).\"rchen\\n\";\nEOF\nchmod +x test.pl\n./test.pl | xargs touch\necho M* | xxd -g1\n------ cut here -------\n\nOn an NFS mounted filesystem, what you will get is this:\n\n0000000: 4d 61 cc 88 72 63 68 65 6e 20 4d c3 a4 72 63 68  Ma..rchen M..rch\n0000010: 65 6e 0a                                         en.\n\nand on an HFS+ mounted filesystem, what you will get is this:\n\n0000000: 4d 61 cc 88 72 63 68 65 6e 0a                    Ma..rchen.\n\nSo this demonstrates that on my MacOS 10.4.11 system, on NFS, MacOS is\ndoing no normalization, as it is creating two files.  On HFS+, MacOS\nis mapping both filenames to the same decomposed name.\n\nMore (or not) surprisingly, given Kevin Ballard's \"reliable source\":\n\n  \"In Mac OS X,  SMB, MSDOS, UDF, ISO 9660 (Joliet), NTFS and ZFS file\n  systems all store in one form -- NFC.  We store in NFC since that what is\n  expected for these files systems.\"\n\nUsing a Sony Reader (which uses an internal FAT filesystem) hooked up\nto a MacOS 10.4.11 system:\n\n% /fs/u1/tmp/test.pl  | xargs touch\n% echo M* | xxd -g1\n0000000: 4d 61 cc 88 72 63 68 65 6e 0a                    Ma..rchen.\n\n.. which is the decomposed form.  So it looks like on FAT/MSDOS\nfilesystems MacOS 10.4.11 normalizes files to NFD, which will *not* do\nthe right thing as far as Windows compatibility is concerned on USB\nsticks, et. al.  Mac OS users would be well advised not to use\nnon-ASCII names in their filesystems if they care about interoperating\nwith other systems.  :-P\n\n\t\t\t\t\t\t\t- Ted\n"},{"id":"66401","messageId":"alpine.LFD.1.00.0801230923450.22568@xanadu.home","threadId":"11645","inReplyTo":"7vsl0pp7u5.fsf@gitster.siamese.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-01-23T14:26:06Z","receivedAt":"2008-01-23T14:26:06Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 22 Jan 2008, Junio C Hamano wrote:\n\n> ebiederm@xmission.com (Eric W. Biederman) writes:\n> \n> > Random thought.  Would it make sense to implement a git paranoid\n> > mode to autodetect name mangling.\n> >\n> > I.e.  After opening or creating a file by name we do a readdir in the\n> > same directory to make certain we can find that same name/inode\n> > combination.  Then on name-mangling systems we can autodetect they\n> > exist and limit ourselves to just what they don't mangle with no\n> > prior knowledge.  By refusing to process names that actively\n> > get mangled.   For small directories that you frequently see in\n> > development it shouldn't even be that slow.\n> \n> Inside init-db where we already check how the filesystem\n> behaves, we could have an autodetection.\n\nI wonder if that is good enough.  Git repositories can be copied over to \ndifferent filesystems.\n\n\nNicolas\n"},{"id":"66408","messageId":"alpine.LFD.1.00.0801230808440.1741@woody.linux-foundation.org","threadId":"11645","inReplyTo":"20080123133802.GC7415@mit.edu","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-23T16:16:33Z","receivedAt":"2008-01-23T16:16:33Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 23 Jan 2008, Theodore Tso wrote:\n> \n> So this demonstrates that on my MacOS 10.4.11 system, on NFS, MacOS is\n> doing no normalization, as it is creating two files.  On HFS+, MacOS\n> is mapping both filenames to the same decomposed name.\n\nWell, it demonstrates that (a) the OS and (b) _perl_ don't mangle \nfilenames on non-HFS+ filesystems.\n\nThe problem is that since most native applications *expect* that name \nmangling, they'll probably do name mangling of their own (internally) just \nto compare the names!\n\nSo I would not be surprised if the globbing libraries, for example, will \ndo NFD-mangling in order to glob \"correctly\", so even programs ported from \nreal Unix might end up getting pathnames subtly changed into NFD as part \nof some hot library-on-library action with UTF hackery inside.\n\nThings like the finder etc, which must be very aware of the fact that \nfilenames get corrupted, would presumably internally always convert \neverything they get into NFD in order to compare names from different \nsources. And as part of that, programs may well corrupt the name before \nthey then use it to create a pathname.\n\nThe fact that your perl program works under NFS, but creates NFD on a VFAT \nvolume, does imply that they probably used at least some of the same \nroutines they use in HFS+ for VFAT. Not entirely surprising: doing case \ninsensitive stuff with Unicode is nasty code, so why not share it (even if \nit's then incorrect for FAT)..\n\nPiece of crap it is, though. Apple has painted themselves into a nasty \ncorner there.\n\n\t\t\tLinus\n"},{"id":"66413","messageId":"769D0E1A-8399-458B-8328-FF3642D833BC@sb.org","threadId":"11645","inReplyTo":"20080123094052.GB6969@glandium.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-23T16:58:15Z","receivedAt":"2008-01-23T16:58:15Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 23, 2008, at 4:40 AM, Mike Hommey wrote:\n\n> On Wed, Jan 23, 2008 at 03:15:02AM -0500, Kevin Ballard  \n> <kevin@sb.org> wrote:\n>> \"In Mac OS X,  SMB, MSDOS, UDF, ISO 9660 (Joliet), NTFS and ZFS file\n>> systems all store in one form -- NFC.  We store in NFC since that  \n>> what\n>> is expected for these files systems.\"\n>\n> That's the point. It's stored in NFC, but what applications see is  \n> NFD.\n\nI was actually asking for you to show this instead of just asserting  \nit, but I realized I have access to an SMB share myself so I just  \ntested.\n\nAnd you're right. That's very curious. I guess they did that because  \nthe entire Carbon stack was written assuming NFD (back at the same  \ntime HFS+ was created), and they wanted to provide a consistent  \ninterface to applications. Since the filesystem already uses NFC,  \nrenormalizing to NFD shouldn't lose anything (want the original  \nrepresentation back? just normalize back to NFC).\n\n>>> - Likewise for Samba shares.\n>>\n>> See above.\n>>\n>>> - When I had my problems with iso9660 rockridge volumes using NFC  \n>>> (you\n>>> can create that just fine with mkisofs), the volume is mounted\n>>> without\n>>> normalisation, i.e. if you get to a shell and want to access files,\n>>> you must use NFC, but at least the Finder does transliteration at\n>>> some\n>>> stage, because going into the mount point and opening some files  \n>>> fail\n>>> because it's trying to open the file with the name transliterated to\n>>> NFD. I just hope the same doesn't happen with other filesystems.\n>>\n>> Can you produce a reproducible set of steps for this? Because the\n>> Finder shouldn't be doing any of this work on its own, all the\n>> normalization stuff happens directly in HFS+.\n>\n> Simple : on a Linux host, create files with NFC names, and create an  \n> iso\n> image with mkisofs, with rockridge but no joliet. Burn this to a  \n> disc, and\n> insert the disc in your OSX host, and try to open files from the  \n> finder.\n> Interestingly, IIRC, Finder is able to copy the files, though.\n>\n> As a bonus, try the same with an iso volume name in NFC, it's even  \n> better :\n> the created mount point is NFD, but it tries to mount on the name in  \n> NFC and\n> fails. And then you just can't eject the CD anymore.\n\nI was actually hoping for something I could test myself.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66415","messageId":"20080123171258.GB32663@mit.edu","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801230808440.1741@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2008-01-23T17:12:58Z","receivedAt":"2008-01-23T17:12:58Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Wed, Jan 23, 2008 at 08:16:33AM -0800, Linus Torvalds wrote:\n> \n> \n> On Wed, 23 Jan 2008, Theodore Tso wrote:\n> > \n> > So this demonstrates that on my MacOS 10.4.11 system, on NFS, MacOS is\n> > doing no normalization, as it is creating two files.  On HFS+, MacOS\n> > is mapping both filenames to the same decomposed name.\n> \n> Well, it demonstrates that (a) the OS and (b) _perl_ don't mangle \n> filenames on non-HFS+ filesystems.\n\nWell \"touch\" actually since that was what was actually creating the\nfiles; I only used perl because it was easist way to gaurantee exactly\nhow the filenames would be generated.\n\n> The problem is that since most native applications *expect* that name \n> mangling, they'll probably do name mangling of their own (internally) just \n> to compare the names!\n> \n> So I would not be surprised if the globbing libraries, for example, will \n> do NFD-mangling in order to glob \"correctly\", so even programs ported from \n> real Unix might end up getting pathnames subtly changed into NFD as part \n> of some hot library-on-library action with UTF hackery inside.\n\nIt's worse than that.  You can specify at format time whether or not\nHFS+ does case-sensitivity or not, and of course, there is UFS, which\nI expect does no Unicode normalization at all, much like NFS.  I\nsuspect what you've pointed out is why certain MacOS programs break\nhorribly when run on non-HFS+ filesystems, though.  And if that is the\ncase, then those same programs might not be reliable if the user's\nhome directory is stored on NFS --- like they would be in an\nenteprise/corproate environment, if Apple ever wants to have any hope\nof penetrating that market.\n\nBecause of this, git code won't be able to just check for HFS+; it\nwill probably have to do a run-time test to see whether or not the\nfilesystem is doing case-folding or not, since that can be turned on\nor off on a per-filesystem basis.  Also unknown, and which should be\ntested, is whether turning off case-folding also turns off Unicode\nnormalization.  It may be that they did this so that HFS+ could be UFS\ncompatible, since Darwin *must* be built on a UFS filesystem,\nreflecting its Mach/BSD heritage.  (I ran across this while doing my\nweb research; apparently HFS+ has been causing Apple headaches\ninternally.  Heh.  :-)\n\n>Things like the finder etc, which must be very aware of the fact that\n>filenames get corrupted, would presumably internally always convert\n>everything they get into NFD in order to compare names from different\n>sources. And as part of that, programs may well corrupt the name before\n>they then use it to create a pathname.\n\nWell, hopefully not everyone inside Apple's OS groups are total\nmorons, and actually use a utf8_str_equiv() routine instead of\nstrcmp() to do their Unicode comparisons.  But then again, maybe\nnot...\n\n> The fact that your perl program works under NFS, but creates NFD on a VFAT \n> volume, does imply that they probably used at least some of the same \n> routines they use in HFS+ for VFAT. Not entirely surprising: doing case \n> insensitive stuff with Unicode is nasty code, so why not share it (even if \n> it's then incorrect for FAT)..\n> \n> Piece of crap it is, though. Apple has painted themselves into a nasty \n> corner there.\n\nNo kidding!!\n\n\t\t\t\t\t\t\t- Ted\n"},{"id":"66417","messageId":"98F90EB6-1930-4643-8C6C-CA11CB123BAA@sb.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801230808440.1741@woody.linux-foundation.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-23T17:19:02Z","receivedAt":"2008-01-23T17:19:02Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 23, 2008, at 11:16 AM, Linus Torvalds wrote:\n\n> On Wed, 23 Jan 2008, Theodore Tso wrote:\n>>\n>> So this demonstrates that on my MacOS 10.4.11 system, on NFS, MacOS  \n>> is\n>> doing no normalization, as it is creating two files.  On HFS+, MacOS\n>> is mapping both filenames to the same decomposed name.\n>\n> Well, it demonstrates that (a) the OS and (b) _perl_ don't mangle\n> filenames on non-HFS+ filesystems.\n>\n> The problem is that since most native applications *expect* that name\n> mangling, they'll probably do name mangling of their own  \n> (internally) just\n> to compare the names!\n\nWell yes, any context in which a string is treated as Unicode instead  \nof an opaque sequence of bytes will probably lead to normalization at  \nsome point (e.g. when searching text, I'm going to want Märchen and  \nMärchen to be treated as the same string). The Mac OS X APIs use NFD,  \nand everybody else uses NFC, but either way it's still normalization.\n\n> So I would not be surprised if the globbing libraries, for example,  \n> will\n> do NFD-mangling in order to glob \"correctly\", so even programs  \n> ported from\n> real Unix might end up getting pathnames subtly changed into NFD as  \n> part\n> of some hot library-on-library action with UTF hackery inside.\n\nWhy would the globbing libraries have to do anything special to  \nunderstand NFD? In fact, I prefer that they don't - it's very handy to  \nbe able to type Ma* and have that match Märchen, as the globbing  \nlibrary sees Ma??rchen and is happy to match the ??rchen against *.  \nWere the filename in NFC, I couldn't do that. Similarly, Ma<tab>  \nautocompletes the name Märchen for me. But the convenience is beside  \nthe point - what I'm trying to show here is that if the globbing  \nlibrary were NFD-aware, it probably would decide Ma* shouldn't match  \nMärchen, right?\n\nI assume globbing libraries et al don't do UTF-8 hackery in Linux,  \nright? And yet using NFC-encoded filenames is fairly common? So why  \nshould it be any different on OS X, especially since HFS+ isn't the  \nonly option here (and thus doing NFD conversion in the library would  \nmess up other filesystems)?\n\nIn fact, probably the biggest reason the NFD-encoding was done at the  \nHFS+ level is because they simply couldn't trust user-level libraries  \nto always do the NFD conversion for pathnames. And I quote:\n\n\"I would prefer that case sensitivity and unicode normalization were  \nnot the responsibility of the file system -- but I realize that we  \ncannot just ignore the problem and let the other layers sort it all  \nout.\"\n\n> Things like the finder etc, which must be very aware of the fact that\n> filenames get corrupted, would presumably internally always convert\n> everything they get into NFD in order to compare names from different\n> sources. And as part of that, programs may well corrupt the name  \n> before\n> they then use it to create a pathname.\n\nI don't get why you're still calling it corruption when, on an HFS+  \nsystem, NFD-encoding is correct. It would be corruption for HFS+ to  \nwrite anything else but NFD.\n\n> The fact that your perl program works under NFS, but creates NFD on  \n> a VFAT\n> volume, does imply that they probably used at least some of the same\n> routines they use in HFS+ for VFAT. Not entirely surprising: doing  \n> case\n> insensitive stuff with Unicode is nasty code, so why not share it  \n> (even if\n> it's then incorrect for FAT)..\n>\n> Piece of crap it is, though. Apple has painted themselves into a nasty\n> corner there.\n\nThere's no reason to assume that OS X is actually storing the NFD on  \nthe volume. In fact, it's quite explicitly not:\n\n\"As far as storing exactly what was passed in,  its not just HFS  \nthat's involved her.  In Mac OS X,  SMB, MSDOS, UDF, ISO 9660  \n(Joliet), NTFS and ZFS file systems all store in one form -- NFC.  We  \nstore in NFC since that what is expected for these files systems.  If  \nwe were to allow KFD to pass through, it would cause problems when  \nthese names were accessed outside of Mac OS X.  So this is not just an  \nHFS issue but an interchange issue for Mac OS X.  We have the legacy  \nNFD use/expectation in our applications and we chose not to ignore the  \nproblem but make a conscience effort to have the appropriate form used  \n(NFD in Mac OS X APIs, NFC elsewhere).  Its not perfect but neither is  \nthe agnostic approach where both forms can be used and you can have  \nduplicate filenames in your file system.\"\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66420","messageId":"alpine.LFD.1.00.0801230930390.1741@woody.linux-foundation.org","threadId":"11645","inReplyTo":"98F90EB6-1930-4643-8C6C-CA11CB123BAA@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-01-23T17:32:11Z","receivedAt":"2008-01-23T17:32:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 23 Jan 2008, Kevin Ballard wrote:\n> \n> Well yes, any context in which a string is treated as Unicode instead of an\n> opaque sequence of bytes will probably lead to normalization at some point\n> (e.g. when searching text, I'm going to want Märchen and Märchen to be treated\n> as the same string).\n\nAs pointed out (multiple times), this is only true if the programmer is a \nmoron.\n\nYou do not need to - and *should* not - convert to a common normalization \nin order to compare to Uncode strings. You should just compare them with a \nUnicode-aware comparison routine. It will be faster, and it will avoid \ncorrupting the input.\n\nSadly, stupid people are much too common.\n\n\t\tLinus\n"},{"id":"66421","messageId":"37fcd2780801230939q6d30488bwc78eaeb9eb01989f@mail.gmail.com","threadId":"11645","inReplyTo":"769D0E1A-8399-458B-8328-FF3642D833BC@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-01-23T17:39:06Z","receivedAt":"2008-01-23T17:39:06Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Jan 23, 2008 7:58 PM, Kevin Ballard <kevin@sb.org> wrote:\n> On Jan 23, 2008, at 4:40 AM, Mike Hommey wrote:\n> >\n> > That's the point. It's stored in NFC, but what applications see is\n> > NFD.\n>\n> I was actually asking for you to show this instead of just asserting\n> it, but I realized I have access to an SMB share myself so I just\n> tested.\n>\n> And you're right. That's very curious. I guess they did that because\n> the entire Carbon stack was written assuming NFD (back at the same\n> time HFS+ was created), and they wanted to provide a consistent\n> interface to applications.\n\nWait, did you tell us some time ago that normalization does not\nmatter and you just need to treat strings \"as text\"? Now, it looks\nlike the Carbon stack does not treat strings \"as text\". How come?\n\nMaybe, you should stop lying and admit that changing Unicode\nstrings does matter even if they remain equivalent.\n\n> Since the filesystem already uses NFC,\n> renormalizing to NFD shouldn't lose anything (want the original\n> representation back? just normalize back to NFC).\n\nOn Windows, you can create two *different* files -- one with NFC\nand the other with NFD name. I wonder, how it is going to work\nwith your renormalization back and force.\n\nDmitry\n"},{"id":"66422","messageId":"7D245A44-AE30-42D7-BD64-AC53B0543C59@sb.org","threadId":"11645","inReplyTo":"37fcd2780801230939q6d30488bwc78eaeb9eb01989f@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-23T17:47:50Z","receivedAt":"2008-01-23T17:47:50Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 23, 2008, at 12:39 PM, Dmitry Potapov wrote:\n\n> On Jan 23, 2008 7:58 PM, Kevin Ballard <kevin@sb.org> wrote:\n>> On Jan 23, 2008, at 4:40 AM, Mike Hommey wrote:\n>>>\n>>> That's the point. It's stored in NFC, but what applications see is\n>>> NFD.\n>>\n>> I was actually asking for you to show this instead of just asserting\n>> it, but I realized I have access to an SMB share myself so I just\n>> tested.\n>>\n>> And you're right. That's very curious. I guess they did that because\n>> the entire Carbon stack was written assuming NFD (back at the same\n>> time HFS+ was created), and they wanted to provide a consistent\n>> interface to applications.\n>\n> Wait, did you tell us some time ago that normalization does not\n> matter and you just need to treat strings \"as text\"? Now, it looks\n> like the Carbon stack does not treat strings \"as text\". How come?\n\nI'm amazed at how badly you manage to misinterpret everything I say.\n\n>> Since the filesystem already uses NFC,\n>> renormalizing to NFD shouldn't lose anything (want the original\n>> representation back? just normalize back to NFC).\n>\n> On Windows, you can create two *different* files -- one with NFC\n> and the other with NFD name. I wonder, how it is going to work\n> with your renormalization back and force.\n\nI'm not sure what you're trying to say here. As near as I can tell,  \nSMB already does encoding conversions itself when talking to different  \nclients, so you can hardly say OS X is doing something bad by  \nconverting between local NFD and NFC on SMB.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66425","messageId":"76718490801231218i53c19e22lda34f2eec88627f8@mail.gmail.com","threadId":"11645","inReplyTo":"98F90EB6-1930-4643-8C6C-CA11CB123BAA@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Jay Soffian","fromEmail":"jaysoffian@gmail.com","sentAt":"2008-01-23T20:18:33Z","receivedAt":"2008-01-23T20:18:33Z","isPatch":false,"sender":{"key":"jaysoffian@gmail.com","avatar":"https://avatars.githubusercontent.com/u/155970?v=4"},"body":"On 1/23/08, Kevin Ballard <kevin@sb.org> wrote:\n>\n> I don't get why you're still calling it corruption when, on an HFS+\n> system, NFD-encoding is correct. It would be corruption for HFS+ to\n> write anything else but NFD.\n\nHow about this: it's lossy. It's lossy in a similar sense that TIFF ->\nJPEG -> TIFF doesn't give you back exactly the same bytes, even though\n(modulo the compression level) the two TIFFs might be visually\nindistinguishable.\n\nYou seem to have an issue with calling this \"corruption\", but to most\nof us, if you have a system where you don't get back *exactly the same\ndata* that you put in, then the data has been corrupted.\n\nNow, please stop trolling this point, agree to disagree, and either\ncontribute some code or be quiet and allow others to make progress.\n\nj.\n"},{"id":"66430","messageId":"7vve5kgrym.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801230923450.22568@xanadu.home","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-23T21:19:45Z","receivedAt":"2008-01-23T21:19:45Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> On Tue, 22 Jan 2008, Junio C Hamano wrote:\n>\n>> ebiederm@xmission.com (Eric W. Biederman) writes:\n>> \n>> > Random thought.  Would it make sense to implement a git paranoid\n>> > mode to autodetect name mangling.\n>> >\n>> > I.e.  After opening or creating a file by name we do a readdir in the\n>> > same directory to make certain we can find that same name/inode\n>> > combination.  Then on name-mangling systems we can autodetect they\n>> > exist and limit ourselves to just what they don't mangle with no\n>> > prior knowledge.  By refusing to process names that actively\n>> > get mangled.   For small directories that you frequently see in\n>> > development it shouldn't even be that slow.\n>> \n>> Inside init-db where we already check how the filesystem\n>> behaves, we could have an autodetection.\n>\n> I wonder if that is good enough.  Git repositories can be copied over to \n> different filesystems.\n\nDo you mean \"cp -a\"?  If I am not mistaken we already have that\nissue, due to core.filemode, when user does that across\nfilesystems with different behaviours.\n\nThere is not much we can do against \"cp -a\" other than telling\nusers that some configurations need to be adjusted.\n"},{"id":"66433","messageId":"46a038f90801231537j6899a1a2t57cd5af9449267f9@mail.gmail.com","threadId":"11645","inReplyTo":"98F90EB6-1930-4643-8C6C-CA11CB123BAA@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-23T23:37:26Z","receivedAt":"2008-01-23T23:37:26Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 24, 2008 6:19 AM, Kevin Ballard <kevin@sb.org> wrote:\n> I don't get why you're still calling it corruption\n\nBecause in a modern Internet-aware world, whoever designs a FS needs\nto acknowledge that they will need to store files from other systems\nthat have other assumptions. That is, if they want to interoperate.\n\nAs you noted not long ago, it is a serious problem if an HFS+\npartition is shared over NFS. If you look at all the apps that have\nproblems with this aspect of HFS+ , they are all apps that transfer\nfiles over the network over diverse protocols. That's why it's a\nproblem with git, because the files may be coming from a different\nmachine, running any arbitrary OS that git supports.\n\nIn such scenario, can you understand why everyone is saying that HFS+\nand the VFS should not mangle names, even if it makes sense to some\nuse cases under OSX? And do you understand why the same applies to\ngit, being a network-sharing-oriented app?\n\nSo -- if OSX was doing things to make it easier for users to find\nmatching files at the Finder level, that'd be _fine_. But the FS has\nto deal with a lot more variety than that. So this is a bad design\ndecision -- perhaps less obvious under OS8/9, but completely\ndisastrous with a network OS such as OSX. Call it \"different\" if you\nwant, but that's a euphemism for \"wrong\".\n\ncheers,\n\n\n\nm\n"},{"id":"66438","messageId":"DE7B2DE6-03B1-4781-92C7-096E591369A1@sb.org","threadId":"11645","inReplyTo":"76718490801231517h6d57e5bfkc19d394d38ad19db@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-24T02:05:50Z","receivedAt":"2008-01-24T02:05:50Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"I hope you don't mind that I'm redirecting this back onto the list.\n\nOn Jan 23, 2008, at 6:17 PM, Jay Soffian wrote:\n\n> On 1/23/08, Kevin Ballard <kevin@sb.org> wrote:\n>\n>> I agree - the argument is fairly worthless. So why does everybody  \n>> else\n>> keep spending time accusing HFS+ of corrupting filenames? Most of\n>> Linus's email was specifically about this point, but apparently  \n>> that's\n>> alright with you while only a *single* line out of my email in direct\n>> response is called trolling?\n>\n> Everyone else considers what HFS+ does as corruption. You, alone in\n> this thread, do not. You are not willing to concede the point, nor let\n> it go. I'm accusing you of trolling because you are the single person\n> defending HFS+'s behavior.\n\nI don't understand how you can possibly think that disagreeing ==  \ntrolling. Similarly, just because I'm the only person *on this list*  \nwho holds my viewpoint doesn't in any way mean I should abandon it. In  \nfact, it makes it much more important that I continue to stand up for  \nwhat I believe, The whole notion of democracy is based on the fact  \nthat every person is important, and that every person has the right to  \ntheir own opinion. I realize this is a mailing list, not a democratic  \nbody, but the same principles should still apply. If your criteria for  \njudging any viewpoint is purely how many people hold that viewpoint,  \nthen you end up ignoring things just because they are different or new.\n\nI may be the single person defending this behavior on this list, but  \nif you were to leave your comfortable linux community and talk to  \npeople elsewhere, you might find yourself in the minority opinion.\n\n> (Also, while it's certainly possible that you are right and everyone\n> else is wrong, most of the other folks have significant experience as\n> kernel, filesystem, or git developers, which leads credence to their\n> point -- reputation matters.)\n\nWhy do you persist in thinking of this as right vs. wrong? I've tried  \nto emphasize, many times, that HFS+ behaves this way not because it's  \n\"right\" and ext4 is \"wrong\", but because HFS+ has a different set of  \nvalues. The developers of HFS+ believed that, for a consumer OS like  \nOS X, it made much more sense to treat visually indistinguishable  \nfilenames as the same file. I, and I'm sure the vast majority of OS X  \nusers, agree. Unfortunately this decision had some drawbacks, but they  \nfelt the trade-off was worth it. I'm well aware that you all don't  \nthink the trade-off was worth it, but like I said, this is a matter of  \nbehaving differently due to a different set of values, not behaving  \n\"right\" or \"wrong\". I've been making an attempt to agree to disagree,  \nbut it seems that you would rather just squash dissent instead of  \naccepting it.\n\n>> Again, you're happy to let everybody else write long paragraphs\n>> accusing HFS+ of bad behavior (and making horrible assumptions which\n>> are generally completely untrue), and you don't think that's noise?\n>\n> They are responding to you. If you let the point drop, so will they.\n\nI did let the point drop. Then you guys resurrected it. You can't pin  \nthis one on me.\n\n>> I do ignore most of it, I'm only getting mad because a few people  \n>> keep\n>> telling me that I'm trolling, or being inflammatory, simply by  \n>> posting\n>> reasoned, factual replies, but everybody who keeps spewing insults\n>> are, apparently, not a problem at all.\n>\n> I understand that you think your replies are reasoned and factual, but\n> everyone else thinks you're wrong. They are getting frustrated\n> defending a point with which you continue to disagree, hence the\n> insults.\n\nDon't you think I'm frustrated at the behavior of everyone else here?  \nBut you don't see me flinging insults.\n\n>> At first, I did. Now it's just tiresome, since he keeps calling me  \n>> and\n>> HFS+ dumb for the exact same reasons he did at the start of the  \n>> thread\n>> no matter how I respond. Apparently he's simply more interested in\n>> keeping his own opinion than in the actual reasons behind HFS+'s\n>> decisions. It's rather frustrating.\n>\n> Have you considered that maybe he's right? In any case, you're not\n> going to convince Linus of anything. From what I can tell, he forms\n> his opinions based on facts he collects himself and his own\n> experience. I gather that anything you say he will consider, *at\n> best*, as hearsay. Besides that, he's actually working to solve the\n> problem, while still taking the time to respond to your points.\n\nCollecting facts yourself is fine, but insulting anybody with a  \ndissenting *opinion* simply because it's different is just plain wrong.\n\n> Really, please, go take a walk outside. Get some fresh air. Maybe stop\n> reading the git list for a week or two. In the grand scheme of things\n> it doesn't matter what git developers think of HFS+ as long as they're\n> willing to make git work with it, which apparently, and in spite of\n> you at this point, they are.\n\nFor the majority of this thread, nobody was making any indication that  \nthey cared at all about fixing this problem - that was my primary  \nmotivation to continue. If I had dropped this the first time someone  \ntold me to, do you think anybody would be working on the problem now?\n\nAs for dropping this conversation now, I'd love to. If you really want  \nto drop it, I urge you to do just that - don't respond to this  \nmessage. Read it, digest it, and then just let it sit. If this is the  \nlast message on the subject, that would be *wonderful*. But if you  \nrespond to this message then you have absolutely no ground to accuse  \nme of refusing to drop it. So please, don't.\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66443","messageId":"7vk5lzc3yr.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"DE7B2DE6-03B1-4781-92C7-096E591369A1@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-24T03:11:40Z","receivedAt":"2008-01-24T03:11:40Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kevin Ballard <kevin@sb.org> writes:\n\n> As for dropping this conversation now, I'd love to. If you really want\n> to drop it, I urge you to do just that - don't respond to this\n> message. Read it, digest it, and then just let it sit. If this is the\n> last message on the subject, that would be *wonderful*. But if you\n> respond to this message then you have absolutely no ground to accuse\n> me of refusing to drop it. So please, don't.\n\nI would not have said that if I were you.  That makes you look\nvery bad.  The impression I get after reading the above is that\nthe only thing you care about is to have the last word in the\nthread.\n\nPeople with opinions different from you could tone their message\ndown and stick to a more neutral sounding statement, \"This patch\nworks around the issue X on HFS+\", but not everybody is always\nnice-and-calm.  But _you_ do not have to counter fire with fire,\nespecially if your goal isn't to flame but is to resolve\ntechnical issues with cool head.  As long as you do not get\nupset and start the flamewar every time whenever somebody says\n\"This patch works around the issue only that broken crap HFS+\nhas due to its stupid filename corruption choice it made\", when\nhe could just have said it in a more neutral way, we can keep\nthe conversation constructive and civilized.\n\nLet me suggest an alternative, as I think this thread raged on\nlong enough.  When you read somebody says \"HFS+ corrupts\", \"HFS+\nis broken\", \"this works around the stupidity of HFS+\", just take\na deep breath, pretend that you did not hear these words that\nmake you feel insulted.  Instead pretend that you heard \"HFS+\nnormalizes\", \"HFS+ is different\", and \"fixes problem on HFS+\".\nDo not respond with \"No it is not a corruption\", \"No, HFS+ is\nnot broken\" and \"No, that is not a work around, but is a fix\"\nwith another long thread.\n\nI can imagine a civilized conversation to go this way:\n\n\tLinus: This patch would hopefully work around the stupid\n\tand broken normalization choice HFS+ people made years ago.\n\n\tYou: Ok, I tested that patch, and it does fix the issue\n\tfor me on HFS+ for most cases, but I still have issues\n\tif I use character X, Y and Z.\n\n\tLinus: Yeah, that is another direct consequence of the\n\tstupidity of HFS+.  At this point I think the previous\n\tpatch bends git backwards enough and I do not know if it\n\tis worth addressing by bending further...\n\n\tYou: How about introducing this new structure so that\n\tthese cases can be handled in a way more friendly to\n\tHFS+, like this patch?\n\n\tLinus: Yeah, I can buy that, it looks ugly but it would\n\tnot hurt people on other systems.\n\nHmm?\n"},{"id":"66447","messageId":"46a038f90801232037t76e103edt1585d49b2ed19862@mail.gmail.com","threadId":"11645","inReplyTo":"7vk5lzc3yr.fsf@gitster.siamese.dyndns.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-24T04:37:58Z","receivedAt":"2008-01-24T04:37:58Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 24, 2008 4:11 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> make you feel insulted.  Instead pretend that you heard \"HFS+\n> normalizes\", \"HFS+ is different\", and \"fixes problem on HFS+\".\n> Do not respond with \"No it is not a corruption\", \"No, HFS+ is\n> not broken\" and \"No, that is not a work around, but is a fix\"\n> with another long thread.\n\nIndeed. And it'd be good if Kevin could consider that this forum is\nfor technical discussion - not democracy but meritocracy,\nbest-solution-cracy and perhaps \"fix-patch-ocracy\". And that people\nthat have written good code in the past, posted amazing patches, and\nwondrous test cases can sometimes get a bit more opinionated. But\nnewcomers needs to earn a bit of respect before lecturing people.\n\nKevin, other people have already started posting nice nuggets of test\ncases. Where are *your* test cases? That would be a nice way to \"have\nthe last word\" on this ;-)\n\ncheers,\n\n\n\nmartin\n"},{"id":"66448","messageId":"C3E61435-AAF8-4F30-ACD6-B083165B5441@sb.org","threadId":"11645","inReplyTo":"46a038f90801232037t76e103edt1585d49b2ed19862@mail.gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-24T05:30:24Z","receivedAt":"2008-01-24T05:30:24Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 23, 2008, at 11:37 PM, Martin Langhoff wrote:\n\n> Kevin, other people have already started posting nice nuggets of test\n> cases. Where are *your* test cases? That would be a nice way to \"have\n> the last word\" on this ;-)\n\n\nI'm planning on devoting time this weekend to learning enough about  \ngit to be able to start hacking. I'm just too busy during the week to  \nbe able to devote the dedicated time necessary to this stuff.  \nHopefully I'll actually be able to start producing stuff this weekend.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66450","messageId":"DADD259A-504E-4D98-9C26-5CD9B35B59C1@zib.de","threadId":"11645","inReplyTo":"C3E61435-AAF8-4F30-ACD6-B083165B5441@sb.org","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Steffen Prohaska","fromEmail":"prohaska@zib.de","sentAt":"2008-01-24T06:39:42Z","receivedAt":"2008-01-24T06:39:42Z","isPatch":false,"sender":{"key":"prohaska@zib.de","avatar":"https://avatars.githubusercontent.com/u/217580?v=4"},"body":"\nOn Jan 24, 2008, at 6:30 AM, Kevin Ballard wrote:\n\n> On Jan 23, 2008, at 11:37 PM, Martin Langhoff wrote:\n>\n>> Kevin, other people have already started posting nice nuggets of test\n>> cases. Where are *your* test cases? That would be a nice way to \"have\n>> the last word\" on this ;-)\n>\n>\n> I'm planning on devoting time this weekend to learning enough about  \n> git to be able to start hacking. I'm just too busy during the week  \n> to be able to devote the dedicated time necessary to this stuff.  \n> Hopefully I'll actually be able to start producing stuff this weekend.\n\nYou do not need to learn much about git to post a test case.\nOnly a few lines of shell code that demonstrate how git fails to\nhandle a specific situation is needed.  To do this, knowledge\nabout the internals of git does not necessary help.  It should be\nsufficient to know how to use git.\n\nYou may start with a simple shell script and send it to the list.\nThough, a real patch would be the preferred way.  For this, you\nshould have a quick look into the t/ subdirectory.  Just open any\nof the tNUMBER*.sh files.  It should be quite obvious how your\nsequence of shell commands could be cast into a git test script.\n\n\tSteffen\n"},{"id":"66499","messageId":"2C31D645-2BB3-4832-BA46-47B441F530EC@gmail.com","threadId":"11645","inReplyTo":"DADD259A-504E-4D98-9C26-5CD9B35B59C1@zib.de","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mitch Tishmack","fromEmail":"mitcht.git@gmail.com","sentAt":"2008-01-24T18:17:05Z","receivedAt":"2008-01-24T18:17:05Z","isPatch":false,"sender":{"key":"mitcht.git@gmail.com","avatar":null},"body":"Here is a start maybe, I was just testing all of the HFS variants for  \nfun. I will write up a test case later tonight when I am out  done  \nwith work.\n\n#!/bin/sh\n#\n# Test git behavior on OSX's multitudes of HFS types\n# So far UFS, HFS, HFS+, HFSX, HFS+J, UFS is the only sane FS  \navailable...\n#\n\n#cloneurl=\"git://git.kernel.org/pub/scm/git/git.git\"\ncloneurl=\"/Volumes/gitufs/git\"\nramdiskdir=\"/private/tmp/gitramdisk\"\nresults=\"/private/tmp/fsresults\"\n\necho \"Creating 100M UFS ramdisk for clone operation.\"\nrawdev=`hdid -nomount ram://102400`\nnewfs $rawdev > /dev/null 2>&1\nmkdir $ramdiskdir > /dev/null 2>&1\nmount -t ufs $rawdev $ramdiskdir > /dev/null 2>&1\ncd $ramdiskdir && git clone $cloneurl > /dev/null 2>&1\ncd $ramdiskdir/git\nif [ -f $results ] ; then\n   echo \"Removing old results.\"\n   rm $results\nfi\n\necho \"Creating HFS image\"\nhdiutil create -size 50m -fs HFS -attach -volname \"hfs\" /tmp/hfs.dmg  \n > /dev/null 2>&1\necho \"Creating HFS+ image\"\nhdiutil create -size 50m -fs HFS+ -attach -volname \"hfsplus\" /tmp/ \nhfsplus.d > /dev/null 2>&1\necho \"Creating HFS+J image\"\nhdiutil create -size 50m -fs HFS+J -attach -volname \"hfsplusJ\" /tmp/ \nhfsplusjournal.dmg > /dev/null 2>&1\necho \"Creating HFSX image\"\nhdiutil create -size 50m -fs HFSX -attach -volname \"hfsx\" /tmp/ \nhfsx.dmg > /dev/null 2>&1\necho \"Creating UFS image\"\nhdiutil create -size 50m -fs UFS -attach -volname \"hfsu\" /tmp/hfsu.dmg  \n > /dev/null 2>&1\n\nfor x in `ls -d /Volumes/hfs*`;do\n   echo \"Testing $x clone.\"\n   echo \"Results for $x:\" >> $results\n   (echo \"-- git clone results --\" && cd ${x} && /usr/bin/time git  \nclone $ramdiskdir/git) >> $results 2>&1\n   (echo \"-- git status results --\" && cd ${x}/git && /usr/bin/time  \ngit status) >> $results 2>&1\n   cd ${x} && perl -CO -e 'print pack(\"U\",0x00E4).\"\\n\"' | xargs touch  \n# umlauted a\n   cd ${x} && perl -CO -e 'print pack(\"U\",0x0061).pack(\"U\", \n0x0308).\"\\n\"' | xargs touch # umlauted a by combining diareses\n   ls -d ${x}/* | xxd >> $results\n   cd && hdiutil eject ${x} > /dev/null 2>&1\ndone\n\n# cleanup\ncd $HOME\numount -f $ramdiskdir > /dev/null 2>&1\nhdiutil detach $rawdev > /dev/null 2>&1\nrm -Rf $ramdiskdir /tmp/hfs*.dmg\nmore $results\n\nMy results on leopard:\n$ cat /tmp/fsresults\nResults for /Volumes/hfs:\n-- git clone results --\nInitialized empty Git repository in /Volumes/hfs/git/.git/\ncpio: Unable to create /Volumes/hfs/git/.git/objects/pack/ \npack-06100ef5fbd98d07358505696e2e0c5600a9b279.pack: Invalid argument\ncpio: Unable to create /Volumes/hfs/git/.git/objects/pack/ \npack-06100ef5fbd98d07358505696e2e0c5600a9b279.idx: Invalid argument\ncpio: Unable to create /Volumes/hfs/git/.git/objects/pack/ \npack-401a5ae571eb23ec896d7e441deae4e313d0de9c.pack: Invalid argument\ncpio: Unable to create /Volumes/hfs/git/.git/objects/pack/pack- \nab8844b63fcb4fc5896e9d75b0d10c566d5ce5bb.pack: Invalid argument\ncpio: Unable to create /Volumes/hfs/git/.git/objects/pack/pack- \nab8844b63fcb4fc5896e9d75b0d10c566d5ce5bb.idx: Invalid argument\ncpio: Unable to create /Volumes/hfs/git/.git/objects/pack/ \npack-9ffbf58084280a496aef6849fe3effe742a99d77.pack: Invalid argument\ncpio: Unable to create /Volumes/hfs/git/.git/objects/pack/ \npack-9ffbf58084280a496aef6849fe3effe742a99d77.idx: Invalid argument\ncpio: Unable to create /Volumes/hfs/git/.git/objects/pack/ \npack-401a5ae571eb23ec896d7e441deae4e313d0de9c.idx: Invalid argument\ncpio: Unable to create /Volumes/hfs/git/.git/objects/ \n03/4ee24912da0a700ba27109825710fd84d64591: Invalid argument\n         0.20 real         0.02 user         0.05 sys\n-- git status results --\ngit_fs.sh: line 39: cd: /Volumes/hfs/git: No such file or directory\n0000000: 2f56 6f6c 756d 6573 2f68 6673 2f61 cc88  /Volumes/hfs/a..\n0000010: 0a                                       .\nResults for /Volumes/hfsplus:\n-- git clone results --\nInitialized empty Git repository in /Volumes/hfsplus/git/.git/\n         2.56 real         0.47 user         0.94 sys\n-- git status results --\n# On branch master\n# Untracked files:\n#   (use \"git add <file>...\" to include in what will be committed)\n#\n#\tgitweb/test/Märchen\nnothing added to commit but untracked files present (use \"git add\" to  \ntrack)\n         0.46 real         0.27 user         0.08 sys\n0000000: 2f56 6f6c 756d 6573 2f68 6673 706c 7573  /Volumes/hfsplus\n0000010: 2f61 cc88 0a2f 566f 6c75 6d65 732f 6866  /a.../Volumes/hf\n0000020: 7370 6c75 732f 6769 740a                 splus/git.\nResults for /Volumes/hfsplusJ:\n-- git clone results --\nInitialized empty Git repository in /Volumes/hfsplusJ/git/.git/\n         2.29 real         0.45 user         0.91 sys\n-- git status results --\n# On branch master\n# Untracked files:\n#   (use \"git add <file>...\" to include in what will be committed)\n#\n#\tgitweb/test/Märchen\nnothing added to commit but untracked files present (use \"git add\" to  \ntrack)\n         0.57 real         0.31 user         0.10 sys\n0000000: 2f56 6f6c 756d 6573 2f68 6673 706c 7573  /Volumes/hfsplus\n0000010: 4a2f 61cc 880a 2f56 6f6c 756d 6573 2f68  J/a.../Volumes/h\n0000020: 6673 706c 7573 4a2f 6769 740a            fsplusJ/git.\nResults for /Volumes/hfsu:\n-- git clone results --\nInitialized empty Git repository in /Volumes/hfsu/git/.git/\n         5.08 real         0.48 user         0.94 sys\n-- git status results --\n# On branch master\nnothing to commit (working directory clean)\n         0.26 real         0.20 user         0.05 sys\n0000000: 2f56 6f6c 756d 6573 2f68 6673 752f 61cc  /Volumes/hfsu/a.\n0000010: 880a 2f56 6f6c 756d 6573 2f68 6673 752f  ../Volumes/hfsu/\n0000020: 6769 740a 2f56 6f6c 756d 6573 2f68 6673  git./Volumes/hfs\n0000030: 752f c3a4 0a                             u/...\nResults for /Volumes/hfsx:\n-- git clone results --\nInitialized empty Git repository in /Volumes/hfsx/git/.git/\n         2.49 real         0.46 user         0.88 sys\n-- git status results --\n# On branch master\n# Untracked files:\n#   (use \"git add <file>...\" to include in what will be committed)\n#\n#\tgitweb/test/Märchen\nnothing added to commit but untracked files present (use \"git add\" to  \ntrack)\n         0.25 real         0.20 user         0.04 sys\n0000000: 2f56 6f6c 756d 6573 2f68 6673 782f 61cc  /Volumes/hfsx/a.\n0000010: 880a 2f56 6f6c 756d 6573 2f68 6673 782f  ../Volumes/hfsx/\n0000020: 6769 740a                                git.\n\n\nOn Jan 24, 2008, at 12:39 AM, Steffen Prohaska wrote:\n\n>\n> On Jan 24, 2008, at 6:30 AM, Kevin Ballard wrote:\n>\n>> On Jan 23, 2008, at 11:37 PM, Martin Langhoff wrote:\n>>\n>>> Kevin, other people have already started posting nice nuggets of  \n>>> test\n>>> cases. Where are *your* test cases? That would be a nice way to  \n>>> \"have\n>>> the last word\" on this ;-)\n>>\n>>\n>> I'm planning on devoting time this weekend to learning enough about  \n>> git to be able to start hacking. I'm just too busy during the week  \n>> to be able to devote the dedicated time necessary to this stuff.  \n>> Hopefully I'll actually be able to start producing stuff this  \n>> weekend.\n>\n> You do not need to learn much about git to post a test case.\n> Only a few lines of shell code that demonstrate how git fails to\n> handle a specific situation is needed.  To do this, knowledge\n> about the internals of git does not necessary help.  It should be\n> sufficient to know how to use git.\n>\n> You may start with a simple shell script and send it to the list.\n> Though, a real patch would be the preferred way.  For this, you\n> should have a quick look into the t/ subdirectory.  Just open any\n> of the tNUMBER*.sh files.  It should be quite obvious how your\n> sequence of shell commands could be cast into a git test script.\n>\n> \tSteffen\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"66507","messageId":"8C467B3C-6C5C-46A2-BB0B-BE689F7D5CAC@gmail.com","threadId":"11645","inReplyTo":"DADD259A-504E-4D98-9C26-5CD9B35B59C1@zib.de","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Mitch Tishmack","fromEmail":"mitcht.git@gmail.com","sentAt":"2008-01-24T18:52:44Z","receivedAt":"2008-01-24T18:52:44Z","isPatch":false,"sender":{"key":"mitcht.git@gmail.com","avatar":null},"body":"Apologies Steffen, I grabbed your CamelCase test and did a search/ \nreplace, wasn't sure what to call it though... But I am on lunch and  \nwanted to be useful. Rip it apart all you want.\n\nFails on hfs* on OSX, works on ufs. I will bother with zfs when it can  \nbe used again.\n\nOn UFS:\n$ /bin/sh ./t0060-normalization.sh\n*   ok 1: setup\n*   ok 2: rename (silent normalization)\n*   ok 3: merge (silent normalization)\n* passed all 3 test(s)\n\n\nOn HFS:\n$ /bin/sh t0060-normalization.sh\n*   ok 1: setup\n* FAIL 2: rename (silent normalization)\n\t\n\t\n\t git mv ä ä &&\n\t git commit -m \"rename\"\n\t\n\t\n* FAIL 3: merge (silent normalization)\n\t\n\t\n\t git reset --hard initial &&\n\t git merge topic\n\t\n\t\n* failed 2 among 3 test(s)\n\nThe test case, it uses perl, assuming only 5.6.1+ will work with this:\ndiff --git a/t/t0060-normalization.sh b/t/t0060-normalization.sh\nnew file mode 100755\nindex 0000000..e012c02\n--- /dev/null\n+++ b/t/t0060-normalization.sh\n@@ -0,0 +1,36 @@\n+#!/bin/sh\n+\n+test_description='Test for silent normalization issues'\n+\n+. ./test-lib.sh\n+\n+auml=`perl -CO -e 'print pack(\"U\",0x00E4)'`\n+aumlcdiar=`perl -CO -e 'print pack(\"U\",0x0061).pack(\"U\",0x0308)'`\n+test_expect_success setup \"\n+  touch $aumlcdiar &&\n+  git add $aumlcdiar &&\n+  git commit -m \\\"initial\\\"\n+  git tag initial &&\n+  git checkout -b topic &&\n+  git mv $aumlcdiar tmp &&\n+  git mv tmp $auml &&\n+  git commit -m \\\"rename\\\" &&\n+  git checkout -f master\n+\n+\"\n+\n+test_expect_success 'rename (silent normalization)' \"\n+\n+ git mv $aumlcdiar $auml &&\n+ git commit -m \\\"rename\\\"\n+\n+\"\n+\n+test_expect_success 'merge (silent normalization)' '\n+\n+ git reset --hard initial &&\n+ git merge topic\n+\n+'\n+\n+test_done\n-- \n1.5.3\n\n\n\n\nOn Jan 24, 2008, at 12:39 AM, Steffen Prohaska wrote:\n\n>\n> On Jan 24, 2008, at 6:30 AM, Kevin Ballard wrote:\n>\n>> On Jan 23, 2008, at 11:37 PM, Martin Langhoff wrote:\n>>\n>>> Kevin, other people have already started posting nice nuggets of  \n>>> test\n>>> cases. Where are *your* test cases? That would be a nice way to  \n>>> \"have\n>>> the last word\" on this ;-)\n>>\n>>\n>> I'm planning on devoting time this weekend to learning enough about  \n>> git to be able to start hacking. I'm just too busy during the week  \n>> to be able to devote the dedicated time necessary to this stuff.  \n>> Hopefully I'll actually be able to start producing stuff this  \n>> weekend.\n>\n> You do not need to learn much about git to post a test case.\n> Only a few lines of shell code that demonstrate how git fails to\n> handle a specific situation is needed.  To do this, knowledge\n> about the internals of git does not necessary help.  It should be\n> sufficient to know how to use git.\n>\n> You may start with a simple shell script and send it to the list.\n> Though, a real patch would be the preferred way.  For this, you\n> should have a quick look into the t/ subdirectory.  Just open any\n> of the tNUMBER*.sh files.  It should be quite obvious how your\n> sequence of shell commands could be cast into a git test script.\n>\n> \tSteffen\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"66514","messageId":"A9803EC0-EAF8-4FA0-ADB0-22B45959C96D@sb.org","threadId":"11645","inReplyTo":"8C467B3C-6C5C-46A2-BB0B-BE689F7D5CAC@gmail.com","subject":"Re: git on MacOSX and files with decomposed utf-8 file names","fromName":"Kevin Ballard","fromEmail":"kevin@sb.org","sentAt":"2008-01-24T19:58:28Z","receivedAt":"2008-01-24T19:58:28Z","isPatch":false,"sender":{"key":"kevin@sb.org","avatar":"https://avatars.githubusercontent.com/u/714?v=4"},"body":"On Jan 24, 2008, at 1:52 PM, Mitch Tishmack wrote:\n\n> Apologies Steffen, I grabbed your CamelCase test and did a search/ \n> replace, wasn't sure what to call it though... But I am on lunch and  \n> wanted to be useful. Rip it apart all you want.\n>\n> [snip]\n\nWell, I was planning on writing my own test case today, but you seem  \nto have beaten me to the punch. I just tested your script and it does  \nindeed fail as expected on HFS+. Thank you for producing this test case.\n\n-Kevin Ballard\n\n-- \nKevin Ballard\nhttp://kevin.sb.org\nkevin@sb.org\nhttp://www.tildesoft.com\n\n\n"},{"id":"66519","messageId":"7vprvr7x8h.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801230930390.1741@woody.linux-foundation.org","subject":"On pathnames","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-24T21:02:54Z","receivedAt":"2008-01-24T21:02:54Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"One of Linus's recent patch introduces an index hashtable so\nthat we can later hash \"equivalent\" names into the same bucket\nto allow us non-byte-by-byte comparison.\n\nBefore going further, I needed to formalize what we are trying\nto achieve.  I learned a few things from the long flamewar\nthread, but it is very inefficient to go back to the thread to\npick only the useful pieces.  The whole flamewar simply did not\nfit a small Panda brain.\n\nThat was the reason for this write-up.\n\nDesign constraints.  In the following, I'll use two names $A and\n$B as an example.  They are a pair of names that are considered\nequivalent in some contexts, such as:\n\n        A=xt_connmark.c  B=xt_CONNMARK.c\n\n (1) Some filesystems prevent you from having these two\n     (confusing) paths in a directory at the same time.  Some do\n     not implement this confusion prevention, and allows both\n     names to exist at the same time.\n\n     Let's call the former \"case insensitive\", and the latter\n     \"case sensitive\".\n\n (2) readdir(3) on some \"case insensitive\" filesystems returns\n     $A, after a successful creat(2) of $B.  Others remember\n     which one of the two \"equivalent\" names were used in\n     creat(2).\n\n     Let's call the former \"case folding\", and the latter \"case\n     preserving\".\n\n     We assume open(2) or lstat(2) of $A or $B will succeed\n     after allowing creat(2) of $B if a case folding filesystem\n     returns $A from readdir(3).\n\n (3) Among the \"case folding\" ones, some filesystems fold the\n     pathname to a form that is less interoperable with other\n     systems, and/or the form that is likely to be different\n     from what the end-user usually enters.\n\n     Such filesystems are \"inconveniently case folding\".\n\nThe last one is not quite apparent with the \"xt_connmark.c\"\nexample, but if you replace $A and $B in the above description\nwith:\n\n        A=Ma\"rchen       B=Märchen\n\nit would hopefully become more clear.\n\nFor example, vfat is generally \"case preserving\".  In that long\nflamewar thread, I think we learned that HFS+ is in general\n\"inconveniently case folding\" with respect to Unicode, by always\nfolding to $A but the keyboard/IM input is more likely to come\nas $B, which happens to be the more interoperable form with\nother systems.\n\n\nIssues with case insensitive filesystems\n----------------------------------------\n\nAt the data structure level, a pathname to git is a sequence of\nbytes terminated with NUL.  This will _not_ change.\n\nBy the way, at the data structure level, a tree entry in git can\nrepresent a blob that is a symbolic link.  A tree entry in git\ncan also represent a blob that is a regular file, and in that\ncase, it can represent if it is executable or not.  These will\nalso not change.\n\nNow, let's think about how we allow use of git on a filesystem\nthat is incapable of symbolic links, and/or a filesystem that\ndoes not have trustable executable bit.\n\nWe do not say \"Symlinks are evil and not supported everywhere,\nso let's introduce a project configuration to disallow addition\nof symlinks\".  We do not say that to the executable bit, either.\n\nInstead, we have fallback methods to allow manipulating symlinks\nand executable bit on such a filesystem that is incapable of\nhandling them natively.\n\nWe should be able to do the same for this \"case sensitivity\"\nissue.  A tree that has xt_connmark.c and xt_CONNMARK.c at the\nsame time cannot be checked out on a case insensitive filesystem.\n\nThe filesystem is simply incapable of it (please just calmly\nrephrase it in your head as \"does not allow such confusing\ncraziness\" instead of starting another flamewar, if you feel the\nexpression \"incapable of\" insults your favorite filesystem).\n\nThat may mean the project should avoid such equivalent names in\nits trees (and having a project wide configuration could be a\ntechnical means to help enforcing that policy), but it does not\nmean the core level of git should prevent them to be created on\nsuch systems.  It just means that there should be a way, that\ncould (and sometimes has to) be different from the \"natural\"\nway, to manipulate such tree entries even on a case insensitive\nfilesystem.\n\nFor example, if I find that RelNotes symlink incorrectly points\nat Documentation/RelNotes-1.5.44.txt and want to fix it and push\nit out immediately, but if I am on the road and the only\nenvironment I can borrow is a git installation on a filesystem\nthat is symlink-challenged, I can still do the fix. On such a\nfilesystem, a symlink is checked out as a regular file but is\nstill marked as a symlink in the index.  The only thing I need\nto do is to edit the file (making sure not to add an extra LF at\nthe end) and add it to the index.  That's certainly different\nfrom the \"natural\" way to do that on a filesystem with symlinks,\nwhich is \"ln -fs Documentation/RelNotse-1.5.4.txt RelNotes\", but\nthe point is that we make it possible.\n\nThe same thing should apply to two files that cannot be checked\nout at the same time on case insensitive filesystems.  Perhaps\nwe could have something like:\n\n\t$ git show :xt_CONNMARK.c >xt_connmark-1.c\n        $ edit xt_connmark-1.c\n\t$ git add --as xt_CONNMARK.c xt_connmark-1.c\n\n\nIssues with case folding filesystems\n------------------------------------\n\nIn addition to the above, case folding filesystems additionally\nhave an issue even when there is no \"confusing\" names in the\ntree.  The project may want to have \"Märchen\" (but not\n\"Ma\"rchen\"), but a checkout (which is creat(2) of \"Märchen\" --\nbecause that is the byte sequence recorded in tree objects and\nthe index) will result in \"Ma\"rchen\" and no \"Märchen\" (hence\nreaddir(3) returns \"Ma\"rchen\").\n\nLinus's patch to use a hashtable that links \"equivalent\" names\ntogether is a step in the right direction to address this.  The\ntree (and the index) has name $B, we check out and the\nfilesystem folds it to $A.  When we get the name $A back from\nthe filesystem (via readdir(3)), we hash the name using a hash\nfunction that would drop names $A and $B into the same bucket,\nand compare that name $A with each hash entry using a comparison\nthat considers $A and $B are equivalent.  If we find one, then\nwe keep the name $B we have already.\n\nIf it is a new file, we won't find any name that is equivalent\nto $A in the index, and we use the name $A obtained from\nreaddir(3).\n\nBUT with a twist.\n\nIf the filesystem is known to be inconveniently case folding, we\nare better off registering $B instead of $A (assuming we can\nconvert from $A to $B).\n\nOne bad issue during development is that we cannot sanely\nemulate case folding behaviour on non case-folding filesystems\nwithout wrapping open(2), lstat(2), and friends, because of the\nassumption we made above in (2) where we defined the term \"case\nfolding\".  This means that the codepath to deal with case\nfolding filesystems inevitably are harder to debug.\n\n\n\nTasks\n-----\n\n - Identify which case folding filesystems need to be supported,\n   and make sure somebody understands its folding logic;\n\n - For each supported case folding logic, these are needed:\n\n   - a hash function that throws \"equivalent\" names in the same\n     bucket, to be used in Linus's patch;\n\n   - a compare function to determine equivalent names;\n\n   - a convert function that takes a possibly inconvenient form\n     of equivalent name (i.e. $A above) as input and returns\n     more convenient form (i.e. $B above)\n\n - Identify places that we use the names obtained from places\n   other than the index and tree.  From these places, we would\n   need to call the convert function to (de)mangle the name\n   before they hit the index.\n\n   Because we may be getting driven by something like:\n\n\t$ find | xargs git-foo\n\n   handling readdir(3) we do ourselves any specially does not\n   make much sense.  Any path from the user is suspect.\n\n - Identify places that we look for a name in the index, and\n   perform equivalent comparison instead of memcmp(3) we\n   traditionally did.  Linus's patch gives scaffolding for this.\n"},{"id":"66525","messageId":"alpine.LFD.1.00.0801241722130.22568@xanadu.home","threadId":"11645","inReplyTo":"7vprvr7x8h.fsf@gitster.siamese.dyndns.org","subject":"Re: On pathnames","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-01-24T22:31:27Z","receivedAt":"2008-01-24T22:31:27Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Thu, 24 Jan 2008, Junio C Hamano wrote:\n\n> If it is a new file, we won't find any name that is equivalent\n> to $A in the index, and we use the name $A obtained from\n> readdir(3).\n> \n> BUT with a twist.\n> \n> If the filesystem is known to be inconveniently case folding, we\n> are better off registering $B instead of $A (assuming we can\n> convert from $A to $B).\n\nWhy?\n\nIf you have no other representation for the file name than $A already, \nthen I don't see why Git would have to play similar evil games and \ncorru^H^H^Hnvert $A into $B.  Just store $A in the index and tree \nobjects and be done with it.\n\n\nNicolas\n"},{"id":"66528","messageId":"BAYC1-PASMTP12CD8892FB04872736D9DEAE380@CEZ.ICE","threadId":"11645","inReplyTo":"7vprvr7x8h.fsf@gitster.siamese.dyndns.org","subject":"Re: On pathnames","fromName":"Sean","fromEmail":"seanlkml@sympatico.ca","sentAt":"2008-01-24T23:56:02Z","receivedAt":"2008-01-24T23:56:02Z","isPatch":false,"sender":{"key":"seanlkml@sympatico.ca","avatar":"https://gravatar.com/avatar/f92923f54fc08c401fc59b71829d4b89e9b8087fbba45ff87c82e6a83aee02ae?d=mp&s=160"},"body":"On Thu, 24 Jan 2008 13:02:54 -0800\nJunio C Hamano <gitster@pobox.com> wrote:\n\n> One bad issue during development is that we cannot sanely\n> emulate case folding behaviour on non case-folding filesystems\n> without wrapping open(2), lstat(2), and friends, because of the\n> assumption we made above in (2) where we defined the term \"case\n> folding\".  This means that the codepath to deal with case\n> folding filesystems inevitably are harder to debug.\n\nAll true.  Though Linux support for creating and using HFS+ volumes\nseems like it may be helpful.  Trying the test case patch[*] posted\nby Mitch Tishmack showed the problem here.  The only slightly\nstrange thing was that there didn't seem to be an issue with the\ngitweb/test/Märchen file after cloning to the HFS volume.\n\nSean.\n\n[*]\n$ dd bs=1M count=250 < /dev/zero > hfs_vol\n  262144000 bytes (262 MB) copied, 6.12703 s, 42.8 MB/s\n\n$ /sbin/mkfs.hfsplus -v Test -n c=4096,e=1024 hfs_vol\n  Initialized hfs_vol as a 250 MB HFS Plus volume\n\n$ mkdir hfs\n$ sudo mount -t hfsplus -o loop hfs_vol hfs\n$ sudo chmod a+rwx hfs\n$ cd hfs\n$ git clone ~/local/sources/git\n  Initialized empty Git repository in ~/hfs/git/.git/\n  49486 blocks\n\n$ cd git\n$ make\n   ...\n\n$ cd t\n$ git apply ~/Mitch_Tishmack.patch\n$ ./t0060-normalization.sh\n  * FAIL 1: setup\n\n\t  touch ä &&\n\t  git add ä &&\n\t  git commit -m \"initial\"\n\t  git tag initial &&\n\t  git checkout -b topic &&\n\t  git mv ä tmp &&\n\t  git mv tmp ä &&\n\t  git commit -m \"rename\" &&\n\t  git checkout -f master\n\t\n  * FAIL 2: rename (silent normalization)\n\t\n\t git mv ä ä &&\n\t git commit -m \"rename\"\n\t\n  * FAIL 3: merge (silent normalization)\n\n\t git reset --hard initial &&\n\t git merge topic\n\t\n  * failed 3 among 3 test(s)\n\nSean.\n"},{"id":"66530","messageId":"alpine.LSU.1.00.0801250007490.5731@racer.site","threadId":"11645","inReplyTo":"7vprvr7x8h.fsf@gitster.siamese.dyndns.org","subject":"Re: On pathnames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-25T00:36:29Z","receivedAt":"2008-01-25T00:36:29Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 24 Jan 2008, Junio C Hamano wrote:\n\n> [A nice, concise, well written and obviously thought-through summary of \n>  the case sensitivity and UTF-8 file name issues.]\n\nThank you Junio.  It must have taken much more time than just sitting \ndown and hacking into the keyboard.  By this thinking before writing, you \ninvested some time that you save all the readers, including me.  I \nappreciate that very much.\n\n> [Goes on to describe what we do with symlinks when the filesystem is not \n>  capable of representing symlinks; compares that situation to the \n>  filenames situation.]\n\nThere is a fundamental difference between the symlinks situation and the \nfilename situation that you should keep in mind:  even if the filesystem \ncannot create symlinks, the nature of filenames as unique keys is not \nchanged.  You cannot have a symlink and a file of the same name.  In a \nway, it takes away a degree of freedom of the _values_ that the _keys_ \npoint to.\n\nThe same is not true for the case-challenged filesystems; they change the \nnature from unique keys to semi-unique keys.  So while other filesystems \ncan discern all different keys, these challenged filesystems cannot; they \ntake away a degree of freedom of the _keys_.\n\nIt is much easier to cope with the lack of degree of freedom in values; \nyou have to store the metadata somewhere else -- in this case the index -- \nbut it is still easily accessible by the key.\n\nBut that is not possible if two different _keys_ are not accepted as \ndifferent by the filesystem.  You can still store the different metadata \nin the index, but the _content_ cannot be in the filesystem under the \ndesired keys; not at the same time, anyway.\n\n> Perhaps we could have something like:\n> \n> \t$ git show :xt_CONNMARK.c >xt_connmark-1.c\n>         $ edit xt_connmark-1.c\n> \t$ git add --as xt_CONNMARK.c xt_connmark-1.c\n\nSomething similar is already possible:\n\n\t$ git checkout xt_CONNMARK.c\n\t$ edit xt_CONNMARK.c\n\t$ git add xt_CONNMARK.c\n\nbut you have to keep in mind that\n\n\t- \"git add -u\" or \"git commit -a\" is a no-no-no, and\n\t- the system will not build, no matter what you change in git\n\non those filesystems.\n\nHaving said that, I think that a config variable/commit hooks for those \nrepositories which _happen_ to live on sane filesystems, but have to be \nchecked out on challenged ones, makes absolute sense.  (The commit hook is \npossible already, but less efficient than the config variable.)\n\n> If it is a new file, we won't find any name that is equivalent to $A in \n> the index, and we use the name $A obtained from readdir(3).\n> \n> BUT with a twist.\n> \n> If the filesystem is known to be inconveniently case folding, we are \n> better off registering $B instead of $A (assuming we can convert from $A \n> to $B).\n\nI tend to agree with Nico.  We should not \"learn\" from the challenged \nfilesystems.\n\n> Tasks\n> -----\n> \n>  - Identify which case folding filesystems need to be supported,\n>    and make sure somebody understands its folding logic;\n> \n>  - For each supported case folding logic, these are needed:\n> \n>    - a hash function that throws \"equivalent\" names in the same\n>      bucket, to be used in Linus's patch;\n\nAFAIR Linus wanted to have one has function to rule them all.  That would \nbe way cool, since it means fewer possibilities for bugs to go undetected.\n\n>    - a compare function to determine equivalent names;\n\nAFAICT we need three functions: strcasecmp(), utf8_strcmp() and \nutf8_strcasecmp().  Although I might be wrong, and the second is not \nneeded.\n\nProbably the answer for this has been buried in many, many lines that I \ndecided not to read.  Maybe I'll ask Randal on IRC, he's usually very \nquick to give me reasonable and concise answers.  And then we trash-talk a \nlittle, just for fun.\n\nCiao,\nDscho\n"},{"id":"66533","messageId":"46a038f90801241955j672ca508q2eb691be5e1d328f@mail.gmail.com","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801241722130.22568@xanadu.home","subject":"Re: On pathnames","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-01-25T03:55:10Z","receivedAt":"2008-01-25T03:55:10Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Jan 25, 2008 11:31 AM, Nicolas Pitre <nico@cam.org> wrote:\n> On Thu, 24 Jan 2008, Junio C Hamano wrote:\n>\n> > If it is a new file, we won't find any name that is equivalent\n> > to $A in the index, and we use the name $A obtained from\n> > readdir(3).\n> >\n> > BUT with a twist.\n> >\n> > If the filesystem is known to be inconveniently case folding, we\n> > are better off registering $B instead of $A (assuming we can\n> > convert from $A to $B).\n>\n> Why?\n>\n> If you have no other representation for the file name than $A already,\n> then I don't see why Git would have to play similar evil games and\n> corru^H^H^Hnvert $A into $B.  Just store $A in the index and tree\n> objects and be done with it.\n\nBecause if you happen to be on a case-challenged filesystem, and you\nneed to add a new file - say xt_CONNBARK.c, you can, even if your FS\nonly has xt_connbark.c to offer to git. Granted, it's harder, but it\ncan be done, and it means that powerusers and wrappers around git have\na hope of dealing with it.\n\nI think this is an excellent plan.\n\nThere is one thing that I don't see in Junio's plan -- which is\nexcellent -- and is\n\n - a warning during checkout if the index contains \"equivalent\" paths\nthat will clobber eachother during checkout.\n - an optional warning/error during add, to be raised if I am adding a\npath that is equivalent to an already-existing path in the index\n\nThe second one is to support project that are known to be developed on\nthese non-filename-preserving platforms. So if I am on a linux host,\nand I add Readme where README already exists, a warning can save me\nand the project a bit of grief - and possibly catch an unintended\nmistake! Because while I agree that users should be able to store any\nfile in git, in practice most instances of Ma\"rchen/Märchen case will\nbe due to user error (or editor/gui error).\n\ncheers,\n\n\nm\n"},{"id":"66535","messageId":"alpine.LNX.1.00.0801242227250.13593@iabervon.org","threadId":"11645","inReplyTo":"7vprvr7x8h.fsf@gitster.siamese.dyndns.org","subject":"Re: On pathnames","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-01-25T04:00:44Z","receivedAt":"2008-01-25T04:00:44Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Thu, 24 Jan 2008, Junio C Hamano wrote:\n\n> The same thing should apply to two files that cannot be checked\n> out at the same time on case insensitive filesystems.  Perhaps\n> we could have something like:\n> \n> \t$ git show :xt_CONNMARK.c >xt_connmark-1.c\n> \t$ edit xt_connmark-1.c\n> \t$ git add --as xt_CONNMARK.c xt_connmark-1.c\n\nI think it would be nicer to have:\n\n$ git checkout branch\nWarning: xt_CONNMARK.c conflicts with xt_connmark.c; not checking it out\n$ git checkout xt_CONNMARK.c --as xt_CONNMARK_caps.c\n$ edit xt_CONNMARK_caps.c\n$ git add xt_CONNMARK_caps.c\n\nWhere the index, when support for filesystems with filename restrictions \nis enabled, keeps track both of the name of the file in the project and \nthe name of the file in the filesystem, with this mapping determined \nentirely by the user asking for problem files to be present under \ndifferent names in the working tree.\n\nOf course, you can already do:\n\n$ git update-index --cacheinfo 100644 $(git hash-object -w xt_connmark-1.c) xt_CONNMARK.c\n\n> If it is a new file, we won't find any name that is equivalent\n> to $A in the index, and we use the name $A obtained from\n> readdir(3).\n> \n> BUT with a twist.\n> \n> If the filesystem is known to be inconveniently case folding, we\n> are better off registering $B instead of $A (assuming we can\n> convert from $A to $B).\n\nIs it not the case that, when a user has a file in the filesystem with the \nname Ma\"rchen, the user will still type:\n\n$ git add Märchen\n\nand so we see filenames which are convenient, and we don't overly care \nwhat readdir(3) returns for new filenames? I suppose there is the case of:\n\n$ touch Märchen\n$ git add .\n\nWhich has to figure out what the files in foo are. But the common case for \na new filename is that it gets provided by the user in argv, and the right \nfile contents come from the one that open(2) returns, and there's no \nobvious way to get the filename that readdir(3) would return for a \nfilename in argv anyway.\n\n\t-Daniel\n*This .sig left intentionally blank*"},{"id":"66536","messageId":"7vy7ae7dcb.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"alpine.LFD.1.00.0801241722130.22568@xanadu.home","subject":"Re: On pathnames","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-25T04:12:36Z","receivedAt":"2008-01-25T04:12:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> On Thu, 24 Jan 2008, Junio C Hamano wrote:\n>\n>> If it is a new file, we won't find any name that is equivalent\n>> to $A in the index, and we use the name $A obtained from\n>> readdir(3).\n>> \n>> BUT with a twist.\n>> \n>> If the filesystem is known to be inconveniently case folding, we\n>> are better off registering $B instead of $A (assuming we can\n>> convert from $A to $B).\n>\n> Why?\n>\n> If you have no other representation for the file name than $A already, \n> then I don't see why Git would have to play similar evil games and \n> corru^H^H^Hnvert $A into $B.  Just store $A in the index and tree \n> objects and be done with it.\n\nBecause this \"conversion\" is limited to the case where the\nfilesystem is known to be inconveniently case folding, I\npersonally do not care about this part of the outline that\ndeeply.  It would not bite _me_ or my friends either way.\n\nBut I would imagine that a person who has to work on HFS+ would\nappreciate it if these two sequences behaved the same way:\n\n    $ edit Märchen ;# assume this is a new file\n    $ git add Märchen ;# we were told that IM gives $B (aka NFC)\n\nvs\n\n    $ edit Märchen ;# assume this is a new file\n    $ git add M*en ;# now readdir(3) gives $A (aka NFD)\n\nIf we always convert $A (less interoperable form) to $B (more\ninteroperable form) on inconveniently case folding filesystems,\nthe new index entry will always be in form $B.  Without the\nconversion, the former will give form $B while the latter will\ngive form $A.  It is, as you said, \"similar evil game to\ncorrupt\", but it is not even a corruption at that point, because\nthe inconveniently case folding filesystem already corrupted the\npathname before we get our hands on it, and it won't make a\ndifference for HFS+ only people anyway.\n\nHowever, if the resulting tree that adds a new file is prepared\non an inconveniently case folding filesystem, the conversion\nprocess, by definition, would make the resulting tree more\ninteroperable with other systems than without.\n\nSo I do not see any downside of doing the conversion on such a\nfilesystem but there is this \"interoperability\" upside.\n"},{"id":"66537","messageId":"7vtzl27d1z.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"46a038f90801241955j672ca508q2eb691be5e1d328f@mail.gmail.com","subject":"Re: On pathnames","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-25T04:18:48Z","receivedAt":"2008-01-25T04:18:48Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n\n> There is one thing that I don't see in Junio's plan...\n>\n>  - a warning during checkout if the index contains \"equivalent\" paths\n> that will clobber eachother during checkout.\n>  - an optional warning/error during add, to be raised if I am adding a\n> path that is equivalent to an already-existing path in the index\n\nThanks.  I think these and many other issues need to be worked\nout.\n\nIn my message, I did not even try to be exhaustive.  I outlined\nthe parts that would most deeply affect the parts I care more\ndeeply about, which is the plumbing.  As I said many times, I do\nnot do Porcelains ;-).\n"},{"id":"66538","messageId":"7vprvq7cy7.fsf@gitster.siamese.dyndns.org","threadId":"11645","inReplyTo":"alpine.LNX.1.00.0801242227250.13593@iabervon.org","subject":"Re: On pathnames","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-25T04:21:04Z","receivedAt":"2008-01-25T04:21:04Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Daniel Barkalow <barkalow@iabervon.org> writes:\n\n> $ git checkout branch\n> Warning: xt_CONNMARK.c conflicts with xt_connmark.c; not checking it out\n> $ git checkout xt_CONNMARK.c --as xt_CONNMARK_caps.c\n> $ edit xt_CONNMARK_caps.c\n> $ git add xt_CONNMARK_caps.c\n\nHeh, I like that very much.\n"},{"id":"66539","messageId":"20080125055910.GB21973@coredump.intra.peff.net","threadId":"11645","inReplyTo":"alpine.LNX.1.00.0801242227250.13593@iabervon.org","subject":"Re: On pathnames","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2008-01-25T05:59:10Z","receivedAt":"2008-01-25T05:59:10Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jan 24, 2008 at 11:00:44PM -0500, Daniel Barkalow wrote:\n\n> I think it would be nicer to have:\n> \n> $ git checkout branch\n> Warning: xt_CONNMARK.c conflicts with xt_connmark.c; not checking it out\n> $ git checkout xt_CONNMARK.c --as xt_CONNMARK_caps.c\n> $ edit xt_CONNMARK_caps.c\n> $ git add xt_CONNMARK_caps.c\n> \n> Where the index, when support for filesystems with filename restrictions \n> is enabled, keeps track both of the name of the file in the project and \n> the name of the file in the filesystem, with this mapping determined \n> entirely by the user asking for problem files to be present under \n> different names in the working tree.\n\nHrm. That makes me think: what if rather than doing utf8-ish\ncomparisons, the index stores a bidirectional mapping for any \"munged\"\nnames, and you can manipulate that mapping?\n\nAs in, the index entry for Märchen has an extra entry saying \"I am\nactually on the filesystem as Ma\"rchen\" (let's call this an alias) and\nthere is a pseudo-entry in the index for Ma\"rchen that says \"I'm not\nreally here. See Märchen\" (let's call this a backref).\n\nThen index-modifying commands like \"git-add\" or \"git-checkout\" can set\nup the mapping, either manually (using --as or similar) or using a\nparticular munging scheme (git config core.filemunge hfs). Any time we\ngive an index path to the filesystem, we use its alias name. Any time we\nlook up an index entry and it ends up being a backref, we dereference\nuntil we get a real entry. Index iterators would need to skip backrefs.\n\nI think all systems would follow the same codepath, there is no penalty\nfor filenames which don't use the mapping, and it would be testable on\nnon-challenged filesystems. But perhaps I am missing some obvious\ndeficiency or impossibility.\n\n-Peff\n"},{"id":"66544","messageId":"D1D9821C-0CB8-40AB-8640-D171E812E428@simplicidade.org","threadId":"11645","inReplyTo":"7vy7ae7dcb.fsf@gitster.siamese.dyndns.org","subject":"Re: On pathnames","fromName":"Pedro Melo","fromEmail":"melo@simplicidade.org","sentAt":"2008-01-25T08:08:08Z","receivedAt":"2008-01-25T08:08:08Z","isPatch":false,"sender":{"key":"melo@simplicidade.org","avatar":"https://gravatar.com/avatar/13ddbb01e300285a93aa1e3739653a81f9b1d3438bd03a4ac36b88e4ffeeafc3?d=mp&s=160"},"body":"Hi,\n\nOn Jan 25, 2008, at 4:12 AM, Junio C Hamano wrote:\n> Nicolas Pitre <nico@cam.org> writes:\n>\n>> On Thu, 24 Jan 2008, Junio C Hamano wrote:\n>>\n>>> If it is a new file, we won't find any name that is equivalent\n>>> to $A in the index, and we use the name $A obtained from\n>>> readdir(3).\n>>>\n>>> BUT with a twist.\n>>>\n>>> If the filesystem is known to be inconveniently case folding, we\n>>> are better off registering $B instead of $A (assuming we can\n>>> convert from $A to $B).\n>>\n>> Why?\n>>\n>> If you have no other representation for the file name than $A  \n>> already,\n>> then I don't see why Git would have to play similar evil games and\n>> corru^H^H^Hnvert $A into $B.  Just store $A in the index and tree\n>> objects and be done with it.\n>\n> Because this \"conversion\" is limited to the case where the\n> filesystem is known to be inconveniently case folding, I\n> personally do not care about this part of the outline that\n> deeply.  It would not bite _me_ or my friends either way.\n>\n> But I would imagine that a person who has to work on HFS+ would\n> appreciate it if these two sequences behaved the same way:\n>\n>     $ edit Märchen ;# assume this is a new file\n>     $ git add Märchen ;# we were told that IM gives $B (aka NFC)\n>\n> vs\n>\n>     $ edit Märchen ;# assume this is a new file\n>     $ git add M*en ;# now readdir(3) gives $A (aka NFD)\n>\n> If we always convert $A (less interoperable form) to $B (more\n> interoperable form) on inconveniently case folding filesystems,\n> the new index entry will always be in form $B.  Without the\n> conversion, the former will give form $B while the latter will\n> give form $A.  It is, as you said, \"similar evil game to\n> corrupt\", but it is not even a corruption at that point, because\n> the inconveniently case folding filesystem already corrupted the\n> pathname before we get our hands on it, and it won't make a\n> difference for HFS+ only people anyway.\n>\n> However, if the resulting tree that adds a new file is prepared\n> on an inconveniently case folding filesystem, the conversion\n> process, by definition, would make the resulting tree more\n> interoperable with other systems than without.\n\nAs a HFS+ user, I would welcome this very very much.\n\nI want my tree to be seen by the rest of the world without problems,  \nand I don't think I should impose my filesystem view of proper naming  \non others.\n\nJunio, amazing write-up, many thanks.\n\nBest regards,\n-- \nPedro Melo\nBlog: http://www.simplicidade.org/notes/\nXMPP ID: melo@simplicidade.org\nUse XMPP!\n"},{"id":"66557","messageId":"alpine.LSU.1.00.0801251134570.5731@racer.site","threadId":"11645","inReplyTo":"7vprvq7cy7.fsf@gitster.siamese.dyndns.org","subject":"Re: On pathnames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-25T11:36:00Z","receivedAt":"2008-01-25T11:36:00Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 24 Jan 2008, Junio C Hamano wrote:\n\n> Daniel Barkalow <barkalow@iabervon.org> writes:\n> \n> > $ git checkout branch\n> > Warning: xt_CONNMARK.c conflicts with xt_connmark.c; not checking it out\n> > $ git checkout xt_CONNMARK.c --as xt_CONNMARK_caps.c\n> > $ edit xt_CONNMARK_caps.c\n> > $ git add xt_CONNMARK_caps.c\n> \n> Heh, I like that very much.\n\nIt would make it easier to test on Linux, too, yes.\n\nBut then, it would break the build process all the same.\n\nAnd the implementation would _need_ the index extension Linus seems to \nresent so.\n\nCiao,\nDscho\n"},{"id":"66559","messageId":"alpine.LSU.1.00.0801251137030.5731@racer.site","threadId":"11645","inReplyTo":"7vy7ae7dcb.fsf@gitster.siamese.dyndns.org","subject":"Re: On pathnames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-25T12:25:22Z","receivedAt":"2008-01-25T12:25:22Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 24 Jan 2008, Junio C Hamano wrote:\n\n>     $ edit Märchen ;# assume this is a new file\n>     $ git add Märchen ;# we were told that IM gives $B (aka NFC)\n\nPlease see the discussion on IRC I started with\n\nhttp://colabti.de/irclogger/irclogger_log/git?date=2008-01-25,Fri&sel=16#l36\n\n(asking for somebody to test the issue if filenames are still mangled when \nthe volume was created _case-sensitive_).  The interesting part is this:\n\nhttp://colabti.de/irclogger/irclogger_log/git?date=2008-01-25,Fri&sel=31#l56\n\n(me asking to git-add the file \"Märchen\" explicitely),\n\nhttp://colabti.de/irclogger/irclogger_log/git?date=2008-01-25,Fri&sel=39#l66\n\n(dsymonds saying that no untracked files are shown), and\n\nhttp://colabti.de/irclogger/irclogger_log/git?date=2008-01-25,Fri&sel=47#l76\n\n(dsymonds showing that the git index contains the mangled filename, _not_ \nwhat we asked for).  The strange thing is that\n\nhttp://colabti.de/irclogger/irclogger_log/git?date=2008-01-25,Fri&sel=57#l89\n\nthe command line seems not to be mangling the name.\n\nSummary:\n\nit seems that for some strange reason, \"git add Märchen\" puts the mangled \nfilename into the index, even if \"echo Märchen\" shows the unmangled \nfilename.\n\nCiao,\nDscho\n"},{"id":"66561","messageId":"85odbaoyqp.fsf@lola.goethe.zz","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801251137030.5731@racer.site","subject":"Re: On pathnames","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-01-25T12:50:38Z","receivedAt":"2008-01-25T12:50:38Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> it seems that for some strange reason, \"git add Märchen\" puts the\n> mangled filename into the index, even if \"echo Märchen\" shows the\n> unmangled filename.\n\necho is likely a shell builtin.  So \"git add Märchen\" goes through exec\nwhile echo doesn't.\n\nWhat does /bin/echo Märchen yield?\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"66564","messageId":"FAC8E3EF-7F18-4601-8B4F-09ED24228C21@wincent.com","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801251137030.5731@racer.site","subject":"Re: On pathnames","fromName":"Wincent Colaiuta","fromEmail":"win@wincent.com","sentAt":"2008-01-25T12:53:54Z","receivedAt":"2008-01-25T12:53:54Z","isPatch":false,"sender":{"key":"greg@hurrell.net","avatar":"https://avatars.githubusercontent.com/u/7074?v=4"},"body":"El 25/1/2008, a las 13:25, Johannes Schindelin escribió:\n\n> The strange thing is that\n>\n> http://colabti.de/irclogger/irclogger_log/git?date=2008-01-25,Fri&sel=57#l89\n>\n> the command line seems not to be mangling the name.\n>\n> Summary:\n>\n> it seems that for some strange reason, \"git add Märchen\" puts the  \n> mangled\n> filename into the index, even if \"echo Märchen\" shows the unmangled\n> filename.\n>\n> Ciao,\n> Dscho\n\n\nNot sure if I grokked the IRC interchange fully but check this out:\n\n$ touch Märchen\n$ echo Märchen | xxd -g1\n0000000: 4d 61 cc 88 72 63 68 65 6e 0a                    Ma..rchen.\n$ echo Märchen | xxd -g1\n0000000: 4d c3 a4 72 63 68 65 6e 0a                       M..rchen.\n\nThe first one shows me creating the file, then typing \"echo M\" and  \nhitting tab so that the shell autocompletes the filename for me based  \non what it sees in the current directory. Note how it's decomposed.\n\nThe second one shows me manually typing the string \"Märchen\" with no  \ntab autocompletion (literally typing ¨ then a), and you'll notice that  \nthis time it is precomposed.\n\nSo that might explain why \"echo Märchen\" is showing an unmangled name;  \nif he just typed it out like I did then that would be the expected  \nresult.\n\nCheers,\nWincent\n"},{"id":"66574","messageId":"alpine.LNX.1.00.0801251111540.13593@iabervon.org","threadId":"11645","inReplyTo":"alpine.LSU.1.00.0801251134570.5731@racer.site","subject":"Re: On pathnames","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-01-25T16:25:05Z","receivedAt":"2008-01-25T16:25:05Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Fri, 25 Jan 2008, Johannes Schindelin wrote:\n\n> Hi,\n> \n> On Thu, 24 Jan 2008, Junio C Hamano wrote:\n> \n> > Daniel Barkalow <barkalow@iabervon.org> writes:\n> > \n> > > $ git checkout branch\n> > > Warning: xt_CONNMARK.c conflicts with xt_connmark.c; not checking it out\n> > > $ git checkout xt_CONNMARK.c --as xt_CONNMARK_caps.c\n> > > $ edit xt_CONNMARK_caps.c\n> > > $ git add xt_CONNMARK_caps.c\n> > \n> > Heh, I like that very much.\n> \n> It would make it easier to test on Linux, too, yes.\n> \n> But then, it would break the build process all the same.\n\nSure, but it would permit a user of a filesystem that can't handle the \nproject to make modifications that generate a commit the filesystem can \nhandle, which is currently pretty difficult.\n\n$ git checkout xt_CONNMARK.c --as xt_CONNMARK_tmp.c\n$ mv xt_CONNMARK_tmp.c xt_connmark_flag.c\n$ edit Makefile\n$ git add xt_connmark_flag.c\n$ git commit -a\n\n(The key thing here being that git will determine that you removed \nxt_CONNMARK.c despite open(xt_CONNMARK.c) returning something unrelated)\n\nRemember that this level of support is to allow users who can't have the \nproject checked out in their filesystems to manipulate the project's data, \nnot to actually make the project work as presented in the filesystem by \ngit.\n\n> And the implementation would _need_ the index extension Linus seems to \n> resent so.\n\nLinus was objecting to having redundant information stored, because it \ncould get skewed. If the information being stored is not redundant (i.e., \nthe normal case is that entries have a flag saying they exist in the \nfilesystem under their own names, and the new cases are that the entry \nisn't present in the filesystem at all or that the entry is present in the \nfilesystem under some other name), that isn't an issue.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"66579","messageId":"alpine.LSU.1.00.0801251733040.5731@racer.site","threadId":"11645","inReplyTo":"alpine.LNX.1.00.0801251111540.13593@iabervon.org","subject":"Re: On pathnames","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-25T17:34:47Z","receivedAt":"2008-01-25T17:34:47Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 25 Jan 2008, Daniel Barkalow wrote:\n\n> On Fri, 25 Jan 2008, Johannes Schindelin wrote:\n> \n> > On Thu, 24 Jan 2008, Junio C Hamano wrote:\n> > \n> > > Daniel Barkalow <barkalow@iabervon.org> writes:\n> > > \n> > > > $ git checkout branch\n> > > > Warning: xt_CONNMARK.c conflicts with xt_connmark.c; not checking it out\n> > > > $ git checkout xt_CONNMARK.c --as xt_CONNMARK_caps.c\n> > > > $ edit xt_CONNMARK_caps.c\n> > > > $ git add xt_CONNMARK_caps.c\n> > > \n> > > Heh, I like that very much.\n> > \n> > It would make it easier to test on Linux, too, yes.\n> > \n> > But then, it would break the build process all the same.\n> \n> Sure, but it would permit a user of a filesystem that can't handle the \n> project to make modifications that generate a commit the filesystem can \n> handle, which is currently pretty difficult.\n> \n> $ git checkout xt_CONNMARK.c --as xt_CONNMARK_tmp.c\n> $ mv xt_CONNMARK_tmp.c xt_connmark_flag.c\n> $ edit Makefile\n> $ git add xt_connmark_flag.c\n> $ git commit -a\n\nAFAICT it is possible right now:\n\n$ git checkout xt_CONNMARK.c\n$ git mv xt_CONNMARK.c xt_connmark_flag.c\n$ git checkout xt_connmark.c\n$ edit Makefile\n$ git add Makefile\n$ git commit\n\nCiao,\nDscho\n"}]}