{"thread":{"id":"11699","subject":"[PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","startedAt":"2008-01-22T04:41:59Z","lastAt":"2008-01-22T10:44:53Z","messageCount":11,"participants":["Sam Vilain","Johannes Schindelin","Junio C Hamano","Mark Junker","Rafael Garcia-Suarez"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"66276","messageId":"20080122050215.DE198200A2@wilber.wgtn.cat-it.co.nz","threadId":"11699","inReplyTo":null,"subject":"[PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Sam Vilain","fromEmail":"sam.vilain@catalyst.net.nz","sentAt":"2008-01-22T04:41:59Z","receivedAt":"2008-01-22T04:41:59Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Some projects may like to enforce a particular encoding is used for\nall filenames in the repository.  Within the UTF-8 encoding, there are\nfour normal forms (see http://unicode.org/reports/tr15/), any of which\nmay be a reasonable repository format choice.  Additionally, some\nfilesystems may have a single encoding that they support when writing\nlocal filenames.  To support this, iconv and a normalization library\nmust have the information they need to perform the correct conversion.\n\nThis is a configuration design proposal, and does not implement any\nchanges.\n---\n   Hi all, I think that restating the problem in these terms might be\n   more productive than the previous discussion, design critiques?\n\n   It is intended that this doesn't impact at all on users with C\n   filesystems without explicit configuration, while adding the feature\n   of allowing projects to specify unicode normalisation (so, eg,\n   Märchen ends up the same as Märchen)\n\n   [apologies if this hits the list twice; I sent the first with a bad\n    content encoding header and assume it got dropped]\n\n Documentation/config.txt        |   16 ++++++++++++++++\n Documentation/gitattributes.txt |   19 +++++++++++++++++++\n Documentation/i18n.txt          |    9 ++++++---\n 3 files changed, 41 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex ee08845..9d2567d 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -146,6 +146,22 @@ core.symlinks::\n \tfile. Useful on filesystems like FAT that do not support\n \tsymbolic links. True by default.\n \n+core.repositoryPathEncoding::\n+\tSpecify the default assumed encoding of repository paths, if\n+\tnot specified in gitlink:gitattributes[3] for that repository.\n+\tThe default value of this is \"C\".\n+\n+core.checkoutPathEncoding::\n+\tSpecify the encoding of local filenames.  The default value of\n+\tthis depends on the platform and filesystem, but for most users\n+\twill be \"C\", indicating no pathname conversion required.\n+\n+core.checkoutPathEncodingFromLocale::\n+\tSpecify whether the checkout path encoding should be\n+\tcontrolled via environment locale variables.  This may have\n+\tsome bizarre side effects if you switch locales between\n+\tworking with a checkout.  False by default.\n+\n core.gitProxy::\n \tA \"proxy command\" to execute (as 'command host port') instead\n \tof establishing direct connection to the remote server when\ndiff --git a/Documentation/gitattributes.txt b/Documentation/gitattributes.txt\nindex cc9c7c5..4136528 100644\n--- a/Documentation/gitattributes.txt\n+++ b/Documentation/gitattributes.txt\n@@ -170,6 +170,25 @@ intent is that if someone unsets the filter driver definition,\n or does not have the appropriate filter program, the project\n should still be usable.\n \n+`encoding`\n+^^^^^^^^^^\n+Specifies the valid encoding for file names (does not affect content)\n+on the specified path.  Git enforces that all filenames are valid in\n+this encoding, and if applicable and possible, will translate from the\n+encoding configured (or, on relevant platform and filesystem\n+combinations, detected) to this encoding.\n+\n+The default value of this is \"C\", which leaves behaviour on\n+filesystems which do not support \"C\" semantics undefined until it is\n+set.  For instance, if your filesystem supports only UTF-8, and you\n+are trying to check out a repository that is in Latin-1, then you will\n+need to configure the repository encoding in `.git/info/attributes` \n+before you can check files out on that system.\n+\n+Valid encodings are currently 'ISO-8859-1' and 'UTF-8'.  'UTF-8' may\n+be followed by '+NFC', '+NFD', '+NFKD' or '+NFKC' to enforce a\n+particular normalization of filenames.\n+\n \n Interaction between checkin/checkout attributes\n ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\ndiff --git a/Documentation/i18n.txt b/Documentation/i18n.txt\nindex b95f99b..fba0407 100644\n--- a/Documentation/i18n.txt\n+++ b/Documentation/i18n.txt\n@@ -1,11 +1,14 @@\n At the core level, git is character encoding agnostic.\n \n  - The pathnames recorded in the index and in the tree objects\n-   are treated as uninterpreted sequences of non-NUL bytes.\n+   are normally treated as uninterpreted sequences of non-NUL bytes.\n    What readdir(2) returns are what are recorded and compared\n    with the data git keeps track of, which in turn are expected\n-   to be what lstat(2) and creat(2) accepts.  There is no such\n-   thing as pathname encoding translation.\n+   to be what lstat(2) and creat(2) accepts.\n+\n+However, if there are configured encodings for the checkout and/or\n+repository, then the defined conversions will occur between the\n+readdir(2) and the index, in both directions.\n \n  - The contents of the blob objects are uninterpreted sequence\n    of bytes.  There is no encoding translation at the core\n-- \n1.5.3.5\n"},{"id":"66277","messageId":"alpine.LSU.1.00.0801220527550.5731@racer.site","threadId":"11699","inReplyTo":"20080122050215.DE198200A2@wilber.wgtn.cat-it.co.nz","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2008-01-22T05:35:40Z","receivedAt":"2008-01-22T05:35:40Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 22 Jan 2008, Sam Vilain wrote:\n\n>  Documentation/gitattributes.txt |   19 +++++++++++++++++++\n\nAs I said on IRC already, I don't think that this is served well as an \n\"attribute\"... it is most likely that the issue either affects _all_ \nfilenames , or _none_.\n\nIn that, it is very similar to the CR/LF issue we encountered.  There \nalso, it depends more on the platform than on the filename if you want to \nenable special handling or not.\n\nI maintain that it is even more obviously a platform issue than CR/LF, \nsince the UTF-8 normalisation takes place in the filesystem driver -- \nregardless if it is needed, or wished for, or not -- whereas CR/LF might \nbe not needed/wished for in one certain project, but might well be wished \nfor in another clone _on the same platform_.\n\nSo I think that this would be a prime candidate for /etc/gitconfig, even \nmore so than core.crlf.\n\nCiao,\nDscho\n"},{"id":"66279","messageId":"7vlk6iv0ik.fsf@gitster.siamese.dyndns.org","threadId":"11699","inReplyTo":"20080122050215.DE198200A2@wilber.wgtn.cat-it.co.nz","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-22T06:26:43Z","receivedAt":"2008-01-22T06:26:43Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sam Vilain <sam.vilain@catalyst.net.nz> writes:\n\n> Some projects may like to enforce a particular encoding is used for\n> all filenames in the repository.  Within the UTF-8 encoding, there are\n> four normal forms (see http://unicode.org/reports/tr15/), any of which\n> may be a reasonable repository format choice.  Additionally, some\n> filesystems may have a single encoding that they support when writing\n> local filenames.  To support this, iconv and a normalization library\n> must have the information they need to perform the correct conversion.\n\nIsn't there a chicken-and-egg problem?  The attributes are by\nnature per-path, and you need to match the pathname string with\na pattern to decide which attribute definition to apply to a\ngiven path.  Before knowing what encoding the pathname you have\njust read from readdir(3), how would you match that pathname\nwith the pattern in the gitattributes file?\n\nI can buy the .git/config (and an in-tree .git-encoding,\nperhaps), though.\n"},{"id":"66280","messageId":"7vejcav01f.fsf@gitster.siamese.dyndns.org","threadId":"11699","inReplyTo":"alpine.LSU.1.00.0801220527550.5731@racer.site","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-22T06:37:00Z","receivedAt":"2008-01-22T06:37:00Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Johannes Schindelin <Johannes.Schindelin@gmx.de> writes:\n\n> On Tue, 22 Jan 2008, Sam Vilain wrote:\n>\n>>  Documentation/gitattributes.txt |   19 +++++++++++++++++++\n>\n> As I said on IRC already, I don't think that this is served well as an \n> \"attribute\"... it is most likely that the issue either affects _all_ \n> filenames , or _none_.\n\nI do not think .gitattributes is the way to go, but I do not\nthink this has to be all or nothing either.\n\nI can well imagine somebody wanting to do:\n\n\tDocumentation/ja/README-spelled-in-Japanese\n\tDocumentation/ja/... other files in Japanese ...\n\tDocumentation/zh/README-spelled-in-Chinese\n\tDocumentation/zh/... other files in Chinese ...\n\nand have all files under Documentation/ja/ in EUC-JP while\nDocumentation/zh/ are BIG5 or whatever (I do not speak nor write\nChinese).\n\nMaybe the project originates from Brasil and the string\n\"Documentation\" itself is spelled as \"Documentação\" and in\nLatin-1 (no, I do not write pt_BR either, and I admit at this\npoint this is a contrived example that I cannot _that_ well\nimagine, but is not so far-fetched).\n\nSo we _could_ have .git-encoding in Documentation/ja/ and\nDocumentation/zh/ each of which says \"this directory and\neverything below are in this encoding, unless overriden\notherwise by a deeper directory\".\n"},{"id":"66283","messageId":"7vr6gatidd.fsf@gitster.siamese.dyndns.org","threadId":"11699","inReplyTo":"7vlk6iv0ik.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-22T07:43:58Z","receivedAt":"2008-01-22T07:43:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Sam Vilain <sam.vilain@catalyst.net.nz> writes:\n>\n>> Some projects may like to enforce a particular encoding is used for\n>> all filenames in the repository.  Within the UTF-8 encoding, there are\n>> four normal forms (see http://unicode.org/reports/tr15/), any of which\n>> may be a reasonable repository format choice.  Additionally, some\n>> filesystems may have a single encoding that they support when writing\n>> local filenames.  To support this, iconv and a normalization library\n>> must have the information they need to perform the correct conversion.\n>\n> Isn't there a chicken-and-egg problem?  The attributes are by\n> nature per-path, and you need to match the pathname string with\n> a pattern to decide which attribute definition to apply to a\n> given path.  Before knowing what encoding the pathname you have\n> just read from readdir(3), how would you match that pathname\n> with the pattern in the gitattributes file?\n>\n> I can buy the .git/config (and an in-tree .git-encoding,\n> perhaps), though.\n\nI admit that Documentação/ja/お読み下さい example was contrived\n(the last component is README-in-Japanese), and if anybody still\nwanted to have such a tree sanely, the only practical\ncross-platform and multi-language way to do so is to have\neverything in UTF-8 at the repository level.\n\nIn that sense, the project does not need to specify anything,\nother than marking that \"all of the pathnames in tree objects\nare in UTF-8 (we could go stronger, and say which kind of\nnormalization we want)\".  As there is no other practical choice\nthan UTF-8-NFC if you want to be cross-platform, compatible, and\nmulti-language, the project can just declare that is what it\nuses and does not have to mark it any specially.\n\nA particular clone of such a project may want to check\neverything out as-is to get an UTF-8 only tree (I'll mention\nHFS+ shortly).  Another clone may want to get mixed legacy\nencodings by running mkdir(utf8_to_latin1(\"Documentação\")) and\ncreat(utf8_to_eucjp(\" お読み下さい\")), but that is purely a\nlocal matter and should not be controlled by anything in-tree,\nbe it .gitattributes or .git-encoding.\n\nOn the other hand, it is not so unusual to see a legacy encoding\nused in the pathnames, especially if your project does not need\nto deal with multi-language issues.  In such a repository, I do\nnot want to enforce that all the paths in tree objects MUST be\nUTF-8.  If all the project participant agree to work with EUC-JP\npathnames in tree objects, we should not make the users always\ngo through double conversion going from readdir(3) to index, and\ncoming from index back to open(2) or creat(2).  Again, that is\ndone by agreement by project participants, so there is nothing\nthat needs to be specified in-tree.\n\nIf the project uses UTF-8-NFC, we would need to adjust check-in\nand check-out codepath like Linus's readdir(3) hack suggested,\nbut that needs to be done only on HFS+.  Of course, the project\nparticipants need to be careful not to create files that HFS+\ncannot handle (two paths that happen to be equivalent strings\nshould not be created), but I do not think that is such a big\nissue as some people seem to make a big deal out of.  If you\nwant to be interoperable with different filesystems, you should\nnot create two paths that are different only in case, and if\nthere are participants who are on such a filesystem, the mistake\nis quickly spotted and corrected.  It happened in git.git to a\nfile other than that infamous Märchen.  It's exactly the same\nissue [*1*].\n\nIn short, initially I did not like Linus's readdir(3) hack very\nmuch, but the more I think about it, I like it the better.\n\nWe pick a reasonable default (i.e. \"no conversion\") at the\ntechnical level, and recommend (but do not pay for the overhead\nof enforcing) a reasonable normalization as the BCP at the human\nlevel.  Only on filesystems that mangle the pathnames, or if you\nwant legacy encodings on the filesystem, we would need to pay\noverhead for conversion and help people with actual code to do\nso.\n\nTo support the above scenarios, I think each instance of\nrepository needs to be able to say \"this path (specified with a\nmatching pattern in the filename encoding) should be converted\nthis way coming in, and that way going out.\"  UTF-8 only project\nwould have NKC<->NKD on HFS+ partition, and nothing on\neverywhere else.  EUC-JP project that checks out as-is would\nspecify nothing either, but people on Shift_JIS platforms would\nlocally specify that EUC-JP <-> Shift_JIS conversion to be made.\n\n\n[Footnote]\n\n*1* This is an important point, especially the breakage was\nabout tests that used files \"a\" and \"A\".  No pathname\nenforcement in git-as-scm would have enforced anything to avoid\nthe breakage.  But there are humans involved in the project and\nthey are an integral part of ensuring interoperability.\n"},{"id":"66288","messageId":"fn48bp$ff8$1@ger.gmane.org","threadId":"11699","inReplyTo":"7vr6gatidd.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Mark Junker","fromEmail":"mjscod@web.de","sentAt":"2008-01-22T08:09:33Z","receivedAt":"2008-01-22T08:09:33Z","isPatch":true,"sender":{"key":"mjscod@web.de","avatar":"https://gravatar.com/avatar/1bd49fe36dddcde665ab9e859d3fd4c9be45dd8ea85b64b6b25f1de179af7af1?d=mp&s=160"},"body":"Junio C Hamano schrieb:\n\n> To support the above scenarios, I think each instance of\n> repository needs to be able to say \"this path (specified with a\n> matching pattern in the filename encoding) should be converted\n> this way coming in, and that way going out.\"  UTF-8 only project\n> would have NKC<->NKD on HFS+ partition, and nothing on\n> everywhere else.  EUC-JP project that checks out as-is would\n> specify nothing either, but people on Shift_JIS platforms would\n> locally specify that EUC-JP <-> Shift_JIS conversion to be made.\n\nJust to sum up what you wrote and to be sure that I understand you \ncorrectly:\n\nLets have two encodings:\n- Encoding for path names stored in the repository\n- Encoding for path names from/to file systems\n\nDo conversion only if they are different. Both encodings are configurable.\n\nRegards,\nMark\n"},{"id":"66292","messageId":"b77c1dce0801220113y3f33c7fjf8fd0cd2c2763274@mail.gmail.com","threadId":"11699","inReplyTo":"7vr6gatidd.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Rafael Garcia-Suarez","fromEmail":"rgarciasuarez@gmail.com","sentAt":"2008-01-22T09:13:12Z","receivedAt":"2008-01-22T09:13:12Z","isPatch":true,"sender":{"key":"rgarciasuarez@gmail.com","avatar":null},"body":"On 22/01/2008, Junio C Hamano wrote:\n> If the project uses UTF-8-NFC, we would need to adjust check-in\n> and check-out codepath like Linus's readdir(3) hack suggested,\n> but that needs to be done only on HFS+.  Of course, the project\n> participants need to be careful not to create files that HFS+\n> cannot handle (two paths that happen to be equivalent strings\n> should not be created), but I do not think that is such a big\n> issue as some people seem to make a big deal out of.  If you\n\nRight, I don't see that as a big issue -- for new files. But we can have\nfiles that were created in the past as non-handleable by HFS+, and later\nrenamed to something more portable.\n\nMore generally, the consensus encoding might change over time. We can\nimagine a project which contains, say, a test file which a latin-1 name,\nthat gets later renamed to a UTF-8 name, (due to a project policy\nchange), but making necessary to adjust the said test. A checkout of the\nearlier version would have that test failing. (But maybe I'm just\nhandwaving towards a non-existent problem here. I'd consider the issue\nas minor anyway.)\n\n> want to be interoperable with different filesystems, you should\n> not create two paths that are different only in case, and if\n> there are participants who are on such a filesystem, the mistake\n> is quickly spotted and corrected.  It happened in git.git to a\n> file other than that infamous Märchen.  It's exactly the same\n> issue [*1*].\n>\n> In short, initially I did not like Linus's readdir(3) hack very\n> much, but the more I think about it, I like it the better.\n>\n> We pick a reasonable default (i.e. \"no conversion\") at the\n> technical level, and recommend (but do not pay for the overhead\n> of enforcing) a reasonable normalization as the BCP at the human\n> level.  Only on filesystems that mangle the pathnames, or if you\n> want legacy encodings on the filesystem, we would need to pay\n> overhead for conversion and help people with actual code to do\n> so.\n>\n> To support the above scenarios, I think each instance of\n> repository needs to be able to say \"this path (specified with a\n> matching pattern in the filename encoding) should be converted\n> this way coming in, and that way going out.\"  UTF-8 only project\n> would have NKC<->NKD on HFS+ partition, and nothing on\n> everywhere else.  EUC-JP project that checks out as-is would\n> specify nothing either, but people on Shift_JIS platforms would\n> locally specify that EUC-JP <-> Shift_JIS conversion to be made.\n\nSounds sane, except maybe the part where you specify paths with a\npattern. Do you really need this layer of complexity? Pattern matching\nin different encodings has proven to be troublesome. Usually that's\nwhere UTF-8 normalisation rules and locale-specific behaviours kick in,\nesp. when you're starting to use \\w or \\d characters classes, or case\ninsensitivity. For example, if you want to do it correctly, \"I\" will\nmatch /i/ case-insensitively, except in Turkish locales... (Sorry, I'm\njust handwaving again here...)\n"},{"id":"66293","messageId":"7vbq7erzj4.fsf@gitster.siamese.dyndns.org","threadId":"11699","inReplyTo":"fn48bp$ff8$1@ger.gmane.org","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-22T09:16:15Z","receivedAt":"2008-01-22T09:16:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Mark Junker <mjscod@web.de> writes:\n\n> Just to sum up what you wrote and to be sure that I understand you\n> correctly:\n>\n> Lets have two encodings:\n> - Encoding for path names stored in the repository\n> - Encoding for path names from/to file systems\n>\n> Do conversion only if they are different. Both encodings are configurable.\n\nNot really.\n\n 1. Encoding for the project does not have to be specified at\n    all.  The project participants are expected to know about it\n    out of band.\n\n 2. Conversion for path names between filesystems and the\n    project (i.e. \"paths in tree objects\") can be specified per\n    repository (i.e. \"a particular clone of the project\").  We\n    could even allow the conversion function to be different\n    per-path-component but I suspect that would be a much\n    future addition that nobody would use in practice.\n\n 3. Suggest use of UTF-8-NFC as the project encoding as a BCP,\n    but never enforce it.  It is a responsibility of the owner\n    of the particular repository to make sure that the\n    conversions used in a particular repository (again, \"a\n    particular clone of the project\") produces the desired\n    encoding in the tree objects.\n\nBut please take these with a moderately large grain of salt, as\nI was more or less handwaving and pretending to know what I was\ntalking about ;-).  I think this should work in theory, but I at\nthe same time suspect that there are many more places than just\nreaddir(3) that need to be wrapped if we take this approach, and\nthe intrusiveness factor might make this infeasible in practice.\n\nThe difference between your version and my 1. and 2. is very\nsubtle, but comes primarily from my desire not to have to use\nthe word \"canonical\".  Yours define \"this canonical encoding is\nused in the repository, and we convert back and forth to that\nlocal encoding\", as opposed to my saying \"here are to and from\nconversion functions\".  The latter is more in line with how we\ndefine smudge/clean filters for blob contents conversion, in\nthat the \"encoding\" used in in-repository blob does not have to\neven have a name.\n"},{"id":"66297","messageId":"4795BE07.4040500@catalyst.net.nz","threadId":"11699","inReplyTo":"7vr6gatidd.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Sam Vilain","fromEmail":"sam.vilain@catalyst.net.nz","sentAt":"2008-01-22T09:57:27Z","receivedAt":"2008-01-22T09:57:27Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> To support the above scenarios, I think each instance of\n> repository needs to be able to say \"this path (specified with a\n> matching pattern in the filename encoding) should be converted\n> this way coming in, and that way going out.\"  UTF-8 only project\n> would have NKC<->NKD on HFS+ partition, and nothing on\n> everywhere else.\n\nI think there is another reason to do this - simple sanity.  Two people\nadding the same filename should not end up with a different tree ID, if\nthey for whatever reason ended up entering a differing equivalent\nvariant of the same Unicode NKC form.\n\nBut, that rule of sanity breaks the C semantics sanity, so it must be a\nper-project setting.  Not a necessity, but a good feature I think.  It\ncan be enforced with external scripts/hooks of course.\n\nWhat happens on the way in and out of the filesystem, I see that as a\nside issue.  Once you define what the normalized form is for the\nproject, then the features should just fall into place without messy\nheuristics.  There is also a correct behaviour when faced with\nfilesystems that have a different idea about who enforces encoding rules\n- so long as you can detect what those ideas are :).  It also means that\nusers can choose to use the same local encoding as their locale, which\nmight interoperate better with other apps.\n\nThe readdir() (case|normalization) tolerance change is good in its own\nright, but it's a slightly different scenario, and an independent\nquestion to what is the normalized form.  Of course, on case folding,\nunicode normalizing filesystems you'd have to have a mixture of these\nsettings for sane operation.\n\nOn the chicken and egg thing, I guess .gitattributes is too late, you're\nright - unless you say that at each directory level, the globbing is\nalways C.  But I haven't thought about that very hard.  I was just\nre-using a mechanism that already exists rather than try to invent\nsomething new.  I do agree with Dscho's point that mixing encodings in a\nrepository is not necessarily a use case worth catering for.\n\nSam.\n"},{"id":"66300","messageId":"7vprvuqh94.fsf@gitster.siamese.dyndns.org","threadId":"11699","inReplyTo":"4795BE07.4040500@catalyst.net.nz","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-01-22T10:36:23Z","receivedAt":"2008-01-22T10:36:23Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sam Vilain <sam.vilain@catalyst.net.nz> writes:\n\n> On the chicken and egg thing, ...\n> ...  I do agree with Dscho's point that mixing encodings in a\n> repository is not necessarily a use case worth catering for.\n\nAre you talking about \"repository\" as in \"a specific clone\", or\n\"a project that can be cloned by many people and checked out to\nsuit cloner's needs\"?  I definitely agree that mixing encodings\nin a project (i.e. \"paths in tree objects\") does not make any\nsense _if_ clones of the projects _may_ want to check things out\nin different pathname encodings from each other.  And if all\nclones would want to check things out the same way, it does not\nreally matter what encoding the paths in tree objects are.\n\nI am not absolutely sure if you are talking about mixing\nencodings depending on parts of the tree in a specific clone (my\nearlier \"Documentação/ja/ お読み下さい\" example).  I would\ncertainly say it would be a very low priority for us to support\nsuch usage, as I imagine that multi-language trees would most\nlikely be checked out in UTF-8 everywhere, but it _might_ be\nsomething people may find real need for.\n"},{"id":"66301","messageId":"4795C925.2090409@catalyst.net.nz","threadId":"11699","inReplyTo":"7vprvuqh94.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH] [RFC] Design for pathname encoding gitattribute [RESEND]","fromName":"Sam Vilain","fromEmail":"sam.vilain@catalyst.net.nz","sentAt":"2008-01-22T10:44:53Z","receivedAt":"2008-01-22T10:44:53Z","isPatch":true,"sender":{"key":"sam@vilain.net","avatar":"https://gravatar.com/avatar/8fc840ca854dbf6f7065b4335e3b934951c1dca3b11db688e95e471901f8f4a8?d=mp&s=160"},"body":"Junio C Hamano wrote:\n> Sam Vilain <sam.vilain@catalyst.net.nz> writes:\n> \n>> On the chicken and egg thing, ...\n>> ...  I do agree with Dscho's point that mixing encodings in a\n>> repository is not necessarily a use case worth catering for.\n> \n> Are you talking about \"repository\" as in \"a specific clone\", or\n> \"a project that can be cloned by many people and checked out to\n> suit cloner's needs\"?  I definitely agree that mixing encodings\n> in a project (i.e. \"paths in tree objects\") does not make any\n> sense _if_ clones of the projects _may_ want to check things out\n> in different pathname encodings from each other.  And if all\n> clones would want to check things out the same way, it does not\n> really matter what encoding the paths in tree objects are.\n\nI'm referring to the normalized form in the object database - ie what\naffects the generated SHA1s - what you check it out to locally is a\ndeveloper's choice, and assuming that they can handle whatever issues\nthey create by doing this, then that should be fine.\n\n> I am not absolutely sure if you are talking about mixing\n> encodings depending on parts of the tree in a specific clone (my\n> earlier \"Documentação/ja/ お読み下さい\" example).  I would\n> certainly say it would be a very low priority for us to support\n> such usage, as I imagine that multi-language trees would most\n> likely be checked out in UTF-8 everywhere, but it _might_ be\n> something people may find real need for.\n\nAgreed - not something you want to condone, but if it's just as easy to\ncome up with a design that doesn't limit to one encoding for a whole\nrepository, it might help some people.\n\nThe use case for mixed encodings I had in mind was when you clone some\nrepository that's got them mixed, and you need to tell git the encoding\nper-path to get the darned thing to behave sensibly for you (presumably\nwhile you write a patch to submit upstream to fix it).\n\nSam.\n"}]}