{"thread":{"id":"2655","subject":"gitweb on kernel.org and UTF-8","startedAt":"2005-11-23T00:03:11Z","lastAt":"2005-11-24T06:19:40Z","messageCount":7,"participants":["Junio C Hamano","H. Peter Anvin","Kay Sievers","Ryan Anderson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"12580","messageId":"7vzmnw9qo0.fsf@assigned-by-dhcp.cox.net","threadId":"2655","inReplyTo":null,"subject":"gitweb on kernel.org and UTF-8","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-23T00:03:11Z","receivedAt":"2005-11-23T00:03:11Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Is it possible that the UTF-8 check in gitweb running on\nkernel.org machine is somehow too strict?\n\nThe following two commits in git.git repository are not showing\nproperly.\n\nI have a track record of getting peoples' names wrong, so I\ndouble checked my commit objects, and as far as I can tell, all\nof them are encoded in UTF-8 properly (or at least I can view\nwhat I expect if I throw raw bytes from the commit objects at my\nFirefox):\n\n        c3df8568424684bbcc7df7722eb3ec34bdae8b2d\n\n        This is from Yoshifuji-san; the third character in\n        author name field is mangled.\n\n\tbb931cf9d73d94d9940b6d0ee56b6c13ad42f1a0\n\n\tThis is from Lukas Sandstr*m; o with Umlaut on top is\n\tshowing a ?.  Incidentally, the blob that records recent\n\tversion of Documentation/git-pack-redundant.txt has his\n\tname in it, which has the same ? problem, but \"plain\"\n\toption shows his name correctly in UTF-8.\n\nInterestingly enough, my name spelled in Japanese\n(Documentatino/git-lost-found.txt) is intact.  Am I getting a\nVIP treatment somehow?\n"},{"id":"12583","messageId":"4383BEE4.1060800@zytor.com","threadId":"2655","inReplyTo":"7vzmnw9qo0.fsf@assigned-by-dhcp.cox.net","subject":"Re: gitweb on kernel.org and UTF-8","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-11-23T00:59:16Z","receivedAt":"2005-11-23T00:59:16Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Junio C Hamano wrote:\n> Is it possible that the UTF-8 check in gitweb running on\n> kernel.org machine is somehow too strict?\n> \n> The following two commits in git.git repository are not showing\n> properly.\n> \n> I have a track record of getting peoples' names wrong, so I\n> double checked my commit objects, and as far as I can tell, all\n> of them are encoded in UTF-8 properly (or at least I can view\n> what I expect if I throw raw bytes from the commit objects at my\n> Firefox):\n> \n>         c3df8568424684bbcc7df7722eb3ec34bdae8b2d\n> \n>         This is from Yoshifuji-san; the third character in\n>         author name field is mangled.\n> \n> \tbb931cf9d73d94d9940b6d0ee56b6c13ad42f1a0\n> \n> \tThis is from Lukas Sandstr*m; o with Umlaut on top is\n> \tshowing a ?.  Incidentally, the blob that records recent\n> \tversion of Documentation/git-pack-redundant.txt has his\n> \tname in it, which has the same ? problem, but \"plain\"\n> \toption shows his name correctly in UTF-8.\n> \n> Interestingly enough, my name spelled in Japanese\n> (Documentatino/git-lost-found.txt) is intact.  Am I getting a\n> VIP treatment somehow?\n> \n\nI think it's missing a \"binmode STDOUT, ':utf8';\" somewhere...\n\nFor what it's worth, I looked at both the above examples and the binary \nencoding in the git repository is undoubtedly correct; the two \ncharacters are U+82F1/E8 8B B1 (英) and U+00F6/C3 B6 (ö) respectively, \nboth of which are 100% valid UTF-8.\n\n\t-hpa\n"},{"id":"12589","messageId":"20051123033526.GA24098@vrfy.org","threadId":"2655","inReplyTo":"4383BEE4.1060800@zytor.com","subject":"Re: gitweb on kernel.org and UTF-8","fromName":"Kay Sievers","fromEmail":"kay.sievers@vrfy.org","sentAt":"2005-11-23T03:35:26Z","receivedAt":"2005-11-23T03:35:26Z","isPatch":false,"sender":{"key":"kay.sievers@vrfy.org","avatar":null},"body":"On Tue, Nov 22, 2005 at 04:59:16PM -0800, H. Peter Anvin wrote:\n> Junio C Hamano wrote:\n> >Is it possible that the UTF-8 check in gitweb running on\n> >kernel.org machine is somehow too strict?\n> >\n> >The following two commits in git.git repository are not showing\n> >properly.\n> >\n> >I have a track record of getting peoples' names wrong, so I\n> >double checked my commit objects, and as far as I can tell, all\n> >of them are encoded in UTF-8 properly (or at least I can view\n> >what I expect if I throw raw bytes from the commit objects at my\n> >Firefox):\n> >\n> >        c3df8568424684bbcc7df7722eb3ec34bdae8b2d\n> >\n> >        This is from Yoshifuji-san; the third character in\n> >        author name field is mangled.\n> >\n> >\tbb931cf9d73d94d9940b6d0ee56b6c13ad42f1a0\n> >\n> >\tThis is from Lukas Sandstr*m; o with Umlaut on top is\n> >\tshowing a ?.  Incidentally, the blob that records recent\n> >\tversion of Documentation/git-pack-redundant.txt has his\n> >\tname in it, which has the same ? problem, but \"plain\"\n> >\toption shows his name correctly in UTF-8.\n> >\n> >Interestingly enough, my name spelled in Japanese\n> >(Documentatino/git-lost-found.txt) is intact.  Am I getting a\n> >VIP treatment somehow?\n> >\n> \n> I think it's missing a \"binmode STDOUT, ':utf8';\" somewhere...\n> \n> For what it's worth, I looked at both the above examples and the binary \n> encoding in the git repository is undoubtedly correct; the two \n> characters are U+82F1/E8 8B B1 (英) and U+00F6/C3 B6 (ö) respectively, \n> both of which are 100% valid UTF-8.\n\nShould be fine now. The escapeHTML() garbled the utf8 \"ö\", and the\ndecode() failed that.\n\nThanks,\nKay\n"},{"id":"12591","messageId":"4383E529.5000608@zytor.com","threadId":"2655","inReplyTo":"20051123033526.GA24098@vrfy.org","subject":"Re: gitweb on kernel.org and UTF-8","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2005-11-23T03:42:33Z","receivedAt":"2005-11-23T03:42:33Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Kay Sievers wrote:\n> \n> Should be fine now. The escapeHTML() garbled the utf8 \"ö\", and the\n> decode() failed that.\n> \n\nIndeed, looks much better.\n\nNow if I could only figure out why both Konsole and Firefox seems to use \na standalone cedilla to represent U+FFFD, instead of something more \nlogical like an inverted question mark or empty box.\n\n\t-hpa\n"},{"id":"12660","messageId":"7vfypm20eh.fsf@assigned-by-dhcp.cox.net","threadId":"2655","inReplyTo":"20051123033526.GA24098@vrfy.org","subject":"Re: gitweb on kernel.org and UTF-8","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-24T03:24:38Z","receivedAt":"2005-11-24T03:24:38Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Kay Sievers <kay.sievers@vrfy.org> writes:\n\n> Should be fine now. The escapeHTML() garbled the utf8 \"ö\", and the\n> decode() failed that.\n\nLooking better.  Thanks.\n\nThis begs for addressing another issue, although I am hesitant\nto open this can of worms at this moment.\n\nIt might be a good idea to have configuration items gitk and\ngitweb can use to get a hint to decide what the commit log\nmessage and the blob data encodings might be.  gitweb already\nadds its own information to the git repository format\n(.git/description), so this _could_ be stored outside just like\nthat (e.g. .git/commit_log_encoding), but using .git/config is\nprobably better.\n\nHow about doing something like this?\n\n\t[i18n]\n        \tcommitEncoding = utf8\n\t\tblobEncoding = utf8\n\nto mean:\n\n\tIf you _have_ to make an assumption on an encoding\n\tcommit and blob objects are in, utf8 is your best bet\n\t(but mistakes can happen, and some blobs can be binary).\n\nThen gitweb and gitk can look at commitEncoding and blobEncoding\nas a hint to base its display defaults on.  Sending everything\nout in utf8 would be sane and safe choice these days for gitweb,\nso if commitEncoding is latin-1 it may need to iconv latin-1 to\nutf8 while reading commits.  For blobs, it might be better off\nasking file(1) or File::MMagic (since you are using Perl in\ngitweb --- sorry I do not know tcl equivalent of that) what they\nare; eventually you would want to be able to show repositories\nfull of jpeg pictures anyway ;-).\n\nOn the commit-producing side, I could have:\n\n\t[i18n]\n        \teditorEncoding = latin-1\n\nand if editorEncoding is different from commitEncoding,\n\"git-commit -c $commit\" would first iconv from utf8 to latin-1\nbefore populating the user's editor, and iconv back from latin-1\nto utf8 before feeding what the user edited to commit-tree.\n\nPathname encoding is the reason why I was hesitant about bring\nthis up.  Although it is too late for 1.0 now, we _could_ have\ndeclared that the paths recorded in git tree objects and index\nfiles are internally utf8, and working tree paths can be in\ndifferent encoding.  As a local repository configuration not\nproject wide configuration, we could have something like:\n\n\t[i18n]\n        \tpathnameEncoding = latin-1\n\nto mean that the filesystem paths returned by readdir(3) and\naccepted by open(2) and friends are in latin-1.  Comparison and\nmovement between working tree files, the index file, and tree\nobjects have to involve iconv and do the right thing.  So if you\nfetch from such a repository into a filesystem that stores\npathnames in utf8, the right thing should happen.\n\nI personally feel any sane project should restrict its pathname\nto ASCII only, so this issue might be moot (or is the right word\n\"mute\"?), but something like this _might_ be useful in later\nversions of git.\n\nBut not in 1.0.\n"},{"id":"12665","messageId":"20051124050104.GC16995@mythryan2.michonline.com","threadId":"2655","inReplyTo":"7vfypm20eh.fsf@assigned-by-dhcp.cox.net","subject":"Re: gitweb on kernel.org and UTF-8","fromName":"Ryan Anderson","fromEmail":"ryan@michonline.com","sentAt":"2005-11-24T05:01:04Z","receivedAt":"2005-11-24T05:01:04Z","isPatch":false,"sender":{"key":"ryan@michonline.com","avatar":null},"body":"On Wed, Nov 23, 2005 at 07:24:38PM -0800, Junio C Hamano wrote:\n> \n> How about doing something like this?\n> \n> \t[i18n]\n>         \tcommitEncoding = utf8\n> \t\tblobEncoding = utf8\n> \n> to mean:\n> \n> \tIf you _have_ to make an assumption on an encoding\n> \tcommit and blob objects are in, utf8 is your best bet\n> \t(but mistakes can happen, and some blobs can be binary).\n\nThe rest of the options help clarify this, but can you make these\noptions 'assumeCommitEncoding' and 'assumeBlobEncoding' to make it clear\nthat these are *assumptions* and not actually controlling what gets\nwritten?\n\n\n-- \n\nRyan Anderson\n  sometimes Pug Majere\n"},{"id":"12667","messageId":"7vsltmy3cz.fsf@assigned-by-dhcp.cox.net","threadId":"2655","inReplyTo":"20051124050104.GC16995@mythryan2.michonline.com","subject":"Re: gitweb on kernel.org and UTF-8","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2005-11-24T06:19:40Z","receivedAt":"2005-11-24T06:19:40Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ryan Anderson <ryan@michonline.com> writes:\n\n> On Wed, Nov 23, 2005 at 07:24:38PM -0800, Junio C Hamano wrote:\n>> \n>> How about doing something like this?\n>> \n>> \t[i18n]\n>>         \tcommitEncoding = utf8\n>> \t\tblobEncoding = utf8\n>> \n>> to mean:\n>> \n>> \tIf you _have_ to make an assumption on an encoding\n>> \tcommit and blob objects are in, utf8 is your best bet\n>> \t(but mistakes can happen, and some blobs can be binary).\n>\n> The rest of the options help clarify this, but can you make these\n> options 'assumeCommitEncoding' and 'assumeBlobEncoding' to make it clear\n> that these are *assumptions* and not actually controlling what gets\n> written?\n\nAs I outlined in the \"editorEncoding\" part, if everything works\nas planned, your latin-1 editing editor would leave latin-1\nmessage for git-commit to pick up (or command line \"-m $msg\"\noption would be encoded in latin-1), and iconv would munge that\nto utf8 to feed commit-tree (because of \"commitEncoding\" being\nutf8). In that sense, commitEncoding is not assumption for the\nwriters.  If everybody, including outside sources we merge from,\nmakes best effort not to screw up, these settings would\nfaithfully describe what encoding logs are in.\n\nBut writers can screw up, and funnily encoded commit messages\nmerge from outside source brings in cannot be fixed after the\nfact, so \"assume\" part must be implied anyway for readers.\n"}]}