{"thread":{"id":"36361","subject":"[PATCH] Unicode: update of combining code points","startedAt":"2014-04-07T19:30:13Z","lastAt":"2014-04-24T09:02:17Z","messageCount":7,"participants":["Torsten Bögershausen","Peter Krefting","Kevin Bracey"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"238503","messageId":"201404072130.15686.tboegi@web.de","threadId":"36361","inReplyTo":null,"subject":"[PATCH] Unicode: update of combining code points","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2014-04-07T19:30:13Z","receivedAt":"2014-04-07T19:30:13Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"Unicode 6.3 defines the following code as combining or accents,\ngit_wcwidth() should return 0.\n\nEarlier unicode standards had defined these code point as \"reserved\":\n\n358 COMBINING DOT ABOVE RIGHT\n359 COMBINING ASTERISK BELOW\n35A COMBINING DOUBLE RING BELOW\n35B COMBINING ZIGZAG ABOVE\n35C COMBINING DOUBLE BREVE BELOW\n487 COMBINING CYRILLIC POKRYTIE\n5A2 HEBREW ACCENT ATNAH HAFUKH,\n5BA HEBREW POINT HOLAM HASER FOR VAV\n5C5 HEBREW MARK LOWER DOT\n5C7 HEBREW POINT QAMATS QATAN\n604 ARABIC SIGN SAMVAT\n616 ARABIC SMALL HIGH LIGATURE ALEF WITH LAM WITH YEH\n617 ARABIC SMALL HIGH ZAIN\n618 ARABIC SMALL FATHA\n619 ARABIC SMALL DAMMA\n61A ARABIC SMALL KASRA\n659 ARABIC ZWARAKAY\n65A ARABIC VOWEL SIGN SMALL V ABOVE\n65B ARABIC VOWEL SIGN INVERTED SMALL V ABOVE\n65C ARABIC VOWEL SIGN DOT BELOW\n65D ARABIC REVERSED DAMMA\n65E ARABIC FATHA WITH TWO DOTS\n65F ARABIC WAVY HAMZA BELOW\n\nThis commit touches only the range 300-6FF, there may be more to be updated.\n\nSigned-off-by: Torsten Bögershausen <tboegi@web.de>\n---\n utf8.c | 9 ++++-----\n 1 file changed, 4 insertions(+), 5 deletions(-)\n\ndiff --git a/utf8.c b/utf8.c\nindex a831d50..77c28d4 100644\n--- a/utf8.c\n+++ b/utf8.c\n@@ -84,11 +84,10 @@ static int git_wcwidth(ucs_char_t ch)\n \t *   \"uniset +cat=Me +cat=Mn +cat=Cf -00AD +1160-11FF +200B c\".\n \t */\n \tstatic const struct interval combining[] = {\n-\t\t{ 0x0300, 0x0357 }, { 0x035D, 0x036F }, { 0x0483, 0x0486 },\n-\t\t{ 0x0488, 0x0489 }, { 0x0591, 0x05A1 }, { 0x05A3, 0x05B9 },\n-\t\t{ 0x05BB, 0x05BD }, { 0x05BF, 0x05BF }, { 0x05C1, 0x05C2 },\n-\t\t{ 0x05C4, 0x05C4 }, { 0x0600, 0x0603 }, { 0x0610, 0x0615 },\n-\t\t{ 0x064B, 0x0658 }, { 0x0670, 0x0670 }, { 0x06D6, 0x06E4 },\n+\t\t{ 0x0300, 0x036F }, { 0x0483, 0x0489 }, { 0x0591, 0x05BD },\n+\t\t{ 0x05BF, 0x05BF }, { 0x05C1, 0x05C2 }, { 0x05C4, 0x05C5 },\n+\t\t{ 0x05C7, 0x05C7 }, { 0x0600, 0x0604 }, { 0x0610, 0x061A },\n+\t\t{ 0x064B, 0x065F }, { 0x0670, 0x0670 }, { 0x06D6, 0x06E4 },\n \t\t{ 0x06E7, 0x06E8 }, { 0x06EA, 0x06ED }, { 0x070F, 0x070F },\n \t\t{ 0x0711, 0x0711 }, { 0x0730, 0x074A }, { 0x07A6, 0x07B0 },\n \t\t{ 0x0901, 0x0902 }, { 0x093C, 0x093C }, { 0x0941, 0x0948 },\n-- \n1.9.0\n"},{"id":"238896","messageId":"alpine.DEB.2.00.1404152009020.29301@ds9.cixit.se","threadId":"36361","inReplyTo":"201404072130.15686.tboegi@web.de","subject":"Re: [PATCH] Unicode: update of combining code points","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2014-04-15T19:10:03Z","receivedAt":"2014-04-15T19:10:03Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Torsten Bögershausen:\n\n> diff --git a/utf8.c b/utf8.c\n> index a831d50..77c28d4 100644\n> --- a/utf8.c\n> +++ b/utf8.c\n\nIs there a script that generates this code from the Unicode database \nfiles, or did you hand-update it?\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"238916","messageId":"534E0B84.6070602@web.de","threadId":"36361","inReplyTo":"alpine.DEB.2.00.1404152009020.29301@ds9.cixit.se","subject":"Re: [PATCH] Unicode: update of combining code points","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2014-04-16T04:48:04Z","receivedAt":"2014-04-16T04:48:04Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 15.04.14 21:10, Peter Krefting wrote:\n> Torsten Bögershausen:\n> \n>> diff --git a/utf8.c b/utf8.c\n>> index a831d50..77c28d4 100644\n>> --- a/utf8.c\n>> +++ b/utf8.c\n> \n> Is there a script that generates this code from the Unicode database files, or did you hand-update it?\n> \nSome of the code points which have \"0 length on the display\" are called\n\"combining\", others are called \"vowels\" or \"accents\".\nE.g. 5BF is not marked any of them, but if you look at the glyph, it should\nbe combining (please correct me if that is wrong).\n\nIf I could have found a file which indicates for each code point, what it\nis, I could write a script.\n\nSo yes, it is updated by hand.\n"},{"id":"238929","messageId":"534E60BF.5020602@bracey.fi","threadId":"36361","inReplyTo":"534E0B84.6070602@web.de","subject":"Re: [PATCH] Unicode: update of combining code points","fromName":"Kevin Bracey","fromEmail":"kevin@bracey.fi","sentAt":"2014-04-16T10:51:43Z","receivedAt":"2014-04-16T10:51:43Z","isPatch":true,"sender":{"key":"kevin@bracey.fi","avatar":"https://avatars.githubusercontent.com/u/96079793?v=4"},"body":"On 16/04/2014 07:48, Torsten Bögershausen wrote:\n> On 15.04.14 21:10, Peter Krefting wrote:\n>> Torsten Bögershausen:\n>>\n>>> diff --git a/utf8.c b/utf8.c\n>>> index a831d50..77c28d4 100644\n>>> --- a/utf8.c\n>>> +++ b/utf8.c\n>> Is there a script that generates this code from the Unicode database files, or did you hand-update it?\n>>\n> Some of the code points which have \"0 length on the display\" are called\n> \"combining\", others are called \"vowels\" or \"accents\".\n> E.g. 5BF is not marked any of them, but if you look at the glyph, it should\n> be combining (please correct me if that is wrong).\n\nIndeed it is combining (more specifically it has General Category \n\"Nonspacing_Mark\" = \"Mn\").\n\n>\n> If I could have found a file which indicates for each code point, what it\n> is, I could write a script.\n>\n\nThe most complete and machine-readable data are in these files:\n\nhttp://www.unicode.org/Public/UCD/latest/ucd/UnicodeData.txt\nhttp://www.unicode.org/Public/UCD/latest/ucd/EastAsianWidth.txt\n\nThe general categories can also be seen more legibly in:\n\nhttp://www.unicode.org/Public/UCD/latest/ucd/extracted/DerivedGeneralCategory.txt\n\nFor docs, see:\n\nhttp://www.unicode.org/reports/tr44/\nhttp://www.unicode.org/reports/tr11/\nhttp://www.unicode.org/ucd/\n\nThe existing utf8.c comments describe the attributes being selected from \nthe tables (general categories \"Cf\",\"Mn\",\"Me\", East Asian Width \"W\", \n\"F\"). And they suggest that the combining character table was originally \nauto-generated from UnicodeData.txt with a \"uniset\" tool. Presumably this?\n\nhttps://github.com/depp/uniset\n\nThe fullwidth-checking code looks like it was done by hand, although \napparently uniset can process EastAsianWidth.txt.\n\nKevin\n"},{"id":"238980","messageId":"534EE0E7.2030608@web.de","threadId":"36361","inReplyTo":"534E60BF.5020602@bracey.fi","subject":"Re: [PATCH] Unicode: update of combining code points","fromName":"Torsten Bögershausen","fromEmail":"tboegi@web.de","sentAt":"2014-04-16T19:58:31Z","receivedAt":"2014-04-16T19:58:31Z","isPatch":true,"sender":{"key":"tboegi@web.de","avatar":"https://avatars.githubusercontent.com/u/7138363?v=4"},"body":"On 2014-04-16 12.51, Kevin Bracey wrote:\n> On 16/04/2014 07:48, Torsten Bögershausen wrote:\n>> On 15.04.14 21:10, Peter Krefting wrote:\n>>> Torsten Bögershausen:\n>>>\n>>>> diff --git a/utf8.c b/utf8.c\n>>>> index a831d50..77c28d4 100644\n>>>> --- a/utf8.c\n>>>> +++ b/utf8.c\n>>> Is there a script that generates this code from the Unicode database files, or did you hand-update it?\n>>>\n>> Some of the code points which have \"0 length on the display\" are called\n>> \"combining\", others are called \"vowels\" or \"accents\".\n>> E.g. 5BF is not marked any of them, but if you look at the glyph, it should\n>> be combining (please correct me if that is wrong).\n> \n> Indeed it is combining (more specifically it has General Category \"Nonspacing_Mark\" = \"Mn\").\n> \n>>\n>> If I could have found a file which indicates for each code point, what it\n>> is, I could write a script.\n>>\n> \n> The most complete and machine-readable data are in these files:\n> \n> http://www.unicode.org/Public/UCD/latest/ucd/UnicodeData.txt\n> http://www.unicode.org/Public/UCD/latest/ucd/EastAsianWidth.txt\n> \n> The general categories can also be seen more legibly in:\n> \n> http://www.unicode.org/Public/UCD/latest/ucd/extracted/DerivedGeneralCategory.txt\n> \n> For docs, see:\n> \n> http://www.unicode.org/reports/tr44/\n> http://www.unicode.org/reports/tr11/\n> http://www.unicode.org/ucd/\n> \n> The existing utf8.c comments describe the attributes being selected from the tables (general categories \"Cf\",\"Mn\",\"Me\", East Asian Width \"W\", \"F\"). And they suggest that the combining character table was originally auto-generated from UnicodeData.txt with a \"uniset\" tool. Presumably this?\n> \n> https://github.com/depp/uniset\n> \n> The fullwidth-checking code looks like it was done by hand, although apparently uniset can process EastAsianWidth.txt.\n> \n> Kevin\nExcellent, thanks for the pointers.\nRunning the script below shows that \n\"0X00AD SOFT HYPHEN\" should have zero length (and some others too).\nI wonder if that is really the case, and which one of the last 2 lines \nin the script is the right one.\n\nWhat does this mean for us:\n\"Cf \tFormat \ta format control character\"\n\n\n#!/bin/sh\n\nif ! test -f UnicodeData.txt; then\n  wget http://www.unicode.org/Public/UCD/latest/ucd/UnicodeData.txt\nfi &&\nif ! test -f EastAsianWidth.txt; then\n  wget http://www.unicode.org/Public/UCD/latest/ucd/EastAsianWidth.txt\nfi\nif ! test -f DerivedGeneralCategory.txt; then\n  wget http://www.unicode.org/Public/UCD/latest/ucd/extracted/DerivedGeneralCategory.txt\nfi &&\nif ! test -d uniset; then\n  git clone https://github.com/tboegi/uniset.git\nfi &&\n(\n  cd uniset &&\n  if ! test -x uniset; then \n    autoreconf -i &&\n    ./configure --enable-warnings=-Werror CFLAGS='-O0 -ggdb'\n  fi &&\n  make\n) &&\nUNICODE_DIR=. ./uniset/uniset --32 cat:Me,Mn,Cf\n#UNICODE_DIR=. ./uniset/uniset --32 cat:Me,Mn\n\n\n\n\n\n\n\n\n\n\n> \n> -- \n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"239007","messageId":"534F7582.9090303@bracey.fi","threadId":"36361","inReplyTo":"534EE0E7.2030608@web.de","subject":"Re: [PATCH] Unicode: update of combining code points","fromName":"Kevin Bracey","fromEmail":"kevin@bracey.fi","sentAt":"2014-04-17T06:32:34Z","receivedAt":"2014-04-17T06:32:34Z","isPatch":true,"sender":{"key":"kevin@bracey.fi","avatar":"https://avatars.githubusercontent.com/u/96079793?v=4"},"body":"On 16/04/2014 22:58, Torsten Bögershausen wrote:\n> Excellent, thanks for the pointers.\n> Running the script below shows that\n> \"0X00AD SOFT HYPHEN\" should have zero length (and some others too).\n> I wonder if that is really the case, and which one of the last 2 lines\n> in the script is the right one.\n>\n> What does this mean for us:\n> \"Cf \tFormat \ta format control character\"\n>\nMaybe dig back through the Git logs to check the original logic, but the \ncomments suggest that \"Cf\" characters have been viewed as zero-width. \nThat makes sense - they're usually markers indicating things like \nbidirectional text flow, so won't be taking space. (Although they may be \ncausing even more extreme layout effects...)\n\nSoft-hyphen is noted as an explicit exception to the rule in the utf8.c \ncomments. As of Unicode 4.0, it's supposed to be a character indicating \na point where a hyphen could be placed if a line-wrap occurs, and if \nthat wrap happens, then it can actually take up 1 space, otherwise not. \nSo its width could be either 0 or 1, depending. Or, quite likely, the \nterminal doesn't treat it specially, and it always just looks like a \nhyphen... Thus we err on the safe side and give it width 1.\n\nSee http://en.wikipedia.org/wiki/Soft_hyphen for background.\n\nThe comments suggest adding \"-00AD +1160-11FF\" to the uniset command \nline for that tweak and for composing Hangul. (The +200B tweak isn't \nnecessary any more - Zero-Width Space U+200B became Cf officially in \nUnicode 4.0.1:\n\nhttp://en.wikipedia.org/wiki/Zero-width_space\nhttp://www.unicode.org/review/resolved-pri.html#pri21\n)\n\nAll of this is only really an approximation - a best-effort attempt to \nfigure out the width of a string without any actual communication with \nthe display device. So it'll never be perfect. The choice between double \nand single width in particular will often be unpredictable, unless you \nhad deeper locale knowledge.\n\nActually, while doing this, I've realised that this was originally \nMarkus Kuhn's implementation, and that is acknowledged at the top of the \nfile:\n\nhttp://www.cl.cam.ac.uk/~mgk25/ucs/wcwidth.c\n\nGood, because he knows what he's doing.\n\nKevin\n"},{"id":"239548","messageId":"alpine.DEB.2.00.1404240949440.28469@ds9.cixit.se","threadId":"36361","inReplyTo":"534E0B84.6070602@web.de","subject":"Re: [PATCH] Unicode: update of combining code points","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2014-04-24T09:02:17Z","receivedAt":"2014-04-24T09:02:17Z","isPatch":true,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Torsten Bögershausen:\n\n> Some of the code points which have \"0 length on the display\" are called\n> \"combining\", others are called \"vowels\" or \"accents\".\n> E.g. 5BF is not marked any of them, but if you look at the glyph, it should\n> be combining (please correct me if that is wrong).\n\nAll combining characters has a non-zero combining class in \nhttp://www.unicode.org/Public/UNIDATA/UnicodeData.txt (fourth field, \ncalled Canonical_Combining_Class in \nhttp://www.unicode.org/reports/tr44/ ). For instance, the aforementioned \nU+05BF is defined as follows:\n\n   05BF;HEBREW POINT RAFE;Mn;23;NSM;;;;;N;;;;;\n\nThe combining class is 23, so this is a combining character.\n\nThere is a difference between non-spacing combining marks (\"Mn\" in the \nthird column (General_Category)) and others (\"Mc\" for spacing marks \nand \"Me\" for enclosing marks), so they might need specifial handling. \nAdditionally, you have the \"zero-width\" characters, such as U+200B \nZero Width Space. These have the \"Cf\" class, although it also contains \nvisible characters IIRC.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"}]}