{"thread":{"id":"19859","subject":"[PATCH] t8005: Nobody writes Russian in shift_jis","startedAt":"2009-06-19T02:18:37Z","lastAt":"2009-06-21T10:07:53Z","messageCount":4,"participants":["Junio C Hamano","Alexander Gavrilov","Brandon Casey"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"116609","messageId":"7vmy85m0ea.fsf@alter.siamese.dyndns.org","threadId":"19859","inReplyTo":null,"subject":"[PATCH] t8005: Nobody writes Russian in shift_jis","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-06-19T02:18:37Z","receivedAt":"2009-06-19T02:18:37Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"The second and third tests of this script expected that Russian strings\nare converted between ISO-8859-5 and Shift_JIS in the \"blame --porcelain\"\nformat output correctly.\n\nSure, many platforms may convert between such a combination, but that is\nonly because one of the base character set of Shift_JIS, JIS X 0208,\ndefines codepoints for Russian characters (among others); I do not think\nanybody uses Shift_JIS when seriously writing Russian, and it is perfectly\nunderstandable if iconv() libraries on some platforms fail converting\nbetween this combination, as it does not matter in reality.\n\nThis patch changes the test to verify Japanese strings are converted\ncorrectly between EUC-JP and Shift_JIS in the same procedure.  The point\nof the test is not about verifying the platform's iconv() library, but to\nsee if \"git blame\" makes correct iconv() library calls when it should.\n\nWe could instead use ISO-8859-5 and KOI8-R as the combination, because\nthey are both meant to represent Russian, in order to make this test\nmeaningful on more platforms, but we already use Shift_JIS vs EUC-JP\ncombinations to test other programs in our test suite, so this combination\nis safer from the point of view of the portability.  Besides, I do not\nread nor write Russian; sorry ;-)\n\nThis change allows tests to pass on my (friend's) Solaris 5.11 box.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n\n * I am Cc'ing Alexander because he originally wrote this test using\n   cp1251 and shift_jis, and I could be wrong in saying that nobody sane\n   writes Russian in shift_jis.\n\n   To allow 7-bit mailpath to pass this patch through, I tentatively\n   dropped this in my t/t8005 directory (the file is not tracked):\n\n\t$ echo '*.txt binary' >t/t8005/.gitattributes\n   \n   before running format-patch on this commit.\n\n t/t8005-blame-i18n.sh |   26 +++++++++++++-------------\n t/t8005/euc-japan.txt |  Bin 0 -> 66 bytes\n t/t8005/iso8859-5.txt |  Bin 74 -> 0 bytes\n t/t8005/sjis.txt      |  Bin 100 -> 56 bytes\n t/t8005/utf8.txt      |  Bin 100 -> 71 bytes\n 5 files changed, 13 insertions(+), 13 deletions(-)\n create mode 100644 t/t8005/euc-japan.txt\n delete mode 100644 t/t8005/iso8859-5.txt\n\ndiff --git a/t/t8005-blame-i18n.sh b/t/t8005-blame-i18n.sh\nindex 9cca14d..cb39055 100755\n--- a/t/t8005-blame-i18n.sh\n+++ b/t/t8005-blame-i18n.sh\n@@ -4,7 +4,7 @@ test_description='git blame encoding conversion'\n . ./test-lib.sh\n \n . \"$TEST_DIRECTORY\"/t8005/utf8.txt\n-. \"$TEST_DIRECTORY\"/t8005/iso8859-5.txt\n+. \"$TEST_DIRECTORY\"/t8005/euc-japan.txt\n . \"$TEST_DIRECTORY\"/t8005/sjis.txt\n \n test_expect_success 'setup the repository' '\n@@ -13,10 +13,10 @@ test_expect_success 'setup the repository' '\n \tgit add file &&\n \tgit commit --author \"$UTF8_NAME <utf8@localhost>\" -m \"$UTF8_MSG\" &&\n \n-\techo \"ISO-8859-5 LINE\" >> file &&\n+\techo \"EUC-JAPAN LINE\" >> file &&\n \tgit add file &&\n-\tgit config i18n.commitencoding ISO8859-5 &&\n-\tgit commit --author \"$ISO8859_5_NAME <iso8859-5@localhost>\" -m \"$ISO8859_5_MSG\" &&\n+\tgit config i18n.commitencoding eucJP &&\n+\tgit commit --author \"$EUC_JAPAN_NAME <euc-japan@localhost>\" -m \"$EUC_JAPAN_MSG\" &&\n \n \techo \"SJIS LINE\" >> file &&\n \tgit add file &&\n@@ -41,17 +41,17 @@ test_expect_success \\\n '\n \n cat >expected <<EOF\n-author $ISO8859_5_NAME\n-summary $ISO8859_5_MSG\n-author $ISO8859_5_NAME\n-summary $ISO8859_5_MSG\n-author $ISO8859_5_NAME\n-summary $ISO8859_5_MSG\n+author $EUC_JAPAN_NAME\n+summary $EUC_JAPAN_MSG\n+author $EUC_JAPAN_NAME\n+summary $EUC_JAPAN_MSG\n+author $EUC_JAPAN_NAME\n+summary $EUC_JAPAN_MSG\n EOF\n \n test_expect_success \\\n \t'blame respects i18n.logoutputencoding' '\n-\tgit config i18n.logoutputencoding ISO8859-5 &&\n+\tgit config i18n.logoutputencoding eucJP &&\n \tgit blame --incremental file | \\\n \t\tegrep \"^(author|summary) \" > actual &&\n \ttest_cmp actual expected\n@@ -76,8 +76,8 @@ test_expect_success \\\n cat >expected <<EOF\n author $SJIS_NAME\n summary $SJIS_MSG\n-author $ISO8859_5_NAME\n-summary $ISO8859_5_MSG\n+author $EUC_JAPAN_NAME\n+summary $EUC_JAPAN_MSG\n author $UTF8_NAME\n summary $UTF8_MSG\n EOF\ndiff --git a/t/t8005/euc-japan.txt b/t/t8005/euc-japan.txt\nnew file mode 100644\nindex 0000000000000000000000000000000000000000..288f040c99f6b61559e3ad964a1247d4b9fd62a3\nGIT binary patch\nliteral 66\nzcmZ<_b&mIP3~=;|_jB}hwN=`^`REaaLkG_9QsQ!jOZf)7+bS)+w)D-yJxd=fIk)uK\nR(w$3BEIGbp=fcHGTmX#39_;`C\n\nliteral 0\nHcmV?d00001\n\ndiff --git a/t/t8005/iso8859-5.txt b/t/t8005/iso8859-5.txt\ndeleted file mode 100644\nindex 2e4b80c8df4da30722561049c46cca778e49cd2f..0000000000000000000000000000000000000000\nGIT binary patch\nliteral 0\nHcmV?d00001\n\nliteral 74\nzcmeYa_P4MwwTw57_jB}hwN=`2>B3!w{Z}77xOeHsbA^L9uG|B%l(;<M%6x;}ZIupP\nZefa3!rF&Nu9^Sim@#WRKH?Asi0RTCvCqV!J\n\ndiff --git a/t/t8005/sjis.txt b/t/t8005/sjis.txt\nindex 2ccfbad207c6e96b1f4f528031d9e4938d364b92..bbdefeaced4b54f98e5d9a85ddd8e0d7346fe7e3 100644\nGIT binary patch\nliteral 56\nzcmWIc@(hmmbM$q!Rq6|xoUAZ$-;78lu3(U;Z?L<qQgdl@Ph)g*L(`e&)aHoh^roXt\nJ+Z&yfxBxx66{!FK\n\nliteral 100\nzcmWIc@(hmmbM$q!Rci5UDQYQbsZ(ePXen)JX=!R{018yLbSkt20jUxo7c8X26%5kk\nk8|)6$6AV<^3{(tK+R##}0OT|PVPQ)*P@)c~tyGB%00x;U@Bjb+\n\ndiff --git a/t/t8005/utf8.txt b/t/t8005/utf8.txt\nindex f46cfc56d80797740c3ec15e166add052f905fcb..4d00dbea7659ee27fda283e7e45cfb2d5f6ea4d1 100644\nGIT binary patch\nliteral 71\nzcmWFyakGf`bM$q!ReHK{<MSyS6rL_w^|HB7i7ON&;~VU5tMs^e+T-RmkDK>AZeH-X\nbaoywQw#Q97A2)YAZe0GjapvQOCM7Na%AzF~\n\nliteral 100\nzcmWFyakGf`bM$q!Rk|?a!lnxwF6>pfF#p2Vi%l0BF6;ve?6}yjaADzv9T&D-*as0(\nt;tB<6@(p$e>RAL-+IX=EtaRUntqK<#fy{juHeT$!u=T=Tpth|_TmU(XJDmUk\n\n-- \n1.6.3.2.316.gda4e4\n"},{"id":"116622","messageId":"bb6f213e0906190325s8fd18d7u8cc29d710cf2e286@mail.gmail.com","threadId":"19859","inReplyTo":"7vmy85m0ea.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] t8005: Nobody writes Russian in shift_jis","fromName":"Alexander Gavrilov","fromEmail":"angavrilov@gmail.com","sentAt":"2009-06-19T10:25:52Z","receivedAt":"2009-06-19T10:25:52Z","isPatch":true,"sender":{"key":"angavrilov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/42666?v=4"},"body":"On Fri, Jun 19, 2009 at 6:18 AM, Junio C Hamano<gitster@pobox.com> wrote:\n>  * I am Cc'ing Alexander because he originally wrote this test using\n>   cp1251 and shift_jis, and I could be wrong in saying that nobody sane\n>   writes Russian in shift_jis.\n\nWell, certainly not intentionally, but I've managed to send a few\nwork-related emails in sjis accidentally (resulting in much confusion\nfor the people on the other side), and thought it is a bit funny :)\n\nI'd guess that almost nobody uses iso8859-5 as well, though. Nobody\nthat I know, anyway.\n\nAlexander\n"},{"id":"116628","messageId":"dfYgk9RFOucTCHxtLQsMXejAeKlGJg-R15sTW_RFemcrjsjqoYD0eg@cipher.nrlssc.navy.mil","threadId":"19859","inReplyTo":"7vmy85m0ea.fsf@alter.siamese.dyndns.org","subject":"Re: [PATCH] t8005: Nobody writes Russian in shift_jis","fromName":"Brandon Casey","fromEmail":"casey@nrlssc.navy.mil","sentAt":"2009-06-19T14:54:56Z","receivedAt":"2009-06-19T14:54:56Z","isPatch":true,"sender":{"key":"drafnel@gmail.com","avatar":"https://avatars.githubusercontent.com/u/921167?v=4"},"body":"Junio C Hamano wrote:\n> The second and third tests of this script expected that Russian strings\n> are converted between ISO-8859-5 and Shift_JIS in the \"blame --porcelain\"\n> format output correctly.\n> \n> Sure, many platforms may convert between such a combination, but that is\n> only because one of the base character set of Shift_JIS, JIS X 0208,\n> defines codepoints for Russian characters (among others); I do not think\n> anybody uses Shift_JIS when seriously writing Russian, and it is perfectly\n> understandable if iconv() libraries on some platforms fail converting\n> between this combination, as it does not matter in reality.\n> \n> This patch changes the test to verify Japanese strings are converted\n> correctly between EUC-JP and Shift_JIS in the same procedure.  The point\n> of the test is not about verifying the platform's iconv() library, but to\n> see if \"git blame\" makes correct iconv() library calls when it should.\n> \n> We could instead use ISO-8859-5 and KOI8-R as the combination, because\n> they are both meant to represent Russian, in order to make this test\n> meaningful on more platforms, but we already use Shift_JIS vs EUC-JP\n> combinations to test other programs in our test suite, so this combination\n> is safer from the point of view of the portability.  Besides, I do not\n> read nor write Russian; sorry ;-)\n> \n> This change allows tests to pass on my (friend's) Solaris 5.11 box.\n\nNo change on my systems.  I can convert eucJP and SJIS from/to UTF-8, but\nI cannot convert between eucJP and SJIS.  So tests 2 and 3 still fail for\nme.  Nothing was broken though.  The fourth test still passes which converts\neach of the encodings to UTF-8.  So this patch is fine with me.\n\n-brandon\n"},{"id":"116710","messageId":"7vfxdtsxvq.fsf@alter.siamese.dyndns.org","threadId":"19859","inReplyTo":"dfYgk9RFOucTCHxtLQsMXejAeKlGJg-R15sTW_RFemcrjsjqoYD0eg@cipher.nrlssc.navy.mil","subject":"Re: [PATCH] t8005: Nobody writes Russian in shift_jis","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-06-21T10:07:53Z","receivedAt":"2009-06-21T10:07:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Brandon Casey <casey@nrlssc.navy.mil> writes:\n\n> No change on my systems.  I can convert eucJP and SJIS from/to UTF-8, but\n> I cannot convert between eucJP and SJIS.\n\nI wonder what's different, but I suspect having lang-support-japanese\npackage on the box perhaps is helping me.\n\n> So tests 2 and 3 still fail for\n> me.  Nothing was broken though.  The fourth test still passes which converts\n> each of the encodings to UTF-8.  So this patch is fine with me.\n\nYikes, so it does not really help by itself.  Taken together with\nAlexander's comment that he did manage to send Russian in Shift_JIS (I\nsomehow do not think Alexander used Solaris for that, though; neither have\nI any clue if the receiving end grokked that), perhaps the patch is\nuseless.\n\nEven though I do not think if any Russian writes in KOI8-R and converts to\nShift_JIS on purpose, converting eucJP directly to SJIS is something\nJapanese people who are on UNIX do quite often, or at least used to before\neverybody moved to UTF-8.\n\nPerhaps we should instead optionally help platform's iconv(3), when it\ncannot convert A to B directly, by pivoting the conversion on UTF-8\n(i.e. A -> UTF-8 -> B)?  That would probably help the real world use cases\nwhile fixing the issue with this test script.\n"}]}