{"thread":{"id":"57071","subject":"[PATCH v2 1/3] git-p4: remove support for Python 2","startedAt":"2021-12-13T00:30:25Z","lastAt":"2021-12-13T00:30:25Z","messageCount":1,"participants":["Tzadik Vanderhoof"],"isPatch":true,"patchVersion":2,"patchTotal":3},"messages":[{"id":"443939","messageId":"CAKu1iLXJQxcjrk1Wny6ccMPPxKqJBpfxzTp+i+EZ90NvFkc-Yg@mail.gmail.com","threadId":"57071","inReplyTo":null,"subject":"[PATCH v2 1/3] git-p4: remove support for Python 2","fromName":"Tzadik Vanderhoof","fromEmail":"tzadik.vanderhoof@gmail.com","sentAt":"2021-12-13T00:30:10Z","receivedAt":"2021-12-13T00:30:25Z","isPatch":true,"sender":{"key":"tzadik.vanderhoof@gmail.com","avatar":null},"body":"> On Sun, Dec 12, 2021, 5:39 PM Ævar Arnfjörð Bjarmason <avarab@gmail.com> wrote:\n>\n> This summary makes sense, i.e. if the original SCM doesn't have a\n> declared or consistent encoding then having no \"encoding\" header etc. in\n> git likewise makes sense, and we should be trying to handle it in our\n> output layer.\n>\n> [Snipped from above]:\n>\n> > It's not clear to me how \"attempt to detect the encoding somehow\" would\n> > work.  The first option therefore seems like the best choice.\n>\n> This really isn't possible to do in the general case, but you can get\n> pretty far with heuristics.\n>\n> I already submitted a patch several months ago to introduce a \"p4.fallbackEncoding\" option. It got merged to at least the lowest branch, but I think it died at that point.\n\nI did considerable research into the possible options at the time, and\nI'm pretty sure the best approach would be:\n\nAdd an optional setting for the user to set the encoding.\n\nWhen decoding, first try UTF-8. If that succeeds, then it's almost\ncertain that the encoding really is UTF-8. The nature of UTF-8 is that\nnon-UTF-8 text almost never just happens to be valid when decoded as\nUTF-8.\n\nIf that fails, use the new setting if present.\n\nThis is what my patch does.\n\nI think it would be better to go beyond that, and if it fails UTF-8,\nand the new setting was not specified, then use some well- accepted\nheuristic library to detect the encoding.\n\nFrankly anything would be better than the current behavior, which is\nto completely crash on the first non UTF-8 character encountered (at\nleast with Python 3).\n"}]}