{"thread":{"id":"52262","subject":"[PATCH 0/1] git-p4.py: Cast byte strings to unicode strings in python3","startedAt":"2019-11-13T21:07:42Z","lastAt":"2019-12-13T20:41:04Z","messageCount":77,"participants":["Ben Keene via GitGitGadget","Junio C Hamano","Luke Diamand","Denton Liu","Ben Keene","Jeff King via GitGitGadget","Jeff King","Yang Zhao"],"isPatch":true,"patchVersion":1,"patchTotal":1},"messages":[{"id":"386140","messageId":"pull.463.git.1573679258.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":null,"subject":"[PATCH 0/1] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-11-13T21:07:37Z","receivedAt":"2019-11-13T21:07:42Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"commit: git-p4.py: Cast byte strings to unicode strings in python3\n\nI tried to run git-p4 under python3 and it failed with an error that it\ncould not connect to the P4 server. This is caused by the return values from\nthe process.popen returning byte strings and the code is failing when it is\ncomparing these with literal strings which are Unicode in Python 3.\n\nTo support this, I added a new function ustring() in the code that\ndetermines if python is natively supporting Unicode (Python 3) or not\n(Python 2). \n\n * If the python version supports Unicode (Python 3), it will cast the text\n   (expected a byte string) to UTF-8. This allows the existing code to match\n   literal strings as expected.\n   \n   \n * If the python version does not natively support Unicode (Python 2) the\n   ustring() function does not change the byte string, maintaining current\n   behavior.\n   \n   \n\nThere are a few notable methods changed:\n\n * pipe functions have their output passed through the ustring() function:\n   \n    * read_pipe_full(c)\n    * p4_has_move_command()\n   \n   \n * p4CmdList has new conditional code to parse the dictionary marshaled from\n   the process call. Both the keys and values are converted to Unicode.\n   \n   \n * gitConfig passes the return value through ustring() so all calls to\n   gitConfig return unicode values.\n   \n   \n\nSigned-off-by: Ben Keene seraphire@gmail.com [seraphire@gmail.com]\n\nBen Keene (1):\n  Cast byte strings to unicode strings in python3\n\n git-p4.py | 26 ++++++++++++++++++++++++--\n 1 file changed, 24 insertions(+), 2 deletions(-)\n\n\nbase-commit: d9f6f3b6195a0ca35642561e530798ad1469bd41\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-463%2Fseraphire%2Fseraphire%2Fp4-python3-unicode-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-463/seraphire/seraphire/p4-python3-unicode-v1\nPull-Request: https://github.com/gitgitgadget/git/pull/463\n-- \ngitgitgadget\n"},{"id":"386141","messageId":"0bca930ff82623bbef172b4cb6c36ef8e5c46098.1573679258.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.git.1573679258.gitgitgadget@gmail.com","subject":"[PATCH 1/1] Cast byte strings to unicode strings in python3","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-11-13T21:07:38Z","receivedAt":"2019-11-13T21:07:43Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <bkeene@partswatch.com>\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 26 ++++++++++++++++++++++++--\n 1 file changed, 24 insertions(+), 2 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 60c73b6a37..6e8b3a26cd 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -36,12 +36,22 @@\n     unicode = str\n     bytes = bytes\n     basestring = (str,bytes)\n+    isunicode = True\n+    def ustring(text):\n+        \"\"\"Returns the byte string as a unicode string\"\"\"\n+        if text == '' or text == b'':\n+            return ''\n+        return unicode(text, \"utf-8\")\n else:\n     # 'unicode' exists, must be Python 2\n     str = str\n     unicode = unicode\n     bytes = str\n     basestring = basestring\n+    isunicode = False\n+    def ustring(text):\n+        \"\"\"Returns the byte string unchanged\"\"\"\n+        return text\n \n try:\n     from subprocess import CalledProcessError\n@@ -196,6 +206,8 @@ def read_pipe_full(c):\n     expand = isinstance(c,basestring)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=expand)\n     (out, err) = p.communicate()\n+    out = ustring(out)\n+    err = ustring(err)\n     return (p.returncode, out, err)\n \n def read_pipe(c, ignore_error=False):\n@@ -263,6 +275,7 @@ def p4_has_move_command():\n     cmd = p4_build_cmd([\"move\", \"-k\", \"@from\", \"@to\"])\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     (out, err) = p.communicate()\n+    err = ustring(err)\n     # return code will be 1 in either case\n     if err.find(\"Invalid option\") >= 0:\n         return False\n@@ -646,10 +659,18 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n             if skip_info:\n                 if 'code' in entry and entry['code'] == 'info':\n                     continue\n+                if b'code' in entry and entry[b'code'] == b'info':\n+                    continue\n             if cb is not None:\n                 cb(entry)\n             else:\n-                result.append(entry)\n+                if isunicode:\n+                    out = {}\n+                    for key, value in entry.items():\n+                        out[ustring(key)] = ustring(value)\n+                    result.append(out)\n+                else:\n+                    result.append(entry)\n     except EOFError:\n         pass\n     exitCode = p4.wait()\n@@ -792,7 +813,7 @@ def gitConfig(key, typeSpecifier=None):\n         cmd += [ key ]\n         s = read_pipe(cmd, ignore_error=True)\n         _gitConfig[key] = s.strip()\n-    return _gitConfig[key]\n+    return ustring(_gitConfig[key])\n \n def gitConfigBool(key):\n     \"\"\"Return a bool, using git config --bool.  It is True only if the\n@@ -860,6 +881,7 @@ def branch_exists(branch):\n     cmd = [ \"git\", \"rev-parse\", \"--symbolic\", \"--verify\", branch ]\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     out, _ = p.communicate()\n+    out = ustring(out)\n     if p.returncode:\n         return False\n     # expect exactly one line of output: the branch name\n-- \ngitgitgadget\n"},{"id":"386154","messageId":"xmqqa78z9ou7.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"pull.463.git.1573679258.gitgitgadget@gmail.com","subject":"Re: [PATCH 0/1] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-11-14T02:25:04Z","receivedAt":"2019-11-14T02:25:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> commit: git-p4.py: Cast byte strings to unicode strings in python3\n\nLuke, this patch [*1*] came in my way, but I am hardly an expert on\nPy2to3 and know nothing about P4.  Could you take a look at them,\nplease?\n\nThanks.\n\n\n[References]\n\n<0bca930ff82623bbef172b4cb6c36ef8e5c46098.1573679258.git.gitgitgadget@gmail.com>\n"},{"id":"386176","messageId":"CAE5ih7_KXJ-4r=hOWdhhWdz9MdZHLrVisZTctGjMYYQQh6Om3Q@mail.gmail.com","threadId":"52262","inReplyTo":"xmqqa78z9ou7.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH 0/1] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Luke Diamand","fromEmail":"luke@diamand.org","sentAt":"2019-11-14T09:46:41Z","receivedAt":"2019-11-14T09:46:33Z","isPatch":true,"sender":{"key":"luke@diamand.org","avatar":"https://avatars.githubusercontent.com/u/5330967?v=4"},"body":"On Thu, 14 Nov 2019 at 02:25, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> \"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n> > commit: git-p4.py: Cast byte strings to unicode strings in python3\n>\n> Luke, this patch [*1*] came in my way, but I am hardly an expert on\n> Py2to3 and know nothing about P4.  Could you take a look at them,\n> please?\n>\n> Thanks.\n>\n>\n> [References]\n>\n> <0bca930ff82623bbef172b4cb6c36ef8e5c46098.1573679258.git.gitgitgadget@gmail.com>\n\nI just quickly tried it, and with git-p4 switched to using python3,\nthe unit tests fail.\n\n$ make -C t T=t98*\n\nBut it looks like a reasonable approach, and with the demise of\nPython2 fast approaching it would be good to get this fully working!\n\nLuke\n"},{"id":"386326","messageId":"pull.463.v2.git.1573828756.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.git.1573679258.gitgitgadget@gmail.com","subject":"[PATCH v2 0/3] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-11-15T14:39:13Z","receivedAt":"2019-11-15T14:39:21Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"git-p4.py: Cast byte strings to unicode strings in python3\n\nI tried to run git-p4 under python3 and it failed with an error that it\ncould not connect to the P4 server. This PR covers updating the git-p4.py\npython script to work with unicode strings in python3.\n\nChanges since v1: Commit: (0435d0e) 2019-11-14\n\nThe problem was caused by the ustring() function being called on a string\nthat had already been cast as a unicode string. This second call to\nustring() would fail with an error of \"decoding str is not supported\"\n\nThe following changes were made to fix this:\n\nThe call to ustring() in the gitConfig() function is actually unnecessary\nbecause the read_pipe() function returns unicode strings so the call has\nbeen removed.\n\nThe ustring() function was given a new conditional test to see if the value\nis already a unicode value. If it is, the value will be returned without any\ncasting.\n\nThese two changes should fix the immediate fail. However, I do not have an\nenvironment that I can run the test suite against so I don't know if another\nerror will be uncovered yet. I'm still working on it.\n\nv1: (Initial Commit)\n\nThis is caused by the return values from the process.popen returning byte\nstrings and the code is failing when it is comparing these with literal\nstrings which are Unicode in Python 3.\n\nTo support this, I added a new function ustring() in the code that\ndetermines if python is natively supporting Unicode (Python 3) or not\n(Python 2). \n\n * If the python version supports Unicode (Python 3), it will cast the text\n   (expected a byte string) to UTF-8. This allows the existing code to match\n   literal strings as expected.\n   \n   \n * If the python version does not natively support Unicode (Python 2) the\n   ustring() function does not change the byte string, maintaining current\n   behavior.\n   \n   \n\nThere are a few notable methods changed:\n\n * pipe functions have their output passed through the ustring() function:\n   \n    * read_pipe_full(c)\n    * p4_has_move_command()\n   \n   \n * p4CmdList has new conditional code to parse the dictionary marshaled from\n   the process call. Both the keys and values are converted to Unicode.\n   \n   \n * gitConfig passes the return value through ustring() so all calls to\n   gitConfig return unicode values.\n   \n   \n\nSigned-off-by: Ben Keene seraphire@gmail.com [seraphire@gmail.com]\n\nBen Keene (3):\n  Cast byte strings to unicode strings in python3\n  FIX: cast as unicode fails when a value is already unicode\n  FIX: wrap return for read_pipe_lines in ustring() and wrap GitLFS read\n    of the pointer file in ustring()\n\n git-p4.py | 38 ++++++++++++++++++++++++++++++++++++--\n 1 file changed, 36 insertions(+), 2 deletions(-)\n\n\nbase-commit: d9f6f3b6195a0ca35642561e530798ad1469bd41\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-463%2Fseraphire%2Fseraphire%2Fp4-python3-unicode-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-463/seraphire/seraphire/p4-python3-unicode-v2\nPull-Request: https://github.com/gitgitgadget/git/pull/463\n\nRange-diff vs v1:\n\n 1:  0bca930ff8 = 1:  0bca930ff8 Cast byte strings to unicode strings in python3\n -:  ---------- > 2:  0435d0e2cb FIX: cast as unicode fails when a value is already unicode\n -:  ---------- > 3:  2288690b94 FIX: wrap return for read_pipe_lines in ustring() and wrap GitLFS read of the pointer file in ustring()\n\n-- \ngitgitgadget\n"},{"id":"386327","messageId":"0bca930ff82623bbef172b4cb6c36ef8e5c46098.1573828756.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v2.git.1573828756.gitgitgadget@gmail.com","subject":"[PATCH v2 1/3] Cast byte strings to unicode strings in python3","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-11-15T14:39:14Z","receivedAt":"2019-11-15T14:39:22Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <bkeene@partswatch.com>\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 26 ++++++++++++++++++++++++--\n 1 file changed, 24 insertions(+), 2 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 60c73b6a37..6e8b3a26cd 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -36,12 +36,22 @@\n     unicode = str\n     bytes = bytes\n     basestring = (str,bytes)\n+    isunicode = True\n+    def ustring(text):\n+        \"\"\"Returns the byte string as a unicode string\"\"\"\n+        if text == '' or text == b'':\n+            return ''\n+        return unicode(text, \"utf-8\")\n else:\n     # 'unicode' exists, must be Python 2\n     str = str\n     unicode = unicode\n     bytes = str\n     basestring = basestring\n+    isunicode = False\n+    def ustring(text):\n+        \"\"\"Returns the byte string unchanged\"\"\"\n+        return text\n \n try:\n     from subprocess import CalledProcessError\n@@ -196,6 +206,8 @@ def read_pipe_full(c):\n     expand = isinstance(c,basestring)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=expand)\n     (out, err) = p.communicate()\n+    out = ustring(out)\n+    err = ustring(err)\n     return (p.returncode, out, err)\n \n def read_pipe(c, ignore_error=False):\n@@ -263,6 +275,7 @@ def p4_has_move_command():\n     cmd = p4_build_cmd([\"move\", \"-k\", \"@from\", \"@to\"])\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     (out, err) = p.communicate()\n+    err = ustring(err)\n     # return code will be 1 in either case\n     if err.find(\"Invalid option\") >= 0:\n         return False\n@@ -646,10 +659,18 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n             if skip_info:\n                 if 'code' in entry and entry['code'] == 'info':\n                     continue\n+                if b'code' in entry and entry[b'code'] == b'info':\n+                    continue\n             if cb is not None:\n                 cb(entry)\n             else:\n-                result.append(entry)\n+                if isunicode:\n+                    out = {}\n+                    for key, value in entry.items():\n+                        out[ustring(key)] = ustring(value)\n+                    result.append(out)\n+                else:\n+                    result.append(entry)\n     except EOFError:\n         pass\n     exitCode = p4.wait()\n@@ -792,7 +813,7 @@ def gitConfig(key, typeSpecifier=None):\n         cmd += [ key ]\n         s = read_pipe(cmd, ignore_error=True)\n         _gitConfig[key] = s.strip()\n-    return _gitConfig[key]\n+    return ustring(_gitConfig[key])\n \n def gitConfigBool(key):\n     \"\"\"Return a bool, using git config --bool.  It is True only if the\n@@ -860,6 +881,7 @@ def branch_exists(branch):\n     cmd = [ \"git\", \"rev-parse\", \"--symbolic\", \"--verify\", branch ]\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     out, _ = p.communicate()\n+    out = ustring(out)\n     if p.returncode:\n         return False\n     # expect exactly one line of output: the branch name\n-- \ngitgitgadget\n\n"},{"id":"386328","messageId":"0435d0e2cb4c62ede5f64b396af262d6bf8b5079.1573828756.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v2.git.1573828756.gitgitgadget@gmail.com","subject":"[PATCH v2 2/3] FIX: cast as unicode fails when a value is already unicode","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-11-15T14:39:15Z","receivedAt":"2019-11-15T14:39:23Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 4 +++-\n 1 file changed, 3 insertions(+), 1 deletion(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 6e8b3a26cd..b088095b15 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -39,6 +39,8 @@\n     isunicode = True\n     def ustring(text):\n         \"\"\"Returns the byte string as a unicode string\"\"\"\n+        if isinstance(text, unicode):\n+            return text\n         if text == '' or text == b'':\n             return ''\n         return unicode(text, \"utf-8\")\n@@ -813,7 +815,7 @@ def gitConfig(key, typeSpecifier=None):\n         cmd += [ key ]\n         s = read_pipe(cmd, ignore_error=True)\n         _gitConfig[key] = s.strip()\n-    return ustring(_gitConfig[key])\n+    return _gitConfig[key]\n \n def gitConfigBool(key):\n     \"\"\"Return a bool, using git config --bool.  It is True only if the\n-- \ngitgitgadget\n\n"},{"id":"386329","messageId":"2288690b94d82a629a3a94e25fa75d24a1c24000.1573828756.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v2.git.1573828756.gitgitgadget@gmail.com","subject":"[PATCH v2 3/3] FIX: wrap return for read_pipe_lines in ustring() and wrap GitLFS read of the pointer file in ustring()","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-11-15T14:39:16Z","receivedAt":"2019-11-15T14:39:23Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 12 +++++++++++-\n 1 file changed, 11 insertions(+), 1 deletion(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex b088095b15..83f59ddca5 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -180,6 +180,11 @@ def die(msg):\n         sys.exit(1)\n \n def write_pipe(c, stdin):\n+    \"\"\"Writes stdin to the command's stdin\n+    Returns the number of bytes written.\n+\n+    Be aware - the byte count may change between \n+    Python2 and Python3\"\"\"\n     if verbose:\n         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n \n@@ -249,6 +254,11 @@ def read_pipe_lines(c):\n     val = pipe.readlines()\n     if pipe.close() or p.wait():\n         die('Command failed: %s' % str(c))\n+    # Unicode conversion from str\n+    # Iterate and fix in-place to avoid a second list in memory.\n+    if isunicode:\n+        for i in range(len(val)):\n+            val[i] = ustring(val[i])\n \n     return val\n \n@@ -1268,7 +1278,7 @@ def generatePointer(self, contentFile):\n             ['git', 'lfs', 'pointer', '--file=' + contentFile],\n             stdout=subprocess.PIPE\n         )\n-        pointerFile = pointerProcess.stdout.read()\n+        pointerFile = ustring(pointerProcess.stdout.read())\n         if pointerProcess.wait():\n             os.remove(contentFile)\n             die('git-lfs pointer command failed. Did you install the extension?')\n-- \ngitgitgadget\n"},{"id":"387408","messageId":"pull.463.v3.git.1575313336.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v2.git.1573828756.gitgitgadget@gmail.com","subject":"[PATCH v3 0/1] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-02T19:02:15Z","receivedAt":"2019-12-02T19:02:23Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"Issue: The current git-p4.py script does not work with python3.\n\nI have attempted to use the P4 integration built into GIT and I was unable\nto get the program to run because I have Python 3.8 installed on my\ncomputer. I was able to get the program to run when I downgraded my python\nto version 2.7. However, python 2 is reaching its end of life.\n\nSubmission: I am submitting a patch for the git-p4.py script that partially\nsupports python 3.8. This code was able to pass the basic tests (t9800) when\nrun against Python3. This provides basic functionality. \n\nIn an attempt to pass the t9822 P4 path-encoding test, a new parameter for\ngit P4 Clone was introduced. \n\n--encoding Format-identifier\n\nThis will create the GIT repository following the current functionality;\nhowever, before importing the files from P4, it will set the\ngit-p4.pathEncoding option so any files or paths that are encoded with\nnon-ASCII/non-UTF-8 formats will import correctly.\n\nTechnical details: The script was updated by futurize (\nhttps://python-future.org/futurize.html) to support Py2/Py3 syntax. The few\nreferences to classes in future were reworked so that future would not be\nrequired. The existing code test for Unicode support was extended to\nnormalize the classes “unicode” and “bytes” to across platforms:\n\n * ‘unicode’ is an alias for ‘str’ in Py3 and is the unicode class in Py2.\n * ‘bytes’ is bytes in Py3 and an alias for ‘str’ in Py2.\n\nNew coercion methods were written for both Python2 and Python3:\n\n * as_string(text) – In Python3, this encodes a bytes object as a UTF-8\n   encoded Unicode string. \n * as_bytes(text) – In Python3, this decodes a Unicode string to an array of\n   bytes.\n\nIn Python2, these functions do not change the data since a ‘str’ object\nfunction in both roles as strings and byte arrays. This reduces the\npotential impact on backward compatibility with Python 2.\n\n * to_unicode(text) – ensures that the supplied data is encoded as a UTF-8\n   string. This function will encode data in both Python2 and Python3. * \n      path_as_string(path) – This function is an extension function that\n      honors the option “git-p4.pathEncoding” to convert a set of bytes or\n      characters to UTF-8. If the str/bytes cannot decode as ASCII, it will\n      use the encodeWithUTF8() method to convert the custom encoded bytes to\n      Unicode in UTF-8.\n   \n   \n\nGenerally speaking, information in the script is converted to Unicode as\nearly as possible and converted back to a byte array just before passing to\nexternal programs or files. The exception to this rule is P4 Repository file\npaths.\n\nPaths are not converted but left as “bytes” so the original file path\nencoding can be preserved. This formatting is required for commands that\ninteract with the P4 file path. When the file path is used by GIT, it is\nconverted with encodeWithUTF8().\n\nSigned-off-by: Ben Keene seraphire@gmail.com [seraphire@gmail.com]\n\nBen Keene (1):\n  Python3 support for t9800 tests. Basic P4/Python3 support\n\n git-p4.py | 825 +++++++++++++++++++++++++++++++++++++++++-------------\n 1 file changed, 628 insertions(+), 197 deletions(-)\n\n\nbase-commit: d9f6f3b6195a0ca35642561e530798ad1469bd41\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-463%2Fseraphire%2Fseraphire%2Fp4-python3-unicode-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-463/seraphire/seraphire/p4-python3-unicode-v3\nPull-Request: https://github.com/gitgitgadget/git/pull/463\n\nRange-diff vs v2:\n\n 1:  0bca930ff8 < -:  ---------- Cast byte strings to unicode strings in python3\n 2:  0435d0e2cb < -:  ---------- FIX: cast as unicode fails when a value is already unicode\n 3:  2288690b94 < -:  ---------- FIX: wrap return for read_pipe_lines in ustring() and wrap GitLFS read of the pointer file in ustring()\n -:  ---------- > 1:  02b3843e9f Python3 support for t9800 tests. Basic P4/Python3 support\n\n-- \ngitgitgadget\n"},{"id":"387409","messageId":"02b3843e9f21105a945335d0b1d78251ddcc8cee.1575313336.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v3.git.1575313336.gitgitgadget@gmail.com","subject":"[PATCH v3 1/1] Python3 support for t9800 tests. Basic P4/Python3 support","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-02T19:02:16Z","receivedAt":"2019-12-02T19:02:30Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 825 +++++++++++++++++++++++++++++++++++++++++-------------\n 1 file changed, 628 insertions(+), 197 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 60c73b6a37..6f82184fe5 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -26,22 +26,87 @@\n import zlib\n import ctypes\n import errno\n+import os.path\n+import codecs\n+import io\n \n # support basestring in python3\n try:\n     unicode = unicode\n except NameError:\n     # 'unicode' is undefined, must be Python 3\n-    str = str\n+    #\n+    # For Python3 which is natively unicode, we will use \n+    # unicode for internal information but all P4 Data\n+    # will remain in bytes\n+    isunicode = True\n     unicode = str\n     bytes = bytes\n-    basestring = (str,bytes)\n+\n+    def as_string(text):\n+        \"\"\"Return a byte array as a unicode string\"\"\"\n+        if text == None:\n+            return None\n+        if isinstance(text, bytes):\n+            return unicode(text, \"utf-8\")\n+        else:\n+            return text\n+\n+    def as_bytes(text):\n+        \"\"\"Return a Unicode string as a byte array\"\"\"\n+        if text == None:\n+            return None\n+        if isinstance(text, bytes):\n+            return text\n+        else:\n+            return bytes(text, \"utf-8\")\n+\n+    def to_unicode(text):\n+        \"\"\"Return a byte array as a unicode string\"\"\"\n+        return as_string(text)    \n+\n+    def path_as_string(path):\n+        \"\"\" Converts a path to the UTF8 encoded string \"\"\"\n+        if isinstance(path, unicode):\n+            return path\n+        return encodeWithUTF8(path).decode('utf-8')\n+    \n else:\n     # 'unicode' exists, must be Python 2\n-    str = str\n+    #\n+    # We will treat the data as:\n+    #   str   -> str\n+    #   bytes -> str\n+    # So for Python2 these functions are no-ops\n+    # and will leave the data in the ambiguious\n+    # string/bytes state\n+    isunicode = False\n     unicode = unicode\n     bytes = str\n-    basestring = basestring\n+\n+    def as_string(text):\n+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n+        return text\n+\n+    def as_bytes(text):\n+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n+        return text\n+\n+    def to_unicode(text):\n+        \"\"\"Return a string as a unicode string\"\"\"\n+        return text.decode('utf-8')\n+    \n+    def path_as_string(path):\n+        \"\"\" Converts a path to the UTF8 encoded bytes \"\"\"\n+        return encodeWithUTF8(path)\n+\n+\n+ \n+# Check for raw_input support\n+try:\n+    raw_input\n+except NameError:\n+    raw_input = input\n \n try:\n     from subprocess import CalledProcessError\n@@ -75,7 +140,11 @@ def p4_build_cmd(cmd):\n     location. It means that hooking into the environment, or other configuration\n     can be done more easily.\n     \"\"\"\n-    real_cmd = [\"p4\"]\n+    # Look for the P4 binary\n+    if (platform.system() == \"Windows\"):\n+        real_cmd = [\"p4.exe\"]    \n+    else:\n+        real_cmd = [\"p4\"]\n \n     user = gitConfig(\"git-p4.user\")\n     if len(user) > 0:\n@@ -105,7 +174,7 @@ def p4_build_cmd(cmd):\n         # Provide a way to not pass this option by setting git-p4.retries to 0\n         real_cmd += [\"-r\", str(retries)]\n \n-    if isinstance(cmd,basestring):\n+    if not isinstance(cmd, list):\n         real_cmd = ' '.join(real_cmd) + ' ' + cmd\n     else:\n         real_cmd += cmd\n@@ -168,10 +237,11 @@ def die(msg):\n         sys.exit(1)\n \n def write_pipe(c, stdin):\n+    \"\"\"Executes the command 'c', passing 'stdin' on the standard input\"\"\"\n     if verbose:\n         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n \n-    expand = isinstance(c,basestring)\n+    expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdin=subprocess.PIPE, shell=expand)\n     pipe = p.stdin\n     val = pipe.write(stdin)\n@@ -179,11 +249,11 @@ def write_pipe(c, stdin):\n     if p.wait():\n         die('Command failed: %s' % str(c))\n \n-    return val\n \n def p4_write_pipe(c, stdin):\n+    \"\"\" Runs a P4 command 'c', passing 'stdin' data to P4\"\"\"\n     real_cmd = p4_build_cmd(c)\n-    return write_pipe(real_cmd, stdin)\n+    write_pipe(real_cmd, stdin)\n \n def read_pipe_full(c):\n     \"\"\" Read output from  command. Returns a tuple\n@@ -193,9 +263,11 @@ def read_pipe_full(c):\n     if verbose:\n         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n \n-    expand = isinstance(c,basestring)\n+    expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=expand)\n     (out, err) = p.communicate()\n+    out = as_string(out)\n+    err = as_string(err)\n     return (p.returncode, out, err)\n \n def read_pipe(c, ignore_error=False):\n@@ -222,19 +294,31 @@ def read_pipe_text(c):\n         return out.rstrip()\n \n def p4_read_pipe(c, ignore_error=False):\n+    \"\"\" Read output from the P4 command 'c'. Returns the output text on\n+        success. On failure, terminates execution, unless\n+        ignore_error is True, when it returns an empty string.\n+    \"\"\"\n     real_cmd = p4_build_cmd(c)\n     return read_pipe(real_cmd, ignore_error)\n \n def read_pipe_lines(c):\n+    \"\"\" Returns a list of text from executing the command 'c'.\n+        The program will die if the command fails to execute.\n+    \"\"\"\n     if verbose:\n         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n \n-    expand = isinstance(c, basestring)\n+    expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, shell=expand)\n     pipe = p.stdout\n     val = pipe.readlines()\n     if pipe.close() or p.wait():\n         die('Command failed: %s' % str(c))\n+    # Unicode conversion from byte-string\n+    # Iterate and fix in-place to avoid a second list in memory.\n+    if isunicode:\n+        for i in range(len(val)):\n+            val[i] = as_string(val[i])\n \n     return val\n \n@@ -263,6 +347,8 @@ def p4_has_move_command():\n     cmd = p4_build_cmd([\"move\", \"-k\", \"@from\", \"@to\"])\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     (out, err) = p.communicate()\n+    out=as_string(out)\n+    err=as_string(err)\n     # return code will be 1 in either case\n     if err.find(\"Invalid option\") >= 0:\n         return False\n@@ -272,7 +358,7 @@ def p4_has_move_command():\n     return True\n \n def system(cmd, ignore_error=False):\n-    expand = isinstance(cmd,basestring)\n+    expand = not isinstance(cmd, list)\n     if verbose:\n         sys.stderr.write(\"executing %s\\n\" % str(cmd))\n     retcode = subprocess.call(cmd, shell=expand)\n@@ -282,9 +368,10 @@ def system(cmd, ignore_error=False):\n     return retcode\n \n def p4_system(cmd):\n-    \"\"\"Specifically invoke p4 as the system command. \"\"\"\n+    \"\"\" Specifically invoke p4 as the system command. \n+    \"\"\"\n     real_cmd = p4_build_cmd(cmd)\n-    expand = isinstance(real_cmd, basestring)\n+    expand = not isinstance(real_cmd, list)\n     retcode = subprocess.call(real_cmd, shell=expand)\n     if retcode:\n         raise CalledProcessError(retcode, real_cmd)\n@@ -390,16 +477,20 @@ def p4_last_change():\n     return int(results[0]['change'])\n \n def p4_describe(change, shelved=False):\n-    \"\"\"Make sure it returns a valid result by checking for\n-       the presence of field \"time\".  Return a dict of the\n-       results.\"\"\"\n+    \"\"\" Returns information about the requested P4 change list.\n+\n+        Data returns is not string encoded (returned as bytes)\n+    \"\"\"\n+    # Make sure it returns a valid result by checking for\n+    #   the presence of field \"time\".  Return a dict of the\n+    #   results.\n \n     cmd = [\"describe\", \"-s\"]\n     if shelved:\n         cmd += [\"-S\"]\n     cmd += [str(change)]\n \n-    ds = p4CmdList(cmd, skip_info=True)\n+    ds = p4CmdList(cmd, skip_info=True, encode_data=False)\n     if len(ds) != 1:\n         die(\"p4 describe -s %d did not return 1 result: %s\" % (change, str(ds)))\n \n@@ -409,21 +500,31 @@ def p4_describe(change, shelved=False):\n         die(\"p4 describe -s %d exited with %d: %s\" % (change, d[\"p4ExitCode\"],\n                                                       str(d)))\n     if \"code\" in d:\n-        if d[\"code\"] == \"error\":\n+        if d[\"code\"] == b\"error\":\n             die(\"p4 describe -s %d returned error code: %s\" % (change, str(d)))\n \n     if \"time\" not in d:\n         die(\"p4 describe -s %d returned no \\\"time\\\": %s\" % (change, str(d)))\n \n+    # Convert depotFile(X) to be UTF-8 encoded, as this is what GIT\n+    # requires. This will also allow us to encode the rest of the text\n+    # at the same time to simplify textual processing later.\n+    keys=d.keys()\n+    for key in keys:\n+        if key.startswith('depotFile'):\n+            d[key]=d[key] #DepotPath(d[key])\n+        elif key == 'path':\n+            d[key]=d[key] #DepotPath(d[key])\n+        else:\n+            d[key] = as_string(d[key])\n+\n     return d\n \n-#\n-# Canonicalize the p4 type and return a tuple of the\n-# base type, plus any modifiers.  See \"p4 help filetypes\"\n-# for a list and explanation.\n-#\n def split_p4_type(p4type):\n-\n+    \"\"\" Canonicalize the p4 type and return a tuple of the\n+        base type, plus any modifiers.  See \"p4 help filetypes\"\n+        for a list and explanation.\n+    \"\"\"\n     p4_filetypes_historical = {\n         \"ctempobj\": \"binary+Sw\",\n         \"ctext\": \"text+C\",\n@@ -452,18 +553,16 @@ def split_p4_type(p4type):\n         mods = s[1]\n     return (base, mods)\n \n-#\n-# return the raw p4 type of a file (text, text+ko, etc)\n-#\n def p4_type(f):\n+    \"\"\" return the raw p4 type of a file (text, text+ko, etc)\n+    \"\"\"\n     results = p4CmdList([\"fstat\", \"-T\", \"headType\", wildcard_encode(f)])\n     return results[0]['headType']\n \n-#\n-# Given a type base and modifier, return a regexp matching\n-# the keywords that can be expanded in the file\n-#\n def p4_keywords_regexp_for_type(base, type_mods):\n+    \"\"\" Given a type base and modifier, return a regexp matching\n+        the keywords that can be expanded in the file\n+    \"\"\"\n     if base in (\"text\", \"unicode\", \"binary\"):\n         kwords = None\n         if \"ko\" in type_mods:\n@@ -482,12 +581,11 @@ def p4_keywords_regexp_for_type(base, type_mods):\n     else:\n         return None\n \n-#\n-# Given a file, return a regexp matching the possible\n-# RCS keywords that will be expanded, or None for files\n-# with kw expansion turned off.\n-#\n def p4_keywords_regexp_for_file(file):\n+    \"\"\" Given a file, return a regexp matching the possible\n+        RCS keywords that will be expanded, or None for files\n+        with kw expansion turned off.\n+    \"\"\"\n     if not os.path.exists(file):\n         return None\n     else:\n@@ -522,7 +620,7 @@ def getP4OpenedType(file):\n # Return the set of all p4 labels\n def getP4Labels(depotPaths):\n     labels = set()\n-    if isinstance(depotPaths,basestring):\n+    if not isinstance(depotPaths, list):\n         depotPaths = [depotPaths]\n \n     for l in p4CmdList([\"labels\"] + [\"%s...\" % p for p in depotPaths]):\n@@ -531,8 +629,8 @@ def getP4Labels(depotPaths):\n \n     return labels\n \n-# Return the set of all git tags\n def getGitTags():\n+    \"\"\"Return the set of all git tags\"\"\"\n     gitTags = set()\n     for line in read_pipe_lines([\"git\", \"tag\"]):\n         tag = line.strip()\n@@ -565,7 +663,7 @@ def parseDiffTreeEntry(entry):\n \n     If the pattern is not matched, None is returned.\"\"\"\n \n-    match = diffTreePattern().next().match(entry)\n+    match = next(diffTreePattern()).match(entry)\n     if match:\n         return {\n             'src_mode': match.group(1),\n@@ -584,6 +682,38 @@ def isModeExec(mode):\n     # otherwise False.\n     return mode[-3:] == \"755\"\n \n+def encodeWithUTF8(path, verbose = False):\n+    \"\"\" Ensure that the path is encoded as a UTF-8 string\n+\n+        Returns bytes(P3)/str(P2)\n+    \"\"\"\n+   \n+    if isunicode:\n+        try:\n+            if isinstance(path, unicode):\n+                # It is already unicode, cast it as a bytes\n+                # that is encoded as utf-8.\n+                return path.encode('utf-8', 'strict')\n+            path.decode('ascii', 'strict')\n+        except:\n+            encoding = 'utf8'\n+            if gitConfig('git-p4.pathEncoding'):\n+                encoding = gitConfig('git-p4.pathEncoding')\n+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n+            if verbose:\n+                print('\\nNOTE:Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, to_unicode(path)))\n+    else:    \n+        try:\n+            path.decode('ascii')\n+        except:\n+            encoding = 'utf8'\n+            if gitConfig('git-p4.pathEncoding'):\n+                encoding = gitConfig('git-p4.pathEncoding')\n+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n+            if verbose:\n+                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n+    return path\n+\n class P4Exception(Exception):\n     \"\"\" Base class for exceptions from the p4 client \"\"\"\n     def __init__(self, exit_code):\n@@ -607,9 +737,25 @@ def isModeExecChanged(src_mode, dst_mode):\n     return isModeExec(src_mode) != isModeExec(dst_mode)\n \n def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n-        errors_as_exceptions=False):\n+        errors_as_exceptions=False, encode_data=True):\n+    \"\"\" Executes a P4 command:  'cmd' optionally passing 'stdin' to the command's\n+        standard input via a temporary file with 'stdin_mode' mode.\n+\n+        Output from the command is optionally passed to the callback function 'cb'.\n+        If 'cb' is None, the response from the command is parsed into a list\n+        of resulting dictionaries. (For each block read from the process pipe.)\n+\n+        If 'skip_info' is true, information in a block read that has a code type of\n+        'info' will be skipped.\n \n-    if isinstance(cmd,basestring):\n+        If 'errors_as_exceptions' is set to true (the default is false) the error\n+        code returned from the execution will generate an exception.\n+\n+        If 'encode_data' is set to true (the default) the data that is returned \n+        by this function will be passed through the \"as_string\" function.\n+    \"\"\"\n+\n+    if not isinstance(cmd, list):\n         cmd = \"-G \" + cmd\n         expand = True\n     else:\n@@ -626,11 +772,11 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n     stdin_file = None\n     if stdin is not None:\n         stdin_file = tempfile.TemporaryFile(prefix='p4-stdin', mode=stdin_mode)\n-        if isinstance(stdin,basestring):\n+        if not isinstance(stdin, list):\n             stdin_file.write(stdin)\n         else:\n             for i in stdin:\n-                stdin_file.write(i + '\\n')\n+                stdin_file.write(as_bytes(i) + b'\\n')\n         stdin_file.flush()\n         stdin_file.seek(0)\n \n@@ -644,12 +790,15 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n         while True:\n             entry = marshal.load(p4.stdout)\n             if skip_info:\n-                if 'code' in entry and entry['code'] == 'info':\n+                if b'code' in entry and entry[b'code'] == b'info':\n                     continue\n             if cb is not None:\n                 cb(entry)\n             else:\n-                result.append(entry)\n+                out = {}\n+                for key, value in entry.items():\n+                    out[as_string(key)] = (as_string(value) if encode_data else value)\n+                result.append(out)\n     except EOFError:\n         pass\n     exitCode = p4.wait()\n@@ -677,6 +826,7 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n     return result\n \n def p4Cmd(cmd):\n+    \"\"\" Executes a P4 command an returns the results in a dictionary\"\"\"\n     list = p4CmdList(cmd)\n     result = {}\n     for entry in list:\n@@ -772,6 +922,7 @@ def extractSettingsGitLog(log):\n     return values\n \n def gitBranchExists(branch):\n+    \"\"\"Checks to see if a given branch exists in the git repo\"\"\"\n     proc = subprocess.Popen([\"git\", \"rev-parse\", branch],\n                             stderr=subprocess.PIPE, stdout=subprocess.PIPE);\n     return proc.wait() == 0;\n@@ -785,20 +936,22 @@ def gitDeleteRef(ref):\n _gitConfig = {}\n \n def gitConfig(key, typeSpecifier=None):\n+    \"\"\" Return a configuration setting from GIT\n+\t\"\"\"\n     if key not in _gitConfig:\n         cmd = [ \"git\", \"config\" ]\n         if typeSpecifier:\n             cmd += [ typeSpecifier ]\n         cmd += [ key ]\n         s = read_pipe(cmd, ignore_error=True)\n-        _gitConfig[key] = s.strip()\n+        _gitConfig[key] = as_string(s).strip()\n     return _gitConfig[key]\n \n def gitConfigBool(key):\n-    \"\"\"Return a bool, using git config --bool.  It is True only if the\n-       variable is set to true, and False if set to false or not present\n-       in the config.\"\"\"\n-\n+    \"\"\" Return a bool, using git config --bool.  It is True only if the\n+        variable is set to true, and False if set to false or not present\n+        in the config.\n+    \"\"\"\n     if key not in _gitConfig:\n         _gitConfig[key] = gitConfig(key, '--bool') == \"true\"\n     return _gitConfig[key]\n@@ -822,6 +975,11 @@ def gitConfigList(key):\n             _gitConfig[key] = []\n     return _gitConfig[key]\n \n+def gitConfigSet(key, value):\n+    \"\"\" Set the git configuration key 'key' to 'value' for this session\n+    \"\"\"\n+    _gitConfig[key] = value\n+\n def p4BranchesInGit(branchesAreInRemotes=True):\n     \"\"\"Find all the branches whose names start with \"p4/\", looking\n        in remotes or heads as specified by the argument.  Return\n@@ -860,6 +1018,7 @@ def branch_exists(branch):\n     cmd = [ \"git\", \"rev-parse\", \"--symbolic\", \"--verify\", branch ]\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     out, _ = p.communicate()\n+    out = as_string(out)\n     if p.returncode:\n         return False\n     # expect exactly one line of output: the branch name\n@@ -869,7 +1028,7 @@ def findUpstreamBranchPoint(head = \"HEAD\"):\n     branches = p4BranchesInGit()\n     # map from depot-path to branch name\n     branchByDepotPath = {}\n-    for branch in branches.keys():\n+    for branch in list(branches.keys()):\n         tip = branches[branch]\n         log = extractLogMessageFromGitCommit(tip)\n         settings = extractSettingsGitLog(log)\n@@ -940,7 +1099,8 @@ def createOrUpdateBranchesFromOrigin(localRefPrefix = \"refs/remotes/p4/\", silent\n             system(\"git update-ref %s %s\" % (remoteHead, originHead))\n \n def originP4BranchesExist():\n-        return gitBranchExists(\"origin\") or gitBranchExists(\"origin/p4\") or gitBranchExists(\"origin/p4/master\")\n+    \"\"\"Checks if origin/p4/master exists\"\"\"\n+    return gitBranchExists(\"origin\") or gitBranchExists(\"origin/p4\") or gitBranchExists(\"origin/p4/master\")\n \n \n def p4ParseNumericChangeRange(parts):\n@@ -1035,7 +1195,7 @@ def p4ChangesForPaths(depotPaths, changeRange, requestedBlockSize):\n     changes = sorted(changes)\n     return changes\n \n-def p4PathStartsWith(path, prefix):\n+def p4PathStartsWith(path, prefix, verbose = False):\n     # This method tries to remedy a potential mixed-case issue:\n     #\n     # If UserA adds  //depot/DirA/file1\n@@ -1043,9 +1203,22 @@ def p4PathStartsWith(path, prefix):\n     #\n     # we may or may not have a problem. If you have core.ignorecase=true,\n     # we treat DirA and dira as the same directory\n+    \n+    # Since we have to deal with mixed encodings for p4 file\n+    # paths, first perform a simple startswith check, this covers\n+    # the case that the formats and path are identical.\n+    if as_bytes(path).startswith(as_bytes(prefix)):\n+        return True\n+    \n+    # attempt to convert the prefix and path both to utf8\n+    path_utf8 = encodeWithUTF8(path)\n+    prefix_utf8 = encodeWithUTF8(prefix)\n+\n     if gitConfigBool(\"core.ignorecase\"):\n-        return path.lower().startswith(prefix.lower())\n-    return path.startswith(prefix)\n+        # Check if we match byte-per-byte.  \n+        \n+        return path_utf8.lower().startswith(prefix_utf8.lower())\n+    return path_utf8.startswith(prefix_utf8)\n \n def getClientSpec():\n     \"\"\"Look at the p4 client spec, create a View() object that contains\n@@ -1063,7 +1236,7 @@ def getClientSpec():\n     client_name = entry[\"Client\"]\n \n     # just the keys that start with \"View\"\n-    view_keys = [ k for k in entry.keys() if k.startswith(\"View\") ]\n+    view_keys = [ k for k in list(entry.keys()) if k.startswith(\"View\") ]\n \n     # hold this new View\n     view = View(client_name)\n@@ -1101,18 +1274,24 @@ def wildcard_decode(path):\n     # Cannot have * in a filename in windows; untested as to\n     # what p4 would do in such a case.\n     if not platform.system() == \"Windows\":\n-        path = path.replace(\"%2A\", \"*\")\n-    path = path.replace(\"%23\", \"#\") \\\n-               .replace(\"%40\", \"@\") \\\n-               .replace(\"%25\", \"%\")\n+        path = path.replace(b\"%2A\", b\"*\")\n+    path = path.replace(b\"%23\", b\"#\") \\\n+               .replace(b\"%40\", b\"@\") \\\n+               .replace(b\"%25\", b\"%\")\n     return path\n \n def wildcard_encode(path):\n     # do % first to avoid double-encoding the %s introduced here\n-    path = path.replace(\"%\", \"%25\") \\\n-               .replace(\"*\", \"%2A\") \\\n-               .replace(\"#\", \"%23\") \\\n-               .replace(\"@\", \"%40\")\n+    if isinstance(path, unicode):\n+        path = path.replace(\"%\", \"%25\") \\\n+                   .replace(\"*\", \"%2A\") \\\n+                   .replace(\"#\", \"%23\") \\\n+                   .replace(\"@\", \"%40\")\n+    else:\n+        path = path.replace(b\"%\", b\"%25\") \\\n+                   .replace(b\"*\", b\"%2A\") \\\n+                   .replace(b\"#\", b\"%23\") \\\n+                   .replace(b\"@\", b\"%40\")\n     return path\n \n def wildcard_present(path):\n@@ -1244,7 +1423,7 @@ def generatePointer(self, contentFile):\n             ['git', 'lfs', 'pointer', '--file=' + contentFile],\n             stdout=subprocess.PIPE\n         )\n-        pointerFile = pointerProcess.stdout.read()\n+        pointerFile = as_string(pointerProcess.stdout.read())\n         if pointerProcess.wait():\n             os.remove(contentFile)\n             die('git-lfs pointer command failed. Did you install the extension?')\n@@ -1305,7 +1484,7 @@ def processContent(self, git_mode, relPath, contents):\n         else:\n             return LargeFileSystem.processContent(self, git_mode, relPath, contents)\n \n-class Command:\n+class Command(object):\n     delete_actions = ( \"delete\", \"move/delete\", \"purge\" )\n     add_actions = ( \"add\", \"branch\", \"move/add\" )\n \n@@ -1320,7 +1499,7 @@ def ensure_value(self, attr, value):\n             setattr(self, attr, value)\n         return getattr(self, attr)\n \n-class P4UserMap:\n+class P4UserMap(object):\n     def __init__(self):\n         self.userMapFromPerforceServer = False\n         self.myP4UserId = None\n@@ -1345,10 +1524,14 @@ def p4UserIsMe(self, p4User):\n             return True\n \n     def getUserCacheFilename(self):\n+        \"\"\" Returns the filename of the username cache \"\"\"\n         home = os.environ.get(\"HOME\", os.environ.get(\"USERPROFILE\"))\n-        return home + \"/.gitp4-usercache.txt\"\n+        return os.path.join(home, \".gitp4-usercache.txt\")\n \n     def getUserMapFromPerforceServer(self):\n+        \"\"\" Creates the usercache from the data in P4.\n+        \"\"\"\n+        \n         if self.userMapFromPerforceServer:\n             return\n         self.users = {}\n@@ -1371,21 +1554,24 @@ def getUserMapFromPerforceServer(self):\n                 self.emails[email] = user\n \n         s = ''\n-        for (key, val) in self.users.items():\n+        for (key, val) in list(self.users.items()):\n             s += \"%s\\t%s\\n\" % (key.expandtabs(1), val.expandtabs(1))\n \n-        open(self.getUserCacheFilename(), \"wb\").write(s)\n+        cache = io.open(self.getUserCacheFilename(), \"wb\")\n+        cache.write(as_bytes(s))\n+        cache.close()\n         self.userMapFromPerforceServer = True\n \n     def loadUserMapFromCache(self):\n+        \"\"\" Reads the P4 username to git email map \"\"\"\n         self.users = {}\n         self.userMapFromPerforceServer = False\n         try:\n-            cache = open(self.getUserCacheFilename(), \"rb\")\n+            cache = io.open(self.getUserCacheFilename(), \"rb\")\n             lines = cache.readlines()\n             cache.close()\n             for line in lines:\n-                entry = line.strip().split(\"\\t\")\n+                entry = as_string(line).strip().split(\"\\t\")\n                 self.users[entry[0]] = entry[1]\n         except IOError:\n             self.getUserMapFromPerforceServer()\n@@ -1585,21 +1771,27 @@ def prepareLogMessage(self, template, message, jobs):\n         return result\n \n     def patchRCSKeywords(self, file, pattern):\n-        # Attempt to zap the RCS keywords in a p4 controlled file matching the given pattern\n+        \"\"\" Attempt to zap the RCS keywords in a p4 \n+            controlled file matching the given pattern\n+        \"\"\"\n+        bSubLine = as_bytes(r'$\\1$')\n         (handle, outFileName) = tempfile.mkstemp(dir='.')\n         try:\n-            outFile = os.fdopen(handle, \"w+\")\n-            inFile = open(file, \"r\")\n-            regexp = re.compile(pattern, re.VERBOSE)\n+            outFile = os.fdopen(handle, \"w+b\")\n+            inFile = open(file, \"rb\")\n+            regexp = re.compile(as_bytes(pattern), re.VERBOSE)\n             for line in inFile.readlines():\n-                line = regexp.sub(r'$\\1$', line)\n+                line = regexp.sub(bSubLine, line)\n                 outFile.write(line)\n             inFile.close()\n             outFile.close()\n+            outFile = None\n             # Forcibly overwrite the original file\n             os.unlink(file)\n             shutil.move(outFileName, file)\n         except:\n+            if outFile != None:\n+                outFile.close()\n             # cleanup our temporary file\n             os.unlink(outFileName)\n             print(\"Failed to strip RCS keywords in %s\" % file)\n@@ -1722,14 +1914,14 @@ def prepareSubmitTemplate(self, changelist=None):\n                 break\n         if not change_entry:\n             die('Failed to decode output of p4 change -o')\n-        for key, value in change_entry.iteritems():\n+        for key, value in list(change_entry.items()):\n             if key.startswith('File'):\n                 if 'depot-paths' in settings:\n                     if not [p for p in settings['depot-paths']\n-                            if p4PathStartsWith(value, p)]:\n+                            if p4PathStartsWith(value, p, self.verbose)]:\n                         continue\n                 else:\n-                    if not p4PathStartsWith(value, self.depotPath):\n+                    if not p4PathStartsWith(value, self.depotPath, self.verbose):\n                         continue\n                 files_list.append(value)\n                 continue\n@@ -1779,7 +1971,8 @@ def edit_template(self, template_file):\n             return True\n \n         while True:\n-            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \")\n+            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \").lower() \\\n+                .strip()[0]\n             if response == 'y':\n                 return True\n             if response == 'n':\n@@ -1817,8 +2010,8 @@ def get_diff_description(self, editedFiles, filesToAdd, symlinks):\n     def applyCommit(self, id):\n         \"\"\"Apply one commit, return True if it succeeded.\"\"\"\n \n-        print(\"Applying\", read_pipe([\"git\", \"show\", \"-s\",\n-                                     \"--format=format:%h %s\", id]))\n+        print((\"Applying\", read_pipe([\"git\", \"show\", \"-s\",\n+                                     \"--format=format:%h %s\", id])))\n \n         (p4User, gitEmail) = self.p4UserForCommit(id)\n \n@@ -1939,8 +2132,23 @@ def applyCommit(self, id):\n                     # disable the read-only bit on windows.\n                     if self.isWindows and file not in editedFiles:\n                         os.chmod(file, stat.S_IWRITE)\n-                    self.patchRCSKeywords(file, kwfiles[file])\n-                    fixed_rcs_keywords = True\n+                    \n+                    try:\n+                        self.patchRCSKeywords(file, kwfiles[file])\n+                        fixed_rcs_keywords = True\n+                    except:\n+                        # We are throwing an exception, undo all open edits\n+                        for f in editedFiles:\n+                            p4_revert(f)\n+                        raise\n+            else:\n+                # They do not have attemptRCSCleanup set, this might be the fail point\n+                # Check to see if the file has RCS keywords and suggest setting the property.\n+                for file in editedFiles | filesToDelete:\n+                    if p4_keywords_regexp_for_file(file) != None:\n+                        print(\"At least one file in this commit has RCS Keywords that may be causing problems. \")\n+                        print(\"Consider:\\ngit config git-p4.attemptRCSCleanup true\")\n+                        break\n \n             if fixed_rcs_keywords:\n                 print(\"Retrying the patch with RCS keywords cleaned up\")\n@@ -1966,7 +2174,7 @@ def applyCommit(self, id):\n             p4_delete(f)\n \n         # Set/clear executable bits\n-        for f in filesToChangeExecBit.keys():\n+        for f in list(filesToChangeExecBit.keys()):\n             mode = filesToChangeExecBit[f]\n             setP4ExecBit(f, mode)\n \n@@ -2003,7 +2211,7 @@ def applyCommit(self, id):\n         tmpFile = os.fdopen(handle, \"w+b\")\n         if self.isWindows:\n             submitTemplate = submitTemplate.replace(\"\\n\", \"\\r\\n\")\n-        tmpFile.write(submitTemplate)\n+        tmpFile.write(as_bytes(submitTemplate))\n         tmpFile.close()\n \n         if self.prepare_p4_only:\n@@ -2053,8 +2261,8 @@ def applyCommit(self, id):\n                 message = tmpFile.read()\n                 tmpFile.close()\n                 if self.isWindows:\n-                    message = message.replace(\"\\r\\n\", \"\\n\")\n-                submitTemplate = message[:message.index(separatorLine)]\n+                    message = message.replace(b\"\\r\\n\", b\"\\n\")\n+                submitTemplate = message[:message.index(as_bytes(separatorLine))]\n \n                 if update_shelve:\n                     p4_write_pipe(['shelve', '-r', '-i'], submitTemplate)\n@@ -2164,6 +2372,50 @@ def exportGitTags(self, gitTags):\n                 if verbose:\n                     print(\"created p4 label for tag %s\" % name)\n \n+    def run_hook(self, hook_name, args = []):\n+        \"\"\" Runs a hook if it is found.\n+\n+            Returns NONE if the hook does not exist\n+            Returns TRUE if the exit code is 0, FALSE for a non-zero exit code.\n+        \"\"\"\n+        hook_file = self.find_hook(hook_name)\n+        if hook_file == None:\n+            if self.verbose:\n+                print(\"Skipping hook: %s\" % hook_name)\n+            return None\n+\n+        if self.verbose:\n+            print(\"hooks_path = %s \" % hooks_path)\n+            print(\"hook_file = %s \" % hook_file)\n+\n+        # Run the hook\n+        # TODO - allow non-list format\n+        cmd = [hook_file] + args\n+        return subprocess.call(cmd) == 0\n+\n+    def find_hook(self, hook_name):\n+        \"\"\" Locates the hook file for the given operating system.\n+        \"\"\"\n+        hooks_path = gitConfig(\"core.hooksPath\")\n+        if len(hooks_path) <= 0:\n+            hooks_path = os.path.join(os.environ.get(\"GIT_DIR\", \".git\"), \"hooks\")\n+\n+        # Look in the obvious place\n+        hook_file = os.path.join(hooks_path, hook_name)\n+        if os.path.isfile(hook_file) and os.access(hook_file, os.X_OK):\n+            return hook_file\n+\n+        # if we are windows, we will also allow them to have the hooks have extensions\n+        if (platform.system() == \"Windows\"):\n+            for ext in ['.exe', '.bat', 'ps1']:\n+                if os.path.isfile(hook_file + ext) and os.access(hook_file + ext, os.X_OK):\n+                    return hook_file + ext\n+\n+        # We didn't find the file\n+        return None\n+\n+\n+\n     def run(self, args):\n         if len(args) == 0:\n             self.master = currentGitBranch()\n@@ -2219,7 +2471,7 @@ def run(self, args):\n             self.clientSpecDirs = getClientSpec()\n \n         # Check for the existence of P4 branches\n-        branchesDetected = (len(p4BranchesInGit().keys()) > 1)\n+        branchesDetected = (len(list(p4BranchesInGit().keys())) > 1)\n \n         if self.useClientSpec and not branchesDetected:\n             # all files are relative to the client spec\n@@ -2314,12 +2566,8 @@ def run(self, args):\n             sys.exit(\"number of commits (%d) must match number of shelved changelist (%d)\" %\n                      (len(commits), num_shelves))\n \n-        hooks_path = gitConfig(\"core.hooksPath\")\n-        if len(hooks_path) <= 0:\n-            hooks_path = os.path.join(os.environ.get(\"GIT_DIR\", \".git\"), \"hooks\")\n-\n-        hook_file = os.path.join(hooks_path, \"p4-pre-submit\")\n-        if os.path.isfile(hook_file) and os.access(hook_file, os.X_OK) and subprocess.call([hook_file]) != 0:\n+        rtn = self.run_hook(\"p4-pre-submit\")\n+        if rtn == False:\n             sys.exit(1)\n \n         #\n@@ -2332,8 +2580,8 @@ def run(self, args):\n         last = len(commits) - 1\n         for i, commit in enumerate(commits):\n             if self.dry_run:\n-                print(\" \", read_pipe([\"git\", \"show\", \"-s\",\n-                                      \"--format=format:%h %s\", commit]))\n+                print((\" \", read_pipe([\"git\", \"show\", \"-s\",\n+                                      \"--format=format:%h %s\", commit])))\n                 ok = True\n             else:\n                 ok = self.applyCommit(commit)\n@@ -2351,7 +2599,7 @@ def run(self, args):\n                         if self.conflict_behavior == \"ask\":\n                             print(\"What do you want to do?\")\n                             response = raw_input(\"[s]kip this commit but apply\"\n-                                                 \" the rest, or [q]uit? \")\n+                                                 \" the rest, or [q]uit? \").lower().strip()[0]\n                             if not response:\n                                 continue\n                         elif self.conflict_behavior == \"skip\":\n@@ -2403,8 +2651,8 @@ def run(self, args):\n                         star = \"*\"\n                     else:\n                         star = \" \"\n-                    print(star, read_pipe([\"git\", \"show\", \"-s\",\n-                                           \"--format=format:%h %s\",  c]))\n+                    print((star, read_pipe([\"git\", \"show\", \"-s\",\n+                                           \"--format=format:%h %s\",  c])))\n                 print(\"You will have to do 'git p4 sync' and rebase.\")\n \n         if gitConfigBool(\"git-p4.exportLabels\"):\n@@ -2533,6 +2781,7 @@ def cloneExcludeCallback(option, opt_str, value, parser):\n     # (\"-//depot/A/...\" becomes \"/depot/A/...\" after option parsing)\n     parser.values.cloneExclude += [\"/\" + re.sub(r\"\\.\\.\\.$\", \"\", value)]\n \n+\n class P4Sync(Command, P4UserMap):\n \n     def __init__(self):\n@@ -2610,7 +2859,7 @@ def __init__(self):\n         self.knownBranches = {}\n         self.initialParents = {}\n \n-        self.tz = \"%+03d%02d\" % (- time.timezone / 3600, ((- time.timezone % 3600) / 60))\n+        self.tz = \"%+03d%02d\" % (- time.timezone // 3600, ((- time.timezone % 3600) // 60))\n         self.labels = {}\n \n     # Force a checkpoint in fast-import and wait for it to finish\n@@ -2624,17 +2873,23 @@ def checkpoint(self):\n     def isPathWanted(self, path):\n         for p in self.cloneExclude:\n             if p.endswith(\"/\"):\n-                if p4PathStartsWith(path, p):\n+                if p4PathStartsWith(path, p, self.verbose):\n                     return False\n             # \"-//depot/file1\" without a trailing \"/\" should only exclude \"file1\", but not \"file111\" or \"file1_dir/file2\"\n             elif path.lower() == p.lower():\n                 return False\n         for p in self.depotPaths:\n-            if p4PathStartsWith(path, p):\n+            if p4PathStartsWith(path, p, self.verbose):\n                 return True\n         return False\n \n     def extractFilesFromCommit(self, commit, shelved=False, shelved_cl = 0):\n+        \"\"\" Generates the list of files to be added in this git commit.\n+\n+            commit     = Unicode[] - data read from the P4 commit\n+            shelved    = Bool      - Is the P4 commit flagged as being shelved.\n+            shelved_cl = Unicode   - Numeric string with the changelist number.\n+        \"\"\"\n         files = []\n         fnum = 0\n         while \"depotFile%s\" % fnum in commit:\n@@ -2676,7 +2931,7 @@ def stripRepoPath(self, path, prefixes):\n             path = self.clientSpecDirs.map_in_client(path)\n             if self.detectBranches:\n                 for b in self.knownBranches:\n-                    if p4PathStartsWith(path, b + \"/\"):\n+                    if p4PathStartsWith(path, b + \"/\", self.verbose):\n                         path = path[len(b)+1:]\n \n         elif self.keepRepoPath:\n@@ -2684,12 +2939,12 @@ def stripRepoPath(self, path, prefixes):\n             # //depot/; just look at first prefix as they all should\n             # be in the same depot.\n             depot = re.sub(\"^(//[^/]+/).*\", r'\\1', prefixes[0])\n-            if p4PathStartsWith(path, depot):\n+            if p4PathStartsWith(path, depot, self.verbose):\n                 path = path[len(depot):]\n \n         else:\n             for p in prefixes:\n-                if p4PathStartsWith(path, p):\n+                if p4PathStartsWith(path, p, self.verbose):\n                     path = path[len(p):]\n                     break\n \n@@ -2697,8 +2952,11 @@ def stripRepoPath(self, path, prefixes):\n         return path\n \n     def splitFilesIntoBranches(self, commit):\n-        \"\"\"Look at each depotFile in the commit to figure out to what\n-           branch it belongs.\"\"\"\n+        \"\"\" Look at each depotFile in the commit to figure out to what\n+            branch it belongs.\n+\n+            Data in the commit will NOT be encoded\n+        \"\"\"\n \n         if self.clientSpecDirs:\n             files = self.extractFilesFromCommit(commit)\n@@ -2727,10 +2985,10 @@ def splitFilesIntoBranches(self, commit):\n             else:\n                 relPath = self.stripRepoPath(path, self.depotPaths)\n \n-            for branch in self.knownBranches.keys():\n+            for branch in list(self.knownBranches.keys()):\n                 # add a trailing slash so that a commit into qt/4.2foo\n                 # doesn't end up in qt/4.2, e.g.\n-                if p4PathStartsWith(relPath, branch + \"/\"):\n+                if p4PathStartsWith(relPath, branch + \"/\", self.verbose):\n                     if branch not in branches:\n                         branches[branch] = []\n                     branches[branch].append(file)\n@@ -2739,36 +2997,34 @@ def splitFilesIntoBranches(self, commit):\n         return branches\n \n     def writeToGitStream(self, gitMode, relPath, contents):\n-        self.gitStream.write('M %s inline %s\\n' % (gitMode, relPath))\n+        \"\"\" Writes the bytes[] 'contents' to the git fast-import\n+            with the given 'gitMode' and 'relPath' as the relative\n+            path.\n+        \"\"\"\n+        self.gitStream.write('M %s inline %s\\n' % (gitMode, as_string(relPath)))\n         self.gitStream.write('data %d\\n' % sum(len(d) for d in contents))\n         for d in contents:\n-            self.gitStream.write(d)\n+            self.gitStreamBytes.write(d)\n         self.gitStream.write('\\n')\n \n-    def encodeWithUTF8(self, path):\n-        try:\n-            path.decode('ascii')\n-        except:\n-            encoding = 'utf8'\n-            if gitConfig('git-p4.pathEncoding'):\n-                encoding = gitConfig('git-p4.pathEncoding')\n-            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n-            if self.verbose:\n-                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n-        return path\n-\n-    # output one file from the P4 stream\n-    # - helper for streamP4Files\n-\n     def streamOneP4File(self, file, contents):\n+        \"\"\" output one file from the P4 stream to the git inbound stream.\n+            helper for streamP4files.\n+\n+            contents should be a bytes (bytes) \n+        \"\"\"\n         relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n-        relPath = self.encodeWithUTF8(relPath)\n+        relPath = encodeWithUTF8(relPath, self.verbose)\n         if verbose:\n             if 'fileSize' in self.stream_file:\n                 size = int(self.stream_file['fileSize'])\n             else:\n                 size = 0 # deleted files don't get a fileSize apparently\n-            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size/1024/1024))\n+            #if isunicode:\n+            #    sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), to_unicode(relPath), size//1024//1024))\n+            #else:\n+            #    sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), relPath, size//1024//1024))\n+            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), as_string(relPath), size//1024//1024))\n             sys.stdout.flush()\n \n         (type_base, type_mods) = split_p4_type(file[\"type\"])\n@@ -2786,7 +3042,7 @@ def streamOneP4File(self, file, contents):\n                 # to nothing.  This causes p4 errors when checking out such\n                 # a change, and errors here too.  Work around it by ignoring\n                 # the bad symlink; hopefully a future change fixes it.\n-                print(\"\\nIgnoring empty symlink in %s\" % file['depotFile'])\n+                print(\"\\nIgnoring empty symlink in %s\" % path_as_string(file['depotFile']))\n                 return\n             elif data[-1] == '\\n':\n                 contents = [data[:-1]]\n@@ -2826,16 +3082,16 @@ def streamOneP4File(self, file, contents):\n             # Ideally, someday, this script can learn how to generate\n             # appledouble files directly and import those to git, but\n             # non-mac machines can never find a use for apple filetype.\n-            print(\"\\nIgnoring apple filetype file %s\" % file['depotFile'])\n+            print(\"\\nIgnoring apple filetype file %s\" % path_as_string(file['depotFile']))\n             return\n \n         # Note that we do not try to de-mangle keywords on utf16 files,\n         # even though in theory somebody may want that.\n-        pattern = p4_keywords_regexp_for_type(type_base, type_mods)\n+        pattern = as_bytes(p4_keywords_regexp_for_type(type_base, type_mods))\n         if pattern:\n             regexp = re.compile(pattern, re.VERBOSE)\n-            text = ''.join(contents)\n-            text = regexp.sub(r'$\\1$', text)\n+            text = b''.join(contents)\n+            text = regexp.sub(as_bytes(r'$\\1$'), text)\n             contents = [ text ]\n \n         if self.largeFileSystem:\n@@ -2845,7 +3101,7 @@ def streamOneP4File(self, file, contents):\n \n     def streamOneP4Deletion(self, file):\n         relPath = self.stripRepoPath(file['path'], self.branchPrefixes)\n-        relPath = self.encodeWithUTF8(relPath)\n+        relPath = encodeWithUTF8(relPath, self.verbose)\n         if verbose:\n             sys.stdout.write(\"delete %s\\n\" % relPath)\n             sys.stdout.flush()\n@@ -2854,21 +3110,25 @@ def streamOneP4Deletion(self, file):\n         if self.largeFileSystem and self.largeFileSystem.isLargeFile(relPath):\n             self.largeFileSystem.removeLargeFile(relPath)\n \n-    # handle another chunk of streaming data\n     def streamP4FilesCb(self, marshalled):\n+        \"\"\" Callback function for recording P4 chunks of data for streaming \n+            into GIT.\n+\n+            marshalled data is bytes[] from the caller\n+        \"\"\"\n \n         # catch p4 errors and complain\n         err = None\n-        if \"code\" in marshalled:\n-            if marshalled[\"code\"] == \"error\":\n-                if \"data\" in marshalled:\n-                    err = marshalled[\"data\"].rstrip()\n+        if b\"code\" in marshalled:\n+            if marshalled[b\"code\"] == b\"error\":\n+                if b\"data\" in marshalled:\n+                    err = marshalled[b\"data\"].rstrip()\n \n         if not err and 'fileSize' in self.stream_file:\n             required_bytes = int((4 * int(self.stream_file[\"fileSize\"])) - calcDiskFree())\n             if required_bytes > 0:\n                 err = 'Not enough space left on %s! Free at least %i MB.' % (\n-                    os.getcwd(), required_bytes/1024/1024\n+                    os.getcwd(), required_bytes//1024//1024\n                 )\n \n         if err:\n@@ -2884,11 +3144,11 @@ def streamP4FilesCb(self, marshalled):\n             # ignore errors, but make sure it exits first\n             self.importProcess.wait()\n             if f:\n-                die(\"Error from p4 print for %s: %s\" % (f, err))\n+                die(\"Error from p4 print for %s: %s\" % (path_as_string(f), err))\n             else:\n                 die(\"Error from p4 print: %s\" % err)\n \n-        if 'depotFile' in marshalled and self.stream_have_file_info:\n+        if b'depotFile' in marshalled and self.stream_have_file_info:\n             # start of a new file - output the old one first\n             self.streamOneP4File(self.stream_file, self.stream_contents)\n             self.stream_file = {}\n@@ -2897,14 +3157,17 @@ def streamP4FilesCb(self, marshalled):\n \n         # pick up the new file information... for the\n         # 'data' field we need to append to our array\n-        for k in marshalled.keys():\n-            if k == 'data':\n+        for k in list(marshalled.keys()):\n+            if k == b'data':\n                 if 'streamContentSize' not in self.stream_file:\n                     self.stream_file['streamContentSize'] = 0\n-                self.stream_file['streamContentSize'] += len(marshalled['data'])\n-                self.stream_contents.append(marshalled['data'])\n+                self.stream_file['streamContentSize'] += len(marshalled[b'data'])\n+                self.stream_contents.append(marshalled[b'data'])\n             else:\n-                self.stream_file[k] = marshalled[k]\n+                if k == b'depotFile':\n+                    self.stream_file[as_string(k)] = marshalled[k]\n+                else:\n+                    self.stream_file[as_string(k)] = as_string(marshalled[k])\n \n         if (verbose and\n             'streamContentSize' in self.stream_file and\n@@ -2912,14 +3175,15 @@ def streamP4FilesCb(self, marshalled):\n             'depotFile' in self.stream_file):\n             size = int(self.stream_file[\"fileSize\"])\n             if size > 0:\n-                progress = 100*self.stream_file['streamContentSize']/size\n-                sys.stdout.write('\\r%s %d%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size/1024/1024)))\n+                progress = 100.0*self.stream_file['streamContentSize']/size\n+                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (path_as_string(self.stream_file['depotFile']), progress, int(size//1024//1024)))\n                 sys.stdout.flush()\n \n         self.stream_have_file_info = True\n \n-    # Stream directly from \"p4 files\" into \"git fast-import\"\n     def streamP4Files(self, files):\n+        \"\"\" Stream directly from \"p4 files\" into \"git fast-import\" \n+        \"\"\"\n         filesForCommit = []\n         filesToRead = []\n         filesToDelete = []\n@@ -2940,7 +3204,7 @@ def streamP4Files(self, files):\n             self.stream_contents = []\n             self.stream_have_file_info = False\n \n-            # curry self argument\n+            # Callback for P4 command to collect file content\n             def streamP4FilesCbSelf(entry):\n                 self.streamP4FilesCb(entry)\n \n@@ -2949,9 +3213,9 @@ def streamP4FilesCbSelf(entry):\n                 if 'shelved_cl' in f:\n                     # Handle shelved CLs using the \"p4 print file@=N\" syntax to print\n                     # the contents\n-                    fileArg = '%s@=%d' % (f['path'], f['shelved_cl'])\n+                    fileArg = b'%s@=%d' % (f['path'], as_bytes(f['shelved_cl']))\n                 else:\n-                    fileArg = '%s#%s' % (f['path'], f['rev'])\n+                    fileArg = b'%s#%s' % (f['path'], as_bytes(f['rev']))\n \n                 fileArgs.append(fileArg)\n \n@@ -2971,7 +3235,7 @@ def make_email(self, userid):\n \n     def streamTag(self, gitStream, labelName, labelDetails, commit, epoch):\n         \"\"\" Stream a p4 tag.\n-        commit is either a git commit, or a fast-import mark, \":<p4commit>\"\n+            commit is either a git commit, or a fast-import mark, \":<p4commit>\"\n         \"\"\"\n \n         if verbose:\n@@ -2994,7 +3258,7 @@ def streamTag(self, gitStream, labelName, labelDetails, commit, epoch):\n \n         gitStream.write(\"tagger %s\\n\" % tagger)\n \n-        print(\"labelDetails=\",labelDetails)\n+        print((\"labelDetails=\",labelDetails))\n         if 'Description' in labelDetails:\n             description = labelDetails['Description']\n         else:\n@@ -3016,7 +3280,7 @@ def hasBranchPrefix(self, path):\n         if not self.branchPrefixes:\n             return True\n         hasPrefix = [p for p in self.branchPrefixes\n-                        if p4PathStartsWith(path, p)]\n+                        if p4PathStartsWith(path, p, self.verbose)]\n         if not hasPrefix and self.verbose:\n             print('Ignoring file outside of prefix: {0}'.format(path))\n         return hasPrefix\n@@ -3043,7 +3307,22 @@ def commit(self, details, files, branch, parent = \"\", allow_empty=False):\n                 .format(details['change']))\n             return\n \n+        # fast-import:\n+        #'commit' SP <ref> LF\n+\t    #mark?\n+\t    #original-oid?\n+\t    #('author' (SP <name>)? SP LT <email> GT SP <when> LF)?\n+\t    #'committer' (SP <name>)? SP LT <email> GT SP <when> LF\n+\t    #('encoding' SP <encoding>)?\n+\t    #data\n+\t    #('from' SP <commit-ish> LF)?\n+\t    #('merge' SP <commit-ish> LF)*\n+\t    #(filemodify | filedelete | filecopy | filerename | filedeleteall | notemodify)*\n+\t    #LF?\n+        \n+        #'commit' - <ref> is the name of the branch to make the commit on\n         self.gitStream.write(\"commit %s\\n\" % branch)\n+        #'mark' SP :<idnum>\n         self.gitStream.write(\"mark :%s\\n\" % details[\"change\"])\n         self.committedChanges.add(int(details[\"change\"]))\n         committer = \"\"\n@@ -3053,19 +3332,29 @@ def commit(self, details, files, branch, parent = \"\", allow_empty=False):\n \n         self.gitStream.write(\"committer %s\\n\" % committer)\n \n-        self.gitStream.write(\"data <<EOT\\n\")\n-        self.gitStream.write(details[\"desc\"])\n+        # Per https://git-scm.com/docs/git-fast-import\n+        # The preferred method for creating the commit message is to supply the \n+        # byte count in the data method and not to use a Delimited format. \n+        # Collect all the text in the commit message into a single string and \n+        # compute the byte count.\n+        commitText = details[\"desc\"]\n         if len(jobs) > 0:\n-            self.gitStream.write(\"\\nJobs: %s\" % (' '.join(jobs)))\n-\n+            commitText += \"\\nJobs: %s\" % (' '.join(jobs))\n         if not self.suppress_meta_comment:\n-            self.gitStream.write(\"\\n[git-p4: depot-paths = \\\"%s\\\": change = %s\" %\n-                                (','.join(self.branchPrefixes), details[\"change\"]))\n-            if len(details['options']) > 0:\n-                self.gitStream.write(\": options = %s\" % details['options'])\n-            self.gitStream.write(\"]\\n\")\n+            # coherce the path to the correct formatting in the branch prefixes as well.\n+            dispPaths = []\n+            for p in self.branchPrefixes:\n+                dispPaths += [path_as_string(p)]\n \n-        self.gitStream.write(\"EOT\\n\\n\")\n+            commitText += (\"\\n[git-p4: depot-paths = \\\"%s\\\": change = %s\" %\n+                                (','.join(dispPaths), details[\"change\"]))\n+            if len(details['options']) > 0:\n+                commitText += (\": options = %s\" % details['options'])\n+            commitText += \"]\"\n+        commitText += \"\\n\" \n+        self.gitStream.write(\"data %s\\n\" % len(as_bytes(commitText)))\n+        self.gitStream.write(commitText)\n+        self.gitStream.write(\"\\n\")\n \n         if len(parent) > 0:\n             if self.verbose:\n@@ -3133,7 +3422,7 @@ def getLabels(self):\n             self.labels[newestChange] = [output, revisions]\n \n         if self.verbose:\n-            print(\"Label changes: %s\" % self.labels.keys())\n+            print(\"Label changes: %s\" % list(self.labels.keys()))\n \n     # Import p4 labels as git tags. A direct mapping does not\n     # exist, so assume that if all the files are at the same revision\n@@ -3234,7 +3523,7 @@ def getBranchMapping(self):\n                 source = paths[0]\n                 destination = paths[1]\n                 ## HACK\n-                if p4PathStartsWith(source, self.depotPaths[0]) and p4PathStartsWith(destination, self.depotPaths[0]):\n+                if p4PathStartsWith(source, self.depotPaths[0], self.verbose) and p4PathStartsWith(destination, self.depotPaths[0], self.verbose):\n                     source = source[len(self.depotPaths[0]):-4]\n                     destination = destination[len(self.depotPaths[0]):-4]\n \n@@ -3276,7 +3565,7 @@ def getBranchMapping(self):\n \n     def getBranchMappingFromGitBranches(self):\n         branches = p4BranchesInGit(self.importIntoRemotes)\n-        for branch in branches.keys():\n+        for branch in list(branches.keys()):\n             if branch == \"master\":\n                 branch = \"main\"\n             else:\n@@ -3388,14 +3677,14 @@ def importChanges(self, changes, origin_revision=0):\n             self.updateOptionDict(description)\n \n             if not self.silent:\n-                sys.stdout.write(\"\\rImporting revision %s (%s%%)\" % (change, cnt * 100 / len(changes)))\n+                sys.stdout.write(\"\\rImporting revision %s (%4.1f%%)\" % (change, cnt * 100 / len(changes)))\n                 sys.stdout.flush()\n             cnt = cnt + 1\n \n             try:\n                 if self.detectBranches:\n                     branches = self.splitFilesIntoBranches(description)\n-                    for branch in branches.keys():\n+                    for branch in list(branches.keys()):\n                         ## HACK  --hwn\n                         branchPrefix = self.depotPaths[0] + branch + \"/\"\n                         self.branchPrefixes = [ branchPrefix ]\n@@ -3464,6 +3753,7 @@ def importChanges(self, changes, origin_revision=0):\n                 sys.exit(1)\n \n     def sync_origin_only(self):\n+        \"\"\" Ensures that the origin has been synchronized if one is set \"\"\"\n         if self.syncWithOrigin:\n             self.hasOrigin = originP4BranchesExist()\n             if self.hasOrigin:\n@@ -3472,30 +3762,35 @@ def sync_origin_only(self):\n                 system(\"git fetch origin\")\n \n     def importHeadRevision(self, revision):\n-        print(\"Doing initial import of %s from revision %s into %s\" % (' '.join(self.depotPaths), revision, self.branch))\n-\n+        # Re-encode depot text\n+        dispPaths = []\n+        utf8Paths = []\n+        for p in self.depotPaths:\n+            dispPaths += [path_as_string(p)]\n+        print(\"Doing initial import of %s from revision %s into %s\" % (' '.join(dispPaths), revision, self.branch))\n         details = {}\n         details[\"user\"] = \"git perforce import user\"\n-        details[\"desc\"] = (\"Initial import of %s from the state at revision %s\\n\"\n-                           % (' '.join(self.depotPaths), revision))\n+        details[\"desc\"] = (\"Initial import of %s from the state at revision %s\\n\" %\n+                           (' '.join(dispPaths), revision))\n         details[\"change\"] = revision\n         newestRevision = 0\n+        del dispPaths\n \n         fileCnt = 0\n         fileArgs = [\"%s...%s\" % (p,revision) for p in self.depotPaths]\n \n-        for info in p4CmdList([\"files\"] + fileArgs):\n+        for info in p4CmdList([\"files\"] + fileArgs, encode_data = False):\n \n-            if 'code' in info and info['code'] == 'error':\n+            if 'code' in info and info['code'] == b'error':\n                 sys.stderr.write(\"p4 returned an error: %s\\n\"\n-                                 % info['data'])\n-                if info['data'].find(\"must refer to client\") >= 0:\n+                                 % as_string(info['data']))\n+                if info['data'].find(b\"must refer to client\") >= 0:\n                     sys.stderr.write(\"This particular p4 error is misleading.\\n\")\n                     sys.stderr.write(\"Perhaps the depot path was misspelled.\\n\");\n                     sys.stderr.write(\"Depot path:  %s\\n\" % \" \".join(self.depotPaths))\n                 sys.exit(1)\n             if 'p4ExitCode' in info:\n-                sys.stderr.write(\"p4 exitcode: %s\\n\" % info['p4ExitCode'])\n+                sys.stderr.write(\"p4 exitcode: %s\\n\" % as_string(info['p4ExitCode']))\n                 sys.exit(1)\n \n \n@@ -3508,8 +3803,10 @@ def importHeadRevision(self, revision):\n                 #fileCnt = fileCnt + 1\n                 continue\n \n+            # Save all the file information, howerver do not translate the depotFile name at \n+            # this time. Leave that as bytes since the encoding may vary.\n             for prop in [\"depotFile\", \"rev\", \"action\", \"type\" ]:\n-                details[\"%s%s\" % (prop, fileCnt)] = info[prop]\n+                details[\"%s%s\" % (prop, fileCnt)] = (info[prop] if prop == \"depotFile\" else as_string(info[prop]))\n \n             fileCnt = fileCnt + 1\n \n@@ -3529,13 +3826,18 @@ def importHeadRevision(self, revision):\n             print(self.gitError.read())\n \n     def openStreams(self):\n+        \"\"\" Opens the fast import pipes.  Note that the git* streams are wrapped\n+            to expect Unicode text.  To send a raw byte Array, use the importProcess\n+            underlying port\n+        \"\"\"\n         self.importProcess = subprocess.Popen([\"git\", \"fast-import\"],\n                                               stdin=subprocess.PIPE,\n                                               stdout=subprocess.PIPE,\n                                               stderr=subprocess.PIPE);\n-        self.gitOutput = self.importProcess.stdout\n-        self.gitStream = self.importProcess.stdin\n-        self.gitError = self.importProcess.stderr\n+        self.gitOutput = Py23File(self.importProcess.stdout, verbose = self.verbose)\n+        self.gitStream = Py23File(self.importProcess.stdin, verbose = self.verbose)\n+        self.gitError = Py23File(self.importProcess.stderr, verbose = self.verbose)\n+        self.gitStreamBytes = self.importProcess.stdin\n \n     def closeStreams(self):\n         self.gitStream.close()\n@@ -3584,13 +3886,13 @@ def run(self, args):\n                 if short in branches:\n                     self.p4BranchesInGit = [ short ]\n             else:\n-                self.p4BranchesInGit = branches.keys()\n+                self.p4BranchesInGit = list(branches.keys())\n \n             if len(self.p4BranchesInGit) > 1:\n                 if not self.silent:\n                     print(\"Importing from/into multiple branches\")\n                 self.detectBranches = True\n-                for branch in branches.keys():\n+                for branch in list(branches.keys()):\n                     self.initialParents[self.refPrefix + branch] = \\\n                         branches[branch]\n \n@@ -3870,19 +4172,25 @@ def __init__(self):\n                                  help=\"where to leave result of the clone\"),\n             optparse.make_option(\"--bare\", dest=\"cloneBare\",\n                                  action=\"store_true\", default=False),\n+            optparse.make_option(\"--encoding\", dest=\"setPathEncoding\",\n+                                 action=\"store\", default=None,\n+                                 help=\"Sets the path encoding for this depot\")\n         ]\n         self.cloneDestination = None\n         self.needsGit = False\n         self.cloneBare = False\n+        self.setPathEncoding = None\n \n     def defaultDestination(self, args):\n+        \"\"\"Returns the last path component as the default git \n+        repository directory name\"\"\"\n         ## TODO: use common prefix of args?\n         depotPath = args[0]\n         depotDir = re.sub(\"(@[^@]*)$\", \"\", depotPath)\n         depotDir = re.sub(\"(#[^#]*)$\", \"\", depotDir)\n         depotDir = re.sub(r\"\\.\\.\\.$\", \"\", depotDir)\n         depotDir = re.sub(r\"/$\", \"\", depotDir)\n-        return os.path.split(depotDir)[1]\n+        return depotDir.split('/')[-1]\n \n     def run(self, args):\n         if len(args) < 1:\n@@ -3894,19 +4202,29 @@ def run(self, args):\n \n         depotPaths = args\n \n+        # If we have an encoding provided, ignore what may already exist\n+        # in the registry. This will ensure we show the displayed values\n+        # using the correct encoding.\n+        if self.setPathEncoding:\n+            gitConfigSet(\"git-p4.pathEncoding\", self.setPathEncoding)\n+\n+        # If more than 1 path element is supplied, the last element\n+        # is the clone destination.\n         if not self.cloneDestination and len(depotPaths) > 1:\n             self.cloneDestination = depotPaths[-1]\n             depotPaths = depotPaths[:-1]\n \n+        dispPaths = []\n         for p in depotPaths:\n             if not p.startswith(\"//\"):\n                 sys.stderr.write('Depot paths must start with \"//\": %s\\n' % p)\n                 return False\n+            dispPaths += [path_as_string(p)]\n \n         if not self.cloneDestination:\n             self.cloneDestination = self.defaultDestination(args)\n \n-        print(\"Importing from %s into %s\" % (', '.join(depotPaths), self.cloneDestination))\n+        print(\"Importing from %s into %s\" % (', '.join(dispPaths), path_as_string(self.cloneDestination)))\n \n         if not os.path.exists(self.cloneDestination):\n             os.makedirs(self.cloneDestination)\n@@ -3919,6 +4237,13 @@ def run(self, args):\n         if retcode:\n             raise CalledProcessError(retcode, init_cmd)\n \n+        # Set the encoding if it was provided command line\n+        if self.setPathEncoding:\n+            init_cmd= [\"git\", \"config\", \"git-p4.pathEncoding\", self.setPathEncoding]\n+            retcode = subprocess.call(init_cmd)\n+            if retcode:\n+                raise CalledProcessError(retcode, init_cmd)\n+\n         if not P4Sync.run(self, depotPaths):\n             return False\n \n@@ -3974,7 +4299,7 @@ def findLastP4Revision(self, starting_point):\n             to find the P4 commit we are based on, and the depot-paths.\n         \"\"\"\n \n-        for parent in (range(65535)):\n+        for parent in (list(range(65535))):\n             log = extractLogMessageFromGitCommit(\"{0}^{1}\".format(starting_point, parent))\n             settings = extractSettingsGitLog(log)\n             if 'change' in settings:\n@@ -4080,6 +4405,107 @@ def run(self, args):\n             print(\"%s <= %s (%s)\" % (branch, \",\".join(settings[\"depot-paths\"]), settings[\"change\"]))\n         return True\n \n+class Py23File():\n+    \"\"\" Python2/3 Unicode File Wrapper \n+    \"\"\"\n+    \n+    stream_handle = None\n+    verbose       = False\n+    debug_handle  = None\n+   \n+    def __init__(self, stream_handle, verbose = False,\n+                 debug_handle = None):\n+        \"\"\" Create a Python3 compliant Unicode to Byte String\n+            Windows compatible wrapper\n+\n+            stream_handle = the underlying file-like handle\n+            verbose       = Boolean if content should be echoed\n+            debug_handle  = A file-like handle data is duplicately written to\n+        \"\"\"\n+        self.stream_handle = stream_handle\n+        self.verbose       = verbose\n+        self.debug_handle  = debug_handle\n+\n+    def write(self, utf8string):\n+        \"\"\" Writes the utf8 encoded string to the underlying \n+            file stream\n+        \"\"\"\n+        self.stream_handle.write(as_bytes(utf8string))\n+        if self.verbose:\n+            sys.stderr.write(\"Stream Output: %s\" % utf8string)\n+            sys.stderr.flush()\n+        if self.debug_handle:\n+            self.debug_handle.write(as_bytes(utf8string))\n+\n+    def read(self, size = None):\n+        \"\"\" Reads int charcters from the underlying stream \n+            and converts it to utf8.\n+\n+            Be aware, the size value is for reading the underlying\n+            bytes so the value may be incorrect. Usage of the size\n+            value is discouraged.\n+        \"\"\"\n+        if size == None:\n+            return as_string(self.stream_handle.read())\n+        else:\n+            return as_string(self.stream_handle.read(size))\n+\n+    def readline(self):\n+        \"\"\" Reads a line from the underlying byte stream \n+            and converts it to utf8\n+        \"\"\"\n+        return as_string(self.stream_handle.readline())\n+\n+    def readlines(self, sizeHint = None):\n+        \"\"\" Returns a list containing lines from the file converted to unicode.\n+\n+            sizehint - Optional. If the optional sizehint argument is \n+            present, instead of reading up to EOF, whole lines totalling \n+            approximately sizehint bytes are read.\n+        \"\"\"\n+        lines = self.stream_handle.readlines(sizeHint)\n+        for i in range(0, len(lines)):\n+            lines[i] = as_string(lines[i])\n+        return lines\n+\n+    def close(self):\n+        \"\"\" Closes the underlying byte stream \"\"\"\n+        self.stream_handle.close()\n+\n+    def flush(self):\n+        \"\"\" Flushes the underlying byte stream \"\"\"\n+        self.stream_handle.flush()\n+\n+class DepotPath():\n+    \"\"\" Describes a DepotPath or File\n+    \"\"\"\n+\n+    raw_path = None\n+    utf8_path = None\n+    bytes_path = None\n+\n+    def __init__(self, path):\n+        \"\"\" Creates a new DepotPath with the path encoded\n+            with by the P4 repository\n+        \"\"\"\n+        raw_path = path\n+\n+    def raw():\n+        \"\"\" Returns the path as it was originally found\n+            in the P4 repository\n+        \"\"\"\n+        return raw_path\n+\n+    def startswith(self, prefix, start = None, end = None):\n+        \"\"\" Return True if string starts with the prefix, otherwise \n+            return False. prefix can also be a tuple of prefixes to \n+            look for. With optional start, test string beginning at \n+            that position. With optional end, stop comparing \n+            string at that position.\n+        \"\"\"\n+        return raw_path.startswith(prefix, start, end)\n+\n+\n class HelpFormatter(optparse.IndentedHelpFormatter):\n     def __init__(self):\n         optparse.IndentedHelpFormatter.__init__(self)\n@@ -4113,7 +4539,7 @@ def printUsage(commands):\n \n def main():\n     if len(sys.argv[1:]) == 0:\n-        printUsage(commands.keys())\n+        printUsage(list(commands.keys()))\n         sys.exit(2)\n \n     cmdName = sys.argv[1]\n@@ -4123,7 +4549,7 @@ def main():\n     except KeyError:\n         print(\"unknown command %s\" % cmdName)\n         print(\"\")\n-        printUsage(commands.keys())\n+        printUsage(list(commands.keys()))\n         sys.exit(2)\n \n     options = cmd.options\n@@ -4140,7 +4566,12 @@ def main():\n                                    description = cmd.description,\n                                    formatter = HelpFormatter())\n \n-    (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n+    try:\n+        (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n+    except:\n+        parser.print_help()\n+        raise\n+\n     global verbose\n     verbose = cmd.verbose\n     if cmd.needsGit:\n@@ -4155,8 +4586,8 @@ def main():\n                         chdir(cdup);\n \n         if not isValidGitDir(cmd.gitdir):\n-            if isValidGitDir(cmd.gitdir + \"/.git\"):\n-                cmd.gitdir += \"/.git\"\n+            if isValidGitDir(os.path.join(cmd.gitdir, \".git\")):\n+                cmd.gitdir = os.path.join(cmd.gitdir, \".git\")\n             else:\n                 die(\"fatal: cannot locate git repository at %s\" % cmd.gitdir)\n \n-- \ngitgitgadget\n"},{"id":"387427","messageId":"20191203001854.GA59841@generichostname","threadId":"52262","inReplyTo":"02b3843e9f21105a945335d0b1d78251ddcc8cee.1575313336.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v3 1/1] Python3 support for t9800 tests. Basic P4/Python3 support","fromName":"Denton Liu","fromEmail":"liu.denton@gmail.com","sentAt":"2019-12-03T00:18:54Z","receivedAt":"2019-12-03T00:19:02Z","isPatch":true,"sender":{"key":"liu.denton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/9620836?v=4"},"body":"Hi Ben,\n\nThanks for the contribution!\n\n> Subject: Python3 support for t9800 tests. Basic P4/Python3 support\n\nIn git.git, the convention for commit subjects is to use \n\"<area>: <summary>\". Perhaps something like, \"git-p4: support Python 3\"?\nAlthough I doubt this patch should remain as is... More below.\n\nOn Mon, Dec 02, 2019 at 07:02:16PM +0000, Ben Keene via GitGitGadget wrote:\n> From: Ben Keene <seraphire@gmail.com>\n\nIt would be nice to have a bit more information about what this patch\ndoes. Could you please fill this in with some more details about the\nwhats and, more importantly, the _whys_ of your change?\n\n> \n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> ---\n>  git-p4.py | 825 +++++++++++++++++++++++++++++++++++++++++-------------\n>  1 file changed, 628 insertions(+), 197 deletions(-)\n\nThis is a very big change to be done in one patch. Could you please\nsplit this into multiple smaller patches that each do one logical\nchange? For example, you could have the following series of changes:\n\n\t1. git-p4: use p4.exe if on Windows\n\t2. git-p4: introduce encoding helper functions # this is to\n\t        introduce the as_string(), as_bytes(), etc. functions\n\t3. git-p4: start using the encoding helper functions\n\t...\n\nThis was just an example and you don't have to follow those literally. I\njust wanted to give you an idea of what I meant.\n\nYou can see Documentation/SubmittingPatches#separate-commits for more\ninformation.\n\nThanks,\n\nDenton\n"},{"id":"387457","messageId":"5c299c98-dda4-768d-7307-666bde85715d@gmail.com","threadId":"52262","inReplyTo":"20191203001854.GA59841@generichostname","subject":"Re: [PATCH v3 1/1] Python3 support for t9800 tests. Basic P4/Python3 support","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-03T16:03:31Z","receivedAt":"2019-12-03T16:03:35Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/2/2019 7:18 PM, Denton Liu wrote:\n> Hi Ben,\n>\n> Thanks for the contribution!\n>\n>> Subject: Python3 support for t9800 tests. Basic P4/Python3 support\n> In git.git, the convention for commit subjects is to use\n> \"<area>: <summary>\". Perhaps something like, \"git-p4: support Python 3\"?\n> Although I doubt this patch should remain as is... More below.\nI didn't realize the email message from gitgitgadget was going to be the \ncommit message, I thought it was the PR message.  I'll work on changing \nthat!\n>\n> On Mon, Dec 02, 2019 at 07:02:16PM +0000, Ben Keene via GitGitGadget wrote:\n>> From: Ben Keene <seraphire@gmail.com>\n> It would be nice to have a bit more information about what this patch\n> does. Could you please fill this in with some more details about the\n> whats and, more importantly, the _whys_ of your change?\nSure, I'll add more detail.\n>> Signed-off-by: Ben Keene <seraphire@gmail.com>\n>> ---\n>>   git-p4.py | 825 +++++++++++++++++++++++++++++++++++++++++-------------\n>>   1 file changed, 628 insertions(+), 197 deletions(-)\n> This is a very big change to be done in one patch. Could you please\n> split this into multiple smaller patches that each do one logical\n> change? For example, you could have the following series of changes:\n>\n> \t1. git-p4: use p4.exe if on Windows\n> \t2. git-p4: introduce encoding helper functions # this is to\n> \t        introduce the as_string(), as_bytes(), etc. functions\n> \t3. git-p4: start using the encoding helper functions\n> \t...\n>\n> This was just an example and you don't have to follow those literally. I\n> just wanted to give you an idea of what I meant.\n>\n> You can see Documentation/SubmittingPatches#separate-commits for more\n> information.\n>\n> Thanks,\n>\n> Denton\nSo my last question would be, should I open a different PR on \ngitgitgadget? I can cherry-pick my changes into another branch and \nrestart my submission?\n"},{"id":"387476","messageId":"20191204061407.GA3381576@generichostname","threadId":"52262","inReplyTo":"5c299c98-dda4-768d-7307-666bde85715d@gmail.com","subject":"Re: [PATCH v3 1/1] Python3 support for t9800 tests. Basic P4/Python3 support","fromName":"Denton Liu","fromEmail":"liu.denton@gmail.com","sentAt":"2019-12-04T06:14:07Z","receivedAt":"2019-12-04T06:51:22Z","isPatch":true,"sender":{"key":"liu.denton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/9620836?v=4"},"body":"On Tue, Dec 03, 2019 at 11:03:31AM -0500, Ben Keene wrote:\n> So my last question would be, should I open a different PR on gitgitgadget?\n> I can cherry-pick my changes into another branch and restart my submission?\n\nYou can reuse the same PR. Just force-push to overwrite your old commits\nand then you'll be able to `/submit` again to send another revision.\n"},{"id":"387531","messageId":"40124269933691796ef57fd8df50f9e740d103b1.1575498577.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 01/11] git-p4: select p4 binary by operating-system","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:27Z","receivedAt":"2019-12-04T22:29:43Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nDepending on the version of GIT and Python installed, the perforce program (p4) may not resolve on Windows without the program extension.\n\nCheck the operating system (platform.system) and if it is reporting that it is Windows, use the full filename of \"p4.exe\" instead of \"p4\"\n\nThe original code unconditionally used \"p4\" as the binary filename.\n\nThis change is Python2 and Python3 compatible.\n\nThanks to: Junio C Hamano <gitster@pobox.com> and  Denton Liu <liu.denton@gmail.com> for patiently explaining proper format for my submissions.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit 9a3a5c4e6d29dbef670072a9605c7a82b3729434)\n---\n git-p4.py | 6 +++++-\n 1 file changed, 5 insertions(+), 1 deletion(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 60c73b6a37..b2ffbc057b 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -75,7 +75,11 @@ def p4_build_cmd(cmd):\n     location. It means that hooking into the environment, or other configuration\n     can be done more easily.\n     \"\"\"\n-    real_cmd = [\"p4\"]\n+    # Look for the P4 binary\n+    if (platform.system() == \"Windows\"):\n+        real_cmd = [\"p4.exe\"]    \n+    else:\n+        real_cmd = [\"p4\"]\n \n     user = gitConfig(\"git-p4.user\")\n     if len(user) > 0:\n-- \ngitgitgadget\n\n"},{"id":"387532","messageId":"0ef2f56b04803cad2e60bf881e86d8bdd69463a6.1575498577.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 02/11] git-p4: change the expansion test from basestring to list","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:28Z","receivedAt":"2019-12-04T22:29:45Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nPython 3+ handles strings differently than Python 2.7.  Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n\nChange references to basestring in the isinstance tests to use list instead. This prepares the code to remove all references to basestring.\n\nThe original code used basestring in a test to determine if a list or literal string was passed into 9 different functions.  This is used to determine if the shell should be evoked when calling subprocess methods.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit 5b1b1c145479b5d5fd242122737a3134890409e6)\n---\n git-p4.py | 18 +++++++++---------\n 1 file changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex b2ffbc057b..0f27996393 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -109,7 +109,7 @@ def p4_build_cmd(cmd):\n         # Provide a way to not pass this option by setting git-p4.retries to 0\n         real_cmd += [\"-r\", str(retries)]\n \n-    if isinstance(cmd,basestring):\n+    if not isinstance(cmd, list):\n         real_cmd = ' '.join(real_cmd) + ' ' + cmd\n     else:\n         real_cmd += cmd\n@@ -175,7 +175,7 @@ def write_pipe(c, stdin):\n     if verbose:\n         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n \n-    expand = isinstance(c,basestring)\n+    expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdin=subprocess.PIPE, shell=expand)\n     pipe = p.stdin\n     val = pipe.write(stdin)\n@@ -197,7 +197,7 @@ def read_pipe_full(c):\n     if verbose:\n         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n \n-    expand = isinstance(c,basestring)\n+    expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=expand)\n     (out, err) = p.communicate()\n     return (p.returncode, out, err)\n@@ -233,7 +233,7 @@ def read_pipe_lines(c):\n     if verbose:\n         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n \n-    expand = isinstance(c, basestring)\n+    expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, shell=expand)\n     pipe = p.stdout\n     val = pipe.readlines()\n@@ -276,7 +276,7 @@ def p4_has_move_command():\n     return True\n \n def system(cmd, ignore_error=False):\n-    expand = isinstance(cmd,basestring)\n+    expand = not isinstance(cmd, list)\n     if verbose:\n         sys.stderr.write(\"executing %s\\n\" % str(cmd))\n     retcode = subprocess.call(cmd, shell=expand)\n@@ -288,7 +288,7 @@ def system(cmd, ignore_error=False):\n def p4_system(cmd):\n     \"\"\"Specifically invoke p4 as the system command. \"\"\"\n     real_cmd = p4_build_cmd(cmd)\n-    expand = isinstance(real_cmd, basestring)\n+    expand = not isinstance(real_cmd, list)\n     retcode = subprocess.call(real_cmd, shell=expand)\n     if retcode:\n         raise CalledProcessError(retcode, real_cmd)\n@@ -526,7 +526,7 @@ def getP4OpenedType(file):\n # Return the set of all p4 labels\n def getP4Labels(depotPaths):\n     labels = set()\n-    if isinstance(depotPaths,basestring):\n+    if not isinstance(depotPaths, list):\n         depotPaths = [depotPaths]\n \n     for l in p4CmdList([\"labels\"] + [\"%s...\" % p for p in depotPaths]):\n@@ -613,7 +613,7 @@ def isModeExecChanged(src_mode, dst_mode):\n def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n         errors_as_exceptions=False):\n \n-    if isinstance(cmd,basestring):\n+    if not isinstance(cmd, list):\n         cmd = \"-G \" + cmd\n         expand = True\n     else:\n@@ -630,7 +630,7 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n     stdin_file = None\n     if stdin is not None:\n         stdin_file = tempfile.TemporaryFile(prefix='p4-stdin', mode=stdin_mode)\n-        if isinstance(stdin,basestring):\n+        if not isinstance(stdin, list):\n             stdin_file.write(stdin)\n         else:\n             for i in stdin:\n-- \ngitgitgadget\n\n"},{"id":"387533","messageId":"f0e658b984ca009c575368e661016f785922f970.1575498577.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 03/11] git-p4: add new helper functions for python3 conversion","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:29Z","receivedAt":"2019-12-04T22:29:47Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nPython 3+ handles strings differently than Python 2.7.  Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n\nChange the existing unicode test add new support functions for python2-python3 support.\n\nDefine the following variables:\n- isunicode - a boolean variable that states if the version of python natively supports unicode (true) or not (false). This is true for Python3 and false for Python2.\n- unicode - a type alias for the datatype that holds a unicode string.  It is assigned to a str under python 3 and the unicode type for Python2.\n- bytes - a type alias for an array of bytes.  It is assigned the native bytes type for Python3 and str for Python2.\n\nAdd the following new functions:\n\n- as_string(text) - A new function that will convert a byte array to a unicode (UTF-8) string under python 3.  Under python 2, this returns the string unchanged.\n- as_bytes(text) - A new function that will convert a unicode string to a byte array under python 3.  Under python 2, this returns the string unchanged.\n- to_unicode(text) - Converts a text string as Unicode(UTF-8) on both Python2 and Python3.\n\nAdd a new function alias raw_input:\nIf raw_input does not exist (it was renamed to input in python 3) alias input as raw_input.\n\nThe AS_STRING and AS_BYTES functions allow for modifying the code with a minimal amount of impact on Python2 support.  When a string is expected, the as_string() will be used to convert \"cast\" the incoming \"bytes\" to a string type. Conversely as_bytes() will be used to convert a \"string\" to a \"byte array\" type. Since Python2 overloads the datatype 'str' to serve both purposes, the Python2 versions of these function do not change the data, since the str functions as both a byte array and a string.\n\nbasestring is removed since its only references are found in tests that were changed in the previous change list.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit 7921aeb3136b07643c1a503c2d9d8b5ada620356)\n---\n git-p4.py | 70 +++++++++++++++++++++++++++++++++++++++++++++++++++----\n 1 file changed, 66 insertions(+), 4 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 0f27996393..93dfd0920a 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -32,16 +32,78 @@\n     unicode = unicode\n except NameError:\n     # 'unicode' is undefined, must be Python 3\n-    str = str\n+    #\n+    # For Python3 which is natively unicode, we will use \n+    # unicode for internal information but all P4 Data\n+    # will remain in bytes\n+    isunicode = True\n     unicode = str\n     bytes = bytes\n-    basestring = (str,bytes)\n+\n+    def as_string(text):\n+        \"\"\"Return a byte array as a unicode string\"\"\"\n+        if text == None:\n+            return None\n+        if isinstance(text, bytes):\n+            return unicode(text, \"utf-8\")\n+        else:\n+            return text\n+\n+    def as_bytes(text):\n+        \"\"\"Return a Unicode string as a byte array\"\"\"\n+        if text == None:\n+            return None\n+        if isinstance(text, bytes):\n+            return text\n+        else:\n+            return bytes(text, \"utf-8\")\n+\n+    def to_unicode(text):\n+        \"\"\"Return a byte array as a unicode string\"\"\"\n+        return as_string(text)    \n+\n+    def path_as_string(path):\n+        \"\"\" Converts a path to the UTF8 encoded string \"\"\"\n+        if isinstance(path, unicode):\n+            return path\n+        return encodeWithUTF8(path).decode('utf-8')\n+    \n else:\n     # 'unicode' exists, must be Python 2\n-    str = str\n+    #\n+    # We will treat the data as:\n+    #   str   -> str\n+    #   bytes -> str\n+    # So for Python2 these functions are no-ops\n+    # and will leave the data in the ambiguious\n+    # string/bytes state\n+    isunicode = False\n     unicode = unicode\n     bytes = str\n-    basestring = basestring\n+\n+    def as_string(text):\n+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n+        return text\n+\n+    def as_bytes(text):\n+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n+        return text\n+\n+    def to_unicode(text):\n+        \"\"\"Return a string as a unicode string\"\"\"\n+        return text.decode('utf-8')\n+    \n+    def path_as_string(path):\n+        \"\"\" Converts a path to the UTF8 encoded bytes \"\"\"\n+        return encodeWithUTF8(path)\n+\n+\n+ \n+# Check for raw_input support\n+try:\n+    raw_input\n+except NameError:\n+    raw_input = input\n \n try:\n     from subprocess import CalledProcessError\n-- \ngitgitgadget\n\n"},{"id":"387534","messageId":"1bf7b073b047ca7625d0861b160a9602135f7baf.1575498578.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 05/11] git-p4: Add new functions in preparation of usage","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:31Z","receivedAt":"2019-12-04T22:29:49Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThis changelist is an intermediate submission for migrating the P4 support from Python2 to Python3. The code needs access to the encodeWithUTF8() for support of non-UTF8 filenames in the clone class as well as the sync class.\n\nMove the function encodeWithUTF8() from the P4Sync class to a stand-alone function.  This will allow other classes to use this function without instanciating the P4Sync class. Change the self.verbose reference to an optional method parameter. Update the existing references to this function to pass the self.verbose since it is no longer available on \"self\" since the function is no longer contained on the P4Sync class.\n\nModify the functions write_pipe() and p4_write_pipe() to remove the return value.  The return value for both functions is the number of bytes, but the meaning is lost under python3 since the count does not match the number of characters that may have been encoded.  Additionally, the return value was never used, so this is removed to avoid future ambiguity.\n\nAdd a new method gitConfigSet(). This method will set a value in the git configuration cache list.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit affe888f432bb6833df78962e8671fccdf76c47a)\n---\n git-p4.py | 60 ++++++++++++++++++++++++++++++++++++++++---------------\n 1 file changed, 44 insertions(+), 16 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex b283ef1029..2659531c2e 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -237,6 +237,8 @@ def die(msg):\n         sys.exit(1)\n \n def write_pipe(c, stdin):\n+    \"\"\" Executes the command 'c', passing 'stdin' on the standard input\n+    \"\"\"\n     if verbose:\n         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n \n@@ -248,11 +250,12 @@ def write_pipe(c, stdin):\n     if p.wait():\n         die('Command failed: %s' % str(c))\n \n-    return val\n \n def p4_write_pipe(c, stdin):\n+    \"\"\" Runs a P4 command 'c', passing 'stdin' data to P4\n+    \"\"\"\n     real_cmd = p4_build_cmd(c)\n-    return write_pipe(real_cmd, stdin)\n+    write_pipe(real_cmd, stdin)\n \n def read_pipe_full(c):\n     \"\"\" Read output from  command. Returns a tuple\n@@ -653,6 +656,38 @@ def isModeExec(mode):\n     # otherwise False.\n     return mode[-3:] == \"755\"\n \n+def encodeWithUTF8(path, verbose = False):\n+    \"\"\" Ensure that the path is encoded as a UTF-8 string\n+\n+        Returns bytes(P3)/str(P2)\n+    \"\"\"\n+   \n+    if isunicode:\n+        try:\n+            if isinstance(path, unicode):\n+                # It is already unicode, cast it as a bytes\n+                # that is encoded as utf-8.\n+                return path.encode('utf-8', 'strict')\n+            path.decode('ascii', 'strict')\n+        except:\n+            encoding = 'utf8'\n+            if gitConfig('git-p4.pathEncoding'):\n+                encoding = gitConfig('git-p4.pathEncoding')\n+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n+            if verbose:\n+                print('\\nNOTE:Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, to_unicode(path)))\n+    else:    \n+        try:\n+            path.decode('ascii')\n+        except:\n+            encoding = 'utf8'\n+            if gitConfig('git-p4.pathEncoding'):\n+                encoding = gitConfig('git-p4.pathEncoding')\n+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n+            if verbose:\n+                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n+    return path\n+\n class P4Exception(Exception):\n     \"\"\" Base class for exceptions from the p4 client \"\"\"\n     def __init__(self, exit_code):\n@@ -891,6 +926,11 @@ def gitConfigList(key):\n             _gitConfig[key] = []\n     return _gitConfig[key]\n \n+def gitConfigSet(key, value):\n+    \"\"\" Set the git configuration key 'key' to 'value' for this session\n+    \"\"\"\n+    _gitConfig[key] = value\n+\n def p4BranchesInGit(branchesAreInRemotes=True):\n     \"\"\"Find all the branches whose names start with \"p4/\", looking\n        in remotes or heads as specified by the argument.  Return\n@@ -2814,24 +2854,12 @@ def writeToGitStream(self, gitMode, relPath, contents):\n             self.gitStream.write(d)\n         self.gitStream.write('\\n')\n \n-    def encodeWithUTF8(self, path):\n-        try:\n-            path.decode('ascii')\n-        except:\n-            encoding = 'utf8'\n-            if gitConfig('git-p4.pathEncoding'):\n-                encoding = gitConfig('git-p4.pathEncoding')\n-            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n-            if self.verbose:\n-                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n-        return path\n-\n     # output one file from the P4 stream\n     # - helper for streamP4Files\n \n     def streamOneP4File(self, file, contents):\n         relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n-        relPath = self.encodeWithUTF8(relPath)\n+        relPath = encodeWithUTF8(relPath, self.verbose)\n         if verbose:\n             if 'fileSize' in self.stream_file:\n                 size = int(self.stream_file['fileSize'])\n@@ -2914,7 +2942,7 @@ def streamOneP4File(self, file, contents):\n \n     def streamOneP4Deletion(self, file):\n         relPath = self.stripRepoPath(file['path'], self.branchPrefixes)\n-        relPath = self.encodeWithUTF8(relPath)\n+        relPath = encodeWithUTF8(relPath, self.verbose)\n         if verbose:\n             sys.stdout.write(\"delete %s\\n\" % relPath)\n             sys.stdout.flush()\n-- \ngitgitgadget\n\n"},{"id":"387535","messageId":"3c41db3e9157e20aeed41d3eff373183c9834bff.1575498577.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 04/11] git-p4: python3 syntax changes","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:30Z","receivedAt":"2019-12-04T22:29:50Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nPython 3+ handles strings differently than Python 2.7.  Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n\nThere are a number of translations suggested by modernize/futureize that should be taken to fix numerous non-string specific issues.\n\nChange references to the X.next() iterator to the function next(X) which is compatible with both Python2 and Python3.\n\nChange references to X.keys() to list(X.keys()) to return a list that can be iterated in both Python2 and Python3.\n\nAdd the literal text (object) to the end of class definitions to be consistent with Python3 class definition.\n\nChange integer divison to use \"//\" instead of \"/\"  Under Both python2 and python3 // will return a floor()ed result which matches existing functionality.\n\nChange the format string for displaying decimal values from %d to %4.1f% when displaying a progress.  This avoids displaying long repeating decimals in user displayed text.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit bde6b83296aa9b3e7a584c5ce2b571c7287d8f9f)\n---\n git-p4.py | 55 +++++++++++++++++++++++++++++--------------------------\n 1 file changed, 29 insertions(+), 26 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 93dfd0920a..b283ef1029 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -26,6 +26,9 @@\n import zlib\n import ctypes\n import errno\n+import os.path\n+import codecs\n+import io\n \n # support basestring in python3\n try:\n@@ -631,7 +634,7 @@ def parseDiffTreeEntry(entry):\n \n     If the pattern is not matched, None is returned.\"\"\"\n \n-    match = diffTreePattern().next().match(entry)\n+    match = next(diffTreePattern()).match(entry)\n     if match:\n         return {\n             'src_mode': match.group(1),\n@@ -935,7 +938,7 @@ def findUpstreamBranchPoint(head = \"HEAD\"):\n     branches = p4BranchesInGit()\n     # map from depot-path to branch name\n     branchByDepotPath = {}\n-    for branch in branches.keys():\n+    for branch in list(branches.keys()):\n         tip = branches[branch]\n         log = extractLogMessageFromGitCommit(tip)\n         settings = extractSettingsGitLog(log)\n@@ -1129,7 +1132,7 @@ def getClientSpec():\n     client_name = entry[\"Client\"]\n \n     # just the keys that start with \"View\"\n-    view_keys = [ k for k in entry.keys() if k.startswith(\"View\") ]\n+    view_keys = [ k for k in list(entry.keys()) if k.startswith(\"View\") ]\n \n     # hold this new View\n     view = View(client_name)\n@@ -1371,7 +1374,7 @@ def processContent(self, git_mode, relPath, contents):\n         else:\n             return LargeFileSystem.processContent(self, git_mode, relPath, contents)\n \n-class Command:\n+class Command(object):\n     delete_actions = ( \"delete\", \"move/delete\", \"purge\" )\n     add_actions = ( \"add\", \"branch\", \"move/add\" )\n \n@@ -1386,7 +1389,7 @@ def ensure_value(self, attr, value):\n             setattr(self, attr, value)\n         return getattr(self, attr)\n \n-class P4UserMap:\n+class P4UserMap(object):\n     def __init__(self):\n         self.userMapFromPerforceServer = False\n         self.myP4UserId = None\n@@ -1437,7 +1440,7 @@ def getUserMapFromPerforceServer(self):\n                 self.emails[email] = user\n \n         s = ''\n-        for (key, val) in self.users.items():\n+        for (key, val) in list(self.users.items()):\n             s += \"%s\\t%s\\n\" % (key.expandtabs(1), val.expandtabs(1))\n \n         open(self.getUserCacheFilename(), \"wb\").write(s)\n@@ -1788,7 +1791,7 @@ def prepareSubmitTemplate(self, changelist=None):\n                 break\n         if not change_entry:\n             die('Failed to decode output of p4 change -o')\n-        for key, value in change_entry.iteritems():\n+        for key, value in list(change_entry.items()):\n             if key.startswith('File'):\n                 if 'depot-paths' in settings:\n                     if not [p for p in settings['depot-paths']\n@@ -2032,7 +2035,7 @@ def applyCommit(self, id):\n             p4_delete(f)\n \n         # Set/clear executable bits\n-        for f in filesToChangeExecBit.keys():\n+        for f in list(filesToChangeExecBit.keys()):\n             mode = filesToChangeExecBit[f]\n             setP4ExecBit(f, mode)\n \n@@ -2285,7 +2288,7 @@ def run(self, args):\n             self.clientSpecDirs = getClientSpec()\n \n         # Check for the existence of P4 branches\n-        branchesDetected = (len(p4BranchesInGit().keys()) > 1)\n+        branchesDetected = (len(list(p4BranchesInGit().keys())) > 1)\n \n         if self.useClientSpec and not branchesDetected:\n             # all files are relative to the client spec\n@@ -2676,7 +2679,7 @@ def __init__(self):\n         self.knownBranches = {}\n         self.initialParents = {}\n \n-        self.tz = \"%+03d%02d\" % (- time.timezone / 3600, ((- time.timezone % 3600) / 60))\n+        self.tz = \"%+03d%02d\" % (- time.timezone // 3600, ((- time.timezone % 3600) // 60))\n         self.labels = {}\n \n     # Force a checkpoint in fast-import and wait for it to finish\n@@ -2793,7 +2796,7 @@ def splitFilesIntoBranches(self, commit):\n             else:\n                 relPath = self.stripRepoPath(path, self.depotPaths)\n \n-            for branch in self.knownBranches.keys():\n+            for branch in list(self.knownBranches.keys()):\n                 # add a trailing slash so that a commit into qt/4.2foo\n                 # doesn't end up in qt/4.2, e.g.\n                 if p4PathStartsWith(relPath, branch + \"/\"):\n@@ -2834,7 +2837,7 @@ def streamOneP4File(self, file, contents):\n                 size = int(self.stream_file['fileSize'])\n             else:\n                 size = 0 # deleted files don't get a fileSize apparently\n-            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size/1024/1024))\n+            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size//1024//1024))\n             sys.stdout.flush()\n \n         (type_base, type_mods) = split_p4_type(file[\"type\"])\n@@ -2934,7 +2937,7 @@ def streamP4FilesCb(self, marshalled):\n             required_bytes = int((4 * int(self.stream_file[\"fileSize\"])) - calcDiskFree())\n             if required_bytes > 0:\n                 err = 'Not enough space left on %s! Free at least %i MB.' % (\n-                    os.getcwd(), required_bytes/1024/1024\n+                    os.getcwd(), required_bytes//1024//1024\n                 )\n \n         if err:\n@@ -2963,7 +2966,7 @@ def streamP4FilesCb(self, marshalled):\n \n         # pick up the new file information... for the\n         # 'data' field we need to append to our array\n-        for k in marshalled.keys():\n+        for k in list(marshalled.keys()):\n             if k == 'data':\n                 if 'streamContentSize' not in self.stream_file:\n                     self.stream_file['streamContentSize'] = 0\n@@ -2978,8 +2981,8 @@ def streamP4FilesCb(self, marshalled):\n             'depotFile' in self.stream_file):\n             size = int(self.stream_file[\"fileSize\"])\n             if size > 0:\n-                progress = 100*self.stream_file['streamContentSize']/size\n-                sys.stdout.write('\\r%s %d%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size/1024/1024)))\n+                progress = 100.0*self.stream_file['streamContentSize']/size\n+                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size//1024//1024)))\n                 sys.stdout.flush()\n \n         self.stream_have_file_info = True\n@@ -3060,7 +3063,7 @@ def streamTag(self, gitStream, labelName, labelDetails, commit, epoch):\n \n         gitStream.write(\"tagger %s\\n\" % tagger)\n \n-        print(\"labelDetails=\",labelDetails)\n+        print((\"labelDetails=\",labelDetails))\n         if 'Description' in labelDetails:\n             description = labelDetails['Description']\n         else:\n@@ -3199,7 +3202,7 @@ def getLabels(self):\n             self.labels[newestChange] = [output, revisions]\n \n         if self.verbose:\n-            print(\"Label changes: %s\" % self.labels.keys())\n+            print(\"Label changes: %s\" % list(self.labels.keys()))\n \n     # Import p4 labels as git tags. A direct mapping does not\n     # exist, so assume that if all the files are at the same revision\n@@ -3342,7 +3345,7 @@ def getBranchMapping(self):\n \n     def getBranchMappingFromGitBranches(self):\n         branches = p4BranchesInGit(self.importIntoRemotes)\n-        for branch in branches.keys():\n+        for branch in list(branches.keys()):\n             if branch == \"master\":\n                 branch = \"main\"\n             else:\n@@ -3454,14 +3457,14 @@ def importChanges(self, changes, origin_revision=0):\n             self.updateOptionDict(description)\n \n             if not self.silent:\n-                sys.stdout.write(\"\\rImporting revision %s (%s%%)\" % (change, cnt * 100 / len(changes)))\n+                sys.stdout.write(\"\\rImporting revision %s (%4.1f%%)\" % (change, cnt * 100 / len(changes)))\n                 sys.stdout.flush()\n             cnt = cnt + 1\n \n             try:\n                 if self.detectBranches:\n                     branches = self.splitFilesIntoBranches(description)\n-                    for branch in branches.keys():\n+                    for branch in list(branches.keys()):\n                         ## HACK  --hwn\n                         branchPrefix = self.depotPaths[0] + branch + \"/\"\n                         self.branchPrefixes = [ branchPrefix ]\n@@ -3650,13 +3653,13 @@ def run(self, args):\n                 if short in branches:\n                     self.p4BranchesInGit = [ short ]\n             else:\n-                self.p4BranchesInGit = branches.keys()\n+                self.p4BranchesInGit = list(branches.keys())\n \n             if len(self.p4BranchesInGit) > 1:\n                 if not self.silent:\n                     print(\"Importing from/into multiple branches\")\n                 self.detectBranches = True\n-                for branch in branches.keys():\n+                for branch in list(branches.keys()):\n                     self.initialParents[self.refPrefix + branch] = \\\n                         branches[branch]\n \n@@ -4040,7 +4043,7 @@ def findLastP4Revision(self, starting_point):\n             to find the P4 commit we are based on, and the depot-paths.\n         \"\"\"\n \n-        for parent in (range(65535)):\n+        for parent in (list(range(65535))):\n             log = extractLogMessageFromGitCommit(\"{0}^{1}\".format(starting_point, parent))\n             settings = extractSettingsGitLog(log)\n             if 'change' in settings:\n@@ -4179,7 +4182,7 @@ def printUsage(commands):\n \n def main():\n     if len(sys.argv[1:]) == 0:\n-        printUsage(commands.keys())\n+        printUsage(list(commands.keys()))\n         sys.exit(2)\n \n     cmdName = sys.argv[1]\n@@ -4189,7 +4192,7 @@ def main():\n     except KeyError:\n         print(\"unknown command %s\" % cmdName)\n         print(\"\")\n-        printUsage(commands.keys())\n+        printUsage(list(commands.keys()))\n         sys.exit(2)\n \n     options = cmd.options\n-- \ngitgitgadget\n\n"},{"id":"387536","messageId":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v3.git.1575313336.gitgitgadget@gmail.com","subject":"[PATCH v4 00/11] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:26Z","receivedAt":"2019-12-04T22:29:50Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"Issue: The current git-p4.py script does not work with python3.\n\nI have attempted to use the P4 integration built into GIT and I was unable\nto get the program to run because I have Python 3.8 installed on my\ncomputer. I was able to get the program to run when I downgraded my python\nto version 2.7. However, python 2 is reaching its end of life.\n\nSubmission: I am submitting a patch for the git-p4.py script that partially \nsupports python 3.8. This code was able to pass the basic tests (t9800) when\nrun against Python3. This provides basic functionality. \n\nIn an attempt to pass the t9822 P4 path-encoding test, a new parameter for\ngit P4 Clone was introduced. \n\n--encoding Format-identifier\n\nThis will create the GIT repository following the current functionality;\nhowever, before importing the files from P4, it will set the\ngit-p4.pathEncoding option so any files or paths that are encoded with\nnon-ASCII/non-UTF-8 formats will import correctly.\n\nTechnical details: The script was updated by futurize (\nhttps://python-future.org/futurize.html) to support Py2/Py3 syntax. The few\nreferences to classes in future were reworked so that future would not be\nrequired. The existing code test for Unicode support was extended to\nnormalize the classes “unicode” and “bytes” to across platforms:\n\n * ‘unicode’ is an alias for ‘str’ in Py3 and is the unicode class in Py2.\n * ‘bytes’ is bytes in Py3 and an alias for ‘str’ in Py2.\n\nNew coercion methods were written for both Python2 and Python3:\n\n * as_string(text) – In Python3, this encodes a bytes object as a UTF-8\n   encoded Unicode string. \n * as_bytes(text) – In Python3, this decodes a Unicode string to an array of\n   bytes.\n\nIn Python2, these functions do not change the data since a ‘str’ object\nfunction in both roles as strings and byte arrays. This reduces the\npotential impact on backward compatibility with Python 2.\n\n * to_unicode(text) – ensures that the supplied data is encoded as a UTF-8\n   string. This function will encode data in both Python2 and Python3. * \n      path_as_string(path) – This function is an extension function that\n      honors the option “git-p4.pathEncoding” to convert a set of bytes or\n      characters to UTF-8. If the str/bytes cannot decode as ASCII, it will\n      use the encodeWithUTF8() method to convert the custom encoded bytes to\n      Unicode in UTF-8.\n   \n   \n\nGenerally speaking, information in the script is converted to Unicode as\nearly as possible and converted back to a byte array just before passing to\nexternal programs or files. The exception to this rule is P4 Repository file\npaths.\n\nPaths are not converted but left as “bytes” so the original file path\nencoding can be preserved. This formatting is required for commands that\ninteract with the P4 file path. When the file path is used by GIT, it is\nconverted with encodeWithUTF8().\n\nSigned-off-by: Ben Keene seraphire@gmail.com [seraphire@gmail.com]\n\nBen Keene (11):\n  git-p4: select p4 binary by operating-system\n  git-p4: change the expansion test from basestring to list\n  git-p4: add new helper functions for python3 conversion\n  git-p4: python3 syntax changes\n  git-p4: Add new functions in preparation of usage\n  git-p4: Fix assumed path separators to be more Windows friendly\n  git-p4: Add a helper class for stream writing\n  git-p4: p4CmdList  - support Unicode encoding\n  git-p4: Add usability enhancements\n  git-p4: Support python3 for basic P4 clone, sync, and submit\n  git-p4: Added --encoding parameter to p4 clone\n\n Documentation/git-p4.txt        |   5 +\n git-p4.py                       | 690 ++++++++++++++++++++++++--------\n t/t9822-git-p4-path-encoding.sh | 101 +++++\n 3 files changed, 629 insertions(+), 167 deletions(-)\n\n\nbase-commit: 228f53135a4a41a37b6be8e4d6e2b6153db4a8ed\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-463%2Fseraphire%2Fseraphire%2Fp4-python3-unicode-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-463/seraphire/seraphire/p4-python3-unicode-v4\nPull-Request: https://github.com/gitgitgadget/git/pull/463\n\nRange-diff vs v3:\n\n  -:  ---------- >  1:  4012426993 git-p4: select p4 binary by operating-system\n  -:  ---------- >  2:  0ef2f56b04 git-p4: change the expansion test from basestring to list\n  -:  ---------- >  3:  f0e658b984 git-p4: add new helper functions for python3 conversion\n  -:  ---------- >  4:  3c41db3e91 git-p4: python3 syntax changes\n  -:  ---------- >  5:  1bf7b073b0 git-p4: Add new functions in preparation of usage\n  -:  ---------- >  6:  8f5752c127 git-p4: Fix assumed path separators to be more Windows friendly\n  -:  ---------- >  7:  10dc059444 git-p4: Add a helper class for stream writing\n  -:  ---------- >  8:  e1a424a955 git-p4: p4CmdList  - support Unicode encoding\n  -:  ---------- >  9:  4fc49313f0 git-p4: Add usability enhancements\n  1:  02b3843e9f ! 10:  04a0aedbaa Python3 support for t9800 tests. Basic P4/Python3 support\n     @@ -1,159 +1,60 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    Python3 support for t9800 tests. Basic P4/Python3 support\n     +    git-p4: Support python3 for basic P4 clone, sync, and submit\n     +\n     +    Issue: Python 3 is still not properly supported for any use with the git-p4 python code.\n     +    Warning - this is a very large atomic commit.  The commit text is also very large.\n     +\n     +    Change the code such that, with the exception of P4 depot paths and depot files, all text read by git-p4 is cast as a string as soon as possible and converted back to bytes as late as possible, following Python2 to Python3 conversion best practices.\n     +\n     +    Important: Do not cast the bytes that contain the p4 depot path or p4 depot file name.  These should be left as bytes until used.\n     +\n     +    These two values should not be converted because the encoding of these values is unknown.  git-p4 supports a configuration value git-p4.pathEncoding that is used by the encodeWithUTF8()  to determine what a UTF8 version of the path and filename should be.  However, since depot path and depot filename need to be sent to P4 in their original encoding, they will be left as byte streams until they are actually used:\n     +\n     +    * When sent to P4, the bytes are literally passed to the p4 command\n     +    * When displayed in text for the user, they should be passed through the path_as_string() function\n     +    * When used by GIT they should be passed through the encodeWithUTF8() function\n     +\n     +    Change all the rest of system calls to cast output (stdin) as_bytes() and input (stdout) as_string().  This retains existing Python 2 support, and adds python 3 support for these functions:\n     +    * read_pipe_full\n     +    * read_pipe_lines\n     +    * p4_has_move_command (used internally)\n     +    * gitConfig\n     +    * branch_exists\n     +    * GitLFS.generatePointer\n     +    * applyCommit - template must be read and written to the temporary file as_bytes() since it is created in memory as a string.\n     +    * streamOneP4File(file, contents) - wrap calls to the depotFile in path_as_string() for display. The file contents must be retained as bytes, so update the RCS changes to be forced to bytes.\n     +    * streamP4Files\n     +    * importHeadRevision(revision) - encode the depotPaths for display separate from the text for processing.\n     +\n     +    Py23File usage -\n     +    Change the P4Sync.OpenStreams() function to cast the gitOutput, gitStream, and gitError streams as Py23File() wrapper classes.  This facilitates taking strings in both python 2 and python 3 and casting them to bytes in the wrapper class instead of having to modify each method. Since the fast-import command also expects a raw byte stream for file content, add a new stream handle - gitStreamBytes which is an unwrapped verison of gitStream.\n     +\n     +    Literal text -\n     +    Depending on context, most literal text does not need casting to unicode or bytes as the text is Python dependent - In python 2, the string is implied as 'str' and python 3 the string is implied as 'unicode'. Under these conditions, they match the rest of the operating text, following best practices.  However, when a literal string is used in functions that are dealing with the raw input from and raw ouput to files streams, literal bytes may be required. Additionally, functions that are dealing with P4 depot paths or P4 depot file names are also dealing with bytes and will require the same casting as bytes.  The following functions cast text as byte strings:\n     +    * wildcard_decode(path) - the path parameter is a P4 depot and is bytes. Cast all the literals to bytes.\n     +    * wildcard_encode(path) - the path parameter is a P4 depot and is bytes. Cast all the literals to bytes.\n     +    * streamP4FilesCb(marshalled) - the marshalled data is in bytes. Cast the literals as bytes. When using this data to manipulate self.stream_file, encode all the marshalled data except for the 'depotFile' name.\n     +    * streamP4Files\n     +\n     +    Special behavior:\n     +    * p4_describe - encoding is disabled for the depotFile(x) and path elements since these are depot path and depo filenames.\n     +    * p4PathStartsWith(path, prefix) - Since P4 depot paths can contain non-UTF-8 encoded strings, change this method to compare paths while supporting the optional encoding.\n     +       - First, perform a byte-to-byte check to see if the path and prefix are both identical text.  There is no need to perform encoding conversions if the text is identical.\n     +       - If the byte check fails, pass both the path and prefix through encodeWithUTF8() to ensure both paths are using the same encoding. Then perform the test as originally written.\n     +    * patchRCSKeywords(file, pattern) - the parameters of file and pattern are both strings. However this function changes the contents of the file itentified by name \"file\". Treat the content of this file as binary to ensure that python does not accidently change the original encoding. The regular expression is cast as_bytes() and run against the file as_bytes(). The P4 keywords are ASCII strings and cannot span lines so iterating over each line of the file is acceptable.\n     +    * writeToGitStream(gitMode, relPath, contents) - Since 'contents' is already bytes data, instead of using the self.gitStream, use the new self.gitStreamBytes - the unwrapped gitStream that does not cast as_bytes() the binary data.\n     +    * commit(details, files, branch, parent = \"\", allow_empty=False) - Changed the encoding for the commit message to the preferred format for fast-import. The number of bytes is sent in the data block instead of using the EOT marker.\n     +    * Change the code for handling the user cache to use binary files. Cast text as_bytes() when writing to the cache and as_string() when reading from the cache.  This makes the reading and writing of the cache determinstic in it's encoding. Unlike file paths, P4 encodes the user names in UTF-8 encoding so no additional string encoding is required.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     +    (cherry picked from commit 65ff0c74ebe62a200b4385ecfd4aa618ce091f48)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n       +++ b/git-p4.py\n      @@\n     - import zlib\n     - import ctypes\n     - import errno\n     -+import os.path\n     -+import codecs\n     -+import io\n     - \n     - # support basestring in python3\n     - try:\n     -     unicode = unicode\n     - except NameError:\n     -     # 'unicode' is undefined, must be Python 3\n     --    str = str\n     -+    #\n     -+    # For Python3 which is natively unicode, we will use \n     -+    # unicode for internal information but all P4 Data\n     -+    # will remain in bytes\n     -+    isunicode = True\n     -     unicode = str\n     -     bytes = bytes\n     --    basestring = (str,bytes)\n     -+\n     -+    def as_string(text):\n     -+        \"\"\"Return a byte array as a unicode string\"\"\"\n     -+        if text == None:\n     -+            return None\n     -+        if isinstance(text, bytes):\n     -+            return unicode(text, \"utf-8\")\n     -+        else:\n     -+            return text\n     -+\n     -+    def as_bytes(text):\n     -+        \"\"\"Return a Unicode string as a byte array\"\"\"\n     -+        if text == None:\n     -+            return None\n     -+        if isinstance(text, bytes):\n     -+            return text\n     -+        else:\n     -+            return bytes(text, \"utf-8\")\n     -+\n     -+    def to_unicode(text):\n     -+        \"\"\"Return a byte array as a unicode string\"\"\"\n     -+        return as_string(text)    \n     -+\n     -+    def path_as_string(path):\n     -+        \"\"\" Converts a path to the UTF8 encoded string \"\"\"\n     -+        if isinstance(path, unicode):\n     -+            return path\n     -+        return encodeWithUTF8(path).decode('utf-8')\n     -+    \n     - else:\n     -     # 'unicode' exists, must be Python 2\n     --    str = str\n     -+    #\n     -+    # We will treat the data as:\n     -+    #   str   -> str\n     -+    #   bytes -> str\n     -+    # So for Python2 these functions are no-ops\n     -+    # and will leave the data in the ambiguious\n     -+    # string/bytes state\n     -+    isunicode = False\n     -     unicode = unicode\n     -     bytes = str\n     --    basestring = basestring\n     -+\n     -+    def as_string(text):\n     -+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n     -+        return text\n     -+\n     -+    def as_bytes(text):\n     -+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n     -+        return text\n     -+\n     -+    def to_unicode(text):\n     -+        \"\"\"Return a string as a unicode string\"\"\"\n     -+        return text.decode('utf-8')\n     -+    \n     -+    def path_as_string(path):\n     -+        \"\"\" Converts a path to the UTF8 encoded bytes \"\"\"\n     -+        return encodeWithUTF8(path)\n     -+\n     -+\n     -+ \n     -+# Check for raw_input support\n     -+try:\n     -+    raw_input\n     -+except NameError:\n     -+    raw_input = input\n     - \n     - try:\n     -     from subprocess import CalledProcessError\n     -@@\n     -     location. It means that hooking into the environment, or other configuration\n     -     can be done more easily.\n     -     \"\"\"\n     --    real_cmd = [\"p4\"]\n     -+    # Look for the P4 binary\n     -+    if (platform.system() == \"Windows\"):\n     -+        real_cmd = [\"p4.exe\"]    \n     -+    else:\n     -+        real_cmd = [\"p4\"]\n     - \n     -     user = gitConfig(\"git-p4.user\")\n     -     if len(user) > 0:\n     -@@\n     -         # Provide a way to not pass this option by setting git-p4.retries to 0\n     -         real_cmd += [\"-r\", str(retries)]\n     - \n     --    if isinstance(cmd,basestring):\n     -+    if not isinstance(cmd, list):\n     -         real_cmd = ' '.join(real_cmd) + ' ' + cmd\n     -     else:\n     -         real_cmd += cmd\n     -@@\n     -         sys.exit(1)\n     - \n     - def write_pipe(c, stdin):\n     -+    \"\"\"Executes the command 'c', passing 'stdin' on the standard input\"\"\"\n     -     if verbose:\n     -         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n     - \n     --    expand = isinstance(c,basestring)\n     -+    expand = not isinstance(c, list)\n     -     p = subprocess.Popen(c, stdin=subprocess.PIPE, shell=expand)\n     -     pipe = p.stdin\n     -     val = pipe.write(stdin)\n     -@@\n     -     if p.wait():\n     -         die('Command failed: %s' % str(c))\n     - \n     --    return val\n     - \n     - def p4_write_pipe(c, stdin):\n     -+    \"\"\" Runs a P4 command 'c', passing 'stdin' data to P4\"\"\"\n     -     real_cmd = p4_build_cmd(c)\n     --    return write_pipe(real_cmd, stdin)\n     -+    write_pipe(real_cmd, stdin)\n     - \n     - def read_pipe_full(c):\n     -     \"\"\" Read output from  command. Returns a tuple\n     -@@\n     -     if verbose:\n     -         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n     - \n     --    expand = isinstance(c,basestring)\n     -+    expand = not isinstance(c, list)\n     +     expand = not isinstance(c, list)\n           p = subprocess.Popen(c, stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=expand)\n           (out, err) = p.communicate()\n      +    out = as_string(out)\n     @@ -179,10 +80,7 @@\n           if verbose:\n               sys.stderr.write('Reading pipe: %s\\n' % str(c))\n       \n     --    expand = isinstance(c, basestring)\n     -+    expand = not isinstance(c, list)\n     -     p = subprocess.Popen(c, stdout=subprocess.PIPE, shell=expand)\n     -     pipe = p.stdout\n     +@@\n           val = pipe.readlines()\n           if pipe.close() or p.wait():\n               die('Command failed: %s' % str(c))\n     @@ -203,28 +101,6 @@\n           # return code will be 1 in either case\n           if err.find(\"Invalid option\") >= 0:\n               return False\n     -@@\n     -     return True\n     - \n     - def system(cmd, ignore_error=False):\n     --    expand = isinstance(cmd,basestring)\n     -+    expand = not isinstance(cmd, list)\n     -     if verbose:\n     -         sys.stderr.write(\"executing %s\\n\" % str(cmd))\n     -     retcode = subprocess.call(cmd, shell=expand)\n     -@@\n     -     return retcode\n     - \n     - def p4_system(cmd):\n     --    \"\"\"Specifically invoke p4 as the system command. \"\"\"\n     -+    \"\"\" Specifically invoke p4 as the system command. \n     -+    \"\"\"\n     -     real_cmd = p4_build_cmd(cmd)\n     --    expand = isinstance(real_cmd, basestring)\n     -+    expand = not isinstance(real_cmd, list)\n     -     retcode = subprocess.call(real_cmd, shell=expand)\n     -     if retcode:\n     -         raise CalledProcessError(retcode, real_cmd)\n      @@\n           return int(results[0]['change'])\n       \n     @@ -234,7 +110,7 @@\n      -       results.\"\"\"\n      +    \"\"\" Returns information about the requested P4 change list.\n      +\n     -+        Data returns is not string encoded (returned as bytes)\n     ++        Data returned is not string encoded (returned as bytes)\n      +    \"\"\"\n      +    # Make sure it returns a valid result by checking for\n      +    #   the presence of field \"time\".  Return a dict of the\n     @@ -261,218 +137,29 @@\n           if \"time\" not in d:\n               die(\"p4 describe -s %d returned no \\\"time\\\": %s\" % (change, str(d)))\n       \n     -+    # Convert depotFile(X) to be UTF-8 encoded, as this is what GIT\n     -+    # requires. This will also allow us to encode the rest of the text\n     -+    # at the same time to simplify textual processing later.\n     ++    # Do not convert 'depotFile(X)' or 'path' to be UTF-8 encoded, however \n     ++    # cast as_string() the rest of the text. \n      +    keys=d.keys()\n      +    for key in keys:\n      +        if key.startswith('depotFile'):\n     -+            d[key]=d[key] #DepotPath(d[key])\n     ++            d[key]=d[key] \n      +        elif key == 'path':\n     -+            d[key]=d[key] #DepotPath(d[key])\n     ++            d[key]=d[key] \n      +        else:\n      +            d[key] = as_string(d[key])\n      +\n           return d\n       \n     --#\n     --# Canonicalize the p4 type and return a tuple of the\n     --# base type, plus any modifiers.  See \"p4 help filetypes\"\n     --# for a list and explanation.\n     --#\n     - def split_p4_type(p4type):\n     --\n     -+    \"\"\" Canonicalize the p4 type and return a tuple of the\n     -+        base type, plus any modifiers.  See \"p4 help filetypes\"\n     -+        for a list and explanation.\n     -+    \"\"\"\n     -     p4_filetypes_historical = {\n     -         \"ctempobj\": \"binary+Sw\",\n     -         \"ctext\": \"text+C\",\n     -@@\n     -         mods = s[1]\n     -     return (base, mods)\n     - \n     --#\n     --# return the raw p4 type of a file (text, text+ko, etc)\n     --#\n     - def p4_type(f):\n     -+    \"\"\" return the raw p4 type of a file (text, text+ko, etc)\n     -+    \"\"\"\n     -     results = p4CmdList([\"fstat\", \"-T\", \"headType\", wildcard_encode(f)])\n     -     return results[0]['headType']\n     - \n     --#\n     --# Given a type base and modifier, return a regexp matching\n     --# the keywords that can be expanded in the file\n     --#\n     - def p4_keywords_regexp_for_type(base, type_mods):\n     -+    \"\"\" Given a type base and modifier, return a regexp matching\n     -+        the keywords that can be expanded in the file\n     -+    \"\"\"\n     -     if base in (\"text\", \"unicode\", \"binary\"):\n     -         kwords = None\n     -         if \"ko\" in type_mods:\n     -@@\n     -     else:\n     -         return None\n     - \n     --#\n     --# Given a file, return a regexp matching the possible\n     --# RCS keywords that will be expanded, or None for files\n     --# with kw expansion turned off.\n     --#\n     - def p4_keywords_regexp_for_file(file):\n     -+    \"\"\" Given a file, return a regexp matching the possible\n     -+        RCS keywords that will be expanded, or None for files\n     -+        with kw expansion turned off.\n     -+    \"\"\"\n     -     if not os.path.exists(file):\n     -         return None\n     -     else:\n     -@@\n     - # Return the set of all p4 labels\n     - def getP4Labels(depotPaths):\n     -     labels = set()\n     --    if isinstance(depotPaths,basestring):\n     -+    if not isinstance(depotPaths, list):\n     -         depotPaths = [depotPaths]\n     - \n     -     for l in p4CmdList([\"labels\"] + [\"%s...\" % p for p in depotPaths]):\n     -@@\n     - \n     -     return labels\n     - \n     --# Return the set of all git tags\n     - def getGitTags():\n     -+    \"\"\"Return the set of all git tags\"\"\"\n     -     gitTags = set()\n     -     for line in read_pipe_lines([\"git\", \"tag\"]):\n     -         tag = line.strip()\n     -@@\n     - \n     -     If the pattern is not matched, None is returned.\"\"\"\n     - \n     --    match = diffTreePattern().next().match(entry)\n     -+    match = next(diffTreePattern()).match(entry)\n     -     if match:\n     -         return {\n     -             'src_mode': match.group(1),\n     -@@\n     -     # otherwise False.\n     -     return mode[-3:] == \"755\"\n     - \n     -+def encodeWithUTF8(path, verbose = False):\n     -+    \"\"\" Ensure that the path is encoded as a UTF-8 string\n     -+\n     -+        Returns bytes(P3)/str(P2)\n     -+    \"\"\"\n     -+   \n     -+    if isunicode:\n     -+        try:\n     -+            if isinstance(path, unicode):\n     -+                # It is already unicode, cast it as a bytes\n     -+                # that is encoded as utf-8.\n     -+                return path.encode('utf-8', 'strict')\n     -+            path.decode('ascii', 'strict')\n     -+        except:\n     -+            encoding = 'utf8'\n     -+            if gitConfig('git-p4.pathEncoding'):\n     -+                encoding = gitConfig('git-p4.pathEncoding')\n     -+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n     -+            if verbose:\n     -+                print('\\nNOTE:Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, to_unicode(path)))\n     -+    else:    \n     -+        try:\n     -+            path.decode('ascii')\n     -+        except:\n     -+            encoding = 'utf8'\n     -+            if gitConfig('git-p4.pathEncoding'):\n     -+                encoding = gitConfig('git-p4.pathEncoding')\n     -+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n     -+            if verbose:\n     -+                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n     -+    return path\n     -+\n     - class P4Exception(Exception):\n     -     \"\"\" Base class for exceptions from the p4 client \"\"\"\n     -     def __init__(self, exit_code):\n     -@@\n     -     return isModeExec(src_mode) != isModeExec(dst_mode)\n     - \n     - def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n     --        errors_as_exceptions=False):\n     -+        errors_as_exceptions=False, encode_data=True):\n     -+    \"\"\" Executes a P4 command:  'cmd' optionally passing 'stdin' to the command's\n     -+        standard input via a temporary file with 'stdin_mode' mode.\n     -+\n     -+        Output from the command is optionally passed to the callback function 'cb'.\n     -+        If 'cb' is None, the response from the command is parsed into a list\n     -+        of resulting dictionaries. (For each block read from the process pipe.)\n     -+\n     -+        If 'skip_info' is true, information in a block read that has a code type of\n     -+        'info' will be skipped.\n     - \n     --    if isinstance(cmd,basestring):\n     -+        If 'errors_as_exceptions' is set to true (the default is false) the error\n     -+        code returned from the execution will generate an exception.\n     -+\n     -+        If 'encode_data' is set to true (the default) the data that is returned \n     -+        by this function will be passed through the \"as_string\" function.\n     -+    \"\"\"\n     -+\n     -+    if not isinstance(cmd, list):\n     -         cmd = \"-G \" + cmd\n     -         expand = True\n     -     else:\n     -@@\n     -     stdin_file = None\n     -     if stdin is not None:\n     -         stdin_file = tempfile.TemporaryFile(prefix='p4-stdin', mode=stdin_mode)\n     --        if isinstance(stdin,basestring):\n     -+        if not isinstance(stdin, list):\n     -             stdin_file.write(stdin)\n     -         else:\n     -             for i in stdin:\n     --                stdin_file.write(i + '\\n')\n     -+                stdin_file.write(as_bytes(i) + b'\\n')\n     -         stdin_file.flush()\n     -         stdin_file.seek(0)\n     - \n     -@@\n     -         while True:\n     -             entry = marshal.load(p4.stdout)\n     -             if skip_info:\n     --                if 'code' in entry and entry['code'] == 'info':\n     -+                if b'code' in entry and entry[b'code'] == b'info':\n     -                     continue\n     -             if cb is not None:\n     -                 cb(entry)\n     -             else:\n     --                result.append(entry)\n     -+                out = {}\n     -+                for key, value in entry.items():\n     -+                    out[as_string(key)] = (as_string(value) if encode_data else value)\n     -+                result.append(out)\n     -     except EOFError:\n     -         pass\n     -     exitCode = p4.wait()\n     + #\n      @@\n           return result\n       \n       def p4Cmd(cmd):\n     -+    \"\"\" Executes a P4 command an returns the results in a dictionary\"\"\"\n     ++    \"\"\" Executes a P4 command and returns the results in a dictionary\n     ++    \"\"\"\n           list = p4CmdList(cmd)\n           result = {}\n           for entry in list:\n     -@@\n     -     return values\n     - \n     - def gitBranchExists(branch):\n     -+    \"\"\"Checks to see if a given branch exists in the git repo\"\"\"\n     -     proc = subprocess.Popen([\"git\", \"rev-parse\", branch],\n     -                             stderr=subprocess.PIPE, stdout=subprocess.PIPE);\n     -     return proc.wait() == 0;\n      @@\n       _gitConfig = {}\n       \n     @@ -490,29 +177,6 @@\n           return _gitConfig[key]\n       \n       def gitConfigBool(key):\n     --    \"\"\"Return a bool, using git config --bool.  It is True only if the\n     --       variable is set to true, and False if set to false or not present\n     --       in the config.\"\"\"\n     --\n     -+    \"\"\" Return a bool, using git config --bool.  It is True only if the\n     -+        variable is set to true, and False if set to false or not present\n     -+        in the config.\n     -+    \"\"\"\n     -     if key not in _gitConfig:\n     -         _gitConfig[key] = gitConfig(key, '--bool') == \"true\"\n     -     return _gitConfig[key]\n     -@@\n     -             _gitConfig[key] = []\n     -     return _gitConfig[key]\n     - \n     -+def gitConfigSet(key, value):\n     -+    \"\"\" Set the git configuration key 'key' to 'value' for this session\n     -+    \"\"\"\n     -+    _gitConfig[key] = value\n     -+\n     - def p4BranchesInGit(branchesAreInRemotes=True):\n     -     \"\"\"Find all the branches whose names start with \"p4/\", looking\n     -        in remotes or heads as specified by the argument.  Return\n      @@\n           cmd = [ \"git\", \"rev-parse\", \"--symbolic\", \"--verify\", branch ]\n           p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     @@ -521,34 +185,6 @@\n           if p.returncode:\n               return False\n           # expect exactly one line of output: the branch name\n     -@@\n     -     branches = p4BranchesInGit()\n     -     # map from depot-path to branch name\n     -     branchByDepotPath = {}\n     --    for branch in branches.keys():\n     -+    for branch in list(branches.keys()):\n     -         tip = branches[branch]\n     -         log = extractLogMessageFromGitCommit(tip)\n     -         settings = extractSettingsGitLog(log)\n     -@@\n     -             system(\"git update-ref %s %s\" % (remoteHead, originHead))\n     - \n     - def originP4BranchesExist():\n     --        return gitBranchExists(\"origin\") or gitBranchExists(\"origin/p4\") or gitBranchExists(\"origin/p4/master\")\n     -+    \"\"\"Checks if origin/p4/master exists\"\"\"\n     -+    return gitBranchExists(\"origin\") or gitBranchExists(\"origin/p4\") or gitBranchExists(\"origin/p4/master\")\n     - \n     - \n     - def p4ParseNumericChangeRange(parts):\n     -@@\n     -     changes = sorted(changes)\n     -     return changes\n     - \n     --def p4PathStartsWith(path, prefix):\n     -+def p4PathStartsWith(path, prefix, verbose = False):\n     -     # This method tries to remedy a potential mixed-case issue:\n     -     #\n     -     # If UserA adds  //depot/DirA/file1\n      @@\n           #\n           # we may or may not have a problem. If you have core.ignorecase=true,\n     @@ -574,15 +210,6 @@\n       \n       def getClientSpec():\n           \"\"\"Look at the p4 client spec, create a View() object that contains\n     -@@\n     -     client_name = entry[\"Client\"]\n     - \n     -     # just the keys that start with \"View\"\n     --    view_keys = [ k for k in entry.keys() if k.startswith(\"View\") ]\n     -+    view_keys = [ k for k in list(entry.keys()) if k.startswith(\"View\") ]\n     - \n     -     # hold this new View\n     -     view = View(client_name)\n      @@\n           # Cannot have * in a filename in windows; untested as to\n           # what p4 would do in such a case.\n     @@ -626,45 +253,16 @@\n                   os.remove(contentFile)\n                   die('git-lfs pointer command failed. Did you install the extension?')\n      @@\n     -         else:\n     -             return LargeFileSystem.processContent(self, git_mode, relPath, contents)\n     - \n     --class Command:\n     -+class Command(object):\n     -     delete_actions = ( \"delete\", \"move/delete\", \"purge\" )\n     -     add_actions = ( \"add\", \"branch\", \"move/add\" )\n     - \n     -@@\n     -             setattr(self, attr, value)\n     -         return getattr(self, attr)\n     - \n     --class P4UserMap:\n     -+class P4UserMap(object):\n     -     def __init__(self):\n     -         self.userMapFromPerforceServer = False\n     -         self.myP4UserId = None\n     -@@\n     -             return True\n     - \n     -     def getUserCacheFilename(self):\n     -+        \"\"\" Returns the filename of the username cache \"\"\"\n     -         home = os.environ.get(\"HOME\", os.environ.get(\"USERPROFILE\"))\n     --        return home + \"/.gitp4-usercache.txt\"\n     -+        return os.path.join(home, \".gitp4-usercache.txt\")\n     +         return os.path.join(home, \".gitp4-usercache.txt\")\n       \n           def getUserMapFromPerforceServer(self):\n      +        \"\"\" Creates the usercache from the data in P4.\n      +        \"\"\"\n     -+        \n               if self.userMapFromPerforceServer:\n                   return\n               self.users = {}\n      @@\n     -                 self.emails[email] = user\n     - \n     -         s = ''\n     --        for (key, val) in self.users.items():\n     -+        for (key, val) in list(self.users.items()):\n     +         for (key, val) in list(self.users.items()):\n                   s += \"%s\\t%s\\n\" % (key.expandtabs(1), val.expandtabs(1))\n       \n      -        open(self.getUserCacheFilename(), \"wb\").write(s)\n     @@ -674,7 +272,8 @@\n               self.userMapFromPerforceServer = True\n       \n           def loadUserMapFromCache(self):\n     -+        \"\"\" Reads the P4 username to git email map \"\"\"\n     ++        \"\"\" Reads the P4 username to git email map \n     ++        \"\"\"\n               self.users = {}\n               self.userMapFromPerforceServer = False\n               try:\n     @@ -721,80 +320,6 @@\n                   # cleanup our temporary file\n                   os.unlink(outFileName)\n                   print(\"Failed to strip RCS keywords in %s\" % file)\n     -@@\n     -                 break\n     -         if not change_entry:\n     -             die('Failed to decode output of p4 change -o')\n     --        for key, value in change_entry.iteritems():\n     -+        for key, value in list(change_entry.items()):\n     -             if key.startswith('File'):\n     -                 if 'depot-paths' in settings:\n     -                     if not [p for p in settings['depot-paths']\n     --                            if p4PathStartsWith(value, p)]:\n     -+                            if p4PathStartsWith(value, p, self.verbose)]:\n     -                         continue\n     -                 else:\n     --                    if not p4PathStartsWith(value, self.depotPath):\n     -+                    if not p4PathStartsWith(value, self.depotPath, self.verbose):\n     -                         continue\n     -                 files_list.append(value)\n     -                 continue\n     -@@\n     -             return True\n     - \n     -         while True:\n     --            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \")\n     -+            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \").lower() \\\n     -+                .strip()[0]\n     -             if response == 'y':\n     -                 return True\n     -             if response == 'n':\n     -@@\n     -     def applyCommit(self, id):\n     -         \"\"\"Apply one commit, return True if it succeeded.\"\"\"\n     - \n     --        print(\"Applying\", read_pipe([\"git\", \"show\", \"-s\",\n     --                                     \"--format=format:%h %s\", id]))\n     -+        print((\"Applying\", read_pipe([\"git\", \"show\", \"-s\",\n     -+                                     \"--format=format:%h %s\", id])))\n     - \n     -         (p4User, gitEmail) = self.p4UserForCommit(id)\n     - \n     -@@\n     -                     # disable the read-only bit on windows.\n     -                     if self.isWindows and file not in editedFiles:\n     -                         os.chmod(file, stat.S_IWRITE)\n     --                    self.patchRCSKeywords(file, kwfiles[file])\n     --                    fixed_rcs_keywords = True\n     -+                    \n     -+                    try:\n     -+                        self.patchRCSKeywords(file, kwfiles[file])\n     -+                        fixed_rcs_keywords = True\n     -+                    except:\n     -+                        # We are throwing an exception, undo all open edits\n     -+                        for f in editedFiles:\n     -+                            p4_revert(f)\n     -+                        raise\n     -+            else:\n     -+                # They do not have attemptRCSCleanup set, this might be the fail point\n     -+                # Check to see if the file has RCS keywords and suggest setting the property.\n     -+                for file in editedFiles | filesToDelete:\n     -+                    if p4_keywords_regexp_for_file(file) != None:\n     -+                        print(\"At least one file in this commit has RCS Keywords that may be causing problems. \")\n     -+                        print(\"Consider:\\ngit config git-p4.attemptRCSCleanup true\")\n     -+                        break\n     - \n     -             if fixed_rcs_keywords:\n     -                 print(\"Retrying the patch with RCS keywords cleaned up\")\n     -@@\n     -             p4_delete(f)\n     - \n     -         # Set/clear executable bits\n     --        for f in filesToChangeExecBit.keys():\n     -+        for f in list(filesToChangeExecBit.keys()):\n     -             mode = filesToChangeExecBit[f]\n     -             setP4ExecBit(f, mode)\n     - \n      @@\n               tmpFile = os.fdopen(handle, \"w+b\")\n               if self.isWindows:\n     @@ -815,179 +340,6 @@\n       \n                       if update_shelve:\n                           p4_write_pipe(['shelve', '-r', '-i'], submitTemplate)\n     -@@\n     -                 if verbose:\n     -                     print(\"created p4 label for tag %s\" % name)\n     - \n     -+    def run_hook(self, hook_name, args = []):\n     -+        \"\"\" Runs a hook if it is found.\n     -+\n     -+            Returns NONE if the hook does not exist\n     -+            Returns TRUE if the exit code is 0, FALSE for a non-zero exit code.\n     -+        \"\"\"\n     -+        hook_file = self.find_hook(hook_name)\n     -+        if hook_file == None:\n     -+            if self.verbose:\n     -+                print(\"Skipping hook: %s\" % hook_name)\n     -+            return None\n     -+\n     -+        if self.verbose:\n     -+            print(\"hooks_path = %s \" % hooks_path)\n     -+            print(\"hook_file = %s \" % hook_file)\n     -+\n     -+        # Run the hook\n     -+        # TODO - allow non-list format\n     -+        cmd = [hook_file] + args\n     -+        return subprocess.call(cmd) == 0\n     -+\n     -+    def find_hook(self, hook_name):\n     -+        \"\"\" Locates the hook file for the given operating system.\n     -+        \"\"\"\n     -+        hooks_path = gitConfig(\"core.hooksPath\")\n     -+        if len(hooks_path) <= 0:\n     -+            hooks_path = os.path.join(os.environ.get(\"GIT_DIR\", \".git\"), \"hooks\")\n     -+\n     -+        # Look in the obvious place\n     -+        hook_file = os.path.join(hooks_path, hook_name)\n     -+        if os.path.isfile(hook_file) and os.access(hook_file, os.X_OK):\n     -+            return hook_file\n     -+\n     -+        # if we are windows, we will also allow them to have the hooks have extensions\n     -+        if (platform.system() == \"Windows\"):\n     -+            for ext in ['.exe', '.bat', 'ps1']:\n     -+                if os.path.isfile(hook_file + ext) and os.access(hook_file + ext, os.X_OK):\n     -+                    return hook_file + ext\n     -+\n     -+        # We didn't find the file\n     -+        return None\n     -+\n     -+\n     -+\n     -     def run(self, args):\n     -         if len(args) == 0:\n     -             self.master = currentGitBranch()\n     -@@\n     -             self.clientSpecDirs = getClientSpec()\n     - \n     -         # Check for the existence of P4 branches\n     --        branchesDetected = (len(p4BranchesInGit().keys()) > 1)\n     -+        branchesDetected = (len(list(p4BranchesInGit().keys())) > 1)\n     - \n     -         if self.useClientSpec and not branchesDetected:\n     -             # all files are relative to the client spec\n     -@@\n     -             sys.exit(\"number of commits (%d) must match number of shelved changelist (%d)\" %\n     -                      (len(commits), num_shelves))\n     - \n     --        hooks_path = gitConfig(\"core.hooksPath\")\n     --        if len(hooks_path) <= 0:\n     --            hooks_path = os.path.join(os.environ.get(\"GIT_DIR\", \".git\"), \"hooks\")\n     --\n     --        hook_file = os.path.join(hooks_path, \"p4-pre-submit\")\n     --        if os.path.isfile(hook_file) and os.access(hook_file, os.X_OK) and subprocess.call([hook_file]) != 0:\n     -+        rtn = self.run_hook(\"p4-pre-submit\")\n     -+        if rtn == False:\n     -             sys.exit(1)\n     - \n     -         #\n     -@@\n     -         last = len(commits) - 1\n     -         for i, commit in enumerate(commits):\n     -             if self.dry_run:\n     --                print(\" \", read_pipe([\"git\", \"show\", \"-s\",\n     --                                      \"--format=format:%h %s\", commit]))\n     -+                print((\" \", read_pipe([\"git\", \"show\", \"-s\",\n     -+                                      \"--format=format:%h %s\", commit])))\n     -                 ok = True\n     -             else:\n     -                 ok = self.applyCommit(commit)\n     -@@\n     -                         if self.conflict_behavior == \"ask\":\n     -                             print(\"What do you want to do?\")\n     -                             response = raw_input(\"[s]kip this commit but apply\"\n     --                                                 \" the rest, or [q]uit? \")\n     -+                                                 \" the rest, or [q]uit? \").lower().strip()[0]\n     -                             if not response:\n     -                                 continue\n     -                         elif self.conflict_behavior == \"skip\":\n     -@@\n     -                         star = \"*\"\n     -                     else:\n     -                         star = \" \"\n     --                    print(star, read_pipe([\"git\", \"show\", \"-s\",\n     --                                           \"--format=format:%h %s\",  c]))\n     -+                    print((star, read_pipe([\"git\", \"show\", \"-s\",\n     -+                                           \"--format=format:%h %s\",  c])))\n     -                 print(\"You will have to do 'git p4 sync' and rebase.\")\n     - \n     -         if gitConfigBool(\"git-p4.exportLabels\"):\n     -@@\n     -     # (\"-//depot/A/...\" becomes \"/depot/A/...\" after option parsing)\n     -     parser.values.cloneExclude += [\"/\" + re.sub(r\"\\.\\.\\.$\", \"\", value)]\n     - \n     -+\n     - class P4Sync(Command, P4UserMap):\n     - \n     -     def __init__(self):\n     -@@\n     -         self.knownBranches = {}\n     -         self.initialParents = {}\n     - \n     --        self.tz = \"%+03d%02d\" % (- time.timezone / 3600, ((- time.timezone % 3600) / 60))\n     -+        self.tz = \"%+03d%02d\" % (- time.timezone // 3600, ((- time.timezone % 3600) // 60))\n     -         self.labels = {}\n     - \n     -     # Force a checkpoint in fast-import and wait for it to finish\n     -@@\n     -     def isPathWanted(self, path):\n     -         for p in self.cloneExclude:\n     -             if p.endswith(\"/\"):\n     --                if p4PathStartsWith(path, p):\n     -+                if p4PathStartsWith(path, p, self.verbose):\n     -                     return False\n     -             # \"-//depot/file1\" without a trailing \"/\" should only exclude \"file1\", but not \"file111\" or \"file1_dir/file2\"\n     -             elif path.lower() == p.lower():\n     -                 return False\n     -         for p in self.depotPaths:\n     --            if p4PathStartsWith(path, p):\n     -+            if p4PathStartsWith(path, p, self.verbose):\n     -                 return True\n     -         return False\n     - \n     -     def extractFilesFromCommit(self, commit, shelved=False, shelved_cl = 0):\n     -+        \"\"\" Generates the list of files to be added in this git commit.\n     -+\n     -+            commit     = Unicode[] - data read from the P4 commit\n     -+            shelved    = Bool      - Is the P4 commit flagged as being shelved.\n     -+            shelved_cl = Unicode   - Numeric string with the changelist number.\n     -+        \"\"\"\n     -         files = []\n     -         fnum = 0\n     -         while \"depotFile%s\" % fnum in commit:\n     -@@\n     -             path = self.clientSpecDirs.map_in_client(path)\n     -             if self.detectBranches:\n     -                 for b in self.knownBranches:\n     --                    if p4PathStartsWith(path, b + \"/\"):\n     -+                    if p4PathStartsWith(path, b + \"/\", self.verbose):\n     -                         path = path[len(b)+1:]\n     - \n     -         elif self.keepRepoPath:\n     -@@\n     -             # //depot/; just look at first prefix as they all should\n     -             # be in the same depot.\n     -             depot = re.sub(\"^(//[^/]+/).*\", r'\\1', prefixes[0])\n     --            if p4PathStartsWith(path, depot):\n     -+            if p4PathStartsWith(path, depot, self.verbose):\n     -                 path = path[len(depot):]\n     - \n     -         else:\n     -             for p in prefixes:\n     --                if p4PathStartsWith(path, p):\n     -+                if p4PathStartsWith(path, p, self.verbose):\n     -                     path = path[len(p):]\n     -                     break\n     - \n      @@\n               return path\n       \n     @@ -1002,19 +354,6 @@\n       \n               if self.clientSpecDirs:\n                   files = self.extractFilesFromCommit(commit)\n     -@@\n     -             else:\n     -                 relPath = self.stripRepoPath(path, self.depotPaths)\n     - \n     --            for branch in self.knownBranches.keys():\n     -+            for branch in list(self.knownBranches.keys()):\n     -                 # add a trailing slash so that a commit into qt/4.2foo\n     -                 # doesn't end up in qt/4.2, e.g.\n     --                if p4PathStartsWith(relPath, branch + \"/\"):\n     -+                if p4PathStartsWith(relPath, branch + \"/\", self.verbose):\n     -                     if branch not in branches:\n     -                         branches[branch] = []\n     -                     branches[branch].append(file)\n      @@\n               return branches\n       \n     @@ -1031,18 +370,6 @@\n      +            self.gitStreamBytes.write(d)\n               self.gitStream.write('\\n')\n       \n     --    def encodeWithUTF8(self, path):\n     --        try:\n     --            path.decode('ascii')\n     --        except:\n     --            encoding = 'utf8'\n     --            if gitConfig('git-p4.pathEncoding'):\n     --                encoding = gitConfig('git-p4.pathEncoding')\n     --            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n     --            if self.verbose:\n     --                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n     --        return path\n     --\n      -    # output one file from the P4 stream\n      -    # - helper for streamP4Files\n      -\n     @@ -1053,18 +380,13 @@\n      +            contents should be a bytes (bytes) \n      +        \"\"\"\n               relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n     --        relPath = self.encodeWithUTF8(relPath)\n     -+        relPath = encodeWithUTF8(relPath, self.verbose)\n     +         relPath = encodeWithUTF8(relPath, self.verbose)\n               if verbose:\n     -             if 'fileSize' in self.stream_file:\n     +@@\n                       size = int(self.stream_file['fileSize'])\n                   else:\n                       size = 0 # deleted files don't get a fileSize apparently\n     --            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size/1024/1024))\n     -+            #if isunicode:\n     -+            #    sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), to_unicode(relPath), size//1024//1024))\n     -+            #else:\n     -+            #    sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), relPath, size//1024//1024))\n     +-            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size//1024//1024))\n      +            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), as_string(relPath), size//1024//1024))\n                   sys.stdout.flush()\n       \n     @@ -1100,15 +422,6 @@\n       \n               if self.largeFileSystem:\n      @@\n     - \n     -     def streamOneP4Deletion(self, file):\n     -         relPath = self.stripRepoPath(file['path'], self.branchPrefixes)\n     --        relPath = self.encodeWithUTF8(relPath)\n     -+        relPath = encodeWithUTF8(relPath, self.verbose)\n     -         if verbose:\n     -             sys.stdout.write(\"delete %s\\n\" % relPath)\n     -             sys.stdout.flush()\n     -@@\n               if self.largeFileSystem and self.largeFileSystem.isLargeFile(relPath):\n                   self.largeFileSystem.removeLargeFile(relPath)\n       \n     @@ -1133,13 +446,6 @@\n       \n               if not err and 'fileSize' in self.stream_file:\n                   required_bytes = int((4 * int(self.stream_file[\"fileSize\"])) - calcDiskFree())\n     -             if required_bytes > 0:\n     -                 err = 'Not enough space left on %s! Free at least %i MB.' % (\n     --                    os.getcwd(), required_bytes/1024/1024\n     -+                    os.getcwd(), required_bytes//1024//1024\n     -                 )\n     - \n     -         if err:\n      @@\n                   # ignore errors, but make sure it exits first\n                   self.importProcess.wait()\n     @@ -1155,12 +461,10 @@\n                   self.streamOneP4File(self.stream_file, self.stream_contents)\n                   self.stream_file = {}\n      @@\n     - \n               # pick up the new file information... for the\n               # 'data' field we need to append to our array\n     --        for k in marshalled.keys():\n     +         for k in list(marshalled.keys()):\n      -            if k == 'data':\n     -+        for k in list(marshalled.keys()):\n      +            if k == b'data':\n                       if 'streamContentSize' not in self.stream_file:\n                           self.stream_file['streamContentSize'] = 0\n     @@ -1178,12 +482,10 @@\n               if (verbose and\n                   'streamContentSize' in self.stream_file and\n      @@\n     -             'depotFile' in self.stream_file):\n                   size = int(self.stream_file[\"fileSize\"])\n                   if size > 0:\n     --                progress = 100*self.stream_file['streamContentSize']/size\n     --                sys.stdout.write('\\r%s %d%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size/1024/1024)))\n     -+                progress = 100.0*self.stream_file['streamContentSize']/size\n     +                 progress = 100.0*self.stream_file['streamContentSize']/size\n     +-                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size//1024//1024)))\n      +                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (path_as_string(self.stream_file['depotFile']), progress, int(size//1024//1024)))\n                       sys.stdout.flush()\n       \n     @@ -1227,24 +529,6 @@\n       \n               if verbose:\n      @@\n     - \n     -         gitStream.write(\"tagger %s\\n\" % tagger)\n     - \n     --        print(\"labelDetails=\",labelDetails)\n     -+        print((\"labelDetails=\",labelDetails))\n     -         if 'Description' in labelDetails:\n     -             description = labelDetails['Description']\n     -         else:\n     -@@\n     -         if not self.branchPrefixes:\n     -             return True\n     -         hasPrefix = [p for p in self.branchPrefixes\n     --                        if p4PathStartsWith(path, p)]\n     -+                        if p4PathStartsWith(path, p, self.verbose)]\n     -         if not hasPrefix and self.verbose:\n     -             print('Ignoring file outside of prefix: {0}'.format(path))\n     -         return hasPrefix\n     -@@\n                       .format(details['change']))\n                   return\n       \n     @@ -1307,58 +591,6 @@\n       \n               if len(parent) > 0:\n                   if self.verbose:\n     -@@\n     -             self.labels[newestChange] = [output, revisions]\n     - \n     -         if self.verbose:\n     --            print(\"Label changes: %s\" % self.labels.keys())\n     -+            print(\"Label changes: %s\" % list(self.labels.keys()))\n     - \n     -     # Import p4 labels as git tags. A direct mapping does not\n     -     # exist, so assume that if all the files are at the same revision\n     -@@\n     -                 source = paths[0]\n     -                 destination = paths[1]\n     -                 ## HACK\n     --                if p4PathStartsWith(source, self.depotPaths[0]) and p4PathStartsWith(destination, self.depotPaths[0]):\n     -+                if p4PathStartsWith(source, self.depotPaths[0], self.verbose) and p4PathStartsWith(destination, self.depotPaths[0], self.verbose):\n     -                     source = source[len(self.depotPaths[0]):-4]\n     -                     destination = destination[len(self.depotPaths[0]):-4]\n     - \n     -@@\n     - \n     -     def getBranchMappingFromGitBranches(self):\n     -         branches = p4BranchesInGit(self.importIntoRemotes)\n     --        for branch in branches.keys():\n     -+        for branch in list(branches.keys()):\n     -             if branch == \"master\":\n     -                 branch = \"main\"\n     -             else:\n     -@@\n     -             self.updateOptionDict(description)\n     - \n     -             if not self.silent:\n     --                sys.stdout.write(\"\\rImporting revision %s (%s%%)\" % (change, cnt * 100 / len(changes)))\n     -+                sys.stdout.write(\"\\rImporting revision %s (%4.1f%%)\" % (change, cnt * 100 / len(changes)))\n     -                 sys.stdout.flush()\n     -             cnt = cnt + 1\n     - \n     -             try:\n     -                 if self.detectBranches:\n     -                     branches = self.splitFilesIntoBranches(description)\n     --                    for branch in branches.keys():\n     -+                    for branch in list(branches.keys()):\n     -                         ## HACK  --hwn\n     -                         branchPrefix = self.depotPaths[0] + branch + \"/\"\n     -                         self.branchPrefixes = [ branchPrefix ]\n     -@@\n     -                 sys.exit(1)\n     - \n     -     def sync_origin_only(self):\n     -+        \"\"\" Ensures that the origin has been synchronized if one is set \"\"\"\n     -         if self.syncWithOrigin:\n     -             self.hasOrigin = originP4BranchesExist()\n     -             if self.hasOrigin:\n      @@\n                       system(\"git fetch origin\")\n       \n     @@ -1439,61 +671,6 @@\n           def closeStreams(self):\n               self.gitStream.close()\n      @@\n     -                 if short in branches:\n     -                     self.p4BranchesInGit = [ short ]\n     -             else:\n     --                self.p4BranchesInGit = branches.keys()\n     -+                self.p4BranchesInGit = list(branches.keys())\n     - \n     -             if len(self.p4BranchesInGit) > 1:\n     -                 if not self.silent:\n     -                     print(\"Importing from/into multiple branches\")\n     -                 self.detectBranches = True\n     --                for branch in branches.keys():\n     -+                for branch in list(branches.keys()):\n     -                     self.initialParents[self.refPrefix + branch] = \\\n     -                         branches[branch]\n     - \n     -@@\n     -                                  help=\"where to leave result of the clone\"),\n     -             optparse.make_option(\"--bare\", dest=\"cloneBare\",\n     -                                  action=\"store_true\", default=False),\n     -+            optparse.make_option(\"--encoding\", dest=\"setPathEncoding\",\n     -+                                 action=\"store\", default=None,\n     -+                                 help=\"Sets the path encoding for this depot\")\n     -         ]\n     -         self.cloneDestination = None\n     -         self.needsGit = False\n     -         self.cloneBare = False\n     -+        self.setPathEncoding = None\n     - \n     -     def defaultDestination(self, args):\n     -+        \"\"\"Returns the last path component as the default git \n     -+        repository directory name\"\"\"\n     -         ## TODO: use common prefix of args?\n     -         depotPath = args[0]\n     -         depotDir = re.sub(\"(@[^@]*)$\", \"\", depotPath)\n     -         depotDir = re.sub(\"(#[^#]*)$\", \"\", depotDir)\n     -         depotDir = re.sub(r\"\\.\\.\\.$\", \"\", depotDir)\n     -         depotDir = re.sub(r\"/$\", \"\", depotDir)\n     --        return os.path.split(depotDir)[1]\n     -+        return depotDir.split('/')[-1]\n     - \n     -     def run(self, args):\n     -         if len(args) < 1:\n     -@@\n     - \n     -         depotPaths = args\n     - \n     -+        # If we have an encoding provided, ignore what may already exist\n     -+        # in the registry. This will ensure we show the displayed values\n     -+        # using the correct encoding.\n     -+        if self.setPathEncoding:\n     -+            gitConfigSet(\"git-p4.pathEncoding\", self.setPathEncoding)\n     -+\n     -+        # If more than 1 path element is supplied, the last element\n     -+        # is the clone destination.\n     -         if not self.cloneDestination and len(depotPaths) > 1:\n                   self.cloneDestination = depotPaths[-1]\n                   depotPaths = depotPaths[:-1]\n       \n     @@ -1512,177 +689,3 @@\n       \n               if not os.path.exists(self.cloneDestination):\n                   os.makedirs(self.cloneDestination)\n     -@@\n     -         if retcode:\n     -             raise CalledProcessError(retcode, init_cmd)\n     - \n     -+        # Set the encoding if it was provided command line\n     -+        if self.setPathEncoding:\n     -+            init_cmd= [\"git\", \"config\", \"git-p4.pathEncoding\", self.setPathEncoding]\n     -+            retcode = subprocess.call(init_cmd)\n     -+            if retcode:\n     -+                raise CalledProcessError(retcode, init_cmd)\n     -+\n     -         if not P4Sync.run(self, depotPaths):\n     -             return False\n     - \n     -@@\n     -             to find the P4 commit we are based on, and the depot-paths.\n     -         \"\"\"\n     - \n     --        for parent in (range(65535)):\n     -+        for parent in (list(range(65535))):\n     -             log = extractLogMessageFromGitCommit(\"{0}^{1}\".format(starting_point, parent))\n     -             settings = extractSettingsGitLog(log)\n     -             if 'change' in settings:\n     -@@\n     -             print(\"%s <= %s (%s)\" % (branch, \",\".join(settings[\"depot-paths\"]), settings[\"change\"]))\n     -         return True\n     - \n     -+class Py23File():\n     -+    \"\"\" Python2/3 Unicode File Wrapper \n     -+    \"\"\"\n     -+    \n     -+    stream_handle = None\n     -+    verbose       = False\n     -+    debug_handle  = None\n     -+   \n     -+    def __init__(self, stream_handle, verbose = False,\n     -+                 debug_handle = None):\n     -+        \"\"\" Create a Python3 compliant Unicode to Byte String\n     -+            Windows compatible wrapper\n     -+\n     -+            stream_handle = the underlying file-like handle\n     -+            verbose       = Boolean if content should be echoed\n     -+            debug_handle  = A file-like handle data is duplicately written to\n     -+        \"\"\"\n     -+        self.stream_handle = stream_handle\n     -+        self.verbose       = verbose\n     -+        self.debug_handle  = debug_handle\n     -+\n     -+    def write(self, utf8string):\n     -+        \"\"\" Writes the utf8 encoded string to the underlying \n     -+            file stream\n     -+        \"\"\"\n     -+        self.stream_handle.write(as_bytes(utf8string))\n     -+        if self.verbose:\n     -+            sys.stderr.write(\"Stream Output: %s\" % utf8string)\n     -+            sys.stderr.flush()\n     -+        if self.debug_handle:\n     -+            self.debug_handle.write(as_bytes(utf8string))\n     -+\n     -+    def read(self, size = None):\n     -+        \"\"\" Reads int charcters from the underlying stream \n     -+            and converts it to utf8.\n     -+\n     -+            Be aware, the size value is for reading the underlying\n     -+            bytes so the value may be incorrect. Usage of the size\n     -+            value is discouraged.\n     -+        \"\"\"\n     -+        if size == None:\n     -+            return as_string(self.stream_handle.read())\n     -+        else:\n     -+            return as_string(self.stream_handle.read(size))\n     -+\n     -+    def readline(self):\n     -+        \"\"\" Reads a line from the underlying byte stream \n     -+            and converts it to utf8\n     -+        \"\"\"\n     -+        return as_string(self.stream_handle.readline())\n     -+\n     -+    def readlines(self, sizeHint = None):\n     -+        \"\"\" Returns a list containing lines from the file converted to unicode.\n     -+\n     -+            sizehint - Optional. If the optional sizehint argument is \n     -+            present, instead of reading up to EOF, whole lines totalling \n     -+            approximately sizehint bytes are read.\n     -+        \"\"\"\n     -+        lines = self.stream_handle.readlines(sizeHint)\n     -+        for i in range(0, len(lines)):\n     -+            lines[i] = as_string(lines[i])\n     -+        return lines\n     -+\n     -+    def close(self):\n     -+        \"\"\" Closes the underlying byte stream \"\"\"\n     -+        self.stream_handle.close()\n     -+\n     -+    def flush(self):\n     -+        \"\"\" Flushes the underlying byte stream \"\"\"\n     -+        self.stream_handle.flush()\n     -+\n     -+class DepotPath():\n     -+    \"\"\" Describes a DepotPath or File\n     -+    \"\"\"\n     -+\n     -+    raw_path = None\n     -+    utf8_path = None\n     -+    bytes_path = None\n     -+\n     -+    def __init__(self, path):\n     -+        \"\"\" Creates a new DepotPath with the path encoded\n     -+            with by the P4 repository\n     -+        \"\"\"\n     -+        raw_path = path\n     -+\n     -+    def raw():\n     -+        \"\"\" Returns the path as it was originally found\n     -+            in the P4 repository\n     -+        \"\"\"\n     -+        return raw_path\n     -+\n     -+    def startswith(self, prefix, start = None, end = None):\n     -+        \"\"\" Return True if string starts with the prefix, otherwise \n     -+            return False. prefix can also be a tuple of prefixes to \n     -+            look for. With optional start, test string beginning at \n     -+            that position. With optional end, stop comparing \n     -+            string at that position.\n     -+        \"\"\"\n     -+        return raw_path.startswith(prefix, start, end)\n     -+\n     -+\n     - class HelpFormatter(optparse.IndentedHelpFormatter):\n     -     def __init__(self):\n     -         optparse.IndentedHelpFormatter.__init__(self)\n     -@@\n     - \n     - def main():\n     -     if len(sys.argv[1:]) == 0:\n     --        printUsage(commands.keys())\n     -+        printUsage(list(commands.keys()))\n     -         sys.exit(2)\n     - \n     -     cmdName = sys.argv[1]\n     -@@\n     -     except KeyError:\n     -         print(\"unknown command %s\" % cmdName)\n     -         print(\"\")\n     --        printUsage(commands.keys())\n     -+        printUsage(list(commands.keys()))\n     -         sys.exit(2)\n     - \n     -     options = cmd.options\n     -@@\n     -                                    description = cmd.description,\n     -                                    formatter = HelpFormatter())\n     - \n     --    (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n     -+    try:\n     -+        (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n     -+    except:\n     -+        parser.print_help()\n     -+        raise\n     -+\n     -     global verbose\n     -     verbose = cmd.verbose\n     -     if cmd.needsGit:\n     -@@\n     -                         chdir(cdup);\n     - \n     -         if not isValidGitDir(cmd.gitdir):\n     --            if isValidGitDir(cmd.gitdir + \"/.git\"):\n     --                cmd.gitdir += \"/.git\"\n     -+            if isValidGitDir(os.path.join(cmd.gitdir, \".git\")):\n     -+                cmd.gitdir = os.path.join(cmd.gitdir, \".git\")\n     -             else:\n     -                 die(\"fatal: cannot locate git repository at %s\" % cmd.gitdir)\n     - \n  -:  ---------- > 11:  883ef45ca5 git-p4: Added --encoding parameter to p4 clone\n\n-- \ngitgitgadget\n"},{"id":"387537","messageId":"e1a424a955071414a634a703a85f1969f968bb0f.1575498578.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 08/11] git-p4: p4CmdList - support Unicode encoding","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:34Z","receivedAt":"2019-12-04T22:29:53Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThe p4CmdList is a commonly used function in the git-p4 code. It is used to execute a command in P4 and return the results of the call in a list.\n\nChange this code to take a new optional parameter, encode_data that will optionally convert the data AS_STRING() that isto be returned by the function.\n\nChange the code so that the key will always be encoded AS_STRING()\n\nData that is passed for standard input (stdin) should be AS_BYTES() to ensure unicode text that is supplied will be written out as bytes.\n\nAdditionally, change literal text prior to conversion to be literal bytes.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit 88306ac269186cbd0f6dc6cfd366b50b28ee4886)\n---\n git-p4.py | 27 +++++++++++++++++++++++----\n 1 file changed, 23 insertions(+), 4 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 0da640be93..f7c0ef0c53 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -711,7 +711,23 @@ def isModeExecChanged(src_mode, dst_mode):\n     return isModeExec(src_mode) != isModeExec(dst_mode)\n \n def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n-        errors_as_exceptions=False):\n+        errors_as_exceptions=False, encode_data=True):\n+    \"\"\" Executes a P4 command:  'cmd' optionally passing 'stdin' to the command's\n+        standard input via a temporary file with 'stdin_mode' mode.\n+\n+        Output from the command is optionally passed to the callback function 'cb'.\n+        If 'cb' is None, the response from the command is parsed into a list\n+        of resulting dictionaries. (For each block read from the process pipe.)\n+\n+        If 'skip_info' is true, information in a block read that has a code type of\n+        'info' will be skipped.\n+\n+        If 'errors_as_exceptions' is set to true (the default is false) the error\n+        code returned from the execution will generate an exception.\n+\n+        If 'encode_data' is set to true (the default) the data that is returned \n+        by this function will be passed through the \"as_string\" function.\n+    \"\"\"\n \n     if not isinstance(cmd, list):\n         cmd = \"-G \" + cmd\n@@ -734,7 +750,7 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n             stdin_file.write(stdin)\n         else:\n             for i in stdin:\n-                stdin_file.write(i + '\\n')\n+                stdin_file.write(as_bytes(i) + b'\\n')\n         stdin_file.flush()\n         stdin_file.seek(0)\n \n@@ -748,12 +764,15 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n         while True:\n             entry = marshal.load(p4.stdout)\n             if skip_info:\n-                if 'code' in entry and entry['code'] == 'info':\n+                if b'code' in entry and entry[b'code'] == b'info':\n                     continue\n             if cb is not None:\n                 cb(entry)\n             else:\n-                result.append(entry)\n+                out = {}\n+                for key, value in entry.items():\n+                    out[as_string(key)] = (as_string(value) if encode_data else value)\n+                result.append(out)\n     except EOFError:\n         pass\n     exitCode = p4.wait()\n-- \ngitgitgadget\n\n"},{"id":"387538","messageId":"10dc059444b965c3db3fda5600de64da32de53b4.1575498578.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 07/11] git-p4: Add a helper class for stream writing","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:33Z","receivedAt":"2019-12-04T22:29:53Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThis is a transtional commit that does not change current behvior.  It adds a new class Py23File.\n\nFollowing the Python recommendation of keeping text as unicode internally and only converting to and from bytes on input and output, this class provides an interface for the methods used for reading and writing files and file like streams.\n\nCreate a class that wraps the input and output functions used by the git-p4.py code for reading and writing to standard file handles.\n\nThe methods of this class should take a Unicode string for writing and return unicode strings in reads.  This class should be a drop-in for existing file like streams\n\nThe following methods should be coded for supporting existing read/write calls:\n* write - this should write a Unicode string to the underlying stream\n* read - this should read from the underlying stream and cast the bytes as a unicode string\n* readline - this should read one line of text from the underlying stream and cast it as a unicode string\n* readline - this should read a number of lines, optionally hinted, and cast each line as a unicode string\n\nThe expression \"cast as a unicode string\" is used because the code should use the AS_BYTES() and AS_UNICODE() functions instead of cohercing the data to actual unicode strings or bytes.  This allows python 2 code to continue to use the internal \"str\" data type instead of converting the data back and forth to actual unicode strings. This retains current python2 support while python3 support may be incomplete.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit 12919111fbaa3e4c0c4c2fdd4f79744cc683d860)\n---\n git-p4.py | 66 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 66 insertions(+)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 7ac8cb42ef..0da640be93 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -4182,6 +4182,72 @@ def run(self, args):\n             print(\"%s <= %s (%s)\" % (branch, \",\".join(settings[\"depot-paths\"]), settings[\"change\"]))\n         return True\n \n+class Py23File():\n+    \"\"\" Python2/3 Unicode File Wrapper \n+    \"\"\"\n+    \n+    stream_handle = None\n+    verbose       = False\n+    debug_handle  = None\n+   \n+    def __init__(self, stream_handle, verbose = False):\n+        \"\"\" Create a Python3 compliant Unicode to Byte String\n+            Windows compatible wrapper\n+\n+            stream_handle = the underlying file-like handle\n+            verbose       = Boolean if content should be echoed\n+        \"\"\"\n+        self.stream_handle = stream_handle\n+        self.verbose       = verbose\n+\n+    def write(self, utf8string):\n+        \"\"\" Writes the utf8 encoded string to the underlying \n+            file stream\n+        \"\"\"\n+        self.stream_handle.write(as_bytes(utf8string))\n+        if self.verbose:\n+            sys.stderr.write(\"Stream Output: %s\" % utf8string)\n+            sys.stderr.flush()\n+\n+    def read(self, size = None):\n+        \"\"\" Reads int charcters from the underlying stream \n+            and converts it to utf8.\n+\n+            Be aware, the size value is for reading the underlying\n+            bytes so the value may be incorrect. Usage of the size\n+            value is discouraged.\n+        \"\"\"\n+        if size == None:\n+            return as_string(self.stream_handle.read())\n+        else:\n+            return as_string(self.stream_handle.read(size))\n+\n+    def readline(self):\n+        \"\"\" Reads a line from the underlying byte stream \n+            and converts it to utf8\n+        \"\"\"\n+        return as_string(self.stream_handle.readline())\n+\n+    def readlines(self, sizeHint = None):\n+        \"\"\" Returns a list containing lines from the file converted to unicode.\n+\n+            sizehint - Optional. If the optional sizehint argument is \n+            present, instead of reading up to EOF, whole lines totalling \n+            approximately sizehint bytes are read.\n+        \"\"\"\n+        lines = self.stream_handle.readlines(sizeHint)\n+        for i in range(0, len(lines)):\n+            lines[i] = as_string(lines[i])\n+        return lines\n+\n+    def close(self):\n+        \"\"\" Closes the underlying byte stream \"\"\"\n+        self.stream_handle.close()\n+\n+    def flush(self):\n+        \"\"\" Flushes the underlying byte stream \"\"\"\n+        self.stream_handle.flush()\n+\n class HelpFormatter(optparse.IndentedHelpFormatter):\n     def __init__(self):\n         optparse.IndentedHelpFormatter.__init__(self)\n-- \ngitgitgadget\n\n"},{"id":"387539","messageId":"8f5752c12737fd861274609fdafac095ad95c519.1575498578.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 06/11] git-p4: Fix assumed path separators to be more Windows friendly","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:32Z","receivedAt":"2019-12-04T22:29:55Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nWhen a computer is configured to use Git for windows and Python for windows, and not a Unix subsystem like cygwin or WSL, the directory separator changes and causes git-p4 to fail to properly determine paths.\n\nFix 3 path separator errors:\n\n1. getUserCacheFilename should not use string concatenation. Change this code to use os.path.join to build an OS tolerant path.\n2. defaultDestiantion used the OS.path.split to split depot paths.  This is incorrect on windows. Change the code to split on a forward slash(/) instead since depot paths use this character regardless  of the operating system.\n3. The call to isvalidGitDir() in the main code also used a literal forward slash. Change the cose to use os.path.join to correctly format the path for the operating system.\n\nThese three changes allow the suggested windows configuration to properly locate files while retaining the existing behavior on non-windows operating systems.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit a5b45c12c3861638a933b05a1ffee0c83978dcb2)\n---\n git-p4.py | 13 +++++++++----\n 1 file changed, 9 insertions(+), 4 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 2659531c2e..7ac8cb42ef 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -1454,8 +1454,10 @@ def p4UserIsMe(self, p4User):\n             return True\n \n     def getUserCacheFilename(self):\n+        \"\"\" Returns the filename of the username cache \n+\t    \"\"\"\n         home = os.environ.get(\"HOME\", os.environ.get(\"USERPROFILE\"))\n-        return home + \"/.gitp4-usercache.txt\"\n+        return os.path.join(home, \".gitp4-usercache.txt\")\n \n     def getUserMapFromPerforceServer(self):\n         if self.userMapFromPerforceServer:\n@@ -3973,13 +3975,16 @@ def __init__(self):\n         self.cloneBare = False\n \n     def defaultDestination(self, args):\n+        \"\"\" Returns the last path component as the default git \n+            repository directory name\n+        \"\"\"\n         ## TODO: use common prefix of args?\n         depotPath = args[0]\n         depotDir = re.sub(\"(@[^@]*)$\", \"\", depotPath)\n         depotDir = re.sub(\"(#[^#]*)$\", \"\", depotDir)\n         depotDir = re.sub(r\"\\.\\.\\.$\", \"\", depotDir)\n         depotDir = re.sub(r\"/$\", \"\", depotDir)\n-        return os.path.split(depotDir)[1]\n+        return depotDir.split('/')[-1]\n \n     def run(self, args):\n         if len(args) < 1:\n@@ -4252,8 +4257,8 @@ def main():\n                         chdir(cdup);\n \n         if not isValidGitDir(cmd.gitdir):\n-            if isValidGitDir(cmd.gitdir + \"/.git\"):\n-                cmd.gitdir += \"/.git\"\n+            if isValidGitDir(os.path.join(cmd.gitdir, \".git\")):\n+                cmd.gitdir = os.path.join(cmd.gitdir, \".git\")\n             else:\n                 die(\"fatal: cannot locate git repository at %s\" % cmd.gitdir)\n \n-- \ngitgitgadget\n\n"},{"id":"387541","messageId":"4fc49313f0d68a913ad19085ddb337ac4c18d0fe.1575498578.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 09/11] git-p4: Add usability enhancements","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:35Z","receivedAt":"2019-12-04T22:29:56Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nIssue: when prompting the user with raw_input, the tests are not forgiving of user input.  For example, on the first query asks for a yes/no response. If the user enters the full word \"yes\" or \"no\" the test will fail. Additionally, offer the suggestion of setting git-p4.attemptRCSCleanup when applying a commit fails because of RCS keywords. Both of these changes are usability enhancement suggestions.\n\nChange the code prompting the user for input to sanitize the user input before checking the response by asking the response as a lower case string, trimming leading/trailing spaces, and returning the first character.\n\nChange the applyCommit() method that when applying a commit fails becasue of the P4 RCS Keywords, the user should consider setting git-p4.attemptRCSCleanup.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit 1fab571664f5b6ad4ef321199f52615a32a9f8c7)\n---\n git-p4.py | 31 ++++++++++++++++++++++++++-----\n 1 file changed, 26 insertions(+), 5 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex f7c0ef0c53..f13e4645a3 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -1909,7 +1909,8 @@ def edit_template(self, template_file):\n             return True\n \n         while True:\n-            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \")\n+            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \").lower() \\\n+                .strip()[0]\n             if response == 'y':\n                 return True\n             if response == 'n':\n@@ -2069,8 +2070,23 @@ def applyCommit(self, id):\n                     # disable the read-only bit on windows.\n                     if self.isWindows and file not in editedFiles:\n                         os.chmod(file, stat.S_IWRITE)\n-                    self.patchRCSKeywords(file, kwfiles[file])\n-                    fixed_rcs_keywords = True\n+                    \n+                    try:\n+                        self.patchRCSKeywords(file, kwfiles[file])\n+                        fixed_rcs_keywords = True\n+                    except:\n+                        # We are throwing an exception, undo all open edits\n+                        for f in editedFiles:\n+                            p4_revert(f)\n+                        raise\n+            else:\n+                # They do not have attemptRCSCleanup set, this might be the fail point\n+                # Check to see if the file has RCS keywords and suggest setting the property.\n+                for file in editedFiles | filesToDelete:\n+                    if p4_keywords_regexp_for_file(file) != None:\n+                        print(\"At least one file in this commit has RCS Keywords that may be causing problems. \")\n+                        print(\"Consider:\\ngit config git-p4.attemptRCSCleanup true\")\n+                        break\n \n             if fixed_rcs_keywords:\n                 print(\"Retrying the patch with RCS keywords cleaned up\")\n@@ -2481,7 +2497,7 @@ def run(self, args):\n                         if self.conflict_behavior == \"ask\":\n                             print(\"What do you want to do?\")\n                             response = raw_input(\"[s]kip this commit but apply\"\n-                                                 \" the rest, or [q]uit? \")\n+                                                 \" the rest, or [q]uit? \").lower().strip()[0]\n                             if not response:\n                                 continue\n                         elif self.conflict_behavior == \"skip\":\n@@ -4327,7 +4343,12 @@ def main():\n                                    description = cmd.description,\n                                    formatter = HelpFormatter())\n \n-    (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n+    try:\n+        (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n+    except:\n+        parser.print_help()\n+        raise\n+\n     global verbose\n     verbose = cmd.verbose\n     if cmd.needsGit:\n-- \ngitgitgadget\n\n"},{"id":"387540","messageId":"04a0aedbaa213f046c49f97b0cb47581962e282c.1575498578.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 10/11] git-p4: Support python3 for basic P4 clone, sync, and submit","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:36Z","receivedAt":"2019-12-04T22:29:57Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nIssue: Python 3 is still not properly supported for any use with the git-p4 python code.\nWarning - this is a very large atomic commit.  The commit text is also very large.\n\nChange the code such that, with the exception of P4 depot paths and depot files, all text read by git-p4 is cast as a string as soon as possible and converted back to bytes as late as possible, following Python2 to Python3 conversion best practices.\n\nImportant: Do not cast the bytes that contain the p4 depot path or p4 depot file name.  These should be left as bytes until used.\n\nThese two values should not be converted because the encoding of these values is unknown.  git-p4 supports a configuration value git-p4.pathEncoding that is used by the encodeWithUTF8()  to determine what a UTF8 version of the path and filename should be.  However, since depot path and depot filename need to be sent to P4 in their original encoding, they will be left as byte streams until they are actually used:\n\n* When sent to P4, the bytes are literally passed to the p4 command\n* When displayed in text for the user, they should be passed through the path_as_string() function\n* When used by GIT they should be passed through the encodeWithUTF8() function\n\nChange all the rest of system calls to cast output (stdin) as_bytes() and input (stdout) as_string().  This retains existing Python 2 support, and adds python 3 support for these functions:\n* read_pipe_full\n* read_pipe_lines\n* p4_has_move_command (used internally)\n* gitConfig\n* branch_exists\n* GitLFS.generatePointer\n* applyCommit - template must be read and written to the temporary file as_bytes() since it is created in memory as a string.\n* streamOneP4File(file, contents) - wrap calls to the depotFile in path_as_string() for display. The file contents must be retained as bytes, so update the RCS changes to be forced to bytes.\n* streamP4Files\n* importHeadRevision(revision) - encode the depotPaths for display separate from the text for processing.\n\nPy23File usage -\nChange the P4Sync.OpenStreams() function to cast the gitOutput, gitStream, and gitError streams as Py23File() wrapper classes.  This facilitates taking strings in both python 2 and python 3 and casting them to bytes in the wrapper class instead of having to modify each method. Since the fast-import command also expects a raw byte stream for file content, add a new stream handle - gitStreamBytes which is an unwrapped verison of gitStream.\n\nLiteral text -\nDepending on context, most literal text does not need casting to unicode or bytes as the text is Python dependent - In python 2, the string is implied as 'str' and python 3 the string is implied as 'unicode'. Under these conditions, they match the rest of the operating text, following best practices.  However, when a literal string is used in functions that are dealing with the raw input from and raw ouput to files streams, literal bytes may be required. Additionally, functions that are dealing with P4 depot paths or P4 depot file names are also dealing with bytes and will require the same casting as bytes.  The following functions cast text as byte strings:\n* wildcard_decode(path) - the path parameter is a P4 depot and is bytes. Cast all the literals to bytes.\n* wildcard_encode(path) - the path parameter is a P4 depot and is bytes. Cast all the literals to bytes.\n* streamP4FilesCb(marshalled) - the marshalled data is in bytes. Cast the literals as bytes. When using this data to manipulate self.stream_file, encode all the marshalled data except for the 'depotFile' name.\n* streamP4Files\n\nSpecial behavior:\n* p4_describe - encoding is disabled for the depotFile(x) and path elements since these are depot path and depo filenames.\n* p4PathStartsWith(path, prefix) - Since P4 depot paths can contain non-UTF-8 encoded strings, change this method to compare paths while supporting the optional encoding.\n   - First, perform a byte-to-byte check to see if the path and prefix are both identical text.  There is no need to perform encoding conversions if the text is identical.\n   - If the byte check fails, pass both the path and prefix through encodeWithUTF8() to ensure both paths are using the same encoding. Then perform the test as originally written.\n* patchRCSKeywords(file, pattern) - the parameters of file and pattern are both strings. However this function changes the contents of the file itentified by name \"file\". Treat the content of this file as binary to ensure that python does not accidently change the original encoding. The regular expression is cast as_bytes() and run against the file as_bytes(). The P4 keywords are ASCII strings and cannot span lines so iterating over each line of the file is acceptable.\n* writeToGitStream(gitMode, relPath, contents) - Since 'contents' is already bytes data, instead of using the self.gitStream, use the new self.gitStreamBytes - the unwrapped gitStream that does not cast as_bytes() the binary data.\n* commit(details, files, branch, parent = \"\", allow_empty=False) - Changed the encoding for the commit message to the preferred format for fast-import. The number of bytes is sent in the data block instead of using the EOT marker.\n* Change the code for handling the user cache to use binary files. Cast text as_bytes() when writing to the cache and as_string() when reading from the cache.  This makes the reading and writing of the cache determinstic in it's encoding. Unlike file paths, P4 encodes the user names in UTF-8 encoding so no additional string encoding is required.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit 65ff0c74ebe62a200b4385ecfd4aa618ce091f48)\n---\n git-p4.py | 287 ++++++++++++++++++++++++++++++++++++++----------------\n 1 file changed, 205 insertions(+), 82 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex f13e4645a3..05db2ec657 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -268,6 +268,8 @@ def read_pipe_full(c):\n     expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=expand)\n     (out, err) = p.communicate()\n+    out = as_string(out)\n+    err = as_string(err)\n     return (p.returncode, out, err)\n \n def read_pipe(c, ignore_error=False):\n@@ -294,10 +296,17 @@ def read_pipe_text(c):\n         return out.rstrip()\n \n def p4_read_pipe(c, ignore_error=False):\n+    \"\"\" Read output from the P4 command 'c'. Returns the output text on\n+        success. On failure, terminates execution, unless\n+        ignore_error is True, when it returns an empty string.\n+    \"\"\"\n     real_cmd = p4_build_cmd(c)\n     return read_pipe(real_cmd, ignore_error)\n \n def read_pipe_lines(c):\n+    \"\"\" Returns a list of text from executing the command 'c'.\n+        The program will die if the command fails to execute.\n+    \"\"\"\n     if verbose:\n         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n \n@@ -307,6 +316,11 @@ def read_pipe_lines(c):\n     val = pipe.readlines()\n     if pipe.close() or p.wait():\n         die('Command failed: %s' % str(c))\n+    # Unicode conversion from byte-string\n+    # Iterate and fix in-place to avoid a second list in memory.\n+    if isunicode:\n+        for i in range(len(val)):\n+            val[i] = as_string(val[i])\n \n     return val\n \n@@ -335,6 +349,8 @@ def p4_has_move_command():\n     cmd = p4_build_cmd([\"move\", \"-k\", \"@from\", \"@to\"])\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     (out, err) = p.communicate()\n+    out=as_string(out)\n+    err=as_string(err)\n     # return code will be 1 in either case\n     if err.find(\"Invalid option\") >= 0:\n         return False\n@@ -462,16 +478,20 @@ def p4_last_change():\n     return int(results[0]['change'])\n \n def p4_describe(change, shelved=False):\n-    \"\"\"Make sure it returns a valid result by checking for\n-       the presence of field \"time\".  Return a dict of the\n-       results.\"\"\"\n+    \"\"\" Returns information about the requested P4 change list.\n+\n+        Data returned is not string encoded (returned as bytes)\n+    \"\"\"\n+    # Make sure it returns a valid result by checking for\n+    #   the presence of field \"time\".  Return a dict of the\n+    #   results.\n \n     cmd = [\"describe\", \"-s\"]\n     if shelved:\n         cmd += [\"-S\"]\n     cmd += [str(change)]\n \n-    ds = p4CmdList(cmd, skip_info=True)\n+    ds = p4CmdList(cmd, skip_info=True, encode_data=False)\n     if len(ds) != 1:\n         die(\"p4 describe -s %d did not return 1 result: %s\" % (change, str(ds)))\n \n@@ -481,12 +501,23 @@ def p4_describe(change, shelved=False):\n         die(\"p4 describe -s %d exited with %d: %s\" % (change, d[\"p4ExitCode\"],\n                                                       str(d)))\n     if \"code\" in d:\n-        if d[\"code\"] == \"error\":\n+        if d[\"code\"] == b\"error\":\n             die(\"p4 describe -s %d returned error code: %s\" % (change, str(d)))\n \n     if \"time\" not in d:\n         die(\"p4 describe -s %d returned no \\\"time\\\": %s\" % (change, str(d)))\n \n+    # Do not convert 'depotFile(X)' or 'path' to be UTF-8 encoded, however \n+    # cast as_string() the rest of the text. \n+    keys=d.keys()\n+    for key in keys:\n+        if key.startswith('depotFile'):\n+            d[key]=d[key] \n+        elif key == 'path':\n+            d[key]=d[key] \n+        else:\n+            d[key] = as_string(d[key])\n+\n     return d\n \n #\n@@ -800,6 +831,8 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n     return result\n \n def p4Cmd(cmd):\n+    \"\"\" Executes a P4 command and returns the results in a dictionary\n+    \"\"\"\n     list = p4CmdList(cmd)\n     result = {}\n     for entry in list:\n@@ -908,13 +941,15 @@ def gitDeleteRef(ref):\n _gitConfig = {}\n \n def gitConfig(key, typeSpecifier=None):\n+    \"\"\" Return a configuration setting from GIT\n+\t\"\"\"\n     if key not in _gitConfig:\n         cmd = [ \"git\", \"config\" ]\n         if typeSpecifier:\n             cmd += [ typeSpecifier ]\n         cmd += [ key ]\n         s = read_pipe(cmd, ignore_error=True)\n-        _gitConfig[key] = s.strip()\n+        _gitConfig[key] = as_string(s).strip()\n     return _gitConfig[key]\n \n def gitConfigBool(key):\n@@ -988,6 +1023,7 @@ def branch_exists(branch):\n     cmd = [ \"git\", \"rev-parse\", \"--symbolic\", \"--verify\", branch ]\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     out, _ = p.communicate()\n+    out = as_string(out)\n     if p.returncode:\n         return False\n     # expect exactly one line of output: the branch name\n@@ -1171,9 +1207,22 @@ def p4PathStartsWith(path, prefix):\n     #\n     # we may or may not have a problem. If you have core.ignorecase=true,\n     # we treat DirA and dira as the same directory\n+    \n+    # Since we have to deal with mixed encodings for p4 file\n+    # paths, first perform a simple startswith check, this covers\n+    # the case that the formats and path are identical.\n+    if as_bytes(path).startswith(as_bytes(prefix)):\n+        return True\n+    \n+    # attempt to convert the prefix and path both to utf8\n+    path_utf8 = encodeWithUTF8(path)\n+    prefix_utf8 = encodeWithUTF8(prefix)\n+\n     if gitConfigBool(\"core.ignorecase\"):\n-        return path.lower().startswith(prefix.lower())\n-    return path.startswith(prefix)\n+        # Check if we match byte-per-byte.  \n+        \n+        return path_utf8.lower().startswith(prefix_utf8.lower())\n+    return path_utf8.startswith(prefix_utf8)\n \n def getClientSpec():\n     \"\"\"Look at the p4 client spec, create a View() object that contains\n@@ -1229,18 +1278,24 @@ def wildcard_decode(path):\n     # Cannot have * in a filename in windows; untested as to\n     # what p4 would do in such a case.\n     if not platform.system() == \"Windows\":\n-        path = path.replace(\"%2A\", \"*\")\n-    path = path.replace(\"%23\", \"#\") \\\n-               .replace(\"%40\", \"@\") \\\n-               .replace(\"%25\", \"%\")\n+        path = path.replace(b\"%2A\", b\"*\")\n+    path = path.replace(b\"%23\", b\"#\") \\\n+               .replace(b\"%40\", b\"@\") \\\n+               .replace(b\"%25\", b\"%\")\n     return path\n \n def wildcard_encode(path):\n     # do % first to avoid double-encoding the %s introduced here\n-    path = path.replace(\"%\", \"%25\") \\\n-               .replace(\"*\", \"%2A\") \\\n-               .replace(\"#\", \"%23\") \\\n-               .replace(\"@\", \"%40\")\n+    if isinstance(path, unicode):\n+        path = path.replace(\"%\", \"%25\") \\\n+                   .replace(\"*\", \"%2A\") \\\n+                   .replace(\"#\", \"%23\") \\\n+                   .replace(\"@\", \"%40\")\n+    else:\n+        path = path.replace(b\"%\", b\"%25\") \\\n+                   .replace(b\"*\", b\"%2A\") \\\n+                   .replace(b\"#\", b\"%23\") \\\n+                   .replace(b\"@\", b\"%40\")\n     return path\n \n def wildcard_present(path):\n@@ -1372,7 +1427,7 @@ def generatePointer(self, contentFile):\n             ['git', 'lfs', 'pointer', '--file=' + contentFile],\n             stdout=subprocess.PIPE\n         )\n-        pointerFile = pointerProcess.stdout.read()\n+        pointerFile = as_string(pointerProcess.stdout.read())\n         if pointerProcess.wait():\n             os.remove(contentFile)\n             die('git-lfs pointer command failed. Did you install the extension?')\n@@ -1479,6 +1534,8 @@ def getUserCacheFilename(self):\n         return os.path.join(home, \".gitp4-usercache.txt\")\n \n     def getUserMapFromPerforceServer(self):\n+        \"\"\" Creates the usercache from the data in P4.\n+        \"\"\"\n         if self.userMapFromPerforceServer:\n             return\n         self.users = {}\n@@ -1504,18 +1561,22 @@ def getUserMapFromPerforceServer(self):\n         for (key, val) in list(self.users.items()):\n             s += \"%s\\t%s\\n\" % (key.expandtabs(1), val.expandtabs(1))\n \n-        open(self.getUserCacheFilename(), \"wb\").write(s)\n+        cache = io.open(self.getUserCacheFilename(), \"wb\")\n+        cache.write(as_bytes(s))\n+        cache.close()\n         self.userMapFromPerforceServer = True\n \n     def loadUserMapFromCache(self):\n+        \"\"\" Reads the P4 username to git email map \n+        \"\"\"\n         self.users = {}\n         self.userMapFromPerforceServer = False\n         try:\n-            cache = open(self.getUserCacheFilename(), \"rb\")\n+            cache = io.open(self.getUserCacheFilename(), \"rb\")\n             lines = cache.readlines()\n             cache.close()\n             for line in lines:\n-                entry = line.strip().split(\"\\t\")\n+                entry = as_string(line).strip().split(\"\\t\")\n                 self.users[entry[0]] = entry[1]\n         except IOError:\n             self.getUserMapFromPerforceServer()\n@@ -1715,21 +1776,27 @@ def prepareLogMessage(self, template, message, jobs):\n         return result\n \n     def patchRCSKeywords(self, file, pattern):\n-        # Attempt to zap the RCS keywords in a p4 controlled file matching the given pattern\n+        \"\"\" Attempt to zap the RCS keywords in a p4 \n+            controlled file matching the given pattern\n+        \"\"\"\n+        bSubLine = as_bytes(r'$\\1$')\n         (handle, outFileName) = tempfile.mkstemp(dir='.')\n         try:\n-            outFile = os.fdopen(handle, \"w+\")\n-            inFile = open(file, \"r\")\n-            regexp = re.compile(pattern, re.VERBOSE)\n+            outFile = os.fdopen(handle, \"w+b\")\n+            inFile = open(file, \"rb\")\n+            regexp = re.compile(as_bytes(pattern), re.VERBOSE)\n             for line in inFile.readlines():\n-                line = regexp.sub(r'$\\1$', line)\n+                line = regexp.sub(bSubLine, line)\n                 outFile.write(line)\n             inFile.close()\n             outFile.close()\n+            outFile = None\n             # Forcibly overwrite the original file\n             os.unlink(file)\n             shutil.move(outFileName, file)\n         except:\n+            if outFile != None:\n+                outFile.close()\n             # cleanup our temporary file\n             os.unlink(outFileName)\n             print(\"Failed to strip RCS keywords in %s\" % file)\n@@ -2149,7 +2216,7 @@ def applyCommit(self, id):\n         tmpFile = os.fdopen(handle, \"w+b\")\n         if self.isWindows:\n             submitTemplate = submitTemplate.replace(\"\\n\", \"\\r\\n\")\n-        tmpFile.write(submitTemplate)\n+        tmpFile.write(as_bytes(submitTemplate))\n         tmpFile.close()\n \n         if self.prepare_p4_only:\n@@ -2199,8 +2266,8 @@ def applyCommit(self, id):\n                 message = tmpFile.read()\n                 tmpFile.close()\n                 if self.isWindows:\n-                    message = message.replace(\"\\r\\n\", \"\\n\")\n-                submitTemplate = message[:message.index(separatorLine)]\n+                    message = message.replace(b\"\\r\\n\", b\"\\n\")\n+                submitTemplate = message[:message.index(as_bytes(separatorLine))]\n \n                 if update_shelve:\n                     p4_write_pipe(['shelve', '-r', '-i'], submitTemplate)\n@@ -2843,8 +2910,11 @@ def stripRepoPath(self, path, prefixes):\n         return path\n \n     def splitFilesIntoBranches(self, commit):\n-        \"\"\"Look at each depotFile in the commit to figure out to what\n-           branch it belongs.\"\"\"\n+        \"\"\" Look at each depotFile in the commit to figure out to what\n+            branch it belongs.\n+\n+            Data in the commit will NOT be encoded\n+        \"\"\"\n \n         if self.clientSpecDirs:\n             files = self.extractFilesFromCommit(commit)\n@@ -2885,16 +2955,22 @@ def splitFilesIntoBranches(self, commit):\n         return branches\n \n     def writeToGitStream(self, gitMode, relPath, contents):\n-        self.gitStream.write('M %s inline %s\\n' % (gitMode, relPath))\n+        \"\"\" Writes the bytes[] 'contents' to the git fast-import\n+            with the given 'gitMode' and 'relPath' as the relative\n+            path.\n+        \"\"\"\n+        self.gitStream.write('M %s inline %s\\n' % (gitMode, as_string(relPath)))\n         self.gitStream.write('data %d\\n' % sum(len(d) for d in contents))\n         for d in contents:\n-            self.gitStream.write(d)\n+            self.gitStreamBytes.write(d)\n         self.gitStream.write('\\n')\n \n-    # output one file from the P4 stream\n-    # - helper for streamP4Files\n-\n     def streamOneP4File(self, file, contents):\n+        \"\"\" output one file from the P4 stream to the git inbound stream.\n+            helper for streamP4files.\n+\n+            contents should be a bytes (bytes) \n+        \"\"\"\n         relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n         relPath = encodeWithUTF8(relPath, self.verbose)\n         if verbose:\n@@ -2902,7 +2978,7 @@ def streamOneP4File(self, file, contents):\n                 size = int(self.stream_file['fileSize'])\n             else:\n                 size = 0 # deleted files don't get a fileSize apparently\n-            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size//1024//1024))\n+            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), as_string(relPath), size//1024//1024))\n             sys.stdout.flush()\n \n         (type_base, type_mods) = split_p4_type(file[\"type\"])\n@@ -2920,7 +2996,7 @@ def streamOneP4File(self, file, contents):\n                 # to nothing.  This causes p4 errors when checking out such\n                 # a change, and errors here too.  Work around it by ignoring\n                 # the bad symlink; hopefully a future change fixes it.\n-                print(\"\\nIgnoring empty symlink in %s\" % file['depotFile'])\n+                print(\"\\nIgnoring empty symlink in %s\" % path_as_string(file['depotFile']))\n                 return\n             elif data[-1] == '\\n':\n                 contents = [data[:-1]]\n@@ -2960,16 +3036,16 @@ def streamOneP4File(self, file, contents):\n             # Ideally, someday, this script can learn how to generate\n             # appledouble files directly and import those to git, but\n             # non-mac machines can never find a use for apple filetype.\n-            print(\"\\nIgnoring apple filetype file %s\" % file['depotFile'])\n+            print(\"\\nIgnoring apple filetype file %s\" % path_as_string(file['depotFile']))\n             return\n \n         # Note that we do not try to de-mangle keywords on utf16 files,\n         # even though in theory somebody may want that.\n-        pattern = p4_keywords_regexp_for_type(type_base, type_mods)\n+        pattern = as_bytes(p4_keywords_regexp_for_type(type_base, type_mods))\n         if pattern:\n             regexp = re.compile(pattern, re.VERBOSE)\n-            text = ''.join(contents)\n-            text = regexp.sub(r'$\\1$', text)\n+            text = b''.join(contents)\n+            text = regexp.sub(as_bytes(r'$\\1$'), text)\n             contents = [ text ]\n \n         if self.largeFileSystem:\n@@ -2988,15 +3064,19 @@ def streamOneP4Deletion(self, file):\n         if self.largeFileSystem and self.largeFileSystem.isLargeFile(relPath):\n             self.largeFileSystem.removeLargeFile(relPath)\n \n-    # handle another chunk of streaming data\n     def streamP4FilesCb(self, marshalled):\n+        \"\"\" Callback function for recording P4 chunks of data for streaming \n+            into GIT.\n+\n+            marshalled data is bytes[] from the caller\n+        \"\"\"\n \n         # catch p4 errors and complain\n         err = None\n-        if \"code\" in marshalled:\n-            if marshalled[\"code\"] == \"error\":\n-                if \"data\" in marshalled:\n-                    err = marshalled[\"data\"].rstrip()\n+        if b\"code\" in marshalled:\n+            if marshalled[b\"code\"] == b\"error\":\n+                if b\"data\" in marshalled:\n+                    err = marshalled[b\"data\"].rstrip()\n \n         if not err and 'fileSize' in self.stream_file:\n             required_bytes = int((4 * int(self.stream_file[\"fileSize\"])) - calcDiskFree())\n@@ -3018,11 +3098,11 @@ def streamP4FilesCb(self, marshalled):\n             # ignore errors, but make sure it exits first\n             self.importProcess.wait()\n             if f:\n-                die(\"Error from p4 print for %s: %s\" % (f, err))\n+                die(\"Error from p4 print for %s: %s\" % (path_as_string(f), err))\n             else:\n                 die(\"Error from p4 print: %s\" % err)\n \n-        if 'depotFile' in marshalled and self.stream_have_file_info:\n+        if b'depotFile' in marshalled and self.stream_have_file_info:\n             # start of a new file - output the old one first\n             self.streamOneP4File(self.stream_file, self.stream_contents)\n             self.stream_file = {}\n@@ -3032,13 +3112,16 @@ def streamP4FilesCb(self, marshalled):\n         # pick up the new file information... for the\n         # 'data' field we need to append to our array\n         for k in list(marshalled.keys()):\n-            if k == 'data':\n+            if k == b'data':\n                 if 'streamContentSize' not in self.stream_file:\n                     self.stream_file['streamContentSize'] = 0\n-                self.stream_file['streamContentSize'] += len(marshalled['data'])\n-                self.stream_contents.append(marshalled['data'])\n+                self.stream_file['streamContentSize'] += len(marshalled[b'data'])\n+                self.stream_contents.append(marshalled[b'data'])\n             else:\n-                self.stream_file[k] = marshalled[k]\n+                if k == b'depotFile':\n+                    self.stream_file[as_string(k)] = marshalled[k]\n+                else:\n+                    self.stream_file[as_string(k)] = as_string(marshalled[k])\n \n         if (verbose and\n             'streamContentSize' in self.stream_file and\n@@ -3047,13 +3130,14 @@ def streamP4FilesCb(self, marshalled):\n             size = int(self.stream_file[\"fileSize\"])\n             if size > 0:\n                 progress = 100.0*self.stream_file['streamContentSize']/size\n-                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size//1024//1024)))\n+                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (path_as_string(self.stream_file['depotFile']), progress, int(size//1024//1024)))\n                 sys.stdout.flush()\n \n         self.stream_have_file_info = True\n \n-    # Stream directly from \"p4 files\" into \"git fast-import\"\n     def streamP4Files(self, files):\n+        \"\"\" Stream directly from \"p4 files\" into \"git fast-import\" \n+        \"\"\"\n         filesForCommit = []\n         filesToRead = []\n         filesToDelete = []\n@@ -3074,7 +3158,7 @@ def streamP4Files(self, files):\n             self.stream_contents = []\n             self.stream_have_file_info = False\n \n-            # curry self argument\n+            # Callback for P4 command to collect file content\n             def streamP4FilesCbSelf(entry):\n                 self.streamP4FilesCb(entry)\n \n@@ -3083,9 +3167,9 @@ def streamP4FilesCbSelf(entry):\n                 if 'shelved_cl' in f:\n                     # Handle shelved CLs using the \"p4 print file@=N\" syntax to print\n                     # the contents\n-                    fileArg = '%s@=%d' % (f['path'], f['shelved_cl'])\n+                    fileArg = b'%s@=%d' % (f['path'], as_bytes(f['shelved_cl']))\n                 else:\n-                    fileArg = '%s#%s' % (f['path'], f['rev'])\n+                    fileArg = b'%s#%s' % (f['path'], as_bytes(f['rev']))\n \n                 fileArgs.append(fileArg)\n \n@@ -3105,7 +3189,7 @@ def make_email(self, userid):\n \n     def streamTag(self, gitStream, labelName, labelDetails, commit, epoch):\n         \"\"\" Stream a p4 tag.\n-        commit is either a git commit, or a fast-import mark, \":<p4commit>\"\n+            commit is either a git commit, or a fast-import mark, \":<p4commit>\"\n         \"\"\"\n \n         if verbose:\n@@ -3177,7 +3261,22 @@ def commit(self, details, files, branch, parent = \"\", allow_empty=False):\n                 .format(details['change']))\n             return\n \n+        # fast-import:\n+        #'commit' SP <ref> LF\n+\t    #mark?\n+\t    #original-oid?\n+\t    #('author' (SP <name>)? SP LT <email> GT SP <when> LF)?\n+\t    #'committer' (SP <name>)? SP LT <email> GT SP <when> LF\n+\t    #('encoding' SP <encoding>)?\n+\t    #data\n+\t    #('from' SP <commit-ish> LF)?\n+\t    #('merge' SP <commit-ish> LF)*\n+\t    #(filemodify | filedelete | filecopy | filerename | filedeleteall | notemodify)*\n+\t    #LF?\n+        \n+        #'commit' - <ref> is the name of the branch to make the commit on\n         self.gitStream.write(\"commit %s\\n\" % branch)\n+        #'mark' SP :<idnum>\n         self.gitStream.write(\"mark :%s\\n\" % details[\"change\"])\n         self.committedChanges.add(int(details[\"change\"]))\n         committer = \"\"\n@@ -3187,19 +3286,29 @@ def commit(self, details, files, branch, parent = \"\", allow_empty=False):\n \n         self.gitStream.write(\"committer %s\\n\" % committer)\n \n-        self.gitStream.write(\"data <<EOT\\n\")\n-        self.gitStream.write(details[\"desc\"])\n+        # Per https://git-scm.com/docs/git-fast-import\n+        # The preferred method for creating the commit message is to supply the \n+        # byte count in the data method and not to use a Delimited format. \n+        # Collect all the text in the commit message into a single string and \n+        # compute the byte count.\n+        commitText = details[\"desc\"]\n         if len(jobs) > 0:\n-            self.gitStream.write(\"\\nJobs: %s\" % (' '.join(jobs)))\n-\n+            commitText += \"\\nJobs: %s\" % (' '.join(jobs))\n         if not self.suppress_meta_comment:\n-            self.gitStream.write(\"\\n[git-p4: depot-paths = \\\"%s\\\": change = %s\" %\n-                                (','.join(self.branchPrefixes), details[\"change\"]))\n-            if len(details['options']) > 0:\n-                self.gitStream.write(\": options = %s\" % details['options'])\n-            self.gitStream.write(\"]\\n\")\n+            # coherce the path to the correct formatting in the branch prefixes as well.\n+            dispPaths = []\n+            for p in self.branchPrefixes:\n+                dispPaths += [path_as_string(p)]\n \n-        self.gitStream.write(\"EOT\\n\\n\")\n+            commitText += (\"\\n[git-p4: depot-paths = \\\"%s\\\": change = %s\" %\n+                                (','.join(dispPaths), details[\"change\"]))\n+            if len(details['options']) > 0:\n+                commitText += (\": options = %s\" % details['options'])\n+            commitText += \"]\"\n+        commitText += \"\\n\" \n+        self.gitStream.write(\"data %s\\n\" % len(as_bytes(commitText)))\n+        self.gitStream.write(commitText)\n+        self.gitStream.write(\"\\n\")\n \n         if len(parent) > 0:\n             if self.verbose:\n@@ -3606,30 +3715,35 @@ def sync_origin_only(self):\n                 system(\"git fetch origin\")\n \n     def importHeadRevision(self, revision):\n-        print(\"Doing initial import of %s from revision %s into %s\" % (' '.join(self.depotPaths), revision, self.branch))\n-\n+        # Re-encode depot text\n+        dispPaths = []\n+        utf8Paths = []\n+        for p in self.depotPaths:\n+            dispPaths += [path_as_string(p)]\n+        print(\"Doing initial import of %s from revision %s into %s\" % (' '.join(dispPaths), revision, self.branch))\n         details = {}\n         details[\"user\"] = \"git perforce import user\"\n-        details[\"desc\"] = (\"Initial import of %s from the state at revision %s\\n\"\n-                           % (' '.join(self.depotPaths), revision))\n+        details[\"desc\"] = (\"Initial import of %s from the state at revision %s\\n\" %\n+                           (' '.join(dispPaths), revision))\n         details[\"change\"] = revision\n         newestRevision = 0\n+        del dispPaths\n \n         fileCnt = 0\n         fileArgs = [\"%s...%s\" % (p,revision) for p in self.depotPaths]\n \n-        for info in p4CmdList([\"files\"] + fileArgs):\n+        for info in p4CmdList([\"files\"] + fileArgs, encode_data = False):\n \n-            if 'code' in info and info['code'] == 'error':\n+            if 'code' in info and info['code'] == b'error':\n                 sys.stderr.write(\"p4 returned an error: %s\\n\"\n-                                 % info['data'])\n-                if info['data'].find(\"must refer to client\") >= 0:\n+                                 % as_string(info['data']))\n+                if info['data'].find(b\"must refer to client\") >= 0:\n                     sys.stderr.write(\"This particular p4 error is misleading.\\n\")\n                     sys.stderr.write(\"Perhaps the depot path was misspelled.\\n\");\n                     sys.stderr.write(\"Depot path:  %s\\n\" % \" \".join(self.depotPaths))\n                 sys.exit(1)\n             if 'p4ExitCode' in info:\n-                sys.stderr.write(\"p4 exitcode: %s\\n\" % info['p4ExitCode'])\n+                sys.stderr.write(\"p4 exitcode: %s\\n\" % as_string(info['p4ExitCode']))\n                 sys.exit(1)\n \n \n@@ -3642,8 +3756,10 @@ def importHeadRevision(self, revision):\n                 #fileCnt = fileCnt + 1\n                 continue\n \n+            # Save all the file information, howerver do not translate the depotFile name at \n+            # this time. Leave that as bytes since the encoding may vary.\n             for prop in [\"depotFile\", \"rev\", \"action\", \"type\" ]:\n-                details[\"%s%s\" % (prop, fileCnt)] = info[prop]\n+                details[\"%s%s\" % (prop, fileCnt)] = (info[prop] if prop == \"depotFile\" else as_string(info[prop]))\n \n             fileCnt = fileCnt + 1\n \n@@ -3663,13 +3779,18 @@ def importHeadRevision(self, revision):\n             print(self.gitError.read())\n \n     def openStreams(self):\n+        \"\"\" Opens the fast import pipes.  Note that the git* streams are wrapped\n+            to expect Unicode text.  To send a raw byte Array, use the importProcess\n+            underlying port\n+        \"\"\"\n         self.importProcess = subprocess.Popen([\"git\", \"fast-import\"],\n                                               stdin=subprocess.PIPE,\n                                               stdout=subprocess.PIPE,\n                                               stderr=subprocess.PIPE);\n-        self.gitOutput = self.importProcess.stdout\n-        self.gitStream = self.importProcess.stdin\n-        self.gitError = self.importProcess.stderr\n+        self.gitOutput = Py23File(self.importProcess.stdout, verbose = self.verbose)\n+        self.gitStream = Py23File(self.importProcess.stdin, verbose = self.verbose)\n+        self.gitError = Py23File(self.importProcess.stderr, verbose = self.verbose)\n+        self.gitStreamBytes = self.importProcess.stdin\n \n     def closeStreams(self):\n         self.gitStream.close()\n@@ -4035,15 +4156,17 @@ def run(self, args):\n             self.cloneDestination = depotPaths[-1]\n             depotPaths = depotPaths[:-1]\n \n+        dispPaths = []\n         for p in depotPaths:\n             if not p.startswith(\"//\"):\n                 sys.stderr.write('Depot paths must start with \"//\": %s\\n' % p)\n                 return False\n+            dispPaths += [path_as_string(p)]\n \n         if not self.cloneDestination:\n             self.cloneDestination = self.defaultDestination(args)\n \n-        print(\"Importing from %s into %s\" % (', '.join(depotPaths), self.cloneDestination))\n+        print(\"Importing from %s into %s\" % (', '.join(dispPaths), path_as_string(self.cloneDestination)))\n \n         if not os.path.exists(self.cloneDestination):\n             os.makedirs(self.cloneDestination)\n-- \ngitgitgadget\n\n"},{"id":"387542","messageId":"883ef45ca5476c6ff412bf4f95b0e7b50c58338b.1575498578.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v4 11/11] git-p4: Added --encoding parameter to p4 clone","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-04T22:29:37Z","receivedAt":"2019-12-04T22:29:59Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThe test t9822 did not have any tests that had encoded a directory name in ISO8859-1.\n\nAdditionally, to make it easier for the user to clone new repositories with a non-UTF-8 encoded path in P4, add a new parameter to p4clone \"--encoding\" that sets the\n\nAdd new tests that use ISO8859-1 encoded text in both the directory and file names.  git-p4.pathEncoding.\n\nUpdate the View class in the git-p4 code to properly cast text as_string() except for depot path and filenames.\n\nUpdate the documentation to include the new command line parameter for p4clone\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n(cherry picked from commit e26f6309d60c6c1615320d4a9071935e23efe6fb)\n---\n Documentation/git-p4.txt        |   5 ++\n git-p4.py                       |  61 +++++++++++++------\n t/t9822-git-p4-path-encoding.sh | 101 ++++++++++++++++++++++++++++++++\n 3 files changed, 149 insertions(+), 18 deletions(-)\n\ndiff --git a/Documentation/git-p4.txt b/Documentation/git-p4.txt\nindex 3494a1db3e..f54af3c917 100644\n--- a/Documentation/git-p4.txt\n+++ b/Documentation/git-p4.txt\n@@ -305,6 +305,11 @@ options described above.\n --bare::\n \tPerform a bare clone.  See linkgit:git-clone[1].\n \n+--encoding <encoding>::\n+    Optionally sets the git-p4.pathEncoding configuration value in \n+\tthe newly created Git repository before files are synchronized \n+\tfrom P4. See git-p4.pathEncoding for more information.\n+\n Submit options\n ~~~~~~~~~~~~~~\n These options can be used to modify 'git p4 submit' behavior.\ndiff --git a/git-p4.py b/git-p4.py\nindex 05db2ec657..1f2e43430a 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -1228,7 +1228,7 @@ def getClientSpec():\n     \"\"\"Look at the p4 client spec, create a View() object that contains\n        all the mappings, and return it.\"\"\"\n \n-    specList = p4CmdList(\"client -o\")\n+    specList = p4CmdList(\"client -o\", encode_data=False)\n     if len(specList) != 1:\n         die('Output from \"client -o\" is %d lines, expecting 1' %\n             len(specList))\n@@ -1237,7 +1237,7 @@ def getClientSpec():\n     entry = specList[0]\n \n     # the //client/ name\n-    client_name = entry[\"Client\"]\n+    client_name = as_string(entry[\"Client\"])\n \n     # just the keys that start with \"View\"\n     view_keys = [ k for k in list(entry.keys()) if k.startswith(\"View\") ]\n@@ -2637,19 +2637,25 @@ def run(self, args):\n         return True\n \n class View(object):\n-    \"\"\"Represent a p4 view (\"p4 help views\"), and map files in a\n-       repo according to the view.\"\"\"\n+    \"\"\" Represent a p4 view (\"p4 help views\"), and map files in a\n+        repo according to the view.\n+    \"\"\"\n \n     def __init__(self, client_name):\n         self.mappings = []\n-        self.client_prefix = \"//%s/\" % client_name\n+        # the client prefix is saved in bytes as it is used for comparison\n+        # against server data.\n+        self.client_prefix = as_bytes(\"//%s/\" % client_name)\n         # cache results of \"p4 where\" to lookup client file locations\n         self.client_spec_path_cache = {}\n \n     def append(self, view_line):\n-        \"\"\"Parse a view line, splitting it into depot and client\n-           sides.  Append to self.mappings, preserving order.  This\n-           is only needed for tag creation.\"\"\"\n+        \"\"\" Parse a view line, splitting it into depot and client\n+            sides.  Append to self.mappings, preserving order.  This\n+            is only needed for tag creation.\n+\n+            view_line should be in bytes (depot path encoding)\n+        \"\"\"\n \n         # Split the view line into exactly two words.  P4 enforces\n         # structure on these lines that simplifies this quite a bit.\n@@ -2662,28 +2668,28 @@ def append(self, view_line):\n         # The line is already white-space stripped.\n         # The two words are separated by a single space.\n         #\n-        if view_line[0] == '\"':\n+        if view_line[0] == b'\"':\n             # First word is double quoted.  Find its end.\n-            close_quote_index = view_line.find('\"', 1)\n+            close_quote_index = view_line.find(b'\"', 1)\n             if close_quote_index <= 0:\n-                die(\"No first-word closing quote found: %s\" % view_line)\n+                die(\"No first-word closing quote found: %s\" % path_as_string(view_line))\n             depot_side = view_line[1:close_quote_index]\n             # skip closing quote and space\n             rhs_index = close_quote_index + 1 + 1\n         else:\n-            space_index = view_line.find(\" \")\n+            space_index = view_line.find(b\" \")\n             if space_index <= 0:\n-                die(\"No word-splitting space found: %s\" % view_line)\n+                die(\"No word-splitting space found: %s\" % path_as_string(view_line))\n             depot_side = view_line[0:space_index]\n             rhs_index = space_index + 1\n \n         # prefix + means overlay on previous mapping\n-        if depot_side.startswith(\"+\"):\n+        if depot_side.startswith(b\"+\"):\n             depot_side = depot_side[1:]\n \n         # prefix - means exclude this path, leave out of mappings\n         exclude = False\n-        if depot_side.startswith(\"-\"):\n+        if depot_side.startswith(b\"-\"):\n             exclude = True\n             depot_side = depot_side[1:]\n \n@@ -2694,7 +2700,7 @@ def convert_client_path(self, clientFile):\n         # chop off //client/ part to make it relative\n         if not clientFile.startswith(self.client_prefix):\n             die(\"No prefix '%s' on clientFile '%s'\" %\n-                (self.client_prefix, clientFile))\n+                (as_string(self.client_prefix)), path_as_string(clientFile))\n         return clientFile[len(self.client_prefix):]\n \n     def update_client_spec_path_cache(self, files):\n@@ -2706,9 +2712,9 @@ def update_client_spec_path_cache(self, files):\n         if len(fileArgs) == 0:\n             return  # All files in cache\n \n-        where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs)\n+        where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs, encode_data=False)\n         for res in where_result:\n-            if \"code\" in res and res[\"code\"] == \"error\":\n+            if \"code\" in res and res[\"code\"] == b\"error\":\n                 # assume error is \"... file(s) not in client view\"\n                 continue\n             if \"clientFile\" not in res:\n@@ -4125,10 +4131,14 @@ def __init__(self):\n                                  help=\"where to leave result of the clone\"),\n             optparse.make_option(\"--bare\", dest=\"cloneBare\",\n                                  action=\"store_true\", default=False),\n+            optparse.make_option(\"--encoding\", dest=\"setPathEncoding\",\n+                                 action=\"store\", default=None,\n+                                 help=\"Sets the path encoding for this depot\")\n         ]\n         self.cloneDestination = None\n         self.needsGit = False\n         self.cloneBare = False\n+        self.setPathEncoding = None\n \n     def defaultDestination(self, args):\n         \"\"\" Returns the last path component as the default git \n@@ -4152,6 +4162,14 @@ def run(self, args):\n \n         depotPaths = args\n \n+        # If we have an encoding provided, ignore what may already exist\n+        # in the registry. This will ensure we show the displayed values\n+        # using the correct encoding.\n+        if self.setPathEncoding:\n+            gitConfigSet(\"git-p4.pathEncoding\", self.setPathEncoding)\n+\n+        # If more than 1 path element is supplied, the last element\n+        # is the clone destination.\n         if not self.cloneDestination and len(depotPaths) > 1:\n             self.cloneDestination = depotPaths[-1]\n             depotPaths = depotPaths[:-1]\n@@ -4179,6 +4197,13 @@ def run(self, args):\n         if retcode:\n             raise CalledProcessError(retcode, init_cmd)\n \n+        # Set the encoding if it was provided command line\n+        if self.setPathEncoding:\n+            init_cmd= [\"git\", \"config\", \"git-p4.pathEncoding\", self.setPathEncoding]\n+            retcode = subprocess.call(init_cmd)\n+            if retcode:\n+                raise CalledProcessError(retcode, init_cmd)\n+\n         if not P4Sync.run(self, depotPaths):\n             return False\n \ndiff --git a/t/t9822-git-p4-path-encoding.sh b/t/t9822-git-p4-path-encoding.sh\nindex 572d395498..cf8a15b2e4 100755\n--- a/t/t9822-git-p4-path-encoding.sh\n+++ b/t/t9822-git-p4-path-encoding.sh\n@@ -4,9 +4,20 @@ test_description='Clone repositories with non ASCII paths'\n \n . ./lib-git-p4.sh\n \n+# lowercase filename\n+# UTF8    - HEX:   a-\\xc3\\xa4_o-\\xc3\\xb6_u-\\xc3\\xbc\n+#         - octal: a-\\303\\244_o-\\303\\266_u-\\303\\274\n+# ISO8859 - HEX:   a-\\xe4_o-\\xf6_u-\\xfc\n UTF8_ESCAPED=\"a-\\303\\244_o-\\303\\266_u-\\303\\274.txt\"\n ISO8859_ESCAPED=\"a-\\344_o-\\366_u-\\374.txt\"\n \n+# lowercase directory\n+# UTF8    - HEX:   dir_a-\\xc3\\xa4_o-\\xc3\\xb6_u-\\xc3\\xbc\n+# ISO8859 - HEX:   dir_a-\\xe4_o-\\xf6_u-\\xfc\n+DIR_UTF8_ESCAPED=\"dir_a-\\303\\244_o-\\303\\266_u-\\303\\274\"\n+DIR_ISO8859_ESCAPED=\"dir_a-\\344_o-\\366_u-\\374\"\n+\n+\n ISO8859=\"$(printf \"$ISO8859_ESCAPED\")\" &&\n echo content123 >\"$ISO8859\" &&\n rm \"$ISO8859\" || {\n@@ -58,6 +69,22 @@ test_expect_success 'Clone repo containing iso8859-1 encoded paths with git-p4.p\n \t)\n '\n \n+test_expect_success 'Clone repo containing iso8859-1 encoded paths with using --encoding parameter' '\n+\ttest_when_finished cleanup_git &&\n+\t(\n+\t\tgit p4 clone --encoding iso8859 --destination=\"$git\" //depot &&\n+\t\tcd \"$git\" &&\n+\t\tUTF8=\"$(printf \"$UTF8_ESCAPED\")\" &&\n+\t\techo \"$UTF8\" >expect &&\n+\t\tgit -c core.quotepath=false ls-files >actual &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\techo content123 >expect &&\n+\t\tcat \"$UTF8\" >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n test_expect_success 'Delete iso8859-1 encoded paths and clone' '\n \t(\n \t\tcd \"$cli\" &&\n@@ -74,4 +101,78 @@ test_expect_success 'Delete iso8859-1 encoded paths and clone' '\n \t)\n '\n \n+# These tests will create a directory with ISO8859-1 characters in both the \n+# directory and the path.  Since it is possible to clone a path instead of using\n+# the whole client-spec.  Check both versions:  client-spec and with a direct\n+# path using --encoding\n+test_expect_success 'Create a repo containing iso8859-1 encoded directory and filename' '\n+\t(\n+\t\tDIR_ISO8859=\"$(printf \"$DIR_ISO8859_ESCAPED\")\" &&\n+\t\tISO8859=\"$(printf \"$ISO8859_ESCAPED\")\" &&\n+\t\tcd \"$cli\" &&\n+\t\tmkdir \"$DIR_ISO8859\" && \n+\t\tcd \"$DIR_ISO8859\" &&\n+\t\techo content123 >\"$ISO8859\" &&\n+\t\tp4 add \"$ISO8859\" &&\n+\t\tp4 submit -d \"test commit (encoded directory)\"\n+\t)\n+'\n+\n+test_expect_success 'Clone repo containing iso8859-1 encoded depot path and files with git-p4.pathEncoding' '\n+\ttest_when_finished cleanup_git &&\n+\t(\n+\t\tDIR_ISO8859=\"$(printf \"$DIR_ISO8859_ESCAPED\")\" &&\n+\t\tDIR_UTF8=\"$(printf \"$DIR_UTF8_ESCAPED\")\" &&\n+\t\tcd \"$git\" &&\n+\t\tgit init . &&\n+\t\tgit config git-p4.pathEncoding iso8859-1 &&\n+\t\tgit p4 clone --use-client-spec --destination=\"$git\" \"//depot/$DIR_ISO8859\" &&\n+\t\tcd \"$DIR_UTF8\" &&\n+\t\tUTF8=\"$(printf \"$UTF8_ESCAPED\")\" &&\n+\t\techo \"$UTF8\" >expect &&\n+\t\tgit -c core.quotepath=false ls-files >actual &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\techo content123 >expect &&\n+\t\tcat \"$UTF8\" >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'Clone repo containing iso8859-1 encoded depot path and files with git-p4.pathEncoding, without --use-client-spec' '\n+\ttest_when_finished cleanup_git &&\n+\t(\n+\t\tDIR_ISO8859=\"$(printf \"$DIR_ISO8859_ESCAPED\")\" &&\n+\t\tcd \"$git\" &&\n+\t\tgit init . &&\n+\t\tgit config git-p4.pathEncoding iso8859-1 &&\n+\t\tgit p4 clone --destination=\"$git\" \"//depot/$DIR_ISO8859\" &&\n+\t\tUTF8=\"$(printf \"$UTF8_ESCAPED\")\" &&\n+\t\techo \"$UTF8\" >expect &&\n+\t\tgit -c core.quotepath=false ls-files >actual &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\techo content123 >expect &&\n+\t\tcat \"$UTF8\" >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'Clone repo containing iso8859-1 encoded depot path and files with using --encoding parameter' '\n+\ttest_when_finished cleanup_git &&\n+\t(\n+\t\tDIR_ISO8859=\"$(printf \"$DIR_ISO8859_ESCAPED\")\" &&\n+\t\tgit p4 clone --encoding iso8859 --destination=\"$git\" \"//depot/$DIR_ISO8859\" &&\n+\t\tcd \"$git\" &&\n+\t\tUTF8=\"$(printf \"$UTF8_ESCAPED\")\" &&\n+\t\techo \"$UTF8\" >expect &&\n+\t\tgit -c core.quotepath=false ls-files >actual &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\techo content123 >expect &&\n+\t\tcat \"$UTF8\" >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n test_done\n-- \ngitgitgadget\n"},{"id":"387545","messageId":"CAE5ih7-6EbEM4z5BtY87=82H_tLypiOPq4WY5mm3190QExTZWQ@mail.gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 00/11] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Luke Diamand","fromEmail":"luke@diamand.org","sentAt":"2019-12-05T09:54:27Z","receivedAt":"2019-12-05T09:54:52Z","isPatch":true,"sender":{"key":"luke@diamand.org","avatar":"https://avatars.githubusercontent.com/u/5330967?v=4"},"body":"On Wed, 4 Dec 2019 at 22:29, Ben Keene via GitGitGadget\n<gitgitgadget@gmail.com> wrote:\n>\n> Issue: The current git-p4.py script does not work with python3.\n>\n> I have attempted to use the P4 integration built into GIT and I was unable\n> to get the program to run because I have Python 3.8 installed on my\n> computer. I was able to get the program to run when I downgraded my python\n> to version 2.7. However, python 2 is reaching its end of life.\n>\n> Submission: I am submitting a patch for the git-p4.py script that partially\n> supports python 3.8. This code was able to pass the basic tests (t9800) when\n> run against Python3. This provides basic functionality.\n>\n> In an attempt to pass the t9822 P4 path-encoding test, a new parameter for\n> git P4 Clone was introduced.\n>\n> --encoding Format-identifier\n>\n> This will create the GIT repository following the current functionality;\n> however, before importing the files from P4, it will set the\n> git-p4.pathEncoding option so any files or paths that are encoded with\n> non-ASCII/non-UTF-8 formats will import correctly.\n>\n> Technical details: The script was updated by futurize (\n> https://python-future.org/futurize.html) to support Py2/Py3 syntax. The few\n> references to classes in future were reworked so that future would not be\n> required. The existing code test for Unicode support was extended to\n> normalize the classes “unicode” and “bytes” to across platforms:\n>\n>  * ‘unicode’ is an alias for ‘str’ in Py3 and is the unicode class in Py2.\n>  * ‘bytes’ is bytes in Py3 and an alias for ‘str’ in Py2.\n>\n> New coercion methods were written for both Python2 and Python3:\n>\n>  * as_string(text) – In Python3, this encodes a bytes object as a UTF-8\n>    encoded Unicode string.\n>  * as_bytes(text) – In Python3, this decodes a Unicode string to an array of\n>    bytes.\n>\n> In Python2, these functions do not change the data since a ‘str’ object\n> function in both roles as strings and byte arrays. This reduces the\n> potential impact on backward compatibility with Python 2.\n>\n>  * to_unicode(text) – ensures that the supplied data is encoded as a UTF-8\n>    string. This function will encode data in both Python2 and Python3. *\n>       path_as_string(path) – This function is an extension function that\n>       honors the option “git-p4.pathEncoding” to convert a set of bytes or\n>       characters to UTF-8. If the str/bytes cannot decode as ASCII, it will\n>       use the encodeWithUTF8() method to convert the custom encoded bytes to\n>       Unicode in UTF-8.\n>\n>\n>\n> Generally speaking, information in the script is converted to Unicode as\n> early as possible and converted back to a byte array just before passing to\n> external programs or files. The exception to this rule is P4 Repository file\n> paths.\n>\n> Paths are not converted but left as “bytes” so the original file path\n> encoding can be preserved. This formatting is required for commands that\n> interact with the P4 file path. When the file path is used by GIT, it is\n> converted with encodeWithUTF8().\n>\n\nAlmost all the tests pass now - nice!\n\n(There's one test that fails for me, t9830-git-p4-symlink-dir.sh).\n\nNitpicking:\n\n- There are some bits of trailing whitespace around - can you strip\nthose out? You can use \"git diff --check\".\n- Also I think the convention for git commits is that they be limited\nto 72 (?) characters.\n- In 10dc commit message, s/behvior/behavior\n- Maybe submit 4fc4 as a separate patch series? It doesn't seem\ndirectly related to your python3 changes.\n- s/howerver/however/\n\nThe comment at line 3261 (showing the fast-import syntax) has wonky\nindentation, and needs a space after the '#'.\n\nThis code looked like we're duplicating stuff:\n\n+    if isinstance(path, unicode):\n+        path = path.replace(\"%\", \"%25\") \\\n+                   .replace(\"*\", \"%2A\") \\\n+                   .replace(\"#\", \"%23\") \\\n+                   .replace(\"@\", \"%40\")\n+    else:\n+        path = path.replace(b\"%\", b\"%25\") \\\n+                   .replace(b\"*\", b\"%2A\") \\\n+                   .replace(b\"#\", b\"%23\") \\\n+                   .replace(b\"@\", b\"%40\")\n\nI wonder if we can have a helper to do this?\n\nIn patchRCSKeywords() you've added code to cleanup outFile. But I\nwonder if we could just use a 'finally' block, or a contextexpr (\"with\nblah as outFile:\")\n\nI don't know if it's worth doing now that you've got it going, but at\none point I tried simplifying code like this:\n\n   path_as_string(file['depotFile'])\nand\n   marshalled[b'data']\n\nby using a dictionary with overloaded operators which would do the\nbytes/string conversion automatically. However, your approach isn't\nactually _that_ invasive, so maybe this is not necessary.\n\nLooks good though, thanks!\nLuke\n\n\n\n\n\n\n> Signed-off-by: Ben Keene seraphire@gmail.com [seraphire@gmail.com]\n>\n> Ben Keene (11):\n>   git-p4: select p4 binary by operating-system\n>   git-p4: change the expansion test from basestring to list\n>   git-p4: add new helper functions for python3 conversion\n>   git-p4: python3 syntax changes\n>   git-p4: Add new functions in preparation of usage\n>   git-p4: Fix assumed path separators to be more Windows friendly\n>   git-p4: Add a helper class for stream writing\n>   git-p4: p4CmdList  - support Unicode encoding\n>   git-p4: Add usability enhancements\n>   git-p4: Support python3 for basic P4 clone, sync, and submit\n>   git-p4: Added --encoding parameter to p4 clone\n>\n>  Documentation/git-p4.txt        |   5 +\n>  git-p4.py                       | 690 ++++++++++++++++++++++++--------\n>  t/t9822-git-p4-path-encoding.sh | 101 +++++\n>  3 files changed, 629 insertions(+), 167 deletions(-)\n>\n>\n> base-commit: 228f53135a4a41a37b6be8e4d6e2b6153db4a8ed\n> Published-As: https://github.com/gitgitgadget/git/releases/tag/pr-463%2Fseraphire%2Fseraphire%2Fp4-python3-unicode-v4\n> Fetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-463/seraphire/seraphire/p4-python3-unicode-v4\n> Pull-Request: https://github.com/gitgitgadget/git/pull/463\n>\n> Range-diff vs v3:\n>\n>   -:  ---------- >  1:  4012426993 git-p4: select p4 binary by operating-system\n>   -:  ---------- >  2:  0ef2f56b04 git-p4: change the expansion test from basestring to list\n>   -:  ---------- >  3:  f0e658b984 git-p4: add new helper functions for python3 conversion\n>   -:  ---------- >  4:  3c41db3e91 git-p4: python3 syntax changes\n>   -:  ---------- >  5:  1bf7b073b0 git-p4: Add new functions in preparation of usage\n>   -:  ---------- >  6:  8f5752c127 git-p4: Fix assumed path separators to be more Windows friendly\n>   -:  ---------- >  7:  10dc059444 git-p4: Add a helper class for stream writing\n>   -:  ---------- >  8:  e1a424a955 git-p4: p4CmdList  - support Unicode encoding\n>   -:  ---------- >  9:  4fc49313f0 git-p4: Add usability enhancements\n>   1:  02b3843e9f ! 10:  04a0aedbaa Python3 support for t9800 tests. Basic P4/Python3 support\n>      @@ -1,159 +1,60 @@\n>       Author: Ben Keene <seraphire@gmail.com>\n>\n>      -    Python3 support for t9800 tests. Basic P4/Python3 support\n>      +    git-p4: Support python3 for basic P4 clone, sync, and submit\n>      +\n>      +    Issue: Python 3 is still not properly supported for any use with the git-p4 python code.\n>      +    Warning - this is a very large atomic commit.  The commit text is also very large.\n>      +\n>      +    Change the code such that, with the exception of P4 depot paths and depot files, all text read by git-p4 is cast as a string as soon as possible and converted back to bytes as late as possible, following Python2 to Python3 conversion best practices.\n>      +\n>      +    Important: Do not cast the bytes that contain the p4 depot path or p4 depot file name.  These should be left as bytes until used.\n>      +\n>      +    These two values should not be converted because the encoding of these values is unknown.  git-p4 supports a configuration value git-p4.pathEncoding that is used by the encodeWithUTF8()  to determine what a UTF8 version of the path and filename should be.  However, since depot path and depot filename need to be sent to P4 in their original encoding, they will be left as byte streams until they are actually used:\n>      +\n>      +    * When sent to P4, the bytes are literally passed to the p4 command\n>      +    * When displayed in text for the user, they should be passed through the path_as_string() function\n>      +    * When used by GIT they should be passed through the encodeWithUTF8() function\n>      +\n>      +    Change all the rest of system calls to cast output (stdin) as_bytes() and input (stdout) as_string().  This retains existing Python 2 support, and adds python 3 support for these functions:\n>      +    * read_pipe_full\n>      +    * read_pipe_lines\n>      +    * p4_has_move_command (used internally)\n>      +    * gitConfig\n>      +    * branch_exists\n>      +    * GitLFS.generatePointer\n>      +    * applyCommit - template must be read and written to the temporary file as_bytes() since it is created in memory as a string.\n>      +    * streamOneP4File(file, contents) - wrap calls to the depotFile in path_as_string() for display. The file contents must be retained as bytes, so update the RCS changes to be forced to bytes.\n>      +    * streamP4Files\n>      +    * importHeadRevision(revision) - encode the depotPaths for display separate from the text for processing.\n>      +\n>      +    Py23File usage -\n>      +    Change the P4Sync.OpenStreams() function to cast the gitOutput, gitStream, and gitError streams as Py23File() wrapper classes.  This facilitates taking strings in both python 2 and python 3 and casting them to bytes in the wrapper class instead of having to modify each method. Since the fast-import command also expects a raw byte stream for file content, add a new stream handle - gitStreamBytes which is an unwrapped verison of gitStream.\n>      +\n>      +    Literal text -\n>      +    Depending on context, most literal text does not need casting to unicode or bytes as the text is Python dependent - In python 2, the string is implied as 'str' and python 3 the string is implied as 'unicode'. Under these conditions, they match the rest of the operating text, following best practices.  However, when a literal string is used in functions that are dealing with the raw input from and raw ouput to files streams, literal bytes may be required. Additionally, functions that are dealing with P4 depot paths or P4 depot file names are also dealing with bytes and will require the same casting as bytes.  The following functions cast text as byte strings:\n>      +    * wildcard_decode(path) - the path parameter is a P4 depot and is bytes. Cast all the literals to bytes.\n>      +    * wildcard_encode(path) - the path parameter is a P4 depot and is bytes. Cast all the literals to bytes.\n>      +    * streamP4FilesCb(marshalled) - the marshalled data is in bytes. Cast the literals as bytes. When using this data to manipulate self.stream_file, encode all the marshalled data except for the 'depotFile' name.\n>      +    * streamP4Files\n>      +\n>      +    Special behavior:\n>      +    * p4_describe - encoding is disabled for the depotFile(x) and path elements since these are depot path and depo filenames.\n>      +    * p4PathStartsWith(path, prefix) - Since P4 depot paths can contain non-UTF-8 encoded strings, change this method to compare paths while supporting the optional encoding.\n>      +       - First, perform a byte-to-byte check to see if the path and prefix are both identical text.  There is no need to perform encoding conversions if the text is identical.\n>      +       - If the byte check fails, pass both the path and prefix through encodeWithUTF8() to ensure both paths are using the same encoding. Then perform the test as originally written.\n>      +    * patchRCSKeywords(file, pattern) - the parameters of file and pattern are both strings. However this function changes the contents of the file itentified by name \"file\". Treat the content of this file as binary to ensure that python does not accidently change the original encoding. The regular expression is cast as_bytes() and run against the file as_bytes(). The P4 keywords are ASCII strings and cannot span lines so iterating over each line of the file is acceptable.\n>      +    * writeToGitStream(gitMode, relPath, contents) - Since 'contents' is already bytes data, instead of using the self.gitStream, use the new self.gitStreamBytes - the unwrapped gitStream that does not cast as_bytes() the binary data.\n>      +    * commit(details, files, branch, parent = \"\", allow_empty=False) - Changed the encoding for the commit message to the preferred format for fast-import. The number of bytes is sent in the data block instead of using the EOT marker.\n>      +    * Change the code for handling the user cache to use binary files. Cast text as_bytes() when writing to the cache and as_string() when reading from the cache.  This makes the reading and writing of the cache determinstic in it's encoding. Unlike file paths, P4 encodes the user names in UTF-8 encoding so no additional string encoding is required.\n>\n>           Signed-off-by: Ben Keene <seraphire@gmail.com>\n>      +    (cherry picked from commit 65ff0c74ebe62a200b4385ecfd4aa618ce091f48)\n>\n>        diff --git a/git-p4.py b/git-p4.py\n>        --- a/git-p4.py\n>        +++ b/git-p4.py\n>       @@\n>      - import zlib\n>      - import ctypes\n>      - import errno\n>      -+import os.path\n>      -+import codecs\n>      -+import io\n>      -\n>      - # support basestring in python3\n>      - try:\n>      -     unicode = unicode\n>      - except NameError:\n>      -     # 'unicode' is undefined, must be Python 3\n>      --    str = str\n>      -+    #\n>      -+    # For Python3 which is natively unicode, we will use\n>      -+    # unicode for internal information but all P4 Data\n>      -+    # will remain in bytes\n>      -+    isunicode = True\n>      -     unicode = str\n>      -     bytes = bytes\n>      --    basestring = (str,bytes)\n>      -+\n>      -+    def as_string(text):\n>      -+        \"\"\"Return a byte array as a unicode string\"\"\"\n>      -+        if text == None:\n>      -+            return None\n>      -+        if isinstance(text, bytes):\n>      -+            return unicode(text, \"utf-8\")\n>      -+        else:\n>      -+            return text\n>      -+\n>      -+    def as_bytes(text):\n>      -+        \"\"\"Return a Unicode string as a byte array\"\"\"\n>      -+        if text == None:\n>      -+            return None\n>      -+        if isinstance(text, bytes):\n>      -+            return text\n>      -+        else:\n>      -+            return bytes(text, \"utf-8\")\n>      -+\n>      -+    def to_unicode(text):\n>      -+        \"\"\"Return a byte array as a unicode string\"\"\"\n>      -+        return as_string(text)\n>      -+\n>      -+    def path_as_string(path):\n>      -+        \"\"\" Converts a path to the UTF8 encoded string \"\"\"\n>      -+        if isinstance(path, unicode):\n>      -+            return path\n>      -+        return encodeWithUTF8(path).decode('utf-8')\n>      -+\n>      - else:\n>      -     # 'unicode' exists, must be Python 2\n>      --    str = str\n>      -+    #\n>      -+    # We will treat the data as:\n>      -+    #   str   -> str\n>      -+    #   bytes -> str\n>      -+    # So for Python2 these functions are no-ops\n>      -+    # and will leave the data in the ambiguious\n>      -+    # string/bytes state\n>      -+    isunicode = False\n>      -     unicode = unicode\n>      -     bytes = str\n>      --    basestring = basestring\n>      -+\n>      -+    def as_string(text):\n>      -+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n>      -+        return text\n>      -+\n>      -+    def as_bytes(text):\n>      -+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n>      -+        return text\n>      -+\n>      -+    def to_unicode(text):\n>      -+        \"\"\"Return a string as a unicode string\"\"\"\n>      -+        return text.decode('utf-8')\n>      -+\n>      -+    def path_as_string(path):\n>      -+        \"\"\" Converts a path to the UTF8 encoded bytes \"\"\"\n>      -+        return encodeWithUTF8(path)\n>      -+\n>      -+\n>      -+\n>      -+# Check for raw_input support\n>      -+try:\n>      -+    raw_input\n>      -+except NameError:\n>      -+    raw_input = input\n>      -\n>      - try:\n>      -     from subprocess import CalledProcessError\n>      -@@\n>      -     location. It means that hooking into the environment, or other configuration\n>      -     can be done more easily.\n>      -     \"\"\"\n>      --    real_cmd = [\"p4\"]\n>      -+    # Look for the P4 binary\n>      -+    if (platform.system() == \"Windows\"):\n>      -+        real_cmd = [\"p4.exe\"]\n>      -+    else:\n>      -+        real_cmd = [\"p4\"]\n>      -\n>      -     user = gitConfig(\"git-p4.user\")\n>      -     if len(user) > 0:\n>      -@@\n>      -         # Provide a way to not pass this option by setting git-p4.retries to 0\n>      -         real_cmd += [\"-r\", str(retries)]\n>      -\n>      --    if isinstance(cmd,basestring):\n>      -+    if not isinstance(cmd, list):\n>      -         real_cmd = ' '.join(real_cmd) + ' ' + cmd\n>      -     else:\n>      -         real_cmd += cmd\n>      -@@\n>      -         sys.exit(1)\n>      -\n>      - def write_pipe(c, stdin):\n>      -+    \"\"\"Executes the command 'c', passing 'stdin' on the standard input\"\"\"\n>      -     if verbose:\n>      -         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n>      -\n>      --    expand = isinstance(c,basestring)\n>      -+    expand = not isinstance(c, list)\n>      -     p = subprocess.Popen(c, stdin=subprocess.PIPE, shell=expand)\n>      -     pipe = p.stdin\n>      -     val = pipe.write(stdin)\n>      -@@\n>      -     if p.wait():\n>      -         die('Command failed: %s' % str(c))\n>      -\n>      --    return val\n>      -\n>      - def p4_write_pipe(c, stdin):\n>      -+    \"\"\" Runs a P4 command 'c', passing 'stdin' data to P4\"\"\"\n>      -     real_cmd = p4_build_cmd(c)\n>      --    return write_pipe(real_cmd, stdin)\n>      -+    write_pipe(real_cmd, stdin)\n>      -\n>      - def read_pipe_full(c):\n>      -     \"\"\" Read output from  command. Returns a tuple\n>      -@@\n>      -     if verbose:\n>      -         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n>      -\n>      --    expand = isinstance(c,basestring)\n>      -+    expand = not isinstance(c, list)\n>      +     expand = not isinstance(c, list)\n>            p = subprocess.Popen(c, stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=expand)\n>            (out, err) = p.communicate()\n>       +    out = as_string(out)\n>      @@ -179,10 +80,7 @@\n>            if verbose:\n>                sys.stderr.write('Reading pipe: %s\\n' % str(c))\n>\n>      --    expand = isinstance(c, basestring)\n>      -+    expand = not isinstance(c, list)\n>      -     p = subprocess.Popen(c, stdout=subprocess.PIPE, shell=expand)\n>      -     pipe = p.stdout\n>      +@@\n>            val = pipe.readlines()\n>            if pipe.close() or p.wait():\n>                die('Command failed: %s' % str(c))\n>      @@ -203,28 +101,6 @@\n>            # return code will be 1 in either case\n>            if err.find(\"Invalid option\") >= 0:\n>                return False\n>      -@@\n>      -     return True\n>      -\n>      - def system(cmd, ignore_error=False):\n>      --    expand = isinstance(cmd,basestring)\n>      -+    expand = not isinstance(cmd, list)\n>      -     if verbose:\n>      -         sys.stderr.write(\"executing %s\\n\" % str(cmd))\n>      -     retcode = subprocess.call(cmd, shell=expand)\n>      -@@\n>      -     return retcode\n>      -\n>      - def p4_system(cmd):\n>      --    \"\"\"Specifically invoke p4 as the system command. \"\"\"\n>      -+    \"\"\" Specifically invoke p4 as the system command.\n>      -+    \"\"\"\n>      -     real_cmd = p4_build_cmd(cmd)\n>      --    expand = isinstance(real_cmd, basestring)\n>      -+    expand = not isinstance(real_cmd, list)\n>      -     retcode = subprocess.call(real_cmd, shell=expand)\n>      -     if retcode:\n>      -         raise CalledProcessError(retcode, real_cmd)\n>       @@\n>            return int(results[0]['change'])\n>\n>      @@ -234,7 +110,7 @@\n>       -       results.\"\"\"\n>       +    \"\"\" Returns information about the requested P4 change list.\n>       +\n>      -+        Data returns is not string encoded (returned as bytes)\n>      ++        Data returned is not string encoded (returned as bytes)\n>       +    \"\"\"\n>       +    # Make sure it returns a valid result by checking for\n>       +    #   the presence of field \"time\".  Return a dict of the\n>      @@ -261,218 +137,29 @@\n>            if \"time\" not in d:\n>                die(\"p4 describe -s %d returned no \\\"time\\\": %s\" % (change, str(d)))\n>\n>      -+    # Convert depotFile(X) to be UTF-8 encoded, as this is what GIT\n>      -+    # requires. This will also allow us to encode the rest of the text\n>      -+    # at the same time to simplify textual processing later.\n>      ++    # Do not convert 'depotFile(X)' or 'path' to be UTF-8 encoded, however\n>      ++    # cast as_string() the rest of the text.\n>       +    keys=d.keys()\n>       +    for key in keys:\n>       +        if key.startswith('depotFile'):\n>      -+            d[key]=d[key] #DepotPath(d[key])\n>      ++            d[key]=d[key]\n>       +        elif key == 'path':\n>      -+            d[key]=d[key] #DepotPath(d[key])\n>      ++            d[key]=d[key]\n>       +        else:\n>       +            d[key] = as_string(d[key])\n>       +\n>            return d\n>\n>      --#\n>      --# Canonicalize the p4 type and return a tuple of the\n>      --# base type, plus any modifiers.  See \"p4 help filetypes\"\n>      --# for a list and explanation.\n>      --#\n>      - def split_p4_type(p4type):\n>      --\n>      -+    \"\"\" Canonicalize the p4 type and return a tuple of the\n>      -+        base type, plus any modifiers.  See \"p4 help filetypes\"\n>      -+        for a list and explanation.\n>      -+    \"\"\"\n>      -     p4_filetypes_historical = {\n>      -         \"ctempobj\": \"binary+Sw\",\n>      -         \"ctext\": \"text+C\",\n>      -@@\n>      -         mods = s[1]\n>      -     return (base, mods)\n>      -\n>      --#\n>      --# return the raw p4 type of a file (text, text+ko, etc)\n>      --#\n>      - def p4_type(f):\n>      -+    \"\"\" return the raw p4 type of a file (text, text+ko, etc)\n>      -+    \"\"\"\n>      -     results = p4CmdList([\"fstat\", \"-T\", \"headType\", wildcard_encode(f)])\n>      -     return results[0]['headType']\n>      -\n>      --#\n>      --# Given a type base and modifier, return a regexp matching\n>      --# the keywords that can be expanded in the file\n>      --#\n>      - def p4_keywords_regexp_for_type(base, type_mods):\n>      -+    \"\"\" Given a type base and modifier, return a regexp matching\n>      -+        the keywords that can be expanded in the file\n>      -+    \"\"\"\n>      -     if base in (\"text\", \"unicode\", \"binary\"):\n>      -         kwords = None\n>      -         if \"ko\" in type_mods:\n>      -@@\n>      -     else:\n>      -         return None\n>      -\n>      --#\n>      --# Given a file, return a regexp matching the possible\n>      --# RCS keywords that will be expanded, or None for files\n>      --# with kw expansion turned off.\n>      --#\n>      - def p4_keywords_regexp_for_file(file):\n>      -+    \"\"\" Given a file, return a regexp matching the possible\n>      -+        RCS keywords that will be expanded, or None for files\n>      -+        with kw expansion turned off.\n>      -+    \"\"\"\n>      -     if not os.path.exists(file):\n>      -         return None\n>      -     else:\n>      -@@\n>      - # Return the set of all p4 labels\n>      - def getP4Labels(depotPaths):\n>      -     labels = set()\n>      --    if isinstance(depotPaths,basestring):\n>      -+    if not isinstance(depotPaths, list):\n>      -         depotPaths = [depotPaths]\n>      -\n>      -     for l in p4CmdList([\"labels\"] + [\"%s...\" % p for p in depotPaths]):\n>      -@@\n>      -\n>      -     return labels\n>      -\n>      --# Return the set of all git tags\n>      - def getGitTags():\n>      -+    \"\"\"Return the set of all git tags\"\"\"\n>      -     gitTags = set()\n>      -     for line in read_pipe_lines([\"git\", \"tag\"]):\n>      -         tag = line.strip()\n>      -@@\n>      -\n>      -     If the pattern is not matched, None is returned.\"\"\"\n>      -\n>      --    match = diffTreePattern().next().match(entry)\n>      -+    match = next(diffTreePattern()).match(entry)\n>      -     if match:\n>      -         return {\n>      -             'src_mode': match.group(1),\n>      -@@\n>      -     # otherwise False.\n>      -     return mode[-3:] == \"755\"\n>      -\n>      -+def encodeWithUTF8(path, verbose = False):\n>      -+    \"\"\" Ensure that the path is encoded as a UTF-8 string\n>      -+\n>      -+        Returns bytes(P3)/str(P2)\n>      -+    \"\"\"\n>      -+\n>      -+    if isunicode:\n>      -+        try:\n>      -+            if isinstance(path, unicode):\n>      -+                # It is already unicode, cast it as a bytes\n>      -+                # that is encoded as utf-8.\n>      -+                return path.encode('utf-8', 'strict')\n>      -+            path.decode('ascii', 'strict')\n>      -+        except:\n>      -+            encoding = 'utf8'\n>      -+            if gitConfig('git-p4.pathEncoding'):\n>      -+                encoding = gitConfig('git-p4.pathEncoding')\n>      -+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n>      -+            if verbose:\n>      -+                print('\\nNOTE:Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, to_unicode(path)))\n>      -+    else:\n>      -+        try:\n>      -+            path.decode('ascii')\n>      -+        except:\n>      -+            encoding = 'utf8'\n>      -+            if gitConfig('git-p4.pathEncoding'):\n>      -+                encoding = gitConfig('git-p4.pathEncoding')\n>      -+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n>      -+            if verbose:\n>      -+                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n>      -+    return path\n>      -+\n>      - class P4Exception(Exception):\n>      -     \"\"\" Base class for exceptions from the p4 client \"\"\"\n>      -     def __init__(self, exit_code):\n>      -@@\n>      -     return isModeExec(src_mode) != isModeExec(dst_mode)\n>      -\n>      - def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n>      --        errors_as_exceptions=False):\n>      -+        errors_as_exceptions=False, encode_data=True):\n>      -+    \"\"\" Executes a P4 command:  'cmd' optionally passing 'stdin' to the command's\n>      -+        standard input via a temporary file with 'stdin_mode' mode.\n>      -+\n>      -+        Output from the command is optionally passed to the callback function 'cb'.\n>      -+        If 'cb' is None, the response from the command is parsed into a list\n>      -+        of resulting dictionaries. (For each block read from the process pipe.)\n>      -+\n>      -+        If 'skip_info' is true, information in a block read that has a code type of\n>      -+        'info' will be skipped.\n>      -\n>      --    if isinstance(cmd,basestring):\n>      -+        If 'errors_as_exceptions' is set to true (the default is false) the error\n>      -+        code returned from the execution will generate an exception.\n>      -+\n>      -+        If 'encode_data' is set to true (the default) the data that is returned\n>      -+        by this function will be passed through the \"as_string\" function.\n>      -+    \"\"\"\n>      -+\n>      -+    if not isinstance(cmd, list):\n>      -         cmd = \"-G \" + cmd\n>      -         expand = True\n>      -     else:\n>      -@@\n>      -     stdin_file = None\n>      -     if stdin is not None:\n>      -         stdin_file = tempfile.TemporaryFile(prefix='p4-stdin', mode=stdin_mode)\n>      --        if isinstance(stdin,basestring):\n>      -+        if not isinstance(stdin, list):\n>      -             stdin_file.write(stdin)\n>      -         else:\n>      -             for i in stdin:\n>      --                stdin_file.write(i + '\\n')\n>      -+                stdin_file.write(as_bytes(i) + b'\\n')\n>      -         stdin_file.flush()\n>      -         stdin_file.seek(0)\n>      -\n>      -@@\n>      -         while True:\n>      -             entry = marshal.load(p4.stdout)\n>      -             if skip_info:\n>      --                if 'code' in entry and entry['code'] == 'info':\n>      -+                if b'code' in entry and entry[b'code'] == b'info':\n>      -                     continue\n>      -             if cb is not None:\n>      -                 cb(entry)\n>      -             else:\n>      --                result.append(entry)\n>      -+                out = {}\n>      -+                for key, value in entry.items():\n>      -+                    out[as_string(key)] = (as_string(value) if encode_data else value)\n>      -+                result.append(out)\n>      -     except EOFError:\n>      -         pass\n>      -     exitCode = p4.wait()\n>      + #\n>       @@\n>            return result\n>\n>        def p4Cmd(cmd):\n>      -+    \"\"\" Executes a P4 command an returns the results in a dictionary\"\"\"\n>      ++    \"\"\" Executes a P4 command and returns the results in a dictionary\n>      ++    \"\"\"\n>            list = p4CmdList(cmd)\n>            result = {}\n>            for entry in list:\n>      -@@\n>      -     return values\n>      -\n>      - def gitBranchExists(branch):\n>      -+    \"\"\"Checks to see if a given branch exists in the git repo\"\"\"\n>      -     proc = subprocess.Popen([\"git\", \"rev-parse\", branch],\n>      -                             stderr=subprocess.PIPE, stdout=subprocess.PIPE);\n>      -     return proc.wait() == 0;\n>       @@\n>        _gitConfig = {}\n>\n>      @@ -490,29 +177,6 @@\n>            return _gitConfig[key]\n>\n>        def gitConfigBool(key):\n>      --    \"\"\"Return a bool, using git config --bool.  It is True only if the\n>      --       variable is set to true, and False if set to false or not present\n>      --       in the config.\"\"\"\n>      --\n>      -+    \"\"\" Return a bool, using git config --bool.  It is True only if the\n>      -+        variable is set to true, and False if set to false or not present\n>      -+        in the config.\n>      -+    \"\"\"\n>      -     if key not in _gitConfig:\n>      -         _gitConfig[key] = gitConfig(key, '--bool') == \"true\"\n>      -     return _gitConfig[key]\n>      -@@\n>      -             _gitConfig[key] = []\n>      -     return _gitConfig[key]\n>      -\n>      -+def gitConfigSet(key, value):\n>      -+    \"\"\" Set the git configuration key 'key' to 'value' for this session\n>      -+    \"\"\"\n>      -+    _gitConfig[key] = value\n>      -+\n>      - def p4BranchesInGit(branchesAreInRemotes=True):\n>      -     \"\"\"Find all the branches whose names start with \"p4/\", looking\n>      -        in remotes or heads as specified by the argument.  Return\n>       @@\n>            cmd = [ \"git\", \"rev-parse\", \"--symbolic\", \"--verify\", branch ]\n>            p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n>      @@ -521,34 +185,6 @@\n>            if p.returncode:\n>                return False\n>            # expect exactly one line of output: the branch name\n>      -@@\n>      -     branches = p4BranchesInGit()\n>      -     # map from depot-path to branch name\n>      -     branchByDepotPath = {}\n>      --    for branch in branches.keys():\n>      -+    for branch in list(branches.keys()):\n>      -         tip = branches[branch]\n>      -         log = extractLogMessageFromGitCommit(tip)\n>      -         settings = extractSettingsGitLog(log)\n>      -@@\n>      -             system(\"git update-ref %s %s\" % (remoteHead, originHead))\n>      -\n>      - def originP4BranchesExist():\n>      --        return gitBranchExists(\"origin\") or gitBranchExists(\"origin/p4\") or gitBranchExists(\"origin/p4/master\")\n>      -+    \"\"\"Checks if origin/p4/master exists\"\"\"\n>      -+    return gitBranchExists(\"origin\") or gitBranchExists(\"origin/p4\") or gitBranchExists(\"origin/p4/master\")\n>      -\n>      -\n>      - def p4ParseNumericChangeRange(parts):\n>      -@@\n>      -     changes = sorted(changes)\n>      -     return changes\n>      -\n>      --def p4PathStartsWith(path, prefix):\n>      -+def p4PathStartsWith(path, prefix, verbose = False):\n>      -     # This method tries to remedy a potential mixed-case issue:\n>      -     #\n>      -     # If UserA adds  //depot/DirA/file1\n>       @@\n>            #\n>            # we may or may not have a problem. If you have core.ignorecase=true,\n>      @@ -574,15 +210,6 @@\n>\n>        def getClientSpec():\n>            \"\"\"Look at the p4 client spec, create a View() object that contains\n>      -@@\n>      -     client_name = entry[\"Client\"]\n>      -\n>      -     # just the keys that start with \"View\"\n>      --    view_keys = [ k for k in entry.keys() if k.startswith(\"View\") ]\n>      -+    view_keys = [ k for k in list(entry.keys()) if k.startswith(\"View\") ]\n>      -\n>      -     # hold this new View\n>      -     view = View(client_name)\n>       @@\n>            # Cannot have * in a filename in windows; untested as to\n>            # what p4 would do in such a case.\n>      @@ -626,45 +253,16 @@\n>                    os.remove(contentFile)\n>                    die('git-lfs pointer command failed. Did you install the extension?')\n>       @@\n>      -         else:\n>      -             return LargeFileSystem.processContent(self, git_mode, relPath, contents)\n>      -\n>      --class Command:\n>      -+class Command(object):\n>      -     delete_actions = ( \"delete\", \"move/delete\", \"purge\" )\n>      -     add_actions = ( \"add\", \"branch\", \"move/add\" )\n>      -\n>      -@@\n>      -             setattr(self, attr, value)\n>      -         return getattr(self, attr)\n>      -\n>      --class P4UserMap:\n>      -+class P4UserMap(object):\n>      -     def __init__(self):\n>      -         self.userMapFromPerforceServer = False\n>      -         self.myP4UserId = None\n>      -@@\n>      -             return True\n>      -\n>      -     def getUserCacheFilename(self):\n>      -+        \"\"\" Returns the filename of the username cache \"\"\"\n>      -         home = os.environ.get(\"HOME\", os.environ.get(\"USERPROFILE\"))\n>      --        return home + \"/.gitp4-usercache.txt\"\n>      -+        return os.path.join(home, \".gitp4-usercache.txt\")\n>      +         return os.path.join(home, \".gitp4-usercache.txt\")\n>\n>            def getUserMapFromPerforceServer(self):\n>       +        \"\"\" Creates the usercache from the data in P4.\n>       +        \"\"\"\n>      -+\n>                if self.userMapFromPerforceServer:\n>                    return\n>                self.users = {}\n>       @@\n>      -                 self.emails[email] = user\n>      -\n>      -         s = ''\n>      --        for (key, val) in self.users.items():\n>      -+        for (key, val) in list(self.users.items()):\n>      +         for (key, val) in list(self.users.items()):\n>                    s += \"%s\\t%s\\n\" % (key.expandtabs(1), val.expandtabs(1))\n>\n>       -        open(self.getUserCacheFilename(), \"wb\").write(s)\n>      @@ -674,7 +272,8 @@\n>                self.userMapFromPerforceServer = True\n>\n>            def loadUserMapFromCache(self):\n>      -+        \"\"\" Reads the P4 username to git email map \"\"\"\n>      ++        \"\"\" Reads the P4 username to git email map\n>      ++        \"\"\"\n>                self.users = {}\n>                self.userMapFromPerforceServer = False\n>                try:\n>      @@ -721,80 +320,6 @@\n>                    # cleanup our temporary file\n>                    os.unlink(outFileName)\n>                    print(\"Failed to strip RCS keywords in %s\" % file)\n>      -@@\n>      -                 break\n>      -         if not change_entry:\n>      -             die('Failed to decode output of p4 change -o')\n>      --        for key, value in change_entry.iteritems():\n>      -+        for key, value in list(change_entry.items()):\n>      -             if key.startswith('File'):\n>      -                 if 'depot-paths' in settings:\n>      -                     if not [p for p in settings['depot-paths']\n>      --                            if p4PathStartsWith(value, p)]:\n>      -+                            if p4PathStartsWith(value, p, self.verbose)]:\n>      -                         continue\n>      -                 else:\n>      --                    if not p4PathStartsWith(value, self.depotPath):\n>      -+                    if not p4PathStartsWith(value, self.depotPath, self.verbose):\n>      -                         continue\n>      -                 files_list.append(value)\n>      -                 continue\n>      -@@\n>      -             return True\n>      -\n>      -         while True:\n>      --            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \")\n>      -+            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \").lower() \\\n>      -+                .strip()[0]\n>      -             if response == 'y':\n>      -                 return True\n>      -             if response == 'n':\n>      -@@\n>      -     def applyCommit(self, id):\n>      -         \"\"\"Apply one commit, return True if it succeeded.\"\"\"\n>      -\n>      --        print(\"Applying\", read_pipe([\"git\", \"show\", \"-s\",\n>      --                                     \"--format=format:%h %s\", id]))\n>      -+        print((\"Applying\", read_pipe([\"git\", \"show\", \"-s\",\n>      -+                                     \"--format=format:%h %s\", id])))\n>      -\n>      -         (p4User, gitEmail) = self.p4UserForCommit(id)\n>      -\n>      -@@\n>      -                     # disable the read-only bit on windows.\n>      -                     if self.isWindows and file not in editedFiles:\n>      -                         os.chmod(file, stat.S_IWRITE)\n>      --                    self.patchRCSKeywords(file, kwfiles[file])\n>      --                    fixed_rcs_keywords = True\n>      -+\n>      -+                    try:\n>      -+                        self.patchRCSKeywords(file, kwfiles[file])\n>      -+                        fixed_rcs_keywords = True\n>      -+                    except:\n>      -+                        # We are throwing an exception, undo all open edits\n>      -+                        for f in editedFiles:\n>      -+                            p4_revert(f)\n>      -+                        raise\n>      -+            else:\n>      -+                # They do not have attemptRCSCleanup set, this might be the fail point\n>      -+                # Check to see if the file has RCS keywords and suggest setting the property.\n>      -+                for file in editedFiles | filesToDelete:\n>      -+                    if p4_keywords_regexp_for_file(file) != None:\n>      -+                        print(\"At least one file in this commit has RCS Keywords that may be causing problems. \")\n>      -+                        print(\"Consider:\\ngit config git-p4.attemptRCSCleanup true\")\n>      -+                        break\n>      -\n>      -             if fixed_rcs_keywords:\n>      -                 print(\"Retrying the patch with RCS keywords cleaned up\")\n>      -@@\n>      -             p4_delete(f)\n>      -\n>      -         # Set/clear executable bits\n>      --        for f in filesToChangeExecBit.keys():\n>      -+        for f in list(filesToChangeExecBit.keys()):\n>      -             mode = filesToChangeExecBit[f]\n>      -             setP4ExecBit(f, mode)\n>      -\n>       @@\n>                tmpFile = os.fdopen(handle, \"w+b\")\n>                if self.isWindows:\n>      @@ -815,179 +340,6 @@\n>\n>                        if update_shelve:\n>                            p4_write_pipe(['shelve', '-r', '-i'], submitTemplate)\n>      -@@\n>      -                 if verbose:\n>      -                     print(\"created p4 label for tag %s\" % name)\n>      -\n>      -+    def run_hook(self, hook_name, args = []):\n>      -+        \"\"\" Runs a hook if it is found.\n>      -+\n>      -+            Returns NONE if the hook does not exist\n>      -+            Returns TRUE if the exit code is 0, FALSE for a non-zero exit code.\n>      -+        \"\"\"\n>      -+        hook_file = self.find_hook(hook_name)\n>      -+        if hook_file == None:\n>      -+            if self.verbose:\n>      -+                print(\"Skipping hook: %s\" % hook_name)\n>      -+            return None\n>      -+\n>      -+        if self.verbose:\n>      -+            print(\"hooks_path = %s \" % hooks_path)\n>      -+            print(\"hook_file = %s \" % hook_file)\n>      -+\n>      -+        # Run the hook\n>      -+        # TODO - allow non-list format\n>      -+        cmd = [hook_file] + args\n>      -+        return subprocess.call(cmd) == 0\n>      -+\n>      -+    def find_hook(self, hook_name):\n>      -+        \"\"\" Locates the hook file for the given operating system.\n>      -+        \"\"\"\n>      -+        hooks_path = gitConfig(\"core.hooksPath\")\n>      -+        if len(hooks_path) <= 0:\n>      -+            hooks_path = os.path.join(os.environ.get(\"GIT_DIR\", \".git\"), \"hooks\")\n>      -+\n>      -+        # Look in the obvious place\n>      -+        hook_file = os.path.join(hooks_path, hook_name)\n>      -+        if os.path.isfile(hook_file) and os.access(hook_file, os.X_OK):\n>      -+            return hook_file\n>      -+\n>      -+        # if we are windows, we will also allow them to have the hooks have extensions\n>      -+        if (platform.system() == \"Windows\"):\n>      -+            for ext in ['.exe', '.bat', 'ps1']:\n>      -+                if os.path.isfile(hook_file + ext) and os.access(hook_file + ext, os.X_OK):\n>      -+                    return hook_file + ext\n>      -+\n>      -+        # We didn't find the file\n>      -+        return None\n>      -+\n>      -+\n>      -+\n>      -     def run(self, args):\n>      -         if len(args) == 0:\n>      -             self.master = currentGitBranch()\n>      -@@\n>      -             self.clientSpecDirs = getClientSpec()\n>      -\n>      -         # Check for the existence of P4 branches\n>      --        branchesDetected = (len(p4BranchesInGit().keys()) > 1)\n>      -+        branchesDetected = (len(list(p4BranchesInGit().keys())) > 1)\n>      -\n>      -         if self.useClientSpec and not branchesDetected:\n>      -             # all files are relative to the client spec\n>      -@@\n>      -             sys.exit(\"number of commits (%d) must match number of shelved changelist (%d)\" %\n>      -                      (len(commits), num_shelves))\n>      -\n>      --        hooks_path = gitConfig(\"core.hooksPath\")\n>      --        if len(hooks_path) <= 0:\n>      --            hooks_path = os.path.join(os.environ.get(\"GIT_DIR\", \".git\"), \"hooks\")\n>      --\n>      --        hook_file = os.path.join(hooks_path, \"p4-pre-submit\")\n>      --        if os.path.isfile(hook_file) and os.access(hook_file, os.X_OK) and subprocess.call([hook_file]) != 0:\n>      -+        rtn = self.run_hook(\"p4-pre-submit\")\n>      -+        if rtn == False:\n>      -             sys.exit(1)\n>      -\n>      -         #\n>      -@@\n>      -         last = len(commits) - 1\n>      -         for i, commit in enumerate(commits):\n>      -             if self.dry_run:\n>      --                print(\" \", read_pipe([\"git\", \"show\", \"-s\",\n>      --                                      \"--format=format:%h %s\", commit]))\n>      -+                print((\" \", read_pipe([\"git\", \"show\", \"-s\",\n>      -+                                      \"--format=format:%h %s\", commit])))\n>      -                 ok = True\n>      -             else:\n>      -                 ok = self.applyCommit(commit)\n>      -@@\n>      -                         if self.conflict_behavior == \"ask\":\n>      -                             print(\"What do you want to do?\")\n>      -                             response = raw_input(\"[s]kip this commit but apply\"\n>      --                                                 \" the rest, or [q]uit? \")\n>      -+                                                 \" the rest, or [q]uit? \").lower().strip()[0]\n>      -                             if not response:\n>      -                                 continue\n>      -                         elif self.conflict_behavior == \"skip\":\n>      -@@\n>      -                         star = \"*\"\n>      -                     else:\n>      -                         star = \" \"\n>      --                    print(star, read_pipe([\"git\", \"show\", \"-s\",\n>      --                                           \"--format=format:%h %s\",  c]))\n>      -+                    print((star, read_pipe([\"git\", \"show\", \"-s\",\n>      -+                                           \"--format=format:%h %s\",  c])))\n>      -                 print(\"You will have to do 'git p4 sync' and rebase.\")\n>      -\n>      -         if gitConfigBool(\"git-p4.exportLabels\"):\n>      -@@\n>      -     # (\"-//depot/A/...\" becomes \"/depot/A/...\" after option parsing)\n>      -     parser.values.cloneExclude += [\"/\" + re.sub(r\"\\.\\.\\.$\", \"\", value)]\n>      -\n>      -+\n>      - class P4Sync(Command, P4UserMap):\n>      -\n>      -     def __init__(self):\n>      -@@\n>      -         self.knownBranches = {}\n>      -         self.initialParents = {}\n>      -\n>      --        self.tz = \"%+03d%02d\" % (- time.timezone / 3600, ((- time.timezone % 3600) / 60))\n>      -+        self.tz = \"%+03d%02d\" % (- time.timezone // 3600, ((- time.timezone % 3600) // 60))\n>      -         self.labels = {}\n>      -\n>      -     # Force a checkpoint in fast-import and wait for it to finish\n>      -@@\n>      -     def isPathWanted(self, path):\n>      -         for p in self.cloneExclude:\n>      -             if p.endswith(\"/\"):\n>      --                if p4PathStartsWith(path, p):\n>      -+                if p4PathStartsWith(path, p, self.verbose):\n>      -                     return False\n>      -             # \"-//depot/file1\" without a trailing \"/\" should only exclude \"file1\", but not \"file111\" or \"file1_dir/file2\"\n>      -             elif path.lower() == p.lower():\n>      -                 return False\n>      -         for p in self.depotPaths:\n>      --            if p4PathStartsWith(path, p):\n>      -+            if p4PathStartsWith(path, p, self.verbose):\n>      -                 return True\n>      -         return False\n>      -\n>      -     def extractFilesFromCommit(self, commit, shelved=False, shelved_cl = 0):\n>      -+        \"\"\" Generates the list of files to be added in this git commit.\n>      -+\n>      -+            commit     = Unicode[] - data read from the P4 commit\n>      -+            shelved    = Bool      - Is the P4 commit flagged as being shelved.\n>      -+            shelved_cl = Unicode   - Numeric string with the changelist number.\n>      -+        \"\"\"\n>      -         files = []\n>      -         fnum = 0\n>      -         while \"depotFile%s\" % fnum in commit:\n>      -@@\n>      -             path = self.clientSpecDirs.map_in_client(path)\n>      -             if self.detectBranches:\n>      -                 for b in self.knownBranches:\n>      --                    if p4PathStartsWith(path, b + \"/\"):\n>      -+                    if p4PathStartsWith(path, b + \"/\", self.verbose):\n>      -                         path = path[len(b)+1:]\n>      -\n>      -         elif self.keepRepoPath:\n>      -@@\n>      -             # //depot/; just look at first prefix as they all should\n>      -             # be in the same depot.\n>      -             depot = re.sub(\"^(//[^/]+/).*\", r'\\1', prefixes[0])\n>      --            if p4PathStartsWith(path, depot):\n>      -+            if p4PathStartsWith(path, depot, self.verbose):\n>      -                 path = path[len(depot):]\n>      -\n>      -         else:\n>      -             for p in prefixes:\n>      --                if p4PathStartsWith(path, p):\n>      -+                if p4PathStartsWith(path, p, self.verbose):\n>      -                     path = path[len(p):]\n>      -                     break\n>      -\n>       @@\n>                return path\n>\n>      @@ -1002,19 +354,6 @@\n>\n>                if self.clientSpecDirs:\n>                    files = self.extractFilesFromCommit(commit)\n>      -@@\n>      -             else:\n>      -                 relPath = self.stripRepoPath(path, self.depotPaths)\n>      -\n>      --            for branch in self.knownBranches.keys():\n>      -+            for branch in list(self.knownBranches.keys()):\n>      -                 # add a trailing slash so that a commit into qt/4.2foo\n>      -                 # doesn't end up in qt/4.2, e.g.\n>      --                if p4PathStartsWith(relPath, branch + \"/\"):\n>      -+                if p4PathStartsWith(relPath, branch + \"/\", self.verbose):\n>      -                     if branch not in branches:\n>      -                         branches[branch] = []\n>      -                     branches[branch].append(file)\n>       @@\n>                return branches\n>\n>      @@ -1031,18 +370,6 @@\n>       +            self.gitStreamBytes.write(d)\n>                self.gitStream.write('\\n')\n>\n>      --    def encodeWithUTF8(self, path):\n>      --        try:\n>      --            path.decode('ascii')\n>      --        except:\n>      --            encoding = 'utf8'\n>      --            if gitConfig('git-p4.pathEncoding'):\n>      --                encoding = gitConfig('git-p4.pathEncoding')\n>      --            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n>      --            if self.verbose:\n>      --                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n>      --        return path\n>      --\n>       -    # output one file from the P4 stream\n>       -    # - helper for streamP4Files\n>       -\n>      @@ -1053,18 +380,13 @@\n>       +            contents should be a bytes (bytes)\n>       +        \"\"\"\n>                relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n>      --        relPath = self.encodeWithUTF8(relPath)\n>      -+        relPath = encodeWithUTF8(relPath, self.verbose)\n>      +         relPath = encodeWithUTF8(relPath, self.verbose)\n>                if verbose:\n>      -             if 'fileSize' in self.stream_file:\n>      +@@\n>                        size = int(self.stream_file['fileSize'])\n>                    else:\n>                        size = 0 # deleted files don't get a fileSize apparently\n>      --            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size/1024/1024))\n>      -+            #if isunicode:\n>      -+            #    sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), to_unicode(relPath), size//1024//1024))\n>      -+            #else:\n>      -+            #    sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), relPath, size//1024//1024))\n>      +-            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size//1024//1024))\n>       +            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), as_string(relPath), size//1024//1024))\n>                    sys.stdout.flush()\n>\n>      @@ -1100,15 +422,6 @@\n>\n>                if self.largeFileSystem:\n>       @@\n>      -\n>      -     def streamOneP4Deletion(self, file):\n>      -         relPath = self.stripRepoPath(file['path'], self.branchPrefixes)\n>      --        relPath = self.encodeWithUTF8(relPath)\n>      -+        relPath = encodeWithUTF8(relPath, self.verbose)\n>      -         if verbose:\n>      -             sys.stdout.write(\"delete %s\\n\" % relPath)\n>      -             sys.stdout.flush()\n>      -@@\n>                if self.largeFileSystem and self.largeFileSystem.isLargeFile(relPath):\n>                    self.largeFileSystem.removeLargeFile(relPath)\n>\n>      @@ -1133,13 +446,6 @@\n>\n>                if not err and 'fileSize' in self.stream_file:\n>                    required_bytes = int((4 * int(self.stream_file[\"fileSize\"])) - calcDiskFree())\n>      -             if required_bytes > 0:\n>      -                 err = 'Not enough space left on %s! Free at least %i MB.' % (\n>      --                    os.getcwd(), required_bytes/1024/1024\n>      -+                    os.getcwd(), required_bytes//1024//1024\n>      -                 )\n>      -\n>      -         if err:\n>       @@\n>                    # ignore errors, but make sure it exits first\n>                    self.importProcess.wait()\n>      @@ -1155,12 +461,10 @@\n>                    self.streamOneP4File(self.stream_file, self.stream_contents)\n>                    self.stream_file = {}\n>       @@\n>      -\n>                # pick up the new file information... for the\n>                # 'data' field we need to append to our array\n>      --        for k in marshalled.keys():\n>      +         for k in list(marshalled.keys()):\n>       -            if k == 'data':\n>      -+        for k in list(marshalled.keys()):\n>       +            if k == b'data':\n>                        if 'streamContentSize' not in self.stream_file:\n>                            self.stream_file['streamContentSize'] = 0\n>      @@ -1178,12 +482,10 @@\n>                if (verbose and\n>                    'streamContentSize' in self.stream_file and\n>       @@\n>      -             'depotFile' in self.stream_file):\n>                    size = int(self.stream_file[\"fileSize\"])\n>                    if size > 0:\n>      --                progress = 100*self.stream_file['streamContentSize']/size\n>      --                sys.stdout.write('\\r%s %d%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size/1024/1024)))\n>      -+                progress = 100.0*self.stream_file['streamContentSize']/size\n>      +                 progress = 100.0*self.stream_file['streamContentSize']/size\n>      +-                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size//1024//1024)))\n>       +                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (path_as_string(self.stream_file['depotFile']), progress, int(size//1024//1024)))\n>                        sys.stdout.flush()\n>\n>      @@ -1227,24 +529,6 @@\n>\n>                if verbose:\n>       @@\n>      -\n>      -         gitStream.write(\"tagger %s\\n\" % tagger)\n>      -\n>      --        print(\"labelDetails=\",labelDetails)\n>      -+        print((\"labelDetails=\",labelDetails))\n>      -         if 'Description' in labelDetails:\n>      -             description = labelDetails['Description']\n>      -         else:\n>      -@@\n>      -         if not self.branchPrefixes:\n>      -             return True\n>      -         hasPrefix = [p for p in self.branchPrefixes\n>      --                        if p4PathStartsWith(path, p)]\n>      -+                        if p4PathStartsWith(path, p, self.verbose)]\n>      -         if not hasPrefix and self.verbose:\n>      -             print('Ignoring file outside of prefix: {0}'.format(path))\n>      -         return hasPrefix\n>      -@@\n>                        .format(details['change']))\n>                    return\n>\n>      @@ -1307,58 +591,6 @@\n>\n>                if len(parent) > 0:\n>                    if self.verbose:\n>      -@@\n>      -             self.labels[newestChange] = [output, revisions]\n>      -\n>      -         if self.verbose:\n>      --            print(\"Label changes: %s\" % self.labels.keys())\n>      -+            print(\"Label changes: %s\" % list(self.labels.keys()))\n>      -\n>      -     # Import p4 labels as git tags. A direct mapping does not\n>      -     # exist, so assume that if all the files are at the same revision\n>      -@@\n>      -                 source = paths[0]\n>      -                 destination = paths[1]\n>      -                 ## HACK\n>      --                if p4PathStartsWith(source, self.depotPaths[0]) and p4PathStartsWith(destination, self.depotPaths[0]):\n>      -+                if p4PathStartsWith(source, self.depotPaths[0], self.verbose) and p4PathStartsWith(destination, self.depotPaths[0], self.verbose):\n>      -                     source = source[len(self.depotPaths[0]):-4]\n>      -                     destination = destination[len(self.depotPaths[0]):-4]\n>      -\n>      -@@\n>      -\n>      -     def getBranchMappingFromGitBranches(self):\n>      -         branches = p4BranchesInGit(self.importIntoRemotes)\n>      --        for branch in branches.keys():\n>      -+        for branch in list(branches.keys()):\n>      -             if branch == \"master\":\n>      -                 branch = \"main\"\n>      -             else:\n>      -@@\n>      -             self.updateOptionDict(description)\n>      -\n>      -             if not self.silent:\n>      --                sys.stdout.write(\"\\rImporting revision %s (%s%%)\" % (change, cnt * 100 / len(changes)))\n>      -+                sys.stdout.write(\"\\rImporting revision %s (%4.1f%%)\" % (change, cnt * 100 / len(changes)))\n>      -                 sys.stdout.flush()\n>      -             cnt = cnt + 1\n>      -\n>      -             try:\n>      -                 if self.detectBranches:\n>      -                     branches = self.splitFilesIntoBranches(description)\n>      --                    for branch in branches.keys():\n>      -+                    for branch in list(branches.keys()):\n>      -                         ## HACK  --hwn\n>      -                         branchPrefix = self.depotPaths[0] + branch + \"/\"\n>      -                         self.branchPrefixes = [ branchPrefix ]\n>      -@@\n>      -                 sys.exit(1)\n>      -\n>      -     def sync_origin_only(self):\n>      -+        \"\"\" Ensures that the origin has been synchronized if one is set \"\"\"\n>      -         if self.syncWithOrigin:\n>      -             self.hasOrigin = originP4BranchesExist()\n>      -             if self.hasOrigin:\n>       @@\n>                        system(\"git fetch origin\")\n>\n>      @@ -1439,61 +671,6 @@\n>            def closeStreams(self):\n>                self.gitStream.close()\n>       @@\n>      -                 if short in branches:\n>      -                     self.p4BranchesInGit = [ short ]\n>      -             else:\n>      --                self.p4BranchesInGit = branches.keys()\n>      -+                self.p4BranchesInGit = list(branches.keys())\n>      -\n>      -             if len(self.p4BranchesInGit) > 1:\n>      -                 if not self.silent:\n>      -                     print(\"Importing from/into multiple branches\")\n>      -                 self.detectBranches = True\n>      --                for branch in branches.keys():\n>      -+                for branch in list(branches.keys()):\n>      -                     self.initialParents[self.refPrefix + branch] = \\\n>      -                         branches[branch]\n>      -\n>      -@@\n>      -                                  help=\"where to leave result of the clone\"),\n>      -             optparse.make_option(\"--bare\", dest=\"cloneBare\",\n>      -                                  action=\"store_true\", default=False),\n>      -+            optparse.make_option(\"--encoding\", dest=\"setPathEncoding\",\n>      -+                                 action=\"store\", default=None,\n>      -+                                 help=\"Sets the path encoding for this depot\")\n>      -         ]\n>      -         self.cloneDestination = None\n>      -         self.needsGit = False\n>      -         self.cloneBare = False\n>      -+        self.setPathEncoding = None\n>      -\n>      -     def defaultDestination(self, args):\n>      -+        \"\"\"Returns the last path component as the default git\n>      -+        repository directory name\"\"\"\n>      -         ## TODO: use common prefix of args?\n>      -         depotPath = args[0]\n>      -         depotDir = re.sub(\"(@[^@]*)$\", \"\", depotPath)\n>      -         depotDir = re.sub(\"(#[^#]*)$\", \"\", depotDir)\n>      -         depotDir = re.sub(r\"\\.\\.\\.$\", \"\", depotDir)\n>      -         depotDir = re.sub(r\"/$\", \"\", depotDir)\n>      --        return os.path.split(depotDir)[1]\n>      -+        return depotDir.split('/')[-1]\n>      -\n>      -     def run(self, args):\n>      -         if len(args) < 1:\n>      -@@\n>      -\n>      -         depotPaths = args\n>      -\n>      -+        # If we have an encoding provided, ignore what may already exist\n>      -+        # in the registry. This will ensure we show the displayed values\n>      -+        # using the correct encoding.\n>      -+        if self.setPathEncoding:\n>      -+            gitConfigSet(\"git-p4.pathEncoding\", self.setPathEncoding)\n>      -+\n>      -+        # If more than 1 path element is supplied, the last element\n>      -+        # is the clone destination.\n>      -         if not self.cloneDestination and len(depotPaths) > 1:\n>                    self.cloneDestination = depotPaths[-1]\n>                    depotPaths = depotPaths[:-1]\n>\n>      @@ -1512,177 +689,3 @@\n>\n>                if not os.path.exists(self.cloneDestination):\n>                    os.makedirs(self.cloneDestination)\n>      -@@\n>      -         if retcode:\n>      -             raise CalledProcessError(retcode, init_cmd)\n>      -\n>      -+        # Set the encoding if it was provided command line\n>      -+        if self.setPathEncoding:\n>      -+            init_cmd= [\"git\", \"config\", \"git-p4.pathEncoding\", self.setPathEncoding]\n>      -+            retcode = subprocess.call(init_cmd)\n>      -+            if retcode:\n>      -+                raise CalledProcessError(retcode, init_cmd)\n>      -+\n>      -         if not P4Sync.run(self, depotPaths):\n>      -             return False\n>      -\n>      -@@\n>      -             to find the P4 commit we are based on, and the depot-paths.\n>      -         \"\"\"\n>      -\n>      --        for parent in (range(65535)):\n>      -+        for parent in (list(range(65535))):\n>      -             log = extractLogMessageFromGitCommit(\"{0}^{1}\".format(starting_point, parent))\n>      -             settings = extractSettingsGitLog(log)\n>      -             if 'change' in settings:\n>      -@@\n>      -             print(\"%s <= %s (%s)\" % (branch, \",\".join(settings[\"depot-paths\"]), settings[\"change\"]))\n>      -         return True\n>      -\n>      -+class Py23File():\n>      -+    \"\"\" Python2/3 Unicode File Wrapper\n>      -+    \"\"\"\n>      -+\n>      -+    stream_handle = None\n>      -+    verbose       = False\n>      -+    debug_handle  = None\n>      -+\n>      -+    def __init__(self, stream_handle, verbose = False,\n>      -+                 debug_handle = None):\n>      -+        \"\"\" Create a Python3 compliant Unicode to Byte String\n>      -+            Windows compatible wrapper\n>      -+\n>      -+            stream_handle = the underlying file-like handle\n>      -+            verbose       = Boolean if content should be echoed\n>      -+            debug_handle  = A file-like handle data is duplicately written to\n>      -+        \"\"\"\n>      -+        self.stream_handle = stream_handle\n>      -+        self.verbose       = verbose\n>      -+        self.debug_handle  = debug_handle\n>      -+\n>      -+    def write(self, utf8string):\n>      -+        \"\"\" Writes the utf8 encoded string to the underlying\n>      -+            file stream\n>      -+        \"\"\"\n>      -+        self.stream_handle.write(as_bytes(utf8string))\n>      -+        if self.verbose:\n>      -+            sys.stderr.write(\"Stream Output: %s\" % utf8string)\n>      -+            sys.stderr.flush()\n>      -+        if self.debug_handle:\n>      -+            self.debug_handle.write(as_bytes(utf8string))\n>      -+\n>      -+    def read(self, size = None):\n>      -+        \"\"\" Reads int charcters from the underlying stream\n>      -+            and converts it to utf8.\n>      -+\n>      -+            Be aware, the size value is for reading the underlying\n>      -+            bytes so the value may be incorrect. Usage of the size\n>      -+            value is discouraged.\n>      -+        \"\"\"\n>      -+        if size == None:\n>      -+            return as_string(self.stream_handle.read())\n>      -+        else:\n>      -+            return as_string(self.stream_handle.read(size))\n>      -+\n>      -+    def readline(self):\n>      -+        \"\"\" Reads a line from the underlying byte stream\n>      -+            and converts it to utf8\n>      -+        \"\"\"\n>      -+        return as_string(self.stream_handle.readline())\n>      -+\n>      -+    def readlines(self, sizeHint = None):\n>      -+        \"\"\" Returns a list containing lines from the file converted to unicode.\n>      -+\n>      -+            sizehint - Optional. If the optional sizehint argument is\n>      -+            present, instead of reading up to EOF, whole lines totalling\n>      -+            approximately sizehint bytes are read.\n>      -+        \"\"\"\n>      -+        lines = self.stream_handle.readlines(sizeHint)\n>      -+        for i in range(0, len(lines)):\n>      -+            lines[i] = as_string(lines[i])\n>      -+        return lines\n>      -+\n>      -+    def close(self):\n>      -+        \"\"\" Closes the underlying byte stream \"\"\"\n>      -+        self.stream_handle.close()\n>      -+\n>      -+    def flush(self):\n>      -+        \"\"\" Flushes the underlying byte stream \"\"\"\n>      -+        self.stream_handle.flush()\n>      -+\n>      -+class DepotPath():\n>      -+    \"\"\" Describes a DepotPath or File\n>      -+    \"\"\"\n>      -+\n>      -+    raw_path = None\n>      -+    utf8_path = None\n>      -+    bytes_path = None\n>      -+\n>      -+    def __init__(self, path):\n>      -+        \"\"\" Creates a new DepotPath with the path encoded\n>      -+            with by the P4 repository\n>      -+        \"\"\"\n>      -+        raw_path = path\n>      -+\n>      -+    def raw():\n>      -+        \"\"\" Returns the path as it was originally found\n>      -+            in the P4 repository\n>      -+        \"\"\"\n>      -+        return raw_path\n>      -+\n>      -+    def startswith(self, prefix, start = None, end = None):\n>      -+        \"\"\" Return True if string starts with the prefix, otherwise\n>      -+            return False. prefix can also be a tuple of prefixes to\n>      -+            look for. With optional start, test string beginning at\n>      -+            that position. With optional end, stop comparing\n>      -+            string at that position.\n>      -+        \"\"\"\n>      -+        return raw_path.startswith(prefix, start, end)\n>      -+\n>      -+\n>      - class HelpFormatter(optparse.IndentedHelpFormatter):\n>      -     def __init__(self):\n>      -         optparse.IndentedHelpFormatter.__init__(self)\n>      -@@\n>      -\n>      - def main():\n>      -     if len(sys.argv[1:]) == 0:\n>      --        printUsage(commands.keys())\n>      -+        printUsage(list(commands.keys()))\n>      -         sys.exit(2)\n>      -\n>      -     cmdName = sys.argv[1]\n>      -@@\n>      -     except KeyError:\n>      -         print(\"unknown command %s\" % cmdName)\n>      -         print(\"\")\n>      --        printUsage(commands.keys())\n>      -+        printUsage(list(commands.keys()))\n>      -         sys.exit(2)\n>      -\n>      -     options = cmd.options\n>      -@@\n>      -                                    description = cmd.description,\n>      -                                    formatter = HelpFormatter())\n>      -\n>      --    (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n>      -+    try:\n>      -+        (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n>      -+    except:\n>      -+        parser.print_help()\n>      -+        raise\n>      -+\n>      -     global verbose\n>      -     verbose = cmd.verbose\n>      -     if cmd.needsGit:\n>      -@@\n>      -                         chdir(cdup);\n>      -\n>      -         if not isValidGitDir(cmd.gitdir):\n>      --            if isValidGitDir(cmd.gitdir + \"/.git\"):\n>      --                cmd.gitdir += \"/.git\"\n>      -+            if isValidGitDir(os.path.join(cmd.gitdir, \".git\")):\n>      -+                cmd.gitdir = os.path.join(cmd.gitdir, \".git\")\n>      -             else:\n>      -                 die(\"fatal: cannot locate git repository at %s\" % cmd.gitdir)\n>      -\n>   -:  ---------- > 11:  883ef45ca5 git-p4: Added --encoding parameter to p4 clone\n>\n> --\n> gitgitgadget\n"},{"id":"387546","messageId":"20191205101935.GA315203@generichostname","threadId":"52262","inReplyTo":"40124269933691796ef57fd8df50f9e740d103b1.1575498577.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 01/11] git-p4: select p4 binary by operating-system","fromName":"Denton Liu","fromEmail":"liu.denton@gmail.com","sentAt":"2019-12-05T10:19:35Z","receivedAt":"2019-12-05T10:19:42Z","isPatch":true,"sender":{"key":"liu.denton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/9620836?v=4"},"body":"Hi Ben,\n\nFirst of all, as a note to you and possibly others, I don't have much\n(read: any) experience with git-p4. I do have experience with Python and\nhow git.git generally does things so I'll be reviewing from that\nperspective.\n\nOn Wed, Dec 04, 2019 at 10:29:27PM +0000, Ben Keene via GitGitGadget wrote:\n> From: Ben Keene <seraphire@gmail.com>\n> \n> Depending on the version of GIT and Python installed, the perforce program (p4) may not resolve on Windows without the program extension.\n\nNit: \"GIT\" should be written as \"Git\" when referring to the whole\nproject and \"git\" when referring to the command. Never in all-caps.\n\nAlso, please wrap your paragraphs at 72 characters. I'll say it once\nhere but it applies to your whole series.\n\n> \n> Check the operating system (platform.system) and if it is reporting that it is Windows, use the full filename of \"p4.exe\" instead of \"p4\"\n> \n> The original code unconditionally used \"p4\" as the binary filename.\n\nAs a rule of thumb, we want to state the problem first before we state\nwhat we did (and why). I'd move this paragraph up.\n\n> \n> This change is Python2 and Python3 compatible.\n> \n> Thanks to: Junio C Hamano <gitster@pobox.com> and  Denton Liu <liu.denton@gmail.com> for patiently explaining proper format for my submissions.\n\nI appreciate the credit but I don't think it's necessary. At _most_, you\ncould include the\n\n\tHelped-by: Junio C Hamano <gitster@pobox.com>\n\tHelped-by: Denton Liu <liu.denton@gmail.com>\n\ntags before your signoff but I don't think we've done anything to\nwarrant it.\n\n> \n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> (cherry picked from commit 9a3a5c4e6d29dbef670072a9605c7a82b3729434)\n\nYou should remove this line in all of your commits. The referenced\ncommit isn't public so the information isn't very useful. Also, try to\nnot include anything after your signoff so if this hypothetically were\nuseful information, you'd include it before your signoff.\n\nIf it's information that's ephemerally useful for current reviewers but\nnot for future readers of your commit in the log message, you can\ninclude it after the three hyphens...\n\n> ---\nlike this and it won't be included as part of the log message.\n\n>  git-p4.py | 6 +++++-\n>  1 file changed, 5 insertions(+), 1 deletion(-)\n> \n> diff --git a/git-p4.py b/git-p4.py\n> index 60c73b6a37..b2ffbc057b 100755\n> --- a/git-p4.py\n> +++ b/git-p4.py\n> @@ -75,7 +75,11 @@ def p4_build_cmd(cmd):\n>      location. It means that hooking into the environment, or other configuration\n>      can be done more easily.\n>      \"\"\"\n> -    real_cmd = [\"p4\"]\n> +    # Look for the P4 binary\n\nI don't think this comment is necessary as the code itself is pretty\nself-explanatory.\n\n> +    if (platform.system() == \"Windows\"):\n> +        real_cmd = [\"p4.exe\"]    \n\nYou have trailing whitespace here. Try to run `git diff --check` before\ncommitting to ensure that you have no whitespace errors.\n\nThanks,\n\nDenton\n\n> +    else:\n> +        real_cmd = [\"p4\"]\n>  \n>      user = gitConfig(\"git-p4.user\")\n>      if len(user) > 0:\n> -- \n> gitgitgadget\n> \n"},{"id":"387547","messageId":"20191205102724.GB315203@generichostname","threadId":"52262","inReplyTo":"0ef2f56b04803cad2e60bf881e86d8bdd69463a6.1575498577.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 02/11] git-p4: change the expansion test from basestring to list","fromName":"Denton Liu","fromEmail":"liu.denton@gmail.com","sentAt":"2019-12-05T10:27:24Z","receivedAt":"2019-12-05T10:27:28Z","isPatch":true,"sender":{"key":"liu.denton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/9620836?v=4"},"body":"Hi Ben,\n\nOn Wed, Dec 04, 2019 at 10:29:28PM +0000, Ben Keene via GitGitGadget wrote:\n> From: Ben Keene <seraphire@gmail.com>\n> \n> Python 3+ handles strings differently than Python 2.7.\n\nDo you mean Python 3?\n\n> Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n\nPython 3.5 doesn't reach EOL until Q4 2020[1]. We should be testing\nthese changes under 3.5 to ensure that we're not accidentally\nintroducing stuff that's not backwards compatible.\n\n> \n> Change references to basestring in the isinstance tests to use list instead. This prepares the code to remove all references to basestring.\n> \n> The original code used basestring in a test to determine if a list or literal string was passed into 9 different functions.  This is used to determine if the shell should be evoked when calling subprocess methods.\n\nOnce again, I'd swap the above two paragraphs. Problem then solution.\n\nAlso, did you mean \"invoked\" instead of \"evoked\"?\n\n> \n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> (cherry picked from commit 5b1b1c145479b5d5fd242122737a3134890409e6)\n> ---\n>  git-p4.py | 18 +++++++++---------\n>  1 file changed, 9 insertions(+), 9 deletions(-)\n\nThe patch itself looks good, though.\n\n[1]: https://devguide.python.org/#branchstatus\n"},{"id":"387548","messageId":"20191205104056.GA1192079@generichostname","threadId":"52262","inReplyTo":"f0e658b984ca009c575368e661016f785922f970.1575498577.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 03/11] git-p4: add new helper functions for python3 conversion","fromName":"Denton Liu","fromEmail":"liu.denton@gmail.com","sentAt":"2019-12-05T10:40:56Z","receivedAt":"2019-12-05T10:41:01Z","isPatch":true,"sender":{"key":"liu.denton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/9620836?v=4"},"body":"On Wed, Dec 04, 2019 at 10:29:29PM +0000, Ben Keene via GitGitGadget wrote:\n> From: Ben Keene <seraphire@gmail.com>\n> \n> Python 3+ handles strings differently than Python 2.7.  Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n> \n> Change the existing unicode test add new support functions for python2-python3 support.\n> \n> Define the following variables:\n> - isunicode - a boolean variable that states if the version of python natively supports unicode (true) or not (false). This is true for Python3 and false for Python2.\n> - unicode - a type alias for the datatype that holds a unicode string.  It is assigned to a str under python 3 and the unicode type for Python2.\n> - bytes - a type alias for an array of bytes.  It is assigned the native bytes type for Python3 and str for Python2.\n> \n> Add the following new functions:\n> \n> - as_string(text) - A new function that will convert a byte array to a unicode (UTF-8) string under python 3.  Under python 2, this returns the string unchanged.\n> - as_bytes(text) - A new function that will convert a unicode string to a byte array under python 3.  Under python 2, this returns the string unchanged.\n> - to_unicode(text) - Converts a text string as Unicode(UTF-8) on both Python2 and Python3.\n> \n> Add a new function alias raw_input:\n> If raw_input does not exist (it was renamed to input in python 3) alias input as raw_input.\n> \n> The AS_STRING and AS_BYTES functions allow for modifying the code with a minimal amount of impact on Python2 support.  When a string is expected, the as_string() will be used to convert \"cast\" the incoming \"bytes\" to a string type. Conversely as_bytes() will be used to convert a \"string\" to a \"byte array\" type. Since Python2 overloads the datatype 'str' to serve both purposes, the Python2 versions of these function do not change the data, since the str functions as both a byte array and a string.\n\nHow come AS_STRING and AS_BYTES are all-caps here?\n\n> \n> basestring is removed since its only references are found in tests that were changed in the previous change list.\n> \n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> (cherry picked from commit 7921aeb3136b07643c1a503c2d9d8b5ada620356)\n> ---\n>  git-p4.py | 70 +++++++++++++++++++++++++++++++++++++++++++++++++++----\n>  1 file changed, 66 insertions(+), 4 deletions(-)\n> \n> diff --git a/git-p4.py b/git-p4.py\n> index 0f27996393..93dfd0920a 100755\n> --- a/git-p4.py\n> +++ b/git-p4.py\n> @@ -32,16 +32,78 @@\n>      unicode = unicode\n>  except NameError:\n>      # 'unicode' is undefined, must be Python 3\n> -    str = str\n> +    #\n> +    # For Python3 which is natively unicode, we will use \n> +    # unicode for internal information but all P4 Data\n> +    # will remain in bytes\n> +    isunicode = True\n>      unicode = str\n>      bytes = bytes\n> -    basestring = (str,bytes)\n> +\n> +    def as_string(text):\n> +        \"\"\"Return a byte array as a unicode string\"\"\"\n> +        if text == None:\n\nNit: use `text is None` instead. Actually, any time you're checking an\nobject to see if it's None, you should use `is` instead of `==` since\nthere's usually only one None reference.\n\n> +            return None\n> +        if isinstance(text, bytes):\n> +            return unicode(text, \"utf-8\")\n> +        else:\n> +            return text\n> +\n> +    def as_bytes(text):\n> +        \"\"\"Return a Unicode string as a byte array\"\"\"\n> +        if text == None:\n> +            return None\n> +        if isinstance(text, bytes):\n> +            return text\n> +        else:\n> +            return bytes(text, \"utf-8\")\n> +\n> +    def to_unicode(text):\n> +        \"\"\"Return a byte array as a unicode string\"\"\"\n> +        return as_string(text)    \n> +\n> +    def path_as_string(path):\n> +        \"\"\" Converts a path to the UTF8 encoded string \"\"\"\n> +        if isinstance(path, unicode):\n> +            return path\n> +        return encodeWithUTF8(path).decode('utf-8')\n> +    \n\nTrailing whitespace.\n\n>  else:\n>      # 'unicode' exists, must be Python 2\n> -    str = str\n> +    #\n> +    # We will treat the data as:\n> +    #   str   -> str\n> +    #   bytes -> str\n> +    # So for Python2 these functions are no-ops\n> +    # and will leave the data in the ambiguious\n> +    # string/bytes state\n> +    isunicode = False\n>      unicode = unicode\n>      bytes = str\n> -    basestring = basestring\n> +\n> +    def as_string(text):\n> +        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n\nI didn't mention this in earlier emails but it's been bothering me a\nlot: is there any reason why you write it as \"Python3\" vs. \"Python 3\"\nsometimes (and Python2 as well)? If there's no difference, then we\nshould probably stick to one variant in both the commit messages and in\nthe code. (I prefer the spaced variant.)\n\n> +        return text\n> +\n> +    def as_bytes(text):\n> +        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n> +        return text\n> +\n> +    def to_unicode(text):\n> +        \"\"\"Return a string as a unicode string\"\"\"\n> +        return text.decode('utf-8')\n> +    \n\nTrailing whitespace.\n\n> +    def path_as_string(path):\n> +        \"\"\" Converts a path to the UTF8 encoded bytes \"\"\"\n> +        return encodeWithUTF8(path)\n> +\n> +\n> + \n\nTrailing whitespace.\n\n> +# Check for raw_input support\n> +try:\n> +    raw_input\n> +except NameError:\n> +    raw_input = input\n>  \n>  try:\n>      from subprocess import CalledProcessError\n> -- \n> gitgitgadget\n> \n"},{"id":"387549","messageId":"20191205105034.GB1192079@generichostname","threadId":"52262","inReplyTo":"1bf7b073b047ca7625d0861b160a9602135f7baf.1575498578.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 05/11] git-p4: Add new functions in preparation of usage","fromName":"Denton Liu","fromEmail":"liu.denton@gmail.com","sentAt":"2019-12-05T10:50:34Z","receivedAt":"2019-12-05T10:50:42Z","isPatch":true,"sender":{"key":"liu.denton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/9620836?v=4"},"body":"> Subject: git-p4: Add new functions in preparation of usage\n\nNit: as a convention, you should lowercase the letter after the colon in\nthe subject. As in \"git-p4: add new functions...\"\n\nThis applies for other patches as well.\n\nOn Wed, Dec 04, 2019 at 10:29:31PM +0000, Ben Keene via GitGitGadget wrote:\n> From: Ben Keene <seraphire@gmail.com>\n> \n> This changelist is an intermediate submission for migrating the P4 support from Python2 to Python3. The code needs access to the encodeWithUTF8() for support of non-UTF8 filenames in the clone class as well as the sync class.\n> \n> Move the function encodeWithUTF8() from the P4Sync class to a stand-alone function.  This will allow other classes to use this function without instanciating the P4Sync class. Change the self.verbose reference to an optional method parameter. Update the existing references to this function to pass the self.verbose since it is no longer available on \"self\" since the function is no longer contained on the P4Sync class.\n\nHmmm, so does the patch before this not actually work since\nencodeWithUTF8() isn't defined yet? When you reroll this series, you\nshould swap the order of the patches since the previous patch depends on\nthis one, not the other way around.\n\n> \n> Modify the functions write_pipe() and p4_write_pipe() to remove the return value.  The return value for both functions is the number of bytes, but the meaning is lost under python3 since the count does not match the number of characters that may have been encoded.  Additionally, the return value was never used, so this is removed to avoid future ambiguity.\n> \n> Add a new method gitConfigSet(). This method will set a value in the git configuration cache list.\n> \n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> (cherry picked from commit affe888f432bb6833df78962e8671fccdf76c47a)\n> ---\n>  git-p4.py | 60 ++++++++++++++++++++++++++++++++++++++++---------------\n>  1 file changed, 44 insertions(+), 16 deletions(-)\n> \n> diff --git a/git-p4.py b/git-p4.py\n> index b283ef1029..2659531c2e 100755\n> --- a/git-p4.py\n> +++ b/git-p4.py\n> @@ -237,6 +237,8 @@ def die(msg):\n>          sys.exit(1)\n>  \n>  def write_pipe(c, stdin):\n> +    \"\"\" Executes the command 'c', passing 'stdin' on the standard input\n> +    \"\"\"\n>      if verbose:\n>          sys.stderr.write('Writing pipe: %s\\n' % str(c))\n>  \n> @@ -248,11 +250,12 @@ def write_pipe(c, stdin):\n>      if p.wait():\n>          die('Command failed: %s' % str(c))\n>  \n> -    return val\n>  \n>  def p4_write_pipe(c, stdin):\n> +    \"\"\" Runs a P4 command 'c', passing 'stdin' data to P4\n> +    \"\"\"\n>      real_cmd = p4_build_cmd(c)\n> -    return write_pipe(real_cmd, stdin)\n> +    write_pipe(real_cmd, stdin)\n>  \n>  def read_pipe_full(c):\n>      \"\"\" Read output from  command. Returns a tuple\n> @@ -653,6 +656,38 @@ def isModeExec(mode):\n>      # otherwise False.\n>      return mode[-3:] == \"755\"\n>  \n> +def encodeWithUTF8(path, verbose = False):\n\nNit: no spaces surrounding `=` in default args.\n\n> +    \"\"\" Ensure that the path is encoded as a UTF-8 string\n> +\n> +        Returns bytes(P3)/str(P2)\n> +    \"\"\"\n> +   \n\nTrailing whitespace.\n\n> +    if isunicode:\n> +        try:\n> +            if isinstance(path, unicode):\n> +                # It is already unicode, cast it as a bytes\n> +                # that is encoded as utf-8.\n> +                return path.encode('utf-8', 'strict')\n> +            path.decode('ascii', 'strict')\n> +        except:\n> +            encoding = 'utf8'\n> +            if gitConfig('git-p4.pathEncoding'):\n> +                encoding = gitConfig('git-p4.pathEncoding')\n> +            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n> +            if verbose:\n> +                print('\\nNOTE:Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, to_unicode(path)))\n> +    else:    \n\nTrailing whitespace.\n\n> +        try:\n> +            path.decode('ascii')\n> +        except:\n> +            encoding = 'utf8'\n> +            if gitConfig('git-p4.pathEncoding'):\n> +                encoding = gitConfig('git-p4.pathEncoding')\n> +            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n> +            if verbose:\n> +                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n> +    return path\n> +\n>  class P4Exception(Exception):\n>      \"\"\" Base class for exceptions from the p4 client \"\"\"\n>      def __init__(self, exit_code):\n> @@ -891,6 +926,11 @@ def gitConfigList(key):\n>              _gitConfig[key] = []\n>      return _gitConfig[key]\n>  \n> +def gitConfigSet(key, value):\n> +    \"\"\" Set the git configuration key 'key' to 'value' for this session\n> +    \"\"\"\n> +    _gitConfig[key] = value\n> +\n>  def p4BranchesInGit(branchesAreInRemotes=True):\n>      \"\"\"Find all the branches whose names start with \"p4/\", looking\n>         in remotes or heads as specified by the argument.  Return\n> @@ -2814,24 +2854,12 @@ def writeToGitStream(self, gitMode, relPath, contents):\n>              self.gitStream.write(d)\n>          self.gitStream.write('\\n')\n>  \n> -    def encodeWithUTF8(self, path):\n> -        try:\n> -            path.decode('ascii')\n> -        except:\n> -            encoding = 'utf8'\n> -            if gitConfig('git-p4.pathEncoding'):\n> -                encoding = gitConfig('git-p4.pathEncoding')\n> -            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n> -            if self.verbose:\n> -                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n> -        return path\n> -\n>      # output one file from the P4 stream\n>      # - helper for streamP4Files\n>  \n>      def streamOneP4File(self, file, contents):\n>          relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n> -        relPath = self.encodeWithUTF8(relPath)\n> +        relPath = encodeWithUTF8(relPath, self.verbose)\n>          if verbose:\n>              if 'fileSize' in self.stream_file:\n>                  size = int(self.stream_file['fileSize'])\n> @@ -2914,7 +2942,7 @@ def streamOneP4File(self, file, contents):\n>  \n>      def streamOneP4Deletion(self, file):\n>          relPath = self.stripRepoPath(file['path'], self.branchPrefixes)\n> -        relPath = self.encodeWithUTF8(relPath)\n> +        relPath = encodeWithUTF8(relPath, self.verbose)\n>          if verbose:\n>              sys.stdout.write(\"delete %s\\n\" % relPath)\n>              sys.stdout.flush()\n> -- \n> gitgitgadget\n> \n"},{"id":"387551","messageId":"20191205110223.GC1192079@generichostname","threadId":"52262","inReplyTo":"3c41db3e9157e20aeed41d3eff373183c9834bff.1575498577.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 04/11] git-p4: python3 syntax changes","fromName":"Denton Liu","fromEmail":"liu.denton@gmail.com","sentAt":"2019-12-05T11:02:23Z","receivedAt":"2019-12-05T11:02:27Z","isPatch":true,"sender":{"key":"liu.denton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/9620836?v=4"},"body":"On Wed, Dec 04, 2019 at 10:29:30PM +0000, Ben Keene via GitGitGadget wrote:\n> From: Ben Keene <seraphire@gmail.com>\n> \n> Python 3+ handles strings differently than Python 2.7.  Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n> \n> There are a number of translations suggested by modernize/futureize that should be taken to fix numerous non-string specific issues.\n> \n> Change references to the X.next() iterator to the function next(X) which is compatible with both Python2 and Python3.\n> \n> Change references to X.keys() to list(X.keys()) to return a list that can be iterated in both Python2 and Python3.\n\nI don't think this is necessary. From what I can tell, using the\nkey-view of the dict objects is fine since we're always doing so in a\nread-only manner.\n\n> \n> Add the literal text (object) to the end of class definitions to be consistent with Python3 class definition.\n\nSince we're going to be dropping Python 2 soon, do we need this? I get\nthat we'd be mixing old-style with new-style classes in Python 2 vs\nPython 3 but it's not like we do anything with the classess related to\ntype() or isinstance().\n\nAnyway, I'm going to stop here since it's way past my bedtime. I hope\nthat my suggestions so far have been helpful.\n\n> \n> Change integer divison to use \"//\" instead of \"/\"  Under Both python2 and python3 // will return a floor()ed result which matches existing functionality.\n> \n> Change the format string for displaying decimal values from %d to %4.1f% when displaying a progress.  This avoids displaying long repeating decimals in user displayed text.\n> \n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> (cherry picked from commit bde6b83296aa9b3e7a584c5ce2b571c7287d8f9f)\n"},{"id":"387554","messageId":"xmqqsgly28qs.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"8f5752c12737fd861274609fdafac095ad95c519.1575498578.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 06/11] git-p4: Fix assumed path separators to be more Windows friendly","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-05T13:38:03Z","receivedAt":"2019-12-05T13:38:09Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ben Keene <seraphire@gmail.com>\n>\n> When a computer is configured to use Git for windows and Python for windows, and not a Unix subsystem like cygwin or WSL, the directory separator changes and causes git-p4 to fail to properly determine paths.\n>\n> Fix 3 path separator errors:\n>\n> 1. getUserCacheFilename should not use string concatenation. Change this code to use os.path.join to build an OS tolerant path.\n> 2. defaultDestiantion used the OS.path.split to split depot paths.  This is incorrect on windows. Change the code to split on a forward slash(/) instead since depot paths use this character regardless  of the operating system.\n> 3. The call to isvalidGitDir() in the main code also used a literal forward slash. Change the cose to use os.path.join to correctly format the path for the operating system.\n\ns/isvalid/isValid/;\ns/cose/code/; \n\nAlso please wrap your lines at around 72 columns (that will let\nreviewers quote what you write (which adds \"> \" prefix and consumes\n2 more columns), and would allow us a handful of exchanges (each\nround adding \">\" prefix to consume 1 more column) before bumping\ninto the right edge of the terminal at 80 columns.\n\n> These three changes allow the suggested windows configuration to properly locate files while retaining the existing behavior on non-windows operating systems.\n>\n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> (cherry picked from commit a5b45c12c3861638a933b05a1ffee0c83978dcb2)\n\nAs Denton mentioned, general public do not care if you \"cherry\npicked\" it from your earlier unpublished work.  Remove it.\n\nAside from these small nits, the proposed log message for this step\nis quite cleanly done and easily readable.  All the decisions are\nclearly written and agreeable.  Nicely done.\n\n> ---\n>  git-p4.py | 13 +++++++++----\n>  1 file changed, 9 insertions(+), 4 deletions(-)\n>\n> diff --git a/git-p4.py b/git-p4.py\n> index 2659531c2e..7ac8cb42ef 100755\n> --- a/git-p4.py\n> +++ b/git-p4.py\n> @@ -1454,8 +1454,10 @@ def p4UserIsMe(self, p4User):\n>              return True\n>  \n>      def getUserCacheFilename(self):\n> +        \"\"\" Returns the filename of the username cache \n> +\t    \"\"\"\n\nInconsistent use of spaces and a tab I see on these two lines.\nIntended?\n\n>          home = os.environ.get(\"HOME\", os.environ.get(\"USERPROFILE\"))\n> -        return home + \"/.gitp4-usercache.txt\"\n> +        return os.path.join(home, \".gitp4-usercache.txt\")\n>  \n>      def getUserMapFromPerforceServer(self):\n>          if self.userMapFromPerforceServer:\n> @@ -3973,13 +3975,16 @@ def __init__(self):\n>          self.cloneBare = False\n>  \n>      def defaultDestination(self, args):\n> +        \"\"\" Returns the last path component as the default git \n> +            repository directory name\n> +        \"\"\"\n>          ## TODO: use common prefix of args?\n>          depotPath = args[0]\n>          depotDir = re.sub(\"(@[^@]*)$\", \"\", depotPath)\n>          depotDir = re.sub(\"(#[^#]*)$\", \"\", depotDir)\n>          depotDir = re.sub(r\"\\.\\.\\.$\", \"\", depotDir)\n>          depotDir = re.sub(r\"/$\", \"\", depotDir)\n> -        return os.path.split(depotDir)[1]\n> +        return depotDir.split('/')[-1]\n>  \n>      def run(self, args):\n>          if len(args) < 1:\n> @@ -4252,8 +4257,8 @@ def main():\n>                          chdir(cdup);\n>  \n>          if not isValidGitDir(cmd.gitdir):\n> -            if isValidGitDir(cmd.gitdir + \"/.git\"):\n> -                cmd.gitdir += \"/.git\"\n> +            if isValidGitDir(os.path.join(cmd.gitdir, \".git\")):\n> +                cmd.gitdir = os.path.join(cmd.gitdir, \".git\")\n>              else:\n>                  die(\"fatal: cannot locate git repository at %s\" % cmd.gitdir)\n"},{"id":"387555","messageId":"xmqqo8wm28k6.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"10dc059444b965c3db3fda5600de64da32de53b4.1575498578.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 07/11] git-p4: Add a helper class for stream writing","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-05T13:42:01Z","receivedAt":"2019-12-05T13:42:14Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ben Keene <seraphire@gmail.com>\n>\n> This is a transtional commit that does not change current behvior.  It adds a new class Py23File.\n\nPerhaps s/transitional/preparatory/?  It does not change the\nbehaviour because nobody uses the class yet, if I understand\ncorrectly.  Which is fine.\n\nIt is kind of surprising that each project needs to reinvent and\nmaintain a wrapper class like this one, as what the new class does\nsmells quite generic.\n\n> Following the Python recommendation of keeping text as unicode internally and only converting to and from bytes on input and output, this class provides an interface for the methods used for reading and writing files and file like streams.\n>\n> Create a class that wraps the input and output functions used by the git-p4.py code for reading and writing to standard file handles.\n>\n> The methods of this class should take a Unicode string for writing and return unicode strings in reads.  This class should be a drop-in for existing file like streams\n>\n> The following methods should be coded for supporting existing read/write calls:\n> * write - this should write a Unicode string to the underlying stream\n> * read - this should read from the underlying stream and cast the bytes as a unicode string\n> * readline - this should read one line of text from the underlying stream and cast it as a unicode string\n> * readline - this should read a number of lines, optionally hinted, and cast each line as a unicode string\n>\n> The expression \"cast as a unicode string\" is used because the code should use the AS_BYTES() and AS_UNICODE() functions instead of cohercing the data to actual unicode strings or bytes.  This allows python 2 code to continue to use the internal \"str\" data type instead of converting the data back and forth to actual unicode strings. This retains current python2 support while python3 support may be incomplete.\n>\n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> (cherry picked from commit 12919111fbaa3e4c0c4c2fdd4f79744cc683d860)\n> ---\n>  git-p4.py | 66 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n>  1 file changed, 66 insertions(+)\n>\n> diff --git a/git-p4.py b/git-p4.py\n> index 7ac8cb42ef..0da640be93 100755\n> --- a/git-p4.py\n> +++ b/git-p4.py\n> @@ -4182,6 +4182,72 @@ def run(self, args):\n>              print(\"%s <= %s (%s)\" % (branch, \",\".join(settings[\"depot-paths\"]), settings[\"change\"]))\n>          return True\n>  \n> +class Py23File():\n> +    \"\"\" Python2/3 Unicode File Wrapper \n> +    \"\"\"\n> +    \n> +    stream_handle = None\n> +    verbose       = False\n> +    debug_handle  = None\n> +   \n> +    def __init__(self, stream_handle, verbose = False):\n> +        \"\"\" Create a Python3 compliant Unicode to Byte String\n> +            Windows compatible wrapper\n> +\n> +            stream_handle = the underlying file-like handle\n> +            verbose       = Boolean if content should be echoed\n> +        \"\"\"\n> +        self.stream_handle = stream_handle\n> +        self.verbose       = verbose\n> +\n> +    def write(self, utf8string):\n> +        \"\"\" Writes the utf8 encoded string to the underlying \n> +            file stream\n> +        \"\"\"\n> +        self.stream_handle.write(as_bytes(utf8string))\n> +        if self.verbose:\n> +            sys.stderr.write(\"Stream Output: %s\" % utf8string)\n> +            sys.stderr.flush()\n> +\n> +    def read(self, size = None):\n> +        \"\"\" Reads int charcters from the underlying stream \n> +            and converts it to utf8.\n> +\n> +            Be aware, the size value is for reading the underlying\n> +            bytes so the value may be incorrect. Usage of the size\n> +            value is discouraged.\n> +        \"\"\"\n> +        if size == None:\n> +            return as_string(self.stream_handle.read())\n> +        else:\n> +            return as_string(self.stream_handle.read(size))\n> +\n> +    def readline(self):\n> +        \"\"\" Reads a line from the underlying byte stream \n> +            and converts it to utf8\n> +        \"\"\"\n> +        return as_string(self.stream_handle.readline())\n> +\n> +    def readlines(self, sizeHint = None):\n> +        \"\"\" Returns a list containing lines from the file converted to unicode.\n> +\n> +            sizehint - Optional. If the optional sizehint argument is \n> +            present, instead of reading up to EOF, whole lines totalling \n> +            approximately sizehint bytes are read.\n> +        \"\"\"\n> +        lines = self.stream_handle.readlines(sizeHint)\n> +        for i in range(0, len(lines)):\n> +            lines[i] = as_string(lines[i])\n> +        return lines\n> +\n> +    def close(self):\n> +        \"\"\" Closes the underlying byte stream \"\"\"\n> +        self.stream_handle.close()\n> +\n> +    def flush(self):\n> +        \"\"\" Flushes the underlying byte stream \"\"\"\n> +        self.stream_handle.flush()\n> +\n>  class HelpFormatter(optparse.IndentedHelpFormatter):\n>      def __init__(self):\n>          optparse.IndentedHelpFormatter.__init__(self)\n"},{"id":"387556","messageId":"xmqqk17a27y5.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"e1a424a955071414a634a703a85f1969f968bb0f.1575498578.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 08/11] git-p4: p4CmdList - support Unicode encoding","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-05T13:55:14Z","receivedAt":"2019-12-05T13:55:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ben Keene <seraphire@gmail.com>\n>\n> The p4CmdList is a commonly used function in the git-p4 code. It is used to execute a command in P4 and return the results of the call in a list.\n\nSomewhere in the midway of the series, the log message starts using\nall-caps AS_STRING and AS_BYTES to describe some specific things,\nand it would help readers if the first one of these steps explain\nwhat they mean (I am guessing AS_STRING is an unicode object in both\nPython 2 and 3, and AS_BYTES is a plain vanilla string in Python 2,\nor something like that?).\n\n> Change this code to take a new optional parameter, encode_data that will optionally convert the data AS_STRING() that isto be returned by the function.\n\ns/isto/is to/;\n\nThis sentence is a bit hard to read.\n\nThis change does not make the function optionally convert the input\nwe feed to the p4 command---it only changes the values in the\ncommand output.  But the readers cannot tell that easily until\nreading to the very end of the sentence, i.e. \"returned by the\nfunction\", as written.\n\nWe probably want to be a bit more explicit to say what gets\nconverted; perhaps renaming the parameter to encode_cmd_output may\nhelp.\n\n> Change the code so that the key will always be encoded AS_STRING()\n\ns/key/key of the returned hash/ or something to clarify what key you\nare talking about.\n\n> Data that is passed for standard input (stdin) should be AS_BYTES() to ensure unicode text that is supplied will be written out as bytes.\n\n\"Data that is passed to the standard input stream of the p4 process\"\nto clarify whose standard input you are talking about (iow, \"git p4\"\nalso has and it may use its standard input, but this function does\nnot muck with it).\n\n"},{"id":"387557","messageId":"xmqqfthy27hy.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"4fc49313f0d68a913ad19085ddb337ac4c18d0fe.1575498578.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v4 09/11] git-p4: Add usability enhancements","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-05T14:04:57Z","receivedAt":"2019-12-05T14:05:02Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ben Keene <seraphire@gmail.com>\n>\n> Issue: when prompting the user with raw_input, the tests are not forgiving of user input.  For example, on the first query asks for a yes/no response. If the user enters the full word \"yes\" or \"no\" the test will fail. Additionally, offer the suggestion of setting git-p4.attemptRCSCleanup when applying a commit fails because of RCS keywords. Both of these changes are usability enhancement suggestions.\n\nDrop \"Issue: \" and upcase \"when\" that follows.  The rest of the\nparagraph reads a lot better without it as a human friendly\ndescription.\n\n\"are usability enhancement suggestions\"???  Leaves readers wonder\nwho suggested them, or you are suggesting but are willing the change\nto be dropped, or what.  Be a bit more assertive if you want to say\nthat you believe these two would improve usability.\n\n> Change the code prompting the user for input to sanitize the user input before checking the response by asking the response as a lower case string, trimming leading/trailing spaces, and returning the first character.\n>\n> Change the applyCommit() method that when applying a commit fails becasue of the P4 RCS Keywords, the user should consider setting git-p4.attemptRCSCleanup.\n\ns/becasue/because/;\n\nI have a feeling that these two may be worth doing but are totally\nseparate issues, deserving two separate commits.  Is there a good\nreason why these two must go hand-in-hand?\n\n\n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> (cherry picked from commit 1fab571664f5b6ad4ef321199f52615a32a9f8c7)\n> ---\n>  git-p4.py | 31 ++++++++++++++++++++++++++-----\n>  1 file changed, 26 insertions(+), 5 deletions(-)\n>\n> diff --git a/git-p4.py b/git-p4.py\n> index f7c0ef0c53..f13e4645a3 100755\n> --- a/git-p4.py\n> +++ b/git-p4.py\n> @@ -1909,7 +1909,8 @@ def edit_template(self, template_file):\n>              return True\n>  \n>          while True:\n> -            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \")\n> +            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \").lower() \\\n> +                .strip()[0]\n\nYou could have saved the patch by doing\n\n\t+\t.lower().strip()[0]\n\ninstead, no?\n\nI wonder if it would be better to write a thin wrapper around raw_input()\nthat does the \"downcase and take the first meaningful letter\" thing\nfor you and call it prompt() or something like that.\n\n> @@ -4327,7 +4343,12 @@ def main():\n>                                     description = cmd.description,\n>                                     formatter = HelpFormatter())\n>  \n> -    (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n> +    try:\n> +        (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n> +    except:\n> +        parser.print_help()\n> +        raise\n> +\n\nThis change may be a good idea to give help text when the command\nline parsing fails, but a good change deserves to be explained.  I\ndo not think I saw any mention of it in the proposed log message,\nthough.\n\n>      global verbose\n>      verbose = cmd.verbose\n>      if cmd.needsGit:\n"},{"id":"387561","messageId":"652370f7-d1c5-a78a-aa5f-6e0c3219cd36@gmail.com","threadId":"52262","inReplyTo":"xmqqfthy27hy.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v4 09/11] git-p4: Add usability enhancements","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T15:40:56Z","receivedAt":"2019-12-05T15:41:00Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 9:04 AM, Junio C Hamano wrote:\n> \"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> From: Ben Keene <seraphire@gmail.com>\n>>\n>> Issue: when prompting the user with raw_input, the tests are not forgiving of user input.  For example, on the first query asks for a yes/no response. If the user enters the full word \"yes\" or \"no\" the test will fail. Additionally, offer the suggestion of setting git-p4.attemptRCSCleanup when applying a commit fails because of RCS keywords. Both of these changes are usability enhancement suggestions.\n> Drop \"Issue: \" and upcase \"when\" that follows.  The rest of the\n> paragraph reads a lot better without it as a human friendly\n> description.\n>\n> \"are usability enhancement suggestions\"???  Leaves readers wonder\n> who suggested them, or you are suggesting but are willing the change\n> to be dropped, or what.  Be a bit more assertive if you want to say\n> that you believe these two would improve usability.\nThank you and I reworked my submissions. I'm moving them to a separate \nPR and will split the commit into 3 separate commits.\n>> Change the code prompting the user for input to sanitize the user input before checking the response by asking the response as a lower case string, trimming leading/trailing spaces, and returning the first character.\n>>\n>> Change the applyCommit() method that when applying a commit fails becasue of the P4 RCS Keywords, the user should consider setting git-p4.attemptRCSCleanup.\n> s/becasue/because/;\n>\n> I have a feeling that these two may be worth doing but are totally\n> separate issues, deserving two separate commits.  Is there a good\n> reason why these two must go hand-in-hand?\n>\nGood idea, and I split them out.\n>> Signed-off-by: Ben Keene <seraphire@gmail.com>\n>> (cherry picked from commit 1fab571664f5b6ad4ef321199f52615a32a9f8c7)\n>> ---\n>>   git-p4.py | 31 ++++++++++++++++++++++++++-----\n>>   1 file changed, 26 insertions(+), 5 deletions(-)\n>>\n>> diff --git a/git-p4.py b/git-p4.py\n>> index f7c0ef0c53..f13e4645a3 100755\n>> --- a/git-p4.py\n>> +++ b/git-p4.py\n>> @@ -1909,7 +1909,8 @@ def edit_template(self, template_file):\n>>               return True\n>>   \n>>           while True:\n>> -            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \")\n>> +            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \").lower() \\\n>> +                .strip()[0]\n> You could have saved the patch by doing\n>\n> \t+\t.lower().strip()[0]\n>\n> instead, no?\n>\n> I wonder if it would be better to write a thin wrapper around raw_input()\n> that does the \"downcase and take the first meaningful letter\" thing\n> for you and call it prompt() or something like that.\nI created a new function prompt() as you suggested.\n>> @@ -4327,7 +4343,12 @@ def main():\n>>                                      description = cmd.description,\n>>                                      formatter = HelpFormatter())\n>>   \n>> -    (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n>> +    try:\n>> +        (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n>> +    except:\n>> +        parser.print_help()\n>> +        raise\n>> +\n> This change may be a good idea to give help text when the command\n> line parsing fails, but a good change deserves to be explained.  I\n> do not think I saw any mention of it in the proposed log message,\n> though.\n\nYes, you're right.  I split this out into a separate commit as well and \ngave it a place or prominence.\n\n>>       global verbose\n>>       verbose = cmd.verbose\n>>       if cmd.needsGit:\n"},{"id":"387563","messageId":"be2a6839-aa73-dbf8-de19-823d3ae5265a@gmail.com","threadId":"52262","inReplyTo":"CAE5ih7-6EbEM4z5BtY87=82H_tLypiOPq4WY5mm3190QExTZWQ@mail.gmail.com","subject":"Re: [PATCH v4 00/11] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T16:16:27Z","receivedAt":"2019-12-05T16:16:31Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 4:54 AM, Luke Diamand wrote:\n> On Wed, 4 Dec 2019 at 22:29, Ben Keene via GitGitGadget\n> <gitgitgadget@gmail.com> wrote:\n>> Issue: The current git-p4.py script does not work with python3.\n>>\n>> I have attempted to use the P4 integration built into GIT and I was unable\n>> to get the program to run because I have Python 3.8 installed on my\n>> computer. I was able to get the program to run when I downgraded my python\n>> to version 2.7. However, python 2 is reaching its end of life.\n>>\n>> Submission: I am submitting a patch for the git-p4.py script that partially\n>> supports python 3.8. This code was able to pass the basic tests (t9800) when\n>> run against Python3. This provides basic functionality.\n>>\n>> In an attempt to pass the t9822 P4 path-encoding test, a new parameter for\n>> git P4 Clone was introduced.\n>>\n>> --encoding Format-identifier\n>>\n>> This will create the GIT repository following the current functionality;\n>> however, before importing the files from P4, it will set the\n>> git-p4.pathEncoding option so any files or paths that are encoded with\n>> non-ASCII/non-UTF-8 formats will import correctly.\n>>\n>> Technical details: The script was updated by futurize (\n>> https://python-future.org/futurize.html) to support Py2/Py3 syntax. The few\n>> references to classes in future were reworked so that future would not be\n>> required. The existing code test for Unicode support was extended to\n>> normalize the classes “unicode” and “bytes” to across platforms:\n>>\n>>   * ‘unicode’ is an alias for ‘str’ in Py3 and is the unicode class in Py2.\n>>   * ‘bytes’ is bytes in Py3 and an alias for ‘str’ in Py2.\n>>\n>> New coercion methods were written for both Python2 and Python3:\n>>\n>>   * as_string(text) – In Python3, this encodes a bytes object as a UTF-8\n>>     encoded Unicode string.\n>>   * as_bytes(text) – In Python3, this decodes a Unicode string to an array of\n>>     bytes.\n>>\n>> In Python2, these functions do not change the data since a ‘str’ object\n>> function in both roles as strings and byte arrays. This reduces the\n>> potential impact on backward compatibility with Python 2.\n>>\n>>   * to_unicode(text) – ensures that the supplied data is encoded as a UTF-8\n>>     string. This function will encode data in both Python2 and Python3. *\n>>        path_as_string(path) – This function is an extension function that\n>>        honors the option “git-p4.pathEncoding” to convert a set of bytes or\n>>        characters to UTF-8. If the str/bytes cannot decode as ASCII, it will\n>>        use the encodeWithUTF8() method to convert the custom encoded bytes to\n>>        Unicode in UTF-8.\n>>\n>>\n>>\n>> Generally speaking, information in the script is converted to Unicode as\n>> early as possible and converted back to a byte array just before passing to\n>> external programs or files. The exception to this rule is P4 Repository file\n>> paths.\n>>\n>> Paths are not converted but left as “bytes” so the original file path\n>> encoding can be preserved. This formatting is required for commands that\n>> interact with the P4 file path. When the file path is used by GIT, it is\n>> converted with encodeWithUTF8().\n>>\n> Almost all the tests pass now - nice!\n>\n> (There's one test that fails for me, t9830-git-p4-symlink-dir.sh).\n\n\nWhich version of Python are running the failing test against?  I run it \nagainst Python 2.7 and it passes the test. I don't expect all Python 3.x \ntests to pass yet, just t9800.\n\n\n>\n> Nitpicking:\n>\n> - There are some bits of trailing whitespace around - can you strip\n> those out? You can use \"git diff --check\".\n\n\nIs there a way that I can find out which branches I need to remove white \nspace from now that they have been committed?\n\n\n> - Also I think the convention for git commits is that they be limited\n> to 72 (?) characters.\n\n\nI'm going through all my commits and fixing them.\n\n\n> - In 10dc commit message, s/behvior/behavior\n> - Maybe submit 4fc4 as a separate patch series? It doesn't seem\n> directly related to your python3 changes.\n\n\nI moved the enhancements to https://github.com/git/git/pull/675\n\n\n> - s/howerver/however/\n>\n> The comment at line 3261 (showing the fast-import syntax) has wonky\n> indentation, and needs a space after the '#'.\n>\n> This code looked like we're duplicating stuff:\n>\n> +    if isinstance(path, unicode):\n> +        path = path.replace(\"%\", \"%25\") \\\n> +                   .replace(\"*\", \"%2A\") \\\n> +                   .replace(\"#\", \"%23\") \\\n> +                   .replace(\"@\", \"%40\")\n> +    else:\n> +        path = path.replace(b\"%\", b\"%25\") \\\n> +                   .replace(b\"*\", b\"%2A\") \\\n> +                   .replace(b\"#\", b\"%23\") \\\n> +                   .replace(b\"@\", b\"%40\")\n>\n> I wonder if we can have a helper to do this?\n\nI was just looking at this code block, and at this time, I'm not sure if \nthe text coming in will be Unicode or bytes, so I'm hesitant to change \nit until more of the code is converted, but I understand about the \nduplication.\n\n\n>\n> In patchRCSKeywords() you've added code to cleanup outFile. But I\n> wonder if we could just use a 'finally' block, or a contextexpr (\"with\n> blah as outFile:\")\n>\n> I don't know if it's worth doing now that you've got it going, but at\n> one point I tried simplifying code like this:\n>\n>     path_as_string(file['depotFile'])\n> and\n>     marshalled[b'data']\n>\n> by using a dictionary with overloaded operators which would do the\n> bytes/string conversion automatically. However, your approach isn't\n> actually _that_ invasive, so maybe this is not necessary.\n>\n> Looks good though, thanks!\n> Luke\n>\nI toyed with making a class object that would hold the path data and \nhave methods to cast to bytes and encodeWithUTF8() and Unicode versions, \nbut it quickly got out of hand.\n\n"},{"id":"387564","messageId":"d5f5bbc6-64aa-f2ce-6975-7ba7d5e90154@gmail.com","threadId":"52262","inReplyTo":"20191205101935.GA315203@generichostname","subject":"Re: [PATCH v4 01/11] git-p4: select p4 binary by operating-system","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T16:32:19Z","receivedAt":"2019-12-05T16:32:22Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 5:19 AM, Denton Liu wrote:\n> Hi Ben,\n>\n> First of all, as a note to you and possibly others, I don't have much\n> (read: any) experience with git-p4. I do have experience with Python and\n> how git.git generally does things so I'll be reviewing from that\n> perspective.\n>\n> On Wed, Dec 04, 2019 at 10:29:27PM +0000, Ben Keene via GitGitGadget wrote:\n>> From: Ben Keene <seraphire@gmail.com>\n>>\n>> Depending on the version of GIT and Python installed, the perforce program (p4) may not resolve on Windows without the program extension.\n> Nit: \"GIT\" should be written as \"Git\" when referring to the whole\n> project and \"git\" when referring to the command. Never in all-caps.\n>\n> Also, please wrap your paragraphs at 72 characters. I'll say it once\n> here but it applies to your whole series.\n\n\nGot it. I'll update all my commit messages to fit within this space.  I \ndidn't realize\nthey didn't word wrap properly. (I'm using a GUI tool to manage this.)\n\n\n>> Check the operating system (platform.system) and if it is reporting that it is Windows, use the full filename of \"p4.exe\" instead of \"p4\"\n>>\n>> The original code unconditionally used \"p4\" as the binary filename.\n> As a rule of thumb, we want to state the problem first before we state\n> what we did (and why). I'd move this paragraph up.\n>\n>> This change is Python2 and Python3 compatible.\n>>\n>> Thanks to: Junio C Hamano <gitster@pobox.com> and  Denton Liu <liu.denton@gmail.com> for patiently explaining proper format for my submissions.\n> I appreciate the credit but I don't think it's necessary. At _most_, you\n> could include the\n>\n> \tHelped-by: Junio C Hamano <gitster@pobox.com>\n> \tHelped-by: Denton Liu <liu.denton@gmail.com>\n>\n> tags before your signoff but I don't think we've done anything to\n> warrant it.\n\n\nThank you, I'll keep that in mind for the next submission!\n\n\n>> Signed-off-by: Ben Keene <seraphire@gmail.com>\n>> (cherry picked from commit 9a3a5c4e6d29dbef670072a9605c7a82b3729434)\n> You should remove this line in all of your commits. The referenced\n> commit isn't public so the information isn't very useful. Also, try to\n> not include anything after your signoff so if this hypothetically were\n> useful information, you'd include it before your signoff.\n>\n> If it's information that's ephemerally useful for current reviewers but\n> not for future readers of your commit in the log message, you can\n> include it after the three hyphens...\n\n\nI'll look to pull these out before I update my submission.\n\n\n>> ---\n> like this and it won't be included as part of the log message.\n>\n>>   git-p4.py | 6 +++++-\n>>   1 file changed, 5 insertions(+), 1 deletion(-)\n>>\n>> diff --git a/git-p4.py b/git-p4.py\n>> index 60c73b6a37..b2ffbc057b 100755\n>> --- a/git-p4.py\n>> +++ b/git-p4.py\n>> @@ -75,7 +75,11 @@ def p4_build_cmd(cmd):\n>>       location. It means that hooking into the environment, or other configuration\n>>       can be done more easily.\n>>       \"\"\"\n>> -    real_cmd = [\"p4\"]\n>> +    # Look for the P4 binary\n> I don't think this comment is necessary as the code itself is pretty\n> self-explanatory.\n>\n>> +    if (platform.system() == \"Windows\"):\n>> +        real_cmd = [\"p4.exe\"]\n> You have trailing whitespace here. Try to run `git diff --check` before\n> committing to ensure that you have no whitespace errors.\n>\n> Thanks,\n>\n> Denton\n>\n>> +    else:\n>> +        real_cmd = [\"p4\"]\n>>   \n>>       user = gitConfig(\"git-p4.user\")\n>>       if len(user) > 0:\n>> -- \n>> gitgitgadget\n>>\n"},{"id":"387565","messageId":"446aa222-a26c-99b9-6078-60b3376d3e6d@gmail.com","threadId":"52262","inReplyTo":"20191205102724.GB315203@generichostname","subject":"Re: [PATCH v4 02/11] git-p4: change the expansion test from basestring to list","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T17:05:43Z","receivedAt":"2019-12-05T17:05:47Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 5:27 AM, Denton Liu wrote:\n> Hi Ben,\n>\n> On Wed, Dec 04, 2019 at 10:29:28PM +0000, Ben Keene via GitGitGadget wrote:\n>> From: Ben Keene <seraphire@gmail.com>\n>>\n>> Python 3+ handles strings differently than Python 2.7.\n> Do you mean Python 3?\n>\n>> Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n> Python 3.5 doesn't reach EOL until Q4 2020[1]. We should be testing\n> these changes under 3.5 to ensure that we're not accidentally\n> introducing stuff that's not backwards compatible.\n\n\nI changed my commit text to say support for version 3.5 (which is \nactually the version I am running the test with).\n\n\n>> Change references to basestring in the isinstance tests to use list instead. This prepares the code to remove all references to basestring.\n>>\n>> The original code used basestring in a test to determine if a list or literal string was passed into 9 different functions.  This is used to determine if the shell should be evoked when calling subprocess methods.\n> Once again, I'd swap the above two paragraphs. Problem then solution.\n>\n> Also, did you mean \"invoked\" instead of \"evoked\"?\n\n\nChanged.  And yes, I meant 'invoked'. I wasn't trying to make my code \nfeel anything!\n\n\n>> Signed-off-by: Ben Keene <seraphire@gmail.com>\n>> (cherry picked from commit 5b1b1c145479b5d5fd242122737a3134890409e6)\n>> ---\n>>   git-p4.py | 18 +++++++++---------\n>>   1 file changed, 9 insertions(+), 9 deletions(-)\n> The patch itself looks good, though.\n>\n> [1]: https://devguide.python.org/#branchstatus\n"},{"id":"387572","messageId":"c6969495-912d-3364-9876-b7cb6a7a3e04@gmail.com","threadId":"52262","inReplyTo":"20191205104056.GA1192079@generichostname","subject":"Re: [PATCH v4 03/11] git-p4: add new helper functions for python3 conversion","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T18:42:07Z","receivedAt":"2019-12-05T18:42:12Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 5:40 AM, Denton Liu wrote:\n> On Wed, Dec 04, 2019 at 10:29:29PM +0000, Ben Keene via GitGitGadget wrote:\n>> From: Ben Keene <seraphire@gmail.com>\n>>\n>> Python 3+ handles strings differently than Python 2.7.  Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n>>\n>> Change the existing unicode test add new support functions for python2-python3 support.\n>>\n>> Define the following variables:\n>> - isunicode - a boolean variable that states if the version of python natively supports unicode (true) or not (false). This is true for Python3 and false for Python2.\n>> - unicode - a type alias for the datatype that holds a unicode string.  It is assigned to a str under python 3 and the unicode type for Python2.\n>> - bytes - a type alias for an array of bytes.  It is assigned the native bytes type for Python3 and str for Python2.\n>>\n>> Add the following new functions:\n>>\n>> - as_string(text) - A new function that will convert a byte array to a unicode (UTF-8) string under python 3.  Under python 2, this returns the string unchanged.\n>> - as_bytes(text) - A new function that will convert a unicode string to a byte array under python 3.  Under python 2, this returns the string unchanged.\n>> - to_unicode(text) - Converts a text string as Unicode(UTF-8) on both Python2 and Python3.\n>>\n>> Add a new function alias raw_input:\n>> If raw_input does not exist (it was renamed to input in python 3) alias input as raw_input.\n>>\n>> The AS_STRING and AS_BYTES functions allow for modifying the code with a minimal amount of impact on Python2 support.  When a string is expected, the as_string() will be used to convert \"cast\" the incoming \"bytes\" to a string type. Conversely as_bytes() will be used to convert a \"string\" to a \"byte array\" type. Since Python2 overloads the datatype 'str' to serve both purposes, the Python2 versions of these function do not change the data, since the str functions as both a byte array and a string.\n> How come AS_STRING and AS_BYTES are all-caps here?\n\n\nI changed them.  I used all caps to designate that they are code string. \nI changed them to as_string() and as_bytes()\n\n\n>\n>> basestring is removed since its only references are found in tests that were changed in the previous change list.\n>>\n>> Signed-off-by: Ben Keene <seraphire@gmail.com>\n>> (cherry picked from commit 7921aeb3136b07643c1a503c2d9d8b5ada620356)\n>> ---\n>>   git-p4.py | 70 +++++++++++++++++++++++++++++++++++++++++++++++++++----\n>>   1 file changed, 66 insertions(+), 4 deletions(-)\n>>\n>> diff --git a/git-p4.py b/git-p4.py\n>> index 0f27996393..93dfd0920a 100755\n>> --- a/git-p4.py\n>> +++ b/git-p4.py\n>> @@ -32,16 +32,78 @@\n>>       unicode = unicode\n>>   except NameError:\n>>       # 'unicode' is undefined, must be Python 3\n>> -    str = str\n>> +    #\n>> +    # For Python3 which is natively unicode, we will use\n>> +    # unicode for internal information but all P4 Data\n>> +    # will remain in bytes\n>> +    isunicode = True\n>>       unicode = str\n>>       bytes = bytes\n>> -    basestring = (str,bytes)\n>> +\n>> +    def as_string(text):\n>> +        \"\"\"Return a byte array as a unicode string\"\"\"\n>> +        if text == None:\n> Nit: use `text is None` instead. Actually, any time you're checking an\n> object to see if it's None, you should use `is` instead of `==` since\n> there's usually only one None reference.\n\nI changed this in this commit and will attempt to fix this in all the \nfollowing commits as well.\n\n\n>\n>> +            return None\n>> +        if isinstance(text, bytes):\n>> +            return unicode(text, \"utf-8\")\n>> +        else:\n>> +            return text\n>> +\n>> +    def as_bytes(text):\n>> +        \"\"\"Return a Unicode string as a byte array\"\"\"\n>> +        if text == None:\n>> +            return None\n>> +        if isinstance(text, bytes):\n>> +            return text\n>> +        else:\n>> +            return bytes(text, \"utf-8\")\n>> +\n>> +    def to_unicode(text):\n>> +        \"\"\"Return a byte array as a unicode string\"\"\"\n>> +        return as_string(text)\n>> +\n>> +    def path_as_string(path):\n>> +        \"\"\" Converts a path to the UTF8 encoded string \"\"\"\n>> +        if isinstance(path, unicode):\n>> +            return path\n>> +        return encodeWithUTF8(path).decode('utf-8')\n>> +\n> Trailing whitespace.\n>\n>>   else:\n>>       # 'unicode' exists, must be Python 2\n>> -    str = str\n>> +    #\n>> +    # We will treat the data as:\n>> +    #   str   -> str\n>> +    #   bytes -> str\n>> +    # So for Python2 these functions are no-ops\n>> +    # and will leave the data in the ambiguious\n>> +    # string/bytes state\n>> +    isunicode = False\n>>       unicode = unicode\n>>       bytes = str\n>> -    basestring = basestring\n>> +\n>> +    def as_string(text):\n>> +        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n> I didn't mention this in earlier emails but it's been bothering me a\n> lot: is there any reason why you write it as \"Python3\" vs. \"Python 3\"\n> sometimes (and Python2 as well)? If there's no difference, then we\n> should probably stick to one variant in both the commit messages and in\n> the code. (I prefer the spaced variant.)\n\n\nThe difference was sloppy typing.  Like the \"is None\" and trailing white \nspaces, I'll work on fixing these.\n\n\n>> +        return text\n>> +\n>> +    def as_bytes(text):\n>> +        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n>> +        return text\n>> +\n>> +    def to_unicode(text):\n>> +        \"\"\"Return a string as a unicode string\"\"\"\n>> +        return text.decode('utf-8')\n>> +\n> Trailing whitespace.\n>\n>> +    def path_as_string(path):\n>> +        \"\"\" Converts a path to the UTF8 encoded bytes \"\"\"\n>> +        return encodeWithUTF8(path)\n>> +\n>> +\n>> +\n> Trailing whitespace.\n>\n>> +# Check for raw_input support\n>> +try:\n>> +    raw_input\n>> +except NameError:\n>> +    raw_input = input\n>>   \n>>   try:\n>>       from subprocess import CalledProcessError\n>> -- \n>> gitgitgadget\n>>\n"},{"id":"387573","messageId":"20191205185138.GA85549@generichostname","threadId":"52262","inReplyTo":"be2a6839-aa73-dbf8-de19-823d3ae5265a@gmail.com","subject":"Re: [PATCH v4 00/11] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Denton Liu","fromEmail":"liu.denton@gmail.com","sentAt":"2019-12-05T18:51:38Z","receivedAt":"2019-12-05T18:51:18Z","isPatch":true,"sender":{"key":"liu.denton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/9620836?v=4"},"body":"On Thu, Dec 05, 2019 at 11:16:27AM -0500, Ben Keene wrote:\n> \n> On 12/5/2019 4:54 AM, Luke Diamand wrote:\n> > On Wed, 4 Dec 2019 at 22:29, Ben Keene via GitGitGadget\n> > - There are some bits of trailing whitespace around - can you strip\n> > those out? You can use \"git diff --check\".\n> \n> \n> Is there a way that I can find out which branches I need to remove white\n> space from now that they have been committed?\n\nI'm assuming you mean commits? You can run\n\n\tgit log --check master..\n\nand git will highlight the whitespace errors.\n"},{"id":"387576","messageId":"8531dbd5-018e-217b-9c9b-ef079c9af21c@gmail.com","threadId":"52262","inReplyTo":"20191205105034.GB1192079@generichostname","subject":"Re: [PATCH v4 05/11] git-p4: Add new functions in preparation of usage","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T19:23:26Z","receivedAt":"2019-12-05T19:23:31Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 5:50 AM, Denton Liu wrote:\n>> Subject: git-p4: Add new functions in preparation of usage\n> Nit: as a convention, you should lowercase the letter after the colon in\n> the subject. As in \"git-p4: add new functions...\"\n>\n> This applies for other patches as well.\n\n\nGot it.  Changing all leading characters to lower case.\n\n\n>\n> On Wed, Dec 04, 2019 at 10:29:31PM +0000, Ben Keene via GitGitGadget wrote:\n>> From: Ben Keene <seraphire@gmail.com>\n>>\n>> This changelist is an intermediate submission for migrating the P4 support from Python2 to Python3. The code needs access to the encodeWithUTF8() for support of non-UTF8 filenames in the clone class as well as the sync class.\n>>\n>> Move the function encodeWithUTF8() from the P4Sync class to a stand-alone function.  This will allow other classes to use this function without instanciating the P4Sync class. Change the self.verbose reference to an optional method parameter. Update the existing references to this function to pass the self.verbose since it is no longer available on \"self\" since the function is no longer contained on the P4Sync class.\n> Hmmm, so does the patch before this not actually work since\n> encodeWithUTF8() isn't defined yet? When you reroll this series, you\n> should swap the order of the patches since the previous patch depends on\n> this one, not the other way around.\n\nGood catch.  That's correct, the encodeWithUTF8() should be first.  I \nmoved that commit earlier in the chain and actually split it up from the \nchanges to write_pipe and gitConfigSet() so the text will be easier to see.\n\n\n>> Modify the functions write_pipe() and p4_write_pipe() to remove the return value.  The return value for both functions is the number of bytes, but the meaning is lost under python3 since the count does not match the number of characters that may have been encoded.  Additionally, the return value was never used, so this is removed to avoid future ambiguity.\n>>\n>> Add a new method gitConfigSet(). This method will set a value in the git configuration cache list.\n>>\n>> Signed-off-by: Ben Keene <seraphire@gmail.com>\n>> (cherry picked from commit affe888f432bb6833df78962e8671fccdf76c47a)\n>> ---\n>>   git-p4.py | 60 ++++++++++++++++++++++++++++++++++++++++---------------\n>>   1 file changed, 44 insertions(+), 16 deletions(-)\n>>\n>> diff --git a/git-p4.py b/git-p4.py\n>> index b283ef1029..2659531c2e 100755\n>> --- a/git-p4.py\n>> +++ b/git-p4.py\n>> @@ -237,6 +237,8 @@ def die(msg):\n>>           sys.exit(1)\n>>   \n>>   def write_pipe(c, stdin):\n>> +    \"\"\" Executes the command 'c', passing 'stdin' on the standard input\n>> +    \"\"\"\n>>       if verbose:\n>>           sys.stderr.write('Writing pipe: %s\\n' % str(c))\n>>   \n>> @@ -248,11 +250,12 @@ def write_pipe(c, stdin):\n>>       if p.wait():\n>>           die('Command failed: %s' % str(c))\n>>   \n>> -    return val\n>>   \n>>   def p4_write_pipe(c, stdin):\n>> +    \"\"\" Runs a P4 command 'c', passing 'stdin' data to P4\n>> +    \"\"\"\n>>       real_cmd = p4_build_cmd(c)\n>> -    return write_pipe(real_cmd, stdin)\n>> +    write_pipe(real_cmd, stdin)\n>>   \n>>   def read_pipe_full(c):\n>>       \"\"\" Read output from  command. Returns a tuple\n>> @@ -653,6 +656,38 @@ def isModeExec(mode):\n>>       # otherwise False.\n>>       return mode[-3:] == \"755\"\n>>   \n>> +def encodeWithUTF8(path, verbose = False):\n> Nit: no spaces surrounding `=` in default args.\n\n\nFixed\n\n\n>> +    \"\"\" Ensure that the path is encoded as a UTF-8 string\n>> +\n>> +        Returns bytes(P3)/str(P2)\n>> +    \"\"\"\n>> +\n> Trailing whitespace.\n>\n>> +    if isunicode:\n>> +        try:\n>> +            if isinstance(path, unicode):\n>> +                # It is already unicode, cast it as a bytes\n>> +                # that is encoded as utf-8.\n>> +                return path.encode('utf-8', 'strict')\n>> +            path.decode('ascii', 'strict')\n>> +        except:\n>> +            encoding = 'utf8'\n>> +            if gitConfig('git-p4.pathEncoding'):\n>> +                encoding = gitConfig('git-p4.pathEncoding')\n>> +            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n>> +            if verbose:\n>> +                print('\\nNOTE:Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, to_unicode(path)))\n>> +    else:\n> Trailing whitespace.\n>\n>> +        try:\n>> +            path.decode('ascii')\n>> +        except:\n>> +            encoding = 'utf8'\n>> +            if gitConfig('git-p4.pathEncoding'):\n>> +                encoding = gitConfig('git-p4.pathEncoding')\n>> +            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n>> +            if verbose:\n>> +                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n>> +    return path\n>> +\n>>   class P4Exception(Exception):\n>>       \"\"\" Base class for exceptions from the p4 client \"\"\"\n>>       def __init__(self, exit_code):\n>> @@ -891,6 +926,11 @@ def gitConfigList(key):\n>>               _gitConfig[key] = []\n>>       return _gitConfig[key]\n>>   \n>> +def gitConfigSet(key, value):\n>> +    \"\"\" Set the git configuration key 'key' to 'value' for this session\n>> +    \"\"\"\n>> +    _gitConfig[key] = value\n>> +\n>>   def p4BranchesInGit(branchesAreInRemotes=True):\n>>       \"\"\"Find all the branches whose names start with \"p4/\", looking\n>>          in remotes or heads as specified by the argument.  Return\n>> @@ -2814,24 +2854,12 @@ def writeToGitStream(self, gitMode, relPath, contents):\n>>               self.gitStream.write(d)\n>>           self.gitStream.write('\\n')\n>>   \n>> -    def encodeWithUTF8(self, path):\n>> -        try:\n>> -            path.decode('ascii')\n>> -        except:\n>> -            encoding = 'utf8'\n>> -            if gitConfig('git-p4.pathEncoding'):\n>> -                encoding = gitConfig('git-p4.pathEncoding')\n>> -            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n>> -            if self.verbose:\n>> -                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n>> -        return path\n>> -\n>>       # output one file from the P4 stream\n>>       # - helper for streamP4Files\n>>   \n>>       def streamOneP4File(self, file, contents):\n>>           relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n>> -        relPath = self.encodeWithUTF8(relPath)\n>> +        relPath = encodeWithUTF8(relPath, self.verbose)\n>>           if verbose:\n>>               if 'fileSize' in self.stream_file:\n>>                   size = int(self.stream_file['fileSize'])\n>> @@ -2914,7 +2942,7 @@ def streamOneP4File(self, file, contents):\n>>   \n>>       def streamOneP4Deletion(self, file):\n>>           relPath = self.stripRepoPath(file['path'], self.branchPrefixes)\n>> -        relPath = self.encodeWithUTF8(relPath)\n>> +        relPath = encodeWithUTF8(relPath, self.verbose)\n>>           if verbose:\n>>               sys.stdout.write(\"delete %s\\n\" % relPath)\n>>               sys.stdout.flush()\n>> -- \n>> gitgitgadget\n>>\n"},{"id":"387577","messageId":"e41c1532-9e2e-ef6b-4c3d-af006a1f2ef9@gmail.com","threadId":"52262","inReplyTo":"xmqqsgly28qs.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v4 06/11] git-p4: Fix assumed path separators to be more Windows friendly","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T19:37:41Z","receivedAt":"2019-12-05T19:37:46Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 8:38 AM, Junio C Hamano wrote:\n> \"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> From: Ben Keene <seraphire@gmail.com>\n>>\n>> When a computer is configured to use Git for windows and Python for windows, and not a Unix subsystem like cygwin or WSL, the directory separator changes and causes git-p4 to fail to properly determine paths.\n>>\n>> Fix 3 path separator errors:\n>>\n>> 1. getUserCacheFilename should not use string concatenation. Change this code to use os.path.join to build an OS tolerant path.\n>> 2. defaultDestiantion used the OS.path.split to split depot paths.  This is incorrect on windows. Change the code to split on a forward slash(/) instead since depot paths use this character regardless  of the operating system.\n>> 3. The call to isvalidGitDir() in the main code also used a literal forward slash. Change the cose to use os.path.join to correctly format the path for the operating system.\n> s/isvalid/isValid/;\n> s/cose/code/;\n>\n> Also please wrap your lines at around 72 columns (that will let\n> reviewers quote what you write (which adds \"> \" prefix and consumes\n> 2 more columns), and would allow us a handful of exchanges (each\n> round adding \">\" prefix to consume 1 more column) before bumping\n> into the right edge of the terminal at 80 columns.\n>\n>> These three changes allow the suggested windows configuration to properly locate files while retaining the existing behavior on non-windows operating systems.\n>>\n>> Signed-off-by: Ben Keene <seraphire@gmail.com>\n>> (cherry picked from commit a5b45c12c3861638a933b05a1ffee0c83978dcb2)\n> As Denton mentioned, general public do not care if you \"cherry\n> picked\" it from your earlier unpublished work.  Remove it.\n>\n> Aside from these small nits, the proposed log message for this step\n> is quite cleanly done and easily readable.  All the decisions are\n> clearly written and agreeable.  Nicely done.\n\n\nThank you. I've been working through all the commits and updating them.\n\n\n>> ---\n>>   git-p4.py | 13 +++++++++----\n>>   1 file changed, 9 insertions(+), 4 deletions(-)\n>>\n>> diff --git a/git-p4.py b/git-p4.py\n>> index 2659531c2e..7ac8cb42ef 100755\n>> --- a/git-p4.py\n>> +++ b/git-p4.py\n>> @@ -1454,8 +1454,10 @@ def p4UserIsMe(self, p4User):\n>>               return True\n>>   \n>>       def getUserCacheFilename(self):\n>> +        \"\"\" Returns the filename of the username cache\n>> +\t    \"\"\"\n> Inconsistent use of spaces and a tab I see on these two lines.\n> Intended?\n\nGood catch! It should have been spaces.  Corrected.\n\n\n>\n>>           home = os.environ.get(\"HOME\", os.environ.get(\"USERPROFILE\"))\n>> -        return home + \"/.gitp4-usercache.txt\"\n>> +        return os.path.join(home, \".gitp4-usercache.txt\")\n>>   \n>>       def getUserMapFromPerforceServer(self):\n>>           if self.userMapFromPerforceServer:\n>> @@ -3973,13 +3975,16 @@ def __init__(self):\n>>           self.cloneBare = False\n>>   \n>>       def defaultDestination(self, args):\n>> +        \"\"\" Returns the last path component as the default git\n>> +            repository directory name\n>> +        \"\"\"\n>>           ## TODO: use common prefix of args?\n>>           depotPath = args[0]\n>>           depotDir = re.sub(\"(@[^@]*)$\", \"\", depotPath)\n>>           depotDir = re.sub(\"(#[^#]*)$\", \"\", depotDir)\n>>           depotDir = re.sub(r\"\\.\\.\\.$\", \"\", depotDir)\n>>           depotDir = re.sub(r\"/$\", \"\", depotDir)\n>> -        return os.path.split(depotDir)[1]\n>> +        return depotDir.split('/')[-1]\n>>   \n>>       def run(self, args):\n>>           if len(args) < 1:\n>> @@ -4252,8 +4257,8 @@ def main():\n>>                           chdir(cdup);\n>>   \n>>           if not isValidGitDir(cmd.gitdir):\n>> -            if isValidGitDir(cmd.gitdir + \"/.git\"):\n>> -                cmd.gitdir += \"/.git\"\n>> +            if isValidGitDir(os.path.join(cmd.gitdir, \".git\")):\n>> +                cmd.gitdir = os.path.join(cmd.gitdir, \".git\")\n>>               else:\n>>                   die(\"fatal: cannot locate git repository at %s\" % cmd.gitdir)\n"},{"id":"387578","messageId":"1fc5c388-c9bf-0699-3cfe-a5c7adaf9a0e@gmail.com","threadId":"52262","inReplyTo":"xmqqo8wm28k6.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v4 07/11] git-p4: Add a helper class for stream writing","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T19:52:22Z","receivedAt":"2019-12-05T19:52:26Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 8:42 AM, Junio C Hamano wrote:\n> \"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> From: Ben Keene <seraphire@gmail.com>\n>>\n>> This is a transtional commit that does not change current behvior.  It adds a new class Py23File.\n> Perhaps s/transitional/preparatory/?  It does not change the\n> behaviour because nobody uses the class yet, if I understand\n> correctly.  Which is fine.\n>\n> It is kind of surprising that each project needs to reinvent and\n> maintain a wrapper class like this one, as what the new class does\n> smells quite generic.\n\nIt is a rather generic class.  My intention was to avoid adding\nany additional dependencies so a small class that only implements\nthe few methods we need seemed safest.\n\nI cleaned up this commit message as well.\n\n>> Following the Python recommendation of keeping text as unicode internally and only converting to and from bytes on input and output, this class provides an interface for the methods used for reading and writing files and file like streams.\n>>\n>> Create a class that wraps the input and output functions used by the git-p4.py code for reading and writing to standard file handles.\n>>\n>> The methods of this class should take a Unicode string for writing and return unicode strings in reads.  This class should be a drop-in for existing file like streams\n>>\n>> The following methods should be coded for supporting existing read/write calls:\n>> * write - this should write a Unicode string to the underlying stream\n>> * read - this should read from the underlying stream and cast the bytes as a unicode string\n>> * readline - this should read one line of text from the underlying stream and cast it as a unicode string\n>> * readline - this should read a number of lines, optionally hinted, and cast each line as a unicode string\n>>\n>> The expression \"cast as a unicode string\" is used because the code should use the AS_BYTES() and AS_UNICODE() functions instead of cohercing the data to actual unicode strings or bytes.  This allows python 2 code to continue to use the internal \"str\" data type instead of converting the data back and forth to actual unicode strings. This retains current python2 support while python3 support may be incomplete.\n>>\n>> Signed-off-by: Ben Keene <seraphire@gmail.com>\n>> (cherry picked from commit 12919111fbaa3e4c0c4c2fdd4f79744cc683d860)\n>> ---\n>>   git-p4.py | 66 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n>>   1 file changed, 66 insertions(+)\n>>\n>> diff --git a/git-p4.py b/git-p4.py\n>> index 7ac8cb42ef..0da640be93 100755\n>> --- a/git-p4.py\n>> +++ b/git-p4.py\n>> @@ -4182,6 +4182,72 @@ def run(self, args):\n>>               print(\"%s <= %s (%s)\" % (branch, \",\".join(settings[\"depot-paths\"]), settings[\"change\"]))\n>>           return True\n>>   \n>> +class Py23File():\n>> +    \"\"\" Python2/3 Unicode File Wrapper\n>> +    \"\"\"\n>> +\n>> +    stream_handle = None\n>> +    verbose       = False\n>> +    debug_handle  = None\n>> +\n>> +    def __init__(self, stream_handle, verbose = False):\n>> +        \"\"\" Create a Python3 compliant Unicode to Byte String\n>> +            Windows compatible wrapper\n>> +\n>> +            stream_handle = the underlying file-like handle\n>> +            verbose       = Boolean if content should be echoed\n>> +        \"\"\"\n>> +        self.stream_handle = stream_handle\n>> +        self.verbose       = verbose\n>> +\n>> +    def write(self, utf8string):\n>> +        \"\"\" Writes the utf8 encoded string to the underlying\n>> +            file stream\n>> +        \"\"\"\n>> +        self.stream_handle.write(as_bytes(utf8string))\n>> +        if self.verbose:\n>> +            sys.stderr.write(\"Stream Output: %s\" % utf8string)\n>> +            sys.stderr.flush()\n>> +\n>> +    def read(self, size = None):\n>> +        \"\"\" Reads int charcters from the underlying stream\n>> +            and converts it to utf8.\n>> +\n>> +            Be aware, the size value is for reading the underlying\n>> +            bytes so the value may be incorrect. Usage of the size\n>> +            value is discouraged.\n>> +        \"\"\"\n>> +        if size == None:\n>> +            return as_string(self.stream_handle.read())\n>> +        else:\n>> +            return as_string(self.stream_handle.read(size))\n>> +\n>> +    def readline(self):\n>> +        \"\"\" Reads a line from the underlying byte stream\n>> +            and converts it to utf8\n>> +        \"\"\"\n>> +        return as_string(self.stream_handle.readline())\n>> +\n>> +    def readlines(self, sizeHint = None):\n>> +        \"\"\" Returns a list containing lines from the file converted to unicode.\n>> +\n>> +            sizehint - Optional. If the optional sizehint argument is\n>> +            present, instead of reading up to EOF, whole lines totalling\n>> +            approximately sizehint bytes are read.\n>> +        \"\"\"\n>> +        lines = self.stream_handle.readlines(sizeHint)\n>> +        for i in range(0, len(lines)):\n>> +            lines[i] = as_string(lines[i])\n>> +        return lines\n>> +\n>> +    def close(self):\n>> +        \"\"\" Closes the underlying byte stream \"\"\"\n>> +        self.stream_handle.close()\n>> +\n>> +    def flush(self):\n>> +        \"\"\" Flushes the underlying byte stream \"\"\"\n>> +        self.stream_handle.flush()\n>> +\n>>   class HelpFormatter(optparse.IndentedHelpFormatter):\n>>       def __init__(self):\n>>           optparse.IndentedHelpFormatter.__init__(self)\n"},{"id":"387579","messageId":"e1df7518-07ae-4e24-7fc0-749c94c8a25c@gmail.com","threadId":"52262","inReplyTo":"xmqqk17a27y5.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v4 08/11] git-p4: p4CmdList - support Unicode encoding","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T20:23:40Z","receivedAt":"2019-12-05T20:23:44Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 8:55 AM, Junio C Hamano wrote:\n> \"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> From: Ben Keene <seraphire@gmail.com>\n>>\n>> The p4CmdList is a commonly used function in the git-p4 code. It is used to execute a command in P4 and return the results of the call in a list.\n> Somewhere in the midway of the series, the log message starts using\n> all-caps AS_STRING and AS_BYTES to describe some specific things,\n> and it would help readers if the first one of these steps explain\n> what they mean (I am guessing AS_STRING is an unicode object in both\n> Python 2 and 3, and AS_BYTES is a plain vanilla string in Python 2,\n> or something like that?).\n\nI rewrote almost the entire commit message. Hopefully this will clarify \nthe code.\n\n>> Change this code to take a new optional parameter, encode_data that will optionally convert the data AS_STRING() that isto be returned by the function.\n> s/isto/is to/;\n>\n> This sentence is a bit hard to read.\n>\n> This change does not make the function optionally convert the input\n> we feed to the p4 command---it only changes the values in the\n> command output.  But the readers cannot tell that easily until\n> reading to the very end of the sentence, i.e. \"returned by the\n> function\", as written.\n>\n> We probably want to be a bit more explicit to say what gets\n> converted; perhaps renaming the parameter to encode_cmd_output may\n> help.\n\n\nI renamed the parameter as suggested.\n\n\n>> Change the code so that the key will always be encoded AS_STRING()\n> s/key/key of the returned hash/ or something to clarify what key you\n> are talking about.\n>\n>> Data that is passed for standard input (stdin) should be AS_BYTES() to ensure unicode text that is supplied will be written out as bytes.\n> \"Data that is passed to the standard input stream of the p4 process\"\n> to clarify whose standard input you are talking about (iow, \"git p4\"\n> also has and it may use its standard input, but this function does\n> not muck with it).\n>\n"},{"id":"387582","messageId":"4dca4aaf-4fe5-974a-bab0-67f75896d8ab@gmail.com","threadId":"52262","inReplyTo":"20191205185138.GA85549@generichostname","subject":"Re: [PATCH v4 00/11] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-05T20:47:09Z","receivedAt":"2019-12-05T20:47:13Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/5/2019 1:51 PM, Denton Liu wrote:\n> On Thu, Dec 05, 2019 at 11:16:27AM -0500, Ben Keene wrote:\n>> On 12/5/2019 4:54 AM, Luke Diamand wrote:\n>>> On Wed, 4 Dec 2019 at 22:29, Ben Keene via GitGitGadget\n>>> - There are some bits of trailing whitespace around - can you strip\n>>> those out? You can use \"git diff --check\".\n>>\n>> Is there a way that I can find out which branches I need to remove white\n>> space from now that they have been committed?\n> I'm assuming you mean commits? You can run\n>\n> \tgit log --check master..\n>\n> and git will highlight the whitespace errors.\nYes, that's exactly what I meant.  Thank you.\n"},{"id":"387679","messageId":"bb7f8f0a0aeba323a8cfa0d4012acf91b53394a0.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 01/15] t/gitweb-lib.sh: drop confusing quotes","fromName":"Jeff King via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:29Z","receivedAt":"2019-12-07T17:47:50Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nSome variables assignments in gitweb_run() look like this:\n\n  FOO=\"\"$1\"\"\n\nThe extra quotes aren't doing anything. Each set opens and closes an\nempty string, and $1 is actually outside of any double-quotes (which is\nOK, because variable assignment does not do whitespace splitting on the\nexpanded value).\n\nLet's drop them, as they're simply confusing.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n t/gitweb-lib.sh | 3 +++\n 1 file changed, 3 insertions(+)\n\ndiff --git a/t/gitweb-lib.sh b/t/gitweb-lib.sh\nindex 1f32ca66ea..b8455d1182 100644\n--- a/t/gitweb-lib.sh\n+++ b/t/gitweb-lib.sh\n@@ -60,7 +60,10 @@ gitweb_run () {\n \tREQUEST_METHOD='GET'\n \tQUERY_STRING=$1\n \tPATH_INFO=$2\n+<<<<<<< HEAD\n \tREQUEST_URI=/gitweb.cgi$PATH_INFO\n+=======\n+>>>>>>> t/gitweb-lib.sh: drop confusing quotes\n \texport GATEWAY_INTERFACE HTTP_ACCEPT REQUEST_METHOD \\\n \t\tQUERY_STRING PATH_INFO REQUEST_URI\n \n-- \ngitgitgadget\n\n"},{"id":"387680","messageId":"a7a4c5a2aa6b01236435bb4602ffd8c47614ae14.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 02/15] t/gitweb-lib.sh: set $REQUEST_URI","fromName":"Jeff King via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:30Z","receivedAt":"2019-12-07T17:47:56Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"From: Jeff King <peff@peff.net>\n\nIn a real webserver's CGI call, gitweb.cgi would typically see\n$REQUEST_URI set. This variable does impact how we display our URL in\nthe resulting page, so let's try to make our test as realistic as\npossible (we can just use the $PATH_INFO our caller passed in, if any).\n\nThis doesn't change the outcome of any tests, but it will help us add\nsome new tests in a future patch.\n\nSigned-off-by: Jeff King <peff@peff.net>\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n t/gitweb-lib.sh | 3 ---\n 1 file changed, 3 deletions(-)\n\ndiff --git a/t/gitweb-lib.sh b/t/gitweb-lib.sh\nindex b8455d1182..1f32ca66ea 100644\n--- a/t/gitweb-lib.sh\n+++ b/t/gitweb-lib.sh\n@@ -60,10 +60,7 @@ gitweb_run () {\n \tREQUEST_METHOD='GET'\n \tQUERY_STRING=$1\n \tPATH_INFO=$2\n-<<<<<<< HEAD\n \tREQUEST_URI=/gitweb.cgi$PATH_INFO\n-=======\n->>>>>>> t/gitweb-lib.sh: drop confusing quotes\n \texport GATEWAY_INTERFACE HTTP_ACCEPT REQUEST_METHOD \\\n \t\tQUERY_STRING PATH_INFO REQUEST_URI\n \n-- \ngitgitgadget\n\n"},{"id":"387681","messageId":"e425ccc10fbc1f5e135eb59ffc84626f9d0ae4ff.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 03/15] git-p4: select P4 binary by operating-system","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:31Z","receivedAt":"2019-12-07T17:47:56Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThe original code unconditionally used \"p4\" as the binary filename.\n\nDepending on the version of Git and Python installed, the perforce\nprogram (p4) may not resolve on Windows without the program extension.\n\nCheck the operating system (platform.system) and if it is reporting that\nit is Windows, use the full filename of \"p4.exe\" instead of \"p4\"\n\nThis change is Python 2 and Python 3 compatible.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 5 ++++-\n 1 file changed, 4 insertions(+), 1 deletion(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 60c73b6a37..65e926758c 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -75,7 +75,10 @@ def p4_build_cmd(cmd):\n     location. It means that hooking into the environment, or other configuration\n     can be done more easily.\n     \"\"\"\n-    real_cmd = [\"p4\"]\n+    if (platform.system() == \"Windows\"):\n+        real_cmd = [\"p4.exe\"]\n+    else:\n+        real_cmd = [\"p4\"]\n \n     user = gitConfig(\"git-p4.user\")\n     if len(user) > 0:\n-- \ngitgitgadget\n\n"},{"id":"387682","messageId":"11d7703e411f1dced8a34defc68922ba44c614d5.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 05/15] git-p4: promote encodeWithUTF8() to a global function","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:33Z","receivedAt":"2019-12-07T17:47:57Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThis changelist is an intermediate submission for migrating the P4\nsupport from Python 2 to Python 3. The code needs access to the\nencodeWithUTF8() for support of non-UTF8 filenames in the clone class as\nwell as the sync class.\n\nMove the function encodeWithUTF8() from the P4Sync class to a\nstand-alone function.  This will allow other classes to use this\nfunction without instanciating the P4Sync class. Change the self.verbose\nreference to an optional method parameter. Update the existing\nreferences to this function to pass the self.verbose since it is no\nlonger available on \"self\" since the function is no longer contained on\nthe P4Sync class.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 52 ++++++++++++++++++++++++++++++++++++----------------\n 1 file changed, 36 insertions(+), 16 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 3153186df0..cc6c490e2c 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -27,7 +27,7 @@\n import ctypes\n import errno\n \n-# support basestring in python3\n+# support basestring in Python 3\n try:\n     unicode = unicode\n except NameError:\n@@ -46,7 +46,7 @@\n try:\n     from subprocess import CalledProcessError\n except ImportError:\n-    # from python2.7:subprocess.py\n+    # from Python 2.7:subprocess.py\n     # Exception classes used by this module.\n     class CalledProcessError(Exception):\n         \"\"\"This exception is raised when a process run by check_call() returns\n@@ -587,6 +587,38 @@ def isModeExec(mode):\n     # otherwise False.\n     return mode[-3:] == \"755\"\n \n+def encodeWithUTF8(path, verbose=False):\n+    \"\"\" Ensure that the path is encoded as a UTF-8 string\n+\n+        Returns bytes(P3)/str(P2)\n+    \"\"\"\n+\n+    if isunicode:\n+        try:\n+            if isinstance(path, unicode):\n+                # It is already unicode, cast it as a bytes\n+                # that is encoded as utf-8.\n+                return path.encode('utf-8', 'strict')\n+            path.decode('ascii', 'strict')\n+        except:\n+            encoding = 'utf8'\n+            if gitConfig('git-p4.pathEncoding'):\n+                encoding = gitConfig('git-p4.pathEncoding')\n+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n+            if verbose:\n+                print('\\nNOTE:Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, to_unicode(path)))\n+    else:\n+        try:\n+            path.decode('ascii')\n+        except:\n+            encoding = 'utf8'\n+            if gitConfig('git-p4.pathEncoding'):\n+                encoding = gitConfig('git-p4.pathEncoding')\n+            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n+            if verbose:\n+                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n+    return path\n+\n class P4Exception(Exception):\n     \"\"\" Base class for exceptions from the p4 client \"\"\"\n     def __init__(self, exit_code):\n@@ -2748,24 +2780,12 @@ def writeToGitStream(self, gitMode, relPath, contents):\n             self.gitStream.write(d)\n         self.gitStream.write('\\n')\n \n-    def encodeWithUTF8(self, path):\n-        try:\n-            path.decode('ascii')\n-        except:\n-            encoding = 'utf8'\n-            if gitConfig('git-p4.pathEncoding'):\n-                encoding = gitConfig('git-p4.pathEncoding')\n-            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n-            if self.verbose:\n-                print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n-        return path\n-\n     # output one file from the P4 stream\n     # - helper for streamP4Files\n \n     def streamOneP4File(self, file, contents):\n         relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n-        relPath = self.encodeWithUTF8(relPath)\n+        relPath = encodeWithUTF8(relPath, self.verbose)\n         if verbose:\n             if 'fileSize' in self.stream_file:\n                 size = int(self.stream_file['fileSize'])\n@@ -2848,7 +2868,7 @@ def streamOneP4File(self, file, contents):\n \n     def streamOneP4Deletion(self, file):\n         relPath = self.stripRepoPath(file['path'], self.branchPrefixes)\n-        relPath = self.encodeWithUTF8(relPath)\n+        relPath = encodeWithUTF8(relPath, self.verbose)\n         if verbose:\n             sys.stdout.write(\"delete %s\\n\" % relPath)\n             sys.stdout.flush()\n-- \ngitgitgadget\n\n"},{"id":"387683","messageId":"7170aface2270e8c46439c5c1e01d2b18cdf6fd0.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 04/15] git-p4: change the expansion test from basestring to list","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:32Z","receivedAt":"2019-12-07T17:47:57Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nPython 3 handles strings differently than Python 2.7.  Since Python 2\nis reaching it's end of life, a series of changes are being submitted to\nenable python 3.5 and following support. The current code fails basic\ntests under python 3.5.\n\nThe original code used 'basestring' in a test to determine if a list or\nliteral string was passed into 9 different functions.  This is used to\ndetermine if the shell should be invoked when calling subprocess\nmethods.\n\nChange references to 'basestring' in the isinstance tests to use 'list'\ninstead. This prepares the code to remove all references to basestring.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 18 +++++++++---------\n 1 file changed, 9 insertions(+), 9 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 65e926758c..3153186df0 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -108,7 +108,7 @@ def p4_build_cmd(cmd):\n         # Provide a way to not pass this option by setting git-p4.retries to 0\n         real_cmd += [\"-r\", str(retries)]\n \n-    if isinstance(cmd,basestring):\n+    if not isinstance(cmd, list):\n         real_cmd = ' '.join(real_cmd) + ' ' + cmd\n     else:\n         real_cmd += cmd\n@@ -174,7 +174,7 @@ def write_pipe(c, stdin):\n     if verbose:\n         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n \n-    expand = isinstance(c,basestring)\n+    expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdin=subprocess.PIPE, shell=expand)\n     pipe = p.stdin\n     val = pipe.write(stdin)\n@@ -196,7 +196,7 @@ def read_pipe_full(c):\n     if verbose:\n         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n \n-    expand = isinstance(c,basestring)\n+    expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=expand)\n     (out, err) = p.communicate()\n     return (p.returncode, out, err)\n@@ -232,7 +232,7 @@ def read_pipe_lines(c):\n     if verbose:\n         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n \n-    expand = isinstance(c, basestring)\n+    expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, shell=expand)\n     pipe = p.stdout\n     val = pipe.readlines()\n@@ -275,7 +275,7 @@ def p4_has_move_command():\n     return True\n \n def system(cmd, ignore_error=False):\n-    expand = isinstance(cmd,basestring)\n+    expand = not isinstance(cmd, list)\n     if verbose:\n         sys.stderr.write(\"executing %s\\n\" % str(cmd))\n     retcode = subprocess.call(cmd, shell=expand)\n@@ -287,7 +287,7 @@ def system(cmd, ignore_error=False):\n def p4_system(cmd):\n     \"\"\"Specifically invoke p4 as the system command. \"\"\"\n     real_cmd = p4_build_cmd(cmd)\n-    expand = isinstance(real_cmd, basestring)\n+    expand = not isinstance(real_cmd, list)\n     retcode = subprocess.call(real_cmd, shell=expand)\n     if retcode:\n         raise CalledProcessError(retcode, real_cmd)\n@@ -525,7 +525,7 @@ def getP4OpenedType(file):\n # Return the set of all p4 labels\n def getP4Labels(depotPaths):\n     labels = set()\n-    if isinstance(depotPaths,basestring):\n+    if not isinstance(depotPaths, list):\n         depotPaths = [depotPaths]\n \n     for l in p4CmdList([\"labels\"] + [\"%s...\" % p for p in depotPaths]):\n@@ -612,7 +612,7 @@ def isModeExecChanged(src_mode, dst_mode):\n def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n         errors_as_exceptions=False):\n \n-    if isinstance(cmd,basestring):\n+    if not isinstance(cmd, list):\n         cmd = \"-G \" + cmd\n         expand = True\n     else:\n@@ -629,7 +629,7 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n     stdin_file = None\n     if stdin is not None:\n         stdin_file = tempfile.TemporaryFile(prefix='p4-stdin', mode=stdin_mode)\n-        if isinstance(stdin,basestring):\n+        if not isinstance(stdin, list):\n             stdin_file.write(stdin)\n         else:\n             for i in stdin:\n-- \ngitgitgadget\n\n"},{"id":"387684","messageId":"e28fe095b41753ecf57c21984913bc2a26f7aa53.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 06/15] git-p4: remove p4_write_pipe() and write_pipe() return values","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:34Z","receivedAt":"2019-12-07T17:47:57Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThe git-p4 functions write_pipe() and p4_write_pipe() originally\nreturn the number of bytes returned from the system call. However,\nthis is a misleading value when this function is run by Python 3.\n\nModify the functions write_pipe() and p4_write_pipe() to remove the\nreturn value.  The return value for both functions is the number of\nbytes, but the meaning is lost under python3 since the count does not\nmatch the number of characters that may have been encoded.\nAdditionally, the return value was never used, so this is removed to\navoid future ambiguity.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 7 +++++--\n 1 file changed, 5 insertions(+), 2 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex cc6c490e2c..e7c24817ad 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -171,6 +171,8 @@ def die(msg):\n         sys.exit(1)\n \n def write_pipe(c, stdin):\n+    \"\"\" Executes the command 'c', passing 'stdin' on the standard input\n+    \"\"\"\n     if verbose:\n         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n \n@@ -182,11 +184,12 @@ def write_pipe(c, stdin):\n     if p.wait():\n         die('Command failed: %s' % str(c))\n \n-    return val\n \n def p4_write_pipe(c, stdin):\n+    \"\"\" Runs a P4 command 'c', passing 'stdin' data to P4\n+    \"\"\"\n     real_cmd = p4_build_cmd(c)\n-    return write_pipe(real_cmd, stdin)\n+    write_pipe(real_cmd, stdin)\n \n def read_pipe_full(c):\n     \"\"\" Read output from  command. Returns a tuple\n-- \ngitgitgadget\n\n"},{"id":"387687","messageId":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v4.git.1575498577.gitgitgadget@gmail.com","subject":"[PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:28Z","receivedAt":"2019-12-07T17:47:57Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"Issue: The current git-p4.py script does not work with python3.\n\nI have attempted to use the P4 integration built into GIT and I was unable\nto get the program to run because I have Python 3.8 installed on my\ncomputer. I was able to get the program to run when I downgraded my python\nto version 2.7. However, python 2 is reaching its end of life.\n\nSubmission: I am submitting a patch for the git-p4.py script that partially \nsupports python 3.8. This code was able to pass the basic tests (t9800) when\nrun against Python3. This provides basic functionality. \n\nIn an attempt to pass the t9822 P4 path-encoding test, a new parameter for\ngit P4 Clone was introduced. \n\n--encoding Format-identifier\n\nThis will create the GIT repository following the current functionality;\nhowever, before importing the files from P4, it will set the\ngit-p4.pathEncoding option so any files or paths that are encoded with\nnon-ASCII/non-UTF-8 formats will import correctly.\n\nTechnical details: The script was updated by futurize (\nhttps://python-future.org/futurize.html) to support Py2/Py3 syntax. The few\nreferences to classes in future were reworked so that future would not be\nrequired. The existing code test for Unicode support was extended to\nnormalize the classes “unicode” and “bytes” to across platforms:\n\n * ‘unicode’ is an alias for ‘str’ in Py3 and is the unicode class in Py2.\n * ‘bytes’ is bytes in Py3 and an alias for ‘str’ in Py2.\n\nNew coercion methods were written for both Python2 and Python3:\n\n * as_string(text) – In Python3, this encodes a bytes object as a UTF-8\n   encoded Unicode string. \n * as_bytes(text) – In Python3, this decodes a Unicode string to an array of\n   bytes.\n\nIn Python2, these functions do not change the data since a ‘str’ object\nfunction in both roles as strings and byte arrays. This reduces the\npotential impact on backward compatibility with Python 2.\n\n * to_unicode(text) – ensures that the supplied data is encoded as a UTF-8\n   string. This function will encode data in both Python2 and Python3. * \n      path_as_string(path) – This function is an extension function that\n      honors the option “git-p4.pathEncoding” to convert a set of bytes or\n      characters to UTF-8. If the str/bytes cannot decode as ASCII, it will\n      use the encodeWithUTF8() method to convert the custom encoded bytes to\n      Unicode in UTF-8.\n   \n   \n\nGenerally speaking, information in the script is converted to Unicode as\nearly as possible and converted back to a byte array just before passing to\nexternal programs or files. The exception to this rule is P4 Repository file\npaths.\n\nPaths are not converted but left as “bytes” so the original file path\nencoding can be preserved. This formatting is required for commands that\ninteract with the P4 file path. When the file path is used by GIT, it is\nconverted with encodeWithUTF8().\n\nSigned-off-by: Ben Keene seraphire@gmail.com [seraphire@gmail.com]\n\nBen Keene (13):\n  git-p4: select P4 binary by operating-system\n  git-p4: change the expansion test from basestring to list\n  git-p4: promote encodeWithUTF8() to a global function\n  git-p4: remove p4_write_pipe() and write_pipe() return values\n  git-p4: add new support function gitConfigSet()\n  git-p4: add casting helper functions for python 3 conversion\n  git-p4: python 3 syntax changes\n  git-p4: fix assumed path separators to be more Windows friendly\n  git-p4: add Py23File() - helper class for stream writing\n  git-p4: p4CmdList - support Unicode encoding\n  git-p4: support Python 3 for basic P4 clone, sync, and submit (t9800)\n  git-p4: added --encoding parameter to p4 clone\n  git-p4: Add depot manipulation functions\n\nJeff King (2):\n  t/gitweb-lib.sh: drop confusing quotes\n  t/gitweb-lib.sh: set $REQUEST_URI\n\n Documentation/git-p4.txt        |   5 +\n git-p4.py                       | 768 +++++++++++++++++++++++++-------\n t/t9822-git-p4-path-encoding.sh | 101 +++++\n 3 files changed, 706 insertions(+), 168 deletions(-)\n\n\nbase-commit: 083378cc35c4dbcc607e4cdd24a5fca440163d17\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-463%2Fseraphire%2Fseraphire%2Fp4-python3-unicode-v5\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-463/seraphire/seraphire/p4-python3-unicode-v5\nPull-Request: https://github.com/gitgitgadget/git/pull/463\n\nRange-diff vs v4:\n\n  -:  ---------- >  1:  bb7f8f0a0a t/gitweb-lib.sh: drop confusing quotes\n  -:  ---------- >  2:  a7a4c5a2aa t/gitweb-lib.sh: set $REQUEST_URI\n  1:  4012426993 !  3:  e425ccc10f git-p4: select p4 binary by operating-system\n     @@ -1,19 +1,18 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: select p4 binary by operating-system\n     -\n     -    Depending on the version of GIT and Python installed, the perforce program (p4) may not resolve on Windows without the program extension.\n     -\n     -    Check the operating system (platform.system) and if it is reporting that it is Windows, use the full filename of \"p4.exe\" instead of \"p4\"\n     +    git-p4: select P4 binary by operating-system\n      \n          The original code unconditionally used \"p4\" as the binary filename.\n      \n     -    This change is Python2 and Python3 compatible.\n     +    Depending on the version of Git and Python installed, the perforce\n     +    program (p4) may not resolve on Windows without the program extension.\n     +\n     +    Check the operating system (platform.system) and if it is reporting that\n     +    it is Windows, use the full filename of \"p4.exe\" instead of \"p4\"\n      \n     -    Thanks to: Junio C Hamano <gitster@pobox.com> and  Denton Liu <liu.denton@gmail.com> for patiently explaining proper format for my submissions.\n     +    This change is Python 2 and Python 3 compatible.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit 9a3a5c4e6d29dbef670072a9605c7a82b3729434)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n     @@ -23,9 +22,8 @@\n           can be done more easily.\n           \"\"\"\n      -    real_cmd = [\"p4\"]\n     -+    # Look for the P4 binary\n      +    if (platform.system() == \"Windows\"):\n     -+        real_cmd = [\"p4.exe\"]    \n     ++        real_cmd = [\"p4.exe\"]\n      +    else:\n      +        real_cmd = [\"p4\"]\n       \n  2:  0ef2f56b04 !  4:  7170aface2 git-p4: change the expansion test from basestring to list\n     @@ -2,14 +2,20 @@\n      \n          git-p4: change the expansion test from basestring to list\n      \n     -    Python 3+ handles strings differently than Python 2.7.  Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n     +    Python 3 handles strings differently than Python 2.7.  Since Python 2\n     +    is reaching it's end of life, a series of changes are being submitted to\n     +    enable python 3.5 and following support. The current code fails basic\n     +    tests under python 3.5.\n      \n     -    Change references to basestring in the isinstance tests to use list instead. This prepares the code to remove all references to basestring.\n     +    The original code used 'basestring' in a test to determine if a list or\n     +    literal string was passed into 9 different functions.  This is used to\n     +    determine if the shell should be invoked when calling subprocess\n     +    methods.\n      \n     -    The original code used basestring in a test to determine if a list or literal string was passed into 9 different functions.  This is used to determine if the shell should be evoked when calling subprocess methods.\n     +    Change references to 'basestring' in the isinstance tests to use 'list'\n     +    instead. This prepares the code to remove all references to basestring.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit 5b1b1c145479b5d5fd242122737a3134890409e6)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n  5:  1bf7b073b0 !  5:  11d7703e41 git-p4: Add new functions in preparation of usage\n     @@ -1,55 +1,53 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: Add new functions in preparation of usage\n     +    git-p4: promote encodeWithUTF8() to a global function\n      \n     -    This changelist is an intermediate submission for migrating the P4 support from Python2 to Python3. The code needs access to the encodeWithUTF8() for support of non-UTF8 filenames in the clone class as well as the sync class.\n     +    This changelist is an intermediate submission for migrating the P4\n     +    support from Python 2 to Python 3. The code needs access to the\n     +    encodeWithUTF8() for support of non-UTF8 filenames in the clone class as\n     +    well as the sync class.\n      \n     -    Move the function encodeWithUTF8() from the P4Sync class to a stand-alone function.  This will allow other classes to use this function without instanciating the P4Sync class. Change the self.verbose reference to an optional method parameter. Update the existing references to this function to pass the self.verbose since it is no longer available on \"self\" since the function is no longer contained on the P4Sync class.\n     -\n     -    Modify the functions write_pipe() and p4_write_pipe() to remove the return value.  The return value for both functions is the number of bytes, but the meaning is lost under python3 since the count does not match the number of characters that may have been encoded.  Additionally, the return value was never used, so this is removed to avoid future ambiguity.\n     -\n     -    Add a new method gitConfigSet(). This method will set a value in the git configuration cache list.\n     +    Move the function encodeWithUTF8() from the P4Sync class to a\n     +    stand-alone function.  This will allow other classes to use this\n     +    function without instanciating the P4Sync class. Change the self.verbose\n     +    reference to an optional method parameter. Update the existing\n     +    references to this function to pass the self.verbose since it is no\n     +    longer available on \"self\" since the function is no longer contained on\n     +    the P4Sync class.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit affe888f432bb6833df78962e8671fccdf76c47a)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n       +++ b/git-p4.py\n      @@\n     -         sys.exit(1)\n     - \n     - def write_pipe(c, stdin):\n     -+    \"\"\" Executes the command 'c', passing 'stdin' on the standard input\n     -+    \"\"\"\n     -     if verbose:\n     -         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n     + import ctypes\n     + import errno\n       \n     +-# support basestring in python3\n     ++# support basestring in Python 3\n     + try:\n     +     unicode = unicode\n     + except NameError:\n      @@\n     -     if p.wait():\n     -         die('Command failed: %s' % str(c))\n     - \n     --    return val\n     - \n     - def p4_write_pipe(c, stdin):\n     -+    \"\"\" Runs a P4 command 'c', passing 'stdin' data to P4\n     -+    \"\"\"\n     -     real_cmd = p4_build_cmd(c)\n     --    return write_pipe(real_cmd, stdin)\n     -+    write_pipe(real_cmd, stdin)\n     - \n     - def read_pipe_full(c):\n     -     \"\"\" Read output from  command. Returns a tuple\n     + try:\n     +     from subprocess import CalledProcessError\n     + except ImportError:\n     +-    # from python2.7:subprocess.py\n     ++    # from Python 2.7:subprocess.py\n     +     # Exception classes used by this module.\n     +     class CalledProcessError(Exception):\n     +         \"\"\"This exception is raised when a process run by check_call() returns\n      @@\n           # otherwise False.\n           return mode[-3:] == \"755\"\n       \n     -+def encodeWithUTF8(path, verbose = False):\n     ++def encodeWithUTF8(path, verbose=False):\n      +    \"\"\" Ensure that the path is encoded as a UTF-8 string\n      +\n      +        Returns bytes(P3)/str(P2)\n      +    \"\"\"\n     -+   \n     ++\n      +    if isunicode:\n      +        try:\n      +            if isinstance(path, unicode):\n     @@ -64,7 +62,7 @@\n      +            path = path.decode(encoding, 'replace').encode('utf8', 'replace')\n      +            if verbose:\n      +                print('\\nNOTE:Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, to_unicode(path)))\n     -+    else:    \n     ++    else:\n      +        try:\n      +            path.decode('ascii')\n      +        except:\n     @@ -79,18 +77,6 @@\n       class P4Exception(Exception):\n           \"\"\" Base class for exceptions from the p4 client \"\"\"\n           def __init__(self, exit_code):\n     -@@\n     -             _gitConfig[key] = []\n     -     return _gitConfig[key]\n     - \n     -+def gitConfigSet(key, value):\n     -+    \"\"\" Set the git configuration key 'key' to 'value' for this session\n     -+    \"\"\"\n     -+    _gitConfig[key] = value\n     -+\n     - def p4BranchesInGit(branchesAreInRemotes=True):\n     -     \"\"\"Find all the branches whose names start with \"p4/\", looking\n     -        in remotes or heads as specified by the argument.  Return\n      @@\n                   self.gitStream.write(d)\n               self.gitStream.write('\\n')\n  9:  4fc49313f0 !  6:  e28fe095b4 git-p4: Add usability enhancements\n     @@ -1,75 +1,44 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: Add usability enhancements\n     +    git-p4: remove p4_write_pipe() and write_pipe() return values\n      \n     -    Issue: when prompting the user with raw_input, the tests are not forgiving of user input.  For example, on the first query asks for a yes/no response. If the user enters the full word \"yes\" or \"no\" the test will fail. Additionally, offer the suggestion of setting git-p4.attemptRCSCleanup when applying a commit fails because of RCS keywords. Both of these changes are usability enhancement suggestions.\n     +    The git-p4 functions write_pipe() and p4_write_pipe() originally\n     +    return the number of bytes returned from the system call. However,\n     +    this is a misleading value when this function is run by Python 3.\n      \n     -    Change the code prompting the user for input to sanitize the user input before checking the response by asking the response as a lower case string, trimming leading/trailing spaces, and returning the first character.\n     -\n     -    Change the applyCommit() method that when applying a commit fails becasue of the P4 RCS Keywords, the user should consider setting git-p4.attemptRCSCleanup.\n     +    Modify the functions write_pipe() and p4_write_pipe() to remove the\n     +    return value.  The return value for both functions is the number of\n     +    bytes, but the meaning is lost under python3 since the count does not\n     +    match the number of characters that may have been encoded.\n     +    Additionally, the return value was never used, so this is removed to\n     +    avoid future ambiguity.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit 1fab571664f5b6ad4ef321199f52615a32a9f8c7)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n       +++ b/git-p4.py\n      @@\n     -             return True\n     +         sys.exit(1)\n       \n     -         while True:\n     --            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \")\n     -+            response = raw_input(\"Submit template unchanged. Submit anyway? [y]es, [n]o (skip this patch) \").lower() \\\n     -+                .strip()[0]\n     -             if response == 'y':\n     -                 return True\n     -             if response == 'n':\n     -@@\n     -                     # disable the read-only bit on windows.\n     -                     if self.isWindows and file not in editedFiles:\n     -                         os.chmod(file, stat.S_IWRITE)\n     --                    self.patchRCSKeywords(file, kwfiles[file])\n     --                    fixed_rcs_keywords = True\n     -+                    \n     -+                    try:\n     -+                        self.patchRCSKeywords(file, kwfiles[file])\n     -+                        fixed_rcs_keywords = True\n     -+                    except:\n     -+                        # We are throwing an exception, undo all open edits\n     -+                        for f in editedFiles:\n     -+                            p4_revert(f)\n     -+                        raise\n     -+            else:\n     -+                # They do not have attemptRCSCleanup set, this might be the fail point\n     -+                # Check to see if the file has RCS keywords and suggest setting the property.\n     -+                for file in editedFiles | filesToDelete:\n     -+                    if p4_keywords_regexp_for_file(file) != None:\n     -+                        print(\"At least one file in this commit has RCS Keywords that may be causing problems. \")\n     -+                        print(\"Consider:\\ngit config git-p4.attemptRCSCleanup true\")\n     -+                        break\n     + def write_pipe(c, stdin):\n     ++    \"\"\" Executes the command 'c', passing 'stdin' on the standard input\n     ++    \"\"\"\n     +     if verbose:\n     +         sys.stderr.write('Writing pipe: %s\\n' % str(c))\n       \n     -             if fixed_rcs_keywords:\n     -                 print(\"Retrying the patch with RCS keywords cleaned up\")\n     -@@\n     -                         if self.conflict_behavior == \"ask\":\n     -                             print(\"What do you want to do?\")\n     -                             response = raw_input(\"[s]kip this commit but apply\"\n     --                                                 \" the rest, or [q]uit? \")\n     -+                                                 \" the rest, or [q]uit? \").lower().strip()[0]\n     -                             if not response:\n     -                                 continue\n     -                         elif self.conflict_behavior == \"skip\":\n      @@\n     -                                    description = cmd.description,\n     -                                    formatter = HelpFormatter())\n     +     if p.wait():\n     +         die('Command failed: %s' % str(c))\n     + \n     +-    return val\n     + \n     + def p4_write_pipe(c, stdin):\n     ++    \"\"\" Runs a P4 command 'c', passing 'stdin' data to P4\n     ++    \"\"\"\n     +     real_cmd = p4_build_cmd(c)\n     +-    return write_pipe(real_cmd, stdin)\n     ++    write_pipe(real_cmd, stdin)\n       \n     --    (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n     -+    try:\n     -+        (cmd, args) = parser.parse_args(sys.argv[2:], cmd);\n     -+    except:\n     -+        parser.print_help()\n     -+        raise\n     -+\n     -     global verbose\n     -     verbose = cmd.verbose\n     -     if cmd.needsGit:\n     + def read_pipe_full(c):\n     +     \"\"\" Read output from  command. Returns a tuple\n  -:  ---------- >  7:  bc7009541b git-p4: add new support function gitConfigSet()\n  3:  f0e658b984 !  8:  1e677781d2 git-p4: add new helper functions for python3 conversion\n     @@ -1,31 +1,54 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: add new helper functions for python3 conversion\n     +    git-p4: add casting helper functions for python 3 conversion\n      \n     -    Python 3+ handles strings differently than Python 2.7.  Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n     +    Python 3 handles strings differently than Python 2.7.  Since Python 2\n     +    is reaching it's end of life, a series of changes are being submitted to\n     +    enable python 3.5 and following support. The current code fails basic\n     +    tests under python 3.5.\n      \n     -    Change the existing unicode test add new support functions for python2-python3 support.\n     +    Change the existing unicode test add new support functions for\n     +    Python 2 - Python 3 support.\n      \n          Define the following variables:\n     -    - isunicode - a boolean variable that states if the version of python natively supports unicode (true) or not (false). This is true for Python3 and false for Python2.\n     -    - unicode - a type alias for the datatype that holds a unicode string.  It is assigned to a str under python 3 and the unicode type for Python2.\n     -    - bytes - a type alias for an array of bytes.  It is assigned the native bytes type for Python3 and str for Python2.\n     +    - isunicode - a boolean variable that states if the version of python\n     +                  natively supports unicode (true) or not (false). This is\n     +                  true for Python 3 and false for Python 2.\n     +    - unicode   - a type alias for the datatype that holds a unicode string.\n     +                  It is assigned to a str under Python 3 and the unicode\n     +                  type for Python 2.\n     +    - bytes     - a type alias for an array of bytes.  It is assigned the\n     +                  native bytes type for Python 3 and str for Python 2.\n      \n          Add the following new functions:\n      \n     -    - as_string(text) - A new function that will convert a byte array to a unicode (UTF-8) string under python 3.  Under python 2, this returns the string unchanged.\n     -    - as_bytes(text) - A new function that will convert a unicode string to a byte array under python 3.  Under python 2, this returns the string unchanged.\n     -    - to_unicode(text) - Converts a text string as Unicode(UTF-8) on both Python2 and Python3.\n     +    - as_string(text)  - A new function that will convert a byte array to a\n     +                         unicode (UTF-8) string under Python 3.  Under\n     +                         Python 2, this returns the string unchanged.\n     +    - as_bytes(text)   - A new function that will convert a unicode string\n     +                         to a byte array under Python 3.  Under Python 2,\n     +                         this returns the string unchanged.\n     +    - to_unicode(text) - Converts a text string as Unicode(UTF-8) on both\n     +                         Python 2 and Python 3.\n      \n          Add a new function alias raw_input:\n     -    If raw_input does not exist (it was renamed to input in python 3) alias input as raw_input.\n     +    If raw_input does not exist (it was renamed to input in Python 3) alias\n     +    input as raw_input.\n      \n     -    The AS_STRING and AS_BYTES functions allow for modifying the code with a minimal amount of impact on Python2 support.  When a string is expected, the as_string() will be used to convert \"cast\" the incoming \"bytes\" to a string type. Conversely as_bytes() will be used to convert a \"string\" to a \"byte array\" type. Since Python2 overloads the datatype 'str' to serve both purposes, the Python2 versions of these function do not change the data, since the str functions as both a byte array and a string.\n     +    The as_string() and as_bytes() functions allow for modifying the code\n     +    with a minimal amount of impact on Python 2 support. When a string is\n     +    expected, the as_string() will be used to \"cast\" the incoming \"bytes\"\n     +    to a string type.\n      \n     -    basestring is removed since its only references are found in tests that were changed in the previous change list.\n     +    Conversely as_bytes() will be used to cast a \"string\" to a \"byte array\"\n     +    type. Since Python 2 overloads the datatype 'str' to serve both purposes,\n     +    the Python 2 versions of these function do not change the data. This\n     +    reduces the regression impact of these code changes.\n     +\n     +    'basestring' is removed since its only references are found in tests\n     +    that were changed in modified in previous commits.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit 7921aeb3136b07643c1a503c2d9d8b5ada620356)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n     @@ -36,7 +59,7 @@\n           # 'unicode' is undefined, must be Python 3\n      -    str = str\n      +    #\n     -+    # For Python3 which is natively unicode, we will use \n     ++    # For Python 3 which is natively unicode, we will use\n      +    # unicode for internal information but all P4 Data\n      +    # will remain in bytes\n      +    isunicode = True\n     @@ -45,8 +68,9 @@\n      -    basestring = (str,bytes)\n      +\n      +    def as_string(text):\n     -+        \"\"\"Return a byte array as a unicode string\"\"\"\n     -+        if text == None:\n     ++        \"\"\" Return a byte array as a unicode string\n     ++        \"\"\"\n     ++        if text is None:\n      +            return None\n      +        if isinstance(text, bytes):\n      +            return unicode(text, \"utf-8\")\n     @@ -54,8 +78,9 @@\n      +            return text\n      +\n      +    def as_bytes(text):\n     -+        \"\"\"Return a Unicode string as a byte array\"\"\"\n     -+        if text == None:\n     ++        \"\"\" Return a Unicode string as a byte array\n     ++        \"\"\"\n     ++        if text is None:\n      +            return None\n      +        if isinstance(text, bytes):\n      +            return text\n     @@ -63,15 +88,17 @@\n      +            return bytes(text, \"utf-8\")\n      +\n      +    def to_unicode(text):\n     -+        \"\"\"Return a byte array as a unicode string\"\"\"\n     -+        return as_string(text)    \n     ++        \"\"\" Return a byte array as a unicode string\n     ++        \"\"\"\n     ++        return as_string(text)\n      +\n      +    def path_as_string(path):\n     -+        \"\"\" Converts a path to the UTF8 encoded string \"\"\"\n     ++        \"\"\" Converts a path to the UTF8 encoded string\n     ++        \"\"\"\n      +        if isinstance(path, unicode):\n      +            return path\n      +        return encodeWithUTF8(path).decode('utf-8')\n     -+    \n     ++\n       else:\n           # 'unicode' exists, must be Python 2\n      -    str = str\n     @@ -79,7 +106,7 @@\n      +    # We will treat the data as:\n      +    #   str   -> str\n      +    #   bytes -> str\n     -+    # So for Python2 these functions are no-ops\n     ++    # So for Python 2 these functions are no-ops\n      +    # and will leave the data in the ambiguious\n      +    # string/bytes state\n      +    isunicode = False\n     @@ -88,23 +115,25 @@\n      -    basestring = basestring\n      +\n      +    def as_string(text):\n     -+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n     ++        \"\"\" Return text unaltered (for Python 3 support)\n     ++        \"\"\"\n      +        return text\n      +\n      +    def as_bytes(text):\n     -+        \"\"\" Return text unaltered (for Python3 support) \"\"\"\n     ++        \"\"\" Return text unaltered (for Python 3 support)\n     ++        \"\"\"\n      +        return text\n      +\n      +    def to_unicode(text):\n     -+        \"\"\"Return a string as a unicode string\"\"\"\n     ++        \"\"\" Return a string as a unicode string\n     ++        \"\"\"\n      +        return text.decode('utf-8')\n     -+    \n     ++\n      +    def path_as_string(path):\n     -+        \"\"\" Converts a path to the UTF8 encoded bytes \"\"\"\n     ++        \"\"\" Converts a path to the UTF8 encoded bytes\n     ++        \"\"\"\n      +        return encodeWithUTF8(path)\n      +\n     -+\n     -+ \n      +# Check for raw_input support\n      +try:\n      +    raw_input\n     @@ -113,3 +142,21 @@\n       \n       try:\n           from subprocess import CalledProcessError\n     +@@\n     +             if data[:space] == depotPath:\n     +                 output = entry\n     +                 break\n     +-    if output == None:\n     ++    if output is None:\n     +         return \"\"\n     +     if output[\"code\"] == \"error\":\n     +         return \"\"\n     +@@\n     +     global verbose\n     +     verbose = cmd.verbose\n     +     if cmd.needsGit:\n     +-        if cmd.gitdir == None:\n     ++        if cmd.gitdir is None:\n     +             cmd.gitdir = os.path.abspath(\".git\")\n     +             if not isValidGitDir(cmd.gitdir):\n     +                 # \"rev-parse --git-dir\" without arguments will try $PWD/.git\n  4:  3c41db3e91 !  9:  a221eb8bb6 git-p4: python3 syntax changes\n     @@ -1,23 +1,33 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: python3 syntax changes\n     +    git-p4: python 3 syntax changes\n      \n     -    Python 3+ handles strings differently than Python 2.7.  Since Python 2 is reaching it's end of life, a series of changes are being submitted to enable python 3.7+ support. The current code fails basic tests under python 3.7.\n     +    Python 3 handles strings differently than Python 2.7.  Since Python 2\n     +    is reaching it's end of life, a series of changes are being submitted to\n     +    enable python 3.5 and following support. The current code fails basic\n     +    tests under python 3.5.\n      \n     -    There are a number of translations suggested by modernize/futureize that should be taken to fix numerous non-string specific issues.\n     +    There are a number of translations suggested by modernize/futureize that\n     +    should be taken to fix numerous non-string specific issues.\n      \n     -    Change references to the X.next() iterator to the function next(X) which is compatible with both Python2 and Python3.\n     +    Change references to the X.next() iterator to the function next(X) which\n     +    is compatible with both Python2 and Python3.\n      \n     -    Change references to X.keys() to list(X.keys()) to return a list that can be iterated in both Python2 and Python3.\n     +    Change references to X.keys() to list(X.keys()) to return a list that\n     +    can be iterated in both Python2 and Python3.\n      \n     -    Add the literal text (object) to the end of class definitions to be consistent with Python3 class definition.\n     +    Add the literal text (object) to the end of class definitions to be\n     +    consistent with Python3 class definition.\n      \n     -    Change integer divison to use \"//\" instead of \"/\"  Under Both python2 and python3 // will return a floor()ed result which matches existing functionality.\n     +    Change integer divison to use \"//\" instead of \"/\"  Under Both Python 2\n     +    and Python 3 // will return a floor()ed result which matches existing\n     +    functionality.\n      \n     -    Change the format string for displaying decimal values from %d to %4.1f% when displaying a progress.  This avoids displaying long repeating decimals in user displayed text.\n     +    Change the format string for displaying decimal values from %d to %4.1f%\n     +    when displaying a progress.  This avoids displaying long repeating\n     +    decimals in user displayed text.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit bde6b83296aa9b3e7a584c5ce2b571c7287d8f9f)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n     @@ -30,7 +40,7 @@\n      +import codecs\n      +import io\n       \n     - # support basestring in python3\n     + # support basestring in Python 3\n       try:\n      @@\n       \n  6:  8f5752c127 ! 10:  b962cce8cd git-p4: Fix assumed path separators to be more Windows friendly\n     @@ -1,19 +1,30 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: Fix assumed path separators to be more Windows friendly\n     +    git-p4: fix assumed path separators to be more Windows friendly\n      \n     -    When a computer is configured to use Git for windows and Python for windows, and not a Unix subsystem like cygwin or WSL, the directory separator changes and causes git-p4 to fail to properly determine paths.\n     +    When a computer is configured to use Git for windows and Python for\n     +    windows, and not a Unix subsystem like cygwin or WSL, the directory\n     +    separator changes and causes git-p4 to fail to properly determine paths.\n      \n          Fix 3 path separator errors:\n      \n     -    1. getUserCacheFilename should not use string concatenation. Change this code to use os.path.join to build an OS tolerant path.\n     -    2. defaultDestiantion used the OS.path.split to split depot paths.  This is incorrect on windows. Change the code to split on a forward slash(/) instead since depot paths use this character regardless  of the operating system.\n     -    3. The call to isvalidGitDir() in the main code also used a literal forward slash. Change the cose to use os.path.join to correctly format the path for the operating system.\n     +    1. getUserCacheFilename() - should not use string concatenation. Change\n     +       this code to use os.path.join to build an OS tolerant path.\n      \n     -    These three changes allow the suggested windows configuration to properly locate files while retaining the existing behavior on non-windows operating systems.\n     +    2. defaultDestiantion used the OS.path.split to split depot paths.  This\n     +       is incorrect on windows. Change the code to split on a forward\n     +       slash(/) instead since depot paths use this character regardless  of\n     +       the operating system.\n     +\n     +    3. The call to isValidGitDir() in the main code also used a literal\n     +       forward slash. Change the code to use os.path.join to correctly\n     +       format the path for the operating system.\n     +\n     +    These three changes allow the suggested windows configuration to\n     +    properly locate files while retaining the existing behavior on\n     +    non-windows operating systems.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit a5b45c12c3861638a933b05a1ffee0c83978dcb2)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n     @@ -22,8 +33,8 @@\n                   return True\n       \n           def getUserCacheFilename(self):\n     -+        \"\"\" Returns the filename of the username cache \n     -+\t    \"\"\"\n     ++        \"\"\" Returns the filename of the username cache\n     ++        \"\"\"\n               home = os.environ.get(\"HOME\", os.environ.get(\"USERPROFILE\"))\n      -        return home + \"/.gitp4-usercache.txt\"\n      +        return os.path.join(home, \".gitp4-usercache.txt\")\n     @@ -34,7 +45,7 @@\n               self.cloneBare = False\n       \n           def defaultDestination(self, args):\n     -+        \"\"\" Returns the last path component as the default git \n     ++        \"\"\" Returns the last path component as the default git\n      +            repository directory name\n      +        \"\"\"\n               ## TODO: use common prefix of args?\n  7:  10dc059444 ! 11:  d22ada1614 git-p4: Add a helper class for stream writing\n     @@ -1,25 +1,43 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: Add a helper class for stream writing\n     +    git-p4: add Py23File() - helper class for stream writing\n      \n     -    This is a transtional commit that does not change current behvior.  It adds a new class Py23File.\n     +    This is a preparatory commit that does not change current behavior.\n     +    It adds a new class Py23File.\n      \n     -    Following the Python recommendation of keeping text as unicode internally and only converting to and from bytes on input and output, this class provides an interface for the methods used for reading and writing files and file like streams.\n     +    Following the Python recommendation of keeping text as unicode\n     +    internally and only converting to and from bytes on input and output,\n     +    this class provides an interface for the methods used for reading and\n     +    writing files and file like streams.\n      \n     -    Create a class that wraps the input and output functions used by the git-p4.py code for reading and writing to standard file handles.\n     +    A new class was implemented to avoid requiring additional dependencies.\n      \n     -    The methods of this class should take a Unicode string for writing and return unicode strings in reads.  This class should be a drop-in for existing file like streams\n     +    Create a class that wraps the input and output functions used by the\n     +    git-p4.py code for reading and writing to standard file handles.\n      \n     -    The following methods should be coded for supporting existing read/write calls:\n     -    * write - this should write a Unicode string to the underlying stream\n     -    * read - this should read from the underlying stream and cast the bytes as a unicode string\n     -    * readline - this should read one line of text from the underlying stream and cast it as a unicode string\n     -    * readline - this should read a number of lines, optionally hinted, and cast each line as a unicode string\n     +    The methods of this class should take a Unicode string for writing and\n     +    return unicode strings in reads.  This class should be a drop-in for\n     +    existing file like streams\n      \n     -    The expression \"cast as a unicode string\" is used because the code should use the AS_BYTES() and AS_UNICODE() functions instead of cohercing the data to actual unicode strings or bytes.  This allows python 2 code to continue to use the internal \"str\" data type instead of converting the data back and forth to actual unicode strings. This retains current python2 support while python3 support may be incomplete.\n     +    The following methods should be coded for supporting existing read/write\n     +    calls:\n     +      * write - this should write a Unicode string to the underlying stream\n     +      * read  - this should read from the underlying stream and cast the\n     +                bytes as a unicode string\n     +      * readline - this should read one line of text from the underlying\n     +                stream and cast it as a unicode string\n     +      * readline - this should read a number of lines, optionally hinted,\n     +                and cast each line as a unicode string\n     +\n     +    The expression \"cast as a unicode string\" is used because the code\n     +    should use the as_bytes() and as_string() functions instead of\n     +    cohercing the data to actual unicode strings or bytes.  This allows\n     +    Python 2 code to continue to use the internal \"str\" data type instead\n     +    of converting the data back and forth to actual unicode strings. This\n     +    retains current Python 2 support while Python 3 support may be\n     +    incomplete.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit 12919111fbaa3e4c0c4c2fdd4f79744cc683d860)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n     @@ -29,13 +47,13 @@\n               return True\n       \n      +class Py23File():\n     -+    \"\"\" Python2/3 Unicode File Wrapper \n     ++    \"\"\" Python2/3 Unicode File Wrapper\n      +    \"\"\"\n     -+    \n     ++\n      +    stream_handle = None\n      +    verbose       = False\n      +    debug_handle  = None\n     -+   \n     ++\n      +    def __init__(self, stream_handle, verbose = False):\n      +        \"\"\" Create a Python3 compliant Unicode to Byte String\n      +            Windows compatible wrapper\n     @@ -47,7 +65,7 @@\n      +        self.verbose       = verbose\n      +\n      +    def write(self, utf8string):\n     -+        \"\"\" Writes the utf8 encoded string to the underlying \n     ++        \"\"\" Writes the utf8 encoded string to the underlying\n      +            file stream\n      +        \"\"\"\n      +        self.stream_handle.write(as_bytes(utf8string))\n     @@ -56,7 +74,7 @@\n      +            sys.stderr.flush()\n      +\n      +    def read(self, size = None):\n     -+        \"\"\" Reads int charcters from the underlying stream \n     ++        \"\"\" Reads int charcters from the underlying stream\n      +            and converts it to utf8.\n      +\n      +            Be aware, the size value is for reading the underlying\n     @@ -69,7 +87,7 @@\n      +            return as_string(self.stream_handle.read(size))\n      +\n      +    def readline(self):\n     -+        \"\"\" Reads a line from the underlying byte stream \n     ++        \"\"\" Reads a line from the underlying byte stream\n      +            and converts it to utf8\n      +        \"\"\"\n      +        return as_string(self.stream_handle.readline())\n     @@ -77,8 +95,8 @@\n      +    def readlines(self, sizeHint = None):\n      +        \"\"\" Returns a list containing lines from the file converted to unicode.\n      +\n     -+            sizehint - Optional. If the optional sizehint argument is \n     -+            present, instead of reading up to EOF, whole lines totalling \n     ++            sizehint - Optional. If the optional sizehint argument is\n     ++            present, instead of reading up to EOF, whole lines totalling\n      +            approximately sizehint bytes are read.\n      +        \"\"\"\n      +        lines = self.stream_handle.readlines(sizeHint)\n  8:  e1a424a955 ! 12:  e97ac0af8a git-p4: p4CmdList  - support Unicode encoding\n     @@ -1,19 +1,55 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: p4CmdList  - support Unicode encoding\n     +    git-p4: p4CmdList - support Unicode encoding\n      \n     -    The p4CmdList is a commonly used function in the git-p4 code. It is used to execute a command in P4 and return the results of the call in a list.\n     +    The p4CmdList is a commonly used function in the git-p4 code. It is used\n     +    to execute a command in P4 and return the results of the call in a list.\n      \n     -    Change this code to take a new optional parameter, encode_data that will optionally convert the data AS_STRING() that isto be returned by the function.\n     +    The problem is that p4CmdList takes bytes as the parameter data and\n     +    returns bytes in the return list.\n      \n     -    Change the code so that the key will always be encoded AS_STRING()\n     +    Add a new optional parameter to the signature, encode_cmd_output, that\n     +    determines if the dictionary values returned in the function output are\n     +    treated as bytes or as strings.\n      \n     -    Data that is passed for standard input (stdin) should be AS_BYTES() to ensure unicode text that is supplied will be written out as bytes.\n     +    Change the code to conditionally pass the output data through the\n     +    as_string() function when encode_cmd_output is true. Otherwise the\n     +    function should return the data as bytes.\n      \n     -    Additionally, change literal text prior to conversion to be literal bytes.\n     +    Change the code so that regardless of the setting of encode_cmd_output,\n     +    the dictionary keys in the return value will always be encoded with\n     +    as_string().\n     +\n     +    as_string(bytes) is a method defined in this project that treats the\n     +    byte data as a string. The word \"string\" is used because the meaning\n     +    varies depending on the version of Python:\n     +\n     +      - Python 2: The \"bytes\" are returned as \"str\", functionally a No-op.\n     +      - Python 3: The \"bytes\" are returned as a Unicode string.\n     +\n     +    The p4CmdList function returns a list of dictionaries that contain\n     +    the result of p4 command. If the callback (cb) is defined, the\n     +    standard output of the p4 command is redirected.\n     +\n     +    Data that is passed to the standard input of the P4 process should be\n     +    as_bytes() to avoid conversion unicode encoding errors.\n     +\n     +    as_bytes(text) is a method defined in this project that treats the text\n     +    data as a string that should be converted to a byte array (bytes). The\n     +    behavior of this function depends on the version of python:\n     +\n     +      - Python 2: The \"text\" is returned as \"str\", functionally a No-op.\n     +      - Python 3: The \"text\" is treated as a UTF-8 encoded Unicode string\n     +            and is decoded to bytes.\n     +\n     +    Additionally, change literal text prior to conversion to be literal\n     +    bytes for the code that is evaluating the standard output from the\n     +    p4 call.\n     +\n     +    Add encode_cmd_output to the p4Cmd since this is a helper function that\n     +    wraps the behavior of p4CmdList.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit 88306ac269186cbd0f6dc6cfd366b50b28ee4886)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n     @@ -23,7 +59,7 @@\n       \n       def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n      -        errors_as_exceptions=False):\n     -+        errors_as_exceptions=False, encode_data=True):\n     ++        errors_as_exceptions=False, encode_cmd_output=True):\n      +    \"\"\" Executes a P4 command:  'cmd' optionally passing 'stdin' to the command's\n      +        standard input via a temporary file with 'stdin_mode' mode.\n      +\n     @@ -37,7 +73,7 @@\n      +        If 'errors_as_exceptions' is set to true (the default is false) the error\n      +        code returned from the execution will generate an exception.\n      +\n     -+        If 'encode_data' is set to true (the default) the data that is returned \n     ++        If 'encode_cmd_output' is set to true (the default) the data that is returned\n      +        by this function will be passed through the \"as_string\" function.\n      +    \"\"\"\n       \n     @@ -65,8 +101,38 @@\n      -                result.append(entry)\n      +                out = {}\n      +                for key, value in entry.items():\n     -+                    out[as_string(key)] = (as_string(value) if encode_data else value)\n     ++                    out[as_string(key)] = (as_string(value) if encode_cmd_output else value)\n      +                result.append(out)\n           except EOFError:\n               pass\n           exitCode = p4.wait()\n     +@@\n     + \n     +     return result\n     + \n     +-def p4Cmd(cmd):\n     +-    list = p4CmdList(cmd)\n     ++def p4Cmd(cmd, encode_cmd_output=True):\n     ++    \"\"\"Executes a P4 command and returns the results in a dictionary\"\"\"\n     ++    list = p4CmdList(cmd, encode_cmd_output=encode_cmd_output)\n     +     result = {}\n     +     for entry in list:\n     +         result.update(entry)\n     +@@\n     +     \"\"\"Look at the p4 client spec, create a View() object that contains\n     +        all the mappings, and return it.\"\"\"\n     + \n     +-    specList = p4CmdList(\"client -o\")\n     ++    specList = p4CmdList(\"client -o\", encode_cmd_output=False)\n     +     if len(specList) != 1:\n     +         die('Output from \"client -o\" is %d lines, expecting 1' %\n     +             len(specList))\n     +@@\n     +         if len(fileArgs) == 0:\n     +             return  # All files in cache\n     + \n     +-        where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs)\n     ++        where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs, encode_cmd_output=False)\n     +         for res in where_result:\n     +             if \"code\" in res and res[\"code\"] == \"error\":\n     +                 # assume error is \"... file(s) not in client view\"\n 10:  04a0aedbaa ! 13:  e7bb92bcd6 git-p4: Support python3 for basic P4 clone, sync, and submit\n     @@ -1,54 +1,130 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: Support python3 for basic P4 clone, sync, and submit\n     +    git-p4: support Python 3 for basic P4 clone, sync, and submit (t9800)\n      \n     -    Issue: Python 3 is still not properly supported for any use with the git-p4 python code.\n     -    Warning - this is a very large atomic commit.  The commit text is also very large.\n     +    NOTE: Python 3 is still not properly supported for any use with the\n     +    git-p4 python code.\n      \n     -    Change the code such that, with the exception of P4 depot paths and depot files, all text read by git-p4 is cast as a string as soon as possible and converted back to bytes as late as possible, following Python2 to Python3 conversion best practices.\n     +    Warning - this is a very large atomic commit.  The commit text is also\n     +    very large.\n      \n     -    Important: Do not cast the bytes that contain the p4 depot path or p4 depot file name.  These should be left as bytes until used.\n     +    Change the code such that, with the exception of P4 depot paths and\n     +    depot files, all text read by git-p4 is cast as a string as soon as\n     +    possible and converted back to bytes as late as possible, following\n     +    Python 2 to Python 3 conversion best practices.\n      \n     -    These two values should not be converted because the encoding of these values is unknown.  git-p4 supports a configuration value git-p4.pathEncoding that is used by the encodeWithUTF8()  to determine what a UTF8 version of the path and filename should be.  However, since depot path and depot filename need to be sent to P4 in their original encoding, they will be left as byte streams until they are actually used:\n     +    Important: Do not cast the bytes that contain the p4 depot path or p4\n     +    depot file name.  These should be left as bytes until used.\n      \n     -    * When sent to P4, the bytes are literally passed to the p4 command\n     -    * When displayed in text for the user, they should be passed through the path_as_string() function\n     -    * When used by GIT they should be passed through the encodeWithUTF8() function\n     +    These two values should not be converted because the encoding of these\n     +    values is unknown.  git-p4 supports a configuration value\n     +    git-p4.pathEncoding that is used by the encodeWithUTF8() to determine\n     +    what a UTF8 version of the path and filename should be. However, since\n     +    depot path and depot filename need to be sent to P4 in their original\n     +    encoding, they will be left as byte streams until they are actually\n     +    used:\n      \n     -    Change all the rest of system calls to cast output (stdin) as_bytes() and input (stdout) as_string().  This retains existing Python 2 support, and adds python 3 support for these functions:\n     -    * read_pipe_full\n     -    * read_pipe_lines\n     -    * p4_has_move_command (used internally)\n     -    * gitConfig\n     -    * branch_exists\n     -    * GitLFS.generatePointer\n     -    * applyCommit - template must be read and written to the temporary file as_bytes() since it is created in memory as a string.\n     -    * streamOneP4File(file, contents) - wrap calls to the depotFile in path_as_string() for display. The file contents must be retained as bytes, so update the RCS changes to be forced to bytes.\n     -    * streamP4Files\n     -    * importHeadRevision(revision) - encode the depotPaths for display separate from the text for processing.\n     +      * When sent to P4, the bytes are literally passed to the p4 command\n     +      * When displayed in text for the user, they should be passed through\n     +        the path_as_string() function\n     +      * When used by GIT they should be passed through the encodeWithUTF8()\n     +        function\n     +\n     +    Change all the rest of system calls to cast output from system calls\n     +    (stdin) as_bytes() and input (stdout) as_string().  This retains\n     +    existing Python 2 support, and adds python 3 support for these\n     +    functions:\n     +\n     +     * read_pipe_full(c)\n     +     * read_pipe_lines(c)\n     +     * p4_has_move_command() - used internally\n     +     * gitConfig(key, typeSpecifier=None)\n     +     * branch_exists(branch)\n     +     * GitLFS.generatePointer(cloneDestination, contentFile)\n     +     * P4Submit.applyCommit(id) - template must be read and written to the\n     +           temporary file as_bytes() since it is created in memory as a\n     +           string.\n     +     * P4Sync.streamOneP4File(file, contents) - wrap calls to the depotFile\n     +           in path_as_string() for display. The file contents must be\n     +           retained as bytes, so update the RCS changes to be forced to\n     +           bytes.\n     +     * P4Sync.streamP4Files(marshalled)\n     +     * P4Sync.importHeadRevision(revision) - encode the depotPaths for\n     +           display separate from the text for processing.\n      \n          Py23File usage -\n     -    Change the P4Sync.OpenStreams() function to cast the gitOutput, gitStream, and gitError streams as Py23File() wrapper classes.  This facilitates taking strings in both python 2 and python 3 and casting them to bytes in the wrapper class instead of having to modify each method. Since the fast-import command also expects a raw byte stream for file content, add a new stream handle - gitStreamBytes which is an unwrapped verison of gitStream.\n     +\n     +    Change the P4Sync.OpenStreams() function to cast the gitOutput,\n     +    gitStream, and gitError streams as Py23File() wrapper classes.\n     +    This facilitates taking strings in both python 2 and python 3 and\n     +    casting them to bytes in the wrapper class instead of having to modify\n     +    each method. Since the fast-import command also expects a raw byte\n     +    stream for file content, add a new stream handle - gitStreamBytes which\n     +    is an unwrapped verison of gitStream.\n      \n          Literal text -\n     -    Depending on context, most literal text does not need casting to unicode or bytes as the text is Python dependent - In python 2, the string is implied as 'str' and python 3 the string is implied as 'unicode'. Under these conditions, they match the rest of the operating text, following best practices.  However, when a literal string is used in functions that are dealing with the raw input from and raw ouput to files streams, literal bytes may be required. Additionally, functions that are dealing with P4 depot paths or P4 depot file names are also dealing with bytes and will require the same casting as bytes.  The following functions cast text as byte strings:\n     -    * wildcard_decode(path) - the path parameter is a P4 depot and is bytes. Cast all the literals to bytes.\n     -    * wildcard_encode(path) - the path parameter is a P4 depot and is bytes. Cast all the literals to bytes.\n     -    * streamP4FilesCb(marshalled) - the marshalled data is in bytes. Cast the literals as bytes. When using this data to manipulate self.stream_file, encode all the marshalled data except for the 'depotFile' name.\n     -    * streamP4Files\n     +    Depending on context, most literal text does not need casting to unicode\n     +    or bytes as the text is Python dependent - In Python 2, the string is\n     +    implied as 'str' and python 3 the string is implied as 'unicode'. Under\n     +    these conditions, they match the rest of the operating text, following\n     +    best practices.  However, when a literal string is used in functions\n     +    that are dealing with the raw input from and raw ouput to files streams,\n     +    literal bytes may be required. Additionally, functions that are dealing\n     +    with P4 depot paths or P4 depot file names are also dealing with bytes\n     +    and will require the same casting as bytes.  The following functions\n     +    cast text as byte strings:\n     +\n     +     * wildcard_decode(path) - the path parameter is a P4 depot and is\n     +           bytes. Cast all the literals to bytes.\n     +     * wildcard_encode(path) - the path parameter is a P4 depot and is\n     +           bytes. Cast all the literals to bytes.\n     +     * P4Sync.streamP4FilesCb(marshalled) - the marshalled data is in bytes.\n     +           Cast the literals as bytes. When using this data to manipulate\n     +           self.stream_file, encode all the marshalled data except for the\n     +           'depotFile' name.\n     +     * P4Sync.streamP4Files(marshalled)\n      \n          Special behavior:\n     -    * p4_describe - encoding is disabled for the depotFile(x) and path elements since these are depot path and depo filenames.\n     -    * p4PathStartsWith(path, prefix) - Since P4 depot paths can contain non-UTF-8 encoded strings, change this method to compare paths while supporting the optional encoding.\n     -       - First, perform a byte-to-byte check to see if the path and prefix are both identical text.  There is no need to perform encoding conversions if the text is identical.\n     -       - If the byte check fails, pass both the path and prefix through encodeWithUTF8() to ensure both paths are using the same encoding. Then perform the test as originally written.\n     -    * patchRCSKeywords(file, pattern) - the parameters of file and pattern are both strings. However this function changes the contents of the file itentified by name \"file\". Treat the content of this file as binary to ensure that python does not accidently change the original encoding. The regular expression is cast as_bytes() and run against the file as_bytes(). The P4 keywords are ASCII strings and cannot span lines so iterating over each line of the file is acceptable.\n     -    * writeToGitStream(gitMode, relPath, contents) - Since 'contents' is already bytes data, instead of using the self.gitStream, use the new self.gitStreamBytes - the unwrapped gitStream that does not cast as_bytes() the binary data.\n     -    * commit(details, files, branch, parent = \"\", allow_empty=False) - Changed the encoding for the commit message to the preferred format for fast-import. The number of bytes is sent in the data block instead of using the EOT marker.\n     -    * Change the code for handling the user cache to use binary files. Cast text as_bytes() when writing to the cache and as_string() when reading from the cache.  This makes the reading and writing of the cache determinstic in it's encoding. Unlike file paths, P4 encodes the user names in UTF-8 encoding so no additional string encoding is required.\n     +\n     +     * p4_describep4_describe(change, shelved=False) - encoding is disabled\n     +           for the depotFile(x) and path elements since these are depot path\n     +           and depo filenames.\n     +     * p4PathStartsWith(path, prefix) - Since P4 depot paths can contain\n     +           non-UTF-8 encoded strings, change this method to compare paths\n     +           while supporting the optional encoding.\n     +\n     +            - First, perform a byte-to-byte check to see if the path and\n     +                  prefix are both identical text.  There is no need to\n     +                  perform encoding conversions if the text is identical.\n     +            - If the byte check fails, pass both the path and prefix through\n     +                  encodeWithUTF8() to ensure both paths are using the same\n     +                  encoding. Then perform the test as originally written.\n     +\n     +     * P4Submit.patchRCSKeywords(file, pattern) - the parameters of file and\n     +           pattern are both strings. However this function changes the\n     +           contents of the file itentified by name \"file\". Treat the content\n     +           of this file as binary to ensure that python does not accidently\n     +           change the original encoding. The regular expression is cast\n     +           as_bytes() and run against the file as_bytes(). The P4 keywords\n     +           are ASCII strings and cannot span lines so iterating over each\n     +           line of the file is acceptable.\n     +     * P4Sync.writeToGitStream(gitMode, relPath, contents) - Since\n     +           'contents' is already bytes data, instead of using the\n     +           self.gitStream, use the new self.gitStreamBytes - the unwrapped\n     +           gitStream that does not cast as_bytes() the binary data.\n     +     * P4Sync.commit(details, files, branch, parent = \"\", allow_empty=False)\n     +           Changed the encoding for the commit message to the preferred\n     +           format for fast-import. The number of bytes is sent in the data\n     +           block instead of using the EOT marker.\n     +\n     +     * Change the code for handling the user cache to use binary files.\n     +           Cast text as_bytes() when writing to the cache and as_string()\n     +           when reading from the cache.  This makes the reading and writing\n     +           of the cache determinstic in it's encoding. Unlike file paths,\n     +           P4 encodes the user names in UTF-8 encoding so no additional\n     +           string encoding is required.\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit 65ff0c74ebe62a200b4385ecfd4aa618ce091f48)\n      \n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n     @@ -122,7 +198,7 @@\n           cmd += [str(change)]\n       \n      -    ds = p4CmdList(cmd, skip_info=True)\n     -+    ds = p4CmdList(cmd, skip_info=True, encode_data=False)\n     ++    ds = p4CmdList(cmd, skip_info=True, encode_cmd_output=False)\n           if len(ds) != 1:\n               die(\"p4 describe -s %d did not return 1 result: %s\" % (change, str(ds)))\n       \n     @@ -137,29 +213,20 @@\n           if \"time\" not in d:\n               die(\"p4 describe -s %d returned no \\\"time\\\": %s\" % (change, str(d)))\n       \n     -+    # Do not convert 'depotFile(X)' or 'path' to be UTF-8 encoded, however \n     -+    # cast as_string() the rest of the text. \n     ++    # Do not convert 'depotFile(X)' or 'path' to be UTF-8 encoded, however\n     ++    # cast as_string() the rest of the text.\n      +    keys=d.keys()\n      +    for key in keys:\n      +        if key.startswith('depotFile'):\n     -+            d[key]=d[key] \n     ++            d[key]=d[key]\n      +        elif key == 'path':\n     -+            d[key]=d[key] \n     ++            d[key]=d[key]\n      +        else:\n      +            d[key] = as_string(d[key])\n      +\n           return d\n       \n       #\n     -@@\n     -     return result\n     - \n     - def p4Cmd(cmd):\n     -+    \"\"\" Executes a P4 command and returns the results in a dictionary\n     -+    \"\"\"\n     -     list = p4CmdList(cmd)\n     -     result = {}\n     -     for entry in list:\n      @@\n       _gitConfig = {}\n       \n     @@ -189,13 +256,13 @@\n           #\n           # we may or may not have a problem. If you have core.ignorecase=true,\n           # we treat DirA and dira as the same directory\n     -+    \n     ++\n      +    # Since we have to deal with mixed encodings for p4 file\n      +    # paths, first perform a simple startswith check, this covers\n      +    # the case that the formats and path are identical.\n      +    if as_bytes(path).startswith(as_bytes(prefix)):\n      +        return True\n     -+    \n     ++\n      +    # attempt to convert the prefix and path both to utf8\n      +    path_utf8 = encodeWithUTF8(path)\n      +    prefix_utf8 = encodeWithUTF8(prefix)\n     @@ -203,8 +270,8 @@\n           if gitConfigBool(\"core.ignorecase\"):\n      -        return path.lower().startswith(prefix.lower())\n      -    return path.startswith(prefix)\n     -+        # Check if we match byte-per-byte.  \n     -+        \n     ++        # Check if we match byte-per-byte.\n     ++\n      +        return path_utf8.lower().startswith(prefix_utf8.lower())\n      +    return path_utf8.startswith(prefix_utf8)\n       \n     @@ -272,7 +339,7 @@\n               self.userMapFromPerforceServer = True\n       \n           def loadUserMapFromCache(self):\n     -+        \"\"\" Reads the P4 username to git email map \n     ++        \"\"\" Reads the P4 username to git email map\n      +        \"\"\"\n               self.users = {}\n               self.userMapFromPerforceServer = False\n     @@ -292,7 +359,7 @@\n       \n           def patchRCSKeywords(self, file, pattern):\n      -        # Attempt to zap the RCS keywords in a p4 controlled file matching the given pattern\n     -+        \"\"\" Attempt to zap the RCS keywords in a p4 \n     ++        \"\"\" Attempt to zap the RCS keywords in a p4\n      +            controlled file matching the given pattern\n      +        \"\"\"\n      +        bSubLine = as_bytes(r'$\\1$')\n     @@ -377,7 +444,7 @@\n      +        \"\"\" output one file from the P4 stream to the git inbound stream.\n      +            helper for streamP4files.\n      +\n     -+            contents should be a bytes (bytes) \n     ++            contents should be a bytes (bytes)\n      +        \"\"\"\n               relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n               relPath = encodeWithUTF8(relPath, self.verbose)\n     @@ -427,7 +494,7 @@\n       \n      -    # handle another chunk of streaming data\n           def streamP4FilesCb(self, marshalled):\n     -+        \"\"\" Callback function for recording P4 chunks of data for streaming \n     ++        \"\"\" Callback function for recording P4 chunks of data for streaming\n      +            into GIT.\n      +\n      +            marshalled data is bytes[] from the caller\n     @@ -493,7 +560,7 @@\n       \n      -    # Stream directly from \"p4 files\" into \"git fast-import\"\n           def streamP4Files(self, files):\n     -+        \"\"\" Stream directly from \"p4 files\" into \"git fast-import\" \n     ++        \"\"\" Stream directly from \"p4 files\" into \"git fast-import\"\n      +        \"\"\"\n               filesForCommit = []\n               filesToRead = []\n     @@ -544,7 +611,7 @@\n      +\t    #('merge' SP <commit-ish> LF)*\n      +\t    #(filemodify | filedelete | filecopy | filerename | filedeleteall | notemodify)*\n      +\t    #LF?\n     -+        \n     ++\n      +        #'commit' - <ref> is the name of the branch to make the commit on\n               self.gitStream.write(\"commit %s\\n\" % branch)\n      +        #'mark' SP :<idnum>\n     @@ -558,9 +625,9 @@\n      -        self.gitStream.write(\"data <<EOT\\n\")\n      -        self.gitStream.write(details[\"desc\"])\n      +        # Per https://git-scm.com/docs/git-fast-import\n     -+        # The preferred method for creating the commit message is to supply the \n     -+        # byte count in the data method and not to use a Delimited format. \n     -+        # Collect all the text in the commit message into a single string and \n     ++        # The preferred method for creating the commit message is to supply the\n     ++        # byte count in the data method and not to use a Delimited format.\n     ++        # Collect all the text in the commit message into a single string and\n      +        # compute the byte count.\n      +        commitText = details[\"desc\"]\n               if len(jobs) > 0:\n     @@ -584,7 +651,7 @@\n      +            if len(details['options']) > 0:\n      +                commitText += (\": options = %s\" % details['options'])\n      +            commitText += \"]\"\n     -+        commitText += \"\\n\" \n     ++        commitText += \"\\n\"\n      +        self.gitStream.write(\"data %s\\n\" % len(as_bytes(commitText)))\n      +        self.gitStream.write(commitText)\n      +        self.gitStream.write(\"\\n\")\n     @@ -617,7 +684,7 @@\n               fileArgs = [\"%s...%s\" % (p,revision) for p in self.depotPaths]\n       \n      -        for info in p4CmdList([\"files\"] + fileArgs):\n     -+        for info in p4CmdList([\"files\"] + fileArgs, encode_data = False):\n     ++        for info in p4CmdList([\"files\"] + fileArgs, encode_cmd_output=False):\n       \n      -            if 'code' in info and info['code'] == 'error':\n      +            if 'code' in info and info['code'] == b'error':\n     @@ -640,7 +707,7 @@\n                       #fileCnt = fileCnt + 1\n                       continue\n       \n     -+            # Save all the file information, howerver do not translate the depotFile name at \n     ++            # Save all the file information, howerver do not translate the depotFile name at\n      +            # this time. Leave that as bytes since the encoding may vary.\n                   for prop in [\"depotFile\", \"rev\", \"action\", \"type\" ]:\n      -                details[\"%s%s\" % (prop, fileCnt)] = info[prop]\n 11:  883ef45ca5 ! 14:  25ad3e23a3 git-p4: Added --encoding parameter to p4 clone\n     @@ -1,19 +1,24 @@\n      Author: Ben Keene <seraphire@gmail.com>\n      \n     -    git-p4: Added --encoding parameter to p4 clone\n     +    git-p4: added --encoding parameter to p4 clone\n      \n     -    The test t9822 did not have any tests that had encoded a directory name in ISO8859-1.\n     +    The test t9822 did not have any tests that had encoded a directory name\n     +    in ISO8859-1.\n      \n     -    Additionally, to make it easier for the user to clone new repositories with a non-UTF-8 encoded path in P4, add a new parameter to p4clone \"--encoding\" that sets the\n     +    Additionally, to make it easier for the user to clone new repositories\n     +    with a non-UTF-8 encoded path in P4, add a new parameter to p4clone\n     +    \"--encoding\" that sets the\n      \n     -    Add new tests that use ISO8859-1 encoded text in both the directory and file names.  git-p4.pathEncoding.\n     +    Add new tests that use ISO8859-1 encoded text in both the directory and\n     +    file names.  git-p4.pathEncoding.\n      \n     -    Update the View class in the git-p4 code to properly cast text as_string() except for depot path and filenames.\n     +    Update the View class in the git-p4 code to properly cast text\n     +    as_string() except for depot path and filenames.\n      \n     -    Update the documentation to include the new command line parameter for p4clone\n     +    Update the documentation to include the new command line parameter for\n     +    p4clone\n      \n          Signed-off-by: Ben Keene <seraphire@gmail.com>\n     -    (cherry picked from commit e26f6309d60c6c1615320d4a9071935e23efe6fb)\n      \n       diff --git a/Documentation/git-p4.txt b/Documentation/git-p4.txt\n       --- a/Documentation/git-p4.txt\n     @@ -23,8 +28,8 @@\n       \tPerform a bare clone.  See linkgit:git-clone[1].\n       \n      +--encoding <encoding>::\n     -+    Optionally sets the git-p4.pathEncoding configuration value in \n     -+\tthe newly created Git repository before files are synchronized \n     ++    Optionally sets the git-p4.pathEncoding configuration value in\n     ++\tthe newly created Git repository before files are synchronized\n      +\tfrom P4. See git-p4.pathEncoding for more information.\n      +\n       Submit options\n     @@ -34,15 +39,6 @@\n       diff --git a/git-p4.py b/git-p4.py\n       --- a/git-p4.py\n       +++ b/git-p4.py\n     -@@\n     -     \"\"\"Look at the p4 client spec, create a View() object that contains\n     -        all the mappings, and return it.\"\"\"\n     - \n     --    specList = p4CmdList(\"client -o\")\n     -+    specList = p4CmdList(\"client -o\", encode_data=False)\n     -     if len(specList) != 1:\n     -         die('Output from \"client -o\" is %d lines, expecting 1' %\n     -             len(specList))\n      @@\n           entry = specList[0]\n       \n     @@ -130,11 +126,8 @@\n       \n           def update_client_spec_path_cache(self, files):\n      @@\n     -         if len(fileArgs) == 0:\n     -             return  # All files in cache\n       \n     --        where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs)\n     -+        where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs, encode_data=False)\n     +         where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs, encode_cmd_output=False)\n               for res in where_result:\n      -            if \"code\" in res and res[\"code\"] == \"error\":\n      +            if \"code\" in res and res[\"code\"] == b\"error\":\n     @@ -155,7 +148,7 @@\n      +        self.setPathEncoding = None\n       \n           def defaultDestination(self, args):\n     -         \"\"\" Returns the last path component as the default git \n     +         \"\"\" Returns the last path component as the default git\n      @@\n       \n               depotPaths = args\n     @@ -246,7 +239,7 @@\n      +\t\tDIR_ISO8859=\"$(printf \"$DIR_ISO8859_ESCAPED\")\" &&\n      +\t\tISO8859=\"$(printf \"$ISO8859_ESCAPED\")\" &&\n      +\t\tcd \"$cli\" &&\n     -+\t\tmkdir \"$DIR_ISO8859\" && \n     ++\t\tmkdir \"$DIR_ISO8859\" &&\n      +\t\tcd \"$DIR_ISO8859\" &&\n      +\t\techo content123 >\"$ISO8859\" &&\n      +\t\tp4 add \"$ISO8859\" &&\n  -:  ---------- > 15:  445dbc59f0 git-p4: Add depot manipulation functions\n\n-- \ngitgitgadget\n"},{"id":"387685","messageId":"d22ada16143ac67551c089ff44ab1832b791dd46.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 11/15] git-p4: add Py23File() - helper class for stream writing","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:39Z","receivedAt":"2019-12-07T17:47:58Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThis is a preparatory commit that does not change current behavior.\nIt adds a new class Py23File.\n\nFollowing the Python recommendation of keeping text as unicode\ninternally and only converting to and from bytes on input and output,\nthis class provides an interface for the methods used for reading and\nwriting files and file like streams.\n\nA new class was implemented to avoid requiring additional dependencies.\n\nCreate a class that wraps the input and output functions used by the\ngit-p4.py code for reading and writing to standard file handles.\n\nThe methods of this class should take a Unicode string for writing and\nreturn unicode strings in reads.  This class should be a drop-in for\nexisting file like streams\n\nThe following methods should be coded for supporting existing read/write\ncalls:\n  * write - this should write a Unicode string to the underlying stream\n  * read  - this should read from the underlying stream and cast the\n            bytes as a unicode string\n  * readline - this should read one line of text from the underlying\n            stream and cast it as a unicode string\n  * readline - this should read a number of lines, optionally hinted,\n            and cast each line as a unicode string\n\nThe expression \"cast as a unicode string\" is used because the code\nshould use the as_bytes() and as_string() functions instead of\ncohercing the data to actual unicode strings or bytes.  This allows\nPython 2 code to continue to use the internal \"str\" data type instead\nof converting the data back and forth to actual unicode strings. This\nretains current Python 2 support while Python 3 support may be\nincomplete.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 66 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 66 insertions(+)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 1838045078..03829f796d 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -4187,6 +4187,72 @@ def run(self, args):\n             print(\"%s <= %s (%s)\" % (branch, \",\".join(settings[\"depot-paths\"]), settings[\"change\"]))\n         return True\n \n+class Py23File():\n+    \"\"\" Python2/3 Unicode File Wrapper\n+    \"\"\"\n+\n+    stream_handle = None\n+    verbose       = False\n+    debug_handle  = None\n+\n+    def __init__(self, stream_handle, verbose = False):\n+        \"\"\" Create a Python3 compliant Unicode to Byte String\n+            Windows compatible wrapper\n+\n+            stream_handle = the underlying file-like handle\n+            verbose       = Boolean if content should be echoed\n+        \"\"\"\n+        self.stream_handle = stream_handle\n+        self.verbose       = verbose\n+\n+    def write(self, utf8string):\n+        \"\"\" Writes the utf8 encoded string to the underlying\n+            file stream\n+        \"\"\"\n+        self.stream_handle.write(as_bytes(utf8string))\n+        if self.verbose:\n+            sys.stderr.write(\"Stream Output: %s\" % utf8string)\n+            sys.stderr.flush()\n+\n+    def read(self, size = None):\n+        \"\"\" Reads int charcters from the underlying stream\n+            and converts it to utf8.\n+\n+            Be aware, the size value is for reading the underlying\n+            bytes so the value may be incorrect. Usage of the size\n+            value is discouraged.\n+        \"\"\"\n+        if size == None:\n+            return as_string(self.stream_handle.read())\n+        else:\n+            return as_string(self.stream_handle.read(size))\n+\n+    def readline(self):\n+        \"\"\" Reads a line from the underlying byte stream\n+            and converts it to utf8\n+        \"\"\"\n+        return as_string(self.stream_handle.readline())\n+\n+    def readlines(self, sizeHint = None):\n+        \"\"\" Returns a list containing lines from the file converted to unicode.\n+\n+            sizehint - Optional. If the optional sizehint argument is\n+            present, instead of reading up to EOF, whole lines totalling\n+            approximately sizehint bytes are read.\n+        \"\"\"\n+        lines = self.stream_handle.readlines(sizeHint)\n+        for i in range(0, len(lines)):\n+            lines[i] = as_string(lines[i])\n+        return lines\n+\n+    def close(self):\n+        \"\"\" Closes the underlying byte stream \"\"\"\n+        self.stream_handle.close()\n+\n+    def flush(self):\n+        \"\"\" Flushes the underlying byte stream \"\"\"\n+        self.stream_handle.flush()\n+\n class HelpFormatter(optparse.IndentedHelpFormatter):\n     def __init__(self):\n         optparse.IndentedHelpFormatter.__init__(self)\n-- \ngitgitgadget\n\n"},{"id":"387686","messageId":"1e677781d2cc75371b5362c7e63ea5ddf824d5da.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 08/15] git-p4: add casting helper functions for python 3 conversion","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:36Z","receivedAt":"2019-12-07T17:47:59Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nPython 3 handles strings differently than Python 2.7.  Since Python 2\nis reaching it's end of life, a series of changes are being submitted to\nenable python 3.5 and following support. The current code fails basic\ntests under python 3.5.\n\nChange the existing unicode test add new support functions for\nPython 2 - Python 3 support.\n\nDefine the following variables:\n- isunicode - a boolean variable that states if the version of python\n              natively supports unicode (true) or not (false). This is\n              true for Python 3 and false for Python 2.\n- unicode   - a type alias for the datatype that holds a unicode string.\n              It is assigned to a str under Python 3 and the unicode\n              type for Python 2.\n- bytes     - a type alias for an array of bytes.  It is assigned the\n              native bytes type for Python 3 and str for Python 2.\n\nAdd the following new functions:\n\n- as_string(text)  - A new function that will convert a byte array to a\n                     unicode (UTF-8) string under Python 3.  Under\n                     Python 2, this returns the string unchanged.\n- as_bytes(text)   - A new function that will convert a unicode string\n                     to a byte array under Python 3.  Under Python 2,\n                     this returns the string unchanged.\n- to_unicode(text) - Converts a text string as Unicode(UTF-8) on both\n                     Python 2 and Python 3.\n\nAdd a new function alias raw_input:\nIf raw_input does not exist (it was renamed to input in Python 3) alias\ninput as raw_input.\n\nThe as_string() and as_bytes() functions allow for modifying the code\nwith a minimal amount of impact on Python 2 support. When a string is\nexpected, the as_string() will be used to \"cast\" the incoming \"bytes\"\nto a string type.\n\nConversely as_bytes() will be used to cast a \"string\" to a \"byte array\"\ntype. Since Python 2 overloads the datatype 'str' to serve both purposes,\nthe Python 2 versions of these function do not change the data. This\nreduces the regression impact of these code changes.\n\n'basestring' is removed since its only references are found in tests\nthat were changed in modified in previous commits.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 80 ++++++++++++++++++++++++++++++++++++++++++++++++++-----\n 1 file changed, 74 insertions(+), 6 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex e020958083..e6f7513384 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -32,16 +32,84 @@\n     unicode = unicode\n except NameError:\n     # 'unicode' is undefined, must be Python 3\n-    str = str\n+    #\n+    # For Python 3 which is natively unicode, we will use\n+    # unicode for internal information but all P4 Data\n+    # will remain in bytes\n+    isunicode = True\n     unicode = str\n     bytes = bytes\n-    basestring = (str,bytes)\n+\n+    def as_string(text):\n+        \"\"\" Return a byte array as a unicode string\n+        \"\"\"\n+        if text is None:\n+            return None\n+        if isinstance(text, bytes):\n+            return unicode(text, \"utf-8\")\n+        else:\n+            return text\n+\n+    def as_bytes(text):\n+        \"\"\" Return a Unicode string as a byte array\n+        \"\"\"\n+        if text is None:\n+            return None\n+        if isinstance(text, bytes):\n+            return text\n+        else:\n+            return bytes(text, \"utf-8\")\n+\n+    def to_unicode(text):\n+        \"\"\" Return a byte array as a unicode string\n+        \"\"\"\n+        return as_string(text)\n+\n+    def path_as_string(path):\n+        \"\"\" Converts a path to the UTF8 encoded string\n+        \"\"\"\n+        if isinstance(path, unicode):\n+            return path\n+        return encodeWithUTF8(path).decode('utf-8')\n+\n else:\n     # 'unicode' exists, must be Python 2\n-    str = str\n+    #\n+    # We will treat the data as:\n+    #   str   -> str\n+    #   bytes -> str\n+    # So for Python 2 these functions are no-ops\n+    # and will leave the data in the ambiguious\n+    # string/bytes state\n+    isunicode = False\n     unicode = unicode\n     bytes = str\n-    basestring = basestring\n+\n+    def as_string(text):\n+        \"\"\" Return text unaltered (for Python 3 support)\n+        \"\"\"\n+        return text\n+\n+    def as_bytes(text):\n+        \"\"\" Return text unaltered (for Python 3 support)\n+        \"\"\"\n+        return text\n+\n+    def to_unicode(text):\n+        \"\"\" Return a string as a unicode string\n+        \"\"\"\n+        return text.decode('utf-8')\n+\n+    def path_as_string(path):\n+        \"\"\" Converts a path to the UTF8 encoded bytes\n+        \"\"\"\n+        return encodeWithUTF8(path)\n+\n+# Check for raw_input support\n+try:\n+    raw_input\n+except NameError:\n+    raw_input = input\n \n try:\n     from subprocess import CalledProcessError\n@@ -740,7 +808,7 @@ def p4Where(depotPath):\n             if data[:space] == depotPath:\n                 output = entry\n                 break\n-    if output == None:\n+    if output is None:\n         return \"\"\n     if output[\"code\"] == \"error\":\n         return \"\"\n@@ -4175,7 +4243,7 @@ def main():\n     global verbose\n     verbose = cmd.verbose\n     if cmd.needsGit:\n-        if cmd.gitdir == None:\n+        if cmd.gitdir is None:\n             cmd.gitdir = os.path.abspath(\".git\")\n             if not isValidGitDir(cmd.gitdir):\n                 # \"rev-parse --git-dir\" without arguments will try $PWD/.git\n-- \ngitgitgadget\n\n"},{"id":"387690","messageId":"a221eb8bb68975030966897a220ba2d328b000f5.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 09/15] git-p4: python 3 syntax changes","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:37Z","receivedAt":"2019-12-07T17:48:00Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nPython 3 handles strings differently than Python 2.7.  Since Python 2\nis reaching it's end of life, a series of changes are being submitted to\nenable python 3.5 and following support. The current code fails basic\ntests under python 3.5.\n\nThere are a number of translations suggested by modernize/futureize that\nshould be taken to fix numerous non-string specific issues.\n\nChange references to the X.next() iterator to the function next(X) which\nis compatible with both Python2 and Python3.\n\nChange references to X.keys() to list(X.keys()) to return a list that\ncan be iterated in both Python2 and Python3.\n\nAdd the literal text (object) to the end of class definitions to be\nconsistent with Python3 class definition.\n\nChange integer divison to use \"//\" instead of \"/\"  Under Both Python 2\nand Python 3 // will return a floor()ed result which matches existing\nfunctionality.\n\nChange the format string for displaying decimal values from %d to %4.1f%\nwhen displaying a progress.  This avoids displaying long repeating\ndecimals in user displayed text.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 55 +++++++++++++++++++++++++++++--------------------------\n 1 file changed, 29 insertions(+), 26 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex e6f7513384..fc6c9406c2 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -26,6 +26,9 @@\n import zlib\n import ctypes\n import errno\n+import os.path\n+import codecs\n+import io\n \n # support basestring in Python 3\n try:\n@@ -639,7 +642,7 @@ def parseDiffTreeEntry(entry):\n \n     If the pattern is not matched, None is returned.\"\"\"\n \n-    match = diffTreePattern().next().match(entry)\n+    match = next(diffTreePattern()).match(entry)\n     if match:\n         return {\n             'src_mode': match.group(1),\n@@ -980,7 +983,7 @@ def findUpstreamBranchPoint(head = \"HEAD\"):\n     branches = p4BranchesInGit()\n     # map from depot-path to branch name\n     branchByDepotPath = {}\n-    for branch in branches.keys():\n+    for branch in list(branches.keys()):\n         tip = branches[branch]\n         log = extractLogMessageFromGitCommit(tip)\n         settings = extractSettingsGitLog(log)\n@@ -1174,7 +1177,7 @@ def getClientSpec():\n     client_name = entry[\"Client\"]\n \n     # just the keys that start with \"View\"\n-    view_keys = [ k for k in entry.keys() if k.startswith(\"View\") ]\n+    view_keys = [ k for k in list(entry.keys()) if k.startswith(\"View\") ]\n \n     # hold this new View\n     view = View(client_name)\n@@ -1416,7 +1419,7 @@ def processContent(self, git_mode, relPath, contents):\n         else:\n             return LargeFileSystem.processContent(self, git_mode, relPath, contents)\n \n-class Command:\n+class Command(object):\n     delete_actions = ( \"delete\", \"move/delete\", \"purge\" )\n     add_actions = ( \"add\", \"branch\", \"move/add\" )\n \n@@ -1431,7 +1434,7 @@ def ensure_value(self, attr, value):\n             setattr(self, attr, value)\n         return getattr(self, attr)\n \n-class P4UserMap:\n+class P4UserMap(object):\n     def __init__(self):\n         self.userMapFromPerforceServer = False\n         self.myP4UserId = None\n@@ -1482,7 +1485,7 @@ def getUserMapFromPerforceServer(self):\n                 self.emails[email] = user\n \n         s = ''\n-        for (key, val) in self.users.items():\n+        for (key, val) in list(self.users.items()):\n             s += \"%s\\t%s\\n\" % (key.expandtabs(1), val.expandtabs(1))\n \n         open(self.getUserCacheFilename(), \"wb\").write(s)\n@@ -1833,7 +1836,7 @@ def prepareSubmitTemplate(self, changelist=None):\n                 break\n         if not change_entry:\n             die('Failed to decode output of p4 change -o')\n-        for key, value in change_entry.iteritems():\n+        for key, value in list(change_entry.items()):\n             if key.startswith('File'):\n                 if 'depot-paths' in settings:\n                     if not [p for p in settings['depot-paths']\n@@ -2077,7 +2080,7 @@ def applyCommit(self, id):\n             p4_delete(f)\n \n         # Set/clear executable bits\n-        for f in filesToChangeExecBit.keys():\n+        for f in list(filesToChangeExecBit.keys()):\n             mode = filesToChangeExecBit[f]\n             setP4ExecBit(f, mode)\n \n@@ -2330,7 +2333,7 @@ def run(self, args):\n             self.clientSpecDirs = getClientSpec()\n \n         # Check for the existence of P4 branches\n-        branchesDetected = (len(p4BranchesInGit().keys()) > 1)\n+        branchesDetected = (len(list(p4BranchesInGit().keys())) > 1)\n \n         if self.useClientSpec and not branchesDetected:\n             # all files are relative to the client spec\n@@ -2721,7 +2724,7 @@ def __init__(self):\n         self.knownBranches = {}\n         self.initialParents = {}\n \n-        self.tz = \"%+03d%02d\" % (- time.timezone / 3600, ((- time.timezone % 3600) / 60))\n+        self.tz = \"%+03d%02d\" % (- time.timezone // 3600, ((- time.timezone % 3600) // 60))\n         self.labels = {}\n \n     # Force a checkpoint in fast-import and wait for it to finish\n@@ -2838,7 +2841,7 @@ def splitFilesIntoBranches(self, commit):\n             else:\n                 relPath = self.stripRepoPath(path, self.depotPaths)\n \n-            for branch in self.knownBranches.keys():\n+            for branch in list(self.knownBranches.keys()):\n                 # add a trailing slash so that a commit into qt/4.2foo\n                 # doesn't end up in qt/4.2, e.g.\n                 if p4PathStartsWith(relPath, branch + \"/\"):\n@@ -2867,7 +2870,7 @@ def streamOneP4File(self, file, contents):\n                 size = int(self.stream_file['fileSize'])\n             else:\n                 size = 0 # deleted files don't get a fileSize apparently\n-            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size/1024/1024))\n+            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size//1024//1024))\n             sys.stdout.flush()\n \n         (type_base, type_mods) = split_p4_type(file[\"type\"])\n@@ -2967,7 +2970,7 @@ def streamP4FilesCb(self, marshalled):\n             required_bytes = int((4 * int(self.stream_file[\"fileSize\"])) - calcDiskFree())\n             if required_bytes > 0:\n                 err = 'Not enough space left on %s! Free at least %i MB.' % (\n-                    os.getcwd(), required_bytes/1024/1024\n+                    os.getcwd(), required_bytes//1024//1024\n                 )\n \n         if err:\n@@ -2996,7 +2999,7 @@ def streamP4FilesCb(self, marshalled):\n \n         # pick up the new file information... for the\n         # 'data' field we need to append to our array\n-        for k in marshalled.keys():\n+        for k in list(marshalled.keys()):\n             if k == 'data':\n                 if 'streamContentSize' not in self.stream_file:\n                     self.stream_file['streamContentSize'] = 0\n@@ -3011,8 +3014,8 @@ def streamP4FilesCb(self, marshalled):\n             'depotFile' in self.stream_file):\n             size = int(self.stream_file[\"fileSize\"])\n             if size > 0:\n-                progress = 100*self.stream_file['streamContentSize']/size\n-                sys.stdout.write('\\r%s %d%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size/1024/1024)))\n+                progress = 100.0*self.stream_file['streamContentSize']/size\n+                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size//1024//1024)))\n                 sys.stdout.flush()\n \n         self.stream_have_file_info = True\n@@ -3093,7 +3096,7 @@ def streamTag(self, gitStream, labelName, labelDetails, commit, epoch):\n \n         gitStream.write(\"tagger %s\\n\" % tagger)\n \n-        print(\"labelDetails=\",labelDetails)\n+        print((\"labelDetails=\",labelDetails))\n         if 'Description' in labelDetails:\n             description = labelDetails['Description']\n         else:\n@@ -3232,7 +3235,7 @@ def getLabels(self):\n             self.labels[newestChange] = [output, revisions]\n \n         if self.verbose:\n-            print(\"Label changes: %s\" % self.labels.keys())\n+            print(\"Label changes: %s\" % list(self.labels.keys()))\n \n     # Import p4 labels as git tags. A direct mapping does not\n     # exist, so assume that if all the files are at the same revision\n@@ -3375,7 +3378,7 @@ def getBranchMapping(self):\n \n     def getBranchMappingFromGitBranches(self):\n         branches = p4BranchesInGit(self.importIntoRemotes)\n-        for branch in branches.keys():\n+        for branch in list(branches.keys()):\n             if branch == \"master\":\n                 branch = \"main\"\n             else:\n@@ -3487,14 +3490,14 @@ def importChanges(self, changes, origin_revision=0):\n             self.updateOptionDict(description)\n \n             if not self.silent:\n-                sys.stdout.write(\"\\rImporting revision %s (%s%%)\" % (change, cnt * 100 / len(changes)))\n+                sys.stdout.write(\"\\rImporting revision %s (%4.1f%%)\" % (change, cnt * 100 / len(changes)))\n                 sys.stdout.flush()\n             cnt = cnt + 1\n \n             try:\n                 if self.detectBranches:\n                     branches = self.splitFilesIntoBranches(description)\n-                    for branch in branches.keys():\n+                    for branch in list(branches.keys()):\n                         ## HACK  --hwn\n                         branchPrefix = self.depotPaths[0] + branch + \"/\"\n                         self.branchPrefixes = [ branchPrefix ]\n@@ -3683,13 +3686,13 @@ def run(self, args):\n                 if short in branches:\n                     self.p4BranchesInGit = [ short ]\n             else:\n-                self.p4BranchesInGit = branches.keys()\n+                self.p4BranchesInGit = list(branches.keys())\n \n             if len(self.p4BranchesInGit) > 1:\n                 if not self.silent:\n                     print(\"Importing from/into multiple branches\")\n                 self.detectBranches = True\n-                for branch in branches.keys():\n+                for branch in list(branches.keys()):\n                     self.initialParents[self.refPrefix + branch] = \\\n                         branches[branch]\n \n@@ -4073,7 +4076,7 @@ def findLastP4Revision(self, starting_point):\n             to find the P4 commit we are based on, and the depot-paths.\n         \"\"\"\n \n-        for parent in (range(65535)):\n+        for parent in (list(range(65535))):\n             log = extractLogMessageFromGitCommit(\"{0}^{1}\".format(starting_point, parent))\n             settings = extractSettingsGitLog(log)\n             if 'change' in settings:\n@@ -4212,7 +4215,7 @@ def printUsage(commands):\n \n def main():\n     if len(sys.argv[1:]) == 0:\n-        printUsage(commands.keys())\n+        printUsage(list(commands.keys()))\n         sys.exit(2)\n \n     cmdName = sys.argv[1]\n@@ -4222,7 +4225,7 @@ def main():\n     except KeyError:\n         print(\"unknown command %s\" % cmdName)\n         print(\"\")\n-        printUsage(commands.keys())\n+        printUsage(list(commands.keys()))\n         sys.exit(2)\n \n     options = cmd.options\n-- \ngitgitgadget\n\n"},{"id":"387688","messageId":"b962cce8cd455d4cc561b14b4f498d5553067985.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 10/15] git-p4: fix assumed path separators to be more Windows friendly","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:38Z","receivedAt":"2019-12-07T17:48:03Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nWhen a computer is configured to use Git for windows and Python for\nwindows, and not a Unix subsystem like cygwin or WSL, the directory\nseparator changes and causes git-p4 to fail to properly determine paths.\n\nFix 3 path separator errors:\n\n1. getUserCacheFilename() - should not use string concatenation. Change\n   this code to use os.path.join to build an OS tolerant path.\n\n2. defaultDestiantion used the OS.path.split to split depot paths.  This\n   is incorrect on windows. Change the code to split on a forward\n   slash(/) instead since depot paths use this character regardless  of\n   the operating system.\n\n3. The call to isValidGitDir() in the main code also used a literal\n   forward slash. Change the code to use os.path.join to correctly\n   format the path for the operating system.\n\nThese three changes allow the suggested windows configuration to\nproperly locate files while retaining the existing behavior on\nnon-windows operating systems.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 13 +++++++++----\n 1 file changed, 9 insertions(+), 4 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex fc6c9406c2..1838045078 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -1459,8 +1459,10 @@ def p4UserIsMe(self, p4User):\n             return True\n \n     def getUserCacheFilename(self):\n+        \"\"\" Returns the filename of the username cache\n+        \"\"\"\n         home = os.environ.get(\"HOME\", os.environ.get(\"USERPROFILE\"))\n-        return home + \"/.gitp4-usercache.txt\"\n+        return os.path.join(home, \".gitp4-usercache.txt\")\n \n     def getUserMapFromPerforceServer(self):\n         if self.userMapFromPerforceServer:\n@@ -3978,13 +3980,16 @@ def __init__(self):\n         self.cloneBare = False\n \n     def defaultDestination(self, args):\n+        \"\"\" Returns the last path component as the default git\n+            repository directory name\n+        \"\"\"\n         ## TODO: use common prefix of args?\n         depotPath = args[0]\n         depotDir = re.sub(\"(@[^@]*)$\", \"\", depotPath)\n         depotDir = re.sub(\"(#[^#]*)$\", \"\", depotDir)\n         depotDir = re.sub(r\"\\.\\.\\.$\", \"\", depotDir)\n         depotDir = re.sub(r\"/$\", \"\", depotDir)\n-        return os.path.split(depotDir)[1]\n+        return depotDir.split('/')[-1]\n \n     def run(self, args):\n         if len(args) < 1:\n@@ -4257,8 +4262,8 @@ def main():\n                         chdir(cdup);\n \n         if not isValidGitDir(cmd.gitdir):\n-            if isValidGitDir(cmd.gitdir + \"/.git\"):\n-                cmd.gitdir += \"/.git\"\n+            if isValidGitDir(os.path.join(cmd.gitdir, \".git\")):\n+                cmd.gitdir = os.path.join(cmd.gitdir, \".git\")\n             else:\n                 die(\"fatal: cannot locate git repository at %s\" % cmd.gitdir)\n \n-- \ngitgitgadget\n\n"},{"id":"387694","messageId":"445dbc59f0cb82fabccc380c0346d65b778d8d1e.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 15/15] git-p4: Add depot manipulation functions","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:43Z","receivedAt":"2019-12-07T17:48:03Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nSince the Depot paths and filenames are encoded according to P4, we need\nto track them in bytes but also have to decode them with different\nencodings (either ASCII or the encoding configured in pathEncoding,\nwhich defaults to UTF-8)\n\nAdd the following functions to support future code conversion actions.\n\n * depot_count_depth         - counts the number of directories in the\n       path\n * depot_remove_leading_path - removes (n) directories from the front\n       of the depot path.\n * depot_Remove_p4_wildcard  - removes \"/...\" from the end of the path\n * depot_encode_utf8         - converts the path from the native\n       encoding to utf8 encoding.  Returns (depot_path, did_decode)\n * depot_encode_restore      - restores the original encoding of the\n       path.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n\n---\nThis code block could use review for the depot_encode_* functions.\n\nShould this code return an absolute Unicode string or a byte array.\n---\n git-p4.py | 93 +++++++++++++++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 93 insertions(+)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 16f29aae41..f82f05632c 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -724,6 +724,99 @@ def encodeWithUTF8(path, verbose=False):\n                 print('Path with non-ASCII characters detected. Used %s to encode: %s ' % (encoding, path))\n     return path\n \n+\n+def depot_count_depth(depot_path):\n+    \"\"\"Counts the number of directories found\n+    in the depot_path. Paths will be decoded \n+    with encodeWithUTF8 to ensure that depot\n+    encoding is repected.\n+\n+    Example:\n+        //depot         = 1\n+        //depot/        = 1\n+        //depot/dir     = 2\n+    \"\"\"\n+    depot_path=encodeWithUTF8(depot_path)\n+    if not depot_path.endswith(b\"/\"):\n+        depot_path+=b\"/\"\n+    return depot_path.count(b\"/\") - 2\n+\n+def depot_remove_leading_path(depot_path, depth):\n+    \"\"\"Remove depth number of directories from \n+    the beginning of the depot_path. This will\n+    be returned in the original encoding.\n+    The leading \"//\" does not count as a directory\n+    and will be automatically stripped.\n+\n+    depot_path should be in bytes\n+\n+    Example:\n+    Given a depot_path of: //depot/main/file.txt\n+    depth: 0        - depot/main/file.txt\n+    depth: 1        - main/file.txt\n+    depth: 2        - file.txt\n+    depth: 3        - (empty string)\n+    \"\"\"\n+\n+    # First, decode the path\n+    [depot_path, did_decode] = depot_encode_utf8(depot_path)\n+\n+    #remove leading //\n+    if depot_path.startswith(b\"//\"):\n+        depot_path=depot_path[2:]\n+    if depth != 0:\n+        segments=depot_path.split(b\"/\")\n+        segments=segments[depth:]\n+        depot_path=b\"/\".join(segments)\n+\n+    if did_decode:\n+        depot_path = depot_encode_restore(depot_path)\n+\n+    return depot_path\n+\n+def depot_remove_p4_wildcard(depot_path):\n+    \"\"\"Removes the \"/...\" from the end of depot\n+    path.\n+\n+    depot_path must be bytes. Bytes are returned.\n+    \"\"\"\n+    # First, decode the path\n+    [path, did_decode] = depot_encode_utf8(depot_path)\n+    \n+    if not path.endswith(b\"/...\"):\n+        return depot_path\n+    path=path[:-4]\n+\n+    if did_decode:\n+        path = depot_encode_restore(path)\n+\n+    return path\n+\n+def depot_encode_utf8(depot_path):\n+    \"\"\"conditionally encodes depot_path\n+    in utf8 using the defined pathEncoding.\n+\n+    Returns a (depot_path, was_encoded)\"\"\"\n+    did_decode=False\n+    encoding = 'utf8'\n+    try:\n+        depot_path.decode('ascii', 'strict')\n+    except:\n+        if gitConfig('git-p4.pathEncoding'):\n+            encoding = gitConfig('git-p4.pathEncoding')\n+        depot_path = depot_path.decode(encoding, 'replace').encode('utf8', 'replace')\n+        did_decode=True\n+    return [depot_path, did_decode]\n+\n+def depot_encode_restore(encoded_depot_path):\n+    \"\"\"Recodes an encoded_depot_path \n+    from utf8 back to the configured \n+    pathEncoding\"\"\"\n+    encoding = 'utf8'\n+    if gitConfig('git-p4.pathEncoding'):\n+        encoding = gitConfig('git-p4.pathEncoding')\n+    return encoded_depot_path.decode('utf8', 'replace').encode(encoding, 'replace')\n+\n class P4Exception(Exception):\n     \"\"\" Base class for exceptions from the p4 client \"\"\"\n     def __init__(self, exit_code):\n-- \ngitgitgadget\n"},{"id":"387691","messageId":"25ad3e23a337b53ef6ca52019899838cc7ec43f7.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 14/15] git-p4: added --encoding parameter to p4 clone","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:42Z","receivedAt":"2019-12-07T17:48:04Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThe test t9822 did not have any tests that had encoded a directory name\nin ISO8859-1.\n\nAdditionally, to make it easier for the user to clone new repositories\nwith a non-UTF-8 encoded path in P4, add a new parameter to p4clone\n\"--encoding\" that sets the\n\nAdd new tests that use ISO8859-1 encoded text in both the directory and\nfile names.  git-p4.pathEncoding.\n\nUpdate the View class in the git-p4 code to properly cast text\nas_string() except for depot path and filenames.\n\nUpdate the documentation to include the new command line parameter for\np4clone\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n Documentation/git-p4.txt        |   5 ++\n git-p4.py                       |  57 +++++++++++++-----\n t/t9822-git-p4-path-encoding.sh | 101 ++++++++++++++++++++++++++++++++\n 3 files changed, 147 insertions(+), 16 deletions(-)\n\ndiff --git a/Documentation/git-p4.txt b/Documentation/git-p4.txt\nindex 3494a1db3e..8fb844fc49 100644\n--- a/Documentation/git-p4.txt\n+++ b/Documentation/git-p4.txt\n@@ -305,6 +305,11 @@ options described above.\n --bare::\n \tPerform a bare clone.  See linkgit:git-clone[1].\n \n+--encoding <encoding>::\n+    Optionally sets the git-p4.pathEncoding configuration value in\n+\tthe newly created Git repository before files are synchronized\n+\tfrom P4. See git-p4.pathEncoding for more information.\n+\n Submit options\n ~~~~~~~~~~~~~~\n These options can be used to modify 'git p4 submit' behavior.\ndiff --git a/git-p4.py b/git-p4.py\nindex 9cf4e94e28..16f29aae41 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -1241,7 +1241,7 @@ def getClientSpec():\n     entry = specList[0]\n \n     # the //client/ name\n-    client_name = entry[\"Client\"]\n+    client_name = as_string(entry[\"Client\"])\n \n     # just the keys that start with \"View\"\n     view_keys = [ k for k in list(entry.keys()) if k.startswith(\"View\") ]\n@@ -2625,19 +2625,25 @@ def run(self, args):\n         return True\n \n class View(object):\n-    \"\"\"Represent a p4 view (\"p4 help views\"), and map files in a\n-       repo according to the view.\"\"\"\n+    \"\"\" Represent a p4 view (\"p4 help views\"), and map files in a\n+        repo according to the view.\n+    \"\"\"\n \n     def __init__(self, client_name):\n         self.mappings = []\n-        self.client_prefix = \"//%s/\" % client_name\n+        # the client prefix is saved in bytes as it is used for comparison\n+        # against server data.\n+        self.client_prefix = as_bytes(\"//%s/\" % client_name)\n         # cache results of \"p4 where\" to lookup client file locations\n         self.client_spec_path_cache = {}\n \n     def append(self, view_line):\n-        \"\"\"Parse a view line, splitting it into depot and client\n-           sides.  Append to self.mappings, preserving order.  This\n-           is only needed for tag creation.\"\"\"\n+        \"\"\" Parse a view line, splitting it into depot and client\n+            sides.  Append to self.mappings, preserving order.  This\n+            is only needed for tag creation.\n+\n+            view_line should be in bytes (depot path encoding)\n+        \"\"\"\n \n         # Split the view line into exactly two words.  P4 enforces\n         # structure on these lines that simplifies this quite a bit.\n@@ -2650,28 +2656,28 @@ def append(self, view_line):\n         # The line is already white-space stripped.\n         # The two words are separated by a single space.\n         #\n-        if view_line[0] == '\"':\n+        if view_line[0] == b'\"':\n             # First word is double quoted.  Find its end.\n-            close_quote_index = view_line.find('\"', 1)\n+            close_quote_index = view_line.find(b'\"', 1)\n             if close_quote_index <= 0:\n-                die(\"No first-word closing quote found: %s\" % view_line)\n+                die(\"No first-word closing quote found: %s\" % path_as_string(view_line))\n             depot_side = view_line[1:close_quote_index]\n             # skip closing quote and space\n             rhs_index = close_quote_index + 1 + 1\n         else:\n-            space_index = view_line.find(\" \")\n+            space_index = view_line.find(b\" \")\n             if space_index <= 0:\n-                die(\"No word-splitting space found: %s\" % view_line)\n+                die(\"No word-splitting space found: %s\" % path_as_string(view_line))\n             depot_side = view_line[0:space_index]\n             rhs_index = space_index + 1\n \n         # prefix + means overlay on previous mapping\n-        if depot_side.startswith(\"+\"):\n+        if depot_side.startswith(b\"+\"):\n             depot_side = depot_side[1:]\n \n         # prefix - means exclude this path, leave out of mappings\n         exclude = False\n-        if depot_side.startswith(\"-\"):\n+        if depot_side.startswith(b\"-\"):\n             exclude = True\n             depot_side = depot_side[1:]\n \n@@ -2682,7 +2688,7 @@ def convert_client_path(self, clientFile):\n         # chop off //client/ part to make it relative\n         if not clientFile.startswith(self.client_prefix):\n             die(\"No prefix '%s' on clientFile '%s'\" %\n-                (self.client_prefix, clientFile))\n+                (as_string(self.client_prefix)), path_as_string(clientFile))\n         return clientFile[len(self.client_prefix):]\n \n     def update_client_spec_path_cache(self, files):\n@@ -2696,7 +2702,7 @@ def update_client_spec_path_cache(self, files):\n \n         where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs, encode_cmd_output=False)\n         for res in where_result:\n-            if \"code\" in res and res[\"code\"] == \"error\":\n+            if \"code\" in res and res[\"code\"] == b\"error\":\n                 # assume error is \"... file(s) not in client view\"\n                 continue\n             if \"clientFile\" not in res:\n@@ -4113,10 +4119,14 @@ def __init__(self):\n                                  help=\"where to leave result of the clone\"),\n             optparse.make_option(\"--bare\", dest=\"cloneBare\",\n                                  action=\"store_true\", default=False),\n+            optparse.make_option(\"--encoding\", dest=\"setPathEncoding\",\n+                                 action=\"store\", default=None,\n+                                 help=\"Sets the path encoding for this depot\")\n         ]\n         self.cloneDestination = None\n         self.needsGit = False\n         self.cloneBare = False\n+        self.setPathEncoding = None\n \n     def defaultDestination(self, args):\n         \"\"\" Returns the last path component as the default git\n@@ -4140,6 +4150,14 @@ def run(self, args):\n \n         depotPaths = args\n \n+        # If we have an encoding provided, ignore what may already exist\n+        # in the registry. This will ensure we show the displayed values\n+        # using the correct encoding.\n+        if self.setPathEncoding:\n+            gitConfigSet(\"git-p4.pathEncoding\", self.setPathEncoding)\n+\n+        # If more than 1 path element is supplied, the last element\n+        # is the clone destination.\n         if not self.cloneDestination and len(depotPaths) > 1:\n             self.cloneDestination = depotPaths[-1]\n             depotPaths = depotPaths[:-1]\n@@ -4167,6 +4185,13 @@ def run(self, args):\n         if retcode:\n             raise CalledProcessError(retcode, init_cmd)\n \n+        # Set the encoding if it was provided command line\n+        if self.setPathEncoding:\n+            init_cmd= [\"git\", \"config\", \"git-p4.pathEncoding\", self.setPathEncoding]\n+            retcode = subprocess.call(init_cmd)\n+            if retcode:\n+                raise CalledProcessError(retcode, init_cmd)\n+\n         if not P4Sync.run(self, depotPaths):\n             return False\n \ndiff --git a/t/t9822-git-p4-path-encoding.sh b/t/t9822-git-p4-path-encoding.sh\nindex 572d395498..8d3fe6c5d1 100755\n--- a/t/t9822-git-p4-path-encoding.sh\n+++ b/t/t9822-git-p4-path-encoding.sh\n@@ -4,9 +4,20 @@ test_description='Clone repositories with non ASCII paths'\n \n . ./lib-git-p4.sh\n \n+# lowercase filename\n+# UTF8    - HEX:   a-\\xc3\\xa4_o-\\xc3\\xb6_u-\\xc3\\xbc\n+#         - octal: a-\\303\\244_o-\\303\\266_u-\\303\\274\n+# ISO8859 - HEX:   a-\\xe4_o-\\xf6_u-\\xfc\n UTF8_ESCAPED=\"a-\\303\\244_o-\\303\\266_u-\\303\\274.txt\"\n ISO8859_ESCAPED=\"a-\\344_o-\\366_u-\\374.txt\"\n \n+# lowercase directory\n+# UTF8    - HEX:   dir_a-\\xc3\\xa4_o-\\xc3\\xb6_u-\\xc3\\xbc\n+# ISO8859 - HEX:   dir_a-\\xe4_o-\\xf6_u-\\xfc\n+DIR_UTF8_ESCAPED=\"dir_a-\\303\\244_o-\\303\\266_u-\\303\\274\"\n+DIR_ISO8859_ESCAPED=\"dir_a-\\344_o-\\366_u-\\374\"\n+\n+\n ISO8859=\"$(printf \"$ISO8859_ESCAPED\")\" &&\n echo content123 >\"$ISO8859\" &&\n rm \"$ISO8859\" || {\n@@ -58,6 +69,22 @@ test_expect_success 'Clone repo containing iso8859-1 encoded paths with git-p4.p\n \t)\n '\n \n+test_expect_success 'Clone repo containing iso8859-1 encoded paths with using --encoding parameter' '\n+\ttest_when_finished cleanup_git &&\n+\t(\n+\t\tgit p4 clone --encoding iso8859 --destination=\"$git\" //depot &&\n+\t\tcd \"$git\" &&\n+\t\tUTF8=\"$(printf \"$UTF8_ESCAPED\")\" &&\n+\t\techo \"$UTF8\" >expect &&\n+\t\tgit -c core.quotepath=false ls-files >actual &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\techo content123 >expect &&\n+\t\tcat \"$UTF8\" >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n test_expect_success 'Delete iso8859-1 encoded paths and clone' '\n \t(\n \t\tcd \"$cli\" &&\n@@ -74,4 +101,78 @@ test_expect_success 'Delete iso8859-1 encoded paths and clone' '\n \t)\n '\n \n+# These tests will create a directory with ISO8859-1 characters in both the \n+# directory and the path.  Since it is possible to clone a path instead of using\n+# the whole client-spec.  Check both versions:  client-spec and with a direct\n+# path using --encoding\n+test_expect_success 'Create a repo containing iso8859-1 encoded directory and filename' '\n+\t(\n+\t\tDIR_ISO8859=\"$(printf \"$DIR_ISO8859_ESCAPED\")\" &&\n+\t\tISO8859=\"$(printf \"$ISO8859_ESCAPED\")\" &&\n+\t\tcd \"$cli\" &&\n+\t\tmkdir \"$DIR_ISO8859\" &&\n+\t\tcd \"$DIR_ISO8859\" &&\n+\t\techo content123 >\"$ISO8859\" &&\n+\t\tp4 add \"$ISO8859\" &&\n+\t\tp4 submit -d \"test commit (encoded directory)\"\n+\t)\n+'\n+\n+test_expect_success 'Clone repo containing iso8859-1 encoded depot path and files with git-p4.pathEncoding' '\n+\ttest_when_finished cleanup_git &&\n+\t(\n+\t\tDIR_ISO8859=\"$(printf \"$DIR_ISO8859_ESCAPED\")\" &&\n+\t\tDIR_UTF8=\"$(printf \"$DIR_UTF8_ESCAPED\")\" &&\n+\t\tcd \"$git\" &&\n+\t\tgit init . &&\n+\t\tgit config git-p4.pathEncoding iso8859-1 &&\n+\t\tgit p4 clone --use-client-spec --destination=\"$git\" \"//depot/$DIR_ISO8859\" &&\n+\t\tcd \"$DIR_UTF8\" &&\n+\t\tUTF8=\"$(printf \"$UTF8_ESCAPED\")\" &&\n+\t\techo \"$UTF8\" >expect &&\n+\t\tgit -c core.quotepath=false ls-files >actual &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\techo content123 >expect &&\n+\t\tcat \"$UTF8\" >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'Clone repo containing iso8859-1 encoded depot path and files with git-p4.pathEncoding, without --use-client-spec' '\n+\ttest_when_finished cleanup_git &&\n+\t(\n+\t\tDIR_ISO8859=\"$(printf \"$DIR_ISO8859_ESCAPED\")\" &&\n+\t\tcd \"$git\" &&\n+\t\tgit init . &&\n+\t\tgit config git-p4.pathEncoding iso8859-1 &&\n+\t\tgit p4 clone --destination=\"$git\" \"//depot/$DIR_ISO8859\" &&\n+\t\tUTF8=\"$(printf \"$UTF8_ESCAPED\")\" &&\n+\t\techo \"$UTF8\" >expect &&\n+\t\tgit -c core.quotepath=false ls-files >actual &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\techo content123 >expect &&\n+\t\tcat \"$UTF8\" >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n+test_expect_success 'Clone repo containing iso8859-1 encoded depot path and files with using --encoding parameter' '\n+\ttest_when_finished cleanup_git &&\n+\t(\n+\t\tDIR_ISO8859=\"$(printf \"$DIR_ISO8859_ESCAPED\")\" &&\n+\t\tgit p4 clone --encoding iso8859 --destination=\"$git\" \"//depot/$DIR_ISO8859\" &&\n+\t\tcd \"$git\" &&\n+\t\tUTF8=\"$(printf \"$UTF8_ESCAPED\")\" &&\n+\t\techo \"$UTF8\" >expect &&\n+\t\tgit -c core.quotepath=false ls-files >actual &&\n+\t\ttest_cmp expect actual &&\n+\n+\t\techo content123 >expect &&\n+\t\tcat \"$UTF8\" >actual &&\n+\t\ttest_cmp expect actual\n+\t)\n+'\n+\n test_done\n-- \ngitgitgadget\n\n"},{"id":"387689","messageId":"bc7009541b3f03c3065a7be2b569cd4bf91f7c05.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 07/15] git-p4: add new support function gitConfigSet()","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:35Z","receivedAt":"2019-12-07T17:48:06Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nAdd a new method gitConfigSet(). This method will set a value in the git\nconfiguration cache list.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 5 +++++\n 1 file changed, 5 insertions(+)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex e7c24817ad..e020958083 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -860,6 +860,11 @@ def gitConfigList(key):\n             _gitConfig[key] = []\n     return _gitConfig[key]\n \n+def gitConfigSet(key, value):\n+    \"\"\" Set the git configuration key 'key' to 'value' for this session\n+    \"\"\"\n+    _gitConfig[key] = value\n+\n def p4BranchesInGit(branchesAreInRemotes=True):\n     \"\"\"Find all the branches whose names start with \"p4/\", looking\n        in remotes or heads as specified by the argument.  Return\n-- \ngitgitgadget\n\n"},{"id":"387692","messageId":"e7bb92bcd635813dda187017378a81fd90eceede.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 13/15] git-p4: support Python 3 for basic P4 clone, sync, and submit (t9800)","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:41Z","receivedAt":"2019-12-07T17:48:06Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nNOTE: Python 3 is still not properly supported for any use with the\ngit-p4 python code.\n\nWarning - this is a very large atomic commit.  The commit text is also\nvery large.\n\nChange the code such that, with the exception of P4 depot paths and\ndepot files, all text read by git-p4 is cast as a string as soon as\npossible and converted back to bytes as late as possible, following\nPython 2 to Python 3 conversion best practices.\n\nImportant: Do not cast the bytes that contain the p4 depot path or p4\ndepot file name.  These should be left as bytes until used.\n\nThese two values should not be converted because the encoding of these\nvalues is unknown.  git-p4 supports a configuration value\ngit-p4.pathEncoding that is used by the encodeWithUTF8() to determine\nwhat a UTF8 version of the path and filename should be. However, since\ndepot path and depot filename need to be sent to P4 in their original\nencoding, they will be left as byte streams until they are actually\nused:\n\n  * When sent to P4, the bytes are literally passed to the p4 command\n  * When displayed in text for the user, they should be passed through\n    the path_as_string() function\n  * When used by GIT they should be passed through the encodeWithUTF8()\n    function\n\nChange all the rest of system calls to cast output from system calls\n(stdin) as_bytes() and input (stdout) as_string().  This retains\nexisting Python 2 support, and adds python 3 support for these\nfunctions:\n\n * read_pipe_full(c)\n * read_pipe_lines(c)\n * p4_has_move_command() - used internally\n * gitConfig(key, typeSpecifier=None)\n * branch_exists(branch)\n * GitLFS.generatePointer(cloneDestination, contentFile)\n * P4Submit.applyCommit(id) - template must be read and written to the\n       temporary file as_bytes() since it is created in memory as a\n       string.\n * P4Sync.streamOneP4File(file, contents) - wrap calls to the depotFile\n       in path_as_string() for display. The file contents must be\n       retained as bytes, so update the RCS changes to be forced to\n       bytes.\n * P4Sync.streamP4Files(marshalled)\n * P4Sync.importHeadRevision(revision) - encode the depotPaths for\n       display separate from the text for processing.\n\nPy23File usage -\n\nChange the P4Sync.OpenStreams() function to cast the gitOutput,\ngitStream, and gitError streams as Py23File() wrapper classes.\nThis facilitates taking strings in both python 2 and python 3 and\ncasting them to bytes in the wrapper class instead of having to modify\neach method. Since the fast-import command also expects a raw byte\nstream for file content, add a new stream handle - gitStreamBytes which\nis an unwrapped verison of gitStream.\n\nLiteral text -\nDepending on context, most literal text does not need casting to unicode\nor bytes as the text is Python dependent - In Python 2, the string is\nimplied as 'str' and python 3 the string is implied as 'unicode'. Under\nthese conditions, they match the rest of the operating text, following\nbest practices.  However, when a literal string is used in functions\nthat are dealing with the raw input from and raw ouput to files streams,\nliteral bytes may be required. Additionally, functions that are dealing\nwith P4 depot paths or P4 depot file names are also dealing with bytes\nand will require the same casting as bytes.  The following functions\ncast text as byte strings:\n\n * wildcard_decode(path) - the path parameter is a P4 depot and is\n       bytes. Cast all the literals to bytes.\n * wildcard_encode(path) - the path parameter is a P4 depot and is\n       bytes. Cast all the literals to bytes.\n * P4Sync.streamP4FilesCb(marshalled) - the marshalled data is in bytes.\n       Cast the literals as bytes. When using this data to manipulate\n       self.stream_file, encode all the marshalled data except for the\n       'depotFile' name.\n * P4Sync.streamP4Files(marshalled)\n\nSpecial behavior:\n\n * p4_describep4_describe(change, shelved=False) - encoding is disabled\n       for the depotFile(x) and path elements since these are depot path\n       and depo filenames.\n * p4PathStartsWith(path, prefix) - Since P4 depot paths can contain\n       non-UTF-8 encoded strings, change this method to compare paths\n       while supporting the optional encoding.\n\n        - First, perform a byte-to-byte check to see if the path and\n              prefix are both identical text.  There is no need to\n              perform encoding conversions if the text is identical.\n        - If the byte check fails, pass both the path and prefix through\n              encodeWithUTF8() to ensure both paths are using the same\n              encoding. Then perform the test as originally written.\n\n * P4Submit.patchRCSKeywords(file, pattern) - the parameters of file and\n       pattern are both strings. However this function changes the\n       contents of the file itentified by name \"file\". Treat the content\n       of this file as binary to ensure that python does not accidently\n       change the original encoding. The regular expression is cast\n       as_bytes() and run against the file as_bytes(). The P4 keywords\n       are ASCII strings and cannot span lines so iterating over each\n       line of the file is acceptable.\n * P4Sync.writeToGitStream(gitMode, relPath, contents) - Since\n       'contents' is already bytes data, instead of using the\n       self.gitStream, use the new self.gitStreamBytes - the unwrapped\n       gitStream that does not cast as_bytes() the binary data.\n * P4Sync.commit(details, files, branch, parent = \"\", allow_empty=False)\n       Changed the encoding for the commit message to the preferred\n       format for fast-import. The number of bytes is sent in the data\n       block instead of using the EOT marker.\n\n * Change the code for handling the user cache to use binary files.\n       Cast text as_bytes() when writing to the cache and as_string()\n       when reading from the cache.  This makes the reading and writing\n       of the cache determinstic in it's encoding. Unlike file paths,\n       P4 encodes the user names in UTF-8 encoding so no additional\n       string encoding is required.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 285 ++++++++++++++++++++++++++++++++++++++----------------\n 1 file changed, 203 insertions(+), 82 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex e8f31339e4..9cf4e94e28 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -273,6 +273,8 @@ def read_pipe_full(c):\n     expand = not isinstance(c, list)\n     p = subprocess.Popen(c, stdout=subprocess.PIPE, stderr=subprocess.PIPE, shell=expand)\n     (out, err) = p.communicate()\n+    out = as_string(out)\n+    err = as_string(err)\n     return (p.returncode, out, err)\n \n def read_pipe(c, ignore_error=False):\n@@ -299,10 +301,17 @@ def read_pipe_text(c):\n         return out.rstrip()\n \n def p4_read_pipe(c, ignore_error=False):\n+    \"\"\" Read output from the P4 command 'c'. Returns the output text on\n+        success. On failure, terminates execution, unless\n+        ignore_error is True, when it returns an empty string.\n+    \"\"\"\n     real_cmd = p4_build_cmd(c)\n     return read_pipe(real_cmd, ignore_error)\n \n def read_pipe_lines(c):\n+    \"\"\" Returns a list of text from executing the command 'c'.\n+        The program will die if the command fails to execute.\n+    \"\"\"\n     if verbose:\n         sys.stderr.write('Reading pipe: %s\\n' % str(c))\n \n@@ -312,6 +321,11 @@ def read_pipe_lines(c):\n     val = pipe.readlines()\n     if pipe.close() or p.wait():\n         die('Command failed: %s' % str(c))\n+    # Unicode conversion from byte-string\n+    # Iterate and fix in-place to avoid a second list in memory.\n+    if isunicode:\n+        for i in range(len(val)):\n+            val[i] = as_string(val[i])\n \n     return val\n \n@@ -340,6 +354,8 @@ def p4_has_move_command():\n     cmd = p4_build_cmd([\"move\", \"-k\", \"@from\", \"@to\"])\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     (out, err) = p.communicate()\n+    out=as_string(out)\n+    err=as_string(err)\n     # return code will be 1 in either case\n     if err.find(\"Invalid option\") >= 0:\n         return False\n@@ -467,16 +483,20 @@ def p4_last_change():\n     return int(results[0]['change'])\n \n def p4_describe(change, shelved=False):\n-    \"\"\"Make sure it returns a valid result by checking for\n-       the presence of field \"time\".  Return a dict of the\n-       results.\"\"\"\n+    \"\"\" Returns information about the requested P4 change list.\n+\n+        Data returned is not string encoded (returned as bytes)\n+    \"\"\"\n+    # Make sure it returns a valid result by checking for\n+    #   the presence of field \"time\".  Return a dict of the\n+    #   results.\n \n     cmd = [\"describe\", \"-s\"]\n     if shelved:\n         cmd += [\"-S\"]\n     cmd += [str(change)]\n \n-    ds = p4CmdList(cmd, skip_info=True)\n+    ds = p4CmdList(cmd, skip_info=True, encode_cmd_output=False)\n     if len(ds) != 1:\n         die(\"p4 describe -s %d did not return 1 result: %s\" % (change, str(ds)))\n \n@@ -486,12 +506,23 @@ def p4_describe(change, shelved=False):\n         die(\"p4 describe -s %d exited with %d: %s\" % (change, d[\"p4ExitCode\"],\n                                                       str(d)))\n     if \"code\" in d:\n-        if d[\"code\"] == \"error\":\n+        if d[\"code\"] == b\"error\":\n             die(\"p4 describe -s %d returned error code: %s\" % (change, str(d)))\n \n     if \"time\" not in d:\n         die(\"p4 describe -s %d returned no \\\"time\\\": %s\" % (change, str(d)))\n \n+    # Do not convert 'depotFile(X)' or 'path' to be UTF-8 encoded, however\n+    # cast as_string() the rest of the text.\n+    keys=d.keys()\n+    for key in keys:\n+        if key.startswith('depotFile'):\n+            d[key]=d[key]\n+        elif key == 'path':\n+            d[key]=d[key]\n+        else:\n+            d[key] = as_string(d[key])\n+\n     return d\n \n #\n@@ -914,13 +945,15 @@ def gitDeleteRef(ref):\n _gitConfig = {}\n \n def gitConfig(key, typeSpecifier=None):\n+    \"\"\" Return a configuration setting from GIT\n+\t\"\"\"\n     if key not in _gitConfig:\n         cmd = [ \"git\", \"config\" ]\n         if typeSpecifier:\n             cmd += [ typeSpecifier ]\n         cmd += [ key ]\n         s = read_pipe(cmd, ignore_error=True)\n-        _gitConfig[key] = s.strip()\n+        _gitConfig[key] = as_string(s).strip()\n     return _gitConfig[key]\n \n def gitConfigBool(key):\n@@ -994,6 +1027,7 @@ def branch_exists(branch):\n     cmd = [ \"git\", \"rev-parse\", \"--symbolic\", \"--verify\", branch ]\n     p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)\n     out, _ = p.communicate()\n+    out = as_string(out)\n     if p.returncode:\n         return False\n     # expect exactly one line of output: the branch name\n@@ -1177,9 +1211,22 @@ def p4PathStartsWith(path, prefix):\n     #\n     # we may or may not have a problem. If you have core.ignorecase=true,\n     # we treat DirA and dira as the same directory\n+\n+    # Since we have to deal with mixed encodings for p4 file\n+    # paths, first perform a simple startswith check, this covers\n+    # the case that the formats and path are identical.\n+    if as_bytes(path).startswith(as_bytes(prefix)):\n+        return True\n+\n+    # attempt to convert the prefix and path both to utf8\n+    path_utf8 = encodeWithUTF8(path)\n+    prefix_utf8 = encodeWithUTF8(prefix)\n+\n     if gitConfigBool(\"core.ignorecase\"):\n-        return path.lower().startswith(prefix.lower())\n-    return path.startswith(prefix)\n+        # Check if we match byte-per-byte.\n+\n+        return path_utf8.lower().startswith(prefix_utf8.lower())\n+    return path_utf8.startswith(prefix_utf8)\n \n def getClientSpec():\n     \"\"\"Look at the p4 client spec, create a View() object that contains\n@@ -1235,18 +1282,24 @@ def wildcard_decode(path):\n     # Cannot have * in a filename in windows; untested as to\n     # what p4 would do in such a case.\n     if not platform.system() == \"Windows\":\n-        path = path.replace(\"%2A\", \"*\")\n-    path = path.replace(\"%23\", \"#\") \\\n-               .replace(\"%40\", \"@\") \\\n-               .replace(\"%25\", \"%\")\n+        path = path.replace(b\"%2A\", b\"*\")\n+    path = path.replace(b\"%23\", b\"#\") \\\n+               .replace(b\"%40\", b\"@\") \\\n+               .replace(b\"%25\", b\"%\")\n     return path\n \n def wildcard_encode(path):\n     # do % first to avoid double-encoding the %s introduced here\n-    path = path.replace(\"%\", \"%25\") \\\n-               .replace(\"*\", \"%2A\") \\\n-               .replace(\"#\", \"%23\") \\\n-               .replace(\"@\", \"%40\")\n+    if isinstance(path, unicode):\n+        path = path.replace(\"%\", \"%25\") \\\n+                   .replace(\"*\", \"%2A\") \\\n+                   .replace(\"#\", \"%23\") \\\n+                   .replace(\"@\", \"%40\")\n+    else:\n+        path = path.replace(b\"%\", b\"%25\") \\\n+                   .replace(b\"*\", b\"%2A\") \\\n+                   .replace(b\"#\", b\"%23\") \\\n+                   .replace(b\"@\", b\"%40\")\n     return path\n \n def wildcard_present(path):\n@@ -1378,7 +1431,7 @@ def generatePointer(self, contentFile):\n             ['git', 'lfs', 'pointer', '--file=' + contentFile],\n             stdout=subprocess.PIPE\n         )\n-        pointerFile = pointerProcess.stdout.read()\n+        pointerFile = as_string(pointerProcess.stdout.read())\n         if pointerProcess.wait():\n             os.remove(contentFile)\n             die('git-lfs pointer command failed. Did you install the extension?')\n@@ -1485,6 +1538,8 @@ def getUserCacheFilename(self):\n         return os.path.join(home, \".gitp4-usercache.txt\")\n \n     def getUserMapFromPerforceServer(self):\n+        \"\"\" Creates the usercache from the data in P4.\n+        \"\"\"\n         if self.userMapFromPerforceServer:\n             return\n         self.users = {}\n@@ -1510,18 +1565,22 @@ def getUserMapFromPerforceServer(self):\n         for (key, val) in list(self.users.items()):\n             s += \"%s\\t%s\\n\" % (key.expandtabs(1), val.expandtabs(1))\n \n-        open(self.getUserCacheFilename(), \"wb\").write(s)\n+        cache = io.open(self.getUserCacheFilename(), \"wb\")\n+        cache.write(as_bytes(s))\n+        cache.close()\n         self.userMapFromPerforceServer = True\n \n     def loadUserMapFromCache(self):\n+        \"\"\" Reads the P4 username to git email map\n+        \"\"\"\n         self.users = {}\n         self.userMapFromPerforceServer = False\n         try:\n-            cache = open(self.getUserCacheFilename(), \"rb\")\n+            cache = io.open(self.getUserCacheFilename(), \"rb\")\n             lines = cache.readlines()\n             cache.close()\n             for line in lines:\n-                entry = line.strip().split(\"\\t\")\n+                entry = as_string(line).strip().split(\"\\t\")\n                 self.users[entry[0]] = entry[1]\n         except IOError:\n             self.getUserMapFromPerforceServer()\n@@ -1721,21 +1780,27 @@ def prepareLogMessage(self, template, message, jobs):\n         return result\n \n     def patchRCSKeywords(self, file, pattern):\n-        # Attempt to zap the RCS keywords in a p4 controlled file matching the given pattern\n+        \"\"\" Attempt to zap the RCS keywords in a p4\n+            controlled file matching the given pattern\n+        \"\"\"\n+        bSubLine = as_bytes(r'$\\1$')\n         (handle, outFileName) = tempfile.mkstemp(dir='.')\n         try:\n-            outFile = os.fdopen(handle, \"w+\")\n-            inFile = open(file, \"r\")\n-            regexp = re.compile(pattern, re.VERBOSE)\n+            outFile = os.fdopen(handle, \"w+b\")\n+            inFile = open(file, \"rb\")\n+            regexp = re.compile(as_bytes(pattern), re.VERBOSE)\n             for line in inFile.readlines():\n-                line = regexp.sub(r'$\\1$', line)\n+                line = regexp.sub(bSubLine, line)\n                 outFile.write(line)\n             inFile.close()\n             outFile.close()\n+            outFile = None\n             # Forcibly overwrite the original file\n             os.unlink(file)\n             shutil.move(outFileName, file)\n         except:\n+            if outFile != None:\n+                outFile.close()\n             # cleanup our temporary file\n             os.unlink(outFileName)\n             print(\"Failed to strip RCS keywords in %s\" % file)\n@@ -2139,7 +2204,7 @@ def applyCommit(self, id):\n         tmpFile = os.fdopen(handle, \"w+b\")\n         if self.isWindows:\n             submitTemplate = submitTemplate.replace(\"\\n\", \"\\r\\n\")\n-        tmpFile.write(submitTemplate)\n+        tmpFile.write(as_bytes(submitTemplate))\n         tmpFile.close()\n \n         if self.prepare_p4_only:\n@@ -2189,8 +2254,8 @@ def applyCommit(self, id):\n                 message = tmpFile.read()\n                 tmpFile.close()\n                 if self.isWindows:\n-                    message = message.replace(\"\\r\\n\", \"\\n\")\n-                submitTemplate = message[:message.index(separatorLine)]\n+                    message = message.replace(b\"\\r\\n\", b\"\\n\")\n+                submitTemplate = message[:message.index(as_bytes(separatorLine))]\n \n                 if update_shelve:\n                     p4_write_pipe(['shelve', '-r', '-i'], submitTemplate)\n@@ -2833,8 +2898,11 @@ def stripRepoPath(self, path, prefixes):\n         return path\n \n     def splitFilesIntoBranches(self, commit):\n-        \"\"\"Look at each depotFile in the commit to figure out to what\n-           branch it belongs.\"\"\"\n+        \"\"\" Look at each depotFile in the commit to figure out to what\n+            branch it belongs.\n+\n+            Data in the commit will NOT be encoded\n+        \"\"\"\n \n         if self.clientSpecDirs:\n             files = self.extractFilesFromCommit(commit)\n@@ -2875,16 +2943,22 @@ def splitFilesIntoBranches(self, commit):\n         return branches\n \n     def writeToGitStream(self, gitMode, relPath, contents):\n-        self.gitStream.write('M %s inline %s\\n' % (gitMode, relPath))\n+        \"\"\" Writes the bytes[] 'contents' to the git fast-import\n+            with the given 'gitMode' and 'relPath' as the relative\n+            path.\n+        \"\"\"\n+        self.gitStream.write('M %s inline %s\\n' % (gitMode, as_string(relPath)))\n         self.gitStream.write('data %d\\n' % sum(len(d) for d in contents))\n         for d in contents:\n-            self.gitStream.write(d)\n+            self.gitStreamBytes.write(d)\n         self.gitStream.write('\\n')\n \n-    # output one file from the P4 stream\n-    # - helper for streamP4Files\n-\n     def streamOneP4File(self, file, contents):\n+        \"\"\" output one file from the P4 stream to the git inbound stream.\n+            helper for streamP4files.\n+\n+            contents should be a bytes (bytes)\n+        \"\"\"\n         relPath = self.stripRepoPath(file['depotFile'], self.branchPrefixes)\n         relPath = encodeWithUTF8(relPath, self.verbose)\n         if verbose:\n@@ -2892,7 +2966,7 @@ def streamOneP4File(self, file, contents):\n                 size = int(self.stream_file['fileSize'])\n             else:\n                 size = 0 # deleted files don't get a fileSize apparently\n-            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (file['depotFile'], relPath, size//1024//1024))\n+            sys.stdout.write('\\r%s --> %s (%i MB)\\n' % (path_as_string(file['depotFile']), as_string(relPath), size//1024//1024))\n             sys.stdout.flush()\n \n         (type_base, type_mods) = split_p4_type(file[\"type\"])\n@@ -2910,7 +2984,7 @@ def streamOneP4File(self, file, contents):\n                 # to nothing.  This causes p4 errors when checking out such\n                 # a change, and errors here too.  Work around it by ignoring\n                 # the bad symlink; hopefully a future change fixes it.\n-                print(\"\\nIgnoring empty symlink in %s\" % file['depotFile'])\n+                print(\"\\nIgnoring empty symlink in %s\" % path_as_string(file['depotFile']))\n                 return\n             elif data[-1] == '\\n':\n                 contents = [data[:-1]]\n@@ -2950,16 +3024,16 @@ def streamOneP4File(self, file, contents):\n             # Ideally, someday, this script can learn how to generate\n             # appledouble files directly and import those to git, but\n             # non-mac machines can never find a use for apple filetype.\n-            print(\"\\nIgnoring apple filetype file %s\" % file['depotFile'])\n+            print(\"\\nIgnoring apple filetype file %s\" % path_as_string(file['depotFile']))\n             return\n \n         # Note that we do not try to de-mangle keywords on utf16 files,\n         # even though in theory somebody may want that.\n-        pattern = p4_keywords_regexp_for_type(type_base, type_mods)\n+        pattern = as_bytes(p4_keywords_regexp_for_type(type_base, type_mods))\n         if pattern:\n             regexp = re.compile(pattern, re.VERBOSE)\n-            text = ''.join(contents)\n-            text = regexp.sub(r'$\\1$', text)\n+            text = b''.join(contents)\n+            text = regexp.sub(as_bytes(r'$\\1$'), text)\n             contents = [ text ]\n \n         if self.largeFileSystem:\n@@ -2978,15 +3052,19 @@ def streamOneP4Deletion(self, file):\n         if self.largeFileSystem and self.largeFileSystem.isLargeFile(relPath):\n             self.largeFileSystem.removeLargeFile(relPath)\n \n-    # handle another chunk of streaming data\n     def streamP4FilesCb(self, marshalled):\n+        \"\"\" Callback function for recording P4 chunks of data for streaming\n+            into GIT.\n+\n+            marshalled data is bytes[] from the caller\n+        \"\"\"\n \n         # catch p4 errors and complain\n         err = None\n-        if \"code\" in marshalled:\n-            if marshalled[\"code\"] == \"error\":\n-                if \"data\" in marshalled:\n-                    err = marshalled[\"data\"].rstrip()\n+        if b\"code\" in marshalled:\n+            if marshalled[b\"code\"] == b\"error\":\n+                if b\"data\" in marshalled:\n+                    err = marshalled[b\"data\"].rstrip()\n \n         if not err and 'fileSize' in self.stream_file:\n             required_bytes = int((4 * int(self.stream_file[\"fileSize\"])) - calcDiskFree())\n@@ -3008,11 +3086,11 @@ def streamP4FilesCb(self, marshalled):\n             # ignore errors, but make sure it exits first\n             self.importProcess.wait()\n             if f:\n-                die(\"Error from p4 print for %s: %s\" % (f, err))\n+                die(\"Error from p4 print for %s: %s\" % (path_as_string(f), err))\n             else:\n                 die(\"Error from p4 print: %s\" % err)\n \n-        if 'depotFile' in marshalled and self.stream_have_file_info:\n+        if b'depotFile' in marshalled and self.stream_have_file_info:\n             # start of a new file - output the old one first\n             self.streamOneP4File(self.stream_file, self.stream_contents)\n             self.stream_file = {}\n@@ -3022,13 +3100,16 @@ def streamP4FilesCb(self, marshalled):\n         # pick up the new file information... for the\n         # 'data' field we need to append to our array\n         for k in list(marshalled.keys()):\n-            if k == 'data':\n+            if k == b'data':\n                 if 'streamContentSize' not in self.stream_file:\n                     self.stream_file['streamContentSize'] = 0\n-                self.stream_file['streamContentSize'] += len(marshalled['data'])\n-                self.stream_contents.append(marshalled['data'])\n+                self.stream_file['streamContentSize'] += len(marshalled[b'data'])\n+                self.stream_contents.append(marshalled[b'data'])\n             else:\n-                self.stream_file[k] = marshalled[k]\n+                if k == b'depotFile':\n+                    self.stream_file[as_string(k)] = marshalled[k]\n+                else:\n+                    self.stream_file[as_string(k)] = as_string(marshalled[k])\n \n         if (verbose and\n             'streamContentSize' in self.stream_file and\n@@ -3037,13 +3118,14 @@ def streamP4FilesCb(self, marshalled):\n             size = int(self.stream_file[\"fileSize\"])\n             if size > 0:\n                 progress = 100.0*self.stream_file['streamContentSize']/size\n-                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (self.stream_file['depotFile'], progress, int(size//1024//1024)))\n+                sys.stdout.write('\\r%s %4.1f%% (%i MB)' % (path_as_string(self.stream_file['depotFile']), progress, int(size//1024//1024)))\n                 sys.stdout.flush()\n \n         self.stream_have_file_info = True\n \n-    # Stream directly from \"p4 files\" into \"git fast-import\"\n     def streamP4Files(self, files):\n+        \"\"\" Stream directly from \"p4 files\" into \"git fast-import\"\n+        \"\"\"\n         filesForCommit = []\n         filesToRead = []\n         filesToDelete = []\n@@ -3064,7 +3146,7 @@ def streamP4Files(self, files):\n             self.stream_contents = []\n             self.stream_have_file_info = False\n \n-            # curry self argument\n+            # Callback for P4 command to collect file content\n             def streamP4FilesCbSelf(entry):\n                 self.streamP4FilesCb(entry)\n \n@@ -3073,9 +3155,9 @@ def streamP4FilesCbSelf(entry):\n                 if 'shelved_cl' in f:\n                     # Handle shelved CLs using the \"p4 print file@=N\" syntax to print\n                     # the contents\n-                    fileArg = '%s@=%d' % (f['path'], f['shelved_cl'])\n+                    fileArg = b'%s@=%d' % (f['path'], as_bytes(f['shelved_cl']))\n                 else:\n-                    fileArg = '%s#%s' % (f['path'], f['rev'])\n+                    fileArg = b'%s#%s' % (f['path'], as_bytes(f['rev']))\n \n                 fileArgs.append(fileArg)\n \n@@ -3095,7 +3177,7 @@ def make_email(self, userid):\n \n     def streamTag(self, gitStream, labelName, labelDetails, commit, epoch):\n         \"\"\" Stream a p4 tag.\n-        commit is either a git commit, or a fast-import mark, \":<p4commit>\"\n+            commit is either a git commit, or a fast-import mark, \":<p4commit>\"\n         \"\"\"\n \n         if verbose:\n@@ -3167,7 +3249,22 @@ def commit(self, details, files, branch, parent = \"\", allow_empty=False):\n                 .format(details['change']))\n             return\n \n+        # fast-import:\n+        #'commit' SP <ref> LF\n+\t    #mark?\n+\t    #original-oid?\n+\t    #('author' (SP <name>)? SP LT <email> GT SP <when> LF)?\n+\t    #'committer' (SP <name>)? SP LT <email> GT SP <when> LF\n+\t    #('encoding' SP <encoding>)?\n+\t    #data\n+\t    #('from' SP <commit-ish> LF)?\n+\t    #('merge' SP <commit-ish> LF)*\n+\t    #(filemodify | filedelete | filecopy | filerename | filedeleteall | notemodify)*\n+\t    #LF?\n+\n+        #'commit' - <ref> is the name of the branch to make the commit on\n         self.gitStream.write(\"commit %s\\n\" % branch)\n+        #'mark' SP :<idnum>\n         self.gitStream.write(\"mark :%s\\n\" % details[\"change\"])\n         self.committedChanges.add(int(details[\"change\"]))\n         committer = \"\"\n@@ -3177,19 +3274,29 @@ def commit(self, details, files, branch, parent = \"\", allow_empty=False):\n \n         self.gitStream.write(\"committer %s\\n\" % committer)\n \n-        self.gitStream.write(\"data <<EOT\\n\")\n-        self.gitStream.write(details[\"desc\"])\n+        # Per https://git-scm.com/docs/git-fast-import\n+        # The preferred method for creating the commit message is to supply the\n+        # byte count in the data method and not to use a Delimited format.\n+        # Collect all the text in the commit message into a single string and\n+        # compute the byte count.\n+        commitText = details[\"desc\"]\n         if len(jobs) > 0:\n-            self.gitStream.write(\"\\nJobs: %s\" % (' '.join(jobs)))\n-\n+            commitText += \"\\nJobs: %s\" % (' '.join(jobs))\n         if not self.suppress_meta_comment:\n-            self.gitStream.write(\"\\n[git-p4: depot-paths = \\\"%s\\\": change = %s\" %\n-                                (','.join(self.branchPrefixes), details[\"change\"]))\n-            if len(details['options']) > 0:\n-                self.gitStream.write(\": options = %s\" % details['options'])\n-            self.gitStream.write(\"]\\n\")\n+            # coherce the path to the correct formatting in the branch prefixes as well.\n+            dispPaths = []\n+            for p in self.branchPrefixes:\n+                dispPaths += [path_as_string(p)]\n \n-        self.gitStream.write(\"EOT\\n\\n\")\n+            commitText += (\"\\n[git-p4: depot-paths = \\\"%s\\\": change = %s\" %\n+                                (','.join(dispPaths), details[\"change\"]))\n+            if len(details['options']) > 0:\n+                commitText += (\": options = %s\" % details['options'])\n+            commitText += \"]\"\n+        commitText += \"\\n\"\n+        self.gitStream.write(\"data %s\\n\" % len(as_bytes(commitText)))\n+        self.gitStream.write(commitText)\n+        self.gitStream.write(\"\\n\")\n \n         if len(parent) > 0:\n             if self.verbose:\n@@ -3596,30 +3703,35 @@ def sync_origin_only(self):\n                 system(\"git fetch origin\")\n \n     def importHeadRevision(self, revision):\n-        print(\"Doing initial import of %s from revision %s into %s\" % (' '.join(self.depotPaths), revision, self.branch))\n-\n+        # Re-encode depot text\n+        dispPaths = []\n+        utf8Paths = []\n+        for p in self.depotPaths:\n+            dispPaths += [path_as_string(p)]\n+        print(\"Doing initial import of %s from revision %s into %s\" % (' '.join(dispPaths), revision, self.branch))\n         details = {}\n         details[\"user\"] = \"git perforce import user\"\n-        details[\"desc\"] = (\"Initial import of %s from the state at revision %s\\n\"\n-                           % (' '.join(self.depotPaths), revision))\n+        details[\"desc\"] = (\"Initial import of %s from the state at revision %s\\n\" %\n+                           (' '.join(dispPaths), revision))\n         details[\"change\"] = revision\n         newestRevision = 0\n+        del dispPaths\n \n         fileCnt = 0\n         fileArgs = [\"%s...%s\" % (p,revision) for p in self.depotPaths]\n \n-        for info in p4CmdList([\"files\"] + fileArgs):\n+        for info in p4CmdList([\"files\"] + fileArgs, encode_cmd_output=False):\n \n-            if 'code' in info and info['code'] == 'error':\n+            if 'code' in info and info['code'] == b'error':\n                 sys.stderr.write(\"p4 returned an error: %s\\n\"\n-                                 % info['data'])\n-                if info['data'].find(\"must refer to client\") >= 0:\n+                                 % as_string(info['data']))\n+                if info['data'].find(b\"must refer to client\") >= 0:\n                     sys.stderr.write(\"This particular p4 error is misleading.\\n\")\n                     sys.stderr.write(\"Perhaps the depot path was misspelled.\\n\");\n                     sys.stderr.write(\"Depot path:  %s\\n\" % \" \".join(self.depotPaths))\n                 sys.exit(1)\n             if 'p4ExitCode' in info:\n-                sys.stderr.write(\"p4 exitcode: %s\\n\" % info['p4ExitCode'])\n+                sys.stderr.write(\"p4 exitcode: %s\\n\" % as_string(info['p4ExitCode']))\n                 sys.exit(1)\n \n \n@@ -3632,8 +3744,10 @@ def importHeadRevision(self, revision):\n                 #fileCnt = fileCnt + 1\n                 continue\n \n+            # Save all the file information, howerver do not translate the depotFile name at\n+            # this time. Leave that as bytes since the encoding may vary.\n             for prop in [\"depotFile\", \"rev\", \"action\", \"type\" ]:\n-                details[\"%s%s\" % (prop, fileCnt)] = info[prop]\n+                details[\"%s%s\" % (prop, fileCnt)] = (info[prop] if prop == \"depotFile\" else as_string(info[prop]))\n \n             fileCnt = fileCnt + 1\n \n@@ -3653,13 +3767,18 @@ def importHeadRevision(self, revision):\n             print(self.gitError.read())\n \n     def openStreams(self):\n+        \"\"\" Opens the fast import pipes.  Note that the git* streams are wrapped\n+            to expect Unicode text.  To send a raw byte Array, use the importProcess\n+            underlying port\n+        \"\"\"\n         self.importProcess = subprocess.Popen([\"git\", \"fast-import\"],\n                                               stdin=subprocess.PIPE,\n                                               stdout=subprocess.PIPE,\n                                               stderr=subprocess.PIPE);\n-        self.gitOutput = self.importProcess.stdout\n-        self.gitStream = self.importProcess.stdin\n-        self.gitError = self.importProcess.stderr\n+        self.gitOutput = Py23File(self.importProcess.stdout, verbose = self.verbose)\n+        self.gitStream = Py23File(self.importProcess.stdin, verbose = self.verbose)\n+        self.gitError = Py23File(self.importProcess.stderr, verbose = self.verbose)\n+        self.gitStreamBytes = self.importProcess.stdin\n \n     def closeStreams(self):\n         self.gitStream.close()\n@@ -4025,15 +4144,17 @@ def run(self, args):\n             self.cloneDestination = depotPaths[-1]\n             depotPaths = depotPaths[:-1]\n \n+        dispPaths = []\n         for p in depotPaths:\n             if not p.startswith(\"//\"):\n                 sys.stderr.write('Depot paths must start with \"//\": %s\\n' % p)\n                 return False\n+            dispPaths += [path_as_string(p)]\n \n         if not self.cloneDestination:\n             self.cloneDestination = self.defaultDestination(args)\n \n-        print(\"Importing from %s into %s\" % (', '.join(depotPaths), self.cloneDestination))\n+        print(\"Importing from %s into %s\" % (', '.join(dispPaths), path_as_string(self.cloneDestination)))\n \n         if not os.path.exists(self.cloneDestination):\n             os.makedirs(self.cloneDestination)\n-- \ngitgitgadget\n\n"},{"id":"387693","messageId":"e97ac0af8a33bc55c32b96d76136aef106d7b337.1575740863.git.gitgitgadget@gmail.com","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"[PATCH v5 12/15] git-p4: p4CmdList - support Unicode encoding","fromName":"Ben Keene via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2019-12-07T17:47:40Z","receivedAt":"2019-12-07T17:48:07Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"From: Ben Keene <seraphire@gmail.com>\n\nThe p4CmdList is a commonly used function in the git-p4 code. It is used\nto execute a command in P4 and return the results of the call in a list.\n\nThe problem is that p4CmdList takes bytes as the parameter data and\nreturns bytes in the return list.\n\nAdd a new optional parameter to the signature, encode_cmd_output, that\ndetermines if the dictionary values returned in the function output are\ntreated as bytes or as strings.\n\nChange the code to conditionally pass the output data through the\nas_string() function when encode_cmd_output is true. Otherwise the\nfunction should return the data as bytes.\n\nChange the code so that regardless of the setting of encode_cmd_output,\nthe dictionary keys in the return value will always be encoded with\nas_string().\n\nas_string(bytes) is a method defined in this project that treats the\nbyte data as a string. The word \"string\" is used because the meaning\nvaries depending on the version of Python:\n\n  - Python 2: The \"bytes\" are returned as \"str\", functionally a No-op.\n  - Python 3: The \"bytes\" are returned as a Unicode string.\n\nThe p4CmdList function returns a list of dictionaries that contain\nthe result of p4 command. If the callback (cb) is defined, the\nstandard output of the p4 command is redirected.\n\nData that is passed to the standard input of the P4 process should be\nas_bytes() to avoid conversion unicode encoding errors.\n\nas_bytes(text) is a method defined in this project that treats the text\ndata as a string that should be converted to a byte array (bytes). The\nbehavior of this function depends on the version of python:\n\n  - Python 2: The \"text\" is returned as \"str\", functionally a No-op.\n  - Python 3: The \"text\" is treated as a UTF-8 encoded Unicode string\n        and is decoded to bytes.\n\nAdditionally, change literal text prior to conversion to be literal\nbytes for the code that is evaluating the standard output from the\np4 call.\n\nAdd encode_cmd_output to the p4Cmd since this is a helper function that\nwraps the behavior of p4CmdList.\n\nSigned-off-by: Ben Keene <seraphire@gmail.com>\n---\n git-p4.py | 36 ++++++++++++++++++++++++++++--------\n 1 file changed, 28 insertions(+), 8 deletions(-)\n\ndiff --git a/git-p4.py b/git-p4.py\nindex 03829f796d..e8f31339e4 100755\n--- a/git-p4.py\n+++ b/git-p4.py\n@@ -716,7 +716,23 @@ def isModeExecChanged(src_mode, dst_mode):\n     return isModeExec(src_mode) != isModeExec(dst_mode)\n \n def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n-        errors_as_exceptions=False):\n+        errors_as_exceptions=False, encode_cmd_output=True):\n+    \"\"\" Executes a P4 command:  'cmd' optionally passing 'stdin' to the command's\n+        standard input via a temporary file with 'stdin_mode' mode.\n+\n+        Output from the command is optionally passed to the callback function 'cb'.\n+        If 'cb' is None, the response from the command is parsed into a list\n+        of resulting dictionaries. (For each block read from the process pipe.)\n+\n+        If 'skip_info' is true, information in a block read that has a code type of\n+        'info' will be skipped.\n+\n+        If 'errors_as_exceptions' is set to true (the default is false) the error\n+        code returned from the execution will generate an exception.\n+\n+        If 'encode_cmd_output' is set to true (the default) the data that is returned\n+        by this function will be passed through the \"as_string\" function.\n+    \"\"\"\n \n     if not isinstance(cmd, list):\n         cmd = \"-G \" + cmd\n@@ -739,7 +755,7 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n             stdin_file.write(stdin)\n         else:\n             for i in stdin:\n-                stdin_file.write(i + '\\n')\n+                stdin_file.write(as_bytes(i) + b'\\n')\n         stdin_file.flush()\n         stdin_file.seek(0)\n \n@@ -753,12 +769,15 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n         while True:\n             entry = marshal.load(p4.stdout)\n             if skip_info:\n-                if 'code' in entry and entry['code'] == 'info':\n+                if b'code' in entry and entry[b'code'] == b'info':\n                     continue\n             if cb is not None:\n                 cb(entry)\n             else:\n-                result.append(entry)\n+                out = {}\n+                for key, value in entry.items():\n+                    out[as_string(key)] = (as_string(value) if encode_cmd_output else value)\n+                result.append(out)\n     except EOFError:\n         pass\n     exitCode = p4.wait()\n@@ -785,8 +804,9 @@ def p4CmdList(cmd, stdin=None, stdin_mode='w+b', cb=None, skip_info=False,\n \n     return result\n \n-def p4Cmd(cmd):\n-    list = p4CmdList(cmd)\n+def p4Cmd(cmd, encode_cmd_output=True):\n+    \"\"\"Executes a P4 command and returns the results in a dictionary\"\"\"\n+    list = p4CmdList(cmd, encode_cmd_output=encode_cmd_output)\n     result = {}\n     for entry in list:\n         result.update(entry)\n@@ -1165,7 +1185,7 @@ def getClientSpec():\n     \"\"\"Look at the p4 client spec, create a View() object that contains\n        all the mappings, and return it.\"\"\"\n \n-    specList = p4CmdList(\"client -o\")\n+    specList = p4CmdList(\"client -o\", encode_cmd_output=False)\n     if len(specList) != 1:\n         die('Output from \"client -o\" is %d lines, expecting 1' %\n             len(specList))\n@@ -2609,7 +2629,7 @@ def update_client_spec_path_cache(self, files):\n         if len(fileArgs) == 0:\n             return  # All files in cache\n \n-        where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs)\n+        where_result = p4CmdList([\"-x\", \"-\", \"where\"], stdin=fileArgs, encode_cmd_output=False)\n         for res in where_result:\n             if \"code\" in res and res[\"code\"] == \"error\":\n                 # assume error is \"... file(s) not in client view\"\n-- \ngitgitgadget\n\n"},{"id":"387697","messageId":"20191207194756.GA43949@coredump.intra.peff.net","threadId":"52262","inReplyTo":"pull.463.v5.git.1575740863.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-12-07T19:47:56Z","receivedAt":"2019-12-07T19:47:59Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Dec 07, 2019 at 05:47:28PM +0000, Ben Keene via GitGitGadget wrote:\n\n> Ben Keene (13):\n>   git-p4: select P4 binary by operating-system\n>   git-p4: change the expansion test from basestring to list\n>   git-p4: promote encodeWithUTF8() to a global function\n>   git-p4: remove p4_write_pipe() and write_pipe() return values\n>   git-p4: add new support function gitConfigSet()\n>   git-p4: add casting helper functions for python 3 conversion\n>   git-p4: python 3 syntax changes\n>   git-p4: fix assumed path separators to be more Windows friendly\n>   git-p4: add Py23File() - helper class for stream writing\n>   git-p4: p4CmdList - support Unicode encoding\n>   git-p4: support Python 3 for basic P4 clone, sync, and submit (t9800)\n>   git-p4: added --encoding parameter to p4 clone\n>   git-p4: Add depot manipulation functions\n> \n> Jeff King (2):\n>   t/gitweb-lib.sh: drop confusing quotes\n>   t/gitweb-lib.sh: set $REQUEST_URI\n\nHmm, looks like rebasing leftovers. :) I think we can probably drop\nthese first two?\n\n-Peff\n"},{"id":"387701","messageId":"95ead4b6-21bb-1aa2-f16f-888e61a4e4c0@gmail.com","threadId":"52262","inReplyTo":"20191207194756.GA43949@coredump.intra.peff.net","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-07T21:27:29Z","receivedAt":"2019-12-07T21:27:33Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"Yes indeed!\n\nI hadn't pulled before I attempted the rebase, and got bit.  Yes those \nshouldn't be there!\n\nOn 12/7/2019 2:47 PM, Jeff King wrote:\n> On Sat, Dec 07, 2019 at 05:47:28PM +0000, Ben Keene via GitGitGadget wrote:\n>\n>> Ben Keene (13):\n>>    git-p4: select P4 binary by operating-system\n>>    git-p4: change the expansion test from basestring to list\n>>    git-p4: promote encodeWithUTF8() to a global function\n>>    git-p4: remove p4_write_pipe() and write_pipe() return values\n>>    git-p4: add new support function gitConfigSet()\n>>    git-p4: add casting helper functions for python 3 conversion\n>>    git-p4: python 3 syntax changes\n>>    git-p4: fix assumed path separators to be more Windows friendly\n>>    git-p4: add Py23File() - helper class for stream writing\n>>    git-p4: p4CmdList - support Unicode encoding\n>>    git-p4: support Python 3 for basic P4 clone, sync, and submit (t9800)\n>>    git-p4: added --encoding parameter to p4 clone\n>>    git-p4: Add depot manipulation functions\n>>\n>> Jeff King (2):\n>>    t/gitweb-lib.sh: drop confusing quotes\n>>    t/gitweb-lib.sh: set $REQUEST_URI\n> Hmm, looks like rebasing leftovers. :) I think we can probably drop\n> these first two?\n>\n> -Peff\n"},{"id":"387827","messageId":"xmqqd0cxuvql.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"e425ccc10fbc1f5e135eb59ffc84626f9d0ae4ff.1575740863.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 03/15] git-p4: select P4 binary by operating-system","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-09T19:47:30Z","receivedAt":"2019-12-09T19:47:38Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ben Keene <seraphire@gmail.com>\n>\n> The original code unconditionally used \"p4\" as the binary filename.\n>\n> Depending on the version of Git and Python installed, the perforce\n> program (p4) may not resolve on Windows without the program extension.\n>\n> Check the operating system (platform.system) and if it is reporting that\n> it is Windows, use the full filename of \"p4.exe\" instead of \"p4\"\n>\n> This change is Python 2 and Python 3 compatible.\n>\n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> ---\n>  git-p4.py | 5 ++++-\n>  1 file changed, 4 insertions(+), 1 deletion(-)\n\nMakes sense.  Ack from somebody on Windows (not required but would\nbe nice to have)?\n\n> diff --git a/git-p4.py b/git-p4.py\n> index 60c73b6a37..65e926758c 100755\n> --- a/git-p4.py\n> +++ b/git-p4.py\n> @@ -75,7 +75,10 @@ def p4_build_cmd(cmd):\n>      location. It means that hooking into the environment, or other configuration\n>      can be done more easily.\n>      \"\"\"\n> -    real_cmd = [\"p4\"]\n> +    if (platform.system() == \"Windows\"):\n> +        real_cmd = [\"p4.exe\"]\n> +    else:\n> +        real_cmd = [\"p4\"]\n>  \n>      user = gitConfig(\"git-p4.user\")\n>      if len(user) > 0:\n"},{"id":"387831","messageId":"xmqq8snlutyx.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"7170aface2270e8c46439c5c1e01d2b18cdf6fd0.1575740863.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 04/15] git-p4: change the expansion test from basestring to list","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-09T20:25:42Z","receivedAt":"2019-12-09T20:25:51Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> The original code used 'basestring' in a test to determine if a list or\n> literal string was passed into 9 different functions.  This is used to\n\ns/literal/a &/, probably, but I do not thin this is about \"literal\"\nat all.  Perhaps \"if a list or a string was passed ...\" is what you\nmeant, as the code seems to have two ways to represent a command\nline in the program, one as a single string with possibly multiple\ntokens on it, separated with IFS and quoted just like you would feed\nto shell, and the other as a list of strings, each element being a\nsingle argv[] element for the command to be invoked.\n\nSo the issue is that isinstance(X, basestring) used to be how you\nare supposed to see if X is a \"string\", and the logic were built\naround \"if we have a string, then keep the command in a string when\nmanipulating, and otherwise what we have must be a command line in a\nlist\".  But because Unicode string is not an instance of basestring,\nthe logic no longer work, so you'd flip the polarity around to use\n\"if it is not a list, then we must have a command in a string\".\n\nWhich makes sense.\n\n> determine if the shell should be invoked when calling subprocess\n> methods.\n\nThis is mostly true, but the use in p4_build_cmd() and the second\nuse among the two uses in p4CmdList() are different.\n\n\tSome codepaths can represent a command line the program\n\tinternally prepares to execute either as a single string\n\t(i.e. each token properly quoted, concatenated with $IFS) or\n\tas a list of argv[] elements, and there are 9 places where\n\twe say \"if X is isinstance(_, basestring), then do this\n\tthing to handle X as a command line in a single string; if\n\tnot, X is a command line in a list form\".\n\n\tThis does not work well with Python 3, as there is no\n\tbasestring (everything is Unicode now), and even with Python\n\t2, it was not an ideal way to tell the two cases apart,\n\tbecause an internally formed command line could have been in\n\ta single Unicode string.\n\n\tFlip the check to say \"if X is not a list, then handle X as\n\ta command line in a single string; otherwise treat it as a\n\tcommand line in a list form\".\n\n\tThis will get rid of references to 'basestring', to migrate\n\tthe code ready for Python 3.\n\nor something like that?\n\n"},{"id":"387950","messageId":"xmqq1rtarf3j.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"11d7703e411f1dced8a34defc68922ba44c614d5.1575740863.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 05/15] git-p4: promote encodeWithUTF8() to a global function","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-11T16:39:44Z","receivedAt":"2019-12-11T16:39:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ben Keene <seraphire@gmail.com>\n>\n> This changelist is an intermediate submission for migrating the P4\n> support from Python 2 to Python 3. The code needs access to the\n> encodeWithUTF8() for support of non-UTF8 filenames in the clone class as\n> well as the sync class.\n>\n> Move the function encodeWithUTF8() from the P4Sync class to a\n> stand-alone function.  This will allow other classes to use this\n> function without instanciating the P4Sync class.\n\nMakes quite a lot of sense, as I do not see a reason why this needs\nto be attached to any specific instance of P4Sync.\n\n"},{"id":"387951","messageId":"xmqqwob2pzty.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"95ead4b6-21bb-1aa2-f16f-888e61a4e4c0@gmail.com","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-11T16:54:49Z","receivedAt":"2019-12-11T16:54:58Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ben Keene <seraphire@gmail.com> writes:\n\n> Yes indeed!\n>\n> I hadn't pulled before I attempted the rebase, and got bit.  Yes those\n> shouldn't be there!\n\nSo, other than that, this is ready to be at least queued on 'pu' if\nnot 'next' at this point?\n\nThanks.\n"},{"id":"387952","messageId":"xmqqsglqpz2z.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"bc7009541b3f03c3065a7be2b569cd4bf91f7c05.1575740863.git.gitgitgadget@gmail.com","subject":"Re: [PATCH v5 07/15] git-p4: add new support function gitConfigSet()","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-11T17:11:00Z","receivedAt":"2019-12-11T17:11:08Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Ben Keene <seraphire@gmail.com>\n>\n> Add a new method gitConfigSet(). This method will set a value in the git\n> configuration cache list.\n>\n> Signed-off-by: Ben Keene <seraphire@gmail.com>\n> ---\n>  git-p4.py | 5 +++++\n>  1 file changed, 5 insertions(+)\n>\n> diff --git a/git-p4.py b/git-p4.py\n> index e7c24817ad..e020958083 100755\n> --- a/git-p4.py\n> +++ b/git-p4.py\n> @@ -860,6 +860,11 @@ def gitConfigList(key):\n>              _gitConfig[key] = []\n>      return _gitConfig[key]\n>  \n> +def gitConfigSet(key, value):\n> +    \"\"\" Set the git configuration key 'key' to 'value' for this session\n> +    \"\"\"\n> +    _gitConfig[key] = value\n> +\n>  def p4BranchesInGit(branchesAreInRemotes=True):\n>      \"\"\"Find all the branches whose names start with \"p4/\", looking\n>         in remotes or heads as specified by the argument.  Return\n\nI am not sure if we want to do this.  The function makes it look as\nif we are not just updating the cached version but also is updating\nthe underlying configuration file, effective even for future use,\nbut that is not what is happening (and you do not want to touch the\nconfiguration file with this helper anyway).  It is misleading.\n\nThis seems to be used only in one place in a later patch (14/15)\n \n         depotPaths = args\n \n+        # If we have an encoding provided, ignore what may already exist\n+        # in the registry. This will ensure we show the displayed values\n+        # using the correct encoding.\n+        if self.setPathEncoding:\n+            gitConfigSet(\"git-p4.pathEncoding\", self.setPathEncoding)\n+\n+        # If more than 1 path element is supplied, the last element\n+        # is the clone destination.\n         if not self.cloneDestination and len(depotPaths) > 1:\n             self.cloneDestination = depotPaths[-1]\n             depotPaths = depotPaths[:-1]\n\nand the reason why it is needed, I am guessing, is because pieces of\ncode that gets the control later in the flow will use \"git-p4.pathEncoding\"\nconfiguration variable to determine how the path need to be encoded.\n\nI think the right fix for that kind of problem is to make sure that\nwe clearly separate (1) what the configured value is, (2) what the\nvalue used to override the configured value for this single shot\ninvocation is, and (3) which value is used.  Perhaps the existing\ncode is fuzzy about the distinction and without allowing the caller\nto override, always uses the configured value, in which  case that\nis what needs to be fixed, perhaps?\n\nI see encodeWithUTF8(self, path) method of P4Sync class (I am\nworking this review on the version in 'master', not with any of the\nprevious steps in this series) unconditionally uses gitConfig() to\ngrab the configured value.  Probably the logical thing to do there\n(i.e. before your step 05/15) would have been to store in the P4Sync\ninstance (i.e. 'self') to hold its preferred path encoding in a new\nfield (say 'self.pathEncoding'), and use that if exists before\nconsulting the config and finally fall back to utf-8.  Or perhaps\nwhen the class is instantiated, populate that configured value to\nthe 'self.pathEncoding' field, and then override it by whatever\ncodepath you call gitConfigSet() in this series to instead override\nthat field.  That way, encodeWithUTF8(self, path) method (by the\nway, that is a horrible name, as the function can use arbitrary\nencoding and not just UTF-8) can always encode in self.pathEncoding\nwhich would make the logic flow simpler, I would imagine.\n\n"},{"id":"387953","messageId":"20191211171356.GA72178@generichostname","threadId":"52262","inReplyTo":"xmqqwob2pzty.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Denton Liu","fromEmail":"liu.denton@gmail.com","sentAt":"2019-12-11T17:13:56Z","receivedAt":"2019-12-11T17:13:02Z","isPatch":true,"sender":{"key":"liu.denton@gmail.com","avatar":"https://avatars.githubusercontent.com/u/9620836?v=4"},"body":"On Wed, Dec 11, 2019 at 08:54:49AM -0800, Junio C Hamano wrote:\n> Ben Keene <seraphire@gmail.com> writes:\n> \n> > Yes indeed!\n> >\n> > I hadn't pulled before I attempted the rebase, and got bit.  Yes those\n> > shouldn't be there!\n> \n> So, other than that, this is ready to be at least queued on 'pu' if\n> not 'next' at this point?\n\nFrom what I can tell, Ben agreed to have this series superseded by Yang\nZhao's competing series[1].\n\nThat being said, I haven't been following along too closely but it seems\nto me this series is further along and has received more review\nfeedback so maybe it should be picked up?\n\n[1]: https://lore.kernel.org/git/afa761cf-9c0e-cdcc-9c32-be88c5507042@gmail.com/\n"},{"id":"387959","messageId":"xmqq1rtapwy1.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"20191211171356.GA72178@generichostname","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-11T17:57:10Z","receivedAt":"2019-12-11T17:57:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Denton Liu <liu.denton@gmail.com> writes:\n\n> On Wed, Dec 11, 2019 at 08:54:49AM -0800, Junio C Hamano wrote:\n>> Ben Keene <seraphire@gmail.com> writes:\n>> \n>> > Yes indeed!\n>> >\n>> > I hadn't pulled before I attempted the rebase, and got bit.  Yes those\n>> > shouldn't be there!\n>> \n>> So, other than that, this is ready to be at least queued on 'pu' if\n>> not 'next' at this point?\n>\n> From what I can tell, Ben agreed to have this series superseded by Yang\n> Zhao's competing series[1].\n\nOK.  Let me not worry about this one, then, at least not yet.\n\n"},{"id":"387973","messageId":"CAE5ih78O4_ZPm1sxA=D9Ff-u3ga5Ax1CbvrFg0_E4KrRdUihDQ@mail.gmail.com","threadId":"52262","inReplyTo":"xmqq1rtapwy1.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Luke Diamand","fromEmail":"luke@diamand.org","sentAt":"2019-12-11T20:19:48Z","receivedAt":"2019-12-11T20:19:46Z","isPatch":true,"sender":{"key":"luke@diamand.org","avatar":"https://avatars.githubusercontent.com/u/5330967?v=4"},"body":"On Wed, 11 Dec 2019 at 17:57, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Denton Liu <liu.denton@gmail.com> writes:\n>\n> > On Wed, Dec 11, 2019 at 08:54:49AM -0800, Junio C Hamano wrote:\n> >> Ben Keene <seraphire@gmail.com> writes:\n> >>\n> >> > Yes indeed!\n> >> >\n> >> > I hadn't pulled before I attempted the rebase, and got bit.  Yes those\n> >> > shouldn't be there!\n> >>\n> >> So, other than that, this is ready to be at least queued on 'pu' if\n> >> not 'next' at this point?\n> >\n> > From what I can tell, Ben agreed to have this series superseded by Yang\n> > Zhao's competing series[1].\n>\n> OK.  Let me not worry about this one, then, at least not yet.\n>\n\nOh, I hadn't seen Yang's python3 changes!\n\nWhat do we need to do to get these ready for merging?\n"},{"id":"387980","messageId":"xmqq5zimo7r3.fsf@gitster-ct.c.googlers.com","threadId":"52262","inReplyTo":"CAE5ih78O4_ZPm1sxA=D9Ff-u3ga5Ax1CbvrFg0_E4KrRdUihDQ@mail.gmail.com","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-12-11T21:46:40Z","receivedAt":"2019-12-11T21:46:52Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Luke Diamand <luke@diamand.org> writes:\n\n> On Wed, 11 Dec 2019 at 17:57, Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> Denton Liu <liu.denton@gmail.com> writes:\n>>\n>> > On Wed, Dec 11, 2019 at 08:54:49AM -0800, Junio C Hamano wrote:\n>> >> Ben Keene <seraphire@gmail.com> writes:\n>> >>\n>> >> > Yes indeed!\n>> >> >\n>> >> > I hadn't pulled before I attempted the rebase, and got bit.  Yes those\n>> >> > shouldn't be there!\n>> >>\n>> >> So, other than that, this is ready to be at least queued on 'pu' if\n>> >> not 'next' at this point?\n>> >\n>> > From what I can tell, Ben agreed to have this series superseded by Yang\n>> > Zhao's competing series[1].\n>>\n>> OK.  Let me not worry about this one, then, at least not yet.\n>>\n>\n> Oh, I hadn't seen Yang's python3 changes!\n\nI haven't been paying attention to them either.  The patches I\nstarted commenting on from Ben were easy to read and understand, and\nI didn't even know until Denton pointed out that Ben's series\nyielded the way.\n\n> What do we need to do to get these ready for merging?\n\nSomebody needs to take the ownership of the topic---we cannot afford\nto have two independently made topics competing reviewers' attention.\n\nIf Ben wants to drop his version and instead wants to use Yang's\nones, that's OK but Ben probably is in a lot better position than\nbystanders like me to review and comment on Yang's to suggest\nimprovements, if he hasn't done so.  The same for those who reviewed\nBen's series earlier.\n\nIt would make sure that the single topic a combined effort to\nproduce the best of both topics.  If there is something Ben's\npatches did that is lacking in Yang's, it may be worth rebuilding it\non top of Yang's series.\n"},{"id":"387987","messageId":"CABvFv3L_oHqiGAcA1QyyQD-YZNeouvNVaqhp65qU2ea3TfuRRQ@mail.gmail.com","threadId":"52262","inReplyTo":"xmqq5zimo7r3.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Yang Zhao","fromEmail":"yang.zhao@skyboxlabs.com","sentAt":"2019-12-11T22:30:50Z","receivedAt":"2019-12-11T22:29:12Z","isPatch":true,"sender":{"key":"yang.zhao@skyboxlabs.com","avatar":"https://avatars.githubusercontent.com/u/45857825?v=4"},"body":"On Wed, Dec 11, 2019 at 1:46 PM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> Luke Diamand <luke@diamand.org> writes:\n>\n> > On Wed, 11 Dec 2019 at 17:57, Junio C Hamano <gitster@pobox.com> wrote:\n> >>\n> >> Denton Liu <liu.denton@gmail.com> writes:\n> >>\n> >> > On Wed, Dec 11, 2019 at 08:54:49AM -0800, Junio C Hamano wrote:\n> >> > From what I can tell, Ben agreed to have this series superseded by Yang\n> >> > Zhao's competing series[1].\n> >>\n> >> OK.  Let me not worry about this one, then, at least not yet.\n> >>\n> >\n> > Oh, I hadn't seen Yang's python3 changes!\n...\n> > What do we need to do to get these ready for merging?\n>\n> Somebody needs to take the ownership of the topic---we cannot afford\n> to have two independently made topics competing reviewers' attention.\n>\n> If Ben wants to drop his version and instead wants to use Yang's\n> ones, that's OK but Ben probably is in a lot better position than\n> bystanders like me to review and comment on Yang's to suggest\n> improvements, if he hasn't done so.  The same for those who reviewed\n> Ben's series earlier.\n>\n> It would make sure that the single topic a combined effort to\n> produce the best of both topics.  If there is something Ben's\n> patches did that is lacking in Yang's, it may be worth rebuilding it\n> on top of Yang's series.\n\nSorry about the bit of communication mess there. I should have paid\nmore attention to who were chiming in to Ben's series and added CCs\nappropriately. The timing was definitely a bit awkward as we were both\nonly dedicating part of work-time to the patchsets.\n\nThe outcome of discussion between Ben and I were that it made the most\nsense to use my set as the base to rebuild his quality-of-life\nchanges. My patchset (plus one missing change I've not sent out yet)\nwill pass all existing tests. I will take ownership of this merge,\nprobably as a separate patchset.\n\n-- \nYang\n"},{"id":"388014","messageId":"7dd1ccdf-7c11-6027-c0b2-0bd95077ff72@gmail.com","threadId":"52262","inReplyTo":"CABvFv3L_oHqiGAcA1QyyQD-YZNeouvNVaqhp65qU2ea3TfuRRQ@mail.gmail.com","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-12T14:13:43Z","receivedAt":"2019-12-12T14:13:46Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/11/2019 5:30 PM, Yang Zhao wrote:\n> On Wed, Dec 11, 2019 at 1:46 PM Junio C Hamano <gitster@pobox.com> wrote:\n>> Luke Diamand <luke@diamand.org> writes:\n>>\n>>> On Wed, 11 Dec 2019 at 17:57, Junio C Hamano <gitster@pobox.com> wrote:\n>>>> Denton Liu <liu.denton@gmail.com> writes:\n>>>>\n>>>>> On Wed, Dec 11, 2019 at 08:54:49AM -0800, Junio C Hamano wrote:\n>>>>>  From what I can tell, Ben agreed to have this series superseded by Yang\n>>>>> Zhao's competing series[1].\n>>>> OK.  Let me not worry about this one, then, at least not yet.\n>>>>\n>>> Oh, I hadn't seen Yang's python3 changes!\n> ...\n>>> What do we need to do to get these ready for merging?\n>> Somebody needs to take the ownership of the topic---we cannot afford\n>> to have two independently made topics competing reviewers' attention.\n>>\n>> If Ben wants to drop his version and instead wants to use Yang's\n>> ones, that's OK but Ben probably is in a lot better position than\n>> bystanders like me to review and comment on Yang's to suggest\n>> improvements, if he hasn't done so.  The same for those who reviewed\n>> Ben's series earlier.\n>>\n>> It would make sure that the single topic a combined effort to\n>> produce the best of both topics.  If there is something Ben's\n>> patches did that is lacking in Yang's, it may be worth rebuilding it\n>> on top of Yang's series.\n> Sorry about the bit of communication mess there. I should have paid\n> more attention to who were chiming in to Ben's series and added CCs\n> appropriately. The timing was definitely a bit awkward as we were both\n> only dedicating part of work-time to the patchsets.\n\n\nSorry for the silence, I was heads down on another issue at work.  I had \ntried to pull Yang's work down to my machine and had trouble getting it \nto run under my configuration but I think I had a mixed environment.  \nI'm going to reset everything on my machine and try again.  Since I'm \nnot a python developer and I won't be able to devote the overall time \nthat Yang will, I'm deferring the changeset to Yang's code. I posted the \ncode so that the work that I had done would be visible, I didn't mean to \ncause the cross-talk!\n\n>\n> The outcome of discussion between Ben and I were that it made the most\n> sense to use my set as the base to rebuild his quality-of-life\n> changes. My patchset (plus one missing change I've not sent out yet)\n> will pass all existing tests. I will take ownership of this merge,\n> probably as a separate patchset.\n\n\n"},{"id":"388116","messageId":"7141d2c7-8b6f-d4a9-f6cf-6660f42fb5e1@gmail.com","threadId":"52262","inReplyTo":"xmqq8snlutyx.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v5 04/15] git-p4: change the expansion test from basestring to list","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-13T14:40:00Z","receivedAt":"2019-12-13T20:37:55Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/9/2019 3:25 PM, Junio C Hamano wrote:\n> \"Ben Keene via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n>\n>> The original code used 'basestring' in a test to determine if a list or\n>> literal string was passed into 9 different functions.  This is used to\n>> This is mostly true, but the use in p4_build_cmd() and the second\n...\n> use among the two uses in p4CmdList() are different.\n>\n> \tSome codepaths can represent a command line the program\n> \tinternally prepares to execute either as a single string\n> \t(i.e. each token properly quoted, concatenated with $IFS) or\n> \tas a list of argv[] elements, and there are 9 places where\n> \twe say \"if X is isinstance(_, basestring), then do this\n> \tthing to handle X as a command line in a single string; if\n> \tnot, X is a command line in a list form\".\n>\n> \tThis does not work well with Python 3, as there is no\n> \tbasestring (everything is Unicode now), and even with Python\n> \t2, it was not an ideal way to tell the two cases apart,\n> \tbecause an internally formed command line could have been in\n> \ta single Unicode string.\n>\n> \tFlip the check to say \"if X is not a list, then handle X as\n> \ta command line in a single string; otherwise treat it as a\n> \tcommand line in a list form\".\n>\n> \tThis will get rid of references to 'basestring', to migrate\n> \tthe code ready for Python 3.\n>\n> or something like that?\n>\nI incorporated this into my suggested submission to Yang's work. Thank you.\n"},{"id":"388151","messageId":"16668dda-391d-275d-8588-cb1affb35f44@gmail.com","threadId":"52262","inReplyTo":"7dd1ccdf-7c11-6027-c0b2-0bd95077ff72@gmail.com","subject":"Re: [PATCH v5 00/15] git-p4.py: Cast byte strings to unicode strings in python3 - Code Review","fromName":"Ben Keene","fromEmail":"seraphire@gmail.com","sentAt":"2019-12-13T19:42:44Z","receivedAt":"2019-12-13T20:41:04Z","isPatch":true,"sender":{"key":"seraphire@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22523774?v=4"},"body":"\nOn 12/12/2019 9:13 AM, Ben Keene wrote:\n>\n> On 12/11/2019 5:30 PM, Yang Zhao wrote:\n>> On Wed, Dec 11, 2019 at 1:46 PM Junio C Hamano <gitster@pobox.com> \n>> wrote:\n>>> Luke Diamand <luke@diamand.org> writes:\n>>>\n>>>> On Wed, 11 Dec 2019 at 17:57, Junio C Hamano <gitster@pobox.com> \n>>>> wrote:\n>>>>> Denton Liu <liu.denton@gmail.com> writes:\n>>>>>\n>>>>>> On Wed, Dec 11, 2019 at 08:54:49AM -0800, Junio C Hamano wrote:\n>>>>>>  From what I can tell, Ben agreed to have this series superseded \n>>>>>> by Yang\n>>>>>> Zhao's competing series[1].\n>>>>> OK.  Let me not worry about this one, then, at least not yet.\n>>>>>\n>>>> Oh, I hadn't seen Yang's python3 changes!\n>> ...\n>>>> What do we need to do to get these ready for merging?\n>>> Somebody needs to take the ownership of the topic---we cannot afford\n>>> to have two independently made topics competing reviewers' attention.\n>>>\n>>> If Ben wants to drop his version and instead wants to use Yang's\n>>> ones, that's OK but Ben probably is in a lot better position than\n>>> bystanders like me to review and comment on Yang's to suggest\n>>> improvements, if he hasn't done so.  The same for those who reviewed\n>>> Ben's series earlier.\n>>>\n>>> It would make sure that the single topic a combined effort to\n>>> produce the best of both topics.  If there is something Ben's\n>>> patches did that is lacking in Yang's, it may be worth rebuilding it\n>>> on top of Yang's series.\n\nI reviewed Yang's changes and added comments to his commits in\nGitHub. Except for using the version number instead of feature\nmapping, all of his changes are much simpler and cleaner than\nI was proposing and I expect will get adoption more quickly.\n\nI recommend the inclusion of the one change I added previously that\nremoves the need for the basestring entirely. This moots some of\nneed for the version checking.\n\nKind regards,\n\nBen\n\n"}]}