{"thread":{"id":"48531","subject":"[RFC/PATCH 0/7] rerere: handle nested conflicts","startedAt":"2018-05-20T21:12:16Z","lastAt":"2018-09-01T09:05:34Z","messageCount":84,"participants":["Thomas Gummerer","Stefan Beller","Junio C Hamano","Simon Ruderich","Ævar Arnfjörð Bjarmason"],"isPatch":true,"patchVersion":1,"patchTotal":7},"messages":[{"id":"348131","messageId":"20180520211210.1248-1-t.gummerer@gmail.com","threadId":"48531","inReplyTo":null,"subject":"[RFC/PATCH 0/7] rerere: handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-20T21:12:03Z","receivedAt":"2018-05-20T21:12:16Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"I started this whole patch series when I did a git rebase, and was too\nlazy to resolve a conflict and just added the file with the conflict\nmarkers and continued.  Once I got nested conflicts in the file, I\ndecided to abort the rebase with 'git rebase --abort' and got a\nsegfault in 'git rerere clear'.\n\nEven if we can't handle the conflict, we shouldn't end with crashing\n'git rerere clear'.  While trying to understand how 'git rerere' works\ninternally I noticed some other improvements that could be made, such\nas marking the strings for translation and adding some docs on how\nrerere works, since I had to find out from the code, and reading some\ndocumentation would definitely have been helpful.\n\nThe next patches are more related to the actual problem I encountered,\nfirst fixing the the possible crashing of 'git rerere clear' when we\ncan't handle conflicts in a file, and then actually trying to handle\nnested conflicts.\n\nI don't know if it's actually worth trying to handle nested conflicts,\nas they are more than likely a very rare use-case, but on the other\nhand resolving such conflicts is especially painful, so only having to\ndo it once would be much nicer.\n\nThis whole patch series is marked as RFC/PATCH, as this is my first\ntime touching the rerere code, so I may well misunderstand some bits\nof the code.\n\nThomas Gummerer (7):\n  rerere: unify error message when read_cache fails\n  rerere: mark strings for translation\n  rerere: add some documentation\n  rerere: fix crash when conflict goes unresolved\n  rerere: only return whether a path has conflicts or not\n  rerere: factor out handle_conflict function\n  rerere: teach rerere to handle nested conflicts\n\n Documentation/technical/rerere.txt |  43 +++++\n rerere.c                           | 244 ++++++++++++++---------------\n t/t4200-rerere.sh                  |  25 +++\n 3 files changed, 186 insertions(+), 126 deletions(-)\n create mode 100644 Documentation/technical/rerere.txt\n\n-- \n2.17.0.588.g4d217cdf8e.dirty\n\n"},{"id":"348132","messageId":"20180520211210.1248-2-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180520211210.1248-1-t.gummerer@gmail.com","subject":"[RFC/PATCH 1/7] rerere: unify error message when read_cache fails","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-20T21:12:04Z","receivedAt":"2018-05-20T21:12:17Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"We have multiple different variants of the error message we show to\nthe user if 'read_cache' fails.  The \"Could not read index\" variant we\nare using in 'rerere.c' is currently not used anywhere in translated\nform.\n\nAs a subsequent commit will mark all output that comes from 'rerere.c'\nfor translation, make the life of the translators a little bit easier\nby using a string that is used elsewhere, and marked for translation\nthere, and thus most likely already translated.\n\n\"index file corrupt\" seems to be the most common error message we show\nwhen 'read_cache' fails, so use that here as well.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n\n\"index file corrupt\" is also what Stefan chose for his series unifying\nthese error messages (and 'die'ing, which I'm not sure is the right\nthing to do here as also mentioned in my reply to [1]).  I'm happy to\ndrop this if we decide to go with that series.\n\n[1]: <20180516222118.233868-8-sbeller@google.com>\n\n rerere.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 18cae2d11c..4b4869662d 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -568,7 +568,7 @@ static int find_conflict(struct string_list *conflict)\n {\n \tint i;\n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -601,7 +601,7 @@ int rerere_remaining(struct string_list *merge_rr)\n \tif (setup_rerere(merge_rr, RERERE_READONLY))\n \t\treturn 0;\n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -1104,7 +1104,7 @@ int rerere_forget(struct pathspec *pathspec)\n \tstruct string_list merge_rr = STRING_LIST_INIT_DUP;\n \n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfd = setup_rerere(&merge_rr, RERERE_NOAUTOUPDATE);\n \tif (fd < 0)\n-- \n2.17.0.588.g4d217cdf8e.dirty\n\n"},{"id":"348133","messageId":"20180520211210.1248-3-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180520211210.1248-1-t.gummerer@gmail.com","subject":"[RFC/PATCH 2/7] rerere: mark strings for translation","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-20T21:12:05Z","receivedAt":"2018-05-20T21:12:17Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"'git rerere' is considered a plumbing command and as such its output\nshould be translated.  Its functionality is also only enabled through\na config setting, so scripts really shouldn't rely on its output\neither way.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 68 ++++++++++++++++++++++++++++----------------------------\n 1 file changed, 34 insertions(+), 34 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 4b4869662d..af5e6179a9 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -212,7 +212,7 @@ static void read_rr(struct string_list *rr)\n \n \t\t/* There has to be the hash, tab, path and then NUL */\n \t\tif (buf.len < 42 || get_sha1_hex(buf.buf, sha1))\n-\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \n \t\tif (buf.buf[40] != '.') {\n \t\t\tvariant = 0;\n@@ -221,10 +221,10 @@ static void read_rr(struct string_list *rr)\n \t\t\terrno = 0;\n \t\t\tvariant = strtol(buf.buf + 41, &path, 10);\n \t\t\tif (errno)\n-\t\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \t\t}\n \t\tif (*(path++) != '\\t')\n-\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \t\tbuf.buf[40] = '\\0';\n \t\tid = new_rerere_id_hex(buf.buf);\n \t\tid->variant = variant;\n@@ -259,12 +259,12 @@ static int write_rr(struct string_list *rr, int out_fd)\n \t\t\t\t    rr->items[i].string, 0);\n \n \t\tif (write_in_full(out_fd, buf.buf, buf.len) < 0)\n-\t\t\tdie(\"unable to write rerere record\");\n+\t\t\tdie(_(\"unable to write rerere record\"));\n \n \t\tstrbuf_release(&buf);\n \t}\n \tif (commit_lock_file(&write_lock) != 0)\n-\t\tdie(\"unable to write rerere record\");\n+\t\tdie(_(\"unable to write rerere record\"));\n \treturn 0;\n }\n \n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"Could not open %s\", path);\n+\t\treturn error_errno(_(\"Could not open %s\"), path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"Could not write %s\", output);\n+\t\t\terror_errno(_(\"Could not write %s\"), output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"There were errors while writing %s (%s)\",\n+\t\terror(_(\"There were errors while writing %s (%s)\"),\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"Failed to flush %s\", path);\n+\t\tio.io.wrerror = error_errno(_(\"Failed to flush %s\"), path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"Could not parse conflict hunks in %s\", path);\n+\t\treturn error(_(\"Could not parse conflict hunks in %s\"), path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -568,7 +568,7 @@ static int find_conflict(struct string_list *conflict)\n {\n \tint i;\n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -601,7 +601,7 @@ int rerere_remaining(struct string_list *merge_rr)\n \tif (setup_rerere(merge_rr, RERERE_READONLY))\n \t\treturn 0;\n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -684,17 +684,17 @@ static int merge(const struct rerere_id *id, const char *path)\n \t * Mark that \"postimage\" was used to help gc.\n \t */\n \tif (utime(rerere_path(id, \"postimage\"), NULL) < 0)\n-\t\twarning_errno(\"failed utime() on %s\",\n+\t\twarning_errno(_(\"failed utime() on %s\"),\n \t\t\t      rerere_path(id, \"postimage\"));\n \n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"Could not open %s\", path);\n+\t\treturn error_errno(_(\"Could not open %s\"), path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"Could not write %s\", path);\n+\t\terror_errno(_(\"Could not write %s\"), path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"Writing %s failed\", path);\n+\t\treturn error_errno(_(\"Writing %s failed\"), path);\n \n out:\n \tfree(cur.ptr);\n@@ -715,13 +715,13 @@ static void update_paths(struct string_list *update)\n \t\tstruct string_list_item *item = &update->items[i];\n \t\tif (add_file_to_cache(item->string, 0))\n \t\t\texit(128);\n-\t\tfprintf(stderr, \"Staged '%s' using previous resolution.\\n\",\n+\t\tfprintf_ln(stderr, _(\"Staged '%s' using previous resolution.\"),\n \t\t\titem->string);\n \t}\n \n \tif (write_locked_index(&the_index, &index_lock,\n \t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED))\n-\t\tdie(\"Unable to write new index file\");\n+\t\tdie(_(\"Unable to write new index file\"));\n }\n \n static void remove_variant(struct rerere_id *id)\n@@ -753,7 +753,7 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \t\tif (!handle_file(path, NULL, NULL)) {\n \t\t\tcopy_file(rerere_path(id, \"postimage\"), path, 0666);\n \t\t\tid->collection->status[variant] |= RR_HAS_POSTIMAGE;\n-\t\t\tfprintf(stderr, \"Recorded resolution for '%s'.\\n\", path);\n+\t\t\tfprintf_ln(stderr, _(\"Recorded resolution for '%s'.\"), path);\n \t\t\tfree_rerere_id(rr_item);\n \t\t\trr_item->util = NULL;\n \t\t\treturn;\n@@ -787,9 +787,9 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \t\tif (rerere_autoupdate)\n \t\t\tstring_list_insert(update, path);\n \t\telse\n-\t\t\tfprintf(stderr,\n-\t\t\t\t\"Resolved '%s' using previous resolution.\\n\",\n-\t\t\t\tpath);\n+\t\t\tfprintf_ln(stderr,\n+\t\t\t\t   _(\"Resolved '%s' using previous resolution.\"),\n+\t\t\t\t   path);\n \t\tfree_rerere_id(rr_item);\n \t\trr_item->util = NULL;\n \t\treturn;\n@@ -803,11 +803,11 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \tif (id->collection->status[variant] & RR_HAS_POSTIMAGE) {\n \t\tconst char *path = rerere_path(id, \"postimage\");\n \t\tif (unlink(path))\n-\t\t\tdie_errno(\"cannot unlink stray '%s'\", path);\n+\t\t\tdie_errno(_(\"cannot unlink stray '%s'\"), path);\n \t\tid->collection->status[variant] &= ~RR_HAS_POSTIMAGE;\n \t}\n \tid->collection->status[variant] |= RR_HAS_PREIMAGE;\n-\tfprintf(stderr, \"Recorded preimage for '%s'\\n\", path);\n+\tfprintf_ln(stderr, _(\"Recorded preimage for '%s'\"), path);\n }\n \n static int do_plain_rerere(struct string_list *rr, int fd)\n@@ -879,7 +879,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"Could not create directory %s\", git_path_rr_cache());\n+\t\tdie(_(\"Could not create directory %s\"), git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1032,7 +1032,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t */\n \tret = handle_cache(path, sha1, NULL);\n \tif (ret < 1)\n-\t\treturn error(\"Could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(_(\"Could not parse conflict hunks in '%s'\"), path);\n \n \t/* Nuke the recorded resolution for the conflict */\n \tid = new_rerere_id(sha1);\n@@ -1050,7 +1050,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t\thandle_cache(path, sha1, rerere_path(id, \"thisimage\"));\n \t\tif (read_mmfile(&cur, rerere_path(id, \"thisimage\"))) {\n \t\t\tfree(cur.ptr);\n-\t\t\terror(\"Failed to update conflicted state in '%s'\", path);\n+\t\t\terror(_(\"Failed to update conflicted state in '%s'\"), path);\n \t\t\tgoto fail_exit;\n \t\t}\n \t\tcleanly_resolved = !try_merge(id, path, &cur, &result);\n@@ -1061,16 +1061,16 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t}\n \n \tif (id->collection->status_nr <= id->variant) {\n-\t\terror(\"no remembered resolution for '%s'\", path);\n+\t\terror(_(\"no remembered resolution for '%s'\"), path);\n \t\tgoto fail_exit;\n \t}\n \n \tfilename = rerere_path(id, \"postimage\");\n \tif (unlink(filename)) {\n \t\tif (errno == ENOENT)\n-\t\t\terror(\"no remembered resolution for %s\", path);\n+\t\t\terror(_(\"no remembered resolution for %s\"), path);\n \t\telse\n-\t\t\terror_errno(\"cannot unlink %s\", filename);\n+\t\t\terror_errno(_(\"cannot unlink %s\"), filename);\n \t\tgoto fail_exit;\n \t}\n \n@@ -1080,7 +1080,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t * the postimage.\n \t */\n \thandle_cache(path, sha1, rerere_path(id, \"preimage\"));\n-\tfprintf(stderr, \"Updated preimage for '%s'\\n\", path);\n+\tfprintf_ln(stderr, _(\"Updated preimage for '%s'\"), path);\n \n \t/*\n \t * And remember that we can record resolution for this\n@@ -1089,7 +1089,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \titem = string_list_insert(rr, path);\n \tfree_rerere_id(item);\n \titem->util = id;\n-\tfprintf(stderr, \"Forgot resolution for %s\\n\", path);\n+\tfprintf_ln(stderr, _(\"Forgot resolution for %s\"), path);\n \treturn 0;\n \n fail_exit:\n@@ -1104,7 +1104,7 @@ int rerere_forget(struct pathspec *pathspec)\n \tstruct string_list merge_rr = STRING_LIST_INIT_DUP;\n \n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfd = setup_rerere(&merge_rr, RERERE_NOAUTOUPDATE);\n \tif (fd < 0)\n@@ -1192,7 +1192,7 @@ void rerere_gc(struct string_list *rr)\n \tgit_config(git_default_config, NULL);\n \tdir = opendir(git_path(\"rr-cache\"));\n \tif (!dir)\n-\t\tdie_errno(\"unable to open rr-cache directory\");\n+\t\tdie_errno(_(\"unable to open rr-cache directory\"));\n \t/* Collect stale conflict IDs ... */\n \twhile ((e = readdir(dir))) {\n \t\tstruct rerere_dir *rr_dir;\n-- \n2.17.0.588.g4d217cdf8e.dirty\n\n"},{"id":"348134","messageId":"20180520211210.1248-4-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180520211210.1248-1-t.gummerer@gmail.com","subject":"[RFC/PATCH 3/7] rerere: add some documentation","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-20T21:12:06Z","receivedAt":"2018-05-20T21:12:20Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add some documentation for the logic behind the conflict normalization\nin rerere.  Also describe a bug that happens because we just linearly\nscan for conflict markers.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n\nThis documents my understanding of the rerere conflict normalization\nand conflict ID computation logic.  Writing this down helped me\nunderstand the logic, and I thought it may be useful to have this as\ndocumentation in Documentation/technical as well.  Junio: as you wrote\nthe original NEEDSWORK comment, did you have something more in mind\nhere that should be documented?\n\n Documentation/technical/rerere.txt | 43 ++++++++++++++++++++++++++++++\n rerere.c                           |  4 ---\n 2 files changed, 43 insertions(+), 4 deletions(-)\n create mode 100644 Documentation/technical/rerere.txt\n\ndiff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\nnew file mode 100644\nindex 0000000000..94cc6a7ef0\n--- /dev/null\n+++ b/Documentation/technical/rerere.txt\n@@ -0,0 +1,43 @@\n+Rerere\n+======\n+\n+This document describes the rerere logic.\n+\n+Conflict normalization\n+----------------------\n+\n+To try and re-do a conflict resolution, even when different merge\n+strategies are used, 'rerere' computes a conflict ID for each\n+conflict in the file.\n+\n+This is done by discarding the common ancestor version in the\n+diff3-style, and re-ordering the two sides of the conflict, in\n+alphabetic order.\n+\n+Using this technique a conflict that looks as follows when for example\n+'master' was merged into a topic branch:\n+\n+    <<<<<<< HEAD\n+    foo\n+    =======\n+    bar\n+    >>>>>>> master\n+\n+and the opposite way when the topic branch is merged into 'master':\n+\n+    <<<<<<< HEAD\n+    bar\n+    =======\n+    foo\n+    >>>>>>> topic\n+\n+can be recognized as the same conflict, and can automatically be\n+re-resolved by 'rerere', as the SHA-1 sum of the two conflicts would\n+be calculated from 'bar<NUL>foo<NUL>' in both cases.\n+\n+If there are multiple conflicts in one file, they are all appended to\n+one another, both in the 'preimage' file as well as in the conflict\n+ID.\n+\n+This is currently implemented by simply scanning through the file and\n+looking for conflict markers.\ndiff --git a/rerere.c b/rerere.c\nindex af5e6179a9..a02a38e072 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -394,10 +394,6 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n  * and NUL concatenated together.\n  *\n  * Return the number of conflict hunks found.\n- *\n- * NEEDSWORK: the logic and theory of operation behind this conflict\n- * normalization may deserve to be documented somewhere, perhaps in\n- * Documentation/technical/rerere.txt.\n  */\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n-- \n2.17.0.588.g4d217cdf8e.dirty\n\n"},{"id":"348135","messageId":"20180520211210.1248-5-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180520211210.1248-1-t.gummerer@gmail.com","subject":"[RFC/PATCH 4/7] rerere: fix crash when conflict goes unresolved","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-20T21:12:07Z","receivedAt":"2018-05-20T21:12:25Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently when a user doesn't resolve a conflict in a file, but\ncommits the file with the conflict markers, and later the file ends up\nin a state in which rerere can't handle it, subsequent rerere\noperations that are interested in that path, such as 'rerere clear' or\n'rerere forget <path>' will fail, or even worse in the case of 'rerere\nclear' segfault.\n\nSuch states include nested conflicts, or an extra conflict marker that\ndoesn't have any match.\n\nThis is because the first 'git rerere' when there was only one\nconflict in the file leaves an entry in the MERGE_RR file behind.  The\nnext 'git rerere' will then pick the rerere ID for that file up, and\nnot assign a new ID as it can't successfully calculate one.  It will\nhowever still try to do the rerere operation, because of the existing\nID.  As the handle_file function fails, it will remove the 'preimage'\nfor the ID in the process, while leaving the ID in the MERGE_RR file.\n\nNow when 'rerere clear' for example is run, it will segfault in\n'has_rerere_resolution', because status is NULL.\n\nTo fix this, remove the rerere ID from the MERGE_RR file in case we\ncan't handle it, and remove the folder for the ID.  Removing it\nunconditionally is fine here, because if the user would have resolved\nthe conflict and ran rerere, the entry would no longer be in the\nMERGE_RR file, so we wouldn't have this problem in the first place,\nwhile if the conflict was not resolved, the only thing that's left in\nthe folder is the 'preimage', which by itself will be regenerated by\ngit if necessary, so the user won't loose any work.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n\nI realize the test here may not be as complete as we would want it to\nbe.  But I first wanted to get some feedback on the approach, before\nspending too much time on a proper test (I did test it manually, and\nthe test does show that the original problem is fixed, but it probably\ndeserves some cleanup).\n\n rerere.c          | 12 +++++++-----\n t/t4200-rerere.sh | 25 +++++++++++++++++++++++++\n 2 files changed, 32 insertions(+), 5 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex a02a38e072..49ace8e108 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -824,10 +824,7 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\tstruct rerere_id *id;\n \t\tunsigned char sha1[20];\n \t\tconst char *path = conflict.items[i].string;\n-\t\tint ret;\n-\n-\t\tif (string_list_has_string(rr, path))\n-\t\t\tcontinue;\n+\t\tint ret, has_string;\n \n \t\t/*\n \t\t * Ask handle_file() to scan and assign a\n@@ -835,7 +832,12 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\t * yet.\n \t\t */\n \t\tret = handle_file(path, sha1, NULL);\n-\t\tif (ret < 1)\n+\t\thas_string = string_list_has_string(rr, path);\n+\t\tif (ret < 0 && has_string) {\n+\t\t\tremove_variant(string_list_lookup(rr, path)->util);\n+\t\t\tstring_list_remove(rr, path, 1);\n+\t\t}\n+\t\tif (ret < 1 || has_string)\n \t\t\tcontinue;\n \n \t\tid = new_rerere_id(sha1);\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex eaf18c81cb..27f8afc0b4 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -580,4 +580,29 @@ test_expect_success 'multiple identical conflicts' '\n \tcount_pre_post 0 0\n '\n \n+test_expect_success 'rerere with extra conflict markers keeps working' '\n+\tgit reset --hard &&\n+\n+\tgit checkout -b branch-1 master &&\n+\techo \"bar\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m two &&\n+\techo \"baz\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m three &&\n+\n+\tgit reset --hard &&\n+\tgit checkout -b branch-2 master &&\n+\techo \"foo\" >test &&\n+\tgit add test &&\n+\tgit commit -q -a -m one &&\n+\n+\ttest_must_fail git merge branch-1~ &&\n+\tgit add test &&\n+\tgit commit -q -m \"will solve conflicts later\" &&\n+\ttest_must_fail git merge branch-1 &&\n+\n+\tgit rerere clear\n+'\n+\n test_done\n-- \n2.17.0.588.g4d217cdf8e.dirty\n\n"},{"id":"348136","messageId":"20180520211210.1248-6-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180520211210.1248-1-t.gummerer@gmail.com","subject":"[RFC/PATCH 5/7] rerere: only return whether a path has conflicts or not","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-20T21:12:08Z","receivedAt":"2018-05-20T21:12:27Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"We currently return the exact number of conflict hunks a certain path\nhas from the 'handle_paths' function.  However all of its callers only\ncare whether there are conflicts or not or if there is an error.\nReturn only that information, and document that only that information\nis returned.  This will simplify the code in the subsequent steps.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 23 ++++++++++++-----------\n 1 file changed, 12 insertions(+), 11 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 49ace8e108..f3e658e374 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -393,12 +393,13 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n  * one side of the conflict, NUL, the other side of the conflict,\n  * and NUL concatenated together.\n  *\n- * Return the number of conflict hunks found.\n+ * Return 1 if conflict hunks are found, 0 if there are no conflict\n+ * hunks and -1 if an error occured.\n  */\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n \tgit_SHA_CTX ctx;\n-\tint hunk_no = 0;\n+\tint has_conflicts = 0;\n \tenum {\n \t\tRR_CONTEXT = 0, RR_SIDE_1, RR_SIDE_2, RR_ORIGINAL\n \t} hunk = RR_CONTEXT;\n@@ -426,7 +427,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\t\t\tgoto bad;\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n-\t\t\thunk_no++;\n+\t\t\thas_conflicts = 1;\n \t\t\thunk = RR_CONTEXT;\n \t\t\trerere_io_putconflict('<', marker_size, io);\n \t\t\trerere_io_putmem(one.buf, one.len, io);\n@@ -462,7 +463,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\tgit_SHA1_Final(sha1, &ctx);\n \tif (hunk != RR_CONTEXT)\n \t\treturn -1;\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n /*\n@@ -471,7 +472,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n  */\n static int handle_file(const char *path, unsigned char *sha1, const char *output)\n {\n-\tint hunk_no = 0;\n+\tint has_conflicts = 0;\n \tstruct rerere_io_file io;\n \tint marker_size = ll_merge_marker_size(path);\n \n@@ -491,7 +492,7 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \t\t}\n \t}\n \n-\thunk_no = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n+\thas_conflicts = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n@@ -500,14 +501,14 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tif (io.io.output && fclose(io.io.output))\n \t\tio.io.wrerror = error_errno(_(\"Failed to flush %s\"), path);\n \n-\tif (hunk_no < 0) {\n+\tif (has_conflicts < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n \t\treturn error(_(\"Could not parse conflict hunks in %s\"), path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n /*\n@@ -955,7 +956,7 @@ static int handle_cache(const char *path, unsigned char *sha1, const char *outpu\n \tmmfile_t mmfile[3] = {{NULL}};\n \tmmbuffer_t result = {NULL, 0};\n \tconst struct cache_entry *ce;\n-\tint pos, len, i, hunk_no;\n+\tint pos, len, i, has_conflicts;\n \tstruct rerere_io_mem io;\n \tint marker_size = ll_merge_marker_size(path);\n \n@@ -1009,11 +1010,11 @@ static int handle_cache(const char *path, unsigned char *sha1, const char *outpu\n \t * Grab the conflict ID and optionally write the original\n \t * contents with conflict markers out.\n \t */\n-\thunk_no = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n+\thas_conflicts = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n \tstrbuf_release(&io.input);\n \tif (io.io.output)\n \t\tfclose(io.io.output);\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n static int rerere_forget_one_path(const char *path, struct string_list *rr)\n-- \n2.17.0.588.g4d217cdf8e.dirty\n\n"},{"id":"348137","messageId":"20180520211210.1248-7-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180520211210.1248-1-t.gummerer@gmail.com","subject":"[RFC/PATCH 6/7] rerere: factor out handle_conflict function","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-20T21:12:09Z","receivedAt":"2018-05-20T21:12:29Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Factor out the handle_conflict function, which handles a single\nconflict in a path.  This is a preparation for the next step, where\nthis function will be re-used.  No functional changes intended.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 143 +++++++++++++++++++++++++------------------------------\n 1 file changed, 65 insertions(+), 78 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex f3e658e374..f3cfd1c09b 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -302,38 +302,6 @@ static void rerere_io_putstr(const char *str, struct rerere_io *io)\n \t\tferr_puts(str, io->output, &io->wrerror);\n }\n \n-/*\n- * Write a conflict marker to io->output (if defined).\n- */\n-static void rerere_io_putconflict(int ch, int size, struct rerere_io *io)\n-{\n-\tchar buf[64];\n-\n-\twhile (size) {\n-\t\tif (size <= sizeof(buf) - 2) {\n-\t\t\tmemset(buf, ch, size);\n-\t\t\tbuf[size] = '\\n';\n-\t\t\tbuf[size + 1] = '\\0';\n-\t\t\tsize = 0;\n-\t\t} else {\n-\t\t\tint sz = sizeof(buf) - 1;\n-\n-\t\t\t/*\n-\t\t\t * Make sure we will not write everything out\n-\t\t\t * in this round by leaving at least 1 byte\n-\t\t\t * for the next round, giving the next round\n-\t\t\t * a chance to add the terminating LF.  Yuck.\n-\t\t\t */\n-\t\t\tif (size <= sz)\n-\t\t\t\tsz -= (sz - size) + 1;\n-\t\t\tmemset(buf, ch, sz);\n-\t\t\tbuf[sz] = '\\0';\n-\t\t\tsize -= sz;\n-\t\t}\n-\t\trerere_io_putstr(buf, io);\n-\t}\n-}\n-\n static void rerere_io_putmem(const char *mem, size_t sz, struct rerere_io *io)\n {\n \tif (io->output)\n@@ -384,37 +352,25 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n \treturn isspace(*buf);\n }\n \n-/*\n- * Read contents a file with conflicts, normalize the conflicts\n- * by (1) discarding the common ancestor version in diff3-style,\n- * (2) reordering our side and their side so that whichever sorts\n- * alphabetically earlier comes before the other one, while\n- * computing the \"conflict ID\", which is just an SHA-1 hash of\n- * one side of the conflict, NUL, the other side of the conflict,\n- * and NUL concatenated together.\n- *\n- * Return 1 if conflict hunks are found, 0 if there are no conflict\n- * hunks and -1 if an error occured.\n- */\n-static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n+static void rerere_strbuf_putconflict(struct strbuf *buf, int ch, size_t size)\n+{\n+\tstrbuf_addchars(buf, ch, size);\n+\tstrbuf_addch(buf, '\\n');\n+}\n+\n+static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n+\t\t\t   int marker_size, git_SHA_CTX *ctx)\n {\n-\tgit_SHA_CTX ctx;\n-\tint has_conflicts = 0;\n \tenum {\n-\t\tRR_CONTEXT = 0, RR_SIDE_1, RR_SIDE_2, RR_ORIGINAL\n-\t} hunk = RR_CONTEXT;\n+\t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n+\t} hunk = RR_SIDE_1;\n \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n \tstruct strbuf buf = STRBUF_INIT;\n-\n-\tif (sha1)\n-\t\tgit_SHA1_Init(&ctx);\n-\n+\tint has_conflicts = 1;\n \twhile (!io->getline(&buf, io)) {\n-\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n-\t\t\tif (hunk != RR_CONTEXT)\n-\t\t\t\tgoto bad;\n-\t\t\thunk = RR_SIDE_1;\n-\t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n+\t\tif (is_cmarker(buf.buf, '<', marker_size))\n+\t\t\tgoto bad;\n+\t\telse if (is_cmarker(buf.buf, '|', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1)\n \t\t\t\tgoto bad;\n \t\t\thunk = RR_ORIGINAL;\n@@ -427,42 +383,73 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\t\t\tgoto bad;\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n-\t\t\thas_conflicts = 1;\n-\t\t\thunk = RR_CONTEXT;\n-\t\t\trerere_io_putconflict('<', marker_size, io);\n-\t\t\trerere_io_putmem(one.buf, one.len, io);\n-\t\t\trerere_io_putconflict('=', marker_size, io);\n-\t\t\trerere_io_putmem(two.buf, two.len, io);\n-\t\t\trerere_io_putconflict('>', marker_size, io);\n-\t\t\tif (sha1) {\n-\t\t\t\tgit_SHA1_Update(&ctx, one.buf ? one.buf : \"\",\n+\t\t\trerere_strbuf_putconflict(out, '<', marker_size);\n+\t\t\tstrbuf_addbuf(out, &one);\n+\t\t\trerere_strbuf_putconflict(out, '=', marker_size);\n+\t\t\tstrbuf_addbuf(out, &two);\n+\t\t\trerere_strbuf_putconflict(out, '>', marker_size);\n+\t\t\tif (ctx) {\n+\t\t\t\tgit_SHA1_Update(ctx, one.buf ? one.buf : \"\",\n \t\t\t\t\t    one.len + 1);\n-\t\t\t\tgit_SHA1_Update(&ctx, two.buf ? two.buf : \"\",\n+\t\t\t\tgit_SHA1_Update(ctx, two.buf ? two.buf : \"\",\n \t\t\t\t\t    two.len + 1);\n \t\t\t}\n-\t\t\tstrbuf_reset(&one);\n-\t\t\tstrbuf_reset(&two);\n+\t\t\tgoto out;\n \t\t} else if (hunk == RR_SIDE_1)\n \t\t\tstrbuf_addbuf(&one, &buf);\n \t\telse if (hunk == RR_ORIGINAL)\n \t\t\t; /* discard */\n \t\telse if (hunk == RR_SIDE_2)\n \t\t\tstrbuf_addbuf(&two, &buf);\n-\t\telse\n-\t\t\trerere_io_putstr(buf.buf, io);\n-\t\tcontinue;\n-\tbad:\n-\t\thunk = 99; /* force error exit */\n-\t\tbreak;\n \t}\n+bad:\n+\thas_conflicts = -1;\n+out:\n \tstrbuf_release(&one);\n \tstrbuf_release(&two);\n \tstrbuf_release(&buf);\n \n+\treturn has_conflicts;\n+}\n+\n+/*\n+ * Read contents a file with conflicts, normalize the conflicts\n+ * by (1) discarding the common ancestor version in diff3-style,\n+ * (2) reordering our side and their side so that whichever sorts\n+ * alphabetically earlier comes before the other one, while\n+ * computing the \"conflict ID\", which is just an SHA-1 hash of\n+ * one side of the conflict, NUL, the other side of the conflict,\n+ * and NUL concatenated together.\n+ *\n+ * Return 1 if conflict hunks are found, 0 if there are no conflict\n+ * hunks and -1 if an error occured.\n+ */\n+static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n+{\n+\tgit_SHA_CTX ctx;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf out = STRBUF_INIT;\n+\tint has_conflicts = 0;\n+\tif (sha1)\n+\t\tgit_SHA1_Init(&ctx);\n+\n+\twhile (!io->getline(&buf, io)) {\n+\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n+\t\t\thas_conflicts = handle_conflict(&out, io, marker_size,\n+\t\t\t\t\t\t\t    sha1 ? &ctx : NULL);\n+\t\t\tif (has_conflicts < 0)\n+\t\t\t\tbreak;\n+\t\t\trerere_io_putmem(out.buf, out.len, io);\n+\t\t\tstrbuf_reset(&out);\n+\t\t} else\n+\t\t\trerere_io_putstr(buf.buf, io);\n+\t}\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&out);\n+\n \tif (sha1)\n \t\tgit_SHA1_Final(sha1, &ctx);\n-\tif (hunk != RR_CONTEXT)\n-\t\treturn -1;\n+\n \treturn has_conflicts;\n }\n \n-- \n2.17.0.588.g4d217cdf8e.dirty\n\n"},{"id":"348138","messageId":"20180520211210.1248-8-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180520211210.1248-1-t.gummerer@gmail.com","subject":"[RFC/PATCH 7/7] rerere: teach rerere to handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-20T21:12:10Z","receivedAt":"2018-05-20T21:12:31Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently rerere can't handle nested conflicts and will error out when\nit encounters such conflicts.  Do that by recursively calling the\n'handle_conflict' function to normalize the conflict.\n\nThe conflict ID calculation here deserves some explanation:\n\nAs we are using the same handle_conflict function, the nested conflict\nis normalized the same way as for non-nested conflicts, which means\nthe ancestor in the diff3 case is stripped out, and the parts of the\nconflict are ordered alphabetically.\n\nThe conflict ID is however is only calculated in the top level\nhandle_conflict call, so it will include the markers that 'rerere'\nadds to the output.  e.g. say there's the following conflict:\n\n    <<<<<<< HEAD\n    1\n    =======\n    <<<<<<< HEAD\n    3\n    =======\n    2\n    >>>>>>> branch-2\n    >>>>>>> branch-3~\n\nit would be reordered as follows in the preimage:\n\n    <<<<<<<\n    1\n    =======\n    <<<<<<<\n    2\n    =======\n    3\n    >>>>>>>\n    >>>>>>>\n\nand the conflict ID would be calculated as\n\n    sha1(1<NUL><<<<<<<\n    2\n    =======\n    3\n    >>>>>>><NUL>)\n\nStripping out vs. leaving the conflict markers in place should have no\npractical impact, but it simplifies the implementation.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n\nNo automated test for this yet.  As mentioned in the cover letter as\nwell, I'm not sure if this is common enough for us to actually\nconsider this use case.  I don't know how nested conflicts could\nactually be created apart from committing a file with conflict\nmarkers, but maybe I'm just lacking imagination, so if someone has an\nexample for that I would be very grateful :)  If we decide to do this,\nit probably also merits a mention in\nDocumentation/technical/rerere.txt.\n\n rerere.c | 14 ++++++++++----\n 1 file changed, 10 insertions(+), 4 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex f3cfd1c09b..45e2bd6ff1 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -365,12 +365,18 @@ static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n \t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n \t} hunk = RR_SIDE_1;\n \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n-\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf buf = STRBUF_INIT, conflict = STRBUF_INIT;\n \tint has_conflicts = 1;\n \twhile (!io->getline(&buf, io)) {\n-\t\tif (is_cmarker(buf.buf, '<', marker_size))\n-\t\t\tgoto bad;\n-\t\telse if (is_cmarker(buf.buf, '|', marker_size)) {\n+\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n+\t\t\tif (handle_conflict(&conflict, io, marker_size, NULL) < 0)\n+\t\t\t\tgoto bad;\n+\t\t\tif (hunk == RR_SIDE_1)\n+\t\t\t\tstrbuf_addbuf(&one, &conflict);\n+\t\t\telse\n+\t\t\t\tstrbuf_addbuf(&two, &conflict);\n+\t\t\tstrbuf_release(&conflict);\n+\t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1)\n \t\t\t\tgoto bad;\n \t\t\thunk = RR_ORIGINAL;\n-- \n2.17.0.588.g4d217cdf8e.dirty\n\n"},{"id":"348241","messageId":"CAGZ79kZnMnG_88YC8bfopfw8o9qwTggC01HPSPddL+LU-QVjFQ@mail.gmail.com","threadId":"48531","inReplyTo":"20180520211210.1248-2-t.gummerer@gmail.com","subject":"Re: [RFC/PATCH 1/7] rerere: unify error message when read_cache fails","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2018-05-21T19:00:31Z","receivedAt":"2018-05-21T19:00:36Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Sun, May 20, 2018 at 2:12 PM, Thomas Gummerer <t.gummerer@gmail.com> wrote:\n> We have multiple different variants of the error message we show to\n> the user if 'read_cache' fails.  The \"Could not read index\" variant we\n> are using in 'rerere.c' is currently not used anywhere in translated\n> form.\n>\n> As a subsequent commit will mark all output that comes from 'rerere.c'\n> for translation, make the life of the translators a little bit easier\n> by using a string that is used elsewhere, and marked for translation\n> there, and thus most likely already translated.\n>\n> \"index file corrupt\" seems to be the most common error message we show\n> when 'read_cache' fails, so use that here as well.\n>\n> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> ---\n>\n> \"index file corrupt\" is also what Stefan chose for his series unifying\n> these error messages (and 'die'ing, which I'm not sure is the right\n> thing to do here as also mentioned in my reply to [1]).  I'm happy to\n> drop this if we decide to go with that series.\n\nAcked-by: <me>\n\nI'd happily have this patch instead of the one in my series.\n\nI was about to ask for translation, but the commit message hints\nat a follow up patch marking this for translation, so I'll read on.\n\nStefan\n"},{"id":"348413","messageId":"xmqq603dseom.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180520211210.1248-3-t.gummerer@gmail.com","subject":"Re: [RFC/PATCH 2/7] rerere: mark strings for translation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-05-24T07:20:09Z","receivedAt":"2018-05-24T07:20:15Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n>  \t\tif (write_in_full(out_fd, buf.buf, buf.len) < 0)\n> -\t\t\tdie(\"unable to write rerere record\");\n> +\t\t\tdie(_(\"unable to write rerere record\"));\n\nAs we'd be adding these new strings to the .po file, perhaps we\nwould want to downcase the first letter in the message to match the\nconvention?\n\n> ...\n> -\t\treturn error_errno(\"Could not open %s\", path);\n> +\t\treturn error_errno(_(\"Could not open %s\"), path);\n> ...\n> -\t\terror(\"There were errors while writing %s (%s)\",\n> +\t\terror(_(\"There were errors while writing %s (%s)\"),\n> ...\n> -\t\tio.io.wrerror = error_errno(\"Failed to flush %s\", path);\n> +\t\tio.io.wrerror = error_errno(_(\"Failed to flush %s\"), path);\n> ...\n> -\t\treturn error(\"Could not parse conflict hunks in %s\", path);\n> +\t\treturn error(_(\"Could not parse conflict hunks in %s\"), path);\n"},{"id":"348414","messageId":"xmqqr2m1quja.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180520211210.1248-4-t.gummerer@gmail.com","subject":"Re: [RFC/PATCH 3/7] rerere: add some documentation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-05-24T09:20:41Z","receivedAt":"2018-05-24T09:20:49Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> +Conflict normalization\n> +----------------------\n> +\n> +To try and re-do a conflict resolution, even when different merge\n> +strategies are used, 'rerere' computes a conflict ID for each\n> +conflict in the file.\n> +\n> +This is done by discarding the common ancestor version in the\n> +diff3-style, and re-ordering the two sides of the conflict, in\n> +alphabetic order.\n\ns/discarding.*-style/normalising the conflicted section to 'merge' style/\n\nThe motivation behind the normalization should probably be given\nupfront in the first paragraph.  It is to ensure the recorded\nresolutions can be looked up from the rerere database for\napplication, even when branches are merged in different order.  I am\nnot sure what you meant by even when different merge stratagies are\nused; I'd drop that if I were writing the paragraph.\n\n> +Using this technique a conflict that looks as follows when for example\n> +'master' was merged into a topic branch:\n> +\n> +    <<<<<<< HEAD\n> +    foo\n> +    =======\n> +    bar\n> +    >>>>>>> master\n> +\n> +and the opposite way when the topic branch is merged into 'master':\n> +\n> +    <<<<<<< HEAD\n> +    bar\n> +    =======\n> +    foo\n> +    >>>>>>> topic\n> +\n> +can be recognized as the same conflict, and can automatically be\n> +re-resolved by 'rerere', as the SHA-1 sum of the two conflicts would\n> +be calculated from 'bar<NUL>foo<NUL>' in both cases.\n\nYou earlier talked about normalizing and reordering, but did not\ntalk about \"concatenate both with NUL in between and hash\", so the\nexplanation in the last two lines are not quite understandable by\nmere mortals, even though I know which part of the code you are\nreferring to.  When you talk about hasing, you may want to make sure\nthe readers understand that the branch label on <<< and >>> lines\nare ignored.\n\n> +If there are multiple conflicts in one file, they are all appended to\n> +one another, both in the 'preimage' file as well as in the conflict\n> +ID.\n\nIn case it was not clear (and I do not think it is to those who only\nread your description and haven't thought things through\nthemselves), this concatenation is why the normalization by\nreordering is helpful.  Imagine that a common ancestor had a file\nwith a line with string \"A\" on it (I'll call such a line \"line A\"\nfor brevity in the following) in its early part, and line X in its\nlate part.  And then you fork four branches that do these things:\n\n    - AB: changes A to B\n    - AC: changes A to C\n    - XY: changes X to Y\n    - XZ: changes X to Z\n\nNow, forking a branch ABAC off of branch AB and then merging AC into\nit, and forking a branch ACAB off of branch AC and then merging AB\ninto it, would yield the conflict in a different order.  The former\nwould say \"A became B or C, what now?\" while the latter would say \"A\nbecame C or B, what now?\"\n\nBut the act of merging AC into ABAC and resolving the conflict to\nleave line D means that you declare: \n\n    After examining what branches AB and AC did, I believe that\n    making line A into line D is the best thing to do that is\n    compatible with what AB and AC wanted to do.\n\nSo the conflict we would see when merging AB into ACAB should be\nresolved the same way---it is the resolution that is in line with\nthat declaration.\n\nImagine that similarly you had previously forked branch XYXZ from\nXY, merged XZ into it, and resolved \"X became Y or Z\" into \"X became\nW\".\n\nNow, if you forked a branch ABXY from AB and then merged XY, then\nABXY would have line B in its early part and line Y in its later\npart.  Such a merge would be quite clean.  We can construct\n4 combinations using these four branches ((AB, AC) x (XY, XZ)).\n\nMerging ABXY and ACXZ would make \"an early A became B or C, a late X\nbecame Y or Z\" conflict, while merging ACXY and ABXZ would make \"an\nearly A became C or B, a late X became Y or Z\".  We can see there\nare 4 combinations of (\"B or C\", \"C or B\") x (\"X or Y\", \"Y or X\").\n\nBy sorting, we can give the conflict its canonical name, namely, \"an\nearly part became B or C, a late part becames X or Y\", and whenever\nany of these four patterns appear, we can get to the same conflict\nand resolution that we saw earlier.  Without the sorting, we will\nhave to somehow find a previous resolution from combinatorial\nexplosion ;-)\n\nThese days post ec34a8b1 (\"Merge branch 'jc/rerere-multi'\",\n2016-05-23), the conflict ID can safely collide, i.e. hash\ncollisions that drops completely different conflicts and their\nresolutions into the same .git/rr-cache/$id directory will not\ninterfere with proper operation of the system, thanks to that\nrerere-multi topic that allows us to store multiple preimage\nconflicts that happens to share the same conflict ID with their\ncorresponding postimage resolutions.\n\nIn theory, we *should* be able to stub out the SHA-1 computation and\ngive every conflict the same ID and rerere should still operate\ncorrectly, even though I haven't tried it yet myself.\n\n"},{"id":"348416","messageId":"xmqqefi1qrpj.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180520211210.1248-8-t.gummerer@gmail.com","subject":"Re: [RFC/PATCH 7/7] rerere: teach rerere to handle nested conflicts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-05-24T10:21:44Z","receivedAt":"2018-05-24T10:21:50Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> No automated test for this yet.  As mentioned in the cover letter as\n> well, I'm not sure if this is common enough for us to actually\n> consider this use case.  I don't know how nested conflicts could\n> actually be created apart from committing a file with conflict\n> markers,\n\nRecursive merge whose inner merge leaves conflict markers?\n\nOne thing that makes me wonder is that the conflict markers may not\n\"nest\" so nicely.  For example, if inner merges had two conflicts\nlike these:\n\n<<<\n <<<<<\n A\n =====\n B\n >>>>>\n===\n <<<<<\n A\n =====\n C\n >>>>>\n>>>\n\nwhere one side made something to A or B, while the other side made\nsomething (or something else) to A or C, I would imagine that the\nouter conflict could be \"optimized\" to produce this instead:\n\n\n <<<<<\n A\n =====\n<<<\n B\n===\n C\n>>>\n >>>>>\n\n"},{"id":"348417","messageId":"xmqqin7dqsl0.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180520211210.1248-6-t.gummerer@gmail.com","subject":"Re: [RFC/PATCH 5/7] rerere: only return whether a path has conflicts or not","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-05-24T10:02:51Z","receivedAt":"2018-05-24T10:22:21Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> We currently return the exact number of conflict hunks a certain path\n> has from the 'handle_paths' function.  However all of its callers only\n> care whether there are conflicts or not or if there is an error.\n> Return only that information, and document that only that information\n> is returned.  This will simplify the code in the subsequent steps.\n\nMakes sense.\n"},{"id":"348420","messageId":"xmqqmuwpqt9s.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180520211210.1248-5-t.gummerer@gmail.com","subject":"Re: [RFC/PATCH 4/7] rerere: fix crash when conflict goes unresolved","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-05-24T09:47:59Z","receivedAt":"2018-05-24T11:45:36Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> To fix this, remove the rerere ID from the MERGE_RR file in case we\n> can't handle it, and remove the folder for the ID.  Removing it\n> unconditionally is fine here, because if the user would have resolved\n> the conflict and ran rerere, the entry would no longer be in the\n> MERGE_RR file, so we wouldn't have this problem in the first place,\n\nI do not think removing the directory and losing _other_ conflicts\nand their resolutions, if they exist, is fine in the modern world\norder post rerere-multi update in 2016.  Well, it is just as safe as\n\"rm -rf .git/rr-cache/\" in the sense that it won't make Git start\nsegfaulting, but it is not fine as it is discarding information of\nconflicts that has nothing to do with the current one that is\nproblematic.\n\n"},{"id":"348464","messageId":"20180524185425.GB18193@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqqmuwpqt9s.fsf@gitster-ct.c.googlers.com","subject":"Re: [RFC/PATCH 4/7] rerere: fix crash when conflict goes unresolved","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-24T18:54:25Z","receivedAt":"2018-05-24T18:54:01Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 05/24, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > To fix this, remove the rerere ID from the MERGE_RR file in case we\n> > can't handle it, and remove the folder for the ID.  Removing it\n> > unconditionally is fine here, because if the user would have resolved\n> > the conflict and ran rerere, the entry would no longer be in the\n> > MERGE_RR file, so we wouldn't have this problem in the first place,\n> \n> I do not think removing the directory and losing _other_ conflicts\n> and their resolutions, if they exist, is fine in the modern world\n> order post rerere-multi update in 2016.  Well, it is just as safe as\n> \"rm -rf .git/rr-cache/\" in the sense that it won't make Git start\n> segfaulting, but it is not fine as it is discarding information of\n> conflicts that has nothing to do with the current one that is\n> problematic.\n\nSorry I botched the description here, and failed to describe what the\ncode is actually doing.  We're actually only removing the variant in\nthe MERGE_RR file, whose path we are now no longer able to handle.\nAnd I think that's fine to do, because if it is still in the MERGE_RR\nfile the conflict hasn't been resolved yet, afaiu.\n\nWill update the commit message.\n"},{"id":"348466","messageId":"20180524190724.GC18193@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqqefi1qrpj.fsf@gitster-ct.c.googlers.com","subject":"Re: [RFC/PATCH 7/7] rerere: teach rerere to handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-05-24T19:07:24Z","receivedAt":"2018-05-24T19:07:00Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 05/24, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > No automated test for this yet.  As mentioned in the cover letter as\n> > well, I'm not sure if this is common enough for us to actually\n> > consider this use case.  I don't know how nested conflicts could\n> > actually be created apart from committing a file with conflict\n> > markers,\n> \n> Recursive merge whose inner merge leaves conflict markers?\n\nThanks, lots of stuff in Git I still have to learn :)\n\n> One thing that makes me wonder is that the conflict markers may not\n> \"nest\" so nicely.  For example, if inner merges had two conflicts\n> like these:\n> \n> <<<\n>  <<<<<\n>  A\n>  =====\n>  B\n>  >>>>>\n> ===\n>  <<<<<\n>  A\n>  =====\n>  C\n>  >>>>>\n> >>>\n> \n> where one side made something to A or B, while the other side made\n> something (or something else) to A or C, I would imagine that the\n> outer conflict could be \"optimized\" to produce this instead:\n> \n> \n>  <<<<<\n>  A\n>  =====\n> <<<\n>  B\n> ===\n>  C\n> >>>\n>  >>>>>\n\nYeah, I do think that would be a nicer merge conflict to solve.  But I\nthink that should be done in a separate patch series if we decide to\ndo so.  When this one lands rerere will be able to handle the conflict\neither way :)\n"},{"id":"348504","messageId":"xmqqtvqwpm3r.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180524185425.GB18193@hank.intra.tgummerer.com","subject":"Re: [RFC/PATCH 4/7] rerere: fix crash when conflict goes unresolved","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-05-25T01:20:24Z","receivedAt":"2018-05-25T01:20:32Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Sorry I botched the description here, and failed to describe what the\n> code is actually doing.  We're actually only removing the variant in\n> the MERGE_RR file, whose path we are now no longer able to handle.\n\nOh, that's absolutely fine, then.  Thanks for a prompt update.\n"},{"id":"349138","messageId":"20180603114128.GD18193@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqqr2m1quja.fsf@gitster-ct.c.googlers.com","subject":"Re: [RFC/PATCH 3/7] rerere: add some documentation","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-03T11:41:28Z","receivedAt":"2018-06-03T11:40:55Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 05/24, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > +Conflict normalization\n> > +----------------------\n> > +\n> > +To try and re-do a conflict resolution, even when different merge\n> > +strategies are used, 'rerere' computes a conflict ID for each\n> > +conflict in the file.\n> > +\n> > +This is done by discarding the common ancestor version in the\n> > +diff3-style, and re-ordering the two sides of the conflict, in\n> > +alphabetic order.\n> \n> s/discarding.*-style/normalising the conflicted section to 'merge' style/\n> \n> The motivation behind the normalization should probably be given\n> upfront in the first paragraph.  It is to ensure the recorded\n> resolutions can be looked up from the rerere database for\n> application, even when branches are merged in different order.  I am\n> not sure what you meant by even when different merge stratagies are\n> used; I'd drop that if I were writing the paragraph.\n\nWhat I meant was when different conflict styles are used, and when the\nbranches are merged in different orders.  But merge strategies is\nobviously not a good word for that.  Will rephrase this.\n\n> > +Using this technique a conflict that looks as follows when for example\n> > +'master' was merged into a topic branch:\n> > +\n> > +    <<<<<<< HEAD\n> > +    foo\n> > +    =======\n> > +    bar\n> > +    >>>>>>> master\n> > +\n> > +and the opposite way when the topic branch is merged into 'master':\n> > +\n> > +    <<<<<<< HEAD\n> > +    bar\n> > +    =======\n> > +    foo\n> > +    >>>>>>> topic\n> > +\n> > +can be recognized as the same conflict, and can automatically be\n> > +re-resolved by 'rerere', as the SHA-1 sum of the two conflicts would\n> > +be calculated from 'bar<NUL>foo<NUL>' in both cases.\n> \n> You earlier talked about normalizing and reordering, but did not\n> talk about \"concatenate both with NUL in between and hash\", so the\n> explanation in the last two lines are not quite understandable by\n> mere mortals, even though I know which part of the code you are\n> referring to.  When you talk about hasing, you may want to make sure\n> the readers understand that the branch label on <<< and >>> lines\n> are ignored.\n> \n> > +If there are multiple conflicts in one file, they are all appended to\n> > +one another, both in the 'preimage' file as well as in the conflict\n> > +ID.\n> \n> In case it was not clear (and I do not think it is to those who only\n> read your description and haven't thought things through\n> themselves), this concatenation is why the normalization by\n> reordering is helpful.  Imagine that a common ancestor had a file\n> with a line with string \"A\" on it (I'll call such a line \"line A\"\n> for brevity in the following) in its early part, and line X in its\n> late part.  And then you fork four branches that do these things:\n> \n>     - AB: changes A to B\n>     - AC: changes A to C\n>     - XY: changes X to Y\n>     - XZ: changes X to Z\n> \n> Now, forking a branch ABAC off of branch AB and then merging AC into\n> it, and forking a branch ACAB off of branch AC and then merging AB\n> into it, would yield the conflict in a different order.  The former\n> would say \"A became B or C, what now?\" while the latter would say \"A\n> became C or B, what now?\"\n> \n> But the act of merging AC into ABAC and resolving the conflict to\n> leave line D means that you declare: \n> \n>     After examining what branches AB and AC did, I believe that\n>     making line A into line D is the best thing to do that is\n>     compatible with what AB and AC wanted to do.\n> \n> So the conflict we would see when merging AB into ACAB should be\n> resolved the same way---it is the resolution that is in line with\n> that declaration.\n> \n> Imagine that similarly you had previously forked branch XYXZ from\n> XY, merged XZ into it, and resolved \"X became Y or Z\" into \"X became\n> W\".\n> \n> Now, if you forked a branch ABXY from AB and then merged XY, then\n> ABXY would have line B in its early part and line Y in its later\n> part.  Such a merge would be quite clean.  We can construct\n> 4 combinations using these four branches ((AB, AC) x (XY, XZ)).\n> \n> Merging ABXY and ACXZ would make \"an early A became B or C, a late X\n> became Y or Z\" conflict, while merging ACXY and ABXZ would make \"an\n> early A became C or B, a late X became Y or Z\".  We can see there\n> are 4 combinations of (\"B or C\", \"C or B\") x (\"X or Y\", \"Y or X\").\n> \n> By sorting, we can give the conflict its canonical name, namely, \"an\n> early part became B or C, a late part becames X or Y\", and whenever\n> any of these four patterns appear, we can get to the same conflict\n> and resolution that we saw earlier.  Without the sorting, we will\n> have to somehow find a previous resolution from combinatorial\n> explosion ;-)\n\nThanks for the in depth explanation!  I'll incorporate this into the\ndocument.\n\n> These days post ec34a8b1 (\"Merge branch 'jc/rerere-multi'\",\n> 2016-05-23), the conflict ID can safely collide, i.e. hash\n> collisions that drops completely different conflicts and their\n> resolutions into the same .git/rr-cache/$id directory will not\n> interfere with proper operation of the system, thanks to that\n> rerere-multi topic that allows us to store multiple preimage\n> conflicts that happens to share the same conflict ID with their\n> corresponding postimage resolutions.\n> \n> In theory, we *should* be able to stub out the SHA-1 computation and\n> give every conflict the same ID and rerere should still operate\n> correctly, even though I haven't tried it yet myself.\n\nI gave this a quick try, and the test suite seems to pass with the\nhash computation giving the same ID to all conflicts.\n"},{"id":"349396","messageId":"20180605215219.28783-1-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180520211210.1248-1-t.gummerer@gmail.com","subject":"[PATCH v2 00/10] rerere: handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:09Z","receivedAt":"2018-06-05T20:47:59Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"The previous round was at\n<20180520211210.1248-1-t.gummerer@gmail.com>.\n\nThanks Junio for the comments on the previous round.\n\nChanges since v2:\n - lowercase the first letter in some error/warning messages before\n   marking them for translation\n - wrap paths in output in single quotes, for consistency, and to make\n   some of the messages the same as ones that are already translated\n - mark messages in builtin/rerere.c for translation as well, which I\n   had previously forgotten.\n - expanded the technical documentation on rerere.  The entire\n   document is basically rewritten.\n - changed the test in 6/10 to just fake a conflict marker inside of\n   one of the hunks instead of using an inner conflict created by a\n   merge.  This is to make sure the codepath is still hit after we\n   handle inner conflicts properly.\n - added tests for handling inner conflict markers\n - added one commit to recalculate the conflict ID when an unresolved\n   conflict is committed, and the subsequent operation conflicts again\n   in the same file.  More explanation in the commit message of that\n   commit.\n\nrange-diff below.  A few commits changed enough for range-diff\nto give up showing the differences in those, they are probably best\nreviewed as the whole patch anyway:\n\n1:  901b638400 ! 1:  2825342cc2 rerere: unify error message when read_cache fails\n    @@ -1,6 +1,6 @@\n     Author: Thomas Gummerer <t.gummerer@gmail.com>\n     \n    -    rerere: unify error message when read_cache fails\n    +    rerere: unify error messages when read_cache fails\n     \n         We have multiple different variants of the error message we show to\n         the user if 'read_cache' fails.  The \"Could not read index\" variant we\n-:  ---------- > 2:  d1500028aa rerere: lowercase error messages\n-:  ---------- > 3:  ed3601ee71 rerere: wrap paths in output in sq\n2:  c48ffededd ! 4:  6ead84a199 rerere: mark strings for translation\n    @@ -9,6 +9,28 @@\n     \n         Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n     \n    +diff --git a/builtin/rerere.c b/builtin/rerere.c\n    +--- a/builtin/rerere.c\n    ++++ b/builtin/rerere.c\n    +@@\n    + \tif (!strcmp(argv[0], \"forget\")) {\n    + \t\tstruct pathspec pathspec;\n    + \t\tif (argc < 2)\n    +-\t\t\twarning(\"'git rerere forget' without paths is deprecated\");\n    ++\t\t\twarning(_(\"'git rerere forget' without paths is deprecated\"));\n    + \t\tparse_pathspec(&pathspec, 0, PATHSPEC_PREFER_CWD,\n    + \t\t\t       prefix, argv + 1);\n    + \t\treturn rerere_forget(&pathspec);\n    +@@\n    + \t\t\tconst char *path = merge_rr.items[i].string;\n    + \t\t\tconst struct rerere_id *id = merge_rr.items[i].util;\n    + \t\t\tif (diff_two(rerere_path(id, \"preimage\"), path, path, path))\n    +-\t\t\t\tdie(\"unable to generate diff for '%s'\", rerere_path(id, NULL));\n    ++\t\t\t\tdie(_(\"unable to generate diff for '%s'\"), rerere_path(id, NULL));\n    + \t\t}\n    + \t} else\n    + \t\tusage_with_options(rerere_usage, options);\n    +\n     diff --git a/rerere.c b/rerere.c\n     --- a/rerere.c\n     +++ b/rerere.c\n    @@ -53,14 +75,14 @@\n      \tio.input = fopen(path, \"r\");\n      \tio.io.wrerror = 0;\n      \tif (!io.input)\n    --\t\treturn error_errno(\"Could not open %s\", path);\n    -+\t\treturn error_errno(_(\"Could not open %s\"), path);\n    +-\t\treturn error_errno(\"could not open '%s'\", path);\n    ++\t\treturn error_errno(_(\"could not open '%s'\"), path);\n      \n      \tif (output) {\n      \t\tio.io.output = fopen(output, \"w\");\n      \t\tif (!io.io.output) {\n    --\t\t\terror_errno(\"Could not write %s\", output);\n    -+\t\t\terror_errno(_(\"Could not write %s\"), output);\n    +-\t\t\terror_errno(\"could not write '%s'\", output);\n    ++\t\t\terror_errno(_(\"could not write '%s'\"), output);\n      \t\t\tfclose(io.input);\n      \t\t\treturn -1;\n      \t\t}\n    @@ -68,18 +90,18 @@\n      \n      \tfclose(io.input);\n      \tif (io.io.wrerror)\n    --\t\terror(\"There were errors while writing %s (%s)\",\n    -+\t\terror(_(\"There were errors while writing %s (%s)\"),\n    +-\t\terror(\"there were errors while writing '%s' (%s)\",\n    ++\t\terror(_(\"there were errors while writing '%s' (%s)\"),\n      \t\t      path, strerror(io.io.wrerror));\n      \tif (io.io.output && fclose(io.io.output))\n    --\t\tio.io.wrerror = error_errno(\"Failed to flush %s\", path);\n    -+\t\tio.io.wrerror = error_errno(_(\"Failed to flush %s\"), path);\n    +-\t\tio.io.wrerror = error_errno(\"failed to flush '%s'\", path);\n    ++\t\tio.io.wrerror = error_errno(_(\"failed to flush '%s'\"), path);\n      \n      \tif (hunk_no < 0) {\n      \t\tif (output)\n      \t\t\tunlink_or_warn(output);\n    --\t\treturn error(\"Could not parse conflict hunks in %s\", path);\n    -+\t\treturn error(_(\"Could not parse conflict hunks in %s\"), path);\n    +-\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n    ++\t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n      \t}\n      \tif (io.io.wrerror)\n      \t\treturn -1;\n    @@ -105,21 +127,21 @@\n      \t * Mark that \"postimage\" was used to help gc.\n      \t */\n      \tif (utime(rerere_path(id, \"postimage\"), NULL) < 0)\n    --\t\twarning_errno(\"failed utime() on %s\",\n    -+\t\twarning_errno(_(\"failed utime() on %s\"),\n    +-\t\twarning_errno(\"failed utime() on '%s'\",\n    ++\t\twarning_errno(_(\"failed utime() on '%s'\"),\n      \t\t\t      rerere_path(id, \"postimage\"));\n      \n      \t/* Update \"path\" with the resolution */\n      \tf = fopen(path, \"w\");\n      \tif (!f)\n    --\t\treturn error_errno(\"Could not open %s\", path);\n    -+\t\treturn error_errno(_(\"Could not open %s\"), path);\n    +-\t\treturn error_errno(\"could not open '%s'\", path);\n    ++\t\treturn error_errno(_(\"could not open '%s'\"), path);\n      \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n    --\t\terror_errno(\"Could not write %s\", path);\n    -+\t\terror_errno(_(\"Could not write %s\"), path);\n    +-\t\terror_errno(\"could not write '%s'\", path);\n    ++\t\terror_errno(_(\"could not write '%s'\"), path);\n      \tif (fclose(f))\n    --\t\treturn error_errno(\"Writing %s failed\", path);\n    -+\t\treturn error_errno(_(\"Writing %s failed\"), path);\n    +-\t\treturn error_errno(\"writing '%s' failed\", path);\n    ++\t\treturn error_errno(_(\"writing '%s' failed\"), path);\n      \n      out:\n      \tfree(cur.ptr);\n    @@ -134,8 +156,8 @@\n      \n      \tif (write_locked_index(&the_index, &index_lock,\n      \t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED))\n    --\t\tdie(\"Unable to write new index file\");\n    -+\t\tdie(_(\"Unable to write new index file\"));\n    +-\t\tdie(\"unable to write new index file\");\n    ++\t\tdie(_(\"unable to write new index file\"));\n      }\n      \n      static void remove_variant(struct rerere_id *id)\n    @@ -179,8 +201,8 @@\n      \t\treturn rr_cache_exists;\n      \n      \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n    --\t\tdie(\"Could not create directory %s\", git_path_rr_cache());\n    -+\t\tdie(_(\"Could not create directory %s\"), git_path_rr_cache());\n    +-\t\tdie(\"could not create directory '%s'\", git_path_rr_cache());\n    ++\t\tdie(_(\"could not create directory '%s'\"), git_path_rr_cache());\n      \treturn 1;\n      }\n      \n    @@ -188,8 +210,8 @@\n      \t */\n      \tret = handle_cache(path, sha1, NULL);\n      \tif (ret < 1)\n    --\t\treturn error(\"Could not parse conflict hunks in '%s'\", path);\n    -+\t\treturn error(_(\"Could not parse conflict hunks in '%s'\"), path);\n    +-\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n    ++\t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n      \n      \t/* Nuke the recorded resolution for the conflict */\n      \tid = new_rerere_id(sha1);\n    @@ -214,11 +236,11 @@\n      \tfilename = rerere_path(id, \"postimage\");\n      \tif (unlink(filename)) {\n      \t\tif (errno == ENOENT)\n    --\t\t\terror(\"no remembered resolution for %s\", path);\n    -+\t\t\terror(_(\"no remembered resolution for %s\"), path);\n    +-\t\t\terror(\"no remembered resolution for '%s'\", path);\n    ++\t\t\terror(_(\"no remembered resolution for '%s'\"), path);\n      \t\telse\n    --\t\t\terror_errno(\"cannot unlink %s\", filename);\n    -+\t\t\terror_errno(_(\"cannot unlink %s\"), filename);\n    +-\t\t\terror_errno(\"cannot unlink '%s'\", filename);\n    ++\t\t\terror_errno(_(\"cannot unlink '%s'\"), filename);\n      \t\tgoto fail_exit;\n      \t}\n      \n    @@ -235,8 +257,8 @@\n      \titem = string_list_insert(rr, path);\n      \tfree_rerere_id(item);\n      \titem->util = id;\n    --\tfprintf(stderr, \"Forgot resolution for %s\\n\", path);\n    -+\tfprintf_ln(stderr, _(\"Forgot resolution for %s\"), path);\n    +-\tfprintf(stderr, \"Forgot resolution for '%s'\\n\", path);\n    ++\tfprintf(stderr, _(\"Forgot resolution for '%s'\\n\"), path);\n      \treturn 0;\n      \n      fail_exit:\n3:  e29449406f < -:  ---------- rerere: add some documentation\n-:  ---------- > 5:  caad276aca rerere: add some documentation\n4:  3b41520b28 ! 6:  ad88a6b8a8 rerere: fix crash when conflict goes unresolved\n    @@ -23,14 +23,18 @@\n         Now when 'rerere clear' for example is run, it will segfault in\n         'has_rerere_resolution', because status is NULL.\n     \n    -    To fix this, remove the rerere ID from the MERGE_RR file in case we\n    -    can't handle it, and remove the folder for the ID.  Removing it\n    -    unconditionally is fine here, because if the user would have resolved\n    -    the conflict and ran rerere, the entry would no longer be in the\n    -    MERGE_RR file, so we wouldn't have this problem in the first place,\n    -    while if the conflict was not resolved, the only thing that's left in\n    -    the folder is the 'preimage', which by itself will be regenerated by\n    -    git if necessary, so the user won't loose any work.\n    +    To fix this, remove the rerere ID from the MERGE_RR file in the case\n    +    when we can't handle it, and remove the corresponding variant from\n    +    .git/rr-cache/.  Removing it unconditionally is fine here, because if\n    +    the user would have resolved the conflict and ran rerere, the entry\n    +    would no longer be in the MERGE_RR file, so we wouldn't have this\n    +    problem in the first place, while if the conflict was not resolved,\n    +    the only thing that's left in the folder is the 'preimage', which by\n    +    itself will be regenerated by git if necessary, so the user won't\n    +    loose any work.\n    +\n    +    Note that other variants that have the same conflict ID will not be\n    +    touched.\n     \n         Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n     \n    @@ -71,16 +75,13 @@\n      \tcount_pre_post 0 0\n      '\n      \n    -+test_expect_success 'rerere with extra conflict markers keeps working' '\n    ++test_expect_success 'rerere with unexpected conflict markers does not crash' '\n     +\tgit reset --hard &&\n     +\n     +\tgit checkout -b branch-1 master &&\n     +\techo \"bar\" >test &&\n     +\tgit add test &&\n     +\tgit commit -q -m two &&\n    -+\techo \"baz\" >test &&\n    -+\tgit add test &&\n    -+\tgit commit -q -m three &&\n     +\n     +\tgit reset --hard &&\n     +\tgit checkout -b branch-2 master &&\n    @@ -88,10 +89,10 @@\n     +\tgit add test &&\n     +\tgit commit -q -a -m one &&\n     +\n    -+\ttest_must_fail git merge branch-1~ &&\n    -+\tgit add test &&\n    -+\tgit commit -q -m \"will solve conflicts later\" &&\n     +\ttest_must_fail git merge branch-1 &&\n    ++\tsed \"s/bar/>>>>>>> a/\" >test.tmp <test &&\n    ++\tmv test.tmp test &&\n    ++\tgit rerere &&\n     +\n     +\tgit rerere clear\n     +'\n5:  411a4ee37e ! 7:  15f9efcba6 rerere: only return whether a path has conflicts or not\n    @@ -67,13 +67,13 @@\n      \tif (io.io.wrerror)\n     @@\n      \tif (io.io.output && fclose(io.io.output))\n    - \t\tio.io.wrerror = error_errno(_(\"Failed to flush %s\"), path);\n    + \t\tio.io.wrerror = error_errno(_(\"failed to flush '%s'\"), path);\n      \n     -\tif (hunk_no < 0) {\n     +\tif (has_conflicts < 0) {\n      \t\tif (output)\n      \t\t\tunlink_or_warn(output);\n    - \t\treturn error(_(\"Could not parse conflict hunks in %s\"), path);\n    + \t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n      \t}\n      \tif (io.io.wrerror)\n      \t\treturn -1;\n6:  fc9f715913 = 8:  1490efaad3 rerere: factor out handle_conflict function\n7:  f7dea09a0a < -:  ---------- rerere: teach rerere to handle nested conflicts\n-:  ---------- > 9:  6619650c42 rerere: teach rerere to handle nested conflicts\n-:  ---------- > 10:  4b11dce7dd rerere: recalculate conflict ID when unresolved conflict is committed\n\nThomas Gummerer (10):\n  rerere: unify error messages when read_cache fails\n  rerere: lowercase error messages\n  rerere: wrap paths in output in sq\n  rerere: mark strings for translation\n  rerere: add some documentation\n  rerere: fix crash when conflict goes unresolved\n  rerere: only return whether a path has conflicts or not\n  rerere: factor out handle_conflict function\n  rerere: teach rerere to handle nested conflicts\n  rerere: recalculate conflict ID when unresolved conflict is committed\n\n Documentation/technical/rerere.txt | 182 +++++++++++++++++++++\n builtin/rerere.c                   |   4 +-\n rerere.c                           | 246 ++++++++++++++---------------\n t/t4200-rerere.sh                  |  67 ++++++++\n 4 files changed, 372 insertions(+), 127 deletions(-)\n create mode 100644 Documentation/technical/rerere.txt\n\n-- \n2.18.0.rc1.242.g61856ae69\n\n"},{"id":"349397","messageId":"20180605215219.28783-2-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 01/10] rerere: unify error messages when read_cache fails","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:10Z","receivedAt":"2018-06-05T20:48:02Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"We have multiple different variants of the error message we show to\nthe user if 'read_cache' fails.  The \"Could not read index\" variant we\nare using in 'rerere.c' is currently not used anywhere in translated\nform.\n\nAs a subsequent commit will mark all output that comes from 'rerere.c'\nfor translation, make the life of the translators a little bit easier\nby using a string that is used elsewhere, and marked for translation\nthere, and thus most likely already translated.\n\n\"index file corrupt\" seems to be the most common error message we show\nwhen 'read_cache' fails, so use that here as well.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 18cae2d11c..4b4869662d 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -568,7 +568,7 @@ static int find_conflict(struct string_list *conflict)\n {\n \tint i;\n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -601,7 +601,7 @@ int rerere_remaining(struct string_list *merge_rr)\n \tif (setup_rerere(merge_rr, RERERE_READONLY))\n \t\treturn 0;\n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -1104,7 +1104,7 @@ int rerere_forget(struct pathspec *pathspec)\n \tstruct string_list merge_rr = STRING_LIST_INIT_DUP;\n \n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfd = setup_rerere(&merge_rr, RERERE_NOAUTOUPDATE);\n \tif (fd < 0)\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"349398","messageId":"20180605215219.28783-4-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 03/10] rerere: wrap paths in output in sq","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:12Z","receivedAt":"2018-06-05T20:48:05Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"It looks like most paths in the output in the git codebase are wrapped\nin single quotes.  Standardize on that in rerere as well.\n\nApart from being more consistent, this also makes some of the strings\nmatch strings that are already translated in other parts of the\ncodebase, thus reducing the work for translators, when the strings are\nmarked for translation in a subsequent commit.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n builtin/rerere.c |  2 +-\n rerere.c         | 26 +++++++++++++-------------\n 2 files changed, 14 insertions(+), 14 deletions(-)\n\ndiff --git a/builtin/rerere.c b/builtin/rerere.c\nindex 0bc40298c2..e0c67c98e9 100644\n--- a/builtin/rerere.c\n+++ b/builtin/rerere.c\n@@ -107,7 +107,7 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \t\t\tconst char *path = merge_rr.items[i].string;\n \t\t\tconst struct rerere_id *id = merge_rr.items[i].util;\n \t\t\tif (diff_two(rerere_path(id, \"preimage\"), path, path, path))\n-\t\t\t\tdie(\"unable to generate diff for %s\", rerere_path(id, NULL));\n+\t\t\t\tdie(\"unable to generate diff for '%s'\", rerere_path(id, NULL));\n \t\t}\n \t} else\n \t\tusage_with_options(rerere_usage, options);\ndiff --git a/rerere.c b/rerere.c\nindex eca182023f..0e5956a51c 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"could not open %s\", path);\n+\t\treturn error_errno(\"could not open '%s'\", path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"could not write %s\", output);\n+\t\t\terror_errno(\"could not write '%s'\", output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"there were errors while writing %s (%s)\",\n+\t\terror(\"there were errors while writing '%s' (%s)\",\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"failed to flush %s\", path);\n+\t\tio.io.wrerror = error_errno(\"failed to flush '%s'\", path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"could not parse conflict hunks in %s\", path);\n+\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -684,17 +684,17 @@ static int merge(const struct rerere_id *id, const char *path)\n \t * Mark that \"postimage\" was used to help gc.\n \t */\n \tif (utime(rerere_path(id, \"postimage\"), NULL) < 0)\n-\t\twarning_errno(\"failed utime() on %s\",\n+\t\twarning_errno(\"failed utime() on '%s'\",\n \t\t\t      rerere_path(id, \"postimage\"));\n \n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"could not open %s\", path);\n+\t\treturn error_errno(\"could not open '%s'\", path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"could not write %s\", path);\n+\t\terror_errno(\"could not write '%s'\", path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"writing %s failed\", path);\n+\t\treturn error_errno(\"writing '%s' failed\", path);\n \n out:\n \tfree(cur.ptr);\n@@ -879,7 +879,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"could not create directory %s\", git_path_rr_cache());\n+\t\tdie(\"could not create directory '%s'\", git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1068,9 +1068,9 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \tfilename = rerere_path(id, \"postimage\");\n \tif (unlink(filename)) {\n \t\tif (errno == ENOENT)\n-\t\t\terror(\"no remembered resolution for %s\", path);\n+\t\t\terror(\"no remembered resolution for '%s'\", path);\n \t\telse\n-\t\t\terror_errno(\"cannot unlink %s\", filename);\n+\t\t\terror_errno(\"cannot unlink '%s'\", filename);\n \t\tgoto fail_exit;\n \t}\n \n@@ -1089,7 +1089,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \titem = string_list_insert(rr, path);\n \tfree_rerere_id(item);\n \titem->util = id;\n-\tfprintf(stderr, \"Forgot resolution for %s\\n\", path);\n+\tfprintf(stderr, \"Forgot resolution for '%s'\\n\", path);\n \treturn 0;\n \n fail_exit:\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"349399","messageId":"20180605215219.28783-6-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 05/10] rerere: add some documentation","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:14Z","receivedAt":"2018-06-05T20:48:08Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add some documentation for the logic behind the conflict normalization\nin rerere.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Documentation/technical/rerere.txt | 142 +++++++++++++++++++++++++++++\n rerere.c                           |   4 -\n 2 files changed, 142 insertions(+), 4 deletions(-)\n create mode 100644 Documentation/technical/rerere.txt\n\ndiff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\nnew file mode 100644\nindex 0000000000..2c517fe0fc\n--- /dev/null\n+++ b/Documentation/technical/rerere.txt\n@@ -0,0 +1,142 @@\n+Rerere\n+======\n+\n+This document describes the rerere logic.\n+\n+Conflict normalization\n+----------------------\n+\n+To ensure recorded conflict resolutions can be looked up in the rerere\n+database, even when branches are merged in a different order,\n+different branches are merged that result in the same conflict, or\n+when different conflict style settings are used, rerere normalizes the\n+conflicts before writing them to the rerere database.\n+\n+Differnt conflict styles and branch names are dealt with by stripping\n+that information from the conflict markers, and removing extraneous\n+information from the `diff3` conflict style.\n+\n+Branches being merged in different order are dealt with by sorting the\n+conflict hunks.  More on each of those parts in the following\n+sections.\n+\n+Once these two normalization operations are applied, a conflict ID is\n+created based on the normalized conflict, which is later used by\n+rerere to look up the conflict in the rerere database.\n+\n+Stripping extraneous information\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+\n+Say we have three branches AB, AC and AC2.  The common ancestor of\n+these branches has a file with with a line with the string \"A\" (for\n+brevity this line is called \"line A\" for brevity in the following) in\n+it.  In branch AB this line is changed to \"B\", in AC, this line is\n+changed to C, and branch AC2 is forked off of AC, after the line was\n+changed to C.\n+\n+Now forking a branch ABAC off of branch AB and then merging AC into it,\n+we'd get a conflict like the following:\n+\n+    <<<<<<< HEAD\n+    B\n+    =======\n+    C\n+    >>>>>>> AC\n+\n+Now doing the analogous with AC2 (forking a branch ABAC2 off of branch\n+AB and then merging branch AC2 into it), maybe using the diff3\n+conflict style, we'd get a conflict like the following:\n+\n+    <<<<<<< HEAD\n+    B\n+    ||||||| merged common ancestors\n+    A\n+    =======\n+    C\n+    >>>>>>> AC2\n+\n+By resolving this conflict, to leave line D, the user declares:\n+\n+    After examining what branches AB and AC did, I believe that making\n+    line A into line D is the best thing to do that is compatible with\n+    what AB and AC wanted to do.\n+\n+As branch AC2 refers to the same commit as AC, the above implies that\n+this is also compatible what AB and AC2 wanted to do.\n+\n+By extension, this means that rerere should recognize that the above\n+conflicts are the same.  To do this, the labels on the conflict\n+markers are stripped, and the diff3 output is removed.  The above\n+examples would both result in the following normalized conflict:\n+\n+    <<<<<<<\n+    B\n+    =======\n+    C\n+    >>>>>>>\n+\n+Sorting hunks\n+~~~~~~~~~~~~~\n+\n+As before, lets imagine that a common ancestor had a file with line A\n+its early part, and line X in its late part.  And then four branches\n+are forked that do these things:\n+\n+    - AB: changes A to B\n+    - AC: changes A to C\n+    - XY: changes X to Y\n+    - XZ: changes X to Z\n+\n+Now, forking a branch ABAC off of branch AB and then merging AC into\n+it, and forking a branch ACAB off of branch AC and then merging AB\n+into it, would yield the conflict in a different order.  The former\n+would say \"A became B or C, what now?\" while the latter would say \"A\n+became C or B, what now?\"\n+\n+As a reminder, the act of merging AC into ABAC and resolving the\n+conflict to leave line D means that the user declares:\n+\n+    After examining what branches AB and AC did, I believe that\n+    making line A into line D is the best thing to do that is\n+    compatible with what AB and AC wanted to do.\n+\n+So the conflict we would see when merging AB into ACAB should be\n+resolved the same way---it is the resolution that is in line with that\n+declaration.\n+\n+Imagine that similarly previously a branch XYXZ was forked from XY,\n+and XZ was merged into it, and resolved \"X became Y or Z\" into \"X\n+became W\".\n+\n+Now, if a branch ABXY was forked from AB and then merged XY, then ABXY\n+would have line B in its early part and line Y in its later part.\n+Such a merge would be quite clean.  We can construct 4 combinations\n+using these four branches ((AB, AC) x (XY, XZ)).\n+\n+Merging ABXY and ACXZ would make \"an early A became B or C, a late X\n+became Y or Z\" conflict, while merging ACXY and ABXZ would make \"an\n+early A became C or B, a late X became Y or Z\".  We can see there are\n+4 combinations of (\"B or C\", \"C or B\") x (\"X or Y\", \"Y or X\").\n+\n+By sorting, the conflict is given its canonical name, namely, \"an\n+early part became B or C, a late part becames X or Y\", and whenever\n+any of these four patterns appear, and we can get to the same conflict\n+and resolution that we saw earlier.\n+\n+Without the sorting, we'd have to somehow find a previous resolution\n+from combinatorial explosion.\n+\n+Conflict ID calculation\n+~~~~~~~~~~~~~~~~~~~~~~~\n+\n+Once the conflict normalization is done, the conflict ID is calculated\n+as the sha1 hash of the conflict hunks appended to each other,\n+separated by <NUL> characters.  The conflict markers are stripped out\n+before the sha1 is calculated.  So in the example above, where we\n+merge branch AC which changes line A to line C, into branch AB, which\n+changes line A to line C, the conflict ID would be\n+SHA1('B<NUL>C<NUL>').\n+\n+If there are multiple conflicts in one file, the sha1 is calculated\n+the same way with all hunks appended to each other, in the order in\n+which they appear in the file, separated by a <NUL> character.\ndiff --git a/rerere.c b/rerere.c\nindex 74ce422634..ef23abe4dd 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -394,10 +394,6 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n  * and NUL concatenated together.\n  *\n  * Return the number of conflict hunks found.\n- *\n- * NEEDSWORK: the logic and theory of operation behind this conflict\n- * normalization may deserve to be documented somewhere, perhaps in\n- * Documentation/technical/rerere.txt.\n  */\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"349400","messageId":"20180605215219.28783-8-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 07/10] rerere: only return whether a path has conflicts or not","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:16Z","receivedAt":"2018-06-05T20:48:11Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"We currently return the exact number of conflict hunks a certain path\nhas from the 'handle_paths' function.  However all of its callers only\ncare whether there are conflicts or not or if there is an error.\nReturn only that information, and document that only that information\nis returned.  This will simplify the code in the subsequent steps.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 23 ++++++++++++-----------\n 1 file changed, 12 insertions(+), 11 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 220020187b..da3744b86b 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -393,12 +393,13 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n  * one side of the conflict, NUL, the other side of the conflict,\n  * and NUL concatenated together.\n  *\n- * Return the number of conflict hunks found.\n+ * Return 1 if conflict hunks are found, 0 if there are no conflict\n+ * hunks and -1 if an error occured.\n  */\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n \tgit_SHA_CTX ctx;\n-\tint hunk_no = 0;\n+\tint has_conflicts = 0;\n \tenum {\n \t\tRR_CONTEXT = 0, RR_SIDE_1, RR_SIDE_2, RR_ORIGINAL\n \t} hunk = RR_CONTEXT;\n@@ -426,7 +427,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\t\t\tgoto bad;\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n-\t\t\thunk_no++;\n+\t\t\thas_conflicts = 1;\n \t\t\thunk = RR_CONTEXT;\n \t\t\trerere_io_putconflict('<', marker_size, io);\n \t\t\trerere_io_putmem(one.buf, one.len, io);\n@@ -462,7 +463,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\tgit_SHA1_Final(sha1, &ctx);\n \tif (hunk != RR_CONTEXT)\n \t\treturn -1;\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n /*\n@@ -471,7 +472,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n  */\n static int handle_file(const char *path, unsigned char *sha1, const char *output)\n {\n-\tint hunk_no = 0;\n+\tint has_conflicts = 0;\n \tstruct rerere_io_file io;\n \tint marker_size = ll_merge_marker_size(path);\n \n@@ -491,7 +492,7 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \t\t}\n \t}\n \n-\thunk_no = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n+\thas_conflicts = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n@@ -500,14 +501,14 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tif (io.io.output && fclose(io.io.output))\n \t\tio.io.wrerror = error_errno(_(\"failed to flush '%s'\"), path);\n \n-\tif (hunk_no < 0) {\n+\tif (has_conflicts < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n \t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n /*\n@@ -955,7 +956,7 @@ static int handle_cache(const char *path, unsigned char *sha1, const char *outpu\n \tmmfile_t mmfile[3] = {{NULL}};\n \tmmbuffer_t result = {NULL, 0};\n \tconst struct cache_entry *ce;\n-\tint pos, len, i, hunk_no;\n+\tint pos, len, i, has_conflicts;\n \tstruct rerere_io_mem io;\n \tint marker_size = ll_merge_marker_size(path);\n \n@@ -1009,11 +1010,11 @@ static int handle_cache(const char *path, unsigned char *sha1, const char *outpu\n \t * Grab the conflict ID and optionally write the original\n \t * contents with conflict markers out.\n \t */\n-\thunk_no = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n+\thas_conflicts = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n \tstrbuf_release(&io.input);\n \tif (io.io.output)\n \t\tfclose(io.io.output);\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n static int rerere_forget_one_path(const char *path, struct string_list *rr)\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"349401","messageId":"20180605215219.28783-10-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 09/10] rerere: teach rerere to handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:18Z","receivedAt":"2018-06-05T20:48:15Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently rerere can't handle nested conflicts and will error out when\nit encounters such conflicts.  Do that by recursively calling the\n'handle_conflict' function to normalize the conflict.\n\nThe conflict ID calculation here deserves some explanation:\n\nAs we are using the same handle_conflict function, the nested conflict\nis normalized the same way as for non-nested conflicts, which means\nthe ancestor in the diff3 case is stripped out, and the parts of the\nconflict are ordered alphabetically.\n\nThe conflict ID is however is only calculated in the top level\nhandle_conflict call, so it will include the markers that 'rerere'\nadds to the output.  e.g. say there's the following conflict:\n\n    <<<<<<< HEAD\n    1\n    =======\n    <<<<<<< HEAD\n    3\n    =======\n    2\n    >>>>>>> branch-2\n    >>>>>>> branch-3~\n\nit would be reordered as follows in the preimage:\n\n    <<<<<<<\n    1\n    =======\n    <<<<<<<\n    2\n    =======\n    3\n    >>>>>>>\n    >>>>>>>\n\nand the conflict ID would be calculated as\n    sha1(1<NUL><<<<<<<\n    2\n    =======\n    3\n    >>>>>>><NUL>)\n\nStripping out vs. leaving the conflict markers in place should have no\npractical impact, but it simplifies the implementation.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n\nI couldn't actually get the conflict markers the right way just using\nmerge-recursive.  But I think that would be fixed either way by\nd694a17986 (\"ll-merge: use a longer conflict marker for internal\nmerge\", 2016-04-14), if I read that correctly.\n\nEither way I still think this can be an improvement for when the user\ncommits merge conflicts (even though they shouldn't do that in the\nfirst place), and for possible other edge cases that I'm not able to\nproduce right now, but I may just not be creative enough for those.\n\n Documentation/technical/rerere.txt | 40 ++++++++++++++++++++++++++++++\n rerere.c                           | 14 ++++++++---\n t/t4200-rerere.sh                  | 38 ++++++++++++++++++++++++++++\n 3 files changed, 88 insertions(+), 4 deletions(-)\n\ndiff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\nindex 2c517fe0fc..7077ab4a08 100644\n--- a/Documentation/technical/rerere.txt\n+++ b/Documentation/technical/rerere.txt\n@@ -140,3 +140,43 @@ SHA1('B<NUL>C<NUL>').\n If there are multiple conflicts in one file, the sha1 is calculated\n the same way with all hunks appended to each other, in the order in\n which they appear in the file, separated by a <NUL> character.\n+\n+Nested conflicts\n+~~~~~~~~~~~~~~~~\n+\n+Nested conflicts are handled very similarly to \"simple\" conflicts.\n+Same as before, labels on conflict markers and diff3 output is\n+stripped, and the conflict hunks are sorted, for both the outer and\n+the inner conflict.\n+\n+The only difference is in how the conflict ID is calculated.  For the\n+inner conflict, the conflict markers themselves are not stripped out\n+before calculating the sha1.\n+\n+Say we have the following conflict for example:\n+\n+    <<<<<<< HEAD\n+    1\n+    =======\n+    <<<<<<< HEAD\n+    3\n+    =======\n+    2\n+    >>>>>>> branch-2\n+    >>>>>>> branch-3~\n+\n+After stripping out the labels of the conflict markers, the conflict\n+would look as follows:\n+\n+    <<<<<<<\n+    1\n+    =======\n+    <<<<<<<\n+    3\n+    =======\n+    2\n+    >>>>>>>\n+    >>>>>>>\n+\n+and finally the conflict ID would be calculated as:\n+`sha1('1<NUL><<<<<<<\\n3\\n=======\\n2\\n>>>>>>><NUL>')`\ndiff --git a/rerere.c b/rerere.c\nindex fac90663b0..f611db7873 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -365,12 +365,18 @@ static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n \t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n \t} hunk = RR_SIDE_1;\n \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n-\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf buf = STRBUF_INIT, conflict = STRBUF_INIT;\n \tint has_conflicts = 1;\n \twhile (!io->getline(&buf, io)) {\n-\t\tif (is_cmarker(buf.buf, '<', marker_size))\n-\t\t\tgoto bad;\n-\t\telse if (is_cmarker(buf.buf, '|', marker_size)) {\n+\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n+\t\t\tif (handle_conflict(&conflict, io, marker_size, NULL) < 0)\n+\t\t\t\tgoto bad;\n+\t\t\tif (hunk == RR_SIDE_1)\n+\t\t\t\tstrbuf_addbuf(&one, &conflict);\n+\t\t\telse\n+\t\t\t\tstrbuf_addbuf(&two, &conflict);\n+\t\t\tstrbuf_release(&conflict);\n+\t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1)\n \t\t\t\tgoto bad;\n \t\t\thunk = RR_ORIGINAL;\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex 5ce411b70d..f433848ccb 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -602,4 +602,42 @@ test_expect_success 'rerere with unexpected conflict markers does not crash' '\n \tgit rerere clear\n '\n \n+test_expect_success 'rerere with inner conflict markers' '\n+\tgit reset --hard &&\n+\n+\tgit checkout -b A master &&\n+\techo \"bar\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m two &&\n+\techo \"baz\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m three &&\n+\n+\tgit reset --hard &&\n+\tgit checkout -b B master &&\n+\techo \"foo\" >test &&\n+\tgit add test &&\n+\tgit commit -q -a -m one &&\n+\n+\ttest_must_fail git merge A~ &&\n+\tgit add test &&\n+\tgit commit -q -m \"will solve conflicts later\" &&\n+\ttest_must_fail git merge A &&\n+\n+\techo \"resolved\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m \"solved conflict\" &&\n+\n+\techo \"resolved\" >expect &&\n+\n+\tgit reset --hard HEAD~~ &&\n+\ttest_must_fail git merge A~ &&\n+\tgit add test &&\n+\tgit commit -q -m \"will solve conflicts later\" &&\n+\ttest_must_fail git merge A &&\n+\tcat test >actual &&\n+\ttest_cmp expect actual\n+'\n+\n+\n test_done\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"349402","messageId":"20180605215219.28783-11-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 10/10] rerere: recalculate conflict ID when unresolved conflict is committed","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:19Z","receivedAt":"2018-06-05T20:48:16Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently when a user doesn't resolve a conflict, commits the results,\nand does an operation which creates another conflict, rerere will use\nthe ID of the previously unresolved conflict for the new conflict.\nThis is because the conflict is kept in the MERGE_RR file, which\n'rerere' reads every time it is invoked.\n\nAfter the new conflict is solved, rerere will record the resolution\nwith the ID of the old conflict.  So in order to replay the conflict,\nboth merges would have to be re-done, instead of just the last one, in\norder for rerere to be able to automatically resolve the conflict.\n\nInstead of that, assign a new conflict ID if there are still conflicts\nin a file and the file had conflicts at a previous step.  This ID\nmatches the conflict we actually resolved at the corresponding step.\n\nNote that there are no backwards compatibility worries here, as rerere\nwould have failed to even normalize the conflict before this patch\nseries.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c          | 7 +++----\n t/t4200-rerere.sh | 7 +++++++\n 2 files changed, 10 insertions(+), 4 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex f611db7873..644f185180 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -818,7 +818,7 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\tstruct rerere_id *id;\n \t\tunsigned char sha1[20];\n \t\tconst char *path = conflict.items[i].string;\n-\t\tint ret, has_string;\n+\t\tint ret;\n \n \t\t/*\n \t\t * Ask handle_file() to scan and assign a\n@@ -826,12 +826,11 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\t * yet.\n \t\t */\n \t\tret = handle_file(path, sha1, NULL);\n-\t\thas_string = string_list_has_string(rr, path);\n-\t\tif (ret < 0 && has_string) {\n+\t\tif (ret != 0 && string_list_has_string(rr, path)) {\n \t\t\tremove_variant(string_list_lookup(rr, path)->util);\n \t\t\tstring_list_remove(rr, path, 1);\n \t\t}\n-\t\tif (ret < 1 || has_string)\n+\t\tif (ret < 1)\n \t\t\tcontinue;\n \n \t\tid = new_rerere_id(sha1);\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex f433848ccb..9578215ff2 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -636,6 +636,13 @@ test_expect_success 'rerere with inner conflict markers' '\n \tgit commit -q -m \"will solve conflicts later\" &&\n \ttest_must_fail git merge A &&\n \tcat test >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tgit add test &&\n+\tgit commit -m \"rerere solved conflict\" &&\n+\tgit reset --hard HEAD~ &&\n+\ttest_must_fail git merge A &&\n+\tcat test >actual &&\n \ttest_cmp expect actual\n '\n \n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"349403","messageId":"20180605215219.28783-3-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 02/10] rerere: lowercase error messages","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:11Z","receivedAt":"2018-06-05T20:48:20Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Documentation/CodingGuidelines mentions that error messages should be\nlowercase.  Prior to marking them for translation follow that pattern\nin rerere as well.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 22 +++++++++++-----------\n 1 file changed, 11 insertions(+), 11 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 4b4869662d..eca182023f 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"Could not open %s\", path);\n+\t\treturn error_errno(\"could not open %s\", path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"Could not write %s\", output);\n+\t\t\terror_errno(\"could not write %s\", output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"There were errors while writing %s (%s)\",\n+\t\terror(\"there were errors while writing %s (%s)\",\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"Failed to flush %s\", path);\n+\t\tio.io.wrerror = error_errno(\"failed to flush %s\", path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"Could not parse conflict hunks in %s\", path);\n+\t\treturn error(\"could not parse conflict hunks in %s\", path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -690,11 +690,11 @@ static int merge(const struct rerere_id *id, const char *path)\n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"Could not open %s\", path);\n+\t\treturn error_errno(\"could not open %s\", path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"Could not write %s\", path);\n+\t\terror_errno(\"could not write %s\", path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"Writing %s failed\", path);\n+\t\treturn error_errno(\"writing %s failed\", path);\n \n out:\n \tfree(cur.ptr);\n@@ -721,7 +721,7 @@ static void update_paths(struct string_list *update)\n \n \tif (write_locked_index(&the_index, &index_lock,\n \t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED))\n-\t\tdie(\"Unable to write new index file\");\n+\t\tdie(\"unable to write new index file\");\n }\n \n static void remove_variant(struct rerere_id *id)\n@@ -879,7 +879,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"Could not create directory %s\", git_path_rr_cache());\n+\t\tdie(\"could not create directory %s\", git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1032,7 +1032,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t */\n \tret = handle_cache(path, sha1, NULL);\n \tif (ret < 1)\n-\t\treturn error(\"Could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n \n \t/* Nuke the recorded resolution for the conflict */\n \tid = new_rerere_id(sha1);\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"349404","messageId":"20180605215219.28783-7-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 06/10] rerere: fix crash when conflict goes unresolved","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:15Z","receivedAt":"2018-06-05T20:48:21Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently when a user doesn't resolve a conflict in a file, but\ncommits the file with the conflict markers, and later the file ends up\nin a state in which rerere can't handle it, subsequent rerere\noperations that are interested in that path, such as 'rerere clear' or\n'rerere forget <path>' will fail, or even worse in the case of 'rerere\nclear' segfault.\n\nSuch states include nested conflicts, or an extra conflict marker that\ndoesn't have any match.\n\nThis is because the first 'git rerere' when there was only one\nconflict in the file leaves an entry in the MERGE_RR file behind.  The\nnext 'git rerere' will then pick the rerere ID for that file up, and\nnot assign a new ID as it can't successfully calculate one.  It will\nhowever still try to do the rerere operation, because of the existing\nID.  As the handle_file function fails, it will remove the 'preimage'\nfor the ID in the process, while leaving the ID in the MERGE_RR file.\n\nNow when 'rerere clear' for example is run, it will segfault in\n'has_rerere_resolution', because status is NULL.\n\nTo fix this, remove the rerere ID from the MERGE_RR file in the case\nwhen we can't handle it, and remove the corresponding variant from\n.git/rr-cache/.  Removing it unconditionally is fine here, because if\nthe user would have resolved the conflict and ran rerere, the entry\nwould no longer be in the MERGE_RR file, so we wouldn't have this\nproblem in the first place, while if the conflict was not resolved,\nthe only thing that's left in the folder is the 'preimage', which by\nitself will be regenerated by git if necessary, so the user won't\nloose any work.\n\nNote that other variants that have the same conflict ID will not be\ntouched.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c          | 12 +++++++-----\n t/t4200-rerere.sh | 22 ++++++++++++++++++++++\n 2 files changed, 29 insertions(+), 5 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex ef23abe4dd..220020187b 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -824,10 +824,7 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\tstruct rerere_id *id;\n \t\tunsigned char sha1[20];\n \t\tconst char *path = conflict.items[i].string;\n-\t\tint ret;\n-\n-\t\tif (string_list_has_string(rr, path))\n-\t\t\tcontinue;\n+\t\tint ret, has_string;\n \n \t\t/*\n \t\t * Ask handle_file() to scan and assign a\n@@ -835,7 +832,12 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\t * yet.\n \t\t */\n \t\tret = handle_file(path, sha1, NULL);\n-\t\tif (ret < 1)\n+\t\thas_string = string_list_has_string(rr, path);\n+\t\tif (ret < 0 && has_string) {\n+\t\t\tremove_variant(string_list_lookup(rr, path)->util);\n+\t\t\tstring_list_remove(rr, path, 1);\n+\t\t}\n+\t\tif (ret < 1 || has_string)\n \t\t\tcontinue;\n \n \t\tid = new_rerere_id(sha1);\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex eaf18c81cb..5ce411b70d 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -580,4 +580,26 @@ test_expect_success 'multiple identical conflicts' '\n \tcount_pre_post 0 0\n '\n \n+test_expect_success 'rerere with unexpected conflict markers does not crash' '\n+\tgit reset --hard &&\n+\n+\tgit checkout -b branch-1 master &&\n+\techo \"bar\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m two &&\n+\n+\tgit reset --hard &&\n+\tgit checkout -b branch-2 master &&\n+\techo \"foo\" >test &&\n+\tgit add test &&\n+\tgit commit -q -a -m one &&\n+\n+\ttest_must_fail git merge branch-1 &&\n+\tsed \"s/bar/>>>>>>> a/\" >test.tmp <test &&\n+\tmv test.tmp test &&\n+\tgit rerere &&\n+\n+\tgit rerere clear\n+'\n+\n test_done\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"349405","messageId":"20180605215219.28783-5-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 04/10] rerere: mark strings for translation","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:13Z","receivedAt":"2018-06-05T20:48:22Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"'git rerere' is considered a plumbing command and as such its output\nshould be translated.  Its functionality is also only enabled through\na config setting, so scripts really shouldn't rely on its output\neither way.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n builtin/rerere.c |  4 +--\n rerere.c         | 68 ++++++++++++++++++++++++------------------------\n 2 files changed, 36 insertions(+), 36 deletions(-)\n\ndiff --git a/builtin/rerere.c b/builtin/rerere.c\nindex e0c67c98e9..5ed941b91f 100644\n--- a/builtin/rerere.c\n+++ b/builtin/rerere.c\n@@ -75,7 +75,7 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \tif (!strcmp(argv[0], \"forget\")) {\n \t\tstruct pathspec pathspec;\n \t\tif (argc < 2)\n-\t\t\twarning(\"'git rerere forget' without paths is deprecated\");\n+\t\t\twarning(_(\"'git rerere forget' without paths is deprecated\"));\n \t\tparse_pathspec(&pathspec, 0, PATHSPEC_PREFER_CWD,\n \t\t\t       prefix, argv + 1);\n \t\treturn rerere_forget(&pathspec);\n@@ -107,7 +107,7 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \t\t\tconst char *path = merge_rr.items[i].string;\n \t\t\tconst struct rerere_id *id = merge_rr.items[i].util;\n \t\t\tif (diff_two(rerere_path(id, \"preimage\"), path, path, path))\n-\t\t\t\tdie(\"unable to generate diff for '%s'\", rerere_path(id, NULL));\n+\t\t\t\tdie(_(\"unable to generate diff for '%s'\"), rerere_path(id, NULL));\n \t\t}\n \t} else\n \t\tusage_with_options(rerere_usage, options);\ndiff --git a/rerere.c b/rerere.c\nindex 0e5956a51c..74ce422634 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -212,7 +212,7 @@ static void read_rr(struct string_list *rr)\n \n \t\t/* There has to be the hash, tab, path and then NUL */\n \t\tif (buf.len < 42 || get_sha1_hex(buf.buf, sha1))\n-\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \n \t\tif (buf.buf[40] != '.') {\n \t\t\tvariant = 0;\n@@ -221,10 +221,10 @@ static void read_rr(struct string_list *rr)\n \t\t\terrno = 0;\n \t\t\tvariant = strtol(buf.buf + 41, &path, 10);\n \t\t\tif (errno)\n-\t\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \t\t}\n \t\tif (*(path++) != '\\t')\n-\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \t\tbuf.buf[40] = '\\0';\n \t\tid = new_rerere_id_hex(buf.buf);\n \t\tid->variant = variant;\n@@ -259,12 +259,12 @@ static int write_rr(struct string_list *rr, int out_fd)\n \t\t\t\t    rr->items[i].string, 0);\n \n \t\tif (write_in_full(out_fd, buf.buf, buf.len) < 0)\n-\t\t\tdie(\"unable to write rerere record\");\n+\t\t\tdie(_(\"unable to write rerere record\"));\n \n \t\tstrbuf_release(&buf);\n \t}\n \tif (commit_lock_file(&write_lock) != 0)\n-\t\tdie(\"unable to write rerere record\");\n+\t\tdie(_(\"unable to write rerere record\"));\n \treturn 0;\n }\n \n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"could not open '%s'\", path);\n+\t\treturn error_errno(_(\"could not open '%s'\"), path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"could not write '%s'\", output);\n+\t\t\terror_errno(_(\"could not write '%s'\"), output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"there were errors while writing '%s' (%s)\",\n+\t\terror(_(\"there were errors while writing '%s' (%s)\"),\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"failed to flush '%s'\", path);\n+\t\tio.io.wrerror = error_errno(_(\"failed to flush '%s'\"), path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -568,7 +568,7 @@ static int find_conflict(struct string_list *conflict)\n {\n \tint i;\n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -601,7 +601,7 @@ int rerere_remaining(struct string_list *merge_rr)\n \tif (setup_rerere(merge_rr, RERERE_READONLY))\n \t\treturn 0;\n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -684,17 +684,17 @@ static int merge(const struct rerere_id *id, const char *path)\n \t * Mark that \"postimage\" was used to help gc.\n \t */\n \tif (utime(rerere_path(id, \"postimage\"), NULL) < 0)\n-\t\twarning_errno(\"failed utime() on '%s'\",\n+\t\twarning_errno(_(\"failed utime() on '%s'\"),\n \t\t\t      rerere_path(id, \"postimage\"));\n \n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"could not open '%s'\", path);\n+\t\treturn error_errno(_(\"could not open '%s'\"), path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"could not write '%s'\", path);\n+\t\terror_errno(_(\"could not write '%s'\"), path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"writing '%s' failed\", path);\n+\t\treturn error_errno(_(\"writing '%s' failed\"), path);\n \n out:\n \tfree(cur.ptr);\n@@ -715,13 +715,13 @@ static void update_paths(struct string_list *update)\n \t\tstruct string_list_item *item = &update->items[i];\n \t\tif (add_file_to_cache(item->string, 0))\n \t\t\texit(128);\n-\t\tfprintf(stderr, \"Staged '%s' using previous resolution.\\n\",\n+\t\tfprintf_ln(stderr, _(\"Staged '%s' using previous resolution.\"),\n \t\t\titem->string);\n \t}\n \n \tif (write_locked_index(&the_index, &index_lock,\n \t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED))\n-\t\tdie(\"unable to write new index file\");\n+\t\tdie(_(\"unable to write new index file\"));\n }\n \n static void remove_variant(struct rerere_id *id)\n@@ -753,7 +753,7 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \t\tif (!handle_file(path, NULL, NULL)) {\n \t\t\tcopy_file(rerere_path(id, \"postimage\"), path, 0666);\n \t\t\tid->collection->status[variant] |= RR_HAS_POSTIMAGE;\n-\t\t\tfprintf(stderr, \"Recorded resolution for '%s'.\\n\", path);\n+\t\t\tfprintf_ln(stderr, _(\"Recorded resolution for '%s'.\"), path);\n \t\t\tfree_rerere_id(rr_item);\n \t\t\trr_item->util = NULL;\n \t\t\treturn;\n@@ -787,9 +787,9 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \t\tif (rerere_autoupdate)\n \t\t\tstring_list_insert(update, path);\n \t\telse\n-\t\t\tfprintf(stderr,\n-\t\t\t\t\"Resolved '%s' using previous resolution.\\n\",\n-\t\t\t\tpath);\n+\t\t\tfprintf_ln(stderr,\n+\t\t\t\t   _(\"Resolved '%s' using previous resolution.\"),\n+\t\t\t\t   path);\n \t\tfree_rerere_id(rr_item);\n \t\trr_item->util = NULL;\n \t\treturn;\n@@ -803,11 +803,11 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \tif (id->collection->status[variant] & RR_HAS_POSTIMAGE) {\n \t\tconst char *path = rerere_path(id, \"postimage\");\n \t\tif (unlink(path))\n-\t\t\tdie_errno(\"cannot unlink stray '%s'\", path);\n+\t\t\tdie_errno(_(\"cannot unlink stray '%s'\"), path);\n \t\tid->collection->status[variant] &= ~RR_HAS_POSTIMAGE;\n \t}\n \tid->collection->status[variant] |= RR_HAS_PREIMAGE;\n-\tfprintf(stderr, \"Recorded preimage for '%s'\\n\", path);\n+\tfprintf_ln(stderr, _(\"Recorded preimage for '%s'\"), path);\n }\n \n static int do_plain_rerere(struct string_list *rr, int fd)\n@@ -879,7 +879,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"could not create directory '%s'\", git_path_rr_cache());\n+\t\tdie(_(\"could not create directory '%s'\"), git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1032,7 +1032,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t */\n \tret = handle_cache(path, sha1, NULL);\n \tif (ret < 1)\n-\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n \n \t/* Nuke the recorded resolution for the conflict */\n \tid = new_rerere_id(sha1);\n@@ -1050,7 +1050,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t\thandle_cache(path, sha1, rerere_path(id, \"thisimage\"));\n \t\tif (read_mmfile(&cur, rerere_path(id, \"thisimage\"))) {\n \t\t\tfree(cur.ptr);\n-\t\t\terror(\"Failed to update conflicted state in '%s'\", path);\n+\t\t\terror(_(\"Failed to update conflicted state in '%s'\"), path);\n \t\t\tgoto fail_exit;\n \t\t}\n \t\tcleanly_resolved = !try_merge(id, path, &cur, &result);\n@@ -1061,16 +1061,16 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t}\n \n \tif (id->collection->status_nr <= id->variant) {\n-\t\terror(\"no remembered resolution for '%s'\", path);\n+\t\terror(_(\"no remembered resolution for '%s'\"), path);\n \t\tgoto fail_exit;\n \t}\n \n \tfilename = rerere_path(id, \"postimage\");\n \tif (unlink(filename)) {\n \t\tif (errno == ENOENT)\n-\t\t\terror(\"no remembered resolution for '%s'\", path);\n+\t\t\terror(_(\"no remembered resolution for '%s'\"), path);\n \t\telse\n-\t\t\terror_errno(\"cannot unlink '%s'\", filename);\n+\t\t\terror_errno(_(\"cannot unlink '%s'\"), filename);\n \t\tgoto fail_exit;\n \t}\n \n@@ -1080,7 +1080,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t * the postimage.\n \t */\n \thandle_cache(path, sha1, rerere_path(id, \"preimage\"));\n-\tfprintf(stderr, \"Updated preimage for '%s'\\n\", path);\n+\tfprintf_ln(stderr, _(\"Updated preimage for '%s'\"), path);\n \n \t/*\n \t * And remember that we can record resolution for this\n@@ -1089,7 +1089,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \titem = string_list_insert(rr, path);\n \tfree_rerere_id(item);\n \titem->util = id;\n-\tfprintf(stderr, \"Forgot resolution for '%s'\\n\", path);\n+\tfprintf(stderr, _(\"Forgot resolution for '%s'\\n\"), path);\n \treturn 0;\n \n fail_exit:\n@@ -1104,7 +1104,7 @@ int rerere_forget(struct pathspec *pathspec)\n \tstruct string_list merge_rr = STRING_LIST_INIT_DUP;\n \n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfd = setup_rerere(&merge_rr, RERERE_NOAUTOUPDATE);\n \tif (fd < 0)\n@@ -1192,7 +1192,7 @@ void rerere_gc(struct string_list *rr)\n \tgit_config(git_default_config, NULL);\n \tdir = opendir(git_path(\"rr-cache\"));\n \tif (!dir)\n-\t\tdie_errno(\"unable to open rr-cache directory\");\n+\t\tdie_errno(_(\"unable to open rr-cache directory\"));\n \t/* Collect stale conflict IDs ... */\n \twhile ((e = readdir(dir))) {\n \t\tstruct rerere_dir *rr_dir;\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"349406","messageId":"20180605215219.28783-9-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v2 08/10] rerere: factor out handle_conflict function","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-06-05T21:52:17Z","receivedAt":"2018-06-05T20:48:26Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Factor out the handle_conflict function, which handles a single\nconflict in a path.  This is a preparation for the next step, where\nthis function will be re-used.  No functional changes intended.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 143 +++++++++++++++++++++++++------------------------------\n 1 file changed, 65 insertions(+), 78 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex da3744b86b..fac90663b0 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -302,38 +302,6 @@ static void rerere_io_putstr(const char *str, struct rerere_io *io)\n \t\tferr_puts(str, io->output, &io->wrerror);\n }\n \n-/*\n- * Write a conflict marker to io->output (if defined).\n- */\n-static void rerere_io_putconflict(int ch, int size, struct rerere_io *io)\n-{\n-\tchar buf[64];\n-\n-\twhile (size) {\n-\t\tif (size <= sizeof(buf) - 2) {\n-\t\t\tmemset(buf, ch, size);\n-\t\t\tbuf[size] = '\\n';\n-\t\t\tbuf[size + 1] = '\\0';\n-\t\t\tsize = 0;\n-\t\t} else {\n-\t\t\tint sz = sizeof(buf) - 1;\n-\n-\t\t\t/*\n-\t\t\t * Make sure we will not write everything out\n-\t\t\t * in this round by leaving at least 1 byte\n-\t\t\t * for the next round, giving the next round\n-\t\t\t * a chance to add the terminating LF.  Yuck.\n-\t\t\t */\n-\t\t\tif (size <= sz)\n-\t\t\t\tsz -= (sz - size) + 1;\n-\t\t\tmemset(buf, ch, sz);\n-\t\t\tbuf[sz] = '\\0';\n-\t\t\tsize -= sz;\n-\t\t}\n-\t\trerere_io_putstr(buf, io);\n-\t}\n-}\n-\n static void rerere_io_putmem(const char *mem, size_t sz, struct rerere_io *io)\n {\n \tif (io->output)\n@@ -384,37 +352,25 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n \treturn isspace(*buf);\n }\n \n-/*\n- * Read contents a file with conflicts, normalize the conflicts\n- * by (1) discarding the common ancestor version in diff3-style,\n- * (2) reordering our side and their side so that whichever sorts\n- * alphabetically earlier comes before the other one, while\n- * computing the \"conflict ID\", which is just an SHA-1 hash of\n- * one side of the conflict, NUL, the other side of the conflict,\n- * and NUL concatenated together.\n- *\n- * Return 1 if conflict hunks are found, 0 if there are no conflict\n- * hunks and -1 if an error occured.\n- */\n-static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n+static void rerere_strbuf_putconflict(struct strbuf *buf, int ch, size_t size)\n+{\n+\tstrbuf_addchars(buf, ch, size);\n+\tstrbuf_addch(buf, '\\n');\n+}\n+\n+static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n+\t\t\t   int marker_size, git_SHA_CTX *ctx)\n {\n-\tgit_SHA_CTX ctx;\n-\tint has_conflicts = 0;\n \tenum {\n-\t\tRR_CONTEXT = 0, RR_SIDE_1, RR_SIDE_2, RR_ORIGINAL\n-\t} hunk = RR_CONTEXT;\n+\t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n+\t} hunk = RR_SIDE_1;\n \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n \tstruct strbuf buf = STRBUF_INIT;\n-\n-\tif (sha1)\n-\t\tgit_SHA1_Init(&ctx);\n-\n+\tint has_conflicts = 1;\n \twhile (!io->getline(&buf, io)) {\n-\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n-\t\t\tif (hunk != RR_CONTEXT)\n-\t\t\t\tgoto bad;\n-\t\t\thunk = RR_SIDE_1;\n-\t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n+\t\tif (is_cmarker(buf.buf, '<', marker_size))\n+\t\t\tgoto bad;\n+\t\telse if (is_cmarker(buf.buf, '|', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1)\n \t\t\t\tgoto bad;\n \t\t\thunk = RR_ORIGINAL;\n@@ -427,42 +383,73 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\t\t\tgoto bad;\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n-\t\t\thas_conflicts = 1;\n-\t\t\thunk = RR_CONTEXT;\n-\t\t\trerere_io_putconflict('<', marker_size, io);\n-\t\t\trerere_io_putmem(one.buf, one.len, io);\n-\t\t\trerere_io_putconflict('=', marker_size, io);\n-\t\t\trerere_io_putmem(two.buf, two.len, io);\n-\t\t\trerere_io_putconflict('>', marker_size, io);\n-\t\t\tif (sha1) {\n-\t\t\t\tgit_SHA1_Update(&ctx, one.buf ? one.buf : \"\",\n+\t\t\trerere_strbuf_putconflict(out, '<', marker_size);\n+\t\t\tstrbuf_addbuf(out, &one);\n+\t\t\trerere_strbuf_putconflict(out, '=', marker_size);\n+\t\t\tstrbuf_addbuf(out, &two);\n+\t\t\trerere_strbuf_putconflict(out, '>', marker_size);\n+\t\t\tif (ctx) {\n+\t\t\t\tgit_SHA1_Update(ctx, one.buf ? one.buf : \"\",\n \t\t\t\t\t    one.len + 1);\n-\t\t\t\tgit_SHA1_Update(&ctx, two.buf ? two.buf : \"\",\n+\t\t\t\tgit_SHA1_Update(ctx, two.buf ? two.buf : \"\",\n \t\t\t\t\t    two.len + 1);\n \t\t\t}\n-\t\t\tstrbuf_reset(&one);\n-\t\t\tstrbuf_reset(&two);\n+\t\t\tgoto out;\n \t\t} else if (hunk == RR_SIDE_1)\n \t\t\tstrbuf_addbuf(&one, &buf);\n \t\telse if (hunk == RR_ORIGINAL)\n \t\t\t; /* discard */\n \t\telse if (hunk == RR_SIDE_2)\n \t\t\tstrbuf_addbuf(&two, &buf);\n-\t\telse\n-\t\t\trerere_io_putstr(buf.buf, io);\n-\t\tcontinue;\n-\tbad:\n-\t\thunk = 99; /* force error exit */\n-\t\tbreak;\n \t}\n+bad:\n+\thas_conflicts = -1;\n+out:\n \tstrbuf_release(&one);\n \tstrbuf_release(&two);\n \tstrbuf_release(&buf);\n \n+\treturn has_conflicts;\n+}\n+\n+/*\n+ * Read contents a file with conflicts, normalize the conflicts\n+ * by (1) discarding the common ancestor version in diff3-style,\n+ * (2) reordering our side and their side so that whichever sorts\n+ * alphabetically earlier comes before the other one, while\n+ * computing the \"conflict ID\", which is just an SHA-1 hash of\n+ * one side of the conflict, NUL, the other side of the conflict,\n+ * and NUL concatenated together.\n+ *\n+ * Return 1 if conflict hunks are found, 0 if there are no conflict\n+ * hunks and -1 if an error occured.\n+ */\n+static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n+{\n+\tgit_SHA_CTX ctx;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf out = STRBUF_INIT;\n+\tint has_conflicts = 0;\n+\tif (sha1)\n+\t\tgit_SHA1_Init(&ctx);\n+\n+\twhile (!io->getline(&buf, io)) {\n+\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n+\t\t\thas_conflicts = handle_conflict(&out, io, marker_size,\n+\t\t\t\t\t\t\t    sha1 ? &ctx : NULL);\n+\t\t\tif (has_conflicts < 0)\n+\t\t\t\tbreak;\n+\t\t\trerere_io_putmem(out.buf, out.len, io);\n+\t\t\tstrbuf_reset(&out);\n+\t\t} else\n+\t\t\trerere_io_putstr(buf.buf, io);\n+\t}\n+\tstrbuf_release(&buf);\n+\tstrbuf_release(&out);\n+\n \tif (sha1)\n \t\tgit_SHA1_Final(sha1, &ctx);\n-\tif (hunk != RR_CONTEXT)\n-\t\treturn -1;\n+\n \treturn has_conflicts;\n }\n \n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"351654","messageId":"20180703210515.GA31234@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"Re: [PATCH v2 00/10] rerere: handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-03T21:05:15Z","receivedAt":"2018-07-03T21:05:22Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 06/05, Thomas Gummerer wrote:\n> The previous round was at\n> <20180520211210.1248-1-t.gummerer@gmail.com>.\n> \n> Thanks Junio for the comments on the previous round.\n> \n> Changes since v2:\n>  - lowercase the first letter in some error/warning messages before\n>    marking them for translation\n>  - wrap paths in output in single quotes, for consistency, and to make\n>    some of the messages the same as ones that are already translated\n>  - mark messages in builtin/rerere.c for translation as well, which I\n>    had previously forgotten.\n>  - expanded the technical documentation on rerere.  The entire\n>    document is basically rewritten.\n>  - changed the test in 6/10 to just fake a conflict marker inside of\n>    one of the hunks instead of using an inner conflict created by a\n>    merge.  This is to make sure the codepath is still hit after we\n>    handle inner conflicts properly.\n>  - added tests for handling inner conflict markers\n>  - added one commit to recalculate the conflict ID when an unresolved\n>    conflict is committed, and the subsequent operation conflicts again\n>    in the same file.  More explanation in the commit message of that\n>    commit.\n\nNow that 2.18 is out (and I'm caught up on the list after being away\nfrom it for a few days), is there any interest in this series? I guess\nit was overlooked as it's been sent in the rc phase for 2.18.\n\nI think the most important bit here is 6/10 which fixes a crash that\ncan happen in \"normal\" usage of git.  The translation bits are also\nnice to have I think, but I could send them in a different series if\nthat's preferred.\n\nThe other patches would be nice to have, but are arguably less\nimportant.\n\n> range-diff below.  A few commits changed enough for range-diff\n> to give up showing the differences in those, they are probably best\n> reviewed as the whole patch anyway:\n>\n> [snip]\n> \n> Thomas Gummerer (10):\n>   rerere: unify error messages when read_cache fails\n>   rerere: lowercase error messages\n>   rerere: wrap paths in output in sq\n>   rerere: mark strings for translation\n>   rerere: add some documentation\n>   rerere: fix crash when conflict goes unresolved\n>   rerere: only return whether a path has conflicts or not\n>   rerere: factor out handle_conflict function\n>   rerere: teach rerere to handle nested conflicts\n>   rerere: recalculate conflict ID when unresolved conflict is committed\n> \n>  Documentation/technical/rerere.txt | 182 +++++++++++++++++++++\n>  builtin/rerere.c                   |   4 +-\n>  rerere.c                           | 246 ++++++++++++++---------------\n>  t/t4200-rerere.sh                  |  67 ++++++++\n>  4 files changed, 372 insertions(+), 127 deletions(-)\n>  create mode 100644 Documentation/technical/rerere.txt\n> \n> -- \n> 2.18.0.rc1.242.g61856ae69\n> \n"},{"id":"351786","messageId":"xmqq1scgmemy.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180703210515.GA31234@hank.intra.tgummerer.com","subject":"Re: [PATCH v2 00/10] rerere: handle nested conflicts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-06T17:56:53Z","receivedAt":"2018-07-06T17:57:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> On 06/05, Thomas Gummerer wrote:\n>> The previous round was at\n>> <20180520211210.1248-1-t.gummerer@gmail.com>.\n>> \n>> Thanks Junio for the comments on the previous round.\n>> \n>> Changes since v2:\n>>  - lowercase the first letter in some error/warning messages before\n>>    marking them for translation\n>>  - wrap paths in output in single quotes, for consistency, and to make\n>>    some of the messages the same as ones that are already translated\n>>  - mark messages in builtin/rerere.c for translation as well, which I\n>>    had previously forgotten.\n>>  - expanded the technical documentation on rerere.  The entire\n>>    document is basically rewritten.\n>>  - changed the test in 6/10 to just fake a conflict marker inside of\n>>    one of the hunks instead of using an inner conflict created by a\n>>    merge.  This is to make sure the codepath is still hit after we\n>>    handle inner conflicts properly.\n>>  - added tests for handling inner conflict markers\n>>  - added one commit to recalculate the conflict ID when an unresolved\n>>    conflict is committed, and the subsequent operation conflicts again\n>>    in the same file.  More explanation in the commit message of that\n>>    commit.\n>\n> Now that 2.18 is out (and I'm caught up on the list after being away\n> from it for a few days), is there any interest in this series? I guess\n> it was overlooked as it's been sent in the rc phase for 2.18.\n\nI deliberately ignored, not because I wasn't interested in it, but\nbecause I'd be distracted during the pre-release feature freeze as\nI'd be heavily intereseted in it.\n\nNow is a good time to repost to stir/re-ignite the interest from\nothers, possibly after rebasing on v2.18.0 and polishing further.\n\nThanks.\n\n>\n> I think the most important bit here is 6/10 which fixes a crash that\n> can happen in \"normal\" usage of git.  The translation bits are also\n> nice to have I think, but I could send them in a different series if\n> that's preferred.\n>\n> The other patches would be nice to have, but are arguably less\n> important.\n>\n>> range-diff below.  A few commits changed enough for range-diff\n>> to give up showing the differences in those, they are probably best\n>> reviewed as the whole patch anyway:\n>>\n>> [snip]\n>> \n>> Thomas Gummerer (10):\n>>   rerere: unify error messages when read_cache fails\n>>   rerere: lowercase error messages\n>>   rerere: wrap paths in output in sq\n>>   rerere: mark strings for translation\n>>   rerere: add some documentation\n>>   rerere: fix crash when conflict goes unresolved\n>>   rerere: only return whether a path has conflicts or not\n>>   rerere: factor out handle_conflict function\n>>   rerere: teach rerere to handle nested conflicts\n>>   rerere: recalculate conflict ID when unresolved conflict is committed\n>> \n>>  Documentation/technical/rerere.txt | 182 +++++++++++++++++++++\n>>  builtin/rerere.c                   |   4 +-\n>>  rerere.c                           | 246 ++++++++++++++---------------\n>>  t/t4200-rerere.sh                  |  67 ++++++++\n>>  4 files changed, 372 insertions(+), 127 deletions(-)\n>>  create mode 100644 Documentation/technical/rerere.txt\n>> \n>> -- \n>> 2.18.0.rc1.242.g61856ae69\n>> \n"},{"id":"352184","messageId":"20180710213742.GA2186@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqq1scgmemy.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v2 00/10] rerere: handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-10T21:37:42Z","receivedAt":"2018-07-10T21:37:48Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 07/06, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > On 06/05, Thomas Gummerer wrote:\n> >> The previous round was at\n> >> <20180520211210.1248-1-t.gummerer@gmail.com>.\n> >> \n> >> Thanks Junio for the comments on the previous round.\n> >> \n> >> Changes since v2:\n> >>  - lowercase the first letter in some error/warning messages before\n> >>    marking them for translation\n> >>  - wrap paths in output in single quotes, for consistency, and to make\n> >>    some of the messages the same as ones that are already translated\n> >>  - mark messages in builtin/rerere.c for translation as well, which I\n> >>    had previously forgotten.\n> >>  - expanded the technical documentation on rerere.  The entire\n> >>    document is basically rewritten.\n> >>  - changed the test in 6/10 to just fake a conflict marker inside of\n> >>    one of the hunks instead of using an inner conflict created by a\n> >>    merge.  This is to make sure the codepath is still hit after we\n> >>    handle inner conflicts properly.\n> >>  - added tests for handling inner conflict markers\n> >>  - added one commit to recalculate the conflict ID when an unresolved\n> >>    conflict is committed, and the subsequent operation conflicts again\n> >>    in the same file.  More explanation in the commit message of that\n> >>    commit.\n> >\n> > Now that 2.18 is out (and I'm caught up on the list after being away\n> > from it for a few days), is there any interest in this series? I guess\n> > it was overlooked as it's been sent in the rc phase for 2.18.\n> \n> I deliberately ignored, not because I wasn't interested in it, but\n> because I'd be distracted during the pre-release feature freeze as\n> I'd be heavily intereseted in it.\n> \n> Now is a good time to repost to stir/re-ignite the interest from\n> others, possibly after rebasing on v2.18.0 and polishing further.\n\nI sometimes find it hard to gauge whether there are no replies because\nnobody is interested in the series, or if it is because it was ignored\nor slipped to the cracks.  I guess I could have inferred it from your\nreplies to the previous iteration though :)\n\nI'll go back and polish my patches, and then send a new iteration,\nthanks! \n\n> Thanks.\n> \n> >\n> > I think the most important bit here is 6/10 which fixes a crash that\n> > can happen in \"normal\" usage of git.  The translation bits are also\n> > nice to have I think, but I could send them in a different series if\n> > that's preferred.\n> >\n> > The other patches would be nice to have, but are arguably less\n> > important.\n> >\n> >> range-diff below.  A few commits changed enough for range-diff\n> >> to give up showing the differences in those, they are probably best\n> >> reviewed as the whole patch anyway:\n> >>\n> >> [snip]\n> >> \n> >> Thomas Gummerer (10):\n> >>   rerere: unify error messages when read_cache fails\n> >>   rerere: lowercase error messages\n> >>   rerere: wrap paths in output in sq\n> >>   rerere: mark strings for translation\n> >>   rerere: add some documentation\n> >>   rerere: fix crash when conflict goes unresolved\n> >>   rerere: only return whether a path has conflicts or not\n> >>   rerere: factor out handle_conflict function\n> >>   rerere: teach rerere to handle nested conflicts\n> >>   rerere: recalculate conflict ID when unresolved conflict is committed\n> >> \n> >>  Documentation/technical/rerere.txt | 182 +++++++++++++++++++++\n> >>  builtin/rerere.c                   |   4 +-\n> >>  rerere.c                           | 246 ++++++++++++++---------------\n> >>  t/t4200-rerere.sh                  |  67 ++++++++\n> >>  4 files changed, 372 insertions(+), 127 deletions(-)\n> >>  create mode 100644 Documentation/technical/rerere.txt\n> >> \n> >> -- \n> >> 2.18.0.rc1.242.g61856ae69\n> >> \n"},{"id":"352561","messageId":"20180714214443.7184-2-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 01/11] rerere: unify error messages when read_cache fails","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:33Z","receivedAt":"2018-07-14T21:44:58Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"We have multiple different variants of the error message we show to\nthe user if 'read_cache' fails.  The \"Could not read index\" variant we\nare using in 'rerere.c' is currently not used anywhere in translated\nform.\n\nAs a subsequent commit will mark all output that comes from 'rerere.c'\nfor translation, make the life of the translators a little bit easier\nby using a string that is used elsewhere, and marked for translation\nthere, and thus most likely already translated.\n\n\"index file corrupt\" seems to be the most common error message we show\nwhen 'read_cache' fails, so use that here as well.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex e0862e2778..473d32a5cd 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -568,7 +568,7 @@ static int find_conflict(struct string_list *conflict)\n {\n \tint i;\n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -601,7 +601,7 @@ int rerere_remaining(struct string_list *merge_rr)\n \tif (setup_rerere(merge_rr, RERERE_READONLY))\n \t\treturn 0;\n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -1103,7 +1103,7 @@ int rerere_forget(struct pathspec *pathspec)\n \tstruct string_list merge_rr = STRING_LIST_INIT_DUP;\n \n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfd = setup_rerere(&merge_rr, RERERE_NOAUTOUPDATE);\n \tif (fd < 0)\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352562","messageId":"20180714214443.7184-4-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 03/11] rerere: wrap paths in output in sq","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:35Z","receivedAt":"2018-07-14T21:44:58Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"It looks like most paths in the output in the git codebase are wrapped\nin single quotes.  Standardize on that in rerere as well.\n\nApart from being more consistent, this also makes some of the strings\nmatch strings that are already translated in other parts of the\ncodebase, thus reducing the work for translators, when the strings are\nmarked for translation in a subsequent commit.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n builtin/rerere.c |  2 +-\n rerere.c         | 26 +++++++++++++-------------\n 2 files changed, 14 insertions(+), 14 deletions(-)\n\ndiff --git a/builtin/rerere.c b/builtin/rerere.c\nindex 0bc40298c2..e0c67c98e9 100644\n--- a/builtin/rerere.c\n+++ b/builtin/rerere.c\n@@ -107,7 +107,7 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \t\t\tconst char *path = merge_rr.items[i].string;\n \t\t\tconst struct rerere_id *id = merge_rr.items[i].util;\n \t\t\tif (diff_two(rerere_path(id, \"preimage\"), path, path, path))\n-\t\t\t\tdie(\"unable to generate diff for %s\", rerere_path(id, NULL));\n+\t\t\t\tdie(\"unable to generate diff for '%s'\", rerere_path(id, NULL));\n \t\t}\n \t} else\n \t\tusage_with_options(rerere_usage, options);\ndiff --git a/rerere.c b/rerere.c\nindex c5d9ea171f..cde1f6e696 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"could not open %s\", path);\n+\t\treturn error_errno(\"could not open '%s'\", path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"could not write %s\", output);\n+\t\t\terror_errno(\"could not write '%s'\", output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"there were errors while writing %s (%s)\",\n+\t\terror(\"there were errors while writing '%s' (%s)\",\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"failed to flush %s\", path);\n+\t\tio.io.wrerror = error_errno(\"failed to flush '%s'\", path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"could not parse conflict hunks in %s\", path);\n+\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -684,17 +684,17 @@ static int merge(const struct rerere_id *id, const char *path)\n \t * Mark that \"postimage\" was used to help gc.\n \t */\n \tif (utime(rerere_path(id, \"postimage\"), NULL) < 0)\n-\t\twarning_errno(\"failed utime() on %s\",\n+\t\twarning_errno(\"failed utime() on '%s'\",\n \t\t\t      rerere_path(id, \"postimage\"));\n \n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"could not open %s\", path);\n+\t\treturn error_errno(\"could not open '%s'\", path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"could not write %s\", path);\n+\t\terror_errno(\"could not write '%s'\", path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"writing %s failed\", path);\n+\t\treturn error_errno(\"writing '%s' failed\", path);\n \n out:\n \tfree(cur.ptr);\n@@ -878,7 +878,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"could not create directory %s\", git_path_rr_cache());\n+\t\tdie(\"could not create directory '%s'\", git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1067,9 +1067,9 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \tfilename = rerere_path(id, \"postimage\");\n \tif (unlink(filename)) {\n \t\tif (errno == ENOENT)\n-\t\t\terror(\"no remembered resolution for %s\", path);\n+\t\t\terror(\"no remembered resolution for '%s'\", path);\n \t\telse\n-\t\t\terror_errno(\"cannot unlink %s\", filename);\n+\t\t\terror_errno(\"cannot unlink '%s'\", filename);\n \t\tgoto fail_exit;\n \t}\n \n@@ -1088,7 +1088,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \titem = string_list_insert(rr, path);\n \tfree_rerere_id(item);\n \titem->util = id;\n-\tfprintf(stderr, \"Forgot resolution for %s\\n\", path);\n+\tfprintf(stderr, \"Forgot resolution for '%s'\\n\", path);\n \treturn 0;\n \n fail_exit:\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352563","messageId":"20180714214443.7184-3-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 02/11] rerere: lowercase error messages","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:34Z","receivedAt":"2018-07-14T21:44:58Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Documentation/CodingGuidelines mentions that error messages should be\nlowercase.  Prior to marking them for translation follow that pattern\nin rerere as well, so translators won't have to translate messages\nthat don't conform to our guidelines.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 24 ++++++++++++------------\n 1 file changed, 12 insertions(+), 12 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 473d32a5cd..c5d9ea171f 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"Could not open %s\", path);\n+\t\treturn error_errno(\"could not open %s\", path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"Could not write %s\", output);\n+\t\t\terror_errno(\"could not write %s\", output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"There were errors while writing %s (%s)\",\n+\t\terror(\"there were errors while writing %s (%s)\",\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"Failed to flush %s\", path);\n+\t\tio.io.wrerror = error_errno(\"failed to flush %s\", path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"Could not parse conflict hunks in %s\", path);\n+\t\treturn error(\"could not parse conflict hunks in %s\", path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -690,11 +690,11 @@ static int merge(const struct rerere_id *id, const char *path)\n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"Could not open %s\", path);\n+\t\treturn error_errno(\"could not open %s\", path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"Could not write %s\", path);\n+\t\terror_errno(\"could not write %s\", path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"Writing %s failed\", path);\n+\t\treturn error_errno(\"writing %s failed\", path);\n \n out:\n \tfree(cur.ptr);\n@@ -720,7 +720,7 @@ static void update_paths(struct string_list *update)\n \n \tif (write_locked_index(&the_index, &index_lock,\n \t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED))\n-\t\tdie(\"Unable to write new index file\");\n+\t\tdie(\"unable to write new index file\");\n }\n \n static void remove_variant(struct rerere_id *id)\n@@ -878,7 +878,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"Could not create directory %s\", git_path_rr_cache());\n+\t\tdie(\"could not create directory %s\", git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1031,7 +1031,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t */\n \tret = handle_cache(path, sha1, NULL);\n \tif (ret < 1)\n-\t\treturn error(\"Could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n \n \t/* Nuke the recorded resolution for the conflict */\n \tid = new_rerere_id(sha1);\n@@ -1049,7 +1049,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t\thandle_cache(path, sha1, rerere_path(id, \"thisimage\"));\n \t\tif (read_mmfile(&cur, rerere_path(id, \"thisimage\"))) {\n \t\t\tfree(cur.ptr);\n-\t\t\terror(\"Failed to update conflicted state in '%s'\", path);\n+\t\t\terror(\"failed to update conflicted state in '%s'\", path);\n \t\t\tgoto fail_exit;\n \t\t}\n \t\tcleanly_resolved = !try_merge(id, path, &cur, &result);\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352564","messageId":"20180714214443.7184-1-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180605215219.28783-1-t.gummerer@gmail.com","subject":"[PATCH v3 00/11] rerere: handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:32Z","receivedAt":"2018-07-14T21:44:58Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"The previous rounds were at\n<20180520211210.1248-1-t.gummerer@gmail.com> and\n<20180605215219.28783-1-t.gummerer@gmail.com>.\n\nThis round is a more polished version of the previous round, as\nsuggested by Junio in <xmqq1scgmemy.fsf@gitster-ct.c.googlers.com>.\nIt's also rebased on v2.18.\n\nThe series grew by one patch, because 8/10 has been split into two\npatches hopefully making it easier to follow.\n\nrange-diff before below:\n\n1:  2825342cc2 = 1:  018bd68a8a rerere: unify error messages when read_cache fails\n2:  d1500028aa ! 2:  281fcbf24f rerere: lowercase error messages\n    @@ -4,7 +4,8 @@\n     \n         Documentation/CodingGuidelines mentions that error messages should be\n         lowercase.  Prior to marking them for translation follow that pattern\n    -    in rerere as well.\n    +    in rerere as well, so translators won't have to translate messages\n    +    that don't conform to our guidelines.\n     \n         Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n     \n    @@ -87,3 +88,12 @@\n      \n      \t/* Nuke the recorded resolution for the conflict */\n      \tid = new_rerere_id(sha1);\n    +@@\n    + \t\thandle_cache(path, sha1, rerere_path(id, \"thisimage\"));\n    + \t\tif (read_mmfile(&cur, rerere_path(id, \"thisimage\"))) {\n    + \t\t\tfree(cur.ptr);\n    +-\t\t\terror(\"Failed to update conflicted state in '%s'\", path);\n    ++\t\t\terror(\"failed to update conflicted state in '%s'\", path);\n    + \t\t\tgoto fail_exit;\n    + \t\t}\n    + \t\tcleanly_resolved = !try_merge(id, path, &cur, &result);\n3:  ed3601ee71 = 3:  b6d5e2e26d rerere: wrap paths in output in sq\n4:  6ead84a199 ! 4:  45f0d7a99f rerere: mark strings for translation\n    @@ -4,7 +4,7 @@\n     \n         'git rerere' is considered a plumbing command and as such its output\n         should be translated.  Its functionality is also only enabled through\n    -    a config setting, so scripts really shouldn't rely on its output\n    +    a config setting, so scripts really shouldn't rely on the output\n         either way.\n     \n         Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n    @@ -219,8 +219,8 @@\n      \t\thandle_cache(path, sha1, rerere_path(id, \"thisimage\"));\n      \t\tif (read_mmfile(&cur, rerere_path(id, \"thisimage\"))) {\n      \t\t\tfree(cur.ptr);\n    --\t\t\terror(\"Failed to update conflicted state in '%s'\", path);\n    -+\t\t\terror(_(\"Failed to update conflicted state in '%s'\"), path);\n    +-\t\t\terror(\"failed to update conflicted state in '%s'\", path);\n    ++\t\t\terror(_(\"failed to update conflicted state in '%s'\"), path);\n      \t\t\tgoto fail_exit;\n      \t\t}\n      \t\tcleanly_resolved = !try_merge(id, path, &cur, &result);\n5:  caad276aca ! 5:  993857a816 rerere: add some documentation\n    @@ -1,6 +1,6 @@\n     Author: Thomas Gummerer <t.gummerer@gmail.com>\n     \n    -    rerere: add some documentation\n    +    rerere: add documentation for conflict normalization\n     \n         Add some documentation for the logic behind the conflict normalization\n         in rerere.\n    @@ -27,30 +27,28 @@\n     +when different conflict style settings are used, rerere normalizes the\n     +conflicts before writing them to the rerere database.\n     +\n    -+Differnt conflict styles and branch names are dealt with by stripping\n    -+that information from the conflict markers, and removing extraneous\n    -+information from the `diff3` conflict style.\n    -+\n    -+Branches being merged in different order are dealt with by sorting the\n    -+conflict hunks.  More on each of those parts in the following\n    -+sections.\n    ++Different conflict styles and branch names are normalized by stripping\n    ++the labels from the conflict markers, and removing extraneous\n    ++information from the `diff3` conflict style. Branches that are merged\n    ++in different order are normalized by sorting the conflict hunks.  More\n    ++on each of those steps in the following sections.\n     +\n     +Once these two normalization operations are applied, a conflict ID is\n    -+created based on the normalized conflict, which is later used by\n    ++calculated based on the normalized conflict, which is later used by\n     +rerere to look up the conflict in the rerere database.\n     +\n     +Stripping extraneous information\n     +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n     +\n     +Say we have three branches AB, AC and AC2.  The common ancestor of\n    -+these branches has a file with with a line with the string \"A\" (for\n    -+brevity this line is called \"line A\" for brevity in the following) in\n    -+it.  In branch AB this line is changed to \"B\", in AC, this line is\n    -+changed to C, and branch AC2 is forked off of AC, after the line was\n    -+changed to C.\n    ++these branches has a file with a line containing the string \"A\" (for\n    ++brevity this is called \"line A\" in the rest of the document).  In\n    ++branch AB this line is changed to \"B\", in AC, this line is changed to\n    ++\"C\", and branch AC2 is forked off of AC, after the line was changed to\n    ++\"C\".\n     +\n    -+Now forking a branch ABAC off of branch AB and then merging AC into it,\n    -+we'd get a conflict like the following:\n    ++Forking a branch ABAC off of branch AB and then merging AC into it, we\n    ++get a conflict like the following:\n     +\n     +    <<<<<<< HEAD\n     +    B\n    @@ -58,9 +56,9 @@\n     +    C\n     +    >>>>>>> AC\n     +\n    -+Now doing the analogous with AC2 (forking a branch ABAC2 off of branch\n    -+AB and then merging branch AC2 into it), maybe using the diff3\n    -+conflict style, we'd get a conflict like the following:\n    ++Doing the analogous with AC2 (forking a branch ABAC2 off of branch AB\n    ++and then merging branch AC2 into it), using the diff3 conflict style,\n    ++we get a conflict like the following:\n     +\n     +    <<<<<<< HEAD\n     +    B\n6:  ad88a6b8a8 = 6:  a7a0f657f3 rerere: fix crash when conflict goes unresolved\n7:  15f9efcba6 = 7:  f1afd4b9a4 rerere: only return whether a path has conflicts or not\n8:  1490efaad3 ! 8:  1f5cef506a rerere: factor out handle_conflict function\n    @@ -3,53 +3,14 @@\n         rerere: factor out handle_conflict function\n     \n         Factor out the handle_conflict function, which handles a single\n    -    conflict in a path.  This is a preparation for the next step, where\n    -    this function will be re-used.  No functional changes intended.\n    +    conflict in a path.  This is in preparation for a subsequent commit,\n    +    where this function will be re-used.  No functional changes intended.\n     \n         Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n     \n     diff --git a/rerere.c b/rerere.c\n     --- a/rerere.c\n     +++ b/rerere.c\n    -@@\n    - \t\tferr_puts(str, io->output, &io->wrerror);\n    - }\n    - \n    --/*\n    -- * Write a conflict marker to io->output (if defined).\n    -- */\n    --static void rerere_io_putconflict(int ch, int size, struct rerere_io *io)\n    --{\n    --\tchar buf[64];\n    --\n    --\twhile (size) {\n    --\t\tif (size <= sizeof(buf) - 2) {\n    --\t\t\tmemset(buf, ch, size);\n    --\t\t\tbuf[size] = '\\n';\n    --\t\t\tbuf[size + 1] = '\\0';\n    --\t\t\tsize = 0;\n    --\t\t} else {\n    --\t\t\tint sz = sizeof(buf) - 1;\n    --\n    --\t\t\t/*\n    --\t\t\t * Make sure we will not write everything out\n    --\t\t\t * in this round by leaving at least 1 byte\n    --\t\t\t * for the next round, giving the next round\n    --\t\t\t * a chance to add the terminating LF.  Yuck.\n    --\t\t\t */\n    --\t\t\tif (size <= sz)\n    --\t\t\t\tsz -= (sz - size) + 1;\n    --\t\t\tmemset(buf, ch, sz);\n    --\t\t\tbuf[sz] = '\\0';\n    --\t\t\tsize -= sz;\n    --\t\t}\n    --\t\trerere_io_putstr(buf, io);\n    --\t}\n    --}\n    --\n    - static void rerere_io_putmem(const char *mem, size_t sz, struct rerere_io *io)\n    - {\n    - \tif (io->output)\n     @@\n      \treturn isspace(*buf);\n      }\n    @@ -67,14 +28,7 @@\n     - * hunks and -1 if an error occured.\n     - */\n     -static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n    -+static void rerere_strbuf_putconflict(struct strbuf *buf, int ch, size_t size)\n    -+{\n    -+\tstrbuf_addchars(buf, ch, size);\n    -+\tstrbuf_addch(buf, '\\n');\n    -+}\n    -+\n    -+static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n    -+\t\t\t   int marker_size, git_SHA_CTX *ctx)\n    ++static int handle_conflict(struct rerere_io *io, int marker_size, git_SHA_CTX *ctx)\n      {\n     -\tgit_SHA_CTX ctx;\n     -\tint has_conflicts = 0;\n    @@ -88,38 +42,39 @@\n     -\n     -\tif (sha1)\n     -\t\tgit_SHA1_Init(&ctx);\n    --\n    -+\tint has_conflicts = 1;\n    ++\tint has_conflicts = -1;\n    + \n      \twhile (!io->getline(&buf, io)) {\n    --\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n    + \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n     -\t\t\tif (hunk != RR_CONTEXT)\n     -\t\t\t\tgoto bad;\n     -\t\t\thunk = RR_SIDE_1;\n    --\t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n    -+\t\tif (is_cmarker(buf.buf, '<', marker_size))\n    -+\t\t\tgoto bad;\n    -+\t\telse if (is_cmarker(buf.buf, '|', marker_size)) {\n    ++\t\t\tbreak;\n    + \t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n      \t\t\tif (hunk != RR_SIDE_1)\n    - \t\t\t\tgoto bad;\n    +-\t\t\t\tgoto bad;\n    ++\t\t\t\tbreak;\n      \t\t\thunk = RR_ORIGINAL;\n    -@@\n    - \t\t\t\tgoto bad;\n    + \t\t} else if (is_cmarker(buf.buf, '=', marker_size)) {\n    + \t\t\tif (hunk != RR_SIDE_1 && hunk != RR_ORIGINAL)\n    +-\t\t\t\tgoto bad;\n    ++\t\t\t\tbreak;\n    + \t\t\thunk = RR_SIDE_2;\n    + \t\t} else if (is_cmarker(buf.buf, '>', marker_size)) {\n    + \t\t\tif (hunk != RR_SIDE_2)\n    +-\t\t\t\tgoto bad;\n    ++\t\t\t\tbreak;\n      \t\t\tif (strbuf_cmp(&one, &two) > 0)\n      \t\t\t\tstrbuf_swap(&one, &two);\n    --\t\t\thas_conflicts = 1;\n    + \t\t\thas_conflicts = 1;\n     -\t\t\thunk = RR_CONTEXT;\n    --\t\t\trerere_io_putconflict('<', marker_size, io);\n    --\t\t\trerere_io_putmem(one.buf, one.len, io);\n    --\t\t\trerere_io_putconflict('=', marker_size, io);\n    --\t\t\trerere_io_putmem(two.buf, two.len, io);\n    --\t\t\trerere_io_putconflict('>', marker_size, io);\n    + \t\t\trerere_io_putconflict('<', marker_size, io);\n    + \t\t\trerere_io_putmem(one.buf, one.len, io);\n    + \t\t\trerere_io_putconflict('=', marker_size, io);\n    + \t\t\trerere_io_putmem(two.buf, two.len, io);\n    + \t\t\trerere_io_putconflict('>', marker_size, io);\n     -\t\t\tif (sha1) {\n     -\t\t\t\tgit_SHA1_Update(&ctx, one.buf ? one.buf : \"\",\n    -+\t\t\trerere_strbuf_putconflict(out, '<', marker_size);\n    -+\t\t\tstrbuf_addbuf(out, &one);\n    -+\t\t\trerere_strbuf_putconflict(out, '=', marker_size);\n    -+\t\t\tstrbuf_addbuf(out, &two);\n    -+\t\t\trerere_strbuf_putconflict(out, '>', marker_size);\n     +\t\t\tif (ctx) {\n     +\t\t\t\tgit_SHA1_Update(ctx, one.buf ? one.buf : \"\",\n      \t\t\t\t\t    one.len + 1);\n    @@ -129,7 +84,7 @@\n      \t\t\t}\n     -\t\t\tstrbuf_reset(&one);\n     -\t\t\tstrbuf_reset(&two);\n    -+\t\t\tgoto out;\n    ++\t\t\tbreak;\n      \t\t} else if (hunk == RR_SIDE_1)\n      \t\t\tstrbuf_addbuf(&one, &buf);\n      \t\telse if (hunk == RR_ORIGINAL)\n    @@ -143,9 +98,6 @@\n     -\t\thunk = 99; /* force error exit */\n     -\t\tbreak;\n      \t}\n    -+bad:\n    -+\thas_conflicts = -1;\n    -+out:\n      \tstrbuf_release(&one);\n      \tstrbuf_release(&two);\n      \tstrbuf_release(&buf);\n    @@ -169,24 +121,20 @@\n     +{\n     +\tgit_SHA_CTX ctx;\n     +\tstruct strbuf buf = STRBUF_INIT;\n    -+\tstruct strbuf out = STRBUF_INIT;\n     +\tint has_conflicts = 0;\n     +\tif (sha1)\n     +\t\tgit_SHA1_Init(&ctx);\n     +\n     +\twhile (!io->getline(&buf, io)) {\n     +\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n    -+\t\t\thas_conflicts = handle_conflict(&out, io, marker_size,\n    -+\t\t\t\t\t\t\t    sha1 ? &ctx : NULL);\n    ++\t\t\thas_conflicts = handle_conflict(io, marker_size,\n    ++\t\t\t\t\t\t\tsha1 ? &ctx : NULL);\n     +\t\t\tif (has_conflicts < 0)\n     +\t\t\t\tbreak;\n    -+\t\t\trerere_io_putmem(out.buf, out.len, io);\n    -+\t\t\tstrbuf_reset(&out);\n     +\t\t} else\n     +\t\t\trerere_io_putstr(buf.buf, io);\n     +\t}\n     +\tstrbuf_release(&buf);\n    -+\tstrbuf_release(&out);\n     +\n      \tif (sha1)\n      \t\tgit_SHA1_Final(sha1, &ctx);\n-:  ---------- > 9:  8ac0d3e903 rerere: return strbuf from handle path\n9:  6619650c42 ! 10:  ef84fdc201 rerere: teach rerere to handle nested conflicts\n    @@ -27,7 +27,7 @@\n             >>>>>>> branch-2\n             >>>>>>> branch-3~\n     \n    -    it would be reordered as follows in the preimage:\n    +    it would be recorde as follows in the preimage:\n     \n             <<<<<<<\n             1\n    @@ -40,14 +40,16 @@\n             >>>>>>>\n     \n         and the conflict ID would be calculated as\n    +\n             sha1(1<NUL><<<<<<<\n             2\n             =======\n             3\n             >>>>>>><NUL>)\n     \n    -    Stripping out vs. leaving the conflict markers in place should have no\n    -    practical impact, but it simplifies the implementation.\n    +    Stripping out vs. leaving the conflict markers in place in the inner\n    +    conflict should have no practical impact, but it simplifies the\n    +    implementation.\n     \n         Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n     \n    @@ -63,9 +65,11 @@\n     +~~~~~~~~~~~~~~~~\n     +\n     +Nested conflicts are handled very similarly to \"simple\" conflicts.\n    -+Same as before, labels on conflict markers and diff3 output is\n    -+stripped, and the conflict hunks are sorted, for both the outer and\n    -+the inner conflict.\n    ++Similar to simple conflicts, the conflict is first normalized by\n    ++stripping the labels from conflict markers, stripping the diff3\n    ++output, and the sorting the conflict hunks, both for the outer and the\n    ++inner conflict.  This is done recursively, so any number of nested\n    ++conflicts can be handled.\n     +\n     +The only difference is in how the conflict ID is calculated.  For the\n     +inner conflict, the conflict markers themselves are not stripped out\n    @@ -83,16 +87,16 @@\n     +    >>>>>>> branch-2\n     +    >>>>>>> branch-3~\n     +\n    -+After stripping out the labels of the conflict markers, the conflict\n    -+would look as follows:\n    ++After stripping out the labels of the conflict markers, and sorting\n    ++the hunks, the conflict would look as follows:\n     +\n     +    <<<<<<<\n     +    1\n     +    =======\n     +    <<<<<<<\n    -+    3\n    -+    =======\n     +    2\n    ++    =======\n    ++    3\n     +    >>>>>>>\n     +    >>>>>>>\n     +\n    @@ -108,23 +112,21 @@\n      \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n     -\tstruct strbuf buf = STRBUF_INIT;\n     +\tstruct strbuf buf = STRBUF_INIT, conflict = STRBUF_INIT;\n    - \tint has_conflicts = 1;\n    + \tint has_conflicts = -1;\n    + \n      \twhile (!io->getline(&buf, io)) {\n    --\t\tif (is_cmarker(buf.buf, '<', marker_size))\n    --\t\t\tgoto bad;\n    --\t\telse if (is_cmarker(buf.buf, '|', marker_size)) {\n    -+\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n    + \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n    +-\t\t\tbreak;\n     +\t\t\tif (handle_conflict(&conflict, io, marker_size, NULL) < 0)\n    -+\t\t\t\tgoto bad;\n    ++\t\t\t\tbreak;\n     +\t\t\tif (hunk == RR_SIDE_1)\n     +\t\t\t\tstrbuf_addbuf(&one, &conflict);\n     +\t\t\telse\n     +\t\t\t\tstrbuf_addbuf(&two, &conflict);\n     +\t\t\tstrbuf_release(&conflict);\n    -+\t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n    + \t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n      \t\t\tif (hunk != RR_SIDE_1)\n    - \t\t\t\tgoto bad;\n    - \t\t\thunk = RR_ORIGINAL;\n    + \t\t\t\tbreak;\n     \n     diff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\n     --- a/t/t4200-rerere.sh\n    @@ -169,6 +171,5 @@\n     +\tcat test >actual &&\n     +\ttest_cmp expect actual\n     +'\n    -+\n     +\n      test_done\n10:  4b11dce7dd = 11:  35a826908f rerere: recalculate conflict ID when unresolved conflict is committed\n\n\nThomas Gummerer (11):\n  rerere: unify error messages when read_cache fails\n  rerere: lowercase error messages\n  rerere: wrap paths in output in sq\n  rerere: mark strings for translation\n  rerere: add documentation for conflict normalization\n  rerere: fix crash when conflict goes unresolved\n  rerere: only return whether a path has conflicts or not\n  rerere: factor out handle_conflict function\n  rerere: return strbuf from handle path\n  rerere: teach rerere to handle nested conflicts\n  rerere: recalculate conflict ID when unresolved conflict is committed\n\n Documentation/technical/rerere.txt | 182 +++++++++++++++++++++\n builtin/rerere.c                   |   4 +-\n rerere.c                           | 243 ++++++++++++++---------------\n t/t4200-rerere.sh                  |  66 ++++++++\n 4 files changed, 366 insertions(+), 129 deletions(-)\n create mode 100644 Documentation/technical/rerere.txt\n\n-- \n2.18.0.233.g985f88cf7e\n"},{"id":"352565","messageId":"20180714214443.7184-5-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 04/11] rerere: mark strings for translation","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:36Z","receivedAt":"2018-07-14T21:44:59Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"'git rerere' is considered a plumbing command and as such its output\nshould be translated.  Its functionality is also only enabled through\na config setting, so scripts really shouldn't rely on the output\neither way.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n builtin/rerere.c |  4 +--\n rerere.c         | 68 ++++++++++++++++++++++++------------------------\n 2 files changed, 36 insertions(+), 36 deletions(-)\n\ndiff --git a/builtin/rerere.c b/builtin/rerere.c\nindex e0c67c98e9..5ed941b91f 100644\n--- a/builtin/rerere.c\n+++ b/builtin/rerere.c\n@@ -75,7 +75,7 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \tif (!strcmp(argv[0], \"forget\")) {\n \t\tstruct pathspec pathspec;\n \t\tif (argc < 2)\n-\t\t\twarning(\"'git rerere forget' without paths is deprecated\");\n+\t\t\twarning(_(\"'git rerere forget' without paths is deprecated\"));\n \t\tparse_pathspec(&pathspec, 0, PATHSPEC_PREFER_CWD,\n \t\t\t       prefix, argv + 1);\n \t\treturn rerere_forget(&pathspec);\n@@ -107,7 +107,7 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \t\t\tconst char *path = merge_rr.items[i].string;\n \t\t\tconst struct rerere_id *id = merge_rr.items[i].util;\n \t\t\tif (diff_two(rerere_path(id, \"preimage\"), path, path, path))\n-\t\t\t\tdie(\"unable to generate diff for '%s'\", rerere_path(id, NULL));\n+\t\t\t\tdie(_(\"unable to generate diff for '%s'\"), rerere_path(id, NULL));\n \t\t}\n \t} else\n \t\tusage_with_options(rerere_usage, options);\ndiff --git a/rerere.c b/rerere.c\nindex cde1f6e696..be98c0afcb 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -212,7 +212,7 @@ static void read_rr(struct string_list *rr)\n \n \t\t/* There has to be the hash, tab, path and then NUL */\n \t\tif (buf.len < 42 || get_sha1_hex(buf.buf, sha1))\n-\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \n \t\tif (buf.buf[40] != '.') {\n \t\t\tvariant = 0;\n@@ -221,10 +221,10 @@ static void read_rr(struct string_list *rr)\n \t\t\terrno = 0;\n \t\t\tvariant = strtol(buf.buf + 41, &path, 10);\n \t\t\tif (errno)\n-\t\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \t\t}\n \t\tif (*(path++) != '\\t')\n-\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \t\tbuf.buf[40] = '\\0';\n \t\tid = new_rerere_id_hex(buf.buf);\n \t\tid->variant = variant;\n@@ -259,12 +259,12 @@ static int write_rr(struct string_list *rr, int out_fd)\n \t\t\t\t    rr->items[i].string, 0);\n \n \t\tif (write_in_full(out_fd, buf.buf, buf.len) < 0)\n-\t\t\tdie(\"unable to write rerere record\");\n+\t\t\tdie(_(\"unable to write rerere record\"));\n \n \t\tstrbuf_release(&buf);\n \t}\n \tif (commit_lock_file(&write_lock) != 0)\n-\t\tdie(\"unable to write rerere record\");\n+\t\tdie(_(\"unable to write rerere record\"));\n \treturn 0;\n }\n \n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"could not open '%s'\", path);\n+\t\treturn error_errno(_(\"could not open '%s'\"), path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"could not write '%s'\", output);\n+\t\t\terror_errno(_(\"could not write '%s'\"), output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"there were errors while writing '%s' (%s)\",\n+\t\terror(_(\"there were errors while writing '%s' (%s)\"),\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"failed to flush '%s'\", path);\n+\t\tio.io.wrerror = error_errno(_(\"failed to flush '%s'\"), path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -568,7 +568,7 @@ static int find_conflict(struct string_list *conflict)\n {\n \tint i;\n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -601,7 +601,7 @@ int rerere_remaining(struct string_list *merge_rr)\n \tif (setup_rerere(merge_rr, RERERE_READONLY))\n \t\treturn 0;\n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -684,17 +684,17 @@ static int merge(const struct rerere_id *id, const char *path)\n \t * Mark that \"postimage\" was used to help gc.\n \t */\n \tif (utime(rerere_path(id, \"postimage\"), NULL) < 0)\n-\t\twarning_errno(\"failed utime() on '%s'\",\n+\t\twarning_errno(_(\"failed utime() on '%s'\"),\n \t\t\t      rerere_path(id, \"postimage\"));\n \n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"could not open '%s'\", path);\n+\t\treturn error_errno(_(\"could not open '%s'\"), path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"could not write '%s'\", path);\n+\t\terror_errno(_(\"could not write '%s'\"), path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"writing '%s' failed\", path);\n+\t\treturn error_errno(_(\"writing '%s' failed\"), path);\n \n out:\n \tfree(cur.ptr);\n@@ -714,13 +714,13 @@ static void update_paths(struct string_list *update)\n \t\tstruct string_list_item *item = &update->items[i];\n \t\tif (add_file_to_cache(item->string, 0))\n \t\t\texit(128);\n-\t\tfprintf(stderr, \"Staged '%s' using previous resolution.\\n\",\n+\t\tfprintf_ln(stderr, _(\"Staged '%s' using previous resolution.\"),\n \t\t\titem->string);\n \t}\n \n \tif (write_locked_index(&the_index, &index_lock,\n \t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED))\n-\t\tdie(\"unable to write new index file\");\n+\t\tdie(_(\"unable to write new index file\"));\n }\n \n static void remove_variant(struct rerere_id *id)\n@@ -752,7 +752,7 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \t\tif (!handle_file(path, NULL, NULL)) {\n \t\t\tcopy_file(rerere_path(id, \"postimage\"), path, 0666);\n \t\t\tid->collection->status[variant] |= RR_HAS_POSTIMAGE;\n-\t\t\tfprintf(stderr, \"Recorded resolution for '%s'.\\n\", path);\n+\t\t\tfprintf_ln(stderr, _(\"Recorded resolution for '%s'.\"), path);\n \t\t\tfree_rerere_id(rr_item);\n \t\t\trr_item->util = NULL;\n \t\t\treturn;\n@@ -786,9 +786,9 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \t\tif (rerere_autoupdate)\n \t\t\tstring_list_insert(update, path);\n \t\telse\n-\t\t\tfprintf(stderr,\n-\t\t\t\t\"Resolved '%s' using previous resolution.\\n\",\n-\t\t\t\tpath);\n+\t\t\tfprintf_ln(stderr,\n+\t\t\t\t   _(\"Resolved '%s' using previous resolution.\"),\n+\t\t\t\t   path);\n \t\tfree_rerere_id(rr_item);\n \t\trr_item->util = NULL;\n \t\treturn;\n@@ -802,11 +802,11 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \tif (id->collection->status[variant] & RR_HAS_POSTIMAGE) {\n \t\tconst char *path = rerere_path(id, \"postimage\");\n \t\tif (unlink(path))\n-\t\t\tdie_errno(\"cannot unlink stray '%s'\", path);\n+\t\t\tdie_errno(_(\"cannot unlink stray '%s'\"), path);\n \t\tid->collection->status[variant] &= ~RR_HAS_POSTIMAGE;\n \t}\n \tid->collection->status[variant] |= RR_HAS_PREIMAGE;\n-\tfprintf(stderr, \"Recorded preimage for '%s'\\n\", path);\n+\tfprintf_ln(stderr, _(\"Recorded preimage for '%s'\"), path);\n }\n \n static int do_plain_rerere(struct string_list *rr, int fd)\n@@ -878,7 +878,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"could not create directory '%s'\", git_path_rr_cache());\n+\t\tdie(_(\"could not create directory '%s'\"), git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1031,7 +1031,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t */\n \tret = handle_cache(path, sha1, NULL);\n \tif (ret < 1)\n-\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n \n \t/* Nuke the recorded resolution for the conflict */\n \tid = new_rerere_id(sha1);\n@@ -1049,7 +1049,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t\thandle_cache(path, sha1, rerere_path(id, \"thisimage\"));\n \t\tif (read_mmfile(&cur, rerere_path(id, \"thisimage\"))) {\n \t\t\tfree(cur.ptr);\n-\t\t\terror(\"failed to update conflicted state in '%s'\", path);\n+\t\t\terror(_(\"failed to update conflicted state in '%s'\"), path);\n \t\t\tgoto fail_exit;\n \t\t}\n \t\tcleanly_resolved = !try_merge(id, path, &cur, &result);\n@@ -1060,16 +1060,16 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t}\n \n \tif (id->collection->status_nr <= id->variant) {\n-\t\terror(\"no remembered resolution for '%s'\", path);\n+\t\terror(_(\"no remembered resolution for '%s'\"), path);\n \t\tgoto fail_exit;\n \t}\n \n \tfilename = rerere_path(id, \"postimage\");\n \tif (unlink(filename)) {\n \t\tif (errno == ENOENT)\n-\t\t\terror(\"no remembered resolution for '%s'\", path);\n+\t\t\terror(_(\"no remembered resolution for '%s'\"), path);\n \t\telse\n-\t\t\terror_errno(\"cannot unlink '%s'\", filename);\n+\t\t\terror_errno(_(\"cannot unlink '%s'\"), filename);\n \t\tgoto fail_exit;\n \t}\n \n@@ -1079,7 +1079,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t * the postimage.\n \t */\n \thandle_cache(path, sha1, rerere_path(id, \"preimage\"));\n-\tfprintf(stderr, \"Updated preimage for '%s'\\n\", path);\n+\tfprintf_ln(stderr, _(\"Updated preimage for '%s'\"), path);\n \n \t/*\n \t * And remember that we can record resolution for this\n@@ -1088,7 +1088,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \titem = string_list_insert(rr, path);\n \tfree_rerere_id(item);\n \titem->util = id;\n-\tfprintf(stderr, \"Forgot resolution for '%s'\\n\", path);\n+\tfprintf(stderr, _(\"Forgot resolution for '%s'\\n\"), path);\n \treturn 0;\n \n fail_exit:\n@@ -1103,7 +1103,7 @@ int rerere_forget(struct pathspec *pathspec)\n \tstruct string_list merge_rr = STRING_LIST_INIT_DUP;\n \n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfd = setup_rerere(&merge_rr, RERERE_NOAUTOUPDATE);\n \tif (fd < 0)\n@@ -1191,7 +1191,7 @@ void rerere_gc(struct string_list *rr)\n \tgit_config(git_default_config, NULL);\n \tdir = opendir(git_path(\"rr-cache\"));\n \tif (!dir)\n-\t\tdie_errno(\"unable to open rr-cache directory\");\n+\t\tdie_errno(_(\"unable to open rr-cache directory\"));\n \t/* Collect stale conflict IDs ... */\n \twhile ((e = readdir(dir))) {\n \t\tstruct rerere_dir *rr_dir;\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352566","messageId":"20180714214443.7184-7-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 06/11] rerere: fix crash when conflict goes unresolved","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:38Z","receivedAt":"2018-07-14T21:45:02Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently when a user doesn't resolve a conflict in a file, but\ncommits the file with the conflict markers, and later the file ends up\nin a state in which rerere can't handle it, subsequent rerere\noperations that are interested in that path, such as 'rerere clear' or\n'rerere forget <path>' will fail, or even worse in the case of 'rerere\nclear' segfault.\n\nSuch states include nested conflicts, or an extra conflict marker that\ndoesn't have any match.\n\nThis is because the first 'git rerere' when there was only one\nconflict in the file leaves an entry in the MERGE_RR file behind.  The\nnext 'git rerere' will then pick the rerere ID for that file up, and\nnot assign a new ID as it can't successfully calculate one.  It will\nhowever still try to do the rerere operation, because of the existing\nID.  As the handle_file function fails, it will remove the 'preimage'\nfor the ID in the process, while leaving the ID in the MERGE_RR file.\n\nNow when 'rerere clear' for example is run, it will segfault in\n'has_rerere_resolution', because status is NULL.\n\nTo fix this, remove the rerere ID from the MERGE_RR file in the case\nwhen we can't handle it, and remove the corresponding variant from\n.git/rr-cache/.  Removing it unconditionally is fine here, because if\nthe user would have resolved the conflict and ran rerere, the entry\nwould no longer be in the MERGE_RR file, so we wouldn't have this\nproblem in the first place, while if the conflict was not resolved,\nthe only thing that's left in the folder is the 'preimage', which by\nitself will be regenerated by git if necessary, so the user won't\nloose any work.\n\nNote that other variants that have the same conflict ID will not be\ntouched.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c          | 12 +++++++-----\n t/t4200-rerere.sh | 22 ++++++++++++++++++++++\n 2 files changed, 29 insertions(+), 5 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex da1ab54027..895ad80c0c 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -823,10 +823,7 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\tstruct rerere_id *id;\n \t\tunsigned char sha1[20];\n \t\tconst char *path = conflict.items[i].string;\n-\t\tint ret;\n-\n-\t\tif (string_list_has_string(rr, path))\n-\t\t\tcontinue;\n+\t\tint ret, has_string;\n \n \t\t/*\n \t\t * Ask handle_file() to scan and assign a\n@@ -834,7 +831,12 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\t * yet.\n \t\t */\n \t\tret = handle_file(path, sha1, NULL);\n-\t\tif (ret < 1)\n+\t\thas_string = string_list_has_string(rr, path);\n+\t\tif (ret < 0 && has_string) {\n+\t\t\tremove_variant(string_list_lookup(rr, path)->util);\n+\t\t\tstring_list_remove(rr, path, 1);\n+\t\t}\n+\t\tif (ret < 1 || has_string)\n \t\t\tcontinue;\n \n \t\tid = new_rerere_id(sha1);\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex 8417e5a4b1..34f0518a5e 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -580,4 +580,26 @@ test_expect_success 'multiple identical conflicts' '\n \tcount_pre_post 0 0\n '\n \n+test_expect_success 'rerere with unexpected conflict markers does not crash' '\n+\tgit reset --hard &&\n+\n+\tgit checkout -b branch-1 master &&\n+\techo \"bar\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m two &&\n+\n+\tgit reset --hard &&\n+\tgit checkout -b branch-2 master &&\n+\techo \"foo\" >test &&\n+\tgit add test &&\n+\tgit commit -q -a -m one &&\n+\n+\ttest_must_fail git merge branch-1 &&\n+\tsed \"s/bar/>>>>>>> a/\" >test.tmp <test &&\n+\tmv test.tmp test &&\n+\tgit rerere &&\n+\n+\tgit rerere clear\n+'\n+\n test_done\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352567","messageId":"20180714214443.7184-6-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 05/11] rerere: add documentation for conflict normalization","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:37Z","receivedAt":"2018-07-14T21:45:02Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add some documentation for the logic behind the conflict normalization\nin rerere.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Documentation/technical/rerere.txt | 140 +++++++++++++++++++++++++++++\n rerere.c                           |   4 -\n 2 files changed, 140 insertions(+), 4 deletions(-)\n create mode 100644 Documentation/technical/rerere.txt\n\ndiff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\nnew file mode 100644\nindex 0000000000..4102cce7aa\n--- /dev/null\n+++ b/Documentation/technical/rerere.txt\n@@ -0,0 +1,140 @@\n+Rerere\n+======\n+\n+This document describes the rerere logic.\n+\n+Conflict normalization\n+----------------------\n+\n+To ensure recorded conflict resolutions can be looked up in the rerere\n+database, even when branches are merged in a different order,\n+different branches are merged that result in the same conflict, or\n+when different conflict style settings are used, rerere normalizes the\n+conflicts before writing them to the rerere database.\n+\n+Different conflict styles and branch names are normalized by stripping\n+the labels from the conflict markers, and removing extraneous\n+information from the `diff3` conflict style. Branches that are merged\n+in different order are normalized by sorting the conflict hunks.  More\n+on each of those steps in the following sections.\n+\n+Once these two normalization operations are applied, a conflict ID is\n+calculated based on the normalized conflict, which is later used by\n+rerere to look up the conflict in the rerere database.\n+\n+Stripping extraneous information\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+\n+Say we have three branches AB, AC and AC2.  The common ancestor of\n+these branches has a file with a line containing the string \"A\" (for\n+brevity this is called \"line A\" in the rest of the document).  In\n+branch AB this line is changed to \"B\", in AC, this line is changed to\n+\"C\", and branch AC2 is forked off of AC, after the line was changed to\n+\"C\".\n+\n+Forking a branch ABAC off of branch AB and then merging AC into it, we\n+get a conflict like the following:\n+\n+    <<<<<<< HEAD\n+    B\n+    =======\n+    C\n+    >>>>>>> AC\n+\n+Doing the analogous with AC2 (forking a branch ABAC2 off of branch AB\n+and then merging branch AC2 into it), using the diff3 conflict style,\n+we get a conflict like the following:\n+\n+    <<<<<<< HEAD\n+    B\n+    ||||||| merged common ancestors\n+    A\n+    =======\n+    C\n+    >>>>>>> AC2\n+\n+By resolving this conflict, to leave line D, the user declares:\n+\n+    After examining what branches AB and AC did, I believe that making\n+    line A into line D is the best thing to do that is compatible with\n+    what AB and AC wanted to do.\n+\n+As branch AC2 refers to the same commit as AC, the above implies that\n+this is also compatible what AB and AC2 wanted to do.\n+\n+By extension, this means that rerere should recognize that the above\n+conflicts are the same.  To do this, the labels on the conflict\n+markers are stripped, and the diff3 output is removed.  The above\n+examples would both result in the following normalized conflict:\n+\n+    <<<<<<<\n+    B\n+    =======\n+    C\n+    >>>>>>>\n+\n+Sorting hunks\n+~~~~~~~~~~~~~\n+\n+As before, lets imagine that a common ancestor had a file with line A\n+its early part, and line X in its late part.  And then four branches\n+are forked that do these things:\n+\n+    - AB: changes A to B\n+    - AC: changes A to C\n+    - XY: changes X to Y\n+    - XZ: changes X to Z\n+\n+Now, forking a branch ABAC off of branch AB and then merging AC into\n+it, and forking a branch ACAB off of branch AC and then merging AB\n+into it, would yield the conflict in a different order.  The former\n+would say \"A became B or C, what now?\" while the latter would say \"A\n+became C or B, what now?\"\n+\n+As a reminder, the act of merging AC into ABAC and resolving the\n+conflict to leave line D means that the user declares:\n+\n+    After examining what branches AB and AC did, I believe that\n+    making line A into line D is the best thing to do that is\n+    compatible with what AB and AC wanted to do.\n+\n+So the conflict we would see when merging AB into ACAB should be\n+resolved the same way---it is the resolution that is in line with that\n+declaration.\n+\n+Imagine that similarly previously a branch XYXZ was forked from XY,\n+and XZ was merged into it, and resolved \"X became Y or Z\" into \"X\n+became W\".\n+\n+Now, if a branch ABXY was forked from AB and then merged XY, then ABXY\n+would have line B in its early part and line Y in its later part.\n+Such a merge would be quite clean.  We can construct 4 combinations\n+using these four branches ((AB, AC) x (XY, XZ)).\n+\n+Merging ABXY and ACXZ would make \"an early A became B or C, a late X\n+became Y or Z\" conflict, while merging ACXY and ABXZ would make \"an\n+early A became C or B, a late X became Y or Z\".  We can see there are\n+4 combinations of (\"B or C\", \"C or B\") x (\"X or Y\", \"Y or X\").\n+\n+By sorting, the conflict is given its canonical name, namely, \"an\n+early part became B or C, a late part becames X or Y\", and whenever\n+any of these four patterns appear, and we can get to the same conflict\n+and resolution that we saw earlier.\n+\n+Without the sorting, we'd have to somehow find a previous resolution\n+from combinatorial explosion.\n+\n+Conflict ID calculation\n+~~~~~~~~~~~~~~~~~~~~~~~\n+\n+Once the conflict normalization is done, the conflict ID is calculated\n+as the sha1 hash of the conflict hunks appended to each other,\n+separated by <NUL> characters.  The conflict markers are stripped out\n+before the sha1 is calculated.  So in the example above, where we\n+merge branch AC which changes line A to line C, into branch AB, which\n+changes line A to line C, the conflict ID would be\n+SHA1('B<NUL>C<NUL>').\n+\n+If there are multiple conflicts in one file, the sha1 is calculated\n+the same way with all hunks appended to each other, in the order in\n+which they appear in the file, separated by a <NUL> character.\ndiff --git a/rerere.c b/rerere.c\nindex be98c0afcb..da1ab54027 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -394,10 +394,6 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n  * and NUL concatenated together.\n  *\n  * Return the number of conflict hunks found.\n- *\n- * NEEDSWORK: the logic and theory of operation behind this conflict\n- * normalization may deserve to be documented somewhere, perhaps in\n- * Documentation/technical/rerere.txt.\n  */\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352568","messageId":"20180714214443.7184-8-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 07/11] rerere: only return whether a path has conflicts or not","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:39Z","receivedAt":"2018-07-14T21:45:04Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"We currently return the exact number of conflict hunks a certain path\nhas from the 'handle_paths' function.  However all of its callers only\ncare whether there are conflicts or not or if there is an error.\nReturn only that information, and document that only that information\nis returned.  This will simplify the code in the subsequent steps.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 23 ++++++++++++-----------\n 1 file changed, 12 insertions(+), 11 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 895ad80c0c..bf803043e2 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -393,12 +393,13 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n  * one side of the conflict, NUL, the other side of the conflict,\n  * and NUL concatenated together.\n  *\n- * Return the number of conflict hunks found.\n+ * Return 1 if conflict hunks are found, 0 if there are no conflict\n+ * hunks and -1 if an error occured.\n  */\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n \tgit_SHA_CTX ctx;\n-\tint hunk_no = 0;\n+\tint has_conflicts = 0;\n \tenum {\n \t\tRR_CONTEXT = 0, RR_SIDE_1, RR_SIDE_2, RR_ORIGINAL\n \t} hunk = RR_CONTEXT;\n@@ -426,7 +427,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\t\t\tgoto bad;\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n-\t\t\thunk_no++;\n+\t\t\thas_conflicts = 1;\n \t\t\thunk = RR_CONTEXT;\n \t\t\trerere_io_putconflict('<', marker_size, io);\n \t\t\trerere_io_putmem(one.buf, one.len, io);\n@@ -462,7 +463,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\tgit_SHA1_Final(sha1, &ctx);\n \tif (hunk != RR_CONTEXT)\n \t\treturn -1;\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n /*\n@@ -471,7 +472,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n  */\n static int handle_file(const char *path, unsigned char *sha1, const char *output)\n {\n-\tint hunk_no = 0;\n+\tint has_conflicts = 0;\n \tstruct rerere_io_file io;\n \tint marker_size = ll_merge_marker_size(path);\n \n@@ -491,7 +492,7 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \t\t}\n \t}\n \n-\thunk_no = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n+\thas_conflicts = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n@@ -500,14 +501,14 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tif (io.io.output && fclose(io.io.output))\n \t\tio.io.wrerror = error_errno(_(\"failed to flush '%s'\"), path);\n \n-\tif (hunk_no < 0) {\n+\tif (has_conflicts < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n \t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n /*\n@@ -954,7 +955,7 @@ static int handle_cache(const char *path, unsigned char *sha1, const char *outpu\n \tmmfile_t mmfile[3] = {{NULL}};\n \tmmbuffer_t result = {NULL, 0};\n \tconst struct cache_entry *ce;\n-\tint pos, len, i, hunk_no;\n+\tint pos, len, i, has_conflicts;\n \tstruct rerere_io_mem io;\n \tint marker_size = ll_merge_marker_size(path);\n \n@@ -1008,11 +1009,11 @@ static int handle_cache(const char *path, unsigned char *sha1, const char *outpu\n \t * Grab the conflict ID and optionally write the original\n \t * contents with conflict markers out.\n \t */\n-\thunk_no = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n+\thas_conflicts = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n \tstrbuf_release(&io.input);\n \tif (io.io.output)\n \t\tfclose(io.io.output);\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n static int rerere_forget_one_path(const char *path, struct string_list *rr)\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352569","messageId":"20180714214443.7184-9-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 08/11] rerere: factor out handle_conflict function","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:40Z","receivedAt":"2018-07-14T21:45:05Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Factor out the handle_conflict function, which handles a single\nconflict in a path.  This is in preparation for a subsequent commit,\nwhere this function will be re-used.  No functional changes intended.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 87 ++++++++++++++++++++++++++++++--------------------------\n 1 file changed, 47 insertions(+), 40 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex bf803043e2..2d62251943 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -384,85 +384,92 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n \treturn isspace(*buf);\n }\n \n-/*\n- * Read contents a file with conflicts, normalize the conflicts\n- * by (1) discarding the common ancestor version in diff3-style,\n- * (2) reordering our side and their side so that whichever sorts\n- * alphabetically earlier comes before the other one, while\n- * computing the \"conflict ID\", which is just an SHA-1 hash of\n- * one side of the conflict, NUL, the other side of the conflict,\n- * and NUL concatenated together.\n- *\n- * Return 1 if conflict hunks are found, 0 if there are no conflict\n- * hunks and -1 if an error occured.\n- */\n-static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n+static int handle_conflict(struct rerere_io *io, int marker_size, git_SHA_CTX *ctx)\n {\n-\tgit_SHA_CTX ctx;\n-\tint has_conflicts = 0;\n \tenum {\n-\t\tRR_CONTEXT = 0, RR_SIDE_1, RR_SIDE_2, RR_ORIGINAL\n-\t} hunk = RR_CONTEXT;\n+\t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n+\t} hunk = RR_SIDE_1;\n \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n \tstruct strbuf buf = STRBUF_INIT;\n-\n-\tif (sha1)\n-\t\tgit_SHA1_Init(&ctx);\n+\tint has_conflicts = -1;\n \n \twhile (!io->getline(&buf, io)) {\n \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n-\t\t\tif (hunk != RR_CONTEXT)\n-\t\t\t\tgoto bad;\n-\t\t\thunk = RR_SIDE_1;\n+\t\t\tbreak;\n \t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1)\n-\t\t\t\tgoto bad;\n+\t\t\t\tbreak;\n \t\t\thunk = RR_ORIGINAL;\n \t\t} else if (is_cmarker(buf.buf, '=', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1 && hunk != RR_ORIGINAL)\n-\t\t\t\tgoto bad;\n+\t\t\t\tbreak;\n \t\t\thunk = RR_SIDE_2;\n \t\t} else if (is_cmarker(buf.buf, '>', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_2)\n-\t\t\t\tgoto bad;\n+\t\t\t\tbreak;\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n \t\t\thas_conflicts = 1;\n-\t\t\thunk = RR_CONTEXT;\n \t\t\trerere_io_putconflict('<', marker_size, io);\n \t\t\trerere_io_putmem(one.buf, one.len, io);\n \t\t\trerere_io_putconflict('=', marker_size, io);\n \t\t\trerere_io_putmem(two.buf, two.len, io);\n \t\t\trerere_io_putconflict('>', marker_size, io);\n-\t\t\tif (sha1) {\n-\t\t\t\tgit_SHA1_Update(&ctx, one.buf ? one.buf : \"\",\n+\t\t\tif (ctx) {\n+\t\t\t\tgit_SHA1_Update(ctx, one.buf ? one.buf : \"\",\n \t\t\t\t\t    one.len + 1);\n-\t\t\t\tgit_SHA1_Update(&ctx, two.buf ? two.buf : \"\",\n+\t\t\t\tgit_SHA1_Update(ctx, two.buf ? two.buf : \"\",\n \t\t\t\t\t    two.len + 1);\n \t\t\t}\n-\t\t\tstrbuf_reset(&one);\n-\t\t\tstrbuf_reset(&two);\n+\t\t\tbreak;\n \t\t} else if (hunk == RR_SIDE_1)\n \t\t\tstrbuf_addbuf(&one, &buf);\n \t\telse if (hunk == RR_ORIGINAL)\n \t\t\t; /* discard */\n \t\telse if (hunk == RR_SIDE_2)\n \t\t\tstrbuf_addbuf(&two, &buf);\n-\t\telse\n-\t\t\trerere_io_putstr(buf.buf, io);\n-\t\tcontinue;\n-\tbad:\n-\t\thunk = 99; /* force error exit */\n-\t\tbreak;\n \t}\n \tstrbuf_release(&one);\n \tstrbuf_release(&two);\n \tstrbuf_release(&buf);\n \n+\treturn has_conflicts;\n+}\n+\n+/*\n+ * Read contents a file with conflicts, normalize the conflicts\n+ * by (1) discarding the common ancestor version in diff3-style,\n+ * (2) reordering our side and their side so that whichever sorts\n+ * alphabetically earlier comes before the other one, while\n+ * computing the \"conflict ID\", which is just an SHA-1 hash of\n+ * one side of the conflict, NUL, the other side of the conflict,\n+ * and NUL concatenated together.\n+ *\n+ * Return 1 if conflict hunks are found, 0 if there are no conflict\n+ * hunks and -1 if an error occured.\n+ */\n+static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n+{\n+\tgit_SHA_CTX ctx;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tint has_conflicts = 0;\n+\tif (sha1)\n+\t\tgit_SHA1_Init(&ctx);\n+\n+\twhile (!io->getline(&buf, io)) {\n+\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n+\t\t\thas_conflicts = handle_conflict(io, marker_size,\n+\t\t\t\t\t\t\tsha1 ? &ctx : NULL);\n+\t\t\tif (has_conflicts < 0)\n+\t\t\t\tbreak;\n+\t\t} else\n+\t\t\trerere_io_putstr(buf.buf, io);\n+\t}\n+\tstrbuf_release(&buf);\n+\n \tif (sha1)\n \t\tgit_SHA1_Final(sha1, &ctx);\n-\tif (hunk != RR_CONTEXT)\n-\t\treturn -1;\n+\n \treturn has_conflicts;\n }\n \n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352570","messageId":"20180714214443.7184-10-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 09/11] rerere: return strbuf from handle path","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:41Z","receivedAt":"2018-07-14T21:45:06Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently we write the conflict to disk directly in the handle_path\nfunction.  To make it re-usable for nested conflicts, instead of\nwriting the conflict out directly, store it in a strbuf and let the\ncaller write it out.\n\nThis does mean some slight increase in memory usage, however that\nincrease is limited to the size of the largest conflict we've\ncurrently processed.  We already keep one copy of the conflict in\nmemory, and it shouldn't be too large, so the increase in memory usage\nseems acceptable.\n\nAs a bonus this lets us get replace the rerere_io_putconflict function\nwith a trivial two line function.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 58 ++++++++++++++++++--------------------------------------\n 1 file changed, 18 insertions(+), 40 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 2d62251943..a35b88916c 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -302,38 +302,6 @@ static void rerere_io_putstr(const char *str, struct rerere_io *io)\n \t\tferr_puts(str, io->output, &io->wrerror);\n }\n \n-/*\n- * Write a conflict marker to io->output (if defined).\n- */\n-static void rerere_io_putconflict(int ch, int size, struct rerere_io *io)\n-{\n-\tchar buf[64];\n-\n-\twhile (size) {\n-\t\tif (size <= sizeof(buf) - 2) {\n-\t\t\tmemset(buf, ch, size);\n-\t\t\tbuf[size] = '\\n';\n-\t\t\tbuf[size + 1] = '\\0';\n-\t\t\tsize = 0;\n-\t\t} else {\n-\t\t\tint sz = sizeof(buf) - 1;\n-\n-\t\t\t/*\n-\t\t\t * Make sure we will not write everything out\n-\t\t\t * in this round by leaving at least 1 byte\n-\t\t\t * for the next round, giving the next round\n-\t\t\t * a chance to add the terminating LF.  Yuck.\n-\t\t\t */\n-\t\t\tif (size <= sz)\n-\t\t\t\tsz -= (sz - size) + 1;\n-\t\t\tmemset(buf, ch, sz);\n-\t\t\tbuf[sz] = '\\0';\n-\t\t\tsize -= sz;\n-\t\t}\n-\t\trerere_io_putstr(buf, io);\n-\t}\n-}\n-\n static void rerere_io_putmem(const char *mem, size_t sz, struct rerere_io *io)\n {\n \tif (io->output)\n@@ -384,7 +352,14 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n \treturn isspace(*buf);\n }\n \n-static int handle_conflict(struct rerere_io *io, int marker_size, git_SHA_CTX *ctx)\n+static void rerere_strbuf_putconflict(struct strbuf *buf, int ch, size_t size)\n+{\n+\tstrbuf_addchars(buf, ch, size);\n+\tstrbuf_addch(buf, '\\n');\n+}\n+\n+static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n+\t\t\t   int marker_size, git_SHA_CTX *ctx)\n {\n \tenum {\n \t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n@@ -410,11 +385,11 @@ static int handle_conflict(struct rerere_io *io, int marker_size, git_SHA_CTX *c\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n \t\t\thas_conflicts = 1;\n-\t\t\trerere_io_putconflict('<', marker_size, io);\n-\t\t\trerere_io_putmem(one.buf, one.len, io);\n-\t\t\trerere_io_putconflict('=', marker_size, io);\n-\t\t\trerere_io_putmem(two.buf, two.len, io);\n-\t\t\trerere_io_putconflict('>', marker_size, io);\n+\t\t\trerere_strbuf_putconflict(out, '<', marker_size);\n+\t\t\tstrbuf_addbuf(out, &one);\n+\t\t\trerere_strbuf_putconflict(out, '=', marker_size);\n+\t\t\tstrbuf_addbuf(out, &two);\n+\t\t\trerere_strbuf_putconflict(out, '>', marker_size);\n \t\t\tif (ctx) {\n \t\t\t\tgit_SHA1_Update(ctx, one.buf ? one.buf : \"\",\n \t\t\t\t\t    one.len + 1);\n@@ -451,21 +426,24 @@ static int handle_conflict(struct rerere_io *io, int marker_size, git_SHA_CTX *c\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n \tgit_SHA_CTX ctx;\n-\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf buf = STRBUF_INIT, out = STRBUF_INIT;\n \tint has_conflicts = 0;\n \tif (sha1)\n \t\tgit_SHA1_Init(&ctx);\n \n \twhile (!io->getline(&buf, io)) {\n \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n-\t\t\thas_conflicts = handle_conflict(io, marker_size,\n+\t\t\thas_conflicts = handle_conflict(&out, io, marker_size,\n \t\t\t\t\t\t\tsha1 ? &ctx : NULL);\n \t\t\tif (has_conflicts < 0)\n \t\t\t\tbreak;\n+\t\t\trerere_io_putmem(out.buf, out.len, io);\n+\t\t\tstrbuf_reset(&out);\n \t\t} else\n \t\t\trerere_io_putstr(buf.buf, io);\n \t}\n \tstrbuf_release(&buf);\n+\tstrbuf_release(&out);\n \n \tif (sha1)\n \t\tgit_SHA1_Final(sha1, &ctx);\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352571","messageId":"20180714214443.7184-11-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:42Z","receivedAt":"2018-07-14T21:45:07Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently rerere can't handle nested conflicts and will error out when\nit encounters such conflicts.  Do that by recursively calling the\n'handle_conflict' function to normalize the conflict.\n\nThe conflict ID calculation here deserves some explanation:\n\nAs we are using the same handle_conflict function, the nested conflict\nis normalized the same way as for non-nested conflicts, which means\nthe ancestor in the diff3 case is stripped out, and the parts of the\nconflict are ordered alphabetically.\n\nThe conflict ID is however is only calculated in the top level\nhandle_conflict call, so it will include the markers that 'rerere'\nadds to the output.  e.g. say there's the following conflict:\n\n    <<<<<<< HEAD\n    1\n    =======\n    <<<<<<< HEAD\n    3\n    =======\n    2\n    >>>>>>> branch-2\n    >>>>>>> branch-3~\n\nit would be recorde as follows in the preimage:\n\n    <<<<<<<\n    1\n    =======\n    <<<<<<<\n    2\n    =======\n    3\n    >>>>>>>\n    >>>>>>>\n\nand the conflict ID would be calculated as\n\n    sha1(1<NUL><<<<<<<\n    2\n    =======\n    3\n    >>>>>>><NUL>)\n\nStripping out vs. leaving the conflict markers in place in the inner\nconflict should have no practical impact, but it simplifies the\nimplementation.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Documentation/technical/rerere.txt | 42 ++++++++++++++++++++++++++++++\n rerere.c                           | 10 +++++--\n t/t4200-rerere.sh                  | 37 ++++++++++++++++++++++++++\n 3 files changed, 87 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\nindex 4102cce7aa..60d48dc4fe 100644\n--- a/Documentation/technical/rerere.txt\n+++ b/Documentation/technical/rerere.txt\n@@ -138,3 +138,45 @@ SHA1('B<NUL>C<NUL>').\n If there are multiple conflicts in one file, the sha1 is calculated\n the same way with all hunks appended to each other, in the order in\n which they appear in the file, separated by a <NUL> character.\n+\n+Nested conflicts\n+~~~~~~~~~~~~~~~~\n+\n+Nested conflicts are handled very similarly to \"simple\" conflicts.\n+Similar to simple conflicts, the conflict is first normalized by\n+stripping the labels from conflict markers, stripping the diff3\n+output, and the sorting the conflict hunks, both for the outer and the\n+inner conflict.  This is done recursively, so any number of nested\n+conflicts can be handled.\n+\n+The only difference is in how the conflict ID is calculated.  For the\n+inner conflict, the conflict markers themselves are not stripped out\n+before calculating the sha1.\n+\n+Say we have the following conflict for example:\n+\n+    <<<<<<< HEAD\n+    1\n+    =======\n+    <<<<<<< HEAD\n+    3\n+    =======\n+    2\n+    >>>>>>> branch-2\n+    >>>>>>> branch-3~\n+\n+After stripping out the labels of the conflict markers, and sorting\n+the hunks, the conflict would look as follows:\n+\n+    <<<<<<<\n+    1\n+    =======\n+    <<<<<<<\n+    2\n+    =======\n+    3\n+    >>>>>>>\n+    >>>>>>>\n+\n+and finally the conflict ID would be calculated as:\n+`sha1('1<NUL><<<<<<<\\n3\\n=======\\n2\\n>>>>>>><NUL>')`\ndiff --git a/rerere.c b/rerere.c\nindex a35b88916c..f78bef80b1 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -365,12 +365,18 @@ static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n \t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n \t} hunk = RR_SIDE_1;\n \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n-\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf buf = STRBUF_INIT, conflict = STRBUF_INIT;\n \tint has_conflicts = -1;\n \n \twhile (!io->getline(&buf, io)) {\n \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n-\t\t\tbreak;\n+\t\t\tif (handle_conflict(&conflict, io, marker_size, NULL) < 0)\n+\t\t\t\tbreak;\n+\t\t\tif (hunk == RR_SIDE_1)\n+\t\t\t\tstrbuf_addbuf(&one, &conflict);\n+\t\t\telse\n+\t\t\t\tstrbuf_addbuf(&two, &conflict);\n+\t\t\tstrbuf_release(&conflict);\n \t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1)\n \t\t\t\tbreak;\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex 34f0518a5e..d63fe2b33b 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -602,4 +602,41 @@ test_expect_success 'rerere with unexpected conflict markers does not crash' '\n \tgit rerere clear\n '\n \n+test_expect_success 'rerere with inner conflict markers' '\n+\tgit reset --hard &&\n+\n+\tgit checkout -b A master &&\n+\techo \"bar\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m two &&\n+\techo \"baz\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m three &&\n+\n+\tgit reset --hard &&\n+\tgit checkout -b B master &&\n+\techo \"foo\" >test &&\n+\tgit add test &&\n+\tgit commit -q -a -m one &&\n+\n+\ttest_must_fail git merge A~ &&\n+\tgit add test &&\n+\tgit commit -q -m \"will solve conflicts later\" &&\n+\ttest_must_fail git merge A &&\n+\n+\techo \"resolved\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m \"solved conflict\" &&\n+\n+\techo \"resolved\" >expect &&\n+\n+\tgit reset --hard HEAD~~ &&\n+\ttest_must_fail git merge A~ &&\n+\tgit add test &&\n+\tgit commit -q -m \"will solve conflicts later\" &&\n+\ttest_must_fail git merge A &&\n+\tcat test >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352572","messageId":"20180714214443.7184-12-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v3 11/11] rerere: recalculate conflict ID when unresolved conflict is committed","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-14T21:44:43Z","receivedAt":"2018-07-14T21:45:10Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently when a user doesn't resolve a conflict, commits the results,\nand does an operation which creates another conflict, rerere will use\nthe ID of the previously unresolved conflict for the new conflict.\nThis is because the conflict is kept in the MERGE_RR file, which\n'rerere' reads every time it is invoked.\n\nAfter the new conflict is solved, rerere will record the resolution\nwith the ID of the old conflict.  So in order to replay the conflict,\nboth merges would have to be re-done, instead of just the last one, in\norder for rerere to be able to automatically resolve the conflict.\n\nInstead of that, assign a new conflict ID if there are still conflicts\nin a file and the file had conflicts at a previous step.  This ID\nmatches the conflict we actually resolved at the corresponding step.\n\nNote that there are no backwards compatibility worries here, as rerere\nwould have failed to even normalize the conflict before this patch\nseries.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c          | 7 +++----\n t/t4200-rerere.sh | 7 +++++++\n 2 files changed, 10 insertions(+), 4 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex f78bef80b1..dd81d09e19 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -815,7 +815,7 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\tstruct rerere_id *id;\n \t\tunsigned char sha1[20];\n \t\tconst char *path = conflict.items[i].string;\n-\t\tint ret, has_string;\n+\t\tint ret;\n \n \t\t/*\n \t\t * Ask handle_file() to scan and assign a\n@@ -823,12 +823,11 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\t * yet.\n \t\t */\n \t\tret = handle_file(path, sha1, NULL);\n-\t\thas_string = string_list_has_string(rr, path);\n-\t\tif (ret < 0 && has_string) {\n+\t\tif (ret != 0 && string_list_has_string(rr, path)) {\n \t\t\tremove_variant(string_list_lookup(rr, path)->util);\n \t\t\tstring_list_remove(rr, path, 1);\n \t\t}\n-\t\tif (ret < 1 || has_string)\n+\t\tif (ret < 1)\n \t\t\tcontinue;\n \n \t\tid = new_rerere_id(sha1);\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex d63fe2b33b..bfb37ed4fc 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -636,6 +636,13 @@ test_expect_success 'rerere with inner conflict markers' '\n \tgit commit -q -m \"will solve conflicts later\" &&\n \ttest_must_fail git merge A &&\n \tcat test >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tgit add test &&\n+\tgit commit -m \"rerere solved conflict\" &&\n+\tgit reset --hard HEAD~ &&\n+\ttest_must_fail git merge A &&\n+\tcat test >actual &&\n \ttest_cmp expect actual\n '\n \n-- \n2.17.0.410.g65aef3a6c4\n\n"},{"id":"352577","messageId":"20180715132421.GA5015@ruderich.org","threadId":"48531","inReplyTo":"20180714214443.7184-5-t.gummerer@gmail.com","subject":"Re: [PATCH v3 04/11] rerere: mark strings for translation","fromName":"Simon Ruderich","fromEmail":"simon@ruderich.org","sentAt":"2018-07-15T13:24:21Z","receivedAt":"2018-07-15T13:28:07Z","isPatch":true,"sender":{"key":"simon@ruderich.org","avatar":"https://avatars.githubusercontent.com/u/390994?v=4"},"body":"On Sat, Jul 14, 2018 at 10:44:36PM +0100, Thomas Gummerer wrote:\n> 'git rerere' is considered a plumbing command and as such its output\n\ns/plumbing/porcelain/?\n\nRegards\nSimon\n-- \n+ privacy is necessary\n+ using gnupg http://gnupg.org\n+ public key id: 0x92FEFDB7E44C32F9\n"},{"id":"352673","messageId":"20180716204017.GB2186@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"20180715132421.GA5015@ruderich.org","subject":"Re: [PATCH v3 04/11] rerere: mark strings for translation","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-16T20:40:17Z","receivedAt":"2018-07-16T20:40:22Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 07/15, Simon Ruderich wrote:\n> On Sat, Jul 14, 2018 at 10:44:36PM +0100, Thomas Gummerer wrote:\n> > 'git rerere' is considered a plumbing command and as such its output\n> \n> s/plumbing/porcelain/?\n\nAh yes indeed.  Thanks for catching!\n\n> Regards\n> Simon\n> -- \n> + privacy is necessary\n> + using gnupg http://gnupg.org\n> + public key id: 0x92FEFDB7E44C32F9\n"},{"id":"353915","messageId":"xmqqzhy8hb2s.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180714214443.7184-11-t.gummerer@gmail.com","subject":"Re: [PATCH v3 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-30T17:45:15Z","receivedAt":"2018-07-30T17:45:20Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Currently rerere can't handle nested conflicts and will error out when\n> it encounters such conflicts.  Do that by recursively calling the\n> 'handle_conflict' function to normalize the conflict.\n>\n> The conflict ID calculation here deserves some explanation:\n>\n> As we are using the same handle_conflict function, the nested conflict\n> is normalized the same way as for non-nested conflicts, which means\n> the ancestor in the diff3 case is stripped out, and the parts of the\n> conflict are ordered alphabetically.\n>\n> The conflict ID is however is only calculated in the top level\n> handle_conflict call, so it will include the markers that 'rerere'\n> adds to the output.  e.g. say there's the following conflict:\n>\n>     <<<<<<< HEAD\n>     1\n>     =======\n>     <<<<<<< HEAD\n>     3\n>     =======\n>     2\n>     >>>>>>> branch-2\n>     >>>>>>> branch-3~\n\nHmph, I vaguely recall that I made inner merges to use the conflict\nmarkers automatically lengthened (by two, if I recall correctly)\nthan its immediate outer merge.  Wouldn't the above look more like\n\n     <<<<<<< HEAD\n     1\n     =======\n     <<<<<<<<< HEAD\n     3\n     =========\n     2\n     >>>>>>>>> branch-2\n     >>>>>>> branch-3~\n    \nPerhaps I am not recalling it correctly.\n\n> it would be recorde as follows in the preimage:\n>\n>     <<<<<<<\n>     1\n>     =======\n>     <<<<<<<\n>     2\n>     =======\n>     3\n>     >>>>>>>\n>     >>>>>>>\n>\n> and the conflict ID would be calculated as\n>\n>     sha1(1<NUL><<<<<<<\n>     2\n>     =======\n>     3\n>     >>>>>>><NUL>)\n>\n> Stripping out vs. leaving the conflict markers in place in the inner\n> conflict should have no practical impact, but it simplifies the\n> implementation.\n>\n> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> ---\n>  Documentation/technical/rerere.txt | 42 ++++++++++++++++++++++++++++++\n>  rerere.c                           | 10 +++++--\n>  t/t4200-rerere.sh                  | 37 ++++++++++++++++++++++++++\n>  3 files changed, 87 insertions(+), 2 deletions(-)\n>\n> diff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\n> index 4102cce7aa..60d48dc4fe 100644\n> --- a/Documentation/technical/rerere.txt\n> +++ b/Documentation/technical/rerere.txt\n> @@ -138,3 +138,45 @@ SHA1('B<NUL>C<NUL>').\n>  If there are multiple conflicts in one file, the sha1 is calculated\n>  the same way with all hunks appended to each other, in the order in\n>  which they appear in the file, separated by a <NUL> character.\n> +\n> +Nested conflicts\n> +~~~~~~~~~~~~~~~~\n> +\n> +Nested conflicts are handled very similarly to \"simple\" conflicts.\n> +Similar to simple conflicts, the conflict is first normalized by\n> +stripping the labels from conflict markers, stripping the diff3\n> +output, and the sorting the conflict hunks, both for the outer and the\n> +inner conflict.  This is done recursively, so any number of nested\n> +conflicts can be handled.\n> +\n> +The only difference is in how the conflict ID is calculated.  For the\n> +inner conflict, the conflict markers themselves are not stripped out\n> +before calculating the sha1.\n> +\n> +Say we have the following conflict for example:\n> +\n> +    <<<<<<< HEAD\n> +    1\n> +    =======\n> +    <<<<<<< HEAD\n> +    3\n> +    =======\n> +    2\n> +    >>>>>>> branch-2\n> +    >>>>>>> branch-3~\n> +\n> +After stripping out the labels of the conflict markers, and sorting\n> +the hunks, the conflict would look as follows:\n> +\n> +    <<<<<<<\n> +    1\n> +    =======\n> +    <<<<<<<\n> +    2\n> +    =======\n> +    3\n> +    >>>>>>>\n> +    >>>>>>>\n> +\n> +and finally the conflict ID would be calculated as:\n> +`sha1('1<NUL><<<<<<<\\n3\\n=======\\n2\\n>>>>>>><NUL>')`\n> diff --git a/rerere.c b/rerere.c\n> index a35b88916c..f78bef80b1 100644\n> --- a/rerere.c\n> +++ b/rerere.c\n> @@ -365,12 +365,18 @@ static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n>  \t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n>  \t} hunk = RR_SIDE_1;\n>  \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n> -\tstruct strbuf buf = STRBUF_INIT;\n> +\tstruct strbuf buf = STRBUF_INIT, conflict = STRBUF_INIT;\n>  \tint has_conflicts = -1;\n>  \n>  \twhile (!io->getline(&buf, io)) {\n>  \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n> -\t\t\tbreak;\n> +\t\t\tif (handle_conflict(&conflict, io, marker_size, NULL) < 0)\n> +\t\t\t\tbreak;\n> +\t\t\tif (hunk == RR_SIDE_1)\n> +\t\t\t\tstrbuf_addbuf(&one, &conflict);\n> +\t\t\telse\n> +\t\t\t\tstrbuf_addbuf(&two, &conflict);\n\nHmph, do we ever see the inner conflict block while we are skipping\nand ignoring the common ancestor version, or it is impossible that\nwe see '<' only while processing either our or their side?\n\n> +\t\t\tstrbuf_release(&conflict);\n>  \t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n>  \t\t\tif (hunk != RR_SIDE_1)\n>  \t\t\t\tbreak;\n> diff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\n> index 34f0518a5e..d63fe2b33b 100755\n> --- a/t/t4200-rerere.sh\n> +++ b/t/t4200-rerere.sh\n> @@ -602,4 +602,41 @@ test_expect_success 'rerere with unexpected conflict markers does not crash' '\n>  \tgit rerere clear\n>  '\n>  \n> +test_expect_success 'rerere with inner conflict markers' '\n> +\tgit reset --hard &&\n> +\n> +\tgit checkout -b A master &&\n> +\techo \"bar\" >test &&\n> +\tgit add test &&\n> +\tgit commit -q -m two &&\n> +\techo \"baz\" >test &&\n> +\tgit add test &&\n> +\tgit commit -q -m three &&\n> +\n> +\tgit reset --hard &&\n> +\tgit checkout -b B master &&\n> +\techo \"foo\" >test &&\n> +\tgit add test &&\n> +\tgit commit -q -a -m one &&\n> +\n> +\ttest_must_fail git merge A~ &&\n> +\tgit add test &&\n> +\tgit commit -q -m \"will solve conflicts later\" &&\n> +\ttest_must_fail git merge A &&\n> +\n> +\techo \"resolved\" >test &&\n> +\tgit add test &&\n> +\tgit commit -q -m \"solved conflict\" &&\n> +\n> +\techo \"resolved\" >expect &&\n> +\n> +\tgit reset --hard HEAD~~ &&\n> +\ttest_must_fail git merge A~ &&\n> +\tgit add test &&\n> +\tgit commit -q -m \"will solve conflicts later\" &&\n> +\ttest_must_fail git merge A &&\n> +\tcat test >actual &&\n> +\ttest_cmp expect actual\n> +'\n> +\n>  test_done\n"},{"id":"353916","messageId":"xmqqpnz4hau1.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"Re: [PATCH v3 00/11] rerere: handle nested conflicts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-30T17:50:30Z","receivedAt":"2018-07-30T17:50:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Thomas Gummerer (11):\n>   rerere: unify error messages when read_cache fails\n>   rerere: lowercase error messages\n>   rerere: wrap paths in output in sq\n>   rerere: mark strings for translation\n>   rerere: add documentation for conflict normalization\n>   rerere: fix crash when conflict goes unresolved\n>   rerere: only return whether a path has conflicts or not\n>   rerere: factor out handle_conflict function\n>   rerere: return strbuf from handle path\n>   rerere: teach rerere to handle nested conflicts\n>   rerere: recalculate conflict ID when unresolved conflict is committed\n\nEven though I am not certain about the last two steps, everything\nbefore them looked trivially correct and good changes (well, the\n\"strbuf\" one's goodness obviously depends on the goodness of the\nlast two, which are helped by it).\n\nSorry for taking so long before getting to the series.\n"},{"id":"353917","messageId":"xmqqin4whatr.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180714214443.7184-6-t.gummerer@gmail.com","subject":"Re: [PATCH v3 05/11] rerere: add documentation for conflict normalization","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-30T17:50:40Z","receivedAt":"2018-07-30T17:50:44Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> +Different conflict styles and branch names are normalized by stripping\n> +the labels from the conflict markers, and removing extraneous\n> +information from the `diff3` conflict style. Branches that are merged\n\ns/extraneous information/commmon ancestor version/ perhaps, to be\nfact-based without passing value judgment?\n\nWe drop the common ancestor version only because we cannot normalize\nfrom `merge` style to `diff3` style by adding one, and not because\nit is extraneous.  It does help humans understand the conflict a lot\nbetter to have that section.\n\n> +By extension, this means that rerere should recognize that the above\n> +conflicts are the same.  To do this, the labels on the conflict\n> +markers are stripped, and the diff3 output is removed.  The above\n\ns/diff3 output/common ancestor version/, as \"diff3 output\" would\nmean the whole thing between <<< and >>> to readers.\n\n> diff --git a/rerere.c b/rerere.c\n> index be98c0afcb..da1ab54027 100644\n> --- a/rerere.c\n> +++ b/rerere.c\n> @@ -394,10 +394,6 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n>   * and NUL concatenated together.\n>   *\n>   * Return the number of conflict hunks found.\n> - *\n> - * NEEDSWORK: the logic and theory of operation behind this conflict\n> - * normalization may deserve to be documented somewhere, perhaps in\n> - * Documentation/technical/rerere.txt.\n>   */\n>  static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n>  {\n\nThanks for finally removing this age-old NEEDSWORK comment.\n"},{"id":"353918","messageId":"xmqqbmaohath.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180714214443.7184-7-t.gummerer@gmail.com","subject":"Re: [PATCH v3 06/11] rerere: fix crash when conflict goes unresolved","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-30T17:50:50Z","receivedAt":"2018-07-30T17:50:55Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Currently when a user doesn't resolve a conflict in a file, but\n> commits the file with the conflict markers, and later the file ends up\n> in a state in which rerere can't handle it, subsequent rerere\n> operations that are interested in that path, such as 'rerere clear' or\n> 'rerere forget <path>' will fail, or even worse in the case of 'rerere\n> clear' segfault.\n>\n> Such states include nested conflicts, or an extra conflict marker that\n> doesn't have any match.\n>\n> This is because the first 'git rerere' when there was only one\n> conflict in the file leaves an entry in the MERGE_RR file behind.  The\n\nI find this sentence, especially the \"only one conflict in the file\"\npart, a bit unclear.  What does the sentence count as one conflict?\nOne block of lines enclosed inside \"<<<\"...\">>>\" pair?  The command\nbehaves differently when there are two such blocks instead?\n\n> next 'git rerere' will then pick the rerere ID for that file up, and\n> not assign a new ID as it can't successfully calculate one.  It will\n> however still try to do the rerere operation, because of the existing\n> ID.  As the handle_file function fails, it will remove the 'preimage'\n> for the ID in the process, while leaving the ID in the MERGE_RR file.\n>\n> Now when 'rerere clear' for example is run, it will segfault in\n> 'has_rerere_resolution', because status is NULL.\n\nI think this \"status\" refers to the collection->status[].  How do we\nget into that state, though?\n\nnew_rerere_id() and new_rerere_id_hex() fills id->collection by\ncalling find_rerere_dir(), which either finds an existing rerere_dir\ninstance or manufactures one with .status==NULL.  The .status[]\narray is later grown by calling fit_variant as we scan and find the\npre/post images, but because there is no pre/post image for a file\nwith unparseable conflicts, it is left NULL.\n\nSo another possible fix could be to make sure that .status[] is only\nread when .status_nr says there is something worth reading.  I am\nnot saying that would be a better fix---I am just thinking out loud\nto make sure I understand the issue correctly.\n\n> To fix this, remove the rerere ID from the MERGE_RR file in the case\n> when we can't handle it, and remove the corresponding variant from\n> .git/rr-cache/.  Removing it unconditionally is fine here, because if\n> the user would have resolved the conflict and ran rerere, the entry\n> would no longer be in the MERGE_RR file, so we wouldn't have this\n> problem in the first place, while if the conflict was not resolved,\n> the only thing that's left in the folder is the 'preimage', which by\n> itself will be regenerated by git if necessary, so the user won't\n> loose any work.\n\ns/loose/lose/\n\n> Note that other variants that have the same conflict ID will not be\n> touched.\n\nNice.  Thanks for a fix.\n\n>\n> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> ---\n>  rerere.c          | 12 +++++++-----\n>  t/t4200-rerere.sh | 22 ++++++++++++++++++++++\n>  2 files changed, 29 insertions(+), 5 deletions(-)\n>\n> diff --git a/rerere.c b/rerere.c\n> index da1ab54027..895ad80c0c 100644\n> --- a/rerere.c\n> +++ b/rerere.c\n> @@ -823,10 +823,7 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n>  \t\tstruct rerere_id *id;\n>  \t\tunsigned char sha1[20];\n>  \t\tconst char *path = conflict.items[i].string;\n> -\t\tint ret;\n> -\n> -\t\tif (string_list_has_string(rr, path))\n> -\t\t\tcontinue;\n> +\t\tint ret, has_string;\n>  \n>  \t\t/*\n>  \t\t * Ask handle_file() to scan and assign a\n> @@ -834,7 +831,12 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n>  \t\t * yet.\n>  \t\t */\n>  \t\tret = handle_file(path, sha1, NULL);\n> -\t\tif (ret < 1)\n> +\t\thas_string = string_list_has_string(rr, path);\n> +\t\tif (ret < 0 && has_string) {\n> +\t\t\tremove_variant(string_list_lookup(rr, path)->util);\n> +\t\t\tstring_list_remove(rr, path, 1);\n> +\t\t}\n> +\t\tif (ret < 1 || has_string)\n>  \t\t\tcontinue;\n\nWe used to say \"if we know about the path we do not do anything\nhere, if we do not see any conflict in the file we do nothing,\notherwise we assign a new id\"; we now say \"see if we can parse\nand also see if we have conflict(s); if we know about the path and\nwe cannot parse, drop it from the rr database (because otherwise the\nentry will cause us trouble elsewhere later).  Otherwise, if we do\nnot have any conflict or we already know about the path, no need to\ndo anything. Otherwise, i.e. a newly discovered path with conflicts\ngets a new id\".\n\nMakes sense.  \"A known path with unparseable conflict gets dropped\"\nis the important change in this hunk.\n\n> diff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\n> index 8417e5a4b1..34f0518a5e 100755\n> --- a/t/t4200-rerere.sh\n> +++ b/t/t4200-rerere.sh\n> @@ -580,4 +580,26 @@ test_expect_success 'multiple identical conflicts' '\n>  \tcount_pre_post 0 0\n>  '\n>  \n> +test_expect_success 'rerere with unexpected conflict markers does not crash' '\n> +\tgit reset --hard &&\n> +\n> +\tgit checkout -b branch-1 master &&\n> +\techo \"bar\" >test &&\n> +\tgit add test &&\n> +\tgit commit -q -m two &&\n> +\n> +\tgit reset --hard &&\n> +\tgit checkout -b branch-2 master &&\n> +\techo \"foo\" >test &&\n> +\tgit add test &&\n> +\tgit commit -q -a -m one &&\n> +\n> +\ttest_must_fail git merge branch-1 &&\n> +\tsed \"s/bar/>>>>>>> a/\" >test.tmp <test &&\n> +\tmv test.tmp test &&\n\nOK, so the \"only one conflict\" in the log message meant just one\nside of the conflict marker.  More generally, the troublesome is\nto have \"conflict marker(s) that cannot be parsed\" in the file.\n\n> +\tgit rerere &&\n> +\n> +\tgit rerere clear\n> +'\n> +\n>  test_done\n"},{"id":"353920","messageId":"xmqq4lgghatc.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180714214443.7184-8-t.gummerer@gmail.com","subject":"Re: [PATCH v3 07/11] rerere: only return whether a path has conflicts or not","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-30T17:50:55Z","receivedAt":"2018-07-30T17:50:59Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> We currently return the exact number of conflict hunks a certain path\n> has from the 'handle_paths' function.  However all of its callers only\n> care whether there are conflicts or not or if there is an error.\n> Return only that information, and document that only that information\n> is returned.  This will simplify the code in the subsequent steps.\n>\n> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> ---\n>  rerere.c | 23 ++++++++++++-----------\n>  1 file changed, 12 insertions(+), 11 deletions(-)\n\nI do recall writing this code without knowing if the actual number\nof conflicts would be useful by callers, but it is apparent that it\nwasn't.  I won't mind losing that bit of info at all.  Besides, we\nwon't risk mistaking a file with 2 billion conflicts with a file\nwhose conflicts cannot be parsed ;-).\n\nThe patch looks good.  Thanks.\n"},{"id":"353921","messageId":"xmqqwotcfw8r.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180714214443.7184-9-t.gummerer@gmail.com","subject":"Re: [PATCH v3 08/11] rerere: factor out handle_conflict function","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-30T17:51:00Z","receivedAt":"2018-07-30T17:51:04Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Factor out the handle_conflict function, which handles a single\n> conflict in a path.  This is in preparation for a subsequent commit,\n> where this function will be re-used.  No functional changes intended.\n>\n> Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> ---\n>  rerere.c | 87 ++++++++++++++++++++++++++++++--------------------------\n>  1 file changed, 47 insertions(+), 40 deletions(-)\n\nRenumbering of the enum made me raise my eyebrow a bit briefly but\nit is merely to keep track of the state locally and invisible from\nthe outside, so it is perfectly fine.\n\n> -\tgit_SHA_CTX ctx;\n> -\tint has_conflicts = 0;\n>  \tenum {\n> -\t\tRR_CONTEXT = 0, RR_SIDE_1, RR_SIDE_2, RR_ORIGINAL\n> -\t} hunk = RR_CONTEXT;\n> +\t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n> +\t} hunk = RR_SIDE_1;\n"},{"id":"353922","messageId":"xmqqpnz4fw7v.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180714214443.7184-10-t.gummerer@gmail.com","subject":"Re: [PATCH v3 09/11] rerere: return strbuf from handle path","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-07-30T17:51:32Z","receivedAt":"2018-07-30T17:51:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Currently we write the conflict to disk directly in the handle_path\n> function.  To make it re-usable for nested conflicts, instead of\n> writing the conflict out directly, store it in a strbuf and let the\n> caller write it out.\n>\n> This does mean some slight increase in memory usage, however that\n> increase is limited to the size of the largest conflict we've\n> currently processed.  We already keep one copy of the conflict in\n> memory, and it shouldn't be too large, so the increase in memory usage\n> seems acceptable.\n>\n> As a bonus this lets us get replace the rerere_io_putconflict function\n> with a trivial two line function.\n\nMakes sense.\n"},{"id":"353949","messageId":"20180730202037.GF9955@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqqzhy8hb2s.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-30T20:20:37Z","receivedAt":"2018-07-30T20:20:42Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 07/30, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > Currently rerere can't handle nested conflicts and will error out when\n> > it encounters such conflicts.  Do that by recursively calling the\n> > 'handle_conflict' function to normalize the conflict.\n> >\n> > The conflict ID calculation here deserves some explanation:\n> >\n> > As we are using the same handle_conflict function, the nested conflict\n> > is normalized the same way as for non-nested conflicts, which means\n> > the ancestor in the diff3 case is stripped out, and the parts of the\n> > conflict are ordered alphabetically.\n> >\n> > The conflict ID is however is only calculated in the top level\n> > handle_conflict call, so it will include the markers that 'rerere'\n> > adds to the output.  e.g. say there's the following conflict:\n> >\n> >     <<<<<<< HEAD\n> >     1\n> >     =======\n> >     <<<<<<< HEAD\n> >     3\n> >     =======\n> >     2\n> >     >>>>>>> branch-2\n> >     >>>>>>> branch-3~\n> \n> Hmph, I vaguely recall that I made inner merges to use the conflict\n> markers automatically lengthened (by two, if I recall correctly)\n> than its immediate outer merge.  Wouldn't the above look more like\n> \n>      <<<<<<< HEAD\n>      1\n>      =======\n>      <<<<<<<<< HEAD\n>      3\n>      =========\n>      2\n>      >>>>>>>>> branch-2\n>      >>>>>>> branch-3~\n>     \n> Perhaps I am not recalling it correctly.\n\nThe only way I could reproduce this is by not resolving a conflict\n(just leaving the conflict markers in place, but running 'git add\nconflicted'), and then merging something else, which produces another\nconflict, where one of the sides was the one with conflict markers\nalready in the file, same as what I did in the test.\n\nSo in that case, the conflict markers of the already existing conflict\nwould just be treated as normal text during the merge I believe, and\nthus the new conflict markers would be the same length.\n\nThe usage of git is really a bit wrong here, so I don't know if it's\nactually worth helping the users at this point.  But trying to\nunderstand how rerere exactly works, I had this written up already, so\nI thought I would include it in this series anyway in case it helps\nsomebody :)\n\n> > it would be recorde as follows in the preimage:\n> >\n> >     <<<<<<<\n> >     1\n> >     =======\n> >     <<<<<<<\n> >     2\n> >     =======\n> >     3\n> >     >>>>>>>\n> >     >>>>>>>\n> >\n> > and the conflict ID would be calculated as\n> >\n> >     sha1(1<NUL><<<<<<<\n> >     2\n> >     =======\n> >     3\n> >     >>>>>>><NUL>)\n> >\n> > Stripping out vs. leaving the conflict markers in place in the inner\n> > conflict should have no practical impact, but it simplifies the\n> > implementation.\n> >\n> > Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> > ---\n> >  Documentation/technical/rerere.txt | 42 ++++++++++++++++++++++++++++++\n> >  rerere.c                           | 10 +++++--\n> >  t/t4200-rerere.sh                  | 37 ++++++++++++++++++++++++++\n> >  3 files changed, 87 insertions(+), 2 deletions(-)\n> >\n> > [..snip..]\n> > \n> > diff --git a/rerere.c b/rerere.c\n> > index a35b88916c..f78bef80b1 100644\n> > --- a/rerere.c\n> > +++ b/rerere.c\n> > @@ -365,12 +365,18 @@ static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n> >  \t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n> >  \t} hunk = RR_SIDE_1;\n> >  \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n> > -\tstruct strbuf buf = STRBUF_INIT;\n> > +\tstruct strbuf buf = STRBUF_INIT, conflict = STRBUF_INIT;\n> >  \tint has_conflicts = -1;\n> >  \n> >  \twhile (!io->getline(&buf, io)) {\n> >  \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n> > -\t\t\tbreak;\n> > +\t\t\tif (handle_conflict(&conflict, io, marker_size, NULL) < 0)\n> > +\t\t\t\tbreak;\n> > +\t\t\tif (hunk == RR_SIDE_1)\n> > +\t\t\t\tstrbuf_addbuf(&one, &conflict);\n> > +\t\t\telse\n> > +\t\t\t\tstrbuf_addbuf(&two, &conflict);\n> \n> Hmph, do we ever see the inner conflict block while we are skipping\n> and ignoring the common ancestor version, or it is impossible that\n> we see '<' only while processing either our or their side?\n\nAs mentioned above, I haven't been able to reproduce creating an inner\nconflict block outside of the case mentioned above, where the user\ncommitted conflict markers, and then did another merge.\n\nI don't think it can appear outside of that case in \"normal\"\noperation.\n\n> > +\t\t\tstrbuf_release(&conflict);\n> >  \t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n> >  \t\t\tif (hunk != RR_SIDE_1)\n> >  \t\t\t\tbreak;\n> > diff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\n> > index 34f0518a5e..d63fe2b33b 100755\n> > --- a/t/t4200-rerere.sh\n> > +++ b/t/t4200-rerere.sh\n> > @@ -602,4 +602,41 @@ test_expect_success 'rerere with unexpected conflict markers does not crash' '\n> >  \tgit rerere clear\n> >  '\n> >  \n> > +test_expect_success 'rerere with inner conflict markers' '\n> > +\tgit reset --hard &&\n> > +\n> > +\tgit checkout -b A master &&\n> > +\techo \"bar\" >test &&\n> > +\tgit add test &&\n> > +\tgit commit -q -m two &&\n> > +\techo \"baz\" >test &&\n> > +\tgit add test &&\n> > +\tgit commit -q -m three &&\n> > +\n> > +\tgit reset --hard &&\n> > +\tgit checkout -b B master &&\n> > +\techo \"foo\" >test &&\n> > +\tgit add test &&\n> > +\tgit commit -q -a -m one &&\n> > +\n> > +\ttest_must_fail git merge A~ &&\n> > +\tgit add test &&\n> > +\tgit commit -q -m \"will solve conflicts later\" &&\n> > +\ttest_must_fail git merge A &&\n> > +\n> > +\techo \"resolved\" >test &&\n> > +\tgit add test &&\n> > +\tgit commit -q -m \"solved conflict\" &&\n> > +\n> > +\techo \"resolved\" >expect &&\n> > +\n> > +\tgit reset --hard HEAD~~ &&\n> > +\ttest_must_fail git merge A~ &&\n> > +\tgit add test &&\n> > +\tgit commit -q -m \"will solve conflicts later\" &&\n> > +\ttest_must_fail git merge A &&\n> > +\tcat test >actual &&\n> > +\ttest_cmp expect actual\n> > +'\n> > +\n> >  test_done\n"},{"id":"353950","messageId":"20180730202140.GG9955@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqqin4whatr.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 05/11] rerere: add documentation for conflict normalization","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-30T20:21:40Z","receivedAt":"2018-07-30T20:21:44Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 07/30, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > +Different conflict styles and branch names are normalized by stripping\n> > +the labels from the conflict markers, and removing extraneous\n> > +information from the `diff3` conflict style. Branches that are merged\n> \n> s/extraneous information/commmon ancestor version/ perhaps, to be\n> fact-based without passing value judgment?\n\nYeah I meant \"extraneous information for rerere\", but common ancester\nversion is better.\n\n> We drop the common ancestor version only because we cannot normalize\n> from `merge` style to `diff3` style by adding one, and not because\n> it is extraneous.  It does help humans understand the conflict a lot\n> better to have that section.\n> \n> > +By extension, this means that rerere should recognize that the above\n> > +conflicts are the same.  To do this, the labels on the conflict\n> > +markers are stripped, and the diff3 output is removed.  The above\n> \n> s/diff3 output/common ancestor version/, as \"diff3 output\" would\n> mean the whole thing between <<< and >>> to readers.\n\nMakes sense, will fix in the re-roll, thanks!\n\n> > diff --git a/rerere.c b/rerere.c\n> > index be98c0afcb..da1ab54027 100644\n> > --- a/rerere.c\n> > +++ b/rerere.c\n> > @@ -394,10 +394,6 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n> >   * and NUL concatenated together.\n> >   *\n> >   * Return the number of conflict hunks found.\n> > - *\n> > - * NEEDSWORK: the logic and theory of operation behind this conflict\n> > - * normalization may deserve to be documented somewhere, perhaps in\n> > - * Documentation/technical/rerere.txt.\n> >   */\n> >  static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n> >  {\n> \n> Thanks for finally removing this age-old NEEDSWORK comment.\n"},{"id":"353952","messageId":"20180730204545.GH9955@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqqbmaohath.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 06/11] rerere: fix crash when conflict goes unresolved","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-30T20:45:45Z","receivedAt":"2018-07-30T20:45:50Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 07/30, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > Currently when a user doesn't resolve a conflict in a file, but\n> > commits the file with the conflict markers, and later the file ends up\n> > in a state in which rerere can't handle it, subsequent rerere\n> > operations that are interested in that path, such as 'rerere clear' or\n> > 'rerere forget <path>' will fail, or even worse in the case of 'rerere\n> > clear' segfault.\n> >\n> > Such states include nested conflicts, or an extra conflict marker that\n> > doesn't have any match.\n> >\n> > This is because the first 'git rerere' when there was only one\n> > conflict in the file leaves an entry in the MERGE_RR file behind.  The\n> \n> I find this sentence, especially the \"only one conflict in the file\"\n> part, a bit unclear.  What does the sentence count as one conflict?\n> One block of lines enclosed inside \"<<<\"...\">>>\" pair?  The command\n> behaves differently when there are two such blocks instead?\n\nYeah as you mentioned below, conflict marker(s) that cannot be parsed\nhere would make more sense.  Will adjust the commit message.\n\n> > next 'git rerere' will then pick the rerere ID for that file up, and\n> > not assign a new ID as it can't successfully calculate one.  It will\n> > however still try to do the rerere operation, because of the existing\n> > ID.  As the handle_file function fails, it will remove the 'preimage'\n> > for the ID in the process, while leaving the ID in the MERGE_RR file.\n> >\n> > Now when 'rerere clear' for example is run, it will segfault in\n> > 'has_rerere_resolution', because status is NULL.\n> \n> I think this \"status\" refers to the collection->status[].  How do we\n> get into that state, though?\n> \n> new_rerere_id() and new_rerere_id_hex() fills id->collection by\n> calling find_rerere_dir(), which either finds an existing rerere_dir\n> instance or manufactures one with .status==NULL.  The .status[]\n> array is later grown by calling fit_variant as we scan and find the\n> pre/post images, but because there is no pre/post image for a file\n> with unparseable conflicts, it is left NULL.\n> \n> So another possible fix could be to make sure that .status[] is only\n> read when .status_nr says there is something worth reading.  I am\n> not saying that would be a better fix---I am just thinking out loud\n> to make sure I understand the issue correctly.\n\nYeah what you are writing above matches my understanding, and that\nshould fix the issue as well.  I haven't actually tried what you're\nproposing above, but I think I find it nicer to just remove the entry\nwe can't do anything with anyway.\n\n> > To fix this, remove the rerere ID from the MERGE_RR file in the case\n> > when we can't handle it, and remove the corresponding variant from\n> > .git/rr-cache/.  Removing it unconditionally is fine here, because if\n> > the user would have resolved the conflict and ran rerere, the entry\n> > would no longer be in the MERGE_RR file, so we wouldn't have this\n> > problem in the first place, while if the conflict was not resolved,\n> > the only thing that's left in the folder is the 'preimage', which by\n> > itself will be regenerated by git if necessary, so the user won't\n> > loose any work.\n> \n> s/loose/lose/\n> \n> > Note that other variants that have the same conflict ID will not be\n> > touched.\n> \n> Nice.  Thanks for a fix.\n> \n> >\n> > Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> > ---\n> >  rerere.c          | 12 +++++++-----\n> >  t/t4200-rerere.sh | 22 ++++++++++++++++++++++\n> >  2 files changed, 29 insertions(+), 5 deletions(-)\n> >\n> > diff --git a/rerere.c b/rerere.c\n> > index da1ab54027..895ad80c0c 100644\n> > --- a/rerere.c\n> > +++ b/rerere.c\n> > @@ -823,10 +823,7 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n> >  \t\tstruct rerere_id *id;\n> >  \t\tunsigned char sha1[20];\n> >  \t\tconst char *path = conflict.items[i].string;\n> > -\t\tint ret;\n> > -\n> > -\t\tif (string_list_has_string(rr, path))\n> > -\t\t\tcontinue;\n> > +\t\tint ret, has_string;\n> >  \n> >  \t\t/*\n> >  \t\t * Ask handle_file() to scan and assign a\n> > @@ -834,7 +831,12 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n> >  \t\t * yet.\n> >  \t\t */\n> >  \t\tret = handle_file(path, sha1, NULL);\n> > -\t\tif (ret < 1)\n> > +\t\thas_string = string_list_has_string(rr, path);\n> > +\t\tif (ret < 0 && has_string) {\n> > +\t\t\tremove_variant(string_list_lookup(rr, path)->util);\n> > +\t\t\tstring_list_remove(rr, path, 1);\n> > +\t\t}\n> > +\t\tif (ret < 1 || has_string)\n> >  \t\t\tcontinue;\n> \n> We used to say \"if we know about the path we do not do anything\n> here, if we do not see any conflict in the file we do nothing,\n> otherwise we assign a new id\"; we now say \"see if we can parse\n> and also see if we have conflict(s); if we know about the path and\n> we cannot parse, drop it from the rr database (because otherwise the\n> entry will cause us trouble elsewhere later).  Otherwise, if we do\n> not have any conflict or we already know about the path, no need to\n> do anything. Otherwise, i.e. a newly discovered path with conflicts\n> gets a new id\".\n> \n> Makes sense.  \"A known path with unparseable conflict gets dropped\"\n> is the important change in this hunk.\n> \n> > diff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\n> > index 8417e5a4b1..34f0518a5e 100755\n> > --- a/t/t4200-rerere.sh\n> > +++ b/t/t4200-rerere.sh\n> > @@ -580,4 +580,26 @@ test_expect_success 'multiple identical conflicts' '\n> >  \tcount_pre_post 0 0\n> >  '\n> >  \n> > +test_expect_success 'rerere with unexpected conflict markers does not crash' '\n> > +\tgit reset --hard &&\n> > +\n> > +\tgit checkout -b branch-1 master &&\n> > +\techo \"bar\" >test &&\n> > +\tgit add test &&\n> > +\tgit commit -q -m two &&\n> > +\n> > +\tgit reset --hard &&\n> > +\tgit checkout -b branch-2 master &&\n> > +\techo \"foo\" >test &&\n> > +\tgit add test &&\n> > +\tgit commit -q -a -m one &&\n> > +\n> > +\ttest_must_fail git merge branch-1 &&\n> > +\tsed \"s/bar/>>>>>>> a/\" >test.tmp <test &&\n> > +\tmv test.tmp test &&\n> \n> OK, so the \"only one conflict\" in the log message meant just one\n> side of the conflict marker.  More generally, the troublesome is\n> to have \"conflict marker(s) that cannot be parsed\" in the file.\n> \n> > +\tgit rerere &&\n> > +\n> > +\tgit rerere clear\n> > +'\n> > +\n> >  test_done\n"},{"id":"353954","messageId":"20180730204743.GI9955@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqq4lgghatc.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 07/11] rerere: only return whether a path has conflicts or not","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-30T20:47:43Z","receivedAt":"2018-07-30T20:47:47Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 07/30, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > We currently return the exact number of conflict hunks a certain path\n> > has from the 'handle_paths' function.  However all of its callers only\n> > care whether there are conflicts or not or if there is an error.\n> > Return only that information, and document that only that information\n> > is returned.  This will simplify the code in the subsequent steps.\n> >\n> > Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n> > ---\n> >  rerere.c | 23 ++++++++++++-----------\n> >  1 file changed, 12 insertions(+), 11 deletions(-)\n> \n> I do recall writing this code without knowing if the actual number\n> of conflicts would be useful by callers, but it is apparent that it\n> wasn't.  I won't mind losing that bit of info at all.  Besides, we\n> won't risk mistaking a file with 2 billion conflicts with a file\n> whose conflicts cannot be parsed ;-).\n\nHah, I would love to see someone actually achieve that ;)\n\n> The patch looks good.  Thanks.\n"},{"id":"353955","messageId":"20180730204918.GJ9955@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqqpnz4hau1.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v3 00/11] rerere: handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-07-30T20:49:18Z","receivedAt":"2018-07-30T20:49:22Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 07/30, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > Thomas Gummerer (11):\n> >   rerere: unify error messages when read_cache fails\n> >   rerere: lowercase error messages\n> >   rerere: wrap paths in output in sq\n> >   rerere: mark strings for translation\n> >   rerere: add documentation for conflict normalization\n> >   rerere: fix crash when conflict goes unresolved\n> >   rerere: only return whether a path has conflicts or not\n> >   rerere: factor out handle_conflict function\n> >   rerere: return strbuf from handle path\n> >   rerere: teach rerere to handle nested conflicts\n> >   rerere: recalculate conflict ID when unresolved conflict is committed\n> \n> Even though I am not certain about the last two steps, everything\n> before them looked trivially correct and good changes (well, the\n> \"strbuf\" one's goodness obviously depends on the goodness of the\n> last two, which are helped by it).\n> \n> Sorry for taking so long before getting to the series.\n\nNo worries, I realize you are busy with a lot of other things.  Thanks\na lot for your review!\n"},{"id":"354552","messageId":"20180805172037.12530-2-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 01/11] rerere: unify error messages when read_cache fails","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:27Z","receivedAt":"2018-08-05T17:20:50Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"We have multiple different variants of the error message we show to\nthe user if 'read_cache' fails.  The \"Could not read index\" variant we\nare using in 'rerere.c' is currently not used anywhere in translated\nform.\n\nAs a subsequent commit will mark all output that comes from 'rerere.c'\nfor translation, make the life of the translators a little bit easier\nby using a string that is used elsewhere, and marked for translation\nthere, and thus most likely already translated.\n\n\"index file corrupt\" seems to be the most common error message we show\nwhen 'read_cache' fails, so use that here as well.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 6 +++---\n 1 file changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex e0862e2778..473d32a5cd 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -568,7 +568,7 @@ static int find_conflict(struct string_list *conflict)\n {\n \tint i;\n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -601,7 +601,7 @@ int rerere_remaining(struct string_list *merge_rr)\n \tif (setup_rerere(merge_rr, RERERE_READONLY))\n \t\treturn 0;\n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -1103,7 +1103,7 @@ int rerere_forget(struct pathspec *pathspec)\n \tstruct string_list merge_rr = STRING_LIST_INIT_DUP;\n \n \tif (read_cache() < 0)\n-\t\treturn error(\"Could not read index\");\n+\t\treturn error(\"index file corrupt\");\n \n \tfd = setup_rerere(&merge_rr, RERERE_NOAUTOUPDATE);\n \tif (fd < 0)\n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354553","messageId":"20180805172037.12530-1-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180714214443.7184-1-t.gummerer@gmail.com","subject":"[PATCH v4 00/11] rerere: handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:26Z","receivedAt":"2018-08-05T17:20:51Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"The previous rounds were at\n<20180520211210.1248-1-t.gummerer@gmail.com>,\n<20180605215219.28783-1-t.gummerer@gmail.com> and\n<20180714214443.7184-1-t.gummerer@gmail.com>.\n\nThanks Junio for the review and Simon for pointing out an error in my\ncommit message.\n\nThe changes in this round are mainly improving the commit messages,\nand polishing the documentation.\n\nIt also simplifies one test case in patch 6/11.\n\nPatches 10 and 11 are still included, however I'm not going to be too\nsad if we decide to not include them, as they really only help in an\nobscure case, which could be considered using git \"wrong\".\n\nI also realized that while I wrote \"no functional changes intended\" in\n7/11, and functional changes were in fact not intended, there still is\na slight functional change.  As I think that's a good change, I\ndocumented it in the commit message.\n\nThomas Gummerer (11):\n  rerere: unify error messages when read_cache fails\n  rerere: lowercase error messages\n  rerere: wrap paths in output in sq\n  rerere: mark strings for translation\n  rerere: add documentation for conflict normalization\n  rerere: fix crash with files rerere can't handle\n  rerere: only return whether a path has conflicts or not\n  rerere: factor out handle_conflict function\n  rerere: return strbuf from handle path\n  rerere: teach rerere to handle nested conflicts\n  rerere: recalculate conflict ID when unresolved conflict is committed\n\n Documentation/technical/rerere.txt | 182 +++++++++++++++++++++\n builtin/rerere.c                   |   4 +-\n rerere.c                           | 243 ++++++++++++++---------------\n t/t4200-rerere.sh                  |  65 ++++++++\n 4 files changed, 365 insertions(+), 129 deletions(-)\n create mode 100644 Documentation/technical/rerere.txt\n\nRange diff below:\n\n 1:  ce876f1b6b =  1:  018bd68a8a rerere: unify error messages when read_cache fails\n 2:  0326503c4a =  2:  281fcbf24f rerere: lowercase error messages\n 3:  a33211e3d3 =  3:  b6d5e2e26d rerere: wrap paths in output in sq\n 4:  3da84604f0 !  4:  6ed390c8f5 rerere: mark strings for translation\n    @@ -2,7 +2,7 @@\n     \n         rerere: mark strings for translation\n     \n    -    'git rerere' is considered a plumbing command and as such its output\n    +    'git rerere' is considered a porcelain command and as such its output\n         should be translated.  Its functionality is also only enabled through\n         a config setting, so scripts really shouldn't rely on the output\n         either way.\n 5:  749d49a625 !  5:  3cef1d57bc rerere: add documentation for conflict normalization\n    @@ -28,8 +28,8 @@\n     +conflicts before writing them to the rerere database.\n     +\n     +Different conflict styles and branch names are normalized by stripping\n    -+the labels from the conflict markers, and removing extraneous\n    -+information from the `diff3` conflict style. Branches that are merged\n    ++the labels from the conflict markers, and removing the common ancestor\n    ++version from the `diff3` conflict style. Branches that are merged\n     +in different order are normalized by sorting the conflict hunks.  More\n     +on each of those steps in the following sections.\n     +\n    @@ -37,8 +37,8 @@\n     +calculated based on the normalized conflict, which is later used by\n     +rerere to look up the conflict in the rerere database.\n     +\n    -+Stripping extraneous information\n    -+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n    ++Removing the common ancestor version\n    ++~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n     +\n     +Say we have three branches AB, AC and AC2.  The common ancestor of\n     +these branches has a file with a line containing the string \"A\" (for\n    @@ -79,7 +79,7 @@\n     +\n     +By extension, this means that rerere should recognize that the above\n     +conflicts are the same.  To do this, the labels on the conflict\n    -+markers are stripped, and the diff3 output is removed.  The above\n    ++markers are stripped, and the common ancestor version is removed.  The above\n     +examples would both result in the following normalized conflict:\n     +\n     +    <<<<<<<\n 6:  d465bd087e !  6:  a02d90157d rerere: fix crash when conflict goes unresolved\n    @@ -1,37 +1,42 @@\n     Author: Thomas Gummerer <t.gummerer@gmail.com>\n     \n    -    rerere: fix crash when conflict goes unresolved\n    +    rerere: fix crash with files rerere can't handle\n     \n    -    Currently when a user doesn't resolve a conflict in a file, but\n    -    commits the file with the conflict markers, and later the file ends up\n    -    in a state in which rerere can't handle it, subsequent rerere\n    -    operations that are interested in that path, such as 'rerere clear' or\n    -    'rerere forget <path>' will fail, or even worse in the case of 'rerere\n    -    clear' segfault.\n    +    Currently when a user does a conflict resolution and ends it (in any\n    +    way that calls 'git rerere' again) with a file 'rerere' can't handle,\n    +    subsequent rerere operations that are interested in that path, such as\n    +    'rerere clear' or 'rerere forget <path>' will fail, or even worse in\n    +    the case of 'rerere clear' segfault.\n     \n    -    Such states include nested conflicts, or an extra conflict marker that\n    +    Such states include nested conflicts, or a conflict marker that\n         doesn't have any match.\n     \n    -    This is because the first 'git rerere' when there was only one\n    -    conflict in the file leaves an entry in the MERGE_RR file behind.  The\n    -    next 'git rerere' will then pick the rerere ID for that file up, and\n    -    not assign a new ID as it can't successfully calculate one.  It will\n    -    however still try to do the rerere operation, because of the existing\n    -    ID.  As the handle_file function fails, it will remove the 'preimage'\n    -    for the ID in the process, while leaving the ID in the MERGE_RR file.\n    +    This is because 'git rerere' calculates a conflict file and writes it\n    +    to the MERGE_RR file.  When the user then changes the file in any way\n    +    rerere can't handle, and then calls 'git rerere' on it again to record\n    +    the conflict resolution, the handle_file function fails, and removes\n    +    the 'preimage' file in the rr-cache in the process, while leaving the\n    +    ID in the MERGE_RR file.\n     \n    -    Now when 'rerere clear' for example is run, it will segfault in\n    -    'has_rerere_resolution', because status is NULL.\n    +    Now when 'rerere clear' is run, it reads the ID from the MERGE_RR\n    +    file, however the 'fit_variant' function for the ID is never called as\n    +    the 'preimage' file does not exist anymore.  This means\n    +    'collection->status' in 'has_rerere_resolution' is NULL, and the\n    +    command will crash.\n     \n         To fix this, remove the rerere ID from the MERGE_RR file in the case\n    -    when we can't handle it, and remove the corresponding variant from\n    -    .git/rr-cache/.  Removing it unconditionally is fine here, because if\n    -    the user would have resolved the conflict and ran rerere, the entry\n    -    would no longer be in the MERGE_RR file, so we wouldn't have this\n    -    problem in the first place, while if the conflict was not resolved,\n    -    the only thing that's left in the folder is the 'preimage', which by\n    -    itself will be regenerated by git if necessary, so the user won't\n    -    loose any work.\n    +    when we can't handle it, just after the 'preimage' file was removed\n    +    and remove the corresponding variant from .git/rr-cache/.  Removing it\n    +    unconditionally is fine here, because if the user would have resolved\n    +    the conflict and ran rerere, the entry would no longer be in the\n    +    MERGE_RR file, so we wouldn't have this problem in the first place,\n    +    while if the conflict was not resolved.\n    +\n    +    Currently there is nothing left in this folder, as the 'preimage'\n    +    was already deleted by the 'handle_file' function, so 'remove_variant'\n    +    is a no-op.  Still call the function, to make sure we clean everything\n    +    up, in case we add some other files corresponding to a variant in the\n    +    future.\n     \n         Note that other variants that have the same conflict ID will not be\n         touched.\n    @@ -90,8 +95,7 @@\n     +\tgit commit -q -a -m one &&\n     +\n     +\ttest_must_fail git merge branch-1 &&\n    -+\tsed \"s/bar/>>>>>>> a/\" >test.tmp <test &&\n    -+\tmv test.tmp test &&\n    ++\techo \"<<<<<<< a\" >test &&\n     +\tgit rerere &&\n     +\n     +\tgit rerere clear\n 7:  fac2b79245 =  7:  49815bee02 rerere: only return whether a path has conflicts or not\n 8:  b5892c1861 !  8:  0c51696d10 rerere: factor out handle_conflict function\n    @@ -4,7 +4,13 @@\n     \n         Factor out the handle_conflict function, which handles a single\n         conflict in a path.  This is in preparation for a subsequent commit,\n    -    where this function will be re-used.  No functional changes intended.\n    +    where this function will be re-used.\n    +\n    +    Note that this does change the behaviour of 'git rerere' slightly.\n    +    Where previously we'd consider all files where an unmatched conflict\n    +    marker is found as invalid, we now only consider files invalid when\n    +    the \"ours\" conflict marker (\"<<<<<<< <text>\") is unmatched, not when\n    +    other conflict markers (e.g. \"=======\") is unmatched.\n     \n         Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n     \n 9:  e8e0ca4db9 =  9:  f604efe05d rerere: return strbuf from handle path\n10:  1fc106ffaa ! 10:  a2393d3424 rerere: teach rerere to handle nested conflicts\n    @@ -6,6 +6,10 @@\n         it encounters such conflicts.  Do that by recursively calling the\n         'handle_conflict' function to normalize the conflict.\n     \n    +    Note that a conflict like this would only be produced if a user\n    +    commits a file with conflict markers, and gets a conflict including\n    +    that in a susbsequent operation.\n    +\n         The conflict ID calculation here deserves some explanation:\n     \n         As we are using the same handle_conflict function, the nested conflict\n    @@ -66,8 +70,8 @@\n     +\n     +Nested conflicts are handled very similarly to \"simple\" conflicts.\n     +Similar to simple conflicts, the conflict is first normalized by\n    -+stripping the labels from conflict markers, stripping the diff3\n    -+output, and the sorting the conflict hunks, both for the outer and the\n    ++stripping the labels from conflict markers, stripping the common ancestor\n    ++version, and the sorting the conflict hunks, both for the outer and the\n     +inner conflict.  This is done recursively, so any number of nested\n     +conflicts can be handled.\n     +\n11:  4463aed2f8 = 11:  371af30766 rerere: recalculate conflict ID when unresolved conflict is committed\n\n-- \n2.18.0.720.gf7a957e2e7\n"},{"id":"354554","messageId":"20180805172037.12530-3-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 02/11] rerere: lowercase error messages","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:28Z","receivedAt":"2018-08-05T17:20:51Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Documentation/CodingGuidelines mentions that error messages should be\nlowercase.  Prior to marking them for translation follow that pattern\nin rerere as well, so translators won't have to translate messages\nthat don't conform to our guidelines.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 24 ++++++++++++------------\n 1 file changed, 12 insertions(+), 12 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 473d32a5cd..c5d9ea171f 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"Could not open %s\", path);\n+\t\treturn error_errno(\"could not open %s\", path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"Could not write %s\", output);\n+\t\t\terror_errno(\"could not write %s\", output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"There were errors while writing %s (%s)\",\n+\t\terror(\"there were errors while writing %s (%s)\",\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"Failed to flush %s\", path);\n+\t\tio.io.wrerror = error_errno(\"failed to flush %s\", path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"Could not parse conflict hunks in %s\", path);\n+\t\treturn error(\"could not parse conflict hunks in %s\", path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -690,11 +690,11 @@ static int merge(const struct rerere_id *id, const char *path)\n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"Could not open %s\", path);\n+\t\treturn error_errno(\"could not open %s\", path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"Could not write %s\", path);\n+\t\terror_errno(\"could not write %s\", path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"Writing %s failed\", path);\n+\t\treturn error_errno(\"writing %s failed\", path);\n \n out:\n \tfree(cur.ptr);\n@@ -720,7 +720,7 @@ static void update_paths(struct string_list *update)\n \n \tif (write_locked_index(&the_index, &index_lock,\n \t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED))\n-\t\tdie(\"Unable to write new index file\");\n+\t\tdie(\"unable to write new index file\");\n }\n \n static void remove_variant(struct rerere_id *id)\n@@ -878,7 +878,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"Could not create directory %s\", git_path_rr_cache());\n+\t\tdie(\"could not create directory %s\", git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1031,7 +1031,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t */\n \tret = handle_cache(path, sha1, NULL);\n \tif (ret < 1)\n-\t\treturn error(\"Could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n \n \t/* Nuke the recorded resolution for the conflict */\n \tid = new_rerere_id(sha1);\n@@ -1049,7 +1049,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t\thandle_cache(path, sha1, rerere_path(id, \"thisimage\"));\n \t\tif (read_mmfile(&cur, rerere_path(id, \"thisimage\"))) {\n \t\t\tfree(cur.ptr);\n-\t\t\terror(\"Failed to update conflicted state in '%s'\", path);\n+\t\t\terror(\"failed to update conflicted state in '%s'\", path);\n \t\t\tgoto fail_exit;\n \t\t}\n \t\tcleanly_resolved = !try_merge(id, path, &cur, &result);\n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354555","messageId":"20180805172037.12530-4-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 03/11] rerere: wrap paths in output in sq","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:29Z","receivedAt":"2018-08-05T17:20:53Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"It looks like most paths in the output in the git codebase are wrapped\nin single quotes.  Standardize on that in rerere as well.\n\nApart from being more consistent, this also makes some of the strings\nmatch strings that are already translated in other parts of the\ncodebase, thus reducing the work for translators, when the strings are\nmarked for translation in a subsequent commit.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n builtin/rerere.c |  2 +-\n rerere.c         | 26 +++++++++++++-------------\n 2 files changed, 14 insertions(+), 14 deletions(-)\n\ndiff --git a/builtin/rerere.c b/builtin/rerere.c\nindex 0bc40298c2..e0c67c98e9 100644\n--- a/builtin/rerere.c\n+++ b/builtin/rerere.c\n@@ -107,7 +107,7 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \t\t\tconst char *path = merge_rr.items[i].string;\n \t\t\tconst struct rerere_id *id = merge_rr.items[i].util;\n \t\t\tif (diff_two(rerere_path(id, \"preimage\"), path, path, path))\n-\t\t\t\tdie(\"unable to generate diff for %s\", rerere_path(id, NULL));\n+\t\t\t\tdie(\"unable to generate diff for '%s'\", rerere_path(id, NULL));\n \t\t}\n \t} else\n \t\tusage_with_options(rerere_usage, options);\ndiff --git a/rerere.c b/rerere.c\nindex c5d9ea171f..cde1f6e696 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"could not open %s\", path);\n+\t\treturn error_errno(\"could not open '%s'\", path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"could not write %s\", output);\n+\t\t\terror_errno(\"could not write '%s'\", output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"there were errors while writing %s (%s)\",\n+\t\terror(\"there were errors while writing '%s' (%s)\",\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"failed to flush %s\", path);\n+\t\tio.io.wrerror = error_errno(\"failed to flush '%s'\", path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"could not parse conflict hunks in %s\", path);\n+\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -684,17 +684,17 @@ static int merge(const struct rerere_id *id, const char *path)\n \t * Mark that \"postimage\" was used to help gc.\n \t */\n \tif (utime(rerere_path(id, \"postimage\"), NULL) < 0)\n-\t\twarning_errno(\"failed utime() on %s\",\n+\t\twarning_errno(\"failed utime() on '%s'\",\n \t\t\t      rerere_path(id, \"postimage\"));\n \n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"could not open %s\", path);\n+\t\treturn error_errno(\"could not open '%s'\", path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"could not write %s\", path);\n+\t\terror_errno(\"could not write '%s'\", path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"writing %s failed\", path);\n+\t\treturn error_errno(\"writing '%s' failed\", path);\n \n out:\n \tfree(cur.ptr);\n@@ -878,7 +878,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"could not create directory %s\", git_path_rr_cache());\n+\t\tdie(\"could not create directory '%s'\", git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1067,9 +1067,9 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \tfilename = rerere_path(id, \"postimage\");\n \tif (unlink(filename)) {\n \t\tif (errno == ENOENT)\n-\t\t\terror(\"no remembered resolution for %s\", path);\n+\t\t\terror(\"no remembered resolution for '%s'\", path);\n \t\telse\n-\t\t\terror_errno(\"cannot unlink %s\", filename);\n+\t\t\terror_errno(\"cannot unlink '%s'\", filename);\n \t\tgoto fail_exit;\n \t}\n \n@@ -1088,7 +1088,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \titem = string_list_insert(rr, path);\n \tfree_rerere_id(item);\n \titem->util = id;\n-\tfprintf(stderr, \"Forgot resolution for %s\\n\", path);\n+\tfprintf(stderr, \"Forgot resolution for '%s'\\n\", path);\n \treturn 0;\n \n fail_exit:\n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354557","messageId":"20180805172037.12530-5-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 04/11] rerere: mark strings for translation","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:30Z","receivedAt":"2018-08-05T17:20:56Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"'git rerere' is considered a porcelain command and as such its output\nshould be translated.  Its functionality is also only enabled through\na config setting, so scripts really shouldn't rely on the output\neither way.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n builtin/rerere.c |  4 +--\n rerere.c         | 68 ++++++++++++++++++++++++------------------------\n 2 files changed, 36 insertions(+), 36 deletions(-)\n\ndiff --git a/builtin/rerere.c b/builtin/rerere.c\nindex e0c67c98e9..5ed941b91f 100644\n--- a/builtin/rerere.c\n+++ b/builtin/rerere.c\n@@ -75,7 +75,7 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \tif (!strcmp(argv[0], \"forget\")) {\n \t\tstruct pathspec pathspec;\n \t\tif (argc < 2)\n-\t\t\twarning(\"'git rerere forget' without paths is deprecated\");\n+\t\t\twarning(_(\"'git rerere forget' without paths is deprecated\"));\n \t\tparse_pathspec(&pathspec, 0, PATHSPEC_PREFER_CWD,\n \t\t\t       prefix, argv + 1);\n \t\treturn rerere_forget(&pathspec);\n@@ -107,7 +107,7 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \t\t\tconst char *path = merge_rr.items[i].string;\n \t\t\tconst struct rerere_id *id = merge_rr.items[i].util;\n \t\t\tif (diff_two(rerere_path(id, \"preimage\"), path, path, path))\n-\t\t\t\tdie(\"unable to generate diff for '%s'\", rerere_path(id, NULL));\n+\t\t\t\tdie(_(\"unable to generate diff for '%s'\"), rerere_path(id, NULL));\n \t\t}\n \t} else\n \t\tusage_with_options(rerere_usage, options);\ndiff --git a/rerere.c b/rerere.c\nindex cde1f6e696..be98c0afcb 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -212,7 +212,7 @@ static void read_rr(struct string_list *rr)\n \n \t\t/* There has to be the hash, tab, path and then NUL */\n \t\tif (buf.len < 42 || get_sha1_hex(buf.buf, sha1))\n-\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \n \t\tif (buf.buf[40] != '.') {\n \t\t\tvariant = 0;\n@@ -221,10 +221,10 @@ static void read_rr(struct string_list *rr)\n \t\t\terrno = 0;\n \t\t\tvariant = strtol(buf.buf + 41, &path, 10);\n \t\t\tif (errno)\n-\t\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \t\t}\n \t\tif (*(path++) != '\\t')\n-\t\t\tdie(\"corrupt MERGE_RR\");\n+\t\t\tdie(_(\"corrupt MERGE_RR\"));\n \t\tbuf.buf[40] = '\\0';\n \t\tid = new_rerere_id_hex(buf.buf);\n \t\tid->variant = variant;\n@@ -259,12 +259,12 @@ static int write_rr(struct string_list *rr, int out_fd)\n \t\t\t\t    rr->items[i].string, 0);\n \n \t\tif (write_in_full(out_fd, buf.buf, buf.len) < 0)\n-\t\t\tdie(\"unable to write rerere record\");\n+\t\t\tdie(_(\"unable to write rerere record\"));\n \n \t\tstrbuf_release(&buf);\n \t}\n \tif (commit_lock_file(&write_lock) != 0)\n-\t\tdie(\"unable to write rerere record\");\n+\t\tdie(_(\"unable to write rerere record\"));\n \treturn 0;\n }\n \n@@ -484,12 +484,12 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tio.input = fopen(path, \"r\");\n \tio.io.wrerror = 0;\n \tif (!io.input)\n-\t\treturn error_errno(\"could not open '%s'\", path);\n+\t\treturn error_errno(_(\"could not open '%s'\"), path);\n \n \tif (output) {\n \t\tio.io.output = fopen(output, \"w\");\n \t\tif (!io.io.output) {\n-\t\t\terror_errno(\"could not write '%s'\", output);\n+\t\t\terror_errno(_(\"could not write '%s'\"), output);\n \t\t\tfclose(io.input);\n \t\t\treturn -1;\n \t\t}\n@@ -499,15 +499,15 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n-\t\terror(\"there were errors while writing '%s' (%s)\",\n+\t\terror(_(\"there were errors while writing '%s' (%s)\"),\n \t\t      path, strerror(io.io.wrerror));\n \tif (io.io.output && fclose(io.io.output))\n-\t\tio.io.wrerror = error_errno(\"failed to flush '%s'\", path);\n+\t\tio.io.wrerror = error_errno(_(\"failed to flush '%s'\"), path);\n \n \tif (hunk_no < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n-\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n@@ -568,7 +568,7 @@ static int find_conflict(struct string_list *conflict)\n {\n \tint i;\n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -601,7 +601,7 @@ int rerere_remaining(struct string_list *merge_rr)\n \tif (setup_rerere(merge_rr, RERERE_READONLY))\n \t\treturn 0;\n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfor (i = 0; i < active_nr;) {\n \t\tint conflict_type;\n@@ -684,17 +684,17 @@ static int merge(const struct rerere_id *id, const char *path)\n \t * Mark that \"postimage\" was used to help gc.\n \t */\n \tif (utime(rerere_path(id, \"postimage\"), NULL) < 0)\n-\t\twarning_errno(\"failed utime() on '%s'\",\n+\t\twarning_errno(_(\"failed utime() on '%s'\"),\n \t\t\t      rerere_path(id, \"postimage\"));\n \n \t/* Update \"path\" with the resolution */\n \tf = fopen(path, \"w\");\n \tif (!f)\n-\t\treturn error_errno(\"could not open '%s'\", path);\n+\t\treturn error_errno(_(\"could not open '%s'\"), path);\n \tif (fwrite(result.ptr, result.size, 1, f) != 1)\n-\t\terror_errno(\"could not write '%s'\", path);\n+\t\terror_errno(_(\"could not write '%s'\"), path);\n \tif (fclose(f))\n-\t\treturn error_errno(\"writing '%s' failed\", path);\n+\t\treturn error_errno(_(\"writing '%s' failed\"), path);\n \n out:\n \tfree(cur.ptr);\n@@ -714,13 +714,13 @@ static void update_paths(struct string_list *update)\n \t\tstruct string_list_item *item = &update->items[i];\n \t\tif (add_file_to_cache(item->string, 0))\n \t\t\texit(128);\n-\t\tfprintf(stderr, \"Staged '%s' using previous resolution.\\n\",\n+\t\tfprintf_ln(stderr, _(\"Staged '%s' using previous resolution.\"),\n \t\t\titem->string);\n \t}\n \n \tif (write_locked_index(&the_index, &index_lock,\n \t\t\t       COMMIT_LOCK | SKIP_IF_UNCHANGED))\n-\t\tdie(\"unable to write new index file\");\n+\t\tdie(_(\"unable to write new index file\"));\n }\n \n static void remove_variant(struct rerere_id *id)\n@@ -752,7 +752,7 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \t\tif (!handle_file(path, NULL, NULL)) {\n \t\t\tcopy_file(rerere_path(id, \"postimage\"), path, 0666);\n \t\t\tid->collection->status[variant] |= RR_HAS_POSTIMAGE;\n-\t\t\tfprintf(stderr, \"Recorded resolution for '%s'.\\n\", path);\n+\t\t\tfprintf_ln(stderr, _(\"Recorded resolution for '%s'.\"), path);\n \t\t\tfree_rerere_id(rr_item);\n \t\t\trr_item->util = NULL;\n \t\t\treturn;\n@@ -786,9 +786,9 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \t\tif (rerere_autoupdate)\n \t\t\tstring_list_insert(update, path);\n \t\telse\n-\t\t\tfprintf(stderr,\n-\t\t\t\t\"Resolved '%s' using previous resolution.\\n\",\n-\t\t\t\tpath);\n+\t\t\tfprintf_ln(stderr,\n+\t\t\t\t   _(\"Resolved '%s' using previous resolution.\"),\n+\t\t\t\t   path);\n \t\tfree_rerere_id(rr_item);\n \t\trr_item->util = NULL;\n \t\treturn;\n@@ -802,11 +802,11 @@ static void do_rerere_one_path(struct string_list_item *rr_item,\n \tif (id->collection->status[variant] & RR_HAS_POSTIMAGE) {\n \t\tconst char *path = rerere_path(id, \"postimage\");\n \t\tif (unlink(path))\n-\t\t\tdie_errno(\"cannot unlink stray '%s'\", path);\n+\t\t\tdie_errno(_(\"cannot unlink stray '%s'\"), path);\n \t\tid->collection->status[variant] &= ~RR_HAS_POSTIMAGE;\n \t}\n \tid->collection->status[variant] |= RR_HAS_PREIMAGE;\n-\tfprintf(stderr, \"Recorded preimage for '%s'\\n\", path);\n+\tfprintf_ln(stderr, _(\"Recorded preimage for '%s'\"), path);\n }\n \n static int do_plain_rerere(struct string_list *rr, int fd)\n@@ -878,7 +878,7 @@ static int is_rerere_enabled(void)\n \t\treturn rr_cache_exists;\n \n \tif (!rr_cache_exists && mkdir_in_gitdir(git_path_rr_cache()))\n-\t\tdie(\"could not create directory '%s'\", git_path_rr_cache());\n+\t\tdie(_(\"could not create directory '%s'\"), git_path_rr_cache());\n \treturn 1;\n }\n \n@@ -1031,7 +1031,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t */\n \tret = handle_cache(path, sha1, NULL);\n \tif (ret < 1)\n-\t\treturn error(\"could not parse conflict hunks in '%s'\", path);\n+\t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n \n \t/* Nuke the recorded resolution for the conflict */\n \tid = new_rerere_id(sha1);\n@@ -1049,7 +1049,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t\thandle_cache(path, sha1, rerere_path(id, \"thisimage\"));\n \t\tif (read_mmfile(&cur, rerere_path(id, \"thisimage\"))) {\n \t\t\tfree(cur.ptr);\n-\t\t\terror(\"failed to update conflicted state in '%s'\", path);\n+\t\t\terror(_(\"failed to update conflicted state in '%s'\"), path);\n \t\t\tgoto fail_exit;\n \t\t}\n \t\tcleanly_resolved = !try_merge(id, path, &cur, &result);\n@@ -1060,16 +1060,16 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t}\n \n \tif (id->collection->status_nr <= id->variant) {\n-\t\terror(\"no remembered resolution for '%s'\", path);\n+\t\terror(_(\"no remembered resolution for '%s'\"), path);\n \t\tgoto fail_exit;\n \t}\n \n \tfilename = rerere_path(id, \"postimage\");\n \tif (unlink(filename)) {\n \t\tif (errno == ENOENT)\n-\t\t\terror(\"no remembered resolution for '%s'\", path);\n+\t\t\terror(_(\"no remembered resolution for '%s'\"), path);\n \t\telse\n-\t\t\terror_errno(\"cannot unlink '%s'\", filename);\n+\t\t\terror_errno(_(\"cannot unlink '%s'\"), filename);\n \t\tgoto fail_exit;\n \t}\n \n@@ -1079,7 +1079,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \t * the postimage.\n \t */\n \thandle_cache(path, sha1, rerere_path(id, \"preimage\"));\n-\tfprintf(stderr, \"Updated preimage for '%s'\\n\", path);\n+\tfprintf_ln(stderr, _(\"Updated preimage for '%s'\"), path);\n \n \t/*\n \t * And remember that we can record resolution for this\n@@ -1088,7 +1088,7 @@ static int rerere_forget_one_path(const char *path, struct string_list *rr)\n \titem = string_list_insert(rr, path);\n \tfree_rerere_id(item);\n \titem->util = id;\n-\tfprintf(stderr, \"Forgot resolution for '%s'\\n\", path);\n+\tfprintf(stderr, _(\"Forgot resolution for '%s'\\n\"), path);\n \treturn 0;\n \n fail_exit:\n@@ -1103,7 +1103,7 @@ int rerere_forget(struct pathspec *pathspec)\n \tstruct string_list merge_rr = STRING_LIST_INIT_DUP;\n \n \tif (read_cache() < 0)\n-\t\treturn error(\"index file corrupt\");\n+\t\treturn error(_(\"index file corrupt\"));\n \n \tfd = setup_rerere(&merge_rr, RERERE_NOAUTOUPDATE);\n \tif (fd < 0)\n@@ -1191,7 +1191,7 @@ void rerere_gc(struct string_list *rr)\n \tgit_config(git_default_config, NULL);\n \tdir = opendir(git_path(\"rr-cache\"));\n \tif (!dir)\n-\t\tdie_errno(\"unable to open rr-cache directory\");\n+\t\tdie_errno(_(\"unable to open rr-cache directory\"));\n \t/* Collect stale conflict IDs ... */\n \twhile ((e = readdir(dir))) {\n \t\tstruct rerere_dir *rr_dir;\n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354556","messageId":"20180805172037.12530-6-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 05/11] rerere: add documentation for conflict normalization","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:31Z","receivedAt":"2018-08-05T17:20:57Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Add some documentation for the logic behind the conflict normalization\nin rerere.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Documentation/technical/rerere.txt | 140 +++++++++++++++++++++++++++++\n rerere.c                           |   4 -\n 2 files changed, 140 insertions(+), 4 deletions(-)\n create mode 100644 Documentation/technical/rerere.txt\n\ndiff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\nnew file mode 100644\nindex 0000000000..3d10dbfa67\n--- /dev/null\n+++ b/Documentation/technical/rerere.txt\n@@ -0,0 +1,140 @@\n+Rerere\n+======\n+\n+This document describes the rerere logic.\n+\n+Conflict normalization\n+----------------------\n+\n+To ensure recorded conflict resolutions can be looked up in the rerere\n+database, even when branches are merged in a different order,\n+different branches are merged that result in the same conflict, or\n+when different conflict style settings are used, rerere normalizes the\n+conflicts before writing them to the rerere database.\n+\n+Different conflict styles and branch names are normalized by stripping\n+the labels from the conflict markers, and removing the common ancestor\n+version from the `diff3` conflict style. Branches that are merged\n+in different order are normalized by sorting the conflict hunks.  More\n+on each of those steps in the following sections.\n+\n+Once these two normalization operations are applied, a conflict ID is\n+calculated based on the normalized conflict, which is later used by\n+rerere to look up the conflict in the rerere database.\n+\n+Removing the common ancestor version\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+\n+Say we have three branches AB, AC and AC2.  The common ancestor of\n+these branches has a file with a line containing the string \"A\" (for\n+brevity this is called \"line A\" in the rest of the document).  In\n+branch AB this line is changed to \"B\", in AC, this line is changed to\n+\"C\", and branch AC2 is forked off of AC, after the line was changed to\n+\"C\".\n+\n+Forking a branch ABAC off of branch AB and then merging AC into it, we\n+get a conflict like the following:\n+\n+    <<<<<<< HEAD\n+    B\n+    =======\n+    C\n+    >>>>>>> AC\n+\n+Doing the analogous with AC2 (forking a branch ABAC2 off of branch AB\n+and then merging branch AC2 into it), using the diff3 conflict style,\n+we get a conflict like the following:\n+\n+    <<<<<<< HEAD\n+    B\n+    ||||||| merged common ancestors\n+    A\n+    =======\n+    C\n+    >>>>>>> AC2\n+\n+By resolving this conflict, to leave line D, the user declares:\n+\n+    After examining what branches AB and AC did, I believe that making\n+    line A into line D is the best thing to do that is compatible with\n+    what AB and AC wanted to do.\n+\n+As branch AC2 refers to the same commit as AC, the above implies that\n+this is also compatible what AB and AC2 wanted to do.\n+\n+By extension, this means that rerere should recognize that the above\n+conflicts are the same.  To do this, the labels on the conflict\n+markers are stripped, and the common ancestor version is removed.  The above\n+examples would both result in the following normalized conflict:\n+\n+    <<<<<<<\n+    B\n+    =======\n+    C\n+    >>>>>>>\n+\n+Sorting hunks\n+~~~~~~~~~~~~~\n+\n+As before, lets imagine that a common ancestor had a file with line A\n+its early part, and line X in its late part.  And then four branches\n+are forked that do these things:\n+\n+    - AB: changes A to B\n+    - AC: changes A to C\n+    - XY: changes X to Y\n+    - XZ: changes X to Z\n+\n+Now, forking a branch ABAC off of branch AB and then merging AC into\n+it, and forking a branch ACAB off of branch AC and then merging AB\n+into it, would yield the conflict in a different order.  The former\n+would say \"A became B or C, what now?\" while the latter would say \"A\n+became C or B, what now?\"\n+\n+As a reminder, the act of merging AC into ABAC and resolving the\n+conflict to leave line D means that the user declares:\n+\n+    After examining what branches AB and AC did, I believe that\n+    making line A into line D is the best thing to do that is\n+    compatible with what AB and AC wanted to do.\n+\n+So the conflict we would see when merging AB into ACAB should be\n+resolved the same way---it is the resolution that is in line with that\n+declaration.\n+\n+Imagine that similarly previously a branch XYXZ was forked from XY,\n+and XZ was merged into it, and resolved \"X became Y or Z\" into \"X\n+became W\".\n+\n+Now, if a branch ABXY was forked from AB and then merged XY, then ABXY\n+would have line B in its early part and line Y in its later part.\n+Such a merge would be quite clean.  We can construct 4 combinations\n+using these four branches ((AB, AC) x (XY, XZ)).\n+\n+Merging ABXY and ACXZ would make \"an early A became B or C, a late X\n+became Y or Z\" conflict, while merging ACXY and ABXZ would make \"an\n+early A became C or B, a late X became Y or Z\".  We can see there are\n+4 combinations of (\"B or C\", \"C or B\") x (\"X or Y\", \"Y or X\").\n+\n+By sorting, the conflict is given its canonical name, namely, \"an\n+early part became B or C, a late part becames X or Y\", and whenever\n+any of these four patterns appear, and we can get to the same conflict\n+and resolution that we saw earlier.\n+\n+Without the sorting, we'd have to somehow find a previous resolution\n+from combinatorial explosion.\n+\n+Conflict ID calculation\n+~~~~~~~~~~~~~~~~~~~~~~~\n+\n+Once the conflict normalization is done, the conflict ID is calculated\n+as the sha1 hash of the conflict hunks appended to each other,\n+separated by <NUL> characters.  The conflict markers are stripped out\n+before the sha1 is calculated.  So in the example above, where we\n+merge branch AC which changes line A to line C, into branch AB, which\n+changes line A to line C, the conflict ID would be\n+SHA1('B<NUL>C<NUL>').\n+\n+If there are multiple conflicts in one file, the sha1 is calculated\n+the same way with all hunks appended to each other, in the order in\n+which they appear in the file, separated by a <NUL> character.\ndiff --git a/rerere.c b/rerere.c\nindex be98c0afcb..da1ab54027 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -394,10 +394,6 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n  * and NUL concatenated together.\n  *\n  * Return the number of conflict hunks found.\n- *\n- * NEEDSWORK: the logic and theory of operation behind this conflict\n- * normalization may deserve to be documented somewhere, perhaps in\n- * Documentation/technical/rerere.txt.\n  */\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354558","messageId":"20180805172037.12530-7-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 06/11] rerere: fix crash with files rerere can't handle","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:32Z","receivedAt":"2018-08-05T17:20:58Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently when a user does a conflict resolution and ends it (in any\nway that calls 'git rerere' again) with a file 'rerere' can't handle,\nsubsequent rerere operations that are interested in that path, such as\n'rerere clear' or 'rerere forget <path>' will fail, or even worse in\nthe case of 'rerere clear' segfault.\n\nSuch states include nested conflicts, or a conflict marker that\ndoesn't have any match.\n\nThis is because 'git rerere' calculates a conflict file and writes it\nto the MERGE_RR file.  When the user then changes the file in any way\nrerere can't handle, and then calls 'git rerere' on it again to record\nthe conflict resolution, the handle_file function fails, and removes\nthe 'preimage' file in the rr-cache in the process, while leaving the\nID in the MERGE_RR file.\n\nNow when 'rerere clear' is run, it reads the ID from the MERGE_RR\nfile, however the 'fit_variant' function for the ID is never called as\nthe 'preimage' file does not exist anymore.  This means\n'collection->status' in 'has_rerere_resolution' is NULL, and the\ncommand will crash.\n\nTo fix this, remove the rerere ID from the MERGE_RR file in the case\nwhen we can't handle it, just after the 'preimage' file was removed\nand remove the corresponding variant from .git/rr-cache/.  Removing it\nunconditionally is fine here, because if the user would have resolved\nthe conflict and ran rerere, the entry would no longer be in the\nMERGE_RR file, so we wouldn't have this problem in the first place,\nwhile if the conflict was not resolved.\n\nCurrently there is nothing left in this folder, as the 'preimage'\nwas already deleted by the 'handle_file' function, so 'remove_variant'\nis a no-op.  Still call the function, to make sure we clean everything\nup, in case we add some other files corresponding to a variant in the\nfuture.\n\nNote that other variants that have the same conflict ID will not be\ntouched.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c          | 12 +++++++-----\n t/t4200-rerere.sh | 21 +++++++++++++++++++++\n 2 files changed, 28 insertions(+), 5 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex da1ab54027..895ad80c0c 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -823,10 +823,7 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\tstruct rerere_id *id;\n \t\tunsigned char sha1[20];\n \t\tconst char *path = conflict.items[i].string;\n-\t\tint ret;\n-\n-\t\tif (string_list_has_string(rr, path))\n-\t\t\tcontinue;\n+\t\tint ret, has_string;\n \n \t\t/*\n \t\t * Ask handle_file() to scan and assign a\n@@ -834,7 +831,12 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\t * yet.\n \t\t */\n \t\tret = handle_file(path, sha1, NULL);\n-\t\tif (ret < 1)\n+\t\thas_string = string_list_has_string(rr, path);\n+\t\tif (ret < 0 && has_string) {\n+\t\t\tremove_variant(string_list_lookup(rr, path)->util);\n+\t\t\tstring_list_remove(rr, path, 1);\n+\t\t}\n+\t\tif (ret < 1 || has_string)\n \t\t\tcontinue;\n \n \t\tid = new_rerere_id(sha1);\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex 8417e5a4b1..23f9c0ca45 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -580,4 +580,25 @@ test_expect_success 'multiple identical conflicts' '\n \tcount_pre_post 0 0\n '\n \n+test_expect_success 'rerere with unexpected conflict markers does not crash' '\n+\tgit reset --hard &&\n+\n+\tgit checkout -b branch-1 master &&\n+\techo \"bar\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m two &&\n+\n+\tgit reset --hard &&\n+\tgit checkout -b branch-2 master &&\n+\techo \"foo\" >test &&\n+\tgit add test &&\n+\tgit commit -q -a -m one &&\n+\n+\ttest_must_fail git merge branch-1 &&\n+\techo \"<<<<<<< a\" >test &&\n+\tgit rerere &&\n+\n+\tgit rerere clear\n+'\n+\n test_done\n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354559","messageId":"20180805172037.12530-8-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 07/11] rerere: only return whether a path has conflicts or not","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:33Z","receivedAt":"2018-08-05T17:20:59Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"We currently return the exact number of conflict hunks a certain path\nhas from the 'handle_paths' function.  However all of its callers only\ncare whether there are conflicts or not or if there is an error.\nReturn only that information, and document that only that information\nis returned.  This will simplify the code in the subsequent steps.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 23 ++++++++++++-----------\n 1 file changed, 12 insertions(+), 11 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 895ad80c0c..bf803043e2 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -393,12 +393,13 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n  * one side of the conflict, NUL, the other side of the conflict,\n  * and NUL concatenated together.\n  *\n- * Return the number of conflict hunks found.\n+ * Return 1 if conflict hunks are found, 0 if there are no conflict\n+ * hunks and -1 if an error occured.\n  */\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n \tgit_SHA_CTX ctx;\n-\tint hunk_no = 0;\n+\tint has_conflicts = 0;\n \tenum {\n \t\tRR_CONTEXT = 0, RR_SIDE_1, RR_SIDE_2, RR_ORIGINAL\n \t} hunk = RR_CONTEXT;\n@@ -426,7 +427,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\t\t\tgoto bad;\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n-\t\t\thunk_no++;\n+\t\t\thas_conflicts = 1;\n \t\t\thunk = RR_CONTEXT;\n \t\t\trerere_io_putconflict('<', marker_size, io);\n \t\t\trerere_io_putmem(one.buf, one.len, io);\n@@ -462,7 +463,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n \t\tgit_SHA1_Final(sha1, &ctx);\n \tif (hunk != RR_CONTEXT)\n \t\treturn -1;\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n /*\n@@ -471,7 +472,7 @@ static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_siz\n  */\n static int handle_file(const char *path, unsigned char *sha1, const char *output)\n {\n-\tint hunk_no = 0;\n+\tint has_conflicts = 0;\n \tstruct rerere_io_file io;\n \tint marker_size = ll_merge_marker_size(path);\n \n@@ -491,7 +492,7 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \t\t}\n \t}\n \n-\thunk_no = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n+\thas_conflicts = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n \n \tfclose(io.input);\n \tif (io.io.wrerror)\n@@ -500,14 +501,14 @@ static int handle_file(const char *path, unsigned char *sha1, const char *output\n \tif (io.io.output && fclose(io.io.output))\n \t\tio.io.wrerror = error_errno(_(\"failed to flush '%s'\"), path);\n \n-\tif (hunk_no < 0) {\n+\tif (has_conflicts < 0) {\n \t\tif (output)\n \t\t\tunlink_or_warn(output);\n \t\treturn error(_(\"could not parse conflict hunks in '%s'\"), path);\n \t}\n \tif (io.io.wrerror)\n \t\treturn -1;\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n /*\n@@ -954,7 +955,7 @@ static int handle_cache(const char *path, unsigned char *sha1, const char *outpu\n \tmmfile_t mmfile[3] = {{NULL}};\n \tmmbuffer_t result = {NULL, 0};\n \tconst struct cache_entry *ce;\n-\tint pos, len, i, hunk_no;\n+\tint pos, len, i, has_conflicts;\n \tstruct rerere_io_mem io;\n \tint marker_size = ll_merge_marker_size(path);\n \n@@ -1008,11 +1009,11 @@ static int handle_cache(const char *path, unsigned char *sha1, const char *outpu\n \t * Grab the conflict ID and optionally write the original\n \t * contents with conflict markers out.\n \t */\n-\thunk_no = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n+\thas_conflicts = handle_path(sha1, (struct rerere_io *)&io, marker_size);\n \tstrbuf_release(&io.input);\n \tif (io.io.output)\n \t\tfclose(io.io.output);\n-\treturn hunk_no;\n+\treturn has_conflicts;\n }\n \n static int rerere_forget_one_path(const char *path, struct string_list *rr)\n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354560","messageId":"20180805172037.12530-9-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 08/11] rerere: factor out handle_conflict function","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:34Z","receivedAt":"2018-08-05T17:21:01Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Factor out the handle_conflict function, which handles a single\nconflict in a path.  This is in preparation for a subsequent commit,\nwhere this function will be re-used.\n\nNote that this does change the behaviour of 'git rerere' slightly.\nWhere previously we'd consider all files where an unmatched conflict\nmarker is found as invalid, we now only consider files invalid when\nthe \"ours\" conflict marker (\"<<<<<<< <text>\") is unmatched, not when\nother conflict markers (e.g. \"=======\") is unmatched.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 87 ++++++++++++++++++++++++++++++--------------------------\n 1 file changed, 47 insertions(+), 40 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex bf803043e2..2d62251943 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -384,85 +384,92 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n \treturn isspace(*buf);\n }\n \n-/*\n- * Read contents a file with conflicts, normalize the conflicts\n- * by (1) discarding the common ancestor version in diff3-style,\n- * (2) reordering our side and their side so that whichever sorts\n- * alphabetically earlier comes before the other one, while\n- * computing the \"conflict ID\", which is just an SHA-1 hash of\n- * one side of the conflict, NUL, the other side of the conflict,\n- * and NUL concatenated together.\n- *\n- * Return 1 if conflict hunks are found, 0 if there are no conflict\n- * hunks and -1 if an error occured.\n- */\n-static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n+static int handle_conflict(struct rerere_io *io, int marker_size, git_SHA_CTX *ctx)\n {\n-\tgit_SHA_CTX ctx;\n-\tint has_conflicts = 0;\n \tenum {\n-\t\tRR_CONTEXT = 0, RR_SIDE_1, RR_SIDE_2, RR_ORIGINAL\n-\t} hunk = RR_CONTEXT;\n+\t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n+\t} hunk = RR_SIDE_1;\n \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n \tstruct strbuf buf = STRBUF_INIT;\n-\n-\tif (sha1)\n-\t\tgit_SHA1_Init(&ctx);\n+\tint has_conflicts = -1;\n \n \twhile (!io->getline(&buf, io)) {\n \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n-\t\t\tif (hunk != RR_CONTEXT)\n-\t\t\t\tgoto bad;\n-\t\t\thunk = RR_SIDE_1;\n+\t\t\tbreak;\n \t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1)\n-\t\t\t\tgoto bad;\n+\t\t\t\tbreak;\n \t\t\thunk = RR_ORIGINAL;\n \t\t} else if (is_cmarker(buf.buf, '=', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1 && hunk != RR_ORIGINAL)\n-\t\t\t\tgoto bad;\n+\t\t\t\tbreak;\n \t\t\thunk = RR_SIDE_2;\n \t\t} else if (is_cmarker(buf.buf, '>', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_2)\n-\t\t\t\tgoto bad;\n+\t\t\t\tbreak;\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n \t\t\thas_conflicts = 1;\n-\t\t\thunk = RR_CONTEXT;\n \t\t\trerere_io_putconflict('<', marker_size, io);\n \t\t\trerere_io_putmem(one.buf, one.len, io);\n \t\t\trerere_io_putconflict('=', marker_size, io);\n \t\t\trerere_io_putmem(two.buf, two.len, io);\n \t\t\trerere_io_putconflict('>', marker_size, io);\n-\t\t\tif (sha1) {\n-\t\t\t\tgit_SHA1_Update(&ctx, one.buf ? one.buf : \"\",\n+\t\t\tif (ctx) {\n+\t\t\t\tgit_SHA1_Update(ctx, one.buf ? one.buf : \"\",\n \t\t\t\t\t    one.len + 1);\n-\t\t\t\tgit_SHA1_Update(&ctx, two.buf ? two.buf : \"\",\n+\t\t\t\tgit_SHA1_Update(ctx, two.buf ? two.buf : \"\",\n \t\t\t\t\t    two.len + 1);\n \t\t\t}\n-\t\t\tstrbuf_reset(&one);\n-\t\t\tstrbuf_reset(&two);\n+\t\t\tbreak;\n \t\t} else if (hunk == RR_SIDE_1)\n \t\t\tstrbuf_addbuf(&one, &buf);\n \t\telse if (hunk == RR_ORIGINAL)\n \t\t\t; /* discard */\n \t\telse if (hunk == RR_SIDE_2)\n \t\t\tstrbuf_addbuf(&two, &buf);\n-\t\telse\n-\t\t\trerere_io_putstr(buf.buf, io);\n-\t\tcontinue;\n-\tbad:\n-\t\thunk = 99; /* force error exit */\n-\t\tbreak;\n \t}\n \tstrbuf_release(&one);\n \tstrbuf_release(&two);\n \tstrbuf_release(&buf);\n \n+\treturn has_conflicts;\n+}\n+\n+/*\n+ * Read contents a file with conflicts, normalize the conflicts\n+ * by (1) discarding the common ancestor version in diff3-style,\n+ * (2) reordering our side and their side so that whichever sorts\n+ * alphabetically earlier comes before the other one, while\n+ * computing the \"conflict ID\", which is just an SHA-1 hash of\n+ * one side of the conflict, NUL, the other side of the conflict,\n+ * and NUL concatenated together.\n+ *\n+ * Return 1 if conflict hunks are found, 0 if there are no conflict\n+ * hunks and -1 if an error occured.\n+ */\n+static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n+{\n+\tgit_SHA_CTX ctx;\n+\tstruct strbuf buf = STRBUF_INIT;\n+\tint has_conflicts = 0;\n+\tif (sha1)\n+\t\tgit_SHA1_Init(&ctx);\n+\n+\twhile (!io->getline(&buf, io)) {\n+\t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n+\t\t\thas_conflicts = handle_conflict(io, marker_size,\n+\t\t\t\t\t\t\tsha1 ? &ctx : NULL);\n+\t\t\tif (has_conflicts < 0)\n+\t\t\t\tbreak;\n+\t\t} else\n+\t\t\trerere_io_putstr(buf.buf, io);\n+\t}\n+\tstrbuf_release(&buf);\n+\n \tif (sha1)\n \t\tgit_SHA1_Final(sha1, &ctx);\n-\tif (hunk != RR_CONTEXT)\n-\t\treturn -1;\n+\n \treturn has_conflicts;\n }\n \n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354561","messageId":"20180805172037.12530-10-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 09/11] rerere: return strbuf from handle path","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:35Z","receivedAt":"2018-08-05T17:21:03Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently we write the conflict to disk directly in the handle_path\nfunction.  To make it re-usable for nested conflicts, instead of\nwriting the conflict out directly, store it in a strbuf and let the\ncaller write it out.\n\nThis does mean some slight increase in memory usage, however that\nincrease is limited to the size of the largest conflict we've\ncurrently processed.  We already keep one copy of the conflict in\nmemory, and it shouldn't be too large, so the increase in memory usage\nseems acceptable.\n\nAs a bonus this lets us get replace the rerere_io_putconflict function\nwith a trivial two line function.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c | 58 ++++++++++++++++++--------------------------------------\n 1 file changed, 18 insertions(+), 40 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex 2d62251943..a35b88916c 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -302,38 +302,6 @@ static void rerere_io_putstr(const char *str, struct rerere_io *io)\n \t\tferr_puts(str, io->output, &io->wrerror);\n }\n \n-/*\n- * Write a conflict marker to io->output (if defined).\n- */\n-static void rerere_io_putconflict(int ch, int size, struct rerere_io *io)\n-{\n-\tchar buf[64];\n-\n-\twhile (size) {\n-\t\tif (size <= sizeof(buf) - 2) {\n-\t\t\tmemset(buf, ch, size);\n-\t\t\tbuf[size] = '\\n';\n-\t\t\tbuf[size + 1] = '\\0';\n-\t\t\tsize = 0;\n-\t\t} else {\n-\t\t\tint sz = sizeof(buf) - 1;\n-\n-\t\t\t/*\n-\t\t\t * Make sure we will not write everything out\n-\t\t\t * in this round by leaving at least 1 byte\n-\t\t\t * for the next round, giving the next round\n-\t\t\t * a chance to add the terminating LF.  Yuck.\n-\t\t\t */\n-\t\t\tif (size <= sz)\n-\t\t\t\tsz -= (sz - size) + 1;\n-\t\t\tmemset(buf, ch, sz);\n-\t\t\tbuf[sz] = '\\0';\n-\t\t\tsize -= sz;\n-\t\t}\n-\t\trerere_io_putstr(buf, io);\n-\t}\n-}\n-\n static void rerere_io_putmem(const char *mem, size_t sz, struct rerere_io *io)\n {\n \tif (io->output)\n@@ -384,7 +352,14 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)\n \treturn isspace(*buf);\n }\n \n-static int handle_conflict(struct rerere_io *io, int marker_size, git_SHA_CTX *ctx)\n+static void rerere_strbuf_putconflict(struct strbuf *buf, int ch, size_t size)\n+{\n+\tstrbuf_addchars(buf, ch, size);\n+\tstrbuf_addch(buf, '\\n');\n+}\n+\n+static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n+\t\t\t   int marker_size, git_SHA_CTX *ctx)\n {\n \tenum {\n \t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n@@ -410,11 +385,11 @@ static int handle_conflict(struct rerere_io *io, int marker_size, git_SHA_CTX *c\n \t\t\tif (strbuf_cmp(&one, &two) > 0)\n \t\t\t\tstrbuf_swap(&one, &two);\n \t\t\thas_conflicts = 1;\n-\t\t\trerere_io_putconflict('<', marker_size, io);\n-\t\t\trerere_io_putmem(one.buf, one.len, io);\n-\t\t\trerere_io_putconflict('=', marker_size, io);\n-\t\t\trerere_io_putmem(two.buf, two.len, io);\n-\t\t\trerere_io_putconflict('>', marker_size, io);\n+\t\t\trerere_strbuf_putconflict(out, '<', marker_size);\n+\t\t\tstrbuf_addbuf(out, &one);\n+\t\t\trerere_strbuf_putconflict(out, '=', marker_size);\n+\t\t\tstrbuf_addbuf(out, &two);\n+\t\t\trerere_strbuf_putconflict(out, '>', marker_size);\n \t\t\tif (ctx) {\n \t\t\t\tgit_SHA1_Update(ctx, one.buf ? one.buf : \"\",\n \t\t\t\t\t    one.len + 1);\n@@ -451,21 +426,24 @@ static int handle_conflict(struct rerere_io *io, int marker_size, git_SHA_CTX *c\n static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)\n {\n \tgit_SHA_CTX ctx;\n-\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf buf = STRBUF_INIT, out = STRBUF_INIT;\n \tint has_conflicts = 0;\n \tif (sha1)\n \t\tgit_SHA1_Init(&ctx);\n \n \twhile (!io->getline(&buf, io)) {\n \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n-\t\t\thas_conflicts = handle_conflict(io, marker_size,\n+\t\t\thas_conflicts = handle_conflict(&out, io, marker_size,\n \t\t\t\t\t\t\tsha1 ? &ctx : NULL);\n \t\t\tif (has_conflicts < 0)\n \t\t\t\tbreak;\n+\t\t\trerere_io_putmem(out.buf, out.len, io);\n+\t\t\tstrbuf_reset(&out);\n \t\t} else\n \t\t\trerere_io_putstr(buf.buf, io);\n \t}\n \tstrbuf_release(&buf);\n+\tstrbuf_release(&out);\n \n \tif (sha1)\n \t\tgit_SHA1_Final(sha1, &ctx);\n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354562","messageId":"20180805172037.12530-11-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:36Z","receivedAt":"2018-08-05T17:21:05Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently rerere can't handle nested conflicts and will error out when\nit encounters such conflicts.  Do that by recursively calling the\n'handle_conflict' function to normalize the conflict.\n\nNote that a conflict like this would only be produced if a user\ncommits a file with conflict markers, and gets a conflict including\nthat in a susbsequent operation.\n\nThe conflict ID calculation here deserves some explanation:\n\nAs we are using the same handle_conflict function, the nested conflict\nis normalized the same way as for non-nested conflicts, which means\nthe ancestor in the diff3 case is stripped out, and the parts of the\nconflict are ordered alphabetically.\n\nThe conflict ID is however is only calculated in the top level\nhandle_conflict call, so it will include the markers that 'rerere'\nadds to the output.  e.g. say there's the following conflict:\n\n    <<<<<<< HEAD\n    1\n    =======\n    <<<<<<< HEAD\n    3\n    =======\n    2\n    >>>>>>> branch-2\n    >>>>>>> branch-3~\n\nit would be recorde as follows in the preimage:\n\n    <<<<<<<\n    1\n    =======\n    <<<<<<<\n    2\n    =======\n    3\n    >>>>>>>\n    >>>>>>>\n\nand the conflict ID would be calculated as\n\n    sha1(1<NUL><<<<<<<\n    2\n    =======\n    3\n    >>>>>>><NUL>)\n\nStripping out vs. leaving the conflict markers in place in the inner\nconflict should have no practical impact, but it simplifies the\nimplementation.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n Documentation/technical/rerere.txt | 42 ++++++++++++++++++++++++++++++\n rerere.c                           | 10 +++++--\n t/t4200-rerere.sh                  | 37 ++++++++++++++++++++++++++\n 3 files changed, 87 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\nindex 3d10dbfa67..e65ba9b0c6 100644\n--- a/Documentation/technical/rerere.txt\n+++ b/Documentation/technical/rerere.txt\n@@ -138,3 +138,45 @@ SHA1('B<NUL>C<NUL>').\n If there are multiple conflicts in one file, the sha1 is calculated\n the same way with all hunks appended to each other, in the order in\n which they appear in the file, separated by a <NUL> character.\n+\n+Nested conflicts\n+~~~~~~~~~~~~~~~~\n+\n+Nested conflicts are handled very similarly to \"simple\" conflicts.\n+Similar to simple conflicts, the conflict is first normalized by\n+stripping the labels from conflict markers, stripping the common ancestor\n+version, and the sorting the conflict hunks, both for the outer and the\n+inner conflict.  This is done recursively, so any number of nested\n+conflicts can be handled.\n+\n+The only difference is in how the conflict ID is calculated.  For the\n+inner conflict, the conflict markers themselves are not stripped out\n+before calculating the sha1.\n+\n+Say we have the following conflict for example:\n+\n+    <<<<<<< HEAD\n+    1\n+    =======\n+    <<<<<<< HEAD\n+    3\n+    =======\n+    2\n+    >>>>>>> branch-2\n+    >>>>>>> branch-3~\n+\n+After stripping out the labels of the conflict markers, and sorting\n+the hunks, the conflict would look as follows:\n+\n+    <<<<<<<\n+    1\n+    =======\n+    <<<<<<<\n+    2\n+    =======\n+    3\n+    >>>>>>>\n+    >>>>>>>\n+\n+and finally the conflict ID would be calculated as:\n+`sha1('1<NUL><<<<<<<\\n3\\n=======\\n2\\n>>>>>>><NUL>')`\ndiff --git a/rerere.c b/rerere.c\nindex a35b88916c..f78bef80b1 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -365,12 +365,18 @@ static int handle_conflict(struct strbuf *out, struct rerere_io *io,\n \t\tRR_SIDE_1 = 0, RR_SIDE_2, RR_ORIGINAL\n \t} hunk = RR_SIDE_1;\n \tstruct strbuf one = STRBUF_INIT, two = STRBUF_INIT;\n-\tstruct strbuf buf = STRBUF_INIT;\n+\tstruct strbuf buf = STRBUF_INIT, conflict = STRBUF_INIT;\n \tint has_conflicts = -1;\n \n \twhile (!io->getline(&buf, io)) {\n \t\tif (is_cmarker(buf.buf, '<', marker_size)) {\n-\t\t\tbreak;\n+\t\t\tif (handle_conflict(&conflict, io, marker_size, NULL) < 0)\n+\t\t\t\tbreak;\n+\t\t\tif (hunk == RR_SIDE_1)\n+\t\t\t\tstrbuf_addbuf(&one, &conflict);\n+\t\t\telse\n+\t\t\t\tstrbuf_addbuf(&two, &conflict);\n+\t\t\tstrbuf_release(&conflict);\n \t\t} else if (is_cmarker(buf.buf, '|', marker_size)) {\n \t\t\tif (hunk != RR_SIDE_1)\n \t\t\t\tbreak;\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex 23f9c0ca45..afaf085e42 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -601,4 +601,41 @@ test_expect_success 'rerere with unexpected conflict markers does not crash' '\n \tgit rerere clear\n '\n \n+test_expect_success 'rerere with inner conflict markers' '\n+\tgit reset --hard &&\n+\n+\tgit checkout -b A master &&\n+\techo \"bar\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m two &&\n+\techo \"baz\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m three &&\n+\n+\tgit reset --hard &&\n+\tgit checkout -b B master &&\n+\techo \"foo\" >test &&\n+\tgit add test &&\n+\tgit commit -q -a -m one &&\n+\n+\ttest_must_fail git merge A~ &&\n+\tgit add test &&\n+\tgit commit -q -m \"will solve conflicts later\" &&\n+\ttest_must_fail git merge A &&\n+\n+\techo \"resolved\" >test &&\n+\tgit add test &&\n+\tgit commit -q -m \"solved conflict\" &&\n+\n+\techo \"resolved\" >expect &&\n+\n+\tgit reset --hard HEAD~~ &&\n+\ttest_must_fail git merge A~ &&\n+\tgit add test &&\n+\tgit commit -q -m \"will solve conflicts later\" &&\n+\ttest_must_fail git merge A &&\n+\tcat test >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"354563","messageId":"20180805172037.12530-12-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-1-t.gummerer@gmail.com","subject":"[PATCH v4 11/11] rerere: recalculate conflict ID when unresolved conflict is committed","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-05T17:20:37Z","receivedAt":"2018-08-05T17:21:05Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"Currently when a user doesn't resolve a conflict, commits the results,\nand does an operation which creates another conflict, rerere will use\nthe ID of the previously unresolved conflict for the new conflict.\nThis is because the conflict is kept in the MERGE_RR file, which\n'rerere' reads every time it is invoked.\n\nAfter the new conflict is solved, rerere will record the resolution\nwith the ID of the old conflict.  So in order to replay the conflict,\nboth merges would have to be re-done, instead of just the last one, in\norder for rerere to be able to automatically resolve the conflict.\n\nInstead of that, assign a new conflict ID if there are still conflicts\nin a file and the file had conflicts at a previous step.  This ID\nmatches the conflict we actually resolved at the corresponding step.\n\nNote that there are no backwards compatibility worries here, as rerere\nwould have failed to even normalize the conflict before this patch\nseries.\n\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n rerere.c          | 7 +++----\n t/t4200-rerere.sh | 7 +++++++\n 2 files changed, 10 insertions(+), 4 deletions(-)\n\ndiff --git a/rerere.c b/rerere.c\nindex f78bef80b1..dd81d09e19 100644\n--- a/rerere.c\n+++ b/rerere.c\n@@ -815,7 +815,7 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\tstruct rerere_id *id;\n \t\tunsigned char sha1[20];\n \t\tconst char *path = conflict.items[i].string;\n-\t\tint ret, has_string;\n+\t\tint ret;\n \n \t\t/*\n \t\t * Ask handle_file() to scan and assign a\n@@ -823,12 +823,11 @@ static int do_plain_rerere(struct string_list *rr, int fd)\n \t\t * yet.\n \t\t */\n \t\tret = handle_file(path, sha1, NULL);\n-\t\thas_string = string_list_has_string(rr, path);\n-\t\tif (ret < 0 && has_string) {\n+\t\tif (ret != 0 && string_list_has_string(rr, path)) {\n \t\t\tremove_variant(string_list_lookup(rr, path)->util);\n \t\t\tstring_list_remove(rr, path, 1);\n \t\t}\n-\t\tif (ret < 1 || has_string)\n+\t\tif (ret < 1)\n \t\t\tcontinue;\n \n \t\tid = new_rerere_id(sha1);\ndiff --git a/t/t4200-rerere.sh b/t/t4200-rerere.sh\nindex afaf085e42..819f6dd672 100755\n--- a/t/t4200-rerere.sh\n+++ b/t/t4200-rerere.sh\n@@ -635,6 +635,13 @@ test_expect_success 'rerere with inner conflict markers' '\n \tgit commit -q -m \"will solve conflicts later\" &&\n \ttest_must_fail git merge A &&\n \tcat test >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tgit add test &&\n+\tgit commit -m \"rerere solved conflict\" &&\n+\tgit reset --hard HEAD~ &&\n+\ttest_must_fail git merge A &&\n+\tcat test >actual &&\n \ttest_cmp expect actual\n '\n \n-- \n2.18.0.720.gf7a957e2e7\n\n"},{"id":"356239","messageId":"CACBZZX6xvsZ4K86b53ura6zENs2p0SBjwYYG=h0TNem3wnEbuQ@mail.gmail.com","threadId":"48531","inReplyTo":"20180805172037.12530-11-t.gummerer@gmail.com","subject":"Re: [PATCH v4 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2018-08-22T11:00:55Z","receivedAt":"2018-08-22T11:01:11Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Sun, Aug 5, 2018 at 7:23 PM Thomas Gummerer <t.gummerer@gmail.com> wrote:\n\nLate reply since I just saw this in next.\n\n> Currently rerere can't handle nested conflicts and will error out when\n> it encounters such conflicts.  Do that by recursively calling the\n> 'handle_conflict' function to normalize the conflict.\n> [...]\n\nMakes sense.\n\n> --- a/Documentation/technical/rerere.txt\n> +++ b/Documentation/technical/rerere.txt\n\nBut why not add this to the git-rerere manpage? These technical docs\nget way less exposure, and in this case we're not describing some\ninterna implementation detail, which the technical docs are for, but\nsomething that's user-visible, let's put that in  the user-visiblee\ndocs.\n"},{"id":"356273","messageId":"xmqqsh365qt0.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"CACBZZX6xvsZ4K86b53ura6zENs2p0SBjwYYG=h0TNem3wnEbuQ@mail.gmail.com","subject":"Re: [PATCH v4 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-22T16:06:35Z","receivedAt":"2018-08-22T16:06:40Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n> But why not add this to the git-rerere manpage? These technical docs\n> get way less exposure, and in this case we're not describing some\n> interna implementation detail, which the technical docs are for, but\n> something that's user-visible, let's put that in  the user-visiblee\n> docs.\n\nI actually consider that the documentation describes low-level\ninternal implementation detail, which the end users do not care nor\nneed to know in order to make use of \"rerere\".  How would it help\nthe end-users to know that the common ancestor portion of diff3\nstyle conflict does not participate in conflict identification,\nsides of conflicts sometimes get swapped for easier indexing of\nconflicts, or conflict shapes are hashed via SHA-1 to determine\nwhich subdirectory of $GIT_DIR/rr-cache/ to use to store it, etc.?\n\nBy the way, I just noticed that what the last section (i.e. nested\nconflicts) says is completely bogus.  Nested conflicts are handled\nby lengthening markers for conflict in inner-merge and paying\nattention only to the outermost merge.  The only case where the\nconflict markers can appear in the way depicted in the section is\nwhen the contents from branches being merged had these conflict\nmarker looking strings from the beginning---that's \"doctor it hurts\nwhen I do this---don't do it then\" situation.  The section may\ndescribe correctly what the code happens to do when it gets thrown\nsuch a garbage at, but I do not think it is a useful piece of\ninformation about a designed behaviour.\n\n"},{"id":"356303","messageId":"20180822203451.GG13316@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqqsh365qt0.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v4 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-22T20:34:51Z","receivedAt":"2018-08-22T20:34:56Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 08/22, Junio C Hamano wrote:\n> Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n> \n> > But why not add this to the git-rerere manpage? These technical docs\n> > get way less exposure, and in this case we're not describing some\n> > interna implementation detail, which the technical docs are for, but\n> > something that's user-visible, let's put that in  the user-visiblee\n> > docs.\n> \n> I actually consider that the documentation describes low-level\n> internal implementation detail, which the end users do not care nor\n> need to know in order to make use of \"rerere\".  How would it help\n> the end-users to know that the common ancestor portion of diff3\n> style conflict does not participate in conflict identification,\n> sides of conflicts sometimes get swapped for easier indexing of\n> conflicts, or conflict shapes are hashed via SHA-1 to determine\n> which subdirectory of $GIT_DIR/rr-cache/ to use to store it, etc.?\n\nAgreed, I don't think this would be very helpful for users.\n\n> By the way, I just noticed that what the last section (i.e. nested\n> conflicts) says is completely bogus.  Nested conflicts are handled\n> by lengthening markers for conflict in inner-merge and paying\n> attention only to the outermost merge.  The only case where the\n> conflict markers can appear in the way depicted in the section is\n> when the contents from branches being merged had these conflict\n> marker looking strings from the beginning---that's \"doctor it hurts\n> when I do this---don't do it then\" situation.  The section may\n> describe correctly what the code happens to do when it gets thrown\n> such a garbage at, but I do not think it is a useful piece of\n> information about a designed behaviour.\n\nHmm, it does describe what happens in the code, which is what this\npatch implements.  Maybe we should rephrase the title here?\n\nOr are you suggesting dropping this patch (and the next one)\ncompletely, as we don't want to try and handle the case where this\nkind of garbage is thrown at 'rerere'?  I don't think it would make\nsense to drop this documentation without dropping the patch itself, as\nit does document how rerere handles this case.  Without this bit of\ndocumentation (but with the code in this patch), the technical\n'rerere' documentation feels incomplete to me.\n"},{"id":"356305","messageId":"xmqq4lfmm7pb.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180822203451.GG13316@hank.intra.tgummerer.com","subject":"Re: [PATCH v4 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-22T21:07:12Z","receivedAt":"2018-08-22T21:07:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Hmm, it does describe what happens in the code, which is what this\n> patch implements.  Maybe we should rephrase the title here?\n>\n> Or are you suggesting dropping this patch (and the next one)\n> completely, as we don't want to try and handle the case where this\n> kind of garbage is thrown at 'rerere'?\n\nI consider these two patches as merely attempting to punt a bit\nbetter.  Once users start committing conflict-marker-looking lines\nin the contents, and getting them involved in actual conflicts, I do\nnot think any approach (including what the original rerere uses\nbefore this patch) that assumes the markers will neatly form set of\nblocks of text enclosed in << == >> will reliably step around such\nbroken contents.  E.g. it is entirely conceivable both branches have\nthe <<< beginning of conflict marker plus contents from the HEAD\nbefore they recorded the marker that are identical, that diverge as\nyou scan the text down and get closer to ===, something like:\n\n        side A                  side B\n        --------------------    --------------------\n\n        shared                  shared\n        <<<<<<<                 <<<<<<<\n        version before          version before\n        these guys merged       these guys merged\n        their ancestor          their ancestor\n        versions                versions.\n        but some                now some\n        lines are different     lines are different\n        =======                 ========\n        and other               totally different\n        contents                contents\n        ...                     ...\n\nAnd a merge of these may make <<< part shared (i.e. outside the\nconflicted region) while lines near and below ==== part of conflict,\nwhich would give us something like\n\n        merge of side A & B\n        -------------------\n\n        shared                  \n        <<<<<<<                 (this is part of contents)\n        version before          \n        these guys merged       \n        their ancestor          \n        <<<<<<< HEAD            (conflict marker)\n        versions\n        but some\n        lines are different\n        =======                 (this is part of contents)\n        and other\n        contents\n        ...\n        =======                 (conflict marker)\n        versions.\n        now some\n        lines are different\n        =======                 (this is part of contents)\n        totally different\n        contents\n        ...\n        >>>>>>> theirs          (conflict marker)\n\nDepending on the shape of the original conflict that was committed,\nwe may have two versions of <<<, together with the real conflict\nmarker, but shared closing >>> marker.  With contents like that,\nthere is no way for us to split these lines into two groups at a\nline '=====' (which one?) and swap to come up with the normalized\nshape.\n\nThe original rerere algorithm would punt when such an unmatched\nmarkers are found, and deals with \"nested conflict\" situation by\navoiding to create such a thing altogether.  I am sure your two\npatches may make the code punt less, but I suspect that is not a\nfoolproof \"solution\" but more of a workaround, as I do not think it\nis solvable, once you allow users to commit conflict-marker looking\nstrings in contents.  As the heuristics used in such a workaround\nare very likely to change, and something the end-users should not\neven rely on, I'd rather not document and promise the exact\nbehaviour---perhaps we should stress \"don't do that\" even stronger\ninstead.\n"},{"id":"356493","messageId":"20180824215619.GH13316@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqq4lfmm7pb.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v4 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-24T21:56:19Z","receivedAt":"2018-08-24T21:56:25Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 08/22, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > Hmm, it does describe what happens in the code, which is what this\n> > patch implements.  Maybe we should rephrase the title here?\n> >\n> > Or are you suggesting dropping this patch (and the next one)\n> > completely, as we don't want to try and handle the case where this\n> > kind of garbage is thrown at 'rerere'?\n> \n> I consider these two patches as merely attempting to punt a bit\n> better.  Once users start committing conflict-marker-looking lines\n> in the contents, and getting them involved in actual conflicts, I do\n> not think any approach (including what the original rerere uses\n> before this patch) that assumes the markers will neatly form set of\n> blocks of text enclosed in << == >> will reliably step around such\n> broken contents.  E.g. it is entirely conceivable both branches have\n> the <<< beginning of conflict marker plus contents from the HEAD\n> before they recorded the marker that are identical, that diverge as\n> you scan the text down and get closer to ===, something like:\n> \n>         side A                  side B\n>         --------------------    --------------------\n> \n>         shared                  shared\n>         <<<<<<<                 <<<<<<<\n>         version before          version before\n>         these guys merged       these guys merged\n>         their ancestor          their ancestor\n>         versions                versions.\n>         but some                now some\n>         lines are different     lines are different\n>         =======                 ========\n>         and other               totally different\n>         contents                contents\n>         ...                     ...\n> \n> And a merge of these may make <<< part shared (i.e. outside the\n> conflicted region) while lines near and below ==== part of conflict,\n> which would give us something like\n> \n>         merge of side A & B\n>         -------------------\n> \n>         shared                  \n>         <<<<<<<                 (this is part of contents)\n>         version before          \n>         these guys merged       \n>         their ancestor          \n>         <<<<<<< HEAD            (conflict marker)\n>         versions\n>         but some\n>         lines are different\n>         =======                 (this is part of contents)\n>         and other\n>         contents\n>         ...\n>         =======                 (conflict marker)\n>         versions.\n>         now some\n>         lines are different\n>         =======                 (this is part of contents)\n>         totally different\n>         contents\n>         ...\n>         >>>>>>> theirs          (conflict marker)\n> \n> Depending on the shape of the original conflict that was committed,\n> we may have two versions of <<<, together with the real conflict\n> marker, but shared closing >>> marker.  With contents like that,\n> there is no way for us to split these lines into two groups at a\n> line '=====' (which one?) and swap to come up with the normalized\n> shape.\n> \n> The original rerere algorithm would punt when such an unmatched\n> markers are found, and deals with \"nested conflict\" situation by\n> avoiding to create such a thing altogether.  I am sure your two\n> patches may make the code punt less, but I suspect that is not a\n> foolproof \"solution\" but more of a workaround, as I do not think it\n> is solvable, once you allow users to commit conflict-marker looking\n> strings in contents.\n\nAgreed.  I think it may be solvable if we'd actually get the\ninformation about what belongs to which side from the merge algorithm\ndirectly.  But that sounds way more involved than what I'm able to\ncommit to for something that I don't forsee running into myself :)\n\n>                       As the heuristics used in such a workaround\n> are very likely to change, and something the end-users should not\n> even rely on, I'd rather not document and promise the exact\n> behaviour---perhaps we should stress \"don't do that\" even stronger\n> instead.\n\nFair enough.  I thought of the technical documentation as something\nthat doesn't promise users anything, but rather describes how the\ninternals work right now, which is what this bit of documentation\nattempted to write down.  But if we are worried about this giving end\nusers ideas then I definitely agree and we should get rid of this bit\nof documentation.  I'll send a patch for that, and for adding a note\nabout \"don't do that\" in the man page.\n"},{"id":"356494","messageId":"20180824221005.5983-1-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180824215619.GH13316@hank.intra.tgummerer.com","subject":"[PATCH 1/2] rerere: remove documentation for \"nested conflicts\"","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-24T22:10:04Z","receivedAt":"2018-08-24T22:10:14Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"4af32207bc (\"rerere: teach rerere to handle nested conflicts\",\n2018-08-05) introduced slightly better behaviour if the user commits\nconflict markers and then gets another conflict in 'git rerere'.\nHowever this is just a heuristic to punt on such conflicts better, and\nthe documentation might be misleading to users, in case we change the\nheuristic in the future.\n\nRemove this documentation to avoid being potentially misleading in the\ndocumentation.\n\nSuggested-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n\nThe original series already made it into 'next', so these patches are\non top of that.  I also see it is marked as \"will merge to master\" in\nthe \"What's cooking\" email, so these two patches would be on top of\nthat.  If you are not planning to merge the series down to master\nbefore 2.19, we could squash this into 10/11, otherwise I'm happy with\nthe patches on top.\n\n Documentation/technical/rerere.txt | 42 ------------------------------\n 1 file changed, 42 deletions(-)\n\ndiff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\nindex e65ba9b0c6..3d10dbfa67 100644\n--- a/Documentation/technical/rerere.txt\n+++ b/Documentation/technical/rerere.txt\n@@ -138,45 +138,3 @@ SHA1('B<NUL>C<NUL>').\n If there are multiple conflicts in one file, the sha1 is calculated\n the same way with all hunks appended to each other, in the order in\n which they appear in the file, separated by a <NUL> character.\n-\n-Nested conflicts\n-~~~~~~~~~~~~~~~~\n-\n-Nested conflicts are handled very similarly to \"simple\" conflicts.\n-Similar to simple conflicts, the conflict is first normalized by\n-stripping the labels from conflict markers, stripping the common ancestor\n-version, and the sorting the conflict hunks, both for the outer and the\n-inner conflict.  This is done recursively, so any number of nested\n-conflicts can be handled.\n-\n-The only difference is in how the conflict ID is calculated.  For the\n-inner conflict, the conflict markers themselves are not stripped out\n-before calculating the sha1.\n-\n-Say we have the following conflict for example:\n-\n-    <<<<<<< HEAD\n-    1\n-    =======\n-    <<<<<<< HEAD\n-    3\n-    =======\n-    2\n-    >>>>>>> branch-2\n-    >>>>>>> branch-3~\n-\n-After stripping out the labels of the conflict markers, and sorting\n-the hunks, the conflict would look as follows:\n-\n-    <<<<<<<\n-    1\n-    =======\n-    <<<<<<<\n-    2\n-    =======\n-    3\n-    >>>>>>>\n-    >>>>>>>\n-\n-and finally the conflict ID would be calculated as:\n-`sha1('1<NUL><<<<<<<\\n3\\n=======\\n2\\n>>>>>>><NUL>')`\n-- \n2.18.0.1088.ge017bf2cd1\n\n"},{"id":"356495","messageId":"20180824221005.5983-2-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180824221005.5983-1-t.gummerer@gmail.com","subject":"[PATCH 2/2] rerere: add not about files with existing conflict markers","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-24T22:10:05Z","receivedAt":"2018-08-24T22:10:21Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"When a file contains lines that look like conflict markers, 'git\nrerere' may fail not be able to record a conflict resolution.\nEmphasize that in the man page.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n\nNot sure if there may be a better place in the man page for this, but\nthis is the best I could come up with.\n\n Documentation/git-rerere.txt | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/Documentation/git-rerere.txt b/Documentation/git-rerere.txt\nindex 031f31fa47..036ea11528 100644\n--- a/Documentation/git-rerere.txt\n+++ b/Documentation/git-rerere.txt\n@@ -211,6 +211,12 @@ would conflict the same way as the test merge you resolved earlier.\n 'git rerere' will be run by 'git rebase' to help you resolve this\n conflict.\n \n+[NOTE]\n+'git rerere' relies on the conflict markers in the file to detect the\n+conflict.  If the file already contains lines that look the same as\n+lines with conflict markers, 'git rerere' may fail to record a\n+conflict resolution.\n+\n GIT\n ---\n Part of the linkgit:git[1] suite\n-- \n2.18.0.1088.ge017bf2cd1\n\n"},{"id":"356594","messageId":"xmqqk1oblnor.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180824215619.GH13316@hank.intra.tgummerer.com","subject":"Re: [PATCH v4 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-27T17:33:08Z","receivedAt":"2018-08-27T17:33:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Agreed.  I think it may be solvable if we'd actually get the\n> information about what belongs to which side from the merge algorithm\n> directly.\n\nThe merge machinery may (eh, rather, \"does\") know, but we do not\nhave a way to express that in the working tree file that becomes the\ninput to the rerere algorithm, without making backward-incompatible\nchanges to the output format.\n\nIn a sense, that is already a solved problem, even though the\nsolution was done a bit differently ;-) If the end users need to\ncommit a half-resolved result with conflict markers (perhaps they\nwant to share it among themselves and work on resolving further),\nwhat they can do is to also say that these are now part of contents,\nnot conflict markers, with conflict-marker-size attribute.  Perhaps\nthey prepare such a half-resolved result with unusual value of the\nattribute, so that later merge of these with standard conflict\nmarker size will not get confused.\n\nThat reminds me of another thing.  I've been running with these in\nmy $GIT_DIR/info/attributes file for the past few years.  Perhaps we\nshould add them to Documentation/.gitattributes and t/.gitattributes\nso that project participants would all benefit?\n\nDocumentation/git-merge.txt\tconflict-marker-size=32\nDocumentation/user-manual.txt\tconflict-marker-size=32\nt/t????-*.sh\t\t\tconflict-marker-size=32\n"},{"id":"356604","messageId":"xmqq4lffk3ez.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180824215619.GH13316@hank.intra.tgummerer.com","subject":"Re: [PATCH v4 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-27T19:36:20Z","receivedAt":"2018-08-27T19:36:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Fair enough.  I thought of the technical documentation as something\n> that doesn't promise users anything, but rather describes how the\n> internals work right now, which is what this bit of documentation\n> attempted to write down.\n\nThat's fine.  I'd rather keep it but perhaps add a reminder to tell\nreaders that it works only when the merging of contents that already\nrecords with nested conflict markers happen to \"cleanly nest\".\n\nThanks.\n"},{"id":"356764","messageId":"20180828212744.18714-1-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180824221005.5983-1-t.gummerer@gmail.com","subject":"[PATCH v2 1/2] rerere: mention caveat about unmatched conflict markers","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-28T21:27:43Z","receivedAt":"2018-08-28T21:27:51Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"4af3220 (\"rerere: teach rerere to handle nested conflicts\",\n2018-08-05) introduced slightly better behaviour if the user commits\nconflict markers and then gets another conflict in 'git rerere'.\n\nHowever this is just a heuristic to punt on such conflicts better, and\ndoesn't deal with any unmatched conflict markers.  Make that clearer\nin the documentation.\n\nSuggested-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n\n> That's fine.  I'd rather keep it but perhaps add a reminder to tell\n> readers that it works only when the merging of contents that already\n> records with nested conflict markers happen to \"cleanly nest\".\n\nYeah that makes sense.  Maybe something like this?\n\n(replying to <xmqq4lffk3ez.fsf@gitster-ct.c.googlers.com> here to keep\nthe patches in one thread)\n\n Documentation/technical/rerere.txt | 4 ++++\n 1 file changed, 4 insertions(+)\n\ndiff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\nindex e65ba9b0c6..8fefe51b00 100644\n--- a/Documentation/technical/rerere.txt\n+++ b/Documentation/technical/rerere.txt\n@@ -149,7 +149,10 @@ version, and the sorting the conflict hunks, both for the outer and the\n inner conflict.  This is done recursively, so any number of nested\n conflicts can be handled.\n \n+Note that this only works for conflict markers that \"cleanly nest\".  If\n+there are any unmatched conflict markers, rerere will fail to handle\n+the conflict and record a conflict resolution.\n+\n The only difference is in how the conflict ID is calculated.  For the\n inner conflict, the conflict markers themselves are not stripped out\n before calculating the sha1.\n-- \n2.18.0.1088.ge017bf2cd1\n\n"},{"id":"356765","messageId":"20180828212744.18714-2-t.gummerer@gmail.com","threadId":"48531","inReplyTo":"20180828212744.18714-1-t.gummerer@gmail.com","subject":"[PATCH v2 2/2] rerere: add note about files with existing conflict markers","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-28T21:27:44Z","receivedAt":"2018-08-28T21:27:53Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"When a file contains lines that look like conflict markers, 'git\nrerere' may fail not be able to record a conflict resolution.\nEmphasize that in the man page, and mention a possible workaround for\nthe issue.\n\nSuggested-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n\nCompared to v1, this now mentions the workaround of setting the\n'conflict-marker-size', as mentioned in\n<xmqqk1oblnor.fsf@gitster-ct.c.googlers.com>\n\n Documentation/git-rerere.txt | 6 ++++++\n 1 file changed, 6 insertions(+)\n\ndiff --git a/Documentation/git-rerere.txt b/Documentation/git-rerere.txt\nindex 031f31fa47..df310d2a58 100644\n--- a/Documentation/git-rerere.txt\n+++ b/Documentation/git-rerere.txt\n@@ -211,6 +211,12 @@ would conflict the same way as the test merge you resolved earlier.\n 'git rerere' will be run by 'git rebase' to help you resolve this\n conflict.\n \n+[NOTE] 'git rerere' relies on the conflict markers in the file to\n+detect the conflict.  If the file already contains lines that look the\n+same as lines with conflict markers, 'git rerere' may fail to record a\n+conflict resolution.  To work around this, the `conflict-marker-size`\n+setting in linkgit:gitattributes[5] can be used.\n+\n GIT\n ---\n Part of the linkgit:git[1] suite\n-- \n2.18.0.1088.ge017bf2cd1\n\n"},{"id":"356776","messageId":"20180828220550.GJ13316@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqqk1oblnor.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v4 10/11] rerere: teach rerere to handle nested conflicts","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-08-28T22:05:50Z","receivedAt":"2018-08-28T22:05:56Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 08/27, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > Agreed.  I think it may be solvable if we'd actually get the\n> > information about what belongs to which side from the merge algorithm\n> > directly.\n> \n> The merge machinery may (eh, rather, \"does\") know, but we do not\n> have a way to express that in the working tree file that becomes the\n> input to the rerere algorithm, without making backward-incompatible\n> changes to the output format.\n\nRight, I was more thinking along the lines of using the stages in the\nindex to redo the merge and get the information that way.  But that\nmay not work as well with using 'git rerere' from the command line,\nand have other backwards compatibility woes, that I didn't quite think\nthrough yet :)\n\n> In a sense, that is already a solved problem, even though the\n> solution was done a bit differently ;-) If the end users need to\n> commit a half-resolved result with conflict markers (perhaps they\n> want to share it among themselves and work on resolving further),\n> what they can do is to also say that these are now part of contents,\n> not conflict markers, with conflict-marker-size attribute.  Perhaps\n> they prepare such a half-resolved result with unusual value of the\n> attribute, so that later merge of these with standard conflict\n> marker size will not get confused.\n\nRight, I wasn't aware of the conflict-marker-size attribute.  Thanks\nfor mentioning it!\n\n> That reminds me of another thing.  I've been running with these in\n> my $GIT_DIR/info/attributes file for the past few years.  Perhaps we\n> should add them to Documentation/.gitattributes and t/.gitattributes\n> so that project participants would all benefit?\n> \n> Documentation/git-merge.txt\tconflict-marker-size=32\n> Documentation/user-manual.txt\tconflict-marker-size=32\n> t/t????-*.sh\t\t\tconflict-marker-size=32\n\nI do think that would be a good idea.  I am wondering what the right\nvalue is though.  Seeing such a long conflict marker before I knew\nabout this setting would have struck me as odd, and probably made me\ntry and track down where it is coming from.  But on the other hand it\nmakes the conflict markers very easy to tell apart from the rest of\nthe lines that kind of look like conflict markers.\n\nI think these tradeoffs probably make it worth setting them to a value\nthis large.\n\nOne other file that I see needs such a treatment is\nDocumentation/gitk.txt, where the first header is 7 \"=\"s, and\ntherefore could confuse 'git rerere' as well.  Arguably that's less\nimportant, as there's unlikely to be a conflict containing that line,\nbut it may be worth including for completeness sake.\n\nMaybe something like this?  Though it may be good for others to chime\nin if they find this helpful or whether they find the long conflict\nmarkers distracting.\n\n--- >8 ---\nSubject: [PATCH] .gitattributes: add conflict-marker-size for relevant files\n\nSome files in git.git contain lines that look like conflict markers,\neither in examples or tests, or in the case of Documentation/gitk.txt\nbecause of the asciidoc heading.\n\nHaving conflict markers the same length as the actual content can be\nconfusing for humans, and is impossible to handle for tools like 'git\nrerere'.  Work around that by setting the 'conflict-marker-size'\nattribute for those files to 32, which makes the conflict markers\nunambiguous.\n\nHelped-by: Junio C Hamano <gitster@pobox.com>\nSigned-off-by: Thomas Gummerer <t.gummerer@gmail.com>\n---\n .gitattributes | 4 ++++\n 1 file changed, 4 insertions(+)\n\ndiff --git a/.gitattributes b/.gitattributes\nindex 1bdc91e282..49b3051641 100644\n--- a/.gitattributes\n+++ b/.gitattributes\n@@ -9,3 +9,7 @@\n /command-list.txt eol=lf\n /GIT-VERSION-GEN eol=lf\n /mergetools/* eol=lf\n+/Documentation/git-merge.txt conflict-marker-size=32\n+/Documentation/gitk.txt conflict-marker-size=32\n+/Documentation/user-manual.txt conflict-marker-size=32\n+/t/t????-*.sh conflict-marker-size=32\n-- \n2.18.0.1088.ge017bf2cd1\n"},{"id":"356905","messageId":"xmqq5zzsajg1.fsf@gitster-ct.c.googlers.com","threadId":"48531","inReplyTo":"20180828212744.18714-1-t.gummerer@gmail.com","subject":"Re: [PATCH v2 1/2] rerere: mention caveat about unmatched conflict markers","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2018-08-29T16:04:35Z","receivedAt":"2018-08-29T22:37:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Thomas Gummerer <t.gummerer@gmail.com> writes:\n\n> Yeah that makes sense.  Maybe something like this?\n>\n> (replying to <xmqq4lffk3ez.fsf@gitster-ct.c.googlers.com> here to keep\n> the patches in one thread)\n>\n>  Documentation/technical/rerere.txt | 4 ++++\n>  1 file changed, 4 insertions(+)\n>\n> diff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\n> index e65ba9b0c6..8fefe51b00 100644\n> --- a/Documentation/technical/rerere.txt\n> +++ b/Documentation/technical/rerere.txt\n> @@ -149,7 +149,10 @@ version, and the sorting the conflict hunks, both for the outer and the\n>  inner conflict.  This is done recursively, so any number of nested\n>  conflicts can be handled.\n>  \n> +Note that this only works for conflict markers that \"cleanly nest\".  If\n> +there are any unmatched conflict markers, rerere will fail to handle\n> +the conflict and record a conflict resolution.\n> +\n>  The only difference is in how the conflict ID is calculated.  For the\n>  inner conflict, the conflict markers themselves are not stripped out\n>  before calculating the sha1.\n\nLooks good to me except for the line count on the @@ line.  The\npreimage ought to have 6 (not 7) lines and adding 4 new lines makes\nit a 10 line postimage.  I wonder who miscounted the hunk---it is\nimmediately followed by the signature cut mark \"-- \\n\" and some\ntools (including Emacs's patch editing mode) are known to\nmisinterpret it as a preimage line that was removed.\n\nWhat is curious is that your 2/2 counts the preimage lines\ncorrectly.\n\nIn any case, both patches look good.  Will apply.\n\nThanks.\n"},{"id":"357156","messageId":"20180901090039.GA6147@hank.intra.tgummerer.com","threadId":"48531","inReplyTo":"xmqq5zzsajg1.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v2 1/2] rerere: mention caveat about unmatched conflict markers","fromName":"Thomas Gummerer","fromEmail":"t.gummerer@gmail.com","sentAt":"2018-09-01T09:00:39Z","receivedAt":"2018-09-01T09:05:34Z","isPatch":true,"sender":{"key":"t.gummerer@gmail.com","avatar":"https://avatars.githubusercontent.com/u/191004?v=4"},"body":"On 08/29, Junio C Hamano wrote:\n> Thomas Gummerer <t.gummerer@gmail.com> writes:\n> \n> > Yeah that makes sense.  Maybe something like this?\n> >\n> > (replying to <xmqq4lffk3ez.fsf@gitster-ct.c.googlers.com> here to keep\n> > the patches in one thread)\n> >\n> >  Documentation/technical/rerere.txt | 4 ++++\n> >  1 file changed, 4 insertions(+)\n> >\n> > diff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt\n> > index e65ba9b0c6..8fefe51b00 100644\n> > --- a/Documentation/technical/rerere.txt\n> > +++ b/Documentation/technical/rerere.txt\n> > @@ -149,7 +149,10 @@ version, and the sorting the conflict hunks, both for the outer and the\n> >  inner conflict.  This is done recursively, so any number of nested\n> >  conflicts can be handled.\n> >  \n> > +Note that this only works for conflict markers that \"cleanly nest\".  If\n> > +there are any unmatched conflict markers, rerere will fail to handle\n> > +the conflict and record a conflict resolution.\n> > +\n> >  The only difference is in how the conflict ID is calculated.  For the\n> >  inner conflict, the conflict markers themselves are not stripped out\n> >  before calculating the sha1.\n> \n> Looks good to me except for the line count on the @@ line.  The\n> preimage ought to have 6 (not 7) lines and adding 4 new lines makes\n> it a 10 line postimage.  I wonder who miscounted the hunk---it is\n> immediately followed by the signature cut mark \"-- \\n\" and some\n> tools (including Emacs's patch editing mode) are known to\n> misinterpret it as a preimage line that was removed.\n\nSorry about that.  Yeah Emacs's patch editing mode doing that would\nexplain it.  I did a round of proof-reading in my editor, and spotted\na typo.  Since it was trivial to fix I just edited the patch\ndirectly, and Emacs changed the line count.  Sorry about that, I'll be\nmore careful about this in the future.\n\n> What is curious is that your 2/2 counts the preimage lines\n> correctly.\n\nI only added some text after the '---' line in 2/2, but did not edit\nthe patch directly.  Emacs's patch editing mode only seems to change\nthe line numbers of the patch that's being edited, not if anything\nsurrounding that is changed, so the line count stayed the same as what\nformat-patch put in the file in the first place.\n\n> In any case, both patches look good.  Will apply.\n\nThanks!\n\n> Thanks.\n"}]}