git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH v3 05/11] rerere: add documentation for conflict normalization

From
Thomas Gummerer <t.gummerer@gmail.com>
Date
Jul 14, 2018, 21:44 UTC
Message-ID
<20180714214443.7184-6-t.gummerer@gmail.com>
In-Reply-To
<20180714214443.7184-1-t.gummerer@gmail.com>

Add some documentation for the logic behind the conflict normalization in rerere.

Helped-by: Junio C Hamano <gitster@pobox.com>
Signed-off-by: Thomas Gummerer <t.gummerer@gmail.com>
---
 Documentation/technical/rerere.txt | 140 +++++++++++++++++++++++++++++
 rerere.c                           |   4 -
 2 files changed, 140 insertions(+), 4 deletions(-)
 create mode 100644 Documentation/technical/rerere.txt
diff --git a/Documentation/technical/rerere.txt b/Documentation/technical/rerere.txt
new file mode 100644
index 0000000000..4102cce7aa
--- /dev/null
+++ b/Documentation/technical/rerere.txt
@@ -0,0 +1,140 @@
+Rerere
+======
+
+This document describes the rerere logic.
+
+Conflict normalization
+----------------------
+
+To ensure recorded conflict resolutions can be looked up in the rerere
+database, even when branches are merged in a different order,
+different branches are merged that result in the same conflict, or
+when different conflict style settings are used, rerere normalizes the
+conflicts before writing them to the rerere database.
+
+Different conflict styles and branch names are normalized by stripping
+the labels from the conflict markers, and removing extraneous
+information from the `diff3` conflict style. Branches that are merged
+in different order are normalized by sorting the conflict hunks.  More
+on each of those steps in the following sections.
+
+Once these two normalization operations are applied, a conflict ID is
+calculated based on the normalized conflict, which is later used by
+rerere to look up the conflict in the rerere database.
+
+Stripping extraneous information
+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
+
+Say we have three branches AB, AC and AC2.  The common ancestor of
+these branches has a file with a line containing the string "A" (for
+brevity this is called "line A" in the rest of the document).  In
+branch AB this line is changed to "B", in AC, this line is changed to
+"C", and branch AC2 is forked off of AC, after the line was changed to
+"C".
+
+Forking a branch ABAC off of branch AB and then merging AC into it, we
+get a conflict like the following:
+
+    <<<<<<< HEAD
+    B
+    =======
+    C
+    >>>>>>> AC
+
+Doing the analogous with AC2 (forking a branch ABAC2 off of branch AB
+and then merging branch AC2 into it), using the diff3 conflict style,
+we get a conflict like the following:
+
+    <<<<<<< HEAD
+    B
+    ||||||| merged common ancestors
+    A
+    =======
+    C
+    >>>>>>> AC2
+
+By resolving this conflict, to leave line D, the user declares:
+
+    After examining what branches AB and AC did, I believe that making
+    line A into line D is the best thing to do that is compatible with
+    what AB and AC wanted to do.
+
+As branch AC2 refers to the same commit as AC, the above implies that
+this is also compatible what AB and AC2 wanted to do.
+
+By extension, this means that rerere should recognize that the above
+conflicts are the same.  To do this, the labels on the conflict
+markers are stripped, and the diff3 output is removed.  The above
+examples would both result in the following normalized conflict:
+
+    <<<<<<<
+    B
+    =======
+    C
+    >>>>>>>
+
+Sorting hunks
+~~~~~~~~~~~~~
+
+As before, lets imagine that a common ancestor had a file with line A
+its early part, and line X in its late part.  And then four branches
+are forked that do these things:
+
+    - AB: changes A to B
+    - AC: changes A to C
+    - XY: changes X to Y
+    - XZ: changes X to Z
+
+Now, forking a branch ABAC off of branch AB and then merging AC into
+it, and forking a branch ACAB off of branch AC and then merging AB
+into it, would yield the conflict in a different order.  The former
+would say "A became B or C, what now?" while the latter would say "A
+became C or B, what now?"
+
+As a reminder, the act of merging AC into ABAC and resolving the
+conflict to leave line D means that the user declares:
+
+    After examining what branches AB and AC did, I believe that
+    making line A into line D is the best thing to do that is
+    compatible with what AB and AC wanted to do.
+
+So the conflict we would see when merging AB into ACAB should be
+resolved the same way---it is the resolution that is in line with that
+declaration.
+
+Imagine that similarly previously a branch XYXZ was forked from XY,
+and XZ was merged into it, and resolved "X became Y or Z" into "X
+became W".
+
+Now, if a branch ABXY was forked from AB and then merged XY, then ABXY
+would have line B in its early part and line Y in its later part.
+Such a merge would be quite clean.  We can construct 4 combinations
+using these four branches ((AB, AC) x (XY, XZ)).
+
+Merging ABXY and ACXZ would make "an early A became B or C, a late X
+became Y or Z" conflict, while merging ACXY and ABXZ would make "an
+early A became C or B, a late X became Y or Z".  We can see there are
+4 combinations of ("B or C", "C or B") x ("X or Y", "Y or X").
+
+By sorting, the conflict is given its canonical name, namely, "an
+early part became B or C, a late part becames X or Y", and whenever
+any of these four patterns appear, and we can get to the same conflict
+and resolution that we saw earlier.
+
+Without the sorting, we'd have to somehow find a previous resolution
+from combinatorial explosion.
+
+Conflict ID calculation
+~~~~~~~~~~~~~~~~~~~~~~~
+
+Once the conflict normalization is done, the conflict ID is calculated
+as the sha1 hash of the conflict hunks appended to each other,
+separated by <NUL> characters.  The conflict markers are stripped out
+before the sha1 is calculated.  So in the example above, where we
+merge branch AC which changes line A to line C, into branch AB, which
+changes line A to line C, the conflict ID would be
+SHA1('B<NUL>C<NUL>').
+
+If there are multiple conflicts in one file, the sha1 is calculated
+the same way with all hunks appended to each other, in the order in
+which they appear in the file, separated by a <NUL> character.
diff --git a/rerere.c b/rerere.c
index be98c0afcb..da1ab54027 100644
--- a/rerere.c
+++ b/rerere.c
@@ -394,10 +394,6 @@ static int is_cmarker(char *buf, int marker_char, int marker_size)
  * and NUL concatenated together.
  *
  * Return the number of conflict hunks found.
- *
- * NEEDSWORK: the logic and theory of operation behind this conflict
- * normalization may deserve to be documented somewhere, perhaps in
- * Documentation/technical/rerere.txt.
  */
 static int handle_path(unsigned char *sha1, struct rerere_io *io, int marker_size)
 {
-- 
2.17.0.410.g65aef3a6c4
Previous: Thomas GummererNext: Junio C Hamano
Message 43 of 84 in “rerere: handle nested conflicts”
  1. 0/7 rerere: handle nested conflictsThomas Gummerer, May 20, 2018
  2. 1/7 rerere: unify error message when read_cache failsThomas Gummerer, May 20, 2018
  3. Stefan BellerMay 21, 2018
  4. 2/7 rerere: mark strings for translationThomas Gummerer, May 20, 2018
  5. Junio C HamanoMay 24, 2018
  6. 3/7 rerere: add some documentationThomas Gummerer, May 20, 2018
  7. Junio C HamanoMay 24, 2018
  8. Thomas GummererJun 3, 2018
  9. 4/7 rerere: fix crash when conflict goes unresolvedThomas Gummerer, May 20, 2018
  10. Junio C HamanoMay 24, 2018
  11. Thomas GummererMay 24, 2018
  12. Junio C HamanoMay 25, 2018
  13. 5/7 rerere: only return whether a path has conflicts or notThomas Gummerer, May 20, 2018
  14. Junio C HamanoMay 24, 2018
  15. 6/7 rerere: factor out handle_conflict functionThomas Gummerer, May 20, 2018
  16. 7/7 rerere: teach rerere to handle nested conflictsThomas Gummerer, May 20, 2018
  17. Junio C HamanoMay 24, 2018
  18. Thomas GummererMay 24, 2018
  19. 00/10 rerere: handle nested conflictsThomas Gummerer, Jun 5, 2018
  20. 01/10 rerere: unify error messages when read_cache failsThomas Gummerer, Jun 5, 2018
  21. 03/10 rerere: wrap paths in output in sqThomas Gummerer, Jun 5, 2018
  22. 05/10 rerere: add some documentationThomas Gummerer, Jun 5, 2018
  23. 07/10 rerere: only return whether a path has conflicts or notThomas Gummerer, Jun 5, 2018
  24. 09/10 rerere: teach rerere to handle nested conflictsThomas Gummerer, Jun 5, 2018
  25. 10/10 rerere: recalculate conflict ID when unresolved conflict is committedThomas Gummerer, Jun 5, 2018
  26. 02/10 rerere: lowercase error messagesThomas Gummerer, Jun 5, 2018
  27. 06/10 rerere: fix crash when conflict goes unresolvedThomas Gummerer, Jun 5, 2018
  28. 04/10 rerere: mark strings for translationThomas Gummerer, Jun 5, 2018
  29. 08/10 rerere: factor out handle_conflict functionThomas Gummerer, Jun 5, 2018
  30. Thomas GummererJul 3, 2018
  31. Junio C HamanoJul 6, 2018
  32. Thomas GummererJul 10, 2018
  33. 00/11 rerere: handle nested conflictsThomas Gummerer, Jul 14, 2018
  34. 01/11 rerere: unify error messages when read_cache failsThomas Gummerer, Jul 14, 2018
  35. 03/11 rerere: wrap paths in output in sqThomas Gummerer, Jul 14, 2018
  36. 02/11 rerere: lowercase error messagesThomas Gummerer, Jul 14, 2018
  37. 04/11 rerere: mark strings for translationThomas Gummerer, Jul 14, 2018
  38. Simon RuderichJul 15, 2018
  39. Thomas GummererJul 16, 2018
  40. 06/11 rerere: fix crash when conflict goes unresolvedThomas Gummerer, Jul 14, 2018
  41. Junio C HamanoJul 30, 2018
  42. Thomas GummererJul 30, 2018
  43. 05/11 rerere: add documentation for conflict normalizationThomas Gummerer, Jul 14, 2018
  44. Junio C HamanoJul 30, 2018
  45. Thomas GummererJul 30, 2018
  46. 07/11 rerere: only return whether a path has conflicts or notThomas Gummerer, Jul 14, 2018
  47. Junio C HamanoJul 30, 2018
  48. Thomas GummererJul 30, 2018
  49. 08/11 rerere: factor out handle_conflict functionThomas Gummerer, Jul 14, 2018
  50. Junio C HamanoJul 30, 2018
  51. 09/11 rerere: return strbuf from handle pathThomas Gummerer, Jul 14, 2018
  52. Junio C HamanoJul 30, 2018
  53. 10/11 rerere: teach rerere to handle nested conflictsThomas Gummerer, Jul 14, 2018
  54. Junio C HamanoJul 30, 2018
  55. Thomas GummererJul 30, 2018
  56. 11/11 rerere: recalculate conflict ID when unresolved conflict is committedThomas Gummerer, Jul 14, 2018
  57. Junio C HamanoJul 30, 2018
  58. Thomas GummererJul 30, 2018
  59. 00/11 rerere: handle nested conflictsThomas Gummerer, Aug 5, 2018
  60. 01/11 rerere: unify error messages when read_cache failsThomas Gummerer, Aug 5, 2018
  61. 02/11 rerere: lowercase error messagesThomas Gummerer, Aug 5, 2018
  62. 03/11 rerere: wrap paths in output in sqThomas Gummerer, Aug 5, 2018
  63. 04/11 rerere: mark strings for translationThomas Gummerer, Aug 5, 2018
  64. 05/11 rerere: add documentation for conflict normalizationThomas Gummerer, Aug 5, 2018
  65. 06/11 rerere: fix crash with files rerere can't handleThomas Gummerer, Aug 5, 2018
  66. 07/11 rerere: only return whether a path has conflicts or notThomas Gummerer, Aug 5, 2018
  67. 08/11 rerere: factor out handle_conflict functionThomas Gummerer, Aug 5, 2018
  68. 09/11 rerere: return strbuf from handle pathThomas Gummerer, Aug 5, 2018
  69. 10/11 rerere: teach rerere to handle nested conflictsThomas Gummerer, Aug 5, 2018
  70. Ævar Arnfjörð BjarmasonAug 22, 2018
  71. Junio C HamanoAug 22, 2018
  72. Thomas GummererAug 22, 2018
  73. Junio C HamanoAug 22, 2018
  74. Thomas GummererAug 24, 2018
  75. 1/2 rerere: remove documentation for "nested conflicts"Thomas Gummerer, Aug 24, 2018
  76. 2/2 rerere: add not about files with existing conflict markersThomas Gummerer, Aug 24, 2018
  77. 1/2 rerere: mention caveat about unmatched conflict markersThomas Gummerer, Aug 28, 2018
  78. 2/2 rerere: add note about files with existing conflict markersThomas Gummerer, Aug 28, 2018
  79. Junio C HamanoAug 29, 2018
  80. Thomas GummererSep 1, 2018
  81. Junio C HamanoAug 27, 2018
  82. Thomas GummererAug 28, 2018
  83. Junio C HamanoAug 27, 2018
  84. 11/11 rerere: recalculate conflict ID when unresolved conflict is committedThomas Gummerer, Aug 5, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.