{"thread":{"id":"65698","subject":"[RFC PATCH 0/3] diff: pair edited lines inside moved blocks","startedAt":"2026-05-27T04:24:14Z","lastAt":"2026-05-27T04:24:17Z","messageCount":4,"participants":["Keita Oda"],"isPatch":true,"patchVersion":1,"patchTotal":3},"messages":[{"id":"544136","messageId":"20260527042402.13607-1-ainsophyao@gmail.com","threadId":"65698","inReplyTo":null,"subject":"[RFC PATCH 0/3] diff: pair edited lines inside moved blocks","fromName":"Keita Oda","fromEmail":"ainsophyao@gmail.com","sentAt":"2026-05-27T04:23:59Z","receivedAt":"2026-05-27T04:24:14Z","isPatch":true,"body":"From: Keita ODA <ainsophyao@gmail.com>\n\nThis is an RFC for a review aid, not a proposed final UI or option name.\n\nThe motivation is the gap between --word-diff and --color-moved.\n--word-diff is very useful when the line-level diff already found useful\nold/new line pairs.  --color-moved is useful when moved lines are exact\nmatches.  But when a block is moved and one line inside the block is edited,\nthe small edit can be buried in a large delete/add region.\n\nThat case matters for review.  A one-line move is usually easy to inspect by\neye.  A ten-line moved block with a one-character change inside it is harder\nto audit.  A small synthetic permission-table example in patch 3 uses this\nshape:\n\n  -#define PERM_RESOURCE_EXPORT       0x0008\n  +#define PERM_RESOURCE_EXPORT       0x0001\n\nThat particular toy example is not meant to show something that\n--color-moved cannot see.  It is meant to make the review question small:\ncan Git expose \"this moved line was also edited\" in a lightweight way?\nThe real-world cases below are less about proving that existing modes are\nblind, and more about making the row-to-row correspondence explicit enough\nthat the small edits are easy to check.\n\nThis series adds an opt-in prototype, --word-diff-align, that post-processes\nthe emitted diff symbols and tries to pair similar deleted and inserted lines.\nIt does not change the underlying diff algorithm, patch semantics, apply, or\nmerge behavior.\nThe prototype is deliberately language-agnostic.  It does not parse source\ncode or build an AST; it only tokenizes diff lines into small text tokens and\nscores local token overlap.  This keeps the experiment applicable to code,\ntests, generated tables, documentation, and other text files.\n\nThe prototype is intentionally split into three pieces:\n\n  * patch 1 adds the candidate retrieval and line-pair scoring, and exposes\n    selected pairs with an RFC/debug comment;\n  * patch 2 adds a small RFC-only renderer that inserts word-diff-like\n    markers on the selected pairs, so that the recovered pairs are easier to\n    inspect;\n  * patch 3 adds a focused test case.\n\nThe current prototype is still larger than I would like, but the split keeps\nthe experimental pieces visible.  The full series is about 1000 inserted\nlines; roughly 800 lines are option plumbing, tokenization, candidate\nretrieval, scoring, pair selection, and debug comments, while about 200 lines\nare temporary rendering code for review.\n\nThe scoring model is:\n\n  S = W + aL\n\nwhere W is a 5-line-window token overlap score and L is a center-line token\nLCS score.  A small 64-bit window fingerprint is used only as a candidate\nretrieval index; candidate pairs are scored again before they are selected.\nTokens repeated in the surrounding small window carry less weight for the\ncenter-line score, which is a local-IDF-like approximation.  This keeps tokens\nsuch as \"import\" or \"#define\" from overwhelming the line-specific identifier.\n\nSome real-world examples that motivated the prototype:\n\n  * CPython opcode/metadata renumbering, where many table rows stay logically\n    paired but their numeric values shift;\n  * CPython test parameterization rewrites such as tuple rows becoming\n    dict(input=..., expected=...) rows;\n  * Git's own expected-output tables, where a column width change adds spaces\n    across many rows and a row insertion shifts the surrounding context;\n  * Git's own remote.c refactoring, where extracted helper code has small\n    identifier changes.\n\nAs a rough trigger-rate sanity check, I ran the prototype over 5734 changed\nfile pairs sampled from recent Git, CPython, and Rust history.  The stricter\n\"crossing and edited\" signal, which ignores the many adjacent row pairs and\nlooks for pairs that cross another recovered pair, appeared in 739 file pairs\n(about 13%).  This is not a gold-label quality number, but it suggests that\nthe mode is not only triggering on the synthetic test.\nA small manual review found both clear wins and loose matches.\n\nI found the problem easiest to inspect with four-way comparisons:\n\n  * git diff --histogram\n  * git diff --histogram --word-diff=plain\n  * git diff --histogram --color-moved=blocks\n  * git diff --histogram --word-diff-align\n\nI put a small set of rendered four-way examples here:\n\n  https://oda.github.io/git-diff-rfc-examples/rfc-word-diff-align/\n\nThese links are supplemental; the patch series is intended to be readable\nwithout them.\n\nKnown limitations:\n\n  * the UI/debug output is not final;\n  * generated or boilerplate-heavy hunks, especially Rust generated test\n    updates, can still produce loose matches;\n  * one-line long-distance pairs are often less useful than block-level pairs;\n  * the prototype intentionally gives local and remote pairs similar treatment\n    for now, to make the recovered pairings visible for discussion;\n  * thresholds and tie-breaking are still experimental.\n\nThe question for this RFC is whether this kind of language-agnostic line-pair\nannotation is worth pursuing in core, and if so whether it should be shaped as\nword-diff plumbing, a color-moved extension, or a separate opt-in mode.\n\nKeita ODA (3):\n  diff: add word-diff-align line pairing\n  diff: render word-diff-align pairs for RFC review\n  t4034: cover moved-and-edited word diff alignment\n\n diff.c                | 996 +++++++++++++++++++++++++++++++++++++++++-\n diff.h                |   1 +\n t/t4034-diff-words.sh |  46 ++\n 3 files changed, 1035 insertions(+), 8 deletions(-)\n\n-- \n2.39.3 (Apple Git-146)\n"},{"id":"544138","messageId":"20260527042402.13607-2-ainsophyao@gmail.com","threadId":"65698","inReplyTo":"20260527042402.13607-1-ainsophyao@gmail.com","subject":"[RFC PATCH 1/3] diff: add word-diff-align line pairing","fromName":"Keita Oda","fromEmail":"ainsophyao@gmail.com","sentAt":"2026-05-27T04:24:00Z","receivedAt":"2026-05-27T04:24:15Z","isPatch":true,"body":"From: Keita ODA <ainsophyao@gmail.com>\n\nAdd an opt-in --word-diff-align mode that post-processes emitted diff\nsymbols and tries to pair similar deleted and inserted lines inside each hunk.\n\nThis is not a replacement for the diff algorithm.  The normal diff is produced\nfirst; the new code only annotates the already-emitted symbols for review.\n\nThe candidate search is intentionally approximate.  Each line is tokenized into\nword-ish runs and single non-space symbols.  Each token gets a hash, each line\ngets a compact 64-bit fingerprint, and each changed line also gets a small\nsurrounding window fingerprint.  Inserted-side windows are placed into bit\nbuckets, and deleted-side windows query those buckets to avoid a full all-pairs\nscan in typical cases.\n\nThe fingerprint is only a retrieval index.  Candidate pairs are scored again\nusing token hashes:\n\n  W = unique token overlap in the small surrounding windows\n  L = center-line token LCS score\n  S = W + 4L\n\nFor the center-line score, tokens repeated in the surrounding window carry less\nweight.  This is a local-IDF-like approximation: in a run of \"import\" or\n\"#define\" lines, the repeated keyword is weak evidence, while the identifier\nspecific to the center line remains strong evidence.\n\nFor this RFC, selected pairs are exposed with a debug-style \"# aligned ...\"\ncomment.  The next patch adds a small renderer to make the selected pairs more\nreadable.\n\n---\n diff.c | 809 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++-\n diff.h |   1 +\n 2 files changed, 802 insertions(+), 8 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex 397e38b41..6b8744920 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -46,6 +46,7 @@\n #include \"setup.h\"\n #include \"strmap.h\"\n #include \"ws.h\"\n+#include \"ewah/ewok.h\"\n \n #ifdef NO_FAST_WORKING_DIRECTORY\n #define FAST_WORKING_DIRECTORY 0\n@@ -1358,6 +1359,774 @@ static void dim_moved_lines(struct diff_options *o)\n \t}\n }\n \n+struct word_diff_align_line_info {\n+\tuint64_t bits;\n+\tint token_start;\n+\tint token_nr;\n+};\n+\n+struct word_diff_align_item {\n+\tint symbol;\n+\tint change_pos;\n+\tint line_no;\n+};\n+\n+struct word_diff_align_candidate {\n+\tint old_pos;\n+\tint new_pos;\n+\tint score;\n+\tint window_score;\n+\tint window_shared;\n+\tint line_score;\n+\tint line_shared;\n+};\n+\n+#define WORD_DIFF_ALIGN_NUMBER_HASH 0x9e3779b9U\n+\n+#define WORD_DIFF_ALIGN_WINDOW_SIZE 2\n+#define WORD_DIFF_ALIGN_FINGERPRINT_BITS 64\n+#define WORD_DIFF_ALIGN_BIT_BUCKET_MAX 256\n+#define WORD_DIFF_ALIGN_WINDOW_MIN_SCORE 35\n+#define WORD_DIFF_ALIGN_WINDOW_MIN_SHARED 4\n+#define WORD_DIFF_ALIGN_LINE_WEIGHT 4\n+#define WORD_DIFF_ALIGN_SCORE_MIN 175\n+\n+static void word_diff_align_window_bounds(int nr, int pos,\n+\t\t\t\t\t  int *start_p, int *end_p)\n+{\n+\tint radius = WORD_DIFF_ALIGN_WINDOW_SIZE;\n+\n+\t*start_p = pos > radius ? pos - radius : 0;\n+\t*end_p = pos + radius + 1 < nr ? pos + radius + 1 : nr;\n+}\n+\n+static int word_diff_align_payload_len(const struct emitted_diff_symbol *e)\n+{\n+\tint len = e->len;\n+\n+\tif (len && e->line[len - 1] == '\\n')\n+\t\tlen--;\n+\tif (len && e->line[len - 1] == '\\r')\n+\t\tlen--;\n+\treturn len;\n+}\n+\n+static int word_diff_align_word_byte(char ch)\n+{\n+\treturn isalnum((unsigned char)ch) || ch == '_';\n+}\n+\n+static int word_diff_align_numeric_token(const char *line, int start, int end)\n+{\n+\tint i;\n+\n+\tfor (i = start; i < end; i++)\n+\t\tif (!isdigit((unsigned char)line[i]))\n+\t\t\treturn 0;\n+\treturn start < end;\n+}\n+\n+static uint64_t word_diff_align_token_bits(unsigned int hash)\n+{\n+\treturn (1ULL << (hash & 63)) | (1ULL << ((hash >> 6) & 63));\n+}\n+\n+static int word_diff_align_next_token(const char *line, int len, int *pos,\n+\t\t\t\t      int *start_p, int *end_p)\n+{\n+\twhile (*pos < len && isspace((unsigned char)line[*pos]))\n+\t\t(*pos)++;\n+\tif (*pos == len)\n+\t\treturn 0;\n+\n+\t*start_p = *pos;\n+\tif (word_diff_align_word_byte(line[*pos]))\n+\t\twhile (*pos < len && word_diff_align_word_byte(line[*pos]))\n+\t\t\t(*pos)++;\n+\telse\n+\t\t(*pos)++;\n+\t*end_p = *pos;\n+\treturn 1;\n+}\n+\n+static void word_diff_align_prepare_line(const struct emitted_diff_symbol *e,\n+\t\t\t\t\t struct word_diff_align_line_info *line,\n+\t\t\t\t\t unsigned int **tokens,\n+\t\t\t\t\t int *tokens_nr,\n+\t\t\t\t\t int *tokens_alloc)\n+{\n+\tint len = word_diff_align_payload_len(e);\n+\tint pos = 0;\n+\tint start, end;\n+\n+\tline->bits = 0;\n+\tline->token_start = *tokens_nr;\n+\twhile (word_diff_align_next_token(e->line, len, &pos, &start, &end)) {\n+\t\tint numeric = word_diff_align_numeric_token(e->line, start, end);\n+\t\tunsigned int hash = numeric ? WORD_DIFF_ALIGN_NUMBER_HASH :\n+\t\t\tmemhash(e->line + start, end - start);\n+\n+\t\tALLOC_GROW(*tokens, *tokens_nr + 1, *tokens_alloc);\n+\t\t(*tokens)[*tokens_nr] = hash;\n+\t\tline->bits |= word_diff_align_token_bits(hash);\n+\t\t(*tokens_nr)++;\n+\t}\n+\tline->token_nr = *tokens_nr - line->token_start;\n+}\n+\n+static uint64_t word_diff_align_window_bits(struct word_diff_align_line_info *lines,\n+\t\t\t\t\t    int nr, int pos)\n+{\n+\tuint64_t bits = 0;\n+\tint start, end;\n+\tint i;\n+\n+\tword_diff_align_window_bounds(nr, pos, &start, &end);\n+\tfor (i = start; i < end; i++)\n+\t\tbits |= lines[i].bits;\n+\treturn bits;\n+}\n+\n+static int word_diff_align_filter_score(uint64_t a, uint64_t b, int *shared_p)\n+{\n+\tint shared = ewah_bit_popcount64(a & b);\n+\tint total = ewah_bit_popcount64(a | b);\n+\n+\t*shared_p = shared;\n+\treturn total ? shared * 100 / total : 0;\n+}\n+\n+static int word_diff_align_line_has_token(const struct word_diff_align_line_info *line,\n+\t\t\t\t\t  const unsigned int *tokens,\n+\t\t\t\t\t  unsigned int hash);\n+\n+static int word_diff_align_window_has_token(const struct word_diff_align_line_info *lines,\n+\t\t\t\t\t    int start, int end,\n+\t\t\t\t\t    const unsigned int *tokens,\n+\t\t\t\t\t    unsigned int hash)\n+{\n+\tint i;\n+\n+\tfor (i = start; i < end; i++)\n+\t\tif (word_diff_align_line_has_token(&lines[i], tokens, hash))\n+\t\t\treturn 1;\n+\treturn 0;\n+}\n+\n+static int word_diff_align_window_seen_token(const struct word_diff_align_line_info *lines,\n+\t\t\t\t\t     int start, int line_pos,\n+\t\t\t\t\t     int token_pos,\n+\t\t\t\t\t     const unsigned int *tokens,\n+\t\t\t\t\t     unsigned int hash)\n+{\n+\tint i;\n+\n+\tfor (i = start; i <= line_pos; i++) {\n+\t\tconst struct word_diff_align_line_info *line = &lines[i];\n+\t\tint end = i == line_pos ? token_pos : line->token_nr;\n+\t\tint j;\n+\n+\t\tfor (j = 0; j < end; j++)\n+\t\t\tif (tokens[line->token_start + j] == hash)\n+\t\t\t\treturn 1;\n+\t}\n+\treturn 0;\n+}\n+\n+static int word_diff_align_window_token_score(const struct word_diff_align_line_info *minus_lines,\n+\t\t\t\t\t      int minus_nr, int minus_pos,\n+\t\t\t\t\t      const unsigned int *minus_tokens,\n+\t\t\t\t\t      const struct word_diff_align_line_info *plus_lines,\n+\t\t\t\t\t      int plus_nr, int plus_pos,\n+\t\t\t\t\t      const unsigned int *plus_tokens,\n+\t\t\t\t\t      int *shared_p)\n+{\n+\tint minus_start, minus_end, plus_start, plus_end;\n+\tint minus_unique = 0, plus_unique = 0, shared = 0;\n+\tint i;\n+\n+\tword_diff_align_window_bounds(minus_nr, minus_pos, &minus_start,\n+\t\t\t\t      &minus_end);\n+\tword_diff_align_window_bounds(plus_nr, plus_pos, &plus_start,\n+\t\t\t\t      &plus_end);\n+\n+\tfor (i = minus_start; i < minus_end; i++) {\n+\t\tconst struct word_diff_align_line_info *line = &minus_lines[i];\n+\t\tint j;\n+\n+\t\tfor (j = 0; j < line->token_nr; j++) {\n+\t\t\tunsigned int hash = minus_tokens[line->token_start + j];\n+\n+\t\t\tif (word_diff_align_window_seen_token(minus_lines,\n+\t\t\t\t\t\t\t      minus_start, i, j,\n+\t\t\t\t\t\t\t      minus_tokens,\n+\t\t\t\t\t\t\t      hash))\n+\t\t\t\tcontinue;\n+\t\t\tminus_unique++;\n+\t\t\tif (word_diff_align_window_has_token(plus_lines,\n+\t\t\t\t\t\t\t     plus_start, plus_end,\n+\t\t\t\t\t\t\t     plus_tokens, hash))\n+\t\t\t\tshared++;\n+\t\t}\n+\t}\n+\n+\tfor (i = plus_start; i < plus_end; i++) {\n+\t\tconst struct word_diff_align_line_info *line = &plus_lines[i];\n+\t\tint j;\n+\n+\t\tfor (j = 0; j < line->token_nr; j++) {\n+\t\t\tunsigned int hash = plus_tokens[line->token_start + j];\n+\n+\t\t\tif (word_diff_align_window_seen_token(plus_lines,\n+\t\t\t\t\t\t\t      plus_start, i, j,\n+\t\t\t\t\t\t\t      plus_tokens,\n+\t\t\t\t\t\t\t      hash))\n+\t\t\t\tcontinue;\n+\t\t\tplus_unique++;\n+\t\t}\n+\t}\n+\n+\t*shared_p = shared;\n+\treturn minus_unique + plus_unique - shared ?\n+\t\tshared * 100 / (minus_unique + plus_unique - shared) : 0;\n+}\n+\n+static int word_diff_align_line_has_token(const struct word_diff_align_line_info *line,\n+\t\t\t\t\t  const unsigned int *tokens,\n+\t\t\t\t\t  unsigned int hash)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i < line->token_nr; i++)\n+\t\tif (tokens[line->token_start + i] == hash)\n+\t\t\treturn 1;\n+\treturn 0;\n+}\n+\n+static int word_diff_align_surrounding_line_freq(const struct word_diff_align_line_info *lines,\n+\t\t\t\t\t\t int nr, int pos,\n+\t\t\t\t\t\t const unsigned int *tokens,\n+\t\t\t\t\t\t unsigned int hash)\n+{\n+\tint radius = WORD_DIFF_ALIGN_WINDOW_SIZE;\n+\tint start = pos > radius ? pos - radius : 0;\n+\tint end = pos + radius + 1 < nr ? pos + radius + 1 : nr;\n+\tint i, freq = 0;\n+\n+\tfor (i = start; i < end; i++) {\n+\t\tif (i == pos)\n+\t\t\tcontinue;\n+\t\tif (word_diff_align_line_has_token(&lines[i], tokens, hash))\n+\t\t\tfreq++;\n+\t}\n+\treturn freq;\n+}\n+\n+static int word_diff_align_shared_token_value(const struct word_diff_align_line_info *minus_lines,\n+\t\t\t\t\t      int minus_nr, int minus_pos,\n+\t\t\t\t\t      const unsigned int *minus_tokens,\n+\t\t\t\t\t      const struct word_diff_align_line_info *plus_lines,\n+\t\t\t\t\t      int plus_nr, int plus_pos,\n+\t\t\t\t\t      const unsigned int *plus_tokens,\n+\t\t\t\t\t      unsigned int hash)\n+{\n+\tint minus_freq = word_diff_align_surrounding_line_freq(minus_lines,\n+\t\t\t\t\t\t\t       minus_nr,\n+\t\t\t\t\t\t\t       minus_pos,\n+\t\t\t\t\t\t\t       minus_tokens,\n+\t\t\t\t\t\t\t       hash);\n+\tint plus_freq = word_diff_align_surrounding_line_freq(plus_lines,\n+\t\t\t\t\t\t\t      plus_nr,\n+\t\t\t\t\t\t\t      plus_pos,\n+\t\t\t\t\t\t\t      plus_tokens,\n+\t\t\t\t\t\t\t      hash);\n+\tint freq = minus_freq > plus_freq ? minus_freq : plus_freq;\n+\tint value = 2 * WORD_DIFF_ALIGN_WINDOW_SIZE - freq;\n+\n+\treturn value > 0 ? value : 0;\n+}\n+\n+static int word_diff_align_unmatched_weight(const struct word_diff_align_line_info *line,\n+\t\t\t\t\t    const unsigned int *tokens,\n+\t\t\t\t\t    const unsigned int *other_tokens,\n+\t\t\t\t\t    int other_start, int other_nr)\n+{\n+\tint i, weight = 0;\n+\n+\tfor (i = 0; i < line->token_nr; i++) {\n+\t\tunsigned int hash = tokens[line->token_start + i];\n+\t\tint j, matched = 0;\n+\n+\t\tfor (j = 0; j < other_nr; j++) {\n+\t\t\tif (hash == other_tokens[other_start + j]) {\n+\t\t\t\tmatched = 1;\n+\t\t\t\tbreak;\n+\t\t\t}\n+\t\t}\n+\t\tif (!matched)\n+\t\t\tweight++;\n+\t}\n+\treturn weight;\n+}\n+\n+static int word_diff_align_line_token_score(const struct word_diff_align_line_info *minus,\n+\t\t\t\t\t    const struct word_diff_align_line_info *minus_lines,\n+\t\t\t\t\t    int minus_nr, int minus_pos,\n+\t\t\t\t\t    const unsigned int *minus_tokens,\n+\t\t\t\t\t    const struct word_diff_align_line_info *plus,\n+\t\t\t\t\t    const struct word_diff_align_line_info *plus_lines,\n+\t\t\t\t\t    int plus_nr, int plus_pos,\n+\t\t\t\t\t    const unsigned int *plus_tokens,\n+\t\t\t\t\t    int *shared_p)\n+{\n+\tint i, j, shared;\n+\tint *lcs;\n+\tint unmatched, total;\n+\n+\tif (!minus->token_nr || !plus->token_nr) {\n+\t\t*shared_p = 0;\n+\t\treturn 0;\n+\t}\n+\n+\tCALLOC_ARRAY(lcs, plus->token_nr + 1);\n+\tfor (i = 0; i < minus->token_nr; i++) {\n+\t\tunsigned int minus_hash = minus_tokens[minus->token_start + i];\n+\t\tint prev = 0;\n+\n+\t\tfor (j = 0; j < plus->token_nr; j++) {\n+\t\t\tunsigned int plus_hash = plus_tokens[plus->token_start + j];\n+\t\t\tint column = j + 1;\n+\t\t\tint saved;\n+\n+\t\t\tsaved = lcs[column];\n+\t\t\tif (minus_hash == plus_hash) {\n+\t\t\t\tint value = word_diff_align_shared_token_value(minus_lines,\n+\t\t\t\t\t\t\t\t\t       minus_nr,\n+\t\t\t\t\t\t\t\t\t       minus_pos,\n+\t\t\t\t\t\t\t\t\t       minus_tokens,\n+\t\t\t\t\t\t\t\t\t       plus_lines,\n+\t\t\t\t\t\t\t\t\t       plus_nr,\n+\t\t\t\t\t\t\t\t\t       plus_pos,\n+\t\t\t\t\t\t\t\t\t       plus_tokens,\n+\t\t\t\t\t\t\t\t\t       minus_hash);\n+\t\t\t\tint best = prev + value;\n+\n+\t\t\t\tif (lcs[column] > best)\n+\t\t\t\t\tbest = lcs[column];\n+\t\t\t\tif (lcs[column - 1] > best)\n+\t\t\t\t\tbest = lcs[column - 1];\n+\t\t\t\tlcs[column] = best;\n+\t\t\t}\n+\t\t\telse if (lcs[column - 1] > lcs[column])\n+\t\t\t\tlcs[column] = lcs[column - 1];\n+\t\t\tprev = saved;\n+\t\t}\n+\t}\n+\n+\tshared = lcs[plus->token_nr];\n+\tfree(lcs);\n+\tunmatched = word_diff_align_unmatched_weight(minus, minus_tokens,\n+\t\t\t\t\t\t     plus_tokens,\n+\t\t\t\t\t\t     plus->token_start,\n+\t\t\t\t\t\t     plus->token_nr) +\n+\t\tword_diff_align_unmatched_weight(plus, plus_tokens,\n+\t\t\t\t\t\t minus_tokens,\n+\t\t\t\t\t\t minus->token_start,\n+\t\t\t\t\t\t minus->token_nr);\n+\ttotal = 2 * shared + unmatched;\n+\t*shared_p = shared;\n+\treturn total ? shared * 200 / total : 0;\n+}\n+\n+static int word_diff_align_pair_score(int window_score,\n+\t\t\t\t      const struct word_diff_align_line_info *minus,\n+\t\t\t\t      const struct word_diff_align_line_info *minus_lines,\n+\t\t\t\t      int minus_nr, int minus_pos,\n+\t\t\t\t      const unsigned int *minus_tokens,\n+\t\t\t\t      const struct word_diff_align_line_info *plus,\n+\t\t\t\t      const struct word_diff_align_line_info *plus_lines,\n+\t\t\t\t      int plus_nr, int plus_pos,\n+\t\t\t\t      const unsigned int *plus_tokens,\n+\t\t\t\t      int *line_score_p,\n+\t\t\t\t      int *line_shared_p)\n+{\n+\tint line_score, line_shared;\n+\n+\tline_score = word_diff_align_line_token_score(minus, minus_lines,\n+\t\t\t\t\t\t      minus_nr, minus_pos,\n+\t\t\t\t\t\t      minus_tokens,\n+\t\t\t\t\t\t      plus, plus_lines,\n+\t\t\t\t\t\t      plus_nr, plus_pos,\n+\t\t\t\t\t\t      plus_tokens, &line_shared);\n+\t*line_score_p = line_score;\n+\t*line_shared_p = line_shared;\n+\n+\t/*\n+\t * The bit fingerprint is only a retrieval approximation.  The window\n+\t * score passed here is computed from exact token hashes in the 5-line\n+\t * windows, while the line score gives the center line an additional\n+\t * identity bonus.\n+\t */\n+\treturn window_score + WORD_DIFF_ALIGN_LINE_WEIGHT * line_score;\n+}\n+\n+static int word_diff_align_candidate_cmp(const void *va, const void *vb)\n+{\n+\tconst struct word_diff_align_candidate *a = va;\n+\tconst struct word_diff_align_candidate *b = vb;\n+\n+\tif (a->score != b->score)\n+\t\treturn b->score - a->score;\n+\tif (a->line_score != b->line_score)\n+\t\treturn b->line_score - a->line_score;\n+\tif (a->window_score != b->window_score)\n+\t\treturn b->window_score - a->window_score;\n+\tif (a->line_shared != b->line_shared)\n+\t\treturn b->line_shared - a->line_shared;\n+\tif (a->window_shared != b->window_shared)\n+\t\treturn b->window_shared - a->window_shared;\n+\tif (a->old_pos != b->old_pos)\n+\t\treturn a->old_pos - b->old_pos;\n+\treturn a->new_pos - b->new_pos;\n+}\n+\n+static void word_diff_align_add_candidate(struct word_diff_align_candidate **candidates,\n+\t\t\t\t\t  int *candidates_nr,\n+\t\t\t\t\t  int *candidates_alloc,\n+\t\t\t\t\t  int old_pos, int new_pos,\n+\t\t\t\t\t  int score, int window_score,\n+\t\t\t\t\t  int window_shared, int line_score,\n+\t\t\t\t\t  int line_shared)\n+{\n+\tstruct word_diff_align_candidate *candidate;\n+\n+\tALLOC_GROW(*candidates, *candidates_nr + 1, *candidates_alloc);\n+\tcandidate = &(*candidates)[(*candidates_nr)++];\n+\tcandidate->old_pos = old_pos;\n+\tcandidate->new_pos = new_pos;\n+\tcandidate->score = score;\n+\tcandidate->window_score = window_score;\n+\tcandidate->window_shared = window_shared;\n+\tcandidate->line_score = line_score;\n+\tcandidate->line_shared = line_shared;\n+}\n+\n+static void word_diff_align_debug_append_comment(struct emitted_diff_symbol *line,\n+\t\t\t\t\t\t const struct strbuf *suffix)\n+{\n+\tstruct strbuf out = STRBUF_INIT;\n+\tchar *old_line = (char *)line->line;\n+\tsize_t new_len;\n+\n+\tif (line->len && line->line[line->len - 1] == '\\n') {\n+\t\tstrbuf_add(&out, line->line, line->len - 1);\n+\t\tstrbuf_addbuf(&out, suffix);\n+\t\tstrbuf_addch(&out, '\\n');\n+\t} else {\n+\t\tstrbuf_add(&out, line->line, line->len);\n+\t\tstrbuf_addbuf(&out, suffix);\n+\t}\n+\tline->line = strbuf_detach(&out, &new_len);\n+\tline->len = (int)new_len;\n+\tfree(old_line);\n+}\n+\n+static void word_diff_align_debug_mark_pair(struct emitted_diff_symbol *minus_line,\n+\t\t\t\t\t    struct emitted_diff_symbol *plus_line,\n+\t\t\t\t\t    int minus_lineno, int plus_lineno,\n+\t\t\t\t\t    int changed, int moved,\n+\t\t\t\t\t\t    int window_score,\n+\t\t\t\t\t\t    int line_score,\n+\t\t\t\t\t\t    int pair_score)\n+{\n+\tstruct strbuf suffix = STRBUF_INIT;\n+\n+\tif (moved) {\n+\t\tminus_line->flags |= DIFF_SYMBOL_MOVED_LINE;\n+\t\tplus_line->flags |= DIFF_SYMBOL_MOVED_LINE;\n+\t}\n+\n+\tstrbuf_addf(&suffix,\n+\t\t    \"        # aligned from %d to %d, %s, W=%d L=%d S=%d\",\n+\t\t    minus_lineno, plus_lineno,\n+\t\t    changed ? \"edited\" : \"unchanged\",\n+\t\t    window_score, line_score, pair_score);\n+\tword_diff_align_debug_append_comment(minus_line, &suffix);\n+\tword_diff_align_debug_append_comment(plus_line, &suffix);\n+\tstrbuf_release(&suffix);\n+}\n+\n+static void word_diff_align_add_item(struct word_diff_align_item **items,\n+\t\t\t\t     int *items_nr, int *items_alloc,\n+\t\t\t\t     int symbol, int change_pos, int line_no)\n+{\n+\tALLOC_GROW(*items, *items_nr + 1, *items_alloc);\n+\t(*items)[*items_nr].symbol = symbol;\n+\t(*items)[*items_nr].change_pos = change_pos;\n+\t(*items)[*items_nr].line_no = line_no;\n+\t(*items_nr)++;\n+}\n+\n+static void word_diff_align_hunk_start(const struct emitted_diff_symbols *e,\n+\t\t\t\t       int start, int *old_lineno,\n+\t\t\t\t       int *new_lineno)\n+{\n+\tconst char *line;\n+\tchar *endp;\n+\n+\t*old_lineno = 0;\n+\t*new_lineno = 0;\n+\tif (start <= 0 || e->buf[start - 1].s != DIFF_SYMBOL_CONTEXT_FRAGINFO)\n+\t\treturn;\n+\tline = e->buf[start - 1].line;\n+\tif (!starts_with(line, \"@@ -\"))\n+\t\treturn;\n+\tline += 4;\n+\tif (!isdigit((unsigned char)*line))\n+\t\treturn;\n+\t*old_lineno = strtol(line, &endp, 10);\n+\tline = endp;\n+\tif (*line == ',') {\n+\t\tline++;\n+\t\tstrtol(line, &endp, 10);\n+\t\tline = endp;\n+\t}\n+\tif (!starts_with(line, \" +\"))\n+\t\treturn;\n+\tline += 2;\n+\tif (!isdigit((unsigned char)*line))\n+\t\treturn;\n+\t*new_lineno = strtol(line, &endp, 10);\n+}\n+\n+static int word_diff_align_local_pair(int old_pos, int new_pos)\n+{\n+\treturn old_pos == new_pos;\n+}\n+\n+static void word_diff_align_hunk(struct emitted_diff_symbols *e, int start, int end)\n+{\n+\tstruct word_diff_align_item *old_items = NULL, *new_items = NULL;\n+\tint old_items_nr = 0, old_items_alloc = 0;\n+\tint new_items_nr = 0, new_items_alloc = 0;\n+\tint minus_nr = 0, plus_nr = 0;\n+\tstruct word_diff_align_line_info *old_lines = NULL, *new_lines = NULL;\n+\tunsigned int *old_tokens = NULL, *new_tokens = NULL;\n+\tint old_tokens_nr = 0, old_tokens_alloc = 0;\n+\tint new_tokens_nr = 0, new_tokens_alloc = 0;\n+\tuint64_t *new_contexts = NULL;\n+\tint *bucket_heads = NULL, *bucket_next = NULL, *bucket_counts = NULL;\n+\tint *candidate_counts = NULL, *candidate_touched = NULL;\n+\tchar *used_old = NULL, *used_plus = NULL;\n+\tstruct word_diff_align_candidate *candidates = NULL;\n+\tint candidates_nr = 0, candidates_alloc = 0;\n+\tint old_lineno, new_lineno;\n+\tint i, j;\n+\n+\tword_diff_align_hunk_start(e, start, &old_lineno, &new_lineno);\n+\tfor (i = start; i < end; i++) {\n+\t\tif (e->buf[i].s == DIFF_SYMBOL_MINUS) {\n+\t\t\tword_diff_align_add_item(&old_items, &old_items_nr,\n+\t\t\t\t\t\t &old_items_alloc, i, minus_nr++,\n+\t\t\t\t\t\t old_lineno++);\n+\t\t} else if (e->buf[i].s == DIFF_SYMBOL_PLUS) {\n+\t\t\tword_diff_align_add_item(&new_items, &new_items_nr,\n+\t\t\t\t\t\t &new_items_alloc, i, plus_nr++,\n+\t\t\t\t\t\t new_lineno++);\n+\t\t} else if (e->buf[i].s == DIFF_SYMBOL_CONTEXT) {\n+\t\t\tword_diff_align_add_item(&old_items, &old_items_nr,\n+\t\t\t\t\t\t &old_items_alloc, i, -1,\n+\t\t\t\t\t\t old_lineno++);\n+\t\t\tword_diff_align_add_item(&new_items, &new_items_nr,\n+\t\t\t\t\t\t &new_items_alloc, i, -1,\n+\t\t\t\t\t\t new_lineno++);\n+\t\t}\n+\t}\n+\n+\tif (!minus_nr || !plus_nr)\n+\t\tgoto cleanup;\n+\n+\tCALLOC_ARRAY(old_lines, old_items_nr);\n+\tCALLOC_ARRAY(new_lines, new_items_nr);\n+\tCALLOC_ARRAY(new_contexts, new_items_nr);\n+\tCALLOC_ARRAY(bucket_heads, WORD_DIFF_ALIGN_FINGERPRINT_BITS);\n+\tCALLOC_ARRAY(bucket_counts, WORD_DIFF_ALIGN_FINGERPRINT_BITS);\n+\tCALLOC_ARRAY(bucket_next, new_items_nr * WORD_DIFF_ALIGN_FINGERPRINT_BITS);\n+\tCALLOC_ARRAY(candidate_counts, new_items_nr);\n+\tCALLOC_ARRAY(candidate_touched, new_items_nr);\n+\tCALLOC_ARRAY(used_old, minus_nr);\n+\tCALLOC_ARRAY(used_plus, plus_nr);\n+\n+\tfor (i = 0; i < old_items_nr; i++)\n+\t\tword_diff_align_prepare_line(&e->buf[old_items[i].symbol], &old_lines[i],\n+\t\t\t\t\t     &old_tokens, &old_tokens_nr,\n+\t\t\t\t\t     &old_tokens_alloc);\n+\tfor (i = 0; i < new_items_nr; i++)\n+\t\tword_diff_align_prepare_line(&e->buf[new_items[i].symbol], &new_lines[i],\n+\t\t\t\t\t     &new_tokens, &new_tokens_nr,\n+\t\t\t\t\t     &new_tokens_alloc);\n+\tfor (i = 0; i < new_items_nr; i++)\n+\t\tnew_contexts[i] = word_diff_align_window_bits(new_lines, new_items_nr, i);\n+\n+\tfor (i = 0; i < new_items_nr; i++) {\n+\t\tif (new_items[i].change_pos < 0)\n+\t\t\tcontinue;\n+\t\tfor (j = 0; j < WORD_DIFF_ALIGN_FINGERPRINT_BITS; j++) {\n+\t\t\tuint64_t bit = 1ULL << j;\n+\n+\t\t\tif (!(new_contexts[i] & bit))\n+\t\t\t\tcontinue;\n+\t\t\tbucket_counts[j]++;\n+\t\t\tbucket_next[i * WORD_DIFF_ALIGN_FINGERPRINT_BITS + j] = bucket_heads[j];\n+\t\t\tbucket_heads[j] = i + 1;\n+\t\t}\n+\t}\n+\n+\tfor (i = 0; i < old_items_nr; i++) {\n+\t\tint touched_nr = 0;\n+\t\tuint64_t old_context;\n+\n+\t\tif (old_items[i].change_pos < 0)\n+\t\t\tcontinue;\n+\t\told_context = word_diff_align_window_bits(old_lines, old_items_nr, i);\n+\t\tfor (j = 0; j < WORD_DIFF_ALIGN_FINGERPRINT_BITS; j++) {\n+\t\t\tuint64_t bit = 1ULL << j;\n+\t\t\tint item;\n+\n+\t\t\tif (!(old_context & bit))\n+\t\t\t\tcontinue;\n+\t\t\tif (bucket_counts[j] > WORD_DIFF_ALIGN_BIT_BUCKET_MAX)\n+\t\t\t\tcontinue;\n+\t\t\tfor (item = bucket_heads[j]; item; item = bucket_next[(item - 1) * WORD_DIFF_ALIGN_FINGERPRINT_BITS + j]) {\n+\t\t\t\tint new_pos = item - 1;\n+\n+\t\t\t\tif (!candidate_counts[new_pos])\n+\t\t\t\t\tcandidate_touched[touched_nr++] = new_pos;\n+\t\t\t\tcandidate_counts[new_pos]++;\n+\t\t\t}\n+\t\t}\n+\n+\t\tfor (j = 0; j < touched_nr; j++) {\n+\t\t\tint new_pos = candidate_touched[j];\n+\t\t\tint filter_score, filter_shared;\n+\t\t\tint window_score, window_shared;\n+\t\t\tint line_score, line_shared, pair_score;\n+\n+\t\t\tcandidate_counts[new_pos] = 0;\n+\t\t\tif (new_items[new_pos].change_pos < 0)\n+\t\t\t\tcontinue;\n+\t\t\tfilter_score = word_diff_align_filter_score(old_context,\n+\t\t\t\t\t\t\t\t    new_contexts[new_pos],\n+\t\t\t\t\t\t\t\t    &filter_shared);\n+\t\t\tif (filter_score < WORD_DIFF_ALIGN_WINDOW_MIN_SCORE ||\n+\t\t\t    filter_shared < WORD_DIFF_ALIGN_WINDOW_MIN_SHARED)\n+\t\t\t\tcontinue;\n+\t\t\twindow_score = word_diff_align_window_token_score(old_lines,\n+\t\t\t\t\t\t\t\t\t old_items_nr,\n+\t\t\t\t\t\t\t\t\t i,\n+\t\t\t\t\t\t\t\t\t old_tokens,\n+\t\t\t\t\t\t\t\t\t new_lines,\n+\t\t\t\t\t\t\t\t\t new_items_nr,\n+\t\t\t\t\t\t\t\t\t new_pos,\n+\t\t\t\t\t\t\t\t\t new_tokens,\n+\t\t\t\t\t\t\t\t\t &window_shared);\n+\t\t\tif (window_score < WORD_DIFF_ALIGN_WINDOW_MIN_SCORE)\n+\t\t\t\tcontinue;\n+\t\t\tpair_score = word_diff_align_pair_score(window_score,\n+\t\t\t\t\t\t\t\t&old_lines[i],\n+\t\t\t\t\t\t\t\told_lines,\n+\t\t\t\t\t\t\t\told_items_nr,\n+\t\t\t\t\t\t\t\ti,\n+\t\t\t\t\t\t\t\told_tokens,\n+\t\t\t\t\t\t\t\t&new_lines[new_pos],\n+\t\t\t\t\t\t\t\tnew_lines,\n+\t\t\t\t\t\t\t\tnew_items_nr,\n+\t\t\t\t\t\t\t\tnew_pos,\n+\t\t\t\t\t\t\t\tnew_tokens,\n+\t\t\t\t\t\t\t\t&line_score,\n+\t\t\t\t\t\t\t\t&line_shared);\n+\t\t\tif (pair_score < WORD_DIFF_ALIGN_SCORE_MIN)\n+\t\t\t\tcontinue;\n+\t\t\tword_diff_align_add_candidate(&candidates, &candidates_nr,\n+\t\t\t\t\t\t      &candidates_alloc, i, new_pos,\n+\t\t\t\t\t\t      pair_score, window_score,\n+\t\t\t\t\t\t      window_shared, line_score,\n+\t\t\t\t\t\t      line_shared);\n+\t\t}\n+\t}\n+\n+\tQSORT(candidates, candidates_nr, word_diff_align_candidate_cmp);\n+\tfor (i = 0; i < candidates_nr; i++) {\n+\t\tint old_pos = candidates[i].old_pos;\n+\t\tint new_pos = candidates[i].new_pos;\n+\t\tint minus_pos = old_items[old_pos].change_pos;\n+\t\tint plus_pos = new_items[new_pos].change_pos;\n+\t\tint old_len, new_len, edited, moved;\n+\n+\t\tif (used_old[minus_pos] || used_plus[plus_pos])\n+\t\t\tcontinue;\n+\t\tused_old[minus_pos] = 1;\n+\t\tused_plus[plus_pos] = 1;\n+\t\told_len = word_diff_align_payload_len(&e->buf[old_items[old_pos].symbol]);\n+\t\tnew_len = word_diff_align_payload_len(&e->buf[new_items[new_pos].symbol]);\n+\t\tedited = old_len != new_len ||\n+\t\t\tmemcmp(e->buf[old_items[old_pos].symbol].line,\n+\t\t\t       e->buf[new_items[new_pos].symbol].line,\n+\t\t\t       old_len);\n+\t\tmoved = !word_diff_align_local_pair(old_pos, new_pos);\n+\t\tword_diff_align_debug_mark_pair(&e->buf[old_items[old_pos].symbol],\n+\t\t\t\t\t\t&e->buf[new_items[new_pos].symbol],\n+\t\t\t\t\t\told_items[old_pos].line_no,\n+\t\t\t\t\t\tnew_items[new_pos].line_no,\n+\t\t\t\t\t\tedited, moved,\n+\t\t\t\t\t\tcandidates[i].window_score,\n+\t\t\t\t\t\tcandidates[i].line_score,\n+\t\t\t\t\t\tcandidates[i].score);\n+\t}\n+\n+cleanup:\n+\tfree(old_items);\n+\tfree(new_items);\n+\tfree(old_lines);\n+\tfree(new_lines);\n+\tfree(old_tokens);\n+\tfree(new_tokens);\n+\tfree(new_contexts);\n+\tfree(bucket_heads);\n+\tfree(bucket_next);\n+\tfree(bucket_counts);\n+\tfree(candidate_counts);\n+\tfree(candidate_touched);\n+\tfree(used_old);\n+\tfree(used_plus);\n+\tfree(candidates);\n+}\n+\n+static void mark_word_diff_align(struct diff_options *o)\n+{\n+\tint i, hunk_start = -1;\n+\tstruct emitted_diff_symbols *e = o->emitted_symbols;\n+\n+\tfor (i = 0; i < e->nr; i++) {\n+\t\tif (e->buf[i].s == DIFF_SYMBOL_CONTEXT_FRAGINFO) {\n+\t\t\tif (hunk_start >= 0)\n+\t\t\t\tword_diff_align_hunk(e, hunk_start, i);\n+\t\t\thunk_start = i + 1;\n+\t\t} else if (hunk_start >= 0 &&\n+\t\t\t   e->buf[i].s != DIFF_SYMBOL_CONTEXT &&\n+\t\t\t   e->buf[i].s != DIFF_SYMBOL_CONTEXT_INCOMPLETE &&\n+\t\t\t   e->buf[i].s != DIFF_SYMBOL_CONTEXT_MARKER &&\n+\t\t\t   e->buf[i].s != DIFF_SYMBOL_PLUS &&\n+\t\t\t   e->buf[i].s != DIFF_SYMBOL_MINUS) {\n+\t\t\tword_diff_align_hunk(e, hunk_start, i);\n+\t\t\thunk_start = -1;\n+\t\t}\n+\t}\n+\n+\tif (hunk_start >= 0)\n+\t\tword_diff_align_hunk(e, hunk_start, e->nr);\n+}\n+\n static void emit_line_ws_markup(struct diff_options *o,\n \t\t\t\tconst char *set_sign, const char *set,\n \t\t\t\tconst char *reset,\n@@ -5245,6 +6014,10 @@ void diff_setup_done(struct diff_options *options)\n \t\tdie(_(\"options '%s' and '%s' cannot be used together, use '%s' with '%s' and '%s'\"),\n \t\t\t\"--pickaxe-all\", \"--find-object\", \"--pickaxe-all\", \"-G\", \"-S\");\n \n+\tif (options->word_diff_align && options->word_diff)\n+\t\tdie(_(\"options '%s' and '%s' cannot be used together\"),\n+\t\t    \"--word-diff-align\", \"--word-diff\");\n+\n \t/*\n \t * Most of the time we can say \"there are changes\"\n \t * only by checking if there are changed paths, but\n@@ -5971,6 +6744,18 @@ static int diff_opt_word_diff_regex(const struct option *opt,\n \treturn 0;\n }\n \n+static int diff_opt_word_diff_align(const struct option *opt,\n+\t\t\t\t    const char *arg, int unset)\n+{\n+\tstruct diff_options *options = opt->value;\n+\n+\tBUG_ON_OPT_ARG(arg);\n+\toptions->word_diff_align = !unset;\n+\tif (!unset)\n+\t\tenable_patch_output(&options->output_format);\n+\treturn 0;\n+}\n+\n static int diff_opt_rotate_to(const struct option *opt, const char *arg, int unset)\n {\n \tstruct diff_options *options = opt->value;\n@@ -6205,6 +6990,9 @@ struct option *add_diff_options(const struct option *opts,\n \t\tOPT_CALLBACK_F(0, \"word-diff-regex\", options, N_(\"<regex>\"),\n \t\t\t       N_(\"use <regex> to decide what a word is\"),\n \t\t\t       PARSE_OPT_NONEG, diff_opt_word_diff_regex),\n+\t\tOPT_CALLBACK_F(0, \"word-diff-align\", options, NULL,\n+\t\t\t       N_(\"annotate similar moved lines\"),\n+\t\t\t       PARSE_OPT_NOARG, diff_opt_word_diff_align),\n \t\tOPT_CALLBACK_F(0, \"color-words\", options, N_(\"<regex>\"),\n \t\t\t       N_(\"equivalent to --word-diff=color --word-diff-regex=<regex>\"),\n \t\t\t       PARSE_OPT_NONEG | PARSE_OPT_OPTARG, diff_opt_color_words),\n@@ -7076,7 +7864,7 @@ static void diff_flush_patch_all_file_pairs(struct diff_options *o)\n \tif (WSEH_NEW & WS_RULE_MASK)\n \t\tBUG(\"WS rules bit mask overlaps with diff symbol flags\");\n \n-\tif (o->color_moved && want_color(o->use_color))\n+\tif ((o->color_moved && want_color(o->use_color)) || o->word_diff_align)\n \t\to->emitted_symbols = &esm;\n \n \tif (o->additional_path_headers)\n@@ -7092,14 +7880,19 @@ static void diff_flush_patch_all_file_pairs(struct diff_options *o)\n \t\tstruct mem_pool entry_pool;\n \t\tstruct moved_entry_list *entry_list;\n \n-\t\tmem_pool_init(&entry_pool, 1024 * 1024);\n-\t\tentry_list = add_lines_to_move_detection(o, &entry_pool);\n-\t\tmark_color_as_moved(o, entry_list);\n-\t\tif (o->color_moved == COLOR_MOVED_ZEBRA_DIM)\n-\t\t\tdim_moved_lines(o);\n+\t\tif (o->word_diff_align)\n+\t\t\tmark_word_diff_align(o);\n+\n+\t\tif (o->color_moved && want_color(o->use_color)) {\n+\t\t\tmem_pool_init(&entry_pool, 1024 * 1024);\n+\t\t\tentry_list = add_lines_to_move_detection(o, &entry_pool);\n+\t\t\tmark_color_as_moved(o, entry_list);\n+\t\t\tif (o->color_moved == COLOR_MOVED_ZEBRA_DIM)\n+\t\t\t\tdim_moved_lines(o);\n \n-\t\tmem_pool_discard(&entry_pool, 0);\n-\t\tfree(entry_list);\n+\t\t\tmem_pool_discard(&entry_pool, 0);\n+\t\t\tfree(entry_list);\n+\t\t}\n \n \t\tfor (i = 0; i < esm.nr; i++)\n \t\t\temit_diff_symbol_from_struct(o, &esm.buf[i]);\ndiff --git a/diff.h b/diff.h\nindex 7eb84aadf..b9a20b7f2 100644\n--- a/diff.h\n+++ b/diff.h\n@@ -351,6 +351,7 @@ struct diff_options {\n \tint stat_count;\n \tconst char *word_regex;\n \tenum diff_words_type word_diff;\n+\tint word_diff_align;\n \tenum diff_submodule_format submodule_format;\n \n \tstruct oidset *objfind;\n-- \n2.39.3 (Apple Git-146)\n"},{"id":"544137","messageId":"20260527042402.13607-3-ainsophyao@gmail.com","threadId":"65698","inReplyTo":"20260527042402.13607-1-ainsophyao@gmail.com","subject":"[RFC PATCH 2/3] diff: render word-diff-align pairs for RFC review","fromName":"Keita Oda","fromEmail":"ainsophyao@gmail.com","sentAt":"2026-05-27T04:24:01Z","receivedAt":"2026-05-27T04:24:16Z","isPatch":true,"body":"From: Keita ODA <ainsophyao@gmail.com>\n\nTeach the RFC prototype to render selected --word-diff-align pairs with\nword-diff-like markers.\n\nThis renderer is deliberately small and local to the RFC.  It exists to make\nthe recovered line pairs inspectable in review output.  It is not meant to be\nthe final UI.  A production version should likely reuse the existing word-diff\nmachinery once the line-pairing question is settled.\n\nThe renderer computes a token LCS for the selected pair and marks the unmatched\nspans with the familiar plain word-diff delimiters:\n\n  [-old-]\n  {+new+}\n\nMoved selected pairs are also marked with DIFF_SYMBOL_MOVED_LINE so that the\ncurrent moved-line coloring can show that the pair came from a moved region.\n\n---\n diff.c | 213 +++++++++++++++++++++++++++++++++++++++++++++++++++++----\n 1 file changed, 200 insertions(+), 13 deletions(-)\n\ndiff --git a/diff.c b/diff.c\nindex 6b8744920..8629d4670 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -1811,6 +1811,104 @@ static void word_diff_align_add_candidate(struct word_diff_align_candidate **can\n \tcandidate->line_shared = line_shared;\n }\n \n+/*\n+ * RFC-only formatter for exposing the selected line pairs.  The final\n+ * presentation should reuse the normal word-diff machinery instead of this\n+ * small debug renderer.\n+ */\n+struct word_diff_align_debug_token {\n+\tint start;\n+\tint end;\n+};\n+\n+static void word_diff_align_debug_collect_tokens(const struct emitted_diff_symbol *line,\n+\t\t\t\t\t\t struct word_diff_align_debug_token **tokens,\n+\t\t\t\t\t\t int *tokens_nr, int *tokens_alloc)\n+{\n+\tint len = word_diff_align_payload_len(line);\n+\tint pos = 0, start, end;\n+\n+\twhile (word_diff_align_next_token(line->line, len, &pos, &start, &end)) {\n+\t\tALLOC_GROW(*tokens, *tokens_nr + 1, *tokens_alloc);\n+\t\t(*tokens)[*tokens_nr].start = start;\n+\t\t(*tokens)[*tokens_nr].end = end;\n+\t\t(*tokens_nr)++;\n+\t}\n+}\n+\n+static void word_diff_align_debug_add_span(struct strbuf *out,\n+\t\t\t\t\t   const char *open,\n+\t\t\t\t\t   const char *line, int len,\n+\t\t\t\t\t   const char *close)\n+{\n+\tif (!len)\n+\t\treturn;\n+\tstrbuf_addstr(out, open);\n+\tstrbuf_add(out, line, len);\n+\tstrbuf_addstr(out, close);\n+}\n+\n+static int word_diff_align_debug_token_eq(const struct emitted_diff_symbol *a,\n+\t\t\t\t\t  const struct word_diff_align_debug_token *a_tok,\n+\t\t\t\t\t  const struct emitted_diff_symbol *b,\n+\t\t\t\t\t  const struct word_diff_align_debug_token *b_tok)\n+{\n+\tint a_len = a_tok->end - a_tok->start;\n+\tint b_len = b_tok->end - b_tok->start;\n+\n+\treturn a_len == b_len &&\n+\t\t!memcmp(a->line + a_tok->start, b->line + b_tok->start, a_len);\n+}\n+\n+static void word_diff_align_debug_rewrite_line(struct emitted_diff_symbol *line,\n+\t\t\t\t\t       struct word_diff_align_debug_token *tokens,\n+\t\t\t\t\t       int tokens_nr, int *match_to,\n+\t\t\t\t\t       const struct emitted_diff_symbol *other,\n+\t\t\t\t\t       struct word_diff_align_debug_token *other_tokens,\n+\t\t\t\t\t       const char *open, const char *close)\n+{\n+\tstruct strbuf out = STRBUF_INIT;\n+\tchar *old_line = (char *)line->line;\n+\tint payload_len = word_diff_align_payload_len(line);\n+\tint other_pos = 0;\n+\tint other_payload_len = word_diff_align_payload_len(other);\n+\tint pos = 0, i;\n+\tsize_t new_len;\n+\n+\tfor (i = 0; i < tokens_nr; i++) {\n+\t\tint other_i = match_to[i];\n+\t\tint gap_len, other_gap_len;\n+\n+\t\tif (other_i < 0)\n+\t\t\tcontinue;\n+\t\tgap_len = tokens[i].start - pos;\n+\t\tother_gap_len = other_tokens[other_i].start - other_pos;\n+\t\tif (gap_len == other_gap_len &&\n+\t\t    !memcmp(line->line + pos, other->line + other_pos, gap_len))\n+\t\t\tstrbuf_add(&out, line->line + pos, gap_len);\n+\t\telse\n+\t\t\tword_diff_align_debug_add_span(&out, open,\n+\t\t\t\t\t\t       line->line + pos,\n+\t\t\t\t\t\t       gap_len, close);\n+\t\tstrbuf_add(&out, line->line + tokens[i].start,\n+\t\t\t   tokens[i].end - tokens[i].start);\n+\t\tpos = tokens[i].end;\n+\t\tother_pos = other_tokens[other_i].end;\n+\t}\n+\tif (payload_len - pos == other_payload_len - other_pos &&\n+\t    !memcmp(line->line + pos, other->line + other_pos,\n+\t\t    payload_len - pos))\n+\t\tstrbuf_add(&out, line->line + pos, payload_len - pos);\n+\telse\n+\t\tword_diff_align_debug_add_span(&out, open, line->line + pos,\n+\t\t\t\t\t       payload_len - pos, close);\n+\tstrbuf_add(&out, line->line + payload_len, line->len - payload_len);\n+\n+\tline->line = strbuf_detach(&out, &new_len);\n+\tline->len = (int)new_len;\n+\tfree(old_line);\n+}\n+\n static void word_diff_align_debug_append_comment(struct emitted_diff_symbol *line,\n \t\t\t\t\t\t const struct strbuf *suffix)\n {\n@@ -1835,25 +1933,114 @@ static void word_diff_align_debug_mark_pair(struct emitted_diff_symbol *minus_li\n \t\t\t\t\t    struct emitted_diff_symbol *plus_line,\n \t\t\t\t\t    int minus_lineno, int plus_lineno,\n \t\t\t\t\t    int changed, int moved,\n-\t\t\t\t\t\t    int window_score,\n-\t\t\t\t\t\t    int line_score,\n-\t\t\t\t\t\t    int pair_score)\n-{\n-\tstruct strbuf suffix = STRBUF_INIT;\n+\t\t\t\t\t    int window_score,\n+\t\t\t\t\t    int line_score,\n+\t\t\t\t\t    int pair_score)\n+{\n+\tstruct word_diff_align_debug_token *minus_tokens = NULL, *plus_tokens = NULL;\n+\tint minus_tokens_nr = 0, minus_tokens_alloc = 0;\n+\tint plus_tokens_nr = 0, plus_tokens_alloc = 0;\n+\tint *minus_match_to = NULL, *plus_match_to = NULL;\n+\tint *lcs = NULL;\n+\tstruct emitted_diff_symbol minus_original = *minus_line;\n+\tstruct emitted_diff_symbol plus_original = *plus_line;\n+\tint i, j, columns;\n \n \tif (moved) {\n \t\tminus_line->flags |= DIFF_SYMBOL_MOVED_LINE;\n \t\tplus_line->flags |= DIFF_SYMBOL_MOVED_LINE;\n \t}\n \n-\tstrbuf_addf(&suffix,\n-\t\t    \"        # aligned from %d to %d, %s, W=%d L=%d S=%d\",\n-\t\t    minus_lineno, plus_lineno,\n-\t\t    changed ? \"edited\" : \"unchanged\",\n-\t\t    window_score, line_score, pair_score);\n-\tword_diff_align_debug_append_comment(minus_line, &suffix);\n-\tword_diff_align_debug_append_comment(plus_line, &suffix);\n-\tstrbuf_release(&suffix);\n+\tminus_original.line = xmemdupz(minus_line->line, minus_line->len);\n+\tplus_original.line = xmemdupz(plus_line->line, plus_line->len);\n+\tif (!changed)\n+\t\tgoto comment;\n+\n+\tword_diff_align_debug_collect_tokens(minus_line, &minus_tokens,\n+\t\t\t\t\t     &minus_tokens_nr,\n+\t\t\t\t\t     &minus_tokens_alloc);\n+\tword_diff_align_debug_collect_tokens(plus_line, &plus_tokens,\n+\t\t\t\t\t     &plus_tokens_nr,\n+\t\t\t\t\t     &plus_tokens_alloc);\n+\tif (!minus_tokens_nr || !plus_tokens_nr)\n+\t\tgoto comment;\n+\n+\tcolumns = plus_tokens_nr + 1;\n+\tCALLOC_ARRAY(lcs, (minus_tokens_nr + 1) * columns);\n+\tALLOC_ARRAY(minus_match_to, minus_tokens_nr);\n+\tALLOC_ARRAY(plus_match_to, plus_tokens_nr);\n+\tfor (i = 0; i < minus_tokens_nr; i++)\n+\t\tminus_match_to[i] = -1;\n+\tfor (j = 0; j < plus_tokens_nr; j++)\n+\t\tplus_match_to[j] = -1;\n+\n+\tfor (i = 1; i <= minus_tokens_nr; i++) {\n+\t\tfor (j = 1; j <= plus_tokens_nr; j++) {\n+\t\t\tif (word_diff_align_debug_token_eq(minus_line,\n+\t\t\t\t\t\t\t   &minus_tokens[i - 1],\n+\t\t\t\t\t\t\t   plus_line,\n+\t\t\t\t\t\t\t   &plus_tokens[j - 1]))\n+\t\t\t\tlcs[i * columns + j] =\n+\t\t\t\t\tlcs[(i - 1) * columns + j - 1] + 1;\n+\t\t\telse if (lcs[(i - 1) * columns + j] >\n+\t\t\t\t lcs[i * columns + j - 1])\n+\t\t\t\tlcs[i * columns + j] = lcs[(i - 1) * columns + j];\n+\t\t\telse\n+\t\t\t\tlcs[i * columns + j] = lcs[i * columns + j - 1];\n+\t\t}\n+\t}\n+\n+\ti = minus_tokens_nr;\n+\tj = plus_tokens_nr;\n+\twhile (i > 0 && j > 0) {\n+\t\tif (lcs[i * columns + j] == lcs[i * columns + j - 1]) {\n+\t\t\tj--;\n+\t\t} else if (lcs[i * columns + j] ==\n+\t\t\t   lcs[(i - 1) * columns + j]) {\n+\t\t\ti--;\n+\t\t} else if (word_diff_align_debug_token_eq(minus_line,\n+\t\t\t\t\t\t\t  &minus_tokens[i - 1],\n+\t\t\t\t\t\t\t  plus_line,\n+\t\t\t\t\t\t\t  &plus_tokens[j - 1])) {\n+\t\t\tminus_match_to[i - 1] = j - 1;\n+\t\t\tplus_match_to[j - 1] = i - 1;\n+\t\t\ti--;\n+\t\t\tj--;\n+\t\t} else {\n+\t\t\tBUG(\"word-diff-align display LCS backtrack failed\");\n+\t\t}\n+\t}\n+\n+\tword_diff_align_debug_rewrite_line(minus_line, minus_tokens,\n+\t\t\t\t\t   minus_tokens_nr, minus_match_to,\n+\t\t\t\t\t   &plus_original, plus_tokens,\n+\t\t\t\t\t   \"[-\", \"-]\");\n+\tword_diff_align_debug_rewrite_line(plus_line, plus_tokens,\n+\t\t\t\t\t   plus_tokens_nr, plus_match_to,\n+\t\t\t\t\t   &minus_original, minus_tokens,\n+\t\t\t\t\t   \"{+\", \"+}\");\n+\n+comment:\n+\t{\n+\t\tstruct strbuf suffix = STRBUF_INIT;\n+\n+\t\tstrbuf_addf(&suffix,\n+\t\t\t    \"        # aligned from %d to %d, %s, W=%d L=%d S=%d\",\n+\t\t\t    minus_lineno, plus_lineno,\n+\t\t\t    changed ? \"edited\" : \"unchanged\",\n+\t\t\t    window_score, line_score, pair_score);\n+\t\tword_diff_align_debug_append_comment(minus_line, &suffix);\n+\t\tword_diff_align_debug_append_comment(plus_line, &suffix);\n+\t\tstrbuf_release(&suffix);\n+\t}\n+\n+\tfree((char *)minus_original.line);\n+\tfree((char *)plus_original.line);\n+\tfree(minus_tokens);\n+\tfree(plus_tokens);\n+\tfree(minus_match_to);\n+\tfree(plus_match_to);\n+\tfree(lcs);\n }\n \n static void word_diff_align_add_item(struct word_diff_align_item **items,\n-- \n2.39.3 (Apple Git-146)\n"},{"id":"544139","messageId":"20260527042402.13607-4-ainsophyao@gmail.com","threadId":"65698","inReplyTo":"20260527042402.13607-1-ainsophyao@gmail.com","subject":"[RFC PATCH 3/3] t4034: cover moved-and-edited word diff alignment","fromName":"Keita Oda","fromEmail":"ainsophyao@gmail.com","sentAt":"2026-05-27T04:24:02Z","receivedAt":"2026-05-27T04:24:17Z","isPatch":true,"body":"From: Keita ODA <ainsophyao@gmail.com>\n\nAdd a focused --word-diff-align test with a moved permission-style block where\none value changes inside the moved block.\n\nThe test is intentionally small.  It checks that the changed line is paired and\nthat the RFC renderer exposes the old and new value with word-diff-like\nmarkers.\n\n---\n t/t4034-diff-words.sh | 46 +++++++++++++++++++++++++++++++++++++++++++\n 1 file changed, 46 insertions(+)\n\ndiff --git a/t/t4034-diff-words.sh b/t/t4034-diff-words.sh\nindex 0be647c2f..7bf696b17 100755\n--- a/t/t4034-diff-words.sh\n+++ b/t/t4034-diff-words.sh\n@@ -148,6 +148,52 @@ test_expect_success '--word-diff=plain' '\n \tword_diff --word-diff=plain --no-color\n '\n \n+test_expect_success '--word-diff-align marks moved-and-edited block lines' '\n+\tcat >old <<-\\EOF &&\n+\t#define LIMIT_PUBLIC_UPLOADS       25\n+\t#define LIMIT_PUBLIC_INVITES       8\n+\n+\t/* Public resource permissions. */\n+\t#define PERM_RESOURCE_READ         0x0001\n+\t#define PERM_RESOURCE_LIST         0x0002\n+\t#define PERM_RESOURCE_COMMENT      0x0004\n+\t#define PERM_RESOURCE_EXPORT       0x0008\n+\t#define PERM_RESOURCE_SHARE        0x0010\n+\t#define PERM_RESOURCE_ARCHIVE      0x0020\n+\t#define PERM_RESOURCE_DELETE       0x0040\n+\t#define PERM_RESOURCE_ADMIN        0x0080\n+\n+\t#define PERM_INTERNAL_DEBUG        0x0100\n+\t#define PERM_INTERNAL_IMPERSONATE  0x0200\n+\n+\t#define AUDIT_POLICY_STRICT        1\n+\t#define AUDIT_POLICY_VERBOSE       2\n+\tEOF\n+\tcat >new <<-\\EOF &&\n+\t#define LIMIT_PUBLIC_UPLOADS       25\n+\t#define LIMIT_PUBLIC_INVITES       8\n+\n+\t#define PERM_INTERNAL_DEBUG        0x0100\n+\t#define PERM_INTERNAL_IMPERSONATE  0x0200\n+\n+\t#define AUDIT_POLICY_STRICT        1\n+\t#define AUDIT_POLICY_VERBOSE       2\n+\n+\t/* Public resource permissions. */\n+\t#define PERM_RESOURCE_READ         0x0001\n+\t#define PERM_RESOURCE_LIST         0x0002\n+\t#define PERM_RESOURCE_COMMENT      0x0004\n+\t#define PERM_RESOURCE_EXPORT       0x0001\n+\t#define PERM_RESOURCE_SHARE        0x0010\n+\t#define PERM_RESOURCE_ARCHIVE      0x0020\n+\t#define PERM_RESOURCE_DELETE       0x0040\n+\t#define PERM_RESOURCE_ADMIN        0x0080\n+\tEOF\n+\ttest_must_fail git diff --no-index --histogram --word-diff-align old new >actual &&\n+\ttest_grep \"^-#define PERM_RESOURCE_EXPORT\\\\[-.*0x0008-\\\\]\" actual &&\n+\ttest_grep \"^+#define PERM_RESOURCE_EXPORT{+.*0x0001+}\" actual\n+'\n+\n test_expect_success '--word-diff=plain --color' '\n \tcat >expect <<-EOF &&\n \t\t<BOLD>diff --git a/pre b/post<RESET>\n-- \n2.39.3 (Apple Git-146)\n"}]}