{"thread":{"id":"64988","subject":"[PATCH] diff --anchored: avoid checking unmatched lines","startedAt":"2026-02-12T15:54:06Z","lastAt":"2026-02-12T15:54:06Z","messageCount":1,"participants":["Phillip Wood"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"535871","messageId":"2a8cc2d6c37f25a58823b501500165d597321749.1770911599.git.phillip.wood@dunelm.org.uk","threadId":"64988","inReplyTo":null,"subject":"[PATCH] diff --anchored: avoid checking unmatched lines","fromName":"Phillip Wood","fromEmail":"phillip.wood123@gmail.com","sentAt":"2026-02-12T15:53:50Z","receivedAt":"2026-02-12T15:54:06Z","isPatch":true,"sender":{"key":"phillip.wood@dunelm.org.uk","avatar":null},"body":"From: Phillip Wood <phillip.wood@dunelm.org.uk>\n\nFor a line to be an anchor it has to appear in each of the files being\ndiffed exactly once. With that in mind lets delay checking whether\na line is an anchor until we know there is exactly one instance of\nthe line in each file. As each line is checked at most once, there\nis no need to cache the result of is_anchor() and we can drop that\nfield from the hashmap entries. When diffing 5000 recent commits in\ngit.git this gives a modest speedup of ~2%. In the (rather extreme)\nexample below that consists largely of deletions the speedup is ~16%.\n\n    seq 0 10000000 >old\n    printf '%s\\n' 300000 100000 200000 >new\n    git diff --no-index --anchored=300000 old new\n\nSigned-off-by: Phillip Wood <phillip.wood@dunelm.org.uk>\n---\nBase-Commit: ea24e2c55433012a0a6c4ae947a87bc66404e484\nPublished-As: https://github.com/phillipwood/git/releases/tag/pw%2Fxdiff-simplify-anchor%2Fv1\nView-Changes-At: https://github.com/phillipwood/git/compare/ea24e2c55...2a8cc2d6c\nFetch-It-Via: git fetch https://github.com/phillipwood/git pw/xdiff-simplify-anchor/v1\n\n xdiff/xpatience.c | 18 ++++++------------\n 1 file changed, 6 insertions(+), 12 deletions(-)\n\ndiff --git a/xdiff/xpatience.c b/xdiff/xpatience.c\nindex 9580d180320..7953490ed0d 100644\n--- a/xdiff/xpatience.c\n+++ b/xdiff/xpatience.c\n@@ -61,12 +61,6 @@ struct hashmap {\n \t\t * initially, \"next\" reflects only the order in file1.\n \t\t */\n \t\tstruct entry *next, *previous;\n-\n-\t\t/*\n-\t\t * If 1, this entry can serve as an anchor. See\n-\t\t * Documentation/diff-options.adoc for more information.\n-\t\t */\n-\t\tunsigned anchor : 1;\n \t} *entries, *first, *last;\n \t/* were common records found? */\n \tunsigned long has_matches;\n@@ -85,8 +79,7 @@ static int is_anchor(xpparam_t const *xpp, const char *line)\n }\n \n /* The argument \"pass\" is 1 for the first file, 2 for the second. */\n-static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n-\t\t\t  int pass)\n+static void insert_record(int line, struct hashmap *map, int pass)\n {\n \txrecord_t *records = pass == 1 ?\n \t\tmap->env->xdf1.recs : map->env->xdf2.recs;\n@@ -121,7 +114,6 @@ static void insert_record(xpparam_t const *xpp, int line, struct hashmap *map,\n \t\treturn;\n \tmap->entries[index].line1 = line;\n \tmap->entries[index].minimal_perfect_hash = record->minimal_perfect_hash;\n-\tmap->entries[index].anchor = is_anchor(xpp, (const char *)map->env->xdf1.recs[line - 1].ptr);\n \tif (!map->first)\n \t\tmap->first = map->entries + index;\n \tif (map->last) {\n@@ -153,11 +145,11 @@ static int fill_hashmap(xpparam_t const *xpp, xdfenv_t *env,\n \n \t/* First, fill with entries from the first file */\n \twhile (count1--)\n-\t\tinsert_record(xpp, line1++, result, 1);\n+\t\tinsert_record(line1++, result, 1);\n \n \t/* Then search for matches in the second file */\n \twhile (count2--)\n-\t\tinsert_record(xpp, line2++, result, 2);\n+\t\tinsert_record(line2++, result, 2);\n \n \treturn 0;\n }\n@@ -194,6 +186,8 @@ static int binary_search(struct entry **sequence, int longest,\n  */\n static int find_longest_common_sequence(struct hashmap *map, struct entry **res)\n {\n+\txpparam_t const *xpp = map->xpp;\n+\txrecord_t const *recs = map->env->xdf2.recs;\n \tstruct entry **sequence;\n \tint longest = 0, i;\n \tstruct entry *entry;\n@@ -220,7 +214,7 @@ static int find_longest_common_sequence(struct hashmap *map, struct entry **res)\n \t\tif (i <= anchor_i)\n \t\t\tcontinue;\n \t\tsequence[i] = entry;\n-\t\tif (entry->anchor) {\n+\t\tif (is_anchor(xpp, (const char*)recs[entry->line2 - 1].ptr)) {\n \t\t\tanchor_i = i;\n \t\t\tlongest = anchor_i + 1;\n \t\t} else if (i == longest) {\n-- \n2.52.0.362.g884e03848a9\n\n"}]}