{"thread":{"id":"51354","subject":"[PATCH v9 0/9] blame: add the ability to ignore commits","startedAt":"2019-06-20T16:38:31Z","lastAt":"2019-07-01T14:16:12Z","messageCount":13,"participants":["Barret Rhoden","SZEDER Gábor","michael@platin.gs"],"isPatch":true,"patchVersion":9,"patchTotal":9},"messages":[{"id":"377666","messageId":"20190620163820.231316-1-brho@google.com","threadId":"51354","inReplyTo":null,"subject":"[PATCH v9 0/9] blame: add the ability to ignore commits","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:11Z","receivedAt":"2019-06-20T16:38:31Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"This patch set adds the ability to ignore a set of commits and their\nchanges when blaming.  This can be used to ignore a commit deemed 'not\ninteresting,' such as reformatting.\n\nv8 -> v9\nv8: https://public-inbox.org/git/20190610153014.42055-1-brho@google.com\n- Fixed tests that had git-blame's output piped to another command\n- Rebased onto master\n\nv7 -> v8\nv7: https://public-inbox.org/git/20190515214503.77162-1-brho@google.com\n- Added Junio's fixes for test portability.\n- Removed unreachable code exposed by Derrick's test coverage report.\n- Added a test to ensure blame_coalesce() is covered.\n- Rebased onto master.\n\nv6 -> v7\nv6: https://public-inbox.org/git/20190410162409.117264-1-brho@google.com\n- Split the init_skiplist commit into two commits: \"change variable\n  names\" then \"move the function\".\n- Fixed the test's usage of grep, from grep \"\\+\" to grep -E \"+\".\n- Fixed comments related to fsck.skipList, added them to\n  config/blame.txt\n- A line in the blame output is either \"ignored\" or \"unblamable\", not\n  both.\n- Changed the way we mark lines.  In particular, we don't zero-out the\n  hash anymore for unblamables, since all zeros already had a meaning.\n  We also distinguish between ignored and unblamable.  Here's the new\n  style:\n\t? for ignored\n\t* for unblamable\n  Both of those markings are controlled by config vars; the discussion\n  on the list shows that no default style works for everyone:\n        if blame.markIgnoredLines\n\t    Line was attributed to a commit that was not the most recent\n\t    to change it (i.e. the ignored commit) and will be marked\n\t    with '?'.  Lines touched by an ignored commit that we could\n\t    not blame on another are unmarked.\n        if blame.markUnblamableLines\n\t    Lines touched by an ignored commit that we could not blame\n\t    on another are marked with *.  I wanted to differentiate\n\t    between Ignored and Unblamable, so a single ? isn't enough.\n- Added Michael's fuzzy fingerprinting code.\n- We guess_line_blames() for an entire chunk, instead of per-blame\n  entry.  A diff chunk can be made up of more than one blame_entry,\n  which made the job of the heuristic unnecessarily difficult.\n- Rebased onto master.\n\nv5 -> v6\nv5: https://public-inbox.org/git/20190403160207.149174-1-brho@google.com/\n- The \"guess\" heuristic can now look anywhere in the parent file for a\n  matching line, instead of just looking in the parent chunk.  The\n  chunks passed to blame_chunk() are smaller than you'd expect: they are\n  just adjacent '-' and '+' sections.  Any diff 'context' is a chunk\n  boundary.\n- Fixed the parent_len calculation.  I had been basing it off of\n  e->num_lines, and treating the blame entry as if it was the target\n  chunk, but the individual blame entries are subsets of the chunk.  I\n  just pass the parent chunk info all the way through now.\n- Use Michael's newest fingerprinting code, which is a large speedup.\n- Made a config option to zero the hash for an ignored line when the\n  heuristic could not find a line in the parent to blame.  Previously,\n  this was always 'on'.\n- Moved the for loop variable declarations out of the for ().\n- Rebased on master.\n\nv4 -> v5\nv4: https://public-inbox.org/git/20190226170648.211847-1-brho@google.com/\n- Changed the handling of blame_entries from ignored commits so that you\n  can use any algorithm you want to map lines from the diff chunk to\n  different parts of the parent commit.\n- fill_origin_blob() optionally can track the offsets of the start of\n  every line, similar to what we do in the scoreboard for the final\n  file.  This can be used by the matching algorithm.  It has no effect\n  if you are not ignoring commits.\n- RFC of a fuzzy/fingerprinting heuristic, based on Michael Platings RFC\n  at https://public-inbox.org/git/20190324235020.49706-2-michael@platin.gs/\n- Made the tests that detect unblamable entries more resilient to\n  different heuristics.\n- Fixed a few bugs:\n\t- tests were not grepping the line number from --line-porcelain\n\t  correctly.\n\t- In the old version, when I passed the \"upper\" part of the\n\t  blame entry to the target and marked unblamable, the suspect\n\t  was incorrectly marked as the parent.  The s_lno was also in\n\t  the parent's address space.\n\nv3 -> v4\nv3: https://public-inbox.org/git/20190212222722.240676-1-brho@google.com/\n- Cleaned up the tests, especially removing usage of sed -i.\n- Squashed the 'tests' commit into the other blame commits.  Let me know\n  if you'd like further squashing.\n\nv2 -> v3\nv2: https://public-inbox.org/git/20190117202919.157326-1-brho@google.com/\n- SHA-1 -> \"object name\", and fixed other comments\n- Changed error string for oidset_parse_file()\n- Adjusted existing fsck tests to handle those string changes\n- Return hash of all zeros for lines we know we cannot identify\n- Allow repeated options for blame.ignoreRevsFile and\n  --ignore-revs-file.  An empty file name resets the list.  Config\n  options are parsed before the command line options.\n- Rebased to master\n- Added regression tests\n\nv1 -> v2\nv1: https://public-inbox.org/git/20190107213013.231514-1-brho@google.com/\n- extracted the skiplist from fsck to avoid duplicating code\n- overhauled the interface and options\n- split out markIgnoredFiles\n- handled merges\n\nBarret Rhoden (8):\n  fsck: rename and touch up init_skiplist()\n  Move oidset_parse_file() to oidset.c\n  blame: use a helper function in blame_chunk()\n  blame: add the ability to ignore commits and their changes\n  blame: add config options for the output of ignored or unblamable\n    lines\n  blame: optionally track line fingerprints during fill_blame_origin()\n  blame: use the fingerprint heuristic to match ignored lines\n  blame: add a test to cover blame_coalesce()\n\nMichael Platings (1):\n  blame: add a fingerprint heuristic to match ignored lines\n\n Documentation/blame-options.txt |   19 +\n Documentation/config/blame.txt  |   16 +\n Documentation/git-blame.txt     |    1 +\n blame.c                         | 1016 +++++++++++++++++++++++++++++--\n blame.h                         |    6 +\n builtin/blame.c                 |   56 ++\n fsck.c                          |   37 +-\n oidset.c                        |   35 ++\n oidset.h                        |    8 +\n t/t5504-fetch-receive-strict.sh |   14 +-\n t/t8003-blame-corner-cases.sh   |   36 ++\n t/t8013-blame-ignore-revs.sh    |  274 +++++++++\n t/t8014-blame-ignore-fuzzy.sh   |  437 +++++++++++++\n 13 files changed, 1856 insertions(+), 99 deletions(-)\n create mode 100755 t/t8013-blame-ignore-revs.sh\n create mode 100755 t/t8014-blame-ignore-fuzzy.sh\n\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"377667","messageId":"20190620163820.231316-2-brho@google.com","threadId":"51354","inReplyTo":"20190620163820.231316-1-brho@google.com","subject":"[PATCH v9 1/9] fsck: rename and touch up init_skiplist()","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:12Z","receivedAt":"2019-06-20T16:38:33Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"init_skiplist() took a file consisting of SHA-1s and comments and added\nthe objects to an oidset.  This functionality is useful for other\ncommands and will be moved to oidset.c in a future commit.\n\nIn preparation for that move, this commit renames it to\noidset_parse_file() to reflect its more generic usage and cleans up a\nfew of the names.\n\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n fsck.c                          | 18 +++++++++---------\n t/t5504-fetch-receive-strict.sh | 14 +++++++-------\n 2 files changed, 16 insertions(+), 16 deletions(-)\n\ndiff --git a/fsck.c b/fsck.c\nindex 4703f5556145..a28cba6b05dd 100644\n--- a/fsck.c\n+++ b/fsck.c\n@@ -181,7 +181,7 @@ static int fsck_msg_type(enum fsck_msg_id msg_id,\n \treturn msg_type;\n }\n \n-static void init_skiplist(struct fsck_options *options, const char *path)\n+void oidset_parse_file(struct oidset *set, const char *path)\n {\n \tFILE *fp;\n \tstruct strbuf sb = STRBUF_INIT;\n@@ -189,26 +189,26 @@ static void init_skiplist(struct fsck_options *options, const char *path)\n \n \tfp = fopen(path, \"r\");\n \tif (!fp)\n-\t\tdie(\"Could not open skip list: %s\", path);\n+\t\tdie(\"could not open object name list: %s\", path);\n \twhile (!strbuf_getline(&sb, fp)) {\n \t\tconst char *p;\n-\t\tconst char *hash;\n+\t\tconst char *name;\n \n \t\t/*\n \t\t * Allow trailing comments, leading whitespace\n \t\t * (including before commits), and empty or whitespace\n \t\t * only lines.\n \t\t */\n-\t\thash = strchr(sb.buf, '#');\n-\t\tif (hash)\n-\t\t\tstrbuf_setlen(&sb, hash - sb.buf);\n+\t\tname = strchr(sb.buf, '#');\n+\t\tif (name)\n+\t\t\tstrbuf_setlen(&sb, name - sb.buf);\n \t\tstrbuf_trim(&sb);\n \t\tif (!sb.len)\n \t\t\tcontinue;\n \n \t\tif (parse_oid_hex(sb.buf, &oid, &p) || *p != '\\0')\n-\t\t\tdie(\"Invalid SHA-1: %s\", sb.buf);\n-\t\toidset_insert(&options->skiplist, &oid);\n+\t\t\tdie(\"invalid object name: %s\", sb.buf);\n+\t\toidset_insert(set, &oid);\n \t}\n \tif (ferror(fp))\n \t\tdie_errno(\"Could not read '%s'\", path);\n@@ -284,7 +284,7 @@ void fsck_set_msg_types(struct fsck_options *options, const char *values)\n \t\tif (!strcmp(buf, \"skiplist\")) {\n \t\t\tif (equal == len)\n \t\t\t\tdie(\"skiplist requires a path\");\n-\t\t\tinit_skiplist(options, buf + equal + 1);\n+\t\t\toidset_parse_file(&options->skiplist, buf + equal + 1);\n \t\t\tbuf += len + 1;\n \t\t\tcontinue;\n \t\t}\ndiff --git a/t/t5504-fetch-receive-strict.sh b/t/t5504-fetch-receive-strict.sh\nindex 7bc706873c5b..fdfe179b1188 100755\n--- a/t/t5504-fetch-receive-strict.sh\n+++ b/t/t5504-fetch-receive-strict.sh\n@@ -164,9 +164,9 @@ test_expect_success 'fsck with unsorted skipList' '\n test_expect_success 'fsck with invalid or bogus skipList input' '\n \tgit -c fsck.skipList=/dev/null -c fsck.missingEmail=ignore fsck &&\n \ttest_must_fail git -c fsck.skipList=does-not-exist -c fsck.missingEmail=ignore fsck 2>err &&\n-\ttest_i18ngrep \"Could not open skip list: does-not-exist\" err &&\n+\ttest_i18ngrep \"could not open.*: does-not-exist\" err &&\n \ttest_must_fail git -c fsck.skipList=.git/config -c fsck.missingEmail=ignore fsck 2>err &&\n-\ttest_i18ngrep \"Invalid SHA-1: \\[core\\]\" err\n+\ttest_i18ngrep \"invalid object name: \\[core\\]\" err\n '\n \n test_expect_success 'fsck with other accepted skipList input (comments & empty lines)' '\n@@ -193,7 +193,7 @@ test_expect_success 'fsck no garbage output from comments & empty lines errors'\n test_expect_success 'fsck with invalid abbreviated skipList input' '\n \techo $commit | test_copy_bytes 20 >SKIP.abbreviated &&\n \ttest_must_fail git -c fsck.skipList=SKIP.abbreviated fsck 2>err-abbreviated &&\n-\ttest_i18ngrep \"^fatal: Invalid SHA-1: \" err-abbreviated\n+\ttest_i18ngrep \"^fatal: invalid object name: \" err-abbreviated\n '\n \n test_expect_success 'fsck with exhaustive accepted skipList input (various types of comments etc.)' '\n@@ -226,10 +226,10 @@ test_expect_success 'push with receive.fsck.skipList' '\n \ttest_must_fail git push --porcelain dst bogus &&\n \tgit --git-dir=dst/.git config receive.fsck.skipList does-not-exist &&\n \ttest_must_fail git push --porcelain dst bogus 2>err &&\n-\ttest_i18ngrep \"Could not open skip list: does-not-exist\" err &&\n+\ttest_i18ngrep \"could not open.*: does-not-exist\" err &&\n \tgit --git-dir=dst/.git config receive.fsck.skipList config &&\n \ttest_must_fail git push --porcelain dst bogus 2>err &&\n-\ttest_i18ngrep \"Invalid SHA-1: \\[core\\]\" err &&\n+\ttest_i18ngrep \"invalid object name: \\[core\\]\" err &&\n \n \tgit --git-dir=dst/.git config receive.fsck.skipList SKIP &&\n \tgit push --porcelain dst bogus\n@@ -255,10 +255,10 @@ test_expect_success 'fetch with fetch.fsck.skipList' '\n \ttest_must_fail git --git-dir=dst/.git fetch \"file://$(pwd)\" $refspec &&\n \tgit --git-dir=dst/.git config fetch.fsck.skipList does-not-exist &&\n \ttest_must_fail git --git-dir=dst/.git fetch \"file://$(pwd)\" $refspec 2>err &&\n-\ttest_i18ngrep \"Could not open skip list: does-not-exist\" err &&\n+\ttest_i18ngrep \"could not open.*: does-not-exist\" err &&\n \tgit --git-dir=dst/.git config fetch.fsck.skipList dst/.git/config &&\n \ttest_must_fail git --git-dir=dst/.git fetch \"file://$(pwd)\" $refspec 2>err &&\n-\ttest_i18ngrep \"Invalid SHA-1: \\[core\\]\" err &&\n+\ttest_i18ngrep \"invalid object name: \\[core\\]\" err &&\n \n \tgit --git-dir=dst/.git config fetch.fsck.skipList dst/.git/SKIP &&\n \tgit --git-dir=dst/.git fetch \"file://$(pwd)\" $refspec\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"377668","messageId":"20190620163820.231316-3-brho@google.com","threadId":"51354","inReplyTo":"20190620163820.231316-1-brho@google.com","subject":"[PATCH v9 2/9] Move oidset_parse_file() to oidset.c","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:13Z","receivedAt":"2019-06-20T16:38:36Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"Signed-off-by: Barret Rhoden <brho@google.com>\n---\n fsck.c   | 35 -----------------------------------\n oidset.c | 35 +++++++++++++++++++++++++++++++++++\n oidset.h |  8 ++++++++\n 3 files changed, 43 insertions(+), 35 deletions(-)\n\ndiff --git a/fsck.c b/fsck.c\nindex a28cba6b05dd..58ff3c4de992 100644\n--- a/fsck.c\n+++ b/fsck.c\n@@ -181,41 +181,6 @@ static int fsck_msg_type(enum fsck_msg_id msg_id,\n \treturn msg_type;\n }\n \n-void oidset_parse_file(struct oidset *set, const char *path)\n-{\n-\tFILE *fp;\n-\tstruct strbuf sb = STRBUF_INIT;\n-\tstruct object_id oid;\n-\n-\tfp = fopen(path, \"r\");\n-\tif (!fp)\n-\t\tdie(\"could not open object name list: %s\", path);\n-\twhile (!strbuf_getline(&sb, fp)) {\n-\t\tconst char *p;\n-\t\tconst char *name;\n-\n-\t\t/*\n-\t\t * Allow trailing comments, leading whitespace\n-\t\t * (including before commits), and empty or whitespace\n-\t\t * only lines.\n-\t\t */\n-\t\tname = strchr(sb.buf, '#');\n-\t\tif (name)\n-\t\t\tstrbuf_setlen(&sb, name - sb.buf);\n-\t\tstrbuf_trim(&sb);\n-\t\tif (!sb.len)\n-\t\t\tcontinue;\n-\n-\t\tif (parse_oid_hex(sb.buf, &oid, &p) || *p != '\\0')\n-\t\t\tdie(\"invalid object name: %s\", sb.buf);\n-\t\toidset_insert(set, &oid);\n-\t}\n-\tif (ferror(fp))\n-\t\tdie_errno(\"Could not read '%s'\", path);\n-\tfclose(fp);\n-\tstrbuf_release(&sb);\n-}\n-\n static int parse_msg_type(const char *str)\n {\n \tif (!strcmp(str, \"error\"))\ndiff --git a/oidset.c b/oidset.c\nindex fe4eb921df81..584be63e520a 100644\n--- a/oidset.c\n+++ b/oidset.c\n@@ -35,3 +35,38 @@ void oidset_clear(struct oidset *set)\n \tkh_release_oid(&set->set);\n \toidset_init(set, 0);\n }\n+\n+void oidset_parse_file(struct oidset *set, const char *path)\n+{\n+\tFILE *fp;\n+\tstruct strbuf sb = STRBUF_INIT;\n+\tstruct object_id oid;\n+\n+\tfp = fopen(path, \"r\");\n+\tif (!fp)\n+\t\tdie(\"could not open object name list: %s\", path);\n+\twhile (!strbuf_getline(&sb, fp)) {\n+\t\tconst char *p;\n+\t\tconst char *name;\n+\n+\t\t/*\n+\t\t * Allow trailing comments, leading whitespace\n+\t\t * (including before commits), and empty or whitespace\n+\t\t * only lines.\n+\t\t */\n+\t\tname = strchr(sb.buf, '#');\n+\t\tif (name)\n+\t\t\tstrbuf_setlen(&sb, name - sb.buf);\n+\t\tstrbuf_trim(&sb);\n+\t\tif (!sb.len)\n+\t\t\tcontinue;\n+\n+\t\tif (parse_oid_hex(sb.buf, &oid, &p) || *p != '\\0')\n+\t\t\tdie(\"invalid object name: %s\", sb.buf);\n+\t\toidset_insert(set, &oid);\n+\t}\n+\tif (ferror(fp))\n+\t\tdie_errno(\"Could not read '%s'\", path);\n+\tfclose(fp);\n+\tstrbuf_release(&sb);\n+}\ndiff --git a/oidset.h b/oidset.h\nindex 14f18f791fea..2dbca84d798f 100644\n--- a/oidset.h\n+++ b/oidset.h\n@@ -61,6 +61,14 @@ int oidset_remove(struct oidset *set, const struct object_id *oid);\n  */\n void oidset_clear(struct oidset *set);\n \n+/**\n+ * Add the contents of the file 'path' to an initialized oidset.  Each line is\n+ * an unabbreviated object name.  Comments begin with '#', and trailing comments\n+ * are allowed.  Leading whitespace and empty or white-space only lines are\n+ * ignored.\n+ */\n+void oidset_parse_file(struct oidset *set, const char *path);\n+\n struct oidset_iter {\n \tkh_oid_t *set;\n \tkhiter_t iter;\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"377669","messageId":"20190620163820.231316-4-brho@google.com","threadId":"51354","inReplyTo":"20190620163820.231316-1-brho@google.com","subject":"[PATCH v9 3/9] blame: use a helper function in blame_chunk()","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:14Z","receivedAt":"2019-06-20T16:38:39Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"The same code for splitting a blame_entry at a particular line was used\ntwice in blame_chunk(), and I'll use the helper again in an upcoming\npatch.\n\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n blame.c | 44 ++++++++++++++++++++++++++++----------------\n 1 file changed, 28 insertions(+), 16 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex 145eaf2faf9c..5369be9a2233 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -839,6 +839,27 @@ static struct blame_entry *reverse_blame(struct blame_entry *head,\n \treturn tail;\n }\n \n+/*\n+ * Splits a blame entry into two entries at 'len' lines.  The original 'e'\n+ * consists of len lines, i.e. [e->lno, e->lno + len), and the second part,\n+ * which is returned, consists of the remainder: [e->lno + len, e->lno +\n+ * e->num_lines).  The caller needs to sort out the reference counting for the\n+ * new entry's suspect.\n+ */\n+static struct blame_entry *split_blame_at(struct blame_entry *e, int len,\n+\t\t\t\t\t  struct blame_origin *new_suspect)\n+{\n+\tstruct blame_entry *n = xcalloc(1, sizeof(struct blame_entry));\n+\n+\tn->suspect = new_suspect;\n+\tn->lno = e->lno + len;\n+\tn->s_lno = e->s_lno + len;\n+\tn->num_lines = e->num_lines - len;\n+\te->num_lines = len;\n+\te->score = 0;\n+\treturn n;\n+}\n+\n /*\n  * Process one hunk from the patch between the current suspect for\n  * blame_entry e and its parent.  This first blames any unfinished\n@@ -865,14 +886,9 @@ static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n \t\t */\n \t\tif (e->s_lno + e->num_lines > tlno) {\n \t\t\t/* Move second half to a new record */\n-\t\t\tint len = tlno - e->s_lno;\n-\t\t\tstruct blame_entry *n = xcalloc(1, sizeof (struct blame_entry));\n-\t\t\tn->suspect = e->suspect;\n-\t\t\tn->lno = e->lno + len;\n-\t\t\tn->s_lno = e->s_lno + len;\n-\t\t\tn->num_lines = e->num_lines - len;\n-\t\t\te->num_lines = len;\n-\t\t\te->score = 0;\n+\t\t\tstruct blame_entry *n;\n+\n+\t\t\tn = split_blame_at(e, tlno - e->s_lno, e->suspect);\n \t\t\t/* Push new record to diffp */\n \t\t\tn->next = diffp;\n \t\t\tdiffp = n;\n@@ -919,14 +935,10 @@ static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n \t\t\t * Move second half to a new record to be\n \t\t\t * processed by later chunks\n \t\t\t */\n-\t\t\tint len = same - e->s_lno;\n-\t\t\tstruct blame_entry *n = xcalloc(1, sizeof (struct blame_entry));\n-\t\t\tn->suspect = blame_origin_incref(e->suspect);\n-\t\t\tn->lno = e->lno + len;\n-\t\t\tn->s_lno = e->s_lno + len;\n-\t\t\tn->num_lines = e->num_lines - len;\n-\t\t\te->num_lines = len;\n-\t\t\te->score = 0;\n+\t\t\tstruct blame_entry *n;\n+\n+\t\t\tn = split_blame_at(e, same - e->s_lno,\n+\t\t\t\t\t   blame_origin_incref(e->suspect));\n \t\t\t/* Push new record to samep */\n \t\t\tn->next = samep;\n \t\t\tsamep = n;\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"377670","messageId":"20190620163820.231316-5-brho@google.com","threadId":"51354","inReplyTo":"20190620163820.231316-1-brho@google.com","subject":"[PATCH v9 4/9] blame: add the ability to ignore commits and their changes","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:15Z","receivedAt":"2019-06-20T16:38:42Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"Commits that make formatting changes or function renames are often not\ninteresting when blaming a file.  A user may deem such a commit as 'not\ninteresting' and want to ignore and its changes it when assigning blame.\n\nFor example, say a file has the following git history / rev-list:\n\n---O---A---X---B---C---D---Y---E---F\n\nCommits X and Y both touch a particular line, and the other commits do\nnot:\n\nX: \"Take a third parameter\"\n-MyFunc(1, 2);\n+MyFunc(1, 2, 3);\n\nY: \"Remove camelcase\"\n-MyFunc(1, 2, 3);\n+my_func(1, 2, 3);\n\ngit-blame will blame Y for the change.  I'd like to be able to ignore Y:\nboth the existence of the commit as well as any changes it made.  This\ndiffers from -S rev-list, which specifies the list of commits to\nprocess for the blame.  We would still process Y, but just don't let the\nblame 'stick.'\n\nThis patch adds the ability for users to ignore a revision with\n--ignore-rev=rev, which may be repeated.  They can specify a set of\nfiles of full object names of revs, e.g. SHA-1 hashes, one per line.  A\nsingle file may be specified with the blame.ignoreRevFile config option\nor with --ignore-rev-file=file.  Both the config option and the command\nline option may be repeated multiple times.  An empty file name \"\" will\nclear the list of revs from previously processed files.  Config options\nare processed before command line options.\n\nFor a typical use case, projects will maintain the file containing\nrevisions for commits that perform mass reformatting, and their users\nhave the option to ignore all of the commits in that file.\n\nAdditionally, a user can use the --ignore-rev option for one-off\ninvestigation.  To go back to the example above, X was a substantive\nchange to the function, but not the change the user is interested in.\nThe user inspected X, but wanted to find the previous change to that\nline - perhaps a commit that introduced that function call.\n\nTo make this work, we can't simply remove all ignored commits from the\nrev-list.  We need to diff the changes introduced by Y so that we can\nignore them.  We let the blames get passed to Y, just like when\nprocessing normally.  When Y is the target, we make sure that Y does not\n*keep* any blames.  Any changes that Y is responsible for get passed to\nits parent.  Note we make one pass through all of the scapegoats\n(parents) to attempt to pass blame normally; we don't know if we *need*\nto ignore the commit until we've checked all of the parents.\n\nThe blame_entry will get passed up the tree until we find a commit that\nhas a diff chunk that affects those lines.\n\nOne issue is that the ignored commit *did* make some change, and there is\nno general solution to finding the line in the parent commit that\ncorresponds to a given line in the ignored commit.  That makes it hard\nto attribute a particular line within an ignored commit's diff\ncorrectly.\n\nFor example, the parent of an ignored commit has this, say at line 11:\n\ncommit-a 11) #include \"a.h\"\ncommit-b 12) #include \"b.h\"\n\nCommit X, which we will ignore, swaps these lines:\n\ncommit-X 11) #include \"b.h\"\ncommit-X 12) #include \"a.h\"\n\nWe can pass that blame entry to the parent, but line 11 will be\nattributed to commit A, even though \"include b.h\" came from commit B.\nThe blame mechanism will be looking at the parent's view of the file at\nline number 11.\n\nignore_blame_entry() is set up to allow alternative algorithms for\nguessing per-line blames.  Any line that is not attributed to the parent\nwill continue to be blamed on the ignored commit as if that commit was\nnot ignored.  Upcoming patches have the ability to detect these lines\nand mark them in the blame output.\n\nThe existing algorithm is simple: blame each line on the corresponding\nline in the parent's diff chunk.  Any lines beyond that stay with the\ntarget.\n\nFor example, the parent of an ignored commit has this, say at line 11:\n\ncommit-a 11) void new_func_1(void *x, void *y);\ncommit-b 12) void new_func_2(void *x, void *y);\ncommit-c 13) some_line_c\ncommit-d 14) some_line_d\n\nAfter a commit 'X', we have:\n\ncommit-X 11) void new_func_1(void *x,\ncommit-X 12)                 void *y);\ncommit-X 13) void new_func_2(void *x,\ncommit-X 14)                 void *y);\ncommit-c 15) some_line_c\ncommit-d 16) some_line_d\n\nCommit X nets two additionally lines: 13 and 14.  The current\nguess_line_blames() algorithm will not attribute these to the parent,\nwhose diff chunk is only two lines - not four.\n\nWhen we ignore with the current algorithm, we get:\n\ncommit-a 11) void new_func_1(void *x,\ncommit-b 12)                 void *y);\ncommit-X 13) void new_func_2(void *x,\ncommit-X 14)                 void *y);\ncommit-c 15) some_line_c\ncommit-d 16) some_line_d\n\nNote that line 12 was blamed on B, though B was the commit for\nnew_func_2(), not new_func_1().  Even when guess_line_blames() finds a\nline in the parent, it may still be incorrect.\n\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n Documentation/blame-options.txt |  14 +++\n Documentation/config/blame.txt  |   7 ++\n Documentation/git-blame.txt     |   1 +\n blame.c                         | 176 +++++++++++++++++++++++++--\n blame.h                         |   2 +\n builtin/blame.c                 |  38 ++++++\n t/t8013-blame-ignore-revs.sh    | 203 ++++++++++++++++++++++++++++++++\n 7 files changed, 432 insertions(+), 9 deletions(-)\n create mode 100755 t/t8013-blame-ignore-revs.sh\n\ndiff --git a/Documentation/blame-options.txt b/Documentation/blame-options.txt\nindex dc41957afab2..2c2d1ceb5653 100644\n--- a/Documentation/blame-options.txt\n+++ b/Documentation/blame-options.txt\n@@ -110,5 +110,19 @@ commit. And the default value is 40. If there are more than one\n `-C` options given, the <num> argument of the last `-C` will\n take effect.\n \n+--ignore-rev <rev>::\n+\tIgnore changes made by the revision when assigning blame, as if the\n+\tchange never happened.  Lines that were changed or added by an ignored\n+\tcommit will be blamed on the previous commit that changed that line or\n+\tnearby lines.  This option may be specified multiple times to ignore\n+\tmore than one revision.\n+\n+--ignore-revs-file <file>::\n+\tIgnore revisions listed in `file`, which must be in the same format as an\n+\t`fsck.skipList`.  This option may be repeated, and these files will be\n+\tprocessed after any files specified with the `blame.ignoreRevsFile` config\n+\toption.  An empty file name, `\"\"`, will clear the list of revs from\n+\tpreviously processed files.\n+\n -h::\n \tShow help message.\ndiff --git a/Documentation/config/blame.txt b/Documentation/config/blame.txt\nindex 67b5c1d1e02a..4da2788f306d 100644\n--- a/Documentation/config/blame.txt\n+++ b/Documentation/config/blame.txt\n@@ -19,3 +19,10 @@ blame.showEmail::\n blame.showRoot::\n \tDo not treat root commits as boundaries in linkgit:git-blame[1].\n \tThis option defaults to false.\n+\n+blame.ignoreRevsFile::\n+\tIgnore revisions listed in the file, one unabbreviated object name per\n+\tline, in linkgit:git-blame[1].  Whitespace and comments beginning with\n+\t`#` are ignored.  This option may be repeated multiple times.  Empty\n+\tfile names will reset the list of ignored revisions.  This option will\n+\tbe handled before the command line option `--ignore-revs-file`.\ndiff --git a/Documentation/git-blame.txt b/Documentation/git-blame.txt\nindex 16323eb80e31..7e8154199635 100644\n--- a/Documentation/git-blame.txt\n+++ b/Documentation/git-blame.txt\n@@ -10,6 +10,7 @@ SYNOPSIS\n [verse]\n 'git blame' [-c] [-b] [-l] [--root] [-t] [-f] [-n] [-s] [-e] [-p] [-w] [--incremental]\n \t    [-L <range>] [-S <revs-file>] [-M] [-C] [-C] [-C] [--since=<date>]\n+\t    [--ignore-rev <rev>] [--ignore-revs-file <file>]\n \t    [--progress] [--abbrev=<n>] [<rev> | --contents <file> | --reverse <rev>..<rev>]\n \t    [--] <file>\n \ndiff --git a/blame.c b/blame.c\nindex 5369be9a2233..290bc97f31db 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -860,6 +860,103 @@ static struct blame_entry *split_blame_at(struct blame_entry *e, int len,\n \treturn n;\n }\n \n+struct blame_line_tracker {\n+\tint is_parent;\n+\tint s_lno;\n+};\n+\n+static int are_lines_adjacent(struct blame_line_tracker *first,\n+\t\t\t      struct blame_line_tracker *second)\n+{\n+\treturn first->is_parent == second->is_parent &&\n+\t       first->s_lno + 1 == second->s_lno;\n+}\n+\n+/*\n+ * This cheap heuristic assigns lines in the chunk to their relative location in\n+ * the parent's chunk.  Any additional lines are left with the target.\n+ */\n+static void guess_line_blames(struct blame_origin *parent,\n+\t\t\t      struct blame_origin *target,\n+\t\t\t      int tlno, int offset, int same, int parent_len,\n+\t\t\t      struct blame_line_tracker *line_blames)\n+{\n+\tint i, best_idx, target_idx;\n+\tint parent_slno = tlno + offset;\n+\n+\tfor (i = 0; i < same - tlno; i++) {\n+\t\ttarget_idx = tlno + i;\n+\t\tbest_idx = target_idx + offset;\n+\t\tif (best_idx < parent_slno + parent_len) {\n+\t\t\tline_blames[i].is_parent = 1;\n+\t\t\tline_blames[i].s_lno = best_idx;\n+\t\t} else {\n+\t\t\tline_blames[i].is_parent = 0;\n+\t\t\tline_blames[i].s_lno = target_idx;\n+\t\t}\n+\t}\n+}\n+\n+/*\n+ * This decides which parts of a blame entry go to the parent (added to the\n+ * ignoredp list) and which stay with the target (added to the diffp list).  The\n+ * actual decision was made in a separate heuristic function, and those answers\n+ * for the lines in 'e' are in line_blames.  This consumes e, essentially\n+ * putting it on a list.\n+ *\n+ * Note that the blame entries on the ignoredp list are not necessarily sorted\n+ * with respect to the parent's line numbers yet.\n+ */\n+static void ignore_blame_entry(struct blame_entry *e,\n+\t\t\t       struct blame_origin *parent,\n+\t\t\t       struct blame_origin *target,\n+\t\t\t       struct blame_entry **diffp,\n+\t\t\t       struct blame_entry **ignoredp,\n+\t\t\t       struct blame_line_tracker *line_blames)\n+{\n+\tint entry_len, nr_lines, i;\n+\n+\t/*\n+\t * We carve new entries off the front of e.  Each entry comes from a\n+\t * contiguous chunk of lines: adjacent lines from the same origin\n+\t * (either the parent or the target).\n+\t */\n+\tentry_len = 1;\n+\tnr_lines = e->num_lines;\t/* e changes in the loop */\n+\tfor (i = 0; i < nr_lines; i++) {\n+\t\tstruct blame_entry *next = NULL;\n+\n+\t\t/*\n+\t\t * We are often adjacent to the next line - only split the blame\n+\t\t * entry when we have to.\n+\t\t */\n+\t\tif (i + 1 < nr_lines) {\n+\t\t\tif (are_lines_adjacent(&line_blames[i],\n+\t\t\t\t\t       &line_blames[i + 1])) {\n+\t\t\t\tentry_len++;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tnext = split_blame_at(e, entry_len,\n+\t\t\t\t\t      blame_origin_incref(e->suspect));\n+\t\t}\n+\t\tif (line_blames[i].is_parent) {\n+\t\t\tblame_origin_decref(e->suspect);\n+\t\t\te->suspect = blame_origin_incref(parent);\n+\t\t\te->s_lno = line_blames[i - entry_len + 1].s_lno;\n+\t\t\te->next = *ignoredp;\n+\t\t\t*ignoredp = e;\n+\t\t} else {\n+\t\t\t/* e->s_lno is already in the target's address space. */\n+\t\t\te->next = *diffp;\n+\t\t\t*diffp = e;\n+\t\t}\n+\t\tassert(e->num_lines == entry_len);\n+\t\te = next;\n+\t\tentry_len = 1;\n+\t}\n+\tassert(!e);\n+}\n+\n /*\n  * Process one hunk from the patch between the current suspect for\n  * blame_entry e and its parent.  This first blames any unfinished\n@@ -869,13 +966,20 @@ static struct blame_entry *split_blame_at(struct blame_entry *e, int len,\n  * -C options may lead to overlapping/duplicate source line number\n  * ranges, all we can rely on from sorting/merging is the order of the\n  * first suspect line number.\n+ *\n+ * tlno: line number in the target where this chunk begins\n+ * same: line number in the target where this chunk ends\n+ * offset: add to tlno to get the chunk starting point in the parent\n+ * parent_len: number of lines in the parent chunk\n  */\n static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n-\t\t\tint tlno, int offset, int same,\n-\t\t\tstruct blame_origin *parent)\n+\t\t\tint tlno, int offset, int same, int parent_len,\n+\t\t\tstruct blame_origin *parent,\n+\t\t\tstruct blame_origin *target, int ignore_diffs)\n {\n \tstruct blame_entry *e = **srcq;\n-\tstruct blame_entry *samep = NULL, *diffp = NULL;\n+\tstruct blame_entry *samep = NULL, *diffp = NULL, *ignoredp = NULL;\n+\tstruct blame_line_tracker *line_blames = NULL;\n \n \twhile (e && e->s_lno < tlno) {\n \t\tstruct blame_entry *next = e->next;\n@@ -924,6 +1028,14 @@ static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n \t */\n \tsamep = NULL;\n \tdiffp = NULL;\n+\n+\tif (ignore_diffs && same - tlno > 0) {\n+\t\tline_blames = xcalloc(sizeof(struct blame_line_tracker),\n+\t\t\t\t      same - tlno);\n+\t\tguess_line_blames(parent, target, tlno, offset, same,\n+\t\t\t\t  parent_len, line_blames);\n+\t}\n+\n \twhile (e && e->s_lno < same) {\n \t\tstruct blame_entry *next = e->next;\n \n@@ -943,10 +1055,29 @@ static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n \t\t\tn->next = samep;\n \t\t\tsamep = n;\n \t\t}\n-\t\te->next = diffp;\n-\t\tdiffp = e;\n+\t\tif (ignore_diffs) {\n+\t\t\tignore_blame_entry(e, parent, target, &diffp, &ignoredp,\n+\t\t\t\t\t   line_blames + e->s_lno - tlno);\n+\t\t} else {\n+\t\t\te->next = diffp;\n+\t\t\tdiffp = e;\n+\t\t}\n \t\te = next;\n \t}\n+\tfree(line_blames);\n+\tif (ignoredp) {\n+\t\t/*\n+\t\t * Note ignoredp is not sorted yet, and thus neither is dstq.\n+\t\t * That list must be sorted before we queue_blames().  We defer\n+\t\t * sorting until after all diff hunks are processed, so that\n+\t\t * guess_line_blames() can pick *any* line in the parent.  The\n+\t\t * slight drawback is that we end up sorting all blame entries\n+\t\t * passed to the parent, including those that are unrelated to\n+\t\t * changes made by the ignored commit.\n+\t\t */\n+\t\t**dstq = reverse_blame(ignoredp, **dstq);\n+\t\t*dstq = &ignoredp->next;\n+\t}\n \t**srcq = reverse_blame(diffp, reverse_blame(samep, e));\n \t/* Move across elements that are in the unblamable portion */\n \tif (diffp)\n@@ -955,7 +1086,9 @@ static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n \n struct blame_chunk_cb_data {\n \tstruct blame_origin *parent;\n+\tstruct blame_origin *target;\n \tlong offset;\n+\tint ignore_diffs;\n \tstruct blame_entry **dstq;\n \tstruct blame_entry **srcq;\n };\n@@ -968,7 +1101,8 @@ static int blame_chunk_cb(long start_a, long count_a,\n \tif (start_a - start_b != d->offset)\n \t\tdie(\"internal error in blame::blame_chunk_cb\");\n \tblame_chunk(&d->dstq, &d->srcq, start_b, start_a - start_b,\n-\t\t    start_b + count_b, d->parent);\n+\t\t    start_b + count_b, count_a, d->parent, d->target,\n+\t\t    d->ignore_diffs);\n \td->offset = start_a + count_a - (start_b + count_b);\n \treturn 0;\n }\n@@ -980,7 +1114,7 @@ static int blame_chunk_cb(long start_a, long count_a,\n  */\n static void pass_blame_to_parent(struct blame_scoreboard *sb,\n \t\t\t\t struct blame_origin *target,\n-\t\t\t\t struct blame_origin *parent)\n+\t\t\t\t struct blame_origin *parent, int ignore_diffs)\n {\n \tmmfile_t file_p, file_o;\n \tstruct blame_chunk_cb_data d;\n@@ -990,7 +1124,9 @@ static void pass_blame_to_parent(struct blame_scoreboard *sb,\n \t\treturn; /* nothing remains for this target */\n \n \td.parent = parent;\n+\td.target = target;\n \td.offset = 0;\n+\td.ignore_diffs = ignore_diffs;\n \td.dstq = &newdest; d.srcq = &target->suspects;\n \n \tfill_origin_blob(&sb->revs->diffopt, parent, &file_p, &sb->num_read_blob);\n@@ -1002,8 +1138,13 @@ static void pass_blame_to_parent(struct blame_scoreboard *sb,\n \t\t    oid_to_hex(&parent->commit->object.oid),\n \t\t    oid_to_hex(&target->commit->object.oid));\n \t/* The rest are the same as the parent */\n-\tblame_chunk(&d.dstq, &d.srcq, INT_MAX, d.offset, INT_MAX, parent);\n+\tblame_chunk(&d.dstq, &d.srcq, INT_MAX, d.offset, INT_MAX, 0,\n+\t\t    parent, target, 0);\n \t*d.dstq = NULL;\n+\tif (ignore_diffs)\n+\t\tnewdest = llist_mergesort(newdest, get_next_blame,\n+\t\t\t\t\t  set_next_blame,\n+\t\t\t\t\t  compare_blame_suspect);\n \tqueue_blames(sb, parent, newdest);\n \n \treturn;\n@@ -1507,11 +1648,28 @@ static void pass_blame(struct blame_scoreboard *sb, struct blame_origin *origin,\n \t\t\tblame_origin_incref(porigin);\n \t\t\torigin->previous = porigin;\n \t\t}\n-\t\tpass_blame_to_parent(sb, origin, porigin);\n+\t\tpass_blame_to_parent(sb, origin, porigin, 0);\n \t\tif (!origin->suspects)\n \t\t\tgoto finish;\n \t}\n \n+\t/*\n+\t * Pass remaining suspects for ignored commits to their parents.\n+\t */\n+\tif (oidset_contains(&sb->ignore_list, &commit->object.oid)) {\n+\t\tfor (i = 0, sg = first_scapegoat(revs, commit, sb->reverse);\n+\t\t     i < num_sg && sg;\n+\t\t     sg = sg->next, i++) {\n+\t\t\tstruct blame_origin *porigin = sg_origin[i];\n+\n+\t\t\tif (!porigin)\n+\t\t\t\tcontinue;\n+\t\t\tpass_blame_to_parent(sb, origin, porigin, 1);\n+\t\t\tif (!origin->suspects)\n+\t\t\t\tgoto finish;\n+\t\t}\n+\t}\n+\n \t/*\n \t * Optionally find moves in parents' files.\n \t */\ndiff --git a/blame.h b/blame.h\nindex d62f80fa74c4..bd2f23ca36cf 100644\n--- a/blame.h\n+++ b/blame.h\n@@ -117,6 +117,8 @@ struct blame_scoreboard {\n \t/* linked list of blames */\n \tstruct blame_entry *ent;\n \n+\tstruct oidset ignore_list;\n+\n \t/* look-up a line in the final buffer */\n \tint num_lines;\n \tint *lineno;\ndiff --git a/builtin/blame.c b/builtin/blame.c\nindex 21cde57e711e..b8ef1e547cae 100644\n--- a/builtin/blame.c\n+++ b/builtin/blame.c\n@@ -53,6 +53,7 @@ static int no_whole_file_rename;\n static int show_progress;\n static char repeated_meta_color[COLOR_MAXLEN];\n static int coloring_mode;\n+static struct string_list ignore_revs_file_list = STRING_LIST_INIT_NODUP;\n \n static struct date_mode blame_date_mode = { DATE_ISO8601 };\n static size_t blame_date_width;\n@@ -696,6 +697,16 @@ static int git_blame_config(const char *var, const char *value, void *cb)\n \t\tparse_date_format(value, &blame_date_mode);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(var, \"blame.ignorerevsfile\")) {\n+\t\tconst char *str;\n+\t\tint ret;\n+\n+\t\tret = git_config_pathname(&str, var, value);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t\tstring_list_insert(&ignore_revs_file_list, str);\n+\t\treturn 0;\n+\t}\n \tif (!strcmp(var, \"color.blame.repeatedlines\")) {\n \t\tif (color_parse_mem(value, strlen(value), repeated_meta_color))\n \t\t\twarning(_(\"invalid color '%s' in color.blame.repeatedLines\"),\n@@ -775,6 +786,27 @@ static int is_a_rev(const char *name)\n \treturn OBJ_NONE < oid_object_info(the_repository, &oid, NULL);\n }\n \n+static void build_ignorelist(struct blame_scoreboard *sb,\n+\t\t\t     struct string_list *ignore_revs_file_list,\n+\t\t\t     struct string_list *ignore_rev_list)\n+{\n+\tstruct string_list_item *i;\n+\tstruct object_id oid;\n+\n+\toidset_init(&sb->ignore_list, 0);\n+\tfor_each_string_list_item(i, ignore_revs_file_list) {\n+\t\tif (!strcmp(i->string, \"\"))\n+\t\t\toidset_clear(&sb->ignore_list);\n+\t\telse\n+\t\t\toidset_parse_file(&sb->ignore_list, i->string);\n+\t}\n+\tfor_each_string_list_item(i, ignore_rev_list) {\n+\t\tif (get_oid_committish(i->string, &oid))\n+\t\t\tdie(_(\"cannot find revision %s to ignore\"), i->string);\n+\t\toidset_insert(&sb->ignore_list, &oid);\n+\t}\n+}\n+\n int cmd_blame(int argc, const char **argv, const char *prefix)\n {\n \tstruct rev_info revs;\n@@ -786,6 +818,7 @@ int cmd_blame(int argc, const char **argv, const char *prefix)\n \tstruct progress_info pi = { NULL, 0 };\n \n \tstruct string_list range_list = STRING_LIST_INIT_NODUP;\n+\tstruct string_list ignore_rev_list = STRING_LIST_INIT_NODUP;\n \tint output_option = 0, opt = 0;\n \tint show_stats = 0;\n \tconst char *revs_file = NULL;\n@@ -807,6 +840,8 @@ int cmd_blame(int argc, const char **argv, const char *prefix)\n \t\tOPT_BIT('s', NULL, &output_option, N_(\"Suppress author name and timestamp (Default: off)\"), OUTPUT_NO_AUTHOR),\n \t\tOPT_BIT('e', \"show-email\", &output_option, N_(\"Show author email instead of name (Default: off)\"), OUTPUT_SHOW_EMAIL),\n \t\tOPT_BIT('w', NULL, &xdl_opts, N_(\"Ignore whitespace differences\"), XDF_IGNORE_WHITESPACE),\n+\t\tOPT_STRING_LIST(0, \"ignore-rev\", &ignore_rev_list, N_(\"rev\"), N_(\"Ignore <rev> when blaming\")),\n+\t\tOPT_STRING_LIST(0, \"ignore-revs-file\", &ignore_revs_file_list, N_(\"file\"), N_(\"Ignore revisions from <file>\")),\n \t\tOPT_BIT(0, \"color-lines\", &output_option, N_(\"color redundant metadata from previous line differently\"), OUTPUT_COLOR_LINE),\n \t\tOPT_BIT(0, \"color-by-age\", &output_option, N_(\"color lines by age\"), OUTPUT_SHOW_AGE_WITH_COLOR),\n \n@@ -1012,6 +1047,9 @@ int cmd_blame(int argc, const char **argv, const char *prefix)\n \tsb.contents_from = contents_from;\n \tsb.reverse = reverse;\n \tsb.repo = the_repository;\n+\tbuild_ignorelist(&sb, &ignore_revs_file_list, &ignore_rev_list);\n+\tstring_list_clear(&ignore_revs_file_list, 0);\n+\tstring_list_clear(&ignore_rev_list, 0);\n \tsetup_scoreboard(&sb, path, &o);\n \tlno = sb.num_lines;\n \ndiff --git a/t/t8013-blame-ignore-revs.sh b/t/t8013-blame-ignore-revs.sh\nnew file mode 100755\nindex 000000000000..fdb2fa879781\n--- /dev/null\n+++ b/t/t8013-blame-ignore-revs.sh\n@@ -0,0 +1,203 @@\n+#!/bin/sh\n+\n+test_description='ignore revisions when blaming'\n+. ./test-lib.sh\n+\n+# Creates:\n+# \tA--B--X\n+# A added line 1 and B added line 2.  X makes changes to those lines.  Sanity\n+# check that X is blamed for both lines.\n+test_expect_success setup '\n+\ttest_commit A file line1 &&\n+\n+\techo line2 >>file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m B &&\n+\tgit tag B &&\n+\n+\ttest_write_lines line-one line-two >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m X &&\n+\tgit tag X &&\n+\n+\tgit blame --line-porcelain file >blame_raw &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse X >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse X >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# Ignore X, make sure A is blamed for line 1 and B for line 2.\n+test_expect_success ignore_rev_changing_lines '\n+\tgit blame --line-porcelain --ignore-rev X file >blame_raw &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse A >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse B >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# For ignored revs that have added 'unblamable' lines, attribute those to the\n+# ignored commit.\n+# \tA--B--X--Y\n+# Where Y changes lines 1 and 2, and adds lines 3 and 4.  The added lines ought\n+# to have nothing in common with \"line-one\" or \"line-two\", to keep any\n+# heuristics from matching them with any lines in the parent.\n+test_expect_success ignore_rev_adding_unblamable_lines '\n+\ttest_write_lines line-one-change line-two-changed y3 y4 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m Y &&\n+\tgit tag Y &&\n+\n+\tgit rev-parse Y >expect &&\n+\tgit blame --line-porcelain file --ignore-rev Y >blame_raw &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 3\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 4\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# Ignore X and Y, both in separate files.  Lines 1 == A, 2 == B.\n+test_expect_success ignore_revs_from_files '\n+\tgit rev-parse X >ignore_x &&\n+\tgit rev-parse Y >ignore_y &&\n+\tgit blame --line-porcelain file --ignore-revs-file ignore_x --ignore-revs-file ignore_y >blame_raw &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse A >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse B >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# Ignore X from the config option, Y from a file.\n+test_expect_success ignore_revs_from_configs_and_files '\n+\tgit config --add blame.ignoreRevsFile ignore_x &&\n+\tgit blame --line-porcelain file --ignore-revs-file ignore_y >blame_raw &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse A >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse B >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# Override blame.ignoreRevsFile (ignore_x) with an empty string.  X should be\n+# blamed now for lines 1 and 2, since we are no longer ignoring X.\n+test_expect_success override_ignore_revs_file '\n+\tgit blame --line-porcelain file --ignore-revs-file \"\" --ignore-revs-file ignore_y >blame_raw &&\n+\tgit rev-parse X >expect &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\ttest_cmp expect actual\n+\t'\n+test_expect_success bad_files_and_revs '\n+\ttest_must_fail git blame file --ignore-rev NOREV 2>err &&\n+\ttest_i18ngrep \"cannot find revision NOREV to ignore\" err &&\n+\n+\ttest_must_fail git blame file --ignore-revs-file NOFILE 2>err &&\n+\ttest_i18ngrep \"could not open.*: NOFILE\" err &&\n+\n+\techo NOREV >ignore_norev &&\n+\ttest_must_fail git blame file --ignore-revs-file ignore_norev 2>err &&\n+\ttest_i18ngrep \"invalid object name: NOREV\" err\n+\t'\n+# The heuristic called by guess_line_blames() tries to find the size of a\n+# blame_entry 'e' in the parent's address space.  Those calculations need to\n+# check for negative or zero values for when a blame entry is completely outside\n+# the window of the parent's version of a file.\n+#\n+# This happens when one commit adds several lines (commit B below).  A later\n+# commit (C) changes one line in the middle of B's change.  Commit C gets blamed\n+# for its change, and that breaks up B's change into multiple blame entries.\n+# When processing B, one of the blame_entries is outside A's window (which was\n+# zero - it had no lines added on its side of the diff).\n+#\n+# A--B--C, ignore B to test the ignore heuristic's boundary checks.\n+test_expect_success ignored_chunk_negative_parent_size '\n+\trm -rf .git/ &&\n+\tgit init &&\n+\n+\ttest_write_lines L1 L2 L7 L8 L9 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m A &&\n+\tgit tag A &&\n+\n+\ttest_write_lines L1 L2 L3 L4 L5 L6 L7 L8 L9 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m B &&\n+\tgit tag B &&\n+\n+\ttest_write_lines L1 L2 L3 L4 xxx L6 L7 L8 L9 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m C &&\n+\tgit tag C &&\n+\n+\tgit blame file --ignore-rev B >blame_raw\n+\t'\n+\n+# Resetting the repo and creating:\n+#\n+# A--B--M\n+#  \\   /\n+#   C-+\n+#\n+# 'A' creates a file.  B changes line 1, and C changes line 9.  M merges.\n+test_expect_success ignore_merge '\n+\trm -rf .git/ &&\n+\tgit init &&\n+\n+\ttest_write_lines L1 L2 L3 L4 L5 L6 L7 L8 L9 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m A &&\n+\tgit tag A &&\n+\n+\ttest_write_lines BB L2 L3 L4 L5 L6 L7 L8 L9 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m B &&\n+\tgit tag B &&\n+\n+\tgit reset --hard A &&\n+\ttest_write_lines L1 L2 L3 L4 L5 L6 L7 L8 CC >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m C &&\n+\tgit tag C &&\n+\n+\ttest_merge M B &&\n+\tgit blame --line-porcelain file --ignore-rev M >blame_raw &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse B >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep -E \"^[0-9a-f]+ [0-9]+ 9\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse C >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+test_done\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"377671","messageId":"20190620163820.231316-6-brho@google.com","threadId":"51354","inReplyTo":"20190620163820.231316-1-brho@google.com","subject":"[PATCH v9 5/9] blame: add config options for the output of ignored or unblamable lines","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:16Z","receivedAt":"2019-06-20T16:38:45Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"When ignoring commits, the commit that is blamed might not be\nresponsible for the change, due to the inaccuracy of our heuristic.\nUsers might want to know when a particular line has a potentially\ninaccurate blame.\n\nFurthermore, guess_line_blames() may fail to find any parent commit for\na given line touched by an ignored commit.  Those 'unblamable' lines\nremain blamed on an ignored commit.  Users might want to know if a line\nis unblamable so that they do not spend time investigating a commit they\nknow is uninteresting.\n\nThis patch adds two config options to mark these two types of lines in\nthe output of blame.\n\nThe first option can identify ignored lines by specifying\nblame.markIgnoredLines.  When this option is set, each blame line that\nwas blamed on a commit other than the ignored commit is marked with a\n'?'.\n\nFor example:\n\t278b6158d6fdb (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\nappears as:\n\t?278b6158d6fd (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\n\nwhere the '?' is placed before the commit, and the hash has one fewer\ncharacters.\n\nSometimes we are unable to even guess at what ancestor commit touched a\nline.  These lines are 'unblamable.'  The second option,\nblame.markUnblamableLines, will mark the line with '*'.\n\nFor example, say we ignore e5e8d36d04cbe, yet we are unable to blame\nthis line on another commit:\n\te5e8d36d04cbe (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\nappears as:\n\t*e5e8d36d04cb (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\n\nWhen these config options are used together, every line touched by an\nignored commit will be marked with either a '?' or a '*'.\n\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n Documentation/blame-options.txt |  7 +++-\n Documentation/config/blame.txt  |  9 +++++\n blame.c                         | 14 ++++++-\n blame.h                         |  2 +\n builtin/blame.c                 | 18 +++++++++\n t/t8013-blame-ignore-revs.sh    | 71 +++++++++++++++++++++++++++++++++\n 6 files changed, 119 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/blame-options.txt b/Documentation/blame-options.txt\nindex 2c2d1ceb5653..5d122db6e9e6 100644\n--- a/Documentation/blame-options.txt\n+++ b/Documentation/blame-options.txt\n@@ -115,7 +115,12 @@ take effect.\n \tchange never happened.  Lines that were changed or added by an ignored\n \tcommit will be blamed on the previous commit that changed that line or\n \tnearby lines.  This option may be specified multiple times to ignore\n-\tmore than one revision.\n+\tmore than one revision.  If the `blame.markIgnoredLines` config option\n+\tis set, then lines that were changed by an ignored commit and attributed to\n+\tanother commit will be marked with a `?` in the blame output.  If the\n+\t`blame.markUnblamableLines` config option is set, then those lines touched\n+\tby an ignored commit that we could not attribute to another revision are\n+\tmarked with a '*'.\n \n --ignore-revs-file <file>::\n \tIgnore revisions listed in `file`, which must be in the same format as an\ndiff --git a/Documentation/config/blame.txt b/Documentation/config/blame.txt\nindex 4da2788f306d..9468e8599c0c 100644\n--- a/Documentation/config/blame.txt\n+++ b/Documentation/config/blame.txt\n@@ -26,3 +26,12 @@ blame.ignoreRevsFile::\n \t`#` are ignored.  This option may be repeated multiple times.  Empty\n \tfile names will reset the list of ignored revisions.  This option will\n \tbe handled before the command line option `--ignore-revs-file`.\n+\n+blame.markUnblamables::\n+\tMark lines that were changed by an ignored revision that we could not\n+\tattribute to another commit with a '*' in the output of\n+\tlinkgit:git-blame[1].\n+\n+blame.markIgnoredLines::\n+\tMark lines that were changed by an ignored revision that we attributed to\n+\tanother commit with a '?' in the output of linkgit:git-blame[1].\ndiff --git a/blame.c b/blame.c\nindex 290bc97f31db..21ae76603f5c 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -480,7 +480,9 @@ void blame_coalesce(struct blame_scoreboard *sb)\n \n \tfor (ent = sb->ent; ent && (next = ent->next); ent = next) {\n \t\tif (ent->suspect == next->suspect &&\n-\t\t    ent->s_lno + ent->num_lines == next->s_lno) {\n+\t\t    ent->s_lno + ent->num_lines == next->s_lno &&\n+\t\t    ent->ignored == next->ignored &&\n+\t\t    ent->unblamable == next->unblamable) {\n \t\t\tent->num_lines += next->num_lines;\n \t\t\tent->next = next->next;\n \t\t\tblame_origin_decref(next->suspect);\n@@ -730,8 +732,14 @@ static void split_overlap(struct blame_entry *split,\n \t\t\t  struct blame_origin *parent)\n {\n \tint chunk_end_lno;\n+\tint i;\n \tmemset(split, 0, sizeof(struct blame_entry [3]));\n \n+\tfor (i = 0; i < 3; i++) {\n+\t\tsplit[i].ignored = e->ignored;\n+\t\tsplit[i].unblamable = e->unblamable;\n+\t}\n+\n \tif (e->s_lno < tlno) {\n \t\t/* there is a pre-chunk part not blamed on parent */\n \t\tsplit[0].suspect = blame_origin_incref(e->suspect);\n@@ -852,6 +860,8 @@ static struct blame_entry *split_blame_at(struct blame_entry *e, int len,\n \tstruct blame_entry *n = xcalloc(1, sizeof(struct blame_entry));\n \n \tn->suspect = new_suspect;\n+\tn->ignored = e->ignored;\n+\tn->unblamable = e->unblamable;\n \tn->lno = e->lno + len;\n \tn->s_lno = e->s_lno + len;\n \tn->num_lines = e->num_lines - len;\n@@ -940,12 +950,14 @@ static void ignore_blame_entry(struct blame_entry *e,\n \t\t\t\t\t      blame_origin_incref(e->suspect));\n \t\t}\n \t\tif (line_blames[i].is_parent) {\n+\t\t\te->ignored = 1;\n \t\t\tblame_origin_decref(e->suspect);\n \t\t\te->suspect = blame_origin_incref(parent);\n \t\t\te->s_lno = line_blames[i - entry_len + 1].s_lno;\n \t\t\te->next = *ignoredp;\n \t\t\t*ignoredp = e;\n \t\t} else {\n+\t\t\te->unblamable = 1;\n \t\t\t/* e->s_lno is already in the target's address space. */\n \t\t\te->next = *diffp;\n \t\t\t*diffp = e;\ndiff --git a/blame.h b/blame.h\nindex bd2f23ca36cf..2458b68f0e22 100644\n--- a/blame.h\n+++ b/blame.h\n@@ -92,6 +92,8 @@ struct blame_entry {\n \t * scanning the lines over and over.\n \t */\n \tunsigned score;\n+\tint ignored;\n+\tint unblamable;\n };\n \n /*\ndiff --git a/builtin/blame.c b/builtin/blame.c\nindex b8ef1e547cae..ce5b0f283843 100644\n--- a/builtin/blame.c\n+++ b/builtin/blame.c\n@@ -54,6 +54,8 @@ static int show_progress;\n static char repeated_meta_color[COLOR_MAXLEN];\n static int coloring_mode;\n static struct string_list ignore_revs_file_list = STRING_LIST_INIT_NODUP;\n+static int mark_unblamable_lines;\n+static int mark_ignored_lines;\n \n static struct date_mode blame_date_mode = { DATE_ISO8601 };\n static size_t blame_date_width;\n@@ -481,6 +483,14 @@ static void emit_other(struct blame_scoreboard *sb, struct blame_entry *ent, int\n \t\t\t}\n \t\t}\n \n+\t\tif (mark_unblamable_lines && ent->unblamable) {\n+\t\t\tlength--;\n+\t\t\tputchar('*');\n+\t\t}\n+\t\tif (mark_ignored_lines && ent->ignored) {\n+\t\t\tlength--;\n+\t\t\tputchar('?');\n+\t\t}\n \t\tprintf(\"%.*s\", length, hex);\n \t\tif (opt & OUTPUT_ANNOTATE_COMPAT) {\n \t\t\tconst char *name;\n@@ -707,6 +717,14 @@ static int git_blame_config(const char *var, const char *value, void *cb)\n \t\tstring_list_insert(&ignore_revs_file_list, str);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(var, \"blame.markunblamablelines\")) {\n+\t\tmark_unblamable_lines = git_config_bool(var, value);\n+\t\treturn 0;\n+\t}\n+\tif (!strcmp(var, \"blame.markignoredlines\")) {\n+\t\tmark_ignored_lines = git_config_bool(var, value);\n+\t\treturn 0;\n+\t}\n \tif (!strcmp(var, \"color.blame.repeatedlines\")) {\n \t\tif (color_parse_mem(value, strlen(value), repeated_meta_color))\n \t\t\twarning(_(\"invalid color '%s' in color.blame.repeatedLines\"),\ndiff --git a/t/t8013-blame-ignore-revs.sh b/t/t8013-blame-ignore-revs.sh\nindex fdb2fa879781..36dc31eb3913 100755\n--- a/t/t8013-blame-ignore-revs.sh\n+++ b/t/t8013-blame-ignore-revs.sh\n@@ -121,6 +121,77 @@ test_expect_success bad_files_and_revs '\n \ttest_must_fail git blame file --ignore-revs-file ignore_norev 2>err &&\n \ttest_i18ngrep \"invalid object name: NOREV\" err\n \t'\n+\n+# For ignored revs that have added 'unblamable' lines, mark those lines with a\n+# '*'\n+# \tA--B--X--Y\n+# Lines 3 and 4 are from Y and unblamable.  This was set up in\n+# ignore_rev_adding_unblamable_lines.\n+test_expect_success mark_unblamable_lines '\n+\tgit config --add blame.markUnblamableLines true &&\n+\n+\tgit blame --ignore-rev Y file >blame_raw &&\n+\techo \"*\" >expect &&\n+\n+\tsed -n \"3p\" blame_raw | cut -c1 >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tsed -n \"4p\" blame_raw | cut -c1 >actual &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# Commit Z will touch the first two lines.  Y touched all four.\n+# \tA--B--X--Y--Z\n+# The blame output when ignoring Z should be:\n+# ?Y ... 1)\n+# ?Y ... 2)\n+# Y  ... 3)\n+# Y  ... 4)\n+# We're checking only the first character\n+test_expect_success mark_ignored_lines '\n+\tgit config --add blame.markIgnoredLines true &&\n+\n+\ttest_write_lines line-one-Z line-two-Z y3 y4 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m Z &&\n+\tgit tag Z &&\n+\n+\tgit blame --ignore-rev Z file >blame_raw &&\n+\techo \"?\" >expect &&\n+\n+\tsed -n \"1p\" blame_raw | cut -c1 >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tsed -n \"2p\" blame_raw | cut -c1 >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tsed -n \"3p\" blame_raw | cut -c1 >actual &&\n+\t! test_cmp expect actual &&\n+\n+\tsed -n \"4p\" blame_raw | cut -c1 >actual &&\n+\t! test_cmp expect actual\n+\t'\n+\n+# For ignored revs that added 'unblamable' lines and more recent commits changed\n+# the blamable lines, mark the unblamable lines with a\n+# '*'\n+# \tA--B--X--Y--Z\n+# Lines 3 and 4 are from Y and unblamable, as set up in\n+# ignore_rev_adding_unblamable_lines.  Z changed lines 1 and 2.\n+test_expect_success mark_unblamable_lines_intermediate '\n+\tgit config --add blame.markUnblamableLines true &&\n+\n+\tgit blame --ignore-rev Y file >blame_raw 2>stderr &&\n+\techo \"*\" >expect &&\n+\n+\tsed -n \"3p\" blame_raw | cut -c1 >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tsed -n \"4p\" blame_raw | cut -c1 >actual &&\n+\ttest_cmp expect actual\n+\t'\n+\n # The heuristic called by guess_line_blames() tries to find the size of a\n # blame_entry 'e' in the parent's address space.  Those calculations need to\n # check for negative or zero values for when a blame entry is completely outside\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"377672","messageId":"20190620163820.231316-7-brho@google.com","threadId":"51354","inReplyTo":"20190620163820.231316-1-brho@google.com","subject":"[PATCH v9 6/9] blame: optionally track line fingerprints during fill_blame_origin()","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:17Z","receivedAt":"2019-06-20T16:38:48Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"fill_blame_origin() is a convenient place to store data that we will use\nthroughout the lifetime of a blame_origin.  Some heuristics for\nignoring commits during a blame session can make use of this storage.\nIn particular, we will calculate a fingerprint for each line of a file\nfor blame_origins involved in an ignored commit.\n\nIn this commit, we only calculate the line_starts, reusing the existing\ncode from the scoreboard's line_starts.  In an upcoming commit, we will\nactually compute the fingerprints.\n\nThis feature will be used when we attempt to pass blame entries to\nparents when we \"ignore\" a commit.  Most uses of fill_blame_origin()\nwill not require this feature, hence the flag parameter.  Multiple calls\nto fill_blame_origin() are idempotent, and any of them can request the\ncreation of the fingerprints structure.\n\nSuggested-by: Michael Platings <michael@platin.gs>\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n blame.c | 90 ++++++++++++++++++++++++++++++++++++++-------------------\n blame.h |  2 ++\n 2 files changed, 62 insertions(+), 30 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex 21ae76603f5c..49698a306e5a 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -311,12 +311,58 @@ static int diff_hunks(mmfile_t *file_a, mmfile_t *file_b,\n \treturn xdi_diff(file_a, file_b, &xpp, &xecfg, &ecb);\n }\n \n+static const char *get_next_line(const char *start, const char *end)\n+{\n+\tconst char *nl = memchr(start, '\\n', end - start);\n+\n+\treturn nl ? nl + 1 : end;\n+}\n+\n+static int find_line_starts(int **line_starts, const char *buf,\n+\t\t\t    unsigned long len)\n+{\n+\tconst char *end = buf + len;\n+\tconst char *p;\n+\tint *lineno;\n+\tint num = 0;\n+\n+\tfor (p = buf; p < end; p = get_next_line(p, end))\n+\t\tnum++;\n+\n+\tALLOC_ARRAY(*line_starts, num + 1);\n+\tlineno = *line_starts;\n+\n+\tfor (p = buf; p < end; p = get_next_line(p, end))\n+\t\t*lineno++ = p - buf;\n+\n+\t*lineno = len;\n+\n+\treturn num;\n+}\n+\n+static void fill_origin_fingerprints(struct blame_origin *o, mmfile_t *file)\n+{\n+\tint *line_starts;\n+\n+\tif (o->fingerprints)\n+\t\treturn;\n+\to->num_lines = find_line_starts(&line_starts, o->file.ptr,\n+\t\t\t\t\to->file.size);\n+\t/* TODO: Will fill in fingerprints in a future commit */\n+\tfree(line_starts);\n+}\n+\n+static void drop_origin_fingerprints(struct blame_origin *o)\n+{\n+}\n+\n /*\n  * Given an origin, prepare mmfile_t structure to be used by the\n  * diff machinery\n  */\n static void fill_origin_blob(struct diff_options *opt,\n-\t\t\t     struct blame_origin *o, mmfile_t *file, int *num_read_blob)\n+\t\t\t     struct blame_origin *o, mmfile_t *file,\n+\t\t\t     int *num_read_blob, int fill_fingerprints)\n {\n \tif (!o->file.ptr) {\n \t\tenum object_type type;\n@@ -340,11 +386,14 @@ static void fill_origin_blob(struct diff_options *opt,\n \t}\n \telse\n \t\t*file = o->file;\n+\tif (fill_fingerprints)\n+\t\tfill_origin_fingerprints(o, file);\n }\n \n static void drop_origin_blob(struct blame_origin *o)\n {\n \tFREE_AND_NULL(o->file.ptr);\n+\tdrop_origin_fingerprints(o);\n }\n \n /*\n@@ -1141,8 +1190,10 @@ static void pass_blame_to_parent(struct blame_scoreboard *sb,\n \td.ignore_diffs = ignore_diffs;\n \td.dstq = &newdest; d.srcq = &target->suspects;\n \n-\tfill_origin_blob(&sb->revs->diffopt, parent, &file_p, &sb->num_read_blob);\n-\tfill_origin_blob(&sb->revs->diffopt, target, &file_o, &sb->num_read_blob);\n+\tfill_origin_blob(&sb->revs->diffopt, parent, &file_p,\n+\t\t\t &sb->num_read_blob, ignore_diffs);\n+\tfill_origin_blob(&sb->revs->diffopt, target, &file_o,\n+\t\t\t &sb->num_read_blob, ignore_diffs);\n \tsb->num_get_patch++;\n \n \tif (diff_hunks(&file_p, &file_o, blame_chunk_cb, &d, sb->xdl_opts))\n@@ -1353,7 +1404,8 @@ static void find_move_in_parent(struct blame_scoreboard *sb,\n \tif (!unblamed)\n \t\treturn; /* nothing remains for this target */\n \n-\tfill_origin_blob(&sb->revs->diffopt, parent, &file_p, &sb->num_read_blob);\n+\tfill_origin_blob(&sb->revs->diffopt, parent, &file_p,\n+\t\t\t &sb->num_read_blob, 0);\n \tif (!file_p.ptr)\n \t\treturn;\n \n@@ -1482,7 +1534,8 @@ static void find_copy_in_parent(struct blame_scoreboard *sb,\n \t\t\tnorigin = get_origin(parent, p->one->path);\n \t\t\toidcpy(&norigin->blob_oid, &p->one->oid);\n \t\t\tnorigin->mode = p->one->mode;\n-\t\t\tfill_origin_blob(&sb->revs->diffopt, norigin, &file_p, &sb->num_read_blob);\n+\t\t\tfill_origin_blob(&sb->revs->diffopt, norigin, &file_p,\n+\t\t\t\t\t &sb->num_read_blob, 0);\n \t\t\tif (!file_p.ptr)\n \t\t\t\tcontinue;\n \n@@ -1822,37 +1875,14 @@ void assign_blame(struct blame_scoreboard *sb, int opt)\n \t}\n }\n \n-static const char *get_next_line(const char *start, const char *end)\n-{\n-\tconst char *nl = memchr(start, '\\n', end - start);\n-\treturn nl ? nl + 1 : end;\n-}\n-\n /*\n  * To allow quick access to the contents of nth line in the\n  * final image, prepare an index in the scoreboard.\n  */\n static int prepare_lines(struct blame_scoreboard *sb)\n {\n-\tconst char *buf = sb->final_buf;\n-\tunsigned long len = sb->final_buf_size;\n-\tconst char *end = buf + len;\n-\tconst char *p;\n-\tint *lineno;\n-\tint num = 0;\n-\n-\tfor (p = buf; p < end; p = get_next_line(p, end))\n-\t\tnum++;\n-\n-\tALLOC_ARRAY(sb->lineno, num + 1);\n-\tlineno = sb->lineno;\n-\n-\tfor (p = buf; p < end; p = get_next_line(p, end))\n-\t\t*lineno++ = p - buf;\n-\n-\t*lineno = len;\n-\n-\tsb->num_lines = num;\n+\tsb->num_lines = find_line_starts(&sb->lineno, sb->final_buf,\n+\t\t\t\t\t sb->final_buf_size);\n \treturn sb->num_lines;\n }\n \ndiff --git a/blame.h b/blame.h\nindex 2458b68f0e22..4a9e1270b036 100644\n--- a/blame.h\n+++ b/blame.h\n@@ -51,6 +51,8 @@ struct blame_origin {\n \t */\n \tstruct blame_entry *suspects;\n \tmmfile_t file;\n+\tint num_lines;\n+\tvoid *fingerprints;\n \tstruct object_id blob_oid;\n \tunsigned short mode;\n \t/* guilty gets set when shipping any suspects to the final\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"377673","messageId":"20190620163820.231316-8-brho@google.com","threadId":"51354","inReplyTo":"20190620163820.231316-1-brho@google.com","subject":"[PATCH v9 7/9] blame: add a fingerprint heuristic to match ignored lines","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:18Z","receivedAt":"2019-06-20T16:38:54Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"From: Michael Platings <michael@platin.gs>\n\nThis algorithm will replace the heuristic used to identify lines from\nignored commits with one that finds likely candidate lines in the\nparent's version of the file.  The actual replacement occurs in an\nupcoming commit.\n\nThe old heuristic simply assigned lines in the target to the same line\nnumber (plus offset) in the parent. The new function uses a\nfingerprinting algorithm to detect similarity between lines.\n\nThe new heuristic is designed to accurately match changes made\nmechanically by formatting tools such as clang-format and clang-tidy.\nThese tools make changes such as breaking up lines to fit within a\ncharacter limit or changing identifiers to fit with a naming convention.\nThe heuristic is not intended to match more extensive refactoring\nchanges and may give misleading results in such cases.\n\nIn most cases formatting tools preserve line ordering, so the heuristic\nis optimised for such cases. (Some types of changes do reorder lines\ne.g. sorting keep the line content identical, the git blame -M option\ncan already be used to address this). The reason that it is advantageous\nto rely on ordering is due to source code repeating the same character\nsequences often e.g. declaring an identifier on one line and using that\nidentifier on several subsequent lines.  This means that lines can look\nvery similar to each other which presents a problem when doing fuzzy\nmatching. Relying on ordering gives us extra clues to point towards the\ntrue match.\n\nThe heuristic operates on a single diff chunk change at a time. It\ncreates a “fingerprint” for each line on each side of the change.\nFingerprints are described in detail in the comment for `struct\nfingerprint`, but essentially are a multiset of the character pairs in a\nline. The heuristic first identifies the line in the target entry whose\nfingerprint is most clearly matched to a line fingerprint in the parent\nentry. Where fingerprints match identically, the position of the lines\nis used as a tie-break. The heuristic locks in the best match, and\nsubtracts the fingerprint of the line in the target entry from the\nfingerprint of the line in the parent entry to prevent other lines being\nmatched on the same parts of that line. It then repeats the process\nrecursively on the section of the chunk before the match, and then the\nsection of the chunk after the match.\n\nHere's an example of the difference the fingerprinting makes. Consider\na file with two commits:\n\n        commit-a 1) void func_1(void *x, void *y);\n        commit-b 2) void func_2(void *x, void *y);\n\nAfter a commit 'X', we have:\n\n        commit-X 1) void func_1(void *x,\n        commit-X 2)             void *y);\n        commit-X 3) void func_2(void *x,\n        commit-X 4)             void *y);\n\nWhen we blame-ignored with the old algorithm, we get:\n\n        commit-a 1) void func_1(void *x,\n        commit-b 2)             void *y);\n        commit-X 3) void func_2(void *x,\n        commit-X 4)             void *y);\n\nWhere commit-b is blamed for 2 instead of 3.  With the fingerprint\nalgorithm, we get:\n\n        commit-a 1) void func_1(void *x,\n        commit-a 2)             void *y);\n        commit-b 3) void func_2(void *x,\n        commit-b 4)             void *y);\n\nNote line 2 could be matched with either commit-a or commit-b as it is\nequally similar to both lines, but is matched with commit-a because its\nposition as a fraction of the new line range is more similar to commit-a\nas a fraction of the old line range. Line 4 is also equally similar to\nboth lines, but as it appears after line 3 which will be matched first\nit cannot be matched with an earlier line.\n\nFor many more examples, see t/t8014-blame-ignore-fuzzy.sh which contains\nexample parent and target files and the line numbers in the parent that\nmust be matched.\n\nSigned-off-by: Michael Platings <michael@platin.gs>\n---\n blame.c                       | 642 ++++++++++++++++++++++++++++++++++\n t/t8014-blame-ignore-fuzzy.sh | 440 +++++++++++++++++++++++\n 2 files changed, 1082 insertions(+)\n create mode 100755 t/t8014-blame-ignore-fuzzy.sh\n\ndiff --git a/blame.c b/blame.c\nindex 49698a306e5a..103838546e07 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -340,6 +340,648 @@ static int find_line_starts(int **line_starts, const char *buf,\n \treturn num;\n }\n \n+struct fingerprint_entry;\n+\n+/* A fingerprint is intended to loosely represent a string, such that two\n+ * fingerprints can be quickly compared to give an indication of the similarity\n+ * of the strings that they represent.\n+ *\n+ * A fingerprint is represented as a multiset of the lower-cased byte pairs in\n+ * the string that it represents. Whitespace is added at each end of the\n+ * string. Whitespace pairs are ignored. Whitespace is converted to '\\0'.\n+ * For example, the string \"Darth   Radar\" will be converted to the following\n+ * fingerprint:\n+ * {\"\\0d\", \"da\", \"da\", \"ar\", \"ar\", \"rt\", \"th\", \"h\\0\", \"\\0r\", \"ra\", \"ad\", \"r\\0\"}\n+ *\n+ * The similarity between two fingerprints is the size of the intersection of\n+ * their multisets, including repeated elements. See fingerprint_similarity for\n+ * examples.\n+ *\n+ * For ease of implementation, the fingerprint is implemented as a map\n+ * of byte pairs to the count of that byte pair in the string, instead of\n+ * allowing repeated elements in a set.\n+ */\n+struct fingerprint {\n+\tstruct hashmap map;\n+\t/* As we know the maximum number of entries in advance, it's\n+\t * convenient to store the entries in a single array instead of having\n+\t * the hashmap manage the memory.\n+\t */\n+\tstruct fingerprint_entry *entries;\n+};\n+\n+/* A byte pair in a fingerprint. Stores the number of times the byte pair\n+ * occurs in the string that the fingerprint represents.\n+ */\n+struct fingerprint_entry {\n+\t/* The hashmap entry - the hash represents the byte pair in its\n+\t * entirety so we don't need to store the byte pair separately.\n+\t */\n+\tstruct hashmap_entry entry;\n+\t/* The number of times the byte pair occurs in the string that the\n+\t * fingerprint represents.\n+\t */\n+\tint count;\n+};\n+\n+/* See `struct fingerprint` for an explanation of what a fingerprint is.\n+ * \\param result the fingerprint of the string is stored here. This must be\n+ * \t\t freed later using free_fingerprint.\n+ * \\param line_begin the start of the string\n+ * \\param line_end the end of the string\n+ */\n+static void get_fingerprint(struct fingerprint *result,\n+\t\t\t    const char *line_begin,\n+\t\t\t    const char *line_end)\n+{\n+\tunsigned int hash, c0 = 0, c1;\n+\tconst char *p;\n+\tint max_map_entry_count = 1 + line_end - line_begin;\n+\tstruct fingerprint_entry *entry = xcalloc(max_map_entry_count,\n+\t\tsizeof(struct fingerprint_entry));\n+\tstruct fingerprint_entry *found_entry;\n+\n+\thashmap_init(&result->map, NULL, NULL, max_map_entry_count);\n+\tresult->entries = entry;\n+\tfor (p = line_begin; p <= line_end; ++p, c0 = c1) {\n+\t\t/* Always terminate the string with whitespace.\n+\t\t * Normalise whitespace to 0, and normalise letters to\n+\t\t * lower case. This won't work for multibyte characters but at\n+\t\t * worst will match some unrelated characters.\n+\t\t */\n+\t\tif ((p == line_end) || isspace(*p))\n+\t\t\tc1 = 0;\n+\t\telse\n+\t\t\tc1 = tolower(*p);\n+\t\thash = c0 | (c1 << 8);\n+\t\t/* Ignore whitespace pairs */\n+\t\tif (hash == 0)\n+\t\t\tcontinue;\n+\t\thashmap_entry_init(entry, hash);\n+\n+\t\tfound_entry = hashmap_get(&result->map, entry, NULL);\n+\t\tif (found_entry) {\n+\t\t\tfound_entry->count += 1;\n+\t\t} else {\n+\t\t\tentry->count = 1;\n+\t\t\thashmap_add(&result->map, entry);\n+\t\t\t++entry;\n+\t\t}\n+\t}\n+}\n+\n+static void free_fingerprint(struct fingerprint *f)\n+{\n+\thashmap_free(&f->map, 0);\n+\tfree(f->entries);\n+}\n+\n+/* Calculates the similarity between two fingerprints as the size of the\n+ * intersection of their multisets, including repeated elements. See\n+ * `struct fingerprint` for an explanation of the fingerprint representation.\n+ * The similarity between \"cat mat\" and \"father rather\" is 2 because \"at\" is\n+ * present twice in both strings while the similarity between \"tim\" and \"mit\"\n+ * is 0.\n+ */\n+static int fingerprint_similarity(struct fingerprint *a, struct fingerprint *b)\n+{\n+\tint intersection = 0;\n+\tstruct hashmap_iter iter;\n+\tconst struct fingerprint_entry *entry_a, *entry_b;\n+\n+\thashmap_iter_init(&b->map, &iter);\n+\n+\twhile ((entry_b = hashmap_iter_next(&iter))) {\n+\t\tif ((entry_a = hashmap_get(&a->map, entry_b, NULL))) {\n+\t\t\tintersection += entry_a->count < entry_b->count ?\n+\t\t\t\t\tentry_a->count : entry_b->count;\n+\t\t}\n+\t}\n+\treturn intersection;\n+}\n+\n+/* Subtracts byte-pair elements in B from A, modifying A in place.\n+ */\n+static void fingerprint_subtract(struct fingerprint *a, struct fingerprint *b)\n+{\n+\tstruct hashmap_iter iter;\n+\tstruct fingerprint_entry *entry_a;\n+\tconst struct fingerprint_entry *entry_b;\n+\n+\thashmap_iter_init(&b->map, &iter);\n+\n+\twhile ((entry_b = hashmap_iter_next(&iter))) {\n+\t\tif ((entry_a = hashmap_get(&a->map, entry_b, NULL))) {\n+\t\t\tif (entry_a->count <= entry_b->count)\n+\t\t\t\thashmap_remove(&a->map, entry_b, NULL);\n+\t\t\telse\n+\t\t\t\tentry_a->count -= entry_b->count;\n+\t\t}\n+\t}\n+}\n+\n+/* Calculate fingerprints for a series of lines.\n+ * Puts the fingerprints in the fingerprints array, which must have been\n+ * preallocated to allow storing line_count elements.\n+ */\n+static void get_line_fingerprints(struct fingerprint *fingerprints,\n+\t\t\t\t  const char *content, const int *line_starts,\n+\t\t\t\t  long first_line, long line_count)\n+{\n+\tint i;\n+\tconst char *linestart, *lineend;\n+\n+\tline_starts += first_line;\n+\tfor (i = 0; i < line_count; ++i) {\n+\t\tlinestart = content + line_starts[i];\n+\t\tlineend = content + line_starts[i + 1];\n+\t\tget_fingerprint(fingerprints + i, linestart, lineend);\n+\t}\n+}\n+\n+static void free_line_fingerprints(struct fingerprint *fingerprints,\n+\t\t\t\t   int nr_fingerprints)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i < nr_fingerprints; i++)\n+\t\tfree_fingerprint(&fingerprints[i]);\n+}\n+\n+/* This contains the data necessary to linearly map a line number in one half\n+ * of a diff chunk to the line in the other half of the diff chunk that is\n+ * closest in terms of its position as a fraction of the length of the chunk.\n+ */\n+struct line_number_mapping {\n+\tint destination_start, destination_length,\n+\t\tsource_start, source_length;\n+};\n+\n+/* Given a line number in one range, offset and scale it to map it onto the\n+ * other range.\n+ * Essentially this mapping is a simple linear equation but the calculation is\n+ * more complicated to allow performing it with integer operations.\n+ * Another complication is that if a line could map onto many lines in the\n+ * destination range then we want to choose the line at the center of those\n+ * possibilities.\n+ * Example: if the chunk is 2 lines long in A and 10 lines long in B then the\n+ * first 5 lines in B will map onto the first line in the A chunk, while the\n+ * last 5 lines will all map onto the second line in the A chunk.\n+ * Example: if the chunk is 10 lines long in A and 2 lines long in B then line\n+ * 0 in B will map onto line 2 in A, and line 1 in B will map onto line 7 in A.\n+ */\n+static int map_line_number(int line_number,\n+\tconst struct line_number_mapping *mapping)\n+{\n+\treturn ((line_number - mapping->source_start) * 2 + 1) *\n+\t       mapping->destination_length /\n+\t       (mapping->source_length * 2) +\n+\t       mapping->destination_start;\n+}\n+\n+/* Get a pointer to the element storing the similarity between a line in A\n+ * and a line in B.\n+ *\n+ * The similarities are stored in a 2-dimensional array. Each \"row\" in the\n+ * array contains the similarities for a line in B. The similarities stored in\n+ * a row are the similarities between the line in B and the nearby lines in A.\n+ * To keep the length of each row the same, it is padded out with values of -1\n+ * where the search range extends beyond the lines in A.\n+ * For example, if max_search_distance_a is 2 and the two sides of a diff chunk\n+ * look like this:\n+ * a | m\n+ * b | n\n+ * c | o\n+ * d | p\n+ * e | q\n+ * Then the similarity array will contain:\n+ * [-1, -1, am, bm, cm,\n+ *  -1, an, bn, cn, dn,\n+ *  ao, bo, co, do, eo,\n+ *  bp, cp, dp, ep, -1,\n+ *  cq, dq, eq, -1, -1]\n+ * Where similarities are denoted either by -1 for invalid, or the\n+ * concatenation of the two lines in the diff being compared.\n+ *\n+ * \\param similarities array of similarities between lines in A and B\n+ * \\param line_a the index of the line in A, in the same frame of reference as\n+ *\tclosest_line_a.\n+ * \\param local_line_b the index of the line in B, relative to the first line\n+ *\t\t       in B that similarities represents.\n+ * \\param closest_line_a the index of the line in A that is deemed to be\n+ *\t\t\t closest to local_line_b. This must be in the same\n+ *\t\t\t frame of reference as line_a. This value defines\n+ *\t\t\t where similarities is centered for the line in B.\n+ * \\param max_search_distance_a maximum distance in lines from the closest line\n+ * \t\t\t\tin A for other lines in A for which\n+ * \t\t\t\tsimilarities may be calculated.\n+ */\n+static int *get_similarity(int *similarities,\n+\t\t\t   int line_a, int local_line_b,\n+\t\t\t   int closest_line_a, int max_search_distance_a)\n+{\n+\tassert(abs(line_a - closest_line_a) <=\n+\t       max_search_distance_a);\n+\treturn similarities + line_a - closest_line_a +\n+\t       max_search_distance_a +\n+\t       local_line_b * (max_search_distance_a * 2 + 1);\n+}\n+\n+#define CERTAIN_NOTHING_MATCHES -2\n+#define CERTAINTY_NOT_CALCULATED -1\n+\n+/* Given a line in B, first calculate its similarities with nearby lines in A\n+ * if not already calculated, then identify the most similar and second most\n+ * similar lines. The \"certainty\" is calculated based on those two\n+ * similarities.\n+ *\n+ * \\param start_a the index of the first line of the chunk in A\n+ * \\param length_a the length in lines of the chunk in A\n+ * \\param local_line_b the index of the line in B, relative to the first line\n+ * \t\t       in the chunk.\n+ * \\param fingerprints_a array of fingerprints for the chunk in A\n+ * \\param fingerprints_b array of fingerprints for the chunk in B\n+ * \\param similarities 2-dimensional array of similarities between lines in A\n+ * \t\t       and B. See get_similarity() for more details.\n+ * \\param certainties array of values indicating how strongly a line in B is\n+ * \t\t      matched with some line in A.\n+ * \\param second_best_result array of absolute indices in A for the second\n+ * \t\t\t     closest match of a line in B.\n+ * \\param result array of absolute indices in A for the closest match of a line\n+ * \t\t in B.\n+ * \\param max_search_distance_a maximum distance in lines from the closest line\n+ * \t\t\t\tin A for other lines in A for which\n+ * \t\t\t\tsimilarities may be calculated.\n+ * \\param map_line_number_in_b_to_a parameter to map_line_number().\n+ */\n+static void find_best_line_matches(\n+\tint start_a,\n+\tint length_a,\n+\tint start_b,\n+\tint local_line_b,\n+\tstruct fingerprint *fingerprints_a,\n+\tstruct fingerprint *fingerprints_b,\n+\tint *similarities,\n+\tint *certainties,\n+\tint *second_best_result,\n+\tint *result,\n+\tconst int max_search_distance_a,\n+\tconst struct line_number_mapping *map_line_number_in_b_to_a)\n+{\n+\n+\tint i, search_start, search_end, closest_local_line_a, *similarity,\n+\t\tbest_similarity = 0, second_best_similarity = 0,\n+\t\tbest_similarity_index = 0, second_best_similarity_index = 0;\n+\n+\t/* certainty has already been calculated so no need to redo the work */\n+\tif (certainties[local_line_b] != CERTAINTY_NOT_CALCULATED)\n+\t\treturn;\n+\n+\tclosest_local_line_a = map_line_number(\n+\t\tlocal_line_b + start_b, map_line_number_in_b_to_a) - start_a;\n+\n+\tsearch_start = closest_local_line_a - max_search_distance_a;\n+\tif (search_start < 0)\n+\t\tsearch_start = 0;\n+\n+\tsearch_end = closest_local_line_a + max_search_distance_a + 1;\n+\tif (search_end > length_a)\n+\t\tsearch_end = length_a;\n+\n+\tfor (i = search_start; i < search_end; ++i) {\n+\t\tsimilarity = get_similarity(similarities,\n+\t\t\t\t\t    i, local_line_b,\n+\t\t\t\t\t    closest_local_line_a,\n+\t\t\t\t\t    max_search_distance_a);\n+\t\tif (*similarity == -1) {\n+\t\t\t/* This value will never exceed 10 but assert just in\n+\t\t\t * case\n+\t\t\t */\n+\t\t\tassert(abs(i - closest_local_line_a) < 1000);\n+\t\t\t/* scale the similarity by (1000 - distance from\n+\t\t\t * closest line) to act as a tie break between lines\n+\t\t\t * that otherwise are equally similar.\n+\t\t\t */\n+\t\t\t*similarity = fingerprint_similarity(\n+\t\t\t\tfingerprints_b + local_line_b,\n+\t\t\t\tfingerprints_a + i) *\n+\t\t\t\t(1000 - abs(i - closest_local_line_a));\n+\t\t}\n+\t\tif (*similarity > best_similarity) {\n+\t\t\tsecond_best_similarity = best_similarity;\n+\t\t\tsecond_best_similarity_index = best_similarity_index;\n+\t\t\tbest_similarity = *similarity;\n+\t\t\tbest_similarity_index = i;\n+\t\t} else if (*similarity > second_best_similarity) {\n+\t\t\tsecond_best_similarity = *similarity;\n+\t\t\tsecond_best_similarity_index = i;\n+\t\t}\n+\t}\n+\n+\tif (best_similarity == 0) {\n+\t\t/* this line definitely doesn't match with anything. Mark it\n+\t\t * with this special value so it doesn't get invalidated and\n+\t\t * won't be recalculated.\n+\t\t */\n+\t\tcertainties[local_line_b] = CERTAIN_NOTHING_MATCHES;\n+\t\tresult[local_line_b] = -1;\n+\t} else {\n+\t\t/* Calculate the certainty with which this line matches.\n+\t\t * If the line matches well with two lines then that reduces\n+\t\t * the certainty. However we still want to prioritise matching\n+\t\t * a line that matches very well with two lines over matching a\n+\t\t * line that matches poorly with one line, hence doubling\n+\t\t * best_similarity.\n+\t\t * This means that if we have\n+\t\t * line X that matches only one line with a score of 3,\n+\t\t * line Y that matches two lines equally with a score of 5,\n+\t\t * and line Z that matches only one line with a score or 2,\n+\t\t * then the lines in order of certainty are X, Y, Z.\n+\t\t */\n+\t\tcertainties[local_line_b] = best_similarity * 2 -\n+\t\t\tsecond_best_similarity;\n+\n+\t\t/* We keep both the best and second best results to allow us to\n+\t\t * check at a later stage of the matching process whether the\n+\t\t * result needs to be invalidated.\n+\t\t */\n+\t\tresult[local_line_b] = start_a + best_similarity_index;\n+\t\tsecond_best_result[local_line_b] =\n+\t\t\tstart_a + second_best_similarity_index;\n+\t}\n+}\n+\n+/*\n+ * This finds the line that we can match with the most confidence, and\n+ * uses it as a partition. It then calls itself on the lines on either side of\n+ * that partition. In this way we avoid lines appearing out of order, and\n+ * retain a sensible line ordering.\n+ * \\param start_a index of the first line in A with which lines in B may be\n+ * \t\t  compared.\n+ * \\param start_b index of the first line in B for which matching should be\n+ * \t\t  done.\n+ * \\param length_a number of lines in A with which lines in B may be compared.\n+ * \\param length_b number of lines in B for which matching should be done.\n+ * \\param fingerprints_a mutable array of fingerprints in A. The first element\n+ * \t\t\t corresponds to the line at start_a.\n+ * \\param fingerprints_b array of fingerprints in B. The first element\n+ * \t\t\t corresponds to the line at start_b.\n+ * \\param similarities 2-dimensional array of similarities between lines in A\n+ * \t\t       and B. See get_similarity() for more details.\n+ * \\param certainties array of values indicating how strongly a line in B is\n+ * \t\t      matched with some line in A.\n+ * \\param second_best_result array of absolute indices in A for the second\n+ * \t\t\t     closest match of a line in B.\n+ * \\param result array of absolute indices in A for the closest match of a line\n+ * \t\t in B.\n+ * \\param max_search_distance_a maximum distance in lines from the closest line\n+ * \t\t\t      in A for other lines in A for which\n+ * \t\t\t      similarities may be calculated.\n+ * \\param max_search_distance_b an upper bound on the greatest possible\n+ * \t\t\t      distance between lines in B such that they will\n+ *                              both be compared with the same line in A\n+ * \t\t\t      according to max_search_distance_a.\n+ * \\param map_line_number_in_b_to_a parameter to map_line_number().\n+ */\n+static void fuzzy_find_matching_lines_recurse(\n+\tint start_a, int start_b,\n+\tint length_a, int length_b,\n+\tstruct fingerprint *fingerprints_a,\n+\tstruct fingerprint *fingerprints_b,\n+\tint *similarities,\n+\tint *certainties,\n+\tint *second_best_result,\n+\tint *result,\n+\tint max_search_distance_a,\n+\tint max_search_distance_b,\n+\tconst struct line_number_mapping *map_line_number_in_b_to_a)\n+{\n+\tint i, invalidate_min, invalidate_max, offset_b,\n+\t\tsecond_half_start_a, second_half_start_b,\n+\t\tsecond_half_length_a, second_half_length_b,\n+\t\tmost_certain_line_a, most_certain_local_line_b = -1,\n+\t\tmost_certain_line_certainty = -1,\n+\t\tclosest_local_line_a;\n+\n+\tfor (i = 0; i < length_b; ++i) {\n+\t\tfind_best_line_matches(start_a,\n+\t\t\t\t       length_a,\n+\t\t\t\t       start_b,\n+\t\t\t\t       i,\n+\t\t\t\t       fingerprints_a,\n+\t\t\t\t       fingerprints_b,\n+\t\t\t\t       similarities,\n+\t\t\t\t       certainties,\n+\t\t\t\t       second_best_result,\n+\t\t\t\t       result,\n+\t\t\t\t       max_search_distance_a,\n+\t\t\t\t       map_line_number_in_b_to_a);\n+\n+\t\tif (certainties[i] > most_certain_line_certainty) {\n+\t\t\tmost_certain_line_certainty = certainties[i];\n+\t\t\tmost_certain_local_line_b = i;\n+\t\t}\n+\t}\n+\n+\t/* No matches. */\n+\tif (most_certain_local_line_b == -1)\n+\t\treturn;\n+\n+\tmost_certain_line_a = result[most_certain_local_line_b];\n+\n+\t/*\n+\t * Subtract the most certain line's fingerprint in B from the matched\n+\t * fingerprint in A. This means that other lines in B can't also match\n+\t * the same parts of the line in A.\n+\t */\n+\tfingerprint_subtract(fingerprints_a + most_certain_line_a - start_a,\n+\t\t\t     fingerprints_b + most_certain_local_line_b);\n+\n+\t/* Invalidate results that may be affected by the choice of most\n+\t * certain line.\n+\t */\n+\tinvalidate_min = most_certain_local_line_b - max_search_distance_b;\n+\tinvalidate_max = most_certain_local_line_b + max_search_distance_b + 1;\n+\tif (invalidate_min < 0)\n+\t\tinvalidate_min = 0;\n+\tif (invalidate_max > length_b)\n+\t\tinvalidate_max = length_b;\n+\n+\t/* As the fingerprint in A has changed, discard previously calculated\n+\t * similarity values with that fingerprint.\n+\t */\n+\tfor (i = invalidate_min; i < invalidate_max; ++i) {\n+\t\tclosest_local_line_a = map_line_number(\n+\t\t\ti + start_b, map_line_number_in_b_to_a) - start_a;\n+\n+\t\t/* Check that the lines in A and B are close enough that there\n+\t\t * is a similarity value for them.\n+\t\t */\n+\t\tif (abs(most_certain_line_a - start_a - closest_local_line_a) >\n+\t\t\tmax_search_distance_a) {\n+\t\t\tcontinue;\n+\t\t}\n+\n+\t\t*get_similarity(similarities, most_certain_line_a - start_a,\n+\t\t\t\ti, closest_local_line_a,\n+\t\t\t\tmax_search_distance_a) = -1;\n+\t}\n+\n+\t/* More invalidating of results that may be affected by the choice of\n+\t * most certain line.\n+\t * Discard the matches for lines in B that are currently matched with a\n+\t * line in A such that their ordering contradicts the ordering imposed\n+\t * by the choice of most certain line.\n+\t */\n+\tfor (i = most_certain_local_line_b - 1; i >= invalidate_min; --i) {\n+\t\t/* In this loop we discard results for lines in B that are\n+\t\t * before most-certain-line-B but are matched with a line in A\n+\t\t * that is after most-certain-line-A.\n+\t\t */\n+\t\tif (certainties[i] >= 0 &&\n+\t\t    (result[i] >= most_certain_line_a ||\n+\t\t     second_best_result[i] >= most_certain_line_a)) {\n+\t\t\tcertainties[i] = CERTAINTY_NOT_CALCULATED;\n+\t\t}\n+\t}\n+\tfor (i = most_certain_local_line_b + 1; i < invalidate_max; ++i) {\n+\t\t/* In this loop we discard results for lines in B that are\n+\t\t * after most-certain-line-B but are matched with a line in A\n+\t\t * that is before most-certain-line-A.\n+\t\t */\n+\t\tif (certainties[i] >= 0 &&\n+\t\t    (result[i] <= most_certain_line_a ||\n+\t\t     second_best_result[i] <= most_certain_line_a)) {\n+\t\t\tcertainties[i] = CERTAINTY_NOT_CALCULATED;\n+\t\t}\n+\t}\n+\n+\t/* Repeat the matching process for lines before the most certain line.\n+\t */\n+\tif (most_certain_local_line_b > 0) {\n+\t\tfuzzy_find_matching_lines_recurse(\n+\t\t\tstart_a, start_b,\n+\t\t\tmost_certain_line_a + 1 - start_a,\n+\t\t\tmost_certain_local_line_b,\n+\t\t\tfingerprints_a, fingerprints_b, similarities,\n+\t\t\tcertainties, second_best_result, result,\n+\t\t\tmax_search_distance_a,\n+\t\t\tmax_search_distance_b,\n+\t\t\tmap_line_number_in_b_to_a);\n+\t}\n+\t/* Repeat the matching process for lines after the most certain line.\n+\t */\n+\tif (most_certain_local_line_b + 1 < length_b) {\n+\t\tsecond_half_start_a = most_certain_line_a;\n+\t\toffset_b = most_certain_local_line_b + 1;\n+\t\tsecond_half_start_b = start_b + offset_b;\n+\t\tsecond_half_length_a =\n+\t\t\tlength_a + start_a - second_half_start_a;\n+\t\tsecond_half_length_b =\n+\t\t\tlength_b + start_b - second_half_start_b;\n+\t\tfuzzy_find_matching_lines_recurse(\n+\t\t\tsecond_half_start_a, second_half_start_b,\n+\t\t\tsecond_half_length_a, second_half_length_b,\n+\t\t\tfingerprints_a + second_half_start_a - start_a,\n+\t\t\tfingerprints_b + offset_b,\n+\t\t\tsimilarities +\n+\t\t\t\toffset_b * (max_search_distance_a * 2 + 1),\n+\t\t\tcertainties + offset_b,\n+\t\t\tsecond_best_result + offset_b, result + offset_b,\n+\t\t\tmax_search_distance_a,\n+\t\t\tmax_search_distance_b,\n+\t\t\tmap_line_number_in_b_to_a);\n+\t}\n+}\n+\n+/* Find the lines in the parent line range that most closely match the lines in\n+ * the target line range. This is accomplished by matching fingerprints in each\n+ * blame_origin, and choosing the best matches that preserve the line ordering.\n+ * See struct fingerprint for details of fingerprint matching, and\n+ * fuzzy_find_matching_lines_recurse for details of preserving line ordering.\n+ *\n+ * The performance is believed to be O(n log n) in the typical case and O(n^2)\n+ * in a pathological case, where n is the number of lines in the target range.\n+ */\n+static int *fuzzy_find_matching_lines(struct blame_origin *parent,\n+\t\t\t\t      struct blame_origin *target,\n+\t\t\t\t      int tlno, int parent_slno, int same,\n+\t\t\t\t      int parent_len)\n+{\n+\t/* We use the terminology \"A\" for the left hand side of the diff AKA\n+\t * parent, and \"B\" for the right hand side of the diff AKA target. */\n+\tint start_a = parent_slno;\n+\tint length_a = parent_len;\n+\tint start_b = tlno;\n+\tint length_b = same - tlno;\n+\n+\tstruct line_number_mapping map_line_number_in_b_to_a = {\n+\t\tstart_a, length_a, start_b, length_b\n+\t};\n+\n+\tstruct fingerprint *fingerprints_a = parent->fingerprints;\n+\tstruct fingerprint *fingerprints_b = target->fingerprints;\n+\n+\tint i, *result, *second_best_result,\n+\t\t*certainties, *similarities, similarity_count;\n+\n+\t/*\n+\t * max_search_distance_a means that given a line in B, compare it to\n+\t * the line in A that is closest to its position, and the lines in A\n+\t * that are no greater than max_search_distance_a lines away from the\n+\t * closest line in A.\n+\t *\n+\t * max_search_distance_b is an upper bound on the greatest possible\n+\t * distance between lines in B such that they will both be compared\n+\t * with the same line in A according to max_search_distance_a.\n+\t */\n+\tint max_search_distance_a = 10, max_search_distance_b;\n+\n+\tif (length_a <= 0)\n+\t\treturn NULL;\n+\n+\tif (max_search_distance_a >= length_a)\n+\t\tmax_search_distance_a = length_a ? length_a - 1 : 0;\n+\n+\tmax_search_distance_b = ((2 * max_search_distance_a + 1) * length_b\n+\t\t\t\t - 1) / length_a;\n+\n+\tresult = xcalloc(sizeof(int), length_b);\n+\tsecond_best_result = xcalloc(sizeof(int), length_b);\n+\tcertainties = xcalloc(sizeof(int), length_b);\n+\n+\t/* See get_similarity() for details of similarities. */\n+\tsimilarity_count = length_b * (max_search_distance_a * 2 + 1);\n+\tsimilarities = xcalloc(sizeof(int), similarity_count);\n+\n+\tfor (i = 0; i < length_b; ++i) {\n+\t\tresult[i] = -1;\n+\t\tsecond_best_result[i] = -1;\n+\t\tcertainties[i] = CERTAINTY_NOT_CALCULATED;\n+\t}\n+\n+\tfor (i = 0; i < similarity_count; ++i)\n+\t\tsimilarities[i] = -1;\n+\n+\tfuzzy_find_matching_lines_recurse(start_a, start_b,\n+\t\t\t\t\t  length_a, length_b,\n+\t\t\t\t\t  fingerprints_a + start_a,\n+\t\t\t\t\t  fingerprints_b + start_b,\n+\t\t\t\t\t  similarities,\n+\t\t\t\t\t  certainties,\n+\t\t\t\t\t  second_best_result,\n+\t\t\t\t\t  result,\n+\t\t\t\t\t  max_search_distance_a,\n+\t\t\t\t\t  max_search_distance_b,\n+\t\t\t\t\t  &map_line_number_in_b_to_a);\n+\n+\tfree(similarities);\n+\tfree(certainties);\n+\tfree(second_best_result);\n+\n+\treturn result;\n+}\n+\n static void fill_origin_fingerprints(struct blame_origin *o, mmfile_t *file)\n {\n \tint *line_starts;\ndiff --git a/t/t8014-blame-ignore-fuzzy.sh b/t/t8014-blame-ignore-fuzzy.sh\nnew file mode 100755\nindex 000000000000..844396615271\n--- /dev/null\n+++ b/t/t8014-blame-ignore-fuzzy.sh\n@@ -0,0 +1,440 @@\n+#!/bin/sh\n+\n+test_description='git blame ignore fuzzy heuristic'\n+. ./test-lib.sh\n+\n+# short circuit until blame has the fuzzy capabilities\n+test_done\n+\n+pick_author='s/^[0-9a-f^]* *(\\([^ ]*\\) .*/\\1/'\n+\n+# Each test is composed of 4 variables:\n+# titleN - the test name\n+# aN - the initial content\n+# bN - the final content\n+# expectedN - the line numbers from aN that we expect git blame\n+#             on bN to identify, or \"Final\" if bN itself should\n+#             be identified as the origin of that line.\n+\n+# We start at test 2 because setup will show as test 1\n+title2=\"Regression test for partially overlapping search ranges\"\n+cat <<EOF >a2\n+1\n+2\n+3\n+abcdef\n+5\n+6\n+7\n+ijkl\n+9\n+10\n+11\n+pqrs\n+13\n+14\n+15\n+wxyz\n+17\n+18\n+19\n+EOF\n+cat <<EOF >b2\n+abcde\n+ijk\n+pqr\n+wxy\n+EOF\n+cat <<EOF >expected2\n+4\n+8\n+12\n+16\n+EOF\n+\n+title3=\"Combine 3 lines into 2\"\n+cat <<EOF >a3\n+if ((maxgrow==0) ||\n+\t( single_line_field && (field->dcols < maxgrow)) ||\n+\t(!single_line_field && (field->drows < maxgrow)))\n+EOF\n+cat <<EOF >b3\n+if ((maxgrow == 0) || (single_line_field && (field->dcols < maxgrow)) ||\n+\t(!single_line_field && (field->drows < maxgrow))) {\n+EOF\n+cat <<EOF >expected3\n+2\n+3\n+EOF\n+\n+title4=\"Add curly brackets\"\n+cat <<EOF >a4\n+\tif (rows) *rows = field->rows;\n+\tif (cols) *cols = field->cols;\n+\tif (frow) *frow = field->frow;\n+\tif (fcol) *fcol = field->fcol;\n+EOF\n+cat <<EOF >b4\n+\tif (rows) {\n+\t\t*rows = field->rows;\n+\t}\n+\tif (cols) {\n+\t\t*cols = field->cols;\n+\t}\n+\tif (frow) {\n+\t\t*frow = field->frow;\n+\t}\n+\tif (fcol) {\n+\t\t*fcol = field->fcol;\n+\t}\n+EOF\n+cat <<EOF >expected4\n+1\n+1\n+Final\n+2\n+2\n+Final\n+3\n+3\n+Final\n+4\n+4\n+Final\n+EOF\n+\n+\n+title5=\"Combine many lines and change case\"\n+cat <<EOF >a5\n+for(row=0,pBuffer=field->buf;\n+\trow<height;\n+\trow++,pBuffer+=width )\n+{\n+\tif ((len = (int)( After_End_Of_Data( pBuffer, width ) - pBuffer )) > 0)\n+\t{\n+\t\twmove( win, row, 0 );\n+\t\twaddnstr( win, pBuffer, len );\n+EOF\n+cat <<EOF >b5\n+for (Row = 0, PBuffer = field->buf; Row < Height; Row++, PBuffer += Width) {\n+\tif ((Len = (int)(afterEndOfData(PBuffer, Width) - PBuffer)) > 0) {\n+\t\twmove(win, Row, 0);\n+\t\twaddnstr(win, PBuffer, Len);\n+EOF\n+cat <<EOF >expected5\n+1\n+5\n+7\n+8\n+EOF\n+\n+title6=\"Rename and combine lines\"\n+cat <<EOF >a6\n+bool need_visual_update = ((form != (FORM *)0)      &&\n+\t(form->status & _POSTED) &&\n+\t(form->current==field));\n+\n+if (need_visual_update)\n+\tSynchronize_Buffer(form);\n+\n+if (single_line_field)\n+{\n+\tgrowth = field->cols * amount;\n+\tif (field->maxgrow)\n+\t\tgrowth = Minimum(field->maxgrow - field->dcols,growth);\n+\tfield->dcols += growth;\n+\tif (field->dcols == field->maxgrow)\n+EOF\n+cat <<EOF >b6\n+bool NeedVisualUpdate = ((Form != (FORM *)0) && (Form->status & _POSTED) &&\n+\t(Form->current == field));\n+\n+if (NeedVisualUpdate) {\n+\tsynchronizeBuffer(Form);\n+}\n+\n+if (SingleLineField) {\n+\tGrowth = field->cols * amount;\n+\tif (field->maxgrow) {\n+\t\tGrowth = Minimum(field->maxgrow - field->dcols, Growth);\n+\t}\n+\tfield->dcols += Growth;\n+\tif (field->dcols == field->maxgrow) {\n+EOF\n+cat <<EOF >expected6\n+1\n+3\n+4\n+5\n+6\n+Final\n+7\n+8\n+10\n+11\n+12\n+Final\n+13\n+14\n+EOF\n+\n+# Both lines match identically so position must be used to tie-break.\n+title7=\"Same line twice\"\n+cat <<EOF >a7\n+abc\n+abc\n+EOF\n+cat <<EOF >b7\n+abcd\n+abcd\n+EOF\n+cat <<EOF >expected7\n+1\n+2\n+EOF\n+\n+title8=\"Enforce line order\"\n+cat <<EOF >a8\n+abcdef\n+ghijkl\n+ab\n+EOF\n+cat <<EOF >b8\n+ghijk\n+abcd\n+EOF\n+cat <<EOF >expected8\n+2\n+3\n+EOF\n+\n+title9=\"Expand lines and rename variables\"\n+cat <<EOF >a9\n+int myFunction(int ArgumentOne, Thing *ArgTwo, Blah XuglyBug) {\n+\tSquiggle FabulousResult = squargle(ArgumentOne, *ArgTwo,\n+\t\tXuglyBug) + EwwwGlobalWithAReallyLongNameYepTooLong;\n+\treturn FabulousResult * 42;\n+}\n+EOF\n+cat <<EOF >b9\n+int myFunction(int argument_one, Thing *arg_asdfgh,\n+\tBlah xugly_bug) {\n+\tSquiggle fabulous_result = squargle(argument_one,\n+\t\t*arg_asdfgh, xugly_bug)\n+\t\t+ g_ewww_global_with_a_really_long_name_yep_too_long;\n+\treturn fabulous_result * 42;\n+}\n+EOF\n+cat <<EOF >expected9\n+1\n+1\n+2\n+3\n+3\n+4\n+5\n+EOF\n+\n+title10=\"Two close matches versus one less close match\"\n+cat <<EOF >a10\n+abcdef\n+abcdef\n+ghijkl\n+EOF\n+cat <<EOF >b10\n+gh\n+abcdefx\n+EOF\n+cat <<EOF >expected10\n+Final\n+2\n+EOF\n+\n+# The first line of b matches best with the last line of a, but the overall\n+# match is better if we match it with the the first line of a.\n+title11=\"Piggy in the middle\"\n+cat <<EOF >a11\n+abcdefg\n+ijklmn\n+abcdefgh\n+EOF\n+cat <<EOF >b11\n+abcdefghx\n+ijklm\n+EOF\n+cat <<EOF >expected11\n+1\n+2\n+EOF\n+\n+title12=\"No trailing newline\"\n+printf \"abc\\ndef\" >a12\n+printf \"abx\\nstu\" >b12\n+cat <<EOF >expected12\n+1\n+Final\n+EOF\n+\n+title13=\"Reorder includes\"\n+cat <<EOF >a13\n+#include \"c.h\"\n+#include \"b.h\"\n+#include \"a.h\"\n+#include \"e.h\"\n+#include \"d.h\"\n+EOF\n+cat <<EOF >b13\n+#include \"a.h\"\n+#include \"b.h\"\n+#include \"c.h\"\n+#include \"d.h\"\n+#include \"e.h\"\n+EOF\n+cat <<EOF >expected13\n+3\n+2\n+1\n+5\n+4\n+EOF\n+\n+last_test=13\n+\n+test_expect_success setup '\n+\t{ for i in $(test_seq 2 $last_test)\n+\tdo\n+\t\t# Append each line in a separate commit to make it easy to\n+\t\t# check which original line the blame output relates to.\n+\n+\t\tline_count=0 &&\n+\t\t{ while IFS= read line\n+\t\tdo\n+\t\t\tline_count=$((line_count+1)) &&\n+\t\t\techo \"$line\" >>\"$i\" &&\n+\t\t\tgit add \"$i\" &&\n+\t\t\ttest_tick &&\n+\t\t\tGIT_AUTHOR_NAME=\"$line_count\" git commit -m \"$line_count\"\n+\t\tdone } <\"a$i\"\n+\tdone } &&\n+\n+\t{ for i in $(test_seq 2 $last_test)\n+\tdo\n+\t\t# Overwrite the files with the final content.\n+\t\tcp b$i $i &&\n+\t\tgit add $i\n+\tdone } &&\n+\ttest_tick &&\n+\n+\t# Commit the final content all at once so it can all be\n+\t# referred to with the same commit ID.\n+\tGIT_AUTHOR_NAME=Final git commit -m Final &&\n+\n+\tIGNOREME=$(git rev-parse HEAD)\n+'\n+\n+for i in $(test_seq 2 $last_test); do\n+\teval title=\"\\$title$i\"\n+\ttest_expect_success \"$title\" \\\n+\t\"git blame -M9 --ignore-rev $IGNOREME $i >output &&\n+\tsed -e \\\"$pick_author\\\" output >actual &&\n+\ttest_cmp expected$i actual\"\n+done\n+\n+# This invoked a null pointer dereference when the chunk callback was called\n+# with a zero length parent chunk and there were no more suspects.\n+test_expect_success 'Diff chunks with no suspects' '\n+\ttest_write_lines xy1 A B C xy1 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=1 git commit -m 1 &&\n+\n+\ttest_write_lines xy2 A B xy2 C xy2 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=2 git commit -m 2 &&\n+\tREV_2=$(git rev-parse HEAD) &&\n+\n+\ttest_write_lines xy3 A >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=3 git commit -m 3 &&\n+\tREV_3=$(git rev-parse HEAD) &&\n+\n+\ttest_write_lines 1 1 >expected &&\n+\n+\tgit blame --ignore-rev $REV_2 --ignore-rev $REV_3 file >output &&\n+\tsed -e \"$pick_author\" output >actual &&\n+\n+\ttest_cmp expected actual\n+\t'\n+\n+test_expect_success 'position matching' '\n+\ttest_write_lines abc def >file2 &&\n+\tgit add file2 &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=1 git commit -m 1 &&\n+\n+\ttest_write_lines abc def abc def >file2 &&\n+\tgit add file2 &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=2 git commit -m 2 &&\n+\n+\ttest_write_lines abcx defx abcx defx >file2 &&\n+\tgit add file2 &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=3 git commit -m 3 &&\n+\tREV_3=$(git rev-parse HEAD) &&\n+\n+\ttest_write_lines abcy defy abcx defx >file2 &&\n+\tgit add file2 &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=4 git commit -m 4 &&\n+\tREV_4=$(git rev-parse HEAD) &&\n+\n+\ttest_write_lines 1 1 2 2 >expected &&\n+\n+\tgit blame --ignore-rev $REV_3 --ignore-rev $REV_4 file2 >output &&\n+\tsed -e \"$pick_author\" output >actual &&\n+\n+\ttest_cmp expected actual\n+\t'\n+\n+# This fails if each blame entry is processed independently instead of\n+# processing each diff change in full.\n+test_expect_success 'preserve order' '\n+\ttest_write_lines bcde >file3 &&\n+\tgit add file3 &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=1 git commit -m 1 &&\n+\n+\ttest_write_lines bcde fghij >file3 &&\n+\tgit add file3 &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=2 git commit -m 2 &&\n+\n+\ttest_write_lines bcde fghij abcd >file3 &&\n+\tgit add file3 &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=3 git commit -m 3 &&\n+\n+\ttest_write_lines abcdx fghijx bcdex >file3 &&\n+\tgit add file3 &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=4 git commit -m 4 &&\n+\tREV_4=$(git rev-parse HEAD) &&\n+\n+\ttest_write_lines abcdx fghijy bcdex >file3 &&\n+\tgit add file3 &&\n+\ttest_tick &&\n+\tGIT_AUTHOR_NAME=5 git commit -m 5 &&\n+\tREV_5=$(git rev-parse HEAD) &&\n+\n+\ttest_write_lines 1 2 3 >expected &&\n+\n+\tgit blame --ignore-rev $REV_4 --ignore-rev $REV_5 file3 >output &&\n+\tsed -e \"$pick_author\" output >actual &&\n+\n+\ttest_cmp expected actual\n+\t'\n+\n+test_done\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"377674","messageId":"20190620163820.231316-9-brho@google.com","threadId":"51354","inReplyTo":"20190620163820.231316-1-brho@google.com","subject":"[PATCH v9 8/9] blame: use the fingerprint heuristic to match ignored lines","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:19Z","receivedAt":"2019-06-20T16:38:55Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"This commit integrates the fuzzy fingerprint heuristic into\nguess_line_blames().\n\nWe actually make two passes.  The first pass uses the fuzzy algorithm to\nfind a match within the current diff chunk.  If that fails, the second\npass searches the entire parent file for the best match.\n\nFor an example of scanning the entire parent for a match, consider:\n\n\tcommit-a 30) #include <sys/header_a.h>\n\tcommit-b 31) #include <header_b.h>\n\tcommit-c 32) #include <header_c.h>\n\nThen commit X alphabetizes them:\n\n\tcommit-X 30) #include <header_b.h>\n\tcommit-X 31) #include <header_c.h>\n\tcommit-X 32) #include <sys/header_a.h>\n\nIf we just check the parent's chunk (i.e. the first pass), we'd get:\n\n\tcommit-b 30) #include <header_b.h>\n\tcommit-c 31) #include <header_c.h>\n\tcommit-X 32) #include <sys/header_a.h>\n\nThat's because commit X actually consists of two chunks: one chunk is\nremoving sys/header_a.h, then some context, and the second chunk is\nadding sys/header_a.h.\n\nIf we scan the entire parent file, we get:\n\n\tcommit-b 30) #include <header_b.h>\n\tcommit-c 31) #include <header_c.h>\n\tcommit-a 32) #include <sys/header_a.h>\n\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n blame.c                       | 60 ++++++++++++++++++++++++++++++++---\n t/t8014-blame-ignore-fuzzy.sh |  3 --\n 2 files changed, 55 insertions(+), 8 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex 103838546e07..f81ec9a8cf80 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -990,12 +990,19 @@ static void fill_origin_fingerprints(struct blame_origin *o, mmfile_t *file)\n \t\treturn;\n \to->num_lines = find_line_starts(&line_starts, o->file.ptr,\n \t\t\t\t\to->file.size);\n-\t/* TODO: Will fill in fingerprints in a future commit */\n+\to->fingerprints = xcalloc(sizeof(struct fingerprint), o->num_lines);\n+\tget_line_fingerprints(o->fingerprints, o->file.ptr, line_starts,\n+\t\t\t      0, o->num_lines);\n \tfree(line_starts);\n }\n \n static void drop_origin_fingerprints(struct blame_origin *o)\n {\n+\tif (o->fingerprints) {\n+\t\tfree_line_fingerprints(o->fingerprints, o->num_lines);\n+\t\to->num_lines = 0;\n+\t\tFREE_AND_NULL(o->fingerprints);\n+\t}\n }\n \n /*\n@@ -1573,9 +1580,34 @@ static int are_lines_adjacent(struct blame_line_tracker *first,\n \t       first->s_lno + 1 == second->s_lno;\n }\n \n+static int scan_parent_range(struct fingerprint *p_fps,\n+\t\t\t     struct fingerprint *t_fps, int t_idx,\n+\t\t\t     int from, int nr_lines)\n+{\n+\tint sim, p_idx;\n+\t#define FINGERPRINT_FILE_THRESHOLD\t10\n+\tint best_sim_val = FINGERPRINT_FILE_THRESHOLD;\n+\tint best_sim_idx = -1;\n+\n+\tfor (p_idx = from; p_idx < from + nr_lines; p_idx++) {\n+\t\tsim = fingerprint_similarity(&t_fps[t_idx], &p_fps[p_idx]);\n+\t\tif (sim < best_sim_val)\n+\t\t\tcontinue;\n+\t\t/* Break ties with the closest-to-target line number */\n+\t\tif (sim == best_sim_val && best_sim_idx != -1 &&\n+\t\t    abs(best_sim_idx - t_idx) < abs(p_idx - t_idx))\n+\t\t\tcontinue;\n+\t\tbest_sim_val = sim;\n+\t\tbest_sim_idx = p_idx;\n+\t}\n+\treturn best_sim_idx;\n+}\n+\n /*\n- * This cheap heuristic assigns lines in the chunk to their relative location in\n- * the parent's chunk.  Any additional lines are left with the target.\n+ * The first pass checks the blame entry (from the target) against the parent's\n+ * diff chunk.  If that fails for a line, the second pass tries to match that\n+ * line to any part of parent file.  That catches cases where a change was\n+ * broken into two chunks by 'context.'\n  */\n static void guess_line_blames(struct blame_origin *parent,\n \t\t\t      struct blame_origin *target,\n@@ -1584,11 +1616,22 @@ static void guess_line_blames(struct blame_origin *parent,\n {\n \tint i, best_idx, target_idx;\n \tint parent_slno = tlno + offset;\n+\tint *fuzzy_matches;\n \n+\tfuzzy_matches = fuzzy_find_matching_lines(parent, target,\n+\t\t\t\t\t\t  tlno, parent_slno, same,\n+\t\t\t\t\t\t  parent_len);\n \tfor (i = 0; i < same - tlno; i++) {\n \t\ttarget_idx = tlno + i;\n-\t\tbest_idx = target_idx + offset;\n-\t\tif (best_idx < parent_slno + parent_len) {\n+\t\tif (fuzzy_matches && fuzzy_matches[i] >= 0) {\n+\t\t\tbest_idx = fuzzy_matches[i];\n+\t\t} else {\n+\t\t\tbest_idx = scan_parent_range(parent->fingerprints,\n+\t\t\t\t\t\t     target->fingerprints,\n+\t\t\t\t\t\t     target_idx, 0,\n+\t\t\t\t\t\t     parent->num_lines);\n+\t\t}\n+\t\tif (best_idx >= 0) {\n \t\t\tline_blames[i].is_parent = 1;\n \t\t\tline_blames[i].s_lno = best_idx;\n \t\t} else {\n@@ -1596,6 +1639,7 @@ static void guess_line_blames(struct blame_origin *parent,\n \t\t\tline_blames[i].s_lno = target_idx;\n \t\t}\n \t}\n+\tfree(fuzzy_matches);\n }\n \n /*\n@@ -2372,6 +2416,12 @@ static void pass_blame(struct blame_scoreboard *sb, struct blame_origin *origin,\n \t\t\tif (!porigin)\n \t\t\t\tcontinue;\n \t\t\tpass_blame_to_parent(sb, origin, porigin, 1);\n+\t\t\t/*\n+\t\t\t * Preemptively drop porigin so we can refresh the\n+\t\t\t * fingerprints if we use the parent again, which can\n+\t\t\t * occur if you ignore back-to-back commits.\n+\t\t\t */\n+\t\t\tdrop_origin_blob(porigin);\n \t\t\tif (!origin->suspects)\n \t\t\t\tgoto finish;\n \t\t}\ndiff --git a/t/t8014-blame-ignore-fuzzy.sh b/t/t8014-blame-ignore-fuzzy.sh\nindex 844396615271..6f1a94caef22 100755\n--- a/t/t8014-blame-ignore-fuzzy.sh\n+++ b/t/t8014-blame-ignore-fuzzy.sh\n@@ -3,9 +3,6 @@\n test_description='git blame ignore fuzzy heuristic'\n . ./test-lib.sh\n \n-# short circuit until blame has the fuzzy capabilities\n-test_done\n-\n pick_author='s/^[0-9a-f^]* *(\\([^ ]*\\) .*/\\1/'\n \n # Each test is composed of 4 variables:\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"377675","messageId":"20190620163820.231316-10-brho@google.com","threadId":"51354","inReplyTo":"20190620163820.231316-1-brho@google.com","subject":"[PATCH v9 9/9] blame: add a test to cover blame_coalesce()","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-06-20T16:38:20Z","receivedAt":"2019-06-20T16:40:04Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"Signed-off-by: Barret Rhoden <brho@google.com>\n---\n t/t8003-blame-corner-cases.sh | 36 +++++++++++++++++++++++++++++++++++\n 1 file changed, 36 insertions(+)\n\ndiff --git a/t/t8003-blame-corner-cases.sh b/t/t8003-blame-corner-cases.sh\nindex c92a47b6d5b1..1c5fb1d1f8c9 100755\n--- a/t/t8003-blame-corner-cases.sh\n+++ b/t/t8003-blame-corner-cases.sh\n@@ -275,4 +275,40 @@ test_expect_success 'blame file with CRLF core.autocrlf=true' '\n \tgrep \"A U Thor\" actual\n '\n \n+# Tests the splitting and merging of blame entries in blame_coalesce().\n+# The output of blame is the same, regardless of whether blame_coalesce() runs\n+# or not, so we'd likely only notice a problem if blame crashes or assigned\n+# blame to the \"splitting\" commit ('SPLIT' below).\n+test_expect_success 'blame coalesce' '\n+\tcat >giraffe <<-\\EOF &&\n+\tABC\n+\tDEF\n+\tEOF\n+\tgit add giraffe &&\n+\tgit commit -m \"original file\" &&\n+\toid=$(git rev-parse HEAD) &&\n+\n+\tcat >giraffe <<-\\EOF &&\n+\tABC\n+\tSPLIT\n+\tDEF\n+\tEOF\n+\tgit add giraffe &&\n+\tgit commit -m \"interior SPLIT line\" &&\n+\n+\tcat >giraffe <<-\\EOF &&\n+\tABC\n+\tDEF\n+\tEOF\n+\tgit add giraffe &&\n+\tgit commit -m \"same contents as original\" &&\n+\n+\tcat >expect <<-EOF &&\n+\t$oid 1) ABC\n+\t$oid 2) DEF\n+\tEOF\n+\tgit -c core.abbrev=40 blame -s giraffe >actual &&\n+\ttest_cmp expect actual\n+'\n+\n test_done\n-- \n2.22.0.410.gd8fdbe21b5-goog\n\n"},{"id":"378329","messageId":"20190629171954.GG21574@szeder.dev","threadId":"51354","inReplyTo":"20190620163820.231316-8-brho@google.com","subject":"Re: [PATCH v9 7/9] blame: add a fingerprint heuristic to match ignored lines","fromName":"SZEDER Gábor","fromEmail":"szeder.dev@gmail.com","sentAt":"2019-06-29T17:19:54Z","receivedAt":"2019-06-29T17:20:02Z","isPatch":true,"sender":{"key":"szeder.dev@gmail.com","avatar":"https://avatars.githubusercontent.com/u/116324?v=4"},"body":"On Thu, Jun 20, 2019 at 12:38:18PM -0400, Barret Rhoden wrote:\n> diff --git a/t/t8014-blame-ignore-fuzzy.sh b/t/t8014-blame-ignore-fuzzy.sh\n> new file mode 100755\n> index 000000000000..844396615271\n> --- /dev/null\n> +++ b/t/t8014-blame-ignore-fuzzy.sh\n> @@ -0,0 +1,440 @@\n\n> +test_expect_success setup '\n> +\t{ for i in $(test_seq 2 $last_test)\n> +\tdo\n> +\t\t# Append each line in a separate commit to make it easy to\n> +\t\t# check which original line the blame output relates to.\n> +\n> +\t\tline_count=0 &&\n> +\t\t{ while IFS= read line\n> +\t\tdo\n> +\t\t\tline_count=$((line_count+1)) &&\n> +\t\t\techo \"$line\" >>\"$i\" &&\n> +\t\t\tgit add \"$i\" &&\n> +\t\t\ttest_tick &&\n> +\t\t\tGIT_AUTHOR_NAME=\"$line_count\" git commit -m \"$line_count\"\n> +\t\tdone } <\"a$i\"\n> +\tdone } &&\n> +\n> +\t{ for i in $(test_seq 2 $last_test)\n> +\tdo\n> +\t\t# Overwrite the files with the final content.\n> +\t\tcp b$i $i &&\n> +\t\tgit add $i\n> +\tdone } &&\n\nAll three loops above have a pair of {} around them...  but why?  I\ndon't think they are are necessary and the test does pass without\nthem.\n\n"},{"id":"378346","messageId":"20190630181732.4128-1-michael@platin.gs","threadId":"51354","inReplyTo":"20190629171954.GG21574@szeder.dev","subject":"[PATCH] t8014: remove unnecessary braces","fromName":"","fromEmail":"michael@platin.gs","sentAt":"2019-06-30T18:17:32Z","receivedAt":"2019-06-30T18:19:01Z","isPatch":true,"sender":{"key":"michael@platin.gs","avatar":"https://avatars.githubusercontent.com/u/1112348?v=4"},"body":"From: Michael Platings <michael@platin.gs>\n\nSigned-off-by: Michael Platings <michael@platin.gs>\n---\n t/t8014-blame-ignore-fuzzy.sh | 12 ++++++------\n 1 file changed, 6 insertions(+), 6 deletions(-)\n\ndiff --git a/t/t8014-blame-ignore-fuzzy.sh b/t/t8014-blame-ignore-fuzzy.sh\nindex 6f1a94caef..6e61882b6f 100755\n--- a/t/t8014-blame-ignore-fuzzy.sh\n+++ b/t/t8014-blame-ignore-fuzzy.sh\n@@ -298,28 +298,28 @@ EOF\n last_test=13\n \n test_expect_success setup '\n-\t{ for i in $(test_seq 2 $last_test)\n+\tfor i in $(test_seq 2 $last_test)\n \tdo\n \t\t# Append each line in a separate commit to make it easy to\n \t\t# check which original line the blame output relates to.\n \n \t\tline_count=0 &&\n-\t\t{ while IFS= read line\n+\t\twhile IFS= read line\n \t\tdo\n \t\t\tline_count=$((line_count+1)) &&\n \t\t\techo \"$line\" >>\"$i\" &&\n \t\t\tgit add \"$i\" &&\n \t\t\ttest_tick &&\n \t\t\tGIT_AUTHOR_NAME=\"$line_count\" git commit -m \"$line_count\"\n-\t\tdone } <\"a$i\"\n-\tdone } &&\n+\t\tdone <\"a$i\"\n+\tdone &&\n \n-\t{ for i in $(test_seq 2 $last_test)\n+\tfor i in $(test_seq 2 $last_test)\n \tdo\n \t\t# Overwrite the files with the final content.\n \t\tcp b$i $i &&\n \t\tgit add $i\n-\tdone } &&\n+\tdone &&\n \ttest_tick &&\n \n \t# Commit the final content all at once so it can all be\n-- \n2.21.0\n\n"},{"id":"378387","messageId":"0835ac2f-59aa-a9bb-5e47-6617dae1f810@google.com","threadId":"51354","inReplyTo":"20190630181732.4128-1-michael@platin.gs","subject":"Re: [PATCH] t8014: remove unnecessary braces","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-07-01T14:16:06Z","receivedAt":"2019-07-01T14:16:12Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"I'll squash this fix in for v10.\n\nOn 6/30/19 2:17 PM, michael@platin.gs wrote:\n> From: Michael Platings <michael@platin.gs>\n> \n> Signed-off-by: Michael Platings <michael@platin.gs>\n> ---\n>   t/t8014-blame-ignore-fuzzy.sh | 12 ++++++------\n>   1 file changed, 6 insertions(+), 6 deletions(-)\n> \n> diff --git a/t/t8014-blame-ignore-fuzzy.sh b/t/t8014-blame-ignore-fuzzy.sh\n> index 6f1a94caef..6e61882b6f 100755\n> --- a/t/t8014-blame-ignore-fuzzy.sh\n> +++ b/t/t8014-blame-ignore-fuzzy.sh\n> @@ -298,28 +298,28 @@ EOF\n>   last_test=13\n>   \n>   test_expect_success setup '\n> -\t{ for i in $(test_seq 2 $last_test)\n> +\tfor i in $(test_seq 2 $last_test)\n>   \tdo\n>   \t\t# Append each line in a separate commit to make it easy to\n>   \t\t# check which original line the blame output relates to.\n>   \n>   \t\tline_count=0 &&\n> -\t\t{ while IFS= read line\n> +\t\twhile IFS= read line\n>   \t\tdo\n>   \t\t\tline_count=$((line_count+1)) &&\n>   \t\t\techo \"$line\" >>\"$i\" &&\n>   \t\t\tgit add \"$i\" &&\n>   \t\t\ttest_tick &&\n>   \t\t\tGIT_AUTHOR_NAME=\"$line_count\" git commit -m \"$line_count\"\n> -\t\tdone } <\"a$i\"\n> -\tdone } &&\n> +\t\tdone <\"a$i\"\n> +\tdone &&\n>   \n> -\t{ for i in $(test_seq 2 $last_test)\n> +\tfor i in $(test_seq 2 $last_test)\n>   \tdo\n>   \t\t# Overwrite the files with the final content.\n>   \t\tcp b$i $i &&\n>   \t\tgit add $i\n> -\tdone } &&\n> +\tdone &&\n>   \ttest_tick &&\n>   \n>   \t# Commit the final content all at once so it can all be\n> \n\n"}]}