{"thread":{"id":"50912","subject":"[PATCH v6 0/6] blame: add the ability to ignore commits","startedAt":"2019-04-10T16:24:27Z","lastAt":"2019-04-16T04:10:42Z","messageCount":25,"participants":["Barret Rhoden","Ævar Arnfjörð Bjarmason","Junio C Hamano","Michael Platings"],"isPatch":true,"patchVersion":6,"patchTotal":6},"messages":[{"id":"373586","messageId":"20190410162409.117264-1-brho@google.com","threadId":"50912","inReplyTo":null,"subject":"[PATCH v6 0/6] blame: add the ability to ignore commits","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-10T16:24:03Z","receivedAt":"2019-04-10T16:24:27Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"This patch set adds the ability to ignore a set of commits and their\nchanges when blaming.  This can be used to ignore a commit deemed 'not\ninteresting,' such as reformatting.\n\nThe last patch in the series changes the heuristic by which ignored\nlines are attributed to specific lines in the parent commit.  This\nincreases the likelihood of blaming the 'right' commit, where 'right' is\nsubjective.  Perhaps we want another algorithm.  I'm using a relatively\nsimple one that uses the basics of Michael's fingerprinting code, but he\nhas another algorithm.\n\nv5 -> v6\nv5: https://public-inbox.org/git/20190403160207.149174-1-brho@google.com/\n- The \"guess\" heuristic can now look anywhere in the parent file for a\n  matching line, instead of just looking in the parent chunk.  The\n  chunks passed to blame_chunk() are smaller than you'd expect: they are\n  just adjacent '-' and '+' sections.  Any diff 'context' is a chunk\n  boundary.\n- Fixed the parent_len calculation.  I had been basing it off of\n  e->num_lines, and treating the blame entry as if it was the target\n  chunk, but the individual blame entries are subsets of the chunk.  I\n  just pass the parent chunk info all the way through now.\n- Use Michael's newest fingerprinting code, which is a large speedup.\n- Made a config option to zero the hash for an ignored line when the\n  heuristic could not find a line in the parent to blame.  Previously,\n  this was always 'on'.\n- Moved the for loop variable declarations out of the for ().\n- Rebased on master.\n\nv4 -> v5\nv4: https://public-inbox.org/git/20190226170648.211847-1-brho@google.com/\n- Changed the handling of blame_entries from ignored commits so that you\n  can use any algorithm you want to map lines from the diff chunk to\n  different parts of the parent commit.\n- fill_origin_blob() optionally can track the offsets of the start of\n  every line, similar to what we do in the scoreboard for the final\n  file.  This can be used by the matching algorithm.  It has no effect\n  if you are not ignoring commits.\n- RFC of a fuzzy/fingerprinting heuristic, based on Michael Platings RFC\n  at https://public-inbox.org/git/20190324235020.49706-2-michael@platin.gs/\n- Made the tests that detect unblamable entries more resilient to\n  different heuristics.\n- Fixed a few bugs:\n\t- tests were not grepping the line number from --line-porcelain\n\t  correctly.\n\t- In the old version, when I passed the \"upper\" part of the\n\t  blame entry to the target and marked unblamable, the suspect\n\t  was incorrectly marked as the parent.  The s_lno was also in\n\t  the parent's address space.\n\nv3 -> v4\nv3: https://public-inbox.org/git/20190212222722.240676-1-brho@google.com/\n- Cleaned up the tests, especially removing usage of sed -i.\n- Squashed the 'tests' commit into the other blame commits.  Let me know\n  if you'd like further squashing.\n\nv2 -> v3\nv2: https://public-inbox.org/git/20190117202919.157326-1-brho@google.com/\n- SHA-1 -> \"object name\", and fixed other comments\n- Changed error string for oidset_parse_file()\n- Adjusted existing fsck tests to handle those string changes\n- Return hash of all zeros for lines we know we cannot identify\n- Allow repeated options for blame.ignoreRevsFile and\n  --ignore-revs-file.  An empty file name resets the list.  Config\n  options are parsed before the command line options.\n- Rebased to master\n- Added regression tests\n\nv1 -> v2\nv1: https://public-inbox.org/git/20190107213013.231514-1-brho@google.com/\n- extracted the skiplist from fsck to avoid duplicating code\n- overhauled the interface and options\n- split out markIgnoredFiles\n- handled merges\n\n\nBarret Rhoden (6):\n  Move init_skiplist() outside of fsck\n  blame: use a helper function in blame_chunk()\n  blame: add the ability to ignore commits and their changes\n  blame: add config options to handle output for ignored lines\n  blame: optionally track line fingerprints during fill_blame_origin()\n  blame: use a fingerprint heuristic to match ignored lines\n\n Documentation/blame-options.txt |  18 ++\n Documentation/config/blame.txt  |  16 ++\n Documentation/git-blame.txt     |   1 +\n blame.c                         | 446 ++++++++++++++++++++++++++++----\n blame.h                         |   6 +\n builtin/blame.c                 |  56 ++++\n fsck.c                          |  37 +--\n oidset.c                        |  35 +++\n oidset.h                        |   8 +\n t/t5504-fetch-receive-strict.sh |  14 +-\n t/t8013-blame-ignore-revs.sh    | 202 +++++++++++++++\n 11 files changed, 740 insertions(+), 99 deletions(-)\n create mode 100755 t/t8013-blame-ignore-revs.sh\n\n-- \n2.21.0.392.gf8f6787159e-goog\n\n"},{"id":"373587","messageId":"20190410162409.117264-2-brho@google.com","threadId":"50912","inReplyTo":"20190410162409.117264-1-brho@google.com","subject":"[PATCH v6 1/6] Move init_skiplist() outside of fsck","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-10T16:24:04Z","receivedAt":"2019-04-10T16:24:31Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"init_skiplist() took a file consisting of SHA-1s and comments and added\nthe objects to an oidset.  This functionality is useful for other\ncommands.\n\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n fsck.c                          | 37 +--------------------------------\n oidset.c                        | 35 +++++++++++++++++++++++++++++++\n oidset.h                        |  8 +++++++\n t/t5504-fetch-receive-strict.sh | 14 ++++++-------\n 4 files changed, 51 insertions(+), 43 deletions(-)\n\ndiff --git a/fsck.c b/fsck.c\nindex 2260adb71e7a..d45534ad90f5 100644\n--- a/fsck.c\n+++ b/fsck.c\n@@ -181,41 +181,6 @@ static int fsck_msg_type(enum fsck_msg_id msg_id,\n \treturn msg_type;\n }\n \n-static void init_skiplist(struct fsck_options *options, const char *path)\n-{\n-\tFILE *fp;\n-\tstruct strbuf sb = STRBUF_INIT;\n-\tstruct object_id oid;\n-\n-\tfp = fopen(path, \"r\");\n-\tif (!fp)\n-\t\tdie(\"Could not open skip list: %s\", path);\n-\twhile (!strbuf_getline(&sb, fp)) {\n-\t\tconst char *p;\n-\t\tconst char *hash;\n-\n-\t\t/*\n-\t\t * Allow trailing comments, leading whitespace\n-\t\t * (including before commits), and empty or whitespace\n-\t\t * only lines.\n-\t\t */\n-\t\thash = strchr(sb.buf, '#');\n-\t\tif (hash)\n-\t\t\tstrbuf_setlen(&sb, hash - sb.buf);\n-\t\tstrbuf_trim(&sb);\n-\t\tif (!sb.len)\n-\t\t\tcontinue;\n-\n-\t\tif (parse_oid_hex(sb.buf, &oid, &p) || *p != '\\0')\n-\t\t\tdie(\"Invalid SHA-1: %s\", sb.buf);\n-\t\toidset_insert(&options->skiplist, &oid);\n-\t}\n-\tif (ferror(fp))\n-\t\tdie_errno(\"Could not read '%s'\", path);\n-\tfclose(fp);\n-\tstrbuf_release(&sb);\n-}\n-\n static int parse_msg_type(const char *str)\n {\n \tif (!strcmp(str, \"error\"))\n@@ -284,7 +249,7 @@ void fsck_set_msg_types(struct fsck_options *options, const char *values)\n \t\tif (!strcmp(buf, \"skiplist\")) {\n \t\t\tif (equal == len)\n \t\t\t\tdie(\"skiplist requires a path\");\n-\t\t\tinit_skiplist(options, buf + equal + 1);\n+\t\t\toidset_parse_file(&options->skiplist, buf + equal + 1);\n \t\t\tbuf += len + 1;\n \t\t\tcontinue;\n \t\t}\ndiff --git a/oidset.c b/oidset.c\nindex fe4eb921df81..878a1b56af1c 100644\n--- a/oidset.c\n+++ b/oidset.c\n@@ -35,3 +35,38 @@ void oidset_clear(struct oidset *set)\n \tkh_release_oid(&set->set);\n \toidset_init(set, 0);\n }\n+\n+void oidset_parse_file(struct oidset *set, const char *path)\n+{\n+\tFILE *fp;\n+\tstruct strbuf sb = STRBUF_INIT;\n+\tstruct object_id oid;\n+\n+\tfp = fopen(path, \"r\");\n+\tif (!fp)\n+\t\tdie(\"Could not open object name list: %s\", path);\n+\twhile (!strbuf_getline(&sb, fp)) {\n+\t\tconst char *p;\n+\t\tconst char *name;\n+\n+\t\t/*\n+\t\t * Allow trailing comments, leading whitespace\n+\t\t * (including before commits), and empty or whitespace\n+\t\t * only lines.\n+\t\t */\n+\t\tname = strchr(sb.buf, '#');\n+\t\tif (name)\n+\t\t\tstrbuf_setlen(&sb, name - sb.buf);\n+\t\tstrbuf_trim(&sb);\n+\t\tif (!sb.len)\n+\t\t\tcontinue;\n+\n+\t\tif (parse_oid_hex(sb.buf, &oid, &p) || *p != '\\0')\n+\t\t\tdie(\"Invalid object name: %s\", sb.buf);\n+\t\toidset_insert(set, &oid);\n+\t}\n+\tif (ferror(fp))\n+\t\tdie_errno(\"Could not read '%s'\", path);\n+\tfclose(fp);\n+\tstrbuf_release(&sb);\n+}\ndiff --git a/oidset.h b/oidset.h\nindex c9d0f6d3cc8b..c4807749df8d 100644\n--- a/oidset.h\n+++ b/oidset.h\n@@ -73,6 +73,14 @@ int oidset_remove(struct oidset *set, const struct object_id *oid);\n  */\n void oidset_clear(struct oidset *set);\n \n+/**\n+ * Add the contents of the file 'path' to an initialized oidset.  Each line is\n+ * an unabbreviated object name.  Comments begin with '#', and trailing comments\n+ * are allowed.  Leading whitespace and empty or white-space only lines are\n+ * ignored.\n+ */\n+void oidset_parse_file(struct oidset *set, const char *path);\n+\n struct oidset_iter {\n \tkh_oid_t *set;\n \tkhiter_t iter;\ndiff --git a/t/t5504-fetch-receive-strict.sh b/t/t5504-fetch-receive-strict.sh\nindex 7bc706873c5b..7184f1d07f90 100755\n--- a/t/t5504-fetch-receive-strict.sh\n+++ b/t/t5504-fetch-receive-strict.sh\n@@ -164,9 +164,9 @@ test_expect_success 'fsck with unsorted skipList' '\n test_expect_success 'fsck with invalid or bogus skipList input' '\n \tgit -c fsck.skipList=/dev/null -c fsck.missingEmail=ignore fsck &&\n \ttest_must_fail git -c fsck.skipList=does-not-exist -c fsck.missingEmail=ignore fsck 2>err &&\n-\ttest_i18ngrep \"Could not open skip list: does-not-exist\" err &&\n+\ttest_i18ngrep \"Could not open object name list: does-not-exist\" err &&\n \ttest_must_fail git -c fsck.skipList=.git/config -c fsck.missingEmail=ignore fsck 2>err &&\n-\ttest_i18ngrep \"Invalid SHA-1: \\[core\\]\" err\n+\ttest_i18ngrep \"Invalid object name: \\[core\\]\" err\n '\n \n test_expect_success 'fsck with other accepted skipList input (comments & empty lines)' '\n@@ -193,7 +193,7 @@ test_expect_success 'fsck no garbage output from comments & empty lines errors'\n test_expect_success 'fsck with invalid abbreviated skipList input' '\n \techo $commit | test_copy_bytes 20 >SKIP.abbreviated &&\n \ttest_must_fail git -c fsck.skipList=SKIP.abbreviated fsck 2>err-abbreviated &&\n-\ttest_i18ngrep \"^fatal: Invalid SHA-1: \" err-abbreviated\n+\ttest_i18ngrep \"^fatal: Invalid object name: \" err-abbreviated\n '\n \n test_expect_success 'fsck with exhaustive accepted skipList input (various types of comments etc.)' '\n@@ -226,10 +226,10 @@ test_expect_success 'push with receive.fsck.skipList' '\n \ttest_must_fail git push --porcelain dst bogus &&\n \tgit --git-dir=dst/.git config receive.fsck.skipList does-not-exist &&\n \ttest_must_fail git push --porcelain dst bogus 2>err &&\n-\ttest_i18ngrep \"Could not open skip list: does-not-exist\" err &&\n+\ttest_i18ngrep \"Could not open object name list: does-not-exist\" err &&\n \tgit --git-dir=dst/.git config receive.fsck.skipList config &&\n \ttest_must_fail git push --porcelain dst bogus 2>err &&\n-\ttest_i18ngrep \"Invalid SHA-1: \\[core\\]\" err &&\n+\ttest_i18ngrep \"Invalid object name: \\[core\\]\" err &&\n \n \tgit --git-dir=dst/.git config receive.fsck.skipList SKIP &&\n \tgit push --porcelain dst bogus\n@@ -255,10 +255,10 @@ test_expect_success 'fetch with fetch.fsck.skipList' '\n \ttest_must_fail git --git-dir=dst/.git fetch \"file://$(pwd)\" $refspec &&\n \tgit --git-dir=dst/.git config fetch.fsck.skipList does-not-exist &&\n \ttest_must_fail git --git-dir=dst/.git fetch \"file://$(pwd)\" $refspec 2>err &&\n-\ttest_i18ngrep \"Could not open skip list: does-not-exist\" err &&\n+\ttest_i18ngrep \"Could not open object name list: does-not-exist\" err &&\n \tgit --git-dir=dst/.git config fetch.fsck.skipList dst/.git/config &&\n \ttest_must_fail git --git-dir=dst/.git fetch \"file://$(pwd)\" $refspec 2>err &&\n-\ttest_i18ngrep \"Invalid SHA-1: \\[core\\]\" err &&\n+\ttest_i18ngrep \"Invalid object name: \\[core\\]\" err &&\n \n \tgit --git-dir=dst/.git config fetch.fsck.skipList dst/.git/SKIP &&\n \tgit --git-dir=dst/.git fetch \"file://$(pwd)\" $refspec\n-- \n2.21.0.392.gf8f6787159e-goog\n\n"},{"id":"373588","messageId":"20190410162409.117264-3-brho@google.com","threadId":"50912","inReplyTo":"20190410162409.117264-1-brho@google.com","subject":"[PATCH v6 2/6] blame: use a helper function in blame_chunk()","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-10T16:24:05Z","receivedAt":"2019-04-10T16:24:34Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"The same code for splitting a blame_entry at a particular line was used\ntwice in blame_chunk(), and I'll use the helper again in an upcoming\npatch.\n\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n blame.c | 44 ++++++++++++++++++++++++++++----------------\n 1 file changed, 28 insertions(+), 16 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex 5c07dec19035..9555a9420836 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -839,6 +839,27 @@ static struct blame_entry *reverse_blame(struct blame_entry *head,\n \treturn tail;\n }\n \n+/*\n+ * Splits a blame entry into two entries at 'len' lines.  The original 'e'\n+ * consists of len lines, i.e. [e->lno, e->lno + len), and the second part,\n+ * which is returned, consists of the remainder: [e->lno + len, e->lno +\n+ * e->num_lines).  The caller needs to sort out the reference counting for the\n+ * new entry's suspect.\n+ */\n+static struct blame_entry *split_blame_at(struct blame_entry *e, int len,\n+\t\t\t\t\t  struct blame_origin *new_suspect)\n+{\n+\tstruct blame_entry *n = xcalloc(1, sizeof(struct blame_entry));\n+\n+\tn->suspect = new_suspect;\n+\tn->lno = e->lno + len;\n+\tn->s_lno = e->s_lno + len;\n+\tn->num_lines = e->num_lines - len;\n+\te->num_lines = len;\n+\te->score = 0;\n+\treturn n;\n+}\n+\n /*\n  * Process one hunk from the patch between the current suspect for\n  * blame_entry e and its parent.  This first blames any unfinished\n@@ -865,14 +886,9 @@ static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n \t\t */\n \t\tif (e->s_lno + e->num_lines > tlno) {\n \t\t\t/* Move second half to a new record */\n-\t\t\tint len = tlno - e->s_lno;\n-\t\t\tstruct blame_entry *n = xcalloc(1, sizeof (struct blame_entry));\n-\t\t\tn->suspect = e->suspect;\n-\t\t\tn->lno = e->lno + len;\n-\t\t\tn->s_lno = e->s_lno + len;\n-\t\t\tn->num_lines = e->num_lines - len;\n-\t\t\te->num_lines = len;\n-\t\t\te->score = 0;\n+\t\t\tstruct blame_entry *n;\n+\n+\t\t\tn = split_blame_at(e, tlno - e->s_lno, e->suspect);\n \t\t\t/* Push new record to diffp */\n \t\t\tn->next = diffp;\n \t\t\tdiffp = n;\n@@ -919,14 +935,10 @@ static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n \t\t\t * Move second half to a new record to be\n \t\t\t * processed by later chunks\n \t\t\t */\n-\t\t\tint len = same - e->s_lno;\n-\t\t\tstruct blame_entry *n = xcalloc(1, sizeof (struct blame_entry));\n-\t\t\tn->suspect = blame_origin_incref(e->suspect);\n-\t\t\tn->lno = e->lno + len;\n-\t\t\tn->s_lno = e->s_lno + len;\n-\t\t\tn->num_lines = e->num_lines - len;\n-\t\t\te->num_lines = len;\n-\t\t\te->score = 0;\n+\t\t\tstruct blame_entry *n;\n+\n+\t\t\tn = split_blame_at(e, same - e->s_lno,\n+\t\t\t\t\t   blame_origin_incref(e->suspect));\n \t\t\t/* Push new record to samep */\n \t\t\tn->next = samep;\n \t\t\tsamep = n;\n-- \n2.21.0.392.gf8f6787159e-goog\n\n"},{"id":"373589","messageId":"20190410162409.117264-4-brho@google.com","threadId":"50912","inReplyTo":"20190410162409.117264-1-brho@google.com","subject":"[PATCH v6 3/6] blame: add the ability to ignore commits and their changes","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-10T16:24:06Z","receivedAt":"2019-04-10T16:24:38Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"Commits that make formatting changes or function renames are often not\ninteresting when blaming a file.  A user may deem such a commit as 'not\ninteresting' and want to ignore and its changes it when assigning blame.\n\nFor example, say a file has the following git history / rev-list:\n\n---O---A---X---B---C---D---Y---E---F\n\nCommits X and Y both touch a particular line, and the other commits do\nnot:\n\nX: \"Take a third parameter\"\n-MyFunc(1, 2);\n+MyFunc(1, 2, 3);\n\nY: \"Remove camelcase\"\n-MyFunc(1, 2, 3);\n+my_func(1, 2, 3);\n\ngit-blame will blame Y for the change.  I'd like to be able to ignore Y:\nboth the existence of the commit as well as any changes it made.  This\ndiffers from -S rev-list, which specifies the list of commits to\nprocess for the blame.  We would still process Y, but just don't let the\nblame 'stick.'\n\nThis patch adds the ability for users to ignore a revision with\n--ignore-rev=rev, which may be repeated.  They can specify a set of\nfiles of full object names of revs, e.g. SHA-1 hashes, one per line.  A\nsingle file may be specified with the blame.ignoreRevFile config option\nor with --ignore-rev-file=file.  Both the config option and the command\nline option may be repeated multiple times.  An empty file name \"\" will\nclear the list of revs from previously processed files.  Config options\nare processed before command line options.\n\nFor a typical use case, projects will maintain the file containing\nrevisions for commits that perform mass reformatting, and their users\nhave the optional to ignore all of the commits in that file.\n\nAdditionally, a user can use the --ignore-rev option for one-off\ninvestigation.  To go back to the example above, X was a substantive\nchange to the function, but not the change the user is interested in.\nThe user inspected X, but wanted to find the previous change to that\nline - perhaps a commit that introduced that function call.\n\nTo make this work, we can't simply remove all ignored commits from the\nrev-list.  We need to diff the changes introduced by Y so that we can\nignore them.  We let the blames get passed to Y, just like when\nprocessing normally.  When Y is the target, we make sure that Y does not\n*keep* any blames.  Any changes that Y is responsible for get passed to\nits parent.  Note we make one pass through all of the scapegoats\n(parents) to attempt to pass blame normally; we don't know if we *need*\nto ignore the commit until we've checked all of the parents.\n\nThe blame_entry will get passed up the tree until we find a commit that\nhas a diff chunk that affects those lines.\n\nOne issue is that the ignored commit *did* make some change, and there is\nno general solution to finding the line in the parent commit that\ncorresponds to a given line in the ignored commit.  That makes it hard\nto attribute a particular line within an ignored commit's diff\ncorrectly.\n\nFor example, the parent of an ignored commit has this, say at line 11:\n\ncommit-a 11) #include \"a.h\"\ncommit-b 12) #include \"b.h\"\n\nCommit X, which we will ignore, swaps these lines:\n\ncommit-X 11) #include \"b.h\"\ncommit-X 12) #include \"a.h\"\n\nWe can pass that blame entry to the parent, but line 11 will be\nattributed to commit A, even though \"include b.h\" came from commit B.\nThe blame mechanism will be looking at the parent's view of the file at\nline number 11.\n\nignore_blame_entry() is set up to allow alternative algorithms for\nguessing per-line blames.  Any line that is not attributed to the parent\nis marked as 'unblamable', and we output a hash of all zeros.\n\nThe existing algorithm is simple: blame each line on the corresponding\nline in the parent's diff chunk.  Any lines beyond that stay with the\ntarget.\n\nFor example, the parent of an ignored commit has this, say at line 11:\n\ncommit-a 11) void new_func_1(void *x, void *y);\ncommit-b 12) void new_func_2(void *x, void *y);\ncommit-c 13) some_line_c\ncommit-d 14) some_line_d\n\nAfter a commit 'X', we have:\n\ncommit-X 11) void new_func_1(void *x,\ncommit-X 12)                 void *y);\ncommit-X 13) void new_func_2(void *x,\ncommit-X 14)                 void *y);\ncommit-c 15) some_line_c\ncommit-d 16) some_line_d\n\nCommit X nets two additionally lines: 13 and 14.  The current\nguess_line_blames() algorithm will not attribute these to the parent,\nwhose diff chunk is only two lines - not four.\n\nWhen we ignore with the current algorithm, we get:\n\ncommit-a 11) void new_func_1(void *x,\ncommit-b 12)                 void *y);\n00000000 13) void new_func_2(void *x,\n00000000 14)                 void *y);\ncommit-c 15) some_line_c\ncommit-d 16) some_line_d\n\nNote that line 12 was blamed on B, though B was the commit for\nnew_func_2(), not new_func_1().  Even when guess_line_blames() finds a\nline in the parent, it may still be incorrect.\n\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n Documentation/blame-options.txt |  14 +++\n Documentation/config/blame.txt  |   7 ++\n Documentation/git-blame.txt     |   1 +\n blame.c                         | 181 ++++++++++++++++++++++++++++++--\n blame.h                         |   3 +\n builtin/blame.c                 |  42 ++++++++\n t/t8013-blame-ignore-revs.sh    | 168 +++++++++++++++++++++++++++++\n 7 files changed, 406 insertions(+), 10 deletions(-)\n create mode 100755 t/t8013-blame-ignore-revs.sh\n\ndiff --git a/Documentation/blame-options.txt b/Documentation/blame-options.txt\nindex dc41957afab2..8f155196c6fe 100644\n--- a/Documentation/blame-options.txt\n+++ b/Documentation/blame-options.txt\n@@ -110,5 +110,19 @@ commit. And the default value is 40. If there are more than one\n `-C` options given, the <num> argument of the last `-C` will\n take effect.\n \n+--ignore-rev <rev>::\n+\tIgnore changes made by the revision when assigning blame, as if the\n+\tchange never happened.  Lines that were changed or added by an ignored\n+\tcommit will be blamed on the previous commit that changed that line or\n+\tnearby lines.  This option may be specified multiple times to ignore\n+\tmore than one revision.\n+\n+--ignore-revs-file <file>::\n+\tIgnore revisions listed in `file`, one unabbreviated object name per line.\n+\tWhitespace and comments beginning with `#` are ignored.  This option may be\n+\trepeated, and these files will be processed after any files specified with\n+\tthe `blame.ignoreRevsFile` config option.  An empty file name, `\"\"`, will\n+\tclear the list of revs from previously processed files.\n+\n -h::\n \tShow help message.\ndiff --git a/Documentation/config/blame.txt b/Documentation/config/blame.txt\nindex 67b5c1d1e02a..4da2788f306d 100644\n--- a/Documentation/config/blame.txt\n+++ b/Documentation/config/blame.txt\n@@ -19,3 +19,10 @@ blame.showEmail::\n blame.showRoot::\n \tDo not treat root commits as boundaries in linkgit:git-blame[1].\n \tThis option defaults to false.\n+\n+blame.ignoreRevsFile::\n+\tIgnore revisions listed in the file, one unabbreviated object name per\n+\tline, in linkgit:git-blame[1].  Whitespace and comments beginning with\n+\t`#` are ignored.  This option may be repeated multiple times.  Empty\n+\tfile names will reset the list of ignored revisions.  This option will\n+\tbe handled before the command line option `--ignore-revs-file`.\ndiff --git a/Documentation/git-blame.txt b/Documentation/git-blame.txt\nindex 16323eb80e31..7e8154199635 100644\n--- a/Documentation/git-blame.txt\n+++ b/Documentation/git-blame.txt\n@@ -10,6 +10,7 @@ SYNOPSIS\n [verse]\n 'git blame' [-c] [-b] [-l] [--root] [-t] [-f] [-n] [-s] [-e] [-p] [-w] [--incremental]\n \t    [-L <range>] [-S <revs-file>] [-M] [-C] [-C] [-C] [--since=<date>]\n+\t    [--ignore-rev <rev>] [--ignore-revs-file <file>]\n \t    [--progress] [--abbrev=<n>] [<rev> | --contents <file> | --reverse <rev>..<rev>]\n \t    [--] <file>\n \ndiff --git a/blame.c b/blame.c\nindex 9555a9420836..0bbb86ad5985 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -480,7 +480,8 @@ void blame_coalesce(struct blame_scoreboard *sb)\n \n \tfor (ent = sb->ent; ent && (next = ent->next); ent = next) {\n \t\tif (ent->suspect == next->suspect &&\n-\t\t    ent->s_lno + ent->num_lines == next->s_lno) {\n+\t\t    ent->s_lno + ent->num_lines == next->s_lno &&\n+\t\t    ent->unblamable == next->unblamable) {\n \t\t\tent->num_lines += next->num_lines;\n \t\t\tent->next = next->next;\n \t\t\tblame_origin_decref(next->suspect);\n@@ -732,6 +733,10 @@ static void split_overlap(struct blame_entry *split,\n \tint chunk_end_lno;\n \tmemset(split, 0, sizeof(struct blame_entry [3]));\n \n+\tsplit[0].unblamable = e->unblamable;\n+\tsplit[1].unblamable = e->unblamable;\n+\tsplit[2].unblamable = e->unblamable;\n+\n \tif (e->s_lno < tlno) {\n \t\t/* there is a pre-chunk part not blamed on parent */\n \t\tsplit[0].suspect = blame_origin_incref(e->suspect);\n@@ -852,6 +857,7 @@ static struct blame_entry *split_blame_at(struct blame_entry *e, int len,\n \tstruct blame_entry *n = xcalloc(1, sizeof(struct blame_entry));\n \n \tn->suspect = new_suspect;\n+\tn->unblamable = e->unblamable;\n \tn->lno = e->lno + len;\n \tn->s_lno = e->s_lno + len;\n \tn->num_lines = e->num_lines - len;\n@@ -860,6 +866,109 @@ static struct blame_entry *split_blame_at(struct blame_entry *e, int len,\n \treturn n;\n }\n \n+struct blame_line_tracker {\n+\tint is_parent;\n+\tint s_lno;\n+};\n+\n+static int are_lines_adjacent(struct blame_line_tracker *first,\n+\t\t\t      struct blame_line_tracker *second)\n+{\n+\treturn first->is_parent == second->is_parent &&\n+\t       first->s_lno + 1 == second->s_lno;\n+}\n+\n+/*\n+ * This cheap heuristic assigns lines in the chunk to their relative location in\n+ * the parent's chunk.  Any additional lines are left with the target.\n+ */\n+static void guess_line_blames(struct blame_entry *e,\n+\t\t\t      struct blame_origin *parent,\n+\t\t\t      struct blame_origin *target,\n+\t\t\t      int offset, int parent_slno, int parent_len,\n+\t\t\t      struct blame_line_tracker *line_blames)\n+{\n+\tint i, parent_idx;\n+\n+\tfor (i = 0; i < e->num_lines; i++) {\n+\t\tparent_idx = e->s_lno + i + offset;\n+\t\tif (parent_slno <= parent_idx &&\n+\t\t    parent_idx < parent_slno + parent_len) {\n+\t\t\tline_blames[i].is_parent = 1;\n+\t\t\tline_blames[i].s_lno = parent_idx;\n+\t\t} else {\n+\t\t\tline_blames[i].is_parent = 0;\n+\t\t\tline_blames[i].s_lno = e->s_lno + i;\n+\t\t}\n+\t}\n+}\n+\n+/*\n+ * This decides which parts of a blame entry go to the parent (added to the\n+ * ignoredp list) and which stay with the target (added to the diffp list).  The\n+ * actual decision is made in a separate heuristic function.  This consumes e,\n+ * essentially putting it on a list.\n+ *\n+ * Note that the blame entries on the ignoredp list are not necessarily sorted\n+ * with respect to the parent's line numbers yet.\n+ */\n+static void ignore_blame_entry(struct blame_entry *e,\n+\t\t\t       struct blame_origin *parent,\n+\t\t\t       struct blame_origin *target,\n+\t\t\t       int offset, int parent_slno, int parent_len,\n+\t\t\t       struct blame_entry **diffp,\n+\t\t\t       struct blame_entry **ignoredp)\n+{\n+\tstruct blame_line_tracker *line_blames;\n+\tint entry_len, nr_lines, i;\n+\n+\tline_blames = xcalloc(sizeof(struct blame_line_tracker),\n+\t\t\t      e->num_lines);\n+\tguess_line_blames(e, parent, target, offset, parent_slno, parent_len,\n+\t\t\t  line_blames);\n+\t/*\n+\t * We carve new entries off the front of e.  Each entry comes from a\n+\t * contiguous chunk of lines: adjacent lines from the same origin\n+\t * (either the parent or the target).\n+\t */\n+\tentry_len = 1;\n+\tnr_lines = e->num_lines;\t// e changes in the loop\n+\tfor (i = 0; i < nr_lines; i++) {\n+\t\tstruct blame_entry *next = NULL;\n+\n+\t\t/*\n+\t\t * We are often adjacent to the next line - only split the blame\n+\t\t * entry when we have to.\n+\t\t */\n+\t\tif (i + 1 < nr_lines) {\n+\t\t\tif (are_lines_adjacent(&line_blames[i],\n+\t\t\t\t\t       &line_blames[i + 1])) {\n+\t\t\t\tentry_len++;\n+\t\t\t\tcontinue;\n+\t\t\t}\n+\t\t\tnext = split_blame_at(e, entry_len,\n+\t\t\t\t\t      blame_origin_incref(e->suspect));\n+\t\t}\n+\t\tif (line_blames[i].is_parent) {\n+\t\t\tblame_origin_decref(e->suspect);\n+\t\t\te->suspect = blame_origin_incref(parent);\n+\t\t\te->s_lno = line_blames[i - entry_len + 1].s_lno;\n+\t\t\te->next = *ignoredp;\n+\t\t\t*ignoredp = e;\n+\t\t} else {\n+\t\t\te->unblamable = 1;\n+\t\t\t/* e->s_lno is already in the target's address space. */\n+\t\t\te->next = *diffp;\n+\t\t\t*diffp = e;\n+\t\t}\n+\t\tassert(e->num_lines == entry_len);\n+\t\te = next;\n+\t\tentry_len = 1;\n+\t}\n+\tassert(!e);\n+\tfree(line_blames);\n+}\n+\n /*\n  * Process one hunk from the patch between the current suspect for\n  * blame_entry e and its parent.  This first blames any unfinished\n@@ -869,13 +978,19 @@ static struct blame_entry *split_blame_at(struct blame_entry *e, int len,\n  * -C options may lead to overlapping/duplicate source line number\n  * ranges, all we can rely on from sorting/merging is the order of the\n  * first suspect line number.\n+ *\n+ * tlno: line number in the target where this chunk begins\n+ * same: line number in the target where this chunk ends\n+ * offset: add to tlno to get the chunk starting point in the parent\n+ * parent_len: number of lines in the parent chunk\n  */\n static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n-\t\t\tint tlno, int offset, int same,\n-\t\t\tstruct blame_origin *parent)\n+\t\t\tint tlno, int offset, int same, int parent_len,\n+\t\t\tstruct blame_origin *parent,\n+\t\t\tstruct blame_origin *target, int ignore_diffs)\n {\n \tstruct blame_entry *e = **srcq;\n-\tstruct blame_entry *samep = NULL, *diffp = NULL;\n+\tstruct blame_entry *samep = NULL, *diffp = NULL, *ignoredp = NULL;\n \n \twhile (e && e->s_lno < tlno) {\n \t\tstruct blame_entry *next = e->next;\n@@ -943,10 +1058,29 @@ static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n \t\t\tn->next = samep;\n \t\t\tsamep = n;\n \t\t}\n-\t\te->next = diffp;\n-\t\tdiffp = e;\n+\t\tif (ignore_diffs) {\n+\t\t\tignore_blame_entry(e, parent, target, offset,\n+\t\t\t\t\t   tlno + offset, parent_len, &diffp,\n+\t\t\t\t\t   &ignoredp);\n+\t\t} else {\n+\t\t\te->next = diffp;\n+\t\t\tdiffp = e;\n+\t\t}\n \t\te = next;\n \t}\n+\tif (ignoredp) {\n+\t\t/*\n+\t\t * Note ignoredp is not sorted yet, and thus neither is dstq.\n+\t\t * That list must be sorted before we queue_blames().  We defer\n+\t\t * sorting until after all diff hunks are processed, so that\n+\t\t * guess_line_blames() can pick *any* line in the parent.  The\n+\t\t * slight drawback is that we end up sorting all blame entries\n+\t\t * passed to the parent, including those that are unrelated to\n+\t\t * changes made by the ignored commit.\n+\t\t */\n+\t\t**dstq = reverse_blame(ignoredp, **dstq);\n+\t\t*dstq = &ignoredp->next;\n+\t}\n \t**srcq = reverse_blame(diffp, reverse_blame(samep, e));\n \t/* Move across elements that are in the unblamable portion */\n \tif (diffp)\n@@ -955,7 +1089,9 @@ static void blame_chunk(struct blame_entry ***dstq, struct blame_entry ***srcq,\n \n struct blame_chunk_cb_data {\n \tstruct blame_origin *parent;\n+\tstruct blame_origin *target;\n \tlong offset;\n+\tint ignore_diffs;\n \tstruct blame_entry **dstq;\n \tstruct blame_entry **srcq;\n };\n@@ -968,7 +1104,8 @@ static int blame_chunk_cb(long start_a, long count_a,\n \tif (start_a - start_b != d->offset)\n \t\tdie(\"internal error in blame::blame_chunk_cb\");\n \tblame_chunk(&d->dstq, &d->srcq, start_b, start_a - start_b,\n-\t\t    start_b + count_b, d->parent);\n+\t\t    start_b + count_b, count_a, d->parent, d->target,\n+\t\t    d->ignore_diffs);\n \td->offset = start_a + count_a - (start_b + count_b);\n \treturn 0;\n }\n@@ -980,7 +1117,7 @@ static int blame_chunk_cb(long start_a, long count_a,\n  */\n static void pass_blame_to_parent(struct blame_scoreboard *sb,\n \t\t\t\t struct blame_origin *target,\n-\t\t\t\t struct blame_origin *parent)\n+\t\t\t\t struct blame_origin *parent, int ignore_diffs)\n {\n \tmmfile_t file_p, file_o;\n \tstruct blame_chunk_cb_data d;\n@@ -990,7 +1127,9 @@ static void pass_blame_to_parent(struct blame_scoreboard *sb,\n \t\treturn; /* nothing remains for this target */\n \n \td.parent = parent;\n+\td.target = target;\n \td.offset = 0;\n+\td.ignore_diffs = ignore_diffs;\n \td.dstq = &newdest; d.srcq = &target->suspects;\n \n \tfill_origin_blob(&sb->revs->diffopt, parent, &file_p, &sb->num_read_blob);\n@@ -1002,8 +1141,13 @@ static void pass_blame_to_parent(struct blame_scoreboard *sb,\n \t\t    oid_to_hex(&parent->commit->object.oid),\n \t\t    oid_to_hex(&target->commit->object.oid));\n \t/* The rest are the same as the parent */\n-\tblame_chunk(&d.dstq, &d.srcq, INT_MAX, d.offset, INT_MAX, parent);\n+\tblame_chunk(&d.dstq, &d.srcq, INT_MAX, d.offset, INT_MAX, 0,\n+\t\t    parent, target, 0);\n \t*d.dstq = NULL;\n+\tif (ignore_diffs)\n+\t\tnewdest = llist_mergesort(newdest, get_next_blame,\n+\t\t\t\t\t  set_next_blame,\n+\t\t\t\t\t  compare_blame_suspect);\n \tqueue_blames(sb, parent, newdest);\n \n \treturn;\n@@ -1507,11 +1651,28 @@ static void pass_blame(struct blame_scoreboard *sb, struct blame_origin *origin,\n \t\t\tblame_origin_incref(porigin);\n \t\t\torigin->previous = porigin;\n \t\t}\n-\t\tpass_blame_to_parent(sb, origin, porigin);\n+\t\tpass_blame_to_parent(sb, origin, porigin, 0);\n \t\tif (!origin->suspects)\n \t\t\tgoto finish;\n \t}\n \n+\t/*\n+\t * Pass remaining suspects for ignored commits to their parents.\n+\t */\n+\tif (oidset_contains(&sb->ignore_list, &commit->object.oid)) {\n+\t\tfor (i = 0, sg = first_scapegoat(revs, commit, sb->reverse);\n+\t\t     i < num_sg && sg;\n+\t\t     sg = sg->next, i++) {\n+\t\t\tstruct blame_origin *porigin = sg_origin[i];\n+\n+\t\t\tif (!porigin)\n+\t\t\t\tcontinue;\n+\t\t\tpass_blame_to_parent(sb, origin, porigin, 1);\n+\t\t\tif (!origin->suspects)\n+\t\t\t\tgoto finish;\n+\t\t}\n+\t}\n+\n \t/*\n \t * Optionally find moves in parents' files.\n \t */\ndiff --git a/blame.h b/blame.h\nindex be3a895043e0..91664913d7c4 100644\n--- a/blame.h\n+++ b/blame.h\n@@ -92,6 +92,7 @@ struct blame_entry {\n \t * scanning the lines over and over.\n \t */\n \tunsigned score;\n+\tint unblamable;\n };\n \n /*\n@@ -117,6 +118,8 @@ struct blame_scoreboard {\n \t/* linked list of blames */\n \tstruct blame_entry *ent;\n \n+\tstruct oidset ignore_list;\n+\n \t/* look-up a line in the final buffer */\n \tint num_lines;\n \tint *lineno;\ndiff --git a/builtin/blame.c b/builtin/blame.c\nindex 177c1022a0c4..b48842d8459b 100644\n--- a/builtin/blame.c\n+++ b/builtin/blame.c\n@@ -52,6 +52,7 @@ static int no_whole_file_rename;\n static int show_progress;\n static char repeated_meta_color[COLOR_MAXLEN];\n static int coloring_mode;\n+static struct string_list ignore_revs_file_list = STRING_LIST_INIT_NODUP;\n \n static struct date_mode blame_date_mode = { DATE_ISO8601 };\n static size_t blame_date_width;\n@@ -346,6 +347,8 @@ static void emit_porcelain(struct blame_scoreboard *sb, struct blame_entry *ent,\n \tchar hex[GIT_MAX_HEXSZ + 1];\n \n \toid_to_hex_r(hex, &suspect->commit->object.oid);\n+\tif (ent->unblamable)\n+\t\tmemset(hex, '0', strlen(hex));\n \tprintf(\"%s %d %d %d\\n\",\n \t       hex,\n \t       ent->s_lno + 1,\n@@ -479,6 +482,8 @@ static void emit_other(struct blame_scoreboard *sb, struct blame_entry *ent, int\n \t\t\t}\n \t\t}\n \n+\t\tif (ent->unblamable)\n+\t\t\tmemset(hex, '0', length);\n \t\tprintf(\"%.*s\", length, hex);\n \t\tif (opt & OUTPUT_ANNOTATE_COMPAT) {\n \t\t\tconst char *name;\n@@ -695,6 +700,16 @@ static int git_blame_config(const char *var, const char *value, void *cb)\n \t\tparse_date_format(value, &blame_date_mode);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(var, \"blame.ignorerevsfile\")) {\n+\t\tconst char *str;\n+\t\tint ret;\n+\n+\t\tret = git_config_pathname(&str, var, value);\n+\t\tif (ret)\n+\t\t\treturn ret;\n+\t\tstring_list_insert(&ignore_revs_file_list, str);\n+\t\treturn 0;\n+\t}\n \tif (!strcmp(var, \"color.blame.repeatedlines\")) {\n \t\tif (color_parse_mem(value, strlen(value), repeated_meta_color))\n \t\t\twarning(_(\"invalid color '%s' in color.blame.repeatedLines\"),\n@@ -774,6 +789,27 @@ static int is_a_rev(const char *name)\n \treturn OBJ_NONE < oid_object_info(the_repository, &oid, NULL);\n }\n \n+static void build_ignorelist(struct blame_scoreboard *sb,\n+\t\t\t     struct string_list *ignore_revs_file_list,\n+\t\t\t     struct string_list *ignore_rev_list)\n+{\n+\tstruct string_list_item *i;\n+\tstruct object_id oid;\n+\n+\toidset_init(&sb->ignore_list, 0);\n+\tfor_each_string_list_item(i, ignore_revs_file_list) {\n+\t\tif (!strcmp(i->string, \"\"))\n+\t\t\toidset_clear(&sb->ignore_list);\n+\t\telse\n+\t\t\toidset_parse_file(&sb->ignore_list, i->string);\n+\t}\n+\tfor_each_string_list_item(i, ignore_rev_list) {\n+\t\tif (get_oid_committish(i->string, &oid))\n+\t\t\tdie(_(\"Cannot find revision %s to ignore\"), i->string);\n+\t\toidset_insert(&sb->ignore_list, &oid);\n+\t}\n+}\n+\n int cmd_blame(int argc, const char **argv, const char *prefix)\n {\n \tstruct rev_info revs;\n@@ -785,6 +821,7 @@ int cmd_blame(int argc, const char **argv, const char *prefix)\n \tstruct progress_info pi = { NULL, 0 };\n \n \tstruct string_list range_list = STRING_LIST_INIT_NODUP;\n+\tstruct string_list ignore_rev_list = STRING_LIST_INIT_NODUP;\n \tint output_option = 0, opt = 0;\n \tint show_stats = 0;\n \tconst char *revs_file = NULL;\n@@ -806,6 +843,8 @@ int cmd_blame(int argc, const char **argv, const char *prefix)\n \t\tOPT_BIT('s', NULL, &output_option, N_(\"Suppress author name and timestamp (Default: off)\"), OUTPUT_NO_AUTHOR),\n \t\tOPT_BIT('e', \"show-email\", &output_option, N_(\"Show author email instead of name (Default: off)\"), OUTPUT_SHOW_EMAIL),\n \t\tOPT_BIT('w', NULL, &xdl_opts, N_(\"Ignore whitespace differences\"), XDF_IGNORE_WHITESPACE),\n+\t\tOPT_STRING_LIST(0, \"ignore-rev\", &ignore_rev_list, N_(\"rev\"), N_(\"Ignore <rev> when blaming\")),\n+\t\tOPT_STRING_LIST(0, \"ignore-revs-file\", &ignore_revs_file_list, N_(\"file\"), N_(\"Ignore revisions from <file>\")),\n \t\tOPT_BIT(0, \"color-lines\", &output_option, N_(\"color redundant metadata from previous line differently\"), OUTPUT_COLOR_LINE),\n \t\tOPT_BIT(0, \"color-by-age\", &output_option, N_(\"color lines by age\"), OUTPUT_SHOW_AGE_WITH_COLOR),\n \n@@ -999,6 +1038,9 @@ int cmd_blame(int argc, const char **argv, const char *prefix)\n \tsb.contents_from = contents_from;\n \tsb.reverse = reverse;\n \tsb.repo = the_repository;\n+\tbuild_ignorelist(&sb, &ignore_revs_file_list, &ignore_rev_list);\n+\tstring_list_clear(&ignore_revs_file_list, 0);\n+\tstring_list_clear(&ignore_rev_list, 0);\n \tsetup_scoreboard(&sb, path, &o);\n \tlno = sb.num_lines;\n \ndiff --git a/t/t8013-blame-ignore-revs.sh b/t/t8013-blame-ignore-revs.sh\nnew file mode 100755\nindex 000000000000..df4993f98682\n--- /dev/null\n+++ b/t/t8013-blame-ignore-revs.sh\n@@ -0,0 +1,168 @@\n+#!/bin/sh\n+\n+test_description='ignore revisions when blaming'\n+. ./test-lib.sh\n+\n+# Creates:\n+# \tA--B--X\n+# A added line 1 and B added line 2.  X makes changes to those lines.  Sanity\n+# check that X is blamed for both lines.\n+test_expect_success setup '\n+\ttest_commit A file line1 &&\n+\n+\techo line2 >>file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m B &&\n+\tgit tag B &&\n+\n+\ttest_write_lines line-one line-two >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m X &&\n+\tgit tag X &&\n+\n+\tgit blame --line-porcelain file >blame_raw &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse X >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse X >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# Ignore X, make sure A is blamed for line 1 and B for line 2.\n+test_expect_success ignore_rev_changing_lines '\n+\tgit blame --line-porcelain --ignore-rev X file >blame_raw &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse A >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse B >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# For ignored revs that have added 'unblamable' lines, blame those lines on an\n+# all-zeros rev.\n+# \tA--B--X--Y\n+# Where Y changes lines 1 and 2, and adds lines 3 and 4.  The added lines ought\n+# to have nothing in common with \"line-one\" or \"line-two\", to keep any\n+# heuristics from matching them with any lines in the parent.\n+test_expect_success ignore_rev_adding_unblamable_lines '\n+\ttest_write_lines line-one-change line-two-changed y3 y4 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m Y &&\n+\tgit tag Y &&\n+\n+\tgit rev-parse Y >y_rev &&\n+\tsed -e \"s/[0-9a-f]/0/g\" y_rev >expect &&\n+\tgit blame --line-porcelain file --ignore-rev Y >blame_raw &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 3\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 4\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# Ignore X and Y, both in separate files.  Lines 1 == A, 2 == B.\n+test_expect_success ignore_revs_from_files '\n+\tgit rev-parse X >ignore_x &&\n+\tgit rev-parse Y >ignore_y &&\n+\tgit blame --line-porcelain file --ignore-revs-file ignore_x --ignore-revs-file ignore_y >blame_raw &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse A >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse B >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# Ignore X from the config option, Y from a file.\n+test_expect_success ignore_revs_from_configs_and_files '\n+\tgit config --add blame.ignoreRevsFile ignore_x &&\n+\tgit blame --line-porcelain file --ignore-revs-file ignore_y >blame_raw &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse A >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse B >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+# Override blame.ignoreRevsFile (ignore_x) with an empty string.  X should be\n+# blamed now for lines 1 and 2, since we are no longer ignoring X.\n+test_expect_success override_ignore_revs_file '\n+\tgit blame --line-porcelain file --ignore-revs-file \"\" --ignore-revs-file ignore_y >blame_raw &&\n+\tgit rev-parse X >expect &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\ttest_cmp expect actual\n+\t'\n+test_expect_success bad_files_and_revs '\n+\ttest_must_fail git blame file --ignore-rev NOREV 2>err &&\n+\ttest_i18ngrep \"Cannot find revision NOREV to ignore\" err &&\n+\n+\ttest_must_fail git blame file --ignore-revs-file NOFILE 2>err &&\n+\ttest_i18ngrep \"Could not open object name list: NOFILE\" err &&\n+\n+\techo NOREV >ignore_norev &&\n+\ttest_must_fail git blame file --ignore-revs-file ignore_norev 2>err &&\n+\ttest_i18ngrep \"Invalid object name: NOREV\" err\n+\t'\n+\n+# Resetting the repo and creating:\n+#\n+# A--B--M\n+#  \\   /\n+#   C-+\n+#\n+# 'A' creates a file.  B changes line 1, and C changes line 9.  M merges.\n+test_expect_success ignore_merge '\n+\trm -rf .git/ &&\n+\tgit init &&\n+\n+\ttest_write_lines L1 L2 L3 L4 L5 L6 L7 L8 L9 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m A &&\n+\tgit tag A &&\n+\n+\ttest_write_lines BB L2 L3 L4 L5 L6 L7 L8 L9 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m B &&\n+\tgit tag B &&\n+\n+\tgit reset --hard A &&\n+\ttest_write_lines L1 L2 L3 L4 L5 L6 L7 L8 CC >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m C &&\n+\tgit tag C &&\n+\n+\ttest_merge M B &&\n+\tgit blame --line-porcelain file --ignore-rev M >blame_raw &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse B >expect &&\n+\ttest_cmp expect actual &&\n+\n+\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 9\" blame_raw | sed -e \"s/ .*//\" >actual &&\n+\tgit rev-parse C >expect &&\n+\ttest_cmp expect actual\n+\t'\n+\n+test_done\n-- \n2.21.0.392.gf8f6787159e-goog\n\n"},{"id":"373590","messageId":"20190410162409.117264-5-brho@google.com","threadId":"50912","inReplyTo":"20190410162409.117264-1-brho@google.com","subject":"[PATCH v6 4/6] blame: add config options to handle output for ignored lines","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-10T16:24:07Z","receivedAt":"2019-04-10T16:24:42Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"When ignoring commits, the commit that is blamed might not be\nresponsible for the change.  Users might want to know when a particular\nline has a potentially inaccurate blame.  Furthermore, they might never\nwant to see the object hash of an ignored commit.\n\nThis patch adds two config options to control the output behavior.\n\nThe first option can identify ignored lines by specifying\nblame.markIgnoredFiles.  When this option is set, each blame line is\nmarked with an '*'.\n\nFor example:\n\t278b6158d6fdb (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\nappears as:\n\t*278b6158d6fd (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\n\nwhere the '*' is placed before the commit, and the hash has one fewer\ncharacters.\n\nSometimes we are unable to even guess at what commit touched a line.\nThese lines are 'unblamable.'  The second option,\nblame.maskIgnoredUnblamables, will zero the hash of any unblamable line.\n\nFor example, say we ignore e5e8d36d04cbe:\n\te5e8d36d04cbe (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\nappears as:\n\t0000000000000 (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\n\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n Documentation/blame-options.txt |  6 +++++-\n Documentation/config/blame.txt  |  9 +++++++++\n blame.c                         |  4 ++++\n blame.h                         |  1 +\n builtin/blame.c                 | 18 +++++++++++++++--\n t/t8013-blame-ignore-revs.sh    | 34 +++++++++++++++++++++++++++++++++\n 6 files changed, 69 insertions(+), 3 deletions(-)\n\ndiff --git a/Documentation/blame-options.txt b/Documentation/blame-options.txt\nindex 8f155196c6fe..e7b8b5e4b87b 100644\n--- a/Documentation/blame-options.txt\n+++ b/Documentation/blame-options.txt\n@@ -115,7 +115,11 @@ take effect.\n \tchange never happened.  Lines that were changed or added by an ignored\n \tcommit will be blamed on the previous commit that changed that line or\n \tnearby lines.  This option may be specified multiple times to ignore\n-\tmore than one revision.\n+\tmore than one revision.  If the `blame.markIgnoredLines` config option\n+\tis set, then lines that were changed by an ignored commit will be\n+\tmarked with a `*` in the blame output.  If the\n+\t`blame.maskIgnoredUnblamables` config option is set, then those lines that\n+\twe could not attribute to another revision are outputted as all zeros.\n \n --ignore-revs-file <file>::\n \tIgnore revisions listed in `file`, one unabbreviated object name per line.\ndiff --git a/Documentation/config/blame.txt b/Documentation/config/blame.txt\nindex 4da2788f306d..bb6674227da1 100644\n--- a/Documentation/config/blame.txt\n+++ b/Documentation/config/blame.txt\n@@ -26,3 +26,12 @@ blame.ignoreRevsFile::\n \t`#` are ignored.  This option may be repeated multiple times.  Empty\n \tfile names will reset the list of ignored revisions.  This option will\n \tbe handled before the command line option `--ignore-revs-file`.\n+\n+blame.maskIgnoredUnblamables::\n+\tOutput an object hash of all zeros for lines that were changed by an ignored\n+\trevision and that we could not attribute to another revision in the output\n+\tof linkgit:git-blame[1].\n+\n+blame.markIgnoredLines::\n+\tMark lines that were changed by an ignored revision with a '*' in the\n+\toutput of linkgit:git-blame[1].\ndiff --git a/blame.c b/blame.c\nindex 0bbb86ad5985..a98ae00e2cfc 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -481,6 +481,7 @@ void blame_coalesce(struct blame_scoreboard *sb)\n \tfor (ent = sb->ent; ent && (next = ent->next); ent = next) {\n \t\tif (ent->suspect == next->suspect &&\n \t\t    ent->s_lno + ent->num_lines == next->s_lno &&\n+\t\t    ent->ignored == next->ignored &&\n \t\t    ent->unblamable == next->unblamable) {\n \t\t\tent->num_lines += next->num_lines;\n \t\t\tent->next = next->next;\n@@ -733,6 +734,7 @@ static void split_overlap(struct blame_entry *split,\n \tint chunk_end_lno;\n \tmemset(split, 0, sizeof(struct blame_entry [3]));\n \n+\tsplit[0].ignored = split[1].ignored = split[2].ignored = e->ignored;\n \tsplit[0].unblamable = e->unblamable;\n \tsplit[1].unblamable = e->unblamable;\n \tsplit[2].unblamable = e->unblamable;\n@@ -857,6 +859,7 @@ static struct blame_entry *split_blame_at(struct blame_entry *e, int len,\n \tstruct blame_entry *n = xcalloc(1, sizeof(struct blame_entry));\n \n \tn->suspect = new_suspect;\n+\tn->ignored = e->ignored;\n \tn->unblamable = e->unblamable;\n \tn->lno = e->lno + len;\n \tn->s_lno = e->s_lno + len;\n@@ -922,6 +925,7 @@ static void ignore_blame_entry(struct blame_entry *e,\n \tstruct blame_line_tracker *line_blames;\n \tint entry_len, nr_lines, i;\n \n+\te->ignored = 1;\n \tline_blames = xcalloc(sizeof(struct blame_line_tracker),\n \t\t\t      e->num_lines);\n \tguess_line_blames(e, parent, target, offset, parent_slno, parent_len,\ndiff --git a/blame.h b/blame.h\nindex 91664913d7c4..53df8b4c5b3f 100644\n--- a/blame.h\n+++ b/blame.h\n@@ -92,6 +92,7 @@ struct blame_entry {\n \t * scanning the lines over and over.\n \t */\n \tunsigned score;\n+\tint ignored;\n \tint unblamable;\n };\n \ndiff --git a/builtin/blame.c b/builtin/blame.c\nindex b48842d8459b..c10a6a802240 100644\n--- a/builtin/blame.c\n+++ b/builtin/blame.c\n@@ -53,6 +53,8 @@ static int show_progress;\n static char repeated_meta_color[COLOR_MAXLEN];\n static int coloring_mode;\n static struct string_list ignore_revs_file_list = STRING_LIST_INIT_NODUP;\n+static int mask_ignored_unblamables;\n+static int mark_ignored_lines;\n \n static struct date_mode blame_date_mode = { DATE_ISO8601 };\n static size_t blame_date_width;\n@@ -347,7 +349,7 @@ static void emit_porcelain(struct blame_scoreboard *sb, struct blame_entry *ent,\n \tchar hex[GIT_MAX_HEXSZ + 1];\n \n \toid_to_hex_r(hex, &suspect->commit->object.oid);\n-\tif (ent->unblamable)\n+\tif (mask_ignored_unblamables && ent->unblamable)\n \t\tmemset(hex, '0', strlen(hex));\n \tprintf(\"%s %d %d %d\\n\",\n \t       hex,\n@@ -482,7 +484,11 @@ static void emit_other(struct blame_scoreboard *sb, struct blame_entry *ent, int\n \t\t\t}\n \t\t}\n \n-\t\tif (ent->unblamable)\n+\t\tif (mark_ignored_lines && ent->ignored) {\n+\t\t\tlength--;\n+\t\t\tputchar('*');\n+\t\t}\n+\t\tif (mask_ignored_unblamables && ent->unblamable)\n \t\t\tmemset(hex, '0', length);\n \t\tprintf(\"%.*s\", length, hex);\n \t\tif (opt & OUTPUT_ANNOTATE_COMPAT) {\n@@ -710,6 +716,14 @@ static int git_blame_config(const char *var, const char *value, void *cb)\n \t\tstring_list_insert(&ignore_revs_file_list, str);\n \t\treturn 0;\n \t}\n+\tif (!strcmp(var, \"blame.maskignoredunblamables\")) {\n+\t\tmask_ignored_unblamables = git_config_bool(var, value);\n+\t\treturn 0;\n+\t}\n+\tif (!strcmp(var, \"blame.markignoredlines\")) {\n+\t\tmark_ignored_lines = git_config_bool(var, value);\n+\t\treturn 0;\n+\t}\n \tif (!strcmp(var, \"color.blame.repeatedlines\")) {\n \t\tif (color_parse_mem(value, strlen(value), repeated_meta_color))\n \t\t\twarning(_(\"invalid color '%s' in color.blame.repeatedLines\"),\ndiff --git a/t/t8013-blame-ignore-revs.sh b/t/t8013-blame-ignore-revs.sh\nindex df4993f98682..cc049a390b0d 100755\n--- a/t/t8013-blame-ignore-revs.sh\n+++ b/t/t8013-blame-ignore-revs.sh\n@@ -53,6 +53,7 @@ test_expect_success ignore_rev_changing_lines '\n # to have nothing in common with \"line-one\" or \"line-two\", to keep any\n # heuristics from matching them with any lines in the parent.\n test_expect_success ignore_rev_adding_unblamable_lines '\n+\tgit config --add blame.maskIgnoredUnblamables true &&\n \ttest_write_lines line-one-change line-two-changed y3 y4 >file &&\n \tgit add file &&\n \ttest_tick &&\n@@ -123,6 +124,39 @@ test_expect_success bad_files_and_revs '\n \ttest_i18ngrep \"Invalid object name: NOREV\" err\n \t'\n \n+# Commit Z will touch the first two lines.  Y touched all four.\n+# \tA--B--X--Y--Z\n+# The blame output when ignoring Z should be:\n+# ^Y ... 1)\n+# ^Y ... 2)\n+# Y  ... 3)\n+# Y  ... 4)\n+# We're checking only the first character\n+test_expect_success mark_ignored_lines '\n+\tgit config --add blame.markIgnoredLines true &&\n+\n+\ttest_write_lines line-one-Z line-two-Z y3 y4 >file &&\n+\tgit add file &&\n+\ttest_tick &&\n+\tgit commit -m Z &&\n+\tgit tag Z &&\n+\n+\tgit blame --ignore-rev Z file >blame_raw &&\n+\techo \"*\" >expect &&\n+\n+\tsed -n \"1p\" blame_raw | cut -c1 >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tsed -n \"2p\" blame_raw | cut -c1 >actual &&\n+\ttest_cmp expect actual &&\n+\n+\tsed -n \"3p\" blame_raw | cut -c1 >actual &&\n+\t! test_cmp expect actual &&\n+\n+\tsed -n \"4p\" blame_raw | cut -c1 >actual &&\n+\t! test_cmp expect actual\n+\t'\n+\n # Resetting the repo and creating:\n #\n # A--B--M\n-- \n2.21.0.392.gf8f6787159e-goog\n\n"},{"id":"373591","messageId":"20190410162409.117264-6-brho@google.com","threadId":"50912","inReplyTo":"20190410162409.117264-1-brho@google.com","subject":"[PATCH v6 5/6] blame: optionally track line fingerprints during fill_blame_origin()","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-10T16:24:08Z","receivedAt":"2019-04-10T16:24:47Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"fill_blame_origin() is a convenient place to store data that we will use\nthroughout the lifetime of a blame_origin.  Some heuristics for\nignoring commits during a blame session can make use of this storage.\nIn particular, we will calculate a fingerprint for each line of a file\nfor blame_origins involved in an ignored commit.\n\nIn this commit, we only calculate the line_starts, reusing the existing\ncode from the scoreboard's line_starts.  In an upcoming commit, we will\nactually compute the fingerprints.\n\nThis feature will be used when we attempt to pass blame entries to\nparents when we \"ignore\" a commit.  Most uses of fill_blame_origin()\nwill not require this feature, hence the flag parameter.  Multiple calls\nto fill_blame_origin() are idempotent, and any of them can request the\ncreation of the fingerprints structure.\n\nSuggested-by: Michael Platings <michael@platin.gs>\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n blame.c | 95 +++++++++++++++++++++++++++++++++++++++------------------\n blame.h |  2 ++\n 2 files changed, 67 insertions(+), 30 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex a98ae00e2cfc..a42dff80b1a5 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -311,12 +311,63 @@ static int diff_hunks(mmfile_t *file_a, mmfile_t *file_b,\n \treturn xdi_diff(file_a, file_b, &xpp, &xecfg, &ecb);\n }\n \n+static const char *get_next_line(const char *start, const char *end)\n+{\n+\tconst char *nl = memchr(start, '\\n', end - start);\n+\n+\treturn nl ? nl + 1 : end;\n+}\n+\n+static int find_line_starts(int **line_starts, const char *buf,\n+\t\t\t    unsigned long len)\n+{\n+\tconst char *end = buf + len;\n+\tconst char *p;\n+\tint *lineno;\n+\tint num = 0;\n+\n+\tfor (p = buf; p < end; p = get_next_line(p, end))\n+\t\tnum++;\n+\n+\tALLOC_ARRAY(*line_starts, num + 1);\n+\tlineno = *line_starts;\n+\n+\tfor (p = buf; p < end; p = get_next_line(p, end))\n+\t\t*lineno++ = p - buf;\n+\n+\t*lineno = len;\n+\n+\treturn num;\n+}\n+\n+static void fill_origin_fingerprints(struct blame_origin *o, mmfile_t *file)\n+{\n+\tint *line_starts;\n+\n+\tif (o->fingerprints)\n+\t\treturn;\n+\to->num_lines = find_line_starts(&line_starts, o->file.ptr,\n+\t\t\t\t\to->file.size);\n+\t/* TODO: Will fill in fingerprints in a future commit */\n+\to->fingerprints = xcalloc(sizeof(struct fingerprint), o->num_lines);\n+\tfree(line_starts);\n+}\n+\n+static void drop_origin_fingerprints(struct blame_origin *o)\n+{\n+\tif (o->fingerprints) {\n+\t\to->num_lines = 0;\n+\t\tFREE_AND_NULL(o->fingerprints);\n+\t}\n+}\n+\n /*\n  * Given an origin, prepare mmfile_t structure to be used by the\n  * diff machinery\n  */\n static void fill_origin_blob(struct diff_options *opt,\n-\t\t\t     struct blame_origin *o, mmfile_t *file, int *num_read_blob)\n+\t\t\t     struct blame_origin *o, mmfile_t *file,\n+\t\t\t     int *num_read_blob, int fill_fingerprints)\n {\n \tif (!o->file.ptr) {\n \t\tenum object_type type;\n@@ -340,11 +391,14 @@ static void fill_origin_blob(struct diff_options *opt,\n \t}\n \telse\n \t\t*file = o->file;\n+\tif (fill_fingerprints)\n+\t\tfill_origin_fingerprints(o, file);\n }\n \n static void drop_origin_blob(struct blame_origin *o)\n {\n \tFREE_AND_NULL(o->file.ptr);\n+\tdrop_origin_fingerprints(o);\n }\n \n /*\n@@ -1136,8 +1190,10 @@ static void pass_blame_to_parent(struct blame_scoreboard *sb,\n \td.ignore_diffs = ignore_diffs;\n \td.dstq = &newdest; d.srcq = &target->suspects;\n \n-\tfill_origin_blob(&sb->revs->diffopt, parent, &file_p, &sb->num_read_blob);\n-\tfill_origin_blob(&sb->revs->diffopt, target, &file_o, &sb->num_read_blob);\n+\tfill_origin_blob(&sb->revs->diffopt, parent, &file_p,\n+\t\t\t &sb->num_read_blob, ignore_diffs);\n+\tfill_origin_blob(&sb->revs->diffopt, target, &file_o,\n+\t\t\t &sb->num_read_blob, ignore_diffs);\n \tsb->num_get_patch++;\n \n \tif (diff_hunks(&file_p, &file_o, blame_chunk_cb, &d, sb->xdl_opts))\n@@ -1348,7 +1404,8 @@ static void find_move_in_parent(struct blame_scoreboard *sb,\n \tif (!unblamed)\n \t\treturn; /* nothing remains for this target */\n \n-\tfill_origin_blob(&sb->revs->diffopt, parent, &file_p, &sb->num_read_blob);\n+\tfill_origin_blob(&sb->revs->diffopt, parent, &file_p,\n+\t\t\t &sb->num_read_blob, 0);\n \tif (!file_p.ptr)\n \t\treturn;\n \n@@ -1477,7 +1534,8 @@ static void find_copy_in_parent(struct blame_scoreboard *sb,\n \t\t\tnorigin = get_origin(parent, p->one->path);\n \t\t\toidcpy(&norigin->blob_oid, &p->one->oid);\n \t\t\tnorigin->mode = p->one->mode;\n-\t\t\tfill_origin_blob(&sb->revs->diffopt, norigin, &file_p, &sb->num_read_blob);\n+\t\t\tfill_origin_blob(&sb->revs->diffopt, norigin, &file_p,\n+\t\t\t\t\t &sb->num_read_blob, 0);\n \t\t\tif (!file_p.ptr)\n \t\t\t\tcontinue;\n \n@@ -1816,37 +1874,14 @@ void assign_blame(struct blame_scoreboard *sb, int opt)\n \t}\n }\n \n-static const char *get_next_line(const char *start, const char *end)\n-{\n-\tconst char *nl = memchr(start, '\\n', end - start);\n-\treturn nl ? nl + 1 : end;\n-}\n-\n /*\n  * To allow quick access to the contents of nth line in the\n  * final image, prepare an index in the scoreboard.\n  */\n static int prepare_lines(struct blame_scoreboard *sb)\n {\n-\tconst char *buf = sb->final_buf;\n-\tunsigned long len = sb->final_buf_size;\n-\tconst char *end = buf + len;\n-\tconst char *p;\n-\tint *lineno;\n-\tint num = 0;\n-\n-\tfor (p = buf; p < end; p = get_next_line(p, end))\n-\t\tnum++;\n-\n-\tALLOC_ARRAY(sb->lineno, num + 1);\n-\tlineno = sb->lineno;\n-\n-\tfor (p = buf; p < end; p = get_next_line(p, end))\n-\t\t*lineno++ = p - buf;\n-\n-\t*lineno = len;\n-\n-\tsb->num_lines = num;\n+\tsb->num_lines = find_line_starts(&sb->lineno, sb->final_buf,\n+\t\t\t\t\t sb->final_buf_size);\n \treturn sb->num_lines;\n }\n \ndiff --git a/blame.h b/blame.h\nindex 53df8b4c5b3f..5dd877bb78fc 100644\n--- a/blame.h\n+++ b/blame.h\n@@ -51,6 +51,8 @@ struct blame_origin {\n \t */\n \tstruct blame_entry *suspects;\n \tmmfile_t file;\n+\tint num_lines;\n+\tvoid *fingerprints;\n \tstruct object_id blob_oid;\n \tunsigned mode;\n \t/* guilty gets set when shipping any suspects to the final\n-- \n2.21.0.392.gf8f6787159e-goog\n\n"},{"id":"373592","messageId":"20190410162409.117264-7-brho@google.com","threadId":"50912","inReplyTo":"20190410162409.117264-1-brho@google.com","subject":"[PATCH v6 6/6] blame: use a fingerprint heuristic to match ignored lines","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-10T16:24:09Z","receivedAt":"2019-04-10T16:24:50Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"This replaces the heuristic used to identify lines from ignored commits\nwith one that finds likely candidate lines in the parent's version of\nthe file.\n\nThe old heuristic simply assigned lines in the target to the same line\nnumber (plus offset) in the parent.  The new function uses a\nfingerprinting algorithm to detect similarity between lines.\n\nThe fingerprint code and the idea to use them for blame came from\nMichael Platings <michael@platin.gs>.\n\nFor each line changed in the target, i.e. in a blame_entry touched by a\ntarget's diff, guess_line_blames() finds the best line in the parent,\nabove a magic threshold.  Ties are broken by proximity of the parent\nline number to the target's line.\n\nWe actually make two passes.  The first pass checks in the diff chunk\nassociated with the blame entry - specifically from blame_chunk().\nOften times, those diff chunks are small; any 'context' in a normal diff\nchunk is broken up into multiple calls to blame_chunk().  We make a\nsecond pass over the entire parent, with a slightly higher threshold.\n\nHere's an example of the difference the fingerprinting makes.  Consider\na file with four commits:\n\n\tcommit-a 11) void new_func_1(void *x, void *y);\n\tcommit-b 12) void new_func_2(void *x, void *y);\n\tcommit-c 13) some_line_c\n\tcommit-d 14) some_line_d\n\nAfter a commit 'X', we have:\n\n\tcommit-X 11) void new_func_1(void *x,\n\tcommit-X 12)                 void *y);\n\tcommit-X 13) void new_func_2(void *x,\n\tcommit-X 14)                 void *y);\n\tcommit-c 15) some_line_c\n\tcommit-d 16) some_line_d\n\nWhen we blame-ignored with the old algorithm, we get:\n\n\tcommit-a 11) void new_func_1(void *x,\n\tcommit-b 12)                 void *y);\n\t00000000 13) void new_func_2(void *x,\n\t00000000 14)                 void *y);\n\tcommit-c 15) some_line_c\n\tcommit-d 16) some_line_d\n\nWhere commit-b is blamed for 12 instead of 13.  With the fingerprint\nalgorithm, we get:\n\n\tcommit-a 11) void new_func_1(void *x,\n\tcommit-b 12)                 void *y);\n\tcommit-b 13) void new_func_2(void *x,\n\tcommit-b 14)                 void *y);\n\tcommit-c 15) some_line_c\n\tcommit-d 16) some_line_d\n\nNote both lines 12 and 14 are given to commit b.  Their match is above\nthe FINGERPRINT_CHUNK_THRESHOLD, and they tied.  Specifically, parent\nlines 11 and 12 both match these lines.  The algorithm chose parent line\n12, since that was closest to the target line numbers of 12 and 14.\n\nIf we increase the threshold, say to 10, those two lines won't match,\nand will be treated as 'unblamable.'\n\nFor an example of scanning the entire parent for a match, consider:\n\n\tcommit-a 30) #include <sys/header_a.h>\n\tcommit-b 31) #include <header_b.h>\n\tcommit-c 32) #include <header_c.h>\n\nThen commit X alphabetizes them:\n\n\tcommit-X 30) #include <header_b.h>\n\tcommit-X 31) #include <header_c.h>\n\tcommit-X 32) #include <sys/header_a.h>\n\nIf we just check the parent's chunk (i.e. the first pass), we'd get:\n\n\tcommit-b 30) #include <header_b.h>\n\tcommit-c 31) #include <header_c.h>\n\t00000000 32) #include <sys/header_a.h>\n\nThat's because commit X consists of two chunks: one chunk is removing\nsys/header_a.h, then some context, and the second chunk is adding\nsys/header_a.h.\n\nIf we scan the entire parent file, we get:\n\n\tcommit-b 30) #include <header_b.h>\n\tcommit-c 31) #include <header_c.h>\n\tcommit-a 32) #include <sys/header_a.h>\n\nSuggested-by: Michael Platings <michael@platin.gs>\nSigned-off-by: Barret Rhoden <brho@google.com>\n---\n blame.c | 140 ++++++++++++++++++++++++++++++++++++++++++++++++++++----\n 1 file changed, 131 insertions(+), 9 deletions(-)\n\ndiff --git a/blame.c b/blame.c\nindex a42dff80b1a5..da2d664b38af 100644\n--- a/blame.c\n+++ b/blame.c\n@@ -340,6 +340,84 @@ static int find_line_starts(int **line_starts, const char *buf,\n \treturn num;\n }\n \n+struct fingerprint {\n+\tstruct hashmap map;\n+\tstruct hashmap_entry *entries;\n+};\n+\n+static void get_fingerprint(struct fingerprint *result,\n+\t\t\t    const char *line_begin,\n+\t\t\t    const char *line_end)\n+{\n+\tunsigned int hash;\n+\tchar c0, c1;\n+\tconst char *p;\n+\tint map_entry_count = line_end - line_begin - 1;\n+\tstruct hashmap_entry *entry = xcalloc(map_entry_count,\n+\t\t\t\t\t      sizeof(struct hashmap_entry));\n+\n+\thashmap_init(&result->map, NULL, NULL, map_entry_count);\n+\tresult->entries = entry;\n+\tfor (p = line_begin; p + 1 < line_end; ++p, ++entry) {\n+\t\tc0 = *p;\n+\t\tc1 = *(p + 1);\n+\t\t/* Ignore whitespace pairs */\n+\t\tif (isspace(c0) && isspace(c1))\n+\t\t\tcontinue;\n+\t\thash = tolower(c0) | (tolower(c1) << 8);\n+\t\thashmap_entry_init(entry, hash);\n+\t\thashmap_put(&result->map, entry);\n+\t}\n+}\n+\n+static void free_fingerprint(struct fingerprint *f)\n+{\n+\thashmap_free(&f->map, 0);\n+\tfree(f->entries);\n+}\n+\n+static int fingerprint_similarity(struct fingerprint *a,\n+\t\t\t\t  struct fingerprint *b)\n+{\n+\tint intersection = 0;\n+\tstruct hashmap_iter iter;\n+\tstruct hashmap_entry *entry;\n+\n+\thashmap_iter_init(&b->map, &iter);\n+\n+\twhile ((entry = hashmap_iter_next(&iter))) {\n+\t\tif (hashmap_get(&a->map, entry, NULL))\n+\t\t\t++intersection;\n+\t}\n+\treturn intersection;\n+}\n+\n+static void get_line_fingerprints(struct fingerprint *fingerprints,\n+\t\t\t\t  const char *content,\n+\t\t\t\t  const int *line_starts,\n+\t\t\t\t  int first_line,\n+\t\t\t\t  int nr_lines)\n+{\n+\tint i;\n+\n+\tline_starts += first_line;\n+\tfor (i = 0; i < nr_lines; ++i) {\n+\t\tconst char *linestart = content + line_starts[i];\n+\t\tconst char *lineend = content + line_starts[i + 1];\n+\n+\t\tget_fingerprint(fingerprints + i, linestart, lineend);\n+\t}\n+}\n+\n+static void free_chunk_fingerprints(struct fingerprint *fingerprints,\n+\t\t\t\t    int nr_fingerprints)\n+{\n+\tint i;\n+\n+\tfor (i = 0; i < nr_fingerprints; i++)\n+\t\tfree_fingerprint(&fingerprints[i]);\n+}\n+\n static void fill_origin_fingerprints(struct blame_origin *o, mmfile_t *file)\n {\n \tint *line_starts;\n@@ -348,14 +426,16 @@ static void fill_origin_fingerprints(struct blame_origin *o, mmfile_t *file)\n \t\treturn;\n \to->num_lines = find_line_starts(&line_starts, o->file.ptr,\n \t\t\t\t\to->file.size);\n-\t/* TODO: Will fill in fingerprints in a future commit */\n \to->fingerprints = xcalloc(sizeof(struct fingerprint), o->num_lines);\n+\tget_line_fingerprints(o->fingerprints, o->file.ptr, line_starts,\n+\t\t\t      0, o->num_lines);\n \tfree(line_starts);\n }\n \n static void drop_origin_fingerprints(struct blame_origin *o)\n {\n \tif (o->fingerprints) {\n+\t\tfree_chunk_fingerprints(o->fingerprints, o->num_lines);\n \t\to->num_lines = 0;\n \t\tFREE_AND_NULL(o->fingerprints);\n \t}\n@@ -935,27 +1015,69 @@ static int are_lines_adjacent(struct blame_line_tracker *first,\n \t       first->s_lno + 1 == second->s_lno;\n }\n \n+static void scan_parent_range(struct fingerprint *p_fps,\n+\t\t\t      struct fingerprint *t_fps, int t_idx,\n+\t\t\t      int from, int nr_lines,\n+\t\t\t      int *best_sim_val, int *best_sim_idx)\n+{\n+\tint sim, p_idx;\n+\n+\tfor (p_idx = from; p_idx < from + nr_lines; p_idx++) {\n+\t\tsim = fingerprint_similarity(&t_fps[t_idx], &p_fps[p_idx]);\n+\t\tif (sim < *best_sim_val)\n+\t\t\tcontinue;\n+\t\t/* Break ties with the closest-to-target line number */\n+\t\tif (sim == *best_sim_val && *best_sim_idx != -1 &&\n+\t\t    abs(*best_sim_idx - t_idx) < abs(p_idx - t_idx))\n+\t\t\tcontinue;\n+\t\t*best_sim_val = sim;\n+\t\t*best_sim_idx = p_idx;\n+\t}\n+}\n+\n /*\n- * This cheap heuristic assigns lines in the chunk to their relative location in\n- * the parent's chunk.  Any additional lines are left with the target.\n+ * The CHUNK threshold is for how similar we must be within a diff chunk, which\n+ * is typically the adjacent '-' and '+' sections in a diff, separated by the\n+ * ' ' context.\n+ *\n+ * We have a greater threshold for similarity for lines in any part of the\n+ * parent's file.  If no line in the parent meets the appropriate threshold,\n+ * then the blame_entry will stay with the target and be considered\n+ * 'unblamable'.\n  */\n+#define FINGERPRINT_CHUNK_THRESHOLD\t1\n+#define FINGERPRINT_FILE_THRESHOLD\t10\n+\n static void guess_line_blames(struct blame_entry *e,\n \t\t\t      struct blame_origin *parent,\n \t\t\t      struct blame_origin *target,\n \t\t\t      int offset, int parent_slno, int parent_len,\n \t\t\t      struct blame_line_tracker *line_blames)\n {\n-\tint i, parent_idx;\n+\tint i, target_idx;\n \n \tfor (i = 0; i < e->num_lines; i++) {\n-\t\tparent_idx = e->s_lno + i + offset;\n-\t\tif (parent_slno <= parent_idx &&\n-\t\t    parent_idx < parent_slno + parent_len) {\n+\t\tint best_val = FINGERPRINT_CHUNK_THRESHOLD;\n+\t\tint best_idx = -1;\n+\n+\t\ttarget_idx = e->s_lno + i;\n+\t\tscan_parent_range(parent->fingerprints,\n+\t\t\t\t  target->fingerprints, target_idx,\n+\t\t\t\t  parent_slno, parent_len,\n+\t\t\t\t  &best_val, &best_idx);\n+\t\tif (best_idx == -1) {\n+\t\t\tbest_val = FINGERPRINT_FILE_THRESHOLD;\n+\t\t\tscan_parent_range(parent->fingerprints,\n+\t\t\t\t\t  target->fingerprints, target_idx,\n+\t\t\t\t\t  0, parent->num_lines,\n+\t\t\t\t\t  &best_val, &best_idx);\n+\t\t}\n+\t\tif (best_idx >= 0) {\n \t\t\tline_blames[i].is_parent = 1;\n-\t\t\tline_blames[i].s_lno = parent_idx;\n+\t\t\tline_blames[i].s_lno = best_idx;\n \t\t} else {\n \t\t\tline_blames[i].is_parent = 0;\n-\t\t\tline_blames[i].s_lno = e->s_lno + i;\n+\t\t\tline_blames[i].s_lno = target_idx;\n \t\t}\n \t}\n }\n-- \n2.21.0.392.gf8f6787159e-goog\n\n"},{"id":"373607","messageId":"878swhfzxb.fsf@evledraar.gmail.com","threadId":"50912","inReplyTo":"20190410162409.117264-4-brho@google.com","subject":"Re: [PATCH v6 3/6] blame: add the ability to ignore commits and their changes","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-04-10T19:00:48Z","receivedAt":"2019-04-10T19:00:55Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Apr 10 2019, Barret Rhoden wrote:\n\n(Just skimming)\n\n> revisions for commits that perform mass reformatting, and their users\n> have the optional to ignore all of the commits in that file.\n\ns/have the optional/have the option/\n\n> +--ignore-revs-file <file>::\n> +\tIgnore revisions listed in `file`, one unabbreviated object name per line.\n> +\tWhitespace and comments beginning with `#` are ignored.\n\nMaybe just say \"Ignore revisions listed in `file`, which is expected to\nbe in the same format as an `fsck.skipList`.\".\n\n> +\tthe `blame.ignoreRevsFile` config option.  An empty file name, `\"\"`, will\n> +\tclear the list of revs from previously processed files.\n\nMaybe I haven't read this carefully enough but the use-case for this\ndoesn't seem to be explained, you need this for the option, but the\nconfig file too? If I want to override fsck.skipList I do\n`fsck.skipList=/dev/zero`. Isn't that enough for this use-case without\nintroducing config state-machine magic?\n\n> +\tsplit[0].unblamable = e->unblamable;\n> +\tsplit[1].unblamable = e->unblamable;\n> +\tsplit[2].unblamable = e->unblamable;\n\nI wonder what the comfort level for people in general is before turning\nthis sort of thing into a for-loop, 4? :)\n\n> +\tnr_lines = e->num_lines;\t// e changes in the loop\n\nA C++-like trailing comment.\n\n> +\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n> +\tgit rev-parse X >expect &&\n> +\ttest_cmp expect actual &&\n> +\n> +\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n> +\tgit rev-parse X >expect &&\n> +\ttest_cmp expect actual\n\nThe grep here is a bug. See my 4abf20f004 (\"tests: fix unportable \"\\?\"\nand \"\\+\" regex syntax\", 2019-02-21).\n"},{"id":"373608","messageId":"877ec1fzqy.fsf@evledraar.gmail.com","threadId":"50912","inReplyTo":"20190410162409.117264-2-brho@google.com","subject":"Re: [PATCH v6 1/6] Move init_skiplist() outside of fsck","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-04-10T19:04:37Z","receivedAt":"2019-04-10T19:04:43Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Apr 10 2019, Barret Rhoden wrote:\n\n> init_skiplist() took a file consisting of SHA-1s and comments and added\n> the objects to an oidset.  This functionality is useful for other\n> commands.\n\nThis change would be much easier to review if you led with a commit\nwhere you s/Invalid SHA-1/invalid object name/ (lower-case while we're\nat it), s/skip list/object name/ etc, and did that rename of the \"hash\"\nto \"name\" variable if you're so inclined.\n\nThen you'd end up with a small refactoring change that changes the tests\n(or even just make the tests grep for e.g. \"Could not open.*:\ndoes-not-exist\" instead), and the moving of the function would be\nentirely caught by the rename detection.\n"},{"id":"373809","messageId":"xmqqo959w8pq.fsf@gitster-ct.c.googlers.com","threadId":"50912","inReplyTo":"20190410162409.117264-5-brho@google.com","subject":"Re: [PATCH v6 4/6] blame: add config options to handle output for ignored lines","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-04-14T03:45:37Z","receivedAt":"2019-04-14T03:45:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Barret Rhoden <brho@google.com> writes:\n\n> Sometimes we are unable to even guess at what commit touched a line.\n> These lines are 'unblamable.'  The second option,\n> blame.maskIgnoredUnblamables, will zero the hash of any unblamable line.\n>\n> For example, say we ignore e5e8d36d04cbe:\n> \te5e8d36d04cbe (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\n> appears as:\n> \t0000000000000 (Barret Rhoden  2016-04-11 13:57:54 -0400 26)\n\nWouldn't this make it impossible to tell between what's done by such\na commit that was marked to be ignored, and what's done locally only\nin the working tree, which the users have long accustomed to see\nwith the ^0*$ object name?  I think it would make a lot more sense\nto show the object name of the \"ignored\" commit, which would be\nrecognizable by the user who fed such an object name to the command\nin the first place.  Alternatively, perhaps the same idea as replacing\none of the hexdigits with '*' used by the other configuration can be\napplied to this as well?\n"},{"id":"373810","messageId":"xmqqk1fxw8ad.fsf@gitster-ct.c.googlers.com","threadId":"50912","inReplyTo":"20190410162409.117264-7-brho@google.com","subject":"Re: [PATCH v6 6/6] blame: use a fingerprint heuristic to match ignored lines","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-04-14T03:54:50Z","receivedAt":"2019-04-14T03:54:56Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Barret Rhoden <brho@google.com> writes:\n\n> This replaces the heuristic used to identify lines from ignored commits\n> with one that finds likely candidate lines in the parent's version of\n> the file.\n>\n> The old heuristic simply assigned lines in the target to the same line\n> number (plus offset) in the parent.  The new function uses a\n> fingerprinting algorithm to detect similarity between lines.\n>\n> The fingerprint code and the idea to use them for blame came from\n> Michael Platings <michael@platin.gs>.\n>\n> For each line changed in the target, i.e. in a blame_entry touched by a\n> target's diff, guess_line_blames() finds the best line in the parent,\n> above a magic threshold.  Ties are broken by proximity of the parent\n> line number to the target's line.\n>\n> We actually make two passes.  The first pass checks in the diff chunk\n> associated with the blame entry - specifically from blame_chunk().\n> Often times, those diff chunks are small; any 'context' in a normal diff\n> chunk is broken up into multiple calls to blame_chunk().  We make a\n> second pass over the entire parent, with a slightly higher threshold.\n\nTwo thoughts.\n\n - Unless the 'old heuristic' is still available as an option after\n   this step, a series that first begins with the 'old heuristic'\n   and then later replaces it with the 'new heuristic' feels\n   somewhat wasteful of reviewer resources, as the 'old heuristic'\n   does not contribute an iota to the end result.\n\n   It is OK while the series is still in RFC/WIP stage, though.  But\n   because I got an impression that this is close to completion, so...\n\n - I wonder if the hash used here can replace what is used in\n   diffcore-delta.c as an improvement (or obviously vice versa), as\n   using two (or more) ad-hoc fingerprinting function without having\n   a clear reason why we need two instead of a unified one feels\n   like a bad idea.\n\n"},{"id":"373814","messageId":"CAJDYR9RHb89mjT65XVERJfo3cTySi++ZAwOFftBtyXkqfC=JOQ@mail.gmail.com","threadId":"50912","inReplyTo":"xmqqk1fxw8ad.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v6 6/6] blame: use a fingerprint heuristic to match ignored lines","fromName":"Michael Platings","fromEmail":"michael@platin.gs","sentAt":"2019-04-14T09:41:26Z","receivedAt":"2019-04-14T09:42:04Z","isPatch":true,"sender":{"key":"michael@platin.gs","avatar":"https://avatars.githubusercontent.com/u/1112348?v=4"},"body":">  - I wonder if the hash used here can replace what is used in\n>    diffcore-delta.c as an improvement (or obviously vice versa), as\n>    using two (or more) ad-hoc fingerprinting function without having\n>    a clear reason why we need two instead of a unified one feels\n>    like a bad idea.\n\nHi Junio,\nIf I understand correctly, the algorithm in diffcore-delta.c is\nintended to match files that contain identical lines (or 64-byte\nchunks). The fingerprinting that Barret & I are talking about is\nintended to match lines that contain identical byte pairs.\nWith significant refactoring, you could make the diffcore-delta\nalgorithm apply in both cases but I think the end result would be\nlonger and more complicated than keeping the two separate.\nUnlike hashing a line, hashing a byte pair is trivial. Unlike hashing\nlines, all except the first and last bytes are included in two\n\"hashes\" - \"hello\" is hashed to \"he\", \"el\", \"ll\", \"lo\".\nSo based on my limited understanding of diffcore-delta.c I think the\ntwo are algorithms are sufficiently different in intent and in\nimplementation that it's appropriate to keep them separate.\n\nRegarding the \"old heuristic\" I think there may still be a use case\nfor that but I'll expand on that later.\n\nThanks,\n-Michael\n"},{"id":"373815","messageId":"CAJDYR9S8XFH=JnQX8WcfgOZ7cr+X6kk45k9g8t3u5aP5wwdu0Q@mail.gmail.com","threadId":"50912","inReplyTo":"xmqqo959w8pq.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v6 4/6] blame: add config options to handle output for ignored lines","fromName":"Michael Platings","fromEmail":"michael@platin.gs","sentAt":"2019-04-14T10:09:55Z","receivedAt":"2019-04-14T10:10:09Z","isPatch":true,"sender":{"key":"michael@platin.gs","avatar":"https://avatars.githubusercontent.com/u/1112348?v=4"},"body":"On Sun, 14 Apr 2019 at 04:45, Junio C Hamano <gitster@pobox.com> wrote:\n> Wouldn't this make it impossible to tell between what's done by such\n> a commit that was marked to be ignored, and what's done locally only\n> in the working tree, which the users have long accustomed to see\n> with the ^0*$ object name?  I think it would make a lot more sense\n> to show the object name of the \"ignored\" commit, which would be\n> recognizable by the user who fed such an object name to the command\n> in the first place.  Alternatively, perhaps the same idea as replacing\n> one of the hexdigits with '*' used by the other configuration can be\n> applied to this as well?\n\nI had the same objection to zeroing out hashes, but this option is off\nby default so I think it's OK.\nIf you enable both blame.markIgnoredLines and\nblame.maskIgnoredUnblamables then the hash does appear as\n\"*0000000000\" like you suggest. I think it's appropriate that the '*'\nis only added if you opt in with the markIgnoredLines option.\n\nIf you only enable blame.markIgnoredLines then the hash for\n\"unblamable\" lines appears as e.g. \"*3252488f5\" - this doesn't seem\nright to me because the commit *wasn't* ignored, it is in fact the\ncommit in which that line was added. I think '*' should denote \"this\ninformation may be inaccurate\" as that's what a typical user needs to\nbe aware of. However given that \"unblamable\" lines tend to be either\nempty or a single character I'm not going to insist :)\n"},{"id":"373816","messageId":"xmqqbm18x4tt.fsf@gitster-ct.c.googlers.com","threadId":"50912","inReplyTo":"CAJDYR9S8XFH=JnQX8WcfgOZ7cr+X6kk45k9g8t3u5aP5wwdu0Q@mail.gmail.com","subject":"Re: [PATCH v6 4/6] blame: add config options to handle output for ignored lines","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-04-14T10:24:14Z","receivedAt":"2019-04-14T10:24:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Michael Platings <michael@platin.gs> writes:\n\n> If you only enable blame.markIgnoredLines then the hash for\n> \"unblamable\" lines appears as e.g. \"*3252488f5\" - this doesn't seem\n> right to me because the commit *wasn't* ignored,\n\nI think you misunderstood me.  I was merely suggesting to use the\napproach to mark the line in a way other than using the NULLed out\nobject name that has been reserved for something totally different,\nand hinting with \"the same *idea*\".\n\nAnd that idea is not even original to this series; the \"^\" marker\nthat is used to say \"the line is attributed to this commit, but that\nmay only be because you blamed with commit range A..B and we reached\nthe bottom of the range---if you dug further, you might find the\nline originates from another commit\" is the origin of the same idea,\nand this topic borrows it and uses a different mark, i.e. '*', for\nthe \"we are not certain---take this with grain of salt\" mark.\n\nIf you ended up hitting the commit the user wanted to ignore,\nperhaps you can find another character that is different from '^' or\n'*' and use that, following the same idea.\n\nThat is what I meant.  So you shouldn't be worried about using the\nsame '*' making the result ambiguous.\n\nBy the way, a configuration only feature is something we usually do\nnot accept.  A feature must be guarded with --command-line-option\nand then optionally can have a corresponding configuration once the\noption proves to be useful enough that it becomes useful to be able\nto say \"in this repository (or to this user), the feature is on by\ndefault\".\n"},{"id":"373818","messageId":"CAJDYR9Q-ixsxWyMrm7aCojTv33SOj3+ALPwJYo9DJE7vLU=DEA@mail.gmail.com","threadId":"50912","inReplyTo":"878swhfzxb.fsf@evledraar.gmail.com","subject":"Re: [PATCH v6 3/6] blame: add the ability to ignore commits and their changes","fromName":"Michael Platings","fromEmail":"michael@platin.gs","sentAt":"2019-04-14T10:42:08Z","receivedAt":"2019-04-14T10:42:21Z","isPatch":true,"sender":{"key":"michael@platin.gs","avatar":"https://avatars.githubusercontent.com/u/1112348?v=4"},"body":"> > +     the `blame.ignoreRevsFile` config option.  An empty file name, `\"\"`, will\n> > +     clear the list of revs from previously processed files.\n>\n> Maybe I haven't read this carefully enough but the use-case for this\n> doesn't seem to be explained, you need this for the option, but the\n> config file too? If I want to override fsck.skipList I do\n> `fsck.skipList=/dev/zero`. Isn't that enough for this use-case without\n> introducing config state-machine magic?\n\nThe difference between blame.ignoreRevsFile and fsck.skipList is that\nignoreRevsFile can be specified repeatedly. This is useful if you have\none file listing reformatting commits, another listing renaming\ncommits etc. Or maybe a checked-in list of commits to ignore, and a\npersonal list of commits to ignore. However sometimes you're going to\nwant to *not* ignore those commits, so you need a way to discard the\npreviously specified options. To accommodate all operating systems an\nempty string seems the best way to do this.\n"},{"id":"373819","messageId":"CAJDYR9TRk99Kwq5S7udVqYsXnupGD=t3o_Ss8ewvwWuTQOy_YQ@mail.gmail.com","threadId":"50912","inReplyTo":"xmqqbm18x4tt.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v6 4/6] blame: add config options to handle output for ignored lines","fromName":"Michael Platings","fromEmail":"michael@platin.gs","sentAt":"2019-04-14T11:27:28Z","receivedAt":"2019-04-14T11:27:42Z","isPatch":true,"sender":{"key":"michael@platin.gs","avatar":"https://avatars.githubusercontent.com/u/1112348?v=4"},"body":"On Sun, 14 Apr 2019 at 11:24, Junio C Hamano <gitster@pobox.com> wrote:\n> > If you only enable blame.markIgnoredLines then the hash for\n> > \"unblamable\" lines appears as e.g. \"*3252488f5\" - this doesn't seem\n> > right to me because the commit *wasn't* ignored,\n>\n> I think you misunderstood me.  I was merely suggesting to use the\n> approach to mark the line in a way other than using the NULLed out\n> object name that has been reserved for something totally different,\n> and hinting with \"the same *idea*\".\n\nHi Junio, that paragraph wasn't targetted at yourself, more a comment\non the functionality as it exists in the latest patch series. Sorry\nfor not making that clear.\n\n> the \"^\" marker\n> that is used to say \"the line is attributed to this commit, but that\n> may only be because you blamed with commit range A..B and we reached\n> the bottom of the range---if you dug further, you might find the\n> line originates from another commit\" is the origin of the same idea,\n> and this topic borrows it and uses a different mark, i.e. '*', for\n> the \"we are not certain---take this with grain of salt\" mark.\n\nSo it sounds like we have many types of blame to consider:\n\n1) This commit is truly the last one to touch this line, and you\ndidn't ask to ignore it.\n2) This commit is truly the last one to touch this line, but you asked\nto ignore it (AKA \"unblamable\").\n3) This commit is at the bottom of the range of commits (^)\n4) The \"true\" commit was ignored but we guess this is the one you're\nactually interested in (*)\n5) The \"true\" commit was ignored and we've reached the bottom of the\nrange of commits (^*)?\n6) This commit is at the bottom of the range of commits, and you asked\nto ignore it.\n\n> If you ended up hitting the commit the user wanted to ignore,\n> perhaps you can find another character that is different from '^' or\n> '*' and use that, following the same idea.\n\nI personally don't find the \"unblamable\" lines interesting enough to\njustify giving them a symbol. But if Barret strongly feels that such\nlines should get a '*' then I won't fight it - these lines tend to be\nas simple as \"}\".\n\n> By the way, a configuration only feature is something we usually do\n> not accept.  A feature must be guarded with --command-line-option\n> and then optionally can have a corresponding configuration once the\n> option proves to be useful enough that it becomes useful to be able\n> to say \"in this repository (or to this user), the feature is on by\n> default\".\n\nIn that case we definitely need a --mark-ignored-lines option to git\nblame, and I would strongly prefer that we also keep the\nblame.markIgnoredLines option as I for one will be switching it on.\n"},{"id":"373829","messageId":"CAJDYR9SL9JCJjdARejV=NCf9GYn72=bfszXx84iDc416sZm31A@mail.gmail.com","threadId":"50912","inReplyTo":"20190410162409.117264-1-brho@google.com","subject":"Re: [PATCH v6 0/6] blame: add the ability to ignore commits","fromName":"Michael Platings","fromEmail":"michael@platin.gs","sentAt":"2019-04-14T21:10:30Z","receivedAt":"2019-04-14T21:10:43Z","isPatch":true,"sender":{"key":"michael@platin.gs","avatar":"https://avatars.githubusercontent.com/u/1112348?v=4"},"body":"Hi Barret,\n\nThis works pretty well for the typical reformatting use case now. I've\nrun it over every commit of every .c file in the git project root,\nboth forwards and backwards with every combination of -w/-M/-C and\ncan't get it to crash so I think it's good in that respect.\n\nHowever, it can still attribute lines to the wrong parent line. See\nhttps://pypi.org/project/autopep8/#usage for an example reformatting\nthat it gets a bit confused on. The patch I submitted handles this\ncase correctly because it uses information about the more similar\nlines to decide how more ambiguous lines should be matched.\n\nYou also gave an example of:\n\n        commit-a 11) void new_func_1(void *x, void *y);\n        commit-b 12) void new_func_2(void *x, void *y);\n\nBeing reformatted to:\n\n        commit-a 11) void new_func_1(void *x,\n        commit-b 12)                 void *y);\n        commit-b 13) void new_func_2(void *x,\n        commit-b 14)                 void *y);\n\nThe patch I submitted handles this case correctly, assigning line 12\nto commit-a because it scales the parent line numbers according to the\nrelative diff chunk sizes instead of assuming a 1-1 mapping.\n\nSo I do ask that you incorporate more of my patch, including the test\ncode. It is more complex but I hope this demonstrates that there are\nreasons for that. Happy to provide more examples or explanation if it\nwould help. On the other hand if you have examples where it falls\nshort then I'd be interested to know.\n\nThe other major use case that I'm interested in is renaming. In this\ncase, the git-hyper-blame approach of mapping line numbers 1-1 works\nperfectly. Here's an example. Before:\n\n        commit-a 11) Position MyClass::location(Offset O) {\n        commit-b 12)    return P + O;\n        commit-c 13) }\n\nAfter:\n\n        commit-a 11) Position MyClass::location(Offset offset) {\n        commit-a 12)    return position + offset;\n        commit-c 13) }\n\nWith the fuzzy matching, line 12 gets incorrectly matched to parent\nline 11 because the similarity of \"position\" and \"offset\" outweighs\nthe similarity of \"return\". I'm considering adding even more\ncomplexity to my patch such that parts of a line that have already\nbeen matched can't be matched again by other lines.\n\nBut the other possibility is that we let the user choose the\nheuristic. For a commit where they know that line numbers haven't\nchanged they could choose 1-1 matching, while for a reformatting\ncommit they could use fuzzy matching. I welcome your thoughts.\n\n-Michael\n"},{"id":"373867","messageId":"9439c697-246f-3bcb-4d34-85099e577e8b@google.com","threadId":"50912","inReplyTo":"CAJDYR9SL9JCJjdARejV=NCf9GYn72=bfszXx84iDc416sZm31A@mail.gmail.com","subject":"Re: [PATCH v6 0/6] blame: add the ability to ignore commits","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-15T13:23:37Z","receivedAt":"2019-04-15T13:23:44Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"Hi Michael -\n\nOn 4/14/19 5:10 PM, Michael Platings wrote:\n> Hi Barret,\n> \n> This works pretty well for the typical reformatting use case now. I've\n> run it over every commit of every .c file in the git project root,\n> both forwards and backwards with every combination of -w/-M/-C and\n> can't get it to crash so I think it's good in that respect.\n> \n> However, it can still attribute lines to the wrong parent line. See\n> https://pypi.org/project/autopep8/#usage for an example reformatting\n> that it gets a bit confused on. The patch I submitted handles this\n> case correctly because it uses information about the more similar\n> lines to decide how more ambiguous lines should be matched.\n\nYeah - I ran your tests against it and noticed a few cases weren't handled.\n\n> You also gave an example of:\n> \n>          commit-a 11) void new_func_1(void *x, void *y);\n>          commit-b 12) void new_func_2(void *x, void *y);\n> \n> Being reformatted to:\n> \n>          commit-a 11) void new_func_1(void *x,\n>          commit-b 12)                 void *y);\n>          commit-b 13) void new_func_2(void *x,\n>          commit-b 14)                 void *y);\n> \n> The patch I submitted handles this case correctly, assigning line 12\n> to commit-a because it scales the parent line numbers according to the\n> relative diff chunk sizes instead of assuming a 1-1 mapping.\n> \n> So I do ask that you incorporate more of my patch, including the test\n> code. It is more complex but I hope this demonstrates that there are\n> reasons for that. Happy to provide more examples or explanation if it\n> would help. On the other hand if you have examples where it falls\n> short then I'd be interested to know.\n\nMy main concerns:\n- Can your version reach outside of a diff chunk?  such as in my \"header \nmoved\" case.  That was a simplified version of something that pops up in \na major file reformatting of mine, where a \"return 0;\" was matched as \ncontext and broke a diff chunk up into two blame_chunk() calls.  I tend \nto think of this as the \"split diff chunk.\"\n\n- Complexity and possibly performance.  The recursive stuff made me \nwonder about it a bit.  It's no reason not to use it, just need to check \nit more closely.\n\nIs the latest version of your stuff still the one you posted last week \nor so?  If we had a patch applied onto this one with something like an \nifdef or a dirt-simple toggle, we can play with both of them in the same \ncodebase.\n\nSimilarly, do you think the \"two pass\" approach I have (check the chunk, \nthen check the parent file) would work with your recursive partitioning \nstyle?  That might make yours able to handle the \"split diff chunk\" case.\n\n> The other major use case that I'm interested in is renaming. In this\n> case, the git-hyper-blame approach of mapping line numbers 1-1 works\n> perfectly. Here's an example. Before:\n> \n>          commit-a 11) Position MyClass::location(Offset O) {\n>          commit-b 12)    return P + O;\n>          commit-c 13) }\n> \n> After:\n> \n>          commit-a 11) Position MyClass::location(Offset offset) {\n>          commit-a 12)    return position + offset;\n>          commit-c 13) }\n> \n> With the fuzzy matching, line 12 gets incorrectly matched to parent\n> line 11 because the similarity of \"position\" and \"offset\" outweighs\n> the similarity of \"return\". I'm considering adding even more\n> complexity to my patch such that parts of a line that have already\n> been matched can't be matched again by other lines.\n> \n> But the other possibility is that we let the user choose the\n> heuristic. For a commit where they know that line numbers haven't\n> changed they could choose 1-1 matching, while for a reformatting\n> commit they could use fuzzy matching. I welcome your thoughts.\n\nNo algorithm will work for all cases.  The one you just gave had the \nsimple heuristic working better than a complex one.  We could make it \nmore complex, but then another example may be worse.  I can live with \nsome inaccuracy in exchange for simplicity.\n\nI ran into something similar with the THRESHOLD #defines.  You want it \nto be able to match certain things, but not other things.  How similar \ndoes something have to be?  Should it depend on how far away the \nmatching line is from the source line?  I went with a \"close enough is \ngood enough\" approach, since we're marking with a '*' or something \nanyways, so the user should know to not trust it 100%.\n\nThanks,\n\nBarret\n\n"},{"id":"373868","messageId":"a742dd62-c84e-1f85-0663-4a3aa4d14989@google.com","threadId":"50912","inReplyTo":"877ec1fzqy.fsf@evledraar.gmail.com","subject":"Re: [PATCH v6 1/6] Move init_skiplist() outside of fsck","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-15T13:32:06Z","receivedAt":"2019-04-15T13:32:11Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"On 4/10/19 3:04 PM, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Wed, Apr 10 2019, Barret Rhoden wrote:\n> \n>> init_skiplist() took a file consisting of SHA-1s and comments and added\n>> the objects to an oidset.  This functionality is useful for other\n>> commands.\n> \n> This change would be much easier to review if you led with a commit\n> where you s/Invalid SHA-1/invalid object name/ (lower-case while we're\n> at it), s/skip list/object name/ etc, and did that rename of the \"hash\"\n> to \"name\" variable if you're so inclined.\n> \n> Then you'd end up with a small refactoring change that changes the tests\n> (or even just make the tests grep for e.g. \"Could not open.*:\n> does-not-exist\" instead), and the moving of the function would be\n> entirely caught by the rename detection.\n> \n\nCan do.  I'll split this up in the next round.\n\nThanks,\n\nBarret\n\n\n"},{"id":"373869","messageId":"7378e4c5-b86c-a7c2-c2df-3beaff0c5970@google.com","threadId":"50912","inReplyTo":"CAJDYR9Q-ixsxWyMrm7aCojTv33SOj3+ALPwJYo9DJE7vLU=DEA@mail.gmail.com","subject":"Re: [PATCH v6 3/6] blame: add the ability to ignore commits and their changes","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-15T13:32:10Z","receivedAt":"2019-04-15T13:32:14Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"On 4/14/19 6:42 AM, Michael Platings wrote:\n>>> +     the `blame.ignoreRevsFile` config option.  An empty file name, `\"\"`, will\n>>> +     clear the list of revs from previously processed files.\n>>\n>> Maybe I haven't read this carefully enough but the use-case for this\n>> doesn't seem to be explained, you need this for the option, but the\n>> config file too? If I want to override fsck.skipList I do\n>> `fsck.skipList=/dev/zero`. Isn't that enough for this use-case without\n>> introducing config state-machine magic?\n> \n> The difference between blame.ignoreRevsFile and fsck.skipList is that\n> ignoreRevsFile can be specified repeatedly. This is useful if you have\n> one file listing reformatting commits, another listing renaming\n> commits etc. Or maybe a checked-in list of commits to ignore, and a\n> personal list of commits to ignore. However sometimes you're going to\n> want to *not* ignore those commits, so you need a way to discard the\n> previously specified options. To accommodate all operating systems an\n> empty string seems the best way to do this.\n> \n\nIn a previous round of reviews[1], this style was recommended.  It's \nbased on what credential.helper does.\n\nThe main thing I've been using the --ignore-revs-file=\"\" for is to turn \noff my default ignore list for debugging.  =)\n\nThanks,\n\nBarret\n\n\n[1] \nhttps://public-inbox.org/git/nycvar.QRO.7.76.6.1901181038540.41@tvgsbejvaqbjf.bet/\n"},{"id":"373870","messageId":"3db6bad3-e7a5-af1d-3fe2-321bd17db2c6@google.com","threadId":"50912","inReplyTo":"878swhfzxb.fsf@evledraar.gmail.com","subject":"Re: [PATCH v6 3/6] blame: add the ability to ignore commits and their changes","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-15T13:34:11Z","receivedAt":"2019-04-15T13:34:16Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"On 4/10/19 3:00 PM, Ævar Arnfjörð Bjarmason wrote:\n[snip]\n\n>> +\tsplit[0].unblamable = e->unblamable;\n>> +\tsplit[1].unblamable = e->unblamable;\n>> +\tsplit[2].unblamable = e->unblamable;\n> \n> I wonder what the comfort level for people in general is before turning\n> this sort of thing into a for-loop, 4? :)\n\n4 sounds good to me.  =)\n\n>> +\tnr_lines = e->num_lines;\t// e changes in the loop\n> \n> A C++-like trailing comment.\n> \n>> +\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 1\" blame_raw | sed -e \"s/ .*//\" >actual &&\n>> +\tgit rev-parse X >expect &&\n>> +\ttest_cmp expect actual &&\n>> +\n>> +\tgrep \"^[0-9a-f]\\+ [0-9]\\+ 2\" blame_raw | sed -e \"s/ .*//\" >actual &&\n>> +\tgit rev-parse X >expect &&\n>> +\ttest_cmp expect actual\n> \n> The grep here is a bug. See my 4abf20f004 (\"tests: fix unportable \"\\?\"\n> and \"\\+\" regex syntax\", 2019-02-21).\n\nThanks - will fix up this stuff in the next round.\n\n"},{"id":"373872","messageId":"1a1b3cd1-5f00-37c9-7382-72de000dd925@google.com","threadId":"50912","inReplyTo":"CAJDYR9TRk99Kwq5S7udVqYsXnupGD=t3o_Ss8ewvwWuTQOy_YQ@mail.gmail.com","subject":"Re: [PATCH v6 4/6] blame: add config options to handle output for ignored lines","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-15T13:51:37Z","receivedAt":"2019-04-15T13:51:43Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"Hi -\n\nOn 4/14/19 7:27 AM, Michael Platings wrote:\n> On Sun, 14 Apr 2019 at 11:24, Junio C Hamano <gitster@pobox.com> wrote:\n>>> If you only enable blame.markIgnoredLines then the hash for\n>>> \"unblamable\" lines appears as e.g. \"*3252488f5\" - this doesn't seem\n>>> right to me because the commit *wasn't* ignored,\n>>\n>> I think you misunderstood me.  I was merely suggesting to use the\n>> approach to mark the line in a way other than using the NULLed out\n>> object name that has been reserved for something totally different,\n>> and hinting with \"the same *idea*\".\n> \n> Hi Junio, that paragraph wasn't targetted at yourself, more a comment\n> on the functionality as it exists in the latest patch series. Sorry\n> for not making that clear.\n> \n>> the \"^\" marker\n>> that is used to say \"the line is attributed to this commit, but that\n>> may only be because you blamed with commit range A..B and we reached\n>> the bottom of the range---if you dug further, you might find the\n>> line originates from another commit\" is the origin of the same idea,\n>> and this topic borrows it and uses a different mark, i.e. '*', for\n>> the \"we are not certain---take this with grain of salt\" mark.\n> \n> So it sounds like we have many types of blame to consider:\n> \n> 1) This commit is truly the last one to touch this line, and you\n> didn't ask to ignore it.\n> 2) This commit is truly the last one to touch this line, but you asked\n> to ignore it (AKA \"unblamable\").\n> 3) This commit is at the bottom of the range of commits (^)\n> 4) The \"true\" commit was ignored but we guess this is the one you're\n> actually interested in (*)\n> 5) The \"true\" commit was ignored and we've reached the bottom of the\n> range of commits (^*)?\n> 6) This commit is at the bottom of the range of commits, and you asked\n> to ignore it.\n> \n>> If you ended up hitting the commit the user wanted to ignore,\n>> perhaps you can find another character that is different from '^' or\n>> '*' and use that, following the same idea.\n> \n> I personally don't find the \"unblamable\" lines interesting enough to\n> justify giving them a symbol. But if Barret strongly feels that such\n> lines should get a '*' then I won't fight it - these lines tend to be\n> as simple as \"}\".\n\nI'm fine with not zeroing the hash, so long as there's some way to mark it.\n\nWe could mark with another *, such that if we mark-ignored and \nmark-unblamable you get \"**hash\".  You can't have an unblamable that \nisn't from an ignored commit, so a single '*' has only one meaning, \nbased on your config options.\n\nIf that works for you all, I can change this to markIgnoredUnblamables \n(instead of 'mask') in the next version.\n\n>> By the way, a configuration only feature is something we usually do\n>> not accept.  A feature must be guarded with --command-line-option\n>> and then optionally can have a corresponding configuration once the\n>> option proves to be useful enough that it becomes useful to be able\n>> to say \"in this repository (or to this user), the feature is on by\n>> default\".\n> \n> In that case we definitely need a --mark-ignored-lines option to git\n> blame, and I would strongly prefer that we also keep the\n> blame.markIgnoredLines option as I for one will be switching it on.\n\nI'd also keep this set.  I think the whole reason for these config \noptions was that everyone has a different preference, but that \npreference rarely changes.  I don't want to have to type \n--mark-ignored-lines every time I run git blame.  If I had to, I'd have \nto alias git blame or something.\n\nI think having config options for these sorts of things is fine, since \nwe know already that for a given user+repo, we want the feature on (or \noff).  But if I have to remove it, then let me know.\n\nThanks,\n\nBarret\n\n"},{"id":"373874","messageId":"4c7bbc6c-805e-da7f-593c-e73989fc37c0@google.com","threadId":"50912","inReplyTo":"xmqqk1fxw8ad.fsf@gitster-ct.c.googlers.com","subject":"Re: [PATCH v6 6/6] blame: use a fingerprint heuristic to match ignored lines","fromName":"Barret Rhoden","fromEmail":"brho@google.com","sentAt":"2019-04-15T14:03:10Z","receivedAt":"2019-04-15T14:03:16Z","isPatch":true,"sender":{"key":"brho@google.com","avatar":null},"body":"On 4/13/19 11:54 PM, Junio C Hamano wrote:\n> Two thoughts.\n> \n>   - Unless the 'old heuristic' is still available as an option after\n>     this step, a series that first begins with the 'old heuristic'\n>     and then later replaces it with the 'new heuristic' feels\n>     somewhat wasteful of reviewer resources, as the 'old heuristic'\n>     does not contribute an iota to the end result.\n> \n>     It is OK while the series is still in RFC/WIP stage, though.  But\n>     because I got an impression that this is close to completion, so...\n\nCan do.  I wasn't sure yet where things were going, but in the final \nversion, I can yank out the old heuristic from the patch set.\n\nThough the old heuristic is pretty basic - really just a couple lines - \nand it may help to see it before looking at a more complicated version. \nEspecially since it helps break the commit up into \"infrastructure to \nignore commits\" and \"brains to find the right commit to blame\" while \nstill being functional between the commits.\n\nThanks,\n\nBarret\n"},{"id":"373916","messageId":"CAJDYR9QSAoYkrbdyJBN1tg0v1x6Do9qGf0+6hYL-n49+HSfu8g@mail.gmail.com","threadId":"50912","inReplyTo":"9439c697-246f-3bcb-4d34-85099e577e8b@google.com","subject":"Re: [PATCH v6 0/6] blame: add the ability to ignore commits","fromName":"Michael Platings","fromEmail":"michael@platin.gs","sentAt":"2019-04-15T21:54:10Z","receivedAt":"2019-04-15T21:54:25Z","isPatch":true,"sender":{"key":"michael@platin.gs","avatar":"https://avatars.githubusercontent.com/u/1112348?v=4"},"body":"> My main concerns:\n> - Can your version reach outside of a diff chunk?\n\nCurrently no. It's optimised for reformatting and renaming, both of\nwhich preserve ordering. I could look into allowing disordered matches\nwhere the similarity is high, while still being biased towards ordered\nmatches. If you can post more examples that would be helpful.\n\n> - Complexity and possibly performance.  The recursive stuff made me\n> wonder about it a bit.  It's no reason not to use it, just need to check\n> it more closely.\n\nComplexity I can't deny, I can only mitigate it with\ndocumentation/comments. I optimised the code pretty heavily and tested\non some contrived worst-case scenarios and the performance was still\ngood so I'm not worried about that.\n\n> Is the latest version of your stuff still the one you posted last week\n> or so?\n\nYes. But reaching outside the chunk might lead to a significantly\ndifferent API in the next version...\n\n> Similarly, do you think the \"two pass\" approach I have (check the chunk,\n> then check the parent file) would work with your recursive partitioning\n> style?  That might make yours able to handle the \"split diff chunk\" case.\n\nYes, should do. I'll see what I can come up with this week.\n\n> No algorithm will work for all cases.  The one you just gave had the\n> simple heuristic working better than a complex one.  We could make it\n> more complex, but then another example may be worse.  I can live with\n> some inaccuracy in exchange for simplicity.\n\nExactly, no algorithm will work for all cases. So what I'm suggesting\nis that it might be best to let the user choose which heuristic is\nappropriate for a given commit. If they know that the simple heuristic\nworks best then perhaps we should let them choose that rather than\nonly offering a one-size-fits-all option. But if we do want to go for\none-size-fits-all then I'm very keen to make sure it at least solves\nthe specific cases that we know about.\n"},{"id":"373926","messageId":"xmqq1s22r3nm.fsf@gitster-ct.c.googlers.com","threadId":"50912","inReplyTo":"4c7bbc6c-805e-da7f-593c-e73989fc37c0@google.com","subject":"Re: [PATCH v6 6/6] blame: use a fingerprint heuristic to match ignored lines","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2019-04-16T04:10:37Z","receivedAt":"2019-04-16T04:10:42Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Barret Rhoden <brho@google.com> writes:\n\n> Though the old heuristic is pretty basic - really just a couple lines\n> - \n> and it may help to see it before looking at a more complicated\n> version.\n\nOK, then.\n\nThanks.\n"}]}