{"thread":{"id":"55244","subject":"Can I convince the diff algorithm to behave better?","startedAt":"2021-03-03T06:42:28Z","lastAt":"2021-03-04T09:54:55Z","messageCount":4,"participants":["Tom Ritter","Thomas Braun","Jonathan Tan","Christian Couder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"418125","messageId":"CA+cU71=FfReSG411Feo=vmkw4MdK4KDgokP1jH6uwOkC_0AbYA@mail.gmail.com","threadId":"55244","inReplyTo":null,"subject":"Can I convince the diff algorithm to behave better?","fromName":"Tom Ritter","fromEmail":"tom@ritter.vg","sentAt":"2021-03-03T02:03:59Z","receivedAt":"2021-03-03T06:42:28Z","isPatch":false,"sender":{"key":"tom@ritter.vg","avatar":null},"body":"(For a specific, nuanced, and personal definition of better...)\n\nI have a frequent behavior that arises when I am copy/pasting chunks\nof code, typically in tests.  Here is an example:\n\nMy Original code:\n\ndef function():\n   line 1\n   line 2\n   line 3\n   line 4\n   line 5\n   line 6\n\n--------------------------------\nI add, after it:\n\ndef function2():\n   line 1\n   line 2\n   line 3\n   line 4\n   line 5\n   line 6\n\n--------------------------------\nMy diff is:\n\n+   line 3\n+   line 4\n+   line 5\n+   line 6\n+\n+def function2():\n+   line 1\n+   line 2\n\n--------------------------------\nI'd like my diff to be\n\n+\n+def function2():\n+   line 1\n+   line 2\n+   line 3\n+   line 4\n+   line 5\n+   line 6\n\n\nObviously there's nothing incorrect about the former diff, I just wish\nit was the latter rather than the former.\n\nI know that git includes four diff algorithms; in my testing patience\nor histogram exacerbated the problem; and none of them improved upon\nit.  If anyone has suggestions I'd be curious to know if there's\nanything that could be done...\n\nThanks,\n-tom\n"},{"id":"418154","messageId":"10c330f1-b3ae-38ab-1a8b-23c0b46f1557@virtuell-zuhause.de","threadId":"55244","inReplyTo":"CA+cU71=FfReSG411Feo=vmkw4MdK4KDgokP1jH6uwOkC_0AbYA@mail.gmail.com","subject":"Re: Can I convince the diff algorithm to behave better?","fromName":"Thomas Braun","fromEmail":"thomas.braun@virtuell-zuhause.de","sentAt":"2021-03-03T12:41:00Z","receivedAt":"2021-03-04T00:23:08Z","isPatch":false,"sender":{"key":"thomas.braun@virtuell-zuhause.de","avatar":"https://avatars.githubusercontent.com/u/1185677?v=4"},"body":"On 3/3/2021 3:03 AM, Tom Ritter wrote:\n\nHi Tom,\n\n> (For a specific, nuanced, and personal definition of better...)\n> \n> I have a frequent behavior that arises when I am copy/pasting chunks\n> of code, typically in tests.  Here is an example:\n> \n> My Original code:\n> \n> def function():\n>    line 1\n>    line 2\n>    line 3\n>    line 4\n>    line 5\n>    line 6\n> \n> --------------------------------\n> I add, after it:\n> \n> def function2():\n>    line 1\n>    line 2\n>    line 3\n>    line 4\n>    line 5\n>    line 6\n> \n> --------------------------------\n> My diff is:\n> \n> +   line 3\n> +   line 4\n> +   line 5\n> +   line 6\n> +\n> +def function2():\n> +   line 1\n> +   line 2\n> \n> --------------------------------\n> I'd like my diff to be\n> \n> +\n> +def function2():\n> +   line 1\n> +   line 2\n> +   line 3\n> +   line 4\n> +   line 5\n> +   line 6\n\nI tried to reproduce and got exactly the diff you wanted to have. I need\nto add a newline after the first \"line 4\" to get the not-sought-for diff.\n\nCommit:\n\n+++ b/test.py\n@@ -0,0 +1,7 @@\n+def function():\n+    line 1\n+    line 2\n+    line 3\n+    line 4\n+    line 5\n+    line 6\n\nand then the following change:\n\n--- a/test.py\n+++ b/test.py\n@@ -3,5 +3,14 @@ def function():\n     line 2\n     line 3\n     line 4\n+\n+    line 5\n+    line 6\n+\n+def function2():\n+    line 1\n+    line 2\n+    line 3\n+    line 4\n     line 5\n     line 6\n\nI usually play around with --anchored when I want to solve an issue like\nthat.\n\nThe documentation of anchored says\n\nIf a line exists in both the source and destination, exists only once,\nand starts with this text, this algorithm attempts to prevent it from\nappearing as a deletion or addition in the output. It uses the \"patience\ndiff\" algorithm internally.\n\nBut I can't get it working here as the \"exists only once\" premise is broken.\n\nStepping back: It might also make sense to rethink the code as repeating\nthe same 6 lines in every function might not be the best possible design.\n\nThomas\n\n[...]\n"},{"id":"418178","messageId":"20210303234530.3122368-1-jonathantanmy@google.com","threadId":"55244","inReplyTo":"CA+cU71=FfReSG411Feo=vmkw4MdK4KDgokP1jH6uwOkC_0AbYA@mail.gmail.com","subject":"Re: Can I convince the diff algorithm to behave better?","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2021-03-03T23:45:30Z","receivedAt":"2021-03-04T00:24:27Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"> I know that git includes four diff algorithms; in my testing patience\n> or histogram exacerbated the problem; and none of them improved upon\n> it.  If anyone has suggestions I'd be curious to know if there's\n> anything that could be done...\n\nIn your particular case, I can't think of anything, but in general, if\none of the lines weren't repeated, you might be able to use the\n--anchored option.\n"},{"id":"418225","messageId":"CAP8UFD1gFA2DyyjyfJ7pKNRyqO2=Y2BOyF0Aeni3GXvgeL8Wtg@mail.gmail.com","threadId":"55244","inReplyTo":"CA+cU71=FfReSG411Feo=vmkw4MdK4KDgokP1jH6uwOkC_0AbYA@mail.gmail.com","subject":"Re: Can I convince the diff algorithm to behave better?","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2021-03-04T09:52:55Z","receivedAt":"2021-03-04T09:54:55Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"On Thu, Mar 4, 2021 at 8:37 AM Tom Ritter <tom@ritter.vg> wrote:\n\n[...]\n\n> Obviously there's nothing incorrect about the former diff, I just wish\n> it was the latter rather than the former.\n>\n> I know that git includes four diff algorithms; in my testing patience\n> or histogram exacerbated the problem; and none of them improved upon\n> it.  If anyone has suggestions I'd be curious to know if there's\n> anything that could be done...\n\nIt's not so easy to implement good diff algorithms. You might want to\ntake a look at the \"v2.11 new diff heuristic?\" article in:\n\nhttps://git.github.io/rev_news/2016/12/14/edition-22/\n\nBest,\nChristian.\n"}]}