{"thread":{"id":"44512","subject":"Fwd: git diff with “--word-diff-regex” extremely slow compared to “--word-diff”?","startedAt":"2016-11-18T23:40:29Z","lastAt":"2016-11-22T19:27:05Z","messageCount":4,"participants":["Matthieu S","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"306185","messageId":"CAEYvigLz3muWD-QFjMZUn=H3RQoxhTYX9EwB6=aiMjWOEN3CBA@mail.gmail.com","threadId":"44512","inReplyTo":"CAEYvigJ14xYDmRG2N0yTgM4spaaB7s9923w0+e9+QQEeFz0NTQ@mail.gmail.com","subject":"Fwd: git diff with “--word-diff-regex” extremely slow compared to “--word-diff”?","fromName":"Matthieu S","fromEmail":"matthieu.stigler@gmail.com","sentAt":"2016-11-18T23:40:22Z","receivedAt":"2016-11-18T23:40:29Z","isPatch":false,"sender":{"key":"matthieu.stigler@gmail.com","avatar":null},"body":"Hi\n\nWhen giving a custom regex to git diff --word-diff-regex= instead of\nusing the default --word-diff (which splits words on whitespace), git\nslows down very considerably... I don't understand why such a speed\ndifference?\n\n(this question was asked on stack overflow, but after two month\nwithout answer, I'm asking it here instead. Post:\nhttp://stackoverflow.com/questions/39027864/git-diff-with-word-diff-regex-extremely-slow-compared-to-word-diff).\n\nExample (sorry, UNIX specific code): create two one-line files, and\ntwo 200000-lines files:\n\necho aaa,bbb ,12,12,15 >file1.txt\necho aaa,bbb ,12,12,16 >file2.txt\n\nawk '{for(i=0;i<200000;i++)print}' file1.txt > file1BIG.txt\nawk '{for(i=0;i<200000;i++)print}' file2.txt > file2BIG.txt\n\nDefault --word-diff has no issues with the BIG files (cannot see time\ndifference):\n\ngit diff --word-diff file1.txt file2.txt\ngit diff --word-diff file1BIG.txt file2BIG.txt\n\nNow use instead --word-diff-regex= argument (with regex from post:\nhttp://stackoverflow.com/questions/10482773/also-use-comma-as-a-word-separator-in-diff\n)\n\ngit diff --word-diff-regex=[^[:space:],] file1.txt file2.txt\ngit diff --word-diff-regex=[^[:space:],] file1BIG.txt file2BIG.txt\n\nWhy is the speed so different if one uses --word-diff instead of\n--word-diff-regex= ? Is it just because my expression is (slightly)\nmore complex than the default one (split on period instead of only\nwhitespace) ? Or is it that the default word-diff is implemented\ndifferently/more efficiently? How can I overcome this speed slowdown?\n\nThanks!!\n\nMatthieu\n\n\nPS: using git 2.7.4 on Ubuntu 16.04\n"},{"id":"306205","messageId":"20161120201744.7ym4gsmjoijw6oow@sigill.intra.peff.net","threadId":"44512","inReplyTo":"CAEYvigLz3muWD-QFjMZUn=H3RQoxhTYX9EwB6=aiMjWOEN3CBA@mail.gmail.com","subject":"Re: Fwd: git diff with “--word-diff-regex” extremely slow compared to “--word-diff”?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-11-20T20:17:44Z","receivedAt":"2016-11-20T20:17:51Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Nov 18, 2016 at 03:40:22PM -0800, Matthieu S wrote:\n\n> Why is the speed so different if one uses --word-diff instead of\n> --word-diff-regex= ? Is it just because my expression is (slightly)\n> more complex than the default one (split on period instead of only\n> whitespace) ? Or is it that the default word-diff is implemented\n> differently/more efficiently? How can I overcome this speed slowdown?\n\nI think it's probably both.\n\nSee diff.c:find_word_boundaries(). If there's no regex, we use a simple\nloop over isspace() to find the boundaries. I don't recall anybody\nmeasuring the performance before, but I'm not surprised to hear that\nmatching a regex is slower.\n\nIf I look at the output of \"perf\", though, it looks like we also spend a\nlot more time in xdl_clean_mmatch(). Which isn't surprising. Your regex\ntreats commas as boundaries, which is going to generate a lot more\nmatches for this particular data set (though the output is the same, I\nthink, because of the nature of the change).\n\nI would have expected \"--word-diff-regex=[^[:space:]]\" to be faster than\nyour regex, though, and it does not seem to be.\n\n-Peff\n"},{"id":"306324","messageId":"CAEYvigLLQq2SK60UsiTPCxpptpjz85_rGtDVugjfu-sCT1juGQ@mail.gmail.com","threadId":"44512","inReplyTo":"20161120201744.7ym4gsmjoijw6oow@sigill.intra.peff.net","subject":"Re: Fwd: git diff with “--word-diff-regex” extremely slow compared to “--word-diff”?","fromName":"Matthieu S","fromEmail":"matthieu.stigler@gmail.com","sentAt":"2016-11-22T18:08:33Z","receivedAt":"2016-11-22T18:09:45Z","isPatch":false,"sender":{"key":"matthieu.stigler@gmail.com","avatar":null},"body":"Thanks Jeff for the answer!\n\nYou are right, I should have compared with the same regex, and indeed,\n--word-diff-regex=[^[:space:]] is also much slower than just\n--word-diff, although they do the same job. Maybe this is a hint that\nthe --word-diff-regex code could be made faster?\n\nI have a small understanding of git, but is git diff computing the\ndiff value for the whole file, and then showing in the terminal the 10\nfirst values? In some cases, it seems to be a lot of unnecessary\ncomputation! Is there any possibility to ask git-diff to only compare\nsay the first 100 lines? Or compute only when necessary, i.e.\nwhen\"enter\" is prompted in the console?\n\nThanks!\n\nMatthieu\n\n2016-11-20 12:17 GMT-08:00 Jeff King <peff@peff.net>:\n> On Fri, Nov 18, 2016 at 03:40:22PM -0800, Matthieu S wrote:\n>\n>> Why is the speed so different if one uses --word-diff instead of\n>> --word-diff-regex= ? Is it just because my expression is (slightly)\n>> more complex than the default one (split on period instead of only\n>> whitespace) ? Or is it that the default word-diff is implemented\n>> differently/more efficiently? How can I overcome this speed slowdown?\n>\n> I think it's probably both.\n>\n> See diff.c:find_word_boundaries(). If there's no regex, we use a simple\n> loop over isspace() to find the boundaries. I don't recall anybody\n> measuring the performance before, but I'm not surprised to hear that\n> matching a regex is slower.\n>\n> If I look at the output of \"perf\", though, it looks like we also spend a\n> lot more time in xdl_clean_mmatch(). Which isn't surprising. Your regex\n> treats commas as boundaries, which is going to generate a lot more\n> matches for this particular data set (though the output is the same, I\n> think, because of the nature of the change).\n>\n> I would have expected \"--word-diff-regex=[^[:space:]]\" to be faster than\n> your regex, though, and it does not seem to be.\n>\n> -Peff\n"},{"id":"306344","messageId":"20161122192658.annsqeokptac3ivv@sigill.intra.peff.net","threadId":"44512","inReplyTo":"CAEYvigLLQq2SK60UsiTPCxpptpjz85_rGtDVugjfu-sCT1juGQ@mail.gmail.com","subject":"Re: Fwd: git diff with “--word-diff-regex” extremely slow compared to “--word-diff”?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-11-22T19:26:58Z","receivedAt":"2016-11-22T19:27:05Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Nov 22, 2016 at 10:08:33AM -0800, Matthieu S wrote:\n\n> You are right, I should have compared with the same regex, and indeed,\n> --word-diff-regex=[^[:space:]] is also much slower than just\n> --word-diff, although they do the same job. Maybe this is a hint that\n> the --word-diff-regex code could be made faster?\n\nMaybe. If most of the time is spent in the regex engine, there may not\nbe much we can do. But perhaps there is something in the surrounding\ncode that can be improved. Looking at find_word_boundaries() (and this\nis the first time I've done so), it does look like we regex-match the\nwhole buffer, and only then find the end-of-line. Now that we have\nregexec_buf(), it might be possible to constrain the regex buffer more.\n\n> I have a small understanding of git, but is git diff computing the\n> diff value for the whole file, and then showing in the terminal the 10\n> first values? In some cases, it seems to be a lot of unnecessary\n> computation! Is there any possibility to ask git-diff to only compare\n> say the first 100 lines? Or compute only when necessary, i.e.\n> when\"enter\" is prompted in the console?\n\nGit always computes the diff for the whole file. The paging is done by\nan external program. So no, there's no easy way to do it incrementally\nas the user interacts with the pager, as the pager does not communicate\nback to git in any way. However, git should generally be streaming out\nresults (and the pager showing them) as they're computed, so in an ideal\nworld you get output immediately, and then the pager buffers the rest of\nit while you're reading the first page.\n\nGit does have to look at the whole file in order to do the initial\nline-by-line diff, so it would be hard to make that incremental. It\ncould do the word-coloring for each hunk incrementally, though. I would\nhave assumed that is already how it is done, though I didn't dig into\nit.\n\n-Peff\n"}]}