{"thread":{"id":"62993","subject":"Diff rename detection performance issues","startedAt":"2025-02-23T10:29:52Z","lastAt":"2025-02-24T17:50:39Z","messageCount":3,"participants":["Devste Devste","Elijah Newren","D. Ben Knoble"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"512882","messageId":"CANM0SV2XOTQ2Mna1B_sX0EF0ffohcrexh1EO5d4G0=sqdmxQtA@mail.gmail.com","threadId":"62993","inReplyTo":null,"subject":"Diff rename detection performance issues","fromName":"Devste Devste","fromEmail":"devstemail@gmail.com","sentAt":"2025-02-23T10:29:38Z","receivedAt":"2025-02-23T10:29:52Z","isPatch":false,"sender":{"key":"devstemail@gmail.com","avatar":null},"body":"I have a merge commit that includes 2 modified (!) files:\nhello/foo/stubs/example.php\nhello/world.php\n\nI want to only get the changes introduced by the merge commit and\nexclude any changes in /foo/stubs/:\ngit diff -l0 --name-status --find-renames \"$sha\"^'!' -- ':!*/foo/stubs/*'\n\nGit takes more than 4 minutes to generate this diff, since\nhello/foo/stubs/example.php is a huge file.\nWhen using --no-renames (instead of --find-renames) it's much, much faster.\nAnd without the example.php file, the diff takes less than 1 second\ninstead of 4+ minutes.\n\nFunnily enough, when I have a merge commit that contains only that 1\nexcluded file, it's the same behavior.\n\n1) if there's only a single file in a commit, why does --find-renames\ncause a slowdown? There's nothing that could have been renamed in that\ncase (probably the same for --find-copies)\n\n2) could rename detection be \"delayed\" to only run/check if there are\nactually additions/deletions (and possibly only check those)? If a\ncommit only contains modifications (unlike in a really, really 0.0001%\nedge case) but no additions+deletions it's extremely unlikely that\nthere's a rename, so detection could be skipped altogether?\n"},{"id":"512917","messageId":"CABPp-BHObCVqxWuBLgeiWghy5gM8-f_qjwYFdBL+=j1bwtPg_A@mail.gmail.com","threadId":"62993","inReplyTo":"CANM0SV2XOTQ2Mna1B_sX0EF0ffohcrexh1EO5d4G0=sqdmxQtA@mail.gmail.com","subject":"Re: Diff rename detection performance issues","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2025-02-24T16:30:00Z","receivedAt":"2025-02-24T16:31:37Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"On Sun, Feb 23, 2025 at 2:30 AM Devste Devste <devstemail@gmail.com> wrote:\n>\n> I have a merge commit that includes 2 modified (!) files:\n\nWhat do you mean that it only includes 2 modified files?  Modified\nrelative to what?  Modified relative to the merge base of its parents?\n Modified relative to its first parent?  to its second parent?\nModified relative to an automatic merge?\n\nAlso, by \"modified\" here do you mean the change type is 'M' in\n--name-status output or could the change type also be 'A' (added) or\n'D'(deleted) or something else?\n\n> hello/foo/stubs/example.php\n> hello/world.php\n>\n> I want to only get the changes introduced by the merge commit and\n> exclude any changes in /foo/stubs/:\n> git diff -l0 --name-status --find-renames \"$sha\"^'!' -- ':!*/foo/stubs/*'\n\nIt's not clear to me from your example what the output of say\n\n   git diff --name-status --no-renames \"$sha\"^'!' | wc -l\n\nwould be, though I would find that very interesting.  I'm also curious\nwhat you'd get from each of\n\n  git diff --diff-filter=D --name-status --no-renames \"$sha\"^'!' | wc -l\n  git diff --diff-filter=A --name-status --no-renames \"$sha\"^'!' | wc -l\n  git diff --diff-filter=M --name-status --no-renames \"$sha\"^'!' | wc -l\n\n(and yes, I am very intentionally leaving off the ':!*/foo/stubs/*'\nnegative refspec; I want the output without that.)\n\n> Git takes more than 4 minutes to generate this diff, since\n> hello/foo/stubs/example.php is a huge file.\n\nHow do you know that is the reason?  Especially since...\n\n> When using --no-renames (instead of --find-renames) it's much, much faster.\n\n...this seems to contradict your statement that the reason for the\nslow diff is that hello/foo/stubs/example.php is a huge file.\n\n> And without the example.php file, the diff takes less than 1 second\n> instead of 4+ minutes.\n\nWhat do you mean without the example.php file?  Did you rewind\nhistory, remove that file, and then redo the merge so that it is no\nlonger included?  Or do you mean something else entirely?  What\nexactly?\n\n> Funnily enough, when I have a merge commit that contains only that 1\n> excluded file, it's the same behavior.\n>\n> 1) if there's only a single file in a commit, why does --find-renames\n> cause a slowdown? There's nothing that could have been renamed in that\n> case (probably the same for --find-copies)\n\nI'm not sure what this has to do with the above; you seem to have\nswitched tracks.  If you have a commit whose toplevel tree has exactly\n1 file, and you're diffing it against some other commit with an\nunspecified number of files, then if that other commit with N files\nhappens to have a file with the same name as the commit with exactly 1\nfile, then --find-renames can't really cause a slowdown.  It'd only\ncause a slowdown when the N files in the other commit were all\ndifferent filenames than the 1 file in your commit you are diffing\nagainst (but of mostly similar filesize).  But I suspect you meant\nsomething other than what you said here.  Could you clarify the actual\nsetup?\n\n> 2) could rename detection be \"delayed\" to only run/check if there are\n> actually additions/deletions (and possibly only check those)? If a\n> commit only contains modifications (unlike in a really, really 0.0001%\n> edge case) but no additions+deletions it's extremely unlikely that\n> there's a rename, so detection could be skipped altogether?\n\nRename detection already does this; in fact, it does better.  Not only\ncan you exit early when additions + deletions are empty, you can also\nexit early when either of the two are empty.\n\n(In fact, there's some other optimizations as well, such as exiting\nearly if either additions or deletions become empty after removing any\npaths involved in exact rename detection, or removing any paths\ninvolved in basename-driven rename matching.)\n\nIf you want to see where this is handled; see the \"if\n(!num_destinations || !num_sources)\" check in diffcore-rename.c.\n\n\nNow, all that said, I suspect you're getting at something with the\nnegative refspecs that is similar to the optimization idea I had for a\nreal --follow-renames, but before I jump into that, I'd need you to\nclarify your setup a fair amount to make sure we're on the same page.\n"},{"id":"512924","messageId":"CALnO6CDpEfugTReF3j_3jefaDg2-YtMB-2XrKg07wD4cofHK7g@mail.gmail.com","threadId":"62993","inReplyTo":"CABPp-BHObCVqxWuBLgeiWghy5gM8-f_qjwYFdBL+=j1bwtPg_A@mail.gmail.com","subject":"Re: Diff rename detection performance issues","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2025-02-24T17:50:26Z","receivedAt":"2025-02-24T17:50:39Z","isPatch":false,"sender":{"key":"ben.knoble@gmail.com","avatar":"https://avatars.githubusercontent.com/u/22802209?v=4"},"body":"On Mon, Feb 24, 2025 at 11:31 AM Elijah Newren <newren@gmail.com> wrote:\n>\n> On Sun, Feb 23, 2025 at 2:30 AM Devste Devste <devstemail@gmail.com> wrote:\n[snip]\n> > Funnily enough, when I have a merge commit that contains only that 1\n> > excluded file, it's the same behavior.\n> >\n> > 1) if there's only a single file in a commit, why does --find-renames\n> > cause a slowdown? There's nothing that could have been renamed in that\n> > case (probably the same for --find-copies)\n>\n> I'm not sure what this has to do with the above; you seem to have\n> switched tracks.  If you have a commit whose toplevel tree has exactly\n> 1 file, and you're diffing it against some other commit with an\n> unspecified number of files, [snip]\n\nI'm only mentioning this in the vein of Elijah's requests for\nclarification: the wording \"only a single file in a commit\" is\nsomething I often see from newcomers who don't yet understand that a\ncommit points to a tree of the entire repo, but the diff between a\ncommit and it's parent might show only one modified file. (Sometimes I\nthink we experts encourage this when we refer to that diff as the\ncommit [1], [2].)\n\nNow, Devste's posted commands indicate they may have more Git\nexperience and didn't fall to this trap, so Elijah's interpretation of\n\"a commit whose toplevel tree has exactly 1 file\" is perfectly\nreasonable—but we'd probably all like to know a bit more to confirm. I\noriginally read \"if there's only a single file in the commit\" (with my\nnewcomer lenses on) as \"if I only changed one file before commiting.\"\nThis is also partly based on a (mis)read of \"a merge commit that\ncontains only that 1 excluded file\" (perhaps OP meant \"modified\").\n\n[1]: https://jvns.ca/blog/2023/11/01/confusing-git-terminology/#commit\n[2]: https://jvns.ca/blog/2024/01/05/do-we-think-of-git-commits-as-diffs--snapshots--or-histories/\n\n-- \nD. Ben Knoble\n"}]}