{"thread":{"id":"65765","subject":"[PATCH] ls-files: filter pathspec before lstat","startedAt":"2026-06-07T15:41:06Z","lastAt":"2026-09-06T12:31:15Z","messageCount":11,"participants":["Tamir Duberstein","Kristoffer Haugsbakk","Junio C Hamano","Jeff King"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"544847","messageId":"20260607-ls-files-pathspec-lstat-v1-1-8cf40b730146@gmail.com","threadId":"65765","inReplyTo":null,"subject":"[PATCH] ls-files: filter pathspec before lstat","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-06-07T15:40:56Z","receivedAt":"2026-06-07T15:41:06Z","isPatch":true,"body":"show_files() checks whether each index entry is deleted or modified\nbefore show_ce() applies the pathspec. prune_index() avoids most of this\nwork for pathspecs with a common directory prefix, but a top-level name\nor leading wildcard leaves every entry to be checked.\n\nMatch the pathspec before lstat() for the deleted and modified modes.\nKeep the later match in show_ce() so --error-unmatch is satisfied only\nby entries that are actually shown.\n\nOn a repository with 859,211 index entries, a 19,931,862-byte index, and\n25,303,439 packed objects occupying 21.13 GiB, I ran the following\ncommand with the parent and patched binaries:\n\n    hyperfine --warmup 0 --runs 3 \\\n        'git -c core.fsmonitor=false ls-files --deleted -- README.md'\n\nThe results were:\n\n             parent       this commit\n  elapsed    60.742 s     1.061 s\n  user        1.117 s     0.963 s\n  system     10.740 s     0.042 s\n\nBoth revisions were built with -O3, -mcpu=native, and ThinLTO using\nApple clang 21.0.0 on macOS 26.5. The machine was a MacBook Pro\n(Mac16,6) with a 16-core Apple M4 Max (12 performance and four\nefficiency cores) and 128 GB RAM.\n\nAssisted-by: Codex gpt-5.5\nSigned-off-by: Tamir Duberstein <tamird@gmail.com>\n---\n builtin/ls-files.c                  |  7 +++++++\n t/meson.build                       |  1 +\n t/perf/p3010-ls-files.sh            | 27 +++++++++++++++++++++++++++\n t/t3010-ls-files-killed-modified.sh | 18 ++++++++++++++++++\n 4 files changed, 53 insertions(+)\n\ndiff --git a/builtin/ls-files.c b/builtin/ls-files.c\nindex e1a22b41b9..702c607183 100644\n--- a/builtin/ls-files.c\n+++ b/builtin/ls-files.c\n@@ -450,6 +450,13 @@ static void show_files(struct repository *repo, struct dir_struct *dir)\n \t\t\tcontinue;\n \t\tif (ce_skip_worktree(ce))\n \t\t\tcontinue;\n+\t\t/* Only entries shown by show_ce() satisfy --error-unmatch. */\n+\t\tif (pathspec.nr &&\n+\t\t    !match_pathspec(repo->index, &pathspec, fullname.buf,\n+\t\t\t\t    fullname.len, max_prefix_len, NULL,\n+\t\t\t\t    S_ISDIR(ce->ce_mode) ||\n+\t\t\t\t    S_ISGITLINK(ce->ce_mode)))\n+\t\t\tcontinue;\n \t\tstat_err = lstat(fullname.buf, &st);\n \t\tif (stat_err && (errno != ENOENT && errno != ENOTDIR))\n \t\t\terror_errno(\"cannot lstat '%s'\", fullname.buf);\ndiff --git a/t/meson.build b/t/meson.build\nindex 2af8d01279..ee8086e6ef 100644\n--- a/t/meson.build\n+++ b/t/meson.build\n@@ -1140,6 +1140,7 @@ benchmarks = [\n   'perf/p1500-graph-walks.sh',\n   'perf/p1501-rev-parse-oneline.sh',\n   'perf/p2000-sparse-operations.sh',\n+  'perf/p3010-ls-files.sh',\n   'perf/p3400-rebase.sh',\n   'perf/p3404-rebase-interactive.sh',\n   'perf/p4000-diff-algorithms.sh',\ndiff --git a/t/perf/p3010-ls-files.sh b/t/perf/p3010-ls-files.sh\nnew file mode 100755\nindex 0000000000..bb80768063\n--- /dev/null\n+++ b/t/perf/p3010-ls-files.sh\n@@ -0,0 +1,27 @@\n+#!/bin/sh\n+\n+test_description='Tests ls-files worktree performance'\n+\n+. ./perf-lib.sh\n+\n+test_perf_large_repo\n+test_checkout_worktree\n+\n+test_expect_success 'select a zero-prefix pathspec' '\n+\ttracked_file=$(git ls-files | sed -n 1p) &&\n+\ttest -n \"$tracked_file\" &&\n+\tpathspec=\"?${tracked_file#?}\" &&\n+\ttest_export pathspec\n+'\n+\n+test_perf 'ls-files --deleted with pathspec' '\n+\tgit -c core.fsmonitor=false ls-files --deleted \\\n+\t\t-- \"$pathspec\" >/dev/null\n+'\n+\n+test_perf 'ls-files --modified with pathspec' '\n+\tgit -c core.fsmonitor=false ls-files --modified \\\n+\t\t-- \"$pathspec\" >/dev/null\n+'\n+\n+test_done\ndiff --git a/t/t3010-ls-files-killed-modified.sh b/t/t3010-ls-files-killed-modified.sh\nindex 7af4532cd1..6e38e10219 100755\n--- a/t/t3010-ls-files-killed-modified.sh\n+++ b/t/t3010-ls-files-killed-modified.sh\n@@ -124,4 +124,22 @@ test_expect_success 'validate git ls-files -m output.' '\n \ttest_cmp .expected .output\n '\n \n+test_expect_success 'worktree modes honor wildcard pathspecs' '\n+\tcat >.expected <<-\\EOF &&\n+\tpath2/file2\n+\tpath3/file3\n+\tEOF\n+\tgit ls-files --deleted -- \"path?/file?\" >.output &&\n+\ttest_cmp .expected .output &&\n+\n+\tcat >.expected <<-\\EOF &&\n+\tpath7\n+\tpath8\n+\tEOF\n+\tgit ls-files --modified --error-unmatch -- \"path[78]\" >.output &&\n+\ttest_cmp .expected .output &&\n+\n+\ttest_must_fail git ls-files --modified --error-unmatch -- path10\n+'\n+\n test_done\n\n---\nbase-commit: 9ac3f193c05c2237e2b14ebaa1149e9fc8a1abe0\nchange-id: 20260607-ls-files-pathspec-lstat-885125a5d644\n\nBest regards,\n--  \nTamir Duberstein <tamird@gmail.com>\n\n"},{"id":"544848","messageId":"8f3bab63-3b37-4492-a39e-95e610a15a07@app.fastmail.com","threadId":"65765","inReplyTo":"20260607-ls-files-pathspec-lstat-v1-1-8cf40b730146@gmail.com","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2026-06-07T16:02:35Z","receivedAt":"2026-06-07T16:02:56Z","isPatch":true,"body":"On Sun, Jun 7, 2026, at 17:40, Tamir Duberstein wrote:\n>[snip]\n> Assisted-by: Codex gpt-5.5\n\nThis is more of a Git for Windows trailer. The Git project doesn’t\ndocument its use.\n\nAn aside here but these trailers attributing specific LLMs feels like\netching “Peter was here” under some table. What benefit for the project\ndoes knowing that it was this version of Codex or Claude or something?\nA link to the prompt/conversation would provide provenance and show how\nthe LLM was used. But three years from now, what information beyond the\nfact that an LLM was involved (any of them) does this offer?\n\nI can understand the benefit for the companies behind these LLMs to have\nthese attributions in OSS projects.\n\nI have done the same thing in our company repo, crediting <LLM> for\nauthoring or co-authoring or helping with a specific thing. Using a\n“people” trailer. But the intent was just to show how some LLM was\ninvolved. So I think I am going to switch to the following trailer for\nour company repo.\n\n    LLM: Yes\n\n> Signed-off-by: Tamir Duberstein <tamird@gmail.com>\n> ---\n>[snip]\n"},{"id":"544849","messageId":"CAJ-ks9nXybntsa5FCJVWSQ2u+hzxaMdrfCdL3D+vmzjO4e21kQ@mail.gmail.com","threadId":"65765","inReplyTo":"8f3bab63-3b37-4492-a39e-95e610a15a07@app.fastmail.com","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-06-07T16:07:21Z","receivedAt":"2026-06-07T16:07:59Z","isPatch":true,"body":"On Sun, Jun 7, 2026 at 12:02 PM Kristoffer Haugsbakk\n<kristofferhaugsbakk@fastmail.com> wrote:\n>\n> On Sun, Jun 7, 2026, at 17:40, Tamir Duberstein wrote:\n> >[snip]\n> > Assisted-by: Codex gpt-5.5\n>\n> This is more of a Git for Windows trailer. The Git project doesn’t\n> document its use.\n>\n> An aside here but these trailers attributing specific LLMs feels like\n> etching “Peter was here” under some table. What benefit for the project\n> does knowing that it was this version of Codex or Claude or something?\n> A link to the prompt/conversation would provide provenance and show how\n> the LLM was used. But three years from now, what information beyond the\n> fact that an LLM was involved (any of them) does this offer?\n>\n> I can understand the benefit for the companies behind these LLMs to have\n> these attributions in OSS projects.\n>\n> I have done the same thing in our company repo, crediting <LLM> for\n> authoring or co-authoring or helping with a specific thing. Using a\n> “people” trailer. But the intent was just to show how some LLM was\n> involved. So I think I am going to switch to the following trailer for\n> our company repo.\n>\n>     LLM: Yes\n\nThis all sounds reasonable to me. The kernel has started asking for\nthis trailer (https://github.com/torvalds/linux/commit/78d979db6cef557c171d6059cbce06c3db89c7ee)\nand I saw precedent in Git as recently as last month\n(https://github.com/git/git/commit/7a094d68a27e321a99c8ab6b700909e503904bd9)\nso I erred on the side of caution.\n\nI am also OK with this trailer being dropped or replaced on apply.\n"},{"id":"544850","messageId":"e42fac49-5037-4eac-b4c8-58bc62857ee2@app.fastmail.com","threadId":"65765","inReplyTo":"CAJ-ks9nXybntsa5FCJVWSQ2u+hzxaMdrfCdL3D+vmzjO4e21kQ@mail.gmail.com","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2026-06-07T16:17:29Z","receivedAt":"2026-06-07T16:17:50Z","isPatch":true,"body":"On Sun, Jun 7, 2026, at 18:07, Tamir Duberstein wrote:\n> On Sun, Jun 7, 2026 at 12:02 PM Kristoffer Haugsbakk\n> <kristofferhaugsbakk@fastmail.com> wrote:\n>>[snip]\n>>\n>> I have done the same thing in our company repo, crediting <LLM> for\n>> authoring or co-authoring or helping with a specific thing. Using a\n>> “people” trailer. But the intent was just to show how some LLM was\n>> involved. So I think I am going to switch to the following trailer for\n>> our company repo.\n>>\n>>     LLM: Yes\n>\n> This all sounds reasonable to me. The kernel has started asking for\n> this trailer\n> (https://github.com/torvalds/linux/commit/78d979db6cef557c171d6059cbce06c3db89c7ee)\n> and I saw precedent in Git as recently as last month\n> (https://github.com/git/git/commit/7a094d68a27e321a99c8ab6b700909e503904bd9)\n> so I erred on the side of caution.\n>\n> I am also OK with this trailer being dropped or replaced on apply.\n\nThe most important thing to be aware of is “Use of Artificial\nIntelligence (AI)” in `Documentation/SubmittingPatches`. :)\n\nThanks\n"},{"id":"544852","messageId":"CAJ-ks9ksEujH2Y1VQ6t8i6MW2umS_6ObCTkkTxgu73yMjHDLkg@mail.gmail.com","threadId":"65765","inReplyTo":"e42fac49-5037-4eac-b4c8-58bc62857ee2@app.fastmail.com","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-06-07T18:24:57Z","receivedAt":"2026-06-07T18:25:36Z","isPatch":true,"body":"On Sun, Jun 7, 2026 at 12:17 PM Kristoffer Haugsbakk\n<kristofferhaugsbakk@fastmail.com> wrote:\n>\n> On Sun, Jun 7, 2026, at 18:07, Tamir Duberstein wrote:\n> > On Sun, Jun 7, 2026 at 12:02 PM Kristoffer Haugsbakk\n> > <kristofferhaugsbakk@fastmail.com> wrote:\n> >>[snip]\n> >>\n> >> I have done the same thing in our company repo, crediting <LLM> for\n> >> authoring or co-authoring or helping with a specific thing. Using a\n> >> “people” trailer. But the intent was just to show how some LLM was\n> >> involved. So I think I am going to switch to the following trailer for\n> >> our company repo.\n> >>\n> >>     LLM: Yes\n> >\n> > This all sounds reasonable to me. The kernel has started asking for\n> > this trailer\n> > (https://github.com/torvalds/linux/commit/78d979db6cef557c171d6059cbce06c3db89c7ee)\n> > and I saw precedent in Git as recently as last month\n> > (https://github.com/git/git/commit/7a094d68a27e321a99c8ab6b700909e503904bd9)\n> > so I erred on the side of caution.\n> >\n> > I am also OK with this trailer being dropped or replaced on apply.\n>\n> The most important thing to be aware of is “Use of Artificial\n> Intelligence (AI)” in `Documentation/SubmittingPatches`. :)\n\nAcknowledge, thanks! Will omit the trailer on future submissions.\n"},{"id":"544919","messageId":"xmqqa4t5yyee.fsf@gitster.g","threadId":"65765","inReplyTo":"20260607-ls-files-pathspec-lstat-v1-1-8cf40b730146@gmail.com","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-08T13:06:33Z","receivedAt":"2026-06-08T13:06:35Z","isPatch":true,"body":"On Sun, Jun 7, 2026 at 11:40, Tamir Duberstein wrote:\n> show_files() checks whether each index entry is deleted or modified\n> before show_ce() applies the pathspec. prune_index() avoids most of this\n> work for pathspecs with a common directory prefix, but a top-level name\n> or leading wildcard leaves every entry to be checked.\n> \n> Match the pathspec before lstat() for the deleted and modified modes.\n> Keep the later match in show_ce() so --error-unmatch is satisfied only\n> by entries that are actually shown.\n\nAdding an extra early `match_pathspec()` check before making slow\nsystem calls like `lstat()` makes sense, especially when most of the\nindex entries need to be skipped.  But if most of them would match,\nthen we would end up doing the same match_pathspec() calls twice for\neach path, and run lstat() anyway, so you may also be able to\nconstruct a perf test that demonstrates a case where this approach\nis not a clear win (or even degradation), perhaps?\n\n> diff --git a/builtin/ls-files.c b/builtin/ls-files.c\n> index e1a22b41b9..702c607183 100644\n> --- a/builtin/ls-files.c\n> +++ b/builtin/ls-files.c\n> @@ -450,6 +450,13 @@ static void show_files(struct repository *repo, struct dir_struct *dir)\n>  \t\t\tcontinue;\n>  \t\tif (ce_skip_worktree(ce))\n>  \t\t\tcontinue;\n> +\t\t/* Only entries shown by show_ce() satisfy --error-unmatch. */\n> +\t\tif (pathspec.nr &&\n> +\t\t    !match_pathspec(repo->index, &pathspec, fullname.buf,\n> +\t\t\t\t    fullname.len, max_prefix_len, NULL,\n> +\t\t\t\t    S_ISDIR(ce->ce_mode) ||\n> +\t\t\t\t    S_ISGITLINK(ce->ce_mode)))\n> +\t\t\tcontinue;\n>  \t\tstat_err = lstat(fullname.buf, &st);\n>  \t\tif (stat_err && (errno != ENOENT && errno != ENOTDIR))\n>  \t\t\terror_errno(\"cannot lstat '%s'\", fullname.buf);\n\nHmph.  In the current code, because there is no such pre-filtering,\nshow_ce() would unconditionally recurse into active submodules when\ntold to with the \"--recurse-submodules\" flag, even if your pathspec\ncoes not match the submodule.  With this change, such a submodule\nwhose path does not match the pathspec would not even be seen by\nshow_ce().  Would it cause a change in behaviour?\n"},{"id":"544963","messageId":"CAJ-ks9njUM4TqHd=3H+aY8TCk6yG4o1yAhiSn2Tfz6oDnML20A@mail.gmail.com","threadId":"65765","inReplyTo":"xmqqa4t5yyee.fsf@gitster.g","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-06-08T19:15:53Z","receivedAt":"2026-06-08T19:16:31Z","isPatch":true,"body":"On Mon, Jun 8, 2026 at 6:06 AM Junio C Hamano <gitster@pobox.com> wrote:\n>\n> On Sun, Jun 7, 2026 at 11:40, Tamir Duberstein wrote:\n> > show_files() checks whether each index entry is deleted or modified\n> > before show_ce() applies the pathspec. prune_index() avoids most of this\n> > work for pathspecs with a common directory prefix, but a top-level name\n> > or leading wildcard leaves every entry to be checked.\n> >\n> > Match the pathspec before lstat() for the deleted and modified modes.\n> > Keep the later match in show_ce() so --error-unmatch is satisfied only\n> > by entries that are actually shown.\n>\n> Adding an extra early `match_pathspec()` check before making slow\n> system calls like `lstat()` makes sense, especially when most of the\n> index entries need to be skipped.  But if most of them would match,\n> then we would end up doing the same match_pathspec() calls twice for\n> each path, and run lstat() anyway, so you may also be able to\n> construct a perf test that demonstrates a case where this approach\n> is not a clear win (or even degradation), perhaps?\n\nYes. I added an all-matching pathspec case to p3010 and ran:\n\n    hyperfine --warmup 0 --runs 3 \\\n        'git -c core.fsmonitor=false ls-files --deleted -- \"*\"'\n\nOn a checkout with 859,940 index entries, I ran the parent and patched\nbinaries in both orders:\n\n                         parent          this commit\n  parent first elapsed    56.807 s        64.618 s\n               user        1.256 s         1.270 s\n               system     10.633 s        11.068 s\n  patched first elapsed   63.361 s        64.316 s\n                user       1.238 s         1.280 s\n                system    10.296 s        11.864 s\n\nThe added match costs 14-42 ms of user time in this case. Elapsed time\nvaries by several seconds with command order, obscuring that CPU cost.\n\nThe later match in show_ce() is reached only for entries actually found\ndeleted or modified. This case therefore exercises the extra match for\nevery index entry while still performing every lstat().\n\n>\n> > diff --git a/builtin/ls-files.c b/builtin/ls-files.c\n> > index e1a22b41b9..702c607183 100644\n> > --- a/builtin/ls-files.c\n> > +++ b/builtin/ls-files.c\n> > @@ -450,6 +450,13 @@ static void show_files(struct repository *repo, struct dir_struct *dir)\n> >                       continue;\n> >               if (ce_skip_worktree(ce))\n> >                       continue;\n> > +             /* Only entries shown by show_ce() satisfy --error-unmatch. */\n> > +             if (pathspec.nr &&\n> > +                 !match_pathspec(repo->index, &pathspec, fullname.buf,\n> > +                                 fullname.len, max_prefix_len, NULL,\n> > +                                 S_ISDIR(ce->ce_mode) ||\n> > +                                 S_ISGITLINK(ce->ce_mode)))\n> > +                     continue;\n> >               stat_err = lstat(fullname.buf, &st);\n> >               if (stat_err && (errno != ENOENT && errno != ENOTDIR))\n> >                       error_errno(\"cannot lstat '%s'\", fullname.buf);\n>\n> Hmph.  In the current code, because there is no such pre-filtering,\n> show_ce() would unconditionally recurse into active submodules when\n> told to with the \"--recurse-submodules\" flag, even if your pathspec\n> coes not match the submodule.  With this change, such a submodule\n> whose path does not match the pathspec would not even be seen by\n> show_ce().  Would it cause a change in behaviour?\n\nThis path cannot affect --recurse-submodules. cmd_ls_files() rejects\n--recurse-submodules together with either --deleted or --modified before\ncalling show_files(), and the new check is reached only for those two\nmodes. Cached and stage output continue to call show_ce() before the new\ncheck. t3007 already verifies that both combinations are rejected.\n\nGiven the 60.742 s to 1.061 s improvement for a selective pathspec, I\nthink this small CPU cost for an all-matching pathspec is a worthwhile\ntradeoff. What do you think?\n\nThanks for the review!\nTamir\n"},{"id":"544977","messageId":"20260608230315.GC340696@coredump.intra.peff.net","threadId":"65765","inReplyTo":"xmqqa4t5yyee.fsf@gitster.g","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-08T23:03:15Z","receivedAt":"2026-06-08T23:03:16Z","isPatch":true,"body":"On Mon, Jun 08, 2026 at 06:06:33AM -0700, Junio C Hamano wrote:\n\n> On Sun, Jun 7, 2026 at 11:40, Tamir Duberstein wrote:\n> > show_files() checks whether each index entry is deleted or modified\n> > before show_ce() applies the pathspec. prune_index() avoids most of this\n> > work for pathspecs with a common directory prefix, but a top-level name\n> > or leading wildcard leaves every entry to be checked.\n> > \n> > Match the pathspec before lstat() for the deleted and modified modes.\n> > Keep the later match in show_ce() so --error-unmatch is satisfied only\n> > by entries that are actually shown.\n> \n> Adding an extra early `match_pathspec()` check before making slow\n> system calls like `lstat()` makes sense, especially when most of the\n> index entries need to be skipped.  But if most of them would match,\n> then we would end up doing the same match_pathspec() calls twice for\n> each path, and run lstat() anyway, so you may also be able to\n> construct a perf test that demonstrates a case where this approach\n> is not a clear win (or even degradation), perhaps?\n\nThe patchspec matching is linear in the number of pathspecs, so it's\neasy to get quadratic-ish results by just asking about:\n\n  git ls-files -- $(git ls-files)\n\nSo that probably provides an easy regression demonstration for this\npatch.\n\nI don't know how much it matters in the real world. That command is\n_already_ slow, and mostly people don't care that much. Long ago I had\npatches to build a trie of literal pathspecs, with the intent that\nblame-tree/last-modified could use the pathspec mechanism, but\nultimately we took the code in a different direction. And nobody really\ncomplained about it since. ;)\n\n-Peff\n"},{"id":"544979","messageId":"20260608232516.GA357822@coredump.intra.peff.net","threadId":"65765","inReplyTo":"20260608230315.GC340696@coredump.intra.peff.net","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-08T23:25:16Z","receivedAt":"2026-06-08T23:25:18Z","isPatch":true,"body":"On Mon, Jun 08, 2026 at 07:03:15PM -0400, Jeff King wrote:\n\n> > Adding an extra early `match_pathspec()` check before making slow\n> > system calls like `lstat()` makes sense, especially when most of the\n> > index entries need to be skipped.  But if most of them would match,\n> > then we would end up doing the same match_pathspec() calls twice for\n> > each path, and run lstat() anyway, so you may also be able to\n> > construct a perf test that demonstrates a case where this approach\n> > is not a clear win (or even degradation), perhaps?\n> \n> The patchspec matching is linear in the number of pathspecs, so it's\n> easy to get quadratic-ish results by just asking about:\n> \n>   git ls-files -- $(git ls-files)\n> \n> So that probably provides an easy regression demonstration for this\n> patch.\n\nAh, yeah, it is easy to demonstrate. Making a repo of size $n like this:\n\n  n=10000\n  git init\n  for i in $(seq $n); do\n    echo $i >file$i\n  done\n  git add .\n  git commit -m foo\n\nIf we then run:\n\n  time git ls-files -- $(git ls-files) >/dev/null\n\nthen n=1000 takes ~15ms for me, but n=10000 takes ~800ms. So that shows\nthe slowdown of the existing pathspec code as the number of pathspecs\ngrows.\n\nWith this patch, starting with n=10000 and adding in \"-m\" (which\ntriggers the code in this patch), like:\n\n  time git ls-files -m -- $(git ls-files) >/dev/null\n\nthe time goes from ~15ms (without the patch) to ~800ms with it. Which\nmakes sense. Nothing is modified, so the current code which puts the\nlstat() check first eliminates each entry before we even consider\npathspecs. So it doesn't hit the slow case at all.\n\nBut after the patch, we do a preliminary pathspec match and\npay the cost.\n\nSo it really is a question of how many items are actually modified, the\ncost of lstat(), and the cost of pathspec matching (which varies with\nthe size of the pathspec).\n\nBut like I said, this is kind of a silly case. If it actually starts to\nmatter in the real world, I think it may be more productive to make the\npathspec code scale better.\n\n-Peff\n"},{"id":"544989","messageId":"CAJ-ks9k5ywxoAuobQpjLUyKt9QJQjkUhfbdwEr2s_yQLVEksDA@mail.gmail.com","threadId":"65765","inReplyTo":"20260608232516.GA357822@coredump.intra.peff.net","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-06-09T00:13:37Z","receivedAt":"2026-06-09T00:14:15Z","isPatch":true,"body":"On Mon, Jun 8, 2026 at 4:25 PM Jeff King <peff@peff.net> wrote:\n>\n> On Mon, Jun 08, 2026 at 07:03:15PM -0400, Jeff King wrote:\n>\n> > > Adding an extra early `match_pathspec()` check before making slow\n> > > system calls like `lstat()` makes sense, especially when most of the\n> > > index entries need to be skipped.  But if most of them would match,\n> > > then we would end up doing the same match_pathspec() calls twice for\n> > > each path, and run lstat() anyway, so you may also be able to\n> > > construct a perf test that demonstrates a case where this approach\n> > > is not a clear win (or even degradation), perhaps?\n> >\n> > The patchspec matching is linear in the number of pathspecs, so it's\n> > easy to get quadratic-ish results by just asking about:\n> >\n> >   git ls-files -- $(git ls-files)\n> >\n> > So that probably provides an easy regression demonstration for this\n> > patch.\n>\n> Ah, yeah, it is easy to demonstrate. Making a repo of size $n like this:\n>\n>   n=10000\n>   git init\n>   for i in $(seq $n); do\n>     echo $i >file$i\n>   done\n>   git add .\n>   git commit -m foo\n>\n> If we then run:\n>\n>   time git ls-files -- $(git ls-files) >/dev/null\n>\n> then n=1000 takes ~15ms for me, but n=10000 takes ~800ms. So that shows\n> the slowdown of the existing pathspec code as the number of pathspecs\n> grows.\n>\n> With this patch, starting with n=10000 and adding in \"-m\" (which\n> triggers the code in this patch), like:\n>\n>   time git ls-files -m -- $(git ls-files) >/dev/null\n>\n> the time goes from ~15ms (without the patch) to ~800ms with it. Which\n> makes sense. Nothing is modified, so the current code which puts the\n> lstat() check first eliminates each entry before we even consider\n> pathspecs. So it doesn't hit the slow case at all.\n>\n> But after the patch, we do a preliminary pathspec match and\n> pay the cost.\n>\n> So it really is a question of how many items are actually modified, the\n> cost of lstat(), and the cost of pathspec matching (which varies with\n> the size of the pathspec).\n>\n> But like I said, this is kind of a silly case. If it actually starts to\n> matter in the real world, I think it may be more productive to make the\n> pathspec code scale better.\n\nYeah, agreed. Still, it exposed an easy-to-avoid downside in this patch,\nso I limited the early match to a single pathspec in v2.\n\nWith 10,000 clean files, hyperfine measured 112.5 ms ± 6.6 ms for the\nparent and 494.1 ms ± 17.2 ms for v1. With the restriction, the patched\nversion took 104.9 ms ± 2.2 ms against 110.1 ms ± 4.1 ms for the parent.\n\nThanks for pointing it out!\n"},{"id":"552070","messageId":"63821ef1-1234-49d5-b359-8631764457e1@app.fastmail.com","threadId":"65765","inReplyTo":"8f3bab63-3b37-4492-a39e-95e610a15a07@app.fastmail.com","subject":"Re: [PATCH] ls-files: filter pathspec before lstat","fromName":"Kristoffer Haugsbakk","fromEmail":"kristofferhaugsbakk@fastmail.com","sentAt":"2026-09-06T12:30:50Z","receivedAt":"2026-09-06T12:31:15Z","isPatch":true,"body":"On Sun, Jun 7, 2026, at 18:02, Kristoffer Haugsbakk wrote:\n> On Sun, Jun 7, 2026, at 17:40, Tamir Duberstein wrote:\n>>[snip]\n>> Assisted-by: Codex gpt-5.5\n>\n> This is more of a Git for Windows trailer. The Git project doesn’t\n> document its use.\n>\n> An aside here but these trailers attributing specific LLMs feels like\n> etching “Peter was here” under some table. What benefit for the project\n> does knowing that it was this version of Codex or Claude or something?\n> A link to the prompt/conversation would provide provenance and show how\n> the LLM was used. But three years from now, what information beyond the\n> fact that an LLM was involved (any of them) does this offer?\n>\n> I can understand the benefit for the companies behind these LLMs to have\n> these attributions in OSS projects.\n>\n> I have done the same thing in our company repo, crediting <LLM> for\n> authoring or co-authoring or helping with a specific thing. Using a\n> “people” trailer. But the intent was just to show how some LLM was\n> involved. So I think I am going to switch to the following trailer for\n> our company repo.\n>\n>     LLM: Yes\n>\n>> [snip]\n\nThe Linux Kernel has since then simplified the mandatory attribution to\njust “LLM”:\n\n    The requirement to identify specific models used in the Assisted-by tag\n    provides free advertising to proprietary software companies while adding\n    little or no useful information.  Change the requirement to simply:\n\n      Assisted-by: LLM\n\n    to capture the fact that an LLM was used without tracking which one.\n\nhttps://lore.kernel.org/workflows/87qzkuahlr.fsf@trenco.lwn.net/#t\n\nSee linux/816d9992 (coding-assistants: simplify attribution,\n2026-07-01).\n"}]}