{"thread":{"id":"65774","subject":"[PATCH v2] describe: limit default ref iteration to tags","startedAt":"2026-06-09T02:32:23Z","lastAt":"2026-06-11T06:49:14Z","messageCount":12,"participants":["Tamir Duberstein","Jeff King","Junio C Hamano","D. Ben Knoble","Patrick Steinhardt"],"isPatch":true,"patchVersion":2,"patchTotal":null},"messages":[{"id":"544996","messageId":"20260608-describe-tag-ref-scope-v2-1-256fd36dca32@gmail.com","threadId":"65774","inReplyTo":null,"subject":"[PATCH v2] describe: limit default ref iteration to tags","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-06-09T02:32:14Z","receivedAt":"2026-06-09T02:32:23Z","isPatch":true,"body":"Unless --all is given, get_name() rejects every ref outside refs/tags/.\nThe rejection happens only after the ref backend has enumerated the ref,\nso repositories with many other refs spend most of a simple describe\ninvocation visiting refs which cannot affect its result.\n\nCommit 8a5a1884e9 (Avoid accessing non-tag refs in git-describe unless\n--all is requested, 2008-02-24) moved this rejection before object\nlookup, but left iteration unscoped. Pass the existing refs/tags/\nrestriction to the iterator unless --all is given so the backend can\navoid unrelated refs.\n\nThe benchmark checkout had 120,532 refs, of which 330 were tags. With\n`$repo` naming the checkout, `$commit` an exactly tagged commit, and\n`$parent` and `$this` the two binaries, I ran:\n\n    hyperfine --warmup 3 --runs 15 \\\n        --command-name parent \\\n        '$parent -C $repo describe --exact-match $commit' \\\n        --command-name 'this commit' \\\n        '$this -C $repo describe --exact-match $commit'\n\nThe results were:\n\n    Benchmark 1: parent\n      Time (mean ± σ):     171.7 ms ±  18.5 ms    [User: 23.9 ms, System: 133.6 ms]\n      Range (min … max):   142.3 ms … 198.3 ms    15 runs\n\n    Benchmark 2: this commit\n      Time (mean ± σ):       9.9 ms ±   1.1 ms    [User: 3.3 ms, System: 4.7 ms]\n      Range (min … max):     8.8 ms …  13.1 ms    15 runs\n\n    Summary\n      this commit ran\n       17.35 ± 2.63 times faster than parent\n\nBoth revisions were built with -O3, -mcpu=native, and ThinLTO using\nApple clang 21.0.0 on macOS 26.5. The machine was a MacBook Pro\n(Mac16,6) with a 16-core Apple M4 Max (12 performance and four\nefficiency cores) and 128 GB RAM.\n\nSigned-off-by: Tamir Duberstein <tamird@gmail.com>\n---\nChanges in v2:\n- Exercise the performance test with both ref backends.\n- Keep the ref count local to its setup test.\n- Report native hyperfine output for an exact-tag lookup.\n- Link to v1: https://patch.msgid.link/20260607-describe-tag-ref-scope-v1-1-653d232b86b5@gmail.com\n---\n builtin/describe.c       |  3 +++\n t/perf/p6100-describe.sh | 15 +++++++++++++++\n 2 files changed, 18 insertions(+)\n\ndiff --git a/builtin/describe.c b/builtin/describe.c\nindex 1c47d7c0b7..3532c8ff22 100644\n--- a/builtin/describe.c\n+++ b/builtin/describe.c\n@@ -740,6 +740,9 @@ int cmd_describe(int argc,\n \t\treturn ret;\n \t}\n \n+\tif (!all)\n+\t\tfor_each_ref_opts.prefix = \"refs/tags/\";\n+\n \thashmap_init(&names, commit_name_neq, NULL, 0);\n \trefs_for_each_ref_ext(get_main_ref_store(the_repository),\n \t\t\t      get_name, NULL, &for_each_ref_opts);\ndiff --git a/t/perf/p6100-describe.sh b/t/perf/p6100-describe.sh\nindex 069f91ce49..ed9f1abe18 100755\n--- a/t/perf/p6100-describe.sh\n+++ b/t/perf/p6100-describe.sh\n@@ -27,4 +27,19 @@ test_perf 'describe HEAD with one tag' '\n \tgit describe --match=new HEAD\n '\n \n+test_expect_success 'set up many unrelated refs' '\n+\tref_count=10000 &&\n+\tgit tag -m tip tip HEAD &&\n+\tfor i in $(test_seq $ref_count)\n+\tdo\n+\t\tprintf \"create refs/heads/describe-perf/%05d HEAD\\n\" $i ||\n+\t\treturn 1\n+\tdone >instructions &&\n+\tgit update-ref --stdin <instructions\n+'\n+\n+test_perf 'describe exact tag with many unrelated refs' '\n+\tgit describe --exact-match HEAD\n+'\n+\n test_done\n\n---\nbase-commit: 9ac3f193c05c2237e2b14ebaa1149e9fc8a1abe0\nchange-id: 20260607-describe-tag-ref-scope-7d00ae140a58\n\nBest regards,\n--  \nTamir Duberstein <tamird@gmail.com>\n\n"},{"id":"545062","messageId":"20260609110957.GB1509396@coredump.intra.peff.net","threadId":"65774","inReplyTo":"20260608-describe-tag-ref-scope-v2-1-256fd36dca32@gmail.com","subject":"Re: [PATCH v2] describe: limit default ref iteration to tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-09T11:09:57Z","receivedAt":"2026-06-09T11:09:59Z","isPatch":true,"body":"On Mon, Jun 08, 2026 at 07:32:14PM -0700, Tamir Duberstein wrote:\n\n> The benchmark checkout had 120,532 refs, of which 330 were tags. With\n> `$repo` naming the checkout, `$commit` an exactly tagged commit, and\n> `$parent` and `$this` the two binaries, I ran:\n> \n>     hyperfine --warmup 3 --runs 15 \\\n>         --command-name parent \\\n>         '$parent -C $repo describe --exact-match $commit' \\\n>         --command-name 'this commit' \\\n>         '$this -C $repo describe --exact-match $commit'\n> \n> The results were:\n> \n>     Benchmark 1: parent\n>       Time (mean ± σ):     171.7 ms ±  18.5 ms    [User: 23.9 ms, System: 133.6 ms]\n>       Range (min … max):   142.3 ms … 198.3 ms    15 runs\n> \n>     Benchmark 2: this commit\n>       Time (mean ± σ):       9.9 ms ±   1.1 ms    [User: 3.3 ms, System: 4.7 ms]\n>       Range (min … max):     8.8 ms …  13.1 ms    15 runs\n> \n>     Summary\n>       this commit ran\n>        17.35 ± 2.63 times faster than parent\n> \n> Both revisions were built with -O3, -mcpu=native, and ThinLTO using\n> Apple clang 21.0.0 on macOS 26.5. The machine was a MacBook Pro\n> (Mac16,6) with a 16-core Apple M4 Max (12 performance and four\n> efficiency cores) and 128 GB RAM.\n\nThis patch looks fine to me, but let me pick a nit for a minute, because\nI think there is a broader conversation to be had.\n\nGiven the discussion in earlier rounds and sibling topics, I assume the\ncommit message here was AI-generated. And it's OK in the sense that it\nis describing what happened and I assume is entirely accurate. But as a\nhuman reader, it feels so much more verbose than what I'd expect, as it\nis full of semi-irrelevant details. Why set --warmup and --runs? Why\nbother with --command-name, which just means you have to show the\ncommands separately anyway? Is the amount of RAM in the machine\nimportant for this test? Surely it could be if it was absurdly tiny, but\nin general, no, I would not expect it to be.\n\nSo while it is perhaps reasonable to document every detail in case\nsomebody later wants to verify or reproduce timings, it is a little\noverwhelming when trying to tell a story, the core of which is:\n\n  In a repo with ~120k refs, ~300 of which were tags, running:\n\n    git describe --exact-match $some_tag\n\n  went from ~170ms to ~10ms, since we no longer needed to iterate all of\n  those other refs.\n\nThat has _way_ less detail, but makes the point succinctly.\n\nI dunno. I am not trying to pick apart your commit in particular, but am\nmore interested in the broader use of AI commit messages going forward.\nThis kind of verbosity is quite common in the output (from my limited\nexperience), and I think creates more work for reviewers. Should we be\nexpecting contributors to make things more concise before submitting\n(either manually or through prompting)? Or do people even agree that the\nshorter version is preferable? I could be the only one.\n\nI have a few other comments on the patch itself below.\n\n> diff --git a/builtin/describe.c b/builtin/describe.c\n> index 1c47d7c0b7..3532c8ff22 100644\n> --- a/builtin/describe.c\n> +++ b/builtin/describe.c\n> @@ -740,6 +740,9 @@ int cmd_describe(int argc,\n>  \t\treturn ret;\n>  \t}\n>  \n> +\tif (!all)\n> +\t\tfor_each_ref_opts.prefix = \"refs/tags/\";\n> +\n>  \thashmap_init(&names, commit_name_neq, NULL, 0);\n>  \trefs_for_each_ref_ext(get_main_ref_store(the_repository),\n>  \t\t\t      get_name, NULL, &for_each_ref_opts);\n\nThe code change looks fine. It creates a bit of a subtle dependency\nbetween what's happening here, and the filtering inside get_name(). But\nI think that's OK for the scope of a single command. It _might_ be\npossible to simplify the top of get_name(), since we'd no longer see\nnon-tag refs in the input. But it also may not, since we have to strip\nout the prefix anyway. It can certainly come on top as a cleanup later\nif we want.\n\n> diff --git a/t/perf/p6100-describe.sh b/t/perf/p6100-describe.sh\n\nIt is a little curious that we add a perf test here, but the commit\nmessage does not even show it off. ;)\n\nI ran it myself here and had trouble showing improvement, simply because\nit is already quite fast! I guess that's because I'm on Linux, where\nwarm-cache filesystem operations are pretty fast. Bumping $ref_count by\na factor of 10 made the \"before\" case 30ms, and after is still sub-1ms.\n\n> +test_expect_success 'set up many unrelated refs' '\n> +\tref_count=10000 &&\n> +\tgit tag -m tip tip HEAD &&\n> +\tfor i in $(test_seq $ref_count)\n> +\tdo\n> +\t\tprintf \"create refs/heads/describe-perf/%05d HEAD\\n\" $i ||\n> +\t\treturn 1\n> +\tdone >instructions &&\n> +\tgit update-ref --stdin <instructions\n> +'\n\nA few things come to mind on reading this.\n\nI have mixed feelings on sticking synthetic constructions in the t/perf\nsuite. Part of the original point was that we'd run it against real\nrepos to see how they perform. But that implies that people running it\nhave some clue about which tests may be interesting on which repos,\nwhich is hopeful at best. So we've turned to this kind of synthetic\nconstruction at times (and this is certainly not the first). It's\nprobably a reasonable tactic here.\n\nI suspect the resulting state is not all that realistic, though. If you\nhave 10,000 refs, you probably didn't make them all at once. And so in\npractice the majority of them would be packed. Sticking \"git pack-refs\n--all\" at the end might give more realistic numbers.\n\nBumping to a larger number of refs shows the effect more clearly, but at\nthe cost of making the setup take a long time (since we have to take a\nlockfile on each ref!). We could sneak around it by generating a\npacked-refs file directly, but now the test really would be\nbackend-specific. Probably better not to go there.\n\nAnd finally, the loop can be written a bit more succinctly these days\nas:\n\ndiff --git a/t/perf/p6100-describe.sh b/t/perf/p6100-describe.sh\nindex ed9f1abe18..b365dc67ee 100755\n--- a/t/perf/p6100-describe.sh\n+++ b/t/perf/p6100-describe.sh\n@@ -30,12 +30,8 @@ test_perf 'describe HEAD with one tag' '\n test_expect_success 'set up many unrelated refs' '\n \tref_count=10000 &&\n \tgit tag -m tip tip HEAD &&\n-\tfor i in $(test_seq $ref_count)\n-\tdo\n-\t\tprintf \"create refs/heads/describe-perf/%05d HEAD\\n\" $i ||\n-\t\treturn 1\n-\tdone >instructions &&\n-\tgit update-ref --stdin <instructions\n+\ttest_seq -f \"create refs/heads/describe-perf/%05d HEAD\" $ref_count |\n+\tgit update-ref --stdin\n '\n \n test_perf 'describe exact tag with many unrelated refs' '\n\n\nProbably not worth re-rolling on its own, though.\n\n-Peff\n"},{"id":"545071","messageId":"xmqqpl1zsv8s.fsf@gitster.g","threadId":"65774","inReplyTo":"20260609110957.GB1509396@coredump.intra.peff.net","subject":"Re: [PATCH v2] describe: limit default ref iteration to tags","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-09T13:23:31Z","receivedAt":"2026-06-09T13:23:34Z","isPatch":true,"body":"Jeff King <peff@peff.net> writes:\n\n> So while it is perhaps reasonable to document every detail in case\n> somebody later wants to verify or reproduce timings, it is a little\n> overwhelming when trying to tell a story, the core of which is:\n>\n>   In a repo with ~120k refs, ~300 of which were tags, running:\n>\n>     git describe --exact-match $some_tag\n>\n>   went from ~170ms to ~10ms, since we no longer needed to iterate all of\n>   those other refs.\n>\n> That has _way_ less detail, but makes the point succinctly.\n>\n> I dunno. I am not trying to pick apart your commit in particular, but am\n> more interested in the broader use of AI commit messages going forward.\n> This kind of verbosity is quite common in the output (from my limited\n> experience), and I think creates more work for reviewers. Should we be\n> expecting contributors to make things more concise before submitting\n> (either manually or through prompting)? Or do people even agree that the\n> shorter version is preferable? I could be the only one.\n\nCount me in.  You are the one who often gives us a patch with 60\nlines that explains a single line change, but I haven't found these\n60 lines are _overly verbose_ in the same way as AI generated log\nmessages.\n"},{"id":"545074","messageId":"CALnO6CB-9a=P4Os90978YzEH=3iYEHwSbG2oLv9sxVBjBfchMA@mail.gmail.com","threadId":"65774","inReplyTo":"20260609110957.GB1509396@coredump.intra.peff.net","subject":"Re: [PATCH v2] describe: limit default ref iteration to tags","fromName":"D. Ben Knoble","fromEmail":"ben.knoble@gmail.com","sentAt":"2026-06-09T13:40:25Z","receivedAt":"2026-06-09T13:40:38Z","isPatch":true,"body":"On Tue, Jun 9, 2026 at 7:10 AM Jeff King <peff@peff.net> wrote:\n>\n> On Mon, Jun 08, 2026 at 07:32:14PM -0700, Tamir Duberstein wrote:\n>\n> > The benchmark checkout had 120,532 refs, of which 330 were tags. With\n> > `$repo` naming the checkout, `$commit` an exactly tagged commit, and\n> > `$parent` and `$this` the two binaries, I ran:\n> >\n> >     hyperfine --warmup 3 --runs 15 \\\n> >         --command-name parent \\\n> >         '$parent -C $repo describe --exact-match $commit' \\\n> >         --command-name 'this commit' \\\n> >         '$this -C $repo describe --exact-match $commit'\n> >\n> > The results were:\n> >\n> >     Benchmark 1: parent\n> >       Time (mean ± σ):     171.7 ms ±  18.5 ms    [User: 23.9 ms, System: 133.6 ms]\n> >       Range (min … max):   142.3 ms … 198.3 ms    15 runs\n> >\n> >     Benchmark 2: this commit\n> >       Time (mean ± σ):       9.9 ms ±   1.1 ms    [User: 3.3 ms, System: 4.7 ms]\n> >       Range (min … max):     8.8 ms …  13.1 ms    15 runs\n> >\n> >     Summary\n> >       this commit ran\n> >        17.35 ± 2.63 times faster than parent\n> >\n> > Both revisions were built with -O3, -mcpu=native, and ThinLTO using\n> > Apple clang 21.0.0 on macOS 26.5. The machine was a MacBook Pro\n> > (Mac16,6) with a 16-core Apple M4 Max (12 performance and four\n> > efficiency cores) and 128 GB RAM.\n>\n> This patch looks fine to me, but let me pick a nit for a minute, because\n> I think there is a broader conversation to be had.\n>\n> Given the discussion in earlier rounds and sibling topics, I assume the\n> commit message here was AI-generated. And it's OK in the sense that it\n> is describing what happened and I assume is entirely accurate. But as a\n> human reader, it feels so much more verbose than what I'd expect, as it\n> is full of semi-irrelevant details. Why set --warmup and --runs? Why\n> bother with --command-name, which just means you have to show the\n> commands separately anyway? Is the amount of RAM in the machine\n> important for this test? Surely it could be if it was absurdly tiny, but\n> in general, no, I would not expect it to be.\n\n[You probably know this] It is common in academic papers to report\nbenchmarks with details about the hardware and how they were run to\ncontextualize the results and help with reproducibility.\n\nOf course, Git's commits do not form an academic paper… so I have no\nreal opinion on what to see here. But I've seen a few other mails\nwhere having perf test outputs or similar was suggested (maybe that\nwas to be reserved for the cover letter? idk).\n\n_If_ we show all the hyperfine details, I think it's reasonable to use\n--command-name to make distinguishing the versions easy, unless it's\nobvious from the path/to/git in each benchmark (which I think I've\nseen from Peff's benchmark reports before?).\n\nSomeone with better lore skills can probably dig up a few exemplars of\nhow to write about performance in a commit message?\n\n-- \nD. Ben Knoble\n"},{"id":"545078","messageId":"CAJ-ks9kz5JGFSF21aOhuXfgsJ+5aa5xE69RPT2Vhn-CRGyHZ6A@mail.gmail.com","threadId":"65774","inReplyTo":"20260609110957.GB1509396@coredump.intra.peff.net","subject":"Re: [PATCH v2] describe: limit default ref iteration to tags","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-06-09T14:44:12Z","receivedAt":"2026-06-09T14:44:51Z","isPatch":true,"body":"On Tue, Jun 9, 2026 at 4:09 AM Jeff King <peff@peff.net> wrote:\n>\n> On Mon, Jun 08, 2026 at 07:32:14PM -0700, Tamir Duberstein wrote:\n>\n> > The benchmark checkout had 120,532 refs, of which 330 were tags. With\n> > `$repo` naming the checkout, `$commit` an exactly tagged commit, and\n> > `$parent` and `$this` the two binaries, I ran:\n> >\n> >     hyperfine --warmup 3 --runs 15 \\\n> >         --command-name parent \\\n> >         '$parent -C $repo describe --exact-match $commit' \\\n> >         --command-name 'this commit' \\\n> >         '$this -C $repo describe --exact-match $commit'\n> >\n> > The results were:\n> >\n> >     Benchmark 1: parent\n> >       Time (mean ± σ):     171.7 ms ±  18.5 ms    [User: 23.9 ms, System: 133.6 ms]\n> >       Range (min … max):   142.3 ms … 198.3 ms    15 runs\n> >\n> >     Benchmark 2: this commit\n> >       Time (mean ± σ):       9.9 ms ±   1.1 ms    [User: 3.3 ms, System: 4.7 ms]\n> >       Range (min … max):     8.8 ms …  13.1 ms    15 runs\n> >\n> >     Summary\n> >       this commit ran\n> >        17.35 ± 2.63 times faster than parent\n> >\n> > Both revisions were built with -O3, -mcpu=native, and ThinLTO using\n> > Apple clang 21.0.0 on macOS 26.5. The machine was a MacBook Pro\n> > (Mac16,6) with a 16-core Apple M4 Max (12 performance and four\n> > efficiency cores) and 128 GB RAM.\n>\n> This patch looks fine to me, but let me pick a nit for a minute, because\n> I think there is a broader conversation to be had.\n\nJust to say from the start: I appreciate you taking the time to discuss this.\n\n>\n> Given the discussion in earlier rounds and sibling topics, I assume the\n> commit message here was AI-generated. And it's OK in the sense that it\n> is describing what happened and I assume is entirely accurate. But as a\n> human reader, it feels so much more verbose than what I'd expect, as it\n> is full of semi-irrelevant details. Why set --warmup and --runs? Why\n> bother with --command-name, which just means you have to show the\n> commands separately anyway? Is the amount of RAM in the machine\n> important for this test? Surely it could be if it was absurdly tiny, but\n> in general, no, I would not expect it to be.\n\nWell, the details matter in case some human reader knows something I\ndon't, or wants to reproduce the findings and observes something\ncompletely different - they aught to be able to reconstruct my\nenvironment.\n\nThe command-name flag is needed; without it the output of the\nhyperfine would include local paths and would require post-processing\nto include in the message.\n\n>\n> So while it is perhaps reasonable to document every detail in case\n> somebody later wants to verify or reproduce timings, it is a little\n> overwhelming when trying to tell a story, the core of which is:\n>\n>   In a repo with ~120k refs, ~300 of which were tags, running:\n>\n>     git describe --exact-match $some_tag\n>\n>   went from ~170ms to ~10ms, since we no longer needed to iterate all of\n>   those other refs.\n>\n> That has _way_ less detail, but makes the point succinctly.\n\nI don't disagree. Ultimately, it is a matter of maintainer preference,\nand I'm happy to follow (and instruct the AI to follow) the\npreferences described in this thread.\n\n>\n> I dunno. I am not trying to pick apart your commit in particular, but am\n> more interested in the broader use of AI commit messages going forward.\n> This kind of verbosity is quite common in the output (from my limited\n> experience), and I think creates more work for reviewers. Should we be\n> expecting contributors to make things more concise before submitting\n> (either manually or through prompting)? Or do people even agree that the\n> shorter version is preferable? I could be the only one.\n\nThe AI does what you tell it; in this case I was telling it to follow\nthe precedent in the repo and to ensure its claims are always cited.\nI'll tune it for succinct output going forward.\n\n>\n> I have a few other comments on the patch itself below.\n>\n> > diff --git a/builtin/describe.c b/builtin/describe.c\n> > index 1c47d7c0b7..3532c8ff22 100644\n> > --- a/builtin/describe.c\n> > +++ b/builtin/describe.c\n> > @@ -740,6 +740,9 @@ int cmd_describe(int argc,\n> >               return ret;\n> >       }\n> >\n> > +     if (!all)\n> > +             for_each_ref_opts.prefix = \"refs/tags/\";\n> > +\n> >       hashmap_init(&names, commit_name_neq, NULL, 0);\n> >       refs_for_each_ref_ext(get_main_ref_store(the_repository),\n> >                             get_name, NULL, &for_each_ref_opts);\n>\n> The code change looks fine. It creates a bit of a subtle dependency\n> between what's happening here, and the filtering inside get_name(). But\n> I think that's OK for the scope of a single command. It _might_ be\n> possible to simplify the top of get_name(), since we'd no longer see\n> non-tag refs in the input. But it also may not, since we have to strip\n> out the prefix anyway. It can certainly come on top as a cleanup later\n> if we want.\n>\n> > diff --git a/t/perf/p6100-describe.sh b/t/perf/p6100-describe.sh\n>\n> It is a little curious that we add a perf test here, but the commit\n> message does not even show it off. ;)\n>\n> I ran it myself here and had trouble showing improvement, simply because\n> it is already quite fast! I guess that's because I'm on Linux, where\n> warm-cache filesystem operations are pretty fast. Bumping $ref_count by\n> a factor of 10 made the \"before\" case 30ms, and after is still sub-1ms.\n>\n> > +test_expect_success 'set up many unrelated refs' '\n> > +     ref_count=10000 &&\n> > +     git tag -m tip tip HEAD &&\n> > +     for i in $(test_seq $ref_count)\n> > +     do\n> > +             printf \"create refs/heads/describe-perf/%05d HEAD\\n\" $i ||\n> > +             return 1\n> > +     done >instructions &&\n> > +     git update-ref --stdin <instructions\n> > +'\n>\n> A few things come to mind on reading this.\n>\n> I have mixed feelings on sticking synthetic constructions in the t/perf\n> suite. Part of the original point was that we'd run it against real\n> repos to see how they perform. But that implies that people running it\n> have some clue about which tests may be interesting on which repos,\n> which is hopeful at best. So we've turned to this kind of synthetic\n> construction at times (and this is certainly not the first). It's\n> probably a reasonable tactic here.\n>\n> I suspect the resulting state is not all that realistic, though. If you\n> have 10,000 refs, you probably didn't make them all at once. And so in\n> practice the majority of them would be packed. Sticking \"git pack-refs\n> --all\" at the end might give more realistic numbers.\n>\n> Bumping to a larger number of refs shows the effect more clearly, but at\n> the cost of making the setup take a long time (since we have to take a\n> lockfile on each ref!). We could sneak around it by generating a\n> packed-refs file directly, but now the test really would be\n> backend-specific. Probably better not to go there.\n>\n> And finally, the loop can be written a bit more succinctly these days\n> as:\n>\n> diff --git a/t/perf/p6100-describe.sh b/t/perf/p6100-describe.sh\n> index ed9f1abe18..b365dc67ee 100755\n> --- a/t/perf/p6100-describe.sh\n> +++ b/t/perf/p6100-describe.sh\n> @@ -30,12 +30,8 @@ test_perf 'describe HEAD with one tag' '\n>  test_expect_success 'set up many unrelated refs' '\n>         ref_count=10000 &&\n>         git tag -m tip tip HEAD &&\n> -       for i in $(test_seq $ref_count)\n> -       do\n> -               printf \"create refs/heads/describe-perf/%05d HEAD\\n\" $i ||\n> -               return 1\n> -       done >instructions &&\n> -       git update-ref --stdin <instructions\n> +       test_seq -f \"create refs/heads/describe-perf/%05d HEAD\" $ref_count |\n> +       git update-ref --stdin\n>  '\n>\n>  test_perf 'describe exact tag with many unrelated refs' '\n>\n>\n> Probably not worth re-rolling on its own, though.\n\nThe suggested changes seem reasonable to me. Certainly I am happy to\nmake them, and re-rolls are cheap. Do let me know explicitly if you'd\nlike that done.\n\n>\n> -Peff\n\nThanks for your time! I really appreciate it.\n"},{"id":"545108","messageId":"aika_Q0rWhcI6eXR@pks.im","threadId":"65774","inReplyTo":"20260609110957.GB1509396@coredump.intra.peff.net","subject":"Re: [PATCH v2] describe: limit default ref iteration to tags","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2026-06-10T08:08:51Z","receivedAt":"2026-06-10T08:09:03Z","isPatch":true,"body":"On Tue, Jun 09, 2026 at 07:09:57AM -0400, Jeff King wrote:\n> On Mon, Jun 08, 2026 at 07:32:14PM -0700, Tamir Duberstein wrote:\n> \n> > The benchmark checkout had 120,532 refs, of which 330 were tags. With\n> > `$repo` naming the checkout, `$commit` an exactly tagged commit, and\n> > `$parent` and `$this` the two binaries, I ran:\n> > \n> >     hyperfine --warmup 3 --runs 15 \\\n> >         --command-name parent \\\n> >         '$parent -C $repo describe --exact-match $commit' \\\n> >         --command-name 'this commit' \\\n> >         '$this -C $repo describe --exact-match $commit'\n> > \n> > The results were:\n> > \n> >     Benchmark 1: parent\n> >       Time (mean ± σ):     171.7 ms ±  18.5 ms    [User: 23.9 ms, System: 133.6 ms]\n> >       Range (min … max):   142.3 ms … 198.3 ms    15 runs\n> > \n> >     Benchmark 2: this commit\n> >       Time (mean ± σ):       9.9 ms ±   1.1 ms    [User: 3.3 ms, System: 4.7 ms]\n> >       Range (min … max):     8.8 ms …  13.1 ms    15 runs\n> > \n> >     Summary\n> >       this commit ran\n> >        17.35 ± 2.63 times faster than parent\n> > \n> > Both revisions were built with -O3, -mcpu=native, and ThinLTO using\n> > Apple clang 21.0.0 on macOS 26.5. The machine was a MacBook Pro\n> > (Mac16,6) with a 16-core Apple M4 Max (12 performance and four\n> > efficiency cores) and 128 GB RAM.\n> \n> This patch looks fine to me, but let me pick a nit for a minute, because\n> I think there is a broader conversation to be had.\n> \n> Given the discussion in earlier rounds and sibling topics, I assume the\n> commit message here was AI-generated. And it's OK in the sense that it\n> is describing what happened and I assume is entirely accurate. But as a\n> human reader, it feels so much more verbose than what I'd expect, as it\n> is full of semi-irrelevant details. Why set --warmup and --runs? Why\n> bother with --command-name, which just means you have to show the\n> commands separately anyway? Is the amount of RAM in the machine\n> important for this test? Surely it could be if it was absurdly tiny, but\n> in general, no, I would not expect it to be.\n\nI agree. Earlier this week I also drafted a message that was going down\nthis angle, but I think I didn't end up sending it to the mailing list.\nOr at least I'm not able to find it anymore.\n\nTo me the biggest problem is not the verbosity, even though it _is_\noverly verbose. The bigger problem though is the incoherence of the\nstory that the commit message is trying to tell where it jumps around\nrandomly. It almost feels like rambling to me, and that makes it\nextremely hard to follow the narrative and figure out what the message\neven wants to tell the reader in the first place.\n\n[snip]\n> I dunno. I am not trying to pick apart your commit in particular, but am\n> more interested in the broader use of AI commit messages going forward.\n> This kind of verbosity is quite common in the output (from my limited\n> experience), and I think creates more work for reviewers. Should we be\n> expecting contributors to make things more concise before submitting\n> (either manually or through prompting)? Or do people even agree that the\n> shorter version is preferable? I could be the only one.\n\nI very much think that we should and even have to expect that\ncontributors adapt, because if we don't we will basically reinforce\nwhatever AI is doing right now and increase the load on reviewers even\nmore.\n\nI also think that we should reserve the right to reject a patch series\ncompletely in case we notice that we're basically just talking to a\nmiddleman that sits between an AI prompt and us (please note that I\ndon't refer to this patch series specifically, this is more of a general\nstatement). My assumption is that this will become more important as AI\ngets established in more workflows. The number of patch series that look\nsane on the surface but that are utter garbage will very likely increase\nquite significantly going forward.\n\nPatrick\n"},{"id":"545176","messageId":"xmqq4ijawatw.fsf@gitster.g","threadId":"65774","inReplyTo":"aika_Q0rWhcI6eXR@pks.im","subject":"Re: [PATCH v2] describe: limit default ref iteration to tags","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-10T17:43:07Z","receivedAt":"2026-06-10T17:43:10Z","isPatch":true,"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> I very much think that we should and even have to expect that\n> contributors adapt, because if we don't we will basically reinforce\n> whatever AI is doing right now and increase the load on reviewers even\n> more.\n>\n> I also think that we should reserve the right to reject a patch series\n> completely in case we notice that we're basically just talking to a\n> middleman that sits between an AI prompt and us (please note that I\n> don't refer to this patch series specifically, this is more of a general\n> statement).\n\nSounds sensible.  Something like this may make a good starting\npoint.\n\ndiff --git c/Documentation/SubmittingPatches w/Documentation/SubmittingPatches\nindex 176567738d..2fd7f6f9e6 100644\n--- c/Documentation/SubmittingPatches\n+++ w/Documentation/SubmittingPatches\n@@ -499,6 +499,12 @@ checking for obvious mistakes, things that can be improved, things\n that don’t match our style, guidelines or our feedback, before sending\n it to us.\n \n+We reserve the right to reject a patch series completely when we\n+notice that we are basically just talking to a middleman that sits\n+between an AI prompt and us, without having a better understanding of\n+the subject matter than the AI agent being used.\n+\n+\n [[git-tools]]\n === Generate your patch using Git tools out of your commits.\n \n"},{"id":"545177","messageId":"xmqqzf12uw53.fsf@gitster.g","threadId":"65774","inReplyTo":"CAJ-ks9kz5JGFSF21aOhuXfgsJ+5aa5xE69RPT2Vhn-CRGyHZ6A@mail.gmail.com","subject":"Re: [PATCH v2] describe: limit default ref iteration to tags","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-06-10T17:45:44Z","receivedAt":"2026-06-10T17:45:46Z","isPatch":true,"body":"Tamir Duberstein <tamird@gmail.com> writes:\n\n> On Tue, Jun 9, 2026 at 4:09 AM Jeff King <peff@peff.net> wrote:\n>>\n>> Probably not worth re-rolling on its own, though.\n>\n> The suggested changes seem reasonable to me. Certainly I am happy to\n> make them, and re-rolls are cheap. Do let me know explicitly if you'd\n> like that done.\n\nLet's see how much more pleasant to read such an updated version ;-)\n\nThanks.\n"},{"id":"545179","messageId":"20260610-describe-tag-ref-scope-v3-1-5aa63ab279f7@gmail.com","threadId":"65774","inReplyTo":"20260608-describe-tag-ref-scope-v2-1-256fd36dca32@gmail.com","subject":"[PATCH v3] describe: limit default ref iteration to tags","fromName":"Tamir Duberstein","fromEmail":"tamird@gmail.com","sentAt":"2026-06-10T18:50:01Z","receivedAt":"2026-06-10T18:50:07Z","isPatch":true,"body":"Without --all, git describe ignores refs outside refs/tags/. Commit\n8a5a1884e9 (Avoid accessing non-tag refs in git-describe unless --all is\nrequested, 2008-02-24) moved this check ahead of object lookup. That\navoided loading objects for irrelevant refs, but the backend still has\nto yield every ref before get_name() can reject it.\n\nPass refs/tags/ to the iterator so the backend can avoid visiting those\nrefs in the first place.\n\nThe new perf test creates 10,000 unrelated packed refs. It measures:\n\n    git describe --exact-match HEAD\n\nThe runtime drops from 0.03(0.01+0.01) to 0.02(0.00+0.00). In a\nrepository with 120,532 refs but only 330 tags, the same command went\nfrom 171.7 ms to 9.9 ms.\n\nSigned-off-by: Tamir Duberstein <tamird@gmail.com>\n---\nChanges in v3:\n- Pack the synthetic refs to better match repositories with many refs.\n- Generate update-ref input with test_seq -f.\n- Shorten the commit message and report the p6100.6 result.\n- Link to v2: https://patch.msgid.link/20260608-describe-tag-ref-scope-v2-1-256fd36dca32@gmail.com\n\nChanges in v2:\n- Exercise the performance test with both ref backends.\n- Keep the ref count local to its setup test.\n- Report native hyperfine output for an exact-tag lookup.\n- Link to v1: https://patch.msgid.link/20260607-describe-tag-ref-scope-v1-1-653d232b86b5@gmail.com\n---\n builtin/describe.c       |  3 +++\n t/perf/p6100-describe.sh | 12 ++++++++++++\n 2 files changed, 15 insertions(+)\n\ndiff --git a/builtin/describe.c b/builtin/describe.c\nindex 1c47d7c0b7..3532c8ff22 100644\n--- a/builtin/describe.c\n+++ b/builtin/describe.c\n@@ -740,6 +740,9 @@ int cmd_describe(int argc,\n \t\treturn ret;\n \t}\n \n+\tif (!all)\n+\t\tfor_each_ref_opts.prefix = \"refs/tags/\";\n+\n \thashmap_init(&names, commit_name_neq, NULL, 0);\n \trefs_for_each_ref_ext(get_main_ref_store(the_repository),\n \t\t\t      get_name, NULL, &for_each_ref_opts);\ndiff --git a/t/perf/p6100-describe.sh b/t/perf/p6100-describe.sh\nindex 069f91ce49..b1c61529bb 100755\n--- a/t/perf/p6100-describe.sh\n+++ b/t/perf/p6100-describe.sh\n@@ -27,4 +27,16 @@ test_perf 'describe HEAD with one tag' '\n \tgit describe --match=new HEAD\n '\n \n+test_expect_success 'set up many unrelated refs' '\n+\tref_count=10000 &&\n+\tgit tag -m tip tip HEAD &&\n+\ttest_seq -f \"create refs/heads/describe-perf/%05d HEAD\" $ref_count |\n+\tgit update-ref --stdin &&\n+\tgit pack-refs --all\n+'\n+\n+test_perf 'describe exact tag with many unrelated refs' '\n+\tgit describe --exact-match HEAD\n+'\n+\n test_done\n\n---\nbase-commit: 9ac3f193c05c2237e2b14ebaa1149e9fc8a1abe0\nchange-id: 20260607-describe-tag-ref-scope-7d00ae140a58\n\nBest regards,\n--  \nTamir Duberstein <tamird@gmail.com>\n\n"},{"id":"545233","messageId":"20260611063711.GA2191159@coredump.intra.peff.net","threadId":"65774","inReplyTo":"CALnO6CB-9a=P4Os90978YzEH=3iYEHwSbG2oLv9sxVBjBfchMA@mail.gmail.com","subject":"Re: [PATCH v2] describe: limit default ref iteration to tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-11T06:37:11Z","receivedAt":"2026-06-11T06:37:12Z","isPatch":true,"body":"On Tue, Jun 09, 2026 at 09:40:25AM -0400, D. Ben Knoble wrote:\n\n> > Given the discussion in earlier rounds and sibling topics, I assume the\n> > commit message here was AI-generated. And it's OK in the sense that it\n> > is describing what happened and I assume is entirely accurate. But as a\n> > human reader, it feels so much more verbose than what I'd expect, as it\n> > is full of semi-irrelevant details. Why set --warmup and --runs? Why\n> > bother with --command-name, which just means you have to show the\n> > commands separately anyway? Is the amount of RAM in the machine\n> > important for this test? Surely it could be if it was absurdly tiny, but\n> > in general, no, I would not expect it to be.\n> \n> [You probably know this] It is common in academic papers to report\n> benchmarks with details about the hardware and how they were run to\n> contextualize the results and help with reproducibility.\n\nYeah, I almost drew the same comparison in my original email. I agree\nthat having every last detail _could_ help with reproducing in the\nfuture. And that's important when producing a high quality dataset or\nacademic paper. But the tradeoff seems worse in a commit message, where\nit is easy to obscure the main point or overwhelm the reader in what is\notherwise a short-ish document.\n\n> Of course, Git's commits do not form an academic paper… so I have no\n> real opinion on what to see here. But I've seen a few other mails\n> where having perf test outputs or similar was suggested (maybe that\n> was to be reserved for the cover letter? idk).\n> \n> _If_ we show all the hyperfine details, I think it's reasonable to use\n> --command-name to make distinguishing the versions easy, unless it's\n> obvious from the path/to/git in each benchmark (which I think I've\n> seen from Peff's benchmark reports before?).\n\nYeah, I tend to copy the various versions to their own executables,\nwhich gives them short names (so you see \"./git.old vs ./git.new\" or\nsomething). That's not always completely obvious either, though.\n\nThe \"short\" example I showed may have been a little hyperbolic. I'm OK\nwith hyperfine output in general, and sometimes show it myself. It is\nkind of verbose, but occasionally the distribution of values, or user vs\nsystem vs clock times are important. I'm even OK with --command-name if\nit makes things more readable.\n\nI guess what I was really responding to is that I think it is helpful\nwhen the incoming data is cut down to the minimal set of useful details.\nThat helps a reader immediately assess what is important to the point\nbeing made. Humans tend to do this naturally because we are lazy and\ndo not want to bother typing or pasting the uninteresting details.\nProgram output (whether AI or just verbose software) has less of that\nimpulse.\n\n> Someone with better lore skills can probably dig up a few exemplars of\n> how to write about performance in a commit message?\n\nProbably searching for emails from René. :)\n\n-Peff\n"},{"id":"545234","messageId":"20260611064117.GB2191159@coredump.intra.peff.net","threadId":"65774","inReplyTo":"aika_Q0rWhcI6eXR@pks.im","subject":"Re: [PATCH v2] describe: limit default ref iteration to tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-11T06:41:17Z","receivedAt":"2026-06-11T06:41:19Z","isPatch":true,"body":"On Wed, Jun 10, 2026 at 10:08:51AM +0200, Patrick Steinhardt wrote:\n\n> > Given the discussion in earlier rounds and sibling topics, I assume the\n> > commit message here was AI-generated. And it's OK in the sense that it\n> > is describing what happened and I assume is entirely accurate. But as a\n> > human reader, it feels so much more verbose than what I'd expect, as it\n> > is full of semi-irrelevant details. Why set --warmup and --runs? Why\n> > bother with --command-name, which just means you have to show the\n> > commands separately anyway? Is the amount of RAM in the machine\n> > important for this test? Surely it could be if it was absurdly tiny, but\n> > in general, no, I would not expect it to be.\n> \n> I agree. Earlier this week I also drafted a message that was going down\n> this angle, but I think I didn't end up sending it to the mailing list.\n> Or at least I'm not able to find it anymore.\n> \n> To me the biggest problem is not the verbosity, even though it _is_\n> overly verbose. The bigger problem though is the incoherence of the\n> story that the commit message is trying to tell where it jumps around\n> randomly. It almost feels like rambling to me, and that makes it\n> extremely hard to follow the narrative and figure out what the message\n> even wants to tell the reader in the first place.\n\nThanks, this hits directly at the point I was trying to make (I have\ntrouble sometimes with verbosity, too!). A commit message should\nprimarily be laying out a narrative about why we are going from the old\nstate to the new, with supporting arguments. Sometimes you need\nback-story for that, sometimes not.  Sometimes you need to discuss\nalternatives, sometimes you need specific details about the platform or\nversions used for testing, and so on.\n\n> I very much think that we should and even have to expect that\n> contributors adapt, because if we don't we will basically reinforce\n> whatever AI is doing right now and increase the load on reviewers even\n> more.\n> \n> I also think that we should reserve the right to reject a patch series\n> completely in case we notice that we're basically just talking to a\n> middleman that sits between an AI prompt and us (please note that I\n> don't refer to this patch series specifically, this is more of a general\n> statement). My assumption is that this will become more important as AI\n> gets established in more workflows. The number of patch series that look\n> sane on the surface but that are utter garbage will very likely increase\n> quite significantly going forward.\n\nYep, agreed.\n\n-Peff\n"},{"id":"545243","messageId":"20260611064912.GC2191159@coredump.intra.peff.net","threadId":"65774","inReplyTo":"20260610-describe-tag-ref-scope-v3-1-5aa63ab279f7@gmail.com","subject":"Re: [PATCH v3] describe: limit default ref iteration to tags","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-06-11T06:49:12Z","receivedAt":"2026-06-11T06:49:14Z","isPatch":true,"body":"On Wed, Jun 10, 2026 at 11:50:01AM -0700, Tamir Duberstein wrote:\n\n> The runtime drops from 0.03(0.01+0.01) to 0.02(0.00+0.00). In a\n> repository with 120,532 refs but only 330 tags, the same command went\n> from 171.7 ms to 9.9 ms.\n> \n> Signed-off-by: Tamir Duberstein <tamird@gmail.com>\n> ---\n> Changes in v3:\n> - Pack the synthetic refs to better match repositories with many refs.\n> - Generate update-ref input with test_seq -f.\n> - Shorten the commit message and report the p6100.6 result.\n> - Link to v2: https://patch.msgid.link/20260608-describe-tag-ref-scope-v2-1-256fd36dca32@gmail.com\n\nThanks, this looks fine to me. I am still puzzled that your 120k ref\nrepo is so slow in the before case. Even bumping the perf test to 100k\nrefs, it ~10ms on my machine. But maybe it's a combination of not being\nwell packed, plus a slower filesystem.\n\nAt any rate, I think it is obvious that doing less work is always going\nto be better, so there's not much need to dig too deeply into the\ndifferences for information that is probably not that useful.\n\n-Peff\n"}]}