{"thread":{"id":"62894","subject":"Continuous Benchmarking","startedAt":"2025-02-03T09:55:08Z","lastAt":"2025-02-21T08:48:25Z","messageCount":4,"participants":["Patrick Steinhardt","Junio C Hamano","Emily Shaffer"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"511725","messageId":"Z6CSc_vyGkn-ozUH@pks.im","threadId":"62894","inReplyTo":null,"subject":"Continuous Benchmarking","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-02-03T09:54:59Z","receivedAt":"2025-02-03T09:55:08Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"Hi,\n\ndue to a couple performance regressions that we have hit over the last\ncouple Git releases at GitLab, we have started to set up an effort to\nimplement continuous benchmarking for the Git project. The intent is to\nhave regular (daily) benchmarking runs against Git's `master` and `next`\nbranches to be able to spot any performance regressions before they make\nit into the next release.\n\nI have started with a relatively simple setup:\n\n  - I have started collection benchmarks that I myself do regularly [1].\n    These benchmarks are built on hyperfine and are thus not part of the\n    Git repository itself.\n\n  - GitLab CI runs on a nightly basis, executing a subset of these\n    benchmarks [2].\n\n  - Results are uploaded with a hyperfine adaptor to Bencher and are\n    summarized in dashboards.\n\nThis at least gives us some visibility in severe performance outliers,\nwhether these are improvements or regressions. Some statistics are\napplied on this data to automatically generate alerts when things are\nsignificantly changing.\n\nThe setup is of course not perfect. It's built on top of CI jobs, which\nare by their very nature not really performing consistent. The scripts\nare hosted outside of Git. And I'm the only one running this.\n\nSo I wonder whether there is a wider interest in the Git community to\nhave this infrastructure part of the Git project itself. This may\ninclude steps like the following:\n\n  - Extending our performance tests we have in \"t/perf\" to cover more\n    benchmarks.\n\n  - Writing an adaptor that is able to upload the data generated from\n    our perf scripts to Bencher.\n\n  - Setting up proper infrastructure to do the benchmarking. We may for\n    now also continue to use GitLab CI, but as said they are quite noisy\n    overall. Dedicated servers would help here.\n\n  - Sending alerts to the Git mailing list.\n\nI'm happy to hear your thoughts on this. Any ideas are welcome,\nincluding \"we're not interested at all\". In that case, we'd simply\ncontinue to maintain the setup ourselves at GitLab.\n\nThanks!\n\nPatrick\n\n[1]: https://gitlab.com/gitlab-org/data-access/git/benchmarks\n[2]: https://gitlab.com/gitlab-org/data-access/git/benchmarks/-/blob/main/.gitlab-ci.yml?ref_type=heads\n[3]: https://bencher.dev/console/projects/git/plots\n"},{"id":"511745","messageId":"xmqqpljz2dk5.fsf@gitster.g","threadId":"62894","inReplyTo":"Z6CSc_vyGkn-ozUH@pks.im","subject":"Re: Continuous Benchmarking","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2025-02-03T16:33:30Z","receivedAt":"2025-02-03T16:33:33Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Patrick Steinhardt <ps@pks.im> writes:\n\n> ... implement continuous benchmarking for the Git project. The intent is to\n> have regular (daily) benchmarking runs against Git's `master` and `next`\n> branches to be able to spot any performance regressions before they make\n> it into the next release.\n\nThis is great.\n\n> I have started with a relatively simple setup:\n>\n>   - I have started collection benchmarks that I myself do regularly [1].\n>     These benchmarks are built on hyperfine and are thus not part of the\n>     Git repository itself.\n>\n>   - GitLab CI runs on a nightly basis, executing a subset of these\n>     benchmarks [2].\n>\n>   - Results are uploaded with a hyperfine adaptor to Bencher and are\n>     summarized in dashboards.\n>\n> This at least gives us some visibility in severe performance outliers,\n> whether these are improvements or regressions. Some statistics are\n> applied on this data to automatically generate alerts when things are\n> significantly changing.\n>\n> The setup is of course not perfect. It's built on top of CI jobs, which\n> are by their very nature not really performing consistent. The scripts\n> are hosted outside of Git. And I'm the only one running this.\n>\n> So I wonder whether there is a wider interest in the Git community to\n> have this infrastructure part of the Git project itself. This may\n> include steps like the following:\n>\n>   - Extending our performance tests we have in \"t/perf\" to cover more\n>     benchmarks.\n>\n>   - Writing an adaptor that is able to upload the data generated from\n>     our perf scripts to Bencher.\n>\n>   - Setting up proper infrastructure to do the benchmarking. We may for\n>     now also continue to use GitLab CI, but as said they are quite noisy\n>     overall. Dedicated servers would help here.\n>\n>   - Sending alerts to the Git mailing list.\n>\n> I'm happy to hear your thoughts on this. Any ideas are welcome,\n> including \"we're not interested at all\". In that case, we'd simply\n> continue to maintain the setup ourselves at GitLab.\n\nElsewhere Peff was talking about his adventure with Coverty running\non 'next'.  The more eyes and tools on the topics before they hit\n'master', the less chance we have to scramble just before the\nrelease.\n\n\n"},{"id":"511913","messageId":"CAJoAoZmJAM--FVmhxs_0sL1A8yrLwNBFULPDYFgV=AtFhn67+g@mail.gmail.com","threadId":"62894","inReplyTo":"Z6CSc_vyGkn-ozUH@pks.im","subject":"Re: Continuous Benchmarking","fromName":"Emily Shaffer","fromEmail":"nasamuffin@google.com","sentAt":"2025-02-05T23:14:21Z","receivedAt":"2025-02-05T23:14:34Z","isPatch":false,"sender":{"key":"nasamuffin@google.com","avatar":"https://avatars.githubusercontent.com/u/1606826?v=4"},"body":"On Mon, Feb 3, 2025 at 1:55 AM Patrick Steinhardt <ps@pks.im> wrote:\n>\n> Hi,\n>\n> due to a couple performance regressions that we have hit over the last\n> couple Git releases at GitLab, we have started to set up an effort to\n> implement continuous benchmarking for the Git project. The intent is to\n> have regular (daily) benchmarking runs against Git's `master` and `next`\n> branches to be able to spot any performance regressions before they make\n> it into the next release.\n>\n> I have started with a relatively simple setup:\n>\n>   - I have started collection benchmarks that I myself do regularly [1].\n>     These benchmarks are built on hyperfine and are thus not part of the\n>     Git repository itself.\n>\n>   - GitLab CI runs on a nightly basis, executing a subset of these\n>     benchmarks [2].\n>\n>   - Results are uploaded with a hyperfine adaptor to Bencher and are\n>     summarized in dashboards.\n>\n> This at least gives us some visibility in severe performance outliers,\n> whether these are improvements or regressions. Some statistics are\n> applied on this data to automatically generate alerts when things are\n> significantly changing.\n>\n> The setup is of course not perfect. It's built on top of CI jobs, which\n> are by their very nature not really performing consistent. The scripts\n> are hosted outside of Git. And I'm the only one running this.\n\nFor the CI \"noisy neighbors\" problem at least, it could be an option\nto try to host in GCE (or some other compute that isn't shared). I\nasked around a little inside Google and it seems like it's possible,\nI'll keep pushing on it and see just how hard it would be. I'd even be\nhappy to trade on-push runs with noisy neighbors for nightly runs with\nno neighbors, which makes it not really a CI thing - guess I will find\nout if that's easier or harder for us to implement. :)\n\n>\n> So I wonder whether there is a wider interest in the Git community to\n> have this infrastructure part of the Git project itself. This may\n> include steps like the following:\n>\n>   - Extending our performance tests we have in \"t/perf\" to cover more\n>     benchmarks.\n\nFolks may be aware that our biggest (in terms of scale) internal\ncustomer at Google is Android project. They are the ones who complain\nto me and my team the most about performance; they are also open to\nsetting up nightly performance regression test. Would it be appealing\nto get reports from such a test upstream? I think it's more compelling\nto our customer team if we run it against the closed-source Android\nrepo, which means the Git project doesn't get to see as much about the\nshape and content of the repos the performance tests are running\nagainst, but we might be able to publish info about the shape without\nthe contents. Would that be useful? What would help to know (# of\ncommits, size of largest object, distribution of object size, # of\nbranches, size of worktree...?) If not having the specifics of the\nrepo-under-test is a dealbreaker we could explore running performance\ntests in public with Android Open Source Project as the\nrepo-under-test instead, but it's much more manageable than full\nAndroid.\n\nMaybe in the long term it would be even better to have some toy\nrepo-under-test, like \"sample repo with massive object store\", \"sample\nrepo with massive history\", etc. to help us pinpoint which ways we're\nscaling well and which ways we aren't. But having a ready made\nrepo-under-test, and a team who's got a very large stake in Git\nperforming well with it (so they can invest their time in setting up\ntests), might be a good enough place to start.\n\n>\n>   - Writing an adaptor that is able to upload the data generated from\n>     our perf scripts to Bencher.\n>\n>   - Setting up proper infrastructure to do the benchmarking. We may for\n>     now also continue to use GitLab CI, but as said they are quite noisy\n>     overall. Dedicated servers would help here.\n>\n>   - Sending alerts to the Git mailing list.\n\nYeah, I'd love to see reports coming to Git mailing list, or at least\nbad news reports (maybe we don't need \"everything ran great!\" every\nnight, but would appreciate \"last night the performance suite ran 50%\nslower than last-6-months average\"). That seems the easiest to\nintegrate with the way the project runs now, and I think we are used\nto list noise :)\n\n>\n> I'm happy to hear your thoughts on this. Any ideas are welcome,\n> including \"we're not interested at all\". In that case, we'd simply\n> continue to maintain the setup ourselves at GitLab.\n\nIn general, though, yes! I am very interested! Google had trouble with\nperformance regressions over the last 3 months or so, I'd love to see\nthe community noticing it more. I think in general we have a sense\nthat performance matters, during code review, but aren't always sure\nwhere it matters most, and a regular performance test that anybody can\nsee the results of would help a lot.\n\n>\n> Thanks!\n>\n> Patrick\n>\n> [1]: https://gitlab.com/gitlab-org/data-access/git/benchmarks\n> [2]: https://gitlab.com/gitlab-org/data-access/git/benchmarks/-/blob/main/.gitlab-ci.yml?ref_type=heads\n> [3]: https://bencher.dev/console/projects/git/plots\n"},{"id":"512812","messageId":"Z7g90CMEiy-skRKK@pks.im","threadId":"62894","inReplyTo":"CAJoAoZmJAM--FVmhxs_0sL1A8yrLwNBFULPDYFgV=AtFhn67+g@mail.gmail.com","subject":"Re: Continuous Benchmarking","fromName":"Patrick Steinhardt","fromEmail":"ps@pks.im","sentAt":"2025-02-21T08:48:16Z","receivedAt":"2025-02-21T08:48:25Z","isPatch":false,"sender":{"key":"ps@pks.im","avatar":"https://avatars.githubusercontent.com/u/4056630?v=4"},"body":"On Wed, Feb 05, 2025 at 03:14:21PM -0800, Emily Shaffer wrote:\n> On Mon, Feb 3, 2025 at 1:55 AM Patrick Steinhardt <ps@pks.im> wrote:\n> >\n> > Hi,\n> >\n> > due to a couple performance regressions that we have hit over the last\n> > couple Git releases at GitLab, we have started to set up an effort to\n> > implement continuous benchmarking for the Git project. The intent is to\n> > have regular (daily) benchmarking runs against Git's `master` and `next`\n> > branches to be able to spot any performance regressions before they make\n> > it into the next release.\n> >\n> > I have started with a relatively simple setup:\n> >\n> >   - I have started collection benchmarks that I myself do regularly [1].\n> >     These benchmarks are built on hyperfine and are thus not part of the\n> >     Git repository itself.\n> >\n> >   - GitLab CI runs on a nightly basis, executing a subset of these\n> >     benchmarks [2].\n> >\n> >   - Results are uploaded with a hyperfine adaptor to Bencher and are\n> >     summarized in dashboards.\n> >\n> > This at least gives us some visibility in severe performance outliers,\n> > whether these are improvements or regressions. Some statistics are\n> > applied on this data to automatically generate alerts when things are\n> > significantly changing.\n> >\n> > The setup is of course not perfect. It's built on top of CI jobs, which\n> > are by their very nature not really performing consistent. The scripts\n> > are hosted outside of Git. And I'm the only one running this.\n> \n> For the CI \"noisy neighbors\" problem at least, it could be an option\n> to try to host in GCE (or some other compute that isn't shared). I\n> asked around a little inside Google and it seems like it's possible,\n> I'll keep pushing on it and see just how hard it would be. I'd even be\n> happy to trade on-push runs with noisy neighbors for nightly runs with\n> no neighbors, which makes it not really a CI thing - guess I will find\n> out if that's easier or harder for us to implement. :)\n\nThat would be awesome.\n\n> > So I wonder whether there is a wider interest in the Git community to\n> > have this infrastructure part of the Git project itself. This may\n> > include steps like the following:\n> >\n> >   - Extending our performance tests we have in \"t/perf\" to cover more\n> >     benchmarks.\n> \n> Folks may be aware that our biggest (in terms of scale) internal\n> customer at Google is Android project. They are the ones who complain\n> to me and my team the most about performance; they are also open to\n> setting up nightly performance regression test. Would it be appealing\n> to get reports from such a test upstream? I think it's more compelling\n> to our customer team if we run it against the closed-source Android\n> repo, which means the Git project doesn't get to see as much about the\n> shape and content of the repos the performance tests are running\n> against, but we might be able to publish info about the shape without\n> the contents. Would that be useful? What would help to know (# of\n> commits, size of largest object, distribution of object size, # of\n> branches, size of worktree...?) If not having the specifics of the\n> repo-under-test is a dealbreaker we could explore running performance\n> tests in public with Android Open Source Project as the\n> repo-under-test instead, but it's much more manageable than full\n> Android.\n\nThe biggest question is whether such regression reports would be\nactionable by the Git community. I often found performance issues to be\nvery specific to the repository at hand, and reconstructing the exact\nsituation tends to be extremely tedious or completely infeasible. I run\ninto the situation way too often where customers come knock at my door\nwith a performance issue, but don't want to provide the underlying data.\nMore often than not I end up not being able to reproduce, so I have to\npush back on such reports.\n\nIdeally, any report should be accompanied by a trivial reproducer that\nany developer can execute on their local machine.\n\n> Maybe in the long term it would be even better to have some toy\n> repo-under-test, like \"sample repo with massive object store\", \"sample\n> repo with massive history\", etc. to help us pinpoint which ways we're\n> scaling well and which ways we aren't. But having a ready made\n> repo-under-test, and a team who's got a very large stake in Git\n> performing well with it (so they can invest their time in setting up\n> tests), might be a good enough place to start.\n\nThat would be great. I guess this wouldn't be a single repository, but a\nset of repositories that have different kinds of characteristics.\n\n> >   - Writing an adaptor that is able to upload the data generated from\n> >     our perf scripts to Bencher.\n> >\n> >   - Setting up proper infrastructure to do the benchmarking. We may for\n> >     now also continue to use GitLab CI, but as said they are quite noisy\n> >     overall. Dedicated servers would help here.\n> >\n> >   - Sending alerts to the Git mailing list.\n> \n> Yeah, I'd love to see reports coming to Git mailing list, or at least\n> bad news reports (maybe we don't need \"everything ran great!\" every\n> night, but would appreciate \"last night the performance suite ran 50%\n> slower than last-6-months average\"). That seems the easiest to\n> integrate with the way the project runs now, and I think we are used\n> to list noise :)\n\nOh, totally, I certainly don't think there's any benefit in reporting\nanything when there is no information. Right now there still are semi-\nfrequent outliers where an alert is generated only because of a flake,\nnot a real performance regression. But my hope would be that we can\naddress this issue once we address the noisy neighbour problem.\n\n> > I'm happy to hear your thoughts on this. Any ideas are welcome,\n> > including \"we're not interested at all\". In that case, we'd simply\n> > continue to maintain the setup ourselves at GitLab.\n> \n> In general, though, yes! I am very interested! Google had trouble with\n> performance regressions over the last 3 months or so, I'd love to see\n> the community noticing it more. I think in general we have a sense\n> that performance matters, during code review, but aren't always sure\n> where it matters most, and a regular performance test that anybody can\n> see the results of would help a lot.\n\nThanks for your input!\n\nPatrick\n"}]}