{"thread":{"id":"61346","subject":"Building with PGO: concurrency and test data","startedAt":"2024-04-21T00:52:53Z","lastAt":"2024-04-23T22:42:33Z","messageCount":3,"participants":["intelfx@intelfx.name","Mike Castle","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"493254","messageId":"65f32df3f49341bf192b606914d44cc937f7971a.camel@intelfx.name","threadId":"61346","inReplyTo":null,"subject":"Building with PGO: concurrency and test data","fromName":"","fromEmail":"intelfx@intelfx.name","sentAt":"2024-04-21T00:52:48Z","receivedAt":"2024-04-21T00:52:53Z","isPatch":false,"sender":{"key":"intelfx@intelfx.name","avatar":"https://gravatar.com/avatar/fb0d6e45051c5d8cda450fd6d02ecba0216f5f3aae1ebcf5de522939ff18eef3?d=mp&s=160"},"body":"Hi!\n\nI'm trying to build Git with PGO (for a private distribution) and I\nhave two questions about the specifics of the profiling process.\n\n\n1. The INSTALL doc says that the profiling pass has to run the test\nsuite using a single CPU, and the Makefile `profile` target also\nencodes this rule:\n\n> As a caveat: a profile-optimized build takes a *lot* longer since the\n> git tree must be built twice, and in order for the profiling\n> measurements to work properly, ccache must be disabled and the test\n> suite has to be run using only a single CPU. <...>\n( https://github.com/git/git/blob/master/INSTALL#L54-L59 )\n\n> profile:: profile-clean\n> \t$(MAKE) PROFILE=GEN all\n> \t$(MAKE) PROFILE=GEN -j1 test\n> \t@if test -n \"$$GIT_PERF_REPO\" || test -d .git; then \\\n> \t\t$(MAKE) PROFILE=GEN -j1 perf; \\\n( https://github.com/git/git/blob/master/Makefile#L2350-L2352 )\n\nHowever, some cursory searching tells me that gcc is equipped to handle\nconcurrent runs of an instrumented program:\n\n> > It is unclear to me if one can safely run multiple processes\nconcurrently.\n> > is there any risk of corruption or overwriting of the various\n\"gcda” files if different processes attempt to write on them?\n>\n> The gcda files are accessed by proper locks, so you should be sa[f]e.\n( https://gcc-help.gcc.gnu.narkive.com/0NItmccw/is-it-safe-to-generate-profiles-from-multiple-concurrent-processes#post1 )\n\nAs far as I understand, the profiling data collected does not include\ntiming information or any performance counters. What am I missing? Why\nis it not possible to run the test suite with parallelism on the\nprofiling pass?\n\n\n2. The performance test suite (t/perf/) uses up to two git repositories\n(\"normal\" and \"large\") as test data to run git commands against. Does\nthe internal organization of these repositories matter? I.e., does it\nmatter if those are \"real-world-used\" repositories with overlapping\npacks, cruft, loose objects, many refs etc., or can I simply use fresh\nclones of git.git and linux.git without loss of profile quality?\n\nThanks,\n\n-- \nIvan Shapovalov / intelfx /\n"},{"id":"493271","messageId":"CA+t9iMwX2anANcpPg15MHwdh1377zMb-k9kCt3jVx_-ggJy=sg@mail.gmail.com","threadId":"61346","inReplyTo":"65f32df3f49341bf192b606914d44cc937f7971a.camel@intelfx.name","subject":"Re: Building with PGO: concurrency and test data","fromName":"Mike Castle","fromEmail":"dalgoda@gmail.com","sentAt":"2024-04-21T15:45:38Z","receivedAt":"2024-04-21T15:45:51Z","isPatch":false,"sender":{"key":"dalgoda@gmail.com","avatar":null},"body":"On Sat, Apr 20, 2024 at 5:53 PM <intelfx@intelfx.name> wrote:\n> I'm trying to build Git with PGO (for a private distribution) and I\n> have two questions about the specifics of the profiling process.\n\nGenerally speaking, there does not need to be a lot of execution to\ngenerate good profiles.\n\nExecute the happy paths and collect data from those.  (Which implies\nthat unittests are usually a bad source of profile data.)\n\nMany folks use performance tests to generate profiles because they are\nalready written, but often, they are overkill.  Depending on what is\ngoing on in the real world, more resources are spent on collecting\ndata than would be saved by the resulting optimizations.\n\nI'd say, don't worry about it, and just go with what is already provided.\n\nFor tools like git, each run is short enough that improvements are not\nlikely to be noticed in day-to-day activities.  It is still likely to\nbe IO bound.  Most perceived performance issues are more likely to be\naddressed by algorithmic improvements (in general, not just git),\nrather than feedback profiles.\n\nNow, for any busy long running servers, this can make a bigger\ndifference, particularly for computationally expensive operations like\nauthentication.  But again, IO is likely to dominate.\n\nmrc\n"},{"id":"493379","messageId":"20240423224232.GE1172807@coredump.intra.peff.net","threadId":"61346","inReplyTo":"65f32df3f49341bf192b606914d44cc937f7971a.camel@intelfx.name","subject":"Re: Building with PGO: concurrency and test data","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2024-04-23T22:42:32Z","receivedAt":"2024-04-23T22:42:33Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Apr 21, 2024 at 02:52:48AM +0200, intelfx@intelfx.name wrote:\n\n> 1. The INSTALL doc says that the profiling pass has to run the test\n> suite using a single CPU, and the Makefile `profile` target also\n> encodes this rule:\n> \n> > As a caveat: a profile-optimized build takes a *lot* longer since the\n> > git tree must be built twice, and in order for the profiling\n> > measurements to work properly, ccache must be disabled and the test\n> > suite has to be run using only a single CPU. <...>\n> ( https://github.com/git/git/blob/master/INSTALL#L54-L59 )\n> [...]\n> However, some cursory searching tells me that gcc is equipped to handle\n> concurrent runs of an instrumented program:\n\nThat text was added quite a while ago, in f2d713fc3e (Fix build problems\nrelated to profile-directed optimization, 2012-02-06). It may be that it\nwas a problem back then, but isn't anymore.\n\n+cc the author of that commit; I don't know offhand how many people\nuse \"make profile\" (now or back then).\n\n> 2. The performance test suite (t/perf/) uses up to two git repositories\n> (\"normal\" and \"large\") as test data to run git commands against. Does\n> the internal organization of these repositories matter? I.e., does it\n> matter if those are \"real-world-used\" repositories with overlapping\n> packs, cruft, loose objects, many refs etc., or can I simply use fresh\n> clones of git.git and linux.git without loss of profile quality?\n\nI'd be surprised if the choice of repository didn't have some impact.\nAfter all, if there are no loose objects, then the routines that\ninteract with them are not going to get a lot of exercise. But how much\ndoes it actually matter in practice? I think you'd have to do a bunch of\ntrial and error measurements to find out.\n\nMy gut is that \"larger is better\" to emphasize the hot loops, but even\nthat might not be true. The main reason we want \"large\" repos in some\nperf scripts is that it makes it easier to measure the thing we are\nspeeding up versus the overhead of starting processes, etc. But PGO\nmight not be as sensitive to that, if it can get what it needs from a\nsmaller number of runs of the sensitive spots.\n\nAll of which is to say \"no idea\". I know that's not very satisfying, but\nI don't recall anybody really discussing PGO much here in the last\ndecade, so I think you're largely on your own.\n\n-Peff\n"}]}