git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 06/18] chainlint.pl: validate test scripts in parallel

From
Jeff King <peff@peff.net>
Date
Sep 6, 2022, 23:26 UTC
Message-ID
<YxfXQ0IJjq/FT2Uh@coredump.intra.peff.net>
In-Reply-To
<CAPig+cSx661-HEr3JcAD5MuYfgHviGQ1cSAftkgw6gj2FgTQVg@mail.gmail.com>
On Tue, Sep 06, 2022 at 06:52:26PM -0400, Eric Sunshine wrote:
Show 12 quoted lines
> On Tue, Sep 6, 2022 at 6:35 PM Eric Wong <e@80x24.org> wrote:
> > Eric Sunshine via GitGitGadget <gitgitgadget@gmail.com> wrote:
> > > +unless ($Config{useithreads} && eval {
> > > +     require threads; threads->import();
> >
> > Fwiw, the threads(3perl) manpage has this since 2014:
> >
> >        The use of interpreter-based threads in perl is officially discouraged.
> 
> Thanks for pointing this out. I did see that, but as no better
> alternative was offered, and since I did want this to work on Windows,
> I went with it.

I did some timings the other night, and I found something quite curious with the thread stuff.

Here's a hyperfine run of "make" in the t/ directory before any of your patches. It uses "prove" to do parallelism under the hood:

  Benchmark 1: make
    Time (mean ± σ):     68.895 s ±  0.840 s    [User: 620.914 s, System: 428.498 s]
    Range (min … max):   67.943 s … 69.531 s    3 runs

So that gives us a baseline. Now the first thing I wondered is how bad it would be to just run chainlint.pl once per script. So I applied up to that patch:

  Benchmark 1: make
    Time (mean ± σ):     71.289 s ±  1.302 s    [User: 673.300 s, System: 417.912 s]
    Range (min … max):   69.788 s … 72.120 s    3 runs

I was quite surprised that it made things slower! It's nice that we're only calling it once per script instead of once per test, but it seems the startup overhead of the script is really high.

And since in this mode we're only feeding it one script at a time, I tried reverting the "chainlint.pl: validate test scripts in parallel" commit. And indeed, now things are much faster:

  Benchmark 1: make
    Time (mean ± σ):     61.544 s ±  3.364 s    [User: 556.486 s, System: 384.001 s]
    Range (min … max):   57.660 s … 63.490 s    3 runs
And you can see the same thing just running chainlint by itself:
  $ time perl chainlint.pl /dev/null
  real	0m0.069s
  user	0m0.042s
  sys	0m0.020s
  $ git revert HEAD^{/validate.test.scripts.in.parallel}
  $ time perl chainlint.pl /dev/null
  real	0m0.014s
  user	0m0.010s
  sys	0m0.004s

I didn't track down the source of the slowness. Maybe it's loading extra modules, or maybe it's opening /proc/cpuinfo, or maybe it's the thread setup. But it's a surprising slowdown.

Now of course your intent is to do a single repo-wide invocation. And that is indeed a bit faster. Here it is without the parallel code:

  Benchmark 1: make
    Time (mean ± σ):     61.727 s ±  2.140 s    [User: 507.712 s, System: 377.753 s]
    Range (min … max):   59.259 s … 63.074 s    3 runs

The wall-clock time didn't improve much, but the CPU time did. Restoring the parallel code does improve the wall-clock time a bit, but at the cost of some extra CPU:

  Benchmark 1: make
    Time (mean ± σ):     59.029 s ±  2.851 s    [User: 515.690 s, System: 380.369 s]
    Range (min … max):   55.736 s … 60.693 s    3 runs

which makes sense. If I do a with/without of just "make test-chainlint", the parallelism is buying a few seconds of wall-clock:

  Benchmark 1: make test-chainlint
    Time (mean ± σ):     900.1 ms ± 102.9 ms    [User: 12049.8 ms, System: 79.7 ms]
    Range (min … max):   704.2 ms … 994.4 ms    10 runs
  Benchmark 1: make test-chainlint
    Time (mean ± σ):      3.778 s ±  0.042 s    [User: 3.756 s, System: 0.023 s]
    Range (min … max):    3.706 s …  3.833 s    10 runs

I'm not sure what it all means. For Linux, I think I'd be just as happy with a single non-parallelized test-chainlint run for each file. But maybe on Windows the startup overhead is worse? OTOH, the whole test run is so much worse there. One process per script is not going to be that much in relative terms either way.

And if we did cache the results and avoid extra invocations via "make", then we'd want all the parallelism to move to there anyway.

Maybe that gives you more food for thought about whether perl's "use threads" is worth having.

-Peff
Previous: Eric SunshineNext: Eric Sunshine
Message 15 of 51 in “make test "linting" more comprehensive”
  1. 00/18 make test "linting" more comprehensiveEric Sunshine via GitGitGadget, Sep 1, 2022
  2. 01/18 t: add skeleton chainlint.plEric Sunshine via GitGitGadget, Sep 1, 2022
  3. Ævar Arnfjörð BjarmasonSep 1, 2022
  4. Eric SunshineSep 2, 2022
  5. 02/18 chainlint.pl: add POSIX shell lexical analyzerEric Sunshine via GitGitGadget, Sep 1, 2022
  6. Ævar Arnfjörð BjarmasonSep 1, 2022
  7. Eric SunshineSep 3, 2022
  8. 04/18 chainlint.pl: add parser to validate testsEric Sunshine via GitGitGadget, Sep 1, 2022
  9. 03/18 chainlint.pl: add POSIX shell parserEric Sunshine via GitGitGadget, Sep 1, 2022
  10. 06/18 chainlint.pl: validate test scripts in parallelEric Sunshine via GitGitGadget, Sep 1, 2022
  11. Ævar Arnfjörð BjarmasonSep 1, 2022
  12. Eric SunshineSep 3, 2022
  13. Eric WongSep 6, 2022
  14. Eric SunshineSep 6, 2022
  15. Jeff KingSep 6, 2022
  16. Eric SunshineNov 21, 2022
  17. Ævar Arnfjörð BjarmasonNov 21, 2022
  18. Eric SunshineNov 21, 2022
  19. Ævar Arnfjörð BjarmasonNov 21, 2022
  20. Eric SunshineNov 21, 2022
  21. Jeff KingNov 21, 2022
  22. Eric SunshineNov 21, 2022
  23. Eric SunshineNov 21, 2022
  24. Jeff KingNov 21, 2022
  25. Eric SunshineNov 21, 2022
  26. Jeff KingNov 21, 2022
  27. Ævar Arnfjörð BjarmasonNov 22, 2022
  28. 05/18 chainlint.pl: add parser to identify test definitionsEric Sunshine via GitGitGadget, Sep 1, 2022
  29. 07/18 chainlint.pl: don't require `return|exit|continue` to end with `&&`Eric Sunshine via GitGitGadget, Sep 1, 2022
  30. 10/18 chainlint.pl: don't flag broken &&-chain if `$?` handled explicitlyEric Sunshine via GitGitGadget, Sep 1, 2022
  31. 12/18 chainlint.pl: complain about loops lacking explicit failure handlingEric Sunshine via GitGitGadget, Sep 1, 2022
  32. 09/18 chainlint.pl: don't require `&` background command to end with `&&`Eric Sunshine via GitGitGadget, Sep 1, 2022
  33. 08/18 t/Makefile: apply chainlint.pl to existing self-testsEric Sunshine via GitGitGadget, Sep 1, 2022
  34. 11/18 chainlint.pl: don't flag broken &&-chain if failure indicated explicitlyEric Sunshine via GitGitGadget, Sep 1, 2022
  35. 13/18 chainlint.pl: allow `|| echo` to signal failure upstream of a pipeEric Sunshine via GitGitGadget, Sep 1, 2022
  36. 14/18 t/chainlint: add more chainlint.pl self-testsEric Sunshine via GitGitGadget, Sep 1, 2022
  37. 15/18 test-lib: retire "lint harder" optimization hackEric Sunshine via GitGitGadget, Sep 1, 2022
  38. 16/18 test-lib: replace chainlint.sed with chainlint.plEric Sunshine via GitGitGadget, Sep 1, 2022
  39. Elijah NewrenSep 3, 2022
  40. Eric SunshineSep 3, 2022
  41. 18/18 t: retire unused chainlint.sedEric Sunshine via GitGitGadget, Sep 1, 2022
  42. Johannes SchindelinSep 2, 2022
  43. Eric SunshineSep 2, 2022
  44. Jeff KingSep 2, 2022
  45. Junio C HamanoSep 2, 2022
  46. 17/18 t/Makefile: teach `make test` and `make prove` to run chainlint.plEric Sunshine via GitGitGadget, Sep 1, 2022
  47. Jeff KingSep 11, 2022
  48. Eric SunshineSep 11, 2022
  49. Jeff KingSep 11, 2022
  50. Eric SunshineSep 12, 2022
  51. Jeff KingSep 13, 2022

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.