git/list[1] front-page[2] threads[3] people[4] search[5] about
 

test suite speedups via some not-so-crazy ideas (was: scripting speedups[...])

From
Ævar Arnfjörð Bjarmason <avarab@gmail.com>
Date
Nov 3, 2021, 09:24 UTC
Message-ID
<211103.864k8t1sma.gmgdl@evledraar.gmail.com>
In-Reply-To
<211030.86ee8246hy.gmgdl@evledraar.gmail.com>
On Sat, Oct 30 2021, Ævar Arnfjörð Bjarmason wrote:
Show 56 quoted lines
> On Tue, Oct 26 2021, Eric Wong wrote:
>
>> Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:
>>> * Test suite is slow. Shell scripts and process forking.
>>> 
>>>    * What if we had a special shell that interpreted the commands in a
>>>      single process?
>>> 
>>>    * Even Git commands like rev-parse and hash-object, as long as that’s
>>>      not the command you’re trying to test
>>
>> This is something I've wanted in a very long time as a scripter.
>> fast-import has been great over the years, as is
>> "cat-file --batch(-check)", but there's gaps should be filled
>> (preferably without fragile linkage of shared libraries into a
>> script process)
>>
>>>    * Dscho wants to slip in a C-based solution
>>> 
>>>    * Jonathan tan commented: going back to your custom shell for tests
>>>      idea, one thing we could do is have a custom command that generates
>>>      the repo commits that we want (and that saves process spawns and
>>>      might make the tests simpler too)
>>
>> Perhaps a not-seriously-proposed patch from 2006 could be
>> modernized for our now-libified internals:
>
> I think something very short of a "C-based solution" could give us most
> of the wins here. Johannes was probably thinking of the scripting being
> slow on Windows aspect of it.
>
> But the main benefit of hypothetical C-based testing is that you can
> connect it to the dependency tree we have in the Makefile, and only
> re-run tests for code you needed to re-compile.
>
> So e.g. we don't need to run tests that invoke "git tag" if the
> dependency graph of builtin/tag.c didn't change.
>
> With COMPUTE_HEADER_DEPENDENCIES we've got access to that dependency
> information for our C code.
>
> With trace2 we could record an initial test run, and know which built-in
> commands are executed by which tests (even down to the sub-test level).
>
> Connecting these two means that we can find all tests that say run "git
> fsck", and if builtin/fsck.c is the only thing that changed in an
> interactive rebase, that's the only tests we need to run.
>
> Of course changes to things like cache.h or t/test-lib.sh would spoil
> that cache entirely, but pretty much the same is true for re-compiling
> things now, so would changing say builtin/init-db.c, as almost every
> test does a "git init" somewhere.
>
> But I think that approch is viable, and should take us from a huge
> hypothetical project like "rewrite all the tests in C" to something
> that's a viable weekend hacking project for someone who's interested.

First to outline some goals: I think saying we'd like to speed up scripts is really getting into the weeds.

Surely we'd like to speed up test runs, and generally speaking our test suite can be parallelized, and it mostly doesn't matter if it runs on your computer or other people's computers, as long as it runs your code. So:

 1. Even for contributors that have a slow system they could benefit from
    the hosted CI (on GitHub or wherever else) being faster.
 2. Our CI takes around 30-60m to finish.
 3. That CI time is almost entirely something that could be sped up by
    throwing hardware at it.
 4. We're currently using "Dv2 and DSv2-series" hosted runners
    (https://docs.github.com/en/actions/using-github-hosted-runners/about-github-hosted-runners)
    we have quite a few people on-list who work for the
    company/companies involved.
    Is it within the realm of possibility to get more CI resources
    assigned to git/git's organization network?
 5. Or, is there willingness to host/pay for hosted runners from
    someone?
    Not wearing PLC hat I'd think that we could speed that up a lot with
    some reasonable money spending, and if pushing to CI made CI run in
    3-5m instead of 60m that would be worthwhile.
 6. Related to #5: I've been able to setup hosted runner jobs, and
    self-hosted runner jobs, but is there a way to do some opportunistic
    mixture of the two? Even one where self-hosted runners could come
    and go, and if they're present contribute resources to git/git's
    network?
 7. We run the various GIT_TEST_* etc. jobs in sequence, is there a
    reason for why we're serializing things in GitHub CI that could be
    parallelized?
    The vs-build and vs-test tests run in parallel, any reason we're not
    doing that trick on the ubuntu runners other than "nobody got to
    it?". We seem to be trying hard to do the exact opposite there..
    At the extreme end we could build git ~once, and have N tests depend
    on that, where N ~= $(ls t/*.sh) x $number_of_test_modes). But
    perhaps runner starting overhead starts to be the limiting factor at
    some point.
 8. To a first approximation, does anyone really care about getting an
    exhaustive list of all failures in a run, or just that we have *a*
    failure? You can always do an exhaustive run later.
 9. On the "no" answer to #8: When I build/test my own git I first run
    those tests that I modified in the relevant branches, and if any of
    those fail I just stop.
    I generally don't need to run the entirety of the rest of the test
    suite to stop and investigate why I have a failure.
    Perhaps our CI could use a similar trick, i.e. first test the set of
    modified test files, and perhaps with some ad-hoc matching of
    filenames, so e.g. if you modify builtin/add.c we'd run t/*add*.sh
    in the first set, and all with --immediate per #8 above.
    If we pass that we'd run the full set, minus that initial set.
Previous: Ævar Arnfjörð BjarmasonNext: Junio C Hamano
Message 6 of 58 in “Notes from the Git Contributors' Summit 2021, virtual, Oct 19/20”
  1. Johannes SchindelinOct 21, 2021
  2. [Summit topic] Crazy (and not so crazy) ideasJohannes Schindelin, Oct 21, 2021
  3. Son Luong NgocOct 21, 2021
  4. scripting speedups [was: [Summit topic] Crazy (and not so crazy) ideas]Eric Wong, Oct 26, 2021
  5. Ævar Arnfjörð BjarmasonOct 30, 2021
  6. test suite speedups via some not-so-crazy ideas (was: scripting speedups[...])Ævar Arnfjörð Bjarmason, Nov 3, 2021
  7. Junio C HamanoNov 3, 2021
  8. Johannes SchindelinNov 2, 2021
  9. [Summit topic] SHA-256 UpdatesJohannes Schindelin, Oct 21, 2021
  10. [Summit topic] Server-side merge/rebase: needs and wants?Johannes Schindelin, Oct 21, 2021
  11. Bagas SanjayaOct 22, 2021
  12. Johannes SchindelinOct 22, 2021
  13. Ævar Arnfjörð BjarmasonOct 23, 2021
  14. Taylor BlauNov 8, 2021
  15. Ævar Arnfjörð BjarmasonNov 9, 2021
  16. Christian CouderNov 30, 2021
  17. [Summit topic] Submodules and how to make them worth usingJohannes Schindelin, Oct 21, 2021
  18. [Summit topic] Sparse checkout behavior and plansJohannes Schindelin, Oct 21, 2021
  19. [Summit topic] The state of getting a reftable backend working in git.gitJohannes Schindelin, Oct 21, 2021
  20. Han-Wen NienhuysOct 25, 2021
  21. Ævar Arnfjörð BjarmasonOct 25, 2021
  22. Han-Wen NienhuysOct 26, 2021
  23. Philip OakleyOct 28, 2021
  24. Philip OakleyOct 26, 2021
  25. [Summit topic] Documentation (translations, FAQ updates, new user-focused, general improvements, etc.)Johannes Schindelin, Oct 21, 2021
  26. Jean-Noël AvilaOct 22, 2021
  27. Ævar Arnfjörð BjarmasonOct 22, 2021
  28. Jean-Noël AvilaOct 27, 2021
  29. Jeff KingOct 27, 2021
  30. [Summit topic] Increasing diversity & inclusion (transition to `main`, etc)Johannes Schindelin, Oct 21, 2021
  31. Son Luong NgocOct 21, 2021
  32. vale check, was Re: [Summit topic] Increasing diversity & inclusion (transition to `main`, etc)Johannes Schindelin, Oct 22, 2021
  33. Johannes SchindelinOct 22, 2021
  34. [Summit topic] Improving Git UXJohannes Schindelin, Oct 21, 2021
  35. changing the experimental 'git switch' (was: [Summit topic] Improving Git UX)Ævar Arnfjörð Bjarmason, Oct 21, 2021
  36. Junio C HamanoOct 21, 2021
  37. Bagas SanjayaOct 22, 2021
  38. martinOct 22, 2021
  39. Ævar Arnfjörð BjarmasonOct 22, 2021
  40. Sergey OrganovOct 22, 2021
  41. martinOct 22, 2021
  42. Sergey OrganovOct 23, 2021
  43. MartinOct 24, 2021
  44. Junio C HamanoOct 24, 2021
  45. Ævar Arnfjörð BjarmasonOct 25, 2021
  46. Junio C HamanoOct 25, 2021
  47. Sergey OrganovOct 25, 2021
  48. Ævar Arnfjörð BjarmasonOct 25, 2021
  49. Sergey OrganovOct 27, 2021
  50. [Summit topic] Improving reviewer quality of life (patchwork, subsystem lists?, etc)Johannes Schindelin, Oct 21, 2021
  51. Konstantin RyabitsevOct 21, 2021
  52. Ævar Arnfjörð BjarmasonOct 22, 2021
  53. Missing notes, was Re: Notes from the Git Contributors' Summit 2021, virtual, Oct 19/20Johannes Schindelin, Oct 22, 2021
  54. Johannes SchindelinOct 22, 2021
  55. Johannes SchindelinOct 22, 2021
  56. Johannes SchindelinOct 22, 2021
  57. Let's have public Git chalk talks, was Re: Notes from the Git Contributors' Summit 2021, virtual, Oct 19/20Johannes Schindelin, Oct 22, 2021
  58. Ævar Arnfjörð BjarmasonOct 25, 2021

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.