git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [ANNOUNCE] Git v2.19.0-rc0

From
Ævar Arnfjörð Bjarmason <avarab@gmail.com>
Date
Aug 24, 2018, 07:57 UTC
Message-ID
<87zhxcdwnb.fsf@evledraar.gmail.com>
In-Reply-To
<20180824065625.GA10556@sigill.intra.peff.net>
On Fri, Aug 24 2018, Jeff King wrote:
Show 47 quoted lines
> On Thu, Aug 23, 2018 at 04:59:27PM -0400, Derrick Stolee wrote:
>
>> Using git/git:
>>
>> Test v2.18.0 v2.19.0-rc0 HEAD
>> -------------------------------------------------------------------------
>> 0001.2: 3.10(3.02+0.08) 3.27(3.17+0.09) +5.5% 3.14(3.02+0.11) +1.3%
>>
>>
>> Using torvalds/linux:
>>
>> Test v2.18.0 v2.19.0-rc0 HEAD
>> ------------------------------------------------------------------------------
>> 0001.2: 56.08(45.91+1.50) 56.60(46.62+1.50) +0.9% 54.61(45.47+1.46) -2.6%
>
> Interesting that these timings aren't as dramatic as the ones you got
> the other day (mine seemed to shift, too; for whatever reason it seems
> like under load the difference is larger).
>
>> Now here is where I get on my soapbox (and create a TODO for myself later).
>> I ran the above with GIT_PERF_REPEAT_COUNT=10, which intuitively suggests
>> that the results should be _more_ accurate than the default of 3. However, I
>> then remember that we only report the *minimum* time from all the runs,
>> which is likely to select an outlier from the distribution. To test this, I
>> ran a few tests manually and found the variation between runs to be larger
>> than 3%.
>
> Yes, I agree it's not a great system. The whole "best of 3" thing is
> OK for throwing out cold-cache warmups, but it's really bad for teasing
> out the significance of small changes, or even understanding how much
> run-to-run noise there is.
>
>> When I choose my own metrics for performance tests, I like to run at least
>> 10 runs, remove the largest AND smallest runs from the samples, and then
>> take the average. I did this manually for 'git rev-list --all --objects' on
>> git/git and got the following results:
>
> I agree that technique is better. I wonder if there's something even
> more statistically rigorous we could do. E.g., to compute the variance
> and throw away outliers based on standard deviations. And also to report
> the variance to give a sense of the significance of any changes.
>
> Obviously more runs gives greater confidence in the results, but 10
> sounds like a lot. Many of these tests take minutes to run. Letting it
> go overnight is OK if you're doing a once-per-release mega-run, but it's
> pretty painful if you just want to generate some numbers to show off
> your commit.

An ex-coworker who's a *lot* smarter about these things than I am wrote this module: https://metacpan.org/pod/Dumbbench

I while ago I dabbled briefly with integrating it into t/perf/ but got distracted by something else:

    The module currently works similar to [more traditional iterative
    benchmark modules], except (in layman terms) it will run the command
    many times, estimate the uncertainty of the result and keep
    iterating until a certain user-defined precision has been
    reached. Then, it calculates the resulting uncertainty and goes
    through some pain to discard bad runs and subtract overhead from the
    timings. The reported timing includes an uncertainty, so that
    multiple benchmarks can more easily be compared.

Details of how it works here: https://metacpan.org/pod/Dumbbench#HOW-IT-WORKS-AND-WHY-IT-DOESN'T

Something like that seems to me to be an inherently better approach. I.e. we have lots of test cases that take 500ms, and some that take maybe 5 minutes (depending on the size of the repository).

Indiscriminately repeating all of those for GIT_PERF_REPEAT_COUNT must be dumber than something like the above method.

We could also speed up the runtime of the perf tests a lot with such a method, by e.g. saying that we're OK with less certainty on tests that take a longer time than those that take a shorter time.

Show 20 quoted lines
>> v2.18.0 v2.19.0-rc0 HEAD
>> --------------------------------
>> 3.126 s 3.308 s 3.170 s
>
> So that's 5% worsening in 2.19, and we reclaim all but 1.4% of it. Those
> numbers match what I expect (and what I was seeing in some of my earlier
> timings).
>
>> I just kicked off a script that will run this test on the Linux repo while I
>> drive home. I'll be able to report a similar table of data easily.
>
> Thanks, I'd expect it to come up with similar percentages. So we'll see
> if that holds true. :)
>
>> My TODO is to consider aggregating the data this way (or with a median)
>> instead of reporting the minimum.
>
> Yes, I think that would be a great improvement for t/perf.
>
> -Peff
Previous: Jeff KingNext: Derrick Stolee
Message 55 of 58 in “[ANNOUNCE] Git v2.19.0-rc0”
  1. Junio C HamanoAug 20, 2018
  2. Stefan BellerAug 20, 2018
  3. Jonathan NiederAug 20, 2018
  4. Jonathan NiederAug 21, 2018
  5. Stefan BellerAug 21, 2018
  6. Derrick StoleeAug 21, 2018
  7. Jeff KingAug 21, 2018
  8. brian m. carlsonAug 22, 2018
  9. Jeff KingAug 22, 2018
  10. Jeff KingAug 22, 2018
  11. Derrick StoleeAug 22, 2018
  12. brian m. carlsonAug 22, 2018
  13. Jeff KingAug 22, 2018
  14. Ævar Arnfjörð BjarmasonAug 22, 2018
  15. Derrick StoleeAug 22, 2018
  16. Jeff KingAug 22, 2018
  17. Duy NguyenAug 22, 2018
  18. Duy NguyenAug 22, 2018
  19. Jeff KingAug 22, 2018
  20. Derrick StoleeAug 22, 2018
  21. Duy NguyenAug 22, 2018
  22. Derrick StoleeAug 22, 2018
  23. Jeff KingAug 22, 2018
  24. Junio C HamanoAug 22, 2018
  25. Jeff KingAug 22, 2018
  26. Derrick StoleeAug 22, 2018
  27. Jeff KingAug 22, 2018
  28. Paul SmithAug 22, 2018
  29. Jeff KingAug 22, 2018
  30. Jonathan NiederAug 23, 2018
  31. Jeff KingAug 23, 2018
  32. Jonathan NiederAug 23, 2018
  33. Jeff KingAug 23, 2018
  34. brian m. carlsonAug 23, 2018
  35. Jonathan NiederAug 23, 2018
  36. Junio C HamanoAug 23, 2018
  37. wide t/perf output, was Re: [ANNOUNCE] Git v2.19.0-rc0Jeff King, Aug 23, 2018
  38. brian m. carlsonAug 23, 2018
  39. Jeff KingAug 23, 2018
  40. Derrick StoleeAug 23, 2018
  41. Junio C HamanoAug 23, 2018
  42. Jeff KingAug 23, 2018
  43. Jacob KellerAug 23, 2018
  44. Jeff KingAug 23, 2018
  45. Jeff KingAug 24, 2018
  46. Jeff KingAug 24, 2018
  47. Jacob KellerAug 24, 2018
  48. Jeff KingAug 24, 2018
  49. Jeff KingAug 24, 2018
  50. Derrick StoleeAug 24, 2018
  51. Junio C HamanoAug 27, 2018
  52. Jeff KingAug 23, 2018
  53. Derrick StoleeAug 23, 2018
  54. Jeff KingAug 24, 2018
  55. Ævar Arnfjörð BjarmasonAug 24, 2018
  56. Derrick StoleeAug 24, 2018
  57. Jeff KingAug 25, 2018
  58. Kaartic SivaraamSep 2, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.