git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH v4 2/5] unpack-trees: add performance tracing

From
Jeff King <peff@peff.net>
Date
Aug 13, 2018, 21:47 UTC
Message-ID
<20180813214702.GA16006@sigill.intra.peff.net>
In-Reply-To
<CACsJy8Cxp9+xiMB6C71Kr63EWyAni-K0ZwVBbpBjUieDbZ+6AA@mail.gmail.com>
On Mon, Aug 13, 2018 at 09:52:41PM +0200, Duy Nguyen wrote:
Show 6 quoted lines
> I don't think I have really fully mastered 'perf'. In this case for
> example, I don't think the default event 'cycles' is the right one
> because we are hit hard by I/O as well. I think at least I now have an
> excuse to try that famous flamegraph out ;-) but if you have time to
> run a quick analysis of this unpack-trees with 'perf', I'd love to
> learn a trick or two from you.

To be honest, I don't feel like I know how to use perf either. ;) But I'll try to contribute what I know.

Usually I'd just use perf to get a callgraph with hot-spots, like:
  perf record -g git ...
  perf report --call-graph=fractal,0.05,caller

But that's not going to show you absolute times, which makes it lousy for comparing run-to-run (if you speed something up, its percentage gets smaller, but it's hard to tell _how much_ you've sped it up). And as you note, it's measuring CPU cycles, not wall-clock.

To get output most similar to what you've shown, I think you'd define some probes at functions of interest:

  for i in unpack_trees cache_tree_update; do
    # Cover both function entrance and return.
    perf probe -x $(which git) $i
    perf probe -x $(which git) ${i}%return
  done
and then record a run looking for those events:
  perf record -e 'probe_git:*' git ...
and then dump the result:
  perf script -F time,event

which gives you the times for each event. If you want elapsed times, you have to compute them yourself:

  perf script -F time,event |
  perl -ne '
    /([0-9.]+):\s+probe_git:(.*):/ or die "confusing: $_";
    my ($t, $func) = ($1, $2);
    if ($func =~ s/__return$//) {
      my $start = pop @stack;
      printf "%0.9f", $t - $start;
      print " s: ";
      print "  " for (0..@stack-1);
      print $func, "\n";
    } else {
      push @stack, $t;
    }
  '

which gives a similar inverted-graph elapsed-time output that your trace output does. One annoying downside is that you have to be root to create or use the dynamic probes. I don't know if there's an easy way around that. Or if there's a perf command which already handles this kind of elapsed stuff (there's a "perf trace" which seems really close, but I couldn't convince it to look at elapsed time for non-syscalls).

Show 11 quoted lines
> > I can buy the argument that it's nice to have some form of profiling
> > that works everywhere, even if it's lowest-common-denominator. I just
> > wonder if we could be investing effort into tooling around existing
> > solutions that will end up more powerful and flexible in the long run.
> 
> I think part of this sprinkling is to highlight the performance
> sensitive spots in the code. And it would be helpful to ask a user to
> enable GIT_TRACE_PERFORMANCE to have a quick breakdown when something
> is reported slow. I don't care that much about other platforms to be
> honest, but perf being largely restricted to root does prevent it from
> replacing GIT_TRACE_PERFORMANCE in this case.

Yeah, this line of reasoning (which is similar to what Stefan said) is compelling to me. GIT_TRACE_* is _most_ useful when we can ask ordinary users to give us output. Even if we scripted the complexity I showed above, it's not guaranteed that perf is even available or that the user has permissions to use it.

-Peff
Previous: Duy NguyenNext: Junio C Hamano
Message 80 of 121 in “[RFC] Speeding up checkout (and merge, rebase, etc)”
  1. 0/3 [RFC] Speeding up checkout (and merge, rebase, etc)Ben Peart, Jul 18, 2018
  2. 1/3 add unbounded Multi-Producer-Multi-Consumer queueBen Peart, Jul 18, 2018
  3. Stefan BellerJul 18, 2018
  4. Junio C HamanoJul 19, 2018
  5. 2/3 add performance tracing around traverse_trees() in unpack_trees()Ben Peart, Jul 18, 2018
  6. 3/3 Add initial parallel version of unpack_trees()Ben Peart, Jul 18, 2018
  7. Junio C HamanoJul 18, 2018
  8. Stefan BellerJul 18, 2018
  9. Jeff KingJul 18, 2018
  10. Ben PeartJul 23, 2018
  11. Duy NguyenJul 23, 2018
  12. Ben PeartJul 23, 2018
  13. Jeff KingJul 24, 2018
  14. Duy NguyenJul 24, 2018
  15. Ben PeartJul 25, 2018
  16. Duy NguyenJul 26, 2018
  17. Duy NguyenJul 26, 2018
  18. Junio C HamanoJul 26, 2018
  19. Duy NguyenJul 27, 2018
  20. Ben PeartJul 27, 2018
  21. Duy NguyenJul 27, 2018
  22. Junio C HamanoJul 27, 2018
  23. Duy NguyenJul 27, 2018
  24. Duy NguyenJul 29, 2018
  25. 0/4 Speed up unpack_trees()Nguyễn Thái Ngọc Duy, Jul 29, 2018
  26. 1/4 unpack-trees.c: add performance tracingNguyễn Thái Ngọc Duy, Jul 29, 2018
  27. Ben PeartJul 30, 2018
  28. 2/4 unpack-trees: optimize walking same trees with cache-treeNguyễn Thái Ngọc Duy, Jul 29, 2018
  29. Ben PeartJul 30, 2018
  30. 3/4 unpack-trees: reduce malloc in cache-tree walkNguyễn Thái Ngọc Duy, Jul 29, 2018
  31. Ben PeartJul 30, 2018
  32. 4/4 unpack-trees: cheaper index update when walking by cache-treeNguyễn Thái Ngọc Duy, Jul 29, 2018
  33. Elijah NewrenAug 8, 2018
  34. Duy NguyenAug 10, 2018
  35. Elijah NewrenAug 10, 2018
  36. Duy NguyenAug 10, 2018
  37. Elijah NewrenAug 10, 2018
  38. Duy NguyenAug 10, 2018
  39. Ben PeartJul 30, 2018
  40. Duy NguyenJul 31, 2018
  41. Ben PeartJul 31, 2018
  42. Ben PeartJul 31, 2018
  43. Duy NguyenAug 1, 2018
  44. Ben PeartAug 8, 2018
  45. Ben PeartAug 9, 2018
  46. Duy NguyenAug 10, 2018
  47. Duy NguyenAug 10, 2018
  48. Ben PeartJul 30, 2018
  49. 0/4 Speed up unpack_trees()Nguyễn Thái Ngọc Duy, Aug 4, 2018
  50. 1/4 unpack-trees: add performance tracingNguyễn Thái Ngọc Duy, Aug 4, 2018
  51. 2/4 unpack-trees: optimize walking same trees with cache-treeNguyễn Thái Ngọc Duy, Aug 4, 2018
  52. Elijah NewrenAug 8, 2018
  53. Duy NguyenAug 10, 2018
  54. Elijah NewrenAug 10, 2018
  55. 3/4 unpack-trees: reduce malloc in cache-tree walkNguyễn Thái Ngọc Duy, Aug 4, 2018
  56. Elijah NewrenAug 8, 2018
  57. 4/4 unpack-trees: cheaper index update when walking by cache-treeNguyễn Thái Ngọc Duy, Aug 4, 2018
  58. Junio C HamanoAug 6, 2018
  59. Duy NguyenAug 6, 2018
  60. Junio C HamanoAug 6, 2018
  61. Ben PeartAug 8, 2018
  62. Junio C HamanoAug 8, 2018
  63. Junio C HamanoAug 8, 2018
  64. Junio C HamanoAug 8, 2018
  65. Duy NguyenAug 10, 2018
  66. 0/5 Speed up unpack_trees()Nguyễn Thái Ngọc Duy, Aug 12, 2018
  67. 3/5 unpack-trees: optimize walking same trees with cache-treeNguyễn Thái Ngọc Duy, Aug 12, 2018
  68. Ben PeartAug 13, 2018
  69. Duy NguyenAug 15, 2018
  70. 1/5 trace.h: support nested performance tracingNguyễn Thái Ngọc Duy, Aug 12, 2018
  71. Ben PeartAug 13, 2018
  72. 2/5 unpack-trees: add performance tracingNguyễn Thái Ngọc Duy, Aug 12, 2018
  73. Thomas AdamAug 12, 2018
  74. Junio C HamanoAug 13, 2018
  75. Ben PeartAug 13, 2018
  76. Jeff KingAug 13, 2018
  77. Stefan BellerAug 13, 2018
  78. Ben PeartAug 13, 2018
  79. Duy NguyenAug 13, 2018
  80. Jeff KingAug 13, 2018
  81. Junio C HamanoAug 13, 2018
  82. Jeff HostetlerAug 14, 2018
  83. Duy NguyenAug 14, 2018
  84. Stefan BellerAug 14, 2018
  85. Duy NguyenAug 14, 2018
  86. Jeff KingAug 14, 2018
  87. Junio C HamanoAug 14, 2018
  88. Duy NguyenAug 15, 2018
  89. Junio C HamanoAug 15, 2018
  90. Jeff HostetlerAug 14, 2018
  91. 4/5 unpack-trees: reduce malloc in cache-tree walkNguyễn Thái Ngọc Duy, Aug 12, 2018
  92. 5/5 unpack-trees: reuse (still valid) cache-tree from src_indexNguyễn Thái Ngọc Duy, Aug 12, 2018
  93. Elijah NewrenAug 13, 2018
  94. Duy NguyenAug 13, 2018
  95. Ben PeartAug 13, 2018
  96. Duy NguyenAug 13, 2018
  97. Ben PeartAug 13, 2018
  98. Junio C HamanoAug 13, 2018
  99. Ben PeartAug 14, 2018
  100. 0/7 Speed up unpack_trees()Nguyễn Thái Ngọc Duy, Aug 18, 2018
  101. 1/7 trace.h: support nested performance tracingNguyễn Thái Ngọc Duy, Aug 18, 2018
  102. 2/7 unpack-trees: add performance tracingNguyễn Thái Ngọc Duy, Aug 18, 2018
  103. 3/7 unpack-trees: optimize walking same trees with cache-treeNguyễn Thái Ngọc Duy, Aug 18, 2018
  104. Ben PeartAug 20, 2018
  105. 5/7 unpack-trees: reuse (still valid) cache-tree from src_indexNguyễn Thái Ngọc Duy, Aug 18, 2018
  106. 6/7 unpack-trees: add missing cache invalidationNguyễn Thái Ngọc Duy, Aug 18, 2018
  107. 4/7 unpack-trees: reduce malloc in cache-tree walkNguyễn Thái Ngọc Duy, Aug 18, 2018
  108. 7/7 cache-tree: verify valid cache-tree in the test suiteNguyễn Thái Ngọc Duy, Aug 18, 2018
  109. Elijah NewrenAug 18, 2018
  110. Elijah NewrenAug 18, 2018
  111. Duy NguyenAug 19, 2018
  112. Document update for nd/unpack-trees-with-cache-treeNguyễn Thái Ngọc Duy, Aug 25, 2018
  113. Martin ÅgrenAug 25, 2018
  114. Document update for nd/unpack-trees-with-cache-treeNguyễn Thái Ngọc Duy, Aug 25, 2018
  115. Ben PeartJul 27, 2018
  116. Duy NguyenJul 26, 2018
  117. Junio C HamanoJul 24, 2018
  118. Duy NguyenJul 24, 2018
  119. Jeff KingJul 24, 2018
  120. Ben PeartJul 25, 2018
  121. Jeff KingJul 24, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.