git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH v1 0/3] [RFC] Speeding up checkout (and merge, rebase, etc)

From
Ben Peart <ben.peart@microsoft.com>
Date
Jul 18, 2018, 20:45 UTC
Message-ID
<20180718204458.20936-1-benpeart@microsoft.com>

When working directories get big, checkout times start to suffer. Even with GVFS virtualization (which limits git to only having to update those files that have been changed locally) we�re seeing P50 times for checkout of 31 seconds and the P80 time is 43 seconds.

Here is a checkout command with tracing turned on to demonstrate where the time is spent. Note, this is somewhat of a �best case� as I�m simply checking out the current commit:

benpeart@gvfs-perf MINGW64 /f/os/src (official/rs_es_debug_dev) $ /usr/src/git/git.exe checkout 12:31:50.419016 read-cache.c:2006 performance: 1.180966800 s: read cache .git/index 12:31:51.184636 name-hash.c:605 performance: 0.664575200 s: initialize name hash 12:31:51.200280 preload-index.c:111 performance: 0.019811600 s: preload index 12:31:51.294012 read-cache.c:1543 performance: 0.094515600 s: refresh index 12:32:29.731344 unpack-trees.c:1358 performance: 33.889840200 s: traverse_trees 12:32:37.512555 read-cache.c:2541 performance: 1.564438300 s: write index, changed mask = 28 12:32:44.918730 unpack-trees.c:1358 performance: 7.243155600 s: traverse_trees 12:32:44.965611 diff-lib.c:527 performance: 7.374729200 s: diff-index Waiting for GVFS to parse index and update placeholder files...Succeeded 12:32:46.824986 trace.c:420 performance: 57.715656000 s: git command: 'C:\git-sdk-64\usr\src\git\git.exe' checkout

Clearly, most of the time (41 seconds) is spent in the traverse_trees() code so the question is, how can we significantly speed up that portion of the command?

I investigated a few options with limited success:

ODB cache ========= Since traverse_trees() hits the ODB for each tree object (of which there are over 500K in this repo) I wrote and tested having an in-memory ODB cache that cached all tree objects. This resulted in a > 50% hit ratio (largely due to the fact we traverse the tree twice during checkout) but resulted in only a minimal savings (1.3 seconds).

Tree Graph File =============== I also considered storing the commit tree in an alternate structure that is faster to load/parse (ala the Commit graph) but the cache results along with the negligible impact of running checkout back to back (thus ensuring the objects were cached in my file system cache) made me believe this would not result in much savings. MIDX has already helped out here given we end up with a lot of pack files of commits and trees.

Sparse tree traversal ===================== We�ve sped up other parts of git by taking advantage of the existing sparse-checkout/excludes logic to limit what files git has to consider to those that have been modified by the user locally. I haven�t been able to think of a way to take advantage of that with unpack-trees() as when you are merging n commits, a change/conflict can occur in any tree object so they must all be traversed. If I�m missing something here and there _is_ a way to entirely skip large parts of the tree, please let me know! Please note that we�re already limiting the files that git needs to update in the working directory via sparse-checkout/excludes but the other/merge logic still executes for the entire tree whether there are files to update or not.

Multi-threading unpack_trees() ============================== The current model of unpack_trees() is that a single thread recursively traverses each tree object as it comes across it. One thought I had was to multi-thread the traversal so that each tree object could be processed in parallel. To test this idea out, I wrote an unbounded Multi-Product-Multi-Consumer queue and then wrote a traverse_trees_parallel() function that would add any new tree objects into the queue where they can be processed by a pool of worker threads. Each thread will wake up when there is work in the queue, remove a tree object, process it adding any additional tree objects it finds.

Multi-threading anything in git is fraught with challenges as much of the code base is not thread safe. To make progress, I wrapped mutexes around code paths that were not thread safe. The end result is that I won�t initially get much parallelization (due to mutexes around all the expensive work) but at least I can test out the idea and resolve any other issues with switching from a serial to a parallel implementation. If this works out, I can update more of the code paths to be thread safe and/or move to more fine grained mutexes around those paths that are difficult to make thread safe.

Final thoughts ==============

The attached set of patches don�t work! For some commands they succeed but I�m including them only to make it explicit what I�m currently investigating. I�d be very interested in design feedback but formatting/spelling/white space errors are less useful at this early stage in the investigation.

When I brought up this idea with some other git contributors they mentioned that multi threading unpack_trees() had been discussed a few years ago on the list but that the idea was discarded. They couldn�t remember exactly why it was discarded and none of us have been able to find the email threads from that earlier discussion. As a result, I decided to write up this RFC and see if the greater git community has ideas, suggestions, or more background/history on whether this is a reasonable path to pursue or if there are other/better ideas on how to speed up checkout especially on large repos.

Base Ref: master
Web-Diff: https://github.com/benpeart/git/commit/a022a91ceb
Checkout: git fetch https://github.com/benpeart/git unpacktrees-v1 && git checkout a022a91ceb
Ben Peart (3):
  add unbounded Multi-Producer-Multi-Consumer queue
  add performance tracing around traverse_trees() in unpack_trees()
  Add initial parallel version of unpack_trees()
 Makefile       |   1 +
 cache.h        |   1 +
 config.c       |   5 +
 environment.c  |   1 +
 mpmcqueue.c    |  49 ++++++++
 mpmcqueue.h    |  80 +++++++++++++
 unpack-trees.c | 314 ++++++++++++++++++++++++++++++++++++++++++++++++-
 unpack-trees.h |  30 +++++
 8 files changed, 480 insertions(+), 1 deletion(-)
 create mode 100644 mpmcqueue.c
 create mode 100644 mpmcqueue.h
base-commit: e3331758f12da22f4103eec7efe1b5304a9be5e9
-- 
2.17.0.gvfs.1.123.g449c066
Next: Ben Peart
Message 1 of 121 in “[RFC] Speeding up checkout (and merge, rebase, etc)”
  1. 0/3 [RFC] Speeding up checkout (and merge, rebase, etc)Ben Peart, Jul 18, 2018
  2. 1/3 add unbounded Multi-Producer-Multi-Consumer queueBen Peart, Jul 18, 2018
  3. Stefan BellerJul 18, 2018
  4. Junio C HamanoJul 19, 2018
  5. 2/3 add performance tracing around traverse_trees() in unpack_trees()Ben Peart, Jul 18, 2018
  6. 3/3 Add initial parallel version of unpack_trees()Ben Peart, Jul 18, 2018
  7. Junio C HamanoJul 18, 2018
  8. Stefan BellerJul 18, 2018
  9. Jeff KingJul 18, 2018
  10. Ben PeartJul 23, 2018
  11. Duy NguyenJul 23, 2018
  12. Ben PeartJul 23, 2018
  13. Jeff KingJul 24, 2018
  14. Duy NguyenJul 24, 2018
  15. Ben PeartJul 25, 2018
  16. Duy NguyenJul 26, 2018
  17. Duy NguyenJul 26, 2018
  18. Junio C HamanoJul 26, 2018
  19. Duy NguyenJul 27, 2018
  20. Ben PeartJul 27, 2018
  21. Duy NguyenJul 27, 2018
  22. Junio C HamanoJul 27, 2018
  23. Duy NguyenJul 27, 2018
  24. Duy NguyenJul 29, 2018
  25. 0/4 Speed up unpack_trees()Nguyễn Thái Ngọc Duy, Jul 29, 2018
  26. 1/4 unpack-trees.c: add performance tracingNguyễn Thái Ngọc Duy, Jul 29, 2018
  27. Ben PeartJul 30, 2018
  28. 2/4 unpack-trees: optimize walking same trees with cache-treeNguyễn Thái Ngọc Duy, Jul 29, 2018
  29. Ben PeartJul 30, 2018
  30. 3/4 unpack-trees: reduce malloc in cache-tree walkNguyễn Thái Ngọc Duy, Jul 29, 2018
  31. Ben PeartJul 30, 2018
  32. 4/4 unpack-trees: cheaper index update when walking by cache-treeNguyễn Thái Ngọc Duy, Jul 29, 2018
  33. Elijah NewrenAug 8, 2018
  34. Duy NguyenAug 10, 2018
  35. Elijah NewrenAug 10, 2018
  36. Duy NguyenAug 10, 2018
  37. Elijah NewrenAug 10, 2018
  38. Duy NguyenAug 10, 2018
  39. Ben PeartJul 30, 2018
  40. Duy NguyenJul 31, 2018
  41. Ben PeartJul 31, 2018
  42. Ben PeartJul 31, 2018
  43. Duy NguyenAug 1, 2018
  44. Ben PeartAug 8, 2018
  45. Ben PeartAug 9, 2018
  46. Duy NguyenAug 10, 2018
  47. Duy NguyenAug 10, 2018
  48. Ben PeartJul 30, 2018
  49. 0/4 Speed up unpack_trees()Nguyễn Thái Ngọc Duy, Aug 4, 2018
  50. 1/4 unpack-trees: add performance tracingNguyễn Thái Ngọc Duy, Aug 4, 2018
  51. 2/4 unpack-trees: optimize walking same trees with cache-treeNguyễn Thái Ngọc Duy, Aug 4, 2018
  52. Elijah NewrenAug 8, 2018
  53. Duy NguyenAug 10, 2018
  54. Elijah NewrenAug 10, 2018
  55. 3/4 unpack-trees: reduce malloc in cache-tree walkNguyễn Thái Ngọc Duy, Aug 4, 2018
  56. Elijah NewrenAug 8, 2018
  57. 4/4 unpack-trees: cheaper index update when walking by cache-treeNguyễn Thái Ngọc Duy, Aug 4, 2018
  58. Junio C HamanoAug 6, 2018
  59. Duy NguyenAug 6, 2018
  60. Junio C HamanoAug 6, 2018
  61. Ben PeartAug 8, 2018
  62. Junio C HamanoAug 8, 2018
  63. Junio C HamanoAug 8, 2018
  64. Junio C HamanoAug 8, 2018
  65. Duy NguyenAug 10, 2018
  66. 0/5 Speed up unpack_trees()Nguyễn Thái Ngọc Duy, Aug 12, 2018
  67. 3/5 unpack-trees: optimize walking same trees with cache-treeNguyễn Thái Ngọc Duy, Aug 12, 2018
  68. Ben PeartAug 13, 2018
  69. Duy NguyenAug 15, 2018
  70. 1/5 trace.h: support nested performance tracingNguyễn Thái Ngọc Duy, Aug 12, 2018
  71. Ben PeartAug 13, 2018
  72. 2/5 unpack-trees: add performance tracingNguyễn Thái Ngọc Duy, Aug 12, 2018
  73. Thomas AdamAug 12, 2018
  74. Junio C HamanoAug 13, 2018
  75. Ben PeartAug 13, 2018
  76. Jeff KingAug 13, 2018
  77. Stefan BellerAug 13, 2018
  78. Ben PeartAug 13, 2018
  79. Duy NguyenAug 13, 2018
  80. Jeff KingAug 13, 2018
  81. Junio C HamanoAug 13, 2018
  82. Jeff HostetlerAug 14, 2018
  83. Duy NguyenAug 14, 2018
  84. Stefan BellerAug 14, 2018
  85. Duy NguyenAug 14, 2018
  86. Jeff KingAug 14, 2018
  87. Junio C HamanoAug 14, 2018
  88. Duy NguyenAug 15, 2018
  89. Junio C HamanoAug 15, 2018
  90. Jeff HostetlerAug 14, 2018
  91. 4/5 unpack-trees: reduce malloc in cache-tree walkNguyễn Thái Ngọc Duy, Aug 12, 2018
  92. 5/5 unpack-trees: reuse (still valid) cache-tree from src_indexNguyễn Thái Ngọc Duy, Aug 12, 2018
  93. Elijah NewrenAug 13, 2018
  94. Duy NguyenAug 13, 2018
  95. Ben PeartAug 13, 2018
  96. Duy NguyenAug 13, 2018
  97. Ben PeartAug 13, 2018
  98. Junio C HamanoAug 13, 2018
  99. Ben PeartAug 14, 2018
  100. 0/7 Speed up unpack_trees()Nguyễn Thái Ngọc Duy, Aug 18, 2018
  101. 1/7 trace.h: support nested performance tracingNguyễn Thái Ngọc Duy, Aug 18, 2018
  102. 2/7 unpack-trees: add performance tracingNguyễn Thái Ngọc Duy, Aug 18, 2018
  103. 3/7 unpack-trees: optimize walking same trees with cache-treeNguyễn Thái Ngọc Duy, Aug 18, 2018
  104. Ben PeartAug 20, 2018
  105. 5/7 unpack-trees: reuse (still valid) cache-tree from src_indexNguyễn Thái Ngọc Duy, Aug 18, 2018
  106. 6/7 unpack-trees: add missing cache invalidationNguyễn Thái Ngọc Duy, Aug 18, 2018
  107. 4/7 unpack-trees: reduce malloc in cache-tree walkNguyễn Thái Ngọc Duy, Aug 18, 2018
  108. 7/7 cache-tree: verify valid cache-tree in the test suiteNguyễn Thái Ngọc Duy, Aug 18, 2018
  109. Elijah NewrenAug 18, 2018
  110. Elijah NewrenAug 18, 2018
  111. Duy NguyenAug 19, 2018
  112. Document update for nd/unpack-trees-with-cache-treeNguyễn Thái Ngọc Duy, Aug 25, 2018
  113. Martin ÅgrenAug 25, 2018
  114. Document update for nd/unpack-trees-with-cache-treeNguyễn Thái Ngọc Duy, Aug 25, 2018
  115. Ben PeartJul 27, 2018
  116. Duy NguyenJul 26, 2018
  117. Junio C HamanoJul 24, 2018
  118. Duy NguyenJul 24, 2018
  119. Jeff KingJul 24, 2018
  120. Ben PeartJul 25, 2018
  121. Jeff KingJul 24, 2018

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.