git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git is not scalable with too many refs/*

From
MFMartin Fick <mfick@codeaurora.org>
Date
Sep 29, 2011, 16:38 UTC
Message-ID
<201109291038.45290.mfick@codeaurora.org>
In-Reply-To
<7c0105c6cca7dd0aa336522f90617fe4@quantumfyre.co.uk>

On Wednesday, September 28, 2011 08:19:16 pm Julian Phillips wrote:

> On Wed, 28 Sep 2011 19:37:18 -0600, Martin Fick wrote:
> > On Wednesday 28 September 2011 18:59:09 Martin Fick 
wrote:
Show 56 quoted lines
> >> Julian Phillips <julian@quantumfyre.co.uk> wrote:
> -- snip --
> 
> >> I've created a test repo with ~100k refs/changes/...
> >> style refs, and ~40000 refs/heads/... style refs, and
> >> checkout can walk the list of ~140k refs seven times
> >> in 85ms user time including doing whatever other
> >> processing is needed for checkout. The real time is
> >> only 114ms - but then my test repo has no real data
> >> in.
> > 
> > If I understand what you are saying, it sounds like you
> > do not have a very good test case. The amount of time
> > it takes for checkout depends on how long it takes to
> > find a ref with the sha1 that you are on. If that sha1
> > is so early in the list of refs that it only took you
> > 7 traversals to find it, then that is not a very good
> > testcase. I think that you should probably try making
> > an orphaned ref (checkout a detached head, commit to
> > it), that is probably the worst testcase since it
> > should then have to search all 140K refs to eventually
> > give up.
> > 
> > Again, if I understand what you are saying, if it took
> > 85ms for 7 traversals, then it takes approximately
> > 10ms per traversal, that's only 100/s! If you have to
> > traverse it 140K times, that should work out to 1400s
> > ~ 23mins.
> 
> Well, it's no more than 10ms per traversal - since the
> rest of the work presumably takes some time too ...
> 
> However, I had forgotten to make the orphaned commit as
> you suggest - and then _bang_ 7N^2, it tries seven
> different variants of each ref (which is silly as they
> are all fully qualified), and with packed refs it has to
> search for them each time, all to turn names into hashes
> that we already know to start with.
> 
> So, yes - it is that list traversal.
> 
> Does the following help?
> 
> diff --git a/builtin/checkout.c b/builtin/checkout.c
> index 5e356a6..f0f4ca1 100644
> --- a/builtin/checkout.c
> +++ b/builtin/checkout.c
> @@ -605,7 +605,7 @@ static int
> add_one_ref_to_rev_list_arg(const char *refname,
>                                         int flags,
>                                         void *cb_data)
>   {
> -       add_one_rev_list_arg(cb_data, refname);
> +       add_one_rev_list_arg(cb_data,
> strdup(sha1_to_hex(sha1))); return 0;
>   }
Yes, but in some strange ways. :)

First, let me clarify that all the tests here involve your "sort fix" from 2 days ago applied first.

In the packed ref repo, it brings the time down to about ~10s (from > 5 mins). In the unpacked ref repo, it brings it down to about the same thing ~10s, but it was only starting at about ~20s.

So, I have to ask, what does that change do, I don't quite understand it? Does it just do only one lookup per ref by normalizing it? Is the list still being traversed, just about 7 time less now? Should the packed_ref list simply be put in an array which could be binary searched instead, it is a fixed list once loaded, no?

I prototyped a packed_ref implementation using the hash.c provided in the git sources and it seemed to speed a checkout up to almost instantaneous, but I was getting a few collisions so the implementation was not good enough. That is when I started to wonder if an array wouldn't be better in this case?

Now I also decided to go back and test a noop fetch (a refetch) of all the changes (since this use case is still taking way longer than I think it should, even with the submodule fix posted earlier). Up until this point, even the sorting fix did not help. So I tried it with this fix. In the unpackref case, it did not seem to change (2~4mins). However, in the packed ref change (which was previously also about 2-4mins), this now only takes about 10-15s!

Any clues as to why the unpacked refs would still be so slow on noop fetches and not be sped up by this?

-Martin
-- 
Employee of Qualcomm Innovation Center, Inc. which is a 
member of Code Aurora Forum
Previous: Julian PhillipsNext: Julian Phillips
Message 46 of 126 in “Git is not scalable with too many refs/*”
  1. NAKAMURA TakumiJun 9, 2011
  2. Sverre RabbelierJun 9, 2011
  3. Shawn PearceJun 9, 2011
  4. A Large Angry SCMJun 9, 2011
  5. Shawn PearceJun 9, 2011
  6. Jeff KingJun 9, 2011
  7. NAKAMURA TakumiJun 10, 2011
  8. Jeff KingJun 13, 2011
  9. Andreas EricssonJun 14, 2011
  10. Jeff KingJun 14, 2011
  11. Junio C HamanoJun 14, 2011
  12. Sverre RabbelierJun 14, 2011
  13. Johan HerlandJun 14, 2011
  14. Sverre RabbelierJun 14, 2011
  15. Jeff KingJun 14, 2011
  16. Shawn PearceJun 14, 2011
  17. Jeff KingJun 14, 2011
  18. Shawn PearceJun 14, 2011
  19. Martin FickSep 8, 2011
  20. Martin FickSep 9, 2011
  21. Thomas RastSep 9, 2011
  22. Thomas RastSep 9, 2011
  23. Jens LehmannSep 9, 2011
  24. Martin FickSep 25, 2011
  25. Christian CouderSep 26, 2011
  26. Martin FickSep 26, 2011
  27. Christian CouderSep 26, 2011
  28. Martin FickSep 30, 2011
  29. Martin FickSep 30, 2011
  30. Martin FickSep 30, 2011
  31. Martin FickSep 30, 2011
  32. Junio C HamanoOct 1, 2011
  33. Michael HaggertyOct 2, 2011
  34. Martin FickOct 3, 2011
  35. Michael HaggertyOct 4, 2011
  36. Martin FickOct 3, 2011
  37. Junio C HamanoOct 3, 2011
  38. Michael HaggertyOct 4, 2011
  39. Martin FickOct 8, 2011
  40. Michael HaggertyOct 9, 2011
  41. Martin FickSep 28, 2011
  42. Martin FickSep 28, 2011
  43. Julian PhillipsSep 29, 2011
  44. Martin FickSep 29, 2011
  45. Julian PhillipsSep 29, 2011
  46. Martin FickSep 29, 2011
  47. Julian PhillipsSep 29, 2011
  48. René ScharfeSep 29, 2011
  49. Junio C HamanoSep 29, 2011
  50. refs: Use binary search to lookup refs fasterJulian Phillips, Sep 29, 2011
  51. Junio C HamanoSep 29, 2011
  52. refs: Use binary search to lookup refs fasterJulian Phillips, Sep 29, 2011
  53. Junio C HamanoSep 29, 2011
  54. refs: Use binary search to lookup refs fasterJulian Phillips, Sep 29, 2011
  55. Junio C HamanoSep 29, 2011
  56. Michael HaggertySep 30, 2011
  57. Junio C HamanoSep 30, 2011
  58. refs: Remove duplicates after sorting with qsortJulian Phillips, Sep 30, 2011
  59. Michael HaggertyOct 2, 2011
  60. Junio C HamanoOct 2, 2011
  61. Junio C HamanoOct 4, 2011
  62. Martin FickSep 30, 2011
  63. Junio C HamanoSep 30, 2011
  64. Julian PhillipsSep 30, 2011
  65. Martin FickSep 30, 2011
  66. Martin FickSep 29, 2011
  67. Julian PhillipsSep 29, 2011
  68. Martin FickSep 29, 2011
  69. René ScharfeSep 30, 2011
  70. Martin FickSep 30, 2011
  71. Junio C HamanoSep 30, 2011
  72. René ScharfeSep 30, 2011
  73. René ScharfeOct 1, 2011
  74. 1/8 checkout: check for "Previous HEAD" notice in t2020René Scharfe, Oct 1, 2011
  75. Sverre RabbelierOct 1, 2011
  76. 2/8 revision: factor out add_pending_sha1René Scharfe, Oct 1, 2011
  77. 3/8 checkout: use add_pending_{object,sha1} in orphan checkRené Scharfe, Oct 1, 2011
  78. 4/8 revision: add leak_pending flagRené Scharfe, Oct 1, 2011
  79. 5/8 bisect: use leak_pending flagRené Scharfe, Oct 1, 2011
  80. 6/8 bundle: use leak_pending flagRené Scharfe, Oct 1, 2011
  81. 7/8 checkout: use leak_pending flagRené Scharfe, Oct 1, 2011
  82. 8/8 commit: factor out clear_commit_marks_for_object_arrayRené Scharfe, Oct 1, 2011
  83. Martin FickSep 26, 2011
  84. Sverre RabbelierSep 26, 2011
  85. Martin FickSep 26, 2011
  86. Sverre RabbelierSep 26, 2011
  87. Martin FickSep 26, 2011
  88. Julian PhillipsSep 26, 2011
  89. Martin FickSep 26, 2011
  90. Julian PhillipsSep 26, 2011
  91. Martin FickSep 26, 2011
  92. Junio C HamanoSep 26, 2011
  93. Julian PhillipsSep 26, 2011
  94. Martin FickSep 26, 2011
  95. Martin FickSep 26, 2011
  96. Julian PhillipsSep 26, 2011
  97. David Michael BarrSep 26, 2011
  98. refs.c: Fix slowness with numerous loose refsDavid Barr, Sep 27, 2011
  99. David Michael BarrSep 27, 2011
  100. Junio C HamanoSep 26, 2011
  101. Don't sort ref_list too earlyJulian Phillips, Sep 27, 2011
  102. Michael HaggertyOct 2, 2011
  103. Martin FickSep 27, 2011
  104. Julian PhillipsSep 27, 2011
  105. Martin FickSep 27, 2011
  106. Julian PhillipsSep 27, 2011
  107. Sverre RabbelierSep 27, 2011
  108. Julian PhillipsSep 27, 2011
  109. Sverre RabbelierSep 27, 2011
  110. Nguyen Thai Ngoc DuySep 27, 2011
  111. Michael HaggertySep 27, 2011
  112. Julian PhillipsSep 27, 2011
  113. Julian PhillipsSep 26, 2011
  114. Michael HaggertySep 26, 2011
  115. Martin FickSep 26, 2011
  116. Thomas RastSep 26, 2011
  117. Michael HaggertySep 9, 2011
  118. Michael HaggertySep 9, 2011
  119. Jens LehmannSep 9, 2011
  120. Andreas EricssonJun 10, 2011
  121. Shawn PearceJun 10, 2011
  122. Jakub NarebskiJun 10, 2011
  123. Jeff KingJun 10, 2011
  124. Andreas EricssonJun 13, 2011
  125. Jakub NarebskiJun 9, 2011
  126. Stephen BashJun 9, 2011

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.