git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: RFC v3: Another proposed hash function transition plan

From
Jeff King <peff@peff.net>
Date
Mar 10, 2017, 19:38 UTC
Message-ID
<20170310193835.t7syswueuu7nmkjz@sigill.intra.peff.net>
In-Reply-To
<20170309202408.GA17847@aiede.mtv.corp.google.com>
On Thu, Mar 09, 2017 at 12:24:08PM -0800, Jonathan Nieder wrote:
Show 8 quoted lines
> > SHA-1 to SHA-3: lookup SHA-1 in .msha1, reverse .idx, find offset to
> > read the SHA-3.
> > SHA-3 to SHA-1: lookup SHA-3 in .idx, and reverse the .msha1 file to
> > translate offset to SHA-1.
> 
> Thanks for this suggestion.  I was initially vaguely nervous about
> lookup times in an idx-style file, but as you say, object reads from a
> packfile already have to deal with this kind of lookup and work fine.

Not exactly. The "reverse .idx" step has to build the reverse mapping on the fly, and it's non-trivial. For instance, try:

  sha1=$(git rev-parse HEAD)
  time echo $sha1 | git cat-file --batch-check='%(objectsize)'
  time echo $sha1 | git cat-file --batch-check='%(objectsize:disk)'

on a large repo (where HEAD is in a big pack). The on-disk size is conceptually simpler, as we only need to look at the offset of the object versus the offset of the object after it. But in practice it takes much longer, because it has to build the revindex on the fly (I get 7ms versus 179ms on linux.git).

The effort is linear in the number of objects (we create the revindex with a radix sort).

The reachability bitmaps suffer from this, too, as they need the revindex to know which object is at which bit position. At GitHub we added an extension to the .bitmap files that stores this "bit cache". Here are timings before and after on linux.git:

  $ time git rev-list --use-bitmap-index --count master
  659371
  real	0m0.182s
  user	0m0.136s
  sys	0m0.044s
  $ time git.gh rev-list --use-bitmap-index --count master
  659371
  real	0m0.016s
  user	0m0.008s
  sys	0m0.004s

It's not a full revindex, but it's enough for bitmap use. You can also use it to generate the revindex slightly more quickly, because you can skip the sorting step (you just insert the entries in the correct order by walking the bit cache and dereferencing the offsets from the .idx portion). So it's still linear, but with a smaller constant factor.

I think for the purposes here, though, we don't actually care about the offsets. For the cost of one uint32_t per object, you can keep a list mapping positions in the sha1 index into the sha3 index. So then you do the log-n binary search to find the sha1, a constant-time lookup in the mapping array, and that gives you the position in the sha3 index, from which you can then access the sha3 (or the actual pack offset, for that matter).

So I think it's solvable, but I suspect we would want an extension to the .idx format to store the mapping array, in order to keep it log-n.

-Peff
Previous: Jonathan NiederNext: Jonathan Nieder
Message 31 of 113 in “RFC: Another proposed hash function transition plan”
  1. Jonathan NiederMar 4, 2017
  2. Linus TorvaldsMar 5, 2017
  3. brian m. carlsonMar 6, 2017
  4. Brandon WilliamsMar 6, 2017
  5. Which hash function to use, was Re: RFC: Another proposed hash function transition planJohannes Schindelin, Jun 15, 2017
  6. Mike HommeyJun 15, 2017
  7. Jeff KingJun 15, 2017
  8. Ævar Arnfjörð BjarmasonJun 15, 2017
  9. Johannes SchindelinJun 15, 2017
  10. Adam LangleyJun 15, 2017
  11. brian m. carlsonJun 15, 2017
  12. Ævar Arnfjörð BjarmasonJun 15, 2017
  13. brian m. carlsonJun 16, 2017
  14. Ævar Arnfjörð BjarmasonJun 16, 2017
  15. Johannes SchindelinJun 16, 2017
  16. Adam LangleyJun 16, 2017
  17. Junio C HamanoJun 16, 2017
  18. Junio C HamanoJun 16, 2017
  19. Jonathan NiederJun 16, 2017
  20. Ævar Arnfjörð BjarmasonJun 16, 2017
  21. Jeff KingJun 16, 2017
  22. Johannes SchindelinJun 19, 2017
  23. Mike HommeyJun 15, 2017
  24. Jeff KingJun 16, 2017
  25. Brandon WilliamsJun 15, 2017
  26. Junio C HamanoJun 15, 2017
  27. Jonathan NiederJun 15, 2017
  28. RFC v3: Another proposed hash function transition planJonathan Nieder, Mar 7, 2017
  29. Shawn PearceMar 9, 2017
  30. Jonathan NiederMar 9, 2017
  31. Jeff KingMar 10, 2017
  32. Jonathan NiederMar 10, 2017
  33. technical doc: add a design doc for hash function transitionJonathan Nieder, Sep 28, 2017
  34. Junio C HamanoSep 29, 2017
  35. Junio C HamanoSep 29, 2017
  36. Jonathan NiederSep 29, 2017
  37. Junio C HamanoOct 2, 2017
  38. Jason CooperOct 2, 2017
  39. Junio C HamanoOct 2, 2017
  40. Jason CooperOct 2, 2017
  41. Junio C HamanoOct 3, 2017
  42. Jason CooperOct 3, 2017
  43. Junio C HamanoOct 4, 2017
  44. Junio C HamanoSep 6, 2017
  45. Junio C HamanoSep 8, 2017
  46. Jeff KingSep 8, 2017
  47. Brandon WilliamsSep 11, 2017
  48. Johannes SchindelinSep 13, 2017
  49. demerphqSep 13, 2017
  50. Jonathan NiederSep 13, 2017
  51. Johannes SchindelinSep 14, 2017
  52. Jonathan NiederSep 14, 2017
  53. Johannes SchindelinSep 14, 2017
  54. Linus TorvaldsSep 13, 2017
  55. Johannes SchindelinSep 14, 2017
  56. Gilles Van AsscheSep 18, 2017
  57. Johannes SchindelinSep 18, 2017
  58. Gilles Van AsscheSep 19, 2017
  59. Johannes SchindelinSep 29, 2017
  60. Joan DaemenSep 29, 2017
  61. Johannes SchindelinSep 29, 2017
  62. Joan DaemenSep 30, 2017
  63. Johannes SchindelinOct 2, 2017
  64. Jonathan NiederSep 18, 2017
  65. Jason CooperSep 26, 2017
  66. Johannes SchindelinSep 26, 2017
  67. technical doc: add a design doc for hash function transitionStefan Beller, Sep 26, 2017
  68. Jonathan NiederSep 26, 2017
  69. Jonathan NiederSep 26, 2017
  70. Jason CooperOct 2, 2017
  71. Brandon WilliamsOct 2, 2017
  72. Jason CooperOct 2, 2017
  73. Linus TorvaldsOct 2, 2017
  74. Jeff KingOct 2, 2017
  75. Jonathan NiederSep 13, 2017
  76. Junio C HamanoSep 13, 2017
  77. Stefan BellerSep 13, 2017
  78. Jonathan NiederSep 13, 2017
  79. Junio C HamanoSep 14, 2017
  80. Johannes SchindelinSep 14, 2017
  81. demerphqSep 14, 2017
  82. Johannes SchindelinSep 14, 2017
  83. Junio C HamanoSep 13, 2017
  84. Jonathan NiederSep 13, 2017
  85. Junio C HamanoSep 14, 2017
  86. Johannes SchindelinSep 14, 2017
  87. Brandon WilliamsSep 14, 2017
  88. Jonathan NiederSep 14, 2017
  89. Philip OakleySep 15, 2017
  90. David LangMar 5, 2017
  91. Jonathan NiederMar 6, 2017
  92. Mike HommeyMar 7, 2017
  93. Jeff KingMar 6, 2017
  94. Junio C HamanoMar 6, 2017
  95. Jonathan TanMar 6, 2017
  96. Linus TorvaldsMar 6, 2017
  97. Brandon WilliamsMar 6, 2017
  98. Junio C HamanoMar 6, 2017
  99. Jeff KingMar 7, 2017
  100. Ian JacksonMar 7, 2017
  101. Linus TorvaldsMar 7, 2017
  102. Ian JacksonMar 8, 2017
  103. Johannes SchindelinMar 8, 2017
  104. Johannes SchindelinMar 8, 2017
  105. Use base32?Jason Hennessey, Mar 20, 2017
  106. Michael SteuerMar 20, 2017
  107. Jacob KellerMar 20, 2017
  108. Michael SteuerMar 21, 2017
  109. The Keccak TeamMar 13, 2017
  110. Jonathan NiederMar 13, 2017
  111. ankostisMar 13, 2017
  112. Johannes SchindelinMar 17, 2017
  113. Jeff KingMar 6, 2017

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.