git/list[1] front-page[2] threads[3] people[4] search[5] about
wed 2026-10-07 18:15 UTC

Re: RFC: Another proposed hash function transition plan

From
Jeff King <peff@peff.net>
Date
Mar 6, 2017, 08:43 UTC
Message-ID
<20170306084353.nrns455dvkdsfgo5@sigill.intra.peff.net>
In-Reply-To
<20170304011251.GA26789@aiede.mtv.corp.google.com>
On Fri, Mar 03, 2017 at 05:12:51PM -0800, Jonathan Nieder wrote:
> This past week we came up with this idea for what a transition to a new
> hash function for Git would look like.  I'd be interested in your
> thoughts (especially if you can make them as comments on the document,
> which makes it easier to address them and update the document).

Overall it's an interesting idea. I thought at first that you were suggesting servers do on-the-fly conversion, but after a more careful reading that isn't the case. And I don't think that would work, because the conversion is expensive.

So this pushes the conversion cost onto the clients who decide to move to SHA-256. That may be a problem for sites which have a lot of clients (like CI hosts). But I guess they would just stick with SHA-1 as long as possible, until the upstream repo switches (and that _is_ a per-repo flag day, because the upstream host isn't going to convert back to SHA-1 on the fly to serve the old clients).

> You can use the doc URL
> 
>  https://goo.gl/gh2Mzc

I'd encourage anybody following along to follow that link. I almost didn't, but there are a ton of comments there (I'm not sure how I feel about splitting the discussion off the list, though).

Show 13 quoted lines
> Goals
> -----
> 1. The transition to SHA256 can be done one local repository at a time.
>    a. Requiring no action by any other party.
>    b. A SHA256 repository can communicate with SHA-1 Git servers and
>       clients (push/fetch).
>    c. Users can use SHA-1 and SHA256 identifiers for objects
>       interchangeably.
>    d. New signed objects make use of a stronger hash function than
>       SHA-1 for their security guarantees.
> 2. Allow a complete transition away from SHA-1.
>    a. Local metadata for SHA-1 compatibility can be dropped in a
>       repository if compatibility with SHA-1 is no longer needed.

I suspect we'll never get away from keeping the mapping table. You'll need at least the sha1->sha256 table if you want to look up names found in historic commit messages, mailing list posts, etc.

And you'll need the sha256->sha1 table if you want to verify the gpg signatures on old tags and commits. That might be something people are willing to drop, though.

Show 14 quoted lines
> After negotiation, the server sends a packfile containing the
> requested objects. We convert the packfile to SHA-256 format using the
> following steps:
> 
> 1. index-pack: inflate each object in the packfile and compute its
>    SHA-1. Objects can contain deltas in OBJ_REF_DELTA format against
>    objects the client has locally. These objects can be looked up using
>    the translation table and their sha1-content read as described above
>    to resolve the deltas.
> 2. topological sort: starting at the "want"s from the negotiation
>    phase, walk through objects in the pack and emit a list of them in
>    topologically sorted order. (This list only contains objects
>    reachable from the "wants". If the pack from the server contained
>    additional extraneous objects, then they will be discarded.)

I don't think we do this right now, but you can actually find the entry (and exit) points of a pack during the index-pack step. Basically:

  1. Keep a hashmap of objects mentioned in the pack.
  2. When we process an object's content (i.e., compute its hash), also
     parse it for any object references. Add entries in the hashmap for
     any object mentioned this way. Mark the entry for the object we
     processed with a "HAVE" bit, and mark any referenced object with a
     "REF" bit.
  3. After processing all objects, anything with a "HAVE" but no "REF"
     is an entry point to the pack (i.e., something that we should have
     asked for with a want). Anything with a "REF" but not a "HAVE" is
     an exit point (i.e., an object that we are expected to already have
     in our repo).
     (I've thought about this before because we could possibly shortcut
     the connectivity check using the exit points. It's complicated by
     the fact that we don't assume the transitive presence of objects
     unless they are reachable).

I don't think using the "want"s as the entry points is unreasonable, though. The server _shouldn't_ generally be sending us other cruft.

I do wonder if you might be able to omit the extra object-graph walk from your step 2, if you could assign "depths" to each object during step 1 instead of HAVE/REF bits. The trouble, of course, is that you're not visiting the nodes in the right order (so given two trees, you're not sure if one might eventually be a child of the other; how do you assign their depths?). I have a feeling there's a proof that it's impossible, but I might just not be clever enough.

Overall the basics of the conversion seem sound to me. The "nohash" things seems more complicated than I think it ought to be, which probably just means I'm missing something. I left a few related comments on the google doc, so I won't repeat them here.

-Peff
Previous: brian m. carlsonNext: Jeff King
Message 5 of 113 in “RFC: Another proposed hash function transition plan”
  1. Jonathan NiederMar 4, 2017
  2. Linus TorvaldsMar 5, 2017
  3. David LangMar 5, 2017
  4. brian m. carlsonMar 6, 2017
  5. Jeff KingMar 6, 2017
  6. Jeff KingMar 6, 2017
  7. Brandon WilliamsMar 6, 2017
  8. Junio C HamanoMar 6, 2017
  9. Jonathan TanMar 6, 2017
  10. Linus TorvaldsMar 6, 2017
  11. Brandon WilliamsMar 6, 2017
  12. Junio C HamanoMar 6, 2017
  13. Jonathan NiederMar 6, 2017
  14. RFC v3: Another proposed hash function transition planJonathan Nieder, Mar 7, 2017
  15. Mike HommeyMar 7, 2017
  16. Jeff KingMar 7, 2017
  17. Linus TorvaldsMar 7, 2017
  18. Ian JacksonMar 7, 2017
  19. Ian JacksonMar 8, 2017
  20. Johannes SchindelinMar 8, 2017
  21. Johannes SchindelinMar 8, 2017
  22. Shawn PearceMar 9, 2017
  23. Jonathan NiederMar 9, 2017
  24. Jeff KingMar 10, 2017
  25. Jonathan NiederMar 10, 2017
  26. The Keccak TeamMar 13, 2017
  27. Jonathan NiederMar 13, 2017
  28. ankostisMar 13, 2017
  29. Johannes SchindelinMar 17, 2017
  30. Use base32?Jason Hennessey, Mar 20, 2017
  31. Michael SteuerMar 20, 2017
  32. Jacob KellerMar 20, 2017
  33. Michael SteuerMar 21, 2017
  34. Which hash function to use, was Re: RFC: Another proposed hash function transition planJohannes Schindelin, Jun 15, 2017
  35. Mike HommeyJun 15, 2017
  36. Jeff KingJun 15, 2017
  37. Ævar Arnfjörð BjarmasonJun 15, 2017
  38. Brandon WilliamsJun 15, 2017
  39. Jonathan NiederJun 15, 2017
  40. Junio C HamanoJun 15, 2017
  41. Johannes SchindelinJun 15, 2017
  42. Mike HommeyJun 15, 2017
  43. Adam LangleyJun 15, 2017
  44. brian m. carlsonJun 15, 2017
  45. Ævar Arnfjörð BjarmasonJun 15, 2017
  46. brian m. carlsonJun 16, 2017
  47. Jeff KingJun 16, 2017
  48. Ævar Arnfjörð BjarmasonJun 16, 2017
  49. Johannes SchindelinJun 16, 2017
  50. Adam LangleyJun 16, 2017
  51. Jeff KingJun 16, 2017
  52. Junio C HamanoJun 16, 2017
  53. Junio C HamanoJun 16, 2017
  54. Jonathan NiederJun 16, 2017
  55. Ævar Arnfjörð BjarmasonJun 16, 2017
  56. Johannes SchindelinJun 19, 2017
  57. Junio C HamanoSep 6, 2017
  58. Junio C HamanoSep 8, 2017
  59. Jeff KingSep 8, 2017
  60. Brandon WilliamsSep 11, 2017
  61. Johannes SchindelinSep 13, 2017
  62. demerphqSep 13, 2017
  63. Jonathan NiederSep 13, 2017
  64. Junio C HamanoSep 13, 2017
  65. Stefan BellerSep 13, 2017
  66. Junio C HamanoSep 13, 2017
  67. Jonathan NiederSep 13, 2017
  68. Jonathan NiederSep 13, 2017
  69. Jonathan NiederSep 13, 2017
  70. Linus TorvaldsSep 13, 2017
  71. Junio C HamanoSep 14, 2017
  72. Junio C HamanoSep 14, 2017
  73. Johannes SchindelinSep 14, 2017
  74. Johannes SchindelinSep 14, 2017
  75. demerphqSep 14, 2017
  76. Brandon WilliamsSep 14, 2017
  77. Johannes SchindelinSep 14, 2017
  78. Jonathan NiederSep 14, 2017
  79. Johannes SchindelinSep 14, 2017
  80. Jonathan NiederSep 14, 2017
  81. Johannes SchindelinSep 14, 2017
  82. Johannes SchindelinSep 14, 2017
  83. Philip OakleySep 15, 2017
  84. Gilles Van AsscheSep 18, 2017
  85. Johannes SchindelinSep 18, 2017
  86. Jonathan NiederSep 18, 2017
  87. Gilles Van AsscheSep 19, 2017
  88. Jason CooperSep 26, 2017
  89. Johannes SchindelinSep 26, 2017
  90. technical doc: add a design doc for hash function transitionStefan Beller, Sep 26, 2017
  91. Jonathan NiederSep 26, 2017
  92. Jonathan NiederSep 26, 2017
  93. technical doc: add a design doc for hash function transitionJonathan Nieder, Sep 28, 2017
  94. Junio C HamanoSep 29, 2017
  95. Junio C HamanoSep 29, 2017
  96. Johannes SchindelinSep 29, 2017
  97. Joan DaemenSep 29, 2017
  98. Jonathan NiederSep 29, 2017
  99. Johannes SchindelinSep 29, 2017
  100. Joan DaemenSep 30, 2017
  101. Junio C HamanoOct 2, 2017
  102. Junio C HamanoOct 2, 2017
  103. Jason CooperOct 2, 2017
  104. Johannes SchindelinOct 2, 2017
  105. Jason CooperOct 2, 2017
  106. Brandon WilliamsOct 2, 2017
  107. Linus TorvaldsOct 2, 2017
  108. Jason CooperOct 2, 2017
  109. Jeff KingOct 2, 2017
  110. Jason CooperOct 2, 2017
  111. Junio C HamanoOct 3, 2017
  112. Jason CooperOct 3, 2017
  113. Junio C HamanoOct 4, 2017

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.