git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: SHA1 collisions found

From
IJIan Jackson <ijackson@chiark.greenend.org.uk>
Date
Mar 3, 2017, 14:54 UTC
Message-ID
<22713.33728.502854.338516@chiark.greenend.org.uk>
In-Reply-To
<20170303111347.6uzuhvmpdwr27qjw@sigill.intra.peff.net>
Jeff King writes ("Re: SHA1 collisions found"):
Show 12 quoted lines
> I think you've read more into my "conversion" than I intended. The old
> history won't get rewritten. It will just be grafted onto the bottom of
> the commit history you've got, and the new trees will all be written
> with the new hash.
> 
> So you still have those old objects hanging around that refer to things
> by their sha1 (not to mention bug trackers, commit messages, etc, which
> all use commit ids). And you want to be able to quickly resolve those
> references.
> 
> What _does_ get rewritten is what's in your ref files, your pack .idx,
> etc. Those are all sha256 (or whatever), and work as sha1's do now.
This all sounds very similar to my proposal.
> Looking up a sha1 reference from an old object just goes through the
> extra level of indirection.

I don't understand why this is a level of indirection, rather than simply a retention of the existing SHA-1 object database (in parallel, but deprecated).

Perhaps I have misunderstood what you mean by "graft". I assume you don't mean info/grafts, because that is not conveyed by transfer protocols.

Stepping back a bit, the key question is what the data structure will look like after the transition.

Specifically, the parent reference in the first post-transition commit has to refer to something. What does it refer to ? The possibilities seem to be:

 1a. It names the SHA1 hash of an old commit object
 1b. It names the BLAKE[1] hash of an old commit object, which
    object of course refers to its parents by SHA1.
 2. It names the BLAKE hash of a translated version of the
    old commit object.
 3. It doesn't name the parent, and the old history is not
    automatically transferred by clone and not readily accessible.

(1a) and (1b) are different variants of something like my mixed hash proposal. Old and new hashes live side by side.

(2) involves rewriting all of the old history, to recursively generate new objects (with BLAKE names, and which in turn refer to other rewritten old objects by BLAKE names). The first time a particular tree needs to look up an object by a BLAKE name, it needs to run a conversion its own entire existing history.

For (2) there would have to be some kind of mapping table in every tree, which allowed object names to be maped in both directions. The object content translation would have to be reversible, so that the actual pre-translation objects would not need to be stored; rather they would be reconstructed from the post-translation objects, when someone asks for a pre-translation object. In principle it would be possible to convert future BLAKE commits to SHA-1 ones, again by recursive rewriting.

I don't think anyone is seriously suggesting (3).
So there is a choice between:

(1) a unified hash tree containing a mixture of different hashes at different reference points, where each object has one identity and one name.

(2) parallel hash tree structures, each using only a single hash, with at least every old object present in both tree structures.

I think (1) is preferable because it provides, to callers of git, the existing object naming semantics. Callers need to be taught to accept an extension to the object name format. Existing object names stored elsewhere than in git remain valid.

Conversely, (2) requires many object names stored elsewhere than in git to be updated. It's possible with (2) to do ad-hoc lookups on object names in mailing list messages or commit messages and so on. Even if it is possible for the new git to answer questions like "does this new branch with BLAKE hash X' contain the commit with SHA1 hash Y" by implicitly looking up the corresponding BLAKE commit Y' and answering the question with reference to Y', this isn't going to help if external code does things like "have git compute the merge base of X and Y' and check that it is equal to Z". Either the external database's record of Z would have to be changed to refer to Z', or the external code would have to be taught to apply an object name normalisation operation to turn Z into Z' each time.

Also, (2) has trouble with shallow clones. This is because it's not possible to translate old objects to new ones without descending to the roots of the object tree and recursively translating everything (or looking up the answer of a previous translation).

Then there is the question of naming syntax.

The distinction between (1) single unified mixed hash tree, and (2) multiple parallel homogenous hash trees, is mostly orthogonal to the command-line (and in-commit-object etc.) representation of new hashes.

The main thing here is that, regardless of the choice between (1) or (2), we need to choose whether object names specified on the git command line, and printed by normal git commands, explicitly identify the hash function.

I think there are roughly three options:
 (A) Decorate all new hashes with a hash function indication
     (sha256:<hex> or blake_<hex> or H<hex>)
 (B) Infer the hash function from the object name length
     (and do some kind of bodge for abbreviated object names).
 (C) Hash function is implicit from context.  (This is compatible with
     (2) only, because (1) requires any object to be able to refer to
     any hash function.)

I think (A) is best because it means everything is unambiguous, and it allows future hash function changes without further difficulty.

(B) is a reasonable possibility although the abbreviated object name bodge would be quite ugly and probably involve thinking about several annoying edge cases.

I think (C) is really bad, because it instantly makes all existing application code which calls git to be buggy. Such application code would need to be adjusted to know for itself which of the object names it has recorded are what hash function, and explicitly specify this to its git operations somehow.

All of these options involve updating many callers of git. In any case any git caller which explicitly checks the object name length will need to be changed. For (a), many git callers which match object names using something like [0-9a-f]+ rather than \w+ will need to be changed - but at least it's a simple change with little semantic import.

(A) has the additional advantage that it becomes possible to make object names syntactically distinguishable from ref names.

The final argument I would make is this:

We don't know what hash function research will look like in 10-20 years. We would like to not have a bunch of pain again. So ideally we would deploy a framework now that would let us switch hash function again without further history-rewriting.

(1)(A) and perhaps (1)(B) are the only options which support this well.

Ian.

[1] I'm going to keep assuming that the bikeshed will be blue, because I think BLAKE2b has is a better choice. It has probably had more serious people looking at it than SHA-3, at least, and it has good performance. The web page has an impressive adoption list - probably wider than SHA-3.

-- 
Ian Jackson <ijackson@chiark.greenend.org.uk>   These opinions are my own.

If I emailed you from an address @fyvzl.net or @evade.org.uk, that is
a private address which bypasses my fierce spamfilter.
Previous: Jeff KingNext: Jeff King
Message 71 of 136 in “SHA1 collisions found”
  1. Joey HessFeb 23, 2017
  2. Junio C HamanoFeb 23, 2017
  3. Junio C HamanoFeb 23, 2017
  4. David LangFeb 23, 2017
  5. Jakub NarębskiFeb 23, 2017
  6. Jeff KingFeb 23, 2017
  7. Joey HessFeb 23, 2017
  8. Linus TorvaldsFeb 23, 2017
  9. Joey HessFeb 23, 2017
  10. Linus TorvaldsFeb 23, 2017
  11. Jeff KingFeb 23, 2017
  12. Linus TorvaldsFeb 23, 2017
  13. Jeff KingFeb 23, 2017
  14. Linus TorvaldsFeb 23, 2017
  15. Jeff KingFeb 23, 2017
  16. Øyvind A. HolmFeb 23, 2017
  17. Joey HessFeb 23, 2017
  18. Joey HessFeb 23, 2017
  19. Morten WelinderFeb 23, 2017
  20. Geert UytterhoevenFeb 24, 2017
  21. Jeff KingFeb 23, 2017
  22. David LangFeb 23, 2017
  23. David LangFeb 23, 2017
  24. David LangFeb 23, 2017
  25. Linus TorvaldsFeb 23, 2017
  26. Linus TorvaldsFeb 23, 2017
  27. Joey HessFeb 23, 2017
  28. Linus TorvaldsFeb 23, 2017
  29. Junio C HamanoFeb 23, 2017
  30. Duy NguyenFeb 24, 2017
  31. brian m. carlsonFeb 25, 2017
  32. René ScharfeFeb 27, 2017
  33. brian m. carlsonFeb 28, 2017
  34. Ian JacksonFeb 24, 2017
  35. ankostisFeb 24, 2017
  36. Junio C HamanoFeb 24, 2017
  37. David LangFeb 24, 2017
  38. Junio C HamanoFeb 24, 2017
  39. Stefan BellerFeb 24, 2017
  40. Junio C HamanoFeb 24, 2017
  41. Junio C HamanoFeb 24, 2017
  42. ankostisFeb 24, 2017
  43. Junio C HamanoFeb 24, 2017
  44. ankostisFeb 25, 2017
  45. Jason CooperFeb 26, 2017
  46. brian m. carlsonFeb 26, 2017
  47. Linus TorvaldsFeb 26, 2017
  48. Ævar Arnfjörð BjarmasonFeb 26, 2017
  49. Jeff KingFeb 26, 2017
  50. Transition plan for git to move to a new hash functionIan Jackson, Feb 27, 2017
  51. Markus TrippelsdorfFeb 27, 2017
  52. Ian JacksonFeb 27, 2017
  53. Tony FinchFeb 27, 2017
  54. brian m. carlsonFeb 28, 2017
  55. Ian JacksonMar 2, 2017
  56. brian m. carlsonMar 4, 2017
  57. Ian JacksonMar 5, 2017
  58. brian m. carlsonMar 5, 2017
  59. Philip OakleyFeb 24, 2017
  60. Jeff KingFeb 24, 2017
  61. Linus TorvaldsFeb 25, 2017
  62. Linus TorvaldsFeb 25, 2017
  63. Jeff KingFeb 25, 2017
  64. Junio C HamanoFeb 26, 2017
  65. Junio C HamanoFeb 25, 2017
  66. Jason CooperFeb 26, 2017
  67. Jeff KingFeb 26, 2017
  68. brian m. carlsonFeb 26, 2017
  69. Brandon WilliamsMar 2, 2017
  70. Jeff KingMar 3, 2017
  71. Ian JacksonMar 3, 2017
  72. Jeff KingMar 3, 2017
  73. Linus TorvaldsMar 2, 2017
  74. Junio C HamanoMar 2, 2017
  75. Linus TorvaldsMar 2, 2017
  76. Joey HessMar 2, 2017
  77. Linus TorvaldsMar 2, 2017
  78. Mike HommeyMar 3, 2017
  79. Linus TorvaldsMar 3, 2017
  80. Jeff KingMar 3, 2017
  81. Stefan BellerMar 3, 2017
  82. David LangFeb 25, 2017
  83. Stefan BellerFeb 25, 2017
  84. Jeff KingFeb 25, 2017
  85. David LangFeb 25, 2017
  86. Jeff KingFeb 25, 2017
  87. David LangFeb 25, 2017
  88. Jacob KellerFeb 25, 2017
  89. Jacob KellerFeb 25, 2017
  90. grarpampFeb 25, 2017
  91. Ian JacksonFeb 24, 2017
  92. Ian JacksonFeb 25, 2017
  93. brian m. carlsonFeb 25, 2017
  94. Jeff KingFeb 25, 2017
  95. Mike HommeyFeb 25, 2017
  96. brian m. carlsonFeb 26, 2017
  97. Jason CooperFeb 24, 2017
  98. ankostisFeb 25, 2017
  99. Jakub NarębskiFeb 24, 2017
  100. Santiago TorresFeb 24, 2017
  101. Jakub NarębskiFeb 24, 2017
  102. Øyvind A. HolmFeb 24, 2017
  103. Jeff KingFeb 24, 2017
  104. Jakub NarębskiFeb 24, 2017
  105. Lars SchneiderFeb 25, 2017
  106. Jeff KingFeb 26, 2017
  107. Junio C HamanoFeb 26, 2017
  108. Thomas BraunFeb 26, 2017
  109. Jeff KingFeb 26, 2017
  110. Geert UytterhoevenFeb 27, 2017
  111. Jeff KingFeb 27, 2017
  112. Morten WelinderFeb 27, 2017
  113. Jeff KingFeb 23, 2017
  114. Linus TorvaldsFeb 23, 2017
  115. Jeff KingFeb 23, 2017
  116. 1/3 add collision-detecting sha1 implementationJeff King, Feb 23, 2017
  117. Stefan BellerFeb 23, 2017
  118. Jeff KingFeb 24, 2017
  119. Linus TorvaldsFeb 24, 2017
  120. Jeff KingFeb 24, 2017
  121. 2/3 sha1dc: adjust header includes for gitJeff King, Feb 23, 2017
  122. 3/3 Makefile: add USE_SHA1DC knobJeff King, Feb 23, 2017
  123. HW42Feb 24, 2017
  124. Jeff KingFeb 24, 2017
  125. Linus TorvaldsFeb 23, 2017
  126. Junio C HamanoFeb 28, 2017
  127. Junio C HamanoFeb 28, 2017
  128. Jeff KingFeb 28, 2017
  129. Dan ShumowMar 1, 2017
  130. Linus TorvaldsFeb 28, 2017
  131. Shawn PearceFeb 28, 2017
  132. Linus TorvaldsFeb 28, 2017
  133. Dan ShumowFeb 28, 2017
  134. Marc StevensFeb 28, 2017
  135. Linus TorvaldsFeb 28, 2017
  136. Jeff KingMar 1, 2017

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.