git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: RFC v3: Another proposed hash function transition plan

From
Johannes Schindelin <johannes.schindelin@gmx.de>
Date
Sep 13, 2017, 12:05 UTC
Message-ID
<alpine.DEB.2.21.1.1709131340030.4132@virtualbox>
In-Reply-To
<20170911185913.GA5869@google.com>
Hi Brandon,
On Mon, 11 Sep 2017, Brandon Williams wrote:
Show 21 quoted lines
> On 09/08, Junio C Hamano wrote:
> > Junio C Hamano <gitster@pobox.com> writes:
> > 
> > > One thing I still do not know how I feel about after re-reading the
> > > thread, and I didn't find the above doc, is Linus's suggestion to
> > > use the objects themselves as NewHash-to-SHA-1 mapper [*1*].  
> > > ...
> > > [Reference]
> > >
> > > *1* <CA+55aFxj7Vtwac64RfAz_u=U4tob4Xg+2pDBDFNpJdmgaTCmxA@mail.gmail.com>
> > 
> > I think this falls into the same category as the often-talked-about
> > addition of the "generation number" field.  It is very tempting to add
> > these "mechanically derivable but expensive to compute" pieces of
> > information to the sha3-content while converting from sha1-content and
> > creating anew.  
> 
> We didn't discuss that in the doc since this particular transition plan
> we made uses an external NewHash-to-SHA1 map instead of an internal one
> because we believe that at some point we would be able to drop
> compatibility with SHA1.

Is there even a question about that? I mean, why would *any* project that switches entirely to SHA-256 want to carry the SHA-1 baggage around?

So even if the code to generate a bidirectional old <-> new hash mapping might be with us forever, it *definitely* should be optional ("optional" at least as in "config setting"), allowing developers who only work with new-hash repositories to save the time and electrons.

> Now I suspect that wont happen for a long time but I think it would be
> preferable over carrying the SHA1 luggage indefinitely.

It should be possible to push back the SHA-1 ginny into a small gin bottle inside Git's source code, so to say, i.e. encapsulate it to the point where it is a compile-time option, in addition to a runtime option.

Of course, that's only unless the SHA-1 calculation is made mandatory as suggested above. I really shudder at the idea of requiring SHA-1 to be required forever. We ignored advice in 2005 against making ourselves too dependent on SHA-1, and I would hope that we would learn from this.

> At some point, then, we would be able to stop hashing objects twice
> (once with SHA1 and once with NewHash) instead of always requiring that
> we hash them with each hash function which was used historically.
Yes, please.
> > Because the "sha1-name" or the "generation number" can mechanically
> > be computed,
... as long as a shallow clone you do not have, of course...
Show 6 quoted lines
> > as long as everybody agrees to _always_ place them in the
> > sha3-content, the same sha1-content will be converted into exactly the
> > same sha3-content without ambiguity, and converting them back to
> > sha1-content while pushing to an older repository will correctly
> > produce the original sha1-content, as it would just be the matter of
> > simply stripping these extra pieces of information.

... or Git would simply handle the absence of the generation number header gracefully, so that sha1-content == sha3-content...

Show 30 quoted lines
> > The same thing could happen if we decide to bake "generation number"
> > in the SHA-3 commit objects.  One possible definition would be that a
> > root commit will have gen #0; a commit with 1 or more parents will get
> > max(parents' gen numbers) + 1 as its gen number.  But somebody may
> > botch the counting and records sum(parents' gen numbers) as its gen
> > number.
> > 
> > In these cases, not just the SHA3-content but also the resulting SHA-3
> > object name would be different from the name of the object that would
> > have recorded the same contents correctly.  So converting back to
> > SHA-1 world from these botched SHA-3 contents may produce the original
> > contents, but we may end up with multiple "plausibly looking" set of
> > SHA-3 objects that (clain to) correspond to a single SHA-1 object,
> > only one of which is a valid one.
> > 
> > Our "git fsck" already treats certain brokenness (like a tree whose
> > entry has mode that is 0-padded to the left) as broken but still
> > tolerate them.  I am not sure if it is sufficient to diagnose and
> > declare broken and invalid when we see sha3-content that records
> > these "mechanically derivable but expensive to compute" pieces of
> > information incorrectly.
> > 
> > I am leaning towards saying "yes, catching in fsck is enough" and
> > suggesting to add generation number to sha3-content of the commit
> > objects, and to add even the "original sha1 name" thing if we find
> > good use of it.  But I cannot shake this nagging feeling off that I
> > am missing some huge problems that adding these fields and opening
> > ourselves to more classes of broken objects.
> > 
> > Thoughts?

Seeing as current Git versions would always ignore the generation number (and therefore work perfectly even with erroneous baked-in generation numbers), and seeing as it would be easy to add a config option to force Git to ignore the embedded generation numbers, I would consider `fsck` catching those problems the best idea.

It seems that every major Git hoster already has some sort of fsck on the fly for newly-pushed objects, so that would be another "line of defense".

Taking a step back, though, it may be a good idea to leave the generation number business for later, as much fun as it is to get side tracked and focus on relatively trivial stuff instead of the far more difficult and complex task to get the transition plan to a new hash ironed out.

For example, I am still in favor of SHA-256 over SHA3-256, after learning some background details from in-house cryptographers: it provides essentially the same level of security, according to my sources, while hardware support seems to be coming to SHA-256 a lot sooner than to SHA3-256.

Which hash algorithm to choose is a tough question to answer, and discussing generation numbers will sadly not help us answer it any quicker.

Ciao, Dscho

Previous: Brandon WilliamsNext: demerphq
Message 48 of 113 in “RFC: Another proposed hash function transition plan”
  1. Jonathan NiederMar 4, 2017
  2. Linus TorvaldsMar 5, 2017
  3. brian m. carlsonMar 6, 2017
  4. Brandon WilliamsMar 6, 2017
  5. Which hash function to use, was Re: RFC: Another proposed hash function transition planJohannes Schindelin, Jun 15, 2017
  6. Mike HommeyJun 15, 2017
  7. Jeff KingJun 15, 2017
  8. Ævar Arnfjörð BjarmasonJun 15, 2017
  9. Johannes SchindelinJun 15, 2017
  10. Adam LangleyJun 15, 2017
  11. brian m. carlsonJun 15, 2017
  12. Ævar Arnfjörð BjarmasonJun 15, 2017
  13. brian m. carlsonJun 16, 2017
  14. Ævar Arnfjörð BjarmasonJun 16, 2017
  15. Johannes SchindelinJun 16, 2017
  16. Adam LangleyJun 16, 2017
  17. Junio C HamanoJun 16, 2017
  18. Junio C HamanoJun 16, 2017
  19. Jonathan NiederJun 16, 2017
  20. Ævar Arnfjörð BjarmasonJun 16, 2017
  21. Jeff KingJun 16, 2017
  22. Johannes SchindelinJun 19, 2017
  23. Mike HommeyJun 15, 2017
  24. Jeff KingJun 16, 2017
  25. Brandon WilliamsJun 15, 2017
  26. Junio C HamanoJun 15, 2017
  27. Jonathan NiederJun 15, 2017
  28. RFC v3: Another proposed hash function transition planJonathan Nieder, Mar 7, 2017
  29. Shawn PearceMar 9, 2017
  30. Jonathan NiederMar 9, 2017
  31. Jeff KingMar 10, 2017
  32. Jonathan NiederMar 10, 2017
  33. technical doc: add a design doc for hash function transitionJonathan Nieder, Sep 28, 2017
  34. Junio C HamanoSep 29, 2017
  35. Junio C HamanoSep 29, 2017
  36. Jonathan NiederSep 29, 2017
  37. Junio C HamanoOct 2, 2017
  38. Jason CooperOct 2, 2017
  39. Junio C HamanoOct 2, 2017
  40. Jason CooperOct 2, 2017
  41. Junio C HamanoOct 3, 2017
  42. Jason CooperOct 3, 2017
  43. Junio C HamanoOct 4, 2017
  44. Junio C HamanoSep 6, 2017
  45. Junio C HamanoSep 8, 2017
  46. Jeff KingSep 8, 2017
  47. Brandon WilliamsSep 11, 2017
  48. Johannes SchindelinSep 13, 2017
  49. demerphqSep 13, 2017
  50. Jonathan NiederSep 13, 2017
  51. Johannes SchindelinSep 14, 2017
  52. Jonathan NiederSep 14, 2017
  53. Johannes SchindelinSep 14, 2017
  54. Linus TorvaldsSep 13, 2017
  55. Johannes SchindelinSep 14, 2017
  56. Gilles Van AsscheSep 18, 2017
  57. Johannes SchindelinSep 18, 2017
  58. Gilles Van AsscheSep 19, 2017
  59. Johannes SchindelinSep 29, 2017
  60. Joan DaemenSep 29, 2017
  61. Johannes SchindelinSep 29, 2017
  62. Joan DaemenSep 30, 2017
  63. Johannes SchindelinOct 2, 2017
  64. Jonathan NiederSep 18, 2017
  65. Jason CooperSep 26, 2017
  66. Johannes SchindelinSep 26, 2017
  67. technical doc: add a design doc for hash function transitionStefan Beller, Sep 26, 2017
  68. Jonathan NiederSep 26, 2017
  69. Jonathan NiederSep 26, 2017
  70. Jason CooperOct 2, 2017
  71. Brandon WilliamsOct 2, 2017
  72. Jason CooperOct 2, 2017
  73. Linus TorvaldsOct 2, 2017
  74. Jeff KingOct 2, 2017
  75. Jonathan NiederSep 13, 2017
  76. Junio C HamanoSep 13, 2017
  77. Stefan BellerSep 13, 2017
  78. Jonathan NiederSep 13, 2017
  79. Junio C HamanoSep 14, 2017
  80. Johannes SchindelinSep 14, 2017
  81. demerphqSep 14, 2017
  82. Johannes SchindelinSep 14, 2017
  83. Junio C HamanoSep 13, 2017
  84. Jonathan NiederSep 13, 2017
  85. Junio C HamanoSep 14, 2017
  86. Johannes SchindelinSep 14, 2017
  87. Brandon WilliamsSep 14, 2017
  88. Jonathan NiederSep 14, 2017
  89. Philip OakleySep 15, 2017
  90. David LangMar 5, 2017
  91. Jonathan NiederMar 6, 2017
  92. Mike HommeyMar 7, 2017
  93. Jeff KingMar 6, 2017
  94. Junio C HamanoMar 6, 2017
  95. Jonathan TanMar 6, 2017
  96. Linus TorvaldsMar 6, 2017
  97. Brandon WilliamsMar 6, 2017
  98. Junio C HamanoMar 6, 2017
  99. Jeff KingMar 7, 2017
  100. Ian JacksonMar 7, 2017
  101. Linus TorvaldsMar 7, 2017
  102. Ian JacksonMar 8, 2017
  103. Johannes SchindelinMar 8, 2017
  104. Johannes SchindelinMar 8, 2017
  105. Use base32?Jason Hennessey, Mar 20, 2017
  106. Michael SteuerMar 20, 2017
  107. Jacob KellerMar 20, 2017
  108. Michael SteuerMar 21, 2017
  109. The Keccak TeamMar 13, 2017
  110. Jonathan NiederMar 13, 2017
  111. ankostisMar 13, 2017
  112. Johannes SchindelinMar 17, 2017
  113. Jeff KingMar 6, 2017

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.