git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Transition plan for git to move to a new hash function

From
IJIan Jackson <ijackson@chiark.greenend.org.uk>
Date
Mar 5, 2017, 13:45 UTC
Message-ID
<22716.5770.95842.704242@chiark.greenend.org.uk>
In-Reply-To
<20170304224936.rqqtkdvfjgyezsht@genre.crustytoothpaste.net>
brian m. carlson writes ("Re: Transition plan for git to move to a new hash function"):
> Instead, I was referring to areas like the notes code.  It has extensive
> use of the last byte as a type of lookup table key.  It's very dependent
> on having exactly one hash, since it will always want to use the last
> byte.

You mean note_tree_search ? (My tree here may be a bit out of date.) This doesn't seem difficult to fix. The nontrivial changes would be mostly confined to SUBTREE_SHA1_PREFIXCMP and GET_NIBBLE.

It's true that like most of git there's a lot of hardcoded `sha1'.

Are you arguing in favour of "replace git with git2 by simply s/20/64/g; s/sha1/blake/g" ? This seems to me to be a poor idea. Takeup of the new `git2' would be very slow because of the pain involved.

Any sensible method of moving to a new hash that isn't "make a completely incompatible new version of git" is going to involve teaching the code we have in git right now to handle new hashes as well as sha1 hashes.

Even if the plan is to try to convert old data, rather than keep it and be able to refer to it from new data, something will have to be able to parse old packfiles, old commits, old tags, old notes, etc. etc. etc. Either that's going to be some separate conversion utility, or it has to be the same code in git that's there already.[1]

The ability to handle both old-format and new-format data can be achieved in the code by doing away with the hardcoded sha1s, so that instead the hash is an abstract data type with operations like "initialise", "compare", "get a nybble", etc. We've already seen patches going in this direction.

[1] I've heard suggestions here that instead we should expect users to "git1 fast-export", which you would presumably feed into "git2 fast-import". But what is `git1' here ? Is it the current git codebase frozen in time ? I don't think it can be. With this conversion strategy, we will need to maintain git1 for decades. It will need portability fixes, security fixes, fixes for new hostile compiler optimisations, and so on. The difficulty of conversion means there will be pressure to backport new features from `git2' to `git1'. (Also this approach means that all signatures are definitively lost during the conversion process.)

So if we want to provide both `git1' and `git2', it's still better to compile `git' and `git2' from the same codebase. But if we do that, the resulting ifdeffery and/or other hash abstractions are most of the work to be hash-agile. It's just the difference between a compile-time and runtime switch.

I think the incompatibile approach is much more work in the medium and long term - and it leads to a longer transition period.

Bear in mind that our objective is not to minimise the time until the new version of git is available. Our objective is to minimise the time until (most) people are using it. An approach which takes longer for the git community to develop, but which is easier to deploy, can easily be better.

Or maybe the objective is to minimise overall effort. In which case more work on git, for an easier transition for all the users, seems like a no-brainer. I think this is arguably true even from the point of view of effort amongst the community of git contributors. git contributors start out as git users - and if git's users are all busy struggling with a difficult transition, they will have less time to improve other stuff and will tend less to get involved upstream. (And they may be less inclined to feel that the git upstream developers understand their needs well.)

The better alternative is to adopt a plan that has a clear and straightforward transition for users, and ask git users to help with implementation.

I think many git users, including sophisticated users and competent organisations, are concerned about sha1. Currently most of those users will find it difficult to help, because it's not clear to them what needs to be done.

Thanks, Ian.

Previous: brian m. carlsonNext: brian m. carlson
Message 57 of 136 in “SHA1 collisions found”
  1. Joey HessFeb 23, 2017
  2. Junio C HamanoFeb 23, 2017
  3. Junio C HamanoFeb 23, 2017
  4. David LangFeb 23, 2017
  5. Jakub NarębskiFeb 23, 2017
  6. Jeff KingFeb 23, 2017
  7. Joey HessFeb 23, 2017
  8. Linus TorvaldsFeb 23, 2017
  9. Joey HessFeb 23, 2017
  10. Linus TorvaldsFeb 23, 2017
  11. Jeff KingFeb 23, 2017
  12. Linus TorvaldsFeb 23, 2017
  13. Jeff KingFeb 23, 2017
  14. Linus TorvaldsFeb 23, 2017
  15. Jeff KingFeb 23, 2017
  16. Øyvind A. HolmFeb 23, 2017
  17. Joey HessFeb 23, 2017
  18. Joey HessFeb 23, 2017
  19. Morten WelinderFeb 23, 2017
  20. Geert UytterhoevenFeb 24, 2017
  21. Jeff KingFeb 23, 2017
  22. David LangFeb 23, 2017
  23. David LangFeb 23, 2017
  24. David LangFeb 23, 2017
  25. Linus TorvaldsFeb 23, 2017
  26. Linus TorvaldsFeb 23, 2017
  27. Joey HessFeb 23, 2017
  28. Linus TorvaldsFeb 23, 2017
  29. Junio C HamanoFeb 23, 2017
  30. Duy NguyenFeb 24, 2017
  31. brian m. carlsonFeb 25, 2017
  32. René ScharfeFeb 27, 2017
  33. brian m. carlsonFeb 28, 2017
  34. Ian JacksonFeb 24, 2017
  35. ankostisFeb 24, 2017
  36. Junio C HamanoFeb 24, 2017
  37. David LangFeb 24, 2017
  38. Junio C HamanoFeb 24, 2017
  39. Stefan BellerFeb 24, 2017
  40. Junio C HamanoFeb 24, 2017
  41. Junio C HamanoFeb 24, 2017
  42. ankostisFeb 24, 2017
  43. Junio C HamanoFeb 24, 2017
  44. ankostisFeb 25, 2017
  45. Jason CooperFeb 26, 2017
  46. brian m. carlsonFeb 26, 2017
  47. Linus TorvaldsFeb 26, 2017
  48. Ævar Arnfjörð BjarmasonFeb 26, 2017
  49. Jeff KingFeb 26, 2017
  50. Transition plan for git to move to a new hash functionIan Jackson, Feb 27, 2017
  51. Markus TrippelsdorfFeb 27, 2017
  52. Ian JacksonFeb 27, 2017
  53. Tony FinchFeb 27, 2017
  54. brian m. carlsonFeb 28, 2017
  55. Ian JacksonMar 2, 2017
  56. brian m. carlsonMar 4, 2017
  57. Ian JacksonMar 5, 2017
  58. brian m. carlsonMar 5, 2017
  59. Philip OakleyFeb 24, 2017
  60. Jeff KingFeb 24, 2017
  61. Linus TorvaldsFeb 25, 2017
  62. Linus TorvaldsFeb 25, 2017
  63. Jeff KingFeb 25, 2017
  64. Junio C HamanoFeb 26, 2017
  65. Junio C HamanoFeb 25, 2017
  66. Jason CooperFeb 26, 2017
  67. Jeff KingFeb 26, 2017
  68. brian m. carlsonFeb 26, 2017
  69. Brandon WilliamsMar 2, 2017
  70. Jeff KingMar 3, 2017
  71. Ian JacksonMar 3, 2017
  72. Jeff KingMar 3, 2017
  73. Linus TorvaldsMar 2, 2017
  74. Junio C HamanoMar 2, 2017
  75. Linus TorvaldsMar 2, 2017
  76. Joey HessMar 2, 2017
  77. Linus TorvaldsMar 2, 2017
  78. Mike HommeyMar 3, 2017
  79. Linus TorvaldsMar 3, 2017
  80. Jeff KingMar 3, 2017
  81. Stefan BellerMar 3, 2017
  82. David LangFeb 25, 2017
  83. Stefan BellerFeb 25, 2017
  84. Jeff KingFeb 25, 2017
  85. David LangFeb 25, 2017
  86. Jeff KingFeb 25, 2017
  87. David LangFeb 25, 2017
  88. Jacob KellerFeb 25, 2017
  89. Jacob KellerFeb 25, 2017
  90. grarpampFeb 25, 2017
  91. Ian JacksonFeb 24, 2017
  92. Ian JacksonFeb 25, 2017
  93. brian m. carlsonFeb 25, 2017
  94. Jeff KingFeb 25, 2017
  95. Mike HommeyFeb 25, 2017
  96. brian m. carlsonFeb 26, 2017
  97. Jason CooperFeb 24, 2017
  98. ankostisFeb 25, 2017
  99. Jakub NarębskiFeb 24, 2017
  100. Santiago TorresFeb 24, 2017
  101. Jakub NarębskiFeb 24, 2017
  102. Øyvind A. HolmFeb 24, 2017
  103. Jeff KingFeb 24, 2017
  104. Jakub NarębskiFeb 24, 2017
  105. Lars SchneiderFeb 25, 2017
  106. Jeff KingFeb 26, 2017
  107. Junio C HamanoFeb 26, 2017
  108. Thomas BraunFeb 26, 2017
  109. Jeff KingFeb 26, 2017
  110. Geert UytterhoevenFeb 27, 2017
  111. Jeff KingFeb 27, 2017
  112. Morten WelinderFeb 27, 2017
  113. Jeff KingFeb 23, 2017
  114. Linus TorvaldsFeb 23, 2017
  115. Jeff KingFeb 23, 2017
  116. 1/3 add collision-detecting sha1 implementationJeff King, Feb 23, 2017
  117. Stefan BellerFeb 23, 2017
  118. Jeff KingFeb 24, 2017
  119. Linus TorvaldsFeb 24, 2017
  120. Jeff KingFeb 24, 2017
  121. 2/3 sha1dc: adjust header includes for gitJeff King, Feb 23, 2017
  122. 3/3 Makefile: add USE_SHA1DC knobJeff King, Feb 23, 2017
  123. HW42Feb 24, 2017
  124. Jeff KingFeb 24, 2017
  125. Linus TorvaldsFeb 23, 2017
  126. Junio C HamanoFeb 28, 2017
  127. Junio C HamanoFeb 28, 2017
  128. Jeff KingFeb 28, 2017
  129. Dan ShumowMar 1, 2017
  130. Linus TorvaldsFeb 28, 2017
  131. Shawn PearceFeb 28, 2017
  132. Linus TorvaldsFeb 28, 2017
  133. Dan ShumowFeb 28, 2017
  134. Marc StevensFeb 28, 2017
  135. Linus TorvaldsFeb 28, 2017
  136. Jeff KingMar 1, 2017

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.