git/list[1] front-page[2] threads[3] people[4] search[5] about
wed 2026-10-07 18:14 UTC

Re: [PATCH v4] technical doc: add a design doc for hash function transition

From
Junio C Hamano <gitster@pobox.com>
Date
Oct 2, 2017, 08:25 UTC
Message-ID
<xmqq3772ot1w.fsf@gitster.mtv.corp.google.com>
In-Reply-To
<20170929173413.GI19555@aiede.mtv.corp.google.com>
Jonathan Nieder <jrnieder@gmail.com> writes:
Show 12 quoted lines
>>> +6. Skip fetching some submodules of a project into a NewHash
>>> +   repository. (This also depends on NewHash support in Git
>>> +   protocol.)
>>
>> It is unclear what this means.  Around submodule support, one thing
>> I can think of is that a NewHash tree in a superproject would record
>> a gitlink that is a NewHash commit object name in it, therefore it
>> cannot refer to an unconverted SHA-1 submodule repository.  But it
>> is unclear if the above description refers to the same issue, or
>> something else.
>
> It refers to that issue.
We may want to find a way to make it clear, then.
Show 8 quoted lines
>> It makes me wonder if we want to add the hashname in this object
>> header.  "length" would be different for non-blob objects anyway,
>> and it is not "compat metadata" we want to avoid baked in, yet it
>> would help diagnose a mistake of attempting to use a "mixed" objects
>> in a single repository.  Not a big issue, though.
>
> Do you mean that adding the hashname into the computation that
> produces the object name would help in some use case?

What I mean is that for SHA-1 objects we keep the object header to be "<type> <length> NUL". For objects in newer world, use the object header to "<type> <hash> <length> NUL", and include the hashname in the object name computation.

Show 10 quoted lines
> For loose objects, it would be nice to name the hash in the file, so
> that "file" can understand what is happening if someone accidentally
> mixes types using "cp".  The only downside is losing the ability to
> copy blobs (which have the same content despite being named using
> different hashes) between repositories after determining their new
> names.  That doesn't seem like a strong downside --- it's pretty
> harmless to include the hash type in loose object files, too.  I think
> I would prefer this to be a "magic number" instead of part of the
> zlib-deflated payload, since this way "file" can discover it more
> easily.
Yeah, thanks for doing pros-and-cons for me ;-)
Show 8 quoted lines
>> If it is a goal to eventually be able to lose SHA-1 compatibility
>> metadata from the objects, then we might want to remove SHA-1 based
>> signature bits (e.g. PGP trailer in signed tag, gpgsig header in the
>> commit object) from NewHash contents, and instead have them stored
>> in a side "metadata" table, only to be used while converting back.
>> I dunno if that is desirable.
>
> I don't consider that desirable.
Agreed.  Let's not go there.
Show 11 quoted lines
>> Hmm, as the corresponding packfile stores object data only in
>> NewHash content format, it is somewhat curious that this table that
>> stores CRC32 of the data appears in the "Tables for each object
>> format" section, as they would be identical, no?  Unless I am
>> grossly misleading the spec, the checksum should either go outside
>> the "Tables for each object format" section but still in .idx, or
>> should be eliminated and become part of the packdata stream instead,
>> perhaps?
>
> It's actually only present for the first object format.  Will find a
> better way to describe this.

I see. One way to do so is to have it upfront before the "after this point, these tables repeat for each of the hashes" part of the file.

Show 11 quoted lines
>> Oy.  So we can go from a short prefix to the pack location by first
>> finding it via binsearch in the short-name table, realize that it is
>> nth object in the object name order, and consulting this table.
>> When we know the pack-order of an object, there is no direct way to
>> go to its location (short of reversing the name-order-to-pack-order
>> table)?
>
> An earlier version of the design also had a pack-order-to-pack-offset
> table, but we weren't able to think of any cases where that would be
> used without also looking up the object name that can be used to
> verify the integrity of the inflated object.

The primary thing I was interested in knowing was if we tried to think of any case where it may be useful and then didn't think of any---I couldn't but I know I am not imaginative enough, and I wanted to know you guys didn't, either.

Previous: Joan DaemenNext: Junio C Hamano
Message 101 of 113 in “RFC: Another proposed hash function transition plan”
  1. Jonathan NiederMar 4, 2017
  2. Linus TorvaldsMar 5, 2017
  3. David LangMar 5, 2017
  4. brian m. carlsonMar 6, 2017
  5. Jeff KingMar 6, 2017
  6. Jeff KingMar 6, 2017
  7. Brandon WilliamsMar 6, 2017
  8. Junio C HamanoMar 6, 2017
  9. Jonathan TanMar 6, 2017
  10. Linus TorvaldsMar 6, 2017
  11. Brandon WilliamsMar 6, 2017
  12. Junio C HamanoMar 6, 2017
  13. Jonathan NiederMar 6, 2017
  14. RFC v3: Another proposed hash function transition planJonathan Nieder, Mar 7, 2017
  15. Mike HommeyMar 7, 2017
  16. Jeff KingMar 7, 2017
  17. Linus TorvaldsMar 7, 2017
  18. Ian JacksonMar 7, 2017
  19. Ian JacksonMar 8, 2017
  20. Johannes SchindelinMar 8, 2017
  21. Johannes SchindelinMar 8, 2017
  22. Shawn PearceMar 9, 2017
  23. Jonathan NiederMar 9, 2017
  24. Jeff KingMar 10, 2017
  25. Jonathan NiederMar 10, 2017
  26. The Keccak TeamMar 13, 2017
  27. Jonathan NiederMar 13, 2017
  28. ankostisMar 13, 2017
  29. Johannes SchindelinMar 17, 2017
  30. Use base32?Jason Hennessey, Mar 20, 2017
  31. Michael SteuerMar 20, 2017
  32. Jacob KellerMar 20, 2017
  33. Michael SteuerMar 21, 2017
  34. Which hash function to use, was Re: RFC: Another proposed hash function transition planJohannes Schindelin, Jun 15, 2017
  35. Mike HommeyJun 15, 2017
  36. Jeff KingJun 15, 2017
  37. Ævar Arnfjörð BjarmasonJun 15, 2017
  38. Brandon WilliamsJun 15, 2017
  39. Jonathan NiederJun 15, 2017
  40. Junio C HamanoJun 15, 2017
  41. Johannes SchindelinJun 15, 2017
  42. Mike HommeyJun 15, 2017
  43. Adam LangleyJun 15, 2017
  44. brian m. carlsonJun 15, 2017
  45. Ævar Arnfjörð BjarmasonJun 15, 2017
  46. brian m. carlsonJun 16, 2017
  47. Jeff KingJun 16, 2017
  48. Ævar Arnfjörð BjarmasonJun 16, 2017
  49. Johannes SchindelinJun 16, 2017
  50. Adam LangleyJun 16, 2017
  51. Jeff KingJun 16, 2017
  52. Junio C HamanoJun 16, 2017
  53. Junio C HamanoJun 16, 2017
  54. Jonathan NiederJun 16, 2017
  55. Ævar Arnfjörð BjarmasonJun 16, 2017
  56. Johannes SchindelinJun 19, 2017
  57. Junio C HamanoSep 6, 2017
  58. Junio C HamanoSep 8, 2017
  59. Jeff KingSep 8, 2017
  60. Brandon WilliamsSep 11, 2017
  61. Johannes SchindelinSep 13, 2017
  62. demerphqSep 13, 2017
  63. Jonathan NiederSep 13, 2017
  64. Junio C HamanoSep 13, 2017
  65. Stefan BellerSep 13, 2017
  66. Junio C HamanoSep 13, 2017
  67. Jonathan NiederSep 13, 2017
  68. Jonathan NiederSep 13, 2017
  69. Jonathan NiederSep 13, 2017
  70. Linus TorvaldsSep 13, 2017
  71. Junio C HamanoSep 14, 2017
  72. Junio C HamanoSep 14, 2017
  73. Johannes SchindelinSep 14, 2017
  74. Johannes SchindelinSep 14, 2017
  75. demerphqSep 14, 2017
  76. Brandon WilliamsSep 14, 2017
  77. Johannes SchindelinSep 14, 2017
  78. Jonathan NiederSep 14, 2017
  79. Johannes SchindelinSep 14, 2017
  80. Jonathan NiederSep 14, 2017
  81. Johannes SchindelinSep 14, 2017
  82. Johannes SchindelinSep 14, 2017
  83. Philip OakleySep 15, 2017
  84. Gilles Van AsscheSep 18, 2017
  85. Johannes SchindelinSep 18, 2017
  86. Jonathan NiederSep 18, 2017
  87. Gilles Van AsscheSep 19, 2017
  88. Jason CooperSep 26, 2017
  89. Johannes SchindelinSep 26, 2017
  90. technical doc: add a design doc for hash function transitionStefan Beller, Sep 26, 2017
  91. Jonathan NiederSep 26, 2017
  92. Jonathan NiederSep 26, 2017
  93. technical doc: add a design doc for hash function transitionJonathan Nieder, Sep 28, 2017
  94. Junio C HamanoSep 29, 2017
  95. Junio C HamanoSep 29, 2017
  96. Johannes SchindelinSep 29, 2017
  97. Joan DaemenSep 29, 2017
  98. Jonathan NiederSep 29, 2017
  99. Johannes SchindelinSep 29, 2017
  100. Joan DaemenSep 30, 2017
  101. Junio C HamanoOct 2, 2017
  102. Junio C HamanoOct 2, 2017
  103. Jason CooperOct 2, 2017
  104. Johannes SchindelinOct 2, 2017
  105. Jason CooperOct 2, 2017
  106. Brandon WilliamsOct 2, 2017
  107. Linus TorvaldsOct 2, 2017
  108. Jason CooperOct 2, 2017
  109. Jeff KingOct 2, 2017
  110. Jason CooperOct 2, 2017
  111. Junio C HamanoOct 3, 2017
  112. Jason CooperOct 3, 2017
  113. Junio C HamanoOct 4, 2017

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.