git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 0/4] faster SHA-1 collision detection

From
Scott Chacon <schacon@gmail.com>
Date
Oct 7, 2026, 18:13 UTC
Message-ID
<CAP2yMaL51H1OAG25nQ0NuLQLb0wevd4CicCG1_ezsJfrZDqfUA@mail.gmail.com>
In-Reply-To
<xmqq5wzda0h6.fsf@gitster.g>
Hey,
On Wed, Oct 7, 2026 at 7:23 PM Junio C Hamano <gitster@pobox.com> wrote:
Show 10 quoted lines
> > This series ports the approach of Sam Reis's sha1dc Rust crate [1],
> > which gitoxide recently switched to [2], to C.
>
> Which means license-wise the original is compatible with us, I
> presume, as they are "Apache2 or MIT, your choice".
>
> How can you/we be sure, with respect to the current AI policy in
> SubmittingPatches (which by the way was vetted by SFC lawyers), that
> your "AI generated" code did not "borrow" from places that gets
> you/us into trouble?

It's a good question. I actually just submitted a proposed update to that policy based on SFC's updated guidelines, but either way, I learned about this from Sam and have talked to him about the port and he seemed excited about it. I can triple check, but I'm fairly confident that he's fine with this and I am fine signing off on it under the terms of the DCO language.

Of course, he in turn used AI tooling to produce _his_ library, but within the guidelines of the updated SFC guidelines. Johannes's alternative series is the original Rust code of Sam that my agent looked at to produce this (in addition to his blog post explaining it), so I'm not sure how that might be materially different.

Show 19 quoted lines
> > The end result hashes roughly 2.7x faster on the Xeon and 2.85x faster
> > on the M5 Max. Single-threaded index-pack of git.git goes from 24.3s to
> > 12.7s on the Xeon, and from 16.1s to 8.7s on the M5 Max.
> >
> > Hashing throughput on the Xeon, in MiB/s:
> >
> >                                 16KiB    1MiB   vs OpenSSL
> >   OpenSSL SHA-1 (no detection)   1234    1129      1.00x
> >   sha1dc/ (today)                 435     450      2.67x
> >   shani+avx2 (default here)      1002     901      1.24x
> >   shani+sse2                     1075    1008      1.13x
> >   portable+avx2                   553     654      1.96x
> >   portable+sse2                   603     681      1.84x
> >   portable                        466     565      2.29x
> >
> > In other words, currently collision detection costs about 1.5–2.5x on
> > top of the hashing itself today, but only about 0.2x with the series.
>
> Thanks for these numbers.

It would have been better had I provided the same relative scale (it should be 1.5-2.5x vs 1.2x, but whatever, you probably get it. It's 20% overhead here vs 50%-150% overhead previously).

> > [1] https://sam.dev/blog/faster-sha1-collision-detection
> > [2] https://github.com/GitoxideLabs/gitoxide/pull/3008
>
> And the pointers to the original sources.

CC'ing Sam (sha1dc rust guy) and Sebastian (Gitoxide) on this, just in case they have an opinion but I'm pretty sure they would be more than happy for this to be integrated.

Scott
Previous: Junio C HamanoNext: Sebastian Thiel
Message 11 of 22 in “faster SHA-1 collision detection”
  1. 0/4 faster SHA-1 collision detectionScott Chacon, Sep 29, 2026
  2. 1/4 sha1dc-accel: add a block loop for sha1dc's SHA1_CTXScott Chacon, Sep 29, 2026
  3. Johannes SchindelinOct 7, 2026
  4. 2/4 sha1dc-accel: vectorize the unavoidable-bitconditions checkScott Chacon, Sep 29, 2026
  5. Johannes SchindelinOct 7, 2026
  6. Junio C HamanoOct 7, 2026
  7. 3/4 sha1dc-accel: compress with SHA-NI on x86-64Scott Chacon, Sep 29, 2026
  8. 4/4 sha1dc-accel: compress with the ARMv8 SHA-1 instructionsScott Chacon, Sep 29, 2026
  9. Johannes SchindelinOct 7, 2026
  10. Junio C HamanoOct 7, 2026
  11. Scott ChaconOct 7, 2026
  12. Sebastian ThielOct 8, 2026
  13. Sam ReisOct 8, 2026
  14. D. Ben KnobleOct 8, 2026
  15. Sam ReisOct 8, 2026
  16. Junio C HamanoOct 8, 2026
  17. D. Ben KnobleOct 8, 2026
  18. Junio C HamanoOct 8, 2026
  19. Todd ZullingerOct 9, 2026
  20. D. Ben KnobleOct 10, 2026
  21. Junio C HamanoOct 9, 2026
  22. Junio C HamanoOct 8, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.