git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 0/7] block-sha1: improved SHA1 hashing

From
ASArtur Skawina <art.08.09@gmail.com>
Date
Aug 6, 2009, 23:19 UTC
Message-ID
<4A7B64F1.2000309@gmail.com>
In-Reply-To
<alpine.LFD.2.01.0908061531310.3390@localhost.localdomain>
Linus Torvalds wrote:
Show 11 quoted lines
> 
> Yeah, verified. Google for
> 
> 	northwood "barrel shifter"
> 
> and you'll find a lot of it.
> 
> Basically, older P4's will I think shift one bit at a time. So while even 
> Prescott is relatively weak in the shifter department, pre-prescott 
> (Willamette and Northwood) are _really_ weak. If your P4 is one of those, 
> you really shouldn't use it to decide on optimizations.

Actually that's even more of a reason to make sure the code doesn't suck :) The difference on less perverse cpus will usually be small, but on P4 it can be huge.

A few years back I found my old ip checksum microbenchmark, and when I ran it on a P4 (prescott iirc) i didn't believe my eyes. The straightforward 32-bit C implementation was running circles around the in-kernel one... And a few tweaks to the assembler version got me another ~100% speedup.[1]

After that the P4 became the very first cpu to test any code on... :)
artur
[1] just reran the benchmark on this p4; true on northwood too:
IACCK 0.9.30  Artur Skawina <...>
[ exec time; lower is better  ] [speed ] [ time ]  [ok?]
TIME-N+S TIME32 TIME33 TIME1480 MBYTES/S TIMEXXXX  CSUM FUNCTION ( rdtsc_overhead=0  null=0 )
   17901    510    557     3010   393.36    59772  56dd csum_partial_cdumb16
    3019    154    156      431  2747.10    43106  56dd csum_partial_c32
    2413    170    177      328  3609.76    37501  56dd csum_partial_c32l
    2437    170    170      328  3609.76    37488  56dd csum_partial_c32i
    5078    205    254      767  1543.68    48117  56dd csum_partial_std
    5612    299    291      851  1391.30    53673  56dd csum_partial_686
    1584     99    127      227  5215.86    14495  56dd csum_partial_586f
    1738    107    121      229  5170.31    14785  56dd csum_partial_586fs
    4893    175    171      759  1559.95    52347  56dd csum_partial_copy_generic_std
    4949    151    189      756  1566.14    67847  56dd csum_partial_copy_generic_686
    2072    110    134      302  3920.53    39061  56dd csum_partial_copy_generic_p4as1
Previous: Linus TorvaldsNext: Linus Torvalds
Message 21 of 37 in “block-sha1: improved SHA1 hashing”
  1. 0/7 block-sha1: improved SHA1 hashingLinus Torvalds, Aug 6, 2009
  2. 1/7 block-sha1: add new optimized C 'block-sha1' routinesLinus Torvalds, Aug 6, 2009
  3. 2/7 block-sha1: try to use rol/ror appropriatelyLinus Torvalds, Aug 6, 2009
  4. 3/7 block-sha1: make the 'ntohl()' part of the first SHA1 loopLinus Torvalds, Aug 6, 2009
  5. 4/7 block-sha1: re-use the temporary array as we calculate the SHA1Linus Torvalds, Aug 6, 2009
  6. 5/7 block-sha1: macroize the rounds a bit furtherLinus Torvalds, Aug 6, 2009
  7. 6/7 block-sha1: Use '(B&C)+(D&(B^C))' instead of '(B&C)|(D&(B|C))' in round 3Linus Torvalds, Aug 6, 2009
  8. 7/7 block-sha1: get rid of redundant 'lenW' contextLinus Torvalds, Aug 6, 2009
  9. Bert WesargAug 6, 2009
  10. Artur SkawinaAug 6, 2009
  11. Linus TorvaldsAug 6, 2009
  12. Artur SkawinaAug 6, 2009
  13. Linus TorvaldsAug 6, 2009
  14. Artur SkawinaAug 6, 2009
  15. Linus TorvaldsAug 6, 2009
  16. Linus TorvaldsAug 6, 2009
  17. Artur SkawinaAug 6, 2009
  18. Artur SkawinaAug 6, 2009
  19. Linus TorvaldsAug 6, 2009
  20. Linus TorvaldsAug 6, 2009
  21. Artur SkawinaAug 6, 2009
  22. Linus TorvaldsAug 6, 2009
  23. Artur SkawinaAug 6, 2009
  24. Linus TorvaldsAug 6, 2009
  25. Linus TorvaldsAug 6, 2009
  26. Linus TorvaldsAug 7, 2009
  27. Artur SkawinaAug 7, 2009
  28. Linus TorvaldsAug 7, 2009
  29. Artur SkawinaAug 7, 2009
  30. Linus TorvaldsAug 7, 2009
  31. Artur SkawinaAug 7, 2009
  32. Linus TorvaldsAug 8, 2009
  33. Artur SkawinaAug 8, 2009
  34. Linus TorvaldsAug 8, 2009
  35. Artur SkawinaAug 8, 2009
  36. Artur SkawinaAug 8, 2009
  37. Artur SkawinaAug 8, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.