{"thread":{"id":"20241","subject":"Re: Performance issue of 'git branch'","startedAt":"2009-07-26T23:21:54Z","lastAt":"2009-08-18T21:50:14Z","messageCount":60,"participants":["George Spelvin","Erik Faye-Lund","Michael J Gruber","Johannes Schindelin","Carlos R. Mafra","Brian Ristuccia","Jakub Narebski","Peter Harris","Jonathan del Strother","Mark Lodato","Linus Torvalds","Jon Smirl","Dmitry Potapov","Junio C Hamano","Nicolas Pitre","Artur Skawina","Andy Polyakov"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"118850","messageId":"20090726232154.29594.qmail@science.horizon.com","threadId":"20241","inReplyTo":null,"subject":"Re: Performance issue of 'git branch'","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-07-26T23:21:54Z","receivedAt":"2009-07-26T23:21:54Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> It's a bit sad, since the _only_ thing we load all of libcrypto for is the \n> (fairly trivial) SHA1 code. \n>\n> But at the same time, last time I benchmarked the different SHA1 \n> libraries, the openssl one was the fastest. I think it has tuned assembly \n> language for most architectures. Our regular mozilla-based C code is \n> perfectly fine, but it doesn't hold a candle to assembler tuning.\n\nActually, openssl only has assembly for x86, x86_64, and ia64.\nTruthfully, once you have 32 registers, SHA1 is awfully easy to\ncompile near-optimally.\n\nGit currently includes some hand-tuned assembly that isn't in OpenSSL:\n- ARM (only 16 registers, and the rotate+op support can be used nicely)\n- PPC (3-way superscalar *without* OO execution benefits from careful\n  scheduling)\n\nFurther, all of the core hand-tuned SHA1 assembly code in OpenSSL is by\nAndy Polyakov and is dual-licensed GPL/3-clause BSD *in addition to*\nthe OpenSSL license.  So we can just import it:\n\nSee http://www.openssl.org/~appro/cryptogams/\nand http://www.openssl.org/~appro/cryptogams/cryptogams-0.tar.gz\n\n(Ooh, look, he has PPC code in there, too.  Does anyone with a PPC machine\nwant to compare it with Git's?)\n\nIt'll take some massaging because that's just the core SHA1_Transform\nfunction and not the wrappers, but it's hardly a heroic effort.\n\nI'm pretty deep in the weeds at $DAY_JOB and can't get to it for a while,\nbut would a patch be appreciated?\n"},{"id":"119249","messageId":"20090731104602.15375.qmail@science.horizon.com","threadId":"20241","inReplyTo":"20090726232154.29594.qmail@science.horizon.com","subject":"Request for benchmarking: x86 SHA1 code","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-07-31T10:46:02Z","receivedAt":"2009-07-31T10:46:02Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"After studying Andy Polyakov's optimized x86 SHA-1 in OpenSSL, I've\ngot a version that's 1.6% slower on a P4 and 15% faster on a Phenom.\nI'm curious about the performance on other CPUs I don't have access to,\nparticularly Core 2 duo and i7.\n\nCould someone do some benchmarking for me?  Old (486/Pentium/P2/P3)\nmachines are also interesting, but I'm optimizing for newer ones.\n\nI haven't packaged this nicely, but it's not that complicated.\n- Download Andy's original code from\n  http://www.openssl.org/~appro/cryptogams/cryptogams-0.tar.gz\n- Unpack and cd to the cryptogams-0/x86 directory\n- \"patch < this_email\" to create \"sha1test.c\", \"sha1-586.h\", \"Makefile\",\n   and \"sha1-x86.pl\".\n- \"make\" \n- Run ./586test (before) and ./x86test (after) and note the timings.\n\nThank you!\n\n--- /dev/null\t2009-05-12 02:55:38.579106460 -0400\n+++ Makefile\t2009-07-31 06:22:42.000000000 -0400\n@@ -0,0 +1,16 @@\n+CC := gcc\n+CFLAGS := -m32 -W -Wall -Os -g\n+ASFLAGS := --32\n+\n+all: 586test x86test\n+\n+586test : sha1test.c sha1-586.o\n+\t$(CC) $(CFLAGS) -o $@ sha1test.c sha1-586.o\n+\n+x86test : sha1test.c sha1-x86.o\n+\t$(CC) $(CFLAGS) -o $@ sha1test.c sha1-x86.o\n+\n+586test x86test : sha1-586.h\n+\n+%.s : %.pl x86asm.pl x86unix.pl\n+\tperl $< elf > $@\n--- /dev/null\t2009-05-12 02:55:38.579106460 -0400\n+++ sha1test.c\t2009-07-28 09:24:09.000000000 -0400\n@@ -0,0 +1,67 @@\n+#include <stdint.h>\n+#include <stdlib.h>\n+#include <stdio.h>\n+#include <sys/time.h>\n+\n+#include \"sha1-586.h\"\n+\n+#define SIZE 1000000\n+\n+#if SIZE % 64\n+# error SIZE must be a multiple of 64!\n+#endif\n+\n+int\n+main(void)\n+{\n+\tuint32_t iv[5] = {\n+\t\t0x67452301, 0xefcdab89, 0x98badcfe, 0x10325476, 0xc3d2e1f0\n+\t};\n+\t/* Simplest known-answer test, \"abc\" */\n+\tstatic uint8_t const testbuf[64] = {\n+\t\t'a','b','c', 0x80, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n+\t\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n+\t\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n+\t\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 24\n+\t};\n+\t/* Expected: A9993E364706816ABA3E25717850C26C9CD0D89D */\n+\tstatic uint32_t const expected[5] = {\n+\t\t0xa9993e36, 0x4706816a, 0xba3e2571, 0x7850c26c, 0x9cd0d89d };\n+\tunsigned i;\n+\tchar *p = malloc(SIZE);\n+\tstruct timeval tv0, tv1;\n+\n+\tif (!p) {\n+\t\tperror(\"malloc\");\n+\t\treturn 1;\n+\t}\n+\n+\tsha1_block_data_order(iv, testbuf, 1);\n+\tprintf(\"Result:  %08x %08x %08x %08x %08x\\n\"\n+\t       \"Expected:%08x %08x %08x %08x %08x\\n\",\n+\t\tiv[0], iv[1], iv[2], iv[3], iv[4], expected[0],\n+\t\texpected[1], expected[2], expected[3], expected[4]);\n+\tfor (i = 0; i < 5; i++)\n+\t\tif (iv[i] != expected[i])\n+\t\t\tprintf(\"MISMATCH in word %u\\n\", i);\n+\n+\tif (gettimeofday(&tv0, NULL) < 0) {\n+\t\tperror(\"gettimeofday\");\n+\t\treturn 1;\n+\t}\n+\tfor (i = 0; i < 500; i++)\n+\t\tsha1_block_data_order(iv, p, SIZE/64);\n+\tif (gettimeofday(&tv1, NULL) < 0) {\n+\t\tperror(\"gettimeofday\");\n+\t\treturn 1;\n+\t}\n+\ttv1.tv_sec -= tv0.tv_sec;\n+\ttv1.tv_usec -= tv0.tv_usec;\n+\tif (tv1.tv_usec < 0) {\n+\t\ttv1.tv_sec--;\n+\t\ttv1.tv_usec += 1000000;\n+\t}\n+\tprintf(\"%u bytes: %u.%06u s\\n\", i * SIZE, (unsigned)tv1.tv_sec,\n+\t\t(unsigned)tv1.tv_usec);\n+\treturn 0;\n+}\n--- /dev/null\t2009-05-12 02:55:38.579106460 -0400\n+++ sha1-586.h\t2009-07-27 09:34:03.000000000 -0400\n@@ -0,0 +1,3 @@\n+#include <stdint.h>\n+\n+void sha1_block_data_order(uint32_t iv[5], void const *in, unsigned len);\n--- /dev/null\t2009-05-12 02:55:38.579106460 -0400\n+++ sha1-x86.pl\t2009-07-31 06:10:18.000000000 -0400\n@@ -0,0 +1,398 @@\n+#!/usr/bin/env perl\n+\n+# ====================================================================\n+# [Re]written by Andy Polyakov <appro@fy.chalmers.se> for the OpenSSL\n+# project. The module is, however, dual licensed under OpenSSL and\n+# CRYPTOGAMS licenses depending on where you obtain it. For further\n+# details see http://www.openssl.org/~appro/cryptogams/.\n+# ====================================================================\n+\n+# \"[Re]written\" was achieved in two major overhauls. In 2004 BODY_*\n+# functions were re-implemented to address P4 performance issue [see\n+# commentary below], and in 2006 the rest was rewritten in order to\n+# gain freedom to liberate licensing terms.\n+\n+# It was noted that Intel IA-32 C compiler generates code which\n+# performs ~30% *faster* on P4 CPU than original *hand-coded*\n+# SHA1 assembler implementation. To address this problem (and\n+# prove that humans are still better than machines:-), the\n+# original code was overhauled, which resulted in following\n+# performance changes:\n+#\n+#\t\tcompared with original\tcompared with Intel cc\n+#\t\tassembler impl.\t\tgenerated code\n+# Pentium\t-16%\t\t\t+48%\n+# PIII/AMD\t+8%\t\t\t+16%\n+# P4\t\t+85%(!)\t\t\t+45%\n+#\n+# As you can see Pentium came out as looser:-( Yet I reckoned that\n+# improvement on P4 outweights the loss and incorporate this\n+# re-tuned code to 0.9.7 and later.\n+# ----------------------------------------------------------------\n+#\t\t\t\t\t<appro@fy.chalmers.se>\n+\n+$0 =~ m/(.*[\\/\\\\])[^\\/\\\\]+$/; $dir=$1;\n+push(@INC,\"${dir}\",\"${dir}../../perlasm\");\n+require \"x86asm.pl\";\n+\n+&asm_init($ARGV[0],\"sha1-586.pl\",$ARGV[$#ARGV] eq \"386\");\n+\n+$A=\"eax\";\n+$B=\"ebx\";\n+$C=\"ecx\";\n+$D=\"edx\";\n+$E=\"ebp\";\n+\n+# Two temporaries\n+$S=\"esi\";\n+$T=\"edi\";\n+\n+# The round constants\n+use constant K1 => 0x5a827999;\n+use constant K2 => 0x6ED9EBA1;\n+use constant K3 => 0x8F1BBCDC;\n+use constant K4 => 0xCA62C1D6;\n+\n+@V=($A,$B,$C,$D,$E);\n+\n+# Given unlimited registers and functional units, it would be\n+# possible to compute SHA-1 at two cycles per round, using 7\n+# operations per round.  Remember, each round computes a new\n+# value for E, which is used as A in the following round and B\n+# in the round after that.  There are two critical paths:\n+# - A must be rotated and added to the output E\n+# - B must go through two boolean operations before being added\n+#   to the result E.  Since this latter addition can't be done\n+#   in the same-cycle as the critical addition of a<<<5, this is\n+#   a total of 2+1+1 = 4 cycles.\n+# Additionally, if you want to avoid copying B, you have to\n+# rotate it soon after use in this round so it is ready for use\n+# as the following round's C.\n+\n+# f = (a <<< 5) + e + K + in[i] + (d^(b&(c^d)))\t\t(0..19)\n+# f = (a <<< 5) + e + K + in[i] + (b^c^d)\t\t(20..39, 60..79)\n+# f = (a <<< 5) + e + K + in[i] + (c&d) + (b&(c^d))\t(40..59)\n+# The hard part is doing this with only two temporary registers.\n+# Let's divide it into 4 parts.  These can be executed in a 7-cycle\n+# loop, assuming triple (quadruple counting xor separately) issue:\n+#\n+#\tin[i]\t\tF1(c,d)\t\tF2(b,c,d)\ta<<<5\n+#\tmov in[i],T\t\t\t(addl S,A)\t(movl B,S)\n+#\txor in[i+1],T\t\t\t\t\t(rorl 5,S)\n+#\txor in[i+2],T\tmovl D,S\t\t\t(addl S,A)\n+#\txor in[i+3],T\tandl C,S\n+#\trotl 1,T\taddl S,E\tmovl D,S\n+#\t(movl T,in[i])                  xorl C,S\n+#\taddl T+K,E\t\t\tandl B,S\trorl 2,B\t//\n+#\t(mov in[i],T)\t\t\taddl S,E\tmovl A,S\n+#\t(xor in[i+1],T)\t\t\t\t\trorl 5,S\n+#\t(xor in[i+2],T)\t(movl C,S)\t\t\taddl S,E\n+#\n+# (The last 3 rounds can omit the store of T.)\n+# The \"addl T+K,E\" line is actually implemented using the lea instruction,\n+# which (on a Pentium) requires that neither T not K was modified on the\n+# previous cycle.\n+#\n+# The other two rounds are a bit simpler, and can therefore be \"pulled in\"\n+# one cycle, to 6.  The bit-select function (0..19):\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\tmov in[i],T\t(xorl E,S)\t\t\t(addl T+K,A)\n+#\txor in[i+1],T\t(addl S,A)\t(movl A,S)\n+#\txor in[i+2],T\t(rorl 2,C)\t(roll 5,S)\n+#\txor in[i+3],T\tmovl D,S\t(addl S,A)\n+#\troll 1,T\txorl C,S\n+#\tmovl T,in[i]\tandl B,S\t\t\trorl 2,B\n+#\taddl T+K,E\txorl D,S\t\t//\t(mov in[i],T)\n+#\t(xor in[i+1],T)\taddl S,E\tmovl A,S\n+#\t(xor in[i+2],T)\t\t\troll 5,S\n+#\t(xor in[i+3],T)\t(movl C,S)\taddl S,E\n+#\n+# And the XOR function (also 6, limited by the in[i] forming) used in\n+# rounds 20..39 and 60..79:\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\tmov in[i],T\t(xorl C,S)\t\t\t(addl T+K,A)\n+#\txor in[i+1],T\t(addl S,A)\t(movl A,S)\n+#\txor in[i+2],T\t\t\t(roll 5,S)\t\n+#\txor in[i+3],T\t\t\t(addl S,A)\n+#\troll 1,T\tmovl D,S\n+#\tmovl T,in[i]\txorl B,S\t\t\trorl 2,B\n+#\taddl T+K,E\txorl C,S\t\t//\t(mov in[i],T)\n+#\t(xor in[i+1],T)\taddl S,E\tmovl A,S\n+#\t(xor in[i+2],T)\t\t\troll 5,S\n+#\t(xor in[i+3],T)\t\t\taddl S,E\n+#\n+# The first 16 rounds don't need to form the in[i] equation, letting\n+# us pull it in another 2 cycles, to 4, after some reassignment of\n+# temporaries:\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\t\t\tmovl D,S\t(roll 5,T)\t(addl S,A)\n+#\tmov in[i],T\txorl C,S\t(addl T,A)\n+#\t\t\tandl B,S\t\t\trorl 2,B\n+#\taddl T+K,E\txorl D,S\tmovl A,T\n+#\t\t\taddl S,E\troll 5,T\t(movl C,S)\n+#\t(mov in[i],T)\t(xorl B,S)\taddl T,E\n+#\n+\n+# The transition between rounds 15 and 16 will be a bit tricky... the best\n+# thing to do is to delay the computation of a<<<5 one cycle and move it back\n+# to the S register.  That way, T is free as early as possible.\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\t(addl T+K,A)\t(xorl E,S)\t(movl A,T)\n+#\t\t\tmovl D,S\t(roll 5,T)\t(addl S,A)\n+#\tmov in[i],T\txorl C,S\t(addl T,A)\n+#\t\t\tandl B,S\t\t\trorl 2,B\n+#\taddl T+K,E\txorl D,S\t\t\t(movl in[1],T)\n+#\t(xor in[i+1],T)\taddl S,E\tmovl A,S\n+#\t(xor in[i+2],T)\trorl 2,B\troll 5,S\n+#\t(xor in[i+3],T)\t(movl C,S)\taddl S,E\n+\n+\n+\n+\n+\n+# This expects the first copy of D to S to have been done already\n+#\t\t\tmovl D,S\t(roll 5,T)\t(addl S,A)\t//\n+#\tmov in[i],T\txorl C,S\t(addl T,A)\n+#\t\t\tandl B,S\t\t\trorl 2,B\n+#\taddl T+K,E\txorl D,S\tmovl A,T\n+#\t\t\taddl S,E\troll 5,T\t(movl C,S)\t//\n+#\t(mov in[i],T)\t(xorl B,S)\taddl T,E\t\n+\n+sub BODY_00_15\n+{\n+\tlocal($n,$a,$b,$c,$d,$e)=@_;\n+\n+\t&comment(\"00_15 $n\");\n+\t\t&mov($S,$d) if ($n == 0);\n+\t&mov($T,&swtmp($n%16));\t\t# Load Xi.\n+\t\t&xor($S,$c);\t\t# Continue computing F() = d^(b&(c^d))\n+\t\t&and($S,$b);\n+\t\t\t&rotr($b,2);\n+\t&lea($e,&DWP(K1,$e,$T));\t# Add Xi and K\n+    if ($n < 15) {\n+\t\t\t&mov($T,$a);\n+\t\t&xor($S,$d);\n+\t\t\t&rotl($T,5);\n+\t\t&add($e,$S);\n+\t\t&mov($S,$c);\t\t# Start of NEXT round's F()\n+\t\t\t&add($e,$T);\n+    } else {\n+\t# This version provides the correct start for BODY_20_39\n+\t&mov($T,&swtmp(($n+1)%16));\t# Start computing mext Xi.\n+\t\t&xor($S,$d);\n+\t&xor($T,&swtmp(($n+3)%16));\n+\t\t&add($e,$S);\t\t# Add F()\n+\t\t\t&mov($S,$a);\t# Start computing a<<<5\n+\t&xor($T,&swtmp(($n+9)%16));\n+\t\t\t&rotl($S,5);\n+    }\n+\n+}\n+\n+# The transition between rounds 15 and 16 will be a bit tricky... the best\n+# thing to do is to delay the computation of a<<<5 one cycle and move it back\n+# to the S register.  That way, T is free as early as possible.\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\t(addl T+K,A)\t(xorl E,S)\t(movl B,T)\n+#\t\t\tmovl D,S\t(roll 5,T)\t(addl S,A)\t//\n+#\tmov in[i],T\txorl C,S\t(addl T,A)\n+#\t\t\tandl B,S\t\t\trorl 2,B\n+#\taddl T+K,E \txorl D,S\t\t\t(movl in[1],T)\n+#\t(xor in[i+1],T)\taddl S,E\tmovl A,S\n+#\t(xor in[i+2],T)\trorl 2,B\troll 5,S\t//\n+#\t(xor in[i+3],T)\t(movl C,S)\taddl S,E\n+\n+\n+# This starts just before starting to compute F(); the Xi should have XORed\n+# the first three values together.  (Break is at //)\n+#\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\tmov in[i],T\t(xorl E,S)\t\t\t(addl T+K,A)\n+#\txor in[i+1],T\t(addl S,A)\t(movl B,S)\n+#\txor in[i+2],T\t\t\t(roll 5,S)\t//\n+#\txor in[i+3],T \tmovl D,S  \t(addl S,A)\n+#\troll 1,T\txorl C,S\n+#\tmovl T,in[i]\tandl B,S\t\t\trorl 2,B\n+#\taddl T+K,E \txorl D,S\t\t\t(mov in[i],T)\n+#\t(xor in[i+1],T)\taddl S,E\tmovl A,S\n+#\t(xor in[i+2],T)\t\t\troll 5,S\t//\n+#\t(xor in[i+3],T)\t(movl C,S)\taddl S,E\n+\n+sub BODY_16_19\n+{\n+\tlocal($n,$a,$b,$c,$d,$e)=@_;\n+\n+\t&comment(\"16_20 $n\");\n+\n+\t&xor($T,&swtmp(($n+13)%16));\n+\t\t\t&add($a,$S);\t# End of previous round\n+\t\t&mov($S,$d);\t\t# Start current round's F()\n+\t&rotl($T,1);\n+\t\t&xor($S,$c);\n+\t&mov(&swtmp($n%16),$T);\t\t# Store computed Xi.\n+\t\t&and($S,$b);\n+\t\t&rotr($b,2);\n+\t&lea($e,&DWP(K1,$e,$T));\t# Add Xi and K\n+\t&mov($T,&swtmp(($n+1)%16));\t# Start computing mext Xi.\n+\t\t&xor($S,$d);\n+\t&xor($T,&swtmp(($n+3)%16));\n+\t\t&add($e,$S);\t\t# Add F()\n+\t\t\t&mov($S,$a);\t# Start computing a<<<5\n+\t&xor($T,&swtmp(($n+9)%16));\n+\t\t\t&rotl($S,5);\n+}\n+\n+# This is just like BODY_16_19, but computes a different F() = b^c^d\n+#\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\tmov in[i],T\t(xorl E,S)\t\t\t(addl T+K,A)\n+#\txor in[i+1],T\t(addl S,A)\t(movl B,S)\n+#\txor in[i+2],T\t\t\t(roll 5,S)\t//\n+#\txor in[i+3],T \t\t\t(addl S,A)\n+#\troll 1,T\tmovl C,S\n+#\tmovl T,in[i]\txorl B,S\t\t\trorl 2,B\n+#\taddl T+K,E \txorl C,S\t\t\t(mov in[i],T)\n+#\t(xor in[i+1],T)\taddl S,E\tmovl A,S\n+#\t(xor in[i+2],T)\t\t\troll 5,S\t//\n+#\t(xor in[i+3],T)\t(movl C,S)\taddl S,E\n+\n+sub BODY_20_39\t# And 61..79\n+{\n+\tlocal($n,$a,$b,$c,$d,$e)=@_;\n+\tlocal $K=($n<40) ? K2 : K4;\n+\n+\t&comment(\"21_30 $n\");\n+\n+\t&xor($T,&swtmp(($n+13)%16));\n+\t\t\t&add($a,$S);\t# End of previous round\n+\t\t&mov($S,$d)\n+\t&rotl($T,1);\n+\t\t&mov($S,$d);\t\t# Start current round's F()\n+\t&mov(&swtmp($n%16),$T) if ($n < 77);\t# Store computed Xi.\n+\t\t&xor($S,$b);\n+\t\t&rotr($b,2);\n+\t&lea($e,&DWP($K,$e,$T));\t# Add Xi and K\n+\t&mov($T,&swtmp(($n+1)%16)) if ($n < 79); # Start computing next Xi.\n+\t\t&xor($S,$c);\n+\t&xor($T,&swtmp(($n+3)%16)) if ($n < 79);\n+\t\t&add($e,$S);\t\t# Add F1()\n+\t\t\t&mov($S,$a);\t# Start computing a<<<5\n+\t&xor($T,&swtmp(($n+9)%16)) if ($n < 79);\n+\t\t\t&rotl($S,5);\n+\n+\t\t\t&add($e,$S) if ($n == 79);\n+}\n+\n+\n+# This starts immediately after the LEA, and expects to need to finish\n+# the previous round. (break is at //)\n+#\n+#\tin[i]\t\tF1(c,d)\t\tF2(b,c,d)\ta<<<5\n+#\t(addl T+K,E)\t\t\t(andl C,S)\t(rorl 2,C)\n+#\tmov in[i],T\t\t\t(addl S,A)\t(movl B,S)\n+#\txor in[i+1],T\t\t\t\t\t(rorl 5,S)\n+#\txor in[i+2],T /\tmovl D,S\t\t\t(addl S,A)\n+#\txor in[i+3],T\tandl C,S\n+#\trotl 1,T\taddl S,E\tmovl D,S\n+#\t(movl T,in[i])                  xorl C,S\n+#\taddl T+K,E\t\t\tandl B,S\trorl 2,B\n+#\t(mov in[i],T)\t\t\taddl S,E\tmovl A,S\n+#\t(xor in[i+1],T)\t\t\t\t\trorl 5,S\n+#\t(xor in[i+2],T)\t// (movl C,S)\t\t\taddl S,E\n+\n+sub BODY_40_59\n+{\n+\tlocal($n,$a,$b,$c,$d,$e)=@_;\n+\n+\t&comment(\"41_59 $n\");\n+\n+\t\t\t&add($a,$S);\t# End of previous round\n+\t\t&mov($S,$d);\t\t# Start current round's F(1)\n+\t&xor($T,&swtmp(($n+13)%16));\n+\t\t&and($S,$c);\n+\t&rotl($T,1);\n+\t\t&add($e,$S);\t\t# Add F1()\n+\t\t&mov($S,$d);\t\t# Start current round's F2()\n+\t&mov(&swtmp($n%16),$T);\t\t# Store computed Xi.\n+\t\t&xor($S,$c);\n+\t&lea($e,&DWP(K3,$e,$T));\n+\t\t&and($S,$b);\n+\t\t&rotr($b,2);\n+\t&mov($T,&swtmp(($n+1)%16));\t# Start computing next Xi.\n+\t\t&add($e,$S);\t\t# Add F2()\n+\t\t&mov($S,$a);\t\t# Start computing a<<<5\n+\t&xor($T,&swtmp(($n+3)%16));\n+\t\t\t&rotl($S,5);\n+\t&xor($T,&swtmp(($n+9)%16));\n+}\n+\n+&function_begin(\"sha1_block_data_order\",16);\n+\t&mov($S,&wparam(0));\t# SHA_CTX *c\n+\t&mov($T,&wparam(1));\t# const void *input\n+\t&mov($A,&wparam(2));\t# size_t num\n+\t&stack_push(16);\t# allocate X[16]\n+\t&shl($A,6);\n+\t&add($A,$T);\n+\t&mov(&wparam(2),$A);\t# pointer beyond the end of input\n+\t&mov($E,&DWP(16,$S));# pre-load E\n+\n+\t&set_label(\"loop\",16);\n+\n+\t# copy input chunk to X, but reversing byte order!\n+\tfor ($i=0; $i<16; $i+=4)\n+\t\t{\n+\t\t&mov($A,&DWP(4*($i+0),$T));\n+\t\t&mov($B,&DWP(4*($i+1),$T));\n+\t\t&mov($C,&DWP(4*($i+2),$T));\n+\t\t&mov($D,&DWP(4*($i+3),$T));\n+\t\t&bswap($A);\n+\t\t&bswap($B);\n+\t\t&bswap($C);\n+\t\t&bswap($D);\n+\t\t&mov(&swtmp($i+0),$A);\n+\t\t&mov(&swtmp($i+1),$B);\n+\t\t&mov(&swtmp($i+2),$C);\n+\t\t&mov(&swtmp($i+3),$D);\n+\t\t}\n+\t&mov(&wparam(1),$T);\t# redundant in 1st spin\n+\n+\t&mov($A,&DWP(0,$S));\t# load SHA_CTX\n+\t&mov($B,&DWP(4,$S));\n+\t&mov($C,&DWP(8,$S));\n+\t&mov($D,&DWP(12,$S));\n+\t# E is pre-loaded\n+\n+\tfor($i=0;$i<16;$i++)\t{ &BODY_00_15($i,@V); unshift(@V,pop(@V)); }\n+\tfor(;$i<16;$i++)\t{ &BODY_15($i,@V);    unshift(@V,pop(@V)); }\n+\tfor(;$i<20;$i++)\t{ &BODY_16_19($i,@V); unshift(@V,pop(@V)); }\n+\tfor(;$i<40;$i++)\t{ &BODY_20_39($i,@V); unshift(@V,pop(@V)); }\n+\tfor(;$i<60;$i++)\t{ &BODY_40_59($i,@V); unshift(@V,pop(@V)); }\n+\tfor(;$i<80;$i++)\t{ &BODY_20_39($i,@V); unshift(@V,pop(@V)); }\n+\n+\t(($V[4] eq $E) and ($V[0] eq $A)) or die;\t# double-check\n+\n+\t&comment(\"Loop trailer\");\n+\n+\t&mov($S,&wparam(0));\t# re-load SHA_CTX*\n+\t&mov($T,&wparam(1));\t# re-load input pointer\n+\n+\t&add($A,&DWP(0,$S));\t# E is last \"A\"...\n+\t&add($B,&DWP(4,$S));\n+\t&add($C,&DWP(8,$S));\n+\t&add($D,&DWP(12,$S));\n+\t&add($E,&DWP(16,$S));\n+\n+\t&mov(&DWP(0,$S),$A);\t# update SHA_CTX\n+\t &add($T,64);\t\t# advance input pointer\n+\t&mov(&DWP(4,$S),$B);\n+\t &cmp($T,&wparam(2));\t# have we reached the end yet?\n+\t&mov(&DWP(8,$S),$C);\n+\t&mov(&DWP(12,$S),$D);\n+\t&mov(&DWP(16,$S),$E);\n+\t&jb(&label(\"loop\"));\n+\n+\t&stack_pop(16);\n+&function_end(\"sha1_block_data_order\");\n+&asciz(\"SHA1 block transform for x86, CRYPTOGAMS by <appro\\@openssl.org>\");\n+\n+&asm_finish();\n"},{"id":"119251","messageId":"40aa078e0907310411k54dc58fbq9a938c489df56b68@mail.gmail.com","threadId":"20241","inReplyTo":"20090731104602.15375.qmail@science.horizon.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@googlemail.com","sentAt":"2009-07-31T11:11:36Z","receivedAt":"2009-07-31T11:11:36Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Fri, Jul 31, 2009 at 12:46 PM, George Spelvin<linux@horizon.com> wrote:\n> I haven't packaged this nicely, but it's not that complicated.\n> - Download Andy's original code from\n>  http://www.openssl.org/~appro/cryptogams/cryptogams-0.tar.gz\n> - Unpack and cd to the cryptogams-0/x86 directory\n> - \"patch < this_email\" to create \"sha1test.c\", \"sha1-586.h\", \"Makefile\",\n>   and \"sha1-x86.pl\".\n> - \"make\"\n\n$ patch < ../sha1-opt.patch.eml\npatching file `Makefile'\npatching file `sha1test.c'\npatching file `sha1-586.h'\npatching file `sha1-x86.pl'\n\n$ make\nmake: *** No rule to make target `sha1-586.o', needed by `586test'.  Stop.\n\nWhat did I do wrong? :)\nWould it be easier if you pushed it out somewhere?\n\n-- \nErik \"kusma\" Faye-Lund\nkusmabite@gmail.com\n(+47) 986 59 656\n"},{"id":"119253","messageId":"4A72D3D7.1080706@drmicha.warpmail.net","threadId":"20241","inReplyTo":"20090731104602.15375.qmail@science.horizon.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2009-07-31T11:21:59Z","receivedAt":"2009-07-31T11:21:59Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"George Spelvin venit, vidit, dixit 31.07.2009 12:46:\n> After studying Andy Polyakov's optimized x86 SHA-1 in OpenSSL, I've\n> got a version that's 1.6% slower on a P4 and 15% faster on a Phenom.\n> I'm curious about the performance on other CPUs I don't have access to,\n> particularly Core 2 duo and i7.\n> \n> Could someone do some benchmarking for me?  Old (486/Pentium/P2/P3)\n> machines are also interesting, but I'm optimizing for newer ones.\n> \n> I haven't packaged this nicely, but it's not that complicated.\n> - Download Andy's original code from\n>   http://www.openssl.org/~appro/cryptogams/cryptogams-0.tar.gz\n> - Unpack and cd to the cryptogams-0/x86 directory\n> - \"patch < this_email\" to create \"sha1test.c\", \"sha1-586.h\", \"Makefile\",\n>    and \"sha1-x86.pl\".\n> - \"make\" \n> - Run ./586test (before) and ./x86test (after) and note the timings.\n> \n> Thank you!\n\nBest of 6 runs:\n./586test\nResult:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\nExpected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n500000000 bytes: 1.642336 s\n./x86test\nResult:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\nExpected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n500000000 bytes: 1.532153 s\n\nSystem:\nuname -a\nLinux localhost.localdomain 2.6.29.6-213.fc11.x86_64 #1 SMP Tue Jul 7\n21:02:57 EDT 2009 x86_64 x86_64 x86_64 GNU/Linux\n\ncat /proc/cpuinfo\nprocessor       : 0\nvendor_id       : GenuineIntel\ncpu family      : 6\nmodel           : 15\nmodel name      : Intel(R) Core(TM)2 Duo CPU     T7500  @ 2.20GHz\nstepping        : 11\ncpu MHz         : 800.000\ncache size      : 4096 KB\nphysical id     : 0\nsiblings        : 2\ncore id         : 0\ncpu cores       : 2\napicid          : 0\ninitial apicid  : 0\nfpu             : yes\nfpu_exception   : yes\ncpuid level     : 10\nwp              : yes\nflags           : fpu vme de pse tsc msr pae mce cx8 apic mtrr pge mca\ncmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall\nnx lm constant_tsc arch_perfmon pebs bts rep_good pni dtes64 monitor\nds_cpl vmx est tm2 ssse3 cx16 xtpr pdcm lahf_lm ida tpr_shadow vnmi\nflexpriority\nbogomips        : 4389.20\nclflush size    : 64\ncache_alignment : 64\naddress sizes   : 36 bits physical, 48 bits virtual\npower management:\n\nprocessor       : 1\nvendor_id       : GenuineIntel\ncpu family      : 6\nmodel           : 15\nmodel name      : Intel(R) Core(TM)2 Duo CPU     T7500  @ 2.20GHz\nstepping        : 11\ncpu MHz         : 800.000\ncache size      : 4096 KB\nphysical id     : 0\nsiblings        : 2\ncore id         : 1\ncpu cores       : 2\napicid          : 1\ninitial apicid  : 1\nfpu             : yes\nfpu_exception   : yes\ncpuid level     : 10\nwp              : yes\nflags           : fpu vme de pse tsc msr pae mce cx8 apic mtrr pge mca\ncmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall\nnx lm constant_tsc arch_perfmon pebs bts rep_good pni dtes64 monitor\nds_cpl vmx est tm2 ssse3 cx16 xtpr pdcm lahf_lm ida tpr_shadow vnmi\nflexpriority\nbogomips        : 4388.78\nclflush size    : 64\ncache_alignment : 64\naddress sizes   : 36 bits physical, 48 bits virtual\npower management:\n"},{"id":"119254","messageId":"4A72D4F6.7050707@drmicha.warpmail.net","threadId":"20241","inReplyTo":"20090731104602.15375.qmail@science.horizon.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2009-07-31T11:26:46Z","receivedAt":"2009-07-31T11:26:46Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"George Spelvin venit, vidit, dixit 31.07.2009 12:46:\n> After studying Andy Polyakov's optimized x86 SHA-1 in OpenSSL, I've\n> got a version that's 1.6% slower on a P4 and 15% faster on a Phenom.\n> I'm curious about the performance on other CPUs I don't have access to,\n> particularly Core 2 duo and i7.\n> \n> Could someone do some benchmarking for me?  Old (486/Pentium/P2/P3)\n> machines are also interesting, but I'm optimizing for newer ones.\n> \n> I haven't packaged this nicely, but it's not that complicated.\n> - Download Andy's original code from\n>   http://www.openssl.org/~appro/cryptogams/cryptogams-0.tar.gz\n> - Unpack and cd to the cryptogams-0/x86 directory\n> - \"patch < this_email\" to create \"sha1test.c\", \"sha1-586.h\", \"Makefile\",\n>    and \"sha1-x86.pl\".\n> - \"make\" \n> - Run ./586test (before) and ./x86test (after) and note the timings.\n> \n> Thank you!\n\nBest of 6 runs:\n./586test\nResult:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\nExpected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n500000000 bytes: 1.258031 s\n./x86test\nResult:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\nExpected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n500000000 bytes: 1.171770 s\n\nSystem:\nuname -a\nLinux whatever 2.6.22-14-generic #1 SMP Tue Feb 12 07:42:25 UTC 2008\ni686 GNU/Linux\ncat /proc/cpuinfo\nprocessor       : 0\nvendor_id       : GenuineIntel\ncpu family      : 6\nmodel           : 23\nmodel name      : Intel(R) Core(TM)2 Duo CPU     E8400  @ 3.00GHz\nstepping        : 10\ncpu MHz         : 2000.000\ncache size      : 6144 KB\nphysical id     : 0\nsiblings        : 2\ncore id         : 0\ncpu cores       : 2\nfdiv_bug        : no\nhlt_bug         : no\nf00f_bug        : no\ncoma_bug        : no\nfpu             : yes\nfpu_exception   : yes\ncpuid level     : 13\nwp              : yes\nflags           : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge\nmca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe nx lm\nconstant_tsc pni monitor ds_cpl vmx smx est tm2 ssse3 cx16 xtpr lahf_lm\nbogomips        : 5988.92\nclflush size    : 64\n\nprocessor       : 1\nvendor_id       : GenuineIntel\ncpu family      : 6\nmodel           : 23\nmodel name      : Intel(R) Core(TM)2 Duo CPU     E8400  @ 3.00GHz\nstepping        : 10\ncpu MHz         : 2000.000\ncache size      : 6144 KB\nphysical id     : 0\nsiblings        : 2\ncore id         : 1\ncpu cores       : 2\nfdiv_bug        : no\nhlt_bug         : no\nf00f_bug        : no\ncoma_bug        : no\nfpu             : yes\nfpu_exception   : yes\ncpuid level     : 13\nwp              : yes\nflags           : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge\nmca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe nx lm\nconstant_tsc pni monitor ds_cpl vmx smx est tm2 ssse3 cx16 xtpr lahf_lm\nbogomips        : 5984.92\nclflush size    : 64\n"},{"id":"119256","messageId":"20090731113100.17584.qmail@science.horizon.com","threadId":"20241","inReplyTo":"40aa078e0907310411k54dc58fbq9a938c489df56b68@mail.gmail.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-07-31T11:31:00Z","receivedAt":"2009-07-31T11:31:00Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> $ make\n> make: *** No rule to make target `sha1-586.o', needed by `586test'.  Stop.\n>\n> What did I do wrong? :)\n> Would it be easier if you pushed it out somewhere?\n\nH'm.... It *should* do\n\nperl sha1-586.pl elf > sha1-586.s\nas --32  -o sha1-586.o sha1-586.s\ngcc -m32 -W -Wall -Os -g -o 586test sha1test.c sha1-586.o\n(And likewise for the \"x86test\" binary.)\n\nwhich is what happened when I tested it.  Obviously I have something\nnon-portable in the Makefile.\n\nYou could try \"make sha1-586.s\" and \"sha1-586.o\" and see which rule is\nf*ed up.\n\nUm, you *are* in a directory which contains a sha1-586.pl file, right?\n\n\nThanks!\n"},{"id":"119257","messageId":"4A72D76D.3050400@drmicha.warpmail.net","threadId":"20241","inReplyTo":"40aa078e0907310411k54dc58fbq9a938c489df56b68@mail.gmail.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2009-07-31T11:37:17Z","receivedAt":"2009-07-31T11:37:17Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Erik Faye-Lund venit, vidit, dixit 31.07.2009 13:11:\n> On Fri, Jul 31, 2009 at 12:46 PM, George Spelvin<linux@horizon.com> wrote:\n>> I haven't packaged this nicely, but it's not that complicated.\n>> - Download Andy's original code from\n>>  http://www.openssl.org/~appro/cryptogams/cryptogams-0.tar.gz\n>> - Unpack and cd to the cryptogams-0/x86 directory\n>> - \"patch < this_email\" to create \"sha1test.c\", \"sha1-586.h\", \"Makefile\",\n>>   and \"sha1-x86.pl\".\n>> - \"make\"\n> \n> $ patch < ../sha1-opt.patch.eml\n> patching file `Makefile'\n> patching file `sha1test.c'\n> patching file `sha1-586.h'\n> patching file `sha1-x86.pl'\n> \n> $ make\n> make: *** No rule to make target `sha1-586.o', needed by `586test'.  Stop.\n> \n> What did I do wrong? :)\n> Would it be easier if you pushed it out somewhere?\n> \n\nYou need to go to the x86 directory, apply the patch and run make there.\n(I made the same mistake.) Also, you i586 (32bit) glibc-devel if you're\non a 64 bit system.\n\nMichael\n"},{"id":"119259","messageId":"40aa078e0907310524x1fe4d84dr858ebc03731ee093@mail.gmail.com","threadId":"20241","inReplyTo":"4A72D76D.3050400@drmicha.warpmail.net","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@googlemail.com","sentAt":"2009-07-31T12:24:28Z","receivedAt":"2009-07-31T12:24:28Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Fri, Jul 31, 2009 at 1:37 PM, Michael J\nGruber<git@drmicha.warpmail.net> wrote:\n>> What did I do wrong? :)\n>\n> You need to go to the x86 directory, apply the patch and run make there.\n> (I made the same mistake.) Also, you i586 (32bit) glibc-devel if you're\n> on a 64 bit system.\n\nAha, thanks :)\n\nNow I'm getting a different error:\n$ make\nas   -o sha1-586.o sha1-586.s\nsha1-586.s: Assembler messages:\nsha1-586.s:4: Warning: .type pseudo-op used outside of .def/.endef ignored.\nsha1-586.s:4: Error: junk at end of line, first unrecognized character is `s'\nsha1-586.s:1438: Warning: .size pseudo-op used outside of .def/.endef ignored.\nsha1-586.s:1438: Error: junk at end of line, first unrecognized character is `s'\n\nmake: *** [sha1-586.o] Error 1\n\nWhat might be relevant, is that I'm trying this on Windows (Vista\n64bit). I'd still think GNU as should be able to assemble the source,\nthough. I've got an i7, so I thought the result might be interresting.\n\n-- \nErik \"kusma\" Faye-Lund\nkusmabite@gmail.com\n(+47) 986 59 656\n"},{"id":"119261","messageId":"alpine.DEB.1.00.0907311428340.4503@intel-tinevez-2-302","threadId":"20241","inReplyTo":"40aa078e0907310524x1fe4d84dr858ebc03731ee093@mail.gmail.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-07-31T12:29:10Z","receivedAt":"2009-07-31T12:29:10Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 31 Jul 2009, Erik Faye-Lund wrote:\n\n> On Fri, Jul 31, 2009 at 1:37 PM, Michael J\n> Gruber<git@drmicha.warpmail.net> wrote:\n> >> What did I do wrong? :)\n> >\n> > You need to go to the x86 directory, apply the patch and run make there.\n> > (I made the same mistake.) Also, you i586 (32bit) glibc-devel if you're\n> > on a 64 bit system.\n> \n> Aha, thanks :)\n> \n> Now I'm getting a different error:\n> $ make\n> as   -o sha1-586.o sha1-586.s\n> sha1-586.s: Assembler messages:\n> sha1-586.s:4: Warning: .type pseudo-op used outside of .def/.endef ignored.\n> sha1-586.s:4: Error: junk at end of line, first unrecognized character is `s'\n> sha1-586.s:1438: Warning: .size pseudo-op used outside of .def/.endef ignored.\n> sha1-586.s:1438: Error: junk at end of line, first unrecognized character is `s'\n> \n> make: *** [sha1-586.o] Error 1\n> \n> What might be relevant, is that I'm trying this on Windows (Vista\n> 64bit).\n\nProbably using msysGit?  Then you're still using the 32-bit environment, \nas MSys is 32-bit only for now.\n\nCiao,\nDscho\n"},{"id":"119263","messageId":"20090731123121.GA4306@Pilar.aei.mpg.de","threadId":"20241","inReplyTo":"20090731104602.15375.qmail@science.horizon.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Carlos R. Mafra","fromEmail":"crmafra2@gmail.com","sentAt":"2009-07-31T12:31:21Z","receivedAt":"2009-07-31T12:31:21Z","isPatch":false,"sender":{"key":"crmafra2@gmail.com","avatar":null},"body":"On Fri 31.Jul'09 at  6:46:02 -0400, George Spelvin wrote:\n> - Run ./586test (before) and ./x86test (after) and note the timings.\n\nFor 8 runs in a Intel(R) Core(TM)2 Duo CPU T7250 @ 2.00GHz,\n\nbefore: 1.75 +- 0.02\nafter: 1.62 +- 0.02\n"},{"id":"119262","messageId":"20090731123203.6160.qmail@science.horizon.com","threadId":"20241","inReplyTo":"40aa078e0907310524x1fe4d84dr858ebc03731ee093@mail.gmail.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-07-31T12:32:03Z","receivedAt":"2009-07-31T12:32:03Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> Now I'm getting a different error:\n> $ make\n> as   -o sha1-586.o sha1-586.s\n> sha1-586.s: Assembler messages:\n> sha1-586.s:4: Warning: .type pseudo-op used outside of .def/.endef ignored.\n> sha1-586.s:4: Error: junk at end of line, first unrecognized character is `s'\n> sha1-586.s:1438: Warning: .size pseudo-op used outside of .def/.endef ignored.\n> sha1-586.s:1438: Error: junk at end of line, first unrecognized character is `s'\n> \n> make: *** [sha1-586.o] Error 1\n\n> What might be relevant, is that I'm trying this on Windows (Vista\n> 64bit). I'd still think GNU as should be able to assemble the source,\n> though. I've got an i7, so I thought the result might be interresting.\n\nAh... what assembler?  the perl proprocessor supports multiple\nassemblers:\n\telf     - Linux, FreeBSD, Solaris x86, etc.\n\ta.out   - DJGPP, elder OpenBSD, etc.\n\tcoff    - GAS/COFF such as Win32 targets\n\twin32n  - Windows 95/Windows NT NASM format\n\tnw-nasm - NetWare NASM format\n\tnw-mwasm- NetWare Metrowerks Assembler\n\nMaybe you need to replace \"elf\" with \"coff\"?\n"},{"id":"119265","messageId":"40aa078e0907310545k64f058b6ja936f583bdaeb120@mail.gmail.com","threadId":"20241","inReplyTo":"20090731123203.6160.qmail@science.horizon.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@googlemail.com","sentAt":"2009-07-31T12:45:15Z","receivedAt":"2009-07-31T12:45:15Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Fri, Jul 31, 2009 at 2:32 PM, George Spelvin<linux@horizon.com> wrote:\n> Maybe you need to replace \"elf\" with \"coff\"?\n\nThat did the trick, thanks!\n\nBest of 6 runs on an Intel Core i7 920 @ 2.67GHz:\n\nbefore (586test): 1.415\nafter (x86test): 1.470\n\n-- \nErik \"kusma\" Faye-Lund\nkusmabite@gmail.com\n(+47) 986 59 656\n"},{"id":"119267","messageId":"20090731130249.27188.qmail@science.horizon.com","threadId":"20241","inReplyTo":"40aa078e0907310545k64f058b6ja936f583bdaeb120@mail.gmail.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-07-31T13:02:49Z","receivedAt":"2009-07-31T13:02:49Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> That did the trick, thanks!\n>\n> Best of 6 runs on an Intel Core i7 920 @ 2.67GHz:\n> \n> before (586test): 1.415\n> after (x86test): 1.470\n\nSo it's slower.  Bummer. :-(  Obviously I have some work to do.\n\nBut thank you very much for the result!\n"},{"id":"119268","messageId":"20090731132702.GY12813@osiris.978.org","threadId":"20241","inReplyTo":"20090731104602.15375.qmail@science.horizon.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Brian Ristuccia","fromEmail":"brian@ristuccia.com","sentAt":"2009-07-31T13:27:02Z","receivedAt":"2009-07-31T13:27:02Z","isPatch":false,"sender":{"key":"brian@ristuccia.com","avatar":null},"body":"The revised code is faster on Intel Atom N270 by around 15% (results below\ntypical of several runs):\n\n$ ./586test ; ./x86test \nResult:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\nExpected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n500000000 bytes: 4.981760 s\nResult:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\nExpected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n500000000 bytes: 4.323324 s\n\n$ cat /proc/cpuinfo\n\nprocessor\t: 0\nvendor_id\t: GenuineIntel\ncpu family\t: 6\nmodel\t\t: 28\nmodel name\t: Intel(R) Atom(TM) CPU N270   @ 1.60GHz\nstepping\t: 2\ncpu MHz\t\t: 800.000\ncache size\t: 512 KB\nphysical id\t: 0\nsiblings\t: 2\ncore id\t\t: 0\ncpu cores\t: 1\napicid\t\t: 0\ninitial apicid\t: 0\nfdiv_bug\t: no\nhlt_bug\t\t: no\nf00f_bug\t: no\ncoma_bug\t: no\nfpu\t\t: yes\nfpu_exception\t: yes\ncpuid level\t: 10\nwp\t\t: yes\nflags\t\t: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca\ncmov pat clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe nx constant_tsc\narch_perfmon pebs bts pni dtes64 monitor ds_cpl est tm2 ssse3 xtpr pdcm\nlahf_lm\nbogomips\t: 3191.59\nclflush size\t: 64\npower management:\n\nprocessor\t: 1\nvendor_id\t: GenuineIntel\ncpu family\t: 6\nmodel\t\t: 28\nmodel name\t: Intel(R) Atom(TM) CPU N270   @ 1.60GHz\nstepping\t: 2\ncpu MHz\t\t: 800.000\ncache size\t: 512 KB\nphysical id\t: 0\nsiblings\t: 2\ncore id\t\t: 0\ncpu cores\t: 1\napicid\t\t: 1\ninitial apicid\t: 1\nfdiv_bug\t: no\nhlt_bug\t\t: no\nf00f_bug\t: no\ncoma_bug\t: no\nfpu\t\t: yes\nfpu_exception\t: yes\ncpuid level\t: 10\nwp\t\t: yes\nflags\t\t: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca\ncmov pat clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe nx constant_tsc\narch_perfmon pebs bts pni dtes64 monitor ds_cpl est tm2 ssse3 xtpr pdcm\nlahf_lm\nbogomips\t: 3191.91\nclflush size\t: 64\npower management:\n\n-- \nBrian Ristuccia\nbrian@ristuccia.com\n"},{"id":"119269","messageId":"m3d47hrniw.fsf@localhost.localdomain","threadId":"20241","inReplyTo":"20090731104602.15375.qmail@science.horizon.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2009-07-31T13:27:07Z","receivedAt":"2009-07-31T13:27:07Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"George Spelvin\" <linux@horizon.com> writes:\n> After studying Andy Polyakov's optimized x86 SHA-1 in OpenSSL, I've\n> got a version that's 1.6% slower on a P4 and 15% faster on a Phenom.\n> I'm curious about the performance on other CPUs I don't have access to,\n> particularly Core 2 duo and i7.\n> \n> Could someone do some benchmarking for me?  Old (486/Pentium/P2/P3)\n> machines are also interesting, but I'm optimizing for newer ones.\n\n----------\n$ [time] ./586test \nResult:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\nExpected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n500000000 bytes: 5.376325 s\n\nreal    0m5.384s\nuser    0m5.108s\nsys     0m0.008s\n\n500000000 bytes: 5.367261 s\n\n5.09user 0.00system 0:05.38elapsed 94%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+378minor)pagefaults 0swaps\n\n----------\n$ [time] ./x86test \nResult:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\nExpected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n500000000 bytes: 5.312238 s\n\nreal    0m5.325s\nuser    0m5.060s\nsys     0m0.008s\n\n500000000 bytes: 5.323783 s\n\n5.06user 0.00system 0:05.34elapsed 94%CPU (0avgtext+0avgdata 0maxresident)k\n0inputs+0outputs (0major+377minor)pagefaults 0swaps\n\n==========\nSystem:\n$ uname -a\nLinux roke 2.6.14-11.1.aur.2 #1 Tue Jan 31 16:05:05 CET 2006 \\\n i686 athlon i386 GNU/Linux\n\n$ cat /proc/cpuinfo\nprocessor\t: 0\nvendor_id\t: AuthenticAMD\ncpu family\t: 6\nmodel\t\t: 4\nmodel name\t: AMD Athlon(tm) processor\nstepping\t: 2\ncpu MHz\t\t: 1000.188\ncache size\t: 256 KB\nfdiv_bug\t: no\nhlt_bug\t\t: no\nf00f_bug\t: no\ncoma_bug\t: no\nfpu\t\t: yes\nfpu_exception\t: yes\ncpuid level\t: 1\nwp\t\t: yes\nflags\t\t: fpu vme de pse tsc msr pae mce cx8 mtrr pge \\\n mca cmov pat pse36 mmx fxsr syscall mmxext 3dnowext 3dnow\nbogomips\t: 2002.43\n\n$ free\n             total       used       free     shared    buffers     cached\nMem:        515616     495812      19804          0       6004     103160\n-/+ buffers/cache:     386648     128968\nSwap:      1052248     279544     772704\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"119273","messageId":"20090731140532.18327.qmail@science.horizon.com","threadId":"20241","inReplyTo":"20090731132702.GY12813@osiris.978.org","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-07-31T14:05:32Z","receivedAt":"2009-07-31T14:05:32Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> The revised code is faster on Intel Atom N270 by around 15% (results below\n> typical of several runs):\n> \n> $ ./586test ; ./x86test \n> Result:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n> Expected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n> 500000000 bytes: 4.981760 s\n> Result:  a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n> Expected:a9993e36 4706816a ba3e2571 7850c26c 9cd0d89d\n> 500000000 bytes: 4.323324 s\n\nCool, thanks!  I hadn't optimized it at all for Atom's\nin-order pipe, so I'm pleasantly surprised.\n"},{"id":"119275","messageId":"eaa105840907310805x452f6b09tf602672359bd1aac@mail.gmail.com","threadId":"20241","inReplyTo":"20090731104602.15375.qmail@science.horizon.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Peter Harris","fromEmail":"git@peter.is-a-geek.org","sentAt":"2009-07-31T15:05:09Z","receivedAt":"2009-07-31T15:05:09Z","isPatch":false,"sender":{"key":"git@peter.is-a-geek.org","avatar":null},"body":"On Fri, Jul 31, 2009 at 6:46 AM, George Spelvin wrote:\n> Could someone do some benchmarking for me?  Old (486/Pentium/P2/P3)\n> machines are also interesting, but I'm optimizing for newer ones.\n\nThe new code appears to be marginally faster on a Pentium 3 Xeon:\n\nBest of five runs:\n586test: 11.658661 s\nx86test: 11.209024 s\n\n$ cat /proc/cpuinfo\nprocessor       : 0\nvendor_id       : GenuineIntel\ncpu family      : 6\nmodel           : 7\nmodel name      : Pentium III (Katmai)\nstepping        : 3\ncpu MHz         : 547.630\ncache size      : 1024 KB\nfdiv_bug        : no\nhlt_bug         : no\nf00f_bug        : no\ncoma_bug        : no\nfpu             : yes\nfpu_exception   : yes\ncpuid level     : 2\nwp              : yes\nflags           : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge\nmca cmov pat pse36 mmx fxsr sse\nbogomips        : 1097.12\nclflush size    : 32\npower management:\n\nprocessor       : 1\nvendor_id       : GenuineIntel\ncpu family      : 6\nmodel           : 7\nmodel name      : Pentium III (Katmai)\nstepping        : 3\ncpu MHz         : 547.630\ncache size      : 1024 KB\nfdiv_bug        : no\nhlt_bug         : no\nf00f_bug        : no\ncoma_bug        : no\nfpu             : yes\nfpu_exception   : yes\ncpuid level     : 2\nwp              : yes\nflags           : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge\nmca cmov pat pse36 mmx fxsr sse\nbogomips        : 1095.37\nclflush size    : 32\npower management:\n"},{"id":"119277","messageId":"eaa105840907310822m57cc29a4q84aa4eb274bd01aa@mail.gmail.com","threadId":"20241","inReplyTo":"20090731104602.15375.qmail@science.horizon.com","subject":"Re: Request for benchmarking: x86 SHA1 code","fromName":"Peter Harris","fromEmail":"git@peter.is-a-geek.org","sentAt":"2009-07-31T15:22:20Z","receivedAt":"2009-07-31T15:22:20Z","isPatch":false,"sender":{"key":"git@peter.is-a-geek.org","avatar":null},"body":"On Fri, Jul 31, 2009 at 6:46 AM, George Spelvin wrote:\n> Could someone do some benchmarking for me?  Old (486/Pentium/P2/P3)\n> machines are also interesting, but I'm optimizing for newer ones.\n\nMy Geode isn't old in age, but I admit it's old in design (and the\nvendor switched to Atom right after I bought it...)\n\nBest of three runs:\n\n586test: 26.536597 s\nx86test: 26.111148 s\n\n$ cat /proc/cpuinfo\nprocessor       : 0\nvendor_id       : AuthenticAMD\ncpu family      : 5\nmodel           : 10\nmodel name      : Geode(TM) Integrated Processor by AMD PCS\nstepping        : 2\ncpu MHz         : 499.927\ncache size      : 128 KB\nfdiv_bug        : no\nhlt_bug         : no\nf00f_bug        : no\ncoma_bug        : no\nfpu             : yes\nfpu_exception   : yes\ncpuid level     : 1\nwp              : yes\nflags           : fpu de pse tsc msr cx8 sep pge cmov clflush mmx\nmmxext 3dnowext 3dnow\nbogomips        : 1001.72\nclflush size    : 32\npower management:\n\nPeter Harris\n"},{"id":"119390","messageId":"20090803034741.23415.qmail@science.horizon.com","threadId":"20241","inReplyTo":"20090731104602.15375.qmail@science.horizon.com","subject":"x86 SHA1: Faster than OpenSSL","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-08-03T03:47:41Z","receivedAt":"2009-08-03T03:47:41Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"(Work in progress, state dump to mailing list archives.)\n\nThis started when discussing git startup overhead due to the dynamic\nlinker.  One big contributor is the openssl library, which is used only\nfor its optimized x86 SHA-1 implementation.  So I took a look at it,\nwith an eye to importing the code directly into the git source tree,\nand decided that I felt like trying to do better.\n\nThe original code was excellent, but it was optimized when the P4 was new.\nAfter a bit of tweaking, I've inflicted a slight (1.4%) slowdown on the\nP4, but a small-but-noticeable speedup on a variety of other processors.\n\nBefore      After       Gain    Processor\n1.585248    1.353314\t+17%\t2500 MHz Phenom\n3.249614    3.295619\t-1.4%\t1594 MHz P4\n1.414512    1.352843\t+4.5%\t2.66 GHz i7\n3.460635    3.284221\t+5.4%\t1596 MHz Athlon XP\n4.077993    3.891826\t+4.8%\t1144 MHz Athlon\n1.912161    1.623212\t+17%\t2100 MHz Athlon 64 X2\n2.956432    2.940210\t+0.55%\t1794 MHz Mobile Celeron (fam 15 model 2)\n\n(Seconds to hash 500x 1 MB, best of 10 runs in all cases.)\n\nThis is based on Andy Polyakov's GPL/BSD licensed cryptogams code, and\n(for now) uses the same perl preprocessor.   To test it, do the following:\n- Download Andy's original code from\n  http://www.openssl.org/~appro/cryptogams/cryptogams-0.tar.gz\n- \"tar xz cryptogams-0.tar.gz\"\n- \"cd cryptogams-0/x86\"\n- \"patch < this_email\" to create \"sha1test.c\", \"sha1-586.h\", \"Makefile\",\n   and \"sha1-x86.pl\".\n- \"make\" \n- Run ./586test (before) and ./x86test (after) and note the timings.\n\nThe code is currenty only the core SHA transform.  Adding the appropriate\ninit/uodate/finish wrappers is straightforward.\n\nAn open question is how to add appropriate CPU detection to the git\nbuild scripts.  (Note that `uname -m`, which it currently uses to select\nthe ARM code, does NOT produce the right answer if you're using a 32-bit\ncompiler on a 64-bit platform.)\n\nI try to explain it in the comments, but with all the software pipelining\nthat makes the rounds overlap (and there are at least 4 different kinds\nof rounds, which overlap with each other), it's a bit intricate.  If you\nfeel really masochistic, make commenting suggestions...\n\n--- /dev/null\t2009-05-12 02:55:38.579106460 -0400\n+++ Makefile\t2009-08-02 06:44:44.000000000 -0400\n@@ -0,0 +1,16 @@\n+CC := gcc\n+CFLAGS := -m32 -W -Wall -Os -g\n+ASFLAGS := --32\n+\n+all: 586test x86test\n+\n+586test : sha1test.c sha1-586.o\n+\t$(CC) $(CFLAGS) -o $@ sha1test.c sha1-586.o\n+\n+x86test : sha1test.c sha1-x86.o\n+\t$(CC) $(CFLAGS) -o $@ sha1test.c sha1-x86.o\n+\n+586test x86test : sha1-586.h\n+\n+%.s : %.pl x86asm.pl x86unix.pl\n+\tperl $< elf > $@\n--- /dev/null\t2009-05-12 02:55:38.579106460 -0400\n+++ sha1-586.h\t2009-08-02 06:44:17.000000000 -0400\n@@ -0,0 +1,3 @@\n+#include <stdint.h>\n+\n+void sha1_block_data_order(uint32_t iv[5], void const *in, unsigned len);\n--- /dev/null\t2009-05-12 02:55:38.579106460 -0400\n+++ sha1test.c\t2009-08-02 08:27:48.449609504 -0400\n@@ -0,0 +1,85 @@\n+#include <stdint.h>\n+#include <stdlib.h>\n+#include <stdio.h>\n+#include <sys/time.h>\n+\n+#include \"sha1-586.h\"\n+\n+#define SIZE 1000000\n+\n+#if SIZE % 64\n+# error SIZE must be a multiple of 64!\n+#endif\n+\n+static unsigned long\n+timing_test(uint32_t iv[5], unsigned iter)\n+{\n+\tunsigned i;\n+\tstruct timeval tv0, tv1;\n+\tstatic char *p;\t/* Very large buffer */\n+\n+\tif (!p) {\n+\t\tp = malloc(SIZE);\n+\t\tif (!p) {\n+\t\t\tperror(\"malloc\");\n+\t\t\texit(1);\n+\t\t}\n+\t}\n+\n+\tif (gettimeofday(&tv0, NULL) < 0) {\n+\t\tperror(\"gettimeofday\");\n+\t\texit(1);\n+\t}\n+\tfor (i = 0; i < iter; i++)\n+\t\tsha1_block_data_order(iv, p, SIZE/64);\n+\tif (gettimeofday(&tv1, NULL) < 0) {\n+\t\tperror(\"gettimeofday\");\n+\t\texit(1);\n+\t}\n+\treturn 1000000ul * (tv1.tv_sec-tv0.tv_sec) + tv1.tv_usec-tv0.tv_usec;\n+}\n+\n+int\n+main(void)\n+{\n+\tuint32_t iv[5] = {\n+\t\t0x67452301, 0xefcdab89, 0x98badcfe, 0x10325476, 0xc3d2e1f0\n+\t};\n+\t/* Simplest known-answer test, \"abc\" */\n+\tstatic uint8_t const testbuf[64] = {\n+\t\t'a','b','c', 0x80, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n+\t\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n+\t\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n+\t\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 24\n+\t};\n+\t/* Expected: A9993E364706816ABA3E25717850C26C9CD0D89D */\n+\tstatic uint32_t const expected[5] = {\n+\t\t0xa9993e36, 0x4706816a, 0xba3e2571, 0x7850c26c, 0x9cd0d89d };\n+\tunsigned i;\n+\tunsigned long min_usec = -1ul;\n+\n+\t/* quick correct-answer test.  silent unless successful. */\n+\tsha1_block_data_order(iv, testbuf, 1);\n+\tfor (i = 0; i < 5; i++) {\n+\t\tif (iv[i] != expected[i]) {\n+\t\t\tprintf(\"Result:  %08x %08x %08x %08x %08x\\n\"\n+\t\t\t       \"Expected:%08x %08x %08x %08x %08x\\n\",\n+\t\t\t\tiv[0], iv[1], iv[2], iv[3], iv[4], expected[0],\n+\t\t\t\texpected[1], expected[2], expected[3],\n+\t\t\t\texpected[4]);\n+\t\t\tbreak;\n+\t\t}\n+\t}\n+\n+\tfor (i = 0; i < 10; i++) {\n+\t\tunsigned long usec = timing_test(iv, 500);\n+\t\tprintf(\"%2u/10: %u.%06u s\\n\", i+1, (unsigned)(usec/1000000),\n+\t\t\t(unsigned)(usec % 1000000));\n+\t\tif (usec < min_usec)\n+\t\t\tmin_usec = usec;\n+\t}\n+\tprintf(\"Minimum time to hash %u bytes: %u.%06u\\n\", \n+\t\t500 * SIZE, (unsigned)(min_usec/1000000),\n+\t\t(unsigned)(min_usec % 1000000));\n+\treturn 0;\n+}\n--- /dev/null\t2009-05-12 02:55:38.579106460 -0400\n+++ sha1-x86.pl\t2009-08-02 08:51:01.069614130 -0400\n@@ -0,0 +1,389 @@\n+#!/usr/bin/env perl\n+\n+# ====================================================================\n+# [Re]written by Andy Polyakov <appro@fy.chalmers.se> for the OpenSSL\n+# project. The module is, however, dual licensed under OpenSSL and\n+# CRYPTOGAMS licenses depending on where you obtain it. For further\n+# details see http://www.openssl.org/~appro/cryptogams/.\n+# ====================================================================\n+\n+# \"[Re]written\" was achieved in two major overhauls. In 2004 BODY_*\n+# functions were re-implemented to address P4 performance issue [see\n+# commentary below], and in 2006 the rest was rewritten in order to\n+# gain freedom to liberate licensing terms.\n+\n+# It was noted that Intel IA-32 C compiler generates code which\n+# performs ~30% *faster* on P4 CPU than original *hand-coded*\n+# SHA1 assembler implementation. To address this problem (and\n+# prove that humans are still better than machines:-), the\n+# original code was overhauled, which resulted in following\n+# performance changes:\n+#\n+#\t\tcompared with original\tcompared with Intel cc\n+#\t\tassembler impl.\t\tgenerated code\n+# Pentium\t-16%\t\t\t+48%\n+# PIII/AMD\t+8%\t\t\t+16%\n+# P4\t\t+85%(!)\t\t\t+45%\n+#\n+# As you can see Pentium came out as looser:-( Yet I reckoned that\n+# improvement on P4 outweights the loss and incorporate this\n+# re-tuned code to 0.9.7 and later.\n+# ----------------------------------------------------------------\n+#\t\t\t\t\t<appro@fy.chalmers.se>\n+\n+$0 =~ m/(.*[\\/\\\\])[^\\/\\\\]+$/; $dir=$1;\n+push(@INC,\"${dir}\",\"${dir}../../perlasm\");\n+require \"x86asm.pl\";\n+\n+&asm_init($ARGV[0],\"sha1-586.pl\",$ARGV[$#ARGV] eq \"386\");\n+\n+$A=\"eax\";\n+$B=\"ebx\";\n+$C=\"ecx\";\n+$D=\"edx\";\n+$E=\"ebp\";\n+\n+# Two temporaries\n+$S=\"edi\";\n+$T=\"esi\";\n+\n+# The round constants\n+use constant K1 => 0x5a827999;\n+use constant K2 => 0x6ED9EBA1;\n+use constant K3 => 0x8F1BBCDC;\n+use constant K4 => 0xCA62C1D6;\n+\n+# Given unlimited registers and functional units, it would be possible to\n+# compute SHA-1 at two cycles per round, using 7 operations per round.\n+# Remember, each round computes a new value for E, which is used as A\n+# in the following round and B in the round after that.  There are two\n+# critical paths:\n+# - A must be rotated and added to the output E (2 cycles between rounds)\n+# - B must go through two boolean operations before being added to\n+#   the result E.  Since this latter addition can't be done in the\n+#   same-cycle as the critical addition of a<<<5, this is a total of\n+#   2+1+1 = 4 cycles per 2 rounds.\n+# Additionally, if you want to avoid copying B, you have to rotate it\n+# soon after use in this round so it is ready for use as the following\n+# round's C.\n+\n+# e += (a <<< 5) + K + in[i] + (d^(b&(c^d)))\t\t(0..19)\n+# e += (a <<< 5) + K + in[i] + (b^c^d)\t\t\t(20..39, 60..79)\n+# e += (a <<< 5) + K + in[i] + (c&d) + (b&(c^d))\t(40..59)\n+#\n+# The hard part is doing this with only two temporary registers.\n+# Taking the most complex F(b,c,d) function, writing it as above\n+# breaks it into two parts which can be accumulated into e separately.\n+# Let's divide it into 4 parts.  These can be executed in a 7-cycle\n+# loop, assuming an in-order triple issue machine\n+# (quadruple counting xor-from-memory as 2):\n+#\n+#\tin[i]\t\tF1(c,d)\t\tF2(b,c,d)\ta<<<5\n+#\tmov in[i],T\t\t\t(addl S,A)\t(movl B,S)\n+#\txor in[i+1],T\t\t\t\t\t(rorl 5,S)\n+#\txor in[i+2],T\tmovl D,S\t\t\t(addl S,A)\n+#\txor in[i+3],T\tandl C,S\n+#\trotl 1,T\taddl S,E\tmovl D,S\n+#\tmovl T,in[i]\t\t\txorl C,S\n+#\taddl T+K,E\t\t\tandl B,S\trorl 2,B\t//\n+#\t(mov in[i],T)\t\t\taddl S,E\tmovl A,S\n+#\t(xor in[i+1],T)\t\t\t\t\trorl 5,S\n+#\t(xor in[i+2],T)\t(movl C,S)\t\t\taddl S,E\n+#\n+# In the above, we routinely read and write a register on the same cycle,\n+# overlapping the beginning of one computation with the end of another.\n+# I've tried to place the reads to the left of the writes, but some of the\n+# overlapping operations from adjacent rounds (given in parentheses)\n+# violate that.\n+#\n+# The \"addl T+K,E\" line is actually implemented using the lea instruction,\n+# which (on a Pentium) requires that neither T not K was modified on the\n+# previous cycle.\n+#\n+# As you can see, in the absence of out-of-order execution, the first\n+# column takes a minimum of 6 cycles (fetch, 3 XORs, rotate, add to E),\n+# and I reserve a seventh cycle before the add to E so that I can use a\n+# Pentium's lea instruction.\n+#\n+# The other three columns take 3, 4 and 3 cycles, respectively.\n+# These can all be overlapped by 1 cycle assuming a superscalar\n+# processor, for a total of 2+2+3 = 7 cycles.\n+#\n+# The other F() functions require 5 and 4 cycles, respectively.\n+# overlapped with the 3-cycle a<<<5 computation, that makes a total of 6\n+# and 5 cycles, respectively.  If we overlap the beginning and end of the\n+# Xi computation, we can get it down to 6 cycles, but below that, we'll\n+# just have to waste a cycle.\n+#\n+# For the first 16 rounds, forming Xi is just a fetch, and the F()\n+# function only requires 5 cycles, so the whole round can be pulled in\n+# to 4 cycles.\n+\n+\n+# Atom pairing rules (not yet optimized):\n+# The Atom has a dial-issue in-order pipeline, similar to the\n+# original Pentium.  However, the issue restrictions are different.\n+# In particular, all memory source operations must use st use port 0,\n+# as must all rotates.\n+#\n+# Given that a round uses 4 fetches and 3 rotates, that's going to\n+# require significant care to pair well.  It may take a completely\n+# different implementation.\n+#\n+# LEA must use port 1, but apparently it has even worse address generation\n+# interlock latency than the Pentium.  Oh well, it's still the best way\n+# to do a 3-way add with a 32-bit immediate.\n+\n+\n+# The first 16 rounds use s simple simplest F(b,c,d) = d^(b&(c^d)), and\n+# don't need to form the in[i] equation, letting us reduce the round to\n+# 4 cycles, after some reassignment of temporaries:\n+\n+#\t\t\tmovl D,S\t(roll 5,T)\t(addl S,A)\t//\n+#\tmov in[i],T\txorl C,S\t(addl T,A)\n+#\t\t\tandl B,S\t\t\trorl 2,B\n+#\taddl T+K,E\txorl D,S\tmovl A,T\n+#\t\t\taddl S,E\troll 5,T\t(movl C,S)\t//\n+#\t(mov in[i],T)\t(xorl B,S)\taddl T,E\t\n+\n+# The // mark where the round function starts.  Each round expects the\n+# first copy of D to S to have been done already.\n+\n+# The transition between rounds 15 and 16 is a bit tricky... the best\n+# thing to do is to delay the computation of a<<<5 one cycle and move it back\n+# to the S register.  That way, T is free as early as possible.\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\t(addl T+K,A)\t(xorl E,S)\t(movl B,T)\n+#\t\t\tmovl D,S\t(roll 5,T)\t(addl S,A)\t//\n+#\tmov in[i],T\txorl C,S\t(addl T,A)\n+#\t\t\tandl B,S\t\t\trorl 2,B\n+#\taddl T+K,E \txorl D,S\t\t\t(movl in[1],T)\n+#\t(xor in[i+1],T)\taddl S,E\tmovl A,S\n+#\t(xor in[i+2],T)\trorl 2,B\troll 5,S\t//\n+#\t(xor in[i+3],T)\t(movl C,S)\taddl S,E\n+\n+sub BODY_00_15\n+{\n+\tlocal($n,$a,$b,$c,$d,$e)=@_;\n+\n+\t&comment(\"00_15 $n\");\n+\t\t&mov($S,$d) if ($n == 0);\n+\t&mov($T,&swtmp($n%16));\t\t#  V Load Xi.\n+\t\t&xor($S,$c);\t\t# U  Continue F() = d^(b&(c^d))\n+\t\t&and($S,$b);\t\t#  V\n+\t\t\t&rotr($b,2);\t# NP\n+\t&lea($e,&DWP(K1,$e,$T));\t# U  Add Xi and K\n+    if ($n < 15) {\n+\t\t\t&mov($T,$a);\t#  V\n+\t\t&xor($S,$d);\t\t# U \n+\t\t\t&rotl($T,5);\t# NP\n+\t\t&add($e,$S);\t\t# U \n+\t\t&mov($S,$c);\t\t#  V Start of NEXT round's F()\n+\t\t\t&add($e,$T);\t# U \n+    } else {\n+\t# This version provides the correct start for BODY_20_39\n+\t\t&xor($S,$d);\t\t#  V\n+\t&mov($T,&swtmp(($n+1)%16));\t# U  Start computing mext Xi.\n+\t\t&add($e,$S);\t\t#  V Add F()\n+\t\t\t&mov($S,$a);\t# U  Start computing a<<<5\n+\t&xor($T,&swtmp(($n+3)%16));\t#  V\n+\t\t\t&rotl($S,5);\t# U \n+\t&xor($T,&swtmp(($n+9)%16));\t#  V\n+    }\n+}\n+\n+\n+# A full round using F(b,c,d) = b^c^d.  6 cycles of dependency chain.\n+# This starts just before starting to compute F(); the Xi should have XORed\n+# the first three values together.  (Break is at //)\n+#\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\tmov in[i],T\t(xorl E,S)\t\t\t(addl T+K,A)\n+#\txor in[i+1],T\t(addl S,A)\t(movl B,S)\n+#\txor in[i+2],T\t\t\t(roll 5,S)\t//\n+#\txor in[i+3],T \tmovl D,S  \t(addl S,A)\n+#\troll 1,T\txorl C,S\n+#\tmovl T,in[i]\tandl B,S\t\t\trorl 2,B\n+#\taddl T+K,E \txorl D,S\t\t\t(mov in[i],T)\n+#\t(xor in[i+1],T)\taddl S,E\tmovl A,S\n+#\t(xor in[i+2],T)\t\t\troll 5,S\t//\n+#\t(xor in[i+3],T)\t(movl C,S)\taddl S,E\n+\n+sub BODY_16_19\n+{\n+\tlocal($n,$a,$b,$c,$d,$e)=@_;\n+\n+\t&comment(\"16_19 $n\");\n+\n+\t&xor($T,&swtmp(($n+13)%16));\t# U \n+\t\t\t&add($a,$S);\t#  V End of previous round\n+\t\t&mov($S,$d);\t\t# U  Start current round's F()\n+\t&rotl($T,1);\t\t\t#  V\n+\t\t&xor($S,$c);\t\t# U \n+\t&mov(&swtmp($n%16),$T);\t\t# U Store computed Xi.\n+\t\t&and($S,$b);\t\t#  V\n+\t\t&rotr($b,2);\t\t# NP\n+\t&lea($e,&DWP(K1,$e,$T));\t# U  Add Xi and K\n+\t&mov($T,&swtmp(($n+1)%16));\t#  V Start computing mext Xi.\n+\t\t&xor($S,$d);\t\t# U\n+\t&xor($T,&swtmp(($n+3)%16));\t#  V\n+\t\t&add($e,$S);\t\t# U  Add F()\n+\t\t\t&mov($S,$a);\t#  V Start computing a<<<5\n+\t&xor($T,&swtmp(($n+9)%16));\t# U\n+\t\t\t&rotl($S,5);\t# NP\n+}\n+\n+# This is just like BODY_16_19, but computes a different F() = b^c^d\n+#\n+#\tin[i]\t\tF(b,c,d)\ta<<<5\n+#\tmov in[i],T\t(xorl E,S)\t\t\t(addl T+K,A)\n+#\txor in[i+1],T\t(addl S,A)\t(movl B,S)\n+#\txor in[i+2],T\t\t\t(roll 5,S)\t//\n+#\txor in[i+3],T \t\t\t(addl S,A)\n+#\troll 1,T\tmovl D,S\n+#\tmovl T,in[i]\txorl B,S\t\t\trorl 2,B\n+#\taddl T+K,E \txorl C,S\t\t\t(mov in[i],T)\n+#\t(xor in[i+1],T)\taddl S,E\tmovl A,S\n+#\t(xor in[i+2],T)\t\t\troll 5,S\t//\n+#\t(xor in[i+3],T)\t(movl C,S)\taddl S,E\n+\n+sub BODY_20_39\t# And 61..79\n+{\n+\tlocal($n,$a,$b,$c,$d,$e)=@_;\n+\tlocal $K=($n<40) ? K2 : K4;\n+\n+\t&comment(\"20_39 $n\");\n+\n+\t&xor($T,&swtmp(($n+13)%16));\t# U \n+\t\t\t&add($a,$S);\t#  V End of previous round\n+\t&rotl($T,1);\t\t\t# U\n+\t\t&mov($S,$d);\t\t#  V Start current round's F()\n+\t&mov(&swtmp($n%16),$T) if ($n < 77);\t# Store computed Xi.\n+\t\t&xor($S,$b);\t\t#  V\n+\t\t&rotr($b,2);\t\t# NP\n+\t&lea($e,&DWP($K,$e,$T));\t# U  Add Xi and K\n+\t&mov($T,&swtmp(($n+1)%16)) if ($n < 79); # Start computing next Xi.\n+\t\t&xor($S,$c);\t\t# U\n+\t&xor($T,&swtmp(($n+3)%16)) if ($n < 79);\n+\t\t&add($e,$S);\t\t# U  Add F1()\n+\t\t\t&mov($S,$a);\t#  V Start computing a<<<5\n+\t&xor($T,&swtmp(($n+9)%16)) if ($n < 79);\n+\t\t\t&rotl($S,5);\t# NP\n+\n+\t\t\t&add($e,$S) if ($n == 79);\n+}\n+\n+\n+# This starts immediately after the LEA, and expects to need to finish\n+# the previous round. (break is at //)\n+#\n+#\tin[i]\t\tF1(c,d)\t\tF2(b,c,d)\ta<<<5\n+#\t(addl T+K,E)\t\t\t(andl C,S)\t(rorl 2,C)\n+#\tmov in[i],T\t\t\t(addl S,A)\t(movl B,S)\n+#\txor in[i+1],T\t\t\t\t\t(rorl 5,S)\n+#\txor in[i+2],T /\tmovl D,S\t\t\t(addl S,A)\n+#\txor in[i+3],T\tandl C,S\n+#\trotl 1,T\taddl S,E\tmovl D,S\n+#\t(movl T,in[i])\t\t\txorl C,S\n+#\taddl T+K,E\t\t\tandl B,S\trorl 2,B\n+#\t(mov in[i],T)\t\t\taddl S,E\tmovl A,S\n+#\t(xor in[i+1],T)\t\t\t\t\trorl 5,S\n+#\t(xor in[i+2],T)\t// (movl C,S)\t\t\taddl S,E\n+\n+sub BODY_40_59\n+{\n+\tlocal($n,$a,$b,$c,$d,$e)=@_;\n+\n+\t&comment(\"40_59 $n\");\n+\n+\t\t\t&add($a,$S);\t#  V End of previous round\n+\t\t&mov($S,$d);\t\t# U  Start current round's F(1)\n+\t&xor($T,&swtmp(($n+13)%16));\t#  V\n+\t\t&and($S,$c);\t\t# U\n+\t&rotl($T,1);\t\t\t# U\tXXX Missed pairing\n+\t\t&add($e,$S);\t\t#  V Add F1()\n+\t\t&mov($S,$d);\t\t# U  Start current round's F2()\n+\t&mov(&swtmp($n%16),$T);\t\t#  V Store computed Xi.\n+\t\t&xor($S,$c);\t\t# U\n+\t&lea($e,&DWP(K3,$e,$T));\t#  V\n+\t\t&and($S,$b);\t\t# U\tXXX Missed pairing\n+\t\t&rotr($b,2);\t\t# NP\n+\t&mov($T,&swtmp(($n+1)%16));\t# U  Start computing next Xi.\n+\t\t&add($e,$S);\t\t#  V Add F2()\n+\t\t&mov($S,$a);\t\t# U  Start computing a<<<5\n+\t&xor($T,&swtmp(($n+3)%16));\t#  V\n+\t\t\t&rotl($S,5);\t# NP\n+\t&xor($T,&swtmp(($n+9)%16));\t# U\n+}\n+# The above code is NOT optimally paired for the Pentium.  (And thus,\n+# presumably, Atom, which has a very similar dual-issue in-order pipeline.)\n+# However, attempts to improve it make it slower on Phenom & i7.\n+\n+&function_begin(\"sha1_block_data_order\",16);\n+\n+\tlocal @V = ($A,$B,$C,$D,$E);\n+\tlocal @W = ($A,$B,$C);\n+\n+\t&mov($S,&wparam(0));\t# SHA_CTX *c\n+\t&mov($T,&wparam(1));\t# const void *input\n+\t&mov($A,&wparam(2));\t# size_t num\n+\t&stack_push(16);\t# allocate X[16]\n+\t&shl($A,6);\n+\t&add($A,$T);\n+\t&mov(&wparam(2),$A);\t# pointer beyond the end of input\n+\t&mov($E,&DWP(16,$S));# pre-load E\n+\t&mov($D,&DWP(12,$S));# pre-load D\n+\n+\t&set_label(\"loop\",16);\n+\n+\t# copy input chunk to X, but reversing byte order!\n+\t&mov($W[2],&DWP(4*(0),$T));\n+\t&mov($W[1],&DWP(4*(1),$T));\n+\t&bswap($W[2]);\n+\tfor ($i=0; $i<14; $i++) {\n+\t\t&mov($W[0],&DWP(4*($i+2),$T));\n+\t\t&bswap($W[1]);\n+\t\t&mov(&swtmp($i+0),$W[2]);\n+\t\tunshift(@W,pop(@W));\n+\t}\n+\t&bswap($W[1]);\n+\t&mov(&swtmp($i+0),$W[2]);\n+\t&mov(&swtmp($i+1),$W[1]);\n+\n+\t&mov(&wparam(1),$T);\t# redundant in 1st spin\n+\n+\t# Reload A, B and C, which we use as temporaries in the copying\n+\t&mov($C,&DWP(8,$S));\n+\t&mov($B,&DWP(4,$S));\n+\t&mov($A,&DWP(0,$S));\n+\n+\tfor($i=0;$i<16;$i++)\t{ &BODY_00_15($i,@V); unshift(@V,pop(@V)); }\n+\tfor(;$i<20;$i++)\t{ &BODY_16_19($i,@V); unshift(@V,pop(@V)); }\n+\tfor(;$i<40;$i++)\t{ &BODY_20_39($i,@V); unshift(@V,pop(@V)); }\n+\tfor(;$i<60;$i++)\t{ &BODY_40_59($i,@V); unshift(@V,pop(@V)); }\n+\tfor(;$i<80;$i++)\t{ &BODY_20_39($i,@V); unshift(@V,pop(@V)); }\n+\n+\t(($V[4] eq $E) and ($V[0] eq $A)) or die;\t# double-check\n+\n+\t&comment(\"Loop trailer\");\n+\n+\t&mov($S,&wparam(0));\t# re-load SHA_CTX*\n+\t&mov($T,&wparam(1));\t# re-load input pointer\n+\n+\t&add($E,&DWP(16,$S));\n+\t&add($D,&DWP(12,$S));\n+\t&add(&DWP(8,$S),$C);\n+\t&add(&DWP(4,$S),$B);\n+\t &add($T,64);\t\t# advance input pointer\n+\t&add(&DWP(0,$S),$A);\n+\t&mov(&DWP(12,$S),$D);\n+\t&mov(&DWP(16,$S),$E);\n+\n+\t&cmp($T,&wparam(2));\t# have we reached the end yet?\n+\t&jb(&label(\"loop\"));\n+\n+\t&stack_pop(16);\n+&function_end(\"sha1_block_data_order\");\n+&asciz(\"SHA1 block transform for x86, CRYPTOGAMS by <appro\\@openssl.org>\");\n+\n+&asm_finish();\n"},{"id":"119396","messageId":"57518fd10908030036l22ccdcf1va2e8ee450c4f5ee5@mail.gmail.com","threadId":"20241","inReplyTo":"20090803034741.23415.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Jonathan del Strother","fromEmail":"maillist@steelskies.com","sentAt":"2009-08-03T07:36:38Z","receivedAt":"2009-08-03T07:36:38Z","isPatch":false,"sender":{"key":"jon.delstrother@bestbefore.tv","avatar":"https://gravatar.com/avatar/754e21ab701c00e2d21fc261187254c34b2a1c0b959d9ee5be1a295990be3081?d=mp&s=160"},"body":"On Mon, Aug 3, 2009 at 4:47 AM, George Spelvin<linux@horizon.com> wrote:\n> (Work in progress, state dump to mailing list archives.)\n>\n> This started when discussing git startup overhead due to the dynamic\n> linker.  One big contributor is the openssl library, which is used only\n> for its optimized x86 SHA-1 implementation.  So I took a look at it,\n> with an eye to importing the code directly into the git source tree,\n> and decided that I felt like trying to do better.\n>\n\nFWIW, this doesn't work on OS X / Darwin.  'as' doesn't take a --32\nflag, it takes an -arch i386 flag.  After changing that, I get:\n\nas -arch i386  -o sha1-586.o sha1-586.s\nsha1-586.s:4:Unknown pseudo-op: .type\nsha1-586.s:4:Rest of line ignored. 1st junk character valued 115 (s).\nsha1-586.s:5:Alignment too large: 15. assumed.\nsha1-586.s:19:Alignment too large: 15. assumed.\nsha1-586.s:1438:Unknown pseudo-op: .size\nsha1-586.s:1438:Rest of line ignored. 1st junk character valued 115 (s).\nmake: *** [sha1-586.o] Error 1\n\n- at which point I have no idea how to fix it.\n"},{"id":"119448","messageId":"ca433830908031840o11efc252r4db65f671deb4913@mail.gmail.com","threadId":"20241","inReplyTo":"20090803034741.23415.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Mark Lodato","fromEmail":"lodatom@gmail.com","sentAt":"2009-08-04T01:40:43Z","receivedAt":"2009-08-04T01:40:43Z","isPatch":false,"sender":{"key":"lodatom@gmail.com","avatar":"https://avatars.githubusercontent.com/u/58860?v=4"},"body":"On Sun, Aug 2, 2009 at 11:47 PM, George Spelvin<linux@horizon.com> wrote:\n> Before      After       Gain    Processor\n> 1.585248    1.353314    +17%    2500 MHz Phenom\n> 3.249614    3.295619    -1.4%   1594 MHz P4\n> 1.414512    1.352843    +4.5%   2.66 GHz i7\n> 3.460635    3.284221    +5.4%   1596 MHz Athlon XP\n> 4.077993    3.891826    +4.8%   1144 MHz Athlon\n> 1.912161    1.623212    +17%    2100 MHz Athlon 64 X2\n> 2.956432    2.940210    +0.55%  1794 MHz Mobile Celeron (fam 15 model 2)\n>\n> (Seconds to hash 500x 1 MB, best of 10 runs in all cases.)\n>\n> This is based on Andy Polyakov's GPL/BSD licensed cryptogams code, and\n> (for now) uses the same perl preprocessor.   To test it, do the following:\n> - Download Andy's original code from\n>  http://www.openssl.org/~appro/cryptogams/cryptogams-0.tar.gz\n> - \"tar xz cryptogams-0.tar.gz\"\n> - \"cd cryptogams-0/x86\"\n> - \"patch < this_email\" to create \"sha1test.c\", \"sha1-586.h\", \"Makefile\",\n>   and \"sha1-x86.pl\".\n> - \"make\"\n> - Run ./586test (before) and ./x86test (after) and note the timings.\n\nNote, to compile this on Ubuntu x86-64, I had to:\n$ sudo apt-get install libc6-dev-i386\n\n$ ./586test\n 1/10: 2.016621 s\n 2/10: 2.030742 s\n 3/10: 2.027333 s\n 4/10: 2.024018 s\n 5/10: 2.022306 s\n 6/10: 2.022418 s\n 7/10: 2.047103 s\n 8/10: 2.035467 s\n 9/10: 2.032237 s\n10/10: 2.029231 s\nMinimum time to hash 500000000 bytes: 2.016621\n$ ./x86test\n 1/10: 1.818661 s\n 2/10: 1.814856 s\n 3/10: 1.816232 s\n 4/10: 1.815208 s\n 5/10: 1.834047 s\n 6/10: 1.843020 s\n 7/10: 1.819564 s\n 8/10: 1.815560 s\n 9/10: 1.824232 s\n10/10: 1.820943 s\nMinimum time to hash 500000000 bytes: 1.814856\n$ python -c 'print 2.016621 / 1.814856'\n1.11117410968\n$ cat /proc/cpuinfo\nprocessor       : 0\nvendor_id       : GenuineIntel\ncpu family      : 6\nmodel           : 15\nmodel name      : Intel(R) Core(TM)2 CPU          6300  @ 1.86GHz\nstepping        : 2\ncpu MHz         : 1861.825\ncache size      : 2048 KB\nphysical id     : 0\nsiblings        : 2\ncore id         : 0\ncpu cores       : 2\napicid          : 0\ninitial apicid  : 0\nfpu             : yes\nfpu_exception   : yes\ncpuid level     : 10\nwp              : yes\nflags           : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov\npat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx lm constant\n_tsc arch_perfmon pebs bts rep_good pni dtes64 monitor ds_cpl vmx est tm2 ssse3\ncx16 xtpr pdcm lahf_lm tpr_shadow\nbogomips        : 3723.65\nclflush size    : 64\ncache_alignment : 64\naddress sizes   : 36 bits physical, 48 bits virtual\npower management:\n\nprocessor       : 1\nvendor_id       : GenuineIntel\ncpu family      : 6\nmodel           : 15\nmodel name      : Intel(R) Core(TM)2 CPU          6300  @ 1.86GHz\nstepping        : 2\ncpu MHz         : 1861.825\ncache size      : 2048 KB\nphysical id     : 0\nsiblings        : 2\ncore id         : 1\ncpu cores       : 2\napicid          : 1\ninitial apicid  : 1\nfpu             : yes\nfpu_exception   : yes\ncpuid level     : 10\nwp              : yes\nflags           : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov\npat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx lm constant\n_tsc arch_perfmon pebs bts rep_good pni dtes64 monitor ds_cpl vmx est tm2 ssse3\ncx16 xtpr pdcm lahf_lm tpr_shadow\nbogomips        : 3724.01\nclflush size    : 64\ncache_alignment : 64\naddress sizes   : 36 bits physical, 48 bits virtual\npower management:\n\n\n\nI imagine that you can get a bigger speedup by making a 64-bit version\n(but maybe not).  Either way, it would be nice if x86-64 users did not\nhave to install an additional package to compile.\n\nCheers,\nMark\n"},{"id":"119453","messageId":"alpine.LFD.2.01.0908031924230.3270@localhost.localdomain","threadId":"20241","inReplyTo":"20090803034741.23415.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-04T02:30:08Z","receivedAt":"2009-08-04T02:30:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 2 Aug 2009, George Spelvin wrote:\n> \n> The original code was excellent, but it was optimized when the P4 was new.\n> After a bit of tweaking, I've inflicted a slight (1.4%) slowdown on the\n> P4, but a small-but-noticeable speedup on a variety of other processors.\n> \n> Before      After       Gain    Processor\n> 1.585248    1.353314\t+17%\t2500 MHz Phenom\n> 3.249614    3.295619\t-1.4%\t1594 MHz P4\n> 1.414512    1.352843\t+4.5%\t2.66 GHz i7\n> 3.460635    3.284221\t+5.4%\t1596 MHz Athlon XP\n> 4.077993    3.891826\t+4.8%\t1144 MHz Athlon\n> 1.912161    1.623212\t+17%\t2100 MHz Athlon 64 X2\n> 2.956432    2.940210\t+0.55%\t1794 MHz Mobile Celeron (fam 15 model 2)\n\nIt would be better to have a more git-centric benchmark that actually \nshows some real git load, rather than a sha1-only microbenchmark.\n\nThe thing that I'd prefer is simply\n\n\tgit fsck --full\n\non the Linux kernel archive. For me (with a fast machine), it takes about \n4m30s with the OpenSSL SHA1, and takes 6m40s with the Mozilla SHA1 (ie \nusing a NO_OPENSSL=1 build).\n\nSo that's an example of a load that is actually very sensitive to SHA1 \nperformance (more so than _most_ git loads, I suspect), and at the same \ntime is a real git load rather than some SHA1-only microbenchmark. It also \nshows very clearly why we default to the OpenSSL version over the Mozilla \none.\n\nNOTE! I didn't do multiple runs to see how stable the numbers are, and \nso it's possible that I exaggerated the OpenSSL advantage over the \nMozilla-SHA1 code. Or vice versa. My point is really only that I don't \nknow how meaningful a \"50 x 1M SHA1\" benchmark is, while I know that a \n\"git fsck\" benchmark has at least _some_ real life value.\n\n\t\tLinus\n"},{"id":"119454","messageId":"alpine.LFD.2.01.0908031938280.3270@localhost.localdomain","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908031924230.3270@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-04T02:51:05Z","receivedAt":"2009-08-04T02:51:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Aug 2009, Linus Torvalds wrote:\n> \n> The thing that I'd prefer is simply\n> \n> \tgit fsck --full\n> \n> on the Linux kernel archive. For me (with a fast machine), it takes about \n> 4m30s with the OpenSSL SHA1, and takes 6m40s with the Mozilla SHA1 (ie \n> using a NO_OPENSSL=1 build).\n> \n> So that's an example of a load that is actually very sensitive to SHA1 \n> performance (more so than _most_ git loads, I suspect), and at the same \n> time is a real git load rather than some SHA1-only microbenchmark. It also \n> shows very clearly why we default to the OpenSSL version over the Mozilla \n> one.\n\n\"perf report --sort comm,dso,symbol\" profiling shows the following for \n'git fsck --full' on the kernel repo, using the Mozilla SHA1:\n\n    47.69%               git  /home/torvalds/git/git     [.] moz_SHA1_Update\n    22.98%               git  /lib64/libz.so.1.2.3       [.] inflate_fast\n     7.32%               git  /lib64/libc-2.10.1.so      [.] __GI_memcpy\n     4.66%               git  /lib64/libz.so.1.2.3       [.] inflate\n     3.76%               git  /lib64/libz.so.1.2.3       [.] adler32\n     2.86%               git  /lib64/libz.so.1.2.3       [.] inflate_table\n     2.41%               git  /home/torvalds/git/git     [.] lookup_object\n     1.31%               git  /lib64/libc-2.10.1.so      [.] _int_malloc\n     0.84%               git  /home/torvalds/git/git     [.] patch_delta\n     0.78%               git  [kernel]                   [k] hpet_next_event\n\nso yeah, SHA1 performance matters. Judging by the OpenSSL numbers, the \nOpenSSL SHA1 implementation must be about twice as fast as the C version \nwe use.\n\nThat said, under \"normal\" git usage models, the SHA1 costs are almost \ninvisible. So git-fsck is definitely a fairly unusual case that stresses \nthe SHA1 performance more than most git lods.\n\n\t\tLinus\n"},{"id":"119455","messageId":"9e4733910908032007td74ef9fp669d0d958df67c1@mail.gmail.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908031938280.3270@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2009-08-04T03:07:21Z","receivedAt":"2009-08-04T03:07:21Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On Mon, Aug 3, 2009 at 10:51 PM, Linus\nTorvalds<torvalds@linux-foundation.org> wrote:\n>\n>\n> On Mon, 3 Aug 2009, Linus Torvalds wrote:\n>>\n>> The thing that I'd prefer is simply\n>>\n>>       git fsck --full\n>>\n>> on the Linux kernel archive. For me (with a fast machine), it takes about\n>> 4m30s with the OpenSSL SHA1, and takes 6m40s with the Mozilla SHA1 (ie\n>> using a NO_OPENSSL=1 build).\n>>\n>> So that's an example of a load that is actually very sensitive to SHA1\n>> performance (more so than _most_ git loads, I suspect), and at the same\n>> time is a real git load rather than some SHA1-only microbenchmark. It also\n>> shows very clearly why we default to the OpenSSL version over the Mozilla\n>> one.\n>\n> \"perf report --sort comm,dso,symbol\" profiling shows the following for\n> 'git fsck --full' on the kernel repo, using the Mozilla SHA1:\n>\n>    47.69%               git  /home/torvalds/git/git     [.] moz_SHA1_Update\n>    22.98%               git  /lib64/libz.so.1.2.3       [.] inflate_fast\n>     7.32%               git  /lib64/libc-2.10.1.so      [.] __GI_memcpy\n>     4.66%               git  /lib64/libz.so.1.2.3       [.] inflate\n>     3.76%               git  /lib64/libz.so.1.2.3       [.] adler32\n>     2.86%               git  /lib64/libz.so.1.2.3       [.] inflate_table\n>     2.41%               git  /home/torvalds/git/git     [.] lookup_object\n>     1.31%               git  /lib64/libc-2.10.1.so      [.] _int_malloc\n>     0.84%               git  /home/torvalds/git/git     [.] patch_delta\n>     0.78%               git  [kernel]                   [k] hpet_next_event\n>\n> so yeah, SHA1 performance matters. Judging by the OpenSSL numbers, the\n> OpenSSL SHA1 implementation must be about twice as fast as the C version\n> we use.\n\nWould there happen to be a SHA1 implementation around that can compute\nthe SHA1 without first decompressing the data? Databases gain a lot of\nspeed by using special algorithms that can directly operate on the\ncompressed data.\n\n>\n> That said, under \"normal\" git usage models, the SHA1 costs are almost\n> invisible. So git-fsck is definitely a fairly unusual case that stresses\n> the SHA1 performance more than most git lods.\n>\n>                Linus\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"119457","messageId":"20090804044842.6792.qmail@science.horizon.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908031924230.3270@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-08-04T04:48:42Z","receivedAt":"2009-08-04T04:48:42Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> It would be better to have a more git-centric benchmark that actually \n> shows some real git load, rather than a sha1-only microbenchmark.\n> \n> The thing that I'd prefer is simply\n>\n>\tgit fsck --full\n>\n> on the Linux kernel archive. For me (with a fast machine), it takes about \n> 4m30s with the OpenSSL SHA1, and takes 6m40s with the Mozilla SHA1 (ie \n> using a NO_OPENSSL=1 build).\n\nThe actual goal of this effort is to address the dynamic linker startup\ntime issues by removing the second-largest contributor after libcurl,\nnamely openssl.  Optimizing the assembly code is just the fun part. ;-)\n\nAnyway, on the git repository:\n\n[1273]$ time x/git-fsck --full\t\t\t(New SHA1 code)\ndangling tree 524973049a7e4593df4af41e0564912f678a41ac\ndangling tree 7da7d73185a1df5c2a477d2ee5599ac8a58cad56\n\nreal    0m59.306s\nuser    0m58.760s\nsys     0m0.550s\n[1274]$ time ./git-fsck --full\t\t\t(OpenSSL)\ndangling tree 524973049a7e4593df4af41e0564912f678a41ac\ndangling tree 7da7d73185a1df5c2a477d2ee5599ac8a58cad56\n\nreal    1m0.364s\nuser    0m59.970s\nsys     0m0.400s\n\n1.6% is a pretty minor difference, especially as the machine is running\na backup at the time (but it's a quad-core, with near-zero CPU usage;\nthe business is all I/O).\n\nOn the full Linux repository, I repacked it first to make sure that\neverything was in RAM, and I have the first result:\n\n[517]$ time ~/git/x/git-fsck --full\t\t(New SHA1 code)\n\nreal    10m12.702s\nuser    9m48.410s\nsys     0m23.350s\n[518]$ time ~/git/git-fsck --full\t\t(OpenSSL)\n\nreal    10m26.083s\nuser    10m2.800s\nsys     0m22.000s\n\nAgain, 2.2% is not a huge improvement.  But my only goal was not to be worse.\n\n> So that's an example of a load that is actually very sensitive to SHA1 \n> performance (more so than _most_ git loads, I suspect), and at the same \n> time is a real git load rather than some SHA1-only microbenchmark. It also \n> shows very clearly why we default to the OpenSSL version over the Mozilla \n> one.\n\nI wasn't questioning *that*.  As I said, I was just doing the fun part\nof importing a heavily-optimized OpenSSL-like SHA1 implementation into\nthe git source tree.\n\n(The un-fun part is modifying the build process to detect the target\nprocessor and include the right asm automatically.)\n\nAnyway, if you want to test it, here's a crude x86_32-only patch to the\ngit tree.  \"make NO_OPENSSL=1\" to use the new code.\n\ndiff --git a/Makefile b/Makefile\nindex daf4296..8531c39 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -1176,8 +1176,10 @@ ifdef ARM_SHA1\n \tLIB_OBJS += arm/sha1.o arm/sha1_arm.o\n else\n ifdef MOZILLA_SHA1\n-\tSHA1_HEADER = \"mozilla-sha1/sha1.h\"\n-\tLIB_OBJS += mozilla-sha1/sha1.o\n+#\tSHA1_HEADER = \"mozilla-sha1/sha1.h\"\n+#\tLIB_OBJS += mozilla-sha1/sha1.o\n+\tSHA1_HEADER = \"x86/sha1.h\"\n+\tLIB_OBJS += x86/sha1.o x86/sha1-x86.o\n else\n \tSHA1_HEADER = <openssl/sha.h>\n \tEXTLIBS += $(LIB_4_CRYPTO)\ndiff --git a/x86/sha1-x86.s b/x86/sha1-x86.s\nnew file mode 100644\nindex 0000000..96796d4\n--- /dev/null\n+++ b/x86/sha1-x86.s\n@@ -0,0 +1,1372 @@\n+.file\t\"sha1-586.s\"\n+.text\n+.globl\tsha1_block_data_order\n+.type\tsha1_block_data_order,@function\n+.align\t16\n+sha1_block_data_order:\n+\tpushl\t%ebp\n+\tpushl\t%ebx\n+\tpushl\t%esi\n+\tpushl\t%edi\n+\tmovl\t20(%esp),%edi\n+\tmovl\t24(%esp),%esi\n+\tmovl\t28(%esp),%eax\n+\tsubl\t$64,%esp\n+\tshll\t$6,%eax\n+\taddl\t%esi,%eax\n+\tmovl\t%eax,92(%esp)\n+\tmovl\t16(%edi),%ebp\n+\tmovl\t12(%edi),%edx\n+.align\t16\n+.L000loop:\n+\tmovl\t(%esi),%ecx\n+\tmovl\t4(%esi),%ebx\n+\tbswap\t%ecx\n+\tmovl\t8(%esi),%eax\n+\tbswap\t%ebx\n+\tmovl\t%ecx,(%esp)\n+\tmovl\t12(%esi),%ecx\n+\tbswap\t%eax\n+\tmovl\t%ebx,4(%esp)\n+\tmovl\t16(%esi),%ebx\n+\tbswap\t%ecx\n+\tmovl\t%eax,8(%esp)\n+\tmovl\t20(%esi),%eax\n+\tbswap\t%ebx\n+\tmovl\t%ecx,12(%esp)\n+\tmovl\t24(%esi),%ecx\n+\tbswap\t%eax\n+\tmovl\t%ebx,16(%esp)\n+\tmovl\t28(%esi),%ebx\n+\tbswap\t%ecx\n+\tmovl\t%eax,20(%esp)\n+\tmovl\t32(%esi),%eax\n+\tbswap\t%ebx\n+\tmovl\t%ecx,24(%esp)\n+\tmovl\t36(%esi),%ecx\n+\tbswap\t%eax\n+\tmovl\t%ebx,28(%esp)\n+\tmovl\t40(%esi),%ebx\n+\tbswap\t%ecx\n+\tmovl\t%eax,32(%esp)\n+\tmovl\t44(%esi),%eax\n+\tbswap\t%ebx\n+\tmovl\t%ecx,36(%esp)\n+\tmovl\t48(%esi),%ecx\n+\tbswap\t%eax\n+\tmovl\t%ebx,40(%esp)\n+\tmovl\t52(%esi),%ebx\n+\tbswap\t%ecx\n+\tmovl\t%eax,44(%esp)\n+\tmovl\t56(%esi),%eax\n+\tbswap\t%ebx\n+\tmovl\t%ecx,48(%esp)\n+\tmovl\t60(%esi),%ecx\n+\tbswap\t%eax\n+\tmovl\t%ebx,52(%esp)\n+\tbswap\t%ecx\n+\tmovl\t%eax,56(%esp)\n+\tmovl\t%ecx,60(%esp)\n+\tmovl\t%esi,88(%esp)\n+\tmovl\t8(%edi),%ecx\n+\tmovl\t4(%edi),%ebx\n+\tmovl\t(%edi),%eax\n+\t/* 00_15 0 */\n+\tmovl\t%edx,%edi\n+\tmovl\t(%esp),%esi\n+\txorl\t%ecx,%edi\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1518500249(%ebp,%esi,1),%ebp\n+\tmovl\t%eax,%esi\n+\txorl\t%edx,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\taddl\t%esi,%ebp\n+\t/* 00_15 1 */\n+\tmovl\t4(%esp),%esi\n+\txorl\t%ebx,%edi\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1518500249(%edx,%esi,1),%edx\n+\tmovl\t%ebp,%esi\n+\txorl\t%ecx,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\taddl\t%esi,%edx\n+\t/* 00_15 2 */\n+\tmovl\t8(%esp),%esi\n+\txorl\t%eax,%edi\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1518500249(%ecx,%esi,1),%ecx\n+\tmovl\t%edx,%esi\n+\txorl\t%ebx,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\taddl\t%esi,%ecx\n+\t/* 00_15 3 */\n+\tmovl\t12(%esp),%esi\n+\txorl\t%ebp,%edi\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1518500249(%ebx,%esi,1),%ebx\n+\tmovl\t%ecx,%esi\n+\txorl\t%eax,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\taddl\t%esi,%ebx\n+\t/* 00_15 4 */\n+\tmovl\t16(%esp),%esi\n+\txorl\t%edx,%edi\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1518500249(%eax,%esi,1),%eax\n+\tmovl\t%ebx,%esi\n+\txorl\t%ebp,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\taddl\t%esi,%eax\n+\t/* 00_15 5 */\n+\tmovl\t20(%esp),%esi\n+\txorl\t%ecx,%edi\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1518500249(%ebp,%esi,1),%ebp\n+\tmovl\t%eax,%esi\n+\txorl\t%edx,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\taddl\t%esi,%ebp\n+\t/* 00_15 6 */\n+\tmovl\t24(%esp),%esi\n+\txorl\t%ebx,%edi\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1518500249(%edx,%esi,1),%edx\n+\tmovl\t%ebp,%esi\n+\txorl\t%ecx,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\taddl\t%esi,%edx\n+\t/* 00_15 7 */\n+\tmovl\t28(%esp),%esi\n+\txorl\t%eax,%edi\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1518500249(%ecx,%esi,1),%ecx\n+\tmovl\t%edx,%esi\n+\txorl\t%ebx,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\taddl\t%esi,%ecx\n+\t/* 00_15 8 */\n+\tmovl\t32(%esp),%esi\n+\txorl\t%ebp,%edi\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1518500249(%ebx,%esi,1),%ebx\n+\tmovl\t%ecx,%esi\n+\txorl\t%eax,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\taddl\t%esi,%ebx\n+\t/* 00_15 9 */\n+\tmovl\t36(%esp),%esi\n+\txorl\t%edx,%edi\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1518500249(%eax,%esi,1),%eax\n+\tmovl\t%ebx,%esi\n+\txorl\t%ebp,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\taddl\t%esi,%eax\n+\t/* 00_15 10 */\n+\tmovl\t40(%esp),%esi\n+\txorl\t%ecx,%edi\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1518500249(%ebp,%esi,1),%ebp\n+\tmovl\t%eax,%esi\n+\txorl\t%edx,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\taddl\t%esi,%ebp\n+\t/* 00_15 11 */\n+\tmovl\t44(%esp),%esi\n+\txorl\t%ebx,%edi\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1518500249(%edx,%esi,1),%edx\n+\tmovl\t%ebp,%esi\n+\txorl\t%ecx,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\taddl\t%esi,%edx\n+\t/* 00_15 12 */\n+\tmovl\t48(%esp),%esi\n+\txorl\t%eax,%edi\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1518500249(%ecx,%esi,1),%ecx\n+\tmovl\t%edx,%esi\n+\txorl\t%ebx,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\taddl\t%esi,%ecx\n+\t/* 00_15 13 */\n+\tmovl\t52(%esp),%esi\n+\txorl\t%ebp,%edi\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1518500249(%ebx,%esi,1),%ebx\n+\tmovl\t%ecx,%esi\n+\txorl\t%eax,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\taddl\t%esi,%ebx\n+\t/* 00_15 14 */\n+\tmovl\t56(%esp),%esi\n+\txorl\t%edx,%edi\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1518500249(%eax,%esi,1),%eax\n+\tmovl\t%ebx,%esi\n+\txorl\t%ebp,%edi\n+\troll\t$5,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\taddl\t%esi,%eax\n+\t/* 00_15 15 */\n+\tmovl\t60(%esp),%esi\n+\txorl\t%ecx,%edi\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1518500249(%ebp,%esi,1),%ebp\n+\txorl\t%edx,%edi\n+\tmovl\t(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t8(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t32(%esp),%esi\n+\t/* 16_19 16 */\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\troll\t$1,%esi\n+\txorl\t%ebx,%edi\n+\tmovl\t%esi,(%esp)\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1518500249(%edx,%esi,1),%edx\n+\tmovl\t4(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t36(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 16_19 17 */\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\troll\t$1,%esi\n+\txorl\t%eax,%edi\n+\tmovl\t%esi,4(%esp)\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1518500249(%ecx,%esi,1),%ecx\n+\tmovl\t8(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t40(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 16_19 18 */\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\troll\t$1,%esi\n+\txorl\t%ebp,%edi\n+\tmovl\t%esi,8(%esp)\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1518500249(%ebx,%esi,1),%ebx\n+\tmovl\t12(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 16_19 19 */\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\troll\t$1,%esi\n+\txorl\t%edx,%edi\n+\tmovl\t%esi,12(%esp)\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1518500249(%eax,%esi,1),%eax\n+\tmovl\t16(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 20 */\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,16(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1859775393(%ebp,%esi,1),%ebp\n+\tmovl\t20(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 21 */\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,20(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1859775393(%edx,%esi,1),%edx\n+\tmovl\t24(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 22 */\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,24(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1859775393(%ecx,%esi,1),%ecx\n+\tmovl\t28(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t60(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 23 */\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,28(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1859775393(%ebx,%esi,1),%ebx\n+\tmovl\t32(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 24 */\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,32(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1859775393(%eax,%esi,1),%eax\n+\tmovl\t36(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t4(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 25 */\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,36(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1859775393(%ebp,%esi,1),%ebp\n+\tmovl\t40(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t8(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 26 */\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,40(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1859775393(%edx,%esi,1),%edx\n+\tmovl\t44(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t12(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 27 */\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,44(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1859775393(%ecx,%esi,1),%ecx\n+\tmovl\t48(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t16(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 28 */\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,48(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1859775393(%ebx,%esi,1),%ebx\n+\tmovl\t52(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t20(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 29 */\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,52(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1859775393(%eax,%esi,1),%eax\n+\tmovl\t56(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t24(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 30 */\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,56(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1859775393(%ebp,%esi,1),%ebp\n+\tmovl\t60(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t28(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 31 */\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,60(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1859775393(%edx,%esi,1),%edx\n+\tmovl\t(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t32(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 32 */\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1859775393(%ecx,%esi,1),%ecx\n+\tmovl\t4(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t36(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 33 */\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,4(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1859775393(%ebx,%esi,1),%ebx\n+\tmovl\t8(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t40(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 34 */\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,8(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1859775393(%eax,%esi,1),%eax\n+\tmovl\t12(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 35 */\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,12(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1859775393(%ebp,%esi,1),%ebp\n+\tmovl\t16(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 36 */\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,16(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1859775393(%edx,%esi,1),%edx\n+\tmovl\t20(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 37 */\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,20(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1859775393(%ecx,%esi,1),%ecx\n+\tmovl\t24(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 38 */\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,24(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1859775393(%ebx,%esi,1),%ebx\n+\tmovl\t28(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t60(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 39 */\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,28(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1859775393(%eax,%esi,1),%eax\n+\tmovl\t32(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 40_59 40 */\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\txorl\t20(%esp),%esi\n+\tandl\t%ecx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,32(%esp)\n+\txorl\t%ecx,%edi\n+\tleal\t2400959708(%ebp,%esi,1),%ebp\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tmovl\t36(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t4(%esp),%esi\n+\t/* 40_59 41 */\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\txorl\t24(%esp),%esi\n+\tandl\t%ebx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,36(%esp)\n+\txorl\t%ebx,%edi\n+\tleal\t2400959708(%edx,%esi,1),%edx\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tmovl\t40(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t8(%esp),%esi\n+\t/* 40_59 42 */\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\txorl\t28(%esp),%esi\n+\tandl\t%eax,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,40(%esp)\n+\txorl\t%eax,%edi\n+\tleal\t2400959708(%ecx,%esi,1),%ecx\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tmovl\t44(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t12(%esp),%esi\n+\t/* 40_59 43 */\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\txorl\t32(%esp),%esi\n+\tandl\t%ebp,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,44(%esp)\n+\txorl\t%ebp,%edi\n+\tleal\t2400959708(%ebx,%esi,1),%ebx\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tmovl\t48(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t16(%esp),%esi\n+\t/* 40_59 44 */\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\txorl\t36(%esp),%esi\n+\tandl\t%edx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,48(%esp)\n+\txorl\t%edx,%edi\n+\tleal\t2400959708(%eax,%esi,1),%eax\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tmovl\t52(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t60(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t20(%esp),%esi\n+\t/* 40_59 45 */\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\txorl\t40(%esp),%esi\n+\tandl\t%ecx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,52(%esp)\n+\txorl\t%ecx,%edi\n+\tleal\t2400959708(%ebp,%esi,1),%ebp\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tmovl\t56(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t24(%esp),%esi\n+\t/* 40_59 46 */\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\txorl\t44(%esp),%esi\n+\tandl\t%ebx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,56(%esp)\n+\txorl\t%ebx,%edi\n+\tleal\t2400959708(%edx,%esi,1),%edx\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tmovl\t60(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t4(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t28(%esp),%esi\n+\t/* 40_59 47 */\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\txorl\t48(%esp),%esi\n+\tandl\t%eax,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,60(%esp)\n+\txorl\t%eax,%edi\n+\tleal\t2400959708(%ecx,%esi,1),%ecx\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tmovl\t(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t8(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t32(%esp),%esi\n+\t/* 40_59 48 */\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\txorl\t52(%esp),%esi\n+\tandl\t%ebp,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,(%esp)\n+\txorl\t%ebp,%edi\n+\tleal\t2400959708(%ebx,%esi,1),%ebx\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tmovl\t4(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t12(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t36(%esp),%esi\n+\t/* 40_59 49 */\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\txorl\t56(%esp),%esi\n+\tandl\t%edx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,4(%esp)\n+\txorl\t%edx,%edi\n+\tleal\t2400959708(%eax,%esi,1),%eax\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tmovl\t8(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t16(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t40(%esp),%esi\n+\t/* 40_59 50 */\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\txorl\t60(%esp),%esi\n+\tandl\t%ecx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,8(%esp)\n+\txorl\t%ecx,%edi\n+\tleal\t2400959708(%ebp,%esi,1),%ebp\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tmovl\t12(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t20(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t44(%esp),%esi\n+\t/* 40_59 51 */\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\txorl\t(%esp),%esi\n+\tandl\t%ebx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,12(%esp)\n+\txorl\t%ebx,%edi\n+\tleal\t2400959708(%edx,%esi,1),%edx\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tmovl\t16(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t24(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t48(%esp),%esi\n+\t/* 40_59 52 */\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\txorl\t4(%esp),%esi\n+\tandl\t%eax,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,16(%esp)\n+\txorl\t%eax,%edi\n+\tleal\t2400959708(%ecx,%esi,1),%ecx\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tmovl\t20(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t28(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t52(%esp),%esi\n+\t/* 40_59 53 */\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\txorl\t8(%esp),%esi\n+\tandl\t%ebp,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,20(%esp)\n+\txorl\t%ebp,%edi\n+\tleal\t2400959708(%ebx,%esi,1),%ebx\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tmovl\t24(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t32(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t56(%esp),%esi\n+\t/* 40_59 54 */\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\txorl\t12(%esp),%esi\n+\tandl\t%edx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,24(%esp)\n+\txorl\t%edx,%edi\n+\tleal\t2400959708(%eax,%esi,1),%eax\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tmovl\t28(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t36(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t60(%esp),%esi\n+\t/* 40_59 55 */\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\txorl\t16(%esp),%esi\n+\tandl\t%ecx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,28(%esp)\n+\txorl\t%ecx,%edi\n+\tleal\t2400959708(%ebp,%esi,1),%ebp\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tmovl\t32(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t40(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t(%esp),%esi\n+\t/* 40_59 56 */\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\txorl\t20(%esp),%esi\n+\tandl\t%ebx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,32(%esp)\n+\txorl\t%ebx,%edi\n+\tleal\t2400959708(%edx,%esi,1),%edx\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tmovl\t36(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t4(%esp),%esi\n+\t/* 40_59 57 */\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\txorl\t24(%esp),%esi\n+\tandl\t%eax,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,36(%esp)\n+\txorl\t%eax,%edi\n+\tleal\t2400959708(%ecx,%esi,1),%ecx\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tmovl\t40(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t8(%esp),%esi\n+\t/* 40_59 58 */\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\txorl\t28(%esp),%esi\n+\tandl\t%ebp,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,40(%esp)\n+\txorl\t%ebp,%edi\n+\tleal\t2400959708(%ebx,%esi,1),%ebx\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tmovl\t44(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t12(%esp),%esi\n+\t/* 40_59 59 */\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\txorl\t32(%esp),%esi\n+\tandl\t%edx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,44(%esp)\n+\txorl\t%edx,%edi\n+\tleal\t2400959708(%eax,%esi,1),%eax\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tmovl\t48(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t16(%esp),%esi\n+\t/* 20_39 60 */\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,48(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t3395469782(%ebp,%esi,1),%ebp\n+\tmovl\t52(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t20(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 61 */\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,52(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t3395469782(%edx,%esi,1),%edx\n+\tmovl\t56(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t24(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 62 */\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,56(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t3395469782(%ecx,%esi,1),%ecx\n+\tmovl\t60(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t28(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 63 */\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,60(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t3395469782(%ebx,%esi,1),%ebx\n+\tmovl\t(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t32(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 64 */\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t3395469782(%eax,%esi,1),%eax\n+\tmovl\t4(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t36(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 65 */\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,4(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t3395469782(%ebp,%esi,1),%ebp\n+\tmovl\t8(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t40(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 66 */\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,8(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t3395469782(%edx,%esi,1),%edx\n+\tmovl\t12(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 67 */\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,12(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t3395469782(%ecx,%esi,1),%ecx\n+\tmovl\t16(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 68 */\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,16(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t3395469782(%ebx,%esi,1),%ebx\n+\tmovl\t20(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 69 */\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,20(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t3395469782(%eax,%esi,1),%eax\n+\tmovl\t24(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 70 */\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,24(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t3395469782(%ebp,%esi,1),%ebp\n+\tmovl\t28(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t60(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 71 */\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,28(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t3395469782(%edx,%esi,1),%edx\n+\tmovl\t32(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 72 */\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,32(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t3395469782(%ecx,%esi,1),%ecx\n+\tmovl\t36(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t4(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 73 */\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,36(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t3395469782(%ebx,%esi,1),%ebx\n+\tmovl\t40(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t8(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 74 */\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,40(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t3395469782(%eax,%esi,1),%eax\n+\tmovl\t44(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t12(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 75 */\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,44(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t3395469782(%ebp,%esi,1),%ebp\n+\tmovl\t48(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t16(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 76 */\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,48(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t3395469782(%edx,%esi,1),%edx\n+\tmovl\t52(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t20(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 77 */\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t3395469782(%ecx,%esi,1),%ecx\n+\tmovl\t56(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t24(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 78 */\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t3395469782(%ebx,%esi,1),%ebx\n+\tmovl\t60(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t28(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 79 */\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t3395469782(%eax,%esi,1),%eax\n+\txorl\t%edx,%edi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%eax\n+\t/* Loop trailer */\n+\tmovl\t84(%esp),%edi\n+\tmovl\t88(%esp),%esi\n+\taddl\t16(%edi),%ebp\n+\taddl\t12(%edi),%edx\n+\taddl\t%ecx,8(%edi)\n+\taddl\t%ebx,4(%edi)\n+\taddl\t$64,%esi\n+\taddl\t%eax,(%edi)\n+\tmovl\t%edx,12(%edi)\n+\tmovl\t%ebp,16(%edi)\n+\tcmpl\t92(%esp),%esi\n+\tjb\t.L000loop\n+\taddl\t$64,%esp\n+\tpopl\t%edi\n+\tpopl\t%esi\n+\tpopl\t%ebx\n+\tpopl\t%ebp\n+\tret\n+.L_sha1_block_data_order_end:\n+.size\tsha1_block_data_order,.L_sha1_block_data_order_end-sha1_block_data_order\n+.byte\t83,72,65,49,32,98,108,111,99,107,32,116,114,97,110,115,102,111,114,109,32,102,111,114,32,120,56,54,44,32,67,82,89,80,84,79,71,65,77,83,32,98,121,32,60,97,112,112,114,111,64,111,112,101,110,115,115,108,46,111,114,103,62,0\ndiff --git a/x86/sha1.c b/x86/sha1.c\nnew file mode 100644\nindex 0000000..4c1a569\n--- /dev/null\n+++ b/x86/sha1.c\n@@ -0,0 +1,81 @@\n+/*\n+ * SHA-1 implementation.\n+ *\n+ * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n+ *\n+ * This version assumes we are running on a big-endian machine.\n+ * It calls an external sha1_core() to process blocks of 64 bytes.\n+ */\n+#include <stdio.h>\n+#include <string.h>\n+#include <arpa/inet.h>\t/* For htonl */\n+#include \"sha1.h\"\n+\n+#define x86_sha1_core sha1_block_data_order\n+extern void x86_sha1_core(uint32_t hash[5], const unsigned char *p,\n+\t\t\t  unsigned int nblocks);\n+\n+void x86_SHA1_Init(x86_SHA_CTX *c)\n+{\n+\t/* Matches prefix of scontext structure */\n+\tstatic struct {\n+\t\tuint32_t hash[5];\n+\t\tuint64_t len;\n+\t} const iv = {\n+\t\t{ 0x67452301, 0xEFCDAB89, 0x98BADCFE, 0x10325476, 0xC3D2E1F0 },\n+\t\t0\n+\t};\n+\n+\tmemcpy(c, &iv, sizeof iv);\n+}\n+\n+void x86_SHA1_Update(x86_SHA_CTX *c, const void *p, unsigned long n)\n+{\n+\tunsigned pos = (unsigned)c->len & 63;\n+\tunsigned long nb;\n+\n+\tc->len += n;\n+\n+\t/* Initial partial block */\n+\tif (pos) {\n+\t\tunsigned space = 64 - pos;\n+\t\tif (space > n)\n+\t\t\tgoto end;\n+\t\tmemcpy(c->buf + pos, p, space);\n+\t\tp += space;\n+\t\tn -= space;\n+\t\tx86_sha1_core(c->hash, c->buf, 1);\n+\t}\n+\n+\t/* The big impressive middle */\n+\tnb = n >> 6;\n+\tif (nb) {\n+\t\tx86_sha1_core(c->hash, p, nb);\n+\t\tp += nb << 6;\n+\t\tn &= 63;\n+\t}\n+\tpos = 0;\n+end:\n+\t/* Final partial block */\n+\tmemcpy(c->buf + pos, p, n);\n+}\n+\n+void x86_SHA1_Final(unsigned char *hash, x86_SHA_CTX *c)\n+{\n+\tunsigned pos = (unsigned)c->len & 63;\n+\n+\tc->buf[pos++] = 0x80;\n+\tif (pos > 56) {\n+\t\tmemset(c->buf + pos, 0, 64 - pos);\n+\t\tx86_sha1_core(c->hash, c->buf, 1);\n+\t\tpos = 0;\n+\t}\n+\tmemset(c->buf + pos, 0, 56 - pos);\n+\t/* Last two words are 64-bit *bit* count */\n+\t*(uint32_t *)(c->buf + 56) = htonl((uint32_t)(c->len >> 29));\n+\t*(uint32_t *)(c->buf + 60) = htonl((uint32_t)c->len << 3);\n+\tx86_sha1_core(c->hash, c->buf, 1);\n+\n+\tfor (pos = 0; pos < 5; pos++)\n+\t\t((uint32_t *)hash)[pos] = htonl(c->hash[pos]);\n+}\ndiff --git a/x86/sha1.h b/x86/sha1.h\nnew file mode 100644\nindex 0000000..8988da9\n--- /dev/null\n+++ b/x86/sha1.h\n@@ -0,0 +1,21 @@\n+/*\n+ * SHA-1 implementation.\n+ *\n+ * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n+ */\n+#include <stdint.h>\n+\n+typedef struct {\n+\tuint32_t hash[5];\n+\tuint64_t len;\n+\tunsigned char buf[64];\t/* Keep this aligned */\n+} x86_SHA_CTX;\n+\n+void x86_SHA1_Init(x86_SHA_CTX *c);\n+void x86_SHA1_Update(x86_SHA_CTX *c, const void *p, unsigned long n);\n+void x86_SHA1_Final(unsigned char *hash, x86_SHA_CTX *c);\n+\n+#define git_SHA_CTX\tx86_SHA_CTX\n+#define git_SHA1_Init\tx86_SHA1_Init\n+#define git_SHA1_Update\tx86_SHA1_Update\n+#define git_SHA1_Final\tx86_SHA1_Final\n"},{"id":"119458","messageId":"20090804050138.13256.qmail@science.horizon.com","threadId":"20241","inReplyTo":"9e4733910908032007td74ef9fp669d0d958df67c1@mail.gmail.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-08-04T05:01:38Z","receivedAt":"2009-08-04T05:01:38Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> Would there happen to be a SHA1 implementation around that can compute\n> the SHA1 without first decompressing the data? Databases gain a lot of\n> speed by using special algorithms that can directly operate on the\n> compressed data.\n\nI can't imagine how.  In general, this requires that the compression\nbe carefully designed to be compatible with the algorithms, and SHA1\nis specifically designed to depend on every bit of the input in\nan un-analyzable way.\n\nAlso, git normally avoids hashing objects that it doesn't need\nuncompressed for some other reason.  git-fsck is a notable exception,\nbut I think the idea of creating special optimized code paths for that\ninterferes with its reliability and robustness goals.\n"},{"id":"119465","messageId":"alpine.LFD.2.01.0908032326440.3270@localhost.localdomain","threadId":"20241","inReplyTo":"20090804044842.6792.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-04T06:30:25Z","receivedAt":"2009-08-04T06:30:25Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 4 Aug 2009, George Spelvin wrote:\n> \n> The actual goal of this effort is to address the dynamic linker startup\n> time issues by removing the second-largest contributor after libcurl,\n> namely openssl.  Optimizing the assembly code is just the fun part. ;-)\n\nNow, I agree that it would be wonderful to get rid of the linker startup, \nbut the startup costs of openssl are very low compared to the equivalent \ncurl ones. So we can't lose _too_ much performance - especially for \nlong-running jobs where startup costs really don't even matter - in the \nquest to get rid of those.\n\nThat said, your numbers are impressive. Improving fsck by 1.1-2.2% is very \ngood. That means that you not only avodied the startup costs, you actually \nimproved on the openssl code. So it's a win-win situation.\n\nThat said, it would be even better if the SHA1 code was also somewhat \nportable to other environments (it looks like your current patch is very \nGNU as specific), and if you had a solution for x86-64 too ;)\n\nYeah, I'm a whiny little b*tch, aren't I?\n\n\t\tLinus\n"},{"id":"119468","messageId":"alpine.LFD.2.01.0908032336460.3270@localhost.localdomain","threadId":"20241","inReplyTo":"20090804044842.6792.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-04T06:40:35Z","receivedAt":"2009-08-04T06:40:35Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 4 Aug 2009, George Spelvin wrote:\n> +sha1_block_data_order:\n> +\tpushl\t%ebp\n> +\tpushl\t%ebx\n> +\tpushl\t%esi\n> +\tpushl\t%edi\n> +\tmovl\t20(%esp),%edi\n> +\tmovl\t24(%esp),%esi\n> +\tmovl\t28(%esp),%eax\n> +\tsubl\t$64,%esp\n> +\tshll\t$6,%eax\n> +\taddl\t%esi,%eax\n> +\tmovl\t%eax,92(%esp)\n> +\tmovl\t16(%edi),%ebp\n> +\tmovl\t12(%edi),%edx\n> +.align\t16\n> +.L000loop:\n> +\tmovl\t(%esi),%ecx\n> +\tmovl\t4(%esi),%ebx\n> +\tbswap\t%ecx\n> +\tmovl\t8(%esi),%eax\n> +\tbswap\t%ebx\n> +\tmovl\t%ecx,(%esp)\n\n...\n\nHmm. Does it really help to do the bswap as a separate initial phase?\n\nAs far as I can tell, you load the result of the bswap just a single time \nfor each value. So the initial \"bswap all 64 bytes\" seems pointless.\n\n> +\t/* 00_15 0 */\n> +\tmovl\t%edx,%edi\n> +\tmovl\t(%esp),%esi\n\nWhy not do the bswap here instead?\n\nIs it because you're running out of registers for scheduling, and want to \nuse the stack pointer rather than the original source?\n\nOr does the data dependency end up being so much better that you're better \noff doing a separate bswap loop?\n\nOr is it just because the code was written that way?\n\nIntriguing, either way.\n\n\t\tLinus\n"},{"id":"119479","messageId":"20090804080147.15201.qmail@science.horizon.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908032326440.3270@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-08-04T08:01:47Z","receivedAt":"2009-08-04T08:01:47Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> Now, I agree that it would be wonderful to get rid of the linker startup, \n> but the startup costs of openssl are very low compared to the equivalent \n> curl ones. So we can't lose _too_ much performance - especially for \n> long-running jobs where startup costs really don't even matter - in the \n> quest to get rid of those.\n>\n> That said, your numbers are impressive. Improving fsck by 1.1-2.2% is very \n> good. That means that you not only avodied the startup costs, you actually \n> improved on the openssl code. So it's a win-win situation.\n\nEr, yes, that *is* what the subject line is advertising.  I started\nwith the OpenSSL core SHA1 code (which is BSD/GPL dual-licensed by its\nauthor) and tweaked it some more for more recent processors.\n\n> That said, it would be even better if the SHA1 code was also somewhat \n> portable to other environments (it looks like your current patch is very \n> GNU as specific), and if you had a solution for x86-64 too ;)\n\nDone and will be done.\n\nThe code is *actually* written (see the first e-mail in this thread)\nin the perl-preprocessor that OpenSSL uses, which can generate quite a\nfew output syntaxes (including Intel).  I just included the preprocessed\nversion to reduce the complexity of the rough-draft patch.\n\nThe one question I have is that currently perl is not a critical\ncompile-time dependency; it's needed for some extra stuff, but AFAIK you\ncan get most of git working without it.  Whether to add that dependency\nor what is a Junio question.\n\nAs for x86-64, I haven't actually *written* it yet, but it'll be a very\nsimple adaptation.  Mostly it's just a matter of using the additional\nregisters effectively.\n\n> Yeah, I'm a whiny little b*tch, aren't I?\n\nNot at all; I expected all of that.  Getting rid of OpenSSL kind of\nrequires those things.\n\n> Hmm. Does it really help to do the bswap as a separate initial phase?\n> \n> As far as I can tell, you load the result of the bswap just a single time \n> for each value. So the initial \"bswap all 64 bytes\" seems pointless.\n\n>> +\t/* 00_15 0 */\n>> +\tmovl\t%edx,%edi\n>> +\tmovl\t(%esp),%esi\n\n> Why not do the bswap here instead?\n>\n> Is it because you're running out of registers for scheduling, and want to \n> use the stack pointer rather than the original source?\n\nExactly.  I looked hard at it, but that means that I'd have to write the\nfirst 16 rounds with only one temp register, because the other is being\nused as an input pointer.\n\nHere's the pipelined loop for the first 16 rounds (when in[i] is the\nstack buffer), showing parallel operations on the same line.\n(Operations in parens belong to adjacent rounds.)\n#                       movl D,S        (roll 5,T)      (addl S,A)      //\n#       mov in[i],T     xorl C,S        (addl T,A)\n#                       andl B,S                        rorl 2,B\n#       addl T+K,E      xorl D,S        movl A,T\n#                       addl S,E        roll 5,T        (movl C,S)      //\n#       (mov in[i],T)   (xorl B,S)      addl T,E        \n\nwhich translates in perl code to:\n\nsub BODY_00_15\n{\n        local($n,$a,$b,$c,$d,$e)=@_;\n\n        &comment(\"00_15 $n\");\n                &mov($S,$d) if ($n == 0);\n        &mov($T,&swtmp($n%16));         #  V Load Xi.\n                &xor($S,$c);            # U  Continue F() = d^(b&(c^d))\n                &and($S,$b);            #  V\n                        &rotr($b,2);    # NP\n        &lea($e,&DWP(K1,$e,$T));        # U  Add Xi and K\n    if ($n < 15) {\n                        &mov($T,$a);    #  V\n                &xor($S,$d);            # U \n                        &rotl($T,5);    # NP\n                &add($e,$S);            # U \n                &mov($S,$c);            #  V Start of NEXT round's F()\n                        &add($e,$T);    # U \n    } else {\n        # This version provides the correct start for BODY_20_39\n                &xor($S,$d);            #  V\n        &mov($T,&swtmp(($n+1)%16));     # U  Start computing mext Xi.\n                &add($e,$S);            #  V Add F()\n                        &mov($S,$a);    # U  Start computing a<<<5\n        &xor($T,&swtmp(($n+3)%16));     #  V\n                        &rotl($S,5);    # U \n        &xor($T,&swtmp(($n+9)%16));     #  V\n    }\n}\n\nAnyway, the round is:\n\n#define K1 0x5a827999\ne += bswap(in[i]) + K1 + (d^(b&(c^d))) + ROTL(a,5).\nb = ROTR(b,2);\n\nNotice how I use one temp (T) for in[i] and ROTL(a,5), and the other (S)\nfor F1(b,c,d) = d^(b&(c^d)).\n\nIf I only had one temporary, I'd have to seriously un-overlap it:\n\tmov\tS[i],T\n\tbswap\tT\n\tmov\tT,in[i]\n\tlea\tK1(T,e),e\n\t  mov\t  d,T\n\t  xor\t  c,T\n\t  and\t  b,T\n\t  xor\t  d,T\n\t  add\t  T,e\n\tmov\ta,T\n\troll\t5,T\n\tadd\tT,e\n\nCurrent processors probably have enough out-of-order scheduling resources to\nfind the parallelism there, but something like an Atom would be doomed.\n\nI just cobbled together a test implementation, and it looks pretty similar\non my Phenom here (minimum of 30 runs):\n\nSeparate copy loop: 1.355603\nIn-line:            1.350444 (+0.4% faster)\n\nA hint of being faster, but not much.\n\nIt is a couple of percent faster on a P4:\nSeparate copy loop: 3.297174\nIn-line:            3.237354 (+1.8% faster)\n\nAnd on an i7:\nSeparate copy loop: 1.353641\nIn-line:            1.336766 (+1.2% faster)\n\nbut I worry about in-order machines.  An Athlon XP:\nSeparate copy loop: 3.252682\nIn-line:            3.313870 (-1.8% slower)\n\nH'm... it's not bad.  And the code is smaller.  Maybe I'll work on\nit a bit.\n\nIf you want to try it, the modified sha1-x86.s file is appended.\n\n--- /dev/null\t2009-05-12 02:55:38.579106460 -0400\n+++ sha1-x86.s\t2009-08-04 03:42:31.073284734 -0400\n@@ -0,0 +1,1359 @@\n+.file\t\"sha1-586.s\"\n+.text\n+.globl\tsha1_block_data_order\n+.type\tsha1_block_data_order,@function\n+.align\t16\n+sha1_block_data_order:\n+\tpushl\t%ebp\n+\tpushl\t%ebx\n+\tpushl\t%esi\n+\tpushl\t%edi\n+\tmovl\t20(%esp),%edi\n+\tmovl\t24(%esp),%esi\n+\tmovl\t28(%esp),%eax\n+\tsubl\t$64,%esp\n+\tshll\t$6,%eax\n+\taddl\t%esi,%eax\n+\tmovl\t%eax,92(%esp)\n+\tmovl\t16(%edi),%ebp\n+\tmovl\t12(%edi),%edx\n+\tmovl\t8(%edi),%ecx\n+\tmovl\t4(%edi),%ebx\n+\tmovl\t(%edi),%eax\n+.align\t16\n+.L000loop:\n+\tmovl\t%esi,88(%esp)\n+\t/* 00_15 0 */\n+\tmovl\t(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,(%esp)\n+\tleal\t1518500249(%ebp,%edi,1),%ebp\n+\tmovl\t%edx,%edi\n+\txorl\t%ecx,%edi\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\txorl\t%edx,%edi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%ebp\n+\t/* 00_15 1 */\n+\tmovl\t4(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,4(%esp)\n+\tleal\t1518500249(%edx,%edi,1),%edx\n+\tmovl\t%ecx,%edi\n+\txorl\t%ebx,%edi\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\txorl\t%ecx,%edi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%edx\n+\t/* 00_15 2 */\n+\tmovl\t8(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,8(%esp)\n+\tleal\t1518500249(%ecx,%edi,1),%ecx\n+\tmovl\t%ebx,%edi\n+\txorl\t%eax,%edi\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\txorl\t%ebx,%edi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%ecx\n+\t/* 00_15 3 */\n+\tmovl\t12(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,12(%esp)\n+\tleal\t1518500249(%ebx,%edi,1),%ebx\n+\tmovl\t%eax,%edi\n+\txorl\t%ebp,%edi\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\txorl\t%eax,%edi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%ebx\n+\t/* 00_15 4 */\n+\tmovl\t16(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,16(%esp)\n+\tleal\t1518500249(%eax,%edi,1),%eax\n+\tmovl\t%ebp,%edi\n+\txorl\t%edx,%edi\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\txorl\t%ebp,%edi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%eax\n+\t/* 00_15 5 */\n+\tmovl\t20(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,20(%esp)\n+\tleal\t1518500249(%ebp,%edi,1),%ebp\n+\tmovl\t%edx,%edi\n+\txorl\t%ecx,%edi\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\txorl\t%edx,%edi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%ebp\n+\t/* 00_15 6 */\n+\tmovl\t24(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,24(%esp)\n+\tleal\t1518500249(%edx,%edi,1),%edx\n+\tmovl\t%ecx,%edi\n+\txorl\t%ebx,%edi\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\txorl\t%ecx,%edi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%edx\n+\t/* 00_15 7 */\n+\tmovl\t28(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,28(%esp)\n+\tleal\t1518500249(%ecx,%edi,1),%ecx\n+\tmovl\t%ebx,%edi\n+\txorl\t%eax,%edi\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\txorl\t%ebx,%edi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%ecx\n+\t/* 00_15 8 */\n+\tmovl\t32(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,32(%esp)\n+\tleal\t1518500249(%ebx,%edi,1),%ebx\n+\tmovl\t%eax,%edi\n+\txorl\t%ebp,%edi\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\txorl\t%eax,%edi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%ebx\n+\t/* 00_15 9 */\n+\tmovl\t36(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,36(%esp)\n+\tleal\t1518500249(%eax,%edi,1),%eax\n+\tmovl\t%ebp,%edi\n+\txorl\t%edx,%edi\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\txorl\t%ebp,%edi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%eax\n+\t/* 00_15 10 */\n+\tmovl\t40(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,40(%esp)\n+\tleal\t1518500249(%ebp,%edi,1),%ebp\n+\tmovl\t%edx,%edi\n+\txorl\t%ecx,%edi\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\txorl\t%edx,%edi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%ebp\n+\t/* 00_15 11 */\n+\tmovl\t44(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,44(%esp)\n+\tleal\t1518500249(%edx,%edi,1),%edx\n+\tmovl\t%ecx,%edi\n+\txorl\t%ebx,%edi\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\txorl\t%ecx,%edi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%edx\n+\t/* 00_15 12 */\n+\tmovl\t48(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,48(%esp)\n+\tleal\t1518500249(%ecx,%edi,1),%ecx\n+\tmovl\t%ebx,%edi\n+\txorl\t%eax,%edi\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\txorl\t%ebx,%edi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%ecx\n+\t/* 00_15 13 */\n+\tmovl\t52(%esi),%edi\n+\tbswap\t%edi\n+\tmovl\t%edi,52(%esp)\n+\tleal\t1518500249(%ebx,%edi,1),%ebx\n+\tmovl\t%eax,%edi\n+\txorl\t%ebp,%edi\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\txorl\t%eax,%edi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%ebx\n+\t/* 00_15 14 */\n+\tmovl\t56(%esi),%edi\n+\tmovl\t60(%esi),%esi\n+\tbswap\t%edi\n+\tmovl\t%edi,56(%esp)\n+\tleal\t1518500249(%eax,%edi,1),%eax\n+\tmovl\t%ebp,%edi\n+\txorl\t%edx,%edi\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\txorl\t%ebp,%edi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%eax\n+\t/* 00_15 15 */\n+\tmovl\t%edx,%edi\n+\tbswap\t%esi\n+\txorl\t%ecx,%edi\n+\tmovl\t%esi,60(%esp)\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\txorl\t%edx,%edi\n+\tleal\t1518500249(%ebp,%esi,1),%ebp\n+\tmovl\t(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t8(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t32(%esp),%esi\n+\t/* 16_19 16 */\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\troll\t$1,%esi\n+\txorl\t%ebx,%edi\n+\tmovl\t%esi,(%esp)\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1518500249(%edx,%esi,1),%edx\n+\tmovl\t4(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t36(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 16_19 17 */\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\troll\t$1,%esi\n+\txorl\t%eax,%edi\n+\tmovl\t%esi,4(%esp)\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1518500249(%ecx,%esi,1),%ecx\n+\tmovl\t8(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t40(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 16_19 18 */\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\troll\t$1,%esi\n+\txorl\t%ebp,%edi\n+\tmovl\t%esi,8(%esp)\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1518500249(%ebx,%esi,1),%ebx\n+\tmovl\t12(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 16_19 19 */\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\troll\t$1,%esi\n+\txorl\t%edx,%edi\n+\tmovl\t%esi,12(%esp)\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1518500249(%eax,%esi,1),%eax\n+\tmovl\t16(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 20 */\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,16(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1859775393(%ebp,%esi,1),%ebp\n+\tmovl\t20(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 21 */\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,20(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1859775393(%edx,%esi,1),%edx\n+\tmovl\t24(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 22 */\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,24(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1859775393(%ecx,%esi,1),%ecx\n+\tmovl\t28(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t60(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 23 */\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,28(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1859775393(%ebx,%esi,1),%ebx\n+\tmovl\t32(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 24 */\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,32(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1859775393(%eax,%esi,1),%eax\n+\tmovl\t36(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t4(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 25 */\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,36(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1859775393(%ebp,%esi,1),%ebp\n+\tmovl\t40(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t8(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 26 */\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,40(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1859775393(%edx,%esi,1),%edx\n+\tmovl\t44(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t12(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 27 */\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,44(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1859775393(%ecx,%esi,1),%ecx\n+\tmovl\t48(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t16(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 28 */\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,48(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1859775393(%ebx,%esi,1),%ebx\n+\tmovl\t52(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t20(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 29 */\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,52(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1859775393(%eax,%esi,1),%eax\n+\tmovl\t56(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t24(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 30 */\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,56(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1859775393(%ebp,%esi,1),%ebp\n+\tmovl\t60(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t28(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 31 */\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,60(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1859775393(%edx,%esi,1),%edx\n+\tmovl\t(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t32(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 32 */\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1859775393(%ecx,%esi,1),%ecx\n+\tmovl\t4(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t36(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 33 */\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,4(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1859775393(%ebx,%esi,1),%ebx\n+\tmovl\t8(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t40(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 34 */\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,8(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1859775393(%eax,%esi,1),%eax\n+\tmovl\t12(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 35 */\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,12(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t1859775393(%ebp,%esi,1),%ebp\n+\tmovl\t16(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 36 */\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,16(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t1859775393(%edx,%esi,1),%edx\n+\tmovl\t20(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 37 */\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,20(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t1859775393(%ecx,%esi,1),%ecx\n+\tmovl\t24(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 38 */\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,24(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t1859775393(%ebx,%esi,1),%ebx\n+\tmovl\t28(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t60(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 39 */\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,28(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t1859775393(%eax,%esi,1),%eax\n+\tmovl\t32(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 40_59 40 */\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\txorl\t20(%esp),%esi\n+\tandl\t%ecx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,32(%esp)\n+\txorl\t%ecx,%edi\n+\tleal\t2400959708(%ebp,%esi,1),%ebp\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tmovl\t36(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t4(%esp),%esi\n+\t/* 40_59 41 */\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\txorl\t24(%esp),%esi\n+\tandl\t%ebx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,36(%esp)\n+\txorl\t%ebx,%edi\n+\tleal\t2400959708(%edx,%esi,1),%edx\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tmovl\t40(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t8(%esp),%esi\n+\t/* 40_59 42 */\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\txorl\t28(%esp),%esi\n+\tandl\t%eax,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,40(%esp)\n+\txorl\t%eax,%edi\n+\tleal\t2400959708(%ecx,%esi,1),%ecx\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tmovl\t44(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t12(%esp),%esi\n+\t/* 40_59 43 */\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\txorl\t32(%esp),%esi\n+\tandl\t%ebp,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,44(%esp)\n+\txorl\t%ebp,%edi\n+\tleal\t2400959708(%ebx,%esi,1),%ebx\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tmovl\t48(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t16(%esp),%esi\n+\t/* 40_59 44 */\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\txorl\t36(%esp),%esi\n+\tandl\t%edx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,48(%esp)\n+\txorl\t%edx,%edi\n+\tleal\t2400959708(%eax,%esi,1),%eax\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tmovl\t52(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t60(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t20(%esp),%esi\n+\t/* 40_59 45 */\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\txorl\t40(%esp),%esi\n+\tandl\t%ecx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,52(%esp)\n+\txorl\t%ecx,%edi\n+\tleal\t2400959708(%ebp,%esi,1),%ebp\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tmovl\t56(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t24(%esp),%esi\n+\t/* 40_59 46 */\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\txorl\t44(%esp),%esi\n+\tandl\t%ebx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,56(%esp)\n+\txorl\t%ebx,%edi\n+\tleal\t2400959708(%edx,%esi,1),%edx\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tmovl\t60(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t4(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t28(%esp),%esi\n+\t/* 40_59 47 */\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\txorl\t48(%esp),%esi\n+\tandl\t%eax,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,60(%esp)\n+\txorl\t%eax,%edi\n+\tleal\t2400959708(%ecx,%esi,1),%ecx\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tmovl\t(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t8(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t32(%esp),%esi\n+\t/* 40_59 48 */\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\txorl\t52(%esp),%esi\n+\tandl\t%ebp,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,(%esp)\n+\txorl\t%ebp,%edi\n+\tleal\t2400959708(%ebx,%esi,1),%ebx\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tmovl\t4(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t12(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t36(%esp),%esi\n+\t/* 40_59 49 */\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\txorl\t56(%esp),%esi\n+\tandl\t%edx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,4(%esp)\n+\txorl\t%edx,%edi\n+\tleal\t2400959708(%eax,%esi,1),%eax\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tmovl\t8(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t16(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t40(%esp),%esi\n+\t/* 40_59 50 */\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\txorl\t60(%esp),%esi\n+\tandl\t%ecx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,8(%esp)\n+\txorl\t%ecx,%edi\n+\tleal\t2400959708(%ebp,%esi,1),%ebp\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tmovl\t12(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t20(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t44(%esp),%esi\n+\t/* 40_59 51 */\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\txorl\t(%esp),%esi\n+\tandl\t%ebx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,12(%esp)\n+\txorl\t%ebx,%edi\n+\tleal\t2400959708(%edx,%esi,1),%edx\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tmovl\t16(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t24(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t48(%esp),%esi\n+\t/* 40_59 52 */\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\txorl\t4(%esp),%esi\n+\tandl\t%eax,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,16(%esp)\n+\txorl\t%eax,%edi\n+\tleal\t2400959708(%ecx,%esi,1),%ecx\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tmovl\t20(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t28(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t52(%esp),%esi\n+\t/* 40_59 53 */\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\txorl\t8(%esp),%esi\n+\tandl\t%ebp,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,20(%esp)\n+\txorl\t%ebp,%edi\n+\tleal\t2400959708(%ebx,%esi,1),%ebx\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tmovl\t24(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t32(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t56(%esp),%esi\n+\t/* 40_59 54 */\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\txorl\t12(%esp),%esi\n+\tandl\t%edx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,24(%esp)\n+\txorl\t%edx,%edi\n+\tleal\t2400959708(%eax,%esi,1),%eax\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tmovl\t28(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t36(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t60(%esp),%esi\n+\t/* 40_59 55 */\n+\taddl\t%edi,%eax\n+\tmovl\t%edx,%edi\n+\txorl\t16(%esp),%esi\n+\tandl\t%ecx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,28(%esp)\n+\txorl\t%ecx,%edi\n+\tleal\t2400959708(%ebp,%esi,1),%ebp\n+\tandl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tmovl\t32(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t40(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t(%esp),%esi\n+\t/* 40_59 56 */\n+\taddl\t%edi,%ebp\n+\tmovl\t%ecx,%edi\n+\txorl\t20(%esp),%esi\n+\tandl\t%ebx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,32(%esp)\n+\txorl\t%ebx,%edi\n+\tleal\t2400959708(%edx,%esi,1),%edx\n+\tandl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tmovl\t36(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t4(%esp),%esi\n+\t/* 40_59 57 */\n+\taddl\t%edi,%edx\n+\tmovl\t%ebx,%edi\n+\txorl\t24(%esp),%esi\n+\tandl\t%eax,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,36(%esp)\n+\txorl\t%eax,%edi\n+\tleal\t2400959708(%ecx,%esi,1),%ecx\n+\tandl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tmovl\t40(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t8(%esp),%esi\n+\t/* 40_59 58 */\n+\taddl\t%edi,%ecx\n+\tmovl\t%eax,%edi\n+\txorl\t28(%esp),%esi\n+\tandl\t%ebp,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,40(%esp)\n+\txorl\t%ebp,%edi\n+\tleal\t2400959708(%ebx,%esi,1),%ebx\n+\tandl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tmovl\t44(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t12(%esp),%esi\n+\t/* 40_59 59 */\n+\taddl\t%edi,%ebx\n+\tmovl\t%ebp,%edi\n+\txorl\t32(%esp),%esi\n+\tandl\t%edx,%edi\n+\troll\t$1,%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,44(%esp)\n+\txorl\t%edx,%edi\n+\tleal\t2400959708(%eax,%esi,1),%eax\n+\tandl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tmovl\t48(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\txorl\t16(%esp),%esi\n+\t/* 20_39 60 */\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,48(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t3395469782(%ebp,%esi,1),%ebp\n+\tmovl\t52(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t20(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 61 */\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,52(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t3395469782(%edx,%esi,1),%edx\n+\tmovl\t56(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t24(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 62 */\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,56(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t3395469782(%ecx,%esi,1),%ecx\n+\tmovl\t60(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t28(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 63 */\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,60(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t3395469782(%ebx,%esi,1),%ebx\n+\tmovl\t(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t32(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 64 */\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t3395469782(%eax,%esi,1),%eax\n+\tmovl\t4(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t36(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 65 */\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,4(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t3395469782(%ebp,%esi,1),%ebp\n+\tmovl\t8(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t40(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 66 */\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,8(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t3395469782(%edx,%esi,1),%edx\n+\tmovl\t12(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t44(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 67 */\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,12(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t3395469782(%ecx,%esi,1),%ecx\n+\tmovl\t16(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t48(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 68 */\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,16(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t3395469782(%ebx,%esi,1),%ebx\n+\tmovl\t20(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t52(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 69 */\n+\txorl\t8(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,20(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t3395469782(%eax,%esi,1),%eax\n+\tmovl\t24(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t56(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 70 */\n+\txorl\t12(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,24(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t3395469782(%ebp,%esi,1),%ebp\n+\tmovl\t28(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t60(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 71 */\n+\txorl\t16(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,28(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t3395469782(%edx,%esi,1),%edx\n+\tmovl\t32(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 72 */\n+\txorl\t20(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\tmovl\t%esi,32(%esp)\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t3395469782(%ecx,%esi,1),%ecx\n+\tmovl\t36(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t4(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 73 */\n+\txorl\t24(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\tmovl\t%esi,36(%esp)\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t3395469782(%ebx,%esi,1),%ebx\n+\tmovl\t40(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t8(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 74 */\n+\txorl\t28(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\tmovl\t%esi,40(%esp)\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t3395469782(%eax,%esi,1),%eax\n+\tmovl\t44(%esp),%esi\n+\txorl\t%edx,%edi\n+\txorl\t52(%esp),%esi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\txorl\t12(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 75 */\n+\txorl\t32(%esp),%esi\n+\taddl\t%edi,%eax\n+\troll\t$1,%esi\n+\tmovl\t%edx,%edi\n+\tmovl\t%esi,44(%esp)\n+\txorl\t%ebx,%edi\n+\trorl\t$2,%ebx\n+\tleal\t3395469782(%ebp,%esi,1),%ebp\n+\tmovl\t48(%esp),%esi\n+\txorl\t%ecx,%edi\n+\txorl\t56(%esp),%esi\n+\taddl\t%edi,%ebp\n+\tmovl\t%eax,%edi\n+\txorl\t16(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 76 */\n+\txorl\t36(%esp),%esi\n+\taddl\t%edi,%ebp\n+\troll\t$1,%esi\n+\tmovl\t%ecx,%edi\n+\tmovl\t%esi,48(%esp)\n+\txorl\t%eax,%edi\n+\trorl\t$2,%eax\n+\tleal\t3395469782(%edx,%esi,1),%edx\n+\tmovl\t52(%esp),%esi\n+\txorl\t%ebx,%edi\n+\txorl\t60(%esp),%esi\n+\taddl\t%edi,%edx\n+\tmovl\t%ebp,%edi\n+\txorl\t20(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 77 */\n+\txorl\t40(%esp),%esi\n+\taddl\t%edi,%edx\n+\troll\t$1,%esi\n+\tmovl\t%ebx,%edi\n+\txorl\t%ebp,%edi\n+\trorl\t$2,%ebp\n+\tleal\t3395469782(%ecx,%esi,1),%ecx\n+\tmovl\t56(%esp),%esi\n+\txorl\t%eax,%edi\n+\txorl\t(%esp),%esi\n+\taddl\t%edi,%ecx\n+\tmovl\t%edx,%edi\n+\txorl\t24(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 78 */\n+\txorl\t44(%esp),%esi\n+\taddl\t%edi,%ecx\n+\troll\t$1,%esi\n+\tmovl\t%eax,%edi\n+\txorl\t%edx,%edi\n+\trorl\t$2,%edx\n+\tleal\t3395469782(%ebx,%esi,1),%ebx\n+\tmovl\t60(%esp),%esi\n+\txorl\t%ebp,%edi\n+\txorl\t4(%esp),%esi\n+\taddl\t%edi,%ebx\n+\tmovl\t%ecx,%edi\n+\txorl\t28(%esp),%esi\n+\troll\t$5,%edi\n+\t/* 20_39 79 */\n+\txorl\t48(%esp),%esi\n+\taddl\t%edi,%ebx\n+\troll\t$1,%esi\n+\tmovl\t%ebp,%edi\n+\txorl\t%ecx,%edi\n+\trorl\t$2,%ecx\n+\tleal\t3395469782(%eax,%esi,1),%eax\n+\txorl\t%edx,%edi\n+\taddl\t%edi,%eax\n+\tmovl\t%ebx,%edi\n+\troll\t$5,%edi\n+\taddl\t%edi,%eax\n+\t/* Loop trailer */\n+\tmovl\t84(%esp),%edi\n+\tmovl\t88(%esp),%esi\n+\taddl\t16(%edi),%ebp\n+\taddl\t12(%edi),%edx\n+\taddl\t8(%edi),%ecx\n+\taddl\t4(%edi),%ebx\n+\taddl\t(%edi),%eax\n+\taddl\t$64,%esi\n+\tmovl\t%ebp,16(%edi)\n+\tmovl\t%edx,12(%edi)\n+\tcmpl\t92(%esp),%esi\n+\tmovl\t%ecx,8(%edi)\n+\tmovl\t%ebx,4(%edi)\n+\tmovl\t%eax,(%edi)\n+\tjb\t.L000loop\n+\taddl\t$64,%esp\n+\tpopl\t%edi\n+\tpopl\t%esi\n+\tpopl\t%ebx\n+\tpopl\t%ebp\n+\tret\n+.L_sha1_block_data_order_end:\n+.size\tsha1_block_data_order,.L_sha1_block_data_order_end-sha1_block_data_order\n+.byte\t83,72,65,49,32,98,108,111,99,107,32,116,114,97,110,115,102,111,114,109,32,102,111,114,32,120,56,54,44,32,67,82,89,80,84,79,71,65,77,83,32,98,121,32,60,97,112,112,114,111,64,111,112,101,110,115,115,108,46,111,114,103,62,0\n"},{"id":"119493","messageId":"9e4733910908040556t477dcb66u86a43e12148e3352@mail.gmail.com","threadId":"20241","inReplyTo":"20090804050138.13256.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2009-08-04T12:56:48Z","receivedAt":"2009-08-04T12:56:48Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On Tue, Aug 4, 2009 at 1:01 AM, George Spelvin<linux@horizon.com> wrote:\n>> Would there happen to be a SHA1 implementation around that can compute\n>> the SHA1 without first decompressing the data? Databases gain a lot of\n>> speed by using special algorithms that can directly operate on the\n>> compressed data.\n>\n> I can't imagine how.  In general, this requires that the compression\n> be carefully designed to be compatible with the algorithms, and SHA1\n> is specifically designed to depend on every bit of the input in\n> an un-analyzable way.\n\nA simple start would be to feed each byte as it is decompressed\ndirectly into the sha code and avoid the intermediate buffer. Removing\nthe buffer reduces cache pressure.\n\n> Also, git normally avoids hashing objects that it doesn't need\n> uncompressed for some other reason.  git-fsck is a notable exception,\n> but I think the idea of creating special optimized code paths for that\n> interferes with its reliability and robustness goals.\n\nAgreed that there is no real need for this, just something to play\nwith if you are trying for a speed record.\n\nI'd much rather have a solution for the rebase problem where one side\nof the diff has moved to a different file and rebase can't figure it\nout.\n\n>\n\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"119500","messageId":"20090804142942.GE10264@dpotapov.dyndns.org","threadId":"20241","inReplyTo":"9e4733910908040556t477dcb66u86a43e12148e3352@mail.gmail.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2009-08-04T14:29:42Z","receivedAt":"2009-08-04T14:29:42Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Tue, Aug 04, 2009 at 08:56:48AM -0400, Jon Smirl wrote:\n> \n> A simple start would be to feed each byte as it is decompressed\n> directly into the sha code and avoid the intermediate buffer. Removing\n> the buffer reduces cache pressure.\n\nFirst, you still have to preserve any decoded byte in the compress\nwindow, which is 32Kb by default. Typical files in Git repositories are\nnot so big, many are under 32Kb and practically all of them fit to L2\ncache of modern processors. Second, complication of assembler code from\nthe coupling of two algorithms will be enormous. It is not sufficient\nregisters on x86 for SHA-1 alone. Third, SHA-1 is very computationally\nintensive and with predictable access pattern (linear), so you do not\nwait for L2, because it will be in L1. So, I don't see where you can\ngain significantly. Perhaps, you can win just from re-writing inflate in\nassembler, but I do not expect any significant gains other than that.\nAnd coupling has obvious disadvantages when it comes to maintenance...\n\n\nDmitry\n"},{"id":"119515","messageId":"7vljlzjorh.fsf@alter.siamese.dyndns.org","threadId":"20241","inReplyTo":"20090804080147.15201.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-04T20:41:06Z","receivedAt":"2009-08-04T20:41:06Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"George Spelvin\" <linux@horizon.com> writes:\n\n> The one question I have is that currently perl is not a critical\n> compile-time dependency; it's needed for some extra stuff, but AFAIK you\n> can get most of git working without it.  Whether to add that dependency\n> or what is a Junio question.\n\nI am actually feel a lot more uneasy to apply a patch signed of by\nsomebody who calls himself George Spelvin, though.\n\nThree classes of people compile git from the source:\n\n * People who want to be on the bleeding edge and compile git for\n   themselves, even though they are on mainstream platforms where they\n   could choose distro-packaged one;\n\n * People who produce binary packages for distribution.\n\n * People who are on minority platforms and have no other way to get git\n   than compiling for themselves;\n\nWe do not have to worry about the first two groups of people.  It won't\nbe too involved for them to install Perl on their system; after all they\nare already coping with asciidoc and xmlto ;-)\n\nWe can continue shipping mozilla one to help the last group.\n\nIn the Makefile, we say:\n\n    # Define NO_OPENSSL environment variable if you do not have OpenSSL.\n    # This also implies MOZILLA_SHA1.\n\nand with your change, we would start implying STANDALONE_OPENSSL_SHA1\ninstead.  But if MOZILLA_SHA1 was given explicitly, we could use that.\n\nIf they really really really want the extra performance out of statically\nlinked OpenSSL derivative, they could prepare a preprocessed assmebly on\nsome other machine and use it as the last resort if they do not have/want\nPerl.  The situation is exactly the same as the documentation set.  They\nare using HTML/man prepared on another machine (namely, mine) as the last\nresort if they do not have/want AsciiDoc toolchain.\n"},{"id":"119649","messageId":"20090805181755.22765.qmail@science.horizon.com","threadId":"20241","inReplyTo":"7vljlzjorh.fsf@alter.siamese.dyndns.org","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-08-05T18:17:55Z","receivedAt":"2009-08-05T18:17:55Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> Three classes of people compile git from the source:\n>\n> * People who want to be on the bleeding edge and compile git for\n>   themselves, even though they are on mainstream platforms where they\n>   could choose distro-packaged one;\n>\n> * People who produce binary packages for distribution.\n>\n> * People who are on minority platforms and have no other way to get git\n>   than compiling for themselves;\n>\n> We do not have to worry about the first two groups of people.  It won't\n> be too involved for them to install Perl on their system; after all they\n> are already coping with asciidoc and xmlto ;-)\n\nActually, I'd get rid of the perl entirely, but I'm not sure how\nnecessary the other-assembler-syntax features are needed by the\nfolks on MacOS X and Windows (msysgit).\n\n> We can continue shipping mozilla one to help the last group.\n\nOf course, we always need a C fallback.  Would you like a faster one?\n\n> In the Makefile, we say:\n>\n>    # Define NO_OPENSSL environment variable if you do not have OpenSSL.\n>    # This also implies MOZILLA_SHA1.\n>\n> and with your change, we would start implying STANDALONE_OPENSSL_SHA1\n> instead.  But if MOZILLA_SHA1 was given explicitly, we could use that.\n\nWell, I'd really like to auto-detect the processor.  Current gcc's\n\"gcc -v\" output includes a \"Target: \" line that will do nicely.  I can,\nof course, fall back to C if it fails, but is there a significant user\nbase using a non-GCC compiler?\n"},{"id":"119680","messageId":"alpine.DEB.1.00.0908052234470.8306@pacific.mpi-cbg.de","threadId":"20241","inReplyTo":"20090805181755.22765.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2009-08-05T20:36:38Z","receivedAt":"2009-08-05T20:36:38Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 5 Aug 2009, George Spelvin wrote:\n\n> > Three classes of people compile git from the source:\n> >\n> > * People who want to be on the bleeding edge and compile git for\n> >   themselves, even though they are on mainstream platforms where they\n> >   could choose distro-packaged one;\n> >\n> > * People who produce binary packages for distribution.\n> >\n> > * People who are on minority platforms and have no other way to get git\n> >   than compiling for themselves;\n> >\n> > We do not have to worry about the first two groups of people.  It won't\n> > be too involved for them to install Perl on their system; after all they\n> > are already coping with asciidoc and xmlto ;-)\n> \n> Actually, I'd get rid of the perl entirely, but I'm not sure how\n> necessary the other-assembler-syntax features are needed by the\n> folks on MacOS X and Windows (msysgit).\n\nDon't worry for MacOSX and msysGit (or Cygwin, for that matter): all of \nthem use GCC.\n\n> > We can continue shipping mozilla one to help the last group.\n> \n> Of course, we always need a C fallback.  Would you like a faster one?\n\nIs that a trick question?\n\n:-)\n\n> > In the Makefile, we say:\n> >\n> >    # Define NO_OPENSSL environment variable if you do not have OpenSSL.\n> >    # This also implies MOZILLA_SHA1.\n> >\n> > and with your change, we would start implying STANDALONE_OPENSSL_SHA1\n> > instead.  But if MOZILLA_SHA1 was given explicitly, we could use that.\n> \n> Well, I'd really like to auto-detect the processor.  Current gcc's\n> \"gcc -v\" output includes a \"Target: \" line that will do nicely.  I can,\n> of course, fall back to C if it fails, but is there a significant user\n> base using a non-GCC compiler?\n\nDo you really want to determine which processor to optimize for at compile \ntime?  Build system and target system are often different...\n\nCiao,\nDscho\n"},{"id":"119683","messageId":"7vzlaeyoqz.fsf@alter.siamese.dyndns.org","threadId":"20241","inReplyTo":"20090805181755.22765.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-05T20:44:36Z","receivedAt":"2009-08-05T20:44:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"George Spelvin\" <linux@horizon.com> writes:\n\n>> We can continue shipping mozilla one to help the last group.\n>\n> Of course, we always need a C fallback.  Would you like a faster one?\n\nNo.  I'd rather keep tested and tried while a better alternative is in\nwork-in-progress state.\n"},{"id":"119689","messageId":"alpine.LFD.2.01.0908051352280.3390@localhost.localdomain","threadId":"20241","inReplyTo":"20090805181755.22765.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-05T20:55:28Z","receivedAt":"2009-08-05T20:55:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Aug 2009, George Spelvin wrote:\n> \n> > We can continue shipping mozilla one to help the last group.\n> \n> Of course, we always need a C fallback.  Would you like a faster one?\n\nI actually looked at code generation (on x86-64) for the C fallback, and \nit should be quite doable to re-write the C one to generate good code on \nx86-64.\n\nOn 32-bit x86, I suspect the register pressures are so intense that it's \nunrealistic to expect gcc to do a good job, but the Mozilla SHA1 C code \nreally seems _designed_ to be slow in stupid ways (that whole \"byte at a \ntime into a word buffer with shifts\" is a really really sucky way to \nhandle the endianness issues).\n\nSo if you'd like to look at the C version, that's definitely worth it. \nMuch bigger bang for the buck than trying to schedule asm language and \nhaving to deal with different assemblers/linkers/whatnot.\n\n\t\tLinus\n"},{"id":"119723","messageId":"alpine.LFD.2.01.0908051545000.3390@localhost.localdomain","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908051352280.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-05T23:13:20Z","receivedAt":"2009-08-05T23:13:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Aug 2009, Linus Torvalds wrote:\n> \n> I actually looked at code generation (on x86-64) for the C fallback, and \n> it should be quite doable to re-write the C one to generate good code on \n> x86-64.\n\nOk, here's a try.\n\nIt's based on the mozilla SHA1 code, but with quite a bit of surgery. \nEnable with \"make BLK_SHA1=1\".\n\nTimings for \"git fsck --full\" on the git directory:\n\n - Mozilla SHA1 portable C-code (sucky sucky): MOZILLA_SHA1=1\n\n\treal\t0m38.194s\n\tuser\t0m37.838s\n\tsys\t0m0.356s\n\n - This code (\"half-portable C code\"): BLK_SHA1=1\n\n\treal\t0m28.120s\n\tuser\t0m27.930s\n\tsys\t0m0.192s\n\n - OpenSSL assembler code:\n\n\treal\t0m26.327s\n\tuser\t0m26.194s\n\tsys\t0m0.136s\n\nie this is slightly slower than the openssh SHA1 routines, but that's only \ntrue on something very SHA1-intensive like \"git fsck\", and this is \n_almost_ portable code. I say \"almost\" because it really does require that \nwe can do unaligned word loads, and do a good job of 'htonl()', and it \nassumes that 'unsigned int' is 32-bit (the latter would be easy to change \nby using 'uint32_t', but since it's not the relevant portability issue, I \ndon't think it matters).\n\nIn other words, unlike the Mozilla SHA1, this one doesn't suck. It's \ncertainly not great either, but it's probably good enough in practice, \nwithout the headaches of actually making people use an assembler version.\n\nAnd maybe somebody can see how to improve it further?\n\n\t\tLinus\n---\nFrom: Linus Torvalds <torvalds@linux-foundation.org>\nSubject: [PATCH] Add new optimized C 'block-sha1' routines\n\nBased on the mozilla SHA1 routine, but doing the input data accesses a\nword at a time and with 'htonl()' instead of loading bytes and shifting.\n\nIt requires an architecture that is ok with unaligned 32-bit loads and a\nfast htonl().\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n Makefile          |    9 +++\n block-sha1/sha1.c |  145 +++++++++++++++++++++++++++++++++++++++++++++++++++++\n block-sha1/sha1.h |   21 ++++++++\n 3 files changed, 175 insertions(+), 0 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex d7669b1..f12024c 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -84,6 +84,10 @@ all::\n # specify your own (or DarwinPort's) include directories and\n # library directories by defining CFLAGS and LDFLAGS appropriately.\n #\n+# Define BLK_SHA1 environment variable if you want the C version\n+# of the SHA1 that assumes you can do unaligned 32-bit loads and\n+# have a fast htonl() function.\n+#\n # Define PPC_SHA1 environment variable when running make to make use of\n # a bundled SHA1 routine optimized for PowerPC.\n #\n@@ -1166,6 +1170,10 @@ ifdef NO_DEFLATE_BOUND\n \tBASIC_CFLAGS += -DNO_DEFLATE_BOUND\n endif\n \n+ifdef BLK_SHA1\n+\tSHA1_HEADER = \"block-sha1/sha1.h\"\n+\tLIB_OBJS += block-sha1/sha1.o\n+else\n ifdef PPC_SHA1\n \tSHA1_HEADER = \"ppc/sha1.h\"\n \tLIB_OBJS += ppc/sha1.o ppc/sha1ppc.o\n@@ -1183,6 +1191,7 @@ else\n endif\n endif\n endif\n+endif\n ifdef NO_PERL_MAKEMAKER\n \texport NO_PERL_MAKEMAKER\n endif\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nnew file mode 100644\nindex 0000000..8fd90b0\n--- /dev/null\n+++ b/block-sha1/sha1.c\n@@ -0,0 +1,145 @@\n+/*\n+ * Based on the Mozilla SHA1 (see mozilla-sha1/sha1.c),\n+ * optimized to do word accesses rather than byte accesses,\n+ * and to avoid unnecessary copies into the context array.\n+ */\n+\n+#include <string.h>\n+#include <arpa/inet.h>\n+\n+#include \"sha1.h\"\n+\n+/* Hash one 64-byte block of data */\n+static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data);\n+\n+void blk_SHA1_Init(blk_SHA_CTX *ctx)\n+{\n+\tctx->lenW = 0;\n+\tctx->size = 0;\n+\n+\t/* Initialize H with the magic constants (see FIPS180 for constants)\n+\t */\n+\tctx->H[0] = 0x67452301;\n+\tctx->H[1] = 0xefcdab89;\n+\tctx->H[2] = 0x98badcfe;\n+\tctx->H[3] = 0x10325476;\n+\tctx->H[4] = 0xc3d2e1f0;\n+}\n+\n+\n+void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, int len)\n+{\n+\tint lenW = ctx->lenW;\n+\n+\tctx->size += len << 3;\n+\n+\t/* Read the data into W and process blocks as they get full\n+\t */\n+\tif (lenW) {\n+\t\tint left = 64 - lenW;\n+\t\tif (len < left)\n+\t\t\tleft = len;\n+\t\tmemcpy(lenW + (char *)ctx->W, data, left);\n+\t\tlenW = (lenW + left) & 63;\n+\t\tlen -= left;\n+\t\tdata += left;\n+\t\tctx->lenW = lenW;\n+\t\tif (lenW)\n+\t\t\treturn;\n+\t\tblk_SHA1Block(ctx, ctx->W);\n+\t}\n+\twhile (len >= 64) {\n+\t\tblk_SHA1Block(ctx, data);\n+\t\tdata += 64;\n+\t\tlen -= 64;\n+\t}\n+\tif (len) {\n+\t\tmemcpy(ctx->W, data, len);\n+\t\tctx->lenW = len;\n+\t}\n+}\n+\n+\n+void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n+{\n+\tstatic const unsigned char pad[64] = { 0x80 };\n+\tunsigned int padlen[2];\n+\tint i;\n+\n+\t/* Pad with a binary 1 (ie 0x80), then zeroes, then length\n+\t */\n+\tpadlen[0] = htonl(ctx->size >> 32);\n+\tpadlen[1] = htonl(ctx->size);\n+\n+\tblk_SHA1_Update(ctx, pad, 1+ (63 & (55 - ctx->lenW)));\n+\tblk_SHA1_Update(ctx, padlen, 8);\n+\n+\t/* Output hash\n+\t */\n+\tfor (i = 0; i < 5; i++)\n+\t\t((unsigned int *)hashout)[i] = htonl(ctx->H[i]);\n+}\n+\n+#define SHA_ROT(X,n) (((X) << (n)) | ((X) >> (32-(n))))\n+\n+static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n+{\n+\tint t;\n+\tunsigned int A,B,C,D,E,TEMP;\n+\tunsigned int W[80];\n+\n+\tfor (t = 0; t < 16; t++)\n+\t\tW[t] = htonl(data[t]);\n+\n+\t/* Unroll it? */\n+\tfor (t = 16; t <= 79; t++)\n+\t\tW[t] = SHA_ROT(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n+\n+\tA = ctx->H[0];\n+\tB = ctx->H[1];\n+\tC = ctx->H[2];\n+\tD = ctx->H[3];\n+\tE = ctx->H[4];\n+\n+#define T_0_19(t) \\\n+\tTEMP = SHA_ROT(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n+\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n+\n+\tT_0_19( 0); T_0_19( 1); T_0_19( 2); T_0_19( 3); T_0_19( 4);\n+\tT_0_19( 5); T_0_19( 6); T_0_19( 7); T_0_19( 8); T_0_19( 9);\n+\tT_0_19(10); T_0_19(11); T_0_19(12); T_0_19(13); T_0_19(14);\n+\tT_0_19(15); T_0_19(16); T_0_19(17); T_0_19(18); T_0_19(19);\n+\n+#define T_20_39(t) \\\n+\tTEMP = SHA_ROT(A,5) + (B^C^D)           + E + W[t] + 0x6ed9eba1; \\\n+\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n+\n+\tT_20_39(20); T_20_39(21); T_20_39(22); T_20_39(23); T_20_39(24);\n+\tT_20_39(25); T_20_39(26); T_20_39(27); T_20_39(28); T_20_39(29);\n+\tT_20_39(30); T_20_39(31); T_20_39(32); T_20_39(33); T_20_39(34);\n+\tT_20_39(35); T_20_39(36); T_20_39(37); T_20_39(38); T_20_39(39);\n+\n+#define T_40_59(t) \\\n+\tTEMP = SHA_ROT(A,5) + ((B&C)|(D&(B|C))) + E + W[t] + 0x8f1bbcdc; \\\n+\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n+\n+\tT_40_59(40); T_40_59(41); T_40_59(42); T_40_59(43); T_40_59(44);\n+\tT_40_59(45); T_40_59(46); T_40_59(47); T_40_59(48); T_40_59(49);\n+\tT_40_59(50); T_40_59(51); T_40_59(52); T_40_59(53); T_40_59(54);\n+\tT_40_59(55); T_40_59(56); T_40_59(57); T_40_59(58); T_40_59(59);\n+\n+#define T_60_79(t) \\\n+\tTEMP = SHA_ROT(A,5) + (B^C^D)           + E + W[t] + 0xca62c1d6; \\\n+\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n+\n+\tT_60_79(60); T_60_79(61); T_60_79(62); T_60_79(63); T_60_79(64);\n+\tT_60_79(65); T_60_79(66); T_60_79(67); T_60_79(68); T_60_79(69);\n+\tT_60_79(70); T_60_79(71); T_60_79(72); T_60_79(73); T_60_79(74);\n+\tT_60_79(75); T_60_79(76); T_60_79(77); T_60_79(78); T_60_79(79);\n+\n+\tctx->H[0] += A;\n+\tctx->H[1] += B;\n+\tctx->H[2] += C;\n+\tctx->H[3] += D;\n+\tctx->H[4] += E;\n+}\ndiff --git a/block-sha1/sha1.h b/block-sha1/sha1.h\nnew file mode 100644\nindex 0000000..dbc719f\n--- /dev/null\n+++ b/block-sha1/sha1.h\n@@ -0,0 +1,21 @@\n+/*\n+ * Based on the Mozilla SHA1 (see mozilla-sha1/sha1.h),\n+ * optimized to do word accesses rather than byte accesses,\n+ * and to avoid unnecessary copies into the context array.\n+ */\n+\n+typedef struct {\n+\tunsigned int H[5];\n+\tunsigned int W[16];\n+\tint lenW;\n+\tunsigned long long size;\n+} blk_SHA_CTX;\n+\n+void blk_SHA1_Init(blk_SHA_CTX *ctx);\n+void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *dataIn, int len);\n+void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx);\n+\n+#define git_SHA_CTX\tblk_SHA_CTX\n+#define git_SHA1_Init\tblk_SHA1_Init\n+#define git_SHA1_Update\tblk_SHA1_Update\n+#define git_SHA1_Final\tblk_SHA1_Final\n"},{"id":"119724","messageId":"alpine.LFD.2.01.0908051800030.3390@localhost.localdomain","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908051545000.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-06T01:18:49Z","receivedAt":"2009-08-06T01:18:49Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Aug 2009, Linus Torvalds wrote:\n> \n> Timings for \"git fsck --full\" on the git directory:\n> \n>  - Mozilla SHA1 portable C-code (sucky sucky): MOZILLA_SHA1=1\n> \n> \treal\t0m38.194s\n> \tuser\t0m37.838s\n> \tsys\t0m0.356s\n> \n>  - This code (\"half-portable C code\"): BLK_SHA1=1\n> \n> \treal\t0m28.120s\n> \tuser\t0m27.930s\n> \tsys\t0m0.192s\n> \n>  - OpenSSL assembler code:\n> \n> \treal\t0m26.327s\n> \tuser\t0m26.194s\n> \tsys\t0m0.136s\n\nOk, I installed the 32-bit libraries too, to see what it looks like for \nthat case. As expected, the compiler is not able to do a great job due to \nit being somewhat register starved, but on the other hand, the old Mozilla \ncode did even worse, so..\n\n - Mozilla SHA:\n\n\treal\t0m47.063s\n\tuser\t0m46.815s\n\tsys\t0m0.252s\n\n - BLK_SHA1=1\n\n\treal\t0m34.705s\n\tuser\t0m34.394s\n\tsys\t0m0.312s\n\n - OPENSSL:\n\n\treal\t0m29.754s\n\tuser\t0m29.446s\n\tsys\t0m0.288s\n\nso the tuned asm from OpenSSL does kick ass, but the C code version isn't \n_that_ far away. It's quite a reasonable alternative if you don't have the \nOpenSSL libraries installed, for example.\n\nI note that MINGW does NO_OPENSSL by default, for example, and maybe the \nMINGW people want to test the patch out and enable BLK_SHA1 rather than \nthe original Mozilla one.\n\nBut while looking at 32-bit issues, I noticed that I really should also \ncast 'len' when shifting it. Otherwise the thing is limited to fairly \nsmall areas (28 bits - 256MB). This is not just a 32-bit problem (\"int\" is \na signed 32-bit thing even in a 64-bit build), but I only noticed it when \nlooking at 32-bit issues.\n\nSo here's an incremental patch to fix that. \n\n\t\tLinus\n\n---\n block-sha1/sha1.c |    4 ++--\n block-sha1/sha1.h |    2 +-\n 2 files changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nindex 8fd90b0..eef32f7 100644\n--- a/block-sha1/sha1.c\n+++ b/block-sha1/sha1.c\n@@ -27,11 +27,11 @@ void blk_SHA1_Init(blk_SHA_CTX *ctx)\n }\n \n \n-void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, int len)\n+void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n {\n \tint lenW = ctx->lenW;\n \n-\tctx->size += len << 3;\n+\tctx->size += (unsigned long long) len << 3;\n \n \t/* Read the data into W and process blocks as they get full\n \t */\ndiff --git a/block-sha1/sha1.h b/block-sha1/sha1.h\nindex dbc719f..7be2d93 100644\n--- a/block-sha1/sha1.h\n+++ b/block-sha1/sha1.h\n@@ -12,7 +12,7 @@ typedef struct {\n } blk_SHA_CTX;\n \n void blk_SHA1_Init(blk_SHA_CTX *ctx);\n-void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *dataIn, int len);\n+void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *dataIn, unsigned long len);\n void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx);\n \n #define git_SHA_CTX\tblk_SHA_CTX\n"},{"id":"119725","messageId":"alpine.LFD.2.00.0908052144430.16073@xanadu.home","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908051800030.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-08-06T01:52:35Z","receivedAt":"2009-08-06T01:52:35Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 5 Aug 2009, Linus Torvalds wrote:\n\n> But while looking at 32-bit issues, I noticed that I really should also \n> cast 'len' when shifting it. Otherwise the thing is limited to fairly \n> small areas (28 bits - 256MB). This is not just a 32-bit problem (\"int\" is \n> a signed 32-bit thing even in a 64-bit build), but I only noticed it when \n> looking at 32-bit issues.\n\nEven better is to not shift len at all in SHA_update() but shift \nctx->size only at the end in SHA_final().  It is not like if \nSHA_update() could operate on partial bytes, so counting total bytes \ninstead of total bits is all you need.  This way you need no cast there \nand make the code slightly faster.\n\n\nNicolas\n"},{"id":"119727","messageId":"7vocqtu286.fsf@alter.siamese.dyndns.org","threadId":"20241","inReplyTo":"alpine.LFD.2.00.0908052144430.16073@xanadu.home","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2009-08-06T02:04:41Z","receivedAt":"2009-08-06T02:04:41Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Nicolas Pitre <nico@cam.org> writes:\n\n> On Wed, 5 Aug 2009, Linus Torvalds wrote:\n>\n>> But while looking at 32-bit issues, I noticed that I really should also \n>> cast 'len' when shifting it. Otherwise the thing is limited to fairly \n>> small areas (28 bits - 256MB). This is not just a 32-bit problem (\"int\" is \n>> a signed 32-bit thing even in a 64-bit build), but I only noticed it when \n>> looking at 32-bit issues.\n>\n> Even better is to not shift len at all in SHA_update() but shift \n> ctx->size only at the end in SHA_final().  It is not like if \n> SHA_update() could operate on partial bytes, so counting total bytes \n> instead of total bits is all you need.  This way you need no cast there \n> and make the code slightly faster.\n\nLike this?\n\nBy the way, Mozilla one calls Init at the end of Final but block-sha1\ndoesn't.  I do not think it matters for our callers, but on the other hand\nFInal is not performance critical part nor Init is heavy, so it may not be\na bad idea to imitate them as well.  Or am I missing something?\n\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nindex eef32f7..8293f7b 100644\n--- a/block-sha1/sha1.c\n+++ b/block-sha1/sha1.c\n@@ -31,7 +31,7 @@ void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n {\n \tint lenW = ctx->lenW;\n \n-\tctx->size += (unsigned long long) len << 3;\n+\tctx->size += (unsigned long long) len;\n \n \t/* Read the data into W and process blocks as they get full\n \t */\n@@ -68,6 +68,7 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n \n \t/* Pad with a binary 1 (ie 0x80), then zeroes, then length\n \t */\n+\tctx->size <<= 3; /* bytes to bits */\n \tpadlen[0] = htonl(ctx->size >> 32);\n \tpadlen[1] = htonl(ctx->size);\n \n"},{"id":"119728","messageId":"alpine.LFD.2.01.0908051902580.3390@localhost.localdomain","threadId":"20241","inReplyTo":"alpine.LFD.2.00.0908052144430.16073@xanadu.home","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-06T02:08:03Z","receivedAt":"2009-08-06T02:08:03Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Aug 2009, Nicolas Pitre wrote:\n> \n> Even better is to not shift len at all in SHA_update() but shift \n> ctx->size only at the end in SHA_final().  It is not like if \n> SHA_update() could operate on partial bytes, so counting total bytes \n> instead of total bits is all you need.  This way you need no cast there \n> and make the code slightly faster.\n\nYeah, I tried it, but it's not noticeable.\n\nThe bigger issue seems to be that it's shifter-limited, or that's what I \ntake away from my profiles. I suspect it's even _more_ shifter-limited on \nsome other micro-architectures, because gcc is being stupid, and generates\n\n\tror $31,%eax\n\nfrom the \"left shift + right shift\" combination. It seems to -always- \ngenerate a \"ror\", rather than trying to generate 'rot' if the shift count \nwould be smaller that way.\n\nAnd I know _some_ old micro-architectures will literally internally loop \non the rol/ror counts, so \"ror $31\" can be _much_ more expensive than \"rol \n$1\".\n\nThat isn't the case on my Nehalem, though. But I can't seem to get gcc to \ngenerate better code without actually using inline asm..\n\n(So to clarify: this patch makes no difference that I can see to \nperformance, but I suspect it could matter on other CPU's like an old \nPentium or maybe an Atom).\n\n\t\tLinus\n\n---\n block-sha1/sha1.c |   36 ++++++++++++++++++++++++------------\n block-sha1/sha1.h |    2 +-\n 2 files changed, 25 insertions(+), 13 deletions(-)\n\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nindex 8fd90b0..a45a3de 100644\n--- a/block-sha1/sha1.c\n+++ b/block-sha1/sha1.c\n@@ -80,7 +80,19 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n \t\t((unsigned int *)hashout)[i] = htonl(ctx->H[i]);\n }\n \n-#define SHA_ROT(X,n) (((X) << (n)) | ((X) >> (32-(n))))\n+#if defined(__i386__) || defined(__x86_64__)\n+\n+#define SHA_ASM(op, x, n) ({ unsigned int __res; asm(op \" %1,%0\":\"=r\" (__res):\"i\" (n), \"0\" (x)); __res; })\n+#define SHA_ROL(x,n)\tSHA_ASM(\"rol\", x, n)\n+#define SHA_ROR(x,n)\tSHA_ASM(\"ror\", x, n)\n+\n+#else\n+\n+#define SHA_ROT(X,n)\t(((X) << (l)) | ((X) >> (r)))\n+#define SHA_ROL(X,n)\tSHA_ROT(X,n,32-(n))\n+#define SHA_ROR(X,n)\tSHA_ROT(X,32-(n),n)\n+\n+#endif\n \n static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n {\n@@ -93,7 +105,7 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \n \t/* Unroll it? */\n \tfor (t = 16; t <= 79; t++)\n-\t\tW[t] = SHA_ROT(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n+\t\tW[t] = SHA_ROL(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n \n \tA = ctx->H[0];\n \tB = ctx->H[1];\n@@ -102,8 +114,8 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \tE = ctx->H[4];\n \n #define T_0_19(t) \\\n-\tTEMP = SHA_ROT(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n-\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n+\tTEMP = SHA_ROL(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n+\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n \n \tT_0_19( 0); T_0_19( 1); T_0_19( 2); T_0_19( 3); T_0_19( 4);\n \tT_0_19( 5); T_0_19( 6); T_0_19( 7); T_0_19( 8); T_0_19( 9);\n@@ -111,8 +123,8 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \tT_0_19(15); T_0_19(16); T_0_19(17); T_0_19(18); T_0_19(19);\n \n #define T_20_39(t) \\\n-\tTEMP = SHA_ROT(A,5) + (B^C^D)           + E + W[t] + 0x6ed9eba1; \\\n-\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n+\tTEMP = SHA_ROL(A,5) + (B^C^D)           + E + W[t] + 0x6ed9eba1; \\\n+\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n \n \tT_20_39(20); T_20_39(21); T_20_39(22); T_20_39(23); T_20_39(24);\n \tT_20_39(25); T_20_39(26); T_20_39(27); T_20_39(28); T_20_39(29);\n@@ -120,8 +132,8 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \tT_20_39(35); T_20_39(36); T_20_39(37); T_20_39(38); T_20_39(39);\n \n #define T_40_59(t) \\\n-\tTEMP = SHA_ROT(A,5) + ((B&C)|(D&(B|C))) + E + W[t] + 0x8f1bbcdc; \\\n-\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n+\tTEMP = SHA_ROL(A,5) + ((B&C)|(D&(B|C))) + E + W[t] + 0x8f1bbcdc; \\\n+\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n \n \tT_40_59(40); T_40_59(41); T_40_59(42); T_40_59(43); T_40_59(44);\n \tT_40_59(45); T_40_59(46); T_40_59(47); T_40_59(48); T_40_59(49);\n@@ -129,8 +141,8 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \tT_40_59(55); T_40_59(56); T_40_59(57); T_40_59(58); T_40_59(59);\n \n #define T_60_79(t) \\\n-\tTEMP = SHA_ROT(A,5) + (B^C^D)           + E + W[t] + 0xca62c1d6; \\\n-\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n+\tTEMP = SHA_ROL(A,5) + (B^C^D)           + E + W[t] + 0xca62c1d6; \\\n+\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n \n \tT_60_79(60); T_60_79(61); T_60_79(62); T_60_79(63); T_60_79(64);\n \tT_60_79(65); T_60_79(66); T_60_79(67); T_60_79(68); T_60_79(69);\n"},{"id":"119729","messageId":"alpine.LFD.2.01.0908051909080.3390@localhost.localdomain","threadId":"20241","inReplyTo":"7vocqtu286.fsf@alter.siamese.dyndns.org","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-06T02:10:18Z","receivedAt":"2009-08-06T02:10:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Aug 2009, Junio C Hamano wrote:\n> \n> Like this?\n\nNo, combine it with the other shifts:\n\nYes:\n\n> -\tctx->size += (unsigned long long) len << 3;\n> +\tctx->size += (unsigned long long) len;\n\nNo:\n\n> +\tctx->size <<= 3; /* bytes to bits */\n>  \tpadlen[0] = htonl(ctx->size >> 32);\n>  \tpadlen[1] = htonl(ctx->size);\n\nDo\n\n\tpadlen[0] = htonl(ctx->size >> 29);\n\tpadlen[1] = htonl(ctx->size << 3);\n\ninstead. Or whatever.\n\n\t\tLinus\n"},{"id":"119730","messageId":"alpine.LFD.2.00.0908052209490.16073@xanadu.home","threadId":"20241","inReplyTo":"7vocqtu286.fsf@alter.siamese.dyndns.org","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2009-08-06T02:20:03Z","receivedAt":"2009-08-06T02:20:03Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 5 Aug 2009, Junio C Hamano wrote:\n\n> Nicolas Pitre <nico@cam.org> writes:\n> \n> > On Wed, 5 Aug 2009, Linus Torvalds wrote:\n> >\n> >> But while looking at 32-bit issues, I noticed that I really should also \n> >> cast 'len' when shifting it. Otherwise the thing is limited to fairly \n> >> small areas (28 bits - 256MB). This is not just a 32-bit problem (\"int\" is \n> >> a signed 32-bit thing even in a 64-bit build), but I only noticed it when \n> >> looking at 32-bit issues.\n> >\n> > Even better is to not shift len at all in SHA_update() but shift \n> > ctx->size only at the end in SHA_final().  It is not like if \n> > SHA_update() could operate on partial bytes, so counting total bytes \n> > instead of total bits is all you need.  This way you need no cast there \n> > and make the code slightly faster.\n> \n> Like this?\n\nAlmost (see below).\n\n> By the way, Mozilla one calls Init at the end of Final but block-sha1\n> doesn't.  I do not think it matters for our callers, but on the other hand\n> FInal is not performance critical part nor Init is heavy, so it may not be\n> a bad idea to imitate them as well.  Or am I missing something?\n\nIt is done only to make sure potentially crypto sensitive information is \nwiped out of the ctx structure instance.  In our case we have no such \nconcerns.\n\n> diff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\n> index eef32f7..8293f7b 100644\n> --- a/block-sha1/sha1.c\n> +++ b/block-sha1/sha1.c\n> @@ -31,7 +31,7 @@ void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n>  {\n>  \tint lenW = ctx->lenW;\n>  \n> -\tctx->size += (unsigned long long) len << 3;\n> +\tctx->size += (unsigned long long) len;\n\nYou can get rid of the cast as well now.\n\n>  \t/* Read the data into W and process blocks as they get full\n>  \t */\n> @@ -68,6 +68,7 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n>  \n>  \t/* Pad with a binary 1 (ie 0x80), then zeroes, then length\n>  \t */\n> +\tctx->size <<= 3; /* bytes to bits */\n>  \tpadlen[0] = htonl(ctx->size >> 32);\n>  \tpadlen[1] = htonl(ctx->size);\n\nInstead, I'd do:\n\n\tpadlen[0] = htonl(ctx->size >> (32 - 3));\n\tpadlen[1] = htonl(ctx->size << 3);\n\nThat would eliminate a redundant write back of ctx->size.\n\n\nNicolas\n"},{"id":"119732","messageId":"4A7A4BC5.7010106@gmail.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908051902580.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2009-08-06T03:19:33Z","receivedAt":"2009-08-06T03:19:33Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> The bigger issue seems to be that it's shifter-limited, or that's what I \n> take away from my profiles. I suspect it's even _more_ shifter-limited on \n> some other micro-architectures, because gcc is being stupid, and generates\n> \n> \tror $31,%eax\n> \n> from the \"left shift + right shift\" combination. It seems to -always- \n> generate a \"ror\", rather than trying to generate 'rot' if the shift count \n> would be smaller that way.\n> \n> And I know _some_ old micro-architectures will literally internally loop \n> on the rol/ror counts, so \"ror $31\" can be _much_ more expensive than \"rol \n> $1\".\n> \n> That isn't the case on my Nehalem, though. But I can't seem to get gcc to \n> generate better code without actually using inline asm..\n\nThe compiler does the right thing w/ something like this:\n\n+#if __GNUC__>1 && defined(__i386)\n+#define SHA_ROT(data,bits) ({ \\\n+  unsigned d = (data); \\\n+  if (bits<16) \\\n+    __asm__ (\"roll %1,%0\" : \"=r\" (d) : \"I\" (bits), \"0\" (d)); \\\n+  else \\\n+    __asm__ (\"rorl %1,%0\" : \"=r\" (d) : \"I\" (32-bits), \"0\" (d)); \\\n+  d; \\\n+  })\n+#else\n #define SHA_ROT(X,n) (((X) << (n)) | ((X) >> (32-(n))))\n+#endif\n \nwhich doesn't obfuscate the code as much.\n(I needed the asm on p4 anyway, as w/o it the mozilla version is even\n slower than an rfc3174 one. rol vs ror makes no measurable difference)\n\n>  static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n>  {\n> @@ -93,7 +105,7 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n>  \n>  \t/* Unroll it? */\n>  \tfor (t = 16; t <= 79; t++)\n> -\t\tW[t] = SHA_ROT(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n> +\t\tW[t] = SHA_ROL(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n\nunrolling this once (but not more) is a win, at least on p4.\n\n>  #define T_0_19(t) \\\n> -\tTEMP = SHA_ROT(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n> -\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n> +\tTEMP = SHA_ROL(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n> +\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n>  \n>  \tT_0_19( 0); T_0_19( 1); T_0_19( 2); T_0_19( 3); T_0_19( 4);\n>  \tT_0_19( 5); T_0_19( 6); T_0_19( 7); T_0_19( 8); T_0_19( 9);\n\nunrolling these otoh is a clear loss (iirc ~10%). \n\nartur\n"},{"id":"119733","messageId":"alpine.LFD.2.01.0908052024081.3390@localhost.localdomain","threadId":"20241","inReplyTo":"4A7A4BC5.7010106@gmail.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-06T03:31:47Z","receivedAt":"2009-08-06T03:31:47Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Aug 2009, Artur Skawina wrote:\n> \n> >  #define T_0_19(t) \\\n> > -\tTEMP = SHA_ROT(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n> > -\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n> > +\tTEMP = SHA_ROL(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n> > +\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n> >  \n> >  \tT_0_19( 0); T_0_19( 1); T_0_19( 2); T_0_19( 3); T_0_19( 4);\n> >  \tT_0_19( 5); T_0_19( 6); T_0_19( 7); T_0_19( 8); T_0_19( 9);\n> \n> unrolling these otoh is a clear loss (iirc ~10%). \n\nI can well imagine. The P4 decode bandwidth is abysmal unless you get \nthings into the trace cache, and the trace cache is of a very limited \nsize.\n\nHowever, on at least Nehalem, unrolling it all is quite a noticeable win.\n\nThe way it's written, I can easily make it do one or the other by just \nturning the macro inside a loop (and we can have a preprocessor flag to \nchoose one or the other), but let me work on it a bit more first.\n\nI'm trying to move the htonl() inside the loops (the same way I suggested \nGeorge do with his assembly), and it seems to help a tiny bit. But I may \nbe measuring noise.\n\nHowever, right now my biggest profile hit is on this irritating loop:\n\n\t/* Unroll it? */\n\tfor (t = 16; t <= 79; t++)\n\t\tW[t] = SHA_ROL(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n\nand I haven't been able to move _that_ into the other iterations yet.\n\nHere's my micro-optimization update. It does the first 16 rounds (of the \nfirst 20-round thing) specially, and takes the data directly from the \ninput array. I'm _this_ close to breaking the 28s second barrier on \ngit-fsck, but not quite yet.\n\n\t\t\tLinus\n\n---\nFrom: Linus Torvalds <torvalds@linux-foundation.org>\nSubject: [PATCH] block-sha1: make the 'ntohl()' part of the first SHA1 loop\n\nThis helps a teeny bit.  But what I -really- want to do is to avoid the\nwhole 80-array loop, and do the xor updates as I go along..\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n block-sha1/sha1.c |   28 ++++++++++++++++------------\n 1 files changed, 16 insertions(+), 12 deletions(-)\n\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nindex a45a3de..39a5bbb 100644\n--- a/block-sha1/sha1.c\n+++ b/block-sha1/sha1.c\n@@ -100,27 +100,31 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \tunsigned int A,B,C,D,E,TEMP;\n \tunsigned int W[80];\n \n-\tfor (t = 0; t < 16; t++)\n-\t\tW[t] = htonl(data[t]);\n-\n-\t/* Unroll it? */\n-\tfor (t = 16; t <= 79; t++)\n-\t\tW[t] = SHA_ROL(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n-\n \tA = ctx->H[0];\n \tB = ctx->H[1];\n \tC = ctx->H[2];\n \tD = ctx->H[3];\n \tE = ctx->H[4];\n \n-#define T_0_19(t) \\\n+#define T_0_15(t) \\\n+\tTEMP = htonl(data[t]); W[t] = TEMP; \\\n+\tTEMP += SHA_ROL(A,5) + (((C^D)&B)^D)     + E + 0x5a827999; \\\n+\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP; \\\n+\n+\tT_0_15( 0); T_0_15( 1); T_0_15( 2); T_0_15( 3); T_0_15( 4);\n+\tT_0_15( 5); T_0_15( 6); T_0_15( 7); T_0_15( 8); T_0_15( 9);\n+\tT_0_15(10); T_0_15(11); T_0_15(12); T_0_15(13); T_0_15(14);\n+\tT_0_15(15);\n+\n+\t/* Unroll it? */\n+\tfor (t = 16; t <= 79; t++)\n+\t\tW[t] = SHA_ROL(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n+\n+#define T_16_19(t) \\\n \tTEMP = SHA_ROL(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n \tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n \n-\tT_0_19( 0); T_0_19( 1); T_0_19( 2); T_0_19( 3); T_0_19( 4);\n-\tT_0_19( 5); T_0_19( 6); T_0_19( 7); T_0_19( 8); T_0_19( 9);\n-\tT_0_19(10); T_0_19(11); T_0_19(12); T_0_19(13); T_0_19(14);\n-\tT_0_19(15); T_0_19(16); T_0_19(17); T_0_19(18); T_0_19(19);\n+\tT_16_19(16); T_16_19(17); T_16_19(18); T_16_19(19);\n \n #define T_20_39(t) \\\n \tTEMP = SHA_ROL(A,5) + (B^C^D)           + E + W[t] + 0x6ed9eba1; \\\n"},{"id":"119735","messageId":"alpine.LFD.2.01.0908052043082.3390@localhost.localdomain","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908052024081.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-06T03:48:20Z","receivedAt":"2009-08-06T03:48:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Aug 2009, Linus Torvalds wrote:\n> \n> However, right now my biggest profile hit is on this irritating loop:\n> \n> \t/* Unroll it? */\n> \tfor (t = 16; t <= 79; t++)\n> \t\tW[t] = SHA_ROL(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n> \n> and I haven't been able to move _that_ into the other iterations yet.\n\nOh yes I have.\n\nHere's the patch that gets me sub-28s git-fsck times. In fact, it gives me \nsub-27s times. In fact, it's really close to the OpenSSL times.\n\nAnd all using plain C.\n\nAgain - this is all on x86-64. I suspect 32-bit code ends up having \nspills due to register pressure. That said, I did get rid of that big \ntemporary array, and it now basically only uses that 512-bit array as one \ncircular queue.\n\n\t\tLinus\n\nPS. Ok, so my definition of \"plain C\" is a bit odd. There's nothing plain \nabout it. It's disgusting C preprocessor misuse. But dang, it's kind of \nfun to abuse the compiler this way.\n\n---\n block-sha1/sha1.c |   28 ++++++++++++++++------------\n 1 files changed, 16 insertions(+), 12 deletions(-)\n\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nindex 39a5bbb..80193d4 100644\n--- a/block-sha1/sha1.c\n+++ b/block-sha1/sha1.c\n@@ -96,9 +96,8 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n \n static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n {\n-\tint t;\n \tunsigned int A,B,C,D,E,TEMP;\n-\tunsigned int W[80];\n+\tunsigned int array[16];\n \n \tA = ctx->H[0];\n \tB = ctx->H[1];\n@@ -107,8 +106,8 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \tE = ctx->H[4];\n \n #define T_0_15(t) \\\n-\tTEMP = htonl(data[t]); W[t] = TEMP; \\\n-\tTEMP += SHA_ROL(A,5) + (((C^D)&B)^D)     + E + 0x5a827999; \\\n+\tTEMP = htonl(data[t]); array[t] = TEMP; \\\n+\tTEMP += SHA_ROL(A,5) + (((C^D)&B)^D) + E + 0x5a827999; \\\n \tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP; \\\n \n \tT_0_15( 0); T_0_15( 1); T_0_15( 2); T_0_15( 3); T_0_15( 4);\n@@ -116,18 +115,21 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \tT_0_15(10); T_0_15(11); T_0_15(12); T_0_15(13); T_0_15(14);\n \tT_0_15(15);\n \n-\t/* Unroll it? */\n-\tfor (t = 16; t <= 79; t++)\n-\t\tW[t] = SHA_ROL(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n+/* This \"rolls\" over the 512-bit array */\n+#define W(x) (array[(x)&15])\n+#define SHA_XOR(t) \\\n+\tTEMP = SHA_ROL(W(t+13) ^ W(t+8) ^ W(t+2) ^ W(t), 1); W(t) = TEMP;\n \n #define T_16_19(t) \\\n-\tTEMP = SHA_ROL(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n-\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n+\tSHA_XOR(t); \\\n+\tTEMP += SHA_ROL(A,5) + (((C^D)&B)^D) + E + 0x5a827999; \\\n+\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP; \\\n \n \tT_16_19(16); T_16_19(17); T_16_19(18); T_16_19(19);\n \n #define T_20_39(t) \\\n-\tTEMP = SHA_ROL(A,5) + (B^C^D)           + E + W[t] + 0x6ed9eba1; \\\n+\tSHA_XOR(t); \\\n+\tTEMP += SHA_ROL(A,5) + (B^C^D) + E + 0x6ed9eba1; \\\n \tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n \n \tT_20_39(20); T_20_39(21); T_20_39(22); T_20_39(23); T_20_39(24);\n@@ -136,7 +138,8 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \tT_20_39(35); T_20_39(36); T_20_39(37); T_20_39(38); T_20_39(39);\n \n #define T_40_59(t) \\\n-\tTEMP = SHA_ROL(A,5) + ((B&C)|(D&(B|C))) + E + W[t] + 0x8f1bbcdc; \\\n+\tSHA_XOR(t); \\\n+\tTEMP += SHA_ROL(A,5) + ((B&C)|(D&(B|C))) + E + 0x8f1bbcdc; \\\n \tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n \n \tT_40_59(40); T_40_59(41); T_40_59(42); T_40_59(43); T_40_59(44);\n@@ -145,7 +148,8 @@ static void blk_SHA1Block(blk_SHA_CTX *ctx, const unsigned int *data)\n \tT_40_59(55); T_40_59(56); T_40_59(57); T_40_59(58); T_40_59(59);\n \n #define T_60_79(t) \\\n-\tTEMP = SHA_ROL(A,5) + (B^C^D)           + E + W[t] + 0xca62c1d6; \\\n+\tSHA_XOR(t); \\\n+\tTEMP += SHA_ROL(A,5) + (B^C^D) + E + 0xca62c1d6; \\\n \tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n \n \tT_60_79(60); T_60_79(61); T_60_79(62); T_60_79(63); T_60_79(64);\n"},{"id":"119737","messageId":"alpine.LFD.2.01.0908052056500.3390@localhost.localdomain","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908052043082.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-06T04:01:26Z","receivedAt":"2009-08-06T04:01:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 5 Aug 2009, Linus Torvalds wrote:\n> \n> Here's the patch that gets me sub-28s git-fsck times. In fact, it gives me \n> sub-27s times. In fact, it's really close to the OpenSSL times.\n\nJust to back that up:\n\n - OpenSSL:\n\n\treal\t0m26.363s\n\tuser\t0m26.174s\n\tsys\t0m0.188s\n\n - This C implementation:\n\n\treal\t0m26.594s\n\tuser\t0m26.310s\n\tsys\t0m0.256s\n\nso I'm still slower, but now you really have to look closely to see the \ndifference. In fact, you have to do multiple runs to make sure, because \nthe error bars are bigger thant he difference - but openssl definitely \nedges my C code out by a small amount, and the above numbers are rairly \nnormal.\n\n\t\tLinus\n"},{"id":"119738","messageId":"4A7A5723.6070704@gmail.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908052024081.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2009-08-06T04:08:03Z","receivedAt":"2009-08-06T04:08:03Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Thu, 6 Aug 2009, Artur Skawina wrote:\n>>>  #define T_0_19(t) \\\n>>> -\tTEMP = SHA_ROT(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n>>> -\tE = D; D = C; C = SHA_ROT(B, 30); B = A; A = TEMP;\n>>> +\tTEMP = SHA_ROL(A,5) + (((C^D)&B)^D)     + E + W[t] + 0x5a827999; \\\n>>> +\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n>>>  \n>>>  \tT_0_19( 0); T_0_19( 1); T_0_19( 2); T_0_19( 3); T_0_19( 4);\n>>>  \tT_0_19( 5); T_0_19( 6); T_0_19( 7); T_0_19( 8); T_0_19( 9);\n>> unrolling these otoh is a clear loss (iirc ~10%). \n> \n> I can well imagine. The P4 decode bandwidth is abysmal unless you get \n> things into the trace cache, and the trace cache is of a very limited \n> size.\n> \n> However, on at least Nehalem, unrolling it all is quite a noticeable win.\n> \n> The way it's written, I can easily make it do one or the other by just \n> turning the macro inside a loop (and we can have a preprocessor flag to \n> choose one or the other), but let me work on it a bit more first.\n\nthat's of course how i measured it.. :)\n\n> I'm trying to move the htonl() inside the loops (the same way I suggested \n> George do with his assembly), and it seems to help a tiny bit. But I may \n> be measuring noise.\n\ni haven't tried your version at all yet (just applied the rol/ror and\nunrolling changes, but neither was a win on p4)\n\n> However, right now my biggest profile hit is on this irritating loop:\n> \n> \t/* Unroll it? */\n> \tfor (t = 16; t <= 79; t++)\n> \t\tW[t] = SHA_ROL(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n> \n> and I haven't been able to move _that_ into the other iterations yet.\n\ni've done that before -- was a small loss -- maybe because of the small\ntrace cache. deleted that attempt while cleaning up the #if mess, so don't\nhave the patch, but it was basically\n\n#define newW(t) (W[t] = SHA_ROL(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1))\n\nand than s/W[t]/newW(t)/ in rounds 16..79.\n\nI've only tested on p4 and there the winner so far is still:\n\n-  for (t = 16; t <= 79; t++)\n+  for (t = 16; t <= 79; t+=2) {\n     ctx->W[t] =\n-      SHA_ROT(ctx->W[t-3] ^ ctx->W[t-8] ^ ctx->W[t-14] ^ ctx->W[t-16], 1);\n+      SHA_ROT(ctx->W[t-16] ^ ctx->W[t-14] ^ ctx->W[t-8] ^ ctx->W[t-3], 1);\n+    ctx->W[t+1] =\n+      SHA_ROT(ctx->W[t-15] ^ ctx->W[t-13] ^ ctx->W[t-7] ^ ctx->W[t-2], 1);\n+  }\n\n> Here's my micro-optimization update. It does the first 16 rounds (of the \n> first 20-round thing) specially, and takes the data directly from the \n> input array. I'm _this_ close to breaking the 28s second barrier on \n> git-fsck, but not quite yet.\n\ntried this before too -- doesn't help. Not much a of a surprise --\nif unrolling didn't help adding another loop (for rounds 17..20) won't.\n\nartur\n"},{"id":"119743","messageId":"alpine.LFD.2.01.0908052120330.3390@localhost.localdomain","threadId":"20241","inReplyTo":"4A7A5723.6070704@gmail.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-06T04:27:05Z","receivedAt":"2009-08-06T04:27:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Aug 2009, Artur Skawina wrote:\n> > \n> > The way it's written, I can easily make it do one or the other by just \n> > turning the macro inside a loop (and we can have a preprocessor flag to \n> > choose one or the other), but let me work on it a bit more first.\n> \n> that's of course how i measured it.. :)\n\nWell, with my \"rolling 512-bit array\" I can't do that easily any more.\n\nNow it actually depends on the compiler being able to statically do that \ncircular list calculation. If I were to turn it back into the chunks of \nloops, my new code would suck, because it would have all those nasty \ndynamic address calculations.\n\n> I've only tested on p4 and there the winner so far is still:\n\nYeah, well, I refuse to touch that crappy micro-architecture any more. I \ncomplained to Intel people for years that their best CPU was only \navailable as a laptop chip (Pentium-M), and I'm really happy to have \ngotten rid of all my horrid P4's.\n\n(Ok, so it was great when the P4 ran at 2x the frequency of the \ncompetition, and then it smoked them all. Except on OS loads, where the P4 \nexception handling took ten times longer than anything else).\n\nSo I'm a big biased against P4. \n\nI'll try it on my Atom's, though. They're pretty crappy CPU's, but they \nhave a fairly good _reason_ to be crappy.\n\n\t\t\tLinus\n"},{"id":"119744","messageId":"4A7A5BE2.5070401@gmail.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908052056500.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2009-08-06T04:28:18Z","receivedAt":"2009-08-06T04:28:18Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Wed, 5 Aug 2009, Linus Torvalds wrote:\n>> Here's the patch that gets me sub-28s git-fsck times. In fact, it gives me \n>> sub-27s times. In fact, it's really close to the OpenSSL times.\n> \n> Just to back that up:\n> \n>  - OpenSSL:\n> \n> \treal\t0m26.363s\n> \tuser\t0m26.174s\n> \tsys\t0m0.188s\n> \n>  - This C implementation:\n> \n> \treal\t0m26.594s\n> \tuser\t0m26.310s\n> \tsys\t0m0.256s\n> \n> so I'm still slower, but now you really have to look closely to see the \n> difference. In fact, you have to do multiple runs to make sure, because \n> the error bars are bigger thant he difference - but openssl definitely \n> edges my C code out by a small amount, and the above numbers are rairly \n> normal.\n\nnice, the p4 microbenchmark #s:\n\n#             TIME[s] SPEED[MB/s]\nrfc3174         1.357       44.99\nrfc3174         1.352       45.13\nmozilla         1.509       40.44\nmozillaas       1.133       53.87\nlinus          0.5818       104.9\n\nso it's more than twice as fast as the mozilla implementation.\n\nartur\n"},{"id":"119746","messageId":"alpine.LFD.2.01.0908052137400.3390@localhost.localdomain","threadId":"20241","inReplyTo":"4A7A5BE2.5070401@gmail.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2009-08-06T04:50:21Z","receivedAt":"2009-08-06T04:50:21Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 6 Aug 2009, Artur Skawina wrote:\n> \n> #             TIME[s] SPEED[MB/s]\n> rfc3174         1.357       44.99\n> rfc3174         1.352       45.13\n> mozilla         1.509       40.44\n> mozillaas       1.133       53.87\n> linus          0.5818       104.9\n> \n> so it's more than twice as fast as the mozilla implementation.\n\nSo that's some general SHA1 benchmark you have?\n\nI hope it tests correctness too. \n\nAlthough I can't imagine it being wrong - I've made mistakes (oh, yes, \nmany mistakes) when trying to convert the code to something efficient, and \neven the smallest mistake results in 'git fsck' immediately complaining \nabout every single object.\n\nBut still. I literally haven't tested it any other way (well, the git \ntest-suite ends up doing a fair amount of testing too, and I _have_ run \nthat).\n\nAs to my atom testing: my poor little atom is a sad little thing, and \nit's almost painful to benchmark that thing. But it's worth it to look at \nhow the 32-bit code compares to the openssl asm code too:\n\n - BLK_SHA1:\n\n\treal\t2m27.160s\n\tuser\t2m23.651s\n\tsys\t0m2.392s\n\n - OpenSSL:\n\n\treal\t2m12.580s\n\tuser\t2m9.998s\n\tsys\t0m1.811s\n\n - Mozilla-SHA1:\n\n\treal\t3m21.836s\n\tuser\t3m18.369s\n\tsys\t0m2.862s\n\nAs expected, the hand-tuned assembly does better (and by a bigger margin). \nProbably partly because scheduling is important when in-order, and partly \nbecause gcc will have a harder time with the small register set.\n\nBut it's still a big improvement over mozilla one.\n\n(This is, as always, 'git fsck --full'. It spends about 50% on that SHA1 \ncalculation, so the SHA1 speedup is larger than you see from just th \nenumbers)\n\n\t\tLinus\n"},{"id":"119747","messageId":"20090806045203.558.qmail@science.horizon.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908052043082.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-08-06T04:52:03Z","receivedAt":"2009-08-06T04:52:03Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"On Wed, 5 Aug 2009, Linus Torvalds wrote:\n> Oh yes I have.\n> \n> Here's the patch that gets me sub-28s git-fsck times. In fact, it gives me \n> sub-27s times. In fact, it's really close to the OpenSSL times.\n> \n> And all using plain C.\n> \n> Again - this is all on x86-64. I suspect 32-bit code ends up having \n> spills due to register pressure. That said, I did get rid of that big \n> temporary array, and it now basically only uses that 512-bit array as one \n> circular queue.\n> \n> \t\tLinus\n> \n> PS. Ok, so my definition of \"plain C\" is a bit odd. There's nothing plain \n> about it. It's disgusting C preprocessor misuse. But dang, it's kind of \n> fun to abuse the compiler this way.\n\nYou're still missing three tricks, which give a slight speedup \non my machine:\n\n1) (major)\n\tInstead of reassigning all those variable all the time,\n\tmake the round function\n\t\tE += SHA_ROL(A,5) + ((B&C)|(D&(B|C))) + W[t] + 0x8f1bbcdc; \\\n\t\tB = SHA_ROR(B, 2);\n\tand rename the variables between rounds.\n\n2 (minor)\n\tOne of the round functions ((B&C)|(D&(B|C))) can be rewritten\n\t\t  (B&C) | (C&D) | (D&B)\n\t\t= (B&C) | (D&(B|C))\n\t\t= (B&C) | (D&(B^C))\n\t\t= (B&C) ^ (D&(B^C))\n\t\t= (B&C) + (D&(B^C))\n\tto expose more associativty (and thus scheduling flexibility)\n\tto the compiler.\n\n3) (minor)\n\tctx->lenW is always simply a copy of the low 6 bits of ctx->size,\n\tso there's no need to bother with it.\n\nActually, looking at the code, GCC manages to figure out the first\n(major) one by itself.  Way to go, GCC authors!\n\nBut getting avoiding the extra temporary in trick 2 also gets rid of\nsome extra REX prefixes, saving 240 bytes in blk_SHA1Block, which is\nkind of nice in an inner loop.\n\nHere's my modified version of your earlier code.  I haven't\nincoporated the W[] formation into the round functions as in\nyour latest version.\n\nI'm sure you can bash the two together in very little time.  Or I'll\nget to it later; I really should attend to $DAY_JOB at the moment.\n\ndiff --git a/Makefile b/Makefile\nindex daf4296..e6df8ec 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -84,6 +84,10 @@ all::\n # specify your own (or DarwinPort's) include directories and\n # library directories by defining CFLAGS and LDFLAGS appropriately.\n #\n+# Define BLK_SHA1 environment variable if you want the C version\n+# of the SHA1 that assumes you can do unaligned 32-bit loads and\n+# have a fast htonl() function.\n+#\n # Define PPC_SHA1 environment variable when running make to make use of\n # a bundled SHA1 routine optimized for PowerPC.\n #\n@@ -1167,6 +1171,10 @@ ifdef NO_DEFLATE_BOUND\n \tBASIC_CFLAGS += -DNO_DEFLATE_BOUND\n endif\n \n+ifdef BLK_SHA1\n+\tSHA1_HEADER = \"block-sha1/sha1.h\"\n+\tLIB_OBJS += block-sha1/sha1.o\n+else\n ifdef PPC_SHA1\n \tSHA1_HEADER = \"ppc/sha1.h\"\n \tLIB_OBJS += ppc/sha1.o ppc/sha1ppc.o\n@@ -1184,6 +1192,7 @@ else\n endif\n endif\n endif\n+endif\n ifdef NO_PERL_MAKEMAKER\n \texport NO_PERL_MAKEMAKER\n endif\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nnew file mode 100644\nindex 0000000..261eae7\n--- /dev/null\n+++ b/block-sha1/sha1.c\n@@ -0,0 +1,141 @@\n+/*\n+ * Based on the Mozilla SHA1 (see mozilla-sha1/sha1.c),\n+ * optimized to do word accesses rather than byte accesses,\n+ * and to avoid unnecessary copies into the context array.\n+ */\n+\n+#include <string.h>\n+#include <arpa/inet.h>\n+\n+#include \"sha1.h\"\n+\n+/* Hash one 64-byte block of data */\n+static void blk_SHA1Block(blk_SHA_CTX *ctx, const uint32_t *data);\n+\n+void blk_SHA1_Init(blk_SHA_CTX *ctx)\n+{\n+\t/* Initialize H with the magic constants (see FIPS180 for constants)\n+\t */\n+\tctx->H[0] = 0x67452301;\n+\tctx->H[1] = 0xefcdab89;\n+\tctx->H[2] = 0x98badcfe;\n+\tctx->H[3] = 0x10325476;\n+\tctx->H[4] = 0xc3d2e1f0;\n+\tctx->size = 0;\n+}\n+\n+\n+void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, int len)\n+{\n+\tint lenW = (int)ctx->size & 63;\n+\n+\tctx->size += len;\n+\n+\t/* Read the data into W and process blocks as they get full\n+\t */\n+\tif (lenW) {\n+\t\tint left = 64 - lenW;\n+\t\tif (len < left)\n+\t\t\tleft = len;\n+\t\tmemcpy(lenW + (char *)ctx->W, data, left);\n+\t\tif (left + lenW != 64)\n+\t\t\treturn;\n+\t\tlen -= left;\n+\t\tdata += left;\n+\t\tblk_SHA1Block(ctx, ctx->W);\n+\t}\n+\twhile (len >= 64) {\n+\t\tblk_SHA1Block(ctx, data);\n+\t\tdata += 64;\n+\t\tlen -= 64;\n+\t}\n+\tmemcpy(ctx->W, data, len);\n+}\n+\n+\n+void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n+{\n+\tint i, lenW = (int)ctx->size & 63;\n+\n+\t/* Pad with a binary 1 (ie 0x80), then zeroes, then length\n+\t */\n+\t((char *)ctx->W)[lenW++] = 0x80;\n+\tif (lenW > 56) {\n+\t\tmemset((char *)ctx->W + lenW, 0, 64 - lenW);\n+\t\tblk_SHA1Block(ctx, ctx->W);\n+\t\tlenW = 0;\n+\t}\n+\tmemset((char *)ctx->W + lenW, 0, 56 - lenW);\n+\tctx->W[14] = htonl(ctx->size >> 29);\n+\tctx->W[15] = htonl((uint32_t)ctx->size << 3);\n+\tblk_SHA1Block(ctx, ctx->W);\n+\n+\t/* Output hash\n+\t */\n+\tfor (i = 0; i < 5; i++)\n+\t\t((unsigned int *)hashout)[i] = htonl(ctx->H[i]);\n+}\n+\n+/* SHA-1 helper macros */\n+#define SHA_ROT(X,n) (((X) << (n)) | ((X) >> (32-(n))))\n+#define F1(b,c,d) (((d^c)&b)^d)\n+#define F2(b,c,d) (b^c^d)\n+/* This version lets the compiler use the fact that + is associative. */\n+#define F3(b,c,d) (c&d) + (b & (c^d))\n+\n+/* The basic SHA-1 round */\n+#define ROUND(a, b, c, d, e, f, k, t) \\\n+\te += SHA_ROT(a,5) + f(b,c,d) + W[t] + k;  b = SHA_ROT(b, 30)\n+/* Five SHA-1 rounds */\n+#define FIVE(f, k, t) \\\n+\tROUND(A, B, C, D, E, f, k, t  ); \\\n+\tROUND(E, A, B, C, D, f, k, t+1); \\\n+\tROUND(D, E, A, B, C, f, k, t+2); \\\n+\tROUND(C, D, E, A, B, f, k, t+3); \\\n+\tROUND(B, C, D, E, A, f, k, t+4)\n+\n+static void blk_SHA1Block(blk_SHA_CTX *ctx, const uint32_t *data)\n+{\n+\tint t;\n+\tuint32_t A,B,C,D,E;\n+\tuint32_t W[80];\n+\n+\tfor (t = 0; t < 16; t++)\n+\t\tW[t] = htonl(data[t]);\n+\n+\t/* Unroll it? */\n+\tfor (t = 16; t <= 79; t++)\n+\t\tW[t] = SHA_ROT(W[t-3] ^ W[t-8] ^ W[t-14] ^ W[t-16], 1);\n+\n+\tA = ctx->H[0];\n+\tB = ctx->H[1];\n+\tC = ctx->H[2];\n+\tD = ctx->H[3];\n+\tE = ctx->H[4];\n+\n+\tFIVE(F1, 0x5a827999,  0);\n+\tFIVE(F1, 0x5a827999,  5);\n+\tFIVE(F1, 0x5a827999, 10);\n+\tFIVE(F1, 0x5a827999, 15);\n+\n+\tFIVE(F2, 0x6ed9eba1, 20);\n+\tFIVE(F2, 0x6ed9eba1, 25);\n+\tFIVE(F2, 0x6ed9eba1, 30);\n+\tFIVE(F2, 0x6ed9eba1, 35);\n+\n+\tFIVE(F3, 0x8f1bbcdc, 40);\n+\tFIVE(F3, 0x8f1bbcdc, 45);\n+\tFIVE(F3, 0x8f1bbcdc, 50);\n+\tFIVE(F3, 0x8f1bbcdc, 55);\n+\n+\tFIVE(F2, 0xca62c1d6, 60);\n+\tFIVE(F2, 0xca62c1d6, 65);\n+\tFIVE(F2, 0xca62c1d6, 70);\n+\tFIVE(F2, 0xca62c1d6, 75);\n+\n+\tctx->H[0] += A;\n+\tctx->H[1] += B;\n+\tctx->H[2] += C;\n+\tctx->H[3] += D;\n+\tctx->H[4] += E;\n+}\ndiff --git a/block-sha1/sha1.h b/block-sha1/sha1.h\nnew file mode 100644\nindex 0000000..c9dc156\n--- /dev/null\n+++ b/block-sha1/sha1.h\n@@ -0,0 +1,21 @@\n+/*\n+ * Based on the Mozilla SHA1 (see mozilla-sha1/sha1.h),\n+ * optimized to do word accesses rather than byte accesses,\n+ * and to avoid unnecessary copies into the context array.\n+ */\n+ #include <stdint.h>\n+\n+typedef struct {\n+\tuint32_t H[5];\n+\tuint64_t size;\n+\tuint32_t W[16];\n+} blk_SHA_CTX;\n+\n+void blk_SHA1_Init(blk_SHA_CTX *ctx);\n+void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *dataIn, int len);\n+void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx);\n+\n+#define git_SHA_CTX\tblk_SHA_CTX\n+#define git_SHA1_Init\tblk_SHA1_Init\n+#define git_SHA1_Update\tblk_SHA1_Update\n+#define git_SHA1_Final\tblk_SHA1_Final\n"},{"id":"119750","messageId":"4A7A67C5.8060109@gmail.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908052137400.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2009-08-06T05:19:01Z","receivedAt":"2009-08-06T05:19:01Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Thu, 6 Aug 2009, Artur Skawina wrote:\n>> #             TIME[s] SPEED[MB/s]\n>> rfc3174         1.357       44.99\n>> rfc3174         1.352       45.13\n>> mozilla         1.509       40.44\n>> mozillaas       1.133       53.87\n>> linus          0.5818       104.9\n>>\n>> so it's more than twice as fast as the mozilla implementation.\n> \n> So that's some general SHA1 benchmark you have?\n> \n> I hope it tests correctness too. \n\nyep, sort of, i just check that all versions return the same result\nwhen hashing some pseudorandom data.\n\n> As to my atom testing: my poor little atom is a sad little thing, and \n> it's almost painful to benchmark that thing. But it's worth it to look at \n> how the 32-bit code compares to the openssl asm code too:\n> \n>  - BLK_SHA1:\n> \treal\t2m27.160s\n>  - OpenSSL:\n> \treal\t2m12.580s\n>  - Mozilla-SHA1:\n> \treal\t3m21.836s\n> \n> As expected, the hand-tuned assembly does better (and by a bigger margin). \n> Probably partly because scheduling is important when in-order, and partly \n> because gcc will have a harder time with the small register set.\n> \n> But it's still a big improvement over mozilla one.\n> \n> (This is, as always, 'git fsck --full'. It spends about 50% on that SHA1 \n> calculation, so the SHA1 speedup is larger than you see from just th \n> enumbers)\n\nI'll start looking at other cpus once i integrate the asm versions into\nmy benchmark. \n\nP4s really are \"special\". Even something as simple as this on top of your\nversion:\n\n@@ -129,8 +133,8 @@\n \n #define T_20_39(t) \\\n        SHA_XOR(t); \\\n-       TEMP += SHA_ROL(A,5) + (B^C^D) + E + 0x6ed9eba1; \\\n-       E = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n+       TEMP += SHA_ROL(A,5) + (B^C^D) + E; \\\n+       E = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP + 0x6ed9eba1;\n \n        T_20_39(20); T_20_39(21); T_20_39(22); T_20_39(23); T_20_39(24);\n        T_20_39(25); T_20_39(26); T_20_39(27); T_20_39(28); T_20_39(29);\n@@ -139,8 +143,8 @@\n \n #define T_40_59(t) \\\n        SHA_XOR(t); \\\n-       TEMP += SHA_ROL(A,5) + ((B&C)|(D&(B|C))) + E + 0x8f1bbcdc; \\\n-       E = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n+       TEMP += SHA_ROL(A,5) + ((B&C)|(D&(B|C))) + E; \\\n+       E = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP + 0x8f1bbcdc;\n \n        T_40_59(40); T_40_59(41); T_40_59(42); T_40_59(43); T_40_59(44);\n        T_40_59(45); T_40_59(46); T_40_59(47); T_40_59(48); T_40_59(49);\n\nsaves another 10% or so:\n\n#Initializing... Rounds: 1000000, size: 62500K, time: 1.421s, speed: 42.97MB/s\n#             TIME[s] SPEED[MB/s]\nrfc3174         1.403        43.5\n# New hash result: b747042d9f4f1fdabd2ac53076f8f830dea7fe0f\nrfc3174         1.403       43.51\nlinus          0.5891       103.6\nlinusas        0.5337       114.4\nmozilla         1.535       39.76\nmozillaas       1.128       54.13\n\n\nartur\n"},{"id":"119751","messageId":"4A7A6DBC.9010107@gmail.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908052120330.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2009-08-06T05:44:28Z","receivedAt":"2009-08-06T05:44:28Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Thu, 6 Aug 2009, Artur Skawina wrote:\n>>> The way it's written, I can easily make it do one or the other by just \n>>> turning the macro inside a loop (and we can have a preprocessor flag to \n>>> choose one or the other), but let me work on it a bit more first.\n>> that's of course how i measured it.. :)\n> \n> Well, with my \"rolling 512-bit array\" I can't do that easily any more.\n> \n> Now it actually depends on the compiler being able to statically do that \n> circular list calculation. If I were to turn it back into the chunks of \n> loops, my new code would suck, because it would have all those nasty \n> dynamic address calculations.\n\ni did try (obvious patch below) and in fact the loops still win on p4:\n\n#Initializing... Rounds: 1000000, size: 62500K, time: 1.428s, speed: 42.76MB/s\n#             TIME[s] SPEED[MB/s]\nrfc3174         1.437       42.47\nrfc3174         1.438       42.45\nlinus          0.5791       105.4\nlinusas        0.5052       120.8\nmozilla         1.525       40.01\nmozillaas       1.192       51.19\n\nartur\n\n--- block-sha1/sha1.c\t2009-08-06 06:45:03.407322970 +0200\n+++ block-sha1/sha1as.c\t2009-08-06 07:36:41.332318683 +0200\n@@ -107,13 +107,17 @@\n \n #define T_0_15(t) \\\n \tTEMP = htonl(data[t]); array[t] = TEMP; \\\n-\tTEMP += SHA_ROL(A,5) + (((C^D)&B)^D) + E + 0x5a827999; \\\n-\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP; \\\n+\tTEMP += SHA_ROL(A,5) + (((C^D)&B)^D) + E; \\\n+\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP + 0x5a827999; \\\n \n+#if UNROLL\n \tT_0_15( 0); T_0_15( 1); T_0_15( 2); T_0_15( 3); T_0_15( 4);\n \tT_0_15( 5); T_0_15( 6); T_0_15( 7); T_0_15( 8); T_0_15( 9);\n \tT_0_15(10); T_0_15(11); T_0_15(12); T_0_15(13); T_0_15(14);\n \tT_0_15(15);\n+#else\n+\tfor (int t = 0; t <= 15; t++) { T_0_15(t); }\n+#endif\n \n /* This \"rolls\" over the 512-bit array */\n #define W(x) (array[(x)&15])\n@@ -125,37 +129,53 @@\n \tTEMP += SHA_ROL(A,5) + (((C^D)&B)^D) + E + 0x5a827999; \\\n \tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP; \\\n \n+#if UNROLL\n \tT_16_19(16); T_16_19(17); T_16_19(18); T_16_19(19);\n+#else\n+\tfor (int t = 16; t <= 19; t++) { T_16_19(t); }\n+#endif\n \n #define T_20_39(t) \\\n \tSHA_XOR(t); \\\n-\tTEMP += SHA_ROL(A,5) + (B^C^D) + E + 0x6ed9eba1; \\\n-\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n+\tTEMP += SHA_ROL(A,5) + (B^C^D) + E; \\\n+\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP + 0x6ed9eba1;\n \n+#if UNROLL\n \tT_20_39(20); T_20_39(21); T_20_39(22); T_20_39(23); T_20_39(24);\n \tT_20_39(25); T_20_39(26); T_20_39(27); T_20_39(28); T_20_39(29);\n \tT_20_39(30); T_20_39(31); T_20_39(32); T_20_39(33); T_20_39(34);\n \tT_20_39(35); T_20_39(36); T_20_39(37); T_20_39(38); T_20_39(39);\n+#else\n+\tfor (int t = 20; t <= 39; t++) { T_20_39(t); }\n+#endif\n \n #define T_40_59(t) \\\n \tSHA_XOR(t); \\\n-\tTEMP += SHA_ROL(A,5) + ((B&C)|(D&(B|C))) + E + 0x8f1bbcdc; \\\n-\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n+\tTEMP += SHA_ROL(A,5) + ((B&C)|(D&(B|C))) + E; \\\n+\tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP + 0x8f1bbcdc;\n \n+#if UNROLL\n \tT_40_59(40); T_40_59(41); T_40_59(42); T_40_59(43); T_40_59(44);\n \tT_40_59(45); T_40_59(46); T_40_59(47); T_40_59(48); T_40_59(49);\n \tT_40_59(50); T_40_59(51); T_40_59(52); T_40_59(53); T_40_59(54);\n \tT_40_59(55); T_40_59(56); T_40_59(57); T_40_59(58); T_40_59(59);\n+#else\n+\tfor (int t = 40; t <= 59; t++) { T_40_59(t); }\n+#endif\n \n #define T_60_79(t) \\\n \tSHA_XOR(t); \\\n \tTEMP += SHA_ROL(A,5) + (B^C^D) + E + 0xca62c1d6; \\\n \tE = D; D = C; C = SHA_ROR(B, 2); B = A; A = TEMP;\n \n+#if UNROLL\n \tT_60_79(60); T_60_79(61); T_60_79(62); T_60_79(63); T_60_79(64);\n \tT_60_79(65); T_60_79(66); T_60_79(67); T_60_79(68); T_60_79(69);\n \tT_60_79(70); T_60_79(71); T_60_79(72); T_60_79(73); T_60_79(74);\n \tT_60_79(75); T_60_79(76); T_60_79(77); T_60_79(78); T_60_79(79);\n+#else\n+\tfor (int t = 60; t <= 79; t++) { T_60_79(t); }\n+#endif\n \n \tctx->H[0] += A;\n \tctx->H[1] += B;\n"},{"id":"119752","messageId":"4A7A7074.1060506@gmail.com","threadId":"20241","inReplyTo":"4A7A6DBC.9010107@gmail.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2009-08-06T05:56:04Z","receivedAt":"2009-08-06T05:56:04Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"Artur Skawina wrote:\n> i did try (obvious patch below) and in fact the loops still win on p4:\n> \n> #Initializing... Rounds: 1000000, size: 62500K, time: 1.428s, speed: 42.76MB/s\n> #             TIME[s] SPEED[MB/s]\n> rfc3174         1.437       42.47\n> rfc3174         1.438       42.45\n> linus          0.5791       105.4\n> linusas        0.5052       120.8\n> mozilla         1.525       40.01\n> mozillaas       1.192       51.19\n\nand my atom seems to like the compact loops too: \n\n#Initializing... Rounds: 1000000, size: 62500K, time: 4.379s, speed: 13.94MB/s\n#             TIME[s] SPEED[MB/s]\nrfc3174         4.429       13.78\nrfc3174         4.414       13.83\nlinus           1.733       35.22\nlinusas           1.5        40.7\nmozilla         2.818       21.66\nmozillaas       2.539       24.04\n\nartur\n"},{"id":"119755","messageId":"20090806070312.13791.qmail@science.horizon.com","threadId":"20241","inReplyTo":"4A7A67C5.8060109@gmail.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"George Spelvin","fromEmail":"linux@horizon.com","sentAt":"2009-08-06T07:03:12Z","receivedAt":"2009-08-06T07:03:12Z","isPatch":false,"sender":{"key":"linux@horizon.com","avatar":null},"body":"> On Thu, 6 Aug 2009, Artur Skawina wrote:\n>> #             TIME[s] SPEED[MB/s]\n>> rfc3174         1.357       44.99\n>> rfc3174         1.352       45.13\n>> mozilla         1.509       40.44\n>> mozillaas       1.133       53.87\n>> linus          0.5818       104.9\n\n> #Initializing... Rounds: 1000000, size: 62500K, time: 1.421s, speed: 42.97MB/s\n> #             TIME[s] SPEED[MB/s]\n> rfc3174         1.403        43.5\n> # New hash result: b747042d9f4f1fdabd2ac53076f8f830dea7fe0f\n> rfc3174         1.403       43.51\n> linus          0.5891       103.6\n> linusas        0.5337       114.4\n> mozilla         1.535       39.76\n> mozillaas       1.128       54.13\n\nI'm trying to absorb what you're learning about P4 performance, but\nI'm getting confused... what is what in these benchmarks?\n\nThe major architectural decisions I see are:\n\n1) Three possible ways to compute the W[] array for rounds 16..79:\n\t1a) Compute W[16..79] in a loop beforehand (you noted that unrolling\n\t    two copies helped significantly.)\n\t1b) Compute W[16..79] as part of hash rounds 16..79.\n\t1c) Compute W[0..15] in-place as part of hash rounds 16..79\n\n2) The main hashing can be rolled up or unrolled:\n\t2a) Four 20-round loops.  (In case of options 1b and 1c, the\n\t    first one might be split into a 16 and a 4.)\n\t2b) Four 4-round loops, each unrolled 5x.  (See the ARM assembly.)\n\t2c) all 80 rounds unrolled.\n\nAs Linus noted, 1c is not friends with options 2a and 2b, because the\nW() indexing math is not longer a compile-time constant.\n\nLinus has posted 1a+2c and 1c+2c.  You posted some code that could be\n2a or 2c depending on an UNROLL preprocessor #define.  Which combinations\nare your \"linus\" and \"linusas\" code?\n\nYou talk about \"and my atom seems to like the compact loops too\", but\nI'm not sure which loops those are.\n\nThanks.\n"},{"id":"119757","messageId":"4A7A89FD.4040800@gmail.com","threadId":"20241","inReplyTo":"4A7A7074.1060506@gmail.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Artur Skawina","fromEmail":"art.08.09@gmail.com","sentAt":"2009-08-06T07:45:01Z","receivedAt":"2009-08-06T07:45:01Z","isPatch":false,"sender":{"key":"art.08.09@gmail.com","avatar":null},"body":"Artur Skawina wrote:\n> \n> and my atom seems to like the compact loops too: \n\nno, that was wrong, i forgot to turn off the ondemand governor...\n\nthe unrolled loops are in fact much faster and the numbers\nlook more reasonable, after a few tweaks even on a P4.\nNow i just need to check how well it does compared to the\nasm implementations...\n\nartur\n\n#             TIME[s] SPEED[MB/s]\n# ATOM\nrfc3174         2.199       27.75\nlinus          0.8642       70.62\nlinusas         1.606       38.01\nlinusas2       0.8763       69.65\nmozilla         2.813        21.7\nmozillaas       2.539       24.04\n# P4\nrfc3174         1.402       43.53\nlinus          0.5835       104.6\nlinusas        0.4625         132\nlinusas2       0.4456         137\nmozilla         1.529       39.91\nmozillaas       1.131       53.96\n# P3\nrfc3174         5.019       12.16\nlinus            1.86       32.81\nlinusas         3.108       19.64\nlinusas2        1.812       33.68\nmozilla         6.431        9.49\nmozillaas       5.868        10.4\n"},{"id":"119813","messageId":"40aa078e0908061149s3d08bcc5qbd86bfa4e5624006@mail.gmail.com","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908051800030.3390@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Erik Faye-Lund","fromEmail":"kusmabite@googlemail.com","sentAt":"2009-08-06T18:49:29Z","receivedAt":"2009-08-06T18:49:29Z","isPatch":false,"sender":{"key":"kusmabite@gmail.com","avatar":"https://avatars.githubusercontent.com/u/47073?v=4"},"body":"On Thu, Aug 6, 2009 at 3:18 AM, Linus\nTorvalds<torvalds@linux-foundation.org> wrote:\n> I note that MINGW does NO_OPENSSL by default, for example, and maybe the\n> MINGW people want to test the patch out and enable BLK_SHA1 rather than\n> the original Mozilla one.\n\nWe recently got OpenSSL in msysgit. The NO_OPENSSL-switch hasn't been\nflipped yet, though. (We did OpenSSL to get https-support in cURL...)\n\n-- \nErik \"kusma\" Faye-Lund\nkusmabite@gmail.com\n(+47) 986 59 656\n"},{"id":"121167","messageId":"4A8B1C77.5080003@fy.chalmers.se","threadId":"20241","inReplyTo":"20090803034741.23415.qmail@science.horizon.com","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Andy Polyakov","fromEmail":"appro@fy.chalmers.se","sentAt":"2009-08-18T21:26:15Z","receivedAt":"2009-08-18T21:26:15Z","isPatch":false,"sender":{"key":"appro@fy.chalmers.se","avatar":null},"body":"George Spelvin wrote:\n> (Work in progress, state dump to mailing list archives.)\n> \n> This started when discussing git startup overhead due to the dynamic\n> linker.  One big contributor is the openssl library, which is used only\n> for its optimized x86 SHA-1 implementation.  So I took a look at it,\n> with an eye to importing the code directly into the git source tree,\n> and decided that I felt like trying to do better.\n> \n> The original code was excellent, but it was optimized when the P4 was new.\n\nEven though last revision took place when \"the P4 was new\" and even\ntriggered by its appearance, *all-round* performance was and will always\nbe the prime goal. This means that improvements on some particular\nmicro-architecture is always weighed against losses on others [and\ncompromise is considered of so required]. Please note that I'm *not*\ntrying to diminish George's effort by saying that proposed code is\ninappropriate, on the contrary I'm nothing but grateful! Thanks, George!\nI'm only saying that it will be given thorough consideration. Well, I've\nactually given the consideration and outcome is already committed:-) See\nhttp://cvs.openssl.org/chngview?cn=18513. I don't deliver +17%, only\n+12%, but at the cost of Intel Atom-specific optimizations. I used this\nopportunity to optimize even for Intel Atom core, something I was\nplanning to do at some point anyway...\n\n>   http://www.openssl.org/~appro/cryptogams/cryptogams-0.tar.gz\n> - \"tar xz cryptogams-0.tar.gz\"\n\nIf there is interest I can pack new tar ball with updated modules.\n\n> An open question is how to add appropriate CPU detection to the git\n> build scripts. (Note that `uname -m`, which it currently uses to select\n> the ARM code, does NOT produce the right answer if you're using a 32-bit\n> compiler on a 64-bit platform.)\n\nIt's not only that. As next subscriber noted problem on MacOS X, it\n[MacOS X] uses slightly different assembler convention and ELF modules\ncan't be compiled on MacOS X. OpenSSL perlasm takes care of several\nassembler flavors and executable formats, including MacOS X. I'm talking\nabout\n\n> +++ Makefile\t2009-08-02 06:44:44.000000000 -0400\n> +%.s : %.pl x86asm.pl x86unix.pl\n> +\tperl $< elf > $@\n                ^^^ this argument.\n\nCheers. A.\n"},{"id":"121174","messageId":"4A8B2216.6080607@fy.chalmers.se","threadId":"20241","inReplyTo":"alpine.LFD.2.01.0908031938280.3270@localhost.localdomain","subject":"Re: x86 SHA1: Faster than OpenSSL","fromName":"Andy Polyakov","fromEmail":"appro@fy.chalmers.se","sentAt":"2009-08-18T21:50:14Z","receivedAt":"2009-08-18T21:50:14Z","isPatch":false,"sender":{"key":"appro@fy.chalmers.se","avatar":null},"body":"> On Mon, 3 Aug 2009, Linus Torvalds wrote:\n>> The thing that I'd prefer is simply\n>>\n>> \tgit fsck --full\n>>\n>> on the Linux kernel archive. For me (with a fast machine), it takes about \n>> 4m30s with the OpenSSL SHA1, and takes 6m40s with the Mozilla SHA1 (ie \n>> using a NO_OPENSSL=1 build).\n>>\n>> So that's an example of a load that is actually very sensitive to SHA1 \n>> performance (more so than _most_ git loads, I suspect), and at the same \n>> time is a real git load rather than some SHA1-only microbenchmark.\n\nI couldn't agree more that real-life benchmarks are of greater value\nthan specific algorithm micro-benchmark. And given the provided\nprofiling data one can argue that +17% (or my +12%) improvement on\nmicro-benchmark aren't really worth bothering about. But it's kind of\nsport [at least for me], so don't judge too harshly:-)\n\n>> It also \n>> shows very clearly why we default to the OpenSSL version over the Mozilla \n>> one.\n\nAs George implicitly mentioned most OpenSSL assembler modules are\navailable under more permissive license and if there is interest I'm\nready to assist...\n\n> \"perf report --sort comm,dso,symbol\" profiling shows the following for \n> 'git fsck --full' on the kernel repo, using the Mozilla SHA1:\n> \n>     47.69%               git  /home/torvalds/git/git     [.] moz_SHA1_Update\n>     22.98%               git  /lib64/libz.so.1.2.3       [.] inflate_fast\n>      7.32%               git  /lib64/libc-2.10.1.so      [.] __GI_memcpy\n>      4.66%               git  /lib64/libz.so.1.2.3       [.] inflate\n>      3.76%               git  /lib64/libz.so.1.2.3       [.] adler32\n>      2.86%               git  /lib64/libz.so.1.2.3       [.] inflate_table\n>      2.41%               git  /home/torvalds/git/git     [.] lookup_object\n>      1.31%               git  /lib64/libc-2.10.1.so      [.] _int_malloc\n>      0.84%               git  /home/torvalds/git/git     [.] patch_delta\n>      0.78%               git  [kernel]                   [k] hpet_next_event\n> \n> so yeah, SHA1 performance matters. Judging by the OpenSSL numbers, the \n> OpenSSL SHA1 implementation must be about twice as fast as the C version \n> we use.\n\nAnd given /lib64 path this is 64-bit C compiler-generated code compared\nto 32-bit assembler? Either way in this context I have extra comment\naddressing previous subscriber, Mark Lodato, who effectively wondered\nhow would 64-bit assembler compare to 32-bit one. First of all there\n*is* even 64-bit assembler version. But as SHA1 is essentially 32-bit\nalgorithm, 64-bit implementation is only nominally faster, +20% at most.\nFaster thanks to larger register bank facilitating more efficient\ninstruction scheduling.\n\nCheers. A.\n"}]}