Re: Revised PPC assembly implementation
- From
Paul Mackerras <paulus@samba.org>
- Date
- Apr 25, 2005, 09:40 UTC
- Message-ID
- <17004.47876.414.756912@cargo.ozlabs.ibm.com>
- In-Reply-To
- <20050425031337.16605.qmail@science.horizon.com>
linux@horizon.com writes:
Show 7 quoted lines
> Three changes: > - Added stack frame as per your description. > - Found two bugs. (Cutting & pasting too fast.) Fixed. > - Minor scheduling improvements. More to come. > > Which lead to three questions: > - Is the stack set properly now?
Not quite; you are saving 20 registers, so you need a 96-byte stack frame, like this:
stwu %r1,-96(%r1) stmw %r13,16(%r1) ... lmw %r13,16(%r1) addi %r1,%r1,96 blr
Since sha1_core is a leaf function, I suppose you could use the lr save area (do stwu %r1,-80(%r1); stmw %r13,0(%r1)) but it seems a bit dodgy.
> - Does it produce the right answer now?
Yes.
> - Is it any faster?
I did 10 repetitions of my program that calls SHA1_Update with a 4096-byte block of zeroes 256,000 times. With my version, the average time was 4.6191 seconds with a standard deviation of 0.0157. With your version, the average was 4.6063 and the standard deviation 0.0148. So I would say that your version is probably just a little faster - of the order of 0.3% faster.
Paul.