git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Revised PPC assembly implementation

From
Llinux@horizon.com <linux@horizon.com>
Date
Apr 26, 2005, 02:30 UTC
Message-ID
<20050426023507.24611.qmail@science.horizon.com>
In-Reply-To
<20050425161746.7d943e62.davem@davemloft.net>

(Sorry about that last e-mail. gnome-terminal crashed and sent the file before I edited it. Here's what I meant to send.)

> Do a block with the integer ALUs in parallel with a block done using
> Altivec :-)  There should be enough spare insn slots so that the loads
> are absorbed properly.

Unfortunately, the blocks are connected by a data dependency. It's basically a large-key block cipher, chained by:

iv[] = fixed_initial_value. iv[] += encrypt(iv, text[0..63]) iv[] += encrypt(iv, text[64..127]) iv[] += encrypt(iv, text[128..191]) iv[] += encrypt(iv, text[192..255]) etc.

There is no coarse-grain parallelism to exploit, unless you want to be hashing two separate files at once. Which would do too much damage to the structure of the source to be worth considering.

> Unlike UltraSPARC's VIS, with altivec you can reasonably do shifts and
> rotates, which is the only reason I'm suggesting this.

I don't quite think it's worth it, though. It's not data-parallel enough.

We could theoretically use it to form the w[] vector, but that's only 4 instructions in registers which are very flexibly schedulable and nicely fill in the cracks between other instructions.

Oh, here's STEPD1+UPDATEW scheduled optimally for the G4. %r5 holds the constant K. Note that t < s <= t+16. W(s) and W((s)-16) are actually the same register.

add RE(t),RE(t),W(t); xor %r0,RD(t),RB(t); xor W(s),W((s)-16),W((s)-3); add RE(t),RE(t),%r5; xor %r0,%r0,RC(t); xor W(s),W(s),W((s)-8); add RE(t),RE(t),%r0; rotlwi %r0,RA(t),5; xor W(s),W(s),W((s)-14); add RE(t),RE(t),%r0; rotlwi RB(t),RB(t),30; rotlwi W(s),W(s),1;

However, whether that can be done in 6 cycles on a G5 is a bit unclear. It can't be 6 consecutive cycles, but with some motion of code across the edges, perhaps...

0: add RE(t),RE(t),W(t); xor %r0,RD(t),RB(t); 1: xor W(s),W((s)-16),W((s)-3); (add) 2: add RE(t),RE(t),%r5; xor %r0,%r0,RC(t); 3: xor W(s),W(s),W((s)-8); (rotlwi) 4: add RE(t),RE(t),%r0; rotlwi %r0,RA(t),5; 5: xor W(s),W(s),W((s)-14); rotlwi RB(t),RB(t),30; 6: 7: add RE(t),RE(t),%r0; 8: 9: rotlwi W(s),W(s),1;

The problem there is forcing that ordering, rather than issuing the final add in cycle 6 and pushing everything else ahead of it.

STEPD0+UPDATEW and STEPD1+UPDATEW are 13 and 14 instructions, respectively, and don't fit into a 3-issue machine as neatly.

Previous: linux@horizon.com
Message 17 of 17 in “Re: [PATCH] PPC assembly implementation of SHA1”
  1. linux@horizon.comApr 23, 2005
  2. linux@horizon.comApr 23, 2005
  3. Benjamin HerrenschmidtApr 24, 2005
  4. Paul MackerrasApr 24, 2005
  5. Wayne ScottApr 24, 2005
  6. linux@horizon.comApr 24, 2005
  7. Revised PPC assembly implementationlinux@horizon.com, Apr 25, 2005
  8. Paul MackerrasApr 25, 2005
  9. linux@horizon.comApr 25, 2005
  10. Paul MackerrasApr 25, 2005
  11. David S. MillerApr 25, 2005
  12. Paul MackerrasApr 26, 2005
  13. linux@horizon.comApr 27, 2005
  14. Paul MackerrasApr 27, 2005
  15. linux@horizon.comApr 27, 2005
  16. linux@horizon.comApr 26, 2005
  17. linux@horizon.comApr 26, 2005

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.