{"thread":{"id":"7550","subject":"[RFC/PATCH] Optimized PowerPC SHA1 generation for Darwin (OS X)","startedAt":"2007-04-06T23:48:26Z","lastAt":"2007-04-10T13:00:50Z","messageCount":7,"participants":["Arjen Laarhoven","Junio C Hamano","Linus Torvalds","Karl Hasselström"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"38821","messageId":"20070406234826.GG3854@regex.yaph.org","threadId":"7550","inReplyTo":null,"subject":"[RFC/PATCH] Optimized PowerPC SHA1 generation for Darwin (OS X)","fromName":"Arjen Laarhoven","fromEmail":"arjen@yaph.org","sentAt":"2007-04-06T23:48:26Z","receivedAt":"2007-04-06T23:48:26Z","isPatch":true,"sender":{"key":"arjen@yaph.org","avatar":"https://gravatar.com/avatar/f776c2c0c5ea62d70827b942eb7d95ce85661a3d70bc3f03cf9773815c599c01?d=mp&s=160"},"body":"The compiler toolchain supplied by Apple's Xcode environment has an old\nversion (1.38) of the GNU assembler.  It cannot assemble the optimized\nppc/sha1ppc.S file.  ppc/sha1ppc.S was rewritten into a Perl script\nwhich outputs the same code, but valid for the Xcode assembler.\n\nSigned-off-by: Arjen Laarhoven <arjen@yaph.org>\n---\n Makefile                     |   15 +++-\n ppc/darwin/darwin_ppc_gen.pl |  211 ++++++++++++++++++++++++++++++++++++++++++\n ppc/{ => linux}/sha1ppc.S    |    0 \n 3 files changed, 223 insertions(+), 3 deletions(-)\n create mode 100755 ppc/darwin/darwin_ppc_gen.pl\n rename ppc/{ => linux}/sha1ppc.S (100%)\n\ndiff --git a/Makefile b/Makefile\nindex b159ffd..a91fa2a 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -587,9 +587,13 @@ ifdef OLD_ICONV\n \tBASIC_CFLAGS += -DOLD_ICONV\n endif\n \n-ifdef PPC_SHA1\n+ifdef PPC_SHA1_LINUX\n \tSHA1_HEADER = \"ppc/sha1.h\"\n-\tLIB_OBJS += ppc/sha1.o ppc/sha1ppc.o\n+\tLIB_OBJS += ppc/sha1.o ppc/linux/sha1ppc.o\n+else\n+ifdef PPC_SHA1_DARWIN\n+\tSHA1_HEADER = \"ppc/sha1.h\"\n+\tLIB_OBJS += ppc/sha1.o ppc/darwin/sha1ppc.o\n else\n ifdef ARM_SHA1\n \tSHA1_HEADER = \"arm/sha1.h\"\n@@ -604,6 +608,7 @@ else\n endif\n endif\n endif\n+endif\n ifdef NO_PERL_MAKEMAKER\n \texport NO_PERL_MAKEMAKER\n endif\n@@ -620,6 +625,7 @@ endif\n ifneq ($(findstring $(MAKEFLAGS),s),s)\n ifndef V\n \tQUIET_CC       = @echo '   ' CC $@;\n+\tQUIET_AS       = @echo '   ' AS $@; \n \tQUIET_AR       = @echo '   ' AR $@;\n \tQUIET_LINK     = @echo '   ' LINK $@;\n \tQUIET_BUILT_IN = @echo '   ' BUILTIN $@;\n@@ -780,6 +786,9 @@ exec_cmd.o: exec_cmd.c GIT-CFLAGS\n builtin-init-db.o: builtin-init-db.c GIT-CFLAGS\n \t$(QUIET_CC)$(CC) -o $*.o -c $(ALL_CFLAGS) -DDEFAULT_GIT_TEMPLATE_DIR='\"$(template_dir_SQ)\"' $<\n \n+ppc/darwin/sha1ppc.S:\n+\t$(QUIET_GEN)$(PERL_PATH) ppc/darwin/darwin_ppc_gen.pl > $@\n+\n http.o: http.c GIT-CFLAGS\n \t$(QUIET_CC)$(CC) -o $*.o -c $(ALL_CFLAGS) -DGIT_USER_AGENT='\"git/$(GIT_VERSION)\"' $<\n \n@@ -962,7 +971,7 @@ dist-doc:\n ### Cleaning rules\n \n clean:\n-\trm -f *.o mozilla-sha1/*.o arm/*.o ppc/*.o compat/*.o xdiff/*.o \\\n+\trm -f *.o mozilla-sha1/*.o arm/*.o ppc/*.o ppc/darwin/*.[os] ppc/linux/*.o compat/*.o xdiff/*.o \\\n \t\ttest-chmtime$X $(LIB_FILE) $(XDIFF_LIB)\n \trm -f $(ALL_PROGRAMS) $(BUILT_INS) git$X\n \trm -f *.spec *.pyc *.pyo */*.pyc */*.pyo common-cmds.h TAGS tags\ndiff --git a/ppc/darwin/darwin_ppc_gen.pl b/ppc/darwin/darwin_ppc_gen.pl\nnew file mode 100755\nindex 0000000..346cd71\n--- /dev/null\n+++ b/ppc/darwin/darwin_ppc_gen.pl\n@@ -0,0 +1,211 @@\n+#!/usr/bin/perl\n+\n+# This script generates the PowerPC assembly code for optimized SHA-1\n+# hash generation on Darwin (Mac OS X).  It is a rewrite of the original\n+# ppc/sha1ppc.S file.\n+#\n+# The original ppc/sha1ppc.S cannot be assembled with the toolchain\n+# supplied with Xcode, as the assembler is (based on) GNU as version\n+# 1.38.  The problem is basically that the 1.38 assembler doesn't\n+# understand the computed register numbers used in the macros and\n+# register numbers without the 'r'.  This script acts as preprocessor\n+# and evaluates the # expressions for the register numbers and outputs\n+# the final correct # assembly for the 1.38 assembler.\n+\n+use strict;\n+use warnings;\n+\n+\n+sub RA { my ($t) = @_; 'r'.((($t)+4)%5+6) }\n+sub RB { my ($t) = @_; 'r'.((($t)+3)%5+6) }\n+sub RC { my ($t) = @_; 'r'.((($t)+2)%5+6) }\n+sub RD { my ($t) = @_; 'r'.((($t)+1)%5+6) }\n+sub RE { my ($t) = @_; 'r'.((($t)+0)%5+6) }\n+sub W  { my ($t) = @_; 'r'.(($t)%16+11)   }\n+\n+sub LOADW { my $s = shift; return \"\\tlwz \".W($s).','.($s)*4 .'(r4)'; }\n+\n+sub STEPD0_LOAD {\n+    my ($t, $s) = @_;\n+\n+    return join \"\\n\",\n+        \"\\tadd \".RE($t).','.RE($t).','.W($t),\n+        \"\\tandc r0,\".RD($t).','.RB($t),\n+        \"\\tand \".W($s).','.RC($t).','.RB($t),\n+        \"\\tadd \".RE($t).','.RE($t).',r0',\n+        \"\\trotlwi r0,\".RA($t).',5',\n+        \"\\trotlwi \".RB($t).','.RB($t).',30',\n+        \"\\tadd \".RE($t).','.RE($t).','.W($s),\n+        \"\\tadd r0,r0,r5\",\n+        \"\\tlwz \".W($s).','.($s)*4 .'(r4)',\n+        \"\\tadd \".RE($t).','.RE($t).',r0';\n+}\n+\n+sub STEPD0_UPDATE {\n+    my ($t, $s, $loadk) = @_;\n+\n+    return join \"\\n\",\n+    \"\\tadd \".RE($t).','.RE($t).','.W($t),\n+    \"\\tandc r0,\".RD($t).','.RB($t),\n+    \"\\txor \".W($s).','.W(($s)-16).','.W(($s)-3),\n+    \"\\tadd \".RE($t).','.RE($t).',r0',\n+    \"\\tand r0,\".RC($t).','.RB($t),\n+    \"\\txor \".W($s).','.W($s).','.W(($s)-8),\n+    \"\\tadd \".RE($t).','.RE($t).',r0',\n+    \"\\trotlwi r0,\".RA($t).',5',\n+    \"\\txor \".W($s).','.W($s).','.W(($s)-14),\n+    \"\\tadd \".RE($t).','.RE($t).',r5',\n+    $loadk || (),\n+    \"\\trotlwi \".RB($t).','.RB($t).',30',\n+    \"\\trotlwi \".W($s).','.W($s).',1',\n+    \"\\tadd \".RE($t).','.RE($t).',r0';\n+}\n+\n+sub STEPD1_UPDATE {\n+    my ($t, $s, $loadk) = @_;\n+\n+    return join \"\\n\",\n+        \"\\tadd \".RE($t).','.RE($t).','.W($t),\n+        \"\\txor r0,\".RD($t).','.RB($t),\n+        \"\\txor \".W($s).','.W(($s)-16).','.W(($s)-3),\n+        \"\\tadd \".RE($t).','.RE($t).',r5',\n+        $loadk || (),\n+        \"\\txor r0,r0,\".RC($t),\n+        \"\\txor \".W($s).','.W($s).','.W(($s)-8),\n+        \"\\tadd \".RE($t).','.RE($t).',r0',\n+        \"\\trotlwi r0,\".RA($t).',5',\n+        \"\\txor \".W($s).','.W($s).','.W(($s)-14),\n+        \"\\tadd \".RE($t).','.RE($t).',r0',\n+        \"\\trotlwi \".RB($t).','.RB($t).',30',\n+        \"\\trotlwi \".W($s).','.W($s).',1';\n+}\n+\n+sub STEPD1 {\n+    my ($t) = @_;\n+\n+    return join \"\\n\",\n+        \"\\tadd \".RE($t).','.RE($t).','.W($t),\n+        \"\\txor r0,\".RD($t).','.RB($t),\n+        \"\\trotlwi \".RB($t).','.RB($t).',30',\n+        \"\\tadd \".RE($t).','.RE($t).',r5',\n+        \"\\txor r0,r0,\".RC($t),\n+        \"\\tadd \".RE($t).','.RE($t).',r0',\n+        \"\\trotlwi r0,\".RA($t).',5',\n+        \"\\tadd \".RE($t).','.RE($t).',r0';\n+}\n+\n+sub STEPD2_UPDATE {\n+    my ($t, $s, $loadk) = @_;\n+\n+    return join \"\\n\",\n+        \"\\tadd \".RE($t).','.RE($t).','.W($t),\n+        \"\\tand r0,\".RD($t).','.RB($t),\n+        \"\\txor \".W($s).','.W(($s)-16).','.W(($s)-3),\n+        \"\\tadd \".RE($t).','.RE($t).',r0',\n+        \"\\txor r0,\".RD($t).','.RB($t),\n+        \"\\txor \".W($s).','.W($s).','.W(($s)-8),\n+        \"\\tadd \".RE($t).','.RE($t).',r5',\n+        $loadk || (),\n+        \"\\tand r0,r0,\".RC($t),\n+        \"\\txor \".W($s).','.W($s).','.W(($s)-14),\n+        \"\\tadd \".RE($t).','.RE($t).',r0',\n+        \"\\trotlwi r0,\".RA($t).',5',\n+        \"\\trotlwi \".W($s).','.W($s).',1',\n+        \"\\tadd \".RE($t).','.RE($t).',r0',\n+        \"\\trotlwi \".RB($t).','.RB($t).',30',\n+}\n+\n+sub STEP0_LOAD4 {\n+    my ($t, $s) = @_;\n+\n+    return join \"\\n\",\n+        STEPD0_LOAD($t, $s),\n+        STEPD0_LOAD($t+1, $s+1),\n+        STEPD0_LOAD($t+2, $s+2),\n+        STEPD0_LOAD($t+3, $s+3);\n+}\n+\n+sub STEPUP4 {\n+    my ($fn, $t, $s, $loadk) = @_;\n+\n+    no strict 'refs';\n+    return join \"\\n\",\n+        &{'STEP' . $fn . '_UPDATE'}($t, $s),\n+        &{'STEP' . $fn . '_UPDATE'}($t+1, $s+1),\n+        &{'STEP' . $fn . '_UPDATE'}($t+2, $s+2),\n+        &{'STEP' . $fn . '_UPDATE'}($t+3, $s+3, $loadk),\n+}\n+\n+sub STEPUP20 {\n+    my ($fn, $t, $s, $loadk) = @_;\n+\n+    return join \"\\n\",\n+        STEPUP4($fn, $t, $s),\n+        STEPUP4($fn, $t+4, $s+4),\n+        STEPUP4($fn, $t+8, $s+8),\n+        STEPUP4($fn, $t+12, $s+12),\n+        STEPUP4($fn, $t+16, $s+16, $loadk),\n+}\n+\n+print <<'EOA';\n+        .globl  _sha1_core\n+_sha1_core:\n+        stwu    r1,-80(r1)\n+        stmw    r13,4(r1)\n+\n+        /* Load up A - E */\n+        lmw     r27,0(r3)\n+\n+        mtctr   r5\n+\n+1:\n+EOA\n+\n+print LOADW(0).\"\\n\";\n+print \"\\tlis r5,0x5a82\\n\";\n+print \"\\tmr \".RE(0).\",r31\\n\";\n+print LOADW(1).\"\\n\";\n+print \"\\tmr \".RD(0).\",r30\\n\";\n+print \"\\tmr \".RC(0).\",r29\\n\";\n+print LOADW(2).\"\\n\";\n+print \"\\tori r5,r5,0x7999\\n\";\n+print \"\\tmr \".RB(0).\",r28\\n\";\n+print LOADW(3).\"\\n\";\n+print \"\\tmr \".RA(0).\",r27\\n\";\n+\n+print STEP0_LOAD4(0, 4).\"\\n\";\n+print STEP0_LOAD4(4, 8).\"\\n\";\n+print STEP0_LOAD4(8, 12).\"\\n\";\n+print STEPUP4(\"D0\", 12, 16,).\"\\n\";\n+print STEPUP4(\"D0\", 16, 20, \"lis r5,0x6ed9\").\"\\n\";\n+\n+print \"\\tori r5,r5,0xeba1\\n\";\n+print STEPUP20(\"D1\", 20, 24, \"lis r5,0x8f1b\").\"\\n\";\n+\n+print \"\\tori r5,r5,0xbcdc\\n\";\n+print STEPUP20(\"D2\", 40, 44, \"lis r5,0xca62\").\"\\n\";\n+\n+print \"\\tori r5,r5,0xc1d6\\n\";\n+print STEPUP4(\"D1\", 60, 64,).\"\\n\";\n+print STEPUP4(\"D1\", 64, 68,).\"\\n\";\n+print STEPUP4(\"D1\", 68, 72,).\"\\n\";\n+print STEPUP4(\"D1\", 72, 76,).\"\\n\";\n+print \"\\taddi r4,r4,64\\n\";\n+print STEPD1(76).\"\\n\";\n+print STEPD1(77).\"\\n\";\n+print STEPD1(78).\"\\n\";\n+print STEPD1(79).\"\\n\";\n+\n+print \"\\tadd r31,r31,\".RE(0).\"\\n\";\n+print \"\\tadd r30,r30,\".RD(0).\"\\n\";\n+print \"\\tadd r29,r29,\".RC(0).\"\\n\";\n+print \"\\tadd r28,r28,\".RB(0).\"\\n\";\n+print \"\\tadd r27,r27,\".RA(0).\"\\n\";\n+\n+print \"\\tbdnz 1b\\n\";\n+\n+print \"\\tstmw r27,0(r3)\\n\";\n+print \"\\tlmw  r13,4(r1)\\n\";\n+print \"\\taddi r1,r1,80\\n\";\n+print \"\\tblr\\n\";\n+\ndiff --git a/ppc/sha1ppc.S b/ppc/linux/sha1ppc.S\nsimilarity index 100%\nrename from ppc/sha1ppc.S\nrename to ppc/linux/sha1ppc.S\n-- \n1.5.1.rc3.29.gd8b6\n"},{"id":"38830","messageId":"7vodm1ggdm.fsf@assigned-by-dhcp.cox.net","threadId":"7550","inReplyTo":"20070406234826.GG3854@regex.yaph.org","subject":"Re: [RFC/PATCH] Optimized PowerPC SHA1 generation for Darwin (OS X)","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-04-07T00:47:17Z","receivedAt":"2007-04-07T00:47:17Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"arjen@yaph.org (Arjen Laarhoven) writes:\n\n> The compiler toolchain supplied by Apple's Xcode environment has an old\n> version (1.38) of the GNU assembler.  It cannot assemble the optimized\n> ppc/sha1ppc.S file.  ppc/sha1ppc.S was rewritten into a Perl script\n> which outputs the same code, but valid for the Xcode assembler.\n>\n> Signed-off-by: Arjen Laarhoven <arjen@yaph.org>\n\nGaah.\n\nWhen there are improvements/fixes to the sha1ppc.S side, how are\nyou going to keep that in sync with darwin_ppc_gen.pl?  If that\nscript *_gen.pl were a postprocessor that munges CPP output from\nsha1ppc.S to make it assemblable with an old assembler, it would\nbe one thing.  But this looks horrible.\n"},{"id":"38831","messageId":"Pine.LNX.4.64.0704061830350.6730@woody.linux-foundation.org","threadId":"7550","inReplyTo":"20070406234826.GG3854@regex.yaph.org","subject":"Re: [RFC/PATCH] Optimized PowerPC SHA1 generation for Darwin (OS X)","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-04-07T01:40:53Z","receivedAt":"2007-04-07T01:40:53Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 7 Apr 2007, Arjen Laarhoven wrote:\n>\n> The compiler toolchain supplied by Apple's Xcode environment has an old\n> version (1.38) of the GNU assembler.  It cannot assemble the optimized\n> ppc/sha1ppc.S file.  ppc/sha1ppc.S was rewritten into a Perl script\n> which outputs the same code, but valid for the Xcode assembler.\n\nUgh. That's just too ugly.\n\nThe Linux version of the GNU assembler can certainly take the same limited \ninput as the old Apple one. \n\nSo how about instea dof having two totally different versions of this \nfile, just having *one*, and having a pre-processor that turns it into \nsomething that is acceptable to both?\n\nAnd yes, it could be your perl script, except your perl script is ugly as \n*hell*. The old C preprocessor code is much nicer than your perl script \nthat does \"print\" statements.\n\nHow about something like the following instead?\n\n (a) make the register macros expand to something easily \n     greppable/parseable\n (b) have a *separate* preprocessor phase that actually then takes that \n     pattern, and evaluates it to a numeric value.\n (c) assemble the end result\n\nThe (a) part is trivial. Just a patch like the appended will make sure \nthat all the registers are now written as \"REG[int-expression]\", and then \nall you need is a perl-script or something that can trigger on the regexp\n\n\t\"REG\\[\\([^]]*\\)\\]\"\n\nand replace that regex with\n\n\t\"%eval(\\1)\"\n\nwhich is somethign that perl should be designed for.\n\nThat way you just have *one* source file (the \"sha1ppc.S\" one), which is \nreadable, and a simple script to then evaluate the register numbers \nstatically instead of expecting that the assembler can do it (since the \nApple one apparently cannot).\n\nSo it would just require somebody who knows perl. What's a one-liner perl \nscript to turn a line like\n\n\tadd REG[((0)+0)%5+6],REG[((0)+0)%5+6],REG[(0)%16+11];\n\ninto\n\n\tadd %6,%6,%11\n\n(ie it just evaluated the expression inside the [] things, and replaced it \nwith the \"%<num>\" string)?\n\n<Taunting mode>Or maybe perl can't do that in a single line!</Taunting mode>\n\n\t\tLinus\n\n---\ndiff --git a/ppc/sha1ppc.S b/ppc/sha1ppc.S\nindex f132696..cc554a4 100644\n--- a/ppc/sha1ppc.S\n+++ b/ppc/sha1ppc.S\n@@ -32,14 +32,14 @@\n  * We use registers 6 - 10 for this.  (Registers 27 - 31 hold\n  * the previous values.)\n  */\n-#define RA(t)\t(((t)+4)%5+6)\n-#define RB(t)\t(((t)+3)%5+6)\n-#define RC(t)\t(((t)+2)%5+6)\n-#define RD(t)\t(((t)+1)%5+6)\n-#define RE(t)\t(((t)+0)%5+6)\n+#define RA(t)\tREG[((t)+4)%5+6]\n+#define RB(t)\tREG[((t)+3)%5+6]\n+#define RC(t)\tREG[((t)+2)%5+6]\n+#define RD(t)\tREG[((t)+1)%5+6]\n+#define RE(t)\tREG[((t)+0)%5+6]\n \n /* We use registers 11 - 26 for the W values */\n-#define W(t)\t((t)%16+11)\n+#define W(t)\tREG[(t)%16+11]\n \n /* Register 5 is used for the constant k */\n \n"},{"id":"38886","messageId":"20070408200939.GL3854@regex.yaph.org","threadId":"7550","inReplyTo":"Pine.LNX.4.64.0704061830350.6730@woody.linux-foundation.org","subject":"Re: [RFC/PATCH] Optimized PowerPC SHA1 generation for Darwin (OS X)","fromName":"Arjen Laarhoven","fromEmail":"arjen@yaph.org","sentAt":"2007-04-08T20:09:39Z","receivedAt":"2007-04-08T20:09:39Z","isPatch":true,"sender":{"key":"arjen@yaph.org","avatar":"https://gravatar.com/avatar/f776c2c0c5ea62d70827b942eb7d95ce85661a3d70bc3f03cf9773815c599c01?d=mp&s=160"},"body":"Hi,\n\nOn Fri, Apr 06, 2007 at 06:40:53PM -0700, Linus Torvalds wrote:\n> \n> \n> On Sat, 7 Apr 2007, Arjen Laarhoven wrote:\n> >\n> > The compiler toolchain supplied by Apple's Xcode environment has an old\n> > version (1.38) of the GNU assembler.  It cannot assemble the optimized\n> > ppc/sha1ppc.S file.  ppc/sha1ppc.S was rewritten into a Perl script\n> > which outputs the same code, but valid for the Xcode assembler.\n> \n> Ugh. That's just too ugly.\n\nYes.  Very.  I should've reworked it before sending it to the list.  Ah\nwell.\n\n> The Linux version of the GNU assembler can certainly take the same limited \n> input as the old Apple one. \n> \n> So how about instea dof having two totally different versions of this \n> file, just having *one*, and having a pre-processor that turns it into \n> something that is acceptable to both?\n\nThat is of course the best way to handle it.  See the patch below for\nthe reworked solution.\n\n[snip excellent pointers]\n\n> So it would just require somebody who knows perl. What's a one-liner perl \n> script to turn a line like\n> \n> \tadd REG[((0)+0)%5+6],REG[((0)+0)%5+6],REG[(0)%16+11];\n> \n> into\n> \n> \tadd %6,%6,%11\n> \n> (ie it just evaluated the expression inside the [] things, and replaced it \n> with the \"%<num>\" string)?\n> \n> <Taunting mode>Or maybe perl can't do that in a single line!</Taunting mode>\n\nOf course it can! :-P\n\nBut there are some other issues like the underscore prefix of the symbol\nin the assembly and the inability of Apple's assembler to handle\nmultiple statements per line.  So for the sake of maintainability I've\nput it in its own file, and even turned on warnings and strict ;-)\n\nI don't have access to a Linux/PPC machine, so it could very well need\nsome tweaking.  Someone with a Linux/PPC box want to give it a try?\n\n---snip---\nOptimized PowerPC SHA-1 calculation for Darwin\n\nThe compiler toolchain from Apple's Xcode environment uses an old\nversion (1.38) of the GNU assembler which cannot assemble the\noptimized SHA-1 calculation in ppc/sha1ppc.S.  The main problem is the\nuse of calculated register numbers which gas 1.38 doesn't understand.\n\nTo create valid assembly code the registers in ppc/sha1ppc.in.S are\nrepresented by R[<register number>].  sha1ppc.in.S is postprocessed by\ngen_sha1ppc.pl to generate valid assembly code for gas 1.38.\n\nSigned-off-by: Arjen Laarhoven <arjen@yaph.org>\n---\n Makefile                        |    7 ++-\n ppc/gen_sha1ppc.pl              |   19 +++++++\n ppc/{sha1ppc.S => sha1ppc.in.S} |  110 +++++++++++++++++++-------------------\n 3 files changed, 79 insertions(+), 57 deletions(-)\n create mode 100644 ppc/gen_sha1ppc.pl\n rename ppc/{sha1ppc.S => sha1ppc.in.S} (70%)\n\ndiff --git a/Makefile b/Makefile\nindex ac29c62..01b69e7 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -825,7 +825,7 @@ git$X git.spec \\\n \n %.o: %.c GIT-CFLAGS\n \t$(QUIET_CC)$(CC) -o $*.o -c $(ALL_CFLAGS) $<\n-%.o: %.S\n+%.o: %.s\n \t$(QUIET_CC)$(CC) -o $*.o -c $(ALL_CFLAGS) $<\n \n exec_cmd.o: exec_cmd.c GIT-CFLAGS\n@@ -836,6 +836,9 @@ builtin-init-db.o: builtin-init-db.c GIT-CFLAGS\n http.o: http.c GIT-CFLAGS\n \t$(QUIET_CC)$(CC) -o $*.o -c $(ALL_CFLAGS) -DGIT_USER_AGENT='\"git/$(GIT_VERSION)\"' $<\n \n+ppc/sha1ppc.s: ppc/sha1ppc.in.S\n+\t$(QUIET_CC)$(CC) -c -E $< | $(PERL_PATH) ppc/gen_sha1ppc.pl > $@\n+\n ifdef NO_EXPAT\n http-fetch.o: http-fetch.c http.h GIT-CFLAGS\n \t$(QUIET_CC)$(CC) -o $*.o -c $(ALL_CFLAGS) -DNO_EXPAT $<\n@@ -1032,7 +1035,7 @@ dist-doc:\n ### Cleaning rules\n \n clean:\n-\trm -f *.o mozilla-sha1/*.o arm/*.o ppc/*.o compat/*.o xdiff/*.o \\\n+\trm -f *.o mozilla-sha1/*.o arm/*.o ppc/*.[so] compat/*.o xdiff/*.o \\\n \t\ttest-chmtime$X $(LIB_FILE) $(XDIFF_LIB)\n \trm -f $(ALL_PROGRAMS) $(BUILT_INS) git$X\n \trm -f *.spec *.pyc *.pyo */*.pyc */*.pyo common-cmds.h TAGS tags\ndiff --git a/ppc/gen_sha1ppc.pl b/ppc/gen_sha1ppc.pl\nnew file mode 100644\nindex 0000000..79ba1a1\n--- /dev/null\n+++ b/ppc/gen_sha1ppc.pl\n@@ -0,0 +1,19 @@\n+#!/usr/bin/perl -w\n+\n+use strict;\n+\n+my %platform = (\n+    # Special extra substitutions that have to be done on this platform\n+    darwin => sub {\n+        s{sha1_core}{_sha1_core};\n+        s{;}{\\n}g;\n+    },\n+);\n+\n+my $extra = exists $platform{$^O} ? $platform{$^O} : sub {};\n+\n+while (<>) {\n+    $extra->();\n+    s{R\\[([^]]+)\\]}{'r'.eval\"$1\"}ge;\n+    print;\n+}\ndiff --git a/ppc/sha1ppc.S b/ppc/sha1ppc.in.S\nsimilarity index 70%\nrename from ppc/sha1ppc.S\nrename to ppc/sha1ppc.in.S\nindex f132696..11bc2e0 100644\n--- a/ppc/sha1ppc.S\n+++ b/ppc/sha1ppc.in.S\n@@ -32,14 +32,14 @@\n  * We use registers 6 - 10 for this.  (Registers 27 - 31 hold\n  * the previous values.)\n  */\n-#define RA(t)\t(((t)+4)%5+6)\n-#define RB(t)\t(((t)+3)%5+6)\n-#define RC(t)\t(((t)+2)%5+6)\n-#define RD(t)\t(((t)+1)%5+6)\n-#define RE(t)\t(((t)+0)%5+6)\n+#define RA(t)\tR[((t)+4)%5+6]\n+#define RB(t)\tR[((t)+3)%5+6]\n+#define RC(t)\tR[((t)+2)%5+6]\n+#define RD(t)\tR[((t)+1)%5+6]\n+#define RE(t)\tR[((t)+0)%5+6]\n \n /* We use registers 11 - 26 for the W values */\n-#define W(t)\t((t)%16+11)\n+#define W(t)\tR[(t)%16+11]\n \n /* Register 5 is used for the constant k */\n \n@@ -86,7 +86,7 @@\n \n /* the initial loads. */\n #define LOADW(s) \\\n-\tlwz\tW(s),(s)*4(%r4)\n+\tlwz\tW(s),(s)*4(R[4])\n \n /*\n  * Perform a step with F0, and load W(s).  Uses W(s) as a temporary\n@@ -97,10 +97,10 @@\n  * second line.)  Thus, two iterations take 7 cycles, 3.5 cycles per round.\n  */\n #define STEPD0_LOAD(t,s) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t);  and    W(s),RC(t),RB(t); \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;      rotlwi RB(t),RB(t),30;   \\\n-add RE(t),RE(t),W(s); add    %r0,%r0,%r5;      lwz    W(s),(s)*4(%r4);  \\\n-add RE(t),RE(t),%r0\n+add RE(t),RE(t),W(t); andc   R[0],RD(t),RB(t); and    W(s),RC(t),RB(t); \\\n+add RE(t),RE(t),R[0]; rotlwi R[0],RA(t),5;     rotlwi RB(t),RB(t),30;   \\\n+add RE(t),RE(t),W(s); add    R[0],R[0],R[5];   lwz    W(s),(s)*4(R[4]); \\\n+add RE(t),RE(t),R[0]\n \n /*\n  * This is likewise awkward, 13 instructions.  However, it can also\n@@ -108,28 +108,28 @@ add RE(t),RE(t),%r0\n  * in 9 cycles, 4.5 cycles/round.\n  */\n #define STEPD0_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  and    %r0,RC(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r5;  loadk; rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1;     \\\n-add RE(t),RE(t),%r0\n+add RE(t),RE(t),W(t); andc   R[0],RD(t),RB(t); xor   W(s),W((s)-16),W((s)-3); \\\n+add RE(t),RE(t),R[0]; and    R[0],RC(t),RB(t); xor   W(s),W(s),W((s)-8);      \\\n+add RE(t),RE(t),R[0]; rotlwi R[0],RA(t),5;     xor   W(s),W(s),W((s)-14);     \\\n+add RE(t),RE(t),R[5]; loadk; rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1;      \\\n+add RE(t),RE(t),R[0]\n \n /* Nicely optimal.  Conveniently, also the most common. */\n #define STEPD1_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r5;  loadk; xor %r0,%r0,RC(t);  xor W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1\n+add RE(t),RE(t),W(t); xor    R[0],RD(t),RB(t);    xor W(s),W((s)-16),W((s)-3); \\\n+add RE(t),RE(t),R[5]; loadk; xor R[0],R[0],RC(t); xor W(s),W(s),W((s)-8);    \\\n+add RE(t),RE(t),R[0]; rotlwi R[0],RA(t),5;   xor    W(s),W(s),W((s)-14);  \\\n+add RE(t),RE(t),R[0]; rotlwi RB(t),RB(t),30; rotlwi W(s),W(s),1\n \n /*\n  * The naked version, no UPDATE, for the last 4 rounds.  3 cycles per.\n  * We could use W(s) as a temp register, but we don't need it.\n  */\n #define STEPD1(t) \\\n-                        add   RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); \\\n-rotlwi RB(t),RB(t),30;  add   RE(t),RE(t),%r5;  xor    %r0,%r0,RC(t);   \\\n-add    RE(t),RE(t),%r0; rotlwi %r0,RA(t),5;     /* spare slot */        \\\n-add    RE(t),RE(t),%r0\n+                        add   RE(t),RE(t),W(t); xor    R[0],RD(t),RB(t); \\\n+rotlwi RB(t),RB(t),30;  add   RE(t),RE(t),R[5]; xor    R[0],R[0],RC(t);   \\\n+add    RE(t),RE(t),R[0]; rotlwi R[0],RA(t),5;     /* spare slot */        \\\n+add    RE(t),RE(t),R[0]\n \n /*\n  * 14 instructions, 5 cycles per.  The majority function is a bit\n@@ -137,11 +137,11 @@ add    RE(t),RE(t),%r0\n  * but it causes a 2-instruction delay, which triggers a stall.\n  */\n #define STEPD2_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); and    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  xor    %r0,RD(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r5;  loadk; and %r0,%r0,RC(t);  xor W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     rotlwi W(s),W(s),1;             \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30\n+add RE(t),RE(t),W(t); and    R[0],RD(t),RB(t); xor  W(s),W((s)-16),W((s)-3); \\\n+add RE(t),RE(t),R[0]; xor    R[0],RD(t),RB(t); xor  W(s),W(s),W((s)-8);      \\\n+add RE(t),RE(t),R[5]; loadk; and R[0],R[0],RC(t);  xor W(s),W(s),W((s)-14);  \\\n+add RE(t),RE(t),R[0]; rotlwi R[0],RA(t),5;     rotlwi W(s),W(s),1;           \\\n+add RE(t),RE(t),R[0]; rotlwi RB(t),RB(t),30\n \n #define STEP0_LOAD4(t,s)\t\t\\\n \tSTEPD0_LOAD(t,s);\t\t\\\n@@ -164,61 +164,61 @@ add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30\n \n \t.globl\tsha1_core\n sha1_core:\n-\tstwu\t%r1,-80(%r1)\n-\tstmw\t%r13,4(%r1)\n+\tstwu\tR[1],-80(R[1])\n+\tstmw\tR[13],4(R[1])\n \n \t/* Load up A - E */\n-\tlmw\t%r27,0(%r3)\n+\tlmw\tR[27],0(R[3])\n \n-\tmtctr\t%r5\n+\tmtctr\tR[5]\n \n 1:\n \tLOADW(0)\n-\tlis\t%r5,0x5a82\n-\tmr\tRE(0),%r31\n+\tlis\tR[5],0x5a82\n+\tmr\tRE(0),R[31]\n \tLOADW(1)\n-\tmr\tRD(0),%r30\n-\tmr\tRC(0),%r29\n+\tmr\tRD(0),R[30]\n+\tmr\tRC(0),R[29]\n \tLOADW(2)\n-\tori\t%r5,%r5,0x7999\t/* K0-19 */\n-\tmr\tRB(0),%r28\n+\tori\tR[5],R[5],0x7999\t/* K0-19 */\n+\tmr\tRB(0),R[28]\n \tLOADW(3)\n-\tmr\tRA(0),%r27\n+\tmr\tRA(0),R[27]\n \n \tSTEP0_LOAD4(0, 4)\n \tSTEP0_LOAD4(4, 8)\n \tSTEP0_LOAD4(8, 12)\n \tSTEPUP4(D0, 12, 16,)\n-\tSTEPUP4(D0, 16, 20, lis %r5,0x6ed9)\n+\tSTEPUP4(D0, 16, 20, lis R[5],0x6ed9)\n \n-\tori\t%r5,%r5,0xeba1\t/* K20-39 */\n-\tSTEPUP20(D1, 20, 24, lis %r5,0x8f1b)\n+\tori\tR[5],R[5],0xeba1\t/* K20-39 */\n+\tSTEPUP20(D1, 20, 24, lis R[5],0x8f1b)\n \n-\tori\t%r5,%r5,0xbcdc\t/* K40-59 */\n-\tSTEPUP20(D2, 40, 44, lis %r5,0xca62)\n+\tori\tR[5],R[5],0xbcdc\t/* K40-59 */\n+\tSTEPUP20(D2, 40, 44, lis R[5],0xca62)\n \n-\tori\t%r5,%r5,0xc1d6\t/* K60-79 */\n+\tori\tR[5],R[5],0xc1d6\t/* K60-79 */\n \tSTEPUP4(D1, 60, 64,)\n \tSTEPUP4(D1, 64, 68,)\n \tSTEPUP4(D1, 68, 72,)\n \tSTEPUP4(D1, 72, 76,)\n-\taddi\t%r4,%r4,64\n+\taddi\tR[4],R[4],64\n \tSTEPD1(76)\n \tSTEPD1(77)\n \tSTEPD1(78)\n \tSTEPD1(79)\n \n \t/* Add results to original values */\n-\tadd\t%r31,%r31,RE(0)\n-\tadd\t%r30,%r30,RD(0)\n-\tadd\t%r29,%r29,RC(0)\n-\tadd\t%r28,%r28,RB(0)\n-\tadd\t%r27,%r27,RA(0)\n+\tadd\tR[31],R[31],RE(0)\n+\tadd\tR[30],R[30],RD(0)\n+\tadd\tR[29],R[29],RC(0)\n+\tadd\tR[28],R[28],RB(0)\n+\tadd\tR[27],R[27],RA(0)\n \n \tbdnz\t1b\n \n \t/* Save final hash, restore registers, and return */\n-\tstmw\t%r27,0(%r3)\n-\tlmw\t%r13,4(%r1)\n-\taddi\t%r1,%r1,80\n+\tstmw\tR[27],0(R[3])\n+\tlmw\tR[13],4(R[1])\n+\taddi\tR[1],R[1],80\n \tblr\n-- \n1.5.1.rc3.29.gd8b6\n"},{"id":"38999","messageId":"20070410094801.GA6148@diana.vm.bytemark.co.uk","threadId":"7550","inReplyTo":"20070408200939.GL3854@regex.yaph.org","subject":"Re: [RFC/PATCH] Optimized PowerPC SHA1 generation for Darwin (OS X)","fromName":"Karl Hasselström","fromEmail":"kha@treskal.com","sentAt":"2007-04-10T09:48:01Z","receivedAt":"2007-04-10T09:48:01Z","isPatch":true,"sender":{"key":"kha@treskal.com","avatar":"https://gravatar.com/avatar/f0120c734b5279b345075a28521e1ac66acb20c9913ffe9bf6ae97e53f7f3f13?d=mp&s=160"},"body":"On 2007-04-08 22:09:39 +0200, Arjen Laarhoven wrote:\n\n>  ppc/{sha1ppc.S => sha1ppc.in.S} |  110 +++++++++++++++++++-------------------\n\nWouldn't it be prettier if this filename was .S.in instead of .in.S?\nAdditional file suffixes are usually added at the end (e.g. .tar.gz),\nand it makes more sense too.\n\n-- \nKarl Hasselström, kha@treskal.com\n      www.treskal.com/kalle\n"},{"id":"39001","messageId":"20070410114507.GA28728@regex.yaph.org","threadId":"7550","inReplyTo":"20070410094801.GA6148@diana.vm.bytemark.co.uk","subject":"Re: [RFC/PATCH] Optimized PowerPC SHA1 generation for Darwin (OS X)","fromName":"Arjen Laarhoven","fromEmail":"arjen@yaph.org","sentAt":"2007-04-10T11:45:07Z","receivedAt":"2007-04-10T11:45:07Z","isPatch":true,"sender":{"key":"arjen@yaph.org","avatar":"https://gravatar.com/avatar/f776c2c0c5ea62d70827b942eb7d95ce85661a3d70bc3f03cf9773815c599c01?d=mp&s=160"},"body":"Hi,\n\nOn Tue, Apr 10, 2007 at 11:48:01AM +0200, Karl Hasselstr?m wrote:\n> On 2007-04-08 22:09:39 +0200, Arjen Laarhoven wrote:\n> \n> >  ppc/{sha1ppc.S => sha1ppc.in.S} |  110 +++++++++++++++++++-------------------\n> \n> Wouldn't it be prettier if this filename was .S.in instead of .in.S?\n> Additional file suffixes are usually added at the end (e.g. .tar.gz),\n> and it makes more sense too.\n\nUsing the .S suffix makes gcc automatically do the right thing. .S.in\nrequires an extra '-x assembler-with-cpp' option to gcc.  Of course,\nit's trivial fix.\n\nArjen\n"},{"id":"39012","messageId":"20070410130050.GA10104@diana.vm.bytemark.co.uk","threadId":"7550","inReplyTo":"20070410114507.GA28728@regex.yaph.org","subject":"Re: [RFC/PATCH] Optimized PowerPC SHA1 generation for Darwin (OS X)","fromName":"Karl Hasselström","fromEmail":"kha@treskal.com","sentAt":"2007-04-10T13:00:50Z","receivedAt":"2007-04-10T13:00:50Z","isPatch":true,"sender":{"key":"kha@treskal.com","avatar":"https://gravatar.com/avatar/f0120c734b5279b345075a28521e1ac66acb20c9913ffe9bf6ae97e53f7f3f13?d=mp&s=160"},"body":"On 2007-04-10 13:45:07 +0200, Arjen Laarhoven wrote:\n\n> On Tue, Apr 10, 2007 at 11:48:01AM +0200, Karl Hasselström wrote:\n>\n> > On 2007-04-08 22:09:39 +0200, Arjen Laarhoven wrote:\n> >\n> > >  ppc/{sha1ppc.S => sha1ppc.in.S} |  110 +++++++++++++++++++-------------------\n> >\n> > Wouldn't it be prettier if this filename was .S.in instead of\n> > .in.S? Additional file suffixes are usually added at the end (e.g.\n> > .tar.gz), and it makes more sense too.\n>\n> Using the .S suffix makes gcc automatically do the right thing.\n> .S.in requires an extra '-x assembler-with-cpp' option to gcc. Of\n> course, it's trivial fix.\n\nI just read the Makefile changes again, a bit slower this time, and\nnoticed that you _first_ feed the .in.S file to gcc, and _then_ to the\nperl script, instead of the other way around like I was expecting.\nWith that arrangement, your naming makes sense, since it reflects\nwhich file format is contained in which. Sorry for the noise.\n\n-- \nKarl Hasselström, kha@treskal.com\n      www.treskal.com/kalle\n"}]}