{"thread":{"id":"52786","subject":"SHA1dc on mac","startedAt":"2020-02-12T08:56:58Z","lastAt":"2022-09-01T15:48:53Z","messageCount":17,"participants":["Mike Hommey","Eric Sunshine","Junio C Hamano","Jeff King","Ævar Arnfjörð Bjarmason","brian m. carlson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"391597","messageId":"20200212085646.hgq3nv2lf4brbb3j@glandium.org","threadId":"52786","inReplyTo":null,"subject":"SHA1dc on mac","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2020-02-12T08:56:46Z","receivedAt":"2020-02-12T08:56:58Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"Hi,\n\nIf I'm not mistaken in my reading of the various files involved, it\nlooks like for some reason, building git on mac leads to using Apple\nCommon Crypto for SHA1, rather than SHA1dc, which seems unfortunate.\nIs that really expected? More generally, at this point, should anything\nother than SHA1dc be supported as a build option at all?\n\nMike\n"},{"id":"391603","messageId":"CAPig+cTfMx_kwUAxBRHp6kNSOtXsdsv=odUQSRYVpV21DnRuvA@mail.gmail.com","threadId":"52786","inReplyTo":"20200212085646.hgq3nv2lf4brbb3j@glandium.org","subject":"Re: SHA1dc on mac","fromName":"Eric Sunshine","fromEmail":"sunshine@sunshineco.com","sentAt":"2020-02-12T16:46:31Z","receivedAt":"2020-02-12T16:46:46Z","isPatch":false,"sender":{"key":"sunshine@sunshineco.com","avatar":"https://avatars.githubusercontent.com/u/163641?v=4"},"body":"On Wed, Feb 12, 2020 at 3:57 AM Mike Hommey <mh@glandium.org> wrote:\n> If I'm not mistaken in my reading of the various files involved, it\n> looks like for some reason, building git on mac leads to using Apple\n> Common Crypto for SHA1, rather than SHA1dc, which seems unfortunate.\n> Is that really expected?\n\nThere was a discussion on this topic a while back[1], and it does seem\nthat the behavior you describe is intentional[2].\n\n> More generally, at this point, should anything\n> other than SHA1dc be supported as a build option at all?\n\nThe conclusion [2,3] was that it likely would make sense to drop\nsupport for Apple's CommonCrypto altogether, although nobody has yet\nstepped up to do the work.\n\n[1]: https://lore.kernel.org/git/CAMYxyaVQyVRQb-b0nVv412tMZ3rEnOfUPRakg2dEREg5_Ba5Ag@mail.gmail.com/T/\n[2]: https://lore.kernel.org/git/20160102234923.GA14424@gmail.com/\n[3]: https://lore.kernel.org/git/CAPig+cQ5kKAt2_RQnqT7Rn=uGmHV9VvxpQ+UgDPOj=D=pq6arg@mail.gmail.com/\n"},{"id":"391627","messageId":"20200212223122.svthiwmh3js7i47b@glandium.org","threadId":"52786","inReplyTo":"CAPig+cTfMx_kwUAxBRHp6kNSOtXsdsv=odUQSRYVpV21DnRuvA@mail.gmail.com","subject":"Re: SHA1dc on mac","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2020-02-12T22:31:22Z","receivedAt":"2020-02-12T22:31:32Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Wed, Feb 12, 2020 at 11:46:31AM -0500, Eric Sunshine wrote:\n> On Wed, Feb 12, 2020 at 3:57 AM Mike Hommey <mh@glandium.org> wrote:\n> > If I'm not mistaken in my reading of the various files involved, it\n> > looks like for some reason, building git on mac leads to using Apple\n> > Common Crypto for SHA1, rather than SHA1dc, which seems unfortunate.\n> > Is that really expected?\n> \n> There was a discussion on this topic a while back[1], and it does seem\n> that the behavior you describe is intentional[2].\n\nThat discussion predates SHA1dc, though.\n\n> > More generally, at this point, should anything\n> > other than SHA1dc be supported as a build option at all?\n> \n> The conclusion [2,3] was that it likely would make sense to drop\n> support for Apple's CommonCrypto altogether, although nobody has yet\n> stepped up to do the work.\n\nI wasn't explicit in my question, but I meant more broadly than Apple\nCommon Crypto. There is still opt-in support for openssl sha1 and PPC\nsha1.\n\nMike\n"},{"id":"391629","messageId":"xmqqk14rtonu.fsf@gitster-ct.c.googlers.com","threadId":"52786","inReplyTo":"20200212223122.svthiwmh3js7i47b@glandium.org","subject":"Re: SHA1dc on mac","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2020-02-12T22:40:37Z","receivedAt":"2020-02-12T22:40:42Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Mike Hommey <mh@glandium.org> writes:\n\n> On Wed, Feb 12, 2020 at 11:46:31AM -0500, Eric Sunshine wrote:\n>> On Wed, Feb 12, 2020 at 3:57 AM Mike Hommey <mh@glandium.org> wrote:\n>> > If I'm not mistaken in my reading of the various files involved, it\n>> > looks like for some reason, building git on mac leads to using Apple\n>> > Common Crypto for SHA1, rather than SHA1dc, which seems unfortunate.\n>> > Is that really expected?\n>> \n>> There was a discussion on this topic a while back[1], and it does seem\n>> that the behavior you describe is intentional[2].\n>\n> That discussion predates SHA1dc, though.\n\nYes, but the essense is the same.  It was phrased as \"is there a\ngood reason to prefer CommonCrypto over block-sha1?\" but it really\nwas \"is there a good reason to prefer CommonCrypto over the best we\noffer?\"  And the best we offer, which used to be block-sha1, is now\nsha1dc.\n"},{"id":"392367","messageId":"20200223223758.120941-1-mh@glandium.org","threadId":"52786","inReplyTo":"xmqqk14rtonu.fsf@gitster-ct.c.googlers.com","subject":"[PATCH] Remove non-SHA1dc sha1 implementations","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2020-02-23T22:37:58Z","receivedAt":"2020-02-23T23:17:48Z","isPatch":true,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"It is 2020, and with the weakening of SHA1 security-wise, there doesn't\nseem to be a reason to support anything else than SHA1dc, with collision\ndetection.\n\nSigned-off-by: Mike Hommey <mh@glandium.org>\n---\n\nNote: I only tested building on Linux.\n\n INSTALL           |   5 -\n Makefile          |  67 ++-----------\n block-sha1/sha1.c | 251 ----------------------------------------------\n block-sha1/sha1.h |  22 ----\n config.mak.uname  |   1 -\n configure.ac      |   3 -\n hash.h            |  24 -----\n ppc/sha1.c        |  72 -------------\n ppc/sha1.h        |  25 -----\n ppc/sha1ppc.S     | 224 -----------------------------------------\n 10 files changed, 6 insertions(+), 688 deletions(-)\n delete mode 100644 block-sha1/sha1.c\n delete mode 100644 block-sha1/sha1.h\n delete mode 100644 ppc/sha1.c\n delete mode 100644 ppc/sha1.h\n delete mode 100644 ppc/sha1ppc.S\n\ndiff --git a/INSTALL b/INSTALL\nindex 22c364f34f..91d649f99e 100644\n--- a/INSTALL\n+++ b/INSTALL\n@@ -133,11 +133,6 @@ Issues of note:\n \t  you are using libcurl older than 7.34.0.  Otherwise you can use\n \t  NO_OPENSSL without losing git-imap-send.\n \n-\t  By default, git uses OpenSSL for SHA1 but it will use its own\n-\t  library (inspired by Mozilla's) with either NO_OPENSSL or\n-\t  BLK_SHA1.  Also included is a version optimized for PowerPC\n-\t  (PPC_SHA1).\n-\n \t- \"libcurl\" library is used by git-http-fetch, git-fetch, and, if\n \t  the curl version >= 7.34.0, for git-imap-send.  You might also\n \t  want the \"curl\" executable for debugging purposes. If you do not\ndiff --git a/Makefile b/Makefile\nindex b7d7374dac..5b4307d332 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -149,37 +149,15 @@ all::\n # specify your own (or DarwinPort's) include directories and\n # library directories by defining CFLAGS and LDFLAGS appropriately.\n #\n-# Define NO_APPLE_COMMON_CRYPTO if you are building on Darwin/Mac OS X\n-# and do not want to use Apple's CommonCrypto library.  This allows you\n-# to provide your own OpenSSL library, for example from MacPorts.\n-#\n-# Define BLK_SHA1 environment variable to make use of the bundled\n-# optimized C SHA1 routine.\n-#\n-# Define PPC_SHA1 environment variable when running make to make use of\n-# a bundled SHA1 routine optimized for PowerPC.\n-#\n-# Define DC_SHA1 to unconditionally enable the collision-detecting sha1\n-# algorithm. This is slower, but may detect attempted collision attacks.\n-# Takes priority over other *_SHA1 knobs.\n-#\n-# Define DC_SHA1_EXTERNAL in addition to DC_SHA1 if you want to build / link\n-# git with the external SHA1 collision-detect library.\n+# Define DC_SHA1_EXTERNAL if you want to build / link git with the\n+# external SHA1 collision-detect library.\n # Without this option, i.e. the default behavior is to build git with its\n # own built-in code (or submodule).\n #\n-# Define DC_SHA1_SUBMODULE in addition to DC_SHA1 to use the\n-# sha1collisiondetection shipped as a submodule instead of the\n-# non-submodule copy in sha1dc/. This is an experimental option used\n-# by the git project to migrate to using sha1collisiondetection as a\n-# submodule.\n-#\n-# Define OPENSSL_SHA1 environment variable when running make to link\n-# with the SHA1 routine from openssl library.\n-#\n-# Define SHA1_MAX_BLOCK_SIZE to limit the amount of data that will be hashed\n-# in one call to the platform's SHA1_Update(). e.g. APPLE_COMMON_CRYPTO\n-# wants 'SHA1_MAX_BLOCK_SIZE=1024L*1024L*1024L' defined.\n+# Define DC_SHA1_SUBMODULE to use the sha1collisiondetection shipped\n+# as a submodule instead of the non-submodule copy in sha1dc/. This is\n+# an experimental option used by the git project to migrate to using\n+# sha1collisiondetection as a submodule.\n #\n # Define BLK_SHA256 to use the built-in SHA-256 routines.\n #\n@@ -1296,11 +1274,6 @@ ifeq ($(uname_S),Darwin)\n \t\t\tBASIC_LDFLAGS += -L/opt/local/lib\n \t\tendif\n \tendif\n-\tifndef NO_APPLE_COMMON_CRYPTO\n-\t\tNO_OPENSSL = YesPlease\n-\t\tAPPLE_COMMON_CRYPTO = YesPlease\n-\t\tCOMPAT_CFLAGS += -DAPPLE_COMMON_CRYPTO\n-\tendif\n \tNO_REGEX = YesPlease\n \tPTHREAD_LIBS =\n endif\n@@ -1430,9 +1403,6 @@ ifdef NEEDS_SSL_WITH_CRYPTO\n else\n \tLIB_4_CRYPTO = $(OPENSSL_LINK) -lcrypto\n endif\n-ifdef APPLE_COMMON_CRYPTO\n-\tLIB_4_CRYPTO += -framework Security -framework CoreFoundation\n-endif\n endif\n ifndef NO_ICONV\n \tifdef NEEDS_LIBICONV\n@@ -1647,27 +1617,6 @@ ifdef NO_POSIX_GOODIES\n \tBASIC_CFLAGS += -DNO_POSIX_GOODIES\n endif\n \n-ifdef APPLE_COMMON_CRYPTO\n-\t# Apple CommonCrypto requires chunking\n-\tSHA1_MAX_BLOCK_SIZE = 1024L*1024L*1024L\n-endif\n-\n-ifdef OPENSSL_SHA1\n-\tEXTLIBS += $(LIB_4_CRYPTO)\n-\tBASIC_CFLAGS += -DSHA1_OPENSSL\n-else\n-ifdef BLK_SHA1\n-\tLIB_OBJS += block-sha1/sha1.o\n-\tBASIC_CFLAGS += -DSHA1_BLK\n-else\n-ifdef PPC_SHA1\n-\tLIB_OBJS += ppc/sha1.o ppc/sha1ppc.o\n-\tBASIC_CFLAGS += -DSHA1_PPC\n-else\n-ifdef APPLE_COMMON_CRYPTO\n-\tCOMPAT_CFLAGS += -DCOMMON_DIGEST_FOR_OPENSSL\n-\tBASIC_CFLAGS += -DSHA1_APPLE\n-else\n \tDC_SHA1 := YesPlease\n \tBASIC_CFLAGS += -DSHA1_DC\n \tLIB_OBJS += sha1dc_git.o\n@@ -1694,10 +1643,6 @@ endif\n \t\t-DSHA1DC_CUSTOM_INCLUDE_SHA1_C=\"\\\"cache.h\\\"\" \\\n \t\t-DSHA1DC_CUSTOM_INCLUDE_UBC_CHECK_C=\"\\\"git-compat-util.h\\\"\"\n endif\n-endif\n-endif\n-endif\n-endif\n \n ifdef OPENSSL_SHA256\n \tEXTLIBS += $(LIB_4_CRYPTO)\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\ndeleted file mode 100644\nindex 22b125cf8c..0000000000\n--- a/block-sha1/sha1.c\n+++ /dev/null\n@@ -1,251 +0,0 @@\n-/*\n- * SHA1 routine optimized to do word accesses rather than byte accesses,\n- * and to avoid unnecessary copies into the context array.\n- *\n- * This was initially based on the Mozilla SHA1 implementation, although\n- * none of the original Mozilla code remains.\n- */\n-\n-/* this is only to get definitions for memcpy(), ntohl() and htonl() */\n-#include \"../git-compat-util.h\"\n-\n-#include \"sha1.h\"\n-\n-#if defined(__GNUC__) && (defined(__i386__) || defined(__x86_64__))\n-\n-/*\n- * Force usage of rol or ror by selecting the one with the smaller constant.\n- * It _can_ generate slightly smaller code (a constant of 1 is special), but\n- * perhaps more importantly it's possibly faster on any uarch that does a\n- * rotate with a loop.\n- */\n-\n-#define SHA_ASM(op, x, n) ({ unsigned int __res; __asm__(op \" %1,%0\":\"=r\" (__res):\"i\" (n), \"0\" (x)); __res; })\n-#define SHA_ROL(x,n)\tSHA_ASM(\"rol\", x, n)\n-#define SHA_ROR(x,n)\tSHA_ASM(\"ror\", x, n)\n-\n-#else\n-\n-#define SHA_ROT(X,l,r)\t(((X) << (l)) | ((X) >> (r)))\n-#define SHA_ROL(X,n)\tSHA_ROT(X,n,32-(n))\n-#define SHA_ROR(X,n)\tSHA_ROT(X,32-(n),n)\n-\n-#endif\n-\n-/*\n- * If you have 32 registers or more, the compiler can (and should)\n- * try to change the array[] accesses into registers. However, on\n- * machines with less than ~25 registers, that won't really work,\n- * and at least gcc will make an unholy mess of it.\n- *\n- * So to avoid that mess which just slows things down, we force\n- * the stores to memory to actually happen (we might be better off\n- * with a 'W(t)=(val);asm(\"\":\"+m\" (W(t))' there instead, as\n- * suggested by Artur Skawina - that will also make gcc unable to\n- * try to do the silly \"optimize away loads\" part because it won't\n- * see what the value will be).\n- *\n- * Ben Herrenschmidt reports that on PPC, the C version comes close\n- * to the optimized asm with this (ie on PPC you don't want that\n- * 'volatile', since there are lots of registers).\n- *\n- * On ARM we get the best code generation by forcing a full memory barrier\n- * between each SHA_ROUND, otherwise gcc happily get wild with spilling and\n- * the stack frame size simply explode and performance goes down the drain.\n- */\n-\n-#if defined(__i386__) || defined(__x86_64__)\n-  #define setW(x, val) (*(volatile unsigned int *)&W(x) = (val))\n-#elif defined(__GNUC__) && defined(__arm__)\n-  #define setW(x, val) do { W(x) = (val); __asm__(\"\":::\"memory\"); } while (0)\n-#else\n-  #define setW(x, val) (W(x) = (val))\n-#endif\n-\n-/* This \"rolls\" over the 512-bit array */\n-#define W(x) (array[(x)&15])\n-\n-/*\n- * Where do we get the source from? The first 16 iterations get it from\n- * the input data, the next mix it from the 512-bit array.\n- */\n-#define SHA_SRC(t) get_be32((unsigned char *) block + (t)*4)\n-#define SHA_MIX(t) SHA_ROL(W((t)+13) ^ W((t)+8) ^ W((t)+2) ^ W(t), 1);\n-\n-#define SHA_ROUND(t, input, fn, constant, A, B, C, D, E) do { \\\n-\tunsigned int TEMP = input(t); setW(t, TEMP); \\\n-\tE += TEMP + SHA_ROL(A,5) + (fn) + (constant); \\\n-\tB = SHA_ROR(B, 2); } while (0)\n-\n-#define T_0_15(t, A, B, C, D, E)  SHA_ROUND(t, SHA_SRC, (((C^D)&B)^D) , 0x5a827999, A, B, C, D, E )\n-#define T_16_19(t, A, B, C, D, E) SHA_ROUND(t, SHA_MIX, (((C^D)&B)^D) , 0x5a827999, A, B, C, D, E )\n-#define T_20_39(t, A, B, C, D, E) SHA_ROUND(t, SHA_MIX, (B^C^D) , 0x6ed9eba1, A, B, C, D, E )\n-#define T_40_59(t, A, B, C, D, E) SHA_ROUND(t, SHA_MIX, ((B&C)+(D&(B^C))) , 0x8f1bbcdc, A, B, C, D, E )\n-#define T_60_79(t, A, B, C, D, E) SHA_ROUND(t, SHA_MIX, (B^C^D) ,  0xca62c1d6, A, B, C, D, E )\n-\n-static void blk_SHA1_Block(blk_SHA_CTX *ctx, const void *block)\n-{\n-\tunsigned int A,B,C,D,E;\n-\tunsigned int array[16];\n-\n-\tA = ctx->H[0];\n-\tB = ctx->H[1];\n-\tC = ctx->H[2];\n-\tD = ctx->H[3];\n-\tE = ctx->H[4];\n-\n-\t/* Round 1 - iterations 0-16 take their input from 'block' */\n-\tT_0_15( 0, A, B, C, D, E);\n-\tT_0_15( 1, E, A, B, C, D);\n-\tT_0_15( 2, D, E, A, B, C);\n-\tT_0_15( 3, C, D, E, A, B);\n-\tT_0_15( 4, B, C, D, E, A);\n-\tT_0_15( 5, A, B, C, D, E);\n-\tT_0_15( 6, E, A, B, C, D);\n-\tT_0_15( 7, D, E, A, B, C);\n-\tT_0_15( 8, C, D, E, A, B);\n-\tT_0_15( 9, B, C, D, E, A);\n-\tT_0_15(10, A, B, C, D, E);\n-\tT_0_15(11, E, A, B, C, D);\n-\tT_0_15(12, D, E, A, B, C);\n-\tT_0_15(13, C, D, E, A, B);\n-\tT_0_15(14, B, C, D, E, A);\n-\tT_0_15(15, A, B, C, D, E);\n-\n-\t/* Round 1 - tail. Input from 512-bit mixing array */\n-\tT_16_19(16, E, A, B, C, D);\n-\tT_16_19(17, D, E, A, B, C);\n-\tT_16_19(18, C, D, E, A, B);\n-\tT_16_19(19, B, C, D, E, A);\n-\n-\t/* Round 2 */\n-\tT_20_39(20, A, B, C, D, E);\n-\tT_20_39(21, E, A, B, C, D);\n-\tT_20_39(22, D, E, A, B, C);\n-\tT_20_39(23, C, D, E, A, B);\n-\tT_20_39(24, B, C, D, E, A);\n-\tT_20_39(25, A, B, C, D, E);\n-\tT_20_39(26, E, A, B, C, D);\n-\tT_20_39(27, D, E, A, B, C);\n-\tT_20_39(28, C, D, E, A, B);\n-\tT_20_39(29, B, C, D, E, A);\n-\tT_20_39(30, A, B, C, D, E);\n-\tT_20_39(31, E, A, B, C, D);\n-\tT_20_39(32, D, E, A, B, C);\n-\tT_20_39(33, C, D, E, A, B);\n-\tT_20_39(34, B, C, D, E, A);\n-\tT_20_39(35, A, B, C, D, E);\n-\tT_20_39(36, E, A, B, C, D);\n-\tT_20_39(37, D, E, A, B, C);\n-\tT_20_39(38, C, D, E, A, B);\n-\tT_20_39(39, B, C, D, E, A);\n-\n-\t/* Round 3 */\n-\tT_40_59(40, A, B, C, D, E);\n-\tT_40_59(41, E, A, B, C, D);\n-\tT_40_59(42, D, E, A, B, C);\n-\tT_40_59(43, C, D, E, A, B);\n-\tT_40_59(44, B, C, D, E, A);\n-\tT_40_59(45, A, B, C, D, E);\n-\tT_40_59(46, E, A, B, C, D);\n-\tT_40_59(47, D, E, A, B, C);\n-\tT_40_59(48, C, D, E, A, B);\n-\tT_40_59(49, B, C, D, E, A);\n-\tT_40_59(50, A, B, C, D, E);\n-\tT_40_59(51, E, A, B, C, D);\n-\tT_40_59(52, D, E, A, B, C);\n-\tT_40_59(53, C, D, E, A, B);\n-\tT_40_59(54, B, C, D, E, A);\n-\tT_40_59(55, A, B, C, D, E);\n-\tT_40_59(56, E, A, B, C, D);\n-\tT_40_59(57, D, E, A, B, C);\n-\tT_40_59(58, C, D, E, A, B);\n-\tT_40_59(59, B, C, D, E, A);\n-\n-\t/* Round 4 */\n-\tT_60_79(60, A, B, C, D, E);\n-\tT_60_79(61, E, A, B, C, D);\n-\tT_60_79(62, D, E, A, B, C);\n-\tT_60_79(63, C, D, E, A, B);\n-\tT_60_79(64, B, C, D, E, A);\n-\tT_60_79(65, A, B, C, D, E);\n-\tT_60_79(66, E, A, B, C, D);\n-\tT_60_79(67, D, E, A, B, C);\n-\tT_60_79(68, C, D, E, A, B);\n-\tT_60_79(69, B, C, D, E, A);\n-\tT_60_79(70, A, B, C, D, E);\n-\tT_60_79(71, E, A, B, C, D);\n-\tT_60_79(72, D, E, A, B, C);\n-\tT_60_79(73, C, D, E, A, B);\n-\tT_60_79(74, B, C, D, E, A);\n-\tT_60_79(75, A, B, C, D, E);\n-\tT_60_79(76, E, A, B, C, D);\n-\tT_60_79(77, D, E, A, B, C);\n-\tT_60_79(78, C, D, E, A, B);\n-\tT_60_79(79, B, C, D, E, A);\n-\n-\tctx->H[0] += A;\n-\tctx->H[1] += B;\n-\tctx->H[2] += C;\n-\tctx->H[3] += D;\n-\tctx->H[4] += E;\n-}\n-\n-void blk_SHA1_Init(blk_SHA_CTX *ctx)\n-{\n-\tctx->size = 0;\n-\n-\t/* Initialize H with the magic constants (see FIPS180 for constants) */\n-\tctx->H[0] = 0x67452301;\n-\tctx->H[1] = 0xefcdab89;\n-\tctx->H[2] = 0x98badcfe;\n-\tctx->H[3] = 0x10325476;\n-\tctx->H[4] = 0xc3d2e1f0;\n-}\n-\n-void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n-{\n-\tunsigned int lenW = ctx->size & 63;\n-\n-\tctx->size += len;\n-\n-\t/* Read the data into W and process blocks as they get full */\n-\tif (lenW) {\n-\t\tunsigned int left = 64 - lenW;\n-\t\tif (len < left)\n-\t\t\tleft = len;\n-\t\tmemcpy(lenW + (char *)ctx->W, data, left);\n-\t\tlenW = (lenW + left) & 63;\n-\t\tlen -= left;\n-\t\tdata = ((const char *)data + left);\n-\t\tif (lenW)\n-\t\t\treturn;\n-\t\tblk_SHA1_Block(ctx, ctx->W);\n-\t}\n-\twhile (len >= 64) {\n-\t\tblk_SHA1_Block(ctx, data);\n-\t\tdata = ((const char *)data + 64);\n-\t\tlen -= 64;\n-\t}\n-\tif (len)\n-\t\tmemcpy(ctx->W, data, len);\n-}\n-\n-void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n-{\n-\tstatic const unsigned char pad[64] = { 0x80 };\n-\tunsigned int padlen[2];\n-\tint i;\n-\n-\t/* Pad with a binary 1 (ie 0x80), then zeroes, then length */\n-\tpadlen[0] = htonl((uint32_t)(ctx->size >> 29));\n-\tpadlen[1] = htonl((uint32_t)(ctx->size << 3));\n-\n-\ti = ctx->size & 63;\n-\tblk_SHA1_Update(ctx, pad, 1 + (63 & (55 - i)));\n-\tblk_SHA1_Update(ctx, padlen, 8);\n-\n-\t/* Output hash */\n-\tfor (i = 0; i < 5; i++)\n-\t\tput_be32(hashout + i * 4, ctx->H[i]);\n-}\ndiff --git a/block-sha1/sha1.h b/block-sha1/sha1.h\ndeleted file mode 100644\nindex 4df6747752..0000000000\n--- a/block-sha1/sha1.h\n+++ /dev/null\n@@ -1,22 +0,0 @@\n-/*\n- * SHA1 routine optimized to do word accesses rather than byte accesses,\n- * and to avoid unnecessary copies into the context array.\n- *\n- * This was initially based on the Mozilla SHA1 implementation, although\n- * none of the original Mozilla code remains.\n- */\n-\n-typedef struct {\n-\tunsigned long long size;\n-\tunsigned int H[5];\n-\tunsigned int W[16];\n-} blk_SHA_CTX;\n-\n-void blk_SHA1_Init(blk_SHA_CTX *ctx);\n-void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *dataIn, unsigned long len);\n-void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx);\n-\n-#define platform_SHA_CTX\tblk_SHA_CTX\n-#define platform_SHA1_Init\tblk_SHA1_Init\n-#define platform_SHA1_Update\tblk_SHA1_Update\n-#define platform_SHA1_Final\tblk_SHA1_Final\ndiff --git a/config.mak.uname b/config.mak.uname\nindex cc8efd95b1..785b9265c3 100644\n--- a/config.mak.uname\n+++ b/config.mak.uname\n@@ -116,7 +116,6 @@ ifeq ($(uname_S),Darwin)\n \t# i.e. \"begins with [15678] and a dot\" means \"10.4.* or older\".\n \tifeq ($(shell expr \"$(uname_R)\" : '[15678]\\.'),2)\n \t\tOLD_ICONV = UnfortunatelyYes\n-\t\tNO_APPLE_COMMON_CRYPTO = YesPlease\n \tendif\n \tifeq ($(shell expr \"$(uname_R)\" : '[15]\\.'),2)\n \t\tNO_STRLCPY = YesPlease\ndiff --git a/configure.ac b/configure.ac\nindex 66aedb9288..dd39b7ecdb 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -237,9 +237,6 @@ AC_MSG_NOTICE([CHECKS for site configuration])\n # tests.  These tests take up a significant amount of the total test time\n # but are not needed unless you plan to talk to SVN repos.\n #\n-# Define PPC_SHA1 environment variable when running make to make use of\n-# a bundled SHA1 routine optimized for PowerPC.\n-#\n # Define NO_OPENSSL environment variable if you do not have OpenSSL.\n #\n # Define OPENSSLDIR=/foo/bar if your openssl header and library files are in\ndiff --git a/hash.h b/hash.h\nindex 52a4f1a3f4..e1a3d00b13 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -3,17 +3,7 @@\n \n #include \"git-compat-util.h\"\n \n-#if defined(SHA1_PPC)\n-#include \"ppc/sha1.h\"\n-#elif defined(SHA1_APPLE)\n-#include <CommonCrypto/CommonDigest.h>\n-#elif defined(SHA1_OPENSSL)\n-#include <openssl/sha.h>\n-#elif defined(SHA1_DC)\n #include \"sha1dc_git.h\"\n-#else /* SHA1_BLK */\n-#include \"block-sha1/sha1.h\"\n-#endif\n \n #if defined(SHA256_GCRYPT)\n #include \"sha256/gcrypt.h\"\n@@ -23,20 +13,6 @@\n #include \"sha256/block/sha256.h\"\n #endif\n \n-#ifndef platform_SHA_CTX\n-/*\n- * platform's underlying implementation of SHA-1; could be OpenSSL,\n- * blk_SHA, Apple CommonCrypto, etc...  Note that the relevant\n- * SHA-1 header may have already defined platform_SHA_CTX for our\n- * own implementations like block-sha1 and ppc-sha1, so we list\n- * the default for OpenSSL compatible SHA-1 implementations here.\n- */\n-#define platform_SHA_CTX\tSHA_CTX\n-#define platform_SHA1_Init\tSHA1_Init\n-#define platform_SHA1_Update\tSHA1_Update\n-#define platform_SHA1_Final    \tSHA1_Final\n-#endif\n-\n #define git_SHA_CTX\t\tplatform_SHA_CTX\n #define git_SHA1_Init\t\tplatform_SHA1_Init\n #define git_SHA1_Update\t\tplatform_SHA1_Update\ndiff --git a/ppc/sha1.c b/ppc/sha1.c\ndeleted file mode 100644\nindex 1b705cee1f..0000000000\n--- a/ppc/sha1.c\n+++ /dev/null\n@@ -1,72 +0,0 @@\n-/*\n- * SHA-1 implementation.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- *\n- * This version assumes we are running on a big-endian machine.\n- * It calls an external sha1_core() to process blocks of 64 bytes.\n- */\n-#include <stdio.h>\n-#include <string.h>\n-#include \"sha1.h\"\n-\n-void ppc_sha1_core(uint32_t *hash, const unsigned char *p,\n-\t\t   unsigned int nblocks);\n-\n-int ppc_SHA1_Init(ppc_SHA_CTX *c)\n-{\n-\tc->hash[0] = 0x67452301;\n-\tc->hash[1] = 0xEFCDAB89;\n-\tc->hash[2] = 0x98BADCFE;\n-\tc->hash[3] = 0x10325476;\n-\tc->hash[4] = 0xC3D2E1F0;\n-\tc->len = 0;\n-\tc->cnt = 0;\n-\treturn 0;\n-}\n-\n-int ppc_SHA1_Update(ppc_SHA_CTX *c, const void *ptr, unsigned long n)\n-{\n-\tunsigned long nb;\n-\tconst unsigned char *p = ptr;\n-\n-\tc->len += (uint64_t) n << 3;\n-\twhile (n != 0) {\n-\t\tif (c->cnt || n < 64) {\n-\t\t\tnb = 64 - c->cnt;\n-\t\t\tif (nb > n)\n-\t\t\t\tnb = n;\n-\t\t\tmemcpy(&c->buf.b[c->cnt], p, nb);\n-\t\t\tif ((c->cnt += nb) == 64) {\n-\t\t\t\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\t\t\t\tc->cnt = 0;\n-\t\t\t}\n-\t\t} else {\n-\t\t\tnb = n >> 6;\n-\t\t\tppc_sha1_core(c->hash, p, nb);\n-\t\t\tnb <<= 6;\n-\t\t}\n-\t\tn -= nb;\n-\t\tp += nb;\n-\t}\n-\treturn 0;\n-}\n-\n-int ppc_SHA1_Final(unsigned char *hash, ppc_SHA_CTX *c)\n-{\n-\tunsigned int cnt = c->cnt;\n-\n-\tc->buf.b[cnt++] = 0x80;\n-\tif (cnt > 56) {\n-\t\tif (cnt < 64)\n-\t\t\tmemset(&c->buf.b[cnt], 0, 64 - cnt);\n-\t\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\t\tcnt = 0;\n-\t}\n-\tif (cnt < 56)\n-\t\tmemset(&c->buf.b[cnt], 0, 56 - cnt);\n-\tc->buf.l[7] = c->len;\n-\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\tmemcpy(hash, c->hash, 20);\n-\treturn 0;\n-}\ndiff --git a/ppc/sha1.h b/ppc/sha1.h\ndeleted file mode 100644\nindex 9b24b32615..0000000000\n--- a/ppc/sha1.h\n+++ /dev/null\n@@ -1,25 +0,0 @@\n-/*\n- * SHA-1 implementation.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- */\n-#include <stdint.h>\n-\n-typedef struct {\n-\tuint32_t hash[5];\n-\tuint32_t cnt;\n-\tuint64_t len;\n-\tunion {\n-\t\tunsigned char b[64];\n-\t\tuint64_t l[8];\n-\t} buf;\n-} ppc_SHA_CTX;\n-\n-int ppc_SHA1_Init(ppc_SHA_CTX *c);\n-int ppc_SHA1_Update(ppc_SHA_CTX *c, const void *p, unsigned long n);\n-int ppc_SHA1_Final(unsigned char *hash, ppc_SHA_CTX *c);\n-\n-#define platform_SHA_CTX\tppc_SHA_CTX\n-#define platform_SHA1_Init\tppc_SHA1_Init\n-#define platform_SHA1_Update\tppc_SHA1_Update\n-#define platform_SHA1_Final\tppc_SHA1_Final\ndiff --git a/ppc/sha1ppc.S b/ppc/sha1ppc.S\ndeleted file mode 100644\nindex 1711eef6e7..0000000000\n--- a/ppc/sha1ppc.S\n+++ /dev/null\n@@ -1,224 +0,0 @@\n-/*\n- * SHA-1 implementation for PowerPC.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- */\n-\n-/*\n- * PowerPC calling convention:\n- * %r0 - volatile temp\n- * %r1 - stack pointer.\n- * %r2 - reserved\n- * %r3-%r12 - Incoming arguments & return values; volatile.\n- * %r13-%r31 - Callee-save registers\n- * %lr - Return address, volatile\n- * %ctr - volatile\n- *\n- * Register usage in this routine:\n- * %r0 - temp\n- * %r3 - argument (pointer to 5 words of SHA state)\n- * %r4 - argument (pointer to data to hash)\n- * %r5 - Constant K in SHA round (initially number of blocks to hash)\n- * %r6-%r10 - Working copies of SHA variables A..E (actually E..A order)\n- * %r11-%r26 - Data being hashed W[].\n- * %r27-%r31 - Previous copies of A..E, for final add back.\n- * %ctr - loop count\n- */\n-\n-\n-/*\n- * We roll the registers for A, B, C, D, E around on each\n- * iteration; E on iteration t is D on iteration t+1, and so on.\n- * We use registers 6 - 10 for this.  (Registers 27 - 31 hold\n- * the previous values.)\n- */\n-#define RA(t)\t(((t)+4)%5+6)\n-#define RB(t)\t(((t)+3)%5+6)\n-#define RC(t)\t(((t)+2)%5+6)\n-#define RD(t)\t(((t)+1)%5+6)\n-#define RE(t)\t(((t)+0)%5+6)\n-\n-/* We use registers 11 - 26 for the W values */\n-#define W(t)\t((t)%16+11)\n-\n-/* Register 5 is used for the constant k */\n-\n-/*\n- * The basic SHA-1 round function is:\n- * E += ROTL(A,5) + F(B,C,D) + W[i] + K;  B = ROTL(B,30)\n- * Then the variables are renamed: (A,B,C,D,E) = (E,A,B,C,D).\n- *\n- * Every 20 rounds, the function F() and the constant K changes:\n- * - 20 rounds of f0(b,c,d) = \"bit wise b ? c : d\" =  (^b & d) + (b & c)\n- * - 20 rounds of f1(b,c,d) = b^c^d = (b^d)^c\n- * - 20 rounds of f2(b,c,d) = majority(b,c,d) = (b&d) + ((b^d)&c)\n- * - 20 more rounds of f1(b,c,d)\n- *\n- * These are all scheduled for near-optimal performance on a G4.\n- * The G4 is a 3-issue out-of-order machine with 3 ALUs, but it can only\n- * *consider* starting the oldest 3 instructions per cycle.  So to get\n- * maximum performance out of it, you have to treat it as an in-order\n- * machine.  Which means interleaving the computation round t with the\n- * computation of W[t+4].\n- *\n- * The first 16 rounds use W values loaded directly from memory, while the\n- * remaining 64 use values computed from those first 16.  We preload\n- * 4 values before starting, so there are three kinds of rounds:\n- * - The first 12 (all f0) also load the W values from memory.\n- * - The next 64 compute W(i+4) in parallel. 8*f0, 20*f1, 20*f2, 16*f1.\n- * - The last 4 (all f1) do not do anything with W.\n- *\n- * Therefore, we have 6 different round functions:\n- * STEPD0_LOAD(t,s) - Perform round t and load W(s).  s < 16\n- * STEPD0_UPDATE(t,s) - Perform round t and compute W(s).  s >= 16.\n- * STEPD1_UPDATE(t,s)\n- * STEPD2_UPDATE(t,s)\n- * STEPD1(t) - Perform round t with no load or update.\n- *\n- * The G5 is more fully out-of-order, and can find the parallelism\n- * by itself.  The big limit is that it has a 2-cycle ALU latency, so\n- * even though it's 2-way, the code has to be scheduled as if it's\n- * 4-way, which can be a limit.  To help it, we try to schedule the\n- * read of RA(t) as late as possible so it doesn't stall waiting for\n- * the previous round's RE(t-1), and we try to rotate RB(t) as early\n- * as possible while reading RC(t) (= RB(t-1)) as late as possible.\n- */\n-\n-/* the initial loads. */\n-#define LOADW(s) \\\n-\tlwz\tW(s),(s)*4(%r4)\n-\n-/*\n- * Perform a step with F0, and load W(s).  Uses W(s) as a temporary\n- * before loading it.\n- * This is actually 10 instructions, which is an awkward fit.\n- * It can execute grouped as listed, or delayed one instruction.\n- * (If delayed two instructions, there is a stall before the start of the\n- * second line.)  Thus, two iterations take 7 cycles, 3.5 cycles per round.\n- */\n-#define STEPD0_LOAD(t,s) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t);  and    W(s),RC(t),RB(t); \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;      rotlwi RB(t),RB(t),30;   \\\n-add RE(t),RE(t),W(s); add    %r0,%r0,%r5;      lwz    W(s),(s)*4(%r4);  \\\n-add RE(t),RE(t),%r0\n-\n-/*\n- * This is likewise awkward, 13 instructions.  However, it can also\n- * execute starting with 2 out of 3 possible moduli, so it does 2 rounds\n- * in 9 cycles, 4.5 cycles/round.\n- */\n-#define STEPD0_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  and    %r0,RC(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r5;  loadk; rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1;     \\\n-add RE(t),RE(t),%r0\n-\n-/* Nicely optimal.  Conveniently, also the most common. */\n-#define STEPD1_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r5;  loadk; xor %r0,%r0,RC(t);  xor W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1\n-\n-/*\n- * The naked version, no UPDATE, for the last 4 rounds.  3 cycles per.\n- * We could use W(s) as a temp register, but we don't need it.\n- */\n-#define STEPD1(t) \\\n-                        add   RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); \\\n-rotlwi RB(t),RB(t),30;  add   RE(t),RE(t),%r5;  xor    %r0,%r0,RC(t);   \\\n-add    RE(t),RE(t),%r0; rotlwi %r0,RA(t),5;     /* spare slot */        \\\n-add    RE(t),RE(t),%r0\n-\n-/*\n- * 14 instructions, 5 cycles per.  The majority function is a bit\n- * awkward to compute.  This can execute with a 1-instruction delay,\n- * but it causes a 2-instruction delay, which triggers a stall.\n- */\n-#define STEPD2_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); and    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  xor    %r0,RD(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r5;  loadk; and %r0,%r0,RC(t);  xor W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     rotlwi W(s),W(s),1;             \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30\n-\n-#define STEP0_LOAD4(t,s)\t\t\\\n-\tSTEPD0_LOAD(t,s);\t\t\\\n-\tSTEPD0_LOAD((t+1),(s)+1);\t\\\n-\tSTEPD0_LOAD((t)+2,(s)+2);\t\\\n-\tSTEPD0_LOAD((t)+3,(s)+3)\n-\n-#define STEPUP4(fn, t, s, loadk...)\t\t\\\n-\tSTEP##fn##_UPDATE(t,s,);\t\t\\\n-\tSTEP##fn##_UPDATE((t)+1,(s)+1,);\t\\\n-\tSTEP##fn##_UPDATE((t)+2,(s)+2,);\t\\\n-\tSTEP##fn##_UPDATE((t)+3,(s)+3,loadk)\n-\n-#define STEPUP20(fn, t, s, loadk...)\t\\\n-\tSTEPUP4(fn, t, s,);\t\t\\\n-\tSTEPUP4(fn, (t)+4, (s)+4,);\t\\\n-\tSTEPUP4(fn, (t)+8, (s)+8,);\t\\\n-\tSTEPUP4(fn, (t)+12, (s)+12,);\t\\\n-\tSTEPUP4(fn, (t)+16, (s)+16, loadk)\n-\n-\t.globl\tppc_sha1_core\n-ppc_sha1_core:\n-\tstwu\t%r1,-80(%r1)\n-\tstmw\t%r13,4(%r1)\n-\n-\t/* Load up A - E */\n-\tlmw\t%r27,0(%r3)\n-\n-\tmtctr\t%r5\n-\n-1:\n-\tLOADW(0)\n-\tlis\t%r5,0x5a82\n-\tmr\tRE(0),%r31\n-\tLOADW(1)\n-\tmr\tRD(0),%r30\n-\tmr\tRC(0),%r29\n-\tLOADW(2)\n-\tori\t%r5,%r5,0x7999\t/* K0-19 */\n-\tmr\tRB(0),%r28\n-\tLOADW(3)\n-\tmr\tRA(0),%r27\n-\n-\tSTEP0_LOAD4(0, 4)\n-\tSTEP0_LOAD4(4, 8)\n-\tSTEP0_LOAD4(8, 12)\n-\tSTEPUP4(D0, 12, 16,)\n-\tSTEPUP4(D0, 16, 20, lis %r5,0x6ed9)\n-\n-\tori\t%r5,%r5,0xeba1\t/* K20-39 */\n-\tSTEPUP20(D1, 20, 24, lis %r5,0x8f1b)\n-\n-\tori\t%r5,%r5,0xbcdc\t/* K40-59 */\n-\tSTEPUP20(D2, 40, 44, lis %r5,0xca62)\n-\n-\tori\t%r5,%r5,0xc1d6\t/* K60-79 */\n-\tSTEPUP4(D1, 60, 64,)\n-\tSTEPUP4(D1, 64, 68,)\n-\tSTEPUP4(D1, 68, 72,)\n-\tSTEPUP4(D1, 72, 76,)\n-\taddi\t%r4,%r4,64\n-\tSTEPD1(76)\n-\tSTEPD1(77)\n-\tSTEPD1(78)\n-\tSTEPD1(79)\n-\n-\t/* Add results to original values */\n-\tadd\t%r31,%r31,RE(0)\n-\tadd\t%r30,%r30,RD(0)\n-\tadd\t%r29,%r29,RC(0)\n-\tadd\t%r28,%r28,RB(0)\n-\tadd\t%r27,%r27,RA(0)\n-\n-\tbdnz\t1b\n-\n-\t/* Save final hash, restore registers, and return */\n-\tstmw\t%r27,0(%r3)\n-\tlmw\t%r13,4(%r1)\n-\taddi\t%r1,%r1,80\n-\tblr\n-- \n2.24.0.424.g559c6fc317.dirty\n\n"},{"id":"392388","messageId":"20200224044732.GK1018190@coredump.intra.peff.net","threadId":"52786","inReplyTo":"20200223223758.120941-1-mh@glandium.org","subject":"Re: [PATCH] Remove non-SHA1dc sha1 implementations","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-02-24T04:47:32Z","receivedAt":"2020-02-24T04:47:34Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Feb 24, 2020 at 07:37:58AM +0900, Mike Hommey wrote:\n\n> It is 2020, and with the weakening of SHA1 security-wise, there doesn't\n> seem to be a reason to support anything else than SHA1dc, with collision\n> detection.\n\nOne possible reason is that they're way faster than sha1dc (block-sha1\nmaybe only a little, but openssl's sha1 is over twice as fast).\n\nTo be clear, I think the slowdown is worth the extra safety, but:\n\n - do we still want to care about people who prefer to make the tradeoff\n   differently?\n\n - when we first switched the default to sha1dc, the idea was raised of\n   continuing to use a faster implementation for non-security checksums\n   (e.g., the checksums at the end of packfiles, index files, etc). I\n   don't think anybody ever implemented that, but it's not a terrible\n   idea. OTOH, if nobody noticed the bottleneck enough to care, maybe\n   it's not worth worrying about.\n\nI'm not convinced the answer to those questions is \"yes\", but I think\nit's worth at least raising them (and arguing against them in the commit\nmessage).\n\nOne thing that compels me is the recent report that we still build with\ncommon crypto by default on macOS, which was definitely _not_ intended.\nThat's a bug that can be fixed, but it wouldn't have happened in the\nfirst place if we only supported sha1dc.\n\n-Peff\n"},{"id":"392389","messageId":"20200224045213.GA1574706@coredump.intra.peff.net","threadId":"52786","inReplyTo":"20200224044732.GK1018190@coredump.intra.peff.net","subject":"Re: [PATCH] Remove non-SHA1dc sha1 implementations","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2020-02-24T04:52:13Z","receivedAt":"2020-02-24T04:52:16Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 23, 2020 at 11:47:32PM -0500, Jeff King wrote:\n\n> One thing that compels me is the recent report that we still build with\n> common crypto by default on macOS, which was definitely _not_ intended.\n> That's a bug that can be fixed, but it wouldn't have happened in the\n> first place if we only supported sha1dc.\n\nI just noticed you were the original reporter there, too. So I guess it\ncompelled you, too. ;)\n\nIf we do want to keep the other implementations around, another thing\nthat might be worth doing is to teach t0013 to complain when the\ncollision-detecting sha1 is not in use (i.e., rather than auto-skipping\nwhen built without DC_SHA1, require the user to set a special\nNO_REALLY_I_CHOOSE_NOT_TO_USE_DC_SHA1_AND_AM_AWARE_OF_THE_IMPLICATIONS\nvariable). That would provide a cross-check on the build flags.\n\n-Peff\n"},{"id":"451658","messageId":"patch-1.1-05dcdca3877-20220319T005952Z-avarab@gmail.com","threadId":"52786","inReplyTo":"20200223223758.120941-1-mh@glandium.org","subject":"[PATCH] ppc: remove custom SHA-1 implementation","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-19T01:02:16Z","receivedAt":"2022-03-19T01:02:41Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Remove the PPC_SHA1 implementation added in a6ef3518f9a ([PATCH] PPC\nassembly implementation of SHA1, 2005-04-22). When this was added\nApple consumer hardware used the PPC architecture, and the\nimplementation was intended to improve SHA-1 speed there.\n\nSince it was added we've moved to DC_SHA1 by default, and anyone\nwanting hard-rolled non-DC SHA-1 implementation can use OpenSSL's via\nthe OPENSSL_SHA1 knob.\n\nI'm unsure if this was ever supposed to work on 64-bit PPC. It clearly\noriginally targeted 32 bit PPC, but there's some mailing list\nreferences to this being tried on G5 (PPC 970). I can't get it to do\nanything but segfault on the BE POWER8 machine in the GCC compile\nfarm. Anyone caring about speed on PPC these days is likely to be\nusing IBM's POWER, not PPC 970.\n\nThere have been proposals to entirely remove non-DC_SHA1\nimplementations from the tree[1]. I think per [2] that would be a bit\noverzealous. I.e. there are various set-ups git's speed is going to be\nmore important than the relatively implausible SHA-1 collision attack,\nor where such attacks are entirely mitigated by other means (e.g. by\nincoming objects being checked with DC_SHA1).\n\nThe main reason for doing so at this point is to simplify follow-up\nMakefile change. Since PPC_SHA1 included the only in-tree *.S assembly\nfile we needed to keep around special support for building objects\nfrom it. By getting rid of it we know we'll always build *.o from *.c\nfiles, which makes the build process simpler.\n\nAs an aside the code being removed here was also throwing warnings\nwith the \"-pedantic\" flag, but let's remove it instead of fixing it,\nas 544d93bc3b4 (block-sha1: remove use of obsolete x86 assembly,\n2022-03-10) did for block-sha1/*.\n\n1. https://lore.kernel.org/git/20200223223758.120941-1-mh@glandium.org/\n2. https://lore.kernel.org/git/20200224044732.GK1018190@coredump.intra.peff.net/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n INSTALL       |   3 +-\n Makefile      |  35 ++++----\n configure.ac  |   3 -\n hash.h        |   6 +-\n ppc/sha1.c    |  72 ----------------\n ppc/sha1.h    |  25 ------\n ppc/sha1ppc.S | 224 --------------------------------------------------\n 7 files changed, 17 insertions(+), 351 deletions(-)\n delete mode 100644 ppc/sha1.c\n delete mode 100644 ppc/sha1.h\n delete mode 100644 ppc/sha1ppc.S\n\ndiff --git a/INSTALL b/INSTALL\nindex 4140a3f5c8b..89b15d71df5 100644\n--- a/INSTALL\n+++ b/INSTALL\n@@ -135,8 +135,7 @@ Issues of note:\n \n \t  By default, git uses OpenSSL for SHA1 but it will use its own\n \t  library (inspired by Mozilla's) with either NO_OPENSSL or\n-\t  BLK_SHA1.  Also included is a version optimized for PowerPC\n-\t  (PPC_SHA1).\n+\t  BLK_SHA1.\n \n \t- \"libcurl\" library is used for fetching and pushing\n \t  repositories over http:// or https://, as well as by\ndiff --git a/Makefile b/Makefile\nindex 70f0a004e75..965e51f773e 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -155,9 +155,6 @@ include shared.mak\n # Define BLK_SHA1 environment variable to make use of the bundled\n # optimized C SHA1 routine.\n #\n-# Define PPC_SHA1 environment variable when running make to make use of\n-# a bundled SHA1 routine optimized for PowerPC.\n-#\n # Define DC_SHA1 to unconditionally enable the collision-detecting sha1\n # algorithm. This is slower, but may detect attempted collision attacks.\n # Takes priority over other *_SHA1 knobs.\n@@ -1770,14 +1767,14 @@ ifdef OPENSSL_SHA1\n \tEXTLIBS += $(LIB_4_CRYPTO)\n \tBASIC_CFLAGS += -DSHA1_OPENSSL\n else\n+ifdef PPC_SHA1\n+$(error PPC_SHA1 has been removed! You should almost definitely remove that \\\n+knob and use the DC_SHA1 default! See INSTALL for more information)\n+endif\n ifdef BLK_SHA1\n \tLIB_OBJS += block-sha1/sha1.o\n \tBASIC_CFLAGS += -DSHA1_BLK\n else\n-ifdef PPC_SHA1\n-\tLIB_OBJS += ppc/sha1.o ppc/sha1ppc.o\n-\tBASIC_CFLAGS += -DSHA1_PPC\n-else\n ifdef APPLE_COMMON_CRYPTO\n \tCOMPAT_CFLAGS += -DCOMMON_DIGEST_FOR_OPENSSL\n \tBASIC_CFLAGS += -DSHA1_APPLE\n@@ -1811,7 +1808,6 @@ endif\n endif\n endif\n endif\n-endif\n \n ifdef OPENSSL_SHA256\n \tEXTLIBS += $(LIB_4_CRYPTO)\n@@ -2509,6 +2505,11 @@ OBJECTS += $(SCALAR_OBJECTS)\n .PHONY: objects\n objects: $(OBJECTS)\n \n+# Derived from $(OBJECTS)\n+OBJECTS_C = $(OBJECTS:%.o=%.c)\n+OBJECTS_S = $(OBJECTS:%.o=%.s)\n+OBJECTS_SP = $(OBJECTS:%.o=%.sp)\n+\n dep_files := $(foreach f,$(OBJECTS),$(dir $f).depend/$(notdir $f).d)\n dep_dirs := $(addsuffix .depend,$(sort $(dir $(OBJECTS))))\n \n@@ -2540,13 +2541,7 @@ missing_compdb_dir =\n compdb_args =\n endif\n \n-ASM_SRC := $(wildcard $(OBJECTS:o=S))\n-ASM_OBJ := $(ASM_SRC:S=o)\n-C_OBJ := $(filter-out $(ASM_OBJ),$(OBJECTS))\n-\n-$(C_OBJ): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n-\t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n-$(ASM_OBJ): %.o: %.S GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n+$(OBJECTS): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n \t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n \n %.s: %.c GIT-CFLAGS FORCE\n@@ -2692,7 +2687,7 @@ XGETTEXT_FLAGS_SH = $(XGETTEXT_FLAGS) --language=Shell \\\n \t--keyword=gettextln --keyword=eval_gettextln\n XGETTEXT_FLAGS_PERL = $(XGETTEXT_FLAGS) --language=Perl \\\n \t--keyword=__ --keyword=N__ --keyword=\"__n:1,2\"\n-LOCALIZED_C = $(C_OBJ:o=c) $(LIB_H) $(GENERATED_H)\n+LOCALIZED_C = $(OBJECTS_C) $(LIB_H) $(GENERATED_H)\n LOCALIZED_SH = $(SCRIPT_SH)\n LOCALIZED_SH += git-sh-setup.sh\n LOCALIZED_PERL = $(SCRIPT_PERL)\n@@ -2939,16 +2934,14 @@ t/helper/test-%$X: t/helper/test-%.o GIT-LDFLAGS $(GITLIBS) $(REFTABLE_TEST_LIB)\n check-sha1:: t/helper/test-tool$X\n \tt/helper/test-sha1.sh\n \n-SP_OBJ = $(patsubst %.o,%.sp,$(C_OBJ))\n-\n-$(SP_OBJ): %.sp: %.c %.o\n+$(OBJECTS_SP): %.sp: %.c %.o\n \t$(QUIET_SP)cgcc -no-compile $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) \\\n \t\t-Wsparse-error \\\n \t\t$(SPARSE_FLAGS) $(SP_EXTRA_FLAGS) $< && \\\n \t>$@\n \n .PHONY: sparse\n-sparse: $(SP_OBJ)\n+sparse: $(OBJECTS_SP)\n \n EXCEPT_HDRS := $(GENERATED_H) unicode-width.h compat/% xdiff/%\n ifndef GCRYPT_SHA256\n@@ -3272,7 +3265,7 @@ clean: profile-clean coverage-clean cocciclean\n \t$(RM) $(ALL_PROGRAMS) $(SCRIPT_LIB) $(BUILT_INS) git$X\n \t$(RM) $(TEST_PROGRAMS)\n \t$(RM) $(FUZZ_PROGRAMS)\n-\t$(RM) $(SP_OBJ)\n+\t$(RM) $(OBJECTS_SP)\n \t$(RM) $(HCC)\n \t$(RM) -r bin-wrappers $(dep_dirs) $(compdb_dir) compile_commands.json\n \t$(RM) -r po/build/\ndiff --git a/configure.ac b/configure.ac\nindex 5ee25ec95c8..9c75b00d3eb 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -237,9 +237,6 @@ AC_MSG_NOTICE([CHECKS for site configuration])\n # tests.  These tests take up a significant amount of the total test time\n # but are not needed unless you plan to talk to SVN repos.\n #\n-# Define PPC_SHA1 environment variable when running make to make use of\n-# a bundled SHA1 routine optimized for PowerPC.\n-#\n # Define NO_OPENSSL environment variable if you do not have OpenSSL.\n #\n # Define OPENSSLDIR=/foo/bar if your openssl header and library files are in\ndiff --git a/hash.h b/hash.h\nindex 5d40368f18a..efc14c5f56d 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -4,9 +4,7 @@\n #include \"git-compat-util.h\"\n #include \"repository.h\"\n \n-#if defined(SHA1_PPC)\n-#include \"ppc/sha1.h\"\n-#elif defined(SHA1_APPLE)\n+#if defined(SHA1_APPLE)\n #include <CommonCrypto/CommonDigest.h>\n #elif defined(SHA1_OPENSSL)\n #include <openssl/sha.h>\n@@ -30,7 +28,7 @@\n  * platform's underlying implementation of SHA-1; could be OpenSSL,\n  * blk_SHA, Apple CommonCrypto, etc...  Note that the relevant\n  * SHA-1 header may have already defined platform_SHA_CTX for our\n- * own implementations like block-sha1 and ppc-sha1, so we list\n+ * own implementations like block-sha1, so we list\n  * the default for OpenSSL compatible SHA-1 implementations here.\n  */\n #define platform_SHA_CTX\tSHA_CTX\ndiff --git a/ppc/sha1.c b/ppc/sha1.c\ndeleted file mode 100644\nindex 1b705cee1fe..00000000000\n--- a/ppc/sha1.c\n+++ /dev/null\n@@ -1,72 +0,0 @@\n-/*\n- * SHA-1 implementation.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- *\n- * This version assumes we are running on a big-endian machine.\n- * It calls an external sha1_core() to process blocks of 64 bytes.\n- */\n-#include <stdio.h>\n-#include <string.h>\n-#include \"sha1.h\"\n-\n-void ppc_sha1_core(uint32_t *hash, const unsigned char *p,\n-\t\t   unsigned int nblocks);\n-\n-int ppc_SHA1_Init(ppc_SHA_CTX *c)\n-{\n-\tc->hash[0] = 0x67452301;\n-\tc->hash[1] = 0xEFCDAB89;\n-\tc->hash[2] = 0x98BADCFE;\n-\tc->hash[3] = 0x10325476;\n-\tc->hash[4] = 0xC3D2E1F0;\n-\tc->len = 0;\n-\tc->cnt = 0;\n-\treturn 0;\n-}\n-\n-int ppc_SHA1_Update(ppc_SHA_CTX *c, const void *ptr, unsigned long n)\n-{\n-\tunsigned long nb;\n-\tconst unsigned char *p = ptr;\n-\n-\tc->len += (uint64_t) n << 3;\n-\twhile (n != 0) {\n-\t\tif (c->cnt || n < 64) {\n-\t\t\tnb = 64 - c->cnt;\n-\t\t\tif (nb > n)\n-\t\t\t\tnb = n;\n-\t\t\tmemcpy(&c->buf.b[c->cnt], p, nb);\n-\t\t\tif ((c->cnt += nb) == 64) {\n-\t\t\t\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\t\t\t\tc->cnt = 0;\n-\t\t\t}\n-\t\t} else {\n-\t\t\tnb = n >> 6;\n-\t\t\tppc_sha1_core(c->hash, p, nb);\n-\t\t\tnb <<= 6;\n-\t\t}\n-\t\tn -= nb;\n-\t\tp += nb;\n-\t}\n-\treturn 0;\n-}\n-\n-int ppc_SHA1_Final(unsigned char *hash, ppc_SHA_CTX *c)\n-{\n-\tunsigned int cnt = c->cnt;\n-\n-\tc->buf.b[cnt++] = 0x80;\n-\tif (cnt > 56) {\n-\t\tif (cnt < 64)\n-\t\t\tmemset(&c->buf.b[cnt], 0, 64 - cnt);\n-\t\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\t\tcnt = 0;\n-\t}\n-\tif (cnt < 56)\n-\t\tmemset(&c->buf.b[cnt], 0, 56 - cnt);\n-\tc->buf.l[7] = c->len;\n-\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\tmemcpy(hash, c->hash, 20);\n-\treturn 0;\n-}\ndiff --git a/ppc/sha1.h b/ppc/sha1.h\ndeleted file mode 100644\nindex 9b24b326159..00000000000\n--- a/ppc/sha1.h\n+++ /dev/null\n@@ -1,25 +0,0 @@\n-/*\n- * SHA-1 implementation.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- */\n-#include <stdint.h>\n-\n-typedef struct {\n-\tuint32_t hash[5];\n-\tuint32_t cnt;\n-\tuint64_t len;\n-\tunion {\n-\t\tunsigned char b[64];\n-\t\tuint64_t l[8];\n-\t} buf;\n-} ppc_SHA_CTX;\n-\n-int ppc_SHA1_Init(ppc_SHA_CTX *c);\n-int ppc_SHA1_Update(ppc_SHA_CTX *c, const void *p, unsigned long n);\n-int ppc_SHA1_Final(unsigned char *hash, ppc_SHA_CTX *c);\n-\n-#define platform_SHA_CTX\tppc_SHA_CTX\n-#define platform_SHA1_Init\tppc_SHA1_Init\n-#define platform_SHA1_Update\tppc_SHA1_Update\n-#define platform_SHA1_Final\tppc_SHA1_Final\ndiff --git a/ppc/sha1ppc.S b/ppc/sha1ppc.S\ndeleted file mode 100644\nindex 1711eef6e71..00000000000\n--- a/ppc/sha1ppc.S\n+++ /dev/null\n@@ -1,224 +0,0 @@\n-/*\n- * SHA-1 implementation for PowerPC.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- */\n-\n-/*\n- * PowerPC calling convention:\n- * %r0 - volatile temp\n- * %r1 - stack pointer.\n- * %r2 - reserved\n- * %r3-%r12 - Incoming arguments & return values; volatile.\n- * %r13-%r31 - Callee-save registers\n- * %lr - Return address, volatile\n- * %ctr - volatile\n- *\n- * Register usage in this routine:\n- * %r0 - temp\n- * %r3 - argument (pointer to 5 words of SHA state)\n- * %r4 - argument (pointer to data to hash)\n- * %r5 - Constant K in SHA round (initially number of blocks to hash)\n- * %r6-%r10 - Working copies of SHA variables A..E (actually E..A order)\n- * %r11-%r26 - Data being hashed W[].\n- * %r27-%r31 - Previous copies of A..E, for final add back.\n- * %ctr - loop count\n- */\n-\n-\n-/*\n- * We roll the registers for A, B, C, D, E around on each\n- * iteration; E on iteration t is D on iteration t+1, and so on.\n- * We use registers 6 - 10 for this.  (Registers 27 - 31 hold\n- * the previous values.)\n- */\n-#define RA(t)\t(((t)+4)%5+6)\n-#define RB(t)\t(((t)+3)%5+6)\n-#define RC(t)\t(((t)+2)%5+6)\n-#define RD(t)\t(((t)+1)%5+6)\n-#define RE(t)\t(((t)+0)%5+6)\n-\n-/* We use registers 11 - 26 for the W values */\n-#define W(t)\t((t)%16+11)\n-\n-/* Register 5 is used for the constant k */\n-\n-/*\n- * The basic SHA-1 round function is:\n- * E += ROTL(A,5) + F(B,C,D) + W[i] + K;  B = ROTL(B,30)\n- * Then the variables are renamed: (A,B,C,D,E) = (E,A,B,C,D).\n- *\n- * Every 20 rounds, the function F() and the constant K changes:\n- * - 20 rounds of f0(b,c,d) = \"bit wise b ? c : d\" =  (^b & d) + (b & c)\n- * - 20 rounds of f1(b,c,d) = b^c^d = (b^d)^c\n- * - 20 rounds of f2(b,c,d) = majority(b,c,d) = (b&d) + ((b^d)&c)\n- * - 20 more rounds of f1(b,c,d)\n- *\n- * These are all scheduled for near-optimal performance on a G4.\n- * The G4 is a 3-issue out-of-order machine with 3 ALUs, but it can only\n- * *consider* starting the oldest 3 instructions per cycle.  So to get\n- * maximum performance out of it, you have to treat it as an in-order\n- * machine.  Which means interleaving the computation round t with the\n- * computation of W[t+4].\n- *\n- * The first 16 rounds use W values loaded directly from memory, while the\n- * remaining 64 use values computed from those first 16.  We preload\n- * 4 values before starting, so there are three kinds of rounds:\n- * - The first 12 (all f0) also load the W values from memory.\n- * - The next 64 compute W(i+4) in parallel. 8*f0, 20*f1, 20*f2, 16*f1.\n- * - The last 4 (all f1) do not do anything with W.\n- *\n- * Therefore, we have 6 different round functions:\n- * STEPD0_LOAD(t,s) - Perform round t and load W(s).  s < 16\n- * STEPD0_UPDATE(t,s) - Perform round t and compute W(s).  s >= 16.\n- * STEPD1_UPDATE(t,s)\n- * STEPD2_UPDATE(t,s)\n- * STEPD1(t) - Perform round t with no load or update.\n- *\n- * The G5 is more fully out-of-order, and can find the parallelism\n- * by itself.  The big limit is that it has a 2-cycle ALU latency, so\n- * even though it's 2-way, the code has to be scheduled as if it's\n- * 4-way, which can be a limit.  To help it, we try to schedule the\n- * read of RA(t) as late as possible so it doesn't stall waiting for\n- * the previous round's RE(t-1), and we try to rotate RB(t) as early\n- * as possible while reading RC(t) (= RB(t-1)) as late as possible.\n- */\n-\n-/* the initial loads. */\n-#define LOADW(s) \\\n-\tlwz\tW(s),(s)*4(%r4)\n-\n-/*\n- * Perform a step with F0, and load W(s).  Uses W(s) as a temporary\n- * before loading it.\n- * This is actually 10 instructions, which is an awkward fit.\n- * It can execute grouped as listed, or delayed one instruction.\n- * (If delayed two instructions, there is a stall before the start of the\n- * second line.)  Thus, two iterations take 7 cycles, 3.5 cycles per round.\n- */\n-#define STEPD0_LOAD(t,s) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t);  and    W(s),RC(t),RB(t); \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;      rotlwi RB(t),RB(t),30;   \\\n-add RE(t),RE(t),W(s); add    %r0,%r0,%r5;      lwz    W(s),(s)*4(%r4);  \\\n-add RE(t),RE(t),%r0\n-\n-/*\n- * This is likewise awkward, 13 instructions.  However, it can also\n- * execute starting with 2 out of 3 possible moduli, so it does 2 rounds\n- * in 9 cycles, 4.5 cycles/round.\n- */\n-#define STEPD0_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  and    %r0,RC(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r5;  loadk; rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1;     \\\n-add RE(t),RE(t),%r0\n-\n-/* Nicely optimal.  Conveniently, also the most common. */\n-#define STEPD1_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r5;  loadk; xor %r0,%r0,RC(t);  xor W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1\n-\n-/*\n- * The naked version, no UPDATE, for the last 4 rounds.  3 cycles per.\n- * We could use W(s) as a temp register, but we don't need it.\n- */\n-#define STEPD1(t) \\\n-                        add   RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); \\\n-rotlwi RB(t),RB(t),30;  add   RE(t),RE(t),%r5;  xor    %r0,%r0,RC(t);   \\\n-add    RE(t),RE(t),%r0; rotlwi %r0,RA(t),5;     /* spare slot */        \\\n-add    RE(t),RE(t),%r0\n-\n-/*\n- * 14 instructions, 5 cycles per.  The majority function is a bit\n- * awkward to compute.  This can execute with a 1-instruction delay,\n- * but it causes a 2-instruction delay, which triggers a stall.\n- */\n-#define STEPD2_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); and    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  xor    %r0,RD(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r5;  loadk; and %r0,%r0,RC(t);  xor W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     rotlwi W(s),W(s),1;             \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30\n-\n-#define STEP0_LOAD4(t,s)\t\t\\\n-\tSTEPD0_LOAD(t,s);\t\t\\\n-\tSTEPD0_LOAD((t+1),(s)+1);\t\\\n-\tSTEPD0_LOAD((t)+2,(s)+2);\t\\\n-\tSTEPD0_LOAD((t)+3,(s)+3)\n-\n-#define STEPUP4(fn, t, s, loadk...)\t\t\\\n-\tSTEP##fn##_UPDATE(t,s,);\t\t\\\n-\tSTEP##fn##_UPDATE((t)+1,(s)+1,);\t\\\n-\tSTEP##fn##_UPDATE((t)+2,(s)+2,);\t\\\n-\tSTEP##fn##_UPDATE((t)+3,(s)+3,loadk)\n-\n-#define STEPUP20(fn, t, s, loadk...)\t\\\n-\tSTEPUP4(fn, t, s,);\t\t\\\n-\tSTEPUP4(fn, (t)+4, (s)+4,);\t\\\n-\tSTEPUP4(fn, (t)+8, (s)+8,);\t\\\n-\tSTEPUP4(fn, (t)+12, (s)+12,);\t\\\n-\tSTEPUP4(fn, (t)+16, (s)+16, loadk)\n-\n-\t.globl\tppc_sha1_core\n-ppc_sha1_core:\n-\tstwu\t%r1,-80(%r1)\n-\tstmw\t%r13,4(%r1)\n-\n-\t/* Load up A - E */\n-\tlmw\t%r27,0(%r3)\n-\n-\tmtctr\t%r5\n-\n-1:\n-\tLOADW(0)\n-\tlis\t%r5,0x5a82\n-\tmr\tRE(0),%r31\n-\tLOADW(1)\n-\tmr\tRD(0),%r30\n-\tmr\tRC(0),%r29\n-\tLOADW(2)\n-\tori\t%r5,%r5,0x7999\t/* K0-19 */\n-\tmr\tRB(0),%r28\n-\tLOADW(3)\n-\tmr\tRA(0),%r27\n-\n-\tSTEP0_LOAD4(0, 4)\n-\tSTEP0_LOAD4(4, 8)\n-\tSTEP0_LOAD4(8, 12)\n-\tSTEPUP4(D0, 12, 16,)\n-\tSTEPUP4(D0, 16, 20, lis %r5,0x6ed9)\n-\n-\tori\t%r5,%r5,0xeba1\t/* K20-39 */\n-\tSTEPUP20(D1, 20, 24, lis %r5,0x8f1b)\n-\n-\tori\t%r5,%r5,0xbcdc\t/* K40-59 */\n-\tSTEPUP20(D2, 40, 44, lis %r5,0xca62)\n-\n-\tori\t%r5,%r5,0xc1d6\t/* K60-79 */\n-\tSTEPUP4(D1, 60, 64,)\n-\tSTEPUP4(D1, 64, 68,)\n-\tSTEPUP4(D1, 68, 72,)\n-\tSTEPUP4(D1, 72, 76,)\n-\taddi\t%r4,%r4,64\n-\tSTEPD1(76)\n-\tSTEPD1(77)\n-\tSTEPD1(78)\n-\tSTEPD1(79)\n-\n-\t/* Add results to original values */\n-\tadd\t%r31,%r31,RE(0)\n-\tadd\t%r30,%r30,RD(0)\n-\tadd\t%r29,%r29,RC(0)\n-\tadd\t%r28,%r28,RB(0)\n-\tadd\t%r27,%r27,RA(0)\n-\n-\tbdnz\t1b\n-\n-\t/* Save final hash, restore registers, and return */\n-\tstmw\t%r27,0(%r3)\n-\tlmw\t%r13,4(%r1)\n-\taddi\t%r1,%r1,80\n-\tblr\n-- \n2.35.1.1438.g8874c8eeb35\n\n"},{"id":"451733","messageId":"xmqq5yo72rbm.fsf@gitster.g","threadId":"52786","inReplyTo":"patch-1.1-05dcdca3877-20220319T005952Z-avarab@gmail.com","subject":"Re: [PATCH] ppc: remove custom SHA-1 implementation","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-21T16:39:09Z","receivedAt":"2022-03-21T16:39:19Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> The main reason for doing so at this point is to simplify follow-up\n> Makefile change. Since PPC_SHA1 included the only in-tree *.S assembly\n> file we needed to keep around special support for building objects\n> from it. By getting rid of it we know we'll always build *.o from *.c\n> files, which makes the build process simpler.\n\nYuck.\n\n> diff --git a/Makefile b/Makefile\n> index 70f0a004e75..965e51f773e 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -155,9 +155,6 @@ include shared.mak\n>  # Define BLK_SHA1 environment variable to make use of the bundled\n>  # optimized C SHA1 routine.\n>  #\n> -# Define PPC_SHA1 environment variable when running make to make use of\n> -# a bundled SHA1 routine optimized for PowerPC.\n> -#\n>  # Define DC_SHA1 to unconditionally enable the collision-detecting sha1\n>  # algorithm. This is slower, but may detect attempted collision attacks.\n>  # Takes priority over other *_SHA1 knobs.\n> @@ -1770,14 +1767,14 @@ ifdef OPENSSL_SHA1\n>  \tEXTLIBS += $(LIB_4_CRYPTO)\n>  \tBASIC_CFLAGS += -DSHA1_OPENSSL\n>  else\n> +ifdef PPC_SHA1\n> +$(error PPC_SHA1 has been removed! You should almost definitely remove that \\\n> +knob and use the DC_SHA1 default! See INSTALL for more information)\n> +endif\n\n\"use the DC_SHA1 default\"?  \n\n\tPPC_SHA1 support is being removed.  Use DC_SHA1 instead, \n\twhich is the default.\n\n\nI am wondering if we can make only these four lines the first step\nof remova, without doing anything else.  It would give us a good\nfeel on how many users we may be inconveniencing (not necessarily\nhurting, as switching to DC_SHA1 would be a good move) by the\nremoval.\n\n> @@ -2509,6 +2505,11 @@ OBJECTS += $(SCALAR_OBJECTS)\n>  .PHONY: objects\n>  objects: $(OBJECTS)\n>  \n> +# Derived from $(OBJECTS)\n> +OBJECTS_C = $(OBJECTS:%.o=%.c)\n> +OBJECTS_S = $(OBJECTS:%.o=%.s)\n> +OBJECTS_SP = $(OBJECTS:%.o=%.sp)\n\nUsually we build objects from sources, so \"derived from\" is a\npuzzling way to call them.  I understand we are deriving a list from\nanother list, but it still feels confusing (see below for ASM_OBJ).\n\nThis seems to have nothing to do with \"we no longer have *.S files\"\nand looks more like \"now we have excuse to touch Makefile, make\nrandom changes that look subjectively good to me, burying them in\nthe noise so that we can sneak them in without justifying them much\nin the proposed log message\".\n\n>  dep_files := $(foreach f,$(OBJECTS),$(dir $f).depend/$(notdir $f).d)\n>  dep_dirs := $(addsuffix .depend,$(sort $(dir $(OBJECTS))))\n>  \n> @@ -2540,13 +2541,7 @@ missing_compdb_dir =\n>  compdb_args =\n>  endif\n>  \n> -ASM_SRC := $(wildcard $(OBJECTS:o=S))\n> -ASM_OBJ := $(ASM_SRC:S=o)\n> -C_OBJ := $(filter-out $(ASM_OBJ),$(OBJECTS))\n\nI tend to agree with this patch that these three lines are ugly in\nmultiple ways.\n\nIt's a confusing construct.  The list of OBJECTS is used as the\nsingle source of truth and others are derived by filtering the list\nand futzing the suffix of the resulting subset of elements; it makes\nme wonder if it should be the other way around (i.e. we have a list\nof source files in various languages, all get turned into objects,\nrather than we have a list of object files, and we see if\ncorresponding source file in possible languages, if any, exists on\ndisk).  It is pleasing that we can see them go.\n\nUnfortunately the new \"Derived from $(OBJECTS) lists are in the same\nspirit as used by ASM_SRC we are removing here, aren't they?\n\n> -$(C_OBJ): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n> -\t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n> -$(ASM_OBJ): %.o: %.S GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n> +$(OBJECTS): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n>  \t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n\nNow we deal only with *.c sources and all objects come from them,\nand the action is the same as the old C_OBJ rule, naturally.\n\n>  %.s: %.c GIT-CFLAGS FORCE\n\nThis is the only remaining rule regarding assembly and it is the\ntarget to generate for debugging, not even used as a source to\ncreate object files.  Do we need OBJECTS_S defined above?  I somehow\ndoubt it.  FWIW, I do not see the need for OBJECTS_SP or OBJECTS_C,\neither.  E.g.\n\n> @@ -2692,7 +2687,7 @@ XGETTEXT_FLAGS_SH = $(XGETTEXT_FLAGS) --language=Shell \\\n>  \t--keyword=gettextln --keyword=eval_gettextln\n>  XGETTEXT_FLAGS_PERL = $(XGETTEXT_FLAGS) --language=Perl \\\n>  \t--keyword=__ --keyword=N__ --keyword=\"__n:1,2\"\n> -LOCALIZED_C = $(C_OBJ:o=c) $(LIB_H) $(GENERATED_H)\n> +LOCALIZED_C = $(OBJECTS_C) $(LIB_H) $(GENERATED_H)\n\nShouldn't it be sufficient to use %(OBJECTS:o=c) here, i.e. equating\nthe old C_OBJ with OBJECTS, now we know all objects come from C?\nThe same question for SP_OBJ below, but I won't repeat.\n"},{"id":"451737","messageId":"patch-v2-1.1-e77fd23a824-20220321T170412Z-avarab@gmail.com","threadId":"52786","inReplyTo":"patch-1.1-05dcdca3877-20220319T005952Z-avarab@gmail.com","subject":"[PATCH v2] ppc: remove custom SHA-1 implementation","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-21T17:06:12Z","receivedAt":"2022-03-21T17:06:36Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Remove the PPC_SHA1 implementation added in a6ef3518f9a ([PATCH] PPC\nassembly implementation of SHA1, 2005-04-22). When this was added\nApple consumer hardware used the PPC architecture, and the\nimplementation was intended to improve SHA-1 speed there.\n\nSince it was added we've moved to DC_SHA1 by default, and anyone\nwanting hard-rolled non-DC SHA-1 implementation can use OpenSSL's via\nthe OPENSSL_SHA1 knob.\n\nI'm unsure if this was ever supposed to work on 64-bit PPC. It clearly\noriginally targeted 32 bit PPC, but there's some mailing list\nreferences to this being tried on G5 (PPC 970). I can't get it to do\nanything but segfault on the BE POWER8 machine in the GCC compile\nfarm. Anyone caring about speed on PPC these days is likely to be\nusing IBM's POWER, not PPC 970.\n\nThere have been proposals to entirely remove non-DC_SHA1\nimplementations from the tree[1]. I think per [2] that would be a bit\noverzealous. I.e. there are various set-ups git's speed is going to be\nmore important than the relatively implausible SHA-1 collision attack,\nor where such attacks are entirely mitigated by other means (e.g. by\nincoming objects being checked with DC_SHA1).\n\nThe main reason for doing so at this point is to simplify follow-up\nMakefile change. Since PPC_SHA1 included the only in-tree *.S assembly\nfile we needed to keep around special support for building objects\nfrom it. By getting rid of it we know we'll always build *.o from *.c\nfiles, which makes the build process simpler.\n\nAs an aside the code being removed here was also throwing warnings\nwith the \"-pedantic\" flag, but let's remove it instead of fixing it,\nas 544d93bc3b4 (block-sha1: remove use of obsolete x86 assembly,\n2022-03-10) did for block-sha1/*.\n\n1. https://lore.kernel.org/git/20200223223758.120941-1-mh@glandium.org/\n2. https://lore.kernel.org/git/20200224044732.GK1018190@coredump.intra.peff.net/\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n\nA more minimal change to the Makefile than that in v1. In v1 I thought\nit was a bit odd to leave in-place the C_OBJ variable that was only\ndeclared for now-removed code. But sure, we can keep it and keep\ndefining SP_OBJ etc. as being derived from it.\n\nRange-diff against v1:\n1:  05dcdca3877 ! 1:  e77fd23a824 ppc: remove custom SHA-1 implementation\n    @@ Makefile: ifdef OPENSSL_SHA1\n      \tBASIC_CFLAGS += -DSHA1_OPENSSL\n      else\n     +ifdef PPC_SHA1\n    -+$(error PPC_SHA1 has been removed! You should almost definitely remove that \\\n    -+knob and use the DC_SHA1 default! See INSTALL for more information)\n    ++$(error PPC_SHA1 has been removed! Use DC_SHA1 instead, which is the default)\n     +endif\n      ifdef BLK_SHA1\n      \tLIB_OBJS += block-sha1/sha1.o\n    @@ Makefile: endif\n      \n      ifdef OPENSSL_SHA256\n      \tEXTLIBS += $(LIB_4_CRYPTO)\n    -@@ Makefile: OBJECTS += $(SCALAR_OBJECTS)\n    - .PHONY: objects\n    - objects: $(OBJECTS)\n    - \n    -+# Derived from $(OBJECTS)\n    -+OBJECTS_C = $(OBJECTS:%.o=%.c)\n    -+OBJECTS_S = $(OBJECTS:%.o=%.s)\n    -+OBJECTS_SP = $(OBJECTS:%.o=%.sp)\n    -+\n    - dep_files := $(foreach f,$(OBJECTS),$(dir $f).depend/$(notdir $f).d)\n    - dep_dirs := $(addsuffix .depend,$(sort $(dir $(OBJECTS))))\n    - \n     @@ Makefile: missing_compdb_dir =\n      compdb_args =\n      endif\n    @@ Makefile: missing_compdb_dir =\n     -ASM_SRC := $(wildcard $(OBJECTS:o=S))\n     -ASM_OBJ := $(ASM_SRC:S=o)\n     -C_OBJ := $(filter-out $(ASM_OBJ),$(OBJECTS))\n    --\n    --$(C_OBJ): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n    --\t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n    --$(ASM_OBJ): %.o: %.S GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n    -+$(OBJECTS): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n    ++C_OBJ = $(OBJECTS)\n    + \n    + $(C_OBJ): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n      \t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n    +-$(ASM_OBJ): %.o: %.S GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n    +-\t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n      \n      %.s: %.c GIT-CFLAGS FORCE\n    -@@ Makefile: XGETTEXT_FLAGS_SH = $(XGETTEXT_FLAGS) --language=Shell \\\n    - \t--keyword=gettextln --keyword=eval_gettextln\n    - XGETTEXT_FLAGS_PERL = $(XGETTEXT_FLAGS) --language=Perl \\\n    - \t--keyword=__ --keyword=N__ --keyword=\"__n:1,2\"\n    --LOCALIZED_C = $(C_OBJ:o=c) $(LIB_H) $(GENERATED_H)\n    -+LOCALIZED_C = $(OBJECTS_C) $(LIB_H) $(GENERATED_H)\n    - LOCALIZED_SH = $(SCRIPT_SH)\n    - LOCALIZED_SH += git-sh-setup.sh\n    - LOCALIZED_PERL = $(SCRIPT_PERL)\n    -@@ Makefile: t/helper/test-%$X: t/helper/test-%.o GIT-LDFLAGS $(GITLIBS) $(REFTABLE_TEST_LIB)\n    - check-sha1:: t/helper/test-tool$X\n    - \tt/helper/test-sha1.sh\n    - \n    --SP_OBJ = $(patsubst %.o,%.sp,$(C_OBJ))\n    --\n    --$(SP_OBJ): %.sp: %.c %.o\n    -+$(OBJECTS_SP): %.sp: %.c %.o\n    - \t$(QUIET_SP)cgcc -no-compile $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) \\\n    - \t\t-Wsparse-error \\\n    - \t\t$(SPARSE_FLAGS) $(SP_EXTRA_FLAGS) $< && \\\n    - \t>$@\n    - \n    - .PHONY: sparse\n    --sparse: $(SP_OBJ)\n    -+sparse: $(OBJECTS_SP)\n    - \n    - EXCEPT_HDRS := $(GENERATED_H) unicode-width.h compat/% xdiff/%\n    - ifndef GCRYPT_SHA256\n    -@@ Makefile: clean: profile-clean coverage-clean cocciclean\n    - \t$(RM) $(ALL_PROGRAMS) $(SCRIPT_LIB) $(BUILT_INS) git$X\n    - \t$(RM) $(TEST_PROGRAMS)\n    - \t$(RM) $(FUZZ_PROGRAMS)\n    --\t$(RM) $(SP_OBJ)\n    -+\t$(RM) $(OBJECTS_SP)\n    - \t$(RM) $(HCC)\n    - \t$(RM) -r bin-wrappers $(dep_dirs) $(compdb_dir) compile_commands.json\n    - \t$(RM) -r po/build/\n    + \t$(QUIET_CC)$(CC) -o $@ -S $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n     \n      ## configure.ac ##\n     @@ configure.ac: AC_MSG_NOTICE([CHECKS for site configuration])\n\n INSTALL       |   3 +-\n Makefile      |  17 +---\n configure.ac  |   3 -\n hash.h        |   6 +-\n ppc/sha1.c    |  72 ----------------\n ppc/sha1.h    |  25 ------\n ppc/sha1ppc.S | 224 --------------------------------------------------\n 7 files changed, 7 insertions(+), 343 deletions(-)\n delete mode 100644 ppc/sha1.c\n delete mode 100644 ppc/sha1.h\n delete mode 100644 ppc/sha1ppc.S\n\ndiff --git a/INSTALL b/INSTALL\nindex 4140a3f5c8b..89b15d71df5 100644\n--- a/INSTALL\n+++ b/INSTALL\n@@ -135,8 +135,7 @@ Issues of note:\n \n \t  By default, git uses OpenSSL for SHA1 but it will use its own\n \t  library (inspired by Mozilla's) with either NO_OPENSSL or\n-\t  BLK_SHA1.  Also included is a version optimized for PowerPC\n-\t  (PPC_SHA1).\n+\t  BLK_SHA1.\n \n \t- \"libcurl\" library is used for fetching and pushing\n \t  repositories over http:// or https://, as well as by\ndiff --git a/Makefile b/Makefile\nindex 70f0a004e75..33c6db5e6c9 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -155,9 +155,6 @@ include shared.mak\n # Define BLK_SHA1 environment variable to make use of the bundled\n # optimized C SHA1 routine.\n #\n-# Define PPC_SHA1 environment variable when running make to make use of\n-# a bundled SHA1 routine optimized for PowerPC.\n-#\n # Define DC_SHA1 to unconditionally enable the collision-detecting sha1\n # algorithm. This is slower, but may detect attempted collision attacks.\n # Takes priority over other *_SHA1 knobs.\n@@ -1770,14 +1767,13 @@ ifdef OPENSSL_SHA1\n \tEXTLIBS += $(LIB_4_CRYPTO)\n \tBASIC_CFLAGS += -DSHA1_OPENSSL\n else\n+ifdef PPC_SHA1\n+$(error PPC_SHA1 has been removed! Use DC_SHA1 instead, which is the default)\n+endif\n ifdef BLK_SHA1\n \tLIB_OBJS += block-sha1/sha1.o\n \tBASIC_CFLAGS += -DSHA1_BLK\n else\n-ifdef PPC_SHA1\n-\tLIB_OBJS += ppc/sha1.o ppc/sha1ppc.o\n-\tBASIC_CFLAGS += -DSHA1_PPC\n-else\n ifdef APPLE_COMMON_CRYPTO\n \tCOMPAT_CFLAGS += -DCOMMON_DIGEST_FOR_OPENSSL\n \tBASIC_CFLAGS += -DSHA1_APPLE\n@@ -1811,7 +1807,6 @@ endif\n endif\n endif\n endif\n-endif\n \n ifdef OPENSSL_SHA256\n \tEXTLIBS += $(LIB_4_CRYPTO)\n@@ -2540,14 +2535,10 @@ missing_compdb_dir =\n compdb_args =\n endif\n \n-ASM_SRC := $(wildcard $(OBJECTS:o=S))\n-ASM_OBJ := $(ASM_SRC:S=o)\n-C_OBJ := $(filter-out $(ASM_OBJ),$(OBJECTS))\n+C_OBJ = $(OBJECTS)\n \n $(C_OBJ): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n \t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n-$(ASM_OBJ): %.o: %.S GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n-\t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n \n %.s: %.c GIT-CFLAGS FORCE\n \t$(QUIET_CC)$(CC) -o $@ -S $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\ndiff --git a/configure.ac b/configure.ac\nindex 5ee25ec95c8..9c75b00d3eb 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -237,9 +237,6 @@ AC_MSG_NOTICE([CHECKS for site configuration])\n # tests.  These tests take up a significant amount of the total test time\n # but are not needed unless you plan to talk to SVN repos.\n #\n-# Define PPC_SHA1 environment variable when running make to make use of\n-# a bundled SHA1 routine optimized for PowerPC.\n-#\n # Define NO_OPENSSL environment variable if you do not have OpenSSL.\n #\n # Define OPENSSLDIR=/foo/bar if your openssl header and library files are in\ndiff --git a/hash.h b/hash.h\nindex 5d40368f18a..efc14c5f56d 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -4,9 +4,7 @@\n #include \"git-compat-util.h\"\n #include \"repository.h\"\n \n-#if defined(SHA1_PPC)\n-#include \"ppc/sha1.h\"\n-#elif defined(SHA1_APPLE)\n+#if defined(SHA1_APPLE)\n #include <CommonCrypto/CommonDigest.h>\n #elif defined(SHA1_OPENSSL)\n #include <openssl/sha.h>\n@@ -30,7 +28,7 @@\n  * platform's underlying implementation of SHA-1; could be OpenSSL,\n  * blk_SHA, Apple CommonCrypto, etc...  Note that the relevant\n  * SHA-1 header may have already defined platform_SHA_CTX for our\n- * own implementations like block-sha1 and ppc-sha1, so we list\n+ * own implementations like block-sha1, so we list\n  * the default for OpenSSL compatible SHA-1 implementations here.\n  */\n #define platform_SHA_CTX\tSHA_CTX\ndiff --git a/ppc/sha1.c b/ppc/sha1.c\ndeleted file mode 100644\nindex 1b705cee1fe..00000000000\n--- a/ppc/sha1.c\n+++ /dev/null\n@@ -1,72 +0,0 @@\n-/*\n- * SHA-1 implementation.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- *\n- * This version assumes we are running on a big-endian machine.\n- * It calls an external sha1_core() to process blocks of 64 bytes.\n- */\n-#include <stdio.h>\n-#include <string.h>\n-#include \"sha1.h\"\n-\n-void ppc_sha1_core(uint32_t *hash, const unsigned char *p,\n-\t\t   unsigned int nblocks);\n-\n-int ppc_SHA1_Init(ppc_SHA_CTX *c)\n-{\n-\tc->hash[0] = 0x67452301;\n-\tc->hash[1] = 0xEFCDAB89;\n-\tc->hash[2] = 0x98BADCFE;\n-\tc->hash[3] = 0x10325476;\n-\tc->hash[4] = 0xC3D2E1F0;\n-\tc->len = 0;\n-\tc->cnt = 0;\n-\treturn 0;\n-}\n-\n-int ppc_SHA1_Update(ppc_SHA_CTX *c, const void *ptr, unsigned long n)\n-{\n-\tunsigned long nb;\n-\tconst unsigned char *p = ptr;\n-\n-\tc->len += (uint64_t) n << 3;\n-\twhile (n != 0) {\n-\t\tif (c->cnt || n < 64) {\n-\t\t\tnb = 64 - c->cnt;\n-\t\t\tif (nb > n)\n-\t\t\t\tnb = n;\n-\t\t\tmemcpy(&c->buf.b[c->cnt], p, nb);\n-\t\t\tif ((c->cnt += nb) == 64) {\n-\t\t\t\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\t\t\t\tc->cnt = 0;\n-\t\t\t}\n-\t\t} else {\n-\t\t\tnb = n >> 6;\n-\t\t\tppc_sha1_core(c->hash, p, nb);\n-\t\t\tnb <<= 6;\n-\t\t}\n-\t\tn -= nb;\n-\t\tp += nb;\n-\t}\n-\treturn 0;\n-}\n-\n-int ppc_SHA1_Final(unsigned char *hash, ppc_SHA_CTX *c)\n-{\n-\tunsigned int cnt = c->cnt;\n-\n-\tc->buf.b[cnt++] = 0x80;\n-\tif (cnt > 56) {\n-\t\tif (cnt < 64)\n-\t\t\tmemset(&c->buf.b[cnt], 0, 64 - cnt);\n-\t\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\t\tcnt = 0;\n-\t}\n-\tif (cnt < 56)\n-\t\tmemset(&c->buf.b[cnt], 0, 56 - cnt);\n-\tc->buf.l[7] = c->len;\n-\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\tmemcpy(hash, c->hash, 20);\n-\treturn 0;\n-}\ndiff --git a/ppc/sha1.h b/ppc/sha1.h\ndeleted file mode 100644\nindex 9b24b326159..00000000000\n--- a/ppc/sha1.h\n+++ /dev/null\n@@ -1,25 +0,0 @@\n-/*\n- * SHA-1 implementation.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- */\n-#include <stdint.h>\n-\n-typedef struct {\n-\tuint32_t hash[5];\n-\tuint32_t cnt;\n-\tuint64_t len;\n-\tunion {\n-\t\tunsigned char b[64];\n-\t\tuint64_t l[8];\n-\t} buf;\n-} ppc_SHA_CTX;\n-\n-int ppc_SHA1_Init(ppc_SHA_CTX *c);\n-int ppc_SHA1_Update(ppc_SHA_CTX *c, const void *p, unsigned long n);\n-int ppc_SHA1_Final(unsigned char *hash, ppc_SHA_CTX *c);\n-\n-#define platform_SHA_CTX\tppc_SHA_CTX\n-#define platform_SHA1_Init\tppc_SHA1_Init\n-#define platform_SHA1_Update\tppc_SHA1_Update\n-#define platform_SHA1_Final\tppc_SHA1_Final\ndiff --git a/ppc/sha1ppc.S b/ppc/sha1ppc.S\ndeleted file mode 100644\nindex 1711eef6e71..00000000000\n--- a/ppc/sha1ppc.S\n+++ /dev/null\n@@ -1,224 +0,0 @@\n-/*\n- * SHA-1 implementation for PowerPC.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- */\n-\n-/*\n- * PowerPC calling convention:\n- * %r0 - volatile temp\n- * %r1 - stack pointer.\n- * %r2 - reserved\n- * %r3-%r12 - Incoming arguments & return values; volatile.\n- * %r13-%r31 - Callee-save registers\n- * %lr - Return address, volatile\n- * %ctr - volatile\n- *\n- * Register usage in this routine:\n- * %r0 - temp\n- * %r3 - argument (pointer to 5 words of SHA state)\n- * %r4 - argument (pointer to data to hash)\n- * %r5 - Constant K in SHA round (initially number of blocks to hash)\n- * %r6-%r10 - Working copies of SHA variables A..E (actually E..A order)\n- * %r11-%r26 - Data being hashed W[].\n- * %r27-%r31 - Previous copies of A..E, for final add back.\n- * %ctr - loop count\n- */\n-\n-\n-/*\n- * We roll the registers for A, B, C, D, E around on each\n- * iteration; E on iteration t is D on iteration t+1, and so on.\n- * We use registers 6 - 10 for this.  (Registers 27 - 31 hold\n- * the previous values.)\n- */\n-#define RA(t)\t(((t)+4)%5+6)\n-#define RB(t)\t(((t)+3)%5+6)\n-#define RC(t)\t(((t)+2)%5+6)\n-#define RD(t)\t(((t)+1)%5+6)\n-#define RE(t)\t(((t)+0)%5+6)\n-\n-/* We use registers 11 - 26 for the W values */\n-#define W(t)\t((t)%16+11)\n-\n-/* Register 5 is used for the constant k */\n-\n-/*\n- * The basic SHA-1 round function is:\n- * E += ROTL(A,5) + F(B,C,D) + W[i] + K;  B = ROTL(B,30)\n- * Then the variables are renamed: (A,B,C,D,E) = (E,A,B,C,D).\n- *\n- * Every 20 rounds, the function F() and the constant K changes:\n- * - 20 rounds of f0(b,c,d) = \"bit wise b ? c : d\" =  (^b & d) + (b & c)\n- * - 20 rounds of f1(b,c,d) = b^c^d = (b^d)^c\n- * - 20 rounds of f2(b,c,d) = majority(b,c,d) = (b&d) + ((b^d)&c)\n- * - 20 more rounds of f1(b,c,d)\n- *\n- * These are all scheduled for near-optimal performance on a G4.\n- * The G4 is a 3-issue out-of-order machine with 3 ALUs, but it can only\n- * *consider* starting the oldest 3 instructions per cycle.  So to get\n- * maximum performance out of it, you have to treat it as an in-order\n- * machine.  Which means interleaving the computation round t with the\n- * computation of W[t+4].\n- *\n- * The first 16 rounds use W values loaded directly from memory, while the\n- * remaining 64 use values computed from those first 16.  We preload\n- * 4 values before starting, so there are three kinds of rounds:\n- * - The first 12 (all f0) also load the W values from memory.\n- * - The next 64 compute W(i+4) in parallel. 8*f0, 20*f1, 20*f2, 16*f1.\n- * - The last 4 (all f1) do not do anything with W.\n- *\n- * Therefore, we have 6 different round functions:\n- * STEPD0_LOAD(t,s) - Perform round t and load W(s).  s < 16\n- * STEPD0_UPDATE(t,s) - Perform round t and compute W(s).  s >= 16.\n- * STEPD1_UPDATE(t,s)\n- * STEPD2_UPDATE(t,s)\n- * STEPD1(t) - Perform round t with no load or update.\n- *\n- * The G5 is more fully out-of-order, and can find the parallelism\n- * by itself.  The big limit is that it has a 2-cycle ALU latency, so\n- * even though it's 2-way, the code has to be scheduled as if it's\n- * 4-way, which can be a limit.  To help it, we try to schedule the\n- * read of RA(t) as late as possible so it doesn't stall waiting for\n- * the previous round's RE(t-1), and we try to rotate RB(t) as early\n- * as possible while reading RC(t) (= RB(t-1)) as late as possible.\n- */\n-\n-/* the initial loads. */\n-#define LOADW(s) \\\n-\tlwz\tW(s),(s)*4(%r4)\n-\n-/*\n- * Perform a step with F0, and load W(s).  Uses W(s) as a temporary\n- * before loading it.\n- * This is actually 10 instructions, which is an awkward fit.\n- * It can execute grouped as listed, or delayed one instruction.\n- * (If delayed two instructions, there is a stall before the start of the\n- * second line.)  Thus, two iterations take 7 cycles, 3.5 cycles per round.\n- */\n-#define STEPD0_LOAD(t,s) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t);  and    W(s),RC(t),RB(t); \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;      rotlwi RB(t),RB(t),30;   \\\n-add RE(t),RE(t),W(s); add    %r0,%r0,%r5;      lwz    W(s),(s)*4(%r4);  \\\n-add RE(t),RE(t),%r0\n-\n-/*\n- * This is likewise awkward, 13 instructions.  However, it can also\n- * execute starting with 2 out of 3 possible moduli, so it does 2 rounds\n- * in 9 cycles, 4.5 cycles/round.\n- */\n-#define STEPD0_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  and    %r0,RC(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r5;  loadk; rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1;     \\\n-add RE(t),RE(t),%r0\n-\n-/* Nicely optimal.  Conveniently, also the most common. */\n-#define STEPD1_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r5;  loadk; xor %r0,%r0,RC(t);  xor W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1\n-\n-/*\n- * The naked version, no UPDATE, for the last 4 rounds.  3 cycles per.\n- * We could use W(s) as a temp register, but we don't need it.\n- */\n-#define STEPD1(t) \\\n-                        add   RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); \\\n-rotlwi RB(t),RB(t),30;  add   RE(t),RE(t),%r5;  xor    %r0,%r0,RC(t);   \\\n-add    RE(t),RE(t),%r0; rotlwi %r0,RA(t),5;     /* spare slot */        \\\n-add    RE(t),RE(t),%r0\n-\n-/*\n- * 14 instructions, 5 cycles per.  The majority function is a bit\n- * awkward to compute.  This can execute with a 1-instruction delay,\n- * but it causes a 2-instruction delay, which triggers a stall.\n- */\n-#define STEPD2_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); and    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  xor    %r0,RD(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r5;  loadk; and %r0,%r0,RC(t);  xor W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     rotlwi W(s),W(s),1;             \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30\n-\n-#define STEP0_LOAD4(t,s)\t\t\\\n-\tSTEPD0_LOAD(t,s);\t\t\\\n-\tSTEPD0_LOAD((t+1),(s)+1);\t\\\n-\tSTEPD0_LOAD((t)+2,(s)+2);\t\\\n-\tSTEPD0_LOAD((t)+3,(s)+3)\n-\n-#define STEPUP4(fn, t, s, loadk...)\t\t\\\n-\tSTEP##fn##_UPDATE(t,s,);\t\t\\\n-\tSTEP##fn##_UPDATE((t)+1,(s)+1,);\t\\\n-\tSTEP##fn##_UPDATE((t)+2,(s)+2,);\t\\\n-\tSTEP##fn##_UPDATE((t)+3,(s)+3,loadk)\n-\n-#define STEPUP20(fn, t, s, loadk...)\t\\\n-\tSTEPUP4(fn, t, s,);\t\t\\\n-\tSTEPUP4(fn, (t)+4, (s)+4,);\t\\\n-\tSTEPUP4(fn, (t)+8, (s)+8,);\t\\\n-\tSTEPUP4(fn, (t)+12, (s)+12,);\t\\\n-\tSTEPUP4(fn, (t)+16, (s)+16, loadk)\n-\n-\t.globl\tppc_sha1_core\n-ppc_sha1_core:\n-\tstwu\t%r1,-80(%r1)\n-\tstmw\t%r13,4(%r1)\n-\n-\t/* Load up A - E */\n-\tlmw\t%r27,0(%r3)\n-\n-\tmtctr\t%r5\n-\n-1:\n-\tLOADW(0)\n-\tlis\t%r5,0x5a82\n-\tmr\tRE(0),%r31\n-\tLOADW(1)\n-\tmr\tRD(0),%r30\n-\tmr\tRC(0),%r29\n-\tLOADW(2)\n-\tori\t%r5,%r5,0x7999\t/* K0-19 */\n-\tmr\tRB(0),%r28\n-\tLOADW(3)\n-\tmr\tRA(0),%r27\n-\n-\tSTEP0_LOAD4(0, 4)\n-\tSTEP0_LOAD4(4, 8)\n-\tSTEP0_LOAD4(8, 12)\n-\tSTEPUP4(D0, 12, 16,)\n-\tSTEPUP4(D0, 16, 20, lis %r5,0x6ed9)\n-\n-\tori\t%r5,%r5,0xeba1\t/* K20-39 */\n-\tSTEPUP20(D1, 20, 24, lis %r5,0x8f1b)\n-\n-\tori\t%r5,%r5,0xbcdc\t/* K40-59 */\n-\tSTEPUP20(D2, 40, 44, lis %r5,0xca62)\n-\n-\tori\t%r5,%r5,0xc1d6\t/* K60-79 */\n-\tSTEPUP4(D1, 60, 64,)\n-\tSTEPUP4(D1, 64, 68,)\n-\tSTEPUP4(D1, 68, 72,)\n-\tSTEPUP4(D1, 72, 76,)\n-\taddi\t%r4,%r4,64\n-\tSTEPD1(76)\n-\tSTEPD1(77)\n-\tSTEPD1(78)\n-\tSTEPD1(79)\n-\n-\t/* Add results to original values */\n-\tadd\t%r31,%r31,RE(0)\n-\tadd\t%r30,%r30,RD(0)\n-\tadd\t%r29,%r29,RC(0)\n-\tadd\t%r28,%r28,RB(0)\n-\tadd\t%r27,%r27,RA(0)\n-\n-\tbdnz\t1b\n-\n-\t/* Save final hash, restore registers, and return */\n-\tstmw\t%r27,0(%r3)\n-\tlmw\t%r13,4(%r1)\n-\taddi\t%r1,%r1,80\n-\tblr\n-- \n2.35.1.1441.g83331fcb493\n\n"},{"id":"451770","messageId":"Yjjr9fkybVmB53M7@camp.crustytoothpaste.net","threadId":"52786","inReplyTo":"patch-v2-1.1-e77fd23a824-20220321T170412Z-avarab@gmail.com","subject":"Re: [PATCH v2] ppc: remove custom SHA-1 implementation","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2022-03-21T21:19:49Z","receivedAt":"2022-03-21T21:20:30Z","isPatch":true,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2022-03-21 at 17:06:12, Ævar Arnfjörð Bjarmason wrote:\n> Remove the PPC_SHA1 implementation added in a6ef3518f9a ([PATCH] PPC\n> assembly implementation of SHA1, 2005-04-22). When this was added\n> Apple consumer hardware used the PPC architecture, and the\n> implementation was intended to improve SHA-1 speed there.\n> \n> Since it was added we've moved to DC_SHA1 by default, and anyone\n> wanting hard-rolled non-DC SHA-1 implementation can use OpenSSL's via\n> the OPENSSL_SHA1 knob.\n> \n> I'm unsure if this was ever supposed to work on 64-bit PPC. It clearly\n> originally targeted 32 bit PPC, but there's some mailing list\n> references to this being tried on G5 (PPC 970). I can't get it to do\n> anything but segfault on the BE POWER8 machine in the GCC compile\n> farm. Anyone caring about speed on PPC these days is likely to be\n> using IBM's POWER, not PPC 970.\n> \n> There have been proposals to entirely remove non-DC_SHA1\n> implementations from the tree[1]. I think per [2] that would be a bit\n> overzealous. I.e. there are various set-ups git's speed is going to be\n> more important than the relatively implausible SHA-1 collision attack,\n> or where such attacks are entirely mitigated by other means (e.g. by\n> incoming objects being checked with DC_SHA1).\n> \n> The main reason for doing so at this point is to simplify follow-up\n> Makefile change. Since PPC_SHA1 included the only in-tree *.S assembly\n> file we needed to keep around special support for building objects\n> from it. By getting rid of it we know we'll always build *.o from *.c\n> files, which makes the build process simpler.\n> \n> As an aside the code being removed here was also throwing warnings\n> with the \"-pedantic\" flag, but let's remove it instead of fixing it,\n> as 544d93bc3b4 (block-sha1: remove use of obsolete x86 assembly,\n> 2022-03-10) did for block-sha1/*.\n\nWhile I don't agree that we shouldn't remove the other non-DC SHA-1\nimplementations, I do agree that we should remove this one.  Given the\ntesting you've done and the fact that almost everyone desiring speed is\nusing a 64-bit machine these days, I think it's unlikely that anyone is\nusing this in the real world.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"462269","messageId":"patch-v3-2.2-cb3bc8b5029-20220831T090744Z-avarab@gmail.com","threadId":"52786","inReplyTo":"cover-v3-0.2-00000000000-20220831T090744Z-avarab@gmail.com","subject":"[PATCH v3 2/2] Makefile: use $(OBJECTS) instead of $(C_OBJ)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-08-31T09:18:44Z","receivedAt":"2022-08-31T09:19:14Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"In the preceding commit $(C_OBJ) added in c373991375a (Makefile: list\ngenerated object files in OBJECTS, 2010-01-26) became synonymous with\n$(OBJECTS). Let's avoid the indirection and use the $(OBJECTS)\nvariable directly instead.\n\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n Makefile | 6 ++----\n 1 file changed, 2 insertions(+), 4 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 7feda7e79be..8956cace8eb 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -2590,9 +2590,7 @@ missing_compdb_dir =\n compdb_args =\n endif\n \n-C_OBJ := $(OBJECTS)\n-\n-$(C_OBJ): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n+$(OBJECTS): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n \t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n \n %.s: %.c GIT-CFLAGS FORCE\n@@ -3084,7 +3082,7 @@ t/helper/test-%$X: t/helper/test-%.o GIT-LDFLAGS $(GITLIBS) $(REFTABLE_TEST_LIB)\n check-sha1:: t/helper/test-tool$X\n \tt/helper/test-sha1.sh\n \n-SP_OBJ = $(patsubst %.o,%.sp,$(C_OBJ))\n+SP_OBJ = $(patsubst %.o,%.sp,$(OBJECTS))\n \n $(SP_OBJ): %.sp: %.c %.o\n \t$(QUIET_SP)cgcc -no-compile $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) \\\n-- \n2.37.3.1406.g184357183a6\n\n"},{"id":"462270","messageId":"patch-v3-1.2-87a204b8937-20220831T090744Z-avarab@gmail.com","threadId":"52786","inReplyTo":"cover-v3-0.2-00000000000-20220831T090744Z-avarab@gmail.com","subject":"[PATCH v3 1/2] Makefile + hash.h: remove PPC_SHA1 implementation","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-08-31T09:18:43Z","receivedAt":"2022-08-31T09:19:18Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"Remove the PPC_SHA1 implementation added in a6ef3518f9a ([PATCH] PPC\nassembly implementation of SHA1, 2005-04-22). When this was added\nApple consumer hardware used the PPC architecture, and the\nimplementation was intended to improve SHA-1 speed there.\n\nSince it was added we've moved to using sha1collisiondetection by\ndefault, and anyone wanting hard-rolled non-DC SHA-1 implementation\ncan use OpenSSL's via the OPENSSL_SHA1 knob.\n\nThe PPC_SHA1 originally originally targeted 32 bit PPC, and later the\n64 bit PPC 970 (a.k.a. Apple PowerPC G5). See 926172c5e48 (block-sha1:\nimprove code on large-register-set machines, 2009-08-10) for a\nreference about the performance on G5 (a comment in block-sha1/sha1.c\nbeing removed here).\n\nI can't get it to do anything but segfault on both the BE and LE POWER\nmachines in the GCC compile farm[1]. Anyone who's concerned about\nperformance on PPC these days is likely to be using the IBM POWER\nprocessors.\n\nThere have been proposals to entirely remove non-sha1collisiondetection\nimplementations from the tree[2]. I think per [3] that would be a bit\noverzealous. I.e. there are various set-ups git's speed is going to be\nmore important than the relatively implausible SHA-1 collision attack,\nor where such attacks are entirely mitigated by other means (e.g. by\nincoming objects being checked with DC_SHA1).\n\nBut that really doesn't apply to PPC_SHA1 in particular, which seems\nto have outlived its usefulness.\n\nAs this gets rid of the only in-tree *.S assembly file we can remove\nthe small bits of logic from the Makefile needed to build objects\nfrom *.S (as opposed to *.c)\n\nThe code being removed here was also throwing warnings with the\n\"-pedantic\" flag, it could have been fixed as 544d93bc3b4 (block-sha1:\nremove use of obsolete x86 assembly, 2022-03-10) did for block-sha1/*,\nbut as noted above let's remove it instead.\n\n1. https://cfarm.tetaneutral.net/machines/list/\n   Tested on gcc{110,112,135,203}, a mixture of POWER [789] ppc64 and\n   ppc64le. All segfault in anything needing object\n   hashing (e.g. t/t1007-hash-object.sh) when compiled with\n   PPC_SHA1=Y.\n2. https://lore.kernel.org/git/20200223223758.120941-1-mh@glandium.org/\n3. https://lore.kernel.org/git/20200224044732.GK1018190@coredump.intra.peff.net/\n\nAcked-by: brian m. carlson\" <sandals@crustytoothpaste.net>\nSigned-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n---\n INSTALL           |   3 +-\n Makefile          |  18 ++--\n block-sha1/sha1.c |   4 -\n configure.ac      |   3 -\n hash.h            |   6 +-\n ppc/sha1.c        |  72 ---------------\n ppc/sha1.h        |  25 ------\n ppc/sha1ppc.S     | 224 ----------------------------------------------\n 8 files changed, 8 insertions(+), 347 deletions(-)\n delete mode 100644 ppc/sha1.c\n delete mode 100644 ppc/sha1.h\n delete mode 100644 ppc/sha1ppc.S\n\ndiff --git a/INSTALL b/INSTALL\nindex 4140a3f5c8b..89b15d71df5 100644\n--- a/INSTALL\n+++ b/INSTALL\n@@ -135,8 +135,7 @@ Issues of note:\n \n \t  By default, git uses OpenSSL for SHA1 but it will use its own\n \t  library (inspired by Mozilla's) with either NO_OPENSSL or\n-\t  BLK_SHA1.  Also included is a version optimized for PowerPC\n-\t  (PPC_SHA1).\n+\t  BLK_SHA1.\n \n \t- \"libcurl\" library is used for fetching and pushing\n \t  repositories over http:// or https://, as well as by\ndiff --git a/Makefile b/Makefile\nindex eac30126e29..7feda7e79be 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -155,9 +155,6 @@ include shared.mak\n # Define BLK_SHA1 environment variable to make use of the bundled\n # optimized C SHA1 routine.\n #\n-# Define PPC_SHA1 environment variable when running make to make use of\n-# a bundled SHA1 routine optimized for PowerPC.\n-#\n # Define DC_SHA1 to unconditionally enable the collision-detecting sha1\n # algorithm. This is slower, but may detect attempted collision attacks.\n # Takes priority over other *_SHA1 knobs.\n@@ -1802,6 +1799,10 @@ ifdef APPLE_COMMON_CRYPTO\n \tSHA1_MAX_BLOCK_SIZE = 1024L*1024L*1024L\n endif\n \n+ifdef PPC_SHA1\n+$(error the PPC_SHA1 flag has been removed along with the PowerPC-specific SHA-1 implementation.)\n+endif\n+\n ifdef OPENSSL_SHA1\n \tEXTLIBS += $(LIB_4_CRYPTO)\n \tBASIC_CFLAGS += -DSHA1_OPENSSL\n@@ -1810,10 +1811,6 @@ ifdef BLK_SHA1\n \tLIB_OBJS += block-sha1/sha1.o\n \tBASIC_CFLAGS += -DSHA1_BLK\n else\n-ifdef PPC_SHA1\n-\tLIB_OBJS += ppc/sha1.o ppc/sha1ppc.o\n-\tBASIC_CFLAGS += -DSHA1_PPC\n-else\n ifdef APPLE_COMMON_CRYPTO\n \tCOMPAT_CFLAGS += -DCOMMON_DIGEST_FOR_OPENSSL\n \tBASIC_CFLAGS += -DSHA1_APPLE\n@@ -1847,7 +1844,6 @@ endif\n endif\n endif\n endif\n-endif\n \n ifdef OPENSSL_SHA256\n \tEXTLIBS += $(LIB_4_CRYPTO)\n@@ -2594,14 +2590,10 @@ missing_compdb_dir =\n compdb_args =\n endif\n \n-ASM_SRC := $(wildcard $(OBJECTS:o=S))\n-ASM_OBJ := $(ASM_SRC:S=o)\n-C_OBJ := $(filter-out $(ASM_OBJ),$(OBJECTS))\n+C_OBJ := $(OBJECTS)\n \n $(C_OBJ): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n \t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n-$(ASM_OBJ): %.o: %.S GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n-\t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n \n %.s: %.c GIT-CFLAGS FORCE\n \t$(QUIET_CC)$(CC) -o $@ -S $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nindex 5974cd7dd3c..80cebd27564 100644\n--- a/block-sha1/sha1.c\n+++ b/block-sha1/sha1.c\n@@ -28,10 +28,6 @@\n  * try to do the silly \"optimize away loads\" part because it won't\n  * see what the value will be).\n  *\n- * Ben Herrenschmidt reports that on PPC, the C version comes close\n- * to the optimized asm with this (ie on PPC you don't want that\n- * 'volatile', since there are lots of registers).\n- *\n  * On ARM we get the best code generation by forcing a full memory barrier\n  * between each SHA_ROUND, otherwise gcc happily get wild with spilling and\n  * the stack frame size simply explode and performance goes down the drain.\ndiff --git a/configure.ac b/configure.ac\nindex 7dcd0482042..38ff86678a0 100644\n--- a/configure.ac\n+++ b/configure.ac\n@@ -237,9 +237,6 @@ AC_MSG_NOTICE([CHECKS for site configuration])\n # tests.  These tests take up a significant amount of the total test time\n # but are not needed unless you plan to talk to SVN repos.\n #\n-# Define PPC_SHA1 environment variable when running make to make use of\n-# a bundled SHA1 routine optimized for PowerPC.\n-#\n # Define NO_OPENSSL environment variable if you do not have OpenSSL.\n #\n # Define OPENSSLDIR=/foo/bar if your openssl header and library files are in\ndiff --git a/hash.h b/hash.h\nindex ea87ae9d92f..36b64165fc9 100644\n--- a/hash.h\n+++ b/hash.h\n@@ -4,9 +4,7 @@\n #include \"git-compat-util.h\"\n #include \"repository.h\"\n \n-#if defined(SHA1_PPC)\n-#include \"ppc/sha1.h\"\n-#elif defined(SHA1_APPLE)\n+#if defined(SHA1_APPLE)\n #include <CommonCrypto/CommonDigest.h>\n #elif defined(SHA1_OPENSSL)\n #include <openssl/sha.h>\n@@ -32,7 +30,7 @@\n  * platform's underlying implementation of SHA-1; could be OpenSSL,\n  * blk_SHA, Apple CommonCrypto, etc...  Note that the relevant\n  * SHA-1 header may have already defined platform_SHA_CTX for our\n- * own implementations like block-sha1 and ppc-sha1, so we list\n+ * own implementations like block-sha1, so we list\n  * the default for OpenSSL compatible SHA-1 implementations here.\n  */\n #define platform_SHA_CTX\tSHA_CTX\ndiff --git a/ppc/sha1.c b/ppc/sha1.c\ndeleted file mode 100644\nindex 1b705cee1fe..00000000000\n--- a/ppc/sha1.c\n+++ /dev/null\n@@ -1,72 +0,0 @@\n-/*\n- * SHA-1 implementation.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- *\n- * This version assumes we are running on a big-endian machine.\n- * It calls an external sha1_core() to process blocks of 64 bytes.\n- */\n-#include <stdio.h>\n-#include <string.h>\n-#include \"sha1.h\"\n-\n-void ppc_sha1_core(uint32_t *hash, const unsigned char *p,\n-\t\t   unsigned int nblocks);\n-\n-int ppc_SHA1_Init(ppc_SHA_CTX *c)\n-{\n-\tc->hash[0] = 0x67452301;\n-\tc->hash[1] = 0xEFCDAB89;\n-\tc->hash[2] = 0x98BADCFE;\n-\tc->hash[3] = 0x10325476;\n-\tc->hash[4] = 0xC3D2E1F0;\n-\tc->len = 0;\n-\tc->cnt = 0;\n-\treturn 0;\n-}\n-\n-int ppc_SHA1_Update(ppc_SHA_CTX *c, const void *ptr, unsigned long n)\n-{\n-\tunsigned long nb;\n-\tconst unsigned char *p = ptr;\n-\n-\tc->len += (uint64_t) n << 3;\n-\twhile (n != 0) {\n-\t\tif (c->cnt || n < 64) {\n-\t\t\tnb = 64 - c->cnt;\n-\t\t\tif (nb > n)\n-\t\t\t\tnb = n;\n-\t\t\tmemcpy(&c->buf.b[c->cnt], p, nb);\n-\t\t\tif ((c->cnt += nb) == 64) {\n-\t\t\t\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\t\t\t\tc->cnt = 0;\n-\t\t\t}\n-\t\t} else {\n-\t\t\tnb = n >> 6;\n-\t\t\tppc_sha1_core(c->hash, p, nb);\n-\t\t\tnb <<= 6;\n-\t\t}\n-\t\tn -= nb;\n-\t\tp += nb;\n-\t}\n-\treturn 0;\n-}\n-\n-int ppc_SHA1_Final(unsigned char *hash, ppc_SHA_CTX *c)\n-{\n-\tunsigned int cnt = c->cnt;\n-\n-\tc->buf.b[cnt++] = 0x80;\n-\tif (cnt > 56) {\n-\t\tif (cnt < 64)\n-\t\t\tmemset(&c->buf.b[cnt], 0, 64 - cnt);\n-\t\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\t\tcnt = 0;\n-\t}\n-\tif (cnt < 56)\n-\t\tmemset(&c->buf.b[cnt], 0, 56 - cnt);\n-\tc->buf.l[7] = c->len;\n-\tppc_sha1_core(c->hash, c->buf.b, 1);\n-\tmemcpy(hash, c->hash, 20);\n-\treturn 0;\n-}\ndiff --git a/ppc/sha1.h b/ppc/sha1.h\ndeleted file mode 100644\nindex 9b24b326159..00000000000\n--- a/ppc/sha1.h\n+++ /dev/null\n@@ -1,25 +0,0 @@\n-/*\n- * SHA-1 implementation.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- */\n-#include <stdint.h>\n-\n-typedef struct {\n-\tuint32_t hash[5];\n-\tuint32_t cnt;\n-\tuint64_t len;\n-\tunion {\n-\t\tunsigned char b[64];\n-\t\tuint64_t l[8];\n-\t} buf;\n-} ppc_SHA_CTX;\n-\n-int ppc_SHA1_Init(ppc_SHA_CTX *c);\n-int ppc_SHA1_Update(ppc_SHA_CTX *c, const void *p, unsigned long n);\n-int ppc_SHA1_Final(unsigned char *hash, ppc_SHA_CTX *c);\n-\n-#define platform_SHA_CTX\tppc_SHA_CTX\n-#define platform_SHA1_Init\tppc_SHA1_Init\n-#define platform_SHA1_Update\tppc_SHA1_Update\n-#define platform_SHA1_Final\tppc_SHA1_Final\ndiff --git a/ppc/sha1ppc.S b/ppc/sha1ppc.S\ndeleted file mode 100644\nindex 1711eef6e71..00000000000\n--- a/ppc/sha1ppc.S\n+++ /dev/null\n@@ -1,224 +0,0 @@\n-/*\n- * SHA-1 implementation for PowerPC.\n- *\n- * Copyright (C) 2005 Paul Mackerras <paulus@samba.org>\n- */\n-\n-/*\n- * PowerPC calling convention:\n- * %r0 - volatile temp\n- * %r1 - stack pointer.\n- * %r2 - reserved\n- * %r3-%r12 - Incoming arguments & return values; volatile.\n- * %r13-%r31 - Callee-save registers\n- * %lr - Return address, volatile\n- * %ctr - volatile\n- *\n- * Register usage in this routine:\n- * %r0 - temp\n- * %r3 - argument (pointer to 5 words of SHA state)\n- * %r4 - argument (pointer to data to hash)\n- * %r5 - Constant K in SHA round (initially number of blocks to hash)\n- * %r6-%r10 - Working copies of SHA variables A..E (actually E..A order)\n- * %r11-%r26 - Data being hashed W[].\n- * %r27-%r31 - Previous copies of A..E, for final add back.\n- * %ctr - loop count\n- */\n-\n-\n-/*\n- * We roll the registers for A, B, C, D, E around on each\n- * iteration; E on iteration t is D on iteration t+1, and so on.\n- * We use registers 6 - 10 for this.  (Registers 27 - 31 hold\n- * the previous values.)\n- */\n-#define RA(t)\t(((t)+4)%5+6)\n-#define RB(t)\t(((t)+3)%5+6)\n-#define RC(t)\t(((t)+2)%5+6)\n-#define RD(t)\t(((t)+1)%5+6)\n-#define RE(t)\t(((t)+0)%5+6)\n-\n-/* We use registers 11 - 26 for the W values */\n-#define W(t)\t((t)%16+11)\n-\n-/* Register 5 is used for the constant k */\n-\n-/*\n- * The basic SHA-1 round function is:\n- * E += ROTL(A,5) + F(B,C,D) + W[i] + K;  B = ROTL(B,30)\n- * Then the variables are renamed: (A,B,C,D,E) = (E,A,B,C,D).\n- *\n- * Every 20 rounds, the function F() and the constant K changes:\n- * - 20 rounds of f0(b,c,d) = \"bit wise b ? c : d\" =  (^b & d) + (b & c)\n- * - 20 rounds of f1(b,c,d) = b^c^d = (b^d)^c\n- * - 20 rounds of f2(b,c,d) = majority(b,c,d) = (b&d) + ((b^d)&c)\n- * - 20 more rounds of f1(b,c,d)\n- *\n- * These are all scheduled for near-optimal performance on a G4.\n- * The G4 is a 3-issue out-of-order machine with 3 ALUs, but it can only\n- * *consider* starting the oldest 3 instructions per cycle.  So to get\n- * maximum performance out of it, you have to treat it as an in-order\n- * machine.  Which means interleaving the computation round t with the\n- * computation of W[t+4].\n- *\n- * The first 16 rounds use W values loaded directly from memory, while the\n- * remaining 64 use values computed from those first 16.  We preload\n- * 4 values before starting, so there are three kinds of rounds:\n- * - The first 12 (all f0) also load the W values from memory.\n- * - The next 64 compute W(i+4) in parallel. 8*f0, 20*f1, 20*f2, 16*f1.\n- * - The last 4 (all f1) do not do anything with W.\n- *\n- * Therefore, we have 6 different round functions:\n- * STEPD0_LOAD(t,s) - Perform round t and load W(s).  s < 16\n- * STEPD0_UPDATE(t,s) - Perform round t and compute W(s).  s >= 16.\n- * STEPD1_UPDATE(t,s)\n- * STEPD2_UPDATE(t,s)\n- * STEPD1(t) - Perform round t with no load or update.\n- *\n- * The G5 is more fully out-of-order, and can find the parallelism\n- * by itself.  The big limit is that it has a 2-cycle ALU latency, so\n- * even though it's 2-way, the code has to be scheduled as if it's\n- * 4-way, which can be a limit.  To help it, we try to schedule the\n- * read of RA(t) as late as possible so it doesn't stall waiting for\n- * the previous round's RE(t-1), and we try to rotate RB(t) as early\n- * as possible while reading RC(t) (= RB(t-1)) as late as possible.\n- */\n-\n-/* the initial loads. */\n-#define LOADW(s) \\\n-\tlwz\tW(s),(s)*4(%r4)\n-\n-/*\n- * Perform a step with F0, and load W(s).  Uses W(s) as a temporary\n- * before loading it.\n- * This is actually 10 instructions, which is an awkward fit.\n- * It can execute grouped as listed, or delayed one instruction.\n- * (If delayed two instructions, there is a stall before the start of the\n- * second line.)  Thus, two iterations take 7 cycles, 3.5 cycles per round.\n- */\n-#define STEPD0_LOAD(t,s) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t);  and    W(s),RC(t),RB(t); \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;      rotlwi RB(t),RB(t),30;   \\\n-add RE(t),RE(t),W(s); add    %r0,%r0,%r5;      lwz    W(s),(s)*4(%r4);  \\\n-add RE(t),RE(t),%r0\n-\n-/*\n- * This is likewise awkward, 13 instructions.  However, it can also\n- * execute starting with 2 out of 3 possible moduli, so it does 2 rounds\n- * in 9 cycles, 4.5 cycles/round.\n- */\n-#define STEPD0_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); andc   %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  and    %r0,RC(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r5;  loadk; rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1;     \\\n-add RE(t),RE(t),%r0\n-\n-/* Nicely optimal.  Conveniently, also the most common. */\n-#define STEPD1_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r5;  loadk; xor %r0,%r0,RC(t);  xor W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     xor    W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30;  rotlwi W(s),W(s),1\n-\n-/*\n- * The naked version, no UPDATE, for the last 4 rounds.  3 cycles per.\n- * We could use W(s) as a temp register, but we don't need it.\n- */\n-#define STEPD1(t) \\\n-                        add   RE(t),RE(t),W(t); xor    %r0,RD(t),RB(t); \\\n-rotlwi RB(t),RB(t),30;  add   RE(t),RE(t),%r5;  xor    %r0,%r0,RC(t);   \\\n-add    RE(t),RE(t),%r0; rotlwi %r0,RA(t),5;     /* spare slot */        \\\n-add    RE(t),RE(t),%r0\n-\n-/*\n- * 14 instructions, 5 cycles per.  The majority function is a bit\n- * awkward to compute.  This can execute with a 1-instruction delay,\n- * but it causes a 2-instruction delay, which triggers a stall.\n- */\n-#define STEPD2_UPDATE(t,s,loadk...) \\\n-add RE(t),RE(t),W(t); and    %r0,RD(t),RB(t); xor    W(s),W((s)-16),W((s)-3); \\\n-add RE(t),RE(t),%r0;  xor    %r0,RD(t),RB(t); xor    W(s),W(s),W((s)-8);      \\\n-add RE(t),RE(t),%r5;  loadk; and %r0,%r0,RC(t);  xor W(s),W(s),W((s)-14);     \\\n-add RE(t),RE(t),%r0;  rotlwi %r0,RA(t),5;     rotlwi W(s),W(s),1;             \\\n-add RE(t),RE(t),%r0;  rotlwi RB(t),RB(t),30\n-\n-#define STEP0_LOAD4(t,s)\t\t\\\n-\tSTEPD0_LOAD(t,s);\t\t\\\n-\tSTEPD0_LOAD((t+1),(s)+1);\t\\\n-\tSTEPD0_LOAD((t)+2,(s)+2);\t\\\n-\tSTEPD0_LOAD((t)+3,(s)+3)\n-\n-#define STEPUP4(fn, t, s, loadk...)\t\t\\\n-\tSTEP##fn##_UPDATE(t,s,);\t\t\\\n-\tSTEP##fn##_UPDATE((t)+1,(s)+1,);\t\\\n-\tSTEP##fn##_UPDATE((t)+2,(s)+2,);\t\\\n-\tSTEP##fn##_UPDATE((t)+3,(s)+3,loadk)\n-\n-#define STEPUP20(fn, t, s, loadk...)\t\\\n-\tSTEPUP4(fn, t, s,);\t\t\\\n-\tSTEPUP4(fn, (t)+4, (s)+4,);\t\\\n-\tSTEPUP4(fn, (t)+8, (s)+8,);\t\\\n-\tSTEPUP4(fn, (t)+12, (s)+12,);\t\\\n-\tSTEPUP4(fn, (t)+16, (s)+16, loadk)\n-\n-\t.globl\tppc_sha1_core\n-ppc_sha1_core:\n-\tstwu\t%r1,-80(%r1)\n-\tstmw\t%r13,4(%r1)\n-\n-\t/* Load up A - E */\n-\tlmw\t%r27,0(%r3)\n-\n-\tmtctr\t%r5\n-\n-1:\n-\tLOADW(0)\n-\tlis\t%r5,0x5a82\n-\tmr\tRE(0),%r31\n-\tLOADW(1)\n-\tmr\tRD(0),%r30\n-\tmr\tRC(0),%r29\n-\tLOADW(2)\n-\tori\t%r5,%r5,0x7999\t/* K0-19 */\n-\tmr\tRB(0),%r28\n-\tLOADW(3)\n-\tmr\tRA(0),%r27\n-\n-\tSTEP0_LOAD4(0, 4)\n-\tSTEP0_LOAD4(4, 8)\n-\tSTEP0_LOAD4(8, 12)\n-\tSTEPUP4(D0, 12, 16,)\n-\tSTEPUP4(D0, 16, 20, lis %r5,0x6ed9)\n-\n-\tori\t%r5,%r5,0xeba1\t/* K20-39 */\n-\tSTEPUP20(D1, 20, 24, lis %r5,0x8f1b)\n-\n-\tori\t%r5,%r5,0xbcdc\t/* K40-59 */\n-\tSTEPUP20(D2, 40, 44, lis %r5,0xca62)\n-\n-\tori\t%r5,%r5,0xc1d6\t/* K60-79 */\n-\tSTEPUP4(D1, 60, 64,)\n-\tSTEPUP4(D1, 64, 68,)\n-\tSTEPUP4(D1, 68, 72,)\n-\tSTEPUP4(D1, 72, 76,)\n-\taddi\t%r4,%r4,64\n-\tSTEPD1(76)\n-\tSTEPD1(77)\n-\tSTEPD1(78)\n-\tSTEPD1(79)\n-\n-\t/* Add results to original values */\n-\tadd\t%r31,%r31,RE(0)\n-\tadd\t%r30,%r30,RD(0)\n-\tadd\t%r29,%r29,RC(0)\n-\tadd\t%r28,%r28,RB(0)\n-\tadd\t%r27,%r27,RA(0)\n-\n-\tbdnz\t1b\n-\n-\t/* Save final hash, restore registers, and return */\n-\tstmw\t%r27,0(%r3)\n-\tlmw\t%r13,4(%r1)\n-\taddi\t%r1,%r1,80\n-\tblr\n-- \n2.37.3.1406.g184357183a6\n\n"},{"id":"462271","messageId":"cover-v3-0.2-00000000000-20220831T090744Z-avarab@gmail.com","threadId":"52786","inReplyTo":"patch-v2-1.1-e77fd23a824-20220321T170412Z-avarab@gmail.com","subject":"[PATCH v3 0/2] Makefile + hash.h: remove PPC_SHA1 implementation","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-08-31T09:18:42Z","receivedAt":"2022-08-31T09:19:20Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"This is a re-roll of [1] sent back in March. It removes the PPC_SHA1\nimplementation, which as noted in in 1/2 is likely completely unused\nat this point, if anyone's using it we have other better tested (and\njust as fast) SHA-1 implementations.\n\nI then included this change in a larger series sent in late April[2],\nwhich stalled for lack of feedback from the OSX crowd[3].\n\nBut this part should be uncontroversial, and is an obvious win in\nterms of the diffstat, and in reducing our complexity surface in this\narea.\n\nChanges since v2:\n\n * Much rewritten commit message (see range-diff)\n * Rebased on upstream changes\n * More PPC_SHA1 removal of comments/docs that I missed the first time\n   around.\n * The s/C_OBJS/OBJECTS/ Makefile change is now split into a new 2/2.\n * Added brian's Acked-by, per [4] (not explicitly an \"Acked-by\", but\n   I thought adding amounted to a fair paraphrasing of the feedback\n   there)\n\n1. https://lore.kernel.org/git/patch-v2-1.1-e77fd23a824-20220321T170412Z-avarab@gmail.com/\n2. https://lore.kernel.org/git/cover-0.5-00000000000-20220422T094624Z-avarab@gmail.com/\n3. https://lore.kernel.org/git/xmqqv8v17xrl.fsf@gitster.g/\n4. https://lore.kernel.org/git/Yjjr9fkybVmB53M7@camp.crustytoothpaste.net/\n\nÆvar Arnfjörð Bjarmason (2):\n  Makefile + hash.h: remove PPC_SHA1 implementation\n  Makefile: use $(OBJECTS) instead of $(C_OBJ)\n\n INSTALL           |   3 +-\n Makefile          |  22 ++---\n block-sha1/sha1.c |   4 -\n configure.ac      |   3 -\n hash.h            |   6 +-\n ppc/sha1.c        |  72 ---------------\n ppc/sha1.h        |  25 ------\n ppc/sha1ppc.S     | 224 ----------------------------------------------\n 8 files changed, 9 insertions(+), 350 deletions(-)\n delete mode 100644 ppc/sha1.c\n delete mode 100644 ppc/sha1.h\n delete mode 100644 ppc/sha1ppc.S\n\nRange-diff against v2:\n1:  e77fd23a824 ! 1:  87a204b8937 ppc: remove custom SHA-1 implementation\n    @@ Metadata\n     Author: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## Commit message ##\n    -    ppc: remove custom SHA-1 implementation\n    +    Makefile + hash.h: remove PPC_SHA1 implementation\n     \n         Remove the PPC_SHA1 implementation added in a6ef3518f9a ([PATCH] PPC\n         assembly implementation of SHA1, 2005-04-22). When this was added\n         Apple consumer hardware used the PPC architecture, and the\n         implementation was intended to improve SHA-1 speed there.\n     \n    -    Since it was added we've moved to DC_SHA1 by default, and anyone\n    -    wanting hard-rolled non-DC SHA-1 implementation can use OpenSSL's via\n    -    the OPENSSL_SHA1 knob.\n    +    Since it was added we've moved to using sha1collisiondetection by\n    +    default, and anyone wanting hard-rolled non-DC SHA-1 implementation\n    +    can use OpenSSL's via the OPENSSL_SHA1 knob.\n     \n    -    I'm unsure if this was ever supposed to work on 64-bit PPC. It clearly\n    -    originally targeted 32 bit PPC, but there's some mailing list\n    -    references to this being tried on G5 (PPC 970). I can't get it to do\n    -    anything but segfault on the BE POWER8 machine in the GCC compile\n    -    farm. Anyone caring about speed on PPC these days is likely to be\n    -    using IBM's POWER, not PPC 970.\n    +    The PPC_SHA1 originally originally targeted 32 bit PPC, and later the\n    +    64 bit PPC 970 (a.k.a. Apple PowerPC G5). See 926172c5e48 (block-sha1:\n    +    improve code on large-register-set machines, 2009-08-10) for a\n    +    reference about the performance on G5 (a comment in block-sha1/sha1.c\n    +    being removed here).\n     \n    -    There have been proposals to entirely remove non-DC_SHA1\n    -    implementations from the tree[1]. I think per [2] that would be a bit\n    +    I can't get it to do anything but segfault on both the BE and LE POWER\n    +    machines in the GCC compile farm[1]. Anyone who's concerned about\n    +    performance on PPC these days is likely to be using the IBM POWER\n    +    processors.\n    +\n    +    There have been proposals to entirely remove non-sha1collisiondetection\n    +    implementations from the tree[2]. I think per [3] that would be a bit\n         overzealous. I.e. there are various set-ups git's speed is going to be\n         more important than the relatively implausible SHA-1 collision attack,\n         or where such attacks are entirely mitigated by other means (e.g. by\n         incoming objects being checked with DC_SHA1).\n     \n    -    The main reason for doing so at this point is to simplify follow-up\n    -    Makefile change. Since PPC_SHA1 included the only in-tree *.S assembly\n    -    file we needed to keep around special support for building objects\n    -    from it. By getting rid of it we know we'll always build *.o from *.c\n    -    files, which makes the build process simpler.\n    +    But that really doesn't apply to PPC_SHA1 in particular, which seems\n    +    to have outlived its usefulness.\n    +\n    +    As this gets rid of the only in-tree *.S assembly file we can remove\n    +    the small bits of logic from the Makefile needed to build objects\n    +    from *.S (as opposed to *.c)\n     \n    -    As an aside the code being removed here was also throwing warnings\n    -    with the \"-pedantic\" flag, but let's remove it instead of fixing it,\n    -    as 544d93bc3b4 (block-sha1: remove use of obsolete x86 assembly,\n    -    2022-03-10) did for block-sha1/*.\n    +    The code being removed here was also throwing warnings with the\n    +    \"-pedantic\" flag, it could have been fixed as 544d93bc3b4 (block-sha1:\n    +    remove use of obsolete x86 assembly, 2022-03-10) did for block-sha1/*,\n    +    but as noted above let's remove it instead.\n     \n    -    1. https://lore.kernel.org/git/20200223223758.120941-1-mh@glandium.org/\n    -    2. https://lore.kernel.org/git/20200224044732.GK1018190@coredump.intra.peff.net/\n    +    1. https://cfarm.tetaneutral.net/machines/list/\n    +       Tested on gcc{110,112,135,203}, a mixture of POWER [789] ppc64 and\n    +       ppc64le. All segfault in anything needing object\n    +       hashing (e.g. t/t1007-hash-object.sh) when compiled with\n    +       PPC_SHA1=Y.\n    +    2. https://lore.kernel.org/git/20200223223758.120941-1-mh@glandium.org/\n    +    3. https://lore.kernel.org/git/20200224044732.GK1018190@coredump.intra.peff.net/\n     \n    +    Acked-by: brian m. carlson\" <sandals@crustytoothpaste.net>\n         Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n     \n      ## INSTALL ##\n    @@ Makefile: include shared.mak\n      # Define DC_SHA1 to unconditionally enable the collision-detecting sha1\n      # algorithm. This is slower, but may detect attempted collision attacks.\n      # Takes priority over other *_SHA1 knobs.\n    -@@ Makefile: ifdef OPENSSL_SHA1\n    - \tEXTLIBS += $(LIB_4_CRYPTO)\n    - \tBASIC_CFLAGS += -DSHA1_OPENSSL\n    - else\n    +@@ Makefile: ifdef APPLE_COMMON_CRYPTO\n    + \tSHA1_MAX_BLOCK_SIZE = 1024L*1024L*1024L\n    + endif\n    + \n     +ifdef PPC_SHA1\n    -+$(error PPC_SHA1 has been removed! Use DC_SHA1 instead, which is the default)\n    ++$(error the PPC_SHA1 flag has been removed along with the PowerPC-specific SHA-1 implementation.)\n     +endif\n    - ifdef BLK_SHA1\n    ++\n    + ifdef OPENSSL_SHA1\n    + \tEXTLIBS += $(LIB_4_CRYPTO)\n    + \tBASIC_CFLAGS += -DSHA1_OPENSSL\n    +@@ Makefile: ifdef BLK_SHA1\n      \tLIB_OBJS += block-sha1/sha1.o\n      \tBASIC_CFLAGS += -DSHA1_BLK\n      else\n    @@ Makefile: missing_compdb_dir =\n     -ASM_SRC := $(wildcard $(OBJECTS:o=S))\n     -ASM_OBJ := $(ASM_SRC:S=o)\n     -C_OBJ := $(filter-out $(ASM_OBJ),$(OBJECTS))\n    -+C_OBJ = $(OBJECTS)\n    ++C_OBJ := $(OBJECTS)\n      \n      $(C_OBJ): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n      \t$(QUIET_CC)$(CC) -o $*.o -c $(dep_args) $(compdb_args) $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n    @@ Makefile: missing_compdb_dir =\n      %.s: %.c GIT-CFLAGS FORCE\n      \t$(QUIET_CC)$(CC) -o $@ -S $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n     \n    + ## block-sha1/sha1.c ##\n    +@@\n    +  * try to do the silly \"optimize away loads\" part because it won't\n    +  * see what the value will be).\n    +  *\n    +- * Ben Herrenschmidt reports that on PPC, the C version comes close\n    +- * to the optimized asm with this (ie on PPC you don't want that\n    +- * 'volatile', since there are lots of registers).\n    +- *\n    +  * On ARM we get the best code generation by forcing a full memory barrier\n    +  * between each SHA_ROUND, otherwise gcc happily get wild with spilling and\n    +  * the stack frame size simply explode and performance goes down the drain.\n    +\n      ## configure.ac ##\n     @@ configure.ac: AC_MSG_NOTICE([CHECKS for site configuration])\n      # tests.  These tests take up a significant amount of the total test time\n-:  ----------- > 2:  cb3bc8b5029 Makefile: use $(OBJECTS) instead of $(C_OBJ)\n-- \n2.37.3.1406.g184357183a6\n\n"},{"id":"462330","messageId":"xmqqmtbkdr5n.fsf@gitster.g","threadId":"52786","inReplyTo":"patch-v3-2.2-cb3bc8b5029-20220831T090744Z-avarab@gmail.com","subject":"Re: [PATCH v3 2/2] Makefile: use $(OBJECTS) instead of $(C_OBJ)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-08-31T21:44:36Z","receivedAt":"2022-08-31T21:44:45Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n\n> In the preceding commit $(C_OBJ) added in c373991375a (Makefile: list\n> generated object files in OBJECTS, 2010-01-26) became synonymous with\n> $(OBJECTS). Let's avoid the indirection and use the $(OBJECTS)\n> variable directly instead.\n>\n> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n> ---\n>  Makefile | 6 ++----\n>  1 file changed, 2 insertions(+), 4 deletions(-)\n\nThis is a declaration that we would never ever build .o files out of\nsources other than .c files.  While it does make sense to have it\noutside the scope of [PATCH 1/2], I am not sure if it even belongs\nto the same series.\n\n"},{"id":"462449","messageId":"220901.867d2njg52.gmgdl@evledraar.gmail.com","threadId":"52786","inReplyTo":"xmqqmtbkdr5n.fsf@gitster.g","subject":"Re: [PATCH v3 2/2] Makefile: use $(OBJECTS) instead of $(C_OBJ)","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-09-01T14:52:05Z","receivedAt":"2022-09-01T14:58:24Z","isPatch":true,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, Aug 31 2022, Junio C Hamano wrote:\n\n> Ævar Arnfjörð Bjarmason  <avarab@gmail.com> writes:\n>\n>> In the preceding commit $(C_OBJ) added in c373991375a (Makefile: list\n>> generated object files in OBJECTS, 2010-01-26) became synonymous with\n>> $(OBJECTS). Let's avoid the indirection and use the $(OBJECTS)\n>> variable directly instead.\n>>\n>> Signed-off-by: Ævar Arnfjörð Bjarmason <avarab@gmail.com>\n>> ---\n>>  Makefile | 6 ++----\n>>  1 file changed, 2 insertions(+), 4 deletions(-)\n>\n> This is a declaration that we would never ever build .o files out of\n> sources other than .c files.  While it does make sense to have it\n> outside the scope of [PATCH 1/2], I am not sure if it even belongs\n> to the same series.\n\nI think it does. Before this the C_OBJ would be:\n\n\tC_OBJ := $(filter-out $(ASM_OBJ),$(OBJECTS))\n\nbut after 1/2 it's the same as $(OBJECTS). An earlier iteration of this\ndid this cleanup \"while we're at it\" (which I do think makes sense as an\natomic change), but I got the feedback that the cleanup wasn't strictly\nnecessary.\n\nBut as 1/2 has removed the ability to build those $(ASM_OBJ), as we had\nonly one of those, I don't think keeping this particular bit of\nindirection makes sense.\n\nOf course it doesn't really matter at all, the real change is the\nremoval of $(ASM_OBJ).\n\nIf we do start building *.o files out of *.S files (or other non-*.c)\nagain we'll need new rules anyway. I think we should just add any such\nvariables back then, and not keep this small bit of dead husk around.\n"},{"id":"462458","messageId":"xmqqsflbayeh.fsf@gitster.g","threadId":"52786","inReplyTo":"220901.867d2njg52.gmgdl@evledraar.gmail.com","subject":"Re: [PATCH v3 2/2] Makefile: use $(OBJECTS) instead of $(C_OBJ)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-09-01T15:48:38Z","receivedAt":"2022-09-01T15:48:53Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n\n>> This is a declaration that we would never ever build .o files out of\n>> sources other than .c files.  While it does make sense to have it\n>> outside the scope of [PATCH 1/2], I am not sure if it even belongs\n>> to the same series.\n>\n> I think it does. Before this the C_OBJ would be:\n>\n> \tC_OBJ := $(filter-out $(ASM_OBJ),$(OBJECTS))\n>\n> but after 1/2 it's the same as $(OBJECTS). An earlier iteration of this\n> did this cleanup \"while we're at it\" (which I do think makes sense as an\n> atomic change), but I got the feedback that the cleanup wasn't strictly\n> necessary.\n>\n> But as 1/2 has removed the ability to build those $(ASM_OBJ), as we had\n> only one of those, I don't think keeping this particular bit of\n> indirection makes sense.\n\nYou are not thinking for longer term to help project maintenance.\n\nThis change removes distinction between C_OBJ and OBJECTS, only\nbecause the sources to the objects we HAPPEN TO have are only C\nfiles.  It is premature and short sighted to declare that it has to\nstay that way forever.  And such a declaration is not something we\nwould casually make \"while at it\" in a topic like this.\n\nWhen we add a source written in another language, say xyzzy, to be\ncompiled into an object file, we'd add $(XYZZY_OBJ), and they will\nbecome part of $(OBJECTS), but the current rule to create $(C_OBJ)\nwill not apply to $(XYZZY_OBJ).  But you do this:\n\n    -$(C_OBJ): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n    +$(OBJECTS): %.o: %.c GIT-CFLAGS $(missing_dep_dirs) $(missing_compdb_dir)\n            $(QUIET_CC)$(CC) -o $*.o -c ... $(ALL_CFLAGS) $(EXTRA_CPPFLAGS) $<\n\nRight now, we know where this patch affected the build procedure,\nbecause the patch highlights what is being changed.  But when future\ndevelopers need to produce some files that belong to $(OBJECTS) out\nof source files that are not .c, they first need to locate the above\nhunk and revert it.  I do not see the benefit of being hostile to\nfuture developers with this patch.  Not before we know that it is\nnot likely that we would add any non-C sources in the future, by\nrunning with 1/2 alone for a year or two.\n"}]}