{"thread":{"id":"45255","subject":"[PATCH] Put sha1dc on a diet","startedAt":"2017-03-01T00:38:46Z","lastAt":"2017-03-16T22:07:51Z","messageCount":36,"participants":["Linus Torvalds","Junio C Hamano","Jeff King","Johannes Schindelin","Dan Shumow","Duy Nguyen","Jeff Hostetler","Marc Stevens"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"312932","messageId":"alpine.LFD.2.20.1702281621050.22202@i7.lan","threadId":"45255","inReplyTo":null,"subject":"[PATCH] Put sha1dc on a diet","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-01T00:30:26Z","receivedAt":"2017-03-01T00:38:46Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\nFrom: Linus Torvalds <torvalds@linux-foundation.org>\nDate: Tue, 28 Feb 2017 16:12:32 -0800\nSubject: [PATCH] Put sha1dc on a diet\n\nThis removes the unnecessary parts of the sha1dc code, shrinking things from\n\n\t[torvalds@i7 git]$ size sha1dc/*.o\n\t   text\t   data\t    bss\t    dec\t    hex\tfilename\n\t 277559\t    640\t      0\t 278199\t  43eb7\tsha1dc/sha1.o\n\t   4438\t  11352\t      0\t  15790\t   3dae\tsha1dc/ubc_check.o\n\nto\n\n\t[torvalds@i7 git]$ size sha1dc/*.o\n\t   text\t   data\t    bss\t    dec\t    hex\tfilename\n\t  13287\t      0\t      0\t  13287\t   33e7\tsha1dc/sha1.o\n\t   4438\t  11352\t      0\t  15790\t   3dae\tsha1dc/ubc_check.o\n\nso the sha1.o text size shrinks from about 271kB to about 13kB.\n\nThe shrinking comes mainly from only generating the recompressio functions \nfor the two rounds that are actually used (58 and 65), but also from \nremoving a couple of other unused functions. The sha1dc library lost its \n\"safe_hash\" parameter to do that, since we check - and refuse to touch - \nthe colliding cases manually.\n\nThe git binary itself is about 2MB of text on my system. For other helper \nbinaries the size reduction is even more noticeable.  A quarter MB here \nand a quarter MB there, and suddenly you have a big binary ;)\n\nThis has been tested with the bad pdf image:\n\n\t[torvalds@i7 git]$ ./t/helper/test-sha1 < ~/Downloads/bad.pdf \n\tfatal: The SHA1 computation detected evidence of a collision attack;\n\trefusing to process the contents.\n\nalthough we obviously still don't have an actual git object to test with.\n\nThe normal git test-suite obviously also passes.\n\nSigned-off-by: Linus Torvalds <torvalds@linux-foundation.org>\n---\n\nI notice that sha1dc is in the 'pu' branch now, so let's put my money \nwhere my mouth is, and send in the sha1dc diet patch.\n\n sha1dc/sha1.c | 356 +++-------------------------------------------------------\n sha1dc/sha1.h |  24 ----\n 2 files changed, 18 insertions(+), 362 deletions(-)\n\ndiff --git a/sha1dc/sha1.c b/sha1dc/sha1.c\nindex 6569b403e..4910f0c35 100644\n--- a/sha1dc/sha1.c\n+++ b/sha1dc/sha1.c\n@@ -9,6 +9,15 @@\n #include \"sha1dc/sha1.h\"\n #include \"sha1dc/ubc_check.h\"\n \n+// function type for sha1_recompression_step_T (uint32_t ihvin[5], uint32_t ihvout[5], const uint32_t me2[80], const uint32_t state[5])\n+// where 0 <= T < 80\n+//       me2 is an expanded message (the expansion of an original message block XOR'ed with a disturbance vector's message block difference)\n+//       state is the internal state (a,b,c,d,e) before step T of the SHA-1 compression function while processing the original message block\n+// the function will return:\n+//       ihvin: the reconstructed input chaining value\n+//       ihvout: the reconstructed output chaining value\n+typedef void(*sha1_recompression_type)(uint32_t*, uint32_t*, const uint32_t*, const uint32_t*);\n+\n #define rotate_right(x,n) (((x)>>(n))|((x)<<(32-(n))))\n #define rotate_left(x,n)  (((x)<<(n))|((x)>>(32-(n))))\n \n@@ -39,212 +48,14 @@\n \n \n \n-void sha1_message_expansion(uint32_t W[80])\n+static void sha1_message_expansion(uint32_t W[80])\n {\n \tunsigned i;\n \tfor (i = 16; i < 80; ++i)\n \t\tW[i] = rotate_left(W[i - 3] ^ W[i - 8] ^ W[i - 14] ^ W[i - 16], 1);\n }\n \n-void sha1_compression(uint32_t ihv[5], const uint32_t m[16])\n-{\n-\tuint32_t W[80];\n-\tuint32_t a, b, c, d, e;\n-\tunsigned i;\n-\n-\tmemcpy(W, m, 16 * 4);\n-\tfor (i = 16; i < 80; ++i)\n-\t\tW[i] = rotate_left(W[i - 3] ^ W[i - 8] ^ W[i - 14] ^ W[i - 16], 1);\n-\n-\ta = ihv[0];\n-\tb = ihv[1];\n-\tc = ihv[2];\n-\td = ihv[3];\n-\te = ihv[4];\n-\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 0);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 1);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 2);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 3);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 4);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 5);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 6);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 7);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 8);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 9);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 10);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 11);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 12);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 13);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 14);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 15);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 16);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 17);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 18);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 19);\n-\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 20);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 21);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 22);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 23);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 24);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 25);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 26);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 27);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 28);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 29);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 30);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 31);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 32);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 33);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 34);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 35);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 36);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 37);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 38);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 39);\n-\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 40);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 41);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 42);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 43);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 44);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 45);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 46);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 47);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 48);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 49);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 50);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 51);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 52);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 53);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 54);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 55);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 56);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 57);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 58);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 59);\n-\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 60);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 61);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 62);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 63);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 64);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 65);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 66);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 67);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 68);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 69);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 70);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 71);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 72);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 73);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 74);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 75);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 76);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 77);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 78);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 79);\n-\n-\tihv[0] += a; ihv[1] += b; ihv[2] += c; ihv[3] += d; ihv[4] += e;\n-}\n-\n-\n-\n-void sha1_compression_W(uint32_t ihv[5], const uint32_t W[80])\n-{\n-\tuint32_t a = ihv[0], b = ihv[1], c = ihv[2], d = ihv[3], e = ihv[4];\n-\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 0);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 1);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 2);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 3);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 4);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 5);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 6);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 7);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 8);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 9);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 10);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 11);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 12);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 13);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 14);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 15);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 16);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 17);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 18);\n-\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 19);\n-\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 20);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 21);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 22);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 23);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 24);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 25);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 26);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 27);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 28);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 29);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 30);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 31);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 32);\n- \tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 33);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 34);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 35);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 36);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 37);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 38);\n-\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 39);\n-\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 40);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 41);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 42);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 43);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 44);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 45);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 46);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 47);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 48);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 49);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 50);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 51);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 52);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 53);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 54);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 55);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 56);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 57);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 58);\n-\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 59);\n-\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 60);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 61);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 62);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 63);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 64);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 65);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 66);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 67);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 68);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 69);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 70);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 71);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 72);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 73);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 74);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 75);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 76);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 77);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 78);\n-\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 79);\n-\n-\tihv[0] += a; ihv[1] += b; ihv[2] += c; ihv[3] += d; ihv[4] += e;\n-}\n-\n-\n-\n-void sha1_compression_states(uint32_t ihv[5], const uint32_t W[80], uint32_t states[80][5])\n+static void sha1_compression_states(uint32_t ihv[5], const uint32_t W[80], uint32_t states[80][5])\n {\n \tuint32_t a = ihv[0], b = ihv[1], c = ihv[2], d = ihv[3], e = ihv[4];\n \n@@ -664,7 +475,7 @@ void sha1_compression_states(uint32_t ihv[5], const uint32_t W[80], uint32_t sta\n \n \n #define SHA1_RECOMPRESS(t) \\\n-void sha1recompress_fast_ ## t (uint32_t ihvin[5], uint32_t ihvout[5], const uint32_t me2[80], const uint32_t state[5]) \\\n+static void sha1recompress_fast_ ## t (uint32_t ihvin[5], uint32_t ihvout[5], const uint32_t me2[80], const uint32_t state[5]) \\\n { \\\n \tuint32_t a = state[0], b = state[1], c = state[2], d = state[3], e = state[4]; \\\n \tif (t > 79) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(b, c, d, e, a, me2, 79); \\\n@@ -832,111 +643,20 @@ void sha1recompress_fast_ ## t (uint32_t ihvin[5], uint32_t ihvout[5], const uin\n \tihvout[0] = ihvin[0] + a; ihvout[1] = ihvin[1] + b; ihvout[2] = ihvin[2] + c; ihvout[3] = ihvin[3] + d; ihvout[4] = ihvin[4] + e; \\\n } \n \n-SHA1_RECOMPRESS(0)\n-SHA1_RECOMPRESS(1)\n-SHA1_RECOMPRESS(2)\n-SHA1_RECOMPRESS(3)\n-SHA1_RECOMPRESS(4)\n-SHA1_RECOMPRESS(5)\n-SHA1_RECOMPRESS(6)\n-SHA1_RECOMPRESS(7)\n-SHA1_RECOMPRESS(8)\n-SHA1_RECOMPRESS(9)\n-\n-SHA1_RECOMPRESS(10)\n-SHA1_RECOMPRESS(11)\n-SHA1_RECOMPRESS(12)\n-SHA1_RECOMPRESS(13)\n-SHA1_RECOMPRESS(14)\n-SHA1_RECOMPRESS(15)\n-SHA1_RECOMPRESS(16)\n-SHA1_RECOMPRESS(17)\n-SHA1_RECOMPRESS(18)\n-SHA1_RECOMPRESS(19)\n-\n-SHA1_RECOMPRESS(20)\n-SHA1_RECOMPRESS(21)\n-SHA1_RECOMPRESS(22)\n-SHA1_RECOMPRESS(23)\n-SHA1_RECOMPRESS(24)\n-SHA1_RECOMPRESS(25)\n-SHA1_RECOMPRESS(26)\n-SHA1_RECOMPRESS(27)\n-SHA1_RECOMPRESS(28)\n-SHA1_RECOMPRESS(29)\n-\n-SHA1_RECOMPRESS(30)\n-SHA1_RECOMPRESS(31)\n-SHA1_RECOMPRESS(32)\n-SHA1_RECOMPRESS(33)\n-SHA1_RECOMPRESS(34)\n-SHA1_RECOMPRESS(35)\n-SHA1_RECOMPRESS(36)\n-SHA1_RECOMPRESS(37)\n-SHA1_RECOMPRESS(38)\n-SHA1_RECOMPRESS(39)\n-\n-SHA1_RECOMPRESS(40)\n-SHA1_RECOMPRESS(41)\n-SHA1_RECOMPRESS(42)\n-SHA1_RECOMPRESS(43)\n-SHA1_RECOMPRESS(44)\n-SHA1_RECOMPRESS(45)\n-SHA1_RECOMPRESS(46)\n-SHA1_RECOMPRESS(47)\n-SHA1_RECOMPRESS(48)\n-SHA1_RECOMPRESS(49)\n-\n-SHA1_RECOMPRESS(50)\n-SHA1_RECOMPRESS(51)\n-SHA1_RECOMPRESS(52)\n-SHA1_RECOMPRESS(53)\n-SHA1_RECOMPRESS(54)\n-SHA1_RECOMPRESS(55)\n-SHA1_RECOMPRESS(56)\n-SHA1_RECOMPRESS(57)\n SHA1_RECOMPRESS(58)\n-SHA1_RECOMPRESS(59)\n-\n-SHA1_RECOMPRESS(60)\n-SHA1_RECOMPRESS(61)\n-SHA1_RECOMPRESS(62)\n-SHA1_RECOMPRESS(63)\n-SHA1_RECOMPRESS(64)\n SHA1_RECOMPRESS(65)\n-SHA1_RECOMPRESS(66)\n-SHA1_RECOMPRESS(67)\n-SHA1_RECOMPRESS(68)\n-SHA1_RECOMPRESS(69)\n-\n-SHA1_RECOMPRESS(70)\n-SHA1_RECOMPRESS(71)\n-SHA1_RECOMPRESS(72)\n-SHA1_RECOMPRESS(73)\n-SHA1_RECOMPRESS(74)\n-SHA1_RECOMPRESS(75)\n-SHA1_RECOMPRESS(76)\n-SHA1_RECOMPRESS(77)\n-SHA1_RECOMPRESS(78)\n-SHA1_RECOMPRESS(79)\n-\n-sha1_recompression_type sha1_recompression_step[80] =\n+\n+static sha1_recompression_type sha1_recompression_step[80] =\n {\n-\tsha1recompress_fast_0, sha1recompress_fast_1, sha1recompress_fast_2, sha1recompress_fast_3, sha1recompress_fast_4, sha1recompress_fast_5, sha1recompress_fast_6, sha1recompress_fast_7, sha1recompress_fast_8, sha1recompress_fast_9,\n-\tsha1recompress_fast_10, sha1recompress_fast_11, sha1recompress_fast_12, sha1recompress_fast_13, sha1recompress_fast_14, sha1recompress_fast_15, sha1recompress_fast_16, sha1recompress_fast_17, sha1recompress_fast_18, sha1recompress_fast_19,\n-\tsha1recompress_fast_20, sha1recompress_fast_21, sha1recompress_fast_22, sha1recompress_fast_23, sha1recompress_fast_24, sha1recompress_fast_25, sha1recompress_fast_26, sha1recompress_fast_27, sha1recompress_fast_28, sha1recompress_fast_29,\n-\tsha1recompress_fast_30, sha1recompress_fast_31, sha1recompress_fast_32, sha1recompress_fast_33, sha1recompress_fast_34, sha1recompress_fast_35, sha1recompress_fast_36, sha1recompress_fast_37, sha1recompress_fast_38, sha1recompress_fast_39,\n-\tsha1recompress_fast_40, sha1recompress_fast_41, sha1recompress_fast_42, sha1recompress_fast_43, sha1recompress_fast_44, sha1recompress_fast_45, sha1recompress_fast_46, sha1recompress_fast_47, sha1recompress_fast_48, sha1recompress_fast_49,\n-\tsha1recompress_fast_50, sha1recompress_fast_51, sha1recompress_fast_52, sha1recompress_fast_53, sha1recompress_fast_54, sha1recompress_fast_55, sha1recompress_fast_56, sha1recompress_fast_57, sha1recompress_fast_58, sha1recompress_fast_59,\n-\tsha1recompress_fast_60, sha1recompress_fast_61, sha1recompress_fast_62, sha1recompress_fast_63, sha1recompress_fast_64, sha1recompress_fast_65, sha1recompress_fast_66, sha1recompress_fast_67, sha1recompress_fast_68, sha1recompress_fast_69,\n-\tsha1recompress_fast_70, sha1recompress_fast_71, sha1recompress_fast_72, sha1recompress_fast_73, sha1recompress_fast_74, sha1recompress_fast_75, sha1recompress_fast_76, sha1recompress_fast_77, sha1recompress_fast_78, sha1recompress_fast_79,\n+\t[58] = sha1recompress_fast_58,\n+\t[65] = sha1recompress_fast_65,\n };\n \n \n \n \n \n-void sha1_process(SHA1_CTX* ctx, const uint32_t block[16]) \n+static void sha1_process(SHA1_CTX* ctx, const uint32_t block[16])\n {\n \tunsigned i, j;\n \tuint32_t ubc_dv_mask[DVMASKSIZE];\n@@ -973,12 +693,6 @@ void sha1_process(SHA1_CTX* ctx, const uint32_t block[16])\n \t\t\t\t\tif (ctx->callback != NULL)\n \t\t\t\t\t\tctx->callback(ctx->total - 64, ctx->ihv1, ctx->ihv2, ctx->m1, ctx->m2);\n \n-\t\t\t\t\tif (ctx->safe_hash) \n-\t\t\t\t\t{\n-\t\t\t\t\t\tsha1_compression_W(ctx->ihv, ctx->m1);\n-\t\t\t\t\t\tsha1_compression_W(ctx->ihv, ctx->m1);\n-\t\t\t\t\t}\n-\n \t\t\t\t\tbreak;\n \t\t\t\t}\n \t\t\t}\n@@ -990,7 +704,7 @@ void sha1_process(SHA1_CTX* ctx, const uint32_t block[16])\n \n \n \n-void swap_bytes(uint32_t val[16]) \n+static void swap_bytes(uint32_t val[16])\n {\n \tunsigned i;\n \tfor (i = 0; i < 16; ++i) \n@@ -1011,7 +725,6 @@ void SHA1DCInit(SHA1_CTX* ctx)\n \tctx->ihv[3] = 0x10325476;\n \tctx->ihv[4] = 0xC3D2E1F0;\n \tctx->found_collision = 0;\n-\tctx->safe_hash = 1;\n \tctx->ubc_check = 1;\n \tctx->detect_coll = 1;\n \tctx->reduced_round_coll = 0;\n@@ -1019,39 +732,6 @@ void SHA1DCInit(SHA1_CTX* ctx)\n \tctx->callback = NULL;\n }\n \n-void SHA1DCSetSafeHash(SHA1_CTX* ctx, int safehash)\n-{\n-\tif (safehash)\n-\t\tctx->safe_hash = 1;\n-\telse\n-\t\tctx->safe_hash = 0;\n-}\n-\n-\n-void SHA1DCSetUseUBC(SHA1_CTX* ctx, int ubc_check)\n-{\n-\tif (ubc_check)\n-\t\tctx->ubc_check = 1;\n-\telse\n-\t\tctx->ubc_check = 0;\n-}\n-\n-void SHA1DCSetUseDetectColl(SHA1_CTX* ctx, int detect_coll)\n-{\n-\tif (detect_coll)\n-\t\tctx->detect_coll = 1;\n-\telse\n-\t\tctx->detect_coll = 0;\n-}\n-\n-void SHA1DCSetDetectReducedRoundCollision(SHA1_CTX* ctx, int reduced_round_coll)\n-{\n-\tif (reduced_round_coll)\n-\t\tctx->reduced_round_coll = 1;\n-\telse\n-\t\tctx->reduced_round_coll = 0;\n-}\n-\n void SHA1DCSetCallback(SHA1_CTX* ctx, collision_block_callback callback)\n {\n \tctx->callback = callback;\ndiff --git a/sha1dc/sha1.h b/sha1dc/sha1.h\nindex 1bb0ace99..f126bf63b 100644\n--- a/sha1dc/sha1.h\n+++ b/sha1dc/sha1.h\n@@ -5,29 +5,6 @@\n * https://opensource.org/licenses/MIT\n ***/\n \n-// uses SHA-1 message expansion to expand the first 16 words of W[] to 80 words\n-void sha1_message_expansion(uint32_t W[80]);\n-\n-// sha-1 compression function; first version takes a message block pre-parsed as 16 32-bit integers, second version takes an already expanded message)\n-void sha1_compression(uint32_t ihv[5], const uint32_t m[16]);\n-void sha1_compression_W(uint32_t ihv[5], const uint32_t W[80]);\n-\n-// same as sha1_compression_W, but additionally store intermediate states\n-// only stores states ii (the state between step ii-1 and step ii) when DOSTORESTATEii is defined in ubc_check.h\n-void sha1_compression_states(uint32_t ihv[5], const uint32_t W[80], uint32_t states[80][5]);\n-\n-// function type for sha1_recompression_step_T (uint32_t ihvin[5], uint32_t ihvout[5], const uint32_t me2[80], const uint32_t state[5])\n-// where 0 <= T < 80\n-//       me2 is an expanded message (the expansion of an original message block XOR'ed with a disturbance vector's message block difference)\n-//       state is the internal state (a,b,c,d,e) before step T of the SHA-1 compression function while processing the original message block\n-// the function will return:\n-//       ihvin: the reconstructed input chaining value\n-//       ihvout: the reconstructed output chaining value\n-typedef void(*sha1_recompression_type)(uint32_t*, uint32_t*, const uint32_t*, const uint32_t*);\n-\n-// table of sha1_recompression_step_0, ... , sha1_recompression_step_79\n-extern sha1_recompression_type sha1_recompression_step[80];\n-\n // a callback function type that can be set to be called when a collision block has been found:\n // void collision_block_callback(uint64_t byteoffset, const uint32_t ihvin1[5], const uint32_t ihvin2[5], const uint32_t m1[80], const uint32_t m2[80])\n typedef void(*collision_block_callback)(uint64_t, const uint32_t*, const uint32_t*, const uint32_t*, const uint32_t*);\n@@ -39,7 +16,6 @@ typedef struct {\n \tunsigned char buffer[64];\n \tint bigendian;\n \tint found_collision;\n-\tint safe_hash;\n \tint detect_coll;\n \tint ubc_check;\n \tint reduced_round_coll;\n-- \n2.12.0.4.g94589516d\n\n"},{"id":"312962","messageId":"xmqq7f48hm8g.fsf@gitster.mtv.corp.google.com","threadId":"45255","inReplyTo":"alpine.LFD.2.20.1702281621050.22202@i7.lan","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-03-01T18:42:55Z","receivedAt":"2017-03-01T18:44:01Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> I notice that sha1dc is in the 'pu' branch now, so let's put my money \n> where my mouth is, and send in the sha1dc diet patch.\n\nI see //c99 comments and also T array[] = { [58] = val } both of\nwhich I think we stay away from (and the former is from the initial\nimport), so some people on other platforms MAY have trouble with\nthis topic.\n\nLet's see what happens by queuing it on 'pu' ;-)\n\nThanks.\n"},{"id":"312963","messageId":"CA+55aFx1wAS-nHS2awuW2waX=cvig4UoZqmN5H3v93yDE7ukyQ@mail.gmail.com","threadId":"45255","inReplyTo":"xmqq7f48hm8g.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-01T18:49:55Z","receivedAt":"2017-03-01T18:57:25Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Wed, Mar 1, 2017 at 10:42 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> I see //c99 comments\n\nsha1dc is already full of // style comments. I just followed the\nexisting practice.\n\n> and also T array[] = { [58] = val } both of\n> which I think we stay away from (and the former is from the initial\n> import), so some people on other platforms MAY have trouble with\n> this topic.\n\nHmm. The \"{ [58] = val; }\" kind of initialization would be easy to\nwork around by just filling in everything else with NULL, but it would\nmake for a pretty nasty readability issue.\n\nThat said, if you mis-count the NULL's, the end result will pretty\nimmediately SIGSEGV, so I guess it wouldn't be much of a maintenance\nproblem.\n\nBut if you're just willing to take the \"let's see\" approach, I think\nthe explicitly numbered initializer is much better.\n\nThe main people who I assume would really want to use the sha1dc\nlibrary are hosting places. And they won't be using crazy compilers\nfrom the last century.\n\nThat said, I think that it would be lovely to just default to\nUSE_SHA1DC and just put the whole attack behind us. Yes, it's slower.\nNo, it doesn't really seem to matter that much in practice.\n\n              Linus\n"},{"id":"312965","messageId":"20170301190711.nwlretmy6w675jvp@sigill.intra.peff.net","threadId":"45255","inReplyTo":"alpine.LFD.2.20.1702281621050.22202@i7.lan","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-01T19:07:11Z","receivedAt":"2017-03-01T19:15:23Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 28, 2017 at 04:30:26PM -0800, Linus Torvalds wrote:\n\n> From: Linus Torvalds <torvalds@linux-foundation.org>\n> Date: Tue, 28 Feb 2017 16:12:32 -0800\n> Subject: [PATCH] Put sha1dc on a diet\n> \n> This removes the unnecessary parts of the sha1dc code, shrinking things from\n> [...]\n\nSo obviously the smaller object size is nice, and the diffstat is\ncertainly satisfying. My only qualm would be whether this conflicts with\nthe optimizations that Dan is working on (probably not conceptually, but\ntextually).\n\n-Peff\n"},{"id":"312970","messageId":"20170301195302.3pybakmjqztosohj@sigill.intra.peff.net","threadId":"45255","inReplyTo":"CA+55aFx1wAS-nHS2awuW2waX=cvig4UoZqmN5H3v93yDE7ukyQ@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-01T19:53:02Z","receivedAt":"2017-03-01T20:02:52Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 01, 2017 at 10:49:55AM -0800, Linus Torvalds wrote:\n\n> That said, I think that it would be lovely to just default to\n> USE_SHA1DC and just put the whole attack behind us. Yes, it's slower.\n> No, it doesn't really seem to matter that much in practice.\n\nMy biggest concern is the index-pack operation. Try this:\n\n  time git clone --no-local --bare linux tmp.git\n\nwith and without USE_SHA1DC. I get:\n\n  [w/ openssl]\n  real\t1m52.307s\n  user\t2m47.928s\n  sys\t0m14.992s\n\n  [w/ sha1dc]\n  real\t3m4.043s\n  user\t6m16.412s\n  sys\t0m13.772s\n\nThat's real latency the user will see. It's hard to break it down,\nthough. The actual \"receiving\" phase is generally going to be network\nbound. The delta-resolution that happens afterwards is totally local and\nCPU-bound (but does run in parallel).\n\nAnd of course this repository tends to the larger side (though certainly\nthere are bigger ones), and you only feel the pain on clone or when\ndoing an initial push, not day-to-day.\n\nSo maybe we just suck it up and accept that it's a bit slower.\n\n-Peff\n"},{"id":"312971","messageId":"xmqq37ewhji1.fsf@gitster.mtv.corp.google.com","threadId":"45255","inReplyTo":"CA+55aFx1wAS-nHS2awuW2waX=cvig4UoZqmN5H3v93yDE7ukyQ@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-03-01T19:41:58Z","receivedAt":"2017-03-01T20:06:34Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> That said, I think that it would be lovely to just default to\n> USE_SHA1DC and just put the whole attack behind us. Yes, it's slower.\n> No, it doesn't really seem to matter that much in practice.\n\nYes.  It would be a very good goal.\n\n"},{"id":"312982","messageId":"20170301203427.e5xa5ej3czli7c3o@sigill.intra.peff.net","threadId":"45255","inReplyTo":"CA+55aFwf3sxKW+dGTMjNAeHMOf=rvctEQohm+rbhEb=e3KLpHw@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-01T20:34:27Z","receivedAt":"2017-03-01T21:35:31Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 01, 2017 at 12:14:34PM -0800, Linus Torvalds wrote:\n\n> > My biggest concern is the index-pack operation. Try this:\n> \n> I'm mobile right now, so I can't test, but I'd this perhaps at least partly\n> due to the full checksum over the pack-file?\n>\n> We have two very different uses of SHA1: the actual object name hash, but\n> also the sha1file checksums that we do on the index file and the pack files.\n>\n> And the checksum code really doesn't need the collision checking at all.\n\nI don't think that helps. The sha1 over the pack-file takes about 1.3s\nwith openssl, and 5s with sha1dc. So we already know the increase there\nis only a few seconds, not a few minutes.\n\nAnd it makes sense if you think about the index-pack operation. It has\nto inflate each object, resolving deltas, and checksum the result. And\nthe number of inflated bytes is _much_ larger than the on-disk bytes.\nYou can see the difference with:\n\n  git cat-file --batch-all-objects \\\n    --batch-check='%(objectsize:disk) %(objectsize)' |\n  perl -alne '\n    $disk += $F[0]; $raw += $F[1];\n    END { print \"$disk $raw\" }\n  '\n\nOn linux.git that yields:\n\n  1210521959 63279680406\n\nThat's over a 50x increase in the bytes we have to sha1 for objects\nversus pack-checksums.\n\n-Peff\n"},{"id":"312986","messageId":"CAPc5daU2GVpc5Nhx13apFMs3XkL+O8_+3uA842vGQouwb4kEAg@mail.gmail.com","threadId":"45255","inReplyTo":"alpine.DEB.2.20.1703012227010.3767@virtualbox","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-03-01T22:05:20Z","receivedAt":"2017-03-01T22:06:13Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"On Wed, Mar 1, 2017 at 1:56 PM, Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n> Hi,\n>\n> On Wed, 1 Mar 2017, Junio C Hamano wrote:\n>\n>> Linus Torvalds <torvalds@linux-foundation.org> writes:\n>>\n>> > That said, I think that it would be lovely to just default to\n>> > USE_SHA1DC and just put the whole attack behind us. Yes, it's slower.\n>> > No, it doesn't really seem to matter that much in practice.\n>>\n>> Yes.  It would be a very good goal.\n>\n> So let me get this straight: not only do we now implicitly want to bump\n> the required C compiler to C99 without any grace period worth mentioning\n> [*1*], we are also all of a sudden no longer worried about a double digit\n> percentage drop of speed [*2*]?\n\nBefore we get the code into shape suitable for 'next', it is more important to\nmake sure it operates correctly, adding necessary features if any (e.g. \"hash\nwith or without check\" knob) while it is in 'pu', and *1* is to allow\nit to progress\nfaster without having to worry about something we could do mechanically\nbefore making it ready for 'next'.\n\nThe performance thing is really \"let's see how well it goes\". With effort to\noptimize still \"just has began\", I think it is too early to tell if\nLinus's \"doesn't\nreally seem to matter\" is the case or not.\n\nQueuing such a topic on 'pu' is one effective way to make sure people are\nworking off of the same codebase.\n"},{"id":"312987","messageId":"alpine.DEB.2.20.1703012227010.3767@virtualbox","threadId":"45255","inReplyTo":"xmqq37ewhji1.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-03-01T21:56:18Z","receivedAt":"2017-03-01T22:06:15Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 1 Mar 2017, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > That said, I think that it would be lovely to just default to\n> > USE_SHA1DC and just put the whole attack behind us. Yes, it's slower.\n> > No, it doesn't really seem to matter that much in practice.\n> \n> Yes.  It would be a very good goal.\n\nSo let me get this straight: not only do we now implicitly want to bump\nthe required C compiler to C99 without any grace period worth mentioning\n[*1*], we are also all of a sudden no longer worried about a double digit\npercentage drop of speed [*2*]?\n\nPuzzled,\nJohannes\n\nFootnote *1*: I know, it is easy to forget that some developers cannot\nchoose their tools, or even their hardware. In the past, we seemed to take\nappropriate care, though.\n\nFootnote *2*: With real-world repositories of notable size, that\nperformance regression hurts. A lot. We just spent time to get the speed\nof SHA-1 down by a couple percent and it was a noticeable improvement here.\n"},{"id":"312989","messageId":"CA+55aFys5oQ0RySQ+Xv0ZDussr-xZNh4_b3+Upx_d9VPWmpM8Q@mail.gmail.com","threadId":"45255","inReplyTo":"alpine.DEB.2.20.1703012227010.3767@virtualbox","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-01T22:16:05Z","receivedAt":"2017-03-01T22:16:40Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Wed, Mar 1, 2017 at 1:56 PM, Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n> Footnote *1*: I know, it is easy to forget that some developers cannot\n> choose their tools, or even their hardware. In the past, we seemed to take\n> appropriate care, though.\n\nI don't think you need to worry about the Windows side. That can\ncontinue to do something else.\n\nWhen I advocated perhaps using  USE_SHA1DC by default, I definitely\ndid not mean it in a \"everywhere, regardless of issues\" manner.\n\nFor example, the conmfig.mak.uname script already explicitly asks for\n\"BLK_SHA1 = YesPlease\" for Windows. Don't bother changing that, it's\nan explicit choice.\n\nBut the Linux rules don't actually specify which SHA1 version to use,\nso the main Makefile currently defaults to just using openssl.\n\nSo that's the \"default\" choice I think we might want to change. Not\nthe \"we're windows, and explicitly want BLK_SHA1 because of\nenvironment and build infrastructure\".\n\n             Linus\n"},{"id":"312997","messageId":"alpine.DEB.2.20.1703012334400.3767@virtualbox","threadId":"45255","inReplyTo":"CA+55aFys5oQ0RySQ+Xv0ZDussr-xZNh4_b3+Upx_d9VPWmpM8Q@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-03-01T22:51:35Z","receivedAt":"2017-03-01T22:59:05Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 1 Mar 2017, Linus Torvalds wrote:\n\n> On Wed, Mar 1, 2017 at 1:56 PM, Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n> > Footnote *1*: I know, it is easy to forget that some developers cannot\n> > choose their tools, or even their hardware. In the past, we seemed to take\n> > appropriate care, though.\n> \n> I don't think you need to worry about the Windows side.\n\nI am not. I build G?t for Windows using GCC.\n\nMy concern is about that unexpected turn \"oh, let's just switch to C99\nbecause, well, because my compiler canehandle it, and everybody else\nshould just switch tn a modern compiler\". That really sounded careless.\n\n> That can continue to do something else.\n> \n> When I advocated perhaps using  USE_SHA1DC by default, I definitely did\n> not mean it in a \"everywhere, regardless of issues\" manner.\n> \n> For example, the conmfig.mak.uname script already explicitly asks for\n> \"BLK_SHA1 = YesPlease\" for Windows. Don't bother changing that, it's an\n> explicit choice.\n\nThat setting is only in git.git's version, not in gxt-for-windows/git.git.\nWe switched to OpenSSL because of speed improvements, in particular with\nrecent Intel processors.\n\n> But the Linux rules don't actually specify which SHA1 version to use,\n> so the main Makefile currently defaults to just using openssl.\n> \n> So that's the \"default\" choice I think we might want to change. Not\n> the \"we're windows, and explicitly want BLK_SHA1 because of\n> environment and build infrastructure\".\n\nSince we switched away from BLOCK_SHA1, any such change would affect Git\nfor Windews.\n\nBut I think bigger than just developers on Windows OS. There are many\ndevelopers out there working on large repositories (yes, much larger than\nLinux). Also using Macs and Linux. I am not at all sure that we want to\ngive them an updated Git they cannot fail to notice to be much slower than\nbefore.\n\nDon't get me wrong: I *hope* that you'll manage to get sha1dc\ncompetitively fast. If you don't, well, then we simply cannot use it by\ndefault for *all* of our calls (you already pointed out that the pack\nindex' checksum does not need collision detection, and in fact, *any*\noperation that works on implicitly trusted data falls into the same court,\ne.g. `git add`).\n\nCiao,\nJohannes\n"},{"id":"313001","messageId":"CA+55aFy9=jBJT36FC2HiAeabJBssY=jE=zLxwrXWzhpiFkMUXg@mail.gmail.com","threadId":"45255","inReplyTo":"alpine.DEB.2.20.1703012334400.3767@virtualbox","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-01T23:05:25Z","receivedAt":"2017-03-01T23:13:10Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Wed, Mar 1, 2017 at 2:51 PM, Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n>\n> But I think bigger than just developers on Windows OS. There are many\n> developers out there working on large repositories (yes, much larger than\n> Linux). Also using Macs and Linux. I am not at all sure that we want to\n> give them an updated Git they cannot fail to notice to be much slower than\n> before.\n\nJohannes, have you *tried* the patches?\n\nI really don't think you have. It is completely unnoticeable in any\nnormal situation. The two cases that it's noticeable is:\n\n - a full fsck is noticeable slower\n\n - a full non-local clone is slower (but not really noticeably so\nsince the network traffic dominates).\n\nIn other words, I think you're making shit up. I don't think you\nunderstand how little the SHA1 performance actually matters. It's\nnoticeable in benchmarks. It's not noticeable in any normal operation.\n\n.. and yes, I've actually been running the patches locally since I\nposted my first version (which apparently didn't go out to the list\nbecause of list size limits) and now running the version in 'pu'.\n\n                Linus\n"},{"id":"313003","messageId":"20170301231302.o2tw6vlu2wqeltzj@sigill.intra.peff.net","threadId":"45255","inReplyTo":"CA+55aFwr1jncrk-cekn0Y8rs_S+zs7RrgQ-Jb-ZbgCvmVrHT_A@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-01T23:13:02Z","receivedAt":"2017-03-01T23:13:21Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 01, 2017 at 12:58:39PM -0800, Linus Torvalds wrote:\n\n>> I don't think that helps. The sha1 over the pack-file takes about 1.3s\n>> with openssl, and 5s with sha1dc. So we already know the increase there\n>> is only a few seconds, not a few minutes.\n> \n> OK. I guess what w could easily do is to just add an argument to\n> git_SHA1_Init() to say whether we want checking or not, and only use the\n> checking functions when receiving objects (which would include creating new\n> objects from files, and obviously deck).\n> \n> You'd still eat the cost on the receiving side of a clone, but that's when\n> you really want the checking anyway. At least it wouldn't be so visible on\n> the sending side, which is all the hosting etc, where there might be server\n> utilization issues.\n> \n> Would that make deployment happier? It should be an easy little flag to\n> add, I think.\n\nI don't think it makes all that big a difference. The sending side\nwastes the extra 2-3 seconds of CPU to checksum the whole packfile, but\nit's not inflating all the object contents in the first place (between\nreachability bitmaps to get the list of objects in the first place, and\nthen verbatim reuse of pack contents).\n\nWhich isn't to say it's not reasonable to limit the checking to a few\nspots (basically anything that's _writing_ objects). But I don't think\nit makes a big difference to the server side of a fetch or clone.\n\n-Peff\n"},{"id":"313005","messageId":"20170301231921.2puf7o7jkrujscwn@sigill.intra.peff.net","threadId":"45255","inReplyTo":"CA+55aFy9=jBJT36FC2HiAeabJBssY=jE=zLxwrXWzhpiFkMUXg@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-01T23:19:21Z","receivedAt":"2017-03-01T23:21:00Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Mar 01, 2017 at 03:05:25PM -0800, Linus Torvalds wrote:\n\n> On Wed, Mar 1, 2017 at 2:51 PM, Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n> >\n> > But I think bigger than just developers on Windows OS. There are many\n> > developers out there working on large repositories (yes, much larger than\n> > Linux). Also using Macs and Linux. I am not at all sure that we want to\n> > give them an updated Git they cannot fail to notice to be much slower than\n> > before.\n> \n> Johannes, have you *tried* the patches?\n> \n> I really don't think you have. It is completely unnoticeable in any\n> normal situation. The two cases that it's noticeable is:\n> \n>  - a full fsck is noticeable slower\n> \n>  - a full non-local clone is slower (but not really noticeably so\n> since the network traffic dominates).\n> \n> In other words, I think you're making shit up. I don't think you\n> understand how little the SHA1 performance actually matters. It's\n> noticeable in benchmarks. It's not noticeable in any normal operation.\n> \n> .. and yes, I've actually been running the patches locally since I\n> posted my first version (which apparently didn't go out to the list\n> because of list size limits) and now running the version in 'pu'.\n\nYou have to remember that some of the Git for Windows users are doing\nhorrific things like using repositories with 450MB .git/index files, and\nthe speed to compute the sha1 during an update is noticeable there.\n\nIMHO that is a good sign that the right approach is to switch to an\nindex format that doesn't require rewriting all 450MB to update one\nentry. But obviously that is a much harder problem than just using an\noptimized sha1 implementation.\n\nI do think that could argue for turning on the collision detection only\nduring object-write operations, which is where it matters. It would be\nreally trivial to flip the \"check collisions\" bit on sha1dc. But I\nsuspect you could go faster still by compiling against two separate\nimplementations: the fast-as-possible one (which could be openssl or\nblk-sha1), and the slower-but-careful sha1dc.\n\n-Peff\n"},{"id":"313008","messageId":"CA+55aFz4ixVKVURki8FeXjL5H51A_cQXsZpzKJ-N9n574Yy1rg@mail.gmail.com","threadId":"45255","inReplyTo":"20170301203427.e5xa5ej3czli7c3o@sigill.intra.peff.net","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-01T23:38:40Z","receivedAt":"2017-03-01T23:46:36Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Wed, Mar 1, 2017 at 12:34 PM, Jeff King <peff@peff.net> wrote:\n>\n> I don't think that helps. The sha1 over the pack-file takes about 1.3s\n> with openssl, and 5s with sha1dc. So we already know the increase there\n> is only a few seconds, not a few minutes.\n\nYeah, I did a few statistics by adding just logging of \"SHA1_Init()\"\ncalls. For that network clone situation, the call distribution is\n\n      1        SHA1: Init at builtin/index-pack.c:326\n 841228        SHA1: Init at builtin/index-pack.c:450\n      2        SHA1: Init at csum-file.c:152\n4415756        SHA1: Init at sha1_file.c:3218\n\n(the line numbers are a bit off from 'pu', because I obviously have\nthe logging code).\n\nThe big number (one for every object) is from\nwrite_sha1_file_prepare(), which we'd want to be the strong collision\nchecking version because those are things we're about to create git\nobjects out of. It's called from\n\n - hash_sha1_file() - doesn't actually write the object, but is used\nto calculate the sha for incoming data after applying the delta, for\nexample.\n\n - write_sha1_file() - many uses, actually writes the object\n\n - hash_sha1_file_literally() - git hash-object\n\nand that index-pack.c:450 is from unpack_entry_data() for the base\nnon-delta objects (which should also be the strong kind).\n\nSo all of them should check against collision attacks, so none of them\nseem to be things you'd want to optimize away..\n\nSo I was wrong in thinking that there were a lot of unnecessary SHA1\ncalculations in that load. They all look like they should be done with\nthe slower checking code.\n\nOh well.\n\n                      Linus\n"},{"id":"313030","messageId":"CY1PR0301MB21073D82F4A6AB0DAD8BF1FCC4280@CY1PR0301MB2107.namprd03.prod.outlook.com","threadId":"45255","inReplyTo":"CA+55aFz4ixVKVURki8FeXjL5H51A_cQXsZpzKJ-N9n574Yy1rg@mail.gmail.com","subject":"RE: [PATCH] Put sha1dc on a diet","fromName":"Dan Shumow","fromEmail":"danshu@microsoft.com","sentAt":"2017-03-02T01:31:17Z","receivedAt":"2017-03-02T03:09:53Z","isPatch":true,"sender":{"key":"danshu@microsoft.com","avatar":null},"body":"I played around tweaking the code a bit more and I got our performance down to a 2.077182x slowdown with check and a 1.055961x slowdown without checking.  However, that slowdown is basically with the check turned off through our API.  If I rip extraneous code for storing states and checking if we are doing collision detection out, I can reach performance parity with the block-sha1 implementation in the Git codebase, which basically tells me that is about as good as I can do for optimizing the C code.\n\nSHA1 is more amenable to assembler implementation because its use of rotations, which are notoriously difficult to access through C code.  And as this happens in the inner loop of the function, the inline asm tends to not cut it.  This is one of the reasons that the OpenSSL SHA-1 runs like a scalded monkey, compared to the C implemenations.  Marc and I have also discussed using SIMD operations to speed up the UBC checks, which could definitely help achieve better performance, but is highly dependent on processor support.  It will take some time to do either a SIMD implementation of the UBC checks or an assembler implementation.\n\nAt this point, I would suggest that I take the C optimizations, clean them up and fold them in with the diet changes Linus has suggested.  The slowdown is still 2x over block-sha1 and more over OpenSSL.  But it is better than nothing.  And then if there is interest Marc and I can investigate other processor specific optimizations like ASM or SIMD and circle back with those performance optimizations at a later date.\n\nAlso, to Johannes Schindelin's point:\n> My concern is about that unexpected turn \"oh, let's just switch to C99 because, well, because my compiler canehandle it, and everybody else should just switch tn a modern compiler\". That really sounded careless.\n\nWhile it will probably be a pain, if it is a requirement, we can modify the code to move away from any c99 specific stuff we have in here, if it makes adopting the code more palatable for Git.\n\nThanks,\nDan\n\n\n"},{"id":"313035","messageId":"xmqq1suge1jn.fsf@gitster.mtv.corp.google.com","threadId":"45255","inReplyTo":"CY1PR0301MB21073D82F4A6AB0DAD8BF1FCC4280@CY1PR0301MB2107.namprd03.prod.outlook.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-03-02T04:38:04Z","receivedAt":"2017-03-02T04:55:02Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Dan Shumow <danshu@microsoft.com> writes:\n\n> At this point, I would suggest that I take the C optimizations,\n> clean them up and fold them in with the diet changes Linus has\n> suggested.  The slowdown is still 2x over block-sha1 and more over\n> OpenSSL.  But it is better than nothing.  And then if there is\n> interest Marc and I can investigate other processor specific\n> optimizations like ASM or SIMD and circle back with those\n> performance optimizations at a later date.\n>\n> Also, to Johannes Schindelin's point:\n>> My concern is about that unexpected turn \"oh, let's just switch\n>> to C99 because, well, because my compiler canehandle it, and\n>> everybody else should just switch tn a modern compiler\". That\n>> really sounded careless.\n>\n> While it will probably be a pain, if it is a requirement, we can\n> modify the code to move away from any c99 specific stuff we have\n> in here, if it makes adopting the code more palatable for Git.\n\nI was assuming that we would treat your code just like how we treat\nany other \"borrowed code from elsewhere\".  The usual way for us to\ndo so is to take code that was released by the \"upstream\" (under a\nlicense that allows us to use it---yours is MIT, which does) in the\nstyle and language of upstream's choice, and then we in the Git\ndevelopment community takes responsiblity for massaging the code to\nmatch our style, for trimming what we won't use and for doing any\nother customization to fit our needs.\n\nAs you and Marc seemed to be still working on speeding up, such a\ncustomization work to fully adjust your code to our codebase was\npremature, so I tentatively queued what we saw on the list as-is on\nour 'pu' branch so that people can have a reference point.  Which\nunfortunately solicited a premature reaction by Johannes.  Please do\nnot worry too much about the comment.\n\nBut if you are willing to help us by getting involved in the\n\"customization\" part, too, that would be a very welcome news to us.\nIn that case, \"welcome to the Git development community\" ;-)\n\nSo,... from my point of view, we are OK either way.  It is OK if you\nare a third-party upstream that is not particularly interested in\nGit project's specific requirement.  We surely would be happier if\nyou and Marc, the upstream authors of the code in question, also act\nas participants in the Git development community.\n\nEither way, thanks for your great help.\n"},{"id":"313039","messageId":"CACsJy8D3h1KAaKi_Esc98za3LqXaB=YeW0Yu+VAV9UnX5vmttg@mail.gmail.com","threadId":"45255","inReplyTo":"20170301231921.2puf7o7jkrujscwn@sigill.intra.peff.net","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2017-03-02T06:10:23Z","receivedAt":"2017-03-02T06:12:44Z","isPatch":true,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Mar 2, 2017 at 6:19 AM, Jeff King <peff@peff.net> wrote:\n> You have to remember that some of the Git for Windows users are doing\n> horrific things like using repositories with 450MB .git/index files, and\n> the speed to compute the sha1 during an update is noticeable there.\n\nWe probably should separate this use case from the object hashing\nanyway. Here we need a better, more reliable crc32 basically, to\ndetect bit flips. Even if we move to SHA-something, we can keep\nstaying with SHA-1 here (and with the fastest implementation)\n-- \nDuy\n"},{"id":"313041","messageId":"CA+55aFzd+QYMCzyxzqucPMfNyazWdoo7XbDgsbdnJQRWHw+xPA@mail.gmail.com","threadId":"45255","inReplyTo":"20170301190711.nwlretmy6w675jvp@sigill.intra.peff.net","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-01T19:10:54Z","receivedAt":"2017-03-02T06:41:24Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Wed, Mar 1, 2017 at 11:07 AM, Jeff King <peff@peff.net> wrote:\n>\n> So obviously the smaller object size is nice, and the diffstat is\n> certainly satisfying. My only qualm would be whether this conflicts with\n> the optimizations that Dan is working on (probably not conceptually, but\n> textually).\n\nYeah. But I'll happily just re-apply the patch on any new version that\nDan posts.  The patch is obviously trivial, even if size-wise it's a\nfair number of lines.\n\nSo I wouldn't suggest using the patched version as some kind of\nstarting point. It's much easier to just take a new version of\nupstream and repeat the diet patch on it.\n\n.. and obviously later versions of the upstream sha1dc code may not\neven need this at all since Dan and Marc are now aware of the issue.\n\n                  Linus\n"},{"id":"313069","messageId":"alpine.DEB.2.20.1703021538060.3767@virtualbox","threadId":"45255","inReplyTo":"20170301231921.2puf7o7jkrujscwn@sigill.intra.peff.net","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-03-02T14:39:27Z","receivedAt":"2017-03-02T14:40:51Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Peff,\n\nOn Wed, 1 Mar 2017, Jeff King wrote:\n\n> I do think that could argue for turning on the collision detection only\n> during object-write operations, which is where it matters. It would be\n> really trivial to flip the \"check collisions\" bit on sha1dc. But I\n> suspect you could go faster still by compiling against two separate\n> implementations: the fast-as-possible one (which could be openssl or\n> blk-sha1), and the slower-but-careful sha1dc.\n\nGiven the speed difference between OpenSSL and sha1dc, it would be a wise\nthing indeed to do sha1dc only where objects enter from possibly untrusted\nsources, and use OpenSSL for all other hashing.\n\nCiao,\nJohannes\n"},{"id":"313071","messageId":"alpine.DEB.2.20.1703021530380.3767@virtualbox","threadId":"45255","inReplyTo":"CA+55aFy9=jBJT36FC2HiAeabJBssY=jE=zLxwrXWzhpiFkMUXg@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-03-02T14:37:00Z","receivedAt":"2017-03-02T14:54:54Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Linus,\n\nOn Wed, 1 Mar 2017, Linus Torvalds wrote:\n\n> On Wed, Mar 1, 2017 at 2:51 PM, Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n> >\n> > But I think bigger than just developers on Windows OS. There are many\n> > developers out there working on large repositories (yes, much larger\n> > than Linux). Also using Macs and Linux. I am not at all sure that we\n> > want to give them an updated Git they cannot fail to notice to be much\n> > slower than before.\n> \n> Johannes, have you *tried* the patches?\n> \n> I really don't think you have. It is completely unnoticeable in any\n> normal situation. The two cases that it's noticeable is:\n> \n>  - a full fsck is noticeable slower\n> \n>  - a full non-local clone is slower (but not really noticeably so\n> since the network traffic dominates).\n> \n> In other words, I think you're making shit up. I don't think you\n> understand how little the SHA1 performance actually matters. It's\n> noticeable in benchmarks. It's not noticeable in any normal operation.\n> \n> .. and yes, I've actually been running the patches locally since I\n> posted my first version (which apparently didn't go out to the list\n> because of list size limits) and now running the version in 'pu'.\n\nIf you think that the Linux repository is a big one, then your reaction is\nunderstandable.\n\nI have zero interest in potty language, therefore my reply is very terse:\nyes, I have been looking ad SHA-1 performance, and yes, it matters. Think\nan index file of 300-400MB.\n\nCiao,\nJohannes\n"},{"id":"313076","messageId":"alpine.DEB.2.20.1703021539330.3767@virtualbox","threadId":"45255","inReplyTo":"CACsJy8D3h1KAaKi_Esc98za3LqXaB=YeW0Yu+VAV9UnX5vmttg@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-03-02T14:45:14Z","receivedAt":"2017-03-02T15:56:48Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Duy,\n\nOn Thu, 2 Mar 2017, Duy Nguyen wrote:\n\n> On Thu, Mar 2, 2017 at 6:19 AM, Jeff King <peff@peff.net> wrote:\n> > You have to remember that some of the Git for Windows users are doing\n> > horrific things like using repositories with 450MB .git/index files,\n> > and the speed to compute the sha1 during an update is noticeable\n> > there.\n> \n> We probably should separate this use case from the object hashing\n> anyway. Here we need a better, more reliable crc32 basically, to detect\n> bit flips. Even if we move to SHA-something, we can keep staying with\n> SHA-1 here (and with the fastest implementation)\n\nI guess it was convenient to use the same hash algorithm for all hashing\npurposes in the beginning. The downside, of course, was that we kept\ntalking about SHA-1s instead of commit hashes and the index checksum (i.e.\nusing labels based on implementation details rather than semantically\nmeaningful names).\n\nIn the meantime, we use different hash algorithms where appropriate, of\ncourse, and we typically encapsulate the exact hash algorithm so that it\nis easy to switch when/if necessary (think the hash functions for strings\nin our hashtables, and the hash functions in xdiff).\n\nIt would probably make sense to switch the index integrity check away from\nSHA-1 because we really only care about detecting bit flips there, and we\nhave no need for the computational overhead of using a full-blown\ncryptographic hash for that purpose.\n\nCiao,\nJohannes\n"},{"id":"313079","messageId":"CA+55aFzscLaviJac-SB65WFYViY=wyAF3EWOnhHSuzSuFLdPTA@mail.gmail.com","threadId":"45255","inReplyTo":"alpine.DEB.2.20.1703021539330.3767@virtualbox","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-02T16:35:36Z","receivedAt":"2017-03-02T16:37:23Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Mar 2, 2017 at 6:45 AM, Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n>\n> It would probably make sense to switch the index integrity check away from\n> SHA-1 because we really only care about detecting bit flips there, and we\n> have no need for the computational overhead of using a full-blown\n> cryptographic hash for that purpose.\n\nWhich index do you actually see as being a problem, btw? The main file\nindex (.git/index) or the pack-file indexes?\n\nWe definitely don't need the checking version of sha1 for either of\nthose, but as Jeff already did the math, at least the pack-file index\nis almost negligible, because the pack-file operations that update it\nend up doing SHA1 over the objects - and the object SHA1 calculations\nare much bigger.\n\nAnd I don't think we even check the pack-file index hashes except on fsck.\n\nNow, if your _file_ index is 300-400MB (and I do think we check the\nSHA fingerprint on that even on just reading it - verify_hdr() in\ndo_read_index()), then that's going to be a somewhat noticeable hit on\nevery normal \"git diff\" etc.\n\nBut I'd have expected the stat() calls of all the files listed by that\nindex to be the _much_ bigger problem in that case. Or do you just\nturn those off with assume-unchanged?\n\nYeah, those stat calls are threaded when preloading, but even so..\n\nAnyway, the file index SHA1 checking could probably just be disabled\nentirely (with a config flag). It's a corruption check that simply\nisn't that important. So if that's your main SHA1 issue, that would be\neasy to fix.\n\nEverything else - like pack-file generation etc for a big clone() may\nend up using a ton of SHA1 too, but the SHA1 costs all scale with the\nother costs that drown them out (ie zlib, network, etc).\n\nI'd love to see a profile if you have one.\n\n                      Linus\n"},{"id":"313084","messageId":"85221b97-759f-b7a9-1256-21515d163cbf@jeffhostetler.com","threadId":"45255","inReplyTo":"CA+55aFzscLaviJac-SB65WFYViY=wyAF3EWOnhHSuzSuFLdPTA@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff Hostetler","fromEmail":"git@jeffhostetler.com","sentAt":"2017-03-02T18:37:27Z","receivedAt":"2017-03-02T18:46:27Z","isPatch":true,"sender":{"key":"git@jeffhostetler.com","avatar":null},"body":"\n\nOn 3/2/2017 11:35 AM, Linus Torvalds wrote:\n> On Thu, Mar 2, 2017 at 6:45 AM, Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n>> It would probably make sense to switch the index integrity check away from\n>> SHA-1 because we really only care about detecting bit flips there, and we\n>> have no need for the computational overhead of using a full-blown\n>> cryptographic hash for that purpose.\n> Which index do you actually see as being a problem, btw? The main file\n> index (.git/index) or the pack-file indexes?\n>\n> We definitely don't need the checking version of sha1 for either of\n> those, but as Jeff already did the math, at least the pack-file index\n> is almost negligible, because the pack-file operations that update it\n> end up doing SHA1 over the objects - and the object SHA1 calculations\n> are much bigger.\n>\n> And I don't think we even check the pack-file index hashes except on fsck.\n>\n> Now, if your _file_ index is 300-400MB (and I do think we check the\n> SHA fingerprint on that even on just reading it - verify_hdr() in\n> do_read_index()), then that's going to be a somewhat noticeable hit on\n> every normal \"git diff\" etc.\n\nYes, the .git/index is 450MB with ~3.1M entries.  verify_hdr() is called \neach time\nwe read it into memory.\n\nWe have been testing a patch in GfW to run the verification in a \nseparate thread\nwhile the main thread parses (and mallocs) the cache_entries.  This does \nhelp\noffset the time.\n         https://github.com/git-for-windows/git/pull/978/files\n\n> But I'd have expected the stat() calls of all the files listed by that\n> index to be the _much_ bigger problem in that case. Or do you just\n> turn those off with assume-unchanged?\n>\n> Yeah, those stat calls are threaded when preloading, but even so..\n\nYes, the stat() calls are more significant percentage of the time (and \nhaving\ncore.fscache and core.preloadindex help that greatly), but the total \ntime for a command\nis just that -- the total -- so using the philosophy of \"every little \nbit helps\", the faster\nroutines help us here.\n\n> Anyway, the file index SHA1 checking could probably just be disabled\n> entirely (with a config flag). It's a corruption check that simply\n> isn't that important. So if that's your main SHA1 issue, that would be\n> easy to fix.\n\nYes, in the GVFS effort, we disabled the verification with a config \nsetting and haven't\nhad any incidents.\n\n\n> Everything else - like pack-file generation etc for a big clone() may\n> end up using a ton of SHA1 too, but the SHA1 costs all scale with the\n> other costs that drown them out (ie zlib, network, etc).\n>\n> I'd love to see a profile if you have one.\n>\n>                        Linus\n\n"},{"id":"313093","messageId":"CA+55aFxbfmW6UL8cd=s43=bbDkTaqUmah0+snhzR3j_Fyv=-gw@mail.gmail.com","threadId":"45255","inReplyTo":"85221b97-759f-b7a9-1256-21515d163cbf@jeffhostetler.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-02T19:04:36Z","receivedAt":"2017-03-02T19:37:07Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Mar 2, 2017 at 10:37 AM, Jeff Hostetler <git@jeffhostetler.com> wrote:\n>>\n>> Now, if your _file_ index is 300-400MB (and I do think we check the\n>> SHA fingerprint on that even on just reading it - verify_hdr() in\n>> do_read_index()), then that's going to be a somewhat noticeable hit on\n>> every normal \"git diff\" etc.\n>\n> Yes, the .git/index is 450MB with ~3.1M entries.  verify_hdr() is called\n> each time we read it into memory.\n\nOk. So that's really just a purely historical artifact.\n\nThe file index is actually the first part of git to have ever been\nwritten. You can't even see it in the history, because the initial\nrevision from Apr 7, 2005, obviously depended on the actual object\nhashing.\n\nBut the file index actually came first. You can _kind_ of see that in\nthe layout of the original git tree, and how the main header file is\nstill called \"cache.h\", and how the original \".git\" directory was\nactually called \".dircache\".\n\nAnd the two biggest files (by a fairly big margin) are \"read-cache.c\"\nand \"update-cache.c\".\n\nSo that file index cache was in many ways _the_ central part of the\noriginal git model. The sha1 file indexing and object database was\njust the backing store for the file index.\n\nBut part of that history is then how much I worried about corruption\nof that index (and, let's face it, general corruption resistance _was_\none of the primary design goals - performance was high up there too,\nbut safety in the face of filesystem corruption was and is a primary\nissue).\n\nBut realistically, I don't think we've *ever* hit anything serious on\nthe index file, and it's obviously not a security issue. It also isn't\neven a compatibility issue, so it would be trivial to just bump the\nversion header and saying that the signature changes the meaning of\nthe checksum.\n\nThat said:\n\n> We have been testing a patch in GfW to run the verification in a separate thread\n> while the main thread parses (and mallocs) the cache_entries.  This does help\n> offset the time.\n\nYeah, that seems an even better solution, honestly.\n\nThe patch would be cleaner without the NO_PTHREADS things.\n\nI wonder how meaningful that thing even is today. Looking at what\nseems to select NO_PTHREADS, I suspect that's all entirely historical.\nFor example, you'll see it for QNX etc, which seems wrong - QNX\ndefinitely has pthreads according to their docs, for example.\n\n                     Linus\n"},{"id":"313252","messageId":"CY1PR0301MB2107112BCC2DECD215E70549C42A0@CY1PR0301MB2107.namprd03.prod.outlook.com","threadId":"45255","inReplyTo":"xmqq1suge1jn.fsf@gitster.mtv.corp.google.com","subject":"RE: [PATCH] Put sha1dc on a diet","fromName":"Dan Shumow","fromEmail":"danshu@microsoft.com","sentAt":"2017-03-04T01:07:16Z","receivedAt":"2017-03-04T01:07:25Z","isPatch":true,"sender":{"key":"danshu@microsoft.com","avatar":null},"body":"From: Junio C Hamano [mailto:gitster@pobox.com] \n\n> As you and Marc seemed to be still working on speeding up, such a customization work to fully adjust your code to our codebase was premature, so I tentatively queued what we saw on the list as-is on our 'pu' branch so that people can have a reference point.  Which unfortunately solicited a premature reaction by Johannes.  Please do not worry too much about the comment.\n\n> But if you are willing to help us by getting involved in the \"customization\" part, too, that would be a very welcome news to us.\n>In that case, \"welcome to the Git development community\" ;-)\n\n> So,... from my point of view, we are OK either way.  It is OK if you are a third-party upstream that is not particularly interested in Git project's specific requirement.  We surely would be happier if you and Marc, the upstream authors of the code in question, also act as participants in the Git development community.\n\n> Either way, thanks for your great help.\n\nYou are very welcome.  Thank you for the warm welcome.  As it turns out, Marc and I are working on the simplifications / removal of c99 and performance upstream in our GitHub repo.  I am happy to help for any GitHub specific customizations that are needed as well.  But for now, lets see if we can get you everything you want upstream -- I think that's the most simple.\n\nThanks again,\nDan\n\n"},{"id":"313895","messageId":"20170313151322.ouryghyb5orkpk5g@sigill.intra.peff.net","threadId":"45255","inReplyTo":"CY1PR0301MB2107112BCC2DECD215E70549C42A0@CY1PR0301MB2107.namprd03.prod.outlook.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-13T15:13:22Z","receivedAt":"2017-03-13T15:13:30Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Mar 04, 2017 at 01:07:16AM +0000, Dan Shumow wrote:\n\n> You are very welcome.  Thank you for the warm welcome.  As it turns\n> out, Marc and I are working on the simplifications / removal of c99\n> and performance upstream in our GitHub repo.  I am happy to help for\n> any GitHub specific customizations that are needed as well.  But for\n> now, lets see if we can get you everything you want upstream -- I\n> think that's the most simple.\n\nI've been watching the repo at:\n\n  https://github.com/cr-marcstevens/sha1collisiondetection\n\nThe work on the feature/performance branch seems to be producing good\nresults. The best timings I got show sha1dc (with checks enabled) at\n1.75x block-sha1, which is pretty good. That was using your\nad744c8b7a841d2afcb2d4c04f8952d9005501be.\n\nCuriously, the performance gets worse after that. Even more curious, the\nbad performance bisects to a merge, and it performs worse than either\nside of the merge.\n\nTry this:\n\n  # mine is a 1.2GB linux packfile, but anything big should do\n  file=/some/large/file\n\n  # the merge with the funny behavior\n  merge=55d1db0980501e582f6cd103a04f493995b1df78\n\n  for i in $merge^ $merge^2 $merge; do\n    git checkout $i &&\n    rm -f bin/* &&\n    make &&\n    time bin/sha1dcsum $file\n  done\n\nI get:\n\n  [$merge^, the feature/performance branch before the merge]\n  real\t0m3.391s\n  user\t0m3.304s\n  sys\t0m0.084s\n\n  [$merge^2, the master branch before the merge]\n  real\t0m5.272s\n  user\t0m5.164s\n  sys\t0m0.096s\n\n  [$merge, the merge of the two]\n  real\t0m7.038s\n  user\t0m6.924s\n  sys\t0m0.104s\n\nSo that's odd. Looking at the diff, I don't see anything that obviously\njumps out as a mis-merge.\n\nFeel free to tell me \"stop looking at that branch; it's a work in\nprogress\". But I think the results from $merge^ (ad744c8b7) are getting\ngood enough to consider moving forward with integrating it into git.\n\n-Peff\n"},{"id":"313938","messageId":"20170313194848.2z2dlgpomu6e3dkh@sigill.intra.peff.net","threadId":"45255","inReplyTo":"CY1PR0301MB2107876B6E47FBCF03AB1EA1C4250@CY1PR0301MB2107.namprd03.prod.outlook.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-13T19:48:49Z","receivedAt":"2017-03-13T19:48:58Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Mar 13, 2017 at 07:42:17PM +0000, Dan Shumow wrote:\n\n> Marc just made a commit this morning fixing problems with the merge.\n> Please give the latest in feature/performance a try, as that seems to\n> eliminate the problem.\n\nYeah, b17728507 makes the problem go away for me. Thanks.\n\nFWIW, I have all sha1s on github.com running through this right now\n(actually, the ad744c8b7 version), and logging any false-positives on\nthe collision detection. Nothing so far, after a few hours.\n\n-Peff\n"},{"id":"313956","messageId":"1e6a592f-7da1-8043-0b29-0bb7c8cda3f3@cwi.nl","threadId":"45255","inReplyTo":"20170313194848.2z2dlgpomu6e3dkh@sigill.intra.peff.net","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Marc Stevens","fromEmail":"marc.stevens@cwi.nl","sentAt":"2017-03-13T20:12:34Z","receivedAt":"2017-03-13T20:13:15Z","isPatch":true,"sender":{"key":"marc.stevens@cwi.nl","avatar":null},"body":"Indeed, I've committed a fix, and a small bug fix for the new code just now.\n\nThe merge incorrectly removed some control logic,\nwhich caused more unnecessary checks to happen.\nI already marked this in the PR, but committed a fix only today.\n\nBTW as noted in the Readme, the theoretic false positive probability is\n<<2^-90, almost non-existent.\n\nBest regards,\nMarc Stevens\n\n\nOn 3/13/2017 8:48 PM, Jeff King wrote:\n> On Mon, Mar 13, 2017 at 07:42:17PM +0000, Dan Shumow wrote:\n>\n>> Marc just made a commit this morning fixing problems with the merge.\n>> Please give the latest in feature/performance a try, as that seems to\n>> eliminate the problem.\n> Yeah, b17728507 makes the problem go away for me. Thanks.\n>\n> FWIW, I have all sha1s on github.com running through this right now\n> (actually, the ad744c8b7 version), and logging any false-positives on\n> the collision detection. Nothing so far, after a few hours.\n>\n> -Peff\n\n"},{"id":"313957","messageId":"CA+55aFyNi2uHwd9nzjy3dOu2L1A0jPN6AD43WKj-05km1GNtRQ@mail.gmail.com","threadId":"45255","inReplyTo":"1e6a592f-7da1-8043-0b29-0bb7c8cda3f3@cwi.nl","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-13T20:20:02Z","receivedAt":"2017-03-13T20:20:09Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Mon, Mar 13, 2017 at 1:12 PM, Marc Stevens <marc.stevens@cwi.nl> wrote:\n> Indeed, I've committed a fix, and a small bug fix for the new code just now.\n\nUnrelated side note: there may be some missing dependencies in the\nbuild infrastructure or something, because when I tried Jeff's script\nthat did that \"test the merge and the two parents\", and used the\npack-file of the kernel for testing, I got:\n\n  5611971c610143e6d38bbdca463f4c9f79a056a0\n/home/torvalds/v2.6/linux/.git/objects/pack/pack-153bb8cd11846cf9a27ef7b1069aa9cb9f5b724f.pack\n\n  real 0m2.432s\n  user 0m2.348s\n  sys 0m0.084s\n\n  5611971c610143e6d38bbdca463f4c9f79a056a0\n/home/torvalds/v2.6/linux/.git/objects/pack/pack-153bb8cd11846cf9a27ef7b1069aa9cb9f5b724f.pack\n\n  real 0m3.747s\n  user 0m3.672s\n  sys 0m0.076s\n\n  5611971c610143e6d38bbdca463f4c9f79a056a0 *coll*\n/home/torvalds/v2.6/linux/.git/objects/pack/pack-153bb8cd11846cf9a27ef7b1069aa9cb9f5b724f.pack\n\n  real 0m5.061s\n  user 0m4.984s\n  sys 0m0.077s\n\nnever mind the performace, notice the *coll* in that last case.\n\nBut doing a \"git clean -dqfx; make -j8\" and re-testing the same tree,\nthe issue is gone.\n\nI suspect some dependency on a header file is broken, causing some\nobject file to not be properly re-built, which in turn then\nincorrectly causes the 'ctx2.found_collision' test to test the wrong\nbit or something.\n\n                 Linus\n"},{"id":"313965","messageId":"161775901.3349663.1489438074825.JavaMail.zimbra@cwi.nl","threadId":"45255","inReplyTo":"CA+55aFyNi2uHwd9nzjy3dOu2L1A0jPN6AD43WKj-05km1GNtRQ@mail.gmail.com","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Marc Stevens","fromEmail":"marc.stevens@cwi.nl","sentAt":"2017-03-13T20:47:54Z","receivedAt":"2017-03-13T20:48:15Z","isPatch":true,"sender":{"key":"marc.stevens@cwi.nl","avatar":null},"body":"Linus:\nI would be surprised, the dependencies should be automatically determined.\n\nBTW Did you make local changes to this perf branch?\nSpecifically did you disable the safe hash mode that is on by default?\nBecause if you did not, it might also be something else as all three hashes below are the same.\n\n-- Marc\n\n----- Original Message -----\nFrom: \"Linus Torvalds\" <torvalds@linux-foundation.org>\nTo: \"Marc Stevens\" <marc.stevens@cwi.nl>\nCc: \"Jeff King\" <peff@peff.net>, \"Dan Shumow\" <danshu@microsoft.com>, \"Junio C Hamano\" <gitster@pobox.com>, \"Git Mailing List\" <git@vger.kernel.org>\nSent: Monday, March 13, 2017 9:20:02 PM\nSubject: Re: [PATCH] Put sha1dc on a diet\n\nOn Mon, Mar 13, 2017 at 1:12 PM, Marc Stevens <marc.stevens@cwi.nl> wrote:\n> Indeed, I've committed a fix, and a small bug fix for the new code just now.\n\nUnrelated side note: there may be some missing dependencies in the\nbuild infrastructure or something, because when I tried Jeff's script\nthat did that \"test the merge and the two parents\", and used the\npack-file of the kernel for testing, I got:\n\n  5611971c610143e6d38bbdca463f4c9f79a056a0\n/home/torvalds/v2.6/linux/.git/objects/pack/pack-153bb8cd11846cf9a27ef7b1069aa9cb9f5b724f.pack\n\n  real 0m2.432s\n  user 0m2.348s\n  sys 0m0.084s\n\n  5611971c610143e6d38bbdca463f4c9f79a056a0\n/home/torvalds/v2.6/linux/.git/objects/pack/pack-153bb8cd11846cf9a27ef7b1069aa9cb9f5b724f.pack\n\n  real 0m3.747s\n  user 0m3.672s\n  sys 0m0.076s\n\n  5611971c610143e6d38bbdca463f4c9f79a056a0 *coll*\n/home/torvalds/v2.6/linux/.git/objects/pack/pack-153bb8cd11846cf9a27ef7b1069aa9cb9f5b724f.pack\n\n  real 0m5.061s\n  user 0m4.984s\n  sys 0m0.077s\n\nnever mind the performace, notice the *coll* in that last case.\n\nBut doing a \"git clean -dqfx; make -j8\" and re-testing the same tree,\nthe issue is gone.\n\nI suspect some dependency on a header file is broken, causing some\nobject file to not be properly re-built, which in turn then\nincorrectly causes the 'ctx2.found_collision' test to test the wrong\nbit or something.\n\n                 Linus\n"},{"id":"313967","messageId":"20170313210023.bumtp6wyw6blmymp@sigill.intra.peff.net","threadId":"45255","inReplyTo":"161775901.3349663.1489438074825.JavaMail.zimbra@cwi.nl","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-13T21:00:23Z","receivedAt":"2017-03-13T21:00:36Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Mar 13, 2017 at 09:47:54PM +0100, Marc Stevens wrote:\n\n> Linus:\n> I would be surprised, the dependencies should be automatically determined.\n> \n> BTW Did you make local changes to this perf branch?\n\nI can reproduce it with:\n\n  cd sha1collisiondetection\n  git clean -dqfx ;# make sure we are starting from scratch\n\n  git checkout 9c8e73cadb35776d3310e3f8ceda7183fa75a39f\n  make\n  bin/sha1dcsum $file\n\n  git checkout 55d1db0980501e582f6cd103a04f493995b1df78\n  make\n  bin/sha1dcsum $file\n\nThe final call to sha1dcsum will report a collision, even though the\nfirst one did not.\n\nIt also reproduces with the original snippet I posted. I didn't notice\nbecause I was just collecting the timings then (and I originally noticed\nthe problem on the versions I had pulled into Git, where it works as\nexpected; but then I am just pulling in the two source files, without\nall of the libtool magic).\n\n-Peff\n"},{"id":"313969","messageId":"1392458356.3351662.1489439723458.JavaMail.zimbra@cwi.nl","threadId":"45255","inReplyTo":"20170313210023.bumtp6wyw6blmymp@sigill.intra.peff.net","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Marc Stevens","fromEmail":"marc.stevens@cwi.nl","sentAt":"2017-03-13T21:15:23Z","receivedAt":"2017-03-13T21:16:00Z","isPatch":true,"sender":{"key":"marc.stevens@cwi.nl","avatar":null},"body":"I think I now understand.\nThe Makefile indeed seems to fail to correctly rebuild when a header has changed.\n\nAs the performance branch has removed the 'int bigendian' from SHA1_CTX in lib/sha1.h,\nthe perf-branch and master-branch are binary incompatible.\nSo the command-line utility does not get fully recompiled \nand instead of the value of found_collision will read a different value of SHA1_CTX.\n\nSo be careful to always do a 'make clean' for now.\n\n-- Marc\n\n----- Original Message -----\nFrom: \"Jeff King\" <peff@peff.net>\nTo: \"Marc Stevens\" <Marc.Stevens@cwi.nl>\nCc: \"Linus Torvalds\" <torvalds@linux-foundation.org>, \"Dan Shumow\" <danshu@microsoft.com>, \"Junio C Hamano\" <gitster@pobox.com>, \"Git Mailing List\" <git@vger.kernel.org>\nSent: Monday, March 13, 2017 10:00:23 PM\nSubject: Re: [PATCH] Put sha1dc on a diet\n\nOn Mon, Mar 13, 2017 at 09:47:54PM +0100, Marc Stevens wrote:\n\n> Linus:\n> I would be surprised, the dependencies should be automatically determined.\n> \n> BTW Did you make local changes to this perf branch?\n\nI can reproduce it with:\n\n  cd sha1collisiondetection\n  git clean -dqfx ;# make sure we are starting from scratch\n\n  git checkout 9c8e73cadb35776d3310e3f8ceda7183fa75a39f\n  make\n  bin/sha1dcsum $file\n\n  git checkout 55d1db0980501e582f6cd103a04f493995b1df78\n  make\n  bin/sha1dcsum $file\n\nThe final call to sha1dcsum will report a collision, even though the\nfirst one did not.\n\nIt also reproduces with the original snippet I posted. I didn't notice\nbecause I was just collecting the timings then (and I originally noticed\nthe problem on the versions I had pulled into Git, where it works as\nexpected; but then I am just pulling in the two source files, without\nall of the libtool magic).\n\n-Peff\n"},{"id":"314238","messageId":"2006239187.136016.1489688534478.JavaMail.zimbra@cwi.nl","threadId":"45255","inReplyTo":"1392458356.3351662.1489439723458.JavaMail.zimbra@cwi.nl","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Marc Stevens","fromEmail":"marc.stevens@cwi.nl","sentAt":"2017-03-16T18:22:14Z","receivedAt":"2017-03-16T18:26:57Z","isPatch":true,"sender":{"key":"marc.stevens@cwi.nl","avatar":null},"body":"Hi all,\n\nToday I merged the perf-branch into master after code review and correctness testing.\nSo master is now more performant and safe to use.\n\n-- Marc\n"},{"id":"314262","messageId":"20170316220620.ihq4ulg4t6m7ktrh@sigill.intra.peff.net","threadId":"45255","inReplyTo":"2006239187.136016.1489688534478.JavaMail.zimbra@cwi.nl","subject":"Re: [PATCH] Put sha1dc on a diet","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-16T22:06:20Z","receivedAt":"2017-03-16T22:06:47Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 16, 2017 at 07:22:14PM +0100, Marc Stevens wrote:\n\n> Today I merged the perf-branch into master after code review and correctness testing.\n> So master is now more performant and safe to use.\n\nGreat, thank you (and Dan) so much for all your work. We're looking at\nintegrating this version in a nearby thread.\n\n-Peff\n"},{"id":"314263","messageId":"CY1PR0301MB2107A6C40342CC8F25D0F661C4260@CY1PR0301MB2107.namprd03.prod.outlook.com","threadId":"45255","inReplyTo":"20170316220620.ihq4ulg4t6m7ktrh@sigill.intra.peff.net","subject":"RE: [PATCH] Put sha1dc on a diet","fromName":"Dan Shumow","fromEmail":"danshu@microsoft.com","sentAt":"2017-03-16T22:07:43Z","receivedAt":"2017-03-16T22:07:51Z","isPatch":true,"sender":{"key":"danshu@microsoft.com","avatar":null},"body":"Great! Keep us posted if there is anything else that you would like from the code.  Or anyway we can make the process go more smoothly.\n\nThanks,\nDan\n\n\n-----Original Message-----\nFrom: Jeff King [mailto:peff@peff.net] \nSent: Thursday, March 16, 2017 3:06 PM\nTo: Marc Stevens <Marc.Stevens@cwi.nl>\nCc: Linus Torvalds <torvalds@linux-foundation.org>; Dan Shumow <danshu@microsoft.com>; Junio C Hamano <gitster@pobox.com>; Git Mailing List <git@vger.kernel.org>\nSubject: Re: [PATCH] Put sha1dc on a diet\n\nOn Thu, Mar 16, 2017 at 07:22:14PM +0100, Marc Stevens wrote:\n\n> Today I merged the perf-branch into master after code review and correctness testing.\n> So master is now more performant and safe to use.\n\nGreat, thank you (and Dan) so much for all your work. We're looking at integrating this version in a nearby thread.\n\n-Peff\n"}]}