{"thread":{"id":"45202","subject":"SHA1 collisions found","startedAt":"2017-02-23T16:52:48Z","lastAt":"2017-03-05T23:45:40Z","messageCount":136,"participants":["Joey Hess","Junio C Hamano","David Lang","Linus Torvalds","Jeff King","Morten Welinder","Øyvind A. Holm","Jakub Narębski","Stefan Beller","Duy Nguyen","Geert Uytterhoeven","Ian Jackson","ankostis","Jason Cooper","HW42","Philip Oakley","Santiago Torres","Jacob Keller","grarpamp","brian m. carlson","Mike Hommey","Lars Schneider","Thomas Braun","Ævar Arnfjörð Bjarmason","René Scharfe","Markus Trippelsdorf","Tony Finch","Shawn Pearce","Dan Shumow","Marc Stevens","Brandon Williams"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"312395","messageId":"20170223164306.spg2avxzukkggrpb@kitenet.net","threadId":"45202","inReplyTo":null,"subject":"SHA1 collisions found","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2017-02-23T16:43:06Z","receivedAt":"2017-02-23T16:52:48Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"https://shattered.io/static/shattered.pdf\nhttps://freedom-to-tinker.com/2017/02/23/rip-sha-1/\n\nIIRC someone has been working on parameterizing git's SHA1 assumptions\nso a repository could eventually use a more secure hash. How far has\nthat gotten? There are still many \"40\" constants in git.git HEAD.\n\nIn the meantime, git commit -S, and checks that commits are signed,\nseems like the only way to mitigate against attacks such as\nthe ones described in the threads at\nhttps://joeyh.name/blog/sha-1/ and\nhttps://joeyh.name/blog/entry/size_of_the_git_sha1_collision_attack_surface/\n\nSince we now have collisions in valid PDF files, collisions in valid git\ncommit and tree objects are probably able to be constructed.\n\n-- \nsee shy jo\n"},{"id":"312396","messageId":"CAPc5daVZ79WWKSw76kxHgDra9a7fSR1AibZa_pvK9aUuuVawLQ@mail.gmail.com","threadId":"45202","inReplyTo":"20170223164306.spg2avxzukkggrpb@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-23T17:02:38Z","receivedAt":"2017-02-23T17:04:59Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"On Thu, Feb 23, 2017 at 8:43 AM, Joey Hess <id@joeyh.name> wrote:\n>\n> Since we now have collisions in valid PDF files, collisions in valid git\n> commit and tree objects are probably able to be constructed.\n\nThat may be true, but\nhttps://public-inbox.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/\n"},{"id":"312398","messageId":"nycvar.QRO.7.75.62.1702230857420.6590@qynat-yncgbc","threadId":"45202","inReplyTo":"20170223164306.spg2avxzukkggrpb@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2017-02-23T17:00:20Z","receivedAt":"2017-02-23T17:20:13Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 23 Feb 2017, Joey Hess wrote:\n\n> https://shattered.io/static/shattered.pdf\n> https://freedom-to-tinker.com/2017/02/23/rip-sha-1/\n>\n> IIRC someone has been working on parameterizing git's SHA1 assumptions\n> so a repository could eventually use a more secure hash. How far has\n> that gotten? There are still many \"40\" constants in git.git HEAD.\n>\n> In the meantime, git commit -S, and checks that commits are signed,\n> seems like the only way to mitigate against attacks such as\n> the ones described in the threads at\n> https://joeyh.name/blog/sha-1/ and\n> https://joeyh.name/blog/entry/size_of_the_git_sha1_collision_attack_surface/\n>\n> Since we now have collisions in valid PDF files, collisions in valid git\n> commit and tree objects are probably able to be constructed.\n\nkeep in mind that there is a huge difference between\n\ncreating a collision between two documents you create, both of which contain a \nhuge amount of arbitrary binary data that can be changed at will without \naffecting the results\n\nand\n\ncreating a collision betwen an existing document that someone else created and a \nnew document that is also valid C code without huge amounts of binary in it.\n\nSo, it's not time to panic, but it is one more push to make the changes to \nsupport something else.\n\nDavid Lang\n"},{"id":"312399","messageId":"CAPc5daVfw50dWPF0rc3kCnx_qCEdDCU5X1w=VJDHGut5FbMbRg@mail.gmail.com","threadId":"45202","inReplyTo":"CAPc5daVZ79WWKSw76kxHgDra9a7fSR1AibZa_pvK9aUuuVawLQ@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-23T17:18:48Z","receivedAt":"2017-02-23T17:20:28Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"On Thu, Feb 23, 2017 at 9:02 AM, Junio C Hamano <gitster@pobox.com> wrote:\n> On Thu, Feb 23, 2017 at 8:43 AM, Joey Hess <id@joeyh.name> wrote:\n>>\n>> Since we now have collisions in valid PDF files, collisions in valid git\n>> commit and tree objects are probably able to be constructed.\n>\n> That may be true, but\n> https://public-inbox.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/\n\nIOW, we want to continue the work to switch from SHA-1, but today's announcement\ndoes not fundamentally change anything and we do not panic.\n"},{"id":"312400","messageId":"CA+55aFxJGDpJXqpcoPnwvzcn_fB-zaggj=w7P2At-TOt4buOqw@mail.gmail.com","threadId":"45202","inReplyTo":"20170223164306.spg2avxzukkggrpb@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-23T17:19:06Z","receivedAt":"2017-02-23T17:21:17Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 8:43 AM, Joey Hess <id@joeyh.name> wrote:\n>\n> IIRC someone has been working on parameterizing git's SHA1 assumptions\n> so a repository could eventually use a more secure hash. How far has\n> that gotten? There are still many \"40\" constants in git.git HEAD.\n\nI don't think you'd necessarily want to change the size of the hash.\nYou can use a different hash and just use the same 160 bits from it.\n\n> Since we now have collisions in valid PDF files, collisions in valid git\n> commit and tree objects are probably able to be constructed.\n\nI haven't seen the attack yet, but git doesn't actually just hash the\ndata, it does prepend a type/length field to it. That usually tends to\nmake collision attacks much harder, because you either have to make\nthe resulting size the same too, or you have to be able to also edit\nthe size field in the header.\n\npdf's don't have that issue, they have a fixed header and you can\nfairly arbitrarily add silent data to the middle that just doesn't get\nshown.\n\nSo pdf's make for a much better attack vector, exactly because they\nare a fairly opaque data format. Git has opaque data in some places\n(we hide things in commit objects intentionally, for example, but by\ndefinition that opaque data is fairly secondary.\n\nPut another way: I doubt the sky is falling for git as a source\ncontrol management tool. Do we want to migrate to another hash? Yes.\nIs it \"game over\" for SHA1 like people want to say? Probably not.\n\nI haven't seen the attack details, but I bet\n\n (a) the fact that we have a separate size encoding makes it much\nharder to do on git objects in the first place\n\n (b) we can probably easily add some extra sanity checks to the opaque\ndata we do have, to make it much harder to do the hiding of random\ndata that these attacks pretty much always depend on.\n\n                Linus\n"},{"id":"312401","messageId":"nycvar.QRO.7.75.62.1702230907340.6590@qynat-yncgbc","threadId":"45202","inReplyTo":"CAPc5daVZ79WWKSw76kxHgDra9a7fSR1AibZa_pvK9aUuuVawLQ@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2017-02-23T17:12:16Z","receivedAt":"2017-02-23T17:28:43Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 23 Feb 2017, Junio C Hamano wrote:\n\n> On Thu, Feb 23, 2017 at 8:43 AM, Joey Hess <id@joeyh.name> wrote:\n>>\n>> Since we now have collisions in valid PDF files, collisions in valid git\n>> commit and tree objects are probably able to be constructed.\n>\n> That may be true, but\n> https://public-inbox.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/\n>\n\nit doesn't help that the Google page on this explicitly says that this shows \nthat it's possible to create two different git repos that have the same hash but \ndifferent contents.\n\nhttps://shattered.it/\n\nHow is GIT affected?\nGIT strongly relies on SHA-1 for the identification and integrity checking of \nall file objects and commits. It is essentially possible to create two GIT \nrepositories with the same head commit hash and different contents, say a benign \nsource code and a backdoored one. An attacker could potentially selectively \nserve either repository to targeted users. This will require attackers to \ncompute their own collision.\n\nDavid Lang\n"},{"id":"312402","messageId":"20170223173547.qljypk7sdqi37oha@kitenet.net","threadId":"45202","inReplyTo":"CAPc5daVZ79WWKSw76kxHgDra9a7fSR1AibZa_pvK9aUuuVawLQ@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2017-02-23T17:35:47Z","receivedAt":"2017-02-23T17:37:12Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Junio C Hamano wrote:\n> On Thu, Feb 23, 2017 at 8:43 AM, Joey Hess <id@joeyh.name> wrote:\n> >\n> > Since we now have collisions in valid PDF files, collisions in valid git\n> > commit and tree objects are probably able to be constructed.\n> \n> That may be true, but\n> https://public-inbox.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/\n\nThat's about someone replacing an valid object in Linus's repository\nwith an invalid random blob they found that collides. This SHA1\nbreak doesn't allow generating such a blob anyway. Linus is right,\nthat's an impractical attack.\n\nAttacks using this SHA1 break will look something more like:\n\n* I push a \"bad\" object to a repo on github I set up under a\n  pseudonym.\n* I publish a \"good\" object in a commit and convince the maintainer to\n  merge it.\n* I wait for the maintainer to push to github.\n* I wait for github to deduplicate and hope they'll replace the good\n  object with the bad one I pre-uploaded, thus silently changing the\n  content of the good commit the maintainer reviewed and pushed.\n* The bad object is pulled from github and deployed.\n* The maintainer still has the good object. They may not notice the bad\n  object is out there for a long time.\n\nOf course, it doesn't need to involve Github, and doesn't need to\nrely on internal details of their deduplication[1]; \nthat only let me publish the bad object under a psydonym.\n\n-- \nsee shy jo\n\n[1] Which I'm only guessing about, but now that we have colliding\n    objects, we can upload them to different repos and see if such\n    dedupication happens.\n"},{"id":"312403","messageId":"CA+55aFxjY7mv7YPLZwit7bEhC3VqpEDk1YSRFwSGOEKVw13x4w@mail.gmail.com","threadId":"45202","inReplyTo":"CA+55aFxJGDpJXqpcoPnwvzcn_fB-zaggj=w7P2At-TOt4buOqw@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-23T17:29:48Z","receivedAt":"2017-02-23T17:40:35Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 9:19 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> I don't think you'd necessarily want to change the size of the hash.\n> You can use a different hash and just use the same 160 bits from it.\n\nSide note: I do believe that in practice you should just change the\nsize of the hash too, I'm just saying that the size of the hash and\nthe choice of the hash algorithm are independent issues.\n\nSo you *could* just use  something like SHA3-256, but then pick the\nfirst 160 bits.\n\nRealistically, changing the few hardcoded sizes internally in git is\nlikely the least problem in switching hashes.\n\nSo what you'd probably do is switch to a 256-bit hash, use that\ninternally and in the native git database, and then by default only\n*show* the hash as a 40-character hex string (kind of like how we\nalready abbreviate things in many situations).\n\nThat way tools around git don't even see the change unless passed in\nsome special \"--full-hash\" argument (or \"--abbrev=64\" or whatever -\nthe default being that we abbreviate to 40).\n\n               Linus\n"},{"id":"312404","messageId":"CA+55aFzFEpi1crykZ33r9f7BsvLt_kiB-CHXOkuCAX=fd4BU-w@mail.gmail.com","threadId":"45202","inReplyTo":"20170223173547.qljypk7sdqi37oha@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-23T17:52:36Z","receivedAt":"2017-02-23T18:01:11Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 9:35 AM, Joey Hess <id@joeyh.name> wrote:\n>\n> Attacks using this SHA1 break will look something more like:\n\nWe don't actually know what the break is, but it's likely that you\ncan't actually do what you think you can do:\n\n> * I push a \"bad\" object to a repo on github I set up under a\n>   pseudonym.\n> * I publish a \"good\" object in a commit and convince the maintainer to\n>   merge it.\n\nIt's not clear that the \"good\" object can be anything sane.\n\nWhat you describe pretty much already requires a pre-image attack,\nwhich the new attack is _not_.\n\nThe new attack doesn't have a controlled \"good\" case, you need two\ndifferent objects that both have \"near-collision\" blocks in the\nmiddle. I don't know what the format of those near-collision blocks\nare, but it's a big problem.\n\nYou blithely just say \"I create a good object\". It's not that simple.\nIf it was, this would be a pre-image attack.\n\nSo basically, the attack needs some kind of random binary garbage in\n*both* objects in the middle.\n\nThat's why pdf's are the classic model for showing these attacks: it's\neasy to insert garbage in the middle of a pdf that is invisible.\n\nIn a psf, you can just define a bitmap that you don't use for printing\n- but you can use them to then make a decision about what to print -\nmaking the printed version of the pdf look radically different in ways\nthat are not so much _directly_ about the invisible block itself.\n\n              Linus\n"},{"id":"312405","messageId":"20170223181018.ns4vyosgzmuoyiva@kitenet.net","threadId":"45202","inReplyTo":"CA+55aFxJGDpJXqpcoPnwvzcn_fB-zaggj=w7P2At-TOt4buOqw@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2017-02-23T18:10:18Z","receivedAt":"2017-02-23T18:10:51Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Linus Torvalds wrote:\n> I haven't seen the attack yet, but git doesn't actually just hash the\n> data, it does prepend a type/length field to it. That usually tends to\n> make collision attacks much harder, because you either have to make\n> the resulting size the same too, or you have to be able to also edit\n> the size field in the header.\n\nI have some sha1 collisions (and other fun along these lines) in \nhttps://github.com/joeyh/supercollider\n\nThat includes two files with the same SHA and size, which do get\ndifferent blobs thanks to the way git prepends the header to the\ncontent.\n\njoey@darkstar:~/tmp/supercollider>sha1sum  bad.pdf good.pdf \nd00bbe65d80f6d53d5c15da7c6b4f0a655c5a86a  bad.pdf\nd00bbe65d80f6d53d5c15da7c6b4f0a655c5a86a  good.pdf\njoey@darkstar:~/tmp/supercollider>git ls-tree HEAD\n100644 blob ca44e9913faf08d625346205e228e2265dd12b65\tbad.pdf\n100644 blob 5f90b67523865ad5b1391cb4a1c010d541c816c1\tgood.pdf\n\nWhile appending identical data to these colliding files does generate\nother collisions, prepending data does not.\n\nIt would cost 6500 CPU years + 100 GPU years to generate valid colliding\ngit objects using the methods of the paper's authors. That might be cost\neffective if it helped get a backdoor into eg, the kernel.\n\n>  (b) we can probably easily add some extra sanity checks to the opaque\n> data we do have, to make it much harder to do the hiding of random\n> data that these attacks pretty much always depend on.\n\nFor example, git fsck does warn about a commit message with opaque\ndata hidden after a NUL. But, git show/merge/pull give no indication\nthat something funky is going on when working with such commits.\n\n-- \nsee shy jo\n"},{"id":"312406","messageId":"nycvar.QRO.7.75.62.1702230950040.6590@qynat-yncgbc","threadId":"45202","inReplyTo":"20170223173547.qljypk7sdqi37oha@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2017-02-23T17:52:47Z","receivedAt":"2017-02-23T18:18:46Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Thu, 23 Feb 2017, Joey Hess wrote:\n\n> Junio C Hamano wrote:\n>> On Thu, Feb 23, 2017 at 8:43 AM, Joey Hess <id@joeyh.name> wrote:\n>>>\n>>> Since we now have collisions in valid PDF files, collisions in valid git\n>>> commit and tree objects are probably able to be constructed.\n>>\n>> That may be true, but\n>> https://public-inbox.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/\n>\n> That's about someone replacing an valid object in Linus's repository\n> with an invalid random blob they found that collides. This SHA1\n> break doesn't allow generating such a blob anyway. Linus is right,\n> that's an impractical attack.\n>\n> Attacks using this SHA1 break will look something more like:\n>\n> * I push a \"bad\" object to a repo on github I set up under a\n>  pseudonym.\n> * I publish a \"good\" object in a commit and convince the maintainer to\n>  merge it.\n> * I wait for the maintainer to push to github.\n> * I wait for github to deduplicate and hope they'll replace the good\n>  object with the bad one I pre-uploaded, thus silently changing the\n>  content of the good commit the maintainer reviewed and pushed.\n> * The bad object is pulled from github and deployed.\n> * The maintainer still has the good object. They may not notice the bad\n>  object is out there for a long time.\n>\n> Of course, it doesn't need to involve Github, and doesn't need to\n> rely on internal details of their deduplication[1];\n> that only let me publish the bad object under a psydonym.\n\nread that e-mail again, it covers the case where a central server gets a blob \nreplaced in it.\n\ntricking a maintainerinto accepting a file that contains huge amounts of binary \ndata in it is going to be a non-trivial task, and even after you trick them into \naccepting one bad file, you then need to replace the file they accepted with a \nnew one (breaking into github or assuming that github is putting both files into \nthe same repo, both of which are fairly unlikely)\n\nDavid Lang\n"},{"id":"312407","messageId":"20170223182147.hbsyxsmyijgkqu75@kitenet.net","threadId":"45202","inReplyTo":"CA+55aFzFEpi1crykZ33r9f7BsvLt_kiB-CHXOkuCAX=fd4BU-w@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2017-02-23T18:21:47Z","receivedAt":"2017-02-23T18:22:03Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Linus Torvalds wrote:\n> What you describe pretty much already requires a pre-image attack,\n> which the new attack is _not_.\n> \n> It's not clear that the \"good\" object can be anything sane.\n\nGenerate a regular commit object; use the entire commit object + NUL as the\nchosen prefix, and use the identical-prefix collision attack to generate\nthe colliding good/bad objects.\n\n(The size in git's object header is a minor complication. Set the size\nfield to something sufficiently large, and then pad out the colliding\nobjects to that size once they're generated.)\n\n-- \nsee shy jo\n"},{"id":"312408","messageId":"CA+55aFz98r7NC_3BW_1HU9-0C-HcrFou3=0gmRcS38=-x8dmVw@mail.gmail.com","threadId":"45202","inReplyTo":"20170223181018.ns4vyosgzmuoyiva@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-23T18:29:09Z","receivedAt":"2017-02-23T18:29:31Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 10:10 AM, Joey Hess <id@joeyh.name> wrote:\n>\n> It would cost 6500 CPU years + 100 GPU years to generate valid colliding\n> git objects using the methods of the paper's authors. That might be cost\n> effective if it helped get a backdoor into eg, the kernel.\n\nI still think it also needs to be interesting enough data, not just\nrandom noise that is then trivial to find with automated tools.\n\nBecause for the kernel, it's not just that an attacker needs to do the\nCPU time. Yes, first he needs the technical resources to just do just\nthe attack and create the situation you described.\n\nBut then he *also* needs to build up the social capital to get the end\nresult pulled into the tree (ie if he depends on the hidden spaces, he\nneeds somebody to actually do a git pull, not just apply a patch).\n\n.. and if we then have a tool that then finds the problem trivially\n(ie \"git fsck\"), he's not only wasted all those technical resources,\nhe's also burned his identity.\n\n>>  (b) we can probably easily add some extra sanity checks to the opaque\n>> data we do have, to make it much harder to do the hiding of random\n>> data that these attacks pretty much always depend on.\n>\n> For example, git fsck does warn about a commit message with opaque\n> data hidden after a NUL. But, git show/merge/pull give no indication\n> that something funky is going on when working with such commits.\n\nI do agree that we might want to do some of the fsck checks\nparticularly at fetch time. That's when doing checks is both relevant\nand cheap.\n\nSo we could do the opaque data checks, but we could/should probably\nalso add the attack pattern (\"disturbance vectors\") checks.\n\nAnd the thing is, adding those checks is really cheap, and basically\nmakes the whole attack vector pointless against git.\n\nBecause unlike some \"signing a pdf\" attack, git doesn't fundamentally\ndepend on the SHA1 as some kind of absolute security.  If we have the\nminimal machinery in git to just notice the attack, the attack\nessentially goes away. Attackers can waste infinite amounts of CPU\ntime, and if it's cheap for us to notice, it completely disarms all\nthat attack work.\n\nAgain, I'm not arguing that people shouldn't work on extending git to\na new (and bigger) hash. I think that's a no-brainer, and we do want\nto have a path to eventually move towards SHA3-256 or whatever.\n\nBut I'm very definitely arguing that the current attack doesn't\nactually sound like it really even _matters_, because it should be so\neasy to mitigate against.\n\n                   Linus\n"},{"id":"312409","messageId":"xmqqk28g92h7.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"20170223181018.ns4vyosgzmuoyiva@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-23T18:38:44Z","receivedAt":"2017-02-23T18:38:51Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Joey Hess <id@joeyh.name> writes:\n\n> For example, git fsck does warn about a commit message with opaque\n> data hidden after a NUL. But, git show/merge/pull give no indication\n> that something funky is going on when working with such commits.\n\nWould\n\n    $ git config transfer.fsckobjects true\n\nhelp?\n"},{"id":"312410","messageId":"CA+55aFxckeEW1ePcebrgG4iN4Lp62A2vU6tA=xnSDC_BnKQiCQ@mail.gmail.com","threadId":"45202","inReplyTo":"20170223182147.hbsyxsmyijgkqu75@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-23T18:40:48Z","receivedAt":"2017-02-23T18:40:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 10:21 AM, Joey Hess <id@joeyh.name> wrote:\n> Linus Torvalds wrote:\n>> What you describe pretty much already requires a pre-image attack,\n>> which the new attack is _not_.\n>>\n>> It's not clear that the \"good\" object can be anything sane.\n>\n> Generate a regular commit object; use the entire commit object + NUL as the\n> chosen prefix, and use the identical-prefix collision attack to generate\n> the colliding good/bad objects.\n\nSo I agree with you that we need to make git check for the opaque\ndata. I think I was the one who brought that whole argument up.\n\nBut even then, what you describe doesn't work. What you describe just\nreplaces the opaque data - that git doesn't actually *use*, and that\nnobody sees - with another piece of opaque data.\n\nYou also need to make the non-opaque data of the bad object be\nsomething that actually encodes valid git data with interesting hashes\nin it (for the parent/tree/whatever pointers).\n\nSo you don't have just that \"chosen prefix\". You actually need to also\nfill in some very specific piece of data *in* the attack parts itself.\nAnd you need to do this in the exact same size (because that's part of\nthe prefix), etc etc.\n\nSo I think it's challenging.\n\n... and then we can discover it trivially.\n\nOk, so \"git fsck\" right now takes a couple of minutes for me and I\ndon't actually run it very often (I used to run it religiously back in\nthe days), but afaik kernel.org actually runs it nightly. So it's\npretty much \"trivially discoverable\" - imagine spending thousands of\nCPU-hours and lots of social capital to get an attack in, and then the\nnext night the kernel.org fsck complains about the strange commit you\nadded?\n\n                  Linus\n"},{"id":"312411","messageId":"20170223184637.xr74k42vc6y2pmse@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CA+55aFxckeEW1ePcebrgG4iN4Lp62A2vU6tA=xnSDC_BnKQiCQ@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T18:46:37Z","receivedAt":"2017-02-23T18:46:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 23, 2017 at 10:40:48AM -0800, Linus Torvalds wrote:\n\n> > Generate a regular commit object; use the entire commit object + NUL as the\n> > chosen prefix, and use the identical-prefix collision attack to generate\n> > the colliding good/bad objects.\n> \n> So I agree with you that we need to make git check for the opaque\n> data. I think I was the one who brought that whole argument up.\n\nWe do already.\n\n> But even then, what you describe doesn't work. What you describe just\n> replaces the opaque data - that git doesn't actually *use*, and that\n> nobody sees - with another piece of opaque data.\n> \n> You also need to make the non-opaque data of the bad object be\n> something that actually encodes valid git data with interesting hashes\n> in it (for the parent/tree/whatever pointers).\n> \n> So you don't have just that \"chosen prefix\". You actually need to also\n> fill in some very specific piece of data *in* the attack parts itself.\n> And you need to do this in the exact same size (because that's part of\n> the prefix), etc etc.\n\nIt's not an identical prefix, but I think collision attacks generally\nare along the lines of selecting two prefixes followed by garbage, and\nthen mutating the garbage on both sides. That would \"work\" in this case\n(modulo the fact that git would complain about the NUL).\n\nI haven't read the paper yet to see if that is the case here, though.\n\nA related case is if you could stick a \"cruft ....\" header at the end of\nthe commit headers, and mutate its value (avoiding newlines). fsck\ndoesn't complain about that.\n\n-Peff\n"},{"id":"312414","messageId":"20170223183105.joxtpbut4wcqfbtu@kitenet.net","threadId":"45202","inReplyTo":"20170223182147.hbsyxsmyijgkqu75@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2017-02-23T18:31:05Z","receivedAt":"2017-02-23T18:57:47Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Joey Hess wrote:\n> Linus Torvalds wrote:\n> > What you describe pretty much already requires a pre-image attack,\n> > which the new attack is _not_.\n> > \n> > It's not clear that the \"good\" object can be anything sane.\n> \n> Generate a regular commit object; use the entire commit object + NUL as the\n> chosen prefix, and use the identical-prefix collision attack to generate\n> the colliding good/bad objects.\n> \n> (The size in git's object header is a minor complication. Set the size\n> field to something sufficiently large, and then pad out the colliding\n> objects to that size once they're generated.)\n\nSorry! While that would work, it's a useless attack because the good and bad\ncommit objects still point to the same tree.\n\nIt would be interesting to have such colliding objects, to see what beaks,\nbut probably not worth $75k to generate them.\n\n-- \nsee shy jo\n"},{"id":"312416","messageId":"20170223184213.rml33rymw6x6pirc@sigill.intra.peff.net","threadId":"45202","inReplyTo":"20170223182147.hbsyxsmyijgkqu75@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T18:42:13Z","receivedAt":"2017-02-23T19:09:01Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 23, 2017 at 02:21:47PM -0400, Joey Hess wrote:\n\n> Linus Torvalds wrote:\n> > What you describe pretty much already requires a pre-image attack,\n> > which the new attack is _not_.\n> > \n> > It's not clear that the \"good\" object can be anything sane.\n> \n> Generate a regular commit object; use the entire commit object + NUL as the\n> chosen prefix, and use the identical-prefix collision attack to generate\n> the colliding good/bad objects.\n\nFWIW, git-fsck complains about those (and transfer.fsck rejects them):\n\n  $ (git cat-file commit HEAD; printf '\\0more stuff') |\n    git hash-object -w --stdin -t commit\n  ecb2e5165c184f9025cb4c49d8f75901f4830354\n\n  $ git fsck\n  warning in commit ecb2e5165c184f9025cb4c49d8f75901f4830354: nulInCommit: NUL byte in the commit object body\n\nSo as long as either your \"good\" or \"evil\" commit has binary junk in it,\nyou are likely to be noticed (not everybody turns on transfer.fsck, but\nGitHub does).\n\n-Peff\n"},{"id":"312417","messageId":"CA+55aFx=0EVfSG2iEKKa78g3hFN_yZ+L_FRm4R749nNAmTGO9w@mail.gmail.com","threadId":"45202","inReplyTo":"20170223184637.xr74k42vc6y2pmse@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-23T19:09:32Z","receivedAt":"2017-02-23T19:09:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 10:46 AM, Jeff King <peff@peff.net> wrote:\n>>\n>> So I agree with you that we need to make git check for the opaque\n>> data. I think I was the one who brought that whole argument up.\n>\n> We do already.\n\nI'm aware of the fsck checks, but I have to admit I wasn't aware of\n'transfer.fsckobjects'. I should turn that on myself.\n\nOr maybe git should just turn it on by default? At least the\nper-object fsck costs should be essentially free compared to the\nnetwork costs when you just apply them to the incoming objects.\n\nI also do think that it would be good to check for the disturbance\nvectors at receive time (and fsck). Not necessarily interesting during\nnormal operations.\n\nAnd in particular, while the *kernel* doesn't generally have critical\nopaque blobs, other projects do. Things like firmware images etc are\nopen to attack, and crazy people put ISO images in repositories etc.\n\nSo I don't think this discussion should focus exclusively on the git metadata.\n\nIt is likely much easier to replace a binary blob than it is to\nreplace a commit or tree (or a source file that has to go through a\ncompiler). And for many projects, that would be a bad thing.\n\n> It's not an identical prefix, but I think collision attacks generally\n> are along the lines of selecting two prefixes followed by garbage, and\n> then mutating the garbage on both sides. That would \"work\" in this case\n> (modulo the fact that git would complain about the NUL).\n\nI think this particular attack depended on an actual identical prefix,\nbut I didn't go back to the paper and check.\n\nBut the attacks tend to very much depend on particular input bit\npatterns that have very particular effects on the resulting\nintermediate hash, and those bit patterns are specific to the hash and\nknown.\n\nSo a very powerful defense is to just look for those bit patterns in\nthe objects, and just warn about them. Those patterns don't tend to\nexist in normal inputs anyway, but particularly if you just warn, it's\na heads-ups that \"ok, something iffy is going on\"\n\nAnd as mentioned, a cheap \"something iffy is going on\" thing is\nbasically a death sentence to SCM attacks.\n\nThe whole _point_ of an SCM is that it isn't about a one-time event,\nbut about continuous history. That also fundamentally means that a\nsuccessful attack needs to work over time, and not be detectable.\n\nIn contrast, many other uses of hashes are \"one-time\" events.  If you\nuse a hash to validate a single piece of data from a source that you\nwouldn't otherwise trust, it's a one-time \"all or nothing\" trust\nsituation.\n\nAnd the attack surface is very different for those \"one-time\" vs\n\"trust over time\" cases. If you can get a bank to trust a session one\ntime, you can empty a bank account and live on a paradise island for\nthe rest of your life. It doesn't matter if it gets detected or not\nafter-the-fact.\n\nBut if you can fool a SCM one time, insert your code, and it gets\ndetected next week, you didn't actually do anything useful. You only\nburned yourself.\n\nSee the difference? One-time vs having a continual interaction makes a\n*fundamntal* difference in game theory.\n\n                Linus\n"},{"id":"312419","messageId":"CANv4PNmSjJUhFgC7GhpuBjiSQhwfAhrP8WxiP_siP2AjjXnrnw@mail.gmail.com","threadId":"45202","inReplyTo":"20170223183105.joxtpbut4wcqfbtu@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Morten Welinder","fromEmail":"mwelinder@gmail.com","sentAt":"2017-02-23T19:13:31Z","receivedAt":"2017-02-23T19:13:48Z","isPatch":false,"sender":{"key":"mwelinder@gmail.com","avatar":null},"body":"The attack seems to generate two 64-bytes blocks, one quarter of which\nis repeated data.  (Table-1 in the paper.)\n\nAssuming the result of that is evenly distributed and that bytes are\nindependent, we can estimate the chances that the result is NUL-free\nas (255/256)^192 = 47% and the probability that the result is NUL and\nnewline free as (254/256)^192 = 22%.  Clearly one should not rely of\nNULs or newlines to save the day.  On  the other hand, the chances of\nan ascii result is something like (95/256)^192 = 10^-83.\n\nThe actual collision in the paper has no newline, but it does have a NUL.\n\nM.\n\n\n\n\nOn Thu, Feb 23, 2017 at 1:31 PM, Joey Hess <id@joeyh.name> wrote:\n> Joey Hess wrote:\n>> Linus Torvalds wrote:\n>> > What you describe pretty much already requires a pre-image attack,\n>> > which the new attack is _not_.\n>> >\n>> > It's not clear that the \"good\" object can be anything sane.\n>>\n>> Generate a regular commit object; use the entire commit object + NUL as the\n>> chosen prefix, and use the identical-prefix collision attack to generate\n>> the colliding good/bad objects.\n>>\n>> (The size in git's object header is a minor complication. Set the size\n>> field to something sufficiently large, and then pad out the colliding\n>> objects to that size once they're generated.)\n>\n> Sorry! While that would work, it's a useless attack because the good and bad\n> commit objects still point to the same tree.\n>\n> It would be interesting to have such colliding objects, to see what beaks,\n> but probably not worth $75k to generate them.\n>\n> --\n> see shy jo\n"},{"id":"312421","messageId":"nycvar.QRO.7.75.62.1702231113450.6590@qynat-yncgbc","threadId":"45202","inReplyTo":"CAPc5daVZ79WWKSw76kxHgDra9a7fSR1AibZa_pvK9aUuuVawLQ@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2017-02-23T19:20:34Z","receivedAt":"2017-02-23T19:22:04Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"\npointers to a little more info\n\n\nhttps://shattered.it/static/\nthe two files are:\n\nhttps://shattered.it/static/shattered-1.pdf\nhttps://shattered.it/static/shattered-2.pdf\n\n422435 shattered-2.pdf\n422435 shattered-1.pdf\n\nidentical length and a lot smaller than I expected (~162K of the 413K file is \nbinary junk)\n\n\n$ sha1sum shattered-*pdf\n38762cf7f55934b34d179ae6a4c80cadccbb7f0a  shattered-1.pdf\n38762cf7f55934b34d179ae6a4c80cadccbb7f0a  shattered-2.pdf\n\n$ sum shattered-*pdf\n62721   413 shattered-1.pdf\n41606   413 shattered-2.pdf\n\n$ md5sum shattered-*pdf\nee4aa52b139d925f8d8884402b0a750c  shattered-1.pdf\n5bd9d8cabc46041579a311230539b8d1  shattered-2.pdf\n\nDavid Lang\n"},{"id":"312426","messageId":"CA+55aFxmr6ntWGbJDa8tOyxXDX3H-yd4TQthgV_Tn1u91yyT8w@mail.gmail.com","threadId":"45202","inReplyTo":"20170223193210.munuqcjltwbrdy22@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-23T19:47:16Z","receivedAt":"2017-02-23T19:47:22Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 11:32 AM, Jeff King <peff@peff.net> wrote:\n>\n> Yeah, they're not expensive. We've discussed enabling them by default.\n> The sticking point is that there is old history with minor bugs which\n> triggers some warnings (e.g., malformed committer names), and it would\n> be annoying to start rejecting that unconditionally.\n>\n> So I think we would need a good review of what is a \"warning\" versus an\n> \"error\", and to only reject on errors (right now the NUL thing is a\n> warning, and it should probably upgraded).\n\nI think even a warning (as opposed to failing the operation) is\nalready a big deal.\n\nIf people start saying \"why do I get this odd warning\", and start\nlooking into it, that's going to be a pretty strong defense against\nbad behavior. SCM attacks depend on flying under the radar.\n\n>> So a very powerful defense is to just look for those bit patterns in\n>> the objects, and just warn about them. Those patterns don't tend to\n>> exist in normal inputs anyway, but particularly if you just warn, it's\n>> a heads-ups that \"ok, something iffy is going on\"\n>\n> Yes, that would be a wonderful hardening to put into Git if we know what\n> those patterns look like. That part isn't clear to me.\n\nThere's actually already code for that, pointed to by the shattered project:\n\n  https://github.com/cr-marcstevens/sha1collisiondetection\n\nthe \"meat\" of that check is in lib/ubc_check.c.\n\n                  Linus\n"},{"id":"312427","messageId":"20170223195753.ppsat2gwd3jq22by@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CA+55aFxmr6ntWGbJDa8tOyxXDX3H-yd4TQthgV_Tn1u91yyT8w@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T19:57:54Z","receivedAt":"2017-02-23T19:58:06Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 23, 2017 at 11:47:16AM -0800, Linus Torvalds wrote:\n\n> On Thu, Feb 23, 2017 at 11:32 AM, Jeff King <peff@peff.net> wrote:\n> >\n> > Yeah, they're not expensive. We've discussed enabling them by default.\n> > The sticking point is that there is old history with minor bugs which\n> > triggers some warnings (e.g., malformed committer names), and it would\n> > be annoying to start rejecting that unconditionally.\n> >\n> > So I think we would need a good review of what is a \"warning\" versus an\n> > \"error\", and to only reject on errors (right now the NUL thing is a\n> > warning, and it should probably upgraded).\n> \n> I think even a warning (as opposed to failing the operation) is\n> already a big deal.\n> \n> If people start saying \"why do I get this odd warning\", and start\n> looking into it, that's going to be a pretty strong defense against\n> bad behavior. SCM attacks depend on flying under the radar.\n\nSorry, I conflated two things there. I agree a warning is better than\nnothing. But right now transfer.fsck croaks even for warnings, and there\nare some warnings that it is not worth croaking for. So before we turn\nit on, we need to stop croaking on warnings (and possibly bump up some\nwarnings to errors).\n\nI think it _is_ important to have dangerous things as errors, though.\nBecause it helps an unattended server (where nobody would see the\nwarning) avoid being a vector for spreading malicious objects to older\nclients which do not do the fsck.\n\n> There's actually already code for that, pointed to by the shattered project:\n> \n>   https://github.com/cr-marcstevens/sha1collisiondetection\n> \n> the \"meat\" of that check is in lib/ubc_check.c.\n\nThanks, I hadn't seen that yet. That doesn't look like it should be hard\nto integrate into Git.\n\n-Peff\n"},{"id":"312428","messageId":"20170223193210.munuqcjltwbrdy22@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CA+55aFx=0EVfSG2iEKKa78g3hFN_yZ+L_FRm4R749nNAmTGO9w@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T19:32:10Z","receivedAt":"2017-02-23T19:58:57Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 23, 2017 at 11:09:32AM -0800, Linus Torvalds wrote:\n\n> On Thu, Feb 23, 2017 at 10:46 AM, Jeff King <peff@peff.net> wrote:\n> >>\n> >> So I agree with you that we need to make git check for the opaque\n> >> data. I think I was the one who brought that whole argument up.\n> >\n> > We do already.\n> \n> I'm aware of the fsck checks, but I have to admit I wasn't aware of\n> 'transfer.fsckobjects'. I should turn that on myself.\n> \n> Or maybe git should just turn it on by default? At least the\n> per-object fsck costs should be essentially free compared to the\n> network costs when you just apply them to the incoming objects.\n\nYeah, they're not expensive. We've discussed enabling them by default.\nThe sticking point is that there is old history with minor bugs which\ntriggers some warnings (e.g., malformed committer names), and it would\nbe annoying to start rejecting that unconditionally.\n\nSo I think we would need a good review of what is a \"warning\" versus an\n\"error\", and to only reject on errors (right now the NUL thing is a\nwarning, and it should probably upgraded).\n\n> And in particular, while the *kernel* doesn't generally have critical\n> opaque blobs, other projects do. Things like firmware images etc are\n> open to attack, and crazy people put ISO images in repositories etc.\n> \n> So I don't think this discussion should focus exclusively on the git metadata.\n> \n> It is likely much easier to replace a binary blob than it is to\n> replace a commit or tree (or a source file that has to go through a\n> compiler). And for many projects, that would be a bad thing.\n\nYes, I'd agree we need to consider both. And no matter what Git does in\nits own data formats, blobs will always be a sequence of bytes. Hiding\ncollision-cruft in them isn't up to us, but rather the data format.\n\nThe nice thing about a blob collision, though, is that you can only\nreplace the opaque files, not, say, C source code. That doesn't make it\na non-issue, but it reduces the scope of an attack.\n\nReplacing a commit or tree wholesale means the attacker has a lot more\nflexibility. So to whatever degree we can make that harder (like\ncomplaining of commits with NULs), the better.\n\n> > It's not an identical prefix, but I think collision attacks generally\n> > are along the lines of selecting two prefixes followed by garbage, and\n> > then mutating the garbage on both sides. That would \"work\" in this case\n> > (modulo the fact that git would complain about the NUL).\n> \n> I think this particular attack depended on an actual identical prefix,\n> but I didn't go back to the paper and check.\n\nThe paper describes the content as:\n\n  SHA-1(P | M1 | M2 | S)\n\nand they replace both \"M1\" and \"M2\", with a near-collision for the\nfirst, and then the final collision for the second. What's not clear to\nme is if part of M1 can be chosen, or if it's perturbed fully into\nrandom garbage.\n\n> But the attacks tend to very much depend on particular input bit\n> patterns that have very particular effects on the resulting\n> intermediate hash, and those bit patterns are specific to the hash and\n> known.\n> \n> So a very powerful defense is to just look for those bit patterns in\n> the objects, and just warn about them. Those patterns don't tend to\n> exist in normal inputs anyway, but particularly if you just warn, it's\n> a heads-ups that \"ok, something iffy is going on\"\n\nYes, that would be a wonderful hardening to put into Git if we know what\nthose patterns look like. That part isn't clear to me.\n\n> The whole _point_ of an SCM is that it isn't about a one-time event,\n> but about continuous history. That also fundamentally means that a\n> successful attack needs to work over time, and not be detectable.\n\nYeah, I'd certainly agree with that. You spend loads of money to\ngenerate a collision, there's a reasonably high chance of detection, and\nthen as soon as one person detects it, your investment is lost.\n\nAccording to the paper, the current cost of the computation for a single\ncollision is ~$670K.\n\nAt least for now, an attacker is much better off using that money to\nbreak into your house and install a keylogger.\n\n-Peff\n"},{"id":"312434","messageId":"20170223204739.f6aqri3l2fydxe2b@sunbase.org","threadId":"45202","inReplyTo":"CA+55aFx=0EVfSG2iEKKa78g3hFN_yZ+L_FRm4R749nNAmTGO9w@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Øyvind A. Holm","fromEmail":"sunny@sunbase.org","sentAt":"2017-02-23T20:47:40Z","receivedAt":"2017-02-23T20:48:24Z","isPatch":false,"sender":{"key":"sunny@sunbase.org","avatar":"https://avatars.githubusercontent.com/u/113445?v=4"},"body":"On 2017-02-23 11:09:32, Linus Torvalds wrote:\n> I'm aware of the fsck checks, but I have to admit I wasn't aware of \n> 'transfer.fsckobjects'. I should turn that on myself.\n>\n> Or maybe git should just turn it on by default?\n\nThe problem with this is that there are many repos with errors out \nthere, for example coreutils.git and nasm.git, which complains about \n\"missingSpaceBeforeDate: invalid author/committer line - missing space \nbefore date\".\n\nThere are also lots of repositories bitten by the Github bug from back \nin 2011 where they zero-padded the file modes, git clone aborts with \n\"zeroPaddedFilemode: contains zero-padded file modes\".\n\nParanoid as I am, I'm using fetch.fsckObjects and receive.fsckObjects \nset to \"true\", but that means I'm not able to clone repositories with \nthese kind of errors, have to use the alias\n\n  fclone = clone -c \"fetch.fsckObjects=false\"\n\nSo enabling them by default will create problems among users. Of course, \none solution would be to turn these kind of errors into warnings so the \nclone isn't aborted.\n\nReagards,\nØyvind\n\n+-| Øyvind A. Holm <sunny@sunbase.org> - N 60.37604° E 5.33339° |-+\n| OpenPGP: 0xFB0CBEE894A506E5 - http://www.sunbase.org/pubkey.asc |\n| Fingerprint: A006 05D6 E676 B319 55E2  E77E FB0C BEE8 94A5 06E5 |\n+------------| c7e47a18-fa06-11e6-ad93-db5caa6d21d3 |-------------+\n"},{"id":"312436","messageId":"e57958d4-7c51-3f5e-6ff5-f863920fd883@gmail.com","threadId":"45202","inReplyTo":"nycvar.QRO.7.75.62.1702230907340.6590@qynat-yncgbc","subject":"Re: SHA1 collisions found","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2017-02-23T20:49:09Z","receivedAt":"2017-02-23T20:49:22Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 23.02.2017 o 18:12, David Lang pisze:\n> On Thu, 23 Feb 2017, Junio C Hamano wrote:\n> \n>> On Thu, Feb 23, 2017 at 8:43 AM, Joey Hess <id@joeyh.name> wrote:\n>>> \n>>> Since we now have collisions in valid PDF files, collisions in\n>>> valid git commit and tree objects are probably able to be\n>>> constructed.\n>> \n>> That may be true, but \n>> https://public-inbox.org/git/Pine.LNX.4.58.0504291221250.18901@ppc970.osdl.org/\n>>\n>\n> it doesn't help that the Google page on this explicitly says that\n> this shows that it's possible to create two different git repos that\n> have the same hash but different contents.\n> \n> https://shattered.it/\n> \n> How is GIT affected? GIT strongly relies on SHA-1 for the\n> identification and integrity checking of all file objects and\n> commits. It is essentially possible to create two GIT repositories\n> with the same head commit hash and different contents, say a benign\n> source code and a backdoored one. An attacker could potentially\n> selectively serve either repository to targeted users. This will\n> require attackers to compute their own collision.\n\nThe attack on SHA-1 presented there is \"identical-prefix\" collision,\nwhich is less powerful than \"chosen-prefix\" collision.  It is the\nlatter that is required to defeat SHA-1 used in object identity.\nObjects in Git _must_ begin with given prefix; the use of zlib\ncompression adds to the difficulty.  'Forged' Git object would\nsimply not validate...\n\nhttps://arstechnica.com/security/2017/02/at-deaths-door-for-years-widely-used-sha1-function-is-now-dead/\n\n"},{"id":"312438","messageId":"20170223204659.prmbuip4s4snwtci@kitenet.net","threadId":"45202","inReplyTo":"20170223184637.xr74k42vc6y2pmse@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2017-02-23T20:46:59Z","receivedAt":"2017-02-23T20:58:13Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Jeff King wrote:\n> It's not an identical prefix, but I think collision attacks generally\n> are along the lines of selecting two prefixes followed by garbage, and\n> then mutating the garbage on both sides. That would \"work\" in this case\n> (modulo the fact that git would complain about the NUL).\n> \n> I haven't read the paper yet to see if that is the case here, though.\n\nThe current attack is an identical-prefix attack, not chosen-prefix, so\nnot quite to that point yet.\n\nThe MD5 chosen-prefix attack was 2^15 harder than the known-prefix attack,\nbut who knows if the numbers will be comprable for SHA1.\n\n> A related case is if you could stick a \"cruft ....\" header at the end of\n> the commit headers, and mutate its value (avoiding newlines). fsck\n> doesn't complain about that.\n\ngit log and git show don't show such cruft headers either.\n\nBTW, the SHA attack only added ~128 bytes to the pdfs, not really a\nhuge amount of garbage.\n\n-- \nsee shy jo\n"},{"id":"312439","messageId":"20170223205739.t4kekrp2kb7zkimv@sigill.intra.peff.net","threadId":"45202","inReplyTo":"e57958d4-7c51-3f5e-6ff5-f863920fd883@gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T20:57:39Z","receivedAt":"2017-02-23T20:58:36Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 23, 2017 at 09:49:09PM +0100, Jakub Narębski wrote:\n\n> > How is GIT affected? GIT strongly relies on SHA-1 for the\n> > identification and integrity checking of all file objects and\n> > commits. It is essentially possible to create two GIT repositories\n> > with the same head commit hash and different contents, say a benign\n> > source code and a backdoored one. An attacker could potentially\n> > selectively serve either repository to targeted users. This will\n> > require attackers to compute their own collision.\n> \n> The attack on SHA-1 presented there is \"identical-prefix\" collision,\n> which is less powerful than \"chosen-prefix\" collision.  It is the\n> latter that is required to defeat SHA-1 used in object identity.\n> Objects in Git _must_ begin with given prefix;\n\nI don't think this helps. The chosen-prefix lets you append hash data to\nan existing file. Here we just have identical prefixes in the two\ncolliding halves. In the real-world example, they used a PDF header. But\nit could have been a PDF header with \"blob 1234\" prepended to it (note\nalso that Git's use of the size doesn't help; the attack files are the\nsame length).\n\n> the use of zlib\n> compression adds to the difficulty.  'Forged' Git object would\n> simply not validate...\n\nNo, zlib doesn't help. The sha1 is computed on the uncompressed data.\n\n-Peff\n"},{"id":"312446","messageId":"20170223224302.joti4zqucme3vqr2@sigill.intra.peff.net","threadId":"45202","inReplyTo":"alpine.LFD.2.20.1702231428540.30435@i7.lan","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T22:43:02Z","receivedAt":"2017-02-23T22:49:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 23, 2017 at 02:38:29PM -0800, Linus Torvalds wrote:\n\n> > Thanks, I hadn't seen that yet. That doesn't look like it should be hard\n> > to integrate into Git.\n> \n> Here's a *very* ugly patch that is absolutely disgusting and should not be \n> used. But it does kind of work (I tested it with a faked-up extra patch \n> that made git accept the broken pdf as a loose object).\n> \n> What do I mean by \"kind of work\"? It uses that ugly and slow checking \n> SHA1 routine from the collision detection project for the SHA1 object \n> verification, and it means that \"git fsck\" ends up being about twice as \n> slow as it used to be.\n\nHeh. I was just putting the finishing touches on a similar patch. Mine\nis much less gross, in that it actually just adds a new USE_SHA1DC knob\n(instead of, say, BLK_SHA1).\n\nHere are the timings I came up with:\n\n  - compute sha1 over whole packfile\n    before: 1.349s\n     after: 5.067s\n    change: +275%\n\n  - rev-list --all\n    before: 5.742s\n     after: 5.730s\n    change: -0.2%\n\n  - rev-list --all --objects\n    before: 33.257s\n     after: 33.392s\n    change: +0.4%\n\n  - index-pack --verify\n    before: 2m20s\n     after: 5m43s\n    change: +145%\n\n  - git log --no-merges -10000 -p\n    before: 9.532s\n     after: 9.683s\n    change: +1.5%\n\nSo overall the sha1 computation is about 3-4x slower. But of\ncourse most operations do more than just sha1. Accessing\ncommits and trees isn't slowed at all (both the +/- changes\nthere are well within the run-to-run noise). Accessing the\nblobs is a little slower, but mostly drowned out by the cost\nof things like actually generating patches.\n\nThe most-affected operation is `index-pack --verify`, which\nis essentially just computing the sha1 on every object. It's\na bit worse than twice as slow, which means every push and\nevery fetch is going to experience that.\n\n> For example, I suspect we could use our (much cleaner) block-sha1 \n> implementation and include just the ubc_check.c code with that, instead of \n> the truly ugly C sha1 implementation that the sha1collisiondetection \n> project uses. \n> \n> But to do that, somebody would have to really know how the unavoidable \n> bit conditions check works with the intermediate hashes. I have only a \n> \"big picture\" mental model of it (read: I'm not competent to do that).\n\nYeah. I started looking at that, but the ubc check happens after the\ninitial expansion. But AFAICT, block-sha1 mixes that expansion in with\nthe rest of the steps for efficiency. So perhaps somebody who really\nunderstands sha1 and the new checks could figure it out, but I'm not at\nall certain that adding it in wouldn't lose some of block-sha1's\nefficiency (on top of the time to actually do the ubc check).\n\n-Peff\n"},{"id":"312447","messageId":"CA+55aFyfWVHYMC+sSyct=uNLK=mHV-NqyNQXXtFY5YX1Uc-tAw@mail.gmail.com","threadId":"45202","inReplyTo":"20170223224302.joti4zqucme3vqr2@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-23T22:50:26Z","receivedAt":"2017-02-23T22:50:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 2:43 PM, Jeff King <peff@peff.net> wrote:\n>\n> Yeah. I started looking at that, but the ubc check happens after the\n> initial expansion.\n\nYes. That's the point where I gave up and just included their ugly sha1.c file.\n\nI suspect it can be done, but it would need somebody to really know\nwhat they are doing.\n\n            Linus\n"},{"id":"312458","messageId":"20170223230507.kuxjqtg3ghcfskc6@sigill.intra.peff.net","threadId":"45202","inReplyTo":"20170223224302.joti4zqucme3vqr2@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T23:05:07Z","receivedAt":"2017-02-23T23:05:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 23, 2017 at 05:43:02PM -0500, Jeff King wrote:\n\n> On Thu, Feb 23, 2017 at 02:38:29PM -0800, Linus Torvalds wrote:\n> \n> > > Thanks, I hadn't seen that yet. That doesn't look like it should be hard\n> > > to integrate into Git.\n> > \n> > Here's a *very* ugly patch that is absolutely disgusting and should not be \n> > used. But it does kind of work (I tested it with a faked-up extra patch \n> > that made git accept the broken pdf as a loose object).\n> > \n> > What do I mean by \"kind of work\"? It uses that ugly and slow checking \n> > SHA1 routine from the collision detection project for the SHA1 object \n> > verification, and it means that \"git fsck\" ends up being about twice as \n> > slow as it used to be.\n> \n> Heh. I was just putting the finishing touches on a similar patch. Mine\n> is much less gross, in that it actually just adds a new USE_SHA1DC knob\n> (instead of, say, BLK_SHA1).\n\nHere's my patches. They _might_ be worth including if only because they\nshouldn't bother anybody unless they enable USE_SHA1DC. So it makes it a\nbit more accessible for people to experiment with (or be paranoid with\nif they like).\n\nThe first one is 98K. Mail headers may bump it over vger's 100K barrier.\nIt's actually the _least_ interesting patch of the 3, because it just\nimports the code wholesale from the other project. But if it doesn't\nmake it, you can fetch the whole series from:\n\n  https://github.com/peff/git jk/sha1dc\n\n(By the way, I don't see your version on the list, Linus, which probably\nmeans it was eaten by the 100K filter).\n\n  [1/3]: add collision-detecting sha1 implementation\n  [2/3]: sha1dc: adjust header includes for git\n  [3/3]: Makefile: add USE_SHA1DC knob\n\n Makefile           |   10 +\n sha1dc/sha1.c      | 1165 ++++++++++++++++++++++++++++++++++++++++++++++++++++\n sha1dc/sha1.h      |  108 +++++\n sha1dc/ubc_check.c |  361 ++++++++++++++++\n sha1dc/ubc_check.h |   33 ++\n 5 files changed, 1677 insertions(+)\n create mode 100644 sha1dc/sha1.c\n create mode 100644 sha1dc/sha1.h\n create mode 100644 sha1dc/ubc_check.c\n create mode 100644 sha1dc/ubc_check.h\n\n-Peff\n"},{"id":"312460","messageId":"20170223230536.tdmtsn46e4lnrimx@sigill.intra.peff.net","threadId":"45202","inReplyTo":"20170223230507.kuxjqtg3ghcfskc6@sigill.intra.peff.net","subject":"[PATCH 1/3] add collision-detecting sha1 implementation","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T23:05:37Z","receivedAt":"2017-02-23T23:05:50Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"This is pulled straight from:\n\n  https://github.com/cr-marcstevens/sha1collisiondetection\n\nwith no modifications yet (though I've pulled in only the\nsubset of files necessary for Git to use).\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1dc/sha1.c      | 1146 ++++++++++++++++++++++++++++++++++++++++++++++++++++\n sha1dc/sha1.h      |   94 +++++\n sha1dc/ubc_check.c |  361 +++++++++++++++++\n sha1dc/ubc_check.h |   35 ++\n 4 files changed, 1636 insertions(+)\n create mode 100644 sha1dc/sha1.c\n create mode 100644 sha1dc/sha1.h\n create mode 100644 sha1dc/ubc_check.c\n create mode 100644 sha1dc/ubc_check.h\n\ndiff --git a/sha1dc/sha1.c b/sha1dc/sha1.c\nnew file mode 100644\nindex 000000000..ed2010911\n--- /dev/null\n+++ b/sha1dc/sha1.c\n@@ -0,0 +1,1146 @@\n+/***\n+* Copyright 2017 Marc Stevens <marc@marc-stevens.nl>, Dan Shumow (danshu@microsoft.com) \n+* Distributed under the MIT Software License.\n+* See accompanying file LICENSE.txt or copy at\n+* https://opensource.org/licenses/MIT\n+***/\n+\n+#include <string.h>\n+#include <memory.h>\n+#include <stdio.h>\n+\n+#include \"sha1.h\"\n+#include \"ubc_check.h\"\n+\n+#define rotate_right(x,n) (((x)>>(n))|((x)<<(32-(n))))\n+#define rotate_left(x,n)  (((x)<<(n))|((x)>>(32-(n))))\n+\n+#define sha1_f1(b,c,d) ((d)^((b)&((c)^(d))))\n+#define sha1_f2(b,c,d) ((b)^(c)^(d))\n+#define sha1_f3(b,c,d) (((b) & ((c)|(d))) | ((c)&(d)))\n+#define sha1_f4(b,c,d) ((b)^(c)^(d))\n+\n+#define HASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, m, t) \\\n+\t{ e += rotate_left(a, 5) + sha1_f1(b,c,d) + 0x5A827999 + m[t]; b = rotate_left(b, 30); }\n+#define HASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, m, t) \\\n+\t{ e += rotate_left(a, 5) + sha1_f2(b,c,d) + 0x6ED9EBA1 + m[t]; b = rotate_left(b, 30); }\n+#define HASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, m, t) \\\n+\t{ e += rotate_left(a, 5) + sha1_f3(b,c,d) + 0x8F1BBCDC + m[t]; b = rotate_left(b, 30); }\n+#define HASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, m, t) \\\n+\t{ e += rotate_left(a, 5) + sha1_f4(b,c,d) + 0xCA62C1D6 + m[t]; b = rotate_left(b, 30); }\n+\n+#define HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(a, b, c, d, e, m, t) \\\n+\t{ b = rotate_right(b, 30); e -= rotate_left(a, 5) + sha1_f1(b,c,d) + 0x5A827999 + m[t]; }\n+#define HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(a, b, c, d, e, m, t) \\\n+\t{ b = rotate_right(b, 30); e -= rotate_left(a, 5) + sha1_f2(b,c,d) + 0x6ED9EBA1 + m[t]; }\n+#define HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(a, b, c, d, e, m, t) \\\n+\t{ b = rotate_right(b, 30); e -= rotate_left(a, 5) + sha1_f3(b,c,d) + 0x8F1BBCDC + m[t]; }\n+#define HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(a, b, c, d, e, m, t) \\\n+\t{ b = rotate_right(b, 30); e -= rotate_left(a, 5) + sha1_f4(b,c,d) + 0xCA62C1D6 + m[t]; }\n+\n+#define SHA1_STORE_STATE(i) states[i][0] = a; states[i][1] = b; states[i][2] = c; states[i][3] = d; states[i][4] = e;\n+\n+\n+\n+void sha1_message_expansion(uint32_t W[80])\n+{\n+\tfor (unsigned i = 16; i < 80; ++i)\n+\t\tW[i] = rotate_left(W[i - 3] ^ W[i - 8] ^ W[i - 14] ^ W[i - 16], 1);\n+}\n+\n+void sha1_compression(uint32_t ihv[5], const uint32_t m[16])\n+{\n+\tuint32_t W[80];\n+\n+\tmemcpy(W, m, 16 * 4);\n+\tfor (unsigned i = 16; i < 80; ++i)\n+\t\tW[i] = rotate_left(W[i - 3] ^ W[i - 8] ^ W[i - 14] ^ W[i - 16], 1);\n+\n+\tuint32_t a = ihv[0], b = ihv[1], c = ihv[2], d = ihv[3], e = ihv[4];\n+\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 0);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 1);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 2);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 3);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 4);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 5);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 6);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 7);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 8);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 9);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 10);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 11);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 12);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 13);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 14);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 15);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 16);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 17);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 18);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 19);\n+\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 20);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 21);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 22);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 23);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 24);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 25);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 26);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 27);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 28);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 29);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 30);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 31);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 32);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 33);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 34);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 35);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 36);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 37);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 38);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 39);\n+\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 40);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 41);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 42);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 43);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 44);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 45);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 46);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 47);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 48);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 49);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 50);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 51);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 52);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 53);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 54);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 55);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 56);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 57);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 58);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 59);\n+\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 60);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 61);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 62);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 63);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 64);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 65);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 66);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 67);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 68);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 69);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 70);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 71);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 72);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 73);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 74);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 75);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 76);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 77);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 78);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 79);\n+\n+\tihv[0] += a; ihv[1] += b; ihv[2] += c; ihv[3] += d; ihv[4] += e;\n+}\n+\n+\n+\n+void sha1_compression_W(uint32_t ihv[5], const uint32_t W[80])\n+{\n+\tuint32_t a = ihv[0], b = ihv[1], c = ihv[2], d = ihv[3], e = ihv[4];\n+\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 0);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 1);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 2);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 3);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 4);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 5);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 6);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 7);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 8);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 9);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 10);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 11);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 12);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 13);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 14);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 15);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 16);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 17);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 18);\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 19);\n+\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 20);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 21);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 22);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 23);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 24);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 25);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 26);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 27);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 28);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 29);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 30);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 31);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 32);\n+ \tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 33);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 34);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 35);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 36);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 37);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 38);\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 39);\n+\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 40);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 41);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 42);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 43);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 44);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 45);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 46);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 47);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 48);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 49);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 50);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 51);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 52);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 53);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 54);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 55);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 56);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 57);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 58);\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 59);\n+\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 60);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 61);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 62);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 63);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 64);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 65);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 66);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 67);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 68);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 69);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 70);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 71);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 72);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 73);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 74);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 75);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 76);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 77);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 78);\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 79);\n+\n+\tihv[0] += a; ihv[1] += b; ihv[2] += c; ihv[3] += d; ihv[4] += e;\n+}\n+\n+\n+\n+void sha1_compression_states(uint32_t ihv[5], const uint32_t W[80], uint32_t states[80][5])\n+{\n+\tuint32_t a = ihv[0], b = ihv[1], c = ihv[2], d = ihv[3], e = ihv[4];\n+\n+#ifdef DOSTORESTATE00\n+\tSHA1_STORE_STATE(0)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 0);\n+\n+#ifdef DOSTORESTATE01\n+\tSHA1_STORE_STATE(1)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 1);\n+\n+#ifdef DOSTORESTATE02\n+\tSHA1_STORE_STATE(2)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 2);\n+\n+#ifdef DOSTORESTATE03\n+\tSHA1_STORE_STATE(3)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 3);\n+\n+#ifdef DOSTORESTATE04\n+\tSHA1_STORE_STATE(4)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 4);\n+\n+#ifdef DOSTORESTATE05\n+\tSHA1_STORE_STATE(5)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 5);\n+\n+#ifdef DOSTORESTATE06\n+\tSHA1_STORE_STATE(6)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 6);\n+\n+#ifdef DOSTORESTATE07\n+\tSHA1_STORE_STATE(7)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 7);\n+\n+#ifdef DOSTORESTATE08\n+\tSHA1_STORE_STATE(8)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 8);\n+\n+#ifdef DOSTORESTATE09\n+\tSHA1_STORE_STATE(9)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 9);\n+\n+#ifdef DOSTORESTATE10\n+\tSHA1_STORE_STATE(10)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 10);\n+\n+#ifdef DOSTORESTATE11\n+\tSHA1_STORE_STATE(11)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 11);\n+\n+#ifdef DOSTORESTATE12\n+\tSHA1_STORE_STATE(12)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 12);\n+\n+#ifdef DOSTORESTATE13\n+\tSHA1_STORE_STATE(13)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 13);\n+\n+#ifdef DOSTORESTATE14\n+\tSHA1_STORE_STATE(14)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 14);\n+\n+#ifdef DOSTORESTATE15\n+\tSHA1_STORE_STATE(15)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, W, 15);\n+\n+#ifdef DOSTORESTATE16\n+\tSHA1_STORE_STATE(16)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, W, 16);\n+\n+#ifdef DOSTORESTATE17\n+\tSHA1_STORE_STATE(17)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, W, 17);\n+\n+#ifdef DOSTORESTATE18\n+\tSHA1_STORE_STATE(18)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, W, 18);\n+\n+#ifdef DOSTORESTATE19\n+\tSHA1_STORE_STATE(19)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, W, 19);\n+\n+\n+\n+#ifdef DOSTORESTATE20\n+\tSHA1_STORE_STATE(20)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 20);\n+\n+#ifdef DOSTORESTATE21\n+\tSHA1_STORE_STATE(21)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 21);\n+\t\n+#ifdef DOSTORESTATE22\n+\tSHA1_STORE_STATE(22)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 22);\n+\t\n+#ifdef DOSTORESTATE23\n+\tSHA1_STORE_STATE(23)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 23);\n+\n+#ifdef DOSTORESTATE24\n+\tSHA1_STORE_STATE(24)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 24);\n+\n+#ifdef DOSTORESTATE25\n+\tSHA1_STORE_STATE(25)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 25);\n+\n+#ifdef DOSTORESTATE26\n+\tSHA1_STORE_STATE(26)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 26);\n+\n+#ifdef DOSTORESTATE27\n+\tSHA1_STORE_STATE(27)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 27);\n+\t\n+#ifdef DOSTORESTATE28\n+\tSHA1_STORE_STATE(28)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 28);\n+\t\n+#ifdef DOSTORESTATE29\n+\tSHA1_STORE_STATE(29)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 29);\n+\t\n+#ifdef DOSTORESTATE30\n+\tSHA1_STORE_STATE(30)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 30);\n+\t\n+#ifdef DOSTORESTATE31\n+\tSHA1_STORE_STATE(31)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 31);\n+\t\n+#ifdef DOSTORESTATE32\n+\tSHA1_STORE_STATE(32)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 32);\n+\n+#ifdef DOSTORESTATE33\n+\tSHA1_STORE_STATE(33)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 33);\n+\n+#ifdef DOSTORESTATE34\n+\tSHA1_STORE_STATE(34)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 34);\n+\n+#ifdef DOSTORESTATE35\n+\tSHA1_STORE_STATE(35)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, W, 35);\n+\t\n+#ifdef DOSTORESTATE36\n+\tSHA1_STORE_STATE(36)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, W, 36);\n+\t\n+#ifdef DOSTORESTATE37\n+\tSHA1_STORE_STATE(37)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, W, 37);\n+\t\n+#ifdef DOSTORESTATE38\n+\tSHA1_STORE_STATE(38)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, W, 38);\n+\t\n+#ifdef DOSTORESTATE39\n+\tSHA1_STORE_STATE(39)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, W, 39);\n+\n+\n+\n+#ifdef DOSTORESTATE40\n+\tSHA1_STORE_STATE(40)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 40);\n+\n+#ifdef DOSTORESTATE41\n+\tSHA1_STORE_STATE(41)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 41);\n+\n+#ifdef DOSTORESTATE42\n+\tSHA1_STORE_STATE(42)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 42);\n+\n+#ifdef DOSTORESTATE43\n+\tSHA1_STORE_STATE(43)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 43);\n+\n+#ifdef DOSTORESTATE44\n+\tSHA1_STORE_STATE(44)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 44);\n+\n+#ifdef DOSTORESTATE45\n+\tSHA1_STORE_STATE(45)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 45);\n+\n+#ifdef DOSTORESTATE46\n+\tSHA1_STORE_STATE(46)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 46);\n+\n+#ifdef DOSTORESTATE47\n+\tSHA1_STORE_STATE(47)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 47);\n+\n+#ifdef DOSTORESTATE48\n+\tSHA1_STORE_STATE(48)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 48);\n+\n+#ifdef DOSTORESTATE49\n+\tSHA1_STORE_STATE(49)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 49);\n+\n+#ifdef DOSTORESTATE50\n+\tSHA1_STORE_STATE(50)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 50);\n+\n+#ifdef DOSTORESTATE51\n+\tSHA1_STORE_STATE(51)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 51);\n+\n+#ifdef DOSTORESTATE52\n+\tSHA1_STORE_STATE(52)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 52);\n+\n+#ifdef DOSTORESTATE53\n+\tSHA1_STORE_STATE(53)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 53);\n+\n+#ifdef DOSTORESTATE54\n+\tSHA1_STORE_STATE(54)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 54);\n+\n+#ifdef DOSTORESTATE55\n+\tSHA1_STORE_STATE(55)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, W, 55);\n+\n+#ifdef DOSTORESTATE56\n+\tSHA1_STORE_STATE(56)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, W, 56);\n+\n+#ifdef DOSTORESTATE57\n+\tSHA1_STORE_STATE(57)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, W, 57);\n+\n+#ifdef DOSTORESTATE58\n+\tSHA1_STORE_STATE(58)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, W, 58);\n+\n+#ifdef DOSTORESTATE59\n+\tSHA1_STORE_STATE(59)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, W, 59);\n+\t\n+\n+\n+\n+#ifdef DOSTORESTATE60\n+\tSHA1_STORE_STATE(60)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 60);\n+\n+#ifdef DOSTORESTATE61\n+\tSHA1_STORE_STATE(61)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 61);\n+\n+#ifdef DOSTORESTATE62\n+\tSHA1_STORE_STATE(62)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 62);\n+\n+#ifdef DOSTORESTATE63\n+\tSHA1_STORE_STATE(63)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 63);\n+\n+#ifdef DOSTORESTATE64\n+\tSHA1_STORE_STATE(64)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 64);\n+\n+#ifdef DOSTORESTATE65\n+\tSHA1_STORE_STATE(65)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 65);\n+\n+#ifdef DOSTORESTATE66\n+\tSHA1_STORE_STATE(66)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 66);\n+\n+#ifdef DOSTORESTATE67\n+\tSHA1_STORE_STATE(67)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 67);\n+\n+#ifdef DOSTORESTATE68\n+\tSHA1_STORE_STATE(68)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 68);\n+\n+#ifdef DOSTORESTATE69\n+\tSHA1_STORE_STATE(69)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 69);\n+\n+#ifdef DOSTORESTATE70\n+\tSHA1_STORE_STATE(70)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 70);\n+\n+#ifdef DOSTORESTATE71\n+\tSHA1_STORE_STATE(71)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 71);\n+\n+#ifdef DOSTORESTATE72\n+\tSHA1_STORE_STATE(72)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 72);\n+\n+#ifdef DOSTORESTATE73\n+\tSHA1_STORE_STATE(73)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 73);\n+\n+#ifdef DOSTORESTATE74\n+\tSHA1_STORE_STATE(74)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 74);\n+\n+#ifdef DOSTORESTATE75\n+\tSHA1_STORE_STATE(75)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, W, 75);\n+\n+#ifdef DOSTORESTATE76\n+\tSHA1_STORE_STATE(76)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, W, 76);\n+\n+#ifdef DOSTORESTATE77\n+\tSHA1_STORE_STATE(77)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, W, 77);\n+\n+#ifdef DOSTORESTATE78\n+\tSHA1_STORE_STATE(78)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, W, 78);\n+\n+#ifdef DOSTORESTATE79\n+\tSHA1_STORE_STATE(79)\n+#endif\n+\tHASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, W, 79);\n+\n+\n+\n+\tihv[0] += a; ihv[1] += b; ihv[2] += c; ihv[3] += d; ihv[4] += e;\n+}\n+\n+\n+\n+\n+#define SHA1_RECOMPRESS(t) \\\n+void sha1recompress_fast_ ## t (uint32_t ihvin[5], uint32_t ihvout[5], const uint32_t me2[80], const uint32_t state[5]) \\\n+{ \\\n+\tuint32_t a = state[0], b = state[1], c = state[2], d = state[3], e = state[4]; \\\n+\tif (t > 79) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(b, c, d, e, a, me2, 79); \\\n+\tif (t > 78) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(c, d, e, a, b, me2, 78); \\\n+\tif (t > 77) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(d, e, a, b, c, me2, 77); \\\n+\tif (t > 76) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(e, a, b, c, d, me2, 76); \\\n+\tif (t > 75) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(a, b, c, d, e, me2, 75); \\\n+\tif (t > 74) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(b, c, d, e, a, me2, 74); \\\n+\tif (t > 73) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(c, d, e, a, b, me2, 73); \\\n+\tif (t > 72) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(d, e, a, b, c, me2, 72); \\\n+\tif (t > 71) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(e, a, b, c, d, me2, 71); \\\n+\tif (t > 70) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(a, b, c, d, e, me2, 70); \\\n+\tif (t > 69) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(b, c, d, e, a, me2, 69); \\\n+\tif (t > 68) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(c, d, e, a, b, me2, 68); \\\n+\tif (t > 67) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(d, e, a, b, c, me2, 67); \\\n+\tif (t > 66) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(e, a, b, c, d, me2, 66); \\\n+\tif (t > 65) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(a, b, c, d, e, me2, 65); \\\n+\tif (t > 64) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(b, c, d, e, a, me2, 64); \\\n+\tif (t > 63) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(c, d, e, a, b, me2, 63); \\\n+\tif (t > 62) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(d, e, a, b, c, me2, 62); \\\n+\tif (t > 61) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(e, a, b, c, d, me2, 61); \\\n+\tif (t > 60) HASHCLASH_SHA1COMPRESS_ROUND4_STEP_BW(a, b, c, d, e, me2, 60); \\\n+\tif (t > 59) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(b, c, d, e, a, me2, 59); \\\n+\tif (t > 58) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(c, d, e, a, b, me2, 58); \\\n+\tif (t > 57) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(d, e, a, b, c, me2, 57); \\\n+\tif (t > 56) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(e, a, b, c, d, me2, 56); \\\n+\tif (t > 55) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(a, b, c, d, e, me2, 55); \\\n+\tif (t > 54) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(b, c, d, e, a, me2, 54); \\\n+\tif (t > 53) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(c, d, e, a, b, me2, 53); \\\n+\tif (t > 52) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(d, e, a, b, c, me2, 52); \\\n+\tif (t > 51) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(e, a, b, c, d, me2, 51); \\\n+\tif (t > 50) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(a, b, c, d, e, me2, 50); \\\n+\tif (t > 49) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(b, c, d, e, a, me2, 49); \\\n+\tif (t > 48) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(c, d, e, a, b, me2, 48); \\\n+\tif (t > 47) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(d, e, a, b, c, me2, 47); \\\n+\tif (t > 46) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(e, a, b, c, d, me2, 46); \\\n+\tif (t > 45) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(a, b, c, d, e, me2, 45); \\\n+\tif (t > 44) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(b, c, d, e, a, me2, 44); \\\n+\tif (t > 43) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(c, d, e, a, b, me2, 43); \\\n+\tif (t > 42) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(d, e, a, b, c, me2, 42); \\\n+\tif (t > 41) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(e, a, b, c, d, me2, 41); \\\n+\tif (t > 40) HASHCLASH_SHA1COMPRESS_ROUND3_STEP_BW(a, b, c, d, e, me2, 40); \\\n+\tif (t > 39) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(b, c, d, e, a, me2, 39); \\\n+\tif (t > 38) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(c, d, e, a, b, me2, 38); \\\n+\tif (t > 37) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(d, e, a, b, c, me2, 37); \\\n+\tif (t > 36) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(e, a, b, c, d, me2, 36); \\\n+\tif (t > 35) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(a, b, c, d, e, me2, 35); \\\n+\tif (t > 34) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(b, c, d, e, a, me2, 34); \\\n+\tif (t > 33) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(c, d, e, a, b, me2, 33); \\\n+\tif (t > 32) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(d, e, a, b, c, me2, 32); \\\n+\tif (t > 31) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(e, a, b, c, d, me2, 31); \\\n+\tif (t > 30) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(a, b, c, d, e, me2, 30); \\\n+\tif (t > 29) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(b, c, d, e, a, me2, 29); \\\n+\tif (t > 28) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(c, d, e, a, b, me2, 28); \\\n+\tif (t > 27) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(d, e, a, b, c, me2, 27); \\\n+\tif (t > 26) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(e, a, b, c, d, me2, 26); \\\n+\tif (t > 25) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(a, b, c, d, e, me2, 25); \\\n+\tif (t > 24) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(b, c, d, e, a, me2, 24); \\\n+\tif (t > 23) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(c, d, e, a, b, me2, 23); \\\n+\tif (t > 22) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(d, e, a, b, c, me2, 22); \\\n+\tif (t > 21) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(e, a, b, c, d, me2, 21); \\\n+\tif (t > 20) HASHCLASH_SHA1COMPRESS_ROUND2_STEP_BW(a, b, c, d, e, me2, 20); \\\n+\tif (t > 19) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(b, c, d, e, a, me2, 19); \\\n+\tif (t > 18) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(c, d, e, a, b, me2, 18); \\\n+\tif (t > 17) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(d, e, a, b, c, me2, 17); \\\n+\tif (t > 16) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(e, a, b, c, d, me2, 16); \\\n+\tif (t > 15) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(a, b, c, d, e, me2, 15); \\\n+\tif (t > 14) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(b, c, d, e, a, me2, 14); \\\n+\tif (t > 13) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(c, d, e, a, b, me2, 13); \\\n+\tif (t > 12) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(d, e, a, b, c, me2, 12); \\\n+\tif (t > 11) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(e, a, b, c, d, me2, 11); \\\n+\tif (t > 10) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(a, b, c, d, e, me2, 10); \\\n+\tif (t > 9) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(b, c, d, e, a, me2, 9); \\\n+\tif (t > 8) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(c, d, e, a, b, me2, 8); \\\n+\tif (t > 7) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(d, e, a, b, c, me2, 7); \\\n+\tif (t > 6) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(e, a, b, c, d, me2, 6); \\\n+\tif (t > 5) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(a, b, c, d, e, me2, 5); \\\n+\tif (t > 4) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(b, c, d, e, a, me2, 4); \\\n+\tif (t > 3) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(c, d, e, a, b, me2, 3); \\\n+\tif (t > 2) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(d, e, a, b, c, me2, 2); \\\n+\tif (t > 1) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(e, a, b, c, d, me2, 1); \\\n+\tif (t > 0) HASHCLASH_SHA1COMPRESS_ROUND1_STEP_BW(a, b, c, d, e, me2, 0); \\\n+\tihvin[0] = a; ihvin[1] = b; ihvin[2] = c; ihvin[3] = d; ihvin[4] = e; \\\n+\ta = state[0]; b = state[1]; c = state[2]; d = state[3]; e = state[4]; \\\n+\tif (t <= 0) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, me2, 0); \\\n+\tif (t <= 1) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, me2, 1); \\\n+\tif (t <= 2) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, me2, 2); \\\n+\tif (t <= 3) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, me2, 3); \\\n+\tif (t <= 4) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, me2, 4); \\\n+\tif (t <= 5) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, me2, 5); \\\n+\tif (t <= 6) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, me2, 6); \\\n+\tif (t <= 7) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, me2, 7); \\\n+\tif (t <= 8) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, me2, 8); \\\n+\tif (t <= 9) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, me2, 9); \\\n+\tif (t <= 10) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, me2, 10); \\\n+\tif (t <= 11) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, me2, 11); \\\n+\tif (t <= 12) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, me2, 12); \\\n+\tif (t <= 13) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, me2, 13); \\\n+\tif (t <= 14) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, me2, 14); \\\n+\tif (t <= 15) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(a, b, c, d, e, me2, 15); \\\n+\tif (t <= 16) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(e, a, b, c, d, me2, 16); \\\n+\tif (t <= 17) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(d, e, a, b, c, me2, 17); \\\n+\tif (t <= 18) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(c, d, e, a, b, me2, 18); \\\n+\tif (t <= 19) HASHCLASH_SHA1COMPRESS_ROUND1_STEP(b, c, d, e, a, me2, 19); \\\n+\tif (t <= 20) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, me2, 20); \\\n+\tif (t <= 21) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, me2, 21); \\\n+\tif (t <= 22) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, me2, 22); \\\n+\tif (t <= 23) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, me2, 23); \\\n+\tif (t <= 24) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, me2, 24); \\\n+\tif (t <= 25) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, me2, 25); \\\n+\tif (t <= 26) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, me2, 26); \\\n+\tif (t <= 27) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, me2, 27); \\\n+\tif (t <= 28) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, me2, 28); \\\n+\tif (t <= 29) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, me2, 29); \\\n+\tif (t <= 30) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, me2, 30); \\\n+\tif (t <= 31) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, me2, 31); \\\n+\tif (t <= 32) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, me2, 32); \\\n+\tif (t <= 33) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, me2, 33); \\\n+\tif (t <= 34) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, me2, 34); \\\n+\tif (t <= 35) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(a, b, c, d, e, me2, 35); \\\n+\tif (t <= 36) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(e, a, b, c, d, me2, 36); \\\n+\tif (t <= 37) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(d, e, a, b, c, me2, 37); \\\n+\tif (t <= 38) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(c, d, e, a, b, me2, 38); \\\n+\tif (t <= 39) HASHCLASH_SHA1COMPRESS_ROUND2_STEP(b, c, d, e, a, me2, 39); \\\n+\tif (t <= 40) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, me2, 40); \\\n+\tif (t <= 41) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, me2, 41); \\\n+\tif (t <= 42) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, me2, 42); \\\n+\tif (t <= 43) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, me2, 43); \\\n+\tif (t <= 44) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, me2, 44); \\\n+\tif (t <= 45) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, me2, 45); \\\n+\tif (t <= 46) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, me2, 46); \\\n+\tif (t <= 47) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, me2, 47); \\\n+\tif (t <= 48) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, me2, 48); \\\n+\tif (t <= 49) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, me2, 49); \\\n+\tif (t <= 50) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, me2, 50); \\\n+\tif (t <= 51) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, me2, 51); \\\n+\tif (t <= 52) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, me2, 52); \\\n+\tif (t <= 53) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, me2, 53); \\\n+\tif (t <= 54) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, me2, 54); \\\n+\tif (t <= 55) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(a, b, c, d, e, me2, 55); \\\n+\tif (t <= 56) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(e, a, b, c, d, me2, 56); \\\n+\tif (t <= 57) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(d, e, a, b, c, me2, 57); \\\n+\tif (t <= 58) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(c, d, e, a, b, me2, 58); \\\n+\tif (t <= 59) HASHCLASH_SHA1COMPRESS_ROUND3_STEP(b, c, d, e, a, me2, 59); \\\n+\tif (t <= 60) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, me2, 60); \\\n+\tif (t <= 61) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, me2, 61); \\\n+\tif (t <= 62) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, me2, 62); \\\n+\tif (t <= 63) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, me2, 63); \\\n+\tif (t <= 64) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, me2, 64); \\\n+\tif (t <= 65) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, me2, 65); \\\n+\tif (t <= 66) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, me2, 66); \\\n+\tif (t <= 67) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, me2, 67); \\\n+\tif (t <= 68) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, me2, 68); \\\n+\tif (t <= 69) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, me2, 69); \\\n+\tif (t <= 70) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, me2, 70); \\\n+\tif (t <= 71) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, me2, 71); \\\n+\tif (t <= 72) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, me2, 72); \\\n+\tif (t <= 73) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, me2, 73); \\\n+\tif (t <= 74) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, me2, 74); \\\n+\tif (t <= 75) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(a, b, c, d, e, me2, 75); \\\n+\tif (t <= 76) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(e, a, b, c, d, me2, 76); \\\n+\tif (t <= 77) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(d, e, a, b, c, me2, 77); \\\n+\tif (t <= 78) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(c, d, e, a, b, me2, 78); \\\n+\tif (t <= 79) HASHCLASH_SHA1COMPRESS_ROUND4_STEP(b, c, d, e, a, me2, 79); \\\n+\tihvout[0] = ihvin[0] + a; ihvout[1] = ihvin[1] + b; ihvout[2] = ihvin[2] + c; ihvout[3] = ihvin[3] + d; ihvout[4] = ihvin[4] + e; \\\n+} \n+\n+SHA1_RECOMPRESS(0)\n+SHA1_RECOMPRESS(1)\n+SHA1_RECOMPRESS(2)\n+SHA1_RECOMPRESS(3)\n+SHA1_RECOMPRESS(4)\n+SHA1_RECOMPRESS(5)\n+SHA1_RECOMPRESS(6)\n+SHA1_RECOMPRESS(7)\n+SHA1_RECOMPRESS(8)\n+SHA1_RECOMPRESS(9)\n+\n+SHA1_RECOMPRESS(10)\n+SHA1_RECOMPRESS(11)\n+SHA1_RECOMPRESS(12)\n+SHA1_RECOMPRESS(13)\n+SHA1_RECOMPRESS(14)\n+SHA1_RECOMPRESS(15)\n+SHA1_RECOMPRESS(16)\n+SHA1_RECOMPRESS(17)\n+SHA1_RECOMPRESS(18)\n+SHA1_RECOMPRESS(19)\n+\n+SHA1_RECOMPRESS(20)\n+SHA1_RECOMPRESS(21)\n+SHA1_RECOMPRESS(22)\n+SHA1_RECOMPRESS(23)\n+SHA1_RECOMPRESS(24)\n+SHA1_RECOMPRESS(25)\n+SHA1_RECOMPRESS(26)\n+SHA1_RECOMPRESS(27)\n+SHA1_RECOMPRESS(28)\n+SHA1_RECOMPRESS(29)\n+\n+SHA1_RECOMPRESS(30)\n+SHA1_RECOMPRESS(31)\n+SHA1_RECOMPRESS(32)\n+SHA1_RECOMPRESS(33)\n+SHA1_RECOMPRESS(34)\n+SHA1_RECOMPRESS(35)\n+SHA1_RECOMPRESS(36)\n+SHA1_RECOMPRESS(37)\n+SHA1_RECOMPRESS(38)\n+SHA1_RECOMPRESS(39)\n+\n+SHA1_RECOMPRESS(40)\n+SHA1_RECOMPRESS(41)\n+SHA1_RECOMPRESS(42)\n+SHA1_RECOMPRESS(43)\n+SHA1_RECOMPRESS(44)\n+SHA1_RECOMPRESS(45)\n+SHA1_RECOMPRESS(46)\n+SHA1_RECOMPRESS(47)\n+SHA1_RECOMPRESS(48)\n+SHA1_RECOMPRESS(49)\n+\n+SHA1_RECOMPRESS(50)\n+SHA1_RECOMPRESS(51)\n+SHA1_RECOMPRESS(52)\n+SHA1_RECOMPRESS(53)\n+SHA1_RECOMPRESS(54)\n+SHA1_RECOMPRESS(55)\n+SHA1_RECOMPRESS(56)\n+SHA1_RECOMPRESS(57)\n+SHA1_RECOMPRESS(58)\n+SHA1_RECOMPRESS(59)\n+\n+SHA1_RECOMPRESS(60)\n+SHA1_RECOMPRESS(61)\n+SHA1_RECOMPRESS(62)\n+SHA1_RECOMPRESS(63)\n+SHA1_RECOMPRESS(64)\n+SHA1_RECOMPRESS(65)\n+SHA1_RECOMPRESS(66)\n+SHA1_RECOMPRESS(67)\n+SHA1_RECOMPRESS(68)\n+SHA1_RECOMPRESS(69)\n+\n+SHA1_RECOMPRESS(70)\n+SHA1_RECOMPRESS(71)\n+SHA1_RECOMPRESS(72)\n+SHA1_RECOMPRESS(73)\n+SHA1_RECOMPRESS(74)\n+SHA1_RECOMPRESS(75)\n+SHA1_RECOMPRESS(76)\n+SHA1_RECOMPRESS(77)\n+SHA1_RECOMPRESS(78)\n+SHA1_RECOMPRESS(79)\n+\n+sha1_recompression_type sha1_recompression_step[80] =\n+{\n+\tsha1recompress_fast_0, sha1recompress_fast_1, sha1recompress_fast_2, sha1recompress_fast_3, sha1recompress_fast_4, sha1recompress_fast_5, sha1recompress_fast_6, sha1recompress_fast_7, sha1recompress_fast_8, sha1recompress_fast_9,\n+\tsha1recompress_fast_10, sha1recompress_fast_11, sha1recompress_fast_12, sha1recompress_fast_13, sha1recompress_fast_14, sha1recompress_fast_15, sha1recompress_fast_16, sha1recompress_fast_17, sha1recompress_fast_18, sha1recompress_fast_19,\n+\tsha1recompress_fast_20, sha1recompress_fast_21, sha1recompress_fast_22, sha1recompress_fast_23, sha1recompress_fast_24, sha1recompress_fast_25, sha1recompress_fast_26, sha1recompress_fast_27, sha1recompress_fast_28, sha1recompress_fast_29,\n+\tsha1recompress_fast_30, sha1recompress_fast_31, sha1recompress_fast_32, sha1recompress_fast_33, sha1recompress_fast_34, sha1recompress_fast_35, sha1recompress_fast_36, sha1recompress_fast_37, sha1recompress_fast_38, sha1recompress_fast_39,\n+\tsha1recompress_fast_40, sha1recompress_fast_41, sha1recompress_fast_42, sha1recompress_fast_43, sha1recompress_fast_44, sha1recompress_fast_45, sha1recompress_fast_46, sha1recompress_fast_47, sha1recompress_fast_48, sha1recompress_fast_49,\n+\tsha1recompress_fast_50, sha1recompress_fast_51, sha1recompress_fast_52, sha1recompress_fast_53, sha1recompress_fast_54, sha1recompress_fast_55, sha1recompress_fast_56, sha1recompress_fast_57, sha1recompress_fast_58, sha1recompress_fast_59,\n+\tsha1recompress_fast_60, sha1recompress_fast_61, sha1recompress_fast_62, sha1recompress_fast_63, sha1recompress_fast_64, sha1recompress_fast_65, sha1recompress_fast_66, sha1recompress_fast_67, sha1recompress_fast_68, sha1recompress_fast_69,\n+\tsha1recompress_fast_70, sha1recompress_fast_71, sha1recompress_fast_72, sha1recompress_fast_73, sha1recompress_fast_74, sha1recompress_fast_75, sha1recompress_fast_76, sha1recompress_fast_77, sha1recompress_fast_78, sha1recompress_fast_79,\n+};\n+\n+\n+\n+\n+\n+void sha1_process(SHA1_CTX* ctx, const uint32_t block[16]) \n+{\n+\tunsigned i, j;\n+\tuint32_t ubc_dv_mask[DVMASKSIZE];\n+\tuint32_t ihvtmp[5];\n+\tfor (i=0; i < DVMASKSIZE; ++i)\n+\t\tubc_dv_mask[i]=0;\n+\tctx->ihv1[0] = ctx->ihv[0];\n+\tctx->ihv1[1] = ctx->ihv[1];\n+\tctx->ihv1[2] = ctx->ihv[2];\n+\tctx->ihv1[3] = ctx->ihv[3];\n+\tctx->ihv1[4] = ctx->ihv[4];\n+\tmemcpy(ctx->m1, block, 64);\n+\tsha1_message_expansion(ctx->m1);\n+\tif (ctx->detect_coll && ctx->ubc_check)\n+\t{\n+\t\tubc_check(ctx->m1, ubc_dv_mask);\n+\t}\n+\tsha1_compression_states(ctx->ihv, ctx->m1, ctx->states);\n+\tif (ctx->detect_coll)\n+\t{\n+\t\tfor (i = 0; sha1_dvs[i].dvType != 0; ++i) \n+\t\t{\n+\t\t\tif ((0 == ctx->ubc_check) || (((uint32_t)(1) << sha1_dvs[i].maskb) & ubc_dv_mask[sha1_dvs[i].maski]))\n+\t\t\t{\n+\t\t\t\tfor (j = 0; j < 80; ++j)\n+\t\t\t\t\tctx->m2[j] = ctx->m1[j] ^ sha1_dvs[i].dm[j];\n+\t\t\t\t(sha1_recompression_step[sha1_dvs[i].testt])(ctx->ihv2, ihvtmp, ctx->m2, ctx->states[sha1_dvs[i].testt]);\n+\t\t\t\t// to verify SHA-1 collision detection code with collisions for reduced-step SHA-1\n+\t\t\t\tif ((ihvtmp[0] == ctx->ihv[0] && ihvtmp[1] == ctx->ihv[1] && ihvtmp[2] == ctx->ihv[2] && ihvtmp[3] == ctx->ihv[3] && ihvtmp[4] == ctx->ihv[4])\n+\t\t\t\t\t|| (ctx->reduced_round_coll && ctx->ihv1[0] == ctx->ihv2[0] && ctx->ihv1[1] == ctx->ihv2[1] && ctx->ihv1[2] == ctx->ihv2[2] && ctx->ihv1[3] == ctx->ihv2[3] && ctx->ihv1[4] == ctx->ihv2[4]))\n+\t\t\t\t{\n+\t\t\t\t\tctx->found_collision = 1;\n+\t\t\t\t\t// TODO: call callback\n+\t\t\t\t\tif (ctx->callback != NULL)\n+\t\t\t\t\t\tctx->callback(ctx->total - 64, ctx->ihv1, ctx->ihv2, ctx->m1, ctx->m2);\n+\n+\t\t\t\t\tif (ctx->safe_hash) \n+\t\t\t\t\t{\n+\t\t\t\t\t\tsha1_compression_W(ctx->ihv, ctx->m1);\n+\t\t\t\t\t\tsha1_compression_W(ctx->ihv, ctx->m1);\n+\t\t\t\t\t}\n+\n+\t\t\t\t\tbreak;\n+\t\t\t\t}\n+\t\t\t}\n+\t\t}\n+\t}\n+}\n+\n+\n+\n+\n+\n+void swap_bytes(uint32_t val[16]) \n+{\n+\tunsigned i;\n+\tfor (i = 0; i < 16; ++i) \n+\t{\n+\t\tval[i] = ((val[i] << 8) & 0xFF00FF00) | ((val[i] >> 8) & 0xFF00FF);\n+\t\tval[i] = (val[i] << 16) | (val[i] >> 16);\n+\t}\n+}\n+\n+void SHA1DCInit(SHA1_CTX* ctx) \n+{\n+\tstatic const union { unsigned char bytes[4]; uint32_t value; } endianness = { { 0, 1, 2, 3 } };\n+\tstatic const uint32_t littleendian = 0x03020100;\n+\tctx->total = 0;\n+\tctx->ihv[0] = 0x67452301;\n+\tctx->ihv[1] = 0xEFCDAB89;\n+\tctx->ihv[2] = 0x98BADCFE;\n+\tctx->ihv[3] = 0x10325476;\n+\tctx->ihv[4] = 0xC3D2E1F0;\n+\tctx->found_collision = 0;\n+\tctx->safe_hash = 1;\n+\tctx->ubc_check = 1;\n+\tctx->detect_coll = 1;\n+\tctx->reduced_round_coll = 0;\n+\tctx->bigendian = (endianness.value != littleendian);\n+\tctx->callback = NULL;\n+}\n+\n+void SHA1DCSetSafeHash(SHA1_CTX* ctx, int safehash)\n+{\n+\tif (safehash)\n+\t\tctx->safe_hash = 1;\n+\telse\n+\t\tctx->safe_hash = 0;\n+}\n+\n+\n+void SHA1DCSetUseUBC(SHA1_CTX* ctx, int ubc_check)\n+{\n+\tif (ubc_check)\n+\t\tctx->ubc_check = 1;\n+\telse\n+\t\tctx->ubc_check = 0;\n+}\n+\n+void SHA1DCSetUseDetectColl(SHA1_CTX* ctx, int detect_coll)\n+{\n+\tif (detect_coll)\n+\t\tctx->detect_coll = 1;\n+\telse\n+\t\tctx->detect_coll = 0;\n+}\n+\n+void SHA1DCSetDetectReducedRoundCollision(SHA1_CTX* ctx, int reduced_round_coll)\n+{\n+\tif (reduced_round_coll)\n+\t\tctx->reduced_round_coll = 1;\n+\telse\n+\t\tctx->reduced_round_coll = 0;\n+}\n+\n+void SHA1DCSetCallback(SHA1_CTX* ctx, collision_block_callback callback)\n+{\n+\tctx->callback = callback;\n+}\n+\n+void SHA1DCUpdate(SHA1_CTX* ctx, const char* buf, unsigned len) \n+{\n+\tunsigned left, fill;\n+\tif (len == 0) \n+\t\treturn;\n+\n+\tleft = ctx->total & 63;\n+\tfill = 64 - left;\n+\n+\tif (left && len >= fill) \n+\t{\n+\t\tctx->total += fill;\n+\t\tmemcpy(ctx->buffer + left, buf, fill);\n+\t\tif (!ctx->bigendian)\n+\t\t\tswap_bytes((uint32_t*)(ctx->buffer));\n+\t\tsha1_process(ctx, (uint32_t*)(ctx->buffer));\n+\t\tbuf += fill;\n+\t\tlen -= fill;\n+\t\tleft = 0;\n+\t}\n+\twhile (len >= 64) \n+\t{\n+\t\tctx->total += 64;\n+\t\tif (!ctx->bigendian) \n+\t\t{\n+\t\t\tmemcpy(ctx->buffer, buf, 64);\n+\t\t\tswap_bytes((uint32_t*)(ctx->buffer));\n+\t\t\tsha1_process(ctx, (uint32_t*)(ctx->buffer));\n+\t\t}\n+\t\telse\n+\t\t\tsha1_process(ctx, (uint32_t*)(buf));\n+\t\tbuf += 64;\n+\t\tlen -= 64;\n+\t}\n+\tif (len > 0) \n+\t{\n+\t\tctx->total += len;\n+\t\tmemcpy(ctx->buffer + left, buf, len);\n+\t}\n+}\n+\n+static const unsigned char sha1_padding[64] =\n+{\n+\t0x80, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n+\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n+\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0,\n+\t0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0\n+};\n+\n+int SHA1DCFinal(unsigned char output[20], SHA1_CTX *ctx)\n+{\n+\tuint32_t last = ctx->total & 63;\n+\tuint32_t padn = (last < 56) ? (56 - last) : (120 - last);\n+\tuint64_t total;\n+\tSHA1DCUpdate(ctx, (const char*)(sha1_padding), padn);\n+\t\n+\ttotal = ctx->total - padn;\n+\ttotal <<= 3;\n+\tctx->buffer[56] = (unsigned char)(total >> 56);\n+\tctx->buffer[57] = (unsigned char)(total >> 48);\n+\tctx->buffer[58] = (unsigned char)(total >> 40);\n+\tctx->buffer[59] = (unsigned char)(total >> 32);\n+\tctx->buffer[60] = (unsigned char)(total >> 24);\n+\tctx->buffer[61] = (unsigned char)(total >> 16);\n+\tctx->buffer[62] = (unsigned char)(total >> 8);\n+\tctx->buffer[63] = (unsigned char)(total);\n+\tif (!ctx->bigendian)\n+\t\tswap_bytes((uint32_t*)(ctx->buffer));\n+\tsha1_process(ctx, (uint32_t*)(ctx->buffer));\n+\toutput[0] = (unsigned char)(ctx->ihv[0] >> 24);\n+\toutput[1] = (unsigned char)(ctx->ihv[0] >> 16);\n+\toutput[2] = (unsigned char)(ctx->ihv[0] >> 8);\n+\toutput[3] = (unsigned char)(ctx->ihv[0]);\n+\toutput[4] = (unsigned char)(ctx->ihv[1] >> 24);\n+\toutput[5] = (unsigned char)(ctx->ihv[1] >> 16);\n+\toutput[6] = (unsigned char)(ctx->ihv[1] >> 8);\n+\toutput[7] = (unsigned char)(ctx->ihv[1]);\n+\toutput[8] = (unsigned char)(ctx->ihv[2] >> 24);\n+\toutput[9] = (unsigned char)(ctx->ihv[2] >> 16);\n+\toutput[10] = (unsigned char)(ctx->ihv[2] >> 8);\n+\toutput[11] = (unsigned char)(ctx->ihv[2]);\n+\toutput[12] = (unsigned char)(ctx->ihv[3] >> 24);\n+\toutput[13] = (unsigned char)(ctx->ihv[3] >> 16);\n+\toutput[14] = (unsigned char)(ctx->ihv[3] >> 8);\n+\toutput[15] = (unsigned char)(ctx->ihv[3]);\n+\toutput[16] = (unsigned char)(ctx->ihv[4] >> 24);\n+\toutput[17] = (unsigned char)(ctx->ihv[4] >> 16);\n+\toutput[18] = (unsigned char)(ctx->ihv[4] >> 8);\n+\toutput[19] = (unsigned char)(ctx->ihv[4]);\n+\treturn ctx->found_collision;\n+}\ndiff --git a/sha1dc/sha1.h b/sha1dc/sha1.h\nnew file mode 100644\nindex 000000000..8b522f9d2\n--- /dev/null\n+++ b/sha1dc/sha1.h\n@@ -0,0 +1,94 @@\n+/***\n+* Copyright 2017 Marc Stevens <marc@marc-stevens.nl>, Dan Shumow <danshu@microsoft.com>\n+* Distributed under the MIT Software License.\n+* See accompanying file LICENSE.txt or copy at\n+* https://opensource.org/licenses/MIT\n+***/\n+\n+#include <stdint.h>\n+\n+// uses SHA-1 message expansion to expand the first 16 words of W[] to 80 words\n+void sha1_message_expansion(uint32_t W[80]);\n+\n+// sha-1 compression function; first version takes a message block pre-parsed as 16 32-bit integers, second version takes an already expanded message)\n+void sha1_compression(uint32_t ihv[5], const uint32_t m[16]);\n+void sha1_compression_W(uint32_t ihv[5], const uint32_t W[80]);\n+\n+// same as sha1_compression_W, but additionally store intermediate states\n+// only stores states ii (the state between step ii-1 and step ii) when DOSTORESTATEii is defined in ubc_check.h\n+void sha1_compression_states(uint32_t ihv[5], const uint32_t W[80], uint32_t states[80][5]);\n+\n+// function type for sha1_recompression_step_T (uint32_t ihvin[5], uint32_t ihvout[5], const uint32_t me2[80], const uint32_t state[5])\n+// where 0 <= T < 80\n+//       me2 is an expanded message (the expansion of an original message block XOR'ed with a disturbance vector's message block difference)\n+//       state is the internal state (a,b,c,d,e) before step T of the SHA-1 compression function while processing the original message block\n+// the function will return:\n+//       ihvin: the reconstructed input chaining value\n+//       ihvout: the reconstructed output chaining value\n+typedef void(*sha1_recompression_type)(uint32_t*, uint32_t*, const uint32_t*, const uint32_t*);\n+\n+// table of sha1_recompression_step_0, ... , sha1_recompression_step_79\n+extern sha1_recompression_type sha1_recompression_step[80];\n+\n+// a callback function type that can be set to be called when a collision block has been found:\n+// void collision_block_callback(uint64_t byteoffset, const uint32_t ihvin1[5], const uint32_t ihvin2[5], const uint32_t m1[80], const uint32_t m2[80])\n+typedef void(*collision_block_callback)(uint64_t, const uint32_t*, const uint32_t*, const uint32_t*, const uint32_t*);\n+\n+// the SHA-1 context\n+typedef struct {\n+\tuint64_t total;\n+\tuint32_t ihv[5];\n+\tunsigned char buffer[64];\n+\tint bigendian;\n+\tint found_collision;\n+\tint safe_hash;\n+\tint detect_coll;\n+\tint ubc_check;\n+\tint reduced_round_coll;\n+\tcollision_block_callback callback;\n+\n+\tuint32_t ihv1[5];\n+\tuint32_t ihv2[5];\n+\tuint32_t m1[80];\n+\tuint32_t m2[80];\n+\tuint32_t states[80][5];\n+} SHA1_CTX;\n+\n+// initialize SHA-1 context\n+void SHA1DCInit(SHA1_CTX*); \n+\n+// function to enable safe SHA-1 hashing:\n+// collision attacks are thwarted by hashing a detected near-collision block 3 times\n+// think of it as extending SHA-1 from 80-steps to 240-steps for such blocks:\n+//   the best collision attacks against SHA-1 have complexity about 2^60, \n+//   thus for 240-steps an immediate lower-bound for the best cryptanalytic attacks would 2^180\n+//   an attacker would be better off using a generic birthday search of complexity 2^80\n+//\n+// enabling safe SHA-1 hashing will result in the correct SHA-1 hash for messages where no collision attack was detected\n+// but it will result in a different SHA-1 hash for messages where a collision attack was detected \n+// this will automatically invalidate SHA-1 based digital signature forgeries\n+// enabled by default\n+void SHA1DCSetSafeHash(SHA1_CTX*, int);\n+\n+// function to disable or enable the use of Unavoidable Bitconditions (provides a significant speed up)\n+// enabled by default\n+void SHA1DCSetUseUBC(SHA1_CTX*, int);\n+\n+// function to disable or enable the use of Collision Detection\n+// enabled by default\n+void SHA1DCSetUseDetectColl(SHA1_CTX* ctx, int detect_coll);\n+\n+// function to disable or enable the detection of reduced-round SHA-1 collisions\n+// disabled by default\n+void SHA1DCSetDetectReducedRoundCollision(SHA1_CTX*, int);\n+\n+// function to set a callback function, pass NULL to disable\n+// by default no callback set\n+void SHA1DCSetCallback(SHA1_CTX*, collision_block_callback);\n+\n+// update SHA-1 context with buffer contents\n+void SHA1DCUpdate(SHA1_CTX*, const char*, unsigned);\n+\n+// obtain SHA-1 hash from SHA-1 context\n+// returns: 0 = no collision detected, otherwise = collision found => warn user for active attack\n+int  SHA1DCFinal(unsigned char[20], SHA1_CTX*); \ndiff --git a/sha1dc/ubc_check.c b/sha1dc/ubc_check.c\nnew file mode 100644\nindex 000000000..556aaf3c5\n--- /dev/null\n+++ b/sha1dc/ubc_check.c\n@@ -0,0 +1,361 @@\n+/***\n+* Copyright 2017 Marc Stevens <marc@marc-stevens.nl>, Dan Shumow <danshu@microsoft.com>\n+* Distributed under the MIT Software License.\n+* See accompanying file LICENSE.txt or copy at\n+* https://opensource.org/licenses/MIT\n+***/\n+\n+// this file was generated by the 'parse_bitrel' program in the tools section\n+// using the data files from directory 'tools/data/3565'\n+//\n+// sha1_dvs contains a list of SHA-1 Disturbance Vectors (DV) to check\n+// dvType, dvK and dvB define the DV: I(K,B) or II(K,B) (see the paper)\n+// dm[80] is the expanded message block XOR-difference defined by the DV\n+// testt is the step to do the recompression from for collision detection\n+// maski and maskb define the bit to check for each DV in the dvmask returned by ubc_check\n+//\n+// ubc_check takes as input an expanded message block and verifies the unavoidable bitconditions for all listed DVs\n+// it returns a dvmask where each bit belonging to a DV is set if all unavoidable bitconditions for that DV have been met\n+// thus one needs to do the recompression check for each DV that has its bit set\n+// \n+// ubc_check is programmatically generated and the unavoidable bitconditions have been hardcoded\n+// a directly verifiable version named ubc_check_verify can be found in ubc_check_verify.c\n+// ubc_check has been verified against ubc_check_verify using the 'ubc_check_test' program in the tools section\n+\n+#include <stdint.h>\n+#include \"ubc_check.h\"\n+\n+static const uint32_t DV_I_43_0_bit \t= (uint32_t)(1) << 0;\n+static const uint32_t DV_I_44_0_bit \t= (uint32_t)(1) << 1;\n+static const uint32_t DV_I_45_0_bit \t= (uint32_t)(1) << 2;\n+static const uint32_t DV_I_46_0_bit \t= (uint32_t)(1) << 3;\n+static const uint32_t DV_I_46_2_bit \t= (uint32_t)(1) << 4;\n+static const uint32_t DV_I_47_0_bit \t= (uint32_t)(1) << 5;\n+static const uint32_t DV_I_47_2_bit \t= (uint32_t)(1) << 6;\n+static const uint32_t DV_I_48_0_bit \t= (uint32_t)(1) << 7;\n+static const uint32_t DV_I_48_2_bit \t= (uint32_t)(1) << 8;\n+static const uint32_t DV_I_49_0_bit \t= (uint32_t)(1) << 9;\n+static const uint32_t DV_I_49_2_bit \t= (uint32_t)(1) << 10;\n+static const uint32_t DV_I_50_0_bit \t= (uint32_t)(1) << 11;\n+static const uint32_t DV_I_50_2_bit \t= (uint32_t)(1) << 12;\n+static const uint32_t DV_I_51_0_bit \t= (uint32_t)(1) << 13;\n+static const uint32_t DV_I_51_2_bit \t= (uint32_t)(1) << 14;\n+static const uint32_t DV_I_52_0_bit \t= (uint32_t)(1) << 15;\n+static const uint32_t DV_II_45_0_bit \t= (uint32_t)(1) << 16;\n+static const uint32_t DV_II_46_0_bit \t= (uint32_t)(1) << 17;\n+static const uint32_t DV_II_46_2_bit \t= (uint32_t)(1) << 18;\n+static const uint32_t DV_II_47_0_bit \t= (uint32_t)(1) << 19;\n+static const uint32_t DV_II_48_0_bit \t= (uint32_t)(1) << 20;\n+static const uint32_t DV_II_49_0_bit \t= (uint32_t)(1) << 21;\n+static const uint32_t DV_II_49_2_bit \t= (uint32_t)(1) << 22;\n+static const uint32_t DV_II_50_0_bit \t= (uint32_t)(1) << 23;\n+static const uint32_t DV_II_50_2_bit \t= (uint32_t)(1) << 24;\n+static const uint32_t DV_II_51_0_bit \t= (uint32_t)(1) << 25;\n+static const uint32_t DV_II_51_2_bit \t= (uint32_t)(1) << 26;\n+static const uint32_t DV_II_52_0_bit \t= (uint32_t)(1) << 27;\n+static const uint32_t DV_II_53_0_bit \t= (uint32_t)(1) << 28;\n+static const uint32_t DV_II_54_0_bit \t= (uint32_t)(1) << 29;\n+static const uint32_t DV_II_55_0_bit \t= (uint32_t)(1) << 30;\n+static const uint32_t DV_II_56_0_bit \t= (uint32_t)(1) << 31;\n+\n+dv_info_t sha1_dvs[] = \n+{\n+  {1,43,0,58,0,0, { 0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012,0x80000202,0x00000018,0x00000164,0x00000408,0x800000e6,0x8000004c,0x00000803,0x80000161,0x80000599 } }\n+, {1,44,0,58,0,1, { 0xb4000008,0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012,0x80000202,0x00000018,0x00000164,0x00000408,0x800000e6,0x8000004c,0x00000803,0x80000161 } }\n+, {1,45,0,58,0,2, { 0xf4000014,0xb4000008,0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012,0x80000202,0x00000018,0x00000164,0x00000408,0x800000e6,0x8000004c,0x00000803 } }\n+, {1,46,0,58,0,3, { 0x2c000010,0xf4000014,0xb4000008,0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012,0x80000202,0x00000018,0x00000164,0x00000408,0x800000e6,0x8000004c } }\n+, {1,46,2,58,0,4, { 0xb0000040,0xd0000053,0xd0000022,0x20000000,0x60000032,0x60000043,0x20000040,0xe0000042,0x60000002,0x80000001,0x00000020,0x00000003,0x40000052,0x40000040,0xe0000052,0xa0000000,0x80000040,0x20000001,0x20000060,0x80000001,0x40000042,0xc0000043,0x40000022,0x00000003,0x40000042,0xc0000043,0xc0000022,0x00000001,0x40000002,0xc0000043,0x40000062,0x80000001,0x40000042,0x40000042,0x40000002,0x00000002,0x00000040,0x80000002,0x80000000,0x80000002,0x80000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000000,0x00000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000101,0x00000009,0x00000012,0x00000202,0x0000001a,0x00000124,0x0000040c,0x00000026,0x0000004a,0x0000080a,0x00000060,0x00000590,0x00001020,0x0000039a,0x00000132 } }\n+, {1,47,0,58,0,5, { 0xc8000010,0x2c000010,0xf4000014,0xb4000008,0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012,0x80000202,0x00000018,0x00000164,0x00000408,0x800000e6 } }\n+, {1,47,2,58,0,6, { 0x20000043,0xb0000040,0xd0000053,0xd0000022,0x20000000,0x60000032,0x60000043,0x20000040,0xe0000042,0x60000002,0x80000001,0x00000020,0x00000003,0x40000052,0x40000040,0xe0000052,0xa0000000,0x80000040,0x20000001,0x20000060,0x80000001,0x40000042,0xc0000043,0x40000022,0x00000003,0x40000042,0xc0000043,0xc0000022,0x00000001,0x40000002,0xc0000043,0x40000062,0x80000001,0x40000042,0x40000042,0x40000002,0x00000002,0x00000040,0x80000002,0x80000000,0x80000002,0x80000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000000,0x00000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000101,0x00000009,0x00000012,0x00000202,0x0000001a,0x00000124,0x0000040c,0x00000026,0x0000004a,0x0000080a,0x00000060,0x00000590,0x00001020,0x0000039a } }\n+, {1,48,0,58,0,7, { 0xb800000a,0xc8000010,0x2c000010,0xf4000014,0xb4000008,0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012,0x80000202,0x00000018,0x00000164,0x00000408 } }\n+, {1,48,2,58,0,8, { 0xe000002a,0x20000043,0xb0000040,0xd0000053,0xd0000022,0x20000000,0x60000032,0x60000043,0x20000040,0xe0000042,0x60000002,0x80000001,0x00000020,0x00000003,0x40000052,0x40000040,0xe0000052,0xa0000000,0x80000040,0x20000001,0x20000060,0x80000001,0x40000042,0xc0000043,0x40000022,0x00000003,0x40000042,0xc0000043,0xc0000022,0x00000001,0x40000002,0xc0000043,0x40000062,0x80000001,0x40000042,0x40000042,0x40000002,0x00000002,0x00000040,0x80000002,0x80000000,0x80000002,0x80000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000000,0x00000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000101,0x00000009,0x00000012,0x00000202,0x0000001a,0x00000124,0x0000040c,0x00000026,0x0000004a,0x0000080a,0x00000060,0x00000590,0x00001020 } }\n+, {1,49,0,58,0,9, { 0x18000000,0xb800000a,0xc8000010,0x2c000010,0xf4000014,0xb4000008,0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012,0x80000202,0x00000018,0x00000164 } }\n+, {1,49,2,58,0,10, { 0x60000000,0xe000002a,0x20000043,0xb0000040,0xd0000053,0xd0000022,0x20000000,0x60000032,0x60000043,0x20000040,0xe0000042,0x60000002,0x80000001,0x00000020,0x00000003,0x40000052,0x40000040,0xe0000052,0xa0000000,0x80000040,0x20000001,0x20000060,0x80000001,0x40000042,0xc0000043,0x40000022,0x00000003,0x40000042,0xc0000043,0xc0000022,0x00000001,0x40000002,0xc0000043,0x40000062,0x80000001,0x40000042,0x40000042,0x40000002,0x00000002,0x00000040,0x80000002,0x80000000,0x80000002,0x80000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000000,0x00000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000101,0x00000009,0x00000012,0x00000202,0x0000001a,0x00000124,0x0000040c,0x00000026,0x0000004a,0x0000080a,0x00000060,0x00000590 } }\n+, {1,50,0,65,0,11, { 0x0800000c,0x18000000,0xb800000a,0xc8000010,0x2c000010,0xf4000014,0xb4000008,0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012,0x80000202,0x00000018 } }\n+, {1,50,2,65,0,12, { 0x20000030,0x60000000,0xe000002a,0x20000043,0xb0000040,0xd0000053,0xd0000022,0x20000000,0x60000032,0x60000043,0x20000040,0xe0000042,0x60000002,0x80000001,0x00000020,0x00000003,0x40000052,0x40000040,0xe0000052,0xa0000000,0x80000040,0x20000001,0x20000060,0x80000001,0x40000042,0xc0000043,0x40000022,0x00000003,0x40000042,0xc0000043,0xc0000022,0x00000001,0x40000002,0xc0000043,0x40000062,0x80000001,0x40000042,0x40000042,0x40000002,0x00000002,0x00000040,0x80000002,0x80000000,0x80000002,0x80000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000000,0x00000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000101,0x00000009,0x00000012,0x00000202,0x0000001a,0x00000124,0x0000040c,0x00000026,0x0000004a,0x0000080a,0x00000060 } }\n+, {1,51,0,65,0,13, { 0xe8000000,0x0800000c,0x18000000,0xb800000a,0xc8000010,0x2c000010,0xf4000014,0xb4000008,0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012,0x80000202 } }\n+, {1,51,2,65,0,14, { 0xa0000003,0x20000030,0x60000000,0xe000002a,0x20000043,0xb0000040,0xd0000053,0xd0000022,0x20000000,0x60000032,0x60000043,0x20000040,0xe0000042,0x60000002,0x80000001,0x00000020,0x00000003,0x40000052,0x40000040,0xe0000052,0xa0000000,0x80000040,0x20000001,0x20000060,0x80000001,0x40000042,0xc0000043,0x40000022,0x00000003,0x40000042,0xc0000043,0xc0000022,0x00000001,0x40000002,0xc0000043,0x40000062,0x80000001,0x40000042,0x40000042,0x40000002,0x00000002,0x00000040,0x80000002,0x80000000,0x80000002,0x80000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000000,0x00000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000101,0x00000009,0x00000012,0x00000202,0x0000001a,0x00000124,0x0000040c,0x00000026,0x0000004a,0x0000080a } }\n+, {1,52,0,65,0,15, { 0x04000010,0xe8000000,0x0800000c,0x18000000,0xb800000a,0xc8000010,0x2c000010,0xf4000014,0xb4000008,0x08000000,0x9800000c,0xd8000010,0x08000010,0xb8000010,0x98000000,0x60000000,0x00000008,0xc0000000,0x90000014,0x10000010,0xb8000014,0x28000000,0x20000010,0x48000000,0x08000018,0x60000000,0x90000010,0xf0000010,0x90000008,0xc0000000,0x90000010,0xf0000010,0xb0000008,0x40000000,0x90000000,0xf0000010,0x90000018,0x60000000,0x90000010,0x90000010,0x90000000,0x80000000,0x00000010,0xa0000000,0x20000000,0xa0000000,0x20000010,0x00000000,0x20000010,0x20000000,0x00000010,0x20000000,0x00000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000040,0x40000002,0x80000004,0x80000080,0x80000006,0x00000049,0x00000103,0x80000009,0x80000012 } }\n+, {2,45,0,58,0,16, { 0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107,0x00000089,0x00000014,0x8000024b,0x0000011b,0x8000016d,0x8000041a,0x000002e4,0x80000054,0x00000967 } }\n+, {2,46,0,58,0,17, { 0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107,0x00000089,0x00000014,0x8000024b,0x0000011b,0x8000016d,0x8000041a,0x000002e4,0x80000054 } }\n+, {2,46,2,58,0,18, { 0x90000070,0xb0000053,0x30000008,0x00000043,0xd0000072,0xb0000010,0xf0000062,0xc0000042,0x00000030,0xe0000042,0x20000060,0xe0000041,0x20000050,0xc0000041,0xe0000072,0xa0000003,0xc0000012,0x60000041,0xc0000032,0x20000001,0xc0000002,0xe0000042,0x60000042,0x80000002,0x00000000,0x00000000,0x80000000,0x00000002,0x00000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000001,0x00000060,0x80000003,0x40000002,0xc0000040,0xc0000002,0x80000000,0x80000000,0x80000002,0x00000040,0x00000002,0x80000000,0x80000000,0x80000000,0x00000002,0x00000040,0x00000000,0x80000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000105,0x00000089,0x00000016,0x0000020b,0x0000011b,0x0000012d,0x0000041e,0x00000224,0x00000050,0x0000092e,0x0000046c,0x000005b6,0x0000106a,0x00000b90,0x00000152 } }\n+, {2,47,0,58,0,19, { 0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107,0x00000089,0x00000014,0x8000024b,0x0000011b,0x8000016d,0x8000041a,0x000002e4 } }\n+, {2,48,0,58,0,20, { 0xbc00001a,0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107,0x00000089,0x00000014,0x8000024b,0x0000011b,0x8000016d,0x8000041a } }\n+, {2,49,0,58,0,21, { 0x3c000004,0xbc00001a,0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107,0x00000089,0x00000014,0x8000024b,0x0000011b,0x8000016d } }\n+, {2,49,2,58,0,22, { 0xf0000010,0xf000006a,0x80000040,0x90000070,0xb0000053,0x30000008,0x00000043,0xd0000072,0xb0000010,0xf0000062,0xc0000042,0x00000030,0xe0000042,0x20000060,0xe0000041,0x20000050,0xc0000041,0xe0000072,0xa0000003,0xc0000012,0x60000041,0xc0000032,0x20000001,0xc0000002,0xe0000042,0x60000042,0x80000002,0x00000000,0x00000000,0x80000000,0x00000002,0x00000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000001,0x00000060,0x80000003,0x40000002,0xc0000040,0xc0000002,0x80000000,0x80000000,0x80000002,0x00000040,0x00000002,0x80000000,0x80000000,0x80000000,0x00000002,0x00000040,0x00000000,0x80000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000105,0x00000089,0x00000016,0x0000020b,0x0000011b,0x0000012d,0x0000041e,0x00000224,0x00000050,0x0000092e,0x0000046c,0x000005b6 } }\n+, {2,50,0,65,0,23, { 0xb400001c,0x3c000004,0xbc00001a,0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107,0x00000089,0x00000014,0x8000024b,0x0000011b } }\n+, {2,50,2,65,0,24, { 0xd0000072,0xf0000010,0xf000006a,0x80000040,0x90000070,0xb0000053,0x30000008,0x00000043,0xd0000072,0xb0000010,0xf0000062,0xc0000042,0x00000030,0xe0000042,0x20000060,0xe0000041,0x20000050,0xc0000041,0xe0000072,0xa0000003,0xc0000012,0x60000041,0xc0000032,0x20000001,0xc0000002,0xe0000042,0x60000042,0x80000002,0x00000000,0x00000000,0x80000000,0x00000002,0x00000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000001,0x00000060,0x80000003,0x40000002,0xc0000040,0xc0000002,0x80000000,0x80000000,0x80000002,0x00000040,0x00000002,0x80000000,0x80000000,0x80000000,0x00000002,0x00000040,0x00000000,0x80000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000105,0x00000089,0x00000016,0x0000020b,0x0000011b,0x0000012d,0x0000041e,0x00000224,0x00000050,0x0000092e,0x0000046c } }\n+, {2,51,0,65,0,25, { 0xc0000010,0xb400001c,0x3c000004,0xbc00001a,0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107,0x00000089,0x00000014,0x8000024b } }\n+, {2,51,2,65,0,26, { 0x00000043,0xd0000072,0xf0000010,0xf000006a,0x80000040,0x90000070,0xb0000053,0x30000008,0x00000043,0xd0000072,0xb0000010,0xf0000062,0xc0000042,0x00000030,0xe0000042,0x20000060,0xe0000041,0x20000050,0xc0000041,0xe0000072,0xa0000003,0xc0000012,0x60000041,0xc0000032,0x20000001,0xc0000002,0xe0000042,0x60000042,0x80000002,0x00000000,0x00000000,0x80000000,0x00000002,0x00000040,0x00000000,0x80000040,0x80000000,0x00000040,0x80000001,0x00000060,0x80000003,0x40000002,0xc0000040,0xc0000002,0x80000000,0x80000000,0x80000002,0x00000040,0x00000002,0x80000000,0x80000000,0x80000000,0x00000002,0x00000040,0x00000000,0x80000040,0x80000002,0x00000000,0x80000000,0x80000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000004,0x00000080,0x00000004,0x00000009,0x00000105,0x00000089,0x00000016,0x0000020b,0x0000011b,0x0000012d,0x0000041e,0x00000224,0x00000050,0x0000092e } }\n+, {2,52,0,65,0,27, { 0x0c000002,0xc0000010,0xb400001c,0x3c000004,0xbc00001a,0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107,0x00000089,0x00000014 } }\n+, {2,53,0,65,0,28, { 0xcc000014,0x0c000002,0xc0000010,0xb400001c,0x3c000004,0xbc00001a,0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107,0x00000089 } }\n+, {2,54,0,65,0,29, { 0x0400001c,0xcc000014,0x0c000002,0xc0000010,0xb400001c,0x3c000004,0xbc00001a,0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b,0x80000107 } }\n+, {2,55,0,65,0,30, { 0x00000010,0x0400001c,0xcc000014,0x0c000002,0xc0000010,0xb400001c,0x3c000004,0xbc00001a,0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046,0x4000004b } }\n+, {2,56,0,65,0,31, { 0x2600001a,0x00000010,0x0400001c,0xcc000014,0x0c000002,0xc0000010,0xb400001c,0x3c000004,0xbc00001a,0x20000010,0x2400001c,0xec000014,0x0c000002,0xc0000010,0xb400001c,0x2c000004,0xbc000018,0xb0000010,0x0000000c,0xb8000010,0x08000018,0x78000010,0x08000014,0x70000010,0xb800001c,0xe8000000,0xb0000004,0x58000010,0xb000000c,0x48000000,0xb0000000,0xb8000010,0x98000010,0xa0000000,0x00000000,0x00000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0x20000000,0x00000010,0x60000000,0x00000018,0xe0000000,0x90000000,0x30000010,0xb0000000,0x20000000,0x20000000,0xa0000000,0x00000010,0x80000000,0x20000000,0x20000000,0x20000000,0x80000000,0x00000010,0x00000000,0x20000010,0xa0000000,0x00000000,0x20000000,0x20000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000000,0x00000001,0x00000020,0x00000001,0x40000002,0x40000041,0x40000022,0x80000005,0xc0000082,0xc0000046 } }\n+, {0,0,0,0,0,0, {0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0,0}}\n+};\n+void ubc_check(const uint32_t W[80], uint32_t dvmask[1])\n+{\n+\tuint32_t mask = ~((uint32_t)(0));\n+\tmask &= (((((W[44]^W[45])>>29)&1)-1) | ~(DV_I_48_0_bit|DV_I_51_0_bit|DV_I_52_0_bit|DV_II_45_0_bit|DV_II_46_0_bit|DV_II_50_0_bit|DV_II_51_0_bit));\n+\tmask &= (((((W[49]^W[50])>>29)&1)-1) | ~(DV_I_46_0_bit|DV_II_45_0_bit|DV_II_50_0_bit|DV_II_51_0_bit|DV_II_55_0_bit|DV_II_56_0_bit));\n+\tmask &= (((((W[48]^W[49])>>29)&1)-1) | ~(DV_I_45_0_bit|DV_I_52_0_bit|DV_II_49_0_bit|DV_II_50_0_bit|DV_II_54_0_bit|DV_II_55_0_bit));\n+\tmask &= ((((W[47]^(W[50]>>25))&(1<<4))-(1<<4)) | ~(DV_I_47_0_bit|DV_I_49_0_bit|DV_I_51_0_bit|DV_II_45_0_bit|DV_II_51_0_bit|DV_II_56_0_bit));\n+\tmask &= (((((W[47]^W[48])>>29)&1)-1) | ~(DV_I_44_0_bit|DV_I_51_0_bit|DV_II_48_0_bit|DV_II_49_0_bit|DV_II_53_0_bit|DV_II_54_0_bit));\n+\tmask &= (((((W[46]>>4)^(W[49]>>29))&1)-1) | ~(DV_I_46_0_bit|DV_I_48_0_bit|DV_I_50_0_bit|DV_I_52_0_bit|DV_II_50_0_bit|DV_II_55_0_bit));\n+\tmask &= (((((W[46]^W[47])>>29)&1)-1) | ~(DV_I_43_0_bit|DV_I_50_0_bit|DV_II_47_0_bit|DV_II_48_0_bit|DV_II_52_0_bit|DV_II_53_0_bit));\n+\tmask &= (((((W[45]>>4)^(W[48]>>29))&1)-1) | ~(DV_I_45_0_bit|DV_I_47_0_bit|DV_I_49_0_bit|DV_I_51_0_bit|DV_II_49_0_bit|DV_II_54_0_bit));\n+\tmask &= (((((W[45]^W[46])>>29)&1)-1) | ~(DV_I_49_0_bit|DV_I_52_0_bit|DV_II_46_0_bit|DV_II_47_0_bit|DV_II_51_0_bit|DV_II_52_0_bit));\n+\tmask &= (((((W[44]>>4)^(W[47]>>29))&1)-1) | ~(DV_I_44_0_bit|DV_I_46_0_bit|DV_I_48_0_bit|DV_I_50_0_bit|DV_II_48_0_bit|DV_II_53_0_bit));\n+\tmask &= (((((W[43]>>4)^(W[46]>>29))&1)-1) | ~(DV_I_43_0_bit|DV_I_45_0_bit|DV_I_47_0_bit|DV_I_49_0_bit|DV_II_47_0_bit|DV_II_52_0_bit));\n+\tmask &= (((((W[43]^W[44])>>29)&1)-1) | ~(DV_I_47_0_bit|DV_I_50_0_bit|DV_I_51_0_bit|DV_II_45_0_bit|DV_II_49_0_bit|DV_II_50_0_bit));\n+\tmask &= (((((W[42]>>4)^(W[45]>>29))&1)-1) | ~(DV_I_44_0_bit|DV_I_46_0_bit|DV_I_48_0_bit|DV_I_52_0_bit|DV_II_46_0_bit|DV_II_51_0_bit));\n+\tmask &= (((((W[41]>>4)^(W[44]>>29))&1)-1) | ~(DV_I_43_0_bit|DV_I_45_0_bit|DV_I_47_0_bit|DV_I_51_0_bit|DV_II_45_0_bit|DV_II_50_0_bit));\n+\tmask &= (((((W[40]^W[41])>>29)&1)-1) | ~(DV_I_44_0_bit|DV_I_47_0_bit|DV_I_48_0_bit|DV_II_46_0_bit|DV_II_47_0_bit|DV_II_56_0_bit));\n+\tmask &= (((((W[54]^W[55])>>29)&1)-1) | ~(DV_I_51_0_bit|DV_II_47_0_bit|DV_II_50_0_bit|DV_II_55_0_bit|DV_II_56_0_bit));\n+\tmask &= (((((W[53]^W[54])>>29)&1)-1) | ~(DV_I_50_0_bit|DV_II_46_0_bit|DV_II_49_0_bit|DV_II_54_0_bit|DV_II_55_0_bit));\n+\tmask &= (((((W[52]^W[53])>>29)&1)-1) | ~(DV_I_49_0_bit|DV_II_45_0_bit|DV_II_48_0_bit|DV_II_53_0_bit|DV_II_54_0_bit));\n+\tmask &= ((((W[50]^(W[53]>>25))&(1<<4))-(1<<4)) | ~(DV_I_50_0_bit|DV_I_52_0_bit|DV_II_46_0_bit|DV_II_48_0_bit|DV_II_54_0_bit));\n+\tmask &= (((((W[50]^W[51])>>29)&1)-1) | ~(DV_I_47_0_bit|DV_II_46_0_bit|DV_II_51_0_bit|DV_II_52_0_bit|DV_II_56_0_bit));\n+\tmask &= ((((W[49]^(W[52]>>25))&(1<<4))-(1<<4)) | ~(DV_I_49_0_bit|DV_I_51_0_bit|DV_II_45_0_bit|DV_II_47_0_bit|DV_II_53_0_bit));\n+\tmask &= ((((W[48]^(W[51]>>25))&(1<<4))-(1<<4)) | ~(DV_I_48_0_bit|DV_I_50_0_bit|DV_I_52_0_bit|DV_II_46_0_bit|DV_II_52_0_bit));\n+\tmask &= (((((W[42]^W[43])>>29)&1)-1) | ~(DV_I_46_0_bit|DV_I_49_0_bit|DV_I_50_0_bit|DV_II_48_0_bit|DV_II_49_0_bit));\n+\tmask &= (((((W[41]^W[42])>>29)&1)-1) | ~(DV_I_45_0_bit|DV_I_48_0_bit|DV_I_49_0_bit|DV_II_47_0_bit|DV_II_48_0_bit));\n+\tmask &= (((((W[40]>>4)^(W[43]>>29))&1)-1) | ~(DV_I_44_0_bit|DV_I_46_0_bit|DV_I_50_0_bit|DV_II_49_0_bit|DV_II_56_0_bit));\n+\tmask &= (((((W[39]>>4)^(W[42]>>29))&1)-1) | ~(DV_I_43_0_bit|DV_I_45_0_bit|DV_I_49_0_bit|DV_II_48_0_bit|DV_II_55_0_bit));\n+\tif (mask & (DV_I_44_0_bit|DV_I_48_0_bit|DV_II_47_0_bit|DV_II_54_0_bit|DV_II_56_0_bit))\n+\t\tmask &= (((((W[38]>>4)^(W[41]>>29))&1)-1) | ~(DV_I_44_0_bit|DV_I_48_0_bit|DV_II_47_0_bit|DV_II_54_0_bit|DV_II_56_0_bit));\n+\tmask &= (((((W[37]>>4)^(W[40]>>29))&1)-1) | ~(DV_I_43_0_bit|DV_I_47_0_bit|DV_II_46_0_bit|DV_II_53_0_bit|DV_II_55_0_bit));\n+\tif (mask & (DV_I_52_0_bit|DV_II_48_0_bit|DV_II_51_0_bit|DV_II_56_0_bit))\n+\t\tmask &= (((((W[55]^W[56])>>29)&1)-1) | ~(DV_I_52_0_bit|DV_II_48_0_bit|DV_II_51_0_bit|DV_II_56_0_bit));\n+\tif (mask & (DV_I_52_0_bit|DV_II_48_0_bit|DV_II_50_0_bit|DV_II_56_0_bit))\n+\t\tmask &= ((((W[52]^(W[55]>>25))&(1<<4))-(1<<4)) | ~(DV_I_52_0_bit|DV_II_48_0_bit|DV_II_50_0_bit|DV_II_56_0_bit));\n+\tif (mask & (DV_I_51_0_bit|DV_II_47_0_bit|DV_II_49_0_bit|DV_II_55_0_bit))\n+\t\tmask &= ((((W[51]^(W[54]>>25))&(1<<4))-(1<<4)) | ~(DV_I_51_0_bit|DV_II_47_0_bit|DV_II_49_0_bit|DV_II_55_0_bit));\n+\tif (mask & (DV_I_48_0_bit|DV_II_47_0_bit|DV_II_52_0_bit|DV_II_53_0_bit))\n+\t\tmask &= (((((W[51]^W[52])>>29)&1)-1) | ~(DV_I_48_0_bit|DV_II_47_0_bit|DV_II_52_0_bit|DV_II_53_0_bit));\n+\tif (mask & (DV_I_46_0_bit|DV_I_49_0_bit|DV_II_45_0_bit|DV_II_48_0_bit))\n+\t\tmask &= (((((W[36]>>4)^(W[40]>>29))&1)-1) | ~(DV_I_46_0_bit|DV_I_49_0_bit|DV_II_45_0_bit|DV_II_48_0_bit));\n+\tif (mask & (DV_I_52_0_bit|DV_II_48_0_bit|DV_II_49_0_bit))\n+\t\tmask &= ((0-(((W[53]^W[56])>>29)&1)) | ~(DV_I_52_0_bit|DV_II_48_0_bit|DV_II_49_0_bit));\n+\tif (mask & (DV_I_50_0_bit|DV_II_46_0_bit|DV_II_47_0_bit))\n+\t\tmask &= ((0-(((W[51]^W[54])>>29)&1)) | ~(DV_I_50_0_bit|DV_II_46_0_bit|DV_II_47_0_bit));\n+\tif (mask & (DV_I_49_0_bit|DV_I_51_0_bit|DV_II_45_0_bit))\n+\t\tmask &= ((0-(((W[50]^W[52])>>29)&1)) | ~(DV_I_49_0_bit|DV_I_51_0_bit|DV_II_45_0_bit));\n+\tif (mask & (DV_I_48_0_bit|DV_I_50_0_bit|DV_I_52_0_bit))\n+\t\tmask &= ((0-(((W[49]^W[51])>>29)&1)) | ~(DV_I_48_0_bit|DV_I_50_0_bit|DV_I_52_0_bit));\n+\tif (mask & (DV_I_47_0_bit|DV_I_49_0_bit|DV_I_51_0_bit))\n+\t\tmask &= ((0-(((W[48]^W[50])>>29)&1)) | ~(DV_I_47_0_bit|DV_I_49_0_bit|DV_I_51_0_bit));\n+\tif (mask & (DV_I_46_0_bit|DV_I_48_0_bit|DV_I_50_0_bit))\n+\t\tmask &= ((0-(((W[47]^W[49])>>29)&1)) | ~(DV_I_46_0_bit|DV_I_48_0_bit|DV_I_50_0_bit));\n+\tif (mask & (DV_I_45_0_bit|DV_I_47_0_bit|DV_I_49_0_bit))\n+\t\tmask &= ((0-(((W[46]^W[48])>>29)&1)) | ~(DV_I_45_0_bit|DV_I_47_0_bit|DV_I_49_0_bit));\n+\tmask &= ((((W[45]^W[47])&(1<<6))-(1<<6)) | ~(DV_I_47_2_bit|DV_I_49_2_bit|DV_I_51_2_bit));\n+\tif (mask & (DV_I_44_0_bit|DV_I_46_0_bit|DV_I_48_0_bit))\n+\t\tmask &= ((0-(((W[45]^W[47])>>29)&1)) | ~(DV_I_44_0_bit|DV_I_46_0_bit|DV_I_48_0_bit));\n+\tmask &= (((((W[44]^W[46])>>6)&1)-1) | ~(DV_I_46_2_bit|DV_I_48_2_bit|DV_I_50_2_bit));\n+\tif (mask & (DV_I_43_0_bit|DV_I_45_0_bit|DV_I_47_0_bit))\n+\t\tmask &= ((0-(((W[44]^W[46])>>29)&1)) | ~(DV_I_43_0_bit|DV_I_45_0_bit|DV_I_47_0_bit));\n+\tmask &= ((0-((W[41]^(W[42]>>5))&(1<<1))) | ~(DV_I_48_2_bit|DV_II_46_2_bit|DV_II_51_2_bit));\n+\tmask &= ((0-((W[40]^(W[41]>>5))&(1<<1))) | ~(DV_I_47_2_bit|DV_I_51_2_bit|DV_II_50_2_bit));\n+\tif (mask & (DV_I_44_0_bit|DV_I_46_0_bit|DV_II_56_0_bit))\n+\t\tmask &= ((0-(((W[40]^W[42])>>4)&1)) | ~(DV_I_44_0_bit|DV_I_46_0_bit|DV_II_56_0_bit));\n+\tmask &= ((0-((W[39]^(W[40]>>5))&(1<<1))) | ~(DV_I_46_2_bit|DV_I_50_2_bit|DV_II_49_2_bit));\n+\tif (mask & (DV_I_43_0_bit|DV_I_45_0_bit|DV_II_55_0_bit))\n+\t\tmask &= ((0-(((W[39]^W[41])>>4)&1)) | ~(DV_I_43_0_bit|DV_I_45_0_bit|DV_II_55_0_bit));\n+\tif (mask & (DV_I_44_0_bit|DV_II_54_0_bit|DV_II_56_0_bit))\n+\t\tmask &= ((0-(((W[38]^W[40])>>4)&1)) | ~(DV_I_44_0_bit|DV_II_54_0_bit|DV_II_56_0_bit));\n+\tif (mask & (DV_I_43_0_bit|DV_II_53_0_bit|DV_II_55_0_bit))\n+\t\tmask &= ((0-(((W[37]^W[39])>>4)&1)) | ~(DV_I_43_0_bit|DV_II_53_0_bit|DV_II_55_0_bit));\n+\tmask &= ((0-((W[36]^(W[37]>>5))&(1<<1))) | ~(DV_I_47_2_bit|DV_I_50_2_bit|DV_II_46_2_bit));\n+\tif (mask & (DV_I_45_0_bit|DV_I_48_0_bit|DV_II_47_0_bit))\n+\t\tmask &= (((((W[35]>>4)^(W[39]>>29))&1)-1) | ~(DV_I_45_0_bit|DV_I_48_0_bit|DV_II_47_0_bit));\n+\tif (mask & (DV_I_48_0_bit|DV_II_48_0_bit))\n+\t\tmask &= ((0-((W[63]^(W[64]>>5))&(1<<0))) | ~(DV_I_48_0_bit|DV_II_48_0_bit));\n+\tif (mask & (DV_I_45_0_bit|DV_II_45_0_bit))\n+\t\tmask &= ((0-((W[63]^(W[64]>>5))&(1<<1))) | ~(DV_I_45_0_bit|DV_II_45_0_bit));\n+\tif (mask & (DV_I_47_0_bit|DV_II_47_0_bit))\n+\t\tmask &= ((0-((W[62]^(W[63]>>5))&(1<<0))) | ~(DV_I_47_0_bit|DV_II_47_0_bit));\n+\tif (mask & (DV_I_46_0_bit|DV_II_46_0_bit))\n+\t\tmask &= ((0-((W[61]^(W[62]>>5))&(1<<0))) | ~(DV_I_46_0_bit|DV_II_46_0_bit));\n+\tmask &= ((0-((W[61]^(W[62]>>5))&(1<<2))) | ~(DV_I_46_2_bit|DV_II_46_2_bit));\n+\tif (mask & (DV_I_45_0_bit|DV_II_45_0_bit))\n+\t\tmask &= ((0-((W[60]^(W[61]>>5))&(1<<0))) | ~(DV_I_45_0_bit|DV_II_45_0_bit));\n+\tif (mask & (DV_II_51_0_bit|DV_II_54_0_bit))\n+\t\tmask &= (((((W[58]^W[59])>>29)&1)-1) | ~(DV_II_51_0_bit|DV_II_54_0_bit));\n+\tif (mask & (DV_II_50_0_bit|DV_II_53_0_bit))\n+\t\tmask &= (((((W[57]^W[58])>>29)&1)-1) | ~(DV_II_50_0_bit|DV_II_53_0_bit));\n+\tif (mask & (DV_II_52_0_bit|DV_II_54_0_bit))\n+\t\tmask &= ((((W[56]^(W[59]>>25))&(1<<4))-(1<<4)) | ~(DV_II_52_0_bit|DV_II_54_0_bit));\n+\tif (mask & (DV_II_51_0_bit|DV_II_52_0_bit))\n+\t\tmask &= ((0-(((W[56]^W[59])>>29)&1)) | ~(DV_II_51_0_bit|DV_II_52_0_bit));\n+\tif (mask & (DV_II_49_0_bit|DV_II_52_0_bit))\n+\t\tmask &= (((((W[56]^W[57])>>29)&1)-1) | ~(DV_II_49_0_bit|DV_II_52_0_bit));\n+\tif (mask & (DV_II_51_0_bit|DV_II_53_0_bit))\n+\t\tmask &= ((((W[55]^(W[58]>>25))&(1<<4))-(1<<4)) | ~(DV_II_51_0_bit|DV_II_53_0_bit));\n+\tif (mask & (DV_II_50_0_bit|DV_II_52_0_bit))\n+\t\tmask &= ((((W[54]^(W[57]>>25))&(1<<4))-(1<<4)) | ~(DV_II_50_0_bit|DV_II_52_0_bit));\n+\tif (mask & (DV_II_49_0_bit|DV_II_51_0_bit))\n+\t\tmask &= ((((W[53]^(W[56]>>25))&(1<<4))-(1<<4)) | ~(DV_II_49_0_bit|DV_II_51_0_bit));\n+\tmask &= ((((W[51]^(W[50]>>5))&(1<<1))-(1<<1)) | ~(DV_I_50_2_bit|DV_II_46_2_bit));\n+\tmask &= ((((W[48]^W[50])&(1<<6))-(1<<6)) | ~(DV_I_50_2_bit|DV_II_46_2_bit));\n+\tif (mask & (DV_I_51_0_bit|DV_I_52_0_bit))\n+\t\tmask &= ((0-(((W[48]^W[55])>>29)&1)) | ~(DV_I_51_0_bit|DV_I_52_0_bit));\n+\tmask &= ((((W[47]^W[49])&(1<<6))-(1<<6)) | ~(DV_I_49_2_bit|DV_I_51_2_bit));\n+\tmask &= ((((W[48]^(W[47]>>5))&(1<<1))-(1<<1)) | ~(DV_I_47_2_bit|DV_II_51_2_bit));\n+\tmask &= ((((W[46]^W[48])&(1<<6))-(1<<6)) | ~(DV_I_48_2_bit|DV_I_50_2_bit));\n+\tmask &= ((((W[47]^(W[46]>>5))&(1<<1))-(1<<1)) | ~(DV_I_46_2_bit|DV_II_50_2_bit));\n+\tmask &= ((0-((W[44]^(W[45]>>5))&(1<<1))) | ~(DV_I_51_2_bit|DV_II_49_2_bit));\n+\tmask &= ((((W[43]^W[45])&(1<<6))-(1<<6)) | ~(DV_I_47_2_bit|DV_I_49_2_bit));\n+\tmask &= (((((W[42]^W[44])>>6)&1)-1) | ~(DV_I_46_2_bit|DV_I_48_2_bit));\n+\tmask &= ((((W[43]^(W[42]>>5))&(1<<1))-(1<<1)) | ~(DV_II_46_2_bit|DV_II_51_2_bit));\n+\tmask &= ((((W[42]^(W[41]>>5))&(1<<1))-(1<<1)) | ~(DV_I_51_2_bit|DV_II_50_2_bit));\n+\tmask &= ((((W[41]^(W[40]>>5))&(1<<1))-(1<<1)) | ~(DV_I_50_2_bit|DV_II_49_2_bit));\n+\tif (mask & (DV_I_52_0_bit|DV_II_51_0_bit))\n+\t\tmask &= ((((W[39]^(W[43]>>25))&(1<<4))-(1<<4)) | ~(DV_I_52_0_bit|DV_II_51_0_bit));\n+\tif (mask & (DV_I_51_0_bit|DV_II_50_0_bit))\n+\t\tmask &= ((((W[38]^(W[42]>>25))&(1<<4))-(1<<4)) | ~(DV_I_51_0_bit|DV_II_50_0_bit));\n+\tif (mask & (DV_I_48_2_bit|DV_I_51_2_bit))\n+\t\tmask &= ((0-((W[37]^(W[38]>>5))&(1<<1))) | ~(DV_I_48_2_bit|DV_I_51_2_bit));\n+\tif (mask & (DV_I_50_0_bit|DV_II_49_0_bit))\n+\t\tmask &= ((((W[37]^(W[41]>>25))&(1<<4))-(1<<4)) | ~(DV_I_50_0_bit|DV_II_49_0_bit));\n+\tif (mask & (DV_II_52_0_bit|DV_II_54_0_bit))\n+\t\tmask &= ((0-((W[36]^W[38])&(1<<4))) | ~(DV_II_52_0_bit|DV_II_54_0_bit));\n+\tmask &= ((0-((W[35]^(W[36]>>5))&(1<<1))) | ~(DV_I_46_2_bit|DV_I_49_2_bit));\n+\tif (mask & (DV_I_51_0_bit|DV_II_47_0_bit))\n+\t\tmask &= ((((W[35]^(W[39]>>25))&(1<<3))-(1<<3)) | ~(DV_I_51_0_bit|DV_II_47_0_bit));\n+if (mask) {\n+\n+\tif (mask & DV_I_43_0_bit)\n+\t\t if (\n+\t\t\t    !((W[61]^(W[62]>>5)) & (1<<1))\n+\t\t\t || !(!((W[59]^(W[63]>>25)) & (1<<5)))\n+\t\t\t || !((W[58]^(W[63]>>30)) & (1<<0))\n+\t\t )  mask &= ~DV_I_43_0_bit;\n+\tif (mask & DV_I_44_0_bit)\n+\t\t if (\n+\t\t\t    !((W[62]^(W[63]>>5)) & (1<<1))\n+\t\t\t || !(!((W[60]^(W[64]>>25)) & (1<<5)))\n+\t\t\t || !((W[59]^(W[64]>>30)) & (1<<0))\n+\t\t )  mask &= ~DV_I_44_0_bit;\n+\tif (mask & DV_I_46_2_bit)\n+\t\tmask &= ((~((W[40]^W[42])>>2)) | ~DV_I_46_2_bit);\n+\tif (mask & DV_I_47_2_bit)\n+\t\t if (\n+\t\t\t    !((W[62]^(W[63]>>5)) & (1<<2))\n+\t\t\t || !(!((W[41]^W[43]) & (1<<6)))\n+\t\t )  mask &= ~DV_I_47_2_bit;\n+\tif (mask & DV_I_48_2_bit)\n+\t\t if (\n+\t\t\t    !((W[63]^(W[64]>>5)) & (1<<2))\n+\t\t\t || !(!((W[48]^(W[49]<<5)) & (1<<6)))\n+\t\t )  mask &= ~DV_I_48_2_bit;\n+\tif (mask & DV_I_49_2_bit)\n+\t\t if (\n+\t\t\t    !(!((W[49]^(W[50]<<5)) & (1<<6)))\n+\t\t\t || !((W[42]^W[50]) & (1<<1))\n+\t\t\t || !(!((W[39]^(W[40]<<5)) & (1<<6)))\n+\t\t\t || !((W[38]^W[40]) & (1<<1))\n+\t\t )  mask &= ~DV_I_49_2_bit;\n+\tif (mask & DV_I_50_0_bit)\n+\t\tmask &= ((((W[36]^W[37])<<7)) | ~DV_I_50_0_bit);\n+\tif (mask & DV_I_50_2_bit)\n+\t\tmask &= ((((W[43]^W[51])<<11)) | ~DV_I_50_2_bit);\n+\tif (mask & DV_I_51_0_bit)\n+\t\tmask &= ((((W[37]^W[38])<<9)) | ~DV_I_51_0_bit);\n+\tif (mask & DV_I_51_2_bit)\n+\t\t if (\n+\t\t\t    !(!((W[51]^(W[52]<<5)) & (1<<6)))\n+\t\t\t || !(!((W[49]^W[51]) & (1<<6)))\n+\t\t\t || !(!((W[37]^(W[37]>>5)) & (1<<1)))\n+\t\t\t || !(!((W[35]^(W[39]>>25)) & (1<<5)))\n+\t\t )  mask &= ~DV_I_51_2_bit;\n+\tif (mask & DV_I_52_0_bit)\n+\t\tmask &= ((((W[38]^W[39])<<11)) | ~DV_I_52_0_bit);\n+\tif (mask & DV_II_46_2_bit)\n+\t\tmask &= ((((W[47]^W[51])<<17)) | ~DV_II_46_2_bit);\n+\tif (mask & DV_II_48_0_bit)\n+\t\t if (\n+\t\t\t    !(!((W[36]^(W[40]>>25)) & (1<<3)))\n+\t\t\t || !((W[35]^(W[40]<<2)) & (1<<30))\n+\t\t )  mask &= ~DV_II_48_0_bit;\n+\tif (mask & DV_II_49_0_bit)\n+\t\t if (\n+\t\t\t    !(!((W[37]^(W[41]>>25)) & (1<<3)))\n+\t\t\t || !((W[36]^(W[41]<<2)) & (1<<30))\n+\t\t )  mask &= ~DV_II_49_0_bit;\n+\tif (mask & DV_II_49_2_bit)\n+\t\t if (\n+\t\t\t    !(!((W[53]^(W[54]<<5)) & (1<<6)))\n+\t\t\t || !(!((W[51]^W[53]) & (1<<6)))\n+\t\t\t || !((W[50]^W[54]) & (1<<1))\n+\t\t\t || !(!((W[45]^(W[46]<<5)) & (1<<6)))\n+\t\t\t || !(!((W[37]^(W[41]>>25)) & (1<<5)))\n+\t\t\t || !((W[36]^(W[41]>>30)) & (1<<0))\n+\t\t )  mask &= ~DV_II_49_2_bit;\n+\tif (mask & DV_II_50_0_bit)\n+\t\t if (\n+\t\t\t    !((W[55]^W[58]) & (1<<29))\n+\t\t\t || !(!((W[38]^(W[42]>>25)) & (1<<3)))\n+\t\t\t || !((W[37]^(W[42]<<2)) & (1<<30))\n+\t\t )  mask &= ~DV_II_50_0_bit;\n+\tif (mask & DV_II_50_2_bit)\n+\t\t if (\n+\t\t\t    !(!((W[54]^(W[55]<<5)) & (1<<6)))\n+\t\t\t || !(!((W[52]^W[54]) & (1<<6)))\n+\t\t\t || !((W[51]^W[55]) & (1<<1))\n+\t\t\t || !((W[45]^W[47]) & (1<<1))\n+\t\t\t || !(!((W[38]^(W[42]>>25)) & (1<<5)))\n+\t\t\t || !((W[37]^(W[42]>>30)) & (1<<0))\n+\t\t )  mask &= ~DV_II_50_2_bit;\n+\tif (mask & DV_II_51_0_bit)\n+\t\t if (\n+\t\t\t    !(!((W[39]^(W[43]>>25)) & (1<<3)))\n+\t\t\t || !((W[38]^(W[43]<<2)) & (1<<30))\n+\t\t )  mask &= ~DV_II_51_0_bit;\n+\tif (mask & DV_II_51_2_bit)\n+\t\t if (\n+\t\t\t    !(!((W[55]^(W[56]<<5)) & (1<<6)))\n+\t\t\t || !(!((W[53]^W[55]) & (1<<6)))\n+\t\t\t || !((W[52]^W[56]) & (1<<1))\n+\t\t\t || !((W[46]^W[48]) & (1<<1))\n+\t\t\t || !(!((W[39]^(W[43]>>25)) & (1<<5)))\n+\t\t\t || !((W[38]^(W[43]>>30)) & (1<<0))\n+\t\t )  mask &= ~DV_II_51_2_bit;\n+\tif (mask & DV_II_52_0_bit)\n+\t\t if (\n+\t\t\t    !(!((W[59]^W[60]) & (1<<29)))\n+\t\t\t || !(!((W[40]^(W[44]>>25)) & (1<<3)))\n+\t\t\t || !(!((W[40]^(W[44]>>25)) & (1<<4)))\n+\t\t\t || !((W[39]^(W[44]<<2)) & (1<<30))\n+\t\t )  mask &= ~DV_II_52_0_bit;\n+\tif (mask & DV_II_53_0_bit)\n+\t\t if (\n+\t\t\t    !((W[58]^W[61]) & (1<<29))\n+\t\t\t || !(!((W[57]^(W[61]>>25)) & (1<<4)))\n+\t\t\t || !(!((W[41]^(W[45]>>25)) & (1<<3)))\n+\t\t\t || !(!((W[41]^(W[45]>>25)) & (1<<4)))\n+\t\t )  mask &= ~DV_II_53_0_bit;\n+\tif (mask & DV_II_54_0_bit)\n+\t\t if (\n+\t\t\t    !(!((W[58]^(W[62]>>25)) & (1<<4)))\n+\t\t\t || !(!((W[42]^(W[46]>>25)) & (1<<3)))\n+\t\t\t || !(!((W[42]^(W[46]>>25)) & (1<<4)))\n+\t\t )  mask &= ~DV_II_54_0_bit;\n+\tif (mask & DV_II_55_0_bit)\n+\t\t if (\n+\t\t\t    !(!((W[59]^(W[63]>>25)) & (1<<4)))\n+\t\t\t || !(!((W[57]^(W[59]>>25)) & (1<<4)))\n+\t\t\t || !(!((W[43]^(W[47]>>25)) & (1<<3)))\n+\t\t\t || !(!((W[43]^(W[47]>>25)) & (1<<4)))\n+\t\t )  mask &= ~DV_II_55_0_bit;\n+\tif (mask & DV_II_56_0_bit)\n+\t\t if (\n+\t\t\t    !(!((W[60]^(W[64]>>25)) & (1<<4)))\n+\t\t\t || !(!((W[44]^(W[48]>>25)) & (1<<3)))\n+\t\t\t || !(!((W[44]^(W[48]>>25)) & (1<<4)))\n+\t\t )  mask &= ~DV_II_56_0_bit;\n+}\n+\n+\tdvmask[0]=mask;\n+}\ndiff --git a/sha1dc/ubc_check.h b/sha1dc/ubc_check.h\nnew file mode 100644\nindex 000000000..27285bdf5\n--- /dev/null\n+++ b/sha1dc/ubc_check.h\n@@ -0,0 +1,35 @@\n+/***\n+* Copyright 2017 Marc Stevens <marc@marc-stevens.nl>, Dan Shumow <danshu@microsoft.com>\n+* Distributed under the MIT Software License.\n+* See accompanying file LICENSE.txt or copy at\n+* https://opensource.org/licenses/MIT\n+***/\n+\n+// this file was generated by the 'parse_bitrel' program in the tools section\n+// using the data files from directory 'tools/data/3565'\n+//\n+// sha1_dvs contains a list of SHA-1 Disturbance Vectors (DV) to check\n+// dvType, dvK and dvB define the DV: I(K,B) or II(K,B) (see the paper)\n+// dm[80] is the expanded message block XOR-difference defined by the DV\n+// testt is the step to do the recompression from for collision detection\n+// maski and maskb define the bit to check for each DV in the dvmask returned by ubc_check\n+//\n+// ubc_check takes as input an expanded message block and verifies the unavoidable bitconditions for all listed DVs\n+// it returns a dvmask where each bit belonging to a DV is set if all unavoidable bitconditions for that DV have been met\n+// thus one needs to do the recompression check for each DV that has its bit set\n+\n+#ifndef UBC_CHECK_H\n+#define UBC_CHECK_H\n+\n+#include <stdint.h>\n+\n+#define DVMASKSIZE 1\n+typedef struct { int dvType; int dvK; int dvB; int testt; int maski; int maskb; uint32_t dm[80]; } dv_info_t;\n+extern dv_info_t sha1_dvs[];\n+void ubc_check(const uint32_t W[80], uint32_t dvmask[DVMASKSIZE]);\n+\n+#define DOSTORESTATE58\n+#define DOSTORESTATE65\n+\n+\n+#endif // UBC_CHECK_H\n-- \n2.12.0.rc2.629.ga7951ed82\n\n"},{"id":"312462","messageId":"20170223230550.7eosg57dzy346jkk@sigill.intra.peff.net","threadId":"45202","inReplyTo":"20170223230507.kuxjqtg3ghcfskc6@sigill.intra.peff.net","subject":"[PATCH 2/3] sha1dc: adjust header includes for git","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T23:05:50Z","receivedAt":"2017-02-23T23:05:57Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"We can replace system includes with git-compat-util.h (and\nshould make sure it is included in all .c files). We can\ndrop includes from headers entirely, as every .c file is\nsupposed to include git-compat-util itself first.\n\nWe also use the full \"sha1dc/\" path for including related\nfiles. This isn't strictly necessary, but makes the expected\nresolution more obvious.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n sha1dc/sha1.c      | 9 +++------\n sha1dc/sha1.h      | 2 --\n sha1dc/ubc_check.c | 4 ++--\n sha1dc/ubc_check.h | 2 --\n 4 files changed, 5 insertions(+), 12 deletions(-)\n\ndiff --git a/sha1dc/sha1.c b/sha1dc/sha1.c\nindex ed2010911..762c6fff8 100644\n--- a/sha1dc/sha1.c\n+++ b/sha1dc/sha1.c\n@@ -5,12 +5,9 @@\n * https://opensource.org/licenses/MIT\n ***/\n \n-#include <string.h>\n-#include <memory.h>\n-#include <stdio.h>\n-\n-#include \"sha1.h\"\n-#include \"ubc_check.h\"\n+#include \"git-compat-util.h\"\n+#include \"sha1dc/sha1.h\"\n+#include \"sha1dc/ubc_check.h\"\n \n #define rotate_right(x,n) (((x)>>(n))|((x)<<(32-(n))))\n #define rotate_left(x,n)  (((x)<<(n))|((x)>>(32-(n))))\ndiff --git a/sha1dc/sha1.h b/sha1dc/sha1.h\nindex 8b522f9d2..ce5390397 100644\n--- a/sha1dc/sha1.h\n+++ b/sha1dc/sha1.h\n@@ -5,8 +5,6 @@\n * https://opensource.org/licenses/MIT\n ***/\n \n-#include <stdint.h>\n-\n // uses SHA-1 message expansion to expand the first 16 words of W[] to 80 words\n void sha1_message_expansion(uint32_t W[80]);\n \ndiff --git a/sha1dc/ubc_check.c b/sha1dc/ubc_check.c\nindex 556aaf3c5..6bccd4f2b 100644\n--- a/sha1dc/ubc_check.c\n+++ b/sha1dc/ubc_check.c\n@@ -22,8 +22,8 @@\n // a directly verifiable version named ubc_check_verify can be found in ubc_check_verify.c\n // ubc_check has been verified against ubc_check_verify using the 'ubc_check_test' program in the tools section\n \n-#include <stdint.h>\n-#include \"ubc_check.h\"\n+#include \"git-compat-util.h\"\n+#include \"sha1dc/ubc_check.h\"\n \n static const uint32_t DV_I_43_0_bit \t= (uint32_t)(1) << 0;\n static const uint32_t DV_I_44_0_bit \t= (uint32_t)(1) << 1;\ndiff --git a/sha1dc/ubc_check.h b/sha1dc/ubc_check.h\nindex 27285bdf5..05ff944eb 100644\n--- a/sha1dc/ubc_check.h\n+++ b/sha1dc/ubc_check.h\n@@ -21,8 +21,6 @@\n #ifndef UBC_CHECK_H\n #define UBC_CHECK_H\n \n-#include <stdint.h>\n-\n #define DVMASKSIZE 1\n typedef struct { int dvType; int dvK; int dvB; int testt; int maski; int maskb; uint32_t dm[80]; } dv_info_t;\n extern dv_info_t sha1_dvs[];\n-- \n2.12.0.rc2.629.ga7951ed82\n\n"},{"id":"312465","messageId":"20170223230621.43anex65ndoqbgnf@sigill.intra.peff.net","threadId":"45202","inReplyTo":"20170223230507.kuxjqtg3ghcfskc6@sigill.intra.peff.net","subject":"[PATCH 3/3] Makefile: add USE_SHA1DC knob","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-23T23:06:21Z","receivedAt":"2017-02-23T23:06:40Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"This knob lets you use the sha1dc implementation from:\n\n      https://github.com/cr-marcstevens/sha1collisiondetection\n\nwhich can detect certain types of collision attacks (even\nwhen we only see half of the colliding pair).\n\nThe big downside is that it's slower than either the openssl\nor block-sha1 implementations.\n\nHere are some timings based off of linux.git:\n\n  - compute sha1 over whole packfile\n    before: 1.349s\n     after: 5.067s\n    change: +275%\n\n  - rev-list --all\n    before: 5.742s\n     after: 5.730s\n    change: -0.2%\n\n  - rev-list --all --objects\n    before: 33.257s\n     after: 33.392s\n    change: +0.4%\n\n  - index-pack --verify\n    before: 2m20s\n     after: 5m43s\n    change: +145%\n\n  - git log --no-merges -10000 -p\n    before: 9.532s\n     after: 9.683s\n    change: +1.5%\n\nSo overall the sha1 computation is about 3-4x slower. But of\ncourse most operations do more than just sha1. Accessing\ncommits and trees isn't slowed at all (both the +/- changes\nthere are well within the run-to-run noise). Accessing the\nblobs is a little slower, but mostly drowned out by the cost\nof things like actually generating patches.\n\nThe most-affected operation is `index-pack --verify`, which\nis essentially just computing the sha1 on every object. It's\na bit worse than twice as slow, which means every push and\nevery fetch is going to experience that.\n\nSigned-off-by: Jeff King <peff@peff.net>\n---\n Makefile      | 10 ++++++++++\n sha1dc/sha1.c | 22 ++++++++++++++++++++++\n sha1dc/sha1.h | 16 ++++++++++++++++\n 3 files changed, 48 insertions(+)\n\ndiff --git a/Makefile b/Makefile\nindex 8e4081e06..7c4906250 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -142,6 +142,10 @@ all::\n # Define PPC_SHA1 environment variable when running make to make use of\n # a bundled SHA1 routine optimized for PowerPC.\n #\n+# Define USE_SHA1DC to unconditionally enable the collision-detecting sha1\n+# algorithm. This is slower, but may detect attempted collision attacks.\n+# Takes priority over other *_SHA1 knobs.\n+#\n # Define SHA1_MAX_BLOCK_SIZE to limit the amount of data that will be hashed\n # in one call to the platform's SHA1_Update(). e.g. APPLE_COMMON_CRYPTO\n # wants 'SHA1_MAX_BLOCK_SIZE=1024L*1024L*1024L' defined.\n@@ -1386,6 +1390,11 @@ ifdef APPLE_COMMON_CRYPTO\n \tSHA1_MAX_BLOCK_SIZE = 1024L*1024L*1024L\n endif\n \n+ifdef USE_SHA1DC\n+\tSHA1_HEADER = \"sha1dc/sha1.h\"\n+\tLIB_OBJS += sha1dc/sha1.o\n+\tLIB_OBJS += sha1dc/ubc_check.o\n+else\n ifdef BLK_SHA1\n \tSHA1_HEADER = \"block-sha1/sha1.h\"\n \tLIB_OBJS += block-sha1/sha1.o\n@@ -1403,6 +1412,7 @@ else\n endif\n endif\n endif\n+endif\n \n ifdef SHA1_MAX_BLOCK_SIZE\n \tLIB_OBJS += compat/sha1-chunked.o\ndiff --git a/sha1dc/sha1.c b/sha1dc/sha1.c\nindex 762c6fff8..1566ec4c7 100644\n--- a/sha1dc/sha1.c\n+++ b/sha1dc/sha1.c\n@@ -1141,3 +1141,25 @@ int SHA1DCFinal(unsigned char output[20], SHA1_CTX *ctx)\n \toutput[19] = (unsigned char)(ctx->ihv[4]);\n \treturn ctx->found_collision;\n }\n+\n+static const char collision_message[] =\n+\"The SHA1 computation detected evidence of a collision attack;\\n\"\n+\"refusing to process the contents.\";\n+\n+void git_SHA1DCFinal(unsigned char hash[20], SHA1_CTX *ctx)\n+{\n+\tif (SHA1DCFinal(hash, ctx))\n+\t\tdie(collision_message);\n+}\n+\n+void git_SHA1DCUpdate(SHA1_CTX *ctx, const void *vdata, unsigned long len)\n+{\n+\tconst char *data = vdata;\n+\t/* We expect an unsigned long, but sha1dc only takes an int */\n+\twhile (len > INT_MAX) {\n+\t\tSHA1DCUpdate(ctx, data, INT_MAX);\n+\t\tdata += INT_MAX;\n+\t\tlen -= INT_MAX;\n+\t}\n+\tSHA1DCUpdate(ctx, data, len);\n+}\ndiff --git a/sha1dc/sha1.h b/sha1dc/sha1.h\nindex ce5390397..1bb0ace99 100644\n--- a/sha1dc/sha1.h\n+++ b/sha1dc/sha1.h\n@@ -90,3 +90,19 @@ void SHA1DCUpdate(SHA1_CTX*, const char*, unsigned);\n // obtain SHA-1 hash from SHA-1 context\n // returns: 0 = no collision detected, otherwise = collision found => warn user for active attack\n int  SHA1DCFinal(unsigned char[20], SHA1_CTX*); \n+\n+\n+/*\n+ * Same as SHA1DCFinal, but convert collision attack case into a verbose die().\n+ */\n+void git_SHA1DCFinal(unsigned char [20], SHA1_CTX *);\n+\n+/*\n+ * Same as SHA1DCUpdate, but adjust types to match git's usual interface.\n+ */\n+void git_SHA1DCUpdate(SHA1_CTX *ctx, const void *data, unsigned long len);\n+\n+#define platform_SHA_CTX SHA1_CTX\n+#define platform_SHA1_Init SHA1DCInit\n+#define platform_SHA1_Update git_SHA1DCUpdate\n+#define platform_SHA1_Final git_SHA1DCFinal\n-- \n2.12.0.rc2.629.ga7951ed82\n"},{"id":"312469","messageId":"CAGZ79kZHPdBTKEqJeAa5xDcsC5v9x4DdUuDOiRNSgOV5aCx9Kw@mail.gmail.com","threadId":"45202","inReplyTo":"20170223230536.tdmtsn46e4lnrimx@sigill.intra.peff.net","subject":"Re: [PATCH 1/3] add collision-detecting sha1 implementation","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2017-02-23T23:15:11Z","receivedAt":"2017-02-23T23:16:39Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Thu, Feb 23, 2017 at 3:05 PM, Jeff King <peff@peff.net> wrote:\n\n> +* Copyright 2017 Marc Stevens <marc@marc-stevens.nl>, Dan Shumow (danshu@microsoft.com)\n> +* Distributed under the MIT Software License.\n> +* See accompanying file LICENSE.txt or copy at\n\nThe accompanying LICENSE file did not make it into this patch,\nthat is more specialized/verbose than the one at\nhttps://opensource.org/licenses/MIT\nw.r.t. copyright notice requirement.\n\nApart from that MIT seems to be compatible with GPL\naccording to the FSF, though IANAL.\n"},{"id":"312471","messageId":"CA+55aFyj2nEec4AL5pqY4Gz-mRuBuhp=+hRjTnhPVEnAU89i=g@mail.gmail.com","threadId":"45202","inReplyTo":"20170223230507.kuxjqtg3ghcfskc6@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-23T23:14:19Z","receivedAt":"2017-02-23T23:23:07Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 3:05 PM, Jeff King <peff@peff.net> wrote:\n>\n> (By the way, I don't see your version on the list, Linus, which probably\n> means it was eaten by the 100K filter).\n\nAhh. I didn't even think about a size filter.\n\nDoesn't matter, your version looks fine.\n\n           Linus\n"},{"id":"312485","messageId":"20170224000143.cate5yncjq74hsys@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CAGZ79kZHPdBTKEqJeAa5xDcsC5v9x4DdUuDOiRNSgOV5aCx9Kw@mail.gmail.com","subject":"Re: [PATCH 1/3] add collision-detecting sha1 implementation","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-24T00:01:43Z","receivedAt":"2017-02-24T00:01:50Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 23, 2017 at 03:15:11PM -0800, Stefan Beller wrote:\n\n> On Thu, Feb 23, 2017 at 3:05 PM, Jeff King <peff@peff.net> wrote:\n> \n> > +* Copyright 2017 Marc Stevens <marc@marc-stevens.nl>, Dan Shumow (danshu@microsoft.com)\n> > +* Distributed under the MIT Software License.\n> > +* See accompanying file LICENSE.txt or copy at\n> \n> The accompanying LICENSE file did not make it into this patch,\n> that is more specialized/verbose than the one at\n> https://opensource.org/licenses/MIT\n> w.r.t. copyright notice requirement.\n\nYou know, I didn't even look at the LICENSE file, since it said MIT and\nhad a link here. It would be trivial to copy it over, too, of course.\n\n> Apart from that MIT seems to be compatible with GPL\n> according to the FSF, though IANAL.\n\nYeah, that's always been my understanding.\n\n-Peff\n"},{"id":"312487","messageId":"CA+55aFx4gXEXqr-z12E4gz5fm2fye-vVio_Td7EbKxymGF2QUw@mail.gmail.com","threadId":"45202","inReplyTo":"20170224000143.cate5yncjq74hsys@sigill.intra.peff.net","subject":"Re: [PATCH 1/3] add collision-detecting sha1 implementation","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-24T00:12:01Z","receivedAt":"2017-02-24T00:12:43Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Feb 23, 2017 at 4:01 PM, Jeff King <peff@peff.net> wrote:\n>\n> You know, I didn't even look at the LICENSE file, since it said MIT and\n> had a link here. It would be trivial to copy it over, too, of course.\n\nYou should do it. It's just good to be careful and clear with\nlicenses, and the license text does require that the copyright notice\nand permission file should be included in copies.\n\nMy patch did it. \"Pats self on head\".\n\n             Linus\n\nPS. And just to be polite, we should probably also just cc at least\nMarc Stevens and Dan Shumow if we take that patch further. Their email\naddresses are in the that LICENSE.txt file.\n"},{"id":"312489","messageId":"20170224001602.vks3lvbrbarcccmk@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CA+55aFx4gXEXqr-z12E4gz5fm2fye-vVio_Td7EbKxymGF2QUw@mail.gmail.com","subject":"Re: [PATCH 1/3] add collision-detecting sha1 implementation","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-24T00:16:02Z","receivedAt":"2017-02-24T00:23:53Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Feb 23, 2017 at 04:12:01PM -0800, Linus Torvalds wrote:\n\n> On Thu, Feb 23, 2017 at 4:01 PM, Jeff King <peff@peff.net> wrote:\n> >\n> > You know, I didn't even look at the LICENSE file, since it said MIT and\n> > had a link here. It would be trivial to copy it over, too, of course.\n> \n> You should do it. It's just good to be careful and clear with\n> licenses, and the license text does require that the copyright notice\n> and permission file should be included in copies.\n> \n> My patch did it. \"Pats self on head\".\n\nAnd that's why yours crossed the 100K barrier. :)\n\nBut yeah, I agree it is better to be safe (and that's we should contact\nthe authors). I'll point them out-of-band to this thread, and cc them if\nit ends up being re-rolled.\n\n-Peff\n"},{"id":"312507","messageId":"CACsJy8AtQG8YXQ+YfSFifUxqtd==THj5weJK5jooyiRN0yamiQ@mail.gmail.com","threadId":"45202","inReplyTo":"20170223164306.spg2avxzukkggrpb@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2017-02-24T09:42:38Z","receivedAt":"2017-02-24T09:43:40Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Thu, Feb 23, 2017 at 11:43 PM, Joey Hess <id@joeyh.name> wrote:\n> IIRC someone has been working on parameterizing git's SHA1 assumptions\n> so a repository could eventually use a more secure hash. How far has\n> that gotten? There are still many \"40\" constants in git.git HEAD.\n\nMichael asked Brian (that \"someone\") the other day and he replied [1]\n\n>> I'm curious; what fraction of the overall convert-to-object_id campaign\n>> do you estimate is done so far? Are you getting close to the promised\n>> land yet?\n>\n> So I think that the current scope left is best estimated by the\n> following command:\n>\n>   git grep -P 'unsigned char\\s+(\\*|.*20)' | grep -v '^Documentation'\n>\n> So there are approximately 1200 call sites left, which is quite a bit of\n> work.  I estimate between the work I've done and other people's\n> refactoring work (such as the refs backend refactor), we're about 40%\n> done.\n\n[1] http://public-inbox.org/git/%3C20170217214513.giua5ksuiqqs2laj@genre.crustytoothpaste.net%3E/\n-- \nDuy\n"},{"id":"312516","messageId":"CAMuHMdWbMFVcp-OK6BU0ey6tHC=QVaSFJWGd_Ge2aUsrAEczxA@mail.gmail.com","threadId":"45202","inReplyTo":"CANv4PNmSjJUhFgC7GhpuBjiSQhwfAhrP8WxiP_siP2AjjXnrnw@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Geert Uytterhoeven","fromEmail":"geert@linux-m68k.org","sentAt":"2017-02-24T15:52:24Z","receivedAt":"2017-02-24T15:55:27Z","isPatch":false,"sender":{"key":"geert@linux-m68k.org","avatar":"https://gravatar.com/avatar/8105b34f653a7b5b98e225e565b11ebcc762ad4ab1a9d905a4663db029a9e6bc?d=mp&s=160"},"body":"On Thu, Feb 23, 2017 at 8:13 PM, Morten Welinder <mwelinder@gmail.com> wrote:\n> The attack seems to generate two 64-bytes blocks, one quarter of which\n> is repeated data.  (Table-1 in the paper.)\n>\n> Assuming the result of that is evenly distributed and that bytes are\n> independent, we can estimate the chances that the result is NUL-free\n> as (255/256)^192 = 47% and the probability that the result is NUL and\n> newline free as (254/256)^192 = 22%.  Clearly one should not rely of\n> NULs or newlines to save the day.  On  the other hand, the chances of\n> an ascii result is something like (95/256)^192 = 10^-83.\n\nGood. So they can replace linux/Documentation/logo.gif, but not actual source\nfiles, not even if they contain hex arrays with \"device parameters\" ;-)\n\nGr{oetje,eeting}s,\n\n                        Geert\n\n--\nGeert Uytterhoeven -- There's lots of Linux beyond ia32 -- geert@linux-m68k.org\n\nIn personal conversations with technical people, I call myself a hacker. But\nwhen I'm talking to journalists I just say \"programmer\" or something like that.\n                                -- Linus Torvalds\n"},{"id":"312517","messageId":"22704.19873.860148.22472@chiark.greenend.org.uk","threadId":"45202","inReplyTo":"20170223164306.spg2avxzukkggrpb@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-02-24T15:13:37Z","receivedAt":"2017-02-24T16:27:42Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"Joey Hess writes (\"SHA1 collisions found\"):\n> https://shattered.io/static/shattered.pdf\n> https://freedom-to-tinker.com/2017/02/23/rip-sha-1/\n> \n> IIRC someone has been working on parameterizing git's SHA1 assumptions\n> so a repository could eventually use a more secure hash. How far has\n> that gotten? There are still many \"40\" constants in git.git HEAD.\n\nI have been thinking about how to do a transition from SHA1 to another\nhash function.\n\nI have concluded that:\n\n * We can should avoid expecting everyone to rewrite all their\n   history.\n\n * Unfortunately, because the data formats (particularly, the commit\n   header) are not in practice extensible (because of the way existing\n   code parses them), it is not useful to try generate new data (new\n   commits etc.) containing both new hashes and old hashes: old\n   clients will mishandle the new data.\n\n * Therefore the transition needs to be done by giving every object\n   two names (old and new hash function).  Objects may refer to each\n   other by either name, but must pick one.  The usual shape of\n   project histories will be a pile of old commits referring to each\n   other by old names, surmounted by new commits referrring to each\n   other by new names.\n\n * It is not possible to solve this problem without extending the\n   object name format.  Therefore all software which calls git and\n   expects to handle object names will need to be updated.\n\nI have been writing a more detailed transition plan.  I hope to post\nthis within a few days.\n\nIan.\n"},{"id":"312520","messageId":"CA+dhYEUfdHAeKwE4Sx4RivX2Vxv55xCxnsPUrPBPZWWbp=d_eA@mail.gmail.com","threadId":"45202","inReplyTo":"22704.19873.860148.22472@chiark.greenend.org.uk","subject":"Re: SHA1 collisions found","fromName":"ankostis","fromEmail":"ankostis@gmail.com","sentAt":"2017-02-24T17:04:12Z","receivedAt":"2017-02-24T17:06:13Z","isPatch":false,"sender":{"key":"ankostis@gmail.com","avatar":"https://gravatar.com/avatar/1f3597ab8ad44cbb0fd782a2773e3aba94a52c03f9f6bbe5ccb7f1fe2581b12b?d=mp&s=160"},"body":"On 24 February 2017 at 16:13, Ian Jackson\n<ijackson@chiark.greenend.org.uk> wrote:\n>\n> Joey Hess writes (\"SHA1 collisions found\"):\n> > https://shattered.io/static/shattered.pdf\n> > https://freedom-to-tinker.com/2017/02/23/rip-sha-1/\n> >\n> > IIRC someone has been working on parameterizing git's SHA1 assumptions\n> > so a repository could eventually use a more secure hash. How far has\n> > that gotten? There are still many \"40\" constants in git.git HEAD.\n>\n> I have been thinking about how to do a transition from SHA1 to another\n> hash function.\n>\n> I have concluded that:\n>\n>  * We can should avoid expecting everyone to rewrite all their\n>    history.\n>\n>  * Unfortunately, because the data formats (particularly, the commit\n>    header) are not in practice extensible (because of the way existing\n>    code parses them), it is not useful to try generate new data (new\n>    commits etc.) containing both new hashes and old hashes: old\n>    clients will mishandle the new data.\n>\n>  * Therefore the transition needs to be done by giving every object\n>    two names (old and new hash function).  Objects may refer to each\n>    other by either name, but must pick one.  The usual shape of\n>    project histories will be a pile of old commits referring to each\n>    other by old names, surmounted by new commits referrring to each\n>    other by new names.\n>\n>  * It is not possible to solve this problem without extending the\n>    object name format.  Therefore all software which calls git and\n>    expects to handle object names will need to be updated.\n>\n> I have been writing a more detailed transition plan.  I hope to post\n> this within a few days.\n\nIt would be great to have a rough plan of the transition to a new hash\nfunction.\n\nWe are writing a git-based application to store electronic-files for\nlegislative purposes for EU.\nAnd one of the great questions we face is about git's SHA-1 validity in 5\nor 20 years of time from now.\n\nIs it possible to have an assessment of the situation for this transition?\n\nBest regards for your efforts,\n  Kostis\n"},{"id":"312521","messageId":"xmqq60jz5wbm.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"22704.19873.860148.22472@chiark.greenend.org.uk","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-24T17:32:13Z","receivedAt":"2017-02-24T17:32:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Ian Jackson <ijackson@chiark.greenend.org.uk> writes:\n\n> I have been thinking about how to do a transition from SHA1 to another\n> hash function.\n\nGood.  I think many of us have also been, too, not necessarily just\nin the past few days in response to shattered, but over the last 10\nyears, yet without coming to a consensus design ;-)\n\n> I have concluded that:\n>\n>  * We can should avoid expecting everyone to rewrite all their\n>    history.\n\nYes.\n\n>  * Unfortunately, because the data formats (particularly, the commit\n>    header) are not in practice extensible (because of the way existing\n>    code parses them), it is not useful to try generate new data (new\n>    commits etc.) containing both new hashes and old hashes: old\n>    clients will mishandle the new data.\n\nYes.\n\n>  * Therefore the transition needs to be done by giving every object\n>    two names (old and new hash function).  Objects may refer to each\n>    other by either name, but must pick one.  The usual shape of\n\nI do not think it is necessrily so.  Existing code may not be able\nto read anything new, but you can make the new code understand\nobject names in both formats, and for a smooth transition, I think\nthe new code needs to.\n\nFor example, a new commit that records a merge of an old and a new\ncommit whose resulting tree happens to be the same as the tree of\nthe old commit may begin like so:\n\n    tree 21b97d4c4f968d1335f16292f954dfdbb91353f0\n    parent 20769079d22a9f8010232bdf6131918c33a1bf6910232bdf6131918c33a1bf69\n    parent 22af6fef9b6538c9e87e147a920be9509acf1ddd\n\nnaming the only object whose name was done with new hash with the\nnew longer hash, while recording the names of the other existing\nobjects with SHA-1.  We would need to extend the object format for\ntag (which would be trivial as the object reference is textual and\nsimilar to a commit) and tree (much harder), of course.\n\nAs long as the reader can tell from the format of object names\nstored in the \"new object format\" object from what era is being\nreferred to in some way [*1*], we can name new objects with only new\nhash, I would think.  \"new refers only to new\" that stratifies\nobjects into older and newer may make things simpler, but I am not\nconvinced yet that it would give our users a smooth enough\ntransition path (but I am open to be educated and pursuaded the\nother way).\n\n\n[Footnote]\n\n*1* In the above toy example, length being 40 vs 64 is used as a\n    sign between SHA-1 and the new hash, and careful readers may\n    wonder if we should use sha-3,20769079d22... or something like\n    that that more explicity identifies what hash is used, so that\n    we can pick a hash whose length is 64 when we transition again.\n\n    I personally do not think such a prefix is necessary during the\n    first transition; we will likely to adopt a new hash again, and\n    at that point that third one can have a prefix to differenciate\n    it from the second one.\n"},{"id":"312523","messageId":"20170224172335.GG11350@io.lakedaemon.net","threadId":"45202","inReplyTo":"22704.19873.860148.22472@chiark.greenend.org.uk","subject":"Re: SHA1 collisions found","fromName":"Jason Cooper","fromEmail":"git@lakedaemon.net","sentAt":"2017-02-24T17:23:35Z","receivedAt":"2017-02-24T17:39:52Z","isPatch":false,"sender":{"key":"git@lakedaemon.net","avatar":null},"body":"Hi Ian,\n\nOn Fri, Feb 24, 2017 at 03:13:37PM +0000, Ian Jackson wrote:\n> Joey Hess writes (\"SHA1 collisions found\"):\n> > https://shattered.io/static/shattered.pdf\n> > https://freedom-to-tinker.com/2017/02/23/rip-sha-1/\n> > \n> > IIRC someone has been working on parameterizing git's SHA1 assumptions\n> > so a repository could eventually use a more secure hash. How far has\n> > that gotten? There are still many \"40\" constants in git.git HEAD.\n> \n> I have been thinking about how to do a transition from SHA1 to another\n> hash function.\n> \n> I have concluded that:\n> \n>  * We can should avoid expecting everyone to rewrite all their\n>    history.\n\nAgreed.\n\n>  * Unfortunately, because the data formats (particularly, the commit\n>    header) are not in practice extensible (because of the way existing\n>    code parses them), it is not useful to try generate new data (new\n>    commits etc.) containing both new hashes and old hashes: old\n>    clients will mishandle the new data.\n\nMy thought here is:\n\n a) re-hash blobs with sha256, hardlink to sha1 objects\n b) create new tree objects which are mirrors of each sha1 tree object,\n    but purely sha256\n c) mirror commits, but they are also purely sha256\n d) future PGP signed tags would sign both hashes (or include both?)\n\nWhich would end up something like:\n\n  .git/\n    \\... #usual files\n    \\objects\n      \\ef\n        \\3c39f7522dc55a24f64da9febcfac71e984366\n    \\objects-sha2_256\n      \\72\n        \\604fd2de5f25c89d692b01081af93bcf00d2af34549d8d1bdeb68bc048932\n    \\info\n      \\...\n    \\info-sha2_256\n      \\refs #uses sha256 commit identifiers\n\nBasically, keep the sha256 stuff out of the way for legacy clients, and\nnew clients will still be able to use it.\n\nThere shouldn't be a need to re-sign old signed tags if the underlying\nobjects are counter-hashed.  There might need to be some transition\ninfo, though.\n\nSay a new client does 'git tag -v tags/v3.16' in the kernel tree.  I would\nexpect it to check the sha1 hashes, verify the PGP signed tag, and then\nalso check the sha256 counter-hashes of the relevant objects.\n\nthx,\n\nJason.\n"},{"id":"312525","messageId":"nycvar.QRO.7.75.62.1702240943540.6590@qynat-yncgbc","threadId":"45202","inReplyTo":"xmqq60jz5wbm.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2017-02-24T17:45:55Z","receivedAt":"2017-02-24T17:49:40Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 24 Feb 2017, Junio C Hamano wrote:\n\n> *1* In the above toy example, length being 40 vs 64 is used as a\n>    sign between SHA-1 and the new hash, and careful readers may\n>    wonder if we should use sha-3,20769079d22... or something like\n>    that that more explicity identifies what hash is used, so that\n>    we can pick a hash whose length is 64 when we transition again.\n>\n>    I personally do not think such a prefix is necessary during the\n>    first transition; we will likely to adopt a new hash again, and\n>    at that point that third one can have a prefix to differenciate\n>    it from the second one.\n\nas the saying goes \"in computer science the interesting numbers are 0, 1, and \nmany\", does it really simplify things much to support 2 hashes vs supporting \nmore so that this issue doesn't have to be revisited? (other than selecting new \nhashes over time)\n\nDavid Lang\n"},{"id":"312529","messageId":"xmqqk28f4fti.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"nycvar.QRO.7.75.62.1702240943540.6590@qynat-yncgbc","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-24T18:14:01Z","receivedAt":"2017-02-24T18:14:13Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"David Lang <david@lang.hm> writes:\n\n> On Fri, 24 Feb 2017, Junio C Hamano wrote:\n>\n>> *1* In the above toy example, length being 40 vs 64 is used as a\n>>    sign between SHA-1 and the new hash, and careful readers may\n>>    wonder if we should use sha-3,20769079d22... or something like\n>>    that that more explicity identifies what hash is used, so that\n>>    we can pick a hash whose length is 64 when we transition again.\n>>\n>>    I personally do not think such a prefix is necessary during the\n>>    first transition; we will likely to adopt a new hash again, and\n>>    at that point that third one can have a prefix to differenciate\n>>    it from the second one.\n>\n> as the saying goes \"in computer science the interesting numbers are 0,\n> 1, and many\", does it really simplify things much to support 2 hashes\n> vs supporting more so that this issue doesn't have to be revisited?\n> (other than selecting new hashes over time)\n\nIt seems that I wasn't clear enough, perhaps?  The scheme I outlined\ndoes not have to revisit this issue at all.  It already declares what\nyou need to do when you add the third one.  \n\nIf it is not 40 or 64 bytes long, you just write it out.  If it is\none of these length, then you add some identifying prefix or\npostfix.  IOW, if the second one is sha-3 and the third one is blake\n(both used at 256-bit), then we would have three kinds of names,\nwritten like so:\n\n    20769079d22a9f8010232bdf6131918c33a1bf69\n    20769079d22a9f8010232bdf6131918c33a1bf6910232bdf6131918c33a1bf69\n    3,20769079d22a9f8010232bdf6131918c33a1bf6910232bdf6131918c33a1bf69\n\nand the readers can well tell that the first one, being 40-chars\nlong, is SHA-1, the second one, being 64-chars long, is SHA-3, and\nthe last one, with the prefix '3' (only because that is the third\none officially supported by Git) and being 64-chars long, is blake,\nfor example.\n\nI do not particularly care if it is prefix or postfix or something\nelse.  A not-so-well-hidden agenda is to avoid inviting people into\nthinking that they can use their choice of random hash functions and\nand claim that their hacked version is still a Git, as long as they\nfollow the object naming convention.  IOW, if you said something\nlike:\n\n * 40-hex is SHA-1 for historical reasons;\n * Others use hash-name, colon, and then N-hex.\n\nyou are inviting people to start using\n\n    md5,54ddf8d47340e048166c45f439ce65fd\n\nas object names.\n"},{"id":"312532","messageId":"20170224185754.bic37suvkiwtvyn5@sigill.intra.peff.net","threadId":"45202","inReplyTo":"16c6d843-a516-9265-d3e7-61b110acbdcf@ipsumj.de","subject":"Re: [PATCH 3/3] Makefile: add USE_SHA1DC knob","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-24T18:57:55Z","receivedAt":"2017-02-24T18:58:02Z","isPatch":true,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 24, 2017 at 06:36:00PM +0000, HW42 wrote:\n\n> > +ifdef USE_SHA1DC\n> > +\tSHA1_HEADER = \"sha1dc/sha1.h\"\n> > +\tLIB_OBJS += sha1dc/sha1.o\n> > +\tLIB_OBJS += sha1dc/ubc_check.o\n> > +else\n> >  ifdef BLK_SHA1\n> >  \tSHA1_HEADER = \"block-sha1/sha1.h\"\n> >  \tLIB_OBJS += block-sha1/sha1.o\n> > @@ -1403,6 +1412,7 @@ else\n> >  endif\n> >  endif\n> >  endif\n> > +endif\n> \n> This sets SHA1_MAX_BLOCK_SIZE and the compiler flags for Apple\n> CommonCrypto even if the user selects USE_SHA1DC. The same happens for\n> BLK_SHA1. Is this intended?\n\nNo, it's not. I suspect that setting BLK_SHA1 has the same problem in\nthe current code, then.\n\n> > +void git_SHA1DCUpdate(SHA1_CTX *ctx, const void *vdata, unsigned long len)\n> > +{\n> > +\tconst char *data = vdata;\n> > +\t/* We expect an unsigned long, but sha1dc only takes an int */\n> > +\twhile (len > INT_MAX) {\n> > +\t\tSHA1DCUpdate(ctx, data, INT_MAX);\n> > +\t\tdata += INT_MAX;\n> > +\t\tlen -= INT_MAX;\n> > +\t}\n> > +\tSHA1DCUpdate(ctx, data, len);\n> > +}\n> \n> I think you can simply change the len parameter from unsigned into\n> size_t (or unsigned long) in SHA1DCUpdate().\n> https://github.com/cr-marcstevens/sha1collisiondetection/pull/6\n\nYeah, I agree that is a cleaner solution. My focus was on changing the\n(presumably tested) sha1dc code as little as possible.\n\n-Peff\n"},{"id":"312533","messageId":"CAGZ79kaZWe-8pMZnQv7uZtr8wXWawFeJjUa68-b0oa4yFo-HcA@mail.gmail.com","threadId":"45202","inReplyTo":"xmqqk28f4fti.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2017-02-24T18:58:28Z","receivedAt":"2017-02-24T18:58:36Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Fri, Feb 24, 2017 at 10:14 AM, Junio C Hamano <gitster@pobox.com> wrote:\n\n> you are inviting people to start using\n>\n>     md5,54ddf8d47340e048166c45f439ce65fd\n>\n> as object names.\n\nwhich might even be okay for specific subsets of operations.\n(e.g. all local work including staging things, making local \"fixup\" commits)\n\nThe addressing scheme should not be too hardcoded, we should rather\ntreat it similar to the cipher schemes in pgp. The additional complexity that\nwe have is the longevity of existence of things, though.\n"},{"id":"312536","messageId":"xmqq7f4f4cqg.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"CAGZ79kaZWe-8pMZnQv7uZtr8wXWawFeJjUa68-b0oa4yFo-HcA@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-24T19:20:39Z","receivedAt":"2017-02-24T19:21:13Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Stefan Beller <sbeller@google.com> writes:\n\n> On Fri, Feb 24, 2017 at 10:14 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n>> you are inviting people to start using\n>>\n>>     md5,54ddf8d47340e048166c45f439ce65fd\n>>\n>> as object names.\n>\n> which might even be okay for specific subsets of operations.\n> (e.g. all local work including staging things, making local \"fixup\" commits)\n>\n> The addressing scheme should not be too hardcoded, we should rather\n> treat it similar to the cipher schemes in pgp. The additional complexity that\n> we have is the longevity of existence of things, though.\n\nThe not-so-well-hidden agenda was exactly that we _SHOULD_ not\nmimick PGP.  They do not have a requirement to encourage everybody\nto use the same thing because each message is encrypted/signed\nindependently, i.e. they do not have to chain things like we do.\n\n"},{"id":"312538","messageId":"16c6d843-a516-9265-d3e7-61b110acbdcf@ipsumj.de","threadId":"45202","inReplyTo":"20170223230621.43anex65ndoqbgnf@sigill.intra.peff.net","subject":"Re: [PATCH 3/3] Makefile: add USE_SHA1DC knob","fromName":"HW42","fromEmail":"hw42@ipsumj.de","sentAt":"2017-02-24T18:36:00Z","receivedAt":"2017-02-24T19:25:04Z","isPatch":true,"sender":{"key":"hw42@ipsumj.de","avatar":null},"body":"Jeff King:\n> diff --git a/Makefile b/Makefile\n> index 8e4081e06..7c4906250 100644\n> --- a/Makefile\n> +++ b/Makefile\n> @@ -1386,6 +1390,11 @@ ifdef APPLE_COMMON_CRYPTO\n>  \tSHA1_MAX_BLOCK_SIZE = 1024L*1024L*1024L\n>  endif\n>  \n> +ifdef USE_SHA1DC\n> +\tSHA1_HEADER = \"sha1dc/sha1.h\"\n> +\tLIB_OBJS += sha1dc/sha1.o\n> +\tLIB_OBJS += sha1dc/ubc_check.o\n> +else\n>  ifdef BLK_SHA1\n>  \tSHA1_HEADER = \"block-sha1/sha1.h\"\n>  \tLIB_OBJS += block-sha1/sha1.o\n> @@ -1403,6 +1412,7 @@ else\n>  endif\n>  endif\n>  endif\n> +endif\n\nThis sets SHA1_MAX_BLOCK_SIZE and the compiler flags for Apple\nCommonCrypto even if the user selects USE_SHA1DC. The same happens for\nBLK_SHA1. Is this intended?\n\n> +void git_SHA1DCUpdate(SHA1_CTX *ctx, const void *vdata, unsigned long len)\n> +{\n> +\tconst char *data = vdata;\n> +\t/* We expect an unsigned long, but sha1dc only takes an int */\n> +\twhile (len > INT_MAX) {\n> +\t\tSHA1DCUpdate(ctx, data, INT_MAX);\n> +\t\tdata += INT_MAX;\n> +\t\tlen -= INT_MAX;\n> +\t}\n> +\tSHA1DCUpdate(ctx, data, len);\n> +}\n\nI think you can simply change the len parameter from unsigned into\nsize_t (or unsigned long) in SHA1DCUpdate().\nhttps://github.com/cr-marcstevens/sha1collisiondetection/pull/6\n\n"},{"id":"312545","messageId":"xmqqpoi71hi7.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"xmqq7f4f4cqg.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-24T20:05:52Z","receivedAt":"2017-02-24T20:06:19Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> The not-so-well-hidden agenda was exactly that we _SHOULD_ not\n> mimick PGP.  They do not have a requirement to encourage everybody\n> to use the same thing because each message is encrypted/signed\n> independently, i.e. they do not have to chain things like we do.\n\nTo put it less succinctly, PGP does not have incentive to encourage\neverybody to converge to the same.  They can afford to say \"You can\nuse whatever you among your circles agree to use and the rest of the\nworld won't care\".  If two groups that have used different ones later\nmeet, both of them can switch to a common one from that point forward,\nbut their past exchanges won't affect the future.\n\nYou cannot say the same thing for Git.  Once you decide to merge two\nhistories from two camps, which may have originated from the same\ncodebase but then decided to use two different ones while they were\nforked, you'd be forced to support all three forever.  We have a lot\nstronger incentive to discourage fragmentation.\n\n\n\n"},{"id":"312546","messageId":"CA+dhYEVOyACM9ARP2deKVLm1hHOVsTah1WfGoNzGGKO6CGrQpw@mail.gmail.com","threadId":"45202","inReplyTo":"xmqq7f4f4cqg.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"ankostis","fromEmail":"ankostis@gmail.com","sentAt":"2017-02-24T20:05:40Z","receivedAt":"2017-02-24T20:06:29Z","isPatch":false,"sender":{"key":"ankostis@gmail.com","avatar":"https://gravatar.com/avatar/1f3597ab8ad44cbb0fd782a2773e3aba94a52c03f9f6bbe5ccb7f1fe2581b12b?d=mp&s=160"},"body":"On 24 February 2017 at 20:20, Junio C Hamano <gitster@pobox.com> wrote:\n> Stefan Beller <sbeller@google.com> writes:\n>\n>> On Fri, Feb 24, 2017 at 10:14 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>>> you are inviting people to start using\n>>>\n>>>     md5,54ddf8d47340e048166c45f439ce65fd\n>>>\n>>> as object names.\n>>\n>> which might even be okay for specific subsets of operations.\n>> (e.g. all local work including staging things, making local \"fixup\" commits)\n>>\n>> The addressing scheme should not be too hardcoded, we should rather\n>> treat it similar to the cipher schemes in pgp. The additional complexity that\n>> we have is the longevity of existence of things, though.\n>\n> The not-so-well-hidden agenda was exactly that we _SHOULD_ not\n> mimick PGP.  They do not have a requirement to encourage everybody\n> to use the same thing because each message is encrypted/signed\n> independently, i.e. they do not have to chain things like we do.\n\nBut there is a scenario where supporting more hashes, in parallel, is\nbeneficial:\n\nLet's assume that git is retroffited to always support the \"default\"\nSHA-3, but support additionally more hash-funcs.\nIf in the future SHA-3 also gets defeated, it would be highly unlikely\nthat the same math would also break e.g. Blake.\nSo certain high-profile repos might choose for extra security 2 or more hashes.\n\nApologies if I'm misusing the list,\n  Kostis\n"},{"id":"312549","messageId":"xmqqh93j1g9n.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"CA+dhYEVOyACM9ARP2deKVLm1hHOVsTah1WfGoNzGGKO6CGrQpw@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-24T20:32:36Z","receivedAt":"2017-02-24T20:32:51Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"ankostis <ankostis@gmail.com> writes:\n\n> Let's assume that git is retroffited to always support the \"default\"\n> SHA-3, but support additionally more hash-funcs.\n> If in the future SHA-3 also gets defeated, it would be highly unlikely\n> that the same math would also break e.g. Blake.\n> So certain high-profile repos might choose for extra security 2 or more hashes.\n\nI think you are conflating two unrelated things.\n\n * How are these \"2 or more hashes\" actually used?  Are you going to\n   add three \"parent \" line to a commit with just one parent, each\n   line storing the different hashes?  How will such a commit object\n   be named---does it have three names and do you plan to have three\n   copies of .git/refs/heads/master somehow, each of which have\n   SHA-1, SHA-3 and Blake, and let any one hash to identify the\n   object?\n\n   I suspect you are not going to do so; instead, you would use a\n   very long string that is a concatenation of these three hashes as\n   if it is an output from a single hash function that produces a\n   long result.\n\n   So I think the most natural way to do the \"2 or more for extra\n   security\" is to allow us to use a very long hash.  It does not\n   help to allow an object to be referred to with any of these 2 or\n   more hashes at the same time.\n\n * If employing 2 or more hashes by combining into one may enhance\n   the security, that is wonderful.  But we want to discourage\n   people from inventing their own combinations left and right and\n   end up fragmenting the world.  If a project that begins with\n   SHA-1 only naming is forked to two (or more) and each fork uses\n   different hashes, merging them back will become harder than\n   necessary unless you support all these hashes forks used.\n\nHaving said all that, the way to figure out the hash used in the way\nwe spell the object name may not be the best place to discourage\npeople from using random hashes of their choice.  But I think we\nwant to avoid doing something that would actively encourage\nfragmentation.\n\n"},{"id":"312550","messageId":"E817E45267184196A3832C3220075D14@PhilipOakley","threadId":"45202","inReplyTo":"CAGZ79kaZWe-8pMZnQv7uZtr8wXWawFeJjUa68-b0oa4yFo-HcA@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":null,"receivedAt":"2017-02-24T20:33:46Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"From: \"Stefan Beller\" <sbeller@google.com>\n> On Fri, Feb 24, 2017 at 10:14 AM, Junio C Hamano <gitster@pobox.com> \n> wrote:\n>\n>> you are inviting people to start using\n>>\n>>     md5,54ddf8d47340e048166c45f439ce65fd\n>>\n>> as object names.\n>\n> which might even be okay for specific subsets of operations.\n> (e.g. all local work including staging things, making local \"fixup\" \n> commits)\n>\n> The addressing scheme should not be too hardcoded, we should rather\n> treat it similar to the cipher schemes in pgp. The additional complexity \n> that\n> we have is the longevity of existence of things, though.\n>\n\nOne potential nicety of using the md5 is that it is a known `toy problem` \nsolution that could be used to explore how things might be made to work, \nwithout any expectation that the temporary code is in any way an \nexperimental part of regular code. Maybe. It's good to have a toy problem to \nwork on.\n\nThere are other issue to be considered as well, such as validating a \ntransition of identical blobs and trees (at some point there will for some \nusers be a forced update of hash of unchanged code), which probably requires \ntwo way traversal.\n\nPhilip \n\n"},{"id":"312575","messageId":"9cedbfa5-4095-15d8-639c-0e3b9b98d6b9@gmail.com","threadId":"45202","inReplyTo":"20170223164306.spg2avxzukkggrpb@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2017-02-24T22:47:46Z","receivedAt":"2017-02-24T22:48:31Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"I have just read on ArsTechnica[1] that while Git repository could be\ncorrupted (though this would require attackers to spend great amount\nof resources creating their own collision, while as said elsewhere\nin this thread allegedly easy to detect), putting two proof-of-concept\ndifferent PDFs with same size and SHA-1 actually *breaks* Subversion.\nRepository can become corrupt, and stop accepting new commits.  \n\nFrom what I understand people tried this, and Git doesn't exhibit\nsuch problem.  I wonder what assumptions SVN made that were broken...\n\nThe https://shattered.io/ page updated their Q&A section with this\ninformation.\n\nBTW. what's with that page use of \"GIT\" instead of \"Git\"??\n\n\n[1]: https://arstechnica.com/security/2017/02/watershed-sha1-collision-just-broke-the-webkit-repository-others-may-follow/ \n     \"Watershed SHA1 collision just broke the WebKit repository, others may follow\"\n"},{"id":"312576","messageId":"20170224225350.xb7rudyhowmsqdbc@LykOS.localdomain","threadId":"45202","inReplyTo":"9cedbfa5-4095-15d8-639c-0e3b9b98d6b9@gmail.com","subject":"Re: SHA1 collisions found","fromName":"Santiago Torres","fromEmail":"santiago@nyu.edu","sentAt":"2017-02-24T22:53:51Z","receivedAt":"2017-02-24T22:53:57Z","isPatch":false,"sender":{"key":"santiago@nyu.edu","avatar":"https://avatars.githubusercontent.com/u/3579933?v=4"},"body":"On Fri, Feb 24, 2017 at 11:47:46PM +0100, Jakub Narębski wrote:\n> I have just read on ArsTechnica[1] that while Git repository could be\n> corrupted (though this would require attackers to spend great amount\n> of resources creating their own collision, while as said elsewhere\n> in this thread allegedly easy to detect), putting two proof-of-concept\n> different PDFs with same size and SHA-1 actually *breaks* Subversion.\n> Repository can become corrupt, and stop accepting new commits.  \n\nFrom what I understood in the thread[1], it was the combination of svn +\ngit-svn together. I think Arstechnica may be a little bit\nsensationalistic here.\n\nCheers!\n-Santiago.\n\n[1] https://bugs.webkit.org/show_bug.cgi?id=168774#c27\n"},{"id":"312577","messageId":"e0ad3c81-aa2c-2eea-eb9e-17591b6b592c@gmail.com","threadId":"45202","inReplyTo":"20170224225350.xb7rudyhowmsqdbc@LykOS.localdomain","subject":"Re: SHA1 collisions found","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2017-02-24T23:05:34Z","receivedAt":"2017-02-24T23:06:12Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 24.02.2017 o 23:53, Santiago Torres pisze:\n> On Fri, Feb 24, 2017 at 11:47:46PM +0100, Jakub Narębski wrote:\n>>\n>> I have just read on ArsTechnica[1] that while Git repository could be\n>> corrupted (though this would require attackers to spend great amount\n>> of resources creating their own collision, while as said elsewhere\n>> in this thread allegedly easy to detect), putting two proof-of-concept\n>> different PDFs with same size and SHA-1 actually *breaks* Subversion.\n>> Repository can become corrupt, and stop accepting new commits.  \n> \n> From what I understood in the thread[1], it was the combination of svn +\n> git-svn together. I think Arstechnica may be a little bit\n> sensationalistic here.\n \n> [1] https://bugs.webkit.org/show_bug.cgi?id=168774#c27\n\nThanks for the link.  It looks like the problem was with svn itself\n(couldn't checkout, couldn't sync), but repository is recovered now,\nthough not protected against the problem occurring again.\n\nWell, anyone with Subversion installed (so not me) can check it\nfor himself/herself... though better do this with separate svnroot.\n\n\nNote that the breakage was an accident, trying to add test case\nfor SHA-1 collision in WebKit cache.\n \nBest regards,\n-- \nJakub Narębski\n"},{"id":"312578","messageId":"20170224230604.nt37uw5y3uehukfd@sigill.intra.peff.net","threadId":"45202","inReplyTo":"9cedbfa5-4095-15d8-639c-0e3b9b98d6b9@gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-24T23:06:05Z","receivedAt":"2017-02-24T23:06:13Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 24, 2017 at 11:47:46PM +0100, Jakub Narębski wrote:\n\n> I have just read on ArsTechnica[1] that while Git repository could be\n> corrupted (though this would require attackers to spend great amount\n> of resources creating their own collision, while as said elsewhere\n> in this thread allegedly easy to detect), putting two proof-of-concept\n> different PDFs with same size and SHA-1 actually *breaks* Subversion.\n> Repository can become corrupt, and stop accepting new commits.  \n> \n> From what I understand people tried this, and Git doesn't exhibit\n> such problem.  I wonder what assumptions SVN made that were broken...\n\nTo be clear, nobody has generated a sha1 collision in Git yet, and you\ncannot blindly use the shattered PDFs to do so. Git's notion of the\nSHA-1 of an object include the header, so somebody would have to do a\nshattered-level collision search for something that starts with the\ncorrect \"blob 1234\\0\" header.\n\nSo we don't actually know how Git would behave in the face of a SHA-1\ncollision. It would be pretty easy to simulate it with something like:\n\n---\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nindex 22b125cf8..1be5b5ba3 100644\n--- a/block-sha1/sha1.c\n+++ b/block-sha1/sha1.c\n@@ -231,6 +231,16 @@ void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n \t\tmemcpy(ctx->W, data, len);\n }\n \n+/* sha1 of blobs containing \"foo\\n\" and \"bar\\n\" */\n+static const unsigned char foo_sha1[] = {\n+\t0x25, 0x7c, 0xc5, 0x64, 0x2c, 0xb1, 0xa0, 0x54, 0xf0, 0x8c,\n+\t0xc8, 0x3f, 0x2d, 0x94, 0x3e, 0x56, 0xfd, 0x3e, 0xbe, 0x99\n+};\n+static const unsigned char bar_sha1[] = {\n+\t0x57, 0x16, 0xca, 0x59, 0x87, 0xcb, 0xf9, 0x7d, 0x6b, 0xb5,\n+\t0x49, 0x20, 0xbe, 0xa6, 0xad, 0xde, 0x24, 0x2d, 0x87, 0xe6\n+};\n+\n void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n {\n \tstatic const unsigned char pad[64] = { 0x80 };\n@@ -248,4 +258,8 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n \t/* Output hash */\n \tfor (i = 0; i < 5; i++)\n \t\tput_be32(hashout + i * 4, ctx->H[i]);\n+\n+\t/* pretend \"foo\" and \"bar\" collide */\n+\tif (!memcmp(hashout, bar_sha1, 20))\n+\t\tmemcpy(hashout, foo_sha1, 20);\n }\n"},{"id":"312579","messageId":"20170224232426.3737imr4qxtlioxd@sunbase.org","threadId":"45202","inReplyTo":"e0ad3c81-aa2c-2eea-eb9e-17591b6b592c@gmail.com","subject":"Re: SHA1 collisions found","fromName":"Øyvind A. Holm","fromEmail":"sunny@sunbase.org","sentAt":"2017-02-24T23:24:28Z","receivedAt":"2017-02-24T23:25:32Z","isPatch":false,"sender":{"key":"sunny@sunbase.org","avatar":"https://avatars.githubusercontent.com/u/113445?v=4"},"body":"On 2017-02-25 00:05:34, Jakub Narębski wrote:\n> W dniu 24.02.2017 o 23:53, Santiago Torres pisze:\n> > On Fri, Feb 24, 2017 at 11:47:46PM +0100, Jakub Narębski wrote:\n> > > I have just read on ArsTechnica[1] that while Git repository could \n> > > be corrupted (though this would require attackers to spend great \n> > > amount of resources creating their own collision, while as said \n> > > elsewhere in this thread allegedly easy to detect), putting two \n> > > proof-of-concept different PDFs with same size and SHA-1 actually \n> > > *breaks* Subversion. Repository can become corrupt, and stop \n> > > accepting new commits.\n> >\n> > From what I understood in the thread[1], it was the combination of \n> > svn + git-svn together. I think Arstechnica may be a little bit \n> > sensationalistic here.\n>\n> > [1] https://bugs.webkit.org/show_bug.cgi?id=168774#c27\n>\n> Thanks for the link.  It looks like the problem was with svn itself \n> (couldn't checkout, couldn't sync), but repository is recovered now, \n> though not protected against the problem occurring again.\n>\n> Well, anyone with Subversion installed (so not me) can check it for \n> himself/herself... though better do this with separate svnroot.\n\nI tested this yesterday by adding the two PDF files to a Subversion \nrepository, and found that it wasn't able to clone (\"checkout\" in svn \nspeak) the repository after the two files had been committed. I posted \nthe results to the svn-dev mailing list, the thread is at \n<https://svn.haxx.se/dev/archive-2017-02/0142.shtml>.\n\nIt seems as it only breaks the working copy because the pristine copies \nare identified with a SHA1 sum, but the FSFS repository backend seems to \ncope with it.\n\nRegards,\nØyvind\n\n+-| Øyvind A. Holm <sunny@sunbase.org> - N 60.37604° E 5.33339° |-+\n| OpenPGP: 0xFB0CBEE894A506E5 - http://www.sunbase.org/pubkey.asc |\n| Fingerprint: A006 05D6 E676 B319 55E2  E77E FB0C BEE8 94A5 06E5 |\n+------------| 41517b2c-fae7-11e6-9521-db5caa6d21d3 |-------------+\n"},{"id":"312581","messageId":"937d395f-77fe-b275-6cbe-f3477e24cd2f@gmail.com","threadId":"45202","inReplyTo":"20170224230604.nt37uw5y3uehukfd@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2017-02-24T23:35:39Z","receivedAt":"2017-02-24T23:36:39Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"W dniu 25.02.2017 o 00:06, Jeff King pisze:\n> On Fri, Feb 24, 2017 at 11:47:46PM +0100, Jakub Narębski wrote:\n> \n>> I have just read on ArsTechnica[1] that while Git repository could be\n>> corrupted (though this would require attackers to spend great amount\n>> of resources creating their own collision, while as said elsewhere\n>> in this thread allegedly easy to detect), putting two proof-of-concept\n>> different PDFs with same size and SHA-1 actually *breaks* Subversion.\n>> Repository can become corrupt, and stop accepting new commits.  \n>>\n>> From what I understand people tried this, and Git doesn't exhibit\n>> such problem.  I wonder what assumptions SVN made that were broken...\n> \n> To be clear, nobody has generated a sha1 collision in Git yet, and you\n> cannot blindly use the shattered PDFs to do so. Git's notion of the\n> SHA-1 of an object include the header, so somebody would have to do a\n> shattered-level collision search for something that starts with the\n> correct \"blob 1234\\0\" header.\n\nWhat I meant by \"Git doesn't exhibit such problem\" (but was not clear\nenough) is that Git doesn't break by just adding SHAttered.io PDFs\n(which somebody had checked), but need customized attack.\n\n> \n> So we don't actually know how Git would behave in the face of a SHA-1\n> collision. It would be pretty easy to simulate it with something like:\n\nYou are right that it would be good to know if such Git-geared customized\nSHA-1 attack would break Git, or would it simply corrupt it (visibly\nor not).\n\n> \n> ---\n> diff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\n> index 22b125cf8..1be5b5ba3 100644\n> --- a/block-sha1/sha1.c\n> +++ b/block-sha1/sha1.c\n> @@ -231,6 +231,16 @@ void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n>  \t\tmemcpy(ctx->W, data, len);\n>  }\n>  \n> +/* sha1 of blobs containing \"foo\\n\" and \"bar\\n\" */\n> +static const unsigned char foo_sha1[] = {\n> +\t0x25, 0x7c, 0xc5, 0x64, 0x2c, 0xb1, 0xa0, 0x54, 0xf0, 0x8c,\n> +\t0xc8, 0x3f, 0x2d, 0x94, 0x3e, 0x56, 0xfd, 0x3e, 0xbe, 0x99\n> +};\n> +static const unsigned char bar_sha1[] = {\n> +\t0x57, 0x16, 0xca, 0x59, 0x87, 0xcb, 0xf9, 0x7d, 0x6b, 0xb5,\n> +\t0x49, 0x20, 0xbe, 0xa6, 0xad, 0xde, 0x24, 0x2d, 0x87, 0xe6\n> +};\n> +\n>  void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n>  {\n>  \tstatic const unsigned char pad[64] = { 0x80 };\n> @@ -248,4 +258,8 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n>  \t/* Output hash */\n>  \tfor (i = 0; i < 5; i++)\n>  \t\tput_be32(hashout + i * 4, ctx->H[i]);\n> +\n> +\t/* pretend \"foo\" and \"bar\" collide */\n> +\tif (!memcmp(hashout, bar_sha1, 20))\n> +\t\tmemcpy(hashout, foo_sha1, 20);\n>  }\n> \n\n"},{"id":"312582","messageId":"20170224233929.p2yckbc6ksyox5nu@sigill.intra.peff.net","threadId":"45202","inReplyTo":"xmqq60jz5wbm.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-24T23:39:29Z","receivedAt":"2017-02-24T23:39:36Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 24, 2017 at 09:32:13AM -0800, Junio C Hamano wrote:\n\n> >  * Therefore the transition needs to be done by giving every object\n> >    two names (old and new hash function).  Objects may refer to each\n> >    other by either name, but must pick one.  The usual shape of\n> \n> I do not think it is necessrily so.  Existing code may not be able\n> to read anything new, but you can make the new code understand\n> object names in both formats, and for a smooth transition, I think\n> the new code needs to.\n> \n> For example, a new commit that records a merge of an old and a new\n> commit whose resulting tree happens to be the same as the tree of\n> the old commit may begin like so:\n> \n>     tree 21b97d4c4f968d1335f16292f954dfdbb91353f0\n>     parent 20769079d22a9f8010232bdf6131918c33a1bf6910232bdf6131918c33a1bf69\n>     parent 22af6fef9b6538c9e87e147a920be9509acf1ddd\n> \n> naming the only object whose name was done with new hash with the\n> new longer hash, while recording the names of the other existing\n> objects with SHA-1.  We would need to extend the object format for\n> tag (which would be trivial as the object reference is textual and\n> similar to a commit) and tree (much harder), of course.\n\nOne thing I worry about in a mixed-hash setting is how often the two\nwill be mixed. That will lead to interoperability complications, but I\nalso think it creates security hazards (if I can convince you somehow to\nrefer to my evil colliding file by its sha1, for example, then I can\nsubvert the strength of the new hash).\n\nSo I'd much rather see strong rules like:\n\n  1. Once a repo has flag-day switched over to the new hash format[1],\n     new references are _always_ done with the new hash. Even ones that\n     point to pre-flag-day objects!\n\n     So you get a \"commit-v2\" object instead of a \"commit\", and it has a\n     distinct hash identity from its \"commit\" counterpart. You can point\n     to a classic \"commit\", but you do so by its new-hash.\n\n     The flag-day switch would probably be a repo config flag based on\n     repositoryformatversion (so old versions would just punt if they\n     see it). Let's call this flag \"newhash\" for lack of a better term.\n\n  2. Repos that have new-hash set will consider the new hash\n     format as primary, and always use it when writing and referring to\n     new objects (e.g., in refs). A (purely local) sha1->new mapping can\n     be maintained for doing old-style object lookups, or for quick\n     equivalence checks (this mapping might need to be bi-directional\n     for some use cases; I haven't thought hard enough about it to say\n     either way).\n\n  3. For protocol interop, the rules would be something like[2]:\n\n      a. If upload-pack is serving a newhash repo, it advertises\n         so in the capabilities.\n\n\t Recent clients know that the rest of the conversation will\n\t involve the new hash format. If they're cloning, they set the\n\t newhash flag in their local config.  If they're fetching, they\n\t probably abort and say \"please enable newhash\" (because for an\n\t existing repo, it probably needs to migrate refs, for example).\n\n\t An old client would fail to send back the newhash capability,\n\t and the server would abort the conversation at that point.\n\n\t A new upload-pack serving a non-newhash repo behaves the same\n\t as now (use sha1, happily interoperate with existing and new\n\t clients).\n\n      b. receive-pack is more or less the mirror image.\n\n         A server for a newhash-flagged repo has a capability for \"this\n\t is a newhash repo\" and advertises newhash refs. An existing\n\t client might still try to push, but the server would reject it\n\t unless it advertises \"newhash\" back to the server.\n\n\t A newhash-enabled client on a non-newhash repo would abort more\n\t gracefully (\"please upgrade your local repo to newhash\").\n\n\t For a newhash-enabled server with a non-newhash repo, it would\n\t probably not advertise anything (not even \"I understand\n\t newhash\"). Because the process for converting to newhash is not\n\t \"just push some newhash objects\", but an out-of-band flag-day\n\t to convert it over.\n\nThat's just a sketch I came up with. There are probably holes. And it\ndefinitely leaves a lot of _possible_ interoperability on the table in\nfavor of the flag-day approach. But I think the flag-day approach is a\nlot easier to reason about. Both in the code, and in terms of the\nsecurity properties.\n\n-Peff\n\n[1] I was intentionally vague on \"new hash format\" here. Obviously there\n    are various contenders like SHA-256. But I think there's also an\n    open question of whether the new format should be a multi-hash\n    format. That would ease further transitions. At the same time, we\n    really _don't_ want people picking bespoke hashes for their\n    repositories. It creates complications in the code, and it destroys\n    a bunch of optimizations (like knowing when we are both talking\n    about the same object based on the hash).\n\n    So I am torn between \"move to SHA-256 (or whatever)\" and \"move to a\n    hash format that encodes the hash-type in the first byte, but refuse\n    to allocate more than one hash for now\".\n\n[2] If we're having a flag-day event, this _might_ be time to consider\n    some of the breaking protocol changes that have been under\n    discussion.  I'm really hesitant to complicate this already-tricky\n    issue by throwing in the kitchen sink. But if there's going to be a\n    flag day where you need to upgrade Git to access certain repos, it\n    might be nice if there's only one. I dunno.\n"},{"id":"312583","messageId":"22704.50445.435156.883001@chiark.greenend.org.uk","threadId":"45202","inReplyTo":"xmqq60jz5wbm.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-02-24T23:43:09Z","receivedAt":"2017-02-24T23:44:07Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"Junio C Hamano writes (\"Re: SHA1 collisions found\"):\n> Ian Jackson <ijackson@chiark.greenend.org.uk> writes:\n> >  * Therefore the transition needs to be done by giving every object\n> >    two names (old and new hash function).  Objects may refer to each\n> >    other by either name, but must pick one.  The usual shape of\n> \n> I do not think it is necessrily so.\n\nIndeed.  And my latest thoughts involve instead having two parallel\nsystems of old and new objects.\n\n> *1* In the above toy example, length being 40 vs 64 is used as a\n>     sign between SHA-1 and the new hash, and careful readers may\n>     wonder if we should use sha-3,20769079d22... or something like\n>     that that more explicity identifies what hash is used, so that\n>     we can pick a hash whose length is 64 when we transition again.\n\nI have an idea for this.  I think we should prefix new hashes with a\nsingle uppercase letter, probably H.\n\nUppercase because: case-only-distinguished ref names are already\ndiscouraged because they do not work properly on case-insensitive\nfilesystems; convention is that ref names are lowercase; so an\nuppercase letter probably won't appear at the start of a ref name\ncomponent even though almost all existing software will treat it as\nlegal.  So the result is that the new object names are unlikely to\ncollide with ref names.\n\n(There is of course no need to store the H as a literal in filenames,\nso the case-insensitive filesystem problem does not apply to ref\nnames.)\n\nWe should definitely not introduce new punctuation into object names.\nThat will cause a great deal of grief for existing software which has\nto handle git object names and may thy to store them in\nrepresentations which assume that they match \\w+.\n\nThe idea of using the length is a neat trick, but it cannot support\nthe dcurrent object name abbreviation approach unworkable.\n\nIan.\n"},{"id":"312590","messageId":"22704.51861.339612.258582@chiark.greenend.org.uk","threadId":"45202","inReplyTo":"22704.50445.435156.883001@chiark.greenend.org.uk","subject":"Re: SHA1 collisions found","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-02-25T00:06:45Z","receivedAt":"2017-02-25T00:06:54Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"Ian Jackson writes (\"Re: SHA1 collisions found\"):\n> The idea of using the length is a neat trick, but it cannot support\n> the dcurrent object name abbreviation approach unworkable.\n\nSorry, it's late here and my grammar seems to have disintegrated !\n\nIan.\n"},{"id":"312591","messageId":"CA+dhYEVwLGNZh-hbcJm+kMR4W45VbwvSVY+7YKt0V9jg_b_M4g@mail.gmail.com","threadId":"45202","inReplyTo":"xmqqh93j1g9n.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"ankostis","fromEmail":"ankostis@gmail.com","sentAt":"2017-02-25T00:31:32Z","receivedAt":"2017-02-25T00:35:18Z","isPatch":false,"sender":{"key":"ankostis@gmail.com","avatar":"https://gravatar.com/avatar/1f3597ab8ad44cbb0fd782a2773e3aba94a52c03f9f6bbe5ccb7f1fe2581b12b?d=mp&s=160"},"body":"On 24 February 2017 at 21:32, Junio C Hamano <gitster@pobox.com> wrote:\n> ankostis <ankostis@gmail.com> writes:\n>\n>> Let's assume that git is retroffited to always support the \"default\"\n>> SHA-3, but support additionally more hash-funcs.\n>> If in the future SHA-3 also gets defeated, it would be highly unlikely\n>> that the same math would also break e.g. Blake.\n>> So certain high-profile repos might choose for extra security 2 or more hashes.\n>\n> I think you are conflating two unrelated things.\n\nI believe the two distinct things you refer to below are these:\n\n  a. storing objects in filesystem and accessing them\n     by name (e.g. from cmdline), and\n\n  b. cross-referencing inside the objects (trees, tags, notes),\n\ncorrect?\n\nIf not, then please ignore my answers, below.\n\n\n>  * How are these \"2 or more hashes\" actually used?  Are you going to\n>    add three \"parent \" line to a commit with just one parent, each\n>    line storing the different hashes?\n\nYes, in all places where references are involved (tags, notes).\nBased on what what the git-hackers have written so far, this might be doable.\n\nTo ensure integrity in the case of crypto-failures, all objects must\ncross-reference each other with multiple hashes.\nOf course this extra security would stop as soon as you reach \"old\"\nhistory (unless you re-write it).\n\n\n>    How will such a commit object\n>    be named---does it have three names and do you plan to have three\n>    copies of .git/refs/heads/master somehow, each of which have\n>    SHA-1, SHA-3 and Blake, and let any one hash to identify the\n>    object?\n\nYes, based on Jason Cooper's idea, above, objects would be stored\nunder all names in the filesystem using hard links (although this\nmight not work nice on Windows).\n\n\n>    I suspect you are not going to do so; instead, you would use a\n>    very long string that is a concatenation of these three hashes as\n>    if it is an output from a single hash function that produces a\n>    long result.\n>\n>    So I think the most natural way to do the \"2 or more for extra\n>    security\" is to allow us to use a very long hash.  It does not\n>    help to allow an object to be referred to with any of these 2 or\n>    more hashes at the same time.\n\nIf hard-linking all names is doable, then most restrictions above are\ngone, correct?\n\n\n>  * If employing 2 or more hashes by combining into one may enhance\n>    the security, that is wonderful.  But we want to discourage\n>    people from inventing their own combinations left and right and\n>    end up fragmenting the world.  If a project that begins with\n>    SHA-1 only naming is forked to two (or more) and each fork uses\n>    different hashes, merging them back will become harder than\n>    necessary unless you support all these hashes forks used.\n\nAgree on discouraging people's inventions.\n\nThat is why I believe that some HASH (e.g. SHA-3) must be the blessed one.\nAll git >= 3.x.x must support at least this one (for naming and\ncross-referencing between objects).\n\n\n> Having said all that, the way to figure out the hash used in the way\n> we spell the object name may not be the best place to discourage\n> people from using random hashes of their choice.  But I think we\n> want to avoid doing something that would actively encourage\n> fragmentation.\n\nI guess the \"blessed SHA-3 will discourage people using the other\nnames., untill the next crypto-crack.\n"},{"id":"312593","messageId":"CA+55aFw6BLjPK-F0RGd9LT7X5xosKOXOxuhmKX65ZHn09r1xow@mail.gmail.com","threadId":"45202","inReplyTo":"20170224233929.p2yckbc6ksyox5nu@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-25T00:39:45Z","receivedAt":"2017-02-25T00:39:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Feb 24, 2017 at 3:39 PM, Jeff King <peff@peff.net> wrote:\n>\n> One thing I worry about in a mixed-hash setting is how often the two\n> will be mixed.\n\nHonestly, I think that a primary goal for a new hash implementation\nabsolutely needs to be to minimize mixing.\n\nNot for security issues, but because of combinatorics. You want to\nhave a model that basically reads old data, but that very aggressively\napproaches \"new data only\" in order to avoid the situation where you\nhave basically the exact same tree state, just _represented_\ndifferently.\n\nFor example, what I would suggest the rules be is something like this:\n\n - introduce new tag2/commit2/tree2/blob2 object type tags that imply\nthat they were hashed using the new hash\n\n - an old type obviously can never contain a pointer to a new type (ie\nyou can't have a \"tree\" object that contains a tree2 object or a blob2\nobject.\n\n - but also make the rule that a *new* type can never contain a\npointer to an old type, with the *very* specific exception that a\ncommit2 can have a parent that is of type \"commit\".\n\nThat way everything \"converges\" towards the new format: the only way\nyou can stay on the old format is if you only have old-format objects,\nand once you have a new-format object all your objects are going to be\nnew format - except for the history.\n\nObviously, if somebody stays in old format, you might end up still\ngetting some object duplication when you continue to merge from him,\nbut that tree can never merge back without converting to new-format,\nso it will be a temporary situation.\n\nSo you will end up with duplicate objects, and that's not good (think\nof what it does to all our full-tree \"diff\" optimizations, for example\n- you no longer get the \"these sub-trees are identical\" across a\nformat change), but realistically you'll have a very limited time of\nthat kind of duplication.\n\nI'd furthermore suggest that from a UI standpoint, we'd\n\n - convert to 64-character hex numbers (32-byte hashes)\n\n - (as mentioned earlier) default to a 40-character abbreviation\n\n - make the old 40-character SHA1's just show up within the same\naddress space (so they'd also be encoded as 32-byte hashes, just with\nthe last 12 bytes zero).\n\n - you'd see in the \"object->type\" whether it's a new or old-style hash.\n\nI suspect it shouldn't be too painful to do it that way.\n\n                Linus\n"},{"id":"312595","messageId":"CA+55aFy6G1QF3Msy2DZbyhFmn974wBeXVuZK78pJ8FgkyeU85g@mail.gmail.com","threadId":"45202","inReplyTo":"CA+55aFw6BLjPK-F0RGd9LT7X5xosKOXOxuhmKX65ZHn09r1xow@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-25T00:54:51Z","receivedAt":"2017-02-25T00:54:59Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Feb 24, 2017 at 4:39 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n>  - you'd see in the \"object->type\" whether it's a new or old-style hash.\n\nActually, I take that back. I think it might be easier to keep\n\"object->type\" as-is, and it would only show the current OBJ_xyz\nfields. Then writing the SHA ends up deciding whether a OBJ_COMMIT\ngets written as \"commit\" or \"commit2\".\n\nWith the reachability rules, you'd never have any ambiguity about which to use.\n\n                Linus\n"},{"id":"312596","messageId":"nycvar.QRO.7.75.62.1702241656010.6590@qynat-yncgbc","threadId":"45202","inReplyTo":"20170224233929.p2yckbc6ksyox5nu@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2017-02-25T01:00:55Z","receivedAt":"2017-02-25T01:02:19Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 24 Feb 2017, Jeff King wrote:\n\n>\n> So I'd much rather see strong rules like:\n>\n>  1. Once a repo has flag-day switched over to the new hash format[1],\n>     new references are _always_ done with the new hash. Even ones that\n>     point to pre-flag-day objects!\n\nhow do you define when a repo has \"switched over\" to the new format in a \ndistributed environment?\n\nso you have one person working on a project that switches their version of git \nto the new one that uses the new format.\n\nBut other people they interact with still use older versions of git\n\nwhat happens when you have someone working on two different projects where one \nhas switched and the other hasn't?\n\nwhat if they are forks of each other? (LEDE and OpenWRT, or just linux-kernel \nand linux-kernel-stable)\n\n\n>     So you get a \"commit-v2\" object instead of a \"commit\", and it has a\n>     distinct hash identity from its \"commit\" counterpart. You can point\n>     to a classic \"commit\", but you do so by its new-hash.\n>\n>     The flag-day switch would probably be a repo config flag based on\n>     repositoryformatversion (so old versions would just punt if they\n>     see it). Let's call this flag \"newhash\" for lack of a better term.\n\nso how do you interact with someone who only expects the old commit instead of \nthe commit-v2?\n\nDavid Lang\n"},{"id":"312597","messageId":"CAGZ79ka+U-TMpjsAOQLmSEfZd0UCi2bzRZ-XsLxpVXTXHfdcLg@mail.gmail.com","threadId":"45202","inReplyTo":"nycvar.QRO.7.75.62.1702241656010.6590@qynat-yncgbc","subject":"Re: SHA1 collisions found","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2017-02-25T01:15:35Z","receivedAt":"2017-02-25T01:15:42Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Fri, Feb 24, 2017 at 5:00 PM, David Lang <david@lang.hm> wrote:\n> On Fri, 24 Feb 2017, Jeff King wrote:\n>\n>>\n>> So I'd much rather see strong rules like:\n>>\n>>  1. Once a repo has flag-day switched over to the new hash format[1],\n>>     new references are _always_ done with the new hash. Even ones that\n>>     point to pre-flag-day objects!\n>\n>\n> how do you define when a repo has \"switched over\" to the new format in a\n> distributed environment?\n>\n> so you have one person working on a project that switches their version of\n> git to the new one that uses the new format.\n>\n> But other people they interact with still use older versions of git\n>\n> what happens when you have someone working on two different projects where\n> one has switched and the other hasn't?\n\nyou get infected by the \"new version requirement\"\nas soon as you pull? (GPL is cancer, anyone? ;)\n\nIf you are using an old version of git that doesn't understand the new version,\nyou're screwed.\n"},{"id":"312598","messageId":"20170225011636.qjlv2luj3zefmrpz@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CA+55aFy6G1QF3Msy2DZbyhFmn974wBeXVuZK78pJ8FgkyeU85g@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-25T01:16:36Z","receivedAt":"2017-02-25T01:16:44Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 24, 2017 at 04:39:45PM -0800, Linus Torvalds wrote:\n\n> For example, what I would suggest the rules be is something like this:\n> \n>  - introduce new tag2/commit2/tree2/blob2 object type tags that imply\n> that they were hashed using the new hash\n> \n>  - an old type obviously can never contain a pointer to a new type (ie\n> you can't have a \"tree\" object that contains a tree2 object or a blob2\n> object.\n> \n>  - but also make the rule that a *new* type can never contain a\n> pointer to an old type, with the *very* specific exception that a\n> commit2 can have a parent that is of type \"commit\".\n\nYeah, this is exactly what I had in mind. That way everybody in\n\"newhash\" mode has no decisions to make. They follow the same rules and\nit's as if sha1 never existed, except when you follow links in\nhistorical objects.\n\n> [in reply...]\n> Actually, I take that back. I think it might be easier to keep\n> \"object->type\" as-is, and it would only show the current OBJ_xyz\n> fields. Then writing the SHA ends up deciding whether a OBJ_COMMIT\n> gets written as \"commit\" or \"commit2\".\n\nYeah, I think there are some data structures with limited bits for the\n\"type\" fields (e.g., the pack format). So sticking with OBJ_COMMIT might\nbe nice. For commits and tags, it would be nice to have an \"I'm v2\"\nheader at the start so there's no confusion about how they are meant to\nbe interpreted.\n\nTrees are more difficult, as they don't have any such field. But a valid\ntree does need to start with a mode, so sticking some non-numeric flag\nat the front of the object would work (it breaks backwards\ncompatibility, but that's kind of the point).\n\nI dunno. Maybe we do not need those markers at all, and could get by\npurely on object-length, or annotating the headers in some way (like\n\"parent sha256:1234abcd\").\n\nIt might just be nice if we could very easily identify objects as one\ntype or the other without having to parse them in detail.\n\n> So you will end up with duplicate objects, and that's not good (think\n> of what it does to all our full-tree \"diff\" optimizations, for example\n> - you no longer get the \"these sub-trees are identical\" across a\n> format change), but realistically you'll have a very limited time of\n> that kind of duplication.\n\nYeah, cross-flag-day diffs will be more expensive. I think that's\nsomething we have to live with. I was thinking originally that the\nsha1->newhash mapping might solve that, but it only works at the blob\nlevel. I.e., you can compare a sha1 and a newhash like:\n\n  if (!hashcmp(sha1_to_newhash(a), b))\n\nwithout having to look at the contents. But it doesn't work recursively,\nbecause the tree-pointing-to-newhash will have different content.\n\n-Peff\n"},{"id":"312603","messageId":"20170225012100.ivfdlwspsqd7bkhf@sigill.intra.peff.net","threadId":"45202","inReplyTo":"nycvar.QRO.7.75.62.1702241656010.6590@qynat-yncgbc","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-25T01:21:00Z","receivedAt":"2017-02-25T01:27:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 24, 2017 at 05:00:55PM -0800, David Lang wrote:\n\n> On Fri, 24 Feb 2017, Jeff King wrote:\n> \n> > \n> > So I'd much rather see strong rules like:\n> > \n> >  1. Once a repo has flag-day switched over to the new hash format[1],\n> >     new references are _always_ done with the new hash. Even ones that\n> >     point to pre-flag-day objects!\n> \n> how do you define when a repo has \"switched over\" to the new format in a\n> distributed environment?\n\nYou don't. It's a decision for each local repo, but the rules push\neverybody towards upgrading (because you forbid them pulling from or\npushing to people who have upgraded).\n\nSo in practice, some centralized distribution point switches, and then\nit floods out from there.\n\n> so you have one person working on a project that switches their version of\n> git to the new one that uses the new format.\n\nThat shouldn't happen when they switch. It should happen when they\ndecide to move their local clone to the new format. So let's assume they\nupgrade _and_ decide to switch.\n\n> But other people they interact with still use older versions of git\n\nThose people get forced to upgrade if they want to continue interacting.\n\n> what happens when you have someone working on two different projects where\n> one has switched and the other hasn't?\n\nSee above. You only flip the flag on for one of the projects.\n\n> what if they are forks of each other? (LEDE and OpenWRT, or just\n> linux-kernel and linux-kernel-stable)\n\nOnce one flips, the other one needs to flip to, or can't interact with\nthem. I know that's harsh, and is likely to create headaches. But in the\nlong run, I think once everything has converged the resulting system is\nless insane.\n\nFor that reason I _wouldn't_ recommend projects like the kernel flip the\nflag immediately. Ideally we write the code and the new versions\npermeate the community. Then somebody (per-project) decides that it's\ntime for the community to start switching.\n\n> >     The flag-day switch would probably be a repo config flag based on\n> >     repositoryformatversion (so old versions would just punt if they\n> >     see it). Let's call this flag \"newhash\" for lack of a better term.\n> \n> so how do you interact with someone who only expects the old commit instead\n> of the commit-v2?\n\nYou ask them to upgrade.\n\n-Peff\n"},{"id":"312605","messageId":"nycvar.QRO.7.75.62.1702241733250.6590@qynat-yncgbc","threadId":"45202","inReplyTo":"20170225012100.ivfdlwspsqd7bkhf@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2017-02-25T01:39:43Z","receivedAt":"2017-02-25T01:42:20Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 24 Feb 2017, Jeff King wrote:\n\n>> what if they are forks of each other? (LEDE and OpenWRT, or just\n>> linux-kernel and linux-kernel-stable)\n>\n> Once one flips, the other one needs to flip to, or can't interact with\n> them. I know that's harsh, and is likely to create headaches. But in the\n> long run, I think once everything has converged the resulting system is\n> less insane.\n>\n> For that reason I _wouldn't_ recommend projects like the kernel flip the\n> flag immediately. Ideally we write the code and the new versions\n> permeate the community. Then somebody (per-project) decides that it's\n> time for the community to start switching.\n\ncan you 'un-flip' the flag? or if you have someone who is a developer flip their \nrepo (because they heard that sha1 is unsafe, and they want to be safe), they \ncan't contribute to the kernel. We don't want to have them loose all their work, \nso how can they convert their local repo back to somthing that's compatible?\n\nhow would submodules work if one module flips and another (or the parent) \ndoesn't?\n\nOpenWRT/LEDE have their core repo, and they pull from many other (unrelated) \nprojects into that repo (and then have 'feeds', which is sort-of-like-submodules \nto pull in other software that's maintained completely independently)\n\nMicrosoft has made lots of money with people being forced to upgrade Word \nbecause one person got a new version and everyone else needed to upgrade to be \ncompatible. There's a LOT of pain during that process. Is that really the best \nway to go?\n\nDavid Lang\n\n"},{"id":"312606","messageId":"20170225014747.f36j2ctlszpebpsy@sigill.intra.peff.net","threadId":"45202","inReplyTo":"nycvar.QRO.7.75.62.1702241733250.6590@qynat-yncgbc","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-25T01:47:48Z","receivedAt":"2017-02-25T01:47:56Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Feb 24, 2017 at 05:39:43PM -0800, David Lang wrote:\n\n> On Fri, 24 Feb 2017, Jeff King wrote:\n> \n> > > what if they are forks of each other? (LEDE and OpenWRT, or just\n> > > linux-kernel and linux-kernel-stable)\n> > \n> > Once one flips, the other one needs to flip to, or can't interact with\n> > them. I know that's harsh, and is likely to create headaches. But in the\n> > long run, I think once everything has converged the resulting system is\n> > less insane.\n> > \n> > For that reason I _wouldn't_ recommend projects like the kernel flip the\n> > flag immediately. Ideally we write the code and the new versions\n> > permeate the community. Then somebody (per-project) decides that it's\n> > time for the community to start switching.\n> \n> can you 'un-flip' the flag? or if you have someone who is a developer flip\n> their repo (because they heard that sha1 is unsafe, and they want to be\n> safe), they can't contribute to the kernel. We don't want to have them loose\n> all their work, so how can they convert their local repo back to somthing\n> that's compatible?\n\nI don't think it would be too hard to write an un-flipper (it's\nbasically just rewriting the newhash bit of history using sha1, and\nconverting your refs back to point at the sha1s).\n\n> how would submodules work if one module flips and another (or the parent)\n> doesn't?\n\nThat's a good question. It's possible that another exception should be\ncarved out for referring to a gitlink via sha1 (we _could_ say \"no,\npoint to a newhash version of the submodule\", but I think that creates a\nlot of hardship for not much gain).\n\n> OpenWRT/LEDE have their core repo, and they pull from many other (unrelated)\n> projects into that repo (and then have 'feeds', which is\n> sort-of-like-submodules to pull in other software that's maintained\n> completely independently)\n\nI think with submodules this should probably still work.  If they are\npulling in with a subtree-ish strategy, then they'd convert the incoming\ntrees to the newhash format as part of that.\n\n> Microsoft has made lots of money with people being forced to upgrade Word\n> because one person got a new version and everyone else needed to upgrade to\n> be compatible. There's a LOT of pain during that process. Is that really the\n> best way to go?\n\nI think there's going to be a lot of pain regardless. Any attempt to\nmitigate that pain and work seamlessly across old and new versions of\ngit is going cause _ongoing_ pain as people quietly rewrite the same\ncontent back and forth with different hashes. The viral-convergence\nstrategy is painful once (when you're forced to upgrade), but after that\njust works.\n\nIf you want to work on a dual-hash strategy, be my guest. I can't\npromise I'll be able to find horrific corner cases in it, but I\ncertainly can't even try to do so until there is a concrete proposal. :)\n\n-Peff\n"},{"id":"312607","messageId":"nycvar.QRO.7.75.62.1702241752350.6590@qynat-yncgbc","threadId":"45202","inReplyTo":"20170225014747.f36j2ctlszpebpsy@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2017-02-25T01:56:27Z","receivedAt":"2017-02-25T01:59:04Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"On Fri, 24 Feb 2017, Jeff King wrote:\n\n>> OpenWRT/LEDE have their core repo, and they pull from many other (unrelated)\n>> projects into that repo (and then have 'feeds', which is\n>> sort-of-like-submodules to pull in other software that's maintained\n>> completely independently)\n>\n> I think with submodules this should probably still work.  If they are\n> pulling in with a subtree-ish strategy, then they'd convert the incoming\n> trees to the newhash format as part of that.\n\nas I understand things, they have two categories of things\n\n1. Feeds, which are completely independent, separate maintainers\n\n2. core, which gets pulled into one repo, I don't know if they use submodules in \nthe process. I know that what downstream users see is a single repo.\n\nI understand and agree with the idea of trying to converge rapidly. I'm just \nlooking at cases where this may be hard (or where there may be holdouts for \nwhatever reason)\n\nDavid Lang\n"},{"id":"312608","messageId":"CA+P7+xqyaWHBxug0BPdCgDVJBLtsvUxbgfgy1uJcoGf3q6xMQg@mail.gmail.com","threadId":"45202","inReplyTo":"20170225012100.ivfdlwspsqd7bkhf@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2017-02-25T02:26:40Z","receivedAt":"2017-02-25T02:27:08Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Fri, Feb 24, 2017 at 5:21 PM, Jeff King <peff@peff.net> wrote:\n> On Fri, Feb 24, 2017 at 05:00:55PM -0800, David Lang wrote:\n>\n>> On Fri, 24 Feb 2017, Jeff King wrote:\n>>\n>> >\n>> > So I'd much rather see strong rules like:\n>> >\n>> >  1. Once a repo has flag-day switched over to the new hash format[1],\n>> >     new references are _always_ done with the new hash. Even ones that\n>> >     point to pre-flag-day objects!\n>>\n>> how do you define when a repo has \"switched over\" to the new format in a\n>> distributed environment?\n>\n> You don't. It's a decision for each local repo, but the rules push\n> everybody towards upgrading (because you forbid them pulling from or\n> pushing to people who have upgraded).\n>\n> So in practice, some centralized distribution point switches, and then\n> it floods out from there.\n\nThis seems like the most reasonable strategy so far. I think that\ntrying to allow long term co-existence is a huge pain that discourages\nswitching, when we actually want to encourage everyone to switch\nsomeone has switched.\n\nI don't think it's sane to try and allow simultaneous use of both\nhashes, since that creates a lot of headaches and discourages\ntransition somewhat.\n\nThanks,\nJake\n"},{"id":"312609","messageId":"CA+P7+xqem_7L3Hyf+vEpkav-JJSvpcyytbendeyLcpwkusE+Zw@mail.gmail.com","threadId":"45202","inReplyTo":"nycvar.QRO.7.75.62.1702241733250.6590@qynat-yncgbc","subject":"Re: SHA1 collisions found","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2017-02-25T02:28:53Z","receivedAt":"2017-02-25T02:35:38Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Fri, Feb 24, 2017 at 5:39 PM, David Lang <david@lang.hm> wrote:\n> On Fri, 24 Feb 2017, Jeff King wrote:\n>\n>>> what if they are forks of each other? (LEDE and OpenWRT, or just\n>>> linux-kernel and linux-kernel-stable)\n>>\n>>\n>> Once one flips, the other one needs to flip to, or can't interact with\n>> them. I know that's harsh, and is likely to create headaches. But in the\n>> long run, I think once everything has converged the resulting system is\n>> less insane.\n>>\n>> For that reason I _wouldn't_ recommend projects like the kernel flip the\n>> flag immediately. Ideally we write the code and the new versions\n>> permeate the community. Then somebody (per-project) decides that it's\n>> time for the community to start switching.\n>\n>\n> can you 'un-flip' the flag? or if you have someone who is a developer flip\n> their repo (because they heard that sha1 is unsafe, and they want to be\n> safe), they can't contribute to the kernel. We don't want to have them loose\n> all their work, so how can they convert their local repo back to somthing\n> that's compatible?\n\nI'd think one of the first things we want is a way to flip *and*\nunflip by re-writing history ala git-filter-branch style. (So if you\nwanted, you could also flip all your old history).\n\nOne unrelated thought I had. When an old client sees the new stuff, it\nwill probably fail in a lot of weird ways. I wonder what we can do so\nthat if we in the future have to switch to an even newer hash, how can\nwe make it so that the old versions give a more clean error\nexperience? Ideally so that it lessens the pain of transition somewhat\nin the future if/when it has to happen again?\n\nThanks,\nJake\n"},{"id":"312611","messageId":"CAD2Ti29v9h5d9mqeQMd5YTi4fDmbAJ8LMT2oX_SagDkC-LBrMQ@mail.gmail.com","threadId":"45202","inReplyTo":"CA+P7+xqyaWHBxug0BPdCgDVJBLtsvUxbgfgy1uJcoGf3q6xMQg@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"grarpamp","fromEmail":"grarpamp@gmail.com","sentAt":"2017-02-25T05:39:13Z","receivedAt":"2017-02-25T05:39:35Z","isPatch":false,"sender":{"key":"grarpamp@gmail.com","avatar":null},"body":"Repos should address keeping / 'fixing' broken sha-1 as needed.\nThey also really need to create new native modes so users can\ninitialize and use repos with (sha-3 / sha-256 / whatever) going forward.\nBackward compatibility with sha-1 or 'fixed sha-1' will be fine. Clients\ncan 'taste' and 'test' repos for which hash mode to use, or add it to\ntheir configs. Make things flexible, modular, configurable, updateable.\nWhat little point is there in 'fixing / caveating' their use of broken sha-1,\nwithout also doing strong (sha-3 / optionals) in the first place, defaulting\nnew init's to whichever strong hash looks good, and letting natural\nmigration to that happen on its own through the default process.\nIntroducing new hash modes also gives good oppurtunity to incorporate\nother generally 'incompatabile with the old' changes to benefit the future.\nOne might argue against mixed mode, after all, export and import,\nas with any other repo migration, is generally possible.  And mixed\nmode tends to prolong the actual endeavour to move to something\nbetter in the init itself. Native and new makes you update to follow.\nA lot of question / wrong ramble here, but the point should be\nconsistant... move, natively, even if only for sake of death of old\nbroken hashes. And attacks only get worse. Thought food is all.\n"},{"id":"312612","messageId":"xmqqinnyztqe.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"CA+55aFw6BLjPK-F0RGd9LT7X5xosKOXOxuhmKX65ZHn09r1xow@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-25T06:10:01Z","receivedAt":"2017-02-25T06:20:22Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> For example, what I would suggest the rules be is something like this:\n>\n>  - introduce new tag2/commit2/tree2/blob2 object type tags that imply\n> that they were hashed using the new hash\n>\n>  - an old type obviously can never contain a pointer to a new type (ie\n> you can't have a \"tree\" object that contains a tree2 object or a blob2\n> object.\n>\n>  - but also make the rule that a *new* type can never contain a\n> pointer to an old type, with the *very* specific exception that a\n> commit2 can have a parent that is of type \"commit\".\n\nOK, I think that is what Peff was suggesting in his message, and I\ndo not have problem with such a transition plan.  Or the *very*\nspecific exception could be that a reference to \"commit\" can use old\nname (which would allow binding a submodule before transition to a\nnew project).\n\nWe probably do not need \"blob2\" object as they do not embed any\npointer to another thing.  A loose blob with old name can be made\navailable on the filesystem also under new name without much \"heavy\"\ntransition, and an in-pack blob can be pointed at with _two_ entries\nin the updated pack index file under old and new names, both for the\nbase (just deflated) representation and also ofs-delta.  A ref-delta\nbased on another blob with old name may need a bit of special\nhandling, but the deltification would not be visible at the \"struct object\"\nlayer, so probably not such a big deal.\n\nWe may also be able to get away without \"commit2\" and \"tag2\" as\ntheir pointers can be widened and parse_{commit,tag}_object() should\nbe able to deal with objects with new names transparently.  \"tree2\"\nmay be a bit tricky, though, but offhand it seems to me that nothing\nis insurmountable.\n\n> That way everything \"converges\" towards the new format: the only way\n> you can stay on the old format is if you only have old-format objects,\n> and once you have a new-format object all your objects are going to be\n> new format - except for the history.\n\nYes.\n\n> So you will end up with duplicate objects, and that's not good (think\n> of what it does to all our full-tree \"diff\" optimizations, for example\n> - you no longer get the \"these sub-trees are identical\" across a\n> format change), but realistically you'll have a very limited time of\n> that kind of duplication.\n>\n> I'd furthermore suggest that from a UI standpoint, we'd\n>\n>  - convert to 64-character hex numbers (32-byte hashes)\n>\n>  - (as mentioned earlier) default to a 40-character abbreviation\n>\n>  - make the old 40-character SHA1's just show up within the same\n> address space (so they'd also be encoded as 32-byte hashes, just with\n> the last 12 bytes zero).\n\nYes to all of the above.\n\n>  - you'd see in the \"object->type\" whether it's a new or old-style hash.\n\nI am not sure if this is needed.  We may need to abstract tree_entry walker\na little bit as a preparatory step, but I suspect that the hash (and\nmore importantly the internal format) can be kept as an internal\nknowledge to the object layer (i.e. {commit,tree,tag}.c).\n\nSo,... thanks for straightening me out.  I was thinking we would\nneed mixed mode support for smoother transition, but it now seems to\nme that the approach to stratify the history into old and new is\nworkable.\n"},{"id":"312642","messageId":"20170225185050.t6e5txrppofgelsf@genre.crustytoothpaste.net","threadId":"45202","inReplyTo":"xmqq60jz5wbm.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-02-25T18:50:50Z","receivedAt":"2017-02-25T18:51:04Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Fri, Feb 24, 2017 at 09:32:13AM -0800, Junio C Hamano wrote:\n> Ian Jackson <ijackson@chiark.greenend.org.uk> writes:\n> \n> > I have been thinking about how to do a transition from SHA1 to another\n> > hash function.\n> \n> Good.  I think many of us have also been, too, not necessarily just\n> in the past few days in response to shattered, but over the last 10\n> years, yet without coming to a consensus design ;-)\n> \n> > I have concluded that:\n> >\n> >  * We can should avoid expecting everyone to rewrite all their\n> >    history.\n> \n> Yes.\n\nThere are security implications for old objects if we mix hashes, but I\nsuppose people who want better security will just rewrite history\nanyway.\n\n> As long as the reader can tell from the format of object names\n> stored in the \"new object format\" object from what era is being\n> referred to in some way [*1*], we can name new objects with only new\n> hash, I would think.  \"new refers only to new\" that stratifies\n> objects into older and newer may make things simpler, but I am not\n> convinced yet that it would give our users a smooth enough\n> transition path (but I am open to be educated and pursuaded the\n> other way).\n\nI would simply use multihash[0] for this purpose.  New-style objects\nserialize data in multihash format, so it's immediately obvious what\nhash we're referring to.  That makes future transitions less\nproblematic.\n\n[0] https://github.com/multiformats/multihash\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"312643","messageId":"20170225190410.anvb7ll7tlhwgm3t@genre.crustytoothpaste.net","threadId":"45202","inReplyTo":"CACsJy8AtQG8YXQ+YfSFifUxqtd==THj5weJK5jooyiRN0yamiQ@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-02-25T19:04:10Z","receivedAt":"2017-02-25T19:12:37Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Fri, Feb 24, 2017 at 04:42:38PM +0700, Duy Nguyen wrote:\n> On Thu, Feb 23, 2017 at 11:43 PM, Joey Hess <id@joeyh.name> wrote:\n> > IIRC someone has been working on parameterizing git's SHA1 assumptions\n> > so a repository could eventually use a more secure hash. How far has\n> > that gotten? There are still many \"40\" constants in git.git HEAD.\n> \n> Michael asked Brian (that \"someone\") the other day and he replied [1]\n> \n> >> I'm curious; what fraction of the overall convert-to-object_id campaign\n> >> do you estimate is done so far? Are you getting close to the promised\n> >> land yet?\n> >\n> > So I think that the current scope left is best estimated by the\n> > following command:\n> >\n> >   git grep -P 'unsigned char\\s+(\\*|.*20)' | grep -v '^Documentation'\n> >\n> > So there are approximately 1200 call sites left, which is quite a bit of\n> > work.  I estimate between the work I've done and other people's\n> > refactoring work (such as the refs backend refactor), we're about 40%\n> > done.\n\nAs a note, I've been working on this pretty much nonstop since the\ncollision announcement was made.  After another 27 commits, I've got it\ndown from 1244 to 1119.\n\nI plan to send another series out sometime after the existing series has\nhit next.  People who are interested can follow the object-id-part*\nbranches at https://github.com/bk2204/git.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"312649","messageId":"20170225192655.l5dbzq42cvk5surl@sigill.intra.peff.net","threadId":"45202","inReplyTo":"20170225185050.t6e5txrppofgelsf@genre.crustytoothpaste.net","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-25T19:26:56Z","receivedAt":"2017-02-25T19:27:05Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Feb 25, 2017 at 06:50:50PM +0000, brian m. carlson wrote:\n\n> > As long as the reader can tell from the format of object names\n> > stored in the \"new object format\" object from what era is being\n> > referred to in some way [*1*], we can name new objects with only new\n> > hash, I would think.  \"new refers only to new\" that stratifies\n> > objects into older and newer may make things simpler, but I am not\n> > convinced yet that it would give our users a smooth enough\n> > transition path (but I am open to be educated and pursuaded the\n> > other way).\n> \n> I would simply use multihash[0] for this purpose.  New-style objects\n> serialize data in multihash format, so it's immediately obvious what\n> hash we're referring to.  That makes future transitions less\n> problematic.\n> \n> [0] https://github.com/multiformats/multihash\n\nI looked at that earlier, because I think it's a reasonable idea for\nfuture-proofing. The first byte is a \"varint\", but I couldn't find where\nthey defined that format.\n\nThe closest I could find is:\n\n  https://github.com/multiformats/unsigned-varint\n\nwhose README says:\n\n  This unsigned varint (VARiable INTeger) format is for the use in all\n  the multiformats.\n\n    - We have not yet decided on a format yet. When we do, this readme\n      will be updated.\n\n    - We have time. All multiformats are far from requiring this varint.\n\nwhich is not exactly confidence inspiring. They also put the length at\nthe front of the hash. That's probably convenient if you're parsing an\nunknown set of hashes, but I'm not sure it's helpful inside Git objects.\nAnd there's an incentive to minimize header data at the front of a hash,\nbecause every byte is one more byte that every single hash will collide\nover, and people will have to type when passing hashes to \"git show\",\netc.\n\nI'd almost rather use something _really_ verbose like\n\n  sha256:1234abcd...\n\nin all of the objects. And then when we get an unadorned hash from the\nuser, we guess it's sha256 (or whatever), and fallback to treating it as\na sha1.\n\nUsing a syntactically-obvious name like that also solves one other\nproblem: there are sha1 hashes whose first bytes will encode as a \"this\nis sha256\" multihash, creating some ambiguity.\n\n-Peff\n"},{"id":"312669","messageId":"20170225220944.fl7fxirtdtcko4xl@glandium.org","threadId":"45202","inReplyTo":"20170225192655.l5dbzq42cvk5surl@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2017-02-25T22:09:44Z","receivedAt":"2017-02-25T22:10:27Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Sat, Feb 25, 2017 at 02:26:56PM -0500, Jeff King wrote:\n> On Sat, Feb 25, 2017 at 06:50:50PM +0000, brian m. carlson wrote:\n> \n> > > As long as the reader can tell from the format of object names\n> > > stored in the \"new object format\" object from what era is being\n> > > referred to in some way [*1*], we can name new objects with only new\n> > > hash, I would think.  \"new refers only to new\" that stratifies\n> > > objects into older and newer may make things simpler, but I am not\n> > > convinced yet that it would give our users a smooth enough\n> > > transition path (but I am open to be educated and pursuaded the\n> > > other way).\n> > \n> > I would simply use multihash[0] for this purpose.  New-style objects\n> > serialize data in multihash format, so it's immediately obvious what\n> > hash we're referring to.  That makes future transitions less\n> > problematic.\n> > \n> > [0] https://github.com/multiformats/multihash\n> \n> I looked at that earlier, because I think it's a reasonable idea for\n> future-proofing. The first byte is a \"varint\", but I couldn't find where\n> they defined that format.\n> \n> The closest I could find is:\n> \n>   https://github.com/multiformats/unsigned-varint\n> \n> whose README says:\n> \n>   This unsigned varint (VARiable INTeger) format is for the use in all\n>   the multiformats.\n> \n>     - We have not yet decided on a format yet. When we do, this readme\n>       will be updated.\n> \n>     - We have time. All multiformats are far from requiring this varint.\n> \n> which is not exactly confidence inspiring. They also put the length at\n> the front of the hash. That's probably convenient if you're parsing an\n> unknown set of hashes, but I'm not sure it's helpful inside Git objects.\n> And there's an incentive to minimize header data at the front of a hash,\n> because every byte is one more byte that every single hash will collide\n> over, and people will have to type when passing hashes to \"git show\",\n> etc.\n> \n> I'd almost rather use something _really_ verbose like\n> \n>   sha256:1234abcd...\n> \n> in all of the objects. And then when we get an unadorned hash from the\n> user, we guess it's sha256 (or whatever), and fallback to treating it as\n> a sha1.\n> \n> Using a syntactically-obvious name like that also solves one other\n> problem: there are sha1 hashes whose first bytes will encode as a \"this\n> is sha256\" multihash, creating some ambiguity.\n\nIndeed, multihash only really is interesting when *all* hashes use it.\nAnd obviously, git can't change the existing sha1s.\n\nMike\n"},{"id":"312671","messageId":"D74A82FF-BF00-481F-9B2A-4AF8EF3D062F@gmail.com","threadId":"45202","inReplyTo":"20170224230604.nt37uw5y3uehukfd@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Lars Schneider","fromEmail":"larsxschneider@gmail.com","sentAt":"2017-02-25T22:35:27Z","receivedAt":"2017-02-25T22:35:34Z","isPatch":false,"sender":{"key":"larsxschneider@gmail.com","avatar":"https://avatars.githubusercontent.com/u/477434?v=4"},"body":"\n> On 25 Feb 2017, at 00:06, Jeff King <peff@peff.net> wrote:\n> \n> On Fri, Feb 24, 2017 at 11:47:46PM +0100, Jakub Narębski wrote:\n> \n>> I have just read on ArsTechnica[1] that while Git repository could be\n>> corrupted (though this would require attackers to spend great amount\n>> of resources creating their own collision, while as said elsewhere\n>> in this thread allegedly easy to detect), putting two proof-of-concept\n>> different PDFs with same size and SHA-1 actually *breaks* Subversion.\n>> Repository can become corrupt, and stop accepting new commits.  \n>> \n>> From what I understand people tried this, and Git doesn't exhibit\n>> such problem.  I wonder what assumptions SVN made that were broken...\n> \n> To be clear, nobody has generated a sha1 collision in Git yet, and you\n> cannot blindly use the shattered PDFs to do so. Git's notion of the\n> SHA-1 of an object include the header, so somebody would have to do a\n> shattered-level collision search for something that starts with the\n> correct \"blob 1234\\0\" header.\n> \n> So we don't actually know how Git would behave in the face of a SHA-1\n> collision. It would be pretty easy to simulate it with something like:\n> \n> ---\n> diff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\n> index 22b125cf8..1be5b5ba3 100644\n> --- a/block-sha1/sha1.c\n> +++ b/block-sha1/sha1.c\n> @@ -231,6 +231,16 @@ void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n> \t\tmemcpy(ctx->W, data, len);\n> }\n> \n> +/* sha1 of blobs containing \"foo\\n\" and \"bar\\n\" */\n> +static const unsigned char foo_sha1[] = {\n> +\t0x25, 0x7c, 0xc5, 0x64, 0x2c, 0xb1, 0xa0, 0x54, 0xf0, 0x8c,\n> +\t0xc8, 0x3f, 0x2d, 0x94, 0x3e, 0x56, 0xfd, 0x3e, 0xbe, 0x99\n> +};\n> +static const unsigned char bar_sha1[] = {\n> +\t0x57, 0x16, 0xca, 0x59, 0x87, 0xcb, 0xf9, 0x7d, 0x6b, 0xb5,\n> +\t0x49, 0x20, 0xbe, 0xa6, 0xad, 0xde, 0x24, 0x2d, 0x87, 0xe6\n> +};\n> +\n> void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n> {\n> \tstatic const unsigned char pad[64] = { 0x80 };\n> @@ -248,4 +258,8 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n> \t/* Output hash */\n> \tfor (i = 0; i < 5; i++)\n> \t\tput_be32(hashout + i * 4, ctx->H[i]);\n> +\n> +\t/* pretend \"foo\" and \"bar\" collide */\n> +\tif (!memcmp(hashout, bar_sha1, 20))\n> +\t\tmemcpy(hashout, foo_sha1, 20);\n> }\n\nThat's a good idea! I wonder if it would make sense to setup an \nadditional job in TravisCI that patches every Git version with some hash \ncollisions and then runs special tests. This way we could ensure Git \nbehaves reasonable in case of a collision. E.g. by printing errors and \nnot crashing or corrupting the repo. Do you think that would be worth \nthe effort?\n\n- Lars"},{"id":"312673","messageId":"CA+dhYEUTZMLKfBXSHU61qx5i9P6a0muUcyAvnNKzc=_E5-z7wA@mail.gmail.com","threadId":"45202","inReplyTo":"20170224172335.GG11350@io.lakedaemon.net","subject":"Re: SHA1 collisions found","fromName":"ankostis","fromEmail":"ankostis@gmail.com","sentAt":"2017-02-25T23:22:10Z","receivedAt":"2017-02-25T23:22:47Z","isPatch":false,"sender":{"key":"ankostis@gmail.com","avatar":"https://gravatar.com/avatar/1f3597ab8ad44cbb0fd782a2773e3aba94a52c03f9f6bbe5ccb7f1fe2581b12b?d=mp&s=160"},"body":"On 24 February 2017 at 18:23, Jason Cooper <git@lakedaemon.net> wrote:\n> Hi Ian,\n>\n> On Fri, Feb 24, 2017 at 03:13:37PM +0000, Ian Jackson wrote:\n>> Joey Hess writes (\"SHA1 collisions found\"):\n>> > https://shattered.io/static/shattered.pdf\n>> > https://freedom-to-tinker.com/2017/02/23/rip-sha-1/\n>> >\n>> > IIRC someone has been working on parameterizing git's SHA1 assumptions\n>> > so a repository could eventually use a more secure hash. How far has\n>> > that gotten? There are still many \"40\" constants in git.git HEAD.\n>>\n>> I have been thinking about how to do a transition from SHA1 to another\n>> hash function.\n>>\n>> I have concluded that:\n>>\n>>  * We can should avoid expecting everyone to rewrite all their\n>>    history.\n>\n> Agreed.\n>\n>>  * Unfortunately, because the data formats (particularly, the commit\n>>    header) are not in practice extensible (because of the way existing\n>>    code parses them), it is not useful to try generate new data (new\n>>    commits etc.) containing both new hashes and old hashes: old\n>>    clients will mishandle the new data.\n>\n> My thought here is:\n>\n>  a) re-hash blobs with sha256, hardlink to sha1 objects\n>  b) create new tree objects which are mirrors of each sha1 tree object,\n>     but purely sha256\n>  c) mirror commits, but they are also purely sha256\n>  d) future PGP signed tags would sign both hashes (or include both?)\n\n\nIMHO that is a great idea that needs more attention.\nYou get to keep 2 or more hash-functions for extra security in a PQ world.\n\nAnd to keep sketches for the future so far,\nSHA-3 must be always one of the new hashes.\nActually, you can get rid of SHA-1 completely, and land on Linus's\ncurrent sketches for the way ahead.\n\nThanks,\n  Kostis\n>\n> Which would end up something like:\n>\n>   .git/\n>     \\... #usual files\n>     \\objects\n>       \\ef\n>         \\3c39f7522dc55a24f64da9febcfac71e984366\n>     \\objects-sha2_256\n>       \\72\n>         \\604fd2de5f25c89d692b01081af93bcf00d2af34549d8d1bdeb68bc048932\n>     \\info\n>       \\...\n>     \\info-sha2_256\n>       \\refs #uses sha256 commit identifiers\n>\n> Basically, keep the sha256 stuff out of the way for legacy clients, and\n> new clients will still be able to use it.\n>\n> There shouldn't be a need to re-sign old signed tags if the underlying\n> objects are counter-hashed.  There might need to be some transition\n> info, though.\n>\n> Say a new client does 'git tag -v tags/v3.16' in the kernel tree.  I would\n> expect it to check the sha1 hashes, verify the PGP signed tag, and then\n> also check the sha256 counter-hashes of the relevant objects.\n>\n> thx,\n>\n> Jason.\n"},{"id":"312674","messageId":"20170226001607.GH11350@io.lakedaemon.net","threadId":"45202","inReplyTo":"CA+dhYEVwLGNZh-hbcJm+kMR4W45VbwvSVY+7YKt0V9jg_b_M4g@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jason Cooper","fromEmail":"git@lakedaemon.net","sentAt":"2017-02-26T00:16:07Z","receivedAt":"2017-02-26T00:16:18Z","isPatch":false,"sender":{"key":"git@lakedaemon.net","avatar":null},"body":"Hi,\n\nOn Sat, Feb 25, 2017 at 01:31:32AM +0100, ankostis wrote:\n> That is why I believe that some HASH (e.g. SHA-3) must be the blessed one.\n> All git >= 3.x.x must support at least this one (for naming and\n> cross-referencing between objects).\n\nI would stress caution here.  SHA3 has survived the NIST competition,\nbut that's about it.  It has *not* received nearly as much scrutiny as\nSHA2.\n\nSHA2 is a similar construction to SHA1 (Merkle–Damgård [1]) so it makes\nsense to be leery of it, but I would argue it's seasoning merits serious\nconsideration.\n\nIdeally, bless SHA2-384 (minimum) as the next hash.  Five or so years\ndown the road, if SHA3 is still in good standing, bless it as the next\nhash.\n\n\nthx,\n\nJason.\n\n[1]\nhttps://en.wikipedia.org/wiki/Merkle%E2%80%93Damg%C3%A5rd_construction\n"},{"id":"312675","messageId":"20170226004657.zowlojdzqrrcalsm@sigill.intra.peff.net","threadId":"45202","inReplyTo":"D74A82FF-BF00-481F-9B2A-4AF8EF3D062F@gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-26T00:46:57Z","receivedAt":"2017-02-26T00:48:27Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sat, Feb 25, 2017 at 11:35:27PM +0100, Lars Schneider wrote:\n\n> > So we don't actually know how Git would behave in the face of a SHA-1\n> > collision. It would be pretty easy to simulate it with something like:\n> [...]\n> \n> That's a good idea! I wonder if it would make sense to setup an \n> additional job in TravisCI that patches every Git version with some hash \n> collisions and then runs special tests. This way we could ensure Git \n> behaves reasonable in case of a collision. E.g. by printing errors and \n> not crashing or corrupting the repo. Do you think that would be worth \n> the effort?\n\nI think it would be interesting to see the results under various\nscenarios. I don't know that it would be all that interesting from an\nongoing CI perspective. But we wouldn't know until somebody actually\nwrites the tests and we see what they do.\n\n-Peff\n"},{"id":"312676","messageId":"20170226011359.GI11350@io.lakedaemon.net","threadId":"45202","inReplyTo":"xmqqinnyztqe.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"Jason Cooper","fromEmail":"git@lakedaemon.net","sentAt":"2017-02-26T01:13:59Z","receivedAt":"2017-02-26T01:14:30Z","isPatch":false,"sender":{"key":"git@lakedaemon.net","avatar":null},"body":"Hi Junio,\n\nOn Fri, Feb 24, 2017 at 10:10:01PM -0800, Junio C Hamano wrote:\n> I was thinking we would need mixed mode support for smoother\n> transition, but it now seems to me that the approach to stratify the\n> history into old and new is workable.\n\nAs someone looking to deploy (and having previously deployed) git in\nunconventional roles, I'd like to add one caveat.  The flag day in the\nhistory is great, but I'd like to be able to confirm the integrity of\nthe old history.\n\n\"Counter-hashing\" the blobs is easy enough, but the trees, commits and\ntags would need to have, iiuc, some sort of cross-reference.  As in my\nprevious example, \"git tag -v v3.16\" also checks the counter hash to\nfurther verify the integrity of the history (yes, it *really* needs to\ncheck all of the old hashes, but I'd like to make sure I can do step one\nfirst).\n\nWould there be opposition to counter-hashing the old commits at the flag\nday?\n\nthx,\n\nJason.\n"},{"id":"312679","messageId":"20170226051834.i37mlqv5wxwz3254@sigill.intra.peff.net","threadId":"45202","inReplyTo":"20170226011359.GI11350@io.lakedaemon.net","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-26T05:18:34Z","receivedAt":"2017-02-26T05:26:58Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 26, 2017 at 01:13:59AM +0000, Jason Cooper wrote:\n\n> On Fri, Feb 24, 2017 at 10:10:01PM -0800, Junio C Hamano wrote:\n> > I was thinking we would need mixed mode support for smoother\n> > transition, but it now seems to me that the approach to stratify the\n> > history into old and new is workable.\n> \n> As someone looking to deploy (and having previously deployed) git in\n> unconventional roles, I'd like to add one caveat.  The flag day in the\n> history is great, but I'd like to be able to confirm the integrity of\n> the old history.\n> \n> \"Counter-hashing\" the blobs is easy enough, but the trees, commits and\n> tags would need to have, iiuc, some sort of cross-reference.  As in my\n> previous example, \"git tag -v v3.16\" also checks the counter hash to\n> further verify the integrity of the history (yes, it *really* needs to\n> check all of the old hashes, but I'd like to make sure I can do step one\n> first).\n> \n> Would there be opposition to counter-hashing the old commits at the flag\n> day?\n\nI don't think a counter-hash needs to be embedded into the git objects\nthemselves. If the \"modern\" repo format stores everything primarily as\nsha-256, say, it will probably need to maintain a (local) mapping table\nof sha1/sha256 equivalence. That table can be generated at any time from\nthe object data (though I suspect we'll keep it up to date as objects\nenter the repository).\n\nAt the flag day[1], you can make a signed tag with the \"correct\" mapping\nin the tag body (so part of the actual GPG signed data, not referenced\nby sha1). Then later you can compare that mapping to the object content\nin the repo (or to the local copy of the mapping based on that data).\n\n-Peff\n\n[1] You don't even need to wait until the flag day. You can do it now.\n    This is conceptually similar to the git-evtag tool, though it just\n    signs the blob contents of the tag's current tree state. Signing the\n    whole mapping lets you verify the entirety of history, but of course\n    that mapping is quite big: 20 + 32 bytes per object for\n    sha1/sha-256, which is ~250MB for the kernel. So you'd probably not\n    want to do it more than once.\n"},{"id":"312688","messageId":"20170226173851.wxv7j7zibmzpo726@genre.crustytoothpaste.net","threadId":"45202","inReplyTo":"20170225220944.fl7fxirtdtcko4xl@glandium.org","subject":"Re: SHA1 collisions found","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-02-26T17:38:51Z","receivedAt":"2017-02-26T17:39:26Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sun, Feb 26, 2017 at 07:09:44AM +0900, Mike Hommey wrote:\n> On Sat, Feb 25, 2017 at 02:26:56PM -0500, Jeff King wrote:\n> > I looked at that earlier, because I think it's a reasonable idea for\n> > future-proofing. The first byte is a \"varint\", but I couldn't find where\n> > they defined that format.\n> > \n> > The closest I could find is:\n> > \n> >   https://github.com/multiformats/unsigned-varint\n> > \n> > whose README says:\n> > \n> >   This unsigned varint (VARiable INTeger) format is for the use in all\n> >   the multiformats.\n> > \n> >     - We have not yet decided on a format yet. When we do, this readme\n> >       will be updated.\n> > \n> >     - We have time. All multiformats are far from requiring this varint.\n> > \n> > which is not exactly confidence inspiring. They also put the length at\n> > the front of the hash. That's probably convenient if you're parsing an\n> > unknown set of hashes, but I'm not sure it's helpful inside Git objects.\n> > And there's an incentive to minimize header data at the front of a hash,\n> > because every byte is one more byte that every single hash will collide\n> > over, and people will have to type when passing hashes to \"git show\",\n> > etc.\n\nThe multihash spec also says that it's not necessary to implement\nvarints until we have 127 hashes, and considering that will be in the\nfar future, I'm quite happy to punt that problem down the road to\nsomeone else[0].\n\n> > I'd almost rather use something _really_ verbose like\n> > \n> >   sha256:1234abcd...\n> > \n> > in all of the objects. And then when we get an unadorned hash from the\n> > user, we guess it's sha256 (or whatever), and fallback to treating it as\n> > a sha1.\n> > \n> > Using a syntactically-obvious name like that also solves one other\n> > problem: there are sha1 hashes whose first bytes will encode as a \"this\n> > is sha256\" multihash, creating some ambiguity.\n> \n> Indeed, multihash only really is interesting when *all* hashes use it.\n> And obviously, git can't change the existing sha1s.\n\nWell, that's why I said in new objects.  If we're going to default to a\nnew hash, we can store it inside the object format, but not actually\nexpose it to the user.\n\nIn other words, if we used SHA-256, a tree object would refer to the SHA-1\nempty blob as 1114e69de29bb2d1d6434b8b29ae775ad8c2e48c5391 and the\nSHA-256 empty blob as\n1220473a0f4c3be8a93681a267e3b1e9a7dcda1185436fe141f7749120a303721813,\nbut user-visible code would parse them as e69d... and 473a... (or as\nsha1:e69d and 473a, or something).\n\nThere's very little code which actually parses objects, so it's easy\nenough to introduce a few new functions to read and write the prefixed\nversions within the objects, and leave the rest to work in the same old\nuser-visible way (or in the way that you've proposed).\n\nNote also that we need some way to distinguish objects in binary form,\nsince if we mix hashes, we need to be able to read data directly from\npack files and other locations where we serialize data that way.\nMultihash would do that, even if we didn't expose that to the user.\n\n[0] And for the record, I'm a maintenance programmer, and I dislike it\nwhen people punt the problem down the road to someone else, because\nthat's usually me.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"312689","messageId":"20170226173810.fp2tqikrm4nzu4uk@genre.crustytoothpaste.net","threadId":"45202","inReplyTo":"20170226001607.GH11350@io.lakedaemon.net","subject":"Re: SHA1 collisions found","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-02-26T17:38:10Z","receivedAt":"2017-02-26T17:39:28Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sun, Feb 26, 2017 at 12:16:07AM +0000, Jason Cooper wrote:\n> Hi,\n> \n> On Sat, Feb 25, 2017 at 01:31:32AM +0100, ankostis wrote:\n> > That is why I believe that some HASH (e.g. SHA-3) must be the blessed one.\n> > All git >= 3.x.x must support at least this one (for naming and\n> > cross-referencing between objects).\n> \n> I would stress caution here.  SHA3 has survived the NIST competition,\n> but that's about it.  It has *not* received nearly as much scrutiny as\n> SHA2.\n> \n> SHA2 is a similar construction to SHA1 (Merkle–Damgård [1]) so it makes\n> sense to be leery of it, but I would argue it's seasoning merits serious\n> consideration.\n> \n> Ideally, bless SHA2-384 (minimum) as the next hash.  Five or so years\n> down the road, if SHA3 is still in good standing, bless it as the next\n> hash.\n\nI don't think we want to be changing hashes that frequently.  Projects\nfrequently last longer than five years.  I think using a 256-bit hash is\nthe right choice because it fits on an 80-column screen in hex format.\n384-bit hashes do not.  This matters because line wrapping makes\ncopy-paste hard, and user experience is important.\n\nI've mentioned this on the list earlier, but here are the contenders in\nmy view:\n\nSHA-256:\n  Common, but cryptanalysis has advanced.  Preimage resistance (which is\n  even more important than collision resistance) has gotten to 52 of 64\n  rounds.  Pseudo-collision attacks are possible against 46 of 64\n  rounds.  Slowest option.\nSHA-3-256:\n  Less common, but has a wide security margin.  Cryptanalysis is\n  ongoing, but has not advanced much.  Somewhat to much faster than\n  SHA-256, unless you have SHA-256 hardware acceleration (which almost\n  nobody does).\nBLAKE2b-256:\n  Lower security margin, but extremely fast (faster than SHA-1 and even\n  MD5).\n\nMy recommendation has been for SHA-3-256, because I think it provides\nthe best tradeoff between security and performance.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"312691","messageId":"xmqqzih8yfpd.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"20170226004657.zowlojdzqrrcalsm@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-26T18:22:54Z","receivedAt":"2017-02-26T18:23:21Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> On Sat, Feb 25, 2017 at 11:35:27PM +0100, Lars Schneider wrote:\n> ...\n>> That's a good idea! I wonder if it would make sense to setup an \n>> additional job in TravisCI that patches every Git version with some hash \n>> collisions and then runs special tests.\n>\n> I think it would be interesting to see the results under various\n> scenarios. I don't know that it would be all that interesting from an\n> ongoing CI perspective.\n\nI had the same thought.  \n\nI view such a test as a very good validation while we are finishing\nup the introduction of new hash and the update to the codepaths that\nneed to handle both hashes, so I'd expect such a test to be a good\nvalidation measure.  But once that work is concluded, I do not know\nif tests in ongoing basis is all that interesting.\n"},{"id":"312692","messageId":"20170226183046.tbpodi5zt4kvrbpg@genre.crustytoothpaste.net","threadId":"45202","inReplyTo":"20170226051834.i37mlqv5wxwz3254@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-02-26T18:30:46Z","receivedAt":"2017-02-26T18:31:06Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sun, Feb 26, 2017 at 12:18:34AM -0500, Jeff King wrote:\n> On Sun, Feb 26, 2017 at 01:13:59AM +0000, Jason Cooper wrote:\n> \n> > On Fri, Feb 24, 2017 at 10:10:01PM -0800, Junio C Hamano wrote:\n> > > I was thinking we would need mixed mode support for smoother\n> > > transition, but it now seems to me that the approach to stratify the\n> > > history into old and new is workable.\n> > \n> > As someone looking to deploy (and having previously deployed) git in\n> > unconventional roles, I'd like to add one caveat.  The flag day in the\n> > history is great, but I'd like to be able to confirm the integrity of\n> > the old history.\n> > \n> > \"Counter-hashing\" the blobs is easy enough, but the trees, commits and\n> > tags would need to have, iiuc, some sort of cross-reference.  As in my\n> > previous example, \"git tag -v v3.16\" also checks the counter hash to\n> > further verify the integrity of the history (yes, it *really* needs to\n> > check all of the old hashes, but I'd like to make sure I can do step one\n> > first).\n> > \n> > Would there be opposition to counter-hashing the old commits at the flag\n> > day?\n> \n> I don't think a counter-hash needs to be embedded into the git objects\n> themselves. If the \"modern\" repo format stores everything primarily as\n> sha-256, say, it will probably need to maintain a (local) mapping table\n> of sha1/sha256 equivalence. That table can be generated at any time from\n> the object data (though I suspect we'll keep it up to date as objects\n> enter the repository).\n\nI really like this look-aside approach.  I think it makes it really easy\nto just rewrite the history internally, but still be able to verify\nsigned commits and signed tags.  We could even synthesize the blobs and\ntrees from the new hash versions if we didn't want to store them.\n\nThis essentially avoids the need for handling competing hashes in the\nsame object (and controversy about multihash or other storage\nfacilities); just specify the new hash in the objects, and look up the\nold one in the database if necessary.\n\nThis also will be the easiest approach to implement, IMHO.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"312694","messageId":"xmqqtw7gye7m.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"20170225011636.qjlv2luj3zefmrpz@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-26T18:55:09Z","receivedAt":"2017-02-26T18:55:46Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> Trees are more difficult, as they don't have any such field. But a valid\n> tree does need to start with a mode, so sticking some non-numeric flag\n> at the front of the object would work (it breaks backwards\n> compatibility, but that's kind of the point).\n\nJust like the object header format does not inherently impose a\nmaximum length the system can handle on our objects or the number of\nmode bits we can use in an entry in the tree object [*1*], the\nformat in which tags and commits refer to other objects does not\nimpose what hash is used for these references [*2*].  \n\nThe object names in the tree format is an oddball; by being a binary\n20-byte field and without any other hint, it does limit us to stick\nto SHA-1.\n\nI think the helper functions in tree-walk.h, namely \n\n\tinit_tree_desc();\n\ttree_entry_extract();\n\tupdate_tree_entry();\n\nand the associated data structures can be updated to read a tree\nobject in a new format without affecting the readers too much.  By\nhaving a \"I am in a new format\" byte at the beginning that cannot be\na valid first byte in the current tree format (non-octal is a good\nthing to use here), init_tree_desc() can set things up in the desc\nstructure to expect that the data that will be read by\ntree_entry_extract() and update_tree_entry() are formatted in a new\nway, and by varying that \"tree-format signature\" byte, we can update\nthe format in the future.\n\nSo at the loose-object format level, we may not even need \"tree2\";\nwe can view this update in a way similar to the change we did when\nwe started supporting submodules/gitlinks.  Older Git would have\nsaid \"There is an object that is not tree or blob recorded\" and\nbarfed but newer one takes such a tree just fine.  This \"we are now\nintroducing a new hash, and a tree can either have objects all named\nby SHA-1 or all new (non SHA-1) hash\" update can be treated the same\nway, methinks.\n\nThe normal flow to write tree objects is (supposed to be) all\ncontained in cache-tree.c.  As long as we can tell from \"struct\nobject\" which hash names the object (i.e. struct object_id may\nbecome an enum and a union), we should be able to use it to convert\nobjects near the tip of the existing history to new hashes\nincrementally. Ideally, the flag-day for one tip of a dag may be\njust a matter of\n\n\tgit commit --allow-empty -m \"object name hash update\"\n\nwithout anything else.  The commit by default would want to name\nitself with the new hash, which requires it to get its tree named\nwith the new hash, which may read the old tree and associated blobs\nall named with SHA-1, but write_index_as_tree() should be able to\n(1) read the tree with its SHA-1 name to learn what is contained;\n(2) read the contents of blobs with their SHA-1 names, and compute\ntheir names with the new hash; and (3) write out a containing tree\nobject in the updated format and named with the new hash.  And that\nwould give us the tree object named with the new hash that the\ncommand can write into the new commit object on its \"tree\" line.\n\n\n[Footnote]\n\n*1* These lengths and mode bits are spelled out in ASCII without any\n    fixed length limit for the number of the bytes in this ASCII\n    string that represents the length.  The current code may happen\n    to read them into unsigned long and unsigned int, which does\n    impose limit on the individual reader in the sense that if your\n    ulong is only 32-bit, you cannot have an object larger than 4GB.\n    But that is not an inherent limit in the format; you can lift it\n    by upgrading the reader.\n\n*2* They are also spelled out in ASCII and there is no length limit.\n    Existing implementation may happen to assume that they are all\n    SHA-1, but the readers and the writers can be updated to allow\n    other hashes to be used in a way that does not break existing\n    code when we are only using SHA-1 by marking a reference that\n    uses new hash distinguishable from SHA-1 references.\n"},{"id":"312695","messageId":"8e98a9f9-a431-9170-df9d-24ad8ec59ed7@virtuell-zuhause.de","threadId":"45202","inReplyTo":"20170224230604.nt37uw5y3uehukfd@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Thomas Braun","fromEmail":"thomas.braun@virtuell-zuhause.de","sentAt":"2017-02-26T18:57:19Z","receivedAt":"2017-02-26T18:57:28Z","isPatch":false,"sender":{"key":"thomas.braun@virtuell-zuhause.de","avatar":"https://avatars.githubusercontent.com/u/1185677?v=4"},"body":"Am 25.02.2017 um 00:06 schrieb Jeff King:\n> So we don't actually know how Git would behave in the face of a SHA-1\n> collision. It would be pretty easy to simulate it with something like:\n>\n> ---\n> diff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\n> index 22b125cf8..1be5b5ba3 100644\n> --- a/block-sha1/sha1.c\n> +++ b/block-sha1/sha1.c\n> @@ -231,6 +231,16 @@ void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n>  \t\tmemcpy(ctx->W, data, len);\n>  }\n>  \n> +/* sha1 of blobs containing \"foo\\n\" and \"bar\\n\" */\n> +static const unsigned char foo_sha1[] = {\n> +\t0x25, 0x7c, 0xc5, 0x64, 0x2c, 0xb1, 0xa0, 0x54, 0xf0, 0x8c,\n> +\t0xc8, 0x3f, 0x2d, 0x94, 0x3e, 0x56, 0xfd, 0x3e, 0xbe, 0x99\n> +};\n> +static const unsigned char bar_sha1[] = {\n> +\t0x57, 0x16, 0xca, 0x59, 0x87, 0xcb, 0xf9, 0x7d, 0x6b, 0xb5,\n> +\t0x49, 0x20, 0xbe, 0xa6, 0xad, 0xde, 0x24, 0x2d, 0x87, 0xe6\n> +};\n> +\n>  void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n>  {\n>  \tstatic const unsigned char pad[64] = { 0x80 };\n> @@ -248,4 +258,8 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n>  \t/* Output hash */\n>  \tfor (i = 0; i < 5; i++)\n>  \t\tput_be32(hashout + i * 4, ctx->H[i]);\n> +\n> +\t/* pretend \"foo\" and \"bar\" collide */\n> +\tif (!memcmp(hashout, bar_sha1, 20))\n> +\t\tmemcpy(hashout, foo_sha1, 20);\n>  }\n\nWhile reading about the subject I came across [1]. The author reduced\nthe hash size to 4bits and then played around with git.\n\nDiff taken from the posting (not my code)\n--- git-2.7.0~rc0+next.20151210.orig/block-sha1/sha1.c\n+++ git-2.7.0~rc0+next.20151210/block-sha1/sha1.c\n@@ -246,6 +246,8 @@ void blk_SHA1_Final(unsigned char hashou\n    blk_SHA1_Update(ctx, padlen, 8);\n\n    /* Output hash */\n-   for (i = 0; i < 5; i++)\n-       put_be32(hashout + i * 4, ctx->H[i]);\n+   for (i = 0; i < 1; i++)\n+       put_be32(hashout + i * 4, (ctx->H[i] & 0xf000000));\n+   for (i = 1; i < 5; i++)\n+       put_be32(hashout + i * 4, 0);\n }\n\nFrom a noob git-dev perspective this sounds more flexibel.\n\n[1]: http://stackoverflow.com/a/34599081\n"},{"id":"312696","messageId":"CA+55aFzJtejiCjV0e43+9oR3QuJK2PiFiLQemytoLpyJWe6P9w@mail.gmail.com","threadId":"45202","inReplyTo":"20170226173810.fp2tqikrm4nzu4uk@genre.crustytoothpaste.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-26T19:11:27Z","receivedAt":"2017-02-26T19:20:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Sun, Feb 26, 2017 at 9:38 AM, brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n>\n> SHA-256:\n>   Common, but cryptanalysis has advanced.  Preimage resistance (which is\n>   even more important than collision resistance) has gotten to 52 of 64\n>   rounds.  Pseudo-collision attacks are possible against 46 of 64\n>   rounds.  Slowest option.\n> SHA-3-256:\n>   Less common, but has a wide security margin.  Cryptanalysis is\n>   ongoing, but has not advanced much.  Somewhat to much faster than\n>   SHA-256, unless you have SHA-256 hardware acceleration (which almost\n>   nobody does).\n> BLAKE2b-256:\n>   Lower security margin, but extremely fast (faster than SHA-1 and even\n>   MD5).\n>\n> My recommendation has been for SHA-3-256, because I think it provides\n> the best tradeoff between security and performance.\n\nI initially was leaning towards SHA256 because of hw acceleration, but\nnoticed that the Intel SHA NI instructions that they've talking about\nso long don't seem to actually exist anywhere (maybe the Goldmont\nAtoms?)\n\nSo SHA256 acceleration is mainly an ARM thing, and nobody develops on\nARM because there's effectively no hardware that is suitable for\ndevelopers. Even ARM people just use PCs (and they won't be Goldmont\nAtoms).\n\nReduced-round SHA256 may have been broken, but on the other hand it's\nbeen around for a lot longer too, so ...\n\nBut yes, SHA3-256 looks like the sane choice. Performance of hashing\nis important in the sense that it shouldn't _suck_, but is largely\nsecondary. All my profiles on real loads (well, *my* real loads) have\nshown that zlib performance is actually much more important than SHA1.\n\nAnyway, I don't think we should make the hash choice based on pure\nperformance concerns - crypto strength first, assuming performance is\n\"not horrible\". SHA3-256 does sound like the best choice.\n\nAnd no, we should not make extensibility a primary concern. It is\nlikely that supporting two hashes will make it easier to support three\nin the future, but I do not think those kinds of worries should even\nbe on the radar.\n\nIt's *much* more important that we don't waste memory and CPU cycles\non being overly \"generic\" than some theoretical \"but but maybe in\nanother fifteen years..\"\n\n              Linus\n"},{"id":"312698","messageId":"20170226213042.rd55ykgymmr37c7n@sigill.intra.peff.net","threadId":"45202","inReplyTo":"8e98a9f9-a431-9170-df9d-24ad8ec59ed7@virtuell-zuhause.de","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-26T21:30:42Z","receivedAt":"2017-02-26T21:30:50Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 26, 2017 at 07:57:19PM +0100, Thomas Braun wrote:\n\n> While reading about the subject I came across [1]. The author reduced\n> the hash size to 4bits and then played around with git.\n> \n> Diff taken from the posting (not my code)\n> --- git-2.7.0~rc0+next.20151210.orig/block-sha1/sha1.c\n> +++ git-2.7.0~rc0+next.20151210/block-sha1/sha1.c\n> @@ -246,6 +246,8 @@ void blk_SHA1_Final(unsigned char hashou\n>     blk_SHA1_Update(ctx, padlen, 8);\n> \n>     /* Output hash */\n> -   for (i = 0; i < 5; i++)\n> -       put_be32(hashout + i * 4, ctx->H[i]);\n> +   for (i = 0; i < 1; i++)\n> +       put_be32(hashout + i * 4, (ctx->H[i] & 0xf000000));\n> +   for (i = 1; i < 5; i++)\n> +       put_be32(hashout + i * 4, 0);\n>  }\n\nYeah, that is a lot more flexible for experimenting. Though I'd think\nyou'd probably want more than 4 bits just to avoid accidental\ncollisions. Something like 24 bits gives you some breathing space (you'd\nexpect a random collision after 4096 objects), but it's still easy to\ndo a preimage attack if you need to.\n\n-Peff\n"},{"id":"312699","messageId":"CACBZZX6fP_JpL+K3XUnke=4m4gZBLu-Afyz5yJkrRnGXHuhR8A@mail.gmail.com","threadId":"45202","inReplyTo":"CA+55aFzJtejiCjV0e43+9oR3QuJK2PiFiLQemytoLpyJWe6P9w@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2017-02-26T21:38:35Z","receivedAt":"2017-02-26T21:39:14Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Sun, Feb 26, 2017 at 8:11 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n> But yes, SHA3-256 looks like the sane choice. Performance of hashing\n> is important in the sense that it shouldn't _suck_, but is largely\n> secondary. All my profiles on real loads (well, *my* real loads) have\n> shown that zlib performance is actually much more important than SHA1.\n\nWhat's the zlib v.s. hash ratio on those profiles? If git is switching\nto another hashing function given the developments in faster\ncompression algorithms (gzip v.s. snappy v.s. zstd v.s. lz4)[1] we'll\nprobably switch to another compression algorithm sooner than later.\n\nWould compression still be the bottleneck by far with zstd, how about with lz4?\n\n1. https://code.facebook.com/posts/1658392934479273/smaller-and-faster-data-compression-with-zstandard/\n"},{"id":"312700","messageId":"20170226215220.jckz6yzgben4zbyz@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CACBZZX6fP_JpL+K3XUnke=4m4gZBLu-Afyz5yJkrRnGXHuhR8A@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-26T21:52:20Z","receivedAt":"2017-02-26T21:52:28Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Feb 26, 2017 at 10:38:35PM +0100, Ævar Arnfjörð Bjarmason wrote:\n\n> On Sun, Feb 26, 2017 at 8:11 PM, Linus Torvalds\n> <torvalds@linux-foundation.org> wrote:\n> > But yes, SHA3-256 looks like the sane choice. Performance of hashing\n> > is important in the sense that it shouldn't _suck_, but is largely\n> > secondary. All my profiles on real loads (well, *my* real loads) have\n> > shown that zlib performance is actually much more important than SHA1.\n> \n> What's the zlib v.s. hash ratio on those profiles? If git is switching\n> to another hashing function given the developments in faster\n> compression algorithms (gzip v.s. snappy v.s. zstd v.s. lz4)[1] we'll\n> probably switch to another compression algorithm sooner than later.\n> \n> Would compression still be the bottleneck by far with zstd, how about with lz4?\n> \n> 1. https://code.facebook.com/posts/1658392934479273/smaller-and-faster-data-compression-with-zstandard/\n\nzstd does help in normal operations that access lots of blobs. Here are\nsome timings:\n\n  http://public-inbox.org/git/20161023080552.lma2v6zxmyaiiqz5@sigill.intra.peff.net/\n\nCompression is part of the on-the-wire packfile format, so it introduces\ncompatibility headaches. Unlike the hash, it _can_ be a local thing\nnegotiated between the two ends, and a server with zstd data could\nconvert on-the-fly to zlib. You just wouldn't want to do so on a server\nbecause it's really expensive (or you double your cache footprint to\nstore both).\n\nIf there were a hash flag day, we _could_ make sure all post-flag-day\nimplementations have zstd, and just start using that (it transparently\nhandles old zlib data, too). I'm just hesitant to through in the kitchen\nsink and make the hash transition harder than it already is.\n\nHash performance doesn't matter much for normal read operations. If your\nimplementation is really _slow_ it does matter for a few operations\n(notably index-pack receiving a large push or fetch). Some timings:\n\n  http://public-inbox.org/git/20170223230621.43anex65ndoqbgnf@sigill.intra.peff.net/\n\nIf the new algorithm is faster than SHA-1, that might be measurable in\nthose operations, too, but obviously less dramatic, as hashing is just a\npercentage of the total operation (so it can balloon the time if it's\nslow, but optimizing it can only save so much).\n\nI don't know if the per-hash setup cost of any of the new algorithms is\nhigher than SHA-1. We care as much about hashing lots of small content\nas we do about sustained throughput of a single hash.\n\n-Peff\n"},{"id":"312710","messageId":"CAMuHMdXZ2ZPsFbPUgmvx8=-xj3GBNBJwLaGAYj+R=Z2zDQJ+hQ@mail.gmail.com","threadId":"45202","inReplyTo":"20170226213042.rd55ykgymmr37c7n@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Geert Uytterhoeven","fromEmail":"geert@linux-m68k.org","sentAt":"2017-02-27T09:57:37Z","receivedAt":"2017-02-27T09:57:56Z","isPatch":false,"sender":{"key":"geert@linux-m68k.org","avatar":"https://gravatar.com/avatar/8105b34f653a7b5b98e225e565b11ebcc762ad4ab1a9d905a4663db029a9e6bc?d=mp&s=160"},"body":"On Sun, Feb 26, 2017 at 10:30 PM, Jeff King <peff@peff.net> wrote:\n> On Sun, Feb 26, 2017 at 07:57:19PM +0100, Thomas Braun wrote:\n>> While reading about the subject I came across [1]. The author reduced\n>> the hash size to 4bits and then played around with git.\n>>\n>> Diff taken from the posting (not my code)\n>> --- git-2.7.0~rc0+next.20151210.orig/block-sha1/sha1.c\n>> +++ git-2.7.0~rc0+next.20151210/block-sha1/sha1.c\n>> @@ -246,6 +246,8 @@ void blk_SHA1_Final(unsigned char hashou\n>>     blk_SHA1_Update(ctx, padlen, 8);\n>>\n>>     /* Output hash */\n>> -   for (i = 0; i < 5; i++)\n>> -       put_be32(hashout + i * 4, ctx->H[i]);\n>> +   for (i = 0; i < 1; i++)\n>> +       put_be32(hashout + i * 4, (ctx->H[i] & 0xf000000));\n>> +   for (i = 1; i < 5; i++)\n>> +       put_be32(hashout + i * 4, 0);\n>>  }\n>\n> Yeah, that is a lot more flexible for experimenting. Though I'd think\n> you'd probably want more than 4 bits just to avoid accidental\n> collisions. Something like 24 bits gives you some breathing space (you'd\n> expect a random collision after 4096 objects), but it's still easy to\n> do a preimage attack if you need to.\n\nJust shortening the hash causes lots of collisions between objects of\ndifferent types. While it's valuable to test git behavior for those cases, you\nprobably want some way to explicitly test collisions that do not change\nthe object type, as they're not trivial to detect.\n\nGr{oetje,eeting}s,\n\n                        Geert\n\n--\nGeert Uytterhoeven -- There's lots of Linux beyond ia32 -- geert@linux-m68k.org\n\nIn personal conversations with technical people, I call myself a hacker. But\nwhen I'm talking to journalists I just say \"programmer\" or something like that.\n                                -- Linus Torvalds\n"},{"id":"312715","messageId":"20170227104338.qfaaktf3or4hwfw7@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CAMuHMdXZ2ZPsFbPUgmvx8=-xj3GBNBJwLaGAYj+R=Z2zDQJ+hQ@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-27T10:43:39Z","receivedAt":"2017-02-27T10:45:08Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Feb 27, 2017 at 10:57:37AM +0100, Geert Uytterhoeven wrote:\n\n> > Yeah, that is a lot more flexible for experimenting. Though I'd think\n> > you'd probably want more than 4 bits just to avoid accidental\n> > collisions. Something like 24 bits gives you some breathing space (you'd\n> > expect a random collision after 4096 objects), but it's still easy to\n> > do a preimage attack if you need to.\n> \n> Just shortening the hash causes lots of collisions between objects of\n> different types. While it's valuable to test git behavior for those cases, you\n> probably want some way to explicitly test collisions that do not change\n> the object type, as they're not trivial to detect.\n\nRight, that's why I'm suggesting to make a longer truncation so that\nyou don't get accidental collisions, but can still find a few specific\nones for your testing.\n\n24 bits is enough to make toy repositories. If you wanted to store a\nreal repository with the truncated sha1s, you might use 36 bits (that's\n9 hex characters, which is enough for git.git to avoid any accidental\ncollisions). But you can still find a collision via brute force in 2^18\ntries, which is not so bad.\n\nI.e., something like:\n\ndiff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\nindex 22b125cf8..9158e39ed 100644\n--- a/block-sha1/sha1.c\n+++ b/block-sha1/sha1.c\n@@ -233,6 +233,10 @@ void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n \n void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n {\n+\t/* copy out only the first 36 bits */\n+\tstatic const uint32_t mask_bits[5] = {\n+\t\t0xffffffff, 0xf0000000\n+\t};\n \tstatic const unsigned char pad[64] = { 0x80 };\n \tunsigned int padlen[2];\n \tint i;\n@@ -247,5 +251,5 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n \n \t/* Output hash */\n \tfor (i = 0; i < 5; i++)\n-\t\tput_be32(hashout + i * 4, ctx->H[i]);\n+\t\tput_be32(hashout + i * 4, ctx->H[i] & mask_bits[i]);\n }\nBuild that and make it available as git.broken, and then feed your repo\ninto it, like:\n\n  git init --bare fake.git\n  git fast-export HEAD | git.broken -C fake.git fast-import\n\nat which point you have an alternate-universe version of the repository,\nwhich you can operate on as usual with your git.broken tool.\n\nAnd then you can come up with collisions via brute force:\n\n  # hack to convince hash-object to do lots of sha1s in a single\n  # invocation\n  N=300000\n  for i in $(seq $N); do\n    echo $i >$i\n  done\n  seq 300000 | git.broken hash-object --stdin-paths >hashes\n\n  for collision in $(sort hashes | uniq -d); do\n\tgrep -n $collision hashes\n  done\n\nThe result is that \"33713\\n\" and \"170653\\n\" collide. So you can now add\nthose to your fake.git repository and watch the chaos ensue.\n\n-Peff\n"},{"id":"312718","messageId":"CANv4PNm4v7-m=UFUtFutdi2JkV7biV68hAaP=AFOYnz9J37VTw@mail.gmail.com","threadId":"45202","inReplyTo":"20170227104338.qfaaktf3or4hwfw7@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Morten Welinder","fromEmail":"mwelinder@gmail.com","sentAt":"2017-02-27T12:39:15Z","receivedAt":"2017-02-27T12:41:30Z","isPatch":false,"sender":{"key":"mwelinder@gmail.com","avatar":null},"body":"Just swap in md5 in place of sha1.  Pad with '0'.  That'll give you\nall the collisions you want and none of those you don't want.\n\nOn Mon, Feb 27, 2017 at 5:43 AM, Jeff King <peff@peff.net> wrote:\n> On Mon, Feb 27, 2017 at 10:57:37AM +0100, Geert Uytterhoeven wrote:\n>\n>> > Yeah, that is a lot more flexible for experimenting. Though I'd think\n>> > you'd probably want more than 4 bits just to avoid accidental\n>> > collisions. Something like 24 bits gives you some breathing space (you'd\n>> > expect a random collision after 4096 objects), but it's still easy to\n>> > do a preimage attack if you need to.\n>>\n>> Just shortening the hash causes lots of collisions between objects of\n>> different types. While it's valuable to test git behavior for those cases, you\n>> probably want some way to explicitly test collisions that do not change\n>> the object type, as they're not trivial to detect.\n>\n> Right, that's why I'm suggesting to make a longer truncation so that\n> you don't get accidental collisions, but can still find a few specific\n> ones for your testing.\n>\n> 24 bits is enough to make toy repositories. If you wanted to store a\n> real repository with the truncated sha1s, you might use 36 bits (that's\n> 9 hex characters, which is enough for git.git to avoid any accidental\n> collisions). But you can still find a collision via brute force in 2^18\n> tries, which is not so bad.\n>\n> I.e., something like:\n>\n> diff --git a/block-sha1/sha1.c b/block-sha1/sha1.c\n> index 22b125cf8..9158e39ed 100644\n> --- a/block-sha1/sha1.c\n> +++ b/block-sha1/sha1.c\n> @@ -233,6 +233,10 @@ void blk_SHA1_Update(blk_SHA_CTX *ctx, const void *data, unsigned long len)\n>\n>  void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n>  {\n> +       /* copy out only the first 36 bits */\n> +       static const uint32_t mask_bits[5] = {\n> +               0xffffffff, 0xf0000000\n> +       };\n>         static const unsigned char pad[64] = { 0x80 };\n>         unsigned int padlen[2];\n>         int i;\n> @@ -247,5 +251,5 @@ void blk_SHA1_Final(unsigned char hashout[20], blk_SHA_CTX *ctx)\n>\n>         /* Output hash */\n>         for (i = 0; i < 5; i++)\n> -               put_be32(hashout + i * 4, ctx->H[i]);\n> +               put_be32(hashout + i * 4, ctx->H[i] & mask_bits[i]);\n>  }\n> Build that and make it available as git.broken, and then feed your repo\n> into it, like:\n>\n>   git init --bare fake.git\n>   git fast-export HEAD | git.broken -C fake.git fast-import\n>\n> at which point you have an alternate-universe version of the repository,\n> which you can operate on as usual with your git.broken tool.\n>\n> And then you can come up with collisions via brute force:\n>\n>   # hack to convince hash-object to do lots of sha1s in a single\n>   # invocation\n>   N=300000\n>   for i in $(seq $N); do\n>     echo $i >$i\n>   done\n>   seq 300000 | git.broken hash-object --stdin-paths >hashes\n>\n>   for collision in $(sort hashes | uniq -d); do\n>         grep -n $collision hashes\n>   done\n>\n> The result is that \"33713\\n\" and \"170653\\n\" collide. So you can now add\n> those to your fake.git repository and watch the chaos ensue.\n>\n> -Peff\n"},{"id":"312720","messageId":"13bb2033-fedd-d7da-f584-21a3142852d3@web.de","threadId":"45202","inReplyTo":"20170225190410.anvb7ll7tlhwgm3t@genre.crustytoothpaste.net","subject":"Re: SHA1 collisions found","fromName":"René Scharfe","fromEmail":"l.s.r@web.de","sentAt":"2017-02-27T13:29:18Z","receivedAt":"2017-02-27T13:39:18Z","isPatch":false,"sender":{"key":"l.s.r@web.de","avatar":"https://avatars.githubusercontent.com/u/26122331?v=4"},"body":"Am 25.02.2017 um 20:04 schrieb brian m. carlson:\n>>> So I think that the current scope left is best estimated by the\n>>> following command:\n>>>\n>>>   git grep -P 'unsigned char\\s+(\\*|.*20)' | grep -v '^Documentation'\n>>>\n>>> So there are approximately 1200 call sites left, which is quite a bit of\n>>> work.  I estimate between the work I've done and other people's\n>>> refactoring work (such as the refs backend refactor), we're about 40%\n>>> done.\n> \n> As a note, I've been working on this pretty much nonstop since the\n> collision announcement was made.  After another 27 commits, I've got it\n> down from 1244 to 1119.\n> \n> I plan to send another series out sometime after the existing series has\n> hit next.  People who are interested can follow the object-id-part*\n> branches at https://github.com/bk2204/git.\n\nPerhaps the following script can help a bit; it converts local and static\nvariables in specified files.  It's just a simplistic parser which can get\nat least shadowing variables, strings and comments wrong, so its results\nneed to be reviewed carefully.\n\nI failed to come up with an equivalent Coccinelle patch so far. :-/\n\nRené\n\n\n#!/bin/sh\nwhile test $# -gt 0\ndo\n\tfile=\"$1\"\n\ttmp=\"$file.new\"\n\ttest -f \"$file\" &&\n\tperl -e '\n\t\tuse strict;\n\t\tmy %indent;\n\t\tmy %old;\n\t\tmy %new;\n\t\tmy $in_struct = 0;\n\t\twhile (<>) {\n\t\t\tif (/^(\\s*)}/) {\n\t\t\t\tmy $len = length $1;\n\t\t\t\tforeach my $key (keys %indent) {\n\t\t\t\t\tif ($len < length($indent{$key})) {\n\t\t\t\t\t\tdelete $indent{$key};\n\t\t\t\t\t\tdelete $old{$key};\n\t\t\t\t\t\tdelete $new{$key};\n\t\t\t\t\t}\n\t\t\t\t}\n\t\t\t\t$in_struct = 0;\n\t\t\t}\n\t\t\tif (!$in_struct and /^(\\s*)(static )?unsigned char (\\w+)\\[20\\];$/) {\n\t\t\t\tmy $prefix = \"$1$2\";\n\t\t\t\tmy $name = $3;\n\t\t\t\t$indent{$.} = $1;\n\t\t\t\t$old{$.} = qr/(?<!->)(?<!\\.)(?<!-)\\b$name\\b/;\n\t\t\t\t$name =~ s/sha1/oid/;\n\t\t\t\tprint $prefix . \"struct object_id \" . $name . \";\\n\";\n\t\t\t\t$new{$.} = $name . \".hash\";\n\t\t\t\tnext;\n\t\t\t}\n\t\t\tif (/^(\\s*)(static )?struct (\\w+ )?\\{$/) {\n\t\t\t\t$in_struct = 1;\n\t\t\t}\n\t\t\tif (!$in_struct and ! /\\/\\*/) {\n\t\t\t\tforeach my $key (keys %indent) {\n\t\t\t\t\ts/$old{$key}/$new{$key}/g;\n\t\t\t\t}\n\t\t\t}\n\t\t\tprint;\n\t\t}\n\t' \"$file\" >\"$tmp\" &&\n\tmv \"$tmp\" \"$file\" ||\n\texit 1\n\tshift\ndone\n"},{"id":"312722","messageId":"22708.8913.864049.452252@chiark.greenend.org.uk","threadId":"45202","inReplyTo":"20170226215220.jckz6yzgben4zbyz@sigill.intra.peff.net","subject":"Transition plan for git to move to a new hash function","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-02-27T13:00:01Z","receivedAt":"2017-02-27T14:19:30Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"I said I was working on a transition plan.  Here it is.  This is\nobviously a draft for review, and I have no official status in the git\nproject.  But I have extensive experience of protocol compatibility\nengineering, and I hope this will be helpful.\n\nIan.\n\n\nSubject: Transition plan for git to move to a new hash function\n\n\nBASIC PRINCIPLES\n================\n\nWe run multiple hashes in parallel.  Each object is named by exactly\none hash.  We define that objects with identical content, but named by\ndifferent hash functions, are different objects.\n\nObjects of one hash may refer to objects named by a different hash\nfunction to their own.  Preference rules arrange that normally, new\nhash objects refer to other new hash objects.\n\nThe intention is that for most projects, the existing SHA-1 based\nhistory will be retained and a new history built on top of it.\n(Rewriting is also possible but means a per-project hard switch.)\n\nWe extend the textual object name syntax to explicitly name the hash\nused.  Every program that invokes git or speaks git protocols will\nneed to understand the extended object name syntax.\n\nPackfiles need to be extended to be able to contain objects named by\nnew hash functions.  Blob objects with identical contents but named by\ndifferent hash functions would ideally share storage.\n\nSafety catches preferent accidental incorporation into a project of\nincompatibly-new objects, or additional deprecatedly-old objects.\nThis allows for incremental deployment.\n\n\nTEXTUAL SYNTAX\n==============\n\nThe object name textual syntax is extended.  The new syntax may be\nused in all textual git objects and protocols (commits, tags, command\nlines, etc.).\n\nWe declare that the object name syntax is henceforth\n  [A-Z]+[0-9a-z]+ | [0-9a-f]+\nand that names [A-Z].* are deprecated as ref name components.\n\n    Rationale:\n\n      Full backwards compatibility (ie, without updating any software\n      that calls git) is impossible, because the hash function needs\n      to be evident in the name, so the new names must be disjoint\n      from all old SHA-1 names.\n\n      We want a short but extensible syntax.  The syntax should impose\n      minimal extra requirements on existing git users.  In most\n      contexts where existing git users use hashes, ASCII alphanumeric\n      object names will fit.  Use of punctuation such as : or even _\n      may give trouble to existing users, who are already using\n      such things as delimiters.\n\n      In existing deployments, refnames that differ only in case are\n      generally avoided (because they are troublesome on\n      case-insensitive filesystems).  And conventionally refnames are\n      lower case.  So names starting with an upper case letter will be\n      disjoint from most existing ref name components.\n\n      (Note that there is no need to write the uppercase letter to a\n      filename in the object store; the object store can use a\n      different naming scheme.)\n\n      Even though we probably want to keep using hex, it is a good\n      idea to reserve the flexibility to use a more compact encoding,\n      while not excessively widening the existing permissible\n      character set.\n\nObject names using SHA-1 are represented, in text, as at present.\n\nObject names starting with uppercase ASCII letters H or later refer to\nnew hash functions.  Programs that use `g<objectname>' should ideally\nbe changed to show `H<hash>' for hash function `H' rather than\n`gH<hash>'.)\n\n    Rationale:\n\n      Object names starting with A-F might look like hex.  G is\n      reserved because of the way that many programs write\n      `g<objectname>'.\n\n      This gives us 19 new hash function values until we have to\n      starting using two-letter hash function prefixes, or decide to\n      use A-F after all.\n\nNew hash names my be abbreviated, by truncation (just like old\nhashes).  The hash function indicator letter must be retained.\n\nInitially we define and assign one new hash function (and textual\nobject name encoding):\n\n  H<hex>    where <hex> is the BLAKE2b hash of the object\n            (in lowercase)\n\n    Note:\n\n      If the git project prefers a different new hash function to\n      BLAKE2b, that's fine.  This proposal can even cope with two new\n      hash functions in parallel.  One could even choose on a\n      per-project basis, or switch back and forth.\n\n      It would also be possible to define a multihash object name,\n      where the full object name is the concatenation of two different\n      hash function values.  One of the hashes would have to be\n      preferred for use when a truncated object name is provided by\n      the human user.\n\nWe also reserve the following syntax for private experiments:\n  E[A-Z]+[0-9a-z]+\nWe declare that public releases of git will never accept such\nobject names.\n\nEverywhere in the git object formats and git protocols, a new object\nname (with hash function indicator) is permitted where an old object\nname is permitted.\n\nA single object may refer to other objects the hash function which\nnames the object itself, or by other hash functions, in any\ncombination.  During git operations, hash function namespace\nboundaries in the object graph are traversed freely.\n\nTwo additional restrictions: a tree object may be referenced only by\nobjects named by the same hash function as the tree object itself;\nand, a tree object may reference only blobs named by the same hash\nfunction.\n\n\nIMPLEMENTATION REQUIREMENTS\n===========================\n\nThe object store will need to store objects named by new hashes,\nalongside SHA-1 objects.\n\nIn binary protocols, where a SHA-1 object name in binary form was\npreviously used, a new codepoint must be allocated in a containing\nstructure (eg a new typecode).  Usually, the new-format binary object\nwill have a new typecode and also an additional name hash indicator,\nand it may also need a length field (as new hashes may be of different\nlengths).\n\nWhenever a new hash function textual syntax is defined, corresponding\nbinary format codepoint(s) are assigned.\n\nImplementation details such as the binary format and protocol\nspecifications, object store layout, and so on, are outside the scope\nof this transition plan.\n\n\nWHILE WE ARE HERE\n=================\n\nWe should audit the git data formats for inextensible parsers.  That\nwill make future changes (for whatever purpose) much less painful.\n\nSpecific instances I'm aware of:\n\n* Permit commits and tags to contain unexpected header lines.  Ie,\n  lines matching /^\\w+\\ / before the first blank like, where the\n  keyword is not recognised.  These should be ignored.\n\n* The signed push certificate format may need reviewing to check that\n  there is space for extension.\n\nThe test suite should contain tests of these extension possibilities.\n\nA registry (a la IANA) should be created for the extensible namespaces\n(eg of header keywords).\n\nSince new objects can be received and understood only by new software,\nanyway, it will be safe to insert extension info whenever we generate\nobjects named by new hash functions.\n\n\nCHOICE OF HASH FUNCTION\n=======================\n\nWhenever objects are created, it is necessary to choose the hash\nfunction which will be used to name it.\n\nHash functions are partially ordered, from `older' to `newer'\n(implicitly, from worse to better).  The ordering is configurable.\nFor details of the defaults, see _Transition Plan_.\n\nEach ref has, optionally, a hash function hint associated with it.\nIe, a dropping in .git which names a particular hash function, with\nthe intent that the next objects made for that ref ought to use the\nspecified hash function.\n\n\nCommits\n-------\n\nA commit is made (by default) as new as the newest of\n (i) each of its parents\n (ii) if applicable, the hash function hint for the ref to which the\n     new commit is to be written\n\nImplicitly this normally means that if HEAD refers to a new commit,\nfurther new commits will be generated on top of it.\n\nThe hash function naming an origin commit is controlled by the hint\nleft in .git for the ref named by HEAD (or for HEAD itself, if HEAD is\ndetached) by git checkout --orphan or git init.\n\nAt boundaries between old and new history, new commit(s) will refer to\nold parent(s).\n\n\nTags\n----\n\nA tag is created (by default) to by named by the same hash function as\nthe object to which it refers.\n\n\nTrees\n-----\n\nTrees are only referenced by objects named by the same hash function\nas the tree.\n\nTo satisfy this rule, occasionally a tree object named by one hash\nmust be recursively rewritten into an equivalent tree named by another\nhash.\n\nWhen a tree refers to a commit (ie, a gitlink), it may refer to a\ncommit named by a different hash function.\n\nTrees generated so that they can be referred to by the index, are\nnamed by the hash function which would name the next commit to be made\non HEAD (see `Commits', above)\n\n    Rationale: we want to avoid new commits and tags relying on weak\n    hashes.  But we should avoid demanding that commits be rewritten.\n\n\nBlobs\n----\n\nBlobs are normally referred to by trees.  Trees always refer to blobs\nnamed by the tree's own hash function.\n\nWhere a blob is created in other circumstances, the caller should\nspecify the hash function.\n\n\nRef hints\n---------\n\nAs noted above, each ref may also have a hash function hint associated\nwith it.  The hint records an intent to switch hash function.\n\nThe hash hint is (by default) copied, when the ref value is copied.\nSo for exmple if `git checkout foo' makes refs/heads/foo out of\nrefs/remotes/origin/foo, it will copy the hash hint (or lack of one)\nfrom refs/remotes/origin/foo.\n\nLikewise, the hash hint is conveyed by `git fetch' (by default) and\ncan be updated with `git push' (though this is not done by default).\n\nThe ref hash hint may be set explicitly.  That is how an individual\nbranch is upgraded.  git checkout --orphan sets it to the hash which\nnames (or the hint of) the previous HEAD.\n\nWhen a commit is made and stored in a ref, the hash hint for that ref\nis removed iff hash naming the commit's is the same as the the hint.\n\n\nCONFIGURATION - OBJECT STORE BEHAVIOUR\n======================================\n\nThe object store has configuration to specify which hash functions are\nenabled.  Each hash function H has a combination of the following\nbehaviours, according to configuration:\n\n* Collision behaviour:\n\n  What to do if we encounter an object we already have (eg as part of\n  a pack, or with hash-object) but with different contents.\n\n  (a) fail: print a scary message and abort operation (on the basis\n    that either (i) the source of the colliding object probably\n    intended the preimage that they provided, in which case proceeding\n    using our own version is wrong, or (ii) the source is conducting\n    (or unwittingly facilitating) an attack).\n\n  (b) tolerate: prefer our own data; print a message, but treat\n    the reference as referring to our version of the object.\n\n  In both cases we keep a copy of the second preimage in our .git, for\n  forensic purposes.\n\n    Rationale:\n\n       This is used as part of a gradual desupport strategy.  Existing\n       history in all existing object stores is safe and cannot be\n       corrupted or modified by receiving colliding objects.\n\n       New trees which receive their initial data from a trustworthy\n       sender over a trustworthy channel will receive correct data.\n       Bad object stores or untrustworthy channels could exploit\n       collisions, but not in new regions of the history which are\n       presumably using new names.  So even with untrustworthy\n       distribution channels, the collisons can only affect\n       archaeology.\n\n       Merging previously-unrelated histories does introduce a\n       collision hazard, but the collision would have had to have been\n       introduced while the colliding hash function was still a live\n       hash function in at least one of the two projects.\n\n\n* Hash function enablement:\n\n  Each hash function is in one of the following states:\n\n  (a) enabled: this hash function is good and available for use\n\n  (b) deprecated (in favour of H2): this hash function is\n     available for use, but newly created objects will use another\n     hash function instead (specifically, when creating an object,\n     this has function is not considered as a candidate; if as a\n     result there are no candidate hash functions, we use the\n     specified replacement H2).\n\n     Existing refs referring to objects with this hash, with no ref\n     hint, are treated as having a ref hint specifying H2.  If no H2\n     is specified, the newest hash is used.\n\n  (c) disabled: existing objects using this hash function can be\n     accessed, but no such objects can be created or received.\n     (again, a replacement may be specified).  This is used both\n     initially to prevent unintended upgrade, and later to block the\n     introduction of vulnerable data generated by badly configured\n     clients.\n\n* Preference ordering:\n\n  As mentioned in `CHOICE OF HASH FUNCTION', there is a configured\n  order on hash functions.  This order should be consistent with the\n  enablement configuration.\n\nDetails of precise configuration option names are beyond the scope of\nthis document.\n\n\nRemote protocol\n---------------\n\nDuring protocol negotation, a receiver needs to specify what hashes it\nunderstands, and whether it is prepared to see only a partial view.\n\nWhen the sender is listing its refs, refs naming objects the receiver\ncannot understand are either elided (if the receiver is content with a\npartial view), or cause an error.\n\n\nEquality testing\n----------------\n\nNote that semantically identical trees (and blobs) may (now) have\ndifferent tree objects because those tree objects might use (and be\nnamed by) different hashes.  So (in some contexts at least) tree\ncomparison cannot any longer be done by comparing object names; rather\nan invocation of git diff is needed, or explicit generation of a tree\nobject named by the right hash.\n\n\nTRANSITION PLAN\n===============\n\nFor brevity I will write `SHA' for hashing with SHA-1, using current\nunqualified object names, and `BLAKE' for hasing with BLAKE2b, using\nH<hex> object names.\n\nY<n> means `Year <n> after the start of the transition'.\nPlease adjust timescales to taste.\n\nI will focus on the default configuration as shipped by git upstream,\nand the recommended configuration for hosting providers.\n\nIndividual projects, and perhaps individual hosting providers, can\nmake their own choices, if they are willing to set appropriate\nconfiguration (on clients, and servers).\n\n\nY0: Implement all of the above.  Test it.\n\n    Default configuration:\n       SHA is enabled\n       BLAKE is disabled in trees without working trees\n       BLAKE is enabled in trees with working trees\n       SHA > BLAKE\n\n    Effects:\n\n    Clients are prepared to process BLAKE data, but it is not\n    generated by default and cannot be pushed to servers.\n\n    All old git clients still work.\n\n    Early adopters can start using the new hashes, at the cost of\n    compatibility.  Projects that want to rewrite history right away\n    can do so.  (In both cases, by setting configuration options.)\n\nY4: BLAKE by default for new projects.\n    Conversion enabled for existing projects.\n    Old git software is going to start rotting.\n\n    Default configuration change:\n       BLAKE > SHA\n       BLAKE enabled (even in trees without working trees)\n\n    Suggested multi-project hosting site configuration change:\n       Newly created projects should get BLAKE enabled\n       Existing projects should retain BLAKE disabled by default\n       Button should be provided to start conversion (see below)\n\n    Effects:\n\n    When creating a new working tree, it starts using BLAKE.\n\n    Servers which have been updated will accept BLAKE.\n\n    Servers which have not been updated to Y4's git will need a small\n    configuration change (enabling BLAKE) to cope with the new\n    projects that are using BLAKE.\n\n    To convert a project, an administrator (or project owner) would\n    set BLAKE to enabled, and SHA to deprecated, on the server(s).  On\n    the next pull the server will provide ref hints naming BLAKE,\n    which will get copied to the client's HEAD.  So the client is\n    infected with BLAKE.\n\n    To convert a project branch-by-branch, the administrator would set\n    BLAKE to enabled but leave SHA enabled.  Then each branch retains\n    its own hash.  A branch can be converted by pushing a BLAKE commit\n    to it, or by setting a ref hint on the server (or on the next\n    client to push).\n\nY6: BLAKE by default for all projects\n    Existing projects start being converted infectiously.\n    It is hard for a project to stop this happening if any of\n     their servers are updated.\n    Old git software is firmly stuffed.\n\n    Default configuration change\n       SHA deprecated in trees without working trees\n\n    Effects:\n\n    Existing projects are, by default, `converted', as described\n    above.\n\nY8: Clients hate SHA\n    Clients insist on trying to convert existing projects\n    It is very hard to stop this happening.\n    Unrepentant servers start being very hard to use.\n\n    Default configuration change\n       SHA deprecated (even in trees without working trees)\n\n    Effects:\n\n    Clients will generate only BLAKE.  Hopefully their server will\n    accept this!\n\nY10: Stop accepting new SHA\n    No-one can manage to make new SHA commits\n\n    Default configuration change\n       SHA disabled in new trees, except during initial\n          `clone', `mirror' and similar\n\n    Effects:\n\n    Existing SHA history is retained, and copied to new clients and\n    servers.  But established clients and servers reject any newly\n    introduced SHA.\n\n\n-- \nIan Jackson <ijackson@chiark.greenend.org.uk>   These opinions are my own.\n\nIf I emailed you from an address @fyvzl.net or @evade.org.uk, that is\na private address which bypasses my fierce spamfilter.\n"},{"id":"312725","messageId":"20170227143747.GB297@x4","threadId":"45202","inReplyTo":"22708.8913.864049.452252@chiark.greenend.org.uk","subject":"Re: Why BLAKE2?","fromName":"Markus Trippelsdorf","fromEmail":"markus@trippelsdorf.de","sentAt":"2017-02-27T14:37:47Z","receivedAt":"2017-02-27T15:04:43Z","isPatch":false,"sender":{"key":"markus@trippelsdorf.de","avatar":null},"body":"On 2017.02.27 at 13:00 +0000, Ian Jackson wrote:\n> \n> For brevity I will write `SHA' for hashing with SHA-1, using current\n> unqualified object names, and `BLAKE' for hasing with BLAKE2b, using\n> H<hex> object names.\n\nWhy do you choose BLAKE2? SHA-2 is generally considered still fine and\nwould be the obvious choice. And if you want to be adventurous then\nSHA-3 (Keccak) would be the next logical candidate.\n\n-- \nMarkus\n"},{"id":"312727","messageId":"22708.18670.790503.11385@chiark.greenend.org.uk","threadId":"45202","inReplyTo":"20170227143747.GB297@x4","subject":"Re: Why BLAKE2?","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-02-27T15:42:38Z","receivedAt":"2017-02-27T16:31:55Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"Markus Trippelsdorf writes (\"Re: Why BLAKE2?\"):\n> On 2017.02.27 at 13:00 +0000, Ian Jackson wrote:\n> > For brevity I will write `SHA' for hashing with SHA-1, using current\n> > unqualified object names, and `BLAKE' for hasing with BLAKE2b, using\n> > H<hex> object names.\n> \n> Why do you choose BLAKE2? SHA-2 is generally considered still fine and\n> would be the obvious choice. And if you want to be adventurous then\n> SHA-3 (Keccak) would be the next logical candidate.\n\nI don't have a strong opinion.  Keccak would be fine too.\nWe should probably avoid SHA-2.\n\nThe main point of my posting was not to argue in favour of a\nparticular hash function :-).\n\nIan.\n\n-- \nIan Jackson <ijackson@chiark.greenend.org.uk>   These opinions are my own.\n\nIf I emailed you from an address @fyvzl.net or @evade.org.uk, that is\na private address which bypasses my fierce spamfilter.\n"},{"id":"312770","messageId":"alpine.DEB.2.11.1702271846040.13590@grey.csi.cam.ac.uk","threadId":"45202","inReplyTo":"22708.8913.864049.452252@chiark.greenend.org.uk","subject":"Re: Transition plan for git to move to a new hash function","fromName":"Tony Finch","fromEmail":"dot@dotat.at","sentAt":"2017-02-27T19:26:54Z","receivedAt":"2017-02-27T20:57:23Z","isPatch":false,"sender":{"key":"dot@dotat.at","avatar":"https://avatars.githubusercontent.com/u/68429?v=4"},"body":"Ian Jackson <ijackson@chiark.greenend.org.uk> wrote:\n\nA few questions and one or two suggestions...\n\n> TEXTUAL SYNTAX\n> ==============\n>\n> We also reserve the following syntax for private experiments:\n>   E[A-Z]+[0-9a-z]+\n> We declare that public releases of git will never accept such\n> object names.\n\nInstead of this I would suggest that experimental hash names should have\nmulti-character prefixes and an easy registration process - rationale:\nhttps://tools.ietf.org/html/rfc6648\n\n> A single object may refer to other objects the hash function which\n> names the object itself, or by other hash functions, in any\n> combination.\n\nIf I understand it correctly, this freedom is greatly restricted later on\nin this document, depending on the object type in question. If so, it's\nprobably worth saying so at this point.\n\n> Commits\n> -------\n>\n> The hash function naming an origin commit is controlled by the hint\n> left in .git for the ref named by HEAD (or for HEAD itself, if HEAD is\n> detached) by git checkout --orphan or git init.\n\nThis confused me for a while - I think you mean \"root commit\"?\n\n> TRANSITION PLAN\n> ===============\n>\n> Y4: BLAKE by default for new projects.\n>\n>     When creating a new working tree, it starts using BLAKE.\n>\n>     Servers which have been updated will accept BLAKE.\n\nWhy not allow newhash pushes before making it the default for new\nprojects? Wouldn't it make sense to get the server side ready some time\nbefore projects start actively using new hashes?\n\nOr is the idea that newhash upgrade is driven from the server?\n\nWhat's the upgrade process for send-email patch exchange?\n\nTony.\n-- \nf.anthony.n.finch  <dot@dotat.at>  http://dotat.at/  -  I xn--zr8h punycode\nFair Isle: Southwest 6 to gale 8, backing east 5 or 6, backing north 6 to gale\n8 later. Rough or very rough. Rain or showers. Moderate or good.\n"},{"id":"312842","messageId":"20170228132545.ahhc6v7zulvlkzeh@genre.crustytoothpaste.net","threadId":"45202","inReplyTo":"13bb2033-fedd-d7da-f584-21a3142852d3@web.de","subject":"Re: SHA1 collisions found","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-02-28T13:25:46Z","receivedAt":"2017-02-28T13:26:43Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Mon, Feb 27, 2017 at 02:29:18PM +0100, René Scharfe wrote:\n> Am 25.02.2017 um 20:04 schrieb brian m. carlson:\n> >>> So I think that the current scope left is best estimated by the\n> >>> following command:\n> >>>\n> >>>   git grep -P 'unsigned char\\s+(\\*|.*20)' | grep -v '^Documentation'\n> >>>\n> >>> So there are approximately 1200 call sites left, which is quite a bit of\n> >>> work.  I estimate between the work I've done and other people's\n> >>> refactoring work (such as the refs backend refactor), we're about 40%\n> >>> done.\n> > \n> > As a note, I've been working on this pretty much nonstop since the\n> > collision announcement was made.  After another 27 commits, I've got it\n> > down from 1244 to 1119.\n> > \n> > I plan to send another series out sometime after the existing series has\n> > hit next.  People who are interested can follow the object-id-part*\n> > branches at https://github.com/bk2204/git.\n> \n> Perhaps the following script can help a bit; it converts local and static\n> variables in specified files.  It's just a simplistic parser which can get\n> at least shadowing variables, strings and comments wrong, so its results\n> need to be reviewed carefully.\n> \n> I failed to come up with an equivalent Coccinelle patch so far. :-/\n> \n> René\n> \n> \n> #!/bin/sh\n> while test $# -gt 0\n> do\n> \tfile=\"$1\"\n> \ttmp=\"$file.new\"\n> \ttest -f \"$file\" &&\n> \tperl -e '\n> \t\tuse strict;\n> \t\tmy %indent;\n> \t\tmy %old;\n> \t\tmy %new;\n> \t\tmy $in_struct = 0;\n> \t\twhile (<>) {\n> \t\t\tif (/^(\\s*)}/) {\n> \t\t\t\tmy $len = length $1;\n> \t\t\t\tforeach my $key (keys %indent) {\n> \t\t\t\t\tif ($len < length($indent{$key})) {\n> \t\t\t\t\t\tdelete $indent{$key};\n> \t\t\t\t\t\tdelete $old{$key};\n> \t\t\t\t\t\tdelete $new{$key};\n> \t\t\t\t\t}\n> \t\t\t\t}\n> \t\t\t\t$in_struct = 0;\n> \t\t\t}\n> \t\t\tif (!$in_struct and /^(\\s*)(static )?unsigned char (\\w+)\\[20\\];$/) {\n> \t\t\t\tmy $prefix = \"$1$2\";\n> \t\t\t\tmy $name = $3;\n> \t\t\t\t$indent{$.} = $1;\n> \t\t\t\t$old{$.} = qr/(?<!->)(?<!\\.)(?<!-)\\b$name\\b/;\n> \t\t\t\t$name =~ s/sha1/oid/;\n> \t\t\t\tprint $prefix . \"struct object_id \" . $name . \";\\n\";\n> \t\t\t\t$new{$.} = $name . \".hash\";\n> \t\t\t\tnext;\n> \t\t\t}\n> \t\t\tif (/^(\\s*)(static )?struct (\\w+ )?\\{$/) {\n> \t\t\t\t$in_struct = 1;\n> \t\t\t}\n> \t\t\tif (!$in_struct and ! /\\/\\*/) {\n> \t\t\t\tforeach my $key (keys %indent) {\n> \t\t\t\t\ts/$old{$key}/$new{$key}/g;\n> \t\t\t\t}\n> \t\t\t}\n> \t\t\tprint;\n> \t\t}\n> \t' \"$file\" >\"$tmp\" &&\n> \tmv \"$tmp\" \"$file\" ||\n> \texit 1\n> \tshift\n> done\n\nI'll see how it works.  I'm currently in New Orleans visiting a friend\nuntil Thursday, so I'll have less time than normal to look at these, but\nI'll definitely give it a try.\n\nMost of the issue is not the actual conversion, but finding the right\norder in which to convert functions.  For example, the object-id-part8\nbranch on my GitHub account converts parse_object, but\nparse_tree_indirect has to be converted before you can do parse_object.\nThat leads to another handful of patches that have to be done.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"312870","messageId":"xmqqefyikvin.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"20170223230507.kuxjqtg3ghcfskc6@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-28T18:41:52Z","receivedAt":"2017-02-28T19:00:11Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> The first one is 98K. Mail headers may bump it over vger's 100K barrier.\n> It's actually the _least_ interesting patch of the 3, because it just\n> imports the code wholesale from the other project. But if it doesn't\n> make it, you can fetch the whole series from:\n>\n>   https://github.com/peff/git jk/sha1dc\n>\n> (By the way, I don't see your version on the list, Linus, which probably\n> means it was eaten by the 100K filter).\n>\n>   [1/3]: add collision-detecting sha1 implementation\n>   [2/3]: sha1dc: adjust header includes for git\n>   [3/3]: Makefile: add USE_SHA1DC knob\n\nI was lazy so I fetched the above and then added this on top before\nI start to play with it.\n\n-- >8 --\nFrom: Junio C Hamano <gitster@pobox.com>\nDate: Tue, 28 Feb 2017 10:39:25 -0800\nSubject: [PATCH] sha1dc: resurrect LICENSE file\n\nThe upstream releases the contents under the MIT license; the\ninitial import accidentally omitted its license file.  \n\nAdd it back.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n sha1dc/LICENSE | 30 ++++++++++++++++++++++++++++++\n 1 file changed, 30 insertions(+)\n create mode 100644 sha1dc/LICENSE\n\ndiff --git a/sha1dc/LICENSE b/sha1dc/LICENSE\nnew file mode 100644\nindex 0000000000..4a3e6a1b15\n--- /dev/null\n+++ b/sha1dc/LICENSE\n@@ -0,0 +1,30 @@\n+MIT License\n+\n+Copyright (c) 2017:\n+    Marc Stevens\n+    Cryptology Group\n+    Centrum Wiskunde & Informatica\n+    P.O. Box 94079, 1090 GB Amsterdam, Netherlands\n+    marc@marc-stevens.nl\n+\n+    Dan Shumow\n+    Microsoft Research\n+    danshu@microsoft.com\n+\n+Permission is hereby granted, free of charge, to any person obtaining a copy\n+of this software and associated documentation files (the \"Software\"), to deal\n+in the Software without restriction, including without limitation the rights\n+to use, copy, modify, merge, publish, distribute, sublicense, and/or sell\n+copies of the Software, and to permit persons to whom the Software is\n+furnished to do so, subject to the following conditions:\n+\n+The above copyright notice and this permission notice shall be included in all\n+copies or substantial portions of the Software.\n+\n+THE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR\n+IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,\n+FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE\n+AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER\n+LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,\n+OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE\n+SOFTWARE.\n-- \n2.12.0-310-g733d1cbbe2\n\n"},{"id":"312871","messageId":"20170228192044.cn56puazsa3wtlkd@sigill.intra.peff.net","threadId":"45202","inReplyTo":"xmqq60jukubq.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-02-28T19:20:44Z","receivedAt":"2017-02-28T19:28:35Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 28, 2017 at 11:07:37AM -0800, Junio C Hamano wrote:\n\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n> >>   [1/3]: add collision-detecting sha1 implementation\n> >>   [2/3]: sha1dc: adjust header includes for git\n> >>   [3/3]: Makefile: add USE_SHA1DC knob\n> >\n> > I was lazy so I fetched the above and then added this on top before\n> > I start to play with it.\n> >\n> > -- >8 --\n> > From: Junio C Hamano <gitster@pobox.com>\n> > Date: Tue, 28 Feb 2017 10:39:25 -0800\n> > Subject: [PATCH] sha1dc: resurrect LICENSE file\n> \n> In a way similar to 8415558f55 (\"sha1dc: avoid c99\n> declaration-after-statement\", 2017-02-24), we would want this on\n> top.\n> \n> -- >8 --\n> Subject: sha1dc: avoid 'for' loop initial decl\n\nYeah, thanks, I had tweaked both that and the license thing locally but\nhad not pushed it out yet. Both are obvious improvements.\n\nFWIW, I've been in touch with Dan Shumow, one of the authors, who has\nbeen looking at whether we can speed up the sha1dc implementation, or\nintegrate the checks into the block-sha1 implementation.\n\nHere are a few notes on the earlier timings I posted that came out in\nour conversation:\n\n  - the timings I showed earlier were actually openssl versus sha1dc.\n    The block-sha1 timings fall somewhere in the middle:\n\n      [running test-sha1 on a fresh linux.git packfile]\n      1.347s openssl\n      2.079s block-sha1\n      4.983s sha1dc\n\n      [index-pack --verify on a fresh git.git packfile]\n       6.919s openssl\n       9.003s block-sha1\n      17.955s sha1dc\n\n    Those are the operations that show off sha1 performance the most.\n    The first one is really not even that interesting; it's raw\n    sha1 performance. The second one is an actual operation users might\n    notice (though not as --verify exactly, but as \"index-pack --stdin\"\n    when receiving a fetch or a push).\n\n    So there's room for improvement, but the gap between block-sha1\n    and sha1dc is not quite as big as I showed earlier.\n\n  - Dan timed the sha1dc implementation with and without the collision\n    detection enabled. The sha1 implementation is only 1.33x slower than\n    block-sha1 (for raw sha1 time). Adding in the detection makes it\n    2.6x slower.\n\n    So there's some potential gain from optimizing the sha1\n    implementation, but ultimately we may be looking at a 2x slowdown to\n    add in the collision detection.\n\n    It doesn't need to happen for _every_ sha1 we compute, but the\n    index-pack case is the one that almost certainly _does_ want it,\n    because that's when we're admitting remote objects into the\n    repository (ditto you'd probably want it for write_sha1_file(),\n    since you could be applying a patch from an untrusted source).\n\n-Peff\n"},{"id":"312872","messageId":"xmqq60jukubq.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"xmqqefyikvin.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-02-28T19:07:37Z","receivedAt":"2017-02-28T19:36:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n>>   [1/3]: add collision-detecting sha1 implementation\n>>   [2/3]: sha1dc: adjust header includes for git\n>>   [3/3]: Makefile: add USE_SHA1DC knob\n>\n> I was lazy so I fetched the above and then added this on top before\n> I start to play with it.\n>\n> -- >8 --\n> From: Junio C Hamano <gitster@pobox.com>\n> Date: Tue, 28 Feb 2017 10:39:25 -0800\n> Subject: [PATCH] sha1dc: resurrect LICENSE file\n\nIn a way similar to 8415558f55 (\"sha1dc: avoid c99\ndeclaration-after-statement\", 2017-02-24), we would want this on\ntop.\n\n-- >8 --\nSubject: sha1dc: avoid 'for' loop initial decl\n\nWe write this:\n\n\ttype i;\n\tfor (i = initial; i < limit; i++)\n\ninstead of this:\n\n\tfor (type i = initial; i < limit; i++)\n\nthe latter of which is from c99.\n\nSigned-off-by: Junio C Hamano <gitster@pobox.com>\n---\n sha1dc/sha1.c | 6 ++++--\n 1 file changed, 4 insertions(+), 2 deletions(-)\n\ndiff --git a/sha1dc/sha1.c b/sha1dc/sha1.c\nindex f4e261ae7a..6569b403e9 100644\n--- a/sha1dc/sha1.c\n+++ b/sha1dc/sha1.c\n@@ -41,7 +41,8 @@\n \n void sha1_message_expansion(uint32_t W[80])\n {\n-\tfor (unsigned i = 16; i < 80; ++i)\n+\tunsigned i;\n+\tfor (i = 16; i < 80; ++i)\n \t\tW[i] = rotate_left(W[i - 3] ^ W[i - 8] ^ W[i - 14] ^ W[i - 16], 1);\n }\n \n@@ -49,9 +50,10 @@ void sha1_compression(uint32_t ihv[5], const uint32_t m[16])\n {\n \tuint32_t W[80];\n \tuint32_t a, b, c, d, e;\n+\tunsigned i;\n \n \tmemcpy(W, m, 16 * 4);\n-\tfor (unsigned i = 16; i < 80; ++i)\n+\tfor (i = 16; i < 80; ++i)\n \t\tW[i] = rotate_left(W[i - 3] ^ W[i - 8] ^ W[i - 14] ^ W[i - 16], 1);\n \n \ta = ihv[0];\n"},{"id":"312874","messageId":"CA+55aFxTWqsTTiDKo4DBZT-8Z9t80bGMD3uijzKONa_bYEZABQ@mail.gmail.com","threadId":"45202","inReplyTo":"xmqq60jukubq.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-28T19:34:14Z","receivedAt":"2017-02-28T19:44:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Tue, Feb 28, 2017 at 11:07 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> In a way similar to 8415558f55 (\"sha1dc: avoid c99\n> declaration-after-statement\", 2017-02-24), we would want this on\n> top.\n\nThere's a few other simplifications that could be done:\n\n (1) make the symbols static that aren't used.\n\n     The sha1.h header ends up declaring several things that shouldn't\nhave been exported.\n\n     I suspect the code may have had some debug mode that got stripped\nout from it before making it public (or that was never there, and was\njust something the generating code could add).\n\n (2) get rid of the \"safe mode\" support.\n\n     That one is meant for non-checking replacements where it\ngenerates a *different* hash for input with the collision fingerpring,\nbut that's pointless for the git use when we abort on a collision\nfingerprint.\n\nI think the first one will show that the sha1_compression() function\nisn't actually used, and with the removal of safe-mode I think\nsha1_compression_W() also is unused.\n\nFinally, only states 58 and 65 (out of all 80 states) are actually\nused, and from what I can tell, the 'maski' value is always 0, so the\nlooping over 80 state masks is really just a loop over two.\n\nThe file has code top *generate* all the 80 sha1_recompression_step()\nfunctions, and I don't think the compiler is smart enough to notice\nthat only two of them matter.\n\nAnd because 'maski' is always zero, thisL\n\n   ubc_dv_mask[sha1_dvs[i].maski]\n\ncode looks like it might as well just use ubc_dv_mask[0] - in fact the\nubc_dv_mask[] \"array\" really is just a single-entry array anyway:\n\n   #define DVMASKSIZE 1\n\nso that code has a few oddities in it. It's generated code, which is\nprobably why.\n\nBasically, some of it could be improved. In particular, the \"generate\ncode for 80 different recompression cases, but only ever use two of\nthem\" really looks like it would blow up the code generation footprint\na lot.\n\nI'm adding Marc Stevens and Dan Shumow to this email (bcc'd, so that\nthey don't get dragged into any unrelated email threads) in case they\nwant to comment.\n\nI'm wondering if they perhaps have a cleaned-up version somewhere, or\nmaybe they can tell me that I'm just full of sh*t and missed\nsomething.\n\n                    Linus\n"},{"id":"312875","messageId":"CAJo=hJuB9JkTZSRbhN2DX0gBqpjddU=Sk8iRV9++TYRv4xKA6Q@mail.gmail.com","threadId":"45202","inReplyTo":"CA+55aFxTWqsTTiDKo4DBZT-8Z9t80bGMD3uijzKONa_bYEZABQ@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2017-02-28T19:52:06Z","receivedAt":"2017-02-28T19:54:31Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Tue, Feb 28, 2017 at 11:34 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n> On Tue, Feb 28, 2017 at 11:07 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>>\n>> In a way similar to 8415558f55 (\"sha1dc: avoid c99\n>> declaration-after-statement\", 2017-02-24), we would want this on\n>> top.\n>\n> There's a few other simplifications that could be done:\n\nYes, I found and did a number of these when I ported sha1dc to Java\nfor JGit[1], and it helped recover some of the lost throughput.\n\n[1] https://git.eclipse.org/r/#/c/91852/\n\n>  (1) make the symbols static that aren't used.\n>\n>      The sha1.h header ends up declaring several things that shouldn't\n> have been exported.\n>\n>      I suspect the code may have had some debug mode that got stripped\n> out from it before making it public (or that was never there, and was\n> just something the generating code could add).\n>\n>  (2) get rid of the \"safe mode\" support.\n>\n>      That one is meant for non-checking replacements where it\n> generates a *different* hash for input with the collision fingerpring,\n> but that's pointless for the git use when we abort on a collision\n> fingerprint.\n>\n> I think the first one will show that the sha1_compression() function\n> isn't actually used, and with the removal of safe-mode I think\n> sha1_compression_W() also is unused.\n\nCorrect.\n\n> Finally, only states 58 and 65 (out of all 80 states) are actually\n> used,\n\nYes, at present only states 58 and 65 are used. I cut out support for\nother states.\n\n> and from what I can tell, the 'maski' value is always 0, so the\n> looping over 80 state masks is really just a loop over two.\n\nActually, look closer at that loop:\n\n  for (i = 0; sha1_dvs[i].dvType != 0; ++i)\n  {\n    if ((0 == ctx->ubc_check) || (((uint32_t)(1) << sha1_dvs[i].maskb)\n& ubc_dv_mask[sha1_dvs[i].maski]))\n\nIts a loop over all 32 bits looking for which bits are set. Most of\nthe time few bits if any are set for most message blocks. Changing\nthis code to find the lowest 1 bit set in ubc_dv_mask[0] provided a\nsignificant improvement in throughput.\n\nThe sha1_dvs array is indexed by maskb, so the code can be reduced to:\n\n  while (ubcDvMask != 0) {\n    int b = numberOfTrailingZeros(lowestOneBit(ubcDvMask));\n    UbcCheck.DvInfo dv = UbcCheck.DV[b];\n\nOr something.\n"},{"id":"312909","messageId":"20170228214724.w7w5f6n4u6ehanzd@genre.crustytoothpaste.net","threadId":"45202","inReplyTo":"22708.8913.864049.452252@chiark.greenend.org.uk","subject":"Re: Transition plan for git to move to a new hash function","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-02-28T21:47:25Z","receivedAt":"2017-02-28T21:49:22Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Mon, Feb 27, 2017 at 01:00:01PM +0000, Ian Jackson wrote:\n> I said I was working on a transition plan.  Here it is.  This is\n> obviously a draft for review, and I have no official status in the git\n> project.  But I have extensive experience of protocol compatibility\n> engineering, and I hope this will be helpful.\n> \n> Ian.\n> \n> \n> Subject: Transition plan for git to move to a new hash function\n> \n> \n> BASIC PRINCIPLES\n> ================\n> \n> We run multiple hashes in parallel.  Each object is named by exactly\n> one hash.  We define that objects with identical content, but named by\n> different hash functions, are different objects.\n\nI think this is fine.\n\n> Objects of one hash may refer to objects named by a different hash\n> function to their own.  Preference rules arrange that normally, new\n> hash objects refer to other new hash objects.\n\nThe existing codebase isn't really intended with that in mind.\n\nIt's not that I am arguing against this because I think it's a bad idea,\nI'm arguing against it because as a contributor, I'm doubtful that this\nis easily achievable given the state of the codebase.\n\n> The intention is that for most projects, the existing SHA-1 based\n> history will be retained and a new history built on top of it.\n> (Rewriting is also possible but means a per-project hard switch.)\n\nI like Peff's suggested approach in which we essentially rewrite history\nunder the hood, but have a lookup table which looks up the old hash\nbased on the new hash.  That allows us to refer to old objects, but not\nhave to share serialized data that mentions both hashes.\n\nObviously only the SHA-1 versions of old tags and commits will be able\nto be validated, but that shouldn't be an issue.  We can hook that code\ninto a conversion routine that can handle on-the-fly object conversion.\n\nWe also can implement (optionally disabled) fallback functionality to\nlook up old SHA-1 hash names based on the new hash.\n\n> We extend the textual object name syntax to explicitly name the hash\n> used.  Every program that invokes git or speaks git protocols will\n> need to understand the extended object name syntax.\n> \n> Packfiles need to be extended to be able to contain objects named by\n> new hash functions.  Blob objects with identical contents but named by\n> different hash functions would ideally share storage.\n> \n> Safety catches preferent accidental incorporation into a project of\n> incompatibly-new objects, or additional deprecatedly-old objects.\n> This allows for incremental deployment.\n\nWe have a compatibility mechanism already in place: if the\nrepositoryFormatVersion option is set to 1, but an unknown extension\nflag is set, Git will bail out.\n\nFor network protocols, we have the server offer a hash=foo extension,\nand make the client echo it back, and either bail or convert on the fly.\nThis makes it fast for new clients, and slow for old clients, which\nencourages migration.\n\nWe could also store old-style packs for easy fetch by clients.\n\n> TEXTUAL SYNTAX\n> ==============\n> \n> The object name textual syntax is extended.  The new syntax may be\n> used in all textual git objects and protocols (commits, tags, command\n> lines, etc.).\n> \n> We declare that the object name syntax is henceforth\n>   [A-Z]+[0-9a-z]+ | [0-9a-f]+\n> and that names [A-Z].* are deprecated as ref name components.\n\nI'd simply say that we have data always be in the new format if it's\navailable, and tag the old SHA-1 versions instead.  Otherwise, as Peff\npointed out, we're going to be stuck typing a bunch of identical stuff\nevery time.  Again, this encourages migration.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"312910","messageId":"CY1PR0301MB21078DDCA8C679983D22821FC4560@CY1PR0301MB2107.namprd03.prod.outlook.com","threadId":"45202","inReplyTo":"CA+55aFxTWqsTTiDKo4DBZT-8Z9t80bGMD3uijzKONa_bYEZABQ@mail.gmail.com","subject":"RE: SHA1 collisions found","fromName":"Dan Shumow","fromEmail":"danshu@microsoft.com","sentAt":"2017-02-28T21:22:49Z","receivedAt":"2017-02-28T21:59:08Z","isPatch":false,"sender":{"key":"danshu@microsoft.com","avatar":null},"body":"[Responses inline]\n\nNo need to keep me \"bcc'd\" (though thanks for the consideration) -- I'm happy to ignore anything I don't want to be pulled into ;-)\n\nHere's a rollup of what needs to be done based on the discussion below:\n\n1) Remove extraneous exports from sha1.h\n2) Remove \"safe mode\" support.\n3) Remove sha1_compression_W if it is not needed by the performance improvements.\n4) Evaluate logic around storing states and generating recompression states.  Remove defines that bloat code footprint.\n\nThanks,\nDan\n\n\n-----Original Message-----\nFrom: linus971@gmail.com [mailto:linus971@gmail.com] On Behalf Of Linus Torvalds\nSent: Tuesday, February 28, 2017 11:34 AM\nTo: Junio C Hamano <gitster@pobox.com>\nCc: Jeff King <peff@peff.net>; Joey Hess <id@joeyh.name>; Git Mailing List <git@vger.kernel.org>\nSubject: Re: SHA1 collisions found\n\nOn Tue, Feb 28, 2017 at 11:07 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> In a way similar to 8415558f55 (\"sha1dc: avoid c99 \n> declaration-after-statement\", 2017-02-24), we would want this on top.\n\nThere's a few other simplifications that could be done:\n\n (1) make the symbols static that aren't used.\n\n     The sha1.h header ends up declaring several things that shouldn't have been exported.\n\n     I suspect the code may have had some debug mode that got stripped out from it before making it public (or that was never there, and was just something the generating code could add).\n\n[danshu] Yes, this is reasonable.  The emphasis of the code, heretofore, had been the illustration of our unavoidable bit condition performance improvement to counter cryptanalysis.  I'm happy to remove the unused stuff from the public header.\n\n (2) get rid of the \"safe mode\" support.\n\n     That one is meant for non-checking replacements where it generates a *different* hash for input with the collision fingerpring, but that's pointless for the git use when we abort on a collision fingerprint.\n\n[danshu] Yes, I agree that if you aren't using this it can be taken out.  I believe Marc has some use cases / potentially consumers of this algorithm in mind.  We can move it into separate header/source files for anyone who wants to use it.\n\nI think the first one will show that the sha1_compression() function isn't actually used, and with the removal of safe-mode I think\nsha1_compression_W() also is unused.\n\n[danshu]  Some of the performance experiments that I've looked at involve putting the sha1_compression_W(...) back in.  Though, that doesn't look like it's helping.  If it is unused after the performance improvements, we'll take it out, or move it into its own file.\n\nFinally, only states 58 and 65 (out of all 80 states) are actually used, and from what I can tell, the 'maski' value is always 0, so the looping over 80 state masks is really just a loop over two.\n\n[danshu]  So, while looking at performance optimizations, I specifically looked at how much removing storing the intermediate states helps -- And I found that it doesn't seem to make a difference for performance.  My cursory hypothesis is because nothing is waiting on those writes to memory, the code moves on quickly.  That said, it is a bunch of code that is essentially doing nothing and removing that is worthwhile.  Though, partially what we're seeing here is that, as you point out below, we're working with generated code that we want to be general.  Specifically, right now, we're checking only disturbance vectors that we know can be used to efficiently attack the compression function.  It may be the case that further cryptanalysis uncovers more.  We want to have a general enough approach that we can add scanning for new disturbance vectors if they're found later.  Over specializing the code makes that more difficult, as currently the algorithm is data driven, and we don't need to write new code, but rather just add more data to check.  One other note -- the \"maski\" field of the  dv_info_t struct is not an index to check the state, but rather an index into the mask generated by the ubc check code, so that doesn't pertain to looping over the states.  More on this below.  \n\nThe file has code top *generate* all the 80 sha1_recompression_step() functions, and I don't think the compiler is smart enough to notice that only two of them matter.\n\n[danshu] That's a good observation -- We should clean up the unused recompression steps, especially because that will generate a ton of object code.  We should add some logic to only compile the functions that are used.\n\nAnd because 'maski' is always zero, thisL\n\n   ubc_dv_mask[sha1_dvs[i].maski]\n\ncode looks like it might as well just use ubc_dv_mask[0] - in fact the ubc_dv_mask[] \"array\" really is just a single-entry array anyway:\n\n   #define DVMASKSIZE 1\n\n[danshu]  The idea here is that we are currently checking 32 disturbance vectors with our bit mask.  We're checking 32 DVs, because we have 32 bits of mask that we can use.  The DVs are ordered by their probability of leading to an attack (which is directly correlated to the complexity of finding a collision.)  Several of those DVs correspond to very low probability / high cost attacks, which we wouldn't expect to see in practice.  We just have the space to check, so why not?  However, improvements in cryptanalysis may make those attacks cheaper, in which case, we would potentially want to add more DVs to check, in which case we would expand the number of DVs and the mask.\n\nso that code has a few oddities in it. It's generated code, which is probably why.\n\n[danshu]  Accurate, we're also just trying to be general enough that we can easily add more DVs later if need be.  I don't know how likely that is, certainly the DVs that we're checking now are based on solid conjectures and rigorous analysis of the problem.  Though we don't want to rule out that there will be subsequent cryptanalytic developments later.  Marc can comment more here.\n\nBasically, some of it could be improved. In particular, the \"generate code for 80 different recompression cases, but only ever use two of them\" really looks like it would blow up the code generation footprint a lot.\n\nI'm adding Marc Stevens and Dan Shumow to this email (bcc'd, so that they don't get dragged into any unrelated email threads) in case they want to comment.\n\nI'm wondering if they perhaps have a cleaned-up version somewhere, or maybe they can tell me that I'm just full of sh*t and missed something.\n\n[danshu]  Naw man, it looks pretty good, modulo a little bit of understandable confusion over 'maski' -- No fake news or alternative facts here ;-)\n\n                    Linus\n"},{"id":"312923","messageId":"CA+55aFxG_5KU54KXZdTMC0p0EF5ixmv0C6ccjnYcPUeN_kDREA@mail.gmail.com","threadId":"45202","inReplyTo":"CAJo=hJuB9JkTZSRbhN2DX0gBqpjddU=Sk8iRV9++TYRv4xKA6Q@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-28T22:56:39Z","receivedAt":"2017-02-28T22:57:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Tue, Feb 28, 2017 at 11:52 AM, Shawn Pearce <spearce@spearce.org> wrote:\n>\n>> and from what I can tell, the 'maski' value is always 0, so the\n>> looping over 80 state masks is really just a loop over two.\n>\n> Actually, look closer at that loop:\n\nNo, sorry, I wasn't clear and took some shortcuts in writing that made\nthat sentence not parse right.\n\nThere's two issues going on. This loop:\n\n>   for (i = 0; sha1_dvs[i].dvType != 0; ++i)\n\nloops over all the dvs - and then inside it has that useless \"maski\"\nthing as part of the test that is always zero.\n\nBut the \"80 state masks\" was not that \"maski' value, but the\n\"ctx->states[5][80]\" thing.\n\nSo we have 80 of those 5-word state values, but only two of them are\nactually used: iterations 58 and 65). You can see how the code\nactually optimizes away (by hand) the SHA1_STORE_STATE() thing by\nusing DOSTORESTATE58 and DOSTORESTATE65, but it does actually generate\nthe code for all of them.\n\nYou can see the \"only two steps\" part in this:\n\n  (sha1_recompression_step[sha1_dvs[i].testt])(...)\n\nif you notice how there are only those two different cases for \"testt\".\n\nSo there is code generated for 80 different recompression step\nfunctions in that array, but there are only two different functions\nthat are actually ever used.\n\nThose are not small functions, either. When I check the build, they\ngenerate about 3.5kB of code each. So it's literally about 250kB of\ncompletely wasted space in the binary.\n\nSee what I'm saying? Two different issues. One is that the code\ngenerates tons of (fairly big) functions, and only uses two of them,\nthe rest are useless and lying around. The other is that it uses a\nvariable that is only ever zero.\n\nSo I think that loop would actually be better not as a loop at all,\nbut as a \"generated code expanded from the dv_data\". It would have\nbeen more obvious. Right now it loads values from the array, and it's\nnot obvious that some of the values it loads are very very limited (to\nthe point of one of them just being the constant \"0\").\n\nAnyway, Dan Shumow already answered and addressed both issues (and the\nsmaller stylistic ones).\n\n               Linus\n"},{"id":"312925","messageId":"259cc328-13f7-2d2d-c86f-8b4fc7da8e34@cwi.nl","threadId":"45202","inReplyTo":"CY1PR0301MB21078DDCA8C679983D22821FC4560@CY1PR0301MB2107.namprd03.prod.outlook.com","subject":"Re: SHA1 collisions found","fromName":"Marc Stevens","fromEmail":"marc.stevens@cwi.nl","sentAt":"2017-02-28T22:50:53Z","receivedAt":"2017-02-28T23:14:56Z","isPatch":false,"sender":{"key":"marc.stevens@cwi.nl","avatar":null},"body":"You can also keep me in this thread, so we can help or answer any\nfurther questions,\nbut I also appreciate the feedback on our project.\n\nLike Dan Shumow said, our main focus on the library has been correctness\nand then performance.\nThe entire files ubc_check.c and ubc_check.h are generated based on\ncryptanalytic data,\nin particular a list of 32 disturbance vectors and their unavoidable\nattack conditions,\nsee https://github.com/cr-marcstevens/sha1collisiondetection-tools.\n\nsha1.c and sha1.h were coded to work for any such generated ubc_check.c\nand ubc_check.h.\nThat means that indeed we might have some superfluous code, once used\nfor testing, or for generality,\nbut nothing that should noticeably impact runtime performance.\n\nBecause we only have 32 disturbance vectors to check, we have DVMASKSIZE\nequal to 1 and maski always 0.\nIn the more general case when we add disturbance vectors this will not\nremain the case.\n\nOf course for dedicated code this can be simplified, and some parts\ncould be further optimized.\n\nRegarding the recompression functions, the ones needed are given in the\nsha1_dvs table,\nbut also via preprocessor defines that are used to actually only store\nneeded internal states:\n#define DOSTORESTATE58\n#define DOSTORESTATE65\nFor each disturbance vector there is a window of which states you can\nstart the recompression from,\nwe've optimized it so there are only 2 unique starting points (58,65)\ninstead of 32.\nThese defines should be easy to use to remove superfluous compiled\nrecompression functions.\n\nNote that as each disturbance vector has its own unique message differences\n(leading to different values for ctx->m2), our code does not loop over\njust 2 items.\nIt loops over 32 distinct computations which have either of the 2\nstarting points.\n\nFinally, thanks for taking a close look at our code,\nthis helps bringing the library in better shape also for other software\nprojects.\n\nBest regards,\nMarc Stevens\n\nOn 2/28/2017 10:22 PM, Dan Shumow wrote:\n> [Responses inline]\n>\n> No need to keep me \"bcc'd\" (though thanks for the consideration) -- I'm happy to ignore anything I don't want to be pulled into ;-)\n>\n> Here's a rollup of what needs to be done based on the discussion below:\n>\n> 1) Remove extraneous exports from sha1.h\n> 2) Remove \"safe mode\" support.\n> 3) Remove sha1_compression_W if it is not needed by the performance improvements.\n> 4) Evaluate logic around storing states and generating recompression states.  Remove defines that bloat code footprint.\n>\n> Thanks,\n> Dan\n>\n>\n> -----Original Message-----\n> From: linus971@gmail.com [mailto:linus971@gmail.com] On Behalf Of Linus Torvalds\n> Sent: Tuesday, February 28, 2017 11:34 AM\n> To: Junio C Hamano <gitster@pobox.com>\n> Cc: Jeff King <peff@peff.net>; Joey Hess <id@joeyh.name>; Git Mailing List <git@vger.kernel.org>\n> Subject: Re: SHA1 collisions found\n>\n> On Tue, Feb 28, 2017 at 11:07 AM, Junio C Hamano <gitster@pobox.com> wrote:\n>> In a way similar to 8415558f55 (\"sha1dc: avoid c99 \n>> declaration-after-statement\", 2017-02-24), we would want this on top.\n> There's a few other simplifications that could be done:\n>\n>  (1) make the symbols static that aren't used.\n>\n>      The sha1.h header ends up declaring several things that shouldn't have been exported.\n>\n>      I suspect the code may have had some debug mode that got stripped out from it before making it public (or that was never there, and was just something the generating code could add).\n>\n> [danshu] Yes, this is reasonable.  The emphasis of the code, heretofore, had been the illustration of our unavoidable bit condition performance improvement to counter cryptanalysis.  I'm happy to remove the unused stuff from the public header.\n>\n>  (2) get rid of the \"safe mode\" support.\n>\n>      That one is meant for non-checking replacements where it generates a *different* hash for input with the collision fingerpring, but that's pointless for the git use when we abort on a collision fingerprint.\n>\n> [danshu] Yes, I agree that if you aren't using this it can be taken out.  I believe Marc has some use cases / potentially consumers of this algorithm in mind.  We can move it into separate header/source files for anyone who wants to use it.\n>\n> I think the first one will show that the sha1_compression() function isn't actually used, and with the removal of safe-mode I think\n> sha1_compression_W() also is unused.\n>\n> [danshu]  Some of the performance experiments that I've looked at involve putting the sha1_compression_W(...) back in.  Though, that doesn't look like it's helping.  If it is unused after the performance improvements, we'll take it out, or move it into its own file.\n>\n> Finally, only states 58 and 65 (out of all 80 states) are actually used, and from what I can tell, the 'maski' value is always 0, so the looping over 80 state masks is really just a loop over two.\n>\n> [danshu]  So, while looking at performance optimizations, I specifically looked at how much removing storing the intermediate states helps -- And I found that it doesn't seem to make a difference for performance.  My cursory hypothesis is because nothing is waiting on those writes to memory, the code moves on quickly.  That said, it is a bunch of code that is essentially doing nothing and removing that is worthwhile.  Though, partially what we're seeing here is that, as you point out below, we're working with generated code that we want to be general.  Specifically, right now, we're checking only disturbance vectors that we know can be used to efficiently attack the compression function.  It may be the case that further cryptanalysis uncovers more.  We want to have a general enough approach that we can add scanning for new disturbance vectors if they're found later.  Over specializing the code makes that more difficult, as currently the algorithm is data driven, and we don't need to write new code, but rather just add more data to check.  One other note -- the \"maski\" field of the  dv_info_t struct is not an index to check the state, but rather an index into the mask generated by the ubc check code, so that doesn't pertain to looping over the states.  More on this below.  \n>\n> The file has code top *generate* all the 80 sha1_recompression_step() functions, and I don't think the compiler is smart enough to notice that only two of them matter.\n>\n> [danshu] That's a good observation -- We should clean up the unused recompression steps, especially because that will generate a ton of object code.  We should add some logic to only compile the functions that are used.\n>\n> And because 'maski' is always zero, thisL\n>\n>    ubc_dv_mask[sha1_dvs[i].maski]\n>\n> code looks like it might as well just use ubc_dv_mask[0] - in fact the ubc_dv_mask[] \"array\" really is just a single-entry array anyway:\n>\n>    #define DVMASKSIZE 1\n>\n> [danshu]  The idea here is that we are currently checking 32 disturbance vectors with our bit mask.  We're checking 32 DVs, because we have 32 bits of mask that we can use.  The DVs are ordered by their probability of leading to an attack (which is directly correlated to the complexity of finding a collision.)  Several of those DVs correspond to very low probability / high cost attacks, which we wouldn't expect to see in practice.  We just have the space to check, so why not?  However, improvements in cryptanalysis may make those attacks cheaper, in which case, we would potentially want to add more DVs to check, in which case we would expand the number of DVs and the mask.\n>\n> so that code has a few oddities in it. It's generated code, which is probably why.\n>\n> [danshu]  Accurate, we're also just trying to be general enough that we can easily add more DVs later if need be.  I don't know how likely that is, certainly the DVs that we're checking now are based on solid conjectures and rigorous analysis of the problem.  Though we don't want to rule out that there will be subsequent cryptanalytic developments later.  Marc can comment more here.\n>\n> Basically, some of it could be improved. In particular, the \"generate code for 80 different recompression cases, but only ever use two of them\" really looks like it would blow up the code generation footprint a lot.\n>\n> I'm adding Marc Stevens and Dan Shumow to this email (bcc'd, so that they don't get dragged into any unrelated email threads) in case they want to comment.\n>\n> I'm wondering if they perhaps have a cleaned-up version somewhere, or maybe they can tell me that I'm just full of sh*t and missed something.\n>\n> [danshu]  Naw man, it looks pretty good, modulo a little bit of understandable confusion over 'maski' -- No fake news or alternative facts here ;-)\n>\n>                     Linus\n\n\n"},{"id":"312929","messageId":"CA+55aFwjzbhYyFm_MqL=cDZZeKbSjqd-jSeb0yW_bJ_WQTzEpA@mail.gmail.com","threadId":"45202","inReplyTo":"259cc328-13f7-2d2d-c86f-8b4fc7da8e34@cwi.nl","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-02-28T23:11:32Z","receivedAt":"2017-02-28T23:43:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Tue, Feb 28, 2017 at 2:50 PM, Marc Stevens <marc.stevens@cwi.nl> wrote:\n>\n> Because we only have 32 disturbance vectors to check, we have DVMASKSIZE\n> equal to 1 and maski always 0.\n> In the more general case when we add disturbance vectors this will not\n> remain the case.\n\nOk, I didn't get why that happened, but it makes sense to me now.\n\n> Of course for dedicated code this can be simplified, and some parts\n> could be further optimized.\n\nSo I'd be worried about changing your tested code too much, since the\nonly test-cases we have are the two pdf files. If we screw up too\nmuch, those will no longer show as collisions, but we could get tons\nof false positives that we wouldn't see, so..\n\nI'm wondering that since the disturbance vector cases themselves are a\nfairly limited number, perhaps the code that generates this could be\nchanged to actually just generate the static calls rather than the\nloop over the sha1_dvs[] array.\n\nIOW, instead of generating this:\n\n        for (i = 0; sha1_dvs[i].dvType != 0; ++i) {\n                .. use sha1_dvs[i] values as indexes into function arrays etc..\n        }\n\nmaybe you could just make the generator generate that loop statically,\nand have 32 function calls with the masks as constant arguments.\n\n.. together with only generating the SHA1_RECOMPRESS() functions for\nthe cases that are actually used.\n\nSo it would still be entirely generated, but it would generate a\nlittle bit more explicit code.\n\nOf course, we could just edit out all the SHA_RECOMPRESS(x) cases by\nhand, and only leave the two that are actually used.\n\nAs it is, the lib/sha1.c code generates about 250kB of code that is\nnever used if I read the code correctly (that's just the sha1.c code -\nentirely ignoring all the avx2 etc versions that I haven't looked at,\nand that I don't think git would use)\n\n> Regarding the recompression functions, the ones needed are given in the\n> sha1_dvs table,\n> but also via preprocessor defines that are used to actually only store\n> needed internal states:\n> #define DOSTORESTATE58\n> #define DOSTORESTATE65\n\nYeah, I guess we could use those #define's to cull the code \"automatically\".\n\nBut I think you could do it at the generation phase easily too, so\nthat we don't then introduce unnecessary differences when we try to\nget rid of the extra fat ;)\n\n\n> Note that as each disturbance vector has its own unique message differences\n> (leading to different values for ctx->m2), our code does not loop over\n> just 2 items.\n> It loops over 32 distinct computations which have either of the 2\n> starting points.\n\nYes, I already had to clarify my badly expressed writing on the git\nlist. My concerns were about the (very much not obvious) limited\nvalues in the dvs array.\n\nSo the code superficially *looks* like it uses all those functions you\ngenerate (and the maski value _looked_ like it was interesting), but\nwhen looking closer it turns out that there's just a two different\nfunction calls that it loops over (but it loops over them multiple\ntimes, I agree).\n\n                           Linus\n"},{"id":"312935","messageId":"CY1PR0301MB210787306B549D00238FD37FC4290@CY1PR0301MB2107.namprd03.prod.outlook.com","threadId":"45202","inReplyTo":"20170228192044.cn56puazsa3wtlkd@sigill.intra.peff.net","subject":"RE: SHA1 collisions found","fromName":"Dan Shumow","fromEmail":"danshu@microsoft.com","sentAt":"2017-03-01T08:57:21Z","receivedAt":"2017-03-01T08:57:30Z","isPatch":false,"sender":{"key":"danshu@microsoft.com","avatar":null},"body":">   - Dan timed the sha1dc implementation with and without the collision\n>     detection enabled. The sha1 implementation is only 1.33x slower than\n>    block-sha1 (for raw sha1 time). Adding in the detection makes it\n>    2.6x slower.\n\n >    So there's some potential gain from optimizing the sha1\n >    implementation, but ultimately we may be looking at a 2x slowdown to\n >     add in the collision detection.\n\nI rearranged our code a little bit and interleaved the message expansion and rounds.  This bring our raw SHA-1 implementation (without collision detection) down to 1.11x slower than the block-sha1 implementation in Git.  Adding the collision detection brings us to 2.12x slower than the block-sha1 implementation.  This was basically attacking the low hanging fruit in optimizing our implementation.  There are some things that I haven't looked into yet, but I'm basically at the point of starting to compare the generated assembler to see what's different between our implementations.\n\nOpenSSL's SHA1 implementation is implemented in assembler, so there's no way we're going to get close to that with just C level coding.\n\nThanks,\nDan\n"},{"id":"312984","messageId":"20170301190524.ocmsi37hv4d22mrc@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CA+55aFwjzbhYyFm_MqL=cDZZeKbSjqd-jSeb0yW_bJ_WQTzEpA@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-01T19:05:24Z","receivedAt":"2017-03-01T21:52:54Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Feb 28, 2017 at 03:11:32PM -0800, Linus Torvalds wrote:\n\n> > Of course for dedicated code this can be simplified, and some parts\n> > could be further optimized.\n> \n> So I'd be worried about changing your tested code too much, since the\n> only test-cases we have are the two pdf files. If we screw up too\n> much, those will no longer show as collisions, but we could get tons\n> of false positives that we wouldn't see, so..\n\nI can probably help with collecting data for that part on GitHub.\n\nI don't have an exact count of how many sha1 computations we do in a\nday, but it's...a lot. Obviously every pushed object gets its sha1\ncomputed, but read operations also cover every commit and tree via\nparse_object() (though I think most of the blob reads do not).\n\nSo it would be trivial to start by swapping out the \"die()\" on collision\nwith something that writes to a log. This is the slow path that we don't\nexpect to trigger at all, so log volume shouldn't be a problem.\n\nI've been waiting to see how speedups develop before deploying it in\nproduction.\n\n-Peff\n"},{"id":"313095","messageId":"CA+55aFwXaSAMF41Dz3u3nS+2S24umdUFv0+k+s18UyPoj+v31g@mail.gmail.com","threadId":"45202","inReplyTo":"CA+55aFw6BLjPK-F0RGd9LT7X5xosKOXOxuhmKX65ZHn09r1xow@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-02T19:55:45Z","receivedAt":"2017-03-02T20:04:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Feb 24, 2017 at 4:39 PM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> Honestly, I think that a primary goal for a new hash implementation\n> absolutely needs to be to minimize mixing.\n>\n> Not for security issues, but because of combinatorics. You want to\n> have a model that basically reads old data, but that very aggressively\n> approaches \"new data only\" in order to avoid the situation where you\n> have basically the exact same tree state, just _represented_\n> differently.\n>\n> For example, what I would suggest the rules be is something like this:\n\nHmm. Having looked at this a fair amount, and in particularly having\nlooked at the code as part of the hash typesafety patch I did, I am\nactually starting to think that it would not be too painful at all to\nhave a totally different approach, which might be a lot easier to do.\n\nSo bear with me, let me try to explain my thinking:\n\n (a) if we want to be backwards compatible and not force people to\nconvert their trees on some flag day, we're going to be stuck with\nhaving to have the SHA1 code around and all the existing object\nparsing for basically forever\n\nNow, this basically means that the _easy_ solution would be that we\njust do the flag day, switch to sha-256, extend everything to 32-byte\nhashes, and just have a \"git2 fast-import\" that makes it easy to\nconvert stuff.\n\nBut it really would be a completely different version of git, with a\nnew pack-file format and no real compatibility. Such a flag-day\napproach would certainly have advantages: it would allow for just\nre-architecting some bad choices:\n\n - make the hashing be something that can be threaded (ie maybe we can\njust block it up in 4MB chunks that you can hash in parallel, and make\nthe git object hash be the hash of hashes)\n\n - replace zlib with something like zstd\n\n - get rid of old pack formats etc.\n\nbut  on the whole, I still think that the compatibility would be worth\nmuch more than the possible technical advantages of a clean slate\nrestart.\n\n (b) the SHA1 hash is actually still quite strong, and the collision\ndetection code effectively means that we don't really have to worry\nabout collisions in the immediate future.\n\nIn other words, the mitigation of the current attack is actually\nreally easy technically (modulo perhaps the performance concerns), and\nthere's still nothing fundamentally wrong with using SHA1 as a content\nhash. It's still a great hash.\n\nNow, my initial reaction (since it's been discussed for so long\nanyway) was obviously \"pick a different hash\". That was everybody's\ninitial reaction, I think.\n\nBut I'm starting to think that maybe that initial obvious reaction was wrong.\n\nThe thing that makes collision attacks so nasty is that our reaction\nto a collision is so deadly.  But that's not necessarily fundamental:\nwe certainly uses hashes with collisions every day, and they work\nfine. And they work fine because the code that uses those hashes is\ndesigned to simply deal gracefully - although very possibly with a\nperformance degradation - with two different things hashing to the\nsame bucket.\n\nSo what if the solution to \"SHA1 has potential collisions\" is \"any\nhash can have collisions in theory, let's just make sure we handle\nthem gracefully\"?\n\nBecause I'm looking at our data structures that have hashes in them,\nand many of them are actually of the type where I go\n\n  \"Hmm..  Replacing the hash entirely is really really painful - but\nit wouldn't necessarily be all that nasty to extend the format to have\nadditional version information\".\n\nand the advantage of that approach is that it actually makes the\ncompatibility part trivial. No more \"have to pick a different hash and\ntwo different formats\", and more of a \"just the same format with\nextended information that might not be there for old objects\".\n\nSo we have a few different types of formats:\n\n - the purely local data structures: the pack index file, the file\nindex, our refs etc\n\n   These we could in change completely, and it wouldn't even be all\nthat painful. The pack index has already gone through versions, and it\ndoesn't affect anything else.\n\n - the actual objects.\n\n   These are fairly painful to change, particularly things like the\n\"tree\" object which is probably the worst designed of the lot. Making\nit contain a fixed-size binary thing was really a nasty mistake. My\nbad.\n\n - the pack format and the protocol to exchange \"I have this\" information\n\n   This is *really* painful to change, because it contains not just\nthe raw object data, but it obviously ends up being the wire format\nfor remote accesses.\n\nand it turns out that *all* of these formats look like they would be\nfairly easy to extend to having extra object version information. Some\nof that extra object version information we already have and don't\nuse, in fact.\n\nEven the tree format, with the annoying fixed-size binary blob. Yes,\nit has that fixed size binary blob, but it has it in _addition_ to the\nASCII textual form that would be really easy to just extend upon. We\nhave that \"tree entry type\" that we've already made extensions with by\nusing it for submodules. It would be quite easy to just say that a\ntree entry also has a \"file version\" field, so that you can have\nmultiple objects that just hash to the same SHA1, and git wouldn't\neven *care*.\n\nThe transfer protocol is the same: yes, we transfer hashes around, but\nit would not be all that hard to extend it to \"transfer hash and\nobject version\".\n\nAnd the difference is that then the \"backwards compatibility\" part\njust means interacting with somebody who didn't know to transfer the\nobject version. So suddenly being backwards compatible isn't a whole\ndifferent object parsing thing, it's just a small extension.\n\nIOW, we could make it so that the SHA1 is just a hash into a list of\nobjects. Even the pack index format wouldn't need to change - right\nnow we assume that an index hit gives us the direct pointer into the\npack file, but we *could* just make it mean that it gives us a direct\npointer to the first object in the pack file with that SHA1 hash.\nExactly like you'd normally use a hash table with linear probing.\n\nLinear probing is usually considered a horrible approach to hash\ntables,. but it's actually a really useful one for the case where\ncollisions are very rare.\n\nAnyway, I do have a suggestion for what the \"object version\" would be,\nbut I'm not even going to mention it, because I want people to first\nthink about the _concept_ and not the implementation.\n\nSo: What do you think about the concept?\n\n               Linus\n"},{"id":"313098","messageId":"22712.24775.714535.313432@chiark.greenend.org.uk","threadId":"45202","inReplyTo":"20170228214724.w7w5f6n4u6ehanzd@genre.crustytoothpaste.net","subject":"Re: Transition plan for git to move to a new hash function","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-03-02T18:13:27Z","receivedAt":"2017-03-02T20:38:45Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"brian m. carlson writes (\"Re: Transition plan for git to move to a new hash function\"):\n> On Mon, Feb 27, 2017 at 01:00:01PM +0000, Ian Jackson wrote:\n> > Objects of one hash may refer to objects named by a different hash\n> > function to their own.  Preference rules arrange that normally, new\n> > hash objects refer to other new hash objects.\n> \n> The existing codebase isn't really intended with that in mind.\n\nYes.  I've seen the attempts to start to replace char* with a hash\nstruct.\n\n> I like Peff's suggested approach in which we essentially rewrite history\n> under the hood, but have a lookup table which looks up the old hash\n> based on the new hash.  That allows us to refer to old objects, but not\n> have to share serialized data that mentions both hashes.\n\nI think this means that the when a project converts, every copy of the\nhistory must be rewritten (separately).  Also, this leaves the whole\nsystem lacking in algorithm agililty.  Meaning we may have to do all\nof this again some time.\n\nI also think that we need to distinguish old hashes from new hashes in\nthe command line interface etc.  Otherwise there is a possible\nambiguity.\n\n> > The object name textual syntax is extended.  The new syntax may be\n> > used in all textual git objects and protocols (commits, tags, command\n> > lines, etc.).\n> > \n> > We declare that the object name syntax is henceforth\n> >   [A-Z]+[0-9a-z]+ | [0-9a-f]+\n> > and that names [A-Z].* are deprecated as ref name components.\n> \n> I'd simply say that we have data always be in the new format if it's\n> available, and tag the old SHA-1 versions instead.  Otherwise, as Peff\n> pointed out, we're going to be stuck typing a bunch of identical stuff\n> every time.  Again, this encourages migration.\n\nThe hash identifier is only one character.  Object names are not\nnormally typed very much anyway.\n\nIf you say we must decorate old hashes, then all existing data\neverywhere in the world which refers to any git objects by object name\nwill become invalid.  I don't mean just data in git here.  I mean CI\nsystems, mailing list archives, commit messages (perhaps in other\nversion control systems), test cases, and so on.\n\nIan.\n"},{"id":"313099","messageId":"xmqqk287be9l.fsf@gitster.mtv.corp.google.com","threadId":"45202","inReplyTo":"CA+55aFwXaSAMF41Dz3u3nS+2S24umdUFv0+k+s18UyPoj+v31g@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-03-02T20:43:50Z","receivedAt":"2017-03-02T20:45:06Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Anyway, I do have a suggestion for what the \"object version\" would be,\n> but I'm not even going to mention it, because I want people to first\n> think about the _concept_ and not the implementation.\n>\n> So: What do you think about the concept?\n\nMy reaction heavily depends on how that \"object version\" thing\nworks.  When I think I have \"variant #1\" of an object and say\n\n  have 860cd699c285f02937a2edbdb78e8231292339a5#1\n\nis there any guarantee that the other end has a (small) set of\ndifferent objects all sharing the same SHA-1 and it thinks it has\n\"variant #1\" only when it has the same thing as I have (otherwise,\nit may have \"variant #2\" that is an unrelated object but happens to\nshare the same hash)?  If so, I think I understand how things would\nwork within your \"concept\".  But otherwise, I am not really sure.\n\nWould \"object version\" be like a truncated SHA-1 over the same data\nbut with different IV or something, i.e. something that guarantees\nanybody would get the same result given the data to be hashed?\n\n\n"},{"id":"313106","messageId":"CA+55aFzdJcFZQdJ78rvZ92nNZRsKfcUKMCiXwTVYR34JAuznrA@mail.gmail.com","threadId":"45202","inReplyTo":"20170302215457.l2zhxgnvhulw2hl5@kitenet.net","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-02T22:27:15Z","receivedAt":"2017-03-02T22:35:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Mar 2, 2017 at 1:54 PM, Joey Hess <id@joeyh.name> wrote:\n>\n> There's a surprising result of combining iterated hash functions, that\n> the combination is no more difficult to attack than the strongest hash\n> function used.\n\nDuh. I should actually have known that. I started reading the paper\nand went \"this seems very familiar\". I'm pretty sure I've been pointed\nat that paper before (or maybe just a similar one), and I just didn't\nreact enough for it to leave a lasting impact.\n\n              Linus\n"},{"id":"313109","messageId":"20170302215457.l2zhxgnvhulw2hl5@kitenet.net","threadId":"45202","inReplyTo":"CA+55aFziZRA29foAMbM-HS5fiup7T0TuYf4XQ1kNT_SR7FfSgw@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2017-03-02T21:54:57Z","receivedAt":"2017-03-03T00:33:15Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Linus Torvalds wrote:\n> So you'd have to be able to attack both the full SHA1, _and_ whatever\n> other different good hash to 128 bits.\n\nThere's a surprising result of combining iterated hash functions, that\nthe combination is no more difficult to attack than the strongest hash\nfunction used.\n\nhttps://www.iacr.org/cryptodb/archive/2004/CRYPTO/1472/1472.pdf\n\nPerhaps you already knew about this, but I had only heard rumors\nthat was the case, until I found that reference recently.\n\n-- \nsee shy jo\n"},{"id":"313110","messageId":"20170302214610.GA86054@google.com","threadId":"45202","inReplyTo":"20170226051834.i37mlqv5wxwz3254@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Brandon Williams","fromEmail":"bmwill@google.com","sentAt":"2017-03-02T21:46:10Z","receivedAt":"2017-03-03T00:45:33Z","isPatch":false,"sender":{"key":"bwilliams.eng@gmail.com","avatar":null},"body":"On 02/26, Jeff King wrote:\n> On Sun, Feb 26, 2017 at 01:13:59AM +0000, Jason Cooper wrote:\n> \n> > On Fri, Feb 24, 2017 at 10:10:01PM -0800, Junio C Hamano wrote:\n> > > I was thinking we would need mixed mode support for smoother\n> > > transition, but it now seems to me that the approach to stratify the\n> > > history into old and new is workable.\n> > \n> > As someone looking to deploy (and having previously deployed) git in\n> > unconventional roles, I'd like to add one caveat.  The flag day in the\n> > history is great, but I'd like to be able to confirm the integrity of\n> > the old history.\n> > \n> > \"Counter-hashing\" the blobs is easy enough, but the trees, commits and\n> > tags would need to have, iiuc, some sort of cross-reference.  As in my\n> > previous example, \"git tag -v v3.16\" also checks the counter hash to\n> > further verify the integrity of the history (yes, it *really* needs to\n> > check all of the old hashes, but I'd like to make sure I can do step one\n> > first).\n> > \n> > Would there be opposition to counter-hashing the old commits at the flag\n> > day?\n> \n> I don't think a counter-hash needs to be embedded into the git objects\n> themselves. If the \"modern\" repo format stores everything primarily as\n> sha-256, say, it will probably need to maintain a (local) mapping table\n> of sha1/sha256 equivalence. That table can be generated at any time from\n> the object data (though I suspect we'll keep it up to date as objects\n> enter the repository).\n> \n> At the flag day[1], you can make a signed tag with the \"correct\" mapping\n> in the tag body (so part of the actual GPG signed data, not referenced\n> by sha1). Then later you can compare that mapping to the object content\n> in the repo (or to the local copy of the mapping based on that data).\n> \n> -Peff\n> \n> [1] You don't even need to wait until the flag day. You can do it now.\n>     This is conceptually similar to the git-evtag tool, though it just\n>     signs the blob contents of the tag's current tree state. Signing the\n>     whole mapping lets you verify the entirety of history, but of course\n>     that mapping is quite big: 20 + 32 bytes per object for\n>     sha1/sha-256, which is ~250MB for the kernel. So you'd probably not\n>     want to do it more than once.\n\nThere were a few of us discussing this sort of approach internally.  We\nalso figured that, given some performance hit, you could maintain your\nrepo in sha256 and do some translation to sha1 if you need to push or\nfetch to a server which has the the repo in a sha1 format.  This way you\ncan convert your repo independently of the rest of the world.\n\nAs for storing the translation table, you should really only need to\nmaintain the table until old clients are phased out and all of the repos\nof a project have experienced flag day and have been converted to\nsha256.\n\n-- \nBrandon Williams\n"},{"id":"313112","messageId":"CA+55aFziZRA29foAMbM-HS5fiup7T0TuYf4XQ1kNT_SR7FfSgw@mail.gmail.com","threadId":"45202","inReplyTo":"xmqqk287be9l.fsf@gitster.mtv.corp.google.com","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-02T21:21:30Z","receivedAt":"2017-03-03T01:43:46Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Mar 2, 2017 at 12:43 PM, Junio C Hamano <gitster@pobox.com> wrote:\n>\n> My reaction heavily depends on how that \"object version\" thing\n> works.\n>\n> Would \"object version\" be like a truncated SHA-1 over the same data\n> but with different IV or something, i.e. something that guarantees\n> anybody would get the same result given the data to be hashed?\n\nYes, it does need to be that in practice. So what I was thinking the\nobject version would be is:\n\n (a) we actually take the object type into account explicitly.\n\n (b) we explicitly add another truncated hash.\n\nThe first part we can already do without any actual data structure\nchanges, since basically all users already know the type of an object\nwhen they look it up.\n\nSo we already have information that we could use to narrow down the\nhash collision case if we saw one.\n\nThere are some (very few) cases where we don't already explicitly have\nthe object type (a tag reference can be any object, for example, and\nexisting scripts might ask for \"give me the type of this SHA1 object\nwith \"git cat-file -t\"), but that just goes back to the whole \"yeah,\nwe'll handle legacy uses and we will look up objects even _without_\nthe extra version data, so it actually integrates well into the whole\nnotion.\n\nBasically, once you accept that \"hey, we'll just have a list of\nobjects with that hash\", it just makes sense to narrow it down by the\nobject type we also already have.\n\nBut yes, the object type is obviously only two bits of information\n(actually, considering the type distribution, probably just one bit),\nand it's already encoded in the first hash, so it doesn't actually\nhelp much as \"collision avoidance\" particularly once you have a\nparticular attack against that hash in place.\n\nIt's just that it *is* extra information that we already have, and\nthat is very natural to use once you start thinking of the hash lookup\nas returning a list of objects. It also mitigates one of the worst\n_confusions_ in git, and so basically mitigates the worst-case\ndownside of an attack basically for free, so it seems like a\nno-brainer.\n\nBut the real new piece of object version would be a truncated second\nhash of the object.\n\nI don't think it matters too much what that second hash is, I would\nsay that we'd just approximate having a total of 256 bits of hash.\n\nSince we already have basically 160 bits of fairly good hashing, and\nroughly 128 bits of that isn't known to be attackable, we'd just use\nanother hash and truncate that to 128 bits. That would be *way*\noverkill in practice, but maybe overkill is what we want. And it\nwouldn't really expand the objects all that much more than just\npicking a new 256-bit hash would do.\n\nSo you'd have to be able to attack both the full SHA1, _and_ whatever\nother different good hash to 128 bits.\n\n                Linus\n\nPS.  if people think that SHA1 is of a good _size_, and only worry\nabout the known weaknesses of the hashing itself, we'd only need to\nget back the bits that the attacks take away from brute force. That's\ncurrently the 80 -> ~63 bits attack, so you'd really only want about\n40 bits of second hash to claw us back back up to 80 bits of brute\nforce (again: brute force is basically sqrt() of the search space, so\nhalf the bits, so adding 40 bits of hash adds 20 bits to the brute\nforce cost and you'd get back up to the 2**80 we started with).\n\nSo 128 bits of secondary hash really is much more than we'd need. 64\nbits would probably be fine.\n"},{"id":"313113","messageId":"20170303015047.p4lpkdzp4hbpz5vi@glandium.org","threadId":"45202","inReplyTo":"CA+55aFzdJcFZQdJ78rvZ92nNZRsKfcUKMCiXwTVYR34JAuznrA@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2017-03-03T01:50:47Z","receivedAt":"2017-03-03T01:51:44Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Thu, Mar 02, 2017 at 02:27:15PM -0800, Linus Torvalds wrote:\n> On Thu, Mar 2, 2017 at 1:54 PM, Joey Hess <id@joeyh.name> wrote:\n> >\n> > There's a surprising result of combining iterated hash functions, that\n> > the combination is no more difficult to attack than the strongest hash\n> > function used.\n> \n> Duh. I should actually have known that. I started reading the paper\n> and went \"this seems very familiar\". I'm pretty sure I've been pointed\n> at that paper before (or maybe just a similar one), and I just didn't\n> react enough for it to leave a lasting impact.\n\nWhat if the \"object version\" is a hash of the content (as opposed to\nheader + content like the normal git hash)?\n\nMike\n"},{"id":"313126","messageId":"CA+55aFyTXodj=EpEu0GNpcNak3GMCMyq7w0_SnzdcU3VKQAsgQ@mail.gmail.com","threadId":"45202","inReplyTo":"20170303015047.p4lpkdzp4hbpz5vi@glandium.org","subject":"Re: SHA1 collisions found","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-03T02:19:07Z","receivedAt":"2017-03-03T02:44:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Thu, Mar 2, 2017 at 5:50 PM, Mike Hommey <mh@glandium.org> wrote:\n>\n> What if the \"object version\" is a hash of the content (as opposed to\n> header + content like the normal git hash)?\n\nIt doesn't actually matter for that attack.\n\nThe concept of the attack is actually fairly simple: generate a lot of\ncollisions in the first hash (and they outline how you only need to\ngenerate 't' serial collisions and turn them into 2**t collisions),\nand then just check those collisions against the second hash.\n\nIf you have enough collisions in the first one, having a collision in\nthe second one is inevitable just from the birthday rule.\n\nNow, *In*practice* that attack is not very easy to do. Each collision\nis still hard to generate. And because of the git object rules (the\nsize has to match), it limits you a bit in what collisions you\ngenerate.\n\nBut the fact that git requires the right header can be considered just\na variation on the initial state to SHA1, and then the additional\nrequirement might be as easy as just saying that your collision\ngeneration function just always needs to generate a fixed-size block\n(which could be just 64 bytes - the SHA1 blocking size).\n\nSo assuming you can arbitrarily generate collisions (brute force:\n2**80 operations) you could make the rule be that you generate\nsomething that starts with one 64-byte block that matches the git\nrules:\n\n   \"blob 6454\\0\"..pad with repeating NUL bytes..\n\nand then you generate 100 pairs of 64-byte SHA1 collisions (where the\nfirst starts with the initial value of that fixed blob prefix, the\nnext with the state after the first block, etc etc).\n\nNow you can generate 2**100 different sequences that all are exactly\n6464 bytes (101 64-byte blocks) and all have the same SHA1 - all all\nshare that first fixed 64-byte block.\n\nYou needed \"only\" on the order of 100 * 2**80 SHA1 operations to do\nthat in theory.\n\nAn since you have 2**100 objects, you know that you will have a likely\nbirthday paradox even if your secondary hash is 200 bits long.\n\nSo all in all, you generated a collision in on the order of 2**100 operations.\n\nSo instead of getting the security of \"160+200\" bits, you only got 200\nbits worth of real collision resistance.\n\nNOTE!! All of the above is very much assuming \"brute force\". If you\ncan brute-force the hash, you can completely break any hash. The input\nblock to SHA1 is 64 bytes, so by definition you have 512 bits of data\nto play with, and you're generating a 160-bit output: there is no\nquestion what-so-ever that you couldn't generate any hash you want if\nyou brute-force things.\n\nThe place where things like having a fixed object header can help is\nwhen the attack in question requires some repeated patterns. For\nexample, if you're not brute-forcing things, your attack on the hash\nwill likely involve using very particular patterns to change a number\nof bits in certain ways, and then combining those particular patterns\nto get the hash collision you wanted.  And *that* is when you may need\nto add some particular data to the middle to make the end result be a\nparticular match.\n\nBut a brute-force attack definitely doesn't need size changes. You can\nmake the size be pretty much anything you want (modulo really small\ninputs, of course - a one-byte input only has 256 different possible\nhashes ;) if you have the resources to just go and try every possible\ncombination until you get the hash you wanted.\n\nI may have overly simplified the paper, but that's the basic idea.\n\n               Linus\n"},{"id":"313160","messageId":"20170303111347.6uzuhvmpdwr27qjw@sigill.intra.peff.net","threadId":"45202","inReplyTo":"20170302214610.GA86054@google.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-03T11:13:48Z","receivedAt":"2017-03-03T11:42:24Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 02, 2017 at 01:46:10PM -0800, Brandon Williams wrote:\n\n> There were a few of us discussing this sort of approach internally.  We\n> also figured that, given some performance hit, you could maintain your\n> repo in sha256 and do some translation to sha1 if you need to push or\n> fetch to a server which has the the repo in a sha1 format.  This way you\n> can convert your repo independently of the rest of the world.\n\nYeah, you definitely _can_ convert between the two. It's just expensive\nto do so on the fly. We'd potentially want to be able to during the\ntransition period just to help people get all the work converted over.\nBut I had assumed that the conversion would be a mix of:\n\n  1. Unpublished work (or work which is otherwise OK to be rewritten)\n     could be converted to the new hash.\n\n  2. Old history could be grafted with a parent pointer that mentions\n     the tip of the old history by its new hash, but the pointed-to\n     parent contains sha1s.\n\n> As for storing the translation table, you should really only need to\n> maintain the table until old clients are phased out and all of the repos\n> of a project have experienced flag day and have been converted to\n> sha256.\n\nI think you've read more into my \"conversion\" than I intended. The old\nhistory won't get rewritten. It will just be grafted onto the bottom of\nthe commit history you've got, and the new trees will all be written\nwith the new hash.\n\nSo you still have those old objects hanging around that refer to things\nby their sha1 (not to mention bug trackers, commit messages, etc, which\nall use commit ids). And you want to be able to quickly resolve those\nreferences.\n\nWhat _does_ get rewritten is what's in your ref files, your pack .idx,\netc. Those are all sha256 (or whatever), and work as sha1's do now.\nLooking up a sha1 reference from an old object just goes through the\nextra level of indirection.\n\n-Peff\n"},{"id":"313181","messageId":"22713.33728.502854.338516@chiark.greenend.org.uk","threadId":"45202","inReplyTo":"20170303111347.6uzuhvmpdwr27qjw@sigill.intra.peff.net","subject":"Re: SHA1 collisions found","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-03-03T14:54:56Z","receivedAt":"2017-03-03T15:21:37Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"Jeff King writes (\"Re: SHA1 collisions found\"):\n> I think you've read more into my \"conversion\" than I intended. The old\n> history won't get rewritten. It will just be grafted onto the bottom of\n> the commit history you've got, and the new trees will all be written\n> with the new hash.\n> \n> So you still have those old objects hanging around that refer to things\n> by their sha1 (not to mention bug trackers, commit messages, etc, which\n> all use commit ids). And you want to be able to quickly resolve those\n> references.\n> \n> What _does_ get rewritten is what's in your ref files, your pack .idx,\n> etc. Those are all sha256 (or whatever), and work as sha1's do now.\n\nThis all sounds very similar to my proposal.\n\n> Looking up a sha1 reference from an old object just goes through the\n> extra level of indirection.\n\nI don't understand why this is a level of indirection, rather than\nsimply a retention of the existing SHA-1 object database (in parallel,\nbut deprecated).\n\nPerhaps I have misunderstood what you mean by \"graft\".  I assume you\ndon't mean info/grafts, because that is not conveyed by transfer\nprotocols.\n\n\nStepping back a bit, the key question is what the data structure will\nlook like after the transition.\n\nSpecifically, the parent reference in the first post-transition commit\nhas to refer to something.  What does it refer to ?  The possibilities\nseem to be:\n\n 1a. It names the SHA1 hash of an old commit object\n 1b. It names the BLAKE[1] hash of an old commit object, which\n    object of course refers to its parents by SHA1.\n\n 2. It names the BLAKE hash of a translated version of the\n    old commit object.\n\n 3. It doesn't name the parent, and the old history is not\n    automatically transferred by clone and not readily accessible.\n\n(1a) and (1b) are different variants of something like my mixed hash\nproposal.  Old and new hashes live side by side.\n\n(2) involves rewriting all of the old history, to recursively generate\nnew objects (with BLAKE names, and which in turn refer to other\nrewritten old objects by BLAKE names).  The first time a particular\ntree needs to look up an object by a BLAKE name, it needs to run a\nconversion its own entire existing history.\n\nFor (2) there would have to be some kind of mapping table in every\ntree, which allowed object names to be maped in both directions.  The\nobject content translation would have to be reversible, so that the\nactual pre-translation objects would not need to be stored; rather\nthey would be reconstructed from the post-translation objects, when\nsomeone asks for a pre-translation object.  In principle it would be\npossible to convert future BLAKE commits to SHA-1 ones, again by\nrecursive rewriting.\n\nI don't think anyone is seriously suggesting (3).\n\n\nSo there is a choice between:\n\n(1) a unified hash tree containing a mixture of different hashes at\ndifferent reference points, where each object has one identity and one\nname.\n\n(2) parallel hash tree structures, each using only a single hash, with\nat least every old object present in both tree structures.\n\nI think (1) is preferable because it provides, to callers of git, the\nexisting object naming semantics.  Callers need to be taught to accept\nan extension to the object name format.  Existing object names stored\nelsewhere than in git remain valid.\n\nConversely, (2) requires many object names stored elsewhere than in\ngit to be updated.  It's possible with (2) to do ad-hoc lookups on\nobject names in mailing list messages or commit messages and so on.\nEven if it is possible for the new git to answer questions like \"does\nthis new branch with BLAKE hash X' contain the commit with SHA1 hash\nY\" by implicitly looking up the corresponding BLAKE commit Y' and\nanswering the question with reference to Y', this isn't going to help\nif external code does things like \"have git compute the merge base of\nX and Y' and check that it is equal to Z\".  Either the external\ndatabase's record of Z would have to be changed to refer to Z', or the\nexternal code would have to be taught to apply an object name\nnormalisation operation to turn Z into Z' each time.\n\nAlso, (2) has trouble with shallow clones.  This is because it's not\npossible to translate old objects to new ones without descending to\nthe roots of the object tree and recursively translating everything\n(or looking up the answer of a previous translation).\n\n\nThen there is the question of naming syntax.\n\nThe distinction between (1) single unified mixed hash tree, and\n(2) multiple parallel homogenous hash trees, is mostly orthogonal to\nthe command-line (and in-commit-object etc.) representation of new\nhashes.\n\nThe main thing here is that, regardless of the choice between (1) or\n(2), we need to choose whether object names specified on the git\ncommand line, and printed by normal git commands, explicitly identify\nthe hash function.\n\nI think there are roughly three options:\n\n (A) Decorate all new hashes with a hash function indication\n     (sha256:<hex> or blake_<hex> or H<hex>)\n\n (B) Infer the hash function from the object name length\n     (and do some kind of bodge for abbreviated object names).\n\n (C) Hash function is implicit from context.  (This is compatible with\n     (2) only, because (1) requires any object to be able to refer to\n     any hash function.)\n\nI think (A) is best because it means everything is unambiguous, and it\nallows future hash function changes without further difficulty.\n\n(B) is a reasonable possibility although the abbreviated object name\nbodge would be quite ugly and probably involve thinking about several\nannoying edge cases.\n\nI think (C) is really bad, because it instantly makes all existing\napplication code which calls git to be buggy.  Such application code\nwould need to be adjusted to know for itself which of the object names\nit has recorded are what hash function, and explicitly specify this to\nits git operations somehow.\n\nAll of these options involve updating many callers of git.  In any\ncase any git caller which explicitly checks the object name length\nwill need to be changed.  For (a), many git callers which match object\nnames using something like [0-9a-f]+ rather than \\w+ will need to be\nchanged - but at least it's a simple change with little semantic\nimport.\n\n(A) has the additional advantage that it becomes possible to make\nobject names syntactically distinguishable from ref names.\n\n\nThe final argument I would make is this:\n\nWe don't know what hash function research will look like in 10-20\nyears.  We would like to not have a bunch of pain again.  So ideally\nwe would deploy a framework now that would let us switch hash function\nagain without further history-rewriting.\n\n(1)(A) and perhaps (1)(B) are the only options which support this\nwell.\n\n\nIan.\n\n[1] I'm going to keep assuming that the bikeshed will be blue, because\nI think BLAKE2b has is a better choice.  It has probably had more\nserious people looking at it than SHA-3, at least, and it has good\nperformance.  The web page has an impressive adoption list - probably\nwider than SHA-3.\n\n-- \nIan Jackson <ijackson@chiark.greenend.org.uk>   These opinions are my own.\n\nIf I emailed you from an address @fyvzl.net or @evade.org.uk, that is\na private address which bypasses my fierce spamfilter.\n"},{"id":"313227","messageId":"20170303110448.se3bstlk5hr4hqv3@sigill.intra.peff.net","threadId":"45202","inReplyTo":"CA+55aFwXaSAMF41Dz3u3nS+2S24umdUFv0+k+s18UyPoj+v31g@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-03T11:04:48Z","receivedAt":"2017-03-03T20:05:57Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 02, 2017 at 11:55:45AM -0800, Linus Torvalds wrote:\n\n> Anyway, I do have a suggestion for what the \"object version\" would be,\n> but I'm not even going to mention it, because I want people to first\n> think about the _concept_ and not the implementation.\n> \n> So: What do you think about the concept?\n\nI think it very much depends on what's in the \"object version\". :)\n\nIMHO, we are best to consider sha1 \"broken\" and not count on any of its\nbytes for cryptographic integrity. I know that's not really the case,\nbut it just makes reasoning about the whole thing simpler. So at that\npoint, it's pretty obvious that the \"object version\" is really just \"an\nintegrity hash\".\n\nAnd that takes us full circle to earlier proposals over the years to do\nsomething like this in the commit header:\n\n  parent ...some sha1...\n  parent-sha256 ...some sha256...\n\nand ditto in tag headers, and trees obviously need to be hackily\nextended as you described to carry the extra hash. And then internally\nwe continue to happily use sha1s, except you can check the\nsha256-validity of any reference if you feel like it.\n\nThis is functionally equivalent to \"just start using sha-256, but keep a\nmapping of old sha1s to sha-256s to handle old references\". The\nadvantage is that it makes the code part of the transition simpler. The\ndisadvantage is that you're effectively carrying a piece of that\nsha1->sha256 mapping around in _every_ object.\n\nAnd that means the same bits of mapping data are repeated over and over.\nGit's pretty good at de-duplicating on the surface. So yeah, every tree\nentry is now 256 bits larger, but deltas mean that we usually only end\nup storing each entry a handful of times. But we still pay the price to\nwalk over the bytes every time we apply a delta, zlib inflate, parse the\ntree, etc. The runtime cost of the transition is carried forward\nforever, even for repositories that are willing to rewrite history, or\nare created after the flag day.\n\nSo I dunno. Maybe I am missing something really clever about your\nproposal. Reading the rest of the thread, it sounds like you had a\nthought that we could get by with a very tiny object version, but the\nhash-adding thing nixed that. If I'm still missing the point, please try\nto sketch it out a bit more concretely, and I'll come back with my\nthinking cap on.\n\n-Peff\n"},{"id":"313240","messageId":"CAGZ79kY+TwN5xp8R7-UHecuErkEmdpR-vnK8gU4zpaekCa1grQ@mail.gmail.com","threadId":"45202","inReplyTo":"CA+55aFwXaSAMF41Dz3u3nS+2S24umdUFv0+k+s18UyPoj+v31g@mail.gmail.com","subject":"Re: SHA1 collisions found","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2017-03-03T21:47:06Z","receivedAt":"2017-03-03T21:48:04Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Thu, Mar 2, 2017 at 11:55 AM, Linus Torvalds\n<torvalds@linux-foundation.org> wrote:\n>\n> So: What do you think about the concept?\n>\n>                Linus\n\nOne of the things I like about working on Git is its pretty\nhigh standard of testing. So we would need to come up with\ngood methods of testing this, e.g. when\nGIT_TEST_WITH_DEGENERATE_HASH is set, we'd use\nan intentionally weak hashing function and have tests for\nthe collisions. These tests would need cover most of the\nworkflows that are currently performed with Git\n(local creation, fetching, pushing). Writing all these\nadditional tests (which consists of creating colliding\nobjects/commits and then performing all these tests),\nsounds about as much work as actually converting to\na new hash function. (First locally and then at a later\npoint in time all the networking related things).\n\nI would not want to go that way.\n\nStefan\n"},{"id":"313241","messageId":"20170303221832.brm3gftekmdcubzi@sigill.intra.peff.net","threadId":"45202","inReplyTo":"22713.33728.502854.338516@chiark.greenend.org.uk","subject":"Re: SHA1 collisions found","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-03T22:18:32Z","receivedAt":"2017-03-03T22:19:08Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Mar 03, 2017 at 02:54:56PM +0000, Ian Jackson wrote:\n\n> > What _does_ get rewritten is what's in your ref files, your pack .idx,\n> > etc. Those are all sha256 (or whatever), and work as sha1's do now.\n> \n> This all sounds very similar to my proposal.\n\nYeah, sorry I haven't reviewed that more carefully yet.\n\n> > Looking up a sha1 reference from an old object just goes through the\n> > extra level of indirection.\n> \n> I don't understand why this is a level of indirection, rather than\n> simply a retention of the existing SHA-1 object database (in parallel,\n> but deprecated).\n\nI just meant that we will need to know both the sha1 and the sha-256 of\nevery object (because the same blob, for example, may be referred to by\neither name depending on whether it is a new or a historical tree).  So\none way to do that is to have a table mapping sha1 to sha-256, and then\na lookup of sha-1 goes through that before looking up the object content\non disk via sha-256.\n\nBut it may also be fine to just keep an index mapping sha1 directly to\nobject contents. That makes a sha1 lookup slightly faster, but it's more\nexpensive to do a sha-256 verification (you have to reverse-index the\nobject location in the sha-256 list).\n\n(And my usual disclaimer: I am using sha-256 as a placeholder; I don't\nhave a strong opinion on the actual hash choice).\n\n> Perhaps I have misunderstood what you mean by \"graft\".  I assume you\n> don't mean info/grafts, because that is not conveyed by transfer\n> protocols.\n\nNo, I just mean there will be a spot in the commit graph (or many spots,\npotentially) where a \"v2\" commit using sha-256 references (by sha-256) a\n\"v1\" commit that is full of sha-1s. I say \"graft\" only to differentiate\nit from the idea of rewriting the content of those commit objects.\n\n> Specifically, the parent reference in the first post-transition commit\n> has to refer to something.  What does it refer to ?  The possibilities\n> seem to be:\n> \n>  1a. It names the SHA1 hash of an old commit object\n>  1b. It names the BLAKE[1] hash of an old commit object, which\n>     object of course refers to its parents by SHA1.\n> \n>  2. It names the BLAKE hash of a translated version of the\n>     old commit object.\n> \n>  3. It doesn't name the parent, and the old history is not\n>     automatically transferred by clone and not readily accessible.\n> \n> (1a) and (1b) are different variants of something like my mixed hash\n> proposal.  Old and new hashes live side by side.\n\nThanks for laying out those options. My proposals have definitely been\nin the (1b) camp.\n\nI think (1a) is not so bad; it just bumps the transition point one\ncommit higher. But I think the \"rules\" for which hash to expect are\neasier if they depend on which object the _pointer_ is using, rather\nthan the _pointee_.\n\n> (2) involves rewriting all of the old history, to recursively generate\n> new objects (with BLAKE names, and which in turn refer to other\n> rewritten old objects by BLAKE names).  The first time a particular\n> tree needs to look up an object by a BLAKE name, it needs to run a\n> conversion its own entire existing history.\n> \n> For (2) there would have to be some kind of mapping table in every\n> tree, which allowed object names to be maped in both directions.  The\n> object content translation would have to be reversible, so that the\n> actual pre-translation objects would not need to be stored; rather\n> they would be reconstructed from the post-translation objects, when\n> someone asks for a pre-translation object.  In principle it would be\n> possible to convert future BLAKE commits to SHA-1 ones, again by\n> recursive rewriting.\n\nHmm. I had initially rejected this as being pretty nasty for accessing\nthe old format on the fly. But as long as you keep the bidirectional\nmapping from the initial expensive conversion, in most cases you only\nneed to convert at most a single object. E.g., the two cases that are\nreally interesting are:\n\n  - I have an old commit sha1 and I want to run \"git show\" on it. We\n    convert the sha1 to the BLAKE name, and then just show the BLAKE\n    contents as usual.\n\n  - I want to verify a commit or tag signature. We need the original\n    bytes for this. So we convert the sha1 to the BLAKE name to get the\n    BLAKE'd contents. Then we rewrite only the _single_ object,\n    converting any of its internal BLAKE references back to sha1s via\n    the mapping.\n\nThat's more appealing than I had originally given it credit for, because\nmost of the other code just happily uses the BLAKE name internally. Even\na diff across the conversion boundary works at full speed, because it's\nusing the same names consistently.\n\nThe big downside is that the mapping is more expensive to generate than\na 1b-style mapping. In 1b, you just compute both hashes over all\nincoming objects and store the sha1 in a side lookup table. But here you\nactually need to _rewrite_ each object to come up with its\nalternate-universe sha1. And you need to do it in reverse-graph order,\nnot the arbitrary order that something like index-pack uses.\n\nI had assumed that local repos would generate these tables themselves,\nso to not have to trust any other repos. And the alternate-universe\nthing is more to ask those local repos to do. But it would also be\nworkable to distribute the mapping out-of-band (e.g., via a signed tag).\n\nYou don't even really need to trust the mapping that much, if it's just\nused for historical name lookups.\n\n> I don't think anyone is seriously suggesting (3).\n\nYeah, agreed.\n\n> Conversely, (2) requires many object names stored elsewhere than in\n> git to be updated.  It's possible with (2) to do ad-hoc lookups on\n> object names in mailing list messages or commit messages and so on.\n> Even if it is possible for the new git to answer questions like \"does\n> this new branch with BLAKE hash X' contain the commit with SHA1 hash\n> Y\" by implicitly looking up the corresponding BLAKE commit Y' and\n> answering the question with reference to Y', this isn't going to help\n> if external code does things like \"have git compute the merge base of\n> X and Y' and check that it is equal to Z\".  Either the external\n> database's record of Z would have to be changed to refer to Z', or the\n> external code would have to be taught to apply an object name\n> normalisation operation to turn Z into Z' each time.\n\nI think the X->X' conversion on input is easy. The universe of objects\nwhose sha1 we are about is not getting any bigger. Outputting sha1 Z\ninstead of BLAKE Z' is a bit harder (and at the very least probably\nneeds the caller to use a specific command-line option).\n\n> Also, (2) has trouble with shallow clones.  This is because it's not\n> possible to translate old objects to new ones without descending to\n> the roots of the object tree and recursively translating everything\n> (or looking up the answer of a previous translation).\n\nYes. Though if we distribute a partial sha1/BLAKE mapping with the clone\n(i.e., just for the objects we're sending) then the client could still\nuse sha1 names. It not be that big a problem in practice, though. If\nyou're shallow, then you don't _have_ the old object names to refer to\nanyway.\n\n> The main thing here is that, regardless of the choice between (1) or\n> (2), we need to choose whether object names specified on the git\n> command line, and printed by normal git commands, explicitly identify\n> the hash function.\n> \n> I think there are roughly three options:\n> \n>  (A) Decorate all new hashes with a hash function indication\n>      (sha256:<hex> or blake_<hex> or H<hex>)\n> \n>  (B) Infer the hash function from the object name length\n>      (and do some kind of bodge for abbreviated object names).\n> \n>  (C) Hash function is implicit from context.  (This is compatible with\n>      (2) only, because (1) requires any object to be able to refer to\n>      any hash function.)\n> \n> I think (A) is best because it means everything is unambiguous, and it\n> allows future hash function changes without further difficulty.\n\nFor input, I think we should definitely _support_ A, but in practice I\nthink people would be happy if an undecorated hash (or partial hash)\nlooks it up as a BLAKE name, and then falls back to the sha1.\n\nIn (2) this is obviously the right thing to do, because all of our\noutput will be BLAKE names.\n\nIn (1) it is less clear if we might output sha1 names for old cases. But\nI think we're still better off using the stronger hash when possible.\n\n> (A) has the additional advantage that it becomes possible to make\n> object names syntactically distinguishable from ref names.\n\nSort of. Our get_sha1() parser accepts a lot of random syntax.  E.g.,\n\"sha256:1234abcd\" is ambiguous with a file inside the tree named by\n\"sha256\". In practice I don't mind carving out a namespace and letting\npeople with a branch named \"sha256\" rot.\n\n> We don't know what hash function research will look like in 10-20\n> years.  We would like to not have a bunch of pain again.  So ideally\n> we would deploy a framework now that would let us switch hash function\n> again without further history-rewriting.\n> \n> (1)(A) and perhaps (1)(B) are the only options which support this\n> well.\n\nYes, I think planning for another migration is a sensible thing.\n\nI just think we should not sacrifice any other properties to the idea\nthat people could flip on their bespoke hashes and interoperate with\nother users. I.e., \"git config core.hash sha3 && git push\" should not be\na use case we care about at all, because it creates all sorts of _other_\nheadaches.\n\nBut I have no objection to making the 20-years-from-now migration less\npainful, and I agree that (1b) is more along those lines.\n\n-Peff\n"},{"id":"313286","messageId":"20170304224936.rqqtkdvfjgyezsht@genre.crustytoothpaste.net","threadId":"45202","inReplyTo":"22712.24775.714535.313432@chiark.greenend.org.uk","subject":"Re: Transition plan for git to move to a new hash function","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-03-04T22:49:37Z","receivedAt":"2017-03-04T22:50:01Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Thu, Mar 02, 2017 at 06:13:27PM +0000, Ian Jackson wrote:\n> brian m. carlson writes (\"Re: Transition plan for git to move to a new hash function\"):\n> > On Mon, Feb 27, 2017 at 01:00:01PM +0000, Ian Jackson wrote:\n> > > Objects of one hash may refer to objects named by a different hash\n> > > function to their own.  Preference rules arrange that normally, new\n> > > hash objects refer to other new hash objects.\n> > \n> > The existing codebase isn't really intended with that in mind.\n> \n> Yes.  I've seen the attempts to start to replace char* with a hash\n> struct.\n\nMy comment actually has nothing to do with the way struct object_id is\nset up.  That actually can be trivially extended with a byte or two of\ntype.\n\nInstead, I was referring to areas like the notes code.  It has extensive\nuse of the last byte as a type of lookup table key.  It's very dependent\non having exactly one hash, since it will always want to use the last\nbyte.\n\nThere are other, more subtle areas of the code that just don't handle\nmultiple hashes well.  Ideally we would remedy this, but I think\neveryone is very eager to move away from SHA-1, and since nobody has\nstepped up to volunteer to do that work, we should probably adopt a\nsolution that doesn't involve doing that.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"313294","messageId":"22716.5770.95842.704242@chiark.greenend.org.uk","threadId":"45202","inReplyTo":"20170304224936.rqqtkdvfjgyezsht@genre.crustytoothpaste.net","subject":"Re: Transition plan for git to move to a new hash function","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-03-05T13:45:46Z","receivedAt":"2017-03-05T14:19:34Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"brian m. carlson writes (\"Re: Transition plan for git to move to a new hash function\"):\n> Instead, I was referring to areas like the notes code.  It has extensive\n> use of the last byte as a type of lookup table key.  It's very dependent\n> on having exactly one hash, since it will always want to use the last\n> byte.\n\nYou mean note_tree_search ?  (My tree here may be a bit out of date.)\nThis doesn't seem difficult to fix.  The nontrivial changes would be\nmostly confined to SUBTREE_SHA1_PREFIXCMP and GET_NIBBLE.\n\nIt's true that like most of git there's a lot of hardcoded `sha1'.\n\n\nAre you arguing in favour of \"replace git with git2 by simply\ns/20/64/g; s/sha1/blake/g\" ?  This seems to me to be a poor idea.\nTakeup of the new `git2' would be very slow because of the pain\ninvolved.\n\nAny sensible method of moving to a new hash that isn't \"make a\ncompletely incompatible new version of git\" is going to involve\nteaching the code we have in git right now to handle new hashes as\nwell as sha1 hashes.\n\nEven if the plan is to try to convert old data, rather than keep it\nand be able to refer to it from new data, something will have to be\nable to parse old packfiles, old commits, old tags, old notes,\netc. etc. etc.  Either that's going to be some separate conversion\nutility, or it has to be the same code in git that's there already.[1]\n\nThe ability to handle both old-format and new-format data can be\nachieved in the code by doing away with the hardcoded sha1s, so that\ninstead the hash is an abstract data type with operations like\n\"initialise\", \"compare\", \"get a nybble\", etc.  We've already seen\npatches going in this direction.\n\n[1] I've heard suggestions here that instead we should expect users to\n\"git1 fast-export\", which you would presumably feed into \"git2\nfast-import\".  But what is `git1' here ?  Is it the current git\ncodebase frozen in time ?  I don't think it can be.  With this\nconversion strategy, we will need to maintain git1 for decades.  It\nwill need portability fixes, security fixes, fixes for new hostile\ncompiler optimisations, and so on.  The difficulty of conversion means\nthere will be pressure to backport new features from `git2' to `git1'.\n(Also this approach means that all signatures are definitively lost\nduring the conversion process.)\n\nSo if we want to provide both `git1' and `git2', it's still better to\ncompile `git' and `git2' from the same codebase.  But if we do that,\nthe resulting ifdeffery and/or other hash abstractions are most of the\nwork to be hash-agile.  It's just the difference between a\ncompile-time and runtime switch.\n\nI think the incompatibile approach is much more work in the medium and\nlong term - and it leads to a longer transition period.\n\n\nBear in mind that our objective is not to minimise the time until the\nnew version of git is available.  Our objective is to minimise the\ntime until (most) people are using it.  An approach which takes longer\nfor the git community to develop, but which is easier to deploy, can\neasily be better.\n\nOr maybe the objective is to minimise overall effort.  In which case\nmore work on git, for an easier transition for all the users, seems\nlike a no-brainer.  I think this is arguably true even from the point\nof view of effort amongst the community of git contributors.  git\ncontributors start out as git users - and if git's users are all busy\nstruggling with a difficult transition, they will have less time to\nimprove other stuff and will tend less to get involved upstream.  (And\nthey may be less inclined to feel that the git upstream developers\nunderstand their needs well.)\n\nThe better alternative is to adopt a plan that has a clear and\nstraightforward transition for users, and ask git users to help with\nimplementation.\n\nI think many git users, including sophisticated users and competent\norganisations, are concerned about sha1.  Currently most of those\nusers will find it difficult to help, because it's not clear to them\nwhat needs to be done.\n\nThanks,\nIan.\n"},{"id":"313303","messageId":"20170305234527.qfsiopua6ygva46p@genre.crustytoothpaste.net","threadId":"45202","inReplyTo":"22716.5770.95842.704242@chiark.greenend.org.uk","subject":"Re: Transition plan for git to move to a new hash function","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-03-05T23:45:27Z","receivedAt":"2017-03-05T23:45:40Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sun, Mar 05, 2017 at 01:45:46PM +0000, Ian Jackson wrote:\n> brian m. carlson writes (\"Re: Transition plan for git to move to a new hash function\"):\n> > Instead, I was referring to areas like the notes code.  It has extensive\n> > use of the last byte as a type of lookup table key.  It's very dependent\n> > on having exactly one hash, since it will always want to use the last\n> > byte.\n> \n> You mean note_tree_search ?  (My tree here may be a bit out of date.)\n> This doesn't seem difficult to fix.  The nontrivial changes would be\n> mostly confined to SUBTREE_SHA1_PREFIXCMP and GET_NIBBLE.\n> \n> It's true that like most of git there's a lot of hardcoded `sha1'.\n\nI'm talking about the entire notes.c file.  There are several different\nuses of \"19\" in there, and they compose at least two separate concepts.\nMy object-id-part9 series tries to split those out into logical\nconstants.\n\nThis code is not going to handle repositories with different-length\nobjects well, which I believe was your initial proposal.  I originally\nthought that mixed-hash repositories would be viable as well, but I no\nlonger do.\n\n> Are you arguing in favour of \"replace git with git2 by simply\n> s/20/64/g; s/sha1/blake/g\" ?  This seems to me to be a poor idea.\n> Takeup of the new `git2' would be very slow because of the pain\n> involved.\n\nI'm arguing that the same binary ought to be able to handle both SHA-1\nand the new hash.  I'm also arguing that a given object have exactly one\nhash and that we not mix hashes in the same object.  A repository will\nbe composed of one type of object, and if that's the new hash, a lookup\ntable will be used to translate SHA-1.  We can synthesize the old\nobjects, should we need them.\n\nThat allows people to use the SHA-1 hashes (in my view, with a prefix,\nsuch as \"sha1:\") in repositories using the new hash.  It also allows\nverifying old tags and commits if need be.\n\nWhat I *would* like to see is an extension to the tag and commit objects\nwhich names the hash that was used to make them.  That makes it easy to\ndetermine which object the signature should be verified over, as it will\nverify over only one of them.\n\n> [1] I've heard suggestions here that instead we should expect users to\n> \"git1 fast-export\", which you would presumably feed into \"git2\n> fast-import\".  But what is `git1' here ?  Is it the current git\n> codebase frozen in time ?  I don't think it can be.  With this\n> conversion strategy, we will need to maintain git1 for decades.  It\n> will need portability fixes, security fixes, fixes for new hostile\n> compiler optimisations, and so on.  The difficulty of conversion means\n> there will be pressure to backport new features from `git2' to `git1'.\n> (Also this approach means that all signatures are definitively lost\n> during the conversion process.)\n\nI'm proposing we have a git hash-convert (the name doesn't matter that\nmuch) that converts in place.  It rebuilds the objects and builds a\nlookup table.  Since the contents of git objects are deterministic, this\nmakes it possible for each individual user to make the transition in\nplace.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"}]}