{"thread":{"id":"914","subject":"[zooko@zooko.com: [Revctrl] colliding md5 hashes of human-meaningful documents]","startedAt":"2005-06-12T08:25:55Z","lastAt":"2005-06-14T02:06:20Z","messageCount":5,"participants":["Petr Baudis","Morten Welinder","Martin Uecker","Linus Torvalds","Daniel Barkalow"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"4890","messageId":"20050612082555.GB6620@pasky.ji.cz","threadId":"914","inReplyTo":null,"subject":"[zooko@zooko.com: [Revctrl] colliding md5 hashes of human-meaningful documents]","fromName":"Petr Baudis","fromEmail":"pasky@ucw.cz","sentAt":"2005-06-12T08:25:55Z","receivedAt":"2005-06-12T08:25:55Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"----- Forwarded message from zooko@zooko.com -----\n\nThere is nothing theoretically surprising about this, but hopefully its\nconcreteness and the accompanying scenario will make an impression on people\non people.  The same technique should work to generate two documents with\nidentical SHA1 hashes.\n\nhttp://www.cits.rub.de/MD5Collisions/\n\n----- End forwarded message -----\n\nI expected the two postscript files differing in some huge binary blob,\nbut it turns out the binary part is very small (about 256 bytes) and\nonly few (about nine) bytes are different, contrary to how people have\npredicted the collisions. This is much more close to finding a collision\nbetween similar pure C files, I think. Rather unsettling.\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\n<Espy> be careful, some twit might quote you out of context..\n"},{"id":"4895","messageId":"118833cc050612061452f0eb3e@mail.gmail.com","threadId":"914","inReplyTo":"20050612082555.GB6620@pasky.ji.cz","subject":"Re: [zooko@zooko.com: [Revctrl] colliding md5 hashes of human-meaningful documents]","fromName":"Morten Welinder","fromEmail":"mwelinder@gmail.com","sentAt":"2005-06-12T13:14:03Z","receivedAt":"2005-06-12T13:14:03Z","isPatch":false,"sender":{"key":"mwelinder@gmail.com","avatar":null},"body":"I looked at this.\n\nIf you can find just *one* colliding pair that does not contain a\nfew select bytes like 0x22 (quote), 0x5c (backslash) and\nperhaps 0x0a then it is trivial to make a colliding pair of\nC programs.\n\nCall the initial pair (A,B) and make the programs\n\n/* Common prefix that happens to have integer block size\n   ending just before the A below. */\nconst char junk[] = \"AB\";\n/* \"AA\" in program 1, \"AB\" in program 2.  */\n...\nint\nmain ()\n{\n  int hsize = sizeof (junk) / 2;\n  if (memcmp (junk, junk + hsize, hsize))\n    return do_program_2 ();\n  else\n    return do_program_1 ();\n}\n"},{"id":"4898","messageId":"20050612145342.GA9882@macavity","threadId":"914","inReplyTo":"20050612082555.GB6620@pasky.ji.cz","subject":"Re: [zooko@zooko.com: [Revctrl] colliding md5 hashes of human-meaningful documents]","fromName":"Martin Uecker","fromEmail":"muecker@gmx.de","sentAt":"2005-06-12T14:53:42Z","receivedAt":"2005-06-12T14:53:42Z","isPatch":false,"sender":{"key":"muecker@gmx.de","avatar":null},"body":"On Sun, Jun 12, 2005 at 10:25:55AM +0200, Petr Baudis wrote:\n> ----- Forwarded message from zooko@zooko.com -----\n> \n> There is nothing theoretically surprising about this, but hopefully its\n> concreteness and the accompanying scenario will make an impression on people\n> on people.  The same technique should work to generate two documents with\n> identical SHA1 hashes.\n> \n> http://www.cits.rub.de/MD5Collisions/\n> \n> ----- End forwarded message -----\n> \n> I expected the two postscript files differing in some huge binary blob,\n> but it turns out the binary part is very small (about 256 bytes) and\n> only few (about nine) bytes are different, contrary to how people have\n> predicted the collisions. This is much more close to finding a collision\n> between similar pure C files, I think. Rather unsettling.\n> \n\nThis attack scenario doesn't demonstrate the danger of hash\ncollisions but the danger of signing documents you do not\nunderstand. The same technique works exactly in the same way\nwith postscript files which are actually identical but produce\ndifferent output under different conditions (time, fonts\ninstalled on the printer whatever).\n\nNever sign anything but plain text or documents which are\ncreated in a controlled way and avoid signing documents\nyou did not create yourself.\n\n\nMartin\n\n-- \nOne night, when little Giana from Milano was fast asleep,\nshe had a strange dream.\n\n"},{"id":"4901","messageId":"Pine.LNX.4.58.0506120949150.2286@ppc970.osdl.org","threadId":"914","inReplyTo":"20050612082555.GB6620@pasky.ji.cz","subject":"Re: [zooko@zooko.com: [Revctrl] colliding md5 hashes of human-meaningful documents]","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-06-12T17:03:28Z","receivedAt":"2005-06-12T17:03:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 12 Jun 2005, Petr Baudis wrote:\n>\n> I expected the two postscript files differing in some huge binary blob,\n> but it turns out the binary part is very small (about 256 bytes) and\n> only few (about nine) bytes are different, contrary to how people have\n> predicted the collisions. This is much more close to finding a collision\n> between similar pure C files, I think. Rather unsettling.\n\nThis is not close at all. The \"small\" binary blob (256 bytes) only encodes \none single bit of information.\n\nIn other words, they've really changed _one_ bit of information by doing a\n256-byte random binary blob. Anybody who calls that \"small\" didn't really\nlook closely.\n\nIs it clever? Yes. But it isn't about making one C file look like another,\nit's using the property of controlling _both_ of the files, and making\nthem contain all the information, and then making the the single-bit\nchange collapse the output into two different modes by using a postscript\ninterpreter to make it print out the same.\n\nIs it a real problem? Yes, because a _lot_ of document formats are\nstructured and are amenable to things like this. But the problem here is\nthe fact that you can fool somebody into signing something without\nrealizing that it has a lot of hidden information thanks to having formats\nthat can hide the blobs.\n\nSo the problem is totally different from the way git uses a hash. In the \ngit model, an attacker by definition cannot control both versions of a \nfile, since if he controls just _one_ version, he doesn't need to do the \nattack in the first place!\n\nPut another way: you could use this exact example for a version of git\nthat uses md5-sums instead of sha1's, but it wouldn't show anything at all \nabout a git vulnerability even so.\n\nThe one thing it does show is that you should probably never sign anything \nbut a nice human-readable ASCII file that you actually opened in your own \neditor.\n\n\t\tLinus\n"},{"id":"4935","messageId":"Pine.LNX.4.21.0506132141250.30848-100000@iabervon.org","threadId":"914","inReplyTo":"Pine.LNX.4.58.0506120949150.2286@ppc970.osdl.org","subject":"Re: [zooko@zooko.com: [Revctrl] colliding md5 hashes of human-meaningful documents]","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2005-06-14T02:06:20Z","receivedAt":"2005-06-14T02:06:20Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Sun, 12 Jun 2005, Linus Torvalds wrote:\n\n> Put another way: you could use this exact example for a version of git\n> that uses md5-sums instead of sha1's, but it wouldn't show anything at all \n> about a git vulnerability even so.\n\nYou couldn't use this exact example for an md5 git; git compresses the\nfiles before hashing, which means that you don't have an md5 block of\narbitrary data you can replace with a different arbitrary block because it\nwouldn't decompress.\n\nOf course, if zlib has a way of saying, \"if bytes 256-511 match 512-767,\ndecompress the first of the two records starting at 768, otherwise\ndecompress the second\" then the attack would work, and we should all by\nworried (and disturbed by zlib in general). Chances are that it would be\nimpractical to find a pair of blocks such that they are both valid in the\nsame part of a zlib record and both leave the compression context such \nthat the same remaining content decompresses successfully and both have\nthe same md5 hash, let alone getting the results in both cases to be valid\nC that depends on the difference between the blocks. It's possible that\nyou could get it to work with only a moderately large number of weak\ncollisions between very similar blocks, but it's not nearly so easy a\ntask.\n\n\t-Daniel\n*This .sig left intentionally blank*\n\n\n"}]}