{"thread":{"id":"13309","subject":"About git and the use of SHA-1","startedAt":"2008-04-28T16:29:07Z","lastAt":"2008-04-30T05:56:58Z","messageCount":38,"participants":["Henrik Austad","Daniel Barkalow","Andreas Ericsson","Russ Dill","Sverre Rabbelier","Dmitry Potapov","Jurko Gospodnetić","Paolo Bonzini","Tom Widmer","Geoffrey Irving","Nicolas Pitre","Matthieu Moy","Fredrik Skolmli","Martin Langhoff","David Brown"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"75409","messageId":"200804281829.11866.henrikau@orakel.ntnu.no","threadId":"13309","inReplyTo":null,"subject":"About git and the use of SHA-1","fromName":"Henrik Austad","fromEmail":"henrikau@orakel.ntnu.no","sentAt":"2008-04-28T16:29:07Z","receivedAt":"2008-04-28T16:29:07Z","isPatch":false,"sender":{"key":"henrikau@orakel.ntnu.no","avatar":null},"body":"Hi list!\n\nAs far as I have gathered, the SHA-1-sum is used as a identifier for commits, \nand that is the primary reason for using sha1.  However, several places \n(including the google tech-talk featuring Linus himself) states that the id's \nare cryptographically secure.\n\nAs discussed in [1], SHA-1 is not as secure as it once was (and this was in \n2005), and I'm wondering - are there any plans for migrating to another \nhash-algorithm? I.e. SHA-2, whirlpool..\n\n[1] http://www.schneier.com/blog/archives/2005/02/cryptanalysis_o.html\n-- \nmvh Henrik Austad\n"},{"id":"75440","messageId":"alpine.LNX.1.00.0804281515480.19665@iabervon.org","threadId":"13309","inReplyTo":"200804281829.11866.henrikau@orakel.ntnu.no","subject":"Re: About git and the use of SHA-1","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-04-28T19:34:50Z","receivedAt":"2008-04-28T19:34:50Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Mon, 28 Apr 2008, Henrik Austad wrote:\n\n> Hi list!\n> \n> As far as I have gathered, the SHA-1-sum is used as a identifier for commits, \n> and that is the primary reason for using sha1.  However, several places \n> (including the google tech-talk featuring Linus himself) states that the id's \n> are cryptographically secure.\n> \n> As discussed in [1], SHA-1 is not as secure as it once was (and this was in \n> 2005), and I'm wondering - are there any plans for migrating to another \n> hash-algorithm? I.e. SHA-2, whirlpool..\n\nNo. The cryptographic security we care about is that it's impractical to \ncome up with another set of content that hashes to the same value as a \ngiven set of content. The known attacks on SHA-1 (and more broken earlier \nhashes in the same general class) only allow the attacker to produce two \nfiles that will collide. Now, it's true that this would allow somebody to \nproduce a commit where some people see the \"good\" blob and some people see \nthe \"evil\" blob, but (a) the \"good\" blob contains some large chunk of \nrandom data, which is a major red flag by itself, and (b) all of these \npeople have to be taking data from the attacker.\n\nIf somebody gives you some source, and it's got some large random chunk in \nit, and the behavior of the object depends on the content of this chunk, \nand it's unspecified where this chunk comes from, you should be aware \nthat they might be able to swap this chunk for a different chunk. But such \na file is pretty blatantly malicious anyway.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"75464","messageId":"200804282329.21336.henrikau@orakel.ntnu.no","threadId":"13309","inReplyTo":"alpine.LNX.1.00.0804281515480.19665@iabervon.org","subject":"Re: About git and the use of SHA-1","fromName":"Henrik Austad","fromEmail":"henrikau@orakel.ntnu.no","sentAt":"2008-04-28T21:29:14Z","receivedAt":"2008-04-28T21:29:14Z","isPatch":false,"sender":{"key":"henrikau@orakel.ntnu.no","avatar":null},"body":"On Monday 28 April 2008 21:34:50 Daniel Barkalow wrote:\n> On Mon, 28 Apr 2008, Henrik Austad wrote:\n> > Hi list!\n> >\n> > As far as I have gathered, the SHA-1-sum is used as a identifier for\n> > commits, and that is the primary reason for using sha1.  However, several\n> > places (including the google tech-talk featuring Linus himself) states\n> > that the id's are cryptographically secure.\n> >\n> > As discussed in [1], SHA-1 is not as secure as it once was (and this was\n> > in 2005), and I'm wondering - are there any plans for migrating to\n> > another hash-algorithm? I.e. SHA-2, whirlpool..\n>\n> No. The cryptographic security we care about is that it's impractical to\n> come up with another set of content that hashes to the same value as a\n> given set of content. The known attacks on SHA-1 (and more broken earlier\n> hashes in the same general class) only allow the attacker to produce two\n> files that will collide. Now, it's true that this would allow somebody to\n> produce a commit where some people see the \"good\" blob and some people see\n> the \"evil\" blob, but (a) the \"good\" blob contains some large chunk of\n> random data, which is a major red flag by itself, and (b) all of these\n> people have to be taking data from the attacker.\n\nyes, I can see that point, but I was thinking more along the line of:\n\n1) clone repo\n2) add malicious code\n3) add a huge block of comment, ifdef-block etc somewhere obscure in the code \nand keep adding random data untill hash matches a well-known release.\n4) publish repo, or even worse, change central repo\n\nMost users, and probably a lot of developers never browse through the *entire* \narchive looking for this, and as long as the hash checks out - why would you? \nYes, it would probably be discovered soon enough, but take the linux kernel \nas an example - if you get, say 100 infected machines due to this, what would \nthis do to the reputation of the kernel?\n\n\n> If somebody gives you some source, and it's got some large random chunk in\n> it, and the behavior of the object depends on the content of this chunk,\n> and it's unspecified where this chunk comes from, you should be aware\n> that they might be able to swap this chunk for a different chunk. But such\n> a file is pretty blatantly malicious anyway.\n\nTrue, but this actually means you have to verify *everything*, even though the \nhash checks out.\n\nbut yes, I can see your point, and it would most likely be infeasible to \ngenerate a collision using this approach, and changing to another \nhashfunction would probably not add much. basically I was just curious and \nplayed ahead with the idea.\n\nThanks for the answer though :)\n-- \nmvh Henrik Austad\n"},{"id":"75473","messageId":"alpine.LNX.1.00.0804281732370.19665@iabervon.org","threadId":"13309","inReplyTo":"200804282329.21336.henrikau@orakel.ntnu.no","subject":"Re: About git and the use of SHA-1","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-04-28T22:15:16Z","receivedAt":"2008-04-28T22:15:16Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Mon, 28 Apr 2008, Henrik Austad wrote:\n\n> On Monday 28 April 2008 21:34:50 Daniel Barkalow wrote:\n> > On Mon, 28 Apr 2008, Henrik Austad wrote:\n> > > Hi list!\n> > >\n> > > As far as I have gathered, the SHA-1-sum is used as a identifier for\n> > > commits, and that is the primary reason for using sha1.  However, several\n> > > places (including the google tech-talk featuring Linus himself) states\n> > > that the id's are cryptographically secure.\n> > >\n> > > As discussed in [1], SHA-1 is not as secure as it once was (and this was\n> > > in 2005), and I'm wondering - are there any plans for migrating to\n> > > another hash-algorithm? I.e. SHA-2, whirlpool..\n> >\n> > No. The cryptographic security we care about is that it's impractical to\n> > come up with another set of content that hashes to the same value as a\n> > given set of content. The known attacks on SHA-1 (and more broken earlier\n> > hashes in the same general class) only allow the attacker to produce two\n> > files that will collide. Now, it's true that this would allow somebody to\n> > produce a commit where some people see the \"good\" blob and some people see\n> > the \"evil\" blob, but (a) the \"good\" blob contains some large chunk of\n> > random data, which is a major red flag by itself, and (b) all of these\n> > people have to be taking data from the attacker.\n> \n> yes, I can see that point, but I was thinking more along the line of:\n> \n> 1) clone repo\n> 2) add malicious code\n> 3) add a huge block of comment, ifdef-block etc somewhere obscure in the code \n> and keep adding random data untill hash matches a well-known release.\n> 4) publish repo, or even worse, change central repo\n\nAll known methods for step 3, even on hashes considered long broken, will \ntake until the heat death of the universe. The latest I can find is that, \nif you use MD4 (which is weak enough that you can find collisions as \nquickly as you can do two hashes), there's a 1 in a quadrillion chance \nthat your message is weak and somebody could find a replacement with the \nsame hash using known techniques. (With a plausible amount of work, an \nattacker could take a file and modify it only slightly, and find a \nreplacement for that, but this again requires the attacker to have some \nnon-trivial input to what gets put in the official tree, which leaves \nthe attacker as the responsible party for that object).\n\nSHA-1 is enough stronger that the latest attacks are still unable to do \nwith the current available computing power in years what can be done to \nMD4 in milliseconds. So it's highly unlikely that somebody will break \nSHA-1 more thoroughly than MD4 is broken any time soon.\n\n> > If somebody gives you some source, and it's got some large random chunk in\n> > it, and the behavior of the object depends on the content of this chunk,\n> > and it's unspecified where this chunk comes from, you should be aware\n> > that they might be able to swap this chunk for a different chunk. But such\n> > a file is pretty blatantly malicious anyway.\n> \n> True, but this actually means you have to verify *everything*, even though the \n> hash checks out.\n\nIf you don't verify *everything* when the hash checks out, the attacker \nwill just send you a properly-constructed commit with a back door in the \ncode. While you're looking for directly-inserted security holes in the \ncode, you can probably notice if there's some big hunk of line noise in a \ncomment that might make the file vulnerable to replacement.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"75496","messageId":"4816C26D.9010304@op5.se","threadId":"13309","inReplyTo":"200804282329.21336.henrikau@orakel.ntnu.no","subject":"Re: About git and the use of SHA-1","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-04-29T06:38:37Z","receivedAt":"2008-04-29T06:38:37Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Henrik Austad wrote:\n> On Monday 28 April 2008 21:34:50 Daniel Barkalow wrote:\n>> On Mon, 28 Apr 2008, Henrik Austad wrote:\n>>> Hi list!\n>>>\n>>> As far as I have gathered, the SHA-1-sum is used as a identifier for\n>>> commits, and that is the primary reason for using sha1.  However, several\n>>> places (including the google tech-talk featuring Linus himself) states\n>>> that the id's are cryptographically secure.\n>>>\n>>> As discussed in [1], SHA-1 is not as secure as it once was (and this was\n>>> in 2005), and I'm wondering - are there any plans for migrating to\n>>> another hash-algorithm? I.e. SHA-2, whirlpool..\n>> No. The cryptographic security we care about is that it's impractical to\n>> come up with another set of content that hashes to the same value as a\n>> given set of content. The known attacks on SHA-1 (and more broken earlier\n>> hashes in the same general class) only allow the attacker to produce two\n>> files that will collide. Now, it's true that this would allow somebody to\n>> produce a commit where some people see the \"good\" blob and some people see\n>> the \"evil\" blob, but (a) the \"good\" blob contains some large chunk of\n>> random data, which is a major red flag by itself, and (b) all of these\n>> people have to be taking data from the attacker.\n> \n> yes, I can see that point, but I was thinking more along the line of:\n> \n> 1) clone repo\n> 2) add malicious code\n> 3) add a huge block of comment, ifdef-block etc somewhere obscure in the code \n> and keep adding random data untill hash matches a well-known release.\n> 4) publish repo, or even worse, change central repo\n> \n\nThis depends greatly on git accepting objects with a colliding object-name,\nwhich it doesn't. Once you have an object with a particular SHA1, it will\nnever get overwritten, ever, as git will believe it's about to do unnecessary\nwork. As such, you'd still have to create a new object, hashing to a new SHA1\nand get that new object added to the kernel.\n\nI think perhaps Andrew Morton and a few other \"high brass\" among the kernel\nhackers can get away with pushing crud like that to Linus' public tree\n(which is the de facto master copy of published kernel sources), but random\nJohn Doe's such as you and me wouldn't stand a chance, as our patches would\nget reviewed by someone who, at the end of the day, makes a living coding\nLinux.\n\n\n> Most users, and probably a lot of developers never browse through the *entire* \n> archive looking for this, and as long as the hash checks out - why would you? \n> Yes, it would probably be discovered soon enough, but take the linux kernel \n> as an example - if you get, say 100 infected machines due to this, what would \n> this do to the reputation of the kernel?\n> \n\nThat depends. If the source of it was Linus' public tree, that would not be\nvery good at all. If the source was a random tarball off a random webpage\nor ftp site (which would be the same as fetching and, unverified, using an\nunchecked git repository), I doubt it would matter much.\n\n> \n>> If somebody gives you some source, and it's got some large random chunk in\n>> it, and the behavior of the object depends on the content of this chunk,\n>> and it's unspecified where this chunk comes from, you should be aware\n>> that they might be able to swap this chunk for a different chunk. But such\n>> a file is pretty blatantly malicious anyway.\n> \n> True, but this actually means you have to verify *everything*, even though the \n> hash checks out.\n> \n\nNot really. What you need to verify is that\na) You cloned from somewhere you trust (kernel.org, fe)\nb) The SHA1 of the commit you want to build from matches the SHA1 of the same\ncommit in the repository you originally cloned from.\n\nColliding objects can never enter a repository. Git is lazy and will reuse the\nalready existing colliding object with the same name instead.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"75502","messageId":"f9d2a5e10804290009p17d291d5wf14e2bb58bedca63@mail.gmail.com","threadId":"13309","inReplyTo":"4816C26D.9010304@op5.se","subject":"Re: About git and the use of SHA-1","fromName":"Russ Dill","fromEmail":"russ.dill@gmail.com","sentAt":"2008-04-29T07:09:24Z","receivedAt":"2008-04-29T07:09:24Z","isPatch":false,"sender":{"key":"russ.dill@gmail.com","avatar":"https://gravatar.com/avatar/989b24f3fa63126a35d7c74069e2626e715a3b1a86616be9962f7ceae04ff9c5?d=mp&s=160"},"body":">  Colliding objects can never enter a repository. Git is lazy and will reuse the\n>  already existing colliding object with the same name instead.\n>\n\nI think you are missing the point. One of the pluses behind originally\nusing SHA-1 and the signed tags is that the system as a whole is\ncryptographically secure. You can verify from the public key of\nwhoever made the tag that yes, this really is the source and history\nthey tagged. Not only can DNS attacks be made, fooling users into\nthinking that they are really connecting to kernel.org, or whatever\nelse server they expect to be connecting to, but also, the server\nitself may be hacked and objects replaced.\n\nI'm just not sure how much time it would take to find a collision.\n"},{"id":"75507","messageId":"4816CC80.9080705@op5.se","threadId":"13309","inReplyTo":"f9d2a5e10804290009p17d291d5wf14e2bb58bedca63@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-04-29T07:21:36Z","receivedAt":"2008-04-29T07:21:36Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Russ Dill wrote:\n>>  Colliding objects can never enter a repository. Git is lazy and will reuse the\n>>  already existing colliding object with the same name instead.\n>>\n> \n> I think you are missing the point. One of the pluses behind originally\n> using SHA-1 and the signed tags is that the system as a whole is\n> cryptographically secure. You can verify from the public key of\n> whoever made the tag that yes, this really is the source and history\n> they tagged. Not only can DNS attacks be made, fooling users into\n> thinking that they are really connecting to kernel.org, or whatever\n> else server they expect to be connecting to, but also, the server\n> itself may be hacked and objects replaced.\n> \n\nIf the server is hacked and objects are replaced, they will either\nno longer match their cryptographic signature, meaning they'll be\nnew objects or git will determine that they are corrupt, or they\n*will* match an existing object, but then that object won't be\npropagated to other repositories since git refuses to overwrite\nalready existing objects. Either way, gits refusal to overwrite\nobjects it already has plays a part in making malicious actions\nfutile, since malicious code is only worth something if it's\npropagated and actually used.\n\n> I'm just not sure how much time it would take to find a collision.\n\nEven crypto-experts are arguing about that, so I'm not surprised.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"75520","messageId":"bd6139dc0804290405w4a7a94a7s15a85285b2122f2f@mail.gmail.com","threadId":"13309","inReplyTo":"4816CC80.9080705@op5.se","subject":"Re: About git and the use of SHA-1","fromName":"Sverre Rabbelier","fromEmail":"alturin@gmail.com","sentAt":"2008-04-29T11:05:46Z","receivedAt":"2008-04-29T11:05:46Z","isPatch":false,"sender":{"key":"alturin@gmail.com","avatar":null},"body":"On Tue, Apr 29, 2008 at 9:21 AM, Andreas Ericsson <ae@op5.se> wrote:\n> Russ Dill wrote:\n>  If the server is hacked and objects are replaced, they will either\n>  no longer match their cryptographic signature, meaning they'll be\n>  new objects or git will determine that they are corrupt, or they\n\nWe were assuming here that once SHA-1 is broken really determined\nhackers will be able to come up with objects that -do- match the\nSHA-1, so the above is not relevant.\n\n>  *will* match an existing object, but then that object won't be\n>  propagated to other repositories since git refuses to overwrite\n>  already existing objects. [...]\n\nWhat about new users cloning the repo? They're just out of luck? I\ndon't think this argument holds, if we want to 'advertise' that git is\ncryptographically secure we can do so only as long as our hashing\nalgorithm is. (As such, should SHA-1 ever be fully broken we'd need to\neither switch to another algorithm or stop advertising being\ncryptographically secure.)\n\n>  [...] Either way, gits refusal to overwrite\n>  objects it already has plays a part in making malicious actions\n>  futile, since malicious code is only worth something if it's\n>  propagated and actually used.\n\nOf course this is true, it makes it a lot harder to do damage, but it\ndoesn't eliminate the problem, it's just a free 'extra protection'.\nYes, malicious code is only worth something if it's propagated and\nactually used, no, it is not impossible to do so in git if/when SHA-1\nturns out to have collisions every other file.\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"75522","messageId":"48171442.4050707@op5.se","threadId":"13309","inReplyTo":"bd6139dc0804290405w4a7a94a7s15a85285b2122f2f@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-04-29T12:27:46Z","receivedAt":"2008-04-29T12:27:46Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Sverre Rabbelier wrote:\n> On Tue, Apr 29, 2008 at 9:21 AM, Andreas Ericsson <ae@op5.se> wrote:\n>> Russ Dill wrote:\n>>  If the server is hacked and objects are replaced, they will either\n>>  no longer match their cryptographic signature, meaning they'll be\n>>  new objects or git will determine that they are corrupt, or they\n> \n> We were assuming here that once SHA-1 is broken really determined\n> hackers will be able to come up with objects that -do- match the\n> SHA-1, so the above is not relevant.\n> \n>>  *will* match an existing object, but then that object won't be\n>>  propagated to other repositories since git refuses to overwrite\n>>  already existing objects. [...]\n> \n> What about new users cloning the repo? They're just out of luck?\n\nOnly until someone who's already cloned the repository fetches\nfrom it, at which point the collision will be detected.\n\n> I\n> don't think this argument holds, if we want to 'advertise' that git is\n> cryptographically secure we can do so only as long as our hashing\n> algorithm is. (As such, should SHA-1 ever be fully broken we'd need to\n> either switch to another algorithm or stop advertising being\n> cryptographically secure.)\n> \n\nTrue. So far though, the only attacks that have been successful requires\nthat the attacker is allowed to create both the colliding data-sets,\nand so far none has been found that would allow the attacker to follow\nany kind of syntactical rules what so ever, so from a practical point\nof view, SHA1 is 100% secure *for sourcecode*.\n\n>From a theoretical point of view, no hash is 100% secure, so changing\nalgorithm buys us nothing.\n\nBesides, \"cryprographically secure\" is not the same as \"will never ever\nbe broken\", because all hashes are obviously susceptible to brute-force\nattacks. \"Cryptographically secure\" means, insofar as I've understood it\nthat given a source-file and a key, it would take such an extremely\nlong time to find a different data-set that hashes to the same key that\nthe result is unusable because the original source is obsolete.\n\nThat is why legal documents are always signed with the \"most secure\"\n(or rather, \"least insecure\") of all available hashes. For our\npurposes, SHA1 suffices until someone comes up with a relatively\ntrivial way of creating a collision within the parameters above.\n\n\n>>  [...] Either way, gits refusal to overwrite\n>>  objects it already has plays a part in making malicious actions\n>>  futile, since malicious code is only worth something if it's\n>>  propagated and actually used.\n> \n> Of course this is true, it makes it a lot harder to do damage, but it\n> doesn't eliminate the problem, it's just a free 'extra protection'.\n> Yes, malicious code is only worth something if it's propagated and\n> actually used, no, it is not impossible to do so in git if/when SHA-1\n> turns out to have collisions every other file.\n> \n\nPoints of fact so far:\n* It possible to create objects with colliding names (SHA1 hash keys).\n  This holds true whichever algorithm we use, although it will be more\n  difficult with a stronger algorithm.\n* It is impossible to distribute the colliding content to already cloned\n  repositories. This also holds true for all hash algorithms.\n\nI've been arguing that the value of the first point is so greatly\ndiminished by the second, that even if SHA1 turns out to be horribly\nbroken, projects using git will still have a decent protection against\nmalicious code entering the repository without the knowledge of one of\nthe authors.\n\nYou've been arguing that SHA1 is not theoretically secure, which is\nobviously true since no hash is theoretically secure.\n\nI can think of one way to make git a lot more resilient to hash\ncollisions, regardless of which hash is used, namely: Add the length\nof the hashed object to the hash.\n\nIn order for an evil-minded hacker to succeed in doing any real harm,\nhe/she now has to create a conflicting file which is valid for its\ntype (be it C, PHP, JPEG, AVI, PDF or whatever) and is also the same\nlength as the original source, without being allowed to create the\noriginal object.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"75523","messageId":"20080429124152.GB6160@dpotapov.dyndns.org","threadId":"13309","inReplyTo":"200804281829.11866.henrikau@orakel.ntnu.no","subject":"Re: About git and the use of SHA-1","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-04-29T12:41:52Z","receivedAt":"2008-04-29T12:41:52Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Mon, Apr 28, 2008 at 06:29:07PM +0200, Henrik Austad wrote:\n> \n> As discussed in [1], SHA-1 is not as secure as it once was (and this was in \n> 2005), and I'm wondering - are there any plans for migrating to another \n> hash-algorithm? I.e. SHA-2, whirlpool..\n\nSHA-1 is broken in the sense that it requires computation less than\nfinding a collision  by brute force (2^80). It is still very costly and\nAFAIK no one yet has found a single collision for SHA-1 yet, but even if\nsuch a collision is found, the question is how it can be exploit?\n\nThis collision cannot be used to replace any existing code in Git. The\nonly way to exploit this collision is to submit a patch based on one\nsequence to the maintainer and it should look legitimate to be accepted\nand then create another blob with malicious code based on the other\nsequence, so the second blob has the same SHA-1 then anyone who pulls\nfrom you will get malicious code.\n\nHowever, it is tricky to create these two blobs -- one which should pass\ninspection and look like as a real improvement but the other one that\nshould do what you want. All what you have is two sequences of 20 bytes\nwith the same SHA-1 and you have no control over them. For some binary\nfiles, it is possible by including both good and bad contents in the\nsubmitted blob and using one sequence in the right place to hide the bad\npart and make only the good one active/visible. Then the other blob will\nbe almost the same but contains the other sequence, which is used to\nactivate the bad part. This can work if the maintainer cannot see\neverything but only the \"visible\" part. However, I don't think you can\ndo anything like that with _source_ code, which is inspect. And if\nsubmitted code is not reviewed, there is nothing that can protect you\nfrom malicious code getting into the repository (and even worse it will\nget directly into the official repository!).\n\nSo, I don't think we have to worry much about possibility a collision\nattack, but only about preimage attacks; and a preimage attack on SHA-1\nis far away from reality.\n\nDmitry\n"},{"id":"75524","messageId":"481718AF.8090000@docte.hr","threadId":"13309","inReplyTo":"f9d2a5e10804290009p17d291d5wf14e2bb58bedca63@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Jurko Gospodnetić","fromEmail":"jurko.gospodnetic@docte.hr","sentAt":"2008-04-29T12:46:39Z","receivedAt":"2008-04-29T12:46:39Z","isPatch":false,"sender":{"key":"jurko.gospodnetic@docte.hr","avatar":null},"body":"> I think you are missing the point. One of the pluses behind originally\n> using SHA-1 and the signed tags is that the system as a whole is\n> cryptographically secure. You can verify from the public key of\n> whoever made the tag that yes, this really is the source and history\n> they tagged.\n\n   I am not really sure I follow this.... how can you 'verify from the \npublic key of whoever made the tag' that the SHA-1 hash is correct!? \nSHA-1 does not have anything do with any externally provided keys or \nhave I managed to get something confused here?\n\n   Best regards,\n     Jurko Gospodnetić\n"},{"id":"75525","messageId":"48171D24.9000104@gnu.org","threadId":"13309","inReplyTo":"48171442.4050707@op5.se","subject":"Re: About git and the use of SHA-1","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-04-29T13:05:40Z","receivedAt":"2008-04-29T13:05:40Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n> I can think of one way to make git a lot more resilient to hash\n> collisions, regardless of which hash is used, namely: Add the length\n> of the hashed object to the hash.\n\nNot really, because most attacks are about collisions, not second \npreimages.  They produce two 64-byte blocks (hence, same length) with \nthe same hash value.\n\nAs such, they allow to change a blob that *the attacker* injected in the \nrepository.  The way the more \"spectacular\" attacks are devised requires \na \"language\" with conditional expressions -- for documents, for example, \nPostscript is used.  If you prepare a postscript file whose code is\n\n    if (AAAA == BBBB)\n      typeset document 1\n    else\n      typeset document 2\n\nwhere AAAA and BBBB are collisions, and you change it to \"if (BBBB == \nBBBB) the hash will be the same, but the outcome will be document 1 \ninstead of document 2.\n\nThe fact that this requires having the two \"behaviors\" in the blob is \nnot a big deal for source code, going in the wrong branch of an \"if\" can \nbe an attack.  On the other hand, it makes adding the length useless for \ncollision attacks.  True, it wouldn't be useless for second preimage \nattacks, but SHA-1 is still secure with respect to those.\n\nPaolo\n"},{"id":"75528","messageId":"481732C0.5020208@op5.se","threadId":"13309","inReplyTo":"48171D24.9000104@gnu.org","subject":"Re: About git and the use of SHA-1","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-04-29T14:37:52Z","receivedAt":"2008-04-29T14:37:52Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Paolo Bonzini wrote:\n> \n>> I can think of one way to make git a lot more resilient to hash\n>> collisions, regardless of which hash is used, namely: Add the length\n>> of the hashed object to the hash.\n> \n> Not really, because most attacks are about collisions, not second \n> preimages.  They produce two 64-byte blocks (hence, same length) with \n> the same hash value.\n> \n> As such, they allow to change a blob that *the attacker* injected in the \n> repository.  The way the more \"spectacular\" attacks are devised requires \n> a \"language\" with conditional expressions -- for documents, for example, \n> Postscript is used.  If you prepare a postscript file whose code is\n> \n>    if (AAAA == BBBB)\n>      typeset document 1\n>    else\n>      typeset document 2\n> \n> where AAAA and BBBB are collisions, and you change it to \"if (BBBB == \n> BBBB) the hash will be the same, but the outcome will be document 1 \n> instead of document 2.\n> \n> The fact that this requires having the two \"behaviors\" in the blob is \n> not a big deal for source code, going in the wrong branch of an \"if\" can \n> be an attack.  On the other hand, it makes adding the length useless for \n> collision attacks.  True, it wouldn't be useless for second preimage \n> attacks, but SHA-1 is still secure with respect to those.\n> \n\nSo what you're saying is that if someone owns a repository and adds a\nfile to it, he can then replace his entire repository with an identical\none where the good file is replaced with a bad one, and this will affect\npeople who clone *after* the file gets replaced.\n\nGee, that's one fiendishly large attack vector, quite apart from the\nfact that said author first has to come up with a program that gets\nwidespread enough that a lot of people all of a sudden wants to use\nit, but not so widespread that anyone would want to review it before\nusing it.\n\nI remain unconvinced as to whether or not SHA1 is, for all practical\npurposes, cryptographically secure for git's uses. Sure, evil programmers\ncan screw you over if you use their software without reviewing it, but\nthat's hardly due to git using a particular cryptographic algorithm.\n\nOtoh, I'm not familiar enough with the nomenclature to say with 100%\ncertainty what's cryprographically secure and what isn't. I just know\nthat there are no collision-less hashes, so whatever \"cryptographically\nsecure\" really means wrt hashes, \"100% collision-free\" isn't it.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"75529","messageId":"481733A3.4010802@op5.se","threadId":"13309","inReplyTo":"20080429124152.GB6160@dpotapov.dyndns.org","subject":"Re: About git and the use of SHA-1","fromName":"Andreas Ericsson","fromEmail":"ae@op5.se","sentAt":"2008-04-29T14:41:39Z","receivedAt":"2008-04-29T14:41:39Z","isPatch":false,"sender":{"key":"ae@op5.se","avatar":"https://gravatar.com/avatar/426e89595c75a8f5252dd0c989e5fabe5bcac616e68557427ad9aef6b0ca342a?d=mp&s=160"},"body":"Dmitry Potapov wrote:\n> On Mon, Apr 28, 2008 at 06:29:07PM +0200, Henrik Austad wrote:\n>> As discussed in [1], SHA-1 is not as secure as it once was (and this was in \n>> 2005), and I'm wondering - are there any plans for migrating to another \n>> hash-algorithm? I.e. SHA-2, whirlpool..\n> \n> SHA-1 is broken in the sense that it requires computation less than\n> finding a collision  by brute force (2^80). It is still very costly and\n> AFAIK no one yet has found a single collision for SHA-1 yet, but even if\n> such a collision is found, the question is how it can be exploit?\n> \n> This collision cannot be used to replace any existing code in Git. The\n> only way to exploit this collision is to submit a patch based on one\n> sequence to the maintainer and it should look legitimate to be accepted\n> and then create another blob with malicious code based on the other\n> sequence, so the second blob has the same SHA-1 then anyone who pulls\n> from you will get malicious code.\n> \n\nBut they won't, because it's impossible to add two objects with the same\nSHA1 hash key to a git repository, since it will lazily re-use the\nexisting one. In practice, this means that in the case of an \"innocent\"\nhash-collision, git will actually break by refusing to store the new\ncontent.\n\n> However, it is tricky to create these two blobs -- one which should pass\n> inspection and look like as a real improvement but the other one that\n> should do what you want. All what you have is two sequences of 20 bytes\n> with the same SHA-1 and you have no control over them. For some binary\n> files, it is possible by including both good and bad contents in the\n> submitted blob and using one sequence in the right place to hide the bad\n> part and make only the good one active/visible. Then the other blob will\n> be almost the same but contains the other sequence, which is used to\n> activate the bad part. This can work if the maintainer cannot see\n> everything but only the \"visible\" part. However, I don't think you can\n> do anything like that with _source_ code, which is inspect. And if\n> submitted code is not reviewed, there is nothing that can protect you\n> from malicious code getting into the repository (and even worse it will\n> get directly into the official repository!).\n> \n> So, I don't think we have to worry much about possibility a collision\n> attack, but only about preimage attacks; and a preimage attack on SHA-1\n> is far away from reality.\n> \n\nRight.\n\n-- \nAndreas Ericsson                   andreas.ericsson@op5.se\nOP5 AB                             www.op5.se\nTel: +46 8-230225                  Fax: +46 8-230231\n"},{"id":"75530","messageId":"48173644.1020306@gnu.org","threadId":"13309","inReplyTo":"481732C0.5020208@op5.se","subject":"Re: About git and the use of SHA-1","fromName":"Paolo Bonzini","fromEmail":"bonzini@gnu.org","sentAt":"2008-04-29T14:52:52Z","receivedAt":"2008-04-29T14:52:52Z","isPatch":false,"sender":{"key":"bonzini@gnu.org","avatar":"https://avatars.githubusercontent.com/u/42082?v=4"},"body":"\n> So what you're saying is that if someone owns a repository and adds a\n> file to it, he can then replace his entire repository with an identical\n> one where the good file is replaced with a bad one, and this will affect\n> people who clone *after* the file gets replaced.\n> \n> Gee, that's one fiendishly large attack vector\n\nI agree (with the irony).\n\nPaolo\n"},{"id":"75531","messageId":"fv7db2$5hi$1@ger.gmane.org","threadId":"13309","inReplyTo":"200804281829.11866.henrikau@orakel.ntnu.no","subject":"Re: About git and the use of SHA-1","fromName":"Tom Widmer","fromEmail":"tom.widmer@googlemail.com","sentAt":"2008-04-29T15:02:57Z","receivedAt":"2008-04-29T15:02:57Z","isPatch":false,"sender":{"key":"tom.widmer@googlemail.com","avatar":null},"body":"Henrik Austad wrote:\n> Hi list!\n> \n> As far as I have gathered, the SHA-1-sum is used as a identifier for commits, \n> and that is the primary reason for using sha1.  However, several places \n> (including the google tech-talk featuring Linus himself) states that the id's \n> are cryptographically secure.\n> \n> As discussed in [1], SHA-1 is not as secure as it once was (and this was in \n> 2005), and I'm wondering - are there any plans for migrating to another \n> hash-algorithm? I.e. SHA-2, whirlpool..\n> \n> [1] http://www.schneier.com/blog/archives/2005/02/cryptanalysis_o.html\n\nWhy not wait until the results of:\n\nare available. That will surely be soon enough.\n\nTom\n"},{"id":"75534","messageId":"7f9d599f0804290834v23da6dfbv47b3ca9058934228@mail.gmail.com","threadId":"13309","inReplyTo":"alpine.LNX.1.00.0804281515480.19665@iabervon.org","subject":"Re: About git and the use of SHA-1","fromName":"Geoffrey Irving","fromEmail":"irving@naml.us","sentAt":"2008-04-29T15:34:11Z","receivedAt":"2008-04-29T15:34:11Z","isPatch":false,"sender":{"key":"irving@naml.us","avatar":"https://gravatar.com/avatar/52d7452fcd134aac0fa12f57a3bb7ef5f3f7e73ca0ab36736d06c6a6132de718?d=mp&s=160"},"body":"On Mon, Apr 28, 2008 at 12:34 PM, Daniel Barkalow <barkalow@iabervon.org> wrote:\n> On Mon, 28 Apr 2008, Henrik Austad wrote:\n>\n>  > Hi list!\n>  >\n>  > As far as I have gathered, the SHA-1-sum is used as a identifier for commits,\n>  > and that is the primary reason for using sha1.  However, several places\n>  > (including the google tech-talk featuring Linus himself) states that the id's\n>  > are cryptographically secure.\n>  >\n>  > As discussed in [1], SHA-1 is not as secure as it once was (and this was in\n>  > 2005), and I'm wondering - are there any plans for migrating to another\n>  > hash-algorithm? I.e. SHA-2, whirlpool..\n>\n>  No. The cryptographic security we care about is that it's impractical to\n>  come up with another set of content that hashes to the same value as a\n>  given set of content. The known attacks on SHA-1 (and more broken earlier\n>  hashes in the same general class) only allow the attacker to produce two\n>  files that will collide. Now, it's true that this would allow somebody to\n>  produce a commit where some people see the \"good\" blob and some people see\n>  the \"evil\" blob, but (a) the \"good\" blob contains some large chunk of\n>  random data, which is a major red flag by itself, and (b) all of these\n>  people have to be taking data from the attacker.\n>\n>  If somebody gives you some source, and it's got some large random chunk in\n>  it, and the behavior of the object depends on the content of this chunk,\n>  and it's unspecified where this chunk comes from, you should be aware\n>  that they might be able to swap this chunk for a different chunk. But such\n>  a file is pretty blatantly malicious anyway.\n\nThis argument is invalid, since the use of git is not limited to\nsource code.  People\ncan and do store unreadable binary data in git, and unless you are completely\nsure that no one would ever care about the security of that data in a\nway that can\nbe attacked with a single collision, git should be secure about those as well.\n\nFor example, I just converted a 20 GB repository to git which, among\nother things,\ncontains pdf files of my tax returns.  I have looked them over, but I\nhave not opened\nthem in a hex editor and looked them over at the binary level, and I\ndon't think git\nshould expect me to.\n\nIncidentally, git was the only version control system I tried except\nfor subversion that\ndidn't choke on that repository.  Mercurial looked at my file renames\nand expanded\nthe size past 45 GB before I killed it, I had to fix a several bugs in\nthe bazaar conversion\nscripts before I realized it was just too slow, and svk turns out to\nbe even more like\nthe Antichrist than subversion itself is (mirroring N repository\ncopies requires an N-fold\nincrease in size).\n\nGeoffrey\n"},{"id":"75535","messageId":"alpine.LFD.1.10.0804291132060.23581@xanadu.home","threadId":"13309","inReplyTo":"481733A3.4010802@op5.se","subject":"Re: About git and the use of SHA-1","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-04-29T15:42:38Z","receivedAt":"2008-04-29T15:42:38Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 29 Apr 2008, Andreas Ericsson wrote:\n\n> But they won't, because it's impossible to add two objects with the same\n> SHA1 hash key to a git repository, since it will lazily re-use the\n> existing one. In practice, this means that in the case of an \"innocent\"\n> hash-collision, git will actually break by refusing to store the new\n> content.\n\nI'd also like to point out that Git usually receive \"untrusted\" new \nobjects via the Git protocol through 'git index-pack'.  If you look at \nsha1_object() in index-pack.c, you'll see that active verification \nagainst hash collision is performed, and the fetch will abruptly be \naborted if ever that happens.\n\nYes, writing a test case for this was tricky.  :-)\n\n\nNicolas\n"},{"id":"75537","messageId":"7f9d599f0804290859y6a579302m5db9f7f827b320a4@mail.gmail.com","threadId":"13309","inReplyTo":"alpine.LFD.1.10.0804291132060.23581@xanadu.home","subject":"Re: About git and the use of SHA-1","fromName":"Geoffrey Irving","fromEmail":"irving@naml.us","sentAt":"2008-04-29T15:59:29Z","receivedAt":"2008-04-29T15:59:29Z","isPatch":false,"sender":{"key":"irving@naml.us","avatar":"https://gravatar.com/avatar/52d7452fcd134aac0fa12f57a3bb7ef5f3f7e73ca0ab36736d06c6a6132de718?d=mp&s=160"},"body":"On Tue, Apr 29, 2008 at 8:42 AM, Nicolas Pitre <nico@cam.org> wrote:\n> On Tue, 29 Apr 2008, Andreas Ericsson wrote:\n>\n>  > But they won't, because it's impossible to add two objects with the same\n>  > SHA1 hash key to a git repository, since it will lazily re-use the\n>  > existing one. In practice, this means that in the case of an \"innocent\"\n>  > hash-collision, git will actually break by refusing to store the new\n>  > content.\n>\n>  I'd also like to point out that Git usually receive \"untrusted\" new\n>  objects via the Git protocol through 'git index-pack'.  If you look at\n>  sha1_object() in index-pack.c, you'll see that active verification\n>  against hash collision is performed, and the fetch will abruptly be\n>  aborted if ever that happens.\n>\n>  Yes, writing a test case for this was tricky.  :-)\n\nHere's the standard scenario for a hash collision attack, with\nparties, A, B, and C:\n\n1. C, the malicious one, computes the standard two pdfs with matching\nsha1 hashes.\n2. C sends the valid pdf to B through a git commit, and B signs it with a tag.\n3. C grabs the signature, and then forwards the \"signed\" commit to A,\nbut substitutes the invalid pdf with the same hash.\n\nThe fact that git will check for hash collisions within one repository\nis nice, but it doesn't significantly increase the security of git\nagainst hash collision attacks.\n\nGeoffrey\n"},{"id":"75545","messageId":"f9d2a5e10804290921y5c961fc5g88d718c40a5ff037@mail.gmail.com","threadId":"13309","inReplyTo":"481718AF.8090000@docte.hr","subject":"Re: About git and the use of SHA-1","fromName":"Russ Dill","fromEmail":"russ.dill@gmail.com","sentAt":"2008-04-29T16:21:19Z","receivedAt":"2008-04-29T16:21:19Z","isPatch":false,"sender":{"key":"russ.dill@gmail.com","avatar":"https://gravatar.com/avatar/989b24f3fa63126a35d7c74069e2626e715a3b1a86616be9962f7ceae04ff9c5?d=mp&s=160"},"body":"On Tue, Apr 29, 2008 at 5:46 AM, Jurko Gospodnetić\n<jurko.gospodnetic@docte.hr> wrote:\n>\n> > I think you are missing the point. One of the pluses behind originally\n> > using SHA-1 and the signed tags is that the system as a whole is\n> > cryptographically secure. You can verify from the public key of\n> > whoever made the tag that yes, this really is the source and history\n> > they tagged.\n> >\n>\n>   I am not really sure I follow this.... how can you 'verify from the public\n> key of whoever made the tag' that the SHA-1 hash is correct!? SHA-1 does not\n> have anything do with any externally provided keys or have I managed to get\n> something confused here?\n>\n\nSorry for the confusion, its about using the signed tag and the SHA-1\nof the parent commits, along with their associated trees and blobs to\nverify the source and history. If you can't trust the signed tag, or\nall of the SHA-1's, you can't trust the source and history.\n\nHowever, as many said, I don't think there is any reason to not trust\nSHA-1 is the context of source control.\n"},{"id":"75543","messageId":"f9d2a5e10804290924y127d6a9frd7e883b8ab0f781f@mail.gmail.com","threadId":"13309","inReplyTo":"481732C0.5020208@op5.se","subject":"Re: About git and the use of SHA-1","fromName":"Russ Dill","fromEmail":"russ.dill@gmail.com","sentAt":"2008-04-29T16:24:54Z","receivedAt":"2008-04-29T16:24:54Z","isPatch":false,"sender":{"key":"russ.dill@gmail.com","avatar":"https://gravatar.com/avatar/989b24f3fa63126a35d7c74069e2626e715a3b1a86616be9962f7ceae04ff9c5?d=mp&s=160"},"body":"On Tue, Apr 29, 2008 at 7:37 AM, Andreas Ericsson <ae@op5.se> wrote:\n>\n> Paolo Bonzini wrote:\n>\n> >\n> >\n> > > I can think of one way to make git a lot more resilient to hash\n> > > collisions, regardless of which hash is used, namely: Add the length\n> > > of the hashed object to the hash.\n> > >\n> >\n> > Not really, because most attacks are about collisions, not second\n> preimages.  They produce two 64-byte blocks (hence, same length) with the\n> same hash value.\n> >\n> > As such, they allow to change a blob that *the attacker* injected in the\n> repository.  The way the more \"spectacular\" attacks are devised requires a\n> \"language\" with conditional expressions -- for documents, for example,\n> Postscript is used.  If you prepare a postscript file whose code is\n> >\n> >   if (AAAA == BBBB)\n> >     typeset document 1\n> >   else\n> >     typeset document 2\n> >\n> > where AAAA and BBBB are collisions, and you change it to \"if (BBBB ==\n> BBBB) the hash will be the same, but the outcome will be document 1 instead\n> of document 2.\n> >\n> > The fact that this requires having the two \"behaviors\" in the blob is not\n> a big deal for source code, going in the wrong branch of an \"if\" can be an\n> attack.  On the other hand, it makes adding the length useless for collision\n> attacks.  True, it wouldn't be useless for second preimage attacks, but\n> SHA-1 is still secure with respect to those.\n> >\n> >\n>\n>  So what you're saying is that if someone owns a repository and adds a\n>  file to it, he can then replace his entire repository with an identical\n>  one where the good file is replaced with a bad one, and this will affect\n>  people who clone *after* the file gets replaced.\n>\n\nNo, if someone 0wnz a repository, not owns (Or really, malicious\nmirror owners could be in on it). Either that or some form of\nredirection attack. When you download a tarball, you can check the\nsigned checksum that is downloadable along with it. When you clone a\nrepo, you depend on signed tags.\n"},{"id":"75544","messageId":"alpine.LNX.1.00.0804291204160.19665@iabervon.org","threadId":"13309","inReplyTo":"7f9d599f0804290834v23da6dfbv47b3ca9058934228@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-04-29T16:27:44Z","receivedAt":"2008-04-29T16:27:44Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Tue, 29 Apr 2008, Geoffrey Irving wrote:\n\n> On Mon, Apr 28, 2008 at 12:34 PM, Daniel Barkalow <barkalow@iabervon.org> wrote:\n> > On Mon, 28 Apr 2008, Henrik Austad wrote:\n> >\n> >  > Hi list!\n> >  >\n> >  > As far as I have gathered, the SHA-1-sum is used as a identifier for commits,\n> >  > and that is the primary reason for using sha1.  However, several places\n> >  > (including the google tech-talk featuring Linus himself) states that the id's\n> >  > are cryptographically secure.\n> >  >\n> >  > As discussed in [1], SHA-1 is not as secure as it once was (and this was in\n> >  > 2005), and I'm wondering - are there any plans for migrating to another\n> >  > hash-algorithm? I.e. SHA-2, whirlpool..\n> >\n> >  No. The cryptographic security we care about is that it's impractical to\n> >  come up with another set of content that hashes to the same value as a\n> >  given set of content. The known attacks on SHA-1 (and more broken earlier\n> >  hashes in the same general class) only allow the attacker to produce two\n> >  files that will collide. Now, it's true that this would allow somebody to\n> >  produce a commit where some people see the \"good\" blob and some people see\n> >  the \"evil\" blob, but (a) the \"good\" blob contains some large chunk of\n> >  random data, which is a major red flag by itself, and (b) all of these\n> >  people have to be taking data from the attacker.\n> >\n> >  If somebody gives you some source, and it's got some large random chunk in\n> >  it, and the behavior of the object depends on the content of this chunk,\n> >  and it's unspecified where this chunk comes from, you should be aware\n> >  that they might be able to swap this chunk for a different chunk. But such\n> >  a file is pretty blatantly malicious anyway.\n> \n> This argument is invalid, since the use of git is not limited to\n> source code.  People\n> can and do store unreadable binary data in git, and unless you are completely\n> sure that no one would ever care about the security of that data in a\n> way that can\n> be attacked with a single collision, git should be secure about those as well.\n>\n> For example, I just converted a 20 GB repository to git which, among\n> other things,\n> contains pdf files of my tax returns.  I have looked them over, but I\n> have not opened\n> them in a hex editor and looked them over at the binary level, and I\n> don't think git\n> should expect me to.\n\nIf you haven't looked over your PDFs with a hex editor, you're depending \non the security of the software generating the PDFs and on what you did in \ngenerating them. (Looking at the resulting image alone may be unwise if, \nfor example, you redacted anything.) In any case, on the basis of your \nactions, you may this commit. Now, anyone receiving the repository can, \ndue to the lack of second preimage attacks, be sure that (a) the document \nis as you committed it; or (b) the document is different from what you \ncommitted, but you made the substitution; or (c) the document is different \nfrom what you committed, and you were tricked into committing a document \ncarefully designed by somebody else to be weak. Additionally, it's \ninfeasible to create a document such that forensics after the fact can't \nturn up both the content as originally shown and the content as swapped \nfrom either document.\n\nI'm also not confident that PDFs are, in general, not vulnerable to an \nattack where they rasterize entirely differently depending on \nenvironmental factors (e.g., the document you're signing says something \nentirely different when printed on A4 paper than what it says printed on \nLetter); if so, it doesn't matter much that the document could be \nreplaced, since an attacker could just control the environment and get the \nsame effect.\n\nIn any case, an attacker can't come along later and make a replacement of \na file that originated in your commit. Also, you know that any sets of \ninterchangable documents had already been created when you get a commit \nthat contains one of them.\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"75546","messageId":"alpine.LFD.1.10.0804291232130.23581@xanadu.home","threadId":"13309","inReplyTo":"7f9d599f0804290859y6a579302m5db9f7f827b320a4@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-04-29T16:39:09Z","receivedAt":"2008-04-29T16:39:09Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 29 Apr 2008, Geoffrey Irving wrote:\n\n> On Tue, Apr 29, 2008 at 8:42 AM, Nicolas Pitre <nico@cam.org> wrote:\n> > On Tue, 29 Apr 2008, Andreas Ericsson wrote:\n> >\n> >  > But they won't, because it's impossible to add two objects with the same\n> >  > SHA1 hash key to a git repository, since it will lazily re-use the\n> >  > existing one. In practice, this means that in the case of an \"innocent\"\n> >  > hash-collision, git will actually break by refusing to store the new\n> >  > content.\n> >\n> >  I'd also like to point out that Git usually receive \"untrusted\" new\n> >  objects via the Git protocol through 'git index-pack'.  If you look at\n> >  sha1_object() in index-pack.c, you'll see that active verification\n> >  against hash collision is performed, and the fetch will abruptly be\n> >  aborted if ever that happens.\n> >\n> >  Yes, writing a test case for this was tricky.  :-)\n> \n> Here's the standard scenario for a hash collision attack, with\n> parties, A, B, and C:\n> \n> 1. C, the malicious one, computes the standard two pdfs with matching\n> sha1 hashes.\n> 2. C sends the valid pdf to B through a git commit, and B signs it with a tag.\n> 3. C grabs the signature, and then forwards the \"signed\" commit to A,\n> but substitutes the invalid pdf with the same hash.\n> \n> The fact that git will check for hash collisions within one repository\n> is nice, but it doesn't significantly increase the security of git\n> against hash collision attacks.\n\nSure.  But this is all complete handwaving until a practical collision \ncan be demonstrated.  So far the demonstration hasn't happened, \npractical or not.\n\n\nNicolas\n"},{"id":"75548","messageId":"fv7kmc$2br$1@ger.gmane.org","threadId":"13309","inReplyTo":"200804281829.11866.henrikau@orakel.ntnu.no","subject":"Re: About git and the use of SHA-1","fromName":"Tom Widmer","fromEmail":"tom.widmer@googlemail.com","sentAt":"2008-04-29T17:08:28Z","receivedAt":"2008-04-29T17:08:28Z","isPatch":false,"sender":{"key":"tom.widmer@googlemail.com","avatar":null},"body":"Henrik Austad wrote:\n> Hi list!\n> \n> As far as I have gathered, the SHA-1-sum is used as a identifier for commits, \n> and that is the primary reason for using sha1.  However, several places \n> (including the google tech-talk featuring Linus himself) states that the id's \n> are cryptographically secure.\n> \n> As discussed in [1], SHA-1 is not as secure as it once was (and this was in \n> 2005), and I'm wondering - are there any plans for migrating to another \n> hash-algorithm? I.e. SHA-2, whirlpool..\n> \n> [1] http://www.schneier.com/blog/archives/2005/02/cryptanalysis_o.html\n\nWhy not wait until the results of:\nhttp://www.csrc.nist.gov/groups/ST/hash/index.html\nare available. That will surely be soon enough (I think 2012 is the\nexpected finish date), and should prevent having to switch again in the\nfuture.\n\nThe necessity or otherwise of improving the hashing will be clearer by\nthen too.\n\nTom\n"},{"id":"75553","messageId":"7f9d599f0804291048n2c706f3amdf159ffe86bdbc8@mail.gmail.com","threadId":"13309","inReplyTo":"alpine.LFD.1.10.0804291232130.23581@xanadu.home","subject":"Re: About git and the use of SHA-1","fromName":"Geoffrey Irving","fromEmail":"irving@naml.us","sentAt":"2008-04-29T17:48:17Z","receivedAt":"2008-04-29T17:48:17Z","isPatch":false,"sender":{"key":"irving@naml.us","avatar":"https://gravatar.com/avatar/52d7452fcd134aac0fa12f57a3bb7ef5f3f7e73ca0ab36736d06c6a6132de718?d=mp&s=160"},"body":"On Tue, Apr 29, 2008 at 9:39 AM, Nicolas Pitre <nico@cam.org> wrote:\n>\n> On Tue, 29 Apr 2008, Geoffrey Irving wrote:\n>\n>  > On Tue, Apr 29, 2008 at 8:42 AM, Nicolas Pitre <nico@cam.org> wrote:\n>  > > On Tue, 29 Apr 2008, Andreas Ericsson wrote:\n>  > >\n>  > >  > But they won't, because it's impossible to add two objects with the same\n>  > >  > SHA1 hash key to a git repository, since it will lazily re-use the\n>  > >  > existing one. In practice, this means that in the case of an \"innocent\"\n>  > >  > hash-collision, git will actually break by refusing to store the new\n>  > >  > content.\n>  > >\n>  > >  I'd also like to point out that Git usually receive \"untrusted\" new\n>  > >  objects via the Git protocol through 'git index-pack'.  If you look at\n>  > >  sha1_object() in index-pack.c, you'll see that active verification\n>  > >  against hash collision is performed, and the fetch will abruptly be\n>  > >  aborted if ever that happens.\n>  > >\n>  > >  Yes, writing a test case for this was tricky.  :-)\n>  >\n>  > Here's the standard scenario for a hash collision attack, with\n>  > parties, A, B, and C:\n>  >\n>  > 1. C, the malicious one, computes the standard two pdfs with matching\n>  > sha1 hashes.\n>  > 2. C sends the valid pdf to B through a git commit, and B signs it with a tag.\n>  > 3. C grabs the signature, and then forwards the \"signed\" commit to A,\n>  > but substitutes the invalid pdf with the same hash.\n>  >\n>  > The fact that git will check for hash collisions within one repository\n>  > is nice, but it doesn't significantly increase the security of git\n>  > against hash collision attacks.\n>\n>  Sure.  But this is all complete handwaving until a practical collision\n>  can be demonstrated.  So far the demonstration hasn't happened,\n>  practical or not.\n\nSorry for the confusion: it would handwaving if I was saying git was insecure,\nbut I'm not.  I'm saying that if or when SHA1 becomes vulnerable to collision\nattacks, git will be insecure.\n\nGeoffrey\n"},{"id":"75554","messageId":"alpine.LFD.1.10.0804291352120.23581@xanadu.home","threadId":"13309","inReplyTo":"7f9d599f0804291048n2c706f3amdf159ffe86bdbc8@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-04-29T17:55:32Z","receivedAt":"2008-04-29T17:55:32Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 29 Apr 2008, Geoffrey Irving wrote:\n\n> Sorry for the confusion: it would handwaving if I was saying git was insecure,\n> but I'm not.  I'm saying that if or when SHA1 becomes vulnerable to collision\n> attacks, git will be insecure.\n\nRight.  And if or when that happens then we'll make Git secure again \nwith a different hash.  In the mean time there is low return for the \neffort involved.\n\n\nNicolas\n"},{"id":"75556","messageId":"7f9d599f0804291102j4a30c344h18d12d03a6d5953b@mail.gmail.com","threadId":"13309","inReplyTo":"alpine.LFD.1.10.0804291352120.23581@xanadu.home","subject":"Re: About git and the use of SHA-1","fromName":"Geoffrey Irving","fromEmail":"irving@naml.us","sentAt":"2008-04-29T18:02:16Z","receivedAt":"2008-04-29T18:02:16Z","isPatch":false,"sender":{"key":"irving@naml.us","avatar":"https://gravatar.com/avatar/52d7452fcd134aac0fa12f57a3bb7ef5f3f7e73ca0ab36736d06c6a6132de718?d=mp&s=160"},"body":"On Tue, Apr 29, 2008 at 10:55 AM, Nicolas Pitre <nico@cam.org> wrote:\n> On Tue, 29 Apr 2008, Geoffrey Irving wrote:\n>\n>\n> > Sorry for the confusion: it would handwaving if I was saying git was insecure,\n>  > but I'm not.  I'm saying that if or when SHA1 becomes vulnerable to collision\n>  > attacks, git will be insecure.\n>\n>  Right.  And if or when that happens then we'll make Git secure again\n>  with a different hash.  In the mean time there is low return for the\n>  effort involved.\n\nYes.  I wasn't trying to advocate switching, just making sure people\nknow that the \"collisions don't matter\" argument is bogus.\n\nOne important thing: when SHA1 becomes vulnerable to collision\nattacks, it will still be secure to trust the repositories and tags\nthat exist *at that moment.*  I.e., the transition period from SHA1 to\nthe next hash will also be secure, assuming that preimage attacks\ndon't become possible simultaneously.  So everything is good.\n\nGeoffrey\n"},{"id":"75558","messageId":"vpqwsmg7cfk.fsf@bauges.imag.fr","threadId":"13309","inReplyTo":"7f9d599f0804290859y6a579302m5db9f7f827b320a4@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Matthieu Moy","fromEmail":"matthieu.moy@imag.fr","sentAt":"2008-04-29T18:17:51Z","receivedAt":"2008-04-29T18:17:51Z","isPatch":false,"sender":{"key":"git@matthieu-moy.fr","avatar":"https://avatars.githubusercontent.com/u/14709?v=4"},"body":"\"Geoffrey Irving\" <irving@naml.us> writes:\n\n> Here's the standard scenario for a hash collision attack, with\n> parties, A, B, and C:\n>\n> 1. C, the malicious one, computes the standard two pdfs with matching\n> sha1 hashes.\n> 2. C sends the valid pdf to B through a git commit, and B signs it with a tag.\n> 3. C grabs the signature, and then forwards the \"signed\" commit to A,\n> but substitutes the invalid pdf with the same hash.\n\nJust to add my 2 cents, examples of this are available on the web,\nlike:\n\nhttp://th.informatik.uni-mannheim.de/People/Lucks/HashCollisions/\n\nSame size, same hash. But that's with md5, not sha1.\n\n-- \nMatthieu\n"},{"id":"75559","messageId":"20080429182325.GC1641@frsk.net","threadId":"13309","inReplyTo":"vpqwsmg7cfk.fsf@bauges.imag.fr","subject":"Re: About git and the use of SHA-1","fromName":"Fredrik Skolmli","fromEmail":"fredrik@frsk.net","sentAt":"2008-04-29T18:23:25Z","receivedAt":"2008-04-29T18:23:25Z","isPatch":false,"sender":{"key":"fredrik@frsk.net","avatar":"https://avatars.githubusercontent.com/u/40261?v=4"},"body":"On Tue, Apr 29, 2008 at 08:17:51PM +0200, Matthieu Moy wrote:\n\n> > Here's the standard scenario for a hash collision attack, with\n> > parties, A, B, and C:\n> >\n> > 1. C, the malicious one, computes the standard two pdfs with matching\n> > sha1 hashes.\n> > 2. C sends the valid pdf to B through a git commit, and B signs it with a tag.\n> > 3. C grabs the signature, and then forwards the \"signed\" commit to A,\n> > but substitutes the invalid pdf with the same hash.\n> \n> Just to add my 2 cents, examples of this are available on the web,\n> like:\n> \n> http://th.informatik.uni-mannheim.de/People/Lucks/HashCollisions/\n> \n> Same size, same hash. But that's with md5, not sha1.\n\nWell yes, but that's still using the methods already mentioned in this\nthread. So you do have to get your \"good\" code approved before replacing it\nwith something nasty.\n\n- Fredrik\n\n-- \nRegards,\nFredrik Skolmli\n"},{"id":"75560","messageId":"alpine.LNX.1.00.0804291410340.19665@iabervon.org","threadId":"13309","inReplyTo":"7f9d599f0804291102j4a30c344h18d12d03a6d5953b@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Daniel Barkalow","fromEmail":"barkalow@iabervon.org","sentAt":"2008-04-29T18:41:31Z","receivedAt":"2008-04-29T18:41:31Z","isPatch":false,"sender":{"key":"barkalow@iabervon.org","avatar":"https://avatars.githubusercontent.com/u/55364219?v=4"},"body":"On Tue, 29 Apr 2008, Geoffrey Irving wrote:\n\n> On Tue, Apr 29, 2008 at 10:55 AM, Nicolas Pitre <nico@cam.org> wrote:\n> > On Tue, 29 Apr 2008, Geoffrey Irving wrote:\n> >\n> >\n> > > Sorry for the confusion: it would handwaving if I was saying git was insecure,\n> >  > but I'm not.  I'm saying that if or when SHA1 becomes vulnerable to collision\n> >  > attacks, git will be insecure.\n> >\n> >  Right.  And if or when that happens then we'll make Git secure again\n> >  with a different hash.  In the mean time there is low return for the\n> >  effort involved.\n> \n> Yes.  I wasn't trying to advocate switching, just making sure people\n> know that the \"collisions don't matter\" argument is bogus.\n\nIt's bogus to say they completely don't matter, but I still claim that \nthey don't matter for the things people actually care about. If people can \ngenerate collisions, they can commit a \"weak\" blob with a conditional that \ncan be switched by replacing the blob. But it's almost always true that \npeople could commit a blob with a conditional that can be switched by \nsomething else under the attacker's more direct control. Using a better \nhash function won't save you from a document like:\n\nif (getdate() < 2009)\n  render_good_text\nelse\n  render_evil_text\n\neven if it does help with:\n\nif (AA == AA)\n  render_good_text\nelse\n  render_evil_text\n\nIf you're not checking your files for the former, you shouldn't worry \nabout the latter, because the former is much easier and more subtle.\n\n(Now, an arbitrary preimage attack would actually be significant, still, \nbecause the attacker could replace an honestly-created \"restrictive \nsecurity policy\" file with garbage that will be ignored, leaving stuff \nunprotected)\n\n\t-Daniel\n*This .sig left intentionally blank*\n"},{"id":"75570","messageId":"7f9d599f0804291331v2f44bee1y29c1580d68a3107a@mail.gmail.com","threadId":"13309","inReplyTo":"alpine.LNX.1.00.0804291410340.19665@iabervon.org","subject":"Re: About git and the use of SHA-1","fromName":"Geoffrey Irving","fromEmail":"irving@naml.us","sentAt":"2008-04-29T20:31:51Z","receivedAt":"2008-04-29T20:31:51Z","isPatch":false,"sender":{"key":"irving@naml.us","avatar":"https://gravatar.com/avatar/52d7452fcd134aac0fa12f57a3bb7ef5f3f7e73ca0ab36736d06c6a6132de718?d=mp&s=160"},"body":"On Tue, Apr 29, 2008 at 11:41 AM, Daniel Barkalow <barkalow@iabervon.org> wrote:\n> On Tue, 29 Apr 2008, Geoffrey Irving wrote:\n>\n>  > On Tue, Apr 29, 2008 at 10:55 AM, Nicolas Pitre <nico@cam.org> wrote:\n>  > > On Tue, 29 Apr 2008, Geoffrey Irving wrote:\n>  > >\n>  > >\n>  > > > Sorry for the confusion: it would handwaving if I was saying git was insecure,\n>  > >  > but I'm not.  I'm saying that if or when SHA1 becomes vulnerable to collision\n>  > >  > attacks, git will be insecure.\n>  > >\n>  > >  Right.  And if or when that happens then we'll make Git secure again\n>  > >  with a different hash.  In the mean time there is low return for the\n>  > >  effort involved.\n>  >\n>  > Yes.  I wasn't trying to advocate switching, just making sure people\n>  > know that the \"collisions don't matter\" argument is bogus.\n>\n>  It's bogus to say they completely don't matter, but I still claim that\n>  they don't matter for the things people actually care about. If people can\n>  generate collisions, they can commit a \"weak\" blob with a conditional that\n>  can be switched by replacing the blob. But it's almost always true that\n>  people could commit a blob with a conditional that can be switched by\n>  something else under the attacker's more direct control. Using a better\n>  hash function won't save you from a document like:\n>\n>  if (getdate() < 2009)\n>   render_good_text\n>  else\n>   render_evil_text\n>\n>  even if it does help with:\n>\n>  if (AA == AA)\n>   render_good_text\n>  else\n>   render_evil_text\n>\n>  If you're not checking your files for the former, you shouldn't worry\n>  about the latter, because the former is much easier and more subtle.\n\nI sincerely hope that pdf/postscript don't allow the internal\nrendering code to branch based on the current date.  That would be an\nabsurd security hole, and would indeed make you entirely correct.  If\nyou actually know that it is possible to write that in postscript, I\nwould very much want to see an example.\n\nIn any case, in a binary document format that isn't insane (examples\nof these at least include black and white .png images of documents), a\nvisual check of the content is sufficient to ensure that the next\nperson who looks at it will see roughly the same visual content.  Git\nshould be (and currently is) a secure method of transferring sane\nbinary documents.\n\nGeoffrey\n"},{"id":"75574","messageId":"20080429205031.GA14547@frsk.net","threadId":"13309","inReplyTo":"7f9d599f0804291331v2f44bee1y29c1580d68a3107a@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Fredrik Skolmli","fromEmail":"fredrik@frsk.net","sentAt":"2008-04-29T20:50:31Z","receivedAt":"2008-04-29T20:50:31Z","isPatch":false,"sender":{"key":"fredrik@frsk.net","avatar":"https://avatars.githubusercontent.com/u/40261?v=4"},"body":"On Tue, Apr 29, 2008 at 01:31:51PM -0700, Geoffrey Irving wrote:\n\n> I sincerely hope that pdf/postscript don't allow the internal\n> rendering code to branch based on the current date.  That would be an\n> absurd security hole, and would indeed make you entirely correct.  If\n> you actually know that it is possible to write that in postscript, I\n> would very much want to see an example.\n\nHave a look at \n\n* http://th.informatik.uni-mannheim.de/People/Lucks/HashCollisions/letter_of_rec.ps\nvs \n* http://th.informatik.uni-mannheim.de/People/Lucks/HashCollisions/order.ps\n\nboth found on a website[1] already mentioned[2] in this thread. :-)\n\n[1]: http://th.informatik.uni-mannheim.de/People/Lucks/HashCollisions/\n[2]: http://marc.info/?l=git&m=120949349923584&w=2\n\n- F\n\n-- \nRegards,\nFredrik Skolmli\n"},{"id":"75588","messageId":"7f9d599f0804291439m6f5dd242jb31b84e1a0205cdc@mail.gmail.com","threadId":"13309","inReplyTo":"20080429205031.GA14547@frsk.net","subject":"Re: About git and the use of SHA-1","fromName":"Geoffrey Irving","fromEmail":"irving@naml.us","sentAt":"2008-04-29T21:39:46Z","receivedAt":"2008-04-29T21:39:46Z","isPatch":false,"sender":{"key":"irving@naml.us","avatar":"https://gravatar.com/avatar/52d7452fcd134aac0fa12f57a3bb7ef5f3f7e73ca0ab36736d06c6a6132de718?d=mp&s=160"},"body":"On Tue, Apr 29, 2008 at 1:50 PM, Fredrik Skolmli <fredrik@frsk.net> wrote:\n> On Tue, Apr 29, 2008 at 01:31:51PM -0700, Geoffrey Irving wrote:\n>\n>  > I sincerely hope that pdf/postscript don't allow the internal\n>  > rendering code to branch based on the current date.  That would be an\n>  > absurd security hole, and would indeed make you entirely correct.  If\n>  > you actually know that it is possible to write that in postscript, I\n>  > would very much want to see an example.\n>\n>  Have a look at\n>\n>  * http://th.informatik.uni-mannheim.de/People/Lucks/HashCollisions/letter_of_rec.ps\n>  vs\n>  * http://th.informatik.uni-mannheim.de/People/Lucks/HashCollisions/order.ps\n>\n>  both found on a website[1] already mentioned[2] in this thread. :-)\n>\n>  [1]: http://th.informatik.uni-mannheim.de/People/Lucks/HashCollisions/\n>  [2]: http://marc.info/?l=git&m=120949349923584&w=2\n\nThis is an example of a hash collision, not conditional rendering\nbased on the current date.  I.e., you didn't actually read my email or\nthe email I was replying to. :)\n\nGeoffrey\n"},{"id":"75592","messageId":"20080429215201.GB14547@frsk.net","threadId":"13309","inReplyTo":"7f9d599f0804291439m6f5dd242jb31b84e1a0205cdc@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Fredrik Skolmli","fromEmail":"fredrik@frsk.net","sentAt":"2008-04-29T21:52:01Z","receivedAt":"2008-04-29T21:52:01Z","isPatch":false,"sender":{"key":"fredrik@frsk.net","avatar":"https://avatars.githubusercontent.com/u/40261?v=4"},"body":"On Tue, Apr 29, 2008 at 02:39:46PM -0700, Geoffrey Irving wrote:\n\n> This is an example of a hash collision, not conditional rendering\n> based on the current date.  I.e., you didn't actually read my email or\n> the email I was replying to. :)\n\nAh, you're right. Didn't notice the part about dates. Sorry ;-)\n\n-- \nRegards,\nFredrik Skolmli\n"},{"id":"75625","messageId":"46a038f90804291958u14eddc49sb54c7fd4a3a10381@mail.gmail.com","threadId":"13309","inReplyTo":"7f9d599f0804291331v2f44bee1y29c1580d68a3107a@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-04-30T02:58:53Z","receivedAt":"2008-04-30T02:58:53Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Wed, Apr 30, 2008 at 8:31 AM, Geoffrey Irving <irving@naml.us> wrote:\n>  I sincerely hope that pdf/postscript don't allow the internal\n>  rendering code to branch based on the current date.  That would be an\n>  absurd security hole, and would indeed make you entirely correct.  If\n\nPS is Turing complete, and does know about dates. So yes, you can make\nsuch conditionals.\n\nThat original md5 paper with the 2 PDF files is mainly a good example\nthat you should trust binary blobs, that's all. The md5 trick is a\nnice demo, but misses the point entirely.\n\nI can't find it now, but someone had written a PDF file that printed\nPi computing in inside the PS VM. The tiny file would keep the printer\nchurning out paper until it ran out of memory. :-)\n\ncheers,\n\n\nm\n-- \n martin.langhoff@gmail.com\n martin@laptop.org -- School Server Architect\n - ask interesting questions\n - don't get distracted with shiny stuff - working code first\n - http://wiki.laptop.org/go/User:Martinlanghoff\n"},{"id":"75634","messageId":"7f9d599f0804292218x7d94d7del20d4d48bbad80fb5@mail.gmail.com","threadId":"13309","inReplyTo":"46a038f90804291958u14eddc49sb54c7fd4a3a10381@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"Geoffrey Irving","fromEmail":"irving@naml.us","sentAt":"2008-04-30T05:18:55Z","receivedAt":"2008-04-30T05:18:55Z","isPatch":false,"sender":{"key":"irving@naml.us","avatar":"https://gravatar.com/avatar/52d7452fcd134aac0fa12f57a3bb7ef5f3f7e73ca0ab36736d06c6a6132de718?d=mp&s=160"},"body":"On Tue, Apr 29, 2008 at 7:58 PM, Martin Langhoff\n<martin.langhoff@gmail.com> wrote:\n> On Wed, Apr 30, 2008 at 8:31 AM, Geoffrey Irving <irving@naml.us> wrote:\n>  >  I sincerely hope that pdf/postscript don't allow the internal\n>  >  rendering code to branch based on the current date.  That would be an\n>  >  absurd security hole, and would indeed make you entirely correct.  If\n>\n>  PS is Turing complete, and does know about dates. So yes, you can make\n>  such conditionals.\n\nI knew postscript was Turing complete, but had (naively) assumed it\nexecuted sandboxed and deterministically and would therefore display\nuniformly barring interpreter bugs.  Looking over the spec, I can't\nfind where it's possible to read the current date, but the\nusertime/realtime variables are sufficient as long as the attacker\nknows how fast the relevant machines are.\n\n>  That original md5 paper with the 2 PDF files is mainly a good example\n>  that you should trust binary blobs, that's all. The md5 trick is a\n>  nice demo, but misses the point entirely.\n>\n>  I can't find it now, but someone had written a PDF file that printed\n>  Pi computing in inside the PS VM. The tiny file would keep the printer\n>  churning out paper until it ran out of memory. :-)\n\nAccording to wikipedia, PDF doesn't have conditionals or loops of any\nkind, so you probably mean a postscript file.\n\nGeoffrey\n"},{"id":"75637","messageId":"20080430054700.GA1345@old.davidb.org","threadId":"13309","inReplyTo":"7f9d599f0804292218x7d94d7del20d4d48bbad80fb5@mail.gmail.com","subject":"Re: About git and the use of SHA-1","fromName":"David Brown","fromEmail":"git@davidb.org","sentAt":"2008-04-30T05:47:00Z","receivedAt":"2008-04-30T05:47:00Z","isPatch":false,"sender":{"key":"git@davidb.org","avatar":"https://gravatar.com/avatar/94c86a2938470a74c2eac5e2b69afc0871f79a660295c02219597aba8cb101c1?d=mp&s=160"},"body":"On Tue, Apr 29, 2008 at 10:18:55PM -0700, Geoffrey Irving wrote:\n\n>>  PS is Turing complete, and does know about dates. So yes, you can make\n>>  such conditionals.\n>\n>I knew postscript was Turing complete, but had (naively) assumed it\n>executed sandboxed and deterministically and would therefore display\n>uniformly barring interpreter bugs.  Looking over the spec, I can't\n>find where it's possible to read the current date, but the\n>usertime/realtime variables are sufficient as long as the attacker\n>knows how fast the relevant machines are.\n\nusertime and realtime are from the start of the invocation of the\npostscript interpreter, not based on the outside world.  So, the\ninterpreter could wait arbitrarily long, but has no way of knowing any\nexternal reference to time.\n\nI could imagine trickery with PDF signatures and their expiration times,\nbut you shouldn't be able to do anything with the information, so it would\nbe an exploit, and would probably be fixed.\n\nDavid\n"},{"id":"75639","messageId":"46a038f90804292256jc99306ck78b8718e0377a5c6@mail.gmail.com","threadId":"13309","inReplyTo":"20080430054700.GA1345@old.davidb.org","subject":"Re: About git and the use of SHA-1","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2008-04-30T05:56:58Z","receivedAt":"2008-04-30T05:56:58Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Wed, Apr 30, 2008 at 5:47 PM, David Brown <git@davidb.org> wrote:\n>  usertime and realtime are from the start of the invocation of the\n>  postscript interpreter, not based on the outside world.  So, the\n\nYou guys are right - I misremembered the spec wrt dates. I had the\ndistinct impression that there was a way to get the epoch.\n\nSorry about the noise.\n\n\n\nmartin\n-- \n martin.langhoff@gmail.com\n martin@laptop.org -- School Server Architect\n - ask interesting questions\n - don't get distracted with shiny stuff - working code first\n - http://wiki.laptop.org/go/User:Martinlanghoff\n"}]}