{"thread":{"id":"5410","subject":"Starting to think about sha-256?","startedAt":"2006-08-27T17:56:07Z","lastAt":"2006-08-29T06:17:02Z","messageCount":19,"participants":["Jeff Garzik","Krzysztof Halasa","Linus Torvalds","Johannes Schindelin","David Lang","Jeff King","Florian Weimer"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"25994","messageId":"44F1DCB7.6020804@garzik.org","threadId":"5410","inReplyTo":null,"subject":"Starting to think about sha-256?","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2006-08-27T17:56:07Z","receivedAt":"2006-08-27T17:56:07Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"\nRecent press[1] is talking about sha-1 collisions again.  Even though \nthe reported attack was against a weakened variant of sha-1 (64, not 80, \npasses), it serves as a useful point to start talking about the future.\n\nI argue that sha-256 is better suited to git's purposes, and to modern \nmachines, than sha-1.\n\nUpsides to sha-256:\n* not just a bit increase, but a stronger algorithm.  there is more \nmixing, doing a more-than-incrementally better job at avoiding collisions.\n* the bit increase itself provides more hash space, theoretically \nreducing collisions.\n* properly aligned, a set of 32-byte hashes won't straddle CPU cachelines.\n\nDownsides to sha-256:\n* git protocol/storage format change implications.\n* increase in storage size (20 to 32 bytes per hash).\n* fewer hand-optimized algorithm variants have been implemented.\n* likely more CPU cycles per hash, though I haven't measured.\n\nWikimedia page has lotsa info: \nhttp://en.wikipedia.org/wiki/Secure_Hash_Algorithm\n\nMaybe sha-256 could be considered for the next major-rev of git?\n\n\tJeff\n\n\n[1] http://www.heise-security.co.uk/news/77244\n"},{"id":"25999","messageId":"m31wr1exbf.fsf@defiant.localdomain","threadId":"5410","inReplyTo":"44F1DCB7.6020804@garzik.org","subject":"Re: Starting to think about sha-256?","fromName":"Krzysztof Halasa","fromEmail":"khc@pm.waw.pl","sentAt":"2006-08-27T20:30:12Z","receivedAt":"2006-08-27T20:30:12Z","isPatch":false,"sender":{"key":"khc@pm.waw.pl","avatar":null},"body":"Jeff Garzik <jeff@garzik.org> writes:\n\n> Downsides to sha-256:\n> * git protocol/storage format change implications.\n\nThe only which really matters, I think.\n\n> Maybe sha-256 could be considered for the next major-rev of git?\n\nNot sure, but _if_ we want it we should do it sooner rather than\nlater.\n-- \nKrzysztof Halasa\n"},{"id":"26000","messageId":"Pine.LNX.4.64.0608271343120.27779@g5.osdl.org","threadId":"5410","inReplyTo":"m31wr1exbf.fsf@defiant.localdomain","subject":"Re: Starting to think about sha-256?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-27T20:46:57Z","receivedAt":"2006-08-27T20:46:57Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 27 Aug 2006, Krzysztof Halasa wrote:\n> \n> > Maybe sha-256 could be considered for the next major-rev of git?\n> \n> Not sure, but _if_ we want it we should do it sooner rather than\n> later.\n\nModifying git-convert-objects.c to rewrite the regular sha1 into a sha256 \nshould be fairly straightforward. It's never been used since the early \ndays (and has limits like a maximum of a million objects etc that can need \nfixing), but it shouldn't be \"fundamentally hard\" per se.\n\n\t\tLinus\n"},{"id":"26001","messageId":"m3hczxdgp1.fsf@defiant.localdomain","threadId":"5410","inReplyTo":"Pine.LNX.4.64.0608271343120.27779@g5.osdl.org","subject":"Re: Starting to think about sha-256?","fromName":"Krzysztof Halasa","fromEmail":"khc@pm.waw.pl","sentAt":"2006-08-27T21:14:34Z","receivedAt":"2006-08-27T21:14:34Z","isPatch":false,"sender":{"key":"khc@pm.waw.pl","avatar":null},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Modifying git-convert-objects.c to rewrite the regular sha1 into a sha256 \n> should be fairly straightforward. It's never been used since the early \n> days (and has limits like a maximum of a million objects etc that can need \n> fixing), but it shouldn't be \"fundamentally hard\" per se.\n\nSure. I was rather thinking of rapidly increasing number of git\nrepositories, each with growing history.\n-- \nKrzysztof Halasa\n"},{"id":"26009","messageId":"Pine.LNX.4.63.0608272341330.28360@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"5410","inReplyTo":"Pine.LNX.4.64.0608271343120.27779@g5.osdl.org","subject":"Re: Starting to think about sha-256?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-08-27T22:02:39Z","receivedAt":"2006-08-27T22:02:39Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Sun, 27 Aug 2006, Linus Torvalds wrote:\n\n> On Sun, 27 Aug 2006, Krzysztof Halasa wrote:\n> > \n> > > Maybe sha-256 could be considered for the next major-rev of git?\n> > \n> > Not sure, but _if_ we want it we should do it sooner rather than\n> > later.\n> \n> Modifying git-convert-objects.c to rewrite the regular sha1 into a sha256 \n> should be fairly straightforward. It's never been used since the early \n> days (and has limits like a maximum of a million objects etc that can need \n> fixing), but it shouldn't be \"fundamentally hard\" per se.\n\nBut what about signed tags? (This issue has come up before, but never has \nbeen adressed.)\n\nI also thought about supporting hybrid hashes, i.e. that older objects \nstill can be hashed with SHA-1. Alas, a simple thought experiment \ndemonstrates how silly that idea is: most of the objects will not change \nbetween two revisions, and they'd have to be rehashed with SHA-256 (or \nwhatever we decide upon) anyway, so hybrids would do no good.\n\nA better idea would be to increment the repository version, and expect \nSHA-1 for version 1, SHA-256 for version >= 2.\n\nHowever, I could imagine that we do not need this huge change (it would \nbreak _many_ setups). The breakthrough was announced last Tuesday, and it \ninvolved 75% payload, i.e. to fake a new -- say -- git.c, one would need \nto enlarge git.c by a factor 4, and you would see a lot of gibberish \ninside some comment. (Note that I did not listen to the talk myself, this \nis all deducted from the scarce information which is available via the \n'net.)\n\nEven if the breakthrough really comes to full SHA-1, you still have to add \n_at least_ 20 bytes of gibberish. Which would be harder to spot, but it \nwould be spotted.\n\nThis made me think about the use of hashes in git. Why do we need a hash \nhere (in no particular order):\n\n1) integrity checking,\n2) fast lookup,\n3) identifying objects (related to (2)),\n4) trust.\n\nExcept for (4), I do not see why SHA-1 -- even if broken -- should not be \nadequate. It is not like somebody found out that all JPGs tend to have \nsimilar hashes so that collisions are more likely.\n\nAnd thinking about trust: The hash is augmented by thinking persons. It is \nnot like you blindly trust a person forever. You build up trust, and once \nyou were failed, the trust is lost, and very hard to build up again. So, \nyou just would try to get all objects again from somebody you still trust, \nand never pull from the loser^H^H^H^H^Huntrusted person again. Ever.\n\nBesides, as has been pointed out several times, a dishonest person could \ntry to sneak bad code into your repository _regardless_ of a secure hash.\n\nSo: Do we really need a secure hash, or do we need an adequate hash, and \njust happen to use one which was intended as a secure hash, but no longer \nis?\n\nCiao,\nDscho\n"},{"id":"26013","messageId":"Pine.LNX.4.64.0608271522390.27779@g5.osdl.org","threadId":"5410","inReplyTo":"Pine.LNX.4.63.0608272341330.28360@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Starting to think about sha-256?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-27T22:35:20Z","receivedAt":"2006-08-27T22:35:20Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 28 Aug 2006, Johannes Schindelin wrote:\n> > \n> > Modifying git-convert-objects.c to rewrite the regular sha1 into a sha256 \n> > should be fairly straightforward. It's never been used since the early \n> > days (and has limits like a maximum of a million objects etc that can need \n> > fixing), but it shouldn't be \"fundamentally hard\" per se.\n> \n> But what about signed tags? (This issue has come up before, but never has \n> been adressed.)\n\nSigned tags fundamentally have to be re-signed. That's by design: if \nsomebody could rewrite an archive and signed tags would still be accepted \nto have the right signature, that would be a _serious_ sign of a totally \nbroken security model.\n\nThe git security model isn't broken.\n\n> I also thought about supporting hybrid hashes, i.e. that older objects \n> still can be hashed with SHA-1. Alas, a simple thought experiment \n> demonstrates how silly that idea is: most of the objects will not change \n> between two revisions, and they'd have to be rehashed with SHA-256 (or \n> whatever we decide upon) anyway, so hybrids would do no good.\n\nIndeed. Hybrids would not only do no good, but they would actually \n_actively_ hurt things, because they'd fundamentally break the notion that \nthe hash being identical means that the object (blob, tree, subtree) is \nthe same.\n\nSo allowing two names for the same object is very fundamentally wrong in \ngit-speak. \n\n> A better idea would be to increment the repository version, and expect \n> SHA-1 for version 1, SHA-256 for version >= 2.\n\nYes. It would be reasonably painful for users, though (as Krzysztof \ncorrectly points out). Every client would have to convert when a \nrepository they track is converted.\n\n> Even if the breakthrough really comes to full SHA-1, you still have to add \n> _at least_ 20 bytes of gibberish. Which would be harder to spot, but it \n> would be spotted.\n\nYeah, I don't think this is at all critical, especially since git really \non a security level doesn't _depend_ on the hashes being cryptographically \nsecure. As I explained early on (ie over a year ago, back when the whole \ndesign of git was being discussed), the _security_ of git actually depends \non not cryptographic hashes, but simply on everybody being able to secure \ntheir own _private_ repository.\n\nSo the only thing git really _requires_ is a hash that is _unique_ for the \ndeveloper (and there we are talking not of an _attacker_, but a benign \nparticipant).\n\nThat said, the cryptographic security of SHA-1 is obviously a real bonus. \nSo I'd be disappointed if SHA-1 can be broken more easily (and I obviously \nalready argued against using MD5, exactly because generating duplicates of \nthat is fairly easy). But it's not \"fundamentally required\" in git per se.\n\n[ The one exception: the \"signed tags\" security does depend on the hashes \n  being cryptographically strong. So again, breaking SHA-1 would not mean \n  that git stops working, but it _would_ potentially mean that if you \n  don't trust your own _private_ repository, the signed tag may no longer \n  protect you entirely ]\n\n> This made me think about the use of hashes in git. Why do we need a hash \n> here (in no particular order):\n> \n> 1) integrity checking,\n> 2) fast lookup,\n> 3) identifying objects (related to (2)),\n> 4) trust.\n> \n> Except for (4), I do not see why SHA-1 -- even if broken -- should not be \n> adequate. It is not like somebody found out that all JPGs tend to have \n> similar hashes so that collisions are more likely.\n\nCorrect. I'm pretty sure we had exactly this discussion around May 2005, \nbut I'm too lazy to search ;)\n\n\t\tLinus\n"},{"id":"26056","messageId":"656C30A1EFC89F6B2082D9B6@localhost","threadId":"5410","inReplyTo":"Pine.LNX.4.64.0608271522390.27779@g5.osdl.org","subject":"Re: Starting to think about sha-256?","fromName":"David Lang","fromEmail":"david.lang@digitalinsight.com","sentAt":"2006-08-28T17:27:29Z","receivedAt":"2006-08-28T17:27:29Z","isPatch":false,"sender":{"key":"david.lang@digitalinsight.com","avatar":null},"body":"--On Sunday, August 27, 2006 03:35:20 PM -0700 Linus Torvalds \n<torvalds@osdl.org> wrote:\n>\n> On Mon, 28 Aug 2006, Johannes Schindelin wrote:\n>> Even if the breakthrough really comes to full SHA-1, you still have to\n>> add  _at least_ 20 bytes of gibberish. Which would be harder to spot,\n>> but it  would be spotted.\n>\n> Yeah, I don't think this is at all critical, especially since git really\n> on a security level doesn't _depend_ on the hashes being\n> cryptographically  secure. As I explained early on (ie over a year ago,\n> back when the whole  design of git was being discussed), the _security_\n> of git actually depends  on not cryptographic hashes, but simply on\n> everybody being able to secure  their own _private_ repository.\n>\n> So the only thing git really _requires_ is a hash that is _unique_ for\n> the  developer (and there we are talking not of an _attacker_, but a\n> benign  participant).\n>\n> That said, the cryptographic security of SHA-1 is obviously a real bonus.\n> So I'd be disappointed if SHA-1 can be broken more easily (and I\n> obviously  already argued against using MD5, exactly because generating\n> duplicates of  that is fairly easy). But it's not \"fundamentally\n> required\" in git per se.\n\n\n>> This made me think about the use of hashes in git. Why do we need a hash\n>> here (in no particular order):\n>>\n>> 1) integrity checking,\n>> 2) fast lookup,\n>> 3) identifying objects (related to (2)),\n>> 4) trust.\n>>\n>> Except for (4), I do not see why SHA-1 -- even if broken -- should not\n>> be  adequate. It is not like somebody found out that all JPGs tend to\n>> have  similar hashes so that collisions are more likely.\n>\n> Correct. I'm pretty sure we had exactly this discussion around May 2005,\n> but I'm too lazy to search ;)\n\njust to double check.\n\nif you already have a file A in git with hash X is there any condition \nwhere a remote file with hash X (but different contents) would overwrite \nthe local version?\n\nwhat would happen if you ended up with two packs that both contained a file \nwith hash X but with different contents and then did a repack on them? \n(either packs from different sources, or packs downloaded through some \nmechanism other then the git protocol are two ways this could happen that I \ncan think of)\n\nDavid Lang\n"},{"id":"26057","messageId":"Pine.LNX.4.64.0608281034440.27779@g5.osdl.org","threadId":"5410","inReplyTo":"656C30A1EFC89F6B2082D9B6@localhost","subject":"Re: Starting to think about sha-256?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-28T17:56:01Z","receivedAt":"2006-08-28T17:56:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 28 Aug 2006, David Lang wrote:\n> \n> just to double check.\n> \n> if you already have a file A in git with hash X is there any condition where a\n> remote file with hash X (but different contents) would overwrite the local\n> version?\n\nNope. If it has the same SHA1, it means that when we receive the object \nfrom the other end, we will _not_ overwrite the object we already have.\n\nSo what happens is that if we ever see a collision, the \"earlier\" object \nin any particular repository will always end up overriding. But note that \n\"earlier\" is obviously per-repository, in the sense that the git object \nnetwork generates a DAG that is not fully ordered, so while different \nrepositories will agree about what is \"earlier\" in the case of direct \nancestry, if the object came through separate and not directly related \nbranches, two different repos may obviously have gotten the two objects in \ndifferent order.\n\nHowever, the \"earlier will override\" is very much what you want from a \nsecurity standpoint: remember that the git model is that you should \nprimarily trust only your _own_ repository. So if you do a \"git pull\", the \nnew incoming objects are by definition less trustworthy than the objects \nyou already have, and as such it would be wrong to allow a new object to \nreplace an old one.\n\nSo you have two cases of collision:\n\n - the inadvertent kind, where you somehow are very very unlucky, and two \n   files end up having the same SHA1. At that point, what happens is that \n   when you commit that file (or do a \"git-update-index\" to move it into \n   the index, but not committed yet), the SHA1 of the new contents will be \n   computed, but since it matches an old object, a new object won't be \n   created, and the commit-or-index ends up pointing to the _old_ object.\n\n   You won't notice immediately (since the index will match the old object \n   SHA1, and that means that something like \"git diff\" will use the \n   checked-out copy), but if you ever do a tree-level diff (or you \n   do a clone or pull, or force a checkout) you'll suddenly notice that \n   that file has changed to something _completely_ different than what you \n   expected. So you would generally notice this kind of collision fairly \n   quickly.\n\n   In related news, the question is what to do about the inadvertent \n   collision.. First off, let me remind people that the inadvertent kind \n   of collision is really really _really_ damn unlikely, so we'll quite \n   likely never ever see it in the full history of the universe. But _if_ \n   it happens, it's not the end of the world: what you'd most likely have \n   to do is just change the file that collided slightly, and just force a \n   new commit with the changed contents (add a comment saying \"/* This \n   line added to avoid collision */\") and then teach git about the magic \n   SHA1 that has been shown to be dangerous.\n\n   So over a couple of million years, maybe we'll have to add one or two \n   \"poisoned\" SHA1 values to git. It's very unlikely to be a maintenance \n   problem ;)\n\n - The attacker kind of collision because somebody broke (or brute-forced) \n   SHA1.\n\n   This one is clearly a _lot_ more likely than the inadvertent kind, but \n   by definition it's always a \"remote\" repository. If the attacker had \n   access to the local repository, he'd have much easier ways to screw you \n   up.\n\n   So in this case, the collision is entirely a non-issue: you'll get a \n   \"bad\" repository that is different from what the attacker intended, but \n   since you'll never actually use his colliding object, it's _literally_ \n   no different from the attacker just not having found a collision at \n   all, but just using the object you already had (ie it's 100% equivalent \n   to the \"trivial\" collision of the identical file generating the same \n   SHA1).\n\n> what would happen if you ended up with two packs that both contained a file\n> with hash X but with different contents and then did a repack on them? (either\n> packs from different sources, or packs downloaded through some mechanism other\n> then the git protocol are two ways this could happen that I can think of)\n\nSee above. The only _dangerous_ kind of collision is the inadvertent kind, \nbut that's obviously also the very very unlikely kind.\n\n\t\t\tLinus\n"},{"id":"26058","messageId":"Pine.LNX.4.64.0608281059380.27779@g5.osdl.org","threadId":"5410","inReplyTo":"Pine.LNX.4.64.0608281034440.27779@g5.osdl.org","subject":"Re: Starting to think about sha-256?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-28T18:06:27Z","receivedAt":"2006-08-28T18:06:27Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 28 Aug 2006, Linus Torvalds wrote:\n> \n>  - The attacker kind of collision because somebody broke (or brute-forced) \n>    SHA1.\n> \n>    This one is clearly a _lot_ more likely than the inadvertent kind, but \n>    by definition it's always a \"remote\" repository. If the attacker had \n>    access to the local repository, he'd have much easier ways to screw you \n>    up.\n> \n>    So in this case, the collision is entirely a non-issue: you'll get a \n>    \"bad\" repository that is different from what the attacker intended, but \n>    since you'll never actually use his colliding object, it's _literally_ \n>    no different from the attacker just not having found a collision at \n>    all, but just using the object you already had (ie it's 100% equivalent \n>    to the \"trivial\" collision of the identical file generating the same \n>    SHA1).\n\nBtw, this is obviously only true for the native git protocol itself.\n\nIf the attacker can fool you into generating the new file _yourself_, he \ncan cause your checked-out copy to not match the git object database any \nmore.\n\nIn other words, one \"interesting\" attack vector is to feed you the \ncolliding SHA1 not through a git-to-git transfer, but by generating a \n_patch_ that when applied will generate the collision, so that when you \nthen commit that patch, you get something else than you expected.\n\nAnd _this_ is where it's important that the hash that git uses be a \nnon-trivial one - ie we don't want people to be able to generate two files \nthat look superficially \"ok\".\n\nSo here's the rule: If you ever get a patch that looks like line-noise, \nespecially from somebody you don't trust, DON'T APPLY IT!\n\nNow, that is obviously something you should never do _regardless_ of any \ngit issues, so I don't think this is really a problem either. If you apply \npatches from people you don't have a good reason to trust without \nsanity-checking them, you deserve whatever you get, and quite frankly, a \nSHA1 hash collision is the _least_ of your problems ;)\n\n(This ends up boiling down to one common issue: it's generally _much_ \neasier to attack a project through _other_ means than through a hash \ncollision. And I pretty much guarantee that that is the case even if we \nwere to use a much weaker hash, like MD5. Hash collisions fundamentally \njust aren't good attack vectors, and it's a hell of a lot easier to try \nto insert bad code by other means)\n\n\t\t\tLinus\n"},{"id":"26060","messageId":"20060828183252.GC2950@coredump.intra.peff.net","threadId":"5410","inReplyTo":"Pine.LNX.4.64.0608281034440.27779@g5.osdl.org","subject":"Re: Starting to think about sha-256?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-08-28T18:32:52Z","receivedAt":"2006-08-28T18:32:52Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Aug 28, 2006 at 10:56:01AM -0700, Linus Torvalds wrote:\n\n> However, the \"earlier will override\" is very much what you want from a \n> security standpoint: remember that the git model is that you should \n> primarily trust only your _own_ repository. So if you do a \"git pull\", the \n\nThis concept breaks down somewhat if you are pulling from two\nrepositories (one good and one evil). If I pull from the evil repo\nfirst, that will become my \"earlier\" object, and I will never get the\ncolliding object from the good repo.\n\nExecuting such an attack might not be that hard, either (once we get\nover that little hump of creating collisions at will!). The owner of\n'evil' has to know a SHA1 that will be in 'good' before it makes it to\n'good'. However, I imagine we frequently see SHA1s migrate from more\ncentral repos (like .../torvalds/linux-2.6.git) to less central ones\n(subsystem / port maintainers, etc).\n\n-Peff\n"},{"id":"26061","messageId":"Pine.LNX.4.64.0608281137530.27779@g5.osdl.org","threadId":"5410","inReplyTo":"20060828183252.GC2950@coredump.intra.peff.net","subject":"Re: Starting to think about sha-256?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-28T18:46:39Z","receivedAt":"2006-08-28T18:46:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 28 Aug 2006, Jeff King wrote:\n>\n> On Mon, Aug 28, 2006 at 10:56:01AM -0700, Linus Torvalds wrote:\n> \n> > However, the \"earlier will override\" is very much what you want from a \n> > security standpoint: remember that the git model is that you should \n> > primarily trust only your _own_ repository. So if you do a \"git pull\", the \n> \n> This concept breaks down somewhat if you are pulling from two\n> repositories (one good and one evil). If I pull from the evil repo\n> first, that will become my \"earlier\" object, and I will never get the\n> colliding object from the good repo.\n\nSure. But if you are pulling from an untrusted source, you'd better at \nleast check the result.\n\nIn fact, that's partly why \"git pull\" will do a diffstat after the pull. \nExactly to force people to at least be minimally aware of what they \npulled. And \"gitk ORIG_HEAD..\" is a great thing to always run when you \npull from somebody you don't know and trust really well.\n\nOf course, that all was done mostly not because I don't \"trust\" the people \nI work with, but more because I didn't always trust that they'd do the \nright thing with git (ie they'd screw up the repo not because they were \nevil, but because they made a mistake).\n\nSo even if you pull from an \"evil\" repo first, and you somehow get a \"bad\" \nobject, the point is, the bad object _should_ be the one that overrides. \n\nWhy? Because once you find out that the evil repo was bad (which you'll \neventually find simply because it caused some bug - if the evil repo only \nhelps you, it's obviously not evil at all), what you need to do is reset \nto _before_ the evil repo happened, do a \"git repack -a -d\" and finally a \n\"git prune\" to clean out all the bad cruft, and then pull the good repo \nwithout pulling the bad one first.\n\nAfter that, you apologize to everybody for screwing up and pulling from \nsomebody you didn't trust, and then ask them to re-clone (or give them the \nappropriate \"git reset\" + \"git repack -a\" + \"git prune\" + \"git pull\" \nsequence so that they can fix their existing repos).\n\nThe point being, a hash attack is really no worse than an attack that \nfools you into applying a really bad diff (regardless of SCM), and it's a \nhell of a lot harder to do. Both a hash attack and a diff attack mean that \nthe person merging data should either trust his source or inspect the end \nresult.\n\nAnybody who just blindly accepts data from untrusted sources is screwed in \nso many other ways that the hash attack simply isn't even on the radar.\n\n\t\tLinus\n"},{"id":"26062","messageId":"20060828190058.GA5027@coredump.intra.peff.net","threadId":"5410","inReplyTo":"Pine.LNX.4.64.0608281137530.27779@g5.osdl.org","subject":"Re: Starting to think about sha-256?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-08-28T19:00:58Z","receivedAt":"2006-08-28T19:00:58Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Aug 28, 2006 at 11:46:39AM -0700, Linus Torvalds wrote:\n\n> Sure. But if you are pulling from an untrusted source, you'd better at \n> least check the result.\n\nI completely agree; however, even discussing \"earlier takes precedence\"\nentails that you are somehow pulling from an untrusted source. I just\nwanted to point out that \"earlier\" does not always mean \"more trusted\nthan the thing you're pulling now\" (since it might have just been pulled\nearlier, not created or verified by you).\n\n> Anybody who just blindly accepts data from untrusted sources is screwed in \n> so many other ways that the hash attack simply isn't even on the radar.\n\nAgreed.\n\n-Peff\n"},{"id":"26064","messageId":"m34pvwfwl5.fsf@defiant.localdomain","threadId":"5410","inReplyTo":"Pine.LNX.4.64.0608281034440.27779@g5.osdl.org","subject":"Re: Starting to think about sha-256?","fromName":"Krzysztof Halasa","fromEmail":"khc@pm.waw.pl","sentAt":"2006-08-28T20:12:54Z","receivedAt":"2006-08-28T20:12:54Z","isPatch":false,"sender":{"key":"khc@pm.waw.pl","avatar":null},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n>    In related news, the question is what to do about the inadvertent \n>    collision.. First off, let me remind people that the inadvertent kind \n>    of collision is really really _really_ damn unlikely, so we'll quite \n>    likely never ever see it in the full history of the universe.\n\nActually I think we may see it when somebody tries to put a real\nexample of conflicting SHA-1 pair into git repository.\n-- \nKrzysztof Halasa\n"},{"id":"26065","messageId":"Pine.LNX.4.64.0608281317070.27779@g5.osdl.org","threadId":"5410","inReplyTo":"m34pvwfwl5.fsf@defiant.localdomain","subject":"Re: Starting to think about sha-256?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-28T20:20:26Z","receivedAt":"2006-08-28T20:20:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 28 Aug 2006, Krzysztof Halasa wrote:\n>\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> >    In related news, the question is what to do about the inadvertent \n> >    collision.. First off, let me remind people that the inadvertent kind \n> >    of collision is really really _really_ damn unlikely, so we'll quite \n> >    likely never ever see it in the full history of the universe.\n> \n> Actually I think we may see it when somebody tries to put a real\n> example of conflicting SHA-1 pair into git repository.\n\nWell, by definition, I wouldn't call that \"inadvertent\" ;)\n\nAnyway, the way to do it (if you want to use git to document SHA1 hash \nmismatches) is to just check the files that have an identical SHA1 in. It \nwill magically work!\n\nWhy? Because a git SHA1 is actually _not_ the SHA1 of the file itself, \nit's the SHA1 of the file _with_the_git_header_added_.\n\nSo if you find two files that have the same SHA1, they would also have to \nhave the same length in order to actually generate the same object name. \nIf they have different lenths, you can just check them into git, and \nthey'll get two different git SHA1 names and you'll have a cool git \narchive that when you check the files out, they checked-out files will \nshare the same SHA1 ;)\n\n\t\tLinus\n"},{"id":"26066","messageId":"m3r6z0ef95.fsf@defiant.localdomain","threadId":"5410","inReplyTo":"Pine.LNX.4.64.0608281317070.27779@g5.osdl.org","subject":"Re: Starting to think about sha-256?","fromName":"Krzysztof Halasa","fromEmail":"khc@pm.waw.pl","sentAt":"2006-08-28T21:12:38Z","receivedAt":"2006-08-28T21:12:38Z","isPatch":false,"sender":{"key":"khc@pm.waw.pl","avatar":null},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> Anyway, the way to do it (if you want to use git to document SHA1 hash \n> mismatches) is to just check the files that have an identical SHA1 in. It \n> will magically work!\n>\n> Why? Because a git SHA1 is actually _not_ the SHA1 of the file itself, \n> it's the SHA1 of the file _with_the_git_header_added_.\n>\n> So if you find two files that have the same SHA1, they would also have to \n> have the same length in order to actually generate the same object name. \n\nWell, conflicting files will most probably have the same size,\nlike with MD5 cases :-)\n-- \nKrzysztof Halasa\n"},{"id":"26068","messageId":"Pine.LNX.4.64.0608281421120.27779@g5.osdl.org","threadId":"5410","inReplyTo":"m3r6z0ef95.fsf@defiant.localdomain","subject":"Re: Starting to think about sha-256?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-28T21:23:08Z","receivedAt":"2006-08-28T21:23:08Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 28 Aug 2006, Krzysztof Halasa wrote:\n> \n> Well, conflicting files will most probably have the same size,\n> like with MD5 cases :-)\n\nThat's only true for the much easier injection case where you generate \n_both_ files together.\n\n>From an external git hash-attack standpoint, that's not a very useful \ncase. It's much more useful if you can make a new file that has a hash \nthat matches a given old file, and in that case, the filelengths are \nlikely not the same.\n\n\t\t\tLinus\n"},{"id":"26071","messageId":"Pine.LNX.4.63.0608290101320.28360@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"5410","inReplyTo":"Pine.LNX.4.64.0608281034440.27779@g5.osdl.org","subject":"Re: Starting to think about sha-256?","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-08-28T23:09:18Z","receivedAt":"2006-08-28T23:09:18Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 28 Aug 2006, Linus Torvalds wrote:\n\n> \n> \n> On Mon, 28 Aug 2006, David Lang wrote:\n> > \n> > just to double check.\n> > \n> > if you already have a file A in git with hash X is there any condition where a\n> > remote file with hash X (but different contents) would overwrite the local\n> > version?\n> \n> Nope. If it has the same SHA1, it means that when we receive the object \n> from the other end, we will _not_ overwrite the object we already have.\n\nThe only notable exception I can think of: \"git fetch -k\". If you then try \nto retrieve the bogus object, it will return the one of whichever pack was \nreturned first be readdir(). (If I read the source correctly.)\n\nNow, the cases are rare where you do both \"git fetch -k\" and \"git repack \n-a -d\" (the latter of which _could_ leave a hole in the directory which \n_could_ make the next fetched pack fill that hole, which in turn _could_ \nmake readdir() return that pack before more \"senior\" packs) in the same \nrepository, but in these cases, yes, you could end up with the copy of the \nremote side.\n\nYou'd need to explicitely use \"git fetch -k\", though.\n\nCiao,\nDscho\n"},{"id":"26072","messageId":"Pine.LNX.4.64.0608281647300.27779@g5.osdl.org","threadId":"5410","inReplyTo":"Pine.LNX.4.63.0608290101320.28360@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Starting to think about sha-256?","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-08-28T23:48:51Z","receivedAt":"2006-08-28T23:48:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 29 Aug 2006, Johannes Schindelin wrote:\n>\n> > Nope. If it has the same SHA1, it means that when we receive the object \n> > from the other end, we will _not_ overwrite the object we already have.\n> \n> The only notable exception I can think of: \"git fetch -k\". If you then try \n> to retrieve the bogus object, it will return the one of whichever pack was \n> returned first be readdir(). (If I read the source correctly.)\n\nGood point.\n\nI didn't even think of \"-k\", since I mentally put that in the \"initial \nclone usage only\" category, but yeah, if people use it for incremental \nupdates too, that could indeed cause ambiguity in which object to use when \nthe other end does something bad.\n\n\t\tLinus\n"},{"id":"26090","messageId":"87fyfg83s1.fsf@mid.deneb.enyo.de","threadId":"5410","inReplyTo":"44F1DCB7.6020804@garzik.org","subject":"Re: Starting to think about sha-256?","fromName":"Florian Weimer","fromEmail":"fw@deneb.enyo.de","sentAt":"2006-08-29T06:17:02Z","receivedAt":"2006-08-29T06:17:02Z","isPatch":false,"sender":{"key":"fw@deneb.enyo.de","avatar":null},"body":"* Jeff Garzik:\n\n> * likely more CPU cycles per hash, though I haven't measured.\n\nAccording to a quick test using \"openssl speed\", it's a factor of two\nto four, depending on the input size (the difference is less\npronounced for small input sizes).\n\n> Maybe sha-256 could be considered for the next major-rev of git?\n\nAnd in 2008, you'd have to rewrite history again, to use the next\n\"stronger\" hash function?  Do you think that's really necessary or\ndesirable?  Most users will have good control over what data enters\ntheir repositories, so they can spot the evil twins thanks to their\nhigh-entropy contents.  Obviously, a second preimage attack would\nmattr, but even for MD5, we aren't close to that one AFAIK.\n"}]}