{"thread":{"id":"42005","subject":"Migrating away from SHA-1?","startedAt":"2016-04-12T22:38:04Z","lastAt":"2016-04-15T02:22:33Z","messageCount":21,"participants":["H. Peter Anvin","Stefan Beller","Jeff King","David Turner","Junio C Hamano","Duy Nguyen","Theodore Ts'o","Joey Hess","brian m. carlson"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"283228","messageId":"570D78CC.9030807@zytor.com","threadId":"42005","inReplyTo":null,"subject":"Migrating away from SHA-1?","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2016-04-12T22:38:04Z","receivedAt":"2016-04-12T22:38:04Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"OK, I'm going to open this can of worms...\n\nAt what point do we migrate from SHA-1?  At this point the \ncryptoanalysis of SHA-1 is most likely a matter of time.\n\nFor existing repositories we will need to have a migration mechanism. \nSince we can't modify objects without completely invalidating the \ncryptographic properties, what I would suggest is that we leave the \nexisting objects as is, with a persistent lookup table from SHA-1 to \n<new hash>, and have that lookup table signed (e.g. GPG) by the person \nresponsible for converting the repository.  This freezes the \ncryptographic status of the existing SHA-1 objects at the time the \nconversion happens.  This is a very good reason to do this before SHA-1 \nis actually broken  In contrast. SHA-2 has been surprisingly resistant \nto cryptoanalysis, to the point that SHA-3 was motivated by performance \nand the desire to have a well-tested function based on entirely \ndifferent principles should a generic attack against the common \nstructure of MD5/SHA-1/SHA-2 would ever be found.\n\n\t-hpa\n"},{"id":"283230","messageId":"CAGZ79kaUN0G7i0GNZgWU7ZzJvWY=k=Rc6tqWvJsTu8gcRhP5bA@mail.gmail.com","threadId":"42005","inReplyTo":"570D78CC.9030807@zytor.com","subject":"Re: Migrating away from SHA-1?","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2016-04-12T23:00:18Z","receivedAt":"2016-04-12T23:00:18Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Tue, Apr 12, 2016 at 3:38 PM, H. Peter Anvin <hpa@zytor.com> wrote:\n> OK, I'm going to open this can of worms...\n>\n> At what point do we migrate from SHA-1?  At this point the cryptoanalysis of\n> SHA-1 is most likely a matter of time.\n\nAnd I thought the cryptographic properties of SHA1 did not matter for\nGits use case.\nWe could employ broken md5 or such as well.\n( see http://stackoverflow.com/questions/28792784/why-does-git-use-a-cryptographic-hash-function\n)\nThat is because security goes on top via gpg signing of tags/commits.\n\nI am not sure if anyone came up with\na counter argument to Linus reasoning there?\n\n>\n> For existing repositories we will need to have a migration mechanism. Since\n> we can't modify objects without completely invalidating the cryptographic\n> properties, what I would suggest is that we leave the existing objects as\n> is, with a persistent lookup table from SHA-1 to <new hash>, and have that\n> lookup table signed (e.g. GPG) by the person responsible for converting the\n> repository.  This freezes the cryptographic status of the existing SHA-1\n> objects at the time the conversion happens.  This is a very good reason to\n> do this before SHA-1 is actually broken  In contrast. SHA-2 has been\n> surprisingly resistant to cryptoanalysis, to the point that SHA-3 was\n> motivated by performance and the desire to have a well-tested function based\n> on entirely different principles should a generic attack against the common\n> structure of MD5/SHA-1/SHA-2 would ever be found.\n\nWhen the kernel moved from BitKeeper to Git, all history was thrown away,\nand started from scratch. The old history could be grafted into the\nrepo, if you cared\nthough.\n\nI'd propose to go that route again and use a sha1 graft history which\nyou can get optionally\nput into your new history for convenience.\n\nStefan\n\n>\n>         -hpa\n>\n> --\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n"},{"id":"283240","messageId":"570D7F8D.9050406@zytor.com","threadId":"42005","inReplyTo":"CAGZ79kaUN0G7i0GNZgWU7ZzJvWY=k=Rc6tqWvJsTu8gcRhP5bA@mail.gmail.com","subject":"Re: Migrating away from SHA-1?","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2016-04-12T23:06:53Z","receivedAt":"2016-04-12T23:06:53Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"On 04/12/16 16:00, Stefan Beller wrote:\n> On Tue, Apr 12, 2016 at 3:38 PM, H. Peter Anvin <hpa@zytor.com> wrote:\n>> OK, I'm going to open this can of worms...\n>>\n>> At what point do we migrate from SHA-1?  At this point the cryptoanalysis of\n>> SHA-1 is most likely a matter of time.\n>\n> And I thought the cryptographic properties of SHA1 did not matter for\n> Gits use case.\n> We could employ broken md5 or such as well.\n> ( see http://stackoverflow.com/questions/28792784/why-does-git-use-a-cryptographic-hash-function\n> )\n> That is because security goes on top via gpg signing of tags/commits.\n>\n> I am not sure if anyone came up with\n> a counter argument to Linus reasoning there?\n>\n\nNot true, because what we are signing is a chain of SHA-1s; the \nsignature is meaningless unless the integrity of the hash chain is \ninviolate.\n\n>>\n>> For existing repositories we will need to have a migration mechanism. Since\n>> we can't modify objects without completely invalidating the cryptographic\n>> properties, what I would suggest is that we leave the existing objects as\n>> is, with a persistent lookup table from SHA-1 to <new hash>, and have that\n>> lookup table signed (e.g. GPG) by the person responsible for converting the\n>> repository.  This freezes the cryptographic status of the existing SHA-1\n>> objects at the time the conversion happens.  This is a very good reason to\n>> do this before SHA-1 is actually broken  In contrast. SHA-2 has been\n>> surprisingly resistant to cryptoanalysis, to the point that SHA-3 was\n>> motivated by performance and the desire to have a well-tested function based\n>> on entirely different principles should a generic attack against the common\n>> structure of MD5/SHA-1/SHA-2 would ever be found.\n>\n> When the kernel moved from BitKeeper to Git, all history was thrown away,\n> and started from scratch. The old history could be grafted into the\n> repo, if you cared\n> though.\n>\n> I'd propose to go that route again and use a sha1 graft history which\n> you can get optionally\n> put into your new history for convenience.\n>\n\nThat was done more for legal reasons than anything else, as far as I \nunderstand.  The userbase of git today is also much, much larger than \nthe userbase for BK ever was.\n\n\t-hpa\n"},{"id":"283243","messageId":"20160412231518.GA2210@sigill.intra.peff.net","threadId":"42005","inReplyTo":"CAGZ79kaUN0G7i0GNZgWU7ZzJvWY=k=Rc6tqWvJsTu8gcRhP5bA@mail.gmail.com","subject":"Re: Migrating away from SHA-1?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-04-12T23:15:19Z","receivedAt":"2016-04-12T23:15:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 12, 2016 at 04:00:18PM -0700, Stefan Beller wrote:\n\n> On Tue, Apr 12, 2016 at 3:38 PM, H. Peter Anvin <hpa@zytor.com> wrote:\n> > OK, I'm going to open this can of worms...\n> >\n> > At what point do we migrate from SHA-1?  At this point the cryptoanalysis of\n> > SHA-1 is most likely a matter of time.\n> \n> And I thought the cryptographic properties of SHA1 did not matter for\n> Gits use case.\n> We could employ broken md5 or such as well.\n> ( see http://stackoverflow.com/questions/28792784/why-does-git-use-a-cryptographic-hash-function\n> )\n> That is because security goes on top via gpg signing of tags/commits.\n> \n> I am not sure if anyone came up with\n> a counter argument to Linus reasoning there?\n\nI have never understood that reasoning at all, nor why it is so often\nrepeated.\n\nThe GPG signature is over a single object, that mentions other objects\nby their sha1 ids. But users don't care that v1.0 is securely mapped to\ntree 1234abcd. They care which files are in 1234abcd, and if sha1 is\nbroken, it means you can't credibly verify the content down to the blob\nlevel.\n\nThere's some additional protection in that git generally prefers objects\nit already has to new ones. So it's hard to reliably distribute your\nevil colliding object, depending on where people might have fetched\nfrom first. But:\n\n  1. I know there's at least once race[1] where a colliding object can\n     still enter the repository. There may be more that have either\n     existed all along, or that have grown over the years. I don't think\n     this is something we've paid attention to and tested.\n\n  2. That helps some people, I guess, but it's little consolation to\n     somebody who runs \"git clone\" followed by verifying the tag.\n\n-Peff\n\n[1] The race I am thinking of is that for performance reasons, we don't\n    re-scan the pack directory when index-pack checks has_sha1_file()\n    on an incoming object and it comes up negative. So if somebody else\n    is repacking, we might skip the collision check in such a case. At\n    least that race is not under control of an attacker, though.\n"},{"id":"283244","messageId":"1460502934.5540.71.camel@twopensource.com","threadId":"42005","inReplyTo":"CAGZ79kaUN0G7i0GNZgWU7ZzJvWY=k=Rc6tqWvJsTu8gcRhP5bA@mail.gmail.com","subject":"Re: Migrating away from SHA-1?","fromName":"David Turner","fromEmail":"dturner@twopensource.com","sentAt":"2016-04-12T23:15:34Z","receivedAt":"2016-04-12T23:15:34Z","isPatch":false,"sender":{"key":"novalis@novalis.org","avatar":"https://avatars.githubusercontent.com/u/77003?v=4"},"body":"On Tue, 2016-04-12 at 16:00 -0700, Stefan Beller wrote:\n> On Tue, Apr 12, 2016 at 3:38 PM, H. Peter Anvin <hpa@zytor.com>\n> wrote:\n> > OK, I'm going to open this can of worms...\n> > \n> > At what point do we migrate from SHA-1?  At this point the\n> > cryptoanalysis of\n> > SHA-1 is most likely a matter of time.\n> \n> And I thought the cryptographic properties of SHA1 did not matter for\n> Gits use case.\n> We could employ broken md5 or such as well.\n> ( see http://stackoverflow.com/questions/28792784/why-does-git-use-a-\n> cryptographic-hash-function\n> )\n> That is because security goes on top via gpg signing of tags/commits.\n> \n> I am not sure if anyone came up with\n> a counter argument to Linus reasoning there?\n\nHere's my reasoning as to why the security of SHA1 matters:\n\nIf SHA-1 is not broken, and someone hacks into e.g. kernel.org, they\ncan't replace an arbitrary blob with anything else without being\ndetected by git's automatic checksumming of objects.  GPG is necessary\nhere because otherwise the HEAD commit could be changed (to point to a\nnew tree that points to the new blob). \n\nIf SHA-1 is broken (in certain ways), someone *can* replace an\narbitrary blob.  GPG does not help in this case, because the signature\nis over the commit object (which points to a tree, which eventually\npoints to the blob), and the commit hasn't changed.  So the GPG\nsignature will still verify.\n\nIt would be possible, of course, to GPG-sign the entire commit's\ntransitive data (rather than just the SHA1s of same).  But as far as I\nknow, that is not ever what is done.\n\nThis is the argument for migration to a more-secure hash.\n"},{"id":"283250","messageId":"20160412234251.GB2210@sigill.intra.peff.net","threadId":"42005","inReplyTo":"570D78CC.9030807@zytor.com","subject":"Re: Migrating away from SHA-1?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-04-12T23:42:52Z","receivedAt":"2016-04-12T23:42:52Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 12, 2016 at 03:38:04PM -0700, H. Peter Anvin wrote:\n\n> For existing repositories we will need to have a migration mechanism. Since\n> we can't modify objects without completely invalidating the cryptographic\n> properties, what I would suggest is that we leave the existing objects as\n> is, with a persistent lookup table from SHA-1 to <new hash>, and have that\n> lookup table signed (e.g. GPG) by the person responsible for converting the\n> repository.  This freezes the cryptographic status of the existing SHA-1\n> objects at the time the conversion happens.  This is a very good reason to\n> do this before SHA-1 is actually broken  In contrast. SHA-2 has been\n> surprisingly resistant to cryptoanalysis, to the point that SHA-3 was\n> motivated by performance and the desire to have a well-tested function based\n> on entirely different principles should a generic attack against the common\n> structure of MD5/SHA-1/SHA-2 would ever be found.\n\nThere are a few threads in the list archive discussing options, if you\nsearch.\n\nA conversion table like you mention seems like a \"step 2\". I think the\nfirst step is figuring out what the new format looks like, and how\nobjects refer to each other.\n\nThe absolute simplest thing that could work is literally replacing sha1\nwith a 160-bit truncation of sha-256, telling everybody to convert their\nrepos, and accepting that existing gpg signatures and external sha1\nreferences are all obsolete. Old versions of git are obsolete, but the\ncode changes are very minor.\n\nThat sucks for a lot of reasons, obviously.\n\nSo a slightly nicer thing is to parameterize the algorithm for every\nobject name reference. So commits look like:\n\n  tree sha256:1234abcd...\n  parent sha256:1234abcd...\n\nand so on. Of course trees don't have any space for this; they have a\nfixed-length for the hash part of each record, which is basically:\n\n  <mode> <name> NUL <20-byte-sha1>\n\nSo we'd probably need a \"treev2\" object type that gives room for an\nalgorithm byte (or we'd have to try to shove it into the mode, but since\nold versions won't know the new algorithm anyway, I don't think it\nsolves that much...). Or you can just define for the whole tree object\n(either implicit in its type, or in a header) that it always uses\nalgorithm X.\n\nAnd then the \"new\" objects can refer to the older sha1 objects directly\n(either via \"sha1:1234abcd\", or we'd probably define a parameter-less\nreference to mean \"sha1:\"), and that essentially grafts the old history\nto the new. You can always walk the old history. And because we're\nreally talk about collision attacks and not pre-image attacks, it\nprobably remains fairly trustworthy for chaining (because nobody is\nmaking _new_ objects and referring to them via sha1).\n\nAnd then if you buy into the collision vs pre-image thing above, there's\nnot much point in caring about the mapping between sha1 and the new\nalgorithm. The old ones are set in stone and probably fine. You might\nwant such a mapping for performance (e.g., so that you can immediately\ntell that an old sha-1 tree and a new sha-2 tree have an empty diff,\neven though they have different ids), but that's purely a local thing.\n\nSo perhaps you were thinking of something in between, or an alternative\nplan altogether.  I haven't been able to think of a scheme that is\nsecure, convenient, and involves less work than the one above.\n\nTransitioning to that would be something like:\n\n  0. Overhaul all of the git code to handle arbitrary-sized object ids.\n\n  1. Decide on the new algorithm and implement it in git.\n\n  2. Recognize parameterized object ids in commits and tags (designing\n     format, implementing the reading side).\n\n  3. Recognize parameterized object ids somehow in trees (designing\n     format, implementing the reading side).\n\n  4. Teach the object database to index objects by the new algorithm (or\n     possibly both algorithms).\n\n  5. Add a protocol extension so that both sides can decide which\n     algorithm is being used when they talk about oids.\n\n  6. Add a config option to write references in objects using the new\n     algorithm.\n\n  7. After a while, flip the config option on. Hopefully the readers\n     from steps 1-5 have percolated to the masses by then, and it's not\n     a horrible flag day.\n\nWe're basically on step 0 right now. I'm sure I'm missing some\nsubtleties in there, too.\n\nThings get simpler if you don't fully parameterize (e.g., just assume\neverything is moved to the new algorithm, and provide a \"legacy\" parent\npointer for connecting to sha1 history). But part of this would be\nfuture-proofing for a day when sha-2 fails.\n\n-Peff\n"},{"id":"283251","messageId":"20160412234406.GC2210@sigill.intra.peff.net","threadId":"42005","inReplyTo":"1460502934.5540.71.camel@twopensource.com","subject":"Re: Migrating away from SHA-1?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-04-12T23:44:07Z","receivedAt":"2016-04-12T23:44:07Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 12, 2016 at 07:15:34PM -0400, David Turner wrote:\n\n> It would be possible, of course, to GPG-sign the entire commit's\n> transitive data (rather than just the SHA1s of same).  But as far as I\n> know, that is not ever what is done.\n\nThere is a project called git-evtag which does this, and you can find\nmention on the list. The problem is just that it's not very efficient.\nThat's maybe OK for tag-signing, which is relatively rare. It wouldn't\nreally work for commit-signing.\n\n-Peff\n"},{"id":"283277","messageId":"xmqqlh4imibd.fsf@gitster.mtv.corp.google.com","threadId":"42005","inReplyTo":"20160412234251.GB2210@sigill.intra.peff.net","subject":"Re: Migrating away from SHA-1?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-04-13T01:03:02Z","receivedAt":"2016-04-13T01:03:02Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> So a slightly nicer thing is to parameterize the algorithm for every\n> object name reference. So commits look like:\n>\n>   tree sha256:1234abcd...\n>   parent sha256:1234abcd...\n>\n> and so on. Of course trees don't have any space for this; they have a\n> fixed-length for the hash part of each record, which is basically:\n>\n>   <mode> <name> NUL <20-byte-sha1>\n>\n> So we'd probably need a \"treev2\" object type that gives room for an\n> algorithm byte (or we'd have to try to shove it into the mode, but since\n> old versions won't know the new algorithm anyway, I don't think it\n> solves that much...). Or you can just define for the whole tree object\n> (either implicit in its type, or in a header) that it always uses\n> algorithm X.\n\nThis will hurt the performance a lot during the transition period as\nit no longer will be possible to rely on \"most of the time a fine\ngrained commit changes only a small part of the tree, and we can\ncheaply avoid descending into trees that haven't changed because we\ncan tell that the corresponding tree objects in the pre- and post-\ntrees have the same object name\" optimization.  But we cannot avoid\nit.\n\n> Transitioning to that would be something like:\n>\n>   0. Overhaul all of the git code to handle arbitrary-sized object ids.\n>\n>   1. Decide on the new algorithm and implement it in git.\n>\n>   2. Recognize parameterized object ids in commits and tags (designing\n>      format, implementing the reading side).\n>\n>   3. Recognize parameterized object ids somehow in trees (designing\n>      format, implementing the reading side).\n>\n>   4. Teach the object database to index objects by the new algorithm (or\n>      possibly both algorithms).\n>\n>   5. Add a protocol extension so that both sides can decide which\n>      algorithm is being used when they talk about oids.\n>\n>   6. Add a config option to write references in objects using the new\n>      algorithm.\n>\n>   7. After a while, flip the config option on. Hopefully the readers\n>      from steps 1-5 have percolated to the masses by then, and it's not\n>      a horrible flag day.\n>\n> We're basically on step 0 right now. I'm sure I'm missing some\n> subtleties in there, too.\n\nOne subtlety is that 7. \"not a flag day\" may not be a good thing.\n\nThere has to be a section of a history that spans the transition,\nset of commits and trees that have pointers to both kinds of object\nnames.  The narrower such a section of the history, the more\npleasant to use the result of the transition would be.\n\nDifferent projects that can have their own flag days at their own\npace is a good thing, so the above observation does not invalidate\nyour transition plan, though.\n"},{"id":"283283","messageId":"20160413013632.GA10656@sigill.intra.peff.net","threadId":"42005","inReplyTo":"xmqqlh4imibd.fsf@gitster.mtv.corp.google.com","subject":"Re: Migrating away from SHA-1?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-04-13T01:36:32Z","receivedAt":"2016-04-13T01:36:32Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Tue, Apr 12, 2016 at 06:03:02PM -0700, Junio C Hamano wrote:\n\n> > So we'd probably need a \"treev2\" object type that gives room for an\n> > algorithm byte (or we'd have to try to shove it into the mode, but since\n> > old versions won't know the new algorithm anyway, I don't think it\n> > solves that much...). Or you can just define for the whole tree object\n> > (either implicit in its type, or in a header) that it always uses\n> > algorithm X.\n> \n> This will hurt the performance a lot during the transition period as\n> it no longer will be possible to rely on \"most of the time a fine\n> grained commit changes only a small part of the tree, and we can\n> cheaply avoid descending into trees that haven't changed because we\n> can tell that the corresponding tree objects in the pre- and post-\n> trees have the same object name\" optimization.  But we cannot avoid\n> it.\n\nYeah. I'd hope in general that there would be a single commit that does\nthe transition, and we'd only pay it when doing diffs across the\nboundary. And even then, I think a local-only cache of aliases could\nmitigate the worst of it.\n\n> >   7. After a while, flip the config option on. Hopefully the readers\n> >      from steps 1-5 have percolated to the masses by then, and it's not\n> >      a horrible flag day.\n> >\n> > We're basically on step 0 right now. I'm sure I'm missing some\n> > subtleties in there, too.\n> \n> One subtlety is that 7. \"not a flag day\" may not be a good thing.\n> \n> There has to be a section of a history that spans the transition,\n> set of commits and trees that have pointers to both kinds of object\n> names.  The narrower such a section of the history, the more\n> pleasant to use the result of the transition would be.\n> \n> Different projects that can have their own flag days at their own\n> pace is a good thing, so the above observation does not invalidate\n> your transition plan, though.\n\nGood point. I do think projects would do well to have a moment where\nthey switch to the new format, and don't freely intermingle. We could\npossibly do some magic there to help things out. For example, if we are\nbuilding on a commit that is sha-2, we automatically use more sha-2\nobjects to point to them. And then the \"flag day\" for a project is\nsimply that somebody pushes to \"master\" using sha-2, and everybody\nelse's git (which learned long ago to speak the new algorithm) just\npicks it up.\n\nOf course that's not exactly a flag day for projects that branch from\nold history for bugfixes. But it might be close enough.\n\n-Peff\n"},{"id":"283285","messageId":"570DA311.3000500@zytor.com","threadId":"42005","inReplyTo":"xmqqlh4imibd.fsf@gitster.mtv.corp.google.com","subject":"Re: Migrating away from SHA-1?","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2016-04-13T01:38:25Z","receivedAt":"2016-04-13T01:38:25Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"On 04/12/16 18:03, Junio C Hamano wrote:\n>>\n>> and so on. Of course trees don't have any space for this; they have a\n>> fixed-length for the hash part of each record, which is basically:\n>>\n>>    <mode> <name> NUL <20-byte-sha1>\n>>\n>> So we'd probably need a \"treev2\" object type that gives room for an\n>> algorithm byte (or we'd have to try to shove it into the mode, but since\n>> old versions won't know the new algorithm anyway, I don't think it\n>> solves that much...). Or you can just define for the whole tree object\n>> (either implicit in its type, or in a header) that it always uses\n>> algorithm X.\n>\n> This will hurt the performance a lot during the transition period as\n> it no longer will be possible to rely on \"most of the time a fine\n> grained commit changes only a small part of the tree, and we can\n> cheaply avoid descending into trees that haven't changed because we\n> can tell that the corresponding tree objects in the pre- and post-\n> trees have the same object name\" optimization.  But we cannot avoid\n> it.\n>\n\nNot really, because you can point to the algoX hash even for the \nexisting objects.\n\nPerhaps the tree object can add a format descriptor at the beginning; \nsomething like:\n\n<invalid mode number> <hash format used>\n\n>> Transitioning to that would be something like:\n>>\n>>    0. Overhaul all of the git code to handle arbitrary-sized object ids.\n>>\n>>    1. Decide on the new algorithm and implement it in git.\n>>\n>>    2. Recognize parameterized object ids in commits and tags (designing\n>>       format, implementing the reading side).\n>>\n>>    3. Recognize parameterized object ids somehow in trees (designing\n>>       format, implementing the reading side).\n>>\n>>    4. Teach the object database to index objects by the new algorithm (or\n>>       possibly both algorithms).\n>>\n>>    5. Add a protocol extension so that both sides can decide which\n>>       algorithm is being used when they talk about oids.\n>>\n>>    6. Add a config option to write references in objects using the new\n>>       algorithm.\n>>\n>>    7. After a while, flip the config option on. Hopefully the readers\n>>       from steps 1-5 have percolated to the masses by then, and it's not\n>>       a horrible flag day.\n>>\n>> We're basically on step 0 right now. I'm sure I'm missing some\n>> subtleties in there, too.\n>\n> One subtlety is that 7. \"not a flag day\" may not be a good thing.\n>\n> There has to be a section of a history that spans the transition,\n> set of commits and trees that have pointers to both kinds of object\n> names.  The narrower such a section of the history, the more\n> pleasant to use the result of the transition would be.\n>\n> Different projects that can have their own flag days at their own\n> pace is a good thing, so the above observation does not invalidate\n> your transition plan, though.\n\nI don't think there is any way this can *not* be by repository and \nsomehow require a manual operation in order to preserve the \ncryptographic integrity.  In some ways, the transition point and the \ntransition table becomes a special kind of tag object.  There may have \nto be more than one in the case of commits in multiple trees.\n"},{"id":"283287","messageId":"CACsJy8DmPw+cbohp-X55bp9NJSbUVN=tsABXoF5Xh-6PgPTbiA@mail.gmail.com","threadId":"42005","inReplyTo":"570D78CC.9030807@zytor.com","subject":"Re: Migrating away from SHA-1?","fromName":"Duy Nguyen","fromEmail":"pclouds@gmail.com","sentAt":"2016-04-13T01:51:12Z","receivedAt":"2016-04-13T01:51:12Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On Wed, Apr 13, 2016 at 5:38 AM, H. Peter Anvin <hpa@zytor.com> wrote:\n> OK, I'm going to open this can of worms...\n>\n> At what point do we migrate from SHA-1?\n\nBrian Carlson has been slowly refactoring git code base, abstracting\nSHA-1 away. Once that work is done, I think we can talk about moving\naway from SHA-1. The process is slow because it likely causes\nconflicts with in-flight topics. A quick grep shows we still have\nabout 300 SHA-1 references, so it'll be quite some time.\n-- \nDuy\n"},{"id":"283288","messageId":"8722D9F3-8A42-4BF7-A945-305F483E8364@zytor.com","threadId":"42005","inReplyTo":"CACsJy8DmPw+cbohp-X55bp9NJSbUVN=tsABXoF5Xh-6PgPTbiA@mail.gmail.com","subject":"Re: Migrating away from SHA-1?","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2016-04-13T01:58:10Z","receivedAt":"2016-04-13T01:58:10Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"On April 12, 2016 6:51:12 PM PDT, Duy Nguyen <pclouds@gmail.com> wrote:\n>On Wed, Apr 13, 2016 at 5:38 AM, H. Peter Anvin <hpa@zytor.com> wrote:\n>> OK, I'm going to open this can of worms...\n>>\n>> At what point do we migrate from SHA-1?\n>\n>Brian Carlson has been slowly refactoring git code base, abstracting\n>SHA-1 away. Once that work is done, I think we can talk about moving\n>away from SHA-1. The process is slow because it likely causes\n>conflicts with in-flight topics. A quick grep shows we still have\n>about 300 SHA-1 references, so it'll be quite some time.\n\nWell, at least it sounds like work is underway.  That is a big deal.\n-- \nSent from my Android device with K-9 Mail. Please excuse brevity and formatting.\n"},{"id":"283434","messageId":"20160414015324.GA16656@thunk.org","threadId":"42005","inReplyTo":"1460502934.5540.71.camel@twopensource.com","subject":"Re: Migrating away from SHA-1?","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2016-04-14T01:53:24Z","receivedAt":"2016-04-14T01:53:24Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Tue, Apr 12, 2016 at 07:15:34PM -0400, David Turner wrote:\n> \n> If SHA-1 is broken (in certain ways), someone *can* replace an\n> arbitrary blob.  GPG does not help in this case, because the signature\n> is over the commit object (which points to a tree, which eventually\n> points to the blob), and the commit hasn't changed.  So the GPG\n> signature will still verify.\n\nThe \"in certain ways\" is the critical bit.  The question is whether\nyou are trying to replace an arbitrary blob, or a blob that was\nsubmitted under your control.\n\nIf you are trying to replace an arbitrary blob under the you need to\ncarry a preimage attack.  That means that given a particular hash, you\nneed to find another blob that has the same hash.  SHA-1 is currently\nresistant against preimage attack (that is, you need to use brute\nforce, so the work factor is 2**159).  \n\nIf you are trying to replace an arbitrary blob which is under your\ncontrol, then all you need is a collision attack, and this is where\nSHA-1 has been weakened.  It is now possible to find a collision with\na work factor of 2**69, instead of the requisite 2**80.\n\nIt was a MD5 collision which was involved with the Flame attack.\nSomeone (in probably the US or Isreali intelligence services)\nsubmitted a Certificate Signing Request (CSR) to the Microsoft\nTerminal Services Licensing server.  That CSR was under the control of\nthe attacker, and it resulted in a certificate where parts of the\ncertificate could be swapped out with the corresponding fields from\nanother CSR (which was not submitted to the Certifiying Authority)\nwhich had the code signing bit set.\n\nSo in order to carry out this attack, not only did the (cough)\n\"unknown\" attackers had to have come up with a collision, but the two\npieces of colliding blobs had to parsable a valid CSR's, one which had\nto pass inspection by the automated CA signing authority, and the\nother which had to contain the desired code signing bits set so the\nattacker could sabotage an Iranian nuclear centrifuge.\n\nOK, so how does this map to git?  First of all, from a collision\nperspective, the two blobs have to map into valid C code, one of which\nhas to be innocuous enough such that any humans who review the patch\nand/or git pull request don't notice anything wrong.  The second has\nto contain whatever security backdoor the attacker is going to try to\nintroduce into the git tree.  Ideally this is also should pass muster\nby humans who are inspecting the code, but if the attack is targetted\nagainst a specific victim which is not likely to look at the code, it\nmight be okay if something like this:\n\n#if 0  /* this is needed to make the hash collision work */\naev2Ein4Hagh8eimshood5aTeteiVo9hOhchohN6jiem6AiNEipeeR3Pie4ePaeJ\nfo8eLa9ateeKie5VeG5eZuu2Sahqu1Ohai9ohGhuAevoot5OtohQuai7koo4IeTh\nohCefae4Ahkah0eiku2Efo0iuHai8ideaRooth8wVahlia0nuu1eeSh5oht1Kaer\naiJi4chunahK9oozpaiWu7viee5aiFahud6Ee2zieich1veKque6PhiaAit1shie\n#endif\n\n... was hidden in the middle of the replacement blob.  One would\n*hope*, though, that if something like this appeared in a blob that\nwas being sent to the upstream repository, that even a sloppy github\npull request reviewer would notice.\n\nThat's because in this scenario, the attacker needs to be able to get\nthe first blob into the git tree first, which means they need to be\ntrusted enough to get the first blob in.  And so the question which\ncomes to mind is if you are that trusted (or if the git pull review\nprocess is that crappy), might it not be easier to simply introduce an\nobfuscated code that has a security weakness?  That is, something from\nthe Underhanded C contest, or an accidental buffer overrun, hopefully\none that isn't noticed by static code checkers.  If you do that, you\ndon't even need to figure out how to create a SHA-1 collision.\n\nDoes that mean that we shouldn't figure out how to migrate to another\nhash function?  No, it's probably worth planning how to do it.  But we\nprobably have a fair amount of time to get this right.\n\nCheers,\n\n\t\t\t\t\t- Ted\n"},{"id":"283457","messageId":"20160414164751.GA3255@kitenet.net","threadId":"42005","inReplyTo":"20160414015324.GA16656@thunk.org","subject":"Re: Migrating away from SHA-1?","fromName":"Joey Hess","fromEmail":"id@joeyh.name","sentAt":"2016-04-14T16:47:51Z","receivedAt":"2016-04-14T16:47:51Z","isPatch":false,"sender":{"key":"id@joeyh.name","avatar":"https://avatars.githubusercontent.com/u/16392?v=4"},"body":"Theodore Ts'o wrote:\n> OK, so how does this map to git?  First of all, from a collision\n> perspective, the two blobs have to map into valid C code\n\nGit provides other places to hide the colliding blobs; the best seems to\nbe as an added header in the commit object, or as trailing data after a \\0\nin the commit message. git is very good at hiding such potentially\ncolliding data from the user, as https://github.com/joeyh/supercollider\ndemonstrates.\n\ncommit 24f30db5790b209fa412ce81c5ef2bf8af5fd4d7\nAuthor: Joey Hess <joey@kitenet.net>\nDate:   Fri Sep 9 11:49:21 2011 -0400\n\n    an innocent commit\n    \n    If this were a sha1 colliding attack, there would be some sort of binary\n    garbage below. Which there isn't. So this can be safely merged.\n\njoey@darkstar:~/tmp/supercollider>git cat-file -p 24f30db5790b209fa412ce81c5ef2bf8af5fd4d7\ntree 735a7633237c07b398856005de3bc9ea00446747\nauthor Joey Hess <joey@kitenet.net> 1315583361 -0400\ncommitter Joey Hess <joey@kitenet.net> 1315583361 -0400\n\nan innocent commit\n\nIf this were a sha1 colliding attack, there would be some sort of binary\ngarbage below. Which there isn't. So this can be safely merged.\n\n\n\n??b???\u001f[?i??ͯ?t?\f2??\u0002????os?\u0014<????h?+,M?mY?e?EW?i\u0013v$???\u0014J??U}n~???L??????f??\u0002?ě??3>?Q??H?޸\u0016*zl\u001a?RA˂q?E\f?\u0006\u0016E7??\u001b?\u0003\\?m???U?\u001e>MU\u000b\tGY?d)?ȼ??'g?~D??ɯhQ?\u0013???/\"E\u0004??X?m???^͸??S?D\u0013??;w6(?`??>?\u0010縘?\u0007AѲ?*!??@v????>?8??2\b?\u0014!??=*?J\t\u001b\r\r???\u0001ynH\u0010???c?w?\\??K7??\u001c?N?6??\u001c???A5?FM?wZ?~?pK\u0002Y?R???s7??(?\u0007ƶ?_\"??m\u0011%????1a??ʀ??K[\rt??\u0011??\u000e!A0?ΈfT.?T?w\u0007?򁛵ƌ\u000b?р???aco?V/2\u0014??nَ?\n?}?6?\u0019_?z?{\n\n\n(The other possibility would be to hide the colliding blob in the tree\nobject, but that seems unlikely.)\n\n-- \nsee shy jo\n"},{"id":"283460","messageId":"1460654583.5540.87.camel@twopensource.com","threadId":"42005","inReplyTo":"20160414015324.GA16656@thunk.org","subject":"Re: Migrating away from SHA-1?","fromName":"David Turner","fromEmail":"dturner@twopensource.com","sentAt":"2016-04-14T17:23:03Z","receivedAt":"2016-04-14T17:23:03Z","isPatch":false,"sender":{"key":"novalis@novalis.org","avatar":"https://avatars.githubusercontent.com/u/77003?v=4"},"body":"On Wed, 2016-04-13 at 21:53 -0400, Theodore Ts'o wrote:\n> On Tue, Apr 12, 2016 at 07:15:34PM -0400, David Turner wrote:\n> > \n> > If SHA-1 is broken (in certain ways), someone *can* replace an\n> > arbitrary blob.  GPG does not help in this case, because the\n> > signature\n> > is over the commit object (which points to a tree, which eventually\n> > points to the blob), and the commit hasn't changed.  So the GPG\n> > signature will still verify.\n> \n> The \"in certain ways\" is the critical bit.  The question is whether\n> you are trying to replace an arbitrary blob, or a blob that was\n> submitted under your control.\n> \n> If you are trying to replace an arbitrary blob under the you need to\n> carry a preimage attack.  That means that given a particular hash,\n> you\n> need to find another blob that has the same hash.  SHA-1 is currently\n> resistant against preimage attack (that is, you need to use brute\n> force, so the work factor is 2**159).  \n> \n> If you are trying to replace an arbitrary blob which is under your\n> control, then all you need is a collision attack, and this is where\n> SHA-1 has been weakened.  It is now possible to find a collision with\n> a work factor of 2**69, instead of the requisite 2**80.\n> \n> It was a MD5 collision which was involved with the Flame attack.\n> Someone (in probably the US or Isreali intelligence services)\n> submitted a Certificate Signing Request (CSR) to the Microsoft\n> Terminal Services Licensing server.  That CSR was under the control\n> of\n> the attacker, and it resulted in a certificate where parts of the\n> certificate could be swapped out with the corresponding fields from\n> another CSR (which was not submitted to the Certifiying Authority)\n> which had the code signing bit set.\n> \n> So in order to carry out this attack, not only did the (cough)\n> \"unknown\" attackers had to have come up with a collision, but the two\n> pieces of colliding blobs had to parsable a valid CSR's, one which\n> had\n> to pass inspection by the automated CA signing authority, and the\n> other which had to contain the desired code signing bits set so the\n> attacker could sabotage an Iranian nuclear centrifuge.\n> \n> OK, so how does this map to git?  First of all, from a collision\n> perspective, the two blobs have to map into valid C code, one of\n> which\n> has to be innocuous enough such that any humans who review the patch\n> and/or git pull request don't notice anything wrong.  \n\nIt looks like Linux contains at least some firmware which would be hard\nto audit.  One random example is:\nfirmware/bnx2x/bnx2x-e1h-6.2.9.0.fw.ihex\n"},{"id":"283461","messageId":"71A5D062-FCCD-42E5-80A8-AA9D8DE20604@zytor.com","threadId":"42005","inReplyTo":"1460654583.5540.87.camel@twopensource.com","subject":"Re: Migrating away from SHA-1?","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2016-04-14T17:28:50Z","receivedAt":"2016-04-14T17:28:50Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"On April 14, 2016 10:23:03 AM PDT, David Turner <dturner@twopensource.com> wrote:\n>On Wed, 2016-04-13 at 21:53 -0400, Theodore Ts'o wrote:\n>> On Tue, Apr 12, 2016 at 07:15:34PM -0400, David Turner wrote:\n>> > \n>> > If SHA-1 is broken (in certain ways), someone *can* replace an\n>> > arbitrary blob.  GPG does not help in this case, because the\n>> > signature\n>> > is over the commit object (which points to a tree, which eventually\n>> > points to the blob), and the commit hasn't changed.  So the GPG\n>> > signature will still verify.\n>> \n>> The \"in certain ways\" is the critical bit.  The question is whether\n>> you are trying to replace an arbitrary blob, or a blob that was\n>> submitted under your control.\n>> \n>> If you are trying to replace an arbitrary blob under the you need to\n>> carry a preimage attack.  That means that given a particular hash,\n>> you\n>> need to find another blob that has the same hash.  SHA-1 is currently\n>> resistant against preimage attack (that is, you need to use brute\n>> force, so the work factor is 2**159).  \n>> \n>> If you are trying to replace an arbitrary blob which is under your\n>> control, then all you need is a collision attack, and this is where\n>> SHA-1 has been weakened.  It is now possible to find a collision with\n>> a work factor of 2**69, instead of the requisite 2**80.\n>> \n>> It was a MD5 collision which was involved with the Flame attack.\n>> Someone (in probably the US or Isreali intelligence services)\n>> submitted a Certificate Signing Request (CSR) to the Microsoft\n>> Terminal Services Licensing server.  That CSR was under the control\n>> of\n>> the attacker, and it resulted in a certificate where parts of the\n>> certificate could be swapped out with the corresponding fields from\n>> another CSR (which was not submitted to the Certifiying Authority)\n>> which had the code signing bit set.\n>> \n>> So in order to carry out this attack, not only did the (cough)\n>> \"unknown\" attackers had to have come up with a collision, but the two\n>> pieces of colliding blobs had to parsable a valid CSR's, one which\n>> had\n>> to pass inspection by the automated CA signing authority, and the\n>> other which had to contain the desired code signing bits set so the\n>> attacker could sabotage an Iranian nuclear centrifuge.\n>> \n>> OK, so how does this map to git?  First of all, from a collision\n>> perspective, the two blobs have to map into valid C code, one of\n>> which\n>> has to be innocuous enough such that any humans who review the patch\n>> and/or git pull request don't notice anything wrong.  \n>\n>It looks like Linux contains at least some firmware which would be hard\n>to audit.  One random example is:\n>firmware/bnx2x/bnx2x-e1h-6.2.9.0.fw.ihex\n\nEither way, I agree with Ted, that we have enough time to do it right, but that is a good reason to do it sooner rather than later (see also my note about freezing the cryptographic properties.)\n-- \nSent from my Android device with K-9 Mail. Please excuse brevity and formatting.\n"},{"id":"283497","messageId":"20160414224051.GD16656@thunk.org","threadId":"42005","inReplyTo":"71A5D062-FCCD-42E5-80A8-AA9D8DE20604@zytor.com","subject":"Re: Migrating away from SHA-1?","fromName":"Theodore Ts'o","fromEmail":"tytso@mit.edu","sentAt":"2016-04-14T22:40:51Z","receivedAt":"2016-04-14T22:40:51Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Thu, Apr 14, 2016 at 10:28:50AM -0700, H. Peter Anvin wrote:\n> \n> Either way, I agree with Ted, that we have enough time to do it\n> right, but that is a good reason to do it sooner rather than later\n> (see also my note about freezing the cryptographic properties.)\n\nSure, I think we should do it as well.  But the fact that the attacker\nwill likely need to get a commit into the tree in order to be able to\ncarry out a collision attack means that it's easier (and probably less\ndetectable) to get some underhanded C code into the tree.  For one\nthing, you just need to introduce it via a patch (\"Hi, I'm super eager\nnewbie Nick, here's a cleanup patch!\"), as opposed to getting a\nsublieutenant to accept a git pull request.\n\nAlso, remember that while we can write programs that look for\nsuspicious git objects that have stuff hidden after the null\nterminator (in fact, maybe that would be a good thing to add to git,\nhmmm?), the state of the art in detecting underhanded C code which is\ndeliberately designed to not be noticed by static code checkers (or\nhumans doing a superficial code review, for that matter) is not\nparticularly encouraging to me.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"283716","messageId":"20160415015004.GB140502@vauxhall.crustytoothpaste.net","threadId":"42005","inReplyTo":"8722D9F3-8A42-4BF7-A945-305F483E8364@zytor.com","subject":"Re: Migrating away from SHA-1?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2016-04-15T01:50:05Z","receivedAt":"2016-04-15T01:50:05Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Tue, Apr 12, 2016 at 06:58:10PM -0700, H. Peter Anvin wrote:\n> On April 12, 2016 6:51:12 PM PDT, Duy Nguyen <pclouds@gmail.com> wrote:\n> >On Wed, Apr 13, 2016 at 5:38 AM, H. Peter Anvin <hpa@zytor.com> wrote:\n> >> OK, I'm going to open this can of worms...\n> >>\n> >> At what point do we migrate from SHA-1?\n> >\n> >Brian Carlson has been slowly refactoring git code base, abstracting\n> >SHA-1 away. Once that work is done, I think we can talk about moving\n> >away from SHA-1. The process is slow because it likely causes\n> >conflicts with in-flight topics. A quick grep shows we still have\n> >about 300 SHA-1 references, so it'll be quite some time.\n> \n> Well, at least it sounds like work is underway.  That is a big deal.\n\nYes, it's a bunch of slow manual refactoring, and I've been busy as\nwe've been doing house- and car-related things recently.  I'll try to\nspend a little more time on it this weekend.\n\nThe first step is to convert all of the individual places that use\nunsigned char [20] to use struct object_id, which can then be extended\nto use different hash algorithms.  There are also constants,\nGIT_SHA1_RAWSZ and GIT_SHA1_HEXSZ, that abstract the 20 and 40 values in\nthe codebase so they can be changed in the future.\n\nWhile this is a project I've been mostly working on, I have no objection\nto other people sending in a patch or series as they feel like it.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"283517","messageId":"20160415021327.GC22112@sigill.intra.peff.net","threadId":"42005","inReplyTo":"20160414224051.GD16656@thunk.org","subject":"Re: Migrating away from SHA-1?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-04-15T02:13:27Z","receivedAt":"2016-04-15T02:13:27Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Apr 14, 2016 at 06:40:51PM -0400, Theodore Ts'o wrote:\n\n> Also, remember that while we can write programs that look for\n> suspicious git objects that have stuff hidden after the null\n> terminator (in fact, maybe that would be a good thing to add to git,\n> hmmm?)[...]\n\nDetecting the hidden bytes is underway elsewhere on the list. And while\nI think it's a good idea to do so, I don't think it really introduces\na meaningful defense against collision attacks.\n\nYou can also hide bytes in arbitrary headers in a git object[1], and\nthey will not be shown by default. Adding the extra bytes at the end is\ncertainly easier if you're micro-optimizing the collision process[2],\nbut I don't think it changes the fundamental equation. It reduces the\nwork you do per-sha1 by a constant factor, but not the number of sha1s\nyou expect to compute.\n\n-Peff\n\n[1] Obviously neither \"extra headers\" nor \"stuff after NUL\" applies to\n    patches sent by email, where everything short of binary-diffs is\n    human-readable. So for the kernel, you're really talking about\n    attacking a lieutenant whose repo gets pulled. But there are plenty\n    of other projects that \"git merge\" from strangers.\n\n[2] Somewhere in the list archive is my patch to find partial\n    collisions like \"git commit --sha1=31337\", and I did in fact use\n    that micro-optimization. That, along with multi-threading, made it\n    feasible to do 6-8 character prefixes, as I recall.\n"},{"id":"283519","messageId":"xmqq37qnehrm.fsf@gitster.mtv.corp.google.com","threadId":"42005","inReplyTo":"20160415021327.GC22112@sigill.intra.peff.net","subject":"Re: Migrating away from SHA-1?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2016-04-15T02:18:53Z","receivedAt":"2016-04-15T02:18:53Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n> [2] Somewhere in the list archive is my patch to find partial\n>     collisions like \"git commit --sha1=31337\", and I did in fact use\n>     that micro-optimization. That, along with multi-threading, made it\n>     feasible to do 6-8 character prefixes, as I recall.\n\nIn our testsuite, we have a test that uses many objects, all of\nwhich have object names that begin with 10 '0' characters.\n"},{"id":"283520","messageId":"20160415022233.GE22112@sigill.intra.peff.net","threadId":"42005","inReplyTo":"xmqq37qnehrm.fsf@gitster.mtv.corp.google.com","subject":"Re: Migrating away from SHA-1?","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2016-04-15T02:22:33Z","receivedAt":"2016-04-15T02:22:33Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Apr 14, 2016 at 07:18:53PM -0700, Junio C Hamano wrote:\n\n> Jeff King <peff@peff.net> writes:\n> \n> > [2] Somewhere in the list archive is my patch to find partial\n> >     collisions like \"git commit --sha1=31337\", and I did in fact use\n> >     that micro-optimization. That, along with multi-threading, made it\n> >     feasible to do 6-8 character prefixes, as I recall.\n> \n> In our testsuite, we have a test that uses many objects, all of\n> which have object names that begin with 10 '0' characters.\n\nCan you give more details on which test? 10 zeroes is 40 bits, which\nmeans that by random chance, only about one in a trillion objects would\nmatch that. We certainly didn't hit that randomly, and it seems like it\nwould be computationally expensive to have come up with the input for\neven one such object, let alone \"many\".\n\n-Peff\n"}]}