{"thread":{"id":"45288","subject":"RFC: Another proposed hash function transition plan","startedAt":"2017-03-04T01:43:24Z","lastAt":"2017-10-04T01:44:21Z","messageCount":113,"participants":["Jonathan Nieder","Linus Torvalds","David Lang","brian m. carlson","Jeff King","Brandon Williams","Junio C Hamano","Jonathan Tan","Mike Hommey","Ian Jackson","Johannes Schindelin","Shawn Pearce","The Keccak Team","ankostis","Jason Hennessey","Michael Steuer","Jacob Keller","Ævar Arnfjörð Bjarmason","Adam Langley","demerphq","Stefan Beller","Philip Oakley","Gilles Van Assche","Jason Cooper","Joan Daemen"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"313254","messageId":"20170304011251.GA26789@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":null,"subject":"RFC: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-03-04T01:12:51Z","receivedAt":"2017-03-04T01:43:24Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nThis past week we came up with this idea for what a transition to a new\nhash function for Git would look like.  I'd be interested in your\nthoughts (especially if you can make them as comments on the document,\nwhich makes it easier to address them and update the document).\n\nThis document is still in flux but I thought it best to send it out\nearly to start getting feedback.\n\nWe tried to incorporate some thoughts from the thread\nhttp://public-inbox.org/git/20170223164306.spg2avxzukkggrpb@kitenet.net\nbut it is a little long so it is easy to imagine we've missed\nsome things already discussed there.\n\nYou can use the doc URL\n\n https://goo.gl/gh2Mzc\n\nto view the latest version and comment.\n\nThoughts welcome, as always.\n\nGit hash function transition\n============================\nStatus: Draft\nLast Updated: 2017-03-03\n\nObjective\n---------\nMigrate Git from SHA-1 to a stronger hash function.\n\nBackground\n----------\nThe Git version control system can be thought of as a content\naddressable filesystem. It uses the SHA-1 hash function to name\ncontent. For example, files, trees, commits are referred to by hash\nvalues unlike in other traditional version control systems where files\nor versions are referred to via sequential numbers. The use of a hash\nfunction to address its content delivers a few advantages:\n\n* Integrity checking is easy. Bit flips, for example, are easily\n  detected, as the hash of corrupted content does not match its name.\n  Lookup of objects is fast.\n\nUsing a cryptographically secure hash function brings additional advantages:\n\n* Object names can be signed and third parties can trust the hash to\n  address the signed object and all objects it references.\n* Communication using Git protocol and out of band communication\n  methods have a short reliable string that can be used to reliably\n  address stored content.\n\nOver time some flaws in SHA-1 have been discovered by security\nresearchers. https://shattered.io demonstrated a practical SHA-1 hash\ncollision. As a result, SHA-1 cannot be considered cryptographically\nsecure any more. This impacts the communication of hash values because\nwe cannot trust that a given hash value represents the known good\nversion of content that the speaker intended.\n\nSHA-1 still possesses the other properties such as fast object lookup\nand safe error checking, but other hash functions are equally suitable\nthat are believed to be cryptographically secure.\n\nGoals\n-----\n1. The transition to SHA256 can be done one local repository at a time.\n   a. Requiring no action by any other party.\n   b. A SHA256 repository can communicate with SHA-1 Git servers and\n      clients (push/fetch).\n   c. Users can use SHA-1 and SHA256 identifiers for objects\n      interchangeably.\n   d. New signed objects make use of a stronger hash function than\n      SHA-1 for their security guarantees.\n2. Allow a complete transition away from SHA-1.\n   a. Local metadata for SHA-1 compatibility can be dropped in a\n      repository if compatibility with SHA-1 is no longer needed.\n3. Maintainability throughout the process.\n   a. The object format is kept simple and consistent.\n   b. Creation of a generalized repository conversion tool.\n\nNon-Goals\n---------\n1. Add SHA256 support to Git protocol. This is valuable and the\n   logical next step but it is out of scope for this initial design.\n2. Transparently improving the security of existing SHA-1 signed\n   objects.\n3. Intermixing objects using multiple hash functions in a single\n   repository.\n4. Taking the opportunity to fix other bugs in git's formats and\n   protocols.\n5. Shallow clones and fetches into a SHA256 repository. (This will\n   change when we add SHA256 support to Git protocol.)\n6. Skip fetching some submodules of a project into a SHA256\n   repository. (This also depends on SHA256 support in Git protocol.)\n\nOverview\n--------\nWe introduce a new repository format extension `sha256`. Repositories\nwith this extension enabled use SHA256 instead of SHA-1 to name their\nobjects. This affects both object names and object content --- both\nthe names of objects and all references to other objects within an\nobject are switched to the new hash function.\n\nsha256 repositories cannot be read by older versions of Git.\n\nAlongside the packfile, a sha256 stores a bidirectional mapping\nbetween sha256 and sha1 object names. The mapping is generated locally\nand can be verified using \"git fsck\". Object lookups use this mapping\nto allow naming objects using either their sha1 and sha256 names\ninterchangeably.\n\n\"git cat-file\" and \"git hash-object\" gain options to display a sha256\nobject in its sha1 form and write a sha256 object given its sha1 form.\nThis requires all objects referenced by that object to be present in\nthe object database so that they can be named using the appropriate\nname (using the bidirectional hash mapping).\n\nFetches from a SHA-1 based server convert the fetched objects into\nsha256 form and record the mapping in the bidirectional mapping table\n(see below for details). Pushes to a SHA-1 based server convert the\nobjects being pushed into sha1 form so the server does not have to be\naware of the hash function the client is using.\n\nDetailed Design\n---------------\nObject names\n~~~~~~~~~~~~\nObjects can be named by their 40 hexadecimal digit sha1-name or 64\nhexadecimal digit sha256-name, plus names derived from those (see\ngitrevisions(7)).\n\nThe sha1-name of an object is the SHA-1 of the concatenation of its\ntype, length, a nul byte, and the object's sha1-content. This is the\ntraditional <sha1> used in Git to name objects.\n\nThe sha256-name of an object is the SHA-256 of the concatenation of\nits type, length, a nul byte, and the object's sha256-content.\n\nObject format\n~~~~~~~~~~~~~\nObjects are stored using a compressed representation of their\nsha256-content. The sha256-content of an object is the same as its\nsha1-content, except that:\n* objects referenced by the object are named using their sha256-names\n  instead of sha1-names\n* signed tags, commits, and merges of signed tags get some additional\n  fields (see below)\n\nThe format allows round-trip conversion between sha256-content and\nsha1-content.\n\nLoose objects use zlib compression and packed objects use the packed\nformat described in Documentation/technical/pack-format.txt, just like\ntoday.\n\nTranslation table\n~~~~~~~~~~~~~~~~~\nA fast bidirectional mapping between sha1-names and sha256-names of\nall local objects in the repository is kept on disk. The exact format\nof that mapping is to be determined.\n\nAll operations that make new objects (e.g., \"git commit\") add the new\nobjects to the translation table.\n\nReading an object's sha1-content\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\nThe sha1-content of an object can be read by converting all\nsha256-names its sha256-content references to sha1-names using the\ntranslation table. There is an additional minor transformation needed\nfor signed tags, commits, and merges (see below).\n\nFetch\n~~~~~\nFetching from a SHA-1 based server requires translating between SHA-1\nand SHA-256 based representations on the fly.\n\nSHA-1s named in the ref advertisement can be translated to SHA-256 and\nlooked up as local objects using the translation table.\n\nNegotiation proceeds as today. Any \"have\"s or \"want\"s generated\nlocally are converted to SHA-1 before being sent to the server, and\nSHA-1s mentioned by the server are converted to SHA-256 when looking\nthem up locally.\n\nAfter negotiation, the server sends a packfile containing the\nrequested objects. We convert the packfile to SHA-256 format using the\nfollowing steps:\n\n1. index-pack: inflate each object in the packfile and compute its\n   SHA-1. Objects can contain deltas in OBJ_REF_DELTA format against\n   objects the client has locally. These objects can be looked up using\n   the translation table and their sha1-content read as described above\n   to resolve the deltas.\n2. topological sort: starting at the \"want\"s from the negotiation\n   phase, walk through objects in the pack and emit a list of them in\n   topologically sorted order. (This list only contains objects\n   reachable from the \"wants\". If the pack from the server contained\n   additional extraneous objects, then they will be discarded.)\n3. convert to sha256: open a new (sha256) packfile. Read the\n   topologically sorted list just generated in reverse order. For each\n   object, inflate its sha1-content, convert to sha256-content, and\n   write it to the sha256 pack. Write an idx file for this pack and\n   include the new sha1<->sha256 mapping entry in the translation\n   table.\n4. clean up: remove the SHA-1 based pack file, index, and\n   topologically sorted list obtained from the server and steps 1 and 2.\n\nStep 3 requires every object referenced by the new object to be in the\ntranslation table. This is why the topological sort step is necessary.\n\nAs an optimization, step 1 can write a file describing what objects\neach object it has inflated from the packfile references. This makes\nthe topological sort in step 2 possible without inflating the objects\nin the packfile for a second time. The objects need to be inflated\nagain in step 3, for a total of two inflations.\n\nPush\n~~~~\nPush is simpler than fetch because the objects referenced by the\npushed objects are already in the translation table. The sha1-content\nof each object being pushed can be read as described in the \"Reading\nan object's sha1-content\" section to generate the pack written by git\nsend-pack.\n\nSigned Objects\n~~~~~~~~~~~~~~\nCommits\n^^^^^^^\nCommits currently have the following sequence of header lines:\n\n\t\"tree\" SP object-name\n\t(\"parent\" SP object-name)*\n\t\"author\" SP ident\n\t\"committer\" SP ident\n\t(\"mergetag\" SP object-content)?\n\t(\"gpgsig\" SP pgp-signature)?\n\nWe introduce new header lines \"hash\" and \"nohash\" that come after the\n\"gpgsig\" field. No \"hash\" lines may appear unless the \"gpgsig\" field\nis present.\n\nHash lines have the form\n\n\t\"hash\" SP hash-function SP field SP alternate-object-name\n\nNohash lines have the form\n\n\t\"nohash\" SP hash-function\n\nThere are only two recognized values of hash-function: \"sha1\" and\n\"sha256\". \"git fsck\" will tolerate values of hash-function it does not\nrecognize, as long as they do not come before either of those two. All\n\"nohash\" lines come before all \"hash\" lines. Any \"hash sha1\" lines\nmust come before all \"hash sha256\" lines, and likewise for nohash. The\nGit project determines any future supported hash-functions that can\ncome after those two and their order.\n\nThere can be at most one \"nohash <hash-function>\" for one hash\nfunction, indicating that this hash function should not be used when\nchecking the commit's signature.\n\nThere is one \"hash <hash-function>\" line for each tree or parent field\nin the commit object header. The hash lines record object names for\nthose trees and parents using the indicated hash function, to be used\nwhen checking the commit's signature.\n\nTODO: simplify signature rules, handle the mergetag field better.\n\nsha256-content of signed commits\n^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nThe sha256-content of a commit with a \"gpgsig\" header can include no\nhash and nohash lines, a \"nohash sha256\" line and \"hash sha1\", or just\na \"hash sha1\" line.\n\nExamples:\n1. tree 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   parent e094bc809626f0a401a40d75c56df478e546902ff812772c4594265203b23980\n   parent 1059dab4748aa33b86dad5ca97357bd322abaa558921255623fbddd066bb3315\n   author A U Thor <author@example.com> 1465982009 +0000\n   committer C O Mitter <committer@example.com> 1465982009 +0000\n   gpgsig ...\n2. tree 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   parent e094bc809626f0a401a40d75c56df478e546902ff812772c4594265203b23980\n   parent 1059dab4748aa33b86dad5ca97357bd322abaa558921255623fbddd066bb3315\n   author A U Thor <author@example.com> 1465982009 +0000\n   committer C O Mitter <committer@example.com> 1465982009 +0000\n   gpgsig ...\n   nohash sha256\n   hash sha1 tree c7b1cff039a93f3600a1d18b82d26688668c7dea\n   hash sha1 parent c33429be94b5f2d3ee9b0adad223f877f174b05d\n   hash sha1 parent 04b871796dc0420f8e7561a895b52484b701d51a\n3. tree 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   parent e094bc809626f0a401a40d75c56df478e546902ff812772c4594265203b23980\n   parent 1059dab4748aa33b86dad5ca97357bd322abaa558921255623fbddd066bb3315\n   author A U Thor <author@example.com> 1465982009 +0000\n   committer C O Mitter <committer@example.com> 1465982009 +0000\n   gpgsig ...\n   hash sha1 tree c7b1cff039a93f3600a1d18b82d26688668c7dea\n   hash sha1 parent c33429be94b5f2d3ee9b0adad223f877f174b05d\n   hash sha1 parent 04b871796dc0420f8e7561a895b52484b701d51a\n\nsha1-content of signed commits\n^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nThe sha1-content of a commit with a \"gpgsig\" header can contain a\n\"nohash sha1\" and \"hash sha256\" line, no hash or nohash lines, or just\na \"hash sha256\" line.\n\nExamples:\n1. tree c7b1cff039a93f3600a1d18b82d26688668c7dea\n   parent c33429be94b5f2d3ee9b0adad223f877f174b05d\n   parent 04b871796dc0420f8e7561a895b52484b701d51a\n   author A U Thor <author@example.com> 1465982009 +0000\n   committer C O Mitter <committer@example.com> 1465982009 +0000\n   gpgsig ...\n   nohash sha1\n   hash sha256 tree 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   hash sha256 parent e094bc809626f0a401a40d75c56df478e546902ff812772c4594265203b23980\n   hash sha256 parent 1059dab4748aa33b86dad5ca97357bd322abaa558921255623fbddd066bb3315\n2. tree c7b1cff039a93f3600a1d18b82d26688668c7dea\n   parent c33429be94b5f2d3ee9b0adad223f877f174b05d\n   parent 04b871796dc0420f8e7561a895b52484b701d51a\n   author A U Thor <author@example.com> 1465982009 +0000\n   committer C O Mitter <committer@example.com> 1465982009 +0000\n   gpgsig ...\n3. tree c7b1cff039a93f3600a1d18b82d26688668c7dea\n   parent c33429be94b5f2d3ee9b0adad223f877f174b05d\n   parent 04b871796dc0420f8e7561a895b52484b701d51a\n   author A U Thor <author@example.com> 1465982009 +0000\n   committer C O Mitter <committer@example.com> 1465982009 +0000\n   gpgsig ...\n   hash sha256 tree 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   hash sha256 parent e094bc809626f0a401a40d75c56df478e546902ff812772c4594265203b23980\n   hash sha256 parent 1059dab4748aa33b86dad5ca97357bd322abaa558921255623fbddd066bb3315\n\nConverting signed commits\n^^^^^^^^^^^^^^^^^^^^^^^^^\nTo convert the sha1-content of a signed commit to its sha256-content:\n\n1. Change \"tree\" and \"parent\" lines to use the sha256-names of\n   referenced objects, as with unsigned commits.\n2. If there is a \"mergetag\" field, convert it from sha1-content to\n   sha256-content, as with unsigned commits with a mergetag (see the\n   \"Mergetag\" section below).\n3. Unless there is a \"nohash sha1\" line, add a full set of \"hash sha1\n   <field> <sha1>\" lines indicating the sha1-names of the tree and\n   parents.\n4. Remove any \"hash sha256 <field> <sha256>\" lines. If no such lines\n   were present, add a \"nohash sha256\" line.\n\nConverting the sha256-content of a signed commit to sha1-content uses\nthe same process with sha1 and sha256 switched.\n\nVerifying signed commit signatures\n^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nIf the commit has a \"hash sha1\" line (or is sha1-content without a\n\"nohash sha1\" line): check that the signature matches the sha1-content\nwith gpgsig field stripped out.\n\nOtherwise: check that the signature matches the sha1-content with\ngpgsig, nohash, tree, and parents fields stripped out.\n\nWith the examples above, the signed payloads are\n1. author A U Thor <author@example.com> 1465982009 +0000\n   committer C O Mitter <committer@example.com> 1465982009 +0000\n   hash sha256 tree 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   hash sha256 parent e094bc809626f0a401a40d75c56df478e546902ff812772c4594265203b23980\n   hash sha256 parent 1059dab4748aa33b86dad5ca97357bd322abaa558921255623fbddd066bb3315\n2. tree c7b1cff039a93f3600a1d18b82d26688668c7dea\n   parent c33429be94b5f2d3ee9b0adad223f877f174b05d\n   parent 04b871796dc0420f8e7561a895b52484b701d51a\n   author A U Thor <author@example.com> 1465982009 +0000\n   committer C O Mitter <committer@example.com> 1465982009 +0000\n3. tree c7b1cff039a93f3600a1d18b82d26688668c7dea\n   parent c33429be94b5f2d3ee9b0adad223f877f174b05d\n   parent 04b871796dc0420f8e7561a895b52484b701d51a\n   author A U Thor <author@example.com> 1465982009 +0000\n   committer C O Mitter <committer@example.com> 1465982009 +0000\n   hash sha1\n   hash sha256 tree 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   hash sha256 parent e094bc809626f0a401a40d75c56df478e546902ff812772c4594265203b23980\n   hash sha256 parent 1059dab4748aa33b86dad5ca97357bd322abaa558921255623fbddd066bb3315\n   \nCurrent versions of \"git verify-commit\" can verify examples (2) and (3)\n(but not (1)).\n\nTags\n~~~~\nTags currently have the following sequence of header lines:\n   \n   \t\"object\" SP object-name\n\t\"type\" SP type\n\t\"tag\" SP identifier\n\t\"tagger\" SP ident\n\nA tag's signature, if it exists, is in the message body.\n\nWe introduce new header lines \"nohash\" and \"hash\" that come after the\n\"tagger\" field. No \"nohash\" or \"hash\" lines may appear unless the\nmessage body contains a PGP signature.\n\nAs with commits, \"nohash\" lines have the form \"nohash\n<hash-function>\", indicating that this hash function should not be\nused when checking the tag's signature.\n\n\"hash\" lines have the form\n\n\t\"hash\" SP hash-function SP alternate-object-name\n\nThis records the pointed-to object name using the indicated hash\nfunction, to be used when checking the tag's signature.\n\nAs with commits, \"sha1\" and \"sha256\" are the only permitted values of\nhash-function and can only appear in that order for a field when they\nappear. There can be at most one \"nohash\" line, and it comes before\nany \"hash\" lines. There can be only one \"hash\" line for a given hash\nfunction.\n\nsha256-content of signed tags\n^^^^^^^^^^^^^^^^^^^^^^^^^^^^^\nThe sha256-content of a signed tag can include no \"hash\" or \"nohash\"\nlines, a \"nohash sha256\" and \"hash sha1 <sha1>\" line, or just a \"hash\nsha1 <sha1>\" line.\n\nExamples:\n1. object 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   type tree\n   tag v1.0\n   tagger C O Mitter <committer@example.com> 1465981006 +0000\n\n   Tag Demo v1.0\n   -----BEGIN PGP SIGNATURE-----\n   Version: GnuPG v1\n\n   iQEcBAABAgAGBQJXYRhOAAoJEGEJLoW3InGJklkIAIcnhL7RwEb/+QeX9enkXhxn\n   ...\n2. object 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   type tree\n   tag v1.0\n   tagger C O Mitter <committer@example.com> 1465981006 +0000\n   nohash sha256\n   hash sha1 c7b1cff039a93f3600a1d18b82d26688668c7dea\n\n   Tag Demo v1.0\n   -----BEGIN PGP SIGNATURE-----\n   ...\n3. object 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n   type tree\n   tag v1.0\n   tagger C O Mitter <committer@example.com> 1465981006 +0000\n   hash sha1 c7b1cff039a93f3600a1d18b82d26688668c7dea\n\n   Tag Demo v1.0\n   ...\n\nsha1-content of signed tags\n^^^^^^^^^^^^^^^^^^^^^^^^^^^\nThe sha1-content of a signed tag can include a \"nohash sha1\" and \"hash\nsha256\" line, no \"nohash\" or \"hash\" lines, or just a \"hash sha256\n<sha256>\" line.\n   \nExamples:\n1. object c7b1cff039a93f3600a1d18b82d26688668c7dea\n   ...\n   tagger C O Mitter <committer@example.com> 1465981006 +0000\n   nohash sha1\n   hash sha256 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n\n   Tag Demo v1.0\n   -----BEGIN PGP SIGNATURE-----\n   ...\n2. object c7b1cff039a93f3600a1d18b82d26688668c7dea\n   ...\n   tagger C O Mitter <committer@example.com> 1465981006 +0000\n\n   Tag Demo v1.0\n   -----BEGIN PGP SIGNATURE-----\n   ...\n3. object c7b1cff039a93f3600a1d18b82d26688668c7dea\n   ...\n   tagger C O Mitter <committer@example.com> 1465981006 +0000\n   hash sha256 98ea6e4f216f2fb4b69fff9b3a44842c38686ca685f3f55dc48c5d3fb1107be4\n\n   Tag Demo v1.0\n   -----BEGIN PGP SIGNATURE-----\n   ...\n\nSigned tags can be converted between sha1-content and sha256-content\nusing the same process as signed commits.\n\nVerifying signed tags\n^^^^^^^^^^^^^^^^^^^^^\nAs with commits, if the tag has a \"hash sha1\" (or is sha1-content\nwithout a \"nohash sha1\" line): check that the signature matches the\nsha1-content with PGP signature stripped out.\n   \nOtherwise: check that the signature matches the sha1-content with\nnohash and object fields and PGP signature stripped out.\n\nMergetag signatures\n~~~~~~~~~~~~~~~~~~~\nThe mergetag field in the sha1-content of a commit contains the\nsha1-content of a tag that was merged by that commit.\n\nThe mergetag field in the sha256-content of the same commit contains\nthe sha256-content of the same tag.\n\nSubmodules\n~~~~~~~~~~\nTo convert recorded submodule pointers, you need to have the converted\nsubmodule repository in place. The bidirectional mapping of the\nsubmodule can be used to look up the new hash.\n\nCaveats\n-------\nShallow clone and submodules\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~\nBecause this requires all referenced objects to be available in the\nlocally generated translation table, this design does not support\nshallow clone or unfetched submodules.\n\nProtocol improvements might allow lifting this restriction.\n\nAlternatives considered\n-----------------------\nUpgrading everyone working on a particular project on a flag day\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\nProjects like the Linux kernel are large and complex enough that\nflipping the switch for all projects based on the repository at once\nis infeasible.\n\nNot only would all developers and server operators supporting\ndevelopers have to switch on the same flag day, but supporting tooling\n(continuous integration, code review, bug trackers, etc) would have to\nbe adapted as well. This also makes it difficult to get early feedback\nfrom some project participants testing before it is time for mass\nadoption.\n\nUsing hash functions in parallel \n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n(e.g. https://public-inbox.org/git/22708.8913.864049.452252@chiark.greenend.org.uk/ )\nObjects newly created would be addressed by the new hash, but inside\nsuch an object (e.g. commit) it is still possible to address objects\nusing the old hash function.\n\n* You cannot trust its history (needed for bisectability) in the\n  future without further work \n* Maintenance burden as the number of supported hash functions grows\n  (they will never go away, so they accumulate). In this proposal, by\n  comparison, converted objects lose all references to SHA-1 except\n  where needed to verify signatures.\n"},{"id":"313287","messageId":"CA+55aFz+gkAsDZ24zmePQuEs1XPS9BP_s8O7Q4wQ7LV7X5-oDA@mail.gmail.com","threadId":"45288","inReplyTo":"20170304011251.GA26789@aiede.mtv.corp.google.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-05T02:35:38Z","receivedAt":"2017-03-05T02:35:45Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Fri, Mar 3, 2017 at 5:12 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n>\n> This document is still in flux but I thought it best to send it out\n> early to start getting feedback.\n\nThis actually looks very reasonable if you can implement it cleanly\nenough. In many ways the \"convert entirely to a new 256-bit hash\" is\nthe cleanest model, and interoperability was at least my personal\nconcern. Maybe your model solves it (devil in the details), in which\ncase I really like it.\n\nI do think that if you end up essentially converting the objects\nwithout really having any true backwards compatibility at the object\nlayer (just the translation code), you should seriously look at doing\nsome other changes at the same time. Like not using zlib compression,\nit really is very slow.\n\nBtw, I do think the particular choice of hash should still be on the\ntable. sha-256 may be the obvious first choice, but there are\ndefinitely a few reasons to consider alternatives, especially if it's\na complete switch-over like this.\n\nOne is large-file behavior - a parallel (or tree) mode could improve\non that noticeably. BLAKE2 does have special support for that, for\nexample. And SHA-256 does have known attacks compared to SHA-3-256 or\nBLAKE2 - whether that is due to age or due to more effort, I can't\nreally judge. But if we're switching away from SHA1 due to known\nattacks, it does feel like we should be careful.\n\n                Linus\n"},{"id":"313289","messageId":"nycvar.QRO.7.75.62.1703050258200.6590@qynat-yncgbc","threadId":"45288","inReplyTo":"20170304011251.GA26789@aiede.mtv.corp.google.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"David Lang","fromEmail":"david@lang.hm","sentAt":"2017-03-05T11:02:25Z","receivedAt":"2017-03-05T11:02:45Z","isPatch":false,"sender":{"key":"david@lang.hm","avatar":null},"body":"> Translation table\n> ~~~~~~~~~~~~~~~~~\n> A fast bidirectional mapping between sha1-names and sha256-names of\n> all local objects in the repository is kept on disk. The exact format\n> of that mapping is to be determined.\n>\n> All operations that make new objects (e.g., \"git commit\") add the new\n> objects to the translation table.\n\nThis seems like a rather nontrival thing to design. It will need to hold \nmillions of mappings, and be quickly searchable from either direction (sha1->new \nand new->sha1) while still be fairly fast to insert new records into.\n\nFor Linux, just the list of hashes recording the commits is going to be in the \nmillions, whiel the list of hashes of individual files for all those commits is \ngoing to be substantially larger.\n\nDavid Lang\n"},{"id":"313305","messageId":"20170306002642.xlatomtcrhxwshzn@genre.crustytoothpaste.net","threadId":"45288","inReplyTo":"CA+55aFz+gkAsDZ24zmePQuEs1XPS9BP_s8O7Q4wQ7LV7X5-oDA@mail.gmail.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-03-06T00:26:42Z","receivedAt":"2017-03-06T00:27:08Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Sat, Mar 04, 2017 at 06:35:38PM -0800, Linus Torvalds wrote:\n> On Fri, Mar 3, 2017 at 5:12 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> >\n> > This document is still in flux but I thought it best to send it out\n> > early to start getting feedback.\n> \n> This actually looks very reasonable if you can implement it cleanly\n> enough. In many ways the \"convert entirely to a new 256-bit hash\" is\n> the cleanest model, and interoperability was at least my personal\n> concern. Maybe your model solves it (devil in the details), in which\n> case I really like it.\n\nIf you think you can do it, I'm all for it.\n\n> Btw, I do think the particular choice of hash should still be on the\n> table. sha-256 may be the obvious first choice, but there are\n> definitely a few reasons to consider alternatives, especially if it's\n> a complete switch-over like this.\n> \n> One is large-file behavior - a parallel (or tree) mode could improve\n> on that noticeably. BLAKE2 does have special support for that, for\n> example. And SHA-256 does have known attacks compared to SHA-3-256 or\n> BLAKE2 - whether that is due to age or due to more effort, I can't\n> really judge. But if we're switching away from SHA1 due to known\n> attacks, it does feel like we should be careful.\n\nI agree with Linus on this.  SHA-256 is the slowest option, and it's the\none with the most advanced cryptanalysis.  SHA-3-256 is faster on 64-bit\nmachines (which, as we've seen on the list, is the overwhelming majority\nof machines using Git), and even BLAKE2b-256 is stronger.\n\nDoing this all over again in another couple years should also be a\nnon-goal.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\n+1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"313313","messageId":"20170306084353.nrns455dvkdsfgo5@sigill.intra.peff.net","threadId":"45288","inReplyTo":"20170304011251.GA26789@aiede.mtv.corp.google.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-06T08:43:53Z","receivedAt":"2017-03-06T09:11:05Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Mar 03, 2017 at 05:12:51PM -0800, Jonathan Nieder wrote:\n\n> This past week we came up with this idea for what a transition to a new\n> hash function for Git would look like.  I'd be interested in your\n> thoughts (especially if you can make them as comments on the document,\n> which makes it easier to address them and update the document).\n\nOverall it's an interesting idea. I thought at first that you were\nsuggesting servers do on-the-fly conversion, but after a more careful\nreading that isn't the case. And I don't think that would work, because\nthe conversion is expensive.\n\nSo this pushes the conversion cost onto the clients who decide to move\nto SHA-256. That may be a problem for sites which have a lot of clients\n(like CI hosts). But I guess they would just stick with SHA-1 as long as\npossible, until the upstream repo switches (and that _is_ a per-repo\nflag day, because the upstream host isn't going to convert back to SHA-1\non the fly to serve the old clients).\n\n> You can use the doc URL\n> \n>  https://goo.gl/gh2Mzc\n\nI'd encourage anybody following along to follow that link. I almost\ndidn't, but there are a ton of comments there (I'm not sure how I feel\nabout splitting the discussion off the list, though).\n\n> Goals\n> -----\n> 1. The transition to SHA256 can be done one local repository at a time.\n>    a. Requiring no action by any other party.\n>    b. A SHA256 repository can communicate with SHA-1 Git servers and\n>       clients (push/fetch).\n>    c. Users can use SHA-1 and SHA256 identifiers for objects\n>       interchangeably.\n>    d. New signed objects make use of a stronger hash function than\n>       SHA-1 for their security guarantees.\n> 2. Allow a complete transition away from SHA-1.\n>    a. Local metadata for SHA-1 compatibility can be dropped in a\n>       repository if compatibility with SHA-1 is no longer needed.\n\nI suspect we'll never get away from keeping the mapping table. You'll\nneed at least the sha1->sha256 table if you want to look up names found\nin historic commit messages, mailing list posts, etc.\n\nAnd you'll need the sha256->sha1 table if you want to verify the gpg\nsignatures on old tags and commits. That might be something people are\nwilling to drop, though.\n\n> After negotiation, the server sends a packfile containing the\n> requested objects. We convert the packfile to SHA-256 format using the\n> following steps:\n> \n> 1. index-pack: inflate each object in the packfile and compute its\n>    SHA-1. Objects can contain deltas in OBJ_REF_DELTA format against\n>    objects the client has locally. These objects can be looked up using\n>    the translation table and their sha1-content read as described above\n>    to resolve the deltas.\n> 2. topological sort: starting at the \"want\"s from the negotiation\n>    phase, walk through objects in the pack and emit a list of them in\n>    topologically sorted order. (This list only contains objects\n>    reachable from the \"wants\". If the pack from the server contained\n>    additional extraneous objects, then they will be discarded.)\n\nI don't think we do this right now, but you can actually find the entry\n(and exit) points of a pack during the index-pack step. Basically:\n\n  1. Keep a hashmap of objects mentioned in the pack.\n\n  2. When we process an object's content (i.e., compute its hash), also\n     parse it for any object references. Add entries in the hashmap for\n     any object mentioned this way. Mark the entry for the object we\n     processed with a \"HAVE\" bit, and mark any referenced object with a\n     \"REF\" bit.\n\n  3. After processing all objects, anything with a \"HAVE\" but no \"REF\"\n     is an entry point to the pack (i.e., something that we should have\n     asked for with a want). Anything with a \"REF\" but not a \"HAVE\" is\n     an exit point (i.e., an object that we are expected to already have\n     in our repo).\n\n     (I've thought about this before because we could possibly shortcut\n     the connectivity check using the exit points. It's complicated by\n     the fact that we don't assume the transitive presence of objects\n     unless they are reachable).\n\nI don't think using the \"want\"s as the entry points is unreasonable,\nthough. The server _shouldn't_ generally be sending us other cruft.\n\nI do wonder if you might be able to omit the extra object-graph walk\nfrom your step 2, if you could assign \"depths\" to each object during\nstep 1 instead of HAVE/REF bits. The trouble, of course, is that you're\nnot visiting the nodes in the right order (so given two trees, you're\nnot sure if one might eventually be a child of the other; how do you\nassign their depths?). I have a feeling there's a proof that it's\nimpossible, but I might just not be clever enough.\n\n\nOverall the basics of the conversion seem sound to me. The \"nohash\"\nthings seems more complicated than I think it ought to be, which\nprobably just means I'm missing something.  I left a few related\ncomments on the google doc, so I won't repeat them here.\n\n-Peff\n"},{"id":"313314","messageId":"20170306094334.whtuvotyppvcom2f@sigill.intra.peff.net","threadId":"45288","inReplyTo":"CA+dhYEXHbQfJ6KUB1tWS9u1MLEOJL81fTYkbxu4XO-i+379LPw@mail.gmail.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-06T09:43:34Z","receivedAt":"2017-03-06T09:45:30Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Mar 06, 2017 at 10:29:33AM +0100, ankostis wrote:\n\n> On 5 March 2017 at 12:02, David Lang <david@lang.hm> wrote:\n> >> Translation table\n> >> ~~~~~~~~~~~~~~~~~\n> >> A fast bidirectional mapping between sha1-names and sha256-names of\n> >> all local objects in the repository is kept on disk. The exact format\n> >> of that mapping is to be determined.\n> >>\n> >> All operations that make new objects (e.g., \"git commit\") add the new\n> >> objects to the translation table.\n> >\n> >\n> > This seems like a rather nontrival thing to design. It will need to hold\n> > millions of mappings, and be quickly searchable from either direction\n> > (sha1->new and new->sha1) while still be fairly fast to insert new records\n> > into.\n> >\n> > For Linux, just the list of hashes recording the commits is going to be in\n> > the millions, whiel the list of hashes of individual files for all those\n> > commits is going to be substantially larger.\n> \n> Apologies if it is a stupid idea, but could we avoid the mappings-table\n> just by\n> hard-linking to the same object from both (or more) hashes?\n> So instead of creating a text-db format, just use the filesystem.\n\nNo, for a few reasons:\n\n  1. Most of these objects will not be in the filesystem at all, but\n     rather in a packfile.\n\n  2. It's not just a different hash over the same bytes. The sha256-name\n     is taken over the sha256-content (which refers to other objects\n     using sha256). So they really are different objects. You probably\n     wouldn't keep the sha1 version around separately, but rather\n     generate it on the fly during a push to a sha1 server.\n\n  3. You really need to be able to take a sha256 name and convert it to\n     a sha1 and vice versa. Hardlinks don't help with that, because they\n     only point in one direction. That get you to the same _content_,\n     but not the other name (and I guess this is where your \"look up the\n     name and then compute the other digest comes in, but that's\n     probably too expensive to be workable).\n\nI do think updating the mapping could potentially be deferred until\ninteracting with a sha1 server. But because it needs to be generated in\nreverse-topological order, it's conceptually easier to do it one object\nat a time.\n\n-Peff\n"},{"id":"313345","messageId":"20170306182423.GB183239@google.com","threadId":"45288","inReplyTo":"20170306002642.xlatomtcrhxwshzn@genre.crustytoothpaste.net","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Brandon Williams","fromEmail":"bmwill@google.com","sentAt":"2017-03-06T18:24:23Z","receivedAt":"2017-03-06T18:31:23Z","isPatch":false,"sender":{"key":"bwilliams.eng@gmail.com","avatar":null},"body":"On 03/06, brian m. carlson wrote:\n> On Sat, Mar 04, 2017 at 06:35:38PM -0800, Linus Torvalds wrote:\n> > On Fri, Mar 3, 2017 at 5:12 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> > >\n> > > This document is still in flux but I thought it best to send it out\n> > > early to start getting feedback.\n> > \n> > This actually looks very reasonable if you can implement it cleanly\n> > enough. In many ways the \"convert entirely to a new 256-bit hash\" is\n> > the cleanest model, and interoperability was at least my personal\n> > concern. Maybe your model solves it (devil in the details), in which\n> > case I really like it.\n> \n> If you think you can do it, I'm all for it.\n> \n> > Btw, I do think the particular choice of hash should still be on the\n> > table. sha-256 may be the obvious first choice, but there are\n> > definitely a few reasons to consider alternatives, especially if it's\n> > a complete switch-over like this.\n> > \n> > One is large-file behavior - a parallel (or tree) mode could improve\n> > on that noticeably. BLAKE2 does have special support for that, for\n> > example. And SHA-256 does have known attacks compared to SHA-3-256 or\n> > BLAKE2 - whether that is due to age or due to more effort, I can't\n> > really judge. But if we're switching away from SHA1 due to known\n> > attacks, it does feel like we should be careful.\n> \n> I agree with Linus on this.  SHA-256 is the slowest option, and it's the\n> one with the most advanced cryptanalysis.  SHA-3-256 is faster on 64-bit\n> machines (which, as we've seen on the list, is the overwhelming majority\n> of machines using Git), and even BLAKE2b-256 is stronger.\n> \n> Doing this all over again in another couple years should also be a\n> non-goal.\n\nI agree that when we decide to move to a new algorithm that we should\nselect one which we plan on using for as long as possible (much longer\nthan a couple years).  While writing the document we simply used\n\"sha256\" because it was more tangible and easier to reference.\n\n> -- \n> brian m. carlson / brian with sandals: Houston, Texas, US\n> +1 832 623 2791 | https://www.crustytoothpaste.net/~bmc | My opinion only\n> OpenPGP: https://keybase.io/bk2204\n\n\n\n-- \nBrandon Williams\n"},{"id":"313346","messageId":"xmqqvarmnt49.fsf@junio-linux.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170306084353.nrns455dvkdsfgo5@sigill.intra.peff.net","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-03-06T18:43:34Z","receivedAt":"2017-03-06T18:48:58Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jeff King <peff@peff.net> writes:\n\n>> You can use the doc URL\n>> \n>>  https://goo.gl/gh2Mzc\n>\n> I'd encourage anybody following along to follow that link. I almost\n> didn't, but there are a ton of comments there (I'm not sure how I feel\n> about splitting the discussion off the list, though).\n\nI am sure how I feel about it---we should really discourage it,\nunless it is an effort to help polishing an early draft for wider\ndistribution and discussion.\n\n> I don't think we do this right now, but you can actually find the entry\n> (and exit) points of a pack during the index-pack step. Basically:\n\nWe have code to do the \"entry point\" computation in index-pack\nalready, I think, in 81a04b01 (\"index-pack: --clone-bundle option\",\n2016-03-03).\n\n> I don't think using the \"want\"s as the entry points is unreasonable,\n> though. The server _shouldn't_ generally be sending us other cruft.\n\nThat's true.\n"},{"id":"313347","messageId":"cdd7779a-acdb-99fd-a685-89b36df65393@google.com","threadId":"45288","inReplyTo":"20170306084353.nrns455dvkdsfgo5@sigill.intra.peff.net","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Jonathan Tan","fromEmail":"jonathantanmy@google.com","sentAt":"2017-03-06T18:39:49Z","receivedAt":"2017-03-06T18:49:54Z","isPatch":false,"sender":{"key":"jonathantanmy@fastmail.com","avatar":null},"body":"On 03/06/2017 12:43 AM, Jeff King wrote:\n> Overall the basics of the conversion seem sound to me. The \"nohash\"\n> things seems more complicated than I think it ought to be, which\n> probably just means I'm missing something.  I left a few related\n> comments on the google doc, so I won't repeat them here.\n\nI think \"nohash\" can be explained in 2 points:\n  1. When creating signed objects, \"nohash\" is almost never written. Just\n     create the object as usual and add \"hash\" lines for every other hash\n     function that you want the signature to cover.\n  2. When converting from function A to function B, add \"nohash B\" if\n     there were no \"hash B\" lines in the original object.\n\nThe \"nohash\" thing was in the hope of requiring only one signature to \nsign all the hashes (in all the functions) that the user wants, while \npreserving round-tripping ability.\n\nMaybe some examples would help to address the apparent complexity. These \nexamples are the same as those in the document. I'll also show future \ncompatibility with a hypothetical NEW hash function, and extend the rule \nabout signing/verification to 'sign in the earliest supported hash \nfunction in ({object's hash function} + {functions in \"hash\" lines} - \n{function in \"nohash\" line})'.\n\nExample 1 (existing signed commit)\n<sha-1 object stuff>  <sha256 object stuff>  <NEW object stuff>\n                       nohash sha256          nohash new\n                       hash sha1 ...          hash sha1 ...\n\nThis object was probably created in a SHA-1 repository with no knowledge \nthat we were going to transition to SHA256 (but there is nothing \npreventing us from creating the middle or right object and then \ntranslating it to the other functions).\n\nExample 2 (recommended way to sign a commit in a SHA256 repo)\n<sha-1 object stuff>  <sha256 object stuff>  <NEW object stuff>\nhash sha256 ...       hash sha1 ...          nohash new\n                                              hash sha1 ...\n                                              hash sha256 ...\n\nThis is the recommended way to create a SHA256 object in a SHA256 repo. \nThe rule about signing/verification (as stated above) is to sign in \nSHA-1, so when signing or verifying, we convert the object to SHA-1 and \nuse that as the payload. Note that the signature covers both the SHA-1 \nand SHA256 hashes, and that existing Git implementations can verify the \nsignature.\n\nExample 3 (a signer that does not care about SHA-1 anymore)\n<sha-1 object stuff>  <sha256 object stuff>  <NEW object stuff>\nnohash sha1                                  nohash new\nhash sha256 ...                              hash sha256 ...\n\nIf we were to create a SHA256 object without any mentions of SHA-1, the \nrule about signing/verification (as stated above) states that the \nsignature payload is the SHA256 object. This means that existing Git \nimplementations cannot verify the signature, but we can still round-trip \nto SHA-1 and back without losing any information (as far as I can tell).\n"},{"id":"313350","messageId":"CA+55aFxj7Vtwac64RfAz_u=U4tob4Xg+2pDBDFNpJdmgaTCmxA@mail.gmail.com","threadId":"45288","inReplyTo":"cdd7779a-acdb-99fd-a685-89b36df65393@google.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-06T19:22:17Z","receivedAt":"2017-03-06T19:22:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Mon, Mar 6, 2017 at 10:39 AM, Jonathan Tan <jonathantanmy@google.com> wrote:\n>\n> I think \"nohash\" can be explained in 2 points:\n\nI do think that that was my least favorite part of the suggestion. Not\njust \"nohash\", but all the special \"hash\" lines too.\n\nI would honestly hope that the design should not be about \"other\nhashes\". If you plan your expectations around the new hash being\nbroken, something is wrong to begin with.\n\nI do wonder if things wouldn't be simpler if the new format just\nincluded the SHA1 object name in the new object. Put it in the\n\"header\" line of the object, so that every time you look up an object,\nyou just _see_ the SHA1 of that object. You can even think of it as an\nadditional protection.\n\nBtw, the multi-collision attack referenced earlier does _not_ work for\nan iterated hash that has a bigger internal state than the final hash.\nWhich is actually a real argument against sha-256: the internal state\nof sha-256 is 256 bits, so if an attack can find collisions due to\nsome weakness, you really can then generate exponential collisions by\nchaining a linear collision search together.\n\nBut for sha3-256 or blake2, the internal hash state is larger than the\nfinal hash, so now you need to generate collisions not in the 256\nbits, but in the much larger search space of the internal hash space\nif you want to generate those exponential collisions.\n\nSo *if* the new object format uses a git header line like\n\n    \"blob <size> <sha1>\\0\"\n\nthen it would inherently contain that mapping from 256-bit hash to the\nSHA1, but it would actually also protect against attacks on the new\nhash. In fact, in particular for objects with internal format that\ndiffers between the two hashing models (ie trees and commits which to\nsome degree are higher-value targets), it would make attacks really\nquite complicated, I suspect.\n\nAnd you wouldn't need those \"hash\" or \"nohash\" things at all. The old\nSHA1 would simply always be there, and cheap to look up (ie you\nwouldn't have to unpack the whole object).\n\nHmm?\n\n                   Linus\n"},{"id":"313354","messageId":"20170306195959.GD183239@google.com","threadId":"45288","inReplyTo":"CA+55aFxj7Vtwac64RfAz_u=U4tob4Xg+2pDBDFNpJdmgaTCmxA@mail.gmail.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Brandon Williams","fromEmail":"bmwill@google.com","sentAt":"2017-03-06T19:59:59Z","receivedAt":"2017-03-06T20:00:12Z","isPatch":false,"sender":{"key":"bwilliams.eng@gmail.com","avatar":null},"body":"On 03/06, Linus Torvalds wrote:\n> On Mon, Mar 6, 2017 at 10:39 AM, Jonathan Tan <jonathantanmy@google.com> wrote:\n> >\n> > I think \"nohash\" can be explained in 2 points:\n> \n> I do think that that was my least favorite part of the suggestion. Not\n> just \"nohash\", but all the special \"hash\" lines too.\n> \n> I would honestly hope that the design should not be about \"other\n> hashes\". If you plan your expectations around the new hash being\n> broken, something is wrong to begin with.\n> \n> I do wonder if things wouldn't be simpler if the new format just\n> included the SHA1 object name in the new object. Put it in the\n> \"header\" line of the object, so that every time you look up an object,\n> you just _see_ the SHA1 of that object. You can even think of it as an\n> additional protection.\n> \n> Btw, the multi-collision attack referenced earlier does _not_ work for\n> an iterated hash that has a bigger internal state than the final hash.\n> Which is actually a real argument against sha-256: the internal state\n> of sha-256 is 256 bits, so if an attack can find collisions due to\n> some weakness, you really can then generate exponential collisions by\n> chaining a linear collision search together.\n> \n> But for sha3-256 or blake2, the internal hash state is larger than the\n> final hash, so now you need to generate collisions not in the 256\n> bits, but in the much larger search space of the internal hash space\n> if you want to generate those exponential collisions.\n> \n> So *if* the new object format uses a git header line like\n> \n>     \"blob <size> <sha1>\\0\"\n> \n> then it would inherently contain that mapping from 256-bit hash to the\n> SHA1, but it would actually also protect against attacks on the new\n> hash. In fact, in particular for objects with internal format that\n> differs between the two hashing models (ie trees and commits which to\n> some degree are higher-value targets), it would make attacks really\n> quite complicated, I suspect.\n> \n> And you wouldn't need those \"hash\" or \"nohash\" things at all. The old\n> SHA1 would simply always be there, and cheap to look up (ie you\n> wouldn't have to unpack the whole object).\n> \n> Hmm?\n\nI'll agree that the \"hash\" \"nohash\" bit isn't my favorite and is really\nonly there to address the signing of tags/commits in this new non-sha1\nworld.  I'm inclined to take a closer look at Jeff's suggestion which\nsimply has a signature for the hash that the signer cares about.\n\nI don't know if keeping around the SHA1 for every object buys you all\nthat much.  It would add an additional layer of protection but you would\nalso need to compute the SHA1 for each object indefinitely (assuming you\ninclude the SHA1 in new objects and not just converted objects).  The\nhope would be that at some point you could not worry about SHA1 at all.\nThat may be difficult for projects with long history with commit msgs\nwhich reference SHA1's of other commits (if you wanted to look up the\nreferenced commit, for example), but projects started in the new\nnon-sha1 world shouldn't have to ever compute a sha1.\n\nAlso, during this transition phase you would still need to maintain the\nsha1<->sha256 translation table to make looking up objects by their sha1\nname in a sha256 repo fast.  Otherwise I think it would take a\nnon-trivial amount of time to search a sha256 repo for a sha1 name.  So\nif you do include the sha1 in the new object format then you would end\nup with some duplicate information, which isn't the end of the world.\n\n-- \nBrandon Williams\n"},{"id":"313375","messageId":"xmqqwpc2m5r0.fsf@junio-linux.mtv.corp.google.com","threadId":"45288","inReplyTo":"CA+55aFxj7Vtwac64RfAz_u=U4tob4Xg+2pDBDFNpJdmgaTCmxA@mail.gmail.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-03-06T21:53:39Z","receivedAt":"2017-03-06T21:53:53Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> So *if* the new object format uses a git header line like\n>\n>     \"blob <size> <sha1>\\0\"\n>\n> then it would inherently contain that mapping from 256-bit hash to the\n> SHA1, but it would actually also protect against attacks on the new\n> hash.\n\nThis is easy for blobs as you only need to hash twice.  I am not\nsure if you can do the same for trees, though.  For that <sha1> to\nbe useful, the hash needs to be over the tree contents whose\nreferences are expressed in <sha1>, which in turn would mean...\n\n... ah, you would read these <sha1> off of the object header in the\nnew world and you do not need to expand the whole thing.  OK, I see\nhow it could work.\n\n> In fact, in particular for objects with internal format that\n> differs between the two hashing models (ie trees and commits which to\n> some degree are higher-value targets), it would make attacks really\n> quite complicated, I suspect.\n>\n> And you wouldn't need those \"hash\" or \"nohash\" things at all. The old\n> SHA1 would simply always be there, and cheap to look up (ie you\n> wouldn't have to unpack the whole object).\n"},{"id":"313384","messageId":"20170306234030.GB26789@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"nycvar.QRO.7.75.62.1703050258200.6590@qynat-yncgbc","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-03-06T23:40:30Z","receivedAt":"2017-03-06T23:50:07Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"David Lang wrote:\n\n>> Translation table\n>> ~~~~~~~~~~~~~~~~~\n>> A fast bidirectional mapping between sha1-names and sha256-names of\n>> all local objects in the repository is kept on disk. The exact format\n>> of that mapping is to be determined.\n>>\n>> All operations that make new objects (e.g., \"git commit\") add the new\n>> objects to the translation table.\n>\n> This seems like a rather nontrival thing to design. It will need to\n> hold millions of mappings, and be quickly searchable from either\n> direction (sha1->new and new->sha1) while still be fairly fast to\n> insert new records into.\n\nI am currently thinking of using LevelDB, since it has the advantages of\nbeing simple, already existing, and having already been ported to Java\n(allowing JGit can read and write the same format).\n\nIf that doesn't work, we'd try some other key-value store like Samba's\ntdb or Kyoto Cabinet.\n\nJonathan\n"},{"id":"313387","messageId":"20170307001709.GC26789@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"CA+55aFz+gkAsDZ24zmePQuEs1XPS9BP_s8O7Q4wQ7LV7X5-oDA@mail.gmail.com","subject":"RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-03-07T00:17:09Z","receivedAt":"2017-03-07T00:27:14Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Linus Torvalds wrote:\n> On Fri, Mar 3, 2017 at 5:12 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n\n>> This document is still in flux but I thought it best to send it out\n>> early to start getting feedback.\n>\n> This actually looks very reasonable if you can implement it cleanly\n> enough.\n\nThanks for the kind words on what had quite a few flaws still.  Here's\na new draft.  I think the next version will be a patch against\nDocumentation/technical/.\n\nAs before, comments welcome, both here and inline at\n\n  https://goo.gl/gh2Mzc\n\nChanges since v2:\n\nUse SHA3-256 instead of SHA2 (thanks, Linus and brian m.\ncarlson).[1][2]\n\nMake sha3-based signatures a separate field, avoiding the need for\n\"hash\" and \"nohash\" fields (thanks to peff[3]).\n\nAdd a sorting phase to fetch (thanks to Junio for noticing the need\nfor this).\n\nOmit blobs from the topological sort during fetch (thanks to peff).\n\nDiscuss alternates, git notes, and git servers in the caveats section\n(thanks to Junio Hamano, brian m. carlson[4], and Shawn Pearce).\n\nClarify language throughout (thanks to various commenters, especially\nJunio).\n\nSincerely,\nJonathan\n\nGit hash function transition\n============================\nStatus: Draft\nLast Updated: 2017-03-06\n\nObjective\n---------\nMigrate Git from SHA-1 to a stronger hash function.\n\nBackground\n----------\nAt its core, the Git version control system is a content addressable\nfilesystem. It uses the SHA-1 hash function to name content. For\nexample, files, directories, and revisions are referred to by hash\nvalues unlike in other traditional version control systems where files\nor versions are referred to via sequential numbers. The use of a hash\nfunction to address its content delivers a few advantages:\n\n* Integrity checking is easy. Bit flips, for example, are easily\n  detected, as the hash of corrupted content does not match its name.\n* Lookup of objects is fast.\n\nUsing a cryptographically secure hash function brings additional\nadvantages:\n\n* Object names can be signed and third parties can trust the hash to\n  address the signed object and all objects it references.\n* Communication using Git protocol and out of band communication\n  methods have a short reliable string that can be used to reliably\n  address stored content.\n\nOver time some flaws in SHA-1 have been discovered by security\nresearchers. https://shattered.io demonstrated a practical SHA-1 hash\ncollision. As a result, SHA-1 cannot be considered cryptographically\nsecure any more. This impacts the communication of hash values because\nwe cannot trust that a given hash value represents the known good\nversion of content that the speaker intended.\n\nSHA-1 still possesses the other properties such as fast object lookup\nand safe error checking, but other hash functions are equally suitable\nthat are believed to be cryptographically secure.\n\nGoals\n-----\n1. The transition to SHA3-256 can be done one local repository at a time.\n   a. Requiring no action by any other party.\n   b. A SHA3-256 repository can communicate with SHA-1 Git servers\n      (push/fetch).\n   c. Users can use SHA-1 and SHA3-256 identifiers for objects\n      interchangeably.\n   d. New signed objects make use of a stronger hash function than\n      SHA-1 for their security guarantees.\n2. Allow a complete transition away from SHA-1.\n   a. Local metadata for SHA-1 compatibility can be removed from a\n      repository if compatibility with SHA-1 is no longer needed.\n3. Maintainability throughout the process.\n   a. The object format is kept simple and consistent.\n   b. Creation of a generalized repository conversion tool.\n\nNon-Goals\n---------\n1. Add SHA3-256 support to Git protocol. This is valuable and the\n   logical next step but it is out of scope for this initial design.\n2. Transparently improving the security of existing SHA-1 signed\n   objects.\n3. Intermixing objects using multiple hash functions in a single\n   repository.\n4. Taking the opportunity to fix other bugs in git's formats and\n   protocols.\n5. Shallow clones and fetches into a SHA3-256 repository. (This will\n   change when we add SHA3-256 support to Git protocol.)\n6. Skip fetching some submodules of a project into a SHA3-256\n   repository. (This also depends on SHA3-256 support in Git\n   protocol.)\n\nOverview\n--------\nWe introduce a new repository format extension `sha3`. Repositories\nwith this extension enabled use SHA3-256 instead of SHA-1 to name\ntheir objects. This affects both object names and object content ---\nboth the names of objects and all references to other objects within\nan object are switched to the new hash function.\n\nsha3 repositories cannot be read by older versions of Git.\n\nAlongside the packfile, a sha3 repository stores a bidirectional\nmapping between sha3 and sha1 object names. The mapping is generated\nlocally and can be verified using \"git fsck\". Object lookups use this\nmapping to allow naming objects using either their sha1 and sha3 names\ninterchangeably.\n\n\"git cat-file\" and \"git hash-object\" gain options to display an object\nin its sha1 form and write an object given its sha1 form. This\nrequires all objects referenced by that object to be present in the\nobject database so that they can be named using the appropriate name\n(using the bidirectional hash mapping).\n\nFetches from a SHA-1 based server convert the fetched objects into\nsha3 form and record the mapping in the bidirectional mapping table\n(see below for details). Pushes to a SHA-1 based server convert the\nobjects being pushed into sha1 form so the server does not have to be\naware of the hash function the client is using.\n\nDetailed Design\n---------------\nObject names\n~~~~~~~~~~~~\nObjects can be named by their 40 hexadecimal digit sha1-name or 64\nhexadecimal digit sha3-name, plus names derived from those (see\ngitrevisions(7)).\n\nThe sha1-name of an object is the SHA-1 of the concatenation of its\ntype, length, a nul byte, and the object's sha1-content. This is the\ntraditional <sha1> used in Git to name objects.\n\nThe sha3-name of an object is the SHA3-256 of the concatenation of its\ntype, length, a nul byte, and the object's sha3-content.\n\nObject format\n~~~~~~~~~~~~~\nThe content as a byte sequence of a tag, commit, or tree object named\nby sha1 and sha3 differ because an object named by sha3-name refers to\nother objects by their sha3-names and an object named by sha1-name\nrefers to other objects by their sha1-names.\n\nThe sha3-content of an object is the same as its sha1-content, except\nthat objects referenced by the object are named using their sha3-names\ninstead of sha1-names. Because a blob object does not refer to any\nother object, its sha1-content and sha3-content are the same.\n\nThe format allows round-trip conversion between sha3-content and\nsha1-content.\n\nObject storage\n~~~~~~~~~~~~~~\nLoose objects use zlib compression and packed objects use the packed\nformat described in Documentation/technical/pack-format.txt, just like\ntoday. The content that is compressed and stored uses sha3-content\ninstead of sha1-content.\n\nTranslation table\n~~~~~~~~~~~~~~~~~\nA fast bidirectional mapping between sha1-names and sha3-names of all\nlocal objects in the repository is kept on disk. The exact format of\nthat mapping is to be determined.\n\nAll operations that make new objects (e.g., \"git commit\") add the new\nobjects to the translation table.\n\n(This work could have been deferred to push time, but that would\nsignificantly complicate and slow down pushes. Calculating the\nsha1-name at object creation time at the same time it is being\nstreamed to disk and having its sha3-name calculated should be an\nacceptable cost.)\n\nReading an object's sha1-content\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\nThe sha1-content of an object can be read by converting all sha3-names\nits sha3-content references to sha1-names using the translation table.\n\nFetch\n~~~~~\nFetching from a SHA-1 based server requires translating between SHA-1\nand SHA3-256 based representations on the fly.\n\nSHA-1s named in the ref advertisement that are present on the client\ncan be translated to SHA3-256 and looked up as local objects using the\ntranslation table.\n\nNegotiation proceeds as today. Any \"have\"s generated locally are\nconverted to SHA-1 before being sent to the server, and SHA-1s\nmentioned by the server are converted to SHA3-256 when looking them up\nlocally.\n\nAfter negotiation, the server sends a packfile containing the\nrequested objects. We convert the packfile to SHA3-256 format using\nthe following steps:\n\n1. index-pack: inflate each object in the packfile and compute its\n   SHA-1. Objects can contain deltas in OBJ_REF_DELTA format against\n   objects the client has locally. These objects can be looked up\n   using the translation table and their sha1-content read as\n   described above to resolve the deltas.\n2. topological sort: starting at the \"want\"s from the negotiation\n   phase, walk through objects in the pack and emit a list of them,\n   excluding blobs, in reverse topologically sorted order, with each\n   object coming later in the list than all objects it references.\n   (This list only contains objects reachable from the \"wants\". If the\n   pack from the server contained additional extraneous objects, then\n   they will be discarded.)\n3. convert to sha3: open a new (sha3) packfile. Read the topologically\n   sorted list just generated. For each object, inflate its\n   sha1-content, convert to sha3-content, and write it to the sha3\n   pack. Include the new sha1<->sha3 mapping entry in the translation\n   table.\n4. sort: reorder entries in the new pack to match the order of objects\n   in the pack the server generated and include blobs. Write a sha3 idx\n   file.\n5. clean up: remove the SHA-1 based pack file, index, and\n   topologically sorted list obtained from the server and steps 1\n   and 2.\n\nStep 3 requires every object referenced by the new object to be in the\ntranslation table. This is why the topological sort step is necessary.\n\nAs an optimization, step 1 could write a file describing what non-blob\nobjects each object it has inflated from the packfile references. This\nmakes the topological sort in step 2 possible without inflating the\nobjects in the packfile for a second time. The objects need to be\ninflated again in step 3, for a total of two inflations.\n\nStep 4 is probably necessary for good read-time performance. \"git\npack-objects\" on the server optimizes the pack file for good data\nlocality (see Documentation/technical/pack-heuristics.txt).\n\nDetails of this process are likely to change. It will take some\nexperimenting to get this to perform well.\n\nPush\n~~~~\nPush is simpler than fetch because the objects referenced by the\npushed objects are already in the translation table. The sha1-content\nof each object being pushed can be read as described in the \"Reading\nan object's sha1-content\" section to generate the pack written by git\nsend-pack.\n\nSigned Commits\n~~~~~~~~~~~~~~\nWe add a new field \"gpgsig-sha3\" to the commit object format to allow\nsigning commits without relying on SHA-1. It is similar to the\nexisting \"gpgsig\" field. Its signed payload is the sha3-content of the\ncommit object with any \"gpgsig\" and \"gpgsig-sha3\" fields removed.\n\nThis means commits can be signed\n1. using SHA-1 only, as in existing signed commit objects\n2. using both SHA-1 and SHA3-256, by using both gpgsig-sha3 and gpgsig\n   fields.\n3. using only SHA3-256, by only using the gpgsig-sha3 field.\n\nOld versions of \"git verify-commit\" can verify the gpgsig signature in\ncases (1) and (2) without modifications and view case (3) as an\nordinary unsigned commit.\n\nSigned Tags\n~~~~~~~~~~~\nWe add a new field \"gpgsig-sha3\" to the tag object format to allow\nsigning tags without relying on SHA-1. Its signed payload is the\nsha3-content of the tag with its gpgsig-sha3 field and \"-----BEGIN PGP\nSIGNATURE-----\" delimited in-body signature removed.\n\nThis means tags can be signed\n1. using SHA-1 only, as in existing signed tag objects\n2. using both SHA-1 and SHA3-256, by using gpgsig-sha3 and an in-body\n   signature.\n3. using only SHA3-256, by only using the gpgsig-sha3 field.\n\nMergetag embedding\n~~~~~~~~~~~~~~~~~~\nThe mergetag field in the sha1-content of a commit contains the\nsha1-content of a tag that was merged by that commit.\n\nThe mergetag field in the sha3-content of the same commit contains the\nsha3-content of the same tag.\n\nSubmodules\n~~~~~~~~~~\nTo convert recorded submodule pointers, you need to have the converted\nsubmodule repository in place. The translation table of the submodule\ncan be used to look up the new hash.\n\nCaveats\n-------\nInvalid objects\n~~~~~~~~~~~~~~~\nThe conversion from sha1-content to sha3-content retains any\nbrokenness in the original object (e.g., tree entry modes encoded with\nleading 0, tree objects whose paths are not sorted correctly, and\ncommit objects without an author or committer). This is a deliberate\nfeature of the design to allow the conversion to round-trip.\n\nMore profoundly broken objects (e.g., a commit with a truncated \"tree\"\nheader line) cannot be converted but were not usable by current Git\nanyway.\n\nShallow clone and submodules\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~\nBecause it requires all referenced objects to be available in the\nlocally generated translation table, this design does not support\nshallow clone or unfetched submodules. Protocol improvements might\nallow lifting this restriction.\n\nAlternates\n~~~~~~~~~~\nFor the same reason, a sha3 repository cannot borrow objects from a\nsha1 repository using objects/info/alternates or\n$GIT_ALTERNATE_OBJECT_REPOSITORIES.\n\ngit notes\n~~~~~~~~~\nThe \"git notes\" tool annotates objects using their sha1-name as key.\nThis design does not describe a way to migrate notes trees to use\nsha3-names. That migration is expected to happen separately (for\nexample using a file at the root of the notes tree to describe which\nhash it uses).\n\nServer-side cost\n~~~~~~~~~~~~~~~~\nUntil Git protocol gains SHA3-256 support, using sha3 based storage on\npublic-facing Git servers is strongly discouraged. Once Git protocol\ngains SHA3-256 support, sha3 based servers are likely not to support\nsha1 compatibility, to avoid what may be a very expensive hash\nreencode during clone and to encourage peers to modernize.\n\nThe design described here allows fetches by SHA-1 clients of a\npersonal SHA256 repository because it's not much more difficult than\nallowing pushes from that repository. This support needs to be guarded\nby a configuration option --- servers like git.kernel.org that serve a\nlarge number of clients would not be expected to bear that cost.\n\nMeaning of signatures\n~~~~~~~~~~~~~~~~~~~~~\nThe signed payload for signed commits and tags does not explicitly\nname the hash used to identify objects. If some day Git adopts a new\nhash function with the same length as the current SHA-1 (40\nhexadecimal digit) or SHA2-256 (64 hexadecimal digit) objects then the\nintent behind the PGP signed payload in an object signature is\nunclear:\n\n\tobject e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7\n\ttype commit\n\ttag v2.12.0\n\ttagger Junio C Hamano <gitster@pobox.com> 1487962205 -0800\n\n\tGit 2.12\n\nDoes this mean Git v2.12.0 is the commit with sha1-name\ne7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7 or the commit with\nnew-40-digit-hash-name e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7?\n\nFortunately SHA3-256 and SHA-1 have different lengths. If Git starts\nusing another hash with the same length to name objects, then it will\nneed to change the format of signed payloads using that hash to\naddress this issue.\n\nAlternatives considered\n-----------------------\nUpgrading everyone working on a particular project on a flag day\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\nProjects like the Linux kernel are large and complex enough that\nflipping the switch for all projects based on the repository at once\nis infeasible.\n\nNot only would all developers and server operators supporting\ndevelopers have to switch on the same flag day, but supporting tooling\n(continuous integration, code review, bug trackers, etc) would have to\nbe adapted as well. This also makes it difficult to get early feedback\nfrom some project participants testing before it is time for mass\nadoption.\n\nUsing hash functions in parallel\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n(e.g. https://public-inbox.org/git/22708.8913.864049.452252@chiark.greenend.org.uk/ )\nObjects newly created would be addressed by the new hash, but inside\nsuch an object (e.g. commit) it is still possible to address objects\nusing the old hash function.\n* You cannot trust its history (needed for bisectability) in the\n  future without further work\n* Maintenance burden as the number of supported hash functions grows\n  (they will never go away, so they accumulate). In this proposal, by\n  comparison, converted objects lose all references to SHA-1.\n\nSigned objects with multiple hashes\n~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\nInstead of introducing the gpgsig-sha3 field in commit and tag objects\nfor sha3-content based signatures, an earlier version of this design\nadded \"hash sha3 <sha3-name>\" fields to strengthen the existing\nsha1-content based signatures.\n\nIn other words, a single signature was used to attest to the object\ncontent using both hash functions. This had some advantages:\n* Using one signature instead of two speeds up the signing process.\n* Having one signed payload with both hashes allows the signer to\n  attest to the sha1-name and sha3-name referring to the same object.\n* All users consume the same signature. Broken signatures are likely\n  to be detected quickly using current versions of git.\n\nHowever, it also came with disadvantages:\n* Verifying a signed object requires access to the sha1-names of all\n  objects it references, even after the transition is complete and\n  translation table is no longer needed for anything else. To support\n  this, the design added fields such as \"hash sha1 tree <sha1-name>\"\n  and \"hash sha1 parent <sha1-name>\" to the sha3-content of a signed\n  commit, complicating the conversion process.\n* Allowing signed objects without a sha1 (for after the transition is\n  complete) complicated the design further, requiring a \"nohash sha1\"\n  field to suppress including \"hash sha1\" fields in the sha3-content\n  and signed payload.\n\nDocument History\n----------------\n\n2017-03-03\nbmwill@google.com, jonathantanmy@google.com, jrnieder@gmail.com,\nsbeller@google.com\n\nInitial version sent to\nhttp://public-inbox.org/git/20170304011251.GA26789@aiede.mtv.corp.google.com\n\n2017-03-03 jrnieder@gmail.com\nIncorporated suggestions from jonathantanmy and sbeller:\n* describe purpose of signed objects with each hash type\n* redefine signed object verification using object content under the\n  first hash function\n\n2017-03-06 jrnieder@gmail.com\n* Use SHA3-256 instead of SHA2 (thanks, Linus and brian m. carlson).[1][2]\n* Make sha3-based signatures a separate field, avoiding the need for\n  \"hash\" and \"nohash\" fields (thanks to peff[3]).\n* Add a sorting phase to fetch (thanks to Junio for noticing the need\n  for this).\n* Omit blobs from the topological sort during fetch (thanks to peff).\n* Discuss alternates, git notes, and git servers in the caveats\n  section (thanks to Junio Hamano, brian m. carlson[4], and Shawn\n  Pearce).\n* Clarify language throughout (thanks to various commenters,\n  especially Junio).\n\n[1] http://public-inbox.org/git/CA+55aFzJtejiCjV0e43+9oR3QuJK2PiFiLQemytoLpyJWe6P9w@mail.gmail.com/\n[2] http://public-inbox.org/git/CA+55aFz+gkAsDZ24zmePQuEs1XPS9BP_s8O7Q4wQ7LV7X5-oDA@mail.gmail.com/\n[3] http://public-inbox.org/git/20170306084353.nrns455dvkdsfgo5@sigill.intra.peff.net/\n[4] http://public-inbox.org/git/20170304224936.rqqtkdvfjgyezsht@genre.crustytoothpaste.net\n"},{"id":"313388","messageId":"20170307000315.6cywmnx35ip7ftmc@glandium.org","threadId":"45288","inReplyTo":"20170306234030.GB26789@aiede.mtv.corp.google.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2017-03-07T00:03:15Z","receivedAt":"2017-03-07T00:27:16Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Mon, Mar 06, 2017 at 03:40:30PM -0800, Jonathan Nieder wrote:\n> David Lang wrote:\n> \n> >> Translation table\n> >> ~~~~~~~~~~~~~~~~~\n> >> A fast bidirectional mapping between sha1-names and sha256-names of\n> >> all local objects in the repository is kept on disk. The exact format\n> >> of that mapping is to be determined.\n> >>\n> >> All operations that make new objects (e.g., \"git commit\") add the new\n> >> objects to the translation table.\n> >\n> > This seems like a rather nontrival thing to design. It will need to\n> > hold millions of mappings, and be quickly searchable from either\n> > direction (sha1->new and new->sha1) while still be fairly fast to\n> > insert new records into.\n> \n> I am currently thinking of using LevelDB, since it has the advantages of\n> being simple, already existing, and having already been ported to Java\n> (allowing JGit can read and write the same format).\n> \n> If that doesn't work, we'd try some other key-value store like Samba's\n> tdb or Kyoto Cabinet.\n\nFWIW, I'm using notes-like data to store mercurial->git mappings in\ngit-cinnabar, (ab)using the commit type in tree items. It's fast enough.\n\nMike\n"},{"id":"313401","messageId":"20170307085926.x4ahakdhyx2vkzlx@sigill.intra.peff.net","threadId":"45288","inReplyTo":"cdd7779a-acdb-99fd-a685-89b36df65393@google.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-07T08:59:26Z","receivedAt":"2017-03-07T09:07:01Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Mar 06, 2017 at 10:39:49AM -0800, Jonathan Tan wrote:\n\n> The \"nohash\" thing was in the hope of requiring only one signature to sign\n> all the hashes (in all the functions) that the user wants, while preserving\n> round-tripping ability.\n\nThanks, this explained it very well.\n\nI understand the tradeoff now, though I am still of the opinion that\nsimplicity is probably a more important goal.\n\nIn practice I'd imagine that anybody doing commit-signing would just\nsign the more-secure hash, and people doing tag releases would probably\ndo a dual-sign to be verifiable by both old and new clients. Those are\ninfrequent enough that the extra computation probably doesn't matter.\nBut that's just my gut feeling.\n\n-Peff\n"},{"id":"313442","messageId":"CA+55aFyyi0vBBApf9grYQzF2PRZMjtCzkB4LzYvLpqQ-Z7QfJQ@mail.gmail.com","threadId":"45288","inReplyTo":"22719.680.730866.781688@chiark.greenend.org.uk","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-03-07T19:15:34Z","receivedAt":"2017-03-07T19:46:00Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Tue, Mar 7, 2017 at 10:57 AM, Ian Jackson\n<ijackson@chiark.greenend.org.uk> wrote:\n>\n> Also I think you need to specify how abbreviated object names are\n> interpreted.\n\nOne option might be to not use hex for the new hash, but base64 encoding.\n\nThat would make the full size ASCII hash encoding length roughly\nsimilar (43 base64 characters rather than 40), which would offset some\nof the new costs (longer filenames in the loose format, for example).\n\nAlso, since 256 isn't evenly divisible by 6, and because you'd want\nsome way to explictly disambiguate the new hashes, the rule *could* be\nthat the ASCII representation of a new hash is the base64 encoding of\nthe 258-bit value that has \"10\" prepended to it as padding.\n\nThat way the first character of the hash would be guaranteed to not be\na hex digit, because it would be in the range [g-v] (indexes 32..47).\n\nOf course, the downside is that base64 encoded hashes can also end up\nlooking very much like real words, and now case would matter too.\n\nThe \"use base64 with a \"10\" two-bit padding prepended\" also means that\nthe natural loose format radix format would remain the first 2\ncharacters of the hash, but due to the first character containing the\npadding, it would be a fan-out of 2**10 rather than 2**12.\n\nOf course, having written that, I now realize how it would cause\nproblems for the usual shit-for-brains case-insensitive filesystems.\nSo I guess base64 encoding doesn't work well for that reason.\n\n                Linus\n"},{"id":"313445","messageId":"22719.680.730866.781688@chiark.greenend.org.uk","threadId":"45288","inReplyTo":"20170304011251.GA26789@aiede.mtv.corp.google.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-03-07T18:57:44Z","receivedAt":"2017-03-07T20:08:35Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"Jonathan Nieder writes (\"RFC: Another proposed hash function transition plan\"):\n> This past week we came up with this idea for what a transition to a new\n> hash function for Git would look like.  I'd be interested in your\n> thoughts (especially if you can make them as comments on the document,\n> which makes it easier to address them and update the document).\n\nThanks for this.\n\nThis is a reasonable plan.  It corresponds to approaches (2) and (B)\nof my survey mail from the other day.  Ie, two parallel homogeneous\nhash trees, rather than a unified but heterogeneous hash tree, with\nold vs new object names distinguished by length.\n\nI still prefer my proposal with the mixed hash tree, mostly because\nthe handling of signatures here is very awkward, and because my\nproposal does not involve altering object ids stored other than in the\ngit object graph (eg CI system databases, etc.)\n\nOne thing you've missed, I think, is notes: notes have to be dealt\nwith in a more complicated way.  Do you intend to rewrite the tree\nobjects for notes commits so that the notes are annotations for the\nnew names for the annotated objects ?  And if so, when ?\n\nAlso I think you need to specify how abbreviated object names are\ninterpreted.\n\nRegards,\nIan.\n"},{"id":"313491","messageId":"22719.59633.269164.986923@chiark.greenend.org.uk","threadId":"45288","inReplyTo":"CA+55aFyyi0vBBApf9grYQzF2PRZMjtCzkB4LzYvLpqQ-Z7QfJQ@mail.gmail.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Ian Jackson","fromEmail":"ijackson@chiark.greenend.org.uk","sentAt":"2017-03-08T11:20:17Z","receivedAt":"2017-03-08T12:52:33Z","isPatch":false,"sender":{"key":"ijackson@chiark.greenend.org.uk","avatar":null},"body":"Linus Torvalds writes (\"Re: RFC: Another proposed hash function transition plan\"):\n> Also, since 256 isn't evenly divisible by 6, and because you'd want\n> some way to explictly disambiguate the new hashes, the rule *could* be\n> that the ASCII representation of a new hash is the base64 encoding of\n> the 258-bit value that has \"10\" prepended to it as padding.\n> \n> That way the first character of the hash would be guaranteed to not be\n> a hex digit, because it would be in the range [g-v] (indexes 32..47).\n\nWe should arrange for this to be an uppercase, not a lowercase,\nletter, for the reasons I explained in my own proposal.  To summarise:\nIt would be undesirable to further increase the overlap between object\nnames and ref names.  Few people use uppercase in ref names because of\nthe case-insensitive filesystem problem; so object names starting with\nuppercase ascii are distinct from most object names.\n\n> Of course, having written that, I now realize how it would cause\n> problems for the usual shit-for-brains case-insensitive filesystems.\n> So I guess base64 encoding doesn't work well for that reason.\n\nAFAIAA object names occur in publicly-visible filenames only in notes\ntree objects, which are manipulated by git internally and do not\nnecessarily need to appear in the filesystem.\n\nThe filenames in .git/objects/ can be in whatever encoding we like, so\nare not an obstacle.\n\nIan.\n\n-- \nIan Jackson <ijackson@chiark.greenend.org.uk>   These opinions are my own.\n\nIf I emailed you from an address @fyvzl.net or @evade.org.uk, that is\na private address which bypasses my fierce spamfilter.\n"},{"id":"313500","messageId":"alpine.DEB.2.20.1703081637030.3767@virtualbox","threadId":"45288","inReplyTo":"22719.59633.269164.986923@chiark.greenend.org.uk","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-03-08T15:37:23Z","receivedAt":"2017-03-08T15:39:48Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Ian,\n\nOn Wed, 8 Mar 2017, Ian Jackson wrote:\n\n> Few people use uppercase in ref names because of the case-insensitive\n> filesystem problem;\n\nNot true.\n\nCiao,\nJohannes\n"},{"id":"313501","messageId":"alpine.DEB.2.20.1703081638010.3767@virtualbox","threadId":"45288","inReplyTo":"22719.59633.269164.986923@chiark.greenend.org.uk","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-03-08T15:40:48Z","receivedAt":"2017-03-08T15:41:42Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Ian,\n\nOn Wed, 8 Mar 2017, Ian Jackson wrote:\n\n> Linus Torvalds writes (\"Re: RFC: Another proposed hash function transition plan\"):\n> > Of course, having written that, I now realize how it would cause\n> > problems for the usual shit-for-brains case-insensitive filesystems.\n> > So I guess base64 encoding doesn't work well for that reason.\n> \n> AFAIAA object names occur in publicly-visible filenames only in notes\n> tree objects, which are manipulated by git internally and do not\n> necessarily need to appear in the filesystem.\n> \n> The filenames in .git/objects/ can be in whatever encoding we like, so\n> are not an obstacle.\n\nGiven that the idea was to encode the new hash in base64 or base85, we\n*are* talking about an encoding. In that respect, yes, it can be whatever\nencoding we like, and Linus just made a good point (with unnecessary foul\nlanguage) of explaining why base64/base85 is not that encoding.\n\nCiao,\nJohannes\n"},{"id":"313668","messageId":"CAJo=hJtoX9=AyLHHpUJS7fueV9ciZ_MNpnEPHUz8Whui6g9F0A@mail.gmail.com","threadId":"45288","inReplyTo":"20170307001709.GC26789@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2017-03-09T19:14:12Z","receivedAt":"2017-03-09T19:14:39Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"On Mon, Mar 6, 2017 at 4:17 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> Linus Torvalds wrote:\n>> On Fri, Mar 3, 2017 at 5:12 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n>\n>>> This document is still in flux but I thought it best to send it out\n>>> early to start getting feedback.\n>>\n>> This actually looks very reasonable if you can implement it cleanly\n>> enough.\n>\n> Thanks for the kind words on what had quite a few flaws still.  Here's\n> a new draft.  I think the next version will be a patch against\n> Documentation/technical/.\n\nFWIW, I like this approach.\n\n> Alongside the packfile, a sha3 repository stores a bidirectional\n> mapping between sha3 and sha1 object names. The mapping is generated\n> locally and can be verified using \"git fsck\". Object lookups use this\n> mapping to allow naming objects using either their sha1 and sha3 names\n> interchangeably.\n\nI saw some discussion about using LevelDB for this mapping table. I\nthink any existing database may be overkill.\n\nFor packs, you may be able to simplify by having only one file\n(pack-*.msha1) that maps SHA-1 to pack offset; idx v2. The CRC32 table\nin v2 is unnecessary, but you need the 64 bit offset support.\n\nSHA-1 to SHA-3: lookup SHA-1 in .msha1, reverse .idx, find offset to\nread the SHA-3.\nSHA-3 to SHA-1: lookup SHA-3 in .idx, and reverse the .msha1 file to\ntranslate offset to SHA-1.\n\n\nFor loose objects, the loose object directories should have only\nO(4000) entries before auto gc is strongly encouraging\npacking/pruning. With 256 shards, each given directory has O(16) loose\nobjects in it. When writing a SHA-3 loose object, Git could also\nappend a line \"$sha3 $sha1\\n\" to objects/${first_byte}/sha1, which\nGC/prune rewrites to remove entries. With O(16) objects in a\ndirectory, these files should only have O(16) entries in them.\n\nSHA-3 to SHA-1: open objects/${sha3_first_byte}/sha1 and scan until a\nmatch is found.\nSHA-1 to SHA-3: brute force read 256 files. Callers performing this\nmapping may load all 256 files into a table in memory.\n"},{"id":"313675","messageId":"20170309202408.GA17847@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"CAJo=hJtoX9=AyLHHpUJS7fueV9ciZ_MNpnEPHUz8Whui6g9F0A@mail.gmail.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-03-09T20:24:08Z","receivedAt":"2017-03-09T20:25:18Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nShawn Pearce wrote:\n> On Mon, Mar 6, 2017 at 4:17 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n\n>> Alongside the packfile, a sha3 repository stores a bidirectional\n>> mapping between sha3 and sha1 object names. The mapping is generated\n>> locally and can be verified using \"git fsck\". Object lookups use this\n>> mapping to allow naming objects using either their sha1 and sha3 names\n>> interchangeably.\n>\n> I saw some discussion about using LevelDB for this mapping table. I\n> think any existing database may be overkill.\n>\n> For packs, you may be able to simplify by having only one file\n> (pack-*.msha1) that maps SHA-1 to pack offset; idx v2. The CRC32 table\n> in v2 is unnecessary, but you need the 64 bit offset support.\n>\n> SHA-1 to SHA-3: lookup SHA-1 in .msha1, reverse .idx, find offset to\n> read the SHA-3.\n> SHA-3 to SHA-1: lookup SHA-3 in .idx, and reverse the .msha1 file to\n> translate offset to SHA-1.\n\nThanks for this suggestion.  I was initially vaguely nervous about\nlookup times in an idx-style file, but as you say, object reads from a\npackfile already have to deal with this kind of lookup and work fine.\n\n> For loose objects, the loose object directories should have only\n> O(4000) entries before auto gc is strongly encouraging\n> packing/pruning. With 256 shards, each given directory has O(16) loose\n> objects in it. When writing a SHA-3 loose object, Git could also\n> append a line \"$sha3 $sha1\\n\" to objects/${first_byte}/sha1, which\n> GC/prune rewrites to remove entries. With O(16) objects in a\n> directory, these files should only have O(16) entries in them.\n\nInsertion time is what worries me.  When writing a small number of\nobjects using a command like \"git commit\", I don't want to have to\nregenerate an entire idx file.  I don't want to move the pain to\nO(loose objects) work at read time, either --- some people disable\nauto gc, and others have a large number of loose objects due to gc\nejecting unreachable objects.\n\nBut some kind of simplification along these lines should be possible.\nI'll experiment.\n\nJonathan\n"},{"id":"313780","messageId":"20170310193835.t7syswueuu7nmkjz@sigill.intra.peff.net","threadId":"45288","inReplyTo":"20170309202408.GA17847@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-03-10T19:38:35Z","receivedAt":"2017-03-10T19:38:42Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Mar 09, 2017 at 12:24:08PM -0800, Jonathan Nieder wrote:\n\n> > SHA-1 to SHA-3: lookup SHA-1 in .msha1, reverse .idx, find offset to\n> > read the SHA-3.\n> > SHA-3 to SHA-1: lookup SHA-3 in .idx, and reverse the .msha1 file to\n> > translate offset to SHA-1.\n> \n> Thanks for this suggestion.  I was initially vaguely nervous about\n> lookup times in an idx-style file, but as you say, object reads from a\n> packfile already have to deal with this kind of lookup and work fine.\n\nNot exactly. The \"reverse .idx\" step has to build the reverse mapping on\nthe fly, and it's non-trivial. For instance, try:\n\n  sha1=$(git rev-parse HEAD)\n  time echo $sha1 | git cat-file --batch-check='%(objectsize)'\n  time echo $sha1 | git cat-file --batch-check='%(objectsize:disk)'\n\non a large repo (where HEAD is in a big pack). The on-disk size is\nconceptually simpler, as we only need to look at the offset of the\nobject versus the offset of the object after it. But in practice it\ntakes much longer, because it has to build the revindex on the fly (I\nget 7ms versus 179ms on linux.git).\n\nThe effort is linear in the number of objects (we create the revindex\nwith a radix sort).\n\nThe reachability bitmaps suffer from this, too, as they need the\nrevindex to know which object is at which bit position. At GitHub we\nadded an extension to the .bitmap files that stores this \"bit cache\".\nHere are timings before and after on linux.git:\n\n  $ time git rev-list --use-bitmap-index --count master\n  659371\n\n  real\t0m0.182s\n  user\t0m0.136s\n  sys\t0m0.044s\n\n  $ time git.gh rev-list --use-bitmap-index --count master\n  659371\n\n  real\t0m0.016s\n  user\t0m0.008s\n  sys\t0m0.004s\n\nIt's not a full revindex, but it's enough for bitmap use. You can also\nuse it to generate the revindex slightly more quickly, because you can\nskip the sorting step (you just insert the entries in the correct order\nby walking the bit cache and dereferencing the offsets from the .idx\nportion). So it's still linear, but with a smaller constant factor.\n\nI think for the purposes here, though, we don't actually care about the\noffsets. For the cost of one uint32_t per object, you can keep a list\nmapping positions in the sha1 index into the sha3 index. So then you do\nthe log-n binary search to find the sha1, a constant-time lookup in the\nmapping array, and that gives you the position in the sha3 index, from\nwhich you can then access the sha3 (or the actual pack offset, for that\nmatter).\n\nSo I think it's solvable, but I suspect we would want an extension to\nthe .idx format to store the mapping array, in order to keep it log-n.\n\n-Peff\n"},{"id":"313786","messageId":"20170310195523.GF26789@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170310193835.t7syswueuu7nmkjz@sigill.intra.peff.net","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-03-10T19:55:24Z","receivedAt":"2017-03-10T19:55:32Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Jeff King wrote:\n> On Thu, Mar 09, 2017 at 12:24:08PM -0800, Jonathan Nieder wrote:\n\n>>> SHA-1 to SHA-3: lookup SHA-1 in .msha1, reverse .idx, find offset to\n>>> read the SHA-3.\n>>> SHA-3 to SHA-1: lookup SHA-3 in .idx, and reverse the .msha1 file to\n>>> translate offset to SHA-1.\n>>\n>> Thanks for this suggestion.  I was initially vaguely nervous about\n>> lookup times in an idx-style file, but as you say, object reads from a\n>> packfile already have to deal with this kind of lookup and work fine.\n>\n> Not exactly. The \"reverse .idx\" step has to build the reverse mapping on\n> the fly, and it's non-trivial.\n\nSure.  To be clear, I was handwaving over that since adding an on-disk\nreverse .idx is a relatively small change.\n\n[...]\n> So I think it's solvable, but I suspect we would want an extension to\n> the .idx format to store the mapping array, in order to keep it log-n.\n\ni.e., this.\n\nThe loose object side is the more worrying bit, since we currently don't\nhave any practical bound on the number of loose objects.\n\nOne way to deal with that is to disallow loose objects completely.\nUse packfiles for new objects, batching the objects produced by a\nsingle process into a single packfile.  Teach \"git gc --auto\" a\nbehavior similar to Martin Fick's \"git exproll\" to combine packfiles\nbetween full gcs to maintain reasonable performance.  For unreachable\nobjects, instead of using loose objects, use \"unreachable garbage\"\npacks explicitly labeled as such, with similar semantics to what\nJGit's DfsRepository backend uses (described in the discussion at\nhttps://git.eclipse.org/r/89455).\n\nThat's a direction that I want in the long term anyway.  I was hoping\nnot to couple such changes with the hash transition but it might be\none of the simpler ways to go.\n\nJonathan\n"},{"id":"313879","messageId":"91a34c5b-7844-3db2-cf29-411df5bcf886@noekeon.org","threadId":"45288","inReplyTo":"20170304011251.GA26789@aiede.mtv.corp.google.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"The Keccak Team","fromEmail":"keccak@noekeon.org","sentAt":"2017-03-13T09:24:25Z","receivedAt":"2017-03-13T09:31:18Z","isPatch":false,"sender":{"key":"keccak@noekeon.org","avatar":null},"body":"Hello,\n\nWe have read your transition plan to move away from SHA-1 and noticed\nyour intent to use SHA3-256 as the new hash function in the new Git\nrepository format and protocol. Although this is a valid choice, we\nthink that the new SHA-3 standard proposes alternatives that may also be\ninteresting for your use cases.  As designers of the Keccak function\nfamily, we thought we could jump in the mail thread and present these\nalternatives.\n\n\nSHA3-256, standardized in FIPS 202 [1], is a fixed-length hash function\nthat provides the same interface and security level as SHA-256 (FIPS\n180-4). SHA3-256's primary goal is to be drop-in compatible with the\nprevious standard, and to allow a fast transition for applications that\nwould already use SHA-256.\n\nSince your application did not use SHA-256, you are free to choose one\nof the alternatives listed below.\n\n\n* SHAKE128\n\n  SHAKE128, defined in FIPS 202, is an eXtendable-Output Function (XOF)\n  that generates digests of any size. In your case, you would use\n  SHAKE128 the same way you would use SHA3-256, just truncating the\n  output at 256 bits. In that case, SHAKE128 provides a security level\n  of 128 bits against all generic attacks, including collisions,\n  preimages, etc. We think this security level is appropriate for your\n  application since this is the maximum you can get with 256-bit tags in\n  the case of collision attacks, and this level is beyond computation\n  reach for any adversary in the foreseeable future.\n\n  The immediate benefit of using SHAKE128 versus SHA3-256 is a\n  performance gain of roughly 20%, both for SW and HW implementations.\n  On Intel Core i5-6500, SHAKE128 throughput is 430MiB/s.\n\n\n* ParallelHash128\n\n  ParallelHash128 (PH128), defined in NIST Special Publication 800-185\n  (SP800-185, SHA-3 Derived Functions [2]), is a XOF implementing a tree\n  hash mode on top of SHAKE128 (in fact cSHAKE128) to provide higher\n  performance for large-file hashing. The tree mode is designed to\n  exploit any available parallelism on the CPU, either through vector\n  instructions or availability of multiple cores. Note that the chosen\n  level of parallelism does not impact the final result, which improves\n  interoperability.\n\n  PH128 offers the same security level and interface as SHAKE128. So\n  likewise, you just truncate the output at 256 bits.\n\n  The net advantage of using PH128 over SHAKE128 is a huge performance\n  boost when hashing big files.  The advantage depends of course on the\n  number of cores used for hashing and their architecture. On an Intel\n  Core i5-6500 (Skylake), with a single-core, PH128 is faster than\n  SHAKE128 by a factor 3 and than SHA-1 by a factor 1.5 over long\n  messages, with a throughput of 1320MiB/s.\n\n\n* KangarooTwelve\n\n  KangarooTwelve (K12) [3] is a very fast parallel and secure XOF we\n  defined for applications that require higher performance that the FIPS\n  202 and SP800-185 functions provide, while retaining the same\n  flexibility and basis of security.\n\n  K12 is very similar to PH128. It uses the same cryptographic primitive\n  (Keccak-p, defined in FIPS 202), the same sponge construction, a\n  similar tree hashing mode, and targets the same generic security level\n  (128 bits). The main differences are the number of rounds for the\n  inner permutation, which is reduced to 12, and the tree mode\n  parameters, which are optimized for both small and long messages.\n\n  Again, the benefit of using K12 over PH128 is performance. K12 is\n  twice as fast as SHAKE128 for short messages, i.e. 820MiB/s on Intel\n  Core i5-6500, and twice as fast as PH128 over long messages, i.e.\n  2500MiB/s on the same platform.\n\n\nIf performance is not your primary concern, we suggest to use SHAKE128\nas the default hash function, and optionally use ParallelHash128 for\nhashing big files. Both functions offer a considerable security margin\nand are standardized algorithms. On the longer term, provided HW\nacceleration, SHAKE128 alone would easily outperform SHA-1 thanks to its\ndesign.\n\nIf however you value first performance, or if you would like to promote\nadoption of the new repository format by offering higher performance,\nthen KangarooTwelve is the right candidate. On modern CPU, K12 offers\nequal performance as SHA-1 for small messages and outperforms it by a\nfactor 3 for long messages.  Regarding security, although K12 offers of\ncourse a smaller security margin than other alternatives, it inherits\nthe security assurance built up for Keccak and the FIPS 202 functions.\nAs of today, the best practical attack broke 6 rounds of Keccak-p, with\n2^50 computation effort. The 12 rounds of K12 offers then a comfortable\nsecurity margin [4].\n\n\nLately, we made a presentation at FOSDEM covering the latest development\nover the Keccak family [5].  You can find reference and optimized\nimplementations of the algorithms listed above in the Keccak Code\nPackage [6]. Also, if you have questions, don't hesitate to contact us.\n\n\nKind regards,\nThe Keccak Team\n\nLinks\n [1]   FIPS 202,\n       http://nvlpubs.nist.gov/nistpubs/FIPS/NIST.FIPS.202.pdf.\n [2]   NIST SP 800-185,\n\nhttp://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-185.pdf.\n [3]   KangarooTwelve, http://keccak.noekeon.org/kangarootwelve.html.\n [4]   Keccak Crunchy Crypto Collision and Pre-image Contest,\n       http://keccak.noekeon.org/crunchy_contest.html.\n [5]   FOSDEM 2017, Portfolio of optimized cryptographic functions based\n       on Keccak, https://fosdem.org/2017/schedule/event/keccak/.\n [6]   Keccak Code Package, https://github.com/gvanas/KeccakCodePackage.\n\n\n"},{"id":"313908","messageId":"20170313174804.GH26789@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"91a34c5b-7844-3db2-cf29-411df5bcf886@noekeon.org","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-03-13T17:48:04Z","receivedAt":"2017-03-13T17:48:14Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nThe Keccak Team wrote:\n\n> We have read your transition plan to move away from SHA-1 and noticed\n> your intent to use SHA3-256 as the new hash function in the new Git\n> repository format and protocol. Although this is a valid choice, we\n> think that the new SHA-3 standard proposes alternatives that may also be\n> interesting for your use cases.  As designers of the Keccak function\n> family, we thought we could jump in the mail thread and present these\n> alternatives.\n\nI indeed had some reservations about SHA3-256's performance.  The main\nhash function we had in mind to compare against is blake2bp-256.  This\noverview of other functions to compare against should end up being\nvery helpful.\n\nThanks for this.  When I have more questions (which I most likely\nwill) I'll keep you posted.\n\nSincerely,\nJonathan\n"},{"id":"313921","messageId":"CA+dhYEViN4-boZLN+5QJyE7RtX+q6a92p0C2O6TA53==BZfTrQ@mail.gmail.com","threadId":"45288","inReplyTo":"20170313174804.GH26789@aiede.mtv.corp.google.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"ankostis","fromEmail":"ankostis@gmail.com","sentAt":"2017-03-13T18:34:16Z","receivedAt":"2017-03-13T18:34:53Z","isPatch":false,"sender":{"key":"ankostis@gmail.com","avatar":"https://gravatar.com/avatar/1f3597ab8ad44cbb0fd782a2773e3aba94a52c03f9f6bbe5ccb7f1fe2581b12b?d=mp&s=160"},"body":"On 13 March 2017 at 18:48, Jonathan Nieder <jrnieder@gmail.com> wrote:\n>\n> Hi,\n>\n> The Keccak Team wrote:\n>\n> > We have read your transition plan to move away from SHA-1 and noticed\n> > your intent to use SHA3-256 as the new hash function in the new Git\n> > repository format and protocol. Although this is a valid choice, we\n> > think that the new SHA-3 standard proposes alternatives that may also be\n> > interesting for your use cases.  As designers of the Keccak function\n> > family, we thought we could jump in the mail thread and present these\n> > alternatives.\n>\n> I indeed had some reservations about SHA3-256's performance.  The main\n> hash function we had in mind to compare against is blake2bp-256.  This\n> overview of other functions to compare against should end up being\n> very helpful.\n\nWhat if some of us need this extra difficulty, and don't mind about\nthe performance tax,\nbecause we need to refer to hashes 10 or 30 years from now,\nor even in the Post Quantum era?\n\nThanks,\n  Kostis\n"},{"id":"314320","messageId":"alpine.DEB.2.20.1703171123350.3767@virtualbox","threadId":"45288","inReplyTo":"CA+dhYEViN4-boZLN+5QJyE7RtX+q6a92p0C2O6TA53==BZfTrQ@mail.gmail.com","subject":"Re: RFC: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-03-17T11:07:48Z","receivedAt":"2017-03-17T11:15:11Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Kostis,\n\nOn Mon, 13 Mar 2017, ankostis wrote:\n\n> On 13 March 2017 at 18:48, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> >\n> > The Keccak Team wrote:\n> >\n> > > We have read your transition plan to move away from SHA-1 and\n> > > noticed your intent to use SHA3-256 as the new hash function in the\n> > > new Git repository format and protocol. Although this is a valid\n> > > choice, we think that the new SHA-3 standard proposes alternatives\n> > > that may also be interesting for your use cases.  As designers of\n> > > the Keccak function family, we thought we could jump in the mail\n> > > thread and present these alternatives.\n> >\n> > I indeed had some reservations about SHA3-256's performance.  The main\n> > hash function we had in mind to compare against is blake2bp-256.  This\n> > overview of other functions to compare against should end up being\n> > very helpful.\n> \n> What if some of us need this extra difficulty, and don't mind about the\n> performance tax, because we need to refer to hashes 10 or 30 years from\n> now, or even in the Post Quantum era?\n\nIf you need this extra difficulty, and if this extra difficulty would\nimply a huge penalty for everybody else, it is safe to assume that that\nextra difficulty would need to be an extra switch, off by default.\n\nIt simply shows that we put too much of a burden on SHA-1: we used it for\nthree separate purposes: to verify data integrity, to allow addressing\nobjects by their own content, and for signing entire commit histories\ncryptographically (more as an afterthought, as I see it: the Linux project\nprovides the context where you never fetch from any untrusted source,\ntherefore cryptographically secure signatures are not quite as important\nas the trust between maintainer and lieutenants).\n\nWe *will* have to separate those concerns, and maybe even switch to\ndifferent algorithms for the different concerns. There are much better\nalgorithms for validating data integrity, for example, including error\ncorrection (which SHA-1 never wanted to do anyway).\n\nIn your case, I could imagine that you would simply require verifiable\ncryptographic signatures (.asc files) to be committed together with the\ndocuments; it would be much harder to find a collision where those\nsignatures still match (or a double collision where the forged document's\nsignature would collide with the non-forget document's signature, in\naddition to the two documents colliding).\n\nAnother idea would be to use Jonathan Nieder's proposed transition plan\nand simply extend it. That transition plan details how the objects would\nbe hashed with two algorithms locally and how to maintain a bidirectional\nmapping between the two. You could simply piggyback on that code and\nprovide patches that allow for a third, configurable algorithm, and that\nalgorithm's hashes would simply be added to the commit objects and fsck\nwould then know to verify those, too. That would be an opt-in feature, of\ncourse, so that only those who need the extra long term security have to\npay the price of a substantially slower hashing.\n\nWhat we cannot do is to pick a super slow hash algorithm just to cater to\nthe use case where legal documents are managed, punishing everybody else\nfor using Git in the intended way: to manage source code.\n\nCiao,\nJohannes\n"},{"id":"314627","messageId":"27077870-76d9-b45a-5727-c339a3d0ffc8@bu.edu","threadId":"45288","inReplyTo":"alpine.DEB.2.20.1703081638010.3767@virtualbox","subject":"Use base32?","fromName":"Jason Hennessey","fromEmail":"henn@bu.edu","sentAt":"2017-03-20T05:21:31Z","receivedAt":"2017-03-20T05:21:45Z","isPatch":false,"sender":{"key":"henn@bu.edu","avatar":null},"body":"\nOn Wed, 8 Mar 2017, Johannes Schindelin wrote:\n> > Linus Torvalds writes (\"Re: RFC: Another proposed hash function transition plan\"): > > Of course, having written that, I now realize how it would cause\n> > > problems for the usual shit-for-brains case-insensitive\n> filesystems. > > So I guess base64 encoding doesn't work well for that\n> reason.\n> Given that the idea was to encode the new hash in base64 or base85, we\n> *are* talking about an encoding. In that respect, yes, it can be whatever\n> encoding we like, and Linus just made a good point (with unnecessary foul\n> language) of explaining why base64/base85 is not that encoding.\n\nSince the hash format is switching anyway, how about using base32\ninstead of hex?\n\nStill get a 20% space savings over hex (minus a little for padding), and\nit's guaranteed to be a single case.\nJason\n\n"},{"id":"314631","messageId":"1c213acb-1deb-8959-b1f8-28f99974640f@constrainttec.com","threadId":"45288","inReplyTo":"27077870-76d9-b45a-5727-c339a3d0ffc8@bu.edu","subject":"Re: Use base32?","fromName":"Michael Steuer","fromEmail":"michael.steuer@constrainttec.com","sentAt":"2017-03-20T05:58:06Z","receivedAt":"2017-03-20T06:07:45Z","isPatch":false,"sender":{"key":"michael.steuer@constrainttec.com","avatar":null},"body":"\nOn 20/03/2017 16:21, Jason Hennessey wrote:\n> On Wed, 8 Mar 2017, Johannes Schindelin wrote:\n>>> Linus Torvalds writes (\"Re: RFC: Another proposed hash function transition plan\"): > > Of course, having written that, I now realize how it would cause\n>>>> problems for the usual shit-for-brains case-insensitive\n>> filesystems. > > So I guess base64 encoding doesn't work well for that\n>> reason.\n>> Given that the idea was to encode the new hash in base64 or base85, we\n>> *are* talking about an encoding. In that respect, yes, it can be whatever\n>> encoding we like, and Linus just made a good point (with unnecessary foul\n>> language) of explaining why base64/base85 is not that encoding.\n> Since the hash format is switching anyway, how about using base32\n> instead of hex?\n>\n> Still get a 20% space savings over hex (minus a little for padding), and\n> it's guaranteed to be a single case.\n> Jason\n>\n\nIf base32 is being considered, I'd suggest the \"base32hex\" variant, \nwhich uses the same amount of space.\n"},{"id":"314633","messageId":"CA+P7+xoCsO=LFw1aSQugmvZz+kjhNXT+Ffwa3DjDUPOdpobzAg@mail.gmail.com","threadId":"45288","inReplyTo":"1c213acb-1deb-8959-b1f8-28f99974640f@constrainttec.com","subject":"Re: Use base32?","fromName":"Jacob Keller","fromEmail":"jacob.keller@gmail.com","sentAt":"2017-03-20T08:05:02Z","receivedAt":"2017-03-20T08:05:35Z","isPatch":false,"sender":{"key":"jacob.keller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/874719?v=4"},"body":"On Sun, Mar 19, 2017 at 10:58 PM, Michael Steuer\n<Michael.Steuer@constrainttec.com> wrote:\n>\n> On 20/03/2017 16:21, Jason Hennessey wrote:\n>>\n>> On Wed, 8 Mar 2017, Johannes Schindelin wrote:\n>>>>\n>>>> Linus Torvalds writes (\"Re: RFC: Another proposed hash function\n>>>> transition plan\"): > > Of course, having written that, I now realize how it\n>>>> would cause\n>>>>>\n>>>>> problems for the usual shit-for-brains case-insensitive\n>>>\n>>> filesystems. > > So I guess base64 encoding doesn't work well for that\n>>> reason.\n>>> Given that the idea was to encode the new hash in base64 or base85, we\n>>> *are* talking about an encoding. In that respect, yes, it can be whatever\n>>> encoding we like, and Linus just made a good point (with unnecessary foul\n>>> language) of explaining why base64/base85 is not that encoding.\n>>\n>> Since the hash format is switching anyway, how about using base32\n>> instead of hex?\n>>\n>> Still get a 20% space savings over hex (minus a little for padding), and\n>> it's guaranteed to be a single case.\n>> Jason\n>>\n>\n> If base32 is being considered, I'd suggest the \"base32hex\" variant, which\n> uses the same amount of space.\n\nI don't see the benefit of adding characters like 0 and 1 which\nconflict with some of the letters? Since there's no need for a human\nto decode the base32 output, it's easier to use the one that's less\nlikely to get screwed up when typing if that ever happens. It's not\nlike we actually need to know what value each character represents.\n(sure the program does, but the human does not).\n\nThanks,\nJake\n"},{"id":"314775","messageId":"f573b82c-591f-22f0-84fb-204c23d1de81@constrainttec.com","threadId":"45288","inReplyTo":"CA+P7+xoCsO=LFw1aSQugmvZz+kjhNXT+Ffwa3DjDUPOdpobzAg@mail.gmail.com","subject":"Re: Use base32?","fromName":"Michael Steuer","fromEmail":"michael.steuer@constrainttec.com","sentAt":"2017-03-21T03:07:16Z","receivedAt":"2017-03-21T03:07:32Z","isPatch":false,"sender":{"key":"michael.steuer@constrainttec.com","avatar":null},"body":"\nOn 20/03/2017 19:05, Jacob Keller wrote:\n> On Sun, Mar 19, 2017 at 10:58 PM, Michael Steuer\n> <Michael.Steuer@constrainttec.com> wrote:\n>> [..]\n>> If base32 is being considered, I'd suggest the \"base32hex\" variant, which\n>> uses the same amount of space.\n> I don't see the benefit of adding characters like 0 and 1 which\n> conflict with some of the letters? Since there's no need for a human\n> to decode the base32 output, it's easier to use the one that's less\n> likely to get screwed up when typing if that ever happens. It's not\n> like we actually need to know what value each character represents.\n> (sure the program does, but the human does not).\n>\n> Thanks,\n> Jake\n\nFair enough and good point. We definitely wouldn't want 0 and O and 1 \nand I mixed together.\n\nCheers,\nMike.\n"},{"id":"322304","messageId":"alpine.DEB.2.21.1.1706151122180.4200@virtualbox","threadId":"45288","inReplyTo":"20170306182423.GB183239@google.com","subject":"Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-06-15T10:30:46Z","receivedAt":"2017-06-15T10:31:36Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nI thought it better to revive this old thread rather than start a new\nthread, so as to automatically reach everybody who chimed in originally.\n\nOn Mon, 6 Mar 2017, Brandon Williams wrote:\n\n> On 03/06, brian m. carlson wrote:\n>\n> > On Sat, Mar 04, 2017 at 06:35:38PM -0800, Linus Torvalds wrote:\n> >\n> > > Btw, I do think the particular choice of hash should still be on the\n> > > table. sha-256 may be the obvious first choice, but there are\n> > > definitely a few reasons to consider alternatives, especially if\n> > > it's a complete switch-over like this.\n> > > \n> > > One is large-file behavior - a parallel (or tree) mode could improve\n> > > on that noticeably. BLAKE2 does have special support for that, for\n> > > example. And SHA-256 does have known attacks compared to SHA-3-256\n> > > or BLAKE2 - whether that is due to age or due to more effort, I\n> > > can't really judge. But if we're switching away from SHA1 due to\n> > > known attacks, it does feel like we should be careful.\n> > \n> > I agree with Linus on this.  SHA-256 is the slowest option, and it's\n> > the one with the most advanced cryptanalysis.  SHA-3-256 is faster on\n> > 64-bit machines (which, as we've seen on the list, is the overwhelming\n> > majority of machines using Git), and even BLAKE2b-256 is stronger.\n> > \n> > Doing this all over again in another couple years should also be a\n> > non-goal.\n> \n> I agree that when we decide to move to a new algorithm that we should\n> select one which we plan on using for as long as possible (much longer\n> than a couple years).  While writing the document we simply used\n> \"sha256\" because it was more tangible and easier to reference.\n\nThe SHA-1 transition *requires* a knob telling Git that the current\nrepository uses a hash function different from SHA-1.\n\nIt would make *a whole of a lot of sense* to make that knob *not* Boolean,\nbut to specify *which* hash function is in use.\n\nThat way, it will be easier to switch another time when it becomes\nnecessary.\n\nAnd it will also make it easier for interested parties to use a different\nhash function in their infrastructure if they want.\n\nAnd it lifts part of that burden that we have to consider *very carefully*\nwhich function to pick. We still should be more careful than in 2005, when\nGit was born, and when, incidentally, when the first attacks on SHA-1\nbecame known, of course. We were just lucky for almost 12 years.\n\nNow, with Dunning-Kruger in mind, I feel that my degree in mathematics\nequips me with *just enough* competence to know just how little *even I*\nknow about cryptography.\n\nThe smart thing to do, hence, was to get involved in this discussion and\nact as Lt Tawney Madison between us Git developers and experts in\ncryptography.\n\nIt just so happens that I work at a company with access to excellent\ncryptographers, and as we own the largest Git repository on the planet, we\nhave a vested interest in ensuring Git's continued success.\n\nAfter a couple of conversations with a couple of experts who I cannot\nthank enough for their time and patience, let alone their knowledge about\nthis matter, it would appear that we may not have had a complete enough\npicture yet to even start to make the decision on the hash function to\nuse.\n\nFrom what I read, pretty much everybody who participated in the discussion\nwas aware that the essential question is: performance vs security.\n\nIt turns out that we can have essentially both.\n\nSHA-256 is most likely the best-studied hash function we currently know\nabout (*maybe* SHA3-256 has been studied slightly more, but only\nslightly). All the experts in the field banged on it with multiple sticks\nand other weapons. And so far, they only found one weakness that does not\neven apply to Git's usage [*1*]. For cryptography experts, this is the\nultimate measure of security: if something has been attacked that\nintensely, by that many experts, for that long, with that little effect,\nit is the best we got at the time.\n\nAnd since SHA-256 has become the standard, and more importantly: since\nSHA-256 was explicitly designed to allow for relatively inexpensive\nhardware acceleration, this is what we will soon have: hardware support in\nthe form of, say, special CPU instructions. (That is what I meant by: we\ncan have performance *and* security.)\n\nThis is a rather important point to stress, by the way: BLAKE's design is\napparently *not* friendly to CPU instruction implementations. Meaning that\nSHA-256 will be faster than BLAKE (and even than BLAKE2) once the Intel\nand AMD CPUs with hardware support for SHA-256 become common.\n\nI also heard something really worrisome about BLAKE2 that makes me want to\nstay away from it (in addition to the difficulty it poses for hardware\nacceleration): to compete in the SHA-3 contest, BLAKE added complexity so\nthat it would be roughly on par with its competitors. To allow for faster\nexecution in software, this complexity was *removed* from BLAKE to create\nBLAKE2, making it weaker than SHA-256.\n\nAnother important point to consider is that SHA-256 implementations are\neverywhere. Think e.g. how difficult we would make it on, say, JGit or\ngo-git if we chose a less common hash function.\n\nAs to KangarooTwelve: it has seen substantially less cryptanalysis than\nSHA-256 and SHA3-256. That does not necessarily mean that it is weaker,\nbut it means that we simply cannot know whether it is as strong. On that\nbasis alone, I would already reject it, and then there are far fewer\nimplementations, too.\n\nWhen it comes to choosing SHA-256 vs SHA3-256, I would like to point out\nthat hardware acceleration is a lot farther in the future than SHA-256\nsupport. And according to the experts I asked, they are roughly equally\nsecure as far as Git's usage is concerned, even if the SHA-3 contest\nprovided SHA3-256 with even fiercer cryptanalysis than SHA-256.\n\nIn short: my takeaway from the conversations with cryptography experts was\nthat SHA-256 would be the best choice for now, and that we should make\nsure that the next switch is not as painful as this one (read: we should\nnot repeat the mistake of hard-coding the new hash function into Git as\nmuch as we hard-coded SHA-1 into it).\n\nCiao,\nJohannes\n\nFootnote *1*: SHA-256, as all hash functions whose output is essentially\nthe entire internal state, are susceptible to a so-called \"length\nextension attack\", where the hash of a secret+message can be used to\ngenerate the hash of secret+message+piggyback without knowing the secret.\nThis is not the case for Git: only visible data are hashed. The type of\nattacks Git has to worry about is very different from the length extension\nattacks, and it is highly unlikely that that weakness of SHA-256 leads to,\nsay, a collision attack.\n"},{"id":"322310","messageId":"20170615110518.ordr43idf2jluips@glandium.org","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1706151122180.4200@virtualbox","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2017-06-15T11:05:18Z","receivedAt":"2017-06-15T11:05:45Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Thu, Jun 15, 2017 at 12:30:46PM +0200, Johannes Schindelin wrote:\n> Footnote *1*: SHA-256, as all hash functions whose output is essentially\n> the entire internal state, are susceptible to a so-called \"length\n> extension attack\", where the hash of a secret+message can be used to\n> generate the hash of secret+message+piggyback without knowing the secret.\n> This is not the case for Git: only visible data are hashed. The type of\n> attacks Git has to worry about is very different from the length extension\n> attacks, and it is highly unlikely that that weakness of SHA-256 leads to,\n> say, a collision attack.\n\nWhat do the experts think or SHA512/256, which completely removes the\nconcerns over length extension attack? (which I'd argue is better than\nsweeping them under the carpet)\n\nMike\n"},{"id":"322318","messageId":"20170615130145.stwbtict7q6oel7e@sigill.intra.peff.net","threadId":"45288","inReplyTo":"20170615110518.ordr43idf2jluips@glandium.org","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-06-15T13:01:45Z","receivedAt":"2017-06-15T13:01:58Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Thu, Jun 15, 2017 at 08:05:18PM +0900, Mike Hommey wrote:\n\n> On Thu, Jun 15, 2017 at 12:30:46PM +0200, Johannes Schindelin wrote:\n> > Footnote *1*: SHA-256, as all hash functions whose output is essentially\n> > the entire internal state, are susceptible to a so-called \"length\n> > extension attack\", where the hash of a secret+message can be used to\n> > generate the hash of secret+message+piggyback without knowing the secret.\n> > This is not the case for Git: only visible data are hashed. The type of\n> > attacks Git has to worry about is very different from the length extension\n> > attacks, and it is highly unlikely that that weakness of SHA-256 leads to,\n> > say, a collision attack.\n> \n> What do the experts think or SHA512/256, which completely removes the\n> concerns over length extension attack? (which I'd argue is better than\n> sweeping them under the carpet)\n\nI don't think it's sweeping them under the carpet. Git does not use the\nhash as a MAC, so length extension attacks aren't a thing (and even if\nwe later wanted to use the same algorithm as a MAC, the HMAC\nconstruction is a well-studied technique for dealing with it).\n\nThat said, SHA-512 is typically a little faster than SHA-256 on 64-bit\nplatforms. I don't know if that will change with the advent of hardware\ninstructions oriented towards SHA-256.\n\n-Peff\n"},{"id":"322360","messageId":"87shj1ciy8.fsf@gmail.com","threadId":"45288","inReplyTo":"20170615130145.stwbtict7q6oel7e@sigill.intra.peff.net","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2017-06-15T16:30:07Z","receivedAt":"2017-06-15T16:30:21Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, Jun 15 2017, Jeff King jotted:\n\n> On Thu, Jun 15, 2017 at 08:05:18PM +0900, Mike Hommey wrote:\n>\n>> On Thu, Jun 15, 2017 at 12:30:46PM +0200, Johannes Schindelin wrote:\n>> > Footnote *1*: SHA-256, as all hash functions whose output is essentially\n>> > the entire internal state, are susceptible to a so-called \"length\n>> > extension attack\", where the hash of a secret+message can be used to\n>> > generate the hash of secret+message+piggyback without knowing the secret.\n>> > This is not the case for Git: only visible data are hashed. The type of\n>> > attacks Git has to worry about is very different from the length extension\n>> > attacks, and it is highly unlikely that that weakness of SHA-256 leads to,\n>> > say, a collision attack.\n>>\n>> What do the experts think or SHA512/256, which completely removes the\n>> concerns over length extension attack? (which I'd argue is better than\n>> sweeping them under the carpet)\n>\n> I don't think it's sweeping them under the carpet. Git does not use the\n> hash as a MAC, so length extension attacks aren't a thing (and even if\n> we later wanted to use the same algorithm as a MAC, the HMAC\n> construction is a well-studied technique for dealing with it).\n>\n> That said, SHA-512 is typically a little faster than SHA-256 on 64-bit\n> platforms. I don't know if that will change with the advent of hardware\n> instructions oriented towards SHA-256.\n\nQuoting my own\nCACBZZX7JRA2niwt9wsGAxnzS+gWS8hTUgzWm8NaY1gs87o8xVQ@mail.gmail.com sent\n~2 weeks ago to the list:\n\n    On Fri, Jun 2, 2017 at 7:54 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n    [...]\n    > 4. When choosing a hash function, people may argue about performance.\n    >    It would be useful for run some benchmarks for git (running\n    >    the test suite, t/perf tests, etc) using a variety of hash\n    >    functions as input to such a discussion.\n\n    To the extent that such benchmarks matter, it seems prudent to heavily\n    weigh them in favor of whatever seems to be likely to be the more\n    common hash function going forward, since those are likely to get\n    faster through future hardware acceleration.\n\n    E.g. Intel announced Goldmont last year which according to one SHA-1\n    implementation improved from 9.5 cycles per byte to 2.7 cpb[1]. They\n    only have acceleration for SHA-1 and SHA-256[2]\n\n    1. https://github.com/weidai11/cryptopp/issues/139#issuecomment-264283385\n\n    2. https://en.wikipedia.org/wiki/Goldmont\n\nMaybe someone else knows of better numbers / benchmarks, but such a\nreduction in CBP likely makes it faster than SHA-512.\n"},{"id":"322367","messageId":"20170615173616.GA176947@google.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1706151122180.4200@virtualbox","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Brandon Williams","fromEmail":"bmwill@google.com","sentAt":"2017-06-15T17:36:16Z","receivedAt":"2017-06-15T17:36:24Z","isPatch":false,"sender":{"key":"bwilliams.eng@gmail.com","avatar":null},"body":"On 06/15, Johannes Schindelin wrote:\n> Hi,\n> \n> I thought it better to revive this old thread rather than start a new\n> thread, so as to automatically reach everybody who chimed in originally.\n> \n> On Mon, 6 Mar 2017, Brandon Williams wrote:\n> \n> > On 03/06, brian m. carlson wrote:\n> >\n> > > On Sat, Mar 04, 2017 at 06:35:38PM -0800, Linus Torvalds wrote:\n> > >\n> > > > Btw, I do think the particular choice of hash should still be on the\n> > > > table. sha-256 may be the obvious first choice, but there are\n> > > > definitely a few reasons to consider alternatives, especially if\n> > > > it's a complete switch-over like this.\n> > > > \n> > > > One is large-file behavior - a parallel (or tree) mode could improve\n> > > > on that noticeably. BLAKE2 does have special support for that, for\n> > > > example. And SHA-256 does have known attacks compared to SHA-3-256\n> > > > or BLAKE2 - whether that is due to age or due to more effort, I\n> > > > can't really judge. But if we're switching away from SHA1 due to\n> > > > known attacks, it does feel like we should be careful.\n> > > \n> > > I agree with Linus on this.  SHA-256 is the slowest option, and it's\n> > > the one with the most advanced cryptanalysis.  SHA-3-256 is faster on\n> > > 64-bit machines (which, as we've seen on the list, is the overwhelming\n> > > majority of machines using Git), and even BLAKE2b-256 is stronger.\n> > > \n> > > Doing this all over again in another couple years should also be a\n> > > non-goal.\n> > \n> > I agree that when we decide to move to a new algorithm that we should\n> > select one which we plan on using for as long as possible (much longer\n> > than a couple years).  While writing the document we simply used\n> > \"sha256\" because it was more tangible and easier to reference.\n> \n> The SHA-1 transition *requires* a knob telling Git that the current\n> repository uses a hash function different from SHA-1.\n> \n> It would make *a whole of a lot of sense* to make that knob *not* Boolean,\n> but to specify *which* hash function is in use.\n\n100% agree on this point.  I believe the current plan is to have the\nhashing function used for a repository be a repository format extension\nwhich would be a value (most likely a string like 'sha1', 'sha256',\n'black2', etc) stored in a repository's .git/config.  This way, upon\nstartup git will die or ignore a repository which uses a hashing\nfunction which it does not recognize or does not compiled to handle.\n\nI hope (and expect) that the end produce of this transition is a nice,\nclean hashing API and interface with sufficient abstractions such that\nif I wanted to switch to a different hashing function I would just need\nto implement the interface with the new hashing function and ensure that\n'verify_repository_format' allows the new function.\n\n> \n> That way, it will be easier to switch another time when it becomes\n> necessary.\n> \n> And it will also make it easier for interested parties to use a different\n> hash function in their infrastructure if they want.\n> \n> And it lifts part of that burden that we have to consider *very carefully*\n> which function to pick. We still should be more careful than in 2005, when\n> Git was born, and when, incidentally, when the first attacks on SHA-1\n> became known, of course. We were just lucky for almost 12 years.\n> \n> Now, with Dunning-Kruger in mind, I feel that my degree in mathematics\n> equips me with *just enough* competence to know just how little *even I*\n> know about cryptography.\n> \n> The smart thing to do, hence, was to get involved in this discussion and\n> act as Lt Tawney Madison between us Git developers and experts in\n> cryptography.\n> \n> It just so happens that I work at a company with access to excellent\n> cryptographers, and as we own the largest Git repository on the planet, we\n> have a vested interest in ensuring Git's continued success.\n> \n> After a couple of conversations with a couple of experts who I cannot\n> thank enough for their time and patience, let alone their knowledge about\n> this matter, it would appear that we may not have had a complete enough\n> picture yet to even start to make the decision on the hash function to\n> use.\n> \n\n-- \nBrandon Williams\n"},{"id":"322375","messageId":"20170615191151.GA45444@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1706151122180.4200@virtualbox","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-06-15T19:13:05Z","receivedAt":"2017-06-15T19:13:33Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi Dscho,\n\nJohannes Schindelin wrote:\n\n> From what I read, pretty much everybody who participated in the discussion\n> was aware that the essential question is: performance vs security.\n\nI don't completely agree with this framing.  The essential question is:\nhow to get the right security properties without abysmal performance.\n\n> It turns out that we can have essentially both.\n>\n> SHA-256 is most likely the best-studied hash function we currently know\n[... etc ...]\n\nThanks for a thoughtful restart to the discussion.  This is much more\nconcrete than your previous objections about process, and that is very\nhelpful.\n\nIn the interest of transparency: here are my current questions for\ncryptographers to whom I have forwarded this thread.  Several of these\nquestions involve predictions or opinions, so in my ideal world we'd\nwant multiple, well reasoned answers to them.  Please feel free to\nforward them to appropriate people or add more.\n\n 1. Now it sounds like SHA-512/256 is the safest choice (see also Mike\n    Hommey's response to Dscho's message).  Please poke holes in my\n    understanding.\n\n 2. Would you be willing to weigh in publicly on the mailing list? I\n    think that would be the most straightforward way to move this\n    forward (and it would give you a chance to ask relevant questions,\n    etc).  Feel free to contact me privately if you have any questions\n    about how this particular mailing list works.\n\n 3. On the speed side, Dscho states \"SHA-256 will be faster than BLAKE\n    (and even than BLAKE2) once the Intel and AMD CPUs with hardware\n    support for SHA-256 become common.\"  Do you agree?\n\n 4. On the security side, Dscho states \"to compete in the SHA-3\n    contest, BLAKE added complexity so that it would be roughly on par\n    with its competitors.  To allow for faster execution in software,\n    this complexity was *removed* from BLAKE to create BLAKE2, making\n    it weaker than SHA-256.\"  Putting aside the historical questions,\n    do you agree with this \"weaker than\" claim?\n\n 5. On the security side, Dscho states, \"The type of attacks Git has to\n    worry about is very different from the length extension attacks,\n    and it is highly unlikely that that weakness of SHA-256 leads to,\n    say, a collision attack\", and Jeff King states, \"Git does not use\n    the hash as a MAC, so length extension attacks aren't a thing (and\n    even if we later wanted to use the same algorithm as a MAC, the\n    HMAC construction is a well-studied technique for dealing with\n    it).\"  Is this correct in spirit?  Is SHA-256 equally strong to\n    SHA-512/256 for Git's purposes, or are the increased bits of\n    internal state (or other differences) relevant?  How would you\n    compare the two functions' properties?\n\n 6. On the speed side, Jeff King states \"That said, SHA-512 is\n    typically a little faster than SHA-256 on 64-bit platforms. I\n    don't know if that will change with the advent of hardware\n    instructions oriented towards SHA-256.\"  Thoughts?\n\n 7. If the answer to (2) is \"no\", do I have permission to quote or\n    paraphrase your replies that were given here?\n\nThanks, sincerely,\nJonathan\n"},{"id":"322378","messageId":"xmqqh8zh12ik.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170615173616.GA176947@google.com","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-06-15T19:20:35Z","receivedAt":"2017-06-15T19:20:42Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Brandon Williams <bmwill@google.com> writes:\n\n>> It would make a whole of a lot of sense to make that knob not Boolean,\n>> but to specify which hash function is in use.\n>\n> 100% agree on this point.  I believe the current plan is to have the\n> hashing function used for a repository be a repository format extension\n> which would be a value (most likely a string like 'sha1', 'sha256',\n> 'black2', etc) stored in a repository's .git/config.  This way, upon\n> startup git will die or ignore a repository which uses a hashing\n> function which it does not recognize or does not compiled to handle.\n>\n> I hope (and expect) that the end produce of this transition is a nice,\n> clean hashing API and interface with sufficient abstractions such that\n> if I wanted to switch to a different hashing function I would just need\n> to implement the interface with the new hashing function and ensure that\n> 'verify_repository_format' allows the new function.\n\nYup.  I thought that part has already been agreed upon, but it is a\ngood thing that somebody is writing it down (perhaps \"again\", if not\n\"for the first time\").\n\nThanks.\n\n"},{"id":"322379","messageId":"alpine.DEB.2.21.1.1706152123060.4200@virtualbox","threadId":"45288","inReplyTo":"87shj1ciy8.fsf@gmail.com","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-06-15T19:34:38Z","receivedAt":"2017-06-15T19:35:31Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 15 Jun 2017, Ævar Arnfjörð Bjarmason wrote:\n\n> On Thu, Jun 15 2017, Jeff King jotted:\n> \n> > On Thu, Jun 15, 2017 at 08:05:18PM +0900, Mike Hommey wrote:\n> >\n> >> On Thu, Jun 15, 2017 at 12:30:46PM +0200, Johannes Schindelin wrote:\n> >>\n> >> > Footnote *1*: SHA-256, as all hash functions whose output is\n> >> > essentially the entire internal state, are susceptible to a\n> >> > so-called \"length extension attack\", where the hash of a\n> >> > secret+message can be used to generate the hash of\n> >> > secret+message+piggyback without knowing the secret.  This is not\n> >> > the case for Git: only visible data are hashed. The type of attacks\n> >> > Git has to worry about is very different from the length extension\n> >> > attacks, and it is highly unlikely that that weakness of SHA-256\n> >> > leads to, say, a collision attack.\n> >>\n> >> What do the experts think or SHA512/256, which completely removes the\n> >> concerns over length extension attack? (which I'd argue is better than\n> >> sweeping them under the carpet)\n> >\n> > I don't think it's sweeping them under the carpet. Git does not use the\n> > hash as a MAC, so length extension attacks aren't a thing (and even if\n> > we later wanted to use the same algorithm as a MAC, the HMAC\n> > construction is a well-studied technique for dealing with it).\n\nI really tried to drive that point home, as it had been made very clear to\nme that the length extension attack is something that Git need not concern\nitself.\n\nThe length extension attack *only* comes into play when there are secrets\nthat are hashed. In that case, one would not want others to be able to\nproduce a valid hash *without* knowing the secrets. And SHA-256 allows to\n\"reconstruct\" the internal state (which is the hash value) in order to\ncontinue at any point, i.e. if the hash for secret+message is known, it is\neasy to calculate the hash for secret+message+addition, without knowing\nthe secret at all.\n\nThat is exactly *not* the case with Git. In Git, what we want to hash is\nknown in its entirety. If the hash value were not identical to the\ninternal state, it would be easy enough to reconstruct, because *there are\nno secrets*.\n\nSo please understand that even the direction that the length extension\nattack takes is completely different than the direction any attack would\nhave to take that weakens SHA-256 for Git's purposes. As far as Git's\nusage is concerned, SHA-256 has no known weaknesses.\n\nIt is *really, really, really* important to understand this before going\non to suggest another hash function such as SHA-512/256 (i.e. SHA-512\ntruncated to 256 bits), based only on that perceived weakness of SHA-256.\n\n> > That said, SHA-512 is typically a little faster than SHA-256 on 64-bit\n> > platforms. I don't know if that will change with the advent of\n> > hardware instructions oriented towards SHA-256.\n> \n> Quoting my own\n> CACBZZX7JRA2niwt9wsGAxnzS+gWS8hTUgzWm8NaY1gs87o8xVQ@mail.gmail.com sent\n> ~2 weeks ago to the list:\n> \n>     On Fri, Jun 2, 2017 at 7:54 PM, Jonathan Nieder <jrnieder@gmail.com>\n>     wrote:\n>     [...]\n>     > 4. When choosing a hash function, people may argue about performance.\n>     >    It would be useful for run some benchmarks for git (running\n>     >    the test suite, t/perf tests, etc) using a variety of hash\n>     >    functions as input to such a discussion.\n> \n>     To the extent that such benchmarks matter, it seems prudent to heavily\n>     weigh them in favor of whatever seems to be likely to be the more\n>     common hash function going forward, since those are likely to get\n>     faster through future hardware acceleration.\n> \n>     E.g. Intel announced Goldmont last year which according to one SHA-1\n>     implementation improved from 9.5 cycles per byte to 2.7 cpb[1]. They\n>     only have acceleration for SHA-1 and SHA-256[2]\n> \n>     1. https://github.com/weidai11/cryptopp/issues/139#issuecomment-264283385\n> \n>     2. https://en.wikipedia.org/wiki/Goldmont\n> \n> Maybe someone else knows of better numbers / benchmarks, but such a\n> reduction in CBP likely makes it faster than SHA-512.\n\nVery, very likely faster than SHA-512.\n\nI'd like to stress explicitly that the Intel SHA extensions do *not* cover\nSHA-512:\n\n\thttps://en.wikipedia.org/wiki/Intel_SHA_extensions\n\nIn other words, once those extensions become commonplace, SHA-256 will be\nfaster than SHA-512, hands down.\n\nCiao,\nDscho"},{"id":"322398","messageId":"20170615211022.vmedlcwmvtdiseqx@glandium.org","threadId":"45288","inReplyTo":"20170615130145.stwbtict7q6oel7e@sigill.intra.peff.net","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2017-06-15T21:10:22Z","receivedAt":"2017-06-15T21:10:57Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Thu, Jun 15, 2017 at 09:01:45AM -0400, Jeff King wrote:\n> On Thu, Jun 15, 2017 at 08:05:18PM +0900, Mike Hommey wrote:\n> \n> > On Thu, Jun 15, 2017 at 12:30:46PM +0200, Johannes Schindelin wrote:\n> > > Footnote *1*: SHA-256, as all hash functions whose output is essentially\n> > > the entire internal state, are susceptible to a so-called \"length\n> > > extension attack\", where the hash of a secret+message can be used to\n> > > generate the hash of secret+message+piggyback without knowing the secret.\n> > > This is not the case for Git: only visible data are hashed. The type of\n> > > attacks Git has to worry about is very different from the length extension\n> > > attacks, and it is highly unlikely that that weakness of SHA-256 leads to,\n> > > say, a collision attack.\n> > \n> > What do the experts think or SHA512/256, which completely removes the\n> > concerns over length extension attack? (which I'd argue is better than\n> > sweeping them under the carpet)\n> \n> I don't think it's sweeping them under the carpet. Git does not use the\n> hash as a MAC, so length extension attacks aren't a thing (and even if\n> we later wanted to use the same algorithm as a MAC, the HMAC\n> construction is a well-studied technique for dealing with it).\n\nAIUI, length extension does make brute force collision attacks (which,\nreally Shattered was) cheaper by allowing one to create the collision\nwith a small message and extend it later.\n\nThis might not be a credible thread against git, but if we go by that\nstandard, post-shattered Sha-1 is still fine for git. As a matter of\nfact, MD5 would also be fine: there is still, to this day, no preimage\nattack against them.\n\nMike\n"},{"id":"322405","messageId":"CAL9PXLzhPyE+geUdcLmd=pidT5P8eFEBbSgX_dS88knz2q_LSw@mail.gmail.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1706152123060.4200@virtualbox","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Adam Langley","fromEmail":"agl@google.com","sentAt":"2017-06-15T21:59:57Z","receivedAt":"2017-06-15T22:00:25Z","isPatch":false,"sender":{"key":"agl@google.com","avatar":null},"body":"(I was asked to comment a few points in public by Jonathan.)\n\nI think this group can safely assume that SHA-256, SHA-512, BLAKE2,\nK12, etc are all secure to the extent that I don't believe that making\ncomparisons between them on that axis is meaningful. Thus I think the\nquestion is primarily concerned with performance and implementation\navailability.\n\nI think any of the above would be reasonable choices. I don't believe\nthat length-extension is a concern here.\n\nSHA-512/256 will be faster than SHA-256 on 64-bit systems in software.\nThe graph at https://blake2.net/ suggests a 50% speedup on Skylake. On\nmy Ivy Bridge system, it's about 20%.\n\n(SHA-512/256 does not enjoy the same availability in common libraries however.)\n\nBoth Intel and ARM have SHA-256 instructions defined. I've not seen\ngood benchmarks of them yet, but they will make SHA-256 faster than\nSHA-512 when available. However, it's very possible that something\nlike BLAKE2bp will still be faster. Of course, BLAKE2bp does not enjoy\nthe ubiquity of SHA-256, but nor do you have to wait years for the CPU\npopulation to advance for high performance.\n\nSo, overall, none of these choices should obviously be excluded. The\nconsiderations at this point are not cryptographic and the tradeoff\nbetween implementation ease and performance is one that the git\ncommunity would have to make.\n\n\nCheers\n\nAGL\n"},{"id":"322410","messageId":"20170615224110.kvrjs3lmwxcoqfaw@genre.crustytoothpaste.net","threadId":"45288","inReplyTo":"CAL9PXLzhPyE+geUdcLmd=pidT5P8eFEBbSgX_dS88knz2q_LSw@mail.gmail.com","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-06-15T22:41:10Z","receivedAt":"2017-06-15T22:41:22Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Thu, Jun 15, 2017 at 02:59:57PM -0700, Adam Langley wrote:\n> (I was asked to comment a few points in public by Jonathan.)\n> \n> I think this group can safely assume that SHA-256, SHA-512, BLAKE2,\n> K12, etc are all secure to the extent that I don't believe that making\n> comparisons between them on that axis is meaningful. Thus I think the\n> question is primarily concerned with performance and implementation\n> availability.\n> \n> I think any of the above would be reasonable choices. I don't believe\n> that length-extension is a concern here.\n> \n> SHA-512/256 will be faster than SHA-256 on 64-bit systems in software.\n> The graph at https://blake2.net/ suggests a 50% speedup on Skylake. On\n> my Ivy Bridge system, it's about 20%.\n> \n> (SHA-512/256 does not enjoy the same availability in common libraries however.)\n> \n> Both Intel and ARM have SHA-256 instructions defined. I've not seen\n> good benchmarks of them yet, but they will make SHA-256 faster than\n> SHA-512 when available. However, it's very possible that something\n> like BLAKE2bp will still be faster. Of course, BLAKE2bp does not enjoy\n> the ubiquity of SHA-256, but nor do you have to wait years for the CPU\n> population to advance for high performance.\n\nSHA-256 acceleration exists for some existing Intel platforms already.\nHowever, they're not practically present on anything but servers at the\nmoment, and so I don't think the acceleration of SHA-256 is a\nsomething we should consider.\n\nThe SUPERCOP benchmarks tell me that generally, on 64-bit systems where\nacceleration is not available, SHA-256 is the slowest, followed by\nSHA3-256.  BLAKE2b is the fastest.\n\nIf our goal is performance, then I would argue BLAKE2b-256 is the best\nchoice.  It is secure and extremely fast.  It does have the benefit that\nwe get to tell people that by moving away from SHA-1, they will get a\nperformance boost, pretty much no matter what the system.\n\nBLAKE2bp may be faster, but it introduces additional implementation\ncomplexity.  I'm not sure crypto libraries will implement it, but then\nagain, OpenSSL only implements BLAKE2b-512 at the moment.  I don't care\nmuch either way, but we should add good tests to exercise the\nimplementation thoroughly.  We're generally going to need to ship our\nown implementation anyway.\n\nI've argued that SHA3-256 probably has the longest life and good\nunaccelerated performance, and for that reason, I've preferred it.  But\nif AGL says that they're all secure (and I generally think he knows\nwhat he's talking about), we could consider performance more.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\nhttps://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"322423","messageId":"CACBZZX5Z3kQHe_5TgOeuJSgzuvpQdaLo6RrgX_EvuZfdz856sA@mail.gmail.com","threadId":"45288","inReplyTo":"20170615224110.kvrjs3lmwxcoqfaw@genre.crustytoothpaste.net","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2017-06-15T23:36:13Z","receivedAt":"2017-06-15T23:36:40Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"On Fri, Jun 16, 2017 at 12:41 AM, brian m. carlson\n<sandals@crustytoothpaste.net> wrote:\n> On Thu, Jun 15, 2017 at 02:59:57PM -0700, Adam Langley wrote:\n>> (I was asked to comment a few points in public by Jonathan.)\n>>\n>> I think this group can safely assume that SHA-256, SHA-512, BLAKE2,\n>> K12, etc are all secure to the extent that I don't believe that making\n>> comparisons between them on that axis is meaningful. Thus I think the\n>> question is primarily concerned with performance and implementation\n>> availability.\n>>\n>> I think any of the above would be reasonable choices. I don't believe\n>> that length-extension is a concern here.\n>>\n>> SHA-512/256 will be faster than SHA-256 on 64-bit systems in software.\n>> The graph at https://blake2.net/ suggests a 50% speedup on Skylake. On\n>> my Ivy Bridge system, it's about 20%.\n>>\n>> (SHA-512/256 does not enjoy the same availability in common libraries however.)\n>>\n>> Both Intel and ARM have SHA-256 instructions defined. I've not seen\n>> good benchmarks of them yet, but they will make SHA-256 faster than\n>> SHA-512 when available. However, it's very possible that something\n>> like BLAKE2bp will still be faster. Of course, BLAKE2bp does not enjoy\n>> the ubiquity of SHA-256, but nor do you have to wait years for the CPU\n>> population to advance for high performance.\n>\n> SHA-256 acceleration exists for some existing Intel platforms already.\n> However, they're not practically present on anything but servers at the\n> moment, and so I don't think the acceleration of SHA-256 is a\n> something we should consider.\n\nWhatever next-gen hash Git ends up with is going to be in use for\ndecades, so what hardware acceleration exists in consumer products\nright now is practically irrelevant, but what acceleration is likely\nto exist for the lifetime of the hash existing *is* relevant.\n\nSo I don't follow the argument that we shouldn't weigh future HW\nacceleration highly just because you can't easily buy a laptop today\nwith these features.\n\nAside from that I think you've got this backwards, it's AMD that's\nadding SHA acceleration to their high-end Ryzen chips[1] but Intel is\nstarting at the lower end this year with Goldmont which'll be in\nlower-end consumer devices[2]. If you read the github issue I linked\nto upthread[3] you can see that the cryptopp devs already tested their\nSHA accelerated code on a consumer Celeron[4] recently.\n\nI don't think Intel has announced the SHA extensions for future Xeon\nreleases, but it seems given that they're going to have it there as\nwell. Have there every been x86 extensions that aren't eventually\nportable across the entire line, or that they've ended up removing\nfrom x86 once introduced?\n\nIn any case, I think by the time we're ready to follow-up the current\nhash refactoring efforts with actually changing the hash\nimplementation many of us are likely to have laptops with these\nextensions, making this easy to test.\n\n1. https://en.wikipedia.org/wiki/Intel_SHA_extensions\n2. https://en.wikipedia.org/wiki/Goldmont\n3. https://github.com/weidai11/cryptopp/issues/139#issuecomment-264283385\n4. https://ark.intel.com/products/95594/Intel-Celeron-Processor-J3455-2M-Cache-up-to-2_3-GHz\n"},{"id":"322428","messageId":"20170616001738.affg4qby7y7yahos@genre.crustytoothpaste.net","threadId":"45288","inReplyTo":"CACBZZX5Z3kQHe_5TgOeuJSgzuvpQdaLo6RrgX_EvuZfdz856sA@mail.gmail.com","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2017-06-16T00:17:38Z","receivedAt":"2017-06-16T00:17:50Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On Fri, Jun 16, 2017 at 01:36:13AM +0200, Ævar Arnfjörð Bjarmason wrote:\n> On Fri, Jun 16, 2017 at 12:41 AM, brian m. carlson\n> <sandals@crustytoothpaste.net> wrote:\n> > SHA-256 acceleration exists for some existing Intel platforms already.\n> > However, they're not practically present on anything but servers at the\n> > moment, and so I don't think the acceleration of SHA-256 is a\n> > something we should consider.\n> \n> Whatever next-gen hash Git ends up with is going to be in use for\n> decades, so what hardware acceleration exists in consumer products\n> right now is practically irrelevant, but what acceleration is likely\n> to exist for the lifetime of the hash existing *is* relevant.\n\nThe life of MD5 was about 23 years (introduction to first document\ncollision).  SHA-1 had about 22.  Decades, yes, but just barely.  SHA-2\nwas introduced in 2001, and by the same estimate, we're a little over\nhalfway through its life.\n\n> So I don't follow the argument that we shouldn't weigh future HW\n> acceleration highly just because you can't easily buy a laptop today\n> with these features.\n> \n> Aside from that I think you've got this backwards, it's AMD that's\n> adding SHA acceleration to their high-end Ryzen chips[1] but Intel is\n> starting at the lower end this year with Goldmont which'll be in\n> lower-end consumer devices[2]. If you read the github issue I linked\n> to upthread[3] you can see that the cryptopp devs already tested their\n> SHA accelerated code on a consumer Celeron[4] recently.\n> \n> I don't think Intel has announced the SHA extensions for future Xeon\n> releases, but it seems given that they're going to have it there as\n> well. Have there every been x86 extensions that aren't eventually\n> portable across the entire line, or that they've ended up removing\n> from x86 once introduced?\n> \n> In any case, I think by the time we're ready to follow-up the current\n> hash refactoring efforts with actually changing the hash\n> implementation many of us are likely to have laptops with these\n> extensions, making this easy to test.\n\nI think you underestimate the life of hardware and software.  I have\nservers running KVM development instances that have been running since\nat least 2012.  Those machines are not scheduled for replacement anytime\nsoon.\n\nWhatever we deploy within the next year is going to run on existing\nhardware for probably a decade, whether we want it to or not.  Most of\nthose machines don't have acceleration.\n\nFurthermore, you need a reasonably modern crypto library to get hardware\nacceleration.  OpenSSL has only recently gained support for it.  RHEL 7\ndoes not currently support it, and probably never will.  That OS is\ngoing to be around for the next 6 years.\n\nIf we're optimizing for performance, I don't want to optimize for the\nlatest, greatest machines.  Those machines are going to outperform\neverything else either way.  I'd rather optimize for something which\nperforms well on the whole everywhere.  There are a lot of developers\nwho have older machines, for cost reasons or otherwise.\n\nHere are some stats (cycles/byte for long messages):\n\n                   SHA-256    BLAKE2b\nRyzen                 1.89       3.06\nKnight's Landing     19.00       5.65\nCortex-A72            1.99       5.48\nCortex-A57           11.81       5.47\nCortex-A7            28.19      15.16\n\nIn other words, BLAKE2b performs well uniformly across a wide variety of\narchitectures even without acceleration.  I'd rather tell people that\nupgrading to a new hash algorithm is a performance win either way, not\njust if they have the latest hardware.\n-- \nbrian m. carlson / brian with sandals: Houston, Texas, US\nhttps://www.crustytoothpaste.net/~bmc | My opinion only\nOpenPGP: https://keybase.io/bk2204\n"},{"id":"322432","messageId":"20170616043040.sofnpqthmt2skdjt@sigill.intra.peff.net","threadId":"45288","inReplyTo":"20170615211022.vmedlcwmvtdiseqx@glandium.org","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-06-16T04:30:41Z","receivedAt":"2017-06-16T04:30:48Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jun 16, 2017 at 06:10:22AM +0900, Mike Hommey wrote:\n\n> > > What do the experts think or SHA512/256, which completely removes the\n> > > concerns over length extension attack? (which I'd argue is better than\n> > > sweeping them under the carpet)\n> > \n> > I don't think it's sweeping them under the carpet. Git does not use the\n> > hash as a MAC, so length extension attacks aren't a thing (and even if\n> > we later wanted to use the same algorithm as a MAC, the HMAC\n> > construction is a well-studied technique for dealing with it).\n> \n> AIUI, length extension does make brute force collision attacks (which,\n> really Shattered was) cheaper by allowing one to create the collision\n> with a small message and extend it later.\n> \n> This might not be a credible thread against git, but if we go by that\n> standard, post-shattered Sha-1 is still fine for git. As a matter of\n> fact, MD5 would also be fine: there is still, to this day, no preimage\n> attack against them.\n\nI think collision attacks are of interest to Git. But I would think\n2^128 would be enough (TBH, 2^80 probably would have been enough for\nSHA-1; it was the weaknesses that brought that down by a factor of a\nmillion that made it a problem).\n\n-Peff\n"},{"id":"322443","messageId":"87y3ss8n4h.fsf@gmail.com","threadId":"45288","inReplyTo":"20170616001738.affg4qby7y7yahos@genre.crustytoothpaste.net","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2017-06-16T06:25:50Z","receivedAt":"2017-06-16T06:26:00Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Jun 16 2017, brian m. carlson jotted:\n\n> On Fri, Jun 16, 2017 at 01:36:13AM +0200, Ævar Arnfjörð Bjarmason wrote:\n>> On Fri, Jun 16, 2017 at 12:41 AM, brian m. carlson\n>> <sandals@crustytoothpaste.net> wrote:\n>> > SHA-256 acceleration exists for some existing Intel platforms already.\n>> > However, they're not practically present on anything but servers at the\n>> > moment, and so I don't think the acceleration of SHA-256 is a\n>> > something we should consider.\n>>\n>> Whatever next-gen hash Git ends up with is going to be in use for\n>> decades, so what hardware acceleration exists in consumer products\n>> right now is practically irrelevant, but what acceleration is likely\n>> to exist for the lifetime of the hash existing *is* relevant.\n>\n> The life of MD5 was about 23 years (introduction to first document\n> collision).  SHA-1 had about 22.  Decades, yes, but just barely.  SHA-2\n> was introduced in 2001, and by the same estimate, we're a little over\n> halfway through its life.\n\nI'm talking about the lifetime of SHA-1 or $newhash's use in Git. As our\ncontinued use of SHA-1 demonstrates the window of practical hash\nfunction use extends well beyond the window from introduction to\npublished breakage.\n\nIt's also telling that SHA-1, which any cryptographer would have waived\nyou off from since around 2011, is just getting widely deployed HW\nacceleration now in 2017. The practical use of hash functions far\nexceeds their recommended use in new projects.\n\n>> So I don't follow the argument that we shouldn't weigh future HW\n>> acceleration highly just because you can't easily buy a laptop today\n>> with these features.\n>>\n>> Aside from that I think you've got this backwards, it's AMD that's\n>> adding SHA acceleration to their high-end Ryzen chips[1] but Intel is\n>> starting at the lower end this year with Goldmont which'll be in\n>> lower-end consumer devices[2]. If you read the github issue I linked\n>> to upthread[3] you can see that the cryptopp devs already tested their\n>> SHA accelerated code on a consumer Celeron[4] recently.\n>>\n>> I don't think Intel has announced the SHA extensions for future Xeon\n>> releases, but it seems given that they're going to have it there as\n>> well. Have there every been x86 extensions that aren't eventually\n>> portable across the entire line, or that they've ended up removing\n>> from x86 once introduced?\n>>\n>> In any case, I think by the time we're ready to follow-up the current\n>> hash refactoring efforts with actually changing the hash\n>> implementation many of us are likely to have laptops with these\n>> extensions, making this easy to test.\n>\n> I think you underestimate the life of hardware and software.  I have\n> servers running KVM development instances that have been running since\n> at least 2012.  Those machines are not scheduled for replacement anytime\n> soon.\n>\n> Whatever we deploy within the next year is going to run on existing\n> hardware for probably a decade, whether we want it to or not.  Most of\n> those machines don't have acceleration.\n\nTo clarify, I'm not dismissing the need to consider existing hardware\nwithout these acceleration functions or future processors without\nthem. I don't think that makes any sense, we need to keep those in mind.\n\nI was replying to a bit in your comment where you (it seems to me) were\nmaking the claim that we shouldn't consider the HW acceleration of\ncertain hash functions either.\n\nClearly both need to be considered.\n\n> Furthermore, you need a reasonably modern crypto library to get hardware\n> acceleration.  OpenSSL has only recently gained support for it.  RHEL 7\n> does not currently support it, and probably never will.  That OS is\n> going to be around for the next 6 years.\n>\n> If we're optimizing for performance, I don't want to optimize for the\n> latest, greatest machines.  Those machines are going to outperform\n> everything else either way.  I'd rather optimize for something which\n> performs well on the whole everywhere.  There are a lot of developers\n> who have older machines, for cost reasons or otherwise.\n\nWe have real data showing that the intersection between people who care\nabout the hash slowing down and those who can't afford the latest\nhardware is pretty much nil.\n\nI.e. in 2.13.0 SHA-1 got slower, and pretty much nobody noticed or cared\nexcept Johannes Schindelin, myself & Christian Couder. This is because\nin practice hashing only becomes a bottleneck on huge monorepos that\nneed to e.g. re-hash the contents of a huge index.\n\n> Here are some stats (cycles/byte for long messages):\n>\n>                    SHA-256    BLAKE2b\n> Ryzen                 1.89       3.06\n> Knight's Landing     19.00       5.65\n> Cortex-A72            1.99       5.48\n> Cortex-A57           11.81       5.47\n> Cortex-A7            28.19      15.16\n>\n> In other words, BLAKE2b performs well uniformly across a wide variety of\n> architectures even without acceleration.  I'd rather tell people that\n> upgrading to a new hash algorithm is a performance win either way, not\n> just if they have the latest hardware.\n\nYup, all of those need to be considered, although given my comment above\nabout big repos a 40% improvement on Ryzen (a processor likely to be\nused for big repos) stands out, where are those numbers from, and is\nthat with or without HW accel for SHA-256 on Ryzen?\n"},{"id":"322458","messageId":"alpine.DEB.2.21.1.1706161438470.4200@virtualbox","threadId":"45288","inReplyTo":"87y3ss8n4h.fsf@gmail.com","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-06-16T13:24:19Z","receivedAt":"2017-06-16T13:25:20Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 16 Jun 2017, Ævar Arnfjörð Bjarmason wrote:\n\n> On Fri, Jun 16 2017, brian m. carlson jotted:\n> \n> > On Fri, Jun 16, 2017 at 01:36:13AM +0200, Ævar Arnfjörð Bjarmason wrote:\n> >\n> >> So I don't follow the argument that we shouldn't weigh future HW\n> >> acceleration highly just because you can't easily buy a laptop today\n> >> with these features.\n> >>\n> >> Aside from that I think you've got this backwards, it's AMD that's\n> >> adding SHA acceleration to their high-end Ryzen chips[1] but Intel is\n> >> starting at the lower end this year with Goldmont which'll be in\n> >> lower-end consumer devices[2]. If you read the github issue I linked\n> >> to upthread[3] you can see that the cryptopp devs already tested\n> >> their SHA accelerated code on a consumer Celeron[4] recently.\n> >>\n> >> I don't think Intel has announced the SHA extensions for future Xeon\n> >> releases, but it seems given that they're going to have it there as\n> >> well. Have there every been x86 extensions that aren't eventually\n> >> portable across the entire line, or that they've ended up removing\n> >> from x86 once introduced?\n> >>\n> >> In any case, I think by the time we're ready to follow-up the current\n> >> hash refactoring efforts with actually changing the hash\n> >> implementation many of us are likely to have laptops with these\n> >> extensions, making this easy to test.\n> >\n> > I think you underestimate the life of hardware and software.  I have\n> > servers running KVM development instances that have been running since\n> > at least 2012.  Those machines are not scheduled for replacement\n> > anytime soon.\n> >\n> > Whatever we deploy within the next year is going to run on existing\n> > hardware for probably a decade, whether we want it to or not.  Most of\n> > those machines don't have acceleration.\n> \n> To clarify, I'm not dismissing the need to consider existing hardware\n> without these acceleration functions or future processors without them.\n> I don't think that makes any sense, we need to keep those in mind.\n> \n> I was replying to a bit in your comment where you (it seems to me) were\n> making the claim that we shouldn't consider the HW acceleration of\n> certain hash functions either.\n\nYes, I also had the impression that it stressed the status quo quite a bit\ntoo much.\n\nWe know for a fact that SHA-256 acceleration is coming to consumer CPUs.\nWe know of no plans for any of the other mentioned hash functions to\nhardware-accelerate them in consumer CPUs.\n\nAnd remember: for those who are affected most (humongous monorepos, source\ncode hosters), upgrading hardware is less of an issue than having a secure\nhash function for the rest of us.\n\nAnd while I am really thankful that Adam chimed in, I think he would agree\nthat BLAKE2 is a purposefully weakened version of BLAKE, for the benefit\nof speed (with the caveat that one of my experts disagrees that BLAKE2b\nwould be faster than hardware-accelerated SHA-256). And while BLAKE has\nseen roughly equivalent cryptanalysis as Keccak (which became SHA-3),\nBLAKE2 has not.\n\nThat makes me *very* uneasy about choosing BLAKE2.\n\n> > Furthermore, you need a reasonably modern crypto library to get hardware\n> > acceleration.  OpenSSL has only recently gained support for it.  RHEL 7\n> > does not currently support it, and probably never will.  That OS is\n> > going to be around for the next 6 years.\n> >\n> > If we're optimizing for performance, I don't want to optimize for the\n> > latest, greatest machines.  Those machines are going to outperform\n> > everything else either way.  I'd rather optimize for something which\n> > performs well on the whole everywhere.  There are a lot of developers\n> > who have older machines, for cost reasons or otherwise.\n> \n> We have real data showing that the intersection between people who care\n> about the hash slowing down and those who can't afford the latest\n> hardware is pretty much nil.\n> \n> I.e. in 2.13.0 SHA-1 got slower, and pretty much nobody noticed or cared\n> except Johannes Schindelin, myself & Christian Couder. This is because\n> in practice hashing only becomes a bottleneck on huge monorepos that\n> need to e.g. re-hash the contents of a huge index.\n\nIndeed. I am still concerned about that. As you mention, though, it really\nonly affects users of ginormous monorepos, and of course source code\nhosters.\n\nThe jury's still out on how much it impacts my colleagues, by the way.\n\nI have no doubt that Visual Studio Team Services, GitHub and Atlassian\nwill eventually end up with FPGAs for hash computation. So that's that.\n\nSide note: BLAKE is actually *not* friendly to hardware acceleration, I\nhave been told by one cryptography expert. In contrast, the Keccak team\nclaims SHA3-256 to be the easiest to hardware-accelerate, making it \"a\ngreen cryptographic primitive\":\nhttp://keccak.noekeon.org/is_sha3_slow.html\n\n> > Here are some stats (cycles/byte for long messages):\n> >\n> >                    SHA-256    BLAKE2b\n> > Ryzen                 1.89       3.06\n> > Knight's Landing     19.00       5.65\n> > Cortex-A72            1.99       5.48\n> > Cortex-A57           11.81       5.47\n> > Cortex-A7            28.19      15.16\n> >\n> > In other words, BLAKE2b performs well uniformly across a wide variety of\n> > architectures even without acceleration.  I'd rather tell people that\n> > upgrading to a new hash algorithm is a performance win either way, not\n> > just if they have the latest hardware.\n> \n> Yup, all of those need to be considered, although given my comment above\n> about big repos a 40% improvement on Ryzen (a processor likely to be\n> used for big repos) stands out, where are those numbers from, and is\n> that with or without HW accel for SHA-256 on Ryzen?\n\nWhen it comes to BLAKE2, I would actually strongly suggest to consider the\namount of attempts to break it. Or rather, how much less attention it got\nthan, say, SHA-256.\n\nIn any case, I have been encouraged to stress the importance of\n\"crypto-agility\", i.e. the ability to switch to another algorithm when the\ncurrent one gets broken \"enough\".\n\nAnd I am delighted that that is exactly the direction we are going. In\nother words, even if I still think (backed up by the experts on whose\nknowledge I lean heavily to form my opinions) that SHA-256 would be the\nbest choice for now, it should be relatively easy to offer BLAKE2b support\nfor (and by [*1*]) those who want it.\n\nCiao,\nDscho\n\nFootnote *1*: I say that the support for BLAKE2b should come from those\nparties who desire it also because it is not as ubiquituous as SHA-256.\nHence, it would add the burden of having a performant and reasonably\nbug-free implementation in Git's source tree. IIUC OpenSSL added BLAKE2b\nsupport only in OpenSSL 1.1.0, the 1.0.2 line (which is still in use in\nmany places, e.g. Git for Windows' SDK) does not, meaning: Git's\nimplementation would be the one *everybody* relies on, with *no*\nfall-back."},{"id":"322470","messageId":"CAL9PXLxMHG1nP5_GQaK_WSJTNKs=_qbaL6V5v2GzVG=9VU2+gA@mail.gmail.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1706161438470.4200@virtualbox","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Adam Langley","fromEmail":"agl@google.com","sentAt":"2017-06-16T17:38:41Z","receivedAt":"2017-06-16T17:39:23Z","isPatch":false,"sender":{"key":"agl@google.com","avatar":null},"body":"On Fri, Jun 16, 2017 at 6:24 AM, Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n>\n> And while I am really thankful that Adam chimed in, I think he would agree\n> that BLAKE2 is a purposefully weakened version of BLAKE, for the benefit\n> of speed\n\nThat is correct.\n\nAlthough worth keeping in mind that the analysis results from the\nSHA-3 process informed this rebalancing. Indeed, NIST proposed[1] to\ndo the same with Keccak before stamping it as SHA-3 (although\nultimately did not in the context of public feeling in late 2013). The\nKeccak team have essentially done the same with K12. Thus there is\nevidence of a fairly widespread belief that the SHA-3 parameters were\nexcessively cautious.\n\n[1] https://docs.google.com/file/d/0BzRYQSHuuMYOQXdHWkRiZXlURVE/edit, slide 48\n\n> (with the caveat that one of my experts disagrees that BLAKE2b\n> would be faster than hardware-accelerated SHA-256).\n\nThe numbers given above for SHA-256 on Ryzen and Cortex-A72 must be\nwith hardware acceleration and I thank Brian Carlson for digging them\nup as I hadn't seen them before.\n\nI suggested above that BLAKE2bp (note the p at the end) might be\nfaster than hardware SHA-256 and that appears to be plausible based on\nbenchmarks[2] of that function. (With the caveat those numbers are for\nHaswell and Skylake and so cannot be directly compared with Ryzen.)\n\nK12 reports similar speeds on Skylake[3] and thus is also plausibly\nfaster than hardware SHA-256.\n\n[2] https://github.com/sneves/blake2-avx2\n[3] http://keccak.noekeon.org/KangarooTwelve.pdf\n\nHowever, as I'm not a git developer, I've no opinion on whether the\ncost of carrying implementations of these functions is worth the speed\nvs using SHA-256, which can be assumed to be supported everywhere\nalready.\n\n\nCheers\n\nAGL\n"},{"id":"322484","messageId":"20170616204208.ak5twydrloxefm42@sigill.intra.peff.net","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1706161438470.4200@virtualbox","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-06-16T20:42:08Z","receivedAt":"2017-06-16T20:42:15Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Jun 16, 2017 at 03:24:19PM +0200, Johannes Schindelin wrote:\n\n> I have no doubt that Visual Studio Team Services, GitHub and Atlassian\n> will eventually end up with FPGAs for hash computation. So that's that.\n\nI actually doubt this from the GitHub side. Hash performance is not even\non our radar as a bottleneck. In most cases the problem is touching\nuncompressed data _at all_, not computing the hash over it (so things\nlike reusing on-disk deltas are really important).\n\n-Peff\n"},{"id":"322485","messageId":"xmqq37azy7ru.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"CAL9PXLxMHG1nP5_GQaK_WSJTNKs=_qbaL6V5v2GzVG=9VU2+gA@mail.gmail.com","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-06-16T20:52:53Z","receivedAt":"2017-06-16T20:53:00Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Adam Langley <agl@google.com> writes:\n\n> However, as I'm not a git developer, I've no opinion on whether the\n> cost of carrying implementations of these functions is worth the speed\n> vs using SHA-256, which can be assumed to be supported everywhere\n> already.\n\nThanks.\n\nMy impression from this thread is that even though fast may be\nbetter than slow, ubiquity trumps it for our use case, as long as\nthe thing is not absurdly and unusably slow, of course.  Which makes\nme lean towards something older/more established like SHA-256, and\nit would be a very nice bonus if it gets hardware acceleration more\nwidely than others ;-)\n\n"},{"id":"322488","messageId":"xmqqr2yjwsb6.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"xmqq37azy7ru.fsf@gitster.mtv.corp.google.com","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-06-16T21:12:13Z","receivedAt":"2017-06-16T21:12:20Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Adam Langley <agl@google.com> writes:\n>\n>> However, as I'm not a git developer, I've no opinion on whether the\n>> cost of carrying implementations of these functions is worth the speed\n>> vs using SHA-256, which can be assumed to be supported everywhere\n>> already.\n>\n> Thanks.\n>\n> My impression from this thread is that even though fast may be\n> better than slow, ubiquity trumps it for our use case, as long as\n> the thing is not absurdly and unusably slow, of course.  Which makes\n> me lean towards something older/more established like SHA-256, and\n> it would be a very nice bonus if it gets hardware acceleration more\n> widely than others ;-)\n\nAh, I recall one thing that was mentioned but not discussed much in\nthe thread: possible use of tree-hashing to exploit multiple cores\nhashing a large-ish payload.  As long as it is OK to pick a sound\ntree hash coding on top of any (secure) underlying hash function,\nI do not think the use of tree-hashing should not affect which exact\nunderlying hash function is to be used, and I also am not convinced\nif we really want tree hashing (some codepaths that deal with a large\npayload wants to stream the data in single pass from head to tail)\nin the context of Git, but I am not a crypto person, so ...\n\n\n"},{"id":"322489","messageId":"20170616212414.GC133952@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"xmqqr2yjwsb6.fsf@gitster.mtv.corp.google.com","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-06-16T21:24:14Z","receivedAt":"2017-06-16T21:24:23Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n>> Adam Langley <agl@google.com> writes:\n\n>>> However, as I'm not a git developer, I've no opinion on whether the\n>>> cost of carrying implementations of these functions is worth the speed\n>>> vs using SHA-256, which can be assumed to be supported everywhere\n>>> already.\n>>\n>> Thanks.\n>>\n>> My impression from this thread is that even though fast may be\n>> better than slow, ubiquity trumps it for our use case, as long as\n>> the thing is not absurdly and unusably slow, of course.  Which makes\n>> me lean towards something older/more established like SHA-256, and\n>> it would be a very nice bonus if it gets hardware acceleration more\n>> widely than others ;-)\n>\n> Ah, I recall one thing that was mentioned but not discussed much in\n> the thread: possible use of tree-hashing to exploit multiple cores\n> hashing a large-ish payload.  As long as it is OK to pick a sound\n> tree hash coding on top of any (secure) underlying hash function,\n> I do not think the use of tree-hashing should not affect which exact\n> underlying hash function is to be used, and I also am not convinced\n> if we really want tree hashing (some codepaths that deal with a large\n> payload wants to stream the data in single pass from head to tail)\n> in the context of Git, but I am not a crypto person, so ...\n\nTree hashing also affects single-core performance because of the\navailability of SIMD instructions.\n\nThat is how software implementations of e.g. blake2bp-256 and\nSHA-256x16[1] are able to have competitive performance with (slightly\nbetter performance than, at least in some cases) hardware\nimplementations of SHA-256.\n\nIt is also satisfying that we have options like these that are faster\nthan SHA-1.\n\nAll that said, SHA-256 seems like a fine choice, despite its worse\nperformance.  The wide availability of reasonable-quality\nimplementations (e.g. in Java you can use\n'MessageDigest.getInstance(\"SHA-256\")') makes it a very tempting one.\n\nPart of the reason I suggested previously that it would be helpful to\ntry to benchmark Git with various hash functions (which didn't go over\nwell, for some reason) is that it makes these comparisons more\nconcrete.  Without measuring, it is hard to get a sense of the\ndistribution of input sizes and how much practical effect the\ndifferences we are talking about have.\n\nThanks,\nJonathan\n\n[1] https://eprint.iacr.org/2012/476.pdf\n"},{"id":"322492","messageId":"87tw3f8vez.fsf@gmail.com","threadId":"45288","inReplyTo":"20170616212414.GC133952@aiede.mtv.corp.google.com","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2017-06-16T21:39:00Z","receivedAt":"2017-06-16T21:39:16Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, Jun 16 2017, Jonathan Nieder jotted:\n> Part of the reason I suggested previously that it would be helpful to\n> try to benchmark Git with various hash functions (which didn't go over\n> well, for some reason) is that it makes these comparisons more\n> concrete.  Without measuring, it is hard to get a sense of the\n> distribution of input sizes and how much practical effect the\n> differences we are talking about have.\n\nIt would be great to have such benchmarks (I probably missed the \"didn't\ngo over well\" part), but FWIW you can get pretty close to this right now\nin git by running various t/perf benchmarks with\nBLKSHA1/OPENSSL/SHA1DC.\n\nBetween the three of those (particularly SHA1DC being slower than\nOpenSSL) you get a similar performance difference as some SHA-1\nv.s. SHA-256 benchmarks I've seen, so to the extent that we have\nexisting performance tests it's revealing to see what's slower & faster.\n\nIt makes a particularly big difference for e.g. p3400-rebase.sh.\n"},{"id":"322553","messageId":"alpine.DEB.2.21.1.1706191125510.57822@virtualbox","threadId":"45288","inReplyTo":"20170616204208.ak5twydrloxefm42@sigill.intra.peff.net","subject":"Re: Which hash function to use, was Re: RFC: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-06-19T09:26:44Z","receivedAt":"2017-06-19T09:27:53Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Peff,\n\nOn Fri, 16 Jun 2017, Jeff King wrote:\n\n> On Fri, Jun 16, 2017 at 03:24:19PM +0200, Johannes Schindelin wrote:\n> \n> > I have no doubt that Visual Studio Team Services, GitHub and Atlassian\n> > will eventually end up with FPGAs for hash computation. So that's\n> > that.\n> \n> I actually doubt this from the GitHub side. Hash performance is not even\n> on our radar as a bottleneck. In most cases the problem is touching\n> uncompressed data _at all_, not computing the hash over it (so things\n> like reusing on-disk deltas are really important).\n\nThanks for pointing that out! As a mainly client-side person, I rarely get\ninsights into the server side...\n\nCiao,\nDscho\n"},{"id":"327635","messageId":"xmqqa828733s.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170307001709.GC26789@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-09-06T06:28:23Z","receivedAt":"2017-09-06T06:28:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> Linus Torvalds wrote:\n>> On Fri, Mar 3, 2017 at 5:12 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n>\n>>> This document is still in flux but I thought it best to send it out\n>>> early to start getting feedback.\n>>\n>> This actually looks very reasonable if you can implement it cleanly\n>> enough.\n>\n> Thanks for the kind words on what had quite a few flaws still.  Here's\n> a new draft.  I think the next version will be a patch against\n> Documentation/technical/.\n\nCan we reboot the discussion and advance this to v4 state?\n\n> As before, comments welcome, both here and inline at\n>\n>   https://goo.gl/gh2Mzc\n\nI think what you have over there looks pretty-much ready as the\nfinal outline.\n\nOne thing I still do not know how I feel about after re-reading the\nthread, and I didn't find the above doc, is Linus's suggestion to\nuse the objects themselves as NewHash-to-SHA-1 mapper [*1*].  \n\nIt does not help the reverse mapping that is needed while pushing\nthings out (the SHA-1 receiver tells us what they have in terms of\nSHA-1 names; we need to figure out where we stop sending based on\nthat).  While it does help maintaining itself (while constructing\nSHA3-content, we'd be required to find out its SHA1 name but the\nSHA3 objects that we refer to all know their SHA-1 names), if it is\nnot useful otherwise, then that does not count as a plus.  Also\nhaving to bake corresponding SHA-1 name in the object would mean\nmistakes can easily propagate and cannot be corrected without\nrewriting the history, which would be a huge downside.  So perhaps\nwe are better off without it, I guess.\n\n\n[Reference]\n\n*1* <CA+55aFxj7Vtwac64RfAz_u=U4tob4Xg+2pDBDFNpJdmgaTCmxA@mail.gmail.com>\n\n\n"},{"id":"327741","messageId":"xmqq1snh29re.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"xmqqa828733s.fsf@gitster.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-09-08T02:40:21Z","receivedAt":"2017-09-08T02:40:32Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> One thing I still do not know how I feel about after re-reading the\n> thread, and I didn't find the above doc, is Linus's suggestion to\n> use the objects themselves as NewHash-to-SHA-1 mapper [*1*].  \n> ...\n> [Reference]\n>\n> *1* <CA+55aFxj7Vtwac64RfAz_u=U4tob4Xg+2pDBDFNpJdmgaTCmxA@mail.gmail.com>\n\nI think this falls into the same category as the often-talked-about\naddition of the \"generation number\" field.  It is very tempting to\nadd these \"mechanically derivable but expensive to compute\" pieces\nof information to the sha3-content while converting from\nsha1-content and creating anew.  \n\nBecause the \"sha1-name\" or the \"generation number\" can mechanically\nbe computed, as long as everybody agrees to _always_ place them in\nthe sha3-content, the same sha1-content will be converted into\nexactly the same sha3-content without ambiguity, and converting them\nback to sha1-content while pushing to an older repository will\ncorrectly produce the original sha1-content, as it would just be the\nmatter of simply stripping these extra pieces of information.\n\nThe reason why I still feel a bit uneasy about adding these things\n(aside from the fact that sha1-name thing will be a baggage we would\nneed to carry forever even after we completely wean ourselves off of\nthe old hash) is because I am not sure what we should do when we\nencounter sha3-content in the wild that has these things _wrong_.\nAn object that exists today in the SHA-1 world is fetched into the\nnew repository and converted to SHA-3 contents, and Linus's extra\n\"original SHA-1 name\" field is added to the object's header while\nrecording the SHA-3 content.  But for whatever reason, the original\nSHA-1 name is recorded incorrectly in the resulting SHA-3 object.\n\nThe same thing could happen if we decide to bake \"generation number\"\nin the SHA-3 commit objects.  One possible definition would be that\na root commit will have gen #0; a commit with 1 or more parents will\nget max(parents' gen numbers) + 1 as its gen number.  But somebody\nmay botch the counting and records sum(parents' gen numbers) as its\ngen number.\n\nIn these cases, not just the SHA3-content but also the resulting\nSHA-3 object name would be different from the name of the object\nthat would have recorded the same contents correctly.  So converting\nback to SHA-1 world from these botched SHA-3 contents may produce\nthe original contents, but we may end up with multiple \"plausibly\nlooking\" set of SHA-3 objects that (clain to) correspond to a single\nSHA-1 object, only one of which is a valid one.\n\nOur \"git fsck\" already treats certain brokenness (like a tree whose\nentry has mode that is 0-padded to the left) as broken but still\ntolerate them.  I am not sure if it is sufficient to diagnose and\ndeclare broken and invalid when we see sha3-content that records\nthese \"mechanically derivable but expensive to compute\" pieces of\ninformation incorrectly.\n\nI am leaning towards saying \"yes, catching in fsck is enough\" and\nsuggesting to add generation number to sha3-content of the commit\nobjects, and to add even the \"original sha1 name\" thing if we find\ngood use of it.  But I cannot shake this nagging feeling off that I\nam missing some huge problems that adding these fields and opening\nourselves to more classes of broken objects.\n\nThoughts?\n\n\n"},{"id":"327743","messageId":"20170908033403.q7e6dj7benasrjes@sigill.intra.peff.net","threadId":"45288","inReplyTo":"xmqq1snh29re.fsf@gitster.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-09-08T03:34:03Z","receivedAt":"2017-09-08T03:34:14Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Sep 08, 2017 at 11:40:21AM +0900, Junio C Hamano wrote:\n\n> Our \"git fsck\" already treats certain brokenness (like a tree whose\n> entry has mode that is 0-padded to the left) as broken but still\n> tolerate them.  I am not sure if it is sufficient to diagnose and\n> declare broken and invalid when we see sha3-content that records\n> these \"mechanically derivable but expensive to compute\" pieces of\n> information incorrectly.\n> \n> I am leaning towards saying \"yes, catching in fsck is enough\" and\n> suggesting to add generation number to sha3-content of the commit\n> objects, and to add even the \"original sha1 name\" thing if we find\n> good use of it.  But I cannot shake this nagging feeling off that I\n> am missing some huge problems that adding these fields and opening\n> ourselves to more classes of broken objects.\n\nI share your nagging feeling.\n\nI have two thoughts on the \"fsck can catch it\" line of reasoning.\n\n  1. It's harder to fsck generation numbers than other syntactic\n     elements of an object, because it inherently depends on the links.\n     So I can't fsck a commit object in isolation. I have to open its\n     parents and check _their_ generation numbers.\n\n     In some sense that isn't a big deal. A real fsck wants to know that\n     we _have_ the parents in the first place. But traditionally we've\n     separated \"is this syntactically valid\" from \"do we have full\n     connectivity\". And features like shallow clones rely on us fudging\n     the latter but not the former. A shallow history could never\n     properly fsck the generation numbers.\n\n     A multiple-hash field doesn't have this problem. It's purely a\n     function of the bytes in the object.\n\n  2. I wouldn't classify the current fsck checks as a wild success in\n     containing breakages. If a buggy implementation produces invalid\n     objects, the same buggy implementation generally lets people (and\n     their colleagues) unwittingly build on top of those objects. It's\n     only later (sometimes much later) that they interact with a\n     non-buggy implementation whose fsck complains.\n\n     And what happens then? If they're lucky, the invalid objects\n     haven't spread far, and the worst thing is that they have to learn\n     to use filter-branch (which itself is punishment enough). But\n     sometimes a significant bit of history has been built on top, and\n     it's awkward or impossible to rewrite it.\n\n     That puts the burden on whoever is running the non-buggy\n     implementation that wants to reject the objects. Do they accept\n     these broken objects? If so, what do they do to mitigate the wrong\n     answers that Git will return?\n\nI'm much more in favor of keeping that data outside the object-hash\ncomputation, and caching the pre-computed results as necessary. Those\ncache can disagree with the objects, of course, but the cost to dropping\nand re-building them is much lower than a history rewrite.\n\nI'm speaking primarily to the generation-number thing, where I really\ndon't think there's any benefit to embedding it in the object beyond the\nobvious \"well, it has to go _somewhere_, and this saves us implementing\na local cache layer\".  I haven't thought hard enough on the\nmultiple-hash thing to know if there's some other benefit to having it\ninside the objects.\n\n-Peff\n"},{"id":"327860","messageId":"20170911185913.GA5869@google.com","threadId":"45288","inReplyTo":"xmqq1snh29re.fsf@gitster.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Brandon Williams","fromEmail":"bmwill@google.com","sentAt":"2017-09-11T18:59:13Z","receivedAt":"2017-09-11T18:59:21Z","isPatch":false,"sender":{"key":"bwilliams.eng@gmail.com","avatar":null},"body":"On 09/08, Junio C Hamano wrote:\n> Junio C Hamano <gitster@pobox.com> writes:\n> \n> > One thing I still do not know how I feel about after re-reading the\n> > thread, and I didn't find the above doc, is Linus's suggestion to\n> > use the objects themselves as NewHash-to-SHA-1 mapper [*1*].  \n> > ...\n> > [Reference]\n> >\n> > *1* <CA+55aFxj7Vtwac64RfAz_u=U4tob4Xg+2pDBDFNpJdmgaTCmxA@mail.gmail.com>\n> \n> I think this falls into the same category as the often-talked-about\n> addition of the \"generation number\" field.  It is very tempting to\n> add these \"mechanically derivable but expensive to compute\" pieces\n> of information to the sha3-content while converting from\n> sha1-content and creating anew.  \n\nWe didn't discuss that in the doc since this particular transition plan\nwe made uses an external NewHash-to-SHA1 map instead of an internal one\nbecause we believe that at some point we would be able to drop\ncompatibility with SHA1.  Now I suspect that wont happen for a long time\nbut I think it would be preferable over carrying the SHA1 luggage\nindefinitely.  At some point, then, we would be able to stop hashing\nobjects twice (once with SHA1 and once with NewHash) instead of always\nrequiring that we hash them with each hash function which was used\nhistorically.\n\n> \n> Because the \"sha1-name\" or the \"generation number\" can mechanically\n> be computed, as long as everybody agrees to _always_ place them in\n> the sha3-content, the same sha1-content will be converted into\n> exactly the same sha3-content without ambiguity, and converting them\n> back to sha1-content while pushing to an older repository will\n> correctly produce the original sha1-content, as it would just be the\n> matter of simply stripping these extra pieces of information.\n> \n> The reason why I still feel a bit uneasy about adding these things\n> (aside from the fact that sha1-name thing will be a baggage we would\n> need to carry forever even after we completely wean ourselves off of\n> the old hash) is because I am not sure what we should do when we\n> encounter sha3-content in the wild that has these things _wrong_.\n> An object that exists today in the SHA-1 world is fetched into the\n> new repository and converted to SHA-3 contents, and Linus's extra\n> \"original SHA-1 name\" field is added to the object's header while\n> recording the SHA-3 content.  But for whatever reason, the original\n> SHA-1 name is recorded incorrectly in the resulting SHA-3 object.\n\nThis wasn't one of the issues that I thought of but it just makes the\nargument against adding sha1's to the sha3 content stronger.\n\n> \n> The same thing could happen if we decide to bake \"generation number\"\n> in the SHA-3 commit objects.  One possible definition would be that\n> a root commit will have gen #0; a commit with 1 or more parents will\n> get max(parents' gen numbers) + 1 as its gen number.  But somebody\n> may botch the counting and records sum(parents' gen numbers) as its\n> gen number.\n> \n> In these cases, not just the SHA3-content but also the resulting\n> SHA-3 object name would be different from the name of the object\n> that would have recorded the same contents correctly.  So converting\n> back to SHA-1 world from these botched SHA-3 contents may produce\n> the original contents, but we may end up with multiple \"plausibly\n> looking\" set of SHA-3 objects that (clain to) correspond to a single\n> SHA-1 object, only one of which is a valid one.\n> \n> Our \"git fsck\" already treats certain brokenness (like a tree whose\n> entry has mode that is 0-padded to the left) as broken but still\n> tolerate them.  I am not sure if it is sufficient to diagnose and\n> declare broken and invalid when we see sha3-content that records\n> these \"mechanically derivable but expensive to compute\" pieces of\n> information incorrectly.\n> \n> I am leaning towards saying \"yes, catching in fsck is enough\" and\n> suggesting to add generation number to sha3-content of the commit\n> objects, and to add even the \"original sha1 name\" thing if we find\n> good use of it.  But I cannot shake this nagging feeling off that I\n> am missing some huge problems that adding these fields and opening\n> ourselves to more classes of broken objects.\n> \n> Thoughts?\n> \n> \n\n-- \nBrandon Williams\n"},{"id":"327925","messageId":"alpine.DEB.2.21.1.1709131340030.4132@virtualbox","threadId":"45288","inReplyTo":"20170911185913.GA5869@google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-13T12:05:23Z","receivedAt":"2017-09-13T12:06:05Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Brandon,\n\nOn Mon, 11 Sep 2017, Brandon Williams wrote:\n\n> On 09/08, Junio C Hamano wrote:\n> > Junio C Hamano <gitster@pobox.com> writes:\n> > \n> > > One thing I still do not know how I feel about after re-reading the\n> > > thread, and I didn't find the above doc, is Linus's suggestion to\n> > > use the objects themselves as NewHash-to-SHA-1 mapper [*1*].  \n> > > ...\n> > > [Reference]\n> > >\n> > > *1* <CA+55aFxj7Vtwac64RfAz_u=U4tob4Xg+2pDBDFNpJdmgaTCmxA@mail.gmail.com>\n> > \n> > I think this falls into the same category as the often-talked-about\n> > addition of the \"generation number\" field.  It is very tempting to add\n> > these \"mechanically derivable but expensive to compute\" pieces of\n> > information to the sha3-content while converting from sha1-content and\n> > creating anew.  \n> \n> We didn't discuss that in the doc since this particular transition plan\n> we made uses an external NewHash-to-SHA1 map instead of an internal one\n> because we believe that at some point we would be able to drop\n> compatibility with SHA1.\n\nIs there even a question about that? I mean, why would *any* project that\nswitches entirely to SHA-256 want to carry the SHA-1 baggage around?\n\nSo even if the code to generate a bidirectional old <-> new hash mapping\nmight be with us forever, it *definitely* should be optional (\"optional\"\nat least as in \"config setting\"), allowing developers who only work with\nnew-hash repositories to save the time and electrons.\n\n> Now I suspect that wont happen for a long time but I think it would be\n> preferable over carrying the SHA1 luggage indefinitely.\n\nIt should be possible to push back the SHA-1 ginny into a small gin bottle\ninside Git's source code, so to say, i.e. encapsulate it to the point\nwhere it is a compile-time option, in addition to a runtime option.\n\nOf course, that's only unless the SHA-1 calculation is made mandatory as\nsuggested above. I really shudder at the idea of requiring SHA-1 to be\nrequired forever. We ignored advice in 2005 against making ourselves too\ndependent on SHA-1, and I would hope that we would learn from this.\n\n> At some point, then, we would be able to stop hashing objects twice\n> (once with SHA1 and once with NewHash) instead of always requiring that\n> we hash them with each hash function which was used historically.\n\nYes, please.\n\n> > Because the \"sha1-name\" or the \"generation number\" can mechanically\n> > be computed,\n\n... as long as a shallow clone you do not have, of course...\n\n> > as long as everybody agrees to _always_ place them in the\n> > sha3-content, the same sha1-content will be converted into exactly the\n> > same sha3-content without ambiguity, and converting them back to\n> > sha1-content while pushing to an older repository will correctly\n> > produce the original sha1-content, as it would just be the matter of\n> > simply stripping these extra pieces of information.\n\n... or Git would simply handle the absence of the generation number header\ngracefully, so that sha1-content == sha3-content...\n\n> > The same thing could happen if we decide to bake \"generation number\"\n> > in the SHA-3 commit objects.  One possible definition would be that a\n> > root commit will have gen #0; a commit with 1 or more parents will get\n> > max(parents' gen numbers) + 1 as its gen number.  But somebody may\n> > botch the counting and records sum(parents' gen numbers) as its gen\n> > number.\n> > \n> > In these cases, not just the SHA3-content but also the resulting SHA-3\n> > object name would be different from the name of the object that would\n> > have recorded the same contents correctly.  So converting back to\n> > SHA-1 world from these botched SHA-3 contents may produce the original\n> > contents, but we may end up with multiple \"plausibly looking\" set of\n> > SHA-3 objects that (clain to) correspond to a single SHA-1 object,\n> > only one of which is a valid one.\n> > \n> > Our \"git fsck\" already treats certain brokenness (like a tree whose\n> > entry has mode that is 0-padded to the left) as broken but still\n> > tolerate them.  I am not sure if it is sufficient to diagnose and\n> > declare broken and invalid when we see sha3-content that records\n> > these \"mechanically derivable but expensive to compute\" pieces of\n> > information incorrectly.\n> > \n> > I am leaning towards saying \"yes, catching in fsck is enough\" and\n> > suggesting to add generation number to sha3-content of the commit\n> > objects, and to add even the \"original sha1 name\" thing if we find\n> > good use of it.  But I cannot shake this nagging feeling off that I\n> > am missing some huge problems that adding these fields and opening\n> > ourselves to more classes of broken objects.\n> > \n> > Thoughts?\n\nSeeing as current Git versions would always ignore the generation number\n(and therefore work perfectly even with erroneous baked-in generation\nnumbers), and seeing as it would be easy to add a config option to force\nGit to ignore the embedded generation numbers, I would consider `fsck`\ncatching those problems the best idea.\n\nIt seems that every major Git hoster already has some sort of fsck on the\nfly for newly-pushed objects, so that would be another \"line of defense\".\n\nTaking a step back, though, it may be a good idea to leave the generation\nnumber business for later, as much fun as it is to get side tracked and\nfocus on relatively trivial stuff instead of the far more difficult and\ncomplex task to get the transition plan to a new hash ironed out.\n\nFor example, I am still in favor of SHA-256 over SHA3-256, after learning\nsome background details from in-house cryptographers: it provides\nessentially the same level of security, according to my sources, while\nhardware support seems to be coming to SHA-256 a lot sooner than to\nSHA3-256.\n\nWhich hash algorithm to choose is a tough question to answer, and\ndiscussing generation numbers will sadly not help us answer it any quicker.\n\nCiao,\nDscho\n"},{"id":"327936","messageId":"CANgJU+Wv1nx79DJTDmYE=O7LUNA3LuRTJhXJn+y0L0C3R+YDEA@mail.gmail.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709131340030.4132@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2017-09-13T13:43:27Z","receivedAt":"2017-09-13T13:43:33Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On 13 September 2017 at 14:05, Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n> For example, I am still in favor of SHA-256 over SHA3-256, after learning\n> some background details from in-house cryptographers: it provides\n> essentially the same level of security, according to my sources, while\n> hardware support seems to be coming to SHA-256 a lot sooner than to\n> SHA3-256.\n\nFWIW, and I know it is not worth much, as far as I can tell there is\nat least some security/math basis to prefer SHA3-256 to SHA-256.\n\nThe SHA1 and SHA-256 hash functions, (iirc along with their older\ncousins MD5 and MD2) all have a common design feature where they mix a\nrelatively large block size into a much smaller state *each block*. So\nfor instance SHA-256 mixes a 512 bit block into a 256 bit state with a\n2:1 \"leverage\" between the block being read and the state. In SHA1\nthis was worse, mixing a 512 bit block into a 160 bit state, closer to\n3:1 leverage.\n\nSHA3 however uses a completely different design where it mixes a 1088\nbit block into a 1600 bit state, for a leverage of 2:3, and the excess\nis *preserved between each block*.\n\nAssuming everything else is equal between SHA-256 and SHA3 this\ndifference alone would seem to justify choosing SHA3 over SHA-256. We\nknow that there MUST be collisions when compressing a 512 bit block\ninto a 256 bit space, however one cannot say the same about mixing\n1088 bits into a 1600 bit state. The excess state which is not\ndirectly modified by the input block makes a big difference when\nreading the next block.\n\nOf course in both cases we end up compressing the entire source\ndocument down to the same number of bits, however SHA3 does that\n*once*, in finalization only, whereas SHA-256 does it *every* block\nread. So it seems to me that the opportunity for collisions is *much*\nhigher in SHA-256 than it is in SHA3-256. (Even if they should be\nvanishingly rare regardless.)\n\nFor this reason if I had a vote I would definitely vote SHA3-256, or\neven for SHA3-512. The latter has an impressive 1:2 leverage between\nblock and state, and much better theoretical security levels.\n\ncheers,\nYves\nNote: I am not a cryptographer, although I am probably pretty well\ninformed as far hobby-hash-function-enthusiasts go.\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"327946","messageId":"20170913163052.GA27425@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709131340030.4132@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-13T16:30:52Z","receivedAt":"2017-09-13T16:31:31Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi Dscho,\n\nJohannes Schindelin wrote:\n\n> So even if the code to generate a bidirectional old <-> new hash mapping\n> might be with us forever, it *definitely* should be optional (\"optional\"\n> at least as in \"config setting\"), allowing developers who only work with\n> new-hash repositories to save the time and electrons.\n\nAgreed.  This is a good reason not to store the sha1 inside the\nsha256-encoded objects.  I think that is exactly what Brandon was saying\nin response to Junio --- did you read it differently?\n\n[...]\n> ... or Git would simply handle the absence of the generation number header\n> gracefully, so that sha1-content == sha3-content...\n\nPart of the sha1-content is references to other objects using their\nsha1-name, so it is not possible to have sha1-content == sha3-content.\n\nThat said, I am also leaning against including generation numbers as\npart of this design.\n\nThere is an argument for including generation numbers.  It is much\nsimpler to have generation numbers in *all* commit objects than only in\nsome, since it means the slop-based heuristics for faking generation\nnumbers using commit timestamp can be completely avoided for a\nrepository using such a format.  Including generation numbers in all\ncommit objects is a painless thing to do during a format change, since\nit can happen without harming round-tripping.\n\nTreating generation numbers as derived data (as in Jeff King's\npreferred design, if I have understood his replies correctly) would\nalso be possible but it does not interact well with shallow clone or\nnarrow clone.\n\nAll that said, for simplicity I still lean against including\ngeneration numbers as part of a hash function transition.  Nothing\nstops us from having another format change later.\n\nThis is a particularly hard decision because I don't have a strong\npreference.  That leads me to err on the side of simplicity.\n\nI will make sure to discuss this issue in my patch to\nDocumentation/technical/, so we don't have to repeat the same\nconversations again and again.\n\n[...]\n> Taking a step back, though, it may be a good idea to leave the generation\n> number business for later, as much fun as it is to get side tracked and\n> focus on relatively trivial stuff instead of the far more difficult and\n> complex task to get the transition plan to a new hash ironed out.\n>\n> For example, I am still in favor of SHA-256 over SHA3-256, after learning\n> some background details from in-house cryptographers: it provides\n> essentially the same level of security, according to my sources, while\n> hardware support seems to be coming to SHA-256 a lot sooner than to\n> SHA3-256.\n>\n> Which hash algorithm to choose is a tough question to answer, and\n> discussing generation numbers will sadly not help us answer it any quicker.\n\nThis is unrelated to Brandon's message, except for his use of SHA3 as\na placeholder for \"the next hash function\".\n\nMy assumption based on previous conversations (and other external\nconversations like [1]) is that we are going to use SHA2-256 and have\na pretty strong consensus for that.  Don't worry!\n\nAs a side note, I am probably misreading, but I found this set of\nparagraphs a bit condescending.  It sounds to me like you are saying\n\"You are making the wrong choice of hash function and everything else\nyou are describing is irrelevant when compared to that monumental\nmistake.  Please stop working on things I don't consider important\".\nWith that reading it is quite demotivating to read.\n\nAn alternative reading is that you are saying that the transition plan\ndescribed in this thread is not ironed out.  Can you spell that out\nmore?  What particular aspect of the transition plan (which is of\ncourse orthogonal to the choice of hash function) are you discontent\nwith?\n\nThanks and hope that helps,\nJonathan\n\n[1] https://www.imperialviolet.org/2017/05/31/skipsha3.html\n"},{"id":"328001","messageId":"xmqq7ex21d2v.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170913163052.GA27425@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-09-13T21:52:08Z","receivedAt":"2017-09-13T21:52:16Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> Treating generation numbers as derived data (as in Jeff King's\n> preferred design, if I have understood his replies correctly) would\n> also be possible but it does not interact well with shallow clone or\n> narrow clone.\n\nJust like we have skewed committer timestamps, there is no reason to\nbelieve that generation numbers embedded in objects are trustable,\nand there is no way for narrow clients to even verify their correctness.\n\nSo I agree with Peff that having generation numbers in object is\npointless; I agree any other derivables like corresponding sha-1\nname is also pointless to have.\n\nThis is a tangent, but it may be fine for a shallow clone to treat\nthe cut-off points in the history as if they are root commits and\ncompute generation numbers locally, just like everybody else does.\nAs generation numbers won't have to be global (because we will not\nbe embedding them in objects), nobody gets hurt if they do not match\nacross repositories---just like often-mentioned rename detection\ncache, it can be kept as a mere local performance aid and does not\nhave to participate in the object model.\n\n> All that said, for simplicity I still lean against including\n> generation numbers as part of a hash function transition.\n\nGood.\n\n> This is unrelated to Brandon's message, except for his use of SHA3 as\n> a placeholder for \"the next hash function\".\n>\n> My assumption based on previous conversations (and other external\n> conversations like [1]) is that we are going to use SHA2-256 and have\n> a pretty strong consensus for that.  Don't worry!\n\nHmph, I actually re-read the thread recently, and my impression was\nthat we didn't quite have a consensus but were leaning towards\nSHA3-256.\n\nI do not personally have a strong preference myself and I would say\nthat anything will do as long as it is with good longevity and\navailability.  SHA2 family would be a fine choice due to its age on\nboth counts, being scrutinized longer and having a chance to be\nimplemented in many places, even though its age itself may have to\nbe subtracted from the longevity factor.\n\nThanks.\n"},{"id":"328011","messageId":"CAGZ79kakGcMJ7HuH+MPsMrvw40uGchr6H-SQw9-p8pgi3Yk_Bw@mail.gmail.com","threadId":"45288","inReplyTo":"xmqq7ex21d2v.fsf@gitster.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2017-09-13T22:07:08Z","receivedAt":"2017-09-13T22:07:16Z","isPatch":false,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"On Wed, Sep 13, 2017 at 2:52 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n>\n>> Treating generation numbers as derived data (as in Jeff King's\n>> preferred design, if I have understood his replies correctly) would\n>> also be possible but it does not interact well with shallow clone or\n>> narrow clone.\n>\n> Just like we have skewed committer timestamps, there is no reason to\n> believe that generation numbers embedded in objects are trustable,\n> and there is no way for narrow clients to even verify their correctness.\n>\n> So I agree with Peff that having generation numbers in object is\n> pointless; I agree any other derivables like corresponding sha-1\n> name is also pointless to have.\n>\n> This is a tangent, but it may be fine for a shallow clone to treat\n> the cut-off points in the history as if they are root commits and\n> compute generation numbers locally, just like everybody else does.\n> As generation numbers won't have to be global (because we will not\n> be embedding them in objects), nobody gets hurt if they do not match\n> across repositories---just like often-mentioned rename detection\n> cache, it can be kept as a mere local performance aid and does not\n> have to participate in the object model.\n\nLocally it helps for some operations such as correct walks.\nFor the network case however, it doesn't really help either.\n\nIf we had global generation numbers, one could imagine that they\nare used in the pack negotiation (server advertises the maximum\ngeneration number or even gen number per branch; client\ncould binary search in there for the fork point)\n\nI wonder if locally generated generation numbers (for the shallow\ncase) could be used somehow to still improve network operations.\n\n\n\n>> My assumption based on previous conversations (and other external\n>> conversations like [1]) is that we are going to use SHA2-256 and have\n>> a pretty strong consensus for that.  Don't worry!\n>\n> Hmph, I actually re-read the thread recently, and my impression was\n> that we didn't quite have a consensus but were leaning towards\n> SHA3-256.\n>\n> I do not personally have a strong preference myself and I would say\n> that anything will do as long as it is with good longevity and\n> availability.  SHA2 family would be a fine choice due to its age on\n> both counts, being scrutinized longer and having a chance to be\n> implemented in many places, even though its age itself may have to\n> be subtracted from the longevity factor.\n\nIf we'd get the transition somewhat right, the next transition will\nbe easier than the current transition, such that I am not that concerned\nabout longevity. I am rather concerned about the complexity that is added\nto the code base (whilst accumulating technical debt instead of clearer\nabstraction layers)\n"},{"id":"328012","messageId":"xmqq377q1c0g.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"xmqq7ex21d2v.fsf@gitster.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-09-13T22:15:11Z","receivedAt":"2017-09-13T22:15:18Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n>\n>> Treating generation numbers as derived data (as in Jeff King's\n>> preferred design, if I have understood his replies correctly) would\n>> also be possible but it does not interact well with shallow clone or\n>> narrow clone.\n>\n> Just like we have skewed committer timestamps, there is no reason to\n> believe that generation numbers embedded in objects are trustable,\n> and there is no way for narrow clients to even verify their correctness.\n>\n> So I agree with Peff that having generation numbers in object is\n> pointless; I agree any other derivables like corresponding sha-1\n> name is also pointless to have.\n>\n> This is a tangent, but it may be fine for a shallow clone to treat\n> the cut-off points in the history as if they are root commits and\n> compute generation numbers locally, just like everybody else does.\n> As generation numbers won't have to be global (because we will not\n> be embedding them in objects), nobody gets hurt if they do not match\n> across repositories---just like often-mentioned rename detection\n> cache, it can be kept as a mere local performance aid and does not\n> have to participate in the object model.\n>\n>> All that said, for simplicity I still lean against including\n>> generation numbers as part of a hash function transition.\n>\n> Good.\n\nIn the proposed transition plan, the treatment of various signatures\n(deliberately) makes the conversion not quite roundtrip.\n\nWhen existing SHA-1 history in individual clones are converted to\nNewHash, we obviously cannot re-sign the corresponding NewHash\ncontents with the same PGP key, so these converted objects will\ncarry only signature on SHA-1 contents.  They can still be validated\nwhen they are exported back to SHA-1 world via the fetch/push\nprotocol, and can be validated locally by converting them back to\nSHA-1 contents and then passing the result to gpgv.\n\nThe plan also states, if I remember what I read correctly, that\nnewly created and signed objects (this includes signed commits and\nsigned tags; mergetags merely carry over what the tag object that\nwas merged was signed with, so we do not have to worry about them\nunless the resulting commit that has mergetag is signed itself, but\nthat is already covered by how we handle signed commits) would be\nsigned both for NewHash contents and its corresponding SHA-1\ncontents (after internally convering it to SHA-1 contents).  That\nwould allow us to strip the signature over NewHash contents and\nderive the SHA-1 contents to be shown to the outside world while\nmigration is going on and I'd imagine it would be a good practice;\nit would allow us to sign something that allows everybody to verify,\nwhen some participants of the project are not yet NewHash capable.\n\nBut the signing over SHA-1 contents has to stop at some point, when\neverybody's Git becomes completely unaware of SHA-1.  We may want to\nhave a guideline in the transition plan to (1) encourage signing for\nboth for quite some time, and (2) the criteria for us to decide when\nto stop.\n\nThanks.\n"},{"id":"328014","messageId":"20170913221854.GP27425@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"CAGZ79kakGcMJ7HuH+MPsMrvw40uGchr6H-SQw9-p8pgi3Yk_Bw@mail.gmail.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-13T22:18:54Z","receivedAt":"2017-09-13T22:19:03Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nStefan Beller wrote:\n> On Wed, Sep 13, 2017 at 2:52 PM, Junio C Hamano <gitster@pobox.com> wrote:\n\n>> This is a tangent, but it may be fine for a shallow clone to treat\n>> the cut-off points in the history as if they are root commits and\n>> compute generation numbers locally, just like everybody else does.\n[...]\n> Locally it helps for some operations such as correct walks.\n> For the network case however, it doesn't really help either.\n>\n> If we had global generation numbers, one could imagine that they\n> are used in the pack negotiation (server advertises the maximum\n> generation number or even gen number per branch; client\n> could binary search in there for the fork point)\n>\n> I wonder if locally generated generation numbers (for the shallow\n> case) could be used somehow to still improve network operations.\n\nI have a different concern about locally generated generation numbers in\na shallow clone.  My concern is that it is slow to recompute them when\ndeepening the shallow clone.\n\nHowever:\n\n 1. That only affects performance and for some use cases could be\n    mitigated e.g. by introducing some laziness, and, more\n    convincingly,\n\n 2. With a small protocol change, the server could communicate the\n    generation numbers for commit objects at the edge of a shallow\n    clone, avoiding this trouble.\n\nSo I am not too concerned.\n\nMore generally, unless there is a very very compelling reason to, I\ndon't want to couple other changes into the hash function transition.\nIf they're worthwhile enough to do, they're worthwhile enough to do\nwhether we're transitioning to a new hash function or not: I have not\nheard a convincing example yet of a \"while at it\" that is worth the\ncomplexity of such coupling.\n\n(That said, if two format changes are worth doing and happen to be\nimplemented at the same time, then we can save users the trouble of\nexperiencing two format change transitions.  That is a kind of\ncoupling from the end user's point of view.  But from the perspective\nof someone writing the code, there is no need to count on that, and it\nis not likely to happen anyway.)\n\n> If we'd get the transition somewhat right, the next transition will\n> be easier than the current transition, such that I am not that concerned\n> about longevity. I am rather concerned about the complexity that is added\n> to the code base (whilst accumulating technical debt instead of clearer\n> abstraction layers)\n\nDuring the transition, users have to suffer reencoding overhead, so it\nis not good for such transitions to need to happen very often.  If the\nnew hash function breaks early, then we have to cope with it and as\nyou say, having the framework in place means we'd be ready for that.\nBut I still don't want the chosen hash function to break early.\n\nIn other words, a long lifetime for the hash absolutely is a design\ngoal.  Coping well with an unexpectedly short lifetime for the hash is\nalso a design goal.\n\nIf the hash function lasts 10 years then I am happy.\n\nThanks,\nJonathan\n"},{"id":"328016","messageId":"20170913222731.GQ27425@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"xmqq377q1c0g.fsf@gitster.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-13T22:27:31Z","receivedAt":"2017-09-13T22:27:39Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n\n> In the proposed transition plan, the treatment of various signatures\n> (deliberately) makes the conversion not quite roundtrip.\n\nThat's not precisely true.  Details below.\n\n> When existing SHA-1 history in individual clones are converted to\n> NewHash, we obviously cannot re-sign the corresponding NewHash\n> contents with the same PGP key, so these converted objects will\n> carry only signature on SHA-1 contents.  They can still be validated\n> when they are exported back to SHA-1 world via the fetch/push\n> protocol, and can be validated locally by converting them back to\n> SHA-1 contents and then passing the result to gpgv.\n\nCorrect.\n\n> The plan also states, if I remember what I read correctly, that\n> newly created and signed objects (this includes signed commits and\n> signed tags; mergetags merely carry over what the tag object that\n> was merged was signed with, so we do not have to worry about them\n> unless the resulting commit that has mergetag is signed itself, but\n> that is already covered by how we handle signed commits) would be\n> signed both for NewHash contents and its corresponding SHA-1\n> contents (after internally convering it to SHA-1 contents).\n\nAlso correct.\n\n> would allow us to strip the signature over NewHash contents and\n> derive the SHA-1 contents to be shown to the outside world while\n> migration is going on and I'd imagine it would be a good practice;\n> it would allow us to sign something that allows everybody to verify,\n> when some participants of the project are not yet NewHash capable.\n\nThe NewHash-based signature is included in the SHA-1 content as well,\nfor the sake of round-tripping.  It is not stripped out.\n\n> But the signing over SHA-1 contents has to stop at some point, when\n> everybody's Git becomes completely unaware of SHA-1.  We may want to\n> have a guideline in the transition plan to (1) encourage signing for\n> both for quite some time, and (2) the criteria for us to decide when\n> to stop.\n\nYes, spelling out a rough schedule is a good idea.  I'll add that.\n\nA version of Git that is aware of NewHash should be able to verify\nNewHash signatures even for users that are using SHA-1 locally for the\nsake of faster fetches and pushes to SHA-1 based peers.\n\nIn addition to a new enough Git, this requires the translation table\nto translate to NewHash to be present.\n\nSo the criterion (2) is largely based on how up-to-date the Git used\nby users wanting to verify signatures is and whether they are willing\nto tolerate the performance implications of having a translation\ntable.  My hope is that when communicating with peers using the same\nhash function, the translation table will not add too much performance\noverhead.\n\nThank you,\nJonathan\n"},{"id":"328019","messageId":"20170913225158.GR27425@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"CANgJU+Wv1nx79DJTDmYE=O7LUNA3LuRTJhXJn+y0L0C3R+YDEA@mail.gmail.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-13T22:51:58Z","receivedAt":"2017-09-13T22:52:08Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nYves wrote:\n> On 13 September 2017 at 14:05, Johannes Schindelin\n\n>> For example, I am still in favor of SHA-256 over SHA3-256, after learning\n>> some background details from in-house cryptographers: it provides\n>> essentially the same level of security, according to my sources, while\n>> hardware support seems to be coming to SHA-256 a lot sooner than to\n>> SHA3-256.\n>\n> FWIW, and I know it is not worth much, as far as I can tell there is\n> at least some security/math basis to prefer SHA3-256 to SHA-256.\n\nThanks for spelling this out.  From my (very cursory) understanding of\nthe math, what you are saying makes sense.  I think there were some\nhints of this topic on-list before, but not made so explicit before.\n\nHere's my summary of the discussion of other aspects of the choice of\nhash functions so far:\n\nMy understanding from asking cryptographers matches what Dscho said.\nOne of the lessons of the history of hash functions is that some kinds\nof attempts to improve the security margin of a hash function do not\nhelp as much as expected once a function is broken.\n\nIn practice, what we are looking for is\n\n- is the algorithm broken, or likely to be broken soon\n- do the algorithm's guarantees match the application\n- is the algorithm fast enough\n- are high quality implementations widely available\n\nOn that first question, every well informed person I have asked has\nassured me that SHA-256, SHA-512, SHA-512/256, SHA-256x16, SHA3-256,\nK12, BLAKE2bp-256, etc are equally likely to be broken in the next 10\nyears.  The main difference for the longevity question is that some of\nthose algorithms have had more scrutiny than others, but all have had\nsignificant scrutiny.  See [1] and the surrounding thread for more\ndiscussion on that.\n\nOn the second question, SHA-256 is vulnerable to length extension\nattacks, which means it would not be usable as a MAC directly (instead\nof using the HMAC construction).  Fortunately Git doesn't use its hash\nfunction that way.\n\nOn the third question, SHA-256 is one of the slower ones, even with\nhardware accelaration, but it should be fast enough.\n\nOn the fourth question, SHA-256 shines.  See [2].  That is where I had\nthought the conversation ended up.\n\nFor what it's worth, I'm pretty happy both with the level of scrutiny\nwe've given to this question and SHA-256 as an answer.  Luckily even\nif at the last minute we learn something that changes the choice of\nhash function, that would not significantly affect the transition\nplan, so we have a chance to learn more.\n\nSee also [3].\n\nThanks,\nJonathan\n\n[1] https://public-inbox.org/git/CAL9PXLzhPyE+geUdcLmd=pidT5P8eFEBbSgX_dS88knz2q_LSw@mail.gmail.com/#t\n[2] https://public-inbox.org/git/xmqq37azy7ru.fsf@gitster.mtv.corp.google.com/\n[3] https://www.imperialviolet.org/2017/05/31/skipsha3.html,\n    https://news.ycombinator.com/item?id=14453622\n"},{"id":"328021","messageId":"CA+55aFwUn0KibpDQK2ZrxzXKOk8-aAub2nJZQqKCpq1ddhDcMQ@mail.gmail.com","threadId":"45288","inReplyTo":"CANgJU+Wv1nx79DJTDmYE=O7LUNA3LuRTJhXJn+y0L0C3R+YDEA@mail.gmail.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-09-13T23:30:07Z","receivedAt":"2017-09-13T23:30:14Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Wed, Sep 13, 2017 at 6:43 AM, demerphq <demerphq@gmail.com> wrote:\n>\n> SHA3 however uses a completely different design where it mixes a 1088\n> bit block into a 1600 bit state, for a leverage of 2:3, and the excess\n> is *preserved between each block*.\n\nYes. And considering that the SHA1 attack was actually predicated on\nthe fact that each block was independent (no extra state between), I\ndo think SHA3 is a better model.\n\nSo I'd rather see SHA3-256 than SHA256.\n\n              Linus\n"},{"id":"328023","messageId":"xmqqtw06yqrq.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170913222731.GQ27425@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-09-14T02:10:01Z","receivedAt":"2017-09-14T02:10:49Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> The NewHash-based signature is included in the SHA-1 content as well,\n> for the sake of round-tripping.  It is not stripped out.\n\nAh, OK, that allays my worries.  We rely on the fact that unknown\nobject headers from the future are ignored.  We use something other\nthan \"gpgsig\" header (say, \"gpgsigN\") to store NewHash based\nsignature on a commit object created in the NewHash world, so that\nSHA-1 clients will ignore it but still include in the signature\ncomputation---is that the idea?\n\nExisting versions of Git that live in the SHA-1 world may still need\nto learn to ignore/drop \"gpgsigN\" while amending a commit that\noriginally was created in the NewHash world.  Or to force upgrade we\nmay freeze the SHA-1 only versions of Git and stop updating them\naltogether.  I dunno.\n\nWe also need to use something other than \"mergetag\" when carrying\nover the contents of a tag being merged in the NewHash world, but\nI'd imagine that you've thought about this already.\n\nThanks.\n\n\n"},{"id":"328024","messageId":"xmqqpoauyqlp.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170913221854.GP27425@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-09-14T02:13:38Z","receivedAt":"2017-09-14T02:13:45Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> In other words, a long lifetime for the hash absolutely is a design\n> goal.  Coping well with an unexpectedly short lifetime for the hash is\n> also a design goal.\n>\n> If the hash function lasts 10 years then I am happy.\n\nAbsolutely.  When two functions have similar expected remaining life\nand are equally widely supported, then faster is better than slower.\nOtherwise our primary goal when picking the function from candidates\nshould be to optimize for its remaining life and wider availability.\n\nThanks.\n"},{"id":"328045","messageId":"alpine.DEB.2.21.1.1709141119140.4132@virtualbox","threadId":"45288","inReplyTo":"20170913163052.GA27425@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-14T12:39:39Z","receivedAt":"2017-09-14T12:40:17Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Jonathan,\n\nOn Wed, 13 Sep 2017, Jonathan Nieder wrote:\n\n> As a side note, I am probably misreading, but I found this set of\n> paragraphs a bit condescending.  It sounds to me like you are saying\n> \"You are making the wrong choice of hash function and everything else\n> you are describing is irrelevant when compared to that monumental\n> mistake.  Please stop working on things I don't consider important\".\n> With that reading it is quite demotivating to read.\n\nI am sorry you read it that way. I did not feel condescending when I wrote\nthat mail, I felt annoyed by the side track, and anxious. In my mind, the\ntransition is too important for side tracking, and I worry that we are not\nfast enough (imagine what would happen if a better attack was discovered\nthat is not as easily detected as the one we know about?).\n\n> An alternative reading is that you are saying that the transition plan\n> described in this thread is not ironed out.  Can you spell that out\n> more?  What particular aspect of the transition plan (which is of\n> course orthogonal to the choice of hash function) are you discontent\n> with?\n\nMy impression from reading Junio's mail was that he does not consider the\ntransition plan ironed out yet, and that he wants to spend time on\ndiscussing generation numbers right now.\n\nI was in particularly frightened by the suggestion to \"reboot\" [*1*].\nHopefully I misunderstand and he meant \"finishing touches\" instead.\n\nAs to *my* opinion: after reading https://goo.gl/gh2Mzc (is it really\ncorrect that its last update has been on March 6th?), my only concern is\nreally that it still talks about SHA3-256 when I think that the\nperformance benefits of SHA-256 (think: \"Git at scale\", and also hardware\nsupport) really make the latter a better choice.\n\nIn order to be \"ironed out\", I think we need to talk about the\nimplementation detail \"Translation table\". This is important. It needs to\nbe *fast*.\n\nSpeaking of *fast*, I could imagine that it would make sense to store the\nSHA-1 objects on disk, still, instead of converting them on the fly. I am\nnot sure whether this is something we need to define in the document,\nthough, as it may very well be premature optimization; Maybe mention that\nwe could do this if necessary?\n\nApart from that, I would *love* to see this document as The Official Plan\nthat I can Show To The Manager so that I can ask to Allocate Time.\n\nCiao,\nDscho\n\nFootnote *1*:\nhttps://public-inbox.org/git/xmqqa828733s.fsf@gitster.mtv.corp.google.com/\n"},{"id":"328061","messageId":"alpine.DEB.2.21.1.1709141722500.4132@virtualbox","threadId":"45288","inReplyTo":"xmqqpoauyqlp.fsf@gitster.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-14T15:23:28Z","receivedAt":"2017-09-14T15:24:13Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Junio,\n\nOn Thu, 14 Sep 2017, Junio C Hamano wrote:\n\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n> \n> > In other words, a long lifetime for the hash absolutely is a design\n> > goal.  Coping well with an unexpectedly short lifetime for the hash is\n> > also a design goal.\n> >\n> > If the hash function lasts 10 years then I am happy.\n> \n> Absolutely.  When two functions have similar expected remaining life\n> and are equally widely supported, then faster is better than slower.\n> Otherwise our primary goal when picking the function from candidates\n> should be to optimize for its remaining life and wider availability.\n\nSHA-256 has been hammered on a lot more than SHA3-256.\n\nThat would be a strong point in favor of SHA2.\n\nCiao,\nDscho\n"},{"id":"328063","messageId":"CANgJU+UpMu82a09h644GjqKLsYzHq-t7Tv8x=+ybTYP-QqyPtQ@mail.gmail.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709141722500.4132@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"demerphq","fromEmail":"demerphq@gmail.com","sentAt":"2017-09-14T15:45:18Z","receivedAt":"2017-09-14T15:45:24Z","isPatch":false,"sender":{"key":"demerphq@gmail.com","avatar":null},"body":"On 14 September 2017 at 17:23, Johannes Schindelin\n<Johannes.Schindelin@gmx.de> wrote:\n> Hi Junio,\n>\n> On Thu, 14 Sep 2017, Junio C Hamano wrote:\n>\n>> Jonathan Nieder <jrnieder@gmail.com> writes:\n>>\n>> > In other words, a long lifetime for the hash absolutely is a design\n>> > goal.  Coping well with an unexpectedly short lifetime for the hash is\n>> > also a design goal.\n>> >\n>> > If the hash function lasts 10 years then I am happy.\n>>\n>> Absolutely.  When two functions have similar expected remaining life\n>> and are equally widely supported, then faster is better than slower.\n>> Otherwise our primary goal when picking the function from candidates\n>> should be to optimize for its remaining life and wider availability.\n>\n> SHA-256 has been hammered on a lot more than SHA3-256.\n\nLast year that was even more true of SHA1 than it is true of SHA-256 today.\n\nAnyway,\nYves\n-- \nperl -Mre=debug -e \"/just|another|perl|hacker/\"\n"},{"id":"328066","messageId":"20170914163645.GA111021@google.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709141119140.4132@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Brandon Williams","fromEmail":"bmwill@google.com","sentAt":"2017-09-14T16:36:45Z","receivedAt":"2017-09-14T16:36:53Z","isPatch":false,"sender":{"key":"bwilliams.eng@gmail.com","avatar":null},"body":"On 09/14, Johannes Schindelin wrote:\n> Hi Jonathan,\n> \n> On Wed, 13 Sep 2017, Jonathan Nieder wrote:\n> \n> > As a side note, I am probably misreading, but I found this set of\n> > paragraphs a bit condescending.  It sounds to me like you are saying\n> > \"You are making the wrong choice of hash function and everything else\n> > you are describing is irrelevant when compared to that monumental\n> > mistake.  Please stop working on things I don't consider important\".\n> > With that reading it is quite demotivating to read.\n> \n> I am sorry you read it that way. I did not feel condescending when I wrote\n> that mail, I felt annoyed by the side track, and anxious. In my mind, the\n> transition is too important for side tracking, and I worry that we are not\n> fast enough (imagine what would happen if a better attack was discovered\n> that is not as easily detected as the one we know about?).\n> \n> > An alternative reading is that you are saying that the transition plan\n> > described in this thread is not ironed out.  Can you spell that out\n> > more?  What particular aspect of the transition plan (which is of\n> > course orthogonal to the choice of hash function) are you discontent\n> > with?\n> \n> My impression from reading Junio's mail was that he does not consider the\n> transition plan ironed out yet, and that he wants to spend time on\n> discussing generation numbers right now.\n> \n> I was in particularly frightened by the suggestion to \"reboot\" [*1*].\n> Hopefully I misunderstand and he meant \"finishing touches\" instead.\n> \n> As to *my* opinion: after reading https://goo.gl/gh2Mzc (is it really\n> correct that its last update has been on March 6th?), my only concern is\n> really that it still talks about SHA3-256 when I think that the\n> performance benefits of SHA-256 (think: \"Git at scale\", and also hardware\n> support) really make the latter a better choice.\n> \n> In order to be \"ironed out\", I think we need to talk about the\n> implementation detail \"Translation table\". This is important. It needs to\n> be *fast*.\n\nAgreed, when that document was written it was hand waved as an\nimplementation detail but once we should probably stare ironing out\nthose details soon so that we have a concrete plan in place.\n\n> \n> Speaking of *fast*, I could imagine that it would make sense to store the\n> SHA-1 objects on disk, still, instead of converting them on the fly. I am\n> not sure whether this is something we need to define in the document,\n> though, as it may very well be premature optimization; Maybe mention that\n> we could do this if necessary?\n> \n> Apart from that, I would *love* to see this document as The Official Plan\n> that I can Show To The Manager so that I can ask to Allocate Time.\n\nSpeaking of having a concrete plan, we discussed in office the other day\nabout finally converting the doc into a Documentation patch.  That was\nalways are intention but after writing up the doc we got busy working on\nother projects.  Getting it in as a patch (with a more concrete road map)\nis probably the next step we'd need to take.\n\nI do want to echo what jonathan has said in other parts of this thread,\nthat the transition plan itself doesn't depend on which hash function we\nend up going with in the end.  I fully expect that for the transition\nplan to succeed that we'll have infrastructure for dropping in different\nhash functions so that we can do some sort of benchmarking before\nselecting one to use.  This would also give us the ability to more\neasily transition to another hash function when the time comes.\n\n-- \nBrandon Williams\n"},{"id":"328073","messageId":"alpine.DEB.2.21.1.1709141754240.4132@virtualbox","threadId":"45288","inReplyTo":"20170913225158.GR27425@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-14T18:26:01Z","receivedAt":"2017-09-14T18:26:37Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Jonathan,\n\nOn Wed, 13 Sep 2017, Jonathan Nieder wrote:\n\n> [3] https://www.imperialviolet.org/2017/05/31/skipsha3.html,\n\nI had read this short after it was published, and had missed the updates.\nOne link in particular caught my eye:\n\n\thttps://eprint.iacr.org/2012/476\n\nEssentially, the authors demonstrate that using SIMD technology can speed\nup computation by factor 2 for longer messages (2kB being considered\n\"long\" already). It is a little bit unclear to me from a cursory look\nwhether their fast algorithm computes SHA-256, or something similar.\n\nAs the author of that paper is also known to have contributed to OpenSSL,\nI had a quick look and it would appear that a comment in\ncrypto/sha/asm/sha256-mb-x86_64.pl speaking about \"lanes\" suggests that\nOpenSSL uses the ideas from the paper, even if b783858654 (x86_64 assembly\npack: add multi-block AES-NI, SHA1 and SHA256., 2013-10-03) does not talk\nabout the paper specifically.\n\nThe numbers shown in\nhttps://github.com/openssl/openssl/blob/master/crypto/sha/asm/keccak1600-x86_64.pl#L28\nand in\nhttps://github.com/openssl/openssl/blob/master/crypto/sha/asm/sha256-mb-x86_64.pl#L17\nare sufficiently satisfying.\n\nCiao,\nDscho\n"},{"id":"328074","messageId":"20170914184022.GB78683@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709141754240.4132@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-14T18:40:22Z","receivedAt":"2017-09-14T18:40:31Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nJohannes Schindelin wrote:\n> On Wed, 13 Sep 2017, Jonathan Nieder wrote:\n\n>> [3] https://www.imperialviolet.org/2017/05/31/skipsha3.html,\n>\n> I had read this short after it was published, and had missed the updates.\n> One link in particular caught my eye:\n>\n> \thttps://eprint.iacr.org/2012/476\n>\n> Essentially, the authors demonstrate that using SIMD technology can speed\n> up computation by factor 2 for longer messages (2kB being considered\n> \"long\" already). It is a little bit unclear to me from a cursory look\n> whether their fast algorithm computes SHA-256, or something similar.\n\nThe latter: that paper is about a variant on SHA-256 called SHA-256x4\n(or SHA-256x16 to take advantage of newer instructions).  It's a\ndifferent hash function.  This is what I was alluding to at [1].\n\n> As the author of that paper is also known to have contributed to OpenSSL,\n> I had a quick look and it would appear that a comment in\n> crypto/sha/asm/sha256-mb-x86_64.pl speaking about \"lanes\" suggests that\n> OpenSSL uses the ideas from the paper, even if b783858654 (x86_64 assembly\n> pack: add multi-block AES-NI, SHA1 and SHA256., 2013-10-03) does not talk\n> about the paper specifically.\n>\n> The numbers shown in\n> https://github.com/openssl/openssl/blob/master/crypto/sha/asm/keccak1600-x86_64.pl#L28\n> and in\n> https://github.com/openssl/openssl/blob/master/crypto/sha/asm/sha256-mb-x86_64.pl#L17\n>\n> are sufficiently satisfying.\n\nThis one is about actual SHA-256, but computing the hash of multiple\nstreams in a single funtion call.  The paper to read is [2].  We could\nprobably take advantage of it for e.g. bulk-checkin and index-pack.\nMost other code paths that compute hashes wouldn't be able to benefit\nfrom it.\n\nThanks,\nJonathan\n\n[1] https://public-inbox.org/git/20170616212414.GC133952@aiede.mtv.corp.google.com/\n[2] https://eprint.iacr.org/2012/371\n"},{"id":"328075","messageId":"alpine.DEB.2.21.1.1709142037490.4132@virtualbox","threadId":"45288","inReplyTo":"CA+55aFwUn0KibpDQK2ZrxzXKOk8-aAub2nJZQqKCpq1ddhDcMQ@mail.gmail.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-14T18:45:35Z","receivedAt":"2017-09-14T18:46:11Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Linus,\n\nOn Wed, 13 Sep 2017, Linus Torvalds wrote:\n\n> On Wed, Sep 13, 2017 at 6:43 AM, demerphq <demerphq@gmail.com> wrote:\n> >\n> > SHA3 however uses a completely different design where it mixes a 1088\n> > bit block into a 1600 bit state, for a leverage of 2:3, and the excess\n> > is *preserved between each block*.\n> \n> Yes. And considering that the SHA1 attack was actually predicated on\n> the fact that each block was independent (no extra state between), I\n> do think SHA3 is a better model.\n> \n> So I'd rather see SHA3-256 than SHA256.\n\nSHA-256 got much more cryptanalysis than SHA3-256, and apart from the\nlength-extension problem that does not affect Git's usage, there are no\nknown weaknesses so far.\n\nIt would seem that the experts I talked to were much more concerned about\nthat amount of attention than the particulars of the algorithm. My\nimpression was that the new features of SHA3 were less studied than the\nwell-known features of SHA2, and that the new-ness of SHA3 is not\nnecessarily a good thing.\n\nYou will have to deal with the fact that I trust the crypto experts'\nopinion on this a lot more than your opinion. Sure, you learned from the\nfact that you had been warned about SHA-1 already seeing theoretical\nattacks in 2005 and still choosing to hard-wire it into Git. And yet, you\nare still no more of a cryptography expert than I am.\n\nCiao,\nDscho\n"},{"id":"328076","messageId":"20170914184915.GC78683@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709141119140.4132@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-14T18:49:15Z","receivedAt":"2017-09-14T18:49:27Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Johannes Schindelin wrote:\n> On Wed, 13 Sep 2017, Jonathan Nieder wrote:\n\n>> As a side note, I am probably misreading, but I found this set of\n>> paragraphs a bit condescending.  It sounds to me like you are saying\n>> \"You are making the wrong choice of hash function and everything else\n>> you are describing is irrelevant when compared to that monumental\n>> mistake.  Please stop working on things I don't consider important\".\n>> With that reading it is quite demotivating to read.\n>\n> I am sorry you read it that way. I did not feel condescending when I wrote\n> that mail, I felt annoyed by the side track, and anxious. In my mind, the\n> transition is too important for side tracking, and I worry that we are not\n> fast enough (imagine what would happen if a better attack was discovered\n> that is not as easily detected as the one we know about?).\n\nThanks for clarifying.  That makes sense.\n\n[...]\n> As to *my* opinion: after reading https://goo.gl/gh2Mzc (is it really\n> correct that its last update has been on March 6th?), my only concern is\n> really that it still talks about SHA3-256 when I think that the\n> performance benefits of SHA-256 (think: \"Git at scale\", and also hardware\n> support) really make the latter a better choice.\n>\n> In order to be \"ironed out\", I think we need to talk about the\n> implementation detail \"Translation table\". This is important. It needs to\n> be *fast*.\n>\n> Speaking of *fast*, I could imagine that it would make sense to store the\n> SHA-1 objects on disk, still, instead of converting them on the fly. I am\n> not sure whether this is something we need to define in the document,\n> though, as it may very well be premature optimization; Maybe mention that\n> we could do this if necessary?\n>\n> Apart from that, I would *love* to see this document as The Official Plan\n> that I can Show To The Manager so that I can ask to Allocate Time.\n\nSounds promising!\n\nThanks much for this feedback.  This is very helpful for knowing what\nv4 of the doc needs.\n\nThe discussion of the translation table in [1] didn't make it to the\ndoc.  You're right that it needs to.\n\nCaching SHA-1 objects (and the pros and cons involved) makes sense to\nmention in an \"ideas for future work\" section.\n\nAn implementation plan with well-defined pieces for people to take on\nand estimates of how much work each involves may be useful for Showing\nTo The Manager.  So I'll include a sketch of that for reviewers to\npoke holes in, too.\n\nAnother thing the doc doesn't currently describe is how Git protocol\nwould work.  That's worth sketching in a \"future work\" section as\nwell.\n\nSorry it has been taking so long to get this out.  I think we should\nhave something ready to send on Monday.\n\nThanks,\nJonathan\n\n[1] https://public-inbox.org/git/CAJo=hJtoX9=AyLHHpUJS7fueV9ciZ_MNpnEPHUz8Whui6g9F0A@mail.gmail.com/\n"},{"id":"328078","messageId":"alpine.DEB.2.21.1.1709142357090.219280@virtualbox","threadId":"45288","inReplyTo":"CANgJU+UpMu82a09h644GjqKLsYzHq-t7Tv8x=+ybTYP-QqyPtQ@mail.gmail.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-14T22:06:32Z","receivedAt":"2017-09-14T22:07:10Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 14 Sep 2017, demerphq wrote:\n\n> On 14 September 2017 at 17:23, Johannes Schindelin\n> <Johannes.Schindelin@gmx.de> wrote:\n> >\n> > SHA-256 has been hammered on a lot more than SHA3-256.\n> \n> Last year that was even more true of SHA1 than it is true of SHA-256\n> today.\n\nI hope you are not deliberately trying to annoy me. I say that because you\nseemed to be interested enough in cryptography to know that the known\nattacks on SHA-256 *today* are unlikely to extend to Git's use case,\nwhereas the known attacks on SHA-1 *in 2005* were already raising doubts.\n\nSo while SHA-1 has been hammered on for longer than SHA-256, the latter\ncame out a lot less scathed than the former.\n\nBesides, you are totally missing the point here that the choice is *not*\nbetween SHA-1 and SHA-256, but between SHA-256 and SHA3-256.\n\nAfter all, we would not consider any hash algorithm with known problems\n(as far as Git's usage is concerned). The amount of scrutiny with which\nthe algorithm was investigated would only be a deciding factor among the\nremaining choices, yes?\n\nIn any case, don't trust me on cryptography (just like I do not trust you\non that matter). Trust the cryptographers. I contacted some of my\ncolleagues who are responsible for crypto, and the two who seem to\ndisagree on pretty much everything agreed on this one thing: that SHA-256\nwould be a good choice for Git (and one of them suggested that it would be\nmuch better than SHA3-256, because SHA-256 saw more cryptanalysis).\n\nCiao,\nJohannes\n"},{"id":"328080","messageId":"alpine.DEB.2.21.1.1709150008390.219280@virtualbox","threadId":"45288","inReplyTo":"20170914184022.GB78683@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-14T22:09:29Z","receivedAt":"2017-09-14T22:10:00Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Jonathan,\n\nOn Thu, 14 Sep 2017, Jonathan Nieder wrote:\n\n> Johannes Schindelin wrote:\n> > On Wed, 13 Sep 2017, Jonathan Nieder wrote:\n> \n> >> [3] https://www.imperialviolet.org/2017/05/31/skipsha3.html,\n> >\n> > I had read this short after it was published, and had missed the updates.\n> > One link in particular caught my eye:\n> >\n> > \thttps://eprint.iacr.org/2012/476\n> >\n> > Essentially, the authors demonstrate that using SIMD technology can speed\n> > up computation by factor 2 for longer messages (2kB being considered\n> > \"long\" already). It is a little bit unclear to me from a cursory look\n> > whether their fast algorithm computes SHA-256, or something similar.\n> \n> The latter: that paper is about a variant on SHA-256 called SHA-256x4\n> (or SHA-256x16 to take advantage of newer instructions).  It's a\n> different hash function.  This is what I was alluding to at [1].\n\nThanks for the explanation!\n\n> > As the author of that paper is also known to have contributed to OpenSSL,\n> > I had a quick look and it would appear that a comment in\n> > crypto/sha/asm/sha256-mb-x86_64.pl speaking about \"lanes\" suggests that\n> > OpenSSL uses the ideas from the paper, even if b783858654 (x86_64 assembly\n> > pack: add multi-block AES-NI, SHA1 and SHA256., 2013-10-03) does not talk\n> > about the paper specifically.\n> >\n> > The numbers shown in\n> > https://github.com/openssl/openssl/blob/master/crypto/sha/asm/keccak1600-x86_64.pl#L28\n> > and in\n> > https://github.com/openssl/openssl/blob/master/crypto/sha/asm/sha256-mb-x86_64.pl#L17\n> >\n> > are sufficiently satisfying.\n> \n> This one is about actual SHA-256, but computing the hash of multiple\n> streams in a single funtion call.  The paper to read is [2].  We could\n> probably take advantage of it for e.g. bulk-checkin and index-pack.\n> Most other code paths that compute hashes wouldn't be able to benefit\n> from it.\n\nAgain, thanks for the explanation.\n\nCiao,\nDscho\n\n> [1] https://public-inbox.org/git/20170616212414.GC133952@aiede.mtv.corp.google.com/\n> [2] https://eprint.iacr.org/2012/371\n> \n"},{"id":"328164","messageId":"12CC12FA3A034D6A9B91695BE1A04641@PhilipOakley","threadId":"45288","inReplyTo":"20170914184915.GC78683@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":"2017-09-15T20:42:18Z","receivedAt":"2017-09-15T20:42:27Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi Jonathan,\n\n\"Jonathan Nieder\" <jrnieder@gmail.com> wrote;\n> Johannes Schindelin wrote:\n>> On Wed, 13 Sep 2017, Jonathan Nieder wrote:\n>\n>>> As a side note, I am probably misreading, but I found this set of\n>>> paragraphs a bit condescending.  It sounds to me like you are saying\n>>> \"You are making the wrong choice of hash function and everything else\n>>> you are describing is irrelevant when compared to that monumental\n>>> mistake.  Please stop working on things I don't consider important\".\n>>> With that reading it is quite demotivating to read.\n>>\n>> I am sorry you read it that way. I did not feel condescending when I \n>> wrote\n>> that mail, I felt annoyed by the side track, and anxious. In my mind, the\n>> transition is too important for side tracking, and I worry that we are \n>> not\n>> fast enough (imagine what would happen if a better attack was discovered\n>> that is not as easily detected as the one we know about?).\n>\n> Thanks for clarifying.  That makes sense.\n>\n> [...]\n>> As to *my* opinion: after reading https://goo.gl/gh2Mzc (is it really\n>> correct that its last update has been on March 6th?), my only concern is\n>> really that it still talks about SHA3-256 when I think that the\n>> performance benefits of SHA-256 (think: \"Git at scale\", and also hardware\n>> support) really make the latter a better choice.\n>>\n>> In order to be \"ironed out\", I think we need to talk about the\n>> implementation detail \"Translation table\". This is important. It needs to\n>> be *fast*.\n>>\n>> Speaking of *fast*, I could imagine that it would make sense to store the\n>> SHA-1 objects on disk, still, instead of converting them on the fly. I am\n>> not sure whether this is something we need to define in the document,\n>> though, as it may very well be premature optimization; Maybe mention that\n>> we could do this if necessary?\n>>\n>> Apart from that, I would *love* to see this document as The Official Plan\n>> that I can Show To The Manager so that I can ask to Allocate Time.\n>\n> Sounds promising!\n>\n> Thanks much for this feedback.  This is very helpful for knowing what\n> v4 of the doc needs.\n>\n> The discussion of the translation table in [1] didn't make it to the\n> doc.  You're right that it needs to.\n>\n> Caching SHA-1 objects (and the pros and cons involved) makes sense to\n> mention in an \"ideas for future work\" section.\n>\n> An implementation plan with well-defined pieces for people to take on\n> and estimates of how much work each involves may be useful for Showing\n> To The Manager.  So I'll include a sketch of that for reviewers to\n> poke holes in, too.\n>\n> Another thing the doc doesn't currently describe is how Git protocol\n> would work.  That's worth sketching in a \"future work\" section as\n> well.\n>\n> Sorry it has been taking so long to get this out.  I think we should\n> have something ready to send on Monday.\n\nI had a look at the current doc  https://goo.gl/gh2Mzc and thought that the \nselection of the \"NewHash\" should be separated out into a section of it's \nown as a 'separation of concerns', so that the general transition plan only \nrefers to the \"NewHash\", so as not to accidentally pre-judge that selection.\n\nI did look up the arguments regarding sha2 (sha256) versus sha3-256 and \nfound these two Q&A items\n\nhttps://security.stackexchange.com/questions/152360/should-we-be-using-sha3-2017\n\nhttps://security.stackexchange.com/questions/86283/how-does-sha3-keccak-shake-compare-to-sha2-should-i-use-non-shake-parameter\n\nwith an onward link to this:\n https://www.imperialviolet.org/2012/10/21/nist.html\n\n\"NIST may not have you in mind (21 Oct 2012)\"\n\n\"A couple of weeks back, NIST announced that Keccak would be SHA-3. Keccak \nhas somewhat disappointing software performance but is a gift to hardware \nimplementations.\"\n\nwhich does appear to cover some of the concerns that dscho had noted, and \nspeed does appear to be a core Git selling point.\n\nIt would be worth at least covering these trade offs in the \"select a \nNewHash\" section of the document, as at the end of the day it will be a \npolitical judgement about what the future might hold regarding the \ncontenders.\n\nWhat may also be worth noting is the fall back plan should the chosen \nNewHash be the first to fail, perhaps spectacularly, as having a ready plan \ncould support the choice at risk.\n\n>\n> Thanks,\n> Jonathan\n>\n> [1] \n> https://public-inbox.org/git/CAJo=hJtoX9=AyLHHpUJS7fueV9ciZ_MNpnEPHUz8Whui6g9F0A@mail.gmail.com/\n\n--\nPhilip \n\n"},{"id":"328284","messageId":"59BFB95D.1030903@st.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709142037490.4132@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Gilles Van Assche","fromEmail":"gilles.vanassche@st.com","sentAt":"2017-09-18T12:17:33Z","receivedAt":"2017-09-18T12:17:19Z","isPatch":false,"sender":{"key":"gilles.vanassche@st.com","avatar":null},"body":"Hi Johannes,\n\n> SHA-256 got much more cryptanalysis than SHA3-256 […].\n\nI do not think this is true. Keccak/SHA-3 actually got (and is still\ngetting) a lot of cryptanalysis, with papers published at renowned\ncrypto conferences [1].\n\nKeccak/SHA-3 is recognized to have a significant safety margin. E.g.,\none can cut the number of rounds in half (as in Keyak or KangarooTwelve)\nand still get a very strong function. I don't think we could say the\nsame for SHA-256 or SHA-512…\n\nKind regards,\nGilles, for the Keccak team\n\n[1] https://keccak.team/third_party.html\n\n"},{"id":"328320","messageId":"alpine.DEB.2.21.1.1709182340350.219280@virtualbox","threadId":"45288","inReplyTo":"59BFB95D.1030903@st.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-18T22:16:32Z","receivedAt":"2017-09-18T22:17:17Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Gilles,\n\nOn Mon, 18 Sep 2017, Gilles Van Assche wrote:\n\n> > SHA-256 got much more cryptanalysis than SHA3-256 […].\n> \n> I do not think this is true.\n\nPlease read what I said again: SHA-256 got much more cryptanalysis than\nSHA3-256.\n\nI never said that SHA3-256 got little cryptanalysis. Personally, I think\nthat SHA3-256 got a ton more cryptanalysis than SHA-1, and that SHA-256\n*still* got more cryptanalysis. But my opinion does not count, really.\nHowever, the two experts I pestered with questions over questions left me\nwith that strong impression, and their opinion does count.\n\n> Keccak/SHA-3 actually got (and is still getting) a lot of cryptanalysis,\n> with papers published at renowned crypto conferences [1].\n> \n> Keccak/SHA-3 is recognized to have a significant safety margin. E.g.,\n> one can cut the number of rounds in half (as in Keyak or KangarooTwelve)\n> and still get a very strong function. I don't think we could say the\n> same for SHA-256 or SHA-512…\n\nAgain, I do not want to criticize SHA3/Keccak. Personally, I have a lot of\nrespect for Keccak.\n\nI also have a lot of respect for everybody who scrutinized the SHA2 family\nof algorithms.\n\nI also respect the fact that there are more implementations of SHA-256,\nand thanks to everybody seeming to demand SHA-256 checksums instead of\nSHA-1 or MD5 for downloads, bugs in those implementations are probably\ndiscovered relatively quickly, and I also cannot ignore the prospect of\nhardware support for SHA-256.\n\nIn any case, having SHA3 as a fallback in case SHA-256 gets broken seems\nlike a very good safety net to me.\n\nCiao,\nJohannes"},{"id":"328321","messageId":"20170918222540.GX27425@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"59BFB95D.1030903@st.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-18T22:25:40Z","receivedAt":"2017-09-18T22:25:49Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nGilles Van Assche wrote:\n> Hi Johannes,\n\n>> SHA-256 got much more cryptanalysis than SHA3-256 […].\n>\n> I do not think this is true. Keccak/SHA-3 actually got (and is still\n> getting) a lot of cryptanalysis, with papers published at renowned\n> crypto conferences [1].\n>\n> Keccak/SHA-3 is recognized to have a significant safety margin. E.g.,\n> one can cut the number of rounds in half (as in Keyak or KangarooTwelve)\n> and still get a very strong function. I don't think we could say the\n> same for SHA-256 or SHA-512…\n\nI just wanted to thank you for paying attention to this conversation\nand weighing in.\n\nMost of the regulars in the git project are not crypto experts.  This\nkind of extra information (and e.g. [2]) is very useful to us.\n\nThanks,\nJonathan\n\n> Kind regards,\n> Gilles, for the Keccak team\n>\n> [1] https://keccak.team/third_party.html\n[2] https://public-inbox.org/git/91a34c5b-7844-3db2-cf29-411df5bcf886@noekeon.org/\n"},{"id":"328394","messageId":"59C149A3.6080506@st.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709182340350.219280@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Gilles Van Assche","fromEmail":"gilles.vanassche@st.com","sentAt":"2017-09-19T16:45:23Z","receivedAt":"2017-09-19T16:45:03Z","isPatch":false,"sender":{"key":"gilles.vanassche@st.com","avatar":null},"body":"Hi Johannes,\n\nThanks for your feedback.\n\nOn 19/09/17 00:16, Johannes Schindelin wrote:\n>>> SHA-256 got much more cryptanalysis than SHA3-256 […]. \n>>\n>> I do not think this is true. \n>\n> Please read what I said again: SHA-256 got much more cryptanalysis\n> than SHA3-256.\n\nIndeed. What I meant is that SHA3-256 got at least as much cryptanalysis\nas SHA-256. :-)\n\n> I never said that SHA3-256 got little cryptanalysis. Personally, I\n> think that SHA3-256 got a ton more cryptanalysis than SHA-1, and that\n> SHA-256 *still* got more cryptanalysis. But my opinion does not count,\n> really. However, the two experts I pestered with questions over\n> questions left me with that strong impression, and their opinion does\n> count.\n\nOK, I respect your opinion and that of your two experts. Yet, the \"much\nmore\" part of your statement, in particular, is something that may\nrequire a bit more explanations.\n\nKind regards,\nGilles\n\n"},{"id":"328951","messageId":"20170926170502.GY31762@io.lakedaemon.net","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709142037490.4132@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jason Cooper","fromEmail":"jason@lakedaemon.net","sentAt":"2017-09-26T17:05:03Z","receivedAt":"2017-09-26T17:21:27Z","isPatch":false,"sender":{"key":"jason@lakedaemon.net","avatar":null},"body":"Hi all,\n\nSorry for late commentary...\n\nOn Thu, Sep 14, 2017 at 08:45:35PM +0200, Johannes Schindelin wrote:\n> On Wed, 13 Sep 2017, Linus Torvalds wrote:\n> > On Wed, Sep 13, 2017 at 6:43 AM, demerphq <demerphq@gmail.com> wrote:\n> > > SHA3 however uses a completely different design where it mixes a 1088\n> > > bit block into a 1600 bit state, for a leverage of 2:3, and the excess\n> > > is *preserved between each block*.\n> > \n> > Yes. And considering that the SHA1 attack was actually predicated on\n> > the fact that each block was independent (no extra state between), I\n> > do think SHA3 is a better model.\n> > \n> > So I'd rather see SHA3-256 than SHA256.\n\nWell, for what it's worth, we need to be aware that SHA3 is *different*.\nIn crypto, \"different\" = \"bugs haven't been found yet\".  :-P\n\nAnd SHA2 is *known*.  So we have a pretty good handle on how it'll\nweaken over time.\n\n> SHA-256 got much more cryptanalysis than SHA3-256, and apart from the\n> length-extension problem that does not affect Git's usage, there are no\n> known weaknesses so far.\n\nWhile I think that statement is true on it's face (particularly when\nincluding post-competition analysis), I don't think it's sufficient\njustification to chose one over the other.\n\n> It would seem that the experts I talked to were much more concerned about\n> that amount of attention than the particulars of the algorithm. My\n> impression was that the new features of SHA3 were less studied than the\n> well-known features of SHA2, and that the new-ness of SHA3 is not\n> necessarily a good thing.\n\nThe only thing I really object to here is the abstract \"experts\".  We're\ntalking about cryptography and integrity here.  It's no longer\nsufficient to cite anonymous experts.  Either they can put their\nthoughts, opinions and analysis on record here, or it shouldn't be\nconsidered.  Sorry.\n\nOther than their anonymity, though, I do agree with your experts\nassessments.\n\nHowever, whether we chose SHA2 or SHA3 doesn't matter.  Moving away from\nSHA1 does.  Once the object_id code is in place to facilitate that\ntransition, the problem is solved from git's perspective.\n\nIf SHA3 is chosen as the successor, it's going to get a *lot* more\nadoption, and thus, a lot more analysis.  If cracks start to show, the\nhard work of making git flexible is already done.  We can migrate to\nSHA4/5/whatever in an orderly fashion with far less effort than the\ntransition away from SHA1.\n\nFor my use cases, as a user of git, I have a plan to maintain provable\nintegrity of existing objects stored in git under sha1 while migrating\naway from sha1.  The same plan works for migrating away from SHA2 or\nSHA3 when the time comes.\n\n\nthx,\n\nJason.\n"},{"id":"328966","messageId":"alpine.DEB.2.21.1.1709262356360.40514@virtualbox","threadId":"45288","inReplyTo":"20170926170502.GY31762@io.lakedaemon.net","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-26T22:11:14Z","receivedAt":"2017-09-26T22:12:26Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Jason,\n\nOn Tue, 26 Sep 2017, Jason Cooper wrote:\n\n> On Thu, Sep 14, 2017 at 08:45:35PM +0200, Johannes Schindelin wrote:\n> > On Wed, 13 Sep 2017, Linus Torvalds wrote:\n> > > On Wed, Sep 13, 2017 at 6:43 AM, demerphq <demerphq@gmail.com> wrote:\n> > > > SHA3 however uses a completely different design where it mixes a 1088\n> > > > bit block into a 1600 bit state, for a leverage of 2:3, and the excess\n> > > > is *preserved between each block*.\n> > > \n> > > Yes. And considering that the SHA1 attack was actually predicated on\n> > > the fact that each block was independent (no extra state between), I\n> > > do think SHA3 is a better model.\n> > > \n> > > So I'd rather see SHA3-256 than SHA256.\n> \n> Well, for what it's worth, we need to be aware that SHA3 is *different*.\n> In crypto, \"different\" = \"bugs haven't been found yet\".  :-P\n> \n> And SHA2 is *known*.  So we have a pretty good handle on how it'll\n> weaken over time.\n\nHere, you seem to agree with me.\n\n> > SHA-256 got much more cryptanalysis than SHA3-256, and apart from the\n> > length-extension problem that does not affect Git's usage, there are no\n> > known weaknesses so far.\n> \n> While I think that statement is true on it's face (particularly when\n> including post-competition analysis), I don't think it's sufficient\n> justification to chose one over the other.\n\nAnd here you don't.\n\nI find that very confusing.\n\n> > It would seem that the experts I talked to were much more concerned about\n> > that amount of attention than the particulars of the algorithm. My\n> > impression was that the new features of SHA3 were less studied than the\n> > well-known features of SHA2, and that the new-ness of SHA3 is not\n> > necessarily a good thing.\n> \n> The only thing I really object to here is the abstract \"experts\".  We're\n> talking about cryptography and integrity here.  It's no longer\n> sufficient to cite anonymous experts.  Either they can put their\n> thoughts, opinions and analysis on record here, or it shouldn't be\n> considered.  Sorry.\n\nSorry, you are asking cryptography experts to spend their time on the Git\nmailing list. I tried to get them to speak out on the Git mailing list.\nThey respectfully declined.\n\nI can't fault them, they have real jobs to do, and none of their managers\nwould be happy for them to educate the Git mailing list on matters of\ncryptography, not after what happened in 2005.\n\n> Other than their anonymity, though, I do agree with your experts\n> assessments.\n\nI know what our in-house cryptography experts have to prove to start\nworking at Microsoft. Forgive me, but you are not a known entity to me.\n\n> However, whether we chose SHA2 or SHA3 doesn't matter.\n\nTo you, it does not matter.\n\nTo me, it matters. To the several thousand developers working on Windows,\nprobably the largest Git repository in active use, it matters. It matters\nbecause the speed difference that has little impact on you has a lot more\nimpact on us.\n\n> Moving away from SHA1 does.  Once the object_id code is in place to\n> facilitate that transition, the problem is solved from git's\n> perspective.\n\nUh oh. You forgot the mapping. And the protocol. And pretty much\neverything except the oid.\n\n> If SHA3 is chosen as the successor, it's going to get a *lot* more\n> adoption, and thus, a lot more analysis.  If cracks start to show, the\n> hard work of making git flexible is already done.  We can migrate to\n> SHA4/5/whatever in an orderly fashion with far less effort than the\n> transition away from SHA1.\n\nSure. And if XYZ789 is chosen, it's going to get a *lot* more adoption,\ntoo.\n\nWe think.\n\nLet's be realistic. Git is pretty important to us, but it is not important\nenough to sway, say, Intel into announcing hardware support for SHA3.\n\nAnd if you try to force through *any* hash function only so that it gets\nmore adoption and hence more support, in the short run you will make life\nharder for developers on more obscure platforms, who may not easily get\nhigh-quality, high-speed implementations of anything but the very\nmainstream (which is, let's face it, MD5, SHA-1 and SHA-256). I know I\nwould have cursed you for such a decision back when I had to work on AIX\nand IRIX.\n\n> For my use cases, as a user of git, I have a plan to maintain provable\n> integrity of existing objects stored in git under sha1 while migrating\n> away from sha1.  The same plan works for migrating away from SHA2 or\n> SHA3 when the time comes.\n\nPlease do not make the mistake of taking your use case to be a template\nfor everybody's use case.\n\nMigrating a large team away from any hash function to another one *will*\nbe painful, and costly.\n\nMigrating will be very costly for hosting companies like GitHub, Microsoft\nand BitBucket, too.\n\nCiao,\nJohannes\n"},{"id":"328968","messageId":"20170926222558.22418-1-sbeller@google.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709262356360.40514@virtualbox","subject":"[PATCH] technical doc: add a design doc for hash function transition","fromName":"Stefan Beller","fromEmail":"sbeller@google.com","sentAt":"2017-09-26T22:25:58Z","receivedAt":"2017-09-26T22:26:25Z","isPatch":true,"sender":{"key":"stefanbeller@gmail.com","avatar":"https://avatars.githubusercontent.com/u/455868?v=4"},"body":"From: Jonathan Nieder <jrn@google.com>\n\nThis is \"RFC v3: Another proposed hash function transition plan\" from\nthe git mailing list.\n\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Brandon Williams <bmwill@google.com>\nSigned-off-by: Stefan Beller <sbeller@google.com>\n---\n\n This takes the original Google Doc[1] and adds it to our history,\n such that the discussion can be on on list and in the commit messages.\n \n * replaced SHA3-256 with NEWHASH, sha3 with newhash\n * added section 'Implementation plan'\n * added section 'Future work'\n * added section 'Agreed-upon criteria for selecting NewHash'\n \n As the discussion restarts again, here is our attempt\n to add value to the discussion, we planned to polish it more, but as the\n discussion is restarting, we might just post it as-is.\n  \n Thanks.\n\n[1] https://docs.google.com/document/d/18hYAQCTsDgaFUo-VJGhT0UqyetL2LbAzkWNK1fYS8R0/edit\n\n Documentation/Makefile                             |   1 +\n .../technical/hash-function-transition.txt         | 571 +++++++++++++++++++++\n 2 files changed, 572 insertions(+)\n create mode 100644 Documentation/technical/hash-function-transition.txt\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 2415e0d657..471bb29725 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -67,6 +67,7 @@ SP_ARTICLES += howto/maintain-git\n API_DOCS = $(patsubst %.txt,%,$(filter-out technical/api-index-skel.txt technical/api-index.txt, $(wildcard technical/api-*.txt)))\n SP_ARTICLES += $(API_DOCS)\n \n+TECH_DOCS += technical/hash-function-transition\n TECH_DOCS += technical/http-protocol\n TECH_DOCS += technical/index-format\n TECH_DOCS += technical/pack-format\ndiff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\nnew file mode 100644\nindex 0000000000..0ac751d600\n--- /dev/null\n+++ b/Documentation/technical/hash-function-transition.txt\n@@ -0,0 +1,571 @@\n+Git hash function transition\n+============================\n+\n+Objective\n+---------\n+Migrate Git from SHA-1 to a stronger hash function.\n+\n+Background\n+----------\n+At its core, the Git version control system is a content addressable\n+filesystem. It uses the SHA-1 hash function to name content. For\n+example, files, directories, and revisions are referred to by hash\n+values unlike in other traditional version control systems where files\n+or versions are referred to via sequential numbers. The use of a hash\n+function to address its content delivers a few advantages:\n+\n+* Integrity checking is easy. Bit flips, for example, are easily\n+  detected, as the hash of corrupted content does not match its name.\n+* Lookup of objects is fast.\n+\n+Using a cryptographically secure hash function brings additional\n+advantages:\n+\n+* Object names can be signed and third parties can trust the hash to\n+  address the signed object and all objects it references.\n+* Communication using Git protocol and out of band communication\n+  methods have a short reliable string that can be used to reliably\n+  address stored content.\n+\n+Over time some flaws in SHA-1 have been discovered by security\n+researchers. https://shattered.io demonstrated a practical SHA-1 hash\n+collision. As a result, SHA-1 cannot be considered cryptographically\n+secure any more. This impacts the communication of hash values because\n+we cannot trust that a given hash value represents the known good\n+version of content that the speaker intended.\n+\n+SHA-1 still possesses the other properties such as fast object lookup\n+and safe error checking, but other hash functions are equally suitable\n+that are believed to be cryptographically secure.\n+\n+Goals\n+-----\n+1. The transition to NEWHASH can be done one local repository at a time.\n+   a. Requiring no action by any other party.\n+   b. A NEWHASH repository can communicate with SHA-1 Git servers\n+      (push/fetch).\n+   c. Users can use SHA-1 and NEWHASH identifiers for objects\n+      interchangeably.\n+   d. New signed objects make use of a stronger hash function than\n+      SHA-1 for their security guarantees.\n+2. Allow a complete transition away from SHA-1.\n+   a. Local metadata for SHA-1 compatibility can be removed from a\n+      repository if compatibility with SHA-1 is no longer needed.\n+3. Maintainability throughout the process.\n+   a. The object format is kept simple and consistent.\n+   b. Creation of a generalized repository conversion tool.\n+\n+Non-Goals\n+---------\n+1. Add NEWHASH support to Git protocol. This is valuable and the\n+   logical next step but it is out of scope for this initial design.\n+2. Transparently improving the security of existing SHA-1 signed\n+   objects.\n+3. Intermixing objects using multiple hash functions in a single\n+   repository.\n+4. Taking the opportunity to fix other bugs in git's formats and\n+   protocols.\n+5. Shallow clones and fetches into a NEWHASH repository. (This will\n+   change when we add NEWHASH support to Git protocol.)\n+6. Skip fetching some submodules of a project into a NEWHASH\n+   repository. (This also depends on NEWHASH support in Git\n+   protocol.)\n+\n+Overview\n+--------\n+We introduce a new repository format extension `newhash`. Repositories\n+with this extension enabled use NEWHASH instead of SHA-1 to name\n+their objects. This affects both object names and object content ---\n+both the names of objects and all references to other objects within\n+an object are switched to the new hash function.\n+\n+newhash repositories cannot be read by older versions of Git.\n+\n+Alongside the packfile, a newhash repository stores a bidirectional\n+mapping between newhash and sha1 object names in a new format of .idx files.\n+The mapping is generated locally and can be verified using \"git fsck\".\n+Object lookups use this mapping to allow naming objects using either\n+their sha1 and newhash names interchangeably.\n+\n+\"git cat-file\" and \"git hash-object\" gain options to display an object\n+in its sha1 form and write an object given its sha1 form. This\n+requires all objects referenced by that object to be present in the\n+object database so that they can be named using the appropriate name\n+(using the bidirectional hash mapping).\n+\n+Fetches from a SHA-1 based server convert the fetched objects into\n+newhash form and record the mapping in the bidirectional mapping table\n+(see below for details). Pushes to a SHA-1 based server convert the\n+objects being pushed into sha1 form so the server does not have to be\n+aware of the hash function the client is using.\n+\n+Detailed Design\n+---------------\n+Object names\n+~~~~~~~~~~~~\n+Objects can be named by their 40 hexadecimal digit sha1-name or <n>\n+hexadecimal digit newhash-name, plus names derived from those (see\n+gitrevisions(7)).\n+\n+The sha1-name of an object is the SHA-1 of the concatenation of its\n+type, length, a nul byte, and the object's sha1-content. This is the\n+traditional <sha1> used in Git to name objects.\n+\n+The newhash-name of an object is the NEWHASH of the concatenation of its\n+type, length, a nul byte, and the object's newhash-content.\n+\n+Object format\n+~~~~~~~~~~~~~\n+The content as a byte sequence of a tag, commit, or tree object named\n+by sha1 and newhash differ because an object named by newhash-name refers to\n+other objects by their newhash-names and an object named by sha1-name\n+refers to other objects by their sha1-names.\n+\n+The newhash-content of an object is the same as its sha1-content, except\n+that objects referenced by the object are named using their newhash-names\n+instead of sha1-names. Because a blob object does not refer to any\n+other object, its sha1-content and newhash-content are the same.\n+\n+The format allows round-trip conversion between newhash-content and\n+sha1-content.\n+\n+Object storage\n+~~~~~~~~~~~~~~\n+Loose objects use zlib compression and packed objects use the packed\n+format described in Documentation/technical/pack-format.txt, just like\n+today. The content that is compressed and stored uses newhash-content\n+instead of sha1-content.\n+\n+Translation table\n+~~~~~~~~~~~~~~~~~\n+A fast bidirectional mapping between sha1-names and newhash-names of all\n+local objects in the repository is kept on disk.\n+\n+For pack files, upgrade the .idx file to be as follows:\n+\n+  4 magic bytes\n+  header, containing pointers to the 3 lists below\n+\n+  list of\n+  abbrev sha1 -> ordinal, sorted by sha1\n+\n+  list of\n+  abbrev newhash -> ordinal, sorted by newhash\n+\n+  list of\n+  ordinal, complete sha1, complete new hash,\n+  sorted by ordinal, such that a lookup can be computed after looking into\n+  one of the first lists.\n+\n+For unpacked objects, keep a simple list\n+  sha1 -> newhash\n+around at $OBJECT_DIR/loose-lookup\n+\n+All operations that make new objects (e.g., \"git commit\") add the new\n+objects to the translation table.\n+\n+(This work could have been deferred to push time, but that would\n+significantly complicate and slow down pushes. Calculating the\n+sha1-name at object creation time at the same time it is being\n+streamed to disk and having its newhash-name calculated should be an\n+acceptable cost.)\n+\n+Reading an object's sha1-content\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+The sha1-content of an object can be read by converting all newhash-names\n+its newhash-content references to sha1-names using the translation table.\n+\n+Fetch\n+~~~~~\n+Fetching from a SHA-1 based server requires translating between SHA-1\n+and NEWHASH based representations on the fly.\n+\n+SHA-1s named in the ref advertisement that are present on the client\n+can be translated to NEWHASH and looked up as local objects using the\n+translation table.\n+\n+Negotiation proceeds as today. Any \"have\"s generated locally are\n+converted to SHA-1 before being sent to the server, and SHA-1s\n+mentioned by the server are converted to NEWHASH when looking them up\n+locally.\n+\n+After negotiation, the server sends a packfile containing the\n+requested objects. We convert the packfile to NEWHASH format using\n+the following steps:\n+\n+1. index-pack: inflate each object in the packfile and compute its\n+   SHA-1. Objects can contain deltas in OBJ_REF_DELTA format against\n+   objects the client has locally. These objects can be looked up\n+   using the translation table and their sha1-content read as\n+   described above to resolve the deltas.\n+2. topological sort: starting at the \"want\"s from the negotiation\n+   phase, walk through objects in the pack and emit a list of them,\n+   excluding blobs, in reverse topologically sorted order, with each\n+   object coming later in the list than all objects it references.\n+   (This list only contains objects reachable from the \"wants\". If the\n+   pack from the server contained additional extraneous objects, then\n+   they will be discarded.)\n+3. convert to newhash: open a new (newhash) packfile. Read the topologically\n+   sorted list just generated. For each object, inflate its\n+   sha1-content, convert to newhash-content, and write it to the newhash\n+   pack. Include the new sha1<->newhash mapping entry in the translation\n+   table.\n+4. sort: reorder entries in the new pack to match the order of objects\n+   in the pack the server generated and include blobs. Write a newhash idx\n+   file.\n+5. clean up: remove the SHA-1 based pack file, index, and\n+   topologically sorted list obtained from the server and steps 1\n+   and 2.\n+\n+Step 3 requires every object referenced by the new object to be in the\n+translation table. This is why the topological sort step is necessary.\n+\n+As an optimization, step 1 could write a file describing what non-blob\n+objects each object it has inflated from the packfile references. This\n+makes the topological sort in step 2 possible without inflating the\n+objects in the packfile for a second time. The objects need to be\n+inflated again in step 3, for a total of two inflations.\n+\n+Step 4 is probably necessary for good read-time performance. \"git\n+pack-objects\" on the server optimizes the pack file for good data\n+locality (see Documentation/technical/pack-heuristics.txt).\n+\n+Details of this process are likely to change. It will take some\n+experimenting to get this to perform well.\n+\n+Push\n+~~~~\n+Push is simpler than fetch because the objects referenced by the\n+pushed objects are already in the translation table. The sha1-content\n+of each object being pushed can be read as described in the \"Reading\n+an object's sha1-content\" section to generate the pack written by git\n+send-pack.\n+\n+Signed Commits\n+~~~~~~~~~~~~~~\n+We add a new field \"gpgsig-newhash\" to the commit object format to allow\n+signing commits without relying on SHA-1. It is similar to the\n+existing \"gpgsig\" field. Its signed payload is the newhash-content of the\n+commit object with any \"gpgsig\" and \"gpgsig-newhash\" fields removed.\n+\n+This means commits can be signed\n+1. using SHA-1 only, as in existing signed commit objects\n+2. using both SHA-1 and NEWHASH, by using both gpgsig-newhash and gpgsig\n+   fields.\n+3. using only NEWHASH, by only using the gpgsig-newhash field.\n+\n+Old versions of \"git verify-commit\" can verify the gpgsig signature in\n+cases (1) and (2) without modifications and view case (3) as an\n+ordinary unsigned commit.\n+\n+Signed Tags\n+~~~~~~~~~~~\n+We add a new field \"gpgsig-newhash\" to the tag object format to allow\n+signing tags without relying on SHA-1. Its signed payload is the\n+newhash-content of the tag with its gpgsig-newhash field and \"-----BEGIN PGP\n+SIGNATURE-----\" delimited in-body signature removed.\n+\n+This means tags can be signed\n+1. using SHA-1 only, as in existing signed tag objects\n+2. using both SHA-1 and NEWHASH, by using gpgsig-newhash and an in-body\n+   signature.\n+3. using only NEWHASH, by only using the gpgsig-newhash field.\n+\n+Mergetag embedding\n+~~~~~~~~~~~~~~~~~~\n+The mergetag field in the sha1-content of a commit contains the\n+sha1-content of a tag that was merged by that commit.\n+\n+The mergetag field in the newhash-content of the same commit contains the\n+newhash-content of the same tag.\n+\n+Submodules\n+~~~~~~~~~~\n+To convert recorded submodule pointers, you need to have the converted\n+submodule repository in place. The translation table of the submodule\n+can be used to look up the new hash.\n+\n+Caveats\n+-------\n+Invalid objects\n+~~~~~~~~~~~~~~~\n+The conversion from sha1-content to newhash-content retains any\n+brokenness in the original object (e.g., tree entry modes encoded with\n+leading 0, tree objects whose paths are not sorted correctly, and\n+commit objects without an author or committer). This is a deliberate\n+feature of the design to allow the conversion to round-trip.\n+\n+More profoundly broken objects (e.g., a commit with a truncated \"tree\"\n+header line) cannot be converted but were not usable by current Git\n+anyway.\n+\n+Shallow clone and submodules\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+Because it requires all referenced objects to be available in the\n+locally generated translation table, this design does not support\n+shallow clone or unfetched submodules. Protocol improvements might\n+allow lifting this restriction.\n+\n+Alternates\n+~~~~~~~~~~\n+For the same reason, a newhash repository cannot borrow objects from a\n+sha1 repository using objects/info/alternates or\n+$GIT_ALTERNATE_OBJECT_REPOSITORIES.\n+\n+git notes\n+~~~~~~~~~\n+The \"git notes\" tool annotates objects using their sha1-name as key.\n+This design does not describe a way to migrate notes trees to use\n+newhash-names. That migration is expected to happen separately (for\n+example using a file at the root of the notes tree to describe which\n+hash it uses).\n+\n+Server-side cost\n+~~~~~~~~~~~~~~~~\n+Until Git protocol gains NEWHASH support, using newhash based storage on\n+public-facing Git servers is strongly discouraged. Once Git protocol\n+gains NEWHASH support, newhash based servers are likely not to support\n+sha1 compatibility, to avoid what may be a very expensive hash\n+reencode during clone and to encourage peers to modernize.\n+\n+The design described here allows fetches by SHA-1 clients of a\n+personal NEWHASH repository because it's not much more difficult than\n+allowing pushes from that repository. This support needs to be guarded\n+by a configuration option --- servers like git.kernel.org that serve a\n+large number of clients would not be expected to bear that cost.\n+\n+Meaning of signatures\n+~~~~~~~~~~~~~~~~~~~~~\n+The signed payload for signed commits and tags does not explicitly\n+name the hash used to identify objects. If some day Git adopts a new\n+hash function with the same length as the current SHA-1 (40\n+hexadecimal digit) or NEWHASH (64 hexadecimal digit) objects then the\n+intent behind the PGP signed payload in an object signature is\n+unclear:\n+\n+\tobject e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7\n+\ttype commit\n+\ttag v2.12.0\n+\ttagger Junio C Hamano <gitster@pobox.com> 1487962205 -0800\n+\n+\tGit 2.12\n+\n+Does this mean Git v2.12.0 is the commit with sha1-name\n+e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7 or the commit with\n+new-40-digit-hash-name e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7?\n+\n+Fortunately NEWHASH and SHA-1 have different lengths. If Git starts\n+using another hash with the same length to name objects, then it will\n+need to change the format of signed payloads using that hash to\n+address this issue.\n+\n+Alternatives considered\n+-----------------------\n+Upgrading everyone working on a particular project on a flag day\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+Projects like the Linux kernel are large and complex enough that\n+flipping the switch for all projects based on the repository at once\n+is infeasible.\n+\n+Not only would all developers and server operators supporting\n+developers have to switch on the same flag day, but supporting tooling\n+(continuous integration, code review, bug trackers, etc) would have to\n+be adapted as well. This also makes it difficult to get early feedback\n+from some project participants testing before it is time for mass\n+adoption.\n+\n+Using hash functions in parallel\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+(e.g. https://public-inbox.org/git/22708.8913.864049.452252@chiark.greenend.org.uk/ )\n+Objects newly created would be addressed by the new hash, but inside\n+such an object (e.g. commit) it is still possible to address objects\n+using the old hash function.\n+* You cannot trust its history (needed for bisectability) in the\n+  future without further work\n+* Maintenance burden as the number of supported hash functions grows\n+  (they will never go away, so they accumulate). In this proposal, by\n+  comparison, converted objects lose all references to SHA-1.\n+\n+Signed objects with multiple hashes\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+Instead of introducing the gpgsig-newhash field in commit and tag objects\n+for newhash-content based signatures, an earlier version of this design\n+added \"hash newhash <newhash-name>\" fields to strengthen the existing\n+sha1-content based signatures.\n+\n+In other words, a single signature was used to attest to the object\n+content using both hash functions. This had some advantages:\n+* Using one signature instead of two speeds up the signing process.\n+* Having one signed payload with both hashes allows the signer to\n+  attest to the sha1-name and newhash-name referring to the same object.\n+* All users consume the same signature. Broken signatures are likely\n+  to be detected quickly using current versions of git.\n+\n+However, it also came with disadvantages:\n+* Verifying a signed object requires access to the sha1-names of all\n+  objects it references, even after the transition is complete and\n+  translation table is no longer needed for anything else. To support\n+  this, the design added fields such as \"hash sha1 tree <sha1-name>\"\n+  and \"hash sha1 parent <sha1-name>\" to the newhash-content of a signed\n+  commit, complicating the conversion process.\n+* Allowing signed objects without a sha1 (for after the transition is\n+  complete) complicated the design further, requiring a \"nohash sha1\"\n+  field to suppress including \"hash sha1\" fields in the newhash-content\n+  and signed payload.\n+\n+\n+Implementation plan\n+-------------------\n+\n+Here's a rough list of some useful tasks, in no particular order:\n+\n+1. bc/object-id: This patch series continues, eliminating assumptions\n+   about the size of object ids by encapsulating them in a struct.\n+   One straightforward way to find code that still needs to be\n+   converted is to grep for \"sha\" --- often the conversion patches\n+   change function and variable names to refer to oid_ where they used\n+   to use sha1_, making the stragglers easier to spot.\n+\n+2. Hard-coded object ids in tests: Many tests beyond t00* make assumptions\n+   about the exact values of object ids.  That's bad for maintainability\n+   for other reasons beyond the hash function transition, too.\n+\n+   It should be possible to suss them out by patching git's sha1\n+   routine to use the ones-complement of sha1 (~sha1) instead and\n+   seeing which tests fail.\n+\n+3. Repository format extension to use a different hash function: we\n+   want git to be able to work with two hash functions: sha1 and\n+   something else.  For interoperability and simplity, it is useful\n+   for a single git binary to support both hash functions.\n+\n+   That means a repository needs to be able to specify what hash\n+   function is used for the objects in that repository.  This can be\n+   configured by setting '[core] repositoryformatversion=1' (to avoid\n+   confusing old versions of git) and\n+   '[extensions] experimentalNewHashFunction = true'.\n+   Documentation/technical/repository-version.txt has more details.\n+\n+   We can start experimenting with this using e.g. the ~sha1 function\n+   described at (2), or the 160-bit hash of the patch author's choice\n+   (e.g. truncated blake2bp-256).\n+\n+4. When choosing a hash function, people may argue about performance.\n+   It would be useful for run some benchmarks for git (running\n+   the test suite, t/perf tests, etc) using a variety of hash\n+   functions as input to such a discussion.\n+\n+5. Longer hash: Even once all object id references in git use struct\n+   object_id (see (1)), we need to tackle other assumptions about\n+   object id size in git and its tests.\n+\n+   It should be possible to suss them out by replacing git's sha1\n+   routine with a 40-byte hash: sha1 with each byte repeated (sha1+sha1)\n+   and seeing what fails.\n+\n+6. Repository format extension for longer hash: As in (3), we could\n+   add a repository format extension to experiment with using the\n+   sha1+sha1 function.\n+\n+7. Avoiding wasted memory from unused hash functions: struct object_id\n+   has definition 'unsigned char hash[GIT_MAX_RAWSZ]', where\n+   GIT_MAX_RAWSZ is the size of the largest supported hash function.\n+   When operating on a repository that only uses sha1, this wastes\n+   memory.\n+\n+   Avoid that by making object identifiers variable-sized.  That is,\n+   something like\n+\n+     struct object_id {\n+        union {\n+           unsigned char hash20[20];\n+           unsigned char hash32[32];\n+        } *hash;\n+     }\n+\n+   or\n+\n+     struct object_id {\n+       unsigned char *hash;\n+     }\n+\n+   The hard part is that allocation and destruction have to be\n+   explicit instead of happening automatically when an object_id is an\n+   automatic variable.\n+\n+8. Implementation of this plan (roughly in order):\n+   - abstract the hash computation to be able to plug in another hash\n+   - make the choice of hash dependant on repository extension\n+   - implement the new .idx format\n+   - implement cat-file's flag to show things in old/new hash\n+   - convert fetch, push\n+\n+9. We can use help from security experts in all of this.  Fuzzing,\n+   analysis of how we use cryptography, security review of other parts\n+   of the design, and information to help choose a hash function are\n+   all appreciated.\n+\n+Agreed-upon criteria for selecting NewHash\n+------------------------------------------\n+\n+The discussion which hash function to use is going in circles, so let's\n+first argree on criteria on how to select the new hash function. These\n+could include:\n+* cryptografic strength\n+* performance\n+* other cryptografic aspects(?)\n+* portability / availability of properly licensed implementations\n+\n+Future work\n+-----------\n+\n+* other compression instead of zlib (this is a stated non goal, though!)\n+* rehash discussion whether to include generation numbers natively\n+  (this is a stated non goal, though!)\n+* describing (1) the possibility of caching translated objects\n+* and (2) protocol changes.\n+* other format changes\n+\n+Document History\n+----------------\n+\n+2017-03-03\n+bmwill@google.com, jonathantanmy@google.com, jrnieder@gmail.com,\n+sbeller@google.com\n+\n+Initial version sent to\n+http://public-inbox.org/git/20170304011251.GA26789@aiede.mtv.corp.google.com\n+\n+2017-03-03 jrnieder@gmail.com\n+Incorporated suggestions from jonathantanmy and sbeller:\n+* describe purpose of signed objects with each hash type\n+* redefine signed object verification using object content under the\n+  first hash function\n+\n+2017-03-06 jrnieder@gmail.com\n+* Use SHA3-256 instead of SHA2 (thanks, Linus and brian m. carlson).[1][2]\n+* Make sha3-based signatures a separate field, avoiding the need for\n+  \"hash\" and \"nohash\" fields (thanks to peff[3]).\n+* Add a sorting phase to fetch (thanks to Junio for noticing the need\n+  for this).\n+* Omit blobs from the topological sort during fetch (thanks to peff).\n+* Discuss alternates, git notes, and git servers in the caveats\n+  section (thanks to Junio Hamano, brian m. carlson[4], and Shawn\n+  Pearce).\n+* Clarify language throughout (thanks to various commenters,\n+  especially Junio).\n+\n+[1] http://public-inbox.org/git/CA+55aFzJtejiCjV0e43+9oR3QuJK2PiFiLQemytoLpyJWe6P9w@mail.gmail.com/\n+[2] http://public-inbox.org/git/CA+55aFz+gkAsDZ24zmePQuEs1XPS9BP_s8O7Q4wQ7LV7X5-oDA@mail.gmail.com/\n+[3] http://public-inbox.org/git/20170306084353.nrns455dvkdsfgo5@sigill.intra.peff.net/\n+[4] http://public-inbox.org/git/20170304224936.rqqtkdvfjgyezsht@genre.crustytoothpaste.net\n+\n+2017-09-25\n+* replaced SHA3-256 with NEWHASH, sha3 with newhash\n+* added section 'Implementation plan'\n+* added section 'Future work'\n+\n+* This version is sent to the list; to be incorporated into git.git, such\n+  that further document history is found using git-log.\n+\n+\n-- \n2.14.0.rc0.3.g6c2e499285\n\n"},{"id":"328974","messageId":"20170926233827.GC19555@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170926222558.22418-1-sbeller@google.com","subject":"Re: [PATCH] technical doc: add a design doc for hash function transition","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-26T23:38:27Z","receivedAt":"2017-09-26T23:38:36Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nStefan Beller wrote:\n\n> From: Jonathan Nieder <jrn@google.com>\n\nI go by jrnieder@gmail.com upstream. :)\n\n> This is \"RFC v3: Another proposed hash function transition plan\" from\n> the git mailing list.\n>\n> Signed-off-by: Jonathan Nieder <jrnieder@gmail.com>\n> Signed-off-by: Jonathan Tan <jonathantanmy@google.com>\n> Signed-off-by: Brandon Williams <bmwill@google.com>\n> Signed-off-by: Stefan Beller <sbeller@google.com>\n\nI hadn't signed-off on this version, but it's not a big deal.\n\n[...]\n> ---\n>\n>  This takes the original Google Doc[1] and adds it to our history,\n>  such that the discussion can be on on list and in the commit messages.\n>\n>  * replaced SHA3-256 with NEWHASH, sha3 with newhash\n>  * added section 'Implementation plan'\n>  * added section 'Future work'\n>  * added section 'Agreed-upon criteria for selecting NewHash'\n\nThanks for sending this out.  I had let it stall too long.\n\nAs a tiny nit, I think NewHash is easier to read than NEWHASH.  Not a\nbig deal.  More importantly, we need some text describing it and\nsaying it's a placeholder.\n\nThe implementation plan included here is out of date.  It comes from\nan email where I was answering a question about what people can do to\nmake progress, before this design had been agreed on.  In the context\nof this design there are other steps we'd want to describe (having to\ndo with implementing the translation table, etc).\n\nI also planned to add a description of the translation table based on\nwhat was discussed previously in this thread.\n\nJonathan\n"},{"id":"328975","messageId":"20170926235158.GD19555@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709262356360.40514@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-26T23:51:58Z","receivedAt":"2017-09-26T23:52:07Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nJohannes Schindelin wrote:\n\n> Sorry, you are asking cryptography experts to spend their time on the Git\n> mailing list. I tried to get them to speak out on the Git mailing list.\n> They respectfully declined.\n>\n> I can't fault them, they have real jobs to do, and none of their managers\n> would be happy for them to educate the Git mailing list on matters of\n> cryptography, not after what happened in 2005.\n\nFortunately we have had a few public comments from crypto specialists:\n\nhttps://public-inbox.org/git/91a34c5b-7844-3db2-cf29-411df5bcf886@noekeon.org/\nhttps://public-inbox.org/git/CAL9PXLzhPyE+geUdcLmd=pidT5P8eFEBbSgX_dS88knz2q_LSw@mail.gmail.com/\nhttps://public-inbox.org/git/CAL9PXLxMHG1nP5_GQaK_WSJTNKs=_qbaL6V5v2GzVG=9VU2+gA@mail.gmail.com/\nhttps://public-inbox.org/git/59BFB95D.1030903@st.com/\nhttps://public-inbox.org/git/59C149A3.6080506@st.com/\n\n[...]\n> Let's be realistic. Git is pretty important to us, but it is not important\n> enough to sway, say, Intel into announcing hardware support for SHA3.\n\nYes, I agree with this.  (Adoption by Git could lead to adoption by\nsome other projects, leading to more work on high quality software\nimplementations in projects like OpenSSL, but I am not convinced that\nthat would be a good thing for the world anyway.  There are downsides\nto a proliferation of too many crypto primitives.  This is the basic\nargument described in more detail at [1].)\n\n[...]\n> On Tue, 26 Sep 2017, Jason Cooper wrote:\n\n>> For my use cases, as a user of git, I have a plan to maintain provable\n>> integrity of existing objects stored in git under sha1 while migrating\n>> away from sha1.  The same plan works for migrating away from SHA2 or\n>> SHA3 when the time comes.\n>\n> Please do not make the mistake of taking your use case to be a template\n> for everybody's use case.\n\nThat said, I'm curious at what plan you are alluding to.  Is it\nsomething that could benefit others on the list?\n\nThanks,\nJonathan\n\n[1] https://www.imperialviolet.org/2017/05/31/skipsha3.html\n"},{"id":"329092","messageId":"20170928044320.GA84719@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"CAJo=hJtoX9=AyLHHpUJS7fueV9ciZ_MNpnEPHUz8Whui6g9F0A@mail.gmail.com","subject":"[PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-28T04:43:21Z","receivedAt":"2017-09-28T04:43:57Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"This document describes what a transition to a new hash function for\nGit would look like.  Add it to Documentation/technical/ as the plan\nof record so that future changes can be recorded as patches.\n\nAlso-by: Brandon Williams <bmwill@google.com>\nAlso-by: Jonathan Tan <jonathantanmy@google.com>\nAlso-by: Stefan Beller <sbeller@google.com>\nSigned-off-by: Jonathan Nieder <jrnieder@gmail.com>\n---\nOn Thu, Mar 09, 2017 at 11:14 AM, Shawn Pearce wrote:\n> On Mon, Mar 6, 2017 at 4:17 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n\n>> Thanks for the kind words on what had quite a few flaws still.  Here's\n>> a new draft.  I think the next version will be a patch against\n>> Documentation/technical/.\n>\n> FWIW, I like this approach.\n\nOkay, here goes.\n\nInstead of sharding the loose object translation tables by first byte,\nwe went for a single table.  It simplifies the design and we need to\nkeep the number of loose objects under control anyway.\n\nWe also included a description of the transition plan and tried to\ninclude a summary of what has been agreed upon so far about the choice\nof hash function.\n\nThanks to Junio for reviving the discussion and in particular to Dscho\nfor pushing this forward and making the missing pieces clearer.\n\nThoughts of all kinds welcome, as always.\n\n Documentation/Makefile                             |   1 +\n .../technical/hash-function-transition.txt         | 797 +++++++++++++++++++++\n 2 files changed, 798 insertions(+)\n create mode 100644 Documentation/technical/hash-function-transition.txt\n\ndiff --git a/Documentation/Makefile b/Documentation/Makefile\nindex 2415e0d657..471bb29725 100644\n--- a/Documentation/Makefile\n+++ b/Documentation/Makefile\n@@ -67,6 +67,7 @@ SP_ARTICLES += howto/maintain-git\n API_DOCS = $(patsubst %.txt,%,$(filter-out technical/api-index-skel.txt technical/api-index.txt, $(wildcard technical/api-*.txt)))\n SP_ARTICLES += $(API_DOCS)\n \n+TECH_DOCS += technical/hash-function-transition\n TECH_DOCS += technical/http-protocol\n TECH_DOCS += technical/index-format\n TECH_DOCS += technical/pack-format\ndiff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\nnew file mode 100644\nindex 0000000000..417ba491d0\n--- /dev/null\n+++ b/Documentation/technical/hash-function-transition.txt\n@@ -0,0 +1,797 @@\n+Git hash function transition\n+============================\n+\n+Objective\n+---------\n+Migrate Git from SHA-1 to a stronger hash function.\n+\n+Background\n+----------\n+At its core, the Git version control system is a content addressable\n+filesystem. It uses the SHA-1 hash function to name content. For\n+example, files, directories, and revisions are referred to by hash\n+values unlike in other traditional version control systems where files\n+or versions are referred to via sequential numbers. The use of a hash\n+function to address its content delivers a few advantages:\n+\n+* Integrity checking is easy. Bit flips, for example, are easily\n+  detected, as the hash of corrupted content does not match its name.\n+* Lookup of objects is fast.\n+\n+Using a cryptographically secure hash function brings additional\n+advantages:\n+\n+* Object names can be signed and third parties can trust the hash to\n+  address the signed object and all objects it references.\n+* Communication using Git protocol and out of band communication\n+  methods have a short reliable string that can be used to reliably\n+  address stored content.\n+\n+Over time some flaws in SHA-1 have been discovered by security\n+researchers. https://shattered.io demonstrated a practical SHA-1 hash\n+collision. As a result, SHA-1 cannot be considered cryptographically\n+secure any more. This impacts the communication of hash values because\n+we cannot trust that a given hash value represents the known good\n+version of content that the speaker intended.\n+\n+SHA-1 still possesses the other properties such as fast object lookup\n+and safe error checking, but other hash functions are equally suitable\n+that are believed to be cryptographically secure.\n+\n+Goals\n+-----\n+Where NewHash is a strong 256-bit hash function to replace SHA-1 (see\n+\"Selection of a New Hash\", below):\n+\n+1. The transition to NewHash can be done one local repository at a time.\n+   a. Requiring no action by any other party.\n+   b. A NewHash repository can communicate with SHA-1 Git servers\n+      (push/fetch).\n+   c. Users can use SHA-1 and NewHash identifiers for objects\n+      interchangeably (see \"Object names on the command line\", below).\n+   d. New signed objects make use of a stronger hash function than\n+      SHA-1 for their security guarantees.\n+2. Allow a complete transition away from SHA-1.\n+   a. Local metadata for SHA-1 compatibility can be removed from a\n+      repository if compatibility with SHA-1 is no longer needed.\n+3. Maintainability throughout the process.\n+   a. The object format is kept simple and consistent.\n+   b. Creation of a generalized repository conversion tool.\n+\n+Non-Goals\n+---------\n+1. Add NewHash support to Git protocol. This is valuable and the\n+   logical next step but it is out of scope for this initial design.\n+2. Transparently improving the security of existing SHA-1 signed\n+   objects.\n+3. Intermixing objects using multiple hash functions in a single\n+   repository.\n+4. Taking the opportunity to fix other bugs in Git's formats and\n+   protocols.\n+5. Shallow clones and fetches into a NewHash repository. (This will\n+   change when we add NewHash support to Git protocol.)\n+6. Skip fetching some submodules of a project into a NewHash\n+   repository. (This also depends on NewHash support in Git\n+   protocol.)\n+\n+Overview\n+--------\n+We introduce a new repository format extension. Repositories with this\n+extension enabled use NewHash instead of SHA-1 to name their objects.\n+This affects both object names and object content --- both the names\n+of objects and all references to other objects within an object are\n+switched to the new hash function.\n+\n+NewHash repositories cannot be read by older versions of Git.\n+\n+Alongside the packfile, a NewHash repository stores a bidirectional\n+mapping between NewHash and SHA-1 object names. The mapping is generated\n+locally and can be verified using \"git fsck\". Object lookups use this\n+mapping to allow naming objects using either their SHA-1 and NewHash names\n+interchangeably.\n+\n+\"git cat-file\" and \"git hash-object\" gain options to display an object\n+in its sha1 form and write an object given its sha1 form. This\n+requires all objects referenced by that object to be present in the\n+object database so that they can be named using the appropriate name\n+(using the bidirectional hash mapping).\n+\n+Fetches from a SHA-1 based server convert the fetched objects into\n+NewHash form and record the mapping in the bidirectional mapping table\n+(see below for details). Pushes to a SHA-1 based server convert the\n+objects being pushed into sha1 form so the server does not have to be\n+aware of the hash function the client is using.\n+\n+Detailed Design\n+---------------\n+Repository format extension\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+A NewHash repository uses repository format version `1` (see\n+Documentation/technical/repository-version.txt) with extensions\n+`objectFormat` and `compatObjectFormat`:\n+\n+\t[core]\n+\t\trepositoryFormatVersion = 1\n+\t[extensions]\n+\t\tobjectFormat = newhash\n+\t\tcompatObjectFormat = sha1\n+\n+Specifying a repository format extension ensures that versions of Git\n+not aware of NewHash do not try to operate on these repositories,\n+instead producing an error message:\n+\n+\t$ git status\n+\tfatal: unknown repository extensions found:\n+\t\tobjectformat\n+\t\tcompatobjectformat\n+\n+See the \"Transition plan\" section below for more details on these\n+repository extensions.\n+\n+Object names\n+~~~~~~~~~~~~\n+Objects can be named by their 40 hexadecimal digit sha1-name or 64\n+hexadecimal digit newhash-name, plus names derived from those (see\n+gitrevisions(7)).\n+\n+The sha1-name of an object is the SHA-1 of the concatenation of its\n+type, length, a nul byte, and the object's sha1-content. This is the\n+traditional <sha1> used in Git to name objects.\n+\n+The newhash-name of an object is the NewHash of the concatenation of its\n+type, length, a nul byte, and the object's newhash-content.\n+\n+Object format\n+~~~~~~~~~~~~~\n+The content as a byte sequence of a tag, commit, or tree object named\n+by sha1 and newhash differ because an object named by newhash-name refers to\n+other objects by their newhash-names and an object named by sha1-name\n+refers to other objects by their sha1-names.\n+\n+The newhash-content of an object is the same as its sha1-content, except\n+that objects referenced by the object are named using their newhash-names\n+instead of sha1-names. Because a blob object does not refer to any\n+other object, its sha1-content and newhash-content are the same.\n+\n+The format allows round-trip conversion between newhash-content and\n+sha1-content.\n+\n+Object storage\n+~~~~~~~~~~~~~~\n+Loose objects use zlib compression and packed objects use the packed\n+format described in Documentation/technical/pack-format.txt, just like\n+today. The content that is compressed and stored uses newhash-content\n+instead of sha1-content.\n+\n+Pack index\n+~~~~~~~~~~\n+Pack index (.idx) files use a new v3 format that supports multiple\n+hash functions. They have the following format (all integers are in\n+network byte order):\n+\n+- A header appears at the beginning and consists of the following:\n+  - The 4-byte pack index signature: '\\377t0c'\n+  - 4-byte version number: 3\n+  - 4-byte length of the header section, including the signature and\n+    version number\n+  - 4-byte number of objects contained in the pack\n+  - 4-byte number of object formats in this pack index: 2\n+  - For each object format:\n+    - 4-byte format identifier (e.g., 'sha1' for SHA-1)\n+    - 4-byte length in bytes of shortened object names. This is the\n+      shortest possible length needed to make names in the shortened\n+      object name table unambiguous.\n+    - 4-byte integer, recording where tables relating to this format\n+      are stored in this index file, as an offset from the beginning.\n+  - 4-byte offset to the trailer from the beginning of this file.\n+  - Zero or more additional key/value pairs (4-byte key, 4-byte\n+    value). Only one key is supported: 'PSRC'. See the \"Loose objects\n+    and unreachable objects\" section for supported values and how this\n+    is used.  All other keys are reserved. Readers must ignore\n+    unrecognized keys.\n+- Zero or more NUL bytes. This can optionally be used to improve the\n+  alignment of the full object name table below.\n+- Tables for the first object format:\n+  - A sorted table of shortened object names.  These are prefixes of\n+    the names of all objects in this pack file, packed together\n+    without offset values to reduce the cache footprint of the binary\n+    search for a specific object name.\n+\n+  - A table of full object names in pack order. This allows resolving\n+    a reference to \"the nth object in the pack file\" (from a\n+    reachability bitmap or from the next table of another object\n+    format) to its object name.\n+\n+  - A table of 4-byte values mapping object name order to pack order.\n+    For an object in the table of sorted shortened object names, the\n+    value at the corresponding index in this table is the index in the\n+    previous table for that same object.\n+\n+    This can be used to look up the object in reachability bitmaps or\n+    to look up its name in another object format.\n+\n+  - A table of 4-byte CRC32 values of the packed object data, in the\n+    order that the objects appear in the pack file. This is to allow\n+    compressed data to be copied directly from pack to pack during\n+    repacking without undetected data corruption.\n+\n+  - A table of 4-byte offset values. For an object in the table of\n+    sorted shortened object names, the value at the corresponding\n+    index in this table indicates where that object can be found in\n+    the pack file. These are usually 31-bit pack file offsets, but\n+    large offsets are encoded as an index into the next table with the\n+    most significant bit set.\n+\n+  - A table of 8-byte offset entries (empty for pack files less than\n+    2 GiB). Pack files are organized with heavily used objects toward\n+    the front, so most object references should not need to refer to\n+    this table.\n+- Zero or more NUL bytes.\n+- Tables for the second object format, with the same layout as above,\n+  up to and not including the table of CRC32 values.\n+- Zero or more NUL bytes.\n+- The trailer consists of the following:\n+  - A copy of the 20-byte NewHash checksum at the end of the\n+    corresponding packfile.\n+\n+  - 20-byte NewHash checksum of all of the above.\n+\n+Loose object index\n+~~~~~~~~~~~~~~~~~~\n+A new file $GIT_OBJECT_DIR/loose-object-idx contains information about\n+all loose objects. Its format is\n+\n+  # loose-object-idx\n+  (newhash-name SP sha1-name LF)*\n+\n+where the object names are in hexadecimal format. The file is not\n+sorted.\n+\n+The loose object index is protected against concurrent writes by a\n+lock file $GIT_OBJECT_DIR/loose-object-idx.lock. To add a new loose\n+object:\n+\n+1. Write the loose object to a temporary file, like today.\n+2. Open loose-object-idx.lock with O_CREAT | O_EXCL to acquire the lock.\n+3. Rename the loose object into place.\n+4. Open loose-object-idx with O_APPEND and write the new object\n+5. Unlink loose-object-idx.lock to release the lock.\n+\n+To remove entries (e.g. in \"git pack-refs\" or \"git-prune\"):\n+\n+1. Open loose-object-idx.lock with O_CREAT | O_EXCL to acquire the\n+   lock.\n+2. Write the new content to loose-object-idx.lock.\n+3. Unlink any loose objects being removed.\n+4. Rename to replace loose-object-idx, releasing the lock.\n+\n+Translation table\n+~~~~~~~~~~~~~~~~~\n+The index files support a bidirectional mapping between sha1-names\n+and newhash-names. The lookup proceeds similarly to ordinary object\n+lookups. For example, to convert a sha1-name to a newhash-name:\n+\n+ 1. Look for the object in idx files. If a match is present in the\n+    idx's sorted list of truncated sha1-names, then:\n+    a. Read the corresponding entry in the sha1-name order to pack\n+       name order mapping.\n+    b. Read the corresponding entry in the full sha1-name table to\n+       verify we found the right object. If it is, then\n+    c. Read the corresponding entry in the full newhash-name table.\n+       That is the object's newhash-name.\n+ 2. Check for a loose object. Read lines from loose-object-idx until\n+    we find a match.\n+\n+Step (1) takes the same amount of time as an ordinary object lookup:\n+O(number of packs * log(objects per pack)). Step (2) takes O(number of\n+loose objects) time. To maintain good performance it will be necessary\n+to keep the number of loose objects low. See the \"Loose objects and\n+unreachable objects\" section below for more details.\n+\n+Since all operations that make new objects (e.g., \"git commit\") add\n+the new objects to the corresponding index, this mapping is possible\n+for all objects in the object store.\n+\n+Reading an object's sha1-content\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+The sha1-content of an object can be read by converting all newhash-names\n+its newhash-content references to sha1-names using the translation table.\n+\n+Fetch\n+~~~~~\n+Fetching from a SHA-1 based server requires translating between SHA-1\n+and NewHash based representations on the fly.\n+\n+SHA-1s named in the ref advertisement that are present on the client\n+can be translated to NewHash and looked up as local objects using the\n+translation table.\n+\n+Negotiation proceeds as today. Any \"have\"s generated locally are\n+converted to SHA-1 before being sent to the server, and SHA-1s\n+mentioned by the server are converted to NewHash when looking them up\n+locally.\n+\n+After negotiation, the server sends a packfile containing the\n+requested objects. We convert the packfile to NewHash format using\n+the following steps:\n+\n+1. index-pack: inflate each object in the packfile and compute its\n+   SHA-1. Objects can contain deltas in OBJ_REF_DELTA format against\n+   objects the client has locally. These objects can be looked up\n+   using the translation table and their sha1-content read as\n+   described above to resolve the deltas.\n+2. topological sort: starting at the \"want\"s from the negotiation\n+   phase, walk through objects in the pack and emit a list of them,\n+   excluding blobs, in reverse topologically sorted order, with each\n+   object coming later in the list than all objects it references.\n+   (This list only contains objects reachable from the \"wants\". If the\n+   pack from the server contained additional extraneous objects, then\n+   they will be discarded.)\n+3. convert to newhash: open a new (newhash) packfile. Read the topologically\n+   sorted list just generated. For each object, inflate its\n+   sha1-content, convert to newhash-content, and write it to the newhash\n+   pack. Record the new sha1<->newhash mapping entry for use in the idx.\n+4. sort: reorder entries in the new pack to match the order of objects\n+   in the pack the server generated and include blobs. Write a newhash idx\n+   file\n+5. clean up: remove the SHA-1 based pack file, index, and\n+   topologically sorted list obtained from the server in steps 1\n+   and 2.\n+\n+Step 3 requires every object referenced by the new object to be in the\n+translation table. This is why the topological sort step is necessary.\n+\n+As an optimization, step 1 could write a file describing what non-blob\n+objects each object it has inflated from the packfile references. This\n+makes the topological sort in step 2 possible without inflating the\n+objects in the packfile for a second time. The objects need to be\n+inflated again in step 3, for a total of two inflations.\n+\n+Step 4 is probably necessary for good read-time performance. \"git\n+pack-objects\" on the server optimizes the pack file for good data\n+locality (see Documentation/technical/pack-heuristics.txt).\n+\n+Details of this process are likely to change. It will take some\n+experimenting to get this to perform well.\n+\n+Push\n+~~~~\n+Push is simpler than fetch because the objects referenced by the\n+pushed objects are already in the translation table. The sha1-content\n+of each object being pushed can be read as described in the \"Reading\n+an object's sha1-content\" section to generate the pack written by git\n+send-pack.\n+\n+Signed Commits\n+~~~~~~~~~~~~~~\n+We add a new field \"gpgsig-newhash\" to the commit object format to allow\n+signing commits without relying on SHA-1. It is similar to the\n+existing \"gpgsig\" field. Its signed payload is the newhash-content of the\n+commit object with any \"gpgsig\" and \"gpgsig-newhash\" fields removed.\n+\n+This means commits can be signed\n+1. using SHA-1 only, as in existing signed commit objects\n+2. using both SHA-1 and NewHash, by using both gpgsig-newhash and gpgsig\n+   fields.\n+3. using only NewHash, by only using the gpgsig-newhash field.\n+\n+Old versions of \"git verify-commit\" can verify the gpgsig signature in\n+cases (1) and (2) without modifications and view case (3) as an\n+ordinary unsigned commit.\n+\n+Signed Tags\n+~~~~~~~~~~~\n+We add a new field \"gpgsig-newhash\" to the tag object format to allow\n+signing tags without relying on SHA-1. Its signed payload is the\n+newhash-content of the tag with its gpgsig-newhash field and \"-----BEGIN PGP\n+SIGNATURE-----\" delimited in-body signature removed.\n+\n+This means tags can be signed\n+1. using SHA-1 only, as in existing signed tag objects\n+2. using both SHA-1 and NewHash, by using gpgsig-newhash and an in-body\n+   signature.\n+3. using only NewHash, by only using the gpgsig-newhash field.\n+\n+Mergetag embedding\n+~~~~~~~~~~~~~~~~~~\n+The mergetag field in the sha1-content of a commit contains the\n+sha1-content of a tag that was merged by that commit.\n+\n+The mergetag field in the newhash-content of the same commit contains the\n+newhash-content of the same tag.\n+\n+Submodules\n+~~~~~~~~~~\n+To convert recorded submodule pointers, you need to have the converted\n+submodule repository in place. The translation table of the submodule\n+can be used to look up the new hash.\n+\n+Loose objects and unreachable objects\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+Fast lookups in the loose-object-idx require that the number of loose\n+objects not grow too high.\n+\n+\"git gc --auto\" currently waits for there to be 6700 loose objects\n+present before consolidating them into a packfile. We will need to\n+measure to find a more appropriate threshold for it to use.\n+\n+\"git gc --auto\" currently waits for there to be 50 packs present\n+before combining packfiles. Packing loose objects more aggressively\n+may cause the number of pack files to grow too quickly. This can be\n+mitigated by using a strategy similar to Martin Fick's exponential\n+rolling garbage collection script:\n+https://gerrit-review.googlesource.com/c/gerrit/+/35215\n+\n+\"git gc\" currently expels any unreachable objects it encounters in\n+pack files to loose objects in an attempt to prevent a race when\n+pruning them (in case another process is simultaneously writing a new\n+object that refers to the about-to-be-deleted object). This leads to\n+an explosion in the number of loose objects present and disk space\n+usage due to the objects in delta form being replaced with independent\n+loose objects.  Worse, the race is still present for loose objects.\n+\n+Instead, \"git gc\" will need to move unreachable objects to a new\n+packfile marked as UNREACHABLE_GARBAGE (using the PSRC field; see\n+below). To avoid the race when writing new objects referring to an\n+about-to-be-deleted object, code paths that write new objects will\n+need to copy any objects from UNREACHABLE_GARBAGE packs that they\n+refer to to new, non-UNREACHABLE_GARBAGE packs (or loose objects).\n+UNREACHABLE_GARBAGE are then safe to delete if their creation time (as\n+indicated by the file's mtime) is long enough ago.\n+\n+To avoid a proliferation of UNREACHABLE_GARBAGE packs, they can be\n+combined under certain circumstances. If \"gc.garbageTtl\" is set to\n+greater than one day, then packs created within a single calendar day,\n+UTC, can be coalesced together. The resulting packfile would have an\n+mtime before midnight on that day, so this makes the effective maximum\n+ttl the garbageTtl + 1 day. If \"gc.garbageTtl\" is less than one day,\n+then we divide the calendar day into intervals one-third of that ttl\n+in duration. Packs created within the same interval can be coalesced\n+together. The resulting packfile would have an mtime before the end of\n+the interval, so this makes the effective maximum ttl equal to the\n+garbageTtl * 4/3.\n+\n+This rule comes from Thirumala Reddy Mutchukota's JGit change\n+https://git.eclipse.org/r/90465.\n+\n+The UNREACHABLE_GARBAGE setting goes in the PSRC field of the pack\n+index. More generally, that field indicates where a pack came from:\n+\n+ - 1 (PACK_SOURCE_RECEIVE) for a pack received over the network\n+ - 2 (PACK_SOURCE_AUTO) for a pack created by a lightweight\n+   \"gc --auto\" operation\n+ - 3 (PACK_SOURCE_GC) for a pack created by a full gc\n+ - 4 (PACK_SOURCE_UNREACHABLE_GARBAGE) for potential garbage\n+   discovered by gc\n+ - 5 (PACK_SOURCE_INSERT) for locally created objects that were\n+   written directly to a pack file, e.g. from \"git add .\"\n+\n+This information can be useful for debugging and for \"gc --auto\" to\n+make appropriate choices about which packs to coalesce.\n+\n+Caveats\n+-------\n+Invalid objects\n+~~~~~~~~~~~~~~~\n+The conversion from sha1-content to newhash-content retains any\n+brokenness in the original object (e.g., tree entry modes encoded with\n+leading 0, tree objects whose paths are not sorted correctly, and\n+commit objects without an author or committer). This is a deliberate\n+feature of the design to allow the conversion to round-trip.\n+\n+More profoundly broken objects (e.g., a commit with a truncated \"tree\"\n+header line) cannot be converted but were not usable by current Git\n+anyway.\n+\n+Shallow clone and submodules\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+Because it requires all referenced objects to be available in the\n+locally generated translation table, this design does not support\n+shallow clone or unfetched submodules. Protocol improvements might\n+allow lifting this restriction.\n+\n+Alternates\n+~~~~~~~~~~\n+For the same reason, a newhash repository cannot borrow objects from a\n+sha1 repository using objects/info/alternates or\n+$GIT_ALTERNATE_OBJECT_REPOSITORIES.\n+\n+git notes\n+~~~~~~~~~\n+The \"git notes\" tool annotates objects using their sha1-name as key.\n+This design does not describe a way to migrate notes trees to use\n+newhash-names. That migration is expected to happen separately (for\n+example using a file at the root of the notes tree to describe which\n+hash it uses).\n+\n+Server-side cost\n+~~~~~~~~~~~~~~~~\n+Until Git protocol gains NewHash support, using NewHash based storage\n+on public-facing Git servers is strongly discouraged. Once Git\n+protocol gains NewHash support, NewHash based servers are likely not\n+to support SHA-1 compatibility, to avoid what may be a very expensive\n+hash reencode during clone and to encourage peers to modernize.\n+\n+The design described here allows fetches by SHA-1 clients of a\n+personal NewHash repository because it's not much more difficult than\n+allowing pushes from that repository. This support needs to be guarded\n+by a configuration option --- servers like git.kernel.org that serve a\n+large number of clients would not be expected to bear that cost.\n+\n+Meaning of signatures\n+~~~~~~~~~~~~~~~~~~~~~\n+The signed payload for signed commits and tags does not explicitly\n+name the hash used to identify objects. If some day Git adopts a new\n+hash function with the same length as the current SHA-1 (40\n+hexadecimal digit) or NewHash (64 hexadecimal digit) objects then the\n+intent behind the PGP signed payload in an object signature is\n+unclear:\n+\n+\tobject e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7\n+\ttype commit\n+\ttag v2.12.0\n+\ttagger Junio C Hamano <gitster@pobox.com> 1487962205 -0800\n+\n+\tGit 2.12\n+\n+Does this mean Git v2.12.0 is the commit with sha1-name\n+e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7 or the commit with\n+new-40-digit-hash-name e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7?\n+\n+Fortunately NewHash and SHA-1 have different lengths. If Git starts\n+using another hash with the same length to name objects, then it will\n+need to change the format of signed payloads using that hash to\n+address this issue.\n+\n+Object names on the command line\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+To support the transition (see Transition plan below), this design\n+supports four different modes of operation:\n+\n+ 1. (\"dark launch\") Treat object names input by the user as SHA-1 and\n+    convert any object names written to output to SHA-1, but store\n+    objects using NewHash.  This allows users to test the code with no\n+    visible behavior change except for performance.  This allows\n+    allows running even tests that assume the SHA-1 hash function, to\n+    sanity-check the behavior of the new mode.\n+\n+ 2. (\"early transition\") Allow both SHA-1 and NewHash object names in\n+    input. Any object names written to output use SHA-1. This allows\n+    users to continue to make use of SHA-1 to communicate with peers\n+    (e.g. by email) that have not migrated yet and prepares for mode 3.\n+\n+ 3. (\"late transition\") Allow both SHA-1 and NewHash object names in\n+    input. Any object names written to output use NewHash. In this\n+    mode, users are using a more secure object naming method by\n+    default.  The disruption is minimal as long as most of their peers\n+    are in mode 2 or mode 3.\n+\n+ 4. (\"post-transition\") Treat object names input by the user as\n+    NewHash and write output using NewHash. This is safer than mode 3\n+    because there is less risk that input is incorrectly interpreted\n+    using the wrong hash function.\n+\n+The mode is specified in configuration.\n+\n+The user can also explicitly specify which format to use for a\n+particular revision specifier and for output, overriding the mode. For\n+example:\n+\n+git --output-format=sha1 log abac87a^{sha1}..f787cac^{newhash}\n+\n+Selection of a New Hash\n+-----------------------\n+In early 2005, around the time that Git was written,  Xiaoyun Wang,\n+Yiqun Lisa Yin, and Hongbo Yu announced an attack finding SHA-1\n+collisions in 2^69 operations. In August they published details.\n+Luckily, no practical demonstrations of a collision in full SHA-1 were\n+published until 10 years later, in 2017.\n+\n+The hash function NewHash to replace SHA-1 should be stronger than\n+SHA-1 was: we would like it to be trustworthy and useful in practice\n+for at least 10 years.\n+\n+Some other relevant properties:\n+\n+1. A 256-bit hash (long enough to match common security practice; not\n+   excessively long to hurt performance and disk usage).\n+\n+2. High quality implementations should be widely available (e.g. in\n+   OpenSSL).\n+\n+3. The hash function's properties should match Git's needs (e.g. Git\n+   requires collision and 2nd preimage resistance and does not require\n+   length extension resistance).\n+\n+4. As a tiebreaker, the hash should be fast to compute (fortunately\n+   many contenders are faster than SHA-1).\n+\n+Some hashes under consideration are SHA-256, SHA-512/256, SHA-256x16,\n+K12, and BLAKE2bp-256.\n+\n+Transition plan\n+---------------\n+Some initial steps can be implemented independently of one another:\n+- adding a hash function API (vtable)\n+- teaching fsck to tolerate the gpgsig-newhash field\n+- excluding gpgsig-* from the fields copied by \"git commit --amend\"\n+- annotating tests that depend on SHA-1 values with a SHA1 test\n+  prerequisite\n+- using \"struct object_id\", GIT_MAX_RAWSZ, and GIT_MAX_HEXSZ\n+  consistently instead of \"unsigned char *\" and the hardcoded\n+  constants 20 and 40.\n+- introducing index v3\n+- adding support for the PSRC field and safer object pruning\n+\n+\n+The first user-visible change is the introduction of the objectFormat\n+extension (without compatObjectFormat). This requires:\n+- implementing the loose-object-idx\n+- teaching fsck about this mode of operation\n+- using the hash function API (vtable) when computing object names\n+- signing objects and verifying signatures\n+- rejecting attempts to fetch from or push to an incompatible\n+  repository\n+\n+Next comes introduction of compatObjectFormat:\n+- translating object names between object formats\n+- translating object content between object formats\n+- generating and verifying signatures in the compat format\n+- adding appropriate index entries when adding a new object to the\n+  object store\n+- --output-format option\n+- ^{sha1} and ^{newhash} revision notation\n+- configuration to specify default input and output format (see\n+  \"Object names on the command line\" above)\n+\n+The next step is supporting fetches and pushes to SHA-1 repositories:\n+- allow pushes to a repository using the compat format\n+- generate a topologically sorted list of the SHA-1 names of fetched\n+  objects\n+- convert the fetched packfile to newhash format and generate an idx\n+  file\n+- re-sort to match the order of objects in the fetched packfile\n+\n+The infrastructure supporting fetch also allows converting an existing\n+repository. In converted repositories and new clones, end users can\n+gain support for the new hash function without any visible change in\n+behavior (see \"dark launch\" in the \"Object names on the command line\"\n+section). In particular this allows users to verify NewHash signatures\n+on objects in the repository, and it should ensure the transition code\n+is stable in production in preparation for using it more widely.\n+\n+Over time projects would encourage their users to adopt the \"early\n+transition\" and then \"late transition\" modes to take advantage of the\n+new, more futureproof NewHash object names.\n+\n+When objectFormat and compatObjectFormat are both set, commands\n+generating signatures would generate both SHA-1 and NewHash signatures\n+by default to support both new and old users.\n+\n+In projects using NewHash heavily, users could be encouraged to adopt\n+the \"post-transition\" mode to avoid accidentally making implicit use\n+of SHA-1 object names.\n+\n+Once a critical mass of users have upgraded to a version of Git that\n+can verify NewHash signatures and have converted their existing\n+repositories to support verifying them, we can add support for a\n+setting to generate only NewHash signatures. This is expected to be at\n+least a year later.\n+\n+That is also a good moment to advertise the ability to convert\n+repositories to use NewHash only, stripping out all SHA-1 related\n+metadata. This improves performance by eliminating translation\n+overhead and security by avoiding the possibility of accidentally\n+relying on the safety of SHA-1.\n+\n+Updating Git's protocols to allow a server to specify which hash\n+functions it supports is also an important part of this transition. It\n+is not discussed in detail in this document but this transition plan\n+assumes it happens. :)\n+\n+Alternatives considered\n+-----------------------\n+Upgrading everyone working on a particular project on a flag day\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+Projects like the Linux kernel are large and complex enough that\n+flipping the switch for all projects based on the repository at once\n+is infeasible.\n+\n+Not only would all developers and server operators supporting\n+developers have to switch on the same flag day, but supporting tooling\n+(continuous integration, code review, bug trackers, etc) would have to\n+be adapted as well. This also makes it difficult to get early feedback\n+from some project participants testing before it is time for mass\n+adoption.\n+\n+Using hash functions in parallel\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+(e.g. https://public-inbox.org/git/22708.8913.864049.452252@chiark.greenend.org.uk/ )\n+Objects newly created would be addressed by the new hash, but inside\n+such an object (e.g. commit) it is still possible to address objects\n+using the old hash function.\n+* You cannot trust its history (needed for bisectability) in the\n+  future without further work\n+* Maintenance burden as the number of supported hash functions grows\n+  (they will never go away, so they accumulate). In this proposal, by\n+  comparison, converted objects lose all references to SHA-1.\n+\n+Signed objects with multiple hashes\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+Instead of introducing the gpgsig-newhash field in commit and tag objects\n+for newhash-content based signatures, an earlier version of this design\n+added \"hash newhash <newhash-name>\" fields to strengthen the existing\n+sha1-content based signatures.\n+\n+In other words, a single signature was used to attest to the object\n+content using both hash functions. This had some advantages:\n+* Using one signature instead of two speeds up the signing process.\n+* Having one signed payload with both hashes allows the signer to\n+  attest to the sha1-name and newhash-name referring to the same object.\n+* All users consume the same signature. Broken signatures are likely\n+  to be detected quickly using current versions of git.\n+\n+However, it also came with disadvantages:\n+* Verifying a signed object requires access to the sha1-names of all\n+  objects it references, even after the transition is complete and\n+  translation table is no longer needed for anything else. To support\n+  this, the design added fields such as \"hash sha1 tree <sha1-name>\"\n+  and \"hash sha1 parent <sha1-name>\" to the newhash-content of a signed\n+  commit, complicating the conversion process.\n+* Allowing signed objects without a sha1 (for after the transition is\n+  complete) complicated the design further, requiring a \"nohash sha1\"\n+  field to suppress including \"hash sha1\" fields in the newhash-content\n+  and signed payload.\n+\n+Lazily populated translation table\n+~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n+Some of the work of building the translation table could be deferred to\n+push time, but that would significantly complicate and slow down pushes.\n+Calculating the sha1-name at object creation time at the same time it is\n+being streamed to disk and having its newhash-name calculated should be\n+an acceptable cost.\n+\n+Document History\n+----------------\n+\n+2017-03-03\n+bmwill@google.com, jonathantanmy@google.com, jrnieder@gmail.com,\n+sbeller@google.com\n+\n+Initial version sent to\n+http://public-inbox.org/git/20170304011251.GA26789@aiede.mtv.corp.google.com\n+\n+2017-03-03 jrnieder@gmail.com\n+Incorporated suggestions from jonathantanmy and sbeller:\n+* describe purpose of signed objects with each hash type\n+* redefine signed object verification using object content under the\n+  first hash function\n+\n+2017-03-06 jrnieder@gmail.com\n+* Use SHA3-256 instead of SHA2 (thanks, Linus and brian m. carlson).[1][2]\n+* Make sha3-based signatures a separate field, avoiding the need for\n+  \"hash\" and \"nohash\" fields (thanks to peff[3]).\n+* Add a sorting phase to fetch (thanks to Junio for noticing the need\n+  for this).\n+* Omit blobs from the topological sort during fetch (thanks to peff).\n+* Discuss alternates, git notes, and git servers in the caveats\n+  section (thanks to Junio Hamano, brian m. carlson[4], and Shawn\n+  Pearce).\n+* Clarify language throughout (thanks to various commenters,\n+  especially Junio).\n+\n+2017-09-27 jrnieder@gmail.com, sbeller@google.com\n+* use placeholder NewHash instead of SHA3-256\n+* describe criteria for picking a hash function.\n+* include a transition plan (thanks especially to Brandon Williams\n+  for fleshing these ideas out)\n+* define the translation table (thanks, Shawn Pearce[5], Jonathan\n+  Tan, and Masaya Suzuki)\n+* avoid loose object overhead by packing more aggressively in\n+  \"git gc --auto\"\n+\n+[1] http://public-inbox.org/git/CA+55aFzJtejiCjV0e43+9oR3QuJK2PiFiLQemytoLpyJWe6P9w@mail.gmail.com/\n+[2] http://public-inbox.org/git/CA+55aFz+gkAsDZ24zmePQuEs1XPS9BP_s8O7Q4wQ7LV7X5-oDA@mail.gmail.com/\n+[3] http://public-inbox.org/git/20170306084353.nrns455dvkdsfgo5@sigill.intra.peff.net/\n+[4] http://public-inbox.org/git/20170304224936.rqqtkdvfjgyezsht@genre.crustytoothpaste.net\n+[5] https://public-inbox.org/git/CAJo=hJtoX9=AyLHHpUJS7fueV9ciZ_MNpnEPHUz8Whui6g9F0A@mail.gmail.com/\n-- \n2.14.2.822.g60be5d43e6-goog\n\n"},{"id":"329153","messageId":"xmqqo9puvy1w.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170928044320.GA84719@aiede.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-09-29T06:06:19Z","receivedAt":"2017-09-29T06:06:27Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> This document describes what a transition to a new hash function for\n> Git would look like.  Add it to Documentation/technical/ as the plan\n> of record so that future changes can be recorded as patches.\n>\n> Also-by: Brandon Williams <bmwill@google.com>\n> Also-by: Jonathan Tan <jonathantanmy@google.com>\n> Also-by: Stefan Beller <sbeller@google.com>\n> Signed-off-by: Jonathan Nieder <jrnieder@gmail.com>\n> ---\n\nShoudln't these all be s-o-b: (with a note immediately before that\nto say all four contributed equally or something)?\n\n> +Background\n> +----------\n> +At its core, the Git version control system is a content addressable\n> +filesystem. It uses the SHA-1 hash function to name content. For\n> +example, files, directories, and revisions are referred to by hash\n> +values unlike in other traditional version control systems where files\n> +or versions are referred to via sequential numbers. The use of a hash\n\nTraditional systems refer to files via numbers???  Perhaps \"where\nversions of files are referred to via sequential numbers\" or\nsomething?\n\n> +function to address its content delivers a few advantages:\n> +\n> +* Integrity checking is easy. Bit flips, for example, are easily\n> +  detected, as the hash of corrupted content does not match its name.\n> +* Lookup of objects is fast.\n\n* There is no ambiguity what the object's name should be, given its\n  content.\n\n* Deduping the same content copied across versions and paths is\n  automatic.\n\n> +SHA-1 still possesses the other properties such as fast object lookup\n> +and safe error checking, but other hash functions are equally suitable\n> +that are believed to be cryptographically secure.\n\ns/secure/more &/, perhaps?\n\n> +Goals\n> +-----\n> +...\n> +   c. Users can use SHA-1 and NewHash identifiers for objects\n> +      interchangeably (see \"Object names on the command line\", below).\n\nMental note.  This needs to extend to the \"index X..Y\" lines in the\npatch output, which is used by \"apply -3\" and \"am -3\".\n\n> +2. Allow a complete transition away from SHA-1.\n> +   a. Local metadata for SHA-1 compatibility can be removed from a\n> +      repository if compatibility with SHA-1 is no longer needed.\n\nI like the emphasis on \"Local\" here.  Metadata for compatiblity that\nis embedded in the objects obviously cannot be removed.\n\nFrom that point of view, one of the goals ought to be \"make sure\nthat as much SHA-1 compatibility metadata as possible is local and\noutside the object\".  This goal may not be able to say more than \"as\nmuch as possible\", as signed objects that came from SHA-1 world\nneeds to carry the compatibility metadata somewhere somehow.  \n\nOr perhaps we could.  There is nothing that says a signed tag\ncreated in the SHA-1 world must have the PGP/SHA-1 signature in the\nNewHash payload---it could be split off of the object data and\nstored in a local metadata cache, to be used only when we need to\nconvert it back to the SHA-1 world.\n\nBut I am getting ahead of myself before reading the proposal\nthrough.\n\n> +Non-Goals\n> +---------\n> ...\n> +6. Skip fetching some submodules of a project into a NewHash\n> +   repository. (This also depends on NewHash support in Git\n> +   protocol.)\n\nIt is unclear what this means.  Around submodule support, one thing\nI can think of is that a NewHash tree in a superproject would record\na gitlink that is a NewHash commit object name in it, therefore it\ncannot refer to an unconverted SHA-1 submodule repository.  But it\nis unclear if the above description refers to the same issue, or\nsomething else.\n\n> +Overview\n> +--------\n> +We introduce a new repository format extension. Repositories with this\n> +extension enabled use NewHash instead of SHA-1 to name their objects.\n> +This affects both object names and object content --- both the names\n> +of objects and all references to other objects within an object are\n> +switched to the new hash function.\n> +\n> +NewHash repositories cannot be read by older versions of Git.\n> +\n> +Alongside the packfile, a NewHash repository stores a bidirectional\n> +mapping between NewHash and SHA-1 object names. The mapping is generated\n> +locally and can be verified using \"git fsck\". Object lookups use this\n> +mapping to allow naming objects using either their SHA-1 and NewHash names\n> +interchangeably.\n> +\n> +\"git cat-file\" and \"git hash-object\" gain options to display an object\n> +in its sha1 form and write an object given its sha1 form.\n\nBoth of these are somewhat unclear.  I am guessing that \"git\ncat-file --convert-to=sha1 <type> <NewHashName>\" would emit the\nobject contents converted from their NewHash payload to SHA-1\npayload (blobs are unchanged, trees, commits and tags get their\noutgoing references converted from NewHash to their SHA-1\ncounterparts), and that is what you mean by \"options to display an\nobject in its sha1 form\".  \n\nI am not sure how \"git hash-object\" with the option would work,\nthough.  Do you give an option \"--hash=sha1 --stdout --stdin -t\n<type>\" to feed a NewHash contents (file, tree, commit or tag) to\nthe command, convert it to the SHA-1 content (hmm, how's that\ndifferent from the cat-file's new option???) and then write out its\nloose object representation suitable to be used in the SHA-1 workd?\nWhere do you write it to?  It won't be in the repository, as we\nrejected mixed repository in our Non-Goals section.\n\n> +Object names\n> +~~~~~~~~~~~~\n> +Objects can be named by their 40 hexadecimal digit sha1-name or 64\n> +hexadecimal digit newhash-name, plus names derived from those (see\n> +gitrevisions(7)).\n> +\n> +The sha1-name of an object is the SHA-1 of the concatenation of its\n> +type, length, a nul byte, and the object's sha1-content. This is the\n> +traditional <sha1> used in Git to name objects.\n> +\n> +The newhash-name of an object is the NewHash of the concatenation of its\n> +type, length, a nul byte, and the object's newhash-content.\n\nIt makes me wonder if we want to add the hashname in this object\nheader.  \"length\" would be different for non-blob objects anyway,\nand it is not \"compat metadata\" we want to avoid baked in, yet it\nwould help diagnose a mistake of attempting to use a \"mixed\" objects\nin a single repository.  Not a big issue, though.\n\n> +The format allows round-trip conversion between newhash-content and\n> +sha1-content.\n\nIf it is a goal to eventually be able to lose SHA-1 compatibility\nmetadata from the objects, then we might want to remove SHA-1 based\nsignature bits (e.g. PGP trailer in signed tag, gpgsig header in the\ncommit object) from NewHash contents, and instead have them stored\nin a side \"metadata\" table, only to be used while converting back.\nI dunno if that is desirable.\n\n> +Pack index\n> +~~~~~~~~~~\n> +Pack index (.idx) files use a new v3 format that supports multiple\n> +hash functions. They have the following format (all integers are in\n> +network byte order):\n> +\n> +- A header appears at the beginning and consists of the following:\n> +  - The 4-byte pack index signature: '\\377t0c'\n> +  - 4-byte version number: 3\n> +  - 4-byte length of the header section, including the signature and\n> +    version number\n> +  - 4-byte number of objects contained in the pack\n> +  - 4-byte number of object formats in this pack index: 2\n> +  - For each object format:\n> +    - 4-byte format identifier (e.g., 'sha1' for SHA-1)\n> +    - 4-byte length in bytes of shortened object names. This is the\n> +      shortest possible length needed to make names in the shortened\n> +      object name table unambiguous.\n> +    - 4-byte integer, recording where tables relating to this format\n> +      are stored in this index file, as an offset from the beginning.\n> +  - 4-byte offset to the trailer from the beginning of this file.\n> +  - Zero or more additional key/value pairs (4-byte key, 4-byte\n> +    value). Only one key is supported: 'PSRC'. See the \"Loose objects\n> +    and unreachable objects\" section for supported values and how this\n> +    is used.  All other keys are reserved. Readers must ignore\n> +    unrecognized keys.\n> +- Zero or more NUL bytes. This can optionally be used to improve the\n> +  alignment of the full object name table below.\n> +- Tables for the first object format:\n> +  - A sorted table of shortened object names.  These are prefixes of\n> +    the names of all objects in this pack file, packed together\n> +    without offset values to reduce the cache footprint of the binary\n> +    search for a specific object name.\n\nI take it to mean that the stride is defined in the \"length in bytes\nof shortened object names\" in the file header.  If so, I can see how\nthis would work.  This \"sorted table\", unlike the next one, does not\nsay how it is sorted, but I assume this is just the object name\norder (as opposed to the pack location order the next table uses)?\n\n> +  - A table of full object names in pack order. This allows resolving\n> +    a reference to \"the nth object in the pack file\" (from a\n> +    reachability bitmap or from the next table of another object\n> +    format) to its object name.\n> +\n> +  - A table of 4-byte values mapping object name order to pack order.\n> +    For an object in the table of sorted shortened object names, the\n> +    value at the corresponding index in this table is the index in the\n> +    previous table for that same object.\n> +\n> +    This can be used to look up the object in reachability bitmaps or\n> +    to look up its name in another object format.\n\nAnd this is a separate table because the short-name table wants to\nbe as compact as possible for binary search?  Otherwise an entry in\nthe short-name table could be <pack order number, n-bytes that is\nshort unique prefix>.\n\n> +  - A table of 4-byte CRC32 values of the packed object data, in the\n> +    order that the objects appear in the pack file. This is to allow\n> +    compressed data to be copied directly from pack to pack during\n> +    repacking without undetected data corruption.\n\nAn obvious alternative would be to have the CRC32 checksum near\n(e.g. immediately before) the object data in the packfile (as\nopposed to the .idx file like this document specifies).  I am not\nsure what the pros and cons are between the two, though, and that is\nwhy I mention the possiblity here.\n\nHmm, as the corresponding packfile stores object data only in\nNewHash content format, it is somewhat curious that this table that\nstores CRC32 of the data appears in the \"Tables for each object\nformat\" section, as they would be identical, no?  Unless I am\ngrossly misleading the spec, the checksum should either go outside\nthe \"Tables for each object format\" section but still in .idx, or\nshould be eliminated and become part of the packdata stream instead,\nperhaps?\n\n> +  - A table of 4-byte offset values. For an object in the table of\n> +    sorted shortened object names, the value at the corresponding\n> +    index in this table indicates where that object can be found in\n> +    the pack file. These are usually 31-bit pack file offsets, but\n> +    large offsets are encoded as an index into the next table with the\n> +    most significant bit set.\n\nOy.  So we can go from a short prefix to the pack location by first\nfinding it via binsearch in the short-name table, realize that it is\nnth object in the object name order, and consulting this table.\nWhen we know the pack-order of an object, there is no direct way to\ngo to its location (short of reversing the name-order-to-pack-order\ntable)?\n\n> +  - A table of 8-byte offset entries (empty for pack files less than\n> +    2 GiB). Pack files are organized with heavily used objects toward\n> +    the front, so most object references should not need to refer to\n> +    this table.\n\n> +- Zero or more NUL bytes.\n\n... for padding/aligning.\n\n> +- Tables for the second object format, with the same layout as above,\n> +  up to and not including the table of CRC32 values.\n> +- Zero or more NUL bytes.\n> +- The trailer consists of the following:\n> +  - A copy of the 20-byte NewHash checksum at the end of the\n> +    corresponding packfile.\n> +\n> +  - 20-byte NewHash checksum of all of the above.\n\nWhen did NewHash shrink to 20-byte suddenly?  I think the above two\nare both \"32-byte\"?\n\n> +Loose object index\n> +~~~~~~~~~~~~~~~~~~\n> +A new file $GIT_OBJECT_DIR/loose-object-idx contains information about\n> +all loose objects. Its format is\n> +\n> +  # loose-object-idx\n> +  (newhash-name SP sha1-name LF)*\n> +\n> +where the object names are in hexadecimal format. The file is not\n> +sorted.\n\nShouldn't the file somehow say what hashes are involved to allow us\nmatch it with extension.{objectFormat,compatObjectFormat}, perhaps\nat the end of the \"# loose-object-idx\" line?\n\n> +The loose object index is protected against concurrent writes by a\n> +lock file $GIT_OBJECT_DIR/loose-object-idx.lock. To add a new loose\n> +object:\n> +\n> +1. Write the loose object to a temporary file, like today.\n> +2. Open loose-object-idx.lock with O_CREAT | O_EXCL to acquire the lock.\n> +3. Rename the loose object into place.\n> +4. Open loose-object-idx with O_APPEND and write the new object\n\n\"write the new entry, fsync and close\"?\n\n> +Translation table\n> +~~~~~~~~~~~~~~~~~\n> +The index files support a bidirectional mapping between sha1-names\n> +and newhash-names. The lookup proceeds similarly to ordinary object\n> +lookups. For example, to convert a sha1-name to a newhash-name:\n> +\n> + 1. Look for the object in idx files. If a match is present in the\n> +    idx's sorted list of truncated sha1-names, then:\n> +    a. Read the corresponding entry in the sha1-name order to pack\n> +       name order mapping.\n> +    b. Read the corresponding entry in the full sha1-name table to\n> +       verify we found the right object. If it is, then\n> +    c. Read the corresponding entry in the full newhash-name table.\n> +       That is the object's newhash-name.\n\nc. is possible because b. and c. are sorted the same way, i.e. the\nindex used to consult the full sha1-name table, which is the pack\norder number, can be used to find its full newhash in the \"full\nnewhash sorted by pack order\" table?\n\n> +Reading an object's sha1-content\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n\nI'd stop here and continue in a separate message.  Thanks for a\ndetailed write-up.\n"},{"id":"329163","messageId":"xmqqk20ivsdb.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"xmqqo9puvy1w.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-09-29T08:09:04Z","receivedAt":"2017-09-29T08:09:24Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Junio C Hamano <gitster@pobox.com> writes:\n\n> Or perhaps we could.  There is nothing that says a signed tag\n> created in the SHA-1 world must have the PGP/SHA-1 signature in the\n> NewHash payload---it could be split off of the object data and\n> stored in a local metadata cache, to be used only when we need to\n> convert it back to the SHA-1 world.\n> ...\n>> +The format allows round-trip conversion between newhash-content and\n>> +sha1-content.\n>\n> If it is a goal to eventually be able to lose SHA-1 compatibility\n> metadata from the objects, then we might want to remove SHA-1 based\n> signature bits (e.g. PGP trailer in signed tag, gpgsig header in the\n> commit object) from NewHash contents, and instead have them stored\n> in a side \"metadata\" table, only to be used while converting back.\n> I dunno if that is desirable.\n\nLet's keep it simple by ignoring all of the above.  Even though\nleaving the sha1-gpgsig and other crufts would etch these\ncompatibility metadata in objects forever, these remain only in\nobjects that originate from SHA-1 world, or in objects created in\nthe NewHash world only while the project participants still care\nabout SHA-1 compatibility.  Strictly speaking, it would be super\nnice if we can do without contaminating these newly created objects\nwith SHA-1 compatibility headers, just like we wish to be able to\ndrop the SHA-1 vs NewHash mapping table after projects participants\nstop careing about SHA-1 compatiblity, it may not be worth it.  Of\ncourse, if we decide to spend a bit more brain cycle to design how\nwe push these out of the object proper, the same solution would\nautomatically allow us to omit SHA-1 compatibility headers from the\nobjects that were converted from SHA-1 world.\n>\n>> +  - A table of 4-byte CRC32 values of the packed object data, in the\n>> +    order that the objects appear in the pack file. This is to allow\n>> +    compressed data to be copied directly from pack to pack during\n>> +    repacking without undetected data corruption.\n>\n> An obvious alternative would be to have the CRC32 checksum near\n> (e.g. immediately before) the object data in the packfile (as\n> opposed to the .idx file like this document specifies).  I am not\n> sure what the pros and cons are between the two, though, and that is\n> why I mention the possiblity here.\n>\n> Hmm, as the corresponding packfile stores object data only in\n> NewHash content format, it is somewhat curious that this table that\n> stores CRC32 of the data appears in the \"Tables for each object\n> format\" section, as they would be identical, no?  Unless I am\n> grossly misleading the spec, the checksum should either go outside\n> the \"Tables for each object format\" section but still in .idx, or\n> should be eliminated and become part of the packdata stream instead,\n> perhaps?\n\nThinking about this a bit more, I think a single table per .idx file\nwould be the right way to go, not a checksum immediately after or\nbefore the object data that is embedded in the pack stream.  In the\nNewHash world (after this initial migration), we would want to be\nable to stream NewHash packstream that comes from the network\nstraight to disk, which would mean these in-line CRC32 data would\nneed to be sent over the wire (i.e. 4-byte per object sent); that is\nan unneeded overhead, as the packstream has its trailing checksum to\nprotect the whole thing anyway.\n"},{"id":"329177","messageId":"alpine.DEB.2.21.1.1709291416290.40514@virtualbox","threadId":"45288","inReplyTo":"59C149A3.6080506@st.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-29T13:17:43Z","receivedAt":"2017-09-29T13:18:27Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Gilles,\n\nOn Tue, 19 Sep 2017, Gilles Van Assche wrote:\n\n> On 19/09/17 00:16, Johannes Schindelin wrote:\n> >>> SHA-256 got much more cryptanalysis than SHA3-256 […].\n> >>\n> >> I do not think this is true.\n> >\n> > Please read what I said again: SHA-256 got much more cryptanalysis\n> > than SHA3-256.\n> \n> Indeed. What I meant is that SHA3-256 got at least as much cryptanalysis\n> as SHA-256. :-)\n\nOh? I got the opposite impression... I got the impression that *everybody*\nin the field banged on all the SHA-2 candidates because everybody was\nworried that SHA-1 would be utterly broken soon, and I got the impression\nthat after this SHA-2 competition, people were less worried?\n\nBesides, I would expect that the difference in age (at *least* 7 years by\nmy humble arithmetic skills) to make a difference...\n\n> > I never said that SHA3-256 got little cryptanalysis. Personally, I\n> > think that SHA3-256 got a ton more cryptanalysis than SHA-1, and that\n> > SHA-256 *still* got more cryptanalysis. But my opinion does not count,\n> > really. However, the two experts I pestered with questions over\n> > questions left me with that strong impression, and their opinion does\n> > count.\n> \n> OK, I respect your opinion and that of your two experts. Yet, the \"much\n> more\" part of your statement, in particular, is something that may\n> require a bit more explanations.\n\nI would also like to point out the ubiquitousness of SHA-256. I have been\nasked to provide SHA-256 checksums for the downloads of Git for Windows,\nbut not SHA3-256...\n\nAnd this is a practically-relevant thing: the more users of an algorithm\nthere are, the more high-quality implementations you can choose from. And\nthis becomes relevant, say, when you have to switch implementations due to\nlicense changes (*cough, cough looking in OpenSSL's direction*). Or when\nyou have to support the biggest Git repository on this planet and have to\neek out 5-10% more performance using the latest hardware. All of a sudden,\nyour consideration cannot only be \"security of the algorithm\" any longer.\n\nHaving said that, I am *really* happy to have SHA3-256 as a valid fallback\noption in case SHA-256 should be broken.\n\nCiao,\nJohannes\n"},{"id":"329178","messageId":"acd96750-c165-650c-c67f-44465f2075f2@noekeon.org","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709291416290.40514@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Joan Daemen","fromEmail":"jda@noekeon.org","sentAt":"2017-09-29T14:54:23Z","receivedAt":"2017-09-29T15:04:15Z","isPatch":false,"sender":{"key":"jda@noekeon.org","avatar":null},"body":"Dear Johannes,\n\nif ever there was a SHA-2 competition, it must have been held inside \nNSA:-) But maybe you are confusing with the SHA-3 competition. In any \ncase, when considering SHA-2 vs SHA-3 for usage in git, you may have a \nlook at arguments we give in the following blogpost:\n\nhttps://keccak.team/2017/open_source_crypto.html\n\nKind regards,\n\nJoan Daemen\n\nOn 29/09/17 15:17, Johannes Schindelin wrote:\n> Hi Gilles,\n>\n> On Tue, 19 Sep 2017, Gilles Van Assche wrote:\n>\n>> On 19/09/17 00:16, Johannes Schindelin wrote:\n>>>>> SHA-256 got much more cryptanalysis than SHA3-256 […].\n>>>> I do not think this is true.\n>>> Please read what I said again: SHA-256 got much more cryptanalysis\n>>> than SHA3-256.\n>> Indeed. What I meant is that SHA3-256 got at least as much cryptanalysis\n>> as SHA-256. :-)\n> Oh? I got the opposite impression... I got the impression that *everybody*\n> in the field banged on all the SHA-2 candidates because everybody was\n> worried that SHA-1 would be utterly broken soon, and I got the impression\n> that after this SHA-2 competition, people were less worried?\n>\n> Besides, I would expect that the difference in age (at *least* 7 years by\n> my humble arithmetic skills) to make a difference...\n>\n>>> I never said that SHA3-256 got little cryptanalysis. Personally, I\n>>> think that SHA3-256 got a ton more cryptanalysis than SHA-1, and that\n>>> SHA-256 *still* got more cryptanalysis. But my opinion does not count,\n>>> really. However, the two experts I pestered with questions over\n>>> questions left me with that strong impression, and their opinion does\n>>> count.\n>> OK, I respect your opinion and that of your two experts. Yet, the \"much\n>> more\" part of your statement, in particular, is something that may\n>> require a bit more explanations.\n> I would also like to point out the ubiquitousness of SHA-256. I have been\n> asked to provide SHA-256 checksums for the downloads of Git for Windows,\n> but not SHA3-256...\n>\n> And this is a practically-relevant thing: the more users of an algorithm\n> there are, the more high-quality implementations you can choose from. And\n> this becomes relevant, say, when you have to switch implementations due to\n> license changes (*cough, cough looking in OpenSSL's direction*). Or when\n> you have to support the biggest Git repository on this planet and have to\n> eek out 5-10% more performance using the latest hardware. All of a sudden,\n> your consideration cannot only be \"security of the algorithm\" any longer.\n>\n> Having said that, I am *really* happy to have SHA3-256 as a valid fallback\n> option in case SHA-256 should be broken.\n>\n> Ciao,\n> Johannes\n\n"},{"id":"329183","messageId":"20170929173413.GI19555@aiede.mtv.corp.google.com","threadId":"45288","inReplyTo":"xmqqo9puvy1w.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2017-09-29T17:34:13Z","receivedAt":"2017-09-29T17:34:23Z","isPatch":true,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Junio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n\n>> This document describes what a transition to a new hash function for\n>> Git would look like.  Add it to Documentation/technical/ as the plan\n>> of record so that future changes can be recorded as patches.\n>>\n>> Also-by: Brandon Williams <bmwill@google.com>\n>> Also-by: Jonathan Tan <jonathantanmy@google.com>\n>> Also-by: Stefan Beller <sbeller@google.com>\n>> Signed-off-by: Jonathan Nieder <jrnieder@gmail.com>\n>> ---\n>\n> Shoudln't these all be s-o-b: (with a note immediately before that\n> to say all four contributed equally or something)?\n\nI don't want to get lost in the weeds in the question of how to\nrepresent such a collaborative effort in git's metadata.\n\nYou're right that I should collect their sign-offs!  Your approach of\nusing text instead of machine-readable data for common authorship also\nseems okay.\n\nIn any event, this is indeed\n\nSigned-off-by: Brandon Williams <bmwill@google.com>\nSigned-off-by: Jonathan Tan <jonathantanmy@google.com>\nSigned-off-by: Stefan Beller <sbeller@google.com>\n\n(I just checked :)).\n\n>> +Background\n>> +----------\n>> +At its core, the Git version control system is a content addressable\n>> +filesystem. It uses the SHA-1 hash function to name content. For\n>> +example, files, directories, and revisions are referred to by hash\n>> +values unlike in other traditional version control systems where files\n>> +or versions are referred to via sequential numbers. The use of a hash\n>\n> Traditional systems refer to files via numbers???  Perhaps \"where\n> versions of files are referred to via sequential numbers\" or\n> something?\n\nGood point.  The wording you suggested will work well.\n\n>> +function to address its content delivers a few advantages:\n>> +\n>> +* Integrity checking is easy. Bit flips, for example, are easily\n>> +  detected, as the hash of corrupted content does not match its name.\n>> +* Lookup of objects is fast.\n>\n> * There is no ambiguity what the object's name should be, given its\n>   content.\n>\n> * Deduping the same content copied across versions and paths is\n>   automatic.\n\n:)  Yep, these are nice too, especially that second one.\n\nIt also is how we make diff-ing fast.\n\n>> +SHA-1 still possesses the other properties such as fast object lookup\n>> +and safe error checking, but other hash functions are equally suitable\n>> +that are believed to be cryptographically secure.\n>\n> s/secure/more &/, perhaps?\n\nWe were looking for a phrase meaning that it should be a cryptographic\nhash function in good standing, which SHA-1 is at least approaching\nnot being.\n\n\"more secure\" should work fine.  Let's go with that.\n\n>> +Goals\n>> +-----\n>> +...\n>> +   c. Users can use SHA-1 and NewHash identifiers for objects\n>> +      interchangeably (see \"Object names on the command line\", below).\n>\n> Mental note.  This needs to extend to the \"index X..Y\" lines in the\n> patch output, which is used by \"apply -3\" and \"am -3\".\n\nWill add a note about this to \"Object names on the command line\".  Stefan\nhad already pointed out that that section should really be renamed to\nsomething like \"Object names in input and output\".\n\n>> +2. Allow a complete transition away from SHA-1.\n>> +   a. Local metadata for SHA-1 compatibility can be removed from a\n>> +      repository if compatibility with SHA-1 is no longer needed.\n>\n> I like the emphasis on \"Local\" here.  Metadata for compatiblity that\n> is embedded in the objects obviously cannot be removed.\n>\n> From that point of view, one of the goals ought to be \"make sure\n> that as much SHA-1 compatibility metadata as possible is local and\n> outside the object\".  This goal may not be able to say more than \"as\n> much as possible\", as signed objects that came from SHA-1 world\n> needs to carry the compatibility metadata somewhere somehow.\n>\n> Or perhaps we could.  There is nothing that says a signed tag\n> created in the SHA-1 world must have the PGP/SHA-1 signature in the\n> NewHash payload---it could be split off of the object data and\n> stored in a local metadata cache, to be used only when we need to\n> convert it back to the SHA-1 world.\n\nThat would break round-tripping and would mean that multiple SHA-1\nobjects could have the same NewHash name.  In other words, from\nmy point of view there is something that says that such data must\nbe preserved.\n\nAnother way to put it: even after removing all SHA-1 compatibility\nmetadata, one nice feature of this design is that it can be recovered\nif I change my mind, from data in the NewHash based repository alone.\n\n[...]\n>> +Non-Goals\n>> +---------\n>> ...\n>> +6. Skip fetching some submodules of a project into a NewHash\n>> +   repository. (This also depends on NewHash support in Git\n>> +   protocol.)\n>\n> It is unclear what this means.  Around submodule support, one thing\n> I can think of is that a NewHash tree in a superproject would record\n> a gitlink that is a NewHash commit object name in it, therefore it\n> cannot refer to an unconverted SHA-1 submodule repository.  But it\n> is unclear if the above description refers to the same issue, or\n> something else.\n\nIt refers to that issue.\n\n[...]\n>> +Overview\n>> +--------\n>> +We introduce a new repository format extension. Repositories with this\n>> +extension enabled use NewHash instead of SHA-1 to name their objects.\n>> +This affects both object names and object content --- both the names\n>> +of objects and all references to other objects within an object are\n>> +switched to the new hash function.\n>> +\n>> +NewHash repositories cannot be read by older versions of Git.\n>> +\n>> +Alongside the packfile, a NewHash repository stores a bidirectional\n>> +mapping between NewHash and SHA-1 object names. The mapping is generated\n>> +locally and can be verified using \"git fsck\". Object lookups use this\n>> +mapping to allow naming objects using either their SHA-1 and NewHash names\n>> +interchangeably.\n>> +\n>> +\"git cat-file\" and \"git hash-object\" gain options to display an object\n>> +in its sha1 form and write an object given its sha1 form.\n>\n> Both of these are somewhat unclear.\n\nI think we can delete this paragraph.  It was written before the\n\"Object names on the command line\" section that goes into such issues\nin more detail.\n\n[...]\n>> +Object names\n>> +~~~~~~~~~~~~\n>> +Objects can be named by their 40 hexadecimal digit sha1-name or 64\n>> +hexadecimal digit newhash-name, plus names derived from those (see\n>> +gitrevisions(7)).\n>> +\n>> +The sha1-name of an object is the SHA-1 of the concatenation of its\n>> +type, length, a nul byte, and the object's sha1-content. This is the\n>> +traditional <sha1> used in Git to name objects.\n>> +\n>> +The newhash-name of an object is the NewHash of the concatenation of its\n>> +type, length, a nul byte, and the object's newhash-content.\n>\n> It makes me wonder if we want to add the hashname in this object\n> header.  \"length\" would be different for non-blob objects anyway,\n> and it is not \"compat metadata\" we want to avoid baked in, yet it\n> would help diagnose a mistake of attempting to use a \"mixed\" objects\n> in a single repository.  Not a big issue, though.\n\nDo you mean that adding the hashname into the computation that\nproduces the object name would help in some use case?\n\nOr do you mean storing the hashname on disk somewhere, even if it\ndoesn't enter into the object name?  For the latter, we store the\nhashname in the .git/config extensions.* configuration and the pack\nindex files.  You also suggested storing the hash name in\n.git/objects/loose-object-idx, which seems to me like a good idea.\n\nWe didn't touch on the .pack format but we probably need to (if only\nbecause of the size of REF_DELTAs and the cksum trailer), and it would\nalso need to name what object format it is using.\n\nFor loose objects, it would be nice to name the hash in the file, so\nthat \"file\" can understand what is happening if someone accidentally\nmixes types using \"cp\".  The only downside is losing the ability to\ncopy blobs (which have the same content despite being named using\ndifferent hashes) between repositories after determining their new\nnames.  That doesn't seem like a strong downside --- it's pretty\nharmless to include the hash type in loose object files, too.  I think\nI would prefer this to be a \"magic number\" instead of part of the\nzlib-deflated payload, since this way \"file\" can discover it more\neasily.\n\n>> +The format allows round-trip conversion between newhash-content and\n>> +sha1-content.\n>\n> If it is a goal to eventually be able to lose SHA-1 compatibility\n> metadata from the objects, then we might want to remove SHA-1 based\n> signature bits (e.g. PGP trailer in signed tag, gpgsig header in the\n> commit object) from NewHash contents, and instead have them stored\n> in a side \"metadata\" table, only to be used while converting back.\n> I dunno if that is desirable.\n\nI don't consider that desirable.\n\nA SHA-1 based signature is still of historical interest even if my\ncenturies-newer version of Git is not able to verify it.\n\n[...]\n> I take it to mean that the stride is defined in the \"length in bytes\n> of shortened object names\" in the file header.  If so, I can see how\n> this would work.  This \"sorted table\", unlike the next one, does not\n> say how it is sorted, but I assume this is just the object name\n> order (as opposed to the pack location order the next table uses)?\n\nYes.  Will clarify.\n\n>> +  - A table of full object names in pack order. This allows resolving\n>> +    a reference to \"the nth object in the pack file\" (from a\n>> +    reachability bitmap or from the next table of another object\n>> +    format) to its object name.\n>> +\n>> +  - A table of 4-byte values mapping object name order to pack order.\n>> +    For an object in the table of sorted shortened object names, the\n>> +    value at the corresponding index in this table is the index in the\n>> +    previous table for that same object.\n>> +\n>> +    This can be used to look up the object in reachability bitmaps or\n>> +    to look up its name in another object format.\n>\n> And this is a separate table because the short-name table wants to\n> be as compact as possible for binary search?  Otherwise an entry in\n> the short-name table could be <pack order number, n-bytes that is\n> short unique prefix>.\n\nYes.  The idx v2 format has a similar design.\n\n>> +  - A table of 4-byte CRC32 values of the packed object data, in the\n>> +    order that the objects appear in the pack file. This is to allow\n>> +    compressed data to be copied directly from pack to pack during\n>> +    repacking without undetected data corruption.\n>\n> An obvious alternative would be to have the CRC32 checksum near\n> (e.g. immediately before) the object data in the packfile (as\n> opposed to the .idx file like this document specifies).  I am not\n> sure what the pros and cons are between the two, though, and that is\n> why I mention the possiblity here.\n\nAs you mentioned under separate cover, it is useful for derived data\nlike this to be outside the packfile.\n\n> Hmm, as the corresponding packfile stores object data only in\n> NewHash content format, it is somewhat curious that this table that\n> stores CRC32 of the data appears in the \"Tables for each object\n> format\" section, as they would be identical, no?  Unless I am\n> grossly misleading the spec, the checksum should either go outside\n> the \"Tables for each object format\" section but still in .idx, or\n> should be eliminated and become part of the packdata stream instead,\n> perhaps?\n\nIt's actually only present for the first object format.  Will find a\nbetter way to describe this.\n\n>> +  - A table of 4-byte offset values. For an object in the table of\n>> +    sorted shortened object names, the value at the corresponding\n>> +    index in this table indicates where that object can be found in\n>> +    the pack file. These are usually 31-bit pack file offsets, but\n>> +    large offsets are encoded as an index into the next table with the\n>> +    most significant bit set.\n>\n> Oy.  So we can go from a short prefix to the pack location by first\n> finding it via binsearch in the short-name table, realize that it is\n> nth object in the object name order, and consulting this table.\n> When we know the pack-order of an object, there is no direct way to\n> go to its location (short of reversing the name-order-to-pack-order\n> table)?\n\nAn earlier version of the design also had a pack-order-to-pack-offset\ntable, but we weren't able to think of any cases where that would be\nused without also looking up the object name that can be used to\nverify the integrity of the inflated object.\n\nDo you have an application in mind?\n\n[...]\n>> +- Tables for the second object format, with the same layout as above,\n>> +  up to and not including the table of CRC32 values.\n>> +- Zero or more NUL bytes.\n>> +- The trailer consists of the following:\n>> +  - A copy of the 20-byte NewHash checksum at the end of the\n>> +    corresponding packfile.\n>> +\n>> +  - 20-byte NewHash checksum of all of the above.\n>\n> When did NewHash shrink to 20-byte suddenly?  I think the above two\n> are both \"32-byte\"?\n\nYes, good catch.\n\n[...]\n>> +Loose object index\n>> +~~~~~~~~~~~~~~~~~~\n>> +A new file $GIT_OBJECT_DIR/loose-object-idx contains information about\n>> +all loose objects. Its format is\n>> +\n>> +  # loose-object-idx\n>> +  (newhash-name SP sha1-name LF)*\n>> +\n>> +where the object names are in hexadecimal format. The file is not\n>> +sorted.\n>\n> Shouldn't the file somehow say what hashes are involved to allow us\n> match it with extension.{objectFormat,compatObjectFormat}, perhaps\n> at the end of the \"# loose-object-idx\" line?\n\nGood idea!\n\n[...]\n>> +The loose object index is protected against concurrent writes by a\n>> +lock file $GIT_OBJECT_DIR/loose-object-idx.lock. To add a new loose\n>> +object:\n>> +\n>> +1. Write the loose object to a temporary file, like today.\n>> +2. Open loose-object-idx.lock with O_CREAT | O_EXCL to acquire the lock.\n>> +3. Rename the loose object into place.\n>> +4. Open loose-object-idx with O_APPEND and write the new object\n>\n> \"write the new entry, fsync and close\"?\n\nYes, I think we do need to fsync. :/\n\n[...]\n>> +Translation table\n>> +~~~~~~~~~~~~~~~~~\n>> +The index files support a bidirectional mapping between sha1-names\n>> +and newhash-names. The lookup proceeds similarly to ordinary object\n>> +lookups. For example, to convert a sha1-name to a newhash-name:\n>> +\n>> + 1. Look for the object in idx files. If a match is present in the\n>> +    idx's sorted list of truncated sha1-names, then:\n>> +    a. Read the corresponding entry in the sha1-name order to pack\n>> +       name order mapping.\n>> +    b. Read the corresponding entry in the full sha1-name table to\n>> +       verify we found the right object. If it is, then\n>> +    c. Read the corresponding entry in the full newhash-name table.\n>> +       That is the object's newhash-name.\n>\n> c. is possible because b. and c. are sorted the same way, i.e. the\n> index used to consult the full sha1-name table, which is the pack\n> order number, can be used to find its full newhash in the \"full\n> newhash sorted by pack order\" table?\n\nYes.\n\n>> +Reading an object's sha1-content\n>> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n>\n> I'd stop here and continue in a separate message.  Thanks for a\n> detailed write-up.\n\nThanks for looking it over.\n\nJonathan\n"},{"id":"329220","messageId":"alpine.DEB.2.21.1.1709292355060.40514@virtualbox","threadId":"45288","inReplyTo":"acd96750-c165-650c-c67f-44465f2075f2@noekeon.org","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-09-29T22:33:33Z","receivedAt":"2017-09-29T22:34:16Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Joan,\n\nOn Fri, 29 Sep 2017, Joan Daemen wrote:\n\n> if ever there was a SHA-2 competition, it must have been held inside NSA:-)\n\nOops. My bad, I indeed got confused about that, as you suggest below (I\nactually thought of the AES competition, but that was obviously not about\nSHA-2). Sorry.\n\n> But maybe you are confusing with the SHA-3 competition. In any case,\n> when considering SHA-2 vs SHA-3 for usage in git, you may have a look at\n> arguments we give in the following blogpost:\n> \n> https://keccak.team/2017/open_source_crypto.html\n\nThanks for the pointer!\n\nSmall nit: the post uses \"its\" in place of \"it's\", twice.\n\nIt does have a good point, of course: the scientific exchange (which you\ncall \"open-source\" in spirit) makes tons of sense.\n\nAs far as Git is concerned, we not only care about the source code of the\nhash algorithm we use, we need to care even more about what you call\n\"executable\": ready-to-use, high quality, well-tested implementations.\n\nWe carry source code for SHA-1 as part of Git's source code, which was\nhand-tuned to be as fast as Linus could get it, which was tricky given\nthat the tuning should be general enough to apply to all common intel\nCPUs.\n\nThis hand-crafted code was blown out of the water by OpenSSL's SHA-1 in\nour tests here at Microsoft, thanks to the fact that OpenSSL does\nvectorized SHA-1 computation now.\n\nTo me, this illustrates why it is not good enough to have only a reference\nimplementation available at our finger tips. Of course, above-mentioned\nOpenSSL supports SHA-256 and SHA3-256, too, and at least recent versions\nvectorize those, too.\n\nAlso, ARM processors have become a lot more popular, so we'll want to have\nhigh-quality implementations of the hash algorithm also for those\nprocessors.\n\nLikewise, in contrast to 2005, nowadays implementations of Git in\nlanguages as obscure as Javascript are not only theoretical but do exist\nin practice (https://github.com/creationix/js-git). I had a *very* quick\nlook for libraries providing crypto in Javascript and immediately found\nthe Standford Javascript Crypto library\n(https://github.com/bitwiseshiftleft/sjcl/) which seems to offer SHA-256\nbut not SHA3-256 computation.\n\nBack to Intel processors: I read some vague hints about extensions\naccelerating SHA-256 computation on future Intel processors, but not\nSHA3-256.\n\nIt would make sense, of course, that more crypto libraries and more\nhardware support would be available for SHA-256 than for SHA3-256 given\nthe time since publication: 16 vs 5 years (I am playing it loose here,\ntaking just the year into account, not the exact date, so please treat\nthat merely as a ballpark figure).\n\nSo from a practical point of view, I wonder what your take is on, say,\nhardware support for SHA3-256. Do you think this will become a focus soon?\n\nAlso, what is your take on the question whether SHA-256 is good enough?\nSHA-1 was broken theoretically already 10 years after it was published\n(which unfortunately did not prevent us from baking it into Git), after\nall, while SHA-256 is 16 years old and the only known weakness does not\napply to Git's usage?\n\nAlso, while I have the attention of somebody who knows a heck more about\ncryptography than Git's top 10 committers combined: how soon do you expect\npractical SHA-1 attacks that are much worse than what we already have\nseen? I am concerned that if we do not move fast enough to a new hash\nalgorithm, and somebody finds a way in the meantime to craft arbitrary\nmessages given a prefix and an SHA-1, then we have a huge problem on\nour hands.\n\nCiao,\nJohannes\n"},{"id":"329259","messageId":"6d96ec902dc1d500ba3fcb11d31b2015@mail.noekeon.org","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709292355060.40514@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Joan Daemen","fromEmail":"jda@noekeon.org","sentAt":"2017-09-30T22:02:08Z","receivedAt":"2017-09-30T22:02:17Z","isPatch":false,"sender":{"key":"jda@noekeon.org","avatar":null},"body":"Dear Johannes,\n\nthanks for your response and taking the effort to express your concerns. \nPlease see below for some feedback.\n\nOn 30/09/17 00:33, Johannes Schindelin wrote:\n> Hi Joan,\n> \n> On Fri, 29 Sep 2017, Joan Daemen wrote:\n> \n>> if ever there was a SHA-2 competition, it must have been held inside \n>> NSA:-)\n> Oops. My bad, I indeed got confused about that, as you suggest below (I\n> actually thought of the AES competition, but that was obviously not \n> about\n> SHA-2). Sorry.\n> \n>> But maybe you are confusing with the SHA-3 competition. In any case,\n>> when considering SHA-2 vs SHA-3 for usage in git, you may have a look \n>> at\n>> arguments we give in the following blogpost:\n>> \n>> https://keccak.team/2017/open_source_crypto.html\n> Thanks for the pointer!\n> \n> Small nit: the post uses \"its\" in place of \"it's\", twice.\n\nThanks, we'll correct that.\n\n> It does have a good point, of course: the scientific exchange (which \n> you\n> call \"open-source\" in spirit) makes tons of sense.\n> \n> As far as Git is concerned, we not only care about the source code of \n> the\n> hash algorithm we use, we need to care even more about what you call\n> \"executable\": ready-to-use, high quality, well-tested implementations.\n> \n> We carry source code for SHA-1 as part of Git's source code, which was\n> hand-tuned to be as fast as Linus could get it, which was tricky given\n> that the tuning should be general enough to apply to all common intel\n> CPUs.\n> \n> This hand-crafted code was blown out of the water by OpenSSL's SHA-1 in\n> our tests here at Microsoft, thanks to the fact that OpenSSL does\n> vectorized SHA-1 computation now.\n> \n> To me, this illustrates why it is not good enough to have only a \n> reference\n> implementation available at our finger tips. Of course, above-mentioned\n> OpenSSL supports SHA-256 and SHA3-256, too, and at least recent \n> versions\n> vectorize those, too.\n\nThere is a lot of high-quality optimized code for all SHA-3 functions \nand many CPUs in the Keccak code package \nhttps://github.com/gvanas/KeccakCodePackage but also OpenSSL contains \nsome good SHA-3 code and then there are all those related to Ethereum.\n\nBy the way, you speak about SHA3-256, but the right choice would be to \nuse SHAKE128. Well, what is exactly the right choice depends on what you \nwant. If you want to have a function in the SHA3 standard (FIPS 202), it \nis SHAKE128. You can boost performance on high-end CPUs by adopting \nParallelhash from NIST SP 800-185, still a NIST standard. You can \nmultiply that performance again by a factor of 2 by adopting \nKangarooTwelve. This is our (Keccak team) proposal for a parallelizable \nKeccak-based hash function that has a safety margin comparable to that \nof the SHA-2 functions. See https://keccak.team/kangarootwelve.html\nMay I also suggest you read https://keccak.team/2017/is_sha3_slow.html\n\n> Also, ARM processors have become a lot more popular, so we'll want to \n> have\n> high-quality implementations of the hash algorithm also for those\n> processors.\n> \n> Likewise, in contrast to 2005, nowadays implementations of Git in\n> languages as obscure as Javascript are not only theoretical but do \n> exist\n> in practice (https://github.com/creationix/js-git). I had a *very* \n> quick\n> look for libraries providing crypto in Javascript and immediately found\n> the Standford Javascript Crypto library\n> (https://github.com/bitwiseshiftleft/sjcl/) which seems to offer \n> SHA-256\n> but not SHA3-256 computation.\n> \n> Back to Intel processors: I read some vague hints about extensions\n> accelerating SHA-256 computation on future Intel processors, but not\n> SHA3-256.\n> \n> It would make sense, of course, that more crypto libraries and more\n> hardware support would be available for SHA-256 than for SHA3-256 given\n> the time since publication: 16 vs 5 years (I am playing it loose here,\n> taking just the year into account, not the exact date, so please treat\n> that merely as a ballpark figure).\n> \n> So from a practical point of view, I wonder what your take is on, say,\n> hardware support for SHA3-256. Do you think this will become a focus \n> soon?\n\nI think this is a chicken-and-egg problem. In any case, hardware support \nfor one SHA3-256 will also work for the other SHA3 and SHAKE functions \nas they all use the same underlying primitive: the Keccak-f permutation. \nThis is not the case for SHA2 because SHA224 and SHA256 use a different \ncompression function than SHA384, SHA512, SHA512/224 and SHA512/256.\n\n> Also, what is your take on the question whether SHA-256 is good enough?\n> SHA-1 was broken theoretically already 10 years after it was published\n> (which unfortunately did not prevent us from baking it into Git), after\n> all, while SHA-256 is 16 years old and the only known weakness does not\n> apply to Git's usage?\n\nSHA-256 is more conservative than SHA-1 and I don't expect it to be \nbroken in the coming decades (unless NSA inserted a backdoor but I don't \nthink that is likely). But looking at the existing cryptanalysis, I \nthink it is even less likely that I SHAKE128, ParallelHash or \nKangarooTwelve will be broken anytime.\n\n> Also, while I have the attention of somebody who knows a heck more \n> about\n> cryptography than Git's top 10 committers combined: how soon do you \n> expect\n> practical SHA-1 attacks that are much worse than what we already have\n> seen? I am concerned that if we do not move fast enough to a new hash\n> algorithm, and somebody finds a way in the meantime to craft arbitrary\n> messages given a prefix and an SHA-1, then we have a huge problem on\n> our hands.\n\nThis is hard to say. To be honest, when witnessing the first MD5 \ncollisions I did not expect them to lead to some real world attacks and \njust a few years later we saw real-world forged certificates based on \nMD5 collisions. And SHA-1 has a lot in common with MD5...\n\nBut let me end with a philosophical note. Independent of all the \narguments for and against, I think this is ultimately about doing the \nright thing. The choice is here between SHA1/SHA2 on the one hand and \nSHA3/Keccak on the other. The former standards are imposed on us by NSA \nand the latter are the best that came out of an open competition \ninvolving all experts in the field worldwide. What would be closest to \nthe philosophy of Git (and by extension Linux or open-source in \ngeneral)?\n\nKind regards,\n\nJoan\n\n\nOn 30/09/17 00:33, Johannes Schindelin wrote:\n> Hi Joan,\n> \n> On Fri, 29 Sep 2017, Joan Daemen wrote:\n> \n>> if ever there was a SHA-2 competition, it must have been held inside \n>> NSA:-)\n> Oops. My bad, I indeed got confused about that, as you suggest below (I\n> actually thought of the AES competition, but that was obviously not \n> about\n> SHA-2). Sorry.\n> \n>> But maybe you are confusing with the SHA-3 competition. In any case,\n>> when considering SHA-2 vs SHA-3 for usage in git, you may have a look \n>> at\n>> arguments we give in the following blogpost:\n>> \n>> https://keccak.team/2017/open_source_crypto.html\n> Thanks for the pointer!\n> \n> Small nit: the post uses \"its\" in place of \"it's\", twice.\n> \n> It does have a good point, of course: the scientific exchange (which \n> you\n> call \"open-source\" in spirit) makes tons of sense.\n> \n> As far as Git is concerned, we not only care about the source code of \n> the\n> hash algorithm we use, we need to care even more about what you call\n> \"executable\": ready-to-use, high quality, well-tested implementations.\n> \n> We carry source code for SHA-1 as part of Git's source code, which was\n> hand-tuned to be as fast as Linus could get it, which was tricky given\n> that the tuning should be general enough to apply to all common intel\n> CPUs.\n> \n> This hand-crafted code was blown out of the water by OpenSSL's SHA-1 in\n> our tests here at Microsoft, thanks to the fact that OpenSSL does\n> vectorized SHA-1 computation now.\n> \n> To me, this illustrates why it is not good enough to have only a \n> reference\n> implementation available at our finger tips. Of course, above-mentioned\n> OpenSSL supports SHA-256 and SHA3-256, too, and at least recent \n> versions\n> vectorize those, too.\n> \n> Also, ARM processors have become a lot more popular, so we'll want to \n> have\n> high-quality implementations of the hash algorithm also for those\n> processors.\n> \n> Likewise, in contrast to 2005, nowadays implementations of Git in\n> languages as obscure as Javascript are not only theoretical but do \n> exist\n> in practice (https://github.com/creationix/js-git). I had a *very* \n> quick\n> look for libraries providing crypto in Javascript and immediately found\n> the Standford Javascript Crypto library\n> (https://github.com/bitwiseshiftleft/sjcl/) which seems to offer \n> SHA-256\n> but not SHA3-256 computation.\n> \n> Back to Intel processors: I read some vague hints about extensions\n> accelerating SHA-256 computation on future Intel processors, but not\n> SHA3-256.\n> \n> It would make sense, of course, that more crypto libraries and more\n> hardware support would be available for SHA-256 than for SHA3-256 given\n> the time since publication: 16 vs 5 years (I am playing it loose here,\n> taking just the year into account, not the exact date, so please treat\n> that merely as a ballpark figure).\n> \n> So from a practical point of view, I wonder what your take is on, say,\n> hardware support for SHA3-256. Do you think this will become a focus \n> soon?\n> \n> Also, what is your take on the question whether SHA-256 is good enough?\n> SHA-1 was broken theoretically already 10 years after it was published\n> (which unfortunately did not prevent us from baking it into Git), after\n> all, while SHA-256 is 16 years old and the only known weakness does not\n> apply to Git's usage?\n> \n> Also, while I have the attention of somebody who knows a heck more \n> about\n> cryptography than Git's top 10 committers combined: how soon do you \n> expect\n> practical SHA-1 attacks that are much worse than what we already have\n> seen? I am concerned that if we do not move fast enough to a new hash\n> algorithm, and somebody finds a way in the meantime to craft arbitrary\n> messages given a prefix and an SHA-1, then we have a huge problem on\n> our hands.\n> \n> Ciao,\n> Johannes\n\n\n\nBegin forwarded message:\n\n From: Gilles Van Assche <gilles.van.assche@noekeon.org>\nSubject: Re: RFC v3: Another proposed hash function transition plan\nDate: 30 Sep 2017 22:20:42 CEST\nTo: Joan Daemen <joan@cs.ru.nl>, keccak@noekeon.org\n\nDag Joan,\n\nAbout the implementations, there are many high-quality implementations \nof Keccak besides the KCP that you could also mention. E.g., those in \nOpenSSL are very good. And there are all those related to Ethereum.\n\nI tend to agree with Guido regarding SHA-1, even if you are right, there \nis no need to reduce/excuse too much the impact of collisions, there \ncould be unexpected use cases. And it's not clean. (And don't \nunderestimate the probability to be quoted on this.)\n\nFinally, just to say that I like your last paragraph.\n\nKind regards,\nGilles\n\n\n\n\nJoan Daemen <joan@cs.ru.nl> wrote:\nwhat about replying with something like this (please have a critical \nlook). I sent this from my Radboud account as I have problems with my \nThunderbird settings. When trying to send a mail, it sometimes works and \nsometimes it says “An error occurred while sending mail: Outgoing server \n(SMTP) error. The server responded:  4.7.1 <joans-mbp.home>: Helo \ncommand rejected: Host not found.\"\nDear Johannes,\nthanks for your response and taking the effort to express your concerns. \nPlease see below for some feedback.\nOn 30/09/17 00:33, Johannes Schindelin wrote:\nHi Joan,\n\nOn Fri, 29 Sep 2017, Joan Daemen wrote:\n\nif ever there was a SHA-2 competition, it must have been held inside \nNSA:-)\nOops. My bad, I indeed got confused about that, as you suggest below (I\nactually thought of the AES competition, but that was obviously not \nabout\nSHA-2). Sorry.\n\nBut maybe you are confusing with the SHA-3 competition. In any case,\nwhen considering SHA-2 vs SHA-3 for usage in git, you may have a look at\narguments we give in the following blogpost:\n\nhttps://keccak.team/2017/open_source_crypto.html\nThanks for the pointer!\n\nSmall nit: the post uses \"its\" in place of \"it's\", twice.\nThanks, we'll correct that.\n\nIt does have a good point, of course: the scientific exchange (which you\ncall \"open-source\" in spirit) makes tons of sense.\n\nAs far as Git is concerned, we not only care about the source code of \nthe\nhash algorithm we use, we need to care even more about what you call\n\"executable\": ready-to-use, high quality, well-tested implementations.\n\nWe carry source code for SHA-1 as part of Git's source code, which was\nhand-tuned to be as fast as Linus could get it, which was tricky given\nthat the tuning should be general enough to apply to all common intel\nCPUs.\n\nThis hand-crafted code was blown out of the water by OpenSSL's SHA-1 in\nour tests here at Microsoft, thanks to the fact that OpenSSL does\nvectorized SHA-1 computation now.\n\nTo me, this illustrates why it is not good enough to have only a \nreference\nimplementation available at our finger tips. Of course, above-mentioned\nOpenSSL supports SHA-256 and SHA3-256, too, and at least recent versions\nvectorize those, too.\nThere is a lot of high-quality optimized code for all SHA-3 functions \nand many CPUs in the Keccak code package \nhttps://github.com/gvanas/KeccakCodePackage\n\nBy the way, you speak about SHA3-256, but the right choice would be to \nuse SHAKE128. Well, what is exactly the right choice depends on what you \nwant. If you want to have a function in the SHA3 standard (FIPS 202), it \nis SHAKE128. You can boost performance on high-end CPUs by adopting \nParallelhash from NIST SP 800-185, still a NIST standard. You can \nmultiply that performance again by a factor of 2 by adopting \nKangarooTwelve. This is our (Keccak team) proposal for a parallelizable \nKeccak-based hash function that has a safety margin comparable to that \nof the SHA-2 functions. See https://keccak.team/kangarootwelve.html\nMay I also suggest you to read \nhttps://keccak.team/2017/is_sha3_slow.html\n\nAlso, ARM processors have become a lot more popular, so we'll want to \nhave\nhigh-quality implementations of the hash algorithm also for those\nprocessors.\n\nLikewise, in contrast to 2005, nowadays implementations of Git in\nlanguages as obscure as Javascript are not only theoretical but do exist\nin practice (https://github.com/creationix/js-git). I had a *very* quick\nlook for libraries providing crypto in Javascript and immediately found\nthe Standford Javascript Crypto library\n(https://github.com/bitwiseshiftleft/sjcl/) which seems to offer SHA-256\nbut not SHA3-256 computation.\n\nBack to Intel processors: I read some vague hints about extensions\naccelerating SHA-256 computation on future Intel processors, but not\nSHA3-256.\n\nIt would make sense, of course, that more crypto libraries and more\nhardware support would be available for SHA-256 than for SHA3-256 given\nthe time since publication: 16 vs 5 years (I am playing it loose here,\ntaking just the year into account, not the exact date, so please treat\nthat merely as a ballpark figure).\n\nSo from a practical point of view, I wonder what your take is on, say,\nhardware support for SHA3-256. Do you think this will become a focus \nsoon?\nI think this is a chicken-and-egg problem. In any case, hardware support \nfor one SHA3-256 will also work for the other SHA3 and SHAKE functions \nas they all use the same underlying primitive: the Keccak-f permutation. \nThis is not the case for SHA2 because SHA224 and SHA256 use a different \ncompression function than SHA384, SHA512, SHA512/224 and SHA512/256.\n\nAlso, what is your take on the question whether SHA-256 is good enough?\nSHA-1 was broken theoretically already 10 years after it was published\n(which unfortunately did not prevent us from baking it into Git), after\nall, while SHA-256 is 16 years old and the only known weakness does not\napply to Git's usage?\nI think even the weakness of SHA-1 will be hard to exploit to do \nsomething bad in Git. SHA-256 is more conservative than SHA-1 and I \ndon't expect it to be broken (unless NSA inserted a backdoor but I don't \nthink that is likely). But I also don't expect SHAKE128, ParallelHash or \nKangarooTwelve to be broken, looking at the existing cryptanalysis.\nAlso, while I have the attention of somebody who knows a heck more about\ncryptography than Git's top 10 committers combined: how soon do you \nexpect\npractical SHA-1 attacks that are much worse than what we already have\nseen? I am concerned that if we do not move fast enough to a new hash\nalgorithm, and somebody finds a way in the meantime to craft arbitrary\nmessages given a prefix and an SHA-1, then we have a huge problem on\nour hands.\nAs said, I don't expect practical SHA-1 attacks soon. But let me end \nwith a philosophical note. Independent of all the arguments for and \nagainst, I think this is about doing the right thing. The choice is here \nbetween SHA1/SHA2 on the one hand and SHA3/Keccak on the other. The \nformer standards are imposed on us by NSA and the latter are the best \nthat came out of an open competition involving all experts worldwide. \nWhat would be closest to the philosophy of Git (and by extension Linux \nor open-source in general)?\n\nKind regards,\n\nJoan\n\n\n"},{"id":"329423","messageId":"xmqq3772ot1w.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170929173413.GI19555@aiede.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-10-02T08:25:15Z","receivedAt":"2017-10-02T08:25:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n>>> +6. Skip fetching some submodules of a project into a NewHash\n>>> +   repository. (This also depends on NewHash support in Git\n>>> +   protocol.)\n>>\n>> It is unclear what this means.  Around submodule support, one thing\n>> I can think of is that a NewHash tree in a superproject would record\n>> a gitlink that is a NewHash commit object name in it, therefore it\n>> cannot refer to an unconverted SHA-1 submodule repository.  But it\n>> is unclear if the above description refers to the same issue, or\n>> something else.\n>\n> It refers to that issue.\n\nWe may want to find a way to make it clear, then.\n\n>> It makes me wonder if we want to add the hashname in this object\n>> header.  \"length\" would be different for non-blob objects anyway,\n>> and it is not \"compat metadata\" we want to avoid baked in, yet it\n>> would help diagnose a mistake of attempting to use a \"mixed\" objects\n>> in a single repository.  Not a big issue, though.\n>\n> Do you mean that adding the hashname into the computation that\n> produces the object name would help in some use case?\n\nWhat I mean is that for SHA-1 objects we keep the object header to\nbe \"<type> <length> NUL\".  For objects in newer world, use the\nobject header to \"<type> <hash> <length> NUL\", and include the\nhashname in the object name computation.\n\n> For loose objects, it would be nice to name the hash in the file, so\n> that \"file\" can understand what is happening if someone accidentally\n> mixes types using \"cp\".  The only downside is losing the ability to\n> copy blobs (which have the same content despite being named using\n> different hashes) between repositories after determining their new\n> names.  That doesn't seem like a strong downside --- it's pretty\n> harmless to include the hash type in loose object files, too.  I think\n> I would prefer this to be a \"magic number\" instead of part of the\n> zlib-deflated payload, since this way \"file\" can discover it more\n> easily.\n\nYeah, thanks for doing pros-and-cons for me ;-)\n\n>> If it is a goal to eventually be able to lose SHA-1 compatibility\n>> metadata from the objects, then we might want to remove SHA-1 based\n>> signature bits (e.g. PGP trailer in signed tag, gpgsig header in the\n>> commit object) from NewHash contents, and instead have them stored\n>> in a side \"metadata\" table, only to be used while converting back.\n>> I dunno if that is desirable.\n>\n> I don't consider that desirable.\n\nAgreed.  Let's not go there.\n\n>> Hmm, as the corresponding packfile stores object data only in\n>> NewHash content format, it is somewhat curious that this table that\n>> stores CRC32 of the data appears in the \"Tables for each object\n>> format\" section, as they would be identical, no?  Unless I am\n>> grossly misleading the spec, the checksum should either go outside\n>> the \"Tables for each object format\" section but still in .idx, or\n>> should be eliminated and become part of the packdata stream instead,\n>> perhaps?\n>\n> It's actually only present for the first object format.  Will find a\n> better way to describe this.\n\nI see.  One way to do so is to have it upfront before the \"after\nthis point, these tables repeat for each of the hashes\" part of the\nfile.\n\n>> Oy.  So we can go from a short prefix to the pack location by first\n>> finding it via binsearch in the short-name table, realize that it is\n>> nth object in the object name order, and consulting this table.\n>> When we know the pack-order of an object, there is no direct way to\n>> go to its location (short of reversing the name-order-to-pack-order\n>> table)?\n>\n> An earlier version of the design also had a pack-order-to-pack-offset\n> table, but we weren't able to think of any cases where that would be\n> used without also looking up the object name that can be used to\n> verify the integrity of the inflated object.\n\nThe primary thing I was interested in knowing was if we tried to\nthink of any case where it may be useful and then didn't think of\nany---I couldn't but I know I am not imaginative enough, and I\nwanted to know you guys didn't, either.\n"},{"id":"329424","messageId":"xmqqefqlorc2.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170928044320.GA84719@aiede.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-10-02T09:02:21Z","receivedAt":"2017-10-02T09:02:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> +Reading an object's sha1-content\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> +The sha1-content of an object can be read by converting all newhash-names\n> +its newhash-content references to sha1-names using the translation table.\n\nSure.\n\n> +Fetch\n> +~~~~~\n> +Fetching from a SHA-1 based server requires translating between SHA-1\n> +and NewHash based representations on the fly.\n> +\n> +SHA-1s named in the ref advertisement that are present on the client\n> +can be translated to NewHash and looked up as local objects using the\n> +translation table.\n> +\n> +Negotiation proceeds as today. Any \"have\"s generated locally are\n> +converted to SHA-1 before being sent to the server, and SHA-1s\n> +mentioned by the server are converted to NewHash when looking them up\n> +locally.\n\nAny of our alternate object store by definition is a NewHash\nrepository--otherwise we'd violate \"no mixing\" rule.  It may or may\nnote have the translation table for its objects.  If it no longer\nhas the translation table (because it migrated to NewHash only world\nbefore we did), then we can still use it as our alternate but we\ncannot use it for the purpose of common ancestore discovery.\n\n> +After negotiation, the server sends a packfile containing the\n> +requested objects.\n\ns/objects.$/& These are all SHA-1 contents./\n\n> +We convert the packfile to NewHash format using\n> +the following steps:\n> +\n> +1. index-pack: inflate each object in the packfile and compute its\n> +   SHA-1. Objects can contain deltas in OBJ_REF_DELTA format against\n> +   objects the client has locally. These objects can be looked up\n> +   using the translation table and their sha1-content read as\n> +   described above to resolve the deltas.\n\nThat procedure would give us the object's SHA-1 contents for\nref-delta objects.  For an ofs-delta object, by definition, its base\nobject should appear in the same packstream, so we should eventually\nbe able to get to the SHA-1 contents of the delta base, and from\nthere we can apply the delta to obtain the SHA-1 contents.  For a\nnon-delta object, we already have its SHA-1 contents in the\npackstream.\n\nSo we can get SHA-1 names and SHA-1 contents of each and every\nobject in the packstream in this step.\n\nAre we actually writing out a .pack/.idx pair that is usable in the\nSHA-1 world at this stage?  Or are we going to read from something\nwe keep in-core in the step #3 below?\n\n> +2. topological sort: starting at the \"want\"s from the negotiation\n> +   phase, walk through objects in the pack and emit a list of them,\n> +   excluding blobs, in reverse topologically sorted order, with each\n> +   object coming later in the list than all objects it references.\n> +   (This list only contains objects reachable from the \"wants\". If the\n> +   pack from the server contained additional extraneous objects, then\n> +   they will be discarded.)\n\nPresumably this is a list of SHA-1 names, as we do not yet have\nenough information to compute NewHash names yet at this point.  May\nwant to spell it out here.\n\nWould it discard the auto-followed tags if we do the \"traverse from\nwants only\"?  Traversing the objects in the packfile to find the\n\"tips\" that are not referenced from any other object in the pack\nmight be necessary, and it shouldn't be too costly, I'd guess.\n\n> +3. convert to newhash: open a new (newhash) packfile. Read the topologically\n> +   sorted list just generated. For each object, inflate its\n> +   sha1-content, convert to newhash-content, and write it to the newhash\n> +   pack. Record the new sha1<->newhash mapping entry for use in the idx.\n\nAre we doing any deltification here?  If we are computing .pack/.idx\npair that can be usable in the SHA-1 world in step #1, then reusing\nblob deltas should be trivial (a good delta-base in the SHA-1 world\nis a good delta-base in the NewHash world, too).  Things that have\noutgoing references like trees, it might be possible that such a\nheuristic may not give us the absolute best delta-base, but I guess\nit would still be a good approximation to reuse the delta/base\nobject relationship in SHA-1 world to NewHash world, assuming that\nthe server did a good job choosing the bases.\n\n> +4. sort: reorder entries in the new pack to match the order of objects\n> +   in the pack the server generated and include blobs. Write a newhash idx\n> +   file\n\nOK.\n\n> +5. clean up: remove the SHA-1 based pack file, index, and\n> +   topologically sorted list obtained from the server in steps 1\n> +   and 2.\n\nAh, OK, so we do write the SHA_1 pack/idx in the first step.  OK.\n\n> +Push\n> +~~~~\n> +Push is simpler than fetch because the objects referenced by the\n> +pushed objects are already in the translation table. The sha1-content\n> +of each object being pushed can be read as described in the \"Reading\n> +an object's sha1-content\" section to generate the pack written by git\n> +send-pack.\n\nOK.\n\n> +Signed Commits\n> +~~~~~~~~~~~~~~\n> +We add a new field \"gpgsig-newhash\" to the commit object format to allow\n> +signing commits without relying on SHA-1. It is similar to the\n> +existing \"gpgsig\" field. Its signed payload is the newhash-content of the\n> +commit object with any \"gpgsig\" and \"gpgsig-newhash\" fields removed.\n\nDo we prepare for newerhash, too?  IOW, should the signed payload be\nthe newhash-contents with any field whose name is \"gpgsig\" or begins\nwith \"gpgsig-\" followed by anything?\n\n> +This means commits can be signed\n> +1. using SHA-1 only, as in existing signed commit objects\n> +2. using both SHA-1 and NewHash, by using both gpgsig-newhash and gpgsig\n> +   fields.\n> +3. using only NewHash, by only using the gpgsig-newhash field.\n> +\n> +Old versions of \"git verify-commit\" can verify the gpgsig signature in\n> +cases (1) and (2) without modifications and view case (3) as an\n> +ordinary unsigned commit.\n\nFor old clients to be able to verify (2), signed payload for SHA-1\nis everything in SHA-1 contents minus \"gpgsig\"; \"gpgsig-newhash\"\nshould not get excluded from the computation.  Am I correct?\n\nI am primarily finding it a bit disturbing that there is a bit of\nasymmetry here.\n\n> +Signed Tags\n> +~~~~~~~~~~~\n\nThis message stops here for now.\n"},{"id":"329444","messageId":"20171002140011.GE31762@io.lakedaemon.net","threadId":"45288","inReplyTo":"alpine.DEB.2.21.1.1709262356360.40514@virtualbox","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jason Cooper","fromEmail":"jason@lakedaemon.net","sentAt":"2017-10-02T14:00:11Z","receivedAt":"2017-10-02T14:00:31Z","isPatch":false,"sender":{"key":"jason@lakedaemon.net","avatar":null},"body":"Hi Johannes,\n\nThanks for the response.  Sorry for the delay.  Had a large deadline for\n$dayjob.\n\nOn Wed, Sep 27, 2017 at 12:11:14AM +0200, Johannes Schindelin wrote:\n> On Tue, 26 Sep 2017, Jason Cooper wrote:\n> > On Thu, Sep 14, 2017 at 08:45:35PM +0200, Johannes Schindelin wrote:\n> > > On Wed, 13 Sep 2017, Linus Torvalds wrote:\n> > > > On Wed, Sep 13, 2017 at 6:43 AM, demerphq <demerphq@gmail.com> wrote:\n> > > > > SHA3 however uses a completely different design where it mixes a 1088\n> > > > > bit block into a 1600 bit state, for a leverage of 2:3, and the excess\n> > > > > is *preserved between each block*.\n> > > > \n> > > > Yes. And considering that the SHA1 attack was actually predicated on\n> > > > the fact that each block was independent (no extra state between), I\n> > > > do think SHA3 is a better model.\n> > > > \n> > > > So I'd rather see SHA3-256 than SHA256.\n> > \n> > Well, for what it's worth, we need to be aware that SHA3 is *different*.\n> > In crypto, \"different\" = \"bugs haven't been found yet\".  :-P\n> > \n> > And SHA2 is *known*.  So we have a pretty good handle on how it'll\n> > weaken over time.\n> \n> Here, you seem to agree with me.\n\nYep.\n\n> > > SHA-256 got much more cryptanalysis than SHA3-256, and apart from the\n> > > length-extension problem that does not affect Git's usage, there are no\n> > > known weaknesses so far.\n> > \n> > While I think that statement is true on it's face (particularly when\n> > including post-competition analysis), I don't think it's sufficient\n> > justification to chose one over the other.\n> \n> And here you don't.\n> \n> I find that very confusing.\n\nWhat I'm saying is that there is more to selecting a hash function for\ngit than just the cryptographic assessment.  In fact I would argue that\nthe primary cryptographic concern for git is \"What is the likelihood\nthat we'll wake up one day to full collisions with no warning?\"\n\nTo that, I'd argue that SHA-256's time in the field and SHA3-256's\ncompetition give them both passing marks in that regard.  fwiw, I'd also\nput Blake and Skein in there as well.\n\nThe chance that any of those will suffer sudden, catastrophic failure is\nminimal.  IOW, we'll have warnings, and time to migrate to the next\nfunction.\n\nNone of us can predict the future, but having a significant amount of\nvetting reduces the chances of catastrophic failure.\n\n> > > It would seem that the experts I talked to were much more concerned about\n> > > that amount of attention than the particulars of the algorithm. My\n> > > impression was that the new features of SHA3 were less studied than the\n> > > well-known features of SHA2, and that the new-ness of SHA3 is not\n> > > necessarily a good thing.\n> > \n> > The only thing I really object to here is the abstract \"experts\".  We're\n> > talking about cryptography and integrity here.  It's no longer\n> > sufficient to cite anonymous experts.  Either they can put their\n> > thoughts, opinions and analysis on record here, or it shouldn't be\n> > considered.  Sorry.\n> \n> Sorry, you are asking cryptography experts to spend their time on the Git\n> mailing list. I tried to get them to speak out on the Git mailing list.\n> They respectfully declined.\n\nOk, fair enough.  Just please understand that it's difficult to place\nmuch weight on statements that we can't discuss with the person who made\nthem.\n\n> > However, whether we chose SHA2 or SHA3 doesn't matter.\n> \n> To you, it does not matter.\n\nWell, I'd say it does not matter for *most* users.\n\n> To me, it matters. To the several thousand developers working on Windows,\n> probably the largest Git repository in active use, it matters. It matters\n> because the speed difference that has little impact on you has a lot more\n> impact on us.\n\nAhhh, so if I understand you correctly, you'd prefer SHA-256 over\nSHA3-256 because it's more performant for your usecase?  Well, that's a\ncompletely different animal that cryptographic suitability.\n\nHave you been able to crunch numbers yet?  Will you be able to share\nsome empirical data?  I'd love to see some comparisons between SHA1,\nSHA-256, SHA512-256, and SHA3-256 for different git operations under\nyour work load.\n\n> > If SHA3 is chosen as the successor, it's going to get a *lot* more\n> > adoption, and thus, a lot more analysis.  If cracks start to show, the\n> > hard work of making git flexible is already done.  We can migrate to\n> > SHA4/5/whatever in an orderly fashion with far less effort than the\n> > transition away from SHA1.\n> \n> Sure. And if XYZ789 is chosen, it's going to get a *lot* more adoption,\n> too.\n> \n> We think.\n> \n> Let's be realistic. Git is pretty important to us, but it is not important\n> enough to sway, say, Intel into announcing hardware support for SHA3.\n> And if you try to force through *any* hash function only so that it gets\n> more adoption and hence more support,\n\nThat's quite a jump from what I was saying.  I would never advise using\ncode in a production setting just to increase adoption.\n\nWhat I /was/ saying: Let's say you don't get what you want, and SHA3-256\nis chosen.  It's not the end of the world from a cryptographic PoV.\nThe hard work of making the git (and libgit2) codebases hash-flexible is\nalready done.  So, if you're correct, and SHA3 was too immature, the\nincreased visibility will help us discover that more quickly.  And, the\ncode will already be in a position to conduct an orderly migration.\n\nWill it still be costly?  Yes.  But I would argue that it's naive to\nthink that we will be using git/sha3-256 or git/sha-256 10 to 15 years\nfrom now.  It might be git, it might not.  But there *will* be another\nmigration of existing data (code, history, etc) from one object storage\nmodel to another.  It might be git/SHA4-512, or hg/sha4-384.\n\nSo, we aren't trying to find the perfect hash function so that we\nnaively think we'll never have to change again.  Rather, we're choosing\nthe next hash function so that we can hold off another migration for as\nlong as possible.  After all, SHA4-512 doesn't exist yet. ;-)\n\n> in the short run you will make life\n> harder for developers on more obscure platforms, who may not easily get\n> high-quality, high-speed implementations of anything but the very\n> mainstream (which is, let's face it, MD5, SHA-1 and SHA-256). I know I\n> would have cursed you for such a decision back when I had to work on AIX\n> and IRIX.\n\nI think you're assuming that all developers on obscure platforms have\na similar git usecase to your current one.  I've not heard of that being\nthe case.\n\n> > For my use cases, as a user of git, I have a plan to maintain provable\n> > integrity of existing objects stored in git under sha1 while migrating\n> > away from sha1.  The same plan works for migrating away from SHA2 or\n> > SHA3 when the time comes.\n> \n> Please do not make the mistake of taking your use case to be a template\n> for everybody's use case.\n\nI wasn't.  But I will argue that my usecase is valid.  Just as yours is.\n\n> Migrating a large team away from any hash function to another one *will*\n> be painful, and costly.\n\nAssuming that it will never happen again would make that doubly costly.\n\n> Migrating will be very costly for hosting companies like GitHub, Microsoft\n> and BitBucket, too.\n\n<with_my_business_hat_on>\nGitHub and BitBucket have git as the core of their business model.  If\nthey aren't keeping an eye on the future path of git and maintaining\nmigration plans, shame on them.\n</with_my_business_hat_on>\n\nThanks,\n\nJason.\n"},{"id":"329447","messageId":"alpine.DEB.2.21.1.1710021601380.40514@virtualbox","threadId":"45288","inReplyTo":"6d96ec902dc1d500ba3fcb11d31b2015@mail.noekeon.org","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2017-10-02T14:26:49Z","receivedAt":"2017-10-02T14:27:39Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi Joan,\n\nOn Sun, 1 Oct 2017, Joan Daemen wrote:\n\n> On 30/09/17 00:33, Johannes Schindelin wrote:\n> \n> > As far as Git is concerned, we not only care about the source code of\n> > the hash algorithm we use, we need to care even more about what you\n> > call \"executable\": ready-to-use, high quality, well-tested\n> > implementations.\n> > \n> > We carry source code for SHA-1 as part of Git's source code, which was\n> > hand-tuned to be as fast as Linus could get it, which was tricky given\n> > that the tuning should be general enough to apply to all common intel\n> > CPUs.\n> > \n> > This hand-crafted code was blown out of the water by OpenSSL's SHA-1\n> > in our tests here at Microsoft, thanks to the fact that OpenSSL does\n> > vectorized SHA-1 computation now.\n> > \n> > To me, this illustrates why it is not good enough to have only a\n> > reference implementation available at our finger tips. Of course,\n> > above-mentioned OpenSSL supports SHA-256 and SHA3-256, too, and at\n> > least recent versions vectorize those, too.\n> \n> There is a lot of high-quality optimized code for all SHA-3 functions\n> and many CPUs in the Keccak code package\n> https://github.com/gvanas/KeccakCodePackage but also OpenSSL contains\n> some good SHA-3 code and then there are all those related to Ethereum.\n> \n> By the way, you speak about SHA3-256, but the right choice would be to\n> use SHAKE128. Well, what is exactly the right choice depends on what you\n> want. If you want to have a function in the SHA3 standard (FIPS 202), it\n> is SHAKE128.  You can boost performance on high-end CPUs by adopting\n> Parallelhash from NIST SP 800-185, still a NIST standard. You can\n> multiply that performance again by a factor of 2 by adopting\n> KangarooTwelve. This is our (Keccak team) proposal for a parallelizable\n> Keccak-based hash function that has a safety margin comparable to that\n> of the SHA-2 functions. See https://keccak.team/kangarootwelve.html May\n> I also suggest you read https://keccak.team/2017/is_sha3_slow.html\n\nThanks.\n\nI have to admit that all those names that do not start with SHA and do not\nend in 256 make me a bit dizzy.\n\n> > Back to Intel processors: I read some vague hints about extensions\n> > accelerating SHA-256 computation on future Intel processors, but not\n> > SHA3-256.\n> > \n> > It would make sense, of course, that more crypto libraries and more\n> > hardware support would be available for SHA-256 than for SHA3-256\n> > given the time since publication: 16 vs 5 years (I am playing it loose\n> > here, taking just the year into account, not the exact date, so please\n> > treat that merely as a ballpark figure).\n> > \n> > So from a practical point of view, I wonder what your take is on, say,\n> > hardware support for SHA3-256. Do you think this will become a focus\n> > soon?\n> \n> I think this is a chicken-and-egg problem. In any case, hardware support\n> for one SHA3-256 will also work for the other SHA3 and SHAKE functions\n> as they all use the same underlying primitive: the Keccak-f permutation.\n> This is not the case for SHA2 because SHA224 and SHA256 use a different\n> compression function than SHA384, SHA512, SHA512/224 and SHA512/256.\n\nOkay.\n\nSo given that Git does not exactly have a big sway on hardware vendors, we\nwould have to hope that some other chicken lays that egg.\n\n> > Also, what is your take on the question whether SHA-256 is good\n> > enough?  SHA-1 was broken theoretically already 10 years after it was\n> > published (which unfortunately did not prevent us from baking it into\n> > Git), after all, while SHA-256 is 16 years old and the only known\n> > weakness does not apply to Git's usage?\n> \n> SHA-256 is more conservative than SHA-1 and I don't expect it to be\n> broken in the coming decades (unless NSA inserted a backdoor but I don't\n> think that is likely). But looking at the existing cryptanalysis, I\n> think it is even less likely that I SHAKE128, ParallelHash or\n> KangarooTwelve will be broken anytime.\n\nThat's reassuring! ;-)\n\n> > Also, while I have the attention of somebody who knows a heck more\n> > about cryptography than Git's top 10 committers combined: how soon do\n> > you expect practical SHA-1 attacks that are much worse than what we\n> > already have seen? I am concerned that if we do not move fast enough\n> > to a new hash algorithm, and somebody finds a way in the meantime to\n> > craft arbitrary messages given a prefix and an SHA-1, then we have a\n> > huge problem on our hands.\n> \n> This is hard to say. To be honest, when witnessing the first MD5\n> collisions I did not expect them to lead to some real world attacks and\n> just a few years later we saw real-world forged certificates based on\n> MD5 collisions. And SHA-1 has a lot in common with MD5...\n\nOh, okay. I did not realize that MD5 and SHA-1 are so similar in design,\nthank you for educating me!\n\n> But let me end with a philosophical note. Independent of all the\n> arguments for and against, I think this is ultimately about doing the\n> right thing. The choice is here between SHA1/SHA2 on the one hand and\n> SHA3/Keccak on the other.  The former standards are imposed on us by NSA\n> and the latter are the best that came out of an open competition\n> involving all experts in the field worldwide.  What would be closest to\n> the philosophy of Git (and by extension Linux or open-source in\n> general)?\n\nHeh. Do you realize that you are talking to a Microsoftie, i.e. one of the\n\"evil company\"? ;-)\n\nSo philosophically, I am much more pragmatic. Or maybe I am not, after\nall, I joined a company at a time when it is arguably going through one of\nthe most dramatic cultural changes any company has seen lately (a year\nago, we became #1 contributor on GitHub according to Business Insider, and\nas far as I can tell, we're not willing to pass that belt to anyone else).\n\nBut when it comes to the philosophy of Git, I fear I have to disappoint\nyou: Git's fundamental concepts were not developed in an open process. Git\neven so much as rejected professional advice *not* to bake SHA-1 into\neverything.\n\nOf course, we are undoing this damage right now, and your input helps\ngreatly, I would think.\n\nWhile I feel reassured by your response that SHA-256 would be \"good\nenough\" and would have some real-life benefits of announced hardware\nsupport, I would now also feel comfortable if my preference was overruled\nin the end, in favor of a hash from the Keccak family. I would understand,\nfor example, if the parallel option turned out to be enticing enough for\nother core Git contributors to aim for, say, K12).\n\nAgain, thank you very much for chiming in,\nJohannes\n"},{"id":"329456","messageId":"20171002145400.GF31762@io.lakedaemon.net","threadId":"45288","inReplyTo":"20170926235158.GD19555@aiede.mtv.corp.google.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jason Cooper","fromEmail":"jason@lakedaemon.net","sentAt":"2017-10-02T14:54:00Z","receivedAt":"2017-10-02T15:10:16Z","isPatch":false,"sender":{"key":"jason@lakedaemon.net","avatar":null},"body":"Hi Jonathan,\n\nOn Tue, Sep 26, 2017 at 04:51:58PM -0700, Jonathan Nieder wrote:\n> Johannes Schindelin wrote:\n> > On Tue, 26 Sep 2017, Jason Cooper wrote:\n> >> For my use cases, as a user of git, I have a plan to maintain provable\n> >> integrity of existing objects stored in git under sha1 while migrating\n> >> away from sha1.  The same plan works for migrating away from SHA2 or\n> >> SHA3 when the time comes.\n> >\n> > Please do not make the mistake of taking your use case to be a template\n> > for everybody's use case.\n> \n> That said, I'm curious at what plan you are alluding to.  Is it\n> something that could benefit others on the list?\n\nWell, it's just a plan at this point.  As there's a lot of other work to\ndo in the mean-time, and there's no possibility of transitioning until\nthe dust has settled on NEWHASH.  :-)\n\nGiven an existing repository that needs to migrate from SHA1 to NEWHASH,\nand maintain backwards compatibility with clients that haven't migrated\nyet, how do we\n\n  a) perform that migration,\n  b) allow non-updated clients to use the data prior to the switch, and\n  c) maintain provable integrity of the old objects as well as the new.\n\nThe primary method is counter-hashing, which re-uses the blobs, and\ncreates parallel, deterministic tree, commit, and tag objects using\nNEWHASH for everything up to flag day.  post-flag-day only uses NEWHASH.\nA PGP \"transition\" key is used to counter-sign the NEWHASH version of\nthe old signed tags.  The transition key is not required to be different\nthan the existing maintainers key.\n\nA critical feature is the ability of entities other than the maintainer\nto migrate to NEWHASH.  For example, let's say that git has fully\nimplemented and tested NEWHASH.  linux.git intends to migrate, but it's\ngoing to take several months (get all the developers herded up).\n\nIn the interim, a security company, relying on Linux for it's products\ncan counter-hash Linus' repo, and continue to do so every time he\nupdates his tree.  This shrinks the attack window for an entity (with an\nundisclosed break of SHA1) down to a few minutes to an hour.  Otherwise,\na check of the counter hashes in the future would reveal the\nsubstitution.\n\nThe deterministic feature is critical here because there is valuable\nintegrity and trust built by counter-hashing quickly after publication.\nSo once Linux migrates to NEWHASH, the hashes calculated by the security\ncompany should be identical.  IOW, use the timestamps that are in the\nSHA1 commit objects for the NEWHASH objects.  Which should be obvious,\nbut it's worth explicitly mentioning that determinism provides great\nvalue.\n\nWe're in the process of writing this up formally, which will provide a\nlot more detail and rationale that this quick stream of thought.  :-)\n\nI'm sure a lot of this has already been discussed on the list.  If so, I\napologize for being repetitive.  Unfortunately, I'm not able to keep up\nwith the MLs like I used to.\n\nthx,\n\nJason.\n"},{"id":"329461","messageId":"20171002165030.GA5189@google.com","threadId":"45288","inReplyTo":"20171002145400.GF31762@io.lakedaemon.net","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Brandon Williams","fromEmail":"bmwill@google.com","sentAt":"2017-10-02T16:50:30Z","receivedAt":"2017-10-02T16:50:39Z","isPatch":false,"sender":{"key":"bwilliams.eng@gmail.com","avatar":null},"body":"On 10/02, Jason Cooper wrote:\n> Hi Jonathan,\n> \n> On Tue, Sep 26, 2017 at 04:51:58PM -0700, Jonathan Nieder wrote:\n> > Johannes Schindelin wrote:\n> > > On Tue, 26 Sep 2017, Jason Cooper wrote:\n> > >> For my use cases, as a user of git, I have a plan to maintain provable\n> > >> integrity of existing objects stored in git under sha1 while migrating\n> > >> away from sha1.  The same plan works for migrating away from SHA2 or\n> > >> SHA3 when the time comes.\n> > >\n> > > Please do not make the mistake of taking your use case to be a template\n> > > for everybody's use case.\n> > \n> > That said, I'm curious at what plan you are alluding to.  Is it\n> > something that could benefit others on the list?\n> \n> Well, it's just a plan at this point.  As there's a lot of other work to\n> do in the mean-time, and there's no possibility of transitioning until\n> the dust has settled on NEWHASH.  :-)\n> \n> Given an existing repository that needs to migrate from SHA1 to NEWHASH,\n> and maintain backwards compatibility with clients that haven't migrated\n> yet, how do we\n> \n>   a) perform that migration,\n>   b) allow non-updated clients to use the data prior to the switch, and\n>   c) maintain provable integrity of the old objects as well as the new.\n> \n> The primary method is counter-hashing, which re-uses the blobs, and\n> creates parallel, deterministic tree, commit, and tag objects using\n> NEWHASH for everything up to flag day.  post-flag-day only uses NEWHASH.\n> A PGP \"transition\" key is used to counter-sign the NEWHASH version of\n> the old signed tags.  The transition key is not required to be different\n> than the existing maintainers key.\n> \n> A critical feature is the ability of entities other than the maintainer\n> to migrate to NEWHASH.  For example, let's say that git has fully\n> implemented and tested NEWHASH.  linux.git intends to migrate, but it's\n> going to take several months (get all the developers herded up).\n> \n> In the interim, a security company, relying on Linux for it's products\n> can counter-hash Linus' repo, and continue to do so every time he\n> updates his tree.  This shrinks the attack window for an entity (with an\n> undisclosed break of SHA1) down to a few minutes to an hour.  Otherwise,\n> a check of the counter hashes in the future would reveal the\n> substitution.\n> \n> The deterministic feature is critical here because there is valuable\n> integrity and trust built by counter-hashing quickly after publication.\n> So once Linux migrates to NEWHASH, the hashes calculated by the security\n> company should be identical.  IOW, use the timestamps that are in the\n> SHA1 commit objects for the NEWHASH objects.  Which should be obvious,\n> but it's worth explicitly mentioning that determinism provides great\n> value.\n> \n> We're in the process of writing this up formally, which will provide a\n> lot more detail and rationale that this quick stream of thought.  :-)\n> \n> I'm sure a lot of this has already been discussed on the list.  If so, I\n> apologize for being repetitive.  Unfortunately, I'm not able to keep up\n> with the MLs like I used to.\n> \n> thx,\n> \n> Jason.\n\nGiven the interests that you've expressed here I'd recommend taking a\nlook at\nhttps://public-inbox.org/git/20170928044320.GA84719@aiede.mtv.corp.google.com/\nwhich is the current version of the transition plan that the community\nhas settled on\n(https://public-inbox.org/git/xmqqlgkyxgvq.fsf@gitster.mtv.corp.google.com/\nshows that it should be merged to 'next' soon).  Once neat aspect of\nthis transition plan is that it doesn't require a flag day but rather\nanyone can migrate to the new hash function and still interact with\nrepositories (via the wire) which are still running SHA1.\n\n-- \nBrandon Williams\n"},{"id":"329465","messageId":"CA+55aFyb=h1V-3tkESY8jkc356k5rcQRmjr6o_8p6ZgKMp=Jag@mail.gmail.com","threadId":"45288","inReplyTo":"20171002140011.GE31762@io.lakedaemon.net","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2017-10-02T17:18:02Z","receivedAt":"2017-10-02T17:18:09Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"On Mon, Oct 2, 2017 at 7:00 AM, Jason Cooper <jason@lakedaemon.net> wrote:\n>\n> Ahhh, so if I understand you correctly, you'd prefer SHA-256 over\n> SHA3-256 because it's more performant for your usecase?  Well, that's a\n> completely different animal that cryptographic suitability.\n\nIn almost all loads I've seen, zlib inflate() cost is a bigger deal\nthan the crypto load. The crypto people talk about cycles per byte,\nbut the deflate code is what usually takes the page faults and cache\nmisses etc, and has bad branch prediction. That ends up easily being\ntens or thousands of cycles, even for small data.\n\nBut it does obviously depend on exactly what you do. The Windows\npeople saw SHA1 as costly mainly due to the index file (which is just\na \"fancy crc\", and not even cryptographically important, and where the\ncache misses actually happen when doing crypto, not decompressing the\ndata).\n\nAnd fsck and big initial checkins can have a very different profile\nthan most \"regular use\" profiles. Again, there the crypto happens\nfirst, and takes the cache misses. And the crypto is almost certainly\n_much_ cheaper than just the act of loading the index file contents in\nthe first place. It may show up on profiles fairly clearly, but that's\nmostly because crypto is *intensive*, not because crypto takes up most\nof the cycles.\n\nEnd result: honestly, the real cost on almost any load is not crypto\nor necessarily even (de)compression, even if those are the things that\nshow up. It's the cache misses and the \"get data into user space\"\n(whether using \"read()\" or page faulting). Worrying about cycles per\nbyte of compression speed is almost certainly missing the real issue.\n\nThe people who benchmark cryptography tend to intentionally avoid the\nactual real work, because they just want to know the crypto costs. So\nwhen you see numbers like \"9 cycles per byte\" vs \"12 cycles per byte\"\nand think that it's a big deal - 30% performance difference! -  it's\nalmost certainly complete garbage. It may be 30%, but it is likely 30%\nout of 10% total, meaning that it's almost in the noise for any but\nsome very special case.\n\n                 Linus\n"},{"id":"329476","messageId":"20171002192333.GH31762@io.lakedaemon.net","threadId":"45288","inReplyTo":"20170928044320.GA84719@aiede.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Jason Cooper","fromEmail":"jason@lakedaemon.net","sentAt":"2017-10-02T19:23:33Z","receivedAt":"2017-10-02T19:23:55Z","isPatch":true,"sender":{"key":"jason@lakedaemon.net","avatar":null},"body":"Hi Jonathan,\n\nOn Wed, Sep 27, 2017 at 09:43:21PM -0700, Jonathan Nieder wrote:\n> This document describes what a transition to a new hash function for\n> Git would look like.  Add it to Documentation/technical/ as the plan\n> of record so that future changes can be recorded as patches.\n> \n> Also-by: Brandon Williams <bmwill@google.com>\n> Also-by: Jonathan Tan <jonathantanmy@google.com>\n> Also-by: Stefan Beller <sbeller@google.com>\n> Signed-off-by: Jonathan Nieder <jrnieder@gmail.com>\n> ---\n> On Thu, Mar 09, 2017 at 11:14 AM, Shawn Pearce wrote:\n> > On Mon, Mar 6, 2017 at 4:17 PM, Jonathan Nieder <jrnieder@gmail.com> wrote:\n> \n> >> Thanks for the kind words on what had quite a few flaws still.  Here's\n> >> a new draft.  I think the next version will be a patch against\n> >> Documentation/technical/.\n> >\n> > FWIW, I like this approach.\n> \n> Okay, here goes.\n> \n> Instead of sharding the loose object translation tables by first byte,\n> we went for a single table.  It simplifies the design and we need to\n> keep the number of loose objects under control anyway.\n> \n> We also included a description of the transition plan and tried to\n> include a summary of what has been agreed upon so far about the choice\n> of hash function.\n> \n> Thanks to Junio for reviving the discussion and in particular to Dscho\n> for pushing this forward and making the missing pieces clearer.\n> \n> Thoughts of all kinds welcome, as always.\n> \n>  Documentation/Makefile                             |   1 +\n>  .../technical/hash-function-transition.txt         | 797 +++++++++++++++++++++\n>  2 files changed, 798 insertions(+)\n>  create mode 100644 Documentation/technical/hash-function-transition.txt\n> \n...\n> diff --git a/Documentation/technical/hash-function-transition.txt b/Documentation/technical/hash-function-transition.txt\n> new file mode 100644\n> index 0000000000..417ba491d0\n> --- /dev/null\n> +++ b/Documentation/technical/hash-function-transition.txt\n> @@ -0,0 +1,797 @@\n> +Git hash function transition\n> +============================\n> +\n> +Objective\n> +---------\n> +Migrate Git from SHA-1 to a stronger hash function.\n> +\n...\n> +Goals\n> +-----\n> +Where NewHash is a strong 256-bit hash function to replace SHA-1 (see\n> +\"Selection of a New Hash\", below):\n\nCould we clarify and say \"a strong hash function with 256-bit output\"?\n\n...\n> +Overview\n> +--------\n> +We introduce a new repository format extension. Repositories with this\n> +extension enabled use NewHash instead of SHA-1 to name their objects.\n> +This affects both object names and object content --- both the names\n> +of objects and all references to other objects within an object are\n> +switched to the new hash function.\n> +\n> +NewHash repositories cannot be read by older versions of Git.\n> +\n> +Alongside the packfile, a NewHash repository stores a bidirectional\n> +mapping between NewHash and SHA-1 object names. The mapping is generated\n> +locally and can be verified using \"git fsck\". Object lookups use this\n> +mapping to allow naming objects using either their SHA-1 and NewHash names\n> +interchangeably.\n\nnit: Are we presuming that abbreviated hashes won't collide?  Or the\nuser needs to specify which hash type?\n\n> +Object format\n> +~~~~~~~~~~~~~\n> +The content as a byte sequence of a tag, commit, or tree object named\n> +by sha1 and newhash differ because an object named by newhash-name refers to\n> +other objects by their newhash-names and an object named by sha1-name\n> +refers to other objects by their sha1-names.\n> +\n> +The newhash-content of an object is the same as its sha1-content, except\n> +that objects referenced by the object are named using their newhash-names\n> +instead of sha1-names. Because a blob object does not refer to any\n> +other object, its sha1-content and newhash-content are the same.\n> +\n> +The format allows round-trip conversion between newhash-content and\n> +sha1-content.\n\nIt would be nice here to explicitly mention deterministic hashing.\nMeaning that anyone who converts a commit from sha1 to newhash shall get\nthe same newhash.\n\n> +\n> +Object storage\n> +~~~~~~~~~~~~~~\n> +Loose objects use zlib compression and packed objects use the packed\n> +format described in Documentation/technical/pack-format.txt, just like\n> +today. The content that is compressed and stored uses newhash-content\n> +instead of sha1-content.\n> +\n> +Pack index\n> +~~~~~~~~~~\n> +Pack index (.idx) files use a new v3 format that supports multiple\n> +hash functions. They have the following format (all integers are in\n> +network byte order):\n> +\n> +- A header appears at the beginning and consists of the following:\n> +  - The 4-byte pack index signature: '\\377t0c'\n> +  - 4-byte version number: 3\n> +  - 4-byte length of the header section, including the signature and\n> +    version number\n> +  - 4-byte number of objects contained in the pack\n> +  - 4-byte number of object formats in this pack index: 2\n> +  - For each object format:\n> +    - 4-byte format identifier (e.g., 'sha1' for SHA-1)\n\nThis seems a little rough to me.  Maybe it would be better to have a 4\nbyte field where 0x01 = SHA-1, 0x02 = NEWHASH?\n\n> +Reading an object's sha1-content\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> +The sha1-content of an object can be read by converting all newhash-names\n> +its newhash-content references to sha1-names using the translation table.\n> +\n> +Fetch\n> +~~~~~\n> +Fetching from a SHA-1 based server requires translating between SHA-1\n> +and NewHash based representations on the fly.\n> +\n> +SHA-1s named in the ref advertisement that are present on the client\n> +can be translated to NewHash and looked up as local objects using the\n> +translation table.\n> +\n> +Negotiation proceeds as today. Any \"have\"s generated locally are\n> +converted to SHA-1 before being sent to the server, and SHA-1s\n> +mentioned by the server are converted to NewHash when looking them up\n> +locally.\n\nBy \"converted\", do you mean \"looked up in the table\" or \"look up\nnewhash, re-calculate sha1, send\" ?  I presume you mean the former, but\nit would be good to clarify.\n\n> +\n> +After negotiation, the server sends a packfile containing the\n> +requested objects. We convert the packfile to NewHash format using\n> +the following steps:\n> +\n> +1. index-pack: inflate each object in the packfile and compute its\n> +   SHA-1. Objects can contain deltas in OBJ_REF_DELTA format against\n> +   objects the client has locally. These objects can be looked up\n> +   using the translation table and their sha1-content read as\n> +   described above to resolve the deltas.\n> +2. topological sort: starting at the \"want\"s from the negotiation\n> +   phase, walk through objects in the pack and emit a list of them,\n> +   excluding blobs, in reverse topologically sorted order, with each\n> +   object coming later in the list than all objects it references.\n> +   (This list only contains objects reachable from the \"wants\". If the\n> +   pack from the server contained additional extraneous objects, then\n> +   they will be discarded.)\n> +3. convert to newhash: open a new (newhash) packfile. Read the topologically\n> +   sorted list just generated. For each object, inflate its\n> +   sha1-content, convert to newhash-content, and write it to the newhash\n> +   pack. Record the new sha1<->newhash mapping entry for use in the idx.\n> +4. sort: reorder entries in the new pack to match the order of objects\n> +   in the pack the server generated and include blobs. Write a newhash idx\n> +   file\n> +5. clean up: remove the SHA-1 based pack file, index, and\n> +   topologically sorted list obtained from the server in steps 1\n> +   and 2.\n\nHow are signed tags (against sha1 commits) to be handled?  See below for\nfurther thoughts.\n\n> +Signed Tags\n> +~~~~~~~~~~~\n> +We add a new field \"gpgsig-newhash\" to the tag object format to allow\n> +signing tags without relying on SHA-1. Its signed payload is the\n> +newhash-content of the tag with its gpgsig-newhash field and \"-----BEGIN PGP\n> +SIGNATURE-----\" delimited in-body signature removed.\n> +\n> +This means tags can be signed\n> +1. using SHA-1 only, as in existing signed tag objects\n> +2. using both SHA-1 and NewHash, by using gpgsig-newhash and an in-body\n> +   signature.\n> +3. using only NewHash, by only using the gpgsig-newhash field.\n\nTo be clear here, \"gpgsig\" = SHA-1, \"gpgsig-SHA-256\" = SHA-256?\n\n> +Caveats\n> +-------\n> +Invalid objects\n> +~~~~~~~~~~~~~~~\n> +The conversion from sha1-content to newhash-content retains any\n> +brokenness in the original object (e.g., tree entry modes encoded with\n> +leading 0, tree objects whose paths are not sorted correctly, and\n> +commit objects without an author or committer). This is a deliberate\n> +feature of the design to allow the conversion to round-trip.\n\nAh, so this is part of the deterministic hashing.\n\n> +Object names on the command line\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> +To support the transition (see Transition plan below), this design\n> +supports four different modes of operation:\n> +\n> + 1. (\"dark launch\") Treat object names input by the user as SHA-1 and\n> +    convert any object names written to output to SHA-1, but store\n> +    objects using NewHash.  This allows users to test the code with no\n> +    visible behavior change except for performance.  This allows\n> +    allows running even tests that assume the SHA-1 hash function, to\n\nnit:  s/allows allows/allows/\n\n> +    sanity-check the behavior of the new mode.\n> +\n> + 2. (\"early transition\") Allow both SHA-1 and NewHash object names in\n> +    input. Any object names written to output use SHA-1. This allows\n> +    users to continue to make use of SHA-1 to communicate with peers\n> +    (e.g. by email) that have not migrated yet and prepares for mode 3.\n> +\n> + 3. (\"late transition\") Allow both SHA-1 and NewHash object names in\n> +    input. Any object names written to output use NewHash. In this\n> +    mode, users are using a more secure object naming method by\n> +    default.  The disruption is minimal as long as most of their peers\n> +    are in mode 2 or mode 3.\n> +\n> + 4. (\"post-transition\") Treat object names input by the user as\n> +    NewHash and write output using NewHash. This is safer than mode 3\n> +    because there is less risk that input is incorrectly interpreted\n> +    using the wrong hash function.\n\nSurely we can error-out if the provided object name is ambiguous?\n\n> +Selection of a New Hash\n> +-----------------------\n> +In early 2005, around the time that Git was written,  Xiaoyun Wang,\n> +Yiqun Lisa Yin, and Hongbo Yu announced an attack finding SHA-1\n> +collisions in 2^69 operations. In August they published details.\n> +Luckily, no practical demonstrations of a collision in full SHA-1 were\n> +published until 10 years later, in 2017.\n> +\n> +The hash function NewHash to replace SHA-1 should be stronger than\n> +SHA-1 was: we would like it to be trustworthy and useful in practice\n> +for at least 10 years.\n> +\n> +Some other relevant properties:\n> +\n> +1. A 256-bit hash (long enough to match common security practice; not\n> +   excessively long to hurt performance and disk usage).\n> +\n> +2. High quality implementations should be widely available (e.g. in\n> +   OpenSSL).\n> +\n> +3. The hash function's properties should match Git's needs (e.g. Git\n> +   requires collision and 2nd preimage resistance and does not require\n> +   length extension resistance).\n\nBased on recent discussion, I would add here, that the candidate hash\nhas had sufficient review.  Such that the likelihood of overnight\ncatastrophic failure is greatly reduced.  This gives git and git users\ntime to migrate away from the now weakening hash function.\n\n> +\n> +4. As a tiebreaker, the hash should be fast to compute (fortunately\n> +   many contenders are faster than SHA-1).\n> +\n> +Some hashes under consideration are SHA-256, SHA-512/256, SHA-256x16,\n> +K12, and BLAKE2bp-256.\n\nIf anyone is counting votes, I prefer either SHA-512/256 or\nBLAKE2bp-256.  But as I've mentioned elsewhere, it's only a preference.\n\n> +\n> +Transition plan\n> +---------------\n...\n> +Once a critical mass of users have upgraded to a version of Git that\n> +can verify NewHash signatures and have converted their existing\n> +repositories to support verifying them, we can add support for a\n> +setting to generate only NewHash signatures. This is expected to be at\n> +least a year later.\n> +\n> +That is also a good moment to advertise the ability to convert\n> +repositories to use NewHash only, stripping out all SHA-1 related\n> +metadata. This improves performance by eliminating translation\n> +overhead and security by avoiding the possibility of accidentally\n> +relying on the safety of SHA-1.\n\nThere is a caveat here regarding old signatures.  Those have value and\nshouldn't be lost.  repos needing to prove the validity of the old\nsha1-only signatures should counter-hash all objects, and then\ncounter-sign the corresponding newhash version of the original sha1-only\ntags.\n\nReviewed-by: Jason Cooper <jason@lakedaemon.net>\n\nthx,\n\nJason.\n"},{"id":"329477","messageId":"20171002193730.2ig5ceaasha47f2y@sigill.intra.peff.net","threadId":"45288","inReplyTo":"CA+55aFyb=h1V-3tkESY8jkc356k5rcQRmjr6o_8p6ZgKMp=Jag@mail.gmail.com","subject":"Re: RFC v3: Another proposed hash function transition plan","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2017-10-02T19:37:31Z","receivedAt":"2017-10-02T19:37:38Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Oct 02, 2017 at 10:18:02AM -0700, Linus Torvalds wrote:\n\n> On Mon, Oct 2, 2017 at 7:00 AM, Jason Cooper <jason@lakedaemon.net> wrote:\n> >\n> > Ahhh, so if I understand you correctly, you'd prefer SHA-256 over\n> > SHA3-256 because it's more performant for your usecase?  Well, that's a\n> > completely different animal that cryptographic suitability.\n> \n> In almost all loads I've seen, zlib inflate() cost is a bigger deal\n> than the crypto load. The crypto people talk about cycles per byte,\n> but the deflate code is what usually takes the page faults and cache\n> misses etc, and has bad branch prediction. That ends up easily being\n> tens or thousands of cycles, even for small data.\n\nIf anyone is interested in the user-visible effects of slower crypto, I\nthink, there are some numbers in 8325e43b82 (Makefile: add DC_SHA1 knob,\n2017-03-16). I don't know how SHA-256 compares to sha1dc exactly, but\ncertainly the latter is a lot slower than normal sha1.\n\nThe only real-world case I found with a noticeable slowdown was\nindex-pack.  Which in the worst case is roughly the same operation as\n\"git fsck\" (inflate and compute the sha1 on every byte), but people tend\nto actually do it a lot more often.\n\nAnd it really _is_ slower for real-world operations; the CPU for\ncomputing the sha1 of an incoming clone of linux.git jumped from ~3\nminutes to ~6 minutes.  But I don't think we've seen a lot of\ncomplaints, probably because that time is lumped in with \"time to\ntransfer a gigabyte of data\", so unless you're on a slow machine on fast\nconnection, you don't even really notice.\n\nFor day-to-day operations in a repository, I never came up with a good\nexample where the speed difference mattered. I think Dscho's giant-index\nexample is an outlier and the right answer there is not \"pick a fast\ncrypto algorithm\" but \"stop using a slow crypto algorithm as a\nchecksum\" (and also, stop routinely reading and writing 400MB for\nday-to-day operations).\n\n-Peff\n"},{"id":"329478","messageId":"20171002194127.GI31762@io.lakedaemon.net","threadId":"45288","inReplyTo":"20170929173413.GI19555@aiede.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Jason Cooper","fromEmail":"jason@lakedaemon.net","sentAt":"2017-10-02T19:41:27Z","receivedAt":"2017-10-02T19:41:47Z","isPatch":true,"sender":{"key":"jason@lakedaemon.net","avatar":null},"body":"On Fri, Sep 29, 2017 at 10:34:13AM -0700, Jonathan Nieder wrote:\n> Junio C Hamano wrote:\n> > Jonathan Nieder <jrnieder@gmail.com> writes:\n...\n> > If it is a goal to eventually be able to lose SHA-1 compatibility\n> > metadata from the objects, then we might want to remove SHA-1 based\n> > signature bits (e.g. PGP trailer in signed tag, gpgsig header in the\n> > commit object) from NewHash contents, and instead have them stored\n> > in a side \"metadata\" table, only to be used while converting back.\n> > I dunno if that is desirable.\n> \n> I don't consider that desirable.\n> \n> A SHA-1 based signature is still of historical interest even if my\n> centuries-newer version of Git is not able to verify it.\n\nAgreed, even a signature made by a now exposed and revoked key still has\nvalidity.  Especially in a commit or merge.  We know it was made prior\nto the key being compromised / revoked.\n\nThis is assuming that the keyholder can definitively say \"Don't trust\nsignatures from this key after this date/time+0000\".  And the signature\nin question is in the git history prior to that cut off.\n\nTags are a different animal because they can be added at any time and\naren't directly incorporated into the history.\n\nthx,\n\nJason.\n"},{"id":"329525","messageId":"xmqqbmlokcvp.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170928044320.GA84719@aiede.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-10-03T05:40:26Z","receivedAt":"2017-10-03T05:40:34Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> +Signed Tags\n> +~~~~~~~~~~~\n> +We add a new field \"gpgsig-newhash\" to the tag object format to allow\n> +signing tags without relying on SHA-1. Its signed payload is the\n> +newhash-content of the tag with its gpgsig-newhash field and \"-----BEGIN PGP\n> +SIGNATURE-----\" delimited in-body signature removed.\n> +\n> +This means tags can be signed\n> +1. using SHA-1 only, as in existing signed tag objects\n> +2. using both SHA-1 and NewHash, by using gpgsig-newhash and an in-body\n> +   signature.\n> +3. using only NewHash, by only using the gpgsig-newhash field.\n\nI have the same issue with signed commit.\n\nThe signed parts for SHA-1 contents exclude the in-body signature\n(obviously) and all the headers including gpgsig-newhash that is not\nknown to our old clients are included.  The signed parts for NewHash\ncontents exclude the in-body signature and gpgsig-newhash header,\nbut all other headers.  I somehow feel that we should just reserve\ngpgsig-* to prepare for the day when we introduce newhash2 and later\nand exclude all of them from the computation.  Treat the difference\nbetween how SHA-1 contents excludes _only_ it knows about and how\nNewHash contents excludes _all_ possible signatures, just like the\ndifferece between where SHA-1 and NewHash contents has the\nsignature.  That is, yes, we didn't know better when we designed\nSHA-1 contents, but now we know better and are correcting the\nmistakes by moving the signature from in-body tail to a header, and\nby excluding anything gpgsig-*, not just the known ones.\n\n> +Mergetag embedding\n> +~~~~~~~~~~~~~~~~~~\n> +The mergetag field in the sha1-content of a commit contains the\n> +sha1-content of a tag that was merged by that commit.\n> +\n> +The mergetag field in the newhash-content of the same commit contains the\n> +newhash-content of the same tag.\n\nOK.  \n\nWe do not have a tool that extracts them and creates a tag object,\nbut if such a tool is invented in the future, it would only have to\nworry about newhash content, as it would be a local operation.\nMakes sense.\n\n> +Submodules\n> +~~~~~~~~~~\n> +To convert recorded submodule pointers, you need to have the converted\n> +submodule repository in place. The translation table of the submodule\n> +can be used to look up the new hash.\n\nOK, I earlier commented on a paragraph that I couldn't tell what it\nwas talking about, but this is a lot more understandable.  Perhaps\nthe earlier one can be removed?\n\nWe saw earlier what happens during \"fetch\".  This seems to hint that\nwe would need to do a \"recursive\" fetch in the bottom-up direction,\nbut without fetching the superproject, you wouldn't know what submodules\nare needed and from where, so there is a bit of chicken-and-egg problem\nwe need to address, as we further make the design more detailed.\n\n> +Loose objects and unreachable objects\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> ...\n> +\"git gc --auto\" currently waits for there to be 50 packs present\n> +before combining packfiles. Packing loose objects more aggressively\n> +may cause the number of pack files to grow too quickly. This can be\n> +mitigated by using a strategy similar to Martin Fick's exponential\n> +rolling garbage collection script:\n> +https://gerrit-review.googlesource.com/c/gerrit/+/35215\n\nYes, concatenating into the latest pack that still is small may be a\nreasonable way, as there won't be many good chances to create good\ndeltas anyway until you have blobs and trees at sufficiently numbers\nof different versions, to do a \"quick GC whose only purpose is to\nkeep the number of loose object down\".\n\n> +To avoid a proliferation of UNREACHABLE_GARBAGE packs, they can be\n> +combined under certain circumstances. If \"gc.garbageTtl\" is set to\n> +greater than one day, then packs created within a single calendar day,\n> +UTC, can be coalesced together. The resulting packfile would have an\n> +mtime before midnight on that day, so this makes the effective maximum\n> +ttl the garbageTtl + 1 day. If \"gc.garbageTtl\" is less than one day,\n> +then we divide the calendar day into intervals one-third of that ttl\n> +in duration. Packs created within the same interval can be coalesced\n> +together. The resulting packfile would have an mtime before the end of\n> +the interval, so this makes the effective maximum ttl equal to the\n> +garbageTtl * 4/3.\n\nOK.  \n\nIs the use of mtime essential, or because packs are \"write once and\nfrom there access read-only\", would a timestamp written somewhere in\nthe header or the trailer of the file, if existed, work equally\nwell?  Not a strong objection, but a mild suggestion that not\nrelying on mtime may be a good idea (it will keep an accidental /\nunintended \"touch\" from keeping garbage alive longer than you want).\n\n> +The UNREACHABLE_GARBAGE setting goes in the PSRC field of the pack\n> +index. More generally, that field indicates where a pack came from:\n> +\n> + - 1 (PACK_SOURCE_RECEIVE) for a pack received over the network\n> + - 2 (PACK_SOURCE_AUTO) for a pack created by a lightweight\n> +   \"gc --auto\" operation\n> + - 3 (PACK_SOURCE_GC) for a pack created by a full gc\n> + - 4 (PACK_SOURCE_UNREACHABLE_GARBAGE) for potential garbage\n> +   discovered by gc\n> + - 5 (PACK_SOURCE_INSERT) for locally created objects that were\n> +   written directly to a pack file, e.g. from \"git add .\"\n> +\n> +This information can be useful for debugging and for \"gc --auto\" to\n> +make appropriate choices about which packs to coalesce.\n\nWould this be the direction we want to take to reduce the number of\nauxiliary files like *.keep, *.promised, etc., or we do not envision\nthese to be useful for anything other than \"gc\"?\n\n> +Caveats\n> +-------\n> +Invalid objects\n> +...\n> +More profoundly broken objects (e.g., a commit with a truncated \"tree\"\n> +header line) cannot be converted but were not usable by current Git\n> +anyway.\n\nFair enough.\n\n> +Shallow clone and submodules\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> +Because it requires all referenced objects to be available in the\n> +locally generated translation table, this design does not support\n> +shallow clone or unfetched submodules. Protocol improvements might\n> +allow lifting this restriction.\n\nOK, I think it is sensible to leave them outside the scope at the\nmoment.  All we need is a reliable way to learn the NewHash name of\nthe objects immediately beyond the cut-off points, but it will have\nto become a huge discussion how to ensure that reliability, without\ntrusting the remote too much.\n\n> +Alternates\n> +~~~~~~~~~~\n> +For the same reason, a newhash repository cannot borrow objects from a\n> +sha1 repository using objects/info/alternates or\n> +$GIT_ALTERNATE_OBJECT_REPOSITORIES.\n\nCorrect.  In addition, if the alternate has already fully migrated\naway from SHA-1 compatiblity, we can only use it for local operation.\n\n    ... goes back and thinks\n\nNo, we cannot use such an alternate even for local operation.  So a\nnewhash repository cannot borrow objects from a SHA-1 repository,\nand from a newhash repository that lost SHA-1 compatiblity if it\nitself wants to retain SHA-1 compatiblity.\n\nWhich again is \"fair enough\", I'd say.\n\n> +git notes\n> +~~~~~~~~~\n> +The \"git notes\" tool annotates objects using their sha1-name as key.\n> +This design does not describe a way to migrate notes trees to use\n> +newhash-names. That migration is expected to happen separately (for\n> +example using a file at the root of the notes tree to describe which\n> +hash it uses).\n\nTo be consistent with the remainder of the design, I think they\nshould also be translated to NewHash, but punting it is OK to limit\nthe scope of the initial migration.\n\n> +Server-side cost\n> +~~~~~~~~~~~~~~~~\n> +Until Git protocol gains NewHash support, using NewHash based storage\n> +on public-facing Git servers is strongly discouraged. Once Git\n> +protocol gains NewHash support, NewHash based servers are likely not\n> +to support SHA-1 compatibility, to avoid what may be a very expensive\n> +hash reencode during clone and to encourage peers to modernize.\n\nI doubt that the first sentence is needed.  We as git-core community\nwill not help people to run Git service backed by NewHash storage\nthat talks SHA-1 over the wire, by limiting the scope to \"NewHash\nGit fetching from SHA-1 Git\" and \"NewHash Git pushing to SHA-1 Git\"\nand not including the other two combinations.  That may be worth\nsaying here.  Masochist server operators are still welcome to build\nand operate such a service and we don't really care.  It's not our\nbusiness.\n\n> +The design described here allows fetches by SHA-1 clients of a\n> +personal NewHash repository because it's not much more difficult than\n> +allowing pushes from that repository.\n\nDoes the design described here really allow that?\n\nI thought what I read was \"everybody talks SHA-1 over the wire, and\nthose who want to use NewHash converts\".  So a user may be able to\npush from a personal NewHash repository to a personal SHA-1\nrepository (to simulate a fetch going in the reverse direction).\n\nIn any case, I do not think I saw conversion issues discussed for a\nfetch from NewHash repository earlier in the document, where\nconversion considerations for other two modes (fetch to NewHash, and\npush from NewHash) were reasonably well described.  If we are to\nallow this third mode, we'd need to make sure \"because it's not much\nmore difficult\" is true.\n\n> This support needs to be guarded\n> +by a configuration option --- servers like git.kernel.org that serve a\n> +large number of clients would not be expected to bear that cost.\n\nYes, of course.  And if these 6 lines are not unintended leftover\nfrom earlier round of the design that we wanted to remove but forget\nto do so, then the first paragraph I doubted its validity of starts\nto make sense.\n\n> +Meaning of signatures\n> +~~~~~~~~~~~~~~~~~~~~~\n> +The signed payload for signed commits and tags does not explicitly\n> +name the hash used to identify objects. If some day Git adopts a new\n> +hash function with the same length as the current SHA-1 (40\n> +hexadecimal digit) or NewHash (64 hexadecimal digit) objects then the\n> +intent behind the PGP signed payload in an object signature is\n> +unclear:\n> +\n> +\tobject e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7\n> +\ttype commit\n> +\ttag v2.12.0\n> +\ttagger Junio C Hamano <gitster@pobox.com> 1487962205 -0800\n> +\n> +\tGit 2.12\n> +\n> +Does this mean Git v2.12.0 is the commit with sha1-name\n> +e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7 or the commit with\n> +new-40-digit-hash-name e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7?\n> +\n> +Fortunately NewHash and SHA-1 have different lengths. If Git starts\n> +using another hash with the same length to name objects, then it will\n> +need to change the format of signed payloads using that hash to\n> +address this issue.\n\nThis is not just signatures, is it?  The reference to parent commits\nand its tree in a commit object would also have ambiguity between\nSHA-1 and new-40-digit-hash.  And the \"no mixed repository\" rule\nresolved that for us---isn't that sufficient for the signed tag (or\ncommit), too?  If such a signed-tag appears in a SHA-1 content of a\ntag, then the \"object\" reference is made with SHA-1.  If the tag is\nin NewHash40 content, \"object\" reference is made with NewHash40, no?\n\n> +Object names on the command line\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> +To support the transition (see Transition plan below), this design\n> +supports four different modes of operation:\n> +\n> + 1. (\"dark launch\") Treat object names input by the user as SHA-1 and\n> +    convert any object names written to output to SHA-1, but store\n> +    objects using NewHash.  This allows users to test the code with no\n> +    visible behavior change except for performance.  This allows\n> +    allows running even tests that assume the SHA-1 hash function, to\n> +    sanity-check the behavior of the new mode.\n\nOooooh.  That's ambitious.\n\n> + 2. (\"early transition\") Allow both SHA-1 and NewHash object names in\n> +    input. Any object names written to output use SHA-1. This allows\n> +    users to continue to make use of SHA-1 to communicate with peers\n> +    (e.g. by email) that have not migrated yet and prepares for mode 3.\n\nThis and others also make sense.\n\n> +Transition plan\n> +---------------\n> +Some initial steps can be implemented independently of one another:\n> +...\n> +- introducing index v3\n\nJust making sure; this is pack .idx v3?\n\n> +The infrastructure supporting fetch also allows converting an existing\n> +repository. In converted repositories and new clones, end users can\n> +gain support for the new hash function without any visible change in\n> +behavior (see \"dark launch\" in the \"Object names on the command line\"\n> +section). In particular this allows users to verify NewHash signatures\n> +on objects in the repository, and it should ensure the transition code\n> +is stable in production in preparation for using it more widely.\n> +\n> +Over time projects would encourage their users to adopt the \"early\n> +transition\" and then \"late transition\" modes to take advantage of the\n> +new, more futureproof NewHash object names.\n> +\n> +When objectFormat and compatObjectFormat are both set, commands\n> +generating signatures would generate both SHA-1 and NewHash signatures\n> +by default to support both new and old users.\n> +\n> +In projects using NewHash heavily, users could be encouraged to adopt\n> +the \"post-transition\" mode to avoid accidentally making implicit use\n> +of SHA-1 object names.\n> +\n> +Once a critical mass of users have upgraded to a version of Git that\n> +can verify NewHash signatures and have converted their existing\n> +repositories to support verifying them, we can add support for a\n> +setting to generate only NewHash signatures. This is expected to be at\n> +least a year later.\n> +\n> +That is also a good moment to advertise the ability to convert\n> +repositories to use NewHash only, stripping out all SHA-1 related\n> +metadata. This improves performance by eliminating translation\n> +overhead and security by avoiding the possibility of accidentally\n> +relying on the safety of SHA-1.\n> +\n> +Updating Git's protocols to allow a server to specify which hash\n> +functions it supports is also an important part of this transition. It\n> +is not discussed in detail in this document but this transition plan\n> +assumes it happens. :)\n\nAll of the above sounds sensible to me.\n\n> +Alternatives considered\n> +-----------------------\n\nThis message stops here...\n"},{"id":"329558","messageId":"20171003130814.GJ31762@io.lakedaemon.net","threadId":"45288","inReplyTo":"xmqqbmlokcvp.fsf@gitster.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Jason Cooper","fromEmail":"jason@lakedaemon.net","sentAt":"2017-10-03T13:08:14Z","receivedAt":"2017-10-03T13:24:29Z","isPatch":true,"sender":{"key":"jason@lakedaemon.net","avatar":null},"body":"On Tue, Oct 03, 2017 at 02:40:26PM +0900, Junio C Hamano wrote:\n> Jonathan Nieder <jrnieder@gmail.com> writes:\n...\n> > +Meaning of signatures\n> > +~~~~~~~~~~~~~~~~~~~~~\n> > +The signed payload for signed commits and tags does not explicitly\n> > +name the hash used to identify objects. If some day Git adopts a new\n> > +hash function with the same length as the current SHA-1 (40\n> > +hexadecimal digit) or NewHash (64 hexadecimal digit) objects then the\n> > +intent behind the PGP signed payload in an object signature is\n> > +unclear:\n> > +\n> > +\tobject e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7\n> > +\ttype commit\n> > +\ttag v2.12.0\n> > +\ttagger Junio C Hamano <gitster@pobox.com> 1487962205 -0800\n> > +\n> > +\tGit 2.12\n> > +\n> > +Does this mean Git v2.12.0 is the commit with sha1-name\n> > +e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7 or the commit with\n> > +new-40-digit-hash-name e7e07d5a4fcc2a203d9873968ad3e6bd4d7419d7?\n> > +\n> > +Fortunately NewHash and SHA-1 have different lengths. If Git starts\n> > +using another hash with the same length to name objects, then it will\n> > +need to change the format of signed payloads using that hash to\n> > +address this issue.\n> \n> This is not just signatures, is it?  The reference to parent commits\n> and its tree in a commit object would also have ambiguity between\n> SHA-1 and new-40-digit-hash.  And the \"no mixed repository\" rule\n> resolved that for us---isn't that sufficient for the signed tag (or\n> commit), too?  If such a signed-tag appears in a SHA-1 content of a\n> tag, then the \"object\" reference is made with SHA-1.  If the tag is\n> in NewHash40 content, \"object\" reference is made with NewHash40, no?\n\nI do hope we adhere to \"no mixed repository\" rule.  Or, at least, \"no\nmixing of hash types\".  Ambiguity opens cracks for uncertainty to creep\nin.\n\nFor our case, where we counter-hash the sha1 commits, and counter-sign\nthe sha1-based signatures, we intend to include the relevant\nsha1<->newhash lookups in the newhash signature body.  afaict, the git\nsha1<->newhash table is not cryptographically secured underneath\nsignatures, and thus can't be used in the verification of objects.\n\nThe advantage to this approach is that we can be as explicit as\nnecessary with \"SHA-1 -> SHA-512/256\" or \"SHA-1 -> SHA3-256\" in the body\nof the message.\n\nthx,\n\nJason.\n"},{"id":"329630","messageId":"xmqqefqjheky.fsf@gitster.mtv.corp.google.com","threadId":"45288","inReplyTo":"20170928044320.GA84719@aiede.mtv.corp.google.com","subject":"Re: [PATCH v4] technical doc: add a design doc for hash function transition","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2017-10-04T01:44:13Z","receivedAt":"2017-10-04T01:44:21Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com> writes:\n\n> +Alternatives considered\n> +-----------------------\n> +Upgrading everyone working on a particular project on a flag day\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> ...\n> +Using hash functions in parallel\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> ...\n\nGood that we are not doing these ;-)\n\n> +Lazily populated translation table\n> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~\n> +Some of the work of building the translation table could be deferred to\n> +push time, but that would significantly complicate and slow down pushes.\n> +Calculating the sha1-name at object creation time at the same time it is\n> +being streamed to disk and having its newhash-name calculated should be\n> +an acceptable cost.\n\nAnd the version described in the body of the document hopefully\nwould be simpler.  It certainly would be, when SHA-1 content and\nNewHash content are the same (i.e. blob).\n\nTHanks.\n"}]}