{"thread":{"id":"55645","subject":"Preserving the ability to have both SHA1 and SHA256 signatures","startedAt":"2021-05-08T02:22:28Z","lastAt":"2021-05-18T05:32:58Z","messageCount":19,"participants":["dwh@linuxprogrammer.org","Christian Couder","Junio C Hamano","Felipe Contreras","Stefan Moch","brian m. carlson","Ævar Arnfjörð Bjarmason","Konstantin Ryabitsev","Jonathan Nieder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"423907","messageId":"20210508022225.GH3986@localhost","threadId":"55645","inReplyTo":null,"subject":"Preserving the ability to have both SHA1 and SHA256 signatures","fromName":"","fromEmail":"dwh@linuxprogrammer.org","sentAt":"2021-05-08T02:22:25Z","receivedAt":"2021-05-08T02:22:28Z","isPatch":false,"sender":{"key":"dwh@linuxprogrammer.org","avatar":null},"body":"Hi Everybody,\n\nI was reading through the\nDocumentation/technical/hash-function-transition.txt doc and realized\nthat the plan is to support allowing BOTH SHA1 and SHA256 signatures to\nexist in a single object:\n\n> Signed Commits\n> 1. using SHA-1 only, as in existing signed commit objects\n> 2. using both SHA-1 and SHA-256, by using both gpgsig-sha256 and gpgsig\n>   fields.\n> 3. using only SHA-256, by only using the gpgsig-sha256 field.\n>\n> Signed Tags\n> 1. using SHA-1 only, as in existing signed tag objects\n> 2. using both SHA-1 and SHA-256, by using gpgsig-sha256 and an in-body\n>   signature.\n> 3. using only SHA-256, by only using the gpgsig-sha256 field.\n\nThe design that I'm working on only supports a single signature that\nuses a combination of fields: one 'signtype', zero or more 'signoption'\nand one 'sign' in objects. I am thinking that the best thing to do is\nreplace the gpgsig-sha256 fields in objects and allow old gpgsig (commits)\nand in-body (tags) signatures to co-exist along side to give the same\nfunctionality.\n\nThat not only paves the way forward but preserves the full backward\ncompatibility that is one of my top requirements.\n\nThoughts?\n\nCheers!\nDave\n"},{"id":"423919","messageId":"CAP8UFD0vp-zZv=Q1+KWv8PHnxTuspTw2aSCUp8QUic0HOSyq4w@mail.gmail.com","threadId":"55645","inReplyTo":"20210508022225.GH3986@localhost","subject":"Re: Preserving the ability to have both SHA1 and SHA256 signatures","fromName":"Christian Couder","fromEmail":"christian.couder@gmail.com","sentAt":"2021-05-08T06:39:28Z","receivedAt":"2021-05-08T06:39:45Z","isPatch":false,"sender":{"key":"christian.couder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/208954?v=4"},"body":"Hi,\n\n(Not sure why, but, when using \"Reply to all\" in Gmail, it doesn't\nactually reply to you (or Cc you), only to the mailing list. I had to\nmanually add your email back.)\n\nOn Sat, May 8, 2021 at 4:25 AM <dwh@linuxprogrammer.org> wrote:\n>\n> Hi Everybody,\n>\n> I was reading through the\n> Documentation/technical/hash-function-transition.txt doc and realized\n> that the plan is to support allowing BOTH SHA1 and SHA256 signatures to\n> exist in a single object:\n>\n> > Signed Commits\n> > 1. using SHA-1 only, as in existing signed commit objects\n> > 2. using both SHA-1 and SHA-256, by using both gpgsig-sha256 and gpgsig\n> >   fields.\n> > 3. using only SHA-256, by only using the gpgsig-sha256 field.\n> >\n> > Signed Tags\n> > 1. using SHA-1 only, as in existing signed tag objects\n> > 2. using both SHA-1 and SHA-256, by using gpgsig-sha256 and an in-body\n> >   signature.\n> > 3. using only SHA-256, by only using the gpgsig-sha256 field.\n>\n> The design that I'm working on only supports a single signature that\n> uses a combination of fields: one 'signtype', zero or more 'signoption'\n> and one 'sign' in objects.\n\nHere I understand that your design doesn't support both a SHA1 and a\nSHA256 signature.\n\n> I am thinking that the best thing to do is\n> replace the gpgsig-sha256 fields in objects and allow old gpgsig (commits)\n> and in-body (tags) signatures to co-exist along side to give the same\n> functionality.\n\nIs this part of your design, or a, maybe temporary, alternative to it?\n\n> That not only paves the way forward but preserves the full backward\n> compatibility that is one of my top requirements.\n\nThere has been patches and discussions quite recently about this, that\nhave been reported on in our Git Rev News newsletter:\n\nhttps://git.github.io/rev_news/2021/02/27/edition-72/\n\nYou can see that, with the latest patches (not sure the documentation\nis up-to-date though), signing both commits and tags\n can now be round-tripped through both SHA-1 and SHA-256 conversions.\nHow isn't that fully backward compatible?\n\nBest,\nChristian.\n"},{"id":"423920","messageId":"xmqqim3tvhlr.fsf@gitster.g","threadId":"55645","inReplyTo":"CAP8UFD0vp-zZv=Q1+KWv8PHnxTuspTw2aSCUp8QUic0HOSyq4w@mail.gmail.com","subject":"Re: Preserving the ability to have both SHA1 and SHA256 signatures","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-05-08T06:56:32Z","receivedAt":"2021-05-08T06:56:39Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Christian Couder <christian.couder@gmail.com> writes:\n\n> Hi,\n>\n> (Not sure why, but, when using \"Reply to all\" in Gmail, it doesn't\n> actually reply to you (or Cc you), only to the mailing list. I had to\n> manually add your email back.)\n\nI am sure why.  DWH, please do not use mail-follow-up-to when\nworking with this list.  It is rude and wastes people's time (like\nthe practice just did by stealing time from Christian).\n\nAlso cf.\nhttps://lore.kernel.org/git/7v63l6f1mc.fsf@gitster.siamese.dyndns.org/\nhttps://lore.kernel.org/git/7vk3zig92n.fsf@alter.siamese.dyndns.org/\n\n\n"},{"id":"423922","messageId":"609645cb11f72_1fc6d208ee@natae.notmuch","threadId":"55645","inReplyTo":"xmqqim3tvhlr.fsf@gitster.g","subject":"Re: Preserving the ability to have both SHA1 and SHA256 signatures","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2021-05-08T08:03:23Z","receivedAt":"2021-05-08T08:03:32Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"Junio C Hamano wrote:\n> Christian Couder <christian.couder@gmail.com> writes:\n> > (Not sure why, but, when using \"Reply to all\" in Gmail, it doesn't\n> > actually reply to you (or Cc you), only to the mailing list. I had to\n> > manually add your email back.)\n> \n> I am sure why.  DWH, please do not use mail-follow-up-to when\n> working with this list.  It is rude and wastes people's time (like\n> the practice just did by stealing time from Christian).\n\nI agree with this, but shouldn't this be written in some kind of mail\netiquiette guideline? Along with a rationale.\n\n-- \nFelipe Contreras\n"},{"id":"423925","messageId":"f4f782c4-3adc-8c1c-428d-8037426fc475@mail.de","threadId":"55645","inReplyTo":"609645cb11f72_1fc6d208ee@natae.notmuch","subject":"Re: Preserving the ability to have both SHA1 and SHA256 signatures","fromName":"Stefan Moch","fromEmail":"stefanmoch@mail.de","sentAt":"2021-05-08T10:11:51Z","receivedAt":"2021-05-08T10:20:00Z","isPatch":false,"sender":{"key":"stefanmoch@mail.de","avatar":null},"body":"Felipe Contreras wrote:\n> Junio C Hamano wrote:\n>> Christian Couder <christian.couder@gmail.com> writes:\n>>> (Not sure why, but, when using \"Reply to all\" in Gmail, it doesn't\n>>> actually reply to you (or Cc you), only to the mailing list. I had to\n>>> manually add your email back.)\n>>\n>> I am sure why.  DWH, please do not use mail-follow-up-to when\n>> working with this list.  It is rude and wastes people's time (like\n>> the practice just did by stealing time from Christian).\n> \n> I agree with this, but shouldn't this be written in some kind of mail\n> etiquiette guideline? Along with a rationale.\n\nGood idea to write this down. How to use the mailing list is only\nsparsely documented. The following files talk about sending to the\nmailing list:\n\n 1. README.md\n 2. Documentation/SubmittingPatches\n 3. Documentation/MyFirstContribution.txt\n 4. MaintNotes (in Junio's “todo” branch, sent out to the list from\n    time to time as “A note from the maintainer”)\n\n2, 3 and 4 mention sending Cc to everyone involved.\n\n2 is about new messages.\n\n3 and 4 specifically talk about keeping everyone in Cc: in replies.\nBoth in the context of “you don't have to be subscribed and you\ndon't need to ask for Cc:”.\n\n\nPlease also note, that mutt sets the “Mail-Followup-To:” header by\ndefault for sending to known mailing lists, unless “followup_to” is\nset to “no”. Whether or not it removes the sender address in this\nheader depends on the list address to be known to be subscribed to\nor simply known to be a mailing list. It also does not set this\nheader if no recipient address is known as a mailing list.\n\nhttp://www.mutt.org/doc/manual/#followup-to\nhttp://www.mutt.org/doc/manual/#using-lists\n"},{"id":"423929","messageId":"xmqqlf8ptr7h.fsf@gitster.g","threadId":"55645","inReplyTo":"f4f782c4-3adc-8c1c-428d-8037426fc475@mail.de","subject":"Re: Preserving the ability to have both SHA1 and SHA256 signatures","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-05-08T11:12:02Z","receivedAt":"2021-05-08T11:12:07Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Stefan Moch <stefanmoch@mail.de> writes:\n\n> Good idea to write this down. How to use the mailing list is only\n> sparsely documented. The following files talk about sending to the\n> mailing list:\n>\n>  1. README.md\n>  2. Documentation/SubmittingPatches\n>  3. Documentation/MyFirstContribution.txt\n>  4. MaintNotes (in Junio's “todo” branch, sent out to the list from\n>     time to time as “A note from the maintainer”)\n>\n> 2, 3 and 4 mention sending Cc to everyone involved.\n>\n> 2 is about new messages.\n>\n> 3 and 4 specifically talk about keeping everyone in Cc: in replies.\n> Both in the context of “you don't have to be subscribed and you\n> don't need to ask for Cc:”.\n\nIn case somebody wants to write a doc, a better pair of references\nthan what I quoted earlier to draw material from are:\n\nhttps://public-inbox.org/git/7v4pndfjym.fsf@assigned-by-dhcp.cox.net/\nhttps://public-inbox.org/git/7vei7zjr3y.fsf@alter.siamese.dyndns.org/\n\n"},{"id":"423951","messageId":"YJcqqYsOerijsxRQ@camp.crustytoothpaste.net","threadId":"55645","inReplyTo":"20210508022225.GH3986@localhost","subject":"Re: Preserving the ability to have both SHA1 and SHA256 signatures","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2021-05-09T00:19:53Z","receivedAt":"2021-05-09T00:19:59Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2021-05-08 at 02:22:25, dwh@linuxprogrammer.org wrote:\n> Hi Everybody,\n> \n> I was reading through the\n> Documentation/technical/hash-function-transition.txt doc and realized\n> that the plan is to support allowing BOTH SHA1 and SHA256 signatures to\n> exist in a single object:\n> \n> > Signed Commits\n> > 1. using SHA-1 only, as in existing signed commit objects\n> > 2. using both SHA-1 and SHA-256, by using both gpgsig-sha256 and gpgsig\n> >   fields.\n> > 3. using only SHA-256, by only using the gpgsig-sha256 field.\n> > \n> > Signed Tags\n> > 1. using SHA-1 only, as in existing signed tag objects\n> > 2. using both SHA-1 and SHA-256, by using gpgsig-sha256 and an in-body\n> >   signature.\n> > 3. using only SHA-256, by only using the gpgsig-sha256 field.\n\nYes, this is the case.  We have tests for this case.\n\n> The design that I'm working on only supports a single signature that\n> uses a combination of fields: one 'signtype', zero or more 'signoption'\n> and one 'sign' in objects. I am thinking that the best thing to do is\n> replace the gpgsig-sha256 fields in objects and allow old gpgsig (commits)\n> and in-body (tags) signatures to co-exist along side to give the same\n> functionality.\n\nYou can't do that.  SHA-256 repositories already exist and that would\nbreak compatibility.\n\n> That not only paves the way forward but preserves the full backward\n> compatibility that is one of my top requirements.\n\nI've reviewed your proposed design and provided feedback that we need to\npreserve this functionality in your new design as well.  People will\nwant to have that functionality.\n-- \nbrian m. carlson (he/him or they/them)\nHouston, Texas, US\n"},{"id":"424026","messageId":"87lf8mu642.fsf@evledraar.gmail.com","threadId":"55645","inReplyTo":"YJcqqYsOerijsxRQ@camp.crustytoothpaste.net","subject":"Is the sha256 object format experimental or not?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-05-10T12:22:00Z","receivedAt":"2021-05-10T12:52:22Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sun, May 09 2021, brian m. carlson wrote:\n\n> [[PGP Signed Part:Undecided]]\n> On 2021-05-08 at 02:22:25, dwh@linuxprogrammer.org wrote:\n>> Hi Everybody,\n>> \n>> I was reading through the\n>> Documentation/technical/hash-function-transition.txt doc and realized\n>> that the plan is to support allowing BOTH SHA1 and SHA256 signatures to\n>> exist in a single object:\n>> \n>> > Signed Commits\n>> > 1. using SHA-1 only, as in existing signed commit objects\n>> > 2. using both SHA-1 and SHA-256, by using both gpgsig-sha256 and gpgsig\n>> >   fields.\n>> > 3. using only SHA-256, by only using the gpgsig-sha256 field.\n>> > \n>> > Signed Tags\n>> > 1. using SHA-1 only, as in existing signed tag objects\n>> > 2. using both SHA-1 and SHA-256, by using gpgsig-sha256 and an in-body\n>> >   signature.\n>> > 3. using only SHA-256, by only using the gpgsig-sha256 field.\n>\n> Yes, this is the case.  We have tests for this case.\n>\n>> The design that I'm working on only supports a single signature that\n>> uses a combination of fields: one 'signtype', zero or more 'signoption'\n>> and one 'sign' in objects. I am thinking that the best thing to do is\n>> replace the gpgsig-sha256 fields in objects and allow old gpgsig (commits)\n>> and in-body (tags) signatures to co-exist along side to give the same\n>> functionality.\n>\n> You can't do that.  SHA-256 repositories already exist and that would\n> break compatibility.\n\nFrom memory this is at least the second time you've brought up this\npoint on-list.\n\nMy feeling is that almost nobody's using sha256 currently, and we have a\nvery prominent ALL CAPS warning saying the format is experimental and\nmay change, see ff233d8dda1 (Documentation: mark\n`--object-format=sha256` as experimental, 2020-08-16).\n\nI agree with the docs as they stand, and don't think we should hold back\non changing the object format for sha256 in general if there's a\ncompelling reason to do so.\n\nWhether this suggested change has a compelling reason is another matter\n(I haven't reviewed it).\n\nBut it seems to me that if the main person pushing the sha256 effort\ndisagrees with the content of\nDocumentation/object-format-disclaimer.txt, we'd be better off at this\npoint discussing a patch to change the wording there to something to the\neffect that we consider the format set in stone at this point.\n"},{"id":"424083","messageId":"YJm23HESQb1Z6h8y@camp.crustytoothpaste.net","threadId":"55645","inReplyTo":"87lf8mu642.fsf@evledraar.gmail.com","subject":"Re: Is the sha256 object format experimental or not?","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2021-05-10T22:42:36Z","receivedAt":"2021-05-10T22:43:13Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2021-05-10 at 12:22:00, Ævar Arnfjörð Bjarmason wrote:\n> \n> On Sun, May 09 2021, brian m. carlson wrote:\n> > You can't do that.  SHA-256 repositories already exist and that would\n> > break compatibility.\n> \n> From memory this is at least the second time you've brought up this\n> point on-list.\n> \n> My feeling is that almost nobody's using sha256 currently, and we have a\n> very prominent ALL CAPS warning saying the format is experimental and\n> may change, see ff233d8dda1 (Documentation: mark\n> `--object-format=sha256` as experimental, 2020-08-16).\n\nYes, I agreed to such text because others thought it was a good idea in\ncase we needed to make a change.  However, we don't need to make an\nincompatible change here, so we should avoid that if possible.\n\nAlmost nobody is using it because the main forges don't yet support it,\nbecause it's going to be just as much work to support it there as it has\nbeen in Git.  We won't be making it easier by making deliberately\nincompatible changes when we don't have to.\n\n> I agree with the docs as they stand, and don't think we should hold back\n> on changing the object format for sha256 in general if there's a\n> compelling reason to do so.\n\nI am using it and I know of other people who are using it.  There are\npeople whose companies cannot use SHA-1 for compliance reasons and are\nalready making use of it.\n\nThe problem here is a chicken and egg: nobody's going to use SHA-256\nsupport if it's experimental and their entire repo might end up totally\nuseless, and it's not going to become stable if nobody uses it.\n\n> But it seems to me that if the main person pushing the sha256 effort\n> disagrees with the content of\n> Documentation/object-format-disclaimer.txt, we'd be better off at this\n> point discussing a patch to change the wording there to something to the\n> effect that we consider the format set in stone at this point.\n\nI've been pretty clear up front that I thought the data was stable and\nwe should avoid making incompatible changes.  It may be that it is still\nexperimental and may change incompatibly, but if we can avoid that\nproblem, we should.\n\nI don't personally intend to send a patch removing the note about it\nbeing experimental until I've finished getting object interop done,\nsince that's the major issue where we might need to make an incompatible\nchange, but that work is moving slowly.\n-- \nbrian m. carlson (he/him or they/them)\nHouston, Texas, US\n"},{"id":"424507","messageId":"20210513202919.GE11882@localhost","threadId":"55645","inReplyTo":"YJm23HESQb1Z6h8y@camp.crustytoothpaste.net","subject":"Re: Is the sha256 object format experimental or not?","fromName":"","fromEmail":"dwh@linuxprogrammer.org","sentAt":"2021-05-13T20:29:19Z","receivedAt":"2021-05-13T20:29:39Z","isPatch":false,"sender":{"key":"dwh@linuxprogrammer.org","avatar":null},"body":"On 10.05.2021 22:42, brian m. carlson wrote:\n>Almost nobody is using it because the main forges don't yet support it,\n>because it's going to be just as much work to support it there as it has\n>been in Git.  We won't be making it easier by making deliberately\n>incompatible changes when we don't have to.\n\nI know that you said there is no reason to make a breaking change to the\nSHA256 implementation now, but because of what you say above, I think we\nstill have the opportunity to make breaking changes. In any case I think\nwe only need to make one breaking change to gain algorithmic agility\ngoing forward and avoid painful, multi-year transitions like the one\nyou've been executing.\n\nMy project to add universal cryptographic signing to Git by using a\nstandard protocol and generalized configuration to support any\ncryptographic signing scheme could also apply to the digests as well,\nand I think it should. If object digests in Git were self-describing\n(i.e. they contain an algorithm identifier as well as the digest) then\nrepos gain \"algorithmic agility\" and can change algorithms at any time\nto keep up as algorithms grow stale and are replaced.\n\nI think Git should externalize the calculation of object digests just\nlike it externalizes the calcualtion of object digital signatures.\nCryptography is very difficult to get correct and the dedicated tools\nfor that (e.g. OpenSSH, OpenSSL, GnuPG, etc) get lots of scrutiny and\nhave the best chance of getting it right. I don't think Git should try\nto do cryptography at all.\n\nObject digests should just be names for objects; Git doesn't really need to\nknow anything more than \"is this the name for that object?\". Answering\nthat question can, and should, be done by an external tool that is\nimplemented correctly and hardened against attack. I think the only\ncounter-argument for this approach is performace related. Pipe-forking a\nchild process and reading/writing over IPC pipes is expensive in terms\nof context switching and process setup/teardown but there are a number\nof mitigations I won't go into here.\n\nI think we should make one last breaking change for digests and not go\nwith the existing SHA-256 implementation but instead switch to\nself-describing digests and digital signatures and rely on external\ntools that Git talks to using a standard protocol. We can maintain full\nbackward compatibility and even support full round tripping using some\nof the similar techniques that Brian came up with. A transitional\nhalf-old/half-new signed tag could look like:\n\n```\nobject 04b871796dc0420f8e7561a895b52484b701d51a\nobj 0ED_zgYrQg584bCrqKPoUvxaQ5aMis0GtnW_NrZFTTxUlHLUOyp77LanoZEGV6ajhYGLGTaTfCIQhryovyeNFJuG\ntype commit\ntag signedtag\ntagger C O Mitter <committer@example.com> 1465981006 +0000\nsigntype openpgp\nsign LS0tLS1CRUdJTiBQR1AgU0lHTkFUVVJFLS0tLS0KVmVyc2lvbjogR251UEcgdjEKCmlRRWN\n CQUFCQWdBR0JRSlhZUmhPQUFvSkVHRUpMb1czSW5HSmtsa0lBSWNuaEw3UndFYi8rUWVYOWVua1\n hoeG4KcnhmZHFydldkMUs4MHNsMlRPdDhCZy9OWXdyVUJ3L1JXSitzZy9oaEhwNFd0dkUxSERHS\n GxrRXozeTExTGt1aAo4dFN4UzNxS1R4WFVHb3p5UEd1RTkwc0pmRXhoWmxXNGtuSVExd3QveVdx\n TSszM0U5cE40aHpQcUx3eXJkb2RzCnE4RldFcVBQVWJTSlhvTWJSUHcwNFM1anJMdFpTc1VXYlJ\n Zam1KQ0h6bGhTZkZXVzRlRmQzN3VxdUlhTFVCUzAKcmtDM0pyeDc0MjBqa0lwZ0ZjVEkyczYwdW\n hTUUx6Z2NDd2RBMnVrU1lJUm5qZy96RGtqOCszaC9HYVJPSjcyeApsWnlJNkhXaXhLSmtXdzhsR\n TlhQU9EOVRtVFc5c0ZKd2NWQXptQXVGWDJrVXJlRFVLTVpkdUdjb1JZR3BEN0U9Cj1qcFhhCi0t\n LS0tRU5EIFBHUCBTSUdOQVRVUkUtLS0tLQo\n\nsigned tag\n\nsigned tag message body\n-----BEGIN PGP SIGNATURE-----\nVersion: GnuPG v1\n\niQEcBAABAgAGBQJXYRhOAAoJEGEJLoW3InGJklkIAIcnhL7RwEb/+QeX9enkXhxn\nrxfdqrvWd1K80sl2TOt8Bg/NYwrUBw/RWJ+sg/hhHp4WtvE1HDGHlkEz3y11Lkuh\n8tSxS3qKTxXUGozyPGuE90sJfExhZlW4knIQ1wt/yWqM+33E9pN4hzPqLwyrdods\nq8FWEqPPUbSJXoMbRPw04S5jrLtZSsUWbRYjmJCHzlhSfFWW4eFd37uquIaLUBS0\nrkC3Jrx7420jkIpgFcTI2s60uhSQLzgcCwdA2ukSYIRnjg/zDkj8+3h/GaROJ72x\nlZyI6HWixKJkWw8lE9aAOD9TmTW9sFJwcVAzmAuFX2kUreDUKMZduGcoRYGpD7E=\n=jpXa\n-----END PGP SIGNATURE-----\n```\n\nI think a good move to make right now would be to add a general function\nfor stripping out any number of named fields from objects and also\nstripping out in-body signatures found in tags. That way we can add\nsupport in today's Git for stripping out fields/data for things like\ncreating/verifying the object digest and/or digital signature.\n\nBTW, in the example above the 'obj' field is a self-describing, URL-safe\nBase64 encoded Blake2b-512 digest encoded using the format described\n[here][1]. The starting '0E' Base64 characters identify the digest as\nBlake2b-512 and also specify that the length of the digest is 64-bytes.\nIf you Base64 decode the 'obj' field value you get 66 bytes, the digest\nvalue is the last 64 bytes of the 66 bytes.\n\nBy going with self-describing digests we can have configuration files\nthat contain 'program' and 'options.*' for each external tool that can\ncreate/validate digests of each type. So in this case there would be\nsomething like:\n\n```\n[digest \"blake2b\"]\n  program = \"blake2bsum\"\n[digest \"blake2b.options\"]\n  length = 64\n```\n\nUsing self-describing cryptographic constructs for digests and\nsignatures and relying on external tools makes it trivial for Git to\nwalk the object graph and enumerate all of the digest types and\nsignature types in a given repo and determine if a user has their\nconfiguration set up correctly to work with that repo. Projects can\ndeclare which types they are using and recommend tools to use for those\ntypes.\n\nCheers!\nDave\n\nTL;DR\n\nLet me try to lay out the case for making a breaking change to sha256\nright now that will future-proof repos going forward.\n\nIt has been known for a few decades now that cryptography has a\nshelf-life. By that I mean as technology and cryptanalisys improves we\nhave had to make keys larger and invent new algorithms that resist the\nnew attacks on cryptography. This has been true digest algorithms (i.e.\nhashes), digital signatures (i.e. non-repudiation), and encryption (i.e.\nconfidentiality). The relevant case here is the fact that sha256 is\nvulnerable to extension attacks and cryptographers have lost some\nconfidence in it after many Davies-Meyer (DM) structure and ARX network\ndesigns based on MD4 were broken 20 years ago. SHA-256 uses DM plus a\nblock cipher based on an ARX network. The end result is that in high\nsecurity software, SHA-256 is being replaced with SHA-3 and Blake2\ndigests.\n\nAnother key thing to think about is that a git repo is a form of a\nprovenance log that could become the primary tool for securing the\nsoftware supply chain if we were to make some careful, well thought out\nchanges arond the digests and digital signatures. What changes exactly?\n\n1. upgrade the digests to something cryptographically secure.\n2. digitally sign all commits/merges/tags using...\n3. key material tracked with cryptographically secure provenance logs\ninside of the repo itself.\n4. switch to \"late binding\", \"self describing\" cryptographic constructs.\n\nLet me go over these and describe how these fit together.\n\n1. SHA-1 is not cryptographically secure and SHA-256 is already not\n   being used in *new* systems and is being replaced in existing, high\n   security systems. I think Git should move to more secure digest\n   algorithms because the hashes in Git repos are used as naming\n   identifiers for Git objects which gives them a higher security\n   burden.\n\n2. Digitally signing all commits/merges/tags is critical to tie\n   contributions to contributors in a non-repudiable fashion. At the\n   very least it is a more secure solution for S-o-b but it also opens\n   up the possibility for cryptographically secure accountability. Banks\n   and governments are already doing know-your-customer (KYC)\n   verifications of identity that can be used to identify contributors\n   and their contributions cryptographically. If privacy is a concern,\n   zero-knowledge proofs, based on the KYC authentic data, can be used\n   to create pseudonymous identities for contributors that can be linked\n   to their real-world identity under judicial order. Essentially a\n   developer can say, \"you don't need to know my real world identity but\n   here's proof that XYZ bank knows who I am and here is a large random\n   number you can use to de-anonymize me with the help of a court if\n   needed\"\n\n3. The key material used for identifying contributors needs to move into\n   the repos themselves for many reasons but the most important two\n   reasons are (1) the repo comes with all of the data necessary to\n   verify all of the digital signatures (i.e. solving the PKI problem\n   for a project) and (2) to track the provenance of the public keys and\n   other related data that each contributor uses. If Git repos contain\n   provenance logs that are controlled and maintained by each\n   contributor, those logs can also contain digital signatures over the\n   code of conduct and the developer certificate of origin and other\n   governing documents for a project that are legally binding (i.e.\n   follow eIDAS and other legal digital signature rules). Solving the\n   PKI problem alone makes digitally signing commits infinitely more\n   useful and will drive adoption. Solving the non-repudiable provenance\n   problem is the raison d'être of organizations like the Linux\n   Foundation. I think Git should align itself with where technology is\n   heading on that front.\n\n4. Currently Git uses \"early-binding\" for all cryptographic material.\n   The digest algorithm is hard coded (SHA-1) and the new SHA-256 is as\n   well. The digital signature algorithm is also hard coded as either\n   GPG or GPGSM. Early-binding makes it very difficult to plan for the\n   obsolescense of cryptographic algorithms. The solution is to move to\n   \"late-binding\"/\"self-describing\" cryptographic constructs. If Git\n   were to switch to self-describing digests and digital signatures,\n   then Git could be entirely agnostic to cryptography and rely entirely\n   upon external crytpographic tools for creating/verifying digests and\n   digital signatures. Instead of the direction we're taking on the\n   SHA256 changeover, I think Git should switch to self-describing\n   digests and digital signatures and use a standard protocol for\n   talking to external cryptographic tools instead of trying to get\n   cryptography correct in its code.\n\n   Secure Scuttlebutt uses late-binding constructs that contain a type\n   \"sigil\", Base-64 encoded key/digest/blob followed by an algorithm\n   decriptor (e.g. \".sha256\" or \".ed25519\"). Other examples exist such\n   as the Multihash encoding scheme for self-describing hashes. All of\n   my work on secure provenance logs uses the emerging consensus\n   encoding described [here][1]. It uses Base64 encoded cryptographic\n   data and it fills what would be the padding bytes with type\n   identifiers. I'm not the only one thinking along these lines. The\n   [KERI project][2] at the Decentralized Identity Foundation as well as\n   [Konstantin][3].\n\n\n[1]: https://github.com/decentralized-identity/keri/blob/master/kids/kid0001.md\n[2]: https://identity.foundation/working-groups/keri.html\n[3]: https://people.kernel.org/monsieuricon/patches-carved-into-developer-sigchains\n"},{"id":"424509","messageId":"20210513204957.5g76czb5bk3thlep@meerkat.local","threadId":"55645","inReplyTo":"20210513202919.GE11882@localhost","subject":"Re: Is the sha256 object format experimental or not?","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2021-05-13T20:49:57Z","receivedAt":"2021-05-13T20:50:02Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Thu, May 13, 2021 at 01:29:19PM -0700, dwh@linuxprogrammer.org wrote:\n> 3. The key material used for identifying contributors needs to move into\n>   the repos themselves for many reasons but the most important two\n>   reasons are (1) the repo comes with all of the data necessary to\n>   verify all of the digital signatures (i.e. solving the PKI problem\n>   for a project) and (2) to track the provenance of the public keys and\n>   other related data that each contributor uses. If Git repos contain\n>   provenance logs that are controlled and maintained by each\n>   contributor, those logs can also contain digital signatures over the\n>   code of conduct and the developer certificate of origin and other\n>   governing documents for a project that are legally binding (i.e.\n>   follow eIDAS and other legal digital signature rules). Solving the\n>   PKI problem alone makes digitally signing commits infinitely more\n>   useful and will drive adoption. Solving the non-repudiable provenance\n>   problem is the raison d'être of organizations like the Linux\n>   Foundation. I think Git should align itself with where technology is\n>   heading on that front.\n\nDave:\n\nCheck out what we're doing as part of patatt and b4:\nhttps://pypi.org/project/patatt/\n\nIt takes your keyring-in-git idea and runs with it -- it would be good to have\nyour input while the project is still young and widely unknown. :)\n\n-K\n"},{"id":"424510","messageId":"xmqqo8de9wis.fsf@gitster.g","threadId":"55645","inReplyTo":"20210513202919.GE11882@localhost","subject":"Re: Is the sha256 object format experimental or not?","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2021-05-13T21:03:23Z","receivedAt":"2021-05-13T21:03:33Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"dwh@linuxprogrammer.org writes:\n\n> I think Git should externalize the calculation of object digests just\n> like it externalizes the calcualtion of object digital signatures.\n\nThe hashing algorithms used to generate object names has\nrequirements fundamentally different from that of digital\nsignatures.  I strongly suspect that that fact would change the\nequation when you rethink what you said above.\n\nWe can \"upgrade\" digital signature algorithms fairly easily---nobody\nwould complain if you suddenly choose different signing algorithm\nover a blob of data, as long as all project participants are aware\n(and self-describing datastream helps here) and are capable of\ngrokking the new algorithm we are adopting.  But because object\nnames are used by one object to refer to another, and most\nimportantly, we do not want a single object to have multiple names,\nwe cannot afford to introduce a new hashing algorithm every time we\nfeel like it.  In other words, diversity of object naming algorithms\nis to be avoided as much as possible, while diversity of signature\nalgorithms is naturally expected.\n"},{"id":"424527","messageId":"20210513232614.GF11882@localhost","threadId":"55645","inReplyTo":"xmqqo8de9wis.fsf@gitster.g","subject":"Re: Is the sha256 object format experimental or not?","fromName":"","fromEmail":"dwh@linuxprogrammer.org","sentAt":"2021-05-13T23:26:14Z","receivedAt":"2021-05-13T23:26:21Z","isPatch":false,"sender":{"key":"dwh@linuxprogrammer.org","avatar":null},"body":"On 14.05.2021 06:03, Junio C Hamano wrote:\n>dwh@linuxprogrammer.org writes:\n>\n>> I think Git should externalize the calculation of object digests just\n>> like it externalizes the calcualtion of object digital signatures.\n>\n>The hashing algorithms used to generate object names has\n>requirements fundamentally different from that of digital\n>signatures.  I strongly suspect that that fact would change the\n>equation when you rethink what you said above.\n\nI agree with you. Object names are exactly that: names. Names for\nresources/data must be persistent, as well as global in scope and\nuniqueness, and autonomously assigned. What this means is that once an\nobject has a name, that name shall never change as long as the object\nremains unchanged. The names must be unique in the scope of all objects\n(e.g. all copies of a repo) and generated without coordination.\n\nCalculating object names using a digest algorithm meets all of these\nrequirements. Choosing a strong digest algorithm creates a strong\ncryptographic binding between the name and the object contents. Using\nself-describing digests allows for a repo to switch digest algorithms at\narbitrary points in the history.\n\nI think that objects named with SHA1 digests should remain named with\nthe SHA1 digest. I do *not* advocate going back and rewriting history\nto change all of the object names to a digest with a different\nalgorithm. Git is a provenance log and history matters. I recommend\npreserving all existing names, even if they were created with known-weak\ndigest algorithms, and making the change to a new algorithm at a\nspecific point in time (e.g. at a tag). Using self-describing digest\nencoding and externalizing digest calculation future-proofs\nrepositories and allows for preservation of history while allowing\nalgorithm agility.\n\nTo illustrate my point, I envision that a repos could have a history\nlike this:\n\nobject 2923f6fa36614586ea09b4424b438915cc1b9b67 (naked SHA1)\n  |\n<many objects named with SHA1>\n  |\nobject 5f167fb6b3e96273b564fff0b041fb94fee4d3de (naked SHA1)\n  |\n<modify Git to ext. digest calculation and self-desc encoding>\n  |\nobject 98c2e1c0965e60b0f137577ac5dd0a5c96ce224d (naked SHA1)\n  |\n<many objects named with SHA1>\n  |\n<a project decides to switch to SHA2-256, maybe marked in a tag>\n  |\nobject IAOdLVxteOxQwKa-xn8yCBUkuPkjAqcuQ2V7fKAlao8o (self-desc.SHA2-256)\n  |\n<many objects named with self-describing SHA2-256 digests>\n  |\n<a project decices to switch to SHA3-256, maybe marked in a tag>\n  |\nobject EK832G0PFhBFf-Dfgr205UKpUMqmVXJX9ltLwQo4Awct (self-desc.SHA3-256)\n  |\n<many objects named with self-descring SHA3-256 digests>\n  .\n  .\n  .\n\nNeither decision to switch to SHA2-256 nor to SHA3-256 would require any\ncode changes. If we continue down the current SHA-256 road, we will have\nto repeat that multi-year effort in the future to switch to SHA3 or\nsomething else. Most importantly, the choice of digest algorithm would\nbe left up to the maintainers of a given repo and not limited to the\nalgorithms we have hard coded into Git.\n\nBrian's work on the SHA-256 switch is valuable. We can leverage a lot of\nit to switch to externalized digest calculation and self-describing\ndigests and never have to worry about doing that again.\n\nCheers!\nDave\n"},{"id":"424529","messageId":"20210513234706.GG11882@localhost","threadId":"55645","inReplyTo":"20210513204957.5g76czb5bk3thlep@meerkat.local","subject":"Re: Is the sha256 object format experimental or not?","fromName":"","fromEmail":"dwh@linuxprogrammer.org","sentAt":"2021-05-13T23:47:06Z","receivedAt":"2021-05-13T23:47:11Z","isPatch":false,"sender":{"key":"dwh@linuxprogrammer.org","avatar":null},"body":"On 13.05.2021 16:49, Konstantin Ryabitsev wrote:\n>Check out what we're doing as part of patatt and b4:\n>https://pypi.org/project/patatt/\n>\n>It takes your keyring-in-git idea and runs with it -- it would be good to have\n>your input while the project is still young and widely unknown. :)\n\nKonstantin:\n\nThat's really clever. I especially love how you're using the list\narchive as the provenance log of old keys developers used. That seems\nlike it would work although I have worries about the security of\nX-Developer-Key and the lack of key history immediately available to\n`git log` because it's in the list archive and not in the repo directly.\nI guess the old keys would still be in your local keyring for `gpg` to\nuse but it would mark signatures created with old revoked keys as\ninvalid even though they are valid.\n\nOld keys--even if revoked or compromised--matter in a world of digitally\nsigned data. As a matter of course, people should rotate their signing\nkeys on a regular basis. It's just good hygiene. That means that there\nwill always be old data signed with old keys and those old keys need to\nbe kept around to validate the old signatures.\n\nMy approach has been to move to cryptographically secure provenance logs\nthat contain key rotation events and commitments to future keys and also\ncryptographically linking to arbitrary metadata (e.g. KYC proofs, etc).\nThe file format is documented using the Community Standard template from\nthe LF. I'm hoping to move Git to use external tools for all digest and\ndigital signature operations. Then I can start putting provenance logs\ninto a \".well-known\" path in Git repos, maybe \".plogs\" or something.\nThen I can write/adapt a signing tool to understand provenance logs\nof public keys in the repo instead of the GPG keyring stuff we have\ntoday.\n\nProvenance logs accumulate the full key history of a developer over\ntime. It represents a second axis of time such that the HEAD of a repo\nwill have the full key history, for every contributor available to\ncryptographic tools for verifying signatures. This makes `git log\n--show-signature` operations maximally efficient because we don't have\nto check out old keyrings from history to recreate the state GPG was in\nwhen the signature was created.\n\nI still like your approach purely for the \"it works right now\" aspect of\nthe solution. Good job. I can't wait to see it in action.\n\nCheers!\nDave\n"},{"id":"424553","messageId":"875yzlsngv.fsf@evledraar.gmail.com","threadId":"55645","inReplyTo":"xmqqo8de9wis.fsf@gitster.g","subject":"Re: Is the sha256 object format experimental or not?","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-05-14T08:49:42Z","receivedAt":"2021-05-14T08:56:25Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Fri, May 14 2021, Junio C Hamano wrote:\n\n> dwh@linuxprogrammer.org writes:\n>\n>> I think Git should externalize the calculation of object digests just\n>> like it externalizes the calcualtion of object digital signatures.\n>\n> The hashing algorithms used to generate object names has\n> requirements fundamentally different from that of digital\n> signatures.  I strongly suspect that that fact would change the\n> equation when you rethink what you said above.\n>\n> We can \"upgrade\" digital signature algorithms fairly easily---nobody\n> would complain if you suddenly choose different signing algorithm\n> over a blob of data, as long as all project participants are aware\n> (and self-describing datastream helps here) and are capable of\n> grokking the new algorithm we are adopting.  But because object\n> names are used by one object to refer to another, and most\n> importantly, we do not want a single object to have multiple names,\n> we cannot afford to introduce a new hashing algorithm every time we\n> feel like it.  In other words, diversity of object naming algorithms\n> is to be avoided as much as possible, while diversity of signature\n> algorithms is naturally expected.\n\nI agree insofar that I don't see a good reason for us to support some\nplethora of hash algorithms, but I wouldn't have objections to adding\nmore if people find them useful for some reason. See e.g. [1] for an\nimplementation.\n\nBut I really don't see how anything you've said would present a\ntechnical hurdle once we have SHA-1<->SHA-256 interop in a good enough\nstate. At that point we'll support re-hashing on arrival of content\nhashed with algorithm X into Y, with a local lookup table between X<=>Y.\n\nSo if somebody wants to maintain content hashed with algorithm Z locally\nwe should easily be able to support that. The \"diversity of naming\"\nwon't matter past that local repository, any mention of Z will be\ntranslated to X or Y on fetch/push.\n\n1. https://lore.kernel.org/git/20191222064809.35667-1-michaeljclark@mac.com/\n"},{"id":"424582","messageId":"20210514134501.3vzgqdfwwejafkq7@meerkat.local","threadId":"55645","inReplyTo":"20210513234706.GG11882@localhost","subject":"Re: Is the sha256 object format experimental or not?","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2021-05-14T13:45:01Z","receivedAt":"2021-05-14T13:45:06Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"On Thu, May 13, 2021 at 04:47:06PM -0700, dwh@linuxprogrammer.org wrote:\n> On 13.05.2021 16:49, Konstantin Ryabitsev wrote:\n> > Check out what we're doing as part of patatt and b4:\n> > https://pypi.org/project/patatt/\n> > \n> > It takes your keyring-in-git idea and runs with it -- it would be good to have\n> > your input while the project is still young and widely unknown. :)\n> \n> Konstantin:\n> \n> That's really clever. I especially love how you're using the list\n> archive as the provenance log of old keys developers used. That seems\n> like it would work although I have worries about the security of\n> X-Developer-Key and the lack of key history immediately available to\n> `git log` because it's in the list archive and not in the repo directly.\n>\n> I guess the old keys would still be in your local keyring for `gpg` to\n> use but it would mark signatures created with old revoked keys as\n> invalid even though they are valid.\n\nThanks for taking a look at it. I don't view this as much of a problem, since\nthe goal for the project is specifically end-to-end patch attestation. For git\ncommits, if they are signed with a key from the in-git keyring, it would\nactually be really straightforward to get the valid key at the time of signing\n-- you just retrieve the keyring using the date of the commit.\n\n> My approach has been to move to cryptographically secure provenance logs\n> that contain key rotation events and commitments to future keys and also\n> cryptographically linking to arbitrary metadata (e.g. KYC proofs, etc).\n> The file format is documented using the Community Standard template from\n> the LF. I'm hoping to move Git to use external tools for all digest and\n> digital signature operations. Then I can start putting provenance logs\n> into a \".well-known\" path in Git repos, maybe \".plogs\" or something.\n> Then I can write/adapt a signing tool to understand provenance logs\n> of public keys in the repo instead of the GPG keyring stuff we have\n> today.\n> \n> Provenance logs accumulate the full key history of a developer over\n> time. It represents a second axis of time such that the HEAD of a repo\n> will have the full key history, for every contributor available to\n> cryptographic tools for verifying signatures. This makes `git log\n> --show-signature` operations maximally efficient because we don't have\n> to check out old keyrings from history to recreate the state GPG was in\n> when the signature was created.\n\nHmm... I'm not sure if it's an inefficient operation in the first place. If\nthe keyring is in the same branch as the commit itself, then you can retrieve\nthe public key using \"git show [commit-sha]:path/to/that/pubkey\". If it's in a\ndifferent branch, then it's slightly more complicated because then you have to\nfind a keyring commit corresponding to the commit-date of the object you're\nchecking. In any case, these are all pretty fast operations in git.\n\n> I still like your approach purely for the \"it works right now\" aspect of\n> the solution. Good job. I can't wait to see it in action.\n\nAs you know, this is my third attempt at getting patch attestation off the\nground. The first one I implemented using detached attestation documents and\nit was clever and neat, but it was too complicated and failed to take off -- I\nthink mostly because a) it wasn't easy to understand what it's doing, and b)\nit required that people adjust their workflows too much.\n\nThe second attempt was better, but I think it was still too complicated,\nbecause it required that we parse patch content, making it fragile and slow on\nvery large patch sets.\n\nI'm hoping that this version resolves the downsides of the previous two\nattempts by both being dumb and simple and by only requiring a simple one-time\nsetup (via the sendemail-validate hook) with no further changes to the usual\ngit-send-email workflow after that.\n\nI've not yet widely promoted this, as patatt is a very new project, but I'm\nhoping to start reaching out to people to trial it out in the next few weeks.\n\nThanks,\n-K\n"},{"id":"424594","messageId":"20210514173909.GA16542@localhost","threadId":"55645","inReplyTo":"20210514134501.3vzgqdfwwejafkq7@meerkat.local","subject":"Re: Is the sha256 object format experimental or not?","fromName":"","fromEmail":"dwh@linuxprogrammer.org","sentAt":"2021-05-14T17:39:09Z","receivedAt":"2021-05-14T17:39:14Z","isPatch":false,"sender":{"key":"dwh@linuxprogrammer.org","avatar":null},"body":"On 14.05.2021 09:45, Konstantin Ryabitsev wrote:\n>As you know, this is my third attempt at getting patch attestation off the\n>ground. \n\nYes, I've been following. It's been a long road.\n\n>I'm hoping that this version resolves the downsides of the previous two\n>attempts by both being dumb and simple and by only requiring a simple one-time\n>setup (via the sendemail-validate hook) with no further changes to the usual\n>git-send-email workflow after that.\n\nI'm very interested in whether this one works. You and I are completely\naligned on this. I don't think I'm paying enough attention to the\nemailed patch attestations as you have. I think I understand the\nrequirements but maybe not all of them. Do you have any threads on\npublic-inbox where you discuss them? I want to make sure that what I'm\ndoing doesn't undermine anything you're trying to do. The end goal is to\nhave an air-tight provenance on all contributions and\naccountable/audtiable software supply chain. We're all working towards\nthat.\n\n>I've not yet widely promoted this, as patatt is a very new project, but I'm\n>hoping to start reaching out to people to trial it out in the next few weeks.\n\nHopefully this approach strikes the right balance.\n\nCheers!\nDave\n"},{"id":"424595","messageId":"20210514181009.GB16542@localhost","threadId":"55645","inReplyTo":"875yzlsngv.fsf@evledraar.gmail.com","subject":"Re: Is the sha256 object format experimental or not?","fromName":"","fromEmail":"dwh@linuxprogrammer.org","sentAt":"2021-05-14T18:10:09Z","receivedAt":"2021-05-14T18:10:14Z","isPatch":false,"sender":{"key":"dwh@linuxprogrammer.org","avatar":null},"body":"On 14.05.2021 10:49, Ævar Arnfjörð Bjarmason wrote:\n>I agree insofar that I don't see a good reason for us to support some\n>plethora of hash algorithms, but I wouldn't have objections to adding\n>more if people find them useful for some reason. See e.g. [1] for an\n>implementation.\n\nI think Git should not try to do any cryptographic operations at all and\nrely on external tools that are implemented properly and hardended.\nImplementing cryptography isn't just about translating the algorithm\ninto code but also getting memory security correct, file handling\ncorrect, input security correct, control flow correct (equal cost\nmulti-path), etc, etc. Most of the cryptography libraries aren't\ndesigned to be misuse resistant. The only one I know of that has that as\na top-line requirement is Hyperledger Ursa [1].\n\nI would like to see us remove all cryptography code (e.g. digests,\ndigital signatures, etc) from Git and rely on external tools entirely.\nIf we store the cryptographic material in a self-describing format that\nidentifies the associated tool as well as the cryptographic data, then\nGit can be completely agnostic.\n\n>But I really don't see how anything you've said would present a\n>technical hurdle once we have SHA-1<->SHA-256 interop in a good enough\n>state. At that point we'll support re-hashing on arrival of content\n>hashed with algorithm X into Y, with a local lookup table between X<=>Y.\n>\n>So if somebody wants to maintain content hashed with algorithm Z locally\n>we should easily be able to support that. The \"diversity of naming\"\n>won't matter past that local repository, any mention of Z will be\n>translated to X or Y on fetch/push.\n\nUsing self-describing formats allows us to honor history and keep old\nobject names as they and eliminate all of this added complications you\ndescribe. I think there is a lot of room for errors to creep in when\ncollaborators have copies of the same repo and they have local mappings\nbetween different hashing algorithms. How is this not setting up for a\ncombinatorial explosion of data? If the canonical repo uses SHA1 and one\ncontributor uses SHA2-512, another uses Blake2b-256, and yet another\nuses SHA3-384, won't they all have to maintain six different translation\ntables for all objects? SHA1 <=> SHA2-512, SHA1 <=> Blake2b-256, SHA1\n<=> SHA3-384, SHA2-512 <=> Blake2b-256, SHA2-512 <=> SHA3-384, and\nBlake2b-256 <=> SHA3-384? I guess that's your motivation for not\nallowing algorithmic agility.\n\nThe way around this is to use self-describing formats and external\ntools. Git repo copies wouldn't be required to have only *one* algorithm\nnaming all objects, requiring the translation tables. Instead Git repos\nwould/could have heterogeneous object names, each one with a single name\ngenerated with a different digest algorithm. Git would simply consider\nthose names as plain strings and validating those strings requires\ntalking to the correct external tool, sending the name string and the\nobject data and reading back the result.\n\nI think this is a much better approach because:\n\n1. It creates algorithmic agility in a way that isn't top-down and heavy\nhanded.\n\n2. It eliminates the need for all of the translation tables and\nround-tripping complexity.\n\n3. It empowers maintainers to decide which algorithms can/must be used\nwhen naming objcts in a given repo. Merge hooks, CI/CD checks and\netiquette guides can be used to enforce this.\n\n4. Git's attack surface becomes smaller (a very good thing) and limited\nto doing IPC to external tools correctly and securely (easy) instead of\ntrying to get cryptography client code correct (very difficult).\n\nOne other thing to consider is that there are new tools being developed\nthat do similar things as Git that do have algorithmic agility and use\nself-describing cryptographic primitives. Late-binding trust is now a\nbest practice and has been for quite some time. Many people rely upon\nGit and I think we should keep up with the best practices.\n\nCheers!\nDave\n"},{"id":"424824","messageId":"YKNRg691mu8x7Pua@google.com","threadId":"55645","inReplyTo":"20210513202919.GE11882@localhost","subject":"Re: Is the sha256 object format experimental or not?","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2021-05-18T05:32:51Z","receivedAt":"2021-05-18T05:32:58Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\ndwh@linuxprogrammer.org wrote:\n\n> I think we should make one last breaking change for digests and not go\n> with the existing SHA-256 implementation but instead switch to\n> self-describing digests and digital signatures and rely on external\n> tools that Git talks to using a standard protocol. We can maintain full\n> backward compatibility and even support full round tripping using some\n> of the similar techniques that Brian came up with.\n\nForgive my ignorance: can you describe what compatibility break you\nmean?  Do you mean _removing_ support for gpgsig-sha256?  If so,\nwhy --- couldn't you get the same benefit by introducing the new\nfunctionality you're describing without getting rid of historical\nfunctionality at the same time?\n\nA nice thing about signatures is that they don't change the semantics\nof the object.  So some future version of Git can remove support for\nverifying them, if they turn out\n\nBy the way, to be clear, the hash-function-transition doc in\nDocumentation/technical/ is not by Brian alone.  It is the result of\ncollaboration by various people on list (see its git history for\ndetails).\n\n[...]\n> object 04b871796dc0420f8e7561a895b52484b701d51a\n> obj 0ED_zgYrQg584bCrqKPoUvxaQ5aMis0GtnW_NrZFTTxUlHLUOyp77LanoZEGV6ajhYGLGTaTfCIQhryovyeNFJuG\n> type commit\n> tag signedtag\n> tagger C O Mitter <committer@example.com> 1465981006 +0000\n> signtype openpgp\n> sign LS0tLS1CRUdJTiBQR1AgU0lHTkFUVVJFLS0tLS0KVmVyc2lvbjogR251UEcgdjEKCmlRRWN\n> CQUFCQWdBR0JRSlhZUmhPQUFvSkVHRUpMb1czSW5HSmtsa0lBSWNuaEw3UndFYi8rUWVYOWVua1\n> hoeG4KcnhmZHFydldkMUs4MHNsMlRPdDhCZy9OWXdyVUJ3L1JXSitzZy9oaEhwNFd0dkUxSERHS\n> GxrRXozeTExTGt1aAo4dFN4UzNxS1R4WFVHb3p5UEd1RTkwc0pmRXhoWmxXNGtuSVExd3QveVdx\n> TSszM0U5cE40aHpQcUx3eXJkb2RzCnE4RldFcVBQVWJTSlhvTWJSUHcwNFM1anJMdFpTc1VXYlJ\n> Zam1KQ0h6bGhTZkZXVzRlRmQzN3VxdUlhTFVCUzAKcmtDM0pyeDc0MjBqa0lwZ0ZjVEkyczYwdW\n> hTUUx6Z2NDd2RBMnVrU1lJUm5qZy96RGtqOCszaC9HYVJPSjcyeApsWnlJNkhXaXhLSmtXdzhsR\n> TlhQU9EOVRtVFc5c0ZKd2NWQXptQXVGWDJrVXJlRFVLTVpkdUdjb1JZR3BEN0U9Cj1qcFhhCi0t\n> LS0tRU5EIFBHUCBTSUdOQVRVUkUtLS0tLQo\n[...]\n> I think a good move to make right now would be to add a general function\n> for stripping out any number of named fields from objects and also\n> stripping out in-body signatures found in tags. That way we can add\n> support in today's Git for stripping out fields/data for things like\n> creating/verifying the object digest and/or digital signature.\n\nCan you say a little more about the user-facing model here?  How does\na user know whether the signature verification result they're looking\nat describes the part of the object they care about or has stripped it\nout?\n\n[...]\n> Let me try to lay out the case for making a breaking change to sha256\n> right now that will future-proof repos going forward.\n>\n> It has been known for a few decades now that cryptography has a\n> shelf-life.\n\nYes, this is a key assumption of the hash function transition.  It is\nmeant to be repeatable, so that we are not stuck on a particular\ncryptographic hash.\n\n[...]\n>                                       The end result is that in high\n> security software, SHA-256 is being replaced with SHA-3 and Blake2\n> digests.\n\nDo you mean that practice is drifting away from the conclusion of\nhttps://www.imperialviolet.org/2017/05/31/skipsha3.html?  Where can I\nread more?\n\nIt took a while to decide on sha256 as the hash for Git to use to\nreplace sha1.  The process involved useful feedback from Keccak team\nand others, and I feel pretty comfortable with how thoroughly it was\ndiscussed, though of course I wouldn't be surprised if the state of\ncryptanalysis has changed in some way since then.\n\nThe front runners were from the SHA2, SHA3, and Blake2 families.  The\nmain factor that led to deciding on SHA2 is the wide availability of\nefficient and trustworthy implementations, in hardware and software.\nSee https://lore.kernel.org/git/alpine.DEB.2.21.1.1706151122180.4200@virtualbox/#t\nand https://lore.kernel.org/git/20180609224913.GC38834@genre.crustytoothpaste.net/#t\nfor some of the discussion that led there.\n\n[...]\n> 4. switch to \"late binding\", \"self describing\" cryptographic constructs.\n\nAs Junio mentioned, Git does not impose a requirement on the signature\nalgorithm used in a signature block, including the digest involved.\nHowever, signing history typically involves signing object names, and\nobject names use a cryptographic hash for other reasons.  If we want\nGit to stop using a content addressable object store, that would be a\nmore fundamental changes to its design.\n\nThanks and hope that helps,\nJonathan\n"}]}