{"thread":{"id":"51137","subject":"RFC: Separate commit identification from Merkle hashing","startedAt":"2019-05-21T01:32:52Z","lastAt":"2019-05-23T21:50:30Z","messageCount":13,"participants":["Eric S. Raymond","Jonathan Nieder","Jakub Narebski","Randall S. Becker","Ævar Arnfjörð Bjarmason"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"375974","messageId":"20190521013250.3506B470485F@snark.thyrsus.com","threadId":"51137","inReplyTo":null,"subject":"RFC: Separate commit identification from Merkle hashing","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-21T01:32:50Z","receivedAt":"2019-05-21T01:32:52Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"I have been thinking hard about the problems raised during my\nrequest for unique timestamps.  I think I've found a better way\nto bust the box I was trying to break out of.  I am therefore\nwithdrawing that proposal and replacing it with this one.\n\nIt's time to separate commit identification from Merkle hashing.\n\nOne reason I am sure of this is the SHA-1 to whatever transition.\nWe can't count on the successor hash to survive attack forever.\nAccordingly, git's design needs to be stable against the possibility\nof having to accommodate multiple future hash algorithms in the\nfuture.\n\nHere's how to do it:\n\n1. Commit IDs and Merkle-tree hashes become separate commit\n   properties in the git filesystem.\n\n2. The data structure representing a Merkle-tree hash becomes\n   a pair consisting of a value and a hash-algorithm tag. An\n   empty tag is interpreted as SHA-1. I will call this entity the\n   \"verification hash\" and avoid unqualified use of \"hash\" in the\n   rest of this proposal.\n\n3. The initial value of a commit's ID in a live repository is a copy\n   of its verification hash, except in one important case.\n\n4. When a repository is exported to a stream, the commit-id is dumped\n   with other commit metadata.  Thus, anything that can read a stream\n   can resolve commit references in its change comments.\n\n5. When a stream is imported, if a commit has a commit-id field it\n   overrides the default assignment of the generated verification hash\n   to that field.\n\n6. Commit IDs are free-format and not interpreted by git except\n   as lookup keys. When git changes verification-hash functions,\n   commit IDs do not change.\n\nNotice several important properties of this design.\n\nA. Git becomes absolutely future-proofed against hash-algorithm\n   changes. It can even support the use of multiple hash types over\n   the lifetime of one repo.\n\nB. All SHA-1 commit references will resolve forever even after git\n   stops generating them.  All future hash-based commit references will\n   also be good forever.\n\nC. The id/verification split will be invisible from clients at start,\n   because initially they coincide and will continue to do so unless\n   an explicit decision changes either the verification-hash algorithm\n   or the way commit-IDs are initialized.\n\nD. My wish for forward-portable unique commit IDs is granted.\n   They're not by default eyeball-friendly, but I can live with that.\n   Furthermore, because they're preserved in streams they can be\n   eternally stable even as hash algorithms and preferred ID\n   formats change.\n\nE. There is now a unique total order on the repo, modulo highly\n   unlikely (and in priciple completely avoidable) commit-ID\n   collisions. It's commit date tie-broken by commit-ID sort order.\n   It too survives hash-function changes.\n\nF. There's no need for timestamp uniqueness any more.\n\nG. When a repository is imported from (say) Subversion, the Subversion\n   IDs *don't have to break*!  They can be used to initialize the\n   commit-ID fields. Many users migrating from other VCSes will be\n   deeply, deeply grateful for this feature.\n\nI believe this solves every problem I walked in with except timestamp\ntruncation.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\nProbably fewer than 2% of handguns and well under 1% of all guns will\never be involved in a violent crime. Thus, the problem of criminal gun\nviolence is concentrated within a very small subset of gun owners,\nindicating that gun control aimed at the general population faces a\nserious needle-in-the-haystack problem.\n\t-- Gary Kleck, \"Point Blank: Handgun Violence In America\"\n"},{"id":"375975","messageId":"20190521015703.GB32230@google.com","threadId":"51137","inReplyTo":"20190521013250.3506B470485F@snark.thyrsus.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2019-05-21T01:57:03Z","receivedAt":"2019-05-21T01:57:07Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi!\n\nEric S. Raymond wrote:\n\n> One reason I am sure of this is the SHA-1 to whatever transition.\n> We can't count on the successor hash to survive attack forever.\n> Accordingly, git's design needs to be stable against the possibility\n> of having to accommodate multiple future hash algorithms in the\n> future.\n\nHave you read through Documentation/technical/hash-function-transition?  It\ntakes the case where the new hash function is found to be weak into account.\n\nHope that helps,\nJonathan\n"},{"id":"375978","messageId":"20190521023832.GA130381@thyrsus.com","threadId":"51137","inReplyTo":"20190521015703.GB32230@google.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-21T02:38:32Z","receivedAt":"2019-05-21T02:38:34Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com>:\n> Hi!\n> \n> Eric S. Raymond wrote:\n> \n> > One reason I am sure of this is the SHA-1 to whatever transition.\n> > We can't count on the successor hash to survive attack forever.\n> > Accordingly, git's design needs to be stable against the possibility\n> > of having to accommodate multiple future hash algorithms in the\n> > future.\n> \n> Have you read through Documentation/technical/hash-function-transition?  It\n> takes the case where the new hash function is found to be weak into account.\n> \n> Hope that helps,\n> Jonathan\n\nReading now...\n\nAt first sight I think it looks pretty compatible with what I am proposing.\nThe goals anyway, some of the implementation tactics would change a bit.\n\nI think it's a weakness, though, that most of it is written as though it\nassumes only one hash transition will be necessary.  (This is me thinking\non long timescales again.)\n\nInstead of having a gpgsig-sha256 field, I would change the code so all\nhash cookies have an delimited optional prefix giving the hash-algorithm\ntype, with an absent prefix interpreted as SHA-1.\n\nI think the idea of mapping future hashes to SHA-1s, which are then\nused as fs lookup keys, is sound.  The same technique (probably the\nsame code!) could be used to map the otherwise uninterpreted\ncommit-IDs I'm proposing to lookup keys.\n\nI should have said in my previous mail that I'm prepared to put\nmy coding fingers into making all this happen. I am pretty sure my\ngramty manager will approve.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375979","messageId":"20190521025813.GA175422@google.com","threadId":"51137","inReplyTo":"20190521023832.GA130381@thyrsus.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2019-05-21T02:58:13Z","receivedAt":"2019-05-21T02:58:18Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nEric S. Raymond wrote:\n> Jonathan Nieder <jrnieder@gmail.com>:\n>> Eric S. Raymond wrote:\n\n>>> One reason I am sure of this is the SHA-1 to whatever transition.\n>>> We can't count on the successor hash to survive attack forever.\n[...]\n>> Have you read through Documentation/technical/hash-function-transition?  It\n>> takes the case where the new hash function is found to be weak into account.\n>>\n>> Hope that helps,\n>> Jonathan\n>\n> Reading now...\n\nTake your time. :)\n\n[...]\n> I think it's a weakness, though, that most of it is written as though it\n> assumes only one hash transition will be necessary.  (This is me thinking\n> on long timescales again.)\n\nHm, can you point to what part of the doc suggested that?  Best to make\nthe text clearer, to avoid confusing the next person.\n\nOn the contrary, the design is very careful to be able to support the\nnext transition.\n\n[...]\n>                                    The same technique (probably the\n> same code!) could be used to map the otherwise uninterpreted\n> commit-IDs I'm proposing to lookup keys.\n\nNo, since Git relies on commit IDs for integrity checking.  The hash\nfunction transition described in that document relies on\nround-tripping ability for the duration of the transition.\n\nJonathan\n"},{"id":"375981","messageId":"20190521033153.GA2909@thyrsus.com","threadId":"51137","inReplyTo":"20190521025813.GA175422@google.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-21T03:31:53Z","receivedAt":"2019-05-21T03:31:54Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com>:\n> > I think it's a weakness, though, that most of it is written as though it\n> > assumes only one hash transition will be necessary.  (This is me thinking\n> > on long timescales again.)\n> \n> Hm, can you point to what part of the doc suggested that?  Best to make\n> the text clearer, to avoid confusing the next person.\n\nI will reread it with an editorial eye and try to come up with\nconcrete suggestions, perhaps a patch. My relative ignorance\nshould actually be helpful here.\n\n> >                                    The same technique (probably the\n> > same code!) could be used to map the otherwise uninterpreted\n> > commit-IDs I'm proposing to lookup keys.\n> \n> No, since Git relies on commit IDs for integrity checking.  The hash\n> function transition described in that document relies on\n> round-tripping ability for the duration of the transition.\n\nI do not quite understand this comment yet. But I don't think it\nmatters that I don't, and I will by the time I write any code.  I\nexpect the worst case is that the separated IDs require a different\nlookup table from the hashes, but will resolve at the same speed.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"376110","messageId":"86h89lq96v.fsf@gmail.com","threadId":"51137","inReplyTo":"20190521013250.3506B470485F@snark.thyrsus.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2019-05-23T19:09:44Z","receivedAt":"2019-05-23T19:09:52Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"esr@thyrsus.com (Eric S. Raymond) writes:\n\n> I have been thinking hard about the problems raised during my\n> request for unique timestamps.  I think I've found a better way\n> to bust the box I was trying to break out of.  I am therefore\n> withdrawing that proposal and replacing it with this one.\n>\n> It's time to separate commit identification from Merkle hashing.\n\nDocumentation/technical/hash-function-transition.txt identifies similar\nproblem, namely that existing signatures in signed tags, signed commits\nand merges of signed tags are signatures of their SHA-1 form.  We want\nto be able to verify those signatures, even if this verification may be\nconsidered less secure now.\n\nYou want both more (stable IDs for all commits, not only those signed)\nand less (you don't need verification down the tree using IDs used for\ncommit ID).\n\n> One reason I am sure of this is the SHA-1 to whatever transition.\n> We can't count on the successor hash to survive attack forever.\n> Accordingly, git's design needs to be stable against the possibility\n> of having to accommodate multiple future hash algorithms in the\n> future.\n>\n> Here's how to do it:\n>\n> 1. Commit IDs and Merkle-tree hashes become separate commit\n>    properties in the git filesystem.\n\nThe issue you need to consider is that for signatures to be secure they\nmust be over verification-hash Merkle-tree.  It is not only commits that\nare identified by hashes, but also trees, blobs and tags.\n\nCommits reference other commits (\"parent\" lines) and a tree (\"tree\");\ntrees reference other trees, blobs and possibly commits (if submodules\nare used).  Tags can reference any object, but most common reference\ncommits.  Blobs, i.e. file contents, do not reference any other\nobjects.  For security, all those references should use most strong hash\nfunction.\n\nChanging referecing hash (e.g. \"parent\" uses SHA-256 instead of \"SHA-1\")\nmeans that the contents of object changes, and thus its hash.\nDocumentation/technical/hash-function-transition.txt therefore talks\nabout SHA-256 and SHA-1 forms and SHA-256 and SHA-1 object names.\n\n \"The sha1-name of an object is the SHA-1 of the concatenation of its\n  type, length, a nul byte, and the object's sha1-content. This is the\n  traditional <sha1> used in Git to name objects.\n\n  The sha256-name of an object is the SHA-256 of the concatenation of its\n  type, length, a nul byte, and the object's sha256-content.\"\n\n\n> 2. The data structure representing a Merkle-tree hash becomes\n>    a pair consisting of a value and a hash-algorithm tag. An\n>    empty tag is interpreted as SHA-1. I will call this entity the\n>    \"verification hash\" and avoid unqualified use of \"hash\" in the\n>    rest of this proposal.\n\nCurrently Git makes use of the fact that SHA-1 and SHA-256 identifiers\nare of different lengths to distinguish them (see section \"Meaning of\nsignatures\") in Documentation/technical/hash-function-transition.txt\n\nThere might be, I think, the problem for \"tree\" objects.  As opposed to\nall other places, \"tree\" objects use binary representation of hash, and\nnot hexadecimal textual representation (some consider that a design\nmistake).\n\n>\n> 3. The initial value of a commit's ID in a live repository is a copy\n>    of its verification hash, except in one important case.\n>\n> 4. When a repository is exported to a stream, the commit-id is dumped\n>    with other commit metadata.  Thus, anything that can read a stream\n>    can resolve commit references in its change comments.\n>\n> 5. When a stream is imported, if a commit has a commit-id field it\n>    overrides the default assignment of the generated verification hash\n>    to that field.\n\nI think Documentation/technical/hash-function-transition.txt misses\nconsiderations for fast-import format (it talks about problem with\nsubmodules, shallow clones, and currently not solved problem of\ntranslating notes; it does not talk about git-replace, either).\n\n>\n> 6. Commit IDs are free-format and not interpreted by git except\n>    as lookup keys. When git changes verification-hash functions,\n>    commit IDs do not change.\n\nAll right.  Looks sensible on first glance.\n\nFor security, all references in Merkle-tree of hashes must use strong\nverification hash.  This means that you need to be able to refer to any\nobject, including commit, by its verification hash name of its\nverification hash form (where all references inside object, like\n\"parent\" and \"tree\" headers in commit objects, use verification hashes).\n\nYou need to store this commit ID somewhere.  Current proposal for\ntransitional period in Documentation/technical/hash-function-transition.txt\ntalks about loose object index ($GIT_OBJECT_DIR/loose-object-idx) with\nthe following format:\n\n  # loose-object-idx\n  (sha256-name SP sha1-name LF)*\n\nIn packfile index contains separate SHA-1 indices and SHA-256 indices\ninto packfile, providing fast mapping from SHA-1 name or SHA-256 name to\nposition (index) of object in the packfile.\n\nSomething similar might have been needed for commit IDs mapping.\n\nOne problem is that neither loose object index, not the packfile index\nare transported alongside with the objects.  So we may need to put\ncommit ID elsewhere...\n\nNote that we cannot put X-hash identifier into X-hash object form, that\nis you cannot add \"id\" header to object (though you might add \"other-id\"\nheader, assuming that if ID is hash based it is on the other-id form\nwithout other-id header).\n\n  id <sha-1 identifier of this object>\n  tree 0fa044a4d161254a3eae0bd06c0452d79e489593\n  parent 6505413ad94ddfc01f9e2f5c1b79ea6b8ffbabbb\n  author A U Thor <author@example.com> 1558619302 +0200\n  committer C O Mitter <committer@example.com> 1558628753 -0500\n\n  fixes\n\n\n> Notice several important properties of this design.\n>\n> A. Git becomes absolutely future-proofed against hash-algorithm\n>    changes. It can even support the use of multiple hash types over\n>    the lifetime of one repo.\n>\n> B. All SHA-1 commit references will resolve forever even after git\n>    stops generating them.  All future hash-based commit references will\n>    also be good forever.\n\nWe might need to be able to distinguish commit IDs from hash-based\nobject identifier of commit on command line, perhaps with something like\n\n  <commit-id>^{id}\n\nThis is similar to proposed\n\n  git --output-format=sha1 log abac87a^{sha1}..f787cac^{sha256}\n\n> C. The id/verification split will be invisible from clients at start,\n>    because initially they coincide and will continue to do so unless\n>    an explicit decision changes either the verification-hash algorithm\n>    or the way commit-IDs are initialized.\n\nThe problem may be with reusing command output for input (to refer to\nobjects and commits).\n\n>\n> D. My wish for forward-portable unique commit IDs is granted.\n>    They're not by default eyeball-friendly, but I can live with that.\n>    Furthermore, because they're preserved in streams they can be\n>    eternally stable even as hash algorithms and preferred ID\n>    formats change.\n\nGood.\n\n>\n> E. There is now a unique total order on the repo, modulo highly\n>    unlikely (and in priciple completely avoidable) commit-ID\n>    collisions. It's commit date tie-broken by commit-ID sort order.\n>    It too survives hash-function changes.\n\nNice.\n\n>\n> F. There's no need for timestamp uniqueness any more.\n>\n> G. When a repository is imported from (say) Subversion, the Subversion\n>    IDs *don't have to break*!  They can be used to initialize the\n>    commit-ID fields. Many users migrating from other VCSes will be\n>    deeply, deeply grateful for this feature.\n\nThere would also need to be some support to retrieve commits using their\n\"commit ID\" stable identifiers.  It may not need to be very fast.\n\n>\n> I believe this solves every problem I walked in with except timestamp\n> truncation.\n\nBest,\n--\nJakub Narębski\n"},{"id":"376116","messageId":"20190523200929.GA70860@google.com","threadId":"51137","inReplyTo":"86h89lq96v.fsf@gmail.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2019-05-23T20:09:29Z","receivedAt":"2019-05-23T20:09:34Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Hi,\n\nJakub Narebski wrote:\n\n> I think Documentation/technical/hash-function-transition.txt misses\n> considerations for fast-import format (it talks about problem with\n> submodules, shallow clones, and currently not solved problem of\n> translating notes; it does not talk about git-replace, either).\n\nHm, can you say more?  I think fast-import is not significantly\ndifferent from other tools that want to pick an appropriate object\nformat for input and an appropriate object format for output.\n\nDo you mean that the fast-import file should have a field for\nexplicitly specifying the input object format, and that that doc\nought to call it out?\n\n[...]\n> For security, all references in Merkle-tree of hashes must use strong\n> verification hash.  This means that you need to be able to refer to any\n> object, including commit, by its verification hash name of its\n> verification hash form (where all references inside object, like\n> \"parent\" and \"tree\" headers in commit objects, use verification hashes).\n\nThis kind of crypto agility weakens any guarantees that rely on\nstrength of a hash function.  The security level would be that of the\nweakest of the supported hash functions.\n\nIn other words, usually the benefit of supporting multiple hash\nfunctions as a reader is that you want the strength of the strongest\nof those hash functions and you need a migration path to get there.\nIf you don't have a way to eventually drop support for the weaker\nhashes, then what benefit do you get from supporting multiple hash\nfunctions?\n\nJonathan\n"},{"id":"376122","messageId":"20190523205009.GA69096@thyrsus.com","threadId":"51137","inReplyTo":"86h89lq96v.fsf@gmail.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-23T20:50:09Z","receivedAt":"2019-05-23T20:50:13Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jakub Narebski <jnareb@gmail.com>:\n> You want both more (stable IDs for all commits, not only those signed)\n> and less (you don't need verification down the tree using IDs used for\n> commit ID).\n\nThat's right.  My assumption is that future VCSes will do their own\nhash chaining in ways we don't really want to try to anticipate or\nconstrain.\n\n> Currently Git makes use of the fact that SHA-1 and SHA-256 identifiers\n> are of different lengths to distinguish them (see section \"Meaning of\n> signatures\") in Documentation/technical/hash-function-transition.txt\n\nThat's the obvious hack.  As a future-proofing issue, though, I think\nit would be unwise to count on all future hashes being of distinguishable\nlengths. Explicit algorithm tagging is better, at least internally.\n\n> There might be, I think, the problem for \"tree\" objects.  As opposed to\n> all other places, \"tree\" objects use binary representation of hash, and\n> not hexadecimal textual representation (some consider that a design\n> mistake).\n\nI'm inclined to agree that it was a mistake.  But whether it gets\nreplaced by a binary struct holding an {algorithm-tag,value} pair or a\ntextual representation of same is not something I care about a lot.\n\n> I think Documentation/technical/hash-function-transition.txt misses\n> considerations for fast-import format\n\nYou can count on me to stay on top of that; fast-import format is utterly\ncritical to how reposurgeon works, so I have a strong incentive to make\nsure it stays healthy.\n\n(Some of you may not know - reposurgeon solves the thorny problems of\nediting repositories by sidestepping to the textual serialized\nrepresentation of them.  It's basically a structure editor for\nfast-import streams that fools the outside world into thinking it\nedits live repositories by having importers and exporters at either\nend of its data flow.)\n\n> All right.  Looks sensible on first glance.\n\nI am very relieved to hear that. My view of git is outside-in; I was quite\nworried I might have missed some crucial issue.\n\n> For security, all references in Merkle-tree of hashes must use strong\n> verification hash.  This means that you need to be able to refer to any\n> object, including commit, by its verification hash name of its\n> verification hash form (where all references inside object, like\n> \"parent\" and \"tree\" headers in commit objects, use verification hashes).\n\nFair enough. One minor way in which my thinking has evolved since\nI wrote the RFC is that I now think it might be fruitful not to throw away\nthe idea of the verification hash as naming a commit, but rather to think\nof the separated commit-ID as an alias for the verification hash.\n\nThis reframing won't make any difference to the code, but it clarifies\nwhat to do if, for example, an import stream declares the same commit\nID for multiple commits, or fails to declare a commit ID at all.  In both\ncases the commit is still uniquely named by its verification hash. Commit-ID\nnamespace-management failures become annoying but not critical.\n\n> You need to store this commit ID somewhere.  Current proposal for\n> transitional period in Documentation/technical/hash-function-transition.txt\n> talks about loose object index ($GIT_OBJECT_DIR/loose-object-idx) with\n> the following format:\n> \n>   # loose-object-idx\n>   (sha256-name SP sha1-name LF)*\n> \n> In packfile index contains separate SHA-1 indices and SHA-256 indices\n> into packfile, providing fast mapping from SHA-1 name or SHA-256 name to\n> position (index) of object in the packfile.\n\nI would generalize this to something like\n\n(hash-algorithm-tag:value SP sha1-name LF)\n\n> Something similar might have been needed for commit IDs mapping.\n\nI think so, yes.\n\n> One problem is that neither loose object index, not the packfile index\n> are transported alongside with the objects.  So we may need to put\n> commit ID elsewhere...\n> \n> Note that we cannot put X-hash identifier into X-hash object form, that\n> is you cannot add \"id\" header to object (though you might add \"other-id\"\n> header, assuming that if ID is hash based it is on the other-id form\n> without other-id header).\n> \n>   id <sha-1 identifier of this object>\n>   tree 0fa044a4d161254a3eae0bd06c0452d79e489593\n>   parent 6505413ad94ddfc01f9e2f5c1b79ea6b8ffbabbb\n>   author A U Thor <author@example.com> 1558619302 +0200\n>   committer C O Mitter <committer@example.com> 1558628753 -0500\n> \n>   fixes\n\nImplementation details. Let's get the design right and properly specified\nbefore worrying too hard about this level of the problem.\n\nI may do another RFC about how to avoid having this problem ever\nagain.  In truth, I think git objects should have open property lists,\nlike bzr, with a property namespace reserved for system\nexpansion. That way, when you need objects to have new semantics, you\ncan do it without having an object-format flag day\n\n> > Notice several important properties of this design.\n> >\n> > A. Git becomes absolutely future-proofed against hash-algorithm\n> >    changes. It can even support the use of multiple hash types over\n> >    the lifetime of one repo.\n> >\n> > B. All SHA-1 commit references will resolve forever even after git\n> >    stops generating them.  All future hash-based commit references will\n> >    also be good forever.\n> \n> We might need to be able to distinguish commit IDs from hash-based\n> object identifier of commit on command line, perhaps with something like\n> \n>   <commit-id>^{id}\n> \n> This is similar to proposed\n> \n>   git --output-format=sha1 log abac87a^{sha1}..f787cac^{sha256}\n\nReasonable.\n\n> > C. The id/verification split will be invisible from clients at start,\n> >    because initially they coincide and will continue to do so unless\n> >    an explicit decision changes either the verification-hash algorithm\n> >    or the way commit-IDs are initialized.\n> \n> The problem may be with reusing command output for input (to refer to\n> objects and commits).\n\nSolvable, I think.\n\n> > D. My wish for forward-portable unique commit IDs is granted.\n> >    They're not by default eyeball-friendly, but I can live with that.\n> >    Furthermore, because they're preserved in streams they can be\n> >    eternally stable even as hash algorithms and preferred ID\n> >    formats change.\n> \n> Good.\n\nOh, man, you have no idea how good yet.  You won't until you've done a\nfew repo conversions yourself.\n\n/me needs a cross-eyed emoji here\n\n> > E. There is now a unique total order on the repo, modulo highly\n> >    unlikely (and in priciple completely avoidable) commit-ID\n> >    collisions. It's commit date tie-broken by commit-ID sort order.\n> >    It too survives hash-function changes.\n> \n> Nice.\n\nOne thing I will commit to do if we get this far is write the fast-export\ncode that does canonical order.  I need this badly for reposurgeon tests.\n\n> > F. There's no need for timestamp uniqueness any more.\n> >\n> > G. When a repository is imported from (say) Subversion, the Subversion\n> >    IDs *don't have to break*!  They can be used to initialize the\n> >    commit-ID fields. Many users migrating from other VCSes will be\n> >    deeply, deeply grateful for this feature.\n> \n> There would also need to be some support to retrieve commits using their\n> \"commit ID\" stable identifiers.  It may not need to be very fast.\n\nAgreed.\n\nOK, what do we do next?  Who needs to sign off on this?  Should I prepare\nan edit for the hash-function-transition.txt describing the splitting off\nof commit IDs?\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"376123","messageId":"20190523205313.GB69096@thyrsus.com","threadId":"51137","inReplyTo":"20190523200929.GA70860@google.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-23T20:53:13Z","receivedAt":"2019-05-23T20:53:15Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com>:\n> In other words, usually the benefit of supporting multiple hash\n> functions as a reader is that you want the strength of the strongest\n> of those hash functions and you need a migration path to get there.\n> If you don't have a way to eventually drop support for the weaker\n> hashes, then what benefit do you get from supporting multiple hash\n> functions?\n\nNot losing the capability to verify old parts of histories up to the\nstrength of the old hash algorithm.  Not perfect, but better than nothing.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"376124","messageId":"20190523205457.GC70860@google.com","threadId":"51137","inReplyTo":"20190523205009.GA69096@thyrsus.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2019-05-23T20:54:57Z","receivedAt":"2019-05-23T20:55:01Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Eric S. Raymond wrote:\n> Jakub Narebski <jnareb@gmail.com>:\n\n>> Currently Git makes use of the fact that SHA-1 and SHA-256 identifiers\n>> are of different lengths to distinguish them (see section \"Meaning of\n>> signatures\") in Documentation/technical/hash-function-transition.txt\n>\n> That's the obvious hack.  As a future-proofing issue, though, I think\n> it would be unwise to count on all future hashes being of distinguishable\n> lengths.\n\nWe're not counting on that.  As discussed in that section, future\nhashes can change the format.\n\n[...]\n>> All right.  Looks sensible on first glance.\n>\n> I am very relieved to hear that. My view of git is outside-in; I was quite\n> worried I might have missed some crucial issue.\n\nHonestly, I do think you have missed some fundamental issues.\nhttps://public-inbox.org/git/ab3222ab-9121-9534-1472-fac790bf08a4@gmail.com/\ndiscusses this further.\n\nRegards,\nJonathan\n"},{"id":"376129","messageId":"20190523211916.GA73150@thyrsus.com","threadId":"51137","inReplyTo":"20190523205457.GC70860@google.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-23T21:19:16Z","receivedAt":"2019-05-23T21:19:18Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jonathan Nieder <jrnieder@gmail.com>:\n> Honestly, I do think you have missed some fundamental issues.\n> https://public-inbox.org/git/ab3222ab-9121-9534-1472-fac790bf08a4@gmail.com/\n> discusses this further.\n\nHave re-read.  That was a different pair of proposals.\n\nI have abandoned the idea of forcing timestamp uniqueness entirely - that was\na hack to define a canonical commit order, and my new RFC describes a better\nway to get this.\n\nI still think finer-grained timestamps would be a good idea, but that is\nmuch less important than the different set of properties we can guarantee\nvia the new RFC.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"376130","messageId":"007a01d511af$f5dc1d50$e19457f0$@nexbridge.com","threadId":"51137","inReplyTo":"20190523211916.GA73150@thyrsus.com","subject":"RE: RFC: Separate commit identification from Merkle hashing","fromName":"Randall S. Becker","fromEmail":"rsbecker@nexbridge.com","sentAt":"2019-05-23T21:39:06Z","receivedAt":"2019-05-23T21:39:21Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On May 23, 2019 17:19, Eric S. Raymond wrote:\n> Jonathan Nieder <jrnieder@gmail.com>:\n> > Honestly, I do think you have missed some fundamental issues.\n> > https://public-inbox.org/git/ab3222ab-9121-9534-1472-\n> fac790bf08a4@gmai\n> > l.com/\n> > discusses this further.\n> \n> Have re-read.  That was a different pair of proposals.\n> \n> I have abandoned the idea of forcing timestamp uniqueness entirely - that\n> was a hack to define a canonical commit order, and my new RFC describes a\n> better way to get this.\n> \n> I still think finer-grained timestamps would be a good idea, but that is\nmuch\n> less important than the different set of properties we can guarantee via\nthe\n> new RFC.\n\nI don't think finer-grained timestamps will help long-term. The faster\nsystems get, the more resolution we need. At this point, I can easily get\ntwo commits within the same microsecond. The weird part is that if the\ncommits are done from two different CPUs on my platform, it is theoretically\npossible (although highly unlikely) that the second commit could be one\nmicrosecond earlier than the first commit, on the same file system, if a\ninter-CPU clock-sync had not been done for the past few seconds. On a\nbroader scale, that is somewhat obvious and assumes global time\nsynchronisation is maintained. It also makes me wonder what happens when git\nruns on a quantum computer and a commit goes to the wrong universe (joke).\n\nJust my $0.014\n\nRandall\n\n"},{"id":"376131","messageId":"87v9y0g7rz.fsf@evledraar.gmail.com","threadId":"51137","inReplyTo":"20190523205457.GC70860@google.com","subject":"Re: RFC: Separate commit identification from Merkle hashing","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-05-23T21:50:24Z","receivedAt":"2019-05-23T21:50:30Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, May 23 2019, Jonathan Nieder wrote:\n\n> Eric S. Raymond wrote:\n>> Jakub Narebski <jnareb@gmail.com>:\n>\n>>> Currently Git makes use of the fact that SHA-1 and SHA-256 identifiers\n>>> are of different lengths to distinguish them (see section \"Meaning of\n>>> signatures\") in Documentation/technical/hash-function-transition.txt\n>>\n>> That's the obvious hack.  As a future-proofing issue, though, I think\n>> it would be unwise to count on all future hashes being of distinguishable\n>> lengths.\n>\n> We're not counting on that.  As discussed in that section, future\n> hashes can change the format.\n\nI think both of you are also missing something that's implicit (but\nunfortunately not very explicitly talked about) in that document, which\nis that such hash transitions are assumed to have an out-of-bounds\ntemporal component to them.\n\nI.e. let's assume that instead of SHA-256 we're switching to SHA-X,\nwhich like SHA-1 is also a 20 byte hash function, so they're the same\nlength.\n\nYou'd then get a git.git with SHA-1 today, next year you'd have A\nSHA-1<->SHA-X mapping table, but the year after that we'd be fully on\nSHA-X for new content.\n\nSo even though we carry code and lookup table for looking up the old\nSHA-1 values we're not going to continue to pointlessly generate that\nbidirectional mapping forever. We'll have some sort of gravestone marker\nwhere we say \"past this point it's SHA-X only\".\n\nThat's not implemented or specified yet, but could e.g. be a magic ref\nof some sort advertised by the server, and the client would enforce that\nsuch a marker could only be made with the stronger hash function.\n\nThus a couple of years after that the SHA-1 -> SHA-X transition someone\ngenerating a colliding tag where a new good SHA-X tag *could* point to\nbad SHA-1 content won't be exploitable in practice. At that point\nclients won't be downloading SHA-1'd content or generating the mapping\ntable anymore.\n\nSo I don't see why a format change for the tags is needed, it would only\nmatter *if* we have a full collision *and* the hashes are the same\nlength (which we have no plan for), *and* if we assume we don't have\nsome other mitigations in play.\n"}]}