{"thread":{"id":"51105","subject":"Finer timestamps and serialization in git","startedAt":"2019-05-15T19:24:03Z","lastAt":"2019-05-21T01:05:24Z","messageCount":33,"participants":["Eric S. Raymond","Derrick Stolee","Ævar Arnfjörð Bjarmason","Jason Pyeron","Jeff King","Philip Oakley","Jakub Narebski","Michal Suchánek","Elijah Newren"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"375640","messageId":"20190515191605.21D394703049@snark.thyrsus.com","threadId":"51105","inReplyTo":null,"subject":"Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-15T19:16:05Z","receivedAt":"2019-05-15T19:24:03Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"The recent increase in vulnerability in SHA-1 means, I hope, that you\nare planning for the day when git needs to change to something like\nan elliptic-curve hash.  This means you're going to have a major\nformat break. Such is life.\n\nSince this is going to have to happen anyway, let me request two\nfunctional changes in git. Neither will be at all difficult, but the\nfirst one is also a thing that cannot be done without a format break,\nwhich is why I have not suggested them before.  They come from lots of\n(often painful) experience with repository conversions via\nreposurgeon.\n\n1. Finer granularity on commit timestamps.\n\n2. Timestamps unique per repository\n\nThe coarse resolution of git timestamps, and the lack of uniqueness,\nare at the bottom of several problems that are persistently irritating\nwhen I do repository conversions and surgery.\n\nThe most obvious issue, though a relatively superficial one, is that I have\nto thow away information whenever I convert a repository from a system with\nfiner-grained time.  Notably this is the case with Subversion, which keeps\ntime to milliseconds. This is probably the only respect in which its data\nmodel remains superior to git's. :-)\n\nThe deeper problem is that I want something from Git that I cannot\nhave with 1-second granularity. That is: a unique timestamp on each\ncommit in a repository. The only way to be certain of this is for git\nto delay accepting integration of a patch until it can issue a unique\ntime mark for it - obviously impractical if the quantum is one second,\nbut not if it's a millisecond or microsecond.\n\nWhy do I want this? There are number of reasons, all related to a\nmathematical concept called \"total ordering\".  At present, commits in\na Git repository only have partial ordering. One consequence is that\naction stamps - the committer/date pairs I use as VCS-independent commit\nidentifications in reposurgeon - are not unique.  When a patch sequence\nis applied, it can easily happen fast enough to give several successive\ncommits the same committer-ID and timestamp.\n\nOf course the commit hash remains a unique commit ID.  But it can't\neasily be parsed and followed by a human, which is a UX problem when\nit's used as a commit stamp in change comments.\n\nMore deeply, the lack of total ordering means that repository graphs\ndon't have a single canonical serialized form.  This sounds abstract\nbut it means there are surgical operations I can't regression-test\nproperly.  My colleague Edward Cree has found cases where git fast-export\ncan issue a stream dump for which git fast-import won't necessarily\nre-color certain interior nodes the same way when it's read back in\nand I'm pretty sure the absence of total ordering on the branch tips\nis at the bottom of that.\n\nI'm willing to write patches if this direction is accepted.  I've figured\nout how to make fast-import streams upward-compatible with finer-grained\ntimestamps.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n"},{"id":"375641","messageId":"ae62476c-1642-0b9c-86a5-c2c8cddf9dfb@gmail.com","threadId":"51105","inReplyTo":"20190515191605.21D394703049@snark.thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-05-15T20:16:00Z","receivedAt":"2019-05-15T20:16:05Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/15/2019 3:16 PM, Eric S. Raymond wrote:\n> The deeper problem is that I want something from Git that I cannot\n> have with 1-second granularity. That is: a unique timestamp on each\n> commit in a repository.\n\nThis is impossible in a distributed version control system like Git\n(where the commits are immutable). No matter your precision, there is\na chance that two machiens commit at the exact same moment on two different\nmachines and then those commits are merged into the same branch. Even\nwhen you specify a committer, there are many environments where a set\nof parallel machines are creating commits with the same identity.\n\n> Why do I want this? There are number of reasons, all related to a\n> mathematical concept called \"total ordering\".  At present, commits in\n> a Git repository only have partial ordering. \n\nThis is true of any directed acyclic graph. If you want a total ordering\nthat is completely unambiguous, then you should think about maintaining\na linear commit history by requiring rebasing instead of merging.\n\n> One consequence is that\n> action stamps - the committer/date pairs I use as VCS-independent commit\n> identifications in reposurgeon - are not unique.  When a patch sequence\n> is applied, it can easily happen fast enough to give several successive\n> commits the same committer-ID and timestamp.\n\nSorting by committer/date pairs sounds like an unhelpful idea, as that\ndoes not take any graph topology into account. It happens that commits\ncan actually have an _earlier_ commit date than its parent.\n\n> More deeply, the lack of total ordering means that repository graphs\n> don't have a single canonical serialized form.  This sounds abstract\n> but it means there are surgical operations I can't regression-test\n> properly.  My colleague Edward Cree has found cases where git fast-export\n> can issue a stream dump for which git fast-import won't necessarily\n> re-color certain interior nodes the same way when it's read back in\n> and I'm pretty sure the absence of total ordering on the branch tips\n> is at the bottom of that.\n\nIf you use `git rev-list --topo-order` with a fixed set of refs to start,\nthen the total ordering given is well-defined (and it is a linear\nextension of the partial order given by the commit graph). However, this\nordering is not stable: adding another merge commit may swap the order between\ntwo commits lower in the order.\n\n> I'm willing to write patches if this direction is accepted.  I've figured\n> out how to make fast-import streams upward-compatible with finer-grained\n> timestamps.\n\nChanging the granularity of timestamps requires changing the commit format,\nwhich is probably a non-starter. More universally-useful suggestions have\nbeen blocked due to keeping the file format consistent.\n\nThanks,\n-Stolee\n\n"},{"id":"375642","messageId":"871s0zwjv0.fsf@evledraar.gmail.com","threadId":"51105","inReplyTo":"20190515191605.21D394703049@snark.thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-05-15T20:20:03Z","receivedAt":"2019-05-15T20:20:08Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 15 2019, Eric S. Raymond wrote:\n\n> The recent increase in vulnerability in SHA-1 means, I hope, that you\n> are planning for the day when git needs to change to something like\n> an elliptic-curve hash.  This means you're going to have a major\n> format break. Such is life.\n\nNote that most users of Git (default build options) won't be vulnerable\nto the latest attack (or SHAttered), see\nhttps://public-inbox.org/git/875zqbx5yz.fsf@evledraar.gmail.com/T/#u\n\nBut yes the plan is to move to SHA-256. See\nhttps://github.com/git/git/blob/next/Documentation/technical/hash-function-transition.txt\n\n> Since this is going to have to happen anyway\n\nThe SHA-1 <-> SHA-256 transition is planned to happen, but there's some\nstrong opinions that this should be *only* for munging the content for\nhashing, not adding new stuff while we're at it (even if optional). See\n: https://public-inbox.org/git/87ftyyedqd.fsf@evledraar.gmail.com/\n\n> let me request two\n> functional changes in git. Neither will be at all difficult, but the\n> first one is also a thing that cannot be done without a format break,\n> which is why I have not suggested them before.  They come from lots of\n> (often painful) experience with repository conversions via\n> reposurgeon.\n>\n> 1. Finer granularity on commit timestamps.\n\nIf you wanted milli/micro/nano-second timestamps for commit objects or\nwhatever other new info then it doesn't need to break the commit header\nformat.\n\nYou put it key-values in the commit message and read it back out via\ngit-interpret-trailers.\n\nOr even put it in the header itself, e.g.:\n\nauthor <name> <epoch> <tz>\ncommitter <name> <epoch> <tz>\nx-author-ns <nanosecond part of author>\nx-committer-ns <nanosecond part of committer>\n\nOf course nobody would understand that new thing from day one, but\nthat's nothing compared to breaking the existing header format.\n\n> 2. Timestamps unique per repository\n>\n> The coarse resolution of git timestamps, and the lack of uniqueness,\n> are at the bottom of several problems that are persistently irritating\n> when I do repository conversions and surgery.\n>\n> The most obvious issue, though a relatively superficial one, is that I have\n> to thow away information whenever I convert a repository from a system with\n> finer-grained time.  Notably this is the case with Subversion, which keeps\n> time to milliseconds. This is probably the only respect in which its data\n> model remains superior to git's. :-)\n\nShould be solved by putting it in the commit as noted above, just not in\nthe very narrow part of the object that's reserved and not going to\nchange.\n\nMore generally plenty of *->git importers write some extra data in the\ncommits, usually in the commit message. Try e.g. cloning a SVN repo with\n\"git svn clone\" and see what it does.\n\n> The deeper problem is that I want something from Git that I cannot\n> have with 1-second granularity. That is: a unique timestamp on each\n> commit in a repository. The only way to be certain of this is for git\n> to delay accepting integration of a patch until it can issue a unique\n> time mark for it - obviously impractical if the quantum is one second,\n> but not if it's a millisecond or microsecond.\n>\n> Why do I want this? There are number of reasons, all related to a\n> mathematical concept called \"total ordering\".  At present, commits in\n> a Git repository only have partial ordering. One consequence is that\n> action stamps - the committer/date pairs I use as VCS-independent commit\n> identifications in reposurgeon - are not unique.  When a patch sequence\n> is applied, it can easily happen fast enough to give several successive\n> commits the same committer-ID and timestamp.\n>\n> Of course the commit hash remains a unique commit ID.  But it can't\n> easily be parsed and followed by a human, which is a UX problem when\n> it's used as a commit stamp in change comments.\n\nYou cannot get a guaranteed \"total order\" of any sort in anything like\ngit's current object model without taking a global lock on all write\noperations.\n\nOtherwise how would two concurrent ref updates / object writes be\nguaranteed not to get the timestamp? Unlikely with nanosecond accuracy,\nbut not impossible.\n\nEven if you solve that, take two such repositories and \"git merge\n--allow-unrelated-histories\" them together. Now what's the order?\n\nThese issues are solved by defining ordering in terms of the graph, and\nwriting this information after-the-fact. That's already part of git. See\nhttps://github.com/git/git/blob/next/Documentation/technical/commit-graph.txt\nand\nhttps://devblogs.microsoft.com/devops/supercharging-the-git-commit-graph-ii-file-format/\n\n> More deeply, the lack of total ordering means that repository graphs\n> don't have a single canonical serialized form.  This sounds abstract\n> but it means there are surgical operations I can't regression-test\n> properly.  My colleague Edward Cree has found cases where git fast-export\n> can issue a stream dump for which git fast-import won't necessarily\n> re-color certain interior nodes the same way when it's read back in\n> and I'm pretty sure the absence of total ordering on the branch tips\n> is at the bottom of that.\n\nCan you clarify what you mean by this? You run fast-import twice and get\ndifferent results, is that it? If so that sounds like a bug.\n\n> I'm willing to write patches if this direction is accepted.  I've figured\n> out how to make fast-import streams upward-compatible with finer-grained\n> timestamps.\n"},{"id":"375645","messageId":"023b01d50b5c$cbd3cd90$637b68b0$@pdinc.us","threadId":"51105","inReplyTo":"ae62476c-1642-0b9c-86a5-c2c8cddf9dfb@gmail.com","subject":"RE: Finer timestamps and serialization in git","fromName":"Jason Pyeron","fromEmail":"jpyeron@pdinc.us","sentAt":"2019-05-15T20:28:45Z","receivedAt":"2019-05-15T21:06:48Z","isPatch":false,"sender":{"key":"jpyeron@pdinc.us","avatar":"https://gravatar.com/avatar/c2e53452caa53d940768a1ffc9cf76196d851b9b534b7a39cd39852a70a0508f?d=mp&s=160"},"body":"(please don’t cc me)\n\n> -----Original Message-----\n> From: Derrick Stolee\n> Sent: Wednesday, May 15, 2019 4:16 PM\n> \n> On 5/15/2019 3:16 PM, Eric S. Raymond wrote:\n\n<snip/> I disagree with many of Eric's reasons - and agree with most of Derrick's refutation. But\n\n> \n> Changing the granularity of timestamps requires changing the commit format,\n> which is probably a non-starter. \n\nis not necessarily true. If we take the below example:\n\ncommitter Name <user@domain> 1557948240 -0400\n\nand we follow the rule that:\n\n1. any trailing zero after the decimal point MUST be omitted\n2. if there are no digits after the decimal point, it MUST be omitted\n\nThis would allow:\n\ncommitter Name <user@domain> 1557948240 -0400\ncommitter Name <user@domain> 1557948240.12 -0400\n\nbut the following are never allowed:\n\ncommitter Name <user@domain> 1557948240. -0400\ncommitter Name <user@domain> 1557948240.000000 -0400\n\nBy following these rules, all previous commits' hash are unchanged. Future commits made on the top of the second will look like old commit formats. Commits coming from \"older\" tools will produce valid and mergeable objects. The loss precision has frustrated us several times as well.\n\n\nRespectfully,\n\n\nJason Pyeron\n\n"},{"id":"375646","messageId":"998895a9-cfbb-c458-cc88-fa1aabed4389@gmail.com","threadId":"51105","inReplyTo":"023b01d50b5c$cbd3cd90$637b68b0$@pdinc.us","subject":"Re: Finer timestamps and serialization in git","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-05-15T21:14:58Z","receivedAt":"2019-05-15T21:15:03Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/15/2019 4:28 PM, Jason Pyeron wrote:\n> (please don’t cc me)\n\nOk. I'll \"To\" you.\n\n> and we follow the rule that:\n> \n> 1. any trailing zero after the decimal point MUST be omitted\n> 2. if there are no digits after the decimal point, it MUST be omitted\n> \n> This would allow:\n> \n> committer Name <user@domain> 1557948240 -0400\n> committer Name <user@domain> 1557948240.12 -0400\n\nThis kind of change would probably break old clients trying to read\ncommits from new clients. Ævar's suggestion [1] of additional headers\nshould not create incompatibilities.\n\n> By following these rules, all previous commits' hash are unchanged. Future commits made on the top of the second will look like old commit formats. Commits coming from \"older\" tools will produce valid and mergeable objects. The loss precision has frustrated us several times as well.\n\nWhat problem are you trying to solve where commit date is important?\nThe only use I have for them is \"how long has it been since someone\nmade this change?\" A question like \"when was this change introduced?\"\nis much less important than \"in which version was this first released?\"\nThis \"in which version\" is a graph reachability question, not a date\nquestion.\n\nI think any attempt to understand Git commits using commit date without\nusing the underling graph topology (commit->parent relationships) is\nfundamentally broken and won't scale to even moderately-sized teams.\nI don't even use \"git log\" without a \"--topo-order\" or \"--graph\" option\nbecause using a date order puts unrelated changes next to each other.\n--topo-order guarantees that a path of commits with only one parent\nand only one child appears in consecutive order.\n\nThanks,\n-Stolee\n\nP.S. All of my (overly strong) opinions on using commit date are made\nmore valid when you realize anyone can set GIT_COMMITTER_DATE to get\nan arbitrary commit date.\n\n[1] https://public-inbox.org/git/871s0zwjv0.fsf@evledraar.gmail.com/T/#t\n"},{"id":"375657","messageId":"87zhnnv0b8.fsf@evledraar.gmail.com","threadId":"51105","inReplyTo":"998895a9-cfbb-c458-cc88-fa1aabed4389@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-05-15T22:07:39Z","receivedAt":"2019-05-15T22:07:45Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Wed, May 15 2019, Derrick Stolee wrote:\n\n> On 5/15/2019 4:28 PM, Jason Pyeron wrote:\n>> (please don’t cc me)\n>\n> Ok. I'll \"To\" you.\n\nI'm a rebel!\n\n>> and we follow the rule that:\n>>\n>> 1. any trailing zero after the decimal point MUST be omitted\n>> 2. if there are no digits after the decimal point, it MUST be omitted\n>>\n>> This would allow:\n>>\n>> committer Name <user@domain> 1557948240 -0400\n>> committer Name <user@domain> 1557948240.12 -0400\n>\n> This kind of change would probably break old clients trying to read\n> commits from new clients. Ævar's suggestion [1] of additional headers\n> should not create incompatibilities.\n\nYes, exactly. Obviously patching git to do this is rather easy, here's\nan initial try:\n\n    diff --git a/date.c b/date.c\n    index 8126146c50..0a97e1d877 100644\n    --- a/date.c\n    +++ b/date.c\n    @@ -762,3 +762,3 @@ static void date_string(timestamp_t date, int offset, struct strbuf *buf)\n            }\n    -       strbuf_addf(buf, \"%\"PRItime\" %c%02d%02d\", date, sign, offset / 60, offset % 60);\n    +       strbuf_addf(buf, \"%\"PRItime\".12345 %c%02d%02d\", date, sign, offset / 60, offset % 60);\n     }\n    diff --git a/usage.c b/usage.c\n    index 2fdb20086b..7760b78cb6 100644\n    --- a/usage.c\n    +++ b/usage.c\n    @@ -267,2 +267,3 @@ NORETURN void BUG_fl(const char *file, int line, const char *fmt, ...)\n            va_list ap;\n    +       return;\n            va_start(ap, fmt);\n\nWe don't need BUG() right? :)\n\nNow let's commit with that git, that gives me a commit object with a\nsub-second timestamp like:\n\n    $ git cat-file -p HEAD\n    tree 4d5fcadc293a348e88f777dc0920f11e7d71441c\n    author Ævar Arnfjörð Bjarmason <avarab@gmail.com> 1557955656.12345 +0200\n    committer Ævar Arnfjörð Bjarmason <avarab@gmail.com> 1557955656.12345 +0200\n\nWorks so far, yay!\n\nAnd now fsck fails:\n\n    error in commit 31b3e9b88c36f75b3375471d9f5b449165c9ff93: badDate: invalid author/committer line - bad date\n\nAnd any sane git hosting site will refuse this, e.g. trying to push this\nto github:\n\n    remote: error: object 31b3e9b88c36f75b3375471d9f5b449165c9ff93: badDate: invalid author/committer line - bad date\n    remote: fatal: fsck error in packed object\n\nAnd that's *just* dealing with the git.git client, any such format\nchanges also need to consider what happens to jgit, libgit2 etc. etc.\n\nOnce you make such changes to the format you've created your own\nversion-control system. It's no longer git.\n\n>> By following these rules, all previous commits' hash are unchanged. Future commits made on the top of the second will look like old commit formats. Commits coming from \"older\" tools will produce valid and mergeable objects. The loss precision has frustrated us several times as well.\n>\n> What problem are you trying to solve where commit date is important?\n> The only use I have for them is \"how long has it been since someone\n> made this change?\" A question like \"when was this change introduced?\"\n> is much less important than \"in which version was this first released?\"\n> This \"in which version\" is a graph reachability question, not a date\n> question.\n>\n> I think any attempt to understand Git commits using commit date without\n> using the underling graph topology (commit->parent relationships) is\n> fundamentally broken and won't scale to even moderately-sized teams.\n> I don't even use \"git log\" without a \"--topo-order\" or \"--graph\" option\n> because using a date order puts unrelated changes next to each other.\n> --topo-order guarantees that a path of commits with only one parent\n> and only one child appears in consecutive order.\n>\n> Thanks,\n> -Stolee\n>\n> P.S. All of my (overly strong) opinions on using commit date are made\n> more valid when you realize anyone can set GIT_COMMITTER_DATE to get\n> an arbitrary commit date.\n>\n> [1] https://public-inbox.org/git/871s0zwjv0.fsf@evledraar.gmail.com/T/#t\n"},{"id":"375668","messageId":"3b8d6a78-bd88-770c-e79b-d732f7e277fd@gmail.com","threadId":"51105","inReplyTo":"20190516002831.GC124956@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-05-16T01:25:46Z","receivedAt":"2019-05-16T01:49:34Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/15/2019 8:28 PM, Eric S. Raymond wrote:\n> Derrick Stolee <stolee@gmail.com>:\n>> What problem are you trying to solve where commit date is important?\n> \n> I don't know what Jason's are.  I know what mine are.\n> \n> A. Portable commit identifiers\n> \n> 1. When I in-migrate a repository from (say) Subversion with\n> reposurgeon, I want to be able to patch change comments so that (say)\n> r2367 becomes a unique reference to its corresponding commit. I do\n> not want the kludge of appending a relic SVN-ID header to be *required*,\n> though some customers may choose that. Requirung that is an orthogonality\n> violation.\n\nInstead of using the free-form nature of a commit message to include links\nto an external VCS, you want a first-class data type in Git to provide this\ndata? Not only is that backwards, it makes the link between the Git repo and\nthe SVN repo weaker. How would you distinguish between a commit generated from\nthe old SVN repo and a commit that was created directly in the Git repo without\nperforming a lookup to the SVN repo based on (committer, timestamp)?\n\n> 2. Because I think in decadal timescales about infrastructure, I want\n> my commit references to be in a format that won't break when the history\n> is forward-migrated to the *next* VCS. That pretty much eliminates any\n> from of opaque hash. (Git itself will have a weaker version of this problem\n> when you change hash formats.)\n> \n> 3. Accordingly, I invented action stamps. This is an action stamp:\n> <esr@thyrsus.com!2019-05-15T20:01:15Z>. One reason I want timestamp\n> uniqueness is for action-stamp uniqueness.\n\nLooks like you have an excellent format for a backwards-facing link.\n\nGerrit uses a commit-msg hook [1] to insert \"Change-Id\" tags into\ncommit messages. You could probably do something similar. If you have\ncontrol over _every_ client interacting with the repo, you could even\nhave this interact with a central authority to give a unique stamp.\n\n> B. Unique canonical form of import-stream representation.\n> \n> Reposurgeon is a very complex piece of software with subtle failure\n> modes.  I have a strong need to be able to regression-test its\n> operation.  Right now there are important cases in which I can't do\n> that because (a) the order in which it writes commits and (b) how it\n> colors branches, are both phase-of-moon dependent.  That is, the\n> algorithms may be deterministic but they're not documented and seem to\n> be dependent on variables that are hidden from me.\n> \n> Before import streams can have a canonical output order without hidden\n> variables (e.g. depending only on visible metadata) in practice, that\n> needs to be possible in principle. I've thought about this a lot and\n> not only are unique commit timestamps the most natural way to make\n> it possible, they're the only way conistent with the reality that\n> commit comments may be altered for various good reasons during\n> repository translation.\n\nIf you are trying to debug or test something, why don't you serialize\nthe input you are using for your test?\n\n>> P.S. All of my (overly strong) opinions on using commit date are made\n>> more valid when you realize anyone can set GIT_COMMITTER_DATE to get\n>> an arbitrary commit date.\n> \n> In the way I would write things, you can *request* that date, but in\n> case of a collision you might actually get one a few microseconds off\n> that preserves its order relationship with your other commits.\n\nAs mentioned above, you need to make this request at the time the commit\nis created, and you'll need to communicate with a central authority. That\ngoes against the distributed nature of Git.\n\nIn my opinion, Git already gives you the flexibility to achieve the goals\nyou are looking for. But changing a core data type to make your goals\nslightly more convenient is not a valuable exercise.\n\n-Stolee\n\n[1] https://gerrit-review.googlesource.com/Documentation/cmd-hook-commit-msg.html\n \n\n"},{"id":"375669","messageId":"ab3222ab-9121-9534-1472-fac790bf08a4@gmail.com","threadId":"51105","inReplyTo":"20190515233230.GA124956@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-05-16T01:14:30Z","receivedAt":"2019-05-16T01:49:37Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/15/2019 7:32 PM, Eric S. Raymond wrote:\n> Derrick Stolee <stolee@gmail.com>:\n>> On 5/15/2019 3:16 PM, Eric S. Raymond wrote:\n>>> The deeper problem is that I want something from Git that I cannot\n>>> have with 1-second granularity. That is: a unique timestamp on each\n>>> commit in a repository.\n>>\n>> This is impossible in a distributed version control system like Git\n>> (where the commits are immutable). No matter your precision, there is\n>> a chance that two machiens commit at the exact same moment on two different\n>> machines and then those commits are merged into the same branch.\n> \n> It's easy to work around that problem. Each git daemon has to single-thread\n> its handling of incoming commits at some level, because you need a lock on the\n> file system to guarantee consistent updates to it.\n> \n> So if a commit comes in that would be the same as the date of the\n> previous commit on the current branch, you bump the incoming commit timestamp.\n\nThis changes the commit, causing it to have a different object id, and\nnow the client that pushed that commit disagrees with your machine on\nthe history.\n\n> That's the simple case. The complicated case is checking for date\n> collisions on *other* branches. But there are ways to make that fast,\n> too. There's a very obvious one involving a presort that is is O(log2\n> n) in the number of commits.\n> \n> I wouldn't have brought this up in the first place if I didn't have a\n> pretty clear idea how to do it in code!\n> \n>> Even when you specify a committer, there are many environments where a set\n>> of parallel machines are creating commits with the same identity.\n> \n> If those commit sets become the same commit in the final graph, this is\n> not a problem for total ordering.\n> \n>>> Why do I want this? There are number of reasons, all related to a\n>>> mathematical concept called \"total ordering\".  At present, commits in\n>>> a Git repository only have partial ordering. \n>>\n>> This is true of any directed acyclic graph. If you want a total ordering\n>> that is completely unambiguous, then you should think about maintaining\n>> a linear commit history by requiring rebasing instead of merging.\n> \n> Excuse me, but your premise is incorrect.  A git DAG isn't just \"any\" DAG.\n> The presence of timestamps makes a total ordering possible.\n> \n> (I was a theoretical mathematician in a former life. This is all very\n> familiar ground to me.)\n\nSame. But you seem to have a fundamental misunderstanding about the immutability\nof commits, which is core to how Git works. If you change a commit, then you\nget a new object id and now distributed copies don't agree on the history.\n\n>>> One consequence is that\n>>> action stamps - the committer/date pairs I use as VCS-independent commit\n>>> identifications in reposurgeon - are not unique.  When a patch sequence\n>>> is applied, it can easily happen fast enough to give several successive\n>>> commits the same committer-ID and timestamp.\n>>\n>> Sorting by committer/date pairs sounds like an unhelpful idea, as that\n>> does not take any graph topology into account. It happens that commits\n>> can actually have an _earlier_ commit date than its parent.\n> \n> Yes, I'm aware of that.  The uniqueness properties that make a total\n> ordering desirable are not actually dependent on timestamp order\n> coinciding with topo order.\n> \n>> Changing the granularity of timestamps requires changing the commit format,\n>> which is probably a non-starter.\n> \n> That's why I started by noting that you're going to have to break the\n> format anyway to move to an ECDSA hash (or whatever you end up using).\n> \n> I'm saying that *since you'll need to do that anyway*, it's a good time\n> to think about making timestamps finer-grained and unique.\n\nThat change is difficult enough as it is. I don't think your goals justify\nmaking this more complicated. You are also not considering:\n\n * The in-memory data type now needs to be a floating-point type, or an\n   even larger integer type using a different set of units.\n\n * This data type now affects our priority queues for commit walks, how\n   we store the commit date in the commit-graph file, how we compute\n   relative dates for 'git log' pretty formats.\n\n-Stolee\n\n"},{"id":"375676","messageId":"20190516003522.GD124956@thyrsus.com","threadId":"51105","inReplyTo":"871s0zwjv0.fsf@evledraar.gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-16T00:35:22Z","receivedAt":"2019-05-16T01:50:17Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com>:\n> You put it key-values in the commit message and read it back out via\n> git-interpret-trailers.\n\nSpeaking as a person who has done a lot of repository migrations, this\nmakes me shudder.  It's fragile, kludgy, and does not maintain proper\nseparation of concerns.\n\nThe feature I *didn't* ask for at the next format break is a user-modifiable\nkey-value store per commit that is *not* in the commit comment.  Bzr\nhas this.  It's useful.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375678","messageId":"20190516002831.GC124956@thyrsus.com","threadId":"51105","inReplyTo":"998895a9-cfbb-c458-cc88-fa1aabed4389@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-16T00:28:31Z","receivedAt":"2019-05-16T01:50:29Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Derrick Stolee <stolee@gmail.com>:\n> What problem are you trying to solve where commit date is important?\n\nI don't know what Jason's are.  I know what mine are.\n\nA. Portable commit identifiers\n\n1. When I in-migrate a repository from (say) Subversion with\nreposurgeon, I want to be able to patch change comments so that (say)\nr2367 becomes a unique reference to its corresponding commit. I do\nnot want the kludge of appending a relic SVN-ID header to be *required*,\nthough some customers may choose that. Requirung that is an orthogonality\nviolation.\n\n2. Because I think in decadal timescales about infrastructure, I want\nmy commit references to be in a format that won't break when the history\nis forward-migrated to the *next* VCS. That pretty much eliminates any\nfrom of opaque hash. (Git itself will have a weaker version of this problem\nwhen you change hash formats.)\n\n3. Accordingly, I invented action stamps. This is an action stamp:\n<esr@thyrsus.com!2019-05-15T20:01:15Z>. One reason I want timestamp\nuniqueness is for action-stamp uniqueness.\n\nB. Unique canonical form of import-stream representation.\n\nReposurgeon is a very complex piece of software with subtle failure\nmodes.  I have a strong need to be able to regression-test its\noperation.  Right now there are important cases in which I can't do\nthat because (a) the order in which it writes commits and (b) how it\ncolors branches, are both phase-of-moon dependent.  That is, the\nalgorithms may be deterministic but they're not documented and seem to\nbe dependent on variables that are hidden from me.\n\nBefore import streams can have a canonical output order without hidden\nvariables (e.g. depending only on visible metadata) in practice, that\nneeds to be possible in principle. I've thought about this a lot and\nnot only are unique commit timestamps the most natural way to make\nit possible, they're the only way conistent with the reality that\ncommit comments may be altered for various good reasons during\nrepository translation.\n\n> P.S. All of my (overly strong) opinions on using commit date are made\n> more valid when you realize anyone can set GIT_COMMITTER_DATE to get\n> an arbitrary commit date.\n\nIn the way I would write things, you can *request* that date, but in\ncase of a collision you might actually get one a few microseconds off\nthat preserves its order relationship with your other commits.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375680","messageId":"20190515233230.GA124956@thyrsus.com","threadId":"51105","inReplyTo":"ae62476c-1642-0b9c-86a5-c2c8cddf9dfb@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-15T23:32:30Z","receivedAt":"2019-05-16T01:50:36Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Derrick Stolee <stolee@gmail.com>:\n> On 5/15/2019 3:16 PM, Eric S. Raymond wrote:\n> > The deeper problem is that I want something from Git that I cannot\n> > have with 1-second granularity. That is: a unique timestamp on each\n> > commit in a repository.\n> \n> This is impossible in a distributed version control system like Git\n> (where the commits are immutable). No matter your precision, there is\n> a chance that two machiens commit at the exact same moment on two different\n> machines and then those commits are merged into the same branch.\n\nIt's easy to work around that problem. Each git daemon has to single-thread\nits handling of incoming commits at some level, because you need a lock on the\nfile system to guarantee consistent updates to it.\n\nSo if a commit comes in that would be the same as the date of the\nprevious commit on the current branch, you bump the incoming commit timestamp.\nThat's the simple case. The complicated case is checking for date\ncollisions on *other* branches. But there are ways to make that fast,\ntoo. There's a very obvious one involving a presort that is is O(log2\nn) in the number of commits.\n\nI wouldn't have brought this up in the first place if I didn't have a\npretty clear idea how to do it in code!\n\n> Even when you specify a committer, there are many environments where a set\n> of parallel machines are creating commits with the same identity.\n\nIf those commit sets become the same commit in the final graph, this is\nnot a problem for total ordering.\n\n> > Why do I want this? There are number of reasons, all related to a\n> > mathematical concept called \"total ordering\".  At present, commits in\n> > a Git repository only have partial ordering. \n> \n> This is true of any directed acyclic graph. If you want a total ordering\n> that is completely unambiguous, then you should think about maintaining\n> a linear commit history by requiring rebasing instead of merging.\n\nExcuse me, but your premise is incorrect.  A git DAG isn't just \"any\" DAG.\nThe presence of timestamps makes a total ordering possible.\n\n(I was a theoretical mathematician in a former life. This is all very\nfamiliar ground to me.)\n\n> > One consequence is that\n> > action stamps - the committer/date pairs I use as VCS-independent commit\n> > identifications in reposurgeon - are not unique.  When a patch sequence\n> > is applied, it can easily happen fast enough to give several successive\n> > commits the same committer-ID and timestamp.\n> \n> Sorting by committer/date pairs sounds like an unhelpful idea, as that\n> does not take any graph topology into account. It happens that commits\n> can actually have an _earlier_ commit date than its parent.\n\nYes, I'm aware of that.  The uniqueness properties that make a total\nordering desirable are not actually dependent on timestamp order\ncoinciding with topo order.\n\n> Changing the granularity of timestamps requires changing the commit format,\n> which is probably a non-starter.\n\nThat's why I started by noting that you're going to have to break the\nformat anyway to move to an ECDSA hash (or whatever you end up using).\n\nI'm saying that *since you'll need to do that anyway*, it's a good time\nto think about making timestamps finer-grained and unique.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375681","messageId":"20190515234034.GB124956@thyrsus.com","threadId":"51105","inReplyTo":"023b01d50b5c$cbd3cd90$637b68b0$@pdinc.us","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-15T23:40:34Z","receivedAt":"2019-05-16T01:50:37Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jason Pyeron <jpyeron@pdinc.us>:\n> If we take the below example:\n> \n> committer Name <user@domain> 1557948240 -0400\n> \n> and we follow the rule that:\n> \n> 1. any trailing zero after the decimal point MUST be omitted\n> 2. if there are no digits after the decimal point, it MUST be omitted\n> \n> This would allow:\n> \n> committer Name <user@domain> 1557948240 -0400\n> committer Name <user@domain> 1557948240.12 -0400\n> \n> but the following are never allowed:\n> \n> committer Name <user@domain> 1557948240. -0400\n> committer Name <user@domain> 1557948240.000000 -0400\n> \n> By following these rules, all previous commits' hash are unchanged. Future commits made on the top of the second will look like old commit formats. Commits coming from \"older\" tools will produce valid and mergeable objects. The loss precision has frustrated us several times as well.\n\nYes, that's almost exactly what I came up with.  I was concerned with upward\ncompatibility in fast-export streams, which reposurgeon ingests and emits.\n\nBut I don't quite understand your claim that there's no format\nbreakage here, unless you're implying to me that timestamps are already\nstored in the git file system as variable-length strings.  Do they\nreally never get translated into time_t?  Good news if so.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375696","messageId":"20190516041444.GG4596@sigill.intra.peff.net","threadId":"51105","inReplyTo":"871s0zwjv0.fsf@evledraar.gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-05-16T04:14:44Z","receivedAt":"2019-05-16T04:14:47Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, May 15, 2019 at 10:20:03PM +0200, Ævar Arnfjörð Bjarmason wrote:\n\n> > Since this is going to have to happen anyway\n> \n> The SHA-1 <-> SHA-256 transition is planned to happen, but there's some\n> strong opinions that this should be *only* for munging the content for\n> hashing, not adding new stuff while we're at it (even if optional). See\n> : https://public-inbox.org/git/87ftyyedqd.fsf@evledraar.gmail.com/\n\nOne reason for this is that the transition plan calls for being able to\nconvert between the sha1 and sha256 representations losslessly (which\nmakes interoperability possible and avoids a flag day). So even if the\nsha256 format understood floating-point timestamps in the committer\nheader, we'd have to have some way of representing that same information\nin the sha1 format. Which implies putting it into a new header, as you\ndescribed below.\n\nAnd if it's in a new header in sha1, then is there any real advantage in\nhaving it somewhere else in the sha256 version? I dunno. Maybe a little,\nas eventually all of the sha1 formats would die off, after everybody has\ntransitioned.\n\n-Peff\n"},{"id":"375712","messageId":"87woiqvic4.fsf@evledraar.gmail.com","threadId":"51105","inReplyTo":"20190515233230.GA124956@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-05-16T09:50:35Z","receivedAt":"2019-05-16T09:50:41Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Thu, May 16 2019, Eric S. Raymond wrote:\n\n> Derrick Stolee <stolee@gmail.com>:\n>> On 5/15/2019 3:16 PM, Eric S. Raymond wrote:\n>> > The deeper problem is that I want something from Git that I cannot\n>> > have with 1-second granularity. That is: a unique timestamp on each\n>> > commit in a repository.\n>>\n>> This is impossible in a distributed version control system like Git\n>> (where the commits are immutable). No matter your precision, there is\n>> a chance that two machiens commit at the exact same moment on two different\n>> machines and then those commits are merged into the same branch.\n>\n> It's easy to work around that problem. Each git daemon has to single-thread\n> its handling of incoming commits at some level, because you need a lock on the\n> file system to guarantee consistent updates to it.\n\nYou don't need a daemon now to write commits to a repository. You can\njust add stuff to the object store, and then later flip the SHA-1 on a\nreference, we lock those indivdiual references, but this sort of thing\nwould require a global write lock. This would introduce huge concurrency\ncaveats that are non-issues now.\n\nDumb clients matter. Now you can e.g. have two libgit2 processes writing\nto ref A and B respectively in the same repo, and they never have to\nknow about each other or care about IPC.\n\nAlso, even if you have daemons accepting pushes they can now be on\ndifferent computers sharing things over e.g. an NFS filesystem. Now you\nneed some FS-based serialization protcol for commits and their\ntimestamps.\n\n> So if a commit comes in that would be the same as the date of the\n> previous commit on the current branch, you bump the incoming commit timestamp.\n> That's the simple case. The complicated case is checking for date\n> collisions on *other* branches. But there are ways to make that fast,\n> too. There's a very obvious one involving a presort that is is O(log2\n> n) in the number of commits.\n\nWhat Derrick mentioned downthread of this \"I rebase your pushes\" being\nfundimentally un-git applies, but let's assume we can somehow get past\nthat for the sake of argument.\n\nThe model you're trying to impose here of \"within a repo I want to\nserialize all X\" just doesn't play with how git views the world. Git\ncares about graphs being serialized, it doesn't care about arbitrary\nsets of graphs.\n\nE.g. let's say I push a commit X to github, and now I want to push the\nsame history to gitlab, I might be twarted because they have some\nside-ref they themselves make (e.g. the PR or MR refs) which conflicts\nwith this \"timestamps must monotonically increase across all branches in\na repo\" view of the world.\n\nThe only thing that matters in git in this regard is how individual refs\nbehave, we then by convention tend to have a 1=1 mapping between those\nsets of refs and a repository, but in a lot of cases it's\nmany=1. E.g. in cases where such a hosting site might have one\nunderlying repo store exposed to multiple users via ref namespace\nprefixes.\n\n> I wouldn't have brought this up in the first place if I didn't have a\n> pretty clear idea how to do it in code!\n>\n>> Even when you specify a committer, there are many environments where a set\n>> of parallel machines are creating commits with the same identity.\n>\n> If those commit sets become the same commit in the final graph, this is\n> not a problem for total ordering.\n>\n>> > Why do I want this? There are number of reasons, all related to a\n>> > mathematical concept called \"total ordering\".  At present, commits in\n>> > a Git repository only have partial ordering.\n>>\n>> This is true of any directed acyclic graph. If you want a total ordering\n>> that is completely unambiguous, then you should think about maintaining\n>> a linear commit history by requiring rebasing instead of merging.\n>\n> Excuse me, but your premise is incorrect.  A git DAG isn't just \"any\" DAG.\n> The presence of timestamps makes a total ordering possible.\n>\n> (I was a theoretical mathematician in a former life. This is all very\n> familiar ground to me.)\n>\n>> > One consequence is that\n>> > action stamps - the committer/date pairs I use as VCS-independent commit\n>> > identifications in reposurgeon - are not unique.  When a patch sequence\n>> > is applied, it can easily happen fast enough to give several successive\n>> > commits the same committer-ID and timestamp.\n>>\n>> Sorting by committer/date pairs sounds like an unhelpful idea, as that\n>> does not take any graph topology into account. It happens that commits\n>> can actually have an _earlier_ commit date than its parent.\n>\n> Yes, I'm aware of that.  The uniqueness properties that make a total\n> ordering desirable are not actually dependent on timestamp order\n> coinciding with topo order.\n>\n>> Changing the granularity of timestamps requires changing the commit format,\n>> which is probably a non-starter.\n>\n> That's why I started by noting that you're going to have to break the\n> format anyway to move to an ECDSA hash (or whatever you end up using).\n>\n> I'm saying that *since you'll need to do that anyway*, it's a good time\n> to think about making timestamps finer-grained and unique.\n\nWe should really discuss proposed format changes separately from tacking\nthem onto the SHA-256 transition, because as I noted upthread your\npremise that you need a format change for this isn't true. *If* this was\na good idea it's something you can add to commit objects.\n\nAnd yeah, git-interpret-trailers is a bit of a kludge, which is why I\nmentioned you can add new headers to the format, this is e.g. how GPG\nsigned commits work.\n\nOf course whether it makes any sense to add such a thing to the format\nis another matter, I'm not at all convinced, but that's a separate\ndiscussion from how it would be done.\n"},{"id":"375863","messageId":"b4b151ba-ab43-445f-6e49-ee8e28b30859@iee.org","threadId":"51105","inReplyTo":"20190515234034.GB124956@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":"2019-05-19T00:16:27Z","receivedAt":"2019-05-19T00:16:32Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"\n\nOn 16/05/2019 00:40, Eric S. Raymond wrote:\n> Jason Pyeron <jpyeron@pdinc.us>:\n>> If we take the below example:\n>>\n>> committer Name <user@domain> 1557948240 -0400\n>>\n>> and we follow the rule that:\n>>\n>> 1. any trailing zero after the decimal point MUST be omitted\n>> 2. if there are no digits after the decimal point, it MUST be omitted\n>>\n>> This would allow:\n>>\n>> committer Name <user@domain> 1557948240 -0400\n>> committer Name <user@domain> 1557948240.12 -0400\n>>\n>> but the following are never allowed:\n>>\n>> committer Name <user@domain> 1557948240. -0400\n>> committer Name <user@domain> 1557948240.000000 -0400\n>>\n>> By following these rules, all previous commits' hash are unchanged. Future commits made on the top of the second will look like old commit formats. Commits coming from \"older\" tools will produce valid and mergeable objects. The loss precision has frustrated us several times as well.\n> Yes, that's almost exactly what I came up with.  I was concerned with upward\n> compatibility in fast-export streams, which reposurgeon ingests and emits.\n>\n> But I don't quite understand your claim that there's no format\n> breakage here, unless you're implying to me that timestamps are already\n> stored in the git file system as variable-length strings.  Do they\n> really never get translated into time_t?  Good news if so.\nMaybe just take some of the object ID bits as being the fractional time \ntimestamp. They are effectively random, so should do a reasonable job of \ndistinguishing commits in a repeatable manner, even with full round \ntripping via older git versions (as long as the sha1 replicates...)\n\nAs I understand it the commit timestamp is actually free text within the \ncommit object (try `git cat-file -p <commit_object>), so the issue is \nwhether the particular git version is ready to accept the additional \n'dot' factional time notation (future versions could be extended, but I \nthink old ones would reject them if I understand the test up thread - \nwhich would compromise backward compatibility and round tripping).\n--\nPhilip\n"},{"id":"375872","messageId":"318a3460-3955-4330-b0bf-5e96e5353178@iee.org","threadId":"51105","inReplyTo":"20190519040902.GA32780@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":"2019-05-19T10:07:28Z","receivedAt":"2019-05-19T16:58:35Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi Eric,\n\nOn 19/05/2019 05:09, Eric S. Raymond wrote:\n> Philip Oakley <philipoakley@iee.org>:\n>>> But I don't quite understand your claim that there's no format\n>>> breakage here, unless you're implying to me that timestamps are already\n>>> stored in the git file system as variable-length strings.  Do they\n>>> really never get translated into time_t?  Good news if so.\n>> Maybe just take some of the object ID bits as being the fractional time\n>> timestamp. They are effectively random, so should do a reasonable job of\n>> distinguishing commits in a repeatable manner, even with full round tripping\n>> via older git versions (as long as the sha1 replicates...)\n> Huh.  That's an interesting idea.  Doesn't absolutely guarantee uniqueness,\n> but even with birthday effect the probability of collisions could be pulled\n> arbitrarily low.\ndepends how many bits are in the 'nano-second' resolution long word ;-)\nsee also\n>\n>> As I understand it the commit timestamp is actually free text within the\n>> commit object (try `git cat-file -p <commit_object>), so the issue is\n>> whether the particular git version is ready to accept the additional 'dot'\n>> factional time notation (future versions could be extended, but I think old\n>> ones would reject them if I understand the test up thread - which would\n>> compromise backward compatibility and round tripping).\n> Nobody seems to want to grapple with the fact that changing hash formats is\n> as large or larger a problem in exactly the same way.\n>\n> I'm not saying that changing the timestamp granularity justifies a format\n> break.  I'm saying that *since you're going to have one anyway*, the option\n> to increase timestamp precision at the same time should not be missed.\nIt is probably the round tripping issue with a non-fixed format (for the \ntime string) that will scupper the idea, plus the focus being primarily \non the DAG as the fundamental lineage (which only gives partial order, \nwhich can be an issue for other VCS systems that are based on \nincremental changes rather than snapshots)\nThe transition is well underway see thread: \nhttps://public-inbox.org/git/20190212012256.1005924-1-sandals@crustytoothpaste.net/ \nfor a patch series.\n\nThe plan is at: \nhttps://github.com/git/git/blob/master/Documentation/technical/hash-function-transition.txt \n<https://github.com/git/git/blob/v2.19.0-rc0/Documentation/technical/hash-function-transition.txt>, \n\nsome discussions at thread: \nhttps://public-inbox.org/git/878t4xfaes.fsf@evledraar.gmail.com/ etc.\n\nThe timestamp problem is known see yesterdays thread: \nhttps://public-inbox.org/git/20190518005412.n45pj5p2rrtm2bfj@glandium.org/\n\nGiven that the object ID should be immutable for a round trip, using \n64bits from the sha1-oid as notional 'nano-second' time does give a \nreasonable birthday attack resistance of ~32 bits (i.e. >1M commits with \nidentical whole second timestamps). [or choose the sha-256 once the \ntransition is well underway]\n--\nPhilip\n\n\n"},{"id":"375874","messageId":"20190519040902.GA32780@thyrsus.com","threadId":"51105","inReplyTo":"b4b151ba-ab43-445f-6e49-ee8e28b30859@iee.org","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-19T04:09:02Z","receivedAt":"2019-05-19T17:04:42Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Philip Oakley <philipoakley@iee.org>:\n> > But I don't quite understand your claim that there's no format\n> > breakage here, unless you're implying to me that timestamps are already\n> > stored in the git file system as variable-length strings.  Do they\n> > really never get translated into time_t?  Good news if so.\n> Maybe just take some of the object ID bits as being the fractional time\n> timestamp. They are effectively random, so should do a reasonable job of\n> distinguishing commits in a repeatable manner, even with full round tripping\n> via older git versions (as long as the sha1 replicates...)\n\nHuh.  That's an interesting idea.  Doesn't absolutely guarantee uniqueness,\nbut even with birthday effect the probability of collisions could be pulled\narbitrarily low.\n\n> As I understand it the commit timestamp is actually free text within the\n> commit object (try `git cat-file -p <commit_object>), so the issue is\n> whether the particular git version is ready to accept the additional 'dot'\n> factional time notation (future versions could be extended, but I think old\n> ones would reject them if I understand the test up thread - which would\n> compromise backward compatibility and round tripping).\n\nNobody seems to want to grapple with the fact that changing hash formats is\nas large or larger a problem in exactly the same way.\n\nI'm not saying that changing the timestamp granularity justifies a format\nbreak.  I'm saying that *since you're going to have one anyway*, the option\nto increase timestamp precision at the same time should not be missed.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375894","messageId":"86woimox24.fsf@gmail.com","threadId":"51105","inReplyTo":"87woiqvic4.fsf@evledraar.gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2019-05-19T23:15:47Z","receivedAt":"2019-05-19T23:16:13Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Ævar Arnfjörð Bjarmason <avarab@gmail.com> writes:\n> On Thu, May 16 2019, Eric S. Raymond wrote:\n>> Derrick Stolee <stolee@gmail.com>:\n>>> On 5/15/2019 3:16 PM, Eric S. Raymond wrote:\n>>>> The deeper problem is that I want something from Git that I cannot\n>>>> have with 1-second granularity. That is: a unique timestamp on each\n>>>> commit in a repository.\n>>>\n>>> This is impossible in a distributed version control system like Git\n>>> (where the commits are immutable). No matter your precision, there is\n>>> a chance that two machines commit at the exact same moment on two different\n>>> machines and then those commits are merged into the same branch.\n>>\n>> It's easy to work around that problem. Each git daemon has to single-thread\n>> its handling of incoming commits at some level, because you need a lock on the\n>> file system to guarantee consistent updates to it.\n\nAs far as I understand it this would slow down receiving new commits\ntremendously.  Currently great care is taken to not have to parse the\ncommit object during fetch or push if it is not necessary (thanks to\nthings such as reachability bitmaps, see e.g. [1]).\n\nWith this restriction you would need to parse each commit to get at\ncommit timestamp and committer, check if the committer+timestamp is\nunique, and bump it if it is not.\n\nAlso, bumping timestamp means that the commit changed, means that its\ncontents-based ID changed, means that all commits that follow it needs\nto have its contents changed...  And now you need to rewrite many\ncommits.  And you also break the assumptions that the same commits have\nthe same contents (including date) and the same ID in different\nrepositories (some of which may include additional branches, some of\nwhich may have been part of network of related repositories, etc.).\n\n[1]: https://github.blog/2015-09-22-counting-objects/\n     http://githubengineering.com/counting-objects/\n\n> You don't need a daemon now to write commits to a repository. You can\n> just add stuff to the object store, and then later flip the SHA-1 on a\n> reference, we lock those indivdiual references, but this sort of thing\n> would require a global write lock. This would introduce huge concurrency\n> caveats that are non-issues now.\n>\n> Dumb clients matter. Now you can e.g. have two libgit2 processes writing\n> to ref A and B respectively in the same repo, and they never have to\n> know about each other or care about IPC.\n>\n> Also, even if you have daemons accepting pushes they can now be on\n> different computers sharing things over e.g. an NFS filesystem. Now you\n> need some FS-based serialization protcol for commits and their\n> timestamps.\n\nAlso, performance matters.  Especially for large repositories, and for\nlarge number of repositories.\n\n>> So if a commit comes in that would be the same as the date of the\n>> previous commit on the current branch, you bump the incoming commit timestamp.\n\nYou do realize that dates may not be monotonic (because of imperfections\nin clock synchronization), thus the fact that the date is different from\nparent does not mean that is different from ancestor.\n\n>> That's the simple case. The complicated case is checking for date\n>> collisions on *other* branches. But there are ways to make that fast,\n>> too. There's a very obvious one involving a presort that is is O(log2\n>> n) in the number of commits.\n\nI don't think performance hit you would get would be acceptable.\n\n[...]\n>>>> Why do I want this? There are number of reasons, all related to a\n>>>> mathematical concept called \"total ordering\".  At present, commits in\n>>>> a Git repository only have partial ordering.\n>>>\n>>> This is true of any directed acyclic graph. If you want a total ordering\n>>> that is completely unambiguous, then you should think about maintaining\n>>> a linear commit history by requiring rebasing instead of merging.\n>>\n>> Excuse me, but your premise is incorrect.  A git DAG isn't just \"any\" DAG.\n>> The presence of timestamps makes a total ordering possible.\n>>\n>> (I was a theoretical mathematician in a former life. This is all very\n>> familiar ground to me.)\n\nMaybe in theory, when all clock are synchronized.  But not in practice.\nShit happens.  Just recently Mike Hommey wrote about the case he has to\ndeal with:\n\nMH> I'm hitting another corner case in some other \"weird\" history, where\nMH> I have 500k commits all with the same date.\n\n[2]: https://public-inbox.org/git/20190518005412.n45pj5p2rrtm2bfj@glandium.org/t/#u\n\n--\nJakub Narębski\n"},{"id":"375896","messageId":"20190520004559.GA41412@thyrsus.com","threadId":"51105","inReplyTo":"86woimox24.fsf@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-20T00:45:59Z","receivedAt":"2019-05-20T00:46:01Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jakub Narebski <jnareb@gmail.com>:\n> As far as I understand it this would slow down receiving new commits\n> tremendously.  Currently great care is taken to not have to parse the\n> commit object during fetch or push if it is not necessary (thanks to\n> things such as reachability bitmaps, see e.g. [1]).\n> \n> With this restriction you would need to parse each commit to get at\n> commit timestamp and committer, check if the committer+timestamp is\n> unique, and bump it if it is not.\n\nSo, I'd want to measure that rather than simply assuming it's a blocker.\nClocks are very cheap these days.\n\n> Also, bumping timestamp means that the commit changed, means that its\n> contents-based ID changed, means that all commits that follow it needs\n> to have its contents changed...  And now you need to rewrite many\n> commits.\n\nWhat \"commits that follow it?\" By hypothesis, the incoming commit's\ntimestamp is bumped (if it's bumped) when it's first added to a branch\nor branches, before there are following commits in the DAG.\n\n>    And you also break the assumptions that the same commits have\n> the same contents (including date) and the same ID in different\n> repositories (some of which may include additional branches, some of\n> which may have been part of network of related repositories, etc.).\n\nWait...unless I completely misunderstand the hash-chain model, doesn't the\nhash of a commit depend on the hashes of its parents?  If that's the case,\ncommits cannot have portable hashes. If it's not, please correct me.\n\nBut if it's not, how does your first objection make sense?\n\n> > You don't need a daemon now to write commits to a repository. You can\n> > just add stuff to the object store, and then later flip the SHA-1 on a\n> > reference, we lock those indivdiual references, but this sort of thing\n> > would require a global write lock. This would introduce huge concurrency\n> > caveats that are non-issues now.\n> >\n> > Dumb clients matter. Now you can e.g. have two libgit2 processes writing\n> > to ref A and B respectively in the same repo, and they never have to\n> > know about each other or care about IPC.\n\nHow do they know they're not writing to the same ref?  What keeps\n*that* operation atomic?\n\n> You do realize that dates may not be monotonic (because of imperfections\n> in clock synchronization), thus the fact that the date is different from\n> parent does not mean that is different from ancestor.\n\nGood point. That means the O(log2 n) version of the check has to be done\nall the time.  Unfortunate.\n\n> >> That's the simple case. The complicated case is checking for date\n> >> collisions on *other* branches. But there are ways to make that fast,\n> >> too. There's a very obvious one involving a presort that is is O(log2\n> >> n) in the number of commits.\n> \n> I don't think performance hit you would get would be acceptable.\n\nAgain, it's bad practice to assume rather than measure. Human intuitions\nabout this sort of thing are notoriously unreliable.\n\n> >> Excuse me, but your premise is incorrect.  A git DAG isn't just \"any\" DAG.\n> >> The presence of timestamps makes a total ordering possible.\n> >>\n> >> (I was a theoretical mathematician in a former life. This is all very\n> >> familiar ground to me.)\n> \n> Maybe in theory, when all clock are synchronized.\n\nMy assertion does not depend on synchronized clocks, because it doesn't have to.\n\nIf the timestamps in your repo are unique, there *is* a total ordering - \nby timestamp. What you don't get is guaranteed consistency with the\ntopo ordering - that is you get no guarantee that a child's timestamp\nis greater than its parents'. That really would require a common\ntimebase.\n\nBut I don't need that stronger property, because the purpose of\ntotally ordering the repo is to guararantee the uniqueness of action\nstamps.  For that, all I need is to be able to generate a unique cookie\nfor each commit that can be inserted in its action stamp.  For my use cases\nthat cookie should *not* be a hash, because hashes always break N years\ndown.  It should be an eternally stable product of the commit metadata.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375901","messageId":"86r28tpikt.fsf@gmail.com","threadId":"51105","inReplyTo":"20190520004559.GA41412@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2019-05-20T09:43:14Z","receivedAt":"2019-05-20T09:43:21Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Eric S. Raymond\" <esr@thyrsus.com> writes:\n> Jakub Narebski <jnareb@gmail.com>:\n\n>> As far as I understand it this would slow down receiving new commits\n>> tremendously.  Currently great care is taken to not have to parse the\n>> commit object during fetch or push if it is not necessary (thanks to\n>> things such as reachability bitmaps, see e.g. [1]).\n>> \n>> With this restriction you would need to parse each commit to get at\n>> commit timestamp and committer, check if the committer+timestamp is\n>> unique, and bump it if it is not.\n>\n> So, I'd want to measure that rather than simply assuming it's a blocker.\n> Clocks are very cheap these days.\n\nClocks may be cheap, but parsing is not.\n\nYou can receive new commits in the repository by creating them, and from\nother repository (via push or fetch).  In the second case you often get\nmany commits at once.\n\nIn [1] it is described how using \"bitmap index\" you can avoid parsing\ncommits when deciding which objects to send to the client; they can be\ndirectly copied to the client (added to the packfile that is sent to\nclient).  Thanks to this reachability bitmap (bit vector) the time to\nclone Linux repository decreased from 57 seconds to 1.6 seconds.\n\nIt is not a direct correspondence, but there most probably would be the\nsame problem with requiring fractional timestamp+committer identity to\nbe unique on the receiving side.\n\n[1]: https://githubengineering.com/counting-objects/\n\n>> Also, bumping timestamp means that the commit changed, means that its\n>> contents-based ID changed, means that all commits that follow it needs\n>> to have its contents changed...  And now you need to rewrite many\n>> commits.\n>\n> What \"commits that follow it?\" By hypothesis, the incoming commit's\n> timestamp is bumped (if it's bumped) when it's first added to a branch\n> or branches, before there are following commits in the DAG.\n\nErrr... the main problem is with distributed nature of Git, i.e. when\ntwo repositories create different commits with the same\ncommitter+timestamp value.  You receive commits on fetch or push, and\nyou receive many commits at once.\n\nSay you have two repositories, and the history looks like this:\n\n repo A:   1<---2<---a<---x<---c<---d      <- master\n\n repo B:   1<---2<---X<---3<---4           <- master\n\nWhen you push from repo A to repo B, or fetch in repo B from repo A you\nwould get the following DAG of revisions\n\n repo B:   1<---2<---X<---3<---4           <- master\n                 \\\n                  \\--a<---x<---c<---d      <- repo_A/master\n\nNow let's assume that commits X and x have the came committer and the\nsame fractional timestamp, while being different commits.  Then you\nwould need to bump timestamp of 'x', changing the commit.  This means\nthat 'c' needs to be rewritten too, and 'd' also:\n\n repo B:   1<---2<---X<---3<---4           <- master\n                 \\\n                  \\--a<---x'<--c'<--d'     <- repo_A/master\n\nAnd now for the final nail in the coffing of the Bazaar-esque idea of\nchanging commits on arrival.  Say that repository A created new commits,\nand pushed them to B.  You would need to rewrite all future commits from\nthis repository too, and you would always fetch all commits starting\nfrom the first \"bumped\"\n\n repo A:   1<---2<---a<---x<---c<---d<---E   <- master\n\ntransfer of [<---x<---c<---d<---E], instead of [<--E], because 'x', 'c',\nand 'd' are missing in repo B.\n\n repo B:   1<---2<---X<---3<---4             <- master\n                 \\\n                  \\--a<---x'<--c'<--d'<--E'  <- repo_A/master\n\nAnd there is yet another problem.  Let's assume that repo B created some\nhistory on top of bump-rewritten commits:\n\n repo B:   1<---2<---X<---3<---4             <- master\n                 \\\n                  \\--a<---x'<--c'<--d'<--E'  <- repo_A/master\n                                \\\n                                 \\--5        <- next\n\nThen if in repo A you fetch from repo B (remember, in Git there is no\nconcept of central repository), you would get the following history\n\n                  /--X'<--3'<--4'            <- repo_B/master\n                 /\n repo A:   1<---2<---a<---x<---c<---d<---E   <- master\n                     \\\n                      \\---x'<--c'\n                                \\\n                                 \\--5        <- repo_B/master\n\n(because 'X' is now incoming, it needs to be \"bumped\", therefore\nchanging 3' and 4').\n\nThe history without all this rewriting looks like this:\n\n                  /--X<---3<---4'            <- repo_B/master\n                 /           \n repo A:   1<---2<---a<---x<---c<---d<---E   <- master\n                                \\\n                                 \\--5        <- repo_B/master\n\nNotice the difference?\n\n>>    And you also break the assumptions that the same commits have\n>> the same contents (including date) and the same ID in different\n>> repositories (some of which may include additional branches, some of\n>> which may have been part of network of related repositories, etc.).\n\nSee repo A and repo B in above example.\n\n> Wait...unless I completely misunderstand the hash-chain model, doesn't the\n> hash of a commit depend on the hashes of its parents?  If that's the case,\n> commits cannot have portable hashes. If it's not, please correct me.\n>\n> But if it's not, how does your first objection make sense?\n\nHash of a commit depend in hashes of its parents (Merkle tree). That is\nwhy signing a commit (or a tag pointing to the commit) signs a whole\nhistory of a commit.\n\n>>> You don't need a daemon now to write commits to a repository. You can\n>>> just add stuff to the object store, and then later flip the SHA-1 on a\n>>> reference, we lock those indivdiual references, but this sort of thing\n>>> would require a global write lock. This would introduce huge concurrency\n>>> caveats that are non-issues now.\n>>>\n>>> Dumb clients matter. Now you can e.g. have two libgit2 processes writing\n>>> to ref A and B respectively in the same repo, and they never have to\n>>> know about each other or care about IPC.\n>\n> How do they know they're not writing to the same ref?  What keeps\n> *that* operation atomic?\n\nBecause different refs are stored in different files (at least for\n\"live\" refs that are stores in loose ref format).  The lock is taken on\nref (to update ref and its reflog in sync), there is no need to take\nglobal lock on all refs.\n\n>> You do realize that dates may not be monotonic (because of imperfections\n>> in clock synchronization), thus the fact that the date is different from\n>> parent does not mean that is different from ancestor.\n>\n> Good point. That means the O(log2 n) version of the check has to be done\n> all the time.  Unfortunate.\n\nEspecially with around 1 million of commits (Linux kernel, Chromium,\nAOSP), or even 3M commits (MS Windows repository).\n\n>>>> That's the simple case. The complicated case is checking for date\n>>>> collisions on *other* branches. But there are ways to make that fast,\n>>>> too. There's a very obvious one involving a presort that is is O(log2\n>>>> n) in the number of commits.\n>> \n>> I don't think performance hit you would get would be acceptable.\n>\n> Again, it's bad practice to assume rather than measure. Human intuitions\n> about this sort of thing are notoriously unreliable.\n\nTechniques created to handle very large repositories (with respect to\nnumber of commits) that make it possible for Git to avoid parsing commit\nobjects, namely bitmap index (for 'git fetch'/'clone') and serialized\ncommit graph (for 'git log') lead to _significant_ performance\nimprovements.\n\nThe performance changes from \"waiting for Git to finish\" to \"done in the\nblink of eye\" (well, almost).\n\n>>>> Excuse me, but your premise is incorrect.  A git DAG isn't just \"any\" DAG.\n>>>> The presence of timestamps makes a total ordering possible.\n>>>>\n>>>> (I was a theoretical mathematician in a former life. This is all very\n>>>> familiar ground to me.)\n>> \n>> Maybe in theory, when all clock are synchronized.\n>\n> My assertion does not depend on synchronized clocks, because it doesn't have to.\n>\n> If the timestamps in your repo are unique, there *is* a total ordering - \n> by timestamp. What you don't get is guaranteed consistency with the\n> topo ordering - that is you get no guarantee that a child's timestamp\n> is greater than its parents'. That really would require a common\n> timebase.\n>\n> But I don't need that stronger property, because the purpose of\n> totally ordering the repo is to guarantee the uniqueness of action\n> stamps.  For that, all I need is to be able to generate a unique cookie\n> for each commit that can be inserted in its action stamp.\n\nFor cookie to be unique among all forks / clones of the same repository\nyou need either centralized naming server, or for the cookie to be based\non contents of the commit (i.e. be a hash function).\n\n>                                                          For my use cases\n> that cookie should *not* be a hash, because hashes always break N years\n> down.  It should be an eternally stable product of the commit metadata.\n\nWell, the idea for SHA-1 <--> NewHash == SHA-256 transition is to avoid\nhaving a flag day, and providing full interoperability between\nrepositories and Git installations using the old hash ad using new\nhash^1.  This will be done internally by using SHA-1 <--> SHA-256\nmapping.  So after the transition all you need is to publish this\nmapping somewhere, be it with Internet Archive or Software Heritage.\nProblem solved.\n\nP.S. Could you explain to me how one can use action stamp, e.g.\n<esr@thyrsus.com!2019-05-15T20:01:15.473209800Z>, to quickly find the\ncommit it refers to?  With SHA-1 id you have either filesystem pathname\nor the index file for pack to find it _fast_.\n\nFootnotes:\n----------\n1. That is why where would be no \"major format break\", thus no place for\n   incompatibile format changes.\n\nBest,\n--\nJakub Narębski\n"},{"id":"375905","messageId":"87k1elv3on.fsf@evledraar.gmail.com","threadId":"51105","inReplyTo":"86r28tpikt.fsf@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2019-05-20T10:08:24Z","receivedAt":"2019-05-20T10:08:29Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, May 20 2019, Jakub Narebski wrote:\n\n> \"Eric S. Raymond\" <esr@thyrsus.com> writes:\n>> Jakub Narebski <jnareb@gmail.com>:\n>\n>>> As far as I understand it this would slow down receiving new commits\n>>> tremendously.  Currently great care is taken to not have to parse the\n>>> commit object during fetch or push if it is not necessary (thanks to\n>>> things such as reachability bitmaps, see e.g. [1]).\n>>>\n>>> With this restriction you would need to parse each commit to get at\n>>> commit timestamp and committer, check if the committer+timestamp is\n>>> unique, and bump it if it is not.\n>>\n>> So, I'd want to measure that rather than simply assuming it's a blocker.\n>> Clocks are very cheap these days.\n>\n> Clocks may be cheap, but parsing is not.\n>\n> You can receive new commits in the repository by creating them, and from\n> other repository (via push or fetch).  In the second case you often get\n> many commits at once.\n>\n> In [1] it is described how using \"bitmap index\" you can avoid parsing\n> commits when deciding which objects to send to the client; they can be\n> directly copied to the client (added to the packfile that is sent to\n> client).  Thanks to this reachability bitmap (bit vector) the time to\n> clone Linux repository decreased from 57 seconds to 1.6 seconds.\n>\n> It is not a direct correspondence, but there most probably would be the\n> same problem with requiring fractional timestamp+committer identity to\n> be unique on the receiving side.\n>\n> [1]: https://githubengineering.com/counting-objects/\n\nWe're in violent agreement about the general viability of ESR's proposed\nplan, but just a side-note on this point. I don't think this is\nright. I.e. I don't think a hypothetical version of git that guarantees\nmonotonically increasing timestamps will be slow in *this* regard.\n\nFor accepting pushes we already unpack all the commits / content / hash\nit to perform fsck checks, which is why screwing with the commit\ntimestamp will fail on push:\nhttps://public-inbox.org/git/87zhnnv0b8.fsf@evledraar.gmail.com/\n\nSame on the client with fetches, although transfer.fsckObjects isn't on\nthere we do most of the work anyway for hashing & basic validation\npurposes.\n\nThe bitmaps wouldn't be affected because they're computed after-the-fact\non the basis of reachability, whereas validating increasing timestamps\nfor a single branch is cheap, you just look at each A..B push\nincrementally and see if the timestamps are increasing and past A's\nparent.\n\nIt's trickier if you're trying to make the same guarantee for *all* ref\nupdates in a given repo (and locking caveats etc. have been discussed\nelsewhere), but not *that* much of a PITA.\n\nWe'd need to compare \"new\" packs/loose objects against the new push, and\nan obvious shortcut in such a schema if you required a global lock\nanyway would be for the process taking the lock to write out \"this is\nthe current max timestamp\" when finished.\n\nIn *this* case that's a long way down the journey into crazytown :)\n\nBut it is intertesting to think about in general, because with e.g. the\ncommit-graph we have a set of commits that are \"optimized\" in some\nside-index, so it becomes useful for many algorithms to be able to ask\n\"what is the current set of unoptimized commits\".\n\nOnce you have that, and can keep the size of it down with \"gc\" many\nalgorithms that require graph traversal become possible, because your\nO(n) of needing to consider the \"n\" unoptimized commits is small enough\nv.s. the bulk of \"optimized\" commits as to not matter.\n\n\n>>> Also, bumping timestamp means that the commit changed, means that its\n>>> contents-based ID changed, means that all commits that follow it needs\n>>> to have its contents changed...  And now you need to rewrite many\n>>> commits.\n>>\n>> What \"commits that follow it?\" By hypothesis, the incoming commit's\n>> timestamp is bumped (if it's bumped) when it's first added to a branch\n>> or branches, before there are following commits in the DAG.\n>\n> Errr... the main problem is with distributed nature of Git, i.e. when\n> two repositories create different commits with the same\n> committer+timestamp value.  You receive commits on fetch or push, and\n> you receive many commits at once.\n>\n> Say you have two repositories, and the history looks like this:\n>\n>  repo A:   1<---2<---a<---x<---c<---d      <- master\n>\n>  repo B:   1<---2<---X<---3<---4           <- master\n>\n> When you push from repo A to repo B, or fetch in repo B from repo A you\n> would get the following DAG of revisions\n>\n>  repo B:   1<---2<---X<---3<---4           <- master\n>                  \\\n>                   \\--a<---x<---c<---d      <- repo_A/master\n>\n> Now let's assume that commits X and x have the came committer and the\n> same fractional timestamp, while being different commits.  Then you\n> would need to bump timestamp of 'x', changing the commit.  This means\n> that 'c' needs to be rewritten too, and 'd' also:\n>\n>  repo B:   1<---2<---X<---3<---4           <- master\n>                  \\\n>                   \\--a<---x'<--c'<--d'     <- repo_A/master\n>\n> And now for the final nail in the coffing of the Bazaar-esque idea of\n> changing commits on arrival.  Say that repository A created new commits,\n> and pushed them to B.  You would need to rewrite all future commits from\n> this repository too, and you would always fetch all commits starting\n> from the first \"bumped\"\n>\n>  repo A:   1<---2<---a<---x<---c<---d<---E   <- master\n>\n> transfer of [<---x<---c<---d<---E], instead of [<--E], because 'x', 'c',\n> and 'd' are missing in repo B.\n>\n>  repo B:   1<---2<---X<---3<---4             <- master\n>                  \\\n>                   \\--a<---x'<--c'<--d'<--E'  <- repo_A/master\n>\n> And there is yet another problem.  Let's assume that repo B created some\n> history on top of bump-rewritten commits:\n>\n>  repo B:   1<---2<---X<---3<---4             <- master\n>                  \\\n>                   \\--a<---x'<--c'<--d'<--E'  <- repo_A/master\n>                                 \\\n>                                  \\--5        <- next\n>\n> Then if in repo A you fetch from repo B (remember, in Git there is no\n> concept of central repository), you would get the following history\n>\n>                   /--X'<--3'<--4'            <- repo_B/master\n>                  /\n>  repo A:   1<---2<---a<---x<---c<---d<---E   <- master\n>                      \\\n>                       \\---x'<--c'\n>                                 \\\n>                                  \\--5        <- repo_B/master\n>\n> (because 'X' is now incoming, it needs to be \"bumped\", therefore\n> changing 3' and 4').\n>\n> The history without all this rewriting looks like this:\n>\n>                   /--X<---3<---4'            <- repo_B/master\n>                  /\n>  repo A:   1<---2<---a<---x<---c<---d<---E   <- master\n>                                 \\\n>                                  \\--5        <- repo_B/master\n>\n> Notice the difference?\n>\n>>>    And you also break the assumptions that the same commits have\n>>> the same contents (including date) and the same ID in different\n>>> repositories (some of which may include additional branches, some of\n>>> which may have been part of network of related repositories, etc.).\n>\n> See repo A and repo B in above example.\n>\n>> Wait...unless I completely misunderstand the hash-chain model, doesn't the\n>> hash of a commit depend on the hashes of its parents?  If that's the case,\n>> commits cannot have portable hashes. If it's not, please correct me.\n>>\n>> But if it's not, how does your first objection make sense?\n>\n> Hash of a commit depend in hashes of its parents (Merkle tree). That is\n> why signing a commit (or a tag pointing to the commit) signs a whole\n> history of a commit.\n>\n>>>> You don't need a daemon now to write commits to a repository. You can\n>>>> just add stuff to the object store, and then later flip the SHA-1 on a\n>>>> reference, we lock those indivdiual references, but this sort of thing\n>>>> would require a global write lock. This would introduce huge concurrency\n>>>> caveats that are non-issues now.\n>>>>\n>>>> Dumb clients matter. Now you can e.g. have two libgit2 processes writing\n>>>> to ref A and B respectively in the same repo, and they never have to\n>>>> know about each other or care about IPC.\n>>\n>> How do they know they're not writing to the same ref?  What keeps\n>> *that* operation atomic?\n>\n> Because different refs are stored in different files (at least for\n> \"live\" refs that are stores in loose ref format).  The lock is taken on\n> ref (to update ref and its reflog in sync), there is no need to take\n> global lock on all refs.\n>\n>>> You do realize that dates may not be monotonic (because of imperfections\n>>> in clock synchronization), thus the fact that the date is different from\n>>> parent does not mean that is different from ancestor.\n>>\n>> Good point. That means the O(log2 n) version of the check has to be done\n>> all the time.  Unfortunate.\n>\n> Especially with around 1 million of commits (Linux kernel, Chromium,\n> AOSP), or even 3M commits (MS Windows repository).\n>\n>>>>> That's the simple case. The complicated case is checking for date\n>>>>> collisions on *other* branches. But there are ways to make that fast,\n>>>>> too. There's a very obvious one involving a presort that is is O(log2\n>>>>> n) in the number of commits.\n>>>\n>>> I don't think performance hit you would get would be acceptable.\n>>\n>> Again, it's bad practice to assume rather than measure. Human intuitions\n>> about this sort of thing are notoriously unreliable.\n>\n> Techniques created to handle very large repositories (with respect to\n> number of commits) that make it possible for Git to avoid parsing commit\n> objects, namely bitmap index (for 'git fetch'/'clone') and serialized\n> commit graph (for 'git log') lead to _significant_ performance\n> improvements.\n>\n> The performance changes from \"waiting for Git to finish\" to \"done in the\n> blink of eye\" (well, almost).\n>\n>>>>> Excuse me, but your premise is incorrect.  A git DAG isn't just \"any\" DAG.\n>>>>> The presence of timestamps makes a total ordering possible.\n>>>>>\n>>>>> (I was a theoretical mathematician in a former life. This is all very\n>>>>> familiar ground to me.)\n>>>\n>>> Maybe in theory, when all clock are synchronized.\n>>\n>> My assertion does not depend on synchronized clocks, because it doesn't have to.\n>>\n>> If the timestamps in your repo are unique, there *is* a total ordering -\n>> by timestamp. What you don't get is guaranteed consistency with the\n>> topo ordering - that is you get no guarantee that a child's timestamp\n>> is greater than its parents'. That really would require a common\n>> timebase.\n>>\n>> But I don't need that stronger property, because the purpose of\n>> totally ordering the repo is to guarantee the uniqueness of action\n>> stamps.  For that, all I need is to be able to generate a unique cookie\n>> for each commit that can be inserted in its action stamp.\n>\n> For cookie to be unique among all forks / clones of the same repository\n> you need either centralized naming server, or for the cookie to be based\n> on contents of the commit (i.e. be a hash function).\n>\n>>                                                          For my use cases\n>> that cookie should *not* be a hash, because hashes always break N years\n>> down.  It should be an eternally stable product of the commit metadata.\n>\n> Well, the idea for SHA-1 <--> NewHash == SHA-256 transition is to avoid\n> having a flag day, and providing full interoperability between\n> repositories and Git installations using the old hash ad using new\n> hash^1.  This will be done internally by using SHA-1 <--> SHA-256\n> mapping.  So after the transition all you need is to publish this\n> mapping somewhere, be it with Internet Archive or Software Heritage.\n> Problem solved.\n>\n> P.S. Could you explain to me how one can use action stamp, e.g.\n> <esr@thyrsus.com!2019-05-15T20:01:15.473209800Z>, to quickly find the\n> commit it refers to?  With SHA-1 id you have either filesystem pathname\n> or the index file for pack to find it _fast_.\n>\n> Footnotes:\n> ----------\n> 1. That is why where would be no \"major format break\", thus no place for\n>    incompatibile format changes.\n>\n> Best,\n"},{"id":"375918","messageId":"20190520124039.GF11212@sigill.intra.peff.net","threadId":"51105","inReplyTo":"86r28tpikt.fsf@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2019-05-20T12:40:39Z","receivedAt":"2019-05-20T12:40:43Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, May 20, 2019 at 11:43:14AM +0200, Jakub Narebski wrote:\n\n> You can receive new commits in the repository by creating them, and from\n> other repository (via push or fetch).  In the second case you often get\n> many commits at once.\n> \n> In [1] it is described how using \"bitmap index\" you can avoid parsing\n> commits when deciding which objects to send to the client; they can be\n> directly copied to the client (added to the packfile that is sent to\n> client).  Thanks to this reachability bitmap (bit vector) the time to\n> clone Linux repository decreased from 57 seconds to 1.6 seconds.\n\nNo, this is mixing up sending and receiving. On the sending side, we try\nvery hard not to open up objects if we can avoid it (using tricks like\nreachability bitmaps helps us quickly decide what to send, and reusing\nthe on-disk packfile data lets us send out objects without decompressing\nthem).\n\nBut on the receiving side, we do not trust the sender at all. The\nprotocol specifically does not send the sha1 of any object. The receiver\ninstead inflates every object it gets and computes the object hash\nitself. And then on top of that, we traverse the commit graph to make\nsure that the server sent us all of the objects we need to have a\ncomplete graph.\n\nSo adding any extra object-quality checks on the receiving side would\nnot really change that equation.\n\nBut I do otherwise agree with your mail that the general idea of having\nthe receiver _change_ the incoming objects is going to lead to a world\nof headaches.\n\n-Peff\n"},{"id":"375925","messageId":"20190520141417.GA83559@thyrsus.com","threadId":"51105","inReplyTo":"86r28tpikt.fsf@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-20T14:14:17Z","receivedAt":"2019-05-20T14:14:21Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jakub Narebski <jnareb@gmail.com>:\n> > What \"commits that follow it?\" By hypothesis, the incoming commit's\n> > timestamp is bumped (if it's bumped) when it's first added to a branch\n> > or branches, before there are following commits in the DAG.\n> \n> Errr... the main problem is with distributed nature of Git, i.e. when\n> two repositories create different commits with the same\n> committer+timestamp value.  You receive commits on fetch or push, and\n> you receive many commits at once.\n> \n> Say you have two repositories, and the history looks like this:\n> \n>  repo A:   1<---2<---a<---x<---c<---d      <- master\n> \n>  repo B:   1<---2<---X<---3<---4           <- master\n> \n> When you push from repo A to repo B, or fetch in repo B from repo A you\n> would get the following DAG of revisions\n> \n>  repo B:   1<---2<---X<---3<---4           <- master\n>                  \\\n>                   \\--a<---x<---c<---d      <- repo_A/master\n> \n> Now let's assume that commits X and x have the came committer and the\n> same fractional timestamp, while being different commits.  Then you\n> would need to bump timestamp of 'x', changing the commit.  This means\n> that 'c' needs to be rewritten too, and 'd' also:\n> \n>  repo B:   1<---2<---X<---3<---4           <- master\n>                  \\\n>                   \\--a<---x'<--c'<--d'     <- repo_A/master\n\nOf course that's true.  But you were talking as though all those commits\nhave to be modified *after they're in the DAG*, and that's not the case.\nIf any timestamp has to be modified, it only has to happen *once*, at the\ntime its commit enters the repo.\n\nActually, in the normal case only x would need to be modified. The only\nway c would need to be modified is if bumping x's timestamp caused an\nactual collision with c's.\n\nI don't see any conceptual problem with this.  You appear to me to be\nconfusing two issues.  Yes, bumping timestamps would mean that all\nhashes downstream in the Merkle tree would be generated differently,\neven when there's no timestamp collision, but so what?  The hash of a\ncommit isn't portable to begin with - it can't be, because AFAIK\nthere's no guarantee that the ancestry parts of the DAG in two\nrepositories where copies of it live contain all the same commits and\ntopo relationships.\n\n> And now for the final nail in the coffing of the Bazaar-esque idea of\n> changing commits on arrival.  Say that repository A created new commits,\n> and pushed them to B.  You would need to rewrite all future commits from\n> this repository too, and you would always fetch all commits starting\n> from the first \"bumped\"\n\nI don't see how the second clause of your last sentence follows from the\nfirst unless commit hashes really are supposed to be portable across\nrepositories.  And I don't see how that can be so given that 'git am'\nexists and a branch can thus be rooted at a different place after\nit is transported and integrated.\n\n> Hash of a commit depend in hashes of its parents (Merkle tree). That is\n> why signing a commit (or a tag pointing to the commit) signs a whole\n> history of a commit.\n\nThat's what I thought.\n\n> > How do they know they're not writing to the same ref?  What keeps\n> > *that* operation atomic?\n> \n> Because different refs are stored in different files (at least for\n> \"live\" refs that are stores in loose ref format).  The lock is taken on\n> ref (to update ref and its reflog in sync), there is no need to take\n> global lock on all refs.\n\nOK, that makes sense.\n\n> For cookie to be unique among all forks / clones of the same repository\n> you need either centralized naming server, or for the cookie to be based\n> on contents of the commit (i.e. be a hash function).\n\nI don't need uniquess across all forks, only uniqueness *within the repo*.\n\nI want this for two reasons: (1) so that action stamps are unique, (2)\nso that there is a unique canonical ordering of commits in a fast export\nstream.\n\n(Without that second property there are surgical cases I can't\nregression-test.)\n\n> >                                                          For my use cases\n> > that cookie should *not* be a hash, because hashes always break N years\n> > down.  It should be an eternally stable product of the commit metadata.\n> \n> Well, the idea for SHA-1 <--> NewHash == SHA-256 transition is to avoid\n> having a flag day, and providing full interoperability between\n> repositories and Git installations using the old hash ad using new\n> hash^1.  This will be done internally by using SHA-1 <--> SHA-256\n> mapping.  So after the transition all you need is to publish this\n> mapping somewhere, be it with Internet Archive or Software Heritage.\n> Problem solved.\n\nI don't see it.  How does this prevent old clients from barfing on new\nrepositories?\n\n> P.S. Could you explain to me how one can use action stamp, e.g.\n> <esr@thyrsus.com!2019-05-15T20:01:15.473209800Z>, to quickly find the\n> commit it refers to?  With SHA-1 id you have either filesystem pathname\n> or the index file for pack to find it _fast_.\n\nFor the purposes that make action stamps important I don't really care\nabout performance much (though there are fairly obvious ways to\nachieve it).  My goal is to ensure that revision histories (e.g. in\ntheir import-stream format) are forward-portable to future VCSes\nwithout requiring any data outside the stream itself.\n\nPlease remember that I'm accustomed to maintaining infrastructure on\ndecadal timescales - I wrote code in the 1980s that is still in wide use\nand I expect some of the code I'm writing now to be still in use thirty\nyears from now.\n\nThis gives me a different perspective on the fragility of things like\nSHA-1 hashes.  From a decadal-scale POV any particular crypto-hash\nformat is unstable garbage, and having them in change comments is a\nmaintainability disaster waiting to happen.\n\nAction stamps are specifically designed so that they're pointers to commits\nthat don't require anything but the target commit's import/export-stream\nmetadata to resolve.  Your idea of an archived hash registry makes me\nextremely nervous; I think it's too fragile to trust.\n\nSo let me back up a step.  I will cheerfully drop advocating bumping\ntimestamps if anyone can tell me how a different way to define a per-commit\nreference cookie that (a) is unique within its repo, and (b) only requires\nmetadata visible in the fast-export representation of the commit.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375926","messageId":"20190520164134.6b35b9f9@kitsune.suse.cz","threadId":"51105","inReplyTo":"20190520141417.GA83559@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2019-05-20T14:41:34Z","receivedAt":"2019-05-20T14:41:40Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Mon, 20 May 2019 10:14:17 -0400\n\"Eric S. Raymond\" <esr@thyrsus.com> wrote:\n\n> Jakub Narebski <jnareb@gmail.com>:\n> > > What \"commits that follow it?\" By hypothesis, the incoming commit's\n> > > timestamp is bumped (if it's bumped) when it's first added to a branch\n> > > or branches, before there are following commits in the DAG.  \n> > \n> > Errr... the main problem is with distributed nature of Git, i.e. when\n> > two repositories create different commits with the same\n> > committer+timestamp value.  You receive commits on fetch or push, and\n> > you receive many commits at once.\n> > \n> > Say you have two repositories, and the history looks like this:\n> > \n> >  repo A:   1<---2<---a<---x<---c<---d      <- master\n> > \n> >  repo B:   1<---2<---X<---3<---4           <- master\n> > \n> > When you push from repo A to repo B, or fetch in repo B from repo A you\n> > would get the following DAG of revisions\n> > \n> >  repo B:   1<---2<---X<---3<---4           <- master\n> >                  \\\n> >                   \\--a<---x<---c<---d      <- repo_A/master\n> > \n> > Now let's assume that commits X and x have the came committer and the\n> > same fractional timestamp, while being different commits.  Then you\n> > would need to bump timestamp of 'x', changing the commit.  This means\n> > that 'c' needs to be rewritten too, and 'd' also:\n> > \n> >  repo B:   1<---2<---X<---3<---4           <- master\n> >                  \\\n> >                   \\--a<---x'<--c'<--d'     <- repo_A/master  \n> \n> Of course that's true.  But you were talking as though all those commits\n> have to be modified *after they're in the DAG*, and that's not the case.\n> If any timestamp has to be modified, it only has to happen *once*, at the\n> time its commit enters the repo.\n\nAnd that's where you get it wrong. Git is *distributed*. There is more\nthan one repository. Each repository has its own DAG that is completely\nunrelated to the other repositories and their DAGs. So when you take\nyour history and push it to another repository and the timestamps\nchange as the result what ends up in the other repository is not the\nhistory you pushed. So the repositories diverge and you no longer know\nwhat is what.\n\n> \n> Actually, in the normal case only x would need to be modified. The only\n> way c would need to be modified is if bumping x's timestamp caused an\n> actual collision with c's.\n> \n> I don't see any conceptual problem with this.  You appear to me to be\n> confusing two issues.  Yes, bumping timestamps would mean that all\n> hashes downstream in the Merkle tree would be generated differently,\n> even when there's no timestamp collision, but so what?  The hash of a\n> commit isn't portable to begin with - it can't be, because AFAIK\n> there's no guarantee that the ancestry parts of the DAG in two\n> repositories where copies of it live contain all the same commits and\n> topo relationships.\n\nIf you push form one repository to another repository now you get exact\nsame history with exact same hashes. So the hashes are portable across\nrepositories that share history. With your proposed change hashes can\nbe modified on push/pull so repositories no longer share history and\nhashes become non-portable. That's why it is a bad idea.\n\nThe commits are currently identified by the hash so it must not change\nduring push/pull. Changing the identifier to something else (eg content\nhas without (some) metadata) might be useful to make the identifier\nmore stable but will bring other problems when you need two\ndifferent identifiers for the same content to include it in two\nunrelated histories.\n\nThanks\n\nMichal\n"},{"id":"375928","messageId":"20190520170518.73ad912b@kitsune.suse.cz","threadId":"51105","inReplyTo":"3b8d6a78-bd88-770c-e79b-d732f7e277fd@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Michal Suchánek","fromEmail":"msuchanek@suse.de","sentAt":"2019-05-20T15:05:18Z","receivedAt":"2019-05-20T15:05:23Z","isPatch":false,"sender":{"key":"msuchanek@suse.de","avatar":"https://avatars.githubusercontent.com/u/787652?v=4"},"body":"On Wed, 15 May 2019 21:25:46 -0400\nDerrick Stolee <stolee@gmail.com> wrote:\n\n> On 5/15/2019 8:28 PM, Eric S. Raymond wrote:\n> > Derrick Stolee <stolee@gmail.com>:  \n> >> What problem are you trying to solve where commit date is important?  \n\n> > B. Unique canonical form of import-stream representation.\n> > \n> > Reposurgeon is a very complex piece of software with subtle failure\n> > modes.  I have a strong need to be able to regression-test its\n> > operation.  Right now there are important cases in which I can't do\n> > that because (a) the order in which it writes commits and (b) how it\n> > colors branches, are both phase-of-moon dependent.  That is, the\n> > algorithms may be deterministic but they're not documented and seem to\n> > be dependent on variables that are hidden from me.\n> > \n> > Before import streams can have a canonical output order without hidden\n> > variables (e.g. depending only on visible metadata) in practice, that\n> > needs to be possible in principle. I've thought about this a lot and\n> > not only are unique commit timestamps the most natural way to make\n> > it possible, they're the only way conistent with the reality that\n> > commit comments may be altered for various good reasons during\n> > repository translation.  \n> \n> If you are trying to debug or test something, why don't you serialize\n> the input you are using for your test?\n\nAnd that's the problem. Serialization of a git repository is not stable\nbecause there is no total ordering on commits. And for testing you need\nto serialize some 'before' and 'after' state and they can be totally\ndifferent. Not because the repository state is totally different but\nbecause the serialization of the state is not stable.\n\nThanks\n\nMichal\n"},{"id":"375932","messageId":"20190520163625.GA99397@thyrsus.com","threadId":"51105","inReplyTo":"20190520170518.73ad912b@kitsune.suse.cz","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-20T16:36:25Z","receivedAt":"2019-05-20T16:36:27Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Michal Suchánek <msuchanek@suse.de>:\n> On Wed, 15 May 2019 21:25:46 -0400\n> Derrick Stolee <stolee@gmail.com> wrote:\n> \n> > On 5/15/2019 8:28 PM, Eric S. Raymond wrote:\n> > > Derrick Stolee <stolee@gmail.com>:  \n> > >> What problem are you trying to solve where commit date is important?  \n> \n> > > B. Unique canonical form of import-stream representation.\n> > > \n> > > Reposurgeon is a very complex piece of software with subtle failure\n> > > modes.  I have a strong need to be able to regression-test its\n> > > operation.  Right now there are important cases in which I can't do\n> > > that because (a) the order in which it writes commits and (b) how it\n> > > colors branches, are both phase-of-moon dependent.  That is, the\n> > > algorithms may be deterministic but they're not documented and seem to\n> > > be dependent on variables that are hidden from me.\n> > > \n> > > Before import streams can have a canonical output order without hidden\n> > > variables (e.g. depending only on visible metadata) in practice, that\n> > > needs to be possible in principle. I've thought about this a lot and\n> > > not only are unique commit timestamps the most natural way to make\n> > > it possible, they're the only way conistent with the reality that\n> > > commit comments may be altered for various good reasons during\n> > > repository translation.  \n> > \n> > If you are trying to debug or test something, why don't you serialize\n> > the input you are using for your test?\n> \n> And that's the problem. Serialization of a git repository is not stable\n> because there is no total ordering on commits. And for testing you need\n> to serialize some 'before' and 'after' state and they can be totally\n> different. Not because the repository state is totally different but\n> because the serialization of the state is not stable.\n\nYes, msuchanek is right - that is exactly the problem.  Very well put.\n\ngit fast-import streams *are* the serialization; they're what reposurgeon\ningests and emits.  The concrete problem I have is that there is no stable\ncorrespondence between a repository and one canonical fast-import\nserialization of it.\n\nThat is a bigger pain in the ass than you will be able to imagine unless\nand until you try writing surgical tools yourself and discover that you\ncan't write tests for them.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375941","messageId":"7e88805c-7e08-2631-599d-b47a098f1ce1@gmail.com","threadId":"51105","inReplyTo":"20190520163625.GA99397@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2019-05-20T17:22:15Z","receivedAt":"2019-05-20T17:22:19Z","isPatch":false,"sender":{"key":"stolee@gmail.com","avatar":"https://avatars.githubusercontent.com/u/570044?v=4"},"body":"On 5/20/2019 12:36 PM, Eric S. Raymond wrote:\n> Michal Suchánek <msuchanek@suse.de>:\n>> On Wed, 15 May 2019 21:25:46 -0400\n>> Derrick Stolee <stolee@gmail.com> wrote:\n>>\n>>> On 5/15/2019 8:28 PM, Eric S. Raymond wrote:\n>>>> Derrick Stolee <stolee@gmail.com>:  \n>>>>> What problem are you trying to solve where commit date is important?  \n>>\n>>>> B. Unique canonical form of import-stream representation.\n>>>>\n>>>> Reposurgeon is a very complex piece of software with subtle failure\n>>>> modes.  I have a strong need to be able to regression-test its\n>>>> operation.  Right now there are important cases in which I can't do\n>>>> that because (a) the order in which it writes commits and (b) how it\n>>>> colors branches, are both phase-of-moon dependent.  That is, the\n>>>> algorithms may be deterministic but they're not documented and seem to\n>>>> be dependent on variables that are hidden from me.\n>>>>\n>>>> Before import streams can have a canonical output order without hidden\n>>>> variables (e.g. depending only on visible metadata) in practice, that\n>>>> needs to be possible in principle. I've thought about this a lot and\n>>>> not only are unique commit timestamps the most natural way to make\n>>>> it possible, they're the only way conistent with the reality that\n>>>> commit comments may be altered for various good reasons during\n>>>> repository translation.  \n>>>\n>>> If you are trying to debug or test something, why don't you serialize\n>>> the input you are using for your test?\n>>\n>> And that's the problem. Serialization of a git repository is not stable\n>> because there is no total ordering on commits. And for testing you need\n>> to serialize some 'before' and 'after' state and they can be totally\n>> different. Not because the repository state is totally different but\n>> because the serialization of the state is not stable.\n> \n> Yes, msuchanek is right - that is exactly the problem.  Very well put.\n> \n> git fast-import streams *are* the serialization; they're what reposurgeon\n> ingests and emits.  The concrete problem I have is that there is no stable\n> correspondence between a repository and one canonical fast-import\n> serialization of it.\n> \n> That is a bigger pain in the ass than you will be able to imagine unless\n> and until you try writing surgical tools yourself and discover that you\n> can't write tests for them.\n\nWhat it sounds like you are doing is piping a 'git fast-import' process into\nreposurgeon, and testing that reposurgeon does the same thing every time.\nOf course this won't be consistent if 'git fast-import' isn't consistent.\n\nBut what you should do instead is store a fixed file from one run of\n'git fast-import' and send that file to reposurgeon for the repeated test.\nDon't rely on fast-import being consistent and instead use fixed input for\nyour test.\n\nIf reposurgeon is providing the input to _and_ consuming the output from\n'git fast-import', then yes you will need to have at least one integration\ntest that runs the full pipeline. But for regression tests covering complicated\nlogic in reposurgeon, you're better off splitting the test (or mocking out\n'git fast-import' with something that provides consistent output given\nfixed input).\n\n-Stolee\n \n\n"},{"id":"375956","messageId":"20190520213203.GA110573@thyrsus.com","threadId":"51105","inReplyTo":"7e88805c-7e08-2631-599d-b47a098f1ce1@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-20T21:32:03Z","receivedAt":"2019-05-20T21:32:05Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Derrick Stolee <stolee@gmail.com>:\n> What it sounds like you are doing is piping a 'git fast-import' process into\n> reposurgeon, and testing that reposurgeon does the same thing every time.\n> Of course this won't be consistent if 'git fast-import' isn't consistent.\n\nIt's not actually import that fails to have consistent behavior, it's export.\n\nThat is, if I fast-import a given stream, I get indistinguishable\nin-core commit DAGs every time. (It would be pretty alarming if this\nweren't true!)\n\nWhat I have no guarantee of is the other direction.  In a multibranch repo,\nfast-export writes out branches in an order I cannot predict and which\nappears from the outside to be randomly variable.\n\n> But what you should do instead is store a fixed file from one run of\n> 'git fast-import' and send that file to reposurgeon for the repeated test.\n> Don't rely on fast-import being consistent and instead use fixed input for\n> your test.\n> \n> If reposurgeon is providing the input to _and_ consuming the output from\n> 'git fast-import', then yes you will need to have at least one integration\n> test that runs the full pipeline. But for regression tests covering complicated\n> logic in reposurgeon, you're better off splitting the test (or mocking out\n> 'git fast-import' with something that provides consistent output given\n> fixed input).\n\nAnd I'd do that... but the problem is more fundamental than you seem to\nunderstand.  git fast-export can't ship a consistent output order because\nit doesn't retain metadata sufficient to totally order child branches.\n\nThis is why I wanted unique timestamps.  That would solve the problem,\nbranch child commits of any node would be ordered by their commit date.\n\nBut I had a realization just now.  A much smaller change would do it.\nSuppose branch creations had creation stamps with a weak uniqueness property;\nfor any given parent node, the creation stamps of all branches originating\nthere are guaranteed to be unique?\n\nIf that were true, there would be an implied total ordering of the\nrepository.  The rules for writing out a totally ordered dump would go\nlike this:\n\n1. At any given step there is a set of active branches and a cursor\non each such branch.  Each cursor points at a commit and caches the\ncreation stamp of the current branch.\n\n2. Look at the set of commits under the cursors.  Write the oldest one.\nIf multiple commits have the same commit date, break ties by their\nbranch creation stamps.\n\n3. Bump that cursor forward. If you're at a branch creation, it\nbecomes multiple cursors, one for each child branch.\nIf you're at a join, some cursors go away.\n\nHere's the clever bit - you make the creation stamp nothing but a\ncounter that says \"This was the Nth branch creation.\"  And it is\nset by these rules:\n\n4. If the branch creation stamp is undefined at branch creation time,\nnumber it in any way you like as long as each stamp is unique. A\ndefined, documented order would be nice but is not necessary for\nstreams to round-trip.\n\n5. When writing an export stream, you always utter a reset at the\npoint of branch creation.\n\n6. When reading an import stream, the ordinal for a new branch is\ndefined as the number of resets you have seen.\n\nRules 5 and 6 together guarantee that branch creation ordinals round-trip\nthrough export streams.  Thus, streams round-trip and I can have my\nregression tests with no change to git's visible interface at all!\n\nI could write this code.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375957","messageId":"CABPp-BHK1N2zZoeBeSgnh12LPqLgZxfbL0DzALj28y97_Q-ahg@mail.gmail.com","threadId":"51105","inReplyTo":"20190520141417.GA83559@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Elijah Newren","fromEmail":"newren@gmail.com","sentAt":"2019-05-20T21:38:20Z","receivedAt":"2019-05-20T21:38:33Z","isPatch":false,"sender":{"key":"newren@gmail.com","avatar":"https://avatars.githubusercontent.com/u/5455730?v=4"},"body":"Hi,\n\nOn Mon, May 20, 2019 at 11:09 AM Eric S. Raymond <esr@thyrsus.com> wrote:\n\n> > For cookie to be unique among all forks / clones of the same repository\n> > you need either centralized naming server, or for the cookie to be based\n> > on contents of the commit (i.e. be a hash function).\n>\n> I don't need uniquess across all forks, only uniqueness *within the repo*.\n\nYou've lost me.  In other places you stated you didn't want to use the\ncommit hash, and now you say this.  If you only care about uniqueness\nwithin the current copy of the repo and don't care about uniqueness\nacross forks (i.e. clones or copies that exist now or in the future --\nincluding copies stored using SHA256), then what's wrong with using\nthe commit hash?\n\n> I want this for two reasons: (1) so that action stamps are unique, (2)\n> so that there is a unique canonical ordering of commits in a fast export\n> stream.\n\nA stable ordering of commits in a fast-export stream might be a cool\nfeature.  But I don't know how to define one, other than perhaps sort\nfirst by commit-depth (maybe optionally adding a few additional\nintermediate sorting criteria), and then finally sort by commit hash\nas a tiebreaker.  Without the fallback to commit hash, you fall back\non normal traversal order which isn't stable (it depends on e.g. order\nof branches listed on the command line to fast-export, or if using\n--all, what new branch you just added that comes alphabetically before\nothers).\n\nI suspect that solution might run afoul of your dislike for commit\nhashes, though, so I'm not sure it'd work for you.\n\n> (Without that second property there are surgical cases I can't\n> regression-test.)\n>\n> > >                                                          For my use case\n> > > that cookie should *not* be a hash, because hashes always break N years\n> > > down.  It should be an eternally stable product of the commit metadata.\n> >\n> > Well, the idea for SHA-1 <--> NewHash == SHA-256 transition is to avoid\n> > having a flag day, and providing full interoperability between\n> > repositories and Git installations using the old hash ad using new\n> > hash^1.  This will be done internally by using SHA-1 <--> SHA-256\n> > mapping.  So after the transition all you need is to publish this\n> > mapping somewhere, be it with Internet Archive or Software Heritage.\n> > Problem solved.\n>\n> I don't see it.  How does this prevent old clients from barfing on new\n> repositories?\n\nDepends on range of time for \"old\".  The plan as I understood it\n(which is suspect): make git version which understand both SHA-1 and\nSHA-256 (which I think is already done, though I haven't followed\nclosely), wait some time, allow people to opt in to converting, allow\nmore time, consider ways of nudging people to switch.\n\nYou are right that clients older than any version that understands\nSHA-256 would barf on the new repositories.\n\n> So let me back up a step.  I will cheerfully drop advocating bumping\n> timestamps if anyone can tell me how a different way to define a per-commit\n> reference cookie that (a) is unique within its repo, and (b) only requires\n> metadata visible in the fast-export representation of the commit.\n\nDoes passing --show-original-ids option to fast-export and using the\nresulting original-oid field as the cookie count?\n"},{"id":"375963","messageId":"2f0ab8c9-adf4-416d-519a-a313de89d5e1@iee.org","threadId":"51105","inReplyTo":"20190520164134.6b35b9f9@kitsune.suse.cz","subject":"Re: Finer timestamps and serialization in git","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.org","sentAt":"2019-05-20T22:18:50Z","receivedAt":"2019-05-20T22:28:04Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"Hi,\n\nOn 20/05/2019 15:41, Michal Suchánek wrote:\n>>   But you were talking as though all those commits\n>> have to be modified*after they're in the DAG*, and that's not the case.\n>> If any timestamp has to be modified, it only has to happen*once*, at the\n>> time its commit enters the repo.\n> And that's where you get it wrong. Git is*distributed*. There is more\n> than one repository. Each repository has its own DAG\nSo far so good. In fact it the change to 'distributed' that has ruined \nEric's Acton stamps that assume that the 'time' came from a single \ncentral server.\n>   that is completely\n> unrelated to the other repositories and their DAGs.\nThis bit will confuse. It is only the new commits in the different \nrepositories that are 'unrelated'. Their common history commits are \nidentical sha1 values, and the DAG links back to their common root commit(s)\n> So when you take\n> your history and push it to another repository and the timestamps\n> change as the result what ends up in the other repository is not the\n> history you pushed. So the repositories diverge and you no longer know\n> what is what.\n>\nIf the sender tweaks their timestamps at commit time, then no one \n'knows'. It's just a minor bit of clock drift/slop. But once they have a \ncascaded history which has been published (and used) you are locked into \nthat.\n\n\nAs noted previously. The significant change is the loss of the central \nserver and the referential nature of it's clock time stamp.\n\n\nIf the action stamp is just a useful temporary intermediary in a \ntransfer then cheats are possible (e.g. some randomising hash of a \ndefinative partr of the commit).\n\n\nBut if the action stamps are meant to be permanent and re-generatable \nfor a round trip between a central server change set based server to \nGit, and then back again, repeatably, without divergence, loss, or \nchange, then it is not going to happen reliably. To do so requires the \ncreation of fixed total order (by design - single clock) from commits \nthat are only partially ordered (by design! - DAG rather than multiple \nunsynchronized user clocks).\n\n\nFor backward compatibility Git only has (and only needs 1 second \nresolution).\n\n\nThe multi-decade/century VCS idea of a master artifact and then near \ncopies (since koalin and linen drawings, blue prints, ..) with central \n_control_ is being replaced by zero cost perfect replication, \nauthentication by hash, with its distribution of control (of artifact \nentry into the VCS) to _users_, from managers. Managers simply select \nand decide on the artifact quality and authorize the use of a hash.\n\n\nMost folks haven't really looked below the surface of what it is that \nmakes GIT and DVCS so successful, and it's not just the Linus effect. \nThe previous certainties (e.g. the idea of a total order to allow \nlogging by change-set) have gone.\n\n--\n\nPhilip\n\n"},{"id":"375966","messageId":"20190520231223.GA117962@thyrsus.com","threadId":"51105","inReplyTo":"CABPp-BHK1N2zZoeBeSgnh12LPqLgZxfbL0DzALj28y97_Q-ahg@mail.gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-20T23:12:23Z","receivedAt":"2019-05-20T23:12:25Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Elijah Newren <newren@gmail.com>:\n> Hi,\n> \n> On Mon, May 20, 2019 at 11:09 AM Eric S. Raymond <esr@thyrsus.com> wrote:\n> \n> > > For cookie to be unique among all forks / clones of the same repository\n> > > you need either centralized naming server, or for the cookie to be based\n> > > on contents of the commit (i.e. be a hash function).\n> >\n> > I don't need uniquess across all forks, only uniqueness *within the repo*.\n> \n> You've lost me.  In other places you stated you didn't want to use the\n> commit hash, and now you say this.  If you only care about uniqueness\n> within the current copy of the repo and don't care about uniqueness\n> across forks (i.e. clones or copies that exist now or in the future --\n> including copies stored using SHA256), then what's wrong with using\n> the commit hash?\n\nBecause it's not self-describing, can't be computed solely from visible\ncommit metadata, and relies on complex external assumptions about how\nthe hash is computed which break when your VCS changes hash algorithms.\n\nThese are dealbreakers because one of my major objectives is forward\nportability of these IDs forever. And I mean *forever*.  It should be\npossible for someone in the year 40,000, in between assaulting planets\nfor the God-Emperor, to look at an import stream and deduce how to\nresolve the cookies to their commits without seeing git's code or\nknowing anything about its hash algorithms.\n\nI think maybe the reason I'm having so much trouble getting this\nacross is that git insiders are used to thinking of import streams as\ntransient things.  Because I do a lot of repo migrations, I have a\nvery different view of them.  I built reposurgeon on the realization\nthat they're a general transport format for revision histories, and\nthat has forward value independent of the existence of git.\n\nIf a stream contained fully forward-portable action stamps, it would be\nforward-portable forever.  Hashes in commit comments are the *only*\nblocker to that.  Take this from a person who has spent way too much time\npatching Subversion IDs like r1234 during repository conversions.\n\nIt would take so little to make this work. Existing stream format is\n*almost there*.\n\n> A stable ordering of commits in a fast-export stream might be a cool\n> feature.  But I don't know how to define one, other than perhaps sort\n> first by commit-depth (maybe optionally adding a few additional\n> intermediate sorting criteria), and then finally sort by commit hash\n> as a tiebreaker. Without the fallback to commit hash, you fall back\n> on normal traversal order which isn't stable (it depends on e.g. order\n> of branches listed on the command line to fast-export, or if using\n> --all, what new branch you just added that comes alphabetically before\n> others).\n>\n> I suspect that solution might run afoul of your dislike for commit\n> hashes, though, so I'm not sure it'd work for you.\n\nIt does. See above.\n\n> > So let me back up a step.  I will cheerfully drop advocating bumping\n> > timestamps if anyone can tell me how a different way to define a per-commit\n> > reference cookie that (a) is unique within its repo, and (b) only requires\n> > metadata visible in the fast-export representation of the commit.\n> \n> Does passing --show-original-ids option to fast-export and using the\n> resulting original-oid field as the cookie count?\n\nI was not aware of this option.  Looking...no wonder, it's not on my\nsystem man page.  Must be recent.\n\nOK. Wow.  That is *useful*, and I am going to upgrade reposurgeon to read\nit.  With that I can do automatic commit-reference rewriting.\n\nI don't consider it a complete solution. The problem is that OID is\na consistent property that can be used to resolve cookies, but there's\nno guaranteed that it's a *preserved* property that survives multiple\nround trips and changes in hash functions.\n\nSo the right way to use it is to pick it up, do reference-cookie\nresolution, and then mung the reference cookies to a format that is\nstable forever.  I don't know what that format should be yet.  I\nhave a message in composition about this.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"},{"id":"375970","messageId":"86v9y4oeiu.fsf@gmail.com","threadId":"51105","inReplyTo":"20190520141417.GA83559@thyrsus.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2019-05-21T00:08:25Z","receivedAt":"2019-05-21T00:08:32Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"\"Eric S. Raymond\" <esr@thyrsus.com> writes:\n> Jakub Narebski <jnareb@gmail.com>:\n\n>>> What \"commits that follow it?\" By hypothesis, the incoming commit's\n>>> timestamp is bumped (if it's bumped) when it's first added to a branch\n>>> or branches, before there are following commits in the DAG.\n>> \n>> Errr... the main problem is with distributed nature of Git, i.e. when\n>> two repositories create different commits with the same\n>> committer+timestamp value.  You receive commits on fetch or push, and\n>> you receive many commits at once.\n>> \n>> Say you have two repositories, and the history looks like this:\n>> \n>>  repo A:   1<---2<---a<---x<---c<---d      <- master\n>> \n>>  repo B:   1<---2<---X<---3<---4           <- master\n>> \n>> When you push from repo A to repo B, or fetch in repo B from repo A you\n>> would get the following DAG of revisions\n>> \n>>  repo B:   1<---2<---X<---3<---4           <- master\n>>                  \\\n>>                   \\--a<---x<---c<---d      <- repo_A/master\n>> \n>> Now let's assume that commits X and x have the came committer and the\n>> same fractional timestamp, while being different commits.  Then you\n>> would need to bump timestamp of 'x', changing the commit.  This means\n>> that 'c' needs to be rewritten too, and 'd' also:\n>> \n>>  repo B:   1<---2<---X<---3<---4           <- master\n>>                  \\\n>>                   \\--a<---x'<--c'<--d'     <- repo_A/master\n>\n> Of course that's true.  But you were talking as though all those commits\n> have to be modified *after they're in the DAG*, and that's not the case.\n> If any timestamp has to be modified, it only has to happen *once*, at the\n> time its commit enters the repo.\n\nThe time commit 'x' was created in repo A there was no need to bump the\ntimestamp.  Same with commit 'X' in repo B (well, unless there is a\ncentral serialization server - which would not fly).  It is only after\npush from repo A to repo B that we have two commits: 'x' and 'X' with\nthe same timestamp.\n\n> Actually, in the normal case only x would need to be modified. The only\n> way c would need to be modified is if bumping x's timestamp caused an\n> actual collision with c's.\n>\n> I don't see any conceptual problem with this.  You appear to me to be\n> confusing two issues.  Yes, bumping timestamps would mean that all\n> hashes downstream in the Merkle tree would be generated differently,\n> even when there's no timestamp collision, but so what?  The hash of a\n> commit isn't portable to begin with - it can't be, because AFAIK\n> there's no guarantee that the ancestry parts of the DAG in two\n> repositories where copies of it live contain all the same commits and\n> topo relationships.\n\nErrr... how did you get that the hash of a commit is not portable???\nSame contents means same hash, i.e. same object identifier.  Two\nrepositories can have part of history in common (for example different\nforks of the same repository, like different \"trees\" of Linux kernel),\nsharing part of DAG.  Same commits, same topo relationships.  That's how\n_distributed_ version control works.\n\n[I think we may have been talking past each other.]\n\n>> And now for the final nail in the coffing of the Bazaar-esque idea of\n>> changing commits on arrival.  Say that repository A created new commits,\n>> and pushed them to B.  You would need to rewrite all future commits from\n>> this repository too, and you would always fetch all commits starting\n>> from the first \"bumped\"\n>\n> I don't see how the second clause of your last sentence follows from the\n> first unless commit hashes really are supposed to be portable across\n> repositories.  And I don't see how that can be so given that 'git am'\n> exists and a branch can thus be rooted at a different place after\n> it is transported and integrated.\n\n'git rebase', 'git rebase --interactive' and 'git am' create diffent\ncommits; that is why their's result is called \"history rewriting\" (it\nactually is creating altered copy, and garbage-collecting old pre-copy\nand pre-change version).  Anyway, the recommended practice is to not\nrewrite published history (where somebody could have bookmarked it).\n\nNote also that this copy preserves author date, not committer date; also\ncommits can be deleted, split and merged during \"rewrite\".\n\nFetch and push do not use 'git am', and they preserve commits and their\nidentities.  That is how they can be effective and peformant.\n\n>> Hash of a commit depend in hashes of its parents (Merkle tree). That is\n>> why signing a commit (or a tag pointing to the commit) signs a whole\n>> history of a commit.\n>\n> That's what I thought.\n\n[...]\n>> For cookie to be unique among all forks / clones of the same repository\n>> you need either centralized naming server, or for the cookie to be based\n>> on contents of the commit (i.e. be a hash function).\n>\n> I don't need uniquess across all forks, only uniqueness *within the repo*.\n\nErr, what?  So the proposed \"action stamp\" identifier is even more\nuseless?  If you can't use <esr@thyrsus.com!2019-05-15T20:01:15.473209800Z>\nto uniquely name revision, so that every person that has that commit can\nknow which commit is it, what's the use?\n\nIs \"action stamp\" meant to be some local identifier, like Mercurial's\nSubversion-like revision number, good only for local repository?\n\n> I want this for two reasons: (1) so that action stamps are unique, (2)\n> so that there is a unique canonical ordering of commits in a fast export\n> stream.\n>\n> (Without that second property there are surgical cases I can't\n> regression-test.)\n\nYou can always use object identifier (hash) for tiebreaking for second\ncase use.\n\n>>>                                                          For my use cases\n>>> that cookie should *not* be a hash, because hashes always break N years\n>>> down.  It should be an eternally stable product of the commit metadata.\n>> \n>> Well, the idea for SHA-1 <--> NewHash == SHA-256 transition is to avoid\n>> having a flag day, and providing full interoperability between\n>> repositories and Git installations using the old hash ad using new\n>> hash^1.  This will be done internally by using SHA-1 <--> SHA-256\n>> mapping.  So after the transition all you need is to publish this\n>> mapping somewhere, be it with Internet Archive or Software Heritage.\n>> Problem solved.\n>\n> I don't see it.  How does this prevent old clients from barfing on new\n> repositories?\n\nThe SHA-1 <--> SHA-256 interoperation is on the client-server level; one\ncan use old Git that uses SHA-1 from repository that uses SHA-256, and\nvice versa.\n\n>> P.S. Could you explain to me how one can use action stamp, e.g.\n>> <esr@thyrsus.com!2019-05-15T20:01:15.473209800Z>, to quickly find the\n>> commit it refers to?  With SHA-1 id you have either filesystem pathname\n>> or the index file for pack to find it _fast_.\n>\n> For the purposes that make action stamps important I don't really care\n> about performance much (though there are fairly obvious ways to\n> achieve it).\n\nWhat ways?\n\n>              My goal is to ensure that revision histories (e.g. in\n> their import-stream format) are forward-portable to future VCSes\n> without requiring any data outside the stream itself.\n\nIn Git you can store \"action stamp\" in extra extension headers in commit\nobjects (as was already proposed in this thread).\n\nBest,\n--\nJakub Narębski\n"},{"id":"375972","messageId":"20190521010523.GA125903@thyrsus.com","threadId":"51105","inReplyTo":"86v9y4oeiu.fsf@gmail.com","subject":"Re: Finer timestamps and serialization in git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2019-05-21T01:05:23Z","receivedAt":"2019-05-21T01:05:24Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jakub Narebski <jnareb@gmail.com>:\n> Errr... how did you get that the hash of a commit is not portable???\n\nOK. You're telling me that premise was wrong.  Thank you,\naccepted.\n\nI've since had a better idea.  Expect mail soon.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\n\n"}]}