{"thread":{"id":"57549","subject":"Dealing with corporate email recycling","startedAt":"2022-03-12T22:44:53Z","lastAt":"2022-03-18T21:29:34Z","messageCount":28,"participants":["Sean Allred","Junio C Hamano","rsbecker@nexbridge.com","Philip Oakley","Ævar Arnfjörð Bjarmason","brian m. carlson","Peter Krefting"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"451227","messageId":"878rtebxk0.fsf@gmail.com","threadId":"57549","inReplyTo":null,"subject":"Dealing with corporate email recycling","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-03-12T22:38:56Z","receivedAt":"2022-03-12T22:44:53Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"Hi all,\n\nWe are currently replaying a 15-year SVN history into Git -- with\ncontributions from thousands of developers -- and are faced with the\nchallenge of corporate email recycling, departures, re-hires, and name\nchanges causing identity issues.\n\n* Background\n\nAs you know (also to validate my own knowledge/assumptions), a Git\ncommit stores identity as a name and an email.  The only means to\nvalidate this information is via signing; commits are otherwise taken\nat face-value.  This seems pretty core to Git's decentralized design.\nSo to identify who is responsible for a commit, you have only the two\nname+email pairs.\n\nThe problem in a nutshell: names and emails change over time.  The\nsimple cases can be handled by gitmailmap, but there are more\nchallenging cases:\n\n  - A commit author might have had some email <one@corp.net>, but then\n    was able to 'upgrade' to <two@corp.net> after a departure.\n\n  - It's even possible that this departure might 'boomerang' and\n    return to their old job, albeit now with a different email (since\n    they forfeited <two@corp.net> upon departure).\n\nIn effect, the email address (or even email+name pair) used by a given\ndeveloper is not enough to identify that developer.\n\nThis issue is exacerbated by the features of some Git forges (e.g.\nGitHub, GitLab, etc.) that will map an email address to a user\naccount.  In the degenerate case, this would cause the forge to\nattribute the commit incorrectly -- causing communication issues as\ndevelopers use the misattribution to reach out to the wrong people.\n\nAs a baseline, we know the following statements are true:\n\n  1. A person's preferred name can change at any time.\n  2. A person's preferred email can change at any time.\n  3. Neither of these pieces of information are necessarily\n     identifying in a given codebase.\n\n* Current Options\n\nSetting aside the dubious practice of email recycling, how should we\nlook at resolving this confusion in a sensible way?  I see three\ngeneral options that are possible today, each with their drawbacks:\n\n  1. Do nothing.  Leave it to the developer to determine the correct\n     contact information without assistance.\n\n     This doesn't really resolve the confusion, but it is technically\n     an option.\n\n  2. Use gitmailmap(5) functionality to resolve historical emails to\n     primary emails.\n\n     Sadly this doesn't actually solve the email recycling problem.\n     Since one email could be used by multiple developers, there's no\n     way (that I can see) to use a single mailmap file to resolve one\n     of these emails to a single person.\n\n  3. Use and require commit signing -- using some separate system to\n     keep track of who used what public key when\n     (valid-before/-after).\n\n     This is an attractive option and very much fits the 'identity'\n     portion of this problem, but evidently, it's not yet supported\n     by git-fast-import.  This becomes a non-starter for at least all\n     our SVN history we're importing over.  Even a potential concept\n     of providing personal private keys to an impersonal import\n     process also seems less than appealing, but I'm open to the\n     possibility that I'm being too paranoid here.\n\nAt this point, I'd like to ask the community what approaches have\nalready proven successful.  I have a design sketch below, but I would\nnot want to propose introducing a new standard if a standard already\nexists.  I tend to think this is not the first time this problem has\ncome up.\n\n* Bad Proposal: Using Mailmap Over-Time\n\nIn the absence of other options, this leaves us to consider another:\nif the mailmap file is tracked in history, we can know who had which\nemail when.  Taking advantage of this fact right now is a bit\nroundabout, but workable.  Using the `mailmap.blob` config and\npointing to the mailmap version as of a given commit yields behavior\nthat *looks* promising: we can resolve an email in a commit to the\nperson who had that email when the commit was made.  Mentally\nextending this concept to do this automatically in git-log and\nfriends, however, shows that this wholesale removes the value of\nmailmap: the ability to change your name as it appears in history\n*after* that history is created.  More details at [1].\n\nThere are of course legitimate reasons for a developer to change their\nname and desire that name to be used throughout the history.  It would\nseem then that current mailmap functionality is insufficient to create\nan external system that solves the problem.\n\n* Proposal: UUIDs\n\nTo get what we want (i.e., the ability to run `git show HEAD~1`, know\nthat Ada wrote it, and report her current contact information), we\nneed some way of tracking identity over time.  A naive solution could\nbe to extend the mailmap format as recognized by Git:\n\n    $ git cat-file blob HEAD~1:.mailmap\n    A. U. Thor <foo@example.com> [uuid A] <ada@example.com>\n\n    $ git cat-file blob HEAD:.mailmap\n    A. U. Thor <ada@example.com> [uuid A]\n    Roy G. Biv <foo@example.com> [uuid B] <roy@example.com>\n\nNow, when I run `git show HEAD~1`, Git would determine the UUID of the\nemail on the commit using the mailmap version in that tree:\n\n    $ git -c mailmap.blob=HEAD~1:.mailmap check-mailmap --uuid \"<foo@example.com>\"\n    A\n\nThen, we can use that UUID to resolve to the current contact information:\n\n    $ git check-mailmap --uuid=A\n    A. U. Thor <ada@example.com>\n\nMailmap-sensitive commands can use this logic internally -- possibly\nguarded under some new config setting.\n\n** Criticisms\n\nMain criticisms of this proposal that I can think of off-hand:\n\n  1. As far as I know, the mailmap format is pretty well-established.\n     I don't know how additions/extensions to the format will be\n     interpreted by other tools.\n\n     This could be mitigated with a config option or even a magic\n     comment in the file itself noting its format version.\n\n  2. This also assumes that any given email can belong to at most one\n     person at a time.  This is true for us, but may not be generally\n     true.  I don't know if this is a new assumption for mailmap.\n\n     I don't know of any mitigation here.\n\n  3. The current mailmap format of 'one line per pair' potentially\n     opens up the format to ambiguities when resolving an email\n     address -- ambiguities that become more apparent with the\n     introduction of a UUID.  For example:\n\n         A. U. Thor <foo@example.com> [uuid A] <ada@example.com>\n         Ada Thor <ada@example.com> [uuid A] <foo@example.com>\n\n     When evaluating `git check-mailmap --uuid=A`, which line is used?\n     Perhaps this is not a new problem, but it is a disconcerting one.\n\n     This might be resolved by allowing many aliases on a single line\n     and restricting UUIDs to be unique in the file.  Since the\n     introduction of UUIDs would be a change to the format that by\n     definition would need to be opted-into (to provide the UUIDs),\n     this would not be a breaking change in and of itself.\n\nI am additionally unsure of how the following current behavior (from\ngitmailmap(5)) should play into this proposal -- mostly because I\ndon't understand the use-case to carve out a specific author/committer\nname for the mapping:\n\n    > [...] and:\n    >\n    >     Proper Name <proper@email.xx> Commit Name <commit@email.xx>\n    >\n    > which allows mailmap to replace both the name and the email of a\n    > commit matching both the specified commit name and email\n    > address.\n\n* Proposal: Valid-Before/Valid-After\n\nAnother potential idea is to record a transition point in the mailmap\nfile:\n\n    A. U. Thor <foo@example.com> <ada@example.com> valid-before=<timestamp>\n    A. U. Thor <ada@example.com>\n\nThis draws on the valid-before/-after pattern used in commit signing.\nWhile I've signed a few commits in my day, I'm admittedly not versed\non that infrastructure/implementation.  I'd be curious if this idea\nwould have a champion to argue for it.  I personally prefer UUIDs.\n\n** Criticisms\n\nThere are a few things I can identify:\n\n  1. While it's potentially good enough to solve the immediate\n     problem, this doesn't actually establish a tangible identity.\n\n  2. Parsing this as a human being could become difficult if there are\n     a few transition points.\n\n* Summary\n\nAll of this is under the assumption that there is no viable approach\nout there, which I'm not convinced of.  This proposal is only a\nsuggestion of how it *could* be standardized with development if no\nstandard practice currently exists.\n\nSo, two questions remain:\n\n  1. Is there a standard approach out there already that was not\n     discussed above?  (Are there statements I made about those\n     approaches that are not true?)\n\n  2. Assuming there is no suitable standard, is the above proposal\n     worth investing development time, or are there fundamental flaws?\n\nThanks,\nSean Allred\n\n---\n\n[1]: See below.  I removed this from the main body to try and control\nits length.\n\n* Failed Plan: using mailmap over-time\n\n(I'm going to say things in the course of this example that are 'wrong'\nand/or 'working as designed', so bear with me!)\n\nConsider the following.  At some point in the past, <foo@example.com>\nbelonged to Ada.  She wrote the following commit:\n\n    $ git cat-file commit HEAD~1\n    ...\n    author A. U. Thor <foo@example.com> ...\n    committer A. U. Thor <foo@example.com> ...\n\n    $ git cat-file blob HEAD~1:.mailmap\n    A. U. Thor <foo@example.com> <ada@example.com>\n\nSomewhere down the line, Ada has left, <foo@example.com> transferred to\nRoy, and he wrote the following commit:\n\n    $ git cat-file commit HEAD\n    ...\n    author Roy G. Biv <foo@example.com> ...\n    committer Roy G. Biv <foo@example.com> ...\n\n    $ git cat-file blob HEAD:.mailmap\n    A. U. Thor <ada@example.com>\n    Roy G. Biv <foo@example.com> <roy@example.com>\n\nIf we check the mailmap to resolve <foo@example.com> from HEAD~1, we\nget the wrong answer:\n\n    $ git show HEAD~1\n    ...\n    Author: Roy G. Biv <foo@example.com>\n    ...\n\nIt's only when we provide the version of the mailmap file that was\nactive at the time do we get the right answer:\n\n    $ git -c mailmap.blob=HEAD~1:.mailmap show HEAD~1\n    ...\n    Author: A. U. Thor <foo@example.com>\n    ...\n\nSo, if we can instruct git-show and friends to check the mailmap\nversion at the time of the commit, we get what appears to be the\ndesired behavior.  Great, right?\n\nNot so fast.  This loses sight of one of the main purposes of\ngitmailmap.  This idea failed.\n\nYou're at the end now; thanks for reading :-)\n\n--\nSean Allred\n"},{"id":"451229","messageId":"xmqq4k42n2g8.fsf@gitster.g","threadId":"57549","inReplyTo":"878rtebxk0.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-13T00:03:35Z","receivedAt":"2022-03-13T00:03:44Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sean Allred <allred.sean@gmail.com> writes:\n\n> As a baseline, we know the following statements are true:\n>\n>   1. A person's preferred name can change at any time.\n>   2. A person's preferred email can change at any time.\n>   3. Neither of these pieces of information are necessarily\n>      identifying in a given codebase.\n\nAnother thing we know is\n\n    4. People know that old e-mail addresses stay in archives and\n       address books of people, and find it wise to avoid reusing an\n       address somebody else (especially well-known ones) has been\n       using, so that they do not get e-mails from total strangers\n       and having to tell them that the intended recipient does not\n       read the mailbox anymore.\n\n>   1. Do nothing.  Leave it to the developer to determine the correct\n>      contact information without assistance.\n>\n>      This doesn't really resolve the confusion, but it is technically\n>      an option.\n>\n>   2. Use gitmailmap(5) functionality to resolve historical emails to\n>      primary emails.\n>\n>      Sadly this doesn't actually solve the email recycling problem.\n>      Since one email could be used by multiple developers, there's no\n>      way (that I can see) to use a single mailmap file to resolve one\n>      of these emails to a single person.\n\n\nGNU arch (tla) had an interesting idea around this area and used\ncombination of time and e-mail address to identify a person.\none@corp--$date referred to the person who had control of the\naddress on the specified date, where $date can be abbreviated to\n2022 or 202201 to mean 20220101.\n\nThe mailmap allows \"Name e-mail\" or \"e-mail\" to be mapped to\ncanonical \"Name e-mail\", but we should be able to coax \"valid time\nrange\" information encoded in each entry of the .mailmap format,\ni.e. \"if you see 'Name e-mail' between time X and Y, map that to...\".\n\n"},{"id":"451230","messageId":"01cc01d83671$0acd4a20$2067de60$@nexbridge.com","threadId":"57549","inReplyTo":"xmqq4k42n2g8.fsf@gitster.g","subject":"RE: Dealing with corporate email recycling","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-03-13T00:26:49Z","receivedAt":"2022-03-13T00:27:03Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On March 12, 2022 7:04 PM, Junio C Hamano wrote:\n>To: Sean Allred <allred.sean@gmail.com>\n>Cc: git@vger.kernel.org; sallred@epic.com; grmason@epic.com;\n>sconrad@epic.com\n>Subject: Re: Dealing with corporate email recycling\n>\n>Sean Allred <allred.sean@gmail.com> writes:\n>\n>> As a baseline, we know the following statements are true:\n>>\n>>   1. A person's preferred name can change at any time.\n>>   2. A person's preferred email can change at any time.\n>>   3. Neither of these pieces of information are necessarily\n>>      identifying in a given codebase.\n>\n>Another thing we know is\n>\n>    4. People know that old e-mail addresses stay in archives and\n>       address books of people, and find it wise to avoid reusing an\n>       address somebody else (especially well-known ones) has been\n>       using, so that they do not get e-mails from total strangers\n>       and having to tell them that the intended recipient does not\n>       read the mailbox anymore.\n>\n>>   1. Do nothing.  Leave it to the developer to determine the correct\n>>      contact information without assistance.\n>>\n>>      This doesn't really resolve the confusion, but it is technically\n>>      an option.\n>>\n>>   2. Use gitmailmap(5) functionality to resolve historical emails to\n>>      primary emails.\n>>\n>>      Sadly this doesn't actually solve the email recycling problem.\n>>      Since one email could be used by multiple developers, there's no\n>>      way (that I can see) to use a single mailmap file to resolve one\n>>      of these emails to a single person.\n>\n>\n>GNU arch (tla) had an interesting idea around this area and used\ncombination of\n>time and e-mail address to identify a person.\n>one@corp--$date referred to the person who had control of the address on\nthe\n>specified date, where $date can be abbreviated to\n>2022 or 202201 to mean 20220101.\n>\n>The mailmap allows \"Name e-mail\" or \"e-mail\" to be mapped to canonical\n\"Name\n>e-mail\", but we should be able to coax \"valid time range\" information\nencoded in\n>each entry of the .mailmap format, i.e. \"if you see 'Name e-mail' between\ntime X\n>and Y, map that to...\".\n\nIs there anything we could do around the new signature infrastructure\nrelating to this? I am NOT a fan of SSH keys without passphrases, but what\nif we could use the coaxing above and map to SSH expiring keys then stitch\nin signatures (a.k.a. sign the commits) to correspond to the users in the\ngiven timeframe - then destroy the private keys to prevent further signing.\nAfter that the Name/email becomes somewhat irrelevant from an integrity\nstandpoint.\n\n"},{"id":"451240","messageId":"27b5bf20-2eb5-7873-4fdd-875aa699862c@iee.email","threadId":"57549","inReplyTo":"878rtebxk0.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-03-13T12:20:34Z","receivedAt":"2022-03-13T12:20:40Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 12/03/2022 22:38, Sean Allred wrote:\n> Hi all,\n>\n> We are currently replaying a 15-year SVN history into Git -- with\n> contributions from thousands of developers -- and are faced with the\n> challenge of corporate email recycling, departures, re-hires, and name\n> changes causing identity issues.\nNaming is a big issue [1,2].\n\nDo you already have a map of those personal name and email name changes\nthat are causing conflicts, or are you hoping for a way of detecting\nsuch changes? If you already know which names produce conflicts you are\nmore than half way there.\n\nIf you do know of the name conflicts, (e.g. when `John Doe` changed to\n`Jane Doe2`, then acquired `Jane Doe`, before being put back to `Jane\nDoe2`), do you have dates for the change over to map into the commit\ndates (assuming no slop or author/committer date slip). At least with\nthe change-over dates you can apply mapping during the history transfer.\n\nAn alternate option is to simply stick with the fact that history is\nmessy, and use internal corporate knowledge for the few case that cause\nthe major issues. It some point it always gets to be a Gödel Grammar\n(needing another rule).\n\nPhilip\n\n[1]\nhttps://www.kalzumeus.com/2010/06/17/falsehoods-programmers-believe-about-names/\n[2] https://acrl.ala.org/techconnect/post/names-are-hard/\n"},{"id":"451242","messageId":"874k42ar61.fsf@gmail.com","threadId":"57549","inReplyTo":"27b5bf20-2eb5-7873-4fdd-875aa699862c@iee.email","subject":"Re: Dealing with corporate email recycling","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-03-13T13:35:00Z","receivedAt":"2022-03-13T14:00:31Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\nPhilip Oakley <philipoakley@iee.email> writes:\n> Do you already have a map of those personal name and email name changes\n> that are causing conflicts, or are you hoping for a way of detecting\n> such changes? If you already know which names produce conflicts you are\n> more than half way there.\n>\n> If you do know of the name conflicts, (e.g. when `John Doe` changed to\n> `Jane Doe2`, then acquired `Jane Doe`, before being put back to `Jane\n> Doe2`), do you have dates for the change over to map into the commit\n> dates (assuming no slop or author/committer date slip). At least with\n> the change-over dates you can apply mapping during the history transfer.\n\nWhether or not we have maintained such a list in real-time remains to be\nseen, but we have been able to put together such a list using both SVN\nhistory and a variety of other internal data sources.  It's worth noting\nthat, if we do have to use this generated mapping, a mapping of over 10k\nentries is surely to have the odd mistake every now and then.  So it's\nnot *totally* trustworthy (and never could be).\n\n> An alternate option is to simply stick with the fact that history is\n> messy, and use internal corporate knowledge for the few case that cause\n> the major issues. It some point it always gets to be a Gödel Grammar\n> (needing another rule).\n\nThis is definitely on the table, though it seems a shame to not attempt\nto use the information we've been able to compile.\n\n--\nSean Allred\n"},{"id":"451243","messageId":"87zglu9c82.fsf@gmail.com","threadId":"57549","inReplyTo":"01cc01d83671$0acd4a20$2067de60$@nexbridge.com","subject":"Re: Dealing with corporate email recycling","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-03-13T14:01:04Z","receivedAt":"2022-03-13T14:08:34Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\n<rsbecker@nexbridge.com> writes:\n> Is there anything we could do around the new signature infrastructure\n> relating to this? I am NOT a fan of SSH keys without passphrases, but what\n> if we could use the coaxing above and map to SSH expiring keys then stitch\n> in signatures (a.k.a. sign the commits) to correspond to the users in the\n> given timeframe - then destroy the private keys to prevent further signing.\n> After that the Name/email becomes somewhat irrelevant from an integrity\n> standpoint.\n\nIs this really possible?  Is it really as straightforward as splicing in\nsome text into the commit message to the effect of 'this commit is\nsigned' along with some signature artifact calculated pre-signing?\n\nThough I'll note I *think* this would only solve the problem for the\ncommitter field -- it's my current understanding that a commit can only\nbe signed by one signature.  (I have heard of systems that generate a\nnew key that is then signed by multiple signatures, then signing with\nthat new key -- but even if this is possible, it seems pretty involved\nfor such a common workflow.  This level of coordination might not be\npossible for us -- especially given the merge workflows we've needed to\ncreate to accommodate our current release process.)\n\n--\nSean Allred\n"},{"id":"451244","messageId":"01f201d836e5$89247c30$9b6d7490$@nexbridge.com","threadId":"57549","inReplyTo":"87zglu9c82.fsf@gmail.com","subject":"RE: Dealing with corporate email recycling","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-03-13T14:20:42Z","receivedAt":"2022-03-13T14:20:58Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On March 13, 2022 10:01 AM, Sean Allred wrote:\n><rsbecker@nexbridge.com> writes:\n>> Is there anything we could do around the new signature infrastructure\n>> relating to this? I am NOT a fan of SSH keys without passphrases, but\n>> what if we could use the coaxing above and map to SSH expiring keys\n>> then stitch in signatures (a.k.a. sign the commits) to correspond to\n>> the users in the given timeframe - then destroy the private keys to\nprevent\n>further signing.\n>> After that the Name/email becomes somewhat irrelevant from an\n>> integrity standpoint.\n>\n>Is this really possible?  Is it really as straightforward as splicing in\nsome text into the\n>commit message to the effect of 'this commit is signed' along with some\nsignature\n>artifact calculated pre-signing?\n>\n>Though I'll note I *think* this would only solve the problem for the\ncommitter field\n>-- it's my current understanding that a commit can only be signed by one\n>signature.  (I have heard of systems that generate a new key that is then\nsigned by\n>multiple signatures, then signing with that new key -- but even if this is\npossible, it\n>seems pretty involved for such a common workflow.  This level of\ncoordination\n>might not be possible for us -- especially given the merge workflows we've\n>needed to create to accommodate our current release process.)\n\n(I am a little nervous about this advice, hoping others will chime in and\ncorrect anything wrong here)\n\nWhile this will change the commit hashes, AFAIK, the other metadata is\npreserved, including date, author, and committer. Set up the specific\nkeys/settings in ssh-agent and the user.signingKey value, then:\n\ngit filter-branch --commit-filter 'git commit-tree -S \"$@\";'\n<FROM-COMMIT>..<TO-COMMIT>\n\nOthers might have a better way of doing this or may tell me this will not\nwork. Test this before you do it. I have not done this operation before. You\ndo need to start from the oldest commit going forward otherwise I think that\nfilter-branch will (should!) invalidate child commits. I suspect this is\ngoing to be a rather lengthy script to build and run.\n\n"},{"id":"451245","messageId":"87v8whap0b.fsf@gmail.com","threadId":"57549","inReplyTo":"01f201d836e5$89247c30$9b6d7490$@nexbridge.com","subject":"Re: Dealing with corporate email recycling","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-03-13T14:41:18Z","receivedAt":"2022-03-13T14:47:07Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\n<rsbecker@nexbridge.com> writes:\n> (I am a little nervous about this advice, hoping others will chime in and\n> correct anything wrong here)\n>\n> While this will change the commit hashes, AFAIK, the other metadata is\n> preserved, including date, author, and committer. Set up the specific\n> keys/settings in ssh-agent and the user.signingKey value, then:\n>\n> git filter-branch --commit-filter 'git commit-tree -S \"$@\";'\n> <FROM-COMMIT>..<TO-COMMIT>\n>\n> Others might have a better way of doing this or may tell me this will not\n> work. Test this before you do it. I have not done this operation before. You\n> do need to start from the oldest commit going forward otherwise I think that\n> filter-branch will (should!) invalidate child commits. I suspect this is\n> going to be a rather lengthy script to build and run.\n\nGiven the size of our history (several orders of magnitude larger than\nlinux.git), using git-filter-branch after the fact is certainly not\nideal.  The replay already takes a week to run (we're IO-bound).  We'd\nrather want to extend git-fast-import to allow signing commits instead\n-- which comes back to our shared 'nervousness' about this approach in\ngeneral: I don't know that Git should endorse this as a standard option.\n\nBut yes -- hoping others can chime in with more thoughts :-)\n\n--\nSean Allred\n"},{"id":"451246","messageId":"01f301d836eb$5c7a6810$156f3830$@nexbridge.com","threadId":"57549","inReplyTo":"87v8whap0b.fsf@gmail.com","subject":"RE: Dealing with corporate email recycling","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-03-13T15:02:24Z","receivedAt":"2022-03-13T15:02:37Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On March 13, 2022 10:41 AM, Sean Allred wrote:\n><rsbecker@nexbridge.com> writes:\n>> (I am a little nervous about this advice, hoping others will chime in\n>> and correct anything wrong here)\n>>\n>> While this will change the commit hashes, AFAIK, the other metadata is\n>> preserved, including date, author, and committer. Set up the specific\n>> keys/settings in ssh-agent and the user.signingKey value, then:\n>>\n>> git filter-branch --commit-filter 'git commit-tree -S \"$@\";'\n>> <FROM-COMMIT>..<TO-COMMIT>\n>>\n>> Others might have a better way of doing this or may tell me this will\n>> not work. Test this before you do it. I have not done this operation\n>> before. You do need to start from the oldest commit going forward\n>> otherwise I think that filter-branch will (should!) invalidate child\n>> commits. I suspect this is going to be a rather lengthy script to build and run.\n>\n>Given the size of our history (several orders of magnitude larger than linux.git),\n>using git-filter-branch after the fact is certainly not ideal.  The replay already takes\n>a week to run (we're IO-bound).  We'd rather want to extend git-fast-import to\n>allow signing commits instead\n>-- which comes back to our shared 'nervousness' about this approach in\n>general: I don't know that Git should endorse this as a standard option.\n>\n>But yes -- hoping others can chime in with more thoughts :-)\n\nI have another reluctant suggestion, but it depends on your industry, regulations, and other factors. In some sectors, there is a requirement to keep only some period of time worth of history. In fact, in some settings, keeping user identifying information beyond, say 7 years, actually is problematic. Pruning your history may be not only an option but required. An alternative is to use filter-branch to essentially tokenize the identities of past authors and keep those in a electronic vault somewhere. I have customers who are interpreting GDPR-like rules just such as situation, where employees gone 7 years ago and cannot be retained, by name, in the repos. I am not personally happy about that, because my own repo-OCD demands that I know exactly who did what until the end of time, but according to them, it actually violates the local regulations. I'm sure you have had conversations with lawyers, yes? ☹\n\n"},{"id":"451247","messageId":"87r175amw2.fsf@gmail.com","threadId":"57549","inReplyTo":"01f301d836eb$5c7a6810$156f3830$@nexbridge.com","subject":"Re: Dealing with corporate email recycling","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-03-13T15:21:04Z","receivedAt":"2022-03-13T15:32:56Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\n<rsbecker@nexbridge.com> writes:\n> I have another reluctant suggestion, but it depends on your industry,\n> regulations, and other factors. In some sectors, there is a\n> requirement to keep only some period of time worth of history. In\n> fact, in some settings, keeping user identifying information beyond,\n> say 7 years, actually is problematic. Pruning your history may be not\n> only an option but required. An alternative is to use filter-branch to\n> essentially tokenize the identities of past authors and keep those in\n> a electronic vault somewhere. I have customers who are interpreting\n> GDPR-like rules just such as situation, where employees gone 7 years\n> ago and cannot be retained, by name, in the repos. I am not personally\n> happy about that, because my own repo-OCD demands that I know exactly\n> who did what until the end of time, but according to them, it actually\n> violates the local regulations. I'm sure you have had conversations\n> with lawyers, yes? ☹\n\nI don't believe we've involved our legal team here (I'll follow up with\nthem internally), but that might be a spin-off discussion for folks who\nknow they're affected.  It would seem that the design of Git makes\npurging history on an ongoing basis problematic -- you would always have\nat least one unresolvable reference to a parent commit.  If this is a\nreal requirement from GDPR-like laws, either 'reasonable' VCS metadata\nneeds to be a specific carve-out in those laws -- but who the heck knows\nwhat is 'reasonable' -- or as a project, Git needs to have an answer to\nthis situation and an ability to truncate history without otherwise\naltering it.\n\nIt's also worth noting that even in the last five years, at our scale,\nwe've definitely run into the email-recycling problem already.\n\nBeing based in the U.S. and not having seen pitchforks about this yet,\nI'd like to assume for the purpose of this discussion that we're keeping\nall our history.\n\nI think if the topic of legal implications of keeping history in\nperpetuity is valuable to continue, we should spin it off into a\nseparate thread.  Personally I'm not seeing what we (Git) could\nrealistically do about it other than provide recommendations and paths\nforward -- which might require considerable development.\n\n--\nSean Allred\n"},{"id":"451248","messageId":"220313.86k0cxg7m2.gmgdl@evledraar.gmail.com","threadId":"57549","inReplyTo":"878rtebxk0.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2022-03-13T15:51:05Z","receivedAt":"2022-03-13T16:06:22Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Sat, Mar 12 2022, Sean Allred wrote:\n\n> We are currently replaying a 15-year SVN history into Git -- with\n> contributions from thousands of developers -- and are faced with the\n> challenge of corporate email recycling, departures, re-hires, and name\n> changes causing identity issues.\n>\n> * Background\n>\n> As you know (also to validate my own knowledge/assumptions), a Git\n> commit stores identity as a name and an email.  The only means to\n> validate this information is via signing; commits are otherwise taken\n> at face-value.  This seems pretty core to Git's decentralized design.\n> So to identify who is responsible for a commit, you have only the two\n> name+email pairs.\n>\n> The problem in a nutshell: names and emails change over time.  The\n> simple cases can be handled by gitmailmap, but there are more\n> challenging cases:\n>\n>   - A commit author might have had some email <one@corp.net>, but then\n>     was able to 'upgrade' to <two@corp.net> after a departure.\n>\n>   - It's even possible that this departure might 'boomerang' and\n>     return to their old job, albeit now with a different email (since\n>     they forfeited <two@corp.net> upon departure).\n> [...skip a bunch of details...]\n> You're at the end now; thanks for reading :-)\n\nAside from technical solutions and twists on mailmap, you haven't\n*really* described what practical problem you're facing here.\n\n> Somewhere down the line, Ada has left, <foo@example.com> transferred to\n> Roy, and he wrote the following commit:\n\nI.e. this, sure, that can happen, but what's the negative effect of that\nin practice?\n\nI've been involved in similar migrations in the past, and the primary\nway to deal with it was to mostly ignore it, especially in a corporate\nsetting.\n\nI.e. sure, you'll have some edge cases here and there, but the value of\nknowing who exactly authored something tend to be proportional to how\nrecent the commit is.\n\nIf someone wrote something 10 years ago they're probably not even\nworking there anymore, or if they are will long since have forgotten\nwhat they need to know to answer any specific questions etc.\n\nThe only people who tend to look at it are developers using \"git blame\"\nor something, and usually humans are smart enough to spot that even if\nit's foo@example.com they were expecting Roy, not Ada, or the other way\naround.\n\nSide note: To the extent that I've had to deal with this (in a corporate\nsetting) I found myself wanting git to have the exact opposite,\ni.e. some feature where we'd just hide the author for anything any work\nthat's >5 years old or whatever.\n\nNot for any privacy reason, but just because some UI's wouldn't really\ncommunicate (in a way that people actually noticed) that the relevant\nwork was ancient, and someone who'd since long-moved-on would get\noccasional interruptions due to ancient code they wrote but weren't\nequipped to currently maintain.\n\nOr similarly, to have anything >N years old \"git blame\" to the team\ncurrently maintaining that thing, not to the person.\n\nBut I digress.\n\nHaving said that I think if you do need such a back-annotated history\nyou should look into \"git notes\" and/or \"git replace\". I.e. you could\nhave some lookup system maintain a mapping from OIDs to current IDs.\n\nI've implemented a system like that in the past (in a MySQL table, but\nwhatever). I'd think this use-case of perfectly annotated old history is\nprobably obscure enough that that's the primary thing we should steer\npeople towards...\n\n>   1. As far as I know, the mailmap format is pretty well-established.\n>      I don't know how additions/extensions to the format will be\n>      interpreted by other tools.\n\nIt's perfectly OK to change parts of that format in\nbackwards-\"incompatible\" ways, i.e. there's enough leeway in the\nexisting format definition and in-the-wild readers to have new readers\npick up new information that old readers will ignore.\n\nI.e. we simply ignore things we can't map now, so one way to do it is to\nstart with something that produces an invalid (but harmless) mapping to\ncurrent readers, another is to borrow a trick from \"/etc/sudoers\" and\n(ab)use the comment syntax.\n"},{"id":"451252","messageId":"Yi4oO+ifSK8OH0Mt@camp.crustytoothpaste.net","threadId":"57549","inReplyTo":"878rtebxk0.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"brian m. carlson","fromEmail":"sandals@crustytoothpaste.net","sentAt":"2022-03-13T17:22:03Z","receivedAt":"2022-03-13T17:22:19Z","isPatch":false,"sender":{"key":"sandals@crustytoothpaste.net","avatar":"https://avatars.githubusercontent.com/u/497054?v=4"},"body":"On 2022-03-12 at 22:38:56, Sean Allred wrote:\n> * Proposal: UUIDs\n> \n> To get what we want (i.e., the ability to run `git show HEAD~1`, know\n> that Ada wrote it, and report her current contact information), we\n> need some way of tracking identity over time.  A naive solution could\n> be to extend the mailmap format as recognized by Git:\n> \n>     $ git cat-file blob HEAD~1:.mailmap\n>     A. U. Thor <foo@example.com> [uuid A] <ada@example.com>\n> \n>     $ git cat-file blob HEAD:.mailmap\n>     A. U. Thor <ada@example.com> [uuid A]\n>     Roy G. Biv <foo@example.com> [uuid B] <roy@example.com>\n> \n> Now, when I run `git show HEAD~1`, Git would determine the UUID of the\n> email on the commit using the mailmap version in that tree:\n> \n>     $ git -c mailmap.blob=HEAD~1:.mailmap check-mailmap --uuid \"<foo@example.com>\"\n>     A\n> \n> Then, we can use that UUID to resolve to the current contact information:\n> \n>     $ git check-mailmap --uuid=A\n>     A. U. Thor <ada@example.com>\n> \n> Mailmap-sensitive commands can use this logic internally -- possibly\n> guarded under some new config setting.\n\nIt's my intention to implement an approach where people's emails are\nidentified by a key fingerprint of some sort and then converted into the\nproper email address by a mailmap that lives outside of the main\nhistory.  That is, my email address might be\nba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad@ssh-sha256.ns.git-scm.com,\nand then we have a mailmap that converts between the two.  If you wanted\nto have a UUID-based one, you could do\n77c747a3-1599-4c8c-9569-f729c17632e6@uuid.ns.git-scm.com (assuming that\nnamespace were registered).\n\nThe benefit to the key part is that you can essentially prove that you\nare who you say you are.  A UUID doesn't have the possibility.\n\nThis was discussed briefly at some sort of contributor summit we had at\nsome point, but I've been busy and haven't gotten to it yet.  It is on\nmy list of projects, however.\n-- \nbrian m. carlson (he/him or they/them)\nToronto, Ontario, CA\n"},{"id":"451255","messageId":"020501d83703$2f8785f0$8e9691d0$@nexbridge.com","threadId":"57549","inReplyTo":"Yi4oO+ifSK8OH0Mt@camp.crustytoothpaste.net","subject":"RE: Dealing with corporate email recycling","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-03-13T17:52:57Z","receivedAt":"2022-03-13T17:53:14Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On March 13, 2022 1:22 PM, brian m. carlson wrote:\n>On 2022-03-12 at 22:38:56, Sean Allred wrote:\n>> * Proposal: UUIDs\n>>\n>> To get what we want (i.e., the ability to run `git show HEAD~1`, know\n>> that Ada wrote it, and report her current contact information), we\n>> need some way of tracking identity over time.  A naive solution could\n>> be to extend the mailmap format as recognized by Git:\n>>\n>>     $ git cat-file blob HEAD~1:.mailmap\n>>     A. U. Thor <foo@example.com> [uuid A] <ada@example.com>\n>>\n>>     $ git cat-file blob HEAD:.mailmap\n>>     A. U. Thor <ada@example.com> [uuid A]\n>>     Roy G. Biv <foo@example.com> [uuid B] <roy@example.com>\n>>\n>> Now, when I run `git show HEAD~1`, Git would determine the UUID of the\n>> email on the commit using the mailmap version in that tree:\n>>\n>>     $ git -c mailmap.blob=HEAD~1:.mailmap check-mailmap --uuid\n>\"<foo@example.com>\"\n>>     A\n>>\n>> Then, we can use that UUID to resolve to the current contact information:\n>>\n>>     $ git check-mailmap --uuid=A\n>>     A. U. Thor <ada@example.com>\n>>\n>> Mailmap-sensitive commands can use this logic internally -- possibly\n>> guarded under some new config setting.\n>\n>It's my intention to implement an approach where people's emails are identified\n>by a key fingerprint of some sort and then converted into the proper email\n>address by a mailmap that lives outside of the main history.  That is, my email\n>address might be\n>ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad@ssh-\n>sha256.ns.git-scm.com,\n>and then we have a mailmap that converts between the two.  If you wanted to\n>have a UUID-based one, you could do 77c747a3-1599-4c8c-9569-\n>f729c17632e6@uuid.ns.git-scm.com (assuming that namespace were registered).\n>\n>The benefit to the key part is that you can essentially prove that you are who you\n>say you are.  A UUID doesn't have the possibility.\n>\n>This was discussed briefly at some sort of contributor summit we had at some\n>point, but I've been busy and haven't gotten to it yet.  It is on my list of projects,\n>however.\n\nThis could require a global and security hardened tokenization or signing approach. Email fingerprints from one organization would have to be able to move to another organization easily - potentially as part of the git repo's metadata. I would not use the same key as is used for signing fingerprints (mostly out of paranoia), but this is conceptually similar to the public side of a key-pair. One would have to have access to the private key in order to be a committer/author. Unfortunately, as it stands today, that may be easily spoofed (--committer, --author), so that part of the code would have to change with safeguards on what can be supplied - something I think would be welcome. Keeping with a distributed philosophy is probably essential. Just my take on it.\n\n"},{"id":"451260","messageId":"020901d83713$2b446ac0$81cd4040$@nexbridge.com","threadId":"57549","inReplyTo":"020501d83703$2f8785f0$8e9691d0$@nexbridge.com","subject":"RE: Dealing with corporate email recycling","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-03-13T19:47:22Z","receivedAt":"2022-03-13T19:47:53Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On March 13, 2022 1:53 PM, I wrote:\n>To: 'brian m. carlson' <sandals@crustytoothpaste.net>; 'Sean Allred'\n><allred.sean@gmail.com>\n>Cc: git@vger.kernel.org; sallred@epic.com; grmason@epic.com;\n>sconrad@epic.com\n>Subject: RE: Dealing with corporate email recycling\n>\n>On March 13, 2022 1:22 PM, brian m. carlson wrote:\n>>On 2022-03-12 at 22:38:56, Sean Allred wrote:\n>>> * Proposal: UUIDs\n>>>\n>>> To get what we want (i.e., the ability to run `git show HEAD~1`, know\n>>> that Ada wrote it, and report her current contact information), we\n>>> need some way of tracking identity over time.  A naive solution could\n>>> be to extend the mailmap format as recognized by Git:\n>>>\n>>>     $ git cat-file blob HEAD~1:.mailmap\n>>>     A. U. Thor <foo@example.com> [uuid A] <ada@example.com>\n>>>\n>>>     $ git cat-file blob HEAD:.mailmap\n>>>     A. U. Thor <ada@example.com> [uuid A]\n>>>     Roy G. Biv <foo@example.com> [uuid B] <roy@example.com>\n>>>\n>>> Now, when I run `git show HEAD~1`, Git would determine the UUID of\n>>> the email on the commit using the mailmap version in that tree:\n>>>\n>>>     $ git -c mailmap.blob=HEAD~1:.mailmap check-mailmap --uuid\n>>\"<foo@example.com>\"\n>>>     A\n>>>\n>>> Then, we can use that UUID to resolve to the current contact information:\n>>>\n>>>     $ git check-mailmap --uuid=A\n>>>     A. U. Thor <ada@example.com>\n>>>\n>>> Mailmap-sensitive commands can use this logic internally -- possibly\n>>> guarded under some new config setting.\n>>\n>>It's my intention to implement an approach where people's emails are\n>>identified by a key fingerprint of some sort and then converted into\n>>the proper email address by a mailmap that lives outside of the main\n>>history.  That is, my email address might be\n>>ba7816bf8f01cfea414140de5dae2223b00361a396177a9cb410ff61f20015ad@ssh-\n>>sha256.ns.git-scm.com,\n>>and then we have a mailmap that converts between the two.  If you\n>>wanted to have a UUID-based one, you could do 77c747a3-1599-4c8c-9569-\n>>f729c17632e6@uuid.ns.git-scm.com (assuming that namespace were registered).\n>>\n>>The benefit to the key part is that you can essentially prove that you\n>>are who you say you are.  A UUID doesn't have the possibility.\n>>\n>>This was discussed briefly at some sort of contributor summit we had at\n>>some point, but I've been busy and haven't gotten to it yet.  It is on\n>>my list of projects, however.\n>\n>This could require a global and security hardened tokenization or signing approach.\n>Email fingerprints from one organization would have to be able to move to\n>another organization easily - potentially as part of the git repo's metadata. I would\n>not use the same key as is used for signing fingerprints (mostly out of paranoia),\n>but this is conceptually similar to the public side of a key-pair. One would have to\n>have access to the private key in order to be a committer/author. Unfortunately,\n>as it stands today, that may be easily spoofed (--committer, --author), so that part\n>of the code would have to change with safeguards on what can be supplied -\n>something I think would be welcome. Keeping with a distributed philosophy is\n>probably essential. Just my take on it.\n\nWhat about abstracting this into a map-email or map-identity hook of some kind? So, whenever there is a need to write an identity (committer, author, signed-off-by, etc.). That way, anyone who wants to, can implement whatever policy they want for replacing emails with some other value in the repo, and back again. It might be good to optimize it so that the hook is only invoked once per identity per request so that git log does not become insanely expensive.\n\nSomething like map-identity from <internal-value>  and map-identity to <external-value>, for example:\n\nmap-identity from \"Randall S. Becker <rsbecker@nexbridge.com>\"      > A056AAB2123\n\nAnd\n\nmap-identity to A056AAB2123      >  Randall S. Becker <rsbecker@nexbridge.com>\n\nAgain, just a notion.\n\n"},{"id":"451267","messageId":"f6ecca05-b669-0e36-302f-a6113571ac12@iee.email","threadId":"57549","inReplyTo":"87r175amw2.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-03-13T19:57:32Z","receivedAt":"2022-03-13T19:57:39Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 13/03/2022 15:21, Sean Allred wrote:\n> <rsbecker@nexbridge.com> writes:\n>> I have another reluctant suggestion, but it depends on your industry,\n>> regulations, and other factors. In some sectors, there is a\n>> requirement to keep only some period of time worth of history. In\n>> fact, in some settings, keeping user identifying information beyond,\n>> say 7 years, actually is problematic. Pruning your history may be not\n>> only an option but required. An alternative is to use filter-branch to\n>> essentially tokenize the identities of past authors and keep those in\n>> a electronic vault somewhere. I have customers who are interpreting\n>> GDPR-like rules just such as situation, where employees gone 7 years\n>> ago and cannot be retained, by name, in the repos. I am not personally\n>> happy about that, because my own repo-OCD demands that I know exactly\n>> who did what until the end of time, but according to them, it actually\n>> violates the local regulations. I'm sure you have had conversations\n>> with lawyers, yes? ☹\n> I don't believe we've involved our legal team here (I'll follow up with\n> them internally), but that might be a spin-off discussion for folks who\n> know they're affected.  It would seem that the design of Git makes\n> purging history on an ongoing basis problematic -- you would always have\n> at least one unresolvable reference to a parent commit.  If this is a\n> real requirement from GDPR-like laws, either 'reasonable' VCS metadata\n> needs to be a specific carve-out in those laws -- but who the heck knows\n> what is 'reasonable' -- or as a project, Git needs to have an answer to\n> this situation and an ability to truncate history without otherwise\n> altering it.\n>\n> It's also worth noting that even in the last five years, at our scale,\n> we've definitely run into the email-recycling problem already.\n>\n> Being based in the U.S. and not having seen pitchforks about this yet,\n> I'd like to assume for the purpose of this discussion that we're keeping\n> all our history.\n>\n> I think if the topic of legal implications of keeping history in\n> perpetuity is valuable to continue, we should spin it off into a\n> separate thread.  Personally I'm not seeing what we (Git) could\n> realistically do about it other than provide recommendations and paths\n> forward -- which might require considerable development.\n>\n>\nThe GDPR isn't as onerous as some suggest, as it isn't a set of black\nand white rules, rather in cases like these you need to have a real\nstrong reason for why data is retained etc, such as being part of the\nverification and validation of the commit data. There have been various\ndiscussions around this in many of the technical journals.\n\nIt maybe that your internal Git version could disable the particular\n`format` option ('%ae'?) for the original name, so only the designated\n('redacted') mailmap entry is shown to casual users (assumes the repo is\ninside the corporate firewall). This would avoid invalidating the repos\nvalidation capability, while meeting the needs of GDPR type regulations.\n\nIn the same vein, a local Git version could, being open source, add\nallowances for your extra mailmap entry details, such as adding a post\nfix \" % <approxidate>\" limits for the use of the particular name/email\ncombo to allow date ranges to emerge.\n\nI noted that all the .mailmap examples in the man page have \">\" as the\nfinal character, but I haven't looked to see if the code always requires\nthat the last element of the entry is an <email> address, or whether it\ncurrently barfs on extra elements.\n\n--\nPhilip\n"},{"id":"451272","messageId":"87mthta3jq.fsf@gmail.com","threadId":"57549","inReplyTo":"020901d83713$2b446ac0$81cd4040$@nexbridge.com","subject":"Re: Dealing with corporate email recycling","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-03-13T22:23:36Z","receivedAt":"2022-03-13T22:30:41Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\n<rsbecker@nexbridge.com> writes:\n> What about abstracting this into a map-email or map-identity hook of\n> some kind? So, whenever there is a need to write an identity\n> (committer, author, signed-off-by, etc.). That way, anyone who wants\n> to, can implement whatever policy they want for replacing emails with\n> some other value in the repo, and back again.\n\nThis is an interesting idea, but I'm afraid it might be difficult for\nforges to implement support for as opposed to something built-in.  If\nthat's not as difficult as I might think it is, then perhaps this is a\nviable option once fleshed out.\n\n> It might be good to optimize it so that the hook is only invoked once\n> per identity per request so that git log does not become insanely\n> expensive.\n\nI'll add that Windows (and our particular environment) makes this\ntroublesome as well.  Currently we see a base cost of 300ms for starting\nup a process.  Given how many identities git-log and friends would be\nchugging through, any hook would need to be capable of staying open --\nfeeding identities through stdin.  I'm not sure run_hooks supports that\nright now.\n\n\n--\nSean Allred\n"},{"id":"451273","messageId":"87ilsha2b7.fsf@gmail.com","threadId":"57549","inReplyTo":"f6ecca05-b669-0e36-302f-a6113571ac12@iee.email","subject":"Re: Dealing with corporate email recycling","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-03-13T22:40:47Z","receivedAt":"2022-03-13T22:58:13Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\nPhilip Oakley <philipoakley@iee.email> writes:\n> The GDPR isn't as onerous as some suggest, as it isn't a set of black\n> and white rules, rather in cases like these you need to have a real\n> strong reason for why data is retained etc, such as being part of the\n> verification and validation of the commit data. There have been various\n> discussions around this in many of the technical journals.\n\nThat's good to hear that this has already been discussed in the\ncommunity (though I'm hardly surprised now that you mention it -- I'm\nsure it was and remains a hot topic!).\n\n> It maybe that your internal Git version could disable the particular\n> `format` option ('%ae'?) for the original name, so only the designated\n> ('redacted') mailmap entry is shown to casual users (assumes the repo is\n> inside the corporate firewall). This would avoid invalidating the repos\n> validation capability, while meeting the needs of GDPR type regulations.\n\nI do want to note that at present we're not primarily concerned with\nGDPR, but I am following up on that internally to see if there are any\nconsiderations we need to make.  This is certainly an interesting tactic\nfor repositories that are hosted in GDPR-effective states, though.\n\n> In the same vein, a local Git version could, being open source, add\n> allowances for your extra mailmap entry details, such as adding a post\n> fix \" % <approxidate>\" limits for the use of the particular name/email\n> combo to allow date ranges to emerge.\n\nI'd prefer the ability to agree on a pattern and merge support for it\nupstream.  This way, forges can pick up support, too.  Bonus points if\nthe forge doesn't necessarily have to do more work than it already does.\n\nYour \" % <approxidate>\" suggestion sounds a lot like the 'Valid-Before/\nValid-After' proposal from my original post in this thread (admittedly\nnot my idea).  Is there a compelling reason to use this approach over\nUUIDs?  I ask not to suggest there isn't a compelling reason, but mostly\nto make sure we consider the best arguments (and drawbacks) for any/all\napproaches.\n\n> I noted that all the .mailmap examples in the man page have \">\" as the\n> final character, but I haven't looked to see if the code always requires\n> that the last element of the entry is an <email> address, or whether it\n> currently barfs on extra elements.\n\nFYI mailmap does support comment syntax (starting with # through \\n).\nIt's worth noting that Ævar suggested earlier that perhaps we could\n\"(ab)use the comment syntax\".  I tend to prefer their other approach,\nthough:\n\n    > I.e. we simply ignore things we can't map now, so one way to do it\n    > is to start with something that produces an invalid (but harmless)\n    > mapping to current readers, [...]\n\nrather than use magic comments :-) Adapting to your suggestion, this\nmight look like the following:\n\n    A. U. Thor <foo@example.com> <ada.example.com> <[ approxidate ]>\n\nWould be curious to know what other mailmap readers exist and how they\nwould react to this.\n\n--\nSean Allred\n"},{"id":"451274","messageId":"xmqqtuc1tpdj.fsf@gitster.g","threadId":"57549","inReplyTo":"87ilsha2b7.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-13T23:16:24Z","receivedAt":"2022-03-13T23:16:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Sean Allred <allred.sean@gmail.com> writes:\n\n> rather than use magic comments :-) Adapting to your suggestion, this\n> might look like the following:\n>\n>     A. U. Thor <foo@example.com> <ada.example.com> <[ approxidate ]>\n\nYou'd probably want a timerange (valid-from and valid-to), instead\nof one single timestamp?\n\nBecause at least three valid forms of mailmap entries should be\nunderstood by the current generation of mailmap readers, i.e.\n\n    Human Readable Name <e-mail@add.re.ss>\n    Right Name <right@add.re.ss> <wrong@add.re.ss>\n    Right Name <right@add.re.ss> Wrong Name <wrong@add.re.ss>\n\nthe extended entry format to record the validity timerange should\nbe chosen to cause parsers that are prepared to take these three\nkinds of lines to barf and ignore.\n"},{"id":"451275","messageId":"021401d83731$62813630$2783a290$@nexbridge.com","threadId":"57549","inReplyTo":"xmqqtuc1tpdj.fsf@gitster.g","subject":"RE: Dealing with corporate email recycling","fromName":"","fromEmail":"rsbecker@nexbridge.com","sentAt":"2022-03-13T23:23:38Z","receivedAt":"2022-03-13T23:23:58Z","isPatch":false,"sender":{"key":"randall.becker@nexbridge.ca","avatar":"https://avatars.githubusercontent.com/u/28956764?v=4"},"body":"On March 13, 2022 7:16 PM, Junio C Hamano wrote:\n>To: Sean Allred <allred.sean@gmail.com>\n>Sean Allred <allred.sean@gmail.com> writes:\n>\n>> rather than use magic comments :-) Adapting to your suggestion, this\n>> might look like the following:\n>>\n>>     A. U. Thor <foo@example.com> <ada.example.com> <[ approxidate ]>\n>\n>You'd probably want a timerange (valid-from and valid-to), instead of one\nsingle\n>timestamp?\n>\n>Because at least three valid forms of mailmap entries should be understood\nby the\n>current generation of mailmap readers, i.e.\n>\n>    Human Readable Name <e-mail@add.re.ss>\n>    Right Name <right@add.re.ss> <wrong@add.re.ss>\n>    Right Name <right@add.re.ss> Wrong Name <wrong@add.re.ss>\n>\n>the extended entry format to record the validity timerange should be chosen\nto\n>cause parsers that are prepared to take these three kinds of lines to barf\nand\n>ignore.\n\nCould we not use SSH's ssh-keygen -V for this purpose when establishing\npersistent identities independent of user/email? We already do this for\nsigned commits.\n\n"},{"id":"451277","messageId":"xmqqlexdtmgd.fsf@gitster.g","threadId":"57549","inReplyTo":"021401d83731$62813630$2783a290$@nexbridge.com","subject":"Re: Dealing with corporate email recycling","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-14T00:19:30Z","receivedAt":"2022-03-14T00:19:37Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"<rsbecker@nexbridge.com> writes:\n\n> Could we not use SSH's ssh-keygen -V for this purpose when establishing\n> persistent identities independent of user/email? We already do this for\n> signed commits.\n\nFingerprint of cryptographic key would be easy to use as an\nidentity, for which the person who claims ownership can easily\nproduce proof of ownership.  Various other \"identitying strings\"\nlike human readable name and e-mail addresses from different\nvalidity periods can be all tied to such an identity.  Taking key\nrevocation into account, keys from different validity period may\nhave to be tied together in a same way.  \"The person who used to\nsign the commits with key A and the person who signs the commits\nwith key B are the same, and in real life, they are known as\nA. U. Thor\"\n\nBut proving that such a mapping is in a meaningful way is much\nharder, I would imagine, but perhaps addresses and human readable\nnames do not matter as much.  Or continuity of identity, for that\nmatter.  I dunno.\n\n\n\n"},{"id":"451305","messageId":"697d8717-bd3f-0871-d5b3-e6303c4ed726@iee.email","threadId":"57549","inReplyTo":"xmqqtuc1tpdj.fsf@gitster.g","subject":"Re: Dealing with corporate email recycling","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-03-14T11:56:17Z","receivedAt":"2022-03-14T11:57:31Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 13/03/2022 23:16, Junio C Hamano wrote:\n> Sean Allred <allred.sean@gmail.com> writes:\n>\n>> rather than use magic comments :-) Adapting to your suggestion, this\n>> might look like the following:\n>>\n>>     A. U. Thor <foo@example.com> <ada.example.com> <[ approxidate ]>\n> You'd probably want a timerange (valid-from and valid-to), instead\n> of one single timestamp?\nI'm not so sure that the date range approach won't bring it's own\nproblems. What happens outside the date range? i.e. Do we then have\nthree identities: Before, During, and After, with only 'During' being\ndefined?\n\nI more see a single date being used as a termination point for an\nexisting email sequence that defines a retrospective end point for the\nmapping of the old email addresses to a single person. Future emails for\nthe same mailbox will be for a different 'current' person. This would\nmatch the single linked list commit history view using the chronology\nheuristic.\n\nThe key here being to have a final identity system in place so that you\ncan uniquely identify the old John Doe, from the newer John Doe`s at the\nrelevant time point in the mailmap.\n\n>\n> Because at least three valid forms of mailmap entries should be\n> understood by the current generation of mailmap readers, i.e.\n>\n>     Human Readable Name <e-mail@add.re.ss>\n>     Right Name <right@add.re.ss> <wrong@add.re.ss>\n>     Right Name <right@add.re.ss> Wrong Name <wrong@add.re.ss>\n>\n> the extended entry format to record the validity timerange should\n> be chosen to cause parsers that are prepared to take these three\n> kinds of lines to barf and ignore.\nThe presence of a _sequence_ of name/email changes isn't well defined.\nAs I remember it we take the name/email updates in sequence and then\napply a last one wins approach. It's not clear what would be done when\nwe have two, or three different John Doe sequences all mixed in.\n\n\nA broader issue for the corporate email mailbox systems is those that\nare allocated to roles. So you may have Traning1@corp.com thru\nTraining9@corp.com (we had) and if that training includes practical low\nhanging fruit examples from a project, it's difficult to disambiguate\nthose commits. More likely is say, having TestPC1 - TestPC9 that\nincluded debug commits, perhaps even with pair programming test & debug\nsessions, so allocation to individuals (rather than mailbox) becomes a\nreal problem. Hopefully that's rare in Sean's case.\n\nPhilip\n\n\n\n \n"},{"id":"451306","messageId":"e7563d48-d0f1-92d0-5f2f-e3d0d2f8d922@iee.email","threadId":"57549","inReplyTo":"874k42ar61.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-03-14T11:59:07Z","receivedAt":"2022-03-14T12:02:42Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 13/03/2022 13:35, Sean Allred wrote:\n> It's worth noting\n> that, if we do have to use this generated mapping, a mapping of over 10k\n> entries is surely to have the odd mistake every now and then.  So it's\n> not *totally* trustworthy (and never could be).\nI'd agree there. Adding role based mailboxes will further complicate\nsuch mappings.\n\nPhilip\n"},{"id":"451348","messageId":"xmqq1qz4p6qn.fsf@gitster.g","threadId":"57549","inReplyTo":"697d8717-bd3f-0871-d5b3-e6303c4ed726@iee.email","subject":"Re: Dealing with corporate email recycling","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2022-03-14T21:24:48Z","receivedAt":"2022-03-14T21:24:55Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Philip Oakley <philipoakley@iee.email> writes:\n\n> On 13/03/2022 23:16, Junio C Hamano wrote:\n>> Sean Allred <allred.sean@gmail.com> writes:\n>>\n>>> rather than use magic comments :-) Adapting to your suggestion, this\n>>> might look like the following:\n>>>\n>>>     A. U. Thor <foo@example.com> <ada.example.com> <[ approxidate ]>\n>> You'd probably want a timerange (valid-from and valid-to), instead\n>> of one single timestamp?\n> I'm not so sure that the date range approach won't bring it's own\n> problems. What happens outside the date range? i.e. Do we then have\n> three identities: Before, During, and After, with only 'During' being\n> defined?\n\nI have been assuming that the default is \"what the commit has is\ncorrect\".\n\n> I more see a single date being used as a termination point for an\n> existing email sequence that defines a retrospective end point for the\n> mapping of the old email addresses to a single person.\n\nImplicitly specifying the valid-from date (which is either the\nbeginning of time, or the newest of valid-until time for the same\nidentifying string that is older than the valid-until date for the\nentry in question) is fine.  I do not see fundamental difference\nbetween the approach you suggest and having an explicit valid-from\ndate.\n"},{"id":"451354","messageId":"e49830fb-aef1-6848-4101-60cbb66704f6@iee.email","threadId":"57549","inReplyTo":"xmqq1qz4p6qn.fsf@gitster.g","subject":"Re: Dealing with corporate email recycling","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-03-14T22:25:20Z","receivedAt":"2022-03-14T22:25:34Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 14/03/2022 21:24, Junio C Hamano wrote:\n> Philip Oakley <philipoakley@iee.email> writes:\n>\n>> On 13/03/2022 23:16, Junio C Hamano wrote:\n>>> Sean Allred <allred.sean@gmail.com> writes:\n>>>\n>>>> rather than use magic comments :-) Adapting to your suggestion, this\n>>>> might look like the following:\n>>>>\n>>>>     A. U. Thor <foo@example.com> <ada.example.com> <[ approxidate ]>\n>>> You'd probably want a timerange (valid-from and valid-to), instead\n>>> of one single timestamp?\n>> I'm not so sure that the date range approach won't bring it's own\n>> problems. What happens outside the date range? i.e. Do we then have\n>> three identities: Before, During, and After, with only 'During' being\n>> defined?\n> I have been assuming that the default is \"what the commit has is\n> correct\".\nThat default is only true when there are no date limitations because of\nemail re-use. i.e. singleton persons with unique emails do fit that\ndefault, which should be the majority.\n\nIf an old email has been reused, then that default becomes false, which\nwas Sean's starting point. In the corporate case, two (or more) distinct\nindividuals have used the same commit|author email address, and the hope\nis, for a way of providing a disambiguation of those persons, based on\ntheir email and the commit date.\n\n>\n>> I more see a single date being used as a termination point for an\n>> existing email sequence that defines a retrospective end point for the\n>> mapping of the old email addresses to a single person.\n> Implicitly specifying the valid-from date (which is either the\n> beginning of time, or the newest of valid-until time for the same\n> identifying string that is older than the valid-until date for the\n> entry in question) is fine.  I do not see fundamental difference\n> between the approach you suggest and having an explicit valid-from\n> date.\nWith the first case we guarantee that we have named cover for all of the\nchronology via bisection, while the trisection can leave gaps without\nany allocation to a person,  or possibly overlaps.\n\nA more convoluted case would be where three persons share the same\nemails in a rollover fashion, so the mailmap's simple name/email\nhandover becomes knotted and intertwined in the handovers\n(Joe3->Joe2->Joe1).\n\nP.\n\n\n\n"},{"id":"451365","messageId":"87y21c2ehl.fsf@gmail.com","threadId":"57549","inReplyTo":"697d8717-bd3f-0871-d5b3-e6303c4ed726@iee.email","subject":"Re: Dealing with corporate email recycling","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-03-15T01:23:57Z","receivedAt":"2022-03-15T01:26:04Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\nPhilip Oakley <philipoakley@iee.email> writes:\n> A broader issue for the corporate email mailbox systems is those that\n> are allocated to roles. So you may have Traning1@corp.com thru\n> Training9@corp.com (we had) and if that training includes practical low\n> hanging fruit examples from a project, it's difficult to disambiguate\n> those commits. More likely is say, having TestPC1 - TestPC9 that\n> included debug commits, perhaps even with pair programming test & debug\n> sessions, so allocation to individuals (rather than mailbox) becomes a\n> real problem. Hopefully that's rare in Sean's case.\n\nYep, this wouldn't happen for us.  Lots of other processes depend on\nthere being an individual making the commit.\n\nI'd also be surprised if this didn't cause process problems for other\nfolks, too.\n\n--\nSean Allred\n"},{"id":"451373","messageId":"87o8282d79.fsf@gmail.com","threadId":"57549","inReplyTo":"878rtebxk0.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"Sean Allred","fromEmail":"allred.sean@gmail.com","sentAt":"2022-03-15T01:27:59Z","receivedAt":"2022-03-15T01:53:52Z","isPatch":false,"sender":{"key":"allred.sean@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2082195?v=4"},"body":"\n(CC others who have been involved in the conversation; I started a\nseparate thread off the original post since this takes an entirely\nseparate direction.)\n\nI wanted to provide a possible approach for other organizations that may\nrun into this issue in the future.  If you have the opportunity that we\ndo in rewriting history, we found a nice workaround by way of\n'unique-ifying' the email address using the '+' syntax supported by most\nemail providers -- notably for us, we tested that it does work with\nExchange.\n\nAs I went through previously, the problem is that our organization\nrecycles emails like <Sean@corp.net>.  If the owner of <Sean@corp.net>\nleaves, I have the opportunity to take ownership of <Sean@corp.net>\nafter a waiting period.  This presents the core problem: <Sean@corp.net>\ncould be two different people at two different points in time.\n\nOur first approach to solving this problem was to collaborate with our\ninternal IT folks to see if we could additionally generate a totally\nunique email address based on that employee's ID.  For various reasons\nwith which I'm not familiar enough to elaborate, this wasn't really\npossible -- we could not be provided a totally static email address for\nany given person.  After bouncing the problem around internally for a\nwhile, we reached out to this list :-) and I think a lot of cool\nconversation is happening that I want to continue!  It seems there is\nstill value in finding a generalizable solution.\n\n---\n\nAs for solutions that are *not* necessarily generalizable, what we were\nable to come up with is to use a 'unique-ified' email of the form\n<Sean+314159@corp.net>, where 314159 might be my employee ID.  That way,\neven if another Sean might inherit <Sean@corp.net>, they can never\ninherit <Sean+314159@corp.net>.  The use of this unique-ified email\naddress will be enforced via pre-receive -- possibly checking for its\nexistence in the mailmap as well.\n\nUsing this unique-ified email address, we can construct a mailmap like\nthe following:\n\n    Sean Allred <Sean@corp.net> <Sean+314159@corp.net>\n\nThis uses the already-built functionality of gitmailmap to interpret the\ncommitter email of <Sean+314159@corp.net> to the friendlier\n<Sean@corp.net> while still maintaining the requirement (for us) that\nthe email on the commit be a real, usable email.\n\nWhen I eventually get hit by a bus and give up <Sean@corp.net>, the\nmailmap can be updated to\n\n    Sean Allred <devnull@corp.net> <Sean+314159@corp.net>\n    Sean Allblue <Sean@corp.net> <Sean+271828@corp.net>\n\n(possibly using some support list instead of <devnull@corp.net>).\n\nThis has proven itself in a few informal design chats already and is\ngoing to be the approach we take for our replay.  I look forward to a\nfuture situation where we might be able to use SSH key validity periods\ninstead :-) but that functionality will surely only be available after\nour migration in a few months' time.\n\nI hope this helps others!\n\n--\nSean Allred\n"},{"id":"451390","messageId":"eecab055-94bc-2174-0490-28b76cf48b0e@iee.email","threadId":"57549","inReplyTo":"87y21c2ehl.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"Philip Oakley","fromEmail":"philipoakley@iee.email","sentAt":"2022-03-15T11:15:18Z","receivedAt":"2022-03-15T11:15:26Z","isPatch":false,"sender":{"key":"philipoakley@iee.email","avatar":"https://avatars.githubusercontent.com/u/914343?v=4"},"body":"On 15/03/2022 01:23, Sean Allred wrote:\n> Philip Oakley <philipoakley@iee.email> writes:\n>> A broader issue for the corporate email mailbox systems is those that\n>> are allocated to roles. So you may have Traning1@corp.com thru\n>> Training9@corp.com (we had) and if that training includes practical low\n>> hanging fruit examples from a project, it's difficult to disambiguate\n>> those commits. More likely is say, having TestPC1 - TestPC9 that\n>> included debug commits, perhaps even with pair programming test & debug\n>> sessions, so allocation to individuals (rather than mailbox) becomes a\n>> real problem. Hopefully that's rare in Sean's case.\n> Yep, this wouldn't happen for us.  Lots of other processes depend on\n> there being an individual making the commit.\n>\n> I'd also be surprised if this didn't cause process problems for other\n> folks, too.\nI was in equipment engineering where independent dedicated test\nequipment was more common, rather than pure software, so the potential\nfor  role based commits from TestPC1 - TestPC9 was far more likely (for\ncases where it's the hardware that needs to be understood, not the\nengineers' choice of code ;-)\n"},{"id":"451646","messageId":"e0b85f26-ab5e-b3ae-a2fb-e7d927c46763@softwolves.pp.se","threadId":"57549","inReplyTo":"878rtebxk0.fsf@gmail.com","subject":"Re: Dealing with corporate email recycling","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2022-03-18T21:22:45Z","receivedAt":"2022-03-18T21:29:34Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"On Sat, 12 Mar 2022, Sean Allred wrote:\n\n> We are currently replaying a 15-year SVN history into Git -- with\n> contributions from thousands of developers -- and are faced with the\n> challenge of corporate email recycling, departures, re-hires, and name\n> changes causing identity issues.\n\nI have performed a couple of imports of old version history into Git, from\nvarious version control systems, some of them with history dating to before\nthe corporation even had e-mail addresses for employees. In those cases I\nfound that the easiest option was just to use whatever user identification\nwas available in the old version control system -- Git does not explicitely\nrequire a valid e-mail address in the author and committer header.\n\nFor Subversion import, for instance, I used \"Name <login>\" where \"login\" was\nthe Subversion committer ID, and \"Name\" was from a mapping file I created\nfor the repository. Where records were sketchy and Name information was not\navailable, I would just use \"<login>\".\n\nWhen it comes to name changes, I have had scripts map login + date to name.\nFor instance, I changed my last name when I married, so I would have my old\n(I don't know what the masculine equivalent of \"maiden name\" is in English)\nmapped up until a specific date, and my current name afterwards.\n\n-- \n\\\\// Peter\n"}]}