{"thread":{"id":"32002","subject":"RFD: fast-import is picky with author names (and maybe it should - but how much so?)","startedAt":"2012-11-02T14:43:24Z","lastAt":"2012-11-13T18:15:59Z","messageCount":24,"participants":["Michael J Gruber","Jeff King","Felipe Contreras","A Large Angry SCM","Junio C Hamano"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"202426","messageId":"5093DC0C.5000603@drmicha.warpmail.net","threadId":"32002","inReplyTo":null,"subject":"RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2012-11-02T14:43:24Z","receivedAt":"2012-11-02T14:43:24Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"It seems that our fast-import is super picky with regards to author\nnames. I've encountered author names like\n\nFoo Bar<foo.bar@dev.null>\nFoo Bar <foo.bar@dev.null\nfoo.bar@dev.null\n\nin the self-hosting repo of some other dvcs, and the question is how to\ntranslate them faithfully into a git author name. In general, we try to do\n\nfullotherdvcsname <none@none>\n\nif the other system's entry does not parse as a git author name, but\nfast-import does not accept either of\n\nFoo Bar<foo.bar@dev.null> <none@none>\n\"Foo Bar<foo.bar@dev.null>\" <none@none>\n\nbecause of the way it parses for <>. While the above could be easily\nturned into\n\nFoo Bar <foo.bar@dev.null>\n\nit would not be a faithful representation of the original commit in the\nother dvcs.\n\nSo the question is:\n\n- How should we represent botched author entries faithfully?\n\nAs a cororollary, fast-import may need to change or not.\n\nMichael\n\nP.S.: Yes, dvcs=hg, and the \"earlier\" remote-hg helper chokes on these.\ngarbage in crash out :(\n"},{"id":"202429","messageId":"5093DD1A.1080104@drmicha.warpmail.net","threadId":"32002","inReplyTo":"5093DC0C.5000603@drmicha.warpmail.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2012-11-02T14:47:54Z","receivedAt":"2012-11-02T14:47:54Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Some additional input:\n\n[mjg@localhost git]$ git commit --author='\"is this<ok@or.not>\"\n<whats@up>' --allow-empty -m test\n[detached HEAD 0734308] test\n Author: is thisok@or.not <whats@up>\n[mjg@localhost git]$ git show\ncommit 0734308b7bf372227bf9f5b9fd6b4b403df33b9e\nAuthor: is thisok@or.not <whats@up>\nDate:   Fri Nov 2 15:45:23 2012 +0100\n\n    test\n"},{"id":"202666","messageId":"20121108200919.GP15560@sigill.intra.peff.net","threadId":"32002","inReplyTo":"5093DC0C.5000603@drmicha.warpmail.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-11-08T20:09:19Z","receivedAt":"2012-11-08T20:09:19Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, Nov 02, 2012 at 03:43:24PM +0100, Michael J Gruber wrote:\n\n> It seems that our fast-import is super picky with regards to author\n> names. I've encountered author names like\n> \n> Foo Bar<foo.bar@dev.null>\n> Foo Bar <foo.bar@dev.null\n> foo.bar@dev.null\n> \n> in the self-hosting repo of some other dvcs, and the question is how to\n> translate them faithfully into a git author name.\n\nIt is not just fast-import. Git's author field looks like an rfc822\naddress, but it's much simpler. It fundamentally does not allow angle\nbrackets in the \"name\" field, regardless of any quoting. As you noted in\nyour followup, we strip them out if you provide them via\nGIT_AUTHOR_NAME.\n\nI doubt this will change anytime soon due to the compatibility fallout.\nSo it is up to generators of fast-import streams to decide how to encode\nwhat they get from another system (you could come up with an encoding\nscheme that represents angle brackets).\n\n> In general, we try to do\n> \n> fullotherdvcsname <none@none>\n> \n> if the other system's entry does not parse as a git author name, but\n> fast-import does not accept either of\n> \n> Foo Bar<foo.bar@dev.null> <none@none>\n> \"Foo Bar<foo.bar@dev.null>\" <none@none>\n> \n> because of the way it parses for <>. While the above could be easily\n> turned into\n> \n> Foo Bar <foo.bar@dev.null>\n> \n> it would not be a faithful representation of the original commit in the\n> other dvcs.\n\nI'd think that if a remote system has names with angle brackets and\nemail-looking things inside them, we would do better to stick them in\nthe email field rather than putting in a useless <none@none>. The latter\nshould only be used for systems that lack the information.\n\nBut that is a quality-of-implementation issue for the import scripts\n(and they may even want to have options, just like git-cvsimport allows\nmapping cvs usernames into full identities).\n\n-Peff\n"},{"id":"202687","messageId":"509CCCBC.8010102@drmicha.warpmail.net","threadId":"32002","inReplyTo":"20121108200919.GP15560@sigill.intra.peff.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2012-11-09T09:28:28Z","receivedAt":"2012-11-09T09:28:28Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Jeff King venit, vidit, dixit 08.11.2012 21:09:\n> On Fri, Nov 02, 2012 at 03:43:24PM +0100, Michael J Gruber wrote:\n> \n>> It seems that our fast-import is super picky with regards to author\n>> names. I've encountered author names like\n>>\n>> Foo Bar<foo.bar@dev.null>\n>> Foo Bar <foo.bar@dev.null\n>> foo.bar@dev.null\n>>\n>> in the self-hosting repo of some other dvcs, and the question is how to\n>> translate them faithfully into a git author name.\n> \n> It is not just fast-import. Git's author field looks like an rfc822\n> address, but it's much simpler. It fundamentally does not allow angle\n> brackets in the \"name\" field, regardless of any quoting. As you noted in\n> your followup, we strip them out if you provide them via\n> GIT_AUTHOR_NAME.\n> \n> I doubt this will change anytime soon due to the compatibility fallout.\n> So it is up to generators of fast-import streams to decide how to encode\n> what they get from another system (you could come up with an encoding\n> scheme that represents angle brackets).\n\nI don't expect our requirements to change. For one thing, I was\nsurprised that git-commit is more tolerant than git-fast-import, but it\nmakes a lot of sense to avoid any behind-the-back conversions in the\nimporter.\n\n>> In general, we try to do\n>>\n>> fullotherdvcsname <none@none>\n>>\n>> if the other system's entry does not parse as a git author name, but\n>> fast-import does not accept either of\n>>\n>> Foo Bar<foo.bar@dev.null> <none@none>\n>> \"Foo Bar<foo.bar@dev.null>\" <none@none>\n>>\n>> because of the way it parses for <>. While the above could be easily\n>> turned into\n>>\n>> Foo Bar <foo.bar@dev.null>\n>>\n>> it would not be a faithful representation of the original commit in the\n>> other dvcs.\n> \n> I'd think that if a remote system has names with angle brackets and\n> email-looking things inside them, we would do better to stick them in\n> the email field rather than putting in a useless <none@none>. The latter\n> should only be used for systems that lack the information.\n> \n> But that is a quality-of-implementation issue for the import scripts\n> (and they may even want to have options, just like git-cvsimport allows\n> mapping cvs usernames into full identities).\n\nThat was more my real concern. In our cvs and svn interfaces, we even\nencourage the use of author maps. For example, if you use an author map,\ngit-svn errors out if it encounters an svn user name which is not in the\nmap. On the other hand, we can map all (most?) svn user names faithfully\nwithout using a map (e.g. to \"username <none@none>\").\n\nHg seems to store just anything in the author field (\"committer\"). The\nvarious interfaces that are floating around do some behind-the-back\nconversion to git format. The more conversions they do, the better they\nseem to work (no erroring out) but I'm wondering whether it's really a\ngood thing, or whether we should encourage a more diligent approach\nwhich requires a user to map non-conforming author names wilfully.\n\nMichael\n"},{"id":"202698","messageId":"CAMP44s3Lhxzcj93=e8TXwqAVvGJBKhZEVX33G8Q=n2+8+UfCww@mail.gmail.com","threadId":"32002","inReplyTo":"509CCCBC.8010102@drmicha.warpmail.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-09T14:34:34Z","receivedAt":"2012-11-09T14:34:34Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Fri, Nov 9, 2012 at 10:28 AM, Michael J Gruber\n<git@drmicha.warpmail.net> wrote:\n\n> Hg seems to store just anything in the author field (\"committer\"). The\n> various interfaces that are floating around do some behind-the-back\n> conversion to git format. The more conversions they do, the better they\n> seem to work (no erroring out) but I'm wondering whether it's really a\n> good thing, or whether we should encourage a more diligent approach\n> which requires a user to map non-conforming author names wilfully.\n\nSo you propose that when somebody does 'git clone hg::hg hg-git' the\nthing should fail. I hope you don't think it's too unbecoming for me\nto say that I disagree.\n\nIMO it should be git fast-import the one that converts these bad\nauthors, not every single tool out there. Maybe throw a warning, but\nthat's all. Or maybe generate a list of bad authors ready to be filled\nout. That way when a project is doing a real conversion, say, when\nmoving to git, they can run the conversion once and see which authors\nare bad and not multiple times, each try taking longer than the next.\n\n-- \nFelipe Contreras\n"},{"id":"202762","messageId":"509E8EB2.7040509@drmicha.warpmail.net","threadId":"32002","inReplyTo":"CAMP44s3Lhxzcj93=e8TXwqAVvGJBKhZEVX33G8Q=n2+8+UfCww@mail.gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2012-11-10T17:28:18Z","receivedAt":"2012-11-10T17:28:18Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Felipe Contreras venit, vidit, dixit 09.11.2012 15:34:\n> On Fri, Nov 9, 2012 at 10:28 AM, Michael J Gruber\n> <git@drmicha.warpmail.net> wrote:\n> \n>> Hg seems to store just anything in the author field (\"committer\"). The\n>> various interfaces that are floating around do some behind-the-back\n>> conversion to git format. The more conversions they do, the better they\n>> seem to work (no erroring out) but I'm wondering whether it's really a\n>> good thing, or whether we should encourage a more diligent approach\n>> which requires a user to map non-conforming author names wilfully.\n> \n> So you propose that when somebody does 'git clone hg::hg hg-git' the\n> thing should fail. I hope you don't think it's too unbecoming for me\n> to say that I disagree.\n\nThere is no need to disagree with a proposal I haven't made. I would\ndisagree with the proposal that I haven't made, too.\n\n> IMO it should be git fast-import the one that converts these bad\n> authors, not every single tool out there. Maybe throw a warning, but\n> that's all. Or maybe generate a list of bad authors ready to be filled\n> out. That way when a project is doing a real conversion, say, when\n> moving to git, they can run the conversion once and see which authors\n> are bad and not multiple times, each try taking longer than the next.\n\nAs Jeff pointed out, git-fast-import expects output conforming to a\ncertain standard, and that's not going to change. import is agnostic to\nwhere its import stream is coming from. Only the producer of that stream\ncan have additional information about the provenience of the stream's\ndata which may aid (possibly together with user input or choices) in\ntransforming that into something conforming.\n\nMichael\n"},{"id":"202767","messageId":"CAMP44s219Zi2NPt2vA+6Od_sVstFK85OXZK-9K1OCFpVh220+A@mail.gmail.com","threadId":"32002","inReplyTo":"509E8EB2.7040509@drmicha.warpmail.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-10T18:43:18Z","receivedAt":"2012-11-10T18:43:18Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sat, Nov 10, 2012 at 6:28 PM, Michael J Gruber\n<git@drmicha.warpmail.net> wrote:\n> Felipe Contreras venit, vidit, dixit 09.11.2012 15:34:\n>> On Fri, Nov 9, 2012 at 10:28 AM, Michael J Gruber\n>> <git@drmicha.warpmail.net> wrote:\n>>\n>>> Hg seems to store just anything in the author field (\"committer\"). The\n>>> various interfaces that are floating around do some behind-the-back\n>>> conversion to git format. The more conversions they do, the better they\n>>> seem to work (no erroring out) but I'm wondering whether it's really a\n>>> good thing, or whether we should encourage a more diligent approach\n>>> which requires a user to map non-conforming author names wilfully.\n>>\n>> So you propose that when somebody does 'git clone hg::hg hg-git' the\n>> thing should fail. I hope you don't think it's too unbecoming for me\n>> to say that I disagree.\n>\n> There is no need to disagree with a proposal I haven't made. I would\n> disagree with the proposal that I haven't made, too.\n\nAll right, we shouldn't encourage a more diligent approach which\nrequires a user to map author names then.\n\n>> IMO it should be git fast-import the one that converts these bad\n>> authors, not every single tool out there. Maybe throw a warning, but\n>> that's all. Or maybe generate a list of bad authors ready to be filled\n>> out. That way when a project is doing a real conversion, say, when\n>> moving to git, they can run the conversion once and see which authors\n>> are bad and not multiple times, each try taking longer than the next.\n>\n> As Jeff pointed out, git-fast-import expects output conforming to a\n> certain standard, and that's not going to change. import is agnostic to\n> where its import stream is coming from. Only the producer of that stream\n> can have additional information about the provenience of the stream's\n> data which may aid (possibly together with user input or choices) in\n> transforming that into something conforming.\n\nWe already know where the import of those streams come from:\nmercurial, bazaar, etc. There's absolutely nothing the tools exporting\ndata from those repositories can do, except try to convert all kind of\nweird names--and many tools do it poorly.\n\nSo, the options are:\n\na) Leave the name conversion to the export tools, and when they miss\nsome weird corner case, like 'Author <email', let the user face the\nconsequences, perhaps after an hour of the process.\n\nWe know there are sources of data that don't have git-formatted author\nnames, so we know every tool out there must do this checking.\n\nIn addition to that, let the export tool decide what to do when one of\nthese bad names appear, which in many cases probably means do nothing,\nso the user would not even see that such a bad name was there, which\nmight not be what they want.\n\nb) Do the name conversion in fast-import itself, perhaps optionally,\nso if a tool missed some weird corner case, the user does not have to\nface the consequences.\n\nThe tool writers don't have to worry about this, so we would not have\ntools out there doing a half-assed job of this.\n\nAnd what happens when such bad names end up being consistent: warning,\na scaffold mapping of bad names, etc.\n\n\nOne is bad for the users, and the tools writers, only disadvantages,\nthe other is good for the users and the tools writers, only\nadvantages.\n\n-- \nFelipe Contreras\n"},{"id":"202777","messageId":"509EAA45.8020005@gmail.com","threadId":"32002","inReplyTo":"CAMP44s219Zi2NPt2vA+6Od_sVstFK85OXZK-9K1OCFpVh220+A@mail.gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2012-11-10T19:25:57Z","receivedAt":"2012-11-10T19:25:57Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 11/10/2012 01:43 PM, Felipe Contreras wrote:\n> On Sat, Nov 10, 2012 at 6:28 PM, Michael J Gruber\n> <git@drmicha.warpmail.net>  wrote:\n>> Felipe Contreras venit, vidit, dixit 09.11.2012 15:34:\n>>> On Fri, Nov 9, 2012 at 10:28 AM, Michael J Gruber\n>>> <git@drmicha.warpmail.net>  wrote:\n>>>\n>>>> Hg seems to store just anything in the author field (\"committer\"). The\n>>>> various interfaces that are floating around do some behind-the-back\n>>>> conversion to git format. The more conversions they do, the better they\n>>>> seem to work (no erroring out) but I'm wondering whether it's really a\n>>>> good thing, or whether we should encourage a more diligent approach\n>>>> which requires a user to map non-conforming author names wilfully.\n>>>\n>>> So you propose that when somebody does 'git clone hg::hg hg-git' the\n>>> thing should fail. I hope you don't think it's too unbecoming for me\n>>> to say that I disagree.\n>>\n>> There is no need to disagree with a proposal I haven't made. I would\n>> disagree with the proposal that I haven't made, too.\n>\n> All right, we shouldn't encourage a more diligent approach which\n> requires a user to map author names then.\n>\n>>> IMO it should be git fast-import the one that converts these bad\n>>> authors, not every single tool out there. Maybe throw a warning, but\n>>> that's all. Or maybe generate a list of bad authors ready to be filled\n>>> out. That way when a project is doing a real conversion, say, when\n>>> moving to git, they can run the conversion once and see which authors\n>>> are bad and not multiple times, each try taking longer than the next.\n>>\n>> As Jeff pointed out, git-fast-import expects output conforming to a\n>> certain standard, and that's not going to change. import is agnostic to\n>> where its import stream is coming from. Only the producer of that stream\n>> can have additional information about the provenience of the stream's\n>> data which may aid (possibly together with user input or choices) in\n>> transforming that into something conforming.\n>\n> We already know where the import of those streams come from:\n> mercurial, bazaar, etc. There's absolutely nothing the tools exporting\n> data from those repositories can do, except try to convert all kind of\n> weird names--and many tools do it poorly.\n>\n> So, the options are:\n>\n> a) Leave the name conversion to the export tools, and when they miss\n> some weird corner case, like 'Author<email', let the user face the\n> consequences, perhaps after an hour of the process.\n>\n> We know there are sources of data that don't have git-formatted author\n> names, so we know every tool out there must do this checking.\n>\n> In addition to that, let the export tool decide what to do when one of\n> these bad names appear, which in many cases probably means do nothing,\n> so the user would not even see that such a bad name was there, which\n> might not be what they want.\n>\n> b) Do the name conversion in fast-import itself, perhaps optionally,\n> so if a tool missed some weird corner case, the user does not have to\n> face the consequences.\n>\n> The tool writers don't have to worry about this, so we would not have\n> tools out there doing a half-assed job of this.\n>\n> And what happens when such bad names end up being consistent: warning,\n> a scaffold mapping of bad names, etc.\n>\n>\n> One is bad for the users, and the tools writers, only disadvantages,\n> the other is good for the users and the tools writers, only\n> advantages.\n>\n\nc) Do the name conversion, and whatever other cleanup and manipulations \nyou're interesting in, in a filter between the exporter and git-fast-import.\n"},{"id":"202823","messageId":"CAMP44s1dsEU=E8tdgMYxWFyFw+F03bstdb5o7Ww_-RCQPd3R0w@mail.gmail.com","threadId":"32002","inReplyTo":"509EAA45.8020005@gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-11T12:41:10Z","receivedAt":"2012-11-11T12:41:10Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sat, Nov 10, 2012 at 8:25 PM, A Large Angry SCM <gitzilla@gmail.com> wrote:\n> On 11/10/2012 01:43 PM, Felipe Contreras wrote:\n\n>> So, the options are:\n>>\n>> a) Leave the name conversion to the export tools, and when they miss\n>> some weird corner case, like 'Author<email', let the user face the\n>> consequences, perhaps after an hour of the process.\n>>\n>> We know there are sources of data that don't have git-formatted author\n>> names, so we know every tool out there must do this checking.\n>>\n>> In addition to that, let the export tool decide what to do when one of\n>> these bad names appear, which in many cases probably means do nothing,\n>> so the user would not even see that such a bad name was there, which\n>> might not be what they want.\n>>\n>> b) Do the name conversion in fast-import itself, perhaps optionally,\n>> so if a tool missed some weird corner case, the user does not have to\n>> face the consequences.\n>>\n>> The tool writers don't have to worry about this, so we would not have\n>> tools out there doing a half-assed job of this.\n>>\n>> And what happens when such bad names end up being consistent: warning,\n>> a scaffold mapping of bad names, etc.\n>>\n>>\n>> One is bad for the users, and the tools writers, only disadvantages,\n>> the other is good for the users and the tools writers, only\n>> advantages.\n>>\n>\n> c) Do the name conversion, and whatever other cleanup and manipulations\n> you're interesting in, in a filter between the exporter and git-fast-import.\n\nSuch a filter would probably be quite complicated, and would decrease\nperformance.\n\n-- \nFelipe Contreras\n"},{"id":"202888","messageId":"509FD9BC.7050204@gmail.com","threadId":"32002","inReplyTo":"CAMP44s1dsEU=E8tdgMYxWFyFw+F03bstdb5o7Ww_-RCQPd3R0w@mail.gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2012-11-11T17:00:44Z","receivedAt":"2012-11-11T17:00:44Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 11/11/2012 07:41 AM, Felipe Contreras wrote:\n> On Sat, Nov 10, 2012 at 8:25 PM, A Large Angry SCM<gitzilla@gmail.com>  wrote:\n>> On 11/10/2012 01:43 PM, Felipe Contreras wrote:\n>\n>>> So, the options are:\n>>>\n>>> a) Leave the name conversion to the export tools, and when they miss\n>>> some weird corner case, like 'Author<email', let the user face the\n>>> consequences, perhaps after an hour of the process.\n>>>\n>>> We know there are sources of data that don't have git-formatted author\n>>> names, so we know every tool out there must do this checking.\n>>>\n>>> In addition to that, let the export tool decide what to do when one of\n>>> these bad names appear, which in many cases probably means do nothing,\n>>> so the user would not even see that such a bad name was there, which\n>>> might not be what they want.\n>>>\n>>> b) Do the name conversion in fast-import itself, perhaps optionally,\n>>> so if a tool missed some weird corner case, the user does not have to\n>>> face the consequences.\n>>>\n>>> The tool writers don't have to worry about this, so we would not have\n>>> tools out there doing a half-assed job of this.\n>>>\n>>> And what happens when such bad names end up being consistent: warning,\n>>> a scaffold mapping of bad names, etc.\n>>>\n>>>\n>>> One is bad for the users, and the tools writers, only disadvantages,\n>>> the other is good for the users and the tools writers, only\n>>> advantages.\n>>>\n>>\n>> c) Do the name conversion, and whatever other cleanup and manipulations\n>> you're interesting in, in a filter between the exporter and git-fast-import.\n>\n> Such a filter would probably be quite complicated, and would decrease\n> performance.\n>\n\nReally?\n\nThe fast import stream protocol is pretty simple. All the filter really \nneeds to do is pass through everything that isn't a 'commit' command. \nAnd for the 'commit' command, it only needs to do something with the \n'author' and 'committer' lines; passing through everything else.\n\nI agree that an additional filter _may_ decrease performance somewhat if \nyou are already CPU constrained. But I suspect that the effect would be \nnegligible compared to the all of the SHA-1 calculations.\n"},{"id":"202897","messageId":"20121111171518.GA20115@sigill.intra.peff.net","threadId":"32002","inReplyTo":"509FD9BC.7050204@gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-11-11T17:15:18Z","receivedAt":"2012-11-11T17:15:18Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Nov 11, 2012 at 12:00:44PM -0500, A Large Angry SCM wrote:\n\n> >>>a) Leave the name conversion to the export tools, and when they miss\n> >>>some weird corner case, like 'Author<email', let the user face the\n> >>>consequences, perhaps after an hour of the process.\n> [...]\n> >>>b) Do the name conversion in fast-import itself, perhaps optionally,\n> >>>so if a tool missed some weird corner case, the user does not have to\n> >>>face the consequences.\n> [...]\n> >>c) Do the name conversion, and whatever other cleanup and manipulations\n> >>you're interesting in, in a filter between the exporter and git-fast-import.\n> >\n> >Such a filter would probably be quite complicated, and would decrease\n> >performance.\n> >\n> \n> Really?\n> \n> The fast import stream protocol is pretty simple. All the filter\n> really needs to do is pass through everything that isn't a 'commit'\n> command. And for the 'commit' command, it only needs to do something\n> with the 'author' and 'committer' lines; passing through everything\n> else.\n> \n> I agree that an additional filter _may_ decrease performance somewhat\n> if you are already CPU constrained. But I suspect that the effect\n> would be negligible compared to the all of the SHA-1 calculations.\n\nIt might be measurable, as you are passing every byte of every version\nof every file in the repo through an extra pipe. But more importantly, I\ndon't think it helps.\n\nIf there is not a standard filter for fixing up names, we do not need to\ncare. The user can use \"sed\" or whatever and pay the performance penalty\n(and deal with the possibility of errors from being lazy about parsing\nthe fast-import stream).\n\nIf there is a standard filter, then what is the advantage in doing it as\na pipe? Why not just teach fast-import the same trick (and possibly make\nit optional)? That would be simpler, more efficient, and it would make\nit easier for remote helpers to turn it on (they use a command-line\nswitch rather than setting up an extra process).\n\nBut what I don't understand is: what would such a standard filter look\nlike? Fast-import (or a filter) would already receive the exporter's\nbest attempt at a git-like ident string. We can clean up and normalize\nthings like whitespace (and we probably should if we do not do so\nalready). But beyond that, we have no context about the name; only the\nexporter has that.\n\nSo if we receive:\n\n  Foo Bar<foo.bar@example.com> <none@none>\n\nor:\n\n  Foo Bar<foo.bar@example.com <none@none>\n\nor:\n\n  Foo Bar<foo.bar@example.com\n\nwhat do we do with it? Is the first part a malformed name/email pair,\nand the second part is crap added by a lazy exporter? Or does the\nexporter want to keep the angle brackets as part of the name field? Is\nthere a malformed email in the last one, or no email at all?\n\nThe exporter is the only program that actually knows where the data came\nfrom, how it should be broken down, and what is appropriate for pulling\ndata out of its particular source system. For that reason, the exporter\nhas to be the place where we come up with a syntactically correct and\nunambiguous ident.\n\nI am not opposed to adding a mailmap-like feature to fast-import to map\nidentities, but it has to start with sane, unambiguous output from the\nexporter.\n\n-Peff\n"},{"id":"202898","messageId":"CAMP44s1pWm_n-SwB5Bi8UxM-oRG=4dGXq7jVKx_E1rcoRaXaHw@mail.gmail.com","threadId":"32002","inReplyTo":"509FD9BC.7050204@gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-11T17:16:43Z","receivedAt":"2012-11-11T17:16:43Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sun, Nov 11, 2012 at 6:00 PM, A Large Angry SCM <gitzilla@gmail.com> wrote:\n> On 11/11/2012 07:41 AM, Felipe Contreras wrote:\n\n>> Such a filter would probably be quite complicated, and would decrease\n>> performance.\n>\n> Really?\n>\n> The fast import stream protocol is pretty simple. All the filter really\n> needs to do is pass through everything that isn't a 'commit' command. And\n> for the 'commit' command, it only needs to do something with the 'author'\n> and 'committer' lines; passing through everything else.\n\nAnd how do you propose to find the commit commands without parsing all\nthe other commands? If you randomly look for lines that begin with\n'commit /refs' you might end up in the middle of a commit message or\nthe contents of a file.\n\n> I agree that an additional filter _may_ decrease performance somewhat if you\n> are already CPU constrained. But I suspect that the effect would be\n> negligible compared to the all of the SHA-1 calculations.\n\nWell. If it's so easy surely you can write one quickly, and I can measure it.\n\nCheers.\n\n-- \nFelipe Contreras\n"},{"id":"202900","messageId":"509FE2EA.3020407@gmail.com","threadId":"32002","inReplyTo":"CAMP44s1pWm_n-SwB5Bi8UxM-oRG=4dGXq7jVKx_E1rcoRaXaHw@mail.gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2012-11-11T17:39:54Z","receivedAt":"2012-11-11T17:39:54Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 11/11/2012 12:16 PM, Felipe Contreras wrote:\n> On Sun, Nov 11, 2012 at 6:00 PM, A Large Angry SCM<gitzilla@gmail.com>  wrote:\n>> On 11/11/2012 07:41 AM, Felipe Contreras wrote:\n>\n>>> Such a filter would probably be quite complicated, and would decrease\n>>> performance.\n>>\n>> Really?\n>>\n>> The fast import stream protocol is pretty simple. All the filter really\n>> needs to do is pass through everything that isn't a 'commit' command. And\n>> for the 'commit' command, it only needs to do something with the 'author'\n>> and 'committer' lines; passing through everything else.\n>\n> And how do you propose to find the commit commands without parsing all\n> the other commands? If you randomly look for lines that begin with\n> 'commit /refs' you might end up in the middle of a commit message or\n> the contents of a file.\n\nI didn't say you didn't have to parse the protocol. I said that the \nprotocol is pretty simple.\n\n>\n>> I agree that an additional filter _may_ decrease performance somewhat if you\n>> are already CPU constrained. But I suspect that the effect would be\n>> negligible compared to the all of the SHA-1 calculations.\n>\n> Well. If it's so easy surely you can write one quickly, and I can measure it.\n\nNot my itch; You care, you do it.\n"},{"id":"202902","messageId":"CAMP44s1mny-fBCxywM0V=AgEoxV5EZdDWc_0NK3gepcKf32nww@mail.gmail.com","threadId":"32002","inReplyTo":"20121111171518.GA20115@sigill.intra.peff.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-11T17:45:32Z","receivedAt":"2012-11-11T17:45:32Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sun, Nov 11, 2012 at 6:15 PM, Jeff King <peff@peff.net> wrote:\n> On Sun, Nov 11, 2012 at 12:00:44PM -0500, A Large Angry SCM wrote:\n\n> If there is a standard filter, then what is the advantage in doing it as\n> a pipe? Why not just teach fast-import the same trick (and possibly make\n> it optional)? That would be simpler, more efficient, and it would make\n> it easier for remote helpers to turn it on (they use a command-line\n> switch rather than setting up an extra process).\n\nRight, but instead of a command-line switch it probably should be\nenabled on the stream:\n\n  feature clean-authors\n\nOr something.\n\n> But what I don't understand is: what would such a standard filter look\n> like? Fast-import (or a filter) would already receive the exporter's\n> best attempt at a git-like ident string.\n\nCurrently, yeah, because there's no other option. It's either try to\nclean it up, or fail.\n\nBut if 'git fast-import' as a superior alternative, I certainly would\nremove my custom code and enable that feature.\n\n> We can clean up and normalize\n> things like whitespace (and we probably should if we do not do so\n> already). But beyond that, we have no context about the name; only the\n> exporter has that.\n\nThere is no context.\n\n> So if we receive:\n>\n>   Foo Bar<foo.bar@example.com> <none@none>\n>\n> or:\n>\n>   Foo Bar<foo.bar@example.com <none@none>\n>\n> or:\n>\n>   Foo Bar<foo.bar@example.com\n>\n> what do we do with it? Is the first part a malformed name/email pair,\n> and the second part is crap added by a lazy exporter? Or does the\n> exporter want to keep the angle brackets as part of the name field? Is\n> there a malformed email in the last one, or no email at all?\n\nThese are exactly the same questions every exporter must answer. And\nthere's no answer, because the field is not a git author, it's a\nmercurial user, or a bazaar committer, or who knows what.\n\n>From whatever source, these all might be valid authors:\njohn\njohn <john@travolta.com> (grease)\n<test@test.com>\ntest@test.com\ntest<test@test.com>\ntest <test@test.com\ntest # a space\ntest < test@test.com >\ntest >test@est.com>\ntest <test <at> test <dot> com>\n<>\n>\n<\nThe first chapter of the LOTR\n\nThere is no context.\n\n> The exporter is the only program that actually knows where the data came\n> from,\n\nIt doesn't matter where it came from, it's not a name/email pair.\n\n> how it should be broken down,\n\nIt cannot be broken down, it's free-form text. Any text.\n\n> and what is appropriate for pulling\n> data out of its particular source system.\n\nThis free-form text is the lowest granularity. There is nothing else.\n\n> For that reason, the exporter\n> has to be the place where we come up with a syntactically correct and\n> unambiguous ident.\n\n*If* the exporter is able to do this, sure, but many don't have any\nmore information.\n\nSee:\n\n% hg commit -u 'Foo Bar<foo.bar@example.com> <none@none>' -m one\n% hg --debug log\nchangeset:   0:5ef37a2c773f02d0e01f1ecdcc59149832d294e8\ntag:         tip\nphase:       draft\nparent:      -1:0000000000000000000000000000000000000000\nparent:      -1:0000000000000000000000000000000000000000\nmanifest:    0:c6d4cd25b9fc2f83b0dd51f4acbea9486fce54d7\nuser:        Foo Bar<foo.bar@example.com> <none@none>\ndate:        Sun Nov 11 18:33:00 2012 +0100\nfiles+:      file\nextra:       branch=default\ndescription:\none\n\nWhat is a hg exporter tool supposed to do with that?\n\nWhat such a tool can do, 'git fast-import' can do.\n\n> I am not opposed to adding a mailmap-like feature to fast-import to map\n> identities, but it has to start with sane, unambiguous output from the\n> exporter.\n\nAnd if that's not possible?\n\n-- \nFelipe Contreras\n"},{"id":"202903","messageId":"CAMP44s2ec-m0Jta2bQeg2McKAFCR6MSSP4nQx3-T=2W=3xUeyw@mail.gmail.com","threadId":"32002","inReplyTo":"509FE2EA.3020407@gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-11T17:49:49Z","receivedAt":"2012-11-11T17:49:49Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sun, Nov 11, 2012 at 6:39 PM, A Large Angry SCM <gitzilla@gmail.com> wrote:\n> On 11/11/2012 12:16 PM, Felipe Contreras wrote:\n\n>> And how do you propose to find the commit commands without parsing all\n>> the other commands? If you randomly look for lines that begin with\n>> 'commit /refs' you might end up in the middle of a commit message or\n>> the contents of a file.\n>\n> I didn't say you didn't have to parse the protocol. I said that the protocol\n> is pretty simple.\n\nParsing is never simple.\n\n>>> I agree that an additional filter _may_ decrease performance somewhat if\n>>> you\n>>> are already CPU constrained. But I suspect that the effect would be\n>>> negligible compared to the all of the SHA-1 calculations.\n>>\n>> Well. If it's so easy surely you can write one quickly, and I can measure\n>> it.\n>\n> Not my itch; You care, you do it.\n\nIt was your idea, I don't care.\n\nIf it's so simple, why don't you do it? Because it's not that simple.\nAnd anyway it will have a performance penalty.\n\n-- \nFelipe Contreras\n"},{"id":"202909","messageId":"20121111181406.GA21654@sigill.intra.peff.net","threadId":"32002","inReplyTo":"CAMP44s1mny-fBCxywM0V=AgEoxV5EZdDWc_0NK3gepcKf32nww@mail.gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-11-11T18:14:06Z","receivedAt":"2012-11-11T18:14:06Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Nov 11, 2012 at 06:45:32PM +0100, Felipe Contreras wrote:\n\n> > If there is a standard filter, then what is the advantage in doing it as\n> > a pipe? Why not just teach fast-import the same trick (and possibly make\n> > it optional)? That would be simpler, more efficient, and it would make\n> > it easier for remote helpers to turn it on (they use a command-line\n> > switch rather than setting up an extra process).\n> \n> Right, but instead of a command-line switch it probably should be\n> enabled on the stream:\n> \n>   feature clean-authors\n> \n> Or something.\n\nYeah, I was thinking it would need a feature switch to the remote helper\nto turn on the command-line, but I forgot that fast-import can take\nfeature lines directly.\n\n> > We can clean up and normalize\n> > things like whitespace (and we probably should if we do not do so\n> > already). But beyond that, we have no context about the name; only the\n> > exporter has that.\n> \n> There is no context.\n\nThere may not be a lot, but there is some:\n\n> These are exactly the same questions every exporter must answer. And\n> there's no answer, because the field is not a git author, it's a\n> mercurial user, or a bazaar committer, or who knows what.\n\nThe exporter knows that the field is a mercurial user (or whatever).\nFast-import does not even know that, and cannot apply any rules or\nheuristics about the format of a mercurial user string, what is common\nin the mercurial world, etc. It may not be a lot of context in some\ncases (I do not know anything about mercurial's formats, so I can't say\nwhat knowledge is available). But at least the exporter has a chance at\ndomain-specific interpretation of the string. Fast-import has no chance,\nbecause it does not know the domain.\n\nI've snipped the rest of your argument, which is basically that\nmercurial does not have any context at all, and knowing that it is a\nmercurial author is useless.  I am not sure that is true; even knowing\nthat it is a free-form field versus something structured (e.g., we know\nCVS authors are usernames on the server server) is useful.\n\nBut I would agree there are probably multiple systems that are like\nmercurial in that the author field is usually something like \"name\n<email>\", but may be arbitrary text (I assume bzr is the same way, but\nyou would know better than me).  So it may make sense to have some stock\nalgorithm to try to convert arbitrary almost-name-and-email text into\nname and email to reduce duplication between exporters, but:\n\n  1. It must be turned on explicitly by the exporter, since we do not\n     want to munge more structured input from clueful exporters.\n\n  2. The exporter should only turn it on after replacing its own munging\n     (e.g., it shouldn't be adding junk like <none@none>; fast-import\n     would need to receive as pristine an input as possible).\n\n  3. Exporters should not use it if they have any broken-down\n     representation at all. Even knowing that the first half is a human\n     name and the second half is something else would give it a better\n     shot at cleaning than fast-import would get.\n\n     Alternatively, the feature could enable the exporter to pass a more\n     structured ident to git.\n\n-Peff\n"},{"id":"202910","messageId":"509FEB7A.2050006@gmail.com","threadId":"32002","inReplyTo":"20121111171518.GA20115@sigill.intra.peff.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"A Large Angry SCM","fromEmail":"gitzilla@gmail.com","sentAt":"2012-11-11T18:16:26Z","receivedAt":"2012-11-11T18:16:26Z","isPatch":false,"sender":{"key":"gitzilla@gmail.com","avatar":"https://gravatar.com/avatar/354625c442439908ff3dd99757dee330e29e9df7847472384faf7a00add247fb?d=mp&s=160"},"body":"On 11/11/2012 12:15 PM, Jeff King wrote:\n> On Sun, Nov 11, 2012 at 12:00:44PM -0500, A Large Angry SCM wrote:\n>\n>>>>> a) Leave the name conversion to the export tools, and when they miss\n>>>>> some weird corner case, like 'Author<email', let the user face the\n>>>>> consequences, perhaps after an hour of the process.\n>> [...]\n>>>>> b) Do the name conversion in fast-import itself, perhaps optionally,\n>>>>> so if a tool missed some weird corner case, the user does not have to\n>>>>> face the consequences.\n>> [...]\n>>>> c) Do the name conversion, and whatever other cleanup and manipulations\n>>>> you're interesting in, in a filter between the exporter and git-fast-import.\n>>>\n>>> Such a filter would probably be quite complicated, and would decrease\n>>> performance.\n>>>\n>>\n>> Really?\n>>\n>> The fast import stream protocol is pretty simple. All the filter\n>> really needs to do is pass through everything that isn't a 'commit'\n>> command. And for the 'commit' command, it only needs to do something\n>> with the 'author' and 'committer' lines; passing through everything\n>> else.\n>>\n>> I agree that an additional filter _may_ decrease performance somewhat\n>> if you are already CPU constrained. But I suspect that the effect\n>> would be negligible compared to the all of the SHA-1 calculations.\n>\n> It might be measurable, as you are passing every byte of every version\n> of every file in the repo through an extra pipe. But more importantly, I\n> don't think it helps.\n>\n> If there is not a standard filter for fixing up names, we do not need to\n> care. The user can use \"sed\" or whatever and pay the performance penalty\n> (and deal with the possibility of errors from being lazy about parsing\n> the fast-import stream).\n>\n> If there is a standard filter, then what is the advantage in doing it as\n> a pipe? Why not just teach fast-import the same trick (and possibly make\n> it optional)? That would be simpler, more efficient, and it would make\n> it easier for remote helpers to turn it on (they use a command-line\n> switch rather than setting up an extra process).\n>\n> But what I don't understand is: what would such a standard filter look\n> like? Fast-import (or a filter) would already receive the exporter's\n> best attempt at a git-like ident string. We can clean up and normalize\n> things like whitespace (and we probably should if we do not do so\n> already). But beyond that, we have no context about the name; only the\n> exporter has that.\n>\n> So if we receive:\n>\n>    Foo Bar<foo.bar@example.com>  <none@none>\n>\n> or:\n>\n>    Foo Bar<foo.bar@example.com<none@none>\n>\n> or:\n>\n>    Foo Bar<foo.bar@example.com\n>\n> what do we do with it? Is the first part a malformed name/email pair,\n> and the second part is crap added by a lazy exporter? Or does the\n> exporter want to keep the angle brackets as part of the name field? Is\n> there a malformed email in the last one, or no email at all?\n>\n> The exporter is the only program that actually knows where the data came\n> from, how it should be broken down, and what is appropriate for pulling\n> data out of its particular source system. For that reason, the exporter\n> has to be the place where we come up with a syntactically correct and\n> unambiguous ident.\n>\n> I am not opposed to adding a mailmap-like feature to fast-import to map\n> identities, but it has to start with sane, unambiguous output from the\n> exporter.\n\nI don't think that there is or can be a standard filter. Cleaning up \nafter a broken exporter is likely to always be a repository unique \nsituation. The example here is about names and email addresses but it \ncould easily be about other things (dates, history, content, etc.). Some \nof which that could possible be fixed using git-filter-branch; some \npossibly not.\n\nFixing the exporter is always the most desirable option, but it may not \nbe the best option for the particular situation. Locally modifying \ngit-fast-import is another option; again, possibly not the best option. \nConvincing the git maintainers to handle your specific situation, though \na good option for you, is not likely to be scalable. A filter in front \nof git-fast-import is always _an_ option and can tailored to the \nparticular situation.\n\nMy preference is to follow the \"Unix philosophy\": the tools are focused \non what they need to do and can be composed with other tools/scripts to \naccomplish the desired result.\n\nd) Another (bad) option is to make git-fast-import very permissive and \nwarn the user to fix things via git-filter-branch before distributing \nthe repository or git's standard repository checks find the problems.\n\nThis isn't my itch so I think I may have exhausted my $0.02 on this subject.\n"},{"id":"202913","messageId":"CAMP44s1m8sAD9D0F-6b=+dm_AvLb_4_f7h=3A_VMYMDUEcTW7g@mail.gmail.com","threadId":"32002","inReplyTo":"20121111181406.GA21654@sigill.intra.peff.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-11T18:48:14Z","receivedAt":"2012-11-11T18:48:14Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Sun, Nov 11, 2012 at 7:14 PM, Jeff King <peff@peff.net> wrote:\n> On Sun, Nov 11, 2012 at 06:45:32PM +0100, Felipe Contreras wrote:\n>\n>> > If there is a standard filter, then what is the advantage in doing it as\n>> > a pipe? Why not just teach fast-import the same trick (and possibly make\n>> > it optional)? That would be simpler, more efficient, and it would make\n>> > it easier for remote helpers to turn it on (they use a command-line\n>> > switch rather than setting up an extra process).\n>>\n>> Right, but instead of a command-line switch it probably should be\n>> enabled on the stream:\n>>\n>>   feature clean-authors\n>>\n>> Or something.\n>\n> Yeah, I was thinking it would need a feature switch to the remote helper\n> to turn on the command-line, but I forgot that fast-import can take\n> feature lines directly.\n>\n>> > We can clean up and normalize\n>> > things like whitespace (and we probably should if we do not do so\n>> > already). But beyond that, we have no context about the name; only the\n>> > exporter has that.\n>>\n>> There is no context.\n>\n> There may not be a lot, but there is some:\n>\n>> These are exactly the same questions every exporter must answer. And\n>> there's no answer, because the field is not a git author, it's a\n>> mercurial user, or a bazaar committer, or who knows what.\n>\n> The exporter knows that the field is a mercurial user (or whatever).\n> Fast-import does not even know that, and cannot apply any rules or\n> heuristics about the format of a mercurial user string, what is common\n> in the mercurial world, etc. It may not be a lot of context in some\n> cases (I do not know anything about mercurial's formats, so I can't say\n> what knowledge is available). But at least the exporter has a chance at\n> domain-specific interpretation of the string. Fast-import has no chance,\n> because it does not know the domain.\n>\n> I've snipped the rest of your argument, which is basically that\n> mercurial does not have any context at all, and knowing that it is a\n> mercurial author is useless.  I am not sure that is true; even knowing\n> that it is a free-form field versus something structured (e.g., we know\n> CVS authors are usernames on the server server) is useful.\n\nIt is useful in the sense that we know we cannot do anything sensible\nabout it. All we can do is try.\n\n> But I would agree there are probably multiple systems that are like\n> mercurial in that the author field is usually something like \"name\n> <email>\", but may be arbitrary text (I assume bzr is the same way, but\n> you would know better than me).  So it may make sense to have some stock\n> algorithm to try to convert arbitrary almost-name-and-email text into\n> name and email to reduce duplication between exporters, but:\n\nYes, bazaar seems to be the same way.\n\n% bzr log\n------------------------------------------------------------\nrevno: 1\ncommitter: Foo Bar<foo.bar@example.com> <none@none\nbranch nick: bzr\ntimestamp: Sun 2012-11-11 19:41:10 +0100\nmessage:\n  one\n\n>   1. It must be turned on explicitly by the exporter, since we do not\n>      want to munge more structured input from clueful exporters.\n\nAgreed.\n\n>   2. The exporter should only turn it on after replacing its own munging\n>      (e.g., it shouldn't be adding junk like <none@none>; fast-import\n>      would need to receive as pristine an input as possible).\n\nAgreed.\n\n>   3. Exporters should not use it if they have any broken-down\n>      representation at all. Even knowing that the first half is a human\n>      name and the second half is something else would give it a better\n>      shot at cleaning than fast-import would get.\n\nI'm not sure what you mean by this. If they have name and email, then\nsure, it's easy.\n\nAnd for the record, I've have encountered this problem also with\nmonotone. There's quite a lot of strategies to convert names to git\nauthors.\n\nCheers.\n\n-- \nFelipe Contreras\n"},{"id":"202975","messageId":"7vwqxqiul3.fsf@alter.siamese.dyndns.org","threadId":"32002","inReplyTo":"509FD9BC.7050204@gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2012-11-12T17:45:12Z","receivedAt":"2012-11-12T17:45:12Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"A Large Angry SCM <gitzilla@gmail.com> writes:\n\n> On 11/11/2012 07:41 AM, Felipe Contreras wrote:\n>> On Sat, Nov 10, 2012 at 8:25 PM, A Large Angry SCM<gitzilla@gmail.com>  wrote:\n>>> On 11/10/2012 01:43 PM, Felipe Contreras wrote:\n>>\n>>>> So, the options are:\n>>>>\n>>>> a) Leave the name conversion to the export tools, and when they miss\n>>>> some weird corner case, like 'Author<email', let the user face the\n>>>> consequences, perhaps after an hour of the process.\n>>>>\n>>>> We know there are sources of data that don't have git-formatted author\n>>>> names, so we know every tool out there must do this checking.\n>>>>\n>>>> In addition to that, let the export tool decide what to do when one of\n>>>> these bad names appear, which in many cases probably means do nothing,\n>>>> so the user would not even see that such a bad name was there, which\n>>>> might not be what they want.\n>>>>\n>>>> b) Do the name conversion in fast-import itself, perhaps optionally,\n>>>> so if a tool missed some weird corner case, the user does not have to\n>>>> face the consequences.\n>>>>\n>>>> The tool writers don't have to worry about this, so we would not have\n>>>> tools out there doing a half-assed job of this.\n>>>>\n>>>> And what happens when such bad names end up being consistent: warning,\n>>>> a scaffold mapping of bad names, etc.\n>>>>\n>>>>\n>>>> One is bad for the users, and the tools writers, only disadvantages,\n>>>> the other is good for the users and the tools writers, only\n>>>> advantages.\n>>>>\n>>>\n>>> c) Do the name conversion, and whatever other cleanup and manipulations\n>>> you're interesting in, in a filter between the exporter and git-fast-import.\n>>\n>> Such a filter would probably be quite complicated, and would decrease\n>> performance.\n>>\n>\n> Really?\n>\n> The fast import stream protocol is pretty simple. All the filter\n> really needs to do is pass through everything that isn't a 'commit'\n> command. And for the 'commit' command, it only needs to do something\n> with the 'author' and 'committer' lines; passing through everything\n> else.\n>\n> I agree that an additional filter _may_ decrease performance somewhat\n> if you are already CPU constrained. But I suspect that the effect\n> would be negligible compared to the all of the SHA-1 calculations.\n\nMore importantly, which do users prefer: quickly produce an\nincorrect result, or spend some more time to get it right?\n\nBecause the exporting tool has a lot more intimate knowledge about\nhow the names are represented in the history of the original SCM,\ncanonicalization of the names, if done at that point, would likely\nto give us more useful results, than a canonicalization done at the\nbeginning of the importer, which lacks SCM specific details.  So in\nthat sense, (a) is more preferrable than (b).\n\nOn the other hand, we would want consistency across the converted\nresults no matter what SCM the history was originally in.  E.g. a\nname without email that came from CVS or SVN would consistently want\nto become \"name <noname@noname>\" or \"name <name>\" or whatever, and\nletting exporting tools responsible for the canonicalization will\nlead them to create their own garbage.  In that sense, (b) can be\nbetter than (a).\n\nI think (c) implements worst of both choices. It cannot exploit\nknowledge specific to the original SCM like (a) would, and while it\ncan enforce consistency the same way as (b) would, it would be a\nseparate program, unlike (b).\n\nSo...\n"},{"id":"202998","messageId":"CAMP44s0HnJvgJc=hKyJ+Jcz4oiN1QS--wPSN_3UsfVM4EWMubg@mail.gmail.com","threadId":"32002","inReplyTo":"7vwqxqiul3.fsf@alter.siamese.dyndns.org","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-12T20:46:38Z","receivedAt":"2012-11-12T20:46:38Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Mon, Nov 12, 2012 at 6:45 PM, Junio C Hamano <gitster@pobox.com> wrote:\n> A Large Angry SCM <gitzilla@gmail.com> writes:\n>\n>> On 11/11/2012 07:41 AM, Felipe Contreras wrote:\n>>> On Sat, Nov 10, 2012 at 8:25 PM, A Large Angry SCM<gitzilla@gmail.com>  wrote:\n>>>> On 11/10/2012 01:43 PM, Felipe Contreras wrote:\n>>>\n>>>>> So, the options are:\n>>>>>\n>>>>> a) Leave the name conversion to the export tools, and when they miss\n>>>>> some weird corner case, like 'Author<email', let the user face the\n>>>>> consequences, perhaps after an hour of the process.\n>>>>>\n>>>>> We know there are sources of data that don't have git-formatted author\n>>>>> names, so we know every tool out there must do this checking.\n>>>>>\n>>>>> In addition to that, let the export tool decide what to do when one of\n>>>>> these bad names appear, which in many cases probably means do nothing,\n>>>>> so the user would not even see that such a bad name was there, which\n>>>>> might not be what they want.\n>>>>>\n>>>>> b) Do the name conversion in fast-import itself, perhaps optionally,\n>>>>> so if a tool missed some weird corner case, the user does not have to\n>>>>> face the consequences.\n>>>>>\n>>>>> The tool writers don't have to worry about this, so we would not have\n>>>>> tools out there doing a half-assed job of this.\n>>>>>\n>>>>> And what happens when such bad names end up being consistent: warning,\n>>>>> a scaffold mapping of bad names, etc.\n>>>>>\n>>>>>\n>>>>> One is bad for the users, and the tools writers, only disadvantages,\n>>>>> the other is good for the users and the tools writers, only\n>>>>> advantages.\n>>>>>\n>>>>\n>>>> c) Do the name conversion, and whatever other cleanup and manipulations\n>>>> you're interesting in, in a filter between the exporter and git-fast-import.\n>>>\n>>> Such a filter would probably be quite complicated, and would decrease\n>>> performance.\n>>>\n>>\n>> Really?\n>>\n>> The fast import stream protocol is pretty simple. All the filter\n>> really needs to do is pass through everything that isn't a 'commit'\n>> command. And for the 'commit' command, it only needs to do something\n>> with the 'author' and 'committer' lines; passing through everything\n>> else.\n>>\n>> I agree that an additional filter _may_ decrease performance somewhat\n>> if you are already CPU constrained. But I suspect that the effect\n>> would be negligible compared to the all of the SHA-1 calculations.\n>\n> More importantly, which do users prefer: quickly produce an\n> incorrect result, or spend some more time to get it right?\n\nWhy not both?\n\nIf I do 'git clone hg::http://selenic.com/hg' I expect it to work, no\nmatter what. Then, if I care about getting it right, like for example\nif the project is moving to git, then check\n.git/hg/origin/bad-authors, and fill them with the right ones.\n\nOf course, the current remote helper framework doesn't have the option\nto map authors, but it could be added. That would be better than\nletting every remote helper tool to have a custom way of mapping\nauthors, and also custom configuration for them.\n\n> Because the exporting tool has a lot more intimate knowledge about\n> how the names are represented in the history of the original SCM,\n> canonicalization of the names, if done at that point, would likely\n> to give us more useful results, than a canonicalization done at the\n> beginning of the importer, which lacks SCM specific details.  So in\n> that sense, (a) is more preferrable than (b).\n\nBut it doesn't have more intimate knowledge. It has exactly the same\ninformation as fast-import; nothing.\n\nWhat intimate knowledge is a tool expected to get from this?\n\n% hg commit -u 'Foo Bar<foo.bar@example.com> <none@none>' -m one\n% hg --debug log\nchangeset:   0:5ef37a2c773f02d0e01f1ecdcc59149832d294e8\ntag:         tip\nphase:       draft\nparent:      -1:0000000000000000000000000000000000000000\nparent:      -1:0000000000000000000000000000000000000000\nmanifest:    0:c6d4cd25b9fc2f83b0dd51f4acbea9486fce54d7\nuser:        Foo Bar<foo.bar@example.com> <none@none>\ndate:        Sun Nov 11 18:33:00 2012 +0100\nfiles+:      file\nextra:       branch=default\ndescription:\none\n\nSome tools might, but if they did, then bad authors wouldn't be a problem.\n\n> On the other hand, we would want consistency across the converted\n> results no matter what SCM the history was originally in.  E.g. a\n> name without email that came from CVS or SVN would consistently want\n> to become \"name <noname@noname>\" or \"name <name>\" or whatever, and\n> letting exporting tools responsible for the canonicalization will\n> lead them to create their own garbage.  In that sense, (b) can be\n> better than (a).\n\nOr 'Unknown <unknown>' or '<none@none>' or '<>', or any of the forms\nconversion tools have been doing for ages.\n\nCheers.\n\n-- \nFelipe Contreras\n"},{"id":"203004","messageId":"20121112214127.GA10531@sigill.intra.peff.net","threadId":"32002","inReplyTo":"CAMP44s1m8sAD9D0F-6b=+dm_AvLb_4_f7h=3A_VMYMDUEcTW7g@mail.gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2012-11-12T21:41:27Z","receivedAt":"2012-11-12T21:41:27Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Sun, Nov 11, 2012 at 07:48:14PM +0100, Felipe Contreras wrote:\n\n> >   3. Exporters should not use it if they have any broken-down\n> >      representation at all. Even knowing that the first half is a human\n> >      name and the second half is something else would give it a better\n> >      shot at cleaning than fast-import would get.\n> \n> I'm not sure what you mean by this. If they have name and email, then\n> sure, it's easy.\n\nBut not as easy as just printing it. What if you have this:\n\n  name=\"Peff <angle brackets> King\"\n  email=\"<peff@peff.net>\"\n\nConcatenating them does not produce a valid git author name. Sending the\nconcatenation through fast-import's cleanup function would lose\ninformation (namely, the location of the boundary between name and\nemail).\n\nSimilarly, one might have other structured data (e.g., CVS username)\nwhere the structure is a useful hint, but some conversion to name+email\nis still necessary.\n\n-Peff\n"},{"id":"203011","messageId":"CAMP44s1gA1P-Lr1M=7RDRqFQmvQAtNnB+yAJfKC1gk3XUjbfCQ@mail.gmail.com","threadId":"32002","inReplyTo":"20121112214127.GA10531@sigill.intra.peff.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-12T22:47:55Z","receivedAt":"2012-11-12T22:47:55Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Mon, Nov 12, 2012 at 10:41 PM, Jeff King <peff@peff.net> wrote:\n> On Sun, Nov 11, 2012 at 07:48:14PM +0100, Felipe Contreras wrote:\n>\n>> >   3. Exporters should not use it if they have any broken-down\n>> >      representation at all. Even knowing that the first half is a human\n>> >      name and the second half is something else would give it a better\n>> >      shot at cleaning than fast-import would get.\n>>\n>> I'm not sure what you mean by this. If they have name and email, then\n>> sure, it's easy.\n>\n> But not as easy as just printing it. What if you have this:\n>\n>   name=\"Peff <angle brackets> King\"\n>   email=\"<peff@peff.net>\"\n>\n> Concatenating them does not produce a valid git author name. Sending the\n> concatenation through fast-import's cleanup function would lose\n> information (namely, the location of the boundary between name and\n> email).\n\nRight. Unfortunately I'm not aware of any DSCM that does that.\n\n> Similarly, one might have other structured data (e.g., CVS username)\n> where the structure is a useful hint, but some conversion to name+email\n> is still necessary.\n\nCVS might be the only one that has such structured data. I think in\nsubversion the username has no meaning. A 'felipec' subversion\nusername is as bad as a mercurial 'felipec' username.\n\nCheers.\n\n-- \nFelipe Contreras\n"},{"id":"203066","messageId":"50A21DB9.7070700@drmicha.warpmail.net","threadId":"32002","inReplyTo":"CAMP44s1gA1P-Lr1M=7RDRqFQmvQAtNnB+yAJfKC1gk3XUjbfCQ@mail.gmail.com","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Michael J Gruber","fromEmail":"git@drmicha.warpmail.net","sentAt":"2012-11-13T10:15:21Z","receivedAt":"2012-11-13T10:15:21Z","isPatch":false,"sender":{"key":"git@grubix.eu","avatar":"https://avatars.githubusercontent.com/u/233215?v=4"},"body":"Felipe Contreras venit, vidit, dixit 12.11.2012 23:47:\n> On Mon, Nov 12, 2012 at 10:41 PM, Jeff King <peff@peff.net> wrote:\n>> On Sun, Nov 11, 2012 at 07:48:14PM +0100, Felipe Contreras wrote:\n>>\n>>>>   3. Exporters should not use it if they have any broken-down\n>>>>      representation at all. Even knowing that the first half is a human\n>>>>      name and the second half is something else would give it a better\n>>>>      shot at cleaning than fast-import would get.\n>>>\n>>> I'm not sure what you mean by this. If they have name and email, then\n>>> sure, it's easy.\n>>\n>> But not as easy as just printing it. What if you have this:\n>>\n>>   name=\"Peff <angle brackets> King\"\n>>   email=\"<peff@peff.net>\"\n>>\n>> Concatenating them does not produce a valid git author name. Sending the\n>> concatenation through fast-import's cleanup function would lose\n>> information (namely, the location of the boundary between name and\n>> email).\n> \n> Right. Unfortunately I'm not aware of any DSCM that does that.\n> \n>> Similarly, one might have other structured data (e.g., CVS username)\n>> where the structure is a useful hint, but some conversion to name+email\n>> is still necessary.\n> \n> CVS might be the only one that has such structured data. I think in\n> subversion the username has no meaning. A 'felipec' subversion\n> username is as bad as a mercurial 'felipec' username.\n\nIn subversion, the username has the clearly defined meaning of being a\nusername on the subversion host. If the host is, e.g., a sourceforge\nsite then I can easily look up the user profile and convert the username\ninto a valid e-mail address (<username>@users.sf.net). That is the\nadvantage that the exporter (together with user knowledge) has over the\nimporter.\n\nIf the initial clone process aborts after every single \"unknown\" user\nit's no fun, of course. On the other hand, if an incremental clone\n(fetch) let's commits with unknown author sneak in it's no fun either\n(because I may want to fetch in crontab and publish that converted beast\nautomatically). That is why I proposed neither approach.\n\nMost conveniently, the export side of a remote helper would\n\n- do \"obvious\" automatic lossless transformations\n- use an author map for other names\n- For names not covered by the above (or having an empty map entry):\nStop exporting commits but continue parsing commits and amend the author\nmap with any unknown usernames (empty entry), and warn the user.\n(crontab script can notify me based on the return code.)\n\nIf the cloning involves a \"foreign clone\" (like the hg clone behind the\nscene) then the runtime of the second pass should be much smaller. In\nprinciple, one could even store all blobs and trees on the first run and\nskip that step on the second, but that would rely on immutability on the\nforeign side, so I dunno. (And to check the sha1, we have to get the\nblob anyways.)\n\nAs for the format for incomplete entries (foo <some@where>), a technical\nguideline should suffice for those that follow guidelines.\n\nMichael\n"},{"id":"203126","messageId":"CAMP44s18diic3KQtH5weCv-sVJXj4Pv-QnAaTeHTbrxk-=+3Gw@mail.gmail.com","threadId":"32002","inReplyTo":"50A21DB9.7070700@drmicha.warpmail.net","subject":"Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)","fromName":"Felipe Contreras","fromEmail":"felipe.contreras@gmail.com","sentAt":"2012-11-13T18:15:59Z","receivedAt":"2012-11-13T18:15:59Z","isPatch":false,"sender":{"key":"felipe.contreras@gmail.com","avatar":"https://avatars.githubusercontent.com/u/8358?v=4"},"body":"On Tue, Nov 13, 2012 at 11:15 AM, Michael J Gruber\n<git@drmicha.warpmail.net> wrote:\n> Felipe Contreras venit, vidit, dixit 12.11.2012 23:47:\n>> On Mon, Nov 12, 2012 at 10:41 PM, Jeff King <peff@peff.net> wrote:\n>>> On Sun, Nov 11, 2012 at 07:48:14PM +0100, Felipe Contreras wrote:\n>>>\n>>>>>   3. Exporters should not use it if they have any broken-down\n>>>>>      representation at all. Even knowing that the first half is a human\n>>>>>      name and the second half is something else would give it a better\n>>>>>      shot at cleaning than fast-import would get.\n>>>>\n>>>> I'm not sure what you mean by this. If they have name and email, then\n>>>> sure, it's easy.\n>>>\n>>> But not as easy as just printing it. What if you have this:\n>>>\n>>>   name=\"Peff <angle brackets> King\"\n>>>   email=\"<peff@peff.net>\"\n>>>\n>>> Concatenating them does not produce a valid git author name. Sending the\n>>> concatenation through fast-import's cleanup function would lose\n>>> information (namely, the location of the boundary between name and\n>>> email).\n>>\n>> Right. Unfortunately I'm not aware of any DSCM that does that.\n>>\n>>> Similarly, one might have other structured data (e.g., CVS username)\n>>> where the structure is a useful hint, but some conversion to name+email\n>>> is still necessary.\n>>\n>> CVS might be the only one that has such structured data. I think in\n>> subversion the username has no meaning. A 'felipec' subversion\n>> username is as bad as a mercurial 'felipec' username.\n>\n> In subversion, the username has the clearly defined meaning of being a\n> username on the subversion host. If the host is, e.g., a sourceforge\n> site then I can easily look up the user profile and convert the username\n> into a valid e-mail address (<username>@users.sf.net). That is the\n> advantage that the exporter (together with user knowledge) has over the\n> importer.\n>\n> If the initial clone process aborts after every single \"unknown\" user\n> it's no fun, of course. On the other hand, if an incremental clone\n> (fetch) let's commits with unknown author sneak in it's no fun either\n> (because I may want to fetch in crontab and publish that converted beast\n> automatically). That is why I proposed neither approach.\n>\n> Most conveniently, the export side of a remote helper would\n>\n> - do \"obvious\" automatic lossless transformations\n> - use an author map for other names\n\nThis should be done by fast-import. It doesn't make any sense that\nevery remote helper and fast-exporter out there have their own way of\nmapping authors (or none).\n\n> - For names not covered by the above (or having an empty map entry):\n> Stop exporting commits but continue parsing commits and amend the author\n> map with any unknown usernames (empty entry), and warn the user.\n> (crontab script can notify me based on the return code.)\n\nStop exporting commits but continue parsing commits? I don't know what\nthat means.\n\nfast-import should try it's best to clean it up, warn the user, sure,\nbut also store the missing entry on a file, so that it can be filed\nlater (if the user so wishes).\n\n> If the cloning involves a \"foreign clone\" (like the hg clone behind the\n> scene) then the runtime of the second pass should be much smaller. In\n> principle, one could even store all blobs and trees on the first run and\n> skip that step on the second, but that would rely on immutability on the\n> foreign side, so I dunno. (And to check the sha1, we have to get the\n> blob anyways.)\n\nNo. There's no concept of partial clones... Either you clone, or you don't.\n\nWait if the remote helper didn't notice that the author was bad?\nfast-import could just just leave everything up to that point, warn\nabut what happened, and exit, but the exporter side would die in the\nmiddle of exporting, and it might end up in a bad state, not saving\nmarks, or who knows what.\n\nIt wouldn't work.\n\nThe cloning should be full, and the bad authors stored in a scaffold author map.\n\n> As for the format for incomplete entries (foo <some@where>), a technical\n> guideline should suffice for those that follow guidelines.\n\nfast-import should do that.\n\nCheers.\n\n-- \nFelipe Contreras\n"}]}