git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)

From
Felipe Contreras <felipe.contreras@gmail.com>
Date
Nov 12, 2012, 20:46 UTC
Message-ID
<CAMP44s0HnJvgJc=hKyJ+Jcz4oiN1QS--wPSN_3UsfVM4EWMubg@mail.gmail.com>
In-Reply-To
<7vwqxqiul3.fsf@alter.siamese.dyndns.org>
On Mon, Nov 12, 2012 at 6:45 PM, Junio C Hamano <gitster@pobox.com> wrote:
Show 57 quoted lines
> A Large Angry SCM <gitzilla@gmail.com> writes:
>
>> On 11/11/2012 07:41 AM, Felipe Contreras wrote:
>>> On Sat, Nov 10, 2012 at 8:25 PM, A Large Angry SCM<gitzilla@gmail.com>  wrote:
>>>> On 11/10/2012 01:43 PM, Felipe Contreras wrote:
>>>
>>>>> So, the options are:
>>>>>
>>>>> a) Leave the name conversion to the export tools, and when they miss
>>>>> some weird corner case, like 'Author<email', let the user face the
>>>>> consequences, perhaps after an hour of the process.
>>>>>
>>>>> We know there are sources of data that don't have git-formatted author
>>>>> names, so we know every tool out there must do this checking.
>>>>>
>>>>> In addition to that, let the export tool decide what to do when one of
>>>>> these bad names appear, which in many cases probably means do nothing,
>>>>> so the user would not even see that such a bad name was there, which
>>>>> might not be what they want.
>>>>>
>>>>> b) Do the name conversion in fast-import itself, perhaps optionally,
>>>>> so if a tool missed some weird corner case, the user does not have to
>>>>> face the consequences.
>>>>>
>>>>> The tool writers don't have to worry about this, so we would not have
>>>>> tools out there doing a half-assed job of this.
>>>>>
>>>>> And what happens when such bad names end up being consistent: warning,
>>>>> a scaffold mapping of bad names, etc.
>>>>>
>>>>>
>>>>> One is bad for the users, and the tools writers, only disadvantages,
>>>>> the other is good for the users and the tools writers, only
>>>>> advantages.
>>>>>
>>>>
>>>> c) Do the name conversion, and whatever other cleanup and manipulations
>>>> you're interesting in, in a filter between the exporter and git-fast-import.
>>>
>>> Such a filter would probably be quite complicated, and would decrease
>>> performance.
>>>
>>
>> Really?
>>
>> The fast import stream protocol is pretty simple. All the filter
>> really needs to do is pass through everything that isn't a 'commit'
>> command. And for the 'commit' command, it only needs to do something
>> with the 'author' and 'committer' lines; passing through everything
>> else.
>>
>> I agree that an additional filter _may_ decrease performance somewhat
>> if you are already CPU constrained. But I suspect that the effect
>> would be negligible compared to the all of the SHA-1 calculations.
>
> More importantly, which do users prefer: quickly produce an
> incorrect result, or spend some more time to get it right?
Why not both?

If I do 'git clone hg::http://selenic.com/hg' I expect it to work, no matter what. Then, if I care about getting it right, like for example if the project is moving to git, then check .git/hg/origin/bad-authors, and fill them with the right ones.

Of course, the current remote helper framework doesn't have the option to map authors, but it could be added. That would be better than letting every remote helper tool to have a custom way of mapping authors, and also custom configuration for them.

Show 6 quoted lines
> Because the exporting tool has a lot more intimate knowledge about
> how the names are represented in the history of the original SCM,
> canonicalization of the names, if done at that point, would likely
> to give us more useful results, than a canonicalization done at the
> beginning of the importer, which lacks SCM specific details.  So in
> that sense, (a) is more preferrable than (b).

But it doesn't have more intimate knowledge. It has exactly the same information as fast-import; nothing.

What intimate knowledge is a tool expected to get from this?

% hg commit -u 'Foo Bar<foo.bar@example.com> <none@none>' -m one % hg --debug log changeset: 0:5ef37a2c773f02d0e01f1ecdcc59149832d294e8 tag: tip phase: draft parent: -1:0000000000000000000000000000000000000000 parent: -1:0000000000000000000000000000000000000000 manifest: 0:c6d4cd25b9fc2f83b0dd51f4acbea9486fce54d7 user: Foo Bar<foo.bar@example.com> <none@none> date: Sun Nov 11 18:33:00 2012 +0100 files+: file extra: branch=default description: one

Some tools might, but if they did, then bad authors wouldn't be a problem.
Show 7 quoted lines
> On the other hand, we would want consistency across the converted
> results no matter what SCM the history was originally in.  E.g. a
> name without email that came from CVS or SVN would consistently want
> to become "name <noname@noname>" or "name <name>" or whatever, and
> letting exporting tools responsible for the canonicalization will
> lead them to create their own garbage.  In that sense, (b) can be
> better than (a).

Or 'Unknown <unknown>' or '<none@none>' or '<>', or any of the forms conversion tools have been doing for ages.

Cheers.
-- 
Felipe Contreras
Previous: Junio C Hamano
Message 24 of 24 in “RFD: fast-import is picky with author names (and maybe it should - but how much so?)”
  1. Michael J GruberNov 2, 2012
  2. Michael J GruberNov 2, 2012
  3. Jeff KingNov 8, 2012
  4. Michael J GruberNov 9, 2012
  5. Felipe ContrerasNov 9, 2012
  6. Michael J GruberNov 10, 2012
  7. Felipe ContrerasNov 10, 2012
  8. A Large Angry SCMNov 10, 2012
  9. Felipe ContrerasNov 11, 2012
  10. A Large Angry SCMNov 11, 2012
  11. Jeff KingNov 11, 2012
  12. Felipe ContrerasNov 11, 2012
  13. Jeff KingNov 11, 2012
  14. Felipe ContrerasNov 11, 2012
  15. Jeff KingNov 12, 2012
  16. Felipe ContrerasNov 12, 2012
  17. Michael J GruberNov 13, 2012
  18. Felipe ContrerasNov 13, 2012
  19. A Large Angry SCMNov 11, 2012
  20. Felipe ContrerasNov 11, 2012
  21. A Large Angry SCMNov 11, 2012
  22. Felipe ContrerasNov 11, 2012
  23. Junio C HamanoNov 12, 2012
  24. Felipe ContrerasNov 12, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.