git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: RFD: fast-import is picky with author names (and maybe it should - but how much so?)

From
Felipe Contreras <felipe.contreras@gmail.com>
Date
Nov 11, 2012, 18:48 UTC
Message-ID
<CAMP44s1m8sAD9D0F-6b=+dm_AvLb_4_f7h=3A_VMYMDUEcTW7g@mail.gmail.com>
In-Reply-To
<20121111181406.GA21654@sigill.intra.peff.net>
On Sun, Nov 11, 2012 at 7:14 PM, Jeff King <peff@peff.net> wrote:
Show 46 quoted lines
> On Sun, Nov 11, 2012 at 06:45:32PM +0100, Felipe Contreras wrote:
>
>> > If there is a standard filter, then what is the advantage in doing it as
>> > a pipe? Why not just teach fast-import the same trick (and possibly make
>> > it optional)? That would be simpler, more efficient, and it would make
>> > it easier for remote helpers to turn it on (they use a command-line
>> > switch rather than setting up an extra process).
>>
>> Right, but instead of a command-line switch it probably should be
>> enabled on the stream:
>>
>>   feature clean-authors
>>
>> Or something.
>
> Yeah, I was thinking it would need a feature switch to the remote helper
> to turn on the command-line, but I forgot that fast-import can take
> feature lines directly.
>
>> > We can clean up and normalize
>> > things like whitespace (and we probably should if we do not do so
>> > already). But beyond that, we have no context about the name; only the
>> > exporter has that.
>>
>> There is no context.
>
> There may not be a lot, but there is some:
>
>> These are exactly the same questions every exporter must answer. And
>> there's no answer, because the field is not a git author, it's a
>> mercurial user, or a bazaar committer, or who knows what.
>
> The exporter knows that the field is a mercurial user (or whatever).
> Fast-import does not even know that, and cannot apply any rules or
> heuristics about the format of a mercurial user string, what is common
> in the mercurial world, etc. It may not be a lot of context in some
> cases (I do not know anything about mercurial's formats, so I can't say
> what knowledge is available). But at least the exporter has a chance at
> domain-specific interpretation of the string. Fast-import has no chance,
> because it does not know the domain.
>
> I've snipped the rest of your argument, which is basically that
> mercurial does not have any context at all, and knowing that it is a
> mercurial author is useless.  I am not sure that is true; even knowing
> that it is a free-form field versus something structured (e.g., we know
> CVS authors are usernames on the server server) is useful.

It is useful in the sense that we know we cannot do anything sensible about it. All we can do is try.

Show 6 quoted lines
> But I would agree there are probably multiple systems that are like
> mercurial in that the author field is usually something like "name
> <email>", but may be arbitrary text (I assume bzr is the same way, but
> you would know better than me).  So it may make sense to have some stock
> algorithm to try to convert arbitrary almost-name-and-email text into
> name and email to reduce duplication between exporters, but:
Yes, bazaar seems to be the same way.
% bzr log
------------------------------------------------------------
revno: 1
committer: Foo Bar<foo.bar@example.com> <none@none
branch nick: bzr
timestamp: Sun 2012-11-11 19:41:10 +0100
message:
  one
>   1. It must be turned on explicitly by the exporter, since we do not
>      want to munge more structured input from clueful exporters.
Agreed.
>   2. The exporter should only turn it on after replacing its own munging
>      (e.g., it shouldn't be adding junk like <none@none>; fast-import
>      would need to receive as pristine an input as possible).
Agreed.
>   3. Exporters should not use it if they have any broken-down
>      representation at all. Even knowing that the first half is a human
>      name and the second half is something else would give it a better
>      shot at cleaning than fast-import would get.

I'm not sure what you mean by this. If they have name and email, then sure, it's easy.

And for the record, I've have encountered this problem also with monotone. There's quite a lot of strategies to convert names to git authors.

Cheers.
-- 
Felipe Contreras
Previous: Jeff KingNext: Jeff King
Message 14 of 24 in “RFD: fast-import is picky with author names (and maybe it should - but how much so?)”
  1. Michael J GruberNov 2, 2012
  2. Michael J GruberNov 2, 2012
  3. Jeff KingNov 8, 2012
  4. Michael J GruberNov 9, 2012
  5. Felipe ContrerasNov 9, 2012
  6. Michael J GruberNov 10, 2012
  7. Felipe ContrerasNov 10, 2012
  8. A Large Angry SCMNov 10, 2012
  9. Felipe ContrerasNov 11, 2012
  10. A Large Angry SCMNov 11, 2012
  11. Jeff KingNov 11, 2012
  12. Felipe ContrerasNov 11, 2012
  13. Jeff KingNov 11, 2012
  14. Felipe ContrerasNov 11, 2012
  15. Jeff KingNov 12, 2012
  16. Felipe ContrerasNov 12, 2012
  17. Michael J GruberNov 13, 2012
  18. Felipe ContrerasNov 13, 2012
  19. A Large Angry SCMNov 11, 2012
  20. Felipe ContrerasNov 11, 2012
  21. A Large Angry SCMNov 11, 2012
  22. Felipe ContrerasNov 11, 2012
  23. Junio C HamanoNov 12, 2012
  24. Felipe ContrerasNov 12, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.