git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Problem with git-cvsimport

From
Michael Haggerty <mhagger@alum.mit.edu>
Date
Oct 31, 2007, 04:42 UTC
Message-ID
<472807A1.8030804@alum.mit.edu>
In-Reply-To
<170fa0d20710301306o6b3798f9k72615eb811d871f2@mail.gmail.com>
Mike Snitzer wrote:
Show 7 quoted lines
> On 10/10/07, Eyvind Bernhardsen <eyvind-git-list@orakel.ntnu.no> wrote:
> ...
>> Thanks for making cvs2svn the best CVS-to-git conversion tool :)  Now
>> if it would only support incremental importing...
> 
> I second this question: is there any chance incremental importing will
> be implemented in cvs2svn?
Unfortunately, no, there is not much chance that I will implement this.
 I wouldn't be interested in a works-most-of-the-time solution, and a
reliable solution would take weeks to implement.

If somebody else wants to implement this feature, I would be happy to help him get started, answer questions, discuss the design, etc. Or if somebody wants to sponsor the work, I might be able to justify working on it myself. But otherwise, I'm afraid it is unlikely to happen.

> I've not used cvs2svn much and when I did it was for svn not git; but
> given that git-cvsimport is known to mess up your git repo (as Eyvind
> pointed out earlier) there doesn't appear to be any reliable tools to
> allow for incrementally importing from cvs to git.

That's because it is quite a tricky problem, especially since CVS allows history to be changed retroactively; for example,

- shift a tag to a different file revision
- add an existing tag to a new file or remove it from an old file
- delete ("obsolete") old revisions
- change files from vendor branches to main line of development
- even nastier server-side repository manipulations like deleting an RCS
file, renaming a file, etc.

These things really happen in the topsy-turvy CVS world; indeed, they are a part of many organizations' standard workflow.

cvs2svn uses repository-wide information in the heuristics that it uses to determine changesets, choose branch parents, fix clock skew, etc. Therefore the naive approach of running a full conversion a second time and just skipping over the revisions that were handled during the first conversion would not even begin to work. (I believe that this is the approach of cvsps, which uses mostly local information to determine changesets.)

I think the correct approach would involve recording the "frontier" of the CVS repository, then at the next incremental conversion:

1. compare the current CVS repository to the recorded information
2. emit "fixup" changesets to reflect any CVS changes that happened
behind the previous "frontier".
3. emit changesets to reflect CVS changes beyond the frontier.

It is step 2 that is IMO the trickiest because it is so open-ended, and modern SCMs don't allow all of the corresponding operations in any straightforward way. Presumably one would have to prohibit some of the nastier CVS tricks and abort the incremental conversion if any are detected.

Furthermore, for many use-cases of incremental conversion the conversion would have to run quickly. Therefore, the incremental conversion code should be written with a strong emphasis on achieving good performance.

Michael
Previous: Aidan Van Dyk
Message 11 of 11 in “Problem with git-cvsimport”
  1. Thomas PaschOct 9, 2007
  2. Jan WielemakerOct 9, 2007
  3. Gerald (Jerry) CarterOct 9, 2007
  4. Eyvind BernhardsenOct 9, 2007
  5. Michael HaggertyOct 10, 2007
  6. Eyvind BernhardsenOct 10, 2007
  7. Mike SnitzerOct 30, 2007
  8. Mike SnitzerOct 30, 2007
  9. Robin RosenbergOct 30, 2007
  10. Aidan Van DykOct 31, 2007
  11. Michael HaggertyOct 31, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.