git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: cvs2svn conversion directly to git ready for experimentation

From
Jon Smirl <jonsmirl@gmail.com>
Date
Aug 2, 2007, 23:44 UTC
Message-ID
<9e4733910708021644q6eba0e78gc2c6bcfba4816012@mail.gmail.com>
In-Reply-To
<f8r09t$qdg$1@sea.gmane.org>
On 8/1/07, Jakub Narebski <jnareb@gmail.com> wrote:
Show 13 quoted lines
> Michael Haggerty wrote:
>
> > I am the maintainer of cvs2svn[1], which is a program for one-time
> > conversions from CVS to Subversion. cvs2svn is very robust against the
> > many peculiarities of CVS and can convert just about every CVS
> > repository we have ever seen.
> >
> > I've been working on a cvs2svn output pass that writes the converted CVS
> > repository directly into git rather than Subversion. The code runs now
> > with at least one repository from our test suite of nasty CVS repositories.
>
> Have you contacted Jon Smirl about his unpublished work on cvs2git,
> cvs2svn based CVS to Git converter?

My converter was derived from Michael's cvs2svn code. The bulk of my work was converting cvs2svn to output in a format that git-fastimport could consume. This was all rather straight forward and there was nothing really interesting in the code.

What it exposed were fundamental issues about the technical complexities of trying to reconstruct a change set history from CVS which didn't record all of the needed info. I was never able to construct a satisfactory git representation of the Mozilla CVS repository. Michael has had a long time to work on the change set detection code and he's probably added some new strategies.

My code did include a CVS file parser for extracting all the revisions from the file in a single pass. Doing that is a major performance benefit. I believe I posted the code to the cvs2svn mailing list. It was about 200 lines of code. Forking off cvs a million times to extract the revisions takes days to run.

Same goes for forking git a million times.git-fastimport uses a pipe to cvs2svn to avoid forking. git-fastimport also uses a technique from the database world for bulk import, it imports everything without indexing it. Indexing is done after the import finishes.

Between parsing the CVS files internally and Shawn's git-fastimport, it was possible to import Mozilla CVS (2.4G) in about 2 hours and generate a 450MB pack file. You need 3GB of RAM to do this - if swap happens the process will take weeks to finish.

Show 38 quoted lines
> Quote from InterfacesFrontendsAndTools page on GIT wiki[1]:
>
>   cvs2git is the unofficial name of Jon Smirl's modifications to cvs2svn.
>   These modifications allow cvs2svn to generate a data stream which is
>   consumed by Shawn Pearce's git-fast-import (now included in git.git).
>   git-fast-import converts its input stream directly into a Git .pack file,
>   minimizing the amount of IO required on large imports.
>
>   Jon Smirl stopped working on cvs2git[2] because first, Mozilla (which was
>   main target of his work) decided that to not to move to git, and second
>   because of troubles with cvs2svn architecture[*] (which it is based on).
>   Jon Smirl has posted his impressions on working on CVS importer in
>   "Some tips for doing a CVS importer" thread[3].
>
> References:
> -----------
> [1] http://git.or.cz/gitwiki/InterfacesFrontendsAndTools#head-23858c2cde0cef60443d8e73e6829a95f8e191ef
> [2] http://msgid.gmane.org/9e4733910611190940y147992b8mbdfac5a51f42e0fe@mail.gmail.com
> [3] http://marc.theaimsgroup.com/?t=116405956000001&r=1&w=2
>
> Footnotes:
> ----------
> [*] If I remember correctly authors of cvs2svn were talking about separating
> the code dealing with disentangling CVS repository structure from the part
> translating it into Subversion repository (with its quirks), and the part
> generating Subversion repository.
>
> --
> Jakub Narebski
> Warsaw, Poland
> ShadeHawk on #git
>
>
> -
> To unsubscribe from this list: send the line "unsubscribe git" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>
-- 
Jon Smirl
jonsmirl@gmail.com
Previous: Michael HaggertyNext: Steffen Prohaska
Message 5 of 40 in “cvs2svn conversion directly to git ready for experimentation”
  1. Michael HaggertyAug 1, 2007
  2. Johannes SchindelinAug 1, 2007
  3. Jakub NarebskiAug 1, 2007
  4. Michael HaggertyAug 2, 2007
  5. Jon SmirlAug 2, 2007
  6. Steffen ProhaskaAug 2, 2007
  7. Michael HaggertyAug 2, 2007
  8. Marko MacekAug 2, 2007
  9. Jon SmirlAug 2, 2007
  10. Oswald BuddenhagenAug 5, 2007
  11. Simon 'corecode' SchubertAug 2, 2007
  12. Steffen ProhaskaAug 2, 2007
  13. Simon 'corecode' SchubertAug 2, 2007
  14. Robin RosenbergAug 2, 2007
  15. Lübbe OnkenAug 2, 2007
  16. Lübbe OnkenAug 2, 2007
  17. Steffen ProhaskaAug 2, 2007
  18. Simon 'corecode' SchubertAug 2, 2007
  19. Michael HaggertyAug 2, 2007
  20. Simon 'corecode' SchubertAug 3, 2007
  21. Steffen ProhaskaAug 4, 2007
  22. Shawn O. PearceAug 3, 2007
  23. Michael HaggertyAug 2, 2007
  24. Linus TorvaldsAug 2, 2007
  25. Michael HaggertyAug 2, 2007
  26. Shawn O. PearceAug 3, 2007
  27. Jon SmirlAug 2, 2007
  28. Michael HaggertyAug 2, 2007
  29. Martin LanghoffAug 2, 2007
  30. Johannes SchindelinAug 3, 2007
  31. Steffen ProhaskaAug 3, 2007
  32. Steffen ProhaskaAug 3, 2007
  33. Michael HaggertyAug 3, 2007
  34. Patwardhan, RajeshAug 3, 2007
  35. Jon SmirlAug 3, 2007
  36. Patwardhan, RajeshAug 3, 2007
  37. Michael HaggertyAug 3, 2007
  38. Jon SmirlAug 3, 2007
  39. Jon SmirlAug 3, 2007
  40. Lübbe OnkenAug 2, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.