git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Some tips for doing a CVS importer

From
Jon Smirl <jonsmirl@gmail.com>
Date
Nov 20, 2006, 21:49 UTC
Message-ID
<9e4733910611201349s4d08b984g772c64982f148bfa@mail.gmail.com>

I have tried all of the available CVS importers. None of them are without problems. If anyone is interested in writing one for git here are some ideas on how to structure it.

1) there is a working lex/yacc for CVS in the parsecvs source code
2) The first time you parse a CVS file record everything and don't
parse it again.
3) When the file is first parsed use the deltas to generate the
revisions and feed them to git-fastimport, just remember the SHA1 or
an id in the import code. This is a critical step to getting decent
performance.
4) If you do #1 and #2 you don't need to store CVS revision numbers
and file names in memory. Because of that you can can easily do a
Mozilla import in 2GB, probably 1GB.
5) When comparing CVS revisions only use the CVS timestamps as a last
resort, instead use the dependency information in the CVS file
6) Match up commits by using an sha1 of the author and commit message
7) After all files are loaded, match up the symbols and insert them
into the dependency chains, if any of the symbols depend on a branch
commit the symbol lies on the branch, otherwise the symbol is on the
trunk,
8) Do a topological sort to build the change set commit tree
9) when you hit a loop in the tree break up delta change sets until
the loop can be removed, don't break up symbol change sets.
10) Mozilla has some large commits that were made over dial up. Commit
change sets can span hours. All of these commits need to be merged
into a single change set.
11) An algorithm needs to be developed for detecting branches merging
back into the trunk
12) cvs2svn has excellent test cases, use them to test the new
importer. The cvs2svn code is quite nice but it doesn't handle #7
-- 
Jon Smirl
Next: Martin Langhoff
Message 1 of 30 in “Some tips for doing a CVS importer”
  1. Jon SmirlNov 20, 2006
  2. Martin LanghoffNov 20, 2006
  3. Jon SmirlNov 20, 2006
  4. Martin LanghoffNov 21, 2006
  5. Carl WorthNov 21, 2006
  6. Jon SmirlNov 21, 2006
  7. Shawn PearceNov 21, 2006
  8. lamikrNov 21, 2006
  9. Shawn PearceNov 21, 2006
  10. Robin RosenbergNov 23, 2006
  11. Shawn PearceNov 25, 2006
  12. Petr BaudisNov 21, 2006
  13. Shawn PearceNov 21, 2006
  14. Johannes SchindelinNov 21, 2006
  15. Johannes SixtNov 23, 2006
  16. Martin LanghoffNov 21, 2006
  17. Jon SmirlNov 21, 2006
  18. Marko MacekNov 26, 2006
  19. Jon SmirlNov 26, 2006
  20. Marko MacekNov 26, 2006
  21. Jon SmirlNov 26, 2006
  22. Michael HaggertyNov 27, 2006
  23. Shawn PearceNov 21, 2006
  24. Michael HaggertyNov 27, 2006
  25. Markus SchiltknechtNov 27, 2006
  26. Michael HaggertyNov 27, 2006
  27. Markus SchiltknechtNov 28, 2006
  28. Michael HaggertyNov 30, 2006
  29. Daniel JacobowitzNov 30, 2006
  30. Jon SmirlNov 27, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.