git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: cvs import

From
Shawn Pearce <spearce@spearce.org>
Date
Sep 16, 2006, 03:39 UTC
Message-ID
<20060916033917.GA24269@spearce.org>
In-Reply-To
<450A581E.2050509@bluegap.ch>
Markus Schiltknecht <markus@bluegap.ch> wrote:
Show 20 quoted lines
> Shawn Pearce wrote:
> >I don't know how the Monotone guys feel about it but I think Git
> >is happy with the data in any order, just so long as the dependency
> >chains aren't fed out of order.  Which I think nearly all changeset
> >based SCMs would have an issue with.  So we should be just fine
> >with the current chronological order produced by cvs2svn.
> 
> I'd vote for splitting into file data (and delta / patches) import and 
> metadata import (author, changelog, DAG).
> 
> Monotone would be happiest if the file data were sent one file after 
> another and (inside each file) in the order of each file's single 
> history. That guarantees good import performance for monotone. I imagine 
> it's about the same for git. And if you have to somehow cache the files 
> anyway, subversion will benefit, too. (Well, at least the cache will 
> thank us with good performance).
>
> After all file data has been delivered, the metadata can be delivered. 
> As neigther monotone nor git care much if they are chronological across 
> branches, I'd vote for doing it that way.

Right. I think that one of the cvs2svn guys had the right idea here. Provide two hooks: one early during the RCS file parse which supplies a backend each full text file revision and another during the very last stage which includes the "file" in the metadata stream for commit.

This would give Git and Monotone a way to grab the full text for each file and stream them out up front, then include only a "token" in the metadata stream which identifies the specific revision. Meanwhile SVN can either cache the file revision during the early part or ignore it, then dump out the full content during the metadata.

As it happens Git doesn't care what order the file revisions come in. If we don't repack the imported data we would prefer to get the revisions in newest->oldest order so we can delta the older versions against the newer versions (like RCS). This is also happens to be the fastest way to extract the revision data from RCS.

On the other hand from what I understand of Monotone it needs the revisions in oldest->newest order, as does SVN.

Doing both orderings in cvs2noncvs is probably ugly. Doing just oldest->newest (since 2/3 backends want that) would be acceptable but would slow down Git imports as the RCS parsing overhead would be much higher.

-- 
Shawn.
Previous: Markus SchiltknechtNext: Oswald Buddenhagen
Message 24 of 38 in “Re: cvs import”
  1. Jon SmirlSep 13, 2006
  2. Martin LanghoffSep 13, 2006
  3. Markus SchiltknechtSep 13, 2006
  4. Oswald BuddenhagenSep 13, 2006
  5. Martin LanghoffSep 13, 2006
  6. Michael HaggertySep 14, 2006
  7. Jon SmirlSep 14, 2006
  8. Michael HaggertySep 14, 2006
  9. Martin LanghoffSep 14, 2006
  10. Michael HaggertySep 14, 2006
  11. Jon SmirlSep 14, 2006
  12. Martin LanghoffSep 14, 2006
  13. Markus SchiltknechtSep 13, 2006
  14. Jon SmirlSep 13, 2006
  15. Michael HaggertySep 14, 2006
  16. Shawn PearceSep 14, 2006
  17. Jakub NarebskiSep 14, 2006
  18. Shawn PearceSep 14, 2006
  19. Jon SmirlSep 14, 2006
  20. Michael HaggertySep 14, 2006
  21. Jakub NarebskiSep 14, 2006
  22. Jon SmirlSep 14, 2006
  23. Markus SchiltknechtSep 15, 2006
  24. Shawn PearceSep 16, 2006
  25. Oswald BuddenhagenSep 16, 2006
  26. Nathaniel SmithSep 16, 2006
  27. Nathaniel SmithSep 13, 2006
  28. Daniel CarosoneSep 13, 2006
  29. Daniel CarosoneSep 13, 2006
  30. Keith PackardSep 13, 2006
  31. Nathaniel SmithSep 14, 2006
  32. Jon SmirlSep 14, 2006
  33. Daniel CarosoneSep 14, 2006
  34. Shawn PearceSep 14, 2006
  35. Daniel CarosoneSep 14, 2006
  36. Petr BaudisSep 14, 2006
  37. Shawn PearceSep 14, 2006
  38. Shawn PearceSep 14, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.