{"thread":{"id":"9022","subject":"CVS -> SVN -> Git","startedAt":"2007-07-13T14:48:40Z","lastAt":"2007-07-20T08:45:50Z","messageCount":31,"participants":["Julian Phillips","Michael Haggerty","Martin Langhoff","Chris Shoemaker","Steffen Prohaska","Eric S. Raymond","Junio C Hamano","Oswald Buddenhagen","Karl Fogel","David Frech","Shawn O. Pearce","Scott Lamb","Markus Schiltknecht","Simon 'corecode' Schubert"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"47269","messageId":"Pine.LNX.4.64.0707131541140.11423@reaper.quantumfyre.co.uk","threadId":"9022","inReplyTo":null,"subject":"CVS -> SVN -> Git","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2007-07-13T14:48:40Z","receivedAt":"2007-07-13T14:48:40Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"\nHas anyone managed to succssfully import a Subversion repository that was \ninitially imported from CVS using cvs2svn using fast-import?\n\nIt looks like cvs2svn has created a rather big mess.   It has created \nsingle commits that change files in more than one branch and/or tag. \nIt also creates tags using more than one commit.  Now I come to try and \nimport the Subversion history into git and I'm having trouble creating a \nsensible stream to feed into fast-import.\n\nI'm trying to use fast-import because git-svnimport creates a incorrect \nrepository that is missing files and even whole directories (I suppose \nthis could be due to the confusion from cvs2svn), and git-svn is a) _way_ \ntoo slow, b) doesn't do merges and c) munges the commit comments.\n\n-- \nJulian\n\n  ---\nRiffle West Virginia is so small that the Boy Scout had to double as the\ntown drunk.\n"},{"id":"47296","messageId":"469804B4.1040509@alum.mit.edu","threadId":"9022","inReplyTo":"Pine.LNX.4.64.0707131541140.11423@reaper.quantumfyre.co.uk","subject":"Re: CVS -> SVN -> Git","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-07-13T23:03:16Z","receivedAt":"2007-07-13T23:03:16Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Julian Phillips wrote:\n> Has anyone managed to succssfully import a Subversion repository that\n> was initially imported from CVS using cvs2svn using fast-import?\n> \n> It looks like cvs2svn has created a rather big mess.   It has created\n> single commits that change files in more than one branch and/or tag. It\n> also creates tags using more than one commit.  Now I come to try and\n> import the Subversion history into git and I'm having trouble creating a\n> sensible stream to feed into fast-import.\n\nI'm the main cvs2svn developer.  Obviously, the tool is intended to\nconvert to Subversion, but there are ways to tune it to make its output\na little bit more git-friendly.\n\n[Please note that both CVS and SVN allow changes to multiple\ntags/branches in a single commit and creating tags using more than one\ncommit.  That is why cvs2svn converts these repository \"features\" 1:1 by\ndefault.]\n\nRelease 2.0.0-rc1 of cvs2svn (released today) has a\n--no-cross-branch-commits option that prevents commits that affect more\nthan one branch.  For multiproject conversions, the\n\"ctx.cross_project_commits\" option might also be useful.  (The latter is\nonly available if you start cvs2svn with an --options file.)\n\nThe new cvs2svn release is also more intelligent about determining the\nmost likely source branch from which a tag/branch was created.  This\ndoes not eliminate the creation of tags from more than one revision, but\nit should reduce its frequency.  If your repository uses any vendor\nbranches, you might also consider --exclude'ing them.  In the new\ncvs2svn version, this causes vendor revisions to be grafted onto trunk\nand thereby eliminates another common cause of multiple-source\nbranches/tags.\n\nIncidentally, now that cvs2svn 2.0.0 is nearly out, I am thinking about\nwhat it would take to write some other back ends for cvs2svn--turning\nit, essentially, into cvs2xxx.  Most of the work that cvs2svn does is\ninferring the most plausible history of the repository from CVS's\nsketchy, incomplete, idiomatic, and often corrupt data.  This work\nshould also be useful for a cvs2git or cvs2hg or cvs2baz or ...\n\nI haven't played with a distributed SCM yet, but if somebody would be\ninterested in working with me on this please let me know.\n\nMichael\n"},{"id":"47313","messageId":"46a038f90707132230n120e6392uaf5cd86ff10b6012@mail.gmail.com","threadId":"9022","inReplyTo":"469804B4.1040509@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-07-14T05:30:33Z","receivedAt":"2007-07-14T05:30:33Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 7/14/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> Incidentally, now that cvs2svn 2.0.0 is nearly out, I am thinking about\n> what it would take to write some other back ends for cvs2svn--turning\n> it, essentially, into cvs2xxx.  Most of the work that cvs2svn does is\n> inferring the most plausible history of the repository from CVS's\n> sketchy, incomplete, idiomatic, and often corrupt data.  This work\n> should also be useful for a cvs2git or cvs2hg or cvs2baz or ...\n\nGreat to hear that. I'm game if we can do something in this direction\n- surely we can make it talk to fastimport ;-)\n\nDoes cvs2svn handle incremental imports, remembering any \"guesses\"\ntaken earlier? Last time I looked at it, it had far better logic than\ncvsps, but it didn't do incremental imports, and repeated imports done\nat different times would \"guess\" different branching points for new\nbranches, so it _really_ didn't support incrementals\n\ncheers,\n\n\n\nm\n"},{"id":"47360","messageId":"4699034A.9090603@alum.mit.edu","threadId":"9022","inReplyTo":"46a038f90707132230n120e6392uaf5cd86ff10b6012@mail.gmail.com","subject":"Re: CVS -> SVN -> Git","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-07-14T17:09:30Z","receivedAt":"2007-07-14T17:09:30Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Martin Langhoff wrote:\n> On 7/14/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> Incidentally, now that cvs2svn 2.0.0 is nearly out, I am thinking about\n>> what it would take to write some other back ends for cvs2svn--turning\n>> it, essentially, into cvs2xxx.  Most of the work that cvs2svn does is\n>> inferring the most plausible history of the repository from CVS's\n>> sketchy, incomplete, idiomatic, and often corrupt data.  This work\n>> should also be useful for a cvs2git or cvs2hg or cvs2baz or ...\n> \n> Great to hear that. I'm game if we can do something in this direction\n> - surely we can make it talk to fastimport ;-)\n\nWe added some hooks to cvs2svn 2.0 to start working in this direction.\nBut I don't really know what information is needed for a git import.\nOne quick-and-dirty idea that I had was to have cvs2svn output\ninformation compatible with cvsps's output, as I believe that several\ntools rely on cvsps to do the dirty work and so could perhaps be\npersuaded to use cvs2svn out of the box.\n\n> Does cvs2svn handle incremental imports, remembering any \"guesses\"\n> taken earlier? Last time I looked at it, it had far better logic than\n> cvsps, but it didn't do incremental imports, and repeated imports done\n> at different times would \"guess\" different branching points for new\n> branches, so it _really_ didn't support incrementals\n\nThat's correct; cvs2svn does not support incremental conversion at all\n(at least not yet).\n\nMichael\n"},{"id":"47362","messageId":"20070714173223.GA25574@pe.Belkin","threadId":"9022","inReplyTo":"4699034A.9090603@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"Chris Shoemaker","fromEmail":"c.shoemaker@cox.net","sentAt":"2007-07-14T17:32:23Z","receivedAt":"2007-07-14T17:32:23Z","isPatch":false,"sender":{"key":"c.shoemaker@cox.net","avatar":null},"body":"On Sat, Jul 14, 2007 at 07:09:30PM +0200, Michael Haggerty wrote:\n> Martin Langhoff wrote:\n> > On 7/14/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> >> Incidentally, now that cvs2svn 2.0.0 is nearly out, I am thinking about\n> >> what it would take to write some other back ends for cvs2svn--turning\n> >> it, essentially, into cvs2xxx.  Most of the work that cvs2svn does is\n> >> inferring the most plausible history of the repository from CVS's\n> >> sketchy, incomplete, idiomatic, and often corrupt data.  This work\n> >> should also be useful for a cvs2git or cvs2hg or cvs2baz or ...\n> > \n> > Great to hear that. I'm game if we can do something in this direction\n> > - surely we can make it talk to fastimport ;-)\n> \n> We added some hooks to cvs2svn 2.0 to start working in this direction.\n> But I don't really know what information is needed for a git import.\n> One quick-and-dirty idea that I had was to have cvs2svn output\n> information compatible with cvsps's output, as I believe that several\n> tools rely on cvsps to do the dirty work and so could perhaps be\n> persuaded to use cvs2svn out of the box.\n\nDepending on how difficult that is, it might be very useful, even\nif it's not the best way to interface with fast-import (which I\nsuspect it's not).  I, for one, would be interested to know how\ncvs2svn's output compared to CVSps's, especially w.r.t. detecting each\nbranch's parent.\n\nPerhaps one is always more correct than the other, but if not, I bet\nthat seeing the differences using the same format would help to\nimprove either one.\n\n-chris\n"},{"id":"47364","messageId":"EC3B307E-EA83-4128-BABD-D9BDB78F987E@zib.de","threadId":"9022","inReplyTo":"4699034A.9090603@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"Steffen Prohaska","fromEmail":"prohaska@zib.de","sentAt":"2007-07-14T18:14:24Z","receivedAt":"2007-07-14T18:14:24Z","isPatch":false,"sender":{"key":"prohaska@zib.de","avatar":"https://avatars.githubusercontent.com/u/217580?v=4"},"body":"\nOn Jul 14, 2007, at 7:09 PM, Michael Haggerty wrote:\n\n> Martin Langhoff wrote:\n>> On 7/14/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>>> Incidentally, now that cvs2svn 2.0.0 is nearly out, I am thinking  \n>>> about\n>>> what it would take to write some other back ends for cvs2svn-- \n>>> turning\n>>> it, essentially, into cvs2xxx.  Most of the work that cvs2svn  \n>>> does is\n>>> inferring the most plausible history of the repository from CVS's\n>>> sketchy, incomplete, idiomatic, and often corrupt data.  This work\n>>> should also be useful for a cvs2git or cvs2hg or cvs2baz or ...\n>>\n>> Great to hear that. I'm game if we can do something in this direction\n>> - surely we can make it talk to fastimport ;-)\n>\n> We added some hooks to cvs2svn 2.0 to start working in this direction.\n> But I don't really know what information is needed for a git import.\n> One quick-and-dirty idea that I had was to have cvs2svn output\n> information compatible with cvsps's output, as I believe that several\n> tools rely on cvsps to do the dirty work and so could perhaps be\n> persuaded to use cvs2svn out of the box.\n\n From my understanding, piping data to git fast-import would be\na sane gateway to git. The input format of fast-import is document\nin [1].\n\nMaybe Shaw Pearce has some comments on that. Shawn did most\n(maybe all) of the work on git-fast-import.\n\nSimon Hausmann wrote a p4 importer that uses fast-import as\nits backend. Maybe, Simon can give hints how to get started.\n\nI have no experience with neither git-fast-import nor the p4\nimporter but would be happy to test any improved way of importing\ncvs to git. I experienced problems using git-cvsimport on a rather\nlarge cvs repository. Hence it would be a real test of the superior\ncapabilities of cvs2svn.\n\n\tSteffen\n\n\n[1] http://www.kernel.org/pub/software/scm/git/docs/git-fast-import.html\n"},{"id":"47373","messageId":"20070714195252.GB11010@thyrsus.com","threadId":"9022","inReplyTo":"4699034A.9090603@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2007-07-14T19:52:52Z","receivedAt":"2007-07-14T19:52:52Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu>:\n> Martin Langhoff wrote:\n> > On 7/14/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> >> Incidentally, now that cvs2svn 2.0.0 is nearly out, I am thinking about\n> >> what it would take to write some other back ends for cvs2svn--turning\n> >> it, essentially, into cvs2xxx.  Most of the work that cvs2svn does is\n> >> inferring the most plausible history of the repository from CVS's\n> >> sketchy, incomplete, idiomatic, and often corrupt data.  This work\n> >> should also be useful for a cvs2git or cvs2hg or cvs2baz or ...\n> > \n> > Great to hear that. I'm game if we can do something in this direction\n> > - surely we can make it talk to fastimport ;-)\n> \n> We added some hooks to cvs2svn 2.0 to start working in this direction.\n\nExcuse me, I missed Michael's Haggerty's original post. But as it\nhappens I've been doing quite a lot of work with VCSes and migration\ntools recently and I have some opinions and experience that I think are\nrelevant.\n\nIn slightly more detail: I just finished forward-porting sccs2rcs to\nPython, I've been moving some Subversion-hosted stuff to Mercurial,\nand I've been thinking about writing a simple rcs2svn because the last\ntime I tried using cvs2svn on a large RCS history (the Jargon File, as\nit happens, a couple years back) it did a very poor job of coalescing\nrelated commits without the CVS metadata.  I'm going to hope 2.0.0 has\nfixed that; I'll experiment and see.\n\nAlso, I'm in the process of rewriting Emacs VC mode and testing it\nwith three different VCSes. One consequence is that the Subversion\nsupport in Emacs is going to cease sucking badly in the very near\nfuture -- the VC-mode rewrite is giving VC mode the ability to make\natomic fileset commits if the underlying VCS will support them.  I'm\nputting the finishing touches on that code today, as it happens.\n\nAnother consequence is that Mercurial and git support will get really\ngood, oh, about ten minutes after the new Subversion backend lands.\nThe blocker on all three was the same weakness in the engine of VC\nmode, for whicch I was (alas) responsible as its original author and\n*which I have now fixed*.\n\nSo, I hear about plans to make cvs2svn generate something other than\nSubversion, and here's my instant reaction:\n\n\t    \t       \t   DON'T DO IT!\n\nThis is not because I think Subversion is some kind of final answer to the\nVCS problem.  Fame from it -- I'm moving towards Mercurial.  No, the\nreal reason I think this would be a waste of time is subtler than that.\n\nSubversion, by design, is very good at capturing the metadata from\nSCCS and RCS and the various CVS variants floating around.  In fact,\nlifting from those into Subversion is basically lossless - the real\nproblems are that (a) as Michael notes, the data you're losslessly\nlifting is scratchy, and (b) as I've noted, you have to use heuristics\nto coalesce file histories into changesets and those don't always make\nthe links they should.\n\nThat being the case, two-step conversion with tools that import CVS to\nSVN and export from SVN to whatever actually works extremely well.\nI'm speaking from direct recent experience here, not just theory. In\nfact, it works so well well that I'm convinced a tool for direct\nconversion from CVS to the third-generation systems would be misplaced\neffort.  \n\nI'd much rather see the effort go into improving import to Subversion\nfrom CVS and older, cruftier systems.  Subversion is like the Heinlein quote\nabout low Earth orbit being halfway to anywhere -- once your code and \nmetadata are there, export to advanced alien VCSes is easy.  So, Michael;\nI know it's nice to think about building space probes that can go direct \nto the aliens -- but please concentrate on building a better heavy-lift \nvehicle, because low earth orbit is the hard part.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"47374","messageId":"46992B9B.1070306@alum.mit.edu","threadId":"9022","inReplyTo":"20070714173223.GA25574@pe.Belkin","subject":"Re: CVS -> SVN -> Git","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-07-14T20:01:31Z","receivedAt":"2007-07-14T20:01:31Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Chris Shoemaker wrote:\n> [...] I, for one, would be interested to know how\n> cvs2svn's output compared to CVSps's, especially w.r.t. detecting each\n> branch's parent.\n\nThe problem, as I'm sure you are aware, is that CVS does not record\nunambiguously the parent of a branch.  For example, the following two\nsituations are indistinguishable from the data stored in CVS:\n\n1. Create BRANCH1 from trunk, then create BRANCH2 from BRANCH1 before\nmaking any commits to BRANCH1\n\n2. Create BRANCH1 from trunk, then create BRANCH2 from the same trunk\nrevision.\n\nOlder versions of cvs2svn would always create both branches from trunk.\n cvs2svn 2.0 gathers statistics about the \"possible parents\" of each\nsymbol across multiple files.  If BRANCH1 and BRANCH2 occur in another\nfile in a context that makes it clear that BRANCH2 was created from\nBRANCH1 (i.e., because a revision was committed to BRANCH1 before\nBRANCH2 was created), then it attempts to use BRANCH1 as the parent of\nBRANCH2 in all files.\n\nSo yes, cvs2svn is somewhat intelligent about determining the correct\nbranch ancestry.\n\nMichael\n"},{"id":"47381","messageId":"7vzm1yg0sj.fsf@assigned-by-dhcp.cox.net","threadId":"9022","inReplyTo":"20070714195252.GB11010@thyrsus.com","subject":"Re: CVS -> SVN -> Git","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2007-07-14T20:58:36Z","receivedAt":"2007-07-14T20:58:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"esr@thyrsus.com (Eric S. Raymond) writes:\n\n> So, I hear about plans to make cvs2svn generate something other than\n> Subversion, and here's my instant reaction:\n>\n> \t    \t       \t   DON'T DO IT!\n>\n> This is not because I think Subversion is some kind of final answer to the\n> VCS problem.  Fame from it -- I'm moving towards Mercurial.  No, the\n> real reason I think this would be a waste of time is subtler than that.\n>\n> Subversion, by design, is very good at capturing the metadata from\n> SCCS and RCS and the various CVS variants floating around.  In fact,\n> lifting from those into Subversion is basically lossless - the real\n> problems are that (a) as Michael notes, the data you're losslessly\n> lifting is scratchy, and (b) as I've noted, you have to use heuristics\n> to coalesce file histories into changesets and those don't always make\n> the links they should.\n\nConverting to Subversion might be lossless, but is it really the\nmost convenient intermediate format for other people to convert\nfurther from?\n\nEven after xxx2svn overcomes the problems (a) and (b) you noted\nabove, my impression has been that svn2yyy needs to work harder\nthan necessary to grok the branches/ and tags/ that artificially\nare flattened, only because Subversion does not do branches nor\ntags, but just represents them as copies.\n"},{"id":"47383","messageId":"20070714215031.GA3833@ugly.local","threadId":"9022","inReplyTo":"20070714195252.GB11010@thyrsus.com","subject":"Re: CVS -> SVN -> Git","fromName":"Oswald Buddenhagen","fromEmail":"ossi@kde.org","sentAt":"2007-07-14T21:50:31Z","receivedAt":"2007-07-14T21:50:31Z","isPatch":false,"sender":{"key":"ossi@kde.org","avatar":"https://avatars.githubusercontent.com/u/812380?v=4"},"body":"On Sat, Jul 14, 2007 at 03:52:52PM -0400, Eric S. Raymond wrote:\n> That being the case, two-step conversion with tools that import CVS to\n> SVN and export from SVN to whatever actually works extremely well.\n>\nwell, yes. hoooowever ... you are missing a few details:\n- conversion time. until we have incremental conversions, this is\n  absolutely critical to many organizations.\n- psychology. cvs2xxx is simpler than cvs2svn + svn2xxx. it's also sort\n  of a mindset thing. don't underestimate this.\n\n-- \nHi! I'm a .signature virus! Copy me into your ~/.signature, please!\n--\nChaos, panic, and disorder - my work here is done.\n"},{"id":"47384","messageId":"46994BDF.6050803@alum.mit.edu","threadId":"9022","inReplyTo":"20070714195252.GB11010@thyrsus.com","subject":"Re: CVS -> SVN -> Git","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-07-14T22:19:11Z","receivedAt":"2007-07-14T22:19:11Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Eric S. Raymond wrote:\n> In slightly more detail: I just finished forward-porting sccs2rcs to\n> Python, I've been moving some Subversion-hosted stuff to Mercurial,\n> and I've been thinking about writing a simple rcs2svn because the last\n> time I tried using cvs2svn on a large RCS history (the Jargon File, as\n> it happens, a couple years back) it did a very poor job of coalescing\n> related commits without the CVS metadata.  I'm going to hope 2.0.0 has\n> fixed that; I'll experiment and see.\n\nCould you give a quick summary of the relevant differences between CVS\nand RCS files in this context?  Then I'd be happy to try to figure out\nhow bad the situation still is today, and whether it can be easily improved.\n\n> [...]\n> So, I hear about plans to make cvs2svn generate something other than\n> Subversion, and here's my instant reaction:\n> \n> \t    \t       \t   DON'T DO IT!\n> \n> This is not because I think Subversion is some kind of final answer to the\n> VCS problem.  Fame from it -- I'm moving towards Mercurial.  No, the\n> real reason I think this would be a waste of time is subtler than that.\n> \n> Subversion, by design, is very good at capturing the metadata from\n> SCCS and RCS and the various CVS variants floating around.  In fact,\n> lifting from those into Subversion is basically lossless - the real\n> problems are that (a) as Michael notes, the data you're losslessly\n> lifting is scratchy, and (b) as I've noted, you have to use heuristics\n> to coalesce file histories into changesets and those don't always make\n> the links they should.\n> \n> That being the case, two-step conversion with tools that import CVS to\n> SVN and export from SVN to whatever actually works extremely well.\n\nOther people have complained about having to convert from SVN to\ndistributed SCMs, because the SVN model doesn't map so easily to their\nfavorite.\n\nYou are basically suggesting that an SVN repository is the best lingua\nfranca of the SCM world, which I don't believe.  The CVS history *does*\nhave to be deformed a bit to fit into SVN, and an svn2xxx converter\nwould have to undo the deformation.\n\nMy idea is not to built (for example) cvs2git; rather, I'd like cvs2svn\nto be split conceptually into two tools:\n\ncvs2<abstract_description_of_cvs_history>, whose job it is to determine\nthe most likely \"true\" CVS history based on the data stored in the CVS\nrepository, and\n\n<abstract_description_of_cvs_history>2svn\n\nThen later write\n\n<abstract_description_of_cvs_history>2git\n<abstract_description_of_cvs_history>2hg\n\netc.\n\nThe first split is partly done in cvs2svn 2.0.  And I naively imagine\nthat writing the new output back ends won't be all that much work.\n\nMichael\n"},{"id":"47387","messageId":"87sl7qzjty.fsf@red-bean.com","threadId":"9022","inReplyTo":"46994BDF.6050803@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"Karl Fogel","fromEmail":"kfogel@red-bean.com","sentAt":"2007-07-14T22:44:41Z","receivedAt":"2007-07-14T22:44:41Z","isPatch":false,"sender":{"key":"kfogel@red-bean.com","avatar":"https://gravatar.com/avatar/2d46620a078001b29f2878f5bfbad8cc4402d32af559f3eab6582315c3f6bb47?d=mp&s=160"},"body":"Michael Haggerty <mhagger@alum.mit.edu> writes:\n> My idea is not to built (for example) cvs2git; rather, I'd like cvs2svn\n> to be split conceptually into two tools:\n>\n> cvs2<abstract_description_of_cvs_history>, whose job it is to determine\n> the most likely \"true\" CVS history based on the data stored in the CVS\n> repository, and\n>\n> <abstract_description_of_cvs_history>2svn\n>\n> Then later write\n>\n> <abstract_description_of_cvs_history>2git\n> <abstract_description_of_cvs_history>2hg\n>\n> etc.\n>\n> The first split is partly done in cvs2svn 2.0.  And I naively imagine\n> that writing the new output back ends won't be all that much work.\n\nI think an intermediate interchange format is the right way to go.\n\nBut, isn't this what VCP / RevML is all about?  Perhaps RevML is\nalready suited to be that interchange format... (Haven't looked at it\nin detail, just pointing out that there has at least been an attempt\nto reinvent this wheel already :-) ).\n\n-Karl\n"},{"id":"47389","messageId":"7154c5c60707141623s3f70e967s226e5da29965a173@mail.gmail.com","threadId":"9022","inReplyTo":"46994BDF.6050803@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"David Frech","fromEmail":"david@nimblemachines.com","sentAt":"2007-07-14T23:23:27Z","receivedAt":"2007-07-14T23:23:27Z","isPatch":false,"sender":{"key":"david@nimblemachines.com","avatar":null},"body":"Now that this party is really rollicking, I think I'll join in. ;-)\n\nI have a modest svn repo (about 800 commits) that contains fifteen or\nso small projects. It started life as a CVS repo, and as the projects\ngrew and changed, and as I learned more about CVS, things got moved\naround. Later, when I got interested in svn (in 2005) I converted the\nrepo, using cvs2svn. It got a few things wrong - mostly, that it\nthought there was one project in the repo, and created toplevel\ntrunk/, branches/, and tags/ directories, and lumped everything below\nthese.\n\nSo, in svn, I moved things around some more.\n\nNow I want to switch to git. I've since added enough to svn that there\nis no option but to use th svn repo as my source. git-svnimport\ndoesn't work for me because its idea of the structure of my repo is\ntoo limited. I looked around, stumbled over fast-import, and got\nhooked on the idea of using it. It seemed simple enough... I wrote a\n350-line Lua (!!) program that parses the svn dump file and creates a\ncommit stream for fast-import.\n\nIt took a day and half to get the svn dump parsing right (it's an\negregiously bad format) but only a couple of hours to write the\nfast-import backend.\n\nThe code \"works\" in the sense that it can read an svn dump and create\na git repo that looks reasonable, but it misses a few things, like\nproperly inferring branch creation from the \"copyfrom\" info in the svn\ndump.\n\nHowever, it's fairly fast (~35 commits/sec) and flexible. I want to,\nin the process of doing this conversion, \"canonicalize\" the structure\nof the repo and throw away all the commits from cvs and svn that just\nmoved things around. This poses another inference challenge, but\nhaving a modest simple tool (ie, a short enough program to easily\nunderstand and modify) helps.\n\nHaving done all this, I realized that this is a good way to go.\nSeparating, as Michael suggests, the \"parsing\" part from the \"commit\ngenerating\" part, not only makes the tools easier to write, but makes\nthem more flexible. If hg or bzr had a git-like fast-import (maybe\nthey do) it would take me about 35 minutes to target that instead. And\nin the process I came across some \"missing features\" in fast-import,\nwhich Shawn Pearce was able to quickly add.\n\nMy repo is tiny, but I still think that speed and flexibility are key\nin this process. If I can write a little script that can be useful to\nsomeone with 100k commits instead of my measly 800, that's great.\n\nFor that matter, fast-import is a fairly short program. It wouldn't be\nhard for other scm projects to do something similar. fast-import could\nbecome a \"standard\" intermediate format. But even if that doesn't\nhappen, the amounts of code we're talking about (to do parsing and\ncommit generation) are reasonably modest and easy to change.\n\nAs soon as I make a bit more progress I'm going to make my code available.\n\nCheers,\n\n- David\n\nOn 7/14/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> My idea is not to built (for example) cvs2git; rather, I'd like cvs2svn\n> to be split conceptually into two tools:\n>\n> cvs2<abstract_description_of_cvs_history>, whose job it is to determine\n> the most likely \"true\" CVS history based on the data stored in the CVS\n> repository, and\n>\n> <abstract_description_of_cvs_history>2svn\n>\n> Then later write\n>\n> <abstract_description_of_cvs_history>2git\n> <abstract_description_of_cvs_history>2hg\n>\n> etc.\n>\n> The first split is partly done in cvs2svn 2.0.  And I naively imagine\n> that writing the new output back ends won't be all that much work.\n>\n> Michael\n>\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n-- \nIf I have not seen farther, it is because I have stood in the\nfootsteps of giants.\n"},{"id":"47392","messageId":"20070715013949.GA20850@thyrsus.com","threadId":"9022","inReplyTo":"46994BDF.6050803@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2007-07-15T01:39:49Z","receivedAt":"2007-07-15T01:39:49Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu>:\n> Could you give a quick summary of the relevant differences between CVS\n> and RCS files in this context?  Then I'd be happy to try to figure out\n> how bad the situation still is today, and whether it can be easily improved.\n\nI found my copy of the bug report, and I misremembered the problem\nslightly.  It turns out to be even more relevant to this \ndiscussion than I thought.\n\nThread begins with <20040810031409.GA25564@thyrsus.com> on\n9 Aug 2004.  The thread title was \"RFC -- enhancing cvs2svn to have a\nnotion of spans of mergeable commits\".  Your mailing-list archive\nsearch can't seem to find it, unfortunately.  I'll repost the query\niseparately\n\n> Other people have complained about having to convert from SVN to\n> distributed SCMs, because the SVN model doesn't map so easily to their\n> favorite.\n\nOK.  But I think that if SVN -> X is hard, CVS -> X is going to be harder.\n\n> You are basically suggesting that an SVN repository is the best lingua\n> franca of the SCM world, which I don't believe.\n\nNot quite.  I'm suggesting it's an appropriate lingua franca for centralized\nVCSes with branching, e.g. everything pre-Arch.\n\n>                                               The CVS history *does*\n> have to be deformed a bit to fit into SVN, and an svn2xxx converter\n> would have to undo the deformation.\n\nThen perhaps the right thing to think about is this: how exactly does\nCVS history need to be deformed, and is there some way to express the\nlost information as conventional properties or tags?\n\n> My idea is not to built (for example) cvs2git; rather, I'd like cvs2svn\n> to be split conceptually into two tools:\n\nWell, that makes more sense.  But how would whatever the first half outputs\nbe different from an svn dump file? \n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"47394","messageId":"20070715022203.GW4436@spearce.org","threadId":"9022","inReplyTo":"EC3B307E-EA83-4128-BABD-D9BDB78F987E@zib.de","subject":"Re: CVS -> SVN -> Git","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-07-15T02:22:03Z","receivedAt":"2007-07-15T02:22:03Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Steffen Prohaska <prohaska@zib.de> wrote:\n> On Jul 14, 2007, at 7:09 PM, Michael Haggerty wrote:\n> >Martin Langhoff wrote:\n> >>On 7/14/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> >>>Incidentally, now that cvs2svn 2.0.0 is nearly out, I am thinking  \n> >>>about\n> >>>what it would take to write some other back ends for cvs2svn-- \n> >>>turning\n> >>>it, essentially, into cvs2xxx.  Most of the work that cvs2svn  \n...\n> >>\n> >>Great to hear that. I'm game if we can do something in this direction\n> >>- surely we can make it talk to fastimport ;-)\n> >\n> >We added some hooks to cvs2svn 2.0 to start working in this direction.\n> >But I don't really know what information is needed for a git import.\n> >One quick-and-dirty idea that I had was to have cvs2svn output\n> >information compatible with cvsps's output, as I believe that several\n> >tools rely on cvsps to do the dirty work and so could perhaps be\n> >persuaded to use cvs2svn out of the box.\n> \n> From my understanding, piping data to git fast-import would be\n> a sane gateway to git. The input format of fast-import is document\n> in [1].\n> \n> Maybe Shaw Pearce has some comments on that. Shawn did most\n> (maybe all) of the work on git-fast-import.\n\nYou must be new to this discussion.  ;-)\n\ngit-fast-import started as a backend for a hacked up version of\ncvs2svn that Jon Smirl was working on to convert the massive Mozilla\nCVS repository into Git.  Jon started from the cvs2svn codebase\nbecause it best handled the damaged RCS files that exist in the\nMozilla repository.  Many emails have been exchanged between myself,\nMichael and Jon on this subject.\n\nSo yes, git-fast-import was designed to act as a backend behind\nsomething like cvs2xxx.  Some of the \"oddities\" of the fast-import\ninput language are the way they are partly because of the way\nthe (older) cvs2svn code generated output in SVN dump format.\nCertain data was available at certain times and not at others,\nso Jon wanted to feed it to git-fast-import when he had it, rather\nthan needing to buffer it or rearrange code.\n\nI'm staying far away from writing fast-import frontends.  Anyone that\nwants/needs a CVS frontend is welcome to implement one, but it\nwon't written be me.  I gave up CVS a long time ago and will never\nreturn to it.  My only VCS is Git, and converting Git->Git is sort\nof stupid.  So I have no need for a fast-import frontend.\n\nBut I do maintain fast-import.  Well over 99% of it was written\nby me.  ;-)\n\n-- \nShawn.\n"},{"id":"47395","messageId":"20070715023035.GX4436@spearce.org","threadId":"9022","inReplyTo":"7154c5c60707141623s3f70e967s226e5da29965a173@mail.gmail.com","subject":"Re: CVS -> SVN -> Git","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-07-15T02:30:35Z","receivedAt":"2007-07-15T02:30:35Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"David Frech <david@nimblemachines.com> wrote:\n> Now I want to switch to git. I've since added enough to svn that there\n> is no option but to use th svn repo as my source. git-svnimport\n> doesn't work for me because its idea of the structure of my repo is\n> too limited. I looked around, stumbled over fast-import, and got\n> hooked on the idea of using it. It seemed simple enough... I wrote a\n> 350-line Lua (!!) program that parses the svn dump file and creates a\n> commit stream for fast-import.\n> \n> It took a day and half to get the svn dump parsing right (it's an\n> egregiously bad format) but only a couple of hours to write the\n> fast-import backend.\n\nWith the 'C' (copy) and 'R' (rename) operators in fast-import I was\nstarting to suspect that an SVN dump->fast-import stream translator\nwasn't going to be that complex.\n\nI wouldn't want to attempt to parse the SVN dump format directly\nin fast-import.  As you said the format is horribly difficult\nto read.  The entire fast-import stream parser is only 624 lines\nof C (the other 1,636 lines of fast-import are for documentation,\nthe in memory tree/branch management and packfile generation).\nI doubt the SVN dump file can be parsed in as few lines of C code.\n\n-- \nShawn.\n"},{"id":"47416","messageId":"469A099E.6060906@alum.mit.edu","threadId":"9022","inReplyTo":"7154c5c60707141623s3f70e967s226e5da29965a173@mail.gmail.com","subject":"Re: CVS -> SVN -> Git","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-07-15T11:48:46Z","receivedAt":"2007-07-15T11:48:46Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"David Frech wrote:\n> I have a modest svn repo (about 800 commits) that contains fifteen or\n> so small projects. It started life as a CVS repo, and as the projects\n> grew and changed, and as I learned more about CVS, things got moved\n> around. Later, when I got interested in svn (in 2005) I converted the\n> repo, using cvs2svn. It got a few things wrong - mostly, that it\n> thought there was one project in the repo, and created toplevel\n> trunk/, branches/, and tags/ directories, and lumped everything below\n> these.\n\nI know this tangential to the main point of your post, but BTW\nmultiproject conversions were added to cvs2svn in release 1.5.\n\n> It took a day and half to get the svn dump parsing right (it's an\n> egregiously bad format) but only a couple of hours to write the\n> fast-import backend.\n\nI'm surprised you think that; I find the svn dump format quite easy and\nstraightforward.  (Of course it assumes some Subversionisms, like easy\ndeep directory copies, which I can imagine would be annoying in other\ncontexts.)  What don't you like about the format?\n\n> Having done all this, I realized that this is a good way to go.\n> Separating, as Michael suggests, the \"parsing\" part from the \"commit\n> generating\" part, not only makes the tools easier to write, but makes\n> them more flexible. If hg or bzr had a git-like fast-import (maybe\n> they do) it would take me about 35 minutes to target that instead. And\n> in the process I came across some \"missing features\" in fast-import,\n> which Shawn Pearce was able to quickly add.\n\nYes, fast-import is a very easy-to-write format and looks to be very\nwell documented.  I don't think that having to write output in\nfast-import format would be any kind of a hindrance for such a tool.\n\nMichael\n"},{"id":"47418","messageId":"469A0D54.8010303@alum.mit.edu","threadId":"9022","inReplyTo":"20070715013949.GA20850@thyrsus.com","subject":"Re: CVS -> SVN -> Git","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-07-15T12:04:36Z","receivedAt":"2007-07-15T12:04:36Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Eric S. Raymond wrote:\n> Michael Haggerty <mhagger@alum.mit.edu>:\n>>                                               The CVS history *does*\n>> have to be deformed a bit to fit into SVN, and an svn2xxx converter\n>> would have to undo the deformation.\n> \n> Then perhaps the right thing to think about is this: how exactly does\n> CVS history need to be deformed, and is there some way to express the\n> lost information as conventional properties or tags?\n\nHmmm, perhaps \"deformed\" was not the best word.  \"Reorganized\" is a\nbetter description.\n\nFor example, cvs2svn internally deduces which files should be added to a\ngiven branch in a given commit.  But the information cannot be output to\nSVN in that form.  Instead, cvs2svn has to figure out which\n*directories* to copy to the branch directory, then which files to\nremove from the copied directory (because they shouldn't have been\ntagged), and which other files to copy from other sources.  This extra\nwork, which is quite time- and space-consuming, is worse than pointless\nwhen converting to git, because git has to invert the process to figure\nout which individual files have to be tagged!\n\n>> My idea is not to built (for example) cvs2git; rather, I'd like cvs2svn\n>> to be split conceptually into two tools:\n> \n> Well, that makes more sense.  But how would whatever the first half outputs\n> be different from an svn dump file? \n\nThe interface between the two halves does not necessarily need to be a\nserialized data stream; it could just as well be via the Python API that\nis used internally by cvs2svn to access the reconstructed commits and\nsupporting databases.  This would require the second half to be written\nin Python, but otherwise would be very flexible and would avoid the need\nto find a be-all serialized format.\n\nMichael\n"},{"id":"47422","messageId":"20070715133655.GA9302@thyrsus.com","threadId":"9022","inReplyTo":"469A0D54.8010303@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2007-07-15T13:36:55Z","receivedAt":"2007-07-15T13:36:55Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu>:\n> For example, cvs2svn internally deduces which files should be added to a\n> given branch in a given commit.  But the information cannot be output to\n> SVN in that form.  Instead, cvs2svn has to figure out which\n> *directories* to copy to the branch directory, then which files to\n> remove from the copied directory (because they shouldn't have been\n> tagged), and which other files to copy from other sources.  This extra\n> work, which is quite time- and space-consuming, is worse than pointless\n> when converting to git, because git has to invert the process to figure\n> out which individual files have to be tagged!\n\nOK, that's a fair point.  I might have known the showstopper would be\nsomewhere near Subversion's tags-are-directories assumption.  And this\nalso neatly explains why I didn't see any problems or poor performance\nduring my recent conversions; the projects I was lifting had no tags.\n\n> The interface between the two halves does not necessarily need to be a\n> serialized data stream; it could just as well be via the Python API that\n> is used internally by cvs2svn to access the reconstructed commits and\n> supporting databases.  This would require the second half to be written\n> in Python, but otherwise would be very flexible and would avoid the need\n> to find a be-all serialized format.\n\nOr...wait for it...the generator for the serialized format could be one\nof the back ends!   Probably a good idea to have for debugging reasons, \nif nothing else.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"47468","messageId":"469AA917.3020401@slamb.org","threadId":"9022","inReplyTo":"4699034A.9090603@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"Scott Lamb","fromEmail":"slamb@slamb.org","sentAt":"2007-07-15T23:09:11Z","receivedAt":"2007-07-15T23:09:11Z","isPatch":false,"sender":{"key":"slamb@slamb.org","avatar":null},"body":"Michael Haggerty wrote:\n> One quick-and-dirty idea that I had was to have cvs2svn output\n> information compatible with cvsps's output, as I believe that several\n> tools rely on cvsps to do the dirty work and so could perhaps be\n> persuaded to use cvs2svn out of the box.\n\nI think this would be an excellent approach. The interface between\ncvs->X (cvsps), Y->git (git-fastimport), and cvs->git glue\n(git-cvsimport) is a great idea for troubleshooting and for code sharing\nwith other converters. (Shawn O. Pearce's attitude is a great example of\nthis - he can maintain the part he cares about and several converters\nbenefit even though he's never used them.)\n\nHowever, I was unhappy to see that cvsps doesn't reuse any cvs2svn code\nor unit tests. I remember seeing a lot of those hairy cases on the\nSubversion list long ago, so a CVS converter without those tests seems\nuntrustworthy. If I maintained an important CVS repository I wanted to\nconvert to git accurately, I would use cvs2svn.py+git-svnimport over\ngit-cvsimport any day.\n\nThey both seem much better than something like Tailor, though. I've\ndiscovered several things that made me realize going through working\ncopies is error-prone (as well as slow).\n\n>> Does cvs2svn handle incremental imports, remembering any \"guesses\"\n>> taken earlier? Last time I looked at it, it had far better logic than\n>> cvsps, but it didn't do incremental imports, and repeated imports done\n>> at different times would \"guess\" different branching points for new\n>> branches, so it _really_ didn't support incrementals\n> \n> That's correct; cvs2svn does not support incremental conversion at all\n> (at least not yet).\n\nThat's an important feature for me. I'm using git-cvsimport to track\nother people's CVS repositories. Initial import is SLOW and\nresource-intensive on the network, client, and server, so I couldn't\nswitch to anything that didn't support incremental use.\n\nBest regards,\nScott\n\n-- \nScott Lamb <http://www.slamb.org/>\n"},{"id":"47487","messageId":"46a038f90707151805j454b57fbvb4d7ed526e1e64ce@mail.gmail.com","threadId":"9022","inReplyTo":"20070715013949.GA20850@thyrsus.com","subject":"Re: CVS -> SVN -> Git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-07-16T01:05:37Z","receivedAt":"2007-07-16T01:05:37Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 7/15/07, Eric S. Raymond <esr@thyrsus.com> wrote:\n> Not quite.  I'm suggesting it's an appropriate lingua franca for centralized\n> VCSes with branching, e.g. everything pre-Arch.\n\nThat's a huge goal that gets in the way of waht we want to do here: we\nare trying to save time, not embark on some huge mission.\n\ncvs2svn has all the \"wtf-did-cvs-mean-by-that\" algorithms that are\nvery hard to write and maintain, and it seems to be the best one at\nthat. Of course, it also writes SVN repos -- but I'm sure that's the\neasiest part.\n\n     We don't need no meta VCS for any of this.\n\nAll we need is to hook into the \"write out a repo based on all the\nstuff we parsed from cvs\". Perhaps it's doable, and if Michael helps\nout abstracting that part a bit, maintainable long term too.\n\ncheers,\n\n\n\nm\n"},{"id":"47488","messageId":"46a038f90707151808u67c4e834lb06ed86c855f58ec@mail.gmail.com","threadId":"9022","inReplyTo":"469A099E.6060906@alum.mit.edu","subject":"Re: CVS -> SVN -> Git","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-07-16T01:08:09Z","receivedAt":"2007-07-16T01:08:09Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 7/15/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> > It took a day and half to get the svn dump parsing right (it's an\n> > egregiously bad format) but only a couple of hours to write the\n> > fast-import backend.\n>\n> I'm surprised you think that; I find the svn dump format quite easy and\n> straightforward.  (Of course it assumes some Subversionisms, like easy\n> deep directory copies, which I can imagine would be annoying in other\n> contexts.)  What don't you like about the format?\n\nIs there good doco and samples for it? I wouldn't mind doing things by\nway of an SVN dump parser.\n\n> Yes, fast-import is a very easy-to-write format and looks to be very\n> well documented.  I don't think that having to write output in\n> fast-import format would be any kind of a hindrance for such a tool.\n\nDamn! You've now figured out that all my volunteering was for the easy\npart of the job ;-)\n\n\n\n\nm\n"},{"id":"47489","messageId":"Pine.LNX.4.64.0707160212260.7851@beast.quantumfyre.co.uk","threadId":"9022","inReplyTo":"46a038f90707151808u67c4e834lb06ed86c855f58ec@mail.gmail.com","subject":"Re: CVS -> SVN -> Git","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2007-07-16T01:13:42Z","receivedAt":"2007-07-16T01:13:42Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Mon, 16 Jul 2007, Martin Langhoff wrote:\n\n> On 7/15/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> >  It took a day and half to get the svn dump parsing right (it's an\n>> >  egregiously bad format) but only a couple of hours to write the\n>> >  fast-import backend.\n>>\n>>  I'm surprised you think that; I find the svn dump format quite easy and\n>>  straightforward.  (Of course it assumes some Subversionisms, like easy\n>>  deep directory copies, which I can imagine would be annoying in other\n>>  contexts.)  What don't you like about the format?\n>\n> Is there good doco and samples for it? I wouldn't mind doing things by\n> way of an SVN dump parser.\n\nI don't know if it classes as what you call good, but it is documented:\n\nhttp://svn.collab.net/repos/svn/trunk/notes/fs_dumprestore.txt\n\n>\n>>  Yes, fast-import is a very easy-to-write format and looks to be very\n>>  well documented.  I don't think that having to write output in\n>>  fast-import format would be any kind of a hindrance for such a tool.\n>\n> Damn! You've now figured out that all my volunteering was for the easy\n> part of the job ;-)\n>\n>\n>\n>\n> m\n>\n>\n\n-- \nJulian\n\n  ---\nThis process can check if this value is zero, and if it is, it does\nsomething child-like.\n \t\t-- Forbes Burkowski, CS 454, University of Washington\n"},{"id":"47492","messageId":"873azpqgnq.fsf@red-bean.com","threadId":"9022","inReplyTo":"46a038f90707151808u67c4e834lb06ed86c855f58ec@mail.gmail.com","subject":"Re: CVS -> SVN -> Git","fromName":"Karl Fogel","fromEmail":"kfogel@red-bean.com","sentAt":"2007-07-16T01:30:17Z","receivedAt":"2007-07-16T01:30:17Z","isPatch":false,"sender":{"key":"kfogel@red-bean.com","avatar":"https://gravatar.com/avatar/2d46620a078001b29f2878f5bfbad8cc4402d32af559f3eab6582315c3f6bb47?d=mp&s=160"},"body":"\"Martin Langhoff\" <martin.langhoff@gmail.com> writes:\n> On 7/15/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> > It took a day and half to get the svn dump parsing right (it's an\n>> > egregiously bad format) but only a couple of hours to write the\n>> > fast-import backend.\n>>\n>> I'm surprised you think that; I find the svn dump format quite easy and\n>> straightforward.  (Of course it assumes some Subversionisms, like easy\n>> deep directory copies, which I can imagine would be annoying in other\n>> contexts.)  What don't you like about the format?\n>\n> Is there good doco and samples for it? I wouldn't mind doing things by\n> way of an SVN dump parser.\n\n   http://svn.collab.net/repos/svn/trunk/notes/dump-load-format.txt\n\nBest,\n-Karl\n"},{"id":"47849","messageId":"469F52BF.8050300@bluegap.ch","threadId":"9022","inReplyTo":"46a038f90707151805j454b57fbvb4d7ed526e1e64ce@mail.gmail.com","subject":"Re: CVS -> SVN -> Git","fromName":"Markus Schiltknecht","fromEmail":"markus@bluegap.ch","sentAt":"2007-07-19T12:02:07Z","receivedAt":"2007-07-19T12:02:07Z","isPatch":false,"sender":{"key":"markus@bluegap.ch","avatar":"https://gravatar.com/avatar/2f3aadbc46c7c942fa9301d0bb9b91da8c7186233f4c7eff9667a8b53c8cc82e?d=mp&s=160"},"body":"Hi,\n\nMartin Langhoff wrote:\n> cvs2svn has all the \"wtf-did-cvs-mean-by-that\" algorithms that are\n> very hard to write and maintain, and it seems to be the best one at\n> that. Of course, it also writes SVN repos -- but I'm sure that's the\n> easiest part.\n> \n>     We don't need no meta VCS for any of this.\n\nSure, we certainly need a meta format of some sort (not a full blown \nVCS, agreed, but somehow we need to represent commits, tags and \nbranches). And IMO, the subversion based format is not a good one, \nbecause it treats branches and tags very different from most other \nsystems (and from what it should be from a users perspective: an atomic \noperation).\n\nWe (Michael, Oswald and me) have discussed joining efforts of my cvs to \nmonotone converter, but I quickly dropped that idea because the cvs2svn \nconverter is too subversion specific. If cvs2svn wants to become a \nuniversal cvs importer, it needs to get rid of those assumptions (and do \nmore work to unify tagging and branching).\n\nRegards\n\nMarkus\n"},{"id":"47894","messageId":"469FB80A.8000001@fs.ei.tum.de","threadId":"9022","inReplyTo":"46a038f90707151805j454b57fbvb4d7ed526e1e64ce@mail.gmail.com","subject":"Re: CVS -> SVN -> Git","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-07-19T19:14:18Z","receivedAt":"2007-07-19T19:14:18Z","isPatch":false,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"[sorry for jumping in so late, didn't read git@vger for a while]\n\nMartin Langhoff wrote:\n> On 7/15/07, Eric S. Raymond <esr@thyrsus.com> wrote:\n>> Not quite.  I'm suggesting it's an appropriate lingua franca for \n>> centralized\n>> VCSes with branching, e.g. everything pre-Arch.\n\nI do not think Eric is right here.  You will allways lose information when converting CVS to svn, and if it is just the uncertainty, the non-atomicity.  This is also information (hidden one, though).\n\n> That's a huge goal that gets in the way of waht we want to do here: we\n> are trying to save time, not embark on some huge mission.\n> \n> cvs2svn has all the \"wtf-did-cvs-mean-by-that\" algorithms that are\n> very hard to write and maintain, and it seems to be the best one at\n> that. Of course, it also writes SVN repos -- but I'm sure that's the\n> easiest part.\n\nTrue.  However, cvs2svn has many assumptions (or at least has had when I last checked) which are targeted to svn, and unsuitable for a generic system (tags + branches).\n\n>     We don't need no meta VCS for any of this.\n\nYes.  I've already done what people want, it is not called cvs2xxx, but fromcvs [1].  I don't think it is necessary to define an output format.  Of course, that's possible, but limiting yourself to a file format means you're losing flexibility, which is needed for efficient, correct and fast repository conversion.\n\ncheers\n  simon\n\n[1] http://ww2.fs.ei.tum.de/~corecode/hg/fromcvs/\n\n-- \nServe - BSD     +++  RENT this banner advert  +++    ASCII Ribbon   /\"\\\nWork - Mac      +++  space for low €€€ NOW!1  +++      Campaign     \\ /\nParty Enjoy Relax   |   http://dragonflybsd.org      Against  HTML   \\\nDude 2c 2 the max   !   http://golden-apple.biz       Mail + News   / \\\n"},{"id":"47901","messageId":"469FB84B.2010909@fs.ei.tum.de","threadId":"9022","inReplyTo":"Pine.LNX.4.64.0707131541140.11423@reaper.quantumfyre.co.uk","subject":"Re: CVS -> SVN -> Git","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-07-19T19:15:23Z","receivedAt":"2007-07-19T19:15:23Z","isPatch":false,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"Julian Phillips wrote:\n> Has anyone managed to succssfully import a Subversion repository that \n> was initially imported from CVS using cvs2svn using fast-import?\n> \n> It looks like cvs2svn has created a rather big mess.   It has created \n> single commits that change files in more than one branch and/or tag. It \n> also creates tags using more than one commit.  Now I come to try and \n> import the Subversion history into git and I'm having trouble creating a \n> sensible stream to feed into fast-import.\n\nDid you try first converting the old CVS repo to git and then adding the svn changes?  That might give you much better results.\n\ncheers\n  simon\n\n-- \nServe - BSD     +++  RENT this banner advert  +++    ASCII Ribbon   /\"\\\nWork - Mac      +++  space for low €€€ NOW!1  +++      Campaign     \\ /\nParty Enjoy Relax   |   http://dragonflybsd.org      Against  HTML   \\\nDude 2c 2 the max   !   http://golden-apple.biz       Mail + News   / \\\n"},{"id":"47902","messageId":"469FB8F4.5090902@fs.ei.tum.de","threadId":"9022","inReplyTo":"46a038f90707132230n120e6392uaf5cd86ff10b6012@mail.gmail.com","subject":"Re: CVS -> SVN -> Git","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-07-19T19:18:12Z","receivedAt":"2007-07-19T19:18:12Z","isPatch":false,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"Martin Langhoff wrote:\n> On 7/14/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> Incidentally, now that cvs2svn 2.0.0 is nearly out, I am thinking about\n>> what it would take to write some other back ends for cvs2svn--turning\n>> it, essentially, into cvs2xxx.  Most of the work that cvs2svn does is\n>> inferring the most plausible history of the repository from CVS's\n>> sketchy, incomplete, idiomatic, and often corrupt data.  This work\n>> should also be useful for a cvs2git or cvs2hg or cvs2baz or ...\n> \n> Great to hear that. I'm game if we can do something in this direction\n> - surely we can make it talk to fastimport ;-)\n\nIn this context I suggest looking at fromcvs [1], my cvs->otherscm converter.  Right now it does git + hg (and sqlite for queries), but it probably is easily extensible for other targets.\n\n> Does cvs2svn handle incremental imports, remembering any \"guesses\"\n> taken earlier? Last time I looked at it, it had far better logic than\n> cvsps, but it didn't do incremental imports, and repeated imports done\n> at different times would \"guess\" different branching points for new\n> branches, so it _really_ didn't support incrementals\n\nfromcvs will also handle incremental imports.  If not, please tell me and I will try to fix it.\n\ncheers\n  simon\n\n[1] http://ww2.fs.ei.tum.de/~corecode/hg/fromcvs/\n\n-- \nServe - BSD     +++  RENT this banner advert  +++    ASCII Ribbon   /\"\\\nWork - Mac      +++  space for low €€€ NOW!1  +++      Campaign     \\ /\nParty Enjoy Relax   |   http://dragonflybsd.org      Against  HTML   \\\nDude 2c 2 the max   !   http://golden-apple.biz       Mail + News   / \\\n"},{"id":"47891","messageId":"87tzrzivgf.fsf@red-bean.com","threadId":"9022","inReplyTo":"469F52BF.8050300@bluegap.ch","subject":"Re: CVS -> SVN -> Git","fromName":"Karl Fogel","fromEmail":"kfogel@red-bean.com","sentAt":"2007-07-20T03:51:28Z","receivedAt":"2007-07-20T03:51:28Z","isPatch":false,"sender":{"key":"kfogel@red-bean.com","avatar":"https://gravatar.com/avatar/2d46620a078001b29f2878f5bfbad8cc4402d32af559f3eab6582315c3f6bb47?d=mp&s=160"},"body":"Markus Schiltknecht <markus@bluegap.ch> writes:\n> Sure, we certainly need a meta format of some sort (not a full blown\n> VCS, agreed, but somehow we need to represent commits, tags and\n> branches). And IMO, the subversion based format is not a good one,\n> because it treats branches and tags very different from most other\n> systems (and from what it should be from a users perspective: an\n> atomic operation).\n\nHuh?  I don't understand what you're saying about atomicity here.\n"},{"id":"47930","messageId":"Pine.LNX.4.64.0707200651530.18125@beast.quantumfyre.co.uk","threadId":"9022","inReplyTo":"469FB84B.2010909@fs.ei.tum.de","subject":"Re: CVS -> SVN -> Git","fromName":"Julian Phillips","fromEmail":"julian@quantumfyre.co.uk","sentAt":"2007-07-20T05:58:09Z","receivedAt":"2007-07-20T05:58:09Z","isPatch":false,"sender":{"key":"julian@quantumfyre.co.uk","avatar":"https://avatars.githubusercontent.com/u/948888?v=4"},"body":"On Thu, 19 Jul 2007, Simon 'corecode' Schubert wrote:\n\n> Julian Phillips wrote:\n>>  Has anyone managed to succssfully import a Subversion repository that was\n>>  initially imported from CVS using cvs2svn using fast-import?\n>>\n>>  It looks like cvs2svn has created a rather big mess.   It has created\n>>  single commits that change files in more than one branch and/or tag. It\n>>  also creates tags using more than one commit.  Now I come to try and\n>>  import the Subversion history into git and I'm having trouble creating a\n>>  sensible stream to feed into fast-import.\n>\n> Did you try first converting the old CVS repo to git and then adding the svn \n> changes?  That might give you much better results.\n\nI thought about it, but there are over 20000 commits sitting on top of the \nconverted history referring back to it - I'm not convinced that I could \nstitch things back together properly, the svn history now really does \nrely on the import done by cvs2svn.  (btw, I blame CVS for the mess not \ncvs2svn, we should have switched _before_ we started using branches \nheavily ...)\n\nThe problem really is that we use branching like it's going out of \nfashion.  We have thousands now, and had at least 10s if not 100s by the \ntime we gave up on CVS.  Similarly with tags.\n\nI think I've managed to get things sorted now with fast-import ... just \nneed to be able to copy blobs from other commits and I think I'll be done. \nIt really is a nice tool.\n\n-- \nJulian\n\n  ---\nThere is nothing so easy but that it becomes difficult when you do it\nreluctantly.\n \t\t-- Publius Terentius Afer (Terence)\n"},{"id":"47944","messageId":"46A0763E.1070402@bluegap.ch","threadId":"9022","inReplyTo":"469FB80A.8000001@fs.ei.tum.de","subject":"Re: CVS -> SVN -> Git","fromName":"Markus Schiltknecht","fromEmail":"markus@bluegap.ch","sentAt":"2007-07-20T08:45:50Z","receivedAt":"2007-07-20T08:45:50Z","isPatch":false,"sender":{"key":"markus@bluegap.ch","avatar":"https://gravatar.com/avatar/2f3aadbc46c7c942fa9301d0bb9b91da8c7186233f4c7eff9667a8b53c8cc82e?d=mp&s=160"},"body":"Simon 'corecode' Schubert wrote:\n> I do not think Eric is right here.  You will allways lose information \n> when converting CVS to svn, and if it is just the uncertainty, the \n> non-atomicity.  This is also information (hidden one, though).\n\nFull ACK.\n\n> Yes.  I've already done what people want, it is not called cvs2xxx, but \n> fromcvs [1].  I don't think it is necessary to define an output format.  \n> Of course, that's possible, but limiting yourself to a file format means \n> you're losing flexibility, which is needed for efficient, correct and \n> fast repository conversion.\n\nHm.. interesting. I'll have a close look.\n\nRegards\n\nMarkus\n"}]}