{"thread":{"id":"43263","subject":"Re: Some tips for doing a CVS importer","startedAt":"2006-11-20T21:49:17Z","lastAt":"2006-11-30T00:45:08Z","messageCount":30,"participants":["Jon Smirl","Daniel Jacobowitz","Martin Langhoff","Johannes Schindelin","Shawn Pearce","Marko Macek","Carl Worth","Robin Rosenberg","Michael Haggerty","Johannes Sixt","Markus Schiltknecht","Petr Baudis","lamikr"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"295403","messageId":"9e4733910611201349s4d08b984g772c64982f148bfa@mail.gmail.com","threadId":"43263","inReplyTo":null,"subject":"Some tips for doing a CVS importer","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-11-20T21:49:17Z","receivedAt":"2006-11-20T21:49:17Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"I have tried all of the available CVS importers. None of them are\nwithout problems. If anyone is interested in writing one for git here\nare some ideas on how to structure it.\n\n1) there is a working lex/yacc for CVS in the parsecvs source code\n2) The first time you parse a CVS file record everything and don't\nparse it again.\n3) When the file is first parsed use the deltas to generate the\nrevisions and feed them to git-fastimport, just remember the SHA1 or\nan id in the import code. This is a critical step to getting decent\nperformance.\n4) If you do #1 and #2 you don't need to store CVS revision numbers\nand file names in memory. Because of that you can can easily do a\nMozilla import in 2GB, probably 1GB.\n5) When comparing CVS revisions only use the CVS timestamps as a last\nresort, instead use the dependency information in the CVS file\n6) Match up commits by using an sha1 of the author and commit message\n7) After all files are loaded, match up the symbols and insert them\ninto the dependency chains, if any of the symbols depend on a branch\ncommit the symbol lies on the branch, otherwise the symbol is on the\ntrunk,\n8) Do a topological sort to build the change set commit tree\n9) when you hit a loop in the tree break up delta change sets until\nthe loop can be removed, don't break up symbol change sets.\n10) Mozilla has some large commits that were made over dial up. Commit\nchange sets can span hours. All of these commits need to be merged\ninto a single change set.\n11) An algorithm needs to be developed for detecting branches merging\nback into the trunk\n12) cvs2svn has excellent test cases, use them to test the new\nimporter. The cvs2svn code is quite nice but it doesn't handle #7\n\n-- \nJon Smirl\n"},{"id":"297541","messageId":"46a038f90611201503m6a63ec8ct347026c635190108@mail.gmail.com","threadId":"43263","inReplyTo":"9e4733910611201349s4d08b984g772c64982f148bfa@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-11-20T23:03:01Z","receivedAt":"2006-11-20T23:03:01Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 11/21/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> I have tried all of the available CVS importers. None of them are\n> without problems. If anyone is interested in writing one for git here\n> are some ideas on how to structure it.\n\nHi Jon,\n\nI gather this means that the cvs2svn track hasn't been as productive\nas expected. Any remaining/unsolvable issues with it? I have been\nchronically busy on my e-learning projects, but don't discard coming\nback to this when I have some time.\n\ncheers,\n\n\n\n"},{"id":"297143","messageId":"9e4733910611201537h30b6c9f4oee9d8df75284c284@mail.gmail.com","threadId":"43263","inReplyTo":"46a038f90611201503m6a63ec8ct347026c635190108@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-11-20T23:37:41Z","receivedAt":"2006-11-20T23:37:41Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 11/20/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> On 11/21/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > I have tried all of the available CVS importers. None of them are\n> > without problems. If anyone is interested in writing one for git here\n> > are some ideas on how to structure it.\n>\n> Hi Jon,\n>\n> I gather this means that the cvs2svn track hasn't been as productive\n> as expected. Any remaining/unsolvable issues with it? I have been\n> chronically busy on my e-learning projects, but don't discard coming\n> back to this when I have some time.\n\nLook in this thread\n[Fwd: Re: What's in git.git]\n\nThere is a message in there that explains a problem that the cvs2svn\npeople aren't going to fix and it kills git.\n\n\n>\n> cheers,\n>\n>\n>\n> martin\n>\n\n\n-- \nJon Smirl\n"},{"id":"294467","messageId":"46a038f90611201629o39f11f42ye07b86159360b66e@mail.gmail.com","threadId":"43263","inReplyTo":"9e4733910611201537h30b6c9f4oee9d8df75284c284@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-11-21T00:29:20Z","receivedAt":"2006-11-21T00:29:20Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 11/21/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > I gather this means that the cvs2svn track hasn't been as productive\n> > as expected. Any remaining/unsolvable issues with it? I have been\n> > chronically busy on my e-learning projects, but don't discard coming\n> > back to this when I have some time.\n>\n> Look in this thread\n> [Fwd: Re: What's in git.git]\n>\n> There is a message in there that explains a problem that the cvs2svn\n> people aren't going to fix and it kills git.\n\nI see - thanks for the pointer. Sorry to hear others in the Moz\nproject weren't so keen on hearing about alternatives to SVN. Long\nterm only something like GIT seems viable for such a large project (in\nterms of community, branches/subprojects and codebase).\n\nTwo remaining questions\n - Where can I get your latest code? :-)\n - I gather the moz cvs repo has some cases that require getting the\nsymbol resolution right. Could this be performed as an extra pass /\ntask?\n\nEventually the Moz crowd will outgrow SVN - perhaps we should be\nparsing the SVN dump format instead ;-)\n\ncheers,\n\n\n"},{"id":"294773","messageId":"87vel9y5x6.wl%cworth@cworth.org","threadId":"43263","inReplyTo":"46a038f90611201629o39f11f42ye07b86159360b66e@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Carl Worth","fromEmail":"cworth@cworth.org","sentAt":"2006-11-21T00:55:33Z","receivedAt":"2006-11-21T00:55:33Z","isPatch":false,"sender":{"key":"cworth@cworth.org","avatar":"https://gravatar.com/avatar/3746dc28cde609bdbd7f939058356e7e2bbd16d21e32274df0725eb3d998bc5b?d=mp&s=160"},"body":"On Tue, 21 Nov 2006 13:29:20 +1300, \"Martin Langhoff\" wrote:\n> I see - thanks for the pointer. Sorry to hear others in the Moz\n> project weren't so keen on hearing about alternatives to SVN. Long\n> term only something like GIT seems viable for such a large project (in\n> terms of community, branches/subprojects and codebase).\n...\n> Eventually the Moz crowd will outgrow SVN - perhaps we should be\n> parsing the SVN dump format instead ;-)\n\nFrom what I understand, mozilla is currently using CVS and is looking\nto replace that. The remaining options being considered are bzr and\nhg, (git having been discarded due to the lack of a \"native\" win32\nclient---the cygwin stuff is apparently not considered viable for\nwhatever reason).\n\n-Carl\n"},{"id":"295637","messageId":"9e4733910611201740i348302e6r84c3c27dc27e5954@mail.gmail.com","threadId":"43263","inReplyTo":"87vel9y5x6.wl%cworth@cworth.org","subject":"Re: Some tips for doing a CVS importer","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-11-21T01:40:40Z","receivedAt":"2006-11-21T01:40:40Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 11/20/06, Carl Worth <cworth@cworth.org> wrote:\n> On Tue, 21 Nov 2006 13:29:20 +1300, \"Martin Langhoff\" wrote:\n> > I see - thanks for the pointer. Sorry to hear others in the Moz\n> > project weren't so keen on hearing about alternatives to SVN. Long\n> > term only something like GIT seems viable for such a large project (in\n> > terms of community, branches/subprojects and codebase).\n> ...\n> > Eventually the Moz crowd will outgrow SVN - perhaps we should be\n> > parsing the SVN dump format instead ;-)\n>\n> From what I understand, mozilla is currently using CVS and is looking\n> to replace that. The remaining options being considered are bzr and\n> hg, (git having been discarded due to the lack of a \"native\" win32\n\nbrendan said SVN is likely for the main Mozilla repo and monotone for\nthe new Mozilla 2 work. No native win32 caused git to be immediately\ndiscarded.\n\n> client---the cygwin stuff is apparently not considered viable for\n> whatever reason).\n>\n> -Carl\n>\n>\n>\n\n\n-- \nJon Smirl\n"},{"id":"296045","messageId":"9e4733910611201753m392b5defpb3eb295a075be789@mail.gmail.com","threadId":"43263","inReplyTo":"46a038f90611201629o39f11f42ye07b86159360b66e@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-11-21T01:53:15Z","receivedAt":"2006-11-21T01:53:15Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 11/20/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> On 11/21/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > > I gather this means that the cvs2svn track hasn't been as productive\n> > > as expected. Any remaining/unsolvable issues with it? I have been\n> > > chronically busy on my e-learning projects, but don't discard coming\n> > > back to this when I have some time.\n> >\n> > Look in this thread\n> > [Fwd: Re: What's in git.git]\n> >\n> > There is a message in there that explains a problem that the cvs2svn\n> > people aren't going to fix and it kills git.\n>\n> I see - thanks for the pointer. Sorry to hear others in the Moz\n> project weren't so keen on hearing about alternatives to SVN. Long\n> term only something like GIT seems viable for such a large project (in\n> terms of community, branches/subprojects and codebase).\n>\n> Two remaining questions\n>  - Where can I get your latest code? :-)\n\nI gave up on my cvs2git code, cvs2svn has been refactored so badly\nthat it was too much trouble tracking. It would be easier to write it\nagain. Most of the smarts from the import process is in the\ngit-fastimport code which Shawn has. cvs2svn underwent a major\nalgorithm change after I wrote the first version of git2svn.\n\nI can probably find the code if you really want it, but it will be\nleading you off in the wrong direction.\n\n>  - I gather the moz cvs repo has some cases that require getting the\n> symbol resolution right. Could this be performed as an extra pass /\n> task?\n\nProcessing the symbols is integral to deciding how to build the change\nsets. Right now cvs2svn ignores the symbol dependency information and\nbuilds the change sets in a way that forces the mini-branches. That\ncauses 60% of the 2,000 symbols in Mozilla CVS to end up as little\nbranches. Look at the three commit example in the other thread to see\nexactly what the problem is.\n\nSVN hides the mini branch by creating a symbol like this:\n\nSymbol XXX, change set 70\ncopy All from change set 50\ncopy file A from change set 55\ncopy file B,C from change set 60\ncopy file D from change set 61\ncopy file E,F,G from change set 63\ncopy file H from change set 67\n\nIt has to do all of those copies because the change sets weren't\nconstructed while taking symbol dependency information into account.\n\nSymbol XXX can't copy from change set 69 because commits from after\nthe symbol was created are included in change sets 51-69.\n\n> Eventually the Moz crowd will outgrow SVN - perhaps we should be\n> parsing the SVN dump format instead ;-)\n>\n> cheers,\n>\n>\n> martin\n>\n\n\n-- \nJon Smirl\n"},{"id":"295489","messageId":"20061121063934.GA3332@spearce.org","threadId":"43263","inReplyTo":"9e4733910611201740i348302e6r84c3c27dc27e5954@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-11-21T06:39:35Z","receivedAt":"2006-11-21T06:39:35Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Jon Smirl <jonsmirl@gmail.com> wrote:\n> brendan said SVN is likely for the main Mozilla repo and monotone for\n> the new Mozilla 2 work. No native win32 caused git to be immediately\n> discarded.\n\nYea, that lack of native win32 seems to be one of a number of\nblockers for people switching their projects onto Git.\n\nI think there's a number of issues that are keeping people from\nswitching to Git and are instead causing them to choose SVN, hg\nor Monotone:\n\n  - No GUI.\n  - No native win32 installation.\n  - CVS import fails on some projects (e.g. Mozilla).\n  - Confusing documentation.\n  - pull/merge debate.\n  - Fear of hash conflicts corrupting a repository.\n\nI think Junio has solved the pull/merge debate issue.  We've talked\nthe hash conflict issue to death, but some new people still haven't\nread those threads (or won't believe them).  I know people are trying\nto work on improving the documentation, but there is obviously still\nroom for improvements.\n\nRight now I'm trying to work on the no GUI problem with git-gui...\n\n-- \n"},{"id":"294563","messageId":"20061121064328.GB3332@spearce.org","threadId":"43263","inReplyTo":"46a038f90611201629o39f11f42ye07b86159360b66e@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-11-21T06:43:28Z","receivedAt":"2006-11-21T06:43:28Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> Eventually the Moz crowd will outgrow SVN - perhaps we should be\n> parsing the SVN dump format instead ;-)\n\nIts a mess.  :)\n\nJon and I considered using the SVN dump format to feed git-fastimport\nbut chose against it.  Its a pretty horrible format.  Especially with\nhow it handles branches and tags, and file data.\n\nFortunately SVN has a C library which parses the file for you.\nWhich means that probably the best way to read the SVN dump format is\nto write a program which links against the SVN library and translates\nit into the datastructures used internally by git-fastimport to\ngenerate an initial pack file, then repack that after the import\nto get good compression.\n\n-- \n"},{"id":"298209","messageId":"456359E2.8010403@cc.jyu.fi","threadId":"43263","inReplyTo":"20061121063934.GA3332@spearce.org","subject":"Re: Some tips for doing a CVS importer","fromName":"lamikr","fromEmail":"lamikr@cc.jyu.fi","sentAt":"2006-11-21T19:56:18Z","receivedAt":"2006-11-21T19:56:18Z","isPatch":false,"sender":{"key":"lamikr@cc.jyu.fi","avatar":null},"body":"Shawn Pearce wrote:\n> Jon Smirl <jonsmirl@gmail.com> wrote:\n>   \n>> brendan said SVN is likely for the main Mozilla repo and monotone for\n>> the new Mozilla 2 work. No native win32 caused git to be immediately\n>> discarded.\n>>     \n>\n> Yea, that lack of native win32 seems to be one of a number of\n> blockers for people switching their projects onto Git.\n>\n> I think there's a number of issues that are keeping people from\n> switching to Git and are instead causing them to choose SVN, hg\n> or Monotone:\n>\n>   - No GUI.\n>   \nQGIT allows using some commands. I plan to try out the GIT eclipse\nplugin in near future myself.\nThis mail list have some discussion and download link to it's repo in\narchives.\n(title: Java GIT/Eclipse GIT version 0.1.1, )\n\n>   - No native win32 installation.\n>   - CVS import fails on some projects (e.g. Mozilla).\n>   \nWell, committing the files from Mozilla cvs to svn has also own problems.\nSVN accepts only a text files which have either a \"Unix\" or DOS style\nline endings.\nIf file contains a both some lines using \"Unix\" way and others using dos\nway SVN roll's\nback the commit and you need to tools like \"dos2unix\" or \"unix2dos\" to\nmanipulate those.\n(And randomly changing all to either of those is propably not a good idea)\n\n"},{"id":"297253","messageId":"20061121200341.GH7201@pasky.or.cz","threadId":"43263","inReplyTo":"20061121063934.GA3332@spearce.org","subject":"Re: Some tips for doing a CVS importer","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-11-21T20:03:41Z","receivedAt":"2006-11-21T20:03:41Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"On Tue, Nov 21, 2006 at 07:39:35AM CET, Shawn Pearce wrote:\n> Jon Smirl <jonsmirl@gmail.com> wrote:\n> > brendan said SVN is likely for the main Mozilla repo and monotone for\n> > the new Mozilla 2 work. No native win32 caused git to be immediately\n> > discarded.\n> \n> Yea, that lack of native win32 seems to be one of a number of\n> blockers for people switching their projects onto Git.\n\nYep. :-(\n\n> I think there's a number of issues that are keeping people from\n> switching to Git and are instead causing them to choose SVN, hg\n> or Monotone:\n> \n>   - No GUI.\n\nIt has been my impression that Git's situation is far better than in\ncase of the other systems (except SVN: TortoiseSVN and RapidSVN). Is\nthat not so?\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nThe meaning of Stonehenge in Traflamadorian, when viewed from above, is:\n\"Replacement part being rushed with all possible speed.\"\n"},{"id":"297125","messageId":"20061121200508.GB22461@spearce.org","threadId":"43263","inReplyTo":"456359E2.8010403@cc.jyu.fi","subject":"Re: Some tips for doing a CVS importer","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-11-21T20:05:08Z","receivedAt":"2006-11-21T20:05:08Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"lamikr <lamikr@cc.jyu.fi> wrote:\n> Shawn Pearce wrote:\n> >   - No GUI.\n> >   \n> QGIT allows using some commands. I plan to try out the GIT eclipse\n> plugin in near future myself.\n> This mail list have some discussion and download link to it's repo in\n> archives.\n> (title: Java GIT/Eclipse GIT version 0.1.1, )\n\nI'm the author of that plugin.  :-)\n\nIts not even capable of making a commit yet.  The underling plumbing\n(aka jgit) can make commits but the Eclipse GUI has no function to\nactually invoke that plumbing and make a commit to the repository.\n\nThe Eclipse plugin has apparently been a low priority for me.\nI haven't worked on it very recently.  Robin Rosenburg has supposedly\ngotten the revision compare interface to work, but its slow as a\nduck in November due to jgit's pack reading code not running as\nfast as it should.\n \n-- \n"},{"id":"298491","messageId":"20061121201505.GC22461@spearce.org","threadId":"43263","inReplyTo":"20061121200341.GH7201@pasky.or.cz","subject":"Re: Some tips for doing a CVS importer","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-11-21T20:15:05Z","receivedAt":"2006-11-21T20:15:05Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Petr Baudis <pasky@suse.cz> wrote:\n> On Tue, Nov 21, 2006 at 07:39:35AM CET, Shawn Pearce wrote:\n> > I think there's a number of issues that are keeping people from\n> > switching to Git and are instead causing them to choose SVN, hg\n> > or Monotone:\n> > \n> >   - No GUI.\n> \n> It has been my impression that Git's situation is far better than in\n> case of the other systems (except SVN: TortoiseSVN and RapidSVN). Is\n> that not so?\n\nHmm.\n\nhg has a browser (hgk).  Its a direct port of gitk.  I don't see\na GUI otherwise, such as qgit or git-gui.  They do however have a\nWindows installer.\n\nMonotone has mtsh and guitone.  Neither appear to be as far along\nas say qgit or even git-gui, which isn't that far along at all.\n\nSo I guess you are right.  Git's situation is better than that\nof hg or Monotone.  Now if only I can finish everything I want\nto put into git-gui, and get it included as part of the core Git\ndistribution.  :)\n\n-- \n"},{"id":"294538","messageId":"Pine.LNX.4.63.0611212116340.26827@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"43263","inReplyTo":"20061121200341.GH7201@pasky.or.cz","subject":"Re: Some tips for doing a CVS importer","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-11-21T20:22:57Z","receivedAt":"2006-11-21T20:22:57Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Tue, 21 Nov 2006, Petr Baudis wrote:\n\n> On Tue, Nov 21, 2006 at 07:39:35AM CET, Shawn Pearce wrote:\n> > \n> > Yea, that lack of native win32 seems to be one of a number of\n> > blockers for people switching their projects onto Git.\n> \n> Yep. :-(\n\nI started playing with MinGW, and got it to compile and run, with some \nfeatures lacking. See\n\nMessage-ID: <Pine.LNX.4.63.0609021724110.28360@wbgn013.biozentrum.uni-wuerzburg.de>\n\nfor details. From TFM\n\n: The two biggest obstacles are fork() and the network stuff (I do not \n: plan on supporting Git.pm there). To overcome the absence of fork() I \n: wanted to use the subprocess stuff in MinGW's port of GNU make.\n\n\n> > I think there's a number of issues that are keeping people from\n> > switching to Git and are instead causing them to choose SVN, hg\n> > or Monotone:\n> > \n> >   - No GUI.\n> \n> It has been my impression that Git's situation is far better than in\n> case of the other systems (except SVN: TortoiseSVN and RapidSVN). Is\n> that not so?\n\nI also started playing with writing a shell extension (this is what custom \ncontext menu entries are called in Windows) using only MinGW, and no \npayware (except, of course, Windows).\n\nSince both of these little projects were sidetracks from what I am really \nsupposed to do, I will not be able to continue on these on a regular \nbasis. Get somebody else interested, though, and I will be glad to help!\n\nCiao,\nDscho\n"},{"id":"297935","messageId":"46a038f90611211240u4e493f46i2cc46ab780e6c49b@mail.gmail.com","threadId":"43263","inReplyTo":"20061121200341.GH7201@pasky.or.cz","subject":"Re: Some tips for doing a CVS importer","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-11-21T20:40:28Z","receivedAt":"2006-11-21T20:40:28Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 11/22/06, Petr Baudis <pasky@suse.cz> wrote:\n> >   - No GUI.\n>\n> It has been my impression that Git's situation is far better than in\n> case of the other systems (except SVN: TortoiseSVN and RapidSVN). Is\n> that not so?\n\nI think GIT is in pretty good shape in all the items mentioned Shawn\nlists except the Win32 port.\n\n     Confusing doco? All of them ;-)\n     Push/pull terminology confusion -- all of them again.\n\nMy only thing is that I continue to teach Cogito instead of GIT\nbecause the index is a great thing for a top-level maintainer of a\nlarge project but it really offers almost next to nothing to a user\nwho wants to make a commit.\n\nbut that hasn't stopped adoption over here...\n\ncheers,\n\n\n\n"},{"id":"295647","messageId":"45656576.E10FA81D@eudaptics.com","threadId":"43263","inReplyTo":"Pine.LNX.4.63.0611212116340.26827@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Some tips for doing a CVS importer","fromName":"Johannes Sixt","fromEmail":"j.sixt@eudaptics.com","sentAt":"2006-11-23T09:10:14Z","receivedAt":"2006-11-23T09:10:14Z","isPatch":false,"sender":{"key":"j6t@kdbg.org","avatar":"https://avatars.githubusercontent.com/u/14810926?v=4"},"body":"Johannes Schindelin wrote:\n> I started playing with MinGW, and got it to compile and run, with some\n> features lacking. See\n> \n> Message-ID: <Pine.LNX.4.63.0609021724110.28360@wbgn013.biozentrum.uni-wuerzburg.de>\n> \n> for details. From TFM\n> \n> : The two biggest obstacles are fork() and the network stuff (I do not\n> : plan on supporting Git.pm there). To overcome the absence of fork() I\n> : wanted to use the subprocess stuff in MinGW's port of GNU make.\n\nI'd like to do something about it. Is your work accessible in some way?\n\nAt the moment I'm limping along with CVS on Windows, which really is the\nwrong tool for my current task (CVS I mean, not Windows ;)\n\n-- Hannes\n"},{"id":"295014","messageId":"200611232045.06974.robin.rosenberg.lists@dewire.com","threadId":"43263","inReplyTo":"20061121200508.GB22461@spearce.org","subject":"Re: Some tips for doing a CVS importer","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2006-11-23T19:45:06Z","receivedAt":"2006-11-23T19:45:06Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"tisdag 21 november 2006 21:05 skrev Shawn Pearce:\n> lamikr <lamikr@cc.jyu.fi> wrote:\n> > Shawn Pearce wrote:\n> > >   - No GUI.\n> >\n> > QGIT allows using some commands. I plan to try out the GIT eclipse\n> > plugin in near future myself.\n> > This mail list have some discussion and download link to it's repo in\n> > archives.\n> > (title: Java GIT/Eclipse GIT version 0.1.1, )\n>\n> I'm the author of that plugin.  :-)\n>\n> Its not even capable of making a commit yet.  The underling plumbing\n> (aka jgit) can make commits but the Eclipse GUI has no function to\n> actually invoke that plumbing and make a commit to the repository.\n>\n> The Eclipse plugin has apparently been a low priority for me.\n> I haven't worked on it very recently.  Robin Rosenburg has supposedly\n> gotten the revision compare interface to work, but its slow as a\n> duck in November due to jgit's pack reading code not running as\n> fast as it should.\n\nSlow it is. It is somewhat usable though, especially the quickdiff. I worked \nthe whole day with help from quickdiff today. The diff is computed against \nHEAD^ (i.e. I get to see the changes that my topmost StGit patch introduces).\n\nThe project contains 20000+ files and six years of history.  Reading the whole \nhistory is out of the question with the current performance so I restrict \nreading to 500 entries which is just about bearable. That's enough for \npractical use with quickdiff and compare though. Improving jgit's speed 50 \ntimes will probably be enough to make jgit shine. \n\nActivating the Git connection seems to be a problem with the egit projects, \ni.e. it works sometimes, but not with my much bigger repo. The only problem \nis that the first time is dog slow. The structure is different though, as my \nrepo has .project at the top, not one level down.\n\n"},{"id":"295771","messageId":"20061125065949.GD4528@spearce.org","threadId":"43263","inReplyTo":"200611232045.06974.robin.rosenberg.lists@dewire.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-11-25T06:59:49Z","receivedAt":"2006-11-25T06:59:49Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Robin Rosenberg <robin.rosenberg.lists@dewire.com> wrote:\n> Slow it is. It is somewhat usable though, especially the quickdiff. I worked \n> the whole day with help from quickdiff today. The diff is computed against \n> HEAD^ (i.e. I get to see the changes that my topmost StGit patch introduces).\n\nThat's good to hear!\n \n> The project contains 20000+ files and six years of history.  Reading the whole \n> history is out of the question with the current performance so I restrict \n> reading to 500 entries which is just about bearable. That's enough for \n> practical use with quickdiff and compare though. Improving jgit's speed 50 \n> times will probably be enough to make jgit shine. \n\nYes.  I have a plan on how to rewrite the pack reading code which\nshould help somewhat here.  There's some fundamental limitations\nof Java though that are going to keep us from performing as well\nas core-Git does (due to the object memory overheads) but I would\nlike to get close.  :-)\n\njgit also has a few quirks still.  For example it assumes everything\nis encoded as UTF-8 but this isn't true.  The encoding is project\nspecific and can be set by any user, which isn't that portable.\nThis is a problem for jgit and I need to go back and refactor the\nparsing code...\n\nI'd like to get back to jgit sometime in mid-Decemeber.  I'm trying\nto push through git-gui first.  :-)\n\n> Activating the Git connection seems to be a problem with the egit projects, \n> i.e. it works sometimes, but not with my much bigger repo. The only problem \n> is that the first time is dog slow. The structure is different though, as my \n> repo has .project at the top, not one level down.\n\nHmm.  That's a bug.  Sounds like a thread timing issue if it works\nsometimes, as the logic should be completely deterministic.\n\n-- \n"},{"id":"294678","messageId":"456969DA.6090702@gmx.net","threadId":"43263","inReplyTo":"9e4733910611201753m392b5defpb3eb295a075be789@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Marko Macek","fromEmail":"marko.macek@gmx.net","sentAt":"2006-11-26T10:18:02Z","receivedAt":"2006-11-26T10:18:02Z","isPatch":false,"sender":{"key":"marko.macek@gmx.net","avatar":null},"body":"Jon Smirl wrote:\n\n> \n> SVN hides the mini branch by creating a symbol like this:\n> \n> Symbol XXX, change set 70\n> copy All from change set 50\n> copy file A from change set 55\n> copy file B,C from change set 60\n> copy file D from change set 61\n> copy file E,F,G from change set 63\n> copy file H from change set 67\n> \n> It has to do all of those copies because the change sets weren't\n> constructed while taking symbol dependency information into account.\n> \n> Symbol XXX can't copy from change set 69 because commits from after\n> the symbol was created are included in change sets 51-69.\n\nSometimes it is not actually possible to have a 'simple' symbol, even \nby following proper symbol dependencies. \n\nSome situations:\n- tags on some files are readjusted later, or tagged separately with an older\n version\n- tag is created with a -D \"date\" and the file times are not in sync\n- tag is created from a mixed-revision working copy\n\nWhile in the cases of 'time warp' the revision sequence should be \nconsidered more important than timestamps, this is not necessarily\ntrue for tags, since it's easily possible to create them on mixed \nrevisions.\n\ncvs2svn also has a problem with vendor branches because it creates\ntags/branches that contain files from vendor branch by copying some\nfiles from the trunk and other files from the vendor branch.\nIf the vendor branch/tag was only used for the initial import, \nit's IMO best to skip them in the conversion (this needs a patch).\nThere are however problems because keyword expansion causes file\ndifferences.\n\nIt seems that mozilla CVS repository has vendor branches/imports in\nsome parts of the tree.\n\nMark\n"},{"id":"295925","messageId":"9e4733910611260735g2b18e9d1p51a0dca153282cc7@mail.gmail.com","threadId":"43263","inReplyTo":"456969DA.6090702@gmx.net","subject":"Re: Some tips for doing a CVS importer","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-11-26T15:35:44Z","receivedAt":"2006-11-26T15:35:44Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 11/26/06, Marko Macek <marko.macek@gmx.net> wrote:\n> Jon Smirl wrote:\n>\n> >\n> > SVN hides the mini branch by creating a symbol like this:\n> >\n> > Symbol XXX, change set 70\n> > copy All from change set 50\n> > copy file A from change set 55\n> > copy file B,C from change set 60\n> > copy file D from change set 61\n> > copy file E,F,G from change set 63\n> > copy file H from change set 67\n> >\n> > It has to do all of those copies because the change sets weren't\n> > constructed while taking symbol dependency information into account.\n> >\n> > Symbol XXX can't copy from change set 69 because commits from after\n> > the symbol was created are included in change sets 51-69.\n>\n> Sometimes it is not actually possible to have a 'simple' symbol, even\n> by following proper symbol dependencies.\n>\n> Some situations:\n> - tags on some files are readjusted later, or tagged separately with an older\n>  version\n> - tag is created with a -D \"date\" and the file times are not in sync\n> - tag is created from a mixed-revision working copy\n\nI agree that there are a few exceptions to making simple symbols. But\nthe current cvs2svn makes no attempt at all to preserve simple\nsymbols. In my attempts at converting Mozilla 60% of the symbols ended\nup as tiny branches. I investigated a couple by hand and was able to\nrearrange things to create simple symbols in every case I looked at.\n\nThis can be dealt with during the topological sort. If there are\ncomplex symbol creations you will end up with loops during the sort\nprocess. At that point you need to start breaking up change sets to\nremove the loops. You would use a heuristic at this point, something\nlike try breaking up to ten commit change sets to preserve a symbol,\nif you can't preserve it with 10 breaks then break the symbol once and\ntry again, repeat until the loop is gone.\n\nThe current cvs2svn code effectively implements a heuristic when the\ncommits are always preserved at the expense of breaking the symbols.\nSince some commit comments are very common comments (blank ones) those\ncommits get combined into bigger change sets and trash the simple\nsymbols.\n\nAnother note for doing a converter. When combining things into change\nsets, for git import the comments in the branches should not be mixed\nbetween branches and the trunk when detecting change set. Git doesn't\nallow simultaneous commits to the trunk and branches.\n\n> While in the cases of 'time warp' the revision sequence should be\n> considered more important than timestamps, this is not necessarily\n> true for tags, since it's easily possible to create them on mixed\n> revisions.\n>\n> cvs2svn also has a problem with vendor branches because it creates\n> tags/branches that contain files from vendor branch by copying some\n> files from the trunk and other files from the vendor branch.\n> If the vendor branch/tag was only used for the initial import,\n> it's IMO best to skip them in the conversion (this needs a patch).\n> There are however problems because keyword expansion causes file\n> differences.\n>\n> It seems that mozilla CVS repository has vendor branches/imports in\n> some parts of the tree.\n\nI never got around to checking out problems with vendor branches in Mozilla.\n\n\n>\n> Mark\n>\n>\n\n\n-- \nJon Smirl\n"},{"id":"296851","messageId":"4569BCB8.9030809@gmx.net","threadId":"43263","inReplyTo":"9e4733910611260735g2b18e9d1p51a0dca153282cc7@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Marko Macek","fromEmail":"marko.macek@gmx.net","sentAt":"2006-11-26T16:11:36Z","receivedAt":"2006-11-26T16:11:36Z","isPatch":false,"sender":{"key":"marko.macek@gmx.net","avatar":null},"body":"Jon Smirl wrote:\n\n> Another note for doing a converter. When combining things into change\n> sets, for git import the comments in the branches should not be mixed\n> between branches and the trunk when detecting change set. Git doesn't\n> allow simultaneous commits to the trunk and branches.\n\nYup, this is the current problem I'm facing now. Even for CVS->SVN conversion,\nI don't want to see multi-branch commits.\n\n"},{"id":"294174","messageId":"9e4733910611260951u59599a16xd9ac6d37272f825f@mail.gmail.com","threadId":"43263","inReplyTo":"4569BCB8.9030809@gmx.net","subject":"Re: Some tips for doing a CVS importer","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-11-26T17:51:54Z","receivedAt":"2006-11-26T17:51:54Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 11/26/06, Marko Macek <marko.macek@gmx.net> wrote:\n> Jon Smirl wrote:\n>\n> > Another note for doing a converter. When combining things into change\n> > sets, for git import the comments in the branches should not be mixed\n> > between branches and the trunk when detecting change set. Git doesn't\n> > allow simultaneous commits to the trunk and branches.\n>\n> Yup, this is the current problem I'm facing now. Even for CVS->SVN conversion,\n> I don't want to see multi-branch commits.\n\nThere is a command line option on cvs2svn to isolate the branches. I\ngot him to add it as part of the attempt at doing git support.\n\n>\n> Mark\n>\n\n\n-- \nJon Smirl\n"},{"id":"295611","messageId":"456ACAF3.1050608@alum.mit.edu","threadId":"43263","inReplyTo":"9e4733910611201349s4d08b984g772c64982f148bfa@mail.gmail.com","subject":"Re: Some tips for doing a CVS importer","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2006-11-27T11:24:35Z","receivedAt":"2006-11-27T11:24:35Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"I am currently the main (and pretty much the only) cvs2svn maintainer.\nDevelopment has been proceeding more slowly lately because (1) I'm very\nbusy with my day job, and (2) nobody has stepped forward to help.\n\nJon Smirl wrote:\n> #1) There needs to be a tool that can accurately import the\n> repository. cvs2svn does not do this. The good programmers working on\n> git could probably whip this out in a week or two if they wanted to.\n> cvs2svn is very close but they refuse to solve the symbol dependency\n> problem.\n\nJon, I wish you wouldn't portray as obstinacy what is simply a lack of\nresources.  I would like very much to support other cvs2svn output\nformats.  I think it would be great if other projects could benefit from\nour work.  Most of the work I've been doing on cvs2svn lately has been\ntowards supporting other output SCMs.\n\nJon Smirl wrote:\n> I gave up on my cvs2git code, cvs2svn has been refactored so badly\n> that it was too much trouble tracking. It would be easier to write it\n> again. Most of the smarts from the import process is in the\n> git-fastimport code which Shawn has. cvs2svn underwent a major\n> algorithm change after I wrote the first version of git2svn.\n\nI hope that by \"badly\" you mean \"extensively\" and not \"poorly\" :-\\  If\nyou mean \"poorly\", then I'd like to hear your feedback/suggestions.\n\nA large amount of refactoring has been needed to make the change to\ndependency-based conversion possible, and a lot more to help support\ndifferent output formats.  I understand that this causes difficulties\nfor people trying to do parallel development, but most of the\nrefactoring was done before your first appearance on the cvs2svn mailing\nlists.  If you had let us know what you were working on, I would have\navoided making conflicting changes (as I did with Oswald Buddenhagen's\ncommit-dependencies changes).\n\nJon Smirl wrote:\n> I have tried all of the available CVS importers. None of them are\n> without problems. If anyone is interested in writing one for git here\n> are some ideas on how to structure it.\n> \n> 1) there is a working lex/yacc for CVS in the parsecvs source code\n> 2) The first time you parse a CVS file record everything and don't\n> parse it again.\n> 3) When the file is first parsed use the deltas to generate the\n> revisions and feed them to git-fastimport, just remember the SHA1 or\n> an id in the import code. This is a critical step to getting decent\n> performance.\n> 4) If you do #1 and #2 you don't need to store CVS revision numbers\n> and file names in memory. Because of that you can can easily do a\n> Mozilla import in 2GB, probably 1GB.\n> 5) When comparing CVS revisions only use the CVS timestamps as a last\n> resort, instead use the dependency information in the CVS file\n> 6) Match up commits by using an sha1 of the author and commit message\n> 7) After all files are loaded, match up the symbols and insert them\n> into the dependency chains, if any of the symbols depend on a branch\n> commit the symbol lies on the branch, otherwise the symbol is on the\n> trunk,\n> 8) Do a topological sort to build the change set commit tree\n> 9) when you hit a loop in the tree break up delta change sets until\n> the loop can be removed, don't break up symbol change sets.\n> 10) Mozilla has some large commits that were made over dial up. Commit\n> change sets can span hours. All of these commits need to be merged\n> into a single change set.\n> 11) An algorithm needs to be developed for detecting branches merging\n> back into the trunk\n> 12) cvs2svn has excellent test cases, use them to test the new\n> importer. The cvs2svn code is quite nice but it doesn't handle #7\n\nMost of this is possible now using cvs2svn, but it is not enough.\n\nBut first there is a problem with your point #9.  It is in general not\npossible to avoid breaking up symbol changesets, even if you are willing\nto massacre the revision changesets.  CVS allows cases like this:\n\nfile1:\n\n    1.1\n    1.2 ----> branch \"A\"\n              1.2.0.1\n              1.2.0.2 ----> branch \"B\"\n\nfile2:\n\n    1.1\n    1.2 ----> branch \"B\"\n              1.2.0.1\n              1.2.0.2 ----> branch \"A\"\n\nClearly there is no way to create symbols \"A\" and \"B\" both in a single\nchangeset.\n\nBut even disallowing cases like the one above, it is often very\nquestionable whether you want to avoid breaking up symbol commits at all\ncosts.  For example, CVS allows\n\n\nJanuary:     file1<1.1>               file2<1.1>\nFebruary:    file1<1.1> tagged \"T\"\nMarch:       file1<1.2>\nNovember:                             file2<1.2>\nDecember:                             file2<1.2> tagged \"T\"\n\nIn such a case, the only way to avoid splitting up the creation of tag\n\"T\" would be to pretend that the commit file1<1.2> didn't occur in March\nbut rather in November.\n\nThe bottom line is that cvs2svn should do a better job of handling\nsymbols, but even then the git importer will necessarily have to deal\nwith some unusual CVS cases.\n\n> Processing the symbols is integral to deciding how to build the change\n> sets. Right now cvs2svn ignores the symbol dependency information and\n> builds the change sets in a way that forces the mini-branches. That\n> causes 60% of the 2,000 symbols in Mozilla CVS to end up as little\n> branches. Look at the three commit example in the other thread to see\n> exactly what the problem is.\n>\n> SVN hides the mini branch by creating a symbol like this:\n>\n> Symbol XXX, change set 70\n> copy All from change set 50\n> copy file A from change set 55\n> copy file B,C from change set 60\n> copy file D from change set 61\n> copy file E,F,G from change set 63\n> copy file H from change set 67\n>\n> It has to do all of those copies because the change sets weren't\n> constructed while taking symbol dependency information into account.\n>\n> Symbol XXX can't copy from change set 69 because commits from after\n> the symbol was created are included in change sets 51-69.\n\nThe vast majority of the mixed-source symbol creations have nothing to\ndo with honoring symbol dependencies, but rather with the fact the\ncvs2svn is not so clever about deducing which branch should be used as\nthe source for a symbol (CVS often does not record this information\nunambiguously).\n\nChanges needed for git import:\n\nThe symbol dependency problem that Jon has focused on is IMO just the\nleast significant of three main changes that have to be made to support\ngit output from cvs2svn:\n\n1. The symbol dependency problem.  Occasionally symbols are created in\nan order that is inconsistent with the CVS dependency graph.  We want to\nfix this in any case (even for SVN).  Work done so far: the symbol\ndependency graph is already generated and recorded when the repository\nis parsed, and the symbol dependencies are carried through the\nconversion (though not yet used).\n\n2. Symbols are often created using multiple branches as sources, when\nthey could be created from a single branch.  This happens because in\nmany cases CVS doesn't record unambiguously which branch was tagged, and\ncvs2svn's heuristics are not especially clever.  A patch has been\nsubmitted to fix this problem, but unfortunately it doesn't apply to\nHEAD anymore.  See\n\nhttp://cvs2svn.tigris.org/servlets/ReadMsg?list=dev&msgNo=1441\n\nfor a discussion.  (The main difficulty with picking better sources for\nsymbols is that the obvious approaches all require tons of intermediate\nstorage.)  I am currently trying to understand symbol handling in\ncvs2svn well enough that I can port the patch to trunk.\n\n3. The default current output format of cvs2svn is a single dump file\nwith file revisions in commit order.  For the distributed SCMs, it is\nusually far more efficient to generate the file revisions file-by-file\n(non-chronologically) during the initial parse of the CVS files, and\nrefer to the revisions by hash for the rest of the conversion.  In\nOctober I added a bunch of hooks to cvs2svn to make this possible.  Work\nremaining: code to reconstruct file text from CVS text + deltas,\nincluding proper handling of line-end conventions and keyword\nexpansion/unexpansion, and of course the code to output the\nreconstructed snapshots in a git-consumable format.\n"},{"id":"295836","messageId":"456ACC02.6090508@alum.mit.edu","threadId":"43263","inReplyTo":"4569BCB8.9030809@gmx.net","subject":"Re: Some tips for doing a CVS importer","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2006-11-27T11:29:06Z","receivedAt":"2006-11-27T11:29:06Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Marko Macek wrote:\n>> Another note for doing a converter. When combining things into change\n>> sets, for git import the comments in the branches should not be mixed\n>> between branches and the trunk when detecting change set. Git doesn't\n>> allow simultaneous commits to the trunk and branches.\n> \n> Yup, this is the current problem I'm facing now. Even for CVS->SVN\n> conversion,\n> I don't want to see multi-branch commits.\n\nTo avoid multi-branch commits, you have to start cvs2svn with an\n--options file, and in the options file set\n\nctx.cross_project_commits = False\n\n"},{"id":"296365","messageId":"456AD137.8060209@bluegap.ch","threadId":"43263","inReplyTo":"456ACAF3.1050608@alum.mit.edu","subject":"Re: Some tips for doing a CVS importer","fromName":"Markus Schiltknecht","fromEmail":"markus@bluegap.ch","sentAt":"2006-11-27T11:51:19Z","receivedAt":"2006-11-27T11:51:19Z","isPatch":false,"sender":{"key":"markus@bluegap.ch","avatar":"https://gravatar.com/avatar/2f3aadbc46c7c942fa9301d0bb9b91da8c7186233f4c7eff9667a8b53c8cc82e?d=mp&s=160"},"body":"Hi,\n\nMichael Haggerty wrote:\n> I am currently the main (and pretty much the only) cvs2svn maintainer.\n> Development has been proceeding more slowly lately because (1) I'm very\n> busy with my day job, and (2) nobody has stepped forward to help.\n\nI understand very well. Same for me here with monotone's cvs_import vs. \nmy day job... and then I also have a life ;-)\n\n> Jon, I wish you wouldn't portray as obstinacy what is simply a lack of\n> resources.  I would like very much to support other cvs2svn output\n> formats.  I think it would be great if other projects could benefit from\n> our work.  Most of the work I've been doing on cvs2svn lately has been\n> towards supporting other output SCMs.\n\nReally? Hm. I'm somehow sorry for not joining cvs2svn but running my own \nthing with monotone. But I really think it took me less time. OTOH, I'm \nfar from finished, yet...\n\nAnyway, I've made an attempt at solving the 'picking better sources for\nsymbols'-problem:\n\nDuring parsing of all the *,v files, where I'm collecting events \n(commits, branching and tagging) into blobs, I do also remember \n'possible parent branches' for all the symbols (tag and branch events).\n\nAfter that and *before* the blob sorting, I check all blobs and try to \nfind one single parent branch for them. In the best case, those symbol \nblobs do have exactly one possible parent branch, then I just pick that \none. If there are multiple possible parents, I try to pick the deepest. \nAs branches are symbols themselves, I have to run that multiple times \nuntil all symbols are resolved.\n\nAn example: having branches ROOT -> A -> B -> C (branched in that order) \nplus a branch D derived from branch A.\n\nThe symbol blob for branch A: has only one possible parent: ROOT. Thus I \nassign A->parent_branch = ROOT.\n\nNext comes the blob for branch C: it has two possible parents: branch B \nand branch A. At that point we know that A is derived from ROOT, but we \ndon't have assigned a parent to B, yet. Thus we can not resolve C this time.\n\nThen comes branch B: one parent: A. Mark it.\n\nNext round, we process C again: this time, we know B is branched from A. \nThus we can remove the possible parent A. Leaving only one possible \nparent branch: B.\n\nNow, say we have a tag 'X', which ended up in a blob having A, B, C and \nD as possible parent branches. I currently remove A and B, as they are \nparents of C. But C and D still remain and conflict. I'm unable to \nresolve that symbol. I'm thinking about leaving such conflicts to the \nuser to resolve.\n\nI've not yet tested this algorithm extensively. Most larger repositories \nseem to fail somewhere, but not necessarily because of that symbol \nresolving algorithm... :-(\n\nAny comments? Questions? Ideas? I hope to have explained clearly...\n\nAnd I wish you all a lot of time for your open source projects and your \nfamilies, friends, wifes, girl-friends, etc...! ;-)\n\nRegards\n\nMarkus\n"},{"id":"296642","messageId":"9e4733910611270720y39767623he59f919b94e70e16@mail.gmail.com","threadId":"43263","inReplyTo":"456ACAF3.1050608@alum.mit.edu","subject":"Re: Some tips for doing a CVS importer","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-11-27T15:20:09Z","receivedAt":"2006-11-27T15:20:09Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 11/27/06, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> I am currently the main (and pretty much the only) cvs2svn maintainer.\n> Development has been proceeding more slowly lately because (1) I'm very\n> busy with my day job, and (2) nobody has stepped forward to help.\n>\n> Jon Smirl wrote:\n> > #1) There needs to be a tool that can accurately import the\n> > repository. cvs2svn does not do this. The good programmers working on\n> > git could probably whip this out in a week or two if they wanted to.\n> > cvs2svn is very close but they refuse to solve the symbol dependency\n> > problem.\n>\n> Jon, I wish you wouldn't portray as obstinacy what is simply a lack of\n> resources.  I would like very much to support other cvs2svn output\n> formats.  I think it would be great if other projects could benefit from\n> our work.  Most of the work I've been doing on cvs2svn lately has been\n> towards supporting other output SCMs.\n\ncvs2avn is a nice piece of code, it is a worthy goal to have a\nuniveral conversion tool.\n\n>\n> Jon Smirl wrote:\n> > I gave up on my cvs2git code, cvs2svn has been refactored so badly\n> > that it was too much trouble tracking. It would be easier to write it\n> > again. Most of the smarts from the import process is in the\n> > git-fastimport code which Shawn has. cvs2svn underwent a major\n> > algorithm change after I wrote the first version of git2svn.\n>\n> I hope that by \"badly\" you mean \"extensively\" and not \"poorly\" :-\\  If\n> you mean \"poorly\", then I'd like to hear your feedback/suggestions.\n\nExtensively, the dependency rewrite changed things some much that my\npatches were basically worthless. I tried merging them and gave up, it\nwould be more efficient to rewrite them or builld hooks in the right\nplaces.\n\n>\n> A large amount of refactoring has been needed to make the change to\n> dependency-based conversion possible, and a lot more to help support\n> different output formats.  I understand that this causes difficulties\n> for people trying to do parallel development, but most of the\n> refactoring was done before your first appearance on the cvs2svn mailing\n> lists.  If you had let us know what you were working on, I would have\n> avoided making conflicting changes (as I did with Oswald Buddenhagen's\n> commit-dependencies changes).\n>\n> Jon Smirl wrote:\n> > I have tried all of the available CVS importers. None of them are\n> > without problems. If anyone is interested in writing one for git here\n> > are some ideas on how to structure it.\n> >\n> > 1) there is a working lex/yacc for CVS in the parsecvs source code\n> > 2) The first time you parse a CVS file record everything and don't\n> > parse it again.\n> > 3) When the file is first parsed use the deltas to generate the\n> > revisions and feed them to git-fastimport, just remember the SHA1 or\n> > an id in the import code. This is a critical step to getting decent\n> > performance.\n> > 4) If you do #1 and #2 you don't need to store CVS revision numbers\n> > and file names in memory. Because of that you can can easily do a\n> > Mozilla import in 2GB, probably 1GB.\n> > 5) When comparing CVS revisions only use the CVS timestamps as a last\n> > resort, instead use the dependency information in the CVS file\n> > 6) Match up commits by using an sha1 of the author and commit message\n> > 7) After all files are loaded, match up the symbols and insert them\n> > into the dependency chains, if any of the symbols depend on a branch\n> > commit the symbol lies on the branch, otherwise the symbol is on the\n> > trunk,\n> > 8) Do a topological sort to build the change set commit tree\n> > 9) when you hit a loop in the tree break up delta change sets until\n> > the loop can be removed, don't break up symbol change sets.\n> > 10) Mozilla has some large commits that were made over dial up. Commit\n> > change sets can span hours. All of these commits need to be merged\n> > into a single change set.\n> > 11) An algorithm needs to be developed for detecting branches merging\n> > back into the trunk\n> > 12) cvs2svn has excellent test cases, use them to test the new\n> > importer. The cvs2svn code is quite nice but it doesn't handle #7\n>\n> Most of this is possible now using cvs2svn, but it is not enough.\n>\n> But first there is a problem with your point #9.  It is in general not\n> possible to avoid breaking up symbol changesets, even if you are willing\n> to massacre the revision changesets.  CVS allows cases like this:\n\nWe don't know how often this case occurs until more alogirthms are\ntried. All I know is that 60% of the Mozilla symbols end up needing\ncopies. And for the few cases I decoded things by hand I was able to\nrearrange things so that copies were not needed. It is likely that\nsome symbols in Mozilla will need copies to construct them, it is a\nquestion of degree, I don't believe copies are required for 60% of the\nsymbols.\n\n>\n> file1:\n>\n>     1.1\n>     1.2 ----> branch \"A\"\n>               1.2.0.1\n>               1.2.0.2 ----> branch \"B\"\n>\n> file2:\n>\n>     1.1\n>     1.2 ----> branch \"B\"\n>               1.2.0.1\n>               1.2.0.2 ----> branch \"A\"\n>\n> Clearly there is no way to create symbols \"A\" and \"B\" both in a single\n> changeset.\n>\n> But even disallowing cases like the one above, it is often very\n> questionable whether you want to avoid breaking up symbol commits at all\n> costs.  For example, CVS allows\n>\n>\n> January:     file1<1.1>               file2<1.1>\n> February:    file1<1.1> tagged \"T\"\n> March:       file1<1.2>\n> November:                             file2<1.2>\n> December:                             file2<1.2> tagged \"T\"\n>\n> In such a case, the only way to avoid splitting up the creation of tag\n> \"T\" would be to pretend that the commit file1<1.2> didn't occur in March\n> but rather in November.\n>\n> The bottom line is that cvs2svn should do a better job of handling\n> symbols, but even then the git importer will necessarily have to deal\n> with some unusual CVS cases.\n\nThe unusal cases can be made into branches. If I remember correctly\nMozilla has about 300 symbols with \"BRANCH\" in the name. But the\nconverted repositories are ending up with over 2,000 branches. When\nyou load this into the git visualization tools it is obvious that the\nbowl of spaghetti caused by 2,000 branches is not a repository a human\nwould have created.\n\n>\n> > Processing the symbols is integral to deciding how to build the change\n> > sets. Right now cvs2svn ignores the symbol dependency information and\n> > builds the change sets in a way that forces the mini-branches. That\n> > causes 60% of the 2,000 symbols in Mozilla CVS to end up as little\n> > branches. Look at the three commit example in the other thread to see\n> > exactly what the problem is.\n> >\n> > SVN hides the mini branch by creating a symbol like this:\n> >\n> > Symbol XXX, change set 70\n> > copy All from change set 50\n> > copy file A from change set 55\n> > copy file B,C from change set 60\n> > copy file D from change set 61\n> > copy file E,F,G from change set 63\n> > copy file H from change set 67\n> >\n> > It has to do all of those copies because the change sets weren't\n> > constructed while taking symbol dependency information into account.\n> >\n> > Symbol XXX can't copy from change set 69 because commits from after\n> > the symbol was created are included in change sets 51-69.\n>\n> The vast majority of the mixed-source symbol creations have nothing to\n> do with honoring symbol dependencies, but rather with the fact the\n> cvs2svn is not so clever about deducing which branch should be used as\n> the source for a symbol (CVS often does not record this information\n> unambiguously).\n>\n> Changes needed for git import:\n>\n> The symbol dependency problem that Jon has focused on is IMO just the\n> least significant of three main changes that have to be made to support\n> git output from cvs2svn:\n>\n> 1. The symbol dependency problem.  Occasionally symbols are created in\n> an order that is inconsistent with the CVS dependency graph.  We want to\n> fix this in any case (even for SVN).  Work done so far: the symbol\n> dependency graph is already generated and recorded when the repository\n> is parsed, and the symbol dependencies are carried through the\n> conversion (though not yet used).\n>\n> 2. Symbols are often created using multiple branches as sources, when\n> they could be created from a single branch.  This happens because in\n> many cases CVS doesn't record unambiguously which branch was tagged, and\n> cvs2svn's heuristics are not especially clever.  A patch has been\n> submitted to fix this problem, but unfortunately it doesn't apply to\n> HEAD anymore.  See\n>\n> http://cvs2svn.tigris.org/servlets/ReadMsg?list=dev&msgNo=1441\n>\n> for a discussion.  (The main difficulty with picking better sources for\n> symbols is that the obvious approaches all require tons of intermediate\n> storage.)  I am currently trying to understand symbol handling in\n> cvs2svn well enough that I can port the patch to trunk.\n\nI'm happy to give new alogorithm a try as they are developed.\n\n>\n> 3. The default current output format of cvs2svn is a single dump file\n> with file revisions in commit order.  For the distributed SCMs, it is\n> usually far more efficient to generate the file revisions file-by-file\n> (non-chronologically) during the initial parse of the CVS files, and\n> refer to the revisions by hash for the rest of the conversion.  In\n> October I added a bunch of hooks to cvs2svn to make this possible.  Work\n> remaining: code to reconstruct file text from CVS text + deltas,\n> including proper handling of line-end conventions and keyword\n> expansion/unexpansion, and of course the code to output the\n> reconstructed snapshots in a git-consumable format.\n\nThis is a major benefit for git conversion, but it hasn't been a big\nissues with the cvs2svn code. Hooks will be helpful.\n\n-- \nJon Smirl\n"},{"id":"295312","messageId":"456B61FE.7060100@alum.mit.edu","threadId":"43263","inReplyTo":"456AD137.8060209@bluegap.ch","subject":"Re: Some tips for doing a CVS importer","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2006-11-27T22:09:02Z","receivedAt":"2006-11-27T22:09:02Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Markus Schiltknecht wrote:\n> Michael Haggerty wrote:\n>> Jon, I wish you wouldn't portray as obstinacy what is simply a lack of\n>> resources.  I would like very much to support other cvs2svn output\n>> formats.  I think it would be great if other projects could benefit from\n>> our work.  Most of the work I've been doing on cvs2svn lately has been\n>> towards supporting other output SCMs.\n> \n> Really? Hm. I'm somehow sorry for not joining cvs2svn but running my own\n> thing with monotone. But I really think it took me less time. OTOH, I'm\n> far from finished, yet...\n\nThere's still time to join forces :-)  \"Far from finished\" on a project\nof this messiness can equal quite a bit of time.\n\nBut even if you want to pursue your own converter, consider visiting\n#cvs2svn on irc.freenode.net if you want to discuss things.\n\n> Anyway, I've made an attempt at solving the 'picking better sources for\n> symbols'-problem:\n\nLet me try to understand this...\n\n> During parsing of all the *,v files, where I'm collecting events\n> (commits, branching and tagging) into blobs, I do also remember\n> 'possible parent branches' for all the symbols (tag and branch events).\n\nThis is the part that can get quite expensive for large repositories, as\nthere can be orders of magnitude more symbol creations than revisions.\nAccording to Daniel Jacobowitz:\n\n> [...] at one point I believe the GCC repository was gaining up\n> to four tags a day (head, two supported release branches, and one\n> vendor branch).  I've been using the principal that the number of tags\n> might be unworkable, but the number of branches generally is not.\n\nThis means that the number of tag events is O(number-of-days *\ntotal-number-of-files-in-repo), where the gcc repo has about 50000\nfiles.  By contrast, only a small fraction of files is typically touched\nin any day.\n\nI've been trying to find a solution that doesn't require quite so much\nspace.  I think that if you allow yourself this much space, the problem\nis not very difficult.\n\n> After that and *before* the blob sorting, I check all blobs and try to\n> find one single parent branch for them. In the best case, those symbol\n> blobs do have exactly one possible parent branch, then I just pick that\n> one. If there are multiple possible parents, I try to pick the deepest.\n> As branches are symbols themselves, I have to run that multiple times\n> until all symbols are resolved.\n> \n> An example: having branches ROOT -> A -> B -> C (branched in that order)\n> plus a branch D derived from branch A.\n\nI assume that you are talking about a situation for which CVS is\nambiguous, like a file with\n\nA = 1.2.2\nB = 1.2.4\nC = 1.2.6\nD = 1.2.2.5.2\n\n> The symbol blob for branch A: has only one possible parent: ROOT. Thus I\n> assign A->parent_branch = ROOT.\n> \n> Next comes the blob for branch C: it has two possible parents: branch B\n> and branch A.\n\nWhy is ROOT not considered as a possible parent of C?\n\n> At that point we know that A is derived from ROOT, but we\n> don't have assigned a parent to B, yet. Thus we can not resolve C this\n> time.\n> \n> Then comes branch B: one parent: A. Mark it.\n> \n> Next round, we process C again: this time, we know B is branched from A.\n> Thus we can remove the possible parent A. Leaving only one possible\n> parent branch: B.\n\nBut the fact that B preceded C chronologically does not mean that C is\nderived from B.  If I branch from ROOT or A after creating branch B, the\nresult as stored in CVS looks exactly the same as if I branch from B\n(unless a file was modified between the creation of the parent branch\nand the creation of the child branch).\n\n> Now, say we have a tag 'X', which ended up in a blob having A, B, C and\n> D as possible parent branches. I currently remove A and B, as they are\n> parents of C. But C and D still remain and conflict. I'm unable to\n> resolve that symbol. I'm thinking about leaving such conflicts to the\n> user to resolve.\n\nFrom your description, this sounds like a tag that cannot be created\nfrom a single parent branch.  Therefore it would have to be cobbled\ntogether from multiple parents.\n\n> I've not yet tested this algorithm extensively. Most larger repositories\n> seem to fail somewhere, but not necessarily because of that symbol\n> resolving algorithm... :-(\n> \n> Any comments? Questions? Ideas? I hope to have explained clearly...\n> \n> And I wish you all a lot of time for your open source projects and your\n> families, friends, wifes, girl-friends, etc...! ;-)\n\n:-) Thanks.  The same to you!\n\n"},{"id":"298299","messageId":"456C5363.6040409@bluegap.ch","threadId":"43263","inReplyTo":"456B61FE.7060100@alum.mit.edu","subject":"Re: Some tips for doing a CVS importer","fromName":"Markus Schiltknecht","fromEmail":"markus@bluegap.ch","sentAt":"2006-11-28T15:18:59Z","receivedAt":"2006-11-28T15:18:59Z","isPatch":false,"sender":{"key":"markus@bluegap.ch","avatar":"https://gravatar.com/avatar/2f3aadbc46c7c942fa9301d0bb9b91da8c7186233f4c7eff9667a8b53c8cc82e?d=mp&s=160"},"body":"Hi,\n\nMichael Haggerty wrote:\n> There's still time to join forces :-)  \"Far from finished\" on a project\n> of this messiness can equal quite a bit of time.\n\nYes. Maybe I'm a little pessimistic ;-)\n\n> But even if you want to pursue your own converter, consider visiting\n> #cvs2svn on irc.freenode.net if you want to discuss things.\n\nThanks, I just happen to not particularly like IRC... I prefer emails\nand mailing lists.\n\n>> During parsing of all the *,v files, where I'm collecting events\n>> (commits, branching and tagging) into blobs, I do also remember\n>> 'possible parent branches' for all the symbols (tag and branch events).\n> \n> This is the part that can get quite expensive for large repositories, as\n> there can be orders of magnitude more symbol creations than revisions.\n> According to Daniel Jacobowitz:\n> \n>> [...] at one point I believe the GCC repository was gaining up\n>> to four tags a day (head, two supported release branches, and one\n>> vendor branch).  I've been using the principal that the number of tags\n>> might be unworkable, but the number of branches generally is not.\n> \n> This means that the number of tag events is O(number-of-days *\n> total-number-of-files-in-repo), where the gcc repo has about 50000\n> files.  By contrast, only a small fraction of files is typically touched\n> in any day.\n\nYeah, 50'000 * 1825 (5 years) * say 100 bytes -> 8GB  sounds like a lot.\nOTOH, I certainly don't need 100 bytes per tag and one tag per day over \nfive years is really a lot. Repositories that large are probably not \nconverted to CVS on an old Pentium III...\n\nI've just tested with the mozilla repository (I don't have the gcc one). \nThe import has been run only through the first two stpes: collecting the \nblobs and symbol resolving. That took almost one and a half hour on my \nCore Duo with 2GB of memory:\n\nreal    85m20.684s\nuser    39m59.082s\nsys     1m32.874s\n\nAnd peak memory consumption was:\n\nVmPeak:  1814024 kB\n\nWhile the mozilla/mozilla cvs repository sums up to 3.1 GB. The monotone \nrepository (which is still lacking the revisions, but has all files and \nfile deltas) is 588MB after that step. I'd guess that once it finishes, \nit would be less than 1GB.\n\n> I've been trying to find a solution that doesn't require quite so much\n> space.  I think that if you allow yourself this much space, the problem\n> is not very difficult.\n\nOkay. As long as I can import it on my laptop I'm fine ;-)\n\n>> After that and *before* the blob sorting, I check all blobs and try to\n>> find one single parent branch for them. In the best case, those symbol\n>> blobs do have exactly one possible parent branch, then I just pick that\n>> one. If there are multiple possible parents, I try to pick the deepest.\n>> As branches are symbols themselves, I have to run that multiple times\n>> until all symbols are resolved.\n>>\n>> An example: having branches ROOT -> A -> B -> C (branched in that order)\n>> plus a branch D derived from branch A.\n> \n> I assume that you are talking about a situation for which CVS is\n> ambiguous, like a file with\n> \n> A = 1.2.2\n> B = 1.2.4\n> C = 1.2.6\n> D = 1.2.2.5.2\n\nWell, almost. I meant a whole repository with these branches. If one\nfile included all the branches it's getting easy to resolve. But for my\nexample, I had something like that in mind:\n\nfileA:\n\nA = 1.2.2\n(no changes for branch B)\nC = 1.2.4      --> makes A a possible parent of branch C\nD = 1.2.2.5.2  --> makes A a possible parent of branch D\nX = 1.2.4      --> makes C a possible parent of tag X\n\nfileB:\n\nA = 1.2.2\nB = 1.2.4      --> makes A a possible parent of branch B\nC = 1.2.6      --> makes B a possible parent of branch C\nD = 1.2.2.5.2  --> makes A a possible parent of branch D\nX = 1.2.2.5.2  --> makes D a possible parent of tag X\n\nfileC:\nA = 1.2.2\nX = 1.2.2      --> makes A a possible parent of tag X\n\nfileD:\nA = 1.2.2\nB = 1.2.4\nX = 1.2.4      --> makes B a possible parent of tag X\n\n>> The symbol blob for branch A: has only one possible parent: ROOT. Thus I\n>> assign A->parent_branch = ROOT.\n>>\n>> Next comes the blob for branch C: it has two possible parents: branch B\n>> and branch A.\n> \n> Why is ROOT not considered as a possible parent of C?\n\nThose were just examples. In my CVS-repository-in-mind, none of the\nfiles were branching from ROOT directly into C.\n\n>> At that point we know that A is derived from ROOT, but we\n>> don't have assigned a parent to B, yet. Thus we can not resolve C this\n>> time.\n>>\n>> Then comes branch B: one parent: A. Mark it.\n>>\n>> Next round, we process C again: this time, we know B is branched from A.\n>> Thus we can remove the possible parent A. Leaving only one possible\n>> parent branch: B.\n> \n> But the fact that B preceded C chronologically does not mean that C is\n> derived from B.\n\nNo. And I don't assume so in any place. Given the files above, I can\nhowever clearly say that C got branched off from B, no?\n\n> If I branch from ROOT or A after creating branch B, the\n> result as stored in CVS looks exactly the same as if I branch from B\n> (unless a file was modified between the creation of the parent branch\n> and the creation of the child branch).\n\nSure. That would result in an unresolvable symbol.\n\n>> Now, say we have a tag 'X', which ended up in a blob having A, B, C and\n>> D as possible parent branches. I currently remove A and B, as they are\n>> parents of C. But C and D still remain and conflict. I'm unable to\n>> resolve that symbol. I'm thinking about leaving such conflicts to the\n>> user to resolve.\n> \n> From your description, this sounds like a tag that cannot be created\n> from a single parent branch.  Therefore it would have to be cobbled\n> together from multiple parents.\n\nRight. I somehow have to cope with those cases, as CVS allows them and\nmonotone does not.\n\nThe main point in my symbol resolving code is trying to uniquely assign\na symbol to one branch wherever possible. And handing cases where this\nis not possible to the user. AFAICT, it does so quite well.\n\nRegards\n\nMarkus\n"},{"id":"298925","messageId":"456E2746.4050707@alum.mit.edu","threadId":"43263","inReplyTo":"456C5363.6040409@bluegap.ch","subject":"Re: Some tips for doing a CVS importer","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2006-11-30T00:35:18Z","receivedAt":"2006-11-30T00:35:18Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Markus Schiltknecht wrote:\n> Michael Haggerty wrote:\n>> This is the part that can get quite expensive for large repositories, as\n>> there can be orders of magnitude more symbol creations than revisions.\n>> According to Daniel Jacobowitz:\n>>\n>>> [...] at one point I believe the GCC repository was gaining up\n>>> to four tags a day (head, two supported release branches, and one\n>>> vendor branch).  I've been using the principal that the number of tags\n>>> might be unworkable, but the number of branches generally is not.\n>>\n>> This means that the number of tag events is O(number-of-days *\n>> total-number-of-files-in-repo), where the gcc repo has about 50000\n>> files.  By contrast, only a small fraction of files is typically touched\n>> in any day.\n> \n> Yeah, 50'000 * 1825 (5 years) * say 100 bytes -> 8GB  sounds like a lot.\n> OTOH, I certainly don't need 100 bytes per tag and one tag per day over\n> five years is really a lot. Repositories that large are probably not\n> converted to CVS on an old Pentium III...\n\n...times 4 (tags per day) -> 32GB.  If I understand correctly, the tags\nwere created nightly by automated scripts.\n\nI admit that this is an extreme example, but the philosophy of the\ncvs2svn project (a philosophy that I inherited from my predecessors, by\nthe way) is to be able to handle the most absurd repositories out there.\n\n> Well, almost. I meant a whole repository with these branches. If one\n> file included all the branches it's getting easy to resolve. But for my\n> example, I had something like that in mind:\n\nI am glad that we are getting into concrete examples.  But your example\nneeds some clarifications (see below).\n\n> fileA:\n> \n> A = 1.2.2\n> (no changes for branch B)\n> C = 1.2.4      --> makes A a possible parent of branch C\n\nIn this case, ROOT can also be C's parent.\n\n> D = 1.2.2.5.2  --> makes A a possible parent of branch D\n\nThis implies that A is *necessarily* the parent of D.  If there were a\nE=1.2.2.5.4, then the parent of E would be ambiguous but the parent of D\nwould still unambiguously be A.\n\n> X = 1.2.4      --> makes C a possible parent of tag X\n\nWait a minute.  A tag always has an even number of integers.  Do you\nmean X=1.2 or X=1.2.4.1?  The same below.\n\n> fileB:\n> \n> A = 1.2.2\n> B = 1.2.4      --> makes A a possible parent of branch B\n\nor ROOT\n\n> C = 1.2.6      --> makes B a possible parent of branch C\n\nor A or ROOT\n\n> D = 1.2.2.5.2  --> makes A a possible parent of branch D\n\nA is unambiguously the parent of D\n\n> X = 1.2.2.5.2  --> makes D a possible parent of tag X\n>\n> fileC:\n> A = 1.2.2\n> X = 1.2.2      --> makes A a possible parent of tag X\n> \n> fileD:\n> A = 1.2.2\n> B = 1.2.4\n> X = 1.2.4      --> makes B a possible parent of tag X\n> \n>>> The symbol blob for branch A: has only one possible parent: ROOT. Thus I\n>>> assign A->parent_branch = ROOT.\n>>>\n>>> Next comes the blob for branch C: it has two possible parents: branch B\n>>> and branch A.\n>>\n>> Why is ROOT not considered as a possible parent of C?\n> \n> Those were just examples. In my CVS-repository-in-mind, none of the\n> files were branching from ROOT directly into C.\n\nIn your example, ROOT *is* a possible parent of C.\n\n>>> At that point we know that A is derived from ROOT, but we\n>>> don't have assigned a parent to B, yet. Thus we can not resolve C this\n>>> time.\n>>>\n>>> Then comes branch B: one parent: A. Mark it.\n\nIn your example, ROOT is also a possible parent of B.\n\n>>> Next round, we process C again: this time, we know B is branched from A.\n>>> Thus we can remove the possible parent A. Leaving only one possible\n>>> parent branch: B.\n>>\n>> But the fact that B preceded C chronologically does not mean that C is\n>> derived from B.\n> \n> No. And I don't assume so in any place. Given the files above, I can\n> however clearly say that C got branched off from B, no?\n\nNo.  C is nowhere unambiguously derived from B, therefore its parent\ncould be ROOT, A, or B.  See my example below.\n\n>> If I branch from ROOT or A after creating branch B, the\n>> result as stored in CVS looks exactly the same as if I branch from B\n>> (unless a file was modified between the creation of the parent branch\n>> and the creation of the child branch).\n> \n> Sure. That would result in an unresolvable symbol.\n> \n>>> Now, say we have a tag 'X', which ended up in a blob having A, B, C and\n>>> D as possible parent branches. I currently remove A and B, as they are\n>>> parents of C. But C and D still remain and conflict. I'm unable to\n>>> resolve that symbol. I'm thinking about leaving such conflicts to the\n>>> user to resolve.\n\nI don't know how to deal with tag X because the numbers that you\nassigned to it above can't be correct.\n\n\nConsider the attached script.  It unambiguously creates branches A1 and\nA2 from ROOT and branch B from A1, then adds tag X on branch B.  But in\nthe files:\n\nfileA symbols\n        X:1.1\n        B:1.1.0.6\n        A2:1.1.0.4\n        A1:1.1.0.2;\n\nfileB symbols\n        X:1.1.2.1\n        B:1.1.2.1.0.2\n        A2:1.1.0.4\n        A1:1.1.0.2;\n\nfileC symbols\n        X:1.1.6.1\n        B:1.1.0.6\n        A2:1.1.0.4\n        A1:1.1.0.2;\n\nfileD symbols\n        X:1.1\n        B:1.1.0.4\n        A2:1.2.0.2\n        A1:1.1.0.2;\n\nNote that from looking at fileA alone, there is no way to tell whether\nA2 was created from ROOT or A1, or whether B was created from ROOT, A1,\nor A2.  And tag X is all over the place, even though for each file it\nwas created from branch B.\n\nIf only information from fileA,v is considered, any of the following\nbranching topologies would give identical fileA,v contents:\n\n      ROOT\n      /|\\\n     / | \\\n    A1 A2 B\n\n\n      ROOT\n      / \\\n     /   \\\n    A1   A2\n    |\n    B\n\n\n      ROOT\n      / \\\n     /   \\\n    A1   A2\n          |\n          B\n\n\n      ROOT\n      / \\\n     /   \\\n    A1    B\n    |\n    A2\n\n\n      ROOT\n       |\n       A1\n      / \\\n     /   \\\n    A2    B\n\n\n      ROOT\n       |\n       A1\n       |\n       A2\n       |\n       B\n\nAnd from the information present in fileA,v, it is not possible to tell\nwhether tag X was applied to ROOT, A1, A2, or B.\n\n(Some topologies *are* ruled out because the revision numbers are\nordered incorrectly; for example:\n\n      ROOT\n       |\n       B\n      / \\\n     /   \\\n    A1   A2\n\n      ROOT\n       |\n       A2\n       |\n       A1\n       |\n       B\n\nare not consistent with fileA,v.)\n\nIf we also consider the information in fileB, it is clear that branch\nB's parent is branch A1, but it is still not clear whether branch A2's\nparent is ROOT or A1, or whether tag X was applied to branch A1, A2, or B.\n\nSimilarly, fileC,v tells us that tag X was applied to branch B, and\nfileD,v tells us that A2's parent is ROOT.\n\nEach file alone is quite ambiguous, but in this case putting the\ninformation from all files together (with the assumption that they have\na mutually-consistent history) is enough to reconstruct the entire\nbranching topology.\n\nWhat's worse in real life?  Each file rules out some possible histories\nand the goal is to find a history that is consistent with all files.  But...\n\n- There can easily be cases where even the total information from all\nfiles is still not enough to choose a unique history.  In such cases we\nneed a way to select between the possible histories.\n\n- Since files in CVS don't necessarily *have* a globally consistent\nbranching/tagging history, heuristics have to be used in such cases to\nfind histories that apply to subsets of the repository in some\nreasonable way (i.e., the one that is most likely considering the way\npeople typically work with CVS).\n\n- \"Unlabeled branches\": often users have removed the label from a\nbranch, but the branch is still used as a source for other branches.\nFiguring out this situation is a real mess.\n\nI imagine that the best results (never mind whether it is practical)\nwould be obtained by recording the topology constraints implied by each\n*,v file, then trying to map the topologies onto each other pair by pair\nto (1) combine the constraints and thereby limit the possible histories\nand (2) deduce which unlabeled branches correspond to one another.  But\nI still don't know how to deal with inconsistent histories.  I think a\nbottom-up approach would be the most sensible, given that people are\nprobably more likely to tag a whole subdirectory rather than files\nscattered here and there.\n\nThe second step is to decide at what point in time a branch or tag\nshould be created, with the goal of being able to create it as a\nsnapshot of the source branch at that moment.  This is not always\npossible, even if the branch topologies are compatible.\n\nMichael\n\n"},{"id":"294363","messageId":"20061130004508.GA22208@nevyn.them.org","threadId":"43263","inReplyTo":"456E2746.4050707@alum.mit.edu","subject":"Re: Some tips for doing a CVS importer","fromName":"Daniel Jacobowitz","fromEmail":"dan@debian.org","sentAt":"2006-11-30T00:45:08Z","receivedAt":"2006-11-30T00:45:08Z","isPatch":false,"sender":{"key":"dan@debian.org","avatar":null},"body":"On Thu, Nov 30, 2006 at 01:35:18AM +0100, Michael Haggerty wrote:\n> ...times 4 (tags per day) -> 32GB.  If I understand correctly, the tags\n> were created nightly by automated scripts.\n\nCorrect.  Remember, checking out a branch from a CVS repository from a\nparticular date was extremely awkward; the tags were the only way to\nhave reproducible snapshots.\n\n-- \nDaniel Jacobowitz\n"}]}