{"thread":{"id":"5549","subject":"Re: cvs import","startedAt":"2006-09-13T19:01:02Z","lastAt":"2006-09-16T06:21:09Z","messageCount":38,"participants":["Jon Smirl","Martin Langhoff","Markus Schiltknecht","Oswald Buddenhagen","Nathaniel Smith","Daniel Carosone","Keith Packard","Shawn Pearce","Michael Haggerty","Jakub Narebski","Petr Baudis"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"26817","messageId":"9e4733910609131201q7f583029r72dac66cd0dd098f@mail.gmail.com","threadId":"5549","inReplyTo":"45084400.1090906@bluegap.ch","subject":"Re: cvs import","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-09-13T19:01:02Z","receivedAt":"2006-09-13T19:01:02Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"Let's copy the git list too and maybe we can come up with one importer\nfor everyone.\n\nOn 9/13/06, Markus Schiltknecht <markus@bluegap.ch> wrote:\n> Hi,\n>\n> I've been trying to understand the cvsimport algorithm used by monotone\n> and wanted to adjust that to be more like the one in cvs2svn.\n>\n> I've had some problems with cvs2svn itself and began to question the\n> algorithm used there. It turned out that the cvs2svn people have\n> discussed an improved algorithms and are about to write a cvs2svn 2.0.\n> The main problem with the current algorithm is that it depends on the\n> timestamp information stored in the CVS repository.\n>\n> Instead, it would be much better to just take the dependencies of the\n> revisions into account. Considering the timestamp an irrelevant (for the\n> import) attribute of the revision.\n>\n> Now, that can be used to convert from CVS to about anything else.\n> Obviously we were discussing about subversion, but then there was git,\n> too. And monotone.\n>\n> I'm beginning to question if one could come up with a generally useful\n> cleaned-and-sane-CVS-changeset-dump-format, which could then be used by\n> importers to all sort of VCSes. This would make monotone's cvsimport\n> function dependent on cvs2svn (and therefore python). But the general\n> try-to-get-something-usefull-from-an-insane-CVS-repository-algorithm\n> would only have to be written once.\n>\n> On the other hand, I see that lots of the cvsimport functionality for\n> monotone has already been written (rcs file parsing, stuffing files,\n> file deltas and complete revisions into the monotone database, etc..).\n> Changing it to a better algorithm does not seem to be _that_ much work\n> anymore. Plus the hard part seems to be to come up with a good\n> algorithm, not implementing it. And we could still exchange our\n> experience with the general algorithm with the cvs2svn people.\n>\n> Plus, the guy who mentioned git pointed out that git needs quite a\n> different dump-format than subversion to do an efficient conversion. I\n> think coming up with a generally-usable dump format would not be that easy.\n>\n> So you see, I'm slightly favoring the second implementation approach\n> with a C++ implementation inside monotone.\n>\n> Thoughts or comments?\n> Sorry, I forgot to mention some pointers:\n>\n> Here is the thread where I've started the discussion about the cvs2svn\n> algorithm:\n> http://cvs2svn.tigris.org/servlets/ReadMsg?list=dev&msgNo=1599\n>\n> And this is a proposal for an algorithm to do cvs imports independant of\n> the timestamp:\n> http://cvs2svn.tigris.org/servlets/ReadMsg?list=dev&msgNo=1451\n>\n> Markus\n>\n> ---------------------------------------------------------------------\n> To unsubscribe, e-mail: dev-unsubscribe@cvs2svn.tigris.org\n> For additional commands, e-mail: dev-help@cvs2svn.tigris.org\n>\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"26822","messageId":"46a038f90609131341se42b2dcne73c017cf757d13a@mail.gmail.com","threadId":"5549","inReplyTo":"9e4733910609131201q7f583029r72dac66cd0dd098f@mail.gmail.com","subject":"Re: cvs import","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-09-13T20:41:29Z","receivedAt":"2006-09-13T20:41:29Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 9/14/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> Let's copy the git list too and maybe we can come up with one importer\n> for everyone.\n\nIt's a really good idea. cvsps has been for a while a (limited, buggy)\nattempt at that. One thing that bothers me in the cvs2svn algorithm is\nthat is not stable in its decisions about where the branching point is\n-- run the import twice at different times and it may tell you that\nthe branching point has moved.\n\nThis is problematic for incremental imports. If we fudged that \"it's\naround *here*\" we better remember we said that and not go changing our\nstory. Git is too smart for that ;-)\n\n\nmartin\n"},{"id":"26828","messageId":"4508724D.2050701@bluegap.ch","threadId":"5549","inReplyTo":"46a038f90609131341se42b2dcne73c017cf757d13a@mail.gmail.com","subject":"Re: cvs import","fromName":"Markus Schiltknecht","fromEmail":"markus@bluegap.ch","sentAt":"2006-09-13T21:04:13Z","receivedAt":"2006-09-13T21:04:13Z","isPatch":false,"sender":{"key":"markus@bluegap.ch","avatar":"https://gravatar.com/avatar/2f3aadbc46c7c942fa9301d0bb9b91da8c7186233f4c7eff9667a8b53c8cc82e?d=mp&s=160"},"body":"Martin Langhoff wrote:\n> On 9/14/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n>> Let's copy the git list too and maybe we can come up with one importer\n>> for everyone.\n> \n> It's a really good idea. cvsps has been for a while a (limited, buggy)\n> attempt at that. One thing that bothers me in the cvs2svn algorithm is\n> that is not stable in its decisions about where the branching point is\n> -- run the import twice at different times and it may tell you that\n> the branching point has moved.\n\nHuh? Really? Why is that? I don't see reasons for such a thing happening \nwhen studying the algorithm.\n\nFor sure the proposed dependency-resolving algorithm which does not rely \non timestamps does not have that problem.\n\nRegards\n\nMarkus\n"},{"id":"26829","messageId":"450872AE.5050409@bluegap.ch","threadId":"5549","inReplyTo":"46a038f90609131341se42b2dcne73c017cf757d13a@mail.gmail.com","subject":"Re: cvs import","fromName":"Markus Schiltknecht","fromEmail":"markus@bluegap.ch","sentAt":"2006-09-13T21:05:50Z","receivedAt":"2006-09-13T21:05:50Z","isPatch":false,"sender":{"key":"markus@bluegap.ch","avatar":"https://gravatar.com/avatar/2f3aadbc46c7c942fa9301d0bb9b91da8c7186233f4c7eff9667a8b53c8cc82e?d=mp&s=160"},"body":"Martin Langhoff wrote:\n> On 9/14/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n>> Let's copy the git list too and maybe we can come up with one importer\n>> for everyone.\n> \n> It's a really good idea. cvsps has been for a while a (limited, buggy)\n> attempt at that.\n\nBTW: good point, I always thought about cvsps. Does anybody know what \n'dump' format that uses?\n\nFor sure it's algorithm isn't that strong. cvs2svn is better, IMHO. The \nproposed dependency resolving algorithm will be even better /me thinks.\n\nRegards\n\nMarkus\n"},{"id":"26832","messageId":"20060913211521.GA3336@ugly.local","threadId":"5549","inReplyTo":"4508724D.2050701@bluegap.ch","subject":"Re: cvs import","fromName":"Oswald Buddenhagen","fromEmail":"ossi@kde.org","sentAt":"2006-09-13T21:15:21Z","receivedAt":"2006-09-13T21:15:21Z","isPatch":false,"sender":{"key":"ossi@kde.org","avatar":"https://avatars.githubusercontent.com/u/812380?v=4"},"body":"On Wed, Sep 13, 2006 at 11:04:13PM +0200, Markus Schiltknecht wrote:\n> Martin Langhoff wrote:\n> >One thing that bothers me in the cvs2svn algorithm is\n> >that is not stable in its decisions about where the branching point is\n> >-- run the import twice at different times and it may tell you that\n> >the branching point has moved.\n> \n> Huh? Really? Why is that? I don't see reasons for such a thing happening \n> when studying the algorithm.\n> \nthat's certainly due to some hash being iterated. python intentionally\nrandomizes this to make wrong assumptions obvious.\nthere is actually a patch pending to improve the branch source selection\ndrastically. maybe this is affected as well.\n\n> For sure the proposed dependency-resolving algorithm which does not rely \n> on timestamps does not have that problem.\n> \ni think that's unrelated.\n\n-- \nHi! I'm a .signature virus! Copy me into your ~/.signature, please!\n--\nChaos, panic, and disorder - my work here is done.\n"},{"id":"26833","messageId":"46a038f90609131416s1a53b53xd12c3661140fec7a@mail.gmail.com","threadId":"5549","inReplyTo":"4508724D.2050701@bluegap.ch","subject":"Re: cvs import","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-09-13T21:16:01Z","receivedAt":"2006-09-13T21:16:01Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 9/14/06, Markus Schiltknecht <markus@bluegap.ch> wrote:\n> Martin Langhoff wrote:\n> > On 9/14/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> >> Let's copy the git list too and maybe we can come up with one importer\n> >> for everyone.\n> >\n> > It's a really good idea. cvsps has been for a while a (limited, buggy)\n> > attempt at that. One thing that bothers me in the cvs2svn algorithm is\n> > that is not stable in its decisions about where the branching point is\n> > -- run the import twice at different times and it may tell you that\n> > the branching point has moved.\n>\n> Huh? Really? Why is that? I don't see reasons for such a thing happening\n> when studying the algorithm.\n>\n> For sure the proposed dependency-resolving algorithm which does not rely\n> on timestamps does not have that problem.\n\nIIRC, it places branch tags as late as possible. I haven't looked at\nit in detail, but an import immediately after the first commit against\nthe branch may yield a different branchpoint from the same import done\na bit later.\n\ncheers,\n\n\nmartin\n"},{"id":"26837","messageId":"9e4733910609131438n686b6d72u4d5799533c7473d7@mail.gmail.com","threadId":"5549","inReplyTo":"450872AE.5050409@bluegap.ch","subject":"Re: cvs import","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-09-13T21:38:11Z","receivedAt":"2006-09-13T21:38:11Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 9/13/06, Markus Schiltknecht <markus@bluegap.ch> wrote:\n> Martin Langhoff wrote:\n> > On 9/14/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> >> Let's copy the git list too and maybe we can come up with one importer\n> >> for everyone.\n> >\n> > It's a really good idea. cvsps has been for a while a (limited, buggy)\n> > attempt at that.\n>\n> BTW: good point, I always thought about cvsps. Does anybody know what\n> 'dump' format that uses?\n\ncvsps has potential but the multiple missing branch labels in the\nMozilla CVS confuse it and its throws away important data. It's\nalgorithm would need reworking too. cvs2svn is the only CVS converter\nthat imported Mozilla CVS on the first try and mostly got things\nright.\n\nPatchset format for cvsps\nhttp://www.cobite.com/cvsps/README\n\nAFAIK none of the CVS converters are using the dependency algorithm.\nSo the proposal on the table is to develop a new converter that uses\nthe dependency data from CVS to form the change sets and then outputs\nthis data in a form that all of the backends can consume. Of course\neach of the backends is going to have to write some code in order to\nconsume this new import format.\n\n>\n> For sure it's algorithm isn't that strong. cvs2svn is better, IMHO. The\n> proposed dependency resolving algorithm will be even better /me thinks.\n>\n> Regards\n>\n> Markus\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"26846","messageId":"20060913225200.GA10186@frances.vorpus.org","threadId":"5549","inReplyTo":"45084400.1090906@bluegap.ch","subject":"Re: cvs import","fromName":"Nathaniel Smith","fromEmail":"njs@pobox.com","sentAt":"2006-09-13T22:52:00Z","receivedAt":"2006-09-13T22:52:00Z","isPatch":false,"sender":{"key":"njs@pobox.com","avatar":null},"body":"On Wed, Sep 13, 2006 at 07:46:40PM +0200, Markus Schiltknecht wrote:\n> Hi,\n> \n> I've been trying to understand the cvsimport algorithm used by monotone \n> and wanted to adjust that to be more like the one in cvs2svn.\n> \n> I've had some problems with cvs2svn itself and began to question the \n> algorithm used there. It turned out that the cvs2svn people have \n> discussed an improved algorithms and are about to write a cvs2svn 2.0. \n> The main problem with the current algorithm is that it depends on the \n> timestamp information stored in the CVS repository.\n> \n> Instead, it would be much better to just take the dependencies of the \n> revisions into account. Considering the timestamp an irrelevant (for the \n> import) attribute of the revision.\n\nI just read over the thread on the cvs2svn list about this -- I have a\nfew random thoughts.  Take them with a grain of salt, since I haven't\nactually tried writing a CVS importer myself...\n\nRegarding the basic dependency-based algorithm, the approach of\nthrowing everything into blobs and then trying to tease them apart\nagain seems backwards.  What I'm thinking is, first we go through and\nbuild the history graph for each file.  Now, advance a frontier across\nthe all of these graphs simultaneously.  Your frontier is basically a\nmap <filename -> CVS revision>, that represents a tree snapshot.  The\nbasic loop is:\n  1) pick some subset of files to advance to their next revision\n  2) slide the frontier one CVS revision forward on each of those\n     files\n  3) snapshot the new frontier (write it to the target VCS as a new\n     tree commit)\n  4) go to step 1\nObviously, this will produce a target VCS history that respects the\nCVS dependency graph, so that's good; it puts a strict limit on how\nbadly whatever heuristics we use can screw us over if they guess wrong\nabout things.  Also, it makes the problem much simpler -- all the\nheuristics are now in step 1, where we are given a bunch of possible\nedits, and we have to pick some subset of them to accept next.\n\nThis isn't trivial problem.  I think the main thing you want to avoid\nis:\n    1  2  3  4\n    |  |  |  |\n  --o--o--o--o----- <-- current frontier\n    |  |  |  |\n    A  B  A  C\n       |\n       A\nsay you have four files named \"1\", \"2\", \"3\", and \"4\".  We want to\nslide the frontier down, and the next edits were originally created by\none of three commits, A, B, or C.  In this situation, we can take\ncommit B, or we can take commit C, but we don't want to take commit A\nuntil _after_ we have taken commit B -- because otherwise we will end\nup splitting A up into two different commits, A1, B, A2.\n\nThere are a lot of approaches one could take here, on up to pulling\nout a full-on optimal constraint satisfaction system (if we can route\nchips, we should be able to pick a good ordering for accepting CVS\nedits, after all).  A really simple heuristic, though, would be to\njust pick the file whose next commit has the earliest timestamp, then\ngroup in all the other \"next commits\" with the same commit message,\nand (maybe) a similar timestamp.  I have a suspicion that this\nheuristic will work really, really, well in practice.  Also, it's\ncheap to apply, and worst case you accidentally split up a commit that\nalready had wacky timestamps, and we already know that we _have_ to do\nthat in some cases.\n\nHandling file additions could potentially be slightly tricky in this\nmodel.  I guess it is not so bad, if you model added files as being\npresent all along (so you never have to add add whole new entries to\nthe frontier), with each file starting out in a pre-birth state, and\nthen addition of the file is the first edit performed on top of that,\nand you treat these edits like any other edits when considering how to\nadvance the frontier.\n\nI have no particular idea on how to handle tags and branches here;\nI've never actually wrapped my head around CVS's model for those :-).\nI'm not seeing any obvious problem with handling them, though.\n\nIn this approach, incremental conversion is cheap, easy, and robust --\nsimply remember what frontier corresponded to the final revision\nimported, and restart the process directly at that frontier.\n\n\nRegarding storing things on disk vs. in memory: we always used to\nstress-test monotone's cvs importer with the gcc history; just a few\nweeks ago someone did a test import of NetBSD's src repo (~180k\ncommits) on a desktop with 2 gigs of RAM.  It takes a pretty big\nhistory to really require disk (and for that matter, people with\nhistories that big likely have a big enough organization that they can\nget access to some big iron to run the conversion on -- and probably\nwill want to anyway, to make it run in reasonable time).\n\n> Now, that can be used to convert from CVS to about anything else. \n> Obviously we were discussing about subversion, but then there was git, \n> too. And monotone.\n> \n> I'm beginning to question if one could come up with a generally useful \n> cleaned-and-sane-CVS-changeset-dump-format, which could then be used by \n> importers to all sort of VCSes. This would make monotone's cvsimport \n> function dependent on cvs2svn (and therefore python). But the general \n> try-to-get-something-usefull-from-an-insane-CVS-repository-algorithm \n> would only have to be written once.\n> \n> On the other hand, I see that lots of the cvsimport functionality for \n> monotone has already been written (rcs file parsing, stuffing files, \n> file deltas and complete revisions into the monotone database, etc..). \n> Changing it to a better algorithm does not seem to be _that_ much work \n> anymore. Plus the hard part seems to be to come up with a good \n> algorithm, not implementing it. And we could still exchange our \n> experience with the general algorithm with the cvs2svn people.\n>\n> Plus, the guy who mentioned git pointed out that git needs quite a \n> different dump-format than subversion to do an efficient conversion. I \n> think coming up with a generally-usable dump format would not be that easy.\n\nProbably the biggest technical advantage of having the converter built\ninto monotone is that it makes it easy to import the file contents.\nSince this data is huge (100x the repo size, maybe?), and the naive\nalgorithm for reconstructing takes time that is quadratic in the depth\nof history, this is very valuable.  I'm not sure what sort of dump\nformat one could come up with that would avoid making this step very\nexpensive.\n\nI also suspect that SVN's dump format is suboptimal at the metadata\nlevel -- we would essentially have to run a lot of branch/tag\ninferencing logic _again_ to go from SVN-style \"one giant tree with\nbranches described as copies, and multiple copies allowed for\nbranches/tags that are built up over time\", to monotone-style\n\"DAG of tree snapshots\".  This would be substantially less annoying\ninferencing logic than that needed to decipher CVS in the first place,\ngranted, and it's stuff we want to write at some point anyway to allow\nSVN importing, but it adds another step where information could be\nlost.  I may be biased because I grok monotone better, but I suspect\nit would be much easier to losslessly convert a monotone-style history\nto an svn-style history than vice versa, possibly a generic dumping\ntool would want to generate output that looks more like monotone's\nmodel?  The biggest stumbling block I see is if it is important to\nbuild up branches and tags by multiple copies out of trunk -- there\nisn't any way to represent that in monotone.  A generic tool could\nalso use some sort of hybrid model (e.g., dag-of-snapshots plus\nsome extra annotations), if that worked better.\n\nIt's also very nice that users don't need any external software to\nimport CVS->monotone, just because it cuts down on hassle, but I would\nrather have a more hasslesome tool that worked then a less hasslesome\ntool that didn't, and I'm not the one volunteering to write the code,\nso :-).\n\nEven if we _do_ end up writing two implementations of the algorithm,\nwe should share a test suite.  Testing cvs importers is way harder\nthan writing them, because there's no ground truth to compare your\nprogram's output to... in fact, having two separate implementations\nand testing them against each other would be useful to increase\nconfidence in each of them.\n\n(I'm only on one of the CC'ed lists, so reply-to-all appreciated)\n\n-- Nathaniel\n\n-- \n\"On arrival in my ward I was immediately served with lunch. `This is\nwhat you ordered yesterday.' I pointed out that I had just arrived,\nonly to be told: `This is what your bed ordered.'\"\n  -- Letter to the Editor, The Times, September 2000\n"},{"id":"26847","messageId":"20060913232139.GU29625@bcd.geek.com.au","threadId":"5549","inReplyTo":"20060913225200.GA10186@frances.vorpus.org","subject":"Re: cvs import","fromName":"Daniel Carosone","fromEmail":"dan@geek.com.au","sentAt":"2006-09-13T23:21:39Z","receivedAt":"2006-09-13T23:21:39Z","isPatch":false,"sender":{"key":"dan@geek.com.au","avatar":null},"body":"On Wed, Sep 13, 2006 at 03:52:00PM -0700, Nathaniel Smith wrote:\n> This isn't trivial problem.  I think the main thing you want to avoid\n> is:\n>     1  2  3  4\n>     |  |  |  |\n>   --o--o--o--o----- <-- current frontier\n>     |  |  |  |\n>     A  B  A  C\n>        |\n>        A\n> There are a lot of approaches one could take here, on up to pulling\n> out a full-on optimal constraint satisfaction system (if we can route\n> chips, we should be able to pick a good ordering for accepting CVS\n> edits, after all).  A really simple heuristic, though, would be to\n> just pick the file whose next commit has the earliest timestamp, then\n> group in all the other \"next commits\" with the same commit message,\n> and (maybe) a similar timestamp.  \n\nPick the earliest first, or more generally: take all the file commits\nimmediately below the frontier.  Find revs further below the frontier\n(up to some small depth or time limit) on other files that might match\nthem, based on changelog etc (the same grouping you describe, and we\ndo now).  Eliminate any of those that are not entirely on the frontier\n(ie, have some other revision in the way, as with file 2).  Commit the\nremaining set in time order. [*]\n\nIf you wind up with an empty set, then you need to split revs, but at\nthis point you have only conflicting revs on the frontier i.e. you've\nalready committed all the other revs you can that might have avoided\nthis need, whereas we currently might be doing this too often).\n\nFor time order, you could look at each rev as having a time window,\nfrom the first to last commit matching.  If the revs windows are\nnon-overlapping, commit them in order.  If the rev windows overlap, at\nthis point we already know the file changes don't overlap - we *could*\ncommit these as parallel heads and merge them, to better model the\noriginal developer's overlapping commits.\n\n> Handling file additions could potentially be slightly tricky in this\n> model.  I guess it is not so bad, if you model added files as being\n> present all along (so you never have to add add whole new entries to\n> the frontier), with each file starting out in a pre-birth state, and\n> then addition of the file is the first edit performed on top of that,\n> and you treat these edits like any other edits when considering how to\n> advance the frontier.\n\nCVS allows resurrections too..\n\n> I have no particular idea on how to handle tags and branches here;\n> I've never actually wrapped my head around CVS's model for those :-).\n> I'm not seeing any obvious problem with handling them, though.\n\nTags could be modelled as another 'event' in the file graph, like a\ncommit. If your frontier advances through both revisions and a 'tag\nthis revision' event, the same sequencing as above would work. If tags\nhad been moved, this would wind up with a sequence whereby commits\ninterceded with tagging, and we'd need to split the commits such that\nwe could end up with a revision matching the tagged content.\n\n> In this approach, incremental conversion is cheap, easy, and robust --\n> simply remember what frontier corresponded to the final revision\n> imported, and restart the process directly at that frontier.\n\nHm. Except for the tagging idea above, because tags can be applied\nbehind a live cvs frontier.\n\n--\nDan.\n\n\n_______________________________________________\nMonotone-devel mailing list\nMonotone-devel@nongnu.org\nhttp://lists.nongnu.org/mailman/listinfo/monotone-devel\n"},{"id":"26848","messageId":"1158190921.29313.175.camel@neko.keithp.com","threadId":"5549","inReplyTo":"20060913225200.GA10186@frances.vorpus.org","subject":"Re: [Monotone-devel] cvs import","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-09-13T23:42:01Z","receivedAt":"2006-09-13T23:42:01Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Wed, 2006-09-13 at 15:52 -0700, Nathaniel Smith wrote:\n\n> Regarding the basic dependency-based algorithm, the approach of\n> throwing everything into blobs and then trying to tease them apart\n> again seems backwards.  What I'm thinking is, first we go through and\n> build the history graph for each file.  Now, advance a frontier across\n> the all of these graphs simultaneously.  Your frontier is basically a\n> map <filename -> CVS revision>, that represents a tree snapshot. \n\nParsecvs does this, except backwards from now into the past; I found it\neasier to identify merge points than branch points (Oh, look, these two\nbranches are the same now, they must have merged).\n\nHowever, this means that parsecvs must hold the entire tree state in\nmemory, which turned out to be its downfall with large repositories.\nWorked great for all of X.org, not so good with Mozilla.\n\n-- \nkeith.packard@intel.com\n"},{"id":"26849","messageId":"20060913235227.GW29625@bcd.geek.com.au","threadId":"5549","inReplyTo":"20060913232139.GU29625@bcd.geek.com.au","subject":"Re: [Monotone-devel] cvs import","fromName":"Daniel Carosone","fromEmail":"dan@geek.com.au","sentAt":"2006-09-13T23:52:27Z","receivedAt":"2006-09-13T23:52:27Z","isPatch":false,"sender":{"key":"dan@geek.com.au","avatar":null},"body":"On Thu, Sep 14, 2006 at 09:21:39AM +1000, Daniel Carosone wrote:\n> > I have no particular idea on how to handle tags and branches here;\n> > I've never actually wrapped my head around CVS's model for those :-).\n> > I'm not seeing any obvious problem with handling them, though.\n> \n> Tags could be modelled as another 'event' in the file graph, like a\n> commit. If your frontier advances through both revisions and a 'tag\n> this revision' event, the same sequencing as above would work.\n\nLikewise, if we had \"file branched\" events in the file lifeline (based\non the rcs id's), then we would be sure to always have a monotone\nrevision that corresponded to the branching event, where we could\nattach the revisions in the branch.\n\nBecause we can't split tags, and can't split branch events, we will\nend up splitting file commits (down to individual commits per file) in\norder to arrive at the revisions we need for those.\n\nBecause tags and branches can be across subsets of the tree, we gain\nsome scheduling flexibility about where in the reconstructed sequence\nthey can come.\n\nMany well-managed CVS repositories will use good practices, such as\nhaving a branch base tag.  If they do, then they will help this\nalgorithm produce correct results.\n\nOnce we have a branch with a base starting revision, we can pretty\nmuch treat it independently from there: make a whole new set of file\nlifelines along the RCS branches and a new frontier for it.\n\n--\nDan.\n"},{"id":"26850","messageId":"20060914003242.GA19228@frances.vorpus.org","threadId":"5549","inReplyTo":"1158190921.29313.175.camel@neko.keithp.com","subject":"Re: cvs import","fromName":"Nathaniel Smith","fromEmail":"njs@pobox.com","sentAt":"2006-09-14T00:32:42Z","receivedAt":"2006-09-14T00:32:42Z","isPatch":false,"sender":{"key":"njs@pobox.com","avatar":null},"body":"On Wed, Sep 13, 2006 at 04:42:01PM -0700, Keith Packard wrote:\n> However, this means that parsecvs must hold the entire tree state in\n> memory, which turned out to be its downfall with large repositories.\n> Worked great for all of X.org, not so good with Mozilla.\n\nDoes anyone know how big Mozilla (or other humonguous repos, like KDE)\nare, in terms of number of files?\n\nA few numbers for repositories I had lying around:\n  Linux kernel -- ~21,000\n  gcc -- ~42,000\n  NetBSD \"src\" repo -- ~100,000\n  uClinux distro -- ~110,000\n\nThese don't seem very indimidating... even if it takes an entire\nkilobyte per CVS revision to store the information about it that we\nneed to make decisions about how to move the frontier... that's only\n110 megabytes for the largest of these repos.  The frontier sweeping\nalgorithm only _needs_ to have available the current frontier, and the\ncurrent frontier+1.  Storing information on every version of every\nfile in memory might be worse; but since the algorithm accesses this\ndata in a linear way, it'd be easy enough to stick those in a\nlookaside table on disk if really necessary, like a bdb or sqlite file\nor something.\n\n(Again, in practice storing all the metadata for the entire 180k\nrevisions of the 100k files in the netbsd repo was possible on a\ndesktop.  Monotone's cvs_import does try somewhat to be frugal about\nmemory, though, interning strings and suchlike.)\n\n-- Nathaniel\n\n-- \nWhen the flush of a new-born sun fell first on Eden's green and gold,\nOur father Adam sat under the Tree and scratched with a stick in the mould;\nAnd the first rude sketch that the world had seen was joy to his mighty heart,\nTill the Devil whispered behind the leaves, \"It's pretty, but is it Art?\"\n  -- The Conundrum of the Workshops, Rudyard Kipling\n"},{"id":"26851","messageId":"9e4733910609131757l7ce4b637oae18b523b1b7f0a4@mail.gmail.com","threadId":"5549","inReplyTo":"20060914003242.GA19228@frances.vorpus.org","subject":"Re: [Monotone-devel] cvs import","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-09-14T00:57:33Z","receivedAt":"2006-09-14T00:57:33Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 9/13/06, Nathaniel Smith <njs@pobox.com> wrote:\n> On Wed, Sep 13, 2006 at 04:42:01PM -0700, Keith Packard wrote:\n> > However, this means that parsecvs must hold the entire tree state in\n> > memory, which turned out to be its downfall with large repositories.\n> > Worked great for all of X.org, not so good with Mozilla.\n>\n> Does anyone know how big Mozilla (or other humonguous repos, like KDE)\n> are, in terms of number of files?\n\nMozilla is 120,000 files. The complexity comes from 10 years worth of\nhistory. A few of the files have around 1,700 revisions. There are\nabout 1,600 branches and 1,000 tags. The branch number is inflated\nbecause cvs2svn is generating extra branches, the real number is\naround 700. The CVS repo takes 4.2GB disk space. cvs2svn turns this\ninto 250,000 commits over about 1M unique revisions.\n\n>\n> A few numbers for repositories I had lying around:\n>   Linux kernel -- ~21,000\n>   gcc -- ~42,000\n>   NetBSD \"src\" repo -- ~100,000\n>   uClinux distro -- ~110,000\n>\n> These don't seem very indimidating... even if it takes an entire\n> kilobyte per CVS revision to store the information about it that we\n> need to make decisions about how to move the frontier... that's only\n> 110 megabytes for the largest of these repos.  The frontier sweeping\n> algorithm only _needs_ to have available the current frontier, and the\n> current frontier+1.  Storing information on every version of every\n> file in memory might be worse; but since the algorithm accesses this\n> data in a linear way, it'd be easy enough to stick those in a\n> lookaside table on disk if really necessary, like a bdb or sqlite file\n> or something.\n>\n> (Again, in practice storing all the metadata for the entire 180k\n> revisions of the 100k files in the netbsd repo was possible on a\n> desktop.  Monotone's cvs_import does try somewhat to be frugal about\n> memory, though, interning strings and suchlike.)\n>\n> -- Nathaniel\n>\n> --\n> When the flush of a new-born sun fell first on Eden's green and gold,\n> Our father Adam sat under the Tree and scratched with a stick in the mould;\n> And the first rude sketch that the world had seen was joy to his mighty heart,\n> Till the Devil whispered behind the leaves, \"It's pretty, but is it Art?\"\n>   -- The Conundrum of the Workshops, Rudyard Kipling\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"26857","messageId":"20060914015324.GX29625@bcd.geek.com.au","threadId":"5549","inReplyTo":"9e4733910609131757l7ce4b637oae18b523b1b7f0a4@mail.gmail.com","subject":"Re: cvs import","fromName":"Daniel Carosone","fromEmail":"dan@geek.com.au","sentAt":"2006-09-14T01:53:24Z","receivedAt":"2006-09-14T01:53:24Z","isPatch":false,"sender":{"key":"dan@geek.com.au","avatar":null},"body":"On Wed, Sep 13, 2006 at 08:57:33PM -0400, Jon Smirl wrote:\n> Mozilla is 120,000 files. The complexity comes from 10 years worth of\n> history. A few of the files have around 1,700 revisions. There are\n> about 1,600 branches and 1,000 tags. The branch number is inflated\n> because cvs2svn is generating extra branches, the real number is\n> around 700. The CVS repo takes 4.2GB disk space. cvs2svn turns this\n> into 250,000 commits over about 1M unique revisions.\n\nThose numbers are pretty close to those in the NetBSD repository, and\nbetween them these probably represent just about the most extensive\npublic CVS test data available. \n\nI've only done imports of individual top-level dirs (what used to be\nmodules), like src and pkgsrc, because they're used independently and\ndon't really overlap.\n\nsrc had about 180k commits over 1M versions of 120k files, 1000 tags\nand 260 branches. pkgsrc had 110k commits over about half as many\nfiles and versions thereof.  We too have a few hot files, one had\n13,625 revisions.  xsrc adds a bunch more files and content, but not\nmany versions; that's mostly vendor branches and only some local\nchanges.  Between them the cvs ,v files take up 4.7G covering about 13\nyears of history.\n\nOne thing that was interesting was that \"src\" used to be several\ndifferent modules, but we rearranged the repository at one point to\nmatch the checkout structure these modules produced (combining them\nall under the src dir).  This doesn't seem to have upset the import at\nall.  Just about every other form of CVS evil has been perpetrated in\nthis repository at some stage or other too, but always very carefully.\n\n--\nDan.\n\n\n_______________________________________________\nMonotone-devel mailing list\nMonotone-devel@nongnu.org\nhttp://lists.nongnu.org/mailman/listinfo/monotone-devel\n"},{"id":"26859","messageId":"20060914023017.GA31889@spearce.org","threadId":"5549","inReplyTo":"20060914015324.GX29625@bcd.geek.com.au","subject":"Re: [Monotone-devel] cvs import","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-09-14T02:30:17Z","receivedAt":"2006-09-14T02:30:17Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Daniel Carosone <dan@geek.com.au> wrote:\n> On Wed, Sep 13, 2006 at 08:57:33PM -0400, Jon Smirl wrote:\n> > Mozilla is 120,000 files. The complexity comes from 10 years worth of\n> > history. A few of the files have around 1,700 revisions. There are\n> > about 1,600 branches and 1,000 tags. The branch number is inflated\n> > because cvs2svn is generating extra branches, the real number is\n> > around 700. The CVS repo takes 4.2GB disk space. cvs2svn turns this\n> > into 250,000 commits over about 1M unique revisions.\n> \n> Those numbers are pretty close to those in the NetBSD repository, and\n> between them these probably represent just about the most extensive\n> public CVS test data available. \n\nI don't know exactly how big it is but the Gentoo CVS repository\nis also considered to be very large (about the size of the Mozilla\nrepository) and just as difficult to import.  Its either crashed or\ntaken about a month to process with the current Git CVS->Git tools.\n\nSince I know that the bulk of the Gentoo CVS repository is the\nportage tree I did a quick find|wc -l in my /usr/portage; its about\n124,500 files.\n\nIts interesting that Gentoo has almost as large of a repository given\nthat its such a young project, compared to NetBSD and Mozilla.  :-)\n\n-- \nShawn.\n"},{"id":"26860","messageId":"20060914023541.GB31889@spearce.org","threadId":"5549","inReplyTo":"1158190921.29313.175.camel@neko.keithp.com","subject":"Re: [Monotone-devel] cvs import","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-09-14T02:35:42Z","receivedAt":"2006-09-14T02:35:42Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Keith Packard <keithp@keithp.com> wrote:\n> On Wed, 2006-09-13 at 15:52 -0700, Nathaniel Smith wrote:\n> \n> > Regarding the basic dependency-based algorithm, the approach of\n> > throwing everything into blobs and then trying to tease them apart\n> > again seems backwards.  What I'm thinking is, first we go through and\n> > build the history graph for each file.  Now, advance a frontier across\n> > the all of these graphs simultaneously.  Your frontier is basically a\n> > map <filename -> CVS revision>, that represents a tree snapshot. \n> \n> Parsecvs does this, except backwards from now into the past; I found it\n> easier to identify merge points than branch points (Oh, look, these two\n> branches are the same now, they must have merged).\n\nWhy not let Git do that?  If two branches are the same in CVS then\nshouldn't they have the same tree SHA1 in Git?  Surely comparing\n20 bytes of SHA1 is faster than almost any other comparsion...\n \n> However, this means that parsecvs must hold the entire tree state in\n> memory, which turned out to be its downfall with large repositories.\n> Worked great for all of X.org, not so good with Mozilla.\n\nAny chance that can be paged in on demand from some sort of work\nfile?  git-fast-import hangs onto a configurable number of tree\nstates (default of 5) but keeps them in an LRU chain and dumps the\nones that aren't current.\n\n-- \nShawn.\n"},{"id":"26861","messageId":"20060914031918.GY29625@bcd.geek.com.au","threadId":"5549","inReplyTo":"20060914023017.GA31889@spearce.org","subject":"Re: cvs import","fromName":"Daniel Carosone","fromEmail":"dan@geek.com.au","sentAt":"2006-09-14T03:19:18Z","receivedAt":"2006-09-14T03:19:18Z","isPatch":false,"sender":{"key":"dan@geek.com.au","avatar":null},"body":"On Wed, Sep 13, 2006 at 10:30:17PM -0400, Shawn Pearce wrote:\n> I don't know exactly how big it is but the Gentoo CVS repository\n> is also considered to be very large (about the size of the Mozilla\n> repository) and just as difficult to import.  Its either crashed or\n> taken about a month to process with the current Git CVS->Git tools.\n\nAh, thanks for the tip.\n\n> Since I know that the bulk of the Gentoo CVS repository is the\n> portage tree I did a quick find|wc -l in my /usr/portage; its about\n> 124,500 files.\n> \n> Its interesting that Gentoo has almost as large of a repository given\n> that its such a young project, compared to NetBSD and Mozilla.  :-)\n\nPortage uses files and thus CVS very differently, though.  Each ebuild\nfor each package revision of each version of a third-party package\n(like, say, monotone 0.28 and 0.29, and -r1, -r2 pkg bumps of those if\nthey were needed) is its own file that's added, maybe edited a couple\nof times, and then deleted again later as new versions are added and\nolder ones retired.  These are copies and renames in the workspace,\nbut are invisible to CVS.  This uses up lots more files than a single\nlong-lived build that gets edited each time; the Attic dirs must have\nhuge numbers of files, way beyond the number that are live now.\n\nThis lets portage keep builds around in a HEAD checkout for multiple\nversions at once, tagged internally with different statuses.\nEffectively, these tags take the place of VCS-based branches and\nreleases, and are more flexible for end users tracking their favourite\napplications while keeping the rest of their system stable.\n\nIf they had a VCS that supported file cloning and/or renaming, and\nused that to follow history between these ebuild files, things would\nbe very different. There are some interesting use cases for VCS tools\nin supporting this behaviour nicely, too.  \n\n--\nDan.\n\n_______________________________________________\nMonotone-devel mailing list\nMonotone-devel@nongnu.org\nhttp://lists.nongnu.org/mailman/listinfo/monotone-devel\n"},{"id":"26863","messageId":"4508D7DA.8000302@alum.mit.edu","threadId":"5549","inReplyTo":"46a038f90609131416s1a53b53xd12c3661140fec7a@mail.gmail.com","subject":"Re: cvs import","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2006-09-14T04:17:30Z","receivedAt":"2006-09-14T04:17:30Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Martin Langhoff wrote:\n> On 9/14/06, Markus Schiltknecht <markus@bluegap.ch> wrote:\n>> Martin Langhoff wrote:\n>> > On 9/14/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n>> >> Let's copy the git list too and maybe we can come up with one importer\n>> >> for everyone.\n>> >\n>> > It's a really good idea. cvsps has been for a while a (limited, buggy)\n>> > attempt at that. One thing that bothers me in the cvs2svn algorithm is\n>> > that is not stable in its decisions about where the branching point is\n>> > -- run the import twice at different times and it may tell you that\n>> > the branching point has moved.\n>>\n>> Huh? Really? Why is that? I don't see reasons for such a thing happening\n>> when studying the algorithm.\n>>\n>> For sure the proposed dependency-resolving algorithm which does not rely\n>> on timestamps does not have that problem.\n> \n> IIRC, it places branch tags as late as possible. I haven't looked at\n> it in detail, but an import immediately after the first commit against\n> the branch may yield a different branchpoint from the same import done\n> a bit later.\n\nThis is correct.  And IMO it makes sense from the standpoint of an\nall-at-once conversion.\n\nBut I was under the impression that this wouldn't matter for\ncontent-indexed-based SCMs.  The content of all possible branching\npoints is identical, and therefore from your point of view the topology\nshould be the same, no?\n\nBut aside from this point, I think an intrinsic part of implementing\nincremental conversion is \"convert the subsequent changes to the CVS\nrepository *subject to the constraints* imposed by decisions made in\nearlier conversion runs.  And the real trick is that things can be done\nin CVS (e.g., line-end changes, manual copying of files in the repo)\nthat (a) are unversioned and (b) have retroactive effects that go\narbitrarily far back in time.  This is the reason that I am pessimistic\nthat incremental conversion will ever work robustly.\n\nMichael\n"},{"id":"26864","messageId":"9e4733910609132134j63857912keed6a42682f69d66@mail.gmail.com","threadId":"5549","inReplyTo":"4508D7DA.8000302@alum.mit.edu","subject":"Re: cvs import","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-09-14T04:34:52Z","receivedAt":"2006-09-14T04:34:52Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 9/14/06, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> But aside from this point, I think an intrinsic part of implementing\n> incremental conversion is \"convert the subsequent changes to the CVS\n> repository *subject to the constraints* imposed by decisions made in\n> earlier conversion runs.  And the real trick is that things can be done\n> in CVS (e.g., line-end changes, manual copying of files in the repo)\n> that (a) are unversioned and (b) have retroactive effects that go\n> arbitrarily far back in time.  This is the reason that I am pessimistic\n> that incremental conversion will ever work robustly.\n\nWe don't need really robust incremental conversion. It just needs to\nwork most of the time. Incremental conversion is usually used to track\nthe main CVS repo with the new tool while people decide if they like\nthe new tool. Commits will still flow to the CVS repo and get\nincrementally copied to the new tool so that it tracks CVS in close to\nreal time.\n\nIf the increment import messes up you can always redo a full import,\nbut a full Mozilla import takes about 2 hours with the git tools. I\nwould always do a full import on the day of the actual cut over.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"26865","messageId":"46a038f90609132140v10118b53q8f8001bcf575263d@mail.gmail.com","threadId":"5549","inReplyTo":"4508D7DA.8000302@alum.mit.edu","subject":"Re: cvs import","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-09-14T04:40:47Z","receivedAt":"2006-09-14T04:40:47Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 9/14/06, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> > IIRC, it places branch tags as late as possible. I haven't looked at\n> > it in detail, but an import immediately after the first commit against\n> > the branch may yield a different branchpoint from the same import done\n> > a bit later.\n>\n> This is correct.  And IMO it makes sense from the standpoint of an\n> all-at-once conversion.\n>\n> But I was under the impression that this wouldn't matter for\n> content-indexed-based SCMs.  The content of all possible branching\n> points is identical, and therefore from your point of view the topology\n> should be the same, no?\n\nExactly. But if you shift the branching point to later, two things change\n\n - it is possible that (in some corner cases) the content itself\nchanges as the branching point could end up being moved a couple of\ncommits \"later\". one of the downsides of cvs not being atomic.\n\n - even if the content does not change, rearranging of history in git\nis a no-no. git relies on history being read-only 100%\n\n> But aside from this point, I think an intrinsic part of implementing\n> incremental conversion is \"convert the subsequent changes to the CVS\n> repository *subject to the constraints* imposed by decisions made in\n> earlier conversion runs.\n\nYes, and that's a fundamental change in the algorithm. That's exactly\nwhy I mentioned it in this thread ;-) Any incremental importer has to\nmake up some parts of history, and then remember what it has made up.\n\nSo part of the process becomes\n - figure our history on top of the history we already parsed\n - check whether the cvs repo now has any 'new' history that affects\nalready-parsed history negatively, and report those as errors\n\nhmmmmmm.\n\n> This is the reason that I am pessimistic\n> that incremental conversion will ever work robustly.\n\nWe all are :) But for a repo that doesn't go through direct tampering,\nwe can improve the algorithm to be more stable.\n\n\n\nmartin\n"},{"id":"26866","messageId":"4508E26B.5000106@alum.mit.edu","threadId":"5549","inReplyTo":"9e4733910609132134j63857912keed6a42682f69d66@mail.gmail.com","subject":"Re: cvs import","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2006-09-14T05:02:35Z","receivedAt":"2006-09-14T05:02:35Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Jon Smirl wrote:\n> On 9/14/06, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> But aside from this point, I think an intrinsic part of implementing\n>> incremental conversion is \"convert the subsequent changes to the CVS\n>> repository *subject to the constraints* imposed by decisions made in\n>> earlier conversion runs.  And the real trick is that things can be done\n>> in CVS (e.g., line-end changes, manual copying of files in the repo)\n>> that (a) are unversioned and (b) have retroactive effects that go\n>> arbitrarily far back in time.  This is the reason that I am pessimistic\n>> that incremental conversion will ever work robustly.\n> \n> We don't need really robust incremental conversion. It just needs to\n> work most of the time. Incremental conversion is usually used to track\n> the main CVS repo with the new tool while people decide if they like\n> the new tool. Commits will still flow to the CVS repo and get\n> incrementally copied to the new tool so that it tracks CVS in close to\n> real time.\n\nI hadn't thought of the idea of using incremental conversion as an\nadvertising method for switching SCM systems :-)  But if changes flow\nback to CVS, doesn't this have to be pretty robust?\n\nIn our trial period, we simply did a single conversion to SVN and let\npeople play with this test repository.  When we decided to switch over\nwe did another full conversion and simply discarded the changes that had\nbeen made in the test SVN repository.\n\nThe use cases that I had considered were:\n\n1. For conversions that take days, one could do a full commit while\nleaving CVS online, then take CVS offline and do only an incremental\nconversion to reduce SCM downtime.  This is of course less of an issue\nif you could bring the conversion time down to a couple hours for even\nthe largest CVS repos.\n\n2. Long-term continuous mirroring (backwards and forwards) between CVS\nand another SCM, to allow people to use their preferred tool.  (I\nactually think that this is a silly idea, but some people seem to like it.)\n\nFor both of these applications, incremental conversion would have to be\nrobust (for 1 it would at least have to give a clear indication of\nunrecoverable errors).\n\n\nMichael\n"},{"id":"26867","messageId":"46a038f90609132221o125c4694r75dbc8f728104832@mail.gmail.com","threadId":"5549","inReplyTo":"4508E26B.5000106@alum.mit.edu","subject":"Re: cvs import","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-09-14T05:21:16Z","receivedAt":"2006-09-14T05:21:16Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 9/14/06, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> 2. Long-term continuous mirroring (backwards and forwards) between CVS\n> and another SCM, to allow people to use their preferred tool.  (I\n> actually think that this is a silly idea, but some people seem to like it.)\n\nCall me silly ;-) I use this all the time to track projects that use\nCVS or SVN, where I either\n\n - Do have write access, but often develop offline (and I have a bunch\nof perl/shell scripts to extract the patches and auto-commit them into\nCVS/SVN).\n\n - Do have write access, but want to experimental work branches\nwithout making much noise in the cvs repo -- and being able to merge\nCVS's HEAD in repeatedly as you'd want.\n\n - Run \"vendor-branch-tracking\" setups for projects where I have a\ncustom branch of a FOSS sofware project, and repeatedly import updates\nfrom upstream. this is the 'killer-app' of DSCMs IMHO.\n\nIt is not as robust as I'd like; with CVS, the git imports eventually\nstray a bit from upstream, and requires manual fixing. But it is\n_good_.\n\ncheers,\n\n\nmartin\n"},{"id":"26868","messageId":"9e4733910609132230p272c205eq45d9da24267396a8@mail.gmail.com","threadId":"5549","inReplyTo":"4508E26B.5000106@alum.mit.edu","subject":"Re: cvs import","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-09-14T05:30:19Z","receivedAt":"2006-09-14T05:30:19Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 9/14/06, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> Jon Smirl wrote:\n> > On 9/14/06, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> >> But aside from this point, I think an intrinsic part of implementing\n> >> incremental conversion is \"convert the subsequent changes to the CVS\n> >> repository *subject to the constraints* imposed by decisions made in\n> >> earlier conversion runs.  And the real trick is that things can be done\n> >> in CVS (e.g., line-end changes, manual copying of files in the repo)\n> >> that (a) are unversioned and (b) have retroactive effects that go\n> >> arbitrarily far back in time.  This is the reason that I am pessimistic\n> >> that incremental conversion will ever work robustly.\n> >\n> > We don't need really robust incremental conversion. It just needs to\n> > work most of the time. Incremental conversion is usually used to track\n> > the main CVS repo with the new tool while people decide if they like\n> > the new tool. Commits will still flow to the CVS repo and get\n> > incrementally copied to the new tool so that it tracks CVS in close to\n> > real time.\n>\n> I hadn't thought of the idea of using incremental conversion as an\n> advertising method for switching SCM systems :-)  But if changes flow\n> back to CVS, doesn't this have to be pretty robust?\n\nChanges flow back to CVS but using the new tool to generate a patch,\napply the patch to your CVS check out and commit it.\n\nThere are too many people working on Mozilla to get agreement to\nswitch in a short amount of time. git may need to mirror CVS for\nseveral months. There are also other people pushing svn, monotone,\nperforce, etc, etc, etc. Bottom line, Mozilla really needs a\ndistributed system because external companies are making large changes\nand want their repos in house.\n\nIn my experience none of the other SCMs are up to taking one Mozilla\nyet. Git has the tools but I can get a clean import.\n\n\nI am using this process on Mozilla right now with git. I have a script\nthat updates my CVS tree overnight and then commits the changes into a\nlocal git repo. I can then work on Mozilla using git but my history is\nall messed up. When a change is ready I generate a diff against last\nnight's check out and apply it to my CVS tree and commit. CVS then\nfinds any merge problems for me.\n\n>\n> In our trial period, we simply did a single conversion to SVN and let\n> people play with this test repository.  When we decided to switch over\n> we did another full conversion and simply discarded the changes that had\n> been made in the test SVN repository.\n>\n> The use cases that I had considered were:\n>\n> 1. For conversions that take days, one could do a full commit while\n> leaving CVS online, then take CVS offline and do only an incremental\n> conversion to reduce SCM downtime.  This is of course less of an issue\n> if you could bring the conversion time down to a couple hours for even\n> the largest CVS repos.\n>\n> 2. Long-term continuous mirroring (backwards and forwards) between CVS\n> and another SCM, to allow people to use their preferred tool.  (I\n> actually think that this is a silly idea, but some people seem to like it.)\n>\n> For both of these applications, incremental conversion would have to be\n> robust (for 1 it would at least have to give a clear indication of\n> unrecoverable errors).\n>\n>\n> Michael\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"26869","messageId":"4508EA0E.80704@alum.mit.edu","threadId":"5549","inReplyTo":"46a038f90609132221o125c4694r75dbc8f728104832@mail.gmail.com","subject":"Re: cvs import","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2006-09-14T05:35:10Z","receivedAt":"2006-09-14T05:35:10Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Martin Langhoff wrote:\n> On 9/14/06, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> 2. Long-term continuous mirroring (backwards and forwards) between CVS\n>> and another SCM, to allow people to use their preferred tool.  (I\n>> actually think that this is a silly idea, but some people seem to like\n>> it.)\n> \n> Call me silly ;-) I use this all the time to track projects that use\n> CVS or SVN, where I either\n>\n> [...]\n\nSorry, I guess I was speaking as a person who prefers and is most\nfamiliar with centralized SCM.  But I see from your response that the\nultimate in decentralized development is that each developer decides\nwhat SCM to use :-) and that incremental conversion makes sense in that\ncontext.\n\nMichael\n"},{"id":"26870","messageId":"4508EA78.5030001@alum.mit.edu","threadId":"5549","inReplyTo":"9e4733910609131438n686b6d72u4d5799533c7473d7@mail.gmail.com","subject":"Re: cvs import","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2006-09-14T05:36:56Z","receivedAt":"2006-09-14T05:36:56Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Jon Smirl wrote:\n> On 9/13/06, Markus Schiltknecht <markus@bluegap.ch> wrote:\n>> Martin Langhoff wrote:\n>> > On 9/14/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n>> >> Let's copy the git list too and maybe we can come up with one importer\n>> >> for everyone.\n\nThat would be great.\n\n> AFAIK none of the CVS converters are using the dependency algorithm.\n> So the proposal on the table is to develop a new converter that uses\n> the dependency data from CVS to form the change sets and then outputs\n> this data in a form that all of the backends can consume. Of course\n> each of the backends is going to have to write some code in order to\n> consume this new import format.\n\nFrankly, I think people are getting the priorities wrong by focusing on\nthe format of the output of cvs2svn.  Hacking a new output format onto\ncvs2svn is a trivial matter of a couple hours of programming.\n\nThe real strength of cvs2svn (and I can say this without bragging\nbecause most of this was done before I got involved in the project) is\nthat it handles dozens of peculiar corner cases and bizarre CVS\nperversions, including a good test suite containing lots of twisted\nlittle example repositories.  This is 90% of the intellectual content of\ncvs2svn.\n\nI've spent many, many hours refactoring and reengineering cvs2svn to\nmake it easy to modify and add new features.  The main thing that I want\nto change is to use the dependency graph (rather than timestamps tweaked\nto reflect dependency ordering) to deduce changesets.  But I would never\nthink of throwing away the \"old\" cvs2svn and starting anew, because then\nI would have to add all the little corner cases again from scratch.\n\nIt would be nice to have a universal dumpfile format, but IMO not\ncritical.  The only difference between our SCMs that might be difficult\nto paper over in a universal dumpfile is that SVN wants its changesets\nin chronological order, whereas I gather that others would prefer the\ndata in dependency order branch by branch.\n\nI say let cvs2svn (or if you like, we can rename it to \"cvs2noncvs\" :-)\n) reconstruct the repository's change sets, then let us build several\nbackends that output the data in the format that is most convenient for\neach project.\n\nMichael\n"},{"id":"26888","messageId":"20060914155003.GB9657@spearce.org","threadId":"5549","inReplyTo":"4508EA78.5030001@alum.mit.edu","subject":"Re: cvs import","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-09-14T15:50:03Z","receivedAt":"2006-09-14T15:50:03Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>  The only difference between our SCMs that might be difficult\n> to paper over in a universal dumpfile is that SVN wants its changesets\n> in chronological order, whereas I gather that others would prefer the\n> data in dependency order branch by branch.\n\nThis really isn't an issue for Git.\n\nOriginally I wanted Jon Smirl to modify the cvs2svn code to emit\nonly one branch at a time as that would be much faster than jumping\naround branches in chronological order.  But it turned out to\nbe too much work to change cvs2svn.  So git-fast-import (the Git\nprogram that consumes the dump stream from Jon's modified cvs2svn)\nmaintains an LRU of the branches in memory and reloads inactive\nbranches as necessary when cvs2svn jumps around.\n\nIt turns out it didn't matter if the git-fast-import maintained 5\nactive branches in the LRU or 60.  Apparently the Mozilla repo didn't\njump around more than 5 branches at a time - most of the time anyway.\n\nBranches in git-fast-import seemed to cost us only 2 MB of memory\nper active branch on the Mozilla repository.  Holding 60 of them at\nonce (120 MB) is peanuts on most machines today.  But really only 5\n(10 MB) were needed for an efficient import.\n\n\nI don't know how the Monotone guys feel about it but I think Git\nis happy with the data in any order, just so long as the dependency\nchains aren't fed out of order.  Which I think nearly all changeset\nbased SCMs would have an issue with.  So we should be just fine\nwith the current chronological order produced by cvs2svn.\n\n-- \nShawn.\n"},{"id":"26891","messageId":"eebuih$u32$1@sea.gmane.org","threadId":"5549","inReplyTo":"20060914155003.GB9657@spearce.org","subject":"Re: cvs import","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-09-14T16:04:57Z","receivedAt":"2006-09-14T16:04:57Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Shawn Pearce wrote:\n\n> Originally I wanted Jon Smirl to modify the cvs2svn (...)\n\nBy the way, will cvs2git (modified cvs2svn) and git-fast-import publicly\navailable?\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"26892","messageId":"20060914161809.GA9885@spearce.org","threadId":"5549","inReplyTo":"eebuih$u32$1@sea.gmane.org","subject":"Re: cvs import","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-09-14T16:18:09Z","receivedAt":"2006-09-14T16:18:09Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Jakub Narebski <jnareb@gmail.com> wrote:\n> Shawn Pearce wrote:\n> \n> > Originally I wanted Jon Smirl to modify the cvs2svn (...)\n> \n> By the way, will cvs2git (modified cvs2svn) and git-fast-import publicly\n> available?\n\nYes.  I want to submit git-fast-import to the main Git project and\nask Junio to bring it in.\n\nHowever right now I feel like the code isn't up-to-snuff and won't\npass peer review on the Git mailing list.  So I wanted to spend a\nlittle bit of time cleaning it up before asking Junio to carry it\nin the main distribution.  My pack mmap window code is actually\npart of that cleanup.\n\nI think the goal of this thread is to try and merge the ideas\nbehind Jon's modified cvs2svn into the core cvs2svn, possibly\ncausing cvs2svn to be renamed to cvs2notcvs (or some such) and\nhaving a slightly more modular output format so Git, Monotone and\nSubversion can all benefit from the difficult-to-do-right changeset\ngeneration logic.\n\n-- \nShawn.\n"},{"id":"26893","messageId":"9e4733910609140927y30ecaa42wae0ff0597b8c3842@mail.gmail.com","threadId":"5549","inReplyTo":"eebuih$u32$1@sea.gmane.org","subject":"Re: cvs import","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-09-14T16:27:40Z","receivedAt":"2006-09-14T16:27:40Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 9/14/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> Shawn Pearce wrote:\n>\n> > Originally I wanted Jon Smirl to modify the cvs2svn (...)\n>\n> By the way, will cvs2git (modified cvs2svn) and git-fast-import publicly\n> available?\n\nIt has some unresolved problems so I wasn't spreading it around everywhere.\n\nIt is based on cvs2svn from August. There has been too much change to\nthe current cvs2svn to merge it anymore. It is going to need\nsignificant rewrite. But cvs2svn will all change again if it converts\nto the dependency model. It is better to get a backend independent\ninterface build into cvs2svn.\n\nIt it not generating an accurate repo. cvs2svn is outputting tags\nbased on multiple revisions, git can't do that. I'm just tossing some\nof the tag data that git can't handle. I base the tag on the fist\nrevision which is not correct.\n\nIf the repo is missing branch tags cvs2svn may turn a single missing\nbranch into hundreds of branches. The Mozilla repo has about 1000\nextra branches because of this.\n\nSometime cvs2svn will partial copy from another rev to generate a new\nrev. Git doesn't do this so I am tossing the copy requests. I need to\nfigure out how to hook into the data before cvs2svn tries to copy\nthings.\n\ncvs2svn makes no attempt to detect merges so gitk will show 1,700\nactive branches when there are really only 10 currently active\nbranches in Mozilla.\n\nThat said 99.9% of Mozilla CVS is in the output git repo, but it isn't\nquite right.\n\nIf you still want the code I'll send it to you.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"26897","messageId":"45098AE0.6030409@alum.mit.edu","threadId":"5549","inReplyTo":"9e4733910609140927y30ecaa42wae0ff0597b8c3842@mail.gmail.com","subject":"Re: cvs import","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2006-09-14T17:01:20Z","receivedAt":"2006-09-14T17:01:20Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Jon Smirl wrote:\n> On 9/14/06, Jakub Narebski <jnareb@gmail.com> wrote:\n>> Shawn Pearce wrote:\n>>\n>> > Originally I wanted Jon Smirl to modify the cvs2svn (...)\n>>\n>> By the way, will cvs2git (modified cvs2svn) and git-fast-import publicly\n>> available?\n> \n> It has some unresolved problems so I wasn't spreading it around everywhere.\n> \n> It is based on cvs2svn from August. There has been too much change to\n> the current cvs2svn to merge it anymore. [...]\n> \n> If the repo is missing branch tags cvs2svn may turn a single missing\n> branch into hundreds of branches. The Mozilla repo has about 1000\n> extra branches because of this.\n\n[To explain to our studio audience:] Currently, if there is an actual\nbranch in CVS but no symbol associated with it, cvs2svn generates branch\nlabels like \"unlabeled-1.2.3\", where \"1.2.3\" is the branch revision\nnumber in CVS for the particular file.  The problem is that the branch\nrevision numbers for files in the same logical branch are usually\ndifferent.  That is why many extra branches are generated.\n\nSuch unnamed branches cannot reasonably be accessed via CVS anyway, and\nsomebody probably made the conscious decision to delete the branch from\nCVS (though without doing it correctly).  Therefore such revisions are\nprobably garbage.  It would be easy to add an option to discard such\nrevisions, and we should probably do so.  (In fact, they can already be\nexcluded with \"--exclude=unlabeled-.*\".)  The only caveat is that it is\npossible for other, named branches to sprout from an unnamed branch.  In\nthis case either the second branch would have to be excluded too, or the\nunlabeled branch would have to be included.\n\nAlternatively, there was a suggestion to add heuristics to guess which\nfiles' \"unlabeled\" branches actually belong in the same original branch.\n This would be a lot of work, and the result would never be very\naccurate (for one thing, there is no evidence of the branch whatsoever\nin files that had no commits on the branch).\n\nOther ideas are welcome.\n\nMichael\n"},{"id":"26899","messageId":"eec2aa$9c9$2@sea.gmane.org","threadId":"5549","inReplyTo":"45098AE0.6030409@alum.mit.edu","subject":"Re: cvs import","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2006-09-14T17:08:49Z","receivedAt":"2006-09-14T17:08:49Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Michael Haggerty wrote:\n\n> Alternatively, there was a suggestion to add heuristics to guess which\n> files' \"unlabeled\" branches actually belong in the same original branch.\n>  This would be a lot of work, and the result would never be very\n> accurate (for one thing, there is no evidence of the branch whatsoever\n> in files that had no commits on the branch).\n> \n> Other ideas are welcome.\n\nInterpolate the state of repository according to timestamps, with some\ncoarse-grainess of course.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"26902","messageId":"9e4733910609141017r37fbbd45q5fcf7d0f39b48cf3@mail.gmail.com","threadId":"5549","inReplyTo":"45098AE0.6030409@alum.mit.edu","subject":"Re: cvs import","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-09-14T17:17:43Z","receivedAt":"2006-09-14T17:17:43Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 9/14/06, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> Jon Smirl wrote:\n> > On 9/14/06, Jakub Narebski <jnareb@gmail.com> wrote:\n> >> Shawn Pearce wrote:\n> >>\n> >> > Originally I wanted Jon Smirl to modify the cvs2svn (...)\n> >>\n> >> By the way, will cvs2git (modified cvs2svn) and git-fast-import publicly\n> >> available?\n> >\n> > It has some unresolved problems so I wasn't spreading it around everywhere.\n> >\n> > It is based on cvs2svn from August. There has been too much change to\n> > the current cvs2svn to merge it anymore. [...]\n> >\n> > If the repo is missing branch tags cvs2svn may turn a single missing\n> > branch into hundreds of branches. The Mozilla repo has about 1000\n> > extra branches because of this.\n>\n> [To explain to our studio audience:] Currently, if there is an actual\n> branch in CVS but no symbol associated with it, cvs2svn generates branch\n> labels like \"unlabeled-1.2.3\", where \"1.2.3\" is the branch revision\n> number in CVS for the particular file.  The problem is that the branch\n> revision numbers for files in the same logical branch are usually\n> different.  That is why many extra branches are generated.\n>\n> Such unnamed branches cannot reasonably be accessed via CVS anyway, and\n> somebody probably made the conscious decision to delete the branch from\n> CVS (though without doing it correctly).  Therefore such revisions are\n> probably garbage.  It would be easy to add an option to discard such\n> revisions, and we should probably do so.  (In fact, they can already be\n> excluded with \"--exclude=unlabeled-.*\".)  The only caveat is that it is\n> possible for other, named branches to sprout from an unnamed branch.  In\n> this case either the second branch would have to be excluded too, or the\n> unlabeled branch would have to be included.\n\nIn MozCVS there are important branches where the first label has been\ndeleted but there are subsequent branches off from the first branch.\nThese subsequent branches are still visible in CVS. Someone else had\nthis same problem on the cvs2svn list. This has happen twice on major\nbranches.\n\nManually looking at one of these it looks like the author wanted to\nchange the branch name. They made a branch with the wrong name,\nbranched again with the new name, and deleted the first branch.\n\n> Alternatively, there was a suggestion to add heuristics to guess which\n> files' \"unlabeled\" branches actually belong in the same original branch.\n>  This would be a lot of work, and the result would never be very\n> accurate (for one thing, there is no evidence of the branch whatsoever\n> in files that had no commits on the branch).\n\nYou wrote up a detailed solution for this a few weeks ago on the\ncvs2svn list. The basic idea is to look at the change sets on the\nunlabeled branches. If change sets span multiple unlabeled branches,\nthere should be one unlabeled branch instead of multiple ones. That\nwould work to reduce the number of unlabeled branches down from 1000\nto the true number which I believe is in the 10-20 range.\n\nWould the dependency based model make these relationships more obvious?\n\n>\n> Other ideas are welcome.\n>\n> Michael\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"26932","messageId":"20060914215728.GL23891@pasky.or.cz","threadId":"5549","inReplyTo":"20060914015324.GX29625@bcd.geek.com.au","subject":"Re: [Monotone-devel] cvs import","fromName":"Petr Baudis","fromEmail":"pasky@suse.cz","sentAt":"2006-09-14T21:57:28Z","receivedAt":"2006-09-14T21:57:28Z","isPatch":false,"sender":{"key":"pasky@ucw.cz","avatar":"https://avatars.githubusercontent.com/u/18439?v=4"},"body":"Dear diary, on Thu, Sep 14, 2006 at 03:53:24AM CEST, I got a letter\nwhere Daniel Carosone <dan@geek.com.au> said that...\n> On Wed, Sep 13, 2006 at 08:57:33PM -0400, Jon Smirl wrote:\n> > Mozilla is 120,000 files. The complexity comes from 10 years worth of\n> > history. A few of the files have around 1,700 revisions. There are\n> > about 1,600 branches and 1,000 tags. The branch number is inflated\n> > because cvs2svn is generating extra branches, the real number is\n> > around 700. The CVS repo takes 4.2GB disk space. cvs2svn turns this\n> > into 250,000 commits over about 1M unique revisions.\n> \n> Those numbers are pretty close to those in the NetBSD repository, and\n> between them these probably represent just about the most extensive\n> public CVS test data available. \n\n  Don't forget OpenOffice. It's just a shame that the OpenOffice CVS\ntree is not available for cloning.\n\n\thttp://wiki.services.openoffice.org/wiki/SVNMigration\n\n-- \n\t\t\t\tPetr \"Pasky\" Baudis\nStuff: http://pasky.or.cz/\nSnow falling on Perl. White noise covering line noise.\nHides all the bugs too. -- J. Putnam\n"},{"id":"26933","messageId":"20060914220443.GB11129@spearce.org","threadId":"5549","inReplyTo":"20060914215728.GL23891@pasky.or.cz","subject":"Re: [Monotone-devel] cvs import","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-09-14T22:04:43Z","receivedAt":"2006-09-14T22:04:43Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Petr Baudis <pasky@suse.cz> wrote:\n>   Don't forget OpenOffice. It's just a shame that the OpenOffice CVS\n> tree is not available for cloning.\n> \n> \thttp://wiki.services.openoffice.org/wiki/SVNMigration\n\nHmm, the KDE repo is even larger than Mozilla: 19 GB in CVS and\n499,367 revisions.  Question is, are those distinct file revisions\nor SVN revisions?  And just what machine did they use that completed\nthat conversion in 38 hours?\n\n-- \nShawn.\n"},{"id":"26949","messageId":"450A581E.2050509@bluegap.ch","threadId":"5549","inReplyTo":"20060914155003.GB9657@spearce.org","subject":"Re: cvs import","fromName":"Markus Schiltknecht","fromEmail":"markus@bluegap.ch","sentAt":"2006-09-15T07:37:02Z","receivedAt":"2006-09-15T07:37:02Z","isPatch":false,"sender":{"key":"markus@bluegap.ch","avatar":"https://gravatar.com/avatar/2f3aadbc46c7c942fa9301d0bb9b91da8c7186233f4c7eff9667a8b53c8cc82e?d=mp&s=160"},"body":"Hi,\n\nShawn Pearce wrote:\n> I don't know how the Monotone guys feel about it but I think Git\n> is happy with the data in any order, just so long as the dependency\n> chains aren't fed out of order.  Which I think nearly all changeset\n> based SCMs would have an issue with.  So we should be just fine\n> with the current chronological order produced by cvs2svn.\n\nI'd vote for splitting into file data (and delta / patches) import and \nmetadata import (author, changelog, DAG).\n\nMonotone would be happiest if the file data were sent one file after \nanother and (inside each file) in the order of each file's single \nhistory. That guarantees good import performance for monotone. I imagine \nit's about the same for git. And if you have to somehow cache the files \nanyway, subversion will benefit, too. (Well, at least the cache will \nthank us with good performance).\n\nAfter all file data has been delivered, the metadata can be delivered. \nAs neigther monotone nor git care much if they are chronological across \nbranches, I'd vote for doing it that way.\n\nRegards\n\nMarkus\n"},{"id":"26991","messageId":"20060916033917.GA24269@spearce.org","threadId":"5549","inReplyTo":"450A581E.2050509@bluegap.ch","subject":"Re: cvs import","fromName":"Shawn Pearce","fromEmail":"spearce@spearce.org","sentAt":"2006-09-16T03:39:18Z","receivedAt":"2006-09-16T03:39:18Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Markus Schiltknecht <markus@bluegap.ch> wrote:\n> Shawn Pearce wrote:\n> >I don't know how the Monotone guys feel about it but I think Git\n> >is happy with the data in any order, just so long as the dependency\n> >chains aren't fed out of order.  Which I think nearly all changeset\n> >based SCMs would have an issue with.  So we should be just fine\n> >with the current chronological order produced by cvs2svn.\n> \n> I'd vote for splitting into file data (and delta / patches) import and \n> metadata import (author, changelog, DAG).\n> \n> Monotone would be happiest if the file data were sent one file after \n> another and (inside each file) in the order of each file's single \n> history. That guarantees good import performance for monotone. I imagine \n> it's about the same for git. And if you have to somehow cache the files \n> anyway, subversion will benefit, too. (Well, at least the cache will \n> thank us with good performance).\n>\n> After all file data has been delivered, the metadata can be delivered. \n> As neigther monotone nor git care much if they are chronological across \n> branches, I'd vote for doing it that way.\n\nRight.  I think that one of the cvs2svn guys had the right idea\nhere.  Provide two hooks: one early during the RCS file parse which\nsupplies a backend each full text file revision and another during\nthe very last stage which includes the \"file\" in the metadata stream\nfor commit.\n\nThis would give Git and Monotone a way to grab the full text for each\nfile and stream them out up front, then include only a \"token\" in the\nmetadata stream which identifies the specific revision.  Meanwhile\nSVN can either cache the file revision during the early part or\nignore it, then dump out the full content during the metadata.\n\n\nAs it happens Git doesn't care what order the file revisions come in.\nIf we don't repack the imported data we would prefer to get the\nrevisions in newest->oldest order so we can delta the older versions\nagainst the newer versions (like RCS).  This is also happens to be\nthe fastest way to extract the revision data from RCS.\n\nOn the other hand from what I understand of Monotone it needs\nthe revisions in oldest->newest order, as does SVN.\n\nDoing both orderings in cvs2noncvs is probably ugly.  Doing just\noldest->newest (since 2/3 backends want that) would be acceptable\nbut would slow down Git imports as the RCS parsing overhead would\nbe much higher.\n\n-- \nShawn.\n"},{"id":"26994","messageId":"20060916060454.GA3769@ugly.local","threadId":"5549","inReplyTo":"20060916033917.GA24269@spearce.org","subject":"Re: cvs import","fromName":"Oswald Buddenhagen","fromEmail":"ossi@kde.org","sentAt":"2006-09-16T06:04:55Z","receivedAt":"2006-09-16T06:04:55Z","isPatch":false,"sender":{"key":"ossi@kde.org","avatar":"https://avatars.githubusercontent.com/u/812380?v=4"},"body":"On Fri, Sep 15, 2006 at 11:39:18PM -0400, Shawn Pearce wrote:\n> On the other hand from what I understand of Monotone it needs\n> the revisions in oldest->newest order, as does SVN.\n> \n> Doing both orderings in cvs2noncvs is probably ugly.\n>\ndon't worry, as i know mike, he'll come up with an abstract, outright\nbeautiful interface that makes you want to implement middle->oldnewest\njust for the sake of doing it. :)\n\n-- \nHi! I'm a .signature virus! Copy me into your ~/.signature, please!\n--\nChaos, panic, and disorder - my work here is done.\n"},{"id":"26996","messageId":"20060916062109.GB1779@frances.vorpus.org","threadId":"5549","inReplyTo":"20060916033917.GA24269@spearce.org","subject":"Re: Re: cvs import","fromName":"Nathaniel Smith","fromEmail":"njs@pobox.com","sentAt":"2006-09-16T06:21:09Z","receivedAt":"2006-09-16T06:21:09Z","isPatch":false,"sender":{"key":"njs@pobox.com","avatar":null},"body":"On Fri, Sep 15, 2006 at 11:39:18PM -0400, Shawn Pearce wrote:\n> On the other hand from what I understand of Monotone it needs\n> the revisions in oldest->newest order, as does SVN.\n\nMonotone stores file deltas new->old, similar to git.  It should be\nreasonably efficient at turning them around if it has to, though -- so\nlong as you give all the versions of a single file at a time, so\nthere's some reasonable locality, instead of jumping all around the\ntree.\n\n-- Nathaniel\n\n-- \n\"...All of this suggests that if we wished to find a modern-day model\nfor British and American speech of the late eighteenth century, we could\nprobably do no better than Yosemite Sam.\"\n"}]}