{"thread":{"id":"9333","subject":"cvs2svn conversion directly to git ready for experimentation","startedAt":"2007-08-01T00:09:34Z","lastAt":"2007-08-05T07:58:00Z","messageCount":40,"participants":["Michael Haggerty","Johannes Schindelin","Jakub Narebski","Steffen Prohaska","Simon 'corecode' Schubert","Marko Macek","Robin Rosenberg","Lübbe Onken","Linus Torvalds","Martin Langhoff","Jon Smirl","Shawn O. Pearce","Patwardhan, Rajesh","Oswald Buddenhagen"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"49253","messageId":"46AFCF3E.5010805@alum.mit.edu","threadId":"9333","inReplyTo":null,"subject":"cvs2svn conversion directly to git ready for experimentation","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-08-01T00:09:34Z","receivedAt":"2007-08-01T00:09:34Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"I am the maintainer of cvs2svn[1], which is a program for one-time\nconversions from CVS to Subversion.  cvs2svn is very robust against the\nmany peculiarities of CVS and can convert just about every CVS\nrepository we have ever seen.\n\nI've been working on a cvs2svn output pass that writes the converted CVS\nrepository directly into git rather than Subversion.  The code runs now\nwith at least one repository from our test suite of nasty CVS repositories.\n\nUnfortunately, I am a complete git newbie, so I would very much\nappreciate help from the git community with feedback and checking\nwhether the conversion output is reasonable and gitlike.\n\nThe git output is very preliminary and virtually untested, and has the\nfollowing limitations (hopefully to be removed in the near future):\n\n- It is rather slow.  Among other things, it still uses RCS or CVS to\nextract the contents of the CVS revisions, which will soon be changed to\nwin a factor of 2 or so.\n\n- CVS allows a branch to be created from arbitrary combinations of\nsource revisions and/or source branches.  cvs2svn tries to create a\nbranch from a single source, but if it can't figure out how to, it\ncreates the branch using \"merge\" from multiple sources.  In pathological\nsituations, the number of merge sources for a branch can be arbitrarily\nlarge.\n\n- It is not very intelligent about creating tags.  When asked to create\na tag, it unconditionally creates a \"tag fixup branch\"[2] with the same\nname and contents as the tag, then tags this branch.  The tag fixup\nbranch is never deleted.\n\n- There are no checks that CVS branch and tag names are legal git names,\nor indeed that any other similar limitations of git are honored.\n\n- The data that should be fed to git-fast-input is written to two files,\nwhich have to be loaded into git-fast-import manually.  Eventually I\nwill add an option to invoke git-fast-import automatically and pipe the\noutput directly into git-fast-import.\n\n- Only single projects can be converted at a time.  I don't think that\nthis will be a significant limitation when outputting to git.\n\n\nTo try it out:\n\n1. Install svn (to be able to check out cvs2svn) and either cvs or rcs.\n\n2. Check out the current trunk version of cvs2svn:\n\n    svn co http://cvs2svn.tigris.org/svn/cvs2svn/trunk cvs2svn-trunk\n    cd cvs2svn-trunk\n    make check # ...optional\n\n3. Configure cvs2svn for your conversion.  This has to be done via the\n\"options-file method\"[3].  See cvs2svn-example.options and\ntest-data/main-cvsrepos/cvs2svn-git.options as examples; the former file\nincludes voluminous documentation.\n\n4. Run cvs2svn.  This outputs two git-fast-import files, with the names\nspecified by your options file.  In the example, these files are named\n'cvs2svn-tmp/git-blob.dat' and 'cvs2svn-tmp/git-dump.dat'.\n\n5. Initialize a git repository, and load the dump files using\ngit-fast-import:\n\n    git-init\n    cat cvs2svn-tmp/git-blob.dat | \\\n        git-fast-import --export-marks=cvs2svn-tmp/git-marks.dat\n    cat cvs2svn-tmp/git-dump.dat | \\\n        git-fast-import --import-marks=cvs2svn-tmp/git-marks.dat\n\n\nI am looking forward to your feedback.  Even better would be if somebody\nwants to join forces on this project.  I would be happy to supply the\ncvs2svn knowledge if you can bring the git experience.\n\nMichael\n\n\n[1] http://cvs2svn.tigris.org/\n[2] http://www.kernel.org/pub/software/scm/git/docs/git-fast-import.html\n[3] http://cvs2svn.tigris.org/cvs2svn.html#cmd-vs-options\n"},{"id":"49261","messageId":"Pine.LNX.4.64.0708010138070.14781@racer.site","threadId":"9333","inReplyTo":"46AFCF3E.5010805@alum.mit.edu","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-08-01T00:41:17Z","receivedAt":"2007-08-01T00:41:17Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Wed, 1 Aug 2007, Michael Haggerty wrote:\n\n> 2. Check out the current trunk version of cvs2svn:\n> \n>     svn co http://cvs2svn.tigris.org/svn/cvs2svn/trunk cvs2svn-trunk\n>     cd cvs2svn-trunk\n>     make check # ...optional\n\nFWIW I tried to clone it with \"git svn\", and needed to prefix the url with \n\"guest\", i.e.\n\n\t$ git clone http://guest@cvs2svn.tigris.org/svn/cvs2svn/trunk\n\nand it still did not work at once.  Somehow I managed to get the \n\"Username\" prompt, input \"guest\", and left the password empty.  Even then, \nonly the second attempt succeeded (I guess somehow that \"password\" got \nstored in $HOME/.subversion/auth/...\n\nCiao,\nDscho\n"},{"id":"49356","messageId":"f8r09t$qdg$1@sea.gmane.org","threadId":"9333","inReplyTo":"46AFCF3E.5010805@alum.mit.edu","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-08-01T22:09:03Z","receivedAt":"2007-08-01T22:09:03Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Michael Haggerty wrote:\n\n> I am the maintainer of cvs2svn[1], which is a program for one-time\n> conversions from CVS to Subversion.  cvs2svn is very robust against the\n> many peculiarities of CVS and can convert just about every CVS\n> repository we have ever seen.\n> \n> I've been working on a cvs2svn output pass that writes the converted CVS\n> repository directly into git rather than Subversion.  The code runs now\n> with at least one repository from our test suite of nasty CVS repositories.\n\nHave you contacted Jon Smirl about his unpublished work on cvs2git,\ncvs2svn based CVS to Git converter?\n\nQuote from InterfacesFrontendsAndTools page on GIT wiki[1]:\n\n  cvs2git is the unofficial name of Jon Smirl's modifications to cvs2svn.\n  These modifications allow cvs2svn to generate a data stream which is\n  consumed by Shawn Pearce's git-fast-import (now included in git.git).\n  git-fast-import converts its input stream directly into a Git .pack file,\n  minimizing the amount of IO required on large imports.\n\n  Jon Smirl stopped working on cvs2git[2] because first, Mozilla (which was\n  main target of his work) decided that to not to move to git, and second\n  because of troubles with cvs2svn architecture[*] (which it is based on).\n  Jon Smirl has posted his impressions on working on CVS importer in \n  \"Some tips for doing a CVS importer\" thread[3].\n\nReferences:\n-----------\n[1] http://git.or.cz/gitwiki/InterfacesFrontendsAndTools#head-23858c2cde0cef60443d8e73e6829a95f8e191ef\n[2] http://msgid.gmane.org/9e4733910611190940y147992b8mbdfac5a51f42e0fe@mail.gmail.com\n[3] http://marc.theaimsgroup.com/?t=116405956000001&r=1&w=2\n\nFootnotes:\n----------\n[*] If I remember correctly authors of cvs2svn were talking about separating\nthe code dealing with disentangling CVS repository structure from the part\ntranslating it into Subversion repository (with its quirks), and the part\ngenerating Subversion repository.\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"49396","messageId":"65F1862F-4DF2-4A52-9FD5-20802AEACDAB@zib.de","threadId":"9333","inReplyTo":"46AFCF3E.5010805@alum.mit.edu","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Steffen Prohaska","fromEmail":"prohaska@zib.de","sentAt":"2007-08-02T08:49:29Z","receivedAt":"2007-08-02T08:49:29Z","isPatch":false,"sender":{"key":"prohaska@zib.de","avatar":"https://avatars.githubusercontent.com/u/217580?v=4"},"body":"Michael,\n\nOn Aug 1, 2007, at 2:09 AM, Michael Haggerty wrote:\n\n> I am looking forward to your feedback.  Even better would be if  \n> somebody\n> wants to join forces on this project.  I would be happy to supply the\n> cvs2svn knowledge if you can bring the git experience.\n\nI tried it with revision trunk@3930 of cvs2svn. The results are as  \nfollows.\n\nsome WARNING: problem encoding log message: [...]\n\ncvs2svn Statistics:\n------------------\nTotal CVS Files:              9578\nTotal CVS Revisions:         66771\nTotal CVS Branches:         229121\nTotal CVS Tags:             371259\nTotal Unique Tags:             112\nTotal Unique Branches:          79\nCVS Repos Size in KB:       210390\nTotal SVN Commits:           18178\nFirst Revision Date:    Fri Jul 23 10:26:11 1999\nLast Revision Date:     Thu Jul 19 17:50:40 2007\n------------------\nTimings (seconds):\n------------------\n3295   pass1    CollectRevsPass\n    0   pass2    CollateSymbolsPass\n3642   pass3    FilterSymbolsPass\n    0   pass4    SortRevisionSummaryPass\n    1   pass5    SortSymbolSummaryPass\n  109   pass6    InitializeChangesetsPass\n   56   pass7    BreakRevisionChangesetCyclesPass\n   66   pass8    RevisionTopologicalSortPass\n   54   pass9    BreakSymbolChangesetCyclesPass\n   99   pass10   BreakAllChangesetCyclesPass\n   92   pass11   TopologicalSortPass\n   46   pass12   CreateRevsPass\n    7   pass13   SortSymbolsPass\n    2   pass14   IndexSymbolsPass\n   70   pass15   OutputPass\n7540   total\n\n\nI checked that CVS head and two other branches match when checked\nout from CVS and from the imported git archive. Everything is ok\n(ignoring some differences introduced by keyword expansion).\nNote, I tried earlier to use cvs2svn to import to svn followed by\ngit-svnimport to import to git. The repository resulting from\nthis two step import not even passed this minimal requirement of\nmatching checkouts from cvs and git.\n\ncvs2svn created a lot of branches that are not present in CVS,\nwith names identical to CVS tags. Apparently these branches are\nused to create a commit matching a certain CVS tag.\n\nI checked one suspicious commit that indicates to me if the root\npoints of branches are right. Note, git-cvsimport fails this check;\nparsecvs and cvs2svn pass the check.\n\nThe branching structure looks, ... hmm ..., interesting. cvs2svn\nmanufactured commits to get the branching points right.\nApparently our CVS has some weired commits like 'unlabeled-1.1.1'\nand two other named tags (maybe vendor branches?) that cause\nthese manufactured commits. In gitk I see long lines running\nparallel to the cvs trunk all down to these weired CVS tags. They\nare not very useful, altough they might be correct. Note,\nparsecvs imports our repository without such basically useless\nlinks.  However, I can't verify if parsecvs gets something wrong.\nOther branches are created over a couple of commits mixing in\nseveral branches (maybe again our weired commits already\nmentioned). See branching1.png, branching2.png, branching3.png.\n[ I have to apologize, our cvs repository contains proprietary\n   information, so I can't publish it's history freely. ]\n\ncvs2svn is the first tool besided parsecvs that worked for me,\nthat is imported the whole repository, passed the basic test of\nmatching checkouts from cvs and git, and got the one suspicious\ncommit right that I'm using for verifying the branching points.\n\n[ I have no time to go into the details of all these tests.\n   Therefore only a very short summary:\n   All tools needed basic cleanup of a few corrupted ,v files and\n      ,v files that were duplicated in Attic.\n   git-cvsimport fails to create branches at the right commit.\n   fromcvs's togit surrendered during the import.\n   fromcvs's tohg accepted more of the history, but finally\n     surrendered as well.\n   parsecvs works for me (crashes on corrupted ,v files).\n   cvs2svn followed by git-svnimport create wrong state at the\n     tips of branches.\n   cvs2svn direct git import works for me (reports corrupted ,v files).\n   ]\n\nRight now, I'd prefer the import by parsecvs because of the\nsimpler history. However, I don't know if I loose history\ninformation by doing so. I'd start by a run of cvs2svn to validate\nthe overall structure of the CVS repository. Dealing with corruption\nin the CVS repository seems to be superior in cvs2svn. It reports\nerrors when parsecvs just crashes.\n\n\n\tSteffen\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n"},{"id":"49436","messageId":"46B1F96B.7050107@alum.mit.edu","threadId":"9333","inReplyTo":"8b65902a0708010438s24d16109k601b52c04cf9c066@mail.gmail.com","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-08-02T15:34:03Z","receivedAt":"2007-08-02T15:34:03Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"[I am CCing this response to the mailing lists.]\n\nGuilhem Bonnefille wrote:\n> On 8/1/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> I am the maintainer of cvs2svn[1], which is a program for one-time\n>> conversions from CVS to Subversion.  cvs2svn is very robust against the\n>> many peculiarities of CVS and can convert just about every CVS\n>> repository we have ever seen.\n> \n> What are the differences with cvsps ( http://www.cobite.com/cvsps/ )?\n\nI'm not extremely familiar with cvsps, and I don't really want to get\ninto a \"my-tool-is-better-than-your-tool\" kind of argument.  Instead I\nwill mention that the goals of the two projects are somewhat different:\n\ncvs2svn is meant for one-time conversions from CVS, and therefore aims\nfor maximum conversion accuracy, robustness even in the presence of some\nkinds of CVS repository corruption, intelligent translation of CVS\nidioms to the idioms of a modern SCM, and scalability to large\nrepositories (by using on-disk databases instead of RAM for intermediate\ndata).  Conversion speed is not a primary goal of cvs2svn, and\nincremental conversions are not supported at all.  cvs2svn requires\nfilesystem access to the CVS repository (it parses the RCS files directly).\n\ncvsps is not a conversion tool at all, though it is used by other\nconversion tools to generate the changesets.  It appears (I hope I am\nnot misinterpreting things) to emphasize speed and incremental\noperation, for example attempting to make changesets consistent from one\nrun to the next, even if the CVS repository has been changed prudently\nbetween runs.  cvsps does not appear to attempt to create atomic branch\nand tag creation commits or handle CVS's special vendorbranch behavior.\n cvsps operates via the CVS protocol; you don't need filesystem access\nto the CVS repository.\n\nI can also point you to a list of cvs2svn features, which includes a\nlist of some of the CVS quirks that it knows how to handle:\n\n    http://cvs2svn.tigris.org/cvs2svn.html#features\n\ncvs2svn includes a large suite of perverse CVS repositories that we use\nfor testing.  Many of them are derived from real-life CVS repositories\nthat people have had problems with.  It would be very interesting to see\nhow other conversion tools handle these repositories, but I don't expect\nto have time to do so in the near future.\n\nMichael\n"},{"id":"49442","messageId":"46B20D1F.1020308@alum.mit.edu","threadId":"9333","inReplyTo":"f8r09t$qdg$1@sea.gmane.org","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-08-02T16:58:07Z","receivedAt":"2007-08-02T16:58:07Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Jakub Narebski wrote:\n> Michael Haggerty wrote:\n> Have you contacted Jon Smirl about his unpublished work on cvs2git,\n> cvs2svn based CVS to Git converter?\n\nYes, I am familiar with Jon Smirl's work, and as soon as he let us know\nwhat he was working on, we tried to help.  Unfortunately the cooperation\nwas not very fruitful.\n\n- While Jon was (unknown to us) working on his git output patch, I was\nworking on a big cvs2svn rewrite to make cvs2svn more robust and easier\nto hack.  By the time he contacted us, his patch did not apply to the\ncvs2svn code.  The refactoring that obsoleted the patch, in fact, was\nlargely to remedy the very same architectural problems that were\nhampering his work.\n\n- In my opinion, Jon misdiagnosed the reason for the \"fragmented branch\ncreation\" problem that he claimed was preventing a clean conversion to\ngit, and he felt that we were not interested in fixing the problem.  In\nfact, I was working on fixing another problem that I believe was the\n*real* reason for the fragmented branch creation.  This fix is\nimplemented in cvs2svn version 2.0.\n\n> Footnotes:\n> ----------\n> [*] If I remember correctly authors of cvs2svn were talking about separating\n> the code dealing with disentangling CVS repository structure from the part\n> translating it into Subversion repository (with its quirks), and the part\n> generating Subversion repository.\n\nYes, this is now done, which was why it was only a couple of days of\nprogramming for me to add a git output option.\n\nMichael\n"},{"id":"49447","messageId":"46B2132D.7090304@alum.mit.edu","threadId":"9333","inReplyTo":"65F1862F-4DF2-4A52-9FD5-20802AEACDAB@zib.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-08-02T17:23:57Z","receivedAt":"2007-08-02T17:23:57Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Steffen Prohaska wrote:\n> On Aug 1, 2007, at 2:09 AM, Michael Haggerty wrote:\n>> I am looking forward to your feedback.  Even better would be if somebody\n>> wants to join forces on this project.  I would be happy to supply the\n>> cvs2svn knowledge if you can bring the git experience.\n> \n> I tried it with revision trunk@3930 of cvs2svn. The results are as follows.\n\nThanks for the feedback!\n\n> cvs2svn created a lot of branches that are not present in CVS,\n> with names identical to CVS tags. Apparently these branches are\n> used to create a commit matching a certain CVS tag.\n\nThat is correct.  This is something that I plan to work on, at least for\ntags that can be created from a single source commit.\n\n> The branching structure looks, ... hmm ..., interesting. cvs2svn\n> manufactured commits to get the branching points right.\n> Apparently our CVS has some weired commits like 'unlabeled-1.1.1'\n> and two other named tags (maybe vendor branches?) that cause\n> these manufactured commits. In gitk I see long lines running\n> parallel to the cvs trunk all down to these weired CVS tags. They\n> are not very useful, altough they might be correct. Note,\n> parsecvs imports our repository without such basically useless\n> links.  However, I can't verify if parsecvs gets something wrong.\n\nBranches with names like \"unlabeled-1.1.1\" come from CVS branches for\nwhich the revisions are still contained in the RCS files but for which\nthe branch name has been deleted.  These wreak havoc on cvs2svn's\nattempt to find simple branch sources and cause a proliferation of\nbasically useless branches.  The main problem is that cvs2svn does not\nattempt to figure out that \"unlabeled-1.2.4\" in one file might be the\nsame as \"unlabeled-1.2.6\" in another etc.\n\nAn \"unlabeled-1.1.1\", in particular, means that the branch whose name\nwas deleted was a vendor branch.  The deletion of a vendor branch name\ncan cause even more mayhem.\n\nIn most cases it makes sense to exclude the unlabeled branches.  After\nall, somebody tried to delete them, so they can't be that important,\nright?  Use --exclude='unlabeled-.*', or add a line like this to your\noptions file:\n\nctx.symbol_strategy.add_rule(ExcludeRegexpStrategyRule(r'unlabeled-.*'))\n\n.  This can of course cause problems if other branches or tags were\ncreated that branched off of the unlabeled branch.  In such cases the\ndependent branches/tags might have to be excluded too.\n\n> Other branches are created over a couple of commits mixing in\n> several branches (maybe again our weired commits already\n> mentioned). See branching1.png, branching2.png, branching3.png.\n> [ I have to apologize, our cvs repository contains proprietary\n>   information, so I can't publish it's history freely. ]\n\nThis can definitely be caused by unlabeled branches.  It can also be\ncaused by branches rooted in a vendor branch.  In many cases, such\nbranches can actually be grafted onto trunk, but cvs2svn does not (yet)\nattempt this.\n\n> cvs2svn is the first tool besided parsecvs that worked for me,\n> that is imported the whole repository, passed the basic test of\n> matching checkouts from cvs and git, and got the one suspicious\n> commit right that I'm using for verifying the branching points.\n> \n> [ I have no time to go into the details of all these tests.\n>   Therefore only a very short summary:\n>   All tools needed basic cleanup of a few corrupted ,v files and\n>      ,v files that were duplicated in Attic.\n>   git-cvsimport fails to create branches at the right commit.\n>   fromcvs's togit surrendered during the import.\n>   fromcvs's tohg accepted more of the history, but finally\n>     surrendered as well.\n>   parsecvs works for me (crashes on corrupted ,v files).\n>   cvs2svn followed by git-svnimport create wrong state at the\n>     tips of branches.\n>   cvs2svn direct git import works for me (reports corrupted ,v files).\n>   ]\n\nThanks very much for this interesting summary.\n\n> Right now, I'd prefer the import by parsecvs because of the\n> simpler history. However, I don't know if I loose history\n> information by doing so. I'd start by a run of cvs2svn to validate\n> the overall structure of the CVS repository. Dealing with corruption\n> in the CVS repository seems to be superior in cvs2svn. It reports\n> errors when parsecvs just crashes.\n\nIf excluding the unlabeled branches does not fix things for you, I\nsuggest checking out the first revision on such a branch, and comparing\nthe results from CVS, from parsecvs, and from cvs2svn.  It *should* be\nthat the version of the file from the vendor branch is included in the\nworking copy.  cvs2svn should handle this correctly.  I am curious\nwhether parsecvs does.\n\nMichael\n"},{"id":"49448","messageId":"46B215E2.8010307@fs.ei.tum.de","threadId":"9333","inReplyTo":"65F1862F-4DF2-4A52-9FD5-20802AEACDAB@zib.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-08-02T17:35:30Z","receivedAt":"2007-08-02T17:35:30Z","isPatch":false,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"Steffen Prohaska wrote:\n>   fromcvs's togit surrendered during the import.\n>   fromcvs's tohg accepted more of the history, but finally\n>     surrendered as well.\n\nWhich repo is it you are converting?  Is this available somewhere?\n\nI'd appreciate any reports concerning \"surrenders\" of fromcvs.  Additionally, it seems strange that tohg should have worked \"better\" than togit, as these are basically just different backends.\n\ncheers\n  simon\n"},{"id":"49459","messageId":"EDE86758-FFD0-4CED-A2C9-033FA13DD3B6@zib.de","threadId":"9333","inReplyTo":"46B215E2.8010307@fs.ei.tum.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Steffen Prohaska","fromEmail":"prohaska@zib.de","sentAt":"2007-08-02T19:13:40Z","receivedAt":"2007-08-02T19:13:40Z","isPatch":false,"sender":{"key":"prohaska@zib.de","avatar":"https://avatars.githubusercontent.com/u/217580?v=4"},"body":"Simon,\n\nOn Aug 2, 2007, at 7:35 PM, Simon 'corecode' Schubert wrote:\n\n> Steffen Prohaska wrote:\n>>   fromcvs's togit surrendered during the import.\n>>   fromcvs's tohg accepted more of the history, but finally\n>>     surrendered as well.\n>\n> Which repo is it you are converting?  Is this available somewhere?\n\nUnfortunately not, the content is a proprietary software package.\n\n\n> I'd appreciate any reports concerning \"surrenders\" of fromcvs.   \n> Additionally, it seems strange that tohg should have worked  \n> \"better\" than togit, as these are basically just different backends.\n\nSome time passed since I did the tests. I had no time to do a\ndetailed investigation then. I'll have more time now and will\nprepare a bug report, which is not easy because I can't sent you\nthe cvs repo, sorry. Any hints what would be most helpful for you?\n\nI remember that togit reported a broken pipe. My feeling was\nthat git-fastimport aborted, which may be reason why tohg\nworked better. I didn't try to understand more details. I never\nread ruby code before and it was already a challenge for me to\nget everything up and running (rcs, rbtree).\n\n\tSteffen\n"},{"id":"49460","messageId":"46B22EF7.80108@gmx.net","threadId":"9333","inReplyTo":"46B2132D.7090304@alum.mit.edu","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Marko Macek","fromEmail":"marko.macek@gmx.net","sentAt":"2007-08-02T19:22:31Z","receivedAt":"2007-08-02T19:22:31Z","isPatch":false,"sender":{"key":"marko.macek@gmx.net","avatar":null},"body":"Michael Haggerty wrote:\n> This can definitely be caused by unlabeled branches.  It can also be\n> caused by branches rooted in a vendor branch.  In many cases, such\n> branches can actually be grafted onto trunk, but cvs2svn does not (yet)\n> attempt this.\n\nIt would be nice to be able to exclude the vendor branch if only \nthe initial commit was made on it (or maybe handle it better, by \nremapping the commits to the main branch when they match).\n\nI have tested this on my repository and currently gitk draws \nlarge 'railroad switching stations' because many tags have the \nvendor branch as a parent (and in some cases also the parent branch, \nin addition to the parent commit).\n\n\tMark\n\n\n---------------------------------------------------------------------\nTo unsubscribe, e-mail: users-unsubscribe@cvs2svn.tigris.org\nFor additional commands, e-mail: users-help@cvs2svn.tigris.org"},{"id":"49461","messageId":"46B2309E.3060804@fs.ei.tum.de","threadId":"9333","inReplyTo":"EDE86758-FFD0-4CED-A2C9-033FA13DD3B6@zib.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-08-02T19:29:34Z","receivedAt":"2007-08-02T19:29:34Z","isPatch":false,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"Steffen Prohaska wrote:\n> I remember that togit reported a broken pipe. My feeling was\n> that git-fastimport aborted, which may be reason why tohg\n> worked better. I didn't try to understand more details. I never\n> read ruby code before and it was already a challenge for me to\n> get everything up and running (rcs, rbtree).\n\nyah, that pretty much tells me it is shawn's bug :)  but without more details, it is very hard to diagnose.  tohg should tell you which rcs revs are the offenders.  be sure to use a recent fromcvs however.\n\ncheers\n  simon\n"},{"id":"49471","messageId":"200708022221.13129.robin.rosenberg.lists@dewire.com","threadId":"9333","inReplyTo":"46B2309E.3060804@fs.ei.tum.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Robin Rosenberg","fromEmail":"robin.rosenberg.lists@dewire.com","sentAt":"2007-08-02T20:21:12Z","receivedAt":"2007-08-02T20:21:12Z","isPatch":false,"sender":{"key":"robin.rosenberg@dewire.com","avatar":"https://avatars.githubusercontent.com/u/46357?v=4"},"body":"torsdag 02 augusti 2007 skrev Simon 'corecode' Schubert:\n> Steffen Prohaska wrote:\n> > I remember that togit reported a broken pipe. My feeling was\n> > that git-fastimport aborted, which may be reason why tohg\n> > worked better. I didn't try to understand more details. I never\n> > read ruby code before and it was already a challenge for me to\n> > get everything up and running (rcs, rbtree).\n> \n> yah, that pretty much tells me it is shawn's bug :)  but without more \ndetails, it is very hard to diagnose.  tohg should tell you which rcs revs \nare the offenders.  be sure to use a recent fromcvs however.\n\nIf the bug is still unfixed and you haven't been able to diagnose for lack of \nrepos, you could try the Eclipse CVS repo.\n\nWhen I converted the Eclipse source to git I had a problem converting the \nwhole repo, i.e. fastimport died. The conversion died so I excluded some \nlarge parts that were effectively forks and some websites. \n\n-- robin\n"},{"id":"49472","messageId":"46B23F0E.2030304@tigris.org","threadId":"9333","inReplyTo":"200708022221.13129.robin.rosenberg.lists-RgPrefM1rjDQT0dZR+AlfA@public.gmane.org","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Lübbe Onken","fromEmail":"luebbe-jqhnx1hy4dsdnm+yrofe0a@public.gmane.org","sentAt":"2007-08-02T20:31:10Z","receivedAt":"2007-08-02T20:31:10Z","isPatch":false,"sender":{"key":"luebbe-jqhnx1hy4dsdnm+yrofe0a@public.gmane.org","avatar":null},"body":"Hi Folks,\n\nI guess that the initial poster sent this message to the TortoiseSVN\nusers list only by mistake, because the subject has nothing at all to do\nwith TortoiseSVN.\n\nCould you please be so kind and remove the TortoiseSVN users list from\nfuture replies to this thread?\n\nthanks\n-Lübbe\n"},{"id":"49474","messageId":"46B23F57.5050102@tigris.org","threadId":"9333","inReplyTo":"200708022221.13129.robin.rosenberg.lists@dewire.com","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Lübbe Onken","fromEmail":"luebbe@tigris.org","sentAt":"2007-08-02T20:32:23Z","receivedAt":"2007-08-02T20:32:23Z","isPatch":false,"sender":{"key":"luebbe@tigris.org","avatar":null},"body":"Hi Folks,\n\nI guess that the initial poster sent this message to the TortoiseSVN\nusers list only by mistake, because the subject has nothing at all to do\nwith TortoiseSVN.\n\nCould you please be so kind and remove the TortoiseSVN users list from\nfuture replies to this thread?\n\nthanks\n-Lübbe\n"},{"id":"49473","messageId":"46B23FB7.2040600@tigris.org","threadId":"9333","inReplyTo":"200708022221.13129.robin.rosenberg.lists@dewire.com","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Lübbe Onken","fromEmail":"l.onken@gmail.com","sentAt":"2007-08-02T20:33:59Z","receivedAt":"2007-08-02T20:33:59Z","isPatch":false,"sender":{"key":"l.onken@gmail.com","avatar":null},"body":"Hi Folks,\n\nI guess that the initial poster sent this message to the TortoiseSVN\nusers list only by mistake, because the subject has nothing at all to do\nwith TortoiseSVN.\n\nCould you please be so kind and remove the TortoiseSVN users list from\nfuture replies to this thread?\n\nthanks\n-Lübbe\n"},{"id":"49476","messageId":"alpine.LFD.0.999.0708021340450.8184@woody.linux-foundation.org","threadId":"9333","inReplyTo":"65F1862F-4DF2-4A52-9FD5-20802AEACDAB@zib.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2007-08-02T20:43:33Z","receivedAt":"2007-08-02T20:43:33Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 2 Aug 2007, Steffen Prohaska wrote:\n> \n> Right now, I'd prefer the import by parsecvs because of the\n> simpler history. However, I don't know if I loose history\n> information by doing so. I'd start by a run of cvs2svn to validate\n> the overall structure of the CVS repository.\n\nWell, once imported, you could just go through the branches and tags, and \njust delete the ones you consider uninteresting, and then do a \"git gc\".\n\nYou'd want to re-pack after a fast-import anyway (regardless of the source \nof the fast-import input), so maybe cvs2svn ends up giving you a bit \nunnecessary info, but it should be easy enough to get rid of \nafter-the-fact.\n\n\t\tLinus\n"},{"id":"49498","messageId":"6715F560-FE69-4F15-8C5F-B5B6071D97ED@zib.de","threadId":"9333","inReplyTo":"46B2309E.3060804@fs.ei.tum.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Steffen Prohaska","fromEmail":"prohaska@zib.de","sentAt":"2007-08-02T22:02:02Z","receivedAt":"2007-08-02T22:02:02Z","isPatch":false,"sender":{"key":"prohaska@zib.de","avatar":"https://avatars.githubusercontent.com/u/217580?v=4"},"body":"Simon,\n\nOn Aug 2, 2007, at 9:29 PM, Simon 'corecode' Schubert wrote:\n\n> Steffen Prohaska wrote:\n>> I remember that togit reported a broken pipe. My feeling was\n>> that git-fastimport aborted, which may be reason why tohg\n>> worked better. I didn't try to understand more details. I never\n>> read ruby code before and it was already a challenge for me to\n>> get everything up and running (rcs, rbtree).\n>\n> yah, that pretty much tells me it is shawn's bug :)  but without  \n> more details, it is very hard to diagnose.\n\nI tried again. Interestingly now togit works but tohg still fails.\n\ntogit starts with reporting\n\nfatal: Not a valid object name\n\nas the first line. But besides that it seems to work fine. What\nconcerns me a bit is that the last line togit reports is\n\ncommitting set 18100/18173\n\nI'd expect it should report 18173/18173.\nThe rest are git-fast-import statistics.\n\nBTW, togit creates much more complex branching patterns than cvs2svn\ndoes. The attached file branching.png displays a small view of a\nbranching pattern that extends downwards over a couple of screens.\nI checked the cvs2svn history again. It doesn't contain anything\nof similar complexity.\n\n\n> tohg should tell you which rcs revs are the offenders.  be sure to  \n> use a recent fromcvs however.\n\ntohg fails (on the same repo that togit imported) with the\nfollowing error\n\nTraceback (most recent call last):\n   File \"./tohg.py\", line 102, in <module>\n     destrepo.dispatch()\n   File \"./tohg.py\", line 98, in dispatch\n     func(*l[1:])\n   File \"./tohg.py\", line 78, in cmd_commit\n     extra = {'branch': branch})\n   File \"/sw/lib/python2.5/site-packages/mercurial/localrepo.py\",  \nline 736, in commit\n     mn = self.manifest.add(m1, tr, linkrev, c1[0], c2[0], (new,  \nremove))\n   File \"/sw/lib/python2.5/site-packages/mercurial/manifest.py\", line  \n191, in add\n     _(\"failed to remove %s from manifest\") % f)\nAssertionError: failed to remove X/Y.cpp from manifest\ntransaction abort!\nrollback completed\n./tohg.rb:200:in `readline': End of file reached while handling set  \n[core/X/Y.cpp,v:1.19,core/X/Z.cpp,v:1.22,core/X/Attic/W,v:1.12]  \n(EOFError)\n         from ./tohg.rb:200:in `_commit'\n         from ./tohg.rb:154:in `commit'\n         from ./fromcvs.rb:894:in `commit'\n         from ./fromcvs.rb:965:in `commit_sets'\n         from ./tohg.rb:228\n\n\nThe versions I used are listed below. I adjusted tohg a bit to use  \npython 2.5\ninstalled by fink. I'm working on Mac OS X.\n\n$ cd fromcvs\n$ hg tip\nchangeset:   103:cccdab84e9e5\ntag:         tip\nuser:        Simon 'corecode' Schubert <corecode@fs.ei.tum.de>\ndate:        Mon Jul 16 23:49:52 2007 +0200\nsummary:     Add error handling on committing sets.\n$ hg diff\ndiff -r cccdab84e9e5 tohg.rb\n--- a/tohg.rb   Mon Jul 16 23:49:52 2007 +0200\n+++ b/tohg.rb   Fri Jul 20 17:06:30 2007 +0200\n@@ -60,7 +60,7 @@ class HGDestRepo\n      @status = status\n      @outs, @ins = \\\n-      Open2.popen2('python', File.join(File.dirname($0), 'tohg.py'),  \nhgroot)\n+      Open2.popen2('python2.5', File.join(File.dirname($0),  \n'tohg.py'), hgroot)\n      @last_date = Time.at(@ins.readline.strip.to_i)\n      @branches = {}\n      while l = @ins.readline do\n\n\n$ cd rcsparse\n$ hg tip\nchangeset:   37:e871e108f2e4\ntag:         tip\nuser:        Simon 'corecode' Schubert <corecode@fs.ei.tum.de>\ndate:        Sun Feb 18 15:46:29 2007 +0100\nsummary:     Return revision date in GMT, like RCS/CVS uses everywhere.\n\nrbtree-0.2.0.tar.gz\n\nruby 1.8.2 (2004-12-25) [universal-darwin8.0]\n\n$ cd git\n$ git describe master\nv1.5.3-rc3-120-g68d4229\n\n$ hg --version\nMercurial Distributed SCM (version 0.9.3)\n\nCopyright (C) 2005, 2006 Matt Mackall <mpm@selenic.com>\nThis is free software; see the source for copying conditions. There  \nis NO\nwarranty; not even for MERCHANTABILITY or FITNESS FOR A PARTICULAR  \nPURPOSE.\n\n$ /sw/bin/python2.5 --version\nPython 2.5.1\n\n\nHope this helps.\n\n\tSteffen\n\n\n\n\n"},{"id":"49506","messageId":"46B25FC3.6000205@fs.ei.tum.de","threadId":"9333","inReplyTo":"6715F560-FE69-4F15-8C5F-B5B6071D97ED@zib.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-08-02T22:50:43Z","receivedAt":"2007-08-02T22:50:43Z","isPatch":false,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"Steffen Prohaska wrote:\n>> yah, that pretty much tells me it is shawn's bug :)  but without more \n>> details, it is very hard to diagnose.\n> \n> I tried again. Interestingly now togit works but tohg still fails.\n> \n> togit starts with reporting\n> \n> fatal: Not a valid object name\n\nthat's fine.\n\n> as the first line. But besides that it seems to work fine. What\n> concerns me a bit is that the last line togit reports is\n> \n> committing set 18100/18173\n> \n> I'd expect it should report 18173/18173.\n\nthat's fine as well.  You only saw multiples of 100, but you didn't consider it would skip the itermediate ones, right? :)\n\n> BTW, togit creates much more complex branching patterns than cvs2svn\n> does. The attached file branching.png displays a small view of a\n> branching pattern that extends downwards over a couple of screens.\n> I checked the cvs2svn history again. It doesn't contain anything\n> of similar complexity.\n\nhaha yea, there is still some issue with duplicate branch names and the branchpoint.  if it doesn't get the branch right, it will always \"pull\" files from the parent branch.\n\ndid you do some manual RCS file copying or manual branch name changing of individual files?  this could be the reason.  I still have to find a simple repo to reproduce this.\n\n> tohg fails (on the same repo that togit imported) with the\n> following error\n[..]\n> AssertionError: failed to remove X/Y.cpp from manifest\n\nThis is a mercurial 0.9.3 error, as far as I can tell from the reports.  This never occured here, and nobody reporting to me could ever reproduce this problem to pinpoint it.\n\ncheers\n  simon\n\n-- \nServe - BSD     +++  RENT this banner advert  +++    ASCII Ribbon   /\"\\\nWork - Mac      +++  space for low €€€ NOW!1  +++      Campaign     \\ /\nParty Enjoy Relax   |   http://dragonflybsd.org      Against  HTML   \\\nDude 2c 2 the max   !   http://golden-apple.biz       Mail + News   / \\\n"},{"id":"49510","messageId":"46a038f90708021608o21480074ybcfada767afc7b04@mail.gmail.com","threadId":"9333","inReplyTo":"46B1F96B.7050107@alum.mit.edu","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-08-02T23:08:25Z","receivedAt":"2007-08-02T23:08:25Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 8/3/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> cvsps is not a conversion tool at all, though it is used by other\n> conversion tools to generate the changesets.  It appears (I hope I am\n> not misinterpreting things) to emphasize speed and incremental\n> operation, for example attempting to make changesets consistent from one\n> run to the next, even if the CVS repository has been changed prudently\n> between runs.  cvsps does not appear to attempt to create atomic branch\n> and tag creation commits or handle CVS's special vendorbranch behavior.\n>  cvsps operates via the CVS protocol; you don't need filesystem access\n> to the CVS repository.\n\n100% in agreement. And though I can't claim to be happy with cvsps, in\nmany scenarios it is mighty useful, in spite of its significant warts.\n The \"does incrementals\" is hugely important these days, as lots of\npeople use git to run \"vendor branches\" of upstream projects that use\nCVS.\n\nTo me, that's *the* killer-app feature of git. Of course, others see\ndifferent aspects of git as their deal-maker. But I'm sure I'm not\nalone on this. Surely enough, others have written git-svn which\naccomplishes this and more for those tracking SVN upstreams.\n\nIs there any way we can run tweak cvs2svn to run incrementals, even if\nnot as fast as cvsps/git-cvsimport? The \"do it remotely\" part can be\nworked around in most cases.\n\ncheers,\n\n\nmartin\n"},{"id":"49512","messageId":"46B26686.3010002@alum.mit.edu","threadId":"9333","inReplyTo":"alpine.LFD.0.999.0708021340450.8184@woody.linux-foundation.org","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-08-02T23:19:34Z","receivedAt":"2007-08-02T23:19:34Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Linus Torvalds wrote:\n> On Thu, 2 Aug 2007, Steffen Prohaska wrote:\n>> Right now, I'd prefer the import by parsecvs because of the\n>> simpler history. However, I don't know if I loose history\n>> information by doing so. I'd start by a run of cvs2svn to validate\n>> the overall structure of the CVS repository.\n> \n> Well, once imported, you could just go through the branches and tags, and \n> just delete the ones you consider uninteresting, and then do a \"git gc\".\n> \n> You'd want to re-pack after a fast-import anyway (regardless of the source \n> of the fast-import input), so maybe cvs2svn ends up giving you a bit \n> unnecessary info, but it should be easy enough to get rid of \n> after-the-fact.\n\nThe real goal is to get cvs2svn to include the useful information and\nexclude the rest. :-)\n\nI definitely want to address the problem of the helper branches used to\ncreate tags.  This problem has has two aspects:\n\n1. The helper branches should be deleted after the tag has been defined.\n I simply couldn't figure out how to do this using git-fast-import, and\ngit-fast-import complained when I tried to use a branch called\n\"TAG_FIXUP\" without the \"refs/head/\" prefix.\n\n2. The helper branch is not needed at all if an existing revision has\nexactly the same contents as needed on the tag.  This requires cvs2svn\nto keep a record of which files exist in the complete file tree on every\nbranch at every revision (which it can already do, though it is\nexpensive), and also to give it the smarts to choose the optimal tag\npoint (which it already does, except that it currently doesn't penalize\nsources that require files to be deleted before making the tag).\n\n\nIf the problem is lots of seemingly-unnecessary merges involving a\nvendor branch, then it is time for me or some other volunteer to add the\noptimization of allowing branches to be grafted from the vendor branch\nto trunk.  I know of the problem and have a good idea how to implement\nit; it is just a matter of finding the time to get it done.\n\n\nIf the problem is unlabeled branches that can't be excluded (because\nother branches or tags depend on them), then the real problem is that it\nis not known which unlabeled branches in individual files correspond to\nthe same project-wide conceptual branch.  I have considered two\npossibilities to improve this situation:\n\n1. Allow unlabeled -- indeed any -- branches to be discarded even if\nother branches or tags depend on them.  This could be done by\nincorporating the content of the source revision (i.e., the revision on\nthe unlabeled branch that is going to be discarded) into the zeroth\nrevision of the daughter branch, then grafting the daughter onto the\nbranch from which the unlabeled branch sprouted.\n\n2. Rename the unlabeled branches by figuring out which unlabeled branch\nin fileA corresponds to which unlabeled branch in fileB, fileC, etc.\nThis would involve a tricky bit of matching file-wise dependency trees\nonto one another to unify unlabeled branch labels, keeping in mind that:\n\n  - The trees have other differences as well.\n  - The unlabeled branch does not necessarily occur in every file.\n  - There may be multiple unlabeled branches per file.\n\nMichael\n"},{"id":"49516","messageId":"46B26ACD.2080405@alum.mit.edu","threadId":"9333","inReplyTo":"EDE86758-FFD0-4CED-A2C9-033FA13DD3B6@zib.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-08-02T23:37:49Z","receivedAt":"2007-08-02T23:37:49Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Steffen Prohaska wrote:\n> On Aug 2, 2007, at 7:35 PM, Simon 'corecode' Schubert wrote:\n>> Steffen Prohaska wrote:\n>>>   fromcvs's togit surrendered during the import.\n>>>   fromcvs's tohg accepted more of the history, but finally\n>>>     surrendered as well.\n>>\n>> Which repo is it you are converting?  Is this available somewhere?\n> \n> Unfortunately not, the content is a proprietary software package.\n> \n>> I'd appreciate any reports concerning \"surrenders\" of fromcvs. \n>> [...]\n> \n> Some time passed since I did the tests. I had no time to do a\n> detailed investigation then. I'll have more time now and will\n> prepare a bug report, which is not easy because I can't sent you\n> the cvs repo, sorry.\n\nI wrote a couple of scripts for dealing with just this situation for\ncvs2svn bug reports, but they should also work for you, and I highly\nrecommend them.  Both scripts are included in the cvs2svn source tree:\n\n1. contrib/destroy_repository.py [1] -- strips almost all of the\ninformation out of a CVS repository, including author names, log\nmessages, and file contents (but not file names, commit dates, or\nbranch/tag names).  Most bugs are not affected by the omission of such\ndata.  Use of this script has the effect of deleting most information\nthat might be considered proprietary and also shrinking the size of the\ntest case considerably.  Use of this script is described in the script\ncomments itself and also in [2].\n\n2. contrib/shrink_test_case.py [2] -- you provide the script with a\ncommand that should \"exit 0\" if the bug you are looking for still\nexists.  It does a kind of \"binary search\" through CVS repository space,\niteratively attempting to delete a chunk of the CVS repository, running\nthe test command, then (depending on whether the test succeeded) either\nreverting or making permanent the deletion.  It can boil most test cases\ndown to just 1-3 files (though presumably not if the \"problem\" is a\n23-way merge).  The things that it will try to delete are:\n\n  - Entire directories and groups of directories\n  - Entire files and groups of files\n  - Branches within individual files\n  - Tags within individual files\n\nIt does this in a somewhat optimal way, trying to minimize the number of\ntimes that the test has to be run.  This script is documented in its own\ncomments and also in [4].\n\nMichael\n\n[1]\nhttp://cvs2svn.tigris.org/svn/cvs2svn/trunk/contrib/destroy_repository.py\n[2] http://cvs2svn.tigris.org/faq.html#reportingbugs\n[3] http://cvs2svn.tigris.org/svn/cvs2svn/trunk/contrib/shrink_test_case.py\n[4] http://cvs2svn.tigris.org/faq.html#testcase\n"},{"id":"49517","messageId":"9e4733910708021644q6eba0e78gc2c6bcfba4816012@mail.gmail.com","threadId":"9333","inReplyTo":"f8r09t$qdg$1@sea.gmane.org","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-02T23:44:37Z","receivedAt":"2007-08-02T23:44:37Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/1/07, Jakub Narebski <jnareb@gmail.com> wrote:\n> Michael Haggerty wrote:\n>\n> > I am the maintainer of cvs2svn[1], which is a program for one-time\n> > conversions from CVS to Subversion. cvs2svn is very robust against the\n> > many peculiarities of CVS and can convert just about every CVS\n> > repository we have ever seen.\n> >\n> > I've been working on a cvs2svn output pass that writes the converted CVS\n> > repository directly into git rather than Subversion. The code runs now\n> > with at least one repository from our test suite of nasty CVS repositories.\n>\n> Have you contacted Jon Smirl about his unpublished work on cvs2git,\n> cvs2svn based CVS to Git converter?\n\nMy converter was derived from Michael's cvs2svn code. The bulk of my\nwork was converting cvs2svn to output in a format that git-fastimport\ncould consume. This was all rather straight forward and there was\nnothing really interesting in the code.\n\nWhat it exposed were fundamental issues about the technical\ncomplexities of trying to reconstruct a change set history from CVS\nwhich didn't record all of the needed info.  I was never able to\nconstruct a satisfactory git representation of the Mozilla CVS\nrepository.  Michael has had a long time to work on the change set\ndetection code and he's probably added some new strategies.\n\nMy code did include a CVS file parser for extracting all the revisions\nfrom the file in a single pass. Doing that is a major performance\nbenefit.  I believe I posted the code to the cvs2svn mailing list. It\nwas about 200 lines of code. Forking off cvs a million times to\nextract the revisions takes days to run.\n\nSame goes for forking git a million times.git-fastimport uses a pipe\nto cvs2svn to avoid forking. git-fastimport also uses a technique from\nthe database world for bulk import, it imports everything without\nindexing it. Indexing is done after the import finishes.\n\nBetween parsing the CVS files internally and Shawn's git-fastimport,\nit was possible to import Mozilla CVS (2.4G) in about 2 hours and\ngenerate a 450MB pack file. You need 3GB of RAM to do this - if swap\nhappens the process will take weeks to finish.\n\n> Quote from InterfacesFrontendsAndTools page on GIT wiki[1]:\n>\n>   cvs2git is the unofficial name of Jon Smirl's modifications to cvs2svn.\n>   These modifications allow cvs2svn to generate a data stream which is\n>   consumed by Shawn Pearce's git-fast-import (now included in git.git).\n>   git-fast-import converts its input stream directly into a Git .pack file,\n>   minimizing the amount of IO required on large imports.\n>\n>   Jon Smirl stopped working on cvs2git[2] because first, Mozilla (which was\n>   main target of his work) decided that to not to move to git, and second\n>   because of troubles with cvs2svn architecture[*] (which it is based on).\n>   Jon Smirl has posted his impressions on working on CVS importer in\n>   \"Some tips for doing a CVS importer\" thread[3].\n>\n> References:\n> -----------\n> [1] http://git.or.cz/gitwiki/InterfacesFrontendsAndTools#head-23858c2cde0cef60443d8e73e6829a95f8e191ef\n> [2] http://msgid.gmane.org/9e4733910611190940y147992b8mbdfac5a51f42e0fe@mail.gmail.com\n> [3] http://marc.theaimsgroup.com/?t=116405956000001&r=1&w=2\n>\n> Footnotes:\n> ----------\n> [*] If I remember correctly authors of cvs2svn were talking about separating\n> the code dealing with disentangling CVS repository structure from the part\n> translating it into Subversion repository (with its quirks), and the part\n> generating Subversion repository.\n>\n> --\n> Jakub Narebski\n> Warsaw, Poland\n> ShadeHawk on #git\n>\n>\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"49519","messageId":"46B26DA8.6010102@alum.mit.edu","threadId":"9333","inReplyTo":"46B25FC3.6000205@fs.ei.tum.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-08-02T23:50:00Z","receivedAt":"2007-08-02T23:50:00Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Simon 'corecode' Schubert wrote:\n> Steffen Prohaska wrote:\n>> BTW, togit creates much more complex branching patterns than cvs2svn\n>> does. The attached file branching.png displays a small view of a\n>> branching pattern that extends downwards over a couple of screens.\n>> I checked the cvs2svn history again. It doesn't contain anything\n>> of similar complexity.\n> \n> haha yea, there is still some issue with duplicate branch names and the\n> branchpoint.  if it doesn't get the branch right, it will always \"pull\"\n> files from the parent branch.\n\nThis sounds very much like the problem reported by Daniel Jacobowitz\n[1].  The problem is that if you create a branch A on a file, then\ncreate branch B from branch A before making a commit on branch A, then\nCVS doesn't record that branch A was the source of branch B.  (It treats\nB as if it sprouted directly from the revision that was the *source* of\nbranch A.)  The same problem exists if \"B\" is a tag.\n\nThe only way to determine the correct branch hierarchy is to consider\nthe branch hierarchy of multiple files at the same time.\n\ncvs2svn 2.0 includes code to choose a \"preferred parent\" of each branch\nand try to use that parent for every file that is on the branch.  It\nhelps simplify branch creation quite a bit.  The main limitation is that\nit still doesn't consider the revision copied back to trunk from a\nvendor branch as the possible parent of a branch whose nominal source\nwas on the vendor branch (a limitation that has come up elsewhere in\nthis thread).\n\nMichael\n\n[1] http://cvs2svn.tigris.org/servlets/ReadMsg?list=dev&msgNo=1441\n"},{"id":"49521","messageId":"9e4733910708021655p34a72428gd0bd33a830faf127@mail.gmail.com","threadId":"9333","inReplyTo":"65F1862F-4DF2-4A52-9FD5-20802AEACDAB@zib.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-02T23:55:28Z","receivedAt":"2007-08-02T23:55:28Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/2/07, Steffen Prohaska <prohaska@zib.de> wrote:\n> Right now, I'd prefer the import by parsecvs because of the\n> simpler history. However, I don't know if I loose history\n> information by doing so. I'd start by a run of cvs2svn to validate\n> the overall structure of the CVS repository. Dealing with corruption\n> in the CVS repository seems to be superior in cvs2svn. It reports\n> errors when parsecvs just crashes.\n\nParsecvs silently throws away things that confuse it. cvs2svn is much\nmore careful about not losing track of anything. For example parsecvs\nis unable to process Mozilla CVS and cvs2svn can. The branching in\nMozilla CVS is too complex for parsecvs to handle.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"49522","messageId":"9e4733910708021659y6e9bb7ddk58817b4de3df26a0@mail.gmail.com","threadId":"9333","inReplyTo":"46B2132D.7090304@alum.mit.edu","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-02T23:59:41Z","receivedAt":"2007-08-02T23:59:41Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/2/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> Branches with names like \"unlabeled-1.1.1\" come from CVS branches for\n> which the revisions are still contained in the RCS files but for which\n> the branch name has been deleted.  These wreak havoc on cvs2svn's\n> attempt to find simple branch sources and cause a proliferation of\n> basically useless branches.  The main problem is that cvs2svn does not\n> attempt to figure out that \"unlabeled-1.2.4\" in one file might be the\n> same as \"unlabeled-1.2.6\" in another etc.\n\nI seem to recall discussing an algorithm  to fix this on the cvs2svn\nmailing list. There was a somewhat simple way to correlate the\n\"unlabeled-1.2.4\" in one file might be the same as \"unlabeled-1.2.6\"\nproblem.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"49536","messageId":"20070803030705.GK20052@spearce.org","threadId":"9333","inReplyTo":"46B2309E.3060804@fs.ei.tum.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-08-03T03:07:05Z","receivedAt":"2007-08-03T03:07:05Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Simon 'corecode' Schubert <corecode@fs.ei.tum.de> wrote:\n> Steffen Prohaska wrote:\n> >I remember that togit reported a broken pipe. My feeling was\n> >that git-fastimport aborted, which may be reason why tohg\n> >worked better.\n> \n> yah, that pretty much tells me it is shawn's bug :)  but without more \n> details, it is very hard to diagnose.  tohg should tell you which rcs revs \n> are the offenders.  be sure to use a recent fromcvs however.\n\nTonight I'm going to try and add crash dump reporting to fast-import.\nOnce that's in it should make debugging some of these failed imports\neasier, as we'll be able to see the immediate commands leading up\nto the crash and the internal state of fast-import when it barfed.\n\nOf course one needs to locate an ugly repository and run on it...\n\n-- \nShawn.\n"},{"id":"49537","messageId":"20070803031239.GL20052@spearce.org","threadId":"9333","inReplyTo":"46B26686.3010002@alum.mit.edu","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-08-03T03:12:39Z","receivedAt":"2007-08-03T03:12:39Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> 1. The helper branches should be deleted after the tag has been defined.\n>  I simply couldn't figure out how to do this using git-fast-import, and\n> git-fast-import complained when I tried to use a branch called\n> \"TAG_FIXUP\" without the \"refs/head/\" prefix.\n\nTwo issues there:\n\n* Deleting branches:\n\n  I currently don't support this in fast-import, but I'll add support\n  for it.  Its actually pretty simple to tell it to drop a branch,\n  especially if the dang thing doesn't actually exist in the git\n  repository yet (because its only in-memory).\n\n* Creating a branch without refs/heads/ prefix:\n\n  This is a bug.  I had good intentions by trying to verify the\n  name was one that didn't contain special reserved characters,\n  but I wound up also requiring you to create branches only in the\n  refs/heads/ namespace.  That was not what I wanted to do.  I'm\n  patching it tonight.\n\n-- \nShawn.\n"},{"id":"49538","messageId":"Pine.LNX.4.64.0708030454200.14781@racer.site","threadId":"9333","inReplyTo":"46a038f90708021608o21480074ybcfada767afc7b04@mail.gmail.com","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-08-03T04:03:10Z","receivedAt":"2007-08-03T04:03:10Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 3 Aug 2007, Martin Langhoff wrote:\n\n> On 8/3/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> > cvsps is not a conversion tool at all, though it is used by other\n> > conversion tools to generate the changesets.  It appears (I hope I am\n> > not misinterpreting things) to emphasize speed and incremental\n> > operation, for example attempting to make changesets consistent from one\n> > run to the next, even if the CVS repository has been changed prudently\n> > between runs.  cvsps does not appear to attempt to create atomic branch\n> > and tag creation commits or handle CVS's special vendorbranch behavior.\n> >  cvsps operates via the CVS protocol; you don't need filesystem access\n> > to the CVS repository.\n> \n> 100% in agreement. And though I can't claim to be happy with cvsps, in\n> many scenarios it is mighty useful, in spite of its significant warts.\n>  The \"does incrementals\" is hugely important these days, as lots of\n> people use git to run \"vendor branches\" of upstream projects that use\n> CVS.\n\nMe too: 100% agreement.  A couple of people seem to be content to proclaim \nthat their incomplete solutions are better, but in the end of the day, \nthey are as bad as the programs they purport to replace: incomplete.\n\nFor the moment, I help myself with tracking the different branches \nindividually, but there, really, git-cvsimport is as good as the other \n\"solutions\", with the further advantage that they are actually hackable, \nand not closed to everybody outside a very small community.\n\nSo I look forward to testing cvs2svn(git-branch) this weekend.\n\nCiao,\nDscho\n"},{"id":"49549","messageId":"2FEBD7A5-9932-4636-955D-F7E258F8E56E@zib.de","threadId":"9333","inReplyTo":"Pine.LNX.4.64.0708030454200.14781@racer.site","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Steffen Prohaska","fromEmail":"prohaska@zib.de","sentAt":"2007-08-03T06:48:13Z","receivedAt":"2007-08-03T06:48:13Z","isPatch":false,"sender":{"key":"prohaska@zib.de","avatar":"https://avatars.githubusercontent.com/u/217580?v=4"},"body":"\nOn Aug 3, 2007, at 6:03 AM, Johannes Schindelin wrote:\n\n> On Fri, 3 Aug 2007, Martin Langhoff wrote:\n>\n>> On 8/3/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>>> cvsps is not a conversion tool at all, though it is used by other\n>>> conversion tools to generate the changesets.  It appears (I hope  \n>>> I am\n>>> not misinterpreting things) to emphasize speed and incremental\n>>> operation, for example attempting to make changesets consistent  \n>>> from one\n>>> run to the next, even if the CVS repository has been changed  \n>>> prudently\n>>> between runs.  cvsps does not appear to attempt to create atomic  \n>>> branch\n>>> and tag creation commits or handle CVS's special vendorbranch  \n>>> behavior.\n>>>  cvsps operates via the CVS protocol; you don't need filesystem  \n>>> access\n>>> to the CVS repository.\n>>\n>> 100% in agreement. And though I can't claim to be happy with  \n>> cvsps, in\n>> many scenarios it is mighty useful, in spite of its significant  \n>> warts.\n>>  The \"does incrementals\" is hugely important these days, as lots of\n>> people use git to run \"vendor branches\" of upstream projects that use\n>> CVS.\n>\n> Me too: 100% agreement.  A couple of people seem to be content to  \n> proclaim\n> that their incomplete solutions are better, but in the end of the day,\n> they are as bad as the programs they purport to replace: incomplete.\n>\n> For the moment, I help myself with tracking the different branches\n> individually, but there, really, git-cvsimport is as good as the other\n> \"solutions\", with the further advantage that they are actually  \n> hackable,\n> and not closed to everybody outside a very small community.\n\nI just want to add a warning. You should be suspicious of branched  \nimported\nusing git-cvsimport (which is based on cvsps). If the time the branch is\ncreated differs from the time of the first commit to the branch git- \ncvsimport\nmay get the branching point wrong. This introduces a race condition.  \nSomeone\nmay have committed changes to a file that is later changed on the  \nbranch. At\nthat point the history of the imported branch is broken and git reports\n_wrong_ changesets.\n\nI ran into this issue and abandoned the use of git-cvsimport. It's  \ntoo dangerous\nfor me. The testcase in [1] illustrates the problem. I still strongly  \nbelieve\nthe warning should be stated in *BOLD* in the documentation.\n\nI'm not saying git-cvsimport is useless. But you should be suspicious  \nabout\nthe result of the import, especially if you plan to rely on  \nchangesets derived\nfrom the imported repo, for example if you plan to do cherry-picking  \nor merging\nin git; or if you plan to blame people for their stupid changes based  \non what\nyou see in gitk (almost happend to me ;).\n\n\tSteffen\n\n[1] http://marc.info/?l=git&m=118260312708709&w=2\n"},{"id":"49553","messageId":"B4109CF3-A849-40C2-BDFD-DD7CA5231FA8@zib.de","threadId":"9333","inReplyTo":"46a038f90708021608o21480074ybcfada767afc7b04@mail.gmail.com","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Steffen Prohaska","fromEmail":"prohaska@zib.de","sentAt":"2007-08-03T07:10:35Z","receivedAt":"2007-08-03T07:10:35Z","isPatch":false,"sender":{"key":"prohaska@zib.de","avatar":"https://avatars.githubusercontent.com/u/217580?v=4"},"body":"\nOn Aug 3, 2007, at 1:08 AM, Martin Langhoff wrote:\n\n> Is there any way we can run tweak cvs2svn to run incrementals, even if\n> not as fast as cvsps/git-cvsimport? The \"do it remotely\" part can be\n> worked around in most cases.\n\nWhat I currently do with parsecvs is to run complete imports again\non the repo. For 'normal' changes to cvs the old import can be fast\nforwarded to the new import. However, if you add or remove files or\ntweak revision in another abnormal way (cvs admin) this might fail.\n\nIn this case I manually search the last common commit and rebase\nnew commits to the old, already imported branch. I need to do this if\nI already publishes the imported branch. Otherwise I can as well just\nreset to the newly imported branch and rebase my work on top of it.\nSome careful validation (git diff-*) is included in my workflow.\n\nA complete run of parsecvs is fine for me because it is so fast. I run\ngit-filter-branch afterwards anyway to cleanup some commit messages\nand author information. This takes most of the time, because it spawns\noff tons of sub processes.\n\nI'd not recommend my approach for incremental imports every hour, but\nyou can run it every day (although I do less often). You only need to\nvalidate the final result (fast forward or not). The rest can be fully\nautomated by some shell scripting.\n\n\tSteffen\n"},{"id":"49562","messageId":"46B2E8F3.30301@alum.mit.edu","threadId":"9333","inReplyTo":"46a038f90708021608o21480074ybcfada767afc7b04@mail.gmail.com","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-08-03T08:36:03Z","receivedAt":"2007-08-03T08:36:03Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Martin Langhoff wrote:\n> Is there any way we can run tweak cvs2svn to run incrementals, even if\n> not as fast as cvsps/git-cvsimport? The \"do it remotely\" part can be\n> worked around in most cases.\n\nI don't see any fundamental reason why not, but I think it would be a\nsignificant amount of work.  There are two main issues:\n\n1. With CVS, it is possible to change things retroactively, such as\nchanging which version of a file is included in a tag, or adding a new\nfile to a tag, or changing whether a file is text vs. binary.  And many\npeople copy and/or rename files within the CVS repository itself (to get\naround CVS's inability to rename a file).  This makes it look like the\nfile has *always* existed under the new name and *never* existed under\nthe old name.  An incremental conversion tool would have to look\ncarefully for such changes and either handle them properly or complain\nloudly and abort.\n\n2. cvs2svn uses a lot of repository-wide information to make decisions\nabout how to group CVSItems into changesets, and a lot of these\ndecisions are based on heuristics.  Incremental conversion would require\nthat the decisions made in one cvs2svn run are recorded and treated as\nunalterable in subsequent runs.\n\nThis hasn't been a priority in the Subversion world, because, frankly,\nwhat reason would a person have to stick with CVS instead of switching\nto Subversion, given that (1) they are intentionally so similar in\nworkflow, an (2) there is no significant competition from other\ncentralized SCMs?  But of course until the distributed SCM playing field\nhas been thinned out a bit, people will probably be reluctant to commit\nto one or the other.\n\nI don't expect to have time to implement incremental conversions in\ncvs2svn in the near future.  (I'd much rather work on output back ends\nto other distributed SCMs.)  But if any volunteers step forward (hint,\nhint) I would be happy to help them get started and answer their\nquestions.  I think that cvs2svn is quite hackable now, so the learning\ncurve is hopefully much less frightening than when I started on the\nproject :-)\n\nMichael\n"},{"id":"49563","messageId":"46B2EA0B.9030805@fs.ei.tum.de","threadId":"9333","inReplyTo":"46B26DA8.6010102@alum.mit.edu","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Simon 'corecode' Schubert","fromEmail":"corecode@fs.ei.tum.de","sentAt":"2007-08-03T08:40:43Z","receivedAt":"2007-08-03T08:40:43Z","isPatch":false,"sender":{"key":"corecode@fs.ei.tum.de","avatar":"https://gravatar.com/avatar/eff9dbf0cdac0d1e6a6cd7ed0e50763edcb376b493b5253a35ff167918ad79e1?d=mp&s=160"},"body":"Michael Haggerty wrote:\n> Simon 'corecode' Schubert wrote:\n>> Steffen Prohaska wrote:\n>>> BTW, togit creates much more complex branching patterns than cvs2svn\n>>> does. The attached file branching.png displays a small view of a\n>>> branching pattern that extends downwards over a couple of screens.\n>>> I checked the cvs2svn history again. It doesn't contain anything\n>>> of similar complexity.\n>> haha yea, there is still some issue with duplicate branch names and the\n>> branchpoint.  if it doesn't get the branch right, it will always \"pull\"\n>> files from the parent branch.\n> \n> This sounds very much like the problem reported by Daniel Jacobowitz\n> [1].  The problem is that if you create a branch A on a file, then\n> create branch B from branch A before making a commit on branch A, then\n> CVS doesn't record that branch A was the source of branch B.  (It treats\n> B as if it sprouted directly from the revision that was the *source* of\n> branch A.)  The same problem exists if \"B\" is a tag.\n\nI think I have covered this case quite well.  I believe \"my\" problem happens when there are files being copied manually within the repository and then branch names being changed (or just branch names being changed).  However, the name change just happens only on a subset of files and branches, so you wind up with a commit which is part of two branches.  Or something like that.  I really should have the time to investigate this.\n\nOne elementary problem with CVS is that you can assign two branch names to the same branch.  During conversion you need to choose one over the other.\n\ncheers\n  simon\n\n-- \nServe - BSD     +++  RENT this banner advert  +++    ASCII Ribbon   /\"\\\nWork - Mac      +++  space for low €€€ NOW!1  +++      Campaign     \\ /\nParty Enjoy Relax   |   http://dragonflybsd.org      Against  HTML   \\\nDude 2c 2 the max   !   http://golden-apple.biz       Mail + News   / \\\n"},{"id":"49600","messageId":"0BB549C6E74E24409FB20B3B1D1B6644029461C0@ATL1EX11.corp.etradegrp.com","threadId":"9333","inReplyTo":"46B2E8F3.30301@alum.mit.edu","subject":"RE: Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Patwardhan, Rajesh","fromEmail":"rajesh.patwardhan@etrade.com","sentAt":"2007-08-03T14:35:31Z","receivedAt":"2007-08-03T14:35:31Z","isPatch":false,"sender":{"key":"rajesh.patwardhan@etrade.com","avatar":null},"body":"\nHello Michael, \nI will explain a scenario (we are passing thru this right now) \n1) you have 10 years worth of cvs data.\n2) We want to move to svn. \n3) The repository move should be in such a way that the development does\nnot get hampered for any 1 work day.   \n4) We have atleast 4 major modules in cvs which takes about 30 - 40\nhours each for conversion currently.\n5) With increamental conversions we can do a few things ... \n\tA) Keep the downtime for hard cutoff minimal \n\tB) try out the svn move for other auxillary tools that are\nneeded by the SCM process. \n\tC) Do some meaningful testing and validation with simulated live\nmoves of changes from cvs to svn before the actual move on a day to day\nbasis. \n\nHopefuly this would substantiate the request \\ need for increamental\nmoves. Or if someone out there has a better suggestion for such\nscenario's please point me in the right direction. \n\nRegards,\nRajesh \n\n-----Original Message-----\nFrom: Michael Haggerty [mailto:mhagger@alum.mit.edu] \nSent: Friday, August 03, 2007 1:36 AM\nTo: Martin Langhoff\nCc: Guilhem Bonnefille; git@vger.kernel.org; users@cvs2svn.tigris.org\nSubject: Re: cvs2svn conversion directly to git ready for\nexperimentation\n\nMartin Langhoff wrote:\n> Is there any way we can run tweak cvs2svn to run incrementals, even if\n\n> not as fast as cvsps/git-cvsimport? The \"do it remotely\" part can be \n> worked around in most cases.\n\nI don't see any fundamental reason why not, but I think it would be a\nsignificant amount of work.  There are two main issues:\n\n1. With CVS, it is possible to change things retroactively, such as\nchanging which version of a file is included in a tag, or adding a new\nfile to a tag, or changing whether a file is text vs. binary.  And many\npeople copy and/or rename files within the CVS repository itself (to get\naround CVS's inability to rename a file).  This makes it look like the\nfile has *always* existed under the new name and *never* existed under\nthe old name.  An incremental conversion tool would have to look\ncarefully for such changes and either handle them properly or complain\nloudly and abort.\n\n2. cvs2svn uses a lot of repository-wide information to make decisions\nabout how to group CVSItems into changesets, and a lot of these\ndecisions are based on heuristics.  Incremental conversion would require\nthat the decisions made in one cvs2svn run are recorded and treated as\nunalterable in subsequent runs.\n\nThis hasn't been a priority in the Subversion world, because, frankly,\nwhat reason would a person have to stick with CVS instead of switching\nto Subversion, given that (1) they are intentionally so similar in\nworkflow, an (2) there is no significant competition from other\ncentralized SCMs?  But of course until the distributed SCM playing field\nhas been thinned out a bit, people will probably be reluctant to commit\nto one or the other.\n\nI don't expect to have time to implement incremental conversions in\ncvs2svn in the near future.  (I'd much rather work on output back ends\nto other distributed SCMs.)  But if any volunteers step forward (hint,\nhint) I would be happy to help them get started and answer their\nquestions.  I think that cvs2svn is quite hackable now, so the learning\ncurve is hopefully much less frightening than when I started on the\nproject :-)\n\nMichael\n\n---------------------------------------------------------------------\nTo unsubscribe, e-mail: users-unsubscribe@cvs2svn.tigris.org\nFor additional commands, e-mail: users-help@cvs2svn.tigris.org\n"},{"id":"49603","messageId":"9e4733910708030841r31175efg4ea4ea41e852ab2@mail.gmail.com","threadId":"9333","inReplyTo":"0BB549C6E74E24409FB20B3B1D1B6644029461C0@ATL1EX11.corp.etradegrp.com","subject":"Re: Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-03T15:41:02Z","receivedAt":"2007-08-03T15:41:02Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/3/07, Patwardhan, Rajesh <rajesh.patwardhan@etrade.com> wrote:\n>\n> Hello Michael,\n> I will explain a scenario (we are passing thru this right now)\n> 1) you have 10 years worth of cvs data.\n> 2) We want to move to svn.\n> 3) The repository move should be in such a way that the development does\n> not get hampered for any 1 work day.\n> 4) We have atleast 4 major modules in cvs which takes about 30 - 40\n> hours each for conversion currently.\n\nThere are known ways (that haven't been implemented) to get the 40 hr\nnumber down to 1/2 hour. Would that be a better approach than doing\nincremental imports?\n\n> 5) With increamental conversions we can do a few things ...\n>         A) Keep the downtime for hard cutoff minimal\n>         B) try out the svn move for other auxillary tools that are\n> needed by the SCM process.\n>         C) Do some meaningful testing and validation with simulated live\n> moves of changes from cvs to svn before the actual move on a day to day\n> basis.\n>\n> Hopefuly this would substantiate the request \\ need for increamental\n> moves. Or if someone out there has a better suggestion for such\n> scenario's please point me in the right direction.\n>\n> Regards,\n> Rajesh\n>\n> -----Original Message-----\n> From: Michael Haggerty [mailto:mhagger@alum.mit.edu]\n> Sent: Friday, August 03, 2007 1:36 AM\n> To: Martin Langhoff\n> Cc: Guilhem Bonnefille; git@vger.kernel.org; users@cvs2svn.tigris.org\n> Subject: Re: cvs2svn conversion directly to git ready for\n> experimentation\n>\n> Martin Langhoff wrote:\n> > Is there any way we can run tweak cvs2svn to run incrementals, even if\n>\n> > not as fast as cvsps/git-cvsimport? The \"do it remotely\" part can be\n> > worked around in most cases.\n>\n> I don't see any fundamental reason why not, but I think it would be a\n> significant amount of work.  There are two main issues:\n>\n> 1. With CVS, it is possible to change things retroactively, such as\n> changing which version of a file is included in a tag, or adding a new\n> file to a tag, or changing whether a file is text vs. binary.  And many\n> people copy and/or rename files within the CVS repository itself (to get\n> around CVS's inability to rename a file).  This makes it look like the\n> file has *always* existed under the new name and *never* existed under\n> the old name.  An incremental conversion tool would have to look\n> carefully for such changes and either handle them properly or complain\n> loudly and abort.\n>\n> 2. cvs2svn uses a lot of repository-wide information to make decisions\n> about how to group CVSItems into changesets, and a lot of these\n> decisions are based on heuristics.  Incremental conversion would require\n> that the decisions made in one cvs2svn run are recorded and treated as\n> unalterable in subsequent runs.\n>\n> This hasn't been a priority in the Subversion world, because, frankly,\n> what reason would a person have to stick with CVS instead of switching\n> to Subversion, given that (1) they are intentionally so similar in\n> workflow, an (2) there is no significant competition from other\n> centralized SCMs?  But of course until the distributed SCM playing field\n> has been thinned out a bit, people will probably be reluctant to commit\n> to one or the other.\n>\n> I don't expect to have time to implement incremental conversions in\n> cvs2svn in the near future.  (I'd much rather work on output back ends\n> to other distributed SCMs.)  But if any volunteers step forward (hint,\n> hint) I would be happy to help them get started and answer their\n> questions.  I think that cvs2svn is quite hackable now, so the learning\n> curve is hopefully much less frightening than when I started on the\n> project :-)\n>\n> Michael\n>\n> ---------------------------------------------------------------------\n> To unsubscribe, e-mail: users-unsubscribe@cvs2svn.tigris.org\n> For additional commands, e-mail: users-help@cvs2svn.tigris.org\n>\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"49605","messageId":"0BB549C6E74E24409FB20B3B1D1B664402946577@ATL1EX11.corp.etradegrp.com","threadId":"9333","inReplyTo":"9e4733910708030841r31175efg4ea4ea41e852ab2@mail.gmail.com","subject":"RE: Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Patwardhan, Rajesh","fromEmail":"rajesh.patwardhan@etrade.com","sentAt":"2007-08-03T16:42:05Z","receivedAt":"2007-08-03T16:42:05Z","isPatch":false,"sender":{"key":"rajesh.patwardhan@etrade.com","avatar":null},"body":"Thank you very much for the email. \nYes if the time for conversion can be brought down to 1/2 hour then it\nwould be really great. \nWe could do a automated cvs2svn everyday for testing and that way\nmaximum lag between cvs and test svn repo would be 1 day. \nPlease do let me know when available.\nRegards,\nRajesh \n\n-----Original Message-----\nFrom: Jon Smirl [mailto:jonsmirl@gmail.com] \nSent: Friday, August 03, 2007 8:41 AM\nTo: Patwardhan, Rajesh\nCc: Michael Haggerty; Martin Langhoff; Guilhem Bonnefille;\ngit@vger.kernel.org; users@cvs2svn.tigris.org\nSubject: Re: Re: cvs2svn conversion directly to git ready for\nexperimentation\n\nOn 8/3/07, Patwardhan, Rajesh <rajesh.patwardhan@etrade.com> wrote:\n>\n> Hello Michael,\n> I will explain a scenario (we are passing thru this right now)\n> 1) you have 10 years worth of cvs data.\n> 2) We want to move to svn.\n> 3) The repository move should be in such a way that the development \n> does not get hampered for any 1 work day.\n> 4) We have atleast 4 major modules in cvs which takes about 30 - 40 \n> hours each for conversion currently.\n\nThere are known ways (that haven't been implemented) to get the 40 hr\nnumber down to 1/2 hour. Would that be a better approach than doing\nincremental imports?\n\n> 5) With increamental conversions we can do a few things ...\n>         A) Keep the downtime for hard cutoff minimal\n>         B) try out the svn move for other auxillary tools that are \n> needed by the SCM process.\n>         C) Do some meaningful testing and validation with simulated \n> live moves of changes from cvs to svn before the actual move on a day \n> to day basis.\n>\n> Hopefuly this would substantiate the request \\ need for increamental \n> moves. Or if someone out there has a better suggestion for such \n> scenario's please point me in the right direction.\n>\n> Regards,\n> Rajesh\n>\n> -----Original Message-----\n> From: Michael Haggerty [mailto:mhagger@alum.mit.edu]\n> Sent: Friday, August 03, 2007 1:36 AM\n> To: Martin Langhoff\n> Cc: Guilhem Bonnefille; git@vger.kernel.org; users@cvs2svn.tigris.org\n> Subject: Re: cvs2svn conversion directly to git ready for \n> experimentation\n>\n> Martin Langhoff wrote:\n> > Is there any way we can run tweak cvs2svn to run incrementals, even \n> > if\n>\n> > not as fast as cvsps/git-cvsimport? The \"do it remotely\" part can be\n\n> > worked around in most cases.\n>\n> I don't see any fundamental reason why not, but I think it would be a \n> significant amount of work.  There are two main issues:\n>\n> 1. With CVS, it is possible to change things retroactively, such as \n> changing which version of a file is included in a tag, or adding a new\n\n> file to a tag, or changing whether a file is text vs. binary.  And \n> many people copy and/or rename files within the CVS repository itself \n> (to get around CVS's inability to rename a file).  This makes it look \n> like the file has *always* existed under the new name and *never* \n> existed under the old name.  An incremental conversion tool would have\n\n> to look carefully for such changes and either handle them properly or \n> complain loudly and abort.\n>\n> 2. cvs2svn uses a lot of repository-wide information to make decisions\n\n> about how to group CVSItems into changesets, and a lot of these \n> decisions are based on heuristics.  Incremental conversion would \n> require that the decisions made in one cvs2svn run are recorded and \n> treated as unalterable in subsequent runs.\n>\n> This hasn't been a priority in the Subversion world, because, frankly,\n\n> what reason would a person have to stick with CVS instead of switching\n\n> to Subversion, given that (1) they are intentionally so similar in \n> workflow, an (2) there is no significant competition from other \n> centralized SCMs?  But of course until the distributed SCM playing \n> field has been thinned out a bit, people will probably be reluctant to\n\n> commit to one or the other.\n>\n> I don't expect to have time to implement incremental conversions in \n> cvs2svn in the near future.  (I'd much rather work on output back ends\n\n> to other distributed SCMs.)  But if any volunteers step forward (hint,\n> hint) I would be happy to help them get started and answer their \n> questions.  I think that cvs2svn is quite hackable now, so the \n> learning curve is hopefully much less frightening than when I started \n> on the project :-)\n>\n> Michael\n>\n> ---------------------------------------------------------------------\n> To unsubscribe, e-mail: users-unsubscribe@cvs2svn.tigris.org\n> For additional commands, e-mail: users-help@cvs2svn.tigris.org\n>\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in the \n> body of a message to majordomo@vger.kernel.org More majordomo info at\n\n> http://vger.kernel.org/majordomo-info.html\n>\n\n\n--\nJon Smirl\njonsmirl@gmail.com\n"},{"id":"49615","messageId":"46B37ADB.8020103@alum.mit.edu","threadId":"9333","inReplyTo":"9e4733910708030841r31175efg4ea4ea41e852ab2@mail.gmail.com","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2007-08-03T18:58:35Z","receivedAt":"2007-08-03T18:58:35Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"[I set followup-to users@cvs2svn.tigris.org, since this has nothing to\ndo with git.]\n\nJon Smirl wrote:\n> On 8/3/07, Patwardhan, Rajesh <rajesh.patwardhan@etrade.com> wrote:\n>> Hello Michael,\n>> I will explain a scenario (we are passing thru this right now)\n>> 1) you have 10 years worth of cvs data.\n>> 2) We want to move to svn.\n>> 3) The repository move should be in such a way that the development does\n>> not get hampered for any 1 work day.\n>> 4) We have atleast 4 major modules in cvs which takes about 30 - 40\n>> hours each for conversion currently.\n> \n> There are known ways (that haven't been implemented) to get the 40 hr\n> number down to 1/2 hour. Would that be a better approach than doing\n> incremental imports?\n\nJon, I would like very much to hear how you propose to get an 60-fold\nspeed increase in cvs2svn.  I've never heard of any plausible way to\naccomplish anything even close to this.\n\nPlease note that the user wants to convert to Subversion, not git.  But\neven converting to git, I don't think that such speeds are possible\nwithout massive changes that would include processing everything in RAM\nand switching large parts of cvs2svn from Python to a compiled language.\n\nMichael\n"},{"id":"49623","messageId":"9e4733910708031316x1b7d2a40n5d0298cedd6cf97c@mail.gmail.com","threadId":"9333","inReplyTo":"46B37ADB.8020103@alum.mit.edu","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-03T20:16:23Z","receivedAt":"2007-08-03T20:16:23Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/3/07, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> [I set followup-to users@cvs2svn.tigris.org, since this has nothing to\n> do with git.]\n>\n> Jon Smirl wrote:\n> > On 8/3/07, Patwardhan, Rajesh <rajesh.patwardhan@etrade.com> wrote:\n> >> Hello Michael,\n> >> I will explain a scenario (we are passing thru this right now)\n> >> 1) you have 10 years worth of cvs data.\n> >> 2) We want to move to svn.\n> >> 3) The repository move should be in such a way that the development does\n> >> not get hampered for any 1 work day.\n> >> 4) We have atleast 4 major modules in cvs which takes about 30 - 40\n> >> hours each for conversion currently.\n> >\n> > There are known ways (that haven't been implemented) to get the 40 hr\n> > number down to 1/2 hour. Would that be a better approach than doing\n> > incremental imports?\n>\n> Jon, I would like very much to hear how you propose to get an 60-fold\n> speed increase in cvs2svn.  I've never heard of any plausible way to\n> accomplish anything even close to this.\n>\n> Please note that the user wants to convert to Subversion, not git.  But\n> even converting to git, I don't think that such speeds are possible\n> without massive changes that would include processing everything in RAM\n> and switching large parts of cvs2svn from Python to a compiled language.\n\nMake a bulk importer for SVN like git-fastimport. I measured some SVN\nimports and the bulk of the time was spent forking off SVN. Before\ngit-fast import it would have taken git two weeks to import Mozilla\nCVS.\n\n>\n> Michael\n>\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"49624","messageId":"9e4733910708031327u7df2205ap56a7ad5430380fb@mail.gmail.com","threadId":"9333","inReplyTo":"9e4733910708031316x1b7d2a40n5d0298cedd6cf97c@mail.gmail.com","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2007-08-03T20:27:49Z","receivedAt":"2007-08-03T20:27:49Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 8/3/07, Jon Smirl <jonsmirl@gmail.com> wrote:\n> Make a bulk importer for SVN like git-fastimport. I measured some SVN\n> imports and the bulk of the time was spent forking off SVN. Before\n> git-fast import it would have taken git two weeks to import Mozilla\n> CVS.\n\nAnd add a CVS parser to cvs2svn. Use the one I posted or write it again.\nFork is not a very fast operation, millions of forks take a week to run.\n\nIn the cvs2git code I did there was one process running cvs2svn and it\nparsed the CVS files internally. A second process ran git-fastimport.\nNothing else was forked.\n\nWhen I first started we were forking both git and cvs. When I ran\noprofile on it 95% of the CPU time was being spent in the kernel.\nLinus helped me figure out what was going on. It was the overhead of\npage table copies associated with millions of forks that was taking so\nlong. The solution is to eliminate the forks.\n\nMy first try with forks for both cvs and git took about a week to\nimport Mozilla CVS. After all the forks were eliminated I could import\nMozilla CVS in four hours.\n\n>\n> >\n> > Michael\n> >\n> >\n>\n>\n> --\n> Jon Smirl\n> jonsmirl@gmail.com\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"49687","messageId":"43D2B97E-6AC5-4A7E-AA86-DDFF4992A284@zib.de","threadId":"9333","inReplyTo":"46B25FC3.6000205@fs.ei.tum.de","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Steffen Prohaska","fromEmail":"prohaska@zib.de","sentAt":"2007-08-04T08:28:04Z","receivedAt":"2007-08-04T08:28:04Z","isPatch":false,"sender":{"key":"prohaska@zib.de","avatar":"https://avatars.githubusercontent.com/u/217580?v=4"},"body":"\nOn Aug 3, 2007, at 12:50 AM, Simon 'corecode' Schubert wrote:\n\n> Steffen Prohaska wrote:\n>>> yah, that pretty much tells me it is shawn's bug :)  but without  \n>>> more details, it is very hard to diagnose.\n>> I tried again. Interestingly now togit works but tohg still fails.\n>> togit starts with reporting\n>> fatal: Not a valid object name\n>\n> that's fine.\n\nLooks a bit scary. Could you hide the message from the user\nif it's fine.\n\n>> as the first line. But besides that it seems to work fine. What\n>> concerns me a bit is that the last line togit reports is\n>> committing set 18100/18173\n>> I'd expect it should report 18173/18173.\n>\n> that's fine as well.  You only saw multiples of 100, but you didn't  \n> consider it would skip the itermediate ones, right? :)\n\nI don't care about the intermediates, but only about the\nlast one. I'd expect that a successful import would report\nas the last line 18173/18173. If the first number is smaller\nthan the second, this indicates to me that there's something\nleft to do.\n\n\n>> BTW, togit creates much more complex branching patterns than cvs2svn\n>> does. The attached file branching.png displays a small view of a\n>> branching pattern that extends downwards over a couple of screens.\n>> I checked the cvs2svn history again. It doesn't contain anything\n>> of similar complexity.\n>\n> haha yea, there is still some issue with duplicate branch names and  \n> the branchpoint.  if it doesn't get the branch right, it will  \n> always \"pull\" files from the parent branch.\n>\n> did you do some manual RCS file copying or manual branch name  \n> changing of individual files?  this could be the reason.  I still  \n> have to find a simple repo to reproduce this.\n\nMaybe, the repo is 8 years old. It started before I joined the\ndevelopment.\n\n\tSteffen\n"},{"id":"49837","messageId":"20070805075800.GA4256@ugly.local","threadId":"9333","inReplyTo":"9e4733910708021659y6e9bb7ddk58817b4de3df26a0@mail.gmail.com","subject":"Re: cvs2svn conversion directly to git ready for experimentation","fromName":"Oswald Buddenhagen","fromEmail":"ossi@kde.org","sentAt":"2007-08-05T07:58:00Z","receivedAt":"2007-08-05T07:58:00Z","isPatch":false,"sender":{"key":"ossi@kde.org","avatar":"https://avatars.githubusercontent.com/u/812380?v=4"},"body":"On Thu, Aug 02, 2007 at 07:59:41PM -0400, Jon Smirl wrote:\n> I seem to recall discussing an algorithm  to fix this on the cvs2svn\n> mailing list. There was a somewhat simple way to correlate the\n> \"unlabeled-1.2.4\" in one file might be the same as \"unlabeled-1.2.6\"\n> problem.\n> \nyes, name them after the first symbol that appears on them. like\nunlabeled-1.2.4 being named __KDE_3_5_RELEASE because of such tag\n(without the underscores, obviously) appearing on it.\nthe naive per-file implementation doesn't get you that far, though.\nagain, one'd have to collect data from all files first, correlate\nit and make a \"majority vote\". very similar to your favorite symbol\nsource problem. ;)\n\n-- \nHi! I'm a .signature virus! Copy me into your ~/.signature, please!\n--\nChaos, panic, and disorder - my work here is done.\n"}]}