{"thread":{"id":"32445","subject":"cvsps, parsecvs, svn2git and the CVS exporter mess","startedAt":"2012-12-22T17:36:48Z","lastAt":"2013-01-06T11:15:24Z","messageCount":11,"participants":["Eric S. Raymond","Heiko Voigt","Michael Haggerty","Martin Langhoff","Max Horn","Jonathan Nieder"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"205423","messageId":"20121222173649.04C5B44119@snark.thyrsus.com","threadId":"32445","inReplyTo":null,"subject":"cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2012-12-22T17:36:48Z","receivedAt":"2012-12-22T17:36:48Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Wanting reposurgeon to be able to read CVS repositories has landed me\nin the middle of a mess.  This note explains my thinking, what I\nintend to do to fix the mess, and how others can help if they are\nmotivated.  I have copied all the individuals who I know have an\ninterest in the problem, and the git list because what I'm planning is\ngoing to be significant for them.\n\nMy requirement is not complicated to describe. For use as a\nreposurgeon front end I need a tool that is basically a\ncvs-fast-export, runnable in either a CVS repository or a checkout\n(either will do, both is not required) and emitting a fast-import\nstream.\n\nThere are three competing tools that might fit this bill.  \n\n* One is Michael Haggerty's cvs2git.  I had bad experiences with the\ncvs2svn code it's derived from in the past, but Michael believes those\nproblems have been fixed and I will accept that - at least until I can\ntest for myself.  Its documented interface is not quite good enough\nyet; as the documentation says, \"The data that should be fed to git\nfast-import are written to two files, which have to be loaded into git\nfast-import manually.\"\n\n* Another is the cvsps code formerly maintained by David Mansfield and\nused by git-cvsimport; he passed the maintainer's baton to me, and I\nhave shipped a 3.0 with a working --fast-export option.\n\n* A third is parsecvs, which Keith Packard and Bart Massey handed off\nto me a week before David invited me to take over cvsps.  While\nparsecvs does not yet have a --fast-export option, I anticipate no\ngreat difficulty in adding one.\n\nIt is pure accident that I now maintain two of these.  Initially I\nwas interested in parsecvs, but it failed to build for me.  By the\ntime Bart Massey sent me a fix patch, I had already been handed \ncvsps, added --fast-export, and shipped 3.0.\n\nHaving three different tools for this job seems to me duplicative and\npointless; two of them should probably be let die an honorable death.\nI don't actually care which of the three survives - and, in\nparticular, if I determine that cvs2git is doing the best job of the\nthree I am quite willing to declare end-of-life for cvsps and\nparsecvs.  It's not like I don't have plenty of other projects to work\non.\n\nTherefore, I think my main focus needs to be developing a really\neffective test suite to triage these tools with and applying it to all\nthree.  I have already made a solid start on this; see\ntests/cvstest.py and tests/basic.tst in the cvsps-3.0 distribution for\nmy test framework.\n\nI presently know of three test suites other than mine. One was built\nby Heiko to test cvsps, another lives in the git t/ directory, and the\nthird is cvs2git's. I haven't looked at cv2git's yet, but the others\nare not in their present form suited to where I am taking cvsps and\nparsecvs.  Heiko's relies on the default human-readable cvsps format,\nwhich I consider obsolete and uninteresting.  The git tests are\ndependent on details of porcelain behavior.  I think it would be\nbetter to test import-stream output.\n\nHere is what I propose.  Let's build a common test suite that cvs2git,\ngit-cvsimport, cvsps, and parsecvs can all use, apply it rigorously,\nand let the best tool win.  (This would mean, among other things, that\ngit can stop carrying things that are essentially cvsps tests in its\ntree.)\n\nThe two people I most need to sign off on this are, I guess, Michael\nHaggerty and either Junio Hamano or whoever specifically owns\ngit-cvsimport and its tests.  Whichever way this comes out, the back\nend of git-cvsimport is going to need some work - I don't plan to put\nany further effort into the output format it's presently using.\n\nIf we can agree on this, I'll start a public repo, and contribute my\nPython framework - it's more capable than any of the shell harnesses\nout there because it can easily drive interleaved operations on multiple \ncheckout directories.\n\nAnybody who is still interested in this problem should contribute\ntests.  Heiko Voigt, I'd particularly like you in on this.  David\nMansfield, if you can spare the few minutes required to write\ngenerators for the \"funky\" and \"invalid\" tag cases, that would be\nreally helpful.  Michael Haggerty, your piece would be moving the\ncvs2git tests to the new framework.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\nThe kind of charity you can force out of people nourishes about as much as\nthe kind of love you can buy --- and spreads even nastier diseases.\n"},{"id":"205452","messageId":"20121223202111.GB29354@book-mint","threadId":"32445","inReplyTo":"20121222173649.04C5B44119@snark.thyrsus.com","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Heiko Voigt","fromEmail":"hvoigt@hvoigt.net","sentAt":"2012-12-23T20:21:11Z","receivedAt":"2012-12-23T20:21:11Z","isPatch":false,"sender":{"key":"hvoigt@hvoigt.net","avatar":"https://avatars.githubusercontent.com/u/184958?v=4"},"body":"Hi,\n\nOn Sat, Dec 22, 2012 at 12:36:48PM -0500, Eric S. Raymond wrote:\n> If we can agree on this, I'll start a public repo, and contribute my\n> Python framework - it's more capable than any of the shell harnesses\n> out there because it can easily drive interleaved operations on multiple \n> checkout directories.\n\nPlease share so we can have a look. BTW, where can I find your cvsps\ncode?\n\n> Anybody who is still interested in this problem should contribute\n> tests.  Heiko Voigt, I'd particularly like you in on this.\n\nIf it does not take to much effort I could port my tests to the new\nframework. Since I currently are not in active need of cvs conversions\nits not of big interest to me anymore. But if it does not take too much\ntime I am happy to help.\n\n>From my past cvs conversion experiences my personal guess is that\ncvs2svn will win this competition.\n\nCheers Heiko\n"},{"id":"205454","messageId":"20121223224542.GA16383@thyrsus.com","threadId":"32445","inReplyTo":"20121223202111.GB29354@book-mint","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2012-12-23T22:45:42Z","receivedAt":"2012-12-23T22:45:42Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Heiko Voigt <hvoigt@hvoigt.net>:\n> Please share so we can have a look. BTW, where can I find your cvsps\n> code?\n\nhttps://gitorious.org/cvsps\n\nDevelopments of the last 48 hours:\n\n1. Andreas Schwab sent me a patch that uses commitids wherever the history\n   has them - this makes all the time-skew problems go away.  I added code\n   to warn if commitids aren't present, so users will get a clear indication\n   of when time-skew problems might bite them versus when that is happily\n   impossible.\n\n2. I've scrapped a lot of obsolete code and options.  The repo head\n   version uses what used to be called cvs-direct mode all the time\n   now; it works, and the effect on performance is major.  This also\n   means that cvsps doesn't need to use any local CVS commands or even\n   have CVS installed where it runs.\n\n> >From my past cvs conversion experiences my personal guess is that\n> cvs2svn will win this competition.\n\nThat could be.  But right now cvsps has one significant advantage over\ncvs2git (which parsecvs might share) - it's *blazingly* fast.  So fast\nthat I scrapped all the local-caching logic; there seems no point to it at\ntoday's network speeds, and that's one less layer of complications to\ngo wrong.\n\nI've removed a couple hundred lines of code and the program works\nbetter and faster than it did before.  That's having a good day!\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"205932","messageId":"50E5A5CF.2070009@alum.mit.edu","threadId":"32445","inReplyTo":"20121222173649.04C5B44119@snark.thyrsus.com","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2013-01-03T15:37:51Z","receivedAt":"2013-01-03T15:37:51Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 12/22/2012 06:36 PM, Eric S. Raymond wrote:\n> * One is Michael Haggerty's cvs2git.  I had bad experiences with the\n> cvs2svn code it's derived from in the past, but Michael believes those\n> problems have been fixed and I will accept that - at least until I can\n> test for myself.  Its documented interface is not quite good enough\n> yet; as the documentation says, \"The data that should be fed to git\n> fast-import are written to two files, which have to be loaded into git\n> fast-import manually.\"\n\nThere are two good reasons that the output is written to two separate files:\n\n1. The files are generated during different passes of cvs2git, and since\nthe cvs2git conversion is restartable pass-by-pass, the first file might\nonly need to be generated once even while the user is iterating on\nadjustments to other conversion options.\n\n2. The first (\"blobfile\") contains blob definitions for file revisions,\nwhich are read out of the RCS files in the order they are held in the\nRCS file.  This is vastly faster than reading the file revisions in the\norder that they are needed for git commits because (1) all revisions for\na file can be computed from one serial read of the RCS file; (2) there\nis no need to jump around from rcsfile to rcsfile.  The second\n(\"dumpfile\") stitches the blobs together into git commits by referring\nto the blobs that are needed.  This file is smaller because it doesn't\ncontain the actual file contents.  Another advantage of this approach is\nthat a blob need only appear once in the blobfile even if it is used\nmultiple times in the git history.\n\nAnyway, surely cat'ing two output files together is not such a difficult\nproblem?\n\nA potentially bigger problem is that if you want to handle such\nblob/dump output, you have to deal with git-fast-import format's \"blob\"\ncommand as opposed to only handling inline blobs.  However, if that is a\nproblem, it is possible to configure cvs2git to write the blobs inline\nwith the rest of the dumpfile (this mode is supported because \"hg\nfast-import\" doesn't support detached blobs).  You would have to create\nan options file that uses GitRevisionInlineWriter, similar to what is\ndone in cvs2hg-example.options.\n\n> [...]\n> Having three different tools for this job seems to me duplicative and\n> pointless; two of them should probably be let die an honorable death.\n> I don't actually care which of the three survives - and, in\n> particular, if I determine that cvs2git is doing the best job of the\n> three I am quite willing to declare end-of-life for cvsps and\n> parsecvs.  It's not like I don't have plenty of other projects to work\n> on.\n\ncvs2git does not currently support incremental conversions; therefore, a\ncvsps-based option (if it would actually work, that is) would have at\nleast one advantage over cvs2git.\n\n> I presently know of three test suites other than mine. One was built\n> by Heiko to test cvsps, another lives in the git t/ directory, and the\n> third is cvs2git's. I haven't looked at cv2git's yet, but the others\n> are not in their present form suited to where I am taking cvsps and\n> parsecvs.  Heiko's relies on the default human-readable cvsps format,\n> which I consider obsolete and uninteresting.  The git tests are\n> dependent on details of porcelain behavior.  I think it would be\n> better to test import-stream output.\n\ncvs2svn has an extensive test suite which includes tests derived from\nbug reports that we have received over the years.  I adapted a few of\nits test repositories to create the git test suite additions that I made\nin Feb 2009, but there are many more in our project.\n\nA lot of our test suite deals with additional conversion features, like:\n\n* Re-encoding filenames, usernames, and log messages from whatever\nhappens to have been used in the CVS repository into UTF-8\n\n* Fixing CVS branches, tags, and mixed branch/tag messes according to\nuser wishes; renaming branches and tags\n\n* Allowing the user to influence the choice of which branch should serve\nas the source for another branch/tag (CVS records this information very\nambiguously)\n\n* Fixing binary vs. text files, expanding/contracting CVS keywords, etc.\n\n* Removing lots of synthetic revisions and other cruft generated by CVS\nto fit within the RCS file format\n\n* Dealing with vendor branches in a sensible way, especially considering\nthat very many users misuse vendor branches for initial imports\n\n* Dealing with various common types of CVS repository corruption\n\nSee our list of features [1] for more details.  Presumably many of these\nfeatures would not be covered by your test framework, and are not\nsupported by the other conversion tools.\n\nUnfortunately, our tests are mostly based on cvs2svn (i.e., not 2git);\nthat is, the conversion is done with cvs2svn and checked by verifying\nthe contents of the resulting Subversion repository.\n\nThe script contrib/verify-cvs2svn.py is another kind of test; it checks\nevery branch and tag out of CVS and the destination repository and\nverifies that their contents are identical.  This script is intended to\nbe used by users to check their own conversion.  Please note that it\ndoesn't check the history, only the branch/tag tips.  But this script\nworks with both Subversion and git (at least it should; it probably\ndoesn't get tested much).\n\n> Here is what I propose.  Let's build a common test suite that cvs2git,\n> git-cvsimport, cvsps, and parsecvs can all use, apply it rigorously,\n> and let the best tool win.  (This would mean, among other things, that\n> git can stop carrying things that are essentially cvsps tests in its\n> tree.)\n\nI think it would be great to have a way to test across tools, though\nplease realize that the inference of the most plausible \"true\" CVS\nhistory is partly objective but also often a matter of heuristics and\ntaste.  Moreover, the choice of how to represent the inferred history in\ngit, which has rather a different model than CVS/Subversion, is also\nnon-obvious and somewhat controversial.  I expect that there will be a\nnumber of simple CVS repositories for which we can all agree about the\ncorrect git output, but not far away will be a vast number for which the\n\"correct\" answer is unclear.  Many of the interesting tests would fall\ninto the latter category.\n\n> The two people I most need to sign off on this are, I guess, Michael\n> Haggerty and either Junio Hamano or whoever specifically owns\n> git-cvsimport and its tests.  [...]\n\nIt's not clear what you want me to sign off on.  I guess you want to\nreplace (or augment?) the cvs2svn test suite with one based on your\nframework?  Right off the top of my head I can think of a few\nconsiderations from the point of view of the cvs2svn project:\n\n* We definitely want to continue testing the Subversion output of\ncvs2svn.  A test suite that only tests the git output could at best be\nan addition to the current test suite, not a replacement for it.  (That\nbeing said, the addition of good tests of the 2git output would be great.)\n\n* A test suite that tests only the easy cases wouldn't really be\ninteresting, because the difficult cases are where the potential\nproblems lie.\n\n* It would be unfortunate if the cvs2svn test suite would grow another\nrun-time dependency or if we would have to invest a lot of time\nsynchronizing with another project, though if the gain were big enough\nwe could consider it.\n\n* The licenses obviously have to be compatible to the extent required by\nthe level of coupling.\n\n* I don't have a lot of time to work on the integration.  cvs2svn has\nlong been at a level of maturity where it doesn't need much care and\nfeeding, and I would like to keep it that way :-)  Nowadays I am far\nmore interested in working on the git project with my little available\nopen-sourcin' time.\n\n\nRereading this email, I realize that it is not clear to me why your new\ntesting project needs the \"signoff\" or cooperation from any of the\nconversions tool projects (git-cvsimport or cvs2svn or parsecvs or ...)\nin the first place.  The essence of your project will be a collection of\nCVS test repositories, and code that can read the conversion output\n(whether via git or as fast-input data) and verify that it matches\nexpectations (right?).  Presumably it will have a place where any of the\nconversion tools could be plugged into it, and perhaps a bit of code\nthat knows how to configure and run the best-known tools (and perhaps\neven to download and build them).\n\nIt would seem natural to me that your project stops there, and stays at\narms-length from the conversion projects.  If your test suite proves\nitself to be obviously better than the cvs2svn test suite, then we might\ntry to integrate it *then* (or not; even then it wouldn't really be\nobligatory).\n\nMichael\n\n[1] http://cvs2svn.tigris.org/features.html\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"205934","messageId":"CACPiFCLpZT6v7tS+D9=hZwpj3kOewndcyFYfUBso62Ry_Y-Eow@mail.gmail.com","threadId":"32445","inReplyTo":"20121222173649.04C5B44119@snark.thyrsus.com","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-01-03T15:51:02Z","receivedAt":"2013-01-03T15:51:02Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Sat, Dec 22, 2012 at 12:36 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> It is pure accident that I now maintain two of these.\n\nMaintainership is always temporary.\n\n> Having three different tools for this job seems to me duplicative and\n> pointless; two of them should probably be let die an honorable death.\n\nPerhaps just maintain the code that serves your goals. That way, you\ndon't need long trolly emails nor approval from anyone.\n\n\n\n\nm\n--\n martin.langhoff@gmail.com\n martin@laptop.org -- Software Architect - OLPC\n - ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n - http://wiki.laptop.org/go/User:Martinlanghoff\n"},{"id":"205953","messageId":"20130103205301.GD26201@thyrsus.com","threadId":"32445","inReplyTo":"50E5A5CF.2070009@alum.mit.edu","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-01-03T20:53:01Z","receivedAt":"2013-01-03T20:53:01Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu>:\n> There are two good reasons that the output is written to two separate files:\n\nThose are good reasons to write to a pair of tempfiles, and I was able\nto deduce in advance most of what your explanation would be from the\nbare fact that you did it that way.\n\nThey are *not* good reasons for having an interface that exposes this\nimplementation detail to the caller - that choice I consider a failure\nof interface-design judgment.  But I know how to fix this in a simple and\nbackward-compatible way, and will do so when I have time to write you\na patch.  Next week or the week after, most likely.\n\nAlso, the cvs2git manual page is still rather half-baked and careless,\nwith several fossil references to cvs2svn that shouldn't be there and\nobviously incomplete feature coverage. Fixing these bugs is also on my\nto-do list for sometime this month.\n\nI'd be willing to put in this work anyway, but it still in the back of\nmy mind that if cvs2git wins the test-suite competition I might\nofficially end-of-life both cvsps and parsecvs.  One of the features\nof the new git-cvsimport is direct support for using cvs2git as a\nconversion engine.\n \n> A potentially bigger problem is that if you want to handle such\n> blob/dump output, you have to deal with git-fast-import format's \"blob\"\n> command as opposed to only handling inline blobs.\n\nNot a problem.  All of the main potential consumers for this output,\nincluding reposurgeon, handle the blob command just fine.\n\n> cvs2git does not currently support incremental conversions; therefore, a\n> cvsps-based option (if it would actually work, that is) would have at\n> least one advantage over cvs2git.\n\nYes. The reason I didn't ship the replacement patch Junio was\nexpecting yesterday is that I don't have test coverage for the\nincremental case.  I'm working on that now.\n\n> cvs2svn has an extensive test suite which includes tests derived from\n> bug reports that we have received over the years.  I adapted a few of\n> its test repositories to create the git test suite additions that I made\n> in Feb 2009, but there are many more in our project.\n\nI've merged those into my tree.\n\n> I think it would be great to have a way to test across tools, though\n> please realize that the inference of the most plausible \"true\" CVS\n> history is partly objective but also often a matter of heuristics and\n> taste.  Moreover, the choice of how to represent the inferred history in\n> git, which has rather a different model than CVS/Subversion, is also\n> non-obvious and somewhat controversial.  I expect that there will be a\n> number of simple CVS repositories for which we can all agree about the\n> correct git output, but not far away will be a vast number for which the\n> \"correct\" answer is unclear.  Many of the interesting tests would fall\n> into the latter category.\n\nI'm aware of the problem.  One of the interesting questions is how much\nfurther into the weird cases everybody can agree on what correct \ntranslation looks like.  We won't know until we push it.\n \n> It's not clear what you want me to sign off on.\n\nIf you're not willing to use the new suite, my spending the effort \nrequired to genericize it gets much less interesting.  I needed \nJunio's agreement because I wanted to move the old git-cvsimport\ntests from the git tree to the new test suite; they're not really\ntests of the wrapper script at all but of the conversion engines.\n\n>                                               I guess you want to\n> replace (or augment?) the cvs2svn test suite with one based on your\n> framework? \n\nAugment, not replace - and just as importantly, commit to writing \nnew tests into the new generic framework when they don't involve a \ntool-specific option.  It would be silly and duplicative for us *not*\nto be sharing as many tests as we can.\n\n> * We definitely want to continue testing the Subversion output of\n> cvs2svn.  A test suite that only tests the git output could at best be\n> an addition to the current test suite, not a replacement for it.  (That\n> being said, the addition of good tests of the 2git output would be great.)\n\nAgreed.\n\n> * A test suite that tests only the easy cases wouldn't really be\n> interesting, because the difficult cases are where the potential\n> problems lie.\n\nYes, I know.  I'm arguing that we should be doing that exploration\njointly rather than separately.\n\n> * It would be unfortunate if the cvs2svn test suite would grow another\n> run-time dependency or if we would have to invest a lot of time\n> synchronizing with another project, though if the gain were big enough\n> we could consider it.\n\nI know how to keep the friction cost low.  You'll see more about this when\nI split off the test suite and announce it.\n\n> * The licenses obviously have to be compatible to the extent required by\n> the level of coupling.\n\nI don't think this will be a problem.  You own the copyright on your tests and\nI own it on mine, so we can relicense under whatever common license we choose.\nI'm not fussy about what we use; ASL 2.0 would be fine by me.\n\n> * I don't have a lot of time to work on the integration.  cvs2svn has\n> long been at a level of maturity where it doesn't need much care and\n> feeding, and I would like to keep it that way :-)  Nowadays I am far\n> more interested in working on the git project with my little available\n> open-sourcin' time.\n\nI don't want to spend the rest of my life on the CVS-lifting problem either.\nMy present plans envision intense work on it for another three weeks or\nso, after which I expect we'll be at a relatively stable and low-maintainance\nstate. \n\nFYI, here are my agenda items in roughly the order I expect to finish them:\n\n1. Write test coverage for incremental imports.\n2. Ship version 2 of the git-cvsimport replacement patch (with the fallback \n   option Junio requested) to the git list.\n3. Get parsecvs to a non-broken state and ship a release\n4. Ship a patch for git-cvsimport that adds the option to use parsecvs \n   as a conversion engine.\n5. Break the test suite out of cvsps, give it its own public repo, document\n   it, and hand you the keys.\n6. Fix the interface-design bug(s) in cvs2git, and its documentation.\n7. Torture-test all three tools (cvsps, parsecvs, cvs2git) against the\n   new suite.\n8. Make a judgement about whether I should EOL cvsps or parsecvs or both.\n\nI have other commitments, so this will take a bit longer than it might\nhave.  I expect to be at step 8 in roughly a month (early February).\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"206015","messageId":"1E7F9F86-F040-42E4-98C4-152B8CCE47CE@quendi.de","threadId":"32445","inReplyTo":"20130103205301.GD26201@thyrsus.com","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Max Horn","fromEmail":"postbox@quendi.de","sentAt":"2013-01-05T08:27:38Z","receivedAt":"2013-01-05T08:27:38Z","isPatch":false,"sender":{"key":"postbox@quendi.de","avatar":null},"body":"\nOn 03.01.2013, at 21:53, Eric S. Raymond wrote:\n\n> Michael Haggerty <mhagger@alum.mit.edu>:\n>> There are two good reasons that the output is written to two separate files:\n> \n> Those are good reasons to write to a pair of tempfiles, and I was able\n> to deduce in advance most of what your explanation would be from the\n> bare fact that you did it that way.\n> \n> They are *not* good reasons for having an interface that exposes this\n> implementation detail to the caller - that choice I consider a failure\n> of interface-design judgment.  But I know how to fix this in a simple and\n> backward-compatible way, and will do so when I have time to write you\n> a patch.  Next week or the week after, most likely.\n> \n> Also, the cvs2git manual page is still rather half-baked and careless,\n> with several fossil references to cvs2svn that shouldn't be there and\n> obviously incomplete feature coverage. Fixing these bugs is also on my\n> to-do list for sometime this month.\n> \n> I'd be willing to put in this work anyway, but it still in the back of\n> my mind that if cvs2git wins the test-suite competition I might\n> officially end-of-life both cvsps and parsecvs.  One of the features\n> of the new git-cvsimport is direct support for using cvs2git as a\n> conversion engine.\n> \n>> A potentially bigger problem is that if you want to handle such\n>> blob/dump output, you have to deal with git-fast-import format's \"blob\"\n>> command as opposed to only handling inline blobs.\n> \n> Not a problem.  All of the main potential consumers for this output,\n> including reposurgeon, handle the blob command just fine.\n\nHm, you snipped this part of Michael's mail:\n\n>> However, if that is a\n>> problem, it is possible to configure cvs2git to write the blobs inline\n>> with the rest of the dumpfile (this mode is supported because \"hg\n>> fast-import\" doesn't support detached blobs).\n\nI would call \"hg fast-import\" a main potential customer, given that there \"cvs2hg\" is another part of the cvs2svn suite. So I can't quite see how you can come to your conclusion above...\n\n\n\nCheers,\nMax"},{"id":"206031","messageId":"20130105151106.GA1938@thyrsus.com","threadId":"32445","inReplyTo":"1E7F9F86-F040-42E4-98C4-152B8CCE47CE@quendi.de","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-01-05T15:11:07Z","receivedAt":"2013-01-05T15:11:07Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Max Horn <postbox@quendi.de>:\n> Hm, you snipped this part of Michael's mail:\n> \n> >> However, if that is a\n> >> problem, it is possible to configure cvs2git to write the blobs inline\n> >> with the rest of the dumpfile (this mode is supported because \"hg\n> >> fast-import\" doesn't support detached blobs).\n> \n> I would call \"hg fast-import\" a main potential customer, given that there \"cvs2hg\" is another part of the cvs2svn suite. So I can't quite see how you can come to your conclusion above...\n\nPerhaps I was unclear.  I consider the interface design error to\nbe not in the fact that all the blobs are written first or detached,\nbut rather that the implementation detail of the two separate journal\nfiles is ever exposed.\n\nI understand why the storage of intermediate results was done this\nway, in order to decrease the tool's working set during the run, but\nfinishing by automatically concatenating the results and streaming\nthem to stdout would surely have been the right thing here.\n \nThe downstream cost of letting the journalling implementation be\nexposed, instead, can be seen in this snippet from the new git-cvsimport\nI've been working on:\n\n    def command(self):\n        \"Emit the command implied by all previous options.\"\n        return \"(cvs2git --username=git-cvsimport --quiet --quiet --blobfile={0} --dumpfile={1} {2} {3} && cat {0} {1} && rm {0} {1})\".format(tempfile.mkstemp()[1], tempfile.mkstemp()[1], self.opts, self.modulepath)\n\nAccording to the documentation, every caller of csv2git must go\nthrough analogous contortions!  This is not the Unix way; if Unix\ndesign principles had been minimally applied, that second line would\njust read like this:\n\n     return \"cvs2git --username=git-cvsimport --quiet --quiet\"\n\nIf Unix design principles had been thoroughly applied, the \"--quiet\n--quiet\" part would be unnecessary too - well-behaved Unix commands\n*default* to being completely quiet unless either (a) they have an\nexceptional condition to report, or (b) their expected running time is\nso long that tasteful silence would leave users in doubt that they're\nworking.\n\n(And yes, I do think violating these principles is a lapse of taste when\ngit tools do it, too.)\n\nMichael Haggerty wants me to trust that cvs2git's analysis stage has\nbeen fixed, but I must say that is a more difficult leap of faith when\ntwo of the most visible things about it are still (a) a conspicuous\ninstance of interface misdesign, and (b) documentation that is careless and\nincomplete.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"206032","messageId":"20130105155804.GB1938@thyrsus.com","threadId":"32445","inReplyTo":"CAA6gtpky9JxFDdpLM6kY9su-9FWX8RoWHU4uptd_Zk+ZJuhrtA@mail.gmail.com","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-01-05T15:58:04Z","receivedAt":"2013-01-05T15:58:04Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Bart Massey <bart@cs.pdx.edu>:\n> I don't know what Eric Raymond \"officially end-of-life\"-ing parsecvs means?\n\nYou and Keith handed me the maintainer's baton.  If I were to EOL it,\nthat would be the successor you two designated judging in public that\nthe code is unsalvageable or has become pointless.  If you wanted to\nexclude the possibility that a successor would make that call, you\nshouldn't have handed it off in a state so broken that I can't even\ntest it properly.\n\nBut I don't in fact think the parsecvs code is pointless. The fact that it\nonly needs the ,v files is nifty and means it could be used as an RCS\nexporter too.  The parsing and topo-analysis stages look like really\ngood work, very crisp and elegant (which is no less than I'd expect\nfrom Keith, actually).\n\nAlas, after wrestling with it I'm beginning to wonder whether the\ncodebase is salvageable by anyone but Keith himself.  The tight coupling\nto the git cache mechanism is the biggest problem.  So far, I can't\nfigure out what tree.c is actually doing in enough detail to fix it or pry\nit loose - the code is opaque and internal documentation is lacking.\n\nMore generally, interfacing to the unstable API of libgit was clearly\na serious mistake, leading directly to the current brokenness.  The\ntool should have emitted an import stream to begin with.  I'm trying\nto fix that, but success is looking doubtful.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"206068","messageId":"20130105225740.GC3247@elie.Belkin","threadId":"32445","inReplyTo":"20130105151106.GA1938@thyrsus.com","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2013-01-05T22:57:40Z","receivedAt":"2013-01-05T22:57:40Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Eric S. Raymond wrote:\n\n> Michael Haggerty wants me to trust that cvs2git's analysis stage has\n> been fixed, but I must say that is a more difficult leap of faith when\n> two of the most visible things about it are still (a) a conspicuous\n> instance of interface misdesign, and (b) documentation that is careless and\n> incomplete.\n\nFor what it's worth, I use cvs2git quite often.  I've found it to work\nwell and its code to be clear and its developers responsive.  But I\ndon't mind if we disagree, and multiple implementations to explore the\ndesign space of importers doesn't seem like a terrible outcome.\n\nThanks for your work,\nJonathan\n"},{"id":"206126","messageId":"50E95CCC.30002@alum.mit.edu","threadId":"32445","inReplyTo":"20130105151106.GA1938@thyrsus.com","subject":"Re: cvsps, parsecvs, svn2git and the CVS exporter mess","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2013-01-06T11:15:24Z","receivedAt":"2013-01-06T11:15:24Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 01/05/2013 04:11 PM, Eric S. Raymond wrote:\n> Perhaps I was unclear.  I consider the interface design error to\n> be not in the fact that all the blobs are written first or detached,\n> but rather that the implementation detail of the two separate journal\n> files is ever exposed.\n> \n> I understand why the storage of intermediate results was done this\n> way, in order to decrease the tool's working set during the run, but\n> finishing by automatically concatenating the results and streaming\n> them to stdout would surely have been the right thing here.\n\ncvs2svn/cvs2git is built to be able to handle very large CVS\nrepositories, not only those that can fit in RAM.  This goal influences\na lot of its design, including the pass-by-pass structure with\nintermediate databases and the resumability of passes.\n\nThe blobfile necessarily contains every version of every file, with no\ndelta-encoding and no compression.  Its size can be a large multiple of\nthe on-disk size of the original CVS repository.  If the \"save to\ntempfiles then cat tempfiles at end of run\" behavior were hard-coded\ninto cvs2git, then there would be no way to avoid requiring enough\ntemporary space to hold the whole blobfile.\n\nWriting the blobfile into a separate file, on the other hand, means that\nfor example the blobfile could be written into a named pipe connected to\nthe standard input of \"git fast-import\" [1].  \"git fast-import\" could\neven be run on a remote server.\n\nI consider these bigger advantages than the ability to pipe the output\nof cvs2git directly into another command.\n\n> The downstream cost of letting the journalling implementation be\n> exposed, instead, can be seen in this snippet from the new git-cvsimport\n> I've been working on:\n> \n>     def command(self):\n>         \"Emit the command implied by all previous options.\"\n>         return \"(cvs2git --username=git-cvsimport --quiet --quiet --blobfile={0} --dumpfile={1} {2} {3} && cat {0} {1} && rm {0} {1})\".format(tempfile.mkstemp()[1], tempfile.mkstemp()[1], self.opts, self.modulepath)\n> \n> According to the documentation, every caller of csv2git must go\n> through analogous contortions!  This is not the Unix way; if Unix\n> design principles had been minimally applied, that second line would\n> just read like this:\n> \n>      return \"cvs2git --username=git-cvsimport --quiet --quiet\"\n\nNever in my worst nightmares did I imagine that my terrible design taste\nwould force you to type an extra two lines of code.  Oh the humanity!\n\nBy the way, patches are welcome.  And you don't need to trumpet their\nimminent arrival [2] or malign the existing code beforehand.  Moreover,\nit would be adequate if you just demonstrate working code and *then* ask\nfor \"sign-in\", rather than the other way around.\n\n> If Unix design principles had been thoroughly applied, the \"--quiet\n> --quiet\" part would be unnecessary too - well-behaved Unix commands\n> *default* to being completely quiet unless either (a) they have an\n> exceptional condition to report, or (b) their expected running time is\n> so long that tasteful silence would leave users in doubt that they're\n> working.\n\ncvs2git is not a command that one uses 100 times a day.  It is a tool\nfor one-shot conversions of CVS repositories to git.  These conversions\ncan take hours or even days of processing time (not to mention the time\nfor configuring the conversion and changing the rest of a project's\ninfrastructure from CVS to git).  So yes, I think we would like to\nappeal to (b) and humbly ask for your permission to give the user some\nfeedback during the conversion.\n\n> (And yes, I do think violating these principles is a lapse of taste when\n> git tools do it, too.)\n> \n> Michael Haggerty wants me to trust that cvs2git's analysis stage has\n> been fixed, but I must say that is a more difficult leap of faith when\n> two of the most visible things about it are still (a) a conspicuous\n> instance of interface misdesign, and (b) documentation that is careless and\n> incomplete.\n\nThe cvs2git documentation is lacking; I admit it (as opposed to the\ncvs2svn documentation, which I think is quite complete).  And the\nprogram itself also has a lot of rough edges, for example its inability\nto convert .cvsignore files into .gitignore files.  Patches are welcome.\n I haven't used cvs2svn for my own purposes in many years and I've\n*never* once had a need to use cvs2git; I maintain these programs purely\nas a service to the community.  Most of the community seems satisfied\nwith the programs as they are, and if not they usually submit courteous\nand concrete bug reports or submit patches.\n\nI request that you follow their example.  I especially ask that you\nrestrain from spreading public FUD about imagined problems based on\nspeculation.  Please do your tests and *then* report any problems that\nyou find.\n\nYours,\nMichael\n\n[1] In fact, the current implementation of generate_blobs.py sometimes\nseeks back to earlier parts of the blob file when it needs the fulltext\nof a revision that has already been output, but this would be easy to\nchange as soon as somebody needs it.\n\n[2] http://comments.gmane.org/gmane.comp.version-control.git/212340\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"}]}