{"thread":{"id":"35524","subject":"I have end-of-lifed cvsps","startedAt":"2013-12-12T00:17:38Z","lastAt":"2013-12-19T16:18:19Z","messageCount":48,"participants":["Eric S. Raymond","Martin Langhoff","Andreas Krey","Jakub Narębski","Johan Herland","Andreas Schwab","Jonathan Nieder","Jeff King","John Keeping","Kent R. Spillner","Michael Haggerty"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"231936","messageId":"20131212001738.996EB38055C@snark.thyrsus.com","threadId":"35524","inReplyTo":null,"subject":"I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-12T00:17:38Z","receivedAt":"2013-12-12T00:17:38Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"On the git tools wiki, the first paragraph of the entry for cvsps now\nreads:\n\n  Warning: this code has been end-of-lifed by its maintainer in favor of\n  cvs-fast-export. Several attempts over the space of a year to repair\n  its deficient branch analysis and tag assignment have failed.  Do not\n  use it unless you are converting a strictly linear repository and\n  cannot get rsync/ssh read access to the repo masters. If you must use\n  it, be prepared to inspect and manually correct the history using\n  reposurgeon.\n\nI tried very hard to salvage this program - the ability to\nremote-fetch CVS repos without rsync access was appealing - but I\nreached my limit earlier today when I actually found time to assemble\na test set of CVS repos and run head-to-head tests comparing cvsps\noutput to cvs-fast-export output.\n\nI've long believed that that cvs-fast-export has a better analyzer\nthan cvsps just from having read the code for both of them, and having\nhad to fix some serious bugs in cvsps that have no analogs in\ncvs-fast-export.  Direct comparison of the stream outputs revealed\nthat the difference in quality was larger than I had prevously grasped.\n\nAlas, I'm afraid the cvsps repo analysis code turns out to be crap all\nthe way down on anything but the simplest linear and near-linear\ncases, and it doesn't do so hot on even those (all this *after* I\nfixed the most obvious bugs in the 2.x version). In retrospect, trying\nto repair it was misdirected effort.\n\nI recommend that git sever its dependency on this tool as soon as\npossible. I have shipped a 3.13 release with deprecation warnings fot\narchival purposes, after which I will cease maintainance and redirect\nanyone inquiring about cvsps to cvs-fast-export.\n\n(I also maintain cvs-fast-export, but credit for the excellent analysis code \ngoes to Keith Packard.  All I did was write the output stage, document\nit, and fix a few minor bugs.)\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n\nYou [should] not examine legislation in the light of the benefits it will\nconvey if properly administered, but in the light of the wrongs it\nwould do and the harm it would cause if improperly administered\n\t-- Lyndon Johnson, former President of the U.S.\n"},{"id":"231940","messageId":"CACPiFCK+Z7dOfO2v29PMKz+Y_fH1++xqMuTquSQ84d8KyjjFeQ@mail.gmail.com","threadId":"35524","inReplyTo":"20131212001738.996EB38055C@snark.thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-12-12T03:38:20Z","receivedAt":"2013-12-12T03:38:20Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Wed, Dec 11, 2013 at 7:17 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> I tried very hard to salvage this program - the ability to\n> remote-fetch CVS repos without rsync access was appealing\n\nIs that the only thing we lose, if we abandon cusps? More to the\npoint, is there today an incremental import option, outside of\ngit-cvsimport+cvsps?\n\n[ I am a bit out of touch with the current codebase but I coded and\nmaintained a good part of it back in the day. However naive/limited\nthe cvsps parser was, it did help a lot of projects make the leap to\ngit... ]\n\nregards,\n\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"231941","messageId":"20131212042624.GB8909@thyrsus.com","threadId":"35524","inReplyTo":"CACPiFCK+Z7dOfO2v29PMKz+Y_fH1++xqMuTquSQ84d8KyjjFeQ@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-12T04:26:24Z","receivedAt":"2013-12-12T04:26:24Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com>:\n> On Wed, Dec 11, 2013 at 7:17 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> > I tried very hard to salvage this program - the ability to\n> > remote-fetch CVS repos without rsync access was appealing\n> \n> Is that the only thing we lose, if we abandon cusps? More to the\n> point, is there today an incremental import option, outside of\n> git-cvsimport+cvsps?\n\nYou'll have to remind me what you mean by \"incremental\" here. Possibly\nit's something cvs-fast-export could support.\n\nBut what I'm trying to tell you is that, even after I've done a dozen\nreleases and fixed the worst problems I could find, cvsps is far too\nlikely to mangle anything that passes through it.  The idea that you\nare preserving *anything* valuable by sticking with it is a mirage.\n\n\"That bear trap!  It's mangling your leg!\"  \"But it's so *shiny*...\"\n\n> [ I am a bit out of touch with the current codebase but I coded and\n> maintained a good part of it back in the day. However naive/limited\n> the cvsps parser was, it did help a lot of projects make the leap to\n> git... ]\n\nI fear those \"lots of projects\" have subtly damaged repository\nhistories, then.  I warned about this problem a year ago; today I\nfound out it is much worse than I knew then, in fact so bad that I\ncannot responsibly do anything but try to get cvsps turfed out of use\n*as soon as possible*.\n\nAnd no, that should *not* wait on cvs-fast-export getting better \nsupport for \"incremental\" or any other legacy feature.  Every week\nthat cvsps remains the git project's choice is another week in which\nsomebody's project history is likely to get trashed.\n\nThis feels very strange and unpleasant.  I've never had to shoot one\nof my own projects through the head before.\n\nI blogged about it: http://esr.ibiblio.org/?p=5167\n\nIgnore the malware warning. It's triggered by something else on ibiblio.org;\nthey're fixing it.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"231944","messageId":"CACPiFC+bopf32cgDcQcVpL5vW=3KxmSP8Oh1see4KduQ1BNcPw@mail.gmail.com","threadId":"35524","inReplyTo":"20131212042624.GB8909@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-12-12T13:42:25Z","receivedAt":"2013-12-12T13:42:25Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Wed, Dec 11, 2013 at 11:26 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> You'll have to remind me what you mean by \"incremental\" here. Possibly\n> it's something cvs-fast-export could support.\n\nUser can\n\n - run a cvs to git import at time T, resulting in repo G\n - make commits to cvs repo\n - run cvs to git import at time T1, pointed to G, and the import tool\nwill only add the new commits found in cvs between T and T1.\n\n> But what I'm trying to tell you is that, even after I've done a dozen\n> releases and fixed the worst problems I could find, cvsps is far too\n> likely to mangle anything that passes through it.  The idea that you\n> are preserving *anything* valuable by sticking with it is a mirage.\n\nThe bugs that lead to a mangled history are real. I acknowledge and\nrespect that.\n\nHowever, with those limitations, the incremental feature has value in\nmany scenarios.\n\nThe two main ones are as follows:\n\n - A developer is tracking his/her own patches on top of a CVS-based\nproject with git. This is often done with git-svn for example. If\nold/convoluted branches in the far past are mangled, this user won't\ncare; as long as HEAD->master and/or the current/recent branch are\nconsistent with reality, the tool fits a need.\n\n - A project plans to transition to git gradually. Experienced\ndevelopers who'd normally work on CVS HEAD start working on git (and\nlanding their work on CVS afterwards). Old/mangled branches and tags\nare of little interest, the big value is CVS HEAD (which is linear)\nand possibly recent release/stable branches. The history captured is\ngood enough for git blame/log/pickaxe along the \"master\" line. At\ntransition time the original CVS repo can be kept around in readonly\nmode, so people can still checkout the exact contents of an old branch\nor tag for example (assuming no destructive \"surgery\" was done in the\nCVS repo).\n\nThe above examples assume that the CVS repos have used \"flying fish\"\napproach in the \"interesting\" (i.e.: recent) parts of their history.\n\n[ Simplifying a bit for non-CVS-geeks -- flying fish is using CVS HEAD\nfor your development, plus 'feature branches' that get landed, plus\nlong-lived 'stable release' branches. Most CVS projects in modern\ntimes use flying fish, which is a lot like what the git project uses\nin its own repo, but tuned to CVS's strengths (interesting commits\nlinearized in CVS HEAD).\n\nOther approaches ('dovetail') tend to end up with unworkable messes\ngiven CVS's weaknesses. ]\n\nThe cvsimport+cvsps combo does a reasonable (though imperfect) job on\n'flying fish' CVS histories _and that is what most projects evolved to\nuse_. If other cvs import tools can handle crazy histories, hats off\nto them. But careful with knifing cvsps!\n\ncheers,\n\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"231945","messageId":"20131212171756.GA6954@inner.h.apk.li","threadId":"35524","inReplyTo":"CACPiFC+bopf32cgDcQcVpL5vW=3KxmSP8Oh1see4KduQ1BNcPw@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Andreas Krey","fromEmail":"a.krey@gmx.de","sentAt":"2013-12-12T17:17:56Z","receivedAt":"2013-12-12T17:17:56Z","isPatch":false,"sender":{"key":"a.krey@gmx.de","avatar":"https://avatars.githubusercontent.com/u/37810?v=4"},"body":"On Thu, 12 Dec 2013 08:42:25 +0000, Martin Langhoff wrote:\n...\n>  - run a cvs to git import at time T, resulting in repo G\n>  - make commits to cvs repo\n>  - run cvs to git import at time T1, pointed to G, and the import tool\n> will only add the new commits found in cvs between T and T1.\n\nI'm pretty sure that being given only G the incremental approach wouldn't\nwork - some extra state would be required.\n\nBut anyway, the replacement question is a) how fast the cvs-fast-export is\nand b) whether its output is stable, that is, if the cvs repo C yields\na git repo G, will then C with a few extra commits yield G' where every\ncommit in G (as identified by its SHA1) is also in G', and G' additionally\ncontains the new commits that were made to the CVS repo.\n\nIf that is the case you effectively have an incremental mode, except that\nit's not quite as fast.\n\nAt least that would be good enough for us - we ended up running a\nfilter-branch on the resulting history, and that takes some time anyway.\n\n...\n> The cvsimport+cvsps combo does a reasonable (though imperfect) job on\n> 'flying fish' CVS histories _and that is what most projects evolved to\n> use_. If other cvs import tools can handle crazy histories, hats off\n> to them. But careful with knifing cvsps!\n\nIt won't magically disappear from your machine, and you have been warned. :-)\n\nAndreas\n\n-- \n\"Totally trivial. Famous last words.\"\nFrom: Linus Torvalds <torvalds@*.org>\nDate: Fri, 22 Jan 2010 07:29:21 -0800\n"},{"id":"231946","messageId":"CACPiFC+CAhh1S6Wt0bO2pDrXWgBW7AEFahGJjq0W7rG9LfTb8A@mail.gmail.com","threadId":"35524","inReplyTo":"20131212171756.GA6954@inner.h.apk.li","subject":"Re: I have end-of-lifed cvsps","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-12-12T17:26:40Z","receivedAt":"2013-12-12T17:26:40Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Thu, Dec 12, 2013 at 12:17 PM, Andreas Krey <a.krey@gmx.de> wrote:\n> But anyway, the replacement question is a) how fast the cvs-fast-export is\n> and b) whether its output is stable\n\nIn my prior work, the \"better\" CVS importers would not have stable\noutput, so were not appropriate for incremental imports.\n\nAnd even the fastest ones were very slow on large repos.\n\nThat is why I am asking the question.\n\n> It won't magically disappear from your machine, and you have been warned. :-)\n\nHowever, esr is making the case that git-cvsimport should stop using\ncvsps. My questions are aimed at understanding whether this actually\nresults in proposing that an important feature is dropped.\n\nPerhaps a better alternative is now available.\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"231947","messageId":"20131212181513.GA16960@thyrsus.com","threadId":"35524","inReplyTo":"CACPiFC+bopf32cgDcQcVpL5vW=3KxmSP8Oh1see4KduQ1BNcPw@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-12T18:15:13Z","receivedAt":"2013-12-12T18:15:13Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com>:\n> On Wed, Dec 11, 2013 at 11:26 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> > You'll have to remind me what you mean by \"incremental\" here. Possibly\n> > it's something cvs-fast-export could support.\n> \n> User can\n> \n>  - run a cvs to git import at time T, resulting in repo G\n>  - make commits to cvs repo\n>  - run cvs to git import at time T1, pointed to G, and the import tool\n> will only add the new commits found in cvs between T and T1.\n\nNo, cvs-fast-export doesn't do that. However, it is fast enough that\nyou can probably just rebuild the whole repo each time you want to\nmove content. \n\nWhen I did the conversion of groff recently I was getting rates of\nabout 150 commits a second - and it will be faster now, because I\nfound an expensive operation in the output stage I could optimize\nout.\n\nNow that you have reminded me of this, I remember implementing a -i\noption for cvsps-3.0 that could be combined with a time restriction \nto output incremental dumps. It's likely I could do the same\nthing for cvs-fast-import.\n\n> The above examples assume that the CVS repos have used \"flying fish\"\n> approach in the \"interesting\" (i.e.: recent) parts of their history.\n> \n> [ Simplifying a bit for non-CVS-geeks -- flying fish is using CVS HEAD\n> for your development, plus 'feature branches' that get landed, plus\n> long-lived 'stable release' branches. Most CVS projects in modern\n> times use flying fish, which is a lot like what the git project uses\n> in its own repo, but tuned to CVS's strengths (interesting commits\n> linearized in CVS HEAD).\n> \n> Other approaches ('dovetail') tend to end up with unworkable messes\n> given CVS's weaknesses. ]\n\nThat terminology -- \"flying fish\" and \"dovetail\" -- is interesting, and\nI have not heard it before.  It might be woth putting in the Jargon File.\nCan you point me at examples of live usage?\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"231948","messageId":"20131212182932.GB16960@thyrsus.com","threadId":"35524","inReplyTo":"20131212171756.GA6954@inner.h.apk.li","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-12T18:29:32Z","receivedAt":"2013-12-12T18:29:32Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Andreas Krey <a.krey@gmx.de>:\n> But anyway, the replacement question is a) how fast the cvs-fast-export is\n> and b) whether its output is stable, that is, if the cvs repo C yields\n> a git repo G, will then C with a few extra commits yield G' where every\n> commit in G (as identified by its SHA1) is also in G', and G' additionally\n> contains the new commits that were made to the CVS repo.\n> \n> If that is the case you effectively have an incremental mode, except that\n> it's not quite as fast.\n\nI am almost certain the output of cvs-fast-export is stable.  I\nbelieve the output of cvsps-3.x was, too.  Not sure about 2.x.\n\nI wrote the output stages for both cvsps-3.x and cvs-fast-export, and\nwent to some effort to verify that they write streams in the same\n\"most natural\" way - marks sequential from :1, blobs always witten as\nlate as possible, fileops in the same sort order the git tools emit,\netc.\n\nI have added writing a regression test test to verify the stability\nproperty to the TODO list. I will have this nailed down before the\nnext point release, in a few days.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"231950","messageId":"20131212183533.GC16960@thyrsus.com","threadId":"35524","inReplyTo":"CACPiFC+CAhh1S6Wt0bO2pDrXWgBW7AEFahGJjq0W7rG9LfTb8A@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-12T18:35:33Z","receivedAt":"2013-12-12T18:35:33Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com>:\n> In my prior work, the \"better\" CVS importers would not have stable\n> output, so were not appropriate for incremental imports.\n\nThat is disturbing.  I would consider lack of stability a severe and\nunacceptable failure mode in such a tool, if only because of the\ndifficulties it creates for proper regression testing.\n\nIf cvs-fast-export does not already have this property I will fix it \nso it does.  And document that fact.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"231951","messageId":"CACPiFCLxC-WkiiwXwLTv4s-1GtbX7GrNVGs94Z10Nz+LW8YCEQ@mail.gmail.com","threadId":"35524","inReplyTo":"20131212181513.GA16960@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-12-12T18:53:27Z","receivedAt":"2013-12-12T18:53:27Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Thu, Dec 12, 2013 at 1:15 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> That terminology -- \"flying fish\" and \"dovetail\" -- is interesting, and\n> I have not heard it before.  It might be woth putting in the Jargon File.\n> Can you point me at examples of live usage?\n\nThe canonical reference would be\nhttp://cvsbook.red-bean.com/cvsbook.html#Going%20Out%20On%20A%20Limb%20(How%20To%20Work%20With%20Branches%20And%20Survive)\n\njust by being on the internet and widely referenced it has probably\neclipsed in google-juice examples of earlier usage. Karl Fogel may\nremember where he got the names from.\n\ncheers,\n\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"231953","messageId":"CACPiFCJ22xiedXAoQktMLd=gASgD0NS24Pya9TvCo9aQP5JaBQ@mail.gmail.com","threadId":"35524","inReplyTo":"20131212182932.GB16960@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-12-12T19:08:33Z","receivedAt":"2013-12-12T19:08:33Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Thu, Dec 12, 2013 at 1:29 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> I am almost certain the output of cvs-fast-export is stable.  I\n> believe the output of cvsps-3.x was, too.  Not sure about 2.x.\n\nIIRC, making the output stable is nontrivial, specially on branches.\nTwo cases are still in my mind, from when I was wrestling with cvsps.\n\n1 - For a history with CVS HEAD and a long-running \"stable release\"\nbranch (\"STABLE\"), which branched at P1...\n\n   a - adding a file only at the tip of STABLE \"retroactively changes\nhistory\"  for P1 and perhaps CVS HEAD\n\n   b - forgetting to properly tag a subset of files with the branch\ntag, and doing it later retroactively changes history\n\n2 - you can create a new branch or tag with files that do not belong\ntogether in any \"commit\". Doing so changes history retroactively\n\n... when I say \"changes history\", I mean that the importers I know\nrevise their guesses of what files were seen together in a 'commit'.\nThis is specially true for history recorded with early cvs versions\nthat did not record a 'commit id'.\n\ncvsps has the strange \"feature\" that it will cache its\nassumptions/guesses, and continue incrementally from there. So if a\nchange in the CVS repo means that the old guess is now invalidated, it\ncontinues the charade instead of forcing a complete rewrite of the git\nhistory.\n\nMaybe the current crop of tools have developed stronger magic than\nwhat was available a few years ago... the task did seem impossible to\nme.\n\ncheers,\n\n\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"231957","messageId":"20131212193918.GA17529@thyrsus.com","threadId":"35524","inReplyTo":"CACPiFCJ22xiedXAoQktMLd=gASgD0NS24Pya9TvCo9aQP5JaBQ@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-12T19:39:18Z","receivedAt":"2013-12-12T19:39:18Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com>:\n> IIRC, making the output stable is nontrivial, specially on branches.\n> Two cases are still in my mind, from when I was wrestling with cvsps.\n> \n> 1 - For a history with CVS HEAD and a long-running \"stable release\"\n> branch (\"STABLE\"), which branched at P1...\n> \n>    a - adding a file only at the tip of STABLE \"retroactively changes\n> history\"  for P1 and perhaps CVS HEAD\n> \n>    b - forgetting to properly tag a subset of files with the branch\n> tag, and doing it later retroactively changes history\n> \n> 2 - you can create a new branch or tag with files that do not belong\n> together in any \"commit\". Doing so changes history retroactively\n> \n> ... when I say \"changes history\", I mean that the importers I know\n> revise their guesses of what files were seen together in a 'commit'.\n> This is specially true for history recorded with early cvs versions\n> that did not record a 'commit id'.\n\nYikes!  That is a much stricter stability criterion than I thought you\nwere specifying.   No, cvs-fast-export probably doesn't satify all of these.\nI think it would handle 1a in a stable way, but 1b and 2 would throw it.\n\nI'm sure it can't be fooled in the presence of commitids, though,\nbecause when it has those it doesn't try to do any similarity\nmatching.  And (this is the important point) it won't match any change\nwith a commit-id to any change without one.\n\nWhat I think this means is that cvs-fast-export is stable if you are\nusing a server/client combination that generates commitids (that is,\nGNU CVS of any version newer than 1.12 of 2004, or CVS-NT). It is\n*not* necessary for stability that the entire history have them.\n\nHere's how the logic works out:\n\n1. Commits grouped by commitid are stable - nothing in CVS ever rewrites\nthose or assigns a duplicate.\n\n2. No file change made with a commitid can destabilize a commit guess\nmade without them, because the similarity checker never tries to put both \nkinds in a single changeset.\n\nCan you detect any flaw in this?\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"231958","messageId":"CACPiFCLXeK9DH=f80ReSmYHJ7zjOn-D2zvs3WmdiV-k=wBGgjA@mail.gmail.com","threadId":"35524","inReplyTo":"20131212193918.GA17529@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-12-12T19:48:44Z","receivedAt":"2013-12-12T19:48:44Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Thu, Dec 12, 2013 at 2:39 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> Yikes!  That is a much stricter stability criterion than I thought you\n> were specifying.\n\n:-) -- cvsps's approach is: if you have a cache, you can remember the\nlies you told earlier.\n\nIt is impossible to be stable purely from the source data in the face\nof these issues.\n\nCVS is truly a PoS.\n\n> I think it would handle 1a in a stable way\n\nthat is pretty important. Files added on a branch not affecting HEAD\nand earlier branch checkout matters.\n\n\n> What I think this means is that cvs-fast-export is stable if you are\n> using a server/client combination that generates commitids (that is,\n> GNU CVS of any version newer than 1.12 of 2004, or CVS-NT). It is\n> *not* necessary for stability that the entire history have them.\n>\n> Here's how the logic works out:\n>\n> 1. Commits grouped by commitid are stable - nothing in CVS ever rewrites\n> those or assigns a duplicate.\n>\n> 2. No file change made with a commitid can destabilize a commit guess\n> made without them, because the similarity checker never tries to put both\n> kinds in a single changeset.\n>\n> Can you detect any flaw in this?\n\nIf someone creates a nonsensical tag or branch point, tagging files\nfrom different commits, how do you handle it?\n\n - without commit ids, does it affect your guesses?\n\n - regardless of commit ids, do you synthesize an artificial commit?\nHow do you define parenthood for that artificial commit?\n\ncurious,\n\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"231961","messageId":"20131212205819.GA18166@thyrsus.com","threadId":"35524","inReplyTo":"CACPiFCLXeK9DH=f80ReSmYHJ7zjOn-D2zvs3WmdiV-k=wBGgjA@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-12T20:58:19Z","receivedAt":"2013-12-12T20:58:19Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com>:\n> If someone creates a nonsensical tag or branch point, tagging files\n> from different commits, how do you handle it?\n> \n>  - without commit ids, does it affect your guesses?\n\nNo.  Tagging is never used to deduce changesets. Look:\n\n/*\n * The heart of the merge operation; detect when two\n * commits are \"the same\"\n */\nstatic bool\nrev_commit_match (rev_commit *a, rev_commit *b)\n{\n    /*\n     * Versions of GNU CVS after 1.12 (2004) place a commitid in\n     * each commit to track patch sets. Use it if present\n     */\n    if (a->commitid && b->commitid)\n\treturn a->commitid == b->commitid;\n    if (a->commitid || b->commitid)\n\treturn false;\n    if (!commit_time_close (a->date, b->date))\n\treturn false;\n    if (a->log != b->log)\n\treturn false;\n    if (a->author != b->author)\n\treturn false;\n    return true;\n}\n\n>  - regardless of commit ids, do you synthesize an artificial commit?\n> How do you define parenthood for that artificial commit?\n\nBecause tagging is never used to deduce changesets, the case does not arise.\n\nI have added an item to my to-do: document what the tool does with\ninconsistent tags.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"231964","messageId":"CACPiFCJDP6OVju2xzm2NWR5gc=bZDeNmXsD_MFH2mgHQru_u6Q@mail.gmail.com","threadId":"35524","inReplyTo":"20131212205819.GA18166@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-12-12T22:51:13Z","receivedAt":"2013-12-12T22:51:13Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Thu, Dec 12, 2013 at 3:58 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n>>  - regardless of commit ids, do you synthesize an artificial commit?\n>> How do you define parenthood for that artificial commit?\n>\n> Because tagging is never used to deduce changesets, the case does not arise.\n\nSo if a branch has a nonsensical branching point, or a tag is\nnonsensical, is it ignored and not imported?\n\ncurious,\n\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"231965","messageId":"20131212230454.GA20054@thyrsus.com","threadId":"35524","inReplyTo":"CACPiFCJDP6OVju2xzm2NWR5gc=bZDeNmXsD_MFH2mgHQru_u6Q@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-12T23:04:54Z","receivedAt":"2013-12-12T23:04:54Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com>:\n> On Thu, Dec 12, 2013 at 3:58 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> >>  - regardless of commit ids, do you synthesize an artificial commit?\n> >> How do you define parenthood for that artificial commit?\n> >\n> > Because tagging is never used to deduce changesets, the case does not arise.\n> \n> So if a branch has a nonsensical branching point, or a tag is\n> nonsensical, is it ignored and not imported?\n\nI don't know what happens when identically-named tags point at changes that\nresolve into two different commits.  I will figure that out and document it.\n\nThere's evidence, in the form of some code that is #ifdefed out, that \nKeith considered trying to make synthetic commits from tag cliques. But\nabandoned the idea because he couldn't figure out how to assign such\ncliques to a branch.\n\nI'm not sure what counts as a nonsensical branching point. I do know that\nKeith left this rather cryptic note in a REAME:\n\n\tDisjoint branch resolution. Branches occurring in a subset of the\n\tfiles are not correctly resolved; instead, an entirely disjoint\n\thistory will be created containing the branch revisions and all\n\tparents back to the root. I'm not sure how to fix this; it seems\n\tto implicitly assume there will be only a single place to attach as\n\tbranch parent, which may not be the case. In any case, the right\n\trevision will have a superset of the revisions present in the\n\toriginal branch parent; perhaps that will suffice.\n\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"231971","messageId":"CACPiFCLJfefNaPYtCXd21fO-ztmaaX6xdz1TtcMdYc0y19t56g@mail.gmail.com","threadId":"35524","inReplyTo":"20131212230454.GA20054@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-12-13T02:35:54Z","receivedAt":"2013-12-13T02:35:54Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Thu, Dec 12, 2013 at 6:04 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> I'm not sure what counts as a nonsensical branching point. I do know that\n> Keith left this rather cryptic note in a REAME:\n\nKeith names exactly what we are talking about. At that time, Keith was\nstruggling with the old xorg cvs repo which these and quite a few\nother nasties. I was also struggling with the mozilla cvs repo with\nits own gremlins.\n\nBetween my earlier explanation and Keith's notes it should be clear to\nyou. It is absolutely trivial in CVS to have an \"inconsistent\"\ncheckout (for example, if you switch branch with the -l parameter\ndisabling recursion, or if you accidentally switch branch in a\nsubdirectory).\n\nOn that inconsistent checkout, nothing prevents you from tagging it,\nnor from creating a new branch.\n\nAn importer with a 'consistent tree mentality' will look at the\nfiles/revs involved in that tag (or branching point) and find no tree\nto match.\n\nCVS repos with that crap exist. x11/xorg did (Jim Gettys challenged me\nto try importing it at an LCA, after the Bazaar NG folks passed on\nit). Mozilla did as well.\n\n\nIMHO it is a valid path to skip importing the tag/branch. As long as\nmain dev work was in HEAD, things end up ok (which goes back to my\nflying fish notes).\n\ncheers,\n\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"231972","messageId":"20131213033833.GB20850@thyrsus.com","threadId":"35524","inReplyTo":"CACPiFCLJfefNaPYtCXd21fO-ztmaaX6xdz1TtcMdYc0y19t56g@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-13T03:38:34Z","receivedAt":"2013-12-13T03:38:34Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Martin Langhoff <martin.langhoff@gmail.com>:\n> On Thu, Dec 12, 2013 at 6:04 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> > I'm not sure what counts as a nonsensical branching point. I do know that\n> > Keith left this rather cryptic note in a REAME:\n> \n> Keith names exactly what we are talking about.\n\nOh, yeah, I figured that much out.  What I wasn't clear on was (a) whether\nthat's a complete description of \"nonsensical branching point\" or whether there\nare other pathologies fundamentally *different* from that one.\n\nI'm also not sure I have the end state of what cvs-fast-export does in that\ncase visualized correctly. When he says: \"an entirely disjoint history will\nbe created containing the branch revisions and all parents back to the\nroot\", I'm visualizing something like this:\n\n  a----b----c----d----e----f----g----h\n                  \\\n                   +----1----2----3---4\n\nSuppose the root is a our pathological branch point is at d, then it\nsounds like he's saying cvs-fast-export will produce a changeset DAG\nthat looks like this:\n\n  a----b'---c'---d'---e----f----g----h\n   \\\n    +----b''---c''---d''----1----2----3----4\n\nWhat I'm not clear on here is how b is related to b' and b'', c to c' and c'',\nand d to d' and d''.  Which file changes go to which commit?  I shall have to\ncraft some broken RCS files to find out.\n\nHave I explained that I'm building a test suite?  I intend to know exactly\nwhat the tool does in these cases and document it.\n\n> Between my earlier explanation and Keith's notes it should be clear to\n> you. It is absolutely trivial in CVS to have an \"inconsistent\"\n> checkout (for example, if you switch branch with the -l parameter\n> disabling recursion, or if you accidentally switch branch in a\n> subdirectory).\n\nThat last one sounds easy to fall into and nasty. \n\n> On that inconsistent checkout, nothing prevents you from tagging it,\n> nor from creating a new branch.\n> \n> An importer with a 'consistent tree mentality' will look at the\n> files/revs involved in that tag (or branching point) and find no tree\n> to match.\n> \n> CVS repos with that crap exist. x11/xorg did (Jim Gettys challenged me\n> to try importing it at an LCA, after the Bazaar NG folks passed on\n> it). Mozilla did as well.\n> \n> \n> IMHO it is a valid path to skip importing the tag/branch. As long as\n> main dev work was in HEAD, things end up ok (which goes back to my\n> flying fish notes).\n\nThe other way to handle it would be to translate the history as though every\nbranch of a file subset had been an attempt to branch eveything.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232094","messageId":"52B02DFF.5010408@gmail.com","threadId":"35524","inReplyTo":"CACPiFC+bopf32cgDcQcVpL5vW=3KxmSP8Oh1see4KduQ1BNcPw@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2013-12-17T10:57:03Z","receivedAt":"2013-12-17T10:57:03Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Martin Langhoff wrote:\n\n> On Wed, Dec 11, 2013 at 11:26 PM, Eric S. Raymond<esr@thyrsus.com>  wrote:\n>> You'll have to remind me what you mean by \"incremental\" here. Possibly\n>> it's something cvs-fast-export could support.\n>\n> User can\n>\n>   - run a cvs to git import at time T, resulting in repo G\n>   - make commits to cvs repo\n>   - run cvs to git import at time T1, pointed to G, and the import tool\n >\n> will only add the new commits found in cvs between T and T1.\n\nI wonder if we can add support for incremental import once, for all\nVCS supporting fast-export, in one place, namely at the remote-helper.\n\nI don't know details, so I don't know if it is possible; certainly\nunstable fast-export output would be a problem, unless some tricks\nare used (like remembering mappings between versions).\n\n-- \nJakub Narębski\n"},{"id":"232095","messageId":"CALKQrgf3kuXRpbWmSp_nk8+zDFYNzkgV+dSBHaBbmUkxqjaDUA@mail.gmail.com","threadId":"35524","inReplyTo":"52B02DFF.5010408@gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2013-12-17T11:18:28Z","receivedAt":"2013-12-17T11:18:28Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Tue, Dec 17, 2013 at 11:57 AM, Jakub Narębski <jnareb@gmail.com> wrote:\n> Martin Langhoff wrote:\n>\n>> On Wed, Dec 11, 2013 at 11:26 PM, Eric S. Raymond<esr@thyrsus.com>  wrote:\n>>>\n>>> You'll have to remind me what you mean by \"incremental\" here. Possibly\n>>> it's something cvs-fast-export could support.\n>>\n>>\n>> User can\n>>\n>>   - run a cvs to git import at time T, resulting in repo G\n>>   - make commits to cvs repo\n>>   - run cvs to git import at time T1, pointed to G, and the import tool\n>\n>>\n>>\n>> will only add the new commits found in cvs between T and T1.\n>\n>\n> I wonder if we can add support for incremental import once, for all\n> VCS supporting fast-export, in one place, namely at the remote-helper.\n>\n> I don't know details, so I don't know if it is possible; certainly\n> unstable fast-export output would be a problem, unless some tricks\n> are used (like remembering mappings between versions).\n\nYou could do this by mapping some CVS revision identifier (like a hash\nover the file:revision pairs if nothing better is available), and that\nwould be useful when trying to match up the git commit from a later\nimport against the existing commits from an earlier import.\n\nHOWEVER, this only solves the \"cheap\" half of the problem. The reason\npeople want incremental CVS import, is to avoid having to repeatedly\nconvert the ENTIRE CVS history. This means that the CVS exporter must\nlearn to start from a given point in the CVS history (identified by\nthe above mapping) and then quickly and efficiently convert only the\n\"new stuff\" without having to consult/convert the rest of the CVS\nhistory. THIS is the hard part of incremental import. And it is much\nharder for systems like CVS - where the starting point has a broken\nconcept of history...\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"232103","messageId":"20131217140746.GB15010@thyrsus.com","threadId":"35524","inReplyTo":"52B02DFF.5010408@gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-17T14:07:46Z","receivedAt":"2013-12-17T14:07:46Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jakub Narębski <jnareb@gmail.com>:\n> I wonder if we can add support for incremental import once, for all\n> VCS supporting fast-export, in one place, namely at the remote-helper.\n\nSomething in the pipeline - either the helper or the exporter - needs to\nhave an equivalent of vc-fast-export's and cvsps's -i option, which\nomits all commits before a specified time and generates cookies like\n\"from refs/heads/master^0\" before each branch root in the incremental\ndump.\n\nThis could be done in the wrapper, but only if the wrapper itself\nincludes an import-stream parser, interprets the output from the\nexporter program, and re-emits it.  Having done similar things\nmyself in reposurgeon, I advise against this strategy; it would\nintroduce a level of complexity to the wrapper that doesn't belong\nthere, and make the exporter+wrapper comnination harder to verify.\n\nFortunately, incremental dump is trivial to implement in the output\nstage of an exporter if you have access to the exporter source code.\nI've done it in two different exporters.  cvs-fast-export now has a\nregression test for this case\n\n> I don't know details, so I don't know if it is possible; certainly\n> unstable fast-export output would be a problem, unless some tricks\n> are used (like remembering mappings between versions).\n\nAbout such tricks I can only say \"That way lies madness\".  The present\nPerl wrapper is buggy because it's over-complex.  The replacement wrapper\nshould do *less*, not more.\n\nStable output and incremental dump are reasonable things to demand of\nyour supported exporters.  cvs-fast-export has incremental dump\nunconditionally, and stability relative to every CVS implementation\nsince 2004.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232104","messageId":"20131217145809.GC15010@thyrsus.com","threadId":"35524","inReplyTo":"CALKQrgf3kuXRpbWmSp_nk8+zDFYNzkgV+dSBHaBbmUkxqjaDUA@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-17T14:58:09Z","receivedAt":"2013-12-17T14:58:09Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Johan Herland <johan@herland.net>:\n> HOWEVER, this only solves the \"cheap\" half of the problem. The reason\n> people want incremental CVS import, is to avoid having to repeatedly\n> convert the ENTIRE CVS history. This means that the CVS exporter must\n> learn to start from a given point in the CVS history (identified by\n> the above mapping) and then quickly and efficiently convert only the\n> \"new stuff\" without having to consult/convert the rest of the CVS\n> history. THIS is the hard part of incremental import. And it is much\n> harder for systems like CVS - where the starting point has a broken\n> concept of history...\n\nI know of *no* importer that solves what you call the \"deep\" part of\nthe problem.  cvsps didn't, cvs-fast-import doesn't, cvs2git doesn't.\nAll take the easy way out; parse the entire history, and limit what\nis emitted in the output stage.\n\nActually, given what I know about delta-file parsing I'd say a \"true\"\nincremental CVS exporter would be so hard that it's really not worth the\nbother.  The problem is the delta-based history representation.\nTrying to interpret that without building a complete set of history\nstates in the process (which is most of the work a whole-history\nexporter does) would be brutally difficult - barely possible in\nprinciple maybe, but I wouldn't care to try it.\n\nIt's much more practical to tune up a whole-history exporter so it's\nacceptably fast, then do incremental dumping by suppressing part of\nthe conversion in the output stage. \n\ncvs-fast-export's benchmark repo is the history of GNU troff.  That's\n3057 commits in 1549 master files; when I reran it just now the\nwhole-history conversion took 49 seconds.  That's 3.7K commits a\nminute, which is plenty fast enough for anything smaller than (say)\none of the *BSD repositories.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232106","messageId":"CALKQrgeegcsO7YVqEmQxD4=HfR4eitodAov0tEh7MRvBxtRKUA@mail.gmail.com","threadId":"35524","inReplyTo":"20131217145809.GC15010@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2013-12-17T17:52:09Z","receivedAt":"2013-12-17T17:52:09Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Tue, Dec 17, 2013 at 3:58 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> Johan Herland <johan@herland.net>:\n>> HOWEVER, this only solves the \"cheap\" half of the problem. The reason\n>> people want incremental CVS import, is to avoid having to repeatedly\n>> convert the ENTIRE CVS history. This means that the CVS exporter must\n>> learn to start from a given point in the CVS history (identified by\n>> the above mapping) and then quickly and efficiently convert only the\n>> \"new stuff\" without having to consult/convert the rest of the CVS\n>> history. THIS is the hard part of incremental import. And it is much\n>> harder for systems like CVS - where the starting point has a broken\n>> concept of history...\n>\n> I know of *no* importer that solves what you call the \"deep\" part of\n> the problem.  cvsps didn't, cvs-fast-import doesn't, cvs2git doesn't.\n> All take the easy way out; parse the entire history, and limit what\n> is emitted in the output stage.\n\nYes, and starting from a non-incremental importer, that's probably the\nonly viable way to approach incrementalism.\n\n> Actually, given what I know about delta-file parsing I'd say a \"true\"\n> incremental CVS exporter would be so hard that it's really not worth the\n> bother.  The problem is the delta-based history representation.\n> Trying to interpret that without building a complete set of history\n> states in the process (which is most of the work a whole-history\n> exporter does) would be brutally difficult - barely possible in\n> principle maybe, but I wouldn't care to try it.\n\nAgreed, you would either have to re-parse the entire ,v-file, or you\nwould have to store some (probably a lot of) intermediate state that\nwould allow you to resolve deltas of new revisions without having to\nparse all the old revisions.\n\n> It's much more practical to tune up a whole-history exporter so it's\n> acceptably fast, then do incremental dumping by suppressing part of\n> the conversion in the output stage.\n>\n> cvs-fast-export's benchmark repo is the history of GNU troff.  That's\n> 3057 commits in 1549 master files; when I reran it just now the\n> whole-history conversion took 49 seconds.  That's 3.7K commits a\n> minute, which is plenty fast enough for anything smaller than (say)\n> one of the *BSD repositories.\n\nThose are impressive numbers, and in that scenario, using a\n\"repurposed\" converter (i.e. whole-history converter that has been\ntaught to do incremental output) is undoubtedly the best solution.\n\nHowever, I fear that you underestimate the number of users that want\nto use Git against CVS repos that are orders of magnitude larger (in\nboth dimensions: #commits and #files) than your example repo. For\nthese repos, running a proper whole-history conversion takes hours -\nor even days - and working incrementally on top of that is simply out\nof the question. Obviously, they still need the whole-history\nconverter for the future point in time when they have collected enough\nmotivation/buy-in to migrate the entire project/company to a better\nVCS, but until then, they want to use Git locally, while enduring CVS\non the server.\n\nAt my previous $DAYJOB, I was one of those people, and I ended up with\na two-pronged \"solution\" to the problem (this is ~5 years ago now, so\nI'm somewhat fuzzy on the details):\n\n 1. Adopt an ad hoc incremental approach for working against the CVS\nserver: Keep a CVS checkout next to my git repo. and maintain a map\nbetween corresponding states/commits in CVS and git. When I update\nfrom CVS, apply the corresponding patch to the \"cvs\" branch in my git\nrepo. Rebase my git-based work on top of that, and use \"git\ncvsexportcommit\" to propagate my Git work back to CVS. This is crude\nand hacky as hell, but it provides me a local git-based workflow.\n\n 2. Start convincing fellow developers and lobby management about\nswitching away from CVS. We got a discussion started, gained momentum,\nand eventually I got to spend most of my time preparing and performing\nthe full-history conversion from CVS to git. This happened mostly\nbefore cvs2svn grew its cvs2git sibling, so I ended up writing a\ncustom converter for our particular variation of insane and demented\nCVS practices. Today, I would probably have gone for cvs2git, or your\nmore recent work.\n\nBut back to my main point:\n\nI believe there are two classes of CVS converters, and I have slowly\ncome to believe that they solve two fundamentally different problems.\nThe first problem is \"how to faithfully recreate the project history\nin a different VCS\", which is solved by the full-history converters.\nCase closed.\n\nThe second problem is somewhat harder to define, but I'll try: \"how to\nallow me to work productively against a CVS server, without having to\ndeal with the icky CVS bits\". Compared to the first problem, the\nparameters differ somewhet:\n\n - Conversion/synchronization time must be short to allow me to stay\nproductive and up-to-date with my colleagues.\n\n - Correctness of \"current state\" is very important. I must be sure\nthat my git working tree is identical to its CVS counterpart, so that\nmy git changes can be reproduced in CVS as faithfully as possible.\n\n - Correctness of \"history\" is less important. I can accept a\nmessy/incorrect Git history, since I can always query the CVS server\nfor the \"correct\" history (whatever that means in a CVS context...).\n\n - As a generic CVS user (not the CVS admin) I don't necessarily have\ndirect access to the ,v files stored on the CVS server.\n\nAlthough a full-history converter with fairly stable output can be\nmade to support this second problem for repos up to a certain size,\nthere will probably still be users that want to work incrementally\nagainst much bigger repos, and I don't think _any_\nfull-history-gone-incremental importer will be able to support the\nbiggest repos.\n\nConsequently I believe that for these big repos it is _impossible_ to\nget both fast incremental workflows and a high degree of (historical)\ncorrectness.\n\ncvsps tried to be all of the above, and failed badly at the\ncorrectness criteria. Therefore I support your decision to \"shoot it\nthrough the head\". I certainly also support any work towards making a\nfull-history converter work in an incremental manner, as it will be\nimmensely useful for smaller CVS repos. But at the same time we should\nrealize that it won't be a solution for incrementally working against\n_large_ CVS repos.\n\nAlthough it should have been made obvious a long time ago, the removal\nof cvsps has now made it abundantly clear that Git currently provides\nno way to support the incremental workflow against large CVS repos.\nMaybe that is ok, and we can ignore that, waiting for the few\nremaining large CVS repos to die? Or maybe we need a new effort to\nfill this niche? Something that is NOT based on a full-history\nconverter, and does NOT try to guarantee a history-correct conversion,\nbut that DOES try to guarantee fast and relatively worry-free two-way\nsynchronization against a CVS server. Unfortunately (or fortunately,\ndepending on POV) I have not had to touch CVS in a long while, and I\ndon't see that changing soon, so it is not my itch to scratch.\n\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"232112","messageId":"20131217184724.GA17709@thyrsus.com","threadId":"35524","inReplyTo":"CALKQrgeegcsO7YVqEmQxD4=HfR4eitodAov0tEh7MRvBxtRKUA@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-17T18:47:24Z","receivedAt":"2013-12-17T18:47:24Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Johan Herland <johan@herland.net>:\n> However, I fear that you underestimate the number of users that want\n> to use Git against CVS repos that are orders of magnitude larger (in\n> both dimensions: #commits and #files) than your example repo.\n\nYou may be right. See below...\n\nI'm working with Alan Barret now on trying to convert the NetBSD\nrepositories. They break cvs-fast-export through sheer bulk of\nmetadata, by running the machine out of core.  This is exactly\nthe kind of huge case that you're talking about.\n\nAlan and I are going to take a good hard whack at modifying cvs-fast-export \nto make this work. Because there really aren't any feasible alternatives.\nThe analysis code in cvsps was never good enough. cvs2git, being written\nin Python, would hit the core limit faster than anything written in C.\n\n> Although a full-history converter with fairly stable output can be\n> made to support this second problem for repos up to a certain size,\n> there will probably still be users that want to work incrementally\n> against much bigger repos, and I don't think _any_\n> full-history-gone-incremental importer will be able to support the\n> biggest repos.\n> \n> Consequently I believe that for these big repos it is _impossible_ to\n> get both fast incremental workflows and a high degree of (historical)\n> correctness.\n> \n> cvsps tried to be all of the above, and failed badly at the\n> correctness criteria. Therefore I support your decision to \"shoot it\n> through the head\". I certainly also support any work towards making a\n> full-history converter work in an incremental manner, as it will be\n> immensely useful for smaller CVS repos. But at the same time we should\n> realize that it won't be a solution for incrementally working against\n> _large_ CVS repos.\n\nIt is certainly the case that a sufficiently large CVS repo will break\nanything, like a star with a mass over the Chandrasekhar limit becoming a \nblack hole :-)\n\nThe question is how common such supermassive cases are. My own guess is that\nthe *BSD repos and a handful of the oldest GNU projects are pretty much the\nwhole set; everybody else converted to Subversion within the last decade. \n \n> Although it should have been made obvious a long time ago, the removal\n> of cvsps has now made it abundantly clear that Git currently provides\n> no way to support the incremental workflow against large CVS repos.\n> Maybe that is ok, and we can ignore that, waiting for the few\n> remaining large CVS repos to die? Or maybe we need a new effort to\n> fill this niche? Something that is NOT based on a full-history\n> converter, and does NOT try to guarantee a history-correct conversion,\n> but that DOES try to guarantee fast and relatively worry-free two-way\n> synchronization against a CVS server. Unfortunately (or fortunately,\n> depending on POV) I have not had to touch CVS in a long while, and I\n> don't see that changing soon, so it is not my itch to scratch.\n\nNor mine.  I find the very idea of writing anything that encourages\nnon-history-correct conversions disturbing and want no part of it.\n\nWhich matters, because right now the set of people working on CVS lifters\nbegins with me and ends with Michael Rafferty (cvs2git), who seems even\nless interested in incremental conversion than I am.  Unless somebody\ncomes out of nowhere and wants to own that problem, it's not going\nto get solved.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232115","messageId":"CANQwDwe8AcbCYG5GZcY1tn9BN0x5KWux_CNQY2OWG+qZJ5rS4Q@mail.gmail.com","threadId":"35524","inReplyTo":"20131217140746.GB15010@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2013-12-17T19:58:18Z","receivedAt":"2013-12-17T19:58:18Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, Dec 17, 2013 at 3:07 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> Jakub Narębski <jnareb@gmail.com>:\n\n>> I wonder if we can add support for incremental import once, for all\n>> VCS supporting fast-export, in one place, namely at the remote-helper.\n>\n> Something in the pipeline - either the helper or the exporter - needs to\n> have an equivalent of vc-fast-export's and cvsps's -i option, which\n> omits all commits before a specified time and generates cookies like\n> \"from refs/heads/master^0\" before each branch root in the incremental\n> dump.\n\nErrr... doesn't cvs-fast-export support --export-marks=<file> to save\nprogress and --import-marks=<file> to continue incremental import?\nI *guess* that 'export' / 'import' capabilities-based remote helpers\nuse 'export-marks <file>' / 'import-marks <file>' capability for incremental\nimport, also known as \"fetch\", isn't it? But I might be mistaken, I don't\nknow enough about remote helpers...\n\nI would check it in cvs-fast-export manpage, but the page seems to\nbe down:\n\n  http://isup.me/www.catb.org\n\n    It's not just you! http://www.catb.org looks down from here.\n\n> This could be done in the wrapper, but only if the wrapper itself\n> includes an import-stream parser, interprets the output from the\n> exporter program, and re-emits it.  Having done similar things\n> myself in reposurgeon, I advise against this strategy; it would\n> introduce a level of complexity to the wrapper that doesn't belong\n> there, and make the exporter+wrapper combination harder to verify.\n\nRight.\n\n> Fortunately, incremental dump is trivial to implement in the output\n> stage of an exporter if you have access to the exporter source code.\n> I've done it in two different exporters.  cvs-fast-export now has a\n> regression test for this case\n\nThis is I guess assuming that information from later commits doesn't\nchange guesses about shape of history from earlier commits...\n\n-- \nJakub Narebski\n"},{"id":"232118","messageId":"20131217210255.GA18217@thyrsus.com","threadId":"35524","inReplyTo":"CANQwDwe8AcbCYG5GZcY1tn9BN0x5KWux_CNQY2OWG+qZJ5rS4Q@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-17T21:02:55Z","receivedAt":"2013-12-17T21:02:55Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jakub Narębski <jnareb@gmail.com>:\n> Errr... doesn't cvs-fast-export support --export-marks=<file> to save\n> progress and --import-marks=<file> to continue incremental import?\n\nNo, cvs-fast-export does not have --export-marks. It doesn't generate the\nSHA1s that would require. Even if it did, it's not clear how that would help.\n\n> I would check it in cvs-fast-export manpage, but the page seems to\n> be down:\n> \n>   http://isup.me/www.catb.org\n> \n>     It's not just you! http://www.catb.org looks down from here.\n\nConfirmed.  Looks like ibiblio is having a bad day.  I'll file a bug report. \n\n> > Fortunately, incremental dump is trivial to implement in the output\n> > stage of an exporter if you have access to the exporter source code.\n> > I've done it in two different exporters.  cvs-fast-export now has a\n> > regression test for this case\n> \n> This is I guess assuming that information from later commits doesn't\n> change guesses about shape of history from earlier commits...\n\nThat's the \"stability\" property that Martin Langhoff and I were discussing\nearlier.\n\ncvs-fast-export conversions are stable under incremental\nlifting providing a commitid-generating version of CVS is in use\nduring each increment.  Portions of the history *before the first\nlift* may lack commitids and will nevertheless remain stable through\nthe whole process.\n\nAll versions of CVS have generated commitids since 2004.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232119","messageId":"CALKQrgeRKosOSOhcbUArkh03mwJLPkcOH-DROCCnmbTdQ8afyg@mail.gmail.com","threadId":"35524","inReplyTo":"20131217184724.GA17709@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2013-12-17T21:26:57Z","receivedAt":"2013-12-17T21:26:57Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Tue, Dec 17, 2013 at 7:47 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> I'm working with Alan Barret now on trying to convert the NetBSD\n> repositories. They break cvs-fast-export through sheer bulk of\n> metadata, by running the machine out of core.  This is exactly\n> the kind of huge case that you're talking about.\n>\n> Alan and I are going to take a good hard whack at modifying cvs-fast-export\n> to make this work. Because there really aren't any feasible alternatives.\n> The analysis code in cvsps was never good enough. cvs2git, being written\n> in Python, would hit the core limit faster than anything written in C.\n\nDepends on how it organizes its data structures. Have you actually\ntried running cvs2git on it? I'm not saying you are wrong, but I had\nsimilar problems with my custom converter (also written in Python),\nand solved them by adding multiple passes/phases instead of trying to\ndo too much work in fewer passes. In the end I ended up storing the\nlargest inter-phase data structures outside of Python (sqlite in my\ncase) to save memory. Obviously it cost a lot in runtime, but it meant\nthat I could actually chew through our largest CVS modules without\nrunning out of memory.\n\n> It is certainly the case that a sufficiently large CVS repo will break\n> anything, like a star with a mass over the Chandrasekhar limit becoming a\n> black hole :-)\n\n:) True, although it's not the sheer size of the files themselves that\nis the actual problem. Most of those bytes are (deltified) file data,\nwhich you can pretty much stream through and convert to a\ncorresponding fast-export stream of blob objects. The code for that\nshould be fairly straightforward (and should also be eminently\nparallelizable, given enough cores and available I/O), resulting in a\ntable mapping CVS file:revision pairs to corresponding Git blob SHA1s,\nand an accompanying (set of) packfile(s) holding said blobs.\n\nThe hard part comes when trying to correlate the metadata for all the\nper-file revisions, and distill that into a consistent sequence/DAG of\nchangesets/commits across the entire CVS repo. And then, of course,\ntrying to fit all the branches and tags into that DAG of commits is\nwhat really drives you mad... ;-)\n\n> The question is how common such supermassive cases are. My own guess is that\n> the *BSD repos and a handful of the oldest GNU projects are pretty much the\n> whole set; everybody else converted to Subversion within the last decade.\n\nYou may be right. At least for the open-source cases. I suspect\nthere's still a considerable number of huge CVS repos within\ncompanies' walls...\n\n> I find the very idea of writing anything that encourages\n> non-history-correct conversions disturbing and want no part of it.\n>\n> Which matters, because right now the set of people working on CVS lifters\n> begins with me and ends with Michael Rafferty (cvs2git),\n\ns/Rafferty/Haggerty/?\n\n> who seems even\n> less interested in incremental conversion than I am.  Unless somebody\n> comes out of nowhere and wants to own that problem, it's not going\n> to get solved.\n\nAgreed. It would be nice to have something to point to for people that\nwant something similar to git-svn for CVS, but without a motivated\nowner, it won't happen.\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"232122","messageId":"20131217224136.GB19511@thyrsus.com","threadId":"35524","inReplyTo":"CALKQrgeRKosOSOhcbUArkh03mwJLPkcOH-DROCCnmbTdQ8afyg@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-17T22:41:36Z","receivedAt":"2013-12-17T22:41:36Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Johan Herland <johan@herland.net>:\n> > Alan and I are going to take a good hard whack at modifying cvs-fast-export\n> > to make this work. Because there really aren't any feasible alternatives.\n> > The analysis code in cvsps was never good enough. cvs2git, being written\n> > in Python, would hit the core limit faster than anything written in C.\n> \n> Depends on how it organizes its data structures. Have you actually\n> tried running cvs2git on it? I'm not saying you are wrong, but I had\n> similar problems with my custom converter (also written in Python),\n> and solved them by adding multiple passes/phases instead of trying to\n> do too much work in fewer passes. In the end I ended up storing the\n> largest inter-phase data structures outside of Python (sqlite in my\n> case) to save memory. Obviously it cost a lot in runtime, but it meant\n> that I could actually chew through our largest CVS modules without\n> running out of memory.\n\nYou make a good point.  cvs2git is descended from cvs2svn, which has\nsuch a multipass organization - it will only have to avoid memory\nlimits per pass.  Alan and I will try that as a fallback if\ncvs-fast-import continues to choke.\n \n> > It is certainly the case that a sufficiently large CVS repo will break\n> > anything, like a star with a mass over the Chandrasekhar limit becoming a\n> > black hole :-)\n> \n> :) True, although it's not the sheer size of the files themselves that\n> is the actual problem. Most of those bytes are (deltified) file data,\n> which you can pretty much stream through and convert to a\n> corresponding fast-export stream of blob objects. The code for that\n> should be fairly straightforward (and should also be eminently\n> parallelizable, given enough cores and available I/O), resulting in a\n> table mapping CVS file:revision pairs to corresponding Git blob SHA1s,\n> and an accompanying (set of) packfile(s) holding said blobs.\n\nAllowing for the fact that cvs-fast-export isn't git and doesn't use\nSHA1s or packfiles, this is in fact how a large portion of\ncvs-fast-export works.  The blob files get created during the walk\nthrough the master file list, before actual topo analysis is done.\n\n> The hard part comes when trying to correlate the metadata for all the\n> per-file revisions, and distill that into a consistent sequence/DAG of\n> changesets/commits across the entire CVS repo. And then, of course,\n> trying to fit all the branches and tags into that DAG of commits is\n> what really drives you mad... ;-)\n\nWell I know this...:-)\n\n> > The question is how common such supermassive cases are. My own guess is that\n> > the *BSD repos and a handful of the oldest GNU projects are pretty much the\n> > whole set; everybody else converted to Subversion within the last decade.\n> \n> You may be right. At least for the open-source cases. I suspect\n> there's still a considerable number of huge CVS repos within\n> companies' walls...\n\nIf people with money want to hire me to slay those beasts, I'm available.\nI'm not proud, I'll use cvs2git if I have to.\n \n> > I find the very idea of writing anything that encourages\n> > non-history-correct conversions disturbing and want no part of it.\n> >\n> > Which matters, because right now the set of people working on CVS lifters\n> > begins with me and ends with Michael Rafferty (cvs2git),\n> \n> s/Rafferty/Haggerty/?\n\nYup, I thinkoed.\n \n> > who seems even\n> > less interested in incremental conversion than I am.  Unless somebody\n> > comes out of nowhere and wants to own that problem, it's not going\n> > to get solved.\n> \n> Agreed. It would be nice to have something to point to for people that\n> want something similar to git-svn for CVS, but without a motivated\n> owner, it won't happen.\n\nI think the fact that it hasn't happened already is a good clue that\nit's not going to. Given the decline curve of CVS usage, writing \ngit-cvs might have looked like a decent investment of time once,\nbut that era probably ended five to eight years ago.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232128","messageId":"CANQwDwdQZGhR=hhFHe7wRAeNej_F5fHspN7+f-LiJu06utwC-w@mail.gmail.com","threadId":"35524","inReplyTo":"20131217210255.GA18217@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2013-12-18T00:02:04Z","receivedAt":"2013-12-18T00:02:04Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Tue, Dec 17, 2013 at 10:02 PM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> Jakub Narębski <jnareb@gmail.com>:\n>>\n>> Errr... doesn't cvs-fast-export support --export-marks=<file> to save\n>> progress and --import-marks=<file> to continue incremental import?\n>\n> No, cvs-fast-export does not have --export-marks. It doesn't generate the\n> SHA1s that would require. Even if it did, it's not clear how that would help.\n\nI was thinking about how the following part of git-fast-export\n`--import-marks=<file>`\n\n  Any commits that have already been marked will not be exported again.\n  If the backend uses a similar --import-marks file, this allows for incremental\n  bidirectional exporting of the repository by keeping the marks the same\n  across runs.\n\nHow cvs-fast-export know where to start exporting from in incremental mode?\n\nBTW. does cvs-fast-export support incremental *output*, or does it\nperform also incremental *work*?\n\nAnyway, that might mean that generic fast-import stream based incremental\n(i.e. supporting proper thin fetch) remote helper is out of question, perhaps\nwriting one for cvs / cvs-fe would bring incremental import from CVS to\ngit?\n\n-- \nJakub Narebski\n"},{"id":"232129","messageId":"87vbynnhwt.fsf@igel.home","threadId":"35524","inReplyTo":"20131217210255.GA18217@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Andreas Schwab","fromEmail":"schwab@linux-m68k.org","sentAt":"2013-12-18T00:04:34Z","receivedAt":"2013-12-18T00:04:34Z","isPatch":false,"sender":{"key":"schwab@linux-m68k.org","avatar":"https://avatars.githubusercontent.com/u/2175493?v=4"},"body":"\"Eric S. Raymond\" <esr@thyrsus.com> writes:\n\n> All versions of CVS have generated commitids since 2004.\n\nThough older versions are still in use, eg. sourceware.org still does\nnot generate commitids.\n\nAndreas.\n\n-- \nAndreas Schwab, schwab@linux-m68k.org\nGPG Key fingerprint = 58CA 54C7 6D53 942B 1756  01D3 44D5 214B 8276 4ED5\n\"And now for something completely different.\"\n"},{"id":"232130","messageId":"20131218002122.GA20152@thyrsus.com","threadId":"35524","inReplyTo":"CANQwDwdQZGhR=hhFHe7wRAeNej_F5fHspN7+f-LiJu06utwC-w@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-18T00:21:22Z","receivedAt":"2013-12-18T00:21:22Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jakub Narębski <jnareb@gmail.com>:\n> > No, cvs-fast-export does not have --export-marks. It doesn't generate the\n> > SHA1s that would require. Even if it did, it's not clear how that would help.\n> \n> I was thinking about how the following part of git-fast-export\n> `--import-marks=<file>`\n> \n>   Any commits that have already been marked will not be exported again.\n>   If the backend uses a similar --import-marks file, this allows for incremental\n>   bidirectional exporting of the repository by keeping the marks the same\n>   across runs.\n\nI understand that. But it's not relevant - cvs-fast-import doesn't know about\ngit SHA1s, and cannot.\n \n> How cvs-fast-export know where to start exporting from in incremental mode?\n\nYou give it a cutoff date. This is the same way cvsps-2.x and 3.x worked,\nand it's what the cvsimport wrapper expects to pass down.\n\n> BTW. does cvs-fast-export support incremental *output*, or does it\n> perform also incremental *work*?\n\nAs I tried to explain previously in my response to John Herland, it's\nincremental output only.  There is *no* CVS exporter known to me, or\nhim, that supports incremental work.  That would be at best be impractically\ndifficult; given CVS's limitations it may be actually impossible. I wouldn't\nbet against impossible.\n\n> Anyway, that might mean that generic fast-import stream based incremental\n> (i.e. supporting proper thin fetch) remote helper is out of question, perhaps\n> writing one for cvs / cvs-fe would bring incremental import from CVS to\n> git?\n\nSorry, I don't understand that.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232131","messageId":"20131218002545.GB20152@thyrsus.com","threadId":"35524","inReplyTo":"87vbynnhwt.fsf@igel.home","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-18T00:25:45Z","receivedAt":"2013-12-18T00:25:45Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Andreas Schwab <schwab@linux-m68k.org>:\n> \"Eric S. Raymond\" <esr@thyrsus.com> writes:\n> \n> > All versions of CVS have generated commitids since 2004.\n> \n> Though older versions are still in use, eg. sourceware.org still does\n> not generate commitids.\n\nThat is awful.  Alas, there is not much anyone can do about stupidity\nthat determined.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232165","messageId":"CANQwDwdgZUWcgyZCWoDni+e9jgQ+8j0Yn_HMxiMn5OHzsRzjwQ@mail.gmail.com","threadId":"35524","inReplyTo":"20131218002122.GA20152@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Jakub Narębski","fromEmail":"jnareb@gmail.com","sentAt":"2013-12-18T15:39:39Z","receivedAt":"2013-12-18T15:39:39Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"On Wed, Dec 18, 2013 at 1:21 AM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> Jakub Narębski <jnareb@gmail.com>:\n\n>>> No, cvs-fast-export does not have --export-marks. It doesn't generate the\n>>> SHA1s that would require. Even if it did, it's not clear how that would help.\n>>\n>> I was thinking about how the following part of git-fast-export\n>> `--import-marks=<file>`\n>>\n>>   Any commits that have already been marked will not be exported again.\n>>   If the backend uses a similar --import-marks file, this allows for incremental\n>>   bidirectional exporting of the repository by keeping the marks the same\n>>   across runs.\n>\n> I understand that. But it's not relevant - cvs-fast-import doesn't know about\n> git SHA1s, and cannot.\n\nIt is a bit strange that markfile has explicitly SHA-1 (\":markid <SHA-1>\"),\ninstead of generic reference to commit, in the case of CVS it would be\ncommitid (what to do for older repositories, though?), in case of Bazaar\nits revision id (GUID), etc.  Can we assume that SCM v1 fast-export and\nSCM v2 fast-import markfile uses compatibile commit names in markfile?\n\n>> How cvs-fast-export know where to start exporting from in incremental mode?\n>\n> You give it a cutoff date. This is the same way cvsps-2.x and 3.x worked,\n> and it's what the cvsimport wrapper expects to pass down.\n\nNice to know.\n\nI think it would be possible for remote-helper for cvs-fast-export to find\nthis cutoff date automatically (perhaps with some safety margin), for\nfetching (incremental import).\n\n>> BTW. does cvs-fast-export support incremental *output*, or does it\n>> perform also incremental *work*?\n>\n> As I tried to explain previously in my response to John Herland, it's\n> incremental output only.  There is *no* CVS exporter known to me, or\n> him, that supports incremental work.  That would be at best be impractically\n> difficult; given CVS's limitations it may be actually impossible. I wouldn't\n> bet against impossible.\n\nEven with saving (or re-calculating from git import) guesses about CVS\nhistory made so far?\n\nAnyway I hope that incremental CVS import would be needed less\nand less as CVS is replaced by any more modern version control system.\n\n>> Anyway, that might mean that generic fast-import stream based incremental\n>> (i.e. supporting proper thin fetch) remote helper is out of question, perhaps\n>> writing one for cvs / cvs-fe would bring incremental import from CVS to\n>> git?\n>\n> Sorry, I don't understand that.\n\nI was thinking about creating remote-helper for cvs-fast-export, so that\ngit can use local CVS repository as \"remote\", using e.g. \"cvsroot::<path>\"\nas repo URL, and using this mechanism for incremental import (aka fetch).\n(Or even \"cvssync::<URL>\" for automatic cvssync + cvs-fast-export).\n\nBut from what I understand this is not as easy as it seems, even with\nremote-helper API having support for fast-import stream.\n\n-- \nJakub Narębski\n"},{"id":"232167","messageId":"20131218162239.GA26668@google.com","threadId":"35524","inReplyTo":"CANQwDwdgZUWcgyZCWoDni+e9jgQ+8j0Yn_HMxiMn5OHzsRzjwQ@mail.gmail.com","subject":"incremental fast-import and marks (Re: I have end-of-lifed cvsps)","fromName":"Jonathan Nieder","fromEmail":"jrnieder@gmail.com","sentAt":"2013-12-18T16:23:34Z","receivedAt":"2013-12-18T16:23:34Z","isPatch":false,"sender":{"key":"jrnieder@gmail.com","avatar":"https://avatars.githubusercontent.com/u/281595?v=4"},"body":"Jakub Narebski wrote:\n\n> It is a bit strange that markfile has explicitly SHA-1 (\":markid <SHA-1>\"),\n> instead of generic reference to commit, in the case of CVS it would be\n> commitid (what to do for older repositories, though?), in case of Bazaar\n> its revision id (GUID), etc.\n\nUsually importers use at least two separate files to save state, one\nmapping between git object names and mark numbers, and the other mapping\nbetween native revision identifiers and mark numbers.  That way,\nwhen the importer uses marks to refer to previously imported commits or\nblobs, fast-import knows what commits or blobs it is talking about.\n"},{"id":"232168","messageId":"20131218162710.GA3573@thyrsus.com","threadId":"35524","inReplyTo":"CANQwDwdgZUWcgyZCWoDni+e9jgQ+8j0Yn_HMxiMn5OHzsRzjwQ@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-18T16:27:10Z","receivedAt":"2013-12-18T16:27:10Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jakub Narębski <jnareb@gmail.com>:\n> It is a bit strange that markfile has explicitly SHA-1 (\":markid <SHA-1>\"),\n> instead of generic reference to commit, in the case of CVS it would be\n> commitid (what to do for older repositories, though?), in case of Bazaar\n> its revision id (GUID), etc.  Can we assume that SCM v1 fast-export and\n> SCM v2 fast-import markfile uses compatibile commit names in markfile?\n\nFor use in reposurgeon I have defined a generic cross-VCS reference to\ncommit I call an \"action stamp\"; it consists of an RFC3339 date followed by \na committer email address. Here's an example:\n\n\t 2013-02-06T09:35:10Z!esr@thyrsus.com\n\nIn any VCS with changesets (git, Subversion, bzr, Mercurial) this\nalmost always suffices to uniquely identify a commit. The \"almost\" is\nbecause in these systems it is possible for a user to do multiple commits\nin the same second.\n\nAnd now you know why I wish git had subsecond timestamp resolution!  If it\ndid, uniqueness of these in a git stream could be guaranteed.\n\nThe implied model completely breaks for CVS, of course.  There you have to \nuse commitids and plain give up when those don't exist.\n \n> I think it would be possible for remote-helper for cvs-fast-export to find\n> this cutoff date automatically (perhaps with some safety margin), for\n> fetching (incremental import).\n\nYes.\n \n> > As I tried to explain previously in my response to John Herland, it's\n> > incremental output only.  There is *no* CVS exporter known to me, or\n> > him, that supports incremental work.  That would be at best be impractically\n> > difficult; given CVS's limitations it may be actually impossible. I wouldn't\n> > bet against impossible.\n> \n> Even with saving (or re-calculating from git import) guesses about CVS\n> history made so far?\n\nEven with that.  cvsps-2.x tried to do something like this.  It was a lose.\n \n> Anyway I hope that incremental CVS import would be needed less\n> and less as CVS is replaced by any more modern version control system.\n\nI agree.  I have never understood why people on this list are attached to it.\n\n> I was thinking about creating remote-helper for cvs-fast-export, so that\n> git can use local CVS repository as \"remote\", using e.g. \"cvsroot::<path>\"\n> as repo URL, and using this mechanism for incremental import (aka fetch).\n> (Or even \"cvssync::<URL>\" for automatic cvssync + cvs-fast-export).\n> \n> But from what I understand this is not as easy as it seems, even with\n> remote-helper API having support for fast-import stream.\n\nIt's a swamp I wouldn't want to walk into.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232171","messageId":"CACPiFC+W-RiO-YL=Wgs7YzV=z-p97ehfA+64j5F2KbayPAQm8w@mail.gmail.com","threadId":"35524","inReplyTo":"20131218162710.GA3573@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2013-12-18T16:53:47Z","receivedAt":"2013-12-18T16:53:47Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On Wed, Dec 18, 2013 at 11:27 AM, Eric S. Raymond <esr@thyrsus.com> wrote:\n>> Anyway I hope that incremental CVS import would be needed less\n>> and less as CVS is replaced by any more modern version control system.\n>\n> I agree.  I have never understood why people on this list are attached to it.\n\nI think I have answered this question already once in this thread, and\na few times in similar threads with Eric in the past.\n\nPeople track CVS repos that they have not control over. Smart\nprogrammers forced to work with a corporate CVS repo. It happens also\nwith SVN, and witness the popularity of git-svn which can sanely\ninteract with an \"active\" svn repo.\n\nThis is a valid use case. Hard (impossible?) to support. But there\nshould be no surprise as to its reasons.\n\ncheers,\n\n\n\nm\n-- \n martin.langhoff@gmail.com\n -  ask interesting questions\n - don't get distracted with shiny stuff  - working code first\n ~ http://docs.moodle.org/en/User:Martin_Langhoff\n"},{"id":"232174","messageId":"20131218174615.GA5597@sigill.intra.peff.net","threadId":"35524","inReplyTo":"20131218162710.GA3573@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2013-12-18T17:46:15Z","receivedAt":"2013-12-18T17:46:15Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Dec 18, 2013 at 11:27:10AM -0500, Eric S. Raymond wrote:\n\n> For use in reposurgeon I have defined a generic cross-VCS reference to\n> commit I call an \"action stamp\"; it consists of an RFC3339 date followed by \n> a committer email address. Here's an example:\n> \n> \t 2013-02-06T09:35:10Z!esr@thyrsus.com\n> \n> In any VCS with changesets (git, Subversion, bzr, Mercurial) this\n> almost always suffices to uniquely identify a commit. The \"almost\" is\n> because in these systems it is possible for a user to do multiple commits\n> in the same second.\n\nFWIW, this has quite a few collisions in git.git:\n\n  $ git log --format='%ct %ce' | sort | uniq -c | sort -rn | head\n     22 1172221032 normalperson@yhbt.net\n     22 1172221031 normalperson@yhbt.net\n     22 1172221029 normalperson@yhbt.net\n     21 1190197351 gitster@pobox.com\n     21 1172221030 normalperson@yhbt.net\n     20 1190197350 gitster@pobox.com\n     17 1172221033 normalperson@yhbt.net\n     15 1263457676 gitster@pobox.com\n     15 1193717011 gitster@pobox.com\n     14 1367447590 gitster@pobox.com\n\nIn git, it may happen quite a bit during \"git am\" or \"git rebase\", in\nwhich a large number of commits are replayed in a tight loop. You can\nuse the author timestamp instead, but it also collides (try \"%at %ae\" in\nthe above command instead).\n\n> And now you know why I wish git had subsecond timestamp resolution!  If it\n> did, uniqueness of these in a git stream could be guaranteed.\n\nIt's still not guaranteed. Even with sufficient resolution that no two\noperations could possibly complete in the same time unit, clocks do not\nalways march forward. They get reset, they may skew from machine to\nmachine, the same operation may happen on different machines, etc. The\nprobability of such collisions is significantly reduced, though, if only\nbecause the extra precision adds an essentially random factor.\n\nBut in some cases you might even see the same commit \"replayed\" on top\nof different parts of the graph, or affecting different paths (e.g., by\nfilter-branch). I.e., no matter what your precision, multiple hacked-up\nviews of the changeset will still always have that same timestamp.\n\n-Peff\n"},{"id":"232183","messageId":"20131218191648.GA4533@thyrsus.com","threadId":"35524","inReplyTo":"20131218174615.GA5597@sigill.intra.peff.net","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-18T19:16:48Z","receivedAt":"2013-12-18T19:16:48Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Jeff King <peff@peff.net>:\n> In git, it may happen quite a bit during \"git am\" or \"git rebase\", in\n> which a large number of commits are replayed in a tight loop.\n\nThat's a good point - a repeatable real-world case in which we can\nexpect that behavior.\n\nThis case could be solved, though, with a slight tweak to the commit generator\nin git (given subsecond timestamps).  It could keep the time of last commit\nand stall by an arbitrary small amount, enough to show up as a timestamp\ndifference. \n\nAction stamps work pretty well inside reposurgeon because they're\nmainly used to identify commits from older VCSes that can't run that\nfast. Collisions are theoretically possible but I'm never seen one in\nthe wild.\n\n>                                                       You can\n> use the author timestamp instead, but it also collides (try \"%at %ae\" in\n> the above command instead).\n\nYes, obviously for the same reason. \n \n> > And now you know why I wish git had subsecond timestamp resolution!  If it\n> > did, uniqueness of these in a git stream could be guaranteed.\n> \n> It's still not guaranteed. Even with sufficient resolution that no two\n> operations could possibly complete in the same time unit, clocks do not\n> always march forward. They get reset, they may skew from machine to\n> machine, the same operation may happen on different machines, etc.\n\nRight...but the *same person* submitting operations from *different\nmachines* within the time window required to be caught by these effects\nis at worst fantastically unlikely.  That case is exactly why action \nstamps have an email part.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232188","messageId":"20131218195450.GK3163@serenity.lan","threadId":"35524","inReplyTo":"CACPiFC+W-RiO-YL=Wgs7YzV=z-p97ehfA+64j5F2KbayPAQm8w@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"John Keeping","fromEmail":"john@keeping.me.uk","sentAt":"2013-12-18T19:54:51Z","receivedAt":"2013-12-18T19:54:51Z","isPatch":false,"sender":{"key":"john@keeping.me.uk","avatar":"https://avatars.githubusercontent.com/u/1702081?v=4"},"body":"On Wed, Dec 18, 2013 at 11:53:47AM -0500, Martin Langhoff wrote:\n> On Wed, Dec 18, 2013 at 11:27 AM, Eric S. Raymond <esr@thyrsus.com> wrote:\n> >> Anyway I hope that incremental CVS import would be needed less\n> >> and less as CVS is replaced by any more modern version control system.\n> >\n> > I agree.  I have never understood why people on this list are attached to it.\n> \n> I think I have answered this question already once in this thread, and\n> a few times in similar threads with Eric in the past.\n> \n> People track CVS repos that they have not control over. Smart\n> programmers forced to work with a corporate CVS repo. It happens also\n> with SVN, and witness the popularity of git-svn which can sanely\n> interact with an \"active\" svn repo.\n> \n> This is a valid use case. Hard (impossible?) to support. But there\n> should be no surprise as to its reasons.\n\nAnd at this point the git-cvsimport manpage says:\n\n   WARNING: git cvsimport uses cvsps version 2, which is considered\n   deprecated; it does not work with cvsps version 3 and later. If you\n   are performing a one-shot import of a CVS repository consider using\n   cvs2git[1] or parsecvs[2].\n\nWhich I think sums up the position nicely; if you're doing a one-shot\nimport then the standalone tools are going to be a better choice, but if\nyou're trying to use Git for your work on top of CVS the only choice is\ncvsps with git-cvsimport.\n"},{"id":"232196","messageId":"20131218202009.GA4935@thyrsus.com","threadId":"35524","inReplyTo":"20131218195450.GK3163@serenity.lan","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-18T20:20:09Z","receivedAt":"2013-12-18T20:20:09Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"John Keeping <john@keeping.me.uk>:\n> Which I think sums up the position nicely; if you're doing a one-shot\n> import then the standalone tools are going to be a better choice, but if\n> you're trying to use Git for your work on top of CVS the only choice is\n> cvsps with git-cvsimport.\n\nWhich will trash your history - the bugs in that are worse than the bugs\nin 3.0, which are bad enough that I *terminated* it.\n\nLovely....\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232198","messageId":"1387399661.014711355@apps.rackspace.com","threadId":"35524","inReplyTo":"20131218202009.GA4935@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Kent R. Spillner","fromEmail":"kspillner@acm.org","sentAt":"2013-12-18T20:47:41Z","receivedAt":"2013-12-18T20:47:41Z","isPatch":false,"sender":{"key":"kspillner@acm.org","avatar":"https://gravatar.com/avatar/050ad1f9bbca4ec2de093b251f225080467f6d8180227caf25b591bd123d140a?d=mp&s=160"},"body":"> Which will trash your history - the bugs in that are worse than the bugs\n> in 3.0, which are bad enough that I *terminated* it.\n\nWhich *might* trash your history.\n\ncvsps v2 and git cvsimport work as advertised with simple, linear CVS\nrepositories.  I maintain a git mirror of an active CVS repo and run git\ncvsimport every few days to sync with the latest upstream changes.  The\nonly problem I encountered so far was when you released cvsps v3 and broke\ngit cvsimport. :)  I had to manually downgrade to cvsps v2.2b1 and configure\nmy package manager to ignore cvsps updates, but I haven't had any problems\nsince.\n"},{"id":"232214","messageId":"52B2335D.2030607@alum.mit.edu","threadId":"35524","inReplyTo":"20131217184724.GA17709@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2013-12-18T23:44:29Z","receivedAt":"2013-12-18T23:44:29Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 12/17/2013 07:47 PM, Eric S. Raymond wrote:\n> Johan Herland <johan@herland.net>:\n>> However, I fear that you underestimate the number of users that want\n>> to use Git against CVS repos that are orders of magnitude larger (in\n>> both dimensions: #commits and #files) than your example repo.\n> \n> You may be right. See below...\n> \n> I'm working with Alan Barret now on trying to convert the NetBSD\n> repositories. They break cvs-fast-export through sheer bulk of\n> metadata, by running the machine out of core.  This is exactly\n> the kind of huge case that you're talking about.\n> \n> Alan and I are going to take a good hard whack at modifying cvs-fast-export \n> to make this work. Because there really aren't any feasible alternatives.\n> The analysis code in cvsps was never good enough. cvs2git, being written\n> in Python, would hit the core limit faster than anything written in C.\n\ncvs2git goes to great lengths to store intermediate data to disk and\nkeep the working set small and therefore (despite the Python overhead) I\nam confident that it scales better than cvs-fast-export.  My usual test\nrepo was gcc:\n\nTotal CVS Files:             25013\nTotal CVS Revisions:        578010\nTotal CVS Branches:        1487929\nTotal CVS Tags:           11435500\nTotal Unique Tags:             814\nTotal Unique Branches:         116\nCVS Repos Size in KB:      2074248\nTotal SVN Commits:           64501\n\nI also regularly converted mozilla (4.2 GB) and emacs (560 MB) for\ntesting purposes.  These could all be converted on a 32-bit computer.\n\nOther projects that cvs2svn/cvs2git could handle: FreeBSD, Gentoo, KDE,\nGNOME, PostgreSQL.  (Though for KDE, which I think was in the 16 GB\nrange, I know that they used a giant machine for the conversion.)\n\nIf you haven't tried cvs2git yet, please start it up somewhere in the\nbackground.  It might take a while but it should have no trouble with\nyour repos, and then you can compare the tools based on experience\nrather than speculation.\n\n> Which matters, because right now the set of people working on CVS lifters\n> begins with me and ends with Michael Rafferty (cvs2git), who seems even\n> less interested in incremental conversion than I am.  Unless somebody\n> comes out of nowhere and wants to own that problem, it's not going\n> to get solved.\n\nA correct incremental converter could be done (as long as the CVS users\ndon't literally change history retroactively) but it would be a lot of\nwork.  Parsing the CVS files isn't the problem; after all, CVS has to do\nthat every time you check out a branch.  The problem is the extra\nbookkeeping that would be needed to keep the overlapping history\nconsistent between runs N and N+1 of the tool.  I sketched out what\nwould be necessary once and it came out to several solid weeks of work.\n\nBut the traffic on the cvs2svn/cvs2git mailing list has trailed off\nessentially to zero, so either the software is perfect already (haha) or\nmost everybody has already converted.  Therefore I don't invest any\nsignificant time in that project these days.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"232224","messageId":"CALKQrgdin=8h9dr=h+VfGjX3suOGRXNsvzzcF=_L9cQDYtKPgg@mail.gmail.com","threadId":"35524","inReplyTo":"52B2335D.2030607@alum.mit.edu","subject":"Re: I have end-of-lifed cvsps","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2013-12-19T01:11:53Z","receivedAt":"2013-12-19T01:11:53Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Thu, Dec 19, 2013 at 12:44 AM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> A correct incremental converter could be done (as long as the CVS users\n> don't literally change history retroactively) but it would be a lot of work.\n\nAlthough I agree with that sentence as it is stated, I also believe\nthat the parenthesized condition rules out a _majority_ of CVS repo of\nnon-trivial size/history. So even though a correct incremental\nconverter could be built, it would be pretty much useless if it did\nnot gracefully handle rewritten history. And in the face of rewritten\nhistory it becomes pretty much impossible to define what a \"correct\"\nconversion should even look like (not to mention the difficulty of\nactually implementing that converter...).\n\nHere are just a couple of things a CVS user can do (and that happened\nfairly regularly at my previous $dayjob) that would make life\ndifficult for an incremental converter (and that also makes stable\noutput from a non-incremental converter hard to solve in practice):\n\n - A user \"deletes\" $file from $branch by simply removing the $branch\nsymbol on $file (cvs tag -B -d $branch $file). CVS stores no record of\nthis. Many non-incremental importers will see $file as never having\nexisted on $branch. An incremental importer starting from a previously\nconverted state, must somehow deal with that previous state no longer\nexisting from the POV of CVS.\n\n - A user moves a release tag on a few files to include a late bugfix\ninto an upcoming release (cvs tag -F -r $new_rev $tag $file). There\nmight be no single point in time where the tagged state existed in the\nrepo, it has become a \"Frankentag\". You could claim user error here,\nand that such shortcuts should not happen, but that doesn't really\nprevent it from ever happening. Recreating the tree state of the\nFrankentag in Git is easy, but what kind of history do you construct\nto lead up to that tree?\n\n - A modularized project develops code on HEAD, and make regular\nreleases of each module by tagging the files in the module dir with\n\"$modulename-$version\". Afterwards a project-wide \"stable\" tag is\nmoved on that subset of files to include the new module release into\nthe \"stable\" tag. (\"stable\" is conceptually a branch, but the CVS\nmechanism used here is still the tag, since CVS branches cannot\n\"follow\" eachother like in Git). This is pretty much the same\nFrankentag scenario as above, except that in this case it might be\nconsidered Best Practice (it was at our $dayjob), and not a\nshortcut/user error made by a single user.\n\n(None of these examples even involve the \"cvs admin\" which allows you\nto do some truly scary and demented things to your CVS history...)\n\nMy point here is that people will use whatever available tools they\nhave to solve whatever problems they are currently having. And when\nCVS is your tool, you will sooner or later end up with a \"solution\"\nthat irrevocably rewrites your CVS history.\n\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"232228","messageId":"20131219040604.GA7654@thyrsus.com","threadId":"35524","inReplyTo":"52B2335D.2030607@alum.mit.edu","subject":"Re: I have end-of-lifed cvsps","fromName":"Eric S. Raymond","fromEmail":"esr@thyrsus.com","sentAt":"2013-12-19T04:06:04Z","receivedAt":"2013-12-19T04:06:04Z","isPatch":false,"sender":{"key":"esr@thyrsus.com","avatar":"https://avatars.githubusercontent.com/u/727961?v=4"},"body":"Michael Haggerty <mhagger@alum.mit.edu>:\n> If you haven't tried cvs2git yet, please start it up somewhere in the\n> background.  It might take a while but it should have no trouble with\n> your repos, and then you can compare the tools based on experience\n> rather than speculation.\n\nThat would be a good thing.\n\nMichael, in case you're wondering why I've continued to work on\ncvs-fast-export when cvs2git exists, there are exactly two reasons:\n(a) it's a whole lot faster on repos that aren't large enough to\ndemand multipass, and (b) the single-whole-dumpfile output makes it a\nbetter reposurgeon front end.\n\n> But the traffic on the cvs2svn/cvs2git mailing list has trailed off\n> essentially to zero, so either the software is perfect already (haha) or\n> most everybody has already converted.  Therefore I don't invest any\n> significant time in that project these days.\n\nReasonable.  I'm doing this as a temporary break from working on GPSD.\nI don't expect to be investing a lot of time in it after I get it\nto a 1.0 state.\n-- \n\t\t<a href=\"http://www.catb.org/~esr/\">Eric S. Raymond</a>\n"},{"id":"232231","messageId":"52B2BCF9.5080300@alum.mit.edu","threadId":"35524","inReplyTo":"CALKQrgdin=8h9dr=h+VfGjX3suOGRXNsvzzcF=_L9cQDYtKPgg@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2013-12-19T09:31:37Z","receivedAt":"2013-12-19T09:31:37Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 12/19/2013 02:11 AM, Johan Herland wrote:\n> On Thu, Dec 19, 2013 at 12:44 AM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> A correct incremental converter could be done (as long as the CVS users\n>> don't literally change history retroactively) but it would be a lot of work.\n> \n> Although I agree with that sentence as it is stated, I also believe\n> that the parenthesized condition rules out a _majority_ of CVS repo of\n> non-trivial size/history. So even though a correct incremental\n> converter could be built, it would be pretty much useless if it did\n> not gracefully handle rewritten history. And in the face of rewritten\n> history it becomes pretty much impossible to define what a \"correct\"\n> conversion should even look like (not to mention the difficulty of\n> actually implementing that converter...).\n\nA correct conversion would, conceptually, take a diff between the old\nCVS history and the new CVS history (I'm talking about the history as a\nwhole, not a diff between two changesets), figure out what had changed,\nand then figure out what Git commits to make to effect the same\nconceptual changes in Git-land.\n\nThis means that the final Git history would have to depend not only on\nthe current entirety of the CVS history, but also on what the CVS\nhistory *was* during previous incremental imports and how the tool chose\nto represent that history in Git the previous rounds.\n\nThere is a tradeoff here.  The smarter the tool is, the fewer\nrestrictions would have to be made on what people can do in CVS.  For\nexample, it wouldn't be unreasonable to impose a rule that people are\nnot allowed to move files within the CVS repository (e.g., to fake\nmove-file-with-history) after the CVS <-> Git bridge is in use.  (Abuses\nof the history that occurred *before* the first incremental conversion,\non the other hand, wouldn't be a problem.)  If the user of the\nincremental tool has *no* influence on how his colleagues use CVS, then\nthe tool would have to be very smart and/or the user would might\nsometimes be forced to do another from-scratch conversion.\n\n> Here are just a couple of things a CVS user can do (and that happened\n> fairly regularly at my previous $dayjob) that would make life\n> difficult for an incremental converter (and that also makes stable\n> output from a non-incremental converter hard to solve in practice):\n> \n>  - A user \"deletes\" $file from $branch by simply removing the $branch\n> symbol on $file (cvs tag -B -d $branch $file). CVS stores no record of\n> this. Many non-incremental importers will see $file as never having\n> existed on $branch. An incremental importer starting from a previously\n> converted state, must somehow deal with that previous state no longer\n> existing from the POV of CVS.\n\nNo problem; the tool could just add a synthetic commit \"git rm\"ming the\nfile from the branch.  It wouldn't know *when* the file was deleted, so\nit would have to pick a plausible date between the time of the last\nincremental conversion and the one that discovers that the branch tag\nhas been removed from the file.  The resulting Git history would contain\nmore complete information than CVS's history.\n\n>  - A user moves a release tag on a few files to include a late bugfix\n> into an upcoming release (cvs tag -F -r $new_rev $tag $file). There\n> might be no single point in time where the tagged state existed in the\n> repo, it has become a \"Frankentag\". You could claim user error here,\n> and that such shortcuts should not happen, but that doesn't really\n> prevent it from ever happening. Recreating the tree state of the\n> Frankentag in Git is easy, but what kind of history do you construct\n> to lead up to that tree?\n\nFrankentags (tags that include file versions that didn't occur\ncontemporaneously) can occur even with one-time CVS->Git conversions.\nThe only way to handle them is to create a Git branch representing the\ntag and base it at a plausible Git commit, and then (on the branch)\nissue a fixup commit that makes the contents of the branch equal to the\ncontents of the CVS branch.  This is a problem that cvs2git already handles.\n\nA hypothetical incremental importer would have to notice the changes in\nthe branch contents between the previous conversion and the current one,\nand create commits on the branch to bring it in line with the current\ncontents.  This is no uglier than what a one-shot conversion already has\nto do.\n\n>  - A modularized project develops code on HEAD, and make regular\n> releases of each module by tagging the files in the module dir with\n> \"$modulename-$version\". Afterwards a project-wide \"stable\" tag is\n> moved on that subset of files to include the new module release into\n> the \"stable\" tag. (\"stable\" is conceptually a branch, but the CVS\n> mechanism used here is still the tag, since CVS branches cannot\n> \"follow\" eachother like in Git). This is pretty much the same\n> Frankentag scenario as above, except that in this case it might be\n> considered Best Practice (it was at our $dayjob), and not a\n> shortcut/user error made by a single user.\n\nSame problem and same solution as above, as far as I can see.\n\n> (None of these examples even involve the \"cvs admin\" which allows you\n> to do some truly scary and demented things to your CVS history...)\n\nEven some of these might be permitted.  For example:\n\n* Obsoleting already-converted revisions: it's a pretty stupid thing to\ndo in most cases and the tool could just ignore such events, retaining\nthe history in Git.  If the revisions were obsoleted because they\ncontained proprietary information or something, then you've got a bigger\nproblem on your hands but one that you would have even if you were using\npure Git.\n\n* Retroactive changes to log messages: would probably have to be ignored\nor handled via notes.\n\n* Changes to the \"default branch\" (another brain-dead CVS feature\nrelated to vendor branches): I'd have to think about it.  But handling\nvendor branches is already difficult for a one-time converter because\nCVS retains too little info (but cvs2git does it except in the most\nambiguous cases).  An incremental importer would have *more* information\nthan a one-shot importer, because it would have a hope of catching the\nchange to the default branch at roughly the time it occurred.\n\n> My point here is that people will use whatever available tools they\n> have to solve whatever problems they are currently having. And when\n> CVS is your tool, you will sooner or later end up with a \"solution\"\n> that irrevocably rewrites your CVS history.\n\nYes, but I maintain that an incremental importer could keep a Git\nhistory that is consistent with the CVS history in the sense that:\n\n1. the result of checking out any branch or tag, right after a run of\nthe importer, gives the same results as checking the same branch or tag\nout of CVS.\n\n2. the Git history from one run is added to (never rewritten) by the\nnext run.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"232232","messageId":"52B2BFBB.5090100@alum.mit.edu","threadId":"35524","inReplyTo":"20131219040604.GA7654@thyrsus.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2013-12-19T09:43:23Z","receivedAt":"2013-12-19T09:43:23Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 12/19/2013 05:06 AM, Eric S. Raymond wrote:\n> Michael Haggerty <mhagger@alum.mit.edu>:\n>> If you haven't tried cvs2git yet, please start it up somewhere in the\n>> background.  It might take a while but it should have no trouble with\n>> your repos, and then you can compare the tools based on experience\n>> rather than speculation.\n> \n> That would be a good thing.\n> \n> Michael, in case you're wondering why I've continued to work on\n> cvs-fast-export when cvs2git exists, there are exactly two reasons:\n> (a) it's a whole lot faster on repos that aren't large enough to\n> demand multipass,\n\nWhat difference does speed make on little repositories?  They are fast\nenough anyway.\n\nIf you are worried about the speed of testing and iterating on your\nreposurgeon configuration, then just write the output of cvs2svn to a\ntemporary file and use the temporary file as input to reposurgeon.\n\n> and (b) the single-whole-dumpfile output makes it a\n> better reposurgeon front end.\n\nI can't believe you are still hung up on this!  OK, just for you, here\nit is: cvs2git-3.0, in gorgeous pipey purity:\n\n    #! /bin/sh\n    blobfile=$(mktemp /tmp/myblobs-XXXXXX.out)\n    dumpfile=$(mktemp /tmp/mydump-XXXXXX.out)\n    cvs2git-2.0 --blobfile=\"$blobfile\" --dumpfile=\"$dumpfile\" \"$@\" 1>&2 &&\n    cat \"$blobfile\" \"$dumpfile\"\n    rm \"$blobfile\" \"$dumpfile\"\n\nI don't think that cvs2git-2.0 outputs any junk to stdout, but just in\ncase it does I've redirected stdout explicitly to stderr to avoid\ncommingling it with the output of this script.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"},{"id":"232236","messageId":"CALKQrgeiVSPhe84xTnKQ6iAmN3UX_Jy77pgp5ieSwFQ21tWPFg@mail.gmail.com","threadId":"35524","inReplyTo":"52B2BCF9.5080300@alum.mit.edu","subject":"Re: I have end-of-lifed cvsps","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2013-12-19T15:26:18Z","receivedAt":"2013-12-19T15:26:18Z","isPatch":false,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Thu, Dec 19, 2013 at 10:31 AM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n> On 12/19/2013 02:11 AM, Johan Herland wrote:\n>> On Thu, Dec 19, 2013 at 12:44 AM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>>> A correct incremental converter could be done (as long as the CVS users\n>>> don't literally change history retroactively) but it would be a lot of work.\n>>\n>> Although I agree with that sentence as it is stated, I also believe\n>> that the parenthesized condition rules out a _majority_ of CVS repo of\n>> non-trivial size/history. So even though a correct incremental\n>> converter could be built, it would be pretty much useless if it did\n>> not gracefully handle rewritten history. And in the face of rewritten\n>> history it becomes pretty much impossible to define what a \"correct\"\n>> conversion should even look like (not to mention the difficulty of\n>> actually implementing that converter...).\n>\n> A correct conversion would, conceptually, take a diff between the old\n> CVS history and the new CVS history (I'm talking about the history as a\n> whole, not a diff between two changesets), figure out what had changed,\n> and then figure out what Git commits to make to effect the same\n> conceptual changes in Git-land.\n>\n> This means that the final Git history would have to depend not only on\n> the current entirety of the CVS history, but also on what the CVS\n> history *was* during previous incremental imports and how the tool chose\n> to represent that history in Git the previous rounds.\n>\n> There is a tradeoff here.  The smarter the tool is, the fewer\n> restrictions would have to be made on what people can do in CVS.  For\n> example, it wouldn't be unreasonable to impose a rule that people are\n> not allowed to move files within the CVS repository (e.g., to fake\n> move-file-with-history) after the CVS <-> Git bridge is in use.  (Abuses\n> of the history that occurred *before* the first incremental conversion,\n> on the other hand, wouldn't be a problem.)  If the user of the\n> incremental tool has *no* influence on how his colleagues use CVS, then\n> the tool would have to be very smart and/or the user would might\n> sometimes be forced to do another from-scratch conversion.\n\nAgreed, but I find it quite ugly how the git history will end up\ndifferent depending on _when_ the incremental conversion is run. It\nmeans that it will be impossible for two users to create the same Git\nrepo (matching SHA1s), unless they carefully synchronize all of their\nconversion runs (at which point it's much simpler to run a single\nconversion and then have both users fetch the result).\n\nThere is a continuum here in incremental converters:\n\nAt one end - given that you're always going to lose _some_ history -\nyou can go \"screw it! let's not care about history at all!\", and do\nthe fastest possible conversion: check out the current CVS version;\ndiff that against the previous CVS version; apply the diff to your Git\nrepo as a single commit. I suspect quite a lot of users would be happy\nwith this solution - at least as a temporary measure while they wait\nfor their surrounding organization to do a proper migraiton off CVS.\n\nAt the other end - you can realize that the CVS storage format on the\nserver is simply too lossy, and you can write a proxy or monitor that\nintercept CVS operations on the server, and replicate those in a\ncompanion Git repo as soon as they occur in CVS. Whether you write a\nCVS server monitor that detects changes to the CVS server files in\nreal time (using e.g. inotify or similar), or you write a CVS server\nproxy that intercepts CVS commands from the user (also forwarding them\nto the _real_ CVS server) is an implementation detail[*]. The\nimportant thing is you should end up with is a real-time stream of\nchanges that can be converted to corresponding changes in a Git repo.\nThat should give you closest possible picture of what really happens\nin a CVS repo, even better than what CVS stores in its on-disk format.\nThis would allow an organization to provide a (read-only) Git mirror\nof their CVS repo.\n\nWhat we have been discussing in this thread (various strategies for\nfixing up broken history in Git) can be considered intermediate points\nbetween the two extremes presented above: You try to recreate as much\nhistory as possible, but realize that you sometimes need to simply\nsynthesize some fake history in order to make everything fit together.\n\n>> Here are just a couple of things a CVS user can do (and that happened\n>> fairly regularly at my previous $dayjob) that would make life\n>> difficult for an incremental converter (and that also makes stable\n>> output from a non-incremental converter hard to solve in practice):\n>>\n>>  - A user \"deletes\" $file from $branch by simply removing the $branch\n>> symbol on $file (cvs tag -B -d $branch $file). CVS stores no record of\n>> this. Many non-incremental importers will see $file as never having\n>> existed on $branch. An incremental importer starting from a previously\n>> converted state, must somehow deal with that previous state no longer\n>> existing from the POV of CVS.\n>\n> No problem; the tool could just add a synthetic commit \"git rm\"ming the\n> file from the branch.  It wouldn't know *when* the file was deleted, so\n> it would have to pick a plausible date between the time of the last\n> incremental conversion and the one that discovers that the branch tag\n> has been removed from the file.  The resulting Git history would contain\n> more complete information than CVS's history.\n\nA server proxy/monitor analyzing CVS operations in real time would\nknow _exactly_ when the file was removed...\n\n>>  - A user moves a release tag on a few files to include a late bugfix\n>> into an upcoming release (cvs tag -F -r $new_rev $tag $file). There\n>> might be no single point in time where the tagged state existed in the\n>> repo, it has become a \"Frankentag\". You could claim user error here,\n>> and that such shortcuts should not happen, but that doesn't really\n>> prevent it from ever happening. Recreating the tree state of the\n>> Frankentag in Git is easy, but what kind of history do you construct\n>> to lead up to that tree?\n>\n> Frankentags (tags that include file versions that didn't occur\n> contemporaneously) can occur even with one-time CVS->Git conversions.\n> The only way to handle them is to create a Git branch representing the\n> tag and base it at a plausible Git commit, and then (on the branch)\n> issue a fixup commit that makes the contents of the branch equal to the\n> contents of the CVS branch.  This is a problem that cvs2git already handles.\n>\n> A hypothetical incremental importer would have to notice the changes in\n> the branch contents between the previous conversion and the current one,\n> and create commits on the branch to bring it in line with the current\n> contents.  This is no uglier than what a one-shot conversion already has\n> to do.\n\nTrue, but analyzing CVS operations in real time, you might be able to\nrecreate the moving (and adding/deleting) of tags as file edits (and\nadds/deletes) in the corresponding Git branch.\n\n>>  - A modularized project develops code on HEAD, and make regular\n>> releases of each module by tagging the files in the module dir with\n>> \"$modulename-$version\". Afterwards a project-wide \"stable\" tag is\n>> moved on that subset of files to include the new module release into\n>> the \"stable\" tag. (\"stable\" is conceptually a branch, but the CVS\n>> mechanism used here is still the tag, since CVS branches cannot\n>> \"follow\" eachother like in Git). This is pretty much the same\n>> Frankentag scenario as above, except that in this case it might be\n>> considered Best Practice (it was at our $dayjob), and not a\n>> shortcut/user error made by a single user.\n>\n> Same problem and same solution as above, as far as I can see.\n>\n>> (None of these examples even involve the \"cvs admin\" which allows you\n>> to do some truly scary and demented things to your CVS history...)\n>\n> Even some of these might be permitted.  For example:\n>\n> * Obsoleting already-converted revisions: it's a pretty stupid thing to\n> do in most cases and the tool could just ignore such events, retaining\n> the history in Git.  If the revisions were obsoleted because they\n> contained proprietary information or something, then you've got a bigger\n> problem on your hands but one that you would have even if you were using\n> pure Git.\n>\n> * Retroactive changes to log messages: would probably have to be ignored\n> or handled via notes.\n>\n> * Changes to the \"default branch\" (another brain-dead CVS feature\n> related to vendor branches): I'd have to think about it.  But handling\n> vendor branches is already difficult for a one-time converter because\n> CVS retains too little info (but cvs2git does it except in the most\n> ambiguous cases).  An incremental importer would have *more* information\n> than a one-shot importer, because it would have a hope of catching the\n> change to the default branch at roughly the time it occurred.\n\nAgreed, but if you want correct metadata (_when_ did these changes\nhappen, _who_ performed them), then you need to actually monitor the\nCVS command stream (or CVS server files) in real time...\n\n>> My point here is that people will use whatever available tools they\n>> have to solve whatever problems they are currently having. And when\n>> CVS is your tool, you will sooner or later end up with a \"solution\"\n>> that irrevocably rewrites your CVS history.\n>\n> Yes, but I maintain that an incremental importer could keep a Git\n> history that is consistent with the CVS history in the sense that:\n>\n> 1. the result of checking out any branch or tag, right after a run of\n> the importer, gives the same results as checking the same branch or tag\n> out of CVS.\n>\n> 2. the Git history from one run is added to (never rewritten) by the\n> next run.\n\nYes, and even my simplest/fastest possible converter described above\ncan meet those criteria. After that, it really becomes a question of\n_how_much_ CVS history you want to retain in your incremental import.\nI have described the two extremes above. Interestingly, _both_ of\nthose extremes would look quite different from the\nwhole-history-gone-incremental converters represented by cvs2git and\ncvs-fast-export, and _both_ of the extremes would probably also\nprovide a converted result quite a bit faster than anything in between\n(one by virtue of depending on a single \"cvs update\" command, and the\nother by monitoring the CVS server and performing the conversion to\nGit in real time).\n\n\n...Johan\n\n\n[*]: That said, I suspect git-cvsserver would be a good starting point\nfor implementing a CVS server proxy, if someone is actually interested\nin looking at this...\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"232237","messageId":"52B31C4B.8080404@alum.mit.edu","threadId":"35524","inReplyTo":"CALKQrgeiVSPhe84xTnKQ6iAmN3UX_Jy77pgp5ieSwFQ21tWPFg@mail.gmail.com","subject":"Re: I have end-of-lifed cvsps","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2013-12-19T16:18:19Z","receivedAt":"2013-12-19T16:18:19Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"On 12/19/2013 04:26 PM, Johan Herland wrote:\n> On Thu, Dec 19, 2013 at 10:31 AM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>> On 12/19/2013 02:11 AM, Johan Herland wrote:\n>>> On Thu, Dec 19, 2013 at 12:44 AM, Michael Haggerty <mhagger@alum.mit.edu> wrote:\n>>>> A correct incremental converter could be done (as long as the CVS users\n>>>> don't literally change history retroactively) but it would be a lot of work.\n>>>\n>>> Although I agree with that sentence as it is stated, I also believe\n>>> that the parenthesized condition rules out a _majority_ of CVS repo of\n>>> non-trivial size/history. So even though a correct incremental\n>>> converter could be built, it would be pretty much useless if it did\n>>> not gracefully handle rewritten history. And in the face of rewritten\n>>> history it becomes pretty much impossible to define what a \"correct\"\n>>> conversion should even look like (not to mention the difficulty of\n>>> actually implementing that converter...).\n>>\n>> A correct conversion would, conceptually, take a diff between the old\n>> CVS history and the new CVS history (I'm talking about the history as a\n>> whole, not a diff between two changesets), figure out what had changed,\n>> and then figure out what Git commits to make to effect the same\n>> conceptual changes in Git-land.\n>>\n>> This means that the final Git history would have to depend not only on\n>> the current entirety of the CVS history, but also on what the CVS\n>> history *was* during previous incremental imports and how the tool chose\n>> to represent that history in Git the previous rounds.\n>>\n>> There is a tradeoff here.  The smarter the tool is, the fewer\n>> restrictions would have to be made on what people can do in CVS.  For\n>> example, it wouldn't be unreasonable to impose a rule that people are\n>> not allowed to move files within the CVS repository (e.g., to fake\n>> move-file-with-history) after the CVS <-> Git bridge is in use.  (Abuses\n>> of the history that occurred *before* the first incremental conversion,\n>> on the other hand, wouldn't be a problem.)  If the user of the\n>> incremental tool has *no* influence on how his colleagues use CVS, then\n>> the tool would have to be very smart and/or the user would might\n>> sometimes be forced to do another from-scratch conversion.\n> \n> Agreed, but I find it quite ugly how the git history will end up\n> different depending on _when_ the incremental conversion is run. It\n> means that it will be impossible for two users to create the same Git\n> repo (matching SHA1s), unless they carefully synchronize all of their\n> conversion runs\n\nEven git-svn doesn't guarantee the same results over time.  The most\nobvious scenario when it fails is when somebody changes an SVN commit's\nmetadata retroactively using something like \"svn propedit --revprop\nsvn:log\".  Consistency over time across two independent conversion\nprocesses (that don't communicate) is not even theoretically possible.\n\n> (at which point it's much simpler to run a single\n> conversion and then have both users fetch the result).\n\nYes.  That is a very reasonable approach.\n\n[Discussion of hypothetical real-time inode-watching or proxy-based\nconverter omitted here...]\n> Agreed, but if you want correct metadata (_when_ did these changes\n> happen, _who_ performed them), then you need to actually monitor the\n> CVS command stream (or CVS server files) in real time...\n\nIn my opinion it is ridiculous to try to design a CVS <-> Git bridge\nthat tries to use back-channels to fill in historical data that even CVS\ndoesn't record.  Such a thing would require an intimate connection to\nthe CVS server from the IT department that is presumably blocking a real\nmove to Git.  So who would ever be able to use it?\n\nThe only reason to record extra information would be to enable the\nbridge to do self-consistent incremental conversions, and in that case\nthe *only* extra information that has to be recorded is the information\nthat would have anyway landed in Git during the previous conversion.\n\n>>> My point here is that people will use whatever available tools they\n>>> have to solve whatever problems they are currently having. And when\n>>> CVS is your tool, you will sooner or later end up with a \"solution\"\n>>> that irrevocably rewrites your CVS history.\n>>\n>> Yes, but I maintain that an incremental importer could keep a Git\n>> history that is consistent with the CVS history in the sense that:\n>>\n>> 1. the result of checking out any branch or tag, right after a run of\n>> the importer, gives the same results as checking the same branch or tag\n>> out of CVS.\n>>\n>> 2. the Git history from one run is added to (never rewritten) by the\n>> next run.\n> \n> Yes, and even my simplest/fastest possible converter described above\n> can meet those criteria. After that, it really becomes a question of\n> _how_much_ CVS history you want to retain in your incremental import.\n\nI think you want enough history to make it pleasant to work with the\nresulting Git repository.  That approximately means that you need some\nsemblance of the CVS commits to be reconstructed, with their correct\nmetadata, on the closest thing to their correct branches that is\nconsistent with the CVS - Git impedance mismatch.\n\n> I have described the two extremes above. Interestingly, _both_ of\n> those extremes would look quite different from the\n> whole-history-gone-incremental converters represented by cvs2git and\n> cvs-fast-export, and _both_ of the extremes would probably also\n> provide a converted result quite a bit faster than anything in between\n> (one by virtue of depending on a single \"cvs update\" command, and the\n> other by monitoring the CVS server and performing the conversion to\n> Git in real time).\n\nI am not an extremist.  And I know how much work it would be to start a\nproject like this from scratch.  After all, what it can do should be a\nstrict superset of what a tool like cvs2git can do, and cvs2svn/cvs2git\n(according to Ohloh's COCOMO estimate) contains the equivalent of 7\nperson-years of effort.\n\nAnyway, this is all just blah blah unless somebody volunteers to work on\nit.  And I think that is highly unlikely, especially given the\ndecreasing number of CVS repositories in the wild.\n\nMichael\n\n-- \nMichael Haggerty\nmhagger@alum.mit.edu\nhttp://softwareswirl.blogspot.com/\n"}]}