{"thread":{"id":"3809","subject":"Fixes to parsecvs","startedAt":"2006-04-06T06:36:32Z","lastAt":"2006-04-09T23:17:54Z","messageCount":13,"participants":["Keith Packard","Jan-Benedict Glaw","Johannes Schindelin","Jim Radford","Martin Langhoff","Francois Romieu"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"18417","messageId":"1144305392.2303.240.camel@neko.keithp.com","threadId":"3809","inReplyTo":null,"subject":"Fixes to parsecvs","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-04-06T06:36:32Z","receivedAt":"2006-04-06T06:36:32Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"note, parsecvs remains available from:\n\n\tgit://git.freedesktop.org/~keithp/parsecvs\n\nI've \"fixed\" the lexer to permit getc/ungetc in the data parsing\nfunctions. This should resolve the flex -l / -X problems.\n\nJim Radford send a patch to add '/' as a legal tag character\n\nI added my custom edit-change-log script for people dealing with\nX.org-style commit messages.\n\nAnd, it deals with import branch revisions that aren't supposed to\nget merged back to the trunk, creating a custom branch name based on the\nbranch revision (which must be global across all files).\n\n5e5f4c012aec2db012a08b1c7ed5219ed5100111\n\n-- \nkeith.packard@intel.com\n\n"},{"id":"18419","messageId":"20060406120812.GO13324@lug-owl.de","threadId":"3809","inReplyTo":"1144305392.2303.240.camel@neko.keithp.com","subject":"Re: Fixes to parsecvs","fromName":"Jan-Benedict Glaw","fromEmail":"jbglaw@lug-owl.de","sentAt":"2006-04-06T12:08:12Z","receivedAt":"2006-04-06T12:08:12Z","isPatch":false,"sender":{"key":"jbglaw@lug-owl.de","avatar":null},"body":"On Wed, 2006-04-05 23:36:32 -0700, Keith Packard <keithp@keithp.com> wrote:\n> note, parsecvs remains available from:\n> \n> \tgit://git.freedesktop.org/~keithp/parsecvs\n\nIt now compiles out-of-the-box for me, nice work.\n\nHowever, it would be nice if you'd add a short description about how\nto use it. Something like this:\n---------------------------------------------------------------------\nThere's still a lot of work to do on parsecvs, but if you want to give\nit a run, first create a copy of the whole CVS tree and go to the base\ndirectory of this copy. (You find a lot of *,v files in this directory\nand all its subdirectories.)\nNow feed all ,v filenames into parsecvs. Keep in mind that a\n`edit-change-log' executable needs to be in your $PATH (a one-line\nscript only exit'ing with 0 will do the job.):\n\n\tfind . -type f -name '*,v' -print | parsecvs\n\nThis will create the .git/ directory and put all the objects, commits\nand tree information into this new git repository.\n---------------------------------------------------------------------\n\nI just ran it against a locally rsync'ed copy of the Binutils ,v\nfiles. Looging at the progress bar, it is bascally ready:\n\n\nLoad:               winsup/configure.in,v ....................* 27704 of 27704\n\n\nBut it seems it now starts to really consume memory:\n\njbglaw@bixie:~/bin$ ps axflwww|egrep '(VSZ|parsecvs)'|grep -v grep\nF   UID   PID  PPID PRI  NI    VSZ   RSS WCHAN  STAT TTY        TIME COMMAND\n0  1000 15564 22879  18   0 2805084 549996 finish T  pts/10    30:51 |       \\_ parsecvs\n\nHow well does this work with even larger repositories?\n\nMfG, JBG\n\n-- \nJan-Benedict Glaw       jbglaw@lug-owl.de    . +49-172-7608481             _ O _\n\"Eine Freie Meinung in  einem Freien Kopf    | Gegen Zensur | Gegen Krieg  _ _ O\n für einen Freien Staat voll Freier Bürger\"  | im Internet! |   im Irak!   O O O\nret = do_actions((curr | FREE_SPEECH) & ~(NEW_COPYRIGHT_LAW | DRM | TCPA));\n"},{"id":"18422","messageId":"1144334896.2303.259.camel@neko.keithp.com","threadId":"3809","inReplyTo":"20060406120812.GO13324@lug-owl.de","subject":"Re: Fixes to parsecvs","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-04-06T14:48:16Z","receivedAt":"2006-04-06T14:48:16Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Thu, 2006-04-06 at 14:08 +0200, Jan-Benedict Glaw wrote:\n> On Wed, 2006-04-05 23:36:32 -0700, Keith Packard <keithp@keithp.com> wrote:\n> > note, parsecvs remains available from:\n> > \n> > \tgit://git.freedesktop.org/~keithp/parsecvs\n> \n> It now compiles out-of-the-box for me, nice work.\n\ncool\n\n> \n> However, it would be nice if you'd add a short description about how\n> to use it. Something like this:\n\nI'd rather just fix the usage to be more sane; that shouldn't take but a\nfew minutes...\n\n> I just ran it against a locally rsync'ed copy of the Binutils ,v\n> files. Looging at the progress bar, it is bascally ready:\n> \n> \n> Load:               winsup/configure.in,v ....................* 27704 of 27704\n\nNow all of the ,v files have been parsed and each revision placed in\nthe .git repository as a blob.\n\n> But it seems it now starts to really consume memory:\n\nYeah, it's doing the change set computation, which is not very space\nefficient; it computes the entire set of files at each commit which can\ntake 'a bit' of space with a large number of files over a long period of\ntime. Obviously computing revision deltas and saving those would make it\nuse a lot less memory.\n\n> jbglaw@bixie:~/bin$ ps axflwww|egrep '(VSZ|parsecvs)'|grep -v grep\n> F   UID   PID  PPID PRI  NI    VSZ   RSS WCHAN  STAT TTY        TIME COMMAND\n> 0  1000 15564 22879  18   0 2805084 549996 finish T  pts/10    30:51 |       \\_ parsecvs\n\nI'd run a large repository on a large machine; I managed to get\npostgresql to run on my laptop (615M CVS with 6000 files), but anything\nlarger I'd probably want to get it onto a big enough machine. The\nquestion is whether it needs to be more efficient so that people can\nconstantly convert repositories or whether moving the repository to a\nsufficiently large machine for the one-time conversion is 'good enough'.\n\n> How well does this work with even larger repositories?\n\npostgresql is the largest I've run; starting with a 615M CVS repository,\nit built a 1.7G .git tree, which packed down to 125M.\n\n-- \nkeith.packard@intel.com\n"},{"id":"18423","messageId":"Pine.LNX.4.63.0604061723410.23681@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"3809","inReplyTo":"1144334896.2303.259.camel@neko.keithp.com","subject":"Re: Fixes to parsecvs","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-04-06T15:26:14Z","receivedAt":"2006-04-06T15:26:14Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Thu, 6 Apr 2006, Keith Packard wrote:\n\n> On Thu, 2006-04-06 at 14:08 +0200, Jan-Benedict Glaw wrote:\n> \n> > But it seems it now starts to really consume memory:\n> \n> The question is whether it needs to be more efficient so that people can \n> constantly convert repositories or whether moving the repository to a \n> sufficiently large machine for the one-time conversion is 'good enough'.\n\nKeep in mind that there are many more valid uses for tracking a CVS \nrepository than to import it once.\n\nCiao,\nDscho\n"},{"id":"18424","messageId":"20060406160921.GU13324@lug-owl.de","threadId":"3809","inReplyTo":"Pine.LNX.4.63.0604061723410.23681@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Fixes to parsecvs","fromName":"Jan-Benedict Glaw","fromEmail":"jbglaw@lug-owl.de","sentAt":"2006-04-06T16:09:21Z","receivedAt":"2006-04-06T16:09:21Z","isPatch":false,"sender":{"key":"jbglaw@lug-owl.de","avatar":null},"body":"On Thu, 2006-04-06 17:26:14 +0200, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:\n> On Thu, 6 Apr 2006, Keith Packard wrote:\n> > On Thu, 2006-04-06 at 14:08 +0200, Jan-Benedict Glaw wrote:\n> > > But it seems it now starts to really consume memory:\n> > The question is whether it needs to be more efficient so that people can \n> > constantly convert repositories or whether moving the repository to a \n> > sufficiently large machine for the one-time conversion is 'good enough'.\n> \n> Keep in mind that there are many more valid uses for tracking a CVS \n> repository than to import it once.\n\nEven the most simplest usage case reveals this. (It's also what I'm\nabout to do the the converted GCC repository.)\n\nGet the repo, locally track the changes (so the importet branches are\nall like \"vendor branches\") and do own work in local branches.\n\nI'll do this eg. to be able to easily re-diff patches, which I want to\nput into GIT, just because it's so much more convenient than SVN.\nHowever, this is only possible because I'm able to keep track of\nupstream SVN changes. They probably won't change their SCM again, just\nafter they've introduced SVN.\n\nMfG, JBG\n\n-- \nJan-Benedict Glaw       jbglaw@lug-owl.de    . +49-172-7608481             _ O _\n\"Eine Freie Meinung in  einem Freien Kopf    | Gegen Zensur | Gegen Krieg  _ _ O\n für einen Freien Staat voll Freier Bürger\"  | im Internet! |   im Irak!   O O O\nret = do_actions((curr | FREE_SPEECH) & ~(NEW_COPYRIGHT_LAW | DRM | TCPA));\n"},{"id":"18426","messageId":"1144344979.2303.263.camel@neko.keithp.com","threadId":"3809","inReplyTo":"Pine.LNX.4.63.0604061723410.23681@wbgn013.biozentrum.uni-wuerzburg.de","subject":"Re: Fixes to parsecvs","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-04-06T17:36:19Z","receivedAt":"2006-04-06T17:36:19Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Thu, 2006-04-06 at 17:26 +0200, Johannes Schindelin wrote:\n\n> Keep in mind that there are many more valid uses for tracking a CVS \n> repository than to import it once.\n\nSure, but we should fix parsecvs to handle incremental CVS tracking if\nthat's one of the goals for this utility. git-cvsimport does this by\nskipping commits earlier than a fixed time; if we did that, we'd\neliminate the huge memory usage except for initial imports. I haven't\nconsidered how this might be done in detail yet; I have no personal need\nfor this functionality.\n\n-- \nkeith.packard@intel.com\n"},{"id":"18428","messageId":"20060406181502.GA15741@blackbean.org","threadId":"3809","inReplyTo":"1144305392.2303.240.camel@neko.keithp.com","subject":"Re: parsecvs tool now creates git repositories","fromName":"Jim Radford","fromEmail":"radford@blackbean.org","sentAt":"2006-04-06T18:15:02Z","receivedAt":"2006-04-06T18:15:02Z","isPatch":false,"sender":{"key":"radford@blackbean.org","avatar":null},"body":"Hi Keith,\n\nHere's one more build patch.  For some reason the Fedora lex doesn't\nwant a space after the -o.\n\nAlmost all of the errors I was seeing in the last version were fixed\nwith your \"branches that don't get merged back to the trunk\" fix.\n\nThanks,\n-Jim\n\ndiff --git a/Makefile b/Makefile\nindex 4ca6ffd..137ed34 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -4,7 +4,7 @@ GCC_WARNINGS3=-Wnested-externs -fno-stri\n GCC_WARNINGS=$(GCC_WARNINGS1) $(GCC_WARNINGS2) $(GCC_WARNINGS3)\n CFLAGS=-O0 -g $(GCC_WARNINGS)\n YFLAGS=-d -l\n-LFLAGS=-l -o lex.c\n+LFLAGS=-l -olex.c\n\n SRCS=gram.y lex.l cvs.h parsecvs.c cvsutil.c \\\n        revlist.c atom.c revcvs.c git.c gitutil.c\n"},{"id":"18430","messageId":"1144354356.2303.270.camel@neko.keithp.com","threadId":"3809","inReplyTo":"20060406181502.GA15741@blackbean.org","subject":"Re: parsecvs tool now creates git repositories","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-04-06T20:12:36Z","receivedAt":"2006-04-06T20:12:36Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Thu, 2006-04-06 at 11:15 -0700, Jim Radford wrote:\n> Hi Keith,\n> \n> Here's one more build patch.  For some reason the Fedora lex doesn't\n> want a space after the -o.\n\nI probably shouldn't even use the -o flag; all it does is change the\n#line directives in the output file to point at lex.c instead of\n<stdout>. I'm sure it'll break something.\n\n> Almost all of the errors I was seeing in the last version were fixed\n> with your \"branches that don't get merged back to the trunk\" fix.\n\nThat's good news at least.\n\n-- \nkeith.packard@intel.com\n"},{"id":"18434","messageId":"46a038f90604061451m4522e3f3qceae2331751a307c@mail.gmail.com","threadId":"3809","inReplyTo":"1144354356.2303.270.camel@neko.keithp.com","subject":"Re: parsecvs tool now creates git repositories","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-04-06T21:51:36Z","receivedAt":"2006-04-06T21:51:36Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 4/7/06, Keith Packard <keithp@keithp.com> wrote:\n> > Almost all of the errors I was seeing in the last version were fixed\n> > with your \"branches that don't get merged back to the trunk\" fix.\n>\n> That's good news at least.\n\nI'm re-running my import of Moodle's cvs (20K commits) with the newer\nparsecvs. The previous attempt looked very good except that\n\n - file additions were recorded with one-commit-per-file. I am not\nsure how rcs is recording these, but hte user does enter a common\nmessage at \"commit\" time. Perhaps the file addition action could be\nignored then?\n\n - some tags made on a branch show up in HEAD. This may be due to\npartial-tree branches, but I am not sure.\n\ncheers\n\n\nm\n"},{"id":"18436","messageId":"1144361968.2303.288.camel@neko.keithp.com","threadId":"3809","inReplyTo":"46a038f90604061451m4522e3f3qceae2331751a307c@mail.gmail.com","subject":"Re: parsecvs tool now creates git repositories","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-04-06T22:19:28Z","receivedAt":"2006-04-06T22:19:28Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Fri, 2006-04-07 at 09:51 +1200, Martin Langhoff wrote:\n\n>  - file additions were recorded with one-commit-per-file. I am not\n> sure how rcs is recording these, but hte user does enter a common\n> message at \"commit\" time. Perhaps the file addition action could be\n> ignored then?\n\nIf the log message is identical, and the dates are in-range, parsecvs\n\"should\" put the adds in the same commit. \n\n>  - some tags made on a branch show up in HEAD. This may be due to\n> partial-tree branches, but I am not sure.\n\nFinding branch points is not perfect; it's complicated by bizzarre\nbehaviour when adding files and casual CVS changes which make precise\nbranch points hard to detect. Can I get at this repository to play with?\nI'd like to see if we can't get the branch point detection more\naccurate.\n\n-- \nkeith.packard@intel.com\n"},{"id":"18437","messageId":"46a038f90604061622s5a7bee4eq6666d9b3796f70f6@mail.gmail.com","threadId":"3809","inReplyTo":"1144361968.2303.288.camel@neko.keithp.com","subject":"Re: parsecvs tool now creates git repositories","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-04-06T23:22:49Z","receivedAt":"2006-04-06T23:22:49Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 4/7/06, Keith Packard <keithp@keithp.com> wrote:\n> On Fri, 2006-04-07 at 09:51 +1200, Martin Langhoff wrote:\n>\n> >  - file additions were recorded with one-commit-per-file. I am not\n> > sure how rcs is recording these, but hte user does enter a common\n> > message at \"commit\" time. Perhaps the file addition action could be\n> > ignored then?\n>\n> If the log message is identical, and the dates are in-range, parsecvs\n> \"should\" put the adds in the same commit.\n\nparsecvs is committing them with the \"added file foo.x\" message, not\nthe actual commit message.\n\n> >  - some tags made on a branch show up in HEAD. This may be due to\n> > partial-tree branches, but I am not sure.\n>\n> Finding branch points is not perfect; it's complicated by bizzarre\n> behaviour when adding files and casual CVS changes which make precise\n> branch points hard to detect. Can I get at this repository to play with?\n\nI fetch it with something along the lines of...\n\nwhile ( true ) ; do\n     wget -qc http://cvs.sourceforge.net/cvstarballs/moodle-cvsroot.tar.bz2 &&\nbreak\n     sleep 5\ndone\n\nand then import the \"moodle\" module.\n\ncheers,\n\n\nm\n"},{"id":"18444","messageId":"1144394697.2303.307.camel@neko.keithp.com","threadId":"3809","inReplyTo":"46a038f90604061622s5a7bee4eq6666d9b3796f70f6@mail.gmail.com","subject":"Re: parsecvs tool now creates git repositories","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-04-07T07:24:57Z","receivedAt":"2006-04-07T07:24:57Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Fri, 2006-04-07 at 11:22 +1200, Martin Langhoff wrote:\n\n> parsecvs is committing them with the \"added file foo.x\" message, not\n> the actual commit message.\n\nheh. my cvs repositories are all so kludged that no files have ever been\nadded, it appears. I'll fix this when I've got a copy of the moodle\nrepository. sf.net is as useful as always.\n\nI suspect the change is as simple as checking the format of the log\nmessage and time time stamps of the commits and then just dropping the\n1.1 revision from the tree entirely.\n\n-- \nkeith.packard@intel.com\n"},{"id":"18522","messageId":"20060409231754.GB13138@electric-eye.fr.zoreil.com","threadId":"3809","inReplyTo":"1144334896.2303.259.camel@neko.keithp.com","subject":"Re: Fixes to parsecvs","fromName":"Francois Romieu","fromEmail":"romieu@fr.zoreil.com","sentAt":"2006-04-09T23:17:54Z","receivedAt":"2006-04-09T23:17:54Z","isPatch":false,"sender":{"key":"romieu@fr.zoreil.com","avatar":null},"body":"Keith Packard <keithp@keithp.com> :\n[...]\n> > How well does this work with even larger repositories?\n> \n> postgresql is the largest I've run; starting with a 615M CVS repository,\n> it built a 1.7G .git tree, which packed down to 125M.\n\nAs a datapoint, I gave parsecvs a try on a local CVS repository.\nThe repository weights 3.28 Go. It contains 53k files (45k non-attic).\n\n.git/objets grew from ~100k files at the end of the first pass to\n199k files (~11k commit). It took 18h on a 3GHz PIV with 2Go RAM.\nAfter 6 hours, 400 Mo were pushed to swap and parsecvs took 1.95 Go\nof RAM for itself. No significant swap activity. Swap grew to 900 Mo\nat end of run. A tarball (5 Mo) containing vmstat + size of objects\nis available at http://www.cogenit.fr/linux/misc/cvsparse-debug.tar.bz2\n\nI have interrupted 'git repack -a -d' after 6 hours.\n\n-- \nUeimor\n"}]}