{"thread":{"id":"4578","subject":"packs and trees","startedAt":"2006-06-20T05:57:31Z","lastAt":"2006-06-21T15:32:40Z","messageCount":10,"participants":["Jon Smirl","Martin Langhoff","Nicolas Pitre","Keith Packard","Linus Torvalds","David Lang"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"22104","messageId":"9e4733910606192257y1516e966t848a3b1e29e5667f@mail.gmail.com","threadId":"4578","inReplyTo":null,"subject":"packs and trees","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-20T05:57:31Z","receivedAt":"2006-06-20T05:57:31Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"Converting from CVS would be a lot more efficient if all of revisions\ncontained in a CVS file were written into git at the same time. So, if\nI extract complete revisions from 100 source files into git objects\nand then ask git to incremental pack, will git find all of the deltas\nand do a good job packing? Some of these files have thousands (50MB)\nof deltas. Also, note that I have not written any tree info into git\nyet.\n\nAfter all of the revisions are into git, I will follow up with the\ntree info and then repack all. How will the pack end up grouped,\nchronologically or will it still be sorted by file? It is not clear to\nme how the tree info interacts with the magic packing sauce.\n\nThe plan is to modify rcs2git from parsecvs to create all of the git\nobjects for the tree. It would be called by the cvs2svn code which\nwould track the object IDs through the changeset generation process.\nAt the end it will write all of the trees connecting the objects\ntogether.\n\ncvs2svn seems to do a good job at generating the trees. I am not\nexactly sure how the changeset detection algorithms in the three apps\ncompare, but cvs2svn is not having any trouble building changesets for\nMozilla. The other two apps have some issues, cvsps throws away some\nof the branches and parsecvs can't complete the analysis.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"22105","messageId":"46a038f90606192313l16b16132r1523f5e05ae1566a@mail.gmail.com","threadId":"4578","inReplyTo":"9e4733910606192257y1516e966t848a3b1e29e5667f@mail.gmail.com","subject":"Re: packs and trees","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-20T06:13:21Z","receivedAt":"2006-06-20T06:13:21Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/20/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> The plan is to modify rcs2git from parsecvs to create all of the git\n> objects for the tree.\n\nSounds like a good plan. Have you seen recent discussions about it\nbeing impossible to repack usefully when you don't have trees (and\nresulting performance problems on ext3).\n\n> cvs2svn seems to do a good job at generating the trees.\n\nNo doubt. Gut the last stage, and use all the data in the intermediate\nDBs to run a git import. It's a great plan, and if you can understand\nthat Python code... all yours ;-)\n\n> exactly sure how the changeset detection algorithms in the three apps\n> compare, but cvs2svn is not having any trouble building changesets for\n> Mozilla. The other two apps have some issues, cvsps throws away some\n> of the branches and parsecvs can't complete the analysis.\n\nHave you tried a recent parsecvs from Keith's tree? There's been quite\na bit of activity there too. And Keith's interested in sorting out\nincremental imports too, which you need for a reasonable Moz\ntransition plan as well.\n\ncheers,\n\n\n\nmartin\n"},{"id":"22140","messageId":"9e4733910606200735u5741a9adr83264ae7d51dd37@mail.gmail.com","threadId":"4578","inReplyTo":"46a038f90606192313l16b16132r1523f5e05ae1566a@mail.gmail.com","subject":"Re: packs and trees","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-20T14:35:49Z","receivedAt":"2006-06-20T14:35:49Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/20/06, Martin Langhoff <martin.langhoff@gmail.com> wrote:\n> On 6/20/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > The plan is to modify rcs2git from parsecvs to create all of the git\n> > objects for the tree.\n>\n> Sounds like a good plan. Have you seen recent discussions about it\n> being impossible to repack usefully when you don't have trees (and\n> resulting performance problems on ext3).\n\nNo, I will look back in the archives.  If needed we can do a repack\nafter each file is added. I would hope that git can handle a repack\nwhen the new stuff is 100% deltas from a single file.\n\nIf I can't pack the exploded deltas need about 35GB disk space. That\nis an awful lot to feed to pack all at once, but it will have trees,\n\n>\n> > cvs2svn seems to do a good job at generating the trees.\n>\n> No doubt. Gut the last stage, and use all the data in the intermediate\n> DBs to run a git import. It's a great plan, and if you can understand\n> that Python code... all yours ;-)\n\nHow hard would it be to adjust cvsps to use cvs2svn's algorithm for\ngrouping the changesets? I'd rather do this in a C app but I haven't\nfigured out the guts of parsecvs or cvsps well enough to change the\nalgorithms. There is no requirement to use external databases, sorting\neverything in RAM is fine.\n\nIf you are interested in changing the cvsps grouping algorithm I can\nlook at moding it to write out the revisions as are they are parsed.\nThen you only need to save the git sha1 in memory instead of the\nfile:rev when sorting.\n\n> > exactly sure how the changeset detection algorithms in the three apps\n> > compare, but cvs2svn is not having any trouble building changesets for\n> > Mozilla. The other two apps have some issues, cvsps throws away some\n> > of the branches and parsecvs can't complete the analysis.\n>\n> Have you tried a recent parsecvs from Keith's tree? There's been quite\n> a bit of activity there too. And Keith's interested in sorting out\n> incremental imports too, which you need for a reasonable Moz\n> transition plan as well.\n\nKeith's parsecvs run ended up in a loop and mine hit a parsecvs error\nand then had memory corruption after about eight hours. That was last\nweek,  I just checked the logs and I don't see any comments about\nfixing it.\n\nEven after spending eight hours building the changeset info iit is\nstill going to take it a couple of days to retrieve the versions one\nat a time and write them to git. Reparsing 50MB delta files n^2/2\ntimes is a major bottleneck for all three programs.\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"22143","messageId":"Pine.LNX.4.64.0606201102410.3377@localhost.localdomain","threadId":"4578","inReplyTo":"46a038f90606192313l16b16132r1523f5e05ae1566a@mail.gmail.com","subject":"Re: packs and trees","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-20T15:03:24Z","receivedAt":"2006-06-20T15:03:24Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Tue, 20 Jun 2006, Martin Langhoff wrote:\n\n> On 6/20/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > The plan is to modify rcs2git from parsecvs to create all of the git\n> > objects for the tree.\n> \n> Sounds like a good plan. Have you seen recent discussions about it\n> being impossible to repack usefully when you don't have trees (and\n> resulting performance problems on ext3).\n\nWhat do you mean?\n\n\nNicolas\n"},{"id":"22144","messageId":"1150816728.5382.27.camel@neko.keithp.com","threadId":"4578","inReplyTo":"9e4733910606200735u5741a9adr83264ae7d51dd37@mail.gmail.com","subject":"Re: packs and trees","fromName":"Keith Packard","fromEmail":"keithp@keithp.com","sentAt":"2006-06-20T15:18:48Z","receivedAt":"2006-06-20T15:18:48Z","isPatch":false,"sender":{"key":"keithp@keithp.com","avatar":"https://gravatar.com/avatar/fa1f479cdd51322fe86215c955a81d296bbf66a1fe625f8a12d87a8ec7faf648?d=mp&s=160"},"body":"On Tue, 2006-06-20 at 10:35 -0400, Jon Smirl wrote:\n\n> Keith's parsecvs run ended up in a loop and mine hit a parsecvs error\n> and then had memory corruption after about eight hours. That was last\n> week,  I just checked the logs and I don't see any comments about\n> fixing it.\n\nYeah, I'm rewriting the tool; the current codebase isn't supportable.\n\n> Even after spending eight hours building the changeset info iit is\n> still going to take it a couple of days to retrieve the versions one\n> at a time and write them to git. Reparsing 50MB delta files n^2/2\n> times is a major bottleneck for all three programs.\n\nThe eight hours in question *were* writing out the deltas and packing\nthe resulting trees. All that remained was to construct actual commit\nobjects and write them out. \n\nThe problem was that parsecvs's internals are structured so that this\nprocesses would take a large amount of memory, so I'm reworking the code\nto free stuff as it goes along.\n\nWith a rewritten parsecvs, I'm hoping to be able to steal the algorithms\nfrom cvs2svn and stick those in place. Then work on truncating the\nhistory so it can deal with incremental updates to the repository, which\nI think will be straightforward if we stick a few breadcrumbs in the git\nrepository to recover state from.\n\n-- \nkeith.packard@intel.com\n"},{"id":"22149","messageId":"9e4733910606200933p2e802954rdf50d5f0ac037677@mail.gmail.com","threadId":"4578","inReplyTo":"1150816728.5382.27.camel@neko.keithp.com","subject":"Re: packs and trees","fromName":"Jon Smirl","fromEmail":"jonsmirl@gmail.com","sentAt":"2006-06-20T16:33:10Z","receivedAt":"2006-06-20T16:33:10Z","isPatch":false,"sender":{"key":"jonsmirl@gmail.com","avatar":"https://gravatar.com/avatar/cff3bf5bfdfa6708b905712ff91f0f9b8aaca161659f38c02b787920d5d28b7e?d=mp&s=160"},"body":"On 6/20/06, Keith Packard <keithp@keithp.com> wrote:\n> > Even after spending eight hours building the changeset info iit is\n> > still going to take it a couple of days to retrieve the versions one\n> > at a time and write them to git. Reparsing 50MB delta files n^2/2\n> > times is a major bottleneck for all three programs.\n>\n> The eight hours in question *were* writing out the deltas and packing\n> the resulting trees. All that remained was to construct actual commit\n> objects and write them out.\n>\n> The problem was that parsecvs's internals are structured so that this\n> processes would take a large amount of memory, so I'm reworking the code\n> to free stuff as it goes along.\n\nHow about writing out all of the revisions from the cvs file using the\nyacc code the first time the file is encountered and parsed. Then you\nonly have to track git IDs and not all of those cumbersome CVS rev\nnumbers. When I was profiling parsecvs the hottest parts of the code\nwere extracting the revisions and comparing cvs rev numbers. Since the\ngit IDs are fixed size they work well in arrays and with pointer\ncompares for sorting. With the right data structure you should be able\nto eliminate the CVS rev numbers that are so slow to deal with.\n\nThere are about 1M revisions in moz cvs. At eight byes for an ID and\neight bytes for a timestamp that is 16MB if ordering is achieved via\narrays. All of the symbols fit into 400K including pointers to their\nrevision. If the revs are written out as they are encountered there is\nno need to save file names, but you do need one rev structure per\nfile. Throw in some more memory for relationship pointers. All of this\nshould fit into less than 100MB RAM.\n\n>\n> With a rewritten parsecvs, I'm hoping to be able to steal the algorithms\n> from cvs2svn and stick those in place. Then work on truncating the\n> history so it can deal with incremental updates to the repository, which\n> I think will be straightforward if we stick a few breadcrumbs in the git\n> repository to recover state from.\n>\n> --\n> keith.packard@intel.com\n>\n>\n> -----BEGIN PGP SIGNATURE-----\n> Version: GnuPG v1.4.3 (GNU/Linux)\n>\n> iD8DBQBEmBHYQp8BWwlsTdMRAvKAAJ9im3xBdUowt9af+/MtoYDXsCHGtACaAtG4\n> GygX7WgiFOamLrnTMzWkIPE=\n> =28dp\n> -----END PGP SIGNATURE-----\n>\n>\n>\n\n\n-- \nJon Smirl\njonsmirl@gmail.com\n"},{"id":"22157","messageId":"46a038f90606201241x3dec242dicde245a24c3ab9ab@mail.gmail.com","threadId":"4578","inReplyTo":"Pine.LNX.4.64.0606201102410.3377@localhost.localdomain","subject":"Re: packs and trees","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2006-06-20T19:41:57Z","receivedAt":"2006-06-20T19:41:57Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 6/21/06, Nicolas Pitre <nico@cam.org> wrote:\n> On Tue, 20 Jun 2006, Martin Langhoff wrote:\n>\n> > On 6/20/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > > The plan is to modify rcs2git from parsecvs to create all of the git\n> > > objects for the tree.\n> >\n> > Sounds like a good plan. Have you seen recent discussions about it\n> > being impossible to repack usefully when you don't have trees (and\n> > resulting performance problems on ext3).\n>\n> What do you mean?\n\nI was thinking of the \"repacking disconnected objects\" thread, but now\nI see it did have a solution in listing all the objects and paths. I\ntake that back.\n\nIf you are asking about the ext3 performance problems, I think Linus\ndiscussed that a while ago, why unpacked repos are slow (in addition\nto huge), and there were some suggestions of using hashed directory\nindexes.\n\ncheers,\n\n\nmartin\n"},{"id":"22161","messageId":"Pine.LNX.4.64.0606201650180.3377@localhost.localdomain","threadId":"4578","inReplyTo":"46a038f90606201241x3dec242dicde245a24c3ab9ab@mail.gmail.com","subject":"Re: packs and trees","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-06-20T20:51:11Z","receivedAt":"2006-06-20T20:51:11Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Wed, 21 Jun 2006, Martin Langhoff wrote:\n\n> On 6/21/06, Nicolas Pitre <nico@cam.org> wrote:\n> > On Tue, 20 Jun 2006, Martin Langhoff wrote:\n> > \n> > > On 6/20/06, Jon Smirl <jonsmirl@gmail.com> wrote:\n> > > > The plan is to modify rcs2git from parsecvs to create all of the git\n> > > > objects for the tree.\n> > >\n> > > Sounds like a good plan. Have you seen recent discussions about it\n> > > being impossible to repack usefully when you don't have trees (and\n> > > resulting performance problems on ext3).\n> > \n> > What do you mean?\n> \n> I was thinking of the \"repacking disconnected objects\" thread, but now\n> I see it did have a solution in listing all the objects and paths. I\n> take that back.\n\nOK.\n\n\nNicolas\n"},{"id":"22172","messageId":"Pine.LNX.4.64.0606202046290.5498@g5.osdl.org","threadId":"4578","inReplyTo":"46a038f90606201241x3dec242dicde245a24c3ab9ab@mail.gmail.com","subject":"Re: packs and trees","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-06-21T03:54:01Z","receivedAt":"2006-06-21T03:54:01Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 21 Jun 2006, Martin Langhoff wrote:\n> \n> If you are asking about the ext3 performance problems, I think Linus\n> discussed that a while ago, why unpacked repos are slow (in addition\n> to huge), and there were some suggestions of using hashed directory\n> indexes.\n\nYes. I think most distros still default to nonhashed directories, but for \nany large-directory case you really want to turn on hashing. \n\nI forget the exact details, it's somethng like\n\n\ttune2fs -O dir_index\n\nor something to turn it on (if I remember correctly, that will only affect \nany directories then created after that, but you can effect that by just \ndoing a \"git repack -a -d\" which will remove all old object directories, \nand now subsequent directories will be done with indexing on).\n\nPersonally, I just ended up using packs extensively, so I think I'm still \nrunning without indexing on all my machines ;)\n\n\t\tLinus\n"},{"id":"22216","messageId":"Pine.LNX.4.63.0606210830140.6305@qynat.qvtvafvgr.pbz","threadId":"4578","inReplyTo":"Pine.LNX.4.64.0606202046290.5498@g5.osdl.org","subject":"Re: packs and trees","fromName":"David Lang","fromEmail":"dlang@digitalinsight.com","sentAt":"2006-06-21T15:32:40Z","receivedAt":"2006-06-21T15:32:40Z","isPatch":false,"sender":{"key":"dlang@digitalinsight.com","avatar":null},"body":"there are performance penalties in some cases when you use directory hashing. I \ndid some tests last year of lots of small files in a directory tree (X/X/X/XXX \nof 1k files) and I found that turning on directory hashing actually slowed down \nthe creation of this tree from a tarball significantly (I also found that in \nthis case the ext2 allocation strategy was significantly better then the ext3 \none), so test on your system.\n\nDavid Lang\n\n\n  On Tue, 20 Jun 2006, Linus Torvalds wrote:\n\n> Date: Tue, 20 Jun 2006 20:54:01 -0700 (PDT)\n> From: Linus Torvalds <torvalds@osdl.org>\n> To: Martin Langhoff <martin.langhoff@gmail.com>\n> Cc: Nicolas Pitre <nico@cam.org>, Jon Smirl <jonsmirl@gmail.com>,\n>     git <git@vger.kernel.org>\n> Subject: Re: packs and trees\n> \n>\n>\n> On Wed, 21 Jun 2006, Martin Langhoff wrote:\n>>\n>> If you are asking about the ext3 performance problems, I think Linus\n>> discussed that a while ago, why unpacked repos are slow (in addition\n>> to huge), and there were some suggestions of using hashed directory\n>> indexes.\n>\n> Yes. I think most distros still default to nonhashed directories, but for\n> any large-directory case you really want to turn on hashing.\n>\n> I forget the exact details, it's somethng like\n>\n> \ttune2fs -O dir_index\n>\n> or something to turn it on (if I remember correctly, that will only affect\n> any directories then created after that, but you can effect that by just\n> doing a \"git repack -a -d\" which will remove all old object directories,\n> and now subsequent directories will be done with indexing on).\n>\n> Personally, I just ended up using packs extensively, so I think I'm still\n> running without indexing on all my machines ;)\n>\n> \t\tLinus\n> -\n> To unsubscribe from this list: send the line \"unsubscribe git\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n>\n"}]}