{"thread":{"id":"19701","subject":"Best way to merge two repos with same content, different history","startedAt":"2009-06-05T16:30:32Z","lastAt":"2009-06-19T09:52:59Z","messageCount":12,"participants":["Kelly F. Hickel","Rostislav Svoboda","Avery Pennarun","Markus Heidelberg","Robin H. Johnson","Michael Haggerty"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"115539","messageId":"63BEA5E623E09F4D92233FB12A9F794303117DBF@emailmn.mqsoftware.com","threadId":"19701","inReplyTo":null,"subject":"Best way to merge two repos with same content, different history","fromName":"Kelly F. Hickel","fromEmail":"kfh@mqsoftware.com","sentAt":"2009-06-05T16:30:32Z","receivedAt":"2009-06-05T16:30:32Z","isPatch":false,"sender":{"key":"kfh@mqsoftware.com","avatar":null},"body":"Hi all,\n\tWe're converting out of CVS after 10 years... The cvs2git\nconversion takes around 4-5 days, and there doesn't seem to be any way\nto speed that up.  So, our current plan is to take the tip of the\nbranches that we need for the next week, and import each of those\nbranches into its own clean git repo (top-skim).  Then we'll work out of\nthose repos while the conversion is in process (presumably creating\nbranches in those repos as needed).  Once the conversion is finished,\nwe'll need to get all the work done in the top-skimmed repos merged into\nthe converted repo so that we end up with little to no developer down\ntime, and all of our history, pre and post conversion in one repo.\n\tI'm testing all of this in advance, of course, and the tricky\npart at the moment is how to \"stitch\" the commits from the top-skim\nrepos back onto the converted repo when the conversion is done.  The\nfile content of the initial commit for the skimmed repo is identical to\nthe last commit for the respective branch in the converted repo, but the\nSHA1s are different, presumably because the history of the content is\ndifferent.\n\n\tStated another way, I have two repositories, \"new\" and \"old\",\nwhere the files in the initial commit on branch \"B1\" in \"new\" have\nexactly the same content as the last commit on branch \"B1\" in \"old\".\nThere also exist various branches in \"new\" based on \"B1\".  I'd like to\nmerge all the commits from \"new\" into \"old\", but the SHA1s are\ndifferent, presumably because the history leading up to those points are\ndifferent.\n\n\tOther than using manually format-patch on every branch in new,\nthen applying the patches (presumably with regular old patch, since the\nancestor commit IDs won't match), is there any \"good\" way to merge \"new\"\ninto \"old\"?\n\nThanks,\t\n\n\n--\n\nKelly F. Hickel\nSenior Product Architect\nMQSoftware, Inc.\n952-345-8677 Office\n952-345-8721 Fax\nkfh@mqsoftware.com\nwww.mqsoftware.com\nCertified IBM SOA Specialty\nYour Full Service Provider for IBM WebSphere\nLearn more at www.mqsoftware.com \n"},{"id":"115541","messageId":"286817520906050953n1afed29cn6c85f219a0c9b8b5@mail.gmail.com","threadId":"19701","inReplyTo":"63BEA5E623E09F4D92233FB12A9F794303117DBF@emailmn.mqsoftware.com","subject":"Re: Best way to merge two repos with same content, different history","fromName":"Rostislav Svoboda","fromEmail":"rostislav.svoboda@gmail.com","sentAt":"2009-06-05T16:53:56Z","receivedAt":"2009-06-05T16:53:56Z","isPatch":false,"sender":{"key":"rostislav.svoboda@gmail.com","avatar":"https://gravatar.com/avatar/9486d31c83feb25557c616b462aaef410dfe21d03bdf08cb62ff4ba9bafcc8eb?d=mp&s=160"},"body":"On Fri, Jun 5, 2009 at 18:30, Kelly F. Hickel<kfh@mqsoftware.com> wrote:\n>        We're converting out of CVS after 10 years... The cvs2git\n> conversion takes around 4-5 days, and there doesn't seem to be any way\n> to speed that up.\n\ntry google:\n\"cvs2git migration - cloning CVS repository\"\n\nBost\n"},{"id":"115543","messageId":"32541b130906051001k1ea4d960m4fcf7679b5b4f740@mail.gmail.com","threadId":"19701","inReplyTo":"63BEA5E623E09F4D92233FB12A9F794303117DBF@emailmn.mqsoftware.com","subject":"Re: Best way to merge two repos with same content, different history","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2009-06-05T17:01:00Z","receivedAt":"2009-06-05T17:01:00Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Fri, Jun 5, 2009 at 12:30 PM, Kelly F. Hickel <kfh@mqsoftware.com> wrote:\n>        Stated another way, I have two repositories, \"new\" and \"old\",\n> where the files in the initial commit on branch \"B1\" in \"new\" have\n> exactly the same content as the last commit on branch \"B1\" in \"old\".\n> There also exist various branches in \"new\" based on \"B1\".  I'd like to\n> merge all the commits from \"new\" into \"old\", but the SHA1s are\n> different, presumably because the history leading up to those points are\n> different.\n>\n>        Other than using manually format-patch on every branch in new,\n> then applying the patches (presumably with regular old patch, since the\n> ancestor commit IDs won't match), is there any \"good\" way to merge \"new\"\n> into \"old\"?\n\nThe usual replacement for \"manually using format-patch\" is to use \"git\nrebase.\"  It does pretty much exactly what you're describing, assuming\nyou don't do too many complicated merges in the meantime.\n\nAnother option is to use the .git/info/grafts file.  Here's a brief\nintro: http://git.or.cz/gitwiki/GraftPoint\n\nYou'd use that to pretend the parent of your top-skimmed branch is\nactually the equivalent commit in your new branch.  Then could run\n\"git filter-branch\" to make the graft permanent, and get all your\nusers to switch to the new repository.\n\nOr you could skip the filter-branch stuff and keep the really hold\nhistory somewhere else, available for use if someone installs the\ngraft in their local repo.  This would lead to a smaller repository in\nthe general case.  (I gather that's what the Linux kernel does for\nper-2.6.11 versions.)\n\nHave fun,\n\nAvery\n"},{"id":"115544","messageId":"63BEA5E623E09F4D92233FB12A9F794303117DC1@emailmn.mqsoftware.com","threadId":"19701","inReplyTo":"286817520906050953n1afed29cn6c85f219a0c9b8b5@mail.gmail.com","subject":"RE: Best way to merge two repos with same content, different history","fromName":"Kelly F. Hickel","fromEmail":"kfh@mqsoftware.com","sentAt":"2009-06-05T17:10:30Z","receivedAt":"2009-06-05T17:10:30Z","isPatch":false,"sender":{"key":"kfh@mqsoftware.com","avatar":null},"body":"\n> -----Original Message-----\n> From: Rostislav Svoboda [mailto:rostislav.svoboda@gmail.com]\n> Sent: Friday, June 05, 2009 11:54 AM\n> To: Kelly F. Hickel\n> Cc: git@vger.kernel.org\n> Subject: Re: Best way to merge two repos with same content, different\n> history\n> \n> On Fri, Jun 5, 2009 at 18:30, Kelly F. Hickel<kfh@mqsoftware.com>\n> wrote:\n> >        We're converting out of CVS after 10 years... The cvs2git\n> > conversion takes around 4-5 days, and there doesn't seem to be any\n> way\n> > to speed that up.\n> \n> try google:\n> \"cvs2git migration - cloning CVS repository\"\n> \n> Bost\n\nBost, \n\tThanks, but I'm already working with a local copy of the CVS repo.  I've corresponded with Michael Haggerty about the time this takes, and there just doesn't seem to be any way to improve the speed, without making some fairly drastic changes to cvs2git.\n\nThanks,\nKelly\n"},{"id":"115545","messageId":"63BEA5E623E09F4D92233FB12A9F794303117DC2@emailmn.mqsoftware.com","threadId":"19701","inReplyTo":"32541b130906051001k1ea4d960m4fcf7679b5b4f740@mail.gmail.com","subject":"RE: Best way to merge two repos with same content, different history","fromName":"Kelly F. Hickel","fromEmail":"kfh@mqsoftware.com","sentAt":"2009-06-05T17:11:55Z","receivedAt":"2009-06-05T17:11:55Z","isPatch":false,"sender":{"key":"kfh@mqsoftware.com","avatar":null},"body":"> -----Original Message-----\n> From: Avery Pennarun [mailto:apenwarr@gmail.com]\n> Sent: Friday, June 05, 2009 12:01 PM\n> To: Kelly F. Hickel\n> Cc: git@vger.kernel.org\n> Subject: Re: Best way to merge two repos with same content, different\n> history\n> \n> On Fri, Jun 5, 2009 at 12:30 PM, Kelly F. Hickel <kfh@mqsoftware.com>\n> wrote:\n> >        Stated another way, I have two repositories, \"new\" and \"old\",\n> > where the files in the initial commit on branch \"B1\" in \"new\" have\n> > exactly the same content as the last commit on branch \"B1\" in \"old\".\n> > There also exist various branches in \"new\" based on \"B1\".  I'd like\n> to\n> > merge all the commits from \"new\" into \"old\", but the SHA1s are\n> > different, presumably because the history leading up to those points\n> are\n> > different.\n> >\n> >        Other than using manually format-patch on every branch in new,\n> > then applying the patches (presumably with regular old patch, since\n> the\n> > ancestor commit IDs won't match), is there any \"good\" way to merge\n> \"new\"\n> > into \"old\"?\n> \n> The usual replacement for \"manually using format-patch\" is to use \"git\n> rebase.\"  It does pretty much exactly what you're describing, assuming\n> you don't do too many complicated merges in the meantime.\n> \n> Another option is to use the .git/info/grafts file.  Here's a brief\n> intro: http://git.or.cz/gitwiki/GraftPoint\n> \n> You'd use that to pretend the parent of your top-skimmed branch is\n> actually the equivalent commit in your new branch.  Then could run\n> \"git filter-branch\" to make the graft permanent, and get all your\n> users to switch to the new repository.\n> \n> Or you could skip the filter-branch stuff and keep the really hold\n> history somewhere else, available for use if someone installs the\n> graft in their local repo.  This would lead to a smaller repository in\n> the general case.  (I gather that's what the Linux kernel does for\n> per-2.6.11 versions.)\n> \n> Have fun,\n> \n> Avery\n\nThanks Avery,\n\tThis appears to be just what I was looking for!  I'll fiddle with it a bit to see if I can convince it to work for me.\n\n\nThanks,\n\tKelly\n"},{"id":"115546","messageId":"200906051915.56115.markus.heidelberg@web.de","threadId":"19701","inReplyTo":"63BEA5E623E09F4D92233FB12A9F794303117DBF@emailmn.mqsoftware.com","subject":"Re: Best way to merge two repos with same content, different history","fromName":"Markus Heidelberg","fromEmail":"markus.heidelberg@web.de","sentAt":"2009-06-05T17:15:55Z","receivedAt":"2009-06-05T17:15:55Z","isPatch":false,"sender":{"key":"markus.heidelberg@web.de","avatar":"https://avatars.githubusercontent.com/u/6334512?v=4"},"body":"Kelly F. Hickel, 05.06.2009:\n> \tOther than using manually format-patch on every branch in new,\n> then applying the patches (presumably with regular old patch, since the\n> ancestor commit IDs won't match), is there any \"good\" way to merge \"new\"\n> into \"old\"?\n\nIf rebasing 'new' on top of 'old' isn't an option, then you could try:\n\n    $ git checkout new\n    $ git merge -s ours old\n\nIt's the other way round (not merging 'new' into 'old', but vice versa),\nbut there is now merge strategy \"theirs\".\n\nMarkus\n"},{"id":"115547","messageId":"286817520906051019k78dc002cg8006987a7258b6de@mail.gmail.com","threadId":"19701","inReplyTo":"63BEA5E623E09F4D92233FB12A9F794303117DC1@emailmn.mqsoftware.com","subject":"Re: Best way to merge two repos with same content, different history","fromName":"Rostislav Svoboda","fromEmail":"rostislav.svoboda@gmail.com","sentAt":"2009-06-05T17:19:51Z","receivedAt":"2009-06-05T17:19:51Z","isPatch":false,"sender":{"key":"rostislav.svoboda@gmail.com","avatar":"https://gravatar.com/avatar/9486d31c83feb25557c616b462aaef410dfe21d03bdf08cb62ff4ba9bafcc8eb?d=mp&s=160"},"body":"On Fri, Jun 5, 2009 at 19:10, Kelly F. Hickel<kfh@mqsoftware.com> wrote:\n> Bost,\n>        Thanks, but I'm already working with a local copy of the CVS repo.  I've corresponded with Michael Haggerty about the time this takes, and there just doesn't seem to be any way to improve the speed, without making some fairly drastic changes to cvs2git.\n\nuhm, see below\n\nBost\n\nOn Tue, Feb 3, 2009 at 14:55, Johannes\nSchindelin<Johannes.Schindelin@gmx.de> wrote:\n>> If you do not have filesystem access to your CVS repository, you might\n>> be able to clone it using CVSSuck [2,3].\n>\n> A substantially faster option would be to go with cvsclone:\n>\n>        http://samba.org/ftp/tridge/rtc/cvsclone.l\n>\n> (in my case, cvsclone was not only faster, but it actually worked, too,\n> which is more than I could say of CVSSuck).\n"},{"id":"115550","messageId":"robbat2-20090605T183716-227340397Z@orbis-terrarum.net","threadId":"19701","inReplyTo":"63BEA5E623E09F4D92233FB12A9F794303117DC1@emailmn.mqsoftware.com","subject":"Re: Best way to merge two repos with same content, different history","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-06-05T18:46:16Z","receivedAt":"2009-06-05T18:46:16Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Fri, Jun 05, 2009 at 12:10:30PM -0500, Kelly F. Hickel wrote:\n> Bost, \n> \tThanks, but I'm already working with a local copy of the CVS repo.\n> \tI've corresponded with Michael Haggerty about the time this takes,\n> \tand there just doesn't seem to be any way to improve the speed,\n> \twithout making some fairly drastic changes to cvs2git.\nI've been working with mhagger lately as it also pertains to the Gentoo\nconversion. We've made some very good progress.\n\nA couple of comments in that regard:\n- Make really sure your box is not short of RAM. Throw some measurement\n  tools onto there to see it. A couple of GiB is worthwhile. After we\n  found this early on, and switched boxes, we dropped from our initial\n  multiple days to 20 hours.\n- His latest ExternalBlobGenerator code (_NOT_ available in SVN yet)\n  reduced our pass1 time from 36204 seconds to 1598 seconds, with\n  a potential to be much faster now, as parallelization of part of that\n  is now trivial.\n- pass9 is still the remaining large time-eater for us. I've started to\n  look at it, but I haven't made any actual developments yet.\n\nWould you mind posting your cvs2svn stats like these?\nhttp://archives.gentoo.org/gentoo-scm/msg_b69b2f6ecee0ec7bb402d31b372b945b.xml\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"115551","messageId":"63BEA5E623E09F4D92233FB12A9F794303117DCB@emailmn.mqsoftware.com","threadId":"19701","inReplyTo":"robbat2-20090605T183716-227340397Z@orbis-terrarum.net","subject":"RE: Best way to merge two repos with same content, differenthistory","fromName":"Kelly F. Hickel","fromEmail":"kfh@mqsoftware.com","sentAt":"2009-06-05T19:06:25Z","receivedAt":"2009-06-05T19:06:25Z","isPatch":false,"sender":{"key":"kfh@mqsoftware.com","avatar":null},"body":"> -----Original Message-----\n> From: git-owner@vger.kernel.org [mailto:git-owner@vger.kernel.org] On\n> Behalf Of Robin H. Johnson\n> Sent: Friday, June 05, 2009 1:46 PM\n> To: Git Mailing List\n> Subject: Re: Best way to merge two repos with same content,\n> differenthistory\n> \n> On Fri, Jun 05, 2009 at 12:10:30PM -0500, Kelly F. Hickel wrote:\n> > Bost,\n> > \tThanks, but I'm already working with a local copy of the CVS\n> repo.\n> > \tI've corresponded with Michael Haggerty about the time this\n> takes,\n> > \tand there just doesn't seem to be any way to improve the speed,\n> > \twithout making some fairly drastic changes to cvs2git.\n> I've been working with mhagger lately as it also pertains to the\nGentoo\n> conversion. We've made some very good progress.\n> \n> A couple of comments in that regard:\n> - Make really sure your box is not short of RAM. Throw some\nmeasurement\n>   tools onto there to see it. A couple of GiB is worthwhile. After we\n>   found this early on, and switched boxes, we dropped from our initial\n>   multiple days to 20 hours.\n> - His latest ExternalBlobGenerator code (_NOT_ available in SVN yet)\n>   reduced our pass1 time from 36204 seconds to 1598 seconds, with\n>   a potential to be much faster now, as parallelization of part of\nthat\n>   is now trivial.\n> - pass9 is still the remaining large time-eater for us. I've started\nto\n>   look at it, but I haven't made any actual developments yet.\n> \n> Would you mind posting your cvs2svn stats like these?\n> http://archives.gentoo.org/gentoo-\n> scm/msg_b69b2f6ecee0ec7bb402d31b372b945b.xml\n> \n> --\n> Robin Hugh Johnson\n> Gentoo Linux Developer & Infra Guy\n> E-Mail     : robbat2@gentoo.org\n> GnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n\nRobin, \n\tThat's all good news, I have an 8 way box with 32gb of ram\nrunning a 64 bit Linux, a box with 4 gb of ram panics during the\nconversion.\n\nMy conversion data is below...\n\nThanks,\nKelly\n\n\ncvs2svn Statistics:\n------------------\nTotal CVS Files:             18488\nTotal CVS Revisions:        225208\nTotal CVS Branches:       15203751\nTotal CVS Tags:           39079236\nTotal Unique Tags:           11453\nTotal Unique Branches:        4364\nCVS Repos Size in KB:      3355895\nTotal SVN Commits:           49967\nFirst Revision Date:    Mon Nov  8 02:26:51 1999\nLast Revision Date:     Wed Apr 22 17:59:44 2009\n------------------\nTimings (seconds):\n------------------\n251546   pass1    CollectRevsPass\n     4   pass2    CleanMetadataPass\n   142   pass3    CollateSymbolsPass\n 53491   pass4    FilterSymbolsPass\n     4   pass5    SortRevisionSummaryPass\n   163   pass6    SortSymbolSummaryPass\n  8825   pass7    InitializeChangesetsPass\n   418   pass8    BreakRevisionChangesetCyclesPass\n   418   pass9    RevisionTopologicalSortPass\n  4256   pass10   BreakSymbolChangesetCyclesPass\n  4914   pass11   BreakAllChangesetCyclesPass\n  4575   pass12   TopologicalSortPass\n  3111   pass13   CreateRevsPass\n   270   pass14   SortSymbolsPass\n   154   pass15   IndexSymbolsPass\n  5517   pass16   OutputPass\n337808   total\n251783.89user 80800.42system 93:50:11elapsed 98%CPU (0avgtext+0avgdata\n0maxresident)k\n0inputs+0outputs (3major+3132264023minor)pagefaults 0swaps\n\n\ngit-fast-import statistics: \n---------------------------------------------------------------------\nAlloc'd objects:     415000\nTotal objects:       410079 (   2628078 duplicates                  )\n      blobs  :       152002 (     68715 duplicates     135125 deltas)\n      trees  :       213636 (   2559363 duplicates     164052 deltas)\n      commits:        44441 (         0 duplicates          0 deltas)\n      tags   :            0 (         0 duplicates          0 deltas)\nTotal branches:       15822 (      6184 loads     )\n      marks:     1073741824 (    265158 unique    )\n      atoms:          14807\nMemory total:         32402 KiB\n       pools:         16192 KiB\n     objects:         16210 KiB\n---------------------------------------------------------------------\npack_report: getpagesize()            =       4096\npack_report: core.packedGitWindowSize = 1073741824\npack_report: core.packedGitLimit      = 8589934592\npack_report: pack_used_ctr            =     311525\npack_report: pack_mmap_calls          =      13303\npack_report: pack_open_windows        =          1 /          1\npack_report: pack_mapped              =  403041230 /  403041230\n---------------------------------------------------------------------\n\n\ngit repack -a -d -f --depth=4000 --window=4000 && git pack-refs --all\nCounting objects: 409458, done.5/119582)   \nCompressing objects: 100% (119582/119582), done.\nWriting objects: 100% (128713/128713), done.\nTotal 128713 (delta 82330), reused 0 (delta 0)\nCompressing objects: 100% (384171/384171), done.\nWriting objects: 100% (409458/409458), done.\nTotal 409458 (delta 309214), reused 0 (delta 0)\n"},{"id":"115557","messageId":"robbat2-20090605T194802-473902673Z@orbis-terrarum.net","threadId":"19701","inReplyTo":"63BEA5E623E09F4D92233FB12A9F794303117DCB@emailmn.mqsoftware.com","subject":"Re: Best way to merge two repos with same content, differenthistory","fromName":"Robin H. Johnson","fromEmail":"robbat2@gentoo.org","sentAt":"2009-06-05T20:02:13Z","receivedAt":"2009-06-05T20:02:13Z","isPatch":false,"sender":{"key":"robbat2@gentoo.org","avatar":"https://avatars.githubusercontent.com/u/373898?v=4"},"body":"On Fri, Jun 05, 2009 at 02:06:25PM -0500, Kelly F. Hickel wrote:\n> Robin, \n> \tThat's all good news, I have an 8 way box with 32gb of ram\n> running a 64 bit Linux, a box with 4 gb of ram panics during the\n> conversion.\nThanks for your data.\n\nFor comparison, our conversion box is also 8-way, but only 16GiB RAM.\n\nI'm surprised at how long pass1 is for you, especially since you've got\na lot less CVS Files and CVS Revisions than the Gentoo repo (I do deduce\nthat your individual revisions are larger, averaging at 15KiB vs. our\n711 bytes).\n\nI think there's something odd in the total CVS branches/tags count\nhowever, as the counts there imply an average of 67 branches and 173\ntags per CVS revision. You might want to dig into that part manually and\nsee about it (not sure of your Python skills). That would probably cut\ndown both your pass1 and pass4 times significantly.\n\nHopefully mhagger will get the external blob stuff committed soon, I was\nworking on validating it's results. \n\nIn doing so discovered a testcase where RCSRevisionReader and\nCVSRevisionReader gave different output themselves, the latter (which is\ndocumented as more accurate otherwise) missing the contents of an entire\nfile. It's on the cvs2svn-dev mailing list now. Tracing that first,\nthereafter comparing it to the new Git side.\n\n> git repack -a -d -f --depth=4000 --window=4000 && git pack-refs --all\nDid those extreme depth/window values actually help size much? The\nGentoo ones actually didn't improve significantly over depth=window=50.\n\n-- \nRobin Hugh Johnson\nGentoo Linux Developer & Infra Guy\nE-Mail     : robbat2@gentoo.org\nGnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n"},{"id":"115559","messageId":"63BEA5E623E09F4D92233FB12A9F794303117DD2@emailmn.mqsoftware.com","threadId":"19701","inReplyTo":"robbat2-20090605T194802-473902673Z@orbis-terrarum.net","subject":"RE: Best way to merge two repos with same content, differenthistory","fromName":"Kelly F. Hickel","fromEmail":"kfh@mqsoftware.com","sentAt":"2009-06-05T20:08:55Z","receivedAt":"2009-06-05T20:08:55Z","isPatch":false,"sender":{"key":"kfh@mqsoftware.com","avatar":null},"body":"> -----Original Message-----\n> From: git-owner@vger.kernel.org [mailto:git-owner@vger.kernel.org] On\n> Behalf Of Robin H. Johnson\n> Sent: Friday, June 05, 2009 3:02 PM\n> To: Git Mailing List\n> Subject: Re: Best way to merge two repos with same content,\n> differenthistory\n> \n> On Fri, Jun 05, 2009 at 02:06:25PM -0500, Kelly F. Hickel wrote:\n> > Robin,\n> > \tThat's all good news, I have an 8 way box with 32gb of ram\n> running a\n> > 64 bit Linux, a box with 4 gb of ram panics during the conversion.\n> Thanks for your data.\n> \n> For comparison, our conversion box is also 8-way, but only 16GiB RAM.\n> \n> I'm surprised at how long pass1 is for you, especially since you've\ngot\n> a lot less CVS Files and CVS Revisions than the Gentoo repo (I do\n> deduce\n> that your individual revisions are larger, averaging at 15KiB vs. our\n> 711 bytes).\n> \n> I think there's something odd in the total CVS branches/tags count\n> however, as the counts there imply an average of 67 branches and 173\n> tags per CVS revision. You might want to dig into that part manually\n> and\n> see about it (not sure of your Python skills). That would probably cut\n> down both your pass1 and pass4 times significantly.\n\nRobin, I'm not much with python, so haven't dug into the code much at\nall. The numbers are high, although we do create a lot of branches (had\nto contribute a fix a year or two to CVS to get the branching time down\nfrom the 2.5 hours it was taking).  At one point I carefully examined\nthe symbol file that cvs2git was outputting and convinced myself that it\nwas doing the right thing, but that was awhile ago.\n\n> \n> Hopefully mhagger will get the external blob stuff committed soon, I\n> was\n> working on validating it's results.\n> \n> In doing so discovered a testcase where RCSRevisionReader and\n> CVSRevisionReader gave different output themselves, the latter (which\n> is\n> documented as more accurate otherwise) missing the contents of an\n> entire\n> file. It's on the cvs2svn-dev mailing list now. Tracing that first,\n> thereafter comparing it to the new Git side.\n> \n> > git repack -a -d -f --depth=4000 --window=4000 && git pack-refs\n--all\n> Did those extreme depth/window values actually help size much? The\n> Gentoo ones actually didn't improve significantly over\ndepth=window=50.\n\nI know that they were still (apparently) improving after the 200 mark,\nit took long enough at 200 that I just decided to crank the numbers way\nup and let it run over the weekend.\n\n> \n> --\n> Robin Hugh Johnson\n> Gentoo Linux Developer & Infra Guy\n> E-Mail     : robbat2@gentoo.org\n> GnuPG FP   : 11AC BA4F 4778 E3F6 E4ED  F38E B27B 944E 3488 4E85\n\nI'll be looking forward to a newer faster cvs2git, although I did just\nget the graft idea working, so not sure if we'll wait that long or not\n(would be nice not to have to muck around with it though).\n\nThanks,\nKelly\n"},{"id":"116619","messageId":"4A3B5FFB.1030001@alum.mit.edu","threadId":"19701","inReplyTo":"robbat2-20090605T194802-473902673Z@orbis-terrarum.net","subject":"Re: Best way to merge two repos with same content, differenthistory","fromName":"Michael Haggerty","fromEmail":"mhagger@alum.mit.edu","sentAt":"2009-06-19T09:52:59Z","receivedAt":"2009-06-19T09:52:59Z","isPatch":false,"sender":{"key":"mhagger@alum.mit.edu","avatar":"https://avatars.githubusercontent.com/u/119718?v=4"},"body":"Robin H. Johnson wrote:\n> [Regarding cvs2git] In doing so discovered a testcase where RCSRevisionReader and\n> CVSRevisionReader gave different output themselves, the latter (which is\n> documented as more accurate otherwise) missing the contents of an entire\n> file. It's on the cvs2svn-dev mailing list now. Tracing that first,\n> thereafter comparing it to the new Git side.\n\nIn case anybody is following this, the issue that Robin found was in\nrelease 2.2.0 of cvs2svn/cvs2git, but had already been fixed in the\ncurrent trunk version...\n\nConclusion 1: Please use the trunk version of cvs2git, not release\n2.2.0.  Trunk is usually the most stable version available, especially\nfor 2git conversions.\n\nConclusion 2: I've really got to get around to making a new release :-/\n\nMichael\n"}]}