{"thread":{"id":"6214","subject":"git-svnimport failed and now git-repack hates me","startedAt":"2007-01-03T23:52:30Z","lastAt":"2007-01-08T02:22:42Z","messageCount":55,"participants":["Chris Lee","Linus Torvalds","Shawn O. Pearce","Eric Wong","Randal L. Schwartz","Junio C Hamano","Sasha Khapyorsky","Johannes Schindelin","alan"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"30781","messageId":"204011cb0701031552j8292d23v950f828279702d3@mail.gmail.com","threadId":"6214","inReplyTo":null,"subject":"git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-03T23:52:30Z","receivedAt":"2007-01-03T23:52:30Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"So I'm using git 1.4.1, and I have been experimenting with importing\nthe KDE sources from Subversion using git-svnimport.\n\nFirst issue I ran into: On a machine with 4GB of RAM, when I tried to\ndo a full import, git-svnimport died after 309906 revisions, saying\nthat it couldn't fork.\n\nChecking `top` and `ps` revealed that there were no git-svnimport\nprocesses doing anything, but all of my 4G of RAM was still marked as\nused by the kernel. I had to do sysctl -w vm.drop_caches=3 to get it\nto free all the RAM that the svn import had used up.\n\nNow, after that, I tried doing `git-repack -a` because I wanted to see\nhow small the packed archive would be (before trying to continue\nimporting the rest of the revisions. There are at least another 100k\nrevisions that I should be able to import, eventually.)\n\nThe repack finished after about nine hours, but when I try to do a\ngit-verify-pack on it, it dies with this error message:\n\nerror: Packfile\n.git/objects/pack/pack-540263fe66ab9398cc796f000d52531a5c6f3df3.pack\nSHA1 mismatch with itself\n\nI get the same message from git-prune.\n\nAny ideas?\n"},{"id":"30783","messageId":"Pine.LNX.4.64.0701031737300.4989@woody.osdl.org","threadId":"6214","inReplyTo":"204011cb0701031552j8292d23v950f828279702d3@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-04T01:59:28Z","receivedAt":"2007-01-04T01:59:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 3 Jan 2007, Chris Lee wrote:\n>\n> So I'm using git 1.4.1, and I have been experimenting with importing\n> the KDE sources from Subversion using git-svnimport.\n\nAs one single _huge_ import? All the sub-projects together? I have to say, \nthat sounds pretty horrid.\n\n> First issue I ran into: On a machine with 4GB of RAM, when I tried to\n> do a full import, git-svnimport died after 309906 revisions, saying\n> that it couldn't fork.\n> \n> Checking `top` and `ps` revealed that there were no git-svnimport\n> processes doing anything, but all of my 4G of RAM was still marked as\n> used by the kernel. I had to do sysctl -w vm.drop_caches=3 to get it\n> to free all the RAM that the svn import had used up.\n\nI think that was just all cached, and all ok. The reason you didn't see \nany git-svnimport was that it had died off already, and all your memory \nwas just caches. You could just have left it alone, and the kernel would \nhave started re-using the memory for other things even without any \n\"drop_caches\". \n\nBut what you did there didn't make anything worse, it was just likely had \nno real impact.\n\nHowever, it does sound like git-svnimport probably acts like git-cvsimport \nused to, and just keeps too much in memory - so it's never going to act \nreally nicely..\n\nIt also looks like git-svnimport never repacks the repo, which is \nabsolutely horrible for performance on all levels. The CVS importer \nrepacks every one thousand commits or something like that.\n\n> Now, after that, I tried doing `git-repack -a` because I wanted to see\n> how small the packed archive would be (before trying to continue\n> importing the rest of the revisions. There are at least another 100k\n> revisions that I should be able to import, eventually.)\n\nI suspect you'd have been better off just re-starting, and using something \nlike\n\n\twhile :\n\tdo\n\t\tgit svnimport -l 1000 <...>\n\t\t.. figure out some way to decide if it's all done ..\n\t\tgit repack -d\n\tdone\n\nwhich would make svnimport act a bit  more sanely, and repack \nincrementally. That should make both the import much faster, _and_ avoid \nany insane big repack at the end (well, you'd still want to do a \"git \nrepack -a -d\" at the end to turn the many smaller packs into a bigger one, \nbut it would be nicer).\n\nHowever, I don't know what the proper magic is for svnimport to do that \nsane \"do it in chunks and tell when you're all done\". Or even better - to \njust make it repack properly and not keep everything in memory.\n\n> The repack finished after about nine hours, but when I try to do a\n> git-verify-pack on it, it dies with this error message:\n> \n> error: Packfile\n> .git/objects/pack/pack-540263fe66ab9398cc796f000d52531a5c6f3df3.pack\n> SHA1 mismatch with itself\n\nThat sounds suspiciously like the bug we had in out POWER sha1 \nimplementation that would generate the wrong SHA1 for any pack-file that \nwas over 512MB in size, due to an overflow in 32 bits (SHA1 does some \ncounting in _bits_, so 512MB is 4G _bits_),\n\nNow, I assume you're not on POWER (and we fixed that bug anyway - and I \nthink long before 1.4.1 too), but I could easily imagine the same bug in \nsome other SHA1 implementation (or perhaps _another_ overflow at the 1GB \nor 2GB mark..). I assume that the pack-file you had was something horrid..\n\nI hope this is with a 64-bit kernel and a 64-bit user space? That should \nlimit _some_ of the issues. But I would still not be surprised if your \nSHA1 libraries had some 32-bit (\"unsigned int\") or 31-bit (\"int\") limits \nin them somewhere - very few people do SHA1's over huge areas, and even \nwhen you do SHA1 on something like a DVD image (which is easily over any \n4GB limit), that tends to be done as many smaller calls to the SHA1 \nlibrary routines.\n\nJunio - I suspect \"pack-check.c\" really shouldn't try to do it as one \nsingle humungous \"SHA1_Update()\" call. It showed one bug on PPC, I \nwouldn't be surprised if it's implicated now on some other architecture. \n\nShawn - does the pack-file-windowing thing already change that? I'm too \nlazy to check..\n\nAs to who knows how to fix git-svnimport to do something saner, I have no \nclue.. Sasha seems to have touched it last. Sasha?\n\n\t\tLinus\n"},{"id":"30784","messageId":"20070104020652.GB18206@spearce.org","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701031737300.4989@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-01-04T02:06:52Z","receivedAt":"2007-01-04T02:06:52Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> wrote:\n> Junio - I suspect \"pack-check.c\" really shouldn't try to do it as one \n> single humungous \"SHA1_Update()\" call. It showed one bug on PPC, I \n> wouldn't be surprised if it's implicated now on some other architecture. \n\nIt used to do it as one big SHA1_Update() call...\n \n> Shawn - does the pack-file-windowing thing already change that? I'm too \n> lazy to check..\n\nBut with the mmap window thing in `next` it does it in window\nunits only.  Which the user could configure to be huge, or could\nconfigure to be sane.  The default when using mmap() is 32 MiB;\n1 MiB when using pread() and git_mmap().\n\n-- \nShawn.\n"},{"id":"30816","messageId":"204011cb0701031816hda8af9bw4d4a469c2b111339@mail.gmail.com","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701031737300.4989@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-04T02:16:51Z","receivedAt":"2007-01-04T02:16:51Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"On 1/3/07, Linus Torvalds <torvalds@osdl.org> wrote:\n> > So I'm using git 1.4.1, and I have been experimenting with importing\n> > the KDE sources from Subversion using git-svnimport.\n>\n> As one single _huge_ import? All the sub-projects together? I have to say,\n> that sounds pretty horrid.\n\nUnfortunately, that's how the KDE repo is organized. (I tried arguing\nagainst this when they were going to do the original import, but I\nlost the argument.) And git-svnimport doesn't appear to have any sort\nof method for splitting a gigantic svn repo into several smaller git\nrepos.\n\n> > First issue I ran into: On a machine with 4GB of RAM, when I tried to\n> > do a full import, git-svnimport died after 309906 revisions, saying\n> > that it couldn't fork.\n> >\n> > Checking `top` and `ps` revealed that there were no git-svnimport\n> > processes doing anything, but all of my 4G of RAM was still marked as\n> > used by the kernel. I had to do sysctl -w vm.drop_caches=3 to get it\n> > to free all the RAM that the svn import had used up.\n>\n> I think that was just all cached, and all ok. The reason you didn't see\n> any git-svnimport was that it had died off already, and all your memory\n> was just caches. You could just have left it alone, and the kernel would\n> have started re-using the memory for other things even without any\n> \"drop_caches\".\n>\n> But what you did there didn't make anything worse, it was just likely had\n> no real impact.\n\nI got the tip about drop_caches from davej. Normally, when a process\ntaking up a huge amount of memory exits, it shows a bunch of free\nmemory in `top` and friends. I was a little bit surprised when that\ndidn't happen this time.\n\n> However, it does sound like git-svnimport probably acts like git-cvsimport\n> used to, and just keeps too much in memory - so it's never going to act\n> really nicely..\n>\n> It also looks like git-svnimport never repacks the repo, which is\n> absolutely horrible for performance on all levels. The CVS importer\n> repacks every one thousand commits or something like that.\n\nYeah. I haven't bothered hacking git-svnimport yet - but it looks like\nhaving it automatically repack every thousand revisions or so would\nprobably be a pretty big win.\n\n> > Now, after that, I tried doing `git-repack -a` because I wanted to see\n> > how small the packed archive would be (before trying to continue\n> > importing the rest of the revisions. There are at least another 100k\n> > revisions that I should be able to import, eventually.)\n>\n> I suspect you'd have been better off just re-starting, and using something\n> like\n>\n>         while :\n>         do\n>                 git svnimport -l 1000 <...>\n>                 .. figure out some way to decide if it's all done ..\n>                 git repack -d\n>         done\n>\n> which would make svnimport act a bit  more sanely, and repack\n> incrementally. That should make both the import much faster, _and_ avoid\n> any insane big repack at the end (well, you'd still want to do a \"git\n> repack -a -d\" at the end to turn the many smaller packs into a bigger one,\n> but it would be nicer).\n>\n> However, I don't know what the proper magic is for svnimport to do that\n> sane \"do it in chunks and tell when you're all done\". Or even better - to\n> just make it repack properly and not keep everything in memory.\n\nYou can pass limits to svnimport to give it a revision to start at and\nanother one to end at, so that wouldn't be too bad - I was thinking\nabout working around it like that (so that i don't have to go poking\naround in the Perl code behind the svn importer).\n\nBy default, if I had, say, one pack with the first 1000 revisions, and\nI imported another 1000, running 'git-repack' on its own would leave\nthe first pack alone and create a new pack with just the second 1000\nrevisions, right?\n\n> > The repack finished after about nine hours, but when I try to do a\n> > git-verify-pack on it, it dies with this error message:\n> >\n> > error: Packfile\n> > .git/objects/pack/pack-540263fe66ab9398cc796f000d52531a5c6f3df3.pack\n> > SHA1 mismatch with itself\n>\n> That sounds suspiciously like the bug we had in out POWER sha1\n> implementation that would generate the wrong SHA1 for any pack-file that\n> was over 512MB in size, due to an overflow in 32 bits (SHA1 does some\n> counting in _bits_, so 512MB is 4G _bits_),\n>\n> Now, I assume you're not on POWER (and we fixed that bug anyway - and I\n> think long before 1.4.1 too), but I could easily imagine the same bug in\n> some other SHA1 implementation (or perhaps _another_ overflow at the 1GB\n> or 2GB mark..). I assume that the pack-file you had was something horrid..\n>\n> I hope this is with a 64-bit kernel and a 64-bit user space? That should\n> limit _some_ of the issues. But I would still not be surprised if your\n> SHA1 libraries had some 32-bit (\"unsigned int\") or 31-bit (\"int\") limits\n> in them somewhere - very few people do SHA1's over huge areas, and even\n> when you do SHA1 on something like a DVD image (which is easily over any\n> 4GB limit), that tends to be done as many smaller calls to the SHA1\n> library routines.\n\nThis is on a dual-CPU dual-core Opteron, running the AMD64 variant of\nUbuntu's Edgy release (64-bit kernel, 64-bit native userland). The\npack-file was around 2.3GB.\n"},{"id":"30785","messageId":"20070104023350.GA1194@localdomain","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701031737300.4989@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Eric Wong","fromEmail":"normalperson@yhbt.net","sentAt":"2007-01-04T02:33:50Z","receivedAt":"2007-01-04T02:33:50Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"Linus Torvalds <torvalds@osdl.org> wrote:\n> On Wed, 3 Jan 2007, Chris Lee wrote:\n> > First issue I ran into: On a machine with 4GB of RAM, when I tried to\n> > do a full import, git-svnimport died after 309906 revisions, saying\n> > that it couldn't fork.\n\nManaging memory with the Perl SVN libraries has been very painful in my\nexperience.\n\nPart of it is Perl, which (as far as I know) never frees allocated\nmemory back to the OS (although Perl can reuse the allocated memory for\nother things).  I'm CC-ing the resident Perl guru on this...\n\nI'm also fairly certain that most higher-level languages have this\nproblem.\n\n> I suspect you'd have been better off just re-starting, and using something \n> like\n> \n> \twhile :\n> \tdo\n> \t\tgit svnimport -l 1000 <...>\n> \t\t.. figure out some way to decide if it's all done ..\n> \t\tgit repack -d\n> \tdone\n\n> However, I don't know what the proper magic is for svnimport to do that \n> sane \"do it in chunks and tell when you're all done\". Or even better - to \n> just make it repack properly and not keep everything in memory.\n\n<shameless self-promotion>\n\tgit-svn already does this chunking internally\n\n\tJust set the repack interval to something smaller than 1000;\n\t(--repack=100) if you experience timeouts.\n</shameless self-promotion>\n\n-- \nEric Wong\n"},{"id":"30786","messageId":"20070104023510.GC18206@spearce.org","threadId":"6214","inReplyTo":"20070104020652.GB18206@spearce.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-01-04T02:35:11Z","receivedAt":"2007-01-04T02:35:11Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"\"Shawn O. Pearce\" <spearce@spearce.org> wrote:\n> Linus Torvalds <torvalds@osdl.org> wrote:\n> > Junio - I suspect \"pack-check.c\" really shouldn't try to do it as one \n> > single humungous \"SHA1_Update()\" call. It showed one bug on PPC, I \n> > wouldn't be surprised if it's implicated now on some other architecture. \n> \n> It used to do it as one big SHA1_Update() call...\n>  \n> > Shawn - does the pack-file-windowing thing already change that? I'm too \n> > lazy to check..\n> \n> But with the mmap window thing in `next` it does it in window\n> units only.  Which the user could configure to be huge, or could\n> configure to be sane.  The default when using mmap() is 32 MiB;\n> 1 MiB when using pread() and git_mmap().\n\nI should also point out that my git-fastimport hack that we used\non the huge Mozilla import may be helpful here.  Its _very_ fast\nas it goes right to a pack file, but there's no SVN frontend for\nit at this time.\n\n-- \nShawn.\n"},{"id":"30787","messageId":"204011cb0701031836w7d33ca8dh5de08984eec9730d@mail.gmail.com","threadId":"6214","inReplyTo":"20070104023510.GC18206@spearce.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-04T02:36:03Z","receivedAt":"2007-01-04T02:36:03Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"On 1/3/07, Shawn O. Pearce <spearce@spearce.org> wrote:\n> I should also point out that my git-fastimport hack that we used\n> on the huge Mozilla import may be helpful here.  Its _very_ fast\n> as it goes right to a pack file, but there's no SVN frontend for\n> it at this time.\n\nI would be *really* interested in playing with that. Where do I get it?\n"},{"id":"30789","messageId":"86ps9vbjlp.fsf@blue.stonehenge.com","threadId":"6214","inReplyTo":"20070104023350.GA1194@localdomain","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Randal L. Schwartz","fromEmail":"merlyn@stonehenge.com","sentAt":"2007-01-04T02:40:02Z","receivedAt":"2007-01-04T02:40:02Z","isPatch":false,"sender":{"key":"merlyn@stonehenge.com","avatar":"https://gravatar.com/avatar/dc528d210743ff0333e6213f9ee7b33b23f1b7bc1f3c5a8c2d819074ecd7ab19?d=mp&s=160"},"body":">>>>> \"Eric\" == Eric Wong <normalperson@yhbt.net> writes:\n\nEric> Part of it is Perl, which (as far as I know) never frees allocated\nEric> memory back to the OS (although Perl can reuse the allocated memory for\nEric> other things).\n\nIt does on Linux, of all things.  That's because Linux has a smarter\nmalloc/free that uses mmap(2) for the large chunks.  On Linux, Perl memory\nsize can apparently grow and shrink nicely.  The \"old school\" advice about\nPerl comes from sbrk(2)-driven malloc/free.\n\nTry:\n\n        $x[1e6] = \"0\";\n        sleep 10; # do a ps here\n        @x = ();\n        sleep 30; # do a ps here\n\nand watch the process on Linux.  If I'm right, this should show a large\nprocess,  then a smaller one.\n\nIf you're getting a growing process though, you probably have a circular data\nreference.  Maybe you have a tree with backpointers, and those backpointers\nshould have been weakened?\n\n-- \nRandal L. Schwartz - Stonehenge Consulting Services, Inc. - +1 503 777 0095\n<merlyn@stonehenge.com> <URL:http://www.stonehenge.com/merlyn/>\nPerl/Unix/security consulting, Technical writing, Comedy, etc. etc.\nSee PerlTraining.Stonehenge.com for onsite and open-enrollment Perl training!\n"},{"id":"30790","messageId":"20070104024523.GD18206@spearce.org","threadId":"6214","inReplyTo":"204011cb0701031836w7d33ca8dh5de08984eec9730d@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-01-04T02:45:23Z","receivedAt":"2007-01-04T02:45:23Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Chris Lee <chris133@gmail.com> wrote:\n> On 1/3/07, Shawn O. Pearce <spearce@spearce.org> wrote:\n> >I should also point out that my git-fastimport hack that we used\n> >on the huge Mozilla import may be helpful here.  Its _very_ fast\n> >as it goes right to a pack file, but there's no SVN frontend for\n> >it at this time.\n> \n> I would be *really* interested in playing with that. Where do I get it?\n\nIts a fork of git.git on repo.or.cz; the gitweb can be seen here:\n\n  http://repo.or.cz/w/git/fastimport.git\n\nthe clone url is:\n\n  git://repo.or.cz/git/fastimport.git\n  http://repo.or.cz/r/git/fastimport.git\n\nThe entire code is in fast-import.c.  The input stream it consumes\ncomes in on STDIN and is documented in a large comment at the top\nof the file.\n\nAll that's needed is to get data from SVN in a way that it can be\nfed into git-fastimport.\n\n-- \nShawn.\n"},{"id":"30791","messageId":"204011cb0701031853xd226683g85f376c206aacf3e@mail.gmail.com","threadId":"6214","inReplyTo":"20070104024523.GD18206@spearce.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-04T02:53:57Z","receivedAt":"2007-01-04T02:53:57Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"On 1/3/07, Shawn O. Pearce <spearce@spearce.org> wrote:\n> Chris Lee <chris133@gmail.com> wrote:\n> > On 1/3/07, Shawn O. Pearce <spearce@spearce.org> wrote:\n> > >I should also point out that my git-fastimport hack that we used\n> > >on the huge Mozilla import may be helpful here.  Its _very_ fast\n> > >as it goes right to a pack file, but there's no SVN frontend for\n> > >it at this time.\n> >\n> > I would be *really* interested in playing with that. Where do I get it?\n>\n> Its a fork of git.git on repo.or.cz; the gitweb can be seen here:\n>\n>   http://repo.or.cz/w/git/fastimport.git\n>\n> the clone url is:\n>\n>   git://repo.or.cz/git/fastimport.git\n>   http://repo.or.cz/r/git/fastimport.git\n>\n> The entire code is in fast-import.c.  The input stream it consumes\n> comes in on STDIN and is documented in a large comment at the top\n> of the file.\n\nNeat. How do I do that?\n"},{"id":"30792","messageId":"20070104025659.GE18206@spearce.org","threadId":"6214","inReplyTo":"204011cb0701031853xd226683g85f376c206aacf3e@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-01-04T02:57:00Z","receivedAt":"2007-01-04T02:57:00Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"[cc: list modified to remove folks who probably aren't immediately\n interested in git-fastimport]\n\nChris Lee <chris133@gmail.com> wrote:\n> On 1/3/07, Shawn O. Pearce <spearce@spearce.org> wrote:\n> >the clone url is:\n> >\n> >  git://repo.or.cz/git/fastimport.git\n> >  http://repo.or.cz/r/git/fastimport.git\n> >\n> >The entire code is in fast-import.c.  The input stream it consumes\n> >comes in on STDIN and is documented in a large comment at the top\n> >of the file.\n> \n> Neat. How do I do that?\n\nI'm not sure I understand the question...\n\n-- \nShawn.\n"},{"id":"30793","messageId":"204011cb0701031858x231df34as424b7f0c0ae4ab8b@mail.gmail.com","threadId":"6214","inReplyTo":"20070104025659.GE18206@spearce.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-04T02:58:52Z","receivedAt":"2007-01-04T02:58:52Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"Uh... somehow, it lost this part:\n\n> All that's needed is to get data from SVN in a way that it can be\n> fed into git-fastimport.\n\nThat's what I meant - I assume that someone already has the\nsvn-repo-to-gfi piece working? Where's that available from?\n"},{"id":"30794","messageId":"20070104030544.GF18206@spearce.org","threadId":"6214","inReplyTo":"204011cb0701031858x231df34as424b7f0c0ae4ab8b@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-01-04T03:05:44Z","receivedAt":"2007-01-04T03:05:44Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Chris Lee <chris133@gmail.com> wrote:\n> Uh... somehow, it lost this part:\n> \n> >All that's needed is to get data from SVN in a way that it can be\n> >fed into git-fastimport.\n> \n> That's what I meant - I assume that someone already has the\n> svn-repo-to-gfi piece working? Where's that available from?\n\nNo.  That hasn't been written.\n\nIn theory someone could take the SVN dump library (its a chunk of\nC code which parses SVN dump files) and write a tool which translates\nit into git-fastimport.\n\nOne could also use the SVN client library to suck data from SVN\nand pump it into git-fastimport.\n\nJon Smirl attempted to create a CVS-->git-fastimport program in\nPython by starting with the cvs2svn codebase, but that doesn't\ndo anything about importing *from* SVN.  Jon was able to import\nthe entire Mozilla CVS repository (250k commits, about 3 GiB\ninput) in 2 hours using his hacked up cvs2svn and git-fastimport.\nThe resulting pack was ~900 MiB.  He recompressed that using\n`git repack -a -d --window=50 --depth=1000` (which is insane) in\nabout an hour.\n\n-- \nShawn.\n"},{"id":"30795","messageId":"204011cb0701031906i30366286q800fba716f9fa725@mail.gmail.com","threadId":"6214","inReplyTo":"204011cb0701031858x231df34as424b7f0c0ae4ab8b@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"clee@kde.org","sentAt":"2007-01-04T03:06:24Z","receivedAt":"2007-01-04T03:06:24Z","isPatch":false,"sender":{"key":"clee@kde.org","avatar":"https://gravatar.com/avatar/c930bdc8cc6465094a5722188409ecb8955e0da2b188d7340137074b08f857e3?d=mp&s=160"},"body":"On 1/3/07, Chris Lee <chris133@gmail.com> wrote:\n> Uh... somehow, it lost this part:\n>\n> > All that's needed is to get data from SVN in a way that it can be\n> > fed into git-fastimport.\n>\n> That's what I meant - I assume that someone already has the\n> svn-repo-to-gfi piece working? Where's that available from?\n\nRight, and I'm an idiot! Awesome.\n\nI obviously didn't comprehend the part where you wrote:\n\n> I should also point out that my git-fastimport hack that we used\n> on the huge Mozilla import may be helpful here.  Its _very_ fast\n> as it goes right to a pack file, but there's no SVN frontend for\n> it at this time.\n\nAnyway. Thanks for the pointers, I'll see if I can't hack something up.\n"},{"id":"30796","messageId":"20070104031340.GA15094@localdomain","threadId":"6214","inReplyTo":"86ps9vbjlp.fsf@blue.stonehenge.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Eric Wong","fromEmail":"normalperson@yhbt.net","sentAt":"2007-01-04T03:13:40Z","receivedAt":"2007-01-04T03:13:40Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"\"Randal L. Schwartz\" <merlyn@stonehenge.com> wrote:\n> >>>>> \"Eric\" == Eric Wong <normalperson@yhbt.net> writes:\n> \n> Eric> Part of it is Perl, which (as far as I know) never frees allocated\n> Eric> memory back to the OS (although Perl can reuse the allocated memory for\n> Eric> other things).\n> \n> It does on Linux, of all things.  That's because Linux has a smarter\n> malloc/free that uses mmap(2) for the large chunks.  On Linux, Perl memory\n> size can apparently grow and shrink nicely.  The \"old school\" advice about\n> Perl comes from sbrk(2)-driven malloc/free.\n> \n> Try:\n> \n>         $x[1e6] = \"0\";\n>         sleep 10; # do a ps here\n>         @x = ();\n>         sleep 30; # do a ps here\n> \n> and watch the process on Linux.  If I'm right, this should show a large\n> process,  then a smaller one.\n\nNope, not happening to me.  I'm using Perl 5.8.8-7 and glibc 2.3.6.ds1-8\non a Debian Etch machine.  The kernel is a vanilla 2.6.18.1 from\nkernel.org.\n\nstrace shows an mmap2 call, but no corresponding mumap.  I've added a\nsleep loop to the end of the above program and had it print\nsomething every 10 seconds; but so far, there's still no munmap.\n\nwhile (1) {\n        print \"hi\\n\" if ((time % 10) == 0);\n\tsleep 1;\n}\n\nTrying to allocate a bigger chunk (1e7) doesn't show anything different,\neither.  I've also conducted similar experiments with Ruby in the past\nand noticed the same things...\n\n-- \nEric Wong\n"},{"id":"30799","messageId":"7v1wmbnw9x.fsf@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701031737300.4989@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-04T06:25:30Z","receivedAt":"2007-01-04T06:25:30Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Wed, 3 Jan 2007, Chris Lee wrote:\n>>\n>> So I'm using git 1.4.1, and I have been experimenting with importing\n>> the KDE sources from Subversion using git-svnimport.\n>\n> As one single _huge_ import? All the sub-projects together? I have to say, \n> that sounds pretty horrid.\n\nThanks -- you said everything I should have said on this issue\nwhile I was in bed ;-).\n\n> Junio - I suspect \"pack-check.c\" really shouldn't try to do it as one \n> single humungous \"SHA1_Update()\" call. It showed one bug on PPC, I \n> wouldn't be surprised if it's implicated now on some other architecture. \n\nIf Chris still has that huge .pack & .idx pair, it would be a\nvery good guinea pig to try a few things on, assuming that this\nproblem is that the pack-check.c feeds a huge blob to SHA-1\nfunction with a single call.\n\n (1) Apply the attached patch on top of \"master\" (the patch\n     should apply to 1.4.1 almost cleanly as well, except that\n     we have hashcmp(a,b) instead of memcmp(a,b,20) since then),\n     and see what it says about the packfile.  If your suspicion\n     is correct, it should complain about your SHA-1\n     implementation.\n\n (2) Try tip of \"next\" to see if its verify-pack passes the\n     check.  Again, if your suspicion is correct, it should, since it\n     uses Shawn's sliding mmap() stuff that will not feed the\n     whole pack in one go.\n\n (3) I suspect that the tip of \"master\" should work except\n     verify-pack.  It may be interesting to see how well the tip\n     of \"master\" and \"next\" performs on the resulting huge pack\n     (say, \"time git log -p HEAD >/dev/null\").  I am hoping this\n     would be another datapoint to judge the runtime penalty of\n     Shawn's sliding mmap() in \"next\" -- I suspect the penalty\n     is either negligible or even negative.\n\ndiff --git a/pack-check.c b/pack-check.c\nindex c0caaee..738a0c5 100644\n--- a/pack-check.c\n+++ b/pack-check.c\n@@ -29,6 +29,28 @@ static int verify_packfile(struct packed_git *p)\n \tpack_base = p->pack_base;\n \tSHA1_Update(&ctx, pack_base, pack_size - 20);\n \tSHA1_Final(sha1, &ctx);\n+\n+\tif (1) {\n+\t\tSHA_CTX another;\n+\t\tunsigned char *data = p->pack_base;\n+\t\tunsigned long size = pack_size - 20;\n+\t\tconst unsigned long batchsize = (1u << 20);\n+\t\tunsigned char another_sha1[20];\n+\n+\t\tSHA1_Init(&another);\n+\t\twhile (size) {\n+\t\t\tunsigned long batch = size;\n+\t\t\tif (batchsize < batch)\n+\t\t\t\tbatch = batchsize;\n+\t\t\tSHA1_Update(&another, data, batch);\n+\t\t\tsize -= batch;\n+\t\t\tdata += batch;\n+\t\t}\n+\t\tSHA1_Final(another_sha1, &another);\n+\t\tif (hashcmp(sha1, another_sha1))\n+\t\t\tdie(\"Your SHA-1 implementation cannot hash %lu bytes correctly at once\", pack_size - 20);\n+\t}\n+\n \tif (hashcmp(sha1, (unsigned char *)pack_base + pack_size - 20))\n \t\treturn error(\"Packfile %s SHA1 mismatch with itself\",\n \t\t\t     p->pack_name);\n"},{"id":"30801","messageId":"7vr6ubmewg.fsf_-_@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"7v1wmbnw9x.fsf@assigned-by-dhcp.cox.net","subject":"[PATCH] pack-check.c::verify_packfile(): don't run SHA-1 update on huge data","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-04T07:26:07Z","receivedAt":"2007-01-04T07:26:07Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Running the SHA1_Update() on the whole packfile in a single call\nrevealed an overflow problem we had in the SHA-1 implementation\non POWER architecture some time ago, which was fixed with commit\nb47f509b (June 19, 2006).  Other SHA-1 implementations may have\na similar problem.\n\nThe sliding mmap() series already makes chunked calls to\nSHA1_Update(), so this patch itself will become moot when it\ngraduates to \"master\", but in the meantime, run the hash\nfunction in smaller chunks to prevent possible future problems.\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n * Chris, if you have a chance could you try this on the huge\n   pack you had trouble with?\n\n   Also whose SHA-1 implementation are you using, if this indeed\n   is the problem, I wonder?\n\n pack-check.c |   20 +++++++++++++++-----\n 1 files changed, 15 insertions(+), 5 deletions(-)\n\ndiff --git a/pack-check.c b/pack-check.c\nindex c0caaee..8e123b7 100644\n--- a/pack-check.c\n+++ b/pack-check.c\n@@ -1,16 +1,18 @@\n #include \"cache.h\"\n #include \"pack.h\"\n \n+#define BATCH (1u<<20)\n+\n static int verify_packfile(struct packed_git *p)\n {\n \tunsigned long index_size = p->index_size;\n \tvoid *index_base = p->index_base;\n \tSHA_CTX ctx;\n \tunsigned char sha1[20];\n-\tunsigned long pack_size = p->pack_size;\n-\tvoid *pack_base;\n \tstruct pack_header *hdr;\n \tint nr_objects, err, i;\n+\tunsigned char *packdata;\n+\tunsigned long datasize;\n \n \t/* Header consistency check */\n \thdr = p->pack_base;\n@@ -25,11 +27,19 @@ static int verify_packfile(struct packed_git *p)\n \t\t\t     \"while idx size expects %d\", nr_objects,\n \t\t\t     num_packed_objects(p));\n \n+\t/* Check integrity of pack data with its SHA-1 checksum */\n \tSHA1_Init(&ctx);\n-\tpack_base = p->pack_base;\n-\tSHA1_Update(&ctx, pack_base, pack_size - 20);\n+\tpackdata = p->pack_base;\n+\tdatasize = p->pack_size - 20;\n+\twhile (datasize) {\n+\t\tunsigned long batch = (datasize < BATCH) ? datasize : BATCH;\n+\t\tSHA1_Update(&ctx, packdata, batch);\n+\t\tdatasize -= batch;\n+\t\tpackdata += batch;\n+\t}\n \tSHA1_Final(sha1, &ctx);\n-\tif (hashcmp(sha1, (unsigned char *)pack_base + pack_size - 20))\n+\n+\tif (hashcmp(sha1, (unsigned char *)(p->pack_base) + p->pack_size - 20))\n \t\treturn error(\"Packfile %s SHA1 mismatch with itself\",\n \t\t\t     p->pack_name);\n \tif (hashcmp(sha1, (unsigned char *)index_base + index_size - 40))\n"},{"id":"30823","messageId":"204011cb0701040956p11ea2cepe3efaaf396056ac0@mail.gmail.com","threadId":"6214","inReplyTo":"204011cb0701031816hda8af9bw4d4a469c2b111339@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"clee@kde.org","sentAt":"2007-01-04T17:56:52Z","receivedAt":"2007-01-04T17:56:52Z","isPatch":false,"sender":{"key":"clee@kde.org","avatar":"https://gravatar.com/avatar/c930bdc8cc6465094a5722188409ecb8955e0da2b188d7340137074b08f857e3?d=mp&s=160"},"body":"Accidentally sent this to just Linus instead of the list...\n\nOn 1/3/07, Linus Torvalds <torvalds@osdl.org> wrote:\n> > So I'm using git 1.4.1, and I have been experimenting with importing\n> > the KDE sources from Subversion using git-svnimport.\n>\n> As one single _huge_ import? All the sub-projects together? I have to say,\n> that sounds pretty horrid.\n\nUnfortunately, that's how the KDE repo is organized. (I tried arguing\nagainst this when they were going to do the original import, but I\nlost the argument.) And git-svnimport doesn't appear to have any sort\nof method for splitting a gigantic svn repo into several smaller git\nrepos.\n\n> > First issue I ran into: On a machine with 4GB of RAM, when I tried to\n> > do a full import, git-svnimport died after 309906 revisions, saying\n> > that it couldn't fork.\n> >\n> > Checking `top` and `ps` revealed that there were no git-svnimport\n> > processes doing anything, but all of my 4G of RAM was still marked as\n> > used by the kernel. I had to do sysctl -w vm.drop_caches=3 to get it\n> > to free all the RAM that the svn import had used up.\n>\n> I think that was just all cached, and all ok. The reason you didn't see\n> any git-svnimport was that it had died off already, and all your memory\n> was just caches. You could just have left it alone, and the kernel would\n> have started re-using the memory for other things even without any\n> \"drop_caches\".\n>\n> But what you did there didn't make anything worse, it was just likely had\n> no real impact.\n\nI got the tip about drop_caches from davej. Normally, when a process\ntaking up a huge amount of memory exits, it shows a bunch of free\nmemory in `top` and friends. I was a little bit surprised when that\ndidn't happen this time.\n\n> However, it does sound like git-svnimport probably acts like git-cvsimport\n> used to, and just keeps too much in memory - so it's never going to act\n> really nicely..\n>\n> It also looks like git-svnimport never repacks the repo, which is\n> absolutely horrible for performance on all levels. The CVS importer\n> repacks every one thousand commits or something like that.\n\nYeah. I haven't bothered hacking git-svnimport yet - but it looks like\nhaving it automatically repack every thousand revisions or so would\nprobably be a pretty big win.\n\n> > Now, after that, I tried doing `git-repack -a` because I wanted to see\n> > how small the packed archive would be (before trying to continue\n> > importing the rest of the revisions. There are at least another 100k\n> > revisions that I should be able to import, eventually.)\n>\n> I suspect you'd have been better off just re-starting, and using something\n> like\n>\n>         while :\n>         do\n>                 git svnimport -l 1000 <...>\n>                 .. figure out some way to decide if it's all done ..\n>                 git repack -d\n>         done\n>\n> which would make svnimport act a bit  more sanely, and repack\n> incrementally. That should make both the import much faster, _and_ avoid\n> any insane big repack at the end (well, you'd still want to do a \"git\n> repack -a -d\" at the end to turn the many smaller packs into a bigger one,\n> but it would be nicer).\n>\n> However, I don't know what the proper magic is for svnimport to do that\n> sane \"do it in chunks and tell when you're all done\". Or even better - to\n> just make it repack properly and not keep everything in memory.\n\nYou can pass limits to svnimport to give it a revision to start at and\nanother one to end at, so that wouldn't be too bad - I was thinking\nabout working around it like that (so that i don't have to go poking\naround in the Perl code behind the svn importer).\n\nBy default, if I had, say, one pack with the first 1000 revisions, and\nI imported another 1000, running 'git-repack' on its own would leave\nthe first pack alone and create a new pack with just the second 1000\nrevisions, right?\n\n> > The repack finished after about nine hours, but when I try to do a\n> > git-verify-pack on it, it dies with this error message:\n> >\n> > error: Packfile\n> > .git/objects/pack/pack-540263fe66ab9398cc796f000d52531a5c6f3df3.pack\n> > SHA1 mismatch with itself\n>\n> That sounds suspiciously like the bug we had in out POWER sha1\n> implementation that would generate the wrong SHA1 for any pack-file that\n> was over 512MB in size, due to an overflow in 32 bits (SHA1 does some\n> counting in _bits_, so 512MB is 4G _bits_),\n>\n> Now, I assume you're not on POWER (and we fixed that bug anyway - and I\n> think long before 1.4.1 too), but I could easily imagine the same bug in\n> some other SHA1 implementation (or perhaps _another_ overflow at the 1GB\n> or 2GB mark..). I assume that the pack-file you had was something horrid..\n>\n> I hope this is with a 64-bit kernel and a 64-bit user space? That should\n> limit _some_ of the issues. But I would still not be surprised if your\n> SHA1 libraries had some 32-bit (\"unsigned int\") or 31-bit (\"int\") limits\n> in them somewhere - very few people do SHA1's over huge areas, and even\n> when you do SHA1 on something like a DVD image (which is easily over any\n> 4GB limit), that tends to be done as many smaller calls to the SHA1\n> library routines.\n\nThis is on a dual-CPU dual-core Opteron, running the AMD64 variant of\nUbuntu's Edgy release (64-bit kernel, 64-bit native userland). The\npack-file was around 2.3GB.\n"},{"id":"30824","messageId":"204011cb0701040958k884b613i8a4639201ae6443b@mail.gmail.com","threadId":"6214","inReplyTo":"7v1wmbnw9x.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-04T17:58:23Z","receivedAt":"2007-01-04T17:58:23Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"> If Chris still has that huge .pack & .idx pair, it would be a\n> very good guinea pig to try a few things on, assuming that this\n> problem is that the pack-check.c feeds a huge blob to SHA-1\n> function with a single call.\n\nI do not still have it, but I can pretty easily regenerate it. Should\nhave it again in another nine hours or so. :)\n\n>  (1) Apply the attached patch on top of \"master\" (the patch\n>      should apply to 1.4.1 almost cleanly as well, except that\n>      we have hashcmp(a,b) instead of memcmp(a,b,20) since then),\n>      and see what it says about the packfile.  If your suspicion\n>      is correct, it should complain about your SHA-1\n>      implementation.\n>\n>  (2) Try tip of \"next\" to see if its verify-pack passes the\n>      check.  Again, if your suspicion is correct, it should, since it\n>      uses Shawn's sliding mmap() stuff that will not feed the\n>      whole pack in one go.\n>\n>  (3) I suspect that the tip of \"master\" should work except\n>      verify-pack.  It may be interesting to see how well the tip\n>      of \"master\" and \"next\" performs on the resulting huge pack\n>      (say, \"time git log -p HEAD >/dev/null\").  I am hoping this\n>      would be another datapoint to judge the runtime penalty of\n>      Shawn's sliding mmap() in \"next\" -- I suspect the penalty\n>      is either negligible or even negative.\n\nI'll try all of this after the pack is regenerated. Thanks!\n"},{"id":"30826","messageId":"Pine.LNX.4.64.0701041016010.3661@woody.osdl.org","threadId":"6214","inReplyTo":"204011cb0701040956p11ea2cepe3efaaf396056ac0@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-04T18:30:43Z","receivedAt":"2007-01-04T18:30:43Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 4 Jan 2007, Chris Lee wrote:\n>\n> Unfortunately, that's how the KDE repo is organized. (I tried arguing\n> against this when they were going to do the original import, but I\n> lost the argument.) And git-svnimport doesn't appear to have any sort\n> of method for splitting a gigantic svn repo into several smaller git\n> repos.\n\nWell, the good news is, I think we could probably split it up from within \ngit. It's not fundamentally hard, although it is pretty damn expensive \n(and it would require the subproject support to do really well).\n\nSo ignore that issue for now. I'd love to see the end result, if only \nbecause it sounds like you have a test-case for git that is four times \nbigger than the mozilla archive - even if it's just because of some really \nreally stupid design decisions from the KDE SVN maintainers ;)\n\n(But I would actually expect that KDE SVN uses SVN subprojects, so \nhopefully it's not _really_ one big repository. Of course, I don't know if \nSVN really does subprojects or how well it does them, so that's just a \ntotal guess).\n\nThe real problem with a SVN import is that I think SVN doesn't do merges \nright, so you can't import merge history properly (well, you can, if you \ndecide that \"properly\" really means \"SVN can't merge, so we can't really \nshow it as merges in git either\").\n\nI think both git-svn and git-svnimport can _guess_ about merges, but it's \njust a heuristic, afaik. Whether it's a good one, I don't know.\n\n> Yeah. I haven't bothered hacking git-svnimport yet - but it looks like\n> having it automatically repack every thousand revisions or so would\n> probably be a pretty big win.\n\nThat, or making it use the same \"fastimport\" that the hacked-up CVS \nimporter was made to use. Either way, somebody who understands SVN \nintimately (and probably perl) would need to work on it. \n\nThat would not be me, so I can't really help ;)\n\n> By default, if I had, say, one pack with the first 1000 revisions, and\n> I imported another 1000, running 'git-repack' on its own would leave\n> the first pack alone and create a new pack with just the second 1000\n> revisions, right?\n\nYes. It's _probably_ better to do a full re-pack every once in a while \n(because if you have a lot of pack-files, eventually that ends up being \nproblematic too), but as a first approximation, it's probably fine to just \ndo a plain \"git repack\" every thousand commits, and then do a full big \nrepack at the end.\n\nThe big repack will still be pretty expensive, but it should be less \npainful than having everything unpacked. And at least the import won't \nhave run with millions and millions of loose objects.\n\nSo doing a \"git repack -a -d\" at the end is a good idea, and _maybe_ it \ncould be done in the middle too for really big packs.\n\nAgain, doing what fastimport does avoids most of the whole issue, since it \njust generates a pack up-front instead. But that requires the importer to \nspecifically understand about that kind of setup.\n\n> This is on a dual-CPU dual-core Opteron, running the AMD64 variant of\n> Ubuntu's Edgy release (64-bit kernel, 64-bit native userland). The\n> pack-file was around 2.3GB.\n\nOk, that should all be fine. A 31-bit thing in OpenSSL would explain it, \nand doesn't sound unlikely. Just somebody using \"int\" somewhere, and it \nwould never have been triggered by any sane user of SHA1_Update(). The git \npack-check.c usage really _is_ very odd, even if it happens to make sense \nin that particular schenario.\n\n\t\tLinus\n"},{"id":"30827","messageId":"204011cb0701041054h76f1f178j3fd7994a01299be8@mail.gmail.com","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701041016010.3661@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"clee@kde.org","sentAt":"2007-01-04T18:54:46Z","receivedAt":"2007-01-04T18:54:46Z","isPatch":false,"sender":{"key":"clee@kde.org","avatar":"https://gravatar.com/avatar/c930bdc8cc6465094a5722188409ecb8955e0da2b188d7340137074b08f857e3?d=mp&s=160"},"body":"On 1/4/07, Linus Torvalds <torvalds@osdl.org> wrote:\n> Well, the good news is, I think we could probably split it up from within\n> git. It's not fundamentally hard, although it is pretty damn expensive\n> (and it would require the subproject support to do really well).\n\nI was hoping that'd be possible at some point. I really want to split\nthe submodules back out into first-class modules - one of my biggest\nmisgivings about the current KDE repository setup is how everything is\npart of one gigantic repository.\n\n> So ignore that issue for now. I'd love to see the end result, if only\n> because it sounds like you have a test-case for git that is four times\n> bigger than the mozilla archive - even if it's just because of some really\n> really stupid design decisions from the KDE SVN maintainers ;)\n\nThe full on-disk size of the KDE SVN repo is about 37GB, last time I\nchecked. It may be up to 38 or 39GB now - I last ran rsync against the\nsvn repo a few weeks ago. I'm only focusing on importing the first\n409k revisions at the moment, because that comprises the commits that\noriginally came from CVS and were imported into SVN. Almost\nimmediately after the CVS import, coolo made some changes - moving all\nof the core KDE modules into /trunk/KDE, and their branches and tags\ninto /branches/KDE and /tags/KDE respectively. This, I suspect will\nend up making things \"fun\" for the other part of the import, which is\nanother 200k revisions, give or take.\n\nSo, yes, I suspect it's quite a bit larger than Mozilla. I'm doing the\nconversion to git as a test so that I can show some numbers to the KDE\nguys; I'm not trying to campaign for a transition to git, but I think\nit's definitely worth exploring what such a world would look like. But\nin order for me to try to make a compelling argument for an eventual\nproject move to git, the git win32 support would need to be really\ngood. (In KDE4, we're supporting Windows and OS X as well as X11 as\nfirst-class platforms.)\n\n> (But I would actually expect that KDE SVN uses SVN subprojects, so\n> hopefully it's not _really_ one big repository. Of course, I don't know if\n> SVN really does subprojects or how well it does them, so that's just a\n> total guess).\n\nI don't think so, but I'll ask coolo (the KDE SVN administrator).\n\n> The real problem with a SVN import is that I think SVN doesn't do merges\n> right, so you can't import merge history properly (well, you can, if you\n> decide that \"properly\" really means \"SVN can't merge, so we can't really\n> show it as merges in git either\").\n>\n> I think both git-svn and git-svnimport can _guess_ about merges, but it's\n> just a heuristic, afaik. Whether it's a good one, I don't know.\n\nNot too worried about the merges right now - as long as I have a rough\napproximation of what the original looked like, I'm pretty happy.\n\n> > Yeah. I haven't bothered hacking git-svnimport yet - but it looks like\n> > having it automatically repack every thousand revisions or so would\n> > probably be a pretty big win.\n>\n> That, or making it use the same \"fastimport\" that the hacked-up CVS\n> importer was made to use. Either way, somebody who understands SVN\n> intimately (and probably perl) would need to work on it.\n>\n> That would not be me, so I can't really help ;)\n\nWell, Shawn pointed me at the fastimport stuff, and I happen to know\nPerl reasonably well (I think) so I'll take a stab at trying it that\nway.\n\n> > By default, if I had, say, one pack with the first 1000 revisions, and\n> > I imported another 1000, running 'git-repack' on its own would leave\n> > the first pack alone and create a new pack with just the second 1000\n> > revisions, right?\n>\n> Yes. It's _probably_ better to do a full re-pack every once in a while\n> (because if you have a lot of pack-files, eventually that ends up being\n> problematic too), but as a first approximation, it's probably fine to just\n> do a plain \"git repack\" every thousand commits, and then do a full big\n> repack at the end.\n\nSounds like a good idea. Also sounds like it would be much less\npainful than the current situation, where it takes over nine hours to\npack up all these revisions. :)\n\n> The big repack will still be pretty expensive, but it should be less\n> painful than having everything unpacked. And at least the import won't\n> have run with millions and millions of loose objects.\n>\n> So doing a \"git repack -a -d\" at the end is a good idea, and _maybe_ it\n> could be done in the middle too for really big packs.\n\nOkay, good to know.\n\n> Again, doing what fastimport does avoids most of the whole issue, since it\n> just generates a pack up-front instead. But that requires the importer to\n> specifically understand about that kind of setup.\n\nI'll definitely be investigating the fastimport option. Looks like\nI'll get to crack open some of my Perl books - haven't had to do that\nin a while. :)\n"},{"id":"30830","messageId":"204011cb0701041124g40440fd4udf1088ab1341c031@mail.gmail.com","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701031737300.4989@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-04T19:24:20Z","receivedAt":"2007-01-04T19:24:20Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"On 1/3/07, Linus Torvalds <torvalds@osdl.org> wrote:\n> > Checking `top` and `ps` revealed that there were no git-svnimport\n> > processes doing anything, but all of my 4G of RAM was still marked as\n> > used by the kernel. I had to do sysctl -w vm.drop_caches=3 to get it\n> > to free all the RAM that the svn import had used up.\n>\n> I think that was just all cached, and all ok. The reason you didn't see\n> any git-svnimport was that it had died off already, and all your memory\n> was just caches. You could just have left it alone, and the kernel would\n> have started re-using the memory for other things even without any\n> \"drop_caches\".\n>\n> But what you did there didn't make anything worse, it was just likely had\n> no real impact.\n\nThought it was worth mentioning this:\n\nWhen I checked top, the numbers it showed me were:\nMem:   4059332k total,  3216480k used,   842852k free,    40824k buffers\nSwap:        0k total,        0k used,        0k free,    37364k cached\n\n40MB in buffers, 37MB in cache, and 3GB used.\n\nSeems like *something* was definitely lost there. The 'used' number\ndidn't go down at all when I started doing other things; it went up as\nthe new programs started, then they used up some RAM, and then when\nthey exited they'd free whatever resources they'd used. However, until\nI did the drop_caches, that number stayed pretty damn big.\n\nThe system has been up since then, doing lots of things, and still\nseems pretty stable, so I think it's okay, but I thought that it was\nworth mentioning that something seemed to be leaky.\n"},{"id":"30834","messageId":"7v1wmalez6.fsf@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"204011cb0701040958k884b613i8a4639201ae6443b@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-04T20:22:05Z","receivedAt":"2007-01-04T20:22:05Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Chris Lee\" <chris133@gmail.com> writes:\n\n> I'll try all of this after the pack is regenerated. Thanks!\n\nThank YOU for helping to make git better.\n"},{"id":"30838","messageId":"Pine.LNX.4.64.0701041300410.3661@woody.osdl.org","threadId":"6214","inReplyTo":"204011cb0701041124g40440fd4udf1088ab1341c031@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-04T21:12:18Z","receivedAt":"2007-01-04T21:12:18Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Thu, 4 Jan 2007, Chris Lee wrote:\n> \n> Seems like *something* was definitely lost there. The 'used' number\n> didn't go down at all when I started doing other things; it went up as\n> the new programs started\n\nThe 'used' number basically _never_ goes down as long as there is memory \nfree. The kernel simply doesn't have any reason to free any of its caches, \neven if those caches end up not being very useful.\n\nWhat happened is almost certainly that with your big unpacked repository, \nthe kernel ended up using a lot of memory on filename caching. In other \nwords, I'd have expected that if you were to do \n\n\tcat /proc/slabinfo\n\nyou'd have seen a _lot_ of memory being used for dentries (\"dentry_cache\") \nand inodes (\"ext3_inode_cache\" assuming you're an ext3 user).\n\nThe kernel can easily drop those caches on demand, but \"free\" isn't quite \nsmart enough to know about them as being caches, so they will just show up \nas \"used\".\n\nThat said, since you didn't want them, dropping them by hand with sysctl \ncertainly didn't hurt. Manual control can often be better than automatic \nheuristics..\n\nSo the reason why repacking is so useful is that it gets rid of all these \nmillions of individual files. They all take up space on the disk, but they \nalso do end up having a lot of caches associated with them.\n\nBtw, you may find that despite your 4GB of RAM, you might still be \nbetter off with a swapfile. It gives the kernel a certain amount of \nfreedom in choosing how to allocate memory, and perhaps more importantly, \neven when the kernel doesn't actively use it, it means that IF the kernel \nruns out of totally free memory (because it has decided to keep a lot of \nstuff in the dentry cache), it gives the kernel choices, and a certain \n\"buffer\" for making the right decision.\n\nWhat often happens is that the memory management heuristics don't make the \n\"perfect\" choice (partly because it's theoretically impossible anyway, but \nlargely just because it's just a damn hard problem to even get all that \n*close* to perfect), and having a swap partition or even a swap file just \nallows the kernel to make some mistakes without it hitting a hard wall of \n\"oh, I can't do anything at all about this particular page\".\n\nSo that buffer zone can be helpful in avoiding bad situations, but it can \nactually also end up improving performance - it doesn't sound like the \ncase in this particular situation, but in some other loads there really \nare a lot of dirty pages that aren't all that useful and where the memory \nreally could be better used for other things if the largely unused dirty \npage could just be written to disk.\n\n\t\t\tLinus\n"},{"id":"30839","messageId":"20070104213142.GE11861@sashak.voltaire.com","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701031737300.4989@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Sasha Khapyorsky","fromEmail":"sashak@voltaire.com","sentAt":"2007-01-04T21:31:42Z","receivedAt":"2007-01-04T21:31:42Z","isPatch":false,"sender":{"key":"sashak@voltaire.com","avatar":null},"body":"On 17:59 Wed 03 Jan     , Linus Torvalds wrote:\n> \n> However, I don't know what the proper magic is for svnimport to do that \n> sane \"do it in chunks and tell when you're all done\". Or even better - to \n> just make it repack properly and not keep everything in memory.\n\n> As to who knows how to fix git-svnimport to do something saner, I have no \n> clue.. Sasha seems to have touched it last. Sasha?\n\nI guess it should not be hard to do svnimport in incrementally with\nrepacking. Like this:\n\n\ndiff --git a/git-svnimport.perl b/git-svnimport.perl\nindex 071777b..afbbe63 100755\n--- a/git-svnimport.perl\n+++ b/git-svnimport.perl\n@@ -31,12 +31,13 @@ $SIG{'PIPE'}=\"IGNORE\";\n $ENV{'TZ'}=\"UTC\";\n \n our($opt_h,$opt_o,$opt_v,$opt_u,$opt_C,$opt_i,$opt_m,$opt_M,$opt_t,$opt_T,\n-    $opt_b,$opt_r,$opt_I,$opt_A,$opt_s,$opt_l,$opt_d,$opt_D,$opt_S,$opt_F,$opt_P);\n+    $opt_b,$opt_r,$opt_I,$opt_A,$opt_s,$opt_l,$opt_d,$opt_D,$opt_S,$opt_F,\n+    $opt_P,$opt_R);\n \n sub usage() {\n \tprint STDERR <<END;\n Usage: ${\\basename $0}     # fetch/update GIT from SVN\n-       [-o branch-for-HEAD] [-h] [-v] [-l max_rev]\n+       [-o branch-for-HEAD] [-h] [-v] [-l max_rev] [-R repack_each_revs]\n        [-C GIT_repository] [-t tagname] [-T trunkname] [-b branchname]\n        [-d|-D] [-i] [-u] [-r] [-I ignorefilename] [-s start_chg]\n        [-m] [-M regex] [-A author_file] [-S] [-F] [-P project_name] [SVN_URL]\n@@ -44,7 +45,7 @@ END\n \texit(1);\n }\n \n-getopts(\"A:b:C:dDFhiI:l:mM:o:rs:t:T:SP:uv\") or usage();\n+getopts(\"A:b:C:dDFhiI:l:mM:o:rs:t:T:SP:R:uv\") or usage();\n usage if $opt_h;\n \n my $tag_name = $opt_t || \"tags\";\n@@ -52,6 +53,7 @@ my $trunk_name = $opt_T || \"trunk\";\n my $branch_name = $opt_b || \"branches\";\n my $project_name = $opt_P || \"\";\n $project_name = \"/\" . $project_name if ($project_name);\n+my $repack_after = $opt_R || 1000;\n \n @ARGV == 1 or @ARGV == 2 or usage();\n \n@@ -938,11 +940,27 @@ if ($opt_l < $current_rev) {\n     exit;\n }\n \n-print \"Fetching from $current_rev to $opt_l ...\\n\" if $opt_v;\n+print \"Processing from $current_rev to $opt_l ...\\n\" if $opt_v;\n \n-my $pool=SVN::Pool->new;\n-$svn->{'svn'}->get_log(\"/\",$current_rev,$opt_l,0,1,1,\\&commit_all,$pool);\n-$pool->clear;\n+my $from_rev;\n+my $to_rev = $current_rev;\n+\n+while ($to_rev < $opt_l) {\n+\t$from_rev = $to_rev;\n+\t$to_rev = $from_rev + $repack_after;\n+\t$to_rev = $opt_l if $opt_l < $to_rev;\n+\tprint \"Fetching from $from_rev to $to_rev ...\\n\" if $opt_v;\n+\tmy $pool=SVN::Pool->new;\n+\t$svn->{'svn'}->get_log(\"/\",$from_rev,$to_rev,0,1,1,\\&commit_all,$pool);\n+\t$pool->clear;\n+\tmy $pid = fork();\n+\tdie \"Fork: $!\\n\" unless defined $pid;\n+\tunless($pid) {\n+\t\texec(\"git-repack\", \"-d\")\n+\t\t\tor die \"Cannot repack: $!\\n\";\n+\t}\n+\twaitpid($pid, 0);\n+}\n \n \n unlink($git_index);\n\n\nChris, it works fine for me with small repository (~9000 revisions), but\nI don't have such huge one as yours. Could you try? Thanks.\n\nSasha\n"},{"id":"30840","messageId":"204011cb0701041404g684525fdm1d057e57a57aca92@mail.gmail.com","threadId":"6214","inReplyTo":"20070104213142.GE11861@sashak.voltaire.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-04T22:04:37Z","receivedAt":"2007-01-04T22:04:37Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"> Chris, it works fine for me with small repository (~9000 revisions), but\n> I don't have such huge one as yours. Could you try? Thanks.\n\nPatch looks like it makes sense. I can definitely try it later.\n\nBack to work for now...\n"},{"id":"30851","messageId":"20070105020955.GA27984@localdomain","threadId":"6214","inReplyTo":"20070104023350.GA1194@localdomain","subject":"[PATCH] git-svn: make --repack work consistently between fetch and multi-fetch","fromName":"Eric Wong","fromEmail":"normalperson@yhbt.net","sentAt":"2007-01-05T02:09:56Z","receivedAt":"2007-01-05T02:09:56Z","isPatch":true,"sender":{"key":"e@80x24.org","avatar":null},"body":"Since fetch reforks itself at most every 1000 revisions, we\nneed to update the counter in the parent process to have a\nworking count if we set our repack interval to be > ~1000\nrevisions.  multi-fetch has always done this correctly\nbecause of an extra process; now fetch uses the extra process;\nas well.\n\nWhile we're at it, only compile the $sha1 regex that checks for\nrepacking once.\n\nSigned-off-by: Eric Wong <normalperson@yhbt.net>\n---\n\nI wrote:\n> \tJust set the repack interval to something smaller than 1000;\n> \t(--repack=100) if you experience timeouts.\n\nChris: you shouldn't get timeouts (at least not across HTTP(s)).\nAlso, don't worry about repack=100 either; there was a bug that\nwas triggered only in 'fetch' not 'multi-fetch' (you should use\n'multi-fetch').  This patch fixes the 'fetch' bug.\n\n git-svn.perl |   10 ++++++----\n 1 files changed, 6 insertions(+), 4 deletions(-)\n\ndiff --git a/git-svn.perl b/git-svn.perl\nindex 0fc386a..5377762 100755\n--- a/git-svn.perl\n+++ b/git-svn.perl\n@@ -102,7 +102,7 @@ my %cmt_opts = ( 'edit|e' => \\$_edit,\n );\n \n my %cmd = (\n-\tfetch => [ \\&fetch, \"Download new revisions from SVN\",\n+\tfetch => [ \\&cmd_fetch, \"Download new revisions from SVN\",\n \t\t\t{ 'revision|r=s' => \\$_revision, %fc_opts } ],\n \tinit => [ \\&init, \"Initialize a repo for tracking\" .\n \t\t\t  \" (requires URL argument)\",\n@@ -293,6 +293,10 @@ sub init {\n \tsetup_git_svn();\n }\n \n+sub cmd_fetch {\n+\tfetch_child_id($GIT_SVN, @_);\n+}\n+\n sub fetch {\n \tcheck_upgrade_needed();\n \t$SVN_URL ||= file_to_s(\"$GIT_SVN_DIR/info/url\");\n@@ -836,7 +840,6 @@ sub fetch_child_id {\n \tmy $ref = \"$GIT_DIR/refs/remotes/$id\";\n \tdefined(my $pid = open my $fh, '-|') or croak $!;\n \tif (!$pid) {\n-\t\t$_repack = undef;\n \t\t$GIT_SVN = $ENV{GIT_SVN_ID} = $id;\n \t\tinit_vars();\n \t\tfetch(@_);\n@@ -844,7 +847,7 @@ sub fetch_child_id {\n \t}\n \twhile (<$fh>) {\n \t\tprint $_;\n-\t\tcheck_repack() if (/^r\\d+ = $sha1/);\n+\t\tcheck_repack() if (/^r\\d+ = $sha1/o);\n \t}\n \tclose $fh or croak $?;\n }\n@@ -1407,7 +1410,6 @@ sub git_commit {\n \n \t# this output is read via pipe, do not change:\n \tprint \"r$log_msg->{revision} = $commit\\n\";\n-\tcheck_repack();\n \treturn $commit;\n }\n \n-- \n1.5.0.rc0.g0d67\n"},{"id":"30891","messageId":"204011cb0701050919w2001105asefe2fd99165dfa95@mail.gmail.com","threadId":"6214","inReplyTo":"7v1wmalez6.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-05T17:19:35Z","receivedAt":"2007-01-05T17:19:35Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"So, first up:\n\nUsing git-verify-pack from master does not fail. It actually does\nverify the pack (after a pretty decent wait.) I should have tried\nmaster first before sending out the first mail. :)\n\nIt takes about eleven minutes for git-verify-pack to complete, but it\ndoes run to completion. So something that changed between 1.4.1 and\nmaster made everything great again.\n\nI haven't tried git-prune yet, but I'll report back with the results\nfrom that next.\n\nJunio: Did you still want me to try those steps with that patch\nanyway, even though it works on master?\n"},{"id":"30894","messageId":"7vbqldfg56.fsf@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"204011cb0701050919w2001105asefe2fd99165dfa95@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-05T19:05:41Z","receivedAt":"2007-01-05T19:05:41Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Chris Lee\" <chris133@gmail.com> writes:\n\n> Using git-verify-pack from master does not fail. It actually does\n> verify the pack (after a pretty decent wait.) I should have tried\n> master first before sending out the first mail. :)\n\nDepends on which \"master\" -- I pushed out the \"chuncked hashing\"\nfix on \"master\" as commit 8977c110 as part of the update last\nnight.\n\n> Junio: Did you still want me to try those steps with that patch\n> anyway, even though it works on master?\n\nIt would give us a confirmation that the above actually fixes\nthe problem, if your 1.4.1 fails to verify that same new pack\nyou just generated, on which you saw that the \"master\" (assuming\nyou mean the one with the above patch) works correctly.\n\nIf your \"master\" before 8977c110 already passes, then there is\nsomething else going on, which would be worrysome.\n"},{"id":"30896","messageId":"204011cb0701051133r1ede14a6gd5093a3e7fa88cb5@mail.gmail.com","threadId":"6214","inReplyTo":"7vbqldfg56.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-05T19:33:45Z","receivedAt":"2007-01-05T19:33:45Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"On 1/5/07, Junio C Hamano <junkio@cox.net> wrote:\n> > Using git-verify-pack from master does not fail. It actually does\n> > verify the pack (after a pretty decent wait.) I should have tried\n> > master first before sending out the first mail. :)\n>\n> Depends on which \"master\" -- I pushed out the \"chuncked hashing\"\n> fix on \"master\" as commit 8977c110 as part of the update last\n> night.\n\nWell, that would definitely explain it. :)\n\nI did a fresh 'git pull' on master last night before I ran the\ngit-verify-pack, and that was around 11PM PST.\n\n> > Junio: Did you still want me to try those steps with that patch\n> > anyway, even though it works on master?\n>\n> It would give us a confirmation that the above actually fixes\n> the problem, if your 1.4.1 fails to verify that same new pack\n> you just generated, on which you saw that the \"master\" (assuming\n> you mean the one with the above patch) works correctly.\n>\n> If your \"master\" before 8977c110 already passes, then there is\n> something else going on, which would be worrysome.\n\nThe 'master' I had definitely included 8977c110. I can try it out with\nthe tip from before that commit, though, if you want.\n\nAlso, 'git-prune' took about 30 minutes to run to completion. Oddly,\ngit-prune didn't remove the older packs - does git-prune ignore packs?\n'git-repack -a -d' did remove them.\n"},{"id":"30898","messageId":"20070105193958.GE8753@spearce.org","threadId":"6214","inReplyTo":"204011cb0701051133r1ede14a6gd5093a3e7fa88cb5@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-01-05T19:39:58Z","receivedAt":"2007-01-05T19:39:58Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Chris Lee <chris133@gmail.com> wrote:\n> Also, 'git-prune' took about 30 minutes to run to completion. Oddly,\n> git-prune didn't remove the older packs - does git-prune ignore packs?\n> 'git-repack -a -d' did remove them.\n\ngit-prune is expensive.  Very expensive on very large projects,\nas it must iterate every object to decide what is needed, before\nit can start to remove objects that aren't needed.\n\nYes, it doesn't deal with removing pack files.  That's what the -d\nto git-repack is for.\n\n-- \nShawn.\n"},{"id":"30902","messageId":"204011cb0701051248xdd9be8ch68db18ea93abd1f6@mail.gmail.com","threadId":"6214","inReplyTo":"20070105193958.GE8753@spearce.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-05T20:48:01Z","receivedAt":"2007-01-05T20:48:01Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"On 1/5/07, Shawn O. Pearce <spearce@spearce.org> wrote:\n> git-prune is expensive.  Very expensive on very large projects,\n> as it must iterate every object to decide what is needed, before\n> it can start to remove objects that aren't needed.\n>\n> Yes, it doesn't deal with removing pack files.  That's what the -d\n> to git-repack is for.\n\nNot nearly as expensive as git-repack, that's for sure. :)\n\nAnd - I originally thought that adding '-d' to git-repack just told it\nto call 'git-prune' afterwards. It does more than that, which is cool.\nHappily importing away - up to r320k now.\n"},{"id":"30909","messageId":"7vtzz5duk1.fsf@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"20070105193958.GE8753@spearce.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-05T21:37:18Z","receivedAt":"2007-01-05T21:37:18Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Subject: [PATCH] builtin-prune: memory diet.\n\nSomehow we forgot to turn save_commit_buffer off while walking\nthe reachable objects.  Releasing the memory for commit object\ndata that we do not use matters for large projects (for example,\nabout 90MB is saved while traversing linux-2.6 history).\n\nSigned-off-by: Junio C Hamano <junkio@cox.net>\n---\n\n * The linux-2.6 history number for me is inflated because I\n   have grafts that connects historical archive behind the\n   current v2.6.12-rc2 based history...\n\n builtin-prune.c |    2 ++\n 1 files changed, 2 insertions(+), 0 deletions(-)\n\ndiff --git a/builtin-prune.c b/builtin-prune.c\nindex 00a53b3..b469c43 100644\n--- a/builtin-prune.c\n+++ b/builtin-prune.c\n@@ -253,6 +253,8 @@ int cmd_prune(int argc, const char **argv, const char *prefix)\n \t\tusage(prune_usage);\n \t}\n \n+\tsave_commit_buffer = 0;\n+\n \t/*\n \t * Set up revision parsing, and mark us as being interested\n \t * in all object types, not just commits.\n"},{"id":"30911","messageId":"Pine.LNX.4.64.0701051354590.3661@woody.osdl.org","threadId":"6214","inReplyTo":"7vtzz5duk1.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-05T21:57:31Z","receivedAt":"2007-01-05T21:57:31Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Jan 2007, Junio C Hamano wrote:\n> \n> Somehow we forgot to turn save_commit_buffer off while walking\n> the reachable objects.  Releasing the memory for commit object\n> data that we do not use matters for large projects (for example,\n> about 90MB is saved while traversing linux-2.6 history).\n\nHeh. Maybe we should just make the default the other way? It's probably \npretty easy to find any users that suddenly start segfaulting ;)\n\n(and just setting it in \"cmd_log_init\" would likely catch quite a number \nof them already).\n\n\t\tLinus\n"},{"id":"31051","messageId":"Pine.LNX.4.64.0701051414140.14017@blackbox.fnordora.org","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701051354590.3661@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"alan","fromEmail":"alan@clueserver.org","sentAt":"2007-01-05T22:18:16Z","receivedAt":"2007-01-05T22:18:16Z","isPatch":false,"sender":{"key":"alan@clueserver.org","avatar":null},"body":"On Fri, 5 Jan 2007, Linus Torvalds wrote:\n\n>\n>\n> On Fri, 5 Jan 2007, Junio C Hamano wrote:\n>>\n>> Somehow we forgot to turn save_commit_buffer off while walking\n>> the reachable objects.  Releasing the memory for commit object\n>> data that we do not use matters for large projects (for example,\n>> about 90MB is saved while traversing linux-2.6 history).\n>\n> Heh. Maybe we should just make the default the other way? It's probably\n> pretty easy to find any users that suddenly start segfaulting ;)\n\nI am trying to import a subversion repository and have yet to be able to \nsuck down the whole thing without segfaulting.  It is a large repository. \nWorks fine until about the last 10% and then runs out of memory.\n\nopen3: fork failed: Cannot allocate memory at /usr/bin/git-svn line 2711\n512 at /usr/bin/git-svn line 446\n         main::fetch_lib() called at /usr/bin/git-svn line 314\n         main::fetch() called at /usr/bin/git-svn line 173\n\nI need to try the \"partial download\" script and see if that helps.\n\n-- \n\"Invoking the supernatural can explain anything, and hence explains nothing.\"\n                   - University of Utah bioengineering professor Gregory Clark\n"},{"id":"30916","messageId":"Pine.LNX.4.64.0701051439060.3661@woody.osdl.org","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701051354590.3661@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-05T22:39:54Z","receivedAt":"2007-01-05T22:39:54Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Jan 2007, Linus Torvalds wrote:\n> \n> Heh. Maybe we should just make the default the other way? It's probably \n> pretty easy to find any users that suddenly start segfaulting ;)\n\nThis seems to pass all the tests, at least.\n\n(But I didn't test the SVN stuff, since I don't have perl::SVN installed)\n\n\t\tLinus\n---\ndiff --git a/builtin-branch.c b/builtin-branch.c\nindex d3df5a5..0b662a8 100644\n--- a/builtin-branch.c\n+++ b/builtin-branch.c\n@@ -441,6 +441,9 @@ int cmd_branch(int argc, const char **argv, const char *prefix)\n \t    (rename && force_create))\n \t\tusage(builtin_branch_usage);\n \n+\tif (verbose)\n+\t\tsave_commit_buffer = 1;\n+\n \thead = xstrdup(resolve_ref(\"HEAD\", head_sha1, 0, NULL));\n \tif (!head)\n \t\tdie(\"Failed to resolve HEAD as a valid ref.\");\ndiff --git a/builtin-diff-tree.c b/builtin-diff-tree.c\nindex 24cb2d7..212ad59 100644\n--- a/builtin-diff-tree.c\n+++ b/builtin-diff-tree.c\n@@ -67,6 +67,7 @@ int cmd_diff_tree(int argc, const char **argv, const char *prefix)\n \tstatic struct rev_info *opt = &log_tree_opt;\n \tint read_stdin = 0;\n \n+\tsave_commit_buffer = 1;\n \tinit_revisions(opt, prefix);\n \tgit_config(git_default_config); /* no \"diff\" UI options */\n \tnr_sha1 = 0;\ndiff --git a/builtin-fmt-merge-msg.c b/builtin-fmt-merge-msg.c\nindex 87d3d63..4053651 100644\n--- a/builtin-fmt-merge-msg.c\n+++ b/builtin-fmt-merge-msg.c\n@@ -251,6 +251,7 @@ int cmd_fmt_merge_msg(int argc, const char **argv, const char *prefix)\n \tunsigned char head_sha1[20];\n \tconst char *current_branch;\n \n+\tsave_commit_buffer = 1;\n \tgit_config(fmt_merge_msg_config);\n \n \twhile (argc > 1) {\ndiff --git a/builtin-log.c b/builtin-log.c\nindex a59b4ac..ac95921 100644\n--- a/builtin-log.c\n+++ b/builtin-log.c\n@@ -22,6 +22,7 @@ static void cmd_log_init(int argc, const char **argv, const char *prefix,\n {\n \tint i;\n \n+\tsave_commit_buffer = 1;\n \trev->abbrev = DEFAULT_ABBREV;\n \trev->commit_format = CMIT_FMT_DEFAULT;\n \trev->verbose_header = 1;\n@@ -372,6 +373,7 @@ int cmd_format_patch(int argc, const char **argv, const char *prefix)\n \trev.ignore_merges = 1;\n \trev.diffopt.msg_sep = \"\";\n \trev.diffopt.recursive = 1;\n+\tsave_commit_buffer = 1;\n \n \trev.extra_headers = extra_headers;\n \n@@ -569,6 +571,7 @@ int cmd_cherry(int argc, const char **argv, const char *prefix)\n \tconst char *limit = NULL;\n \tint verbose = 0;\n \n+\tsave_commit_buffer = 1;\n \tif (argc > 1 && !strcmp(argv[1], \"-v\")) {\n \t\tverbose = 1;\n \t\targc--;\ndiff --git a/builtin-show-branch.c b/builtin-show-branch.c\nindex c67f2fa..53d1b29 100644\n--- a/builtin-show-branch.c\n+++ b/builtin-show-branch.c\n@@ -586,6 +586,7 @@ int cmd_show_branch(int ac, const char **av, const char *prefix)\n \tint dense = 1;\n \tint reflog = 0;\n \n+\tsave_commit_buffer = 1;\n \tgit_config(git_show_branch_config);\n \n \t/* If nothing is specified, try the default first */\ndiff --git a/commit.c b/commit.c\nindex 2a58175..660d365 100644\n--- a/commit.c\n+++ b/commit.c\n@@ -4,7 +4,7 @@\n #include \"pkt-line.h\"\n #include \"utf8.h\"\n \n-int save_commit_buffer = 1;\n+int save_commit_buffer = 0;\n \n struct sort_node\n {\ndiff --git a/merge-recursive.c b/merge-recursive.c\nindex bac16f5..b98ed1a 100644\n--- a/merge-recursive.c\n+++ b/merge-recursive.c\n@@ -1286,6 +1286,7 @@ int main(int argc, char *argv[])\n \tconst char *branch1, *branch2;\n \tstruct commit *result, *h1, *h2;\n \n+\tsave_commit_buffer = 1;\n \tgit_config(git_default_config); /* core.filemode */\n \toriginal_index_file = getenv(INDEX_ENVIRONMENT);\n \ndiff --git a/revision.c b/revision.c\nindex 6e4ec46..aa10088 100644\n--- a/revision.c\n+++ b/revision.c\n@@ -737,6 +737,7 @@ static void add_grep(struct rev_info *revs, const char *ptn, enum grep_pat_token\n \t\topt->pattern_tail = &(opt->pattern_list);\n \t\topt->regflags = REG_NEWLINE;\n \t\trevs->grep_filter = opt;\n+\t\tsave_commit_buffer = 1;\n \t}\n \tappend_grep_pattern(revs->grep_filter, ptn,\n \t\t\t    \"command line\", 0, what);\n"},{"id":"30918","messageId":"7vac0xdr97.fsf@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701051439060.3661@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-05T22:48:36Z","receivedAt":"2007-01-05T22:48:36Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Fri, 5 Jan 2007, Linus Torvalds wrote:\n>> \n>> Heh. Maybe we should just make the default the other way? It's probably \n>> pretty easy to find any users that suddenly start segfaulting ;)\n>\n> This seems to pass all the tests, at least.\n>\n> (But I didn't test the SVN stuff, since I don't have perl::SVN installed)\n\nI do not think we have too many branch refs (builtin-branch and\nbuiltin-show-branch) for this patch to make any practical\ndifference, but I wonder why this is needed...\n\n> diff --git a/merge-recursive.c b/merge-recursive.c\n> index bac16f5..b98ed1a 100644\n> --- a/merge-recursive.c\n> +++ b/merge-recursive.c\n> @@ -1286,6 +1286,7 @@ int main(int argc, char *argv[])\n>  \tconst char *branch1, *branch2;\n>  \tstruct commit *result, *h1, *h2;\n>  \n> +\tsave_commit_buffer = 1;\n>  \tgit_config(git_default_config); /* core.filemode */\n>  \toriginal_index_file = getenv(INDEX_ENVIRONMENT);\n\nAh, there are those annoying \"using this as the merge base whose\ncommit log is...\" business.  I wonder if anybody is actually\nreading them (I once considered squelching that output).\n"},{"id":"30920","messageId":"Pine.LNX.4.64.0701051457020.3661@woody.osdl.org","threadId":"6214","inReplyTo":"7vac0xdr97.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-05T23:00:52Z","receivedAt":"2007-01-05T23:00:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Jan 2007, Junio C Hamano wrote:\n> \n> I do not think we have too many branch refs (builtin-branch and\n> builtin-show-branch) for this patch to make any practical\n> difference\n\nYeah, it's mainly a \"safety thing\" - have the default be the \"don't waste \nmemory\".\n\n> but I wonder why this is needed...\n> \n> > diff --git a/merge-recursive.c b/merge-recursive.c\n> > index bac16f5..b98ed1a 100644\n> > --- a/merge-recursive.c\n> > +++ b/merge-recursive.c\n> > @@ -1286,6 +1286,7 @@ int main(int argc, char *argv[])\n> >  \tconst char *branch1, *branch2;\n> >  \tstruct commit *result, *h1, *h2;\n> >  \n> > +\tsave_commit_buffer = 1;\n> >  \tgit_config(git_default_config); /* core.filemode */\n> >  \toriginal_index_file = getenv(INDEX_ENVIRONMENT);\n> \n> Ah, there are those annoying \"using this as the merge base whose\n> commit log is...\" business.  I wonder if anybody is actually\n> reading them (I once considered squelching that output).\n\n\"output_commit_title()\" used it. Not just for the merge base, but for the \nregular \"merging X and Y\" messages, I think.\n\n\t\tLinus\n"},{"id":"30921","messageId":"Pine.LNX.4.64.0701051501030.3661@woody.osdl.org","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701051457020.3661@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-05T23:02:05Z","receivedAt":"2007-01-05T23:02:05Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Jan 2007, Linus Torvalds wrote:\n> \n> Yeah, it's mainly a \"safety thing\" - have the default be the \"don't waste \n> memory\".\n\nBtw, I'm not at all certain whether it's necessary or a good thing. I just \ndecided to see how many people really seem to use the commit messages at \nall. So feel free to throw the patch away if you don't think this is \nworthwhile, I won't push it.\n\n\t\tLinus\n"},{"id":"30923","messageId":"204011cb0701051503m3a431e07qc12662eecc08884f@mail.gmail.com","threadId":"6214","inReplyTo":"7vtzz5duk1.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-05T23:03:26Z","receivedAt":"2007-01-05T23:03:26Z","isPatch":false,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"On 1/5/07, Junio C Hamano <junkio@cox.net> wrote:\n> Subject: [PATCH] builtin-prune: memory diet.\n>\n> Somehow we forgot to turn save_commit_buffer off while walking\n> the reachable objects.  Releasing the memory for commit object\n> data that we do not use matters for large projects (for example,\n> about 90MB is saved while traversing linux-2.6 history).\n\nIs git-verify-pack supposed to mmap the entire packfile? Because the\nversion I have maps 2.3GB into RAM and keeps it there until it's done.\n"},{"id":"30924","messageId":"7v64bldqas.fsf@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"204011cb0701051503m3a431e07qc12662eecc08884f@mail.gmail.com","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-05T23:09:15Z","receivedAt":"2007-01-05T23:09:15Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"\"Chris Lee\" <chris133@gmail.com> writes:\n\n> On 1/5/07, Junio C Hamano <junkio@cox.net> wrote:\n>> Subject: [PATCH] builtin-prune: memory diet.\n>>\n>> Somehow we forgot to turn save_commit_buffer off while walking\n>> the reachable objects.  Releasing the memory for commit object\n>> data that we do not use matters for large projects (for example,\n>> about 90MB is saved while traversing linux-2.6 history).\n>\n> Is git-verify-pack supposed to mmap the entire packfile? Because the\n> version I have maps 2.3GB into RAM and keeps it there until it's done.\n\nYes -- we need to hash the whole thing as well as doing other\nchecks on it.  Sliding mmap() in \"next\" will mmap that in chunks\nof 32MB or 1GB, but its needing to read every byte of it does\nnot change.\n\nThe problem Linus pointed out was that your SHA1_Update()\nimplementations may not be prepared to hash the whole 2.3GB in\none go.  The one in \"master\" (and \"maint\", although I haven't\ndone a v1.4.4.4 maintenance release yet) calls SHA1_Update()\nin chunks to work around that potential issue.\n"},{"id":"30927","messageId":"Pine.LNX.4.64.0701051515000.3661@woody.osdl.org","threadId":"6214","inReplyTo":"7v64bldqas.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-05T23:17:30Z","receivedAt":"2007-01-05T23:17:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Jan 2007, Junio C Hamano wrote:\n> \n> The problem Linus pointed out was that your SHA1_Update()\n> implementations may not be prepared to hash the whole 2.3GB in\n> one go.  The one in \"master\" (and \"maint\", although I haven't\n> done a v1.4.4.4 maintenance release yet) calls SHA1_Update()\n> in chunks to work around that potential issue.\n\nWell, I think Chris is worried about having it all mapped at the same \ntime.\n\nIt does actually end up forcing the kernel to do more work (it's harder to \nre-use a mapped page than it is to reuse one that isn't), and in that \nsense, if you have less than <n> GB of RAM and can't just keep it all in \nmemory at the same time, doing one large mmap is possibly more expensive \nthan chunking things up.\n\nThat said, I doubt it's a huge problem. If you can't fit the whole file in \nmemory, your real performance issue is going to be the IO, not the fact \nthat the kernel has to work a bit harder at unmapping pages ;)\n\n\t\tLinus\n"},{"id":"30932","messageId":"7virflca43.fsf@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701051457020.3661@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-05T23:44:12Z","receivedAt":"2007-01-05T23:44:12Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n>> Ah, there are those annoying \"using this as the merge base whose\n>> commit log is...\" business.  I wonder if anybody is actually\n>> reading them (I once considered squelching that output).\n>\n> \"output_commit_title()\" used it. Not just for the merge base, but for the \n> regular \"merging X and Y\" messages, I think.\n\nYes, what I really was wondering were (1) if the messages are\nuseful, and (2) if so should that belong to git-merge not\ngit-merge-recursive.\n"},{"id":"30934","messageId":"7vac0xc9g8.fsf@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701051515000.3661@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-05T23:58:31Z","receivedAt":"2007-01-05T23:58:31Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> It does actually end up forcing the kernel to do more work (it's harder to \n> re-use a mapped page than it is to reuse one that isn't), and in that \n> sense, if you have less than <n> GB of RAM and can't just keep it all in \n> memory at the same time, doing one large mmap is possibly more expensive \n> than chunking things up.\n\nEven if it is a read-only private mapping?  Would MAP_SHARED\nhelp?\n"},{"id":"30935","messageId":"Pine.LNX.4.64.0701051559020.3661@woody.osdl.org","threadId":"6214","inReplyTo":"7virflca43.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-05T23:59:51Z","receivedAt":"2007-01-05T23:59:51Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Jan 2007, Junio C Hamano wrote:\n> \n> Yes, what I really was wondering were (1) if the messages are\n> useful, and (2) if so should that belong to git-merge not\n> git-merge-recursive.\n\nI kind of like them, but I don't really look _too_ much at them, so .. \n\nI guess it would make more sense to do that at a higher level, and have \nthe low-level merger just do the actual merge itself.\n\n\t\tLinus\n"},{"id":"30936","messageId":"Pine.LNX.4.63.0701060103190.22628@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"6214","inReplyTo":"7virflca43.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2007-01-06T00:06:58Z","receivedAt":"2007-01-06T00:06:58Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 5 Jan 2007, Junio C Hamano wrote:\n\n> Linus Torvalds <torvalds@osdl.org> writes:\n> \n> >> Ah, there are those annoying \"using this as the merge base whose\n> >> commit log is...\" business.  I wonder if anybody is actually\n> >> reading them (I once considered squelching that output).\n> >\n> > \"output_commit_title()\" used it. Not just for the merge base, but for the \n> > regular \"merging X and Y\" messages, I think.\n> \n> Yes, what I really was wondering were (1) if the messages are\n> useful, and (2) if so should that belong to git-merge not\n> git-merge-recursive.\n\nSince recursive merge performs possibly more than one merge, it belongs \ninto merge-recursive.c, _if_ we want that message.\n\nI found it helpful for \"debugging\" failed _recursive_ merges. I.e. I knew \nwhich of the recursive merges introduced the many, many conflicts. But I \ncannot remember off-hand if that was a test merge, and if it was before, \nor after, I sorted the merge bases by date.\n\nSince the conflict markers now say which commit the conflicts came from, I \nam okay with removing the message, though.\n\nCiao,\nDscho\n"},{"id":"30937","messageId":"Pine.LNX.4.64.0701051610290.3661@woody.osdl.org","threadId":"6214","inReplyTo":"7vac0xc9g8.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-06T00:11:25Z","receivedAt":"2007-01-06T00:11:25Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Jan 2007, Junio C Hamano wrote:\n> \n> Even if it is a read-only private mapping?  Would MAP_SHARED\n> help?\n\nmmap is mmap, and it all boils down to having to remove it from the page \ntables.\n\nBut it really shouldn't be a problem. \n\n\t\tLinus\n"},{"id":"30938","messageId":"Pine.LNX.4.64.0701051611550.3661@woody.osdl.org","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701051610290.3661@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-06T00:15:28Z","receivedAt":"2007-01-06T00:15:28Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Jan 2007, Linus Torvalds wrote:\n> \n> But it really shouldn't be a problem. \n\nBasically, this boils down to the same old issue: if you have a fixed \naccess pattern (like SHA1_Update() over the whole buffer), you're actually \nlikely to perform better with a loop of read() calls than with mmap.\n\nSo if we ONLY did the SHA1 thing, we shouldn't do mmap, we should just \nchunk things up into 16kB buffers or something, and read them.\n\nBut the mmap in pack-check _also_ ends up being for the subsequent object \nchecking (with unpacking etc), so the mmap here actually is probably the \nright thing to do. I really wouldn't worry, unless we get people who \nreport real problems (and I think the problems with svn-import of the huge \nKDE repos are all elsewhere, notably in teh SVN import itself, not in any \npack handling ;)\n\n\t\tLinus\n"},{"id":"30941","messageId":"7v3b6pc89y.fsf@assigned-by-dhcp.cox.net","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701051611550.3661@woody.osdl.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-06T00:23:53Z","receivedAt":"2007-01-06T00:23:53Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> On Fri, 5 Jan 2007, Linus Torvalds wrote:\n>> \n>> But it really shouldn't be a problem. \n>\n> Basically, this boils down to the same old issue: if you have a fixed \n> access pattern (like SHA1_Update() over the whole buffer), you're actually \n> likely to perform better with a loop of read() calls than with mmap.\n>\n> So if we ONLY did the SHA1 thing, we shouldn't do mmap, we should just \n> chunk things up into 16kB buffers or something, and read them.\n\nWhile I have your attention, there is a patch for the sliding\nmmap() thing that raises the mmap window to 1GB (which means a\npack smaller than that is mmap'ed in its entirety, whle 2.3GB\npack will be mapped perhaps as three separate chunks) and the\ntotal mmap window to 8GB (and any overflows we LRU out) on\nplaces where sizeof(void*) == 8 (i.e. git compiled for 64-bit).\n\nCurrently these limits are 32MB and 256MB respectively on\nplatforms with real mmap().\n\nDo you have any comments on it?\n"},{"id":"30946","messageId":"Pine.LNX.4.64.0701051721340.3661@woody.osdl.org","threadId":"6214","inReplyTo":"7v3b6pc89y.fsf@assigned-by-dhcp.cox.net","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-06T01:22:16Z","receivedAt":"2007-01-06T01:22:16Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Fri, 5 Jan 2007, Junio C Hamano wrote:\n>\n> While I have your attention, there is a patch for the sliding\n> mmap() thing that raises the mmap window to 1GB (which means a\n> pack smaller than that is mmap'ed in its entirety, whle 2.3GB\n> pack will be mapped perhaps as three separate chunks) and the\n> total mmap window to 8GB (and any overflows we LRU out) on\n> places where sizeof(void*) == 8 (i.e. git compiled for 64-bit).\n> \n> Currently these limits are 32MB and 256MB respectively on\n> platforms with real mmap().\n> \n> Do you have any comments on it?\n\nI think it's fine. Most \"normal\" mmap users hopefully will only use a \nsmall portion of the mapped space, adn if they use it all, it means that \nthey needed it all, so..\n\n\t\tLinus\n"},{"id":"30988","messageId":"20070107001719.GB16771@sashak.voltaire.com","threadId":"6214","inReplyTo":"204011cb0701041404g684525fdm1d057e57a57aca92@mail.gmail.com","subject":"[PATCH] git-svnimport: support for incremental import","fromName":"Sasha Khapyorsky","fromEmail":"sashak@voltaire.com","sentAt":"2007-01-07T00:17:19Z","receivedAt":"2007-01-07T00:17:19Z","isPatch":true,"sender":{"key":"sashak@voltaire.com","avatar":null},"body":"This adds ability to do import \"in chunks\" (default 1000 revisions),\nafter each chunk git repo will be repacked. The option -R is used to\nchange default value of chunk size (or how often repository will\nrepacked).\n\nSigned-off-by: Sasha Khapyorsky <sashak@voltaire.com>\n---\n\nChris reported successful test with this patch.\n\n Documentation/git-svnimport.txt |   10 +++++++++-\n git-svnimport.perl              |   32 +++++++++++++++++++++++++-------\n 2 files changed, 34 insertions(+), 8 deletions(-)\n\ndiff --git a/Documentation/git-svnimport.txt b/Documentation/git-svnimport.txt\nindex 2c7c7da..b166cf3 100644\n--- a/Documentation/git-svnimport.txt\n+++ b/Documentation/git-svnimport.txt\n@@ -15,7 +15,7 @@ SYNOPSIS\n \t\t[ -b branch_subdir ] [ -T trunk_subdir ] [ -t tag_subdir ]\n \t\t[ -s start_chg ] [ -m ] [ -r ] [ -M regex ]\n \t\t[ -I <ignorefile_name> ] [ -A <author_file> ]\n-\t\t[ -P <path_from_trunk> ]\n+\t\t[ -R <repack_each_revs>] [ -P <path_from_trunk> ]\n \t\t<SVN_repository_URL> [ <path> ]\n \n \n@@ -108,6 +108,14 @@ repository without -A.\n Formerly, this option controlled how many revisions to pull,\n due to SVN memory leaks. (These have been worked around.)\n \n+-R <repack_each_revs>::\n+\tSpecify how often git repository should be repacked.\n++\n+The default value is 1000. git-svnimport will do import in chunks of 1000\n+revisions, after each chunk git repository will be repacked. To disable\n+this behavior specify some big value here which is mote than number of\n+revisions to import.\n+\n -P <path_from_trunk>::\n \tPartial import of the SVN tree.\n +\ndiff --git a/git-svnimport.perl b/git-svnimport.perl\nindex 071777b..afbbe63 100755\n--- a/git-svnimport.perl\n+++ b/git-svnimport.perl\n@@ -31,12 +31,13 @@ $SIG{'PIPE'}=\"IGNORE\";\n $ENV{'TZ'}=\"UTC\";\n \n our($opt_h,$opt_o,$opt_v,$opt_u,$opt_C,$opt_i,$opt_m,$opt_M,$opt_t,$opt_T,\n-    $opt_b,$opt_r,$opt_I,$opt_A,$opt_s,$opt_l,$opt_d,$opt_D,$opt_S,$opt_F,$opt_P);\n+    $opt_b,$opt_r,$opt_I,$opt_A,$opt_s,$opt_l,$opt_d,$opt_D,$opt_S,$opt_F,\n+    $opt_P,$opt_R);\n \n sub usage() {\n \tprint STDERR <<END;\n Usage: ${\\basename $0}     # fetch/update GIT from SVN\n-       [-o branch-for-HEAD] [-h] [-v] [-l max_rev]\n+       [-o branch-for-HEAD] [-h] [-v] [-l max_rev] [-R repack_each_revs]\n        [-C GIT_repository] [-t tagname] [-T trunkname] [-b branchname]\n        [-d|-D] [-i] [-u] [-r] [-I ignorefilename] [-s start_chg]\n        [-m] [-M regex] [-A author_file] [-S] [-F] [-P project_name] [SVN_URL]\n@@ -44,7 +45,7 @@ END\n \texit(1);\n }\n \n-getopts(\"A:b:C:dDFhiI:l:mM:o:rs:t:T:SP:uv\") or usage();\n+getopts(\"A:b:C:dDFhiI:l:mM:o:rs:t:T:SP:R:uv\") or usage();\n usage if $opt_h;\n \n my $tag_name = $opt_t || \"tags\";\n@@ -52,6 +53,7 @@ my $trunk_name = $opt_T || \"trunk\";\n my $branch_name = $opt_b || \"branches\";\n my $project_name = $opt_P || \"\";\n $project_name = \"/\" . $project_name if ($project_name);\n+my $repack_after = $opt_R || 1000;\n \n @ARGV == 1 or @ARGV == 2 or usage();\n \n@@ -938,11 +940,27 @@ if ($opt_l < $current_rev) {\n     exit;\n }\n \n-print \"Fetching from $current_rev to $opt_l ...\\n\" if $opt_v;\n+print \"Processing from $current_rev to $opt_l ...\\n\" if $opt_v;\n \n-my $pool=SVN::Pool->new;\n-$svn->{'svn'}->get_log(\"/\",$current_rev,$opt_l,0,1,1,\\&commit_all,$pool);\n-$pool->clear;\n+my $from_rev;\n+my $to_rev = $current_rev;\n+\n+while ($to_rev < $opt_l) {\n+\t$from_rev = $to_rev;\n+\t$to_rev = $from_rev + $repack_after;\n+\t$to_rev = $opt_l if $opt_l < $to_rev;\n+\tprint \"Fetching from $from_rev to $to_rev ...\\n\" if $opt_v;\n+\tmy $pool=SVN::Pool->new;\n+\t$svn->{'svn'}->get_log(\"/\",$from_rev,$to_rev,0,1,1,\\&commit_all,$pool);\n+\t$pool->clear;\n+\tmy $pid = fork();\n+\tdie \"Fork: $!\\n\" unless defined $pid;\n+\tunless($pid) {\n+\t\texec(\"git-repack\", \"-d\")\n+\t\t\tor die \"Cannot repack: $!\\n\";\n+\t}\n+\twaitpid($pid, 0);\n+}\n \n \n unlink($git_index);\n-- \n1.5.0.rc0.g2484-dirty\n"},{"id":"30991","messageId":"20070107003654.GA12551@localdomain","threadId":"6214","inReplyTo":"Pine.LNX.4.64.0701051414140.14017@blackbox.fnordora.org","subject":"Re: git-svnimport failed and now git-repack hates me","fromName":"Eric Wong","fromEmail":"normalperson@yhbt.net","sentAt":"2007-01-07T00:36:54Z","receivedAt":"2007-01-07T00:36:54Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"alan <alan@clueserver.org> wrote:\n> I am trying to import a subversion repository and have yet to be able to \n> suck down the whole thing without segfaulting.  It is a large repository. \n> Works fine until about the last 10% and then runs out of memory.\n> \n> open3: fork failed: Cannot allocate memory at /usr/bin/git-svn line 2711\n> 512 at /usr/bin/git-svn line 446\n>         main::fetch_lib() called at /usr/bin/git-svn line 314\n>         main::fetch() called at /usr/bin/git-svn line 173\n> \n> I need to try the \"partial download\" script and see if that helps.\n\nWhich version of git-svn is this?  If it's a public repository I'd\nlike to have a look.\n\ngit-svn memory usage should be bounded by:\n\tmax(max(commit-message size),\n\t    max(number of files changed per revision))\n\nI'm not sure if the size of the files changed per-revision or if the\nsize of the deltas is an issue with git-svn.  But if you have a repo\nwith big files and big changes to them, let me know so I can take a\nlook.\n\nCan you also try lowering $inc in git-svn to something lower (perhaps\n100)? (my $inc = 1000; in the fetch_lib function) and see if that helps\nthings?  Thanks.\n\n-- \nEric Wong\n"},{"id":"31062","messageId":"204011cb0701071012g30cb69a5h4622d94574d10521@mail.gmail.com","threadId":"6214","inReplyTo":"20070107001719.GB16771@sashak.voltaire.com","subject":"Re: [PATCH] git-svnimport: support for incremental import","fromName":"Chris Lee","fromEmail":"chris133@gmail.com","sentAt":"2007-01-07T18:12:24Z","receivedAt":"2007-01-07T18:12:24Z","isPatch":true,"sender":{"key":"chris133@gmail.com","avatar":null},"body":"On 1/6/07, Sasha Khapyorsky <sashak@voltaire.com> wrote:\n> This adds ability to do import \"in chunks\" (default 1000 revisions),\n> after each chunk git repo will be repacked. The option -R is used to\n> change default value of chunk size (or how often repository will\n> repacked).\n\nActually, I just noticed an issue here with this - it appears to be\ndouble-importing the edge revisions.\n\nSo if I started with -s 349000 and tell it to repack every 1000\nrevisions, it's now importing every thousandth revision twice.\n\nOff-by-one?\n"},{"id":"31064","messageId":"20070107185906.GD18379@sashak.voltaire.com","threadId":"6214","inReplyTo":"204011cb0701071012g30cb69a5h4622d94574d10521@mail.gmail.com","subject":"Re: [PATCH] git-svnimport: support for incremental import","fromName":"Sasha Khapyorsky","fromEmail":"sashak@voltaire.com","sentAt":"2007-01-07T18:59:06Z","receivedAt":"2007-01-07T18:59:06Z","isPatch":true,"sender":{"key":"sashak@voltaire.com","avatar":null},"body":"On 10:12 Sun 07 Jan     , Chris Lee wrote:\n> On 1/6/07, Sasha Khapyorsky <sashak@voltaire.com> wrote:\n> >This adds ability to do import \"in chunks\" (default 1000 revisions),\n> >after each chunk git repo will be repacked. The option -R is used to\n> >change default value of chunk size (or how often repository will\n> >repacked).\n> \n> Actually, I just noticed an issue here with this - it appears to be\n> double-importing the edge revisions.\n> \n> So if I started with -s 349000 and tell it to repack every 1000\n> revisions, it's now importing every thousandth revision twice.\n\nIndeed. There is the fix:\n\n\ndiff --git a/git-svnimport.perl b/git-svnimport.perl\nindex afbbe63..f1f1a7d 100755\n--- a/git-svnimport.perl\n+++ b/git-svnimport.perl\n@@ -943,10 +943,10 @@ if ($opt_l < $current_rev) {\n print \"Processing from $current_rev to $opt_l ...\\n\" if $opt_v;\n \n my $from_rev;\n-my $to_rev = $current_rev;\n+my $to_rev = $current_rev - 1;\n \n while ($to_rev < $opt_l) {\n-\t$from_rev = $to_rev;\n+\t$from_rev = $to_rev + 1;\n \t$to_rev = $from_rev + $repack_after;\n \t$to_rev = $opt_l if $opt_l < $to_rev;\n \tprint \"Fetching from $from_rev to $to_rev ...\\n\" if $opt_v;\n\n\nSasha\n"},{"id":"31105","messageId":"20070108022242.GA19217@sashak.voltaire.com","threadId":"6214","inReplyTo":"20070107185906.GD18379@sashak.voltaire.com","subject":"[PATCH] git-svnimport: fix edge revisions double importing","fromName":"Sasha Khapyorsky","fromEmail":"sashak@voltaire.com","sentAt":"2007-01-08T02:22:42Z","receivedAt":"2007-01-08T02:22:42Z","isPatch":true,"sender":{"key":"sashak@voltaire.com","avatar":null},"body":"This fixes newly introduced bug when the incremental cycle edge revisions\nare imported twice.\n\nSigned-off-by: Sasha Khapyorsky <sashak@voltaire.com>\n---\n git-svnimport.perl |    4 ++--\n 1 files changed, 2 insertions(+), 2 deletions(-)\n\ndiff --git a/git-svnimport.perl b/git-svnimport.perl\nindex afbbe63..f1f1a7d 100755\n--- a/git-svnimport.perl\n+++ b/git-svnimport.perl\n@@ -943,10 +943,10 @@ if ($opt_l < $current_rev) {\n print \"Processing from $current_rev to $opt_l ...\\n\" if $opt_v;\n \n my $from_rev;\n-my $to_rev = $current_rev;\n+my $to_rev = $current_rev - 1;\n \n while ($to_rev < $opt_l) {\n-\t$from_rev = $to_rev;\n+\t$from_rev = $to_rev + 1;\n \t$to_rev = $from_rev + $repack_after;\n \t$to_rev = $opt_l if $opt_l < $to_rev;\n \tprint \"Fetching from $from_rev to $to_rev ...\\n\" if $opt_v;\n-- \n1.5.0.rc0.g2484-dirty\n"}]}