{"thread":{"id":"16713","subject":"Optimizing cloning of a high object count repository","startedAt":"2008-12-13T15:24:56Z","lastAt":"2008-12-13T21:50:52Z","messageCount":7,"participants":["Resul Cetin","Nguyen Thai Ngoc Duy","Jean-Luc Herren","Nicolas Pitre"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"97796","messageId":"200812131624.57618.Resul-Cetin@gmx.net","threadId":"16713","inReplyTo":null,"subject":"Optimizing cloning of a high object count repository","fromName":"Resul Cetin","fromEmail":"resul-cetin@gmx.net","sentAt":"2008-12-13T15:24:56Z","receivedAt":"2008-12-13T15:24:56Z","isPatch":false,"sender":{"key":"resul-cetin@gmx.net","avatar":null},"body":"Hi,\nthere are currently different ideas to move gentoo's cvs repository to an \nother scm. Current tests showed that svn will not make anything better (it \ngets in most perfomance and size based benchmarks even worse). Another idea is \nto move to git. It looks really promising in size based benchmarks but cloning \nseems nearly impossible. The current test repository is available at \ngit://git.overlays.gentoo.org/exp/gentoo-x86.git and is around 900MB in size \nand has 4696137 objects. It really takes ages to do the counting of the \nobjects on the server and compressing takes much longer.\nThe size of the linux repository seems to be smaller but in the same range \nobject count and repository size but clones are much much faster. Is there any \nway to optimize the server operations like counting and compressing of objects \nto get the same speed as we get from git.kernel.org (which does it in nearly \nno time and the only limiting factor seems to be my bandwith)?\nThe only other information I have is that Robin H. Johnson made a single \n~910MiB pack for the whole repository.\n\nThx in advance,\n\tResul\n"},{"id":"97797","messageId":"fcaeb9bf0812130746l38a12f37wde26f31d5fa0d2a2@mail.gmail.com","threadId":"16713","inReplyTo":"200812131624.57618.Resul-Cetin@gmx.net","subject":"Re: Optimizing cloning of a high object count repository","fromName":"Nguyen Thai Ngoc Duy","fromEmail":"pclouds@gmail.com","sentAt":"2008-12-13T15:46:50Z","receivedAt":"2008-12-13T15:46:50Z","isPatch":false,"sender":{"key":"pclouds@gmail.com","avatar":"https://avatars.githubusercontent.com/u/720?v=4"},"body":"On 12/13/08, Resul Cetin <Resul-Cetin@gmx.net> wrote:\n> Hi,\n>  there are currently different ideas to move gentoo's cvs repository to an\n>  other scm. Current tests showed that svn will not make anything better (it\n>  gets in most perfomance and size based benchmarks even worse). Another idea is\n>  to move to git. It looks really promising in size based benchmarks but cloning\n>  seems nearly impossible. The current test repository is available at\n>  git://git.overlays.gentoo.org/exp/gentoo-x86.git and is around 900MB in size\n>  and has 4696137 objects. It really takes ages to do the counting of the\n>  objects on the server and compressing takes much longer.\n>  The size of the linux repository seems to be smaller but in the same range\n>  object count and repository size but clones are much much faster. Is there any\n>  way to optimize the server operations like counting and compressing of objects\n>  to get the same speed as we get from git.kernel.org (which does it in nearly\n>  no time and the only limiting factor seems to be my bandwith)?\n>  The only other information I have is that Robin H. Johnson made a single\n>  ~910MiB pack for the whole repository.\n\nMake yearly packed repository snapshots and publish them via http.\nPeople can wget the latest snapshot, then pull updates later.\n-- \nDuy\n"},{"id":"97798","messageId":"200812131714.05472.Resul-Cetin@gmx.net","threadId":"16713","inReplyTo":"fcaeb9bf0812130746l38a12f37wde26f31d5fa0d2a2@mail.gmail.com","subject":"Re: Optimizing cloning of a high object count repository","fromName":"Resul Cetin","fromEmail":"resul-cetin@gmx.net","sentAt":"2008-12-13T16:14:04Z","receivedAt":"2008-12-13T16:14:04Z","isPatch":false,"sender":{"key":"resul-cetin@gmx.net","avatar":null},"body":"On Saturday 13 December 2008 16:46:50 you wrote:\n[...]\n> >  The size of the linux repository seems to be smaller but in the same\n> > range object count and repository size but clones are much much faster.\n> > Is there any way to optimize the server operations like counting and\n> > compressing of objects to get the same speed as we get from\n> > git.kernel.org (which does it in nearly no time and the only limiting\n> > factor seems to be my bandwith)?\n> >  The only other information I have is that Robin H. Johnson made a single\n> >  ~910MiB pack for the whole repository.\n>\n> Make yearly packed repository snapshots and publish them via http.\n> People can wget the latest snapshot, then pull updates later.\nThat would be a workaround but it doesn't explain why git.kernel.org deliveres \ntorvalds repository without any notable counting and compressing time. Maybe \nit has something todo with the config I found inside the repository:\nhttp://git.overlays.gentoo.org/gitroot/exp/gentoo-x86.git/config\nIt says that it isnt a bare repository.\nBefore I forget. I was wrong that it is a single 910mb file. Somebody seems to \nhave repacked it into 7 single packs.\n\nRegards,\n\tResul\n"},{"id":"97799","messageId":"4943E657.9040204@gmx.ch","threadId":"16713","inReplyTo":"200812131714.05472.Resul-Cetin@gmx.net","subject":"Re: Optimizing cloning of a high object count repository","fromName":"Jean-Luc Herren","fromEmail":"jlh@gmx.ch","sentAt":"2008-12-13T16:44:07Z","receivedAt":"2008-12-13T16:44:07Z","isPatch":false,"sender":{"key":"jlh@gmx.ch","avatar":null},"body":"Resul Cetin wrote:\n> That would be a workaround but it doesn't explain why git.kernel.org deliveres \n> torvalds repository without any notable counting and compressing time.\n\nIf I remember right, git.kernel.org is a quite beefy machine.  But\nthen again it has a lot more traffic too.  It might be interesting\nto know what machine you're on, compared to git.kernel.org.\n\njlh\n"},{"id":"97803","messageId":"200812131920.50736.Resul-Cetin@gmx.net","threadId":"16713","inReplyTo":"4943E657.9040204@gmx.ch","subject":"Re: Optimizing cloning of a high object count repository","fromName":"Resul Cetin","fromEmail":"resul-cetin@gmx.net","sentAt":"2008-12-13T18:20:50Z","receivedAt":"2008-12-13T18:20:50Z","isPatch":false,"sender":{"key":"resul-cetin@gmx.net","avatar":null},"body":"On Saturday 13 December 2008 17:44:07 you wrote:\n> Resul Cetin wrote:\n> > That would be a workaround but it doesn't explain why git.kernel.org\n> > deliveres torvalds repository without any notable counting and\n> > compressing time.\n>\n> If I remember right, git.kernel.org is a quite beefy machine.  But\n> then again it has a lot more traffic too.  It might be interesting\n> to know what machine you're on, compared to git.kernel.org.\nI dont know what type of machine git.overlay.g.o is but my athlon64 3500+ with \n4GB ram has exactly the same problem without any other load. I made a clone  \nover http and did no other changes to the repository until now.\n\nhttp://git.overlays.gentoo.org/gitroot/exp/gentoo-x86.git/ is the http clone \nurl.\n\nI will try some stuff to reduce the time spend before sending anything..... If \nanyone has some ideas how to do that....\n\nRegards,\n\tResul\n"},{"id":"97804","messageId":"alpine.LFD.2.00.0812131347130.30035@xanadu.home","threadId":"16713","inReplyTo":"200812131714.05472.Resul-Cetin@gmx.net","subject":"Re: Optimizing cloning of a high object count repository","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-12-13T18:56:19Z","receivedAt":"2008-12-13T18:56:19Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 13 Dec 2008, Resul Cetin wrote:\n\n> On Saturday 13 December 2008 16:46:50 you wrote:\n> [...]\n> > >  The size of the linux repository seems to be smaller but in the same\n> > > range object count and repository size but clones are much much faster.\n> > > Is there any way to optimize the server operations like counting and\n> > > compressing of objects to get the same speed as we get from\n> > > git.kernel.org (which does it in nearly no time and the only limiting\n> > > factor seems to be my bandwith)?\n> > >  The only other information I have is that Robin H. Johnson made a single\n> > >  ~910MiB pack for the whole repository.\n> >\n> > Make yearly packed repository snapshots and publish them via http.\n> > People can wget the latest snapshot, then pull updates later.\n> That would be a workaround but it doesn't explain why git.kernel.org deliveres \n> torvalds repository without any notable counting and compressing time. Maybe \n> it has something todo with the config I found inside the repository:\n> http://git.overlays.gentoo.org/gitroot/exp/gentoo-x86.git/config\n> It says that it isnt a bare repository.\n\nThat's not relevant.\n\nThe counting time is a bit unfortunate (although I have plans to speed \nthat up, if only I can find the time).\n\nYou should be able to skip the compression time entirely though, if you \ndo repack the repository first.  And you want it to be as tightly packed \nas possible for public access.  I'm currently cloning it and the \ncounting phase is not _that_ bad compared to the compression phase.  Try \nsomething like 'git repack -a -f -d --window=200' and let it run \novernight if necessary.  You need to do this only once, and preferably \non a machine with lots of RAM, and preferably on a 64-bit machine.  Once \nthis is done then things should go much more smoothly afterwards.\n\n\nNicolas\n"},{"id":"97818","messageId":"alpine.LFD.2.00.0812131636330.30035@xanadu.home","threadId":"16713","inReplyTo":"alpine.LFD.2.00.0812131347130.30035@xanadu.home","subject":"Re: Optimizing cloning of a high object count repository","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2008-12-13T21:50:52Z","receivedAt":"2008-12-13T21:50:52Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sat, 13 Dec 2008, Nicolas Pitre wrote:\n\n> On Sat, 13 Dec 2008, Resul Cetin wrote:\n> \n> > On Saturday 13 December 2008 16:46:50 you wrote:\n> > [...]\n> > > >  The size of the linux repository seems to be smaller but in the same\n> > > > range object count and repository size but clones are much much faster.\n> > > > Is there any way to optimize the server operations like counting and\n> > > > compressing of objects to get the same speed as we get from\n> > > > git.kernel.org (which does it in nearly no time and the only limiting\n> > > > factor seems to be my bandwith)?\n> > > >  The only other information I have is that Robin H. Johnson made a single\n> > > >  ~910MiB pack for the whole repository.\n> > >\n> > > Make yearly packed repository snapshots and publish them via http.\n> > > People can wget the latest snapshot, then pull updates later.\n> > That would be a workaround but it doesn't explain why git.kernel.org deliveres \n> > torvalds repository without any notable counting and compressing time. Maybe \n> > it has something todo with the config I found inside the repository:\n> > http://git.overlays.gentoo.org/gitroot/exp/gentoo-x86.git/config\n> > It says that it isnt a bare repository.\n> \n> That's not relevant.\n> \n> The counting time is a bit unfortunate (although I have plans to speed \n> that up, if only I can find the time).\n> \n> You should be able to skip the compression time entirely though, if you \n> do repack the repository first.  And you want it to be as tightly packed \n> as possible for public access.  I'm currently cloning it and the \n> counting phase is not _that_ bad compared to the compression phase.  Try \n> something like 'git repack -a -f -d --window=200' and let it run \n> overnight if necessary.  You need to do this only once, and preferably \n> on a machine with lots of RAM, and preferably on a 64-bit machine.  Once \n> this is done then things should go much more smoothly afterwards.\n\nFYI, I repacked that repository after cloning it, and that operation \nrequired around 2.5G of resident memory.  Given the address space \nfragmentation, it is possible that a full repack cannot be performed on \na 32-bit machine.\n\nI did 'git repack -a -f -d --window=500 --depth=100'.  This took less \nthan an hour on a quad core machine.  The resulting pack is 695MB in \nsize.  That's the amount of data that would be transfered during a \nclone of this repository, and nothing would have to be compressed during \nthe clone as everything is already fully compressed.\n\n\nNicolas\n"}]}