git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Optimizing cloning of a high object count repository

From
Nicolas Pitre <nico@cam.org>
Date
Dec 13, 2008, 18:56 UTC
Message-ID
<alpine.LFD.2.00.0812131347130.30035@xanadu.home>
In-Reply-To
<200812131714.05472.Resul-Cetin@gmx.net>
On Sat, 13 Dec 2008, Resul Cetin wrote:
Show 18 quoted lines
> On Saturday 13 December 2008 16:46:50 you wrote:
> [...]
> > >  The size of the linux repository seems to be smaller but in the same
> > > range object count and repository size but clones are much much faster.
> > > Is there any way to optimize the server operations like counting and
> > > compressing of objects to get the same speed as we get from
> > > git.kernel.org (which does it in nearly no time and the only limiting
> > > factor seems to be my bandwith)?
> > >  The only other information I have is that Robin H. Johnson made a single
> > >  ~910MiB pack for the whole repository.
> >
> > Make yearly packed repository snapshots and publish them via http.
> > People can wget the latest snapshot, then pull updates later.
> That would be a workaround but it doesn't explain why git.kernel.org deliveres 
> torvalds repository without any notable counting and compressing time. Maybe 
> it has something todo with the config I found inside the repository:
> http://git.overlays.gentoo.org/gitroot/exp/gentoo-x86.git/config
> It says that it isnt a bare repository.
That's not relevant.

The counting time is a bit unfortunate (although I have plans to speed that up, if only I can find the time).

You should be able to skip the compression time entirely though, if you do repack the repository first. And you want it to be as tightly packed as possible for public access. I'm currently cloning it and the counting phase is not _that_ bad compared to the compression phase. Try something like 'git repack -a -f -d --window=200' and let it run overnight if necessary. You need to do this only once, and preferably on a machine with lots of RAM, and preferably on a 64-bit machine. Once this is done then things should go much more smoothly afterwards.

Nicolas
Previous: Resul CetinNext: Nicolas Pitre
Message 6 of 7 in “Optimizing cloning of a high object count repository”
  1. Resul CetinDec 13, 2008
  2. Nguyen Thai Ngoc DuyDec 13, 2008
  3. Resul CetinDec 13, 2008
  4. Jean-Luc HerrenDec 13, 2008
  5. Resul CetinDec 13, 2008
  6. Nicolas PitreDec 13, 2008
  7. Nicolas PitreDec 13, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.