threads / discuss / 20371

git gc expanding packed data?

Subject: git gc expanding packed data?

## tl;dr

7 messages between Aug 4, 2009 and Aug 13, 2009.

replies: 6people: 2as markdown or json

Hin-Tak Leung· Aug 4, 2009, 20:25 UTC · lore

I cloned gcc's git about a week ago to work on some problems I have with gcc on minor platforms, just plain 'git clone git://gcc.gnu.org/git/gcc.git gcc' .and ran gcc fetch about daily, and 'git rebase origin' from time to time. I don't have local changes, just following and monitoring what's going on in gcc. So after a week, I thought I'd do a git gc . Then it goes very bizarre.

Before I start 'git gc', .The whole of .git was about 700MB and git/objects/pack was a bit under 600MB, with a few other directories under .git/objects at 10's of K's and a few 30000-40000K's, and the checkout was, well, the size of gcc source code. But after I started git gc, the message stays in the 'counting objects' at about 900,000 for a long time, while a lot of directories under .git/objects/ gets a bit large, and .git blows up to at least 7GB with a lot of small files under .git/objects/*/, before seeing as I will run out of disk space, I kill the whole lot and ran git clone again, since I don't have any local change and there is nothing to lose.

I am running git version 1.6.2.5 (fedora 11). Is there any reason why 'git gc' does that?

Nicolas Pitre· Aug 5, 2009, 22:39 UTC · re: Hin-Tak Leung · lore

Re: git gc expanding packed data?

On Tue, 4 Aug 2009, Hin-Tak Leung wrote:
Show 20 quoted lines
> I cloned gcc's git about a week ago to work on some problems I have
> with gcc on minor platforms, just plain 'git clone
> git://gcc.gnu.org/git/gcc.git gcc' .and ran gcc fetch about daily, and
> 'git rebase origin' from time to time. I don't have local changes,
> just following and monitoring what's going on in gcc. So after a week,
> I thought I'd do a git gc . Then it goes very bizarre.
> 
> Before I start 'git gc', .The whole of .git was about 700MB and
> git/objects/pack was a bit under 600MB, with a few other directories
> under .git/objects at 10's of K's and a few 30000-40000K's, and the
> checkout was, well, the size of gcc source code. But after I started
> git gc, the message stays in the 'counting objects' at about 900,000
> for a long time, while a lot of directories under .git/objects/ gets a
> bit large, and .git blows up to at least 7GB with a lot of small files
> under .git/objects/*/, before seeing as I will run out of disk space,
> I kill the whole lot and ran git clone again, since I don't have any
> local change and there is nothing to lose.
> 
> I am running git version 1.6.2.5 (fedora 11). Is there any reason why
> 'git gc' does that?
There is probably a reason, although a bad one for sure.
Well... OK.

It appears that the git installation serving clone requests for git://gcc.gnu.org/git/gcc.git generates lots of unreferenced objects. I just cloned it and the pack I was sent contains 1383356 objects (can be determined with 'git show-index < .git/objects/pack/*.idx | wc -l'). However, there are only 978501 actually referenced objects in that cloned repository ( 'git rev-list --all --objects | wc -l'). That makes for 404855 useless objects in the cloned repository.

Now git has a safety mechanism to _not_ delete unreferenced objects right away when running 'git gc'. By default unreferenced objects are kept around for a period of 2 weeks. This is to make it easy for you to recover accidentally deleted branches or commits, or to avoid a race where a just-created object in the process of being but not yet referenced could be deleted by a 'git gc' process running in parallel.

So to give that grace period to packed but unreferenced objects, the repack process pushes those unreferenced objects out of the pack into their loose form so they can be aged and eventually pruned. Objects becoming unreferenced are usually not that many though. Having 404855 unreferenced objects is quite a lot, and being sent those objects in the first place via a clone is stupid and a complete waste of network bandwidth.

Anyone has an idea of the git version running on gcc.gnu.org? It is certainly buggy and needs fixing.

Anyway... To solve your problem, you simply need to run 'git gc' with the --prune=now argument to disable that grace period and get rid of those unreferenced objects right away (safe only if no other git activities are taking place at the same time which should be easy to ensure on a workstation). The resulting .git/objects directory size will shrink to about 441 MB. If the gcc.gnu.org git server was doing its job properly, the size of the clone transfer would also be significantly smaller, meaning around 414 MB instead of the current 600+ MB.

And BTW, using 'git gc --aggressive' with a later git version (or 'git repack -a -f -d --window=250 --depth=250') gives me a .git/objects directory size of 310 MB, meaning that the actual repository with all the trunk history is _smaller_ than the actual source checkout. If that repository was properly repacked on the server, the clone data transfer would be 283 MB. This is less than half the current clone transfer size.

Nicolas
Hin-Tak Leung· Aug 11, 2009, 10:17 UTC · re: Nicolas Pitre · lore

Re: git gc expanding packed data?

On Wed, Aug 5, 2009 at 11:39 PM, Nicolas Pitre<nico@cam.org> wrote:
Show 74 quoted lines
> On Tue, 4 Aug 2009, Hin-Tak Leung wrote:
>
>> I cloned gcc's git about a week ago to work on some problems I have
>> with gcc on minor platforms, just plain 'git clone
>> git://gcc.gnu.org/git/gcc.git gcc' .and ran gcc fetch about daily, and
>> 'git rebase origin' from time to time. I don't have local changes,
>> just following and monitoring what's going on in gcc. So after a week,
>> I thought I'd do a git gc . Then it goes very bizarre.
>>
>> Before I start 'git gc', .The whole of .git was about 700MB and
>> git/objects/pack was a bit under 600MB, with a few other directories
>> under .git/objects at 10's of K's and a few 30000-40000K's, and the
>> checkout was, well, the size of gcc source code. But after I started
>> git gc, the message stays in the 'counting objects' at about 900,000
>> for a long time, while a lot of directories under .git/objects/ gets a
>> bit large, and .git blows up to at least 7GB with a lot of small files
>> under .git/objects/*/, before seeing as I will run out of disk space,
>> I kill the whole lot and ran git clone again, since I don't have any
>> local change and there is nothing to lose.
>>
>> I am running git version 1.6.2.5 (fedora 11). Is there any reason why
>> 'git gc' does that?
>
> There is probably a reason, although a bad one for sure.
>
> Well... OK.
>
> It appears that the git installation serving clone requests for
> git://gcc.gnu.org/git/gcc.git generates lots of unreferenced objects. I
> just cloned it and the pack I was sent contains 1383356 objects (can be
> determined with 'git show-index < .git/objects/pack/*.idx | wc -l').
> However, there are only 978501 actually referenced objects in that
> cloned repository ( 'git rev-list --all --objects | wc -l').  That makes
> for 404855 useless objects in the cloned repository.
>
> Now git has a safety mechanism to _not_ delete unreferenced objects
> right away when running 'git gc'.  By default unreferenced objects are
> kept around for a period of 2 weeks.  This is to make it easy for you to
> recover accidentally deleted branches or commits, or to avoid a race
> where a just-created object in the process of being but not yet
> referenced could be deleted by a 'git gc' process running in parallel.
>
> So to give that grace period to packed but unreferenced objects, the
> repack process pushes those unreferenced objects out of the pack into
> their loose form so they can be aged and eventually pruned.  Objects
> becoming unreferenced are usually not that many though.  Having 404855
> unreferenced objects is quite a lot, and being sent those objects in the
> first place via a clone is stupid and a complete waste of network
> bandwidth.
>
> Anyone has an idea of the git version running on gcc.gnu.org?  It is
> certainly buggy and needs fixing.
>
> Anyway... To solve your problem, you simply need to run 'git gc' with
> the --prune=now argument to disable that grace period and get rid of
> those unreferenced objects right away (safe only if no other git
> activities are taking place at the same time which should be easy to
> ensure on a workstation).  The resulting .git/objects directory size
> will shrink to about 441 MB.  If the gcc.gnu.org git server was doing
> its job properly, the size of the clone transfer would also be
> significantly smaller, meaning around 414 MB instead of the current 600+
> MB.
>
> And BTW, using 'git gc --aggressive' with a later git version (or
> 'git repack -a -f -d --window=250 --depth=250') gives me a .git/objects
> directory size of 310 MB, meaning that the actual repository with all
> the trunk history is _smaller_ than the actual source checkout.  If that
> repository was properly repacked on the server, the clone data transfer
> would be 283 MB.  This is less than half the current clone transfer
> size.
>
>
> Nicolas
>

'git gc --prune=now' does work, but 'git gc --prune=now --aggressive' (before) and 'git gc --aggressive' (after) both create very large (>2GB; I stopped it) packs from the ~400MB-600MB packed objects. I noted that you specifically wrote 'with a later git version' - presumably there is a some sort of a known and fixed issue there? Just curious.

I guess --aggressive doesn't always save space...
Hin-Tak
Nicolas Pitre· Aug 11, 2009, 21:33 UTC · re: Hin-Tak Leung · lore

Re: git gc expanding packed data?

On Tue, 11 Aug 2009, Hin-Tak Leung wrote:
Show 29 quoted lines
> On Wed, Aug 5, 2009 at 11:39 PM, Nicolas Pitre<nico@cam.org> wrote:
> > Anyway... To solve your problem, you simply need to run 'git gc' with
> > the --prune=now argument to disable that grace period and get rid of
> > those unreferenced objects right away (safe only if no other git
> > activities are taking place at the same time which should be easy to
> > ensure on a workstation).  The resulting .git/objects directory size
> > will shrink to about 441 MB.  If the gcc.gnu.org git server was doing
> > its job properly, the size of the clone transfer would also be
> > significantly smaller, meaning around 414 MB instead of the current 600+
> > MB.
> >
> > And BTW, using 'git gc --aggressive' with a later git version (or
> > 'git repack -a -f -d --window=250 --depth=250') gives me a .git/objects
> > directory size of 310 MB, meaning that the actual repository with all
> > the trunk history is _smaller_ than the actual source checkout.  If that
> > repository was properly repacked on the server, the clone data transfer
> > would be 283 MB.  This is less than half the current clone transfer
> > size.
> >
> >
> > Nicolas
> >
> 
> 'git gc --prune=now' does work, but 'git gc --prune=now --aggressive'
> (before) and 'git gc --aggressive' (after) both create very large
> (>2GB; I stopped it) packs from the ~400MB-600MB packed objects. I
> noted that you specifically wrote 'with a later git version' -
> presumably there is a some sort of a known and fixed issue there? Just
> curious.
>From git v1.6.3 the --aggressive switch makes for 'git repack' to be 
called with --window=250 --depth=250, meaning the equivalent of:
	git repack -a -d -f --window=250 --depth=250
Do you still get a huge pack with the above?
> I guess --aggressive doesn't always save space...
If so that is (and was) a bug.
Nicolas
Hin-Tak Leung· Aug 12, 2009, 14:45 UTC · re: Nicolas Pitre · lore

Re: git gc expanding packed data?

On Tue, Aug 11, 2009 at 10:33 PM, Nicolas Pitre<nico@cam.org> wrote: <snipped>

Show 10 quoted lines
> From git v1.6.3 the --aggressive switch makes for 'git repack' to be
> called with --window=250 --depth=250, meaning the equivalent of:
>
>        git repack -a -d -f --window=250 --depth=250
>
> Do you still get a huge pack with the above?
>
>> I guess --aggressive doesn't always save space...
>
> If so that is (and was) a bug.

I tried 'git repack -a -d -f --window=250 --depth=250' with 1.6.2.5 (fc11.x86_64) and it took half a day, swallowed up all the memory - 3GB virtual & 1.3GB resident - and finally the kernel oom killer killed it at a last message of (601460/957910). Left no temp files. Would git 1.6.3 use less memory? :-(

Hin-Tak
Nicolas Pitre· Aug 12, 2009, 15:35 UTC · re: Hin-Tak Leung · lore

Re: git gc expanding packed data?

On Wed, 12 Aug 2009, Hin-Tak Leung wrote:
Show 19 quoted lines
> On Tue, Aug 11, 2009 at 10:33 PM, Nicolas Pitre<nico@cam.org> wrote:
> <snipped>
> 
> > From git v1.6.3 the --aggressive switch makes for 'git repack' to be
> > called with --window=250 --depth=250, meaning the equivalent of:
> >
> >        git repack -a -d -f --window=250 --depth=250
> >
> > Do you still get a huge pack with the above?
> >
> >> I guess --aggressive doesn't always save space...
> >
> > If so that is (and was) a bug.
> 
> I tried 'git repack -a -d -f --window=250 --depth=250' with 1.6.2.5
> (fc11.x86_64) and it took half a day, swallowed up all the memory -
> 3GB virtual & 1.3GB resident - and finally the kernel oom killer
> killed it at a last message of (601460/957910). Left no temp files.
> Would git 1.6.3 use less memory? :-(
Probably not.  However you should try:
	git config pack.deltaCacheSize 1

That limits the delta cache size to one byte (effectively disabling it) instead of the default of 0 which means unlimited. With that I'm able to repack that repository using the above git repack command on an x86-64 system with 4GB of RAM and using 4 threads (this is a quad core). Resident memory usage grows to nearly 3.3GB though.

If your machine is SMP and you don't have sufficient RAM then you can reduce the number of threads to only one:

	git config pack.threads 1

Additionally, you can further limit memory usage with the --window-memory argument to 'git repack'. For example, using --window-memory=128M should keep a reasonable upper bound on the delta search memory usage although this can result in less optimal delta match if the repo contains lots of large files (and I think this is the case for the gcc repo).

Nicolas
Hin-Tak Leung· Aug 13, 2009, 17:31 UTC · re: Nicolas Pitre · lore

Re: git gc expanding packed data?

On Wed, Aug 12, 2009 at 4:35 PM, Nicolas Pitre<nico@cam.org> wrote:
Show 47 quoted lines
> On Wed, 12 Aug 2009, Hin-Tak Leung wrote:
>
>> On Tue, Aug 11, 2009 at 10:33 PM, Nicolas Pitre<nico@cam.org> wrote:
>> <snipped>
>>
>> > From git v1.6.3 the --aggressive switch makes for 'git repack' to be
>> > called with --window=250 --depth=250, meaning the equivalent of:
>> >
>> >        git repack -a -d -f --window=250 --depth=250
>> >
>> > Do you still get a huge pack with the above?
>> >
>> >> I guess --aggressive doesn't always save space...
>> >
>> > If so that is (and was) a bug.
>>
>> I tried 'git repack -a -d -f --window=250 --depth=250' with 1.6.2.5
>> (fc11.x86_64) and it took half a day, swallowed up all the memory -
>> 3GB virtual & 1.3GB resident - and finally the kernel oom killer
>> killed it at a last message of (601460/957910). Left no temp files.
>> Would git 1.6.3 use less memory? :-(
>
> Probably not.  However you should try:
>
>        git config pack.deltaCacheSize 1
>
> That limits the delta cache size to one byte (effectively disabling it)
> instead of the default of 0 which means unlimited.  With that I'm able
> to repack that repository using the above git repack command on an
> x86-64 system with 4GB of RAM and using 4 threads (this is a quad core).
> Resident memory usage grows to nearly 3.3GB though.
>
> If your machine is SMP and you don't have sufficient RAM then you can
> reduce the number of threads to only one:
>
>        git config pack.threads 1
>
> Additionally, you can further limit memory usage with the
> --window-memory argument to 'git repack'.  For example, using
> --window-memory=128M should keep a reasonable upper bound on the delta
> search memory usage although this can result in less optimal delta match
> if the repo contains lots of large files (and I think this is the case
> for the gcc repo).
>
>
> Nicolas
>
Thanks.  I used the two git config pack.* commands, and
   git repack -a -d -f --window=250 --depth=250
finished after 8 hours (dual core Turion, 2GB RAM + 2GB swap). The
pack directory went from 457MB to 308MB.

Thanks a lot for the advice - learned a few interesting things about git on the way :-).

Hin-Tak

← back to recent threads