# why is my repo size ten times bigger than what I expected?

10 messages from 2011-03-05 to 2011-03-09. Participants: Ruben Laguna, Pascal Obry, Jonathan del Strother, Andreas Schwab, Phillip Susi, Tor Arvid Lund.
Thread: https://gitlist.dev/t/26664

## Ruben Laguna, 2011-03-05 10:05

Subject: why is my repo size ten times bigger than what I expected?
Message-ID: <AANLkTi=V3nEamocowbHovvV0U69nZgD70fysu1CQOwrR@mail.gmail.com>
URL: https://gitlist.dev/e/AANLkTi%3DV3nEamocowbHovvV0U69nZgD70fysu1CQOwrR%40mail.gmail.com
In-Reply-To: <AANLkTimi+OnpdX+Y7jx1JaOmGbZc_XEgJFeK0PKLpu2o@mail.gmail.com>

```
Hi,

I had a repo which was big 143MB because it contained a bunch of jar
files. So I decided to remove those completely from the history.

In short I used the git-large-blob [1] to find all the jars and used
the git-remove-history script [2] which does the filter-branch thing,
prune, etc.

I did this on all branches (that I know of) and now I can see that the
jars are gone because I can't find them with git-large-blob.  and the
repo size has dropped from 143Mb to 87Mb.

My concern is that 87Mb is still really big taking into account he
size of the project.  in fact if I run "git diff-tree -r -p $commit
|wc -c" for each commit and sum all I get 5.5Mb.


I also ran the git-rev-size [3] script that I found in this mailing
list and I only see that the size grows steadly from commit to commit
up to 1482731 bytes. So again how come the .git directory is 87MB?


So, Can anybody tell me if this repository size is "normal" for a
project with 1.4MB source and 352 commits?
Is there a better way to calculate the size (in bytes) of each commit?

Is there any other thing I could do to reduce and audit  the repository size?


Thanks in advance!
Rubén

---
[1] http://stackoverflow.com/questions/298314/find-files-in-git-repo-over-x-megabytes-that-dont-exist-in-head
[2] http://dound.com/2009/04/git-forever-remove-files-or-folders-from-history/
[3] http://markmail.org/message/762zzg5zckbiq2i7

```

## Pascal Obry, 2011-03-05 10:49

Subject: Re: why is my repo size ten times bigger than what I expected?
Message-ID: <4D721552.4090205@obry.net>
URL: https://gitlist.dev/e/4D721552.4090205%40obry.net
In-Reply-To: <AANLkTi=V3nEamocowbHovvV0U69nZgD70fysu1CQOwrR@mail.gmail.com>

```
Le 05/03/2011 11:05, Ruben Laguna a écrit :
> Is there any other thing I could do to reduce and audit  the repository size?

    $ git gc

?

-- 

--|------------------------------------------------------
--| Pascal Obry                           Team-Ada Member
--| 45, rue Gabriel Peri - 78114 Magny Les Hameaux FRANCE
--|------------------------------------------------------
--|    http://www.obry.net  -  http://v2p.fr.eu.org
--| "The best way to travel is by means of imagination"
--|
--| gpg --keyserver keys.gnupg.net --recv-key F949BD3B

```

## Jonathan del Strother, 2011-03-05 10:49

Subject: Re: why is my repo size ten times bigger than what I expected?
Message-ID: <AANLkTimp8B5Lv15qhGOwOzh+kqOS0g3Xwvgib8vyk+m+@mail.gmail.com>
URL: https://gitlist.dev/e/AANLkTimp8B5Lv15qhGOwOzh%2BkqOS0g3Xwvgib8vyk%2Bm%2B%40mail.gmail.com
In-Reply-To: <AANLkTi=V3nEamocowbHovvV0U69nZgD70fysu1CQOwrR@mail.gmail.com>

```
On 5 March 2011 10:05, Ruben Laguna <ruben.laguna@gmail.com> wrote:
> Hi,
>
> I had a repo which was big 143MB because it contained a bunch of jar
> files. So I decided to remove those completely from the history.
>
> In short I used the git-large-blob [1] to find all the jars and used
> the git-remove-history script [2] which does the filter-branch thing,
> prune, etc.
>
> I did this on all branches (that I know of) and now I can see that the
> jars are gone because I can't find them with git-large-blob.  and the
> repo size has dropped from 143Mb to 87Mb.
>
> My concern is that 87Mb is still really big taking into account he
> size of the project.  in fact if I run "git diff-tree -r -p $commit
> |wc -c" for each commit and sum all I get 5.5Mb.
>
>
> I also ran the git-rev-size [3] script that I found in this mailing
> list and I only see that the size grows steadly from commit to commit
> up to 1482731 bytes. So again how come the .git directory is 87MB?
>
>
> So, Can anybody tell me if this repository size is "normal" for a
> project with 1.4MB source and 352 commits?
> Is there a better way to calculate the size (in bytes) of each commit?
>
> Is there any other thing I could do to reduce and audit  the repository size?
>
>
> Thanks in advance!
> Rubén
>
> ---
> [1] http://stackoverflow.com/questions/298314/find-files-in-git-repo-over-x-megabytes-that-dont-exist-in-head
> [2] http://dound.com/2009/04/git-forever-remove-files-or-folders-from-history/
> [3] http://markmail.org/message/762zzg5zckbiq2i7

What happens if you clone that repo?
git-gc will only pruned unused objects that're older than 2 weeks by
default, so it's possible that your repo size will suddenly shrink in
2 weeks time (or sooner, if you run git-gc with the appropriate
options)

```

## Ruben Laguna, 2011-03-05 11:41

Subject: Re: why is my repo size ten times bigger than what I expected?
Message-ID: <AANLkTinwHMULqPZSguYtJztuA4Oy6-s6Ah3_tcVVO7D9@mail.gmail.com>
URL: https://gitlist.dev/e/AANLkTinwHMULqPZSguYtJztuA4Oy6-s6Ah3_tcVVO7D9%40mail.gmail.com
In-Reply-To: <AANLkTimp8B5Lv15qhGOwOzh+kqOS0g3Xwvgib8vyk+m+@mail.gmail.com>

```
well, the git-remove-history script does

rm -rf .git/refs/original/
git reflog expire --expire=now --all
git fsck --unreachable
git gc --prune=now
git gc --aggressive --prune=now


after filter-branch so I don't think it's that.

also cloning the repo doesn't change a thing

$ git clone en4j en4j_xx
Cloning into en4j_xx...
done.
$ cd en4j_xx
$ du -sh .git
 87M    .git

any other idea?

On Sat, Mar 5, 2011 at 11:49 AM, Jonathan del Strother
<maillist@steelskies.com> wrote:
> On 5 March 2011 10:05, Ruben Laguna <ruben.laguna@gmail.com> wrote:
>> Hi,
>>
>> I had a repo which was big 143MB because it contained a bunch of jar
>> files. So I decided to remove those completely from the history.
>>
>> In short I used the git-large-blob [1] to find all the jars and used
>> the git-remove-history script [2] which does the filter-branch thing,
>> prune, etc.
>>
>> I did this on all branches (that I know of) and now I can see that the
>> jars are gone because I can't find them with git-large-blob.  and the
>> repo size has dropped from 143Mb to 87Mb.
>>
>> My concern is that 87Mb is still really big taking into account he
>> size of the project.  in fact if I run "git diff-tree -r -p $commit
>> |wc -c" for each commit and sum all I get 5.5Mb.
>>
>>
>> I also ran the git-rev-size [3] script that I found in this mailing
>> list and I only see that the size grows steadly from commit to commit
>> up to 1482731 bytes. So again how come the .git directory is 87MB?
>>
>>
>> So, Can anybody tell me if this repository size is "normal" for a
>> project with 1.4MB source and 352 commits?
>> Is there a better way to calculate the size (in bytes) of each commit?
>>
>> Is there any other thing I could do to reduce and audit  the repository size?
>>
>>
>> Thanks in advance!
>> Rubén
>>
>> ---
>> [1] http://stackoverflow.com/questions/298314/find-files-in-git-repo-over-x-megabytes-that-dont-exist-in-head
>> [2] http://dound.com/2009/04/git-forever-remove-files-or-folders-from-history/
>> [3] http://markmail.org/message/762zzg5zckbiq2i7
>
> What happens if you clone that repo?
> git-gc will only pruned unused objects that're older than 2 weeks by
> default, so it's possible that your repo size will suddenly shrink in
> 2 weeks time (or sooner, if you run git-gc with the appropriate
> options)
>



-- 
/Rubén

```

## Andreas Schwab, 2011-03-05 12:57

Subject: Re: why is my repo size ten times bigger than what I expected?
Message-ID: <m2zkp9wwqe.fsf@igel.home>
URL: https://gitlist.dev/e/m2zkp9wwqe.fsf%40igel.home
In-Reply-To: <AANLkTinwHMULqPZSguYtJztuA4Oy6-s6Ah3_tcVVO7D9@mail.gmail.com>

```
Ruben Laguna <ruben.laguna@gmail.com> writes:

> also cloning the repo doesn't change a thing
>
> $ git clone en4j en4j_xx
> Cloning into en4j_xx...
> done.
> $ cd en4j_xx
> $ du -sh .git
>  87M    .git
>
> any other idea?

Please use file://$PWD/en4j as URL, otherwise git clone just hard links
everything.

Andreas.

-- 
Andreas Schwab, schwab@linux-m68k.org
GPG Key fingerprint = 58CA 54C7 6D53 942B 1756  01D3 44D5 214B 8276 4ED5
"And now for something completely different."

```

## Ruben Laguna, 2011-03-07 09:59

Subject: Re: why is my repo size ten times bigger than what I expected?
Message-ID: <AANLkTinJZFGVNSQVfCipo33h5uPpK0pFY10E203oTfhU@mail.gmail.com>
URL: https://gitlist.dev/e/AANLkTinJZFGVNSQVfCipo33h5uPpK0pFY10E203oTfhU%40mail.gmail.com
In-Reply-To: <m2zkp9wwqe.fsf@igel.home>

```
Cloning it that way didn't help either,

But I have more info

If I set a bare repo and push my four branches to it (master, develop,
gh-pages and experimental) the total size of the repo is 2.4MB
(instead of 87MB)


$ git init --bare en4j_xx
$ cd en4j
$ git checkout master
$ git push file://$PWD/../en4j_xx master
$ git checkout develop
$ git push file://$PWD/../en4j_xx develop
$ git checkout experimental
$ git push file://$PWD/../en4j_xx experimental
$ git checkout gh-pages
$ git push file://$PWD/../en4j_xx gh-pages
$ $ du -sh ../en4j_xx
2.3M	../en4j_xx

So, how can I find the contents present in en4j that are not present in en4j_xx?






On Sat, Mar 5, 2011 at 1:57 PM, Andreas Schwab <schwab@linux-m68k.org> wrote:
> Ruben Laguna <ruben.laguna@gmail.com> writes:
>
>> also cloning the repo doesn't change a thing
>>
>> $ git clone en4j en4j_xx
>> Cloning into en4j_xx...
>> done.
>> $ cd en4j_xx
>> $ du -sh .git
>>  87M    .git
>>
>> any other idea?
>
> Please use file://$PWD/en4j as URL, otherwise git clone just hard links
> everything.
>
> Andreas.
>
> --
> Andreas Schwab, schwab@linux-m68k.org
> GPG Key fingerprint = 58CA 54C7 6D53 942B 1756  01D3 44D5 214B 8276 4ED5
> "And now for something completely different."
>



-- 
/Rubén

```

## Phillip Susi, 2011-03-08 21:25

Subject: Re: why is my repo size ten times bigger than what I expected?
Message-ID: <4D769EBE.40008@cfl.rr.com>
URL: https://gitlist.dev/e/4D769EBE.40008%40cfl.rr.com
In-Reply-To: <AANLkTinJZFGVNSQVfCipo33h5uPpK0pFY10E203oTfhU@mail.gmail.com>

```
On 3/7/2011 4:59 AM, Ruben Laguna wrote:
> Cloning it that way didn't help either,
> 
> But I have more info
> 
> If I set a bare repo and push my four branches to it (master, develop,
> gh-pages and experimental) the total size of the repo is 2.4MB
> (instead of 87MB)

Then there are other branches using that space.  Run git branch -a and
see what else is there.

```

## Tor Arvid Lund, 2011-03-08 21:44

Subject: Re: why is my repo size ten times bigger than what I expected?
Message-ID: <AANLkTikuuzHZ897kOY2u0Sv=0JTDffo0UhcxkyynVQAZ@mail.gmail.com>
URL: https://gitlist.dev/e/AANLkTikuuzHZ897kOY2u0Sv%3D0JTDffo0UhcxkyynVQAZ%40mail.gmail.com
In-Reply-To: <AANLkTi=V3nEamocowbHovvV0U69nZgD70fysu1CQOwrR@mail.gmail.com>

```
On Sat, Mar 5, 2011 at 11:05 AM, Ruben Laguna <ruben.laguna@gmail.com> wrote:
> Hi,
>
> I had a repo which was big 143MB because it contained a bunch of jar
> files. So I decided to remove those completely from the history.
>
> In short I used the git-large-blob [1] to find all the jars and used
> the git-remove-history script [2] which does the filter-branch thing,
> prune, etc.
>
> I did this on all branches (that I know of) and now I can see that the
> jars are gone because I can't find them with git-large-blob.  and the
> repo size has dropped from 143Mb to 87Mb.

I just thought I'd mention that the git-remove-history script that you
mention does filter-branch on HEAD, and not using the --all parameter.
I thought --all was the best way to "catch all" branches in one go...

    -- Tor Arvid

> My concern is that 87Mb is still really big taking into account he
> size of the project.  in fact if I run "git diff-tree -r -p $commit
> |wc -c" for each commit and sum all I get 5.5Mb.
>
>
> I also ran the git-rev-size [3] script that I found in this mailing
> list and I only see that the size grows steadly from commit to commit
> up to 1482731 bytes. So again how come the .git directory is 87MB?
>
>
> So, Can anybody tell me if this repository size is "normal" for a
> project with 1.4MB source and 352 commits?
> Is there a better way to calculate the size (in bytes) of each commit?
>
> Is there any other thing I could do to reduce and audit  the repository size?
>
>
> Thanks in advance!
> Rubén
>
> ---
> [1] http://stackoverflow.com/questions/298314/find-files-in-git-repo-over-x-megabytes-that-dont-exist-in-head
> [2] http://dound.com/2009/04/git-forever-remove-files-or-folders-from-history/
> [3] http://markmail.org/message/762zzg5zckbiq2i7
> --
> To unsubscribe from this list: send the line "unsubscribe git" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
>

```

## Ruben Laguna, 2011-03-09 16:35

Subject: Re: why is my repo size ten times bigger than what I expected?
Message-ID: <AANLkTikxPPo4pGqrvo4rdDUOwp1PYYEdcERVfXfLnsVh@mail.gmail.com>
URL: https://gitlist.dev/e/AANLkTikxPPo4pGqrvo4rdDUOwp1PYYEdcERVfXfLnsVh%40mail.gmail.com
In-Reply-To: <AANLkTikuuzHZ897kOY2u0Sv=0JTDffo0UhcxkyynVQAZ@mail.gmail.com>

```
> I just thought I'd mention that the git-remove-history script that you
> mention does filter-branch on HEAD, and not using the --all parameter.
> I thought --all was the best way to "catch all" branches in one go...
>
>    -- Tor Arvid
>

Much faster this way, thanks Tor,

But it still gives the same result 88MB


$ git branch -a
* develop
  master
  remotes/origin/HEAD -> origin/develop
  remotes/origin/develop
  remotes/origin/experimental
  remotes/origin/gh-pages
  remotes/origin/master

Finally I have deleted my public repo on github, created a new one and
pushed master and develop to the new empty one.


-- 
/Rubén

```

## Tor Arvid Lund, 2011-03-09 21:06

Subject: Re: why is my repo size ten times bigger than what I expected?
Message-ID: <AANLkTimYE8qCz98u3r2HRmeCx7k0wd_cWp-tR+tSoKbD@mail.gmail.com>
URL: https://gitlist.dev/e/AANLkTimYE8qCz98u3r2HRmeCx7k0wd_cWp-tR%2BtSoKbD%40mail.gmail.com
In-Reply-To: <AANLkTikxPPo4pGqrvo4rdDUOwp1PYYEdcERVfXfLnsVh@mail.gmail.com>

```
On Wed, Mar 9, 2011 at 5:35 PM, Ruben Laguna <ruben.laguna@gmail.com> wrote:
>> I just thought I'd mention that the git-remove-history script that you
>> mention does filter-branch on HEAD, and not using the --all parameter.
>> I thought --all was the best way to "catch all" branches in one go...
>>
>>    -- Tor Arvid
>>
>
> Much faster this way, thanks Tor,
>
> But it still gives the same result 88MB
>
>
> $ git branch -a
> * develop
>  master
>  remotes/origin/HEAD -> origin/develop
>  remotes/origin/develop
>  remotes/origin/experimental
>  remotes/origin/gh-pages
>  remotes/origin/master
>
> Finally I have deleted my public repo on github, created a new one and
> pushed master and develop to the new empty one.

Ah, that's why I got only 3.6M when i cloned just now ;)

FWIW (if you still want to figure it out...) - Whatever refs that your
origin branches point to - their history and objects will *not* get
deleted by git gc/prune/whatever. So if they point to commits which
have these big jars in the history, that may be the cause. Also, when
I do filter-branch, it saves the old refs in .git/refs/original so
that I can revert it all those times when I screw it up ;)

Basically - since your "new" repo is so small, there is something in
your original repo that refers to your large objects.

Have a good night.

    -- Tor Arvid

```
