threads / discuss / 28342

git repository size / compression

Subject: git repository size / compression

## tl;dr

10 messages between Sep 9, 2011 and Sep 9, 2011.

replies: 9people: 6as markdown or json

neubyr· Sep 9, 2011, 02:37 UTC · lore

I have a test git repository with just two files in it. One of the file in it has a set of two lines that is repeated n times. e.g.: {{{ $ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt && cat ./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt; done }}}

I ran above command few times and performed commit after each run. Now disk usage of this repository directory is mentioned below. The 419M is working directory size and 2.7M is git repository/database size.

{{{ $ du -h -d 1 . 2.7M ./.git 419M .

}}}

Is it because of the compression performed by git before storing data (or before sending commit)??

Following were results with subversion:

Subversion client (redundant(?) copy exists in .svn/text-base/ directory, hence double size in client): {{{ $ du -h -d 1 416M ./.svn 832M . }}}

Subversion repo/server:
{{{
$ du -h -d 1
 12K    ./conf
1.2M    ./db
 36K    ./hooks
8.0K    ./locks
1.2M    .
}}}

-- neuby.r

Carlos Martín Nieto· Sep 9, 2011, 08:23 UTC · re: neubyr · lore

Re: git repository size / compression

On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:
Show 8 quoted lines
> I have a test git repository with just two files in it. One of the
> file in it has a set of two lines that is repeated n times.
> e.g.:
> {{{
> $ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt &&  cat
> ./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done
> }}}
> 

So you've just created some data that can be compressed quite efficiently.

Show 14 quoted lines
> I ran above command few times and performed commit after each run. Now
> disk usage of this repository directory is mentioned below. The 419M
> is working directory size and 2.7M is git repository/database size.
> 
> {{{
> $ du -h -d 1 .
> 2.7M    ./.git
> 419M    .
> 
> }}}
> 
> Is it because of the compression performed by git before storing data
> (or before sending commit)??
> 

Yes. Git stores its objects (the commit, the snapshot of the files, etc.) compressed. When these objects are stored in a pack, the size can be further reduced by storing some objects as deltas which describe the difference between itself and some other object in the object-db.

Show 9 quoted lines
> Following were results with subversion:
> 
> Subversion client (redundant(?) copy exists in .svn/text-base/
> directory, hence double size in client):
> {{{
> $ du -h -d 1
> 416M    ./.svn
> 832M    .
> }}}

Subversion stores the "pristines" (which is the status of the files in the latest revision) inside the .svn directory. I wouldn't call this copy redundant, though, as it allows you to run diff locally. The pristines are stored uncompressed, which is why you half of the space is taken up by the .svn directory.

Show 10 quoted lines
> 
> Subversion repo/server:
> {{{
> $ du -h -d 1
>  12K    ./conf
> 1.2M    ./db
>  36K    ./hooks
> 8.0K    ./locks
> 1.2M    .
> }}}

I don't know how the repository is stored in Subversion, but it may also be compressed. You may be able to reduced your git repository size by (re)generating packs with 'git repack' and doing some cleanups with 'git gc', but the repository size is not often a concern.

   cmn
neubyr· Sep 9, 2011, 14:04 UTC · re: Carlos Martín Nieto · lore

Re: git repository size / compression

On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:
Show 33 quoted lines
> On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:
>> I have a test git repository with just two files in it. One of the
>> file in it has a set of two lines that is repeated n times.
>> e.g.:
>> {{{
>> $ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt &&  cat
>> ./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done
>> }}}
>>
>
> So you've just created some data that can be compressed quite
> efficiently.
>
>> I ran above command few times and performed commit after each run. Now
>> disk usage of this repository directory is mentioned below. The 419M
>> is working directory size and 2.7M is git repository/database size.
>>
>> {{{
>> $ du -h -d 1 .
>> 2.7M    ./.git
>> 419M    .
>>
>> }}}
>>
>> Is it because of the compression performed by git before storing data
>> (or before sending commit)??
>>
>
> Yes. Git stores its objects (the commit, the snapshot of the files,
> etc.) compressed. When these objects are stored in a pack, the size can
> be further reduced by storing some objects as deltas which describe the
> difference between itself and some other object in the object-db.
>

Does git store deltas for some files? I thought it uses snapshots (exact copy of staged files) only.

Show 36 quoted lines
>> Following were results with subversion:
>>
>> Subversion client (redundant(?) copy exists in .svn/text-base/
>> directory, hence double size in client):
>> {{{
>> $ du -h -d 1
>> 416M    ./.svn
>> 832M    .
>> }}}
>
> Subversion stores the "pristines" (which is the status of the files in
> the latest revision) inside the .svn directory. I wouldn't call this
> copy redundant, though, as it allows you to run diff locally. The
> pristines are stored uncompressed, which is why you half of the space is
> taken up by the .svn directory.
>
>>
>> Subversion repo/server:
>> {{{
>> $ du -h -d 1
>>  12K    ./conf
>> 1.2M    ./db
>>  36K    ./hooks
>> 8.0K    ./locks
>> 1.2M    .
>> }}}
>
> I don't know how the repository is stored in Subversion, but it may also
> be compressed. You may be able to reduced your git repository size by
> (re)generating packs with 'git repack' and doing some cleanups with 'git
> gc', but the repository size is not often a concern.
>
>   cmn
>
>
>
that's helpful. thanks.

-- neuby.r

Sverre Rabbelier· Sep 9, 2011, 14:25 UTC · re: neubyr · lore

Re: git repository size / compression

Heya,
On Fri, Sep 9, 2011 at 16:04, neubyr <neubyr@gmail.com> wrote:
> Does git store deltas for some files? I thought it uses snapshots
> (exact copy of staged files) only.
In packs, yes, it will try to delta objects as efficient as possible.
-- 
Cheers,

Sverre Rabbelier
Carlos Martín Nieto· Sep 9, 2011, 14:28 UTC · re: neubyr · lore

Re: git repository size / compression

On Fri, 2011-09-09 at 09:04 -0500, neubyr wrote:
Show 37 quoted lines
> On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:
> > On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:
> >> I have a test git repository with just two files in it. One of the
> >> file in it has a set of two lines that is repeated n times.
> >> e.g.:
> >> {{{
> >> $ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt &&  cat
> >> ./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done
> >> }}}
> >>
> >
> > So you've just created some data that can be compressed quite
> > efficiently.
> >
> >> I ran above command few times and performed commit after each run. Now
> >> disk usage of this repository directory is mentioned below. The 419M
> >> is working directory size and 2.7M is git repository/database size.
> >>
> >> {{{
> >> $ du -h -d 1 .
> >> 2.7M    ./.git
> >> 419M    .
> >>
> >> }}}
> >>
> >> Is it because of the compression performed by git before storing data
> >> (or before sending commit)??
> >>
> >
> > Yes. Git stores its objects (the commit, the snapshot of the files,
> > etc.) compressed. When these objects are stored in a pack, the size can
> > be further reduced by storing some objects as deltas which describe the
> > difference between itself and some other object in the object-db.
> >
> 
> Does git store deltas for some files? I thought it uses snapshots
> (exact copy of staged files) only.

Yes and no. The data model for git is to always store snapshots, and it always expects to have the full files available. In a packfile, however, in order to save space, some objects are stored as deltas to other objects in the same file.

http://progit.org/book/ch9-4.html
Show 44 quoted lines
> 
> 
> >> Following were results with subversion:
> >>
> >> Subversion client (redundant(?) copy exists in .svn/text-base/
> >> directory, hence double size in client):
> >> {{{
> >> $ du -h -d 1
> >> 416M    ./.svn
> >> 832M    .
> >> }}}
> >
> > Subversion stores the "pristines" (which is the status of the files in
> > the latest revision) inside the .svn directory. I wouldn't call this
> > copy redundant, though, as it allows you to run diff locally. The
> > pristines are stored uncompressed, which is why you half of the space is
> > taken up by the .svn directory.
> >
> >>
> >> Subversion repo/server:
> >> {{{
> >> $ du -h -d 1
> >>  12K    ./conf
> >> 1.2M    ./db
> >>  36K    ./hooks
> >> 8.0K    ./locks
> >> 1.2M    .
> >> }}}
> >
> > I don't know how the repository is stored in Subversion, but it may also
> > be compressed. You may be able to reduced your git repository size by
> > (re)generating packs with 'git repack' and doing some cleanups with 'git
> > gc', but the repository size is not often a concern.
> >
> >   cmn
> >
> >
> >
> 
> that's helpful. thanks.
> 
> --
> neuby.r
> 
neubyr· Sep 9, 2011, 15:07 UTC · re: Carlos Martín Nieto · lore

Re: git repository size / compression

On Fri, Sep 9, 2011 at 9:28 AM, Carlos Martín Nieto <cmn@elego.de> wrote:
Show 46 quoted lines
> On Fri, 2011-09-09 at 09:04 -0500, neubyr wrote:
>> On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:
>> > On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:
>> >> I have a test git repository with just two files in it. One of the
>> >> file in it has a set of two lines that is repeated n times.
>> >> e.g.:
>> >> {{{
>> >> $ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt &&  cat
>> >> ./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done
>> >> }}}
>> >>
>> >
>> > So you've just created some data that can be compressed quite
>> > efficiently.
>> >
>> >> I ran above command few times and performed commit after each run. Now
>> >> disk usage of this repository directory is mentioned below. The 419M
>> >> is working directory size and 2.7M is git repository/database size.
>> >>
>> >> {{{
>> >> $ du -h -d 1 .
>> >> 2.7M    ./.git
>> >> 419M    .
>> >>
>> >> }}}
>> >>
>> >> Is it because of the compression performed by git before storing data
>> >> (or before sending commit)??
>> >>
>> >
>> > Yes. Git stores its objects (the commit, the snapshot of the files,
>> > etc.) compressed. When these objects are stored in a pack, the size can
>> > be further reduced by storing some objects as deltas which describe the
>> > difference between itself and some other object in the object-db.
>> >
>>
>> Does git store deltas for some files? I thought it uses snapshots
>> (exact copy of staged files) only.
>
> Yes and no. The data model for git is to always store snapshots, and it
> always expects to have the full files available. In a packfile, however,
> in order to save space, some objects are stored as deltas to other
> objects in the same file.
>
> http://progit.org/book/ch9-4.html
>
Excellent.. That explains compression and deltas really well. Thanks again..

-- neuby.r

Jakub Narebski· Sep 9, 2011, 14:54 UTC · re: neubyr · lore

Re: git repository size / compression

neubyr <neubyr@gmail.com> writes:
> On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:
> > On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:
Show 22 quoted lines
>>> I have a test git repository with just two files in it. One of the
>>> file in it has a set of two lines that is repeated n times.
>>> e.g.:
>>> {{{
>>> $ for i in {1..5}; do cat ./lexico.txt>> lexico1.txt &&  cat
>>> ./lexico.txt>> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done
>>> }}}
>>>
>>
>> So you've just created some data that can be compressed quite
>> efficiently.
>>
>>> I ran above command few times and performed commit after each run. Now
>>> disk usage of this repository directory is mentioned below. The 419M
>>> is working directory size and 2.7M is git repository/database size.
>>>
>>> {{{
>>> $ du -h -d 1 .
>>> 2.7M    ./.git
>>> 419M    .
>>>
>>> }}}
Have you tried the same but with
   $ git gc --prune=now
before running `du`?
Show 10 quoted lines
>>> Is it because of the compression performed by git before storing data
>>> (or before sending commit)??
>>
>> Yes. Git stores its objects (the commit, the snapshot of the files,
>> etc.) compressed. When these objects are stored in a pack, the size can
>> be further reduced by storing some objects as deltas which describe the
>> difference between itself and some other object in the object-db.
> 
> Does git store deltas for some files? I thought it uses snapshots
> (exact copy of staged files) only.

When creating packfile from loose objects (e.g. via `git gc`), it does perform delta compression.

-- 
Jakub Narębski
neubyr· Sep 9, 2011, 15:09 UTC · re: Jakub Narebski · lore

Re: git repository size / compression

2011/9/9 Jakub Narebski <jnareb@gmail.com>:
Show 33 quoted lines
> neubyr <neubyr@gmail.com> writes:
>> On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:
>> > On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:
>
>>>> I have a test git repository with just two files in it. One of the
>>>> file in it has a set of two lines that is repeated n times.
>>>> e.g.:
>>>> {{{
>>>> $ for i in {1..5}; do cat ./lexico.txt>> lexico1.txt &&  cat
>>>> ./lexico.txt>> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done
>>>> }}}
>>>>
>>>
>>> So you've just created some data that can be compressed quite
>>> efficiently.
>>>
>>>> I ran above command few times and performed commit after each run. Now
>>>> disk usage of this repository directory is mentioned below. The 419M
>>>> is working directory size and 2.7M is git repository/database size.
>>>>
>>>> {{{
>>>> $ du -h -d 1 .
>>>> 2.7M    ./.git
>>>> 419M    .
>>>>
>>>> }}}
>
> Have you tried the same but with
>
>   $ git gc --prune=now
>
> before running `du`?
>

Nope, I hadn't run git gc before. Here are du results after running git gc command. That's about 55% less space now.. Great!

{{{ $ du -d 1 -h 924K ./.git 417M . }}}

Show 17 quoted lines
>>>> Is it because of the compression performed by git before storing data
>>>> (or before sending commit)??
>>>
>>> Yes. Git stores its objects (the commit, the snapshot of the files,
>>> etc.) compressed. When these objects are stored in a pack, the size can
>>> be further reduced by storing some objects as deltas which describe the
>>> difference between itself and some other object in the object-db.
>>
>> Does git store deltas for some files? I thought it uses snapshots
>> (exact copy of staged files) only.
>
> When creating packfile from loose objects (e.g. via `git gc`), it
> does perform delta compression.
>
> --
> Jakub Narębski
>
thank you everyone for explaining in detail..

-- neuby.r

John Szakmeister· Sep 9, 2011, 16:05 UTC · re: Carlos Martín Nieto · lore

Re: git repository size / compression

On Fri, Sep 9, 2011 at 4:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote: [snip]

Show 14 quoted lines
>> Subversion repo/server:
>> {{{
>> $ du -h -d 1
>>  12K    ./conf
>> 1.2M    ./db
>>  36K    ./hooks
>> 8.0K    ./locks
>> 1.2M    .
>> }}}
>
> I don't know how the repository is stored in Subversion, but it may also
> be compressed. You may be able to reduced your git repository size by
> (re)generating packs with 'git repack' and doing some cleanups with 'git
> gc', but the repository size is not often a concern.

It is stored compressed in Subversion, and it also generates deltas against previous versions. IIRC, the delta algorithm in an xdelta based one, and then the data is run through compression. Subversion will at times choose to self-compress the file, instead of doing a delta and compressing. IIRC, there is some heuristics in there for determining when to do that, but I forget the exact method.

HTH!
-John
Andreas Krey· Sep 9, 2011, 17:49 UTC · re: John Szakmeister · lore

Re: git repository size / compression

On Fri, 09 Sep 2011 12:05:03 +0000, John Szakmeister wrote: ...

> will at times choose to self-compress the file, instead of doing a
> delta and compressing.  IIRC, there is some heuristics in there for
> determining when to do that, but I forget the exact method.

Don't know about the compression part, but subversion does a delta of the nth version of a file (not the global revision number n) against the version m, where m is (n & (n-1)), or the least significant '1' bit flipped to '0'. That way, there are only O(log(n)) instead of O(n) deltas to apply to get at a specific version.

[Was on the svn users list just then. They described it differently,
 but in essence it's that.]
Andreas
-- 
"Totally trivial. Famous last words."
From: Linus Torvalds <torvalds@*.org>
Date: Fri, 22 Jan 2010 07:29:21 -0800

← back to recent threads