{"thread":{"id":"28342","subject":"git repository size / compression","startedAt":"2011-09-09T02:37:42Z","lastAt":"2011-09-09T17:49:37Z","messageCount":10,"participants":["neubyr","Carlos Martín Nieto","Sverre Rabbelier","Jakub Narebski","John Szakmeister","Andreas Krey"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"175147","messageId":"CALFxCvzVjC+u=RDkDCQp0QqPETsv8ROE8tY=37tmMWxmQoJOEw@mail.gmail.com","threadId":"28342","inReplyTo":null,"subject":"git repository size / compression","fromName":"neubyr","fromEmail":"neubyr@gmail.com","sentAt":"2011-09-09T02:37:42Z","receivedAt":"2011-09-09T02:37:42Z","isPatch":false,"sender":{"key":"neubyr@gmail.com","avatar":null},"body":"I have a test git repository with just two files in it. One of the\nfile in it has a set of two lines that is repeated n times.\ne.g.:\n{{{\n$ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt &&  cat\n./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done\n}}}\n\nI ran above command few times and performed commit after each run. Now\ndisk usage of this repository directory is mentioned below. The 419M\nis working directory size and 2.7M is git repository/database size.\n\n{{{\n$ du -h -d 1 .\n2.7M    ./.git\n419M    .\n\n}}}\n\nIs it because of the compression performed by git before storing data\n(or before sending commit)??\n\nFollowing were results with subversion:\n\nSubversion client (redundant(?) copy exists in .svn/text-base/\ndirectory, hence double size in client):\n{{{\n$ du -h -d 1\n416M    ./.svn\n832M    .\n}}}\n\nSubversion repo/server:\n{{{\n$ du -h -d 1\n 12K    ./conf\n1.2M    ./db\n 36K    ./hooks\n8.0K    ./locks\n1.2M    .\n}}}\n\n--\nneuby.r\n"},{"id":"175155","messageId":"1315556595.2019.11.camel@bee.lab.cmartin.tk","threadId":"28342","inReplyTo":"CALFxCvzVjC+u=RDkDCQp0QqPETsv8ROE8tY=37tmMWxmQoJOEw@mail.gmail.com","subject":"Re: git repository size / compression","fromName":"Carlos Martín Nieto","fromEmail":"cmn@elego.de","sentAt":"2011-09-09T08:23:13Z","receivedAt":"2011-09-09T08:23:13Z","isPatch":false,"sender":{"key":"cmn@elego.de","avatar":"https://avatars.githubusercontent.com/u/335443?v=4"},"body":"On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:\n> I have a test git repository with just two files in it. One of the\n> file in it has a set of two lines that is repeated n times.\n> e.g.:\n> {{{\n> $ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt &&  cat\n> ./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done\n> }}}\n> \n\nSo you've just created some data that can be compressed quite\nefficiently.\n\n> I ran above command few times and performed commit after each run. Now\n> disk usage of this repository directory is mentioned below. The 419M\n> is working directory size and 2.7M is git repository/database size.\n> \n> {{{\n> $ du -h -d 1 .\n> 2.7M    ./.git\n> 419M    .\n> \n> }}}\n> \n> Is it because of the compression performed by git before storing data\n> (or before sending commit)??\n> \n\nYes. Git stores its objects (the commit, the snapshot of the files,\netc.) compressed. When these objects are stored in a pack, the size can\nbe further reduced by storing some objects as deltas which describe the\ndifference between itself and some other object in the object-db.\n\n> Following were results with subversion:\n> \n> Subversion client (redundant(?) copy exists in .svn/text-base/\n> directory, hence double size in client):\n> {{{\n> $ du -h -d 1\n> 416M    ./.svn\n> 832M    .\n> }}}\n\nSubversion stores the \"pristines\" (which is the status of the files in\nthe latest revision) inside the .svn directory. I wouldn't call this\ncopy redundant, though, as it allows you to run diff locally. The\npristines are stored uncompressed, which is why you half of the space is\ntaken up by the .svn directory.\n\n> \n> Subversion repo/server:\n> {{{\n> $ du -h -d 1\n>  12K    ./conf\n> 1.2M    ./db\n>  36K    ./hooks\n> 8.0K    ./locks\n> 1.2M    .\n> }}}\n\nI don't know how the repository is stored in Subversion, but it may also\nbe compressed. You may be able to reduced your git repository size by\n(re)generating packs with 'git repack' and doing some cleanups with 'git\ngc', but the repository size is not often a concern.\n\n   cmn\n\n\n"},{"id":"175189","messageId":"CALFxCvxmPN_O_3xpkrGUYtdkVfz5nr7eaucMrAYQ3uvi820FBg@mail.gmail.com","threadId":"28342","inReplyTo":"1315556595.2019.11.camel@bee.lab.cmartin.tk","subject":"Re: git repository size / compression","fromName":"neubyr","fromEmail":"neubyr@gmail.com","sentAt":"2011-09-09T14:04:07Z","receivedAt":"2011-09-09T14:04:07Z","isPatch":false,"sender":{"key":"neubyr@gmail.com","avatar":null},"body":"On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:\n> On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:\n>> I have a test git repository with just two files in it. One of the\n>> file in it has a set of two lines that is repeated n times.\n>> e.g.:\n>> {{{\n>> $ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt &&  cat\n>> ./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done\n>> }}}\n>>\n>\n> So you've just created some data that can be compressed quite\n> efficiently.\n>\n>> I ran above command few times and performed commit after each run. Now\n>> disk usage of this repository directory is mentioned below. The 419M\n>> is working directory size and 2.7M is git repository/database size.\n>>\n>> {{{\n>> $ du -h -d 1 .\n>> 2.7M    ./.git\n>> 419M    .\n>>\n>> }}}\n>>\n>> Is it because of the compression performed by git before storing data\n>> (or before sending commit)??\n>>\n>\n> Yes. Git stores its objects (the commit, the snapshot of the files,\n> etc.) compressed. When these objects are stored in a pack, the size can\n> be further reduced by storing some objects as deltas which describe the\n> difference between itself and some other object in the object-db.\n>\n\nDoes git store deltas for some files? I thought it uses snapshots\n(exact copy of staged files) only.\n\n\n>> Following were results with subversion:\n>>\n>> Subversion client (redundant(?) copy exists in .svn/text-base/\n>> directory, hence double size in client):\n>> {{{\n>> $ du -h -d 1\n>> 416M    ./.svn\n>> 832M    .\n>> }}}\n>\n> Subversion stores the \"pristines\" (which is the status of the files in\n> the latest revision) inside the .svn directory. I wouldn't call this\n> copy redundant, though, as it allows you to run diff locally. The\n> pristines are stored uncompressed, which is why you half of the space is\n> taken up by the .svn directory.\n>\n>>\n>> Subversion repo/server:\n>> {{{\n>> $ du -h -d 1\n>>  12K    ./conf\n>> 1.2M    ./db\n>>  36K    ./hooks\n>> 8.0K    ./locks\n>> 1.2M    .\n>> }}}\n>\n> I don't know how the repository is stored in Subversion, but it may also\n> be compressed. You may be able to reduced your git repository size by\n> (re)generating packs with 'git repack' and doing some cleanups with 'git\n> gc', but the repository size is not often a concern.\n>\n>   cmn\n>\n>\n>\n\nthat's helpful. thanks.\n\n--\nneuby.r\n"},{"id":"175192","messageId":"CAGdFq_iZSRuvNP6Z+Gao+TSVwDaEAREjCKYgiPVAXeSWbzq2EA@mail.gmail.com","threadId":"28342","inReplyTo":"CALFxCvxmPN_O_3xpkrGUYtdkVfz5nr7eaucMrAYQ3uvi820FBg@mail.gmail.com","subject":"Re: git repository size / compression","fromName":"Sverre Rabbelier","fromEmail":"srabbelier@gmail.com","sentAt":"2011-09-09T14:25:41Z","receivedAt":"2011-09-09T14:25:41Z","isPatch":false,"sender":{"key":"srabbelier@gmail.com","avatar":"https://avatars.githubusercontent.com/u/3098?v=4"},"body":"Heya,\n\nOn Fri, Sep 9, 2011 at 16:04, neubyr <neubyr@gmail.com> wrote:\n> Does git store deltas for some files? I thought it uses snapshots\n> (exact copy of staged files) only.\n\nIn packs, yes, it will try to delta objects as efficient as possible.\n\n-- \nCheers,\n\nSverre Rabbelier\n"},{"id":"175193","messageId":"1315578547.4377.2.camel@centaur.lab.cmartin.tk","threadId":"28342","inReplyTo":"CALFxCvxmPN_O_3xpkrGUYtdkVfz5nr7eaucMrAYQ3uvi820FBg@mail.gmail.com","subject":"Re: git repository size / compression","fromName":"Carlos Martín Nieto","fromEmail":"cmn@elego.de","sentAt":"2011-09-09T14:28:54Z","receivedAt":"2011-09-09T14:28:54Z","isPatch":false,"sender":{"key":"cmn@elego.de","avatar":"https://avatars.githubusercontent.com/u/335443?v=4"},"body":"On Fri, 2011-09-09 at 09:04 -0500, neubyr wrote:\n> On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:\n> > On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:\n> >> I have a test git repository with just two files in it. One of the\n> >> file in it has a set of two lines that is repeated n times.\n> >> e.g.:\n> >> {{{\n> >> $ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt &&  cat\n> >> ./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done\n> >> }}}\n> >>\n> >\n> > So you've just created some data that can be compressed quite\n> > efficiently.\n> >\n> >> I ran above command few times and performed commit after each run. Now\n> >> disk usage of this repository directory is mentioned below. The 419M\n> >> is working directory size and 2.7M is git repository/database size.\n> >>\n> >> {{{\n> >> $ du -h -d 1 .\n> >> 2.7M    ./.git\n> >> 419M    .\n> >>\n> >> }}}\n> >>\n> >> Is it because of the compression performed by git before storing data\n> >> (or before sending commit)??\n> >>\n> >\n> > Yes. Git stores its objects (the commit, the snapshot of the files,\n> > etc.) compressed. When these objects are stored in a pack, the size can\n> > be further reduced by storing some objects as deltas which describe the\n> > difference between itself and some other object in the object-db.\n> >\n> \n> Does git store deltas for some files? I thought it uses snapshots\n> (exact copy of staged files) only.\n\nYes and no. The data model for git is to always store snapshots, and it\nalways expects to have the full files available. In a packfile, however,\nin order to save space, some objects are stored as deltas to other\nobjects in the same file.\n\nhttp://progit.org/book/ch9-4.html\n\n> \n> \n> >> Following were results with subversion:\n> >>\n> >> Subversion client (redundant(?) copy exists in .svn/text-base/\n> >> directory, hence double size in client):\n> >> {{{\n> >> $ du -h -d 1\n> >> 416M    ./.svn\n> >> 832M    .\n> >> }}}\n> >\n> > Subversion stores the \"pristines\" (which is the status of the files in\n> > the latest revision) inside the .svn directory. I wouldn't call this\n> > copy redundant, though, as it allows you to run diff locally. The\n> > pristines are stored uncompressed, which is why you half of the space is\n> > taken up by the .svn directory.\n> >\n> >>\n> >> Subversion repo/server:\n> >> {{{\n> >> $ du -h -d 1\n> >>  12K    ./conf\n> >> 1.2M    ./db\n> >>  36K    ./hooks\n> >> 8.0K    ./locks\n> >> 1.2M    .\n> >> }}}\n> >\n> > I don't know how the repository is stored in Subversion, but it may also\n> > be compressed. You may be able to reduced your git repository size by\n> > (re)generating packs with 'git repack' and doing some cleanups with 'git\n> > gc', but the repository size is not often a concern.\n> >\n> >   cmn\n> >\n> >\n> >\n> \n> that's helpful. thanks.\n> \n> --\n> neuby.r\n> \n\n\n"},{"id":"175195","messageId":"m339g5u5pm.fsf@localhost.localdomain","threadId":"28342","inReplyTo":"CALFxCvxmPN_O_3xpkrGUYtdkVfz5nr7eaucMrAYQ3uvi820FBg@mail.gmail.com","subject":"Re: git repository size / compression","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2011-09-09T14:54:55Z","receivedAt":"2011-09-09T14:54:55Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"neubyr <neubyr@gmail.com> writes:\n> On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:\n> > On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:\n\n>>> I have a test git repository with just two files in it. One of the\n>>> file in it has a set of two lines that is repeated n times.\n>>> e.g.:\n>>> {{{\n>>> $ for i in {1..5}; do cat ./lexico.txt>> lexico1.txt &&  cat\n>>> ./lexico.txt>> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done\n>>> }}}\n>>>\n>>\n>> So you've just created some data that can be compressed quite\n>> efficiently.\n>>\n>>> I ran above command few times and performed commit after each run. Now\n>>> disk usage of this repository directory is mentioned below. The 419M\n>>> is working directory size and 2.7M is git repository/database size.\n>>>\n>>> {{{\n>>> $ du -h -d 1 .\n>>> 2.7M    ./.git\n>>> 419M    .\n>>>\n>>> }}}\n\nHave you tried the same but with\n\n   $ git gc --prune=now\n\nbefore running `du`?\n\n>>> Is it because of the compression performed by git before storing data\n>>> (or before sending commit)??\n>>\n>> Yes. Git stores its objects (the commit, the snapshot of the files,\n>> etc.) compressed. When these objects are stored in a pack, the size can\n>> be further reduced by storing some objects as deltas which describe the\n>> difference between itself and some other object in the object-db.\n> \n> Does git store deltas for some files? I thought it uses snapshots\n> (exact copy of staged files) only.\n\nWhen creating packfile from loose objects (e.g. via `git gc`), it\ndoes perform delta compression.\n\n-- \nJakub Narębski\n"},{"id":"175196","messageId":"CALFxCvyptAMU_uOzv06pe+N2W9L7ysBzcMj3YSoXH-2GGBsf-A@mail.gmail.com","threadId":"28342","inReplyTo":"1315578547.4377.2.camel@centaur.lab.cmartin.tk","subject":"Re: git repository size / compression","fromName":"neubyr","fromEmail":"neubyr@gmail.com","sentAt":"2011-09-09T15:07:38Z","receivedAt":"2011-09-09T15:07:38Z","isPatch":false,"sender":{"key":"neubyr@gmail.com","avatar":null},"body":"On Fri, Sep 9, 2011 at 9:28 AM, Carlos Martín Nieto <cmn@elego.de> wrote:\n> On Fri, 2011-09-09 at 09:04 -0500, neubyr wrote:\n>> On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:\n>> > On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:\n>> >> I have a test git repository with just two files in it. One of the\n>> >> file in it has a set of two lines that is repeated n times.\n>> >> e.g.:\n>> >> {{{\n>> >> $ for i in {1..5}; do cat ./lexico.txt >> lexico1.txt &&  cat\n>> >> ./lexico.txt >> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done\n>> >> }}}\n>> >>\n>> >\n>> > So you've just created some data that can be compressed quite\n>> > efficiently.\n>> >\n>> >> I ran above command few times and performed commit after each run. Now\n>> >> disk usage of this repository directory is mentioned below. The 419M\n>> >> is working directory size and 2.7M is git repository/database size.\n>> >>\n>> >> {{{\n>> >> $ du -h -d 1 .\n>> >> 2.7M    ./.git\n>> >> 419M    .\n>> >>\n>> >> }}}\n>> >>\n>> >> Is it because of the compression performed by git before storing data\n>> >> (or before sending commit)??\n>> >>\n>> >\n>> > Yes. Git stores its objects (the commit, the snapshot of the files,\n>> > etc.) compressed. When these objects are stored in a pack, the size can\n>> > be further reduced by storing some objects as deltas which describe the\n>> > difference between itself and some other object in the object-db.\n>> >\n>>\n>> Does git store deltas for some files? I thought it uses snapshots\n>> (exact copy of staged files) only.\n>\n> Yes and no. The data model for git is to always store snapshots, and it\n> always expects to have the full files available. In a packfile, however,\n> in order to save space, some objects are stored as deltas to other\n> objects in the same file.\n>\n> http://progit.org/book/ch9-4.html\n>\n\nExcellent.. That explains compression and deltas really well. Thanks again..\n\n--\nneuby.r\n"},{"id":"175197","messageId":"CALFxCvxAn9tEaOWM6r2A8UiDjkTrrzky1Q10VtqCiA-vhQxrug@mail.gmail.com","threadId":"28342","inReplyTo":"m339g5u5pm.fsf@localhost.localdomain","subject":"Re: git repository size / compression","fromName":"neubyr","fromEmail":"neubyr@gmail.com","sentAt":"2011-09-09T15:09:23Z","receivedAt":"2011-09-09T15:09:23Z","isPatch":false,"sender":{"key":"neubyr@gmail.com","avatar":null},"body":"2011/9/9 Jakub Narebski <jnareb@gmail.com>:\n> neubyr <neubyr@gmail.com> writes:\n>> On Fri, Sep 9, 2011 at 3:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:\n>> > On Thu, 2011-09-08 at 21:37 -0500, neubyr wrote:\n>\n>>>> I have a test git repository with just two files in it. One of the\n>>>> file in it has a set of two lines that is repeated n times.\n>>>> e.g.:\n>>>> {{{\n>>>> $ for i in {1..5}; do cat ./lexico.txt>> lexico1.txt &&  cat\n>>>> ./lexico.txt>> lexico1.txt && mv ./lexico1.txt ./lexico.txt;  done\n>>>> }}}\n>>>>\n>>>\n>>> So you've just created some data that can be compressed quite\n>>> efficiently.\n>>>\n>>>> I ran above command few times and performed commit after each run. Now\n>>>> disk usage of this repository directory is mentioned below. The 419M\n>>>> is working directory size and 2.7M is git repository/database size.\n>>>>\n>>>> {{{\n>>>> $ du -h -d 1 .\n>>>> 2.7M    ./.git\n>>>> 419M    .\n>>>>\n>>>> }}}\n>\n> Have you tried the same but with\n>\n>   $ git gc --prune=now\n>\n> before running `du`?\n>\n\nNope, I hadn't run git gc before. Here are du results after running\ngit gc command. That's about 55% less space now.. Great!\n\n{{{\n$ du -d 1 -h\n924K    ./.git\n417M    .\n}}}\n\n\n>>>> Is it because of the compression performed by git before storing data\n>>>> (or before sending commit)??\n>>>\n>>> Yes. Git stores its objects (the commit, the snapshot of the files,\n>>> etc.) compressed. When these objects are stored in a pack, the size can\n>>> be further reduced by storing some objects as deltas which describe the\n>>> difference between itself and some other object in the object-db.\n>>\n>> Does git store deltas for some files? I thought it uses snapshots\n>> (exact copy of staged files) only.\n>\n> When creating packfile from loose objects (e.g. via `git gc`), it\n> does perform delta compression.\n>\n> --\n> Jakub Narębski\n>\n\nthank you everyone for explaining in detail..\n\n--\nneuby.r\n"},{"id":"175210","messageId":"CAEBDL5U5-1nBGbWtb6+CfBrESoy8+p0Qqw1t1n_5EKFmpq9NhA@mail.gmail.com","threadId":"28342","inReplyTo":"1315556595.2019.11.camel@bee.lab.cmartin.tk","subject":"Re: git repository size / compression","fromName":"John Szakmeister","fromEmail":"john@szakmeister.net","sentAt":"2011-09-09T16:05:03Z","receivedAt":"2011-09-09T16:05:03Z","isPatch":false,"sender":{"key":"john@szakmeister.net","avatar":"https://avatars.githubusercontent.com/u/448087?v=4"},"body":"On Fri, Sep 9, 2011 at 4:23 AM, Carlos Martín Nieto <cmn@elego.de> wrote:\n[snip]\n>> Subversion repo/server:\n>> {{{\n>> $ du -h -d 1\n>>  12K    ./conf\n>> 1.2M    ./db\n>>  36K    ./hooks\n>> 8.0K    ./locks\n>> 1.2M    .\n>> }}}\n>\n> I don't know how the repository is stored in Subversion, but it may also\n> be compressed. You may be able to reduced your git repository size by\n> (re)generating packs with 'git repack' and doing some cleanups with 'git\n> gc', but the repository size is not often a concern.\n\nIt is stored compressed in Subversion, and it also generates deltas\nagainst previous versions.  IIRC, the delta algorithm in an xdelta\nbased one, and then the data is run through compression.  Subversion\nwill at times choose to self-compress the file, instead of doing a\ndelta and compressing.  IIRC, there is some heuristics in there for\ndetermining when to do that, but I forget the exact method.\n\nHTH!\n\n-John\n"},{"id":"175219","messageId":"20110909174937.GA6057@inner.h.iocl.org","threadId":"28342","inReplyTo":"CAEBDL5U5-1nBGbWtb6+CfBrESoy8+p0Qqw1t1n_5EKFmpq9NhA@mail.gmail.com","subject":"Re: git repository size / compression","fromName":"Andreas Krey","fromEmail":"a.krey@gmx.de","sentAt":"2011-09-09T17:49:37Z","receivedAt":"2011-09-09T17:49:37Z","isPatch":false,"sender":{"key":"a.krey@gmx.de","avatar":"https://avatars.githubusercontent.com/u/37810?v=4"},"body":"On Fri, 09 Sep 2011 12:05:03 +0000, John Szakmeister wrote:\n...\n> will at times choose to self-compress the file, instead of doing a\n> delta and compressing.  IIRC, there is some heuristics in there for\n> determining when to do that, but I forget the exact method.\n\nDon't know about the compression part, but subversion does a delta of the nth\nversion of a file (not the global revision number n) against the version m, where\nm is (n & (n-1)), or the least significant '1' bit flipped to '0'. That way, there\nare only O(log(n)) instead of O(n) deltas to apply to get at a specific version.\n\n[Was on the svn users list just then. They described it differently,\n but in essence it's that.]\n\nAndreas\n\n-- \n\"Totally trivial. Famous last words.\"\nFrom: Linus Torvalds <torvalds@*.org>\nDate: Fri, 22 Jan 2010 07:29:21 -0800\n"}]}