{"thread":{"id":"4748","subject":"Compression speed for large files","startedAt":"2006-07-03T11:13:34Z","lastAt":"2006-07-08T02:10:45Z","messageCount":21,"participants":["Joachim B Haga","Alex Riesen","Elrond","Joachim Berdal Haga","Nicolas Pitre","Yakov Lerner","Johannes Schindelin","Linus Torvalds","Junio C Hamano","Jeff King","David Lang"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"23099","messageId":"loom.20060703T124601-969@post.gmane.org","threadId":"4748","inReplyTo":null,"subject":"Compression speed for large files","fromName":"Joachim B Haga","fromEmail":"cjhaga@fys.uio.no","sentAt":"2006-07-03T11:13:34Z","receivedAt":"2006-07-03T11:13:34Z","isPatch":false,"sender":{"key":"cjhaga@fys.uio.no","avatar":null},"body":"I'm looking at doing version control of data files, potentially very large,\noften binary. In git, committing of large files is very slow; I have tested with\na 45MB file, which takes about 1 minute to check in (on an intel core-duo 2GHz).\n\nNow, most of the time is spent in compressing the file. Would it be a good idea\nto change the Z_BEST_COMPRESSION flag to zlib, at least for large files? I have\nmeasured the time spent by git-commit with different flags in sha1_file.c:\n\n  method                 time (s)  object size (kB)\n  Z_BEST_COMPRESSION     62.0      17136\n  Z_DEFAULT_COMPRESSION  10.4      16536\n  Z_BEST_SPEED            4.8      17071\n\nIn this case Z_BEST_COMPRESSION also compresses worse, but that's not the major\nissue: the time is. Here's a couple of other data points, measured with gzip -9,\n-6 and -1 (comparable to the Z_ flags above):\n\n129MB ascii data file\n  method    time (s)  object size (kB)\n  gzip -9   158       23066\n  gzip -6    18       23619\n  gzip -1     6       32304\n\n3MB ascii data file\n  gzip -9   2.2        887\n  gzip -6   0.7        912\n  gzip -1   0.3       1134\n\nSo: is it a good idea to change to faster compression, at least for larger\nfiles? From my (limited) testing I would suggest using Z_BEST_COMPRESSION only\nfor small files (perhaps <1MB?) and Z_DEFAULT_COMPRESSION/Z_BEST_SPEED for\nlarger ones.\n\n\n-j.\n"},{"id":"23101","messageId":"81b0412b0607030503p63b4ee31v7776bd155d3dab29@mail.gmail.com","threadId":"4748","inReplyTo":"loom.20060703T124601-969@post.gmane.org","subject":"Re: Compression speed for large files","fromName":"Alex Riesen","fromEmail":"raa.lkml@gmail.com","sentAt":"2006-07-03T12:03:43Z","receivedAt":"2006-07-03T12:03:43Z","isPatch":false,"sender":{"key":"raa.lkml@gmail.com","avatar":"https://avatars.githubusercontent.com/u/324101?v=4"},"body":"On 7/3/06, Joachim B Haga <cjhaga@fys.uio.no> wrote:\n> So: is it a good idea to change to faster compression, at least for larger\n> files? From my (limited) testing I would suggest using Z_BEST_COMPRESSION only\n> for small files (perhaps <1MB?) and Z_DEFAULT_COMPRESSION/Z_BEST_SPEED for\n> larger ones.\n\nProbably yes, as a per-repo config option.\n"},{"id":"23102","messageId":"loom.20060703T143544-407@post.gmane.org","threadId":"4748","inReplyTo":"81b0412b0607030503p63b4ee31v7776bd155d3dab29@mail.gmail.com","subject":"Re: Compression speed for large files","fromName":"Elrond","fromEmail":"elrond+kernel.org@samba-tng.org","sentAt":"2006-07-03T12:42:20Z","receivedAt":"2006-07-03T12:42:20Z","isPatch":false,"sender":{"key":"elrond+kernel.org@samba-tng.org","avatar":null},"body":"Joachim B Haga <cjhaga <at> fys.uio.no> writes:\n[...]\n>   method                 time (s)  object size (kB)\n>   Z_BEST_COMPRESSION     62.0      17136\n>   Z_DEFAULT_COMPRESSION  10.4      16536\n>   Z_BEST_SPEED            4.8      17071\n> \n> In this case Z_BEST_COMPRESSION also compresses worse,\n[...]\n\nI personally find that very interesting, is this a known \"issue\" with zlib?\nIt suggests, that with different options, it's possible to create smaller\nrepositories, despite the 'advertised' (by zlib, not git) \"best\" compression.\n\n\nAlex Riesen <raa.lkml <at> gmail.com> writes:\n[...]\n> Probably yes, as a per-repo config option.\n\nThe option probably should be the size for which to start using\n\"default\" compression.\n\n\n    Elrond\n"},{"id":"23104","messageId":"44A91C7A.6090902@fys.uio.no","threadId":"4748","inReplyTo":"81b0412b0607030503p63b4ee31v7776bd155d3dab29@mail.gmail.com","subject":"Re: Compression speed for large files","fromName":"Joachim Berdal Haga","fromEmail":"cjhaga@student.matnat.uio.no","sentAt":"2006-07-03T13:32:42Z","receivedAt":"2006-07-03T13:32:42Z","isPatch":false,"sender":{"key":"cjhaga@student.matnat.uio.no","avatar":null},"body":"Alex Riesen wrote:\n> On 7/3/06, Joachim B Haga <cjhaga@fys.uio.no> wrote:\n>> So: is it a good idea to change to faster compression, at least for \n>> larger files? From my (limited) testing I would suggest using \n>> Z_BEST_COMPRESSION only for small files (perhaps <1MB?) and \n>> Z_DEFAULT_COMPRESSION/Z_BEST_SPEED for\n>> larger ones.\n> \n> Probably yes, as a per-repo config option.\n\nI can send a patch later. If it's to be a per-repo option, it's probably \ntoo confusing with several values. Is it ok with\n\ncore.compression = [-1..9]\n\nwhere the numbers are the zlib/gzip constants,\n   -1 = zlib default (currently 6)\n    0 = no compression\n1..9 = various speed/size tradeoffs (9 is git default)\n\nBtw; I just tested the kernel sources. With gzip only, but files \ncompressed individually:\n   time find . -type f | xargs gzip -9 -c | wc -c\n\nI found the space saving from -6 to -9 to be under 0.6%, at double the \nCPU time. So perhaps Z_DEFAULT_COMPRESSION would be good as default.\n\n-j\n"},{"id":"23105","messageId":"loom.20060703T153349-582@post.gmane.org","threadId":"4748","inReplyTo":"loom.20060703T143544-407@post.gmane.org","subject":"Re: Compression speed for large files","fromName":"Joachim B Haga","fromEmail":"cjhaga@fys.uio.no","sentAt":"2006-07-03T13:44:46Z","receivedAt":"2006-07-03T13:44:46Z","isPatch":false,"sender":{"key":"cjhaga@fys.uio.no","avatar":null},"body":"Elrond <elrond+kernel.org <at> samba-tng.org> writes:\n\n> \n> Joachim B Haga <cjhaga <at> fys.uio.no> writes:\n> [...]\n> > In this case Z_BEST_COMPRESSION also compresses worse,\n> [...]\n> \n> I personally find that very interesting, is this a known \"issue\" with zlib?\n> It suggests, that with different options, it's possible to create smaller\n> repositories, despite the 'advertised' (by zlib, not git) \"best\" compression.\n\nThere are also other tunables in zlib, such as the balance between Huffman\ncoding (good for data files) and string matching (good for text files). So with\nmore knowledge of the data it should be possible to compress even better. I'm\nnot advocating tuning this in git though ;)\n\n> \n> Alex Riesen <raa.lkml <at> gmail.com> writes:\n> [...]\n> > Probably yes, as a per-repo config option.\n> \n> The option probably should be the size for which to start using\n> \"default\" compression.\n\nThat is possible, too. I'm open to any decision or consensus, as long as I get\nmy commits in less than 10s :)\n\n-j.\n"},{"id":"23108","messageId":"Pine.LNX.4.64.0607031030150.1213@localhost.localdomain","threadId":"4748","inReplyTo":"44A91C7A.6090902@fys.uio.no","subject":"Re: Compression speed for large files","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2006-07-03T14:33:14Z","receivedAt":"2006-07-03T14:33:14Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Mon, 3 Jul 2006, Joachim Berdal Haga wrote:\n\n> Alex Riesen wrote:\n> > On 7/3/06, Joachim B Haga <cjhaga@fys.uio.no> wrote:\n> > > So: is it a good idea to change to faster compression, at least for larger\n> > > files? From my (limited) testing I would suggest using Z_BEST_COMPRESSION\n> > > only for small files (perhaps <1MB?) and\n> > > Z_DEFAULT_COMPRESSION/Z_BEST_SPEED for\n> > > larger ones.\n> > \n> > Probably yes, as a per-repo config option.\n> \n> I can send a patch later. If it's to be a per-repo option, it's probably too\n> confusing with several values. Is it ok with\n> \n> core.compression = [-1..9]\n> \n> where the numbers are the zlib/gzip constants,\n>   -1 = zlib default (currently 6)\n>    0 = no compression\n> 1..9 = various speed/size tradeoffs (9 is git default)\n\nI think this makes a lot of sense, although IMHO I'd simply use \nZ_DEFAULT_COMPRESSION everywhere and be done with it without extra \ncomplexity which aren't worth the size difference.\n\n\nNicolas\n"},{"id":"23110","messageId":"f36b08ee0607030754k4d10548pfb71dc62c6ee0b21@mail.gmail.com","threadId":"4748","inReplyTo":"Pine.LNX.4.64.0607031030150.1213@localhost.localdomain","subject":"Re: Compression speed for large files","fromName":"Yakov Lerner","fromEmail":"iler.ml@gmail.com","sentAt":"2006-07-03T14:54:01Z","receivedAt":"2006-07-03T14:54:01Z","isPatch":false,"sender":{"key":"iler.ml@gmail.com","avatar":null},"body":"On 7/3/06, Nicolas Pitre <nico@cam.org> wrote:\n> On Mon, 3 Jul 2006, Joachim Berdal Haga wrote:\n>\n> > Alex Riesen wrote:\n> > > On 7/3/06, Joachim B Haga <cjhaga@fys.uio.no> wrote:\n> > > > So: is it a good idea to change to faster compression, at least for larger\n> > > > files? From my (limited) testing I would suggest using Z_BEST_COMPRESSION\n> > > > only for small files (perhaps <1MB?) and\n> > > > Z_DEFAULT_COMPRESSION/Z_BEST_SPEED for\n> > > > larger ones.\n> > >\n> > > Probably yes, as a per-repo config option.\n> >\n> > I can send a patch later. If it's to be a per-repo option, it's probably too\n> > confusing with several values. Is it ok with\n> >\n> > core.compression = [-1..9]\n> >\n> > where the numbers are the zlib/gzip constants,\n> >   -1 = zlib default (currently 6)\n> >    0 = no compression\n> > 1..9 = various speed/size tradeoffs (9 is git default)\n\nIt would be arguable whether, say, 10% better compression is worth\nx(3-8) slower compression. But 3-4% better compression at the cost of\nx(3-8) slower compression time as data suggest ? I think this begs\nfor switching the default to Z_DEFAULT_COMPRESSION\n\nYakov\n"},{"id":"23111","messageId":"Pine.LNX.4.63.0607031702420.29667@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"4748","inReplyTo":"f36b08ee0607030754k4d10548pfb71dc62c6ee0b21@mail.gmail.com","subject":"Re: Compression speed for large files","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-07-03T15:17:46Z","receivedAt":"2006-07-03T15:17:46Z","isPatch":false,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Mon, 3 Jul 2006, Yakov Lerner wrote:\n\n> It would be arguable whether, say, 10% better compression is worth \n> x(3-8) slower compression. But 3-4% better compression at the cost of \n> x(3-8) slower compression time as data suggest ? I think this begs for \n> switching the default to Z_DEFAULT_COMPRESSION\n\nThe real problem, of course, is that you cannot know before you tried, if \nyour data is really well compressible or not.\n\nCiao,\nDscho\n"},{"id":"23118","messageId":"Pine.LNX.4.64.0607030929490.12404@g5.osdl.org","threadId":"4748","inReplyTo":"Pine.LNX.4.64.0607031030150.1213@localhost.localdomain","subject":"Re: Compression speed for large files","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-03T16:31:55Z","receivedAt":"2006-07-03T16:31:55Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Nicolas Pitre wrote:\n\n> On Mon, 3 Jul 2006, Joachim Berdal Haga wrote:\n> > \n> > I can send a patch later. If it's to be a per-repo option, it's probably too\n> > confusing with several values. Is it ok with\n> > \n> > core.compression = [-1..9]\n> > \n> > where the numbers are the zlib/gzip constants,\n> >   -1 = zlib default (currently 6)\n> >    0 = no compression\n> > 1..9 = various speed/size tradeoffs (9 is git default)\n> \n> I think this makes a lot of sense, although IMHO I'd simply use \n> Z_DEFAULT_COMPRESSION everywhere and be done with it without extra \n> complexity which aren't worth the size difference.\n\nI think Z_DEFAULT_COMPRESSION is fine too - we've long since started \nrelying on pack-files and the delta compression for the _real_ size \nimprovements, and as such, the zlib compression is less important.\n\nThat said, the \"core.compression\" thing sounds good to me, and gives \npeople the ability to tune things for their loads.\n\n\t\tLinus\n"},{"id":"23120","messageId":"85d5cm8qfn.fsf_-_@lupus.ig3.net","threadId":"4748","inReplyTo":"Pine.LNX.4.64.0607030929490.12404@g5.osdl.org","subject":"[PATCH] Make zlib compression level configurable, and change default.","fromName":"Joachim B Haga","fromEmail":"cjhaga@fys.uio.no","sentAt":"2006-07-03T18:59:56Z","receivedAt":"2006-07-03T18:59:56Z","isPatch":true,"sender":{"key":"cjhaga@fys.uio.no","avatar":null},"body":"Make zlib compression level configurable, and change the default.\n\nWith the change in default, \"git add .\" on kernel dir is about\ntwice as fast as before, with only minimal (0.5%) change in\nobject size. The speed difference is even more noticeable\nwhen committing large files, which is now up to 8 times faster.\n\nThe configurability is through setting core.compression = [-1..9]\nwhich maps to the zlib constants; -1 is the default, 0 is no\ncompression, and 1..9 are various speed/size tradeoffs, 9\nbeing slowest.\n\nSigned-off-by: Joachim B Haga (cjhaga@fys.uio.no)\n---\n Documentation/config.txt |    6 ++++++\n cache.h                  |    1 +\n config.c                 |    5 +++++\n environment.c            |    1 +\n sha1_file.c              |    4 ++--\n 5 files changed, 15 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex a04c5ad..ac89be7 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -91,6 +91,12 @@ core.warnAmbiguousRefs::\n        If true, git will warn you if the ref name you passed it is ambiguous\n        and might match multiple refs in the .git/refs/ tree. True by default.\n \n+core.compression:\n+       An integer -1..9, indicating the compression level for objects that\n+       are not in a pack file. -1 is the zlib and git default. 0 means no \n+       compression, and 1..9 are various speed/size tradeoffs, 9 being\n+       slowest.\n+\n alias.*::\n        Command aliases for the gitlink:git[1] command wrapper - e.g.\n        after defining \"alias.last = cat-file commit HEAD\", the invocation\ndiff --git a/cache.h b/cache.h\nindex 8719939..84770bf 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -183,6 +183,7 @@ extern int log_all_ref_updates;\n extern int warn_ambiguous_refs;\n extern int shared_repository;\n extern const char *apply_default_whitespace;\n+extern int zlib_compression_level;\n \n #define GIT_REPO_VERSION 0\n extern int repository_format_version;\ndiff --git a/config.c b/config.c\nindex ec44827..61563be 100644\n--- a/config.c\n+++ b/config.c\n@@ -279,6 +279,11 @@ int git_default_config(const char *var, \n                return 0;\n        }\n \n+       if (!strcmp(var, \"core.compression\")) {\n+               zlib_compression_level = git_config_int(var, value);\n+               return 0;\n+       }\n+\n        if (!strcmp(var, \"user.name\")) {\n                strlcpy(git_default_name, value, sizeof(git_default_name));\n                return 0;\ndiff --git a/environment.c b/environment.c\nindex 3de8eb3..1d8ceef 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -20,6 +20,7 @@ int repository_format_version = 0;\n char git_commit_encoding[MAX_ENCODING_LENGTH] = \"utf-8\";\n int shared_repository = PERM_UMASK;\n const char *apply_default_whitespace = NULL;\n+int zlib_compression_level = -1;\n \n static char *git_dir, *git_object_dir, *git_index_file, *git_refs_dir,\n        *git_graft_file;\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 8179630..bc35808 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1458,7 +1458,7 @@ int write_sha1_file(void *buf, unsigned \n \n        /* Set it up */\n        memset(&stream, 0, sizeof(stream));\n-       deflateInit(&stream, Z_BEST_COMPRESSION);\n+       deflateInit(&stream, zlib_compression_level);\n        size = deflateBound(&stream, len+hdrlen);\n        compressed = xmalloc(size);\n \n@@ -1511,7 +1511,7 @@ static void *repack_object(const unsigne\n \n        /* Set it up */\n        memset(&stream, 0, sizeof(stream));\n-       deflateInit(&stream, Z_BEST_COMPRESSION);\n+       deflateInit(&stream, zlib_compression_level);\n        size = deflateBound(&stream, len + hdrlen);\n        buf = xmalloc(size);\n \n-- \n1.4.1.g8fced-dirty\n"},{"id":"23121","messageId":"8564ie8qbe.fsf_-_@lupus.ig3.net","threadId":"4748","inReplyTo":"Pine.LNX.4.64.0607030929490.12404@g5.osdl.org","subject":"[PATCH] Use configurable zlib compression level everywhere.","fromName":"Joachim B Haga","fromEmail":"cjhaga@fys.uio.no","sentAt":"2006-07-03T19:02:29Z","receivedAt":"2006-07-03T19:02:29Z","isPatch":true,"sender":{"key":"cjhaga@fys.uio.no","avatar":null},"body":"This one I'm not so sure about, it's for completeness. But I don't actually use\ngit and haven't tested beyond the git add / git commit stage. Still...\n\nSigned-off-by: Joachim B Haga (cjhaga@fys.uio.no)\n---\n csum-file.c |    2 +-\n diff.c      |    2 +-\n http-push.c |    2 +-\n 3 files changed, 3 insertions(+), 3 deletions(-)\n\ndiff --git a/csum-file.c b/csum-file.c\nindex ebaad03..6a7b40f 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -122,7 +122,7 @@ int sha1write_compressed(struct sha1file\n        void *out;\n \n        memset(&stream, 0, sizeof(stream));\n-       deflateInit(&stream, Z_DEFAULT_COMPRESSION);\n+       deflateInit(&stream, zlib_compression_level);\n        maxsize = deflateBound(&stream, size);\n        out = xmalloc(maxsize);\n \ndiff --git a/diff.c b/diff.c\nindex 5a71489..428ff78 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -583,7 +583,7 @@ static unsigned char *deflate_it(char *d\n        z_stream stream;\n \n        memset(&stream, 0, sizeof(stream));\n-       deflateInit(&stream, Z_BEST_COMPRESSION);\n+       deflateInit(&stream, zlib_compression_level);\n        bound = deflateBound(&stream, size);\n        deflated = xmalloc(bound);\n        stream.next_out = deflated;\ndiff --git a/http-push.c b/http-push.c\nindex e281f70..f761584 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -492,7 +492,7 @@ static void start_put(struct transfer_re\n \n        /* Set it up */\n        memset(&stream, 0, sizeof(stream));\n-       deflateInit(&stream, Z_BEST_COMPRESSION);\n+       deflateInit(&stream, zlib_compression_level);\n        size = deflateBound(&stream, len + hdrlen);\n        request->buffer.buffer = xmalloc(size);\n \n-- \n1.4.1.g8fced-dirty\n"},{"id":"23122","messageId":"Pine.LNX.4.64.0607031226370.12404@g5.osdl.org","threadId":"4748","inReplyTo":"85d5cm8qfn.fsf_-_@lupus.ig3.net","subject":"Re: [PATCH] Make zlib compression level configurable, and change default.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-03T19:33:03Z","receivedAt":"2006-07-03T19:33:03Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Joachim B Haga wrote:\n> \n> The configurability is through setting core.compression = [-1..9]\n> which maps to the zlib constants; -1 is the default, 0 is no\n> compression, and 1..9 are various speed/size tradeoffs, 9\n> being slowest.\n\nMy only worry is that this encodes \"Z_DEFAULT_COMPRESSION\" as being -1, \nwhich happens to be /true/, but I don't think that's a documented \ninterface (you're supposed to use the Z_DEFAULT_COMPRESSION macro, which \ncould have any value, and just _happens_ to be -1).\n\nIs it likely to ever change from that -1? Probably not. So I think your \npatch is technically correct, but it might just be nicer if it did \nsomething like\n\n\t..\n\tif (!strcmp(var, \"core.compression\")) {\n\t\tint level = git_config_int(var, value);\n\t\tif (level == -1)\n\t\t\tlevel = Z_DEFAULT_COMPRESSION;\n\t\telse if (level < 0 || level > Z_BEST_COMPRESSION)\n\t\t\tdie(\"bad zlib compression level %d\", level);\n\t\tzlib_compression_level = level;\n\t\treturn 0;\n\t}\n\t..\n\nwhich would be safer, and a smart compiler might notice that the -1 case \nends up being a no-op, and then just generate code AS IF we just had a\n\n\tif (level < -1 || level > Z_BEST_COMPRESSION)\n\t\tdie(...\n\nthere.\n\nOh, and for all the same reasons, we should use\n\n\tint zlib_compression_level = Z_BEST_COMPRESSION;\n\nfor the default initializer.\n\nHmm?\n\n\t\tLinus\n"},{"id":"23123","messageId":"7v4pxyscdt.fsf@assigned-by-dhcp.cox.net","threadId":"4748","inReplyTo":"8564ie8qbe.fsf_-_@lupus.ig3.net","subject":"Re: [PATCH] Use configurable zlib compression level everywhere.","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2006-07-03T19:43:10Z","receivedAt":"2006-07-03T19:43:10Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Joachim B Haga <cjhaga@fys.uio.no> writes:\n\n> This one I'm not so sure about, it's for completeness. But I don't actually use\n> git and haven't tested beyond the git add / git commit stage. Still...\n>\n> Signed-off-by: Joachim B Haga (cjhaga@fys.uio.no)\n\nYou made a good judgement to notice that these three are\ndifferent.\n\n * sha1write_compressed() in csum-file.c is for producing packs\n   and most of the things we compress there are deltas and less\n   compressible, so even when core.compression is set to high we\n   might be better off using faster compression.\n\n * diff's deflate_it() is about producing binary diffs (later\n   encoded in base85) for textual transfer.  Again it is almost\n   always used to compress deltas so the same comment as above\n   apply to this.\n\n * http-push uses it to send compressed whole object, and this\n   is only used over the network, so it is plausible that the\n   user would want to use different compression level than the\n   usual core.compression.\n\nIt is fine by me to use the same core.compression to these\nthree.  If somebody comes up with a workload that benefits from\nhaving different settings for them, we can add separate\nvariables, falling back on the default core.compression if there\nisn't one, as needed.\n\nThanks for the patches.\n"},{"id":"23125","messageId":"Pine.LNX.4.64.0607031248320.12404@g5.osdl.org","threadId":"4748","inReplyTo":"Pine.LNX.4.64.0607031226370.12404@g5.osdl.org","subject":"Re: [PATCH] Make zlib compression level configurable, and change default.","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-03T19:50:23Z","receivedAt":"2006-07-03T19:50:23Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 3 Jul 2006, Linus Torvalds wrote:\n> \n> Oh, and for all the same reasons, we should use\n> \n> \tint zlib_compression_level = Z_BEST_COMPRESSION;\n\nThat should be Z_DEFAULT_COMPRESSION, of course.\n\nAnyway, I think the patches are ok as-is, and my suggestion to avoid the \n\"-1\" and use Z_DEFAULT_COMPRESSION is really just an additional comment, \nnot anything fundamental.\n\nSo Junio, feel free to add an\n\n\tAcked-by: Linus Torvalds <torvalds@osdl.org>\n\nregardless of whether also doing that.\n\n\t\tLinus\n"},{"id":"23128","messageId":"85u05y78jg.fsf@lupus.ig3.net","threadId":"4748","inReplyTo":"Pine.LNX.4.64.0607031226370.12404@g5.osdl.org","subject":"Re: [PATCH] Make zlib compression level configurable, and change default.","fromName":"Joachim B Haga","fromEmail":"cjhaga@fys.uio.no","sentAt":"2006-07-03T20:11:47Z","receivedAt":"2006-07-03T20:11:47Z","isPatch":true,"sender":{"key":"cjhaga@fys.uio.no","avatar":null},"body":"Linus Torvalds <torvalds@osdl.org> writes:\n\n> [snip suggested improvements]\n\nYes, that would be more... thorough. And, especially for the\n\n  int zlib_compression_level = Z_DEFAULT_COMPRESSION;\n\nline, more self-explanatory, too. So here's an updated patch\n(replacing the previous) including your suggestions.\n\n-j.\n\n-\n\nMake zlib compression level configurable, and change default.\n\nWith the change in default, \"git add .\" on kernel dir is about\ntwice as fast as before, with only minimal (0.5%) change in\nobject size. The speed difference is even more noticeable\nwhen committing large files, which is now up to 8 times faster.\n\nThe configurability is through setting core.compression = [-1..9]\nwhich maps to the zlib constants; -1 is the default, 0 is no\ncompression, and 1..9 are various speed/size tradeoffs, 9\nbeing slowest.\n\nSigned-off-by: Joachim B Haga (cjhaga@fys.uio.no)\n---\n Documentation/config.txt |    6 ++++++\n cache.h                  |    1 +\n config.c                 |   10 ++++++++++\n environment.c            |    1 +\n sha1_file.c              |    4 ++--\n 5 files changed, 20 insertions(+), 2 deletions(-)\n\ndiff --git a/Documentation/config.txt b/Documentation/config.txt\nindex a04c5ad..ac89be7 100644\n--- a/Documentation/config.txt\n+++ b/Documentation/config.txt\n@@ -91,6 +91,12 @@ core.warnAmbiguousRefs::\n        If true, git will warn you if the ref name you passed it is ambiguous\n        and might match multiple refs in the .git/refs/ tree. True by default.\n \n+core.compression:\n+       An integer -1..9, indicating the compression level for objects that\n+       are not in a pack file. -1 is the zlib and git default. 0 means no \n+       compression, and 1..9 are various speed/size tradeoffs, 9 being\n+       slowest.\n+\n alias.*::\n        Command aliases for the gitlink:git[1] command wrapper - e.g.\n        after defining \"alias.last = cat-file commit HEAD\", the invocation\ndiff --git a/cache.h b/cache.h\nindex 8719939..84770bf 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -183,6 +183,7 @@ extern int log_all_ref_updates;\n extern int warn_ambiguous_refs;\n extern int shared_repository;\n extern const char *apply_default_whitespace;\n+extern int zlib_compression_level;\n \n #define GIT_REPO_VERSION 0\n extern int repository_format_version;\ndiff --git a/config.c b/config.c\nindex ec44827..b23f4bf 100644\n--- a/config.c\n+++ b/config.c\n@@ -279,6 +279,16 @@ int git_default_config(const char *var, \n                return 0;\n        }\n \n+       if (!strcmp(var, \"core.compression\")) {\n+               int level = git_config_int(var, value);\n+               if (level == -1)\n+                       level = Z_DEFAULT_COMPRESSION;\n+               else if (level < 0 || level > Z_BEST_COMPRESSION)\n+                       die(\"bad zlib compression level %d\", level);\n+               zlib_compression_level = level;\n+               return 0;\n+       }\n+\n        if (!strcmp(var, \"user.name\")) {\n                strlcpy(git_default_name, value, sizeof(git_default_name));\n                return 0;\ndiff --git a/environment.c b/environment.c\nindex 3de8eb3..43823ff 100644\n--- a/environment.c\n+++ b/environment.c\n@@ -20,6 +20,7 @@ int repository_format_version = 0;\n char git_commit_encoding[MAX_ENCODING_LENGTH] = \"utf-8\";\n int shared_repository = PERM_UMASK;\n const char *apply_default_whitespace = NULL;\n+int zlib_compression_level = Z_DEFAULT_COMPRESSION;\n \n static char *git_dir, *git_object_dir, *git_index_file, *git_refs_dir,\n        *git_graft_file;\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 8179630..bc35808 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -1458,7 +1458,7 @@ int write_sha1_file(void *buf, unsigned \n \n        /* Set it up */\n        memset(&stream, 0, sizeof(stream));\n-       deflateInit(&stream, Z_BEST_COMPRESSION);\n+       deflateInit(&stream, zlib_compression_level);\n        size = deflateBound(&stream, len+hdrlen);\n        compressed = xmalloc(size);\n \n@@ -1511,7 +1511,7 @@ static void *repack_object(const unsigne\n \n        /* Set it up */\n        memset(&stream, 0, sizeof(stream));\n-       deflateInit(&stream, Z_BEST_COMPRESSION);\n+       deflateInit(&stream, zlib_compression_level);\n        size = deflateBound(&stream, len + hdrlen);\n        buf = xmalloc(size);\n \n-- \n1.4.1.g8fced-dirty\n"},{"id":"23145","messageId":"20060703214503.GA3897@coredump.intra.peff.net","threadId":"4748","inReplyTo":"loom.20060703T124601-969@post.gmane.org","subject":"Re: Compression speed for large files","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2006-07-03T21:45:03Z","receivedAt":"2006-07-03T21:45:03Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Jul 03, 2006 at 11:13:34AM +0000, Joachim B Haga wrote:\n\n> often binary. In git, committing of large files is very slow; I have\n> tested with a 45MB file, which takes about 1 minute to check in (on an\n> intel core-duo 2GHz).\n\nI know this has already been somewhat solved, but I found your numbers\ncuriously high. I work quite a bit with git and large files and I\nhaven't noticed this slowdown. Can you be more specific about your load?\nAre you sure it is zlib?\n\nOn my 1.8Ghz Athlon, compressing 45MB of zeros into 20K takes about 2s.\nCompressing 45MB of random data into a 45MB object takes 6.3s. In either\ncase, the commit takes only about 0.5s (since cogito stores the object\nduring the cg-add).\n\nIs there some specific file pattern which is slow to compress? \n\n-Peff\n"},{"id":"23148","messageId":"44A99961.8090504@fys.uio.no","threadId":"4748","inReplyTo":"20060703214503.GA3897@coredump.intra.peff.net","subject":"Re: Compression speed for large files","fromName":"Joachim Berdal Haga","fromEmail":"c.j.b.haga@fys.uio.no","sentAt":"2006-07-03T22:25:37Z","receivedAt":"2006-07-03T22:25:37Z","isPatch":false,"sender":{"key":"c.j.b.haga@fys.uio.no","avatar":null},"body":"Jeff King wrote:\n> On Mon, Jul 03, 2006 at 11:13:34AM +0000, Joachim B Haga wrote:\n> \n>> often binary. In git, committing of large files is very slow; I have\n>> tested with a 45MB file, which takes about 1 minute to check in (on an\n>> intel core-duo 2GHz).\n> \n> I know this has already been somewhat solved, but I found your numbers\n> curiously high. I work quite a bit with git and large files and I\n> haven't noticed this slowdown. Can you be more specific about your load?\n> Are you sure it is zlib?\n\nQuite sure: at least to the extent that it is fixed by lowering the\ncompression level. But the wording was inexact: it's during object\ncreation, which happens at initial \"git add\" and then later during \"git\ncommit\".\n\nBut...\n\n> y 1.8Ghz Athlon, compressing 45MB of zeros into 20K takes about 2s.\n> Compressing 45MB of random data into a 45MB object takes 6.3s. In either\n> case, the commit takes only about 0.5s (since cogito stores the object\n> during the cg-add).\n> \n> Is there some specific file pattern which is slow to compress? \n\nyes, it seems so. At least the effect is much more pronounced for my\nfiles than for random/null data. \"My\" files are in this context generated\ndata files, binary or ascii.\n\nHere's a test with \"time gzip -[169] -c file >/dev/null\". Random data\nfrom /dev/urandom, kernel headers are concatenation of *.h in kernel\nsources. All times in seconds, on my puny home computer (1GHz Via Nehemiah)\n\n       random (23MB)  data (23MB)   headers (44MB)\n-9     10.2           72.5          38.5\n-6     10.2           13.5          12.9\n-1      9.9            4.1           7.0\n\nSo... data dependent, yes. But it hits even for normal source code.\n\n(Btw; the default (-6) seems to be less data dependent than the other\nvalues. Maybe that's on purpose.)\n\nIf you want to look at a highly-variable dataset (the one above), try\nhttp://lupus.ig3.net/SIMULATION.dx.gz (5MB, slow server), but that's just\nan example, I see the same variability for example also on binary data files.\n\n-j.\n"},{"id":"23152","messageId":"Pine.LNX.4.64.0607031556480.12404@g5.osdl.org","threadId":"4748","inReplyTo":"44A99961.8090504@fys.uio.no","subject":"Re: Compression speed for large files","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2006-07-03T23:02:39Z","receivedAt":"2006-07-03T23:02:39Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 4 Jul 2006, Joachim Berdal Haga wrote:\n> \n> Here's a test with \"time gzip -[169] -c file >/dev/null\". Random data\n> from /dev/urandom, kernel headers are concatenation of *.h in kernel\n> sources. All times in seconds, on my puny home computer (1GHz Via Nehemiah)\n\nThat \"Via Nehemiah\" is probably a big part of it.\n\nI think the VIA Nehemiah just has a 64kB L2 cache, and I bet performance \nplummets if the tables end up being used past that. \n\nAnd I think a large part of the higher compressions is that they allow the \ncompression window and tables to grow bigger.\n\n\t\tLinus\n"},{"id":"23181","messageId":"44A9FFAB.1010708@fys.uio.no","threadId":"4748","inReplyTo":"Pine.LNX.4.64.0607031556480.12404@g5.osdl.org","subject":"Re: Compression speed for large files","fromName":"Joachim Berdal Haga","fromEmail":"c.j.b.haga@fys.uio.no","sentAt":"2006-07-04T05:42:03Z","receivedAt":"2006-07-04T05:42:03Z","isPatch":false,"sender":{"key":"c.j.b.haga@fys.uio.no","avatar":null},"body":"Linus Torvalds wrote:\n> \n> On Tue, 4 Jul 2006, Joachim Berdal Haga wrote:\n>> Here's a test with \"time gzip -[169] -c file >/dev/null\". Random data\n>> from /dev/urandom, kernel headers are concatenation of *.h in kernel\n>> sources. All times in seconds, on my puny home computer (1GHz Via Nehemiah)\n> \n> That \"Via Nehemiah\" is probably a big part of it.\n> \n> I think the VIA Nehemiah just has a 64kB L2 cache, and I bet performance \n> plummets if the tables end up being used past that. \n\nNot really. The numbers in my original post were from a Intel core-duo,\nthey were: 158/18/6 s for comparable (but larger) data.\n\nAnd on a P4 1.8GHz with 512kB L2, the same 23MB data file compresses in\n28.1/5.9/1.3 s. That's a factor 22 slowest/fastest; the VIA was only\nfactor 18, so the difference is actually *larger*.\n\n-j.\n"},{"id":"23397","messageId":"Pine.LNX.4.63.0607071451430.1836@qynat.qvtvafvgr.pbz","threadId":"4748","inReplyTo":"7v4pxyscdt.fsf@assigned-by-dhcp.cox.net","subject":"Re: [PATCH] Use configurable zlib compression level everywhere.","fromName":"David Lang","fromEmail":"dlang@digitalinsight.com","sentAt":"2006-07-07T21:53:34Z","receivedAt":"2006-07-07T21:53:34Z","isPatch":true,"sender":{"key":"dlang@digitalinsight.com","avatar":null},"body":"On Mon, 3 Jul 2006, Junio C Hamano wrote:\n\n> * sha1write_compressed() in csum-file.c is for producing packs\n>   and most of the things we compress there are deltas and less\n>   compressible, so even when core.compression is set to high we\n>   might be better off using faster compression.\n\nwhy would deltas have poor compression? I'd expect them to have about the same \nas the files they are deltas of (or slightly better due to the fact that the \ndeta metainfo is highly repetitive)\n\nDavid Lang\n"},{"id":"23401","messageId":"Pine.LNX.4.63.0607080406370.29667@wbgn013.biozentrum.uni-wuerzburg.de","threadId":"4748","inReplyTo":"Pine.LNX.4.63.0607071451430.1836@qynat.qvtvafvgr.pbz","subject":"Re: [PATCH] Use configurable zlib compression level everywhere.","fromName":"Johannes Schindelin","fromEmail":"johannes.schindelin@gmx.de","sentAt":"2006-07-08T02:10:45Z","receivedAt":"2006-07-08T02:10:45Z","isPatch":true,"sender":{"key":"johannes.schindelin@gmx.de","avatar":"https://avatars.githubusercontent.com/u/127790?v=4"},"body":"Hi,\n\nOn Fri, 7 Jul 2006, David Lang wrote:\n\n> On Mon, 3 Jul 2006, Junio C Hamano wrote:\n> \n> > * sha1write_compressed() in csum-file.c is for producing packs\n> >   and most of the things we compress there are deltas and less\n> >   compressible, so even when core.compression is set to high we\n> >   might be better off using faster compression.\n> \n> why would deltas have poor compression? I'd expect them to have about the same\n> as the files they are deltas of (or slightly better due to the fact that the\n> deta metainfo is highly repetitive)\n\nDeltas should have poor compression by definition, because compression \ntries to encode those parts of the file more efficiently, which do not \nbear much information (think entropy).\n\nIf you have deltas which really make sense, they are almost _pure_ \ninformation, i.e. they do not contain much redundancy, as compared to real \nfiles. So, the compression (which does not know anything about the \ncharacteristics of deltas in particular) cannot take much redundancy out \nof the delta. Therefore, the entropy is very high, and the compression \nrate is low.\n\nHope this makes sense to you,\nDscho\n"}]}