git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Compression speed for large files

From
JHJoachim Berdal Haga <c.j.b.haga@fys.uio.no>
Date
Jul 3, 2006, 22:25 UTC
Message-ID
<44A99961.8090504@fys.uio.no>
In-Reply-To
<20060703214503.GA3897@coredump.intra.peff.net>
Jeff King wrote:
Show 10 quoted lines
> On Mon, Jul 03, 2006 at 11:13:34AM +0000, Joachim B Haga wrote:
> 
>> often binary. In git, committing of large files is very slow; I have
>> tested with a 45MB file, which takes about 1 minute to check in (on an
>> intel core-duo 2GHz).
> 
> I know this has already been somewhat solved, but I found your numbers
> curiously high. I work quite a bit with git and large files and I
> haven't noticed this slowdown. Can you be more specific about your load?
> Are you sure it is zlib?

Quite sure: at least to the extent that it is fixed by lowering the compression level. But the wording was inexact: it's during object creation, which happens at initial "git add" and then later during "git commit".

But...
Show 6 quoted lines
> y 1.8Ghz Athlon, compressing 45MB of zeros into 20K takes about 2s.
> Compressing 45MB of random data into a 45MB object takes 6.3s. In either
> case, the commit takes only about 0.5s (since cogito stores the object
> during the cg-add).
> 
> Is there some specific file pattern which is slow to compress? 

yes, it seems so. At least the effect is much more pronounced for my files than for random/null data. "My" files are in this context generated data files, binary or ascii.

Here's a test with "time gzip -[169] -c file >/dev/null". Random data from /dev/urandom, kernel headers are concatenation of *.h in kernel sources. All times in seconds, on my puny home computer (1GHz Via Nehemiah)

       random (23MB)  data (23MB)   headers (44MB)
-9     10.2           72.5          38.5
-6     10.2           13.5          12.9
-1      9.9            4.1           7.0
So... data dependent, yes. But it hits even for normal source code.

(Btw; the default (-6) seems to be less data dependent than the other values. Maybe that's on purpose.)

If you want to look at a highly-variable dataset (the one above), try http://lupus.ig3.net/SIMULATION.dx.gz (5MB, slow server), but that's just an example, I see the same variability for example also on binary data files.

-j.
Previous: Jeff KingNext: Linus Torvalds
Message 19 of 21 in “Compression speed for large files”
  1. Joachim B HagaJul 3, 2006
  2. Alex RiesenJul 3, 2006
  3. ElrondJul 3, 2006
  4. Joachim B HagaJul 3, 2006
  5. Joachim Berdal HagaJul 3, 2006
  6. Nicolas PitreJul 3, 2006
  7. Yakov LernerJul 3, 2006
  8. Johannes SchindelinJul 3, 2006
  9. Linus TorvaldsJul 3, 2006
  10. Make zlib compression level configurable, and change default.Joachim B Haga, Jul 3, 2006
  11. Linus TorvaldsJul 3, 2006
  12. Linus TorvaldsJul 3, 2006
  13. Joachim B HagaJul 3, 2006
  14. Use configurable zlib compression level everywhere.Joachim B Haga, Jul 3, 2006
  15. Junio C HamanoJul 3, 2006
  16. David LangJul 7, 2006
  17. Johannes SchindelinJul 8, 2006
  18. Jeff KingJul 3, 2006
  19. Joachim Berdal HagaJul 3, 2006
  20. Linus TorvaldsJul 3, 2006
  21. Joachim Berdal HagaJul 4, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.