git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Decompression speed: zip vs lzo

From
Nicolas Pitre <nico@cam.org>
Date
Jan 10, 2008, 20:39 UTC
Message-ID
<alpine.LFD.1.00.0801101332150.3054@xanadu.home>
In-Reply-To
<20080110091607.GA17944@artemis.madism.org>
On Thu, 10 Jan 2008, Pierre Habouzit wrote:
Show 8 quoted lines
> Well, lzma is excellent for *big* chunks of data, but not that impressive for
> small files:
> 
> $ ll git.c git.c.gz git.c.lzma git.c.lzop
> -rw-r--r-- 1 madcoder madcoder 12915 2008-01-09 13:47 git.c
> -rw-r--r-- 1 madcoder madcoder  4225 2008-01-10 10:00 git.c.gz
> -rw-r--r-- 1 madcoder madcoder  4094 2008-01-10 10:00 git.c.lzma
> -rw-r--r-- 1 madcoder madcoder  5068 2008-01-10 09:59 git.c.lzop

This is really the big point here. Git uses _lots_ of *small* objects, usually much smaller than 12KB. For example, my copy of the gcc repository has an average of 270 _bytes_ per compressed object, and objects must be individually compressed.

Performance with really small objects should be the basis for any Git compression algorithm comparison.

Show 6 quoted lines
> Though I don't agree with you (and some others) about the fact that 
> gzip is fast enough. It's clearly a bottleneck in many log related 
> commands where you would expect it to be rather IO bound than CPU 
> bound.  LZO seems like a fairer choice, especially since what it makes 
> gain is basically the compression of the biggest blobs, aka the delta 
> chains heads.

The delta heads, though, are far from being the most frequently accessed objects. First they're clearly in minority, and often cached in the delta base cache.

> It's really unclear to me if we really gain in 
> compressing the deltas, trees, and other smallish informations.

Remember that delta objects represent the vast majority of all objects. For example, my kernel repo currently has 555015 delta objects out of 677073 objects, or 82% of the total. There is actually only 25869 non deltified blob objects which are likely to be the larger objects, but they represent only 4% of the total.

But just let's try not compressing delta objects so to check your assertion with the following hack:

diff --git a/builtin-pack-objects.c b/builtin-pack-objects.c
index a39cb82..252b03e 100644
--- a/builtin-pack-objects.c
+++ b/builtin-pack-objects.c
@@ -433,7 +433,10 @@ static unsigned long write_object(struct sha1file *f,
 		}
 		/* compress the data to store and put compressed length in datalen */
 		memset(&stream, 0, sizeof(stream));
-		deflateInit(&stream, pack_compression_level);
+		if (obj_type == OBJ_REF_DELTA || obj_type == OBJ_OFS_DELTA)
+			deflateInit(&stream, 0);
+		else
+			deflateInit(&stream, pack_compression_level);
 		maxsize = deflateBound(&stream, size);
 		out = xmalloc(maxsize);
 		/* Compress it */

You then only need to run 'git repack -a -f -d' with and without the 
above patch.

Here's my rather surprising results:

My kernel repo pack size without the patch:	184275401 bytes
Same repo with the above patch applied:		205204930 bytes

So it is only 11% larger.  I was expecting much more.

I'll let someone else do profiling/timing comparisons.

> What is obvious to me is that lzop seems to take 10% more space than gzip,
> while being around 1.5 to 2 times faster. Of course this is very sketchy and a
> real test with git will be better.

Right.  Abstracting the zlib code and having different compression 
algorithms tested in the Git context is the only way to do meaningful 
comparisons.


Nicolas
Previous: Pierre HabouzitNext: Linus Torvalds
Message 8 of 39 in “Decompression speed: zip vs lzo”
  1. Marco CostalbaJan 9, 2008
  2. Junio C HamanoJan 9, 2008
  3. Sam VilainJan 9, 2008
  4. Johannes SchindelinJan 9, 2008
  5. Sam VilainJan 10, 2008
  6. Sam VilainJan 10, 2008
  7. Pierre HabouzitJan 10, 2008
  8. Nicolas PitreJan 10, 2008
  9. Linus TorvaldsJan 10, 2008
  10. Nicolas PitreJan 10, 2008
  11. Pierre HabouzitJan 11, 2008
  12. Sam VilainJan 10, 2008
  13. Linus TorvaldsJan 10, 2008
  14. Sam VilainJan 10, 2008
  15. Linus TorvaldsJan 10, 2008
  16. Sam VilainJan 11, 2008
  17. Linus TorvaldsJan 11, 2008
  18. Sam VilainJan 11, 2008
  19. Sam VilainJan 11, 2008
  20. Linus TorvaldsJan 11, 2008
  21. Sam VilainJan 12, 2008
  22. Nicolas PitreJan 12, 2008
  23. Sam VilainJan 12, 2008
  24. Nicolas PitreJan 12, 2008
  25. Johannes SchindelinJan 12, 2008
  26. Junio C HamanoJan 12, 2008
  27. Marco CostalbaJan 10, 2008
  28. Sam VilainJan 10, 2008
  29. Nicolas PitreJan 10, 2008
  30. Pierre HabouzitJan 11, 2008
  31. Nicolas PitreJan 11, 2008
  32. Morten WelinderJan 11, 2008
  33. Nicolas PitreJan 10, 2008
  34. Marco CostalbaJan 10, 2008
  35. Marco CostalbaJan 10, 2008
  36. Johannes SchindelinJan 10, 2008
  37. Marco CostalbaJan 10, 2008
  38. Dana HowJan 10, 2008
  39. Junio C HamanoJan 9, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.