git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Decompression speed: zip vs lzo

From
Junio C Hamano <gitster@pobox.com>
Date
Jan 12, 2008, 04:46 UTC
Message-ID
<7vzlvby7li.fsf@gitster.siamese.dyndns.org>
In-Reply-To
<47881D44.9060105@vilain.net>
Sam Vilain <sam@vilain.net> writes:
> If the uncompressed objects are clustered in the pack, then they might
> stream compress a lot better, should they be tranmitted over a http
> transport with gzip encoding.

That would only have been a sensible optimization in older native pack protocol, where we always exploded the transferred packfile. However, these days, we tend to keep the packfile and re-index at the receiving end (http transport never exploded the packfile and it still doesn't). When used that way, choosing object layout in packfile in such a way to ignore recency order and cluster objects by their delta chain, which you are advocating to reduce the transfer overhead, is a bad tradeoff. Your packs will be kept in the form you chose for transport, which is a layout that hurts the runtime performance. And you keep using that suboptimal packs number of times, getting hurt every time.

Show 9 quoted lines
> @@ -433,7 +434,7 @@ static unsigned long write_object(struct sha1file *f,
>  		}
>  		/* compress the data to store and put compressed length in datalen */
>  		memset(&stream, 0, sizeof(stream));
> -		deflateInit(&stream, pack_compression_level);
> +		deflateInit(&stream, size >= compression_min_size ? pack_compression_level : 0);
>  		maxsize = deflateBound(&stream, size);
>  		out = xmalloc(maxsize);
>  		/* Compress it */

I very much like the simplicity of the patch. If such a simple approach can give us a clear performance gain, I am all for it.

Benchmarks on different repositories need to back that up, though.

Previous: Johannes SchindelinNext: Marco Costalba
Message 26 of 39 in “Decompression speed: zip vs lzo”
  1. Marco CostalbaJan 9, 2008
  2. Junio C HamanoJan 9, 2008
  3. Sam VilainJan 9, 2008
  4. Johannes SchindelinJan 9, 2008
  5. Sam VilainJan 10, 2008
  6. Sam VilainJan 10, 2008
  7. Pierre HabouzitJan 10, 2008
  8. Nicolas PitreJan 10, 2008
  9. Linus TorvaldsJan 10, 2008
  10. Nicolas PitreJan 10, 2008
  11. Pierre HabouzitJan 11, 2008
  12. Sam VilainJan 10, 2008
  13. Linus TorvaldsJan 10, 2008
  14. Sam VilainJan 10, 2008
  15. Linus TorvaldsJan 10, 2008
  16. Sam VilainJan 11, 2008
  17. Linus TorvaldsJan 11, 2008
  18. Sam VilainJan 11, 2008
  19. Sam VilainJan 11, 2008
  20. Linus TorvaldsJan 11, 2008
  21. Sam VilainJan 12, 2008
  22. Nicolas PitreJan 12, 2008
  23. Sam VilainJan 12, 2008
  24. Nicolas PitreJan 12, 2008
  25. Johannes SchindelinJan 12, 2008
  26. Junio C HamanoJan 12, 2008
  27. Marco CostalbaJan 10, 2008
  28. Sam VilainJan 10, 2008
  29. Nicolas PitreJan 10, 2008
  30. Pierre HabouzitJan 11, 2008
  31. Nicolas PitreJan 11, 2008
  32. Morten WelinderJan 11, 2008
  33. Nicolas PitreJan 10, 2008
  34. Marco CostalbaJan 10, 2008
  35. Marco CostalbaJan 10, 2008
  36. Johannes SchindelinJan 10, 2008
  37. Marco CostalbaJan 10, 2008
  38. Dana HowJan 10, 2008
  39. Junio C HamanoJan 9, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.