git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: obnoxious CLI complaints

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Sep 11, 2009, 22:16 UTC
Message-ID
<alpine.LFD.2.01.0909111510520.3654@localhost.localdomain>
In-Reply-To
<4AAAC8CE.8020302@lsrfire.ath.cx>
On Sat, 12 Sep 2009, René Scharfe wrote:
Show 13 quoted lines
> 
> But what has bugged me since I added zip support is this result:
> 
> 	# git v1.6.5-rc0
> 	$ time git archive --format=zip -6 v2.6.31 >/dev/null
> 
> 	real	0m16.471s
> 	user	0m16.340s
> 	sys	0m0.128s
> 
> I'd have expected this to be the slowest case, because it's compressing
> all files separately, i.e. it needs to create and flush the compression
> context lots of times instead of only once as in the two cases above.
Oh no, I think it's easily explained.

Compressing many small files really is often cheaper than compressing one large one.

With lots of small files, you end up being very limited in the search-space, so the compression decisions get simpler. Compression in general is not O(n), it's some non-linear factor, often something like O(n**2).

Of course, all compression libraries have an upper bound on the non-linearity (often expressed as a "window size"), so a particular compression algorithm may end up being close to O(n) (with a huge constant). But that upper bound will only kick in for large files, small files that fit entirely into the compression window will still see the underlying O(n**2) or whatever.

But I have no actual numbers to back up the above blathering. But feel free to try to compress 10 small files and compare it to compressing one file that is as big as the sum. I bet you'll see it.

			Linus
Previous: René ScharfeNext: Dmitry Potapov
Message 25 of 46 in “obnoxious CLI complaints”
  1. Brendan MillerSep 9, 2009
  2. Jakub NarebskiSep 9, 2009
  3. Wincent ColaiutaSep 9, 2009
  4. Jakub NarebskiSep 10, 2009
  5. Junio C HamanoSep 10, 2009
  6. René ScharfeSep 10, 2009
  7. Björn SteinbrinkSep 11, 2009
  8. John TapsellSep 10, 2009
  9. Sverre RabbelierSep 10, 2009
  10. Jakub NarebskiSep 10, 2009
  11. John TapsellSep 10, 2009
  12. Junio C HamanoSep 10, 2009
  13. demerphqSep 10, 2009
  14. Junio C HamanoSep 11, 2009
  15. John TapsellSep 11, 2009
  16. Junio C HamanoSep 11, 2009
  17. Brendan MillerSep 10, 2009
  18. Todd ZullingerSep 10, 2009
  19. Jakub NarebskiSep 10, 2009
  20. Eric SchaeferSep 10, 2009
  21. Sverre RabbelierSep 10, 2009
  22. René ScharfeSep 10, 2009
  23. Linus TorvaldsSep 11, 2009
  24. René ScharfeSep 11, 2009
  25. Linus TorvaldsSep 11, 2009
  26. Dmitry PotapovSep 12, 2009
  27. John TapsellSep 12, 2009
  28. Dmitry PotapovSep 12, 2009
  29. John TapsellSep 12, 2009
  30. A Large Angry SCMSep 12, 2009
  31. Dmitry PotapovSep 12, 2009
  32. John TapsellSep 12, 2009
  33. Junio C HamanoSep 13, 2009
  34. 1/2 git-archive: add '-o' as a alias for '--output'Dmitry Potapov, Sep 13, 2009
  35. 2/2 teach git-archive to auto detect the output formatDmitry Potapov, Sep 13, 2009
  36. Junio C HamanoSep 13, 2009
  37. 2/2 teach git-archive to auto detect the output formatDmitry Potapov, Sep 13, 2009
  38. Junio C HamanoSep 13, 2009
  39. Junio C HamanoSep 13, 2009
  40. 1/2 git-archive: add '-o' as a alias for '--output'Dmitry Potapov, Sep 13, 2009
  41. Brendan MillerSep 17, 2009
  42. Junio C HamanoSep 17, 2009
  43. Sverre RabbelierSep 9, 2009
  44. Pierre HabouzitSep 9, 2009
  45. Björn SteinbrinkSep 10, 2009
  46. Matthieu MoySep 10, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.