git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: cleaner/better zlib sources?

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Mar 16, 2007, 17:51 UTC
Message-ID
<Pine.LNX.4.64.0703161026220.3816@woody.linux-foundation.org>
In-Reply-To
<alpine.LFD.0.83.0703161236180.5518@xanadu.home>
On Fri, 16 Mar 2007, Nicolas Pitre wrote:
Show 9 quoted lines
> On Fri, 16 Mar 2007, Linus Torvalds wrote:
> 
> > The most performance-critical objects for uncompression are commits and 
> > trees. At least for the kernel, the average size of a tree object is 678
> > bytes. And that's ignoring the fact that most of them are then deltified, 
> > so about 80% of them are likely just a ~60-byte delta.
> 
> This is why in pack v4 there will be an alternate tree object 
> representation which is not deflated at all.
Well, the thing is, for things that really don't compress, zlib shouldn't 
add much of an overhead on uncompression. It *should* just end up being a 
single "memcpy()" after you've done:
 - check the header for size and mode ("plain data")
 - check the adler checksum (which is *really* nice - we've found real 
   corruption this way!).

The adler32 checksumming may sound unnecessary when you already have the SHA1 checksum, but the thing is, we normally don't actually *check* the SHA1 except when doing a full fsck. So I actually like the fact that object unpacking always checks at least the adler32 checksum at each stage, which you get "for free" when you use zlib.

So not using compression at all actually not only gets rid of the compression, it gets rid of a good safety valve - something that may not be immediately obvious when you don't think about what all zlib entails.

People think of zlib as just compressing, but I think the checksumming is almost as important, which is why it isn't an obviously good thing to not compress small objects just because you don't win on size!

Remember: stability and safety of the data is *the* #1 objective here. The 
git SHA1 checksums guarantees that we can find any corruption, but in 
every-day git usage, the adler32 checksum is the one that generally would 
*notice* the corruption and cause us to say "uhhuh, need to fsck".

Everything else is totally secondary to the goal of "your data is secure". Yes, performance is a primary goal too, but it's always "performance with correctness guarantees"!

But I just traced through a simple 60-byte incompressible zlib thing. It's painful. This should be *the* simplest case, and it should really just be the memcpy and the adler32 check. But:

	[torvalds@woody ~]$ grep '<inflate' trace | wc -l
	460
	[torvalds@woody ~]$ grep '<adler32' trace | wc -l
	403
	[torvalds@woody ~]$ grep '<memcpy' trace | wc -l
	59

ie we spend *more* instructions on just the stupid setup in "inflate()" than we spend on the adler32 (or, obviously, on the actual 60-byte memcpy of the actual incompressible data)

I dunno. I don't mind the adler32 that much. The rest seems to be pretty annoying, though.

		Linus
Previous: Nicolas PitreNext: Nicolas Pitre
Message 76 of 79 in “cleaner/better zlib sources?”
  1. Linus TorvaldsMar 16, 2007
  2. Shawn O. PearceMar 16, 2007
  3. Jeff GarzikMar 16, 2007
  4. Matt MackallMar 16, 2007
  5. Linus TorvaldsMar 16, 2007
  6. Linus TorvaldsMar 16, 2007
  7. Davide LibenziMar 16, 2007
  8. Linus TorvaldsMar 16, 2007
  9. Davide LibenziMar 16, 2007
  10. Linus TorvaldsMar 16, 2007
  11. Davide LibenziMar 16, 2007
  12. Linus TorvaldsMar 16, 2007
  13. Davide LibenziMar 16, 2007
  14. Linus TorvaldsMar 17, 2007
  15. Linus TorvaldsMar 17, 2007
  16. Nicolas PitreMar 17, 2007
  17. Shawn O. PearceMar 17, 2007
  18. Linus TorvaldsMar 17, 2007
  19. Linus TorvaldsMar 17, 2007
  20. 1/2 Make trivial wrapper functions around delta base generation and freeingLinus Torvalds, Mar 17, 2007
  21. 2/2 Implement a simple delta_base cacheLinus Torvalds, Mar 17, 2007
  22. Linus TorvaldsMar 17, 2007
  23. Junio C HamanoMar 17, 2007
  24. Linus TorvaldsMar 17, 2007
  25. Linus TorvaldsMar 17, 2007
  26. Nicolas PitreMar 18, 2007
  27. Junio C HamanoMar 18, 2007
  28. Junio C HamanoMar 17, 2007
  29. Linus TorvaldsMar 17, 2007
  30. Jon SmirlMar 17, 2007
  31. Morten WelinderMar 18, 2007
  32. Linus TorvaldsMar 18, 2007
  33. Nicolas PitreMar 18, 2007
  34. Linus TorvaldsMar 18, 2007
  35. Nicolas PitreMar 18, 2007
  36. Linus TorvaldsMar 18, 2007
  37. Nicolas PitreMar 18, 2007
  38. Linus TorvaldsMar 18, 2007
  39. Julian PhillipsMar 18, 2007
  40. Linus TorvaldsMar 18, 2007
  41. Robin RosenbergMar 18, 2007
  42. Linus TorvaldsMar 18, 2007
  43. Robin RosenbergMar 18, 2007
  44. Shawn O. PearceMar 18, 2007
  45. David BrodskyMar 19, 2007
  46. Robin RosenbergMar 20, 2007
  47. David BrodskyMar 20, 2007
  48. Linus TorvaldsMar 21, 2007
  49. Nicolas PitreMar 21, 2007
  50. 3/2 Avoid unnecessary strlen() callsLinus Torvalds, Mar 18, 2007
  51. Junio C HamanoMar 18, 2007
  52. Linus TorvaldsMar 18, 2007
  53. Linus TorvaldsMar 18, 2007
  54. Shawn O. PearceMar 18, 2007
  55. Linus TorvaldsMar 18, 2007
  56. Johannes SchindelinMar 20, 2007
  57. Shawn O. PearceMar 20, 2007
  58. Shawn O. PearceMar 20, 2007
  59. Linus TorvaldsMar 20, 2007
  60. Shawn O. PearceMar 20, 2007
  61. Linus TorvaldsMar 20, 2007
  62. Junio C HamanoMar 20, 2007
  63. Junio C HamanoMar 20, 2007
  64. Linus TorvaldsMar 20, 2007
  65. Shawn O. PearceMar 20, 2007
  66. Linus TorvaldsMar 20, 2007
  67. Linus TorvaldsMar 18, 2007
  68. Avi KivityMar 18, 2007
  69. Linus TorvaldsMar 17, 2007
  70. Jeff GarzikMar 16, 2007
  71. Matt MackallMar 16, 2007
  72. Linus TorvaldsMar 16, 2007
  73. Nicolas PitreMar 16, 2007
  74. Shawn O. PearceMar 16, 2007
  75. Nicolas PitreMar 16, 2007
  76. Linus TorvaldsMar 16, 2007
  77. Nicolas PitreMar 16, 2007
  78. Davide LibenziMar 16, 2007
  79. Davide LibenziMar 16, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.