git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: cleaner/better zlib sources?

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
Mar 16, 2007, 01:46 UTC
Message-ID
<Pine.LNX.4.64.0703151822490.3816@woody.linux-foundation.org>
In-Reply-To
<45F9EED5.3070706@garzik.org>
On Thu, 15 Mar 2007, Jeff Garzik wrote:
> 
> ISTR there are a bunch of state transitions per byte, which would make sense
> that it shows up on profiles.

Yeah. I only looked at the kernel sources (which include a cleaned-up older version of zlib), but they seemed to match the disassembly that I saw when doing the instruction-level profiling.

The code *sometimes* falls through from one state to another, ie there's a case where zlib does:

	...
        case DICTID:
            NEEDBITS(32);
            strm->adler = state->check = REVERSE(hold);
            INITBITS();
            state->mode = DICT;
        case DICT:
	...

which will obviously generate fine code. There's a few other examples of that, *BUT* most of the stuff is just horrid, like

        case HEAD:
            if (state->wrap == 0) {
                state->mode = TYPEDO;
                break;
            }

where the "break" (which simply breaks out of the case-statement, and thus just loops back to it:

    for (;;)
        switch (state->mode) {
thing is nasty nasty nasty).
A trivial thing to do is to just replace such
	state->mode = xxx;
	break;
with
	state->mode = xxx;
	goto xxx_mode;

and all of that complicated run-time code *just*goes*away* and is replaced by a relative no-op (ie an unconditional direct jump).

Some of them are slightly more complex, like
	state->mode = hold & 0x200 ? DICTID : TYPE;
	INITBITS();
	break;
which would need to be rewritten as
	old_hold = hold;
	INITBITS();
	if (old_hold & 0x200) {
		state->mode = DICTID;
		goto dictid_mode;
	}
	state->mode = TYPE;
	goto dictid_mode;

but while that looks more complicated on a source code level it's a *lot* easier for a CPU to actually execute.

Same obvious performance problems go for
	case STORED:
	   ...
        case COPY:
            copy = state->length;
            if (copy) {
		...
		state->length -= copy;
		break;
	    }
	    Tracev((stderr, "inflate:       stored end\n"));
	    state->mode = TYPE;
	    break;

notice how when you get to the COPY state it will actually have nicely fallen through from STORED (one of the places where it does that), but then it will go through the expensive indirect jump not just once, but at least *TWICE* just to get to the TYPE thing, because you'll first have to re-enter COPY with a count of zero to get to the place where it sets TYPE, and then does the indirect jump immediately again!

In other words, the code is *incredibly* badly written from a performance angle.

Yeah, a perfect compiler could do this all for us even with unchanged zlib source code, but (a) gcc isn't perfect and (b) I don't really think anything else is either, although if things like this happen as part of gzip in SpecInt, I wouldn't be surprised by compilers doing things like this just to look good ;)

Especially with something like git, where we know that most of the time we have the full input buffer (so inflate() generally won't be called in the "middle" of a inflate event), we probably have a very well-defined start state too, so we'd actually predict perfectlyon the *first* indirect jump if we just never did any more of them.

> I could have sworn that either Matt Mackall or Ben LaHaise had cleaned up the
> existing zlib so much that it was practically a new implementation.  I'm not
> aware of any open source implementations independent of zlib (except maybe
> that C++ behemoth, 7zip).

I looked at the latest release (1.2.3), so either Matt/Ben hasn't been merged, or this wasn't part of the thing they looked at.

			Linus
Previous: Matt MackallNext: Linus Torvalds
Message 5 of 79 in “cleaner/better zlib sources?”
  1. Linus TorvaldsMar 16, 2007
  2. Shawn O. PearceMar 16, 2007
  3. Jeff GarzikMar 16, 2007
  4. Matt MackallMar 16, 2007
  5. Linus TorvaldsMar 16, 2007
  6. Linus TorvaldsMar 16, 2007
  7. Davide LibenziMar 16, 2007
  8. Linus TorvaldsMar 16, 2007
  9. Davide LibenziMar 16, 2007
  10. Linus TorvaldsMar 16, 2007
  11. Davide LibenziMar 16, 2007
  12. Linus TorvaldsMar 16, 2007
  13. Davide LibenziMar 16, 2007
  14. Linus TorvaldsMar 17, 2007
  15. Linus TorvaldsMar 17, 2007
  16. Nicolas PitreMar 17, 2007
  17. Shawn O. PearceMar 17, 2007
  18. Linus TorvaldsMar 17, 2007
  19. Linus TorvaldsMar 17, 2007
  20. 1/2 Make trivial wrapper functions around delta base generation and freeingLinus Torvalds, Mar 17, 2007
  21. 2/2 Implement a simple delta_base cacheLinus Torvalds, Mar 17, 2007
  22. Linus TorvaldsMar 17, 2007
  23. Junio C HamanoMar 17, 2007
  24. Linus TorvaldsMar 17, 2007
  25. Linus TorvaldsMar 17, 2007
  26. Nicolas PitreMar 18, 2007
  27. Junio C HamanoMar 18, 2007
  28. Junio C HamanoMar 17, 2007
  29. Linus TorvaldsMar 17, 2007
  30. Jon SmirlMar 17, 2007
  31. Morten WelinderMar 18, 2007
  32. Linus TorvaldsMar 18, 2007
  33. Nicolas PitreMar 18, 2007
  34. Linus TorvaldsMar 18, 2007
  35. Nicolas PitreMar 18, 2007
  36. Linus TorvaldsMar 18, 2007
  37. Nicolas PitreMar 18, 2007
  38. Linus TorvaldsMar 18, 2007
  39. Julian PhillipsMar 18, 2007
  40. Linus TorvaldsMar 18, 2007
  41. Robin RosenbergMar 18, 2007
  42. Linus TorvaldsMar 18, 2007
  43. Robin RosenbergMar 18, 2007
  44. Shawn O. PearceMar 18, 2007
  45. David BrodskyMar 19, 2007
  46. Robin RosenbergMar 20, 2007
  47. David BrodskyMar 20, 2007
  48. Linus TorvaldsMar 21, 2007
  49. Nicolas PitreMar 21, 2007
  50. 3/2 Avoid unnecessary strlen() callsLinus Torvalds, Mar 18, 2007
  51. Junio C HamanoMar 18, 2007
  52. Linus TorvaldsMar 18, 2007
  53. Linus TorvaldsMar 18, 2007
  54. Shawn O. PearceMar 18, 2007
  55. Linus TorvaldsMar 18, 2007
  56. Johannes SchindelinMar 20, 2007
  57. Shawn O. PearceMar 20, 2007
  58. Shawn O. PearceMar 20, 2007
  59. Linus TorvaldsMar 20, 2007
  60. Shawn O. PearceMar 20, 2007
  61. Linus TorvaldsMar 20, 2007
  62. Junio C HamanoMar 20, 2007
  63. Junio C HamanoMar 20, 2007
  64. Linus TorvaldsMar 20, 2007
  65. Shawn O. PearceMar 20, 2007
  66. Linus TorvaldsMar 20, 2007
  67. Linus TorvaldsMar 18, 2007
  68. Avi KivityMar 18, 2007
  69. Linus TorvaldsMar 17, 2007
  70. Jeff GarzikMar 16, 2007
  71. Matt MackallMar 16, 2007
  72. Linus TorvaldsMar 16, 2007
  73. Nicolas PitreMar 16, 2007
  74. Shawn O. PearceMar 16, 2007
  75. Nicolas PitreMar 16, 2007
  76. Linus TorvaldsMar 16, 2007
  77. Nicolas PitreMar 16, 2007
  78. Davide LibenziMar 16, 2007
  79. Davide LibenziMar 16, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.