Re: Compression and dictionaries
- From
Shawn Pearce <spearce@spearce.org>
- Date
- Aug 14, 2006, 04:17 UTC
- Message-ID
- <20060814041705.GD18667@spearce.org>
- In-Reply-To
- <9e4733910608132107j7bca0271g360de3447febbf51@mail.gmail.com>
Jon Smirl <jonsmirl@gmail.com> wrote:
Show 5 quoted lines
> The zlib doc says to put your most common strings into the fixed > dictionary. If a string isn't in the fixed dictionary it will get > handled with an internal dictionary entry. By default zlib runs with > an empty fixed dictionary and handles everything with the internal > dictionary.
> Since we are encoding C many strings will always be present (if, > static, define, const, char, include, int, void, while, continue, > etc). Do you have any tools to identify the top 500 strings in C > code? The fixed dictionary would get hardcoded into the git apps.
Actually GIT itself may also benefit from other strings beyond those common found in C-like languages:
'10644 ' '40000 ' 'parent ' 'tree ' 'author ' 'committer '
as these occur frequently in trees and commits.
> A fixed dictionary could conceivably take 5-10% off the size of each entry.
Could be an interesting experiment to see if that's really true for common loads (e.g. the kernel repo). I don't think anyone has tried it.
-- Shawn.