git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Compression and dictionaries

From
Jon Smirl <jonsmirl@gmail.com>
Date
Aug 14, 2006, 16:15 UTC
Message-ID
<9e4733910608140915i728004c1p216bf3d74fcc6ab7@mail.gmail.com>
In-Reply-To
<Pine.LNX.4.63.0608141641330.28360@wbgn013.biozentrum.uni-wuerzburg.de>
On 8/14/06, Johannes Schindelin <Johannes.Schindelin@gmx.de> wrote:
> I still think that this is important to think through: Is it worth a
> couple of kilobytes (I doubt that it would be as much as 1MB in _total_),
> and be on the unsafe side?

The maps look something like this: void => 10101010 char => 111101 int => 1110101 tree => 11010111 commit => 101011001

These maps are repeated in every one of my 1M revisions including the deltas. I have 1GB pack files with 1M entries in them - 1K each entry. Each byte saved out of a zlib entry take 1MB off my pack.

Note that the current internal maps aren't the same in each of each of the zlib blobs since the algorithm that builds the internal maps depends on the order the identifiers were encountered.

If the git tools add a global dictionary the tools would still be able to read existing packs. If old tools try to read a new dictionary based pack they will get the zlib NEED_DICT error.

If the entire file was one big zlib blob there would only be one dictionary and adding a fixed dictionary wouldn't make any difference. But since it is 1M little zlib blobs it is has 1M dictionaries.

The only "unsafe" aspect I see to this is if the global dictionary doesn't contain any of the words in the documents being encoded. In that case the global dictionary will occupy the short huffman keys forcing longer internal keys. The keys for the words in the document would be longer by a about a bit on average.

A solution for making this work over time would be to store the global dictionary at the front of the pack file and for the unpack tools to use the stored copy. This would let us change the global dictionary in the pack tool with no downside, you could even support multiple dictionaries in the pack tool.

If someone wants to get fancy you could write a tool that would scan a pack file and compute an optimal fixed dictionary. Store it at the front of the pack file and repack using it.

Global dictionaries are common in full text searching. I seem to recall an article stating that Google's global dictionary has about 250K entries in it. If git packs switch to a global dictionary model it's not a big leap to add a full text search index. You just need objects for each word in the dictionary pointing to the revisions that contain it.

-- 
Jon Smirl
jonsmirl@gmail.com
Previous: Johannes SchindelinNext: David Lang
Message 10 of 20 in “Compression and dictionaries”
  1. Jon SmirlAug 14, 2006
  2. Shawn PearceAug 14, 2006
  3. Jon SmirlAug 14, 2006
  4. Shawn PearceAug 14, 2006
  5. Alex RiesenAug 14, 2006
  6. Erik MouwAug 14, 2006
  7. Johannes SchindelinAug 14, 2006
  8. Jon SmirlAug 14, 2006
  9. Johannes SchindelinAug 14, 2006
  10. Jon SmirlAug 14, 2006
  11. David LangAug 14, 2006
  12. Jakub NarebskiAug 14, 2006
  13. Jeff GarzikAug 14, 2006
  14. David LangAug 14, 2006
  15. Jeff GarzikAug 14, 2006
  16. Jon SmirlAug 14, 2006
  17. David LangAug 14, 2006
  18. Johannes SchindelinAug 14, 2006
  19. Alex RiesenAug 14, 2006
  20. Johannes SchindelinAug 14, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.