git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 2/2] index-v4: document the entry format

From
Thomas Rast <trast@student.ethz.ch>
Date
Apr 30, 2012, 17:20 UTC
Message-ID
<87vckhuofj.fsf@thomas.inf.ethz.ch>
In-Reply-To
<xmqqpqas93sa.fsf_-_@junio.mtv.corp.google.com>
Hi Junio,
I seem to have completely missed the earlier series at
  http://thread.gmane.org/gmane.comp.version-control.git/194660
My bad.

Thomas has been working on a prototype converter over the past few days, with results similar to (but not quite as good as) your numbers

    $ ls -l .git/index*
    -rw-r----- 1 jch eng 25586488 2012-04-03 15:27 .git/index
    -rw-r----- 1 jch eng 14654328 2012-04-03 15:38 .git/index-4
while taking a different approach with different tradeoffs.
Nevertheless...
Show 9 quoted lines
> +  (Version 4) In version 4, the entry path name is prefix-compressed
> +    relative to the path name for the previous entry (the very first
> +    entry is encoded as if the path name for the previous entry is an
> +    empty string).  At the beginning of an entry, an integer N in the
> +    variable width encoding (the same encoding as the offset is encoded
> +    for OFS_DELTA pack entries; see pack-format.txt) is stored, followed
> +    by a NUL-terminated string S.  Removing N bytes from the end of the
> +    path name for the previous entry, and replacing it with the string S
> +    yields the path name for this entry.
[..]
> +  (Version 4) In version 4, the padding after the pathname does not
> +  exist.
I think there are actually several separate ideas here:
* The prefix compression.  Thomas is not using this idea; we've been
  toying with making the index bisectable (within each directory) for
  fast single-entry lookups, which inherently conflicts with this.  The
  directory-like layout partially achieves the same (elides common path
  components).
* The varint encoding (or offset encoding, but "varint" is something you
  can google :-).  David suggested using it on stat() data, combined
  with zigzag encoding and delta against the first entry in the
  directory, which gives some good compression results.  Profiling will
  have to say whether the extra decoding effort is worth the space
  savings.
* The lack of variable padding, which is a good idea -- in any case I
  seem to remember Shawn complaining about it.
-- 
Thomas Rast
trast@{inf,student}.ethz.ch
Previous: Junio C HamanoNext: Junio C Hamano
Message 22 of 27 in “Prefix-compress on-disk index entries”
  1. 0/9 Prefix-compress on-disk index entriesJunio C Hamano, Apr 3, 2012
  2. 1/9 varint: make it available outside the context of packJunio C Hamano, Apr 3, 2012
  3. 2/9 cache.h: hide on-disk index detailsJunio C Hamano, Apr 3, 2012
  4. 3/9 read-cache.c: allow unaligned mapping of the index fileJunio C Hamano, Apr 3, 2012
  5. 4/9 read-cache.c: make create_from_disk() report number of bytes it consumedJunio C Hamano, Apr 3, 2012
  6. 5/9 read-cache.c: report the header version we do not understandJunio C Hamano, Apr 3, 2012
  7. 6/9 read-cache.c: move code to copy ondisk to incore cache to a helper functionJunio C Hamano, Apr 3, 2012
  8. 7/9 read-cache.c: move code to copy incore to ondisk cache to a helper functionJunio C Hamano, Apr 3, 2012
  9. 8/9 read-cache.c: read prefix-compressed names in index on-disk version v4Junio C Hamano, Apr 3, 2012
  10. 9/9 read-cache.c: write index v4 formatJunio C Hamano, Apr 3, 2012
  11. David BarrApr 4, 2012
  12. Junio C HamanoApr 4, 2012
  13. Junio C HamanoApr 4, 2012
  14. 2/2 update-index: upgrade/downgrade on-disk index versionJunio C Hamano, Apr 4, 2012
  15. Nguyen Thai Ngoc DuyApr 4, 2012
  16. Junio C HamanoApr 4, 2012
  17. David BarrApr 6, 2012
  18. Nguyen Thai Ngoc DuyMay 2, 2012
  19. David BarrMay 2, 2012
  20. 1/2 unpack-trees: preserve the index file version of originalJunio C Hamano, Apr 27, 2012
  21. 2/2 index-v4: document the entry formatJunio C Hamano, Apr 27, 2012
  22. Thomas RastApr 30, 2012
  23. Junio C HamanoMay 1, 2012
  24. Thomas RastMay 1, 2012
  25. Shawn PearceMay 2, 2012
  26. Junio C HamanoMay 2, 2012
  27. Shawn PearceMay 2, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.