git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: [PATCH 3/6] Stop producing index version 2

From
Thomas Rast <trast@inf.ethz.ch>
Date
Feb 7, 2012, 17:25 UTC
Message-ID
<874nv2o8rs.fsf@thomas.inf.ethz.ch>
In-Reply-To
<CAJo=hJvtRnmvALcn3vKpYTr3j6ada8iboPjWN3cQnwwKzRvrDA@mail.gmail.com>
Shawn Pearce <spearce@spearce.org> writes:
Show 6 quoted lines
> I have long wanted to scrap the current index format. I unfortunately
> don't have the time to do it myself. But I suspect there may be a lot
> of gains by making the index format match the canonical tree format
> better by keeping the tree structure within a single file stream,
> nesting entries below their parent directory, and keeping tree SHA-1
> data along with the directory entry.

If I may add to this: the one thing that I would like to see fixed about the index is that it's flat out impossible to change a single thing in it without re"writing" it from scratch.

I'm saying "writing" because it is possible to change a few things around, but recomputing the trailing SHA1 swamps that by a large margin unless you are writing to a floppy disk, so it doesn't matter. I'm sure using a CRC32 helps here, but if we're going to make an incompatible change, why not go all the way?

A tree layout can fix that if it is properly arranged so that if you 'git add path/to/file', it only updates the SHA1s for path/to/file, path/to and path. For this to work, the checks would have to correspond to the trees, perhaps even directly use the actual tree SHA1. This would at least be natural in some sense; getting to actual log(n) complexity for hilariously large directories would require dynamically splitting directories where appropriate.

Along the same lines the format should allow for changing the extension data for a single extension while only rehashing the new data.

When I worked on cache-tree, I considered making a change to the latter effect, but thought the impact too great for a little gain. Now from this thread, I'm getting the impression that such a change would be ok, even if users would have to scrap the index if they downgrade. Is that right?

-- 
Thomas Rast
trast@{inf,student}.ethz.ch
Previous: Junio C HamanoNext: Nguyễn Thái Ngọc Duy
Message 9 of 20 in “read-cache: use sha1file for sha1 calculation”
  1. 1/6 read-cache: use sha1file for sha1 calculationNguyễn Thái Ngọc Duy, Feb 6, 2012
  2. 2/6 csum-file: make sha1 calculation optionalNguyễn Thái Ngọc Duy, Feb 6, 2012
  3. 3/6 Stop producing index version 2Nguyễn Thái Ngọc Duy, Feb 6, 2012
  4. Junio C HamanoFeb 6, 2012
  5. Shawn PearceFeb 7, 2012
  6. Nguyen Thai Ngoc DuyFeb 7, 2012
  7. Nguyen Thai Ngoc DuyFeb 7, 2012
  8. Junio C HamanoFeb 7, 2012
  9. Thomas RastFeb 7, 2012
  10. 4/6 Introduce index version 4 with global flagsNguyễn Thái Ngọc Duy, Feb 6, 2012
  11. 5/6 Allow to use crc32 as a lighter checksum on indexNguyễn Thái Ngọc Duy, Feb 6, 2012
  12. Shawn PearceFeb 7, 2012
  13. Dave ZarzyckiFeb 7, 2012
  14. Dave ZarzyckiFeb 7, 2012
  15. 6/6 Automatically switch to crc32 checksum for index when it's too largeNguyễn Thái Ngọc Duy, Feb 6, 2012
  16. Dave ZarzyckiFeb 6, 2012
  17. Nguyen Thai Ngoc DuyFeb 6, 2012
  18. Dave ZarzyckiFeb 6, 2012
  19. Junio C HamanoFeb 6, 2012
  20. Nguyen Thai Ngoc DuyFeb 6, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.