git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git's database structure

From
KMKyle Moffett <mrmacman_g4@mac.com>
Date
Sep 6, 2007, 01:27 UTC
Message-ID
<ED15F422-C83D-4749-8B0A-24F862AB1940@mac.com>
In-Reply-To
<Pine.LNX.4.64.0709051823470.26016@reaper.quantumfyre.co.uk>
On Sep 05, 2007, at 13:31:43, Julian Phillips wrote:
Show 10 quoted lines
> And this is advantaged by having the path in the blob how?  The  
> important information here is knowing which commits touched the  
> file - this information is expensive in git because it is snapshot  
> based.  You have to go back through all the commits looking for  
> changes to the given path. The information you might want to cache  
> is which commits touched the file, which you could do without  
> changing the current data storage. Presumably you are suggesting  
> that such a cache would be cleaner with the filename in the blob?   
> Or do you think that it would somehow be faster to create?  If so,  
> how?

The only possible reason I can think of for moving data into the blob would be to make a POSIX-compliant git-like filesystem, and EVEN THEN you would NOT move the path out of the tree objects. In order to have somewhat consistent inodes (and also for performance when changing 4 bytes in a 40GB file) you would want to have 3 different types of "inode" objects:

1)  4-64k of (metadata + filedata)
2)  4-64k of (metadata + list of 4-64k filedata blobs)
3)  4-64k of (metadata + list of 4-64k lists of filedata blobs)

On the other hand... that isn't GIT, it's something completely different with a very different usage pattern and set of requirements. And you still don't put the path name in the objects, just the permissions and other attributes/metadata.

<Random Thought Experiment> You would of course want to better define those 4-64k limits for allocation and performance reasons, but a double-indirect table of SHA128s with 64kb chunks lets you address up to 1TB of file data, and for each additional power-of-two increase in the chunk size you get 8 times the storage space. Furthermore, the actual double-indirect tables for an 8TB file using 128k chunks would be all of 64MB, for a more reasonable 4GB file with 32k tables (max of 128GB) it would be maybe 128kB of indirect SHA1 hash tables. </Random Thought Experiment>

Cheers, Kyle Moffett

Previous: Julian PhillipsNext: Mike Hommey
Message 28 of 39 in “Git's database structure”
  1. Jon SmirlSep 4, 2007
  2. Andreas EricssonSep 4, 2007
  3. Mike HommeySep 4, 2007
  4. Andreas EricssonSep 4, 2007
  5. Jon SmirlSep 4, 2007
  6. Andreas EricssonSep 4, 2007
  7. Jeff KingSep 4, 2007
  8. David TweedSep 4, 2007
  9. Junio C HamanoSep 4, 2007
  10. Jon SmirlSep 4, 2007
  11. Andreas EricssonSep 4, 2007
  12. Jon SmirlSep 4, 2007
  13. Andreas EricssonSep 4, 2007
  14. Junio C HamanoSep 4, 2007
  15. Jon SmirlSep 4, 2007
  16. Mike HommeySep 4, 2007
  17. Reece DunnSep 4, 2007
  18. Junio C HamanoSep 4, 2007
  19. Theodore TsoSep 4, 2007
  20. Jon SmirlSep 4, 2007
  21. Andreas EricssonSep 5, 2007
  22. Jon SmirlSep 5, 2007
  23. Andreas EricssonSep 5, 2007
  24. Jon SmirlSep 5, 2007
  25. Julian PhillipsSep 5, 2007
  26. Jon SmirlSep 5, 2007
  27. Julian PhillipsSep 5, 2007
  28. Kyle MoffettSep 6, 2007
  29. Mike HommeySep 5, 2007
  30. Andreas EricssonSep 6, 2007
  31. Junio C HamanoSep 6, 2007
  32. Wincent ColaiutaSep 6, 2007
  33. Johannes SchindelinSep 6, 2007
  34. Steven GrimmSep 6, 2007
  35. Martin LanghoffSep 7, 2007
  36. Andy ParkinsSep 5, 2007
  37. Julian PhillipsSep 4, 2007
  38. Jon SmirlSep 4, 2007
  39. Andreas EricssonSep 4, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.