git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Calculating tree nodes

From
Jon Smirl <jonsmirl@gmail.com>
Date
Sep 4, 2007, 03:26 UTC
Message-ID
<9e4733910709032026s7f94eed9h25d5165840cc38d2@mail.gmail.com>
In-Reply-To
<20070904025153.GS18160@spearce.org>
Are tree objects really needed?
1) Make the path an attribute of the file object.
2) Commits are simply a list of all the objects that make up the commit.
Sort the SHAs in the commit and delta them.

This is something that has always bugged me about file systems. File systems force hierarchical naming due to their directory structure. There is no reason they have to work that way. Google is an example of a giant file system that works just fine without hierarchical directories. The full path should be just another attribute on the file. If you want a hierarchical index into the file system you can generate it by walking the files or using triggers. But you could also delete the hierarchical directory and replace it with something else like a full text index. Directories would become a computationally generated cache, not a critical part of the file system. But this is a git list so I shouldn't go too far off into file system design.

Git has picked up the hierarchical storage scheme since it was built on a hierarchical file system. I don't this this is necessarily a good thing moving forward.

If we really need tree objects they could become a new class of computationally generated objects that could be deleted out of the database at any time and recreated. For example if you think of the file objects as being in a table, inserting a new row into this table would compute new tree objects (an index).

Index is the key here, we may want other kinds of indexes in the future. It was the mail about auto-generating the Maintainers list that caused me to think about this. If file objects are a table with triggers, building a hierarchical index for the Maintainers field doesn't make sense.

These are just some initial thoughts on a different way to view the data git is storing. Thinking about it as a database with fields and indexes built via triggers may change the way we want to structure things.

-- 
Jon Smirl
jonsmirl@gmail.com
Previous: Shawn O. PearceNext: Johannes Schindelin
Message 3 of 27 in “Calculating tree nodes”
  1. Jon SmirlSep 4, 2007
  2. Shawn O. PearceSep 4, 2007
  3. Jon SmirlSep 4, 2007
  4. Johannes SchindelinSep 4, 2007
  5. Jon SmirlSep 4, 2007
  6. Martin LanghoffSep 4, 2007
  7. Jon SmirlSep 4, 2007
  8. Andreas EricssonSep 4, 2007
  9. Johannes SchindelinSep 4, 2007
  10. Jon SmirlSep 4, 2007
  11. Johannes SchindelinSep 4, 2007
  12. Andreas EricssonSep 4, 2007
  13. Martin LanghoffSep 4, 2007
  14. Junio C HamanoSep 4, 2007
  15. Jon SmirlSep 4, 2007
  16. David TweedSep 4, 2007
  17. Jon SmirlSep 4, 2007
  18. Andreas EricssonSep 4, 2007
  19. Shawn O. PearceSep 4, 2007
  20. Jon SmirlSep 4, 2007
  21. Andreas EricssonSep 4, 2007
  22. David TweedSep 4, 2007
  23. Shawn O. PearceSep 4, 2007
  24. Junio C HamanoSep 4, 2007
  25. Shawn O. PearceSep 6, 2007
  26. Junio C HamanoSep 6, 2007
  27. Daniel HulmeSep 4, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.