git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Calculating tree nodes

From
Jon Smirl <jonsmirl@gmail.com>
Date
Sep 4, 2007, 05:50 UTC
Message-ID
<9e4733910709032250r1198379cmafd4e14fd330513b@mail.gmail.com>
In-Reply-To
<7vmyw3835y.fsf@gitster.siamese.dyndns.org>
On 9/4/07, Junio C Hamano <gitster@pobox.com> wrote:
Show 21 quoted lines
> "Jon Smirl" <jonsmirl@gmail.com> writes:
>
> >> Yes.  For performance reasons, since a simple commit would kill you in any
> >> reasonably sized repo.
> >
> > That's not an obvious conclusion. A new commit is just a series of
> > edits to the previous commit. Start with the previous commit, edit it,
> > delta it and store it. Storing of the file objects is the same. Why
> > isn't this scheme fast than the current one?
>
> I think you seem to be forgetting about tree comparison.
>
> With a large project that has a reasonable directory structure
> (i.e. not insanely narrow), a commit touches isolated subparts
> of the whole tree.  Think of an architecture specific patch to
> the Linux kernel touching only include/asm-i386 and arch/i386
> directories.
>
> Being able to cull an entire subdirectory (e.g. drivers/ which
> has 5700 files underneath) by only looking at the tree SHA-1 of
> the containing tree is a _HUGE_ win.

In my scheme you have all of the SHAs for the commit in RAM because the are contained in the commit and you have the commit in RAM. It take microseconds to compare these two lists in RAM.

The current scheme is doing disk accesses to get those tree nodes so of course it is a win to cull the 5700 files.

Show 12 quoted lines
>
> And this is not just about two tree comparison.  When you say:
>
>         git log v2.6.20 -- arch/i386/
>
> what you are seeing is a simplified history that consists of
> commits that touch only these paths.  How would we determine if
> a commit touch these paths efficiently?  By comparing the "i386"
> entry in tree objects for $commit^:arch and $commit:arch.  You
> do not have to look inside arch/i386/ trees to see if any of the
> 330 files in it is different.  You just check a single SHA-1
> pair.

It's more than just comparing a SHA, you have to do disk accesses to retrieve the SHA.

I'm proposing that we only really need commit and file objects. I also mentioned that if you think of the file objects as a table you could use triggers to build cached indexes. To get performance back to the current level we may want to construct some of these indexes. We need to explore the scheme more before we can figure out the best cached indexes to build.

Right now we only have a single index type, the tree nodes. And it's a permanent part of the storage not cached. A hierarchical index is not very useful of indexing non file name attributes.

-- 
Jon Smirl
jonsmirl@gmail.com
Previous: Junio C HamanoNext: David Tweed
Message 15 of 27 in “Calculating tree nodes”
  1. Jon SmirlSep 4, 2007
  2. Shawn O. PearceSep 4, 2007
  3. Jon SmirlSep 4, 2007
  4. Johannes SchindelinSep 4, 2007
  5. Jon SmirlSep 4, 2007
  6. Martin LanghoffSep 4, 2007
  7. Jon SmirlSep 4, 2007
  8. Andreas EricssonSep 4, 2007
  9. Johannes SchindelinSep 4, 2007
  10. Jon SmirlSep 4, 2007
  11. Johannes SchindelinSep 4, 2007
  12. Andreas EricssonSep 4, 2007
  13. Martin LanghoffSep 4, 2007
  14. Junio C HamanoSep 4, 2007
  15. Jon SmirlSep 4, 2007
  16. David TweedSep 4, 2007
  17. Jon SmirlSep 4, 2007
  18. Andreas EricssonSep 4, 2007
  19. Shawn O. PearceSep 4, 2007
  20. Jon SmirlSep 4, 2007
  21. Andreas EricssonSep 4, 2007
  22. David TweedSep 4, 2007
  23. Shawn O. PearceSep 4, 2007
  24. Junio C HamanoSep 4, 2007
  25. Shawn O. PearceSep 6, 2007
  26. Junio C HamanoSep 6, 2007
  27. Daniel HulmeSep 4, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.