git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: jgit performance update

From
Shawn Pearce <spearce@spearce.org>
Date
Dec 3, 2006, 22:59 UTC
Message-ID
<20061203225947.GD15965@spearce.org>
In-Reply-To
<200612031455.48032.robin.rosenberg.lists@dewire.com>
Robin Rosenberg <robin.rosenberg.lists@dewire.com> wrote:
Show 5 quoted lines
> So, just go on to the next case. I added filtering on filenames (yes, 
> CVS-induced brain damage, I should track the content. next version. filenames 
> are so much handier to work with). That gives me 4.5s to retrieve a filtered 
> history (from 10800 commits).Half of the time is spent in re-sorting tree 
> entries. Is that really necessary?

Yea, I was looking at that code while doing the other performance improvements and thought it might start to become a bottleneck. I guess I was right.

What is happening here is jgit wants to store the items in the tree in name ordering, but Git stores the items in the tree sorted such that subtrees sort with a '/' on the end of their name. This is a different ordering...

The reason I'm resorting them is so we can find an entry without
knowing what its type is first.  Looks like that's going to have
to change somehow.
 
> Most of java's slowness comes from the programmers using it. (Lutz Prechelt. 
> Technical opinion: comparing Java vs. C/C++ efficiency differences to 
> interpersonal differences. ACM, Vol 42,#10, 1999)
Yes, that was clearly the case here with jgit!  :-)
_This_ programmer made jgit slow.  Learned from the mistake, and
made it faster.
 
Show 12 quoted lines
> > One of the biggest annoyances has been the fact that although Java 
> > 1.4 offers a way to mmap a file into the process, the overhead to
> > access that data seems to be far higher than just reading the file
> > content into a very large byte array, especially if we are going
> > to access that file content multiple times.  So jgit performs worse
> > than core Git early on while it copies everything from the OS buffer
> > cache into the Java process, but then performs reasonably well once
> > the internal cache is hot.  On the other hand using the mmap call
> > reduces early latency but hurts the access times so much that we're
> > talking closer to 3s average read times for the same log operation.
> 
> Have you tried that with difference JVM's?

No, I'm on Mac OS X so I don't have a huge JVM selection (that I know of). And I haven't tried jgit or egit on any other system yet.

Previous: Shawn PearceNext: Linus Torvalds
Message 6 of 16 in “jgit performance update”
  1. Shawn PearceDec 3, 2006
  2. Robin RosenbergDec 3, 2006
  3. Jakub NarebskiDec 3, 2006
  4. Robin RosenbergDec 3, 2006
  5. Shawn PearceDec 3, 2006
  6. Shawn PearceDec 3, 2006
  7. Linus TorvaldsDec 3, 2006
  8. Jakub NarebskiDec 3, 2006
  9. Juergen StuberDec 3, 2006
  10. Robin RosenbergDec 3, 2006
  11. Jakub NarebskiDec 3, 2006
  12. Shawn PearceDec 4, 2006
  13. Juergen StuberDec 4, 2006
  14. Shawn PearceDec 3, 2006
  15. sfDec 3, 2006
  16. Shawn PearceDec 3, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.