git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Why Git is so fast

From
Shawn O. Pearce <spearce@spearce.org>
Date
Apr 30, 2009, 18:52 UTC
Message-ID
<20090430185244.GR23604@spearce.org>
In-Reply-To
<200904301728.06989.jnareb@gmail.com>
Jakub Narebski <jnareb@gmail.com> wrote:
> Let's rephrase question a bit then: what low-level operation were needed
> for good performance in JGit? 
Aside from the message I just posted:
- Avoid String, its too expensive most of the time.  Stick with
  byte[], and better, stick with data that is a triplet of (byte[],
  int start, int end) to define a region of data.  Yes, its annoying,
  as its 3 values you need to pass around instead of just 1, but
  its makes a big difference in running time.
- Avoid allocating byte[] for SHA-1s, instead we convert to 5 ints,
  which can be inlined into an object allocation.
- Subclass instead of contain references.  We extend ObjectId to
  attach application data, rather than contain a reference to an
  ObjectId.  Classical Java programming techniques would say this
  is a violation of encapsulatio.  But it gets us the same memory
  impact that C Git gets by saying:
    struct appdata {
      unsigned char[20] sha1;
      ....
	}
- We're hurting dearly for not having more efficient access to the
  pack-*.pack file data.  mmap in Java is crap.  We implement our
  own page buffer, reading in blocks of 8192 bytes at a time and
  holding them in our own cache.
  Really, we should write our own mmap library as an optional JNI
  thing, and tie it into libz so we can efficiently run inflate()
  off the pack data directly.
- We're hurting dearly for not having more efficient access to the
  pack-*.idx files.  Again, with no mmap we read the entire bloody
  index into memory.  But since you won't touch most of it we keep
  it in large byte[], but since you are searching with an ObjectId
  (5 ints) we pay a conversion price on every search step where
  we have to copy from the large byte[] to 5 local variable ints,
  and then compare to the ObjectId.  Its an overhead C git doesn't
  have to deal with.
Anyway.

I'm still just amazed at how well JGit runs given these limitations. I guess that's Moore's Law for you. 10 years ago, JGit wouldn't have been practical.

-- 
Shawn.
Previous: Jakub NarebskiNext: Kjetil Barvik
Message 14 of 39 in “Eric Sink's blog - notes on git, dscms and a "whole product" approach”
  1. Martin LanghoffApr 27, 2009
  2. Cross-Platform Version Control (was: Eric Sink's blog - notes on git, dscms and a "whole product" approach)Jakub Narebski, Apr 28, 2009
  3. Robin RosenbergApr 28, 2009
  4. Martin LanghoffApr 29, 2009
  5. Jeff KingApr 29, 2009
  6. Markus HeidelbergApr 29, 2009
  7. Jakub NarebskiApr 29, 2009
  8. Martin LanghoffApr 29, 2009
  9. Jakub NarebskiApr 28, 2009
  10. Sitaram ChamartyApr 29, 2009
  11. Why Git is so fast (was: Re: Eric Sink's blog - notes on git, dscms and a "whole product" approach)Jakub Narebski, Apr 30, 2009
  12. Michael WittenApr 30, 2009
  13. Jakub NarebskiApr 30, 2009
  14. Shawn O. PearceApr 30, 2009
  15. Kjetil BarvikApr 30, 2009
  16. Shawn O. PearceApr 30, 2009
  17. Kjetil BarvikApr 30, 2009
  18. Steven NoonanMay 1, 2009
  19. James PickensMay 1, 2009
  20. Kjetil BarvikMay 1, 2009
  21. Mike HommeyMay 1, 2009
  22. Kjetil BarvikMay 1, 2009
  23. Tony FinchMay 1, 2009
  24. Dmitry PotapovMay 1, 2009
  25. Mike HommeyMay 1, 2009
  26. Dmitry PotapovMay 1, 2009
  27. Shawn O. PearceApr 30, 2009
  28. Jeff KingApr 30, 2009
  29. Linus TorvaldsMay 1, 2009
  30. Jeff KingMay 1, 2009
  31. david@lang.hmMay 1, 2009
  32. Nicolas PitreMay 1, 2009
  33. Daniel BarkalowMay 1, 2009
  34. Linus TorvaldsMay 1, 2009
  35. david@lang.hmMay 1, 2009
  36. Nicolas PitreApr 30, 2009
  37. Alex RiesenApr 30, 2009
  38. Andreas EricssonMay 4, 2009
  39. Jakub NarebskiApr 30, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.