git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Git performance results on a large repository

From
Sam Vilain <sam@vilain.net>
Date
Feb 6, 2012, 21:17 UTC
Message-ID
<4F30435B.5070709@vilain.net>
In-Reply-To
<243C23AF01622E49BEA3F28617DBF0AD5912CA85@SC-MBX02-5.TheFacebook.com>
 > Sam Vilain: Thanks for the pointer, i didn't realize that
 > fast-import was bi-directional.  I used it for generating the
 > synthetic repo.  Will look into using it the other way around.
 > Though that still won't speed up things like git-blame,
 > presumably?

It could, because blame is an operation which primarily works on the source history with little reference to the working copy. Of course this will depend on the quality of the implementation server-side. Blame should suit distribution over a cluster, as it is mostly involved with scanning candidate revisions for string matches which is the compute intensive part. Coming up with candidate revisions has its own cost and can probably also be distributed, but just working on the lowest loop level might be a good place to start.

What it doesn't help with is local filesystem operations. For this I think a different approach is required, if you can tie into fam or a similar inode change notification system, then you should be able to avoid the entire recursive stat on 'git status'. I'm not sure --assume-unchanged on its own is a good idea, you could easily miss things. Those stat's are useful.

Making the index able to hold just changes to the checked-out tree, as others have mentioned, would also save the massive reads and writes you've identified. Perhaps a more high performance back-end could be developed.

 > The sparse-checkout issue you mention is a good one.

It's actually been on the table since at least GitTogether 2008; there's been some design discussion on it and I think it's just one of those features which doesn't have enough demand yet for it to be built. It keeps coming up but not from anyone with the inclination or resources to make it happen. There is a protocol issue, but this should be able to fit into the current extension system.

 > There is a good question of how to support quick checkout,
 > branch switching, clone, push and so forth.

Sure. It will be much more network intensive as you are replacing the part which normally has a very fast link through the buffercache to pack files etc. A hybrid approach is also possible, where objects are fetched individually via fast-import and cached in a local .git repo. And I have a hunch that LZOP compression of the stream may also be a win, but as with all of these ideas, it would be after profiling identifies it as a choke point than just because it sounds good.

 > I'll look into the approaches you suggest.  One consideration
 > is coming up with a high-leverage approach - i.e. not doing
 > heavy dev work if we can avoid it.

Right. You don't actually need to port the whole of git to Hadoop initially, to begin with it can just pass through all commands to a server-side git fast-import process. When you find specific operations which are slow then these specific operations can be implemented using a Hadoop back-end, and the rest backed to the standard git. If done using a useful plug-in system, these systems could be accepted by the core project as an enterprise scaling option.

This could let you get going with the knowledge that the scaling option is there should it come out.

 > On the other hand, it would be nice if we (including the entire
 > community:) ) improve git in areas that others that share
 > similar issues benefit from as well.
Like I say, a lot of people have run into this already...

HTH, Sam

Previous: david@lang.hmNext: Joshua Redstone
Message 26 of 34 in “Git performance results on a large repository”
  1. Joshua RedstoneFeb 3, 2012
  2. Ævar Arnfjörð BjarmasonFeb 3, 2012
  3. Joshua RedstoneFeb 3, 2012
  4. Sam VilainFeb 3, 2012
  5. Sam VilainFeb 3, 2012
  6. Nguyen Thai Ngoc DuyFeb 7, 2012
  7. Matt GrahamFeb 3, 2012
  8. Evgeny SazhinFeb 4, 2012
  9. Chris LeeFeb 3, 2012
  10. Zeki MokhtarzadaFeb 4, 2012
  11. Joey HessFeb 4, 2012
  12. Nguyen Thai Ngoc DuyFeb 4, 2012
  13. Joshua RedstoneFeb 4, 2012
  14. Nguyen Thai Ngoc DuyFeb 5, 2012
  15. Joey HessFeb 6, 2012
  16. Nguyen Thai Ngoc DuyFeb 7, 2012
  17. Joshua RedstoneFeb 9, 2012
  18. Nguyen Thai Ngoc DuyFeb 10, 2012
  19. Christian CouderFeb 10, 2012
  20. Nguyen Thai Ngoc DuyFeb 10, 2012
  21. David MohsFeb 6, 2012
  22. Matt GrahamFeb 6, 2012
  23. Joshua RedstoneFeb 6, 2012
  24. Greg TroxelFeb 6, 2012
  25. david@lang.hmFeb 7, 2012
  26. Sam VilainFeb 6, 2012
  27. Joshua RedstoneFeb 4, 2012
  28. Tomas CarneckyFeb 5, 2012
  29. Nguyen Thai Ngoc DuyFeb 5, 2012
  30. slinkyFeb 4, 2012
  31. Greg TroxelFeb 4, 2012
  32. david@lang.hmFeb 5, 2012
  33. David BarrFeb 5, 2012
  34. Emanuele ZattinFeb 7, 2012

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.