git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Handling large files with GIT

From
Linus Torvalds <torvalds@osdl.org>
Date
Feb 15, 2006, 17:16 UTC
Message-ID
<Pine.LNX.4.64.0602150904310.3691@g5.osdl.org>
In-Reply-To
<Pine.LNX.4.64.0602150715470.3691@g5.osdl.org>

Btw, some actual numbers: I did the recent kernel networking merge (which is a trivial in-index merge) with the standard three-way

	git-read-tree -m <base> <branch> <branch>
and with the new git-merge-tree to compare performance.
Doing git-read-tree takes ~0.35s, while git-merge-tree took 0.015s.

Now, that's not a really fair comparison, because the end result is very different: the git-read-tree has populated the index, ready for a git-writet-ree, while the git-merge-tree has not.

However, the interesting part is that especially for a trivial merge, we don't actually _want_ to necessarily populate the index, because doing a "git-write-tree" is actually a pretty expensive operation (on the kernel, it will try to write 1000+ directory trees, most of which already exist. Admittedly we don't actually have to write the objects, since we figure out that they already exist, but we have to do the SHA1 calculations to do so).

So if we made the git-merge-tree based merge work entirely on trees all the way, and never even necessarily populate the index at all (unless it has to, due to actual data conflicts that want to be fixed up), that would actually be another performance advantage. The only downside there is that we would literally have to write the resulting tree objects by hand (ie we'd need a new helper for doing that, and another thing to validate).

Anyway, that should almost certainly make it possible to scale up git merges to hundreds of thousands of files without huge performance problems (still, that depends a bit on layout - again, flat directory structures won't scale as well, so it might not be enough for maildir handling).

But just at a guess, I think there's at least an order of magnitude to be had there. So if a maildir merge currently takes an hour, at least we should be able to get it down to a few minutes.

Ben, are you interested in trying this out in your maildir experiments?
		Linus
Previous: Linus TorvaldsNext: Linus Torvalds
Message 29 of 39 in “Handling large files with GIT”
  1. Martin LanghoffFeb 8, 2006
  2. Johannes SchindelinFeb 8, 2006
  3. Linus TorvaldsFeb 8, 2006
  4. Linus TorvaldsFeb 8, 2006
  5. Junio C HamanoFeb 8, 2006
  6. Florian WeimerFeb 8, 2006
  7. Martin LanghoffFeb 8, 2006
  8. Ben CliffordFeb 13, 2006
  9. Linus TorvaldsFeb 13, 2006
  10. Linus TorvaldsFeb 13, 2006
  11. Linus TorvaldsFeb 13, 2006
  12. Ian MoltonFeb 13, 2006
  13. Martin LanghoffFeb 13, 2006
  14. Johannes SchindelinFeb 14, 2006
  15. Linus TorvaldsFeb 14, 2006
  16. Sam VilainFeb 14, 2006
  17. Linus TorvaldsFeb 14, 2006
  18. Junio C HamanoFeb 14, 2006
  19. Sam VilainFeb 15, 2006
  20. Junio C HamanoFeb 15, 2006
  21. Sam VilainFeb 15, 2006
  22. Martin LanghoffFeb 15, 2006
  23. Linus TorvaldsFeb 15, 2006
  24. Linus TorvaldsFeb 15, 2006
  25. Linus TorvaldsFeb 15, 2006
  26. Linus TorvaldsFeb 15, 2006
  27. Junio C HamanoFeb 15, 2006
  28. Linus TorvaldsFeb 15, 2006
  29. Linus TorvaldsFeb 15, 2006
  30. Linus TorvaldsFeb 16, 2006
  31. Junio C HamanoFeb 16, 2006
  32. Fredrik KuivinenFeb 16, 2006
  33. Jeff GarzikFeb 13, 2006
  34. Keith PackardFeb 13, 2006
  35. Martin LanghoffFeb 14, 2006
  36. Linus TorvaldsFeb 13, 2006
  37. Martin LanghoffFeb 13, 2006
  38. Greg KHFeb 9, 2006
  39. Martin LanghoffFeb 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.