git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Handling large files with GIT

From
Linus Torvalds <torvalds@osdl.org>
Date
Feb 13, 2006, 16:19 UTC
Message-ID
<Pine.LNX.4.64.0602130806070.3691@g5.osdl.org>
In-Reply-To
<43F01F5A.5020808@pobox.com>
On Mon, 13 Feb 2006, Jeff Garzik wrote:
Show 7 quoted lines
>
> Linus Torvalds wrote:
> > I've never used maildir layout, but if it is a couple of large _flat_
> > subdirectories,
> 
> That's what it is :/   One directory per mail folder, with each email an
> individual file in that dir.
Ok.

Anyway, I double-checked, and I'm wrong anyway. While the "static directories" thing is a huge performance optimization for doing many things (diffing trees, file history in git-rev-list, etc etc), for merging it doesn't help. We always end up expanding the whole tree.

Which is kind of sad.

It's inevitable in one sense: we do the merge in the index, after all, and the index - unlike the tree structures - is a flat file (like the "manifest" in mercurial or monotone). It's also represented that way in memory.

However, it is a total and complete waste in other cases.

Thinking more about it, this is also why merging causes all the horrible index performance: not only do we (unnecessarily) read the same trees over and over again only to collapse them back to stage0 later when they are the same, but because we keep the index in a linear format, when we read the other trees, we'll have to move things around with memmove() (just the pointers, but still).

We'd actually be a _lot_ better off if we split "git-read-tree" up into two phases: one that did the recursive tree operation (which can optimize the "same tree everywhere" case), and the second stage that actually populated the index.

I'll have to think about this. It would be an absolutely _huge_ optimization for merging in certain patterns, it just doesn't matter for something like the kernel with "just" 18,000 files and not a lot of strange merging going on.

In contrast, I can see a mail archive easily having hundreds of thousands of individual emails. At which time it's horribly stupid to read them all in three times (for a merge - base, origin, new) and do so in a pretty inefficient manner.

Ho humm. It doesn't look _hard_ per se, and I think the two-stage git-read-tree is actually also what the recursive merge strategy wants anyway (it can't use the index - it really just wants to get a list of conflict information). So this definitely sounds like the RightThing(tm) to do anyway, and it fits the git data structures really well.

So no downsides. Except that this is some rather core code, and you can't afford to get it wrong. And the fact that I'm a lazy bastard, of course.

			Linus
Previous: Martin LanghoffNext: Martin Langhoff
Message 36 of 39 in “Handling large files with GIT”
  1. Martin LanghoffFeb 8, 2006
  2. Johannes SchindelinFeb 8, 2006
  3. Linus TorvaldsFeb 8, 2006
  4. Linus TorvaldsFeb 8, 2006
  5. Junio C HamanoFeb 8, 2006
  6. Florian WeimerFeb 8, 2006
  7. Martin LanghoffFeb 8, 2006
  8. Ben CliffordFeb 13, 2006
  9. Linus TorvaldsFeb 13, 2006
  10. Linus TorvaldsFeb 13, 2006
  11. Linus TorvaldsFeb 13, 2006
  12. Ian MoltonFeb 13, 2006
  13. Martin LanghoffFeb 13, 2006
  14. Johannes SchindelinFeb 14, 2006
  15. Linus TorvaldsFeb 14, 2006
  16. Sam VilainFeb 14, 2006
  17. Linus TorvaldsFeb 14, 2006
  18. Junio C HamanoFeb 14, 2006
  19. Sam VilainFeb 15, 2006
  20. Junio C HamanoFeb 15, 2006
  21. Sam VilainFeb 15, 2006
  22. Martin LanghoffFeb 15, 2006
  23. Linus TorvaldsFeb 15, 2006
  24. Linus TorvaldsFeb 15, 2006
  25. Linus TorvaldsFeb 15, 2006
  26. Linus TorvaldsFeb 15, 2006
  27. Junio C HamanoFeb 15, 2006
  28. Linus TorvaldsFeb 15, 2006
  29. Linus TorvaldsFeb 15, 2006
  30. Linus TorvaldsFeb 16, 2006
  31. Junio C HamanoFeb 16, 2006
  32. Fredrik KuivinenFeb 16, 2006
  33. Jeff GarzikFeb 13, 2006
  34. Keith PackardFeb 13, 2006
  35. Martin LanghoffFeb 14, 2006
  36. Linus TorvaldsFeb 13, 2006
  37. Martin LanghoffFeb 13, 2006
  38. Greg KHFeb 9, 2006
  39. Martin LanghoffFeb 9, 2006

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.