git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Achieving efficient storage of weirdly structured repos

From
Jeff King <peff@peff.net>
Date
Apr 6, 2008, 16:10 UTC
Message-ID
<20080406161003.GA24358@coredump.intra.peff.net>
In-Reply-To
<alpine.LFD.1.00.0804051729230.11277@woody.linux-foundation.org>
On Sat, Apr 05, 2008 at 05:48:43PM -0700, Linus Torvalds wrote:
Show 10 quoted lines
> One thing that the git model sucks at is how it's not very good at 
> handling large objects. I've often wondered if I should have made "object" 
> be more fine-grained and tried to build up large files from multiple 
> smaller objects.
> 
> [ That said, I think git does the right thing - for source code. The 
>   blocking-up of files would cause a rather more complex model, and one of 
>   the great things about git is how simple the basic model is. But the
>   large-file thing does mean that git potentially sucks really badly for 
>   some other loads ]

I have considered something like this for one of my repos, which is full of images. The large image data very rarely changes, but the small EXIF tags do.

My thought was something like:
  - add a new object type, multiblob; a multiblob contains zero or more
    "child" sha1s, each of which is another multiblob or a blob. The
    data in the multiblob is an in-order concatenation of its children.
  - you would create multiblobs with a "smart" git-add that understands
    the filetype and splits the file accordingly (in my case, probably a
    chunk of headers and EXIF data, and then a chunk with the image
    data).
  - in most of git, whenever you need a blob, you just "unwrap" the
    multiblob to get the original blob data
  - because they're separate objects, pack-objects automagically does
    the right thing
  - a few places would benefit from handling multiblobs specially. In
    particular:
      - the diff machinery could do much more efficient comparisons for
        some inexact renames. E.g., multiblob "1234\n5678" and multiblob
        "abcd\n5678" could ignore the "5678" id.
      - the diff machinery could show diffs that were more human
        readable (e.g., even without understanding what the chunks of
        the multiblob _mean_, it can still say "most of this image
        didn't change, but this textual part did").
Of course there are a few drawbacks:
  - one of git's strengths is that content is the same no matter who
    adds it or how. Now the same file has a different sha1 as a
    multiblob versus a regular blob.
  - it breaks the git model of "we store state in the simplest way, and
    figure everything out afterwards." IOW, you are stuck with whatever
    crappy multiblob split you did when you added or updated the file.
    The usual pattern in git is "dumb add, smart view". Now maybe it is
    worth breaking this for two reasons:
      - dumb add, smart view is often very resource intensive; we can
        get smaller packs and faster rename detection out of this
      - we might be losing information; in the case of renames, we can
        justify not explicitly recording because we can figure out
        later what actually happened. I don't know if there is a
        multiblob split that would encapsulate useful user input.
        My EXIF example doesn't; with a little more CPU time, you could
        just do the automated split at diff or delta time.

So it's an approach that I think would work, but I'm not sure it's worth the effort unless somebody comes up with a compelling reason that you can't just split the blobs up after the fact (and maybe the right approach is that pack v5 can split blobs intelligently to get better deltas, so they are still blobs, but we just store them differently).

-Peff
Previous: Linus TorvaldsNext: Nicolas Pitre
Message 11 of 14 in “Achieving efficient storage of weirdly structured repos”
  1. Roman ShaposhnikApr 3, 2008
  2. Linus TorvaldsApr 3, 2008
  3. Jakub NarebskiApr 4, 2008
  4. Nicolas PitreApr 4, 2008
  5. Pieter de BieApr 4, 2008
  6. Shawn O. PearceApr 5, 2008
  7. Roman ShaposhnikApr 4, 2008
  8. Linus TorvaldsApr 4, 2008
  9. Roman ShaposhnikApr 6, 2008
  10. Linus TorvaldsApr 6, 2008
  11. Jeff KingApr 6, 2008
  12. Nicolas PitreApr 7, 2008
  13. Jeff KingApr 7, 2008
  14. Nicolas PitreApr 7, 2008

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.