git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: Multiblobs

From
Avery Pennarun <apenwarr@gmail.com>
Date
Apr 30, 2010, 17:32 UTC
Message-ID
<z2p32541b131004301032jd28b4b0azbb600880f4e15871@mail.gmail.com>
In-Reply-To
<4BDA9F5C.2080808@itaapy.com>
2010/4/30 Hervé Cauwelier <herve@itaapy.com>:
Show 15 quoted lines
> I'll obviously let the Git experts answer you, but I can answer about
> OpenDocument itself.
>
> In a presentation each slide is a <draw:page/> inside a single content.xml.
> So if you change one slide, the whole XML will serialize with a different
> SHA.
>
> And maybe you'll add style to that slide, or probably OpenOffice.org will
> generate an automatic style, so styles.xml will also change. Adding an image
> also changes manifest.xml, along with storing the image itself. OOo will
> surely record the last slide displayed when closing the application, so
> settings.xml will change too.
>
> So, all in all, for a single slide, 30 to 80 % of the Zip content may
> change.

Sure. But if you name the chunks consistently, git's delta compression can deal with tiny changes like those very easily.

The question is whether it'll work equally well, or better, or worse, with a one-big-file format. I think we won't know this without doing some actual tests.

(Normally, you could assume that one-big-file is the most space-efficient storage format, because then xdelta and gzip have the most data to work with. But if you have a lot of *duplicated* content inside the same file, and the distance between duplications is outside the gzip window, you could find that more unusual methods - like the method used by bup - results in better compression. I know this is true for VM images, so it may be true for other things. I haven't tested everything :))

> You may also be interested in the git-bigfiles project that was mentioned
> last week.
>
> http://caca.zoy.org/wiki/git-bigfiles

git-bigfiles is a worthwhile project. Its goal of "make life bearable" is aiming kind of low, though. Basically they seem to be aiming simply to make git not die horribly when given lots of large files. This is commendable, but the resulting repo will be very space inefficient when your large files change frequently in small ways. So I think it doesn't solve the problem Sergio brought up.

Have fun,
Avery
Previous: Hervé CauwelierNext: Michael Witten
Message 13 of 21 in “Multiblobs”
  1. Sergio CallegariApr 28, 2010
  2. Avery PennarunApr 28, 2010
  3. Sergio CallegariApr 28, 2010
  4. Avery PennarunApr 28, 2010
  5. Michael WittenApr 28, 2010
  6. SergioApr 28, 2010
  7. Avery PennarunApr 29, 2010
  8. Peter KreftingApr 29, 2010
  9. Avery PennarunApr 29, 2010
  10. Peter KreftingApr 30, 2010
  11. Avery PennarunApr 30, 2010
  12. Hervé CauwelierApr 30, 2010
  13. Avery PennarunApr 30, 2010
  14. Michael WittenApr 30, 2010
  15. Hervé CauwelierApr 30, 2010
  16. Geert BoschApr 28, 2010
  17. Mike HommeyApr 29, 2010
  18. Jeff KingMay 6, 2010
  19. Sergio CallegariMay 6, 2010
  20. Jeff KingMay 10, 2010
  21. Sergio CallegariMay 10, 2010

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.