{"thread":{"id":"23630","subject":"Multiblobs","startedAt":"2010-04-28T15:12:07Z","lastAt":"2010-05-10T13:58:10Z","messageCount":21,"participants":["Sergio Callegari","Avery Pennarun","Geert Bosch","Michael Witten","Sergio","Peter Krefting","Hervé Cauwelier","Mike Hommey","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"140592","messageId":"loom.20100428T164432-954@post.gmane.org","threadId":"23630","inReplyTo":null,"subject":"Multiblobs","fromName":"Sergio Callegari","fromEmail":"sergio.callegari@gmail.com","sentAt":"2010-04-28T15:12:07Z","receivedAt":"2010-04-28T15:12:07Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"Hi,\n\nit happened to me to read an older post by Jeff King about \"multiblobs\"\n(http://kerneltrap.org/mailarchive/git/2008/4/6/1360014) and I was wandering\nwhether the idea has been abandoned for some reason or just put on hold.\n\nApparently, this would marvellously help on\n- storing large binary blobs (the split could happen with a rolling checksum\napproach)\n- storing \"structured files\", such as the many zip-based file formats\n(Opendocument, Docx, Jar files, zip files themselves), tars (including\ncompressed tars), pdfs, etc, whose number is rising day after day...\n- storing binary files with textual tags, where the tags could go on a separate\nblob, greatly simplifying their readout without any need for caching them on a\nnote tree.\n- etc...\n\nFurthermore, this could also\n- help the management of upstream trees. This could be simplified since the\n\"pristine tree\" distributed as a tar.gz file and the exploded repo could share\ntheir blobs making commands such as pristine-tree unnecessary.\n- help projects such as bup that currently need to provide split mechanisms of\ntheir own.\n- be used to add \"different representations\" to objects... for instance, when\nstoring a pdf one could use a fake split to store in a separate blob the\ncorresponding text, making the git-diff of pdfs almost instantaneous.\n\n>From Jeff's post, I guess that the major issue could be that the same file could\nget a different sha1 as a multiblob versus a regular blob, but maybe it could be\npossible to make the multiblob take the same sha1 of the \"equivalent plain blob\"\nrather than its real hash.\n\nFor the moment, I am just very curious about the idea and the possible pros and\ncons... can someone (maybe Jeff himself) tell me a little more? Also I wonder\nabout the two possibilities (implement it in git vs implement it \"on top of\"\ngit).\n\nSergio\n"},{"id":"140597","messageId":"k2y32541b131004281107u6d15ed4ex54b5e5c138cc0e24@mail.gmail.com","threadId":"23630","inReplyTo":"loom.20100428T164432-954@post.gmane.org","subject":"Re: Multiblobs","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-04-28T18:07:02Z","receivedAt":"2010-04-28T18:07:02Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Apr 28, 2010 at 11:12 AM, Sergio Callegari\n<sergio.callegari@gmail.com> wrote:\n> - storing \"structured files\", such as the many zip-based file formats\n> (Opendocument, Docx, Jar files, zip files themselves), tars (including\n> compressed tars), pdfs, etc, whose number is rising day after day...\n\nI'm not sure it would help very much for these sorts of files.  The\nproblem is that compressed files tend to change a lot even if only a\nfew bytes of the original data have changed.\n\nFor things like opendocument, or uncompressed tars, you'd be better\noff to decompress them (or recompress with zip -0) using\n.gitattributes.  Generally these files aren't *so* large that they\nreally need to be chunked; what you want to do is improve the deltas,\nwhich decompressing will do.\n\n> - storing binary files with textual tags, where the tags could go on a separate\n> blob, greatly simplifying their readout without any need for caching them on a\n> note tree.\n\nThat sounds complicated and error prone, and is suspiciously like\nApple's \"resource forks,\" which even Apple has mostly realized were a\nbad idea.\n\n> - help the management of upstream trees. This could be simplified since the\n> \"pristine tree\" distributed as a tar.gz file and the exploded repo could share\n> their blobs making commands such as pristine-tree unnecessary.\n\nSharing the blobs of a tarball with a checked-out tree would require a\ntar-specific chunking algorithm.  Not impossible, but a pain, and you\nmight have a hard time getting it accepted into git since it's\nobviously not something you really need for a normal \"source code\"\ntracking system.\n\n> - help projects such as bup that currently need to provide split mechanisms of\n> their own.\n\nSince bup is so awesome that it will soon rule the world of file\nsplitting backup systems, and bup already has a working implemention,\nthis reason by itself probably isn't enough to integrate the feature\ninto git.\n\n> - be used to add \"different representations\" to objects... for instance, when\n> storing a pdf one could use a fake split to store in a separate blob the\n> corresponding text, making the git-diff of pdfs almost instantaneous.\n\nAie, files that have different content depending how you look at them?\n You'll make a lot of enemies with such a patch :)\n\n> From Jeff's post, I guess that the major issue could be that the same file could\n> get a different sha1 as a multiblob versus a regular blob, but maybe it could be\n> possible to make the multiblob take the same sha1 of the \"equivalent plain blob\"\n> rather than its real hash.\n\nI think that's actually not a very important problem.  Files that are\ndifferent will still always have differing sha1s, which is the\nimportant part.  Files that are the same might not have the same sha1,\nwhich is a bit weird, but it's unlikely that any algorithm in git\ndepends fundamentally on the fact that the sha1s match.\n\nStoring files as split does have a lot of usefulness for calculating\ndiffs, however: because you can walk through the tree of hashes and\nshort entire circuit subtrees with identical sha1s, you can diff even\n20GB files really rapidly.\n\n> For the moment, I am just very curious about the idea and the possible pros and\n> cons... can someone (maybe Jeff himself) tell me a little more? Also I wonder\n> about the two possibilities (implement it in git vs implement it \"on top of\"\n> git).\n\n\"on top of\" git has one major advantage, which is that it's easy: for\nexample, bup already does it.  The disadvantage is that checking out\nthe resulting repository won't be smart enough to re-merge the data\nagain, so you have a bunch of tiny chunk files you have to concatenate\nby hand.\n\nImplementing inside git could be done in one of two ways: add support\nfor a new 'multiblob' data type (which is really more like a tree\nobject, but gets checked out as a single file), or implement chunking\nat the packfile level, so that higher-level tools never have to know\nabout multiblobs.\n\nThe latter would probably be easier and more backward-compatibility,\nbut you'd probably lose the ability to do really fast diffs between\nmultiblobs, since diff happens at the higher level.\n\nOverall, I'm not sure git would benefit much from supporting large\nfiles in this way; at least not yet.  As soon as you supported this,\nyou'd start running into other problems... such as the fact that\nshallow repos don't really work very well, and you obviously don't\nwant to clone every single copy of a 100MB file just so you can edit\nthe most recent version.  So you might want to make sure shallow repos\n/ sparse checkouts are fully up to speed first.\n\nHave fun,\n\nAvery\n"},{"id":"140599","messageId":"AAA4B50E-9539-450D-9B6C-E67856D3D5BC@adacore.com","threadId":"23630","inReplyTo":"loom.20100428T164432-954@post.gmane.org","subject":"Re: Multiblobs","fromName":"Geert Bosch","fromEmail":"bosch@adacore.com","sentAt":"2010-04-28T18:34:02Z","receivedAt":"2010-04-28T18:34:02Z","isPatch":false,"sender":{"key":"bosch@adacore.com","avatar":null},"body":"\nOn Apr 28, 2010, at 11:12, Sergio Callegari wrote:\n\n> Hi,\n> \n> it happened to me to read an older post by Jeff King about \"multiblobs\"\n> (http://kerneltrap.org/mailarchive/git/2008/4/6/1360014) and I was wandering\n> whether the idea has been abandoned for some reason or just put on hold.\n> \n> Apparently, this would marvellously help on\n> - storing large binary blobs (the split could happen with a rolling checksum\n> approach)\n> - storing \"structured files\", such as the many zip-based file formats\n> (Opendocument, Docx, Jar files, zip files themselves), tars (including\n> compressed tars), pdfs, etc, whose number is rising day after day...\n> - storing binary files with textual tags, where the tags could go on a separate\n> blob, greatly simplifying their readout without any need for caching them on a\n> note tree.\n> - etc...\n\nIn the early days of GIT I once implemented a \"git pipe\" command that would\nallow an unbounded stream of data to be stored in GIT. The stream would be\nbroken up in small segments using context-sensitive break points (essentially\npoints in the code where a hash H of the last N bytes modulo P is equal to some Q).\nThe average segment length will then be about P bytes long.\nMultiple segments would be put in a tree with each tree entry's name being the\ncumulative length of the segment or subtree it references, with enough leading\nzeros to accomodate for the largest length in the tree.\n\nThis works well and allows efficient diff operations or updates of arbitrarily\nlarge files. In particular, all operations take a time proportional to the\nsize of the change rather than the size of the file.\n\nThe draw backs are:\n\n  - All of the variables H, N, P and Q above influence the final hash\n    that is computed for an object, so the values picked must work well.\n  - You'd only want to use this method for largish files, but because\n    this threshold influences final hashes, it again should be picked with care.\n  - more complex than having just simple straight blobs.\n\nOne of the nice aspects of this representation is that extracting the tree\ninto the local filesystem and concatenating all files in the directory\ntree in alphabetical order does yield the original file.\n\n  -Geert"},{"id":"140601","messageId":"loom.20100428T204406-308@post.gmane.org","threadId":"23630","inReplyTo":"k2y32541b131004281107u6d15ed4ex54b5e5c138cc0e24@mail.gmail.com","subject":"Re: Multiblobs","fromName":"Sergio Callegari","fromEmail":"sergio.callegari@gmail.com","sentAt":"2010-04-28T19:13:42Z","receivedAt":"2010-04-28T19:13:42Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"Avery Pennarun <apenwarr <at> gmail.com> writes:\n\n> \n> On Wed, Apr 28, 2010 at 11:12 AM, Sergio Callegari\n> <sergio.callegari <at> gmail.com> wrote:\n> > - storing \"structured files\", such as the many zip-based file formats\n> > (Opendocument, Docx, Jar files, zip files themselves), tars (including\n> > compressed tars), pdfs, etc, whose number is rising day after day...\n> \n> I'm not sure it would help very much for these sorts of files.  The\n> problem is that compressed files tend to change a lot even if only a\n> few bytes of the original data have changed.\n\nProbably I have not provided enough elements... My idea is the following:\n\nIf you store a structured file as a multiblob, you can use a blob for each\nuncompressed element of content.  For instance, when storing an opendocument\nfile you could use a blob for manifest.xml, one for content.xml, etc... (try\nunzip -l on an odt or odp file to get an idea). When you edit your file only a\nfew of these change. For instance, if we talk about a presentation, each slide\nhas its own content.xml, so changing one slide only that changes.\n\nThe same for PDF files, if you split them using a blob for each uncompressed\nstream, little variations of the pdf file will touch only a blob.\n\nIn other terms, to benefit from multiblobs you should use a different splitting\nstrategy for PDFs (1 blob per uncompressed stream + 1 header blob telling how\nstreams should be put together), Zip files (1 blob per uncompressed file + 1\nheader blob also containing metadata), long unstructured binary files (1 blob\nper chunk + 1 header blob), etc.\n\n> For things like opendocument, or uncompressed tars, you'd be better\n> off to decompress them (or recompress with zip -0) using\n> .gitattributes.  Generally these files aren't *so* large that they\n> really need to be chunked; what you want to do is improve the deltas,\n> which decompressing will do.\n\nThis is what I currently do.  But using multiblobs would be a definite\nimprovement over this.\n \n> > - storing binary files with textual tags, where the tags could go on a\nseparate\n> > blob, greatly simplifying their readout without any need for caching them on\na\n> > note tree.\n> \n> That sounds complicated and error prone, and is suspiciously like\n> Apple's \"resource forks,\" which even Apple has mostly realized were a\n> bad idea.\n\nI did not mean the Apple way... Suppose that you need to store images with exif\ntags.  In order to diff them you would tipically set a textconv attribute, to\nsee only the tags.  However, this kind of filter needs to read the whole file\n(expensive). BTW this is why a caching mechanism involving notes has recently\nbeen proposed. Now suppose that you can set up a rule so that image files with\ntags are stored as a multiblob. You can use 3 blobs... 1 as a header, one for\nthe raw image data and one for the tags.  Now your textconv filter only needs to\nlook at the content of the tags blob.\n\n> > - help the management of upstream trees. This could be simplified since the\n> > \"pristine tree\" distributed as a tar.gz file and the exploded repo could\nshare\n> > their blobs making commands such as pristine-tree unnecessary.\n\nSimilar... Right now to do package management with git, you need to use pristine\ntar. This is because when you check in the upstream tar you only check in its\nelements, not the whole tar.gz.  So you need pristine tar to recreate the\nupstream tar.gz whenever needed. But with multiblob you could store both the\ncontent /and/ the upstream tar and there would be minimal overlap since the\nblobs would be the same. \n \n> Sharing the blobs of a tarball with a checked-out tree would require a\n> tar-specific chunking algorithm.  Not impossible, but a pain, and you\n> might have a hard time getting it accepted into git since it's\n> obviously not something you really need for a normal \"source code\"\n> tracking system.\n\nI agree... but there could be just a mere couple of gitattributes multiblobsplit\nand multiblobcompose, so that one could provide his own splitting and composing\nmethods for the types of files he is interested in (and maybe contribute them to\nthe community).\n\n> > - help projects such as bup that currently need to provide split mechanisms\nof\n> > their own.\n> \n> Since bup is so awesome that it will soon rule the world of file\n> splitting backup systems, and bup already has a working implemention,\n> this reason by itself probably isn't enough to integrate the feature\n> into git.\n\nOn this I tend to agree!\n\n > > - be used to add \"different representations\" to objects... for instance,\nwhen\n> > storing a pdf one could use a fake split to store in a separate blob the\n> > corresponding text, making the git-diff of pdfs almost instantaneous.\n> \n> Aie, files that have different content depending how you look at them?\n>  You'll make a lot of enemies with such a patch :)\n\nI would not consider it as different content... rather as a way to cache data\nyou might need.  But I agree this is probably going too far.\n \n> Overall, I'm not sure git would benefit much from supporting large\n> files in this way; at least not yet.  As soon as you supported this,\n> you'd start running into other problems... such as the fact that\n> shallow repos don't really work very well, and you obviously don't\n> want to clone every single copy of a 100MB file just so you can edit\n> the most recent version.  So you might want to make sure shallow repos\n> / sparse checkouts are fully up to speed first.\n\nI am not really thinking that much about large binary files (that would anyway\ncome as a bonus - an many people often talk about them on the list), but of\nstructured files that currently do not pack well.  My personal issue is with\nopendocument files, since I need to check in lots of documentation and\npresentation material.\n"},{"id":"140610","messageId":"k2x32541b131004281427o2101720at3d324f5e94f05327@mail.gmail.com","threadId":"23630","inReplyTo":"loom.20100428T204406-308@post.gmane.org","subject":"Re: Multiblobs","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-04-28T21:27:32Z","receivedAt":"2010-04-28T21:27:32Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Apr 28, 2010 at 3:13 PM, Sergio Callegari\n<sergio.callegari@gmail.com> wrote:\n> Avery Pennarun <apenwarr <at> gmail.com> writes:\n>> I'm not sure it would help very much for these sorts of files.  The\n>> problem is that compressed files tend to change a lot even if only a\n>> few bytes of the original data have changed.\n>\n> Probably I have not provided enough elements... My idea is the following:\n>\n> If you store a structured file as a multiblob, you can use a blob for each\n> uncompressed element of content.  For instance, when storing an opendocument\n> file you could use a blob for manifest.xml, one for content.xml, etc... (try\n> unzip -l on an odt or odp file to get an idea). When you edit your file only a\n> few of these change. For instance, if we talk about a presentation, each slide\n> has its own content.xml, so changing one slide only that changes.\n\nBut why not use a .gitattributes filter to recompress the zip/odp file\nwith no compression, as I suggested?  Then you can just dump the whole\nthing into git directly.  When you change the file, only the changes\nneed to be stored thanks to delta compression.  Unless your\npresentation is hundreds of megs in size, git should be able to handle\nthat just fine already.\n\n> The same for PDF files, if you split them using a blob for each uncompressed\n> stream, little variations of the pdf file will touch only a blob.\n\nBut then you're digging around inside the pdf file by hand, which is a\nlot of pdf-specific work that probably doesn't belong inside git.\nWorse, because compression programs don't always produce the same\noutput, this operation would most likely actually *change* the hash of\nyour pdf file as you do it.  (That's also true for openoffice files,\nbut at least those are just plain zip files, and zip files are\nsomewhat less of a special case.)\n\n>> For things like opendocument, or uncompressed tars, you'd be better\n>> off to decompress them (or recompress with zip -0) using\n>> .gitattributes.  Generally these files aren't *so* large that they\n>> really need to be chunked; what you want to do is improve the deltas,\n>> which decompressing will do.\n>\n> This is what I currently do.  But using multiblobs would be a definite\n> improvement over this.\n\nIn what way?  I doubt you'd get more efficient storage, at least.\nGit's deltas are awfully hard to beat.\n\n>> That sounds complicated and error prone, and is suspiciously like\n>> Apple's \"resource forks,\" which even Apple has mostly realized were a\n>> bad idea.\n>\n> I did not mean the Apple way... Suppose that you need to store images with exif\n> tags.  In order to diff them you would tipically set a textconv attribute, to\n> see only the tags.  However, this kind of filter needs to read the whole file\n> (expensive). BTW this is why a caching mechanism involving notes has recently\n> been proposed. Now suppose that you can set up a rule so that image files with\n> tags are stored as a multiblob. You can use 3 blobs... 1 as a header, one for\n> the raw image data and one for the tags.  Now your textconv filter only needs to\n> look at the content of the tags blob.\n\nA resource fork by any other name is still a resource fork, and it's\nstill ugly.  If you really need something like this, just cache the\nattributes in a file alongside the big file, and store both files in\nthe git repo.\n\n> Similar... Right now to do package management with git, you need to use pristine\n> tar. This is because when you check in the upstream tar you only check in its\n> elements, not the whole tar.gz.  So you need pristine tar to recreate the\n> upstream tar.gz whenever needed. But with multiblob you could store both the\n> content /and/ the upstream tar and there would be minimal overlap since the\n> blobs would be the same.\n\nI guess.  For something like that, though, Debian's pristine-tarball\ntool seems to already solve the problem and works with any VCS, not\njust git.\n\n>> Sharing the blobs of a tarball with a checked-out tree would require a\n>> tar-specific chunking algorithm.  Not impossible, but a pain, and you\n>> might have a hard time getting it accepted into git since it's\n>> obviously not something you really need for a normal \"source code\"\n>> tracking system.\n>\n> I agree... but there could be just a mere couple of gitattributes multiblobsplit\n> and multiblobcompose, so that one could provide his own splitting and composing\n> methods for the types of files he is interested in (and maybe contribute them to\n> the community).\n\nI guess this would be mostly harmless; the implementation could mirror\nthe filter stuff.\n\n> I am not really thinking that much about large binary files (that would anyway\n> come as a bonus - an many people often talk about them on the list), but of\n> structured files that currently do not pack well.  My personal issue is with\n> opendocument files, since I need to check in lots of documentation and\n> presentation material.\n\nIn that case, I'd like to see some comparisons of real numbers\n(memory, disk usage, CPU usage) when storing your openoffice documents\n(using the .gitattributes filter, of course).  I can't really imagine\nhow splitting the files into more pieces would really improve disk\nspace usage, at least.\n\nHaving done some tests while writing bup, my experience has been that\nchunking-without-deltas is great for these situations:\n1) you have the same data shared across *multiple* files (eg. the same\nimages in lots of openoffice documents with different filenames);\n2) you have the same data *repeated* in the same file at large\ndistances (so that gzip compression doesn't catch it; eg. VMware\nimages)\n3) your file is too big to work with the delta compressor (eg. VMware images).\n\nHowever, in my experience #1 is pretty rare and #2 and #3 aren't in\nyour use case.  And deltas-between-chunks is not very easy to do,\nsince it's hard to guess which chunks might be \"similar\" to which\nother chunks.\n\nPersonally, I think it would be great if git could natively handle\nlarge numbers of large binary files efficiently, because there are a\nfew use cases I would have for it.  But whenever I start investigating\nmy use cases, it always turns out that just \"supporting large files\"\nis just the tip of the iceberg, and there's a huge submerged mass of\niceberg that becomes obvious as soon as you start crashing into it.\n\nThe bup use case (write-once, read-almost-never, incremental backups)\nis a rare exception in which fixing *only* the file size problem has\nproduced useful results.\n\nHave fun,\n\nAvery\n"},{"id":"140614","messageId":"o2ub4087cc51004281610sdaba4276u9726cfaca6bff5ad@mail.gmail.com","threadId":"23630","inReplyTo":"k2x32541b131004281427o2101720at3d324f5e94f05327@mail.gmail.com","subject":"Re: Multiblobs","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2010-04-28T23:10:38Z","receivedAt":"2010-04-28T23:10:38Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"On Wed, Apr 28, 2010 at 16:27, Avery Pennarun <apenwarr@gmail.com> wrote:\n>\n> But then you're digging around inside the pdf file by hand, which is a\n> lot of pdf-specific work that probably doesn't belong inside git.\n\nCore git could provide just the mechanisms for easily defining\n'plugins' for handling different formats.\n"},{"id":"140615","messageId":"loom.20100429T010742-199@post.gmane.org","threadId":"23630","inReplyTo":"k2x32541b131004281427o2101720at3d324f5e94f05327@mail.gmail.com","subject":"Re: Multiblobs","fromName":"Sergio","fromEmail":"sergio.callegari@gmail.com","sentAt":"2010-04-28T23:26:27Z","receivedAt":"2010-04-28T23:26:27Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"Avery Pennarun <apenwarr <at> gmail.com> writes:\n\n> But why not use a .gitattributes filter to recompress the zip/odp file\n> with no compression, as I suggested?  Then you can just dump the whole\n> thing into git directly.  When you change the file, only the changes\n> need to be stored thanks to delta compression.  Unless your\n> presentation is hundreds of megs in size, git should be able to handle\n> that just fine already.\n\nActually, I'm doing so...  But in some occasions odf file that share many\ncomponents do not delta, even when passed through a filter that uncompresses\nthem. Multiblobs are like taking advantage of a known structure to get better\ndeltas.\n\n> But then you're digging around inside the pdf file by hand, which is a\n> lot of pdf-specific work that probably doesn't belong inside git.\n\nI perfectly agree that git should not know about the inner structure of things\nlike PDFs, Zips, Tars, Jars, whatever. But having an infrastructure allowing\nmultiblobs and attributes like clean/smudge to trigger creation and use of\nmultiblobs with user provided split/unsplit drivers could be nice.\n\n> Worse, because compression programs don't always produce the same\n> output, this operation would most likely actually *change* the hash of\n> your pdf file as you do it. \n\nThis should depend on the split/unsplit driver that you write. If your driver\nstores a sufficient amount of metadata about the streams and their order, you\nshould be able to recreate the original file.\n\n> In what way?  I doubt you'd get more efficient storage, at least.\n> Git's deltas are awfully hard to beat.\n\nUsing the known structure of the file, you automatically identify the bits that\nare identical and you save the need to find a delta altogether.\n\n\n> > I agree... but there could be just a mere couple of gitattributes\nmultiblobsplit\n> > and multiblobcompose, so that one could provide his own splitting and\ncomposing\n> > methods for the types of files he is interested in (and maybe contribute\nthem to\n> > the community).\n> \n> I guess this would be mostly harmless; the implementation could mirror\n> the filter stuff.\n\nThis is exactly what I was thinking of: multiblobs as a generalization of the\nfilter infrastructure.\n\n> In that case, I'd like to see some comparisons of real numbers\n> (memory, disk usage, CPU usage) when storing your openoffice documents\n> (using the .gitattributes filter, of course).  I can't really imagine\n> how splitting the files into more pieces would really improve disk\n> space usage, at least.\n\nI'll try to isolate test cases, making test repos:\n\na) with 1 odf file changing a little on each checkin\nb) the same storing the odf file with no compression with a suitable filter\nc) the same storing the tree inside the odf file.\n\n> Having done some tests while writing bup, my experience has been that\n> chunking-without-deltas is great for these situations:\n> 1) you have the same data shared across *multiple* files (eg. the same\n> images in lots of openoffice documents with different filenames);\n> 2) you have the same data *repeated* in the same file at large\n> distances (so that gzip compression doesn't catch it; eg. VMware\n> images)\n> 3) your file is too big to work with the delta compressor (eg. VMware images).\n\nAn aside: bup is great!!! Thanks!\n \nAnd thanks for all your comments, of course!\n\nSergio\n"},{"id":"140618","messageId":"h2w32541b131004281744xea800f1eq9459bbe462ba3a1e@mail.gmail.com","threadId":"23630","inReplyTo":"loom.20100429T010742-199@post.gmane.org","subject":"Re: Multiblobs","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-04-29T00:44:07Z","receivedAt":"2010-04-29T00:44:07Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Wed, Apr 28, 2010 at 7:26 PM, Sergio <sergio.callegari@gmail.com> wrote:\n> Avery Pennarun <apenwarr <at> gmail.com> writes:\n>> But why not use a .gitattributes filter to recompress the zip/odp file\n>> with no compression, as I suggested?  Then you can just dump the whole\n>> thing into git directly.  When you change the file, only the changes\n>> need to be stored thanks to delta compression.  Unless your\n>> presentation is hundreds of megs in size, git should be able to handle\n>> that just fine already.\n>\n> Actually, I'm doing so...  But in some occasions odf file that share many\n> components do not delta, even when passed through a filter that uncompresses\n> them. Multiblobs are like taking advantage of a known structure to get better\n> deltas.\n\nHmm, it might be a good idea to investigate the specific reasons why\nthat's not working.  Fixing it may be easier (and help more people)\nthan introducing a whole new infrastructure for these multiblobs.\n\n>> But then you're digging around inside the pdf file by hand, which is a\n>> lot of pdf-specific work that probably doesn't belong inside git.\n>\n> I perfectly agree that git should not know about the inner structure of things\n> like PDFs, Zips, Tars, Jars, whatever. But having an infrastructure allowing\n> multiblobs and attributes like clean/smudge to trigger creation and use of\n> multiblobs with user provided split/unsplit drivers could be nice.\n\nYes, it could.  Sorry to be playing the devil's advocate :)\n\n>> Worse, because compression programs don't always produce the same\n>> output, this operation would most likely actually *change* the hash of\n>> your pdf file as you do it.\n>\n> This should depend on the split/unsplit driver that you write. If your driver\n> stores a sufficient amount of metadata about the streams and their order, you\n> should be able to recreate the original file.\n\nAlmost.  The one thing you can't count on replicating reliably is\ncompression.  If you use git-zlib the first time, and git-zlib the\nsecond time with the same settings, of course the results will be\nidentical each time.  But if the original file used Acrobat-zlib, and\nyour new one uses git-zlib, the most likely situation is the files\nwill be functionally identical but not the same stream of bytes, and\nthat could be a problem.  (Then again, maybe it's not a problem in\nsome use cases.)\n\nAnother danger of this method is that different versions of git may\nhave slightly different versions of zlib that compress slightly\ndifferently.  In that case, you'd (rather surprisingly) end up with\ndifferent output files depending which version of git you use to check\nthem out.  Maybe that's manageable, though.\n\n>> In what way?  I doubt you'd get more efficient storage, at least.\n>> Git's deltas are awfully hard to beat.\n>\n> Using the known structure of the file, you automatically identify the bits that\n> are identical and you save the need to find a delta altogether.\n\nbup avoids the need to find a delta altogether.  This isn't entirely a\ngood thing; it's a necessity because it processes huge amounts of data\nand doing deltas across it all would be ungodly slow.\n\nHowever, in all my tests (except with massively self-redundant files\nlike VMware images) deltas are at least somewhat smaller than bup\ndeduplication.  This isn't surprising, since deltas can eliminate\nduplication on a byte-by-byte level, while bup chunks have a much\nlarger threshold (around 8k).\n\nSo I question the idea that this method would actually save any space\nover git's existing deltas.  CPU time, yes, but only really during gc,\nand you can run gc overnight while you're not waiting for it.\n\n>> In that case, I'd like to see some comparisons of real numbers\n>> (memory, disk usage, CPU usage) when storing your openoffice documents\n>> (using the .gitattributes filter, of course).  I can't really imagine\n>> how splitting the files into more pieces would really improve disk\n>> space usage, at least.\n>\n> I'll try to isolate test cases, making test repos:\n>\n> a) with 1 odf file changing a little on each checkin\n> b) the same storing the odf file with no compression with a suitable filter\n> c) the same storing the tree inside the odf file.\n\nThis sounds like it would be quite interesting to see.  I would also\nbe interested in d) the test from (b) using bup instead of git.\n\nYou might also want to compare results with 'git gc' vs. 'git gc --aggressive'.\n\n>> Having done some tests while writing bup, my experience has been that\n>> chunking-without-deltas is great for these situations:\n>> 1) you have the same data shared across *multiple* files (eg. the same\n>> images in lots of openoffice documents with different filenames);\n>> 2) you have the same data *repeated* in the same file at large\n>> distances (so that gzip compression doesn't catch it; eg. VMware\n>> images)\n>> 3) your file is too big to work with the delta compressor (eg. VMware images).\n>\n> An aside: bup is great!!! Thanks!\n\nGlad you like it :)\n\nHave fun,\n\nAvery\n"},{"id":"140647","messageId":"20100429065541.GB3268@glandium.org","threadId":"23630","inReplyTo":"loom.20100428T164432-954@post.gmane.org","subject":"Re: Multiblobs","fromName":"Mike Hommey","fromEmail":"mh@glandium.org","sentAt":"2010-04-29T06:55:41Z","receivedAt":"2010-04-29T06:55:41Z","isPatch":false,"sender":{"key":"mh@glandium.org","avatar":"https://avatars.githubusercontent.com/u/1038527?v=4"},"body":"On Wed, Apr 28, 2010 at 03:12:07PM +0000, Sergio Callegari wrote:\n> Hi,\n> \n> it happened to me to read an older post by Jeff King about \"multiblobs\"\n> (http://kerneltrap.org/mailarchive/git/2008/4/6/1360014) and I was wandering\n> whether the idea has been abandoned for some reason or just put on hold.\n> \n> Apparently, this would marvellously help on\n> - storing large binary blobs (the split could happen with a rolling checksum\n> approach)\n> - storing \"structured files\", such as the many zip-based file formats\n> (Opendocument, Docx, Jar files, zip files themselves), tars (including\n> compressed tars), pdfs, etc, whose number is rising day after day...\n> - storing binary files with textual tags, where the tags could go on a separate\n> blob, greatly simplifying their readout without any need for caching them on a\n> note tree.\n> - etc...\n\nThis sounds very much like what I've had in mind for a while, but I\nalways thought that git as a VCS doesn't need that, and that it could be\na feature of a new program, for which the git object database would be a\nspecial case. That is, a program using the git object database format\nfor individual objects and packs, but with additional object types.\n\nMike\n"},{"id":"140644","messageId":"alpine.DEB.2.00.1004291231410.16241@ds9.cixit.se","threadId":"23630","inReplyTo":"k2x32541b131004281427o2101720at3d324f5e94f05327@mail.gmail.com","subject":"Re: Multiblobs","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2010-04-29T11:34:59Z","receivedAt":"2010-04-29T11:34:59Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Avery Pennarun:\n\n> But why not use a .gitattributes filter to recompress the zip/odp file \n> with no compression, as I suggested?  Then you can just dump the whole \n> thing into git directly.\n\nThe advantage would be that you could look at the history of the individual \ncomponents of the zip/openoffice file and follow changes. When looking at \nthe entire zip file (even if using no compression), it is still a compound \nfile.\n\nThe few times I need to version control zip or openoffice files, I only need \nto version control it *as* a zipped file, I don't need the version control \nto ensure that I get exactly the file out that I put in, just that it is \nzipped in both ends. If Git could do that by unzipping and storing the \nindividual components itself, that would be great.\n\nOr if someone could create a \"zgit\" that would allow me to version control \nsuch a file by internally unzipping it and storing it in git, and then \nzipping it up on checkout.\n\nHaving support for merging files inside the zip file would also be a \nwonderful feature to have, especially if the zip file holds mostly-text data.\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"140645","messageId":"x2v32541b131004290828ua9c2d194o1280177360dd231e@mail.gmail.com","threadId":"23630","inReplyTo":"alpine.DEB.2.00.1004291231410.16241@ds9.cixit.se","subject":"Re: Multiblobs","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-04-29T15:28:20Z","receivedAt":"2010-04-29T15:28:20Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Thu, Apr 29, 2010 at 7:34 AM, Peter Krefting <peter@softwolves.pp.se> wrote:\n> Avery Pennarun:\n>> But why not use a .gitattributes filter to recompress the zip/odp file\n>> with no compression, as I suggested?  Then you can just dump the whole thing\n>> into git directly.\n>\n> The advantage would be that you could look at the history of the individual\n> components of the zip/openoffice file and follow changes.\n\nThis use case seems to be converging more and more on the\n\"clean/smudge filter like\" idea, which might be ok.  But I think it\nwould be a kind of messy if the git index/worktree shows only one\nfile, but the actual object shows up as a tree, though.  What should\n'git show HEAD:filename.odt' do?  How about 'git cat-file\nHEAD:filename.odt'?  What if I *do* want to check out one of the\nindividual components?\n\nIt might be saner to just write some wrapper scripts on top of git,\nand cleanly just check in the individual components.  Then just build\na Makefile and run something like 'make extract' before checkin (to\nmake sure all the .odp files/etc are broken into components) and 'make\nassemble' after checkout.\n\nHave fun,\n\nAvery\n"},{"id":"140626","messageId":"alpine.DEB.2.00.1004300918310.24359@ds9.cixit.se","threadId":"23630","inReplyTo":"x2v32541b131004290828ua9c2d194o1280177360dd231e@mail.gmail.com","subject":"Re: Multiblobs","fromName":"Peter Krefting","fromEmail":"peter@softwolves.pp.se","sentAt":"2010-04-30T08:20:39Z","receivedAt":"2010-04-30T08:20:39Z","isPatch":false,"sender":{"key":"peter@softwolves.pp.se","avatar":"https://avatars.githubusercontent.com/u/990764?v=4"},"body":"Avery Pennarun:\n\n(I seem to have been unsubscribed from the list, and can't subscribe again; \nplease keep cc's to me for the time being).\n\n> This use case seems to be converging more and more on the \"clean/smudge \n> filter like\" idea, which might be ok.\n\nThat's what I am using now (recompressing files), but that approach is a bit \nfragile (it suddenly broke on my Mac install, and it only works \nintermittently on Windows).\n\n> It might be saner to just write some wrapper scripts on top of git, and \n> cleanly just check in the individual components.\n\nYeah, that was my thought to (thus the \"zgit\" idea).\n\n-- \n\\\\// Peter - http://www.softwolves.pp.se/\n"},{"id":"140628","messageId":"4BDA9F5C.2080808@itaapy.com","threadId":"23630","inReplyTo":"loom.20100428T204406-308@post.gmane.org","subject":"Re: Multiblobs","fromName":"Hervé Cauwelier","fromEmail":"herve@itaapy.com","sentAt":"2010-04-30T09:14:04Z","receivedAt":"2010-04-30T09:14:04Z","isPatch":false,"sender":{"key":"herve@itaapy.com","avatar":null},"body":"On 04/28/10 21:13, Sergio Callegari wrote:\n> If you store a structured file as a multiblob, you can use a blob for each\n> uncompressed element of content.  For instance, when storing an opendocument\n> file you could use a blob for manifest.xml, one for content.xml, etc... (try\n> unzip -l on an odt or odp file to get an idea). When you edit your file only a\n> few of these change. For instance, if we talk about a presentation, each slide\n> has its own content.xml, so changing one slide only that changes.\n\nI'll obviously let the Git experts answer you, but I can answer about \nOpenDocument itself.\n\nIn a presentation each slide is a <draw:page/> inside a single \ncontent.xml. So if you change one slide, the whole XML will serialize \nwith a different SHA.\n\nAnd maybe you'll add style to that slide, or probably OpenOffice.org \nwill generate an automatic style, so styles.xml will also change. Adding \nan image also changes manifest.xml, along with storing the image itself. \nOOo will surely record the last slide displayed when closing the \napplication, so settings.xml will change too.\n\nSo, all in all, for a single slide, 30 to 80 % of the Zip content may \nchange.\n\nUnless you are talking about a dedicated application to store and \ngenerate on-the-fly office documents, built on top of Git, you're better \nnot touching the contents the user is entrusting git to store, and write \na .gitattribute not to compress them in a pack.\n\nYou may also be interested in the git-bigfiles project that was \nmentioned last week.\n\nhttp://caca.zoy.org/wiki/git-bigfiles\n\n-- \nHervé Cauwelier - ITAAPY - 9 rue Darwin 75018 Paris\nTél. 01 42 23 67 45 - Fax 01 53 28 27 88\nhttp://www.itaapy.com/ - http://www.cms-migration.com\n"},{"id":"140631","messageId":"k2g32541b131004301026n5e8d20b0pc74d22507d4f23d4@mail.gmail.com","threadId":"23630","inReplyTo":"alpine.DEB.2.00.1004300918310.24359@ds9.cixit.se","subject":"Re: Multiblobs","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-04-30T17:26:19Z","receivedAt":"2010-04-30T17:26:19Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"On Fri, Apr 30, 2010 at 4:20 AM, Peter Krefting <peter@softwolves.pp.se> wrote:\n> Avery Pennarun:\n>> This use case seems to be converging more and more on the \"clean/smudge\n>> filter like\" idea, which might be ok.\n>\n> That's what I am using now (recompressing files), but that approach is a bit\n> fragile (it suddenly broke on my Mac install, and it only works\n> intermittently on Windows).\n\nIn general, if you find that existing features have bugs, the correct\nsolution is not to add more buggy features, but to fix the ones that\nalready exist :)\n\nAvery\n"},{"id":"140632","messageId":"z2p32541b131004301032jd28b4b0azbb600880f4e15871@mail.gmail.com","threadId":"23630","inReplyTo":"4BDA9F5C.2080808@itaapy.com","subject":"Re: Multiblobs","fromName":"Avery Pennarun","fromEmail":"apenwarr@gmail.com","sentAt":"2010-04-30T17:32:39Z","receivedAt":"2010-04-30T17:32:39Z","isPatch":false,"sender":{"key":"apenwarr@gmail.com","avatar":"https://avatars.githubusercontent.com/u/20592?v=4"},"body":"2010/4/30 Hervé Cauwelier <herve@itaapy.com>:\n> I'll obviously let the Git experts answer you, but I can answer about\n> OpenDocument itself.\n>\n> In a presentation each slide is a <draw:page/> inside a single content.xml.\n> So if you change one slide, the whole XML will serialize with a different\n> SHA.\n>\n> And maybe you'll add style to that slide, or probably OpenOffice.org will\n> generate an automatic style, so styles.xml will also change. Adding an image\n> also changes manifest.xml, along with storing the image itself. OOo will\n> surely record the last slide displayed when closing the application, so\n> settings.xml will change too.\n>\n> So, all in all, for a single slide, 30 to 80 % of the Zip content may\n> change.\n\nSure.  But if you name the chunks consistently, git's delta\ncompression can deal with tiny changes like those very easily.\n\nThe question is whether it'll work equally well, or better, or worse,\nwith a one-big-file format.  I think we won't know this without doing\nsome actual tests.\n\n(Normally, you could assume that one-big-file is the most\nspace-efficient storage format, because then xdelta and gzip have the\nmost data to work with.  But if you have a lot of *duplicated* content\ninside the same file, and the distance between duplications is outside\nthe gzip window, you could find that more unusual methods - like the\nmethod used by bup - results in better compression.  I know this is\ntrue for VM images, so it may be true for other things.  I haven't\ntested everything :))\n\n> You may also be interested in the git-bigfiles project that was mentioned\n> last week.\n>\n> http://caca.zoy.org/wiki/git-bigfiles\n\ngit-bigfiles is a worthwhile project.  Its goal of \"make life\nbearable\" is aiming kind of low, though.  Basically they seem to be\naiming simply to make git not die horribly when given lots of large\nfiles.  This is commendable, but the resulting repo will be very space\ninefficient when your large files change frequently in small ways.  So\nI think it doesn't solve the problem Sergio brought up.\n\nHave fun,\n\nAvery\n"},{"id":"140635","messageId":"u2lb4087cc51004301116t17ba0efamf4c9b38842bad409@mail.gmail.com","threadId":"23630","inReplyTo":"4BDA9F5C.2080808@itaapy.com","subject":"Re: Multiblobs","fromName":"Michael Witten","fromEmail":"mfwitten@gmail.com","sentAt":"2010-04-30T18:16:55Z","receivedAt":"2010-04-30T18:16:55Z","isPatch":false,"sender":{"key":"mfwitten@gmail.com","avatar":"https://avatars.githubusercontent.com/u/597101?v=4"},"body":"2010/4/30 Hervé Cauwelier <herve@itaapy.com>:\n>\n> Unless you are talking about a dedicated application to store and generate\n> on-the-fly office documents, built on top of Git, you're better not touching\n> the contents the user is entrusting git to store, and write a .gitattribute\n> not to compress them in a pack.\n\nDoesn't OOo provide at least some library of official code for\nhandling such files, so that other programs might be able to\ninteroperate?\n\nIf so, then it would be almost trivial for an OpenDocument 'plugin' to\nbe 'built on top of Git'.\n\nIf not, then OOo is crap.\n"},{"id":"140642","messageId":"4BDB2A22.40400@itaapy.com","threadId":"23630","inReplyTo":"u2lb4087cc51004301116t17ba0efamf4c9b38842bad409@mail.gmail.com","subject":"Re: Multiblobs","fromName":"Hervé Cauwelier","fromEmail":"herve@itaapy.com","sentAt":"2010-04-30T19:06:10Z","receivedAt":"2010-04-30T19:06:10Z","isPatch":false,"sender":{"key":"herve@itaapy.com","avatar":null},"body":"On 04/30/10 20:16, Michael Witten wrote:\n> 2010/4/30 Hervé Cauwelier<herve@itaapy.com>:\n>>\n>> Unless you are talking about a dedicated application to store and generate\n>> on-the-fly office documents, built on top of Git, you're better not touching\n>> the contents the user is entrusting git to store, and write a .gitattribute\n>> not to compress them in a pack.\n>\n> Doesn't OOo provide at least some library of official code for\n> handling such files, so that other programs might be able to\n> interoperate?\n\nI'm not sure what you mean but the only way to interoperate with OOo is \nto run it in \"server mode\" with at least a framebuffer xorg in the \nbackground. Then you connect a client and use their RPC/Corba-like API.\n\nOpenDocument libraries all start from scratch, or at least the RelaxNG \nschema to generate validating code.\n\nIf the chunks are Zip parts, you're almost done. If you want smarter \nsplitting logic like slides in a presentation, sheets in a spreadsheet, \nand pages... no, there is no page in a text; well, you need to go \nthrough the XML layer or better use a OpenDocument library that \nabstracts it. Other parts in the Zip like styles and metadata are easier \nto split since they are basically a linear collection of objects.\n\n> If so, then it would be almost trivial for an OpenDocument 'plugin' to\n> be 'built on top of Git'.\n>\n> If not, then OOo is crap.\n\nI already had reasons to conclude this. But hopefully OD is an open \nstandard, not restricted to OOo.\n\n-- \nHervé Cauwelier - ITAAPY - 9 rue Darwin 75018 Paris\nTél. 01 42 23 67 45 - Fax 01 53 28 27 88\nhttp://www.itaapy.com/ - http://www.cms-migration.com\n"},{"id":"141033","messageId":"20100506062644.GB16151@coredump.intra.peff.net","threadId":"23630","inReplyTo":"loom.20100428T164432-954@post.gmane.org","subject":"Re: Multiblobs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2010-05-06T06:26:44Z","receivedAt":"2010-05-06T06:26:44Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Wed, Apr 28, 2010 at 03:12:07PM +0000, Sergio Callegari wrote:\n\n> it happened to me to read an older post by Jeff King about \"multiblobs\"\n> (http://kerneltrap.org/mailarchive/git/2008/4/6/1360014) and I was wandering\n> whether the idea has been abandoned for some reason or just put on hold.\n\nI am a little late getting to this thread, and I agree with a lot of\nwhat Avery said elsewhere, so I won't repeat what's been said. But after\nreading my own message that you linked and the rest of this thread, I\nwanted to note a few things.\n\nOne is that many of the applications for these multiblobs are extremely\nvaried, and many of them are vague and hand-waving. I think you really\nhave to look at each application individually to see how a solution\nwould fit. In my original email, I mentioned linear chunking of large\nblobs for:\n\n  1. faster inexact rename detection\n\n  2. better diffs of binary files\n\nI think (2) is now obsolete. Since that message, we now have textconv\nfilters, which allow simple and fast diffs of large objects (in my\nexample, I talked about exif tags on images. I now textconv the images\ninto a text representation of the exif tags and diff those). And with\ntextconv caching, we can do it on the fly without impacting how we\nrepresent the object in git (we don't even have to pull the original\nlarge blob out of storage at all, as the cache provide a look-aside\ntable keyed by the object name).\n\nI also mentioned in that email that in theory we could diff individual\nchunks even if we don't understand their semantic meaning. In practice,\nI don't think this works. Most binary formats are going to involve not\njust linear chunking, but decoding the binary chunks into some\nhuman-readable form. So smart chunking isn't enough; you need a decoder,\nwhich is what a textconv filter does.\n\nFor item (1), this is closely related to faster (and possibly better)\ndelta compression. I say only possibly better, because in theory our\ndelta algorithm should be finding something as simple as my example\nalready.\n\nAnd for both of those cases, the upside is a speed increase, but the\ndownside is a breakage of the user-visible git model (i.e., blobs get\ndifferent sha1's depending on how they've been split). But being two\nyears wiser than when I wrote the original message, I don't think that\nbreakage is justified. Instead, you should retain the simple git object\nmodel, and consider on-the-fly content-specific splits. In other words,\nat rename (or delta) time notice that blob 123abc is a PDF, and that it\ncan be intelligently split into several chunks, and then look for other\nfiles which share chunks with it. As a bonus, this sort of scheme is\nvery easy to cache, just as textconv is. You cache the smart-split of\nthe blob, which is immutable for some blob/split-scheme combination. And\nthen you can even do rename detection on large blob 123abc without even\nretrieving it from storage.\n\nAnother benefit is that you still _store_ the original (you just don't\nlook at it as often). Which means there is no annoyance with perfectly\nreconstructing a file. I had originally envisioned straight splitting,\nwith concatenation as the reverse operation. But I have seen things like\nzip and tar files mentioned in this thread. They are quite challenging,\nbecause it is difficult to reproduce them byte-for-byte. But if you take\nthe splitting out of the git data model, then that problem just goes\naway.\n\nThe other application I saw in this thread is structured files where you\nactually _want_ to see all of the innards as individual files (e.g.,\nbeing able to do \"git show HEAD:foo.zip/file.txt\"). And for those, I\ndon't think any sort of automated chunking is really desirable. If you\nwant git to store and process those files individually, then you should\nprovide them to git individually. In other words, there is no need for\ngit to know or care at all that \"foo.zip\" exists, but you should simply\nfeed it a directory containing the files. The right place to do that\nconversion is either totally outside of git, or at the edges of git\n(i.e., git-add and when git places the file in the repository). Our\ncurrent hooks may not be sufficient, but that means those hooks should\nbe improved, which to me is much more favorable than a scheme that\nalters the core of the git data model.\n\nSo no, reading my original message, I don't think it was a good idea. :)\nThe things people want to accomplish are reasonable goals, but there are\nbetter ways to go about it.\n\n-Peff\n"},{"id":"142084","messageId":"4BE3493B.8010409@gmail.com","threadId":"23630","inReplyTo":"20100506062644.GB16151@coredump.intra.peff.net","subject":"Re: Multiblobs","fromName":"Sergio Callegari","fromEmail":"sergio.callegari@gmail.com","sentAt":"2010-05-06T22:56:59Z","receivedAt":"2010-05-06T22:56:59Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"Many thanks for the clear and evidently very well thought answer.\nI wonder if I can take another minute of your (and Avery, and anybody \nelse who is interested) time to feed a little more my curiosity.\nAnd I apologize in advance for possible mistakes in my understanding of \ngit internals.\n\nJeff King wrote:\n\n> And for both of those cases, the upside is a speed increase, but the\n> downside is a breakage of the user-visible git model (i.e., blobs get\n> different sha1's depending on how they've been split).\nIs this different from what happens with clean/smudge filters? I wonder \nwhat hash does a cleanable object get. The hash of its cleaned version \nor its original hash? If it is the first case, the hash can change if \nthe filter is used/not-used or slightly modified, so I wonder if an \nenhanced \"clean\" filter capable of splitting an object into a multiblob \nwould be different in this sense. If it gets the original hash, again I \nwonder if an enhanced \"clean\" filter capable of splitting an object into \na multiblob could not do the same.\n>  But being two\n> years wiser than when I wrote the original message, I don't think that\n> breakage is justified. Instead, you should retain the simple git object\n> model, and consider on-the-fly content-specific splits. In other words,\n> at rename (or delta) time notice that blob 123abc is a PDF, and that it\n> can be intelligently split into several chunks, and then look for other\n> files which share chunks with it. As a bonus, this sort of scheme is\n> very easy to cache, just as textconv is. You cache the smart-split of\n> the blob, which is immutable for some blob/split-scheme combination. And\n> then you can even do rename detection on large blob 123abc without even\n> retrieving it from storage.\n>   \nNow I see why for things like diffing, showing textual representations \nor rename detection caching can be much more practical.\nMy initial list of \"potential applications\" was definitely too wide and \nvague.\n> Another benefit is that you still _store_ the original (you just don't\n> look at it as often). \n... but of course if you keep storing the original, I guess there is no \nadvantage in storage efficiency.\n> Which means there is no annoyance with perfectly\n> reconstructing a file. I had originally envisioned straight splitting,\n> with concatenation as the reverse operation. But I have seen things like\n> zip and tar files mentioned in this thread. They are quite challenging,\n> because it is difficult to reproduce them byte-for-byte.\nI agree, but this is already being done. For instance on odf and zip \nfiles, by using clean filters capable of removing compression you can \ngreatly improve the storage efficiency of the delta machinery included \nin git. And of course, to re-create the original file is potentially \nchallenging. But most time, it does not really matter. For instance, \nwhen I use this technique with odf files, I do not need to care if the \nsmudge filter recreates the original file or not, the important thing is \nthat it recreates a file that can then be cleaned to the same thing (and \nthis makes me think that cleanable objects get the sha1 of the cleaned \nblob, see above).\n\nIn other terms, all the time we underline that git is about tracking \n/content/. However, when you have a structured file, and you want to \ntrack its /content/, most time you are not interested at all at the \n/envelope/ (e.g. the compression level of the odf/zip file): the content \nis what is inside (typically a tree-structured thing). And maybe scms \ncould be made better at tracking structured files, by providing an easy \nway to tell the scm how to discard the envelope.\n\nIn fact, having the hash of the structured file only depend on its real \ncontent (the inner tree or list of files/streams/whatever), seems to me \nto be completely respectful of the git model. This is why I originally \nthought that having enhanced filters enabling the storage of the the \ninner matter of a structured file as a multiblob could make sense.\n> The other application I saw in this thread is structured files where you\n> actually _want_ to see all of the innards as individual files (e.g.,\n> being able to do \"git show HEAD:foo.zip/file.txt\"). And for those, I\n> don't think any sort of automated chunking is really desirable. If you\n> want git to store and process those files individually, then you should\n> provide them to git individually. In other words, there is no need for\n> git to know or care at all that \"foo.zip\" exists, but you should simply\n> feed it a directory containing the files. The right place to do that\n> conversion is either totally outside of git, or at the edges of git\n> (i.e., git-add and when git places the file in the repository).\nOriginally, I thought of creating wrappers for some git commands. \nHowever, things like \"status\" or \"commit -a\" appeared to me quite \ncomplicated to be done in a wrapper.\n>  Our\n> current hooks may not be sufficient, but that means those hooks should\n> be improved, which to me is much more favorable than a scheme that\n> alters the core of the git data model.\n>   \nHaving a sufficient number of hooks could help a lot. However, if I \nremember properly, one of the reasons why the clean/smudge filters were \nintroduced was to avoid the need to implement a similar functionality \nwith hooks.\n\n\nThanks in advance for the further explanations that might come!\n\nSergio\n"},{"id":"141349","messageId":"20100510063618.GD13340@coredump.intra.peff.net","threadId":"23630","inReplyTo":"4BE3493B.8010409@gmail.com","subject":"Re: Multiblobs","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2010-05-10T06:36:18Z","receivedAt":"2010-05-10T06:36:18Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Fri, May 07, 2010 at 12:56:59AM +0200, Sergio Callegari wrote:\n\n> >And for both of those cases, the upside is a speed increase, but the\n> >downside is a breakage of the user-visible git model (i.e., blobs get\n> >different sha1's depending on how they've been split).\n> Is this different from what happens with clean/smudge filters? I\n> wonder what hash does a cleanable object get. The hash of its cleaned\n> version or its original hash? If it is the first case, the hash can\n\nIt gets the cleaned version. The idea is that the sha1 in the repository\nis the \"official\" version, and anything else is simply a representation\nsuitable for use on your platform.\n\nSo in that sense, clean/smudge filters are very visible. Splitting into\nmultiple blobs would mean that as far as git was concerned, your data\n_is_ multiple blobs. And it would diff and merge them as separate\nentities. That makes sense for something where that breakdown happens\nalong user-visible lines, and is useful to the user. For example,\nautomatically breaking down a tarfile into its constituent files might\nbe a more desirable representation for git to diff and merge (though the\ncurrent implementation of clean/smudge filters does not allow breaking\nthe file into multiple blobs).\n\nBut as I argued later in my email, I think that is not the right model\nfor performance-oriented multiblobs. Splitting a file at certain length\nboundaries simply because it is large is going to be awkward when you\nwant to look at it as a whole item.\n\n> >Another benefit is that you still _store_ the original (you just don't\n> >look at it as often).\n> ... but of course if you keep storing the original, I guess there is\n> no advantage in storage efficiency.\n\nYes and no. If you are storing some set of N bytes, then you need to\nstore N bytes whether they are in a single blob or multiple blobs. The\nonly way that multiple blobs can improve on that is if you can find\nbetter delta candidates by doing so.  Which means that you are just as\nwell off by splitting the large blob when looking for delta candidates\nas you are in splitting it in storage.\n\n> I agree, but this is already being done. For instance on odf and zip\n> files, by using clean filters capable of removing compression you can\n> greatly improve the storage efficiency of the delta machinery\n> included in git. And of course, to re-create the original file is\n> potentially challenging. But most time, it does not really matter.\n> For instance, when I use this technique with odf files, I do not need\n> to care if the smudge filter recreates the original file or not, the\n> important thing is that it recreates a file that can then be cleaned\n> to the same thing (and this makes me think that cleanable objects get\n> the sha1 of the cleaned blob, see above).\n\nSure. And for those cases, I think clean/smudge filters are perhaps\nalready doing the job.\n\nAs an aside, I don't think that _git_ cares about pristine tars. It is\nthat people want to store compressed tarfiles in git that have a\nparticular checksum because they are interacting with some _other_\nsystem that cares about the tarfile.  In your case, where you don't care\nabout the particular byte pattern of the odf file, it is much simpler.\nSo clean/smudge filters are even easier there.\n\n> In other terms, all the time we underline that git is about tracking\n> /content/. However, when you have a structured file, and you want to\n> track its /content/, most time you are not interested at all at the\n> /envelope/ (e.g. the compression level of the odf/zip file): the\n> content is what is inside (typically a tree-structured thing). And\n> maybe scms could be made better at tracking structured files, by\n> providing an easy way to tell the scm how to discard the envelope.\n\nRight. The question is how the structured contents are handled\ninternally by the SCM. Git's choice is to leave contents as opaque as\npossible, and let you handle conversion at the boundaries: textconv (or\na custom external diff) for viewing diffs, and clean/smudge for working\ntree files.\n\n> In fact, having the hash of the structured file only depend on its\n> real content (the inner tree or list of files/streams/whatever),\n> seems to me to be completely respectful of the git model. This is why\n\nYes, and that is how it works with clean/smudge filters.\n\n> I originally thought that having enhanced filters enabling the\n> storage of the the inner matter of a structured file as a multiblob\n> could make sense.\n\nI do think it makes sense, but only for some applications. But for those\napplications, rather than a multiblob, I think creating a tree structure\nis a natural fit, and works well with existing git tools. But again,\nthat isn't really implemented. Blobs must stay as blobs. So the closest\nyou can come is saying:\n\n  - an ODF file may be a collection of structured text, but we will\n    store it marshalled as a single binary data stream\n\n  - we don't want it compressed for performance reasons, so we won't use\n    the native marshalling format. Instead, we'll clean/smudge it as an\n    uncompressed collection format inside of git (e.g., a zip without\n    compression, or a tarball).\n\n  - even though git doesn't understand the structure, we _do_ want to\n    see the structure when doing diffs or merges. For that, we define\n    custom diff/merge drivers which can operate on the file. They can\n    unpack the structure as necessary.\n\nwhich is really not too bad, and it means git can remain blissfully\nunaware of the details of any format.\n\n> >provide them to git individually. In other words, there is no need for\n> >git to know or care at all that \"foo.zip\" exists, but you should simply\n> >feed it a directory containing the files. The right place to do that\n> >conversion is either totally outside of git, or at the edges of git\n> >(i.e., git-add and when git places the file in the repository).\n> Originally, I thought of creating wrappers for some git commands.\n> However, things like \"status\" or \"commit -a\" appeared to me quite\n> complicated to be done in a wrapper.\n\nYes, I would just do it manually. But in theory a clean/smudge filter\ncould be the right sort of place for that, if somebody made an\nimplementation that handle exploding a single file into an arbitrary\ntree/blob hierarchy. I think it was discussed when filters were\nintroduced, but the complexity (both in terms of implementation, and\nin meeting user expectations) prevented anyone from moving it forward.\n\n-Peff\n"},{"id":"141399","messageId":"4BE810F2.3080107@gmail.com","threadId":"23630","inReplyTo":"20100510063618.GD13340@coredump.intra.peff.net","subject":"Re: Multiblobs","fromName":"Sergio Callegari","fromEmail":"sergio.callegari@gmail.com","sentAt":"2010-05-10T13:58:10Z","receivedAt":"2010-05-10T13:58:10Z","isPatch":false,"sender":{"key":"sergio.callegari@gmail.com","avatar":"https://gravatar.com/avatar/c98f41317e0422c1e630385de0e3970227b8e5ad15f35ba8586066467cc833bc?d=mp&s=160"},"body":"On 10/05/2010 08:36, Jeff King wrote:\n> Sure. And for those cases, I think clean/smudge filters are perhaps\n> already doing the job.\n>\n>    \nAs a matter of fact, my idea was exactly to think of a multiblob as a \ngit tree (maybe plus a signature).\n\nWith this, one can set up a \"multiclean\" filter, triggered by a filename \npattern as for a normal filter or by the invocation of \"file\" to look at \ninner magic.\n\nWhen this filter is invoked it should take the file to be cleaned as the \nstdin and output a tree at the output, while (as a side effect) \npopulating the git storage by the normal blobs pointed to by the tree.\n\nIn a complementary fashion, the multismudge filter should receive the \n\"multiblob\" tree on stdin, and output on stdout the smudged file, \ninspecting the blobs pointed to by the tree to do the work.\n\nThis would require having trees as tree entries, and (I guess) also some \nupdate to the git package machinery, but apart from that should fit well \nwith the current clean/smudge approach, nor significantly alter the git \nmodel.\n\nSergio\n"}]}