From: Nicolas Pitre Date: Wed, 18 May 2005 18:41:54 GMT Subject: Re: [PATCH] improved delta support for git Message-ID: In-Reply-To: On Wed, 18 May 2005, Dan Holmsand wrote: > Nicolas Pitre wrote: > > One thing I've been wondering about is whether gzipping small deltas is > > actually a gain. For very small files it seems that gzip is adding more > > overhead making the compressed file actually larger. Might be worth storing > > some deltas uncompressed if the compressed version turns out to be larger. > > It's probably better to skip deltafication of very small files altogether. Big > pain for small gain, and all that. No, that's not what I mean. Suppose a large source file that may change only one line between two versions. The delta may therefore end up being only a few bytes long. Compressing a few bytes with zlib creates a _larger_ file than the original few bytes. > > Well, any delta object smaller than its original object saves space, even if > > it's 75% of the original size. But... > > That's not true if you want to keep the delta chain length down (and thus > performance up). Sure. That's why I added the -d switch to mkdelta. But if you can fit a delta which is 75% the size of its original object size then you still save 25% of the space, regardless of the delta chain length. > But in this case, the trick is to know when to stop deltafying against one > base file, and start over with another. If you switch to a new keyframe too > often, you obviously lose some potential savings. But if you don't switch > often enough, you end up repeating the same data in too many delta files. That's why multiple combinations should be tried. And to keep things under control then a new argument specifying the delta "distance" might limit the number of trials. > A maximum delta size of 10% turned out to be ideal for at least the "fs" > tree. 8% was significantly worse, as was 15%. (The ideal size depends on how > big the average change is: the smaller the average change, the smaller the max > delta size should be). In fact it seems that deltas might be significantly harder to compress. Therefore a test on the resulting file should probably be done as well to make sure we don't end up with a delta larger than the original object. > > ... but then the ultimate solution is to try out all possible references > > within a given list. My git-deltafy-script already finds out the list of > > objects belonging to the same file. Maybe git-mkdelta should try all > > combinations between them. This way a deeper delta chain could be allowed > > for maximum space saving. > > Yeah. But then you lose the ability to do incremental deltafication, or > deltafication on-the-fly. Not at all. Nothing prevents you from making the latest revision of a file be the reference object and the previous revision turned into a delta against that latest revision, even if it was itself a reference object before. The only thing that must be avoided is a delta loop and current mkdelta code takes care of that already. Nicolas