Re: [PATCH] improved delta support for git
- From
Nicolas Pitre <nico@cam.org>
- Date
- May 18, 2005, 18:41 UTC
- Message-ID
- <Pine.LNX.4.62.0505181428170.20274@localhost.localdomain>
- In-Reply-To
- <d6evrk$jv2$1@sea.gmane.org>
On Wed, 18 May 2005, Dan Holmsand wrote:
Show 8 quoted lines
> Nicolas Pitre wrote: > > One thing I've been wondering about is whether gzipping small deltas is > > actually a gain. For very small files it seems that gzip is adding more > > overhead making the compressed file actually larger. Might be worth storing > > some deltas uncompressed if the compressed version turns out to be larger. > > It's probably better to skip deltafication of very small files altogether. Big > pain for small gain, and all that.
No, that's not what I mean.
Suppose a large source file that may change only one line between two versions. The delta may therefore end up being only a few bytes long. Compressing a few bytes with zlib creates a _larger_ file than the original few bytes.
Show 5 quoted lines
> > Well, any delta object smaller than its original object saves space, even if > > it's 75% of the original size. But... > > That's not true if you want to keep the delta chain length down (and thus > performance up).
Sure. That's why I added the -d switch to mkdelta. But if you can fit a delta which is 75% the size of its original object size then you still save 25% of the space, regardless of the delta chain length.
> But in this case, the trick is to know when to stop deltafying against one > base file, and start over with another. If you switch to a new keyframe too > often, you obviously lose some potential savings. But if you don't switch > often enough, you end up repeating the same data in too many delta files.
That's why multiple combinations should be tried. And to keep things under control then a new argument specifying the delta "distance" might limit the number of trials.
> A maximum delta size of 10% turned out to be ideal for at least the "fs" > tree. 8% was significantly worse, as was 15%. (The ideal size depends on how > big the average change is: the smaller the average change, the smaller the max > delta size should be).
In fact it seems that deltas might be significantly harder to compress. Therefore a test on the resulting file should probably be done as well to make sure we don't end up with a delta larger than the original object.
Show 8 quoted lines
> > ... but then the ultimate solution is to try out all possible references > > within a given list. My git-deltafy-script already finds out the list of > > objects belonging to the same file. Maybe git-mkdelta should try all > > combinations between them. This way a deeper delta chain could be allowed > > for maximum space saving. > > Yeah. But then you lose the ability to do incremental deltafication, or > deltafication on-the-fly.
Not at all. Nothing prevents you from making the latest revision of a file be the reference object and the previous revision turned into a delta against that latest revision, even if it was itself a reference object before. The only thing that must be avoided is a delta loop and current mkdelta code takes care of that already.
Nicolas