From: Dan Holmsand Date: Tue, 03 May 2005 12:48:14 GMT Subject: Re: RFC: adding xdelta compression to git Message-ID: In-Reply-To: Linus Torvalds wrote: > Also, the fact is, since git saves things as separate files, you'd not win > as much as you would with some other backing store. So the second step is > to start packing the objects etc. I think there is actually a very steep > complexity edge here - not because any of the individual steps necessarily > add a whole lot, but because they all lead to the "next step". Actually, you can win quite a lot. I've just been playing with storing the entire linux-2.4.0-to-2.6.12-rc2-patchset as xdelta patches in git. The entire thing ended up being 577M (instead of some 3.5G), according to du -sh --apparent-size. Considering that that's some 800M of patches, that's not too bad. I used a very simple scheme: I stored a delta to the previous version of every file if that delta was less than 20% in size of the new file (otherwise, the whole file was stored as usual). If the previous version was already in delta form, the delta was computed from that versions "parent". I didn't even care to look at files less than 4k in size. In other words, I didn't have to use any delta "chains", and still got quite massive storage size gains. And this scheme could easily be used on the fly; I'm guessing that it would be performance neutral (or even a slight gain, since less has to be compressed and written to disk on commits, and there might be less to read when diffing, since two "delta-blobs" might share the same parent). Or, with careful tuning, a repo might be "deltaified" later on (assuming that delta blobs are addressed using the expanded files' hash). /dan