Re: RFC: adding xdelta compression to git
- From
Dan Holmsand <holmsand@gmail.com>
- Date
- May 3, 2005, 12:48 UTC
- Message-ID
- <d57rip$ojm$1@sea.gmane.org>
- In-Reply-To
- <Pine.LNX.4.58.0505022131380.3594@ppc970.osdl.org>
Linus Torvalds wrote:
Show 5 quoted lines
> Also, the fact is, since git saves things as separate files, you'd not win > as much as you would with some other backing store. So the second step is > to start packing the objects etc. I think there is actually a very steep > complexity edge here - not because any of the individual steps necessarily > add a whole lot, but because they all lead to the "next step".
Actually, you can win quite a lot.
I've just been playing with storing the entire linux-2.4.0-to-2.6.12-rc2-patchset as xdelta patches in git. The entire thing ended up being 577M (instead of some 3.5G), according to du -sh --apparent-size. Considering that that's some 800M of patches, that's not too bad.
I used a very simple scheme: I stored a delta to the previous version of every file if that delta was less than 20% in size of the new file (otherwise, the whole file was stored as usual). If the previous version was already in delta form, the delta was computed from that versions "parent". I didn't even care to look at files less than 4k in size.
In other words, I didn't have to use any delta "chains", and still got quite massive storage size gains. And this scheme could easily be used on the fly; I'm guessing that it would be performance neutral (or even a slight gain, since less has to be compressed and written to disk on commits, and there might be less to read when diffing, since two "delta-blobs" might share the same parent).
Or, with careful tuning, a repo might be "deltaified" later on (assuming that delta blobs are addressed using the expanded files' hash).
/dan