Re: Handling large files with GIT
- From
Greg KH <greg@kroah.com>
- Date
- Feb 9, 2006, 04:54 UTC
- Message-ID
- <20060209045420.GB15924@kroah.com>
- In-Reply-To
- <87slqty2c8.fsf@mid.deneb.enyo.de>
On Wed, Feb 08, 2006 at 10:20:39PM +0100, Florian Weimer wrote:
Show 14 quoted lines
> * Martin Langhoff: > > > SVN does reasonably well tracking his >1GB mbox file. Now, I don't > > know if I like the idea of putting my own mbox file under version > > control, but it looks like projects with large and slow-changing files > > would be in trouble with GIT. > > To my surprise, it's not that bad. The Debian testing-security team > uses a single 1.8 MB file (400 KB compressed) to keep vulnerability > data. Most changes to that file involve just a few lines. But even > in this extreme case, git doesn't compare too badly against Subversion > if you pack regularly (but not too often). Disk usage is actually > *below* Subversion FSFS even with --depth=10 (the default, > unfortunately a bit hard to override).
I have a project that has 2.5Mb files, and git handles them just fine, even on my old slow laptop.
But when I tried to use it to backup my old email archive a few months ago, running about 2Gb in about 300 files, it took forever. Luckily I only archive stuff off every other month or so, otherwise it would be unusable.
However, I did notice one problem. When cloning from one machine to another, for a project that is already fully packed, it seems that the whole project is packed again before sending it accross the wire. With an archive this big, that takes over an hour for my slow old fileserver. I ended up just rsyncing over the files and pointing the parent back to the original. Is there anyway to not repack everything if it's not needed?
thanks,
greg k-h