From: Linus Torvalds Date: Wed, 01 Feb 2006 02:04:31 GMT Subject: Re: [Census] So who uses git? Message-ID: In-Reply-To: <20060201013901.GA16832@mail.com> On Tue, 31 Jan 2006, Ray Lehtiniemi wrote: > > for what it's worth, it's certainly true here... i'm using git to help > me manage a similar project where i work. Hmm. We _could_ actually fairly easily add a flag to the index which means "don't even bother comparing - assume same", and then have specific operations to clear that flag. That would allow people with slow filesystems (not just Windows: even under Linux, the cold-cache case is always going to be pretty slow) to have a _choice_: they could continue to use git it is done now (explicit checks), _or_ they could mark all their index caches as "implicitly up-to-date" and use a separate program to mark them as being potentially edited. We still have one unused bit in the cache-entry "ce_flags", so we wouldn't even need to break any existing index files with it. We'd just need to have two new (fast) operations: - mark one or more files as being "implicitly up-to-date" "git checkout" would do this if the proper flag was set in the .git/config file. "git-update-index --refresh" would do this for files that weren't already implicitly up-to-date _and_ the refresh actually showed it to match (and the .git/config file said so). - mark one or more files as _not_ being implicitly up-to-date: people would do this by hand when editing a file (or when just deciding that they want git to re-check everything again) They're fast, because they are purely in the cache (well, git-update-index obviously isn't, but the new op wouldn't be any _slower_ than the old one). Looks simple enough. The big thing to remember is to clear that "implicitly up-to-date" flag whenever we make changes (ie we'd probably make "add_cache_entry()" always clear it, possibly with a flag to add it as "pre-verified" which would set it). Comments? Junio, what do you think? > we're working on a vendor supplied tree which is also hacked upon > by various VAR companies. the tree in question has ~20,000 files > totalling nearly 1.4 GB of source files, ms word docs, binary-only > libraries for a wide array of processor variants, windows exe > files, video clips, etc. (however, the amount of actual source code > interspersed in there is only about 6000 files totaling about 112 MB) > > here's a repo sitting on the local linux filesystem with cold cache: > > reiserfs$ time git update-index --refresh > real 0m17.422s > user 0m0.025s > sys 0m0.320s .. somewhat painful, but with enough memory this is hopefully a pretty rare case. > and with hot cache > > reiserfs$ time git update-index --refresh > real 0m0.151s > user 0m0.020s > sys 0m0.067s This is how it _should_ look. But: > for comparison, one of our sandboxes is sitting on an NTFS file system, > accessed via SMB: > > smbfs$ time git update-index --refresh > real 11m36.502s > user 0m6.830s > sys 0m5.086s Ouch, ouch, ouch. Sounds like every single stat() will go out the wire. I forget what the Linux NFS client does, but I _think_ it has a metadata timeout that avoids this. But it might be as bad under NFS. Has anybody used git over NFS? If it's this bad (or even close to), I guess the "mark files as up-to-date in the index" approach is a really good idea.. Of course, the whole point of git is that you should keep your repository close, but sometimes NFS - or similar - is enforced upon you by other issues, like the fact that the powers-that-be want anonymous workstations and everybody should work with a home-directory automounted over NFS.. Linus