Re: [Census] So who uses git?
- From
Linus Torvalds <torvalds@osdl.org>
- Date
- Feb 1, 2006, 02:04 UTC
- Message-ID
- <Pine.LNX.4.64.0601311747360.7301@g5.osdl.org>
- In-Reply-To
- <20060201013901.GA16832@mail.com>
On Tue, 31 Jan 2006, Ray Lehtiniemi wrote:
> > for what it's worth, it's certainly true here... i'm using git to help > me manage a similar project where i work.
Hmm.
We _could_ actually fairly easily add a flag to the index which means "don't even bother comparing - assume same", and then have specific operations to clear that flag.
That would allow people with slow filesystems (not just Windows: even under Linux, the cold-cache case is always going to be pretty slow) to have a _choice_: they could continue to use git it is done now (explicit checks), _or_ they could mark all their index caches as "implicitly up-to-date" and use a separate program to mark them as being potentially edited.
We still have one unused bit in the cache-entry "ce_flags", so we wouldn't even need to break any existing index files with it.
We'd just need to have two new (fast) operations:
- mark one or more files as being "implicitly up-to-date"
"git checkout" would do this if the proper flag was set in the .git/config file.
"git-update-index --refresh" would do this for files that weren't already implicitly up-to-date _and_ the refresh actually showed it to match (and the .git/config file said so).
- mark one or more files as _not_ being implicitly up-to-date:
people would do this by hand when editing a file (or when just deciding that they want git to re-check everything again)
They're fast, because they are purely in the cache (well, git-update-index obviously isn't, but the new op wouldn't be any _slower_ than the old one).
Looks simple enough. The big thing to remember is to clear that "implicitly up-to-date" flag whenever we make changes (ie we'd probably make "add_cache_entry()" always clear it, possibly with a flag to add it as "pre-verified" which would set it).
Comments? Junio, what do you think?
Show 13 quoted lines
> we're working on a vendor supplied tree which is also hacked upon > by various VAR companies. the tree in question has ~20,000 files > totalling nearly 1.4 GB of source files, ms word docs, binary-only > libraries for a wide array of processor variants, windows exe > files, video clips, etc. (however, the amount of actual source code > interspersed in there is only about 6000 files totaling about 112 MB) > > here's a repo sitting on the local linux filesystem with cold cache: > > reiserfs$ time git update-index --refresh > real 0m17.422s > user 0m0.025s > sys 0m0.320s
.. somewhat painful, but with enough memory this is hopefully a pretty rare case.
Show 6 quoted lines
> and with hot cache > > reiserfs$ time git update-index --refresh > real 0m0.151s > user 0m0.020s > sys 0m0.067s
This is how it _should_ look.
But:
Show 7 quoted lines
> for comparison, one of our sandboxes is sitting on an NTFS file system, > accessed via SMB: > > smbfs$ time git update-index --refresh > real 11m36.502s > user 0m6.830s > sys 0m5.086s
Ouch, ouch, ouch.
Sounds like every single stat() will go out the wire. I forget what the Linux NFS client does, but I _think_ it has a metadata timeout that avoids this. But it might be as bad under NFS.
Has anybody used git over NFS? If it's this bad (or even close to), I guess the "mark files as up-to-date in the index" approach is a really good idea..
Of course, the whole point of git is that you should keep your repository close, but sometimes NFS - or similar - is enforced upon you by other issues, like the fact that the powers-that-be want anonymous workstations and everybody should work with a home-directory automounted over NFS..
Linus