{"thread":{"id":"13180","subject":"Git performance on OS X","startedAt":"2008-04-19T19:28:20Z","lastAt":"2008-04-21T20:06:35Z","messageCount":39,"participants":["Pieter de Bie","Linus Torvalds","Jakub Narebski","Roman Shaposhnik","Dmitry Potapov","Junio C Hamano","Luciano Rocha","David Kastrup","Johan Herland"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"74765","messageId":"1208633300-74603-1-git-send-email-pdebie@ai.rug.nl","threadId":"13180","inReplyTo":null,"subject":"Git performance on OS X","fromName":"Pieter de Bie","fromEmail":"pdebie@ai.rug.nl","sentAt":"2008-04-19T19:28:20Z","receivedAt":"2008-04-19T19:28:20Z","isPatch":false,"sender":{"key":"pdebie@ai.rug.nl","avatar":null},"body":"Hi Git mailing list,\n\nI have done some tests regarding git's performance on the OS X platform. We\nnoticed that mercurial is a lot faster than git in the \"git status\" command,\nespecially on the webkit repository. This repository has 45k files, so one\nwould expect it to be slow because of OS X's slow lstat. However, mercurial is\na lot faster (usually 6.2 seconds for git vs ~4 seconds for hg).\n\nFor a reference to the statistics below, `git status' in the webkit repo takes\nabout 6.21 seconds with a std dev of 0.26.\n\n1. 10k empty files\n\nFirst off, I started with the most simple case: a repository with 10k empty\nfiles in a flat repo.\n\nGit add times\n\nIt appears that on large initial imports, git add * is a lot slower than git\nadd .. This test was performed on a directory with 10000 empty files in it.\n\nResults\n=========================================================\nCommand                               Mean     Std\nrm -rf .git && git init && git add .  0.617    0.153\nrm -rf .git && git init && git add *  43.383   0.419\nrm -rf .hg && hg init && hg add .     0.926    0.027\nrm -rf .hg && hg init && hg add *     4.312    0.013\n=========================================================\n\nSampling this, it appears that git add spends a lot of time in fnmatch.\ntop function calls in 4 second sample:\n\n        fnmatch$UNIX2003  2452\n        fnmatch1          310\n        strlen            292\n        mbrtowc_l         188\n\nprobably because git is performing its own glob expansion. This is expensive\non 10,000 supplied files. Of course, this is an uncommon scenario, but still\nMercurial seems to do things differently (I don't know how to sample python,\nunfortunately).\n\nGit status on these 10k files takes about 0.111 seconds:\n\nResults\n================================================\nCommand                     Mean     Std\ngit status                   0.112    0.006\nhg status                    0.317    0.005\n================================================\n\nThis all seems very acceptable. Now we scale up to 50,000 files.\n\n2. 50k empty files\n\nUnfortunately, this was too much for my system to pass as arguments:\n\n  sh: /opt/local/bin/hg: Argument list too long\n\nTherefore, only part of the git adds can be compared\n\nResults\n======================================================================\nCommand                                            Mean     Std    \nrm -rf .git .hg && git init && git add .           6.239   0.184\nrm -rf .hg .git && hg init && hg add .             11.059  0.342\n======================================================================\n\nGit is still faster than Mercurial on adding files.. so far so good. Now the\ngit / hg status test:\n\nResults\n======================================================================\nCommand                                            Mean     Std    \nhg status                                          4.984  0.249\ngit status                                         3.709  0.150\n======================================================================\n\nSo, git takes a bit less time than hg in this case. These are mostly system\ncalls:\n\n    Vienna:perf pieter$ time git status\n    # On branch master\n    nothing to commit (working directory clean)\n\n    real  0m3.705s\n    user  0m0.212s\n    sys   0m3.256s\n    \nSo it's not git's fault here that the status is slow.\n\n\n3. A more complex directory structure.\n\nWe now use Webkit's directory and file structure and see what happens. This\ntest repository has exactly the same files and structure as the webkit repo,\nbut all files are empty.\n\nResults\n======================================================================\nCommand                                            Mean     Std    \nrm -rf .git .hg && git init && git add .           6.014  0.523\nrm -rf .git .hg && git init && git add *           6.198  0.228\nrm -rf .hg .git && hg init && hg add .             7.707  0.519\nrm -rf .hg .git && hg init && hg add *             7.632  0.405\n======================================================================\n\nFunnily enough, Mercurial is faster with this structure than with the\none-directory structure. Git shows linear scaling. Also, with a real\nstructure, the * vs . problem in git goes away.\n\nNow we can look at the \"git status\" commands and compare them to the actual\nstatus' of the actual webkit repository.\n\nResults\n======================================================================\nCommand                                            Mean     Std    \ngit status                                         4.573  0.514\ngit status .                                       13.515  0.448\nhg status                                          4.411  1.594\nhg status .                                        4.903  0.171\n======================================================================\n\nThere's no significant difference between the git and hg status things.\nRemember that in the webkit repo, \"git status\" takes about 6.2 seconds, which\nis a lot slower than we see here.\n\nTherefore, it is interesting to look at what happens if we import the whole\nwebkit branch.\n\n4. A new webkit repository\n\nThis test was done by creating a new clone of the webkit repository.\nBasically, I did a git archive | tar x and did a git add on that.\n\nThis is where some interesting stuff happens. I haven't done the git add\nthing, as that should be clear by now and takes a lot of time. The status\ncommand, however:\n\nResults\n======================================================================\nCommand                                            Mean     Std    \ngit status                                         4.428   0.486\ngit status .                                       13.508  1.451\nhg status                                          4.285   1.681\nhg status .                                        4.930   0.165\n======================================================================\n\nAgain, git shows similar performance to mercurial. Furthermore, the status\ntime hasn't changed since last time. Apparently, the increased file size and\nincreased number of objects didn't matter. So, why is there such a big\ndifference between the real webkit repository and this fresh one?\n\n5. A repacked shallow webkit repo\n\nOne thing that could be it is that the webkit repo is heavily packed. To test\nthis, I created a new clone and repacked this one and (21 minutes later):\n\nResults\n======================================================================\nCommand                                            Mean     Std    \n(Pre-GC): git status                               4.470   0.423\n(Pre-GC): git status .                             13.355  1.025\n(Post-GC): git status                              4.910   0.324\n(Post-GC): git status .                            11.265  0.222\n======================================================================\n\nWhen run with 10 tests in the pre and post case, there is a significant\ndifference according to a t-test (df=18, p << 0.01). Therefore, I compared the\nreal, user and system times pre and post of git status. I also included the\nreal webkit again, and also a shallow clone of that repository.\n\n          Pre-GC       Post-GC      (shallow)   (real webkit) \nreal     4.36 (0.06)   4.61 (0.06)  5.72 (0.25)  6.21 (0.28)  \nuser     0.39 (0.01)   0.37 (0.00)  0.36 (0.00)  0.39 (0.00)  \nsys      3.28 (0.04)   2.86 (0.03)  3.21 (0.09)  2.90 (0.01)  \n\nThe system time seems to jump up and down sometimes, but the real times\ndefinitely keep getting higher. This isn't due to system or user time. Where\ndoes this extra time come from?\n\nI hope anyone can explain this. I tried profiling the commands, but `sample'\noften doesn't want to show symbols (sometimes it does, though) and gprof\ndoesn't show most functions. My profiling skills aren't that high, so if\nanyone has suggestions, I'll be glad to help.\n\n- Pieter\n"},{"id":"74768","messageId":"alpine.LFD.1.10.0804191341210.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"1208633300-74603-1-git-send-email-pdebie@ai.rug.nl","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-19T21:22:38Z","receivedAt":"2008-04-19T21:22:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 19 Apr 2008, Pieter de Bie wrote:\n> \n> It appears that on large initial imports, git add * is a lot slower than git\n> add .\n\n\"git add *\" is actually fundamentally different from \"git add .\", and \nyeah, you should generally use the latter.\n\nThe reason? The argument list is actually something different from what \nyou think it is. For git, it's a \"pathspec\", so what actualy happens is \nthat in *both* cases, it will really traverse the whole tree, and then \nmatch every file it finds against the pathspec.\n\nSo think of the arguments not as a file list, but as a random bunch of \npatterns to match against the files you have!\n\nWhich is why the cost is actually approximately O(n*m), where \"n\" is the \nsize of the working tree, and \"m\" is the number of pathspecs.\n\nSo the reason \"git add .\" is fast is actually that \"m\" in that case is \njust 1 (just one trivial pattern), and then \"git add *\" is slow because \n\"m\" is large (lots of complicated patterns). In both cases, 'n' is the \nsame (== the whole set of files in your working tree).\n\nNow, it has some trivial optimizations, so that if the all the patterns in \nthe pathspec begin with the same base directory, it will only look at that \nbase directory, which is why\n\n\tgit add drivers/block/*\n\nis much faster than\n\n\tgit add *\n\neven though they both have a fair number of pattners (ie 'm' is roughly \nthe same, but now it has basically artificially limited 'n' to just a \nsubset of the tree). When you do \"git add *\", there is obviously no \nsuch trivial subset - '*' is not going to limit the pattern space to just \na small part of the subtree!\n\nI do agree that the git behavior is kind of odd, but it's very consistent \nwith all the other uses of pathspecs in git, and we've never optimized it \na lot simply because nobody normally _should_ do \"git add *\" with a lot of \nfiles.\n\nIf you want to see the worst-case, do\n\n\tgit add $(git ls-files)\n\nwhich basically means that it's O(n^2) in the number of files we track \n(regardless of depth). Because remember: the arguments to git add are not \nsomething we just iterate over - we always iterate over the whole tree, \nand then the arguments are just patterns that we then match that tree \nagainst.\n\nAnd yeah, we could create a few other optimization heuristics that would \nalmost certainly speed up those worst cases by a huge amount. The logic is \nall in \"match_pathspec()\" (where the \"prefix\" count is just the common \nprefix that we don't even need to match because of the trivial \"under this \ntree\" optimization).\n\nNotice how match_pathspec() just walks over the whole pathspec (that's \nyour argument list), and how we call this for every single file we find \n(after we've done .gitignore handling etc).\n\nAnyway, here's a trivial patch that doesn't change this fundamental fact, \nbut that avoids doing anything *expensive* until we've done some cheap \ninitial tests. It may or may not help your test-case, but it's pretty \nsimple and it matches the other git optimizations in this area (ie \n\"conceptually handle the general case, but optimize the simple cases where \nwe can exit early\")\n\nNotice how this patch doesn' actually change the fundamental O(n^2) \nbehaviour, but it makes it much cheaper by generally avoiding the \nexpensive 'fnmatch' and 'strlen/strncmp' when they are obviously not \nneeded.\n\n\t\tLinus\n\n---\n dir.c |   22 ++++++++++++++++++++--\n 1 files changed, 20 insertions(+), 2 deletions(-)\n\ndiff --git a/dir.c b/dir.c\nindex 63715c9..8d45321 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -52,6 +52,11 @@ int common_prefix(const char **pathspec)\n \treturn prefix;\n }\n \n+static inline int special_char(unsigned char c1)\n+{\n+\treturn !c1 || c1 == '*' || c1 == '[' || c1 == '?';\n+}\n+\n /*\n  * Does 'match' matches the given name?\n  * A match is found if\n@@ -69,14 +74,27 @@ static int match_one(const char *match, const char *name, int namelen)\n \tint matchlen;\n \n \t/* If the match was just the prefix, we matched */\n-\tmatchlen = strlen(match);\n-\tif (!matchlen)\n+\tif (!*match)\n \t\treturn MATCHED_RECURSIVELY;\n \n+\tfor (;;) {\n+\t\tunsigned char c1 = *match;\n+\t\tunsigned char c2 = *name;\n+\t\tif (special_char(c1))\n+\t\t\tbreak;\n+\t\tif (c1 != c2)\n+\t\t\treturn 0;\n+\t\tmatch++;\n+\t\tname++;\n+\t\tnamelen--;\n+\t}\n+\t\n+\n \t/*\n \t * If we don't match the matchstring exactly,\n \t * we need to match by fnmatch\n \t */\n+\tmatchlen = strlen(match);\n \tif (strncmp(match, name, matchlen))\n \t\treturn !fnmatch(match, name, 0) ? MATCHED_FNMATCH : 0;\n \n"},{"id":"74769","messageId":"alpine.LFD.1.10.0804191422480.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191341210.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-19T21:29:56Z","receivedAt":"2008-04-19T21:29:56Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 19 Apr 2008, Linus Torvalds wrote:\n> \n> Notice how this patch doesn' actually change the fundamental O(n^2) \n> behaviour, but it makes it much cheaper by generally avoiding the \n> expensive 'fnmatch' and 'strlen/strncmp' when they are obviously not \n> needed.\n\nSide note: on the kenrel tree, it makes the (insane!) operation \n\n\tgit add $(git ls-files)\n\ngo from 49 seconds down to 17 sec. So it does make a huge difference for \nme, but I also want to point out that this really isn't a sane operation \nto do (I also think that 17 sec is totally unacceptable, but I cannot find \nit in me to care, since I don't think this is an operation that anybody \nshould ever do!)\n\nThe optimization is probably worth doing just to avoid the bad worst case, \nbut we should teach people not to do \"git add *\" (or that insane ls-files \nthing), and instead do \"git add .\" and \"git add -u\".\n\nBut in the absense of teaching people that, the patch should at least \nmakes that bad pattern behavior be slightly more acceptable for git, even \nif it's still not very nice.\n\n(Btw, we need to stop using \"fnmatch()\" entirely some day, if only because \nwe can't ever use it for case-insensitive stuff. That patch doesn't help \nus with that, but it doesn't hurt either, and conceptually it's moving in \nthe direction of doing more in \"native\" git code than in \"fnmatch()\")\n\n\t\t\t\tLinus\n"},{"id":"74770","messageId":"alpine.LFD.1.10.0804191443550.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"1208633300-74603-1-git-send-email-pdebie@ai.rug.nl","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-19T21:54:40Z","receivedAt":"2008-04-19T21:54:40Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nBtw, the \"git status\" issue is totally different.\n\nOn Sat, 19 Apr 2008, Pieter de Bie wrote:\n>\n> Now we can look at the \"git status\" commands and compare them to the actual\n> status' of the actual webkit repository.\n> \n> Results\n> ======================================================================\n> Command                                            Mean     Std    \n> git status                                         4.573  0.514\n> git status .                                       13.515  0.448\n> hg status                                          4.411  1.594\n> hg status .                                        4.903  0.171\n\nThe reason \"git status .\" is slower has nothing to do with the pathspec \nmatching, and everything to do with the fact that \"git status\" with a \npathspec means soemthing different again.\n\nRemember: \"git status\" is basically shorthand for \"what would happen if I \ndid a \"git commit\" with these arguments.\n\nWhich means that \"git status .\" basically is something similar to a \nprivate invocation of \"git add -u .\" in addition to the regular git \nstatus.\n\nSo try it out: change some file (let's say the top-level Makefile) and do \nthe two operations, and see how the _output_ is totally different:\n\nWithout the \".\", you should see something like:\n\n\t# On branch master\n\t# Changed but not updated:\n\t#   (use \"git add <file>...\" to update what will be committed)\n\t#\n\t#       modified:   Makefile\n\t#\n\tno changes added to commit (use \"git add\" and/or \"git commit -a\")\n\nie there is nothing to commit, but Makefile is modified and _could_ be \ncommitted.\n\nNow, don't look down at the answer, but instead try to think it through: \nwhat is \"git status .\" going to say?\n\nAnswer: it's going to show something totally different, because \"git \ncommit .\" is going to add that changed Makefile to the commit, so by the \nlogic that \"git status\" is supposed to show what commit will do, you'll \n*not* see that \"no changes\" line at all, but instead you'll see\n\n\t# On branch master\n\t# Changes to be committed:\n\t#   (use \"git reset HEAD <file>...\" to unstage)\n\t#\n\t#       modified:   Makefile\n\nie now we *would* commit that Makefile change!\n\nSo the reason \"git status .\" is more expensive is that it's doing \nsomething else.\n\nI don't know what \"hg status .\" means, but I suspect that it's more of a \n\"same thing as 'hg status', but limited to '.'\".\n\nAnd yes, most of the time in \"git status .\" is going to be the lstat() \ncalls. Which are expensive on OS X. And yes, we do too many of them. I'll \nlook at seeing if we can avoid some.\n\n\t\tLinus\n"},{"id":"74771","messageId":"FEFAB19F-742A-452E-87C1-CD55AD0996DB@ai.rug.nl","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191443550.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Pieter de Bie","fromEmail":"pdebie@ai.rug.nl","sentAt":"2008-04-19T22:00:58Z","receivedAt":"2008-04-19T22:00:58Z","isPatch":false,"sender":{"key":"pdebie@ai.rug.nl","avatar":null},"body":"\nOn 19 apr 2008, at 23:54, Linus Torvalds wrote:\n>\n>\n> Btw, the \"git status\" issue is totally different.\n>\n\nYes, I was aware of that, but still found it interesting to test.\n>\n> And yes, most of the time in \"git status .\" is going to be the lstat()\n> calls. Which are expensive on OS X. And yes, we do too many of them.  \n> I'll\n> look at seeing if we can avoid some.\n\nI just tested this. \"git status .\" does 428815 (400k!) lstats, almost  \n10x as many as there are files in the repository. I'd agree that this  \nis the reason it's slow on OS X :).\n\n- Pieter\n"},{"id":"74772","messageId":"A80AA52C-4254-4580-83AA-1B3134B74430@ai.rug.nl","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191422480.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Pieter de Bie","fromEmail":"pdebie@ai.rug.nl","sentAt":"2008-04-19T22:08:15Z","receivedAt":"2008-04-19T22:08:15Z","isPatch":false,"sender":{"key":"pdebie@ai.rug.nl","avatar":null},"body":"\nOn 19 apr 2008, at 23:29, Linus Torvalds wrote:\n> Side note: on the kenrel tree, it makes the (insane!) operation\n>\n> \tgit add $(git ls-files)\n>\n> go from 49 seconds down to 17 sec. So it does make a huge difference  \n> for\n> me, but I also want to point out that this really isn't a sane  \n> operation\n> to do (I also think that 17 sec is totally unacceptable, but I  \n> cannot find\n> it in me to care, since I don't think this is an operation that  \n> anybody\n> should ever do!)\n\nYes, using the patch decreases the time for my from ~40 seconds to  \njust about 4 seconds.\n\n- Pieter\n"},{"id":"74775","messageId":"alpine.LFD.1.10.0804191515120.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"FEFAB19F-742A-452E-87C1-CD55AD0996DB@ai.rug.nl","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-19T22:39:37Z","receivedAt":"2008-04-19T22:39:37Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Apr 2008, Pieter de Bie wrote:\n>\n> I just tested this. \"git status .\" does 428815 (400k!) lstats, almost 10x as\n> many as there are files in the repository. I'd agree that this is the reason\n> it's slow on OS X :).\n\nYeah. I didn't look any further, but we do a total of *nine* 'lstat()' \ncalls for each file we know about that is dirty, and *seven* when they are \nclean. Plus maybe a few more.\n\nI had some patches that cut that down a lot, but some of it had to be \nreverted because of the subtle interactions with different internal copies \nof the index, and I think the case with a partial commit (which is what \nyou have when using a pathspec) was the case that got reverted.\n\nMaybe I should revisit it, now that we have internal support for actually \nkeeping multiple indexes in place and being able to merge them.\n\nAppended, in case somebody is interested, are the callchains for the \ndifferent lstat() callers (for the case of a single file that was clean).\n\n(This is with a non-pristine git source-base, so the line numbers may not \nmatch 100%, but the changes are pretty small, so it should still be an \ninteresting set of callers)\n\nAt least two of them are due to ce_smudge_racily_clean_entry(), which in \nturn is because we trigger the is_racy_timestamp() test. Hmm.\n\nAnd four of them are because of two cases of the pattern\n\n\tif (file_exists())\n\t\tadd_file_to_cache(p->path, 0);\n\nwhere the \"file_exists()\" first does an lstat() to see if it exists, and \nthen \"add_file_to_cache()\" does an lstat() to get the stat info..\n\n\t\t\tLinus\n\n---\nFirst:\n\t#1  0x0000000000481fd7 in file_exists (f=0x1ac3f10 \"Makefile\") at dir.c:742\n\t#2  0x00000000004196bd in add_remove_files (list=0x7fff87223310) at builtin-commit.c:179\n\t#3  0x0000000000419aa1 in prepare_index (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:308\n\t#4  0x000000000041b3b2 in cmd_status (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:781\n\t#5  0x000000000040482e in run_command (p=0x704eb8, argc=2, argv=0x7fff872235d0) at git.c:264\n\t#6  0x00000000004049db in handle_internal_command (argc=2, argv=0x7fff872235d0) at git.c:394\n\t#7  0x0000000000404b47 in main (argc=2, argv=0x7fff872235d0) at git.c:458\n\nSecond:\n\t#1  0x00000000004973a0 in add_file_to_index (istate=0x7494c0, path=0x1ac3f10 \"Makefile\", verbose=0) at read-cache.c:471\n\t#2  0x00000000004196d7 in add_remove_files (list=0x7fff87223310) at builtin-commit.c:180\n\t#3  0x0000000000419aa1 in prepare_index (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:308\n\t#4  0x000000000041b3b2 in cmd_status (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:781\n\t#5  0x000000000040482e in run_command (p=0x704eb8, argc=2, argv=0x7fff872235d0) at git.c:264\n\t#6  0x00000000004049db in handle_internal_command (argc=2, argv=0x7fff872235d0) at git.c:394\n\t#7  0x0000000000404b47 in main (argc=2, argv=0x7fff872235d0) at git.c:458\n\nThird:\n\t#1  0x00000000004992a4 in ce_smudge_racily_clean_entry (ce=0x1ac4000) at read-cache.c:1267\n\t#2  0x000000000049956d in write_index (istate=0x7494c0, newfd=6) at read-cache.c:1348\n\t#3  0x0000000000419ac7 in prepare_index (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:310\n\t#4  0x000000000041b3b2 in cmd_status (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:781\n\t#5  0x000000000040482e in run_command (p=0x704eb8, argc=2, argv=0x7fff872235d0) at git.c:264\n\t#6  0x00000000004049db in handle_internal_command (argc=2, argv=0x7fff872235d0) at git.c:394\n\t#7  0x0000000000404b47 in main (argc=2, argv=0x7fff872235d0) at git.c:458\n\nFourth:\n\t#1  0x0000000000481fd7 in file_exists (f=0x1ac3f10 \"Makefile\") at dir.c:742\n\t#2  0x00000000004196bd in add_remove_files (list=0x7fff87223310) at builtin-commit.c:179\n\t#3  0x0000000000419b21 in prepare_index (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:318\n\t#4  0x000000000041b3b2 in cmd_status (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:781\n\t#5  0x000000000040482e in run_command (p=0x704eb8, argc=2, argv=0x7fff872235d0) at git.c:264\n\t#6  0x00000000004049db in handle_internal_command (argc=2, argv=0x7fff872235d0) at git.c:394\n\t#7  0x0000000000404b47 in main (argc=2, argv=0x7fff872235d0) at git.c:458\n\nFifth:\n\t#1  0x00000000004973a0 in add_file_to_index (istate=0x7494c0, path=0x1ac3f10 \"Makefile\", verbose=0) at read-cache.c:471\n\t#2  0x00000000004196d7 in add_remove_files (list=0x7fff87223310) at builtin-commit.c:180\n\t#3  0x0000000000419b21 in prepare_index (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:318\n\t#4  0x000000000041b3b2 in cmd_status (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:781\n\t#5  0x000000000040482e in run_command (p=0x704eb8, argc=2, argv=0x7fff872235d0) at git.c:264\n\t#6  0x00000000004049db in handle_internal_command (argc=2, argv=0x7fff872235d0) at git.c:394\n\t#7  0x0000000000404b47 in main (argc=2, argv=0x7fff872235d0) at git.c:458\n\nSixth:\n\t#1  0x00000000004992a4 in ce_smudge_racily_clean_entry (ce=0x1ac4740) at read-cache.c:1267\n\t#2  0x000000000049956d in write_index (istate=0x7494c0, newfd=6) at read-cache.c:1348\n\t#3  0x0000000000419b47 in prepare_index (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:321\n\t#4  0x000000000041b3b2 in cmd_status (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:781\n\t#5  0x000000000040482e in run_command (p=0x704eb8, argc=2, argv=0x7fff872235d0) at git.c:264\n\t#6  0x00000000004049db in handle_internal_command (argc=2, argv=0x7fff872235d0) at git.c:394\n\t#7  0x0000000000404b47 in main (argc=2, argv=0x7fff872235d0) at git.c:458\n\nSeventh:\n\t#1  0x0000000000475a37 in check_work_tree_entity (ce=0x1ac4870, st=0x7fff87221ea0, symcache=0x7fff87221f30 \"\") at diff-lib.c:343\n\t#2  0x0000000000475eab in run_diff_files (revs=0x7fff87222fc0, option=0) at diff-lib.c:466\n\t#3  0x00000000004bd62e in wt_status_print_changed (s=0x7fff872232f0) at wt-status.c:220\n\t#4  0x00000000004bdb1d in wt_status_print (s=0x7fff872232f0) at wt-status.c:310\n\t#5  0x0000000000419c0b in run_status (fp=0x37d8751760, index_file=0x711ef1 \"/home/torvalds/git-test/.git/next-index-14646.lock\", prefix=0x0, \n\t    nowarn=0) at builtin-commit.c:349\n\t#6  0x000000000041b3cf in cmd_status (argc=1, argv=0x7fff872235d0, prefix=0x0) at builtin-commit.c:783\n\t#7  0x000000000040482e in run_command (p=0x704eb8, argc=2, argv=0x7fff872235d0) at git.c:264\n\t#8  0x00000000004049db in handle_internal_command (argc=2, argv=0x7fff872235d0) at git.c:394\n\t#9  0x0000000000404b47 in main (argc=2, argv=0x7fff872235d0) at git.c:458\n"},{"id":"74776","messageId":"m3od85qxcl.fsf@localhost.localdomain","threadId":"13180","inReplyTo":"FEFAB19F-742A-452E-87C1-CD55AD0996DB@ai.rug.nl","subject":"Re: Git performance on OS X","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2008-04-19T22:44:05Z","receivedAt":"2008-04-19T22:44:05Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Pieter de Bie <pdebie@ai.rug.nl> writes:\n\n> On 19 apr 2008, at 23:54, Linus Torvalds wrote:\n> >\n> > And yes, most of the time in \"git status .\" is going to be the lstat()\n> > calls. Which are expensive on OS X. And yes, we do too many of them.\n> > I'll\n> > look at seeing if we can avoid some.\n> \n> I just tested this. \"git status .\" does 428815 (400k!) lstats, almost\n> 10x as many as there are files in the repository. I'd agree that this\n> is the reason it's slow on OS X :).\n\nBy the way, what version of git do you use? Because in RelNotes for\n1.5.5 there is:\n\n * \"git commit\" does not run lstat(2) more than necessary\n   anymore.\n\nwhich I guess also apply to git status.  This change was written by\nLinus if I remember correctly...\n\n-- \nJakub Narebski\nPoland\nShadeHawk on #git\n"},{"id":"74778","messageId":"alpine.LFD.1.10.0804191547320.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"m3od85qxcl.fsf@localhost.localdomain","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-19T22:50:29Z","receivedAt":"2008-04-19T22:50:29Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 19 Apr 2008, Jakub Narebski wrote:\n> \n> By the way, what version of git do you use? Because in RelNotes for\n> 1.5.5 there is:\n> \n>  * \"git commit\" does not run lstat(2) more than necessary\n>    anymore.\n> \n> which I guess also apply to git status.  This change was written by\n> Linus if I remember correctly...\n\nThat's only true for a plain \"git status\" (and even there it does *one* \nsuperfluous lstat()).\n\nIf you do \"git status .\" it still does a _lot_ of unnecessary commits.\n\nThis patch will help. It removes the one superfluous lstat() from \"git \nstatus\", and for \"git status .\" it removes two of them. Basically the old \npattern\n\n\tif (file_exists(..))\n\t\tadd_file_to_cache(..)\n\nnow uses just one lstat() and shares the data between the two.\n\n\t\tLinus\n\n---\n builtin-commit.c |    6 ++++--\n cache.h          |    2 ++\n read-cache.c     |   29 +++++++++++++++++------------\n 3 files changed, 23 insertions(+), 14 deletions(-)\n\ndiff --git a/builtin-commit.c b/builtin-commit.c\nindex bcb7aaa..fb3e359 100644\n--- a/builtin-commit.c\n+++ b/builtin-commit.c\n@@ -175,9 +175,11 @@ static void add_remove_files(struct path_list *list)\n {\n \tint i;\n \tfor (i = 0; i < list->nr; i++) {\n+\t\tstruct stat st;\n \t\tstruct path_list_item *p = &(list->items[i]);\n-\t\tif (file_exists(p->path))\n-\t\t\tadd_file_to_cache(p->path, 0);\n+\n+\t\tif (!lstat(p->path, &st))\n+\t\t\tadd_to_cache(p->path, &st, 0);\n \t\telse\n \t\t\tremove_file_from_cache(p->path);\n \t}\ndiff --git a/cache.h b/cache.h\nindex 1b66cc0..c058125 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -261,6 +261,7 @@ static inline void remove_name_hash(struct cache_entry *ce)\n #define add_cache_entry(ce, option) add_index_entry(&the_index, (ce), (option))\n #define remove_cache_entry_at(pos) remove_index_entry_at(&the_index, (pos))\n #define remove_file_from_cache(path) remove_file_from_index(&the_index, (path))\n+#define add_to_cache(path, st, verbose) add_to_index(&the_index, (path), (st), (verbose))\n #define add_file_to_cache(path, verbose) add_file_to_index(&the_index, (path), (verbose))\n #define refresh_cache(flags) refresh_index(&the_index, (flags), NULL, NULL)\n #define ce_match_stat(ce, st, options) ie_match_stat(&the_index, (ce), (st), (options))\n@@ -364,6 +365,7 @@ extern int add_index_entry(struct index_state *, struct cache_entry *ce, int opt\n extern struct cache_entry *refresh_cache_entry(struct cache_entry *ce, int really);\n extern int remove_index_entry_at(struct index_state *, int pos);\n extern int remove_file_from_index(struct index_state *, const char *path);\n+extern int add_to_index(struct index_state *, const char *path, struct stat *, int verbose);\n extern int add_file_to_index(struct index_state *, const char *path, int verbose);\n extern struct cache_entry *make_cache_entry(unsigned int mode, const unsigned char *sha1, const char *path, int stage, int refresh);\n extern int ce_same_name(struct cache_entry *a, struct cache_entry *b);\ndiff --git a/read-cache.c b/read-cache.c\nindex 15d3d72..4c1a9f4 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -461,21 +461,18 @@ static struct cache_entry *create_alias_ce(struct cache_entry *ce, struct cache_\n \treturn new;\n }\n \n-int add_file_to_index(struct index_state *istate, const char *path, int verbose)\n+int add_to_index(struct index_state *istate, const char *path, struct stat *st, int verbose)\n {\n \tint size, namelen;\n-\tstruct stat st;\n+\tmode_t st_mode = st->st_mode;\n \tstruct cache_entry *ce, *alias;\n \tunsigned ce_option = CE_MATCH_IGNORE_VALID|CE_MATCH_RACY_IS_DIRTY;\n \n-\tif (lstat(path, &st))\n-\t\tdie(\"%s: unable to stat (%s)\", path, strerror(errno));\n-\n-\tif (!S_ISREG(st.st_mode) && !S_ISLNK(st.st_mode) && !S_ISDIR(st.st_mode))\n+\tif (!S_ISREG(st_mode) && !S_ISLNK(st_mode) && !S_ISDIR(st_mode))\n \t\tdie(\"%s: can only add regular files, symbolic links or git-directories\", path);\n \n \tnamelen = strlen(path);\n-\tif (S_ISDIR(st.st_mode)) {\n+\tif (S_ISDIR(st_mode)) {\n \t\twhile (namelen && path[namelen-1] == '/')\n \t\t\tnamelen--;\n \t}\n@@ -483,10 +480,10 @@ int add_file_to_index(struct index_state *istate, const char *path, int verbose)\n \tce = xcalloc(1, size);\n \tmemcpy(ce->name, path, namelen);\n \tce->ce_flags = namelen;\n-\tfill_stat_cache_info(ce, &st);\n+\tfill_stat_cache_info(ce, st);\n \n \tif (trust_executable_bit && has_symlinks)\n-\t\tce->ce_mode = create_ce_mode(st.st_mode);\n+\t\tce->ce_mode = create_ce_mode(st_mode);\n \telse {\n \t\t/* If there is an existing entry, pick the mode bits and type\n \t\t * from it, otherwise assume unexecutable regular file.\n@@ -495,18 +492,18 @@ int add_file_to_index(struct index_state *istate, const char *path, int verbose)\n \t\tint pos = index_name_pos_also_unmerged(istate, path, namelen);\n \n \t\tent = (0 <= pos) ? istate->cache[pos] : NULL;\n-\t\tce->ce_mode = ce_mode_from_stat(ent, st.st_mode);\n+\t\tce->ce_mode = ce_mode_from_stat(ent, st_mode);\n \t}\n \n \talias = index_name_exists(istate, ce->name, ce_namelen(ce), ignore_case);\n-\tif (alias && !ce_stage(alias) && !ie_match_stat(istate, alias, &st, ce_option)) {\n+\tif (alias && !ce_stage(alias) && !ie_match_stat(istate, alias, st, ce_option)) {\n \t\t/* Nothing changed, really */\n \t\tfree(ce);\n \t\tce_mark_uptodate(alias);\n \t\talias->ce_flags |= CE_ADDED;\n \t\treturn 0;\n \t}\n-\tif (index_path(ce->sha1, path, &st, 1))\n+\tif (index_path(ce->sha1, path, st, 1))\n \t\tdie(\"unable to index file %s\", path);\n \tif (ignore_case && alias && different_name(ce, alias))\n \t\tce = create_alias_ce(ce, alias);\n@@ -518,6 +515,14 @@ int add_file_to_index(struct index_state *istate, const char *path, int verbose)\n \treturn 0;\n }\n \n+int add_file_to_index(struct index_state *istate, const char *path, int verbose)\n+{\n+\tstruct stat st;\n+\tif (lstat(path, &st))\n+\t\tdie(\"%s: unable to stat (%s)\", path, strerror(errno));\n+\treturn add_to_index(istate, path, &st, verbose);\n+}\n+\n struct cache_entry *make_cache_entry(unsigned int mode,\n \t\tconst unsigned char *sha1, const char *path, int stage,\n \t\tint refresh)\n"},{"id":"74779","messageId":"alpine.LFD.1.10.0804191551540.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191547320.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-19T22:54:04Z","receivedAt":"2008-04-19T22:54:04Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 19 Apr 2008, Linus Torvalds wrote:\n> \n> This patch will help.\n\nPieter? Assuming the lstat() cost is the dominant one, it should cut down \nyour \"git status .\" cost by about 15-20% or so. Can you confirm?\n\n\t\tLinus\n"},{"id":"74780","messageId":"alpine.LFD.1.10.0804191603020.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191547320.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-19T23:04:19Z","receivedAt":"2008-04-19T23:04:19Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 19 Apr 2008, Linus Torvalds wrote:\n> \n> If you do \"git status .\" it still does a _lot_ of unnecessary commits.\n                                                                ^^^^^^^\n\nI meant 'lstat()'s, of course.\n\nI think I need to take my alzheimer medication now.\n\n\t\t\tLinus\n"},{"id":"74781","messageId":"0BE9BBE3-EA9D-4A66-A086-A2A1B289B0DD@ai.rug.nl","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191551540.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Pieter de Bie","fromEmail":"pdebie@ai.rug.nl","sentAt":"2008-04-19T23:10:55Z","receivedAt":"2008-04-19T23:10:55Z","isPatch":false,"sender":{"key":"pdebie@ai.rug.nl","avatar":null},"body":"\nOn 20 apr 2008, at 00:54, Linus Torvalds wrote:\n> Pieter? Assuming the lstat() cost is the dominant one, it should cut  \n> down\n> your \"git status .\" cost by about 15-20% or so. Can you confirm?\n\nThe number of lstats are cut down by your patch: 428761 vs. 338091.  \nFunnily enough, there is no significant difference in run-time:\n\nCommand                                            Mean     Std\ngit status .                                       13.970  1.298\n/Users/pieter/projects/External/git/git-status .   13.759  0.321\n\n\nSystem times are also approximately the same (10.79s vs 10.43s).\n\n- Pieter\n"},{"id":"74782","messageId":"alpine.LFD.1.10.0804191619240.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"0BE9BBE3-EA9D-4A66-A086-A2A1B289B0DD@ai.rug.nl","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-19T23:26:19Z","receivedAt":"2008-04-19T23:26:19Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Apr 2008, Pieter de Bie wrote:\n> \n> The number of lstats are cut down by your patch: 428761 vs. 338091. Funnily\n> enough, there is no significant difference in run-time:\n\nIt may be that the problem with OS X is a sucky pathname cache mechanism.\n\nThe trivial patch cut down the number of stat() calls by a fair amount, \nbut the calls that got removed were all of the \"do two 'lstat()' calls on \nthe exact same pathname consecutively\" type.\n\nMaybe OS X has some very limited pathname caching that catches that, or \neven if not, it just ends up being very nice in the D$, so it's not a big \ndeal. And then the real suckiness happens only with bigger workloads.\n\nIt may also be that the bulk of the OS X cost isn't in lstat() at all, but \nin the VM. That was true for some other OS X load.\n\n> Command                                            Mean     Std\n> git status .                                       13.970  1.298\n> /Users/pieter/projects/External/git/git-status .   13.759  0.321\n\nThis is the WebKit archive, right?\n\nFor me, doing a \"time git status .\" on the WebKit thing I just cloned from \ngit://git.webkit.org/WebKit.git is much faster: 1.264s (and it goes down \nby maybe 5-10% with my lstat-avoidance patch).\n\nIs there any system-level profiler for OS X to get a clue where that cost \nis, in case it's not the lstat() at all?\n\n\t\t\tLinus\n"},{"id":"74783","messageId":"A07C1A99-084B-4DFC-90CB-B8BDAF7E72EF@sun.com","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191619240.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Roman Shaposhnik","fromEmail":"rvs@sun.com","sentAt":"2008-04-19T23:35:55Z","receivedAt":"2008-04-19T23:35:55Z","isPatch":false,"sender":{"key":"rvs@sun.com","avatar":null},"body":"On Apr 19, 2008, at 4:26 PM, Linus Torvalds wrote:\n> This is the WebKit archive, right?\n>\n> For me, doing a \"time git status .\" on the WebKit thing I just  \n> cloned from\n> git://git.webkit.org/WebKit.git is much faster: 1.264s (and it goes  \n> down\n> by maybe 5-10% with my lstat-avoidance patch).\n>\n> Is there any system-level profiler for OS X to get a clue where  \n> that cost\n> is, in case it's not the lstat() at all?\n\nIf it happens on Leopard, DTrace would be a perfect way to query the  \nsystem:\n\n   $ dtrace -n 'syscall::*:entry /pid==$target/ { @[probefunc] = count \n(); }' -c \"git <do stuff>\"\n\nE.g.:\n\n$ dtrace -n 'syscall::*:entry /pid==$target/ { @[probefunc] = count \n(); }' -c \"echo Hello World\"\ndtrace: description 'syscall::*:entry ' matched 234 probes\nHello World\ndtrace: pid 1325 has exited\n\n   fstat64                                                           1\n   getpid                                                            1\n   getrlimit                                                         1\n   ioctl                                                             1\n   mmap                                                              1\n   munmap                                                            1\n   rexit                                                             1\n   sysi86                                                            1\n   setcontext                                                        2\n   write                                                             2\n\n\nThanks,\nRoman.\n"},{"id":"74784","messageId":"2F8F3BF2-66F9-473C-BE82-8F784E1FF9A4@ai.rug.nl","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191619240.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Pieter de Bie","fromEmail":"pdebie@ai.rug.nl","sentAt":"2008-04-19T23:56:19Z","receivedAt":"2008-04-19T23:56:19Z","isPatch":false,"sender":{"key":"pdebie@ai.rug.nl","avatar":null},"body":"\nOn 20 apr 2008, at 01:26, Linus Torvalds wrote:\n>\n> It may be that the problem with OS X is a sucky pathname cache  \n> mechanism.\n>\n> The trivial patch cut down the number of stat() calls by a fair  \n> amount,\n> but the calls that got removed were all of the \"do two 'lstat()'  \n> calls on\n> the exact same pathname consecutively\" type.\n>\n> Maybe OS X has some very limited pathname caching that catches that,  \n> or\n> even if not, it just ends up being very nice in the D$, so it's not  \n> a big\n> deal. And then the real suckiness happens only with bigger workloads.\n>\n\nYes, I just tested this.\n\n\tfor (int i = 0; i < 50000; i++) {\n\t\tsprintf(s, \"/Users/pieter/test/perf/%i\", i);\n\t\tint ret = lstat(s, a);\n\t}\n\nThis loop needs about 3 seconds to  run. Replacing the i with 10 in  \nthe sprintf reduces it to 0.24seconds.\n\n>> Command                                            Mean     Std\n>> git status .                                       13.970  1.298\n>> /Users/pieter/projects/External/git/git-status .   13.759  0.321\n>\n> This is the WebKit archive, right?\n>\n> For me, doing a \"time git status .\" on the WebKit thing I just  \n> cloned from\n> git://git.webkit.org/WebKit.git is much faster: 1.264s (and it goes  \n> down\n> by maybe 5-10% with my lstat-avoidance patch).\n>\n> Is there any system-level profiler for OS X to get a clue where that  \n> cost\n> is, in case it's not the lstat() at all?\n\nYes, that was the webkit repo (the test above was in a dir with 50k  \nfiles).\n\nAlas, I tried to create a nice profiling for the \"git status .\". In  \nthe Instruments application I can create a sampler, but I see no way  \nto export it. The option to export the script as a dtrace script is  \ngreyed out in the menu.\n\n From the sampler, it appears that the lstat calls still account for  \nmost of the time. I have uploaded a screenshot to http://ss.frim.nl/==759.png \n. It actually shows quite nicely when the lstats are being done --  \nit's when the CPU is idle. Next to the lstats, the read_tree_recursive  \nis also called often.\n\n- Pieter\n"},{"id":"74785","messageId":"AAECC98F-2B6A-4785-8BED-BB4B86F2F2D6@ai.rug.nl","threadId":"13180","inReplyTo":"A07C1A99-084B-4DFC-90CB-B8BDAF7E72EF@sun.com","subject":"Re: Git performance on OS X","fromName":"Pieter de Bie","fromEmail":"pdebie@ai.rug.nl","sentAt":"2008-04-19T23:57:41Z","receivedAt":"2008-04-19T23:57:41Z","isPatch":false,"sender":{"key":"pdebie@ai.rug.nl","avatar":null},"body":"\nOn 20 apr 2008, at 01:35, Roman Shaposhnik wrote:\n> If it happens on Leopard, DTrace would be a perfect way to query the  \n> system:\n>\n>  $ dtrace -n 'syscall::*:entry /pid==$target/ { @[probefunc] =  \n> count(); }' -c \"git <do stuff>\"\n\nYes, but this only shows syscalls and only entries. If you have enough  \ndtrace-fu to create a script that samples all functions (and thus  \nshows which functions are being active most of the time), I will  \ngladly run it.\n\n- Pieter\n"},{"id":"74786","messageId":"alpine.LFD.1.10.0804191658430.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"A07C1A99-084B-4DFC-90CB-B8BDAF7E72EF@sun.com","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-20T00:06:52Z","receivedAt":"2008-04-20T00:06:52Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 19 Apr 2008, Roman Shaposhnik wrote:\n> > \n> > Is there any system-level profiler for OS X to get a clue where that cost\n> > is, in case it's not the lstat() at all?\n> \n> If it happens on Leopard, DTrace would be a perfect way to query the system:\n\nWell, I'd really like to see a traditional _time_ profile, not system \ncall counts.\n\nThe system call profile is trivial - it's generally going to be pretty \nsimilar under OS X and Linux (modulo library differences, but git doesn't \nreally use any really complex libraries that would do system calls).\n\nThe problem we've had in the past is that Linux is simply an order of \nmagnitude faster (sometimes more) at some operations than OS X is, so \nissues that show up on OS X don't even show up on Linux.\n\nThis was the case for doing lots of small \"mmap()/munmap()\" calls, for \nexample, where we literally had a load where OS X was two orders of \nmagnitude slower. We switched over to reading the files with \"pread()\" \ninstead of mmap(), and that fixed that particular issue.\n\nSo a real system profile would be nice.\n\n\t\t\tLinus\n"},{"id":"74787","messageId":"F9D9143D-D849-4454-91CD-024D691652C3@sun.com","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191658430.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Roman Shaposhnik","fromEmail":"rvs@sun.com","sentAt":"2008-04-20T00:21:13Z","receivedAt":"2008-04-20T00:21:13Z","isPatch":false,"sender":{"key":"rvs@sun.com","avatar":null},"body":"On Apr 19, 2008, at 5:06 PM, Linus Torvalds wrote:\n> On Sat, 19 Apr 2008, Roman Shaposhnik wrote:\n>>>\n>>> Is there any system-level profiler for OS X to get a clue where  \n>>> that cost\n>>> is, in case it's not the lstat() at all?\n>>\n>> If it happens on Leopard, DTrace would be a perfect way to query  \n>> the system:\n>\n> Well, I'd really like to see a traditional _time_ profile, not system\n> call counts.\n\nGood point. I just thought your original question was simply about  \nconfirming\na hunch of an enormous # of lstat syscalls. Now, if by _time_ profile\nyou mean how much time gets spent in each syscall than the following\nshould help (time is in nanoseconds):\n\n$ dtrace -n 'syscall::*:entry /pid==$target/ { self->ts=timestamp; }  \nsyscall::*:return /pid==$target/ { @[probefunc]=sum(timestamp-self- \n >ts); }' -c \"echo Hello World\"\ndtrace: description 'syscall::*:entry ' matched 468 probes\nHello World\ndtrace: pid 1400 has exited\n\n   getpid                                                         1392\n   sysi86                                                         2799\n   getrlimit                                                      2918\n   setcontext                                                     4273\n   fstat64                                                        6220\n   mmap                                                          15419\n   munmap                                                        22593\n   write                                                         27860\n   ioctl                                                         30750\n\nThe script can be modified slightly to also profile all of the libc.\n\nOn the other hand, if by real system profile you mean a full fledged\nsampling of the application itself than I can suggest running Git\nunder Shark, and not under Instruments. Although I'm extremly\ncurious about *why* would you need a full fledged *application*\nlevel profile of Git as opposed to a profile of how Git interacts\nwith an OS.\n\n> The system call profile is trivial - it's generally going to be pretty\n> similar under OS X and Linux (modulo library differences, but git  \n> doesn't\n> really use any really complex libraries that would do system calls).\n>\n> The problem we've had in the past is that Linux is simply an order of\n> magnitude faster (sometimes more) at some operations than OS X is, so\n> issues that show up on OS X don't even show up on Linux.\n>\n> This was the case for doing lots of small \"mmap()/munmap()\" calls, for\n> example, where we literally had a load where OS X was two orders of\n> magnitude slower. We switched over to reading the files with \"pread()\"\n> instead of mmap(), and that fixed that particular issue.\n\nTotally understood!\n\nThanks,\nRoman.\n"},{"id":"74788","messageId":"alpine.LFD.1.10.0804191727270.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"2F8F3BF2-66F9-473C-BE82-8F784E1FF9A4@ai.rug.nl","subject":"Re: Git performance on OS X","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-20T00:31:41Z","receivedAt":"2008-04-20T00:31:41Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Apr 2008, Pieter de Bie wrote:\n> \n> Yes, I just tested this.\n> \n> \tfor (int i = 0; i < 50000; i++) {\n> \t\tsprintf(s, \"/Users/pieter/test/perf/%i\", i);\n> \t\tint ret = lstat(s, a);\n> \t}\n> \n> This loop needs about 3 seconds to  run. Replacing the i with 10 in the\n> sprintf reduces it to 0.24seconds.\n\nOk.\n\nOn my machine, that's\n\n\treal    0m0.090s\n\nwith the 50,000 different files, and with the same filename it's\n\n\treal    0m0.081s\n\nso yes, we're looking at another case of Linux performance just being in a \nclass of its own.\n\nTaking three seconds for the warm-cache case for just 50,000 files is \nludicrous. That's about an order-and-a-half slower than what I see.\n\nMaybe my CPU is faster too (2.66GHz Core 2), but the thing is, Linux \nreally does tend to outperform others at a lot of these kinds of loads. \nSystem calls are fast to begin with, and the Linux directory cache kicks \nass, if I do say so myself.\n\nOS X doth suck. \n\n\t\t\tLinus\n"},{"id":"74791","messageId":"20080420012314.GS3133@dpotapov.dyndns.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191727270.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-04-20T01:23:14Z","receivedAt":"2008-04-20T01:23:14Z","isPatch":false,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Sat, Apr 19, 2008 at 05:31:41PM -0700, Linus Torvalds wrote:\n> \n> Maybe my CPU is faster too (2.66GHz Core 2), but the thing is, Linux \n> really does tend to outperform others at a lot of these kinds of loads. \n> System calls are fast to begin with, and the Linux directory cache kicks \n> ass, if I do say so myself.\n\nYes, Linux is really fast. Even with relatively old AMD Sempron 1.8 GHz,\nI got the following numbers:\nreal    0m0.177s\nreal    0m0.154s\n\nDmitry\n"},{"id":"74796","messageId":"7viqydduxk.fsf@gitster.siamese.dyndns.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191515120.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-04-20T04:14:31Z","receivedAt":"2008-04-20T04:14:31Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> At least two of them are due to ce_smudge_racily_clean_entry(), which in\n> turn is because we trigger the is_racy_timestamp() test. Hmm.\n\nThere is one change that has been held back in 'next' for quite some time.\nPerhaps it would help?\n"},{"id":"74800","messageId":"20080420111346.GA13411@bit.office.eurotux.com","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191515120.2779@woody.linux-foundation.org","subject":"[PATCH 01/02/RFC] implement a stat cache","fromName":"Luciano Rocha","fromEmail":"luciano@eurotux.com","sentAt":"2008-04-20T11:13:46Z","receivedAt":"2008-04-20T11:13:46Z","isPatch":true,"sender":{"key":"luciano@eurotux.com","avatar":null},"body":"An implementation of stat(2) and lstat(2) caching. Both the return code\nand returned information are cached.\n\nSigned-off-by: Luciano Rocha <strange@nsk.no-ip.org>\n---\nOn Sat, Apr 19, 2008 at 03:39:37PM -0700, Linus Torvalds wrote:\n> Yeah. I didn't look any further, but we do a total of *nine* 'lstat()' \n> calls for each file we know about that is dirty, and *seven* when they are \n> clean. Plus maybe a few more.\n\nThat's a lot. Why not use a stat cache?\n\nWith these changes, my git status . in WebKit changes from 28.215s to\n15.414s.\n\n stat-cache.c |   69 ++++++++++++++++++++++++++++++++++++++++++++++++++++++++++\n stat-cache.h |    9 +++++++\n 2 files changed, 78 insertions(+), 0 deletions(-)\n create mode 100644 stat-cache.c\n create mode 100644 stat-cache.h\n\ndiff --git a/stat-cache.c b/stat-cache.c\nnew file mode 100644\nindex 0000000..6a33cec\n--- /dev/null\n+++ b/stat-cache.c\n@@ -0,0 +1,69 @@\n+/*\n+ * Cache (l)stat operations\n+ */\n+\n+#include \"stat-cache.h\"\n+#include \"hash.h\"\n+#include \"path-list.h\";\n+\n+static struct hash_table stat_cache;\n+static struct hash_table lstat_cache;\n+\n+struct stat_result {\n+\tstruct stat st;\n+\tint ret;\n+};\n+\n+/* based on hash_name from read_cache.c */\n+static unsigned int hash_path(const char *path)\n+{\n+\tunsigned int hash = 0x123;\n+\n+\twhile (*path)\n+\t\thash = hash*101 + *path++;\n+\treturn hash;\n+}\n+\n+/* cache is HASH->PATH-LIST->(return code, struct stat) */\n+static int cached_stat(int (*f)(const char *, struct stat *),\n+\t\tstruct hash_table *ht, const char *path, struct stat *buf)\n+{\n+\tunsigned int hash;\n+\tstruct path_list *list;\n+\tstruct path_list_item *cached;\n+\tstruct stat_result *result;\n+\n+\thash = hash_path(path);\n+\n+\tlist = lookup_hash(hash, ht);\n+\n+\tif (!list) {\n+\t\tlist = xcalloc(1, sizeof *list);\n+\t\tlist->strdup_paths = 1;\n+\t\tinsert_hash(hash, list, ht);\n+\t}\n+\n+\tcached = path_list_lookup(path, list);\n+\n+\tif (cached) {\n+\t\tresult = cached->util;\n+\t} else {\n+\t\tresult = xmalloc(sizeof *result);\n+\t\tresult->ret = f(path, &result->st);\n+\t\tpath_list_insert(path, list)->util = result;\n+\t}\n+\n+\tif (result->ret == 0)\n+\t\tmemcpy(buf, &result->st, sizeof *buf);\n+\treturn result->ret;\n+}\n+\n+int cstat(const char *path, struct stat *buf)\n+{\n+\treturn cached_stat(stat, &stat_cache, path, buf);\n+}\n+\n+int clstat(const char *path, struct stat *buf)\n+{\n+\treturn cached_stat(lstat, &lstat_cache, path, buf);\n+}\ndiff --git a/stat-cache.h b/stat-cache.h\nnew file mode 100644\nindex 0000000..754348f\n--- /dev/null\n+++ b/stat-cache.h\n@@ -0,0 +1,9 @@\n+#ifndef STAT_CACHE_H\n+#define STAT_CACHE_H\n+\n+#include \"git-compat-util.h\"\n+\n+int cstat(const char *path, struct stat *buf);\n+int clstat(const char *path, struct stat *buf);\n+\n+#endif /* STAT_CACHE_H */\n-- \n1.5.5.76.gbb45.dirty\n"},{"id":"74801","messageId":"20080420111531.GB13411@bit.office.eurotux.com","threadId":"13180","inReplyTo":"20080420111346.GA13411@bit.office.eurotux.com","subject":"[PATCH 02/02/RFC] make use of the stat cache","fromName":"Luciano Rocha","fromEmail":"luciano@eurotux.com","sentAt":"2008-04-20T11:15:31Z","receivedAt":"2008-04-20T11:15:31Z","isPatch":true,"sender":{"key":"luciano@eurotux.com","avatar":null},"body":"Replace stat/lstat calls with cstat/clstat.\n\nSigned-off-by: Luciano Rocha <strange@nsk.no-ip.org>\n\n---\n Makefile                  |    2 ++\n builtin-apply.c           |   13 +++++++------\n builtin-blame.c           |    7 ++++---\n builtin-clean.c           |    3 ++-\n builtin-commit.c          |    7 ++++---\n builtin-count-objects.c   |    3 ++-\n builtin-diff.c            |    3 ++-\n builtin-fetch-pack.c      |    5 +++--\n builtin-grep.c            |    3 ++-\n builtin-init-db.c         |   11 ++++++-----\n builtin-ls-files.c        |    3 ++-\n builtin-mailsplit.c       |    3 ++-\n builtin-merge-recursive.c |    3 ++-\n builtin-mv.c              |    9 +++++----\n builtin-pack-objects.c    |    3 ++-\n builtin-prune.c           |    5 +++--\n builtin-rerere.c          |   15 ++++++++-------\n builtin-rm.c              |    3 ++-\n builtin-update-index.c    |    3 ++-\n check-racy.c              |    3 ++-\n combine-diff.c            |    3 ++-\n daemon.c                  |    3 ++-\n diff-lib.c                |    5 +++--\n diff.c                    |    9 +++++----\n dir.c                     |    7 ++++---\n entry.c                   |   11 ++++++-----\n help.c                    |    5 +++--\n http-push.c               |    3 ++-\n http-walker.c             |    3 ++-\n path.c                    |    9 +++++----\n read-cache.c              |    7 ++++---\n refs.c                    |   11 ++++++-----\n setup.c                   |    5 +++--\n sha1_file.c               |   15 ++++++++-------\n sha1_name.c               |    5 +++--\n symlinks.c                |    3 ++-\n test-chmtime.c            |    3 ++-\n transport.c               |    3 ++-\n unpack-trees.c            |    7 ++++---\n xdiff-interface.c         |    3 ++-\n 40 files changed, 134 insertions(+), 93 deletions(-)\n\ndiff --git a/Makefile b/Makefile\nindex 2cf38f0..7f01b71 100644\n--- a/Makefile\n+++ b/Makefile\n@@ -374,6 +374,7 @@ LIB_H += tree.h\n LIB_H += tree-walk.h\n LIB_H += unpack-trees.h\n LIB_H += utf8.h\n+LIB_H += stat-cache.h\n \n LIB_OBJS += alias.o\n LIB_OBJS += alloc.o\n@@ -465,6 +466,7 @@ LIB_OBJS += write_or_die.o\n LIB_OBJS += ws.o\n LIB_OBJS += wt-status.o\n LIB_OBJS += xdiff-interface.o\n+LIB_OBJS += stat-cache.o\n \n BUILTIN_OBJS += builtin-add.o\n BUILTIN_OBJS += builtin-annotate.o\ndiff --git a/builtin-apply.c b/builtin-apply.c\nindex caa3f2a..bedc7f4 100644\n--- a/builtin-apply.c\n+++ b/builtin-apply.c\n@@ -12,6 +12,7 @@\n #include \"blob.h\"\n #include \"delta.h\"\n #include \"builtin.h\"\n+#include \"stat-cache.h\"\n \n /*\n  *  --check turns on checking that the working tree matches the\n@@ -2237,7 +2238,7 @@ static int apply_data(struct patch *patch, struct stat *st, struct cache_entry *\n static int check_to_create_blob(const char *new_name, int ok_if_exists)\n {\n \tstruct stat nst;\n-\tif (!lstat(new_name, &nst)) {\n+\tif (!clstat(new_name, &nst)) {\n \t\tif (S_ISDIR(nst.st_mode) || ok_if_exists)\n \t\t\treturn 0;\n \t\t/*\n@@ -2289,7 +2290,7 @@ static int check_patch(struct patch *patch, struct patch *prev_patch)\n \t\tunsigned st_mode = 0;\n \n \t\tif (!cached)\n-\t\t\tstat_ret = lstat(old_name, &st);\n+\t\t\tstat_ret = clstat(old_name, &st);\n \t\tif (check_index) {\n \t\t\tint pos = cache_name_pos(old_name, strlen(old_name));\n \t\t\tif (pos < 0)\n@@ -2311,7 +2312,7 @@ static int check_patch(struct patch *patch, struct patch *prev_patch)\n \t\t\t\tif (checkout_entry(ce,\n \t\t\t\t\t\t   &costate,\n \t\t\t\t\t\t   NULL) ||\n-\t\t\t\t    lstat(old_name, &st))\n+\t\t\t\t    clstat(old_name, &st))\n \t\t\t\t\treturn -1;\n \t\t\t}\n \t\t\tif (!cached && verify_index_match(ce, &st))\n@@ -2632,7 +2633,7 @@ static void add_index_file(const char *path, unsigned mode, void *buf, unsigned\n \t\t\tdie(\"corrupt patch for subproject %s\", path);\n \t} else {\n \t\tif (!cached) {\n-\t\t\tif (lstat(path, &st) < 0)\n+\t\t\tif (clstat(path, &st) < 0)\n \t\t\t\tdie(\"unable to stat newly created file %s\",\n \t\t\t\t    path);\n \t\t\tfill_stat_cache_info(ce, &st);\n@@ -2651,7 +2652,7 @@ static int try_create_file(const char *path, unsigned int mode, const char *buf,\n \n \tif (S_ISGITLINK(mode)) {\n \t\tstruct stat st;\n-\t\tif (!lstat(path, &st) && S_ISDIR(st.st_mode))\n+\t\tif (!clstat(path, &st) && S_ISDIR(st.st_mode))\n \t\t\treturn 0;\n \t\treturn mkdir(path, 0777);\n \t}\n@@ -2703,7 +2704,7 @@ static void create_one_file(char *path, unsigned mode, const char *buf, unsigned\n \t\t * used to be.\n \t\t */\n \t\tstruct stat st;\n-\t\tif (!lstat(path, &st) && (!S_ISDIR(st.st_mode) || !rmdir(path)))\n+\t\tif (!clstat(path, &st) && (!S_ISDIR(st.st_mode) || !rmdir(path)))\n \t\t\terrno = EEXIST;\n \t}\n \ndiff --git a/builtin-blame.c b/builtin-blame.c\nindex bfd562d..ca52fa8 100644\n--- a/builtin-blame.c\n+++ b/builtin-blame.c\n@@ -18,6 +18,7 @@\n #include \"cache-tree.h\"\n #include \"path-list.h\"\n #include \"mailmap.h\"\n+#include \"stat-cache.h\"\n \n static char blame_usage[] =\n \"git-blame [-c] [-b] [-l] [--root] [-t] [-f] [-n] [-s] [-p] [-w] [-L n,m] [-S <revs-file>] [-M] [-C] [-C] [--contents <filename>] [--incremental] [commit] [--] file\\n\"\n@@ -1879,7 +1880,7 @@ static void sanity_check_refcnt(struct scoreboard *sb)\n static int has_path_in_work_tree(const char *path)\n {\n \tstruct stat st;\n-\treturn !lstat(path, &st);\n+\treturn !clstat(path, &st);\n }\n \n static unsigned parse_score(const char *arg)\n@@ -2038,12 +2039,12 @@ static struct commit *fake_working_tree_commit(const char *path, const char *con\n \t\tunsigned long fin_size;\n \n \t\tif (contents_from) {\n-\t\t\tif (stat(contents_from, &st) < 0)\n+\t\t\tif (cstat(contents_from, &st) < 0)\n \t\t\t\tdie(\"Cannot stat %s\", contents_from);\n \t\t\tread_from = contents_from;\n \t\t}\n \t\telse {\n-\t\t\tif (lstat(path, &st) < 0)\n+\t\t\tif (clstat(path, &st) < 0)\n \t\t\t\tdie(\"Cannot lstat %s\", path);\n \t\t\tread_from = path;\n \t\t}\ndiff --git a/builtin-clean.c b/builtin-clean.c\nindex 6778a03..97a8ec6 100644\n--- a/builtin-clean.c\n+++ b/builtin-clean.c\n@@ -11,6 +11,7 @@\n #include \"dir.h\"\n #include \"parse-options.h\"\n #include \"quote.h\"\n+#include \"stat-cache.h\"\n \n static int force = -1; /* unset */\n \n@@ -123,7 +124,7 @@ int cmd_clean(int argc, const char **argv, const char *prefix)\n \t\t * recursive directory removal, so lstat() here could\n \t\t * fail with ENOENT.\n \t\t */\n-\t\tif (lstat(ent->name, &st))\n+\t\tif (clstat(ent->name, &st))\n \t\t\tcontinue;\n \n \t\tif (pathspec) {\ndiff --git a/builtin-commit.c b/builtin-commit.c\nindex bcb7aaa..f5ce6f1 100644\n--- a/builtin-commit.c\n+++ b/builtin-commit.c\n@@ -23,6 +23,7 @@\n #include \"parse-options.h\"\n #include \"path-list.h\"\n #include \"unpack-trees.h\"\n+#include \"stat-cache.h\"\n \n static const char * const builtin_commit_usage[] = {\n \t\"git-commit [options] [--] <filepattern>...\",\n@@ -430,15 +431,15 @@ static int prepare_to_commit(const char *index_file, const char *prefix)\n \t\tstrbuf_add(&sb, buffer + 2, strlen(buffer + 2));\n \t\thook_arg1 = \"commit\";\n \t\thook_arg2 = use_message;\n-\t} else if (!stat(git_path(\"MERGE_MSG\"), &statbuf)) {\n+\t} else if (!cstat(git_path(\"MERGE_MSG\"), &statbuf)) {\n \t\tif (strbuf_read_file(&sb, git_path(\"MERGE_MSG\"), 0) < 0)\n \t\t\tdie(\"could not read MERGE_MSG: %s\", strerror(errno));\n \t\thook_arg1 = \"merge\";\n-\t} else if (!stat(git_path(\"SQUASH_MSG\"), &statbuf)) {\n+\t} else if (!cstat(git_path(\"SQUASH_MSG\"), &statbuf)) {\n \t\tif (strbuf_read_file(&sb, git_path(\"SQUASH_MSG\"), 0) < 0)\n \t\t\tdie(\"could not read SQUASH_MSG: %s\", strerror(errno));\n \t\thook_arg1 = \"squash\";\n-\t} else if (template_file && !stat(template_file, &statbuf)) {\n+\t} else if (template_file && !cstat(template_file, &statbuf)) {\n \t\tif (strbuf_read_file(&sb, template_file, 0) < 0)\n \t\t\tdie(\"could not read %s: %s\",\n \t\t\t    template_file, strerror(errno));\ndiff --git a/builtin-count-objects.c b/builtin-count-objects.c\nindex f00306f..fd4832e 100644\n--- a/builtin-count-objects.c\n+++ b/builtin-count-objects.c\n@@ -7,6 +7,7 @@\n #include \"cache.h\"\n #include \"builtin.h\"\n #include \"parse-options.h\"\n+#include \"stat-cache.h\"\n \n static void count_objects(DIR *d, char *path, int len, int verbose,\n \t\t\t  unsigned long *loose,\n@@ -40,7 +41,7 @@ static void count_objects(DIR *d, char *path, int len, int verbose,\n \t\t\tmemcpy(path + len + 3, ent->d_name, 38);\n \t\t\tpath[len + 2] = '/';\n \t\t\tpath[len + 41] = 0;\n-\t\t\tif (lstat(path, &st) || !S_ISREG(st.st_mode))\n+\t\t\tif (clstat(path, &st) || !S_ISREG(st.st_mode))\n \t\t\t\tbad = 1;\n \t\t\telse\n \t\t\t\t(*loose_size) += xsize_t(st.st_blocks);\ndiff --git a/builtin-diff.c b/builtin-diff.c\nindex 7c2a841..fab2f79 100644\n--- a/builtin-diff.c\n+++ b/builtin-diff.c\n@@ -13,6 +13,7 @@\n #include \"revision.h\"\n #include \"log-tree.h\"\n #include \"builtin.h\"\n+#include \"stat-cache.h\"\n \n struct blobinfo {\n \tunsigned char sha1[20];\n@@ -69,7 +70,7 @@ static int builtin_diff_b_f(struct rev_info *revs,\n \tif (argc > 1)\n \t\tusage(builtin_diff_usage);\n \n-\tif (lstat(path, &st))\n+\tif (clstat(path, &st))\n \t\tdie(\"'%s': %s\", path, strerror(errno));\n \tif (!(S_ISREG(st.st_mode) || S_ISLNK(st.st_mode)))\n \t\tdie(\"'%s': not a regular file or symlink\", path);\ndiff --git a/builtin-fetch-pack.c b/builtin-fetch-pack.c\nindex 65350ca..b3aba25 100644\n--- a/builtin-fetch-pack.c\n+++ b/builtin-fetch-pack.c\n@@ -9,6 +9,7 @@\n #include \"fetch-pack.h\"\n #include \"remote.h\"\n #include \"run-command.h\"\n+#include \"stat-cache.h\"\n \n static int transfer_unpack_limit = -1;\n static int fetch_unpack_limit = -1;\n@@ -780,7 +781,7 @@ struct ref *fetch_pack(struct fetch_pack_args *my_args,\n \tfetch_pack_setup();\n \tmemcpy(&args, my_args, sizeof(args));\n \tif (args.depth > 0) {\n-\t\tif (stat(git_path(\"shallow\"), &st))\n+\t\tif (cstat(git_path(\"shallow\"), &st))\n \t\t\tst.st_mtime = 0;\n \t}\n \n@@ -801,7 +802,7 @@ struct ref *fetch_pack(struct fetch_pack_args *my_args,\n #ifdef USE_NSEC\n \t\tmtime.usec = st.st_mtim.usec;\n #endif\n-\t\tif (stat(shallow, &st)) {\n+\t\tif (cstat(shallow, &st)) {\n \t\t\tif (mtime.sec)\n \t\t\t\tdie(\"shallow file was removed during fetch\");\n \t\t} else if (st.st_mtime != mtime.sec\ndiff --git a/builtin-grep.c b/builtin-grep.c\nindex ef29910..1fd1c58 100644\n--- a/builtin-grep.c\n+++ b/builtin-grep.c\n@@ -11,6 +11,7 @@\n #include \"tree-walk.h\"\n #include \"builtin.h\"\n #include \"grep.h\"\n+#include \"stat-cache.h\"\n \n #ifndef NO_EXTERNAL_GREP\n #ifdef __unix__\n@@ -132,7 +133,7 @@ static int grep_file(struct grep_opt *opt, const char *filename)\n \tchar *data;\n \tsize_t sz;\n \n-\tif (lstat(filename, &st) < 0) {\n+\tif (clstat(filename, &st) < 0) {\n \terr_ret:\n \t\tif (errno != ENOENT)\n \t\t\terror(\"'%s': %s\", filename, strerror(errno));\ndiff --git a/builtin-init-db.c b/builtin-init-db.c\nindex 2854868..3060c1a 100644\n--- a/builtin-init-db.c\n+++ b/builtin-init-db.c\n@@ -6,6 +6,7 @@\n #include \"cache.h\"\n #include \"builtin.h\"\n #include \"exec_cmd.h\"\n+#include \"stat-cache.h\"\n \n #ifndef DEFAULT_GIT_TEMPLATE_DIR\n #define DEFAULT_GIT_TEMPLATE_DIR \"/usr/share/git-core/templates\"\n@@ -56,14 +57,14 @@ static void copy_templates_1(char *path, int baselen,\n \t\t\tdie(\"insanely long template name %s\", de->d_name);\n \t\tmemcpy(path + baselen, de->d_name, namelen+1);\n \t\tmemcpy(template + template_baselen, de->d_name, namelen+1);\n-\t\tif (lstat(path, &st_git)) {\n+\t\tif (clstat(path, &st_git)) {\n \t\t\tif (errno != ENOENT)\n \t\t\t\tdie(\"cannot stat %s\", path);\n \t\t}\n \t\telse\n \t\t\texists = 1;\n \n-\t\tif (lstat(template, &st_template))\n+\t\tif (clstat(template, &st_template))\n \t\t\tdie(\"cannot stat template %s\", template);\n \n \t\tif (S_ISDIR(st_template.st_mode)) {\n@@ -235,10 +236,10 @@ static int create_default_files(const char *git_dir, const char *template_path)\n \n \t/* Check filemode trustability */\n \tfilemode = TEST_FILEMODE;\n-\tif (TEST_FILEMODE && !lstat(path, &st1)) {\n+\tif (TEST_FILEMODE && !clstat(path, &st1)) {\n \t\tstruct stat st2;\n \t\tfilemode = (!chmod(path, st1.st_mode ^ S_IXUSR) &&\n-\t\t\t\t!lstat(path, &st2) &&\n+\t\t\t\t!clstat(path, &st2) &&\n \t\t\t\tst1.st_mode != st2.st_mode);\n \t}\n \tgit_config_set(\"core.filemode\", filemode ? \"true\" : \"false\");\n@@ -262,7 +263,7 @@ static int create_default_files(const char *git_dir, const char *template_path)\n \t\tif (!close(xmkstemp(path)) &&\n \t\t    !unlink(path) &&\n \t\t    !symlink(\"testing\", path) &&\n-\t\t    !lstat(path, &st1) &&\n+\t\t    !clstat(path, &st1) &&\n \t\t    S_ISLNK(st1.st_mode))\n \t\t\tunlink(path); /* good */\n \t\telse\ndiff --git a/builtin-ls-files.c b/builtin-ls-files.c\nindex dc7eab8..be4a0fa 100644\n--- a/builtin-ls-files.c\n+++ b/builtin-ls-files.c\n@@ -10,6 +10,7 @@\n #include \"dir.h\"\n #include \"builtin.h\"\n #include \"tree.h\"\n+#include \"stat-cache.h\"\n \n static int abbrev;\n static int show_deleted;\n@@ -256,7 +257,7 @@ static void show_files(struct dir_struct *dir, const char *prefix)\n \t\t\tint dtype = ce_to_dtype(ce);\n \t\t\tif (excluded(dir, ce->name, &dtype) != dir->show_ignored)\n \t\t\t\tcontinue;\n-\t\t\terr = lstat(ce->name, &st);\n+\t\t\terr = clstat(ce->name, &st);\n \t\t\tif (show_deleted && err)\n \t\t\t\tshow_ce_entry(tag_removed, ce);\n \t\t\tif (show_modified && ce_modified(ce, &st, 0))\ndiff --git a/builtin-mailsplit.c b/builtin-mailsplit.c\nindex 46b27cd..838c52f 100644\n--- a/builtin-mailsplit.c\n+++ b/builtin-mailsplit.c\n@@ -7,6 +7,7 @@\n #include \"cache.h\"\n #include \"builtin.h\"\n #include \"path-list.h\"\n+#include \"stat-cache.h\"\n \n static const char git_mailsplit_usage[] =\n \"git-mailsplit [-d<prec>] [-f<n>] [-b] -o<directory> <mbox>|<Maildir>...\";\n@@ -278,7 +279,7 @@ int cmd_mailsplit(int argc, const char **argv, const char *prefix)\n \t\t\tcontinue;\n \t\t}\n \n-\t\tif (stat(arg, &argstat) == -1) {\n+\t\tif (cstat(arg, &argstat) == -1) {\n \t\t\terror(\"cannot stat %s (%s)\", arg, strerror(errno));\n \t\t\treturn 1;\n \t\t}\ndiff --git a/builtin-merge-recursive.c b/builtin-merge-recursive.c\nindex 910c0d2..5eb2407 100644\n--- a/builtin-merge-recursive.c\n+++ b/builtin-merge-recursive.c\n@@ -19,6 +19,7 @@\n #include \"interpolate.h\"\n #include \"attr.h\"\n #include \"merge-recursive.h\"\n+#include \"stat-cache.h\"\n \n static int subtree_merge;\n \n@@ -470,7 +471,7 @@ static char *unique_path(const char *path, const char *branch)\n \t\t\t*p = '_';\n \twhile (path_list_has_path(&current_file_set, newpath) ||\n \t       path_list_has_path(&current_directory_set, newpath) ||\n-\t       lstat(newpath, &st) == 0)\n+\t       clstat(newpath, &st) == 0)\n \t\tsprintf(p, \"_%d\", suffix++);\n \n \tpath_list_insert(newpath, &current_file_set);\ndiff --git a/builtin-mv.c b/builtin-mv.c\nindex 94f6dd2..807984a 100644\n--- a/builtin-mv.c\n+++ b/builtin-mv.c\n@@ -9,6 +9,7 @@\n #include \"cache-tree.h\"\n #include \"path-list.h\"\n #include \"parse-options.h\"\n+#include \"stat-cache.h\"\n \n static const char * const builtin_mv_usage[] = {\n \t\"git-mv [options] <source>... <destination>\",\n@@ -98,7 +99,7 @@ int cmd_mv(int argc, const char **argv, const char *prefix)\n \tif (dest_path[0][0] == '\\0')\n \t\t/* special case: \".\" was normalized to \"\" */\n \t\tdestination = copy_pathspec(dest_path[0], argv, argc, 1);\n-\telse if (!lstat(dest_path[0], &st) &&\n+\telse if (!clstat(dest_path[0], &st) &&\n \t\t\tS_ISDIR(st.st_mode)) {\n \t\tdest_path[0] = add_slash(dest_path[0]);\n \t\tdestination = copy_pathspec(dest_path[0], argv, argc, 1);\n@@ -118,13 +119,13 @@ int cmd_mv(int argc, const char **argv, const char *prefix)\n \t\t\tprintf(\"Checking rename of '%s' to '%s'\\n\", src, dst);\n \n \t\tlength = strlen(src);\n-\t\tif (lstat(src, &st) < 0)\n+\t\tif (clstat(src, &st) < 0)\n \t\t\tbad = \"bad source\";\n \t\telse if (!strncmp(src, dst, length) &&\n \t\t\t\t(dst[length] == 0 || dst[length] == '/')) {\n \t\t\tbad = \"can not move directory into itself\";\n \t\t} else if ((src_is_dir = S_ISDIR(st.st_mode))\n-\t\t\t\t&& lstat(dst, &st) == 0)\n+\t\t\t\t&& clstat(dst, &st) == 0)\n \t\t\tbad = \"cannot move directory over file\";\n \t\telse if (src_is_dir) {\n \t\t\tconst char *src_w_slash = add_slash(src);\n@@ -177,7 +178,7 @@ int cmd_mv(int argc, const char **argv, const char *prefix)\n \t\t\t\t}\n \t\t\t\targc += last - first;\n \t\t\t}\n-\t\t} else if (lstat(dst, &st) == 0) {\n+\t\t} else if (clstat(dst, &st) == 0) {\n \t\t\tbad = \"destination exists\";\n \t\t\tif (force) {\n \t\t\t\t/*\ndiff --git a/builtin-pack-objects.c b/builtin-pack-objects.c\nindex 777f272..50da2fa 100644\n--- a/builtin-pack-objects.c\n+++ b/builtin-pack-objects.c\n@@ -16,6 +16,7 @@\n #include \"list-objects.h\"\n #include \"progress.h\"\n #include \"refs.h\"\n+#include \"stat-cache.h\"\n \n #ifdef THREADED_DELTA_SEARCH\n #include \"thread-utils.h\"\n@@ -530,7 +531,7 @@ static void write_pack_file(void)\n \t\t\t * packs then we should modify the mtime of later ones\n \t\t\t * to preserve this property.\n \t\t\t */\n-\t\t\tif (stat(tmpname, &st) < 0) {\n+\t\t\tif (cstat(tmpname, &st) < 0) {\n \t\t\t\twarning(\"failed to stat %s: %s\",\n \t\t\t\t\ttmpname, strerror(errno));\n \t\t\t} else if (!last_mtime) {\ndiff --git a/builtin-prune.c b/builtin-prune.c\nindex 25f9304..ca4f636 100644\n--- a/builtin-prune.c\n+++ b/builtin-prune.c\n@@ -5,6 +5,7 @@\n #include \"builtin.h\"\n #include \"reachable.h\"\n #include \"parse-options.h\"\n+#include \"stat-cache.h\"\n \n static const char * const prune_usage[] = {\n \t\"git-prune [-n] [--expire <time>] [--] [<head>...]\",\n@@ -18,7 +19,7 @@ static int prune_object(char *path, const char *filename, const unsigned char *s\n \tconst char *fullpath = mkpath(\"%s/%s\", path, filename);\n \tif (expire) {\n \t\tstruct stat st;\n-\t\tif (lstat(fullpath, &st))\n+\t\tif (clstat(fullpath, &st))\n \t\t\treturn error(\"Could not stat '%s'\", fullpath);\n \t\tif (st.st_mtime > expire)\n \t\t\treturn 0;\n@@ -114,7 +115,7 @@ static void remove_temporary_files(void)\n \t\t\t\tcontinue;\n \t\t\tif (expire) {\n \t\t\t\tstruct stat st;\n-\t\t\t\tif (stat(name, &st) != 0 || st.st_mtime >= expire)\n+\t\t\t\tif (cstat(name, &st) != 0 || st.st_mtime >= expire)\n \t\t\t\t\tcontinue;\n \t\t\t}\n \t\t\tprintf(\"Removing stale temporary file %s\\n\", name);\ndiff --git a/builtin-rerere.c b/builtin-rerere.c\nindex c607aad..ff18ea9 100644\n--- a/builtin-rerere.c\n+++ b/builtin-rerere.c\n@@ -3,6 +3,7 @@\n #include \"path-list.h\"\n #include \"xdiff/xdiff.h\"\n #include \"xdiff-interface.h\"\n+#include \"stat-cache.h\"\n \n #include <time.h>\n \n@@ -219,11 +220,11 @@ static void garbage_collect(struct path_list *rr)\n \t\t\tcontinue;\n \t\ti = snprintf(buf + len, sizeof(buf) - len, \"%s\", name);\n \t\tstrlcpy(buf + len + i, \"/preimage\", sizeof(buf) - len - i);\n-\t\tif (stat(buf, &st))\n+\t\tif (cstat(buf, &st))\n \t\t\tcontinue;\n \t\tthen = st.st_mtime;\n \t\tstrlcpy(buf + len + i, \"/postimage\", sizeof(buf) - len - i);\n-\t\tcutoff = stat(buf, &st) ? cutoff_noresolve : cutoff_resolve;\n+\t\tcutoff = cstat(buf, &st) ? cutoff_noresolve : cutoff_resolve;\n \t\tif (then < now - cutoff * 86400) {\n \t\t\tbuf[len + i] = '\\0';\n \t\t\tpath_list_insert(xstrdup(name), &to_remove);\n@@ -311,8 +312,8 @@ static int do_plain_rerere(struct path_list *rr, int fd)\n \t\tconst char *path = rr->items[i].path;\n \t\tconst char *name = (const char *)rr->items[i].util;\n \n-\t\tif (!stat(rr_path(name, \"preimage\"), &st) &&\n-\t\t\t\t!stat(rr_path(name, \"postimage\"), &st)) {\n+\t\tif (!cstat(rr_path(name, \"preimage\"), &st) &&\n+\t\t\t\t!cstat(rr_path(name, \"postimage\"), &st)) {\n \t\t\tif (!merge(name, path)) {\n \t\t\t\tfprintf(stderr, \"Resolved '%s' using \"\n \t\t\t\t\t\t\"previous resolution.\\n\", path);\n@@ -362,7 +363,7 @@ static int is_rerere_enabled(void)\n \t\treturn 0;\n \n \trr_cache = git_path(\"rr-cache\");\n-\trr_cache_exists = !stat(rr_cache, &st) && S_ISDIR(st.st_mode);\n+\trr_cache_exists = !cstat(rr_cache, &st) && S_ISDIR(st.st_mode);\n \tif (rerere_enabled < 0)\n \t\treturn rr_cache_exists;\n \n@@ -412,9 +413,9 @@ int cmd_rerere(int argc, const char **argv, const char *prefix)\n \t\tfor (i = 0; i < merge_rr.nr; i++) {\n \t\t\tstruct stat st;\n \t\t\tconst char *name = (const char *)merge_rr.items[i].util;\n-\t\t\tif (!stat(git_path(\"rr-cache/%s\", name), &st) &&\n+\t\t\tif (!cstat(git_path(\"rr-cache/%s\", name), &st) &&\n \t\t\t\t\tS_ISDIR(st.st_mode) &&\n-\t\t\t\t\tstat(rr_path(name, \"postimage\"), &st))\n+\t\t\t\t\tcstat(rr_path(name, \"postimage\"), &st))\n \t\t\t\tunlink_rr_item(name);\n \t\t}\n \t\tunlink(merge_rr_path);\ndiff --git a/builtin-rm.c b/builtin-rm.c\nindex c0a8bb6..4039a44 100644\n--- a/builtin-rm.c\n+++ b/builtin-rm.c\n@@ -9,6 +9,7 @@\n #include \"cache-tree.h\"\n #include \"tree-walk.h\"\n #include \"parse-options.h\"\n+#include \"stat-cache.h\"\n \n static const char * const builtin_rm_usage[] = {\n \t\"git-rm [options] [--] <file>...\",\n@@ -76,7 +77,7 @@ static int check_local_mod(unsigned char *head, int index_only)\n \t\t\tcontinue; /* removing unmerged entry */\n \t\tce = active_cache[pos];\n \n-\t\tif (lstat(ce->name, &st) < 0) {\n+\t\tif (clstat(ce->name, &st) < 0) {\n \t\t\tif (errno != ENOENT)\n \t\t\t\tfprintf(stderr, \"warning: '%s': %s\",\n \t\t\t\t\tce->name, strerror(errno));\ndiff --git a/builtin-update-index.c b/builtin-update-index.c\nindex a8795d3..e88cd95 100644\n--- a/builtin-update-index.c\n+++ b/builtin-update-index.c\n@@ -9,6 +9,7 @@\n #include \"tree-walk.h\"\n #include \"builtin.h\"\n #include \"refs.h\"\n+#include \"stat-cache.h\"\n \n /*\n  * Default to not allowing changes to the list of files. The\n@@ -198,7 +199,7 @@ static int process_path(const char *path)\n \t * First things first: get the stat information, to decide\n \t * what to do about the pathname!\n \t */\n-\tif (lstat(path, &st) < 0)\n+\tif (clstat(path, &st) < 0)\n \t\treturn process_lstat_error(path, errno);\n \n \tlen = strlen(path);\ndiff --git a/check-racy.c b/check-racy.c\nindex 00d92a1..a33029f 100644\n--- a/check-racy.c\n+++ b/check-racy.c\n@@ -1,4 +1,5 @@\n #include \"cache.h\"\n+#include \"stat-cache.h\"\n \n int main(int ac, char **av)\n {\n@@ -11,7 +12,7 @@ int main(int ac, char **av)\n \t\tstruct cache_entry *ce = active_cache[i];\n \t\tstruct stat st;\n \n-\t\tif (lstat(ce->name, &st)) {\n+\t\tif (clstat(ce->name, &st)) {\n \t\t\terror(\"lstat(%s): %s\", ce->name, strerror(errno));\n \t\t\tcontinue;\n \t\t}\ndiff --git a/combine-diff.c b/combine-diff.c\nindex 0e19cba..1e0af00 100644\n--- a/combine-diff.c\n+++ b/combine-diff.c\n@@ -6,6 +6,7 @@\n #include \"quote.h\"\n #include \"xdiff-interface.h\"\n #include \"log-tree.h\"\n+#include \"stat-cache.h\"\n \n static struct combine_diff_path *intersect_paths(struct combine_diff_path *curr, int n, int num_parent)\n {\n@@ -683,7 +684,7 @@ static void show_patch_diff(struct combine_diff_path *elem, int num_parent,\n \t\tstruct stat st;\n \t\tint fd = -1;\n \n-\t\tif (lstat(elem->path, &st) < 0)\n+\t\tif (clstat(elem->path, &st) < 0)\n \t\t\tgoto deleted_file;\n \n \t\tif (S_ISLNK(st.st_mode)) {\ndiff --git a/daemon.c b/daemon.c\nindex 2b4a6f1..ac4943b 100644\n--- a/daemon.c\n+++ b/daemon.c\n@@ -2,6 +2,7 @@\n #include \"pkt-line.h\"\n #include \"exec_cmd.h\"\n #include \"interpolate.h\"\n+#include \"stat-cache.h\"\n \n #include <syslog.h>\n \n@@ -1187,7 +1188,7 @@ int main(int argc, char **argv)\n \tif (base_path) {\n \t\tstruct stat st;\n \n-\t\tif (stat(base_path, &st) || !S_ISDIR(st.st_mode))\n+\t\tif (cstat(base_path, &st) || !S_ISDIR(st.st_mode))\n \t\t\tdie(\"base-path '%s' does not exist or \"\n \t\t\t    \"is not a directory\", base_path);\n \t}\ndiff --git a/diff-lib.c b/diff-lib.c\nindex 069e450..be594ba 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -11,6 +11,7 @@\n #include \"path-list.h\"\n #include \"unpack-trees.h\"\n #include \"refs.h\"\n+#include \"stat-cache.h\"\n \n /*\n  * diff-files\n@@ -40,7 +41,7 @@ static int get_mode(const char *path, int *mode)\n \t\t*mode = 0;\n \telse if (!strcmp(path, \"-\"))\n \t\t*mode = create_ce_mode(0666);\n-\telse if (stat(path, &st))\n+\telse if (cstat(path, &st))\n \t\treturn error(\"Could not access '%s'\", path);\n \telse\n \t\t*mode = st.st_mode;\n@@ -340,7 +341,7 @@ int run_diff_files_cmd(struct rev_info *revs, int argc, const char **argv)\n  */\n static int check_work_tree_entity(const struct cache_entry *ce, struct stat *st, char *symcache)\n {\n-\tif (lstat(ce->name, st) < 0) {\n+\tif (clstat(ce->name, st) < 0) {\n \t\tif (errno != ENOENT && errno != ENOTDIR)\n \t\t\treturn -1;\n \t\treturn 1;\ndiff --git a/diff.c b/diff.c\nindex 8022e67..4508000 100644\n--- a/diff.c\n+++ b/diff.c\n@@ -11,6 +11,7 @@\n #include \"attr.h\"\n #include \"run-command.h\"\n #include \"utf8.h\"\n+#include \"stat-cache.h\"\n \n #ifdef NO_FAST_WORKING_DIRECTORY\n #define FAST_WORKING_DIRECTORY 0\n@@ -1624,7 +1625,7 @@ static int reuse_worktree_file(const char *name, const unsigned char *sha1, int\n \t * If ce matches the file in the work tree, we can reuse it.\n \t */\n \tif (ce_uptodate(ce) ||\n-\t    (!lstat(name, &st) && !ce_match_stat(ce, &st, 0)))\n+\t    (!clstat(name, &st) && !ce_match_stat(ce, &st, 0)))\n \t\treturn 1;\n \n \treturn 0;\n@@ -1694,7 +1695,7 @@ int diff_populate_filespec(struct diff_filespec *s, int size_only)\n \t\tif (!strcmp(s->path, \"-\"))\n \t\t\treturn populate_from_stdin(s);\n \n-\t\tif (lstat(s->path, &st) < 0) {\n+\t\tif (clstat(s->path, &st) < 0) {\n \t\t\tif (errno == ENOENT) {\n \t\t\terr_empty:\n \t\t\t\terr = -1;\n@@ -1810,7 +1811,7 @@ static void prepare_temp_file(const char *name,\n \tif (!one->sha1_valid ||\n \t    reuse_worktree_file(name, one->sha1, 1)) {\n \t\tstruct stat st;\n-\t\tif (lstat(name, &st) < 0) {\n+\t\tif (clstat(name, &st) < 0) {\n \t\t\tif (errno == ENOENT)\n \t\t\t\tgoto not_a_valid_file;\n \t\t\tdie(\"stat(%s): %s\", name, strerror(errno));\n@@ -1996,7 +1997,7 @@ static void diff_fill_sha1_info(struct diff_filespec *one)\n \t\t\t\thashcpy(one->sha1, null_sha1);\n \t\t\t\treturn;\n \t\t\t}\n-\t\t\tif (lstat(one->path, &st) < 0)\n+\t\t\tif (clstat(one->path, &st) < 0)\n \t\t\t\tdie(\"stat %s\", one->path);\n \t\t\tif (index_path(one->sha1, one->path, &st, 0))\n \t\t\t\tdie(\"cannot hash %s\\n\", one->path);\ndiff --git a/dir.c b/dir.c\nindex d79762c..6bdab38 100644\n--- a/dir.c\n+++ b/dir.c\n@@ -8,6 +8,7 @@\n #include \"cache.h\"\n #include \"dir.h\"\n #include \"refs.h\"\n+#include \"stat-cache.h\"\n \n struct path_simplify {\n \tint len;\n@@ -538,7 +539,7 @@ static int get_dtype(struct dirent *de, const char *path)\n \n \tif (dtype != DT_UNKNOWN)\n \t\treturn dtype;\n-\tif (lstat(path, &st))\n+\tif (clstat(path, &st))\n \t\treturn dtype;\n \tif (S_ISREG(st.st_mode))\n \t\treturn DT_REG;\n@@ -721,7 +722,7 @@ int read_directory(struct dir_struct *dir, const char *path, const char *base, i\n int file_exists(const char *f)\n {\n \tstruct stat sb;\n-\treturn lstat(f, &sb) == 0;\n+\treturn clstat(f, &sb) == 0;\n }\n \n /*\n@@ -788,7 +789,7 @@ int remove_dir_recursively(struct strbuf *path, int only_empty)\n \n \t\tstrbuf_setlen(path, len);\n \t\tstrbuf_addstr(path, e->d_name);\n-\t\tif (lstat(path->buf, &st))\n+\t\tif (clstat(path->buf, &st))\n \t\t\t; /* fall thru */\n \t\telse if (S_ISDIR(st.st_mode)) {\n \t\t\tif (!remove_dir_recursively(path, only_empty))\ndiff --git a/entry.c b/entry.c\nindex 222aaa3..6d31ac3 100644\n--- a/entry.c\n+++ b/entry.c\n@@ -1,5 +1,6 @@\n #include \"cache.h\"\n #include \"blob.h\"\n+#include \"stat-cache.h\"\n \n static void create_directories(const char *path, const struct checkout *state)\n {\n@@ -21,13 +22,13 @@ static void create_directories(const char *path, const struct checkout *state)\n \t\t\t * allowed to be a symlink to an existing\n \t\t\t * directory.\n \t\t\t */\n-\t\t\tstat_status = stat(buf, &st);\n+\t\t\tstat_status = cstat(buf, &st);\n \t\telse\n \t\t\t/*\n \t\t\t * if there currently is a symlink, we would\n \t\t\t * want to replace it with a real directory.\n \t\t\t */\n-\t\t\tstat_status = lstat(buf, &st);\n+\t\t\tstat_status = clstat(buf, &st);\n \n \t\tif (!stat_status && S_ISDIR(st.st_mode))\n \t\t\tcontinue; /* ok, it is already a directory. */\n@@ -67,7 +68,7 @@ static void remove_subtree(const char *path)\n \t\t     ((de->d_name[1] == '.') && de->d_name[2] == 0)))\n \t\t\tcontinue;\n \t\tstrcpy(name, de->d_name);\n-\t\tif (lstat(pathbuf, &st))\n+\t\tif (clstat(pathbuf, &st))\n \t\t\tdie(\"cannot lstat %s (%s)\", pathbuf, strerror(errno));\n \t\tif (S_ISDIR(st.st_mode))\n \t\t\tremove_subtree(pathbuf);\n@@ -184,7 +185,7 @@ static int write_entry(struct cache_entry *ce, char *path, const struct checkout\n \n \tif (state->refresh_cache) {\n \t\tstruct stat st;\n-\t\tlstat(ce->name, &st);\n+\t\tclstat(ce->name, &st);\n \t\tfill_stat_cache_info(ce, &st);\n \t}\n \treturn 0;\n@@ -202,7 +203,7 @@ int checkout_entry(struct cache_entry *ce, const struct checkout *state, char *t\n \tmemcpy(path, state->base_dir, len);\n \tstrcpy(path + len, ce->name);\n \n-\tif (!lstat(path, &st)) {\n+\tif (!clstat(path, &st)) {\n \t\tunsigned changed = ce_match_stat(ce, &st, CE_MATCH_IGNORE_VALID);\n \t\tif (!changed)\n \t\t\treturn 0;\ndiff --git a/help.c b/help.c\nindex 10298fb..6616f25 100644\n--- a/help.c\n+++ b/help.c\n@@ -9,6 +9,7 @@\n #include \"common-cmds.h\"\n #include \"parse-options.h\"\n #include \"run-command.h\"\n+#include \"stat-cache.h\"\n \n static struct man_viewer_list {\n \tvoid (*exec)(const char *);\n@@ -299,7 +300,7 @@ static unsigned int list_commands_in_dir(struct cmdnames *cmds,\n \t\tif (prefixcmp(de->d_name, prefix))\n \t\t\tcontinue;\n \n-\t\tif (stat(de->d_name, &st) || /* stat, not lstat */\n+\t\tif (cstat(de->d_name, &st) || /* stat, not lstat */\n \t\t    !S_ISREG(st.st_mode) ||\n \t\t    !(st.st_mode & S_IXUSR))\n \t\t\tcontinue;\n@@ -479,7 +480,7 @@ static void get_html_page_path(struct strbuf *page_path, const char *page)\n \tstruct stat st;\n \n \t/* Check that we have a git documentation directory. */\n-\tif (stat(GIT_HTML_PATH \"/git.html\", &st) || !S_ISREG(st.st_mode))\n+\tif (cstat(GIT_HTML_PATH \"/git.html\", &st) || !S_ISREG(st.st_mode))\n \t\tdie(\"'%s': not a documentation directory.\", GIT_HTML_PATH);\n \n \tstrbuf_init(page_path, 0);\ndiff --git a/http-push.c b/http-push.c\nindex 5b23038..f536a8b 100644\n--- a/http-push.c\n+++ b/http-push.c\n@@ -9,6 +9,7 @@\n #include \"revision.h\"\n #include \"exec_cmd.h\"\n #include \"remote.h\"\n+#include \"stat-cache.h\"\n \n #include <expat.h>\n \n@@ -732,7 +733,7 @@ static void finish_request(struct transfer_request *request)\n \n \t\tif (request->curl_result != CURLE_OK &&\n \t\t    request->http_code != 416) {\n-\t\t\tif (stat(request->tmpfile, &st) == 0) {\n+\t\t\tif (cstat(request->tmpfile, &st) == 0) {\n \t\t\t\tif (st.st_size == 0)\n \t\t\t\t\tunlink(request->tmpfile);\n \t\t\t}\ndiff --git a/http-walker.c b/http-walker.c\nindex 7bda34d..cd16036 100644\n--- a/http-walker.c\n+++ b/http-walker.c\n@@ -3,6 +3,7 @@\n #include \"pack.h\"\n #include \"walker.h\"\n #include \"http.h\"\n+#include \"stat-cache.h\"\n \n #define PREV_BUF_SIZE 4096\n #define RANGE_HEADER_SIZE 30\n@@ -237,7 +238,7 @@ static void finish_object_request(struct object_request *obj_req)\n \tif (obj_req->http_code == 416) {\n \t\tfprintf(stderr, \"Warning: requested range invalid; we may already have all the data.\\n\");\n \t} else if (obj_req->curl_result != CURLE_OK) {\n-\t\tif (stat(obj_req->tmpfile, &st) == 0)\n+\t\tif (cstat(obj_req->tmpfile, &st) == 0)\n \t\t\tif (st.st_size == 0)\n \t\t\t\tunlink(obj_req->tmpfile);\n \t\treturn;\ndiff --git a/path.c b/path.c\nindex f4ed979..401ab0f 100644\n--- a/path.c\n+++ b/path.c\n@@ -11,6 +11,7 @@\n  * which is what it's designed for.\n  */\n #include \"cache.h\"\n+#include \"stat-cache.h\"\n \n static char bad_path[] = \"/bad-path/\";\n \n@@ -93,7 +94,7 @@ int validate_headref(const char *path)\n \tunsigned char sha1[20];\n \tint len, fd;\n \n-\tif (lstat(path, &st) < 0)\n+\tif (clstat(path, &st) < 0)\n \t\treturn -1;\n \n \t/* Make sure it is a \"refs/..\" symlink */\n@@ -263,7 +264,7 @@ int adjust_shared_perm(const char *path)\n \n \tif (!shared_repository)\n \t\treturn 0;\n-\tif (lstat(path, &st) < 0)\n+\tif (clstat(path, &st) < 0)\n \t\treturn -1;\n \tmode = st.st_mode;\n \tif (mode & S_IRUSR)\n@@ -306,7 +307,7 @@ const char *make_absolute_path(const char *path)\n \t\tdie (\"Too long path: %.*s\", 60, path);\n \n \twhile (depth--) {\n-\t\tif (stat(buf, &st) || !S_ISDIR(st.st_mode)) {\n+\t\tif (cstat(buf, &st) || !S_ISDIR(st.st_mode)) {\n \t\t\tchar *last_slash = strrchr(buf, '/');\n \t\t\tif (last_slash) {\n \t\t\t\t*last_slash = '\\0';\n@@ -338,7 +339,7 @@ const char *make_absolute_path(const char *path)\n \t\t\tlast_elem = NULL;\n \t\t}\n \n-\t\tif (!lstat(buf, &st) && S_ISLNK(st.st_mode)) {\n+\t\tif (!clstat(buf, &st) && S_ISLNK(st.st_mode)) {\n \t\t\tlen = readlink(buf, next_buf, PATH_MAX);\n \t\t\tif (len < 0)\n \t\t\t\tdie (\"Invalid symlink: %s\", buf);\ndiff --git a/read-cache.c b/read-cache.c\nindex a92b25b..a80af56 100644\n--- a/read-cache.c\n+++ b/read-cache.c\n@@ -8,6 +8,7 @@\n #include \"cache-tree.h\"\n #include \"refs.h\"\n #include \"dir.h\"\n+#include \"stat-cache.h\"\n \n /* Index extensions.\n  *\n@@ -495,7 +496,7 @@ int add_file_to_index(struct index_state *istate, const char *path, int verbose)\n \tstruct cache_entry *ce;\n \tunsigned ce_option = CE_MATCH_IGNORE_VALID|CE_MATCH_RACY_IS_DIRTY;\n \n-\tif (lstat(path, &st))\n+\tif (clstat(path, &st))\n \t\tdie(\"%s: unable to stat (%s)\", path, strerror(errno));\n \n \tif (!S_ISREG(st.st_mode) && !S_ISLNK(st.st_mode) && !S_ISDIR(st.st_mode))\n@@ -902,7 +903,7 @@ static struct cache_entry *refresh_cache_ent(struct index_state *istate,\n \tif (ce_uptodate(ce))\n \t\treturn ce;\n \n-\tif (lstat(ce->name, &st) < 0) {\n+\tif (clstat(ce->name, &st) < 0) {\n \t\tif (err)\n \t\t\t*err = errno;\n \t\treturn NULL;\n@@ -1290,7 +1291,7 @@ static void ce_smudge_racily_clean_entry(struct cache_entry *ce)\n \t */\n \tstruct stat st;\n \n-\tif (lstat(ce->name, &st) < 0)\n+\tif (clstat(ce->name, &st) < 0)\n \t\treturn;\n \tif (ce_match_stat_basic(ce, &st))\n \t\treturn;\ndiff --git a/refs.c b/refs.c\nindex 1b0050e..35d2f2b 100644\n--- a/refs.c\n+++ b/refs.c\n@@ -3,6 +3,7 @@\n #include \"object.h\"\n #include \"tag.h\"\n #include \"dir.h\"\n+#include \"stat-cache.h\"\n \n /* ISSYMREF=01 and ISPACKED=02 are public interfaces */\n #define REF_KNOWS_PEELED 04\n@@ -256,7 +257,7 @@ static struct ref_list *get_ref_dir(const char *base, struct ref_list *list)\n \t\t\tif (has_extension(de->d_name, \".lock\"))\n \t\t\t\tcontinue;\n \t\t\tmemcpy(ref + baselen, de->d_name, namelen+1);\n-\t\t\tif (stat(git_path(\"%s\", ref), &st) < 0)\n+\t\t\tif (cstat(git_path(\"%s\", ref), &st) < 0)\n \t\t\t\tcontinue;\n \t\t\tif (S_ISDIR(st.st_mode)) {\n \t\t\t\tlist = get_ref_dir(ref, list);\n@@ -391,7 +392,7 @@ const char *resolve_ref(const char *ref, unsigned char *sha1, int reading, int *\n \t\t * born.  It is NOT OK if we are resolving for\n \t\t * reading.\n \t\t */\n-\t\tif (lstat(path, &st) < 0) {\n+\t\tif (clstat(path, &st) < 0) {\n \t\t\tstruct ref_list *list = get_packed_refs();\n \t\t\twhile (list) {\n \t\t\t\tif (!strcmp(ref, list->name)) {\n@@ -805,7 +806,7 @@ static struct ref_lock *lock_ref_sha1_basic(const char *ref, const unsigned char\n \tlock->ref_name = xstrdup(ref);\n \tlock->orig_ref_name = xstrdup(orig_ref);\n \tref_file = git_path(\"%s\", ref);\n-\tif (lstat(ref_file, &st) && errno == ENOENT)\n+\tif (clstat(ref_file, &st) && errno == ENOENT)\n \t\tlock->force_write = 1;\n \tif ((flags & REF_NODEREF) && (type & REF_ISSYMREF))\n \t\tlock->force_write = 1;\n@@ -924,7 +925,7 @@ int rename_ref(const char *oldref, const char *newref, const char *logmsg)\n \tint flag = 0, logmoved = 0;\n \tstruct ref_lock *lock;\n \tstruct stat loginfo;\n-\tint log = !lstat(git_path(\"logs/%s\", oldref), &loginfo);\n+\tint log = !clstat(git_path(\"logs/%s\", oldref), &loginfo);\n \n \tif (S_ISLNK(loginfo.st_mode))\n \t\treturn error(\"reflog for %s is a symlink\", oldref);\n@@ -1465,7 +1466,7 @@ static int do_for_each_reflog(const char *base, each_ref_fn fn, void *cb_data)\n \t\t\tif (has_extension(de->d_name, \".lock\"))\n \t\t\t\tcontinue;\n \t\t\tmemcpy(log + baselen, de->d_name, namelen+1);\n-\t\t\tif (stat(git_path(\"logs/%s\", log), &st) < 0)\n+\t\t\tif (cstat(git_path(\"logs/%s\", log), &st) < 0)\n \t\t\t\tcontinue;\n \t\t\tif (S_ISDIR(st.st_mode)) {\n \t\t\t\tretval = do_for_each_reflog(log, fn, cb_data);\ndiff --git a/setup.c b/setup.c\nindex 3d2d958..e72b48e 100644\n--- a/setup.c\n+++ b/setup.c\n@@ -1,5 +1,6 @@\n #include \"cache.h\"\n #include \"dir.h\"\n+#include \"stat-cache.h\"\n \n static int inside_git_dir = -1;\n static int inside_work_tree = -1;\n@@ -148,7 +149,7 @@ void verify_filename(const char *prefix, const char *arg)\n \tif (*arg == '-')\n \t\tdie(\"bad flag '%s' used after filename\", arg);\n \tname = prefix ? prefix_filename(prefix, strlen(prefix), arg) : arg;\n-\tif (!lstat(name, &st))\n+\tif (!clstat(name, &st))\n \t\treturn;\n \tif (errno == ENOENT)\n \t\tdie(\"ambiguous argument '%s': unknown revision or path not in the working tree.\\n\"\n@@ -171,7 +172,7 @@ void verify_non_filename(const char *prefix, const char *arg)\n \tif (*arg == '-')\n \t\treturn; /* flag */\n \tname = prefix ? prefix_filename(prefix, strlen(prefix), arg) : arg;\n-\tif (!lstat(name, &st))\n+\tif (!clstat(name, &st))\n \t\tdie(\"ambiguous argument '%s': both revision and filename\\n\"\n \t\t    \"Use '--' to separate filenames from revisions\", arg);\n \tif (errno != ENOENT && errno != ENOTDIR)\ndiff --git a/sha1_file.c b/sha1_file.c\nindex 445a871..777f47d 100644\n--- a/sha1_file.c\n+++ b/sha1_file.c\n@@ -15,6 +15,7 @@\n #include \"tree.h\"\n #include \"refs.h\"\n #include \"pack-revindex.h\"\n+#include \"stat-cache.h\"\n \n #ifndef O_NOATIME\n #if defined(__linux__) && (defined(__i386__) || defined(__PPC__))\n@@ -97,7 +98,7 @@ int safe_create_leading_directories(char *path)\n \t\tif (!pos)\n \t\t\tbreak;\n \t\t*pos = 0;\n-\t\tif (!stat(path, &st)) {\n+\t\tif (!cstat(path, &st)) {\n \t\t\t/* path exists */\n \t\t\tif (!S_ISDIR(st.st_mode)) {\n \t\t\t\t*pos = '/';\n@@ -278,7 +279,7 @@ static int link_alt_odb_entry(const char * entry, int len, const char * relative\n \tent->base[pfxlen] = ent->base[entlen-1] = 0;\n \n \t/* Detect cases where alternate disappeared */\n-\tif (stat(ent->base, &st) || !S_ISDIR(st.st_mode)) {\n+\tif (cstat(ent->base, &st) || !S_ISDIR(st.st_mode)) {\n \t\terror(\"object directory %s does not exist; \"\n \t\t      \"check .git/objects/info/alternates.\",\n \t\t      ent->base);\n@@ -400,13 +401,13 @@ static char *find_sha1_file(const unsigned char *sha1, struct stat *st)\n \tchar *name = sha1_file_name(sha1);\n \tstruct alternate_object_database *alt;\n \n-\tif (!stat(name, st))\n+\tif (!cstat(name, st))\n \t\treturn name;\n \tprepare_alt_odb();\n \tfor (alt = alt_odb_list; alt; alt = alt->next) {\n \t\tname = alt->name;\n \t\tfill_sha1_path(name, sha1);\n-\t\tif (!stat(alt->base, st))\n+\t\tif (!cstat(alt->base, st))\n \t\t\treturn alt->base;\n \t}\n \treturn NULL;\n@@ -804,7 +805,7 @@ struct packed_git *add_packed_git(const char *path, int path_len, int local)\n \t\treturn NULL;\n \tmemcpy(p->pack_name, path, path_len);\n \tstrcpy(p->pack_name + path_len, \".pack\");\n-\tif (stat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n+\tif (cstat(p->pack_name, &st) || !S_ISREG(st.st_mode)) {\n \t\tfree(p);\n \t\treturn NULL;\n \t}\n@@ -2303,7 +2304,7 @@ int write_sha1_from_fd(const unsigned char *sha1, int fd, char *buffer,\n int has_pack_index(const unsigned char *sha1)\n {\n \tstruct stat st;\n-\tif (stat(sha1_pack_index_name(sha1), &st))\n+\tif (cstat(sha1_pack_index_name(sha1), &st))\n \t\treturn 0;\n \treturn 1;\n }\n@@ -2311,7 +2312,7 @@ int has_pack_index(const unsigned char *sha1)\n int has_pack_file(const unsigned char *sha1)\n {\n \tstruct stat st;\n-\tif (stat(sha1_pack_name(sha1), &st))\n+\tif (cstat(sha1_pack_name(sha1), &st))\n \t\treturn 0;\n \treturn 1;\n }\ndiff --git a/sha1_name.c b/sha1_name.c\nindex 491d2e7..df8b36c 100644\n--- a/sha1_name.c\n+++ b/sha1_name.c\n@@ -5,6 +5,7 @@\n #include \"blob.h\"\n #include \"tree-walk.h\"\n #include \"refs.h\"\n+#include \"stat-cache.h\"\n \n static int find_short_object_filename(int len, const char *name, unsigned char *sha1)\n {\n@@ -276,11 +277,11 @@ int dwim_log(const char *str, int len, unsigned char *sha1, char **log)\n \t\tref = resolve_ref(path, hash, 0, NULL);\n \t\tif (!ref)\n \t\t\tcontinue;\n-\t\tif (!stat(git_path(\"logs/%s\", path), &st) &&\n+\t\tif (!cstat(git_path(\"logs/%s\", path), &st) &&\n \t\t    S_ISREG(st.st_mode))\n \t\t\tit = path;\n \t\telse if (strcmp(ref, path) &&\n-\t\t\t !stat(git_path(\"logs/%s\", ref), &st) &&\n+\t\t\t !cstat(git_path(\"logs/%s\", ref), &st) &&\n \t\t\t S_ISREG(st.st_mode))\n \t\t\tit = ref;\n \t\telse\ndiff --git a/symlinks.c b/symlinks.c\nindex be9ace6..2a167a2 100644\n--- a/symlinks.c\n+++ b/symlinks.c\n@@ -1,4 +1,5 @@\n #include \"cache.h\"\n+#include \"stat-cache.h\"\n \n int has_symlink_leading_path(const char *name, char *last_symlink)\n {\n@@ -32,7 +33,7 @@ int has_symlink_leading_path(const char *name, char *last_symlink)\n \t\tmemcpy(dp, sp, len);\n \t\tdp[len] = 0;\n \n-\t\tif (lstat(path, &st))\n+\t\tif (clstat(path, &st))\n \t\t\treturn 0;\n \t\tif (S_ISLNK(st.st_mode)) {\n \t\t\tif (last_symlink)\ndiff --git a/test-chmtime.c b/test-chmtime.c\nindex 90da448..fec5ae2 100644\n--- a/test-chmtime.c\n+++ b/test-chmtime.c\n@@ -1,4 +1,5 @@\n #include \"git-compat-util.h\"\n+#include \"stat-cache.h\"\n #include <utime.h>\n \n static const char usage_str[] = \"(+|=|=+|=-|-)<seconds> <file>...\";\n@@ -37,7 +38,7 @@ int main(int argc, const char *argv[])\n \t\tstruct stat sb;\n \t\tstruct utimbuf utb;\n \n-\t\tif (stat(argv[i], &sb) < 0) {\n+\t\tif (cstat(argv[i], &sb) < 0) {\n \t\t\tfprintf(stderr, \"Failed to stat %s: %s\\n\",\n \t\t\t        argv[i], strerror(errno));\n \t\t\treturn -1;\ndiff --git a/transport.c b/transport.c\nindex 393e0e8..aace948 100644\n--- a/transport.c\n+++ b/transport.c\n@@ -11,6 +11,7 @@\n #include \"bundle.h\"\n #include \"dir.h\"\n #include \"refs.h\"\n+#include \"stat-cache.h\"\n \n /* rsync support */\n \n@@ -703,7 +704,7 @@ static int is_local(const char *url)\n static int is_file(const char *url)\n {\n \tstruct stat buf;\n-\tif (stat(url, &buf))\n+\tif (cstat(url, &buf))\n \t\treturn 0;\n \treturn S_ISREG(buf.st_mode);\n }\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex a59f475..ef7ac26 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -7,6 +7,7 @@\n #include \"unpack-trees.h\"\n #include \"progress.h\"\n #include \"refs.h\"\n+#include \"stat-cache.h\"\n \n static void add_entry(struct unpack_trees_options *o, struct cache_entry *ce,\n \tunsigned int set, unsigned int clear)\n@@ -409,7 +410,7 @@ static int verify_uptodate(struct cache_entry *ce,\n \tif (o->index_only || o->reset)\n \t\treturn 0;\n \n-\tif (!lstat(ce->name, &st)) {\n+\tif (!clstat(ce->name, &st)) {\n \t\tunsigned changed = ie_match_stat(o->src_index, ce, &st, CE_MATCH_IGNORE_VALID);\n \t\tif (!changed)\n \t\t\treturn 0;\n@@ -535,7 +536,7 @@ static int verify_absent(struct cache_entry *ce, const char *action,\n \tif (has_symlink_leading_path(ce->name, NULL))\n \t\treturn 0;\n \n-\tif (!lstat(ce->name, &st)) {\n+\tif (!clstat(ce->name, &st)) {\n \t\tint cnt;\n \t\tint dtype = ce_to_dtype(ce);\n \n@@ -931,7 +932,7 @@ int oneway_merge(struct cache_entry **src, struct unpack_trees_options *o)\n \t\tint update = 0;\n \t\tif (o->reset) {\n \t\t\tstruct stat st;\n-\t\t\tif (lstat(old->name, &st) ||\n+\t\t\tif (clstat(old->name, &st) ||\n \t\t\t    ie_match_stat(o->src_index, old, &st, CE_MATCH_IGNORE_VALID))\n \t\t\t\tupdate |= CE_UPDATE;\n \t\t}\ndiff --git a/xdiff-interface.c b/xdiff-interface.c\nindex 61dc5c5..7eb4a9e 100644\n--- a/xdiff-interface.c\n+++ b/xdiff-interface.c\n@@ -1,5 +1,6 @@\n #include \"cache.h\"\n #include \"xdiff-interface.h\"\n+#include \"stat-cache.h\"\n \n static int parse_num(char **cp_p, int *num_p)\n {\n@@ -147,7 +148,7 @@ int read_mmfile(mmfile_t *ptr, const char *filename)\n \tFILE *f;\n \tsize_t sz;\n \n-\tif (stat(filename, &st))\n+\tif (cstat(filename, &st))\n \t\treturn error(\"Could not stat %s\", filename);\n \tif ((f = fopen(filename, \"rb\")) == NULL)\n \t\treturn error(\"Could not open %s\", filename);\n-- \n1.5.5.76.gbb45.dirty\n"},{"id":"74802","messageId":"20080420111847.GC13411@bit.office.eurotux.com","threadId":"13180","inReplyTo":"20080420111346.GA13411@bit.office.eurotux.com","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Luciano Rocha","fromEmail":"luciano@eurotux.com","sentAt":"2008-04-20T11:18:47Z","receivedAt":"2008-04-20T11:18:47Z","isPatch":true,"sender":{"key":"luciano@eurotux.com","avatar":null},"body":"On Sun, Apr 20, 2008 at 12:13:46PM +0100, Luciano Rocha wrote:\n> An implementation of stat(2) and lstat(2) caching. Both the return code\n> and returned information are cached.\n> \n> Signed-off-by: Luciano Rocha <strange@nsk.no-ip.org>\n> ---\n> On Sat, Apr 19, 2008 at 03:39:37PM -0700, Linus Torvalds wrote:\n> > Yeah. I didn't look any further, but we do a total of *nine* 'lstat()' \n> > calls for each file we know about that is dirty, and *seven* when they are \n> > clean. Plus maybe a few more.\n> \n> That's a lot. Why not use a stat cache?\n> \n> With these changes, my git status . in WebKit changes from 28.215s to\n> 15.414s.\n\ngit status . in git changes from 0.477s to 0.412s.\n\nAll tests under OS X.\n\n-- \nLuciano Rocha <luciano@eurotux.com>\nEurotux Informática, S.A. <http://www.eurotux.com/>\n"},{"id":"74808","messageId":"alpine.LFD.1.10.0804200836310.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"20080420111346.GA13411@bit.office.eurotux.com","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-20T16:03:13Z","receivedAt":"2008-04-20T16:03:13Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Apr 2008, Luciano Rocha wrote:\n> \n> That's a lot. Why not use a stat cache?\n\nWell, the thing is, the OS _does_ a stat cache for us, and the one that \nthe OS maintains is a lot better, in that it works across processes and is \ncoherent with other processes changing things.\n\nAnd the thing is, your stat cache makes the *common* cases slower. I \ndidn't do a whole lot of testing, but on my machine, doing just a \"git \nstatus\" with and without your stat cache shows\n\n\tCurrent git 'master':\n\t\treal    0m0.302s\n\t\treal    0m0.308s\n\t\treal    0m0.314s\n\n\tWith your patch:\n\t\treal    0m0.352s\n\t\treal    0m0.354s\n\t\treal    0m0.355s\n\niow, it slowed down the case that I think matters more (the one you're \n*supposed* to use, and people most commonly do) by 15%.\n\nNow, admittedly, I also do think that we should generally optimize the \nslow cases more than we should care about things that are already very \nfast, so I do not think that it's wrong to say \"ok, let's make the really \nfast case a bit slower, in order to not be so slow in the bad case\", so in \nthat sense I do not think the slowdown is disastrous.\n\nBUT. \n\nI really dislike adding a cache that is there just because we do something \nstupid. We can fix the over-abundance of lstat() calls by just being \nsmarter. And the smarter we are, the less the cache will help, and the \nmore it will hurt. Which is the real reason why I think the cache is a \nreally really bad idea: it optimizes for the wrong kind of behavior.\n\nSo we have other caches and hashes we use, like the index itself, or the \nname lookup hash into the index, or the delta cache. Maintaining those \ncaches takes some effort too, but those caches aren't there because we're \ndoing something stupid, they are there because they allow us to do \nsomething smart.\n\nFor example, the index itself actually has really important semantic \ncharacteristics. And while the name hashing actually improves on index \nlookup performance, I'd never have implemented it if it wasn't for the \nfact that it was also designed to allow us to do case-insensitive lookups. \nAnd the delta cache is not hiding stupidity, it's literally avoiding very \nexpensive work that we can't avoid by being smarter.\n\nSo the stat cache is not horribly bad, but I think it's the wrong path to \ngo down. \n\n> With these changes, my git status . in WebKit changes from 28.215s to\n> 15.414s.\n\nOf course, one reason I don't think it's such a great idea is that on \nLinux, your stat cache doesn't even then end up helping _nearly_ as much \nas it does on OS X. You see an almost 50% improvement, so the 15% \n*deprovement* may not sound like much to you. But under Linux, the numbers \nare quite different:\n\n\"git status .\" with your patch:\n\n\treal    0m1.043s\n\treal    0m1.009s\n\treal    0m0.972s\n\nWith my trivial patch that just removed 2 of the 9 lstat calls:\n\n\treal    0m1.116s\n\treal    0m1.115s\n\treal    0m1.119s\n\nIOW, it does help the \".\" case on Linux, but only by a fairly small \namount. In fact, the improvement seems slightly smaller than the \npeformance degradation (~12% vs ~15%), but that is probably within the \nmargin of noise, so...\n\nSo another reason to avoid the stat cache is that it's really just working \naround an OS X deficiency.\n\nI'd rather work at avoiding more lstat calls. I know we can do it.\n\n\t\t\tLinus\n"},{"id":"74810","messageId":"85tzhwv6tt.fsf@lola.goethe.zz","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191422480.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-04-20T16:17:50Z","receivedAt":"2008-04-20T16:17:50Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Sat, 19 Apr 2008, Linus Torvalds wrote:\n>> \n>> Notice how this patch doesn' actually change the fundamental O(n^2) \n>> behaviour, but it makes it much cheaper by generally avoiding the \n>> expensive 'fnmatch' and 'strlen/strncmp' when they are obviously not \n>> needed.\n>\n> Side note: on the kenrel tree, it makes the (insane!) operation \n>\n> \tgit add $(git ls-files)\n>\n> go from 49 seconds down to 17 sec. So it does make a huge difference\n> for me, but I also want to point out that this really isn't a sane\n> operation to do (I also think that 17 sec is totally unacceptable, but\n> I cannot find it in me to care, since I don't think this is an\n> operation that anybody should ever do!)\n\nIt is my opinion that git should likely presort the patterns (not just\nhere), and should traverse the trees alphabetically.  In that case, a\nmerge-like algorithm will pretty much do the trick in O(n), with O(n lg\nn) preprocessing cost.\n\nPresorting can only be done approximately in the case of wildcards: for\nthose, we have two relevant points in the sort order: one where it can\nstart matching, one where it can't match anymore.\n\nThe easiest way to make this more efficient would be to retain the\nO(n*m) algorithm, but presort the patterns and let them trickle\nhead-first into the O(m) pattern list only when they start having a\nchance of matching, and remove them from the O(m) list once a non-match\nhas passed them alphabetically for good.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"74811","messageId":"85prskv6l8.fsf@lola.goethe.zz","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804191727270.2779@woody.linux-foundation.org","subject":"Re: Git performance on OS X","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-04-20T16:22:59Z","receivedAt":"2008-04-20T16:22:59Z","isPatch":false,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Taking three seconds for the warm-cache case for just 50,000 files is \n> ludicrous. That's about an order-and-a-half slower than what I see.\n>\n> Maybe my CPU is faster too (2.66GHz Core 2), but the thing is, Linux \n> really does tend to outperform others at a lot of these kinds of loads. \n> System calls are fast to begin with, and the Linux directory cache kicks \n> ass, if I do say so myself.\n>\n> OS X doth suck. \n\nOh, but OS X does utf8 unification and case folding: so it is actually\ndoing quite a bit more than Linux.\n\nOf course, doing that sucks for even more reasons, but it is not an\naccident.  It is intentional brain damage.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"74835","messageId":"20080420215700.GA18626@bit.office.eurotux.com","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804200836310.2779@woody.linux-foundation.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Luciano Rocha","fromEmail":"luciano@eurotux.com","sentAt":"2008-04-20T22:04:02Z","receivedAt":"2008-04-20T22:04:02Z","isPatch":true,"sender":{"key":"luciano@eurotux.com","avatar":null},"body":"On Sun, Apr 20, 2008 at 09:03:13AM -0700, Linus Torvalds wrote:\n> \n> \n> On Sun, 20 Apr 2008, Luciano Rocha wrote:\n> > \n> > That's a lot. Why not use a stat cache?\n> \n> Well, the thing is, the OS _does_ a stat cache for us, and the one that \n> the OS maintains is a lot better, in that it works across processes and is \n> coherent with other processes changing things.\n\nSure. I am even unsure if the cache didn't break any sanity check (did a\nfile change after ...? Did someone chdir(2)?).\n\n> And the thing is, your stat cache makes the *common* cases slower. I \n> didn't do a whole lot of testing, but on my machine, doing just a \"git \n> status\" with and without your stat cache shows\n<snip>\n\nWell, it can be improved. The memcpy can be avoided by using the stored\ndata directly, and a _or_die can be added for the common case.\n\n> Now, admittedly, I also do think that we should generally optimize the \n> slow cases more than we should care about things that are already very \n> fast, so I do not think that it's wrong to say \"ok, let's make the really \n> fast case a bit slower, in order to not be so slow in the bad case\", so in \n> that sense I do not think the slowdown is disastrous.\n> \n> BUT. \n> \n> I really dislike adding a cache that is there just because we do something \n> stupid. We can fix the over-abundance of lstat() calls by just being \n> smarter. And the smarter we are, the less the cache will help, and the \n> more it will hurt. Which is the real reason why I think the cache is a \n> really really bad idea: it optimizes for the wrong kind of behavior.\n\nI agree completly. If we can reduce the number of (l)stat calls to a\nsingle one per file, then we'll all be happier. But that kind of change\nis beyond my current understanding of git internals. ;)\n\nRegards,\nLuciano Rocha\n\n-- \nLuciano Rocha <luciano@eurotux.com>\nEurotux Informática, S.A. <http://www.eurotux.com/>\n"},{"id":"293737","messageId":"alpine.LFD.1.10.0804201520370.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"20080420215700.GA18626@bit.office.eurotux.com","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-20T22:29:02Z","receivedAt":"2008-04-20T22:29:02Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Apr 2008, Luciano Rocha wrote:\n> \n> Well, it can be improved. The memcpy can be avoided by using the stored\n> data directly, and a _or_die can be added for the common case.\n\nAgreed. I think a big part of the overhead is also the allocation cost, \nand that can probably be obliterated (or at least minimized) by using a \nspecial and much faster allocator (see \"alloc.c\") since none of the \nallocations will ever be free'd.\n\nThe one thing I liked about your patch was that I think we could be better \noff with wrapping \"lstat()\" for other reasons - the same way we wrap \nread/write calls in our own \"write_in_full()\" simplified library \nfunctions. I hate tracing them, for example, and a wrapper around lstat() \nwould have helped my efforts to avoid some of the unnecessary ones.\n\nSo I do think your stat cache could be improved, but for the reasons I \noutlined I would much prefer to make it unimportant instead.\n\nI do agree that actually actively removing stat calls requires a lot more \nsubtle interactions. We almost always *have* the stat information in the \nindex, but the problem with \"git status .\" is that we re-read the index so \nmany times (and then have to re-validate the stat info).\n\n"},{"id":"74847","messageId":"alpine.LFD.1.10.0804201556290.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804201520370.2779@woody.linux-foundation.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-20T23:07:35Z","receivedAt":"2008-04-20T23:07:35Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Apr 2008, Linus Torvalds wrote:\n> \n> I do agree that actually actively removing stat calls requires a lot more \n> subtle interactions. We almost always *have* the stat information in the \n> index, but the problem with \"git status .\" is that we re-read the index so \n> many times (and then have to re-validate the stat info).\n\nActually, looking closer, one of the issues seems to be not just the fact \nthat we throw out the index by re-reading it, but run_diff_files() does\n\n\t\t...\n                if (ce_uptodate(ce))\n                        continue;\n\n                changed = check_work_tree_entity(ce, &st, symcache);\n                if (changed) {\n\t\t\t...\n\nwhere that \"check_work_tree_entity()\" check is very expensive for deep \ndirectory structures, because it ends up checking the stat() information \nfo every single directory leading up to it.\n\nThere's some bug there, because it really shouldn't do that.\n\nThis causes lstat() patterns like\n\n\t..\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma/Boolean/15.6.4.2-2.js\", {st_mode=S_IFREG|0664, st_size=3197, ...}) = 0\n\tlstat(\"JavaScriptCore\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma/Boolean\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\t..\n\nie instead of doing just *one* lstat (on that file), it does six: the file \nitself, and the five directories leading up to it!\n\nThis is the *real* cause of WebKit having ~7 lstat's per file in the \nrepository - if it wasn't for this braindamage, we'd have just three \nlstat's per file for \"git status .\".\n\nWhat's really sad is how we do this for every file in a directory, so the \npattern actually ends up looking like\n\n\t...\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma/Boolean/15.6.4.1.js\", {st_mode=S_IFREG|0664, st_size=2164, ...}) = 0\n\tlstat(\"JavaScriptCore\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma/Boolean\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma/Boolean/15.6.4.2-1.js\", {st_mode=S_IFREG|0664, st_size=5219, ...}) = 0\n\tlstat(\"JavaScriptCore\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma/Boolean\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma/Boolean/15.6.4.2-2.js\", {st_mode=S_IFREG|0664, st_size=3197, ...}) = 0\n\tlstat(\"JavaScriptCore\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\tlstat(\"JavaScriptCore/tests/mozilla/ecma/Boolean\", {st_mode=S_IFDIR|0775, st_size=4096, ...}) = 0\n\t...\n\nie for deep directories with lots of files in them, we end up doing an \nlstat() on all the directories leading up to that directory oevr and over \nand over again - for each file in that directory.\n\nOops.\n\nWe're supposed to have that \"char *symcache\" thing to not do that, but it \ndoesn't actually work that way.\n\nJunio, what was the logic for that whole \"has_symlink_leading_path()\" \nthing? I forget. Whatever, it's broken. \n\n\t\tLinus\n"},{"id":"74854","messageId":"20080421005340.GA2631@dpotapov.dyndns.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804201556290.2779@woody.linux-foundation.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Dmitry Potapov","fromEmail":"dpotapov@gmail.com","sentAt":"2008-04-21T00:53:40Z","receivedAt":"2008-04-21T00:53:40Z","isPatch":true,"sender":{"key":"dpotapov@gmail.com","avatar":"https://avatars.githubusercontent.com/u/6568595?v=4"},"body":"On Sun, Apr 20, 2008 at 04:07:35PM -0700, Linus Torvalds wrote:\n>\n> Junio, what was the logic for that whole \"has_symlink_leading_path()\"\n> thing? I forget. Whatever, it's broken.\n\n===\ncommit f859c846e90b385c7ef873df22403529208ade50\nAuthor: Junio C Hamano <junkio@cox.net>\nDate:   Fri May 11 22:11:07 2007 -0700\n\n    Add has_symlink_leading_path() function.\n\n    When we are applying a patch that creates a blob at a path, or\n    when we are switching from a branch that does not have a blob at\n    the path to another branch that has one, we need to make sure\n    that there is nothing at the path in the working tree, as such a\n    file is a local modification made by the user that would be lost\n    by the operation.\n\n    Normally, lstat() on the path and making sure ENOENT is returned\n    is good enough for that purpose.  However there is a twist.  We\n    may be creating a regular file arch/x86_64/boot/Makefile, while\n    removing an existing symbolic link at arch/x86_64/boot that\n    points at existing ../i386/boot directory that has Makefile in\n    it.  We always first check without touching filesystem and then\n    perform the actual operation, so when we verify the new file,\n    arch/x86_64/boot/Makefile, does not exist, we haven't removed\n    the symbolic link arc/x86_64/boot symbolic link yet.  lstat() on\n    the file sees through the symbolic link and reports the file is\n    there, which is not what we want.\n\n    The function has_symlink_leading_path() function takes a path,\n    and sees if any of the leading directory component is a symbolic\n    link.\n\n    When files in a new directory are created, we tend to process\n    them together because both index and tree are sorted.  The\n    function takes advantage of this and allows the caller to cache\n    and reuse which symbolic link on the filesystem caused the\n    function to return true.\n\n    The calling sequence would be:\n\n        char last_symlink[PATH_MAX];\n\n            *last_symlink = '\\0';\n            for each index entry {\n                if (!lose)\n                        continue;\n                if (lstat(it))\n                        if (errno == ENOENT)\n                                ; /* happy */\n                        else\n                                error;\n                else if (has_symlink_leading_path(it, last_symlink))\n                        ; /* happy */\n                else\n                        error; /* would lose local changes */\n                unlink_entry(it, last_symlink);\n        }\n===\n\nAnd there are some cases where stat() on path is desirable:\nhttp://www.spinics.net/lists/git/msg63988.html\n\nSo while stat information for regular files is cached in the index,\nstat information for directories is not cached, and that appears to\nbe wrong. Maybe, Lucano's cache makes sense if it stores only stat\ninformation for directories.\n\nIIRC, some time ago, an otherwise reasonable patch for .gitignore was\nrejected just because it would drive the number calls to lstat() up as\nthese calls on directories are not cached in the index.\n\nDmitry\n"},{"id":"74855","messageId":"7vk5isatpe.fsf@gitster.siamese.dyndns.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804201556290.2779@woody.linux-foundation.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-04-21T01:21:33Z","receivedAt":"2008-04-21T01:21:33Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Junio, what was the logic for that whole \"has_symlink_leading_path()\" \n> thing?\n\nIf you have a tracked path a/b/c/d/e, and you changed your work tree to\nmake a/b to a symlink that points at a random directory, potentially\neven outside work tree, that has c/d/e in it, we should not be fooled by\nthe fact that lstat(\"a/b/c/d/e\") says \"yup, the file exists\".  As far as\ngit is concerned, that path does _not_ exist, as \"a/b\" is a symlink now.\n"},{"id":"74856","messageId":"alpine.LFD.1.10.0804201959590.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"7vk5isatpe.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-21T03:15:20Z","receivedAt":"2008-04-21T03:15:20Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Apr 2008, Junio C Hamano wrote:\n> Linus Torvalds <torvalds@linux-foundation.org> writes:\n> \n> > Junio, what was the logic for that whole \"has_symlink_leading_path()\" \n> > thing?\n> \n> If you have a tracked path a/b/c/d/e, and you changed your work tree to\n> make a/b to a symlink that points at a random directory, potentially\n> even outside work tree, that has c/d/e in it, we should not be fooled by\n> the fact that lstat(\"a/b/c/d/e\") says \"yup, the file exists\".  As far as\n> git is concerned, that path does _not_ exist, as \"a/b\" is a symlink now.\n\nOk, I can see the logic behind that, but the code is really dense and hard \nto read. And obviously very inefficient.\n\nHere's a trial balloon patch that totally revamps how that whole function \nworks. Instead of passing in a \"symlink_cache\" thing that it modifies for \nthe caller, it just has its totally *internal* cache of where it found the \nlast symlink, and what the last directory it found last time was.\n\nSo now the logic becomes:\n\n - if a pathname that is passed in matches the last known symlink prefix, \n   we don't even need to do anything else - it is known to have a symlink \n   prefix.\n\n - if the pathname that is passed in matches the last known directory \n   prefix, we start looking just from that point onward (since we know \n   that the leading part is a directory without symlinks)\n\nand this not only speeds things up regardless, it also cuts down lstat() \ncalls by a huge amount.\n\nOn that WebKit repo, and a \"git status .\", it used to do 338132 lstat() \ncalls. With this patch, it only does 141411. Which is still three per \npathname we know about, plus roughly one per directory we look at, but \nthat's a *lot* better.\n\nIt also improves performance from 1.125s to under one second for me on \nLinux. \n\nBut more fundamentally, I think it's more readable.\n\nCaveat: I do think we should add a way to invalidate the pathname caches \nwhen we turn a symlink into a directory or vice versa, so this patch isn't \nreally complete as-is, but I think it's a good start.\n\nAnd once we do that, I think the code is actually understandable. It was \nreally hard to see what the point of that \"last_symlink\" thing was.\n\nHmm?\n\n\t\tLinus\n---\n builtin-apply.c |    2 +-\n cache.h         |    2 +-\n diff-lib.c      |   10 +++---\n symlinks.c      |   78 +++++++++++++++++++++++++++++++++----------------------\n unpack-trees.c  |   12 +++-----\n 5 files changed, 59 insertions(+), 45 deletions(-)\n\ndiff --git a/builtin-apply.c b/builtin-apply.c\nindex caa3f2a..1103625 100644\n--- a/builtin-apply.c\n+++ b/builtin-apply.c\n@@ -2247,7 +2247,7 @@ static int check_to_create_blob(const char *new_name, int ok_if_exists)\n \t\t * In such a case, path \"new_name\" does not exist as\n \t\t * far as git is concerned.\n \t\t */\n-\t\tif (has_symlink_leading_path(new_name, NULL))\n+\t\tif (has_symlink_leading_path(strlen(new_name), new_name))\n \t\t\treturn 0;\n \n \t\treturn error(\"%s: already exists in working directory\", new_name);\ndiff --git a/cache.h b/cache.h\nindex c058125..6dc6543 100644\n--- a/cache.h\n+++ b/cache.h\n@@ -587,7 +587,7 @@ struct checkout {\n };\n \n extern int checkout_entry(struct cache_entry *ce, const struct checkout *state, char *topath);\n-extern int has_symlink_leading_path(const char *name, char *last_symlink);\n+extern int has_symlink_leading_path(int len, const char *name);\n \n extern struct alternate_object_database {\n \tstruct alternate_object_database *next;\ndiff --git a/diff-lib.c b/diff-lib.c\nindex 069e450..6a26b53 100644\n--- a/diff-lib.c\n+++ b/diff-lib.c\n@@ -338,14 +338,14 @@ int run_diff_files_cmd(struct rev_info *revs, int argc, const char **argv)\n  * See if work tree has an entity that can be staged.  Return 0 if so,\n  * return 1 if not and return -1 if error.\n  */\n-static int check_work_tree_entity(const struct cache_entry *ce, struct stat *st, char *symcache)\n+static int check_work_tree_entity(const struct cache_entry *ce, struct stat *st)\n {\n \tif (lstat(ce->name, st) < 0) {\n \t\tif (errno != ENOENT && errno != ENOTDIR)\n \t\t\treturn -1;\n \t\treturn 1;\n \t}\n-\tif (has_symlink_leading_path(ce->name, symcache))\n+\tif (has_symlink_leading_path(ce_namelen(ce), ce->name))\n \t\treturn 1;\n \tif (S_ISDIR(st->st_mode)) {\n \t\tunsigned char sub[20];\n@@ -399,7 +399,7 @@ int run_diff_files(struct rev_info *revs, unsigned int option)\n \t\t\tmemset(&(dpath->parent[0]), 0,\n \t\t\t       sizeof(struct combine_diff_parent)*5);\n \n-\t\t\tchanged = check_work_tree_entity(ce, &st, symcache);\n+\t\t\tchanged = check_work_tree_entity(ce, &st);\n \t\t\tif (!changed)\n \t\t\t\tdpath->mode = ce_mode_from_stat(ce, st.st_mode);\n \t\t\telse {\n@@ -463,7 +463,7 @@ int run_diff_files(struct rev_info *revs, unsigned int option)\n \t\tif (ce_uptodate(ce))\n \t\t\tcontinue;\n \n-\t\tchanged = check_work_tree_entity(ce, &st, symcache);\n+\t\tchanged = check_work_tree_entity(ce, &st);\n \t\tif (changed) {\n \t\t\tif (changed < 0) {\n \t\t\t\tperror(ce->name);\n@@ -521,7 +521,7 @@ static int get_stat_data(struct cache_entry *ce,\n \tif (!cached) {\n \t\tint changed;\n \t\tstruct stat st;\n-\t\tchanged = check_work_tree_entity(ce, &st, cbdata->symcache);\n+\t\tchanged = check_work_tree_entity(ce, &st);\n \t\tif (changed < 0)\n \t\t\treturn -1;\n \t\telse if (changed) {\ndiff --git a/symlinks.c b/symlinks.c\nindex be9ace6..04ce2d4 100644\n--- a/symlinks.c\n+++ b/symlinks.c\n@@ -1,48 +1,64 @@\n #include \"cache.h\"\n \n-int has_symlink_leading_path(const char *name, char *last_symlink)\n-{\n+struct pathname {\n+\tint len;\n \tchar path[PATH_MAX];\n-\tconst char *sp, *ep;\n-\tchar *dp;\n+};\n \n-\tsp = name;\n-\tdp = path;\n+/* Return matching pathname prefix length, or zero if not matching */\n+static inline int match_pathname(int len, const char *name, struct pathname *match)\n+{\n+\tint match_len = match->len;\n+\treturn (len > match_len &&\n+\t\tname[match_len] == '/' &&\n+\t\t!memcmp(name, match->path, match_len)) ? match_len : 0;\n+}\n \n-\tif (last_symlink && *last_symlink) {\n-\t\tsize_t last_len = strlen(last_symlink);\n-\t\tsize_t len = strlen(name);\n-\t\tif (last_len < len &&\n-\t\t    !strncmp(name, last_symlink, last_len) &&\n-\t\t    name[last_len] == '/')\n-\t\t\treturn 1;\n-\t\t*last_symlink = '\\0';\n+static inline void set_pathname(int len, const char *name, struct pathname *match)\n+{\n+\tif (len < PATH_MAX) {\n+\t\tmatch->len = len;\n+\t\tmemcpy(match->path, name, len);\n+\t\tmatch->path[len] = 0;\n \t}\n+}\n \n-\twhile (1) {\n-\t\tsize_t len;\n-\t\tstruct stat st;\n+int has_symlink_leading_path(int len, const char *name)\n+{\n+\tstatic struct pathname link, nonlink;\n+\tchar path[PATH_MAX];\n+\tstruct stat st;\n+\tchar *sp;\n+\tint known_dir;\n+\n+\t/*\n+\t * See if the last known symlink cache matches.\n+\t */\n+\tif (match_pathname(len, name, &link))\n+\t\treturn 1;\n \n-\t\tep = strchr(sp, '/');\n-\t\tif (!ep)\n-\t\t\tbreak;\n-\t\tlen = ep - sp;\n-\t\tif (PATH_MAX <= dp + len - path + 2)\n-\t\t\treturn 0; /* new name is longer than that??? */\n-\t\tmemcpy(dp, sp, len);\n-\t\tdp[len] = 0;\n+\t/*\n+\t * Get rid of the last known directory part\n+\t */\n+\tknown_dir = match_pathname(len, name, &nonlink);\n+\t    \t\n+\twhile ((sp = strchr(name + known_dir + 1, '/')) != NULL) {\n+\t\tint thislen = sp - name ;\n+\t\tmemcpy(path, name, thislen);\n+\t\tpath[thislen] = 0;\n \n \t\tif (lstat(path, &st))\n \t\t\treturn 0;\n+\t\tif (S_ISDIR(st.st_mode)) {\n+\t\t\tset_pathname(thislen, path, &nonlink);\n+\t\t\tknown_dir = thislen;\n+\t\t\tcontinue;\n+\t\t}\n \t\tif (S_ISLNK(st.st_mode)) {\n-\t\t\tif (last_symlink)\n-\t\t\t\tstrcpy(last_symlink, path);\n+\t\t\tset_pathname(thislen, path, &link);\n \t\t\treturn 1;\n \t\t}\n-\n-\t\tdp[len++] = '/';\n-\t\tdp = dp + len;\n-\t\tsp = ep + 1;\n+\t\tbreak;\n \t}\n \treturn 0;\n }\ndiff --git a/unpack-trees.c b/unpack-trees.c\nindex feae846..1ab28fd 100644\n--- a/unpack-trees.c\n+++ b/unpack-trees.c\n@@ -26,11 +26,12 @@ static void add_entry(struct unpack_trees_options *o, struct cache_entry *ce,\n  * directories, in case this unlink is the removal of the\n  * last entry in the directory -- empty directories are removed.\n  */\n-static void unlink_entry(char *name, char *last_symlink)\n+static void unlink_entry(struct cache_entry *ce)\n {\n \tchar *cp, *prev;\n+\tchar *name = ce->name;\n \n-\tif (has_symlink_leading_path(name, last_symlink))\n+\tif (has_symlink_leading_path(ce_namelen(ce), ce->name))\n \t\treturn;\n \tif (unlink(name))\n \t\treturn;\n@@ -58,7 +59,6 @@ static int check_updates(struct unpack_trees_options *o)\n {\n \tunsigned cnt = 0, total = 0;\n \tstruct progress *progress = NULL;\n-\tchar last_symlink[PATH_MAX];\n \tstruct index_state *index = &o->result;\n \tint i;\n \tint errs = 0;\n@@ -75,14 +75,13 @@ static int check_updates(struct unpack_trees_options *o)\n \t\tcnt = 0;\n \t}\n \n-\t*last_symlink = '\\0';\n \tfor (i = 0; i < index->cache_nr; i++) {\n \t\tstruct cache_entry *ce = index->cache[i];\n \n \t\tif (ce->ce_flags & CE_REMOVE) {\n \t\t\tdisplay_progress(progress, ++cnt);\n \t\t\tif (o->update)\n-\t\t\t\tunlink_entry(ce->name, last_symlink);\n+\t\t\t\tunlink_entry(ce);\n \t\t\tremove_index_entry_at(&o->result, i);\n \t\t\ti--;\n \t\t\tcontinue;\n@@ -97,7 +96,6 @@ static int check_updates(struct unpack_trees_options *o)\n \t\t\tce->ce_flags &= ~CE_UPDATE;\n \t\t\tif (o->update) {\n \t\t\t\terrs |= checkout_entry(ce, &state, NULL);\n-\t\t\t\t*last_symlink = '\\0';\n \t\t\t}\n \t\t}\n \t}\n@@ -553,7 +551,7 @@ static int verify_absent(struct cache_entry *ce, const char *action,\n \tif (o->index_only || o->reset || !o->update)\n \t\treturn 0;\n \n-\tif (has_symlink_leading_path(ce->name, NULL))\n+\tif (has_symlink_leading_path(ce_namelen(ce), ce->name))\n \t\treturn 0;\n \n \tif (!lstat(ce->name, &st)) {\n"},{"id":"74857","messageId":"alpine.LFD.1.10.0804202016080.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804201959590.2779@woody.linux-foundation.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-21T03:20:09Z","receivedAt":"2008-04-21T03:20:09Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 20 Apr 2008, Linus Torvalds wrote:\n> On Sun, 20 Apr 2008, Junio C Hamano wrote:\n> > \n> > If you have a tracked path a/b/c/d/e, and you changed your work tree to\n> > make a/b to a symlink that points at a random directory, potentially\n> > even outside work tree, that has c/d/e in it, we should not be fooled by\n> > the fact that lstat(\"a/b/c/d/e\") says \"yup, the file exists\".  As far as\n> > git is concerned, that path does _not_ exist, as \"a/b\" is a symlink now.\n> \n> Ok, I can see the logic behind that, but the code is really dense and hard \n> to read. And obviously very inefficient.\n\nOne more note: I think that if we really care about this, we should do \nthis inside \"ce_match_stat()\", so that we catch it in *all* the cases \nwhere we match against the stat information. \n\nAs it is, the \"diff\" mechanism (and \"apply\") knows to check whether a \ndirectory has changed into a symlink, but it looks like doing a simple \n\"git update-index --refresh\" will never even test it, so it will never \nnotice that the index isn't actually up-to-date if a directory has been \nmoved and the old directory has been replaced by a symlink to the new \nlocation.\n\nHmm?\n\n\t\tLinus\n"},{"id":"74860","messageId":"200804211041.41760.johan@herland.net","threadId":"13180","inReplyTo":"20080421005340.GA2631@dpotapov.dyndns.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Johan Herland","fromEmail":"johan@herland.net","sentAt":"2008-04-21T08:41:41Z","receivedAt":"2008-04-21T08:41:41Z","isPatch":true,"sender":{"key":"johan@herland.net","avatar":"https://avatars.githubusercontent.com/u/547031?v=4"},"body":"On Monday 21 April 2008, Dmitry Potapov wrote:\n> On Sun, Apr 20, 2008 at 04:07:35PM -0700, Linus Torvalds wrote:\n> > Junio, what was the logic for that whole \"has_symlink_leading_path()\"\n> > thing? I forget. Whatever, it's broken.\n> \n> ===\n> commit f859c846e90b385c7ef873df22403529208ade50\n> Author: Junio C Hamano <junkio@cox.net>\n> Date:   Fri May 11 22:11:07 2007 -0700\n> \n[snip snip]\n> ===\n> \n> And there are some cases where stat() on path is desirable:\n> http://www.spinics.net/lists/git/msg63988.html\n> \n> So while stat information for regular files is cached in the index,\n> stat information for directories is not cached, and that appears to\n> be wrong. Maybe, Lucano's cache makes sense if it stores only stat\n> information for directories.\n> \n> IIRC, some time ago, an otherwise reasonable patch for .gitignore was\n> rejected just because it would drive the number calls to lstat() up as\n> these calls on directories are not cached in the index.\n\nPardon me for butting in (and I'm honestly NOT trying to start a flamewar),\nbut I'm wondering if this could be solved by tracking directories in the\nindex. AFAICS it would:\n\n- Help bring the number of lstat() calls down (since we can cache the\n  lstat() results for directories like we currently do for regular files)\n\n- More easily detect complicated cases like \"add across symlinks\" (see\n  Junio's email at the spinics.net link above)\n\n- (less important) When discussing empty directory support several months\n  ago, ISTR one of the biggest hurdles being that directories were not\n  tracked in the index\n\n\nI don't know much about how the index is implemented (few do, I think), so\nif there is a glaringly obvious reason why tracking directories in the\nindex is a bad idea, please enlighten me.\n\n\n...Johan\n\n-- \nJohan Herland, <johan@herland.net>\nwww.herland.net\n"},{"id":"74864","messageId":"85d4ojseve.fsf@lola.goethe.zz","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804201520370.2779@woody.linux-foundation.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"David Kastrup","fromEmail":"dak@gnu.org","sentAt":"2008-04-21T10:04:37Z","receivedAt":"2008-04-21T10:04:37Z","isPatch":true,"sender":{"key":"dak@gnu.org","avatar":"https://avatars.githubusercontent.com/u/52141349?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> So I do think your stat cache could be improved, but for the reasons I\n> outlined I would much prefer to make it unimportant instead.\n\nUsing a cache for a single algorithmic task is probably a mistake: a\ncache tries to keep some data around on the assumption that it might get\nused.  So it tends to either waste lots of memory or keep the wrong\ndata.  And the reloads increase with the size of the processed data.\n\nUsing a sorted-traverse-and-merge algorithm instead never needs to\nreload data and relinquishes it as soon as it is no longer needed.\n\nA stat cache is fine for an operating system which has no clue about\nwhat access patterns to except next.\n\nBut in this case, our application has the whole task outlines in\nadvance, and it makes sense organizing it in the best manner.\n\n-- \nDavid Kastrup, Kriemhildstr. 15, 44793 Bochum\n"},{"id":"74879","messageId":"7v3apfawry.fsf@gitster.siamese.dyndns.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804201959590.2779@woody.linux-foundation.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-04-21T18:27:29Z","receivedAt":"2008-04-21T18:27:29Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> Here's a trial balloon patch that totally revamps how that whole function \n> works. Instead of passing in a \"symlink_cache\" thing that it modifies for \n> the caller, it just has its totally *internal* cache of where it found the \n> last symlink, and what the last directory it found last time was.\n>\n> So now the logic becomes:\n>\n>  - if a pathname that is passed in matches the last known symlink prefix, \n>    we don't even need to do anything else - it is known to have a symlink \n>    prefix.\n>\n>  - if the pathname that is passed in matches the last known directory \n>    prefix, we start looking just from that point onward (since we know \n>    that the leading part is a directory without symlinks)\n\nThat makes sense.\n\n> Caveat: I do think we should add a way to invalidate the pathname caches \n> when we turn a symlink into a directory or vice versa, so this patch isn't \n> really complete as-is, but I think it's a good start.\n\nTrue.\n\nThere are a few patches in flight that are not in 'master' (Dmitry quoted\none of them), that use more has_symlink_leading_path() calls.  In\nretrospect, the function was misnamed.  It describes what it checks\n(i.e. \"does the path have leading component that is a symlink?\") but I\nprobably should have named it after what it really wants to tell\n(i.e. \"lstat(2) says this exists, but does it really, from the point of\nview of git?\")\n\nDoesn't it become very tempting to replace lstat() calls we make to check\nthe status of a work tree path, with a function git_wtstat() that is:\n\n        int git_wtstat(const char *path, struct stat *st)\n        {\n                int status = lstat(path, st);\n\n                if (status)\n                        return status;\n\n                if (!has_symlink_leading_path(path, strlen(path)))\n                        return 0;\n\n                /*\n                 * As far as git is concerned, this does not exist in\n                 * the work tree!\n                 */\n                errno = ENOENT;\n                return -1;\n        }\n\nThis unfortunately is not enough to hide the need for has_symlink calls\nfrom outside callers.  When we check out a new path \"a/b/c/d/e\", for\nexample, if we naively checked if we creat(2) \"a/b/c/d/e\" (and otherwise\nwe try the equivalent of \"mkdir -p\"), we would be tricked by a symlink\n\"a/b\" that points at some random place that has \"c/d\" subdirectory in it,\nand we need to unlink \"a/b\" first, and the above git_wtstat() does not\nreally help such codepath.\n"},{"id":"74881","messageId":"alpine.LFD.1.10.0804211203460.2779@woody.linux-foundation.org","threadId":"13180","inReplyTo":"7v3apfawry.fsf@gitster.siamese.dyndns.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Linus Torvalds","fromEmail":"torvalds@linux-foundation.org","sentAt":"2008-04-21T19:09:43Z","receivedAt":"2008-04-21T19:09:43Z","isPatch":true,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Mon, 21 Apr 2008, Junio C Hamano wrote:\n> \n> Doesn't it become very tempting to replace lstat() calls we make to check\n> the status of a work tree path, with a function git_wtstat() that is:\n\nYes.\n\nThat looks like a very good abstraction.\n\n>                 /*\n>                  * As far as git is concerned, this does not exist in\n>                  * the work tree!\n>                  */\n>                 errno = ENOENT;\n>                 return -1;\n>         }\n\nWell, how about returning something else than \"ENOENT\" here? \n\nAs you point out, git doesn't actually think this is a \"does not exist\" \ncase, but something else that may require more work:\n\n> This unfortunately is not enough to hide the need for has_symlink calls\n> from outside callers.  When we check out a new path \"a/b/c/d/e\", for\n> example, if we naively checked if we creat(2) \"a/b/c/d/e\" (and otherwise\n> we try the equivalent of \"mkdir -p\"), we would be tricked by a symlink\n> \"a/b\" that points at some random place that has \"c/d\" subdirectory in it,\n> and we need to unlink \"a/b\" first, and the above git_wtstat() does not\n> really help such codepath.\n\nMaybe ENOTDIR would be a better error return? That would conceptually be \nwhat an OS that refuses to follow symlinks at path walk time (because it \ndoesn't support symlinks as such) would return: the symlink component \nwould not be a directory, so it's as if you were trying to use a path \na/b/c/d/e where \"a/b\" isn't even a directory.\n\nIn fact, even on Linux, ENOTDIR is what an lstat() would return if \"b\" had \nbeen turned from a directory into a regular file - which is conceptually \n(for git) _exactly_ the same as \"b\" being a symlink.\n\nNo?\n\n\t\t\tLinus\n"},{"id":"74883","messageId":"7vprsj9dmc.fsf@gitster.siamese.dyndns.org","threadId":"13180","inReplyTo":"alpine.LFD.1.10.0804211203460.2779@woody.linux-foundation.org","subject":"Re: [PATCH 01/02/RFC] implement a stat cache","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2008-04-21T20:06:35Z","receivedAt":"2008-04-21T20:06:35Z","isPatch":true,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Linus Torvalds <torvalds@linux-foundation.org> writes:\n\n> On Mon, 21 Apr 2008, Junio C Hamano wrote:\n>> \n>> Doesn't it become very tempting to replace lstat() calls we make to check\n>> the status of a work tree path, with a function git_wtstat() that is:\n>\n> Yes.\n>\n> That looks like a very good abstraction.\n>\n>>                 /*\n>>                  * As far as git is concerned, this does not exist in\n>>                  * the work tree!\n>>                  */\n>>                 errno = ENOENT;\n>>                 return -1;\n>>         }\n>\n> Well, how about returning something else than \"ENOENT\" here? \n>\n> As you point out, git doesn't actually think this is a \"does not exist\" \n> case, but something else that may require more work:\n>\n>> This unfortunately is not enough to hide the need for has_symlink calls\n>> from outside callers.  When we check out a new path \"a/b/c/d/e\", for\n>> example, if we naively checked if we creat(2) \"a/b/c/d/e\" (and otherwise\n>> we try the equivalent of \"mkdir -p\"), we would be tricked by a symlink\n>> \"a/b\" that points at some random place that has \"c/d\" subdirectory in it,\n>> and we need to unlink \"a/b\" first, and the above git_wtstat() does not\n>> really help such codepath.\n>\n> Maybe ENOTDIR would be a better error return?\n\nYeah, and we could return which component in the given path is the\noffending one at the same time.\n\nIn the above example, we would say \"No, a/b/c/d/e does not exist because\na/b is a symlink\".  But would that be enough, I have to wonder.  lstat(2)\nmay have already said \"There is no a/b/c/d/f\" in the same example, but we\nstill need to know \"a/b\" is an unwanted symbolic link if the reason we are\nasking that question is because we would want to check out \"a/b/c/d/f\".\n\nSo the answer need to be \"a/b/c/d/f\" (does not exist|exists in the work\ntree), and it cannot exist because \"a/b\" is a symlink for such a caller.\n\nBut when we are trying to git-add, we simply do not care such\ndistinction.\n"}]}