git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re:

From
Linus Torvalds <torvalds@linux-foundation.org>
Date
May 8, 2009, 21:49 UTC
Message-ID
<alpine.LFD.2.01.0905081432150.4983@localhost.localdomain>
In-Reply-To
<OWEdfN5mNBoNl1TcdOvhhNfi_nLsao-aFrHkz_rNtuX_4lqXHisfcQ@cipher.nrlssc.navy.mil>
On Fri, 8 May 2009, Brandon Casey wrote:
Show 10 quoted lines
> 
> Before (cold cache):
> % time     seconds  usecs/call     calls    errors syscall
> ------ ----------- ----------- --------- --------- ----------------
>  98.60    6.365501         111     57432           lstat64
> 
> After (cold cache, no lstat fix, just cache_preload):
> % time     seconds  usecs/call     calls    errors syscall
> ------ ----------- ----------- --------- --------- ----------------
>  90.90   23.717981         413     57432           lstat64

Yes, interesting. I really smells like it's all fixed performance and there is a single lock around it. That 111us -> 413us increase is very consistent with four cores all serializing on the same lock. So it parallelizes to all four cores, but then will take exactly as long in total.

Quite frankly, 2.6.9 is so old that I have absolutely _no_ memory of what we used to do back then. Not that I follow NFS all that much even now - I did some of the original page cache and dentry work on the Linux NFS client way back when, but that was when I actually used NFS (and we were converting everything to the page cache).

I've long since forgotten everything I knew, and I'm just as happy about that. But clearly something is bad, and equally clearly it worked much better for you a couple of months ago. Which does imply that there's probably some centos issues.

Can you ask your MIS people if it would be possible to at least _test_ a new kernel? In 2.6.9, I'm quite frankly inclined to just say "it will likely never get fixed unless centos knows what it is", but if you test a more modern kernel and see similar issues, then I'll be intrigued.

It's kind of sad, but at the same time, NFS was using the BKL up into 2.6.26 or something like that (about a year ago). And your kernel is based on something _much_ older.

That said, even with the BKL, NFS should allow all the actual IO to be done in parallel (since the BKL is dropped on scheduling). But it's really wasting a _lot_ of CPU time, and that hurts you enormously, even though the cold-cache case still seems to win, judging by your other email:

Show 9 quoted lines
> Best without patch: 6.02 (systime 1.57)
> 
>   0.43user 1.57system 0:06.02elapsed 33%CPU (0avgtext+0avgdata 0maxresident)k
>   5336inputs+0outputs (12major+15472minor)pagefaults 0swaps
> 
> Best with patch (preload_cache,lstat reduction): 2.69 (systime 10.47)
> 
>   0.45user 10.47system 0:02.69elapsed 405%CPU (0avgtext+0avgdata 0maxresident)k
>   5336inputs+0outputs (12major+13985minor)pagefaults 0swaps

so there's a _huge_ increase in system time (again), but the change from 33% CPU -> 405% CPU makes up for it and you get lower elapsed times.

But that 7x increase in system time really is sad. I do suspect it's likely due to spinning on the BKL. And if so, then a modern kernel should fix it.

			Linus
Previous: Brandon CaseyNext: Brandon Casey
Message 29 of 34 in “(unknown)”
  1. Bevan WatkissMay 7, 2009
  2. Alex RiesenMay 7, 2009
  3. Bevan WatkissMay 7, 2009
  4. Alex RiesenMay 7, 2009
  5. Bevan WatkissMay 7, 2009
  6. Björn SteinbrinkMay 7, 2009
  7. Linus TorvaldsMay 7, 2009
  8. Bevan WatkissMay 7, 2009
  9. Linus TorvaldsMay 7, 2009
  10. Linus TorvaldsMay 7, 2009
  11. Junio C HamanoMay 7, 2009
  12. Linus TorvaldsMay 7, 2009
  13. Linus TorvaldsMay 7, 2009
  14. david@lang.hmMay 7, 2009
  15. Linus TorvaldsMay 7, 2009
  16. david@lang.hmMay 7, 2009
  17. Linus TorvaldsMay 7, 2009
  18. david@lang.hmMay 7, 2009
  19. Linus TorvaldsMay 7, 2009
  20. david@lang.hmMay 7, 2009
  21. Johan HerlandMay 7, 2009
  22. Bevan WatkissMay 8, 2009
  23. Alex RiesenMay 8, 2009
  24. Linus TorvaldsMay 8, 2009
  25. Brandon CaseyMay 8, 2009
  26. Linus TorvaldsMay 8, 2009
  27. Brandon CaseyMay 8, 2009
  28. Brandon CaseyMay 8, 2009
  29. Linus TorvaldsMay 8, 2009
  30. Brandon CaseyMay 8, 2009
  31. Linus TorvaldsMay 9, 2009
  32. Linus TorvaldsMay 8, 2009
  33. 'git checkout' and unlink() calls (was: Re: )Kjetil Barvik, May 8, 2009
  34. Linus TorvaldsMay 8, 2009

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.