{"thread":{"id":"6261","subject":"Re: [KORG] Re: kernel.org lies about latest -mm kernel","startedAt":"2007-01-07T04:22:31Z","lastAt":"2007-01-12T10:54:43Z","messageCount":52,"participants":["Jeff Garzik","Linus Torvalds","H. Peter Anvin","Willy Tarreau","Andrew Morton","Rene Herman","Christoph Hellwig","Jan Engelhardt","Robert Fitzsimons","Krzysztof Halasa","Randy Dunlap","J.H.","Greg KH","Shawn O. Pearce","Junio C Hamano","Martin Langhoff","Jakub Narebski","Suparna Bhattacharya","Theodore Tso","Pavel Machek","Johannes Stezenbach","Nicolas Pitre","Paul Jackson","Jeremy Higdon","Fengguang Wu","Nigel Cunningham"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"31006","messageId":"45A07587.3080503@garzik.org","threadId":"6261","inReplyTo":"1168140954.2153.1.camel@nigel.suspend2.net","subject":"Re: [KORG] Re: kernel.org lies about latest -mm kernel","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2007-01-07T04:22:31Z","receivedAt":"2007-01-07T04:22:31Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"> On Tue, 2006-12-26 at 08:49 -0800, H. Peter Anvin wrote:\n>> Not really.  In fact, it would hardly help at all.\n>>\n>> The two things git users can do to help is:\n>>\n>> 1. Make sure your alternatives file is set up correctly;\n>> 2. Keep your trees packed and pruned, to keep the file count down.\n>>\n>> If you do this, the load imposed by a single git tree is fairly negible.\n\n\nWould kernel hackers be amenable to having their trees auto-repacked, \nand linked via alternatives to Linus's linux-2.6.git?\n\nLooking through kernel.org, we have a ton of repositories, however \npacked, that carrying their own copies of the linux-2.6.git repo.\n\nAlso, I wonder if \"git push\" will push only the non-linux-2.6.git \nobjects, if both local and remote sides have the proper alternatives set up?\n\n\tJeff\n"},{"id":"31007","messageId":"Pine.LNX.4.64.0701062029170.3661@woody.osdl.org","threadId":"6261","inReplyTo":"45A07587.3080503@garzik.org","subject":"Re: [KORG] Re: kernel.org lies about latest -mm kernel","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-07T04:29:26Z","receivedAt":"2007-01-07T04:29:26Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 6 Jan 2007, Jeff Garzik wrote:\n> \n> Also, I wonder if \"git push\" will push only the non-linux-2.6.git objects, if\n> both local and remote sides have the proper alternatives set up?\n\nYes.\n\n\t\tLinus\n"},{"id":"31008","messageId":"45A083F2.5000000@zytor.com","threadId":"6261","inReplyTo":"45A08269.4050504@zytor.com","subject":"How git affects kernel.org performance","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2007-01-07T05:24:02Z","receivedAt":"2007-01-07T05:24:02Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Some more data on how git affects kernel.org...\n\nDuring extremely high load, it appears that what slows kernel.org down \nmore than anything else is the time that each individual getdents() call \ntakes.  When I've looked this I've observed times from 200 ms to almost \n2 seconds!  Since an unpacked *OR* unpruned git tree adds 256 \ndirectories to a cleanly packed tree, you can do the math yourself.\n\nI have tried reducing vm.vfs_cache_pressure down to 1 on the kernel.org \nmachines in order to improve the situation, but even at that point it \nappears the kernel doesn't readily hold the entire directory hierarchy \nin memory, even though there is space to do so.  I have suggested that \nwe might want to add a sysctl to change the denominator from the default \n100.\n\nThe one thing that we need done locally is to have a smart uploader, \ninstead of relying on rsync.  That, unfortunately, is a fairly sizable \nproject.\n\n\t-hpa\n"},{"id":"31009","messageId":"Pine.LNX.4.64.0701062130260.3661@woody.osdl.org","threadId":"6261","inReplyTo":"45A083F2.5000000@zytor.com","subject":"Re: How git affects kernel.org performance","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-07T05:39:42Z","receivedAt":"2007-01-07T05:39:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sat, 6 Jan 2007, H. Peter Anvin wrote:\n> \n> During extremely high load, it appears that what slows kernel.org down more\n> than anything else is the time that each individual getdents() call takes.\n> When I've looked this I've observed times from 200 ms to almost 2 seconds!\n> Since an unpacked *OR* unpruned git tree adds 256 directories to a cleanly\n> packed tree, you can do the math yourself.\n\n\"getdents()\" is totally serialized by the inode semaphore. It's one of the \nmost expensive system calls in Linux, partly because of that, and partly \nbecause it has to call all the way down into the filesystem in a way that \nalmost no other common system call has to (99% of all filesystem calls can \nbe handled basically at the VFS layer with generic caches - but not \ngetdents()).\n\nSo if there are concurrent readdirs on the same directory, they get \nserialized. If there is any file creation/deletion activity in the \ndirectory, it serializes getdents(). \n\nTo make matters worse, I don't think it has any read-ahead at all when you \nuse hashed directory entries. So if you have cold-cache case, you'll read \nevery single block totally individually, and serialized. One block at a \ntime (I think the non-hashed case is likely also suspect, but that's a \nseparate issue)\n\nIn other words, I'm not at all surprised it hits on filldir time. \nEspecially on ext3.\n\n\t\tLinus\n"},{"id":"31019","messageId":"20070107085526.GR24090@1wt.eu","threadId":"6261","inReplyTo":"Pine.LNX.4.64.0701062130260.3661@woody.osdl.org","subject":"Re: How git affects kernel.org performance","fromName":"Willy Tarreau","fromEmail":"w@1wt.eu","sentAt":"2007-01-07T08:55:26Z","receivedAt":"2007-01-07T08:55:26Z","isPatch":false,"sender":{"key":"w@1wt.eu","avatar":"https://avatars.githubusercontent.com/u/8141789?v=4"},"body":"On Sat, Jan 06, 2007 at 09:39:42PM -0800, Linus Torvalds wrote:\n> \n> \n> On Sat, 6 Jan 2007, H. Peter Anvin wrote:\n> > \n> > During extremely high load, it appears that what slows kernel.org down more\n> > than anything else is the time that each individual getdents() call takes.\n> > When I've looked this I've observed times from 200 ms to almost 2 seconds!\n> > Since an unpacked *OR* unpruned git tree adds 256 directories to a cleanly\n> > packed tree, you can do the math yourself.\n> \n> \"getdents()\" is totally serialized by the inode semaphore. It's one of the \n> most expensive system calls in Linux, partly because of that, and partly \n> because it has to call all the way down into the filesystem in a way that \n> almost no other common system call has to (99% of all filesystem calls can \n> be handled basically at the VFS layer with generic caches - but not \n> getdents()).\n> \n> So if there are concurrent readdirs on the same directory, they get \n> serialized. If there is any file creation/deletion activity in the \n> directory, it serializes getdents(). \n> \n> To make matters worse, I don't think it has any read-ahead at all when you \n> use hashed directory entries. So if you have cold-cache case, you'll read \n> every single block totally individually, and serialized. One block at a \n> time (I think the non-hashed case is likely also suspect, but that's a \n> separate issue)\n> \n> In other words, I'm not at all surprised it hits on filldir time. \n> Especially on ext3.\n\nAt work, we had the same problem on a file server with ext3. We use rsync\nto make backups to a local IDE disk, and we noticed that getdents() took\nabout the same time as Peter reports (0.2 to 2 seconds), especially in\nmaildir directories. We tried many things to fix it with no result,\nincluding enabling dirindexes. Finally, we made a full backup, and switched\nover to XFS and the problem totally disappeared. So it seems that the\nfilesystem matters a lot here when there are lots of entries in a\ndirectory, and that ext3 is not suitable for usages with thousands\nof entries in directories with millions of files on disk. I'm not\ncertain it would be that easy to try other filesystems on kernel.org\nthough :-/\n\nWilly\n"},{"id":"31020","messageId":"45A0B63E.2020803@zytor.com","threadId":"6261","inReplyTo":"20070107085526.GR24090@1wt.eu","subject":"Re: How git affects kernel.org performance","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2007-01-07T08:58:38Z","receivedAt":"2007-01-07T08:58:38Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Willy Tarreau wrote:\n> \n> At work, we had the same problem on a file server with ext3. We use rsync\n> to make backups to a local IDE disk, and we noticed that getdents() took\n> about the same time as Peter reports (0.2 to 2 seconds), especially in\n> maildir directories. We tried many things to fix it with no result,\n> including enabling dirindexes. Finally, we made a full backup, and switched\n> over to XFS and the problem totally disappeared. So it seems that the\n> filesystem matters a lot here when there are lots of entries in a\n> directory, and that ext3 is not suitable for usages with thousands\n> of entries in directories with millions of files on disk. I'm not\n> certain it would be that easy to try other filesystems on kernel.org\n> though :-/\n> \n\nChanging filesystems would mean about a week of downtime for a server. \nIt's painful, but it's doable; however, if we get a traffic spike during \nthat time it'll hurt like hell.\n\nHowever, if there is credible reasons to believe XFS will help, I'd be \ninclined to try it out.\n\n\t-hpa\n"},{"id":"31021","messageId":"20070107090336.GA7741@1wt.eu","threadId":"6261","inReplyTo":"45A0B63E.2020803@zytor.com","subject":"Re: How git affects kernel.org performance","fromName":"Willy Tarreau","fromEmail":"w@1wt.eu","sentAt":"2007-01-07T09:03:36Z","receivedAt":"2007-01-07T09:03:36Z","isPatch":false,"sender":{"key":"w@1wt.eu","avatar":"https://avatars.githubusercontent.com/u/8141789?v=4"},"body":"On Sun, Jan 07, 2007 at 12:58:38AM -0800, H. Peter Anvin wrote:\n> Willy Tarreau wrote:\n> >\n> >At work, we had the same problem on a file server with ext3. We use rsync\n> >to make backups to a local IDE disk, and we noticed that getdents() took\n> >about the same time as Peter reports (0.2 to 2 seconds), especially in\n> >maildir directories. We tried many things to fix it with no result,\n> >including enabling dirindexes. Finally, we made a full backup, and switched\n> >over to XFS and the problem totally disappeared. So it seems that the\n> >filesystem matters a lot here when there are lots of entries in a\n> >directory, and that ext3 is not suitable for usages with thousands\n> >of entries in directories with millions of files on disk. I'm not\n> >certain it would be that easy to try other filesystems on kernel.org\n> >though :-/\n> >\n> \n> Changing filesystems would mean about a week of downtime for a server. \n> It's painful, but it's doable; however, if we get a traffic spike during \n> that time it'll hurt like hell.\n> \n> However, if there is credible reasons to believe XFS will help, I'd be \n> inclined to try it out.\n\nThe problem is that I have no sufficient FS knowledge to argument why\nit helps here. It was a desperate attempt to fix the problem for us\nand it definitely worked well.\n\nHmmm I'm thinking about something very dirty : would it be possible\nto reduce the current FS size to get more space to create another\nFS ? Supposing you create a XX GB/TB XFS after the current ext3,\nyou would be able to mount it in some directories with --bind and\nslowly switch some parts to it. The problem with this approach is\nthat it will never be 100% converted, but as an experiment it might\nbe worth it, no ?\n\nWilly\n"},{"id":"31022","messageId":"20070107011542.3496bc76.akpm@osdl.org","threadId":"6261","inReplyTo":"20070107085526.GR24090@1wt.eu","subject":"Re: How git affects kernel.org performance","fromName":"Andrew Morton","fromEmail":"akpm@osdl.org","sentAt":"2007-01-07T09:15:42Z","receivedAt":"2007-01-07T09:15:42Z","isPatch":false,"sender":{"key":"akpm@osdl.org","avatar":null},"body":"On Sun, 7 Jan 2007 09:55:26 +0100\nWilly Tarreau <w@1wt.eu> wrote:\n\n> On Sat, Jan 06, 2007 at 09:39:42PM -0800, Linus Torvalds wrote:\n> > \n> > \n> > On Sat, 6 Jan 2007, H. Peter Anvin wrote:\n> > > \n> > > During extremely high load, it appears that what slows kernel.org down more\n> > > than anything else is the time that each individual getdents() call takes.\n> > > When I've looked this I've observed times from 200 ms to almost 2 seconds!\n> > > Since an unpacked *OR* unpruned git tree adds 256 directories to a cleanly\n> > > packed tree, you can do the math yourself.\n> > \n> > \"getdents()\" is totally serialized by the inode semaphore. It's one of the \n> > most expensive system calls in Linux, partly because of that, and partly \n> > because it has to call all the way down into the filesystem in a way that \n> > almost no other common system call has to (99% of all filesystem calls can \n> > be handled basically at the VFS layer with generic caches - but not \n> > getdents()).\n> > \n> > So if there are concurrent readdirs on the same directory, they get \n> > serialized. If there is any file creation/deletion activity in the \n> > directory, it serializes getdents(). \n> > \n> > To make matters worse, I don't think it has any read-ahead at all when you \n> > use hashed directory entries. So if you have cold-cache case, you'll read \n> > every single block totally individually, and serialized. One block at a \n> > time (I think the non-hashed case is likely also suspect, but that's a \n> > separate issue)\n> > \n> > In other words, I'm not at all surprised it hits on filldir time. \n> > Especially on ext3.\n> \n> At work, we had the same problem on a file server with ext3. We use rsync\n> to make backups to a local IDE disk, and we noticed that getdents() took\n> about the same time as Peter reports (0.2 to 2 seconds), especially in\n> maildir directories. We tried many things to fix it with no result,\n> including enabling dirindexes. Finally, we made a full backup, and switched\n> over to XFS and the problem totally disappeared. So it seems that the\n> filesystem matters a lot here when there are lots of entries in a\n> directory, and that ext3 is not suitable for usages with thousands\n> of entries in directories with millions of files on disk. I'm not\n> certain it would be that easy to try other filesystems on kernel.org\n> though :-/\n> \n\nYeah, slowly-growing directories will get splattered all over the disk.\n\nPossible short-term fixes would be to just allocate up to (say) eight\nblocks when we grow a directory by one block.  Or teach the\ndirectory-growth code to use ext3 reservations.\n\nLonger-term people are talking about things like on-disk rerservations. \nBut I expect directories are being forgotten about in all of that.\n"},{"id":"31027","messageId":"45A0BF8C.4040508@gmail.com","threadId":"6261","inReplyTo":"20070107011542.3496bc76.akpm@osdl.org","subject":"Re: How git affects kernel.org performance","fromName":"Rene Herman","fromEmail":"rene.herman@gmail.com","sentAt":"2007-01-07T09:38:20Z","receivedAt":"2007-01-07T09:38:20Z","isPatch":false,"sender":{"key":"rene.herman@gmail.com","avatar":null},"body":"On 01/07/2007 10:15 AM, Andrew Morton wrote:\n\n> Yeah, slowly-growing directories will get splattered all over the\n> disk.\n> \n> Possible short-term fixes would be to just allocate up to (say) eight\n>  blocks when we grow a directory by one block.  Or teach the \n> directory-growth code to use ext3 reservations.\n> \n> Longer-term people are talking about things like on-disk\n> rerservations. But I expect directories are being forgotten about in\n> all of that.\n\nI wish people would just talk about de2fsrag... ;-\\\n\nRene\n"},{"id":"31031","messageId":"20070107102853.GB26849@infradead.org","threadId":"6261","inReplyTo":"20070107090336.GA7741@1wt.eu","subject":"Re: How git affects kernel.org performance","fromName":"Christoph Hellwig","fromEmail":"hch@infradead.org","sentAt":"2007-01-07T10:28:53Z","receivedAt":"2007-01-07T10:28:53Z","isPatch":false,"sender":{"key":"hch@infradead.org","avatar":null},"body":"On Sun, Jan 07, 2007 at 10:03:36AM +0100, Willy Tarreau wrote:\n> The problem is that I have no sufficient FS knowledge to argument why\n> it helps here. It was a desperate attempt to fix the problem for us\n> and it definitely worked well.\n\nXFS does rather efficient btree directories, and it does sophisticated\nreadahead for directories.  I suspect that's what is helping you there.\n"},{"id":"31035","messageId":"Pine.LNX.4.61.0701071141580.4365@yvahk01.tjqt.qr","threadId":"6261","inReplyTo":"20070107090336.GA7741@1wt.eu","subject":"Re: How git affects kernel.org performance","fromName":"Jan Engelhardt","fromEmail":"jengelh@linux01.gwdg.de","sentAt":"2007-01-07T10:50:57Z","receivedAt":"2007-01-07T10:50:57Z","isPatch":false,"sender":{"key":"jengelh@linux01.gwdg.de","avatar":null},"body":"\nOn Jan 7 2007 10:03, Willy Tarreau wrote:\n>On Sun, Jan 07, 2007 at 12:58:38AM -0800, H. Peter Anvin wrote:\n>> >[..]\n>> >entries in directories with millions of files on disk. I'm not\n>> >certain it would be that easy to try other filesystems on\n>> >kernel.org though :-/\n>> \n>> Changing filesystems would mean about a week of downtime for a server. \n>> It's painful, but it's doable; however, if we get a traffic spike during \n>> that time it'll hurt like hell.\n\nThen make sure noone releases a kernel ;-)\n\n>> However, if there is credible reasons to believe XFS will help, I'd be \n>> inclined to try it out.\n>\n>Hmmm I'm thinking about something very dirty : would it be possible\n>to reduce the current FS size to get more space to create another\n>FS ? Supposing you create a XX GB/TB XFS after the current ext3,\n>you would be able to mount it in some directories with --bind and\n>slowly switch some parts to it. The problem with this approach is\n>that it will never be 100% converted, but as an experiment it might\n>be worth it, no ?\n\nMuch better: rsync from /oldfs to /newfs, stop all ftp uploads, rsync\nagain to catch any new files that have been added until the ftp\nupload was closed, then do _one_ (technically two) mountpoint moves\n(as opposed to Willy's idea of \"some directories\") in a mere second\nalong the lines of\n\n  mount --move /oldfs /older; mount --move /newfs /oldfs.\n\nlet old transfers that still use files in /older complete (lsof or\nfuser -m), then disconnect the old volume. In case /newfs (now\n/oldfs) is a volume you borrowed from someone and need to return it,\nwell, I guess you need to rsync back somehow.\n\n\n\t-`J'\n-- \n"},{"id":"31033","messageId":"20070107105230.GA8345@1wt.eu","threadId":"6261","inReplyTo":"20070107102853.GB26849@infradead.org","subject":"Re: How git affects kernel.org performance","fromName":"Willy Tarreau","fromEmail":"w@1wt.eu","sentAt":"2007-01-07T10:52:30Z","receivedAt":"2007-01-07T10:52:30Z","isPatch":false,"sender":{"key":"w@1wt.eu","avatar":"https://avatars.githubusercontent.com/u/8141789?v=4"},"body":"On Sun, Jan 07, 2007 at 10:28:53AM +0000, Christoph Hellwig wrote:\n> On Sun, Jan 07, 2007 at 10:03:36AM +0100, Willy Tarreau wrote:\n> > The problem is that I have no sufficient FS knowledge to argument why\n> > it helps here. It was a desperate attempt to fix the problem for us\n> > and it definitely worked well.\n> \n> XFS does rather efficient btree directories, and it does sophisticated\n> readahead for directories.  I suspect that's what is helping you there.\n\nOk. Do you too think it might help (or even solve) the problem on\nkernel.org ?\n\nWilly\n"},{"id":"31053","messageId":"20070107145730.GB24706@localhost","threadId":"6261","inReplyTo":"45A083F2.5000000@zytor.com","subject":"Re: How git affects kernel.org performance","fromName":"Robert Fitzsimons","fromEmail":"robfitz@273k.net","sentAt":"2007-01-07T14:57:30Z","receivedAt":"2007-01-07T14:57:30Z","isPatch":false,"sender":{"key":"robfitz@273k.net","avatar":null},"body":"> Some more data on how git affects kernel.org...\n\nI have a quick question about the gitweb configuration, does the\n$projects_list config entry point to a directory or a file?\n\nWhen it is a directory gitweb ends up doing the equivalent of a 'find\n$project_list' to find all the available projects, so it really should\nbe changed to a projects list file.\n\nRobert\n"},{"id":"31054","messageId":"m3odpazxit.fsf@defiant.localdomain","threadId":"6261","inReplyTo":"45A083F2.5000000@zytor.com","subject":"Re: How git affects kernel.org performance","fromName":"Krzysztof Halasa","fromEmail":"khc@pm.waw.pl","sentAt":"2007-01-07T15:06:50Z","receivedAt":"2007-01-07T15:06:50Z","isPatch":false,"sender":{"key":"khc@pm.waw.pl","avatar":null},"body":"\"H. Peter Anvin\" <hpa@zytor.com> writes:\n\n> During extremely high load, it appears that what slows kernel.org down\n> more than anything else is the time that each individual getdents()\n> call takes.  When I've looked this I've observed times from 200 ms to\n> almost 2 seconds!  Since an unpacked *OR* unpruned git tree adds 256\n> directories to a cleanly packed tree, you can do the math yourself.\n\nHmm... Perhaps it should be possible to push git updates as a pack\nfile only? I mean, the pack file would stay packed = never individual\nfiles and never 256 directories?\n\nPeople aren't doing commit/etc. activity there, right?\n-- \nKrzysztof Halasa\n"},{"id":"31063","messageId":"Pine.LNX.4.64.0701070957080.3661@woody.osdl.org","threadId":"6261","inReplyTo":"20070107102853.GB26849@infradead.org","subject":"Re: How git affects kernel.org performance","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-07T18:17:38Z","receivedAt":"2007-01-07T18:17:38Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 7 Jan 2007, Christoph Hellwig wrote:\n>\n> On Sun, Jan 07, 2007 at 10:03:36AM +0100, Willy Tarreau wrote:\n> > The problem is that I have no sufficient FS knowledge to argument why\n> > it helps here. It was a desperate attempt to fix the problem for us\n> > and it definitely worked well.\n> \n> XFS does rather efficient btree directories, and it does sophisticated\n> readahead for directories.  I suspect that's what is helping you there.\n\nThe sad part is that this is a long-standing issue, and the directory \nreading code in ext3 really _should_ be able to do ok. \n\nA year or two ago I did a totally half-assed code for the non-hashed \nreaddir that improved performance by an order of magnitude for ext3 for a \ntest-case of mine, but it was subtly buggy and didn't do the hashed case \nAT ALL. Andrew fixed it up so that it at least wasn't subtly buggy any \nmore, but in the process it also lost all capability of doing fragmented \ndirectories (so it doesn't help very much any more under exactly the \nsituation that is the worst case), and it still doesn't do the hashed \ndirectory case.\n\nIt's my personal pet peeve with ext3 (as Andrew can attest). And it's \nreally sad, because I don't think it is fundamental per se, but the way \nthe directory handling and jdb are done, it's apparently very hard to fix.\n\n(It's clearly not _impossible_ to do: I think that it should be possible \nto treat ext3 directories the same way we treat files, except they would \nalways be in \"data=journal\" mode. But I understand ext2, not ext3 (and \nabsolutely not jbd), so I'm not going to be able to do anything about it \npersonally).\n\nAnyway, I think that disabling hashing can actually help. And I suspect \nthat even with hashing enabled, there should be some quick hack for making \nthe directory reading at least be able to do multiple outstanding reads in \nparallel, instead of reading the blocks totally synchronously (\"read five \nblocks, then wait for the one we care\" rather than the current \"read one \nblock at a time, wait for it, read the next one, wait for it..\" \nsituation).\n\n\t\t\tLinus\n"},{"id":"31065","messageId":"20070107104943.ee2c5e6f.randy.dunlap@oracle.com","threadId":"6261","inReplyTo":"Pine.LNX.4.61.0701071141580.4365@yvahk01.tjqt.qr","subject":"Re: How git affects kernel.org performance","fromName":"Randy Dunlap","fromEmail":"randy.dunlap@oracle.com","sentAt":"2007-01-07T18:49:43Z","receivedAt":"2007-01-07T18:49:43Z","isPatch":false,"sender":{"key":"randy.dunlap@oracle.com","avatar":null},"body":"On Sun, 7 Jan 2007 11:50:57 +0100 (MET) Jan Engelhardt wrote:\n\n> \n> On Jan 7 2007 10:03, Willy Tarreau wrote:\n> >On Sun, Jan 07, 2007 at 12:58:38AM -0800, H. Peter Anvin wrote:\n> >> >[..]\n> >> >entries in directories with millions of files on disk. I'm not\n> >> >certain it would be that easy to try other filesystems on\n> >> >kernel.org though :-/\n> >> \n> >> Changing filesystems would mean about a week of downtime for a server. \n> >> It's painful, but it's doable; however, if we get a traffic spike during \n> >> that time it'll hurt like hell.\n> \n> Then make sure noone releases a kernel ;-)\n\nmaybe the week of LCA ?\n\n> >> However, if there is credible reasons to believe XFS will help, I'd be \n> >> inclined to try it out.\n> >\n> >Hmmm I'm thinking about something very dirty : would it be possible\n> >to reduce the current FS size to get more space to create another\n> >FS ? Supposing you create a XX GB/TB XFS after the current ext3,\n> >you would be able to mount it in some directories with --bind and\n> >slowly switch some parts to it. The problem with this approach is\n> >that it will never be 100% converted, but as an experiment it might\n> >be worth it, no ?\n> \n> Much better: rsync from /oldfs to /newfs, stop all ftp uploads, rsync\n> again to catch any new files that have been added until the ftp\n> upload was closed, then do _one_ (technically two) mountpoint moves\n> (as opposed to Willy's idea of \"some directories\") in a mere second\n> along the lines of\n> \n>   mount --move /oldfs /older; mount --move /newfs /oldfs.\n> \n> let old transfers that still use files in /older complete (lsof or\n> fuser -m), then disconnect the old volume. In case /newfs (now\n> /oldfs) is a volume you borrowed from someone and need to return it,\n> well, I guess you need to rsync back somehow.\n\n---\n~Randy\n"},{"id":"31070","messageId":"Pine.LNX.4.61.0701072004290.4365@yvahk01.tjqt.qr","threadId":"6261","inReplyTo":"20070107104943.ee2c5e6f.randy.dunlap@oracle.com","subject":"Re: How git affects kernel.org performance","fromName":"Jan Engelhardt","fromEmail":"jengelh@linux01.gwdg.de","sentAt":"2007-01-07T19:07:43Z","receivedAt":"2007-01-07T19:07:43Z","isPatch":false,"sender":{"key":"jengelh@linux01.gwdg.de","avatar":null},"body":"\nOn Jan 7 2007 10:49, Randy Dunlap wrote:\n>On Sun, 7 Jan 2007 11:50:57 +0100 (MET) Jan Engelhardt wrote:\n>> On Jan 7 2007 10:03, Willy Tarreau wrote:\n>> >On Sun, Jan 07, 2007 at 12:58:38AM -0800, H. Peter Anvin wrote:\n>> >> >[..]\n>> >> >entries in directories with millions of files on disk. I'm not\n>> >> >certain it would be that easy to try other filesystems on\n>> >> >kernel.org though :-/\n>> >> \n>> >> Changing filesystems would mean about a week of downtime for a server. \n>> >> It's painful, but it's doable; however, if we get a traffic spike during \n>> >> that time it'll hurt like hell.\n>> \n>> Then make sure noone releases a kernel ;-)\n>\n>maybe the week of LCA ?\n\nI don't know that acronym, but if you ask me when it should happen:\n_Before_ the next big thing is released, e.g. before 2.6.20-final.\nReason: You never know how long they're chewing [downloading] on 2.6.20.\nExcluding other projects on kernel.org from my hypothesis, I'd suppose the\nlowest bandwidth usage the longer no new files have been released. (Because\neveryone has them then more or less.)\n\n\n\t-`J'\n-- \n"},{"id":"31067","messageId":"1168197145.14963.1.camel@localhost.localdomain","threadId":"6261","inReplyTo":"20070107145730.GB24706@localhost","subject":"Re: How git affects kernel.org performance","fromName":"J.H.","fromEmail":"warthog9@kernel.org","sentAt":"2007-01-07T19:12:25Z","receivedAt":"2007-01-07T19:12:25Z","isPatch":false,"sender":{"key":"warthog9@kernel.org","avatar":"https://avatars.githubusercontent.com/u/2334704?v=4"},"body":"With my gitweb caching changes this isn't as big of a deal as the front\npage is only generated once every 10 minutes or so (and with the changes\nI'm working on today that timeout will be variable)\n\n- John\n\nOn Sun, 2007-01-07 at 14:57 +0000, Robert Fitzsimons wrote:\n> > Some more data on how git affects kernel.org...\n> \n> I have a quick question about the gitweb configuration, does the\n> $projects_list config entry point to a directory or a file?\n> \n> When it is a directory gitweb ends up doing the equivalent of a 'find\n> $project_list' to find all the available projects, so it really should\n> be changed to a projects list file.\n> \n> Robert\n"},{"id":"31068","messageId":"Pine.LNX.4.64.0701071028450.3661@woody.osdl.org","threadId":"6261","inReplyTo":"Pine.LNX.4.64.0701070957080.3661@woody.osdl.org","subject":"Re: How git affects kernel.org performance","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-07T19:13:06Z","receivedAt":"2007-01-07T19:13:06Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 7 Jan 2007, Linus Torvalds wrote:\n> \n> A year or two ago I did a totally half-assed code for the non-hashed \n> readdir that improved performance by an order of magnitude for ext3 for a \n> test-case of mine, but it was subtly buggy and didn't do the hashed case \n> AT ALL.\n\nBtw, this isn't the test-case, but it's a half-way re-creation of \nsomething like it. It's _really_ stupid, but here's what you can do:\n\n - compile and run this idiotic program. It creates a directory called \n   \"throwaway\" that is ~44kB in size, and if I did things right, it should \n   not be totally contiguous on disk with the current ext3 allocation \n   logic.\n\n - as root, do \"echo 3 > /proc/sys/vm/drop_caches\" to get a cache-cold \n   schenario.\n\n - do \"time ls throwaway > /dev/null\".\n\nI don't know what people consider to be reasonable performance, but for \nme, it takes about half a second to do a simple \"ls\". NOTE! This is _not_ \nreading inode stat information or anything like that. It literally takes \n0.3-0.4 seconds to read ~44kB off the disk. That's a whopping 125kB/s \nthroughput on a reasonably fast modern disk.\n\nThat's what we in the industry call  \"sad\".\n\nAnd that's on a totally unloaded machine. There was _nothing_ else going \non. No IO congestion, no nothing. Just the cost of synchronously doing \nten or eleven disk reads.\n\nThe fix?\n\n - proper read-ahead. Right now, even if the directory is totally \n   contiguous on disk (just remove the thing that writes data to the \n   files, so that you'll have empty files instead of 8kB files), I think \n   we do those reads totally synchronously if the filesystem was mounted \n   with directory hashing enabled.\n\n   Without hashing, the directory will be much smaller too, so readdir() \n   will have less data to read. And it _should_ do some readahead, \n   although in my testing, the best I could do was still 0.185s for a (now \n   shrunken) 28kB directory. \n\n - better directory block allocation patterns would likely help a lot, \n   rather than single blocks. That's true even without any read-ahead (at \n   least the disk wouldn't need to seek, and any on-disk track buffers etc \n   would work better), but with read-ahead and contiguous blocks it should \n   be just a couple of IO's (the indirect stuff means that it's more than \n   one), and so you should see much better IO patterns because the \n   elevator can try to help too.\n\nMaybe I just have unrealistic expectations, but I really don't like how a \nfairly small 50kB directory takes an appreciable fraction of a second to \nread.\n\nOnce it's cached, it still takes too long, but at least at that point the \nindividual getdents calls take just tens of microseconds.\n\nHere's cold-cache numbers (notice: 34 msec for the first one, and 17 msec \nin the middle.. The 5-6ms range indicates a single IO for the intermediate \nones, which basically says that each call does roughly one IO, except the \nfirst one that does ~5 (probably the indirect index blocks), and two in \nthe middle who are able to fill up the buffer from the IO done by the \nprevious one (4kB buffers, so if the previous getdents() happened to just \nread the beginning of a block, the next one might be able to fill \neverything from that block without having to do IO).\n\n\tgetdents(3, /* 103 entries */, 4096)    = 4088 <0.034830>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.006703>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.006719>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000354>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000017>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.005302>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.016957>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000017>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.003530>\n\tgetdents(3, /* 83 entries */, 4096)     = 3320 <0.000296>\n\tgetdents(3, /* 0 entries */, 4096)      = 0 <0.000006>\n\nHere's the pure CPU overhead: still pretty high (200 usec! For a single \nsystem call! That's disgusting! In contrast, a 4kB read() call takes 7 \nusec on this machine, so the overhead of doing things one dentry at a \ntime, and calling down to several layers of filesystem is quite high):\n\n\tgetdents(3, /* 103 entries */, 4096)    = 4088 <0.000204>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000122>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000112>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000153>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000018>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000103>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000217>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000018>\n\tgetdents(3, /* 102 entries */, 4096)    = 4080 <0.000095>\n\tgetdents(3, /* 83 entries */, 4096)     = 3320 <0.000089>\n\tgetdents(3, /* 0 entries */, 4096)      = 0 <0.000006>\n\nbut you can see the difference.. The real cost is obviously the IO.\n\n\t\tLinus\n\n----\n#include <stdio.h>\n#include <stdlib.h>\n#include <unistd.h>\n#include <fcntl.h>\n#include <sys/stat.h>\n#include <sys/types.h>\n\nstatic char buffer[8192];\n\nstatic int create_file(const char *name)\n{\n\tint fd = open(name, O_RDWR | O_CREAT | O_TRUNC, 0666);\n\tif (fd < 0)\n\t\treturn fd;\n\n\twrite(fd, buffer, sizeof(buffer));\n\tclose(fd);\n\treturn 0;\n}\n\nint main(int argc, char **argv)\n{\n\tint i;\n\tchar name[256];\n\n\t/* Fill up the buffer with some random garbage */\n\tfor (i = 0; i < sizeof(buffer); i++)\n\t\tbuffer[i] = \"abcdefghijklmnopqrstuvwxyz\\n\"[i % 27];\n\n\tif (mkdir(\"throwaway\", 0777) < 0 || chdir(\"throwaway\") < 0) {\n\t\tperror(\"throwaway\");\n\t\texit(1);\n\t}\n\n\t/*\n\t * Create a reasonably big directory by having a number\n\t * of files with non-trivial filenames, and with some\n\t * real content to fragment the directory blocks..\n\t */\n\tfor (i = 0; i < 1000; i++) {\n\t\tsnprintf(name, sizeof(name),\n\t\t\t\"file-name-%d-%d-%d-%d\",\n\t\t\ti / 1000,\n\t\t\t(i / 100) % 10,\n\t\t\t(i / 10) % 10,\n\t\t\t(i / 1) % 10);\n\t\tcreate_file(name);\n\t}\n\treturn 0;\n}\n"},{"id":"31071","messageId":"20070107112834.a8746a98.randy.dunlap@oracle.com","threadId":"6261","inReplyTo":"Pine.LNX.4.61.0701072004290.4365@yvahk01.tjqt.qr","subject":"Re: How git affects kernel.org performance","fromName":"Randy Dunlap","fromEmail":"randy.dunlap@oracle.com","sentAt":"2007-01-07T19:28:34Z","receivedAt":"2007-01-07T19:28:34Z","isPatch":false,"sender":{"key":"randy.dunlap@oracle.com","avatar":null},"body":"On Sun, 7 Jan 2007 20:07:43 +0100 (MET) Jan Engelhardt wrote:\n\n> \n> On Jan 7 2007 10:49, Randy Dunlap wrote:\n> >On Sun, 7 Jan 2007 11:50:57 +0100 (MET) Jan Engelhardt wrote:\n> >> On Jan 7 2007 10:03, Willy Tarreau wrote:\n> >> >On Sun, Jan 07, 2007 at 12:58:38AM -0800, H. Peter Anvin wrote:\n> >> >> >[..]\n> >> >> >entries in directories with millions of files on disk. I'm not\n> >> >> >certain it would be that easy to try other filesystems on\n> >> >> >kernel.org though :-/\n> >> >> \n> >> >> Changing filesystems would mean about a week of downtime for a server. \n> >> >> It's painful, but it's doable; however, if we get a traffic spike during \n> >> >> that time it'll hurt like hell.\n> >> \n> >> Then make sure noone releases a kernel ;-)\n> >\n> >maybe the week of LCA ?\n\nSorry, it means Linux.conf.au (Australia):\n  http://lca2007.linux.org.au/\nJan. 15-20, 2007\n\n> I don't know that acronym, but if you ask me when it should happen:\n> _Before_ the next big thing is released, e.g. before 2.6.20-final.\n> Reason: You never know how long they're chewing [downloading] on 2.6.20.\n> Excluding other projects on kernel.org from my hypothesis, I'd suppose the\n> lowest bandwidth usage the longer no new files have been released. (Because\n> everyone has them then more or less.)\n\nISTM that Linus is trying to make 2.6.20-final before LCA.  We'll see.\n\n---\n~Randy\n"},{"id":"31072","messageId":"Pine.LNX.4.64.0701071132450.3661@woody.osdl.org","threadId":"6261","inReplyTo":"9e4733910701071126r7931042eldfb73060792f4f41@mail.gmail.com","subject":"Re: How git affects kernel.org performance","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-07T19:35:40Z","receivedAt":"2007-01-07T19:35:40Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 7 Jan 2007, Jon Smirl wrote:\n> > \n> >  - proper read-ahead. Right now, even if the directory is totally\n> >    contiguous on disk (just remove the thing that writes data to the\n> >    files, so that you'll have empty files instead of 8kB files), I think\n> >    we do those reads totally synchronously if the filesystem was mounted\n> >    with directory hashing enabled.\n> \n> What's the status on the Adaptive Read-ahead patch from Wu Fengguang\n> <wfg@mail.ustc.edu.cn> ? That patch really helped with read ahead\n> problems I was having with mmap. It was in mm forever and I've lost\n> track of it.\n\nWon't help. ext3 does NO readahead at all. It doesn't use the general VFS \nhelper routines to read data (because it doesn't use the page cache), it \njust does the raw buffer-head IO directly.\n\n(In the non-indexed case, it does do some read-ahead, and it uses the \ngeneric routines for it, but because it does everything by physical \naddress, even the generic routines will decide that it's just doing random \nreading if the directory isn't physically contiguous - and stop reading \nahead).\n\n(I may have missed some case where it does do read-ahead in the index \nroutines, so don't take my word as being unquestionably true. I'm _fairly_ \nsure, but..)\n\n\t\t\tLinus\n"},{"id":"31073","messageId":"Pine.LNX.4.64.0701071136110.3661@woody.osdl.org","threadId":"6261","inReplyTo":"20070107112834.a8746a98.randy.dunlap@oracle.com","subject":"Re: How git affects kernel.org performance","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-07T19:37:42Z","receivedAt":"2007-01-07T19:37:42Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 7 Jan 2007, Randy Dunlap wrote:\n> \n> ISTM that Linus is trying to make 2.6.20-final before LCA.  We'll see.\n\nNo. Hopefully \"final -rc\" before LCA, but I'll do the actual 2.6.20 \nrelease afterwards. I don't want to have a merge window during LCA, as I \nand many others will all be out anyway. So it's much better to have LCA \nhappen during the end of the stabilization phase when there's hopefully \nnot a lot going on.\n\n(Of course, often at the end of the stabilization phase there is all the \n\"ok, what about regression XyZ?\" panic)\n\n\t\tLinus\n"},{"id":"31075","messageId":"20070107201146.GA21956@suse.de","threadId":"6261","inReplyTo":"45A07587.3080503@garzik.org","subject":"Re: [KORG] Re: kernel.org lies about latest -mm kernel","fromName":"Greg KH","fromEmail":"gregkh@suse.de","sentAt":"2007-01-07T20:11:46Z","receivedAt":"2007-01-07T20:11:46Z","isPatch":false,"sender":{"key":"gregkh@suse.de","avatar":"https://gravatar.com/avatar/e52bfe8b8ad890236109deb2ce59960a1584dc070d18487f7a2b5273d709fcdc?d=mp&s=160"},"body":"On Sat, Jan 06, 2007 at 11:22:31PM -0500, Jeff Garzik wrote:\n> >On Tue, 2006-12-26 at 08:49 -0800, H. Peter Anvin wrote:\n> >>Not really.  In fact, it would hardly help at all.\n> >>\n> >>The two things git users can do to help is:\n> >>\n> >>1. Make sure your alternatives file is set up correctly;\n> >>2. Keep your trees packed and pruned, to keep the file count down.\n> >>\n> >>If you do this, the load imposed by a single git tree is fairly negible.\n> \n> \n> Would kernel hackers be amenable to having their trees auto-repacked, \n> and linked via alternatives to Linus's linux-2.6.git?\n> \n> Looking through kernel.org, we have a ton of repositories, however \n> packed, that carrying their own copies of the linux-2.6.git repo.\n\nWell, I create my repos by doing a:\n\tgit clone -l --bare\nwhich makes a hardlink from Linus's tree.\n\nBut then it gets copied over to the public server, which probably severs\nthat hardlink :(\n\nAny shortcut to clone or set up a repo using \"alternatives\" so that we\ndon't have this issue at all?\n\nthanks,\n\ngreg k-h\n"},{"id":"31077","messageId":"20070107203120.GA4970@spearce.org","threadId":"6261","inReplyTo":"m3odpazxit.fsf@defiant.localdomain","subject":"Re: How git affects kernel.org performance","fromName":"Shawn O. Pearce","fromEmail":"spearce@spearce.org","sentAt":"2007-01-07T20:31:20Z","receivedAt":"2007-01-07T20:31:20Z","isPatch":false,"sender":{"key":"spearce@spearce.org","avatar":"https://avatars.githubusercontent.com/u/34844?v=4"},"body":"Krzysztof Halasa <khc@pm.waw.pl> wrote:\n> Hmm... Perhaps it should be possible to push git updates as a pack\n> file only? I mean, the pack file would stay packed = never individual\n> files and never 256 directories?\n\nLatest Git does this.  If the server is later than 1.4.3.3 then\nthe receive-pack process can actually store the pack file rather\nthan unpacking it into loose objects.  The downside is that it will\ncopy any missing base objects onto the end of a thin pack to make\nit not-thin.\n\nThere's actually a limit that controls when to keep the pack and when\nnot to (receive.unpackLimit).  In 1.4.3.3 this defaulted to 5000\nobjects, which meant all but the largest pushes will be exploded\ninto loose objects.  In 1.5.0-rc0 that limit changed from 5000 to\n100, though Nico did a lot of study and discovered that the optimum\nis likely 3.  But that tends to create too many pack files so 100\nwas arbitrarily chosen.\n\nSo if the user pushes <100 objects to a 1.5.0-rc0 server we unpack\nto loose; >= 100 we keep the pack file.  Perhaps this would help\nkernel.org.\n \n-- \nShawn.\n"},{"id":"31082","messageId":"45A1668D.2010203@zytor.com","threadId":"6261","inReplyTo":"20070107201146.GA21956@suse.de","subject":"Re: [KORG] Re: kernel.org lies about latest -mm kernel","fromName":"H. Peter Anvin","fromEmail":"hpa@zytor.com","sentAt":"2007-01-07T21:30:53Z","receivedAt":"2007-01-07T21:30:53Z","isPatch":false,"sender":{"key":"hpa@zytor.com","avatar":null},"body":"Greg KH wrote:\n\n> \n> Well, I create my repos by doing a:\n> \tgit clone -l --bare\n> which makes a hardlink from Linus's tree.\n> \n> But then it gets copied over to the public server, which probably severs\n> that hardlink :(\n> \n> Any shortcut to clone or set up a repo using \"alternatives\" so that we\n> don't have this issue at all?\n> \n\nUse the -s option to git clone.\n\n\t-hpa\n"},{"id":"31083","messageId":"7v4pr21p0o.fsf@assigned-by-dhcp.cox.net","threadId":"6261","inReplyTo":"20070107201146.GA21956@suse.de","subject":"Re: [KORG] Re: kernel.org lies about latest -mm kernel","fromName":"Junio C Hamano","fromEmail":"junkio@cox.net","sentAt":"2007-01-07T21:54:31Z","receivedAt":"2007-01-07T21:54:31Z","isPatch":false,"sender":{"key":"gitster@pobox.com","avatar":"https://avatars.githubusercontent.com/u/54884?v=4"},"body":"Greg KH <gregkh@suse.de> writes:\n\n> Any shortcut to clone or set up a repo using \"alternatives\" so that we\n> don't have this issue at all?\n\n\"clone -l -s\" has been there for quote a long time (since mid Aug\n2005).  Because -s implies -l since end of November 2005, you\nshould be able to say\n\n\tgit clone --bare -s ..../torvalds/linux-2.6.git stable-queue.git\n"},{"id":"31084","messageId":"45A1727B.7070302@garzik.org","threadId":"6261","inReplyTo":"7v4pr21p0o.fsf@assigned-by-dhcp.cox.net","subject":"Re: [KORG] Re: kernel.org lies about latest -mm kernel","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2007-01-07T22:21:47Z","receivedAt":"2007-01-07T22:21:47Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Junio C Hamano wrote:\n> Greg KH <gregkh@suse.de> writes:\n> \n>> Any shortcut to clone or set up a repo using \"alternatives\" so that we\n>> don't have this issue at all?\n> \n> \"clone -l -s\" has been there for quote a long time (since mid Aug\n> 2005).  Because -s implies -l since end of November 2005, you\n> should be able to say\n> \n> \tgit clone --bare -s ..../torvalds/linux-2.6.git stable-queue.git\n\nYes but what about existing trees?\n\nCan you add an alternatives file, then prune, and get the same result as \nif you had done a clone -s ?\n\n\tJeff\n"},{"id":"31085","messageId":"Pine.LNX.4.64.0701071452300.3661@woody.osdl.org","threadId":"6261","inReplyTo":"45A1727B.7070302@garzik.org","subject":"Re: [KORG] Re: kernel.org lies about latest -mm kernel","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-07T22:53:30Z","receivedAt":"2007-01-07T22:53:30Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Sun, 7 Jan 2007, Jeff Garzik wrote:\n> \n> Yes but what about existing trees?\n> \n> Can you add an alternatives file, then prune, and get the same result as if\n> you had done a clone -s ?\n\nYes. Also do\n\n\tgit repack -a -d -l\n\nwhere the \"-l\" flag is the magic (it says to repack only objects that \naren't already packed in the alternate repository)\n\n\t\tLinus\n"},{"id":"31086","messageId":"46a038f90701071532o6d55b92eu4c8399ed6149b2e@mail.gmail.com","threadId":"6261","inReplyTo":"Pine.LNX.4.64.0701071452300.3661@woody.osdl.org","subject":"Re: [KORG] Re: kernel.org lies about latest -mm kernel","fromName":"Martin Langhoff","fromEmail":"martin.langhoff@gmail.com","sentAt":"2007-01-07T23:32:26Z","receivedAt":"2007-01-07T23:32:26Z","isPatch":false,"sender":{"key":"martin.langhoff@gmail.com","avatar":"https://gravatar.com/avatar/1e3f311b6c4c15836501901ca58f8c0b0667246488084ba524d8bc9867e22fd9?d=mp&s=160"},"body":"On 1/8/07, Linus Torvalds <torvalds@osdl.org> wrote:\n> On Sun, 7 Jan 2007, Jeff Garzik wrote:\n> > Yes but what about existing trees?\n> > Can you add an alternatives file, then prune, and get the same result as if\n> > you had done a clone -s ?\n> Yes. Also do\n>         git repack -a -d -l\n>\n> where the \"-l\" flag is the magic (it says to repack only objects that\n> aren't already packed in the alternate repository)\n\nIf all kernel.org repos get git-repack -a -d -l and git-pack-refs,\ngitweb will see a significant speedup, as some up-to-date checks\nbecome extremely cheap.\n\ncheers\n\n\n\nmartin\n"},{"id":"31101","messageId":"ens836$jr2$1@sea.gmane.org","threadId":"6261","inReplyTo":"20070107145730.GB24706@localhost","subject":"Re: How git affects kernel.org performance","fromName":"Jakub Narebski","fromEmail":"jnareb@gmail.com","sentAt":"2007-01-08T01:51:42Z","receivedAt":"2007-01-08T01:51:42Z","isPatch":false,"sender":{"key":"jnareb@gmail.com","avatar":"https://avatars.githubusercontent.com/u/2706?v=4"},"body":"Robert Fitzsimons wrote:\n\n>> Some more data on how git affects kernel.org...\n> \n> I have a quick question about the gitweb configuration, does the\n> $projects_list config entry point to a directory or a file?\n\nIt can point to both. Usually it is either unset, and then we\ndo find over $projectroot, or it is a file (URI escaped path\nrelative to $projectroot, SPACE, and URI escaped owner of a project;\nyou can get the file clicking on TXT on projects_list page).\n\n-- \nJakub Narebski\nWarsaw, Poland\nShadeHawk on #git\n"},{"id":"31111","messageId":"20070108030555.GA7289@in.ibm.com","threadId":"6261","inReplyTo":"20070107011542.3496bc76.akpm@osdl.org","subject":"Re: How git affects kernel.org performance","fromName":"Suparna Bhattacharya","fromEmail":"suparna@in.ibm.com","sentAt":"2007-01-08T03:05:55Z","receivedAt":"2007-01-08T03:05:55Z","isPatch":false,"sender":{"key":"suparna@in.ibm.com","avatar":null},"body":"On Sun, Jan 07, 2007 at 01:15:42AM -0800, Andrew Morton wrote:\n> On Sun, 7 Jan 2007 09:55:26 +0100\n> Willy Tarreau <w@1wt.eu> wrote:\n> \n> > On Sat, Jan 06, 2007 at 09:39:42PM -0800, Linus Torvalds wrote:\n> > >\n> > >\n> > > On Sat, 6 Jan 2007, H. Peter Anvin wrote:\n> > > >\n> > > > During extremely high load, it appears that what slows kernel.org down more\n> > > > than anything else is the time that each individual getdents() call takes.\n> > > > When I've looked this I've observed times from 200 ms to almost 2 seconds!\n> > > > Since an unpacked *OR* unpruned git tree adds 256 directories to a cleanly\n> > > > packed tree, you can do the math yourself.\n> > >\n> > > \"getdents()\" is totally serialized by the inode semaphore. It's one of the\n> > > most expensive system calls in Linux, partly because of that, and partly\n> > > because it has to call all the way down into the filesystem in a way that\n> > > almost no other common system call has to (99% of all filesystem calls can\n> > > be handled basically at the VFS layer with generic caches - but not\n> > > getdents()).\n> > >\n> > > So if there are concurrent readdirs on the same directory, they get\n> > > serialized. If there is any file creation/deletion activity in the\n> > > directory, it serializes getdents().\n> > >\n> > > To make matters worse, I don't think it has any read-ahead at all when you\n> > > use hashed directory entries. So if you have cold-cache case, you'll read\n> > > every single block totally individually, and serialized. One block at a\n> > > time (I think the non-hashed case is likely also suspect, but that's a\n> > > separate issue)\n> > >\n> > > In other words, I'm not at all surprised it hits on filldir time.\n> > > Especially on ext3.\n> >\n> > At work, we had the same problem on a file server with ext3. We use rsync\n> > to make backups to a local IDE disk, and we noticed that getdents() took\n> > about the same time as Peter reports (0.2 to 2 seconds), especially in\n> > maildir directories. We tried many things to fix it with no result,\n> > including enabling dirindexes. Finally, we made a full backup, and switched\n> > over to XFS and the problem totally disappeared. So it seems that the\n> > filesystem matters a lot here when there are lots of entries in a\n> > directory, and that ext3 is not suitable for usages with thousands\n> > of entries in directories with millions of files on disk. I'm not\n> > certain it would be that easy to try other filesystems on kernel.org\n> > though :-/\n> >\n> \n> Yeah, slowly-growing directories will get splattered all over the disk.\n> \n> Possible short-term fixes would be to just allocate up to (say) eight\n> blocks when we grow a directory by one block.  Or teach the\n> directory-growth code to use ext3 reservations.\n> \n> Longer-term people are talking about things like on-disk rerservations.\n> But I expect directories are being forgotten about in all of that.\n\nBy on-disk reservations, do you mean persistent file preallocation ? (that\nis explicit preallocation of blocks to a given file) If so, you are\nright, we haven't really given any thought to the possibility of directories\nneeding that feature.\n\nRegards\nSuparna\n\n> \n> -\n> To unsubscribe from this list: send the line \"unsubscribe linux-ext4\" in\n> the body of a message to majordomo@vger.kernel.org\n> More majordomo info at  http://vger.kernel.org/majordomo-info.html\n\n-- \nSuparna Bhattacharya (suparna@in.ibm.com)\nLinux Technology Center\nIBM Software Lab, India\n"},{"id":"31135","messageId":"20070108125819.GA32756@thunk.org","threadId":"6261","inReplyTo":"20070108030555.GA7289@in.ibm.com","subject":"Re: How git affects kernel.org performance","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-01-08T12:58:19Z","receivedAt":"2007-01-08T12:58:19Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jan 08, 2007 at 08:35:55AM +0530, Suparna Bhattacharya wrote:\n> > Yeah, slowly-growing directories will get splattered all over the disk.\n> > \n> > Possible short-term fixes would be to just allocate up to (say) eight\n> > blocks when we grow a directory by one block.  Or teach the\n> > directory-growth code to use ext3 reservations.\n> > \n> > Longer-term people are talking about things like on-disk rerservations.\n> > But I expect directories are being forgotten about in all of that.\n> \n> By on-disk reservations, do you mean persistent file preallocation ? (that\n> is explicit preallocation of blocks to a given file) If so, you are\n> right, we haven't really given any thought to the possibility of directories\n> needing that feature.\n\nThe fastest and probably most important thing to add is some readahead\nsmarts to directories --- both to the htree and non-htree cases.  If\nyou're using some kind of b-tree structure, such as XFS does for\ndirectories, preallocation doesn't help you much.  Delayed allocation\ncan save you if your delayed allocator knows how to structure disk\nblocks so that a btree-traversal is efficient, but I'm guessing the\nbiggest reason why we are losing is because we don't have sufficient\nreadahead.  This also has the advantage that it will help without\nneeding to doing a backup/restore to improve layout.\n\nAllocating some number of empty blocks when we grow the directory\nwould be a quick hack that I'd probably do as a 2nd priority.  It\nwon't help pre-existing directories, but combined with readahead\nlogic, should help us out greatly in the non-btree case.  \n\n\t\t\t\t\t\t- Ted\n"},{"id":"31145","messageId":"20070108134147.GB5291@linuxtv.org","threadId":"6261","inReplyTo":"20070108125819.GA32756@thunk.org","subject":"Re: How git affects kernel.org performance","fromName":"Johannes Stezenbach","fromEmail":"js@linuxtv.org","sentAt":"2007-01-08T13:41:47Z","receivedAt":"2007-01-08T13:41:47Z","isPatch":false,"sender":{"key":"js@linuxtv.org","avatar":null},"body":"On Mon, Jan 08, 2007 at 07:58:19AM -0500, Theodore Tso wrote:\n> \n> The fastest and probably most important thing to add is some readahead\n> smarts to directories --- both to the htree and non-htree cases.  If\n> you're using some kind of b-tree structure, such as XFS does for\n> directories, preallocation doesn't help you much.  Delayed allocation\n> can save you if your delayed allocator knows how to structure disk\n> blocks so that a btree-traversal is efficient, but I'm guessing the\n> biggest reason why we are losing is because we don't have sufficient\n> readahead.  This also has the advantage that it will help without\n> needing to doing a backup/restore to improve layout.\n\nWould e2fsck -D help? What kind of optimization\ndoes it perform?\n\n\nThanks,\nJohannes\n"},{"id":"31140","messageId":"45A24A65.1070706@garzik.org","threadId":"6261","inReplyTo":"20070108125819.GA32756@thunk.org","subject":"Re: How git affects kernel.org performance","fromName":"Jeff Garzik","fromEmail":"jeff@garzik.org","sentAt":"2007-01-08T13:43:01Z","receivedAt":"2007-01-08T13:43:01Z","isPatch":false,"sender":{"key":"jeff@garzik.org","avatar":null},"body":"Theodore Tso wrote:\n> The fastest and probably most important thing to add is some readahead\n> smarts to directories --- both to the htree and non-htree cases.  If\n> you're using some kind of b-tree structure, such as XFS does for\n> directories, preallocation doesn't help you much.  Delayed allocation\n> can save you if your delayed allocator knows how to structure disk\n> blocks so that a btree-traversal is efficient, but I'm guessing the\n> biggest reason why we are losing is because we don't have sufficient\n> readahead.  This also has the advantage that it will help without\n> needing to doing a backup/restore to improve layout.\n\n\nSomething I just thought of:  ATA and SCSI hard disks do their own \nread-ahead.  Seeking all over the place to pick up bits of directory \nwill hurt even more with the disk reading and throwing away data (albeit \nin its internal elevator and cache).\n\n\tJeff\n"},{"id":"31142","messageId":"20070108135622.GD32756@thunk.org","threadId":"6261","inReplyTo":"20070108134147.GB5291@linuxtv.org","subject":"Re: How git affects kernel.org performance","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-01-08T13:56:22Z","receivedAt":"2007-01-08T13:56:22Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jan 08, 2007 at 02:41:47PM +0100, Johannes Stezenbach wrote:\n> \n> Would e2fsck -D help? What kind of optimization\n> does it perform?\n\nIt will help a little; e2fsck -D compresses the logical view of the\ndirectory, but it doesn't optimize the physical layout on disk at all,\nand of course, it won't help with the lack of readahead logic.  It's\npossible to improve how e2fsck -D works, at the moment, it's not\ntrying to make the directory be contiguous on disk.  What it should\nprobably do is to pull a list of all of the blocks used by the\ndirectory, sort them, and then try to see if it can improve on the\nlist by allocating some new blocks that would make the directory more\ncontiguous on disk.  I suspect any improvements that would be seen by\ndoing this would be second order effects at most, though.\n\n\t\t\t\t\t\t- Ted\n"},{"id":"31143","messageId":"20070108135952.GF25857@elf.ucw.cz","threadId":"6261","inReplyTo":"20070108135622.GD32756@thunk.org","subject":"Re: How git affects kernel.org performance","fromName":"Pavel Machek","fromEmail":"pavel@ucw.cz","sentAt":"2007-01-08T13:59:52Z","receivedAt":"2007-01-08T13:59:52Z","isPatch":false,"sender":{"key":"pavel@ucw.cz","avatar":null},"body":"Hi!\n\n> > Would e2fsck -D help? What kind of optimization\n> > does it perform?\n> \n> It will help a little; e2fsck -D compresses the logical view of the\n> directory, but it doesn't optimize the physical layout on disk at all,\n> and of course, it won't help with the lack of readahead logic.  It's\n> possible to improve how e2fsck -D works, at the moment, it's not\n> trying to make the directory be contiguous on disk.  What it should\n> probably do is to pull a list of all of the blocks used by the\n> directory, sort them, and then try to see if it can improve on the\n> list by allocating some new blocks that would make the directory more\n> contiguous on disk.  I suspect any improvements that would be seen by\n> doing this would be second order effects at most, though.\n\n...sounds like a job for e2defrag, not e2fsck...\n\t\t\t\t\t\t\t\t\tPavel\n-- \n(english) http://www.livejournal.com/~pavelmachek\n(cesky, pictures) http://atrey.karlin.mff.cuni.cz/~pavel/picture/horses/blog.html\n"},{"id":"31146","messageId":"20070108141755.GF32756@thunk.org","threadId":"6261","inReplyTo":"20070108135952.GF25857@elf.ucw.cz","subject":"Re: How git affects kernel.org performance","fromName":"Theodore Tso","fromEmail":"tytso@mit.edu","sentAt":"2007-01-08T14:17:55Z","receivedAt":"2007-01-08T14:17:55Z","isPatch":false,"sender":{"key":"tytso@mit.edu","avatar":"https://avatars.githubusercontent.com/u/51416?v=4"},"body":"On Mon, Jan 08, 2007 at 02:59:52PM +0100, Pavel Machek wrote:\n> Hi!\n> \n> > > Would e2fsck -D help? What kind of optimization\n> > > does it perform?\n> > \n> > It will help a little; e2fsck -D compresses the logical view of the\n> > directory, but it doesn't optimize the physical layout on disk at all,\n> > and of course, it won't help with the lack of readahead logic.  It's\n> > possible to improve how e2fsck -D works, at the moment, it's not\n> > trying to make the directory be contiguous on disk.  What it should\n> > probably do is to pull a list of all of the blocks used by the\n> > directory, sort them, and then try to see if it can improve on the\n> > list by allocating some new blocks that would make the directory more\n> > contiguous on disk.  I suspect any improvements that would be seen by\n> > doing this would be second order effects at most, though.\n> \n> ...sounds like a job for e2defrag, not e2fsck...\n\nI wasn't proposing to move other data blocks around in order make the\ndirectory be contiguous, but just a \"quick and dirty\" try to make\nthings better.  But yes, in order to really fix layout issues you\nwould have to do a full defrag, and it's probably more important that\nwe try to fix things so that defragmentation runs aren't necessary in\nthe first place....\n\n\t\t\t\t\t\t- Ted\n"},{"id":"31148","messageId":"Pine.LNX.4.64.0701080943191.4964@xanadu.home","threadId":"6261","inReplyTo":"20070107203120.GA4970@spearce.org","subject":"Re: How git affects kernel.org performance","fromName":"Nicolas Pitre","fromEmail":"nico@cam.org","sentAt":"2007-01-08T14:46:41Z","receivedAt":"2007-01-08T14:46:41Z","isPatch":false,"sender":{"key":"nico@fluxnic.net","avatar":"https://avatars.githubusercontent.com/u/702790?v=4"},"body":"On Sun, 7 Jan 2007, Shawn O. Pearce wrote:\n\n> Krzysztof Halasa <khc@pm.waw.pl> wrote:\n> > Hmm... Perhaps it should be possible to push git updates as a pack\n> > file only? I mean, the pack file would stay packed = never individual\n> > files and never 256 directories?\n> \n> Latest Git does this.  If the server is later than 1.4.3.3 then\n> the receive-pack process can actually store the pack file rather\n> than unpacking it into loose objects.  The downside is that it will\n> copy any missing base objects onto the end of a thin pack to make\n> it not-thin.\n\nNo.  There are no thin packs for pushes.  And IMHO it should stay that \nway exactly to avoid this little inconvenience on servers.\n\nThe fetch case is a different story of course.\n\n\nNicolas\n"},{"id":"31175","messageId":"20070108170934.dafc5b81.pj@sgi.com","threadId":"6261","inReplyTo":"45A24A65.1070706@garzik.org","subject":"Re: How git affects kernel.org performance","fromName":"Paul Jackson","fromEmail":"pj@sgi.com","sentAt":"2007-01-09T01:09:34Z","receivedAt":"2007-01-09T01:09:34Z","isPatch":false,"sender":{"key":"pj@sgi.com","avatar":null},"body":"Jeff wrote:\n> Something I just thought of:  ATA and SCSI hard disks do their own\n> read-ahead.\n\nProbably this is wishful thinking on my part, but I would have hoped\nthat most of the read-ahead they did was for stuff that happened to be\non the cylinder they were reading anyway.  So long as their read-ahead\ndoesn't cause much extra or delayed disk head motion, what does it\nmatter?\n\n-- \n                  I won't rest till it's the best ...\n                  Programmer, Linux Scalability\n                  Paul Jackson <pj@sgi.com> 1.925.600.0401\n"},{"id":"31179","messageId":"20070109021812.GE44262@sgi.com","threadId":"6261","inReplyTo":"20070108170934.dafc5b81.pj@sgi.com","subject":"Re: How git affects kernel.org performance","fromName":"Jeremy Higdon","fromEmail":"jeremy@sgi.com","sentAt":"2007-01-09T02:18:12Z","receivedAt":"2007-01-09T02:18:12Z","isPatch":false,"sender":{"key":"jeremy@sgi.com","avatar":null},"body":"On Mon, Jan 08, 2007 at 05:09:34PM -0800, Paul Jackson wrote:\n> Jeff wrote:\n> > Something I just thought of:  ATA and SCSI hard disks do their own\n> > read-ahead.\n> \n> Probably this is wishful thinking on my part, but I would have hoped\n> that most of the read-ahead they did was for stuff that happened to be\n> on the cylinder they were reading anyway.  So long as their read-ahead\n> doesn't cause much extra or delayed disk head motion, what does it\n> matter?\n\n\nAnd they usually won't readahead if there is another command to\nprocess, though they can be set up to read unrequested data in\nspite of outstanding commands.\n\nWhen they are reading ahead, they'll only fetch LBAs beyond the last\nrequest until a buffer fills or the readahead gets interrupted.\n\njeremy\n"},{"id":"31206","messageId":"368329554.17014@ustc.edu.cn","threadId":"6261","inReplyTo":"20070108125819.GA32756@thunk.org","subject":"Re: How git affects kernel.org performance","fromName":"Fengguang Wu","fromEmail":"fengguang.wu@gmail.com","sentAt":"2007-01-09T07:59:46Z","receivedAt":"2007-01-09T07:59:46Z","isPatch":false,"sender":{"key":"fengguang.wu@gmail.com","avatar":null},"body":"On Mon, Jan 08, 2007 at 07:58:19AM -0500, Theodore Tso wrote:\n> On Mon, Jan 08, 2007 at 08:35:55AM +0530, Suparna Bhattacharya wrote:\n> > > Yeah, slowly-growing directories will get splattered all over the disk.\n> > > \n> > > Possible short-term fixes would be to just allocate up to (say) eight\n> > > blocks when we grow a directory by one block.  Or teach the\n> > > directory-growth code to use ext3 reservations.\n> > > \n> > > Longer-term people are talking about things like on-disk rerservations.\n> > > But I expect directories are being forgotten about in all of that.\n> > \n> > By on-disk reservations, do you mean persistent file preallocation ? (that\n> > is explicit preallocation of blocks to a given file) If so, you are\n> > right, we haven't really given any thought to the possibility of directories\n> > needing that feature.\n> \n> The fastest and probably most important thing to add is some readahead\n> smarts to directories --- both to the htree and non-htree cases.  If\n\nHere's is a quick hack to practice the directory readahead idea.\nComments are welcome, it's a freshman's work :)\n\nRegards,\nWu\n---\n fs/ext3/dir.c   |   22 ++++++++++++++++++++++\n fs/ext3/inode.c |    2 +-\n 2 files changed, 23 insertions(+), 1 deletion(-)\n\n--- linux.orig/fs/ext3/dir.c\n+++ linux/fs/ext3/dir.c\n@@ -94,6 +94,25 @@ int ext3_check_dir_entry (const char * f\n \treturn error_msg == NULL ? 1 : 0;\n }\n \n+int ext3_get_block(struct inode *inode, sector_t iblock,\n+\t\t\tstruct buffer_head *bh_result, int create);\n+\n+static void ext3_dir_readahead(struct file * filp)\n+{\n+\tstruct inode *inode = filp->f_path.dentry->d_inode;\n+\tstruct address_space *mapping = inode->i_sb->s_bdev->bd_inode->i_mapping;\n+\tunsigned long sector;\n+\tunsigned long blk;\n+\tpgoff_t offset;\n+\n+\tfor (blk = 0; blk < inode->i_blocks; blk++) {\n+\t\tsector = blk << (inode->i_blkbits - 9);\n+\t\tsector = generic_block_bmap(inode->i_mapping, sector, ext3_get_block);\n+\t\toffset = sector >> (PAGE_CACHE_SHIFT - 9);\n+\t\tdo_page_cache_readahead(mapping, filp, offset, 1);\n+\t}\n+}\n+\n static int ext3_readdir(struct file * filp,\n \t\t\t void * dirent, filldir_t filldir)\n {\n@@ -108,6 +127,9 @@ static int ext3_readdir(struct file * fi\n \n \tsb = inode->i_sb;\n \n+\tif (!filp->f_pos)\n+\t\text3_dir_readahead(filp);\n+\n #ifdef CONFIG_EXT3_INDEX\n \tif (EXT3_HAS_COMPAT_FEATURE(inode->i_sb,\n \t\t\t\t    EXT3_FEATURE_COMPAT_DIR_INDEX) &&\n--- linux.orig/fs/ext3/inode.c\n+++ linux/fs/ext3/inode.c\n@@ -945,7 +945,7 @@ out:\n \n #define DIO_CREDITS (EXT3_RESERVE_TRANS_BLOCKS + 32)\n \n-static int ext3_get_block(struct inode *inode, sector_t iblock,\n+int ext3_get_block(struct inode *inode, sector_t iblock,\n \t\t\tstruct buffer_head *bh_result, int create)\n {\n \thandle_t *handle = journal_current_handle();\n"},{"id":"31207","messageId":"20070109075945.GA8799__45841.3940541961$1168330087$gmane$org@mail.ustc.edu.cn","threadId":"6261","inReplyTo":"20070108125819.GA32756@thunk.org","subject":"Re: How git affects kernel.org performance","fromName":"Fengguang Wu","fromEmail":"fengguang.wu@gmail.com","sentAt":"2007-01-09T07:59:46Z","receivedAt":"2007-01-09T07:59:46Z","isPatch":false,"sender":{"key":"fengguang.wu@gmail.com","avatar":null},"body":"On Mon, Jan 08, 2007 at 07:58:19AM -0500, Theodore Tso wrote:\n> On Mon, Jan 08, 2007 at 08:35:55AM +0530, Suparna Bhattacharya wrote:\n> > > Yeah, slowly-growing directories will get splattered all over the disk.\n> > > \n> > > Possible short-term fixes would be to just allocate up to (say) eight\n> > > blocks when we grow a directory by one block.  Or teach the\n> > > directory-growth code to use ext3 reservations.\n> > > \n> > > Longer-term people are talking about things like on-disk rerservations.\n> > > But I expect directories are being forgotten about in all of that.\n> > \n> > By on-disk reservations, do you mean persistent file preallocation ? (that\n> > is explicit preallocation of blocks to a given file) If so, you are\n> > right, we haven't really given any thought to the possibility of directories\n> > needing that feature.\n> \n> The fastest and probably most important thing to add is some readahead\n> smarts to directories --- both to the htree and non-htree cases.  If\n\nHere's is a quick hack to practice the directory readahead idea.\nComments are welcome, it's a freshman's work :)\n\nRegards,\nWu\n---\n fs/ext3/dir.c   |   22 ++++++++++++++++++++++\n fs/ext3/inode.c |    2 +-\n 2 files changed, 23 insertions(+), 1 deletion(-)\n\n--- linux.orig/fs/ext3/dir.c\n+++ linux/fs/ext3/dir.c\n@@ -94,6 +94,25 @@ int ext3_check_dir_entry (const char * f\n \treturn error_msg == NULL ? 1 : 0;\n }\n \n+int ext3_get_block(struct inode *inode, sector_t iblock,\n+\t\t\tstruct buffer_head *bh_result, int create);\n+\n+static void ext3_dir_readahead(struct file * filp)\n+{\n+\tstruct inode *inode = filp->f_path.dentry->d_inode;\n+\tstruct address_space *mapping = inode->i_sb->s_bdev->bd_inode->i_mapping;\n+\tunsigned long sector;\n+\tunsigned long blk;\n+\tpgoff_t offset;\n+\n+\tfor (blk = 0; blk < inode->i_blocks; blk++) {\n+\t\tsector = blk << (inode->i_blkbits - 9);\n+\t\tsector = generic_block_bmap(inode->i_mapping, sector, ext3_get_block);\n+\t\toffset = sector >> (PAGE_CACHE_SHIFT - 9);\n+\t\tdo_page_cache_readahead(mapping, filp, offset, 1);\n+\t}\n+}\n+\n static int ext3_readdir(struct file * filp,\n \t\t\t void * dirent, filldir_t filldir)\n {\n@@ -108,6 +127,9 @@ static int ext3_readdir(struct file * fi\n \n \tsb = inode->i_sb;\n \n+\tif (!filp->f_pos)\n+\t\text3_dir_readahead(filp);\n+\n #ifdef CONFIG_EXT3_INDEX\n \tif (EXT3_HAS_COMPAT_FEATURE(inode->i_sb,\n \t\t\t\t    EXT3_FEATURE_COMPAT_DIR_INDEX) &&\n--- linux.orig/fs/ext3/inode.c\n+++ linux/fs/ext3/inode.c\n@@ -945,7 +945,7 @@ out:\n \n #define DIO_CREDITS (EXT3_RESERVE_TRANS_BLOCKS + 32)\n \n-static int ext3_get_block(struct inode *inode, sector_t iblock,\n+int ext3_get_block(struct inode *inode, sector_t iblock,\n \t\t\tstruct buffer_head *bh_result, int create)\n {\n \thandle_t *handle = journal_current_handle();\n"},{"id":"31208","messageId":"20070109075945.GA8799__43201.4260594316$1168330087$gmane$org@mail.ustc.edu.cn","threadId":"6261","inReplyTo":"20070108125819.GA32756@thunk.org","subject":"Re: How git affects kernel.org performance","fromName":"Fengguang Wu","fromEmail":"fengguang.wu@gmail.com","sentAt":"2007-01-09T07:59:46Z","receivedAt":"2007-01-09T07:59:46Z","isPatch":false,"sender":{"key":"fengguang.wu@gmail.com","avatar":null},"body":"On Mon, Jan 08, 2007 at 07:58:19AM -0500, Theodore Tso wrote:\n> On Mon, Jan 08, 2007 at 08:35:55AM +0530, Suparna Bhattacharya wrote:\n> > > Yeah, slowly-growing directories will get splattered all over the disk.\n> > > \n> > > Possible short-term fixes would be to just allocate up to (say) eight\n> > > blocks when we grow a directory by one block.  Or teach the\n> > > directory-growth code to use ext3 reservations.\n> > > \n> > > Longer-term people are talking about things like on-disk rerservations.\n> > > But I expect directories are being forgotten about in all of that.\n> > \n> > By on-disk reservations, do you mean persistent file preallocation ? (that\n> > is explicit preallocation of blocks to a given file) If so, you are\n> > right, we haven't really given any thought to the possibility of directories\n> > needing that feature.\n> \n> The fastest and probably most important thing to add is some readahead\n> smarts to directories --- both to the htree and non-htree cases.  If\n\nHere's is a quick hack to practice the directory readahead idea.\nComments are welcome, it's a freshman's work :)\n\nRegards,\nWu\n---\n fs/ext3/dir.c   |   22 ++++++++++++++++++++++\n fs/ext3/inode.c |    2 +-\n 2 files changed, 23 insertions(+), 1 deletion(-)\n\n--- linux.orig/fs/ext3/dir.c\n+++ linux/fs/ext3/dir.c\n@@ -94,6 +94,25 @@ int ext3_check_dir_entry (const char * f\n \treturn error_msg == NULL ? 1 : 0;\n }\n \n+int ext3_get_block(struct inode *inode, sector_t iblock,\n+\t\t\tstruct buffer_head *bh_result, int create);\n+\n+static void ext3_dir_readahead(struct file * filp)\n+{\n+\tstruct inode *inode = filp->f_path.dentry->d_inode;\n+\tstruct address_space *mapping = inode->i_sb->s_bdev->bd_inode->i_mapping;\n+\tunsigned long sector;\n+\tunsigned long blk;\n+\tpgoff_t offset;\n+\n+\tfor (blk = 0; blk < inode->i_blocks; blk++) {\n+\t\tsector = blk << (inode->i_blkbits - 9);\n+\t\tsector = generic_block_bmap(inode->i_mapping, sector, ext3_get_block);\n+\t\toffset = sector >> (PAGE_CACHE_SHIFT - 9);\n+\t\tdo_page_cache_readahead(mapping, filp, offset, 1);\n+\t}\n+}\n+\n static int ext3_readdir(struct file * filp,\n \t\t\t void * dirent, filldir_t filldir)\n {\n@@ -108,6 +127,9 @@ static int ext3_readdir(struct file * fi\n \n \tsb = inode->i_sb;\n \n+\tif (!filp->f_pos)\n+\t\text3_dir_readahead(filp);\n+\n #ifdef CONFIG_EXT3_INDEX\n \tif (EXT3_HAS_COMPAT_FEATURE(inode->i_sb,\n \t\t\t\t    EXT3_FEATURE_COMPAT_DIR_INDEX) &&\n--- linux.orig/fs/ext3/inode.c\n+++ linux/fs/ext3/inode.c\n@@ -945,7 +945,7 @@ out:\n \n #define DIO_CREDITS (EXT3_RESERVE_TRANS_BLOCKS + 32)\n \n-static int ext3_get_block(struct inode *inode, sector_t iblock,\n+int ext3_get_block(struct inode *inode, sector_t iblock,\n \t\t\tstruct buffer_head *bh_result, int create)\n {\n \thandle_t *handle = journal_current_handle();\n"},{"id":"31257","messageId":"Pine.LNX.4.64.0701090821550.3661@woody.osdl.org","threadId":"6261","inReplyTo":"368329554.17014@ustc.edu.cn","subject":"Re: How git affects kernel.org performance","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2007-01-09T16:23:32Z","receivedAt":"2007-01-09T16:23:32Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Tue, 9 Jan 2007, Fengguang Wu wrote:\n> > \n> > The fastest and probably most important thing to add is some readahead\n> > smarts to directories --- both to the htree and non-htree cases.  If\n> \n> Here's is a quick hack to practice the directory readahead idea.\n> Comments are welcome, it's a freshman's work :)\n\nWell, I'd probably have done it differently, but more important is whether \nthis actually makes a difference performance-wise. Have you benchmarked it \nat all?\n\nDoing an\n\n\techo 3 > /proc/sys/vm/drop_caches\n\nis your friend for testing things like this, to force cold-cache \nbehaviour..\n\n\t\tLinus\n"},{"id":"31350","messageId":"368394226.19365@ustc.edu.cn","threadId":"6261","inReplyTo":"Pine.LNX.4.64.0701090821550.3661@woody.osdl.org","subject":"Re: How git affects kernel.org performance","fromName":"Fengguang Wu","fromEmail":"fengguang.wu@gmail.com","sentAt":"2007-01-10T01:57:39Z","receivedAt":"2007-01-10T01:57:39Z","isPatch":false,"sender":{"key":"fengguang.wu@gmail.com","avatar":null},"body":"On Tue, Jan 09, 2007 at 08:23:32AM -0800, Linus Torvalds wrote:\n>\n>\n> On Tue, 9 Jan 2007, Fengguang Wu wrote:\n> > >\n> > > The fastest and probably most important thing to add is some readahead\n> > > smarts to directories --- both to the htree and non-htree cases.  If\n> >\n> > Here's is a quick hack to practice the directory readahead idea.\n> > Comments are welcome, it's a freshman's work :)\n>\n> Well, I'd probably have done it differently, but more important is whether\n> this actually makes a difference performance-wise. Have you benchmarked it\n> at all?\n\nYes, a trivial test shows a marginal improvement, on a minimal debian system:\n\n# find / | wc -l\n13641\n\n# time find / > /dev/null\n\nreal    0m10.000s\nuser    0m0.210s\nsys     0m4.370s\n\n# time find / > /dev/null\n\nreal    0m9.890s\nuser    0m0.160s\nsys     0m3.270s\n\n> Doing an\n>\n> \techo 3 > /proc/sys/vm/drop_caches\n>\n> is your friend for testing things like this, to force cold-cache\n> behaviour..\n\nThanks, I'll work out numbers on large/concurrent dir accesses soon.\n\nRegards,\nWu\n"},{"id":"31351","messageId":"20070110015739.GA26978__35540.4359674596$1168394253$gmane$org@mail.ustc.edu.cn","threadId":"6261","inReplyTo":"Pine.LNX.4.64.0701090821550.3661@woody.osdl.org","subject":"Re: How git affects kernel.org performance","fromName":"Fengguang Wu","fromEmail":"fengguang.wu@gmail.com","sentAt":"2007-01-10T01:57:39Z","receivedAt":"2007-01-10T01:57:39Z","isPatch":false,"sender":{"key":"fengguang.wu@gmail.com","avatar":null},"body":"On Tue, Jan 09, 2007 at 08:23:32AM -0800, Linus Torvalds wrote:\n>\n>\n> On Tue, 9 Jan 2007, Fengguang Wu wrote:\n> > >\n> > > The fastest and probably most important thing to add is some readahead\n> > > smarts to directories --- both to the htree and non-htree cases.  If\n> >\n> > Here's is a quick hack to practice the directory readahead idea.\n> > Comments are welcome, it's a freshman's work :)\n>\n> Well, I'd probably have done it differently, but more important is whether\n> this actually makes a difference performance-wise. Have you benchmarked it\n> at all?\n\nYes, a trivial test shows a marginal improvement, on a minimal debian system:\n\n# find / | wc -l\n13641\n\n# time find / > /dev/null\n\nreal    0m10.000s\nuser    0m0.210s\nsys     0m4.370s\n\n# time find / > /dev/null\n\nreal    0m9.890s\nuser    0m0.160s\nsys     0m3.270s\n\n> Doing an\n>\n> \techo 3 > /proc/sys/vm/drop_caches\n>\n> is your friend for testing things like this, to force cold-cache\n> behaviour..\n\nThanks, I'll work out numbers on large/concurrent dir accesses soon.\n\nRegards,\nWu\n"},{"id":"31352","messageId":"20070110015739.GA26978__11276.4408383102$1168394293$gmane$org@mail.ustc.edu.cn","threadId":"6261","inReplyTo":"Pine.LNX.4.64.0701090821550.3661@woody.osdl.org","subject":"Re: How git affects kernel.org performance","fromName":"Fengguang Wu","fromEmail":"fengguang.wu@gmail.com","sentAt":"2007-01-10T01:57:39Z","receivedAt":"2007-01-10T01:57:39Z","isPatch":false,"sender":{"key":"fengguang.wu@gmail.com","avatar":null},"body":"On Tue, Jan 09, 2007 at 08:23:32AM -0800, Linus Torvalds wrote:\n>\n>\n> On Tue, 9 Jan 2007, Fengguang Wu wrote:\n> > >\n> > > The fastest and probably most important thing to add is some readahead\n> > > smarts to directories --- both to the htree and non-htree cases.  If\n> >\n> > Here's is a quick hack to practice the directory readahead idea.\n> > Comments are welcome, it's a freshman's work :)\n>\n> Well, I'd probably have done it differently, but more important is whether\n> this actually makes a difference performance-wise. Have you benchmarked it\n> at all?\n\nYes, a trivial test shows a marginal improvement, on a minimal debian system:\n\n# find / | wc -l\n13641\n\n# time find / > /dev/null\n\nreal    0m10.000s\nuser    0m0.210s\nsys     0m4.370s\n\n# time find / > /dev/null\n\nreal    0m9.890s\nuser    0m0.160s\nsys     0m3.270s\n\n> Doing an\n>\n> \techo 3 > /proc/sys/vm/drop_caches\n>\n> is your friend for testing things like this, to force cold-cache\n> behaviour..\n\nThanks, I'll work out numbers on large/concurrent dir accesses soon.\n\nRegards,\nWu\n"},{"id":"31356","messageId":"1168399249.2585.6.camel@nigel.suspend2.net","threadId":"6261","inReplyTo":"20070110015739.GA26978@mail.ustc.edu.cn","subject":"Re: How git affects kernel.org performance","fromName":"Nigel Cunningham","fromEmail":"nigel@nigel.suspend2.net","sentAt":"2007-01-10T03:20:49Z","receivedAt":"2007-01-10T03:20:49Z","isPatch":false,"sender":{"key":"nigel@nigel.suspend2.net","avatar":null},"body":"Hi.\n\nOn Wed, 2007-01-10 at 09:57 +0800, Fengguang Wu wrote:\n> On Tue, Jan 09, 2007 at 08:23:32AM -0800, Linus Torvalds wrote:\n> >\n> >\n> > On Tue, 9 Jan 2007, Fengguang Wu wrote:\n> > > >\n> > > > The fastest and probably most important thing to add is some readahead\n> > > > smarts to directories --- both to the htree and non-htree cases.  If\n> > >\n> > > Here's is a quick hack to practice the directory readahead idea.\n> > > Comments are welcome, it's a freshman's work :)\n> >\n> > Well, I'd probably have done it differently, but more important is whether\n> > this actually makes a difference performance-wise. Have you benchmarked it\n> > at all?\n> \n> Yes, a trivial test shows a marginal improvement, on a minimal debian system:\n> \n> # find / | wc -l\n> 13641\n> \n> # time find / > /dev/null\n> \n> real    0m10.000s\n> user    0m0.210s\n> sys     0m4.370s\n> \n> # time find / > /dev/null\n> \n> real    0m9.890s\n> user    0m0.160s\n> sys     0m3.270s\n> \n> > Doing an\n> >\n> > \techo 3 > /proc/sys/vm/drop_caches\n> >\n> > is your friend for testing things like this, to force cold-cache\n> > behaviour..\n> \n> Thanks, I'll work out numbers on large/concurrent dir accesses soon.\n\nI gave it a try, and I'm afraid the results weren't pretty.\n\nI did:\n\ntime find /usr/src | wc -l\n\non current git with (3 times) and without (5 times) the patch, and got\n\nwith:\nreal   54.306, 54.327, 53.742s\nusr    0.324, 0.284, 0.234s\nsys    2.432, 2.484, 2.592s\n\nwithout:\nreal   24.413, 24.616, 24.080s\nusr    0.208, 0.316, 0.312s\nsys:   2.496, 2.440, 2.540s\n\nSubsequent runs without dropping caches did give a significant\nimprovement in both cases (1.821/.188/1.632 is one result I wrote with\nthe patch applied).\n\nRegards,\n\nNigel\n"},{"id":"31395","messageId":"368438013.19600@ustc.edu.cn","threadId":"6261","inReplyTo":"1168399249.2585.6.camel@nigel.suspend2.net","subject":"Re: How git affects kernel.org performance","fromName":"Fengguang Wu","fromEmail":"fengguang.wu@gmail.com","sentAt":"2007-01-10T14:07:30Z","receivedAt":"2007-01-10T14:07:30Z","isPatch":false,"sender":{"key":"fengguang.wu@gmail.com","avatar":null},"body":"On Wed, Jan 10, 2007 at 02:20:49PM +1100, Nigel Cunningham wrote:\n> Hi.\n> \n> On Wed, 2007-01-10 at 09:57 +0800, Fengguang Wu wrote:\n> > On Tue, Jan 09, 2007 at 08:23:32AM -0800, Linus Torvalds wrote:\n> > >\n> > >\n> > > On Tue, 9 Jan 2007, Fengguang Wu wrote:\n> > > > >\n> > > > > The fastest and probably most important thing to add is some readahead\n> > > > > smarts to directories --- both to the htree and non-htree cases.  If\n> > > >\n> > > > Here's is a quick hack to practice the directory readahead idea.\n> > > > Comments are welcome, it's a freshman's work :)\n> > >\n> > > Well, I'd probably have done it differently, but more important is whether\n> > > this actually makes a difference performance-wise. Have you benchmarked it\n> > > at all?\n> > \n> > Yes, a trivial test shows a marginal improvement, on a minimal debian system:\n> > \n> > # find / | wc -l\n> > 13641\n> > \n> > # time find / > /dev/null\n> > \n> > real    0m10.000s\n> > user    0m0.210s\n> > sys     0m4.370s\n> > \n> > # time find / > /dev/null\n> > \n> > real    0m9.890s\n> > user    0m0.160s\n> > sys     0m3.270s\n> > \n> > > Doing an\n> > >\n> > > \techo 3 > /proc/sys/vm/drop_caches\n> > >\n> > > is your friend for testing things like this, to force cold-cache\n> > > behaviour..\n> > \n> > Thanks, I'll work out numbers on large/concurrent dir accesses soon.\n> \n> I gave it a try, and I'm afraid the results weren't pretty.\n> \n> I did:\n> \n> time find /usr/src | wc -l\n> \n> on current git with (3 times) and without (5 times) the patch, and got\n> \n> with:\n> real   54.306, 54.327, 53.742s\n> usr    0.324, 0.284, 0.234s\n> sys    2.432, 2.484, 2.592s\n> \n> without:\n> real   24.413, 24.616, 24.080s\n> usr    0.208, 0.316, 0.312s\n> sys:   2.496, 2.440, 2.540s\n> \n> Subsequent runs without dropping caches did give a significant\n> improvement in both cases (1.821/.188/1.632 is one result I wrote with\n> the patch applied).\n\nThanks, Nigel.\nBut I'm very sorry that the calculation in the patch was wrong.\n\nWould you give this new patch a run?\n\nIt produced pretty numbers here:\n\n#!/bin/zsh\n\nROOT=/mnt/mnt\nTIMEFMT=\"%E clock  %S kernel  %U user  %w+%c cs  %J\"\n\necho 3 > /proc/sys/vm/drop_caches\n\n# 49: enable dir readahead\n# 50: disable\necho ${1:-50} > /proc/sys/vm/readahead_ratio\n\n# time find $ROOT/a > /dev/null\n\ntime find /etch > /dev/null\n\n# time find $ROOT/a > /dev/null&\n# time grep -r asdf $ROOT/b > /dev/null&\n# time cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null&\n\nexit 0\n\n# collected results on a SATA disk:\n# ./test-parallel-dir-reada.sh 49\n4.18s clock  0.08s kernel  0.04s user  418+0 cs  find $ROOT/a > /dev/null\n4.09s clock  0.10s kernel  0.02s user  410+1 cs  find $ROOT/a > /dev/null\n\n# ./test-parallel-dir-reada.sh 50\n12.18s clock  0.15s kernel  0.07s user  1520+4 cs  find $ROOT/a > /dev/null\n11.99s clock  0.13s kernel  0.04s user  1558+6 cs  find $ROOT/a > /dev/null\n\n\n# ./test-parallel-dir-reada.sh 49\n4.01s clock  0.06s kernel  0.01s user  1567+2 cs  find /etch > /dev/null\n4.08s clock  0.07s kernel  0.00s user  1568+0 cs  find /etch > /dev/null\n\n# ./test-parallel-dir-reada.sh 50\n4.10s clock  0.09s kernel  0.01s user  1578+1 cs  find /etch > /dev/null\n4.19s clock  0.08s kernel  0.03s user  1578+0 cs  find /etch > /dev/null\n\n\n# ./test-parallel-dir-reada.sh 49\n7.73s clock  0.11s kernel  0.06s user  438+2 cs  find $ROOT/a > /dev/null\n18.92s clock  0.43s kernel  0.02s user  1246+13 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n32.91s clock  4.20s kernel  1.55s user  103564+51 cs  grep -r asdf $ROOT/b > /dev/null\n\n8.47s clock  0.10s kernel  0.02s user  442+4 cs  find $ROOT/a > /dev/null\n19.24s clock  0.53s kernel  0.03s user  1250+23 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n29.93s clock  4.18s kernel  1.61s user  100425+47 cs  grep -r asdf $ROOT/b > /dev/null\n\n# ./test-parallel-dir-reada.sh 50\n17.87s clock  0.57s kernel  0.02s user  1244+21 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n21.30s clock  0.08s kernel  0.05s user  1517+5 cs  find $ROOT/a > /dev/null\n49.68s clock  3.94s kernel  1.67s user  101520+57 cs  grep -r asdf $ROOT/b > /dev/null\n\n15.66s clock  0.51s kernel  0.00s user  1248+25 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n22.15s clock  0.15s kernel  0.04s user  1520+5 cs  find $ROOT/a > /dev/null\n46.14s clock  4.08s kernel  1.68s user  101517+63 cs  grep -r asdf $ROOT/b > /dev/null\n\nThanks,\nWu\n---\n\nSubject: ext3 readdir readahead\n\nDo readahead for ext3_readdir().\n\nReasons to be aggressive:\n- readdir() users are likely to traverse the whole directory,\n  so readahead miss is not a concern.\n- most dirs are small, so slow start is not good\n- the htree indexing introduces some randomness,\n  which can be helped by the aggressiveness.\n\nSo we do 128K sized readaheads, at twice the speed of reads.\n\nThe following actual readahead pages are collected for a dir with\n110000 entries:\n\t32 31 30 31 28 29 29 28 27 25 29 22 25 30 24 15 19\nThat means a readahead hit ratio of\n\t454/541 = 84%\n\nThe performance is marginally better for a minimal debian system:\n\tcommand:\tfind /\n\tbaseline:\t4.10s\t4.19s\n\tpatched:\t4.01s\t4.08s\n\nAnd considerably better for 100 directories, each with 1000 8K files:\n\tcommand:\tfind /throwaways\n\tbaseline:\t12.18s\t11.99s\n\tpatched:\t 4.18s\t 4.09s\n\nAnd also noticable better for parallel operations:\n\t\t\t\t\tbaseline\tpatched\n\tfind /throwaways &\t\t21.30s 22.15s    7.73s  8.47s \n\tgrep -r asdf /throwaways2 &     49.68s 46.14s   32.91s 29.93s \n\tcp /KNOPPIX_CD.iso /dev/null &  17.87s 15.66s   18.92s 19.24s \n\nSigned-off-by: Fengguang Wu <wfg@mail.ustc.edu.cn>\n---\n fs/ext3/dir.c           |   33 +++++++++++++++++++++++++++++++++\n fs/ext3/inode.c         |    2 +-\n include/linux/ext3_fs.h |    2 ++\n 3 files changed, 36 insertions(+), 1 deletion(-)\n\n--- linux.orig/fs/ext3/dir.c\n+++ linux/fs/ext3/dir.c\n@@ -94,6 +94,28 @@ int ext3_check_dir_entry (const char * f\n \treturn error_msg == NULL ? 1 : 0;\n }\n \n+#define DIR_READAHEAD_BYTES  (128*1024)\n+#define DIR_READAHEAD_PGMASK ((DIR_READAHEAD_BYTES >> PAGE_CACHE_SHIFT) - 1)\n+\n+static void ext3_dir_readahead(struct file * filp)\n+{\n+\tstruct inode *inode = filp->f_path.dentry->d_inode;\n+\tstruct address_space *mapping = inode->i_sb->s_bdev->bd_inode->i_mapping;\n+\tint bbits = inode->i_blkbits;\n+\tunsigned long blk, end;\n+\n+\tblk = filp->f_ra.prev_page << (PAGE_CACHE_SHIFT - bbits);\n+\tend = min(inode->i_blocks >> (bbits - 9),\n+\t\t  blk + (DIR_READAHEAD_BYTES >> bbits));\n+\n+\tfor (; blk < end; blk++) {\n+\t\tpgoff_t phy;\n+\t\tphy = generic_block_bmap(inode->i_mapping, blk, ext3_get_block)\n+\t\t\t\t>> (PAGE_CACHE_SHIFT - bbits);\n+\t\tdo_page_cache_readahead(mapping, filp, phy, 1);\n+\t}\n+}\n+\n static int ext3_readdir(struct file * filp,\n \t\t\t void * dirent, filldir_t filldir)\n {\n@@ -108,6 +130,17 @@ static int ext3_readdir(struct file * fi\n \n \tsb = inode->i_sb;\n \n+\t/*\n+\t * Reading-ahead at 2x the page fault rate, in hope of reducing\n+\t * readahead misses caused by the partially random htree order.\n+\t */\n+\tfilp->f_ra.prev_page += 2;\n+\tfilp->f_ra.prev_page &= ~1;\n+\n+\tif (!(filp->f_ra.prev_page & DIR_READAHEAD_PGMASK) &&\n+\t\tfilp->f_ra.prev_page < (inode->i_blocks >> (PAGE_CACHE_SHIFT-9)))\n+\t\text3_dir_readahead(filp);\n+\n #ifdef CONFIG_EXT3_INDEX\n \tif (EXT3_HAS_COMPAT_FEATURE(inode->i_sb,\n \t\t\t\t    EXT3_FEATURE_COMPAT_DIR_INDEX) &&\n--- linux.orig/fs/ext3/inode.c\n+++ linux/fs/ext3/inode.c\n@@ -945,7 +945,7 @@ out:\n \n #define DIO_CREDITS (EXT3_RESERVE_TRANS_BLOCKS + 32)\n \n-static int ext3_get_block(struct inode *inode, sector_t iblock,\n+int ext3_get_block(struct inode *inode, sector_t iblock,\n \t\t\tstruct buffer_head *bh_result, int create)\n {\n \thandle_t *handle = journal_current_handle();\n--- linux.orig/include/linux/ext3_fs.h\n+++ linux/include/linux/ext3_fs.h\n@@ -814,6 +814,8 @@ struct buffer_head * ext3_bread (handle_\n int ext3_get_blocks_handle(handle_t *handle, struct inode *inode,\n \tsector_t iblock, unsigned long maxblocks, struct buffer_head *bh_result,\n \tint create, int extend_disksize);\n+extern int ext3_get_block(struct inode *inode, sector_t iblock,\n+\t\t\tstruct buffer_head *bh_result, int create);\n \n extern void ext3_read_inode (struct inode *);\n extern int  ext3_write_inode (struct inode *, int);\n"},{"id":"31396","messageId":"20070110140730.GA986__4640.65993805907$1168438064$gmane$org@mail.ustc.edu.cn","threadId":"6261","inReplyTo":"1168399249.2585.6.camel@nigel.suspend2.net","subject":"Re: How git affects kernel.org performance","fromName":"Fengguang Wu","fromEmail":"fengguang.wu@gmail.com","sentAt":"2007-01-10T14:07:30Z","receivedAt":"2007-01-10T14:07:30Z","isPatch":false,"sender":{"key":"fengguang.wu@gmail.com","avatar":null},"body":"On Wed, Jan 10, 2007 at 02:20:49PM +1100, Nigel Cunningham wrote:\n> Hi.\n> \n> On Wed, 2007-01-10 at 09:57 +0800, Fengguang Wu wrote:\n> > On Tue, Jan 09, 2007 at 08:23:32AM -0800, Linus Torvalds wrote:\n> > >\n> > >\n> > > On Tue, 9 Jan 2007, Fengguang Wu wrote:\n> > > > >\n> > > > > The fastest and probably most important thing to add is some readahead\n> > > > > smarts to directories --- both to the htree and non-htree cases.  If\n> > > >\n> > > > Here's is a quick hack to practice the directory readahead idea.\n> > > > Comments are welcome, it's a freshman's work :)\n> > >\n> > > Well, I'd probably have done it differently, but more important is whether\n> > > this actually makes a difference performance-wise. Have you benchmarked it\n> > > at all?\n> > \n> > Yes, a trivial test shows a marginal improvement, on a minimal debian system:\n> > \n> > # find / | wc -l\n> > 13641\n> > \n> > # time find / > /dev/null\n> > \n> > real    0m10.000s\n> > user    0m0.210s\n> > sys     0m4.370s\n> > \n> > # time find / > /dev/null\n> > \n> > real    0m9.890s\n> > user    0m0.160s\n> > sys     0m3.270s\n> > \n> > > Doing an\n> > >\n> > > \techo 3 > /proc/sys/vm/drop_caches\n> > >\n> > > is your friend for testing things like this, to force cold-cache\n> > > behaviour..\n> > \n> > Thanks, I'll work out numbers on large/concurrent dir accesses soon.\n> \n> I gave it a try, and I'm afraid the results weren't pretty.\n> \n> I did:\n> \n> time find /usr/src | wc -l\n> \n> on current git with (3 times) and without (5 times) the patch, and got\n> \n> with:\n> real   54.306, 54.327, 53.742s\n> usr    0.324, 0.284, 0.234s\n> sys    2.432, 2.484, 2.592s\n> \n> without:\n> real   24.413, 24.616, 24.080s\n> usr    0.208, 0.316, 0.312s\n> sys:   2.496, 2.440, 2.540s\n> \n> Subsequent runs without dropping caches did give a significant\n> improvement in both cases (1.821/.188/1.632 is one result I wrote with\n> the patch applied).\n\nThanks, Nigel.\nBut I'm very sorry that the calculation in the patch was wrong.\n\nWould you give this new patch a run?\n\nIt produced pretty numbers here:\n\n#!/bin/zsh\n\nROOT=/mnt/mnt\nTIMEFMT=\"%E clock  %S kernel  %U user  %w+%c cs  %J\"\n\necho 3 > /proc/sys/vm/drop_caches\n\n# 49: enable dir readahead\n# 50: disable\necho ${1:-50} > /proc/sys/vm/readahead_ratio\n\n# time find $ROOT/a > /dev/null\n\ntime find /etch > /dev/null\n\n# time find $ROOT/a > /dev/null&\n# time grep -r asdf $ROOT/b > /dev/null&\n# time cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null&\n\nexit 0\n\n# collected results on a SATA disk:\n# ./test-parallel-dir-reada.sh 49\n4.18s clock  0.08s kernel  0.04s user  418+0 cs  find $ROOT/a > /dev/null\n4.09s clock  0.10s kernel  0.02s user  410+1 cs  find $ROOT/a > /dev/null\n\n# ./test-parallel-dir-reada.sh 50\n12.18s clock  0.15s kernel  0.07s user  1520+4 cs  find $ROOT/a > /dev/null\n11.99s clock  0.13s kernel  0.04s user  1558+6 cs  find $ROOT/a > /dev/null\n\n\n# ./test-parallel-dir-reada.sh 49\n4.01s clock  0.06s kernel  0.01s user  1567+2 cs  find /etch > /dev/null\n4.08s clock  0.07s kernel  0.00s user  1568+0 cs  find /etch > /dev/null\n\n# ./test-parallel-dir-reada.sh 50\n4.10s clock  0.09s kernel  0.01s user  1578+1 cs  find /etch > /dev/null\n4.19s clock  0.08s kernel  0.03s user  1578+0 cs  find /etch > /dev/null\n\n\n# ./test-parallel-dir-reada.sh 49\n7.73s clock  0.11s kernel  0.06s user  438+2 cs  find $ROOT/a > /dev/null\n18.92s clock  0.43s kernel  0.02s user  1246+13 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n32.91s clock  4.20s kernel  1.55s user  103564+51 cs  grep -r asdf $ROOT/b > /dev/null\n\n8.47s clock  0.10s kernel  0.02s user  442+4 cs  find $ROOT/a > /dev/null\n19.24s clock  0.53s kernel  0.03s user  1250+23 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n29.93s clock  4.18s kernel  1.61s user  100425+47 cs  grep -r asdf $ROOT/b > /dev/null\n\n# ./test-parallel-dir-reada.sh 50\n17.87s clock  0.57s kernel  0.02s user  1244+21 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n21.30s clock  0.08s kernel  0.05s user  1517+5 cs  find $ROOT/a > /dev/null\n49.68s clock  3.94s kernel  1.67s user  101520+57 cs  grep -r asdf $ROOT/b > /dev/null\n\n15.66s clock  0.51s kernel  0.00s user  1248+25 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n22.15s clock  0.15s kernel  0.04s user  1520+5 cs  find $ROOT/a > /dev/null\n46.14s clock  4.08s kernel  1.68s user  101517+63 cs  grep -r asdf $ROOT/b > /dev/null\n\nThanks,\nWu\n---\n\nSubject: ext3 readdir readahead\n\nDo readahead for ext3_readdir().\n\nReasons to be aggressive:\n- readdir() users are likely to traverse the whole directory,\n  so readahead miss is not a concern.\n- most dirs are small, so slow start is not good\n- the htree indexing introduces some randomness,\n  which can be helped by the aggressiveness.\n\nSo we do 128K sized readaheads, at twice the speed of reads.\n\nThe following actual readahead pages are collected for a dir with\n110000 entries:\n\t32 31 30 31 28 29 29 28 27 25 29 22 25 30 24 15 19\nThat means a readahead hit ratio of\n\t454/541 = 84%\n\nThe performance is marginally better for a minimal debian system:\n\tcommand:\tfind /\n\tbaseline:\t4.10s\t4.19s\n\tpatched:\t4.01s\t4.08s\n\nAnd considerably better for 100 directories, each with 1000 8K files:\n\tcommand:\tfind /throwaways\n\tbaseline:\t12.18s\t11.99s\n\tpatched:\t 4.18s\t 4.09s\n\nAnd also noticable better for parallel operations:\n\t\t\t\t\tbaseline\tpatched\n\tfind /throwaways &\t\t21.30s 22.15s    7.73s  8.47s \n\tgrep -r asdf /throwaways2 &     49.68s 46.14s   32.91s 29.93s \n\tcp /KNOPPIX_CD.iso /dev/null &  17.87s 15.66s   18.92s 19.24s \n\nSigned-off-by: Fengguang Wu <wfg@mail.ustc.edu.cn>\n---\n fs/ext3/dir.c           |   33 +++++++++++++++++++++++++++++++++\n fs/ext3/inode.c         |    2 +-\n include/linux/ext3_fs.h |    2 ++\n 3 files changed, 36 insertions(+), 1 deletion(-)\n\n--- linux.orig/fs/ext3/dir.c\n+++ linux/fs/ext3/dir.c\n@@ -94,6 +94,28 @@ int ext3_check_dir_entry (const char * f\n \treturn error_msg == NULL ? 1 : 0;\n }\n \n+#define DIR_READAHEAD_BYTES  (128*1024)\n+#define DIR_READAHEAD_PGMASK ((DIR_READAHEAD_BYTES >> PAGE_CACHE_SHIFT) - 1)\n+\n+static void ext3_dir_readahead(struct file * filp)\n+{\n+\tstruct inode *inode = filp->f_path.dentry->d_inode;\n+\tstruct address_space *mapping = inode->i_sb->s_bdev->bd_inode->i_mapping;\n+\tint bbits = inode->i_blkbits;\n+\tunsigned long blk, end;\n+\n+\tblk = filp->f_ra.prev_page << (PAGE_CACHE_SHIFT - bbits);\n+\tend = min(inode->i_blocks >> (bbits - 9),\n+\t\t  blk + (DIR_READAHEAD_BYTES >> bbits));\n+\n+\tfor (; blk < end; blk++) {\n+\t\tpgoff_t phy;\n+\t\tphy = generic_block_bmap(inode->i_mapping, blk, ext3_get_block)\n+\t\t\t\t>> (PAGE_CACHE_SHIFT - bbits);\n+\t\tdo_page_cache_readahead(mapping, filp, phy, 1);\n+\t}\n+}\n+\n static int ext3_readdir(struct file * filp,\n \t\t\t void * dirent, filldir_t filldir)\n {\n@@ -108,6 +130,17 @@ static int ext3_readdir(struct file * fi\n \n \tsb = inode->i_sb;\n \n+\t/*\n+\t * Reading-ahead at 2x the page fault rate, in hope of reducing\n+\t * readahead misses caused by the partially random htree order.\n+\t */\n+\tfilp->f_ra.prev_page += 2;\n+\tfilp->f_ra.prev_page &= ~1;\n+\n+\tif (!(filp->f_ra.prev_page & DIR_READAHEAD_PGMASK) &&\n+\t\tfilp->f_ra.prev_page < (inode->i_blocks >> (PAGE_CACHE_SHIFT-9)))\n+\t\text3_dir_readahead(filp);\n+\n #ifdef CONFIG_EXT3_INDEX\n \tif (EXT3_HAS_COMPAT_FEATURE(inode->i_sb,\n \t\t\t\t    EXT3_FEATURE_COMPAT_DIR_INDEX) &&\n--- linux.orig/fs/ext3/inode.c\n+++ linux/fs/ext3/inode.c\n@@ -945,7 +945,7 @@ out:\n \n #define DIO_CREDITS (EXT3_RESERVE_TRANS_BLOCKS + 32)\n \n-static int ext3_get_block(struct inode *inode, sector_t iblock,\n+int ext3_get_block(struct inode *inode, sector_t iblock,\n \t\t\tstruct buffer_head *bh_result, int create)\n {\n \thandle_t *handle = journal_current_handle();\n--- linux.orig/include/linux/ext3_fs.h\n+++ linux/include/linux/ext3_fs.h\n@@ -814,6 +814,8 @@ struct buffer_head * ext3_bread (handle_\n int ext3_get_blocks_handle(handle_t *handle, struct inode *inode,\n \tsector_t iblock, unsigned long maxblocks, struct buffer_head *bh_result,\n \tint create, int extend_disksize);\n+extern int ext3_get_block(struct inode *inode, sector_t iblock,\n+\t\t\tstruct buffer_head *bh_result, int create);\n \n extern void ext3_read_inode (struct inode *);\n extern int  ext3_write_inode (struct inode *, int);\n"},{"id":"31397","messageId":"20070110140730.GA986__11582.1861465976$1168438162$gmane$org@mail.ustc.edu.cn","threadId":"6261","inReplyTo":"1168399249.2585.6.camel@nigel.suspend2.net","subject":"Re: How git affects kernel.org performance","fromName":"Fengguang Wu","fromEmail":"fengguang.wu@gmail.com","sentAt":"2007-01-10T14:07:30Z","receivedAt":"2007-01-10T14:07:30Z","isPatch":false,"sender":{"key":"fengguang.wu@gmail.com","avatar":null},"body":"On Wed, Jan 10, 2007 at 02:20:49PM +1100, Nigel Cunningham wrote:\n> Hi.\n> \n> On Wed, 2007-01-10 at 09:57 +0800, Fengguang Wu wrote:\n> > On Tue, Jan 09, 2007 at 08:23:32AM -0800, Linus Torvalds wrote:\n> > >\n> > >\n> > > On Tue, 9 Jan 2007, Fengguang Wu wrote:\n> > > > >\n> > > > > The fastest and probably most important thing to add is some readahead\n> > > > > smarts to directories --- both to the htree and non-htree cases.  If\n> > > >\n> > > > Here's is a quick hack to practice the directory readahead idea.\n> > > > Comments are welcome, it's a freshman's work :)\n> > >\n> > > Well, I'd probably have done it differently, but more important is whether\n> > > this actually makes a difference performance-wise. Have you benchmarked it\n> > > at all?\n> > \n> > Yes, a trivial test shows a marginal improvement, on a minimal debian system:\n> > \n> > # find / | wc -l\n> > 13641\n> > \n> > # time find / > /dev/null\n> > \n> > real    0m10.000s\n> > user    0m0.210s\n> > sys     0m4.370s\n> > \n> > # time find / > /dev/null\n> > \n> > real    0m9.890s\n> > user    0m0.160s\n> > sys     0m3.270s\n> > \n> > > Doing an\n> > >\n> > > \techo 3 > /proc/sys/vm/drop_caches\n> > >\n> > > is your friend for testing things like this, to force cold-cache\n> > > behaviour..\n> > \n> > Thanks, I'll work out numbers on large/concurrent dir accesses soon.\n> \n> I gave it a try, and I'm afraid the results weren't pretty.\n> \n> I did:\n> \n> time find /usr/src | wc -l\n> \n> on current git with (3 times) and without (5 times) the patch, and got\n> \n> with:\n> real   54.306, 54.327, 53.742s\n> usr    0.324, 0.284, 0.234s\n> sys    2.432, 2.484, 2.592s\n> \n> without:\n> real   24.413, 24.616, 24.080s\n> usr    0.208, 0.316, 0.312s\n> sys:   2.496, 2.440, 2.540s\n> \n> Subsequent runs without dropping caches did give a significant\n> improvement in both cases (1.821/.188/1.632 is one result I wrote with\n> the patch applied).\n\nThanks, Nigel.\nBut I'm very sorry that the calculation in the patch was wrong.\n\nWould you give this new patch a run?\n\nIt produced pretty numbers here:\n\n#!/bin/zsh\n\nROOT=/mnt/mnt\nTIMEFMT=\"%E clock  %S kernel  %U user  %w+%c cs  %J\"\n\necho 3 > /proc/sys/vm/drop_caches\n\n# 49: enable dir readahead\n# 50: disable\necho ${1:-50} > /proc/sys/vm/readahead_ratio\n\n# time find $ROOT/a > /dev/null\n\ntime find /etch > /dev/null\n\n# time find $ROOT/a > /dev/null&\n# time grep -r asdf $ROOT/b > /dev/null&\n# time cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null&\n\nexit 0\n\n# collected results on a SATA disk:\n# ./test-parallel-dir-reada.sh 49\n4.18s clock  0.08s kernel  0.04s user  418+0 cs  find $ROOT/a > /dev/null\n4.09s clock  0.10s kernel  0.02s user  410+1 cs  find $ROOT/a > /dev/null\n\n# ./test-parallel-dir-reada.sh 50\n12.18s clock  0.15s kernel  0.07s user  1520+4 cs  find $ROOT/a > /dev/null\n11.99s clock  0.13s kernel  0.04s user  1558+6 cs  find $ROOT/a > /dev/null\n\n\n# ./test-parallel-dir-reada.sh 49\n4.01s clock  0.06s kernel  0.01s user  1567+2 cs  find /etch > /dev/null\n4.08s clock  0.07s kernel  0.00s user  1568+0 cs  find /etch > /dev/null\n\n# ./test-parallel-dir-reada.sh 50\n4.10s clock  0.09s kernel  0.01s user  1578+1 cs  find /etch > /dev/null\n4.19s clock  0.08s kernel  0.03s user  1578+0 cs  find /etch > /dev/null\n\n\n# ./test-parallel-dir-reada.sh 49\n7.73s clock  0.11s kernel  0.06s user  438+2 cs  find $ROOT/a > /dev/null\n18.92s clock  0.43s kernel  0.02s user  1246+13 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n32.91s clock  4.20s kernel  1.55s user  103564+51 cs  grep -r asdf $ROOT/b > /dev/null\n\n8.47s clock  0.10s kernel  0.02s user  442+4 cs  find $ROOT/a > /dev/null\n19.24s clock  0.53s kernel  0.03s user  1250+23 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n29.93s clock  4.18s kernel  1.61s user  100425+47 cs  grep -r asdf $ROOT/b > /dev/null\n\n# ./test-parallel-dir-reada.sh 50\n17.87s clock  0.57s kernel  0.02s user  1244+21 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n21.30s clock  0.08s kernel  0.05s user  1517+5 cs  find $ROOT/a > /dev/null\n49.68s clock  3.94s kernel  1.67s user  101520+57 cs  grep -r asdf $ROOT/b > /dev/null\n\n15.66s clock  0.51s kernel  0.00s user  1248+25 cs  cp /etch/KNOPPIX_V5.0.1CD-2006-06-01-EN.iso /dev/null\n22.15s clock  0.15s kernel  0.04s user  1520+5 cs  find $ROOT/a > /dev/null\n46.14s clock  4.08s kernel  1.68s user  101517+63 cs  grep -r asdf $ROOT/b > /dev/null\n\nThanks,\nWu\n---\n\nSubject: ext3 readdir readahead\n\nDo readahead for ext3_readdir().\n\nReasons to be aggressive:\n- readdir() users are likely to traverse the whole directory,\n  so readahead miss is not a concern.\n- most dirs are small, so slow start is not good\n- the htree indexing introduces some randomness,\n  which can be helped by the aggressiveness.\n\nSo we do 128K sized readaheads, at twice the speed of reads.\n\nThe following actual readahead pages are collected for a dir with\n110000 entries:\n\t32 31 30 31 28 29 29 28 27 25 29 22 25 30 24 15 19\nThat means a readahead hit ratio of\n\t454/541 = 84%\n\nThe performance is marginally better for a minimal debian system:\n\tcommand:\tfind /\n\tbaseline:\t4.10s\t4.19s\n\tpatched:\t4.01s\t4.08s\n\nAnd considerably better for 100 directories, each with 1000 8K files:\n\tcommand:\tfind /throwaways\n\tbaseline:\t12.18s\t11.99s\n\tpatched:\t 4.18s\t 4.09s\n\nAnd also noticable better for parallel operations:\n\t\t\t\t\tbaseline\tpatched\n\tfind /throwaways &\t\t21.30s 22.15s    7.73s  8.47s \n\tgrep -r asdf /throwaways2 &     49.68s 46.14s   32.91s 29.93s \n\tcp /KNOPPIX_CD.iso /dev/null &  17.87s 15.66s   18.92s 19.24s \n\nSigned-off-by: Fengguang Wu <wfg@mail.ustc.edu.cn>\n---\n fs/ext3/dir.c           |   33 +++++++++++++++++++++++++++++++++\n fs/ext3/inode.c         |    2 +-\n include/linux/ext3_fs.h |    2 ++\n 3 files changed, 36 insertions(+), 1 deletion(-)\n\n--- linux.orig/fs/ext3/dir.c\n+++ linux/fs/ext3/dir.c\n@@ -94,6 +94,28 @@ int ext3_check_dir_entry (const char * f\n \treturn error_msg == NULL ? 1 : 0;\n }\n \n+#define DIR_READAHEAD_BYTES  (128*1024)\n+#define DIR_READAHEAD_PGMASK ((DIR_READAHEAD_BYTES >> PAGE_CACHE_SHIFT) - 1)\n+\n+static void ext3_dir_readahead(struct file * filp)\n+{\n+\tstruct inode *inode = filp->f_path.dentry->d_inode;\n+\tstruct address_space *mapping = inode->i_sb->s_bdev->bd_inode->i_mapping;\n+\tint bbits = inode->i_blkbits;\n+\tunsigned long blk, end;\n+\n+\tblk = filp->f_ra.prev_page << (PAGE_CACHE_SHIFT - bbits);\n+\tend = min(inode->i_blocks >> (bbits - 9),\n+\t\t  blk + (DIR_READAHEAD_BYTES >> bbits));\n+\n+\tfor (; blk < end; blk++) {\n+\t\tpgoff_t phy;\n+\t\tphy = generic_block_bmap(inode->i_mapping, blk, ext3_get_block)\n+\t\t\t\t>> (PAGE_CACHE_SHIFT - bbits);\n+\t\tdo_page_cache_readahead(mapping, filp, phy, 1);\n+\t}\n+}\n+\n static int ext3_readdir(struct file * filp,\n \t\t\t void * dirent, filldir_t filldir)\n {\n@@ -108,6 +130,17 @@ static int ext3_readdir(struct file * fi\n \n \tsb = inode->i_sb;\n \n+\t/*\n+\t * Reading-ahead at 2x the page fault rate, in hope of reducing\n+\t * readahead misses caused by the partially random htree order.\n+\t */\n+\tfilp->f_ra.prev_page += 2;\n+\tfilp->f_ra.prev_page &= ~1;\n+\n+\tif (!(filp->f_ra.prev_page & DIR_READAHEAD_PGMASK) &&\n+\t\tfilp->f_ra.prev_page < (inode->i_blocks >> (PAGE_CACHE_SHIFT-9)))\n+\t\text3_dir_readahead(filp);\n+\n #ifdef CONFIG_EXT3_INDEX\n \tif (EXT3_HAS_COMPAT_FEATURE(inode->i_sb,\n \t\t\t\t    EXT3_FEATURE_COMPAT_DIR_INDEX) &&\n--- linux.orig/fs/ext3/inode.c\n+++ linux/fs/ext3/inode.c\n@@ -945,7 +945,7 @@ out:\n \n #define DIO_CREDITS (EXT3_RESERVE_TRANS_BLOCKS + 32)\n \n-static int ext3_get_block(struct inode *inode, sector_t iblock,\n+int ext3_get_block(struct inode *inode, sector_t iblock,\n \t\t\tstruct buffer_head *bh_result, int create)\n {\n \thandle_t *handle = journal_current_handle();\n--- linux.orig/include/linux/ext3_fs.h\n+++ linux/include/linux/ext3_fs.h\n@@ -814,6 +814,8 @@ struct buffer_head * ext3_bread (handle_\n int ext3_get_blocks_handle(handle_t *handle, struct inode *inode,\n \tsector_t iblock, unsigned long maxblocks, struct buffer_head *bh_result,\n \tint create, int extend_disksize);\n+extern int ext3_get_block(struct inode *inode, sector_t iblock,\n+\t\t\tstruct buffer_head *bh_result, int create);\n \n extern void ext3_read_inode (struct inode *);\n extern int  ext3_write_inode (struct inode *, int);\n"},{"id":"31549","messageId":"1168599283.2744.5.camel@nigel.suspend2.net","threadId":"6261","inReplyTo":"20070110140730.GA986@mail.ustc.edu.cn","subject":"Re: How git affects kernel.org performance","fromName":"Nigel Cunningham","fromEmail":"nigel@nigel.suspend2.net","sentAt":"2007-01-12T10:54:43Z","receivedAt":"2007-01-12T10:54:43Z","isPatch":false,"sender":{"key":"nigel@nigel.suspend2.net","avatar":null},"body":"Hi.\n\nOn Wed, 2007-01-10 at 22:07 +0800, Fengguang Wu wrote:\n> Thanks, Nigel.\n> But I'm very sorry that the calculation in the patch was wrong.\n> \n> Would you give this new patch a run?\n\nSorry for my slowness. I just did\n\ntime find /usr/src | wc -l\n\nagain:\n\nWithout patch: 35.137, 35.104, 35.351 seconds\nWith patch: 34.518, 34.376, 34.489 seconds\n\nSo there's about .8 seconds saved.\n\nRegards,\n\nNigel\n"}]}