git/list[1] front-page[2] threads[3] people[4] search[5] about
 

Re: How git affects kernel.org performance

From
SBSuparna Bhattacharya <suparna@in.ibm.com>
Date
Jan 8, 2007, 03:05 UTC
Message-ID
<20070108030555.GA7289@in.ibm.com>
In-Reply-To
<20070107011542.3496bc76.akpm@osdl.org>
On Sun, Jan 07, 2007 at 01:15:42AM -0800, Andrew Morton wrote:
Show 55 quoted lines
> On Sun, 7 Jan 2007 09:55:26 +0100
> Willy Tarreau <w@1wt.eu> wrote:
> 
> > On Sat, Jan 06, 2007 at 09:39:42PM -0800, Linus Torvalds wrote:
> > >
> > >
> > > On Sat, 6 Jan 2007, H. Peter Anvin wrote:
> > > >
> > > > During extremely high load, it appears that what slows kernel.org down more
> > > > than anything else is the time that each individual getdents() call takes.
> > > > When I've looked this I've observed times from 200 ms to almost 2 seconds!
> > > > Since an unpacked *OR* unpruned git tree adds 256 directories to a cleanly
> > > > packed tree, you can do the math yourself.
> > >
> > > "getdents()" is totally serialized by the inode semaphore. It's one of the
> > > most expensive system calls in Linux, partly because of that, and partly
> > > because it has to call all the way down into the filesystem in a way that
> > > almost no other common system call has to (99% of all filesystem calls can
> > > be handled basically at the VFS layer with generic caches - but not
> > > getdents()).
> > >
> > > So if there are concurrent readdirs on the same directory, they get
> > > serialized. If there is any file creation/deletion activity in the
> > > directory, it serializes getdents().
> > >
> > > To make matters worse, I don't think it has any read-ahead at all when you
> > > use hashed directory entries. So if you have cold-cache case, you'll read
> > > every single block totally individually, and serialized. One block at a
> > > time (I think the non-hashed case is likely also suspect, but that's a
> > > separate issue)
> > >
> > > In other words, I'm not at all surprised it hits on filldir time.
> > > Especially on ext3.
> >
> > At work, we had the same problem on a file server with ext3. We use rsync
> > to make backups to a local IDE disk, and we noticed that getdents() took
> > about the same time as Peter reports (0.2 to 2 seconds), especially in
> > maildir directories. We tried many things to fix it with no result,
> > including enabling dirindexes. Finally, we made a full backup, and switched
> > over to XFS and the problem totally disappeared. So it seems that the
> > filesystem matters a lot here when there are lots of entries in a
> > directory, and that ext3 is not suitable for usages with thousands
> > of entries in directories with millions of files on disk. I'm not
> > certain it would be that easy to try other filesystems on kernel.org
> > though :-/
> >
> 
> Yeah, slowly-growing directories will get splattered all over the disk.
> 
> Possible short-term fixes would be to just allocate up to (say) eight
> blocks when we grow a directory by one block.  Or teach the
> directory-growth code to use ext3 reservations.
> 
> Longer-term people are talking about things like on-disk rerservations.
> But I expect directories are being forgotten about in all of that.

By on-disk reservations, do you mean persistent file preallocation ? (that is explicit preallocation of blocks to a given file) If so, you are right, we haven't really given any thought to the possibility of directories needing that feature.

Regards Suparna

Show 5 quoted lines
> 
> -
> To unsubscribe from this list: send the line "unsubscribe linux-ext4" in
> the body of a message to majordomo@vger.kernel.org
> More majordomo info at  http://vger.kernel.org/majordomo-info.html
-- 
Suparna Bhattacharya (suparna@in.ibm.com)
Linux Technology Center
IBM Software Lab, India
Previous: Rene HermanNext: Theodore Tso
Message 25 of 52 in “Re: [KORG] Re: kernel.org lies about latest -mm kernel”
  1. Jeff GarzikJan 7, 2007
  2. Linus TorvaldsJan 7, 2007
  3. Greg KHJan 7, 2007
  4. H. Peter AnvinJan 7, 2007
  5. Junio C HamanoJan 7, 2007
  6. Jeff GarzikJan 7, 2007
  7. Linus TorvaldsJan 7, 2007
  8. Martin LanghoffJan 7, 2007
  9. How git affects kernel.org performanceH. Peter Anvin, Jan 7, 2007
  10. Linus TorvaldsJan 7, 2007
  11. Willy TarreauJan 7, 2007
  12. H. Peter AnvinJan 7, 2007
  13. Willy TarreauJan 7, 2007
  14. Christoph HellwigJan 7, 2007
  15. Willy TarreauJan 7, 2007
  16. Linus TorvaldsJan 7, 2007
  17. Linus TorvaldsJan 7, 2007
  18. Jan EngelhardtJan 7, 2007
  19. Randy DunlapJan 7, 2007
  20. Jan EngelhardtJan 7, 2007
  21. Randy DunlapJan 7, 2007
  22. Linus TorvaldsJan 7, 2007
  23. Andrew MortonJan 7, 2007
  24. Rene HermanJan 7, 2007
  25. Suparna BhattacharyaJan 8, 2007
  26. Theodore TsoJan 8, 2007
  27. Johannes StezenbachJan 8, 2007
  28. Theodore TsoJan 8, 2007
  29. Pavel MachekJan 8, 2007
  30. Theodore TsoJan 8, 2007
  31. Jeff GarzikJan 8, 2007
  32. Paul JacksonJan 9, 2007
  33. Jeremy HigdonJan 9, 2007
  34. Fengguang WuJan 9, 2007
  35. Linus TorvaldsJan 9, 2007
  36. Fengguang WuJan 10, 2007
  37. Fengguang WuJan 10, 2007
  38. Fengguang WuJan 10, 2007
  39. Fengguang WuJan 9, 2007
  40. Fengguang WuJan 9, 2007
  41. Robert FitzsimonsJan 7, 2007
  42. J.H.Jan 7, 2007
  43. Jakub NarebskiJan 8, 2007
  44. Krzysztof HalasaJan 7, 2007
  45. Shawn O. PearceJan 7, 2007
  46. Nicolas PitreJan 8, 2007
  47. Linus TorvaldsJan 7, 2007
  48. Nigel CunninghamJan 10, 2007
  49. Fengguang WuJan 10, 2007
  50. Fengguang WuJan 10, 2007
  51. Fengguang WuJan 10, 2007
  52. Nigel CunninghamJan 12, 2007

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.