{"thread":{"id":"350","subject":"Darcs-git pulling from the Linux repo: a Linux VM question","startedAt":"2005-04-27T13:10:25Z","lastAt":"2005-04-29T11:25:47Z","messageCount":7,"participants":["Juliusz Chroboczek","Linus Torvalds","David Roundy"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"1860","messageId":"7i7jionz5q.fsf@lanthane.pps.jussieu.fr","threadId":"350","inReplyTo":null,"subject":"Darcs-git pulling from the Linux repo: a Linux VM question","fromName":"Juliusz Chroboczek","fromEmail":"juliusz.chroboczek@pps.jussieu.fr","sentAt":"2005-04-27T13:10:25Z","receivedAt":"2005-04-27T13:10:25Z","isPatch":false,"sender":{"key":"juliusz.chroboczek@pps.jussieu.fr","avatar":null},"body":"Hi,\n\nIf you are one of the few initiated who can tune the Linux VM, please\nskip to the end of this mail and give me some advice.  If you are one\nof the even fewer initiated who understand Darcs' memory usage, read\nthe whole of this message and send me a patch.  Otherwise, press D.\n\nNow that I've got a Darcs that groks Git repos, I can play with a\nfairly large tree -- the Linux 2.6 one.  All the experiments described\nbelow were done on a 1.4 MHz Pentium-M with 640 MB of memory, running\nLinux 2.6.9 (Debian branded) over Reiserfs. \n\nAll the commands that don't need to actually read the underlying blobs\nare instantaneous; for example, ``darcs changes'' takes 0.4s.\nCommands that require reading the blobs but allow discarding them\nstraight away are reasonable enough -- ``darcs changes -s'' on all but\nthe initial import takes a very reasonable 15s, ``darcs changes -s''\nincluding the initial import takes 2m30s real time, (50s CPU time).\n\nThe trouble, of course, is with commands that need to read a full tree\nand keep it in memory.  This is, unfortunately, the case with pull of\nthe initial commit, which is over 200MB in size.  Darcs behaviour when\npulling this initial commit is as follows.\n\nAs I'm currently reading the git repository eagerly, Darcs starts by\nreading the whole of the initial tree into memory; this takes roughly\n2 minutes of real time (at less than 10% CPU), reads 18987 Git files\n(blobs and trees), of which 18512 are unique (meaning that less than\n500 were read two times or more -- yes, I should be keeping track of\nthe blobs I've already read).  When that is done, Darcs' VMEM usage is\nbeneath 300MB.\n\nAt that point, Darcs stops doing I/O, and starts trying to interpret\nthe data.  It runs between 80% and 100% of CPU, and grows steadily\nuntil its VMEM reaches 550MB.  At that point, the system starts\nswapping very lightly (no more than 200kB/s or so), and Darcs' VMEM\nusage grows up to 720MB after 5 minutes CPU, 8 minutes real time.\n\nWhen Darcs has grokked the fullness of the Linux kernel, it decides to\nwrite out a patch.  So it starts touching all of its memory while\nsimultaneously writing out data to a patch file at a fairly sustained\nrate.  It gets pretty close to the end -- over 200MB of patch are\nwritten --, when suddenly the system appears to freeze for a second,\nthen the OOM killer triggers and kills the Darcs process.\n\nNow obviously there is a problem with Darcs -- it shouldn't be needing\n720MB of virtual memory just to grok a 250 MB import --, but there's\nalso a problem with the VM.  A 720 MB process should be reasonable on\na machine with 640 MB, and there's no apparent reason why the kernel\ncouldn't go more heavily into swap.  My completely uninformed guess\nwould be that the heavy I/O activity generated by Darcs in the final\nstage causes a shortage of some resource (probably buffers) that is\nessential for the VM to perform the swapping, and that the only way\nthe kernel sees to get itself out of the tight spot is to invoke the\nOOM killer on the process that's causing the I/O activity.\n\nSo yes, in the longer term we need to fix Darcs.  For now, does anyone\nknow how I can tune the Linux VM to get a 720 MB process to run\nreliably in 640 MB of main memory?  Obviously, adding swap or tuning\nthe overcommit policy doesn't help (the issue is precisely that the VM\nrefuses to dig into the swap early enough).  I don't understand what\n``swappinness'' is, but it doesn't appear to help.  The\n``min_free_kbytes'' and ``dirty_*'' knobs look promising, but nobody\nseems to know what they mean.\n\nSo what was it you said about self-tuning VM systems?\n\n                                        Juliusz\n"},{"id":"1866","messageId":"Pine.LNX.4.58.0504270823480.18901@ppc970.osdl.org","threadId":"350","inReplyTo":"7i7jionz5q.fsf@lanthane.pps.jussieu.fr","subject":"Re: Darcs-git pulling from the Linux repo: a Linux VM question","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-27T15:31:37Z","receivedAt":"2005-04-27T15:31:37Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, Juliusz Chroboczek wrote:\n> \n> So yes, in the longer term we need to fix Darcs.  For now, does anyone\n> know how I can tune the Linux VM to get a 720 MB process to run\n> reliably in 640 MB of main memory?\n\nI really think you're screwed. The only way you have even a _chance_ of\ngetting it to work well is that if you have very nice access patterns to\nthat 720MB, but my guess is that that simply isn't the case. You probably\nread most of it in once (and write out changes once, but I hope you at\nleast notice the case of \"nothing changed\" so that probably is the smaller\nof your problems), and the fact is, you're going to have absolutely\n_horrible_ access patterns, since you'll end up not just with a 720MB\nprocess that doesn't have much locality, you'll end up with another 720MB\nthat you needed to have in the page cache for the IO.\n\nThe only way I can see to fix it short-term is to try to use \"mmap()\"  \ninstead of \"read()\" to read the file data, and then try to avoid touching\nthe mapping unless you _have_ to. In other words: if you actually need to\n_compare_ the data (which obviously reads from the mapping), you're\nscrewed.\n\nUsing mmap() will at least mean that the system can re-use the page cache \npages, though, so it should improve memory pressure a bit.\n\n> So what was it you said about self-tuning VM systems?\n\nThe kernel tries to tune itself in the sense that it automatically \nallocates the memory to user processes vs caching (page cache, directory \ncaching etc) and tunes itself quite well that way.\n\nBut there's no way to tune for crappy access patterns and working sets\nbigger than the amount of RAM. Sorry. You really need to fix darcs.\n\nYou _really_ shouldn't read in files that you don't absolutely need.  \nThat's really the biggest point of git: using the sha1 for naming the\nobjects is really all about \"descrive the contents using 20 bytes instead\nof by reading the contents\". Because reading the content _will_ be\nexpensive. Even if you have 2GB of memory and you can keep it all cached,\nit will be horribly expensive.\n\n\t\tLinus\n"},{"id":"1869","messageId":"7iu0lskyfb.fsf@lanthane.pps.jussieu.fr","threadId":"350","inReplyTo":"Pine.LNX.4.58.0504270823480.18901@ppc970.osdl.org","subject":"Re: Re: Darcs-git pulling from the Linux repo: a Linux VM question","fromName":"Juliusz Chroboczek","fromEmail":"juliusz.chroboczek@pps.jussieu.fr","sentAt":"2005-04-27T15:54:32Z","receivedAt":"2005-04-27T15:54:32Z","isPatch":false,"sender":{"key":"juliusz.chroboczek@pps.jussieu.fr","avatar":null},"body":">> For now, does anyone know how I can tune the Linux VM to get a 720\n>> MB process to run reliably in 640 MB of main memory?\n\n> I really think you're screwed.\n\nThanks, that's what I needed to know.\n\n> You _really_ shouldn't read in files that you don't absolutely need.\n\nAhem... you don't expect me to embark on hacking Git without at least\nunderstanding that, do you?\n\n> That's really the biggest point of git: using the sha1 for naming the\n> objects is really all about \"descrive the contents using 20 bytes instead\n> of by reading the contents\".\n\nHere we're speaking about the initial import.  Committed on 17 April\n2005 by Linus Torvalds, with the comment ``Let it rip''.  220 MB of\nchanged files in a single commit.  2 minutes real time just to read\nall the files, never mind doing anything useful with them.\n\nTo put it mildly, Darcs is not optimised for that sort of usage.\n\n> Sorry.  You really need to fix darcs.\n\nThat's exactly why we're so interested in your repository.\n\n                                        Juliusz\n"},{"id":"1871","messageId":"Pine.LNX.4.58.0504270910510.18901@ppc970.osdl.org","threadId":"350","inReplyTo":"7iu0lskyfb.fsf@lanthane.pps.jussieu.fr","subject":"Re: Re: Darcs-git pulling from the Linux repo: a Linux VM question","fromName":"Linus Torvalds","fromEmail":"torvalds@osdl.org","sentAt":"2005-04-27T16:16:03Z","receivedAt":"2005-04-27T16:16:03Z","isPatch":false,"sender":{"key":"torvalds@linux-foundation.org","avatar":"https://avatars.githubusercontent.com/u/1024025?v=4"},"body":"\n\nOn Wed, 27 Apr 2005, Juliusz Chroboczek wrote:\n> \n> Here we're speaking about the initial import.  Committed on 17 April\n> 2005 by Linus Torvalds, with the comment ``Let it rip''.  220 MB of\n> changed files in a single commit.  2 minutes real time just to read\n> all the files, never mind doing anything useful with them.\n\nI think you may well want to consider the initial commit special. In many \nways it is - it has no parents etc, so even apart from the fact that the \ninitial commit obviously tends to be a lot bigger than any other commit, \nit actually fundamnetally is _technically_ different too.\n\n> To put it mildly, Darcs is not optimised for that sort of usage.\n\nIt shouldn't be. Make the initial one a special case, and import things \nfile-by-file for that one special case.\n\nAfterwards, you should be able to handle other commits as \"diffs\", and\nthen it's entirely reasonable to have the difference all in memory. If\nsomebody really does end up having a 220MB diff, and darcs sucks at it,\nthen at that point I don't think it's darcs' problem any more, it's the\nproject that you're trying to track that is doing something wrong..\n\nSo if you _just_ consider the initial git commit special (and it's easy to \nnotice by just looking at the lack of parents), then you may not need to \nchange darcs in the other cases.\n\nAnd almost all SCM's consider the initial state a special case anyway. The \nfact that GIT doesn't is just a result of the strange way of representing \ndata, which doesn't care. I don't think you should emulate git in that \nrespect.\n\n\t\tLinus\n"},{"id":"2027","messageId":"20050428113947.GC9422@abridgegame.org","threadId":"350","inReplyTo":"Pine.LNX.4.58.0504270910510.18901@ppc970.osdl.org","subject":"Re: [darcs-devel] Re: Darcs-git pulling from the Linux repo: a Linux VM question","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-28T11:39:52Z","receivedAt":"2005-04-28T11:39:52Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Wed, Apr 27, 2005 at 09:16:03AM -0700, Linus Torvalds wrote:\n> On Wed, 27 Apr 2005, Juliusz Chroboczek wrote:\n> > Here we're speaking about the initial import.  Committed on 17 April\n> > 2005 by Linus Torvalds, with the comment ``Let it rip''.  220 MB of\n> > changed files in a single commit.  2 minutes real time just to read\n> > all the files, never mind doing anything useful with them.\n> \n> I think you may well want to consider the initial commit special. In many \n> ways it is - it has no parents etc, so even apart from the fact that the \n> initial commit obviously tends to be a lot bigger than any other commit, \n> it actually fundamnetally is _technically_ different too.\n\nThis has been discussed, and while I'm not opposed to special-casing the\ninitial commit, mostly we've taken the stance so far of not special-casing.\nIt's much nicer if we can make darcs efficient enough to perform the\ninitial commit without a special case, which has the nice side-effect of\nalso improving other cases.\n\nWhen we're desperate, we'll special-case the initial commit, but currently\nI'm sure we can pretty easily adjust things by making the git-tree-reading\nlazy, which should pretty well address both the memory and speed\nconcerns--and also improve performance of other commands.  Perhaps more to\nthe point, it will also ensure that the same optimizations that work for\nworking with darcs repos will help when dealing with git repos.\n-- \nDavid Roundy\nhttp://www.darcs.net\n"},{"id":"2043","messageId":"7ir7guj4m6.fsf@lanthane.pps.jussieu.fr","threadId":"350","inReplyTo":"20050428113947.GC9422@abridgegame.org","subject":"Re: [darcs-devel] Re: Darcs-git pulling from the Linux repo: a Linux VM question","fromName":"Juliusz Chroboczek","fromEmail":"juliusz.chroboczek@pps.jussieu.fr","sentAt":"2005-04-28T15:36:01Z","receivedAt":"2005-04-28T15:36:01Z","isPatch":false,"sender":{"key":"juliusz.chroboczek@pps.jussieu.fr","avatar":null},"body":"> When we're desperate, we'll special-case the initial commit, but currently\n> I'm sure we can pretty easily adjust things by making the git-tree-reading\n> lazy,\n\nJust to make it clear: reading the git tree is lazy.  The problem is\nsomewhere in the higher layers, probably in pull_cmd.\n\nThere's also another problem: reading the git tree takes 220MB.  Then\nDarcs allocates a further 500MB without calling my code at all.  (Some\nof it is doubtless due to linesPS, that should be more than a handful\nof megabytes.)\n\n                                        Juliusz\n"},{"id":"2125","messageId":"20050429112541.GA15698@abridgegame.org","threadId":"350","inReplyTo":"7ir7guj4m6.fsf@lanthane.pps.jussieu.fr","subject":"Re: [darcs-devel] Re: Darcs-git pulling from the Linux repo: a Linux VM question","fromName":"David Roundy","fromEmail":"droundy@abridgegame.org","sentAt":"2005-04-29T11:25:47Z","receivedAt":"2005-04-29T11:25:47Z","isPatch":false,"sender":{"key":"droundy@abridgegame.org","avatar":"https://gravatar.com/avatar/e8bcfd76f63303732bdfcdba6fc8ac6ccdff8f5a224a25ffaca24bd8a4c4571f?d=mp&s=160"},"body":"On Thu, Apr 28, 2005 at 05:36:01PM +0200, Juliusz Chroboczek wrote:\n> > When we're desperate, we'll special-case the initial commit, but\n> > currently I'm sure we can pretty easily adjust things by making the\n> > git-tree-reading lazy,\n> \n> Just to make it clear: reading the git tree is lazy.  The problem is\n> somewhere in the higher layers, probably in pull_cmd.\n\nI guess really that's the issue.  Get itself is a special case that we've\nalready optimized for the initial get.  You'll also run into trouble using\npull to grab an entire plain old darcs repository with a large initial\ncommit.  We can (and should) also optimize pull, but it's not going to ever\nbe as efficient as a get is, for the case where you start with an empty\nrepository.\n\n> There's also another problem: reading the git tree takes 220MB.  Then\n> Darcs allocates a further 500MB without calling my code at all.  (Some\n> of it is doubtless due to linesPS, that should be more than a handful\n> of megabytes.)\n\nThe output of linesPS actually does take a huge amount of space.\n-- \nDavid Roundy\nhttp://www.darcs.net\n"}]}