{"thread":{"id":"55518","subject":"Strategy to deal with slow cloners","startedAt":"2021-04-19T12:46:33Z","lastAt":"2021-04-23T10:02:09Z","messageCount":6,"participants":["Konstantin Ryabitsev","Eric Wong","Thomas Braun","Ævar Arnfjörð Bjarmason","Jeff King"],"isPatch":false,"patchVersion":null,"patchTotal":null},"messages":[{"id":"422309","messageId":"20210419124623.wwps2s35x2mrrhi6@nitro.local","threadId":"55518","inReplyTo":null,"subject":"Strategy to deal with slow cloners","fromName":"Konstantin Ryabitsev","fromEmail":"konstantin@linuxfoundation.org","sentAt":"2021-04-19T12:46:23Z","receivedAt":"2021-04-19T12:46:33Z","isPatch":false,"sender":{"key":"konstantin@linuxfoundation.org","avatar":"https://gravatar.com/avatar/7cb8827c6de56e1bd2dea16508c6708aa43feed3bf3813bcdacecdf96ceadd79?d=mp&s=160"},"body":"Hello:\n\nI try to keep repositories routinely repacked and optimized for clones, in\nhopes that most operations needing lots of objects would be sending packs\nstraight from disk. However, every now and again a client from a slow\nconnection requests a large clone and then takes half a day downloading it,\nresulting in gigabytes of RAM being occupied by a temporary pack.\n\nAre there any strategies to reduce RAM usage in such cases, other than\nvm.swappiness (which I'm not sure would work, since it's not a sleeping\nprocess)? Is there a way to write large temporary packs somewhere to disk\nbefore sendfile'ing them?\n\n-K\n"},{"id":"422320","messageId":"20210419180803.GA10171@dcvr","threadId":"55518","inReplyTo":"20210419124623.wwps2s35x2mrrhi6@nitro.local","subject":"Re: Strategy to deal with slow cloners","fromName":"Eric Wong","fromEmail":"e@80x24.org","sentAt":"2021-04-19T18:08:03Z","receivedAt":"2021-04-19T18:08:05Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"Konstantin Ryabitsev <konstantin@linuxfoundation.org> wrote:\n> Hello:\n> \n> I try to keep repositories routinely repacked and optimized for clones, in\n> hopes that most operations needing lots of objects would be sending packs\n> straight from disk. However, every now and again a client from a slow\n> connection requests a large clone and then takes half a day downloading it,\n> resulting in gigabytes of RAM being occupied by a temporary pack.\n\nYeah, I'm familiar with the problem.\n\n> Are there any strategies to reduce RAM usage in such cases, other than\n> vm.swappiness (which I'm not sure would work, since it's not a sleeping\n> process)? Is there a way to write large temporary packs somewhere to disk\n> before sendfile'ing them?\n\npublic-inbox-httpd actually switched buffering strategies in\n2019 to favor hitting ENOSPC instead of ENOMEM :)\n\n  https://public-inbox.org/meta/20190629195951.32160-11-e@80x24.org/\n\nIt doesn't support sendfile, currently (I didn't want separate\nHTTPS vs HTTP code paths), but that's probably not too big of a\ndeal, especially with slow clients.\n\nIt's capable of serving non-public-inbox coderepos (and running\ncgit).  Instead of configuring every [coderepo \"...\"] manually,\npublicinbox.cgitrc can be set in ~/.public-inbox/config to\nmass-configure [coderepo] sections.  It's only lightly-tested\nfor my setup atm, though.\n\nMapping publicinbox.<name>.coderepo to [coderepo \"...\"]\nentries for solver (blob reconstruction) isn't required;\nit's a bit of a pain at a large scale and I haven't figured\nout how to make it easier.\n"},{"id":"422441","messageId":"f3b406a7-90cd-89c1-c532-d9a7c0f71599@virtuell-zuhause.de","threadId":"55518","inReplyTo":"20210419124623.wwps2s35x2mrrhi6@nitro.local","subject":"Re: Strategy to deal with slow cloners","fromName":"Thomas Braun","fromEmail":"thomas.braun@virtuell-zuhause.de","sentAt":"2021-04-20T14:52:39Z","receivedAt":"2021-04-20T15:17:30Z","isPatch":false,"sender":{"key":"thomas.braun@virtuell-zuhause.de","avatar":"https://avatars.githubusercontent.com/u/1185677?v=4"},"body":"On 19.04.2021 14:46, Konstantin Ryabitsev wrote:\n\n> I try to keep repositories routinely repacked and optimized for clones, in\n> hopes that most operations needing lots of objects would be sending packs\n> straight from disk. However, every now and again a client from a slow\n> connection requests a large clone and then takes half a day downloading it,\n> resulting in gigabytes of RAM being occupied by a temporary pack.\n> \n> Are there any strategies to reduce RAM usage in such cases, other than\n> vm.swappiness (which I'm not sure would work, since it's not a sleeping\n> process)? Is there a way to write large temporary packs somewhere to disk\n> before sendfile'ing them?\n\nThere is the packfile-uris feature which allows protocol v2 servers to\nadvertise static packfiles via http/https. But clients must explicitly\nenable it via fetch.uriprotocols. So this does only work for newish\nclients which explicitly ask for it. See\nDocumentation/technical/packfile-uri.txt.\n\nFrom my limited understanding one clone/fetch the server can only send\none packfile at most.\n\nWhat is the advertised git clone command on the website? Maybe something\nlike git clone --depth=$num would help reduce the load? Usually not\neveryone needs the whole history.\n\n"},{"id":"422631","messageId":"20210421200816.GA13772@dcvr","threadId":"55518","inReplyTo":"20210419180803.GA10171@dcvr","subject":"Re: Strategy to deal with slow cloners","fromName":"Eric Wong","fromEmail":"e@80x24.org","sentAt":"2021-04-21T20:08:16Z","receivedAt":"2021-04-21T20:08:18Z","isPatch":false,"sender":{"key":"e@80x24.org","avatar":null},"body":"Eric Wong <e@80x24.org> wrote:\n> Konstantin Ryabitsev <konstantin@linuxfoundation.org> wrote:\n> > Hello:\n> > \n> > I try to keep repositories routinely repacked and optimized for clones, in\n> > hopes that most operations needing lots of objects would be sending packs\n> > straight from disk. However, every now and again a client from a slow\n> > connection requests a large clone and then takes half a day downloading it,\n> > resulting in gigabytes of RAM being occupied by a temporary pack.\n> \n> Yeah, I'm familiar with the problem.\n\nAlso, AFAIK nginx has \"proxy_buffering on\" by default.  However,\nI seem to recall that prevents clients from seeing a single byte\nuntil the pack is completely generated.  It's been many years\nsince I've used nginx myself, so my knowledge about it could be\nout-of-date.\n"},{"id":"422670","messageId":"87y2da1xx2.fsf@evledraar.gmail.com","threadId":"55518","inReplyTo":"20210419124623.wwps2s35x2mrrhi6@nitro.local","subject":"Re: Strategy to deal with slow cloners","fromName":"Ævar Arnfjörð Bjarmason","fromEmail":"avarab@gmail.com","sentAt":"2021-04-22T09:16:57Z","receivedAt":"2021-04-22T09:17:03Z","isPatch":false,"sender":{"key":"avarab@gmail.com","avatar":"https://avatars.githubusercontent.com/u/45301?v=4"},"body":"\nOn Mon, Apr 19 2021, Konstantin Ryabitsev wrote:\n\n> Hello:\n>\n> I try to keep repositories routinely repacked and optimized for clones, in\n> hopes that most operations needing lots of objects would be sending packs\n> straight from disk. However, every now and again a client from a slow\n> connection requests a large clone and then takes half a day downloading it,\n> resulting in gigabytes of RAM being occupied by a temporary pack.\n>\n> Are there any strategies to reduce RAM usage in such cases, other than\n> vm.swappiness (which I'm not sure would work, since it's not a sleeping\n> process)? Is there a way to write large temporary packs somewhere to disk\n> before sendfile'ing them?\n\nAside from any Git-specific solutions, perhaps the right kernel settings\n+ a cron script re-nicing such processes that have been active for more\nthan X amount of time will help?\n\nI'm not familiar with the guts of Linux's swapping algorithm, but some\nresults online seem to suggest that it takes the nice level into account\nwhen deciding what to swap out, i.e. with the right level it might give\npreference to swapping out this mostly idle process.\n"},{"id":"422743","messageId":"YIKbHWzxB5Q0Pe0E@coredump.intra.peff.net","threadId":"55518","inReplyTo":"20210419124623.wwps2s35x2mrrhi6@nitro.local","subject":"Re: Strategy to deal with slow cloners","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2021-04-23T10:02:05Z","receivedAt":"2021-04-23T10:02:09Z","isPatch":false,"sender":{"key":"peff@peff.net","avatar":"https://avatars.githubusercontent.com/u/45925?v=4"},"body":"On Mon, Apr 19, 2021 at 08:46:23AM -0400, Konstantin Ryabitsev wrote:\n\n> I try to keep repositories routinely repacked and optimized for clones, in\n> hopes that most operations needing lots of objects would be sending packs\n> straight from disk. However, every now and again a client from a slow\n> connection requests a large clone and then takes half a day downloading it,\n> resulting in gigabytes of RAM being occupied by a temporary pack.\n> \n> Are there any strategies to reduce RAM usage in such cases, other than\n> vm.swappiness (which I'm not sure would work, since it's not a sleeping\n> process)?\n\n\nDo you know where the RAM is going? I.e., heap or mmap'd files in block\ncache? Do you have recent reachability bitmaps built?\n\nTraditionally, most of the heap usage in pack-objects went to:\n\n  - the set of object structs used for traversal; likewise, internal\n    caches like the delta-base cache that get filled during the\n    traversal\n\n  - the big book-keeping array of all of the objects we are planning to\n    send (and all their metadata)\n\n  - the reverse index we load in memory to find object offsets and\n    sizes within the packfile\n\nBut with bitmaps, we can skip most of the traversal entirely. And\nthere's a \"pack reuse\" mechanism that tries to avoid even adding objects\nto the book-keeping array when we are just sending the first chunk of\nthe pack verbatim anyway.\n\nE.g., on a clone of torvalds/linux, running:\n\n  git for-each-ref --format='%(objectname)' refs/heads/ refs/tags/ |\n  valgrind --tool=massif git pack-objects --revs --delta-base-offset --stdout |\n  wc -c\n\nhits a peak heap of 1.9GB without bitmaps enabled but only 326MB with.\n\nOn top of that, if you have Git v2.31, try enabling pack.writeReverseIndex\nand repacking. That drops the heap to just 23MB! (though note there's\nsome cheating here; we're mmap-ing 31MB of .rev file plus 47MB of\n.bitmap file).\n\nFrom previous conversations, I expect you're already using bitmaps, but\nyou might double-check that things are kicking in as you'd expect (you\ncan get a rough read on heap of running processes by subtracting shared\nmemory from rss). And probably you aren't using on-disk revindexes yet,\nbecause they're not enabled by default.\n\nIf your problem is block cache (i.e., it's the total rss that's the\nproblem, not just the heap parts), that's harder. If you have a lot of\nrelated repositories (say, forks of the kernel), your best bet is to use\nalternates to share the storage. That opens up a whole other can of\ncomplexity worms that I won't get into here.\n\nGetting back to your other question:\n\n> Is there a way to write large temporary packs somewhere to disk\n> before sendfile'ing them?\n\nThe uploadpack.packObjectsHook config would let you wrap pack-objects\nwith a script that writes to a temporary file, and then just uses\nsomething simple like \"cat\" to feed it back to upload-pack. You can't\nuse sendfile(), because there's some protocol framing that happens in\nupload-pack, but its memory use is relatively low.\n\nIf block cache is your problem (and not heap), this _could_ make things\nslightly worse, as now you're writing the same pack data out to an extra\nfile which isn't shared by multiple processes. So if you have multiple\nclients hitting the same repository, you may increase your working set\nsize. The OS may be able to handle it better, though (e.g., the linear\nread through the file by cat makes it obvious that it can drop pages\nfrom earlier parts of the file under memory pressure).\n\nYour wrapper of course can get more clever about such things, too. E.g.,\nyou can coalesce identical requests to use the same cached copy,\nskipping even the extra call to pack-objects in the first place. We do\nsomething like that at GitHub (unfortunately not with an open tool that\nI can share at this point).\n\n-Peff\n"}]}