Re: [PATCH v2 03/10] upload-pack: reduce lock contention when writing packfile data
- From
Jeff King <peff@peff.net>
- Date
- Mar 5, 2026, 01:16 UTC
- Message-ID
- <20260305011638.GC4943@coredump.intra.peff.net>
- In-Reply-To
- <20260303-pks-upload-pack-write-contention-v2-3-7321830f08fe@pks.im>
On Tue, Mar 03, 2026 at 04:00:18PM +0100, Patrick Steinhardt wrote:
> Extend our use of the buffering infrastructure so that we soak up bytes > until the buffer is filled up at least 2/3rds of its capacity. The > change is relatively simple to implement as we already know to flush the > buffer in `create_pack_file()` after git-pack-objects(1) has finished.
This 2/3rds feels kind of arbitrary. Isn't our best bet to try to fill pkt-lines? Later you say:
Show 5 quoted lines
> Now we could of course go even further and make sure that we always fill > up the whole buffer. But this might cause an increase in read(3p) > syscalls, and some tests show that this only reduces the number of > write(3p) syscalls from 130,000 to 100,000. So overall this doesn't seem > worth it.
But I am not clear how it increases the number of read() calls. I guess you are concerned that we'll get 50k, and then do a read for the remaining 14k, and then read 50k, and then 14k, and so on. But I'm unconvinced that 2/3 is really any better here, as it depends on the buffering patterns of the upstream writer. They could be writing 1 byte less than 2/3, and we'd wait to buffer, then read half their next packet, write it, then read the second of of their next packet, wait to buffer, and so on.
Even just doing:
git clone --upload-pack='strace -e write git-upload-pack' --bare --no-local . foo.git
with this patch (and not the later one to increase the buffer size of pack-objects), I see an interesting flip-flop between packets of size 65515 and 61461. But we never send a single full-size one, even though pack-objects should be outpacing us (because we're slowed by running under strace). That's probably an OK loss of efficiency in practice, but it's very dependent on pack-objects buffering.
I'm still a little bit negative on the whole concept of buffering in upload-pack, just because the interactions between buffering layers can be so subtle. But I guess I'm not really making any argument that I didn't make in v1, and you kept this in v2, so I suppose you are not swayed by it. ;)
If we are going to buffer in upload-pack, there is an obvious optimization that I didn't see in your series. When we send a keepalive, we should just send whatever we have in os->buffer (even if it is nothing). If we are wasting 5 bytes of pkt-line header and a write() call to send the keepalive, we may as well send what data we do have.
I don't know how much it would help in practice, though. Most keepalives will come before the pack data starts, as once pack-objects starts producing data, it tends to do so pretty consistently. And of course we can't send os->buffer before we see the PACK header, because the whole point is to buffer the early bit waiting for packfile uris.
So it might not be worth adding.
-Peff