git/list[1] front-page[2] threads[3] people[4] search[5] about
 

[PATCH 3/5] wrapper: properly handle MAX_IO_SIZE in writev(3p)

From
Patrick Steinhardt <ps@pks.im>
Date
Jul 16, 2026, 07:52 UTC
Message-ID
<20260716-pks-reintroduce-writev-v1-3-ea9038c884bc@pks.im>
In-Reply-To
<20260716-pks-reintroduce-writev-v1-0-ea9038c884bc@pks.im>

Some systems like NonStop set a comparatively small `MAX_IO_SIZE`, which limits the maximum number of bytes we're allowed to write in a single call. We already handle this limit properly in `xwrite()`, but we have recently introduced wrappers for writev(3p) where we don't. This will cause the syscall to return EINVAL in case somebody passes an iovec entry to writev(3p) that is larger than `MAX_IO_SIZE`.

Introduce a new function `xwritev()` that is similar to `xwrite()` in that it handles such platform-specific nuances:

  - We only pass the leading iovec entries to writev(3p) that fit into
    `MAX_IO_SIZE`, pretending that the underlying syscall performed a
    short write. This mirrors how `xwrite()` chomps overly large
    requests before handing them to write(3p). As a consequence, callers
    will never see writev(3p)'s EINVAL error for requests whose summed
    length would overflow an ssize_t, but observe a short write instead.
  - If already the first iovec entry exceeds the limit we instead punt
    to `xwrite()`, which knows to handle this case for us.
  - We restart the underlying syscall on EINTR and EAGAIN, just like
    `xwrite()` does for write(3p).

Adapt `writev_in_full()` to use this new wrapper. With the retry logic now living in `xwritev()`, the calling loop becomes the exact mirror image of `write_in_full()`, which also retains the responsibility of translating a zero-length write into ENOSPC.

Reported-by: Randall Becker <randall.becker@nexbridge.ca>
Helped-by: Jeff King <peff@peff.net>
Helped-by: Junio C Hamano <gitster@pobox.com>
Signed-off-by: Patrick Steinhardt <ps@pks.im>
---
 wrapper.c | 47 ++++++++++++++++++++++++++++++++++++++++++-----
 wrapper.h |  1 +
 2 files changed, 43 insertions(+), 5 deletions(-)
diff --git a/wrapper.c b/wrapper.c
index be8fa575e6..561f9ee9c9 100644
--- a/wrapper.c
+++ b/wrapper.c
@@ -323,17 +323,54 @@ ssize_t write_in_full(int fd, const void *buf, size_t count)
 	return total;
 }
 
+ssize_t xwritev(int fd, struct iovec *iov, int iovcnt)
+{
+	size_t allowed = MAX_IO_SIZE;
+	int i;
+
+	/*
+	 * Some platforms define a comparatively small `MAX_IO_SIZE` that
+	 * limits how many bytes can be written with a single call to
+	 * write(3p) or writev(3p); exceeding that limit causes the syscall to
+	 * fail with EINVAL. Just like xwrite() chomps overly large requests
+	 * for write(3p), pretend that the underlying writev(3p) performed a
+	 * short write by only passing along the leading iovec entries that
+	 * fit into that limit.
+	 */
+	for (i = 0; i < iovcnt; i++) {
+		if (iov[i].iov_len > allowed) {
+			/*
+			 * If the first buffer is larger than MAX_IO_SIZE,
+			 * let xwrite() deal with it.
+			 */
+			if (!i)
+				return xwrite(fd, iov->iov_base, iov->iov_len);
+			break;
+		}
+		allowed -= iov[i].iov_len;
+	}
+
+	while (1) {
+		ssize_t bytes_written = writev(fd, iov, i);
+		if (bytes_written < 0) {
+			if (errno == EINTR)
+				continue;
+			if (handle_nonblock(fd, POLLOUT, errno))
+				continue;
+		}
+
+		return bytes_written;
+	}
+}
+
 ssize_t writev_in_full(int fd, struct iovec *iov, int iovcnt)
 {
 	ssize_t total_written = 0;
 
 	while (iovcnt) {
-		ssize_t bytes_written = writev(fd, iov, iovcnt);
-		if (bytes_written < 0) {
-			if (errno == EINTR || errno == EAGAIN)
-				continue;
+		ssize_t bytes_written = xwritev(fd, iov, iovcnt);
+		if (bytes_written < 0)
 			return -1;
-		}
 		if (!bytes_written) {
 			errno = ENOSPC;
 			return -1;
diff --git a/wrapper.h b/wrapper.h
index 27519b32d1..a6287d7f4d 100644
--- a/wrapper.h
+++ b/wrapper.h
@@ -16,6 +16,7 @@ void *xmmap_gently(void *start, size_t length, int prot, int flags, int fd, off_
 int xopen(const char *path, int flags, ...);
 ssize_t xread(int fd, void *buf, size_t len);
 ssize_t xwrite(int fd, const void *buf, size_t len);
+ssize_t xwritev(int fd, struct iovec *iov, int iovcnt);
 ssize_t xpread(int fd, void *buf, size_t len, off_t offset);
 int xdup(int fd);
 FILE *xfopen(const char *path, const char *mode);
-- 
2.55.0.313.g8d093f411d.dirty
Previous: Patrick SteinhardtNext: Patrick Steinhardt
Message 8 of 27 in “Reintroduce writev(3p)”
  1. 0/5 Reintroduce writev(3p)Patrick Steinhardt, Jul 16, 2026
  2. 1/5 compat/posix: introduce writev(3p) wrapperPatrick Steinhardt, Jul 16, 2026
  3. Simon RichterJul 16, 2026
  4. Junio C HamanoJul 16, 2026
  5. Junio C HamanoJul 16, 2026
  6. Patrick SteinhardtAug 5, 2026
  7. 2/5 wrapper: introduce writev(3p) wrappersPatrick Steinhardt, Jul 16, 2026
  8. 3/5 wrapper: properly handle MAX_IO_SIZE in writev(3p)Patrick Steinhardt, Jul 16, 2026
  9. 4/5 sideband: use writev(3p) to send pktlinesPatrick Steinhardt, Jul 16, 2026
  10. 5/5 fast-import: use writev(3p) to send cat-blob responsesPatrick Steinhardt, Jul 16, 2026
  11. Johannes SixtJul 16, 2026
  12. Junio C HamanoJul 27, 2026
  13. Patrick SteinhardtAug 5, 2026
  14. Junio C HamanoAug 5, 2026
  15. Johannes SixtAug 5, 2026
  16. Junio C HamanoAug 5, 2026
  17. Johannes SixtAug 5, 2026
  18. Junio C HamanoAug 5, 2026
  19. Patrick SteinhardtAug 6, 2026
  20. Junio C HamanoAug 6, 2026
  21. Patrick SteinhardtAug 7, 2026
  22. 0/5 Reintroduce writev(3p)Patrick Steinhardt, Aug 7, 2026
  23. 1/5 compat/posix: introduce writev(3p) wrapperPatrick Steinhardt, Aug 7, 2026
  24. 2/5 wrapper: introduce writev(3p) wrappersPatrick Steinhardt, Aug 7, 2026
  25. 3/5 wrapper: properly handle MAX_IO_SIZE in writev(3p)Patrick Steinhardt, Aug 7, 2026
  26. 4/5 sideband: use writev(3p) to send pktlinesPatrick Steinhardt, Aug 7, 2026
  27. 5/5 fast-import: use writev(3p) to send cat-blob responsesPatrick Steinhardt, Aug 7, 2026

Read the whole thread, see it on lore, or plain text.

$ cat FOOTERMessages come from the public archive at lore.kernel.org/git, fetched every hour. The front page is chosen and written each morning by an AI editor and can be wrong; the threads themselves are the record. About and API. For agents: an MCP server at https://gitlist.dev/mcp, and any thread, story or person page as Markdown by adding .md to its URL (or sending Accept: text/markdown). Details in /llms.txt.