{"thread":{"id":"65551","subject":"[PATCH] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","startedAt":"2026-04-24T19:15:01Z","lastAt":"2026-05-12T05:51:20Z","messageCount":12,"participants":["Scott Bauersfeld via GitGitGadget","Junio C Hamano","Derrick Stolee","Jeff King"],"isPatch":true,"patchVersion":1,"patchTotal":null},"messages":[{"id":"542270","messageId":"pull.2282.git.git.1777058098756.gitgitgadget@gmail.com","threadId":"65551","inReplyTo":null,"subject":"[PATCH] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Scott Bauersfeld via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-24T19:14:58Z","receivedAt":"2026-04-24T19:15:01Z","isPatch":true,"body":"From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n\nBoth index-pack and unpack-objects read pack data from stdin through\na 4 KiB static buffer (input_buffer[4096]). On each fill(), consumed\nbytes are flushed to the output pack file via write_or_die(), so\nevery write(2) moves at most 4 KiB.\n\nOn FUSE-backed filesystems every write(2) is a synchronous round\ntrip through the FUSE protocol (userspace -> kernel -> userspace ->\nback), so the 4 KiB buffer turns a clone into many unnecessary tiny\nwrites with noticeable latency overhead.\n\nIncrease the buffer from 4 KiB to 128 KiB, matching the default\nalready used by the hashfile layer in csum-file.c.\n\nTesting with strace on HTTPS clones of git/git (~296 MB pack, 5 runs\nper variant, isolated builds from the same v2.54.0 source) shows:\n\n  index-pack pack file writes: 72,465 -> 24,943 avg (66% reduction)\n  total write() syscalls:     310,192 -> 259,530 avg (17% reduction)\n  writes of exactly 4096 bytes: ~40,077 -> 0 (eliminated)\n\nAll clones produce identical HEAD, file count, and pass fsck.\n\nSigned-off-by: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n---\n    index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB\n    \n    Both index-pack and unpack-objects read pack data from stdin through a 4\n    KiB static buffer (input_buffer[4096]). On each fill(), consumed bytes\n    are flushed to the output pack file via write_or_die(), so every\n    write(2) moves at most 4 KiB.\n    \n    On FUSE-backed filesystems every write(2) is a synchronous round trip\n    through the FUSE protocol (userspace → kernel → userspace → back), so\n    the 4 KiB buffer turns a clone into many unnecessary tiny writes with\n    noticeable latency overhead.\n    \n    This change increase the buffer from 4 KiB to 128 KiB, matching the\n    default already used by the hashfile layer in csum-file.c.\n    \n    Benchmarked with 5 HTTPS clones per version of\n    https://github.com/sbauersfeld/git.git (~296 MB pack), using strace -f\n    to count write() syscalls. Both binaries built from the same v2.54.0\n    source tree in isolated directories to ensure the bin-wrappers resolve\n    to the correct binary.\n    \n    Correctness verified via git fsck --no-dangling, rev-parse HEAD, and\n    working tree file count — all 10 clones match.\n    \n    Results:\n    \n    Metric Unpatched (4 KiB) Patched (128 KiB) Change index-pack writes to\n    pack file 72,465 avg 24,943 avg −66% Total write() syscalls (all\n    processes) 310,192 avg 259,530 avg −17% Writes of exactly 4096 bytes\n    ~40,077 avg 0 eliminated HEAD / file count / fsck ✓ ✓ None\n    \n    Raw data:\n    \n    unpatched (input_buffer[4096]): run 1: total_writes=311787\n    ip_pack_writes=72353 ip_4k=35311 run 2: total_writes=310252\n    ip_pack_writes=72348 ip_4k=38024 run 3: total_writes=309737\n    ip_pack_writes=72303 ip_4k=43003 run 4: total_writes=309801\n    ip_pack_writes=72661 ip_4k=42349 run 5: total_writes=309383\n    ip_pack_writes=72662 ip_4k=41702\n    \n    patched (input_buffer[128 * 1024]): run 1: total_writes=264659\n    ip_pack_writes=26605 ip_4k=0 run 2: total_writes=264276\n    ip_pack_writes=26568 ip_4k=0 run 3: total_writes=227796 ip_pack_writes=\n    9762 ip_4k=0 run 4: total_writes=262464 ip_pack_writes=27830 ip_4k=0 run\n    5: total_writes=278455 ip_pack_writes=33952 ip_4k=0\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2282%2Fsbauersfeld%2Fsb%2Fincrease-index-pack-input-buffer-v1\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2282/sbauersfeld/sb/increase-index-pack-input-buffer-v1\nPull-Request: https://github.com/git/git/pull/2282\n\n builtin/index-pack.c     | 4 ++--\n builtin/unpack-objects.c | 4 ++--\n 2 files changed, 4 insertions(+), 4 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex ca7784dc2c..81a628bf34 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -145,8 +145,8 @@ static int check_self_contained_and_connected;\n \n static struct progress *progress;\n \n-/* We always read in 4kB chunks. */\n-static unsigned char input_buffer[4096];\n+#define INPUT_BUFFER_SIZE (128 * 1024)\n+static unsigned char input_buffer[INPUT_BUFFER_SIZE];\n static unsigned int input_offset, input_len;\n static off_t consumed_bytes;\n static off_t max_input_size;\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex e01cf6e360..535c019f82 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -23,8 +23,8 @@\n static int dry_run, quiet, recover, has_errors, strict;\n static const char unpack_usage[] = \"git unpack-objects [-n] [-q] [-r] [--strict]\";\n \n-/* We always read in 4kB chunks. */\n-static unsigned char buffer[4096];\n+#define INPUT_BUFFER_SIZE (128 * 1024)\n+static unsigned char buffer[INPUT_BUFFER_SIZE];\n static unsigned int offset, len;\n static off_t consumed_bytes;\n static off_t max_input_size;\n\nbase-commit: 94f057755b7941b321fd11fec1b2e3ca5313a4e0\n-- \ngitgitgadget\n"},{"id":"542287","messageId":"xmqqldeb9w8e.fsf@gitster.g","threadId":"65551","inReplyTo":"pull.2282.git.git.1777058098756.gitgitgadget@gmail.com","subject":"Re: [PATCH] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-25T10:21:05Z","receivedAt":"2026-04-25T10:21:07Z","isPatch":true,"body":"\"Scott Bauersfeld via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n>\n> Both index-pack and unpack-objects read pack data from stdin through\n> a 4 KiB static buffer (input_buffer[4096]). On each fill(), consumed\n> bytes are flushed to the output pack file via write_or_die(), so\n> every write(2) moves at most 4 KiB.\n\nMicronit.  Output of unpack-objects obviously does not get flushed\nto \"the output pack file\".\n\n> On FUSE-backed filesystems every write(2) is a synchronous round\n> trip through the FUSE protocol (userspace -> kernel -> userspace ->\n> back), so the 4 KiB buffer turns a clone into many unnecessary tiny\n> writes with noticeable latency overhead.\n>\n> Increase the buffer from 4 KiB to 128 KiB, matching the default\n> already used by the hashfile layer in csum-file.c.\n\nQuite sensible reasoning presented very nicely.\n\nIt may probably be a #leftoverbit but these three instances of (128\n* 1024) may want to have a common symbolic constant, like\n\n    #define DEFAULT_IOBUFFER_SIZE_IN_BYTES (128 * 1024)\n\nin a bit more central header file.  Especially for the one in\ncsum-file.c where there is no symbolic constant used for that\npurpose.\n\n> Testing with strace on HTTPS clones of git/git (~296 MB pack, 5 runs\n> per variant, isolated builds from the same v2.54.0 source) shows:\n>\n>   index-pack pack file writes: 72,465 -> 24,943 avg (66% reduction)\n>   total write() syscalls:     310,192 -> 259,530 avg (17% reduction)\n>   writes of exactly 4096 bytes: ~40,077 -> 0 (eliminated)\n\nHmph, I would have expected more like (1 - 4/128) ~ 97% reduction.\nThe difference between that and 66% is coming from where?  There are\ninherently short writes that do not utilize the new larger buffer\nbeyond 4kB?  If so, another number of interest might be the number\nof writes smaller than 4096 bytes, perhaps?\n"},{"id":"542368","messageId":"c19a0e29-1218-4239-a362-df514153b5ff@gmail.com","threadId":"65551","inReplyTo":"xmqqldeb9w8e.fsf@gitster.g","subject":"Re: [PATCH] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-04-27T12:36:40Z","receivedAt":"2026-04-27T12:36:42Z","isPatch":true,"body":"On 4/25/2026 6:21 AM, Junio C Hamano wrote:\n> \"Scott Bauersfeld via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n> \n>> From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n>>\n>> On FUSE-backed filesystems every write(2) is a synchronous round\n>> trip through the FUSE protocol (userspace -> kernel -> userspace ->\n>> back), so the 4 KiB buffer turns a clone into many unnecessary tiny\n>> writes with noticeable latency overhead.\n>>\n>> Increase the buffer from 4 KiB to 128 KiB, matching the default\n>> already used by the hashfile layer in csum-file.c.\n> \n> Quite sensible reasoning presented very nicely.\n> \n> It may probably be a #leftoverbit but these three instances of (128\n> * 1024) may want to have a common symbolic constant, like\n> \n>     #define DEFAULT_IOBUFFER_SIZE_IN_BYTES (128 * 1024)\n> \n> in a bit more central header file.  Especially for the one in\n> csum-file.c where there is no symbolic constant used for that\n> purpose.\n\nI also had this thought. Would environment.h be the best place? \n>> Testing with strace on HTTPS clones of git/git (~296 MB pack, 5 runs\n>> per variant, isolated builds from the same v2.54.0 source) shows:\n>>\n>>   index-pack pack file writes: 72,465 -> 24,943 avg (66% reduction)\n>>   total write() syscalls:     310,192 -> 259,530 avg (17% reduction)\n>>   writes of exactly 4096 bytes: ~40,077 -> 0 (eliminated)\n> \n> Hmph, I would have expected more like (1 - 4/128) ~ 97% reduction.\n> The difference between that and 66% is coming from where?  There are\n> inherently short writes that do not utilize the new larger buffer\n> beyond 4kB?  If so, another number of interest might be the number\n> of writes smaller than 4096 bytes, perhaps?\n \nOne way to reword what you're asking is to measure \"number of writes\nnot using the whole buffer\" which is basically going to be \"the\nnumber of flush events from the application layer\". Every time the\napplication intends to flush, the current buffer is likely to not\nbe exactly full. I would expect this number to not change between\nimplementations in real experiments.\n\nThe improvement here comes from the reduced number of flushes due\nto buffer limits. I see that this can be measured in the number of\nsystem-level events, but what impact does this have on the end-to-\nend time of 'git index-pack' or 'git unpack-objects'? Is there a\nt/perf/ test that can demonstrate this improvement for a variety\nof real repos using GIT_PERF_REPO?\n\nThanks,\n-Stolee\n\n"},{"id":"542391","messageId":"pull.2282.v2.git.git.1777306114914.gitgitgadget@gmail.com","threadId":"65551","inReplyTo":"pull.2282.git.git.1777058098756.gitgitgadget@gmail.com","subject":"[PATCH v2] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Scott Bauersfeld via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-27T16:08:34Z","receivedAt":"2026-04-27T16:08:38Z","isPatch":true,"body":"From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n\nindex-pack and unpack-objects both read pack data from stdin through\na 4 KiB static buffer. In index-pack, each fill() flushes consumed\nbytes to the pack file via write_or_die(), capping every write(2)\nat 4 KiB. unpack-objects uses the same buffer pattern for reads.\n\nOn FUSE-backed filesystems every write(2) is a synchronous round\ntrip through the FUSE protocol (userspace -> kernel -> userspace ->\nback), so the 4 KiB buffer turns a clone into many unnecessary tiny\nwrites with noticeable latency overhead.\n\nIncrease the buffer from 4 KiB to 128 KiB. Introduce a shared\nDEFAULT_PACKFILE_BUFFER_SIZE constant in git-compat-util.h (next to\nMAX_IO_SIZE) and use it in index-pack, unpack-objects, and the\nhashfile layer in csum-file (which already used 128 KiB but\nhardcoded the value).\n\nSyscall counts via strace on HTTPS clones of git/git (~296 MB pack,\n5 runs per variant, isolated builds from the same v2.54.0 source):\n\n  index-pack pack file writes: 72,465 -> 24,943 avg (65% fewer)\n  total write() syscalls:     310,192 -> 259,530 avg (16% fewer)\n  writes of exactly 4096 bytes: ~40,077 -> 0\n\nWall-clock time of git clone over HTTPS onto a FUSE passthrough\nfilesystem with writeback caching disabled, 3 runs per variant:\n\n  vscode (~1.26 GB pack): 84.5s -> 75.7s avg (10% faster)\n  git/git (~306 MB pack):  22.6s -> 20.0s avg (11% faster)\n\nSigned-off-by: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n---\n    index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB\n    \n    index-pack and unpack-objects read pack data from stdin through a 4 KiB\n    static buffer. In index-pack, each fill() flushes consumed bytes to the\n    pack file via write_or_die(), capping every write(2) at 4 KiB.\n    unpack-objects uses the same buffer pattern for reads.\n    \n    On FUSE-backed filesystems every write(2) is a synchronous round trip\n    through the FUSE protocol (userspace → kernel → userspace → back), so\n    the 4 KiB buffer turns a clone into many unnecessary tiny writes with\n    noticeable latency overhead.\n    \n    Increase the buffer from 4 KiB to 128 KiB. Introduce a shared\n    DEFAULT_PACKFILE_BUFFER_SIZE constant in git-compat-util.h (next to\n    MAX_IO_SIZE) and use it in index-pack, unpack-objects, and the hashfile\n    layer in csum-file (which already used 128 KiB but hardcoded the value).\n    \n    \n    Syscall reduction\n    =================\n    \n    Measured via strace -f on HTTPS clones of git/git (~296 MB pack, 5 runs\n    per variant, isolated builds from the same v2.54.0 source):\n    \n    Metric Unpatched (4 KiB) Patched (128 KiB) Change index-pack writes to\n    pack file 72,465 avg 24,943 avg −65% Total write() syscalls (all\n    processes) 310,192 avg 259,530 avg −16% Writes of exactly 4096 bytes\n    ~40,077 avg 0 eliminated HEAD / file count / fsck ✓ ✓ identical\n    \n    \n    Wall-clock time on FUSE\n    =======================\n    \n    Measured wall-clock time of git clone over HTTPS onto a FUSE passthrough\n    filesystem with writeback caching disabled. 3 runs per variant:\n    \n    Repo Unpatched avg Patched avg Change microsoft/vscode (~1.26 GB pack)\n    84.5s 75.7s −10% git/git (~306 MB pack) 22.6s 20.0s −11%\n    \n    \n    Changes since v1\n    ================\n    \n     * Introduced shared DEFAULT_PACKFILE_BUFFER_SIZE constant in\n       git-compat-util.h (next to MAX_IO_SIZE), replacing per-file #define\n       and the hardcoded value in csum-file.c. Placed here rather than\n       environment.h since it is an I/O buffer size, not an environment\n       variable or repo config.\n     * Added wall-clock timing on a FUSE filesystem.\n     * Cleaned up the commit description a bit.\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2282%2Fsbauersfeld%2Fsb%2Fincrease-index-pack-input-buffer-v2\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2282/sbauersfeld/sb/increase-index-pack-input-buffer-v2\nPull-Request: https://github.com/git/git/pull/2282\n\nRange-diff vs v1:\n\n 1:  c388e1dc2f ! 1:  ac2559ccb5 index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB\n     @@ Metadata\n       ## Commit message ##\n          index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB\n      \n     -    Both index-pack and unpack-objects read pack data from stdin through\n     -    a 4 KiB static buffer (input_buffer[4096]). On each fill(), consumed\n     -    bytes are flushed to the output pack file via write_or_die(), so\n     -    every write(2) moves at most 4 KiB.\n     +    index-pack and unpack-objects both read pack data from stdin through\n     +    a 4 KiB static buffer. In index-pack, each fill() flushes consumed\n     +    bytes to the pack file via write_or_die(), capping every write(2)\n     +    at 4 KiB. unpack-objects uses the same buffer pattern for reads.\n      \n          On FUSE-backed filesystems every write(2) is a synchronous round\n          trip through the FUSE protocol (userspace -> kernel -> userspace ->\n          back), so the 4 KiB buffer turns a clone into many unnecessary tiny\n          writes with noticeable latency overhead.\n      \n     -    Increase the buffer from 4 KiB to 128 KiB, matching the default\n     -    already used by the hashfile layer in csum-file.c.\n     +    Increase the buffer from 4 KiB to 128 KiB. Introduce a shared\n     +    DEFAULT_PACKFILE_BUFFER_SIZE constant in git-compat-util.h (next to\n     +    MAX_IO_SIZE) and use it in index-pack, unpack-objects, and the\n     +    hashfile layer in csum-file (which already used 128 KiB but\n     +    hardcoded the value).\n      \n     -    Testing with strace on HTTPS clones of git/git (~296 MB pack, 5 runs\n     -    per variant, isolated builds from the same v2.54.0 source) shows:\n     +    Syscall counts via strace on HTTPS clones of git/git (~296 MB pack,\n     +    5 runs per variant, isolated builds from the same v2.54.0 source):\n      \n     -      index-pack pack file writes: 72,465 -> 24,943 avg (66% reduction)\n     -      total write() syscalls:     310,192 -> 259,530 avg (17% reduction)\n     -      writes of exactly 4096 bytes: ~40,077 -> 0 (eliminated)\n     +      index-pack pack file writes: 72,465 -> 24,943 avg (65% fewer)\n     +      total write() syscalls:     310,192 -> 259,530 avg (16% fewer)\n     +      writes of exactly 4096 bytes: ~40,077 -> 0\n      \n     -    All clones produce identical HEAD, file count, and pass fsck.\n     +    Wall-clock time of git clone over HTTPS onto a FUSE passthrough\n     +    filesystem with writeback caching disabled, 3 runs per variant:\n     +\n     +      vscode (~1.26 GB pack): 84.5s -> 75.7s avg (10% faster)\n     +      git/git (~306 MB pack):  22.6s -> 20.0s avg (11% faster)\n      \n          Signed-off-by: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n      \n     @@ builtin/index-pack.c: static int check_self_contained_and_connected;\n       \n      -/* We always read in 4kB chunks. */\n      -static unsigned char input_buffer[4096];\n     -+#define INPUT_BUFFER_SIZE (128 * 1024)\n     -+static unsigned char input_buffer[INPUT_BUFFER_SIZE];\n     ++static unsigned char input_buffer[DEFAULT_PACKFILE_BUFFER_SIZE];\n       static unsigned int input_offset, input_len;\n       static off_t consumed_bytes;\n       static off_t max_input_size;\n     @@ builtin/unpack-objects.c\n       \n      -/* We always read in 4kB chunks. */\n      -static unsigned char buffer[4096];\n     -+#define INPUT_BUFFER_SIZE (128 * 1024)\n     -+static unsigned char buffer[INPUT_BUFFER_SIZE];\n     ++static unsigned char buffer[DEFAULT_PACKFILE_BUFFER_SIZE];\n       static unsigned int offset, len;\n       static off_t consumed_bytes;\n       static off_t max_input_size;\n     +\n     + ## csum-file.c ##\n     +@@ csum-file.c: struct hashfile *hashfd_ext(const struct git_hash_algo *algop,\n     + \tf->algop = unsafe_hash_algo(algop);\n     + \tf->algop->init_fn(&f->ctx);\n     + \n     +-\tf->buffer_len = opts->buffer_len ? opts->buffer_len : 128 * 1024;\n     ++\tf->buffer_len = opts->buffer_len ? opts->buffer_len : DEFAULT_PACKFILE_BUFFER_SIZE;\n     + \tf->buffer = xmalloc(f->buffer_len);\n     + \tf->check_buffer = NULL;\n     + \n     +\n     + ## git-compat-util.h ##\n     +@@ git-compat-util.h: static inline uint64_t u64_add(uint64_t a, uint64_t b)\n     + # endif\n     + #endif\n     + \n     ++/*\n     ++ * Default buffer size for buffered I/O in pack file operations (index-pack,\n     ++ * unpack-objects) and the hashfile layer in csum-file.\n     ++ */\n     ++#define DEFAULT_PACKFILE_BUFFER_SIZE (128 * 1024)\n     ++\n     + #ifdef HAVE_ALLOCA_H\n     + # include <alloca.h>\n     + # define xalloca(size)      (alloca(size))\n\n\n builtin/index-pack.c     | 3 +--\n builtin/unpack-objects.c | 3 +--\n csum-file.c              | 2 +-\n git-compat-util.h        | 6 ++++++\n 4 files changed, 9 insertions(+), 5 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex ca7784dc2c..d86476676f 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -145,8 +145,7 @@ static int check_self_contained_and_connected;\n \n static struct progress *progress;\n \n-/* We always read in 4kB chunks. */\n-static unsigned char input_buffer[4096];\n+static unsigned char input_buffer[DEFAULT_PACKFILE_BUFFER_SIZE];\n static unsigned int input_offset, input_len;\n static off_t consumed_bytes;\n static off_t max_input_size;\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex e01cf6e360..da8ec83d9f 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -23,8 +23,7 @@\n static int dry_run, quiet, recover, has_errors, strict;\n static const char unpack_usage[] = \"git unpack-objects [-n] [-q] [-r] [--strict]\";\n \n-/* We always read in 4kB chunks. */\n-static unsigned char buffer[4096];\n+static unsigned char buffer[DEFAULT_PACKFILE_BUFFER_SIZE];\n static unsigned int offset, len;\n static off_t consumed_bytes;\n static off_t max_input_size;\ndiff --git a/csum-file.c b/csum-file.c\nindex 9558177a11..c1aeaf587a 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -178,7 +178,7 @@ struct hashfile *hashfd_ext(const struct git_hash_algo *algop,\n \tf->algop = unsafe_hash_algo(algop);\n \tf->algop->init_fn(&f->ctx);\n \n-\tf->buffer_len = opts->buffer_len ? opts->buffer_len : 128 * 1024;\n+\tf->buffer_len = opts->buffer_len ? opts->buffer_len : DEFAULT_PACKFILE_BUFFER_SIZE;\n \tf->buffer = xmalloc(f->buffer_len);\n \tf->check_buffer = NULL;\n \ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex ae1bdc90a4..a2f037811c 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -712,6 +712,12 @@ static inline uint64_t u64_add(uint64_t a, uint64_t b)\n # endif\n #endif\n \n+/*\n+ * Default buffer size for buffered I/O in pack file operations (index-pack,\n+ * unpack-objects) and the hashfile layer in csum-file.\n+ */\n+#define DEFAULT_PACKFILE_BUFFER_SIZE (128 * 1024)\n+\n #ifdef HAVE_ALLOCA_H\n # include <alloca.h>\n # define xalloca(size)      (alloca(size))\n\nbase-commit: 94f057755b7941b321fd11fec1b2e3ca5313a4e0\n-- \ngitgitgadget\n"},{"id":"542393","messageId":"5498637e-178f-48aa-8cdc-adc38b100627@gmail.com","threadId":"65551","inReplyTo":"pull.2282.v2.git.git.1777306114914.gitgitgadget@gmail.com","subject":"Re: [PATCH v2] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-04-27T17:23:06Z","receivedAt":"2026-04-27T17:23:10Z","isPatch":true,"body":"On 4/27/2026 12:08 PM, Scott Bauersfeld via GitGitGadget wrote:\n> From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n...\n> Wall-clock time of git clone over HTTPS onto a FUSE passthrough\n> filesystem with writeback caching disabled, 3 runs per variant:\n> \n>   vscode (~1.26 GB pack): 84.5s -> 75.7s avg (10% faster)\n>   git/git (~306 MB pack):  22.6s -> 20.0s avg (11% faster)\n\nWow! This is much higher than I expected. Great find.\n\nI imagine that other platforms or non-FUSE setups will not\nhave the same benefits. As long as they aren't _regressions_\nthen this is a great find.\n\n> -/* We always read in 4kB chunks. */\n> -static unsigned char input_buffer[4096];\n> +static unsigned char input_buffer[DEFAULT_PACKFILE_BUFFER_SIZE];\n\n> -/* We always read in 4kB chunks. */\n> -static unsigned char buffer[4096];\n> +static unsigned char buffer[DEFAULT_PACKFILE_BUFFER_SIZE];\n\nThese changes are what I expected in v2.\n\n> diff --git a/csum-file.c b/csum-file.c\n> index 9558177a11..c1aeaf587a 100644\n> --- a/csum-file.c\n> +++ b/csum-file.c\n> @@ -178,7 +178,7 @@ struct hashfile *hashfd_ext(const struct git_hash_algo *algop,\n>  \tf->algop = unsafe_hash_algo(algop);\n>  \tf->algop->init_fn(&f->ctx);\n>  \n> -\tf->buffer_len = opts->buffer_len ? opts->buffer_len : 128 * 1024;\n> +\tf->buffer_len = opts->buffer_len ? opts->buffer_len : DEFAULT_PACKFILE_BUFFER_SIZE;\n>  \tf->buffer = xmalloc(f->buffer_len);\n>  \tf->check_buffer = NULL;\n\nThis one surprised me, as this hunk wasn't in your v1 patch.\n\nI think using this replacement makes sense, since it _is_ an\nexact value. It did make me think as to how we landed on 128K\nfor this example.\n\nThe previous line is due to a1118c0a446 (csum-file: introduce\n`hashfd_ext()`, 2026-03-13), but it only moved the 128K default\nfrom hashfd(). Notably, hashfd_throughput() still uses an 8K\nsetting in opt->buffer_len.\n\nHilariously, I went spelunking for the original reason for the\n128K and it was 2ca245f8be5 (csum-file.h: increase hashfile\nbuffer size, 2021-05-18) written by...me. The motivation was\ndue to using the hashfile logic for the .git/index file which\nalso used 128K buffers in  f279894 (read-cache: make the index\nwrite buffer size 128K, 2021-02-18).\n\nAll this is to say that we now have two constants of identical\nvalue, where WRITE_BUFFER_SIZE in read-cache.c could be replaced\nwith your new DEFAULT_PACKFILE_BUFFER_SIZE.\n\nThis does make me think that maybe DEFAULT_PACKFILE_BUFFER_SIZE\nis misnamed? Should it be DEFAULT_HASHFILE_BUFFER_SIZE or\nDEFAULT_FILESYSTEM_BUFFER_SIZE to better fit this size value\nbeing used in both packfiles and index files?\n\n> diff --git a/git-compat-util.h b/git-compat-util.h\n> index ae1bdc90a4..a2f037811c 100644\n> --- a/git-compat-util.h\n> +++ b/git-compat-util.h\n> @@ -712,6 +712,12 @@ static inline uint64_t u64_add(uint64_t a, uint64_t b)\n>  # endif\n>  #endif\n>  \n> +/*\n> + * Default buffer size for buffered I/O in pack file operations (index-pack,\n> + * unpack-objects) and the hashfile layer in csum-file.\n> + */\n> +#define DEFAULT_PACKFILE_BUFFER_SIZE (128 * 1024)\n> +\nI see. Putting this in git-compat-util.h makes the rest\nof the changes good without any need to add a new include.\n\nThanks,\n-Stolee\n"},{"id":"542397","messageId":"pull.2282.v3.git.git.1777317998098.gitgitgadget@gmail.com","threadId":"65551","inReplyTo":"pull.2282.v2.git.git.1777306114914.gitgitgadget@gmail.com","subject":"[PATCH v3] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Scott Bauersfeld via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-27T19:26:38Z","receivedAt":"2026-04-27T19:26:41Z","isPatch":true,"body":"From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n\nindex-pack and unpack-objects both read pack data from stdin through\na 4 KiB static buffer. In index-pack, each fill() flushes consumed\nbytes to the pack file via write_or_die(), capping every write(2)\nat 4 KiB. unpack-objects uses the same buffer pattern for reads.\n\nOn FUSE-backed filesystems every write(2) is a synchronous round\ntrip through the FUSE protocol (userspace -> kernel -> userspace ->\nback), so the 4 KiB buffer turns a clone into many unnecessary tiny\nwrites with noticeable latency overhead.\n\nIncrease the buffer from 4 KiB to 128 KiB. Introduce a shared\nDEFAULT_IO_BUFFER_SIZE constant in git-compat-util.h (next to\nMAX_IO_SIZE) and use it in index-pack, unpack-objects, and the\nhashfile layer in csum-file (which already used 128 KiB but\nhardcoded the value).\n\nSyscall counts via strace on HTTPS clones of git/git (~296 MB pack,\n5 runs per variant, isolated builds from the same v2.54.0 source):\n\n  index-pack pack file writes: 72,465 -> 24,943 avg (65% fewer)\n  total write() syscalls:     310,192 -> 259,530 avg (16% fewer)\n  writes of exactly 4096 bytes: ~40,077 -> 0\n\nWall-clock time of git clone over HTTPS onto a FUSE passthrough\nfilesystem with writeback caching disabled, 3 runs per variant:\n\n  vscode (~1.26 GB pack): 84.5s -> 75.7s avg (10% faster)\n  git/git (~306 MB pack):  22.6s -> 20.0s avg (11% faster)\n\nSigned-off-by: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n---\n    index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB\n    \n    index-pack and unpack-objects read pack data from stdin through a 4 KiB\n    static buffer. In index-pack, each fill() flushes consumed bytes to the\n    pack file via write_or_die(), capping every write(2) at 4 KiB.\n    unpack-objects uses the same buffer pattern for reads.\n    \n    On FUSE-backed filesystems every write(2) is a synchronous round trip\n    through the FUSE protocol (userspace → kernel → userspace → back), so\n    the 4 KiB buffer turns a clone into many unnecessary tiny writes with\n    noticeable latency overhead.\n    \n    Increase the buffer from 4 KiB to 128 KiB. Introduce a shared\n    DEFAULT_IO_BUFFER_SIZE constant in git-compat-util.h (next to\n    MAX_IO_SIZE) and use it in index-pack, unpack-objects, and the hashfile\n    layer in csum-file (which already used 128 KiB but hardcoded the value).\n    \n    \n    Syscall reduction\n    =================\n    \n    Measured via strace -f on HTTPS clones of git/git (~296 MB pack, 5 runs\n    per variant, isolated builds from the same v2.54.0 source):\n    \n    Metric Unpatched (4 KiB) Patched (128 KiB) Change index-pack writes to\n    pack file 72,465 avg 24,943 avg −65% Total write() syscalls (all\n    processes) 310,192 avg 259,530 avg −16% Writes of exactly 4096 bytes\n    ~40,077 avg 0 eliminated HEAD / file count / fsck ✓ ✓ identical\n    \n    \n    Wall-clock time on FUSE\n    =======================\n    \n    Measured wall-clock time of git clone over HTTPS onto a FUSE passthrough\n    filesystem with writeback caching disabled. 3 runs per variant:\n    \n    Repo Unpatched avg Patched avg Change microsoft/vscode (~1.26 GB pack)\n    84.5s 75.7s −10% git/git (~306 MB pack) 22.6s 20.0s −11%\n    \n    \n    Changes since v2\n    ================\n    \n     * Renamed DEFAULT_PACKFILE_BUFFER_SIZE → DEFAULT_IO_BUFFER_SIZE per\n       Stolee's feedback. The constant is not packfile-specific, since it is\n       also used by the hashfile layer.\n     * Stolee noted that WRITE_BUFFER_SIZE in read-cache.c could be\n       consolidated. That constant was already removed in f6e2cd0625\n       (\"read-cache: delete unused hashing methods\", 2021-05-18) when\n       read-cache.c was converted to use the hashfile API, so there is\n       nothing left to unify. The rename to DEFAULT_IO_BUFFER_SIZE helps\n       account for the multiple usages of this constant.\n    \n    \n    Changes since v1\n    ================\n    \n     * Introduced shared DEFAULT_PACKFILE_BUFFER_SIZE constant in\n       git-compat-util.h (next to MAX_IO_SIZE), replacing per-file #define\n       and the hardcoded value in csum-file.c. Placed here rather than\n       environment.h since it is an I/O buffer size, not an environment\n       variable or repo config.\n     * Added wall-clock timing on a FUSE filesystem.\n     * Cleaned up the commit description a bit.\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2282%2Fsbauersfeld%2Fsb%2Fincrease-index-pack-input-buffer-v3\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2282/sbauersfeld/sb/increase-index-pack-input-buffer-v3\nPull-Request: https://github.com/git/git/pull/2282\n\nRange-diff vs v2:\n\n 1:  ac2559ccb5 ! 1:  df754ac879 index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB\n     @@ Commit message\n          writes with noticeable latency overhead.\n      \n          Increase the buffer from 4 KiB to 128 KiB. Introduce a shared\n     -    DEFAULT_PACKFILE_BUFFER_SIZE constant in git-compat-util.h (next to\n     +    DEFAULT_IO_BUFFER_SIZE constant in git-compat-util.h (next to\n          MAX_IO_SIZE) and use it in index-pack, unpack-objects, and the\n          hashfile layer in csum-file (which already used 128 KiB but\n          hardcoded the value).\n     @@ builtin/index-pack.c: static int check_self_contained_and_connected;\n       \n      -/* We always read in 4kB chunks. */\n      -static unsigned char input_buffer[4096];\n     -+static unsigned char input_buffer[DEFAULT_PACKFILE_BUFFER_SIZE];\n     ++static unsigned char input_buffer[DEFAULT_IO_BUFFER_SIZE];\n       static unsigned int input_offset, input_len;\n       static off_t consumed_bytes;\n       static off_t max_input_size;\n     @@ builtin/unpack-objects.c\n       \n      -/* We always read in 4kB chunks. */\n      -static unsigned char buffer[4096];\n     -+static unsigned char buffer[DEFAULT_PACKFILE_BUFFER_SIZE];\n     ++static unsigned char buffer[DEFAULT_IO_BUFFER_SIZE];\n       static unsigned int offset, len;\n       static off_t consumed_bytes;\n       static off_t max_input_size;\n     @@ csum-file.c: struct hashfile *hashfd_ext(const struct git_hash_algo *algop,\n       \tf->algop->init_fn(&f->ctx);\n       \n      -\tf->buffer_len = opts->buffer_len ? opts->buffer_len : 128 * 1024;\n     -+\tf->buffer_len = opts->buffer_len ? opts->buffer_len : DEFAULT_PACKFILE_BUFFER_SIZE;\n     ++\tf->buffer_len = opts->buffer_len ? opts->buffer_len : DEFAULT_IO_BUFFER_SIZE;\n       \tf->buffer = xmalloc(f->buffer_len);\n       \tf->check_buffer = NULL;\n       \n     @@ git-compat-util.h: static inline uint64_t u64_add(uint64_t a, uint64_t b)\n       #endif\n       \n      +/*\n     -+ * Default buffer size for buffered I/O in pack file operations (index-pack,\n     -+ * unpack-objects) and the hashfile layer in csum-file.\n     ++ * Default buffer size for buffered I/O in index-pack, unpack-objects,\n     ++ * and the hashfile layer in csum-file.\n      + */\n     -+#define DEFAULT_PACKFILE_BUFFER_SIZE (128 * 1024)\n     ++#define DEFAULT_IO_BUFFER_SIZE (128 * 1024)\n      +\n       #ifdef HAVE_ALLOCA_H\n       # include <alloca.h>\n\n\n builtin/index-pack.c     | 3 +--\n builtin/unpack-objects.c | 3 +--\n csum-file.c              | 2 +-\n git-compat-util.h        | 6 ++++++\n 4 files changed, 9 insertions(+), 5 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex ca7784dc2c..bb3639641c 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -145,8 +145,7 @@ static int check_self_contained_and_connected;\n \n static struct progress *progress;\n \n-/* We always read in 4kB chunks. */\n-static unsigned char input_buffer[4096];\n+static unsigned char input_buffer[DEFAULT_IO_BUFFER_SIZE];\n static unsigned int input_offset, input_len;\n static off_t consumed_bytes;\n static off_t max_input_size;\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex e01cf6e360..af67d1a1d3 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -23,8 +23,7 @@\n static int dry_run, quiet, recover, has_errors, strict;\n static const char unpack_usage[] = \"git unpack-objects [-n] [-q] [-r] [--strict]\";\n \n-/* We always read in 4kB chunks. */\n-static unsigned char buffer[4096];\n+static unsigned char buffer[DEFAULT_IO_BUFFER_SIZE];\n static unsigned int offset, len;\n static off_t consumed_bytes;\n static off_t max_input_size;\ndiff --git a/csum-file.c b/csum-file.c\nindex 9558177a11..d7a682c2b6 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -178,7 +178,7 @@ struct hashfile *hashfd_ext(const struct git_hash_algo *algop,\n \tf->algop = unsafe_hash_algo(algop);\n \tf->algop->init_fn(&f->ctx);\n \n-\tf->buffer_len = opts->buffer_len ? opts->buffer_len : 128 * 1024;\n+\tf->buffer_len = opts->buffer_len ? opts->buffer_len : DEFAULT_IO_BUFFER_SIZE;\n \tf->buffer = xmalloc(f->buffer_len);\n \tf->check_buffer = NULL;\n \ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex ae1bdc90a4..5024814bd4 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -712,6 +712,12 @@ static inline uint64_t u64_add(uint64_t a, uint64_t b)\n # endif\n #endif\n \n+/*\n+ * Default buffer size for buffered I/O in index-pack, unpack-objects,\n+ * and the hashfile layer in csum-file.\n+ */\n+#define DEFAULT_IO_BUFFER_SIZE (128 * 1024)\n+\n #ifdef HAVE_ALLOCA_H\n # include <alloca.h>\n # define xalloca(size)      (alloca(size))\n\nbase-commit: 94f057755b7941b321fd11fec1b2e3ca5313a4e0\n-- \ngitgitgadget\n"},{"id":"542398","messageId":"469a26e8-4309-4221-abac-e9a09e3f743d@gmail.com","threadId":"65551","inReplyTo":"pull.2282.v3.git.git.1777317998098.gitgitgadget@gmail.com","subject":"Re: [PATCH v3] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Derrick Stolee","fromEmail":"stolee@gmail.com","sentAt":"2026-04-27T20:12:55Z","receivedAt":"2026-04-27T20:12:57Z","isPatch":true,"body":"On 4/27/2026 3:26 PM, Scott Bauersfeld via GitGitGadget wrote:\n> From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n\n>     Changes since v2\n>     ================\n>     \n>      * Renamed DEFAULT_PACKFILE_BUFFER_SIZE → DEFAULT_IO_BUFFER_SIZE per\n>        Stolee's feedback. The constant is not packfile-specific, since it is\n>        also used by the hashfile layer.\n>      * Stolee noted that WRITE_BUFFER_SIZE in read-cache.c could be\n>        consolidated. That constant was already removed in f6e2cd0625\n>        (\"read-cache: delete unused hashing methods\", 2021-05-18) when\n>        read-cache.c was converted to use the hashfile API, so there is\n>        nothing left to unify. The rename to DEFAULT_IO_BUFFER_SIZE helps\n>        account for the multiple usages of this constant.\n\nThank you for discovering this context which made my recommendation\nnon-actionable. I was looking at the commit that added the 128K limit,\nwhich had that in its context, but not at the latest code. My mistake!\n\nI'm very happy with this version and look forward to the performance\nbenefits!\n\nThanks,\n-Stolee\n\n"},{"id":"542405","messageId":"xmqqecjz26wr.fsf@gitster.g","threadId":"65551","inReplyTo":"c19a0e29-1218-4239-a362-df514153b5ff@gmail.com","subject":"Re: [PATCH] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-28T01:46:44Z","receivedAt":"2026-04-28T01:46:46Z","isPatch":true,"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n>> The difference between that and 66% is coming from where?  There are\n>> inherently short writes that do not utilize the new larger buffer\n>> beyond 4kB?  If so, another number of interest might be the number\n>> of writes smaller than 4096 bytes, perhaps?\n>  \n> One way to reword what you're asking is to measure \"number of writes\n> not using the whole buffer\" which is basically going to be \"the\n> number of flush events from the application layer\".\n\nI do not think it would differ between the old and the new\nimplementation.\n\n> Every time the\n> application intends to flush, the current buffer is likely to not\n> be exactly full. I would expect this number to not change between\n> implementations in real experiments.\n\nYes, I agree.\n\nBut what I was trying to get at was a bit different.  \n\nThe application may have produced only 2kB before it issues a\n\"flush\".  Whether the buffer size is 4kB or 128kB, such a flush will\nonly write out 2kB, and the larger buffer size does not help at all.\nBut if the application has produced 90kB before it issues a \"flush\",\nthe larger buffer size would give us a great improvement.  With 4kB\nbuffer, before such an application level \"flush\", we would have seen\n22 = floor(90/4) calls of write(2) to flush the buffer, plus a 2kB\nwrite(2).  With 128kB buffer, we would see a single 90kB write(2).\n\nSo the apparently lower improvement than I naively have expected may\nbe attributable to the fact that many application level \"flush\" was\nnot large enough to benefit from 128kB buffer?  How much of the\ntotal number of bytes written came in large batches, vs tiny ones?\n\n> The improvement here comes from the reduced number of flushes due\n> to buffer limits.\n\nYes.\n\n> I see that this can be measured in the number of\n> system-level events, but what impact does this have on the end-to-\n> end time of 'git index-pack' or 'git unpack-objects'? Is there a\n> t/perf/ test that can demonstrate this improvement for a variety\n> of real repos using GIT_PERF_REPO?\n\nInteresting thought, but the number of system-level events (or the\nnumber of write(2) system calls) is not reduced by 97% because we\napparently are issuing too many of them, and the reason is?  I\nsuspect the reason why we still issue too many write(2) is because\nwe often do not send enough data between application-level flushes.\n"},{"id":"542406","messageId":"xmqq8qa726w9.fsf@gitster.g","threadId":"65551","inReplyTo":"469a26e8-4309-4221-abac-e9a09e3f743d@gmail.com","subject":"Re: [PATCH v3] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-04-28T01:47:02Z","receivedAt":"2026-04-28T01:47:04Z","isPatch":true,"body":"Derrick Stolee <stolee@gmail.com> writes:\n\n> On 4/27/2026 3:26 PM, Scott Bauersfeld via GitGitGadget wrote:\n>> From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n>\n>>     Changes since v2\n>>     ================\n>>     \n>>      * Renamed DEFAULT_PACKFILE_BUFFER_SIZE → DEFAULT_IO_BUFFER_SIZE per\n>>        Stolee's feedback. The constant is not packfile-specific, since it is\n>>        also used by the hashfile layer.\n>>      * Stolee noted that WRITE_BUFFER_SIZE in read-cache.c could be\n>>        consolidated. That constant was already removed in f6e2cd0625\n>>        (\"read-cache: delete unused hashing methods\", 2021-05-18) when\n>>        read-cache.c was converted to use the hashfile API, so there is\n>>        nothing left to unify. The rename to DEFAULT_IO_BUFFER_SIZE helps\n>>        account for the multiple usages of this constant.\n>\n> Thank you for discovering this context which made my recommendation\n> non-actionable. I was looking at the commit that added the 128K limit,\n> which had that in its context, but not at the latest code. My mistake!\n>\n> I'm very happy with this version and look forward to the performance\n> benefits!\n>\n> Thanks,\n> -Stolee\n\nYes, this version was very pleasant to read.  Thanks both.\n"},{"id":"542409","messageId":"20260428020933.GA660154@coredump.intra.peff.net","threadId":"65551","inReplyTo":"xmqqecjz26wr.fsf@gitster.g","subject":"Re: [PATCH] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Jeff King","fromEmail":"peff@peff.net","sentAt":"2026-04-28T02:09:33Z","receivedAt":"2026-04-28T02:09:41Z","isPatch":true,"body":"On Tue, Apr 28, 2026 at 10:46:44AM +0900, Junio C Hamano wrote:\n\n> The application may have produced only 2kB before it issues a\n> \"flush\".  Whether the buffer size is 4kB or 128kB, such a flush will\n> only write out 2kB, and the larger buffer size does not help at all.\n> But if the application has produced 90kB before it issues a \"flush\",\n> the larger buffer size would give us a great improvement.  With 4kB\n> buffer, before such an application level \"flush\", we would have seen\n> 22 = floor(90/4) calls of write(2) to flush the buffer, plus a 2kB\n> write(2).  With 128kB buffer, we would see a single 90kB write(2).\n> \n> So the apparently lower improvement than I naively have expected may\n> be attributable to the fact that many application level \"flush\" was\n> not large enough to benefit from 128kB buffer?  How much of the\n> total number of bytes written came in large batches, vs tiny ones?\n\nThe input to index-pack in a fetch is going to be the demuxing of the\nsideband via git-fetch. So it's probably flushing 64k or less each time\n(because that's the max size of a packet), and unless index-pack is\ngoing much slower than the input, that maximizes how much it will read.\n\nDepending on the source, though, it may be possible to go faster than\nindex-pack (which has to at least update the pack checksum for every\nbyte, and may even zlib inflate and hash the object itself if it's a\nnon-delta). In which case the sideband demuxer would start filling the\npipe and index-pack may get larger reads.\n\nWe could actually reduce the number of syscalls further if index-pack\ndid the demuxing itself, and we just handed it the descriptor. That\nprobably doesn't help all that much in this case, though, if the problem\nis not raw reads/writes on pipes, but rather ones that go to the slow\nFUSE filesystem. And as long as those pipe reads/writes are \"wide\"\n(allowing the eventual filesystem writes to also be wide), then the\nexact number may not be as important.\n\nBut the demuxing may also explain why the total number of writes did not\ndecrease as much as you expected, since those ones will probably not be\nreduced by the patch in question. So the improvement is a percentage of\nonly a smaller portion of the total (but not necessarily half, because\nthey may have been larger writes in the first place).\n\n-Peff\n"},{"id":"542431","messageId":"pull.2282.v4.git.git.1777387660841.gitgitgadget@gmail.com","threadId":"65551","inReplyTo":"pull.2282.v3.git.git.1777317998098.gitgitgadget@gmail.com","subject":"[PATCH v4] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Scott Bauersfeld via GitGitGadget","fromEmail":"gitgitgadget@gmail.com","sentAt":"2026-04-28T14:47:40Z","receivedAt":"2026-04-28T14:47:43Z","isPatch":true,"body":"From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n\nindex-pack and unpack-objects both read pack data from stdin through\na 4 KiB static buffer. In index-pack, each fill() flushes consumed\nbytes to the pack file via write_or_die(), capping every write(2)\nat 4 KiB. unpack-objects uses the same buffer pattern for reads.\n\nOn FUSE-backed filesystems every write(2) is a synchronous round\ntrip through the FUSE protocol (userspace -> kernel -> userspace ->\nback), so the 4 KiB buffer turns a clone into many unnecessary tiny\nwrites with noticeable latency overhead.\n\nIncrease the buffer from 4 KiB to 128 KiB. Introduce a shared\nDEFAULT_IO_BUFFER_SIZE constant in git-compat-util.h (next to\nMAX_IO_SIZE) and use it in index-pack, unpack-objects, and the\nhashfile layer in csum-file (which already used 128 KiB but\nhardcoded the value).\n\nPack file writes to a FUSE filesystem with writeback caching\ndisabled during HTTPS clones of git/git (~293 MB pack):\n\n  74,958 -> 4,687 (94% fewer)\n\nWall-clock time of git clone over HTTPS onto a FUSE passthrough\nfilesystem with writeback caching disabled, 3 runs per variant:\n\n  vscode (~1.26 GB pack): 84.5s -> 75.7s avg (10% faster)\n  git/git (~306 MB pack):  22.6s -> 20.0s avg (11% faster)\n\nSigned-off-by: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n---\n    index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB\n    \n    index-pack and unpack-objects read pack data from stdin through a 4 KiB\n    static buffer. In index-pack, each fill() flushes consumed bytes to the\n    pack file via write_or_die(), capping every write(2) at 4 KiB.\n    unpack-objects uses the same buffer pattern for reads.\n    \n    On FUSE-backed filesystems every write(2) is a synchronous round trip\n    through the FUSE protocol (userspace → kernel → userspace → back), so\n    the 4 KiB buffer turns a clone into many unnecessary tiny writes with\n    noticeable latency overhead.\n    \n    Increase the buffer from 4 KiB to 128 KiB. Introduce a shared\n    DEFAULT_IO_BUFFER_SIZE constant in git-compat-util.h (next to\n    MAX_IO_SIZE) and use it in index-pack, unpack-objects, and the hashfile\n    layer in csum-file (which already used 128 KiB but hardcoded the value).\n    \n    \n    Pack file write reduction\n    =========================\n    \n    Pack file writes to a FUSE filesystem with writeback caching disabled\n    during HTTPS clones of git/git (~293 MB pack):\n    \n    Unpatched avg Patched avg Change 74,958 4,687 −94%\n    \n    Write counts measured by logging writes in a FUSE passthrough daemon\n    (libfuse 3.10.5, writeback cache off).\n    \n    \n    Wall-clock time on FUSE\n    =======================\n    \n    Measured wall-clock time of git clone over HTTPS onto a FUSE passthrough\n    filesystem with writeback caching disabled. 3 runs per variant:\n    \n    Repo Unpatched avg Patched avg Change microsoft/vscode (~1.26 GB pack)\n    84.5s 75.7s −10% git/git (~306 MB pack) 22.6s 20.0s −11%\n    \n    \n    Changes since v3\n    ================\n    \n     * Replaced strace-based syscall measurements with FUSE daemon write\n       logging. The earlier strace numbers (72,465 → 24,943, 65% reduction)\n       were distorted: strace -f ptrace intercepts every syscall in all\n       traced processes and added enough overhead to distort the\n       measurements. The FUSE daemon logging captures write sizes without\n       perturbing the traced processes, showing the true reduction is 94%\n       (74,958 → 4,687).\n     * Note: Why 4,687 writes instead of ~2k writes as would be expected\n       with a 128 KiB buffer size? It appears that fill() is calling xread()\n       on a pipe and the linux default buffer size for pipes is 64KiB. I\n       also tested using fcntl(F_SETPIPE_SZ) to increase the pipe's buffer\n       size to 128KiB, which does indeed reduce total pack file writes to\n       ~2.4K.\n    \n    \n    Changes since v2\n    ================\n    \n     * Renamed DEFAULT_PACKFILE_BUFFER_SIZE → DEFAULT_IO_BUFFER_SIZE per\n       Stolee's feedback. The constant is not packfile-specific, since it is\n       also used by the hashfile layer.\n     * Stolee noted that WRITE_BUFFER_SIZE in read-cache.c could be\n       consolidated. That constant was already removed in f6e2cd0625\n       (\"read-cache: delete unused hashing methods\", 2021-05-18) when\n       read-cache.c was converted to use the hashfile API, so there is\n       nothing left to unify. The rename to DEFAULT_IO_BUFFER_SIZE helps\n       account for the multiple usages of this constant.\n    \n    \n    Changes since v1\n    ================\n    \n     * Introduced shared DEFAULT_PACKFILE_BUFFER_SIZE constant in\n       git-compat-util.h (next to MAX_IO_SIZE), replacing per-file #define\n       and the hardcoded value in csum-file.c. Placed here rather than\n       environment.h since it is an I/O buffer size, not an environment\n       variable or repo config.\n     * Added wall-clock timing on a FUSE filesystem.\n     * Cleaned up the commit description a bit.\n\nPublished-As: https://github.com/gitgitgadget/git/releases/tag/pr-git-2282%2Fsbauersfeld%2Fsb%2Fincrease-index-pack-input-buffer-v4\nFetch-It-Via: git fetch https://github.com/gitgitgadget/git pr-git-2282/sbauersfeld/sb/increase-index-pack-input-buffer-v4\nPull-Request: https://github.com/git/git/pull/2282\n\nRange-diff vs v3:\n\n 1:  df754ac879 ! 1:  146b1846a5 index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB\n     @@ Commit message\n          hashfile layer in csum-file (which already used 128 KiB but\n          hardcoded the value).\n      \n     -    Syscall counts via strace on HTTPS clones of git/git (~296 MB pack,\n     -    5 runs per variant, isolated builds from the same v2.54.0 source):\n     +    Pack file writes to a FUSE filesystem with writeback caching\n     +    disabled during HTTPS clones of git/git (~293 MB pack):\n      \n     -      index-pack pack file writes: 72,465 -> 24,943 avg (65% fewer)\n     -      total write() syscalls:     310,192 -> 259,530 avg (16% fewer)\n     -      writes of exactly 4096 bytes: ~40,077 -> 0\n     +      74,958 -> 4,687 (94% fewer)\n      \n          Wall-clock time of git clone over HTTPS onto a FUSE passthrough\n          filesystem with writeback caching disabled, 3 runs per variant:\n\n\n builtin/index-pack.c     | 3 +--\n builtin/unpack-objects.c | 3 +--\n csum-file.c              | 2 +-\n git-compat-util.h        | 6 ++++++\n 4 files changed, 9 insertions(+), 5 deletions(-)\n\ndiff --git a/builtin/index-pack.c b/builtin/index-pack.c\nindex ca7784dc2c..bb3639641c 100644\n--- a/builtin/index-pack.c\n+++ b/builtin/index-pack.c\n@@ -145,8 +145,7 @@ static int check_self_contained_and_connected;\n \n static struct progress *progress;\n \n-/* We always read in 4kB chunks. */\n-static unsigned char input_buffer[4096];\n+static unsigned char input_buffer[DEFAULT_IO_BUFFER_SIZE];\n static unsigned int input_offset, input_len;\n static off_t consumed_bytes;\n static off_t max_input_size;\ndiff --git a/builtin/unpack-objects.c b/builtin/unpack-objects.c\nindex e01cf6e360..af67d1a1d3 100644\n--- a/builtin/unpack-objects.c\n+++ b/builtin/unpack-objects.c\n@@ -23,8 +23,7 @@\n static int dry_run, quiet, recover, has_errors, strict;\n static const char unpack_usage[] = \"git unpack-objects [-n] [-q] [-r] [--strict]\";\n \n-/* We always read in 4kB chunks. */\n-static unsigned char buffer[4096];\n+static unsigned char buffer[DEFAULT_IO_BUFFER_SIZE];\n static unsigned int offset, len;\n static off_t consumed_bytes;\n static off_t max_input_size;\ndiff --git a/csum-file.c b/csum-file.c\nindex 9558177a11..d7a682c2b6 100644\n--- a/csum-file.c\n+++ b/csum-file.c\n@@ -178,7 +178,7 @@ struct hashfile *hashfd_ext(const struct git_hash_algo *algop,\n \tf->algop = unsafe_hash_algo(algop);\n \tf->algop->init_fn(&f->ctx);\n \n-\tf->buffer_len = opts->buffer_len ? opts->buffer_len : 128 * 1024;\n+\tf->buffer_len = opts->buffer_len ? opts->buffer_len : DEFAULT_IO_BUFFER_SIZE;\n \tf->buffer = xmalloc(f->buffer_len);\n \tf->check_buffer = NULL;\n \ndiff --git a/git-compat-util.h b/git-compat-util.h\nindex ae1bdc90a4..5024814bd4 100644\n--- a/git-compat-util.h\n+++ b/git-compat-util.h\n@@ -712,6 +712,12 @@ static inline uint64_t u64_add(uint64_t a, uint64_t b)\n # endif\n #endif\n \n+/*\n+ * Default buffer size for buffered I/O in index-pack, unpack-objects,\n+ * and the hashfile layer in csum-file.\n+ */\n+#define DEFAULT_IO_BUFFER_SIZE (128 * 1024)\n+\n #ifdef HAVE_ALLOCA_H\n # include <alloca.h>\n # define xalloca(size)      (alloca(size))\n\nbase-commit: 94f057755b7941b321fd11fec1b2e3ca5313a4e0\n-- \ngitgitgadget\n"},{"id":"543135","messageId":"xmqqy0hpnpkb.fsf@gitster.g","threadId":"65551","inReplyTo":"pull.2282.v4.git.git.1777387660841.gitgitgadget@gmail.com","subject":"Re: [PATCH v4] index-pack, unpack-objects: increase input buffer from 4 KiB to 128 KiB","fromName":"Junio C Hamano","fromEmail":"gitster@pobox.com","sentAt":"2026-05-12T05:51:16Z","receivedAt":"2026-05-12T05:51:20Z","isPatch":true,"body":"\"Scott Bauersfeld via GitGitGadget\" <gitgitgadget@gmail.com> writes:\n\n> From: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n>\n> index-pack and unpack-objects both read pack data from stdin through\n> a 4 KiB static buffer. In index-pack, each fill() flushes consumed\n> bytes to the pack file via write_or_die(), capping every write(2)\n> at 4 KiB. unpack-objects uses the same buffer pattern for reads.\n>\n> On FUSE-backed filesystems every write(2) is a synchronous round\n> trip through the FUSE protocol (userspace -> kernel -> userspace ->\n> back), so the 4 KiB buffer turns a clone into many unnecessary tiny\n> writes with noticeable latency overhead.\n>\n> Increase the buffer from 4 KiB to 128 KiB. Introduce a shared\n> DEFAULT_IO_BUFFER_SIZE constant in git-compat-util.h (next to\n> MAX_IO_SIZE) and use it in index-pack, unpack-objects, and the\n> hashfile layer in csum-file (which already used 128 KiB but\n> hardcoded the value).\n>\n> Pack file writes to a FUSE filesystem with writeback caching\n> disabled during HTTPS clones of git/git (~293 MB pack):\n>\n>   74,958 -> 4,687 (94% fewer)\n>\n> Wall-clock time of git clone over HTTPS onto a FUSE passthrough\n> filesystem with writeback caching disabled, 3 runs per variant:\n>\n>   vscode (~1.26 GB pack): 84.5s -> 75.7s avg (10% faster)\n>   git/git (~306 MB pack):  22.6s -> 20.0s avg (11% faster)\n>\n> Signed-off-by: Scott Bauersfeld <sbauersfeld@g.ucla.edu>\n> ---\n>...\n>     \n>     Changes since v3\n>     ================\n>     \n>      * Replaced strace-based syscall measurements with FUSE daemon write\n>        logging. The earlier strace numbers (72,465 → 24,943, 65% reduction)\n>        were distorted: strace -f ptrace intercepts every syscall in all\n>        traced processes and added enough overhead to distort the\n>        measurements. The FUSE daemon logging captures write sizes without\n>        perturbing the traced processes, showing the true reduction is 94%\n>        (74,958 → 4,687).\n>      * Note: Why 4,687 writes instead of ~2k writes as would be expected\n>        with a 128 KiB buffer size? It appears that fill() is calling xread()\n>        on a pipe and the linux default buffer size for pipes is 64KiB. I\n>        also tested using fcntl(F_SETPIPE_SZ) to increase the pipe's buffer\n>        size to 128KiB, which does indeed reduce total pack file writes to\n>        ~2.4K.\n\nIt seems that everybody was happy with v3 already, so let's merge it\ndown to 'next'.\n\nThanks.\n"}]}